LD based tagSNP Selection System for Large-scale SNP Haplotype Dataset

KIPS Transactions on Computer and Communication Systems, Vol. 13, No. 1, pp. 79-86, Jan. 2006
10.3745/KIPSTA.2006.13.1.79, Full Text:


Recently the tagSNP selection problem has been researched for reducing the cost of association studies between human''s diversities and SNPs. General approach for this problem is that all of SNPs are separated into appropriate blocks and then tagSNPs are chosen in each block. MarSel in this paper is the system that involved the concept of linkage disequilibrium for overcoming the problem that the existing block partitioning approaches have short of biological meanings. In most approaches, the contiguous regions, which recombinations have ot been occurred in, may be separated into several blocks. Otherwise, in MarSel, the each contiguous region is kept into one block using LD coefficient D'' and then tagSNP selection step is performed. And MarSel guarantees the minimum tagSNP selection using entropy-based optimal selection algorithm when tagSNPs are chosen in each block, and enables chromosome-level association studies using efficient memory management technique when input is very large-scale dataset that is impossible to be processed in the existing systems.

