Unknown

Dataset Information

0

A New Fast Phasing Method Based On Haplotype Subtraction.


ABSTRACT: We developed a novel phasing approach, based solely on molecules and genotype frequency, that does not rely on inference of new alleles. We initiated the project because of errors that were detected in the phased 1000 Genomes Project data. The algorithm first combined identical genotypes into clusters and ranked them by descending frequency. Using alleles defined in homozygotes, it combined them to produce expected genotypes that were dismissed and subtracted them from remaining genotypes to define additional new putative alleles. Putative alleles had to be confirmed by identifying them in independent genotypes, and the process was iterated until all alleles were identified. The new approach was validated using single-molecule sequencing of eight loci, 145 (8 to 35 per locus) alleles were identified, and an average 98.2% (range, 95.0% to 99.9%) of 1000 genome individuals at these loci were explained. The accuracy of the new method was compared with that from PHASE and SHAPEIT2 to the experimentally determined genotypes based on single-molecule sequencing. Our method was comparable to PHASE and SHAPEIT2 in accuracy but was, on average, 14.6- and 10.8-fold faster, respectively.

SUBMITTER: Mocci E 

PROVIDER: S-EPMC6504677 | biostudies-literature | 2019 May

REPOSITORIES: biostudies-literature

altmetric image

Publications

A New Fast Phasing Method Based On Haplotype Subtraction.

Mocci Evelina E   Debeljak Marija M   Klein Alison P AP   Eshleman James R JR  

The Journal of molecular diagnostics : JMD 20190311 3


We developed a novel phasing approach, based solely on molecules and genotype frequency, that does not rely on inference of new alleles. We initiated the project because of errors that were detected in the phased 1000 Genomes Project data. The algorithm first combined identical genotypes into clusters and ranked them by descending frequency. Using alleles defined in homozygotes, it combined them to produce expected genotypes that were dismissed and subtracted them from remaining genotypes to def  ...[more]

Similar Datasets

| EGAS00001001853 | EGA
| S-EPMC5096458 | biostudies-literature
| S-EPMC4756330 | biostudies-literature
| S-EPMC7310859 | biostudies-literature
| S-EPMC6822470 | biostudies-literature
| S-EPMC4517485 | biostudies-literature
| S-EPMC6721696 | biostudies-literature
| S-EPMC1459002 | biostudies-literature
| S-EPMC6022575 | biostudies-literature
| S-EPMC7504856 | biostudies-literature