Separation and assembly of deep sequencing data into discrete sub-population genomes.
Ontology highlight
ABSTRACT: Sequence heterogeneity is a common characteristic of RNA viruses that is often referred to as sub-populations or quasispecies. Traditional techniques used for assembly of short sequence reads produced by deep sequencing, such as de-novo assemblers, ignore the underlying diversity. Here, we introduce a novel algorithm that simultaneously assembles discrete sequences of multiple genomes present in populations. Using in silico data we were able to detect populations at as low as 0.1% frequency with complete global genome reconstruction and in a single sample detected 16 resolved sequences with no mismatches. We also applied the algorithm to high throughput sequencing data obtained for viruses present in sewage samples and successfully detected multiple sub-populations and recombination events
SUBMITTER: Karagiannis K
PROVIDER: S-EPMC5737798 | biostudies-literature | 2017 Nov
REPOSITORIES: biostudies-literature
ACCESS DATA