Dataset Information

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.

ABSTRACT: We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows-Wheeler transform (BWT) and FM-index have been widely employed as a full-text searchable index for read alignment and de novo assembly. We introduce the concept of a population BWT and use it to store and index the sequencing reads of 2705 samples from the 1000 Genomes Project. A key feature is that, as more genomes are added, identical read sequences are increasingly observed, and compression becomes more efficient. We assess the support in the 1000 Genomes read data for every base position of two human reference assembly versions, identifying that 3.2 Mbp with population support was lost in the transition from GRCh37 with 13.7 Mbp added to GRCh38. We show that the vast majority of variant alleles can be uniquely described by overlapping 31-mers and show how rapid and accurate SNP and indel genotyping can be carried out across the genomes in the population BWT. We use the population BWT to carry out nonreference queries to search for the presence of all known viral genomes and discover human T-lymphotropic virus 1 integrations in six samples in a recognized epidemiological distribution.

SUBMITTER: Dolle DD

PROVIDER: S-EPMC5287235 | biostudies-literature | 2017 Feb

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.

Dolle Dirk D DD Liu Zhicheng Z Cotten Matthew M Simpson Jared T JT Iqbal Zamin Z Durbin Richard R McCarthy Shane A SA Keane Thomas M TM

Genome research 20161216 2

We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows-Wheeler transform (BWT) and FM-index have been widely employed as a full-text searchable index for read ...[more]

PMID: 27986821

Dataset Information

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.

Publications

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

The Nubeam reference-free approach to analyze metagenomic sequencing reads.
| S-EPMC7545149 | biostudies-literature

Data-dependent bucketing improves reference-free compression of sequencing reads.
| S-EPMC4547610 | biostudies-literature

Reference-free Association Mapping from Sequencing Reads Using k-mers.
| S-EPMC7842384 | biostudies-literature

Efficient de novo assembly of large genomes using compressed data structures.
| S-EPMC3290790 | biostudies-literature

Sequencing thousands of single-cell genomes with combinatorial indexing.
| S-EPMC5908213 | biostudies-literature

RNA-CODE: a noncoding RNA classification tool for short reads in NGS data lacking reference genomes.
| S-EPMC3808423 | biostudies-literature

Alignment of 1000 Genomes Project reads to reference assembly GRCh38.
| S-EPMC5522380 | biostudies-literature

GenomeScope: fast reference-free genome profiling from short reads.
| S-EPMC5870704 | biostudies-literature

qc3C: Reference-free quality control for Hi-C sequencing data.
| S-EPMC8530316 | biostudies-literature

ReorientExpress: reference-free orientation of nanopore cDNA reads with deep learning.
| S-EPMC6883653 | biostudies-literature