Unknown

Dataset Information

0

DeBWT: parallel construction of Burrows-Wheeler Transform for large collection of genomes with de Bruijn-branch encoding.


ABSTRACT:

Motivation

With the development of high-throughput sequencing, the number of assembled genomes continues to rise. It is critical to well organize and index many assembled genomes to promote future genomics studies. Burrows-Wheeler Transform (BWT) is an important data structure of genome indexing, which has many fundamental applications; however, it is still non-trivial to construct BWT for large collection of genomes, especially for highly similar or repetitive genomes. Moreover, the state-of-the-art approaches cannot well support scalable parallel computing owing to their incremental nature, which is a bottleneck to use modern computers to accelerate BWT construction.

Results

We propose de Bruijn branch-based BWT constructor (deBWT), a novel parallel BWT construction approach. DeBWT innovatively represents and organizes the suffixes of input sequence with a novel data structure, de Bruijn branch encoding. This data structure takes the advantage of de Bruijn graph to facilitate the comparison between the suffixes with long common prefix, which breaks the bottleneck of the BWT construction of repetitive genomic sequences. Meanwhile, deBWT also uses the structure of de Bruijn graph for reducing unnecessary comparisons between suffixes. The benchmarking suggests that, deBWT is efficient and scalable to construct BWT for large dataset by parallel computing. It is well-suited to index many genomes, such as a collection of individual human genomes, with multiple-core servers or clusters.

Availability and implementation

deBWT is implemented in C language, the source code is available at https://github.com/hitbc/deBWT or https://github.com/DixianZhu/deBWTContact: ydwang@hit.edu.cn

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Liu B 

PROVIDER: S-EPMC4908350 | biostudies-literature | 2016 Jun

REPOSITORIES: biostudies-literature

altmetric image

Publications

deBWT: parallel construction of Burrows-Wheeler Transform for large collection of genomes with de Bruijn-branch encoding.

Liu Bo B   Zhu Dixian D   Wang Yadong Y  

Bioinformatics (Oxford, England) 20160601 12


<h4>Motivation</h4>With the development of high-throughput sequencing, the number of assembled genomes continues to rise. It is critical to well organize and index many assembled genomes to promote future genomics studies. Burrows-Wheeler Transform (BWT) is an important data structure of genome indexing, which has many fundamental applications; however, it is still non-trivial to construct BWT for large collection of genomes, especially for highly similar or repetitive genomes. Moreover, the sta  ...[more]

Similar Datasets

| S-EPMC7493373 | biostudies-literature
| S-EPMC7704051 | biostudies-literature
| S-EPMC9628837 | biostudies-literature
| S-EPMC2705234 | biostudies-literature
| S-EPMC2828108 | biostudies-literature
| S-EPMC4708002 | biostudies-literature
| S-EPMC5834436 | biostudies-literature
| S-EPMC10005600 | biostudies-literature
| S-EPMC7499882 | biostudies-literature
| S-BSST673 | biostudies-other