Unknown

Dataset Information

0

Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data.


ABSTRACT:

Summary

Tandem DNA repeats can be sequenced with long-read technologies, but cannot be accurately deciphered due to the lack of computational tools taking high error rates of these technologies into account. Here we introduce Noise-Cancelling Repeat Finder (NCRF) to uncover putative tandem repeats of specified motifs in noisy long reads produced by Pacific Biosciences and Oxford Nanopore sequencers. Using simulations, we validated the use of NCRF to locate tandem repeats with motifs of various lengths and demonstrated its superior performance as compared to two alternative tools. Using real human whole-genome sequencing data, NCRF identified long arrays of the (AATGG)n repeat involved in heat shock stress response.

Availability and implementation

NCRF is implemented in C, supported by several python scripts, and is available in bioconda and at https://github.com/makovalab-psu/NoiseCancellingRepeatFinder.

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Harris RS 

PROVIDER: S-EPMC6853708 | biostudies-literature | 2019 Nov

REPOSITORIES: biostudies-literature

altmetric image

Publications

Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data.

Harris Robert S RS   Cechova Monika M   Makova Kateryna D KD  

Bioinformatics (Oxford, England) 20191101 22


<h4>Summary</h4>Tandem DNA repeats can be sequenced with long-read technologies, but cannot be accurately deciphered due to the lack of computational tools taking high error rates of these technologies into account. Here we introduce Noise-Cancelling Repeat Finder (NCRF) to uncover putative tandem repeats of specified motifs in noisy long reads produced by Pacific Biosciences and Oxford Nanopore sequencers. Using simulations, we validated the use of NCRF to locate tandem repeats with motifs of v  ...[more]

Similar Datasets

| S-EPMC148217 | biostudies-other
| S-EPMC7539535 | biostudies-literature
| S-EPMC7305381 | biostudies-literature
| S-EPMC6288141 | biostudies-literature
| S-EPMC7688782 | biostudies-literature
| S-EPMC4546671 | biostudies-literature
| S-EPMC6857246 | biostudies-literature
| S-EPMC8145836 | biostudies-literature
| S-EPMC8562783 | biostudies-literature
| S-EPMC3609794 | biostudies-literature