Unknown

Dataset Information

0

DACCOR-Detection, characterization, and reconstruction of repetitive regions in bacterial genomes.


ABSTRACT: The reconstruction of genomes using mapping-based approaches with short reads experiences difficulties when resolving repetitive regions. These repetitive regions in genomes result in low mapping qualities of the respective reads, which in turn lead to many unresolved bases. Currently, the reconstruction of these regions is often based on modified references in which the repetitive regions are masked. However, for many references, such masked genomes are not available or are based on repetitive regions of other genomes. Our idea is to identify repetitive regions in the reference genome de novo. These regions can then be used to reconstruct them separately using short read sequencing data. Afterward, the reconstructed repetitive sequence can be inserted into the reconstructed genome. We present the program detection, characterization, and reconstruction of repetitive regions, which performs these steps automatically. Our results show an increased base pair resolution of the repetitive regions in the reconstruction of Treponema pallidum samples, resulting in fewer unresolved bases.

SUBMITTER: Seitz A 

PROVIDER: S-EPMC5983011 | biostudies-literature | 2018

REPOSITORIES: biostudies-literature

altmetric image

Publications

DACCOR-Detection, characterization, and reconstruction of repetitive regions in bacterial genomes.

Seitz Alexander A   Hanssen Friederike F   Nieselt Kay K  

PeerJ 20180529


The reconstruction of genomes using mapping-based approaches with short reads experiences difficulties when resolving repetitive regions. These repetitive regions in genomes result in low mapping qualities of the respective reads, which in turn lead to many unresolved bases. Currently, the reconstruction of these regions is often based on modified references in which the repetitive regions are masked. However, for many references, such masked genomes are not available or are based on repetitive  ...[more]

Similar Datasets

| S-EPMC6052550 | biostudies-literature
| PRJNA610909 | ENA
| S-EPMC4828583 | biostudies-literature
| S-EPMC4032127 | biostudies-literature
| S-EPMC7605250 | biostudies-literature
| S-EPMC2817692 | biostudies-literature
| S-EPMC1274298 | biostudies-literature
| S-EPMC1769404 | biostudies-literature
| S-EPMC3792089 | biostudies-literature
| S-EPMC208921 | biostudies-other