Dataset Information

Increased yields of duplex sequencing data by a series of quality control tools.

ABSTRACT: Duplex sequencing is currently the most reliable method to identify ultra-low frequency DNA variants by grouping sequence reads derived from the same DNA molecule into families with information on the forward and reverse strand. However, only a small proportion of reads are assembled into duplex consensus sequences (DCS), and reads with potentially valuable information are discarded at different steps of the bioinformatics pipeline, especially reads without a family. We developed a bioinformatics toolset that analyses the tag and family composition with the purpose to understand data loss and implement modifications to maximize the data output for the variant calling. Specifically, our tools show that tags contain polymerase chain reaction and sequencing errors that contribute to data loss and lower DCS yields. Our tools also identified chimeras, which likely reflect barcode collisions. Finally, we also developed a tool that re-examines variant calls from raw reads and provides different summary data that categorizes the confidence level of a variant call by a tier-based system. With this tool, we can include reads without a family and check the reliability of the call, that increases substantially the sequencing depth for variant calling, a particular important advantage for low-input samples or low-coverage regions.

SUBMITTER: Povysil G

PROVIDER: S-EPMC7872198 | biostudies-literature | 2021 Mar

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Increased yields of duplex sequencing data by a series of quality control tools.

Povysil Gundula G Heinzl Monika M Salazar Renato R Stoler Nicholas N Nekrutenko Anton A Tiemann-Boege Irene I

NAR genomics and bioinformatics 20210209 1

Duplex sequencing is currently the most reliable method to identify ultra-low frequency DNA variants by grouping sequence reads derived from the same DNA molecule into families with information on the forward and reverse strand. However, only a small proportion of reads are assembled into duplex consensus sequences (DCS), and reads with potentially valuable information are discarded at different steps of the bioinformatics pipeline, especially reads without a family. We developed a bioinformatic ...[more]

PMID: 33575654

Dataset Information

Increased yields of duplex sequencing data by a series of quality control tools.

Publications

Increased yields of duplex sequencing data by a series of quality control tools.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Quality control of next-generation sequencing data without a reference.
| S-EPMC4018527 | biostudies-literature

Three-stage quality control strategies for DNA re-sequencing data.
| S-EPMC4492405 | biostudies-literature

qc3C: Reference-free quality control for Hi-C sequencing data.
| S-EPMC8530316 | biostudies-literature

Streamlined analysis of duplex sequencing data with Du Novo.
| S-EPMC5000403 | biostudies-literature

Multi-perspective quality control of Illumina exome sequencing data using QC3.
| S-EPMC5755963 | biostudies-literature

Falco: high-speed FastQC emulation for quality control of sequencing data.
| S-EPMC7845152 | biostudies-literature

Kraken: a set of tools for quality control and analysis of high-throughput sequence data.
| S-EPMC3991327 | biostudies-literature

seqQscorer: automated quality control of next-generation sequencing data using machine learning.
| S-EPMC7934511 | biostudies-literature

LongQC: A Quality Control Tool for Third Generation Sequencing Long Read Data.
| S-EPMC7144081 | biostudies-literature

Qualimap 2: advanced multi-sample quality control for high-throughput sequencing data.
| S-EPMC4708105 | biostudies-literature