Unknown

Dataset Information

0

TAP: a targeted clinical genomics pipeline for detecting transcript variants using RNA-seq data.


ABSTRACT: BACKGROUND:RNA-seq is a powerful and cost-effective technology for molecular diagnostics of cancer and other diseases, and it can reach its full potential when coupled with validated clinical-grade informatics tools. Despite recent advances in long-read sequencing, transcriptome assembly of short reads remains a useful and cost-effective methodology for unveiling transcript-level rearrangements and novel isoforms. One of the major concerns for adopting the proven de novo assembly approach for RNA-seq data in clinical settings has been the analysis turnaround time. To address this concern, we have developed a targeted approach to expedite assembly and analysis of RNA-seq data. RESULTS:Here we present our Targeted Assembly Pipeline (TAP), which consists of four stages: 1) alignment-free gene-level classification of RNA-seq reads using BioBloomTools, 2) de novo assembly of individual targets using Trans-ABySS, 3) alignment of assembled contigs to the reference genome and transcriptome with GMAP and BWA and 4) structural and splicing variant detection using PAVFinder. We show that PAVFinder is a robust gene fusion detection tool when compared to established methods such as Tophat-Fusion and deFuse on simulated data of 448 events. Using the Leucegene acute myeloid leukemia (AML) RNA-seq data and a set of 580 COSMIC target genes, TAP identified a wide range of hallmark molecular anomalies including gene fusions, tandem duplications, insertions and deletions in agreement with published literature results. Moreover, also in this dataset, TAP captured AML-specific splicing variants such as skipped exons and novel splice sites reported in studies elsewhere. Running time of TAP on 100-150 million read pairs and a 580-gene set is one to 2 hours on a 48-core machine. CONCLUSIONS:We demonstrated that TAP is a fast and robust RNA-seq variant detection pipeline that is potentially amenable to clinical applications. TAP is available at http://www.bcgsc.ca/platform/bioinfo/software/pavfinder.

SUBMITTER: Chiu R 

PROVIDER: S-EPMC6131862 | biostudies-literature | 2018 Sep

REPOSITORIES: biostudies-literature

altmetric image

Publications

TAP: a targeted clinical genomics pipeline for detecting transcript variants using RNA-seq data.

Chiu Readman R   Nip Ka Ming KM   Chu Justin J   Birol Inanc I  

BMC medical genomics 20180910 1


<h4>Background</h4>RNA-seq is a powerful and cost-effective technology for molecular diagnostics of cancer and other diseases, and it can reach its full potential when coupled with validated clinical-grade informatics tools. Despite recent advances in long-read sequencing, transcriptome assembly of short reads remains a useful and cost-effective methodology for unveiling transcript-level rearrangements and novel isoforms. One of the major concerns for adopting the proven de novo assembly approac  ...[more]

Similar Datasets

| S-EPMC4739097 | biostudies-literature
| S-EPMC3597146 | biostudies-literature
| S-EPMC8044432 | biostudies-literature
| S-EPMC5716224 | biostudies-literature
| S-EPMC3467745 | biostudies-literature
| S-EPMC3051320 | biostudies-literature
| S-EPMC3460195 | biostudies-literature
| S-EPMC6213952 | biostudies-literature
| S-EPMC6373869 | biostudies-other
| S-EPMC4818202 | biostudies-other