Dataset Information

Towards accurate detection and genotyping of expressed variants from whole transcriptome sequencing data.

ABSTRACT: BACKGROUND: Massively parallel transcriptome sequencing (RNA-Seq) is becoming the method of choice for studying functional effects of genetic variability and establishing causal relationships between genetic variants and disease. However, RNA-Seq poses new technical and computational challenges compared to genome sequencing. In particular, mapping transcriptome reads onto the genome is more challenging than mapping genomic reads due to splicing. Furthermore, detection and genotyping of single nucleotide variants (SNVs) requires statistical models that are robust to variability in read coverage due to unequal transcript expression levels. RESULTS: In this paper we present a strategy to more reliably map transcriptome reads by taking advantage of the availability of both the genome reference sequence and transcript databases such as CCDS. We also present a novel Bayesian model for SNV discovery and genotyping based on quality scores. CONCLUSIONS: Experimental results on RNA-Seq data generated from blood cell tissue of three Hapmap individuals show that our methods yield increased accuracy compared to several widely used methods. The open source code implementing our methods, released under the GNU General Public License, is available at http://dna.engr.uconn.edu/software/NGSTools/.

SUBMITTER: Duitama J

PROVIDER: S-EPMC3394419 | biostudies-literature | 2012

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Towards accurate detection and genotyping of expressed variants from whole transcriptome sequencing data.

Duitama Jorge J Srivastava Pramod K PK Măndoiu Ion I II

BMC genomics 20120412

<h4>Background</h4>Massively parallel transcriptome sequencing (RNA-Seq) is becoming the method of choice for studying functional effects of genetic variability and establishing causal relationships between genetic variants and disease. However, RNA-Seq poses new technical and computational challenges compared to genome sequencing. In particular, mapping transcriptome reads onto the genome is more challenging than mapping genomic reads due to splicing. Furthermore, detection and genotyping of si ...[more]

PMID: 22537301

Dataset Information

Towards accurate detection and genotyping of expressed variants from whole transcriptome sequencing data.

Publications

Towards accurate detection and genotyping of expressed variants from whole transcriptome sequencing data.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Cyrius: accurate CYP2D6 genotyping using whole-genome sequencing data.
| S-EPMC7997805 | biostudies-literature

Variant analysis pipeline for accurate detection of genomic variants from transcriptome sequencing data.
| S-EPMC6756534 | biostudies-literature

Accurate detection and genotyping of SNPs utilizing population sequencing data.
| S-EPMC2847757 | biostudies-literature

XCAVATOR: accurate detection and genotyping of copy number variants from second and third generation whole-genome sequencing experiments.
| S-EPMC5609061 | biostudies-literature

Accurate detection of mosaic variants in sequencing data without matched controls.
| S-EPMC7065972 | biostudies-literature

NPSV: A simulation-driven approach to genotyping structural variants in whole-genome sequencing data.
| S-EPMC8246072 | biostudies-literature

Detection of minor variants in Mycobacterium tuberculosis whole genome sequencing data.
| S-EPMC8769888 | biostudies-literature

Enhanced copy number variants detection from whole-exome sequencing data using EXCAVATOR2.
| S-EPMC5175347 | biostudies-literature

Accurate multi-cancer detection using cell-free DNA shallow whole-genome sequencing data
2022-07-03 | E-MTAB-10934 | biostudies-arrayexpress

Simple, rapid and accurate genotyping-by-sequencing from aligned whole genomes with ArrayMaker.
| S-EPMC4325546 | biostudies-literature