Unknown,Transcriptomics,Genomics,Proteomics

Dataset Information

0

Using RNA-Seq to create sample-specific proteomic databases that enable mass spectrometric discovery of splice junction peptides


ABSTRACT: Many new alternative splice forms have been detected at the transcript level using next generation sequencing (NGS) methods, especially RNA-Seq, but it is not known how many of these transcripts are being translated. Leveraging the unprecedented capabilities of NGS, we collected RNA-Seq and proteomics data from the same cell population (Jurkat cells) and created a bioinformatics pipeline that builds customized databases for the discovery of novel splice-junction peptides. Results: Eighty million paired-end Illumina reads and ~500,000 tandem mass spectra were used to identify 12,873 transcripts (19,320 including isoforms) and 6,810 proteins. We developed a bioinformatics workflow to retrieve high-confidence, novel splice junction sequences from the RNA data, translate these sequences into the analogous polypeptide sequence, and create a customized splice junction database for MS searching. Jurkat T-cell mRNA was analyzed on an Illumina HiSeq2000. ~80 million paired end reads (2x200bp, ~350bp lengths) were collected.

ORGANISM(S): Homo sapiens

SUBMITTER: Gloria Sheynkman 

PROVIDER: E-GEOD-45428 | biostudies-arrayexpress |

REPOSITORIES: biostudies-arrayexpress

Similar Datasets

2013-05-17 | GSE45428 | GEO
2011-11-06 | E-GEOD-30738 | biostudies-arrayexpress
2010-12-09 | E-GEOD-25494 | biostudies-arrayexpress
2011-11-06 | E-GEOD-30734 | biostudies-arrayexpress
2016-06-27 | E-GEOD-83777 | biostudies-arrayexpress
2015-06-18 | MSV000079166 | MassIVE
2017-09-30 | E-MTAB-5978 | biostudies-arrayexpress
2010-04-03 | E-GEOD-21132 | biostudies-arrayexpress
2011-11-06 | E-GEOD-30736 | biostudies-arrayexpress
2010-03-24 | E-MEXP-2627 | biostudies-arrayexpress