Unknown

Dataset Information

0

Comparative proteomics reveals a significant bias toward alternative protein isoforms with conserved structure and function.


ABSTRACT: Advances in high-throughput mass spectrometry are making proteomics an increasingly important tool in genome annotation projects. Peptides detected in mass spectrometry experiments can be used to validate gene models and verify the translation of putative coding sequences (CDSs). Here, we have identified peptides that cover 35% of the genes annotated by the GENCODE consortium for the human genome as part of a comprehensive analysis of experimental spectra from two large publicly available mass spectrometry databases. We detected the translation to protein of "novel" and "putative" protein-coding transcripts as well as transcripts annotated as pseudogenes and nonsense-mediated decay targets. We provide a detailed overview of the population of alternatively spliced protein isoforms that are detectable by peptide identification methods. We found that 150 genes expressed multiple alternative protein isoforms. This constitutes the largest set of reliably confirmed alternatively spliced proteins yet discovered. Three groups of genes were highly overrepresented. We detected alternative isoforms for 10 of the 25 possible heterogeneous nuclear ribonucleoproteins, proteins with a key role in the splicing process. Alternative isoforms generated from interchangeable homologous exons and from short indels were also significantly enriched, both in human experiments and in parallel analyses of mouse and Drosophila proteomics experiments. Our results show that a surprisingly high proportion (almost 25%) of the detected alternative isoforms are only subtly different from their constitutive counterparts. Many of the alternative splicing events that give rise to these alternative isoforms are conserved in mouse. It was striking that very few of these conserved splicing events broke Pfam functional domains or would damage globular protein structures. This evidence of a strong bias toward subtle differences in CDS and likely conserved cellular function and structure is remarkable and strongly suggests that the translation of alternative transcripts may be subject to selective constraints.

SUBMITTER: Ezkurdia I 

PROVIDER: S-EPMC3424414 | biostudies-literature | 2012 Sep

REPOSITORIES: biostudies-literature

altmetric image

Publications

Comparative proteomics reveals a significant bias toward alternative protein isoforms with conserved structure and function.

Ezkurdia Iakes I   del Pozo Angela A   Frankish Adam A   Rodriguez Jose Manuel JM   Harrow Jennifer J   Ashman Keith K   Valencia Alfonso A   Tress Michael L ML  

Molecular biology and evolution 20120322 9


Advances in high-throughput mass spectrometry are making proteomics an increasingly important tool in genome annotation projects. Peptides detected in mass spectrometry experiments can be used to validate gene models and verify the translation of putative coding sequences (CDSs). Here, we have identified peptides that cover 35% of the genes annotated by the GENCODE consortium for the human genome as part of a comprehensive analysis of experimental spectra from two large publicly available mass s  ...[more]

Similar Datasets

| S-EPMC2614494 | biostudies-literature
| S-EPMC8722536 | biostudies-literature
| S-EPMC29329 | biostudies-literature
| S-EPMC2745633 | biostudies-literature
| S-EPMC3518137 | biostudies-literature
| S-EPMC9677521 | biostudies-literature
| S-EPMC177548 | biostudies-other
| S-EPMC4354845 | biostudies-literature
| S-EPMC310876 | biostudies-literature
| S-EPMC7814404 | biostudies-literature