Dataset Information

Integrated Identification and Quantification Error Probabilities for Shotgun Proteomics.

ABSTRACT: Protein quantification by label-free shotgun proteomics experiments is plagued by a multitude of error sources. Typical pipelines for identifying differential proteins use intermediate filters to control the error rate. However, they often ignore certain error sources and, moreover, regard filtered lists as completely correct in subsequent steps. These two indiscretions can easily lead to a loss of control of the false discovery rate (FDR). We propose a probabilistic graphical model, Triqler, that propagates error information through all steps, employing distributions in favor of point estimates, most notably for missing value imputation. The model outputs posterior probabilities for fold changes between treatment groups, highlighting uncertainty rather than hiding it. We analyzed 3 engineered data sets and achieved FDR control and high sensitivity, even for truly absent proteins. In a bladder cancer clinical data set we discovered 35 proteins at 5% FDR, whereas the original study discovered 1 and MaxQuant/Perseus 4 proteins at this threshold. Compellingly, these 35 proteins showed enrichment for functional annotation terms, whereas the top ranked proteins reported by MaxQuant/Perseus showed no enrichment. The model executes in minutes and is freely available at https://pypi.org/project/triqler/.

SUBMITTER: The M

PROVIDER: S-EPMC6398204 | biostudies-literature | 2019 Mar

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Integrated Identification and Quantification Error Probabilities for Shotgun Proteomics.

The Matthew M Käll Lukas L

Molecular & cellular proteomics : MCP 20181127 3

Protein quantification by label-free shotgun proteomics experiments is plagued by a multitude of error sources. Typical pipelines for identifying differential proteins use intermediate filters to control the error rate. However, they often ignore certain error sources and, moreover, regard filtered lists as completely correct in subsequent steps. These two indiscretions can easily lead to a loss of control of the false discovery rate (FDR). We propose a probabilistic graphical model, <i>Triqler< ...[more]

PMID: 30482846

Similar Datasets

Project description:The results of analysis of shotgun proteomics mass spectrometry data can be greatly affected by the selection of the reference protein sequence database against which the spectra are matched. For many species there are multiple sources from which somewhat different sequence sets can be obtained. This can lead to confusion about which database is best in which circumstances-a problem especially acute in human sample analysis. All sequence databases are genome-based, with sequences for the predicted gene and their protein translation products compiled. Our goal is to create a set of primary sequence databases that comprise the union of sequences from many of the different available sources and make the result easily available to the community. We have compiled a set of four sequence databases of varying sizes, from a small database consisting of only the ∼20,000 primary isoforms plus contaminants to a very large database that includes almost all nonredundant protein sequences from several sources. This set of tiered, increasingly complete human protein sequence databases suitable for mass spectrometry proteomics sequence database searching is called the Tiered Human Integrated Search Proteome set. In order to evaluate the utility of these databases, we have analyzed two different data sets, one from the HeLa cell line and the other from normal human liver tissue, with each of the four tiers of database complexity. The result is that approximately 0.8%, 1.1%, and 1.5% additional peptides can be identified for Tiers 2, 3, and 4, respectively, as compared with the Tier 1 database, at substantially increasing computational cost. This increase in computational cost may be worth bearing if the identification of sequence variants or the discovery of sequences that are not present in the reviewed knowledge base entries is an important goal of the study. We find that it is useful to search a data set against a simpler database, and then check the uniqueness of the discovered peptides against a more complex database. We have set up an automated system that downloads all the source databases on the first of each month and automatically generates a new set of search databases and makes them available for download at http://www.peptideatlas.org/thisp/ .

Project description:BackgroundThe cornea is a specialized transparent connective tissue responsible for the majority of light refraction and image focus for the retina. There are three main layers of the cornea: the epithelium that is exposed and acts as a protective barrier for the eye, the center stroma consisting of parallel collagen fibrils that refract light, and the endothelium that is responsible for hydration of the cornea from the aqueous humor. Normal cornea is an immunologically privileged tissue devoid of blood vessels, but injury can produce a loss of these conditions causing invasion of other processes that degrade the homeostatic properties resulting in a decrease in the amount of light refracted onto the retina. Determining a measure and drift of phenotypic cornea state from normal to an injured or diseased state requires knowledge of the existing protein signature within the tissue. In the study of corneal proteins, proteomics procedures have typically involved the pulverization of the entire cornea prior to analysis. Separation of the epithelium and endothelium from the core stroma and performing separate shotgun proteomics using liquid chromatography/mass spectrometry results in identification of many more proteins than previously employed methods using complete pulverized cornea.ResultsRabbit corneas were purchased, the epithelium and endothelium regions were removed, proteins processed and separately analyzed using liquid chromatography/mass spectrometry. Proteins identified from separate layers were compared against results from complete corneal samples. Protein digests were separated using a six hour liquid chromatographic gradient and ion-trap mass spectrometry used for detection of eluted peptide fractions. The SEQUEST database search results were filtered to allow only proteins with match probabilities of equal or better than 10-3 and peptides with a probability of 10-2 or less with at least two unique peptides isolated within the run along with default Xcorr values. These parameters resulted in the identification of over 350 proteins, including over 225 new proteins not previously detected in the cornea by mass spectrometry. In addition, corneal layer separation resulted in identification of nearly every protein that was identified in the complete cornea assay. The epithelium and endothelium each revealed many unique proteomes specific to each layer. In the endothelium, the protein olfactomedin-like 3 was identified for the first time in the cornea by this analysis. Olfactomedin-3 is a neuronal expressed protein also known as optimedin that stimulates formation of cell adherent and cell-cell tight junctions and its expression modulates cytoskeleton organization and cell migration. However, the function of this protein in rabbit corneal endothelium is currently unknown.ConclusionThis manuscript presents a description of a more comprehensive proteomic profile for mammalian cornea compared to past methods. The use of simple dissection procedures of the tissue and the application of long chromatographic gradients, many more proteins can be identified.

Dataset Information

Integrated Identification and Quantification Error Probabilities for Shotgun Proteomics.

Publications

Integrated Identification and Quantification Error Probabilities for Shotgun Proteomics.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets