Dataset Information

Wrangling phosphoproteomic data to elucidate cancer signaling pathways.

ABSTRACT: The interpretation of biological data sets is essential for generating hypotheses that guide research, yet modern methods of global analysis challenge our ability to discern meaningful patterns and then convey results in a way that can be easily appreciated. Proteomic data is especially challenging because mass spectrometry detectors often miss peptides in complex samples, resulting in sparsely populated data sets. Using the R programming language and techniques from the field of pattern recognition, we have devised methods to resolve and evaluate clusters of proteins related by their pattern of expression in different samples in proteomic data sets. We examined tyrosine phosphoproteomic data from lung cancer samples. We calculated dissimilarities between the proteins based on Pearson or Spearman correlations and on Euclidean distances, whilst dealing with large amounts of missing data. The dissimilarities were then used as feature vectors in clustering and visualization algorithms. The quality of the clusterings and visualizations were evaluated internally based on the primary data and externally based on gene ontology and protein interaction networks. The results show that t-distributed stochastic neighbor embedding (t-SNE) followed by minimum spanning tree methods groups sparse proteomic data into meaningful clusters more effectively than other methods such as k-means and classical multidimensional scaling. Furthermore, our results show that using a combination of Spearman correlation and Euclidean distance as a dissimilarity representation increases the resolution of clusters. Our analyses show that many clusters contain one or more tyrosine kinases and include known effectors as well as proteins with no known interactions. Visualizing these clusters as networks elucidated previously unknown tyrosine kinase signal transduction pathways that drive cancer. Our approach can be applied to other data types, and can be easily adopted because open source software packages are employed.

SUBMITTER: Grimes ML

PROVIDER: S-EPMC3536783 | biostudies-literature | 2013

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Wrangling phosphoproteomic data to elucidate cancer signaling pathways.

Grimes Mark L ML Lee Wan-Jui WJ van der Maaten Laurens L Shannon Paul P

PloS one 20130103 1

The interpretation of biological data sets is essential for generating hypotheses that guide research, yet modern methods of global analysis challenge our ability to discern meaningful patterns and then convey results in a way that can be easily appreciated. Proteomic data is especially challenging because mass spectrometry detectors often miss peptides in complex samples, resulting in sparsely populated data sets. Using the R programming language and techniques from the field of pattern recogni ...[more]

PMID: 23300999

Similar Datasets

Project description:Breast cancer is the leading cause of cancer-related mortality in women worldwide, with an estimated 1.7 million new cases and 522,000 deaths around the world in 2012 alone. Cancer stem cells (CSCs) are essential for tumor reoccurrence and metastasis which is the major source of cancer lethality. G protein-coupled receptor chemokine (C-X-C motif) receptor 4 (CXCR4) is critical for tumor metastasis. However, stromal cell-derived factor 1 (SDF-1)/CXCR4-mediated signaling pathways in breast CSCs are largely unknown. Using isotope reductive dimethylation and large-scale MS-based quantitative phosphoproteome analysis, we examined protein phosphorylation induced by SDF-1/CXCR4 signaling in breast CSCs. We quantified more than 11,000 phosphorylation sites in 2,500 phosphoproteins. Of these phosphosites, 87% were statistically unchanged in abundance in response to SDF-1/CXCR4 stimulation. In contrast, 545 phosphosites in 266 phosphoproteins were significantly increased, whereas 113 phosphosites in 74 phosphoproteins were significantly decreased. SDF-1/CXCR4 increases phosphorylation in 60 cell migration- and invasion-related proteins, of them 43 (>70%) phosphoproteins are unrecognized. In addition, SDF-1/CXCR4 upregulates the phosphorylation of 44 previously uncharacterized kinases, 8 phosphatases, and 1 endogenous phosphatase inhibitor. Using computational approaches, we performed system-based analyses examining SDF-1/CXCR4-mediated phosphoproteome, including construction of kinase-substrate network and feedback regulation loops downstream of SDF-1/CXCR4 signaling in breast CSCs. We identified a previously unidentified SDF-1/CXCR4-PKA-MAP2K2-ERK signaling pathway and demonstrated the feedback regulation on MEK, ERK1/2, δ-catenin, and PPP1Cα in SDF-1/CXCR4 signaling in breast CSCs. This study gives a system-wide view of phosphorylation events downstream of SDF-1/CXCR4 signaling in breast CSCs, providing a resource for the study of CSC-targeted cancer therapy.

Dataset Information

Wrangling phosphoproteomic data to elucidate cancer signaling pathways.

Publications

Wrangling phosphoproteomic data to elucidate cancer signaling pathways.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets