Dataset Information

Subclass mapping: identifying common subtypes in independent disease data sets.

ABSTRACT: Whole genome expression profiles are widely used to discover molecular subtypes of diseases. A remaining challenge is to identify the correspondence or commonality of subtypes found in multiple, independent data sets generated on various platforms. While model-based supervised learning is often used to make these connections, the models can be biased to the training data set and thus miss inherent, relevant substructure in the test data. Here we describe an unsupervised subclass mapping method (SubMap), which reveals common subtypes between independent data sets. The subtypes within a data set can be determined by unsupervised clustering or given by predetermined phenotypes before applying SubMap. We define a measure of correspondence for subtypes and evaluate its significance building on our previous work on gene set enrichment analysis. The strength of the SubMap method is that it does not impose the structure of one data set upon another, but rather uses a bi-directional approach to highlight the common substructures in both. We show how this method can reveal the correspondence between several cancer-related data sets. Notably, it identifies common subtypes of breast cancer associated with estrogen receptor status, and a subgroup of lymphoma patients who share similar survival patterns, thus improving the accuracy of a clinical outcome predictor.

SUBMITTER: Hoshida Y

PROVIDER: S-EPMC2065909 | biostudies-literature | 2007

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Subclass mapping: identifying common subtypes in independent disease data sets.

Hoshida Yujin Y Brunet Jean-Philippe JP Tamayo Pablo P Golub Todd R TR Mesirov Jill P JP

PloS one 20071121 11

Whole genome expression profiles are widely used to discover molecular subtypes of diseases. A remaining challenge is to identify the correspondence or commonality of subtypes found in multiple, independent data sets generated on various platforms. While model-based supervised learning is often used to make these connections, the models can be biased to the training data set and thus miss inherent, relevant substructure in the test data. Here we describe an unsupervised subclass mapping method ( ...[more]

PMID: 18030330

Dataset Information

Subclass mapping: identifying common subtypes in independent disease data sets.

Publications

Subclass mapping: identifying common subtypes in independent disease data sets.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Repeated observation of breast tumor subtypes in independent gene expression data sets.
| S-EPMC166244 | biostudies-literature

Repeated observation of breast tumor subtypes in independent gene expression data sets
2006-03-07 | E-GEOD-4382 | biostudies-arrayexpress

Repeated observation of breast tumor subtypes in independent gene expression data sets
2006-03-08 | GSE4382 | GEO

Common sleep data pipeline for combined data sets.
| S-EPMC11302895 | biostudies-literature

MetaPhinder-Identifying Bacteriophage Sequences in Metagenomic Data Sets.
| S-EPMC5042410 | biostudies-other

Identifying Sepsis Subtypes from Routine Clinical Data.
| S-EPMC6680309 | biostudies-literature

Identifying Diagnoses of Schizophrenia Spectrum Disorder in Large Data Sets.
| S-EPMC9582046 | biostudies-literature

Identifying stably expressed genes from multiple RNA-Seq data sets.
| S-EPMC5178351 | biostudies-literature

MTIOT: Identifying HPV subtypes from multiple infection data.
| S-EPMC11755069 | biostudies-literature

Identifying proteomic LC-MS/MS data sets with Bumbershoot and IDPicker.
| S-EPMC4547833 | biostudies-literature