Dataset Information

Measuring similarities between gene expression profiles through new data transformations.

ABSTRACT:

Background

Clustering methods are widely used on gene expression data to categorize genes with similar expression profiles. Finding an appropriate (dis)similarity measure is critical to the analysis. In our study, we developed a new measure for clustering the genes when the key factor is the shape of the profile, and when the expression magnitude should also be accounted for in determining the gene relationship. This is achieved by modeling the shape and magnitude parameters separately in a gene expression profile, and then using the estimated shape and magnitude parameters to define a measure in a new feature space.

Results

We explored several different transformation schemes to construct the feature spaces that include a space whose features are determined by the mutual differences of the original expression components, a space derived from a parametric covariance matrix, and the principal component space in traditional PCA analysis. The former two are the newly proposed and the latter is explored for comparison purposes. The new measures we defined in these feature spaces were employed in a K-means clustering procedure to perform analyses. Applying these algorithms to a simulation dataset, a developing mouse retina SAGE dataset, a small yeast sporulation cDNA dataset, and a maize root affymetrix microarray dataset, we found from the results that the algorithm associated with the first feature space, named TransChisq, showed clear advantages over other methods.

Conclusion

The proposed TransChisq is very promising in capturing meaningful gene expression clusters. This study also demonstrates the importance of data transformations in defining an efficient distance measure. Our method should provide new insights in analyzing gene expression data. The clustering algorithms are available upon request.

SUBMITTER: Kim K

PROVIDER: S-EPMC1804284 | biostudies-literature | 2007 Jan

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Measuring similarities between gene expression profiles through new data transformations.

Kim Kyungpil K Zhang Shibo S Jiang Keni K Cai Li L Lee In-Beum IB Feldman Lewis J LJ Huang Haiyan H

BMC bioinformatics 20070127

<h4>Background</h4>Clustering methods are widely used on gene expression data to categorize genes with similar expression profiles. Finding an appropriate (dis)similarity measure is critical to the analysis. In our study, we developed a new measure for clustering the genes when the key factor is the shape of the profile, and when the expression magnitude should also be accounted for in determining the gene relationship. This is achieved by modeling the shape and magnitude parameters separately i ...[more]

PMID: 17257435

Similar Datasets

Project description:Alcoholic liver disease (ALD) is a leading cause of cirrhosis in the United States, which is characterized by extensive deposition of extracellular matrix proteins and formation of a fibrous scar. Hepatic stellate cells (HSCs) are the major source of collagen type 1 producing myofibroblasts in ALD fibrosis. However, the mechanism of alcohol-induced activation of human and mouse HSCs is not fully understood. We compared the gene-expression profiles of primary cultured human HSCs (hHSCs) isolated from patients with ALD (n = 3) or without underlying liver disease (n = 4) using RNA-sequencing analysis. Furthermore, the gene-expression profile of ALD hHSCs was compared with that of alcohol-activated mHSCs (isolated from intragastric alcohol-fed mice) or CCl4-activated mouse HSCs (mHSCs). Comparative transcriptome analysis revealed that ALD hHSCs, in addition to alcohol-activated and CCl4-activated mHSCs, share the expression of common HSC activation (Col1a1 [collagen type I alpha 1 chain], Acta1 [actin alpha 1, skeletal muscle], PAI1 [plasminogen activator inhibitor-1], TIMP1 [tissue inhibitor of metalloproteinase 1], and LOXL2 [lysyl oxidase homolog 2]), indicating that a common mechanism underlies the activation of human and mouse HSCs. Furthermore, alcohol-activated mHSCs most closely recapitulate the gene-expression profile of ALD hHSCs. We identified the genes that are similarly and uniquely up-regulated in primary cultured alcohol-activated hHSCs and freshly isolated mHSCs, which include CSF1R (macrophage colony-stimulating factor 1 receptor), PLEK (pleckstrin), LAPTM5 (lysosmal-associated transmembrane protein 5), CD74 (class I transactivator, the invariant chain), CD53, MMP9 (matrix metallopeptidase 9), CD14, CTSS (cathepsin S), TYROBP (TYRO protein tyrosine kinase-binding protein), and ITGB2 (integrin beta-2), and other genes (compared with CCl4-activated mHSCs). Conclusion: We identified genes in alcohol-activated mHSCs from intragastric alcohol-fed mice that are largely consistent with the gene-expression profile of primary cultured hHSCs from patients with ALD. These genes are unique to alcohol-induced HSC activation in two species, and therefore may become targets or readout for antifibrotic therapy in experimental models of ALD.

Dataset Information

Measuring similarities between gene expression profiles through new data transformations.

Background

Results

Conclusion

Publications

Measuring similarities between gene expression profiles through new data transformations.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets