Unknown

Dataset Information

0

Semi-supervised consensus clustering for gene expression data analysis.


ABSTRACT: BACKGROUND:Simple clustering methods such as hierarchical clustering and k-means are widely used for gene expression data analysis; but they are unable to deal with noise and high dimensionality associated with the microarray gene expression data. Consensus clustering appears to improve the robustness and quality of clustering results. Incorporating prior knowledge in clustering process (semi-supervised clustering) has been shown to improve the consistency between the data partitioning and domain knowledge. METHODS:We proposed semi-supervised consensus clustering (SSCC) to integrate the consensus clustering with semi-supervised clustering for analyzing gene expression data. We investigated the roles of consensus clustering and prior knowledge in improving the quality of clustering. SSCC was compared with one semi-supervised clustering algorithm, one consensus clustering algorithm, and k-means. Experiments on eight gene expression datasets were performed using h-fold cross-validation. RESULTS:Using prior knowledge improved the clustering quality by reducing the impact of noise and high dimensionality in microarray data. Integration of consensus clustering with semi-supervised clustering improved performance as compared to using consensus clustering or semi-supervised clustering separately. Our SSCC method outperformed the others tested in this paper.

SUBMITTER: Wang Y 

PROVIDER: S-EPMC4036113 | biostudies-literature | 2014

REPOSITORIES: biostudies-literature

altmetric image

Publications

Semi-supervised consensus clustering for gene expression data analysis.

Wang Yunli Y   Pan Youlian Y  

BioData mining 20140508


<h4>Background</h4>Simple clustering methods such as hierarchical clustering and k-means are widely used for gene expression data analysis; but they are unable to deal with noise and high dimensionality associated with the microarray gene expression data. Consensus clustering appears to improve the robustness and quality of clustering results. Incorporating prior knowledge in clustering process (semi-supervised clustering) has been shown to improve the consistency between the data partitioning a  ...[more]

Similar Datasets

2008-08-30 | GSE12627 | GEO
| S-EPMC387275 | biostudies-literature
| S-EPMC10008853 | biostudies-literature
| S-EPMC4556708 | biostudies-literature
| S-EPMC4845510 | biostudies-other
| S-EPMC2666814 | biostudies-literature
| S-EPMC6455938 | biostudies-literature
| S-ECPF-GEOD-12627 | biostudies-other
| S-EPMC7096458 | biostudies-literature
| S-EPMC6548328 | biostudies-literature