Dataset Information

PIntMF: Penalized Integrative Matrix Factorization method for multi-omics data.

ABSTRACT:

Motivation

It is more and more common to perform multi-omics analyses to explore the genome at diverse levels and not only at a single level. Through integrative statistical methods, multi-omics data have the power to reveal new biological processes, potential biomarkers and subgroups in a cohort. Matrix factorization (MF) is an unsupervised statistical method that allows a clustering of individuals, but also reveals relevant omics variables from the various blocks.

Results

Here, we present PIntMF (Penalized Integrative Matrix Factorization), an MF model with sparsity, positivity and equality constraints. To induce sparsity in the model, we used a classical Lasso penalization on variable and individual matrices. For the matrix of samples, sparsity helps in the clustering, while normalization (matching an equality constraint) of inferred coefficients is added to improve interpretation. Moreover, we added an automatic tuning of the sparsity parameters using the famous glmnet package. We also proposed three criteria to help the user to choose the number of latent variables. PIntMF was compared with other state-of-the-art integrative methods including feature selection techniques in both synthetic and real data. PIntMF succeeds in finding relevant clusters as well as variables in two types of simulated data (correlated and uncorrelated). Next, PIntMF was applied to two real datasets (Diet and cancer), and it revealed interpretable clusters linked to available clinical data. Our method outperforms the existing ones on two criteria (clustering and variable selection). We show that PIntMF is an easy, fast and powerful tool to extract patterns and cluster samples from multi-omics data.

Availability and implementation

An R package is available at https://github.com/mpierrejean/pintmf.

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Pierre-Jean M

PROVIDER: S-EPMC8796362 | biostudies-literature | 2022 Jan

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

PIntMF: Penalized Integrative Matrix Factorization method for multi-omics data.

Pierre-Jean Morgane M Mauger Florence F Deleuze Jean-François JF Le Floch Edith E

Bioinformatics (Oxford, England) 20220101 4

<h4>Motivation</h4>It is more and more common to perform multi-omics analyses to explore the genome at diverse levels and not only at a single level. Through integrative statistical methods, multi-omics data have the power to reveal new biological processes, potential biomarkers and subgroups in a cohort. Matrix factorization (MF) is an unsupervised statistical method that allows a clustering of individuals, but also reveals relevant omics variables from the various blocks.<h4>Results</h4>Here, ...[more]

PMID: 34849583

Dataset Information

PIntMF: Penalized Integrative Matrix Factorization method for multi-omics data.

Motivation

Results

Availability and implementation

Supplementary information

Publications

PIntMF: Penalized Integrative Matrix Factorization method for multi-omics data.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

JSNMFuP: a unsupervised method for the integrative analysis of single-cell multi-omics data based on non-negative matrix factorization.
| S-EPMC11924690 | biostudies-literature

A non-negative matrix factorization method for detecting modules in heterogeneous omics multi-modal data.
| S-EPMC5006236 | biostudies-literature

HONMF: integration analysis of multi-omics microbiome data via matrix factorization and hypergraph.
| S-EPMC10243929 | biostudies-literature

Integrative clustering of multi-level 'omic data based on non-negative matrix factorization algorithm.
| S-EPMC5411077 | biostudies-literature

A Unified Bayesian Framework for Bi-overlapping-Clustering Multi-omics Data via Sparse Matrix Factorization.
| S-EPMC10766378 | biostudies-literature

Survival analysis by penalized regression and matrix factorization.
| S-EPMC3655687 | biostudies-literature

TiMEG: an integrative statistical method for partially missing multi-omics data.
| S-EPMC8674330 | biostudies-literature

Multi-omics assessment of dilated cardiomyopathy using non-negative matrix factorization.
| S-EPMC9387871 | biostudies-literature

TriTan: an efficient triple nonnegative matrix factorization method for integrative analysis of single-cell multiomics data.
| S-EPMC11586128 | biostudies-literature

A probabilistic multi-omics data matching method for detecting sample errors in integrative analysis.
| S-EPMC6615984 | biostudies-literature