Dataset Information

Data integration by multi-tuning parameter elastic net regression.

ABSTRACT: BACKGROUND:To integrate molecular features from multiple high-throughput platforms in prediction, a regression model that penalizes features from all platforms equally is commonly used. However, data from different platforms are likely to differ in effect sizes, the proportion of predictive features, and correlations structures. Subtle but important features may be missed by shrinking all features equally. RESULTS:We propose an Elastic net (EN) model with separate tuning parameter penalties for each platform that is fit using standard software. In a comprehensive simulation study, we evaluated the performance of EN logistic regression with multiple tuning penalties. We found that when the number of informative features differs among the platforms, and when there is no notable correlation between the features from different platforms, the multi-tuning parameter EN yields more predictive models. Moreover, the multi-tuning parameter EN is robust, in the sense that there is no loss of predictivity relative to a single tuning parameter EN when features across all platforms have similar effects. We also investigated the performance of multi-tuning parameter EN using real cancer datasets. CONCLUSION:The proposed multi-tuning parameter EN model, fit using standard penalized regression software, can achieve better prediction in sample classification when integrating multiple genomic platforms, compared to the traditional method where a single penalty parameter is used for all features in different platforms.

SUBMITTER: Liu J

PROVIDER: S-EPMC6180486 | biostudies-literature | 2018 Oct

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Data integration by multi-tuning parameter elastic net regression.

Liu Jie J Liang Gangning G Siegmund Kimberly D KD Lewinger Juan Pablo JP

BMC bioinformatics 20181010 1

<h4>Background</h4>To integrate molecular features from multiple high-throughput platforms in prediction, a regression model that penalizes features from all platforms equally is commonly used. However, data from different platforms are likely to differ in effect sizes, the proportion of predictive features, and correlations structures. Subtle but important features may be missed by shrinking all features equally.<h4>Results</h4>We propose an Elastic net (EN) model with separate tuning parameter ...[more]

PMID: 30305021

Similar Datasets

Project description:BackgroundHigh-dimensional omics data integration has emerged as a prominent avenue within the healthcare industry, presenting substantial potential to improve predictive models. However, the data integration process faces several challenges, including data heterogeneity, priority sequence in which data blocks are prioritized for rendering predictive information contained in multiple blocks, assessing the flow of information from one omics level to the other and multicollinearity.MethodsWe propose the Priority-Elastic net algorithm, a hierarchical regression method extending Priority-Lasso for the binary logistic regression model by incorporating a priority order for blocks of variables while fitting Elastic-net models sequentially for each block. The fitted values from each step are then used as an offset in the subsequent step. Additionally, we considered the adaptive elastic-net penalty within our priority framework to compare the results.ResultsThe Priority-Elastic net and Priority-Adaptive Elastic net algorithms were evaluated on a brain tumor dataset available from The Cancer Genome Atlas (TCGA), accounting for transcriptomics, proteomics, and clinical information measured over two glioma types: Lower-grade glioma (LGG) and glioblastoma (GBM).ConclusionOur findings suggest that the Priority-Elastic net is a highly advantageous choice for a wide range of applications. It offers moderate computational complexity, flexibility in integrating prior knowledge while introducing a hierarchical modeling perspective, and, importantly, improved stability and accuracy in predictions, making it superior to the other methods discussed. This evolution marks a significant step forward in predictive modeling, offering a sophisticated tool for navigating the complexities of multi-omics datasets in pursuit of precision medicine's ultimate goal: personalized treatment optimization based on a comprehensive array of patient-specific data. This framework can be generalized to time-to-event, Cox proportional hazards regression and multicategorical outcomes. A practical implementation of this method is available upon request in R script, complete with an example to facilitate its application.

Project description:ObjectivesCOVID-19 has been at the forefront of global concern since its emergence in December of 2019. Determining the social factors that drive case incidence is paramount to mitigating disease spread. We gathered data from the Social Vulnerability Index (SVI) along with Democratic voting percentage to attempt to understand which county-level sociodemographic metrics had a significant correlation with case rate for COVID-19.MethodsWe used elastic net regression due to issues with variable collinearity and model overfitting. Our modelling framework included using the ten Health and Human Services regions as submodels for the two time periods 22 March 2020 to 15 June 2021 (prior to the Delta time period) and 15 June 2021 to 1 November 2021 (the Delta time period).ResultsStatistically, elastic net improved prediction when compared to multiple regression, as almost every HHS model consistently had a lower root mean square error (RMSE) and satisfactory R2 coefficients. These analyses show that the percentage of minorities, disabled individuals, individuals living in group quarters, and individuals who voted Democratic correlated significantly with COVID-19 attack rate as determined by Variable Importance Plots (VIPs).ConclusionsThe percentage of minorities per county correlated positively with cases in the earlier time period and negatively in the later time period, which complements previous research. In contrast, higher percentages of disabled individuals per county correlated negatively in the earlier time period. Counties with an above average percentage of group quarters experienced a high attack rate early which then diminished in significance after the primary vaccine rollout. Higher Democratic voting consistently correlated negatively with cases, coinciding with previous findings regarding a partisan divide in COVID-19 cases at the county level. Our findings can assist regional policymakers in distributing resources to more vulnerable counties in future pandemics based on SVI.

Dataset Information

Data integration by multi-tuning parameter elastic net regression.

Publications

Data integration by multi-tuning parameter elastic net regression.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets