Dataset Information

Variable importance for sustaining macrophyte presence via random forests: data imputation and model settings.

ABSTRACT: Data sets plagued with missing data and performance-affecting model parameters represent recurrent issues within the field of data mining. Via random forests, the influence of data reduction, outlier and correlated variable removal and missing data imputation technique on the performance of habitat suitability models for three macrophytes (Lemna minor, Spirodela polyrhiza and Nuphar lutea) was assessed. Higher performances (Cohen's kappa values around 0.2-0.3) were obtained for a high degree of data reduction, without outlier or correlated variable removal and with imputation of the median value. Moreover, the influence of model parameter settings on the performance of random forest trained on this data set was investigated along a range of individual trees (ntree), while the number of variables to be considered (mtry), was fixed at two. Altering the number of individual trees did not have a uniform effect on model performance, but clearly changed the required computation time. Combining both criteria provided an ntree value of 100, with the overall effect of ntree on performance being relatively limited. Temperature, pH and conductivity remained as variables and showed to affect the likelihood of L. minor, S. polyrhiza and N. lutea being present. Generally, high likelihood values were obtained when temperature is high (>20 °C), conductivity is intermediately low (50-200 mS m^-1) or pH is intermediate (6.9-8), thereby also highlighting that a multivariate management approach for supporting macrophyte presence remains recommended. Yet, as our conclusions are only based on a single freshwater data set, they should be further tested for other data sets.

SUBMITTER: Van Echelpoel W

PROVIDER: S-EPMC6162213 | biostudies-literature | 2018 Sep

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Variable importance for sustaining macrophyte presence via random forests: data imputation and model settings.

Van Echelpoel Wout W Goethals Peter L M PLM

Scientific reports 20180928 1

Data sets plagued with missing data and performance-affecting model parameters represent recurrent issues within the field of data mining. Via random forests, the influence of data reduction, outlier and correlated variable removal and missing data imputation technique on the performance of habitat suitability models for three macrophytes (Lemna minor, Spirodela polyrhiza and Nuphar lutea) was assessed. Higher performances (Cohen's kappa values around 0.2-0.3) were obtained for a high degree of ...[more]

PMID: 30266931

Dataset Information

Variable importance for sustaining macrophyte presence via random forests: data imputation and model settings.

Publications

Variable importance for sustaining macrophyte presence via random forests: data imputation and model settings.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Variable importance-weighted Random Forests.
| S-EPMC6051549 | biostudies-literature

Variable selection in the presence of missing data: resampling and imputation.
| S-EPMC5156376 | biostudies-literature

Evaluation of variable selection methods for random forests and omics data sets.
| S-EPMC6433899 | biostudies-literature

Maximal conditional chi-square importance in random forests.
| S-EPMC2832825 | biostudies-literature

Exploring the variable importance in random forests under correlations: a general concept applied to donor organ quality in post-transplant survival.
| S-EPMC10507897 | biostudies-literature

Mood Disorder Detection in Adolescents by Classification Trees, Random Forests and XGBoost in Presence of Missing Data.
| S-EPMC8468933 | biostudies-literature

Accuracy of random-forest-based imputation of missing data in the presence of non-normality, non-linearity, and interaction.
| S-EPMC7382855 | biostudies-literature

Surrogate minimal depth as an importance measure for variables in random forests.
| S-EPMC6761946 | biostudies-literature

Block Forests: random forests for blocks of clinical and omics covariate data.
| S-EPMC6598279 | biostudies-literature

The use of classification and regression algorithms using the random forests method with presence-only data to model species' distribution.
| S-EPMC6812352 | biostudies-literature