Dataset Information

The effect of noise on the predictive limit of QSAR models.

ABSTRACT: A key challenge in the field of Quantitative Structure Activity Relationships (QSAR) is how to effectively treat experimental error in the training and evaluation of computational models. It is often assumed in the field of QSAR that models cannot produce predictions which are more accurate than their training data. Additionally, it is implicitly assumed, by necessity, that data points in test sets or validation sets do not contain error, and that each data point is a population mean. This work proposes the hypothesis that QSAR models can make predictions which are more accurate than their training data and that the error-free test set assumption leads to a significant misevaluation of model performance. This work used 8 datasets with six different common QSAR endpoints, because different endpoints should have different amounts of experimental error associated with varying complexity of the measurements. Up to 15 levels of simulated Gaussian distributed random error was added to the datasets, and models were built on the error laden datasets using five different algorithms. The models were trained on the error laden data, evaluated on error-laden test sets, and evaluated on error-free test sets. The results show that for each level of added error, the RMSE for evaluation on the error free test sets was always better. The results support the hypothesis that, at least under the conditions of Gaussian distributed random error, QSAR models can make predictions which are more accurate than their training data, and that the evaluation of models on error laden test and validation sets may give a flawed measure of model performance. These results have implications for how QSAR models are evaluated, especially for disciplines where experimental error is very large, such as in computational toxicology.

SUBMITTER: Kolmar SS

PROVIDER: S-EPMC8613965 | biostudies-literature | 2021 Nov

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

The effect of noise on the predictive limit of QSAR models.

Kolmar Scott S SS Grulke Christopher M CM

Journal of cheminformatics 20211125 1

A key challenge in the field of Quantitative Structure Activity Relationships (QSAR) is how to effectively treat experimental error in the training and evaluation of computational models. It is often assumed in the field of QSAR that models cannot produce predictions which are more accurate than their training data. Additionally, it is implicitly assumed, by necessity, that data points in test sets or validation sets do not contain error, and that each data point is a population mean. This work ...[more]

PMID: 34823605

Dataset Information

The effect of noise on the predictive limit of QSAR models.

Publications

The effect of noise on the predictive limit of QSAR models.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Predictive QSAR Models for the Toxicity of Disinfection Byproducts.
| S-EPMC6151816 | biostudies-literature

On two novel parameters for validation of predictive QSAR models.
| S-EPMC6254296 | biostudies-literature

Integration of QSAR and SAR methods for the mechanistic interpretation of predictive models for carcinogenicity.
| S-EPMC3962111 | biostudies-literature

The Low Noise Limit in Gene Expression.
| S-EPMC4619080 | biostudies-literature

A new structure-based QSAR method affords both descriptive and predictive models for phosphodiesterase-4 inhibitors.
| S-EPMC2803435 | biostudies-literature

Predictive Capability of QSAR Models Based on the CompTox Zebrafish Embryo Assays: An Imbalanced Classification Problem.
| S-EPMC7998177 | biostudies-literature

3D-ALMOND-QSAR Models to Predict the Antidepressant Effect of Some Natural Compounds.
| S-EPMC8470101 | biostudies-literature

High predictive QSAR models for predicting the SARS coronavirus main protease inhibition activity of ketone-based covalent inhibitors
| S-EPMC8547569 | biostudies-literature

Noise reduction in genome-wide perturbation screens using linear mixed-effect models.
| S-EPMC3150043 | biostudies-literature

Benchmarks for interpretation of QSAR models.
| S-EPMC8157407 | biostudies-literature