Dataset Information

Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.

ABSTRACT: Deep learning is currently the most successful machine learning technique in a wide range of application areas and has recently been applied successfully in drug discovery research to predict potential drug targets and to screen for active molecules. However, due to (1) the lack of large-scale studies, (2) the compound series bias that is characteristic of drug discovery datasets and (3) the hyperparameter selection bias that comes with the high number of potential deep learning architectures, it remains unclear whether deep learning can indeed outperform existing computational methods in drug discovery tasks. We therefore assessed the performance of several deep learning methods on a large-scale drug discovery dataset and compared the results with those of other machine learning and target prediction methods. To avoid potential biases from hyperparameter selection or compound series, we used a nested cluster-cross-validation strategy. We found (1) that deep learning methods significantly outperform all competing methods and (2) that the predictive performance of deep learning is in many cases comparable to that of tests performed in wet labs (i.e., in vitro assays).

SUBMITTER: Mayr A

PROVIDER: S-EPMC6011237 | biostudies-literature | 2018 Jun

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.

Mayr Andreas A Klambauer Günter G Unterthiner Thomas T Steijaert Marvin M Wegner Jörg K JK Ceulemans Hugo H Clevert Djork-Arné DA Hochreiter Sepp S

Chemical science 20180606 24

Deep learning is currently the most successful machine learning technique in a wide range of application areas and has recently been applied successfully in drug discovery research to predict potential drug targets and to screen for active molecules. However, due to (1) the lack of large-scale studies, (2) the compound series bias that is characteristic of drug discovery datasets and (3) the hyperparameter selection bias that comes with the high number of potential deep learning architectures, i ...[more]

PMID: 30155234

Dataset Information

Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.

Publications

Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Large-scale comparison of machine learning methods for profiling prediction of kinase inhibitors.
| S-EPMC10829268 | biostudies-literature

ChEMBL: a large-scale bioactivity database for drug discovery.
| S-EPMC3245175 | biostudies-literature

Large-scale prediction of activity cliffs using machine and deep learning methods of increasing complexity.
| S-EPMC9825040 | biostudies-literature

Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery.
| S-EPMC6690068 | biostudies-literature

Validating the validation: reanalyzing a large-scale comparison of deep learning and machine learning models for bioactivity prediction.
| S-EPMC7292817 | biostudies-literature

Alzheimer's disease risk assessment using large-scale machine learning methods.
| S-EPMC3826736 | biostudies-literature

Machine Learning-Enabled Pipeline for Large-Scale Virtual Drug Screening.
| S-EPMC8478848 | biostudies-literature

Similarity-Based Methods and Machine Learning Approaches for Target Prediction in Early Drug Discovery: Performance and Scope.
| S-EPMC7279241 | biostudies-literature

Gene prediction in metagenomic fragments: a large scale machine learning approach.
| S-EPMC2409338 | biostudies-literature

Comparison of Machine Learning Methods towards Developing Interpretable Polyamide Property Prediction.
| S-EPMC8587315 | biostudies-literature