Unknown

Dataset Information

0

Shallow Representation Learning via Kernel PCA Improves QSAR Modelability.


ABSTRACT: Linear models offer a robust, flexible, and computationally efficient set of tools for modeling quantitative structure-activity relationships (QSARs) but have been eclipsed in performance by nonlinear methods. Support vector machines (SVMs) and neural networks are currently among the most popular and accurate QSAR methods because they learn new representations of the data that greatly improve modelability. In this work, we use shallow representation learning to improve the accuracy of L1 regularized logistic regression (LASSO) and meet the performance of Tanimoto SVM. We embedded chemical fingerprints in Euclidean space using Tanimoto (a.k.a. Jaccard) similarity kernel principal component analysis (KPCA) and compared the effects on LASSO and SVM model performance for predicting the binding activities of chemical compounds against 102 virtual screening targets. We observed similar performance and patterns of improvement for LASSO and SVM. We also empirically measured model training and cross-validation times to show that KPCA used in concert with LASSO classification is significantly faster than linear SVM over a wide range of training set sizes. Our work shows that powerful linear QSAR methods can match nonlinear methods and demonstrates a modular approach to nonlinear classification that greatly enhances QSAR model prototyping facility, flexibility, and transferability.

SUBMITTER: Rensi SE 

PROVIDER: S-EPMC5942586 | biostudies-literature | 2017 Aug

REPOSITORIES: biostudies-literature

altmetric image

Publications

Shallow Representation Learning via Kernel PCA Improves QSAR Modelability.

Rensi Stefano E SE   Altman Russ B RB  

Journal of chemical information and modeling 20170807 8


Linear models offer a robust, flexible, and computationally efficient set of tools for modeling quantitative structure-activity relationships (QSARs) but have been eclipsed in performance by nonlinear methods. Support vector machines (SVMs) and neural networks are currently among the most popular and accurate QSAR methods because they learn new representations of the data that greatly improve modelability. In this work, we use shallow representation learning to improve the accuracy of L1 regular  ...[more]

Similar Datasets

| S-EPMC5526982 | biostudies-other
2021-06-22 | GSE175456 | GEO
| S-EPMC10337068 | biostudies-literature
| S-EPMC4177829 | biostudies-other
| S-EPMC9489795 | biostudies-literature
| S-EPMC5994931 | biostudies-literature
| S-EPMC8271634 | biostudies-literature
| S-EPMC10569512 | biostudies-literature
| S-EPMC9576150 | biostudies-literature
| S-EPMC7267840 | biostudies-literature