Dataset Information

Deep Learning-Based Structure-Activity Relationship Modeling for Multi-Category Toxicity Classification: A Case Study of 10K Tox21 Chemicals With High-Throughput Cell-Based Androgen Receptor Bioassay Data.

ABSTRACT: Deep learning (DL) has attracted the attention of computational toxicologists as it offers a potentially greater power for in silico predictive toxicology than existing shallow learning algorithms. However, contradicting reports have been documented. To further explore the advantages of DL over shallow learning, we conducted this case study using two cell-based androgen receptor (AR) activity datasets with 10K chemicals generated from the Tox21 program. A nested double-loop cross-validation approach was adopted along with a stratified sampling strategy for partitioning chemicals of multiple AR activity classes (i.e., agonist, antagonist, inactive, and inconclusive) at the same distribution rates amongst the training, validation and test subsets. Deep neural networks (DNN) and random forest (RF), representing deep and shallow learning algorithms, respectively, were chosen to carry out structure-activity relationship-based chemical toxicity prediction. Results suggest that DNN significantly outperformed RF (p < 0.001, ANOVA) by 22-27% for four metrics (precision, recall, F-measure, and AUPRC) and by 11% for another (AUROC). Further in-depth analyses of chemical scaffolding shed insights on structural alerts for AR agonists/antagonists and inactive/inconclusive compounds, which may aid in future drug discovery and improvement of toxicity prediction modeling.

SUBMITTER: Idakwo G

PROVIDER: S-EPMC6700714 | biostudies-literature | 2019

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Deep Learning-Based Structure-Activity Relationship Modeling for Multi-Category Toxicity Classification: A Case Study of 10K Tox21 Chemicals With High-Throughput Cell-Based Androgen Receptor Bioassay Data.

Idakwo Gabriel G Thangapandian Sundar S Luttrell Joseph J Zhou Zhaoxian Z Zhang Chaoyang C Gong Ping P

Frontiers in physiology 20190813

Deep learning (DL) has attracted the attention of computational toxicologists as it offers a potentially greater power for <i>in silico</i> predictive toxicology than existing shallow learning algorithms. However, contradicting reports have been documented. To further explore the advantages of DL over shallow learning, we conducted this case study using two cell-based androgen receptor (AR) activity datasets with 10K chemicals generated from the Tox21 program. A nested double-loop cross-validati ...[more]

PMID: 31456700

Similar Datasets

Project description:Cytotoxicity is a commonly used in vitro endpoint for evaluating chemical toxicity. In support of the U.S. Tox21 screening program, the cytotoxicity of ~10K chemicals was interrogated at 0, 8, 16, 24, 32, & 40 hours of exposure in a concentration dependent fashion in two cell lines (HEK293, HepG2) using two multiplexed, real-time assay technologies. One technology measures the metabolic activity of cells (i.e., cell viability, glo) while the other evaluates cell membrane integrity (i.e., cell death, flor). Using glo technology, more actives and greater temporal variations were seen in HEK293 cells, while results for the flor technology were more similar across the two cell types. Chemicals were grouped into classes based on their cytotoxicity kinetics profiles and these classes were evaluated for their associations with activity in the Tox21 nuclear receptor and stress response pathway assays. Some pathways, such as the activation of H2AX, were associated with the fast-responding cytotoxicity classes, while others, such as activation of TP53, were associated with the slow-responding cytotoxicity classes. By clustering pathways based on their degree of association to the different cytotoxicity kinetics labels, we identified clusters of pathways where active chemicals presented similar kinetics of cytotoxicity. Such linkages could be due to shared underlying biological processes between pathways, for example, activation of H2AX and heat shock factor. Others involving nuclear receptor activity are likely due to shared chemical structures rather than pathway level interactions. Based on the linkage between androgen receptor antagonism and Nrf2 activity, we surmise that a subclass of androgen receptor antagonists cause cytotoxicity via oxidative stress that is associated with Nrf2 activation. In summary, the real-time cytotoxicity screen provides informative chemical cytotoxicity kinetics data related to their cytotoxicity mechanisms, and with our analysis, it is possible to formulate mechanism-based hypotheses on the cytotoxic properties of the tested chemicals.

Project description:Since 2009, the Tox21 project has screened ∼8500 chemicals in more than 70 high-throughput assays, generating upward of 100 million data points, with all data publicly available through partner websites at the United States Environmental Protection Agency (EPA), National Center for Advancing Translational Sciences (NCATS), and National Toxicology Program (NTP). Underpinning this public effort is the largest compound library ever constructed specifically for improving understanding of the chemical basis of toxicity across research and regulatory domains. Each Tox21 federal partner brought specialized resources and capabilities to the partnership, including three approximately equal-sized compound libraries. All Tox21 data generated to date have resulted from a confluence of ideas, technologies, and expertise used to design, screen, and analyze the Tox21 10K library. The different programmatic objectives of the partners led to three distinct, overlapping compound libraries that, when combined, not only covered a diversity of chemical structures, use-categories, and properties but also incorporated many types of compound replicates. The history of development of the Tox21 "10K" chemical library and data workflows implemented to ensure quality chemical annotations and allow for various reproducibility assessments are described. Cheminformatics profiling demonstrates how the three partner libraries complement one another to expand the reach of each individual library, as reflected in coverage of regulatory lists, predicted toxicity end points, and physicochemical properties. ToxPrint chemotypes (CTs) and enrichment approaches further demonstrate how the combined partner libraries amplify structure-activity patterns that would otherwise not be detected. Finally, CT enrichments are used to probe global patterns of activity in combined ToxCast and Tox21 activity data sets relative to test-set size and chemical versus biological end point diversity, illustrating the power of CT approaches to discern patterns in chemical-activity data sets. These results support a central premise of the Tox21 program: A collaborative merging of programmatically distinct compound libraries would yield greater rewards than could be achieved separately.

Dataset Information

Deep Learning-Based Structure-Activity Relationship Modeling for Multi-Category Toxicity Classification: A Case Study of 10K Tox21 Chemicals With High-Throughput Cell-Based Androgen Receptor Bioassay Data.

Publications

Deep Learning-Based Structure-Activity Relationship Modeling for Multi-Category Toxicity Classification: A Case Study of 10K Tox21 Chemicals With High-Throughput Cell-Based Androgen Receptor Bioassay Data.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets