Skip to main content
Environmental Health Perspectives logoLink to Environmental Health Perspectives
. 2026 Apr 28;134(3):361–374. doi: 10.1021/EHP.6c00052

Expanded Tox21 Biological Assay Panel for the Prediction of Drug-Induced Liver Injury and Cardiotoxicity

Tuan Xu †, Masato Ooka †, Jinghua Zhao †, Srilatha Sakamuru †, Deborah K Ngan †, Li Zhang †, Shu Yang †, Jameson Travers †, Menghang Xia †, Tongan Zhao †, Carleen Klumpp-Thomas †, Hu Zhu †, Mathew D Hall †, Stephen Ferguson ‡, Natalie D Shaw §, David M Reif ‡, Anton Simeonov †, Ruili Huang †,*
PMCID: PMC13347643  PMID: 42428255

Abstract

BACKGROUND: Toxicology in the 21st Century (Tox21) assay data provide a valuable resource for the prediction of in vivo toxicity using machine learning models. However, the performances of these models previously developed using the pre-existing Tox21 assay data were less than ideal, likely due to insufficient coverage of the biological response space by the assay targets. OBJECTIVES: This study aimed to assess whether expanding the Tox21 portfolio with new assays that probe under-represented targets/pathways related to unanticipated adverse drug effects could improve the predictive capacity of in vitro assay data for in vivo toxicity such as drug-induced liver injury (DILI) and cardiotoxicity (DICT). METHODS: Models were constructed using data from the pre-existing panel of 36 assay targets and the expanded panel of 49 assay targets. A feature selection approach was used to determine the optimal number of assays needed for each model. The models were then applied to predict the potential hepatotoxicity and cardiotoxicity of compounds in the Tox21 10K compound library. RESULTS: For both DILI and DICT prediction, the best-performing models developed using the expanded assay panel required a smaller number of assays to achieve the same level of performance compared to those based on the pre-existing assays. Models constructed by combining both assay data (pre-existing + expanded) and chemical structure consistently outperformed those constructed based on assay data alone but showed similar performance to those constructed based on chemical structure. The compounds predicted to have the highest toxic potential were experimentally verified to demonstrate the effectiveness of our models in identifying new potentially toxic compounds. DISCUSSION: The expansion of the Tox21 assay panel has significantly enhanced the predictive capacity of assay data for predicting the DILI and DICT potential. This improvement underscores the importance of a diverse and comprehensive in vitro assay portfolio in advancing safety assessment.

Introduction

Drug-induced liver injury (DILI) and cardiotoxicity (DICT) represent the most severe and frequently encountered adverse drug reactions (ADRs) that can occur throughout the process of drug development and postmarketing surveillance, as they can result in the premature termination of drug development programs, the restricted use or withdrawal of previously approved and marketed medications, and, most critically, can lead to serious health risks for patients. For instance, the antidiabetic drug troglitazone was withdrawn because it was found to be hepatotoxic, causing mitochondrial damage and oxidative stress. The use of the antiarrhythmic drug dronedarone was restricted because it was found to increase rates of heart failure, stroke, and death from cardiovascular causes in patients with permanent atrial fibrillation. In terms of clinical risks, DILI can manifest across a spectrum of clinical presentations, ranging from mild transaminase elevation to acute liver failure, while DICT can lead to serious complications, such as arrhythmias, cardiomyopathy, and heart failure. Therefore, an accurate assessment of the DILI and DICT potential at an early stage of new drug development is imperative.

Conventional in vivo animal models used to evaluate DILI or DICT potential are costly and time-consuming, and also suffer from inherent limitations in their extrapolative capacity to human physiology due to interspecies variability. − Endeavors to predict DILI or DICT potential based on chemical structures often lack interpretability to support robust decision-making in drug development, as they frequently fail to account for the complex biological mechanisms involved in the toxicological risks. Machine learning models trained on in vitro assay data could shed light on the underlying biological mechanisms behind their predictions, providing valuable clues for molecular design and drug repositioning. , For example, the interpretability analyses of machine learning models confirmed that the nuclear factor, erythroid 2-related factor 2, was found to be a regulator of the onset and progression of chronic diseases, such as liver injury, and the lipid-activated transcription factor, peroxisome proliferator-activated receptor gamma (PPARγ), exhibited cardiovascular risks through both PPARγ-dependent and independent mechanisms.

The Toxicology in the 21st Century (Tox21) program is a toxicological research project jointly initiated by U.S. government agencies, aiming to establish high-throughput in vitro methods to efficiently evaluate the toxicity of chemicals and support risk assessment. − To date, the Tox21 program has developed a panel of in vitro assays based on toxicity related targets/pathways (e.g., nuclear receptor signaling and stress response pathways) to test a collection of approximately 10,000 environmental chemicals and drugs (Tox21 10K compound library) in a quantitative high-throughput screening (qHTS) format, generating over 120 million public data points to date. , We previously attempted to build machine learning models using Tox21 assay data and chemical structure information to predict DILI and DICT potential. Although the addition of assay data helped to improve the model performance, the models based on assay data alone did not show sufficient predictive power.

The less-than-ideal predictive performance of the assay data could be attributed to the insufficient coverage of the biological response space relevant for toxicity. The pre-existing Tox21 assay panel focused primarily on selected nuclear receptors and stress response pathways. This relatively limited focus suggests that activity in other toxicity pathways has not been adequately assessed, making it challenging to comprehensively reflect the complex in vivo mechanisms and processes of DILI and DICT. Therefore, expanding the coverage of biological responses by adding assays that probe under-represented pathways may improve the predictivity of Tox21 assay data. Toward this goal, we systematically identified these under-represented pathways in a data-driven approach for common adverse drug effects. For DILI and DICT, cytochrome P450 metabolic enzymes, G protein-coupled receptors, and several other stress response pathways were identified as targets of interest. As a proof-of-concept study, we prioritized several targets in the aforementioned families, with assays amenable for HTS to test their capacity in improving the prediction of DILI and DICT.

This study aims to compare the differences between the pre-existing assays and the expanded Tox21 assay panel in predicting the DILI and DICT potential. The results will provide insight into the key molecular mechanisms leading to DILI and DICT, that will help guide the optimization of the Tox21 assays for better in vivo toxicity prediction. The robust prediction models developed through this process have the promise of saving time and cost in assessing the adverse effects of new drugs and environmental chemicals.

Methods

Reactive Oxygen Species (ROS) Assay

For experimental validation, an in vitro cell-based ROS assay was used to test the compounds predicted as hepatotoxic in this study. HepaRG cells (Lonza, Basel, Switzerland) were plated into a white solid bottom 1536 well plate at 2,000 cells/well in 5 μL of Basal medium (Lonza, cat. no. MH100) containing Basal Medium Supplement (Lonza), Thawing and Plating Additive (Lonza), 20U/mL penicillin, and 20 μg/mL streptomycin. The next day, the medium was replaced with 5 μL of Basal medium containing Basal Medium Supplement, Preinduction and Tox Additive (Lonza), Serum-free Induction Additive (Lonza), 20U/mL penicillin, and 20 μg/mL streptomycin. After 3-day incubation at 37 °C, the medium was replaced with fresh induction medium, and the plate was incubated for 3 h at 37 °C. Twenty-three nL of the candidate compounds were transferred using the Pintool station (Wako Automation, San Diego, CA). The cells were incubated at 37 °C for 16 h. One μL of H2O2 substrate (Promega, Madison, WI) was added to each well, followed by 6 h incubation at 37 °C. Four μL of ROS-Glo detection reagent (Promega) were added to each well. Immediately after the detection reagent was added, the luminescence signal was read using the ViewLux plate reader (PerkinElmer, Shelton, CT). Compounds were tested as 11-point 1:3 titrations, starting from various stock concentrations (ranging from 20.6 to 125 μM, depending on solubility), in this assay. Prior to experimental use, compounds were subjected to analytical quality control (QC) testing that evaluates the compounds’ purity, identity, concentration and stability to ensure compliance with required standards. To measure cell viability, HepaRG cells were precultured as described above. Twenty-three nanoliters of the compounds were transferred using the Pintool station (Wako Automation, San Diego, CA). The cells were incubated at 37 °C for 24 h. Four μL of CellTiter-Glo reagent (Promega) was added to each well, and the plate was incubated for 30 min at RT. The luminescence signal was read using the ViewLux plate reader (PerkinElmer, Shelton, CT).

Mitochondrial Membrane Potential (MMP) Assay

An in vitro cell-based MMP assay was used to experimentally validate the compounds predicted to be cardiotoxic in this study. AC16 human cardiomyocyte cell line, a proliferating cell line that was derived from the fusion of primary cells from adult human ventricular heart tissues with SV40 transformed, was obtained from Sigma (Burlington, MA). Cells were cultured and maintained as per the manufacturer’s protocol. Membrane Potential Dye Kit was purchased from Dexorgen Inc. (Rockville, MD). The Mitochondrial Membrane Potential Assay is a fluorescence-based assay that quantifies the mitochondrial membrane potential changes. As previously described, , AC16 human cardiomyocyte cells were dispensed at 1000cells/5 μL/well in collagen-coated 1536-well black-clear bottom assay plate. After the assay plates were incubated at 37 °C for 18h, 23 nL of compounds dissolved in DMSO, and positive and negative controls were transferred to the assay plates using a Pintool station (Wako Automation, San Diego, CA). The assay plates were incubated for 1h at 37 °C, followed by the addition of 5 μL/well of MMP dye loading solution. After the assay plates were incubated at 37 °C for 30 min, the fluorescence intensities at 490 nm excitation and 535 nm emission and 540 nm excitation and 590 nm emission were measured using an Envision plate reader (PerkinElmer, Shelton, CT). Data was expressed as the ratio of 590 nm/535 nm emissions values. The cell viability of AC16 human cardiomyocyte cells was measured using a CellTiter-Glo reagent after 48h of compound treatment using ViewLux plate reader (PerkinElmer, Shelton, CT). Compounds were tested as 11-point 1:3 titrations, starting at various stock concentrations (ranging from 20.6 to 92.2 μM, depending on solubility), in this assay. Prior to experimental use, compounds were subjected to analytical QC testing to verify their purity, identity, concentration, and stability to ensure compliance with required standards.

In Vivo Toxicity Data

Details on the collection of in vivo DILI and DICT data have been described previously. Briefly, the in vivo toxicity data used in this study were compiled from six distinct sources to establish comprehensive reference lists for DILI and DICT. ChemIDplus (https://chem.nlm.nih.gov/chemidplus/), a publicly accessible database managed by the U.S. National Library of Medicine, aggregates chemical records, including toxicity data, from over 100 sources. The PharmaPendium database (https://www.pharmapendium.com) provided curated adverse drug effect data extracted from the U.S. Food and Drug Administration (FDA) and the European Medicines Agency (EMA) approval documents. ChemIDplus and PharmaPendium were used to retrieve drug annotations on both DILI and DICT. The FDA Center for Drug Evaluation and Research (CDER) (https://www.fda.gov/about-fda/fda-organization/center-drug-evaluation-and-research-cder) provided a list of compounds with DICT annotations. The FDA National Center for Toxicological Research (NCTR) (https://www.fda.gov/about-fda/office-chief-scientist/national-center-toxicological-research) provided reference compound lists for both DILI and DICT. A list of 130 compounds was obtained from the Enzo Life Sciences Cardiotoxicity Library (https://www.enzolifesciences.com/), which are compounds with well-documented DICT, including mitochondrial dysfunction, ion channel blockade, and fibrosis. In addition, we collected 474 compounds from the Side Effect Resource (SIDER) database (http://sideeffects.embl.de/), hosted by the European Molecular Biology Laboratory (EMBL), which had no reported cardiotoxicity, as negative controls in the DICT data set. Adverse effects in SIDER were extracted from drug package inserts and public documents to ensure reliable annotations. To standardize the data, each compound was assigned a toxicity score based on the frequency and nature of toxicity reports across these sources. ChemIDplus and PharmaPendium compounds received scores of N/10 (the score was set to 1 if N > 10), where N represented the number of toxicity reports. NCTR compounds were assigned a score of 1 for high toxicity concern, 0.5 for moderate concern, and 0 for no concern. Compounds listed as toxic in the Enzo Life Sciences Cardiotoxicity Library, CDER, and SIDER were assigned a score of 1, while nontoxic compounds were given a score of 0. For each compound, the scores from all sources were averaged to yield the final DILI or DICT score, reflecting the compound’s overall likelihood of being toxic. For DILI, compounds with scores >0.4 were classified as toxic, while those with scores ≤ 0.4 were classified as nontoxic. For DICT, compounds with scores ≥ 0.5 were classified as toxic, those with scores of 0 as nontoxic, and compounds with scores between 0 and 0.5 were excluded as inconclusive. Ultimately, a comprehensive data set comprising 1,190 compounds was curated (Excel Table S1). Of these, 1,078 compounds (541 nontoxic and 537 toxic) were assembled for the in vivo DILI reference list, while 604 compounds (235 nontoxic and 369 toxic) were compiled for the in vivo DICT reference list.

Tox21 In Vitro Assay Data

The Tox21 qHTS assay data generated on the Tox21 10K compound library were retrieved from the National Center for Advancing Translational Sciences (NCATS) Tox21 Data Browser (https://tripod.nih.gov/pubdata/, accessed on March 31, 2023). In vitro screens of the Tox21 10K library were conducted on the Tox21 robotic platform. The qHTS data were analyzed using custom software developed at NCATS. Concentration–response data were fit to a four-parameter Hill equation to determine the half-maximal activity concentration (AC50) and maximal response (efficacy). Compounds were categorized into classes 1–4 based on the shape and quality of their concentration–response curves. Compounds that showed activation were assigned class 1.1, 1.2, 1.3, 1.4, 2.1, 2.2, 2.3, 2.4, and 3 curves, while compounds that showed inhibitory effects were assigned class −1.1, −1.2, −1.3, −1.4, −2.1, −2.2, −2.3, −2.4, and −3 curves. Compounds without significant concentration responses were classified as inactive (class 4). To facilitate analysis, each curve class was associated with an efficacy cutoff and converted into a numeric curve rank. Higher ranks indicated compounds with greater potency, efficacy, and curve quality. Curve ranks ranged from −9 to 9, with inhibitors assigned negative ranks to reflect suppressive effects and activators assigned positive ranks to indicate stimulatory effects. A curve rank of 0 represented no activity observed above the background. Curve ranks from replicate compound runs were averaged. An absolute average curve rank of ≤ 0.5 indicates that the compound did not show any significant concentration response most of the time. For modeling purposes, compounds with an absolute curve rank >0.5 were labeled as active (1), while those with a curve rank ≤ 0.5 were labeled as inactive (0).

Generation of Molecular Fingerprints for Structural Representation

Compound structures were converted to Extended-Connectivity Fingerprints 4 (ECFP4), which are molecular representations that encode the structural features of each compound in a 1024-bit vector. These fingerprints were generated utilizing the Chemistry Development Kit, an open-source cheminformatics library, integrated within the Konstanz Information Miner (KNIME) software (version 4.7.1), a data analytics platform. The 2D structural representations of the compounds, encoded in the Simplified Molecular Input Line Entry System (SMILES) format, were processed to generate the corresponding ECFP4 fingerprints. Each bit in the ECFP4 fingerprint corresponds to a distinct molecular feature, with a value of 1 indicating its presence and 0 denoting its absence within the molecular structure.

Feature Selection

Four feature selection methods were performed to identify the most informative features for building machine learning models according to our previously described methods with slight modifications. ,− For the Fisher’s exact test method, features were selected at ten different p-value cutoffs, ranging from 0.01 to 0.1 with an increment of 0.01. This method allowed for the selection of features across varying degrees of statistical significance. For the area under the receiver operating characteristic curve (AUC-ROC) method, features were selected at seven different cutoff thresholds, ranging from 0.51 to 0.57 with an increment of 0.01 using the ″pROC″ package in R version 4.2.1, facilitating the evaluation of features based on their predictive power. For the Random Forest (RF) method and the eXtreme Gradient Boosting (XGBoost) method, the ″Random Forest″ package and ″xgboost″ package in R version 4.2.1 were employed, respectively. Feature importance scores, such as Gini importance or Gain scores, were calculated, and features were selected at 20 intervals from the top 20 to the top 200 ranked features. The resulting feature sets were subsequently utilized to build and evaluate machine learning models.

Machine Learning Modeling

The machine learning models were constructed using in vitro assays and/or ECFP4 fingerprints as input features, with in vivo toxicity data (DILI or DICT) serving as the toxicity end point. The modeling process followed the methodologies outlined in our previous studies. ,− Briefly, five classification models were constructed, including Naïve Bayes (NB) and Support Vector Machine (SVM) utilizing the ″e1071″ package, Neural Networks (NNET) via the ″nnet″ package, RF employing the ″Random Forest″ package, and XGBoost with the ″xgboost″ package. The implementation of the NB classifier was adapted with the incorporation of Laplace smoothing to address issues related to zero probabilities. Furthermore, the SVM classifier employed the Gaussian radial basis function kernel. The remaining parameters of the SVM, RF, and NNET classifiers were set to their default values. For the XGBoost classifier, the following parameters were specified: a maximum depth of a tree set to 3, a learning rate of 0.01 to control the step size during the optimization process, and a subsample ratio of columns of 0.5 when constructing each tree. For model evaluation, the data set was randomly split into training (70%) and test (30%) sets. This process was repeated 20 times with different random seeds to assess the model stability. The evaluation of model performance was conducted by calculating the AUC-ROC value and balanced accuracy (BA) using the “pROC” packages, and the Matthews correlation coefficient (MCC) using the “mltools” package. The entire machine learning modeling process was executed in R version 4.2.1.

DILI and DICT Potential Prediction of the Tox21 10K Compound Library

The best-performing models were applied to predict the hepatotoxicity and cardiotoxicity potentials of compounds in the Tox21 10K compound library. To improve the reliability and accuracy, consensus predictions from multiple individual models were utilized. Specifically, compounds predicted as positive by multiple models were first assigned consensus scores. The consensus score for a given compound was calculated as the sum of its probability scores from multiple models weighted by the respective AUC-ROC value of each model. This consensus scoring strategy was designed to leverage the complementary strengths of the models while minimizing biases or limitations inherent in individual predictions. To identify a structurally diverse set of candidate compounds for experimental validation, the compounds were clustered based on structural similarity using the k-means algorithm (k = 50), which partitioned the compounds into 50 clusters. Representative compounds from each cluster were then selected, prioritizing those exhibiting the highest consensus scores within their cluster. This clustering approach was applied to maximize the chemical space coverage with a limited set of compounds. The prioritized candidates were then subjected to experimental validation using in vitro assays.

Statistical Analysis

qHTS assay data from experimental validation were analyzed following the methodologies outlined in previous studies. ,, Concentration–response curves were fit using the four-parameter logistic regression model. In this model, the percentage of assay activity served as the dependent variable, and the log-transformed compound concentration (base 10) was the independent variable. These analyses were conducted using the ″drc″ statistical package within R version 4.2.1. Visualization of the data was accomplished through plots generated by the ″ggplot2″ package in R version 4.2.1. Additionally, representative chemical structures were illustrated using the ChemDraw Professional software, version 22.0. The comprehensive workflow, including data collection, model development, performance evaluation, and experimental validation, is illustrated in Figure S1.

Results

Composition of the Tox21 In Vitro Assays and Their Corresponding Biological Targets

The current Tox21 assay panel contains 49 targets with 262 assay readouts, 227 of which were from old (pre-existing) assays and 35 were from new assays (newly added to expand the assay panel) (Excel Table S2). As each target may have multiple assays and each assay may have multiple readouts (e.g., the target adrenoceptor beta 2 (ADRB2) had two probing assays: tox21-adrb2-antagonist-p1 and tox21-adrb2-agonist-p1; and each ADRB2 assay had three readouts: tox21-adrb2-agonist-p1_ratio, tox21-adrb2-agonist-p1_ch1, and tox21-adrb2-agonist-p1_ch2), there were 36 biological targets covered by the old assays, and 13 additional targets covered by the new assays, as shown in Figure A and Excel Table S2. Of these, the biological targets covered by the new assays included ADRB2, cholinergic receptor muscarinic 1 (CHRM1), cAMP responsive element binding protein 1 (CREB1), dopamine receptor D2 (DRD2), gonadotropin-releasing hormone receptor (GnRHR), human ether-à-go-go-related gene (hERG), 5-hydroxytryptamine receptor 2A (HTR2A), kisspeptin 1 receptor (KISS1R), cytochrome P450 2C9 (CYP2C9), cytochrome P450 2D6 (CYP2D6), cytochrome P450 3A4 (CYP3A4), cytochrome P450 1A2 (CYP1A2), and cytochrome 450 2C19 (CYP2C19). Figure B and C show the biological target spaces covered by the old and new assay panels, respectively. The old assay panel comprised eight distinct target categories, with nuclear receptor (NR) signaling as the most predominant category (52.1%) followed by stress response (SR) pathways (12.3%; Figure B). In comparison, the new assay panel incorporated two new categories (Cardiotoxicity and Metabolism) and had a marked increase in the number of GPCR targets (from 5.5% to 18.5%), making it the second largest category following NR (41.3%) (Figure C). This resulted in a substantial improvement in drug space coverage, represented by the GPCRs, in the Tox21 assays.

1.

1

Overview of Tox21 assay targets. (A) Composition of Tox21 assay targets, with the term ″old″ referring to pre-existing assay targets and the term ″new″ referring to newly introduced assay targets. (B) Target categories of the pre-existing Tox21 assays. (C) Target categories of the expanded Tox21 assays. The numerical values displayed on the pie chart represent the number of distinct assays within each respective target category. Summary data is available in Excel Table S2. Note: NR, nuclear receptor; SR, stress response; GPCR, G protein-coupled receptor.

Best-Performing Models for Predicting DILI/DICT Potential

Ninety-five best-performing models for predicting DILI or DICT potential were generated by employing a pairwise combination of five machine learning algorithms and four feature selection methods based on five input data types, namely, assay only (old), assay only (new), structure only, assay (old) + structure, and assay (new) + structure (Excel Tables S3, S4, and S5). Among them, five best-performing models derived from each input data type for the prediction of DILI and DICT potentials are shown in Figures A and A, respectively. The term “old” here referred to the pre-existing assays, and the term ″ new″ here referred to the expanded assay panel that encompassed the pre-existing assays and the newly introduced assays. For the prediction of DILI potential, models that were built based on chemical structures with either old or new assay data, utilizing the XGBoost algorithm, demonstrated superior performance with AUC-ROC of 0.74 ± 0.03, BA of 0.70 ± 0.02, and MCC of 0.41 ± 0.04. This performance surpassed the results obtained from models that were built based on chemical structures only (AUC-ROC = 0.73 ± 0.03, BA = 0.70 ± 0.02, and MCC = 0.40 ± 0.04), or assay only (old/new) (AUC-ROC of 0.65 ± 0.03, BA of 0.62 ± 0.02, and MCC of 0.26 ± 0.04). For the prediction of DICT potential, models that were built based on chemical structures with either old or new assay data, utilizing the XGBoost algorithm, exhibited good performance with AUC-ROC of 0.82 ± 0.02, BA of 0.76 ± 0.02, and MCC of 0.53 ± 0.05. This performance was improved compared to the results obtained from models that were built based on chemical structures alone (AUC-ROC = 0.80 ± 0.04, BA = 0.76 ± 0.03, and MCC = 0.51 ± 0.06), or assay data alone (old/new) (AUC-ROC of 0.69 ± 0.02, BA of 0.66 ± 0.02, and MCC of 0.34 ± 0.04). In addition to exhibiting robust predictive power, the XGBoost algorithm demonstrated exceptional performance as a feature selection method across five distinct feature categories, such as assay only (old/new), structure only, or a combination of both (structure + assay data; Figures A and A, respectively). The parameters and data composition used to generate the best-performing models for both DILI and DICT prediction are summarized in Table S4, with a detailed list of features utilized to construct the models provided in Table S5. Comparing model performance with and without feature selection (Excel Table S6) showed that feature selection significantly improved performance by identifying the most relevant features from the assay only, structure only, and assay plus structure feature sets.

2.

2

Optimal models for predicting DILI potential. (A) Model performances with their parameter settings. Results are presented as mean ± standard deviation (SD), and the error bars represent the SD of 20 model iterations. Summary data is available in Excel Table S3 and S4. (B) Details of biological targets in old and new Tox21 assay panels. The term “old” here refers to the pre-existing Tox21 assays, and the term ″new″ here refers to the expanded Tox21 assay panel that encompasses the pre-existing assays and the newly introduced assays. Summary data is available in Excel Table S5. Note: AUC-ROC, area under the receiver operating characteristic curve; BA, balanced accuracy; MCC, Matthews correlation coefficient; RF, random forest; XGBoost, eXtreme gradient boosting.

3.

3

Optimal models for predicting DICT potential. (A) Model performances with their parameter settings. Results are presented as mean ± standard deviation (SD), and the error bars represent the SD of 20 model iterations. Summary data is available in Excel Table S3 and S4. (B) Details of biological targets in old and new Tox21 assay panels. The term “old” here refers to the pre-existing Tox21 assays, and the term ″ new″ here refers to the expanded Tox21 assay panel that encompasses the pre-existing assays and the newly introduced assays. Summary data is available in Excel Table S5. Note: AUC-ROC, area under the receiver operating characteristic curve; BA, balanced accuracy; MCC, Matthews correlation coefficient; NB, Naïve Bayes; XGBoost, eXtreme gradient boosting.

The performances of the best-performing models based on the old and new assay panels were comparable, with the new assay panel exhibiting slightly better metrics. However, a notable difference was observed in the number of assays required to generate the best-performing models based on the new compared to old assay data, both for predicting DILI and DICT potential (Figures B and B, respectively). For the prediction of DILI potential, 40 assay readouts from the new assay panel were needed to produce the best-performing model (i.e., the XGBoost algorithm) with AUC-ROC = 0.65 ± 0.03, BA = 0.62 ± 0.02, and MCC = 0.26 ± 0.04, while a much larger number of assay readouts (60 in total) were required to produce the best-performing model using the old assay panel (i.e., RF algorithm) with AUC-ROC = 0.62 ± 0.03, BA = 0.61 ± 0.02, and MCC = 0.23 ± 0.05. Compared to the old assay panel, the additional assay targets required for the best-performing model based on the new assay panel included HTR2A, ADRB2, CHRM1, CYP1A2, CYP2C9, CYP3A4, hERG, and KISS1R (Figure B). For the prediction of DICT potential, 20 assay readouts in the new assay panel were needed for the best-performing model (i.e., XGBoost algorithm) with AUC-ROC = 0.69 ± 0.02, BA = 0.66 ± 0.02, and MCC = 0.34 ± 0.04, while a much larger number of assay readouts (60 in total) were required for the best-performing model based on the old assay data (i.e., the XGBoost algorithm) with AUC-ROC = 0.66 ± 0.04, BA = 0.65 ± 0.03, and MCC = 0.30 ± 0.07. Compared to the old assay panel, the additional assay targets required for the best-performing model based on the new assay panel included ADRB2, CREB1, CYP1A2, CYP2C19, CYP2D6, CYP3A4, DRD2, and GnRHR (Figure B).

Experimental Validation of Predicted Hepatotoxic Compounds

A total of 5,570 compounds within the Tox21 compound library was predicted using the best-performing models developed based on the “assay (new) + structure” data set (Excel Table S7). Among them, a group of 234 structurally diverse compounds exhibiting the highest consensus scores was prioritized for experimental validation. The consensus score for each compound was obtained using the predicted probabilities derived from the NB algorithm and the XGBoost algorithm multiplied by their respective AUC-ROCs. The 234 selected compounds were tested for reactive oxygen species (ROS) induction in HepaRG cells, 135 of which were identified as toxic, yielding a confirmation rate of 58%. In addition, 40 compounds were identified as nontoxic, while 59 compounds showed inconclusive results (Excel Table S8). Nine representative hepatotoxic compounds were categorized based on their IC50 values using 2 and 10 μM as the thresholds (Figure ). The compound that showed the most potent ROS activity (IC50 < 2 μM) was darbufelone mesylate (NCGC00254170, IC50 = 0.83 ± 0.54 μM, efficacy = −65.64 ± 41.15%, consensus score = 1.44). Compounds demonstrating moderate to strong ROS activity (2 μM < IC50 < 10 μM) included canadine (NCGC00016486, IC50 = 2.69 ± 1.46 μM, efficacy = −73.85 ± 51.44%, consensus score = 1.46), C.I. basic violet 14 (NCGC00166015, IC50 = 4.96 ± 3.24 μM, efficacy = −52.22 ± 76.88%, consensus score = 1.43), meprednisone (NCGC00182851, IC50 = 5.27 ± 4.48 μM, efficacy = −53.63 ± 71.23%, consensus score = 1.45), nefiracetam (NCGC00186006, IC50 = 5.93 ± 2.23 μM, efficacy = −89.82 ± 11.90%, consensus score = 1.45), netilmicin sulfate (NCGC00016881, IC50 = 6.21 ± 2.81 μM, efficacy = −91.39 ± 28.05%, consensus score = 1.44), and nocodazole (NCGC00015647, IC50 = 9.61 ± 5.86 μM, efficacy = −105.97 ± 13.33%, consensus score = 1.46). Lastly, a subset of compounds exhibited weak ROS activity (IC50 > 10 μM), comprising acibenzolar-s-methyl (NCGC00254768, IC50 = 10.67 ± 6.41 μM, efficacy = −87.36 ± 16.86%, consensus score = 1.43), and ciglitizone (NCGC00025104, IC50 = 11.14 ± 6.26 μM, efficacy = −82.56 ± 15.56%, consensus score = 1.45).

4.

4

Concentration–response curves of representative predicted hepatotoxic compounds with ROS activity in HepaRG cells. Results are presented as mean ± standard deviation (SD), and the error bars represent the SD of six independent experiments. Summary data is available in Excel Table S8.

Experimental Validation of Predicted Cardiotoxic Compounds

A total of 5,996 compounds within the Tox21 compound library was predicted using the best-performing models developed based on the “assay (new) + structure” data set (Excel Table S9). Among them, a group of 247 structurally diverse compounds exhibiting the highest consensus scores was prioritized for experimental validation. The consensus score for each compound was obtained using the predicted probabilities derived from the NB algorithm and the XGBoost algorithm multiplied by their respective AUC-ROCs. The 247 selected compounds were tested for mitochondrial membrane potential (MMP) disruption using the AC16 human cardiomyocyte cell line, and 61 compounds (25%) decreased MMP. In addition, 123 compounds were inactive, while 63 compounds showed inconclusive activities (Excel Table S10). Nine representative compounds were categorized based on their IC50 values according to thresholds of 2 and 5 μM (Figure ). The compounds that showed the most potent MMP reduction (IC50 < 2 μM) included tributyltin acrylate (NCGC00256111, IC50 = 0.48 ± 0.00 μM, efficacy = −50.14 ± 1.47%, consensus score = 1.63), azacyclonol (NCGC00016366, IC50 = 0.69 ± 0.45 μM, efficacy = −47.39 ± 3.07%, consensus score = 1.62), and frentizole (NCGC00160657, IC50 = 1.73 ± 0.12 μM, efficacy = −90.03 ± 1.16%, consensus score = 1.56). Compounds demonstrating moderate to strong MMP inhibition (2 μM < IC50 < 5 μM) included azelnidipine (NCGC00167436, IC50 = 2.56 ± 0.44 μM, efficacy = −56.40 ± 4.68%, consensus score = 0.82), lapatinib (NCGC00167507, IC50 = 2.65 ± 0.30 μM, efficacy = −105.33 ± 2.33%, consensus score = 1.63), and 4-nonylphenol (NCGC00090918, IC50 = 4.35 ± 0.29 μM, efficacy = −83.89 ± 0.68%, consensus score = 1.63). Lastly, a subset of compounds exhibited weak MMP inhibition (IC50 > 10 μM), comprising 2-ethylhexylparaben (NCGC00256492, IC50 = 11.64 ± 0.76 μM, efficacy = −101.59 ± 2.42%, consensus score = 1.61), riboflavin (NCGC00017291, IC50 = 16.63 ± 0.00 μM, efficacy = −116.43 ± 3.01%, consensus score = 1.63), and hexylparaben (NCGC00257292, IC50 = 17.62 ± 5.67 μM, efficacy = −96.47 ± 12.76%, consensus score = 1.63).

5.

5

Concentration–response curves of representative predicted cardiotoxic compounds that inhibited MMP in AC16 human cardiomyocytes. Results are presented as mean ± standard deviation (SD), and the error bars represent the SD of three independent experiments. Summary data is available in Excel Table S10.

Discussion

To explore whether expanding Tox21 assay panel with more relevant targets could improve the predictive capacity of in vitro assay data for in vivo DILI and DICT or help deconvolute the molecular mechanisms of DILI and DICT, we compared the performances of machine learning models that were built based on the old (pre-existing) and new (expanded) Tox21 assay panels, respectively. In addition, the compounds identified by the best-performing models were experimentally validated using the ROS assay for hepatotoxicity and the MMP assay for cardiotoxicity. Through this study, we aimed to evaluate the role and value of the expanded Tox21 in vitro assay panel in optimizing predictive models for DILI and DICT potential.

The pre-existing Tox21 assays have been previously utilized in the construction of machine learning models to predict in vivo human organ level toxicity of compounds and to elucidate their underlying molecular mechanisms. , For example, our previous study demonstrated the effectiveness of assay-only data in toxicity prediction across organs, achieving strong performance for endocrine (AUC-ROC = 0.83 ± 0.02) and peripheral nerve toxicity (AUC-ROC = 0.79 ± 0.02), with moderate success in musculoskeletal (AUC-ROC = 0.72 ± 0.03) and urological systems (AUC-ROC = 0.72 ± 0.02). Additionally, in this previous study, we found that the hypoxia response element (HRE) assay emerged as a significant contributor to vascular toxicity, where hypoxia inducible factor 1 alpha (HIF-1α) binding to HRE enhances vascular endothelial growth factor (VEGF) transcription, a crucial growth factor driving angiogenesis and vascularization. While useful for mechanism deconvolution, the models constructed on assay data alone showed moderate predictive performance for predicting DILI and DICT. To address this issue, we systematically identified additional targets and pathways that were previously under-explored to improve the biological response space coverage of the Tox21 assay panel. The newly identified targets and pathways covered mostly the drug target space, which was barely represented by the old assay panel. In this study, we evaluated the expanded Tox21 assay panel with the addition of new assays probing 13 biological targets spanning four crucial categories of biological functions (Figure ). These categories encompassed G protein-coupled receptors (GPCRs) (including ADRB2, CHRM1, DRD2, GnRHR, HTR2A, and KISS1R), which are important drug targets nonspecific activity on which can lead to undesirable side-effects and other liabilities, transcription factors (including CREB1) regulating gene expression in response to cellular stress, ion channels (including hERG) crucial for neuronal and cardiovascular function, inhibition of which could result in cardiotoxicity, and the family of cytochrome P450 enzymes (including CYP2C9, CYP2D6, CYP3A4, CYP1A2, and CYP2C19), which are responsible for the metabolism of approximately two-thirds of known drugs in humans.

The best-performing models based on assay data alone showed slightly improved performance for predicting DILI or DICT when the expanded assay panel was used, though the performance was still less than ideal (Figures A and A), indicating that the few targets added so far may not provide sufficient coverage of the key toxicity mechanisms involved in DILI and DICT. Although the improvement gained by integrating assay data with structural information is modest compared to using structural data alone, the inclusion of assay data offers critical mechanistic insights by highlighting the contribution of specific assay targets to model performance, i.e., their role in the prediction of in vivo toxicity. For example, the Aryl hydrocarbon Receptor (AhR), a ligand-activated transcription factor expressed in various tissues such as the liver and heart, emerged in our study as a significant predictor of DILI and DICT potential (Excel Table S5). This finding aligns with previous evidence that persistent or excessive activation of AhR, particularly by environmental contaminants like dioxins, can trigger hepatotoxicity and cardiotoxicity through reactive metabolite formation and mitochondrial dysfunction. In vitro assays are often subject to variability arising from changes in experimental conditions, technical noise, and batch effects, introducing uncertainties that can affect the reliability of model predictions. In contrast, QSAR-based models rely exclusively on chemical structure information, making them less susceptible to such variability and potentially more robust under these conditions. This could be a contributing factor to the marginal performance improvement observed when assay data were introduced to the models based on chemical structure alone.

Nonetheless, models based on the expanded assay panel required far fewer assay readouts (40 or 20 assay readouts for predicting DILI and DICT potential, respectively) to achieve performance comparable to or better than that of the models based on the pre-existing assay panel that required 60 assay readouts for predicting either DILI or DICT potential. These results suggested that the expanded Tox21 assay panel is more efficient in capturing the relevant mechanisms for DILI and DICT compared to the pre-existing assay panel. Eight new targets involved in the expanded assay panel were selected for constructing the best-performing model for predicting DILI, including HTR2A, ADRB2, CHRM1, hERG, KISS1R, CYP1A2, CYP2C9, and CYP3A4 (Figure B), all of which have been implicated in liver function. For example, Choi et al. found that HTR2A mediated the lipogenic action of gut-derived serotonin in the liver. Wu et al. found that ADRB2 signaling negatively regulated autophagy, leading to hypoxia-inducible factor-1α stabilization, and reprogramming of hepatocellular carcinoma cells glucose metabolism. Rachakonda et al. found that deficiencies in M1 muscarinic receptors (encoded by CHRM1) modulated chronic liver injury by activating antioxidant responses, and attenuated hepatocyte death by reducing caspase-3 activation. Xia et al. found that potassium ions could inhibit tumorigenesis by upregulating the potassium ion transport channel protein hERG and voltage-dependent anion-selective channel protein 1. Guzman et al. found that hepatic KISS1R signaling activated the master energy regulator, AMP-activated protein kinase, to thereby decrease lipogenesis and progression to nonalcoholic steatohepatitis. Klyushova et al. summarized that the hepatic CYP enzymes (including CYP1A2, CYP2C9, and CYP3A4) could metabolize arachidonic acid into epoxyeicosatrienoic acids that play an important part in blood pressure regulation, inflammation, and cell proliferation. In addition to DILI, eight new targets in the expanded assay panel were selected for constructing the best-performing model for predicting DICT, including ADRB2, CREB1, DRD2, GnRHR, CYP1A2, CYP2C19, CYP2D6, and CYP3A4 (Figure B), all of which have been reported to have connections to the heart. For example, Sarpeshkar et al. found that the polymorphisms of ADRB2 receptors could influence the cardiovascular, respiratory, metabolic and musculoskeletal systems to enhance endurance phenotypes. Matus et al. found that the transcription factor CREB1 played a critical role in regulating gene expression in response to activation of the cAMP-dependent signaling pathway, which is implicated in the pathophysiology of heart failure. Dehelean et al. found that at the cardiac level, DRD2 blockade decreased the inotropic effect and favored the installation of dilated cardiomyopathy with risk of myocardial infarction or severe rhythm disorders, and even sudden death. Wang et al. found that sevoflurane might exert important roles in cardioprotective effects, at least in part, through targeting growth hormone secretagogue receptor and GnRHR. Fatunde et al. summarized that the CYP enzymes (e.g., CYP1A2, CYP2C19, CYP2D6, and CYP3A4) were primarily located in the liver, but could also be found in the small intestine, lung, kidney and even the heart.

Since the assay data used to train the models originated from the Tox21 high-throughput screening program, a key limitation is that the models’ predictive performance may be constrained to the chemical space covered by the Tox21 compound library, potentially reducing their generalizability to chemicals beyond this space. To determine the applicability domain (AD) of the structure-based models, we identified the closest structural analogue in the training set for each compound to be predicted. Structural similarity was quantified using the Tanimoto coefficient based on ECFP4 fingerprints, calculated as the ratio of shared structural features to unique features between two compounds. The coefficient ranged from 0 (no shared features) to 1 (identical compounds). For this study, compounds with maximum Tanimoto similarity (Tmax) ≥ 0.4 to any training set compound were considered within the model’s AD. This threshold was selected to reflect sufficient structural similarity for reliable predictions while acknowledging that compounds outside this domain may carry greater uncertainty. A model is considered ″robust” in this study when it demonstrated good predictive performance on compounds that represent diverse chemical space. In addition, the Uniform Manifold Approximation and Projection (UMAP) analysis of the ECFP4 fingerprints of the active and inactive compounds in both the DILI and DICT training sets (Figure S2) revealed that the active compounds are not structurally distinct from the inactive compounds, indicating unbiased chemical space coverage and supporting the robustness of our models.

To further assess the robustness of the models, Y-randomization was conducted by randomly shuffling the response variables 20 times. The randomized data yielded poor model performance, with AUC-ROC values ranging from 0.51 to 0.54, close to the performance of a random classifier (0.5; Excel Table S11). These results confirmed that our best-performing model’s predictive power stems from genuine patterns rather than chance correlation. Hyperparameter tuning was conducted for the best-performing DICT and DILI models, including NNET, RF and SVM, based on the ″assay (new) + structure″ data set. A grid search was used to identify the best-performing configurations for each model. The results indicated that additional tuning of parameters did not further improve the models’ performance compared to the default configurations for DILI and DICT prediction used in this study (Excel Table S12, Figures A and A).

Consistent with our previous study, the combination of assay data (either new or old) and structural information can produce robust models for DILI and DICT prediction (Figures and ), showing the promise to improve the efficiency of the drug development process. In addition, we compared the experimental validation results with the predictions of the best-performing models based on the ″assay (old) + structure″ for both DILI and DICT potential. For DILI, the ″assay (old) + structure″ consensus model, incorporating NB and XGBoost algorithms, predicted 201 out of the 234 compounds selected by the ″assay (new) + structure″ model as positive. Experimental validation confirmed 121 of these as toxic, yielding a confirmation rate of 60%. Among the 33 compounds predicted as negative by the ″assay (old) + structure″ consensus model, 14 (42%) were experimentally confirmed as toxic (Excel Table S13). For DICT, the same consensus model predicted 216 out of the 247 compounds selected by the ″assay (new) + structure″ model as positive. Of these, 53 were validated as toxic, resulting in a confirmation rate of 25%. Among the 31 compounds predicted as negative by the ″assay (old) + structure″ model, 8 (26%) were experimentally confirmed as toxic (Excel Table S14). These results show that the ″assay (new) + structure″ and ″assay (old) + structure″ models achieved comparable positive predictive values, whereas the ″assay (new) + structure″ model appeared to be more sensitive (i.e., lower false negative rates) in detecting potentially toxic compounds.

Consensus models, which combine the predictions of multiple individual models, have been demonstrated to yield superior performance compared to the constituent models employed in isolation. , In this study, the compounds in the Tox21 compound library were assessed for their hepatotoxicity and cardiotoxicity potential using consensus models that combined the NB algorithm and the XGBoost algorithm. The two algorithms could leverage their respective strengths, enhancing the overall predictive accuracy for the critical tasks of the DILI and DICT assessment. Consensus models often outperformed individual models through their ability to integrate multiple perspectives and approaches. This advantage has been demonstrated in the “Tox21 Data Challenge 2014”, where consensus models not only performed well but also outperformed some of the winning individual models, exemplifying the “wisdom of the crowd” principle. Similar conclusions were drawn from other studies including the Collaborative Modeling Project for Androgen Receptor Activity (CoMPARA) and the Collaborative Estrogen Receptor Activity Prediction Project (CERAPP). Based on the positive observations from these studies, we employed the consensus modeling approach in the current study to identify potential hepatotoxic and cardiotoxic compounds for experimental validation.

The best-performing DILI and DICT models used different subsets of the Tox21 assays, and different assays tested slightly different numbers of compounds due to updates to the Tox21 library composition overtime. Hepatotoxicity and cardiotoxicity predictions were made only on compounds with no missing data in the subsets of assays employed by the respective models. For example, 518 compounds had complete data in the assays used to build the DICT model but had missing data in some of the assays used by the DILI model (Excel Table S1). As a result, the numbers of compounds with hepatotoxicity and cardiotoxicity predictions are different (5,570 for hepatotoxicity and 5,996 for cardiotoxicity). A total of 462 predicted toxic compounds were selected to experimentally assess their hepatotoxicity or cardiotoxicity potential using the ROS assay and the MMP assay, respectively (Figures and , Excel Tables S8 and S10). Oxidative stress induced by ROS has been implicated in causing varying degrees of hepatocellular damage, potentially playing a crucial role in the pathogenesis of various liver injuries. Some compounds predicted as hepatotoxic in this study have been previously reported to be associated with ROS production. For instance, Signoretto et al. demonstrated that nocodazole (NCGC00015647), a microtubule assembly inhibitor, induced ROS formation in human erythrocytes, suggesting its potential role in oxidative stress-mediated cellular damage. Sakata et al. discovered that acibenzolar-s-methyl (NCGC00254768), a plant activator, systemically activated stomatal-based defense in Japanese radish by inducing peroxidase-dependent ROS production. Furthermore, Kwon et al. revealed that ciglitazone (NCGC00025104) induced apoptosis via activation of p38 mitogen-activated protein kinases and apoptosis inducing factor mediated by ROS and Ca2+ in opossum kidney cells, emphasizing the role of ROS in cellular toxicity and apoptotic pathways. MMP plays a pivotal role in maintaining cellular function, particularly in highly energy-demanding cells like cardiomyocytes. Disruption of MMP by drugs or other agents could lead to cellular dysfunction and potentially irreversible damage, including cardiotoxicity. Some predicted cardiotoxic compounds in this study were previously reported to be associated with MMP disruption. In a study conducted by Roos et al., lapatinib (NCGC00167507) was found to impair mitochondrial function in a hepatocellular carcinoma-derived cell line. This impairment was accompanied by the accumulation of ROS and the release of cytochrome C from mitochondria into the cytosol, ultimately inducing apoptosis. Another example of a compound that affected the MMP is 4-nonylphenol (NCGC00090918). Li et al. demonstrated that treatment with 4-nonylphenol decreased the MMP levels of pancreatic islets in a dose-dependent manner. Furthermore, their study revealed that 4-nonylphenol-induced pancreatic damage in rats through mitochondrial dysfunction and oxidative stress, leading to disruption of glucose tolerance and a decrease in insulin secretion. Other compounds that have not been documented (e.g., dabafenone mesylate (IC50 = 0.83 ± 0.54 μM in the ROS assay), and tributyltin acrylate (IC50 = 0.48 ± 0.00 μM in the MMP assay)) should be prioritized for further evaluation of their potential hepatotoxicity or cardiotoxicity.

Drug-induced toxicity necessitates a thorough and multifaceted assessment due to its complex nature, which encompasses various organ-specific toxicities, such as hepatotoxicity and cardiotoxicity. Although in vitro assays, including the ROS assay and MMP assay, could discern toxic compounds that act through certain mechanisms, they are not able to capture all aspects of chemical-induced toxicity and their predictive value for in vivo toxicity remains limited. The extrapolation of in vitro findings to in vivo scenarios is significantly influenced by various factors, such as compound metabolism, distribution, and elimination, which are not adequately captured in isolated cell-based systems. , Furthermore, the intricate interplay between different cell types and organ systems, as well as the complex physiological and pathophysiological processes involved, significantly influences the manifestation of toxicity in vivo. As a result, many of the compounds predicted as toxic by our models could not be validated by either the ROS assay for hepatotoxicity or the MMP assay for cardiotoxicity. There is no one in vitro assay that can serve as the surrogate for an in vivo toxicity end point. This discordance implied that the compounds predicted as hepatotoxic or cardiotoxic in this study necessitate additional studies to accurately ascertain their DILI or DICT potential.

Conclusions

We developed robust models for predicting DILI and DICT potentials that can be applied to identify potential hepatotoxic or cardiotoxic compounds. Experimental validation was performed on selected prediction results. A much smaller number of assays were needed in producing the best-performing models for predicting DILI or DICT when new assay targets were added to the Tox21 assay panel. The expanded Tox21 assay panel with improved coverage of the biological response space with targets relevant for the under-represented toxicity mechanisms shows promise for more efficient prediction of in vivo toxicity such as DILI and DICT. The results of this study serve as a proof of concept that sufficient coverage of the biological response space is essential in enhancing the predictive capacity of in vitro data for in vivo toxicity.

Supplementary Material

hp6c00052_si_001.docx (885KB, docx)
hp6c00052_si_002.xlsx (2MB, xlsx)

Acknowledgments

This work was supported by the Intramural Research Programs of the National Toxicology Program (Interagency agreement #Y2-ES-7020–01), National Institute of Environmental Health Sciences (NIEHS) and the National Center for Advancing Translational Sciences (NCATS), National Institutes of Health (NIH). The authors would like to thank Caitlin Lynch for performing CAR assay screening, Shuaizhang Li for performing AChE assay screening, and Hsiu-Ling Lin and Abhinav Asthana for compound management. The contributions of the NIH author(s) were made as part of their official duties as NIH federal employees, are in compliance with agency policy requirements, and are considered Works of the United States Government. However, the findings and conclusions presented in this paper are those of the authors and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.

All data needed to evaluate the conclusions in the paper are present in the paper and/or the . The codes used for modeling have been deposited in GitHub at https://github.com/TX-2017/machine-learning.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/EHP.6c00052.

  • Comprehensive workflow of the study; UMAP analysis of ECFP4 fingerprints illustrating the chemical space coverage and distribution of active and inactive compounds in both DILI and DICT training sets (DOCX)

  • Comprehensive composition of the Tox21 assay dataset and in vivo toxicity profiles; distribution analysis of new (newly developed) and old (pre-existing) Tox21 assays; optimal AUC-ROC values, balanced accuracy (BA), and Matthews correlation coefficient (MCC) for each endpoint; parameters and data composition of the best-performing model; feature combinations used to generate the best-performing models; model performance obtained by modeling with all features (no feature selection); prediction of DILI potential within the Tox21 10K compound library utilizing assay (new) and structural data; experimental validation of model-predicted hepatotoxic compounds using the ROS assay; prediction of DICT potential within the Tox21 10K compound library utilizing assay (new) and structural data; experimental validation of model-predicted cardiotoxic compounds using the MMP assay; model performance after y-randomization; model performance under different hyperparameter configurations; prediction of DILI potential within the Tox21 10K compound library utilizing assay (old) and structural data; prediction of DICT potential within the Tox21 10K compound library utilizing assay (old) and structural data (XLSX)

T.X. performed statistical analysis of all data. M.O., J.Z., S.S., L.Z., S.Y., and H.Z. conducted the experiments. D.K.N. and T.Z. aided data collection and analysis. J.T. and C.K.-T. aided experimental data generation. T.X. and R.H. wrote the manuscript. M.X., N.D.S., and M.D.H. designed experiments and directed the generation of experimental data. R.H., S.F., D.R. and A.S. designed the study and directed the project. All authors reviewed the manuscript.

The authors declare no competing financial interest.

References

  1. Patton K., Borshoff D.. Adverse drug reactions. Anaesthesia. 2018;73:76–84. doi: 10.1111/anae.14143. [DOI] [PubMed] [Google Scholar]
  2. Funk C., Roth A.. Current limitations and future opportunities for prediction of DILI from in vitro . Arch. Toxicol. 2017;91(1):131–142. doi: 10.1007/s00204-016-1874-9. [DOI] [PubMed] [Google Scholar]
  3. Connolly S. J., Camm A. J., Halperin J. L., Joyner C., Alings M., Amerena J.. et al. Dronedarone in high-risk permanent atrial fibrillation. N Engl J. Med. 2011;365(24):2268–2276. doi: 10.1056/NEJMoa1109867. [DOI] [PubMed] [Google Scholar]
  4. Andrade R. J., Chalasani N., Björnsson E. S., Suzuki A., Kullak-Ublick G. A., Watkins P. B.. et al. Drug-induced liver injury. Nat. Rev. Dis Primers. 2019;5(1):58. doi: 10.1038/s41572-019-0105-0. [DOI] [PubMed] [Google Scholar]
  5. Varga Z. V., Ferdinandy P., Liaudet L., Pacher P.. Drug-induced mitochondrial dysfunction and cardiotoxicity. Am. J. Physiol Heart Circ Physiol. 2015;309(9):H1453–H1467. doi: 10.1152/ajpheart.00554.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. McGill M. R., Jaeschke H.. Animal models of drug-induced liver injury. Biochim Biophys Acta Mol. Basis Dis. 2019;1865(5):1031–1039. doi: 10.1016/j.bbadis.2018.08.037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Meng C., Fan L., Wang X., Wang Y., Li Y., Pang S.. et al. Preparation and evaluation of animal models of cardiotoxicity in antineoplastic therapy. Oxid Med. Cell Longev. 2022;2022(1):3820591. doi: 10.1155/2022/3820591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Olson H., Betton G., Robinson D., Thomas K., Monro A., Kolaja G.. et al. Concordance of the toxicity of pharmaceuticals in humans and in animals. Regul. Toxicol. Pharmacol. 2000;32(1):56–67. doi: 10.1006/rtph.2000.1399. [DOI] [PubMed] [Google Scholar]
  9. Ye L., Ngan D. K., Xu T., Liu Z., Zhao J., Sakamuru S.. et al. Prediction of drug-induced liver injury and cardiotoxicity using chemical structure and in vitro assay data. Toxicol. Appl. Pharmacol. 2022;454:116250. doi: 10.1016/j.taap.2022.116250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Huang R., Xia M., Sakamuru S., Zhao J., Shahane S. A., Attene-Ramos M.. et al. Modelling the Tox21 10 K chemical profiles for in vivo toxicity prediction and mechanism characterization. Nat. Commun. 2016;7(1):10425. doi: 10.1038/ncomms10425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Xu, T. ; Xia, M. ; Huang, R. . Modeling Tox21 Data for Toxicity Prediction and Mechanism Deconvolution. In Machine Learning and Deep Learning in Computational Toxicology; Springer: 2023; pp 463–477. [Google Scholar]
  12. Richard A. M., Huang R., Waidyanatha S., Shinn P., Collins B. J., Thillainadarajah I.. et al. The Tox21 10K Compound Library: Collaborative Chemistry Advancing Toxicology. Chem. Res. Toxicol. 2021;34(2):189–216. doi: 10.1021/acs.chemrestox.0c00264. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Gorshkov K., Sima N., Sun W., Lu B., Huang W., Travers J.. et al. Quantitative Chemotherapeutic Profiling of Gynecologic Cancer Cell Lines Using Approved Drugs and Bioactive Compounds. Transl Oncol. 2019;12(3):441–452. doi: 10.1016/j.tranon.2018.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Huang R., Xia M., Sakamuru S., Zhao J., Lynch C., Zhao T.. et al. Expanding biological space coverage enhances the prediction of drug adverse effects in human using in vitro activity profiles. Sci. Rep. 2018;8(1):3783. doi: 10.1038/s41598-018-22046-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Richard A. M., Tao D., LeClair C. A., Leister W., Tretyakov K. V., White E.. et al. Analytical Quality Evaluation of the Tox21 Compound Library. Chem. Res. Toxicol. 2025;38:15. doi: 10.1021/acs.chemrestox.4c00330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Wei Z., Zhao J., Niebler J., Hao J.-J., Merrick B. A., Xia M.. Quantitative proteomic profiling of mitochondrial toxicants in a human cardiomyocyte cell line. Front Genet. 2020;11:719. doi: 10.3389/fgene.2020.00719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Sakamuru S., Attene-Ramos M. S., Xia M.. Mitochondrial membrane potential assay. High-throughput screening assays in toxicology. 2016;1473:17–22. doi: 10.1007/978-1-4939-6346-1_2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Attene-Ramos M. S., Miller N., Huang R., Michael S., Itkin M., Kavlock R. J.. et al. The Tox21 robotic platform for the assessment of environmental chemicals–from vision to reality. Drug Discov Today. 2013;18(15–16):716–723. doi: 10.1016/j.drudis.2013.05.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Huang R.. A Quantitative High-Throughput Screening Data Analysis Pipeline for Activity Profiling. Methods Mol. Biol. 2022;2474:133–145. doi: 10.1007/978-1-0716-2213-1_13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Xu T., Xu M., Zhu W., Chen C. Z., Zhang Q., Zheng W.. et al. Efficient Identification of Anti-SARS-CoV-2 Compounds Using Chemical Structure-and Biological Activity-Based Modeling. J. Med. Chem. 2022;65:4590. doi: 10.1021/acs.jmedchem.1c01372. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Xu T., Ngan D. K., Ye L., Xia M., Xie H. Q., Zhao B.. et al. Predictive Models for Human Organ Toxicity Based on In Vitro Bioactivity Data and Chemical Structure. Chem. Res. Toxicol. 2020;33(3):731–741. doi: 10.1021/acs.chemrestox.9b00305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Xu T., Wu L., Xia M., Simeonov A., Huang R.. Systematic identification of molecular targets and pathways related to human organ level toxicity. Chem. Res. Toxicol. 2021;34(2):412–421. doi: 10.1021/acs.chemrestox.0c00305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Xu T., Li S., Li A. J., Zhao J., Sakamuru S., Huang W.. et al. Identification of Potent and Selective Acetylcholinesterase/Butyrylcholinesterase Inhibitors by Virtual Screening. J. Chem. Inf Model. 2023;63(8):2321–2330. doi: 10.1021/acs.jcim.3c00230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Xu T., Kabir M., Sakamuru S., Shah P., Padilha E. C., Ngan D. K.. et al. Predictive models for human cytochrome P450 3A7 selective inhibitors and substrates. J. Chem. Inf Model. 2023;63(3):846–855. doi: 10.1021/acs.jcim.2c01516. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Wei Z., Xu T., Strickland J., Zhang L., Fang Y., Tao D.. et al. Use of in vitro methods combined with in silico analysis to identify potential skin sensitizers in the Tox21 10K compound library. Front Toxicol. 2024;6:1321857. doi: 10.3389/ftox.2024.1321857. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Inglese J., Auld D. S., Jadhav A., Johnson R. L., Simeonov A., Yasgar A.. et al. Quantitative high-throughput screening: a titration-based approach that efficiently identifies biological activities in large chemical libraries. Proc. Natl. Acad. Sci. U. S. A. 2006;103(31):11473–11478. doi: 10.1073/pnas.0604348103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Wang Y., Jadhav A., Southal N., Huang R., Nguyen D. T.. A grid algorithm for high throughput fitting of dose-response curve data. Curr. Chem. Genomics. 2010;4:57–66. doi: 10.2174/1875397301004010057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Yang D., Zhou Q., Labroska V., Qin S., Darbalaei S., Wu Y.. et al. G protein-coupled receptors: structure-and function-based drug discovery. Signal Transduct Target Ther. 2021;6(1):7. doi: 10.1038/s41392-020-00435-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Wang L., Nie Q., Gao M., Yang L., Xiang J.-W., Xiao Y.. et al. The transcription factor CREB acts as an important regulator mediating oxidative stress-induced apoptosis by suppressing αB-Crystallin expression. Aging (Albany NY) 2020;12(13):13594. doi: 10.18632/aging.103474. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Rampe D., Brown A. M.. A history of the role of the hERG channel in cardiac risk assessment. J. Pharmacol Toxicol Methods. 2013;68(1):13–22. doi: 10.1016/j.vascn.2013.03.005. [DOI] [PubMed] [Google Scholar]
  31. Zhao M., Ma J., Li M., Zhang Y., Jiang B., Zhao X.. et al. Cytochrome P450 enzymes and drug metabolism in humans. Int. J. Mol. Sci. 2021;22(23):12808. doi: 10.3390/ijms222312808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Brown D. R., Clark B. W., Garner L. V., Di Giulio R. T.. Zebrafish cardiotoxicity: the effects of CYP1A inhibition and AHR2 knockdown following exposure to weak aryl hydrocarbon receptor agonists. Environ. Sci. Pollut Res. 2015;22:8329–8338. doi: 10.1007/s11356-014-3969-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Kim Y. S., Ko B., Kim D. J., Tak J., Han C. Y., Cho J.-Y.. et al. Induction of the hepatic aryl hydrocarbon receptor by alcohol dysregulates autophagy and phospholipid metabolism via PPP2R2D. Nat. Commun. 2022;13(1):6080. doi: 10.1038/s41467-022-33749-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Choi W., Namkung J., Hwang I., Kim H., Lim A., Park H. J.. et al. Serotonin signals through a gut-liver axis to regulate hepatic steatosis. Nat. Commun. 2018;9(1):4824. doi: 10.1038/s41467-018-07287-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Wu F.-Q., Fang T., Yu L.-X., Lv G.-S., Lv H.-W., Liang D.. et al. ADRB2 signaling promotes HCC progression and sorafenib resistance by inhibiting autophagic degradation of HIF1α. J. Hepatol. 2016;65(2):314–324. doi: 10.1016/j.jhep.2016.04.019. [DOI] [PubMed] [Google Scholar]
  36. Rachakonda V., Jadeja R. N., Urrunaga N. H., Shah N., Ahmad D., Cheng K.. et al. M1 muscarinic receptor deficiency attenuates azoxymethane-induced chronic liver injury in mice. Sci. Rep. 2015;5(1):14110. doi: 10.1038/srep14110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Xia Z., Huang X., Chen K., Wang H., Xiao J., He K.. et al. Proapoptotic role of potassium ions in liver cells. Biomed Res. Int. 2016;2016(1):1729135. doi: 10.1155/2016/1729135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Guzman S., Dragan M., Kwon H., de Oliveira V., Rao S., Bhatt V.. Targeting hepatic kisspeptin receptor ameliorates nonalcoholic fatty liver disease in a mouse model. J. Clin Invest. 2022;132(10):e145889. doi: 10.1172/JCI145889. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Klyushova L. S., Perepechaeva M. L., Grishanova A. Y.. The role of CYP3A in health and disease. Biomedicines. 2022;10(11):2686. doi: 10.3390/biomedicines10112686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Sarpeshkar V., Bentley D. J.. Adrenergic-β2 receptor polymorphism and athletic performance. J. Hum Genet. 2010;55(8):479–485. doi: 10.1038/jhg.2010.42. [DOI] [PubMed] [Google Scholar]
  41. Matus M., Lewin G., Stümpel F., Buchwalow I. B., Schneider M. D., Schütz G.. et al. Cardiomyocyte-specific inactivation of transcription factor CREB in mice. FASEB J. 2007;21(8):1884–1892. doi: 10.1096/fj.06-7915com. [DOI] [PubMed] [Google Scholar]
  42. Dehelean L., Marinescu I., Stovicek P. O., Romoşan A. M., Romoşan RŞ, Bálint R.. et al. Impairment of the cardiac ejection fraction by blocking dopamine D2 receptors induced by long-acting injectable antipsychotic treatment. Rom J. Morphol Embryol. 2022;62(2):497. doi: 10.47162/RJME.62.2.16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Wang J., Cheng J., Zhang C., Li X.. Cardioprotection effects of sevoflurane by regulating the pathway of neuroactive ligand-receptor interaction in patients undergoing coronary artery bypass graft surgery. Comput. Math Methods Med. 2017;2017:1. doi: 10.1155/2017/3618213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Fatunde O. A., Brown S.-A.. The role of CYP450 drug metabolism in precision cardio-oncology. Int. J. Mol. Sci. 2020;21(2):604. doi: 10.3390/ijms21020604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Mansouri K., Abdelaziz A., Rybacka A., Roncaglioni A., Tropsha A., Varnek A.. et al. CERAPP: collaborative estrogen receptor activity prediction project. Environ. Health Perspect. 2016;124(7):1023–1033. doi: 10.1289/ehp.1510267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Mansouri K., Kleinstreuer N., Abdelaziz A. M., Alberga D., Alves V. M., Andersson P. L.. et al. CoMPARA: collaborative modeling project for androgen receptor activity. Environ. Health Perspect. 2020;128(2):027002. doi: 10.1289/EHP5580. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Huang R., Xia M., Nguyen D.-T., Zhao T., Sakamuru S., Zhao J.. et al. Tox21Challenge to build predictive models of nuclear receptor and stress response pathways as mediated by exposure to environmental chemicals and drugs. Front Environ. Sci. 2016;3:85. doi: 10.3389/fenvs.2015.00085. [DOI] [Google Scholar]
  48. Signoretto E., Honisch S., Briglia M., Faggio C., Castagna M., Lang F.. Nocodazole induced suicidal death of human erythrocytes. Cell Physiol Biochem. 2016;38(1):379–392. doi: 10.1159/000438638. [DOI] [PubMed] [Google Scholar]
  49. Sakata N., Ishiga T., Taniguchi S., Ishiga Y.. Acibenzolar-S-methyl activates stomatal-based defense systemically in Japanese radish by inducing peroxidase-dependent reactive oxygen species production. Front Plant Sci. 2020;11:565745. doi: 10.3389/fpls.2020.565745. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Kwon C. H., Park J. Y., Kim T. H., Woo J. S., Kim Y. K.. Ciglitazone induces apoptosis via activation of p38 MAPK and AIF nuclear translocation mediated by reactive oxygen species and Ca2+ in opossum kidney cells. Toxicology. 2009;257(1–2):1–9. doi: 10.1016/j.tox.2008.11.019. [DOI] [PubMed] [Google Scholar]
  51. Roos N. J., Aliu D., Bouitbir J., Krähenbühl S.. Lapatinib activates the Kelch-like ECH-associated protein 1-nuclear factor erythroid 2-related factor 2 pathway in HepG2 cells. Front Pharmacol. 2020;11:544552. doi: 10.3389/fphar.2020.00944. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Li X., Zhou L., Ni Y., Wang A., Hu M., Lin Y.. et al. Nonylphenol induces pancreatic damage in rats through mitochondrial dysfunction and oxidative stress. Toxicol Res. 2017;6(3):353–360. doi: 10.1039/C6TX00450D. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Bell S. M., Chang X., Wambaugh J. F., Allen D. G., Bartels M., Brouwer K. L.. et al. In vitro to in vivo extrapolation for high throughput prioritization and decision making. Toxicol In Vitro. 2018;47:213–227. doi: 10.1016/j.tiv.2017.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Algharably E. A., Di Consiglio E., Testai E., Pistollato F., Mielke H., Gundert-Remy U.. In Vitro–In Vivo Extrapolation by Physiologically Based Kinetic Modeling: Experience With Three Case Studies and Lessons Learned. Front Toxicol. 2022;4:885843. doi: 10.3389/ftox.2022.885843. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

hp6c00052_si_001.docx (885KB, docx)
hp6c00052_si_002.xlsx (2MB, xlsx)

Data Availability Statement

All data needed to evaluate the conclusions in the paper are present in the paper and/or the . The codes used for modeling have been deposited in GitHub at https://github.com/TX-2017/machine-learning.


Articles from Environmental Health Perspectives are provided here courtesy of American Chemical Society

RESOURCES