Skip to main content
JACS Au logoLink to JACS Au
. 2026 Aug 31;6(9):5420–5432. doi: 10.1021/jacsau.6c01105

MacroTox: A Macroscopic Graph Topology-Based Multimodal Learning Framework for Robust Molecular Toxicity Prediction

He Huang †, Qinyi Wang †, Manzhan Zhang §, Pan Dou †, Xiaobo Yang †, Honglin Li †,*, Shiliang Li †,‡,*
PMCID: PMC13625743  PMID: 42819392

Abstract

Despite breakthroughs in predicting acute organ toxicities, advanced computational models continue to struggle with complex, insidious end points like drug-induced bone toxicity. A fundamental limitation of most supervised pipelines is their failure to explicitly model intermolecular similarities during task-specific training. Neglecting this macroscopic topology of the chemical space leads to fragmented latent representations, poor generalization, and high false negative rates. To address this, we propose MacroTox, a macroscopic topology-driven multimodal deep learning framework. Beyond intramolecular fusion, its core innovation is the dynamic construction of an intrabatch drug–drug similarity graph. Guided by a topology-aware synergistic optimization, this mechanism captures latent network correlations, maximizing the utilization of the chemical space to resolve representational fragmentation and enhance the predictive capacity under data scarcity. Evaluated on a rigorously curated bone toxicity data set, MacroTox achieved an area under the receiver operating characteristic curve of 0.93 and a Matthews correlation coefficient of 0.73. Crucially, it effectively balances the sensitivity-specificity trade-off (SEN: 0.88, SPE: 0.86), substantially mitigating the underreporting risks in bone toxicity screening. Notably, ablation studies confirm that these predictive enhancements stem from the synergistic effect of dynamic graph topology and edge loss regularization. Extensive benchmarking of MoleculeNet further confirms its robust transferability. Furthermore, via multilevel feature attribution, MacroTox elucidates chemically intuitive structure–activity relationships. By pinpointing specific toxicophores and resolving complex activity cliffs, the framework proves that it captures authentic toxicological mechanisms rather than memorizing superficial data set biases. Ultimately, MacroTox offers an reliable, generalizable, and interpretable virtual screening engine for early stage drug discovery.

Keywords: dynamic similarity macroscopic graph, computational toxicology, multimodal deep learning, topology-aware synergistic optimization, drug-induced bone toxicity


graphic file with name au6c01105_0010.webp


graphic file with name au6c01105_0008.webp

1. Introduction

In recent years, in silico toxicology has significantly transformed the safety assessment paradigm in early stage drug discovery. − Driven by increases in computational power and the rise of deep learning, particularly the introduction of Graph Neural Networks (GNNs) and multimodal architectures, recent computational frameworks have yielded remarkable improvements in anticipating acute and chronic toxicity end points, including cardiotoxicity and hepatotoxicity. −

Nevertheless, while computational methods have made progress in predicting acute organ toxicities, certain complex end points characterized by insidious onset and slow disease progression often receive less research attention. Moreover, because these late-onset and chronic toxicities generally lack high-throughput screening metrics compared to acute risks, they are frequently limited by a scarcity of experimental data. This data deficiency intrinsically restricts the model’s capacity to explore the broader chemical space, making it exceedingly difficult to extract critical molecular representations. Without these generalized deep latent features, conventional computational models struggle to transfer learned knowledge to novel chemical entities, leading to a notable decline in overall predictive performance. ,

Drug-induced bone toxicity represents a prime example of this data-scarcity dilemma. This condition refers to the skeletal damage caused by drugs that disrupt the balance of bone remodeling. − Such skeletal degradation typically progresses silently until severe adverse events, such as atypical femoral fractures or medication-related osteonecrosis of the jaw, manifest clinically. Although these adverse skeletal events have become critical factors leading to late-stage clinical trial failures and postmarket drug withdrawals, they systematically receive much less attention than organ toxicities associated with acute lethal risks. − Exacerbated by the global aging population and the prevalence of lifelong drug intervention programs, the cumulative risk of iatrogenic bone diseases has evolved into a critical yet underestimated public health challenge. Therefore, an urgent need exists to develop reliable early prediction tools during the drug discovery phase to assess bone safety. ,

Because of the limited data availability, computational frameworks based solely on traditional features often struggle to extract robust and generalizable features from the limited chemical space, resulting in insufficient model training and poor predictive performance. ,,, To navigate the predictive challenges, current strategies rely on molecule-based multimodal representation learning to maximize the informational yield of chemical structures. In typical computational pipelines, molecules are deconstructed into diverse, complementary representations: textual annotations capturing high-level semantic abstractions through linear sequences, , microscopic molecular graphs preserving spatial connectivity, , molecular fingerprints condensing global topological features, and explicit physicochemical descriptors characterizing fundamental molecular properties.

Recently, advanced multimodal methods have successfully integrated the physical fidelity of graphs, the semantic richness of texts, and the representational capacity of fingerprints through cross-modal interactions and prototype-guided alignment strategies, enhancing the recognition of complex pharmacophores. − Nevertheless, when applied to data-scarce end points like bone toxicity, most supervised pipelines exhibit a fundamental limitation: they often fail to explicitly model intermolecular similarities during task-specific training. Consequently, they focus almost exclusively on intramolecular feature alignment, neglecting the macroscopic topological relationships across the broader chemical space. This macroscopic topology, defined as the overarching network of intermolecular similarities and global structural relationships connecting individual compounds, is crucial for establishing global semantic anchors. Overlooking it exacerbates the fragmentation of latent representations and limits predictive capacity, rendering such models inadequate for providing reliable guidance in practical early stage safety screening. ,−

To address these limitations, we propose MacroTox (macroscopic topology-driven toxicity prediction framework, Scheme ), a framework based on macroscopic graph topology and multimodal representation learning. MacroTox comprehensively extracts multiview features by integrating four complementary chemical modalities: descriptors, fingerprints, sequences, and molecular graphs. Building upon this intramolecular fusion, the model dynamically constructs a macroscopic drug–drug similarity graph within the training batch. We introduce a topology-aware edge loss regularization term into the overall optimization objective. This mechanism guides network parameters based on the global geometry of the data by explicitly incorporating structural properties like edge connectivity and internode similarity. Ultimately, driving the model to align shared features among analogous drugs in the latent space empowers MacroTox to gather information across the chemical space manifold, thereby compensating for limited training data sets. We comprehensively evaluate this architecture on a rigorously curated drug-induced bone toxicity data set. The results indicate that MacroTox achieves competitive predictive performance and alleviates the constraints of data scarcity, balancing sensitivity (SEN) and specificity (SPE) to reduce false-negative rates (underreporting risk). Furthermore, extensive evaluations across multiple MoleculeNet benchmark data sets confirm its generalization capabilities. Beyond addressing bone toxicity prediction, this framework offers a transferable computational paradigm for other molecular property predictions that are constrained by data scarcity.

1. Overall Architecture of MacroTox. (a) Multimodal Feature Extraction and Prediction: Canonical Compound SMILES Strings are Parsed to Extract Four Complementary Modalities: Explicit Mordred Descriptors, Sparse MACCS Fingerprints, MolFormer-Encoded Contextual Sequence Embeddings, and Pseudo-3D Molecular Graphs Processed by a Graph Isomorphism Network (GIN). These Heterogeneous Features are Aligned and Concatenated Into a Unified Multidimensional Latent Representation (h i ), which is Subsequently Propagated Through a Dynamic Similarity Graph Module (DSGM) and a Linear Prediction Layer. (b) Dynamic Similarity Graph Module: Within Each Mini-batch, an Intra-batch Macroscopic Relational Graph is Dynamically Constructed. Pairwise Cosine Similarities (S ij ) are Computed, Followed by Strict Lower Triangular (Tril) Extraction to Eliminate Self-Loops. Thresholding is Then Applied to Form a Sparse Topological Adjacency Matrix. DropEdge is Applied Exclusively During the Training Phase to Enhance Robustness. Subsequently, a Graph Attention Network (GAT) Selectively Aggregates Analogous Features, with the Message-passing Process Iteratively to Deeply Refine the Latent Embeddings. (c) Joint Optimization: The Entire Architecture is Optimized End-To-End Guided by a Composite Objective. The Task-specific Loss (Ltask) Establishes the Decision Boundary for Accurate Toxicity Prediction, while the Topology-Aware Edge Loss (Ledge) Acts as a Structural Regularizer, Forcing Dissimilar Instances Apart in the Latent Space while Preserving Meaningful Intra-class Chemical Diversity.

1

2. Results

2.1. Rationale for the Drug–Drug Similarity Graph

Visual analysis of the raw chemical space highlights the necessity of the macroscopic drug–drug similarity graph (Figure ). The MACCS-based Tanimoto heatmap (Figure a) exhibits sparse but localized structural clustering, while the initial 3D t-SNE projections (Figure b) reveal highly intertwined toxic and nontoxic samples without discernible decision boundaries. These factors severely limit the conventional isolated instance learning.

1.

1

Chemical space analysis supporting the dynamic macroscopic similarity graph learning strategy. (a) Tanimoto similarity heatmap computed from MACCS fingerprints across the 1163 compounds in the curated bone toxicity data set. The heterogeneous distribution of similarities highlights the potential of utilizing intermolecular relationships to establish message-passing bridges across discrete molecular scaffolds. (b) 3D t-SNE projection of the initial high-dimensional feature space. Positive (toxic, orange) and negative (nontoxic, blue) samples exhibit extensive nonlinear overlap with no discernible decision boundary. This profound distribution overlap underscores the insufficiency of relying solely on isolated intramolecular features, explicitly motivating the integration of the topology-aware edge loss to impose structural constraints and enhance class separability in the latent space.

To alleviate this, MacroTox constructs a dynamic intrabatch similarity graph combined with topology-aware edge loss regularization. This relational inductive bias establishes message-passing bridges across analogous clusters, explicitly constraining the latent space to align structurally similar compounds while repelling dissimilar ones. 2D t-SNE visualizations confirm this mechanism’s efficacy: the initial overlapping state (Figure a) undergoes distinct topological reorganization, stratifying into strictly separable toxicity clusters post-training (Figure b). This demonstrates that the macroscopic graph effectively transforms a fragmented chemical manifold into highly discriminative representations with well-defined toxicological boundaries.

2.

2

t-SNE visualization of the high-dimensional latent space distributions before and after MacroTox optimization. (a) The pretraining feature space. (b) The post-training latent space optimized by the full MacroTox framework. Distinct colors denote the specific toxicity classes (blue: nontoxic; red: toxic), demonstrating the enhanced class separability and topological reorganization achieved by the model.

2.2. Comparison with Baseline Methods

To objectively evaluate the predictive efficacy of MacroTox, we compared our model against a diverse array of advanced deep learning and traditional machine learning baselines, including BTP-MFFGNN, MorganFP-MLP, WeaveGNN, MolCLR [graph isomorphism network (GIN) and GCN variants], SVC, KNN, D-MPNN, GROVER, GraphMVP, and ChemBERTa. The detailed comparative results across all metrics are summarized in Table .

1. Predictive Performance of MacroTox and Baseline Methods on the Bone Toxicity Dataset .

methods AUC ACC SEN SPE PRE F1 BAC MCC
MacroTox 0.93 0.87 0.88 0.86 0.81 0.84 0.87 0.73
BTP-MFFGNN 0.92 0.85 0.83 0.87 0.82 0.84 0.84 0.69
MorganFP-MLP 0.73 0.73 0.66 0.80 0.76 0.71 0.73 0.47
WeaveGNN 0.74 0.74 0.77 0.71 0.71 0.74 0.74 0.48
SVC 0.82 0.73 0.71 0.75 0.77 0.74 0.73 0.46
KNN 0.81 0.75 0.79 0.72 0.69 0.74 0.75 0.51
MolCLR_gin 0.85 0.74 0.52 0.89 0.61 0.71 0.75 0.45
MolCLR_gcn 0.78 0.69 0.41 0.89 0.51 0.65 0.71 0.35
GROVER 0.87 0.82 0.78 0.84 0.80 0.79 0.81 0.63
D-MPNN 0.88 0.77 0.76 0.80 0.81 0.79 0.78 0.55
GraphMVP 0.85 0.78 0.85 0.77 0.43 0.57 0.81 0.49
ChemBERTa 0.81 0.69 0.50 0.86 0.78 0.61 0.68 0.40
a

The best results in each column are colored in bold.

As demonstrated in Table and Figure a, the proposed MacroTox framework consistently achieves highly competitive performance across all primary evaluation metrics compared to the baseline models. Specifically, MacroTox achieved an area under the receiver operating characteristic curve (AUC) value of 0.93 and an accuracy (ACC) value of 0.87. This represents relative improvements of 1.09% (AUC), 2.4% (ACC), and 5.8% [Matthews correlation coefficient (MCC)] over the leading baseline, BTP-MFFGNN, with margins expanding up to 27.4%, 26.1%, and 108.6% against those of other models, respectively.

3.

3

Comprehensive performance comparison and stability analysis of MacroTox on the bone toxicity data set. (a) Radar chart illustrating the multidimensional performance of MacroTox against four representative baseline models (BTP-MFFGNN, D-MPNN, GROVER, and ChemBERTa) across eight evaluation metrics. The extensive coverage area of MacroTox indicates its robust and balanced performance across all evaluated metrics. (b) Scatter plot evaluating the trade-off between sensitivity (SEN) and specificity (SPE) for MacroTox and its ablation variants. The diagonal dashed line represents the optimal balance (SEN = SPE), and the light gray shaded region denotes the highly balanced zone (Δ ≤ 0.02), which was determined empirically based on the data set’s specific performance distribution to mathematically delineate the most robust balance. Horizontal and vertical error bars indicate the standard deviations of SPE and SEN, respectively. The concentrated distribution and minimal error bars of the full MacroTox framework demonstrate its superior capability to balance SEN and SPE while maintaining high predictive stability.

In computational toxicology, minimizing false negatives (i.e., failing to identify a toxic compound) is practically more vital than avoiding false positives. MacroTox successfully mitigates the majority-class bias inherent in existing models. While competitors like BTP-MFFGNN and MolCLR favor SPE (0.87 and 0.89) at the severe expense of SEN (it drops to 0.83 and 0.41–0.52), MacroTox maintains a highly robust balance (SEN: 0.88, SPE: 0.86). Coupled with leading F1-score (F1) and balanced accuracy (BAC) values, this balanced metric profile confirms the framework’s stability against distributional bias, effectively minimizing the dangerous underreporting of potentially toxic molecules.

2.3. Ablation Experiments

2.3.1. Contributions of Different Modules

To evaluate the independent and synergistic contributions of MacroTox modules, we evaluated three variants via five-fold cross-validation.

  • Without Edge Loss: Retains the dynamic macroscopic similarity graph learning strategy but removes the topology-aware edge loss.

  • w/o SimGraph: Replaces the similarity graph and topological message-passing with fully connected layers while retaining the edge loss.

  • Base model: Strips both modules, relying entirely on multimodal features and the task-specific loss.

As detailed in Table , AUC and ACC remain highly consistent across all variants (fluctuations ≤0.01). This baseline stability confirms that the multimodal composite representations intrinsically possess substantial discriminative power. Despite stable overall ACC, the variants diverge in SEN, SPE, and MCC. The Base Model exhibits majority class bias (SPE 0.88, SEN 0.79). Employing only the similarity graph (w/o Edge Loss) narrows the gap between SEN and SPE to 0.04. Conversely, utilizing only the edge loss (w/o SimGraph) pulls sparse positive samples closer in the latent space, raising SEN to 0.82. Crucially, integrating both modules (MacroTox) compresses the gap between SEN and SPE to 0.01 (0.86 vs 0.85, Figure b). Furthermore, this synergistic integration acts as a powerful structural regularizer, significantly dropping cross-validation variance to approximately ±0.02. Consequently, this mechanism maximizes the utilization of the chemical space, directly addressing the limitations of existing models to yield improvements in generalization capacity and stable predictive robustness.

2. Comparative Analysis of MacroTox and its Variants on the Bone Toxicity Dataset.
variants AUC ACC SEN SPE PRE F1 BAC MCC
MacroTox 0.92 ± 0.02 0.85 ± 0.01 0.86 ± 0.02 0.85 ± 0.02 0.79 ± 0.02 0.82 ± 0.02 0.85 ± 0.01 0.70 ± 0.03
w/o edge loss 0.91 ± 0.01 0.84 ± 0.02 0.87 ± 0.06 0.83 ± 0.05 0.77 ± 0.04 0.81 ± 0.02 0.85 ± 0.02 0.68 ± 0.03
w/o SimGraph 0.92 ± 0.01 0.85 ± 0.02 0.82 ± 0.04 0.88 ± 0.04 0.82 ± 0.04 0.82 ± 0.02 0.85 ± 0.02 0.70 ± 0.03
base model 0.91 ± 0.01 0.84 ± 0.02 0.79 ± 0.05 0.88 ± 0.02 0.81 ± 0.02 0.80 ± 0.03 0.83 ± 0.02 0.67 ± 0.04

2.3.2. Representational Advantages of Multimodal Fusion

To visually evaluate the multimodal fusion strategy, we applied t-SNE and kernel density estimation (KDE) to the latent features and prediction scores (Figure ).

4.

4

Evolution of latent feature representations and prediction score distributions in MacroTox across progressive modality integrations. The left panels display the t-SNE projections of the learned latent space, while the right panels illustrate the KDE of the prediction scores. (a) Single modality (SMILES only); (b) dual modalities (SMILES + Graph); (c) triple modalities (MD + SMILES + Graph); (d) integration of all four modalities. Red and blue colors represent positive (toxic) and negative (nontoxic) samples, respectively. The progression from (a) to (d) visually demonstrates a transition from highly entangled feature spaces to distinctly separated class boundaries, thereby validating the effectiveness of the multimodal fusion strategy in markedly enhancing the model’s discriminative power.

Relying on a single modality (e.g., SMILES, Figure a) results in severe topological entanglement and regional overlap, lacking sufficient discriminative power. Progressively integrating multiple features (Figure b,c) yields rudimentary clustering, yet hard examples persistently bridge the decision boundary. Specifically, dual modalities like descriptors and fingerprints (Figure S1) display a diffuse KDE distribution with a pronounced heavy tail, indicating insufficient global constraints. Conversely, fusing all multidimensional features with macroscopic graph representations (Figure d) naturally stratifies instances into two clearly bounded, highly separable clusters. The corresponding KDE distributions exhibit markedly enhanced kurtosis and reduced variance, substantially narrowing the dispersion of prediction scores. This visual evidence confirms that comprehensive multimodal fusion successfully bridges local structural cohesiveness and global prediction performance, amplifying the discriminative power across the latent manifold.

2.4. Evaluation of Generalization Performance on MoleculeNet Data Sets

To validate broad applicability, we evaluated MacroTox on seven MoleculeNet benchmarks against pure graph-driven and multimodal baselines (Table ), comprising Uni-Mol, S-CGIB, MOLGT, MoAMa, Tri-SGD, MMFRL, MDFCL, and ProtoMol. Results reveal that multimodal networks consistently outperform unimodal graph methods, as integrating complementary modalities alleviates the representational bottlenecks of relying solely on molecular topology. Building on this advantage, MacroTox achieved highly competitive performance across BBBP, SIDER, ESOL, and FreeSolv without task-specific hyperparameter tuning.

3. Prediction Results of MacroTox and Eight Baseline Models on the MoleculeNet Datasets .

  classification (ROC–AUC% ↑)
regression (RMSE ↓)
method BBBP BACE Tox21 SIDER ESOL freesolv lipophilicity
Uni-Mol 85.7 ± 0.2 72.9 ± 0.6 79.6 ± 0.5 65.9 ± 1.3 0.788 ± 0.029 1.620 ± 0.035 0.603 ± 0.010
S-CGIB 86.5 ± 0.8 88.8 ± 0.5 80.9 ± 0.2 64.0 ± 1.0 0.816 ± 0.019 1.648 ± 0.074 0.762 ± 0.042
MOLGT 73.7 ± 1.0 84.5 ± 0.1 75.8 ± 0.2 65.4 ± 0.4 0.839 ± 0.006 - 0.788 ± 0.013
MoAMa 81.3 ± 1.1 85.9 ± 0.6 78.3 ± 0.6 62.7 ± 0.4 1.125 ± 0.029 2.072 ± 0.053 1.085 ± 0.024
Tri_SGD 90.3 ± 1.0 89.4 ± 0.5 80.7 ± 0.6 67.8 ± 0.8 0.654 ± 0.106 1.607 ± 0.114 0.606 ± 0.023
MMFRL 91.6 ± 5.0 94.3 ± 2.4 85.2 ± 0.2 66.4 ± 1.9 1.037 ± 0.170 2.093 ± 0.090 0.607 ± 0.034
MDFCL 86.4 ± 1.0 78.4 ± 1.0 80.5 ± 0.7 67.7 ± 0.6 0.663 ± 0.026 1.607 ± 0.032 0.608 ± 0.037
ProtoMol 91.4 ± 0.3 90.3 ± 0.6 81.2 ± 0.3 68.1 ± 0.5 0.629 ± 0.014 1.522 ± 0.044 0.583 ± 0.039
MacroTox 92.5 ± 0.8 89.6 ± 0.7 79.0 ± 0.9 73.3 ± 0.9 0.538 ± 0.013 1.368 ± 0.424 0.748 ± 0.017
a

The best results in each column are highlighted in bold.

b

“-” indicates that the result was not reported.

This exceptional generalization is fundamentally attributed to the macroscopic similarity graph, which facilitates deep batchwise feature alignment and markedly amplifies class separability. Performance variations highlight distinct mechanistic trade-offs regarding the task types and data scales. On extremely small regression data sets (e.g., FreeSolv and ESOL), limited batch diversity restricts macroscopic graph construction. Furthermore, the topology-aware edge loss acts as a spatial smoothing penalty in continuous domains, potentially hindering the prediction of sharp-activity cliffs. Despite these adverse conditions, MacroTox still achieves robust predictive precision [root-mean-square errors (RMSEs) of 0.538 on ESOL and 1.368 on FreeSolv]. Conversely, massive data sets like Tox21 inherently favor large-scale self-supervised pretraining paradigms (e.g., MMFRL). Ultimately, these observations confirm that MacroTox is a highly reliable, transferable framework uniquely optimized for data-scarce molecular property predictions.

2.5. Model Interpretability

2.5.1. Topological Evolution of the Macroscopic Similarity Graph

We visualized the layer-wise structural evolution of the test set to elucidate the dynamic macroscopic similarity graph’s recognition capability (Figure a–c). In the initial feature space (Layer 0), toxic and nontoxic samples exhibit severe spatial entanglement with uniformly low relational weights. Following the first message-passing layer (Layer 1), emergent clustering patterns appear as intraclass attention weights intensify, though marginal interclass entanglement remains. By the second layer (Layer 2), driven by the topology-aware edge loss, the embeddings achieve pronounced spatial segregation. The attention mechanism dynamically allocates dominant weights to analogous compounds, significantly compressing latent intraclass distances. This evolution empirically validates that the dynamic macroscopic similarity graph learning strategy effectively maximizes interclass variance while minimizing intraclass variance.

5.

5

Layer-wise topological evolution of the macroscopic similarity graph in MacroTox. (a–c) Progressive structural reorganization of the latent feature space across two successive Dynamic Similarity Graph Module (DSGM) modules (corresponding to Layer 0, Layer 1, and Layer 2, respectively). The left column depicts the spatial distribution of the nodes, where individual drug molecules are color-coded according to their predicted toxicity score. While the right column illustrates the dynamically constructed network connectivity. Edges denote the dynamically learned cosine similarities between compounds, where darker (deep blue) lines indicate higher similarity weights. The progression demonstrates the framework’s efficacy in enhancing interclass separability and fostering cohesive intraclass clustering.

2.5.2. Mechanistic Interpretability via Gradient-Based SHAP and Molecular Heatmap Visualizations

To dissect the predictive rationale, we conducted multilevel feature attribution. Directional SHAP analysis of MACCS keys (Figure a) delineates toxicophores from protective moieties. Positive SHAP values highlight structural alerts, notably seven-membered rings (+1.40), primary amines (+0.99), and strained cyclobutanes (+0.98), aligning with the propensity of terminal amines to act as reactive metabolic centers. Conversely, terminal bromides (−1.79), imine linkers (−1.21), sulfonyl groups (−0.80), and rigid cyclopropanes (−0.94) strongly correlate with nontoxic classifications, offering viable derisking strategies.

6.

6

Mechanistic interpretability of MacroTox. (a) Directional SHAP analysis of the top-10 MACCS substructures driving the model’s predictions. Blue bars (negative SHAP values) represent protective moieties (e.g., R–Br, CN) that mitigate predicted risk, while red bars (positive SHAP values) indicate structural alerts or toxicophores (e.g., 7-membered rings, –NH2) that increase the predicted bone toxicity probability. (b) SHAP attribution of the top-10 Mordred physicochemical descriptors, illustrating the directional impact of 2D spatial autocorrelations and electronic distributions on toxicity prediction. (c) Gradient-based atom-level molecular heatmaps for representative nontoxic (Label: 0, left) and toxic (Label: 1, right) compounds. Red activation halos highlight the critical atomic substructures driving the specific predictions. The local spatial activations strongly corroborate the global SHAP analysis, with the model accurately focusing on protective halogens for nontoxic instances and reactive primary amines for toxic instances. (d) Comparative activation profiles demonstrating the fine-grained discrimination of structurally similar compound pairs (activity cliffs) with conflicting toxicity labels. The framework successfully isolates the precise atomic substitutions (highlighted by dashed circles) that trigger opposite toxicological outcomes.

Complementing this, Mordred descriptor SHAP attribution (Figure b) emphasizes 2D spatial autocorrelations. Descriptors like BCUTc-1l (+0.16) and GATS3pe (+0.12) positively correlate with toxicity, with ATSC3i (+0.10) specifically associating toxic potential with electronegativity gradients spanning three bonds. Conversely, GATS4c (−0.20) and AATSC1are (−0.12) function as physicochemical safety indicators.

To bridge global attribution and local interpretation, we generated atom-level molecular heatmaps using gradient-based activation weights from the GIN module (Figure c). Corroborating the SHAP findings, activation halos for nontoxic compounds highlight protective moieties like terminal bromines, whereas attention for toxic drugs sharply focuses on reactive centers such as terminal primary amines. Crucially, MacroTox accurately distinguishes structural analogues with conflicting toxicity labels (Figure d) by isolating the precise atomic substitutions driving the toxicological shift. Although the predictive margin slightly narrows for complex scaffolds due to global pooling dilution, the overall alignment between global SHAP, microscopic heatmaps, and analog resolution robustly proves that the framework captures chemically intuitive structure–activity relationships (SAR) rather than memorizing superficial data set biases.

3. Discussion and Conclusion

In this study, we introduce MacroTox, a multimodal representation learning framework driven by a macroscopic graph topology. This framework deeply integrates four heterogeneous chemical modalities, specifically physicochemical descriptors, molecular fingerprints, contextual sequence representations, and microlevel molecular graphs, to maximize the extraction of multiview features from compounds. Building upon this multidimensional feature fusion, the model dynamically constructs a macroscopic drug–drug similarity graph within each training batch. By incorporating a topology-aware edge loss regularization term into the overall optimization objective, MacroTox explicitly models the latent network associations among drugs, enabling efficient aggregation and deep alignment of shared features among analogous drugs in the latent space.

Evaluated on a rigorously curated drug-induced bone toxicity data set, MacroTox demonstrates superior and robust predictive performance. Notably, the synergistic graph learning mechanism successfully balances the trade-off between SEN and SPE, achieving reliable identification of high-risk molecules while minimizing the false-negative risks that typically plague data-scarce scenarios. Furthermore, extensive benchmarking across diverse MoleculeNet data sets corroborates the framework’s robust generalization capabilities and broad applicability in complex molecular property prediction.

This high generalization capacity is fundamentally explained by the model’s intrinsic interpretability. Feature attribution via SHAP analysis and atom-level molecular heatmaps empirically validates that the framework avoids superficial pattern memorization, confirming that it captures chemically intuitive structure–activity relationships, isolates critical toxicophores such as reactive primary amines, and achieves fine-grained discrimination of structurally similar analogs to offer actionable mechanistic insights for lead optimization.

To further validate the structural robustness of our framework and address the hypersensitivity to topological noise, a critical limitation often observed in existing graph-based models, we conducted comprehensive ablation studies focusing on similarity thresholds (Figure S2) and dynamic edge-dropping (DropEdge) strategies (Figure S3). By employing a fixed-number edge-dropping mechanism, the model achieves its maximum intrinsic discriminative capacity (AUC) and maintains optimal topological integrity at a specific sparse threshold. The consistently stable performance observed across broad threshold ranges and alternative dropping mechanisms explicitly demonstrates that the predictive efficacy stems from authentic macroscopic topology learning, rather than hyperparameter overfitting.

Despite these promising outcomes, we acknowledge certain methodological limitations that warrant further investigation. First, the macroscopic similarity graph is dynamically constructed at the minibatch level, making its scale, sample diversity, and topological receptive field inherently dependent on batch size and sampling variance. While an optimal batch size effectively balances global topological awareness and SEN, smaller batch sizes constrain the chemical space, whereas overly large batches risk structural noise, oversmoothing, and minor trade-offs in SPE (Section 1.3, Table S2, and Figure S5 of the Supporting Information). Future iterations will explore cross-batch graph memory banks to approximate the global data manifold more comprehensively without compromising the class-specific precision. Second, the topology-aware edge loss, originally optimized for discrete classification margins, can act as an unintended spatial smoothing penalty in continuous domains, occasionally restricting prediction precision near sharp toxicological activity cliffs on exceptionally small regression data sets. Third, although the framework excels in data-scarce environments by relying exclusively on chemical modalities, it currently lacks the integration of complex in vivo pharmacokinetic parameters (e.g., systemic ADME profiles), which remain pivotal to the ultimate clinical manifestation of toxicity.

In summary, MacroTox establishes a dependable computational paradigm for predicting complex toxicological end points constrained by a severe lack of experimental data. As a robust virtual screening tool, it holds translational potential for accelerating the safety assessment of early stage drug candidates. Looking forward, the inherent scalability of this macroscopic graph architecture provides a versatile foundation. It is well positioned to be extended to tackle other insidious and complex organ toxicities that suffer from similar data scarcity challenges.

4. Materials and Methods

4.1. Data Collection

4.1.1. Bone Toxicity Data Set

This study utilizes a high-quality bone toxicity data set curated by Chen et al., comprising 563 toxic compounds (sourced from FAERS, SIDER, and CTD) and 600 nontoxic controls (randomly sampled from approved DrugBank small molecules lacking documented skeletal adverse events). As per the original data set curation, all molecular structures were standardized using RDKit. For our experiments, the 1163 compounds were randomly partitioned into training, validation, and test sets at an 8:1:1 ratio utilizing a global random seed of 0 to ensure computational reproducibility.

4.1.2. Benchmark Evaluations on MoleculeNet

To validate the generalizability and robustness of the proposed model architecture, we selected seven benchmark data sets from MoleculeNet, encompassing diverse tasks ranging from molecular properties to biological activities. To maintain strict methodological consistency with the primary bone toxicity prediction task, all benchmark data sets were partitioned using the random splitting strategy. Crucially, to ensure a rigorous evaluation and mitigate the stochastic variance inherent in data partitioning, model performance on these benchmarks was assessed via five-fold cross-validation, with the initial random seed fixed at 0 to guarantee strict computational reproducibility.

  • Classification Tasks: To assess the model’s direct transferability, we evaluated it on the BBBP, BACE, Tox21, and SIDER data sets. During this process, we retained the exact hyperparameter configuration of the bone toxicity prediction model without any task-specific parameter tuning. This rigorous setup was designed to evaluate the baseline performance of the model and the universality of its multimodal feature extraction mechanism without additional adaptation. The primary evaluation metric for the classification tasks was the AUC.

  • Regression Tasks: We further evaluated the model on the ESOL, FreeSolv, and Lipophilicity data sets. Given the fundamental shift from discrete labels to continuous numerical distributions and the heightened susceptibility of optimization trajectories to numerical scales, minimal fine-tuning was necessary. Specifically, modifications were restricted exclusively to the learning rate, the similarity threshold, and the loss function weight parameter to accommodate the continuous numerical domain (Table ). The core model architecture and all other training parameters remain strictly consistent with the bone toxicity model (Table S1). The predictive performance of these regression tasks is evaluated using the RMSE.

4. Key Hyperparameter of MacroTox Across Different Prediction Tasks.
  classification
regression
hyperparameters all tasks ESOL freesolv lipophilicity
learning rate (lr) 0.0005 0.0001 0.0005 0.005
similarity threshold (τ) 0.60 0.95 0.95 0.75
edge loss weight (λ) 10 0.02 0.02 0.02

4.2. Model Framework

Scheme illustrates the overall architecture of the MacroTox framework designed for drug-induced bone toxicity prediction. The framework comprises three core modules: molecular representation generation, multimodal feature fusion, and similarity graph learning. Specifically, MacroTox first parses canonical SMILES strings to extract four complementary chemical modalities: Mordred-derived explicit physicochemical descriptors, global MACCS structural fingerprints, MolFormer-encoded contextual sequence representations, and pseudo-3D molecular graphs. ,−

For the microscopic molecular graphs, a GIN is employed to capture fine-grained local interactions between atoms and bonds, followed by a global pooling layer to generate graph-level embeddings. Instead of direct concatenation, these four heterogeneous representations are first projected into a unified latent space via domain-specific linear transformations to align their dimensions and scales. Subsequently, they were fused to construct a comprehensive multimodal molecular profile.

To address the inability of conventional pipelines to explicitly model intermolecular similarities, the model incorporates a dynamic similarity learning module. Within each training minibatch, the module computes the pairwise cosine similarity among the fused multimodal features. By applying a predefined similarity threshold and a masking operation, it dynamically constructs a macroscopic intrabatch drug–drug similarity graph. To mitigate overfitting and enhance the robustness of the message-passing process, a stochastic edge-dropping mechanism (DropEdge) is integrated during training. Utilizing a graph attention network (GAT), this module adaptively assigns attention weights to analogous drugs, facilitating the alignment and aggregation of shared features across the chemical space.

Finally, the topology-optimized composite features are fed into a fully connected neural network classifier. The entire architecture is optimized end-to-end under the joint guidance of a task-specific loss (Ltask) and a topology-aware edge loss (Ledge) . As established by the synergistic effect of these modules, this joint optimization mechanism effectively captures latent network correlations, enabling the robust modeling of complex, nonlinear structure-toxicity relationships through the synergy of microlevel molecular features and macrolevel topological graphs.

4.3. Feature Representation

4.3.1. Multimodal Chemical Features

To extract multidimensional features from canonical SMILES strings, we employed three computational strategies. First, 167-dimensional MACCS fingerprints were computed via RDKit (v2025.03.6) to encode the core substructural features. Second, 1613 physicochemical descriptors were calculated using Mordred (v1.2.0). To prevent data leakage, an eXtreme Gradient Boosting (XGBoost) algorithm fitted exclusively on the training set selected the top 100 critical descriptors, which were consistently applied across all data sets. Full details regarding the feature selection rationale, descriptor rankings, and importance evaluations are provided in the Supporting Information (Section 1.2 and Figure S4). Third, contextual sequence representations were extracted using the pretrained MolFormer (MoLFormer-XL-both-10pct), projecting SMILES sequences into 768-dimensional vectors. All independent modalities were condensed via a fully connected layer before the fusion.

4.3.2. Molecular Graph Features

Canonical SMILES strings were converted to 2D molecules with explicit hydrogens by using RDKit. To integrate spatial constraints efficiently, pseudo-3D information was incorporated into the topological graphs. The initial 3D conformations were generated via the ETKDG algorithm and optimized using the MMFF94 force field. Each compound was then abstracted as an undirected graph where nodes represent atoms and edges represent bonds. Following Lee et al., node features included the atomic number, hybridization state, Gasteiger partial charges, valences, and chirality. Edge features comprised the bond type, ring inclusion status, and 3D bond length (calculated as the Euclidean distance between optimized 3D atomic coordinates). These attributes equip the downstream GIN with crucial relational biases.

4.4. Dynamic Macroscopic Similarity Graph Learning Strategy

To achieve deep fusion and feature enhancement of multimodal representations, we introduced a dynamic macroscopic similarity graph learning strategy (Scheme b). Unlike conventional multimodal architectures limited to shallow feature fusion, the proposed approach leveraged an end-to-end learning paradigm to explicitly capture the hidden topological dependencies among chemical instances within the data set. Specifically, during each forward pass, we dynamically constructed a macroscopic drug–drug similarity graph at the mini-batch level. For compounds within the same batch, the model first computed the pairwise cosine similarity S ij based on their concatenated multimodal composite feature vectors, h i and h j

Sij=hiΤ·hj∥hi∥∥hj∥ 1

Because h i and h j were dynamically updated during backpropagation, S ij serves as a differentiable dynamic edge weight that continuously evolves during training. To mitigate computational redundancy and prevent the self-reinforcement bias of node features, we extracted the strict lower triangular portion of the similarity matrix, which explicitly eliminates self-loops and duplicate undirected edges. Subsequently, a predefined similarity threshold τ was applied to prune loosely correlated node pairs, thereby constructing an initial sparse weighted adjacency matrix. The edge set ε is formally defined as follows

ε={(i,j)|Sij>τ,i<j} 2

To alleviate overfitting and improve structural robustness during training, a uniform DropEdge was employed. Given a predefined hyperparameter n pairs specifying the target number of edges to discard, the model uniformly and randomly sampled a subset εdrop from the initial edge set ε for removal. Crucially, to preserve the essential structural connectivity of the graph, the mechanism strictly ensures the retention of at least one edge. The cardinality of the removed subset, |εdrop|, is mathematically constrained as follows

|εdrop|=min(npairs,|ε|−1) 3

Consequently, the refined sparse edge set ε̃ employed for the subsequent macro-level message-passing phase was formally defined as

ε̃=ε\εdrop 4

During the inference phase, the DropEdge was strictly disabled. The model directly utilizes the complete sparse edge set ε, determined solely by the threshold τ, for feature aggregation. This decoupled mechanism between training and inference facilitates an adaptive, discriminative learning paradigm, substantially enhancing the model’s generalization capability on unseen data.

4.5. Macroscopic GNN-Based Multimodal Feature Fusion

Following the construction of the batch-wise similarity graph, we deployed a GAT to perform the final macro-level multimodal feature fusion. In this relational topology, each node represents a complete compound instance, and the explicit connections were dictated by the filtered edge set ε̃. In contrast to standard graph convolutions that employ isotropic aggregation, GAT incorporates a self-attention mechanism. During the message-passing phase, the attention coefficient was dynamically computed to assign varying importance weights to topologically connected neighbors. This attention-driven paradigm enabled the model to update the latent representation of compound i by selectively aggregating multidimensional attributes from its analogous neighbors within the batch. By utilizing the topological connectivity defined by the thresholding mechanism, the model forced deep feature interaction among drugs sharing structural or physicochemical similarities. Ultimately, the GAT outputs the refined, topology-aware node embeddings, which serve as the final global compound representations. These discriminative representations effectively bridged the gap between the high-dimensional multimodal feature space and the downstream task-specific prediction modules.

4.6. Loss Function

The framework is optimized end-to-end using a composite objective comprising a task-specific prediction loss (Ltask) and a topology-aware edge loss (Ledge) .

4.6.1. Task-Specific Prediction Loss

MacroTox maps updated node embeddings from the macroscopic similarity graph to the output space via a linear layer. We employ binary cross-entropy (BCE) for classification tasks and MSE for regression tasks

Ltask=−1N∑i=1N[yilog(ŷi)+(1−yi)log(1−ŷi)] 5
Ltask=1N∑i=1N(yi−ŷi)2 6

where N is the batch size, y i represents the ground-truth label or continuous value, and ŷi is the predicted value.

4.6.2. Topology-Aware Edge Loss

To enhance discriminative capacity, we introduce Ledge as a structural regularizer. This loss maximizes interclass variance while mitigating representation collapse, thus preserving meaningful intraclass structural diversity. It is formally defined as

Ledge=1|E|∑(i,j)∈ESij·∥yi−yj∥2 7

where S ij represents the dynamic edge weight connecting nodes i and j in the latent space and y i , y j denote their ground-truth labels or continuous values. Utilizing ground truth labels rather than predicted probabilities intentionally nullifies the penalty between identical-class samples. This maximizes the distance between dissimilar samples without overcompressing analogous intraclass instances, thereby preserving their distinct chemical identities. However, in regression tasks involving continuous values, this formulation may cause the model to overfit the training data, thereby compromising its generalization capabilities. To mitigate this risk, we reduced the weight coefficient of the edge loss in regression tasks.

4.6.3. Overall Training Objective

The final objective balances the specific prediction task and graph structural learning

Ltotal=Ltask+λLedge 8

where λ is a tunable hyperparameter controlling the edge loss contribution, and Ltask alternates between BCE and MSE depending on the objective. This unified strategy jointly enables discriminative latent representation learning and robust downstream prediction.

4.7. Evaluation Metrics and Experimental Setup

To comprehensively evaluate the classification performance of MacroTox, we employed eight standard evaluation metrics: ACC, AUC, BAC, F1, precision, MCC, SPE, and SEN. The detailed mathematical definitions and formula derivations for these evaluation metrics are provided in Section 1.1 of the Supporting Information. Among these metrics, AUC was established as the primary optimization criterion to evaluate the overall discriminative capacity across the chemical space, while SEN was prioritized to minimize the false-negative rate, in alignment with clinical and toxicological screening requirements.

The entire MacroTox framework was developed and executed within a Python (v3.12) environment using PyTorch (v2.8.0) and PyTorch Geometric (v2.7.0). All local training and inference procedures were conducted on a high-performance computing node equipped with an Intel­(R) Xeon­(R) Platinum 8369B CPU at 2.90 GHz, 512GB RAM, and an NVIDIA A100 80GB PCIe GPU, leveraging CUDA 12.6 for hardware acceleration to guarantee absolute computational reproducibility.

Supplementary Material

au6c01105_si_001.pdf (711.5KB, pdf)

Acknowledgments

This work was supported in part by the National Natural Science Foundation of China (82425104 to H.L. and 82622110 to S.L.), the National Key R&D Program of China (2022YFC3400504), and the Shanghai Science and Technology Program (26JS2830100 and 26JS2830102).

The original raw bone toxicity data utilized in this study were collected and curated by Chen et al. To ensure strict computational reproducibility and facilitate community deployment, the precise code for standardized data set splitting (executed under a fixed random seed), the final pretrained model weights, and the underlying source code for the proposed framework are fully open-sourced and freely available in the GitHub repository at https://github.com/hhuang-hub/MacroTox.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jacsau.6c01105.

  • Mathematical definitions of eight standard evaluation metrics, rationale and details for descriptor selection, and the impact of batch sizes on macroscopic graph construction; comprehensive visualization of latent feature representations and prediction score distributions, hyperparameter sensitivity analyses, feature importance ranking of the top 10 Mordred descriptors, and hyperparameter configurations (PDF)

H.H.: Conceptualization, methodology, software, validation, visualization, formal analysis, writingoriginal draft, and writingreview and editing; Q.W.: writingreview and editing; M.Z.: data curation and validation; D.P. and X.Y.: visualization; H.L.: resources and supervision; S.L.: conceptualization, resources, writingreview and editing, and supervision.

The authors declare no competing financial interest.

References

  1. Mullowney M. W.. et al. Artificial Intelligence for Natural Product Drug Discovery. Nat. Rev. Drug Discovery. 2023;22:895–916. doi: 10.1038/s41573-023-00774-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Pognan F., Beilmann M., Boonen H. C. M., Czich A., Dear G., Hewitt P., Mow T., Oinonen T., Roth A., Steger-Hartmann T., Valentin J.-P., Van Goethem F., Weaver R. J., Newham P.. The Evolving Role of Investigative Toxicology in the Pharmaceutical Industry. Nat. Rev. Drug Discovery. 2023;22:317–335. doi: 10.1038/s41573-022-00633-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bai C., Wu L., Li R., Cao Y., He S., Bo X.. Machine Learning-Enabled Drug-Induced Toxicity Prediction. Adv. Sci. 2025;12:2413405. doi: 10.1002/advs.202413405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Liu J.-W., Liu K.-Y., Deng Y.-C., Shi S.-H., Fu X.-Z., Jiang Y.-P., Fang J., Zhang Q., Jiang D.-J., Liu S., Cao D.-S.. Uncertainty-Aware Deep Learning and Structural Feature Analysis for Reliable Nephrotoxicity Prediction. J. Chem. Inf. Model. 2025;65(17):9082–9096. doi: 10.1021/acs.jcim.5c01532. [DOI] [PubMed] [Google Scholar]
  5. Zhu Y., Zhang Y., Li X., Wang L.. 3MTox: A Motif-Level Graph-Based Multi-View Chemical Language Model for Toxicity Identification with Deep Interpretation. J. Hazard. Mater. 2024;476:135114. doi: 10.1016/j.jhazmat.2024.135114. [DOI] [PubMed] [Google Scholar]
  6. Ha S., Bang D., Kim S.. Fate-Tox: Fragment Attention Transformer for E(3)-Equivariant Multi-Organ Toxicity Prediction. Chem. inf. 2025;17:74. doi: 10.1186/s13321-025-01012-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Lee T., Posma J. M.. Improving Drug-Induced Liver Injury Prediction Using Graph Neural Networks with Augmented Graph Features from Molecular Optimisation. Chem. inf. 2025;17:124. doi: 10.1186/s13321-025-01068-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Feng L., Fu X., Du Z., Guo Y., Zhuo L., Yang Y., Cao D., Yao X.. MultiCTox: Empowering Accurate Cardiotoxicity Prediction through Adaptive Multimodal Learning. J. Chem. Inf. Model. 2025;65:3517–3528. doi: 10.1021/acs.jcim.5c00022. [DOI] [PubMed] [Google Scholar]
  9. Liu K., Cui H., Yu X., Li W., Han W.. Predicting Cardiotoxicity in Drug Development: A Deep Learning Approach. J. Pharm. Anal. 2025;15:101263. doi: 10.1016/j.jpha.2025.101263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Tan S., Ding Y., Wang W., Rao J., Cheng F., Zhang Q., Xu T., Hu T., Hu Q., Ye Z.. et al. Development of an AI Model for DILI-level Prediction Using Liver Organoid Brightfield Images. Commun. Biol. 2025;8:886. doi: 10.1038/s42003-025-08205-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Kumar A., Ojha P. K., Roy K.. Chemometric Modeling of the Lowest Observed Effect Level (LOEL) and No Observed Effect Level (NOEL) for Rat Toxicity. Environ. Sci.:Adv. 2024;3:686–705. doi: 10.1039/D3VA00265A. [DOI] [Google Scholar]
  12. Zhang J., Li H., Zhang Y., Huang J., Ren L., Zhang C., Zou Q., Zhang Y.. Computational Toxicology in Drug Discovery: Applications of Artificial Intelligence in ADMET and Toxicity Prediction. Briefings Bioinf. 2025;26:bbaf533. doi: 10.1093/bib/bbaf533. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Olali A. Z., Carpenter K. A., Myers M., Sharma A., Yin M. T., Al-Harthi L., Ross R. D.. Bone Quality in Relation to HIV and Antiretroviral Drugs. Curr. HIV/AIDS Rep. 2022;19:312–327. doi: 10.1007/s11904-022-00613-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Hofbauer L. C., Compston J. E., Saag K. G., Rauner M., Tsourdi E.. Glucocorticoid-Induced Osteoporosis: Novel Concepts and Clinical Implications. Lancet Diabetes Endocrinol. 2025;13:964–979. doi: 10.1016/S2213-8587(25)00251-7. [DOI] [PubMed] [Google Scholar]
  15. Tuladhar A., Guffroy M., Finnema S. J., Christmann R., Van Vleet T. R., Mayana S. A. K., Fossey S.. Investigation of Bone Toxicity in Drug Development: Review of Current and Emerging Technologies. Toxicol. Sci. 2025;208:207–224. doi: 10.1093/toxsci/kfaf131. [DOI] [PubMed] [Google Scholar]
  16. Ribeiro A. C., Gerheim P. S. A. S., Chebli J. M. F., Nascimento J. W. L., de Faria Pinto P.. The Role of Pharmacogenetics in the Therapeutic Response to Thiopurines in the Treatment of Inflammatory Bowel Disease: A Systematic Review. J. Clin. Med. 2023;12:6742. doi: 10.3390/jcm12216742. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Khosla S., Hofbauer L. C.. Osteoporosis Treatment: Recent Developments and Ongoing Challenges. Lancet Diabetes Endocrinol. 2017;5:898–907. doi: 10.1016/S2213-8587(17)30188-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Fernández I., Gavaia P. J., Laizé V., Cancela M. L.. Fish as a Model to Assess Chemical Toxicity in Bone. Aquat. Toxicol. 2018;194:208–226. doi: 10.1016/j.aquatox.2017.11.015. [DOI] [PubMed] [Google Scholar]
  19. Cui H., He Y., Wang Z., Liu K., Li W., Han W.. Unveiling Drug-Induced Osteotoxicity: A Machine Learning Approach and Webserver. J. Hazard. Mater. 2025;492:138044. doi: 10.1016/j.jhazmat.2025.138044. [DOI] [PubMed] [Google Scholar]
  20. Pan Z., Yang X., Wang C., Han T., Zhao Q.. Multimodal Feature Fusion for Bone Toxicity Prediction and Local Platform. J. Chem. Inf. Model. 2025;65:12799–12810. doi: 10.1021/acs.jcim.5c02280. [DOI] [PubMed] [Google Scholar]
  21. Gagnon M.-E., Talbot D., Tremblay F., Desforges K., Sirois C.. Polypharmacy and Risk of Fractures in Older Adults: A Systematic Review. J. Evid.-Based Med. 2024;17:145–171. doi: 10.1111/jebm.12593. [DOI] [PubMed] [Google Scholar]
  22. Fang E. F.. et al. A Research Agenda for Ageing in China in the 21st Century (2nd Edition): Focusing on Basic and Translational Research, Long-Term Care, Policy and Social Networks. Ageing Res. Rev. 2020;64:101174. doi: 10.1016/j.arr.2020.101174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Li T., Qu Y., Chen A., Thakkar S., Li D., Tong W.. Beyond QSARs: Quantitative Knowledge–Activity Relationships (QKARs) for Enhanced Drug Toxicity Prediction. Toxicol. Sci. 2025;208:269–278. doi: 10.1093/toxsci/kfaf135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Chen Y.-Q., Yu T., Song Z.-Q., Wang C.-Y., Luo J.-T., Xiao Y., Qiu H., Wang Q.-Q., Jin H.-M.. Application of Large Language Models in Drug-Induced Osteotoxicity Prediction. J. Chem. Inf. Model. 2025;65:3370–3379. doi: 10.1021/acs.jcim.5c00275. [DOI] [PubMed] [Google Scholar]
  25. Weininger D.. SMILES, a Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules. J. Chem. Inf. Comput. Sci. 1988;28:31–36. doi: 10.1021/ci00057a005. [DOI] [Google Scholar]
  26. Krenn M., Häse F., Nigam A., Friederich P., Aspuru-Guzik A.. Self-Referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation. Mach. Learn.: Sci. Technol. 2020;1:045024. doi: 10.1088/2632-2153/aba947. [DOI] [Google Scholar]
  27. Duvenaud, D. ; Maclaurin, D. ; Aguilera-Iparraguirre, J. ; Gómez-Bombarelli, R. ; Hirzel, T. ; Aspuru-Guzik, A. ; Adams, R. P. In Proceedings of the 29th International Conference on Neural Information Processing Systems - Vol. 2; MIT Press: Cambridge, MA, USA, 2015; Vol. 2, pp 2224–2232. [Google Scholar]
  28. Boulougouri M., Vandergheynst P., Probst D.. Molecular Set Representation Learning. Nat. Mach. Intell. 2024;6:754–763. doi: 10.1038/s42256-024-00856-0. [DOI] [Google Scholar]
  29. Zagidullin B., Wang Z., Guan Y., Pitkänen E., Tang J.. Comparative Analysis of Molecular Fingerprints in Prediction of Drug Combination Effects. Briefings Bioinf. 2021;22:bbab291. doi: 10.1093/bib/bbab291. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. David L., Thakkar A., Mercado R., Engkvist O.. Molecular Representations in AI-driven Drug Discovery: A Review and Practical Guide. Chem. inf. 2020;12:56. doi: 10.1186/s13321-020-00460-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Gong X., Liu M., Liu Q., Guo Y., Wang G.. MDFCL: Multimodal Data Fusion-Based Graph Contrastive Learning Framework for Molecular Property Prediction. Pattern Recognit. 2025;163:111463. doi: 10.1016/j.patcog.2025.111463. [DOI] [Google Scholar]
  32. Zhou, G. ; Janarthanan, S. ; Lu, Y. ; Hu, P. . CL-MFAP: A Contrastive Learning-based Multimodal Foundation Model for Molecular Property Prediction and Antibiotic Screening, 2025. [Google Scholar]
  33. Rollins Z. A., Cheng A. C., Metwally E.. MolPROP: Molecular Property Prediction with Multimodal Language and Graph Fusion. Chem. inf. 2024;16:56. doi: 10.1186/s13321-024-00846-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Lu X., Xie L., Xu L., Mao R., Xu X., Chang S.. Multimodal Fused Deep Learning for Drug Property Prediction: Integrating Chemical Language and Molecular Graph. Comput. Struct. Biotechnol. J. 2024;23:1666–1679. doi: 10.1016/j.csbj.2024.04.030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Wang Y., Zhang K., Huang J., Yin N., Liu S., Segal E.. ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning. Briefings Bioinf. 2025;26:bbaf629. doi: 10.1093/bib/bbaf629. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Wang B., Li J., Zhou D., Li L., Li J., Wang E., Hao J., Shi L., Lu C., Qiu J., Hou T., Cao D., Chen G., Heng P. A.. Unified and Explainable Molecular Representation Learning for Imperfectly Annotated Data from the Hypergraph View. Nat. Commun. 2025;16:8717. doi: 10.1038/s41467-025-63730-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Keiser M. J., Roth B. L., Armbruster B. N., Ernsberger P., Irwin J. J., Shoichet B. K.. Relating Protein Pharmacology by Ligand Chemistry. Nat. Biotechnol. 2007;25:197–206. doi: 10.1038/nbt1284. [DOI] [PubMed] [Google Scholar]
  38. Lounkine E., Keiser M. J., Whitebread S., Mikhailov D., Hamon J., Jenkins J. L., Lavan P., Weber E., Doak A. K., Côté S., Shoichet B. K., Urban L.. Large-Scale Prediction and Testing of Drug Activity on Side-Effect Targets. Nature. 2012;486:361–367. doi: 10.1038/nature11159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Liu R., AbdulHameed M. D. M., Xu Z., Clancy B., Desai V., Wallqvist A.. Rapid Screening of Chemicals for Their Potential to Cause Specific Toxidromes. Front. Drug Discovery. 2024;4:1324564. doi: 10.3389/fddsv.2024.1324564. [DOI] [Google Scholar]
  40. Yang H., Xiu J., Yan W., Liu K., Cui H., Wang Z., He Q., Gao Y., Han W.. Large Language Models as Tools for Molecular Toxicity Prediction: AI Insights into Cardiotoxicity. J. Chem. Inf. Model. 2025;65:2268–2282. doi: 10.1021/acs.jcim.4c01371. [DOI] [PubMed] [Google Scholar]
  41. Wang Y., Wang J., Cao Z., Barati Farimani A.. Molecular Contrastive Learning of Representations via Graph Neural Networks. Nat. Mach. Intell. 2022;4:279–287. doi: 10.1038/s42256-022-00447-x. [DOI] [Google Scholar]
  42. Cortes C., Vapnik V.. Support-Vector Networks. Mach. Learn. 1995;20:273–297. doi: 10.1023/a:1022627411411. [DOI] [Google Scholar]
  43. Altman N. S.. An Introduction to Kernel and Nearest-Neighbor Nonparametric Regression. Am. Stat. 1992;46:175–185. doi: 10.1080/00031305.1992.10475879. [DOI] [Google Scholar]
  44. Liao H.-C., Lin Y.-H., Peng C.-H., Li Y.-P.. Directed Message Passing Neural Networks for Accurate Prediction of Polymer–Solvent Interaction Parameters. ACS Eng. Au. 2025;5:530–539. doi: 10.1021/acsengineeringau.5c00027. [DOI] [Google Scholar]
  45. Rong, Y. ; Bian, Y. ; Xu, T. ; Xie, W. ; Wei, Y. ; Huang, W. ; Huang, J. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: 2020; Vol. 33, pp 12559–12571. [Google Scholar]
  46. Liu, S. ; Wang, H. ; Liu, W. ; Lasenby, J. ; Guo, H. ; Tang, J. . Pre-Training Molecular Graph Representation with 3D Geometry, 2022. [Google Scholar]
  47. Ahmad, W. ; Simon, E. ; Chithrananda, S. ; Grand, G. ; Ramsundar, B. . ChemBERTa-2: Towards Chemical Foundation Models, 2022. [Google Scholar]
  48. Zhou, G. ; Gao, Z. ; Ding, Q. ; Zheng, H. ; Xu, H. ; Wei, Z. ; Zhang, L. ; Ke, G. In The Eleventh International Conference on Learning Representations, 2022. [Google Scholar]
  49. Hoang, V. T. ; Lee, O.-J. In Proceedings of the 39th AAAI Conference on Artificial Intelligence; AAAI Press: 2025; Vol. 39, pp 17204–17213. [Google Scholar]
  50. Chen R., Li C., Wang L., Liu M., Chen S., Yang J., Zeng X.. Pretraining Graph Transformer for Molecular Representation with Fusion of Multimodal Information. Inf. Fusion. 2025;115:102784. doi: 10.1016/j.inffus.2024.102784. [DOI] [Google Scholar]
  51. Inae, E. ; Liu, G. ; Jiang, M. In Proceedings of the Third Learning on Graphs Conference; PMLR: 2025, 41:1–41:15. [Google Scholar]
  52. Zhou Z., Li Y., Hong P., Xu H.. Multimodal Fusion with Relational Learning for Molecular Property Prediction. Commun. Chem. 2025;8:200. doi: 10.1038/s42004-025-01586-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Wishart D. S.. et al. DrugBank 5.0: A Major Update to the DrugBank Database for 2018. Nucleic Acids Res. 2018;46:D1074–D1082. doi: 10.1093/nar/gkx1037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Wu Z., Ramsundar B., Feinberg E. N., Gomes J., Geniesse C., Pappu A. S., Leswing K., Pande V.. MoleculeNet: A Benchmark for Molecular Machine Learning. Chem. Sci. 2018;9:513–530. doi: 10.1039/C7SC02664A. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Moriwaki H., Tian Y.-S., Kawashita N., Takagi T. M.. A Molecular Descriptor Calculator. Chem. inf. 2018;10:4. doi: 10.1186/s13321-018-0258-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Kuwahara H., Gao X.. Analysis of the Effects of Related Fingerprints on Molecular Similarity Using an Eigenvalue Entropy Approach. Chem. inf. 2021;13:27. doi: 10.1186/s13321-021-00506-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Ross J., Belgodere B., Chenthamarakshan V., Padhi I., Mroueh Y., Das P.. Large-Scale Chemical Language Representations Capture Molecular Structure and Properties. Nat. Mach. Intell. 2022;4:1256–1264. doi: 10.1038/s42256-022-00580-7. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

au6c01105_si_001.pdf (711.5KB, pdf)

Data Availability Statement

The original raw bone toxicity data utilized in this study were collected and curated by Chen et al. To ensure strict computational reproducibility and facilitate community deployment, the precise code for standardized data set splitting (executed under a fixed random seed), the final pretrained model weights, and the underlying source code for the proposed framework are fully open-sourced and freely available in the GitHub repository at https://github.com/hhuang-hub/MacroTox.


Articles from JACS Au are provided here courtesy of American Chemical Society

RESOURCES