Skip to main content
ACS Omega logoLink to ACS Omega
. 2026 Jun 3;11(23):33930–33942. doi: 10.1021/acsomega.6c00663

Explainable Bidirectional Long Short-Term Memory Networks Learn Chemistry from SMILES for Predicting Toxicity of Androgen and Estrogen Receptor Chemicals

Francesca Cutropia †, Fabrizio Mastrolorito †, Nicola Gambacorta ‡, Maria Vittoria Togo †, Vincenzo Amenduni †, Valentina Belgiovine †, Anna Rita Tondo †, Lydia Siragusa §,∥, Daniela Trisciuzzi †, Fulvio Ciriaco †,⊥,*, Nicola Amoroso †, Orazio Nicolotti †
PMCID: PMC13281003  PMID: 42326690

Abstract

Endocrine disruption remains a major concern in predictive toxicology, demanding accurate and interpretable models to assess molecular interactions with hormonal pathways. Here, we present a descriptor-free bidirectional long short-term memory (BiLSTM) framework designed to predict the toxicity of chemicals toward androgen (AR) and estrogen receptors (ER), two of the most complex and biologically relevant end points in toxicology. The model operates directly on SMILES strings, which are tokenized and converted into one-hot encoded sequences, enabling the automatic extraction of chemically meaningful representations without reliance on handcrafted molecular descriptors or fingerprints. To enhance interpretability, we introduce a novel explainable artificial intelligence (XAI) approach that aggregates character-level attribution scores into color-coded substructures, revealing features that drive or reduce toxicity and offering mechanistic insight into receptor-mediated effects. The models were trained on publicly available, high-quality data sets comprising 1664 and 1529 chemicals with experimental binary labels for AR and ER, respectively. Employing cross-validation analyses, based on 20% randomly stratified resampling iterated 10 times, the proposed workflow returned accuracy equal to 0.75 ± 0.08 and 0.81 ± 0.05, sensitivity equal to 0.66 ± 0.36 and 0.69 ± 0.17, and specificity equal to 0.76 ± 0.14 and 0.82 ± 0.06 for AR and ER end points, respectively. Our descriptor-free models ensure highly transparent results with a substructure level interpretability. These findings demonstrate the potential of deep learning directly on molecular textual representations to advance predictive toxicology and to support mechanistic understanding in chemical risk assessment.


graphic file with name ao6c00663_0010.jpg


graphic file with name ao6c00663_0008.jpg

1. Introduction

In recent years, computational approaches have become extremely effective for the prediction of critical human health and environmental end points, revealing tight links between toxicity patterns and structural motifs. − The importance of developing and implementing in silico alternative methods, known as New Approach Methodologies (NAM), − as reported in the REACH (Registration, Evaluation, Authorization and Restriction of Chemicals) regulation, lies in reducing the costs associated with chemical testing and the risk posed to humans, animals and the environment. − Although progress has been very significant, many proposed models still have limitations in terms of generalizability and interpretability.

Within computational toxicology, the employment of Natural Language Processing (NLP) can open new ways to encode and analyze molecular structures. NLP provides new options for exploring large collections of textual data, with the goal of uncovering and structuring latent knowledge. Advances in this domain have significantly enhanced various tasks that rely on language-based data to derive meaningful insights. Broadly, NLP involves the computational modeling, analysis, and manipulation of linguistic structures to facilitate interaction between human language and automated systems. As far as life science is concerned, NLP is becoming increasingly important for: (i) processing natural language in scientific articles, patents, and web resources; (ii) analyzing domain-specific sequence “languages” such as genomes and proteins; (iii) treating chemical notations, such as Simplified Molecular Input Line Entry Specification (SMILES), as symbolic languages for learning. In this respect, NLP can be applied to extract an unprecedented wealth of knowledge from curated databases such as PubChem, ChEMBL, UniProt, or DrugBank, whose chemical data are reported in textual form. Remarkably, NLP-based approaches can be also employed to empower de novo drug design strategies − by improving the exploration of the chemical spaces and the prediction of protein–ligand interactions (PLI) by capturing sequential motifs and chemical patterns relevant to binding specificity.

Among the various textual notations used to represent chemical structures, SMILES is by far the most widely adopted. SMILES strings provide a linear text-based representation that can be processed as linguistic sequences, enabling models to learn directly from strings. This perspective has led to the birth of the Chemical Language Processing (CLP), where molecules are processed analogously to natural language, atoms or fragments as “words” and entire molecules as “sentences.” Thus, CLP can be applied to assess toxicological end points such as endocrine disruption mediated by estrogen and androgen receptor (ER/AR) pathway activity, stress response pathways and cytotoxicity. , Recurrent neural networks (RNNs) are among the first architectures designed to capture sequential dependencies and have been widely employed to model temporal or symbolic data, as they maintain a memory of preceding inputs across the sequence. Their ability to propagate contextual information along the entire sequence makes them particularly well-suited for symbolic chemical languages such as SMILES, enabling models to identify structural relationships that extend beyond local contexts. Among different RNN architectures, Long Short-Term Memory (LSTM) networks are particularly effective in handling the syntactic complexity of SMILES, capturing long-range dependencies, contextual information, and character sequences. Although recurring architectures have been widely used for drug discovery purposes, their use in toxicity prediction remains underexplored. In this respect, the present model relies on a Bidirectional Long Short-Term Memory (BiLSTM) network, an advanced form of the LSTM architecture, as it learns molecular patterns directly from SMILES strings without relying on predefined descriptors or fingerprints. , By processing molecular input sequences in both forward and backward directions simultaneously, the BiLSTM enhances the ability to capture correlations within the sequence and directly learns toxicity relevant features from tokenized SMILES.

The growing availability of biochemical data in public repositories (e.g., PubChem, ChEMBL, UniProt, and VEGA) has further accelerated the development and applicability of such data-driven models, fostering strong interest within the scientific community. These data sets not only provide the bases for training robust CLP models but also enable systematic exploration of structure–function relationships at unprecedented scale and detail. Nevertheless, interpretability remains a pivotal challenge in AI-based methodologies. In this context, Explainable Artificial Intelligence (XAI) tools, such as SHAP, LIME, CAPTUM are used to provide mechanistic insights, highlight toxic substructures as structural alerts, and support decision-making in cheminformatics pipelines. −

Unlike feature-based representations widely used in QSAR modeling, , we present a descriptor-free approach which can learn directly from the SMILES strings and whose high interpretability is derived by translating token-level attributions into chemically meaningful substructures. In this scenario, we introduce a descriptor-free BiLSTM framework designed to predict the toxicity of chemicals affecting AR and ER. , The models are trained directly on SMILES representations that are tokenized and transformed into one-hot encoded sequences, enabling them to learn molecular features without the need for handcrafted descriptors. In addition, we introduce a novel XAI approach that aggregates character attributions into toxic and nontoxic substructure level maps with respect to AR and ER end points. The proposed framework improves interpretability relative to descriptor-driven models by resolving molecular substructure contributions.

2. Methods

2.1. Data Gathering

The data were collected from the open-access platform VEGA HUB, a widely recognized resource in the field of predictive toxicology and QSAR modeling. VEGA HUB contains more than 100 in silico classification and regression models and provides a substantial collection of curated data sets covering toxicological, ecotoxicological, physicochemical and toxicokinetic end points. , In particular, this work is focused on the AR and ER mediated effects, both central to the assessment of endocrine-disrupting chemicals (EDCs), which are known to interfere with hormonal activity by altering synthesis, transport, metabolism, or binding to endogenous hormone receptors. The data sets used in this study provide binary classification labels derived from experimental assays, distinguishing chemicals as toxic or nontoxic. Here, these terms are used solely to indicate AR/ER activity (e.g., active vs inactive) as defined in the VEGA AR and ER data sets; this end point specific terminology does not imply general systemic toxicity or overall safety, nor does it distinguish agonists from antagonists. The AR data set comprises 1664 chemicals, including 197 toxic and 1467 nontoxic, while the ER data set contains 1529 chemicals, with 89 toxic and 1440 nontoxic. AR and ER data sets are available as File_S1.csv and File_S2.csv in the Supporting Information. Such data sets were employed to train BiLSTM models to predict the toxicity of chemicals affecting AR and ER, using only tokenized SMILES representations as input.

2.2. Data Processing and Augmentation

SMILES strings were used as text-based molecular representations, capturing the atomic composition and bonding patterns of each chemical. Each SMILES was tokenized into chemically meaningful units (“tokens”) to facilitate computational processing while preserving the underlying structural information. These tokens serve as the primary input for our BiLSTM deep learning architecture. The vocabulary was built separately for AR and ER end points, including all unique characters found in the SMILES strings for each data set, such as common syntactic elements like double bonds (=), triple bonds (#), parentheses, and ring-closing digits. Multicharacter atom symbols (e.g., Cl, Br, or [nH]) were tokenized as single unique tokens to enhance model interpretability. In addition, special tokens, such as < START>, < END>, and < PAD>, were introduced to indicate the beginning and end of each sequence and to normalize input lengths for batch training, respectively.

The token sequences were then converted into numerical inputs using one-hot encoding, where each token is represented as a binary vector. Each token was represented as a one-hot vector, where a single bit was assigned the value 1 while all others were set to 0. These vectors served as input, thereby enabling the learning of relationships between the chemical information encoded in string representations and a given molecular property of interest, such as AR and ER toxicity.

To enhance both the quality and diversity of the training data, a well-established augmentation strategy was applied. ,, For each chemical in the training set, one canonical SMILES and four randomly generated noncanonical SMILES were created by using the RDKit library, providing alternative yet chemically equivalent representations of the same chemical. This approach promotes a more robust and generalizable representation of molecular information. By generating alternative yet equivalent molecular representations, SMILES augmentation enriches the training data set while avoiding the introduction of structural noise. AR and ER augmented data sets are available as File_S3.csv and File_S4.csv in the Supporting Information.

2.3. RNN Model

2.3.1. BiLSTM Model

We implemented a descriptor-free classifier based on a BiLSTM network in the PyTorch Lightning framework for AR and ER end point. The model was designed to classify chemicals directly from tokenized and one-hot encoded SMILES sequences without the use of predefined molecular descriptors. Input sequences were padded with a special <PAD> token to a fixed maximum length and processed using PyTorch’s pack_padded_sequence function, which permits the network to ignore the padding positions during sequence elaboration. Before this step, the tensors were transposed into the format required by PyTorch’s LSTM layer, namely (sequence length, batch size, features). The RNN architecture was based on BiLSTM cells to capture sequential dependencies and obtain hidden state representations of tokenized SMILES sequences. During the forward pass, the final hidden states from the forward and backward directions were concatenated to form a fixed-length vector. Following, a dense linear layer was applied to the resulting high-dimensional hidden state to gradually reduce dimensionality, configuration adopted by sentimental analysis models to improve efficiency. This processed representation was then mapped to a single output unit. During the training, the raw output scores were passed directly into the BCEWithLogitsLoss function, which internally applies the sigmoid transformation. At prediction time, the scores were converted into classification scores using a sigmoid function, and chemicals were classified as toxic or nontoxic based on a 0.5 threshold. The loss function was strengthened by the integration of a pos_weight parameter, which aimed to solve the class imbalance problem by assigning more weight to toxic chemicals and the Adam optimizer was adopted to update model weights. To obtain robust performance estimates, we employed a cross-validation strategy with Stratified ShuffleSplit consisting of ten iterations, in which the initial data set was divided in a training and cross-validation split, including the 80 and 20% of data, respectively. This configuration represents a widely accepted practice in deep learning, as it maximizes the amount of data available for model learning while retaining a sufficient portion for reliable performance evaluation. Given the high unbalance between the two classes of AR and ER data sets, the stratification was used to maintain the same proportion of toxic and nontoxic chemicals. The model was trained using the training split, and SMILES augmentation was applied only to this part. , The model was then evaluated using the independent validation split. The final results are given as the mean and standard deviation of the performance metrics across the ten iterations. A random hyperparameter search was conducted to identify the optimal configuration that minimizes validation loss. Training employed early stopping based on the first convergence plateau of the loss curve. The best performing setup consisted of a single BiLSTM layer with 32 hidden units, followed by a dimension reduction to 8 units. Dropout rate was set to 0.4 to reduce the overfitting, the learning rate was set to 0.001 and the batch size to 256. Furthermore, a scaffold-based splitting strategy was applied by computing Bemis–Murcko scaffolds for both investigated data sets. In this setting, compounds sharing the same scaffold were grouped together, and each scaffold group was iteratively used as an external test set, while the remaining compounds were used for training.

2.3.2. Performance Metrics

The classification performances were evaluated by using a cross-validation analysis, based on 20% randomly stratified resampling and iterated for 10 times, to obtain robust predictions and corresponding uncertainties. The quality of the models was measured by several key metrics. In particular, the accuracy (ACC), sensitivity or recall (sens), specificity (spec), balanced accuracy (bACC), Matthews Correlation Coefficient (MCC), positive predictive value (PPV) and negative predictive value (NPV), defined as follows:

ACC=TP+TNTP+TN+FP+FN
spec=TNTN+FP
sens=TPTP+FN
bACC=sens+spec2
MCC=TP×TN−FP×FN(TP+FP)(TP+FN)(TN+FP)+(TN+FN)
PPV=TPTP+FP
NPV=TNTN+FN

where TP, TN, FP, and FN are the number of true positives, true negatives, false positives and false negatives, respectively. For the sake of completeness, MCC computes the quality of predictions in the range from −1 to 1. In addition, we reported also the Area under the Receiver Operating Characteristic Curve (AUC-ROC).

2.4. Explainability Implementation

Deep neural networks are widely regarded as black-box models, given that their internal decision-making mechanisms lack transparency and are often challenging to interpret. To overcome this limitation, we implemented an explainability strategy in line with recent advances in XAI methods that aim to enhance transparency without compromising predictive accuracy. In this respect, we employed Captum, PyTorch’s unified interpretability library, and applied the Integrated Gradients algorithm to compute attribution scores for each SMILES token, thereby quantifying their relative contributions to the final classification, as shown in Figure . This approach transforms numerical gradients into chemically interpretable maps, thus closing the gap between abstract attributions and human understanding.

1.

1

Integrated Gradients algorithm was applied to compute attribution scores for each SMILES token, quantifying their local contribution to the model prediction. Scores were then projected from SMILES strings (a) to tokens (b), mapped and aggregated into contiguous substructures sharing the same attribution sign (c), thereby generating chemically interpretable color maps (d). Toxic and nontoxic substructures were separated according to the sign of their attribution. For syntactic tokens (e.g., parentheses, ring indices), highlighted with black rings in panel (a), attribution values were equally redistributed to the adjacent atoms.

For the ease of investigation, we carried out the explainability analysis on an expanded cross-validation randomly stratified resampling split, thereby including the 50% of the initial data set set for AR and ER end points. The attribution scores were first calculated at the SMILES token level and then projected onto the atoms of the molecular graph. This step was pivotal as it allowed us to transform numerical values into chemical information by mapping them from textual to structural representation of the chemicals.

To define substructure boundaries, contiguous atoms with the same attribution sign (either positive or negative) were grouped together. Attribution scores associated with syntactic elements, such as parentheses indicating ring closures, bond symbols, or ring indices, were redistributed evenly between the neighboring tokens to preserve their contextual contribution. The substructures were labeled based on their positive or negative computed contributions and a final score for classification was obtained by aggregating all the substructure attribution values. The substructures were generated by fragmentation based on the RDKit’s “Chem.FragmentOnBonds” function, which properly breaks bonds at the boundaries between SMILES tokens of opposite sign. When no bonds were available for cleavage, the entire chemical was considered as a single substructure. The resulting substructures were then canonicalized to ensure consistent SMILES representation. RDKit flags the cleavage points by inserting dummy atoms reported as *. To avoid chemically incoherent substructures, additional corrections were applied whenever excess dummy atoms were generated or when information related to valence, aromaticity, or hybridization was lost. In such cases, dummy atoms were removed and replaced with a single jolly atom connected to the appropriate bond type, thus providing chemically robust and consistent SMILES substructures. In particular, a number of 1070/749 initial substructures were generated for the AR/ER end point starting from 832/765 chemicals taken from the corresponding expanded cross-validation splits. In addition, a custom colormap was developed to visualize the results, with V max and V min corresponding to the maximum and minimum attribution values in the entire data set. This mapping produced color coded molecular highlights directly on the 2D structure and this allowed to easily visualize substructures associated with positive (in red, associated with toxicity) and negative (in green, associated with nontoxicity) contributions. The entire explainability framework is summarized in Figure .

Finally, we performed a consistency analysis of the substructure obtained from the attribution-based fragmentation phase, matching them back to the SMILES of the expanded cross-validation split for AR and ER end point, as shown in Figure . The rationale of this phase was to extract chemically meaningful substructure from their molecular context and to investigate their recurrence across the expanded cross-validation split. This approach enabled us to examine how the influence of a specific substructure on the classification result depended on its molecular surroundings, thereby evaluating the consistency of the decision model process. Importantly, this analysis allowed as not only to quantify the distribution and frequency of the selected substructures but also to investigate if they were part of larger substructures undetected in the fragmentation phase.

2.

2

(a) Attribution-based fragmentation analysis; (b) Curation phase and consistency analysis.

In this respect, we carried out a consistency analysis, by curating the substructures generated after the fragmentation by paying attention to avoid overlaps and redundancies. Specifically, dummy atoms at the beginning or end of a given substructure, as well as before branching brackets (e.g., “(O*)”), were removed, since they correspond to dummy terminals in branched chains. Subsequently, the information concerning the connectivity notation was removed (e.g., “C”). This procedure guaranteed that only syntactically and structurally valid substructures were stored, thereby allowing for a consistent and reliable match with the data set. After the curation phase, duplicates were removed and all substructures were canonicalized, thus obtaining unique and chemically valid substructures reported as SMARTS. In this way, motifs that were originally represented in multiple syntactic forms (e.g., “*C*”, “*C”, “C*”, “C”) were reduced to a single canonical representation (“C”), thus generating a chemically coherent SMARTS dictionary, whose elements represented univocal substructures formed by single or multiple atoms.

In particular, the AR and ER expanded cross-validation split allowed to generate two SMART dictionaries including 614 and 369 substructures. These curated substructures were then matched back to the chemicals in the expanded cross-validation split to evaluate their occurrence and clarify their role in toxic versus nontoxic classifications.

The substructures were prioritized by ranking them based on their absolute distance from the midpoint of the frequency distribution:

d=|N2−fi|

where N is the total number of SMILES in the expanded cross-validation split and f i is the number of SMILES in which substructure i occurs. This criterium awards substructures including most meaningful structural information and penalizes those that are rare or ubiquitous.

3. Results and Discussion

3.1. Model Performance

Model performance was evaluated based on cross-validation analysis, based on 20% randomly stratified resampling and iterated for 10 times. The training set was expanded with SMILES augmentation by using the canonical representations and 4 randomized equivalents per chemical. For clarity, the learning curves are provided in the Supporting Information (Figure S1). Training loss decreases consistently, while validation loss levels off after the initial epochs, suggesting limited improvements in generalization after the first few epochs. The overall model performance was quantified in terms of sensitivity, specificity, balanced accuracy, accuracy, MCC, PPV, NPV and AUC-ROC.

As shown in Table , BLSTMs returned satisfactory and robust results despite the complexity of AR and ER end points and their strongly imbalanced data. For the sake of completeness, the mean ROC curves from the 10-fold cross-validation are reported in the Supporting Information (Figure S2), together with the corresponding AUC-ROC mean ± standard deviation. For the sake of completeness, although a straight and fair statistical comparison is unfeasible being very different the amount of data and the employed algorithms, Tables S1 and S2 provide a comparative analyses against benchmarking models available in other papers published elsewhere. ,, Finally, we also carried out a 50% randomly stratified resampling, iterated for 10 times, for intentionally reducing the training set size, to evaluate the robustness of the model under data scarcity conditions. The results shown as Table S3 of the Supporting Information indicated that data augmentation was effective in ensuring nearly constant statistics irrespective of the proportion of excluded samples (20% vs 50%) in cross-validation analysis, thereby proving the reliability of the proposed models.

1. Performances of BiLSTM Model Based on Cross-Validation Analysis, Based On 20% Randomly Stratified Resampling and Iterated for 10 Times .

  sens spec bACC ACC MCC PPV NPV AUC ROC
AR 0.66 ± 0.36 0.76 ± 0.14 0.71 ± 0.11 0.75 ± 0.08 0.29 ± 0.15 0.22 ± 0.12 0.95 ± 0.04 0.76 ± 0.14
ER 0.69 ± 0.17 0.82 ± 0.06 0.75 ± 0.06 0.81 ± 0.05 0.30 ± 0.06 0.20 ± 0.03 0.98 ± 0.01 0.82 ± 0.04
a

Median and standard deviation values are also reported.

To provide a comprehensive overview, we also implemented a scaffold-based splitting strategy and the resulting metrics are provided in Tables S4 and S5. Finally, while large-scale pretrained chemical language models have demonstrated improved generalization across molecular property prediction tasks, we intentionally adopted a from-scratch training strategy to preserve deterministic token-level representations and enable a transparent mapping between attribution scores and chemically meaningful substructures.

3.2. Assessment of Concordance across Predicted Augmented SMILES

We carried out an augmentation experiment by using a random external data set of 1000 chemicals taken from ChEMBL with molecular weights in the range from 300 to 400 Da with the aim of assessing the concordance in prediction of multiple equivalent SMILES. Each chemical underwent the same preprocessing procedure earlier described. For each chemical, the predictions of the five equivalent SMILES were compared by calculating both the mean and the standard deviation of the predicted class. As shown in Figure , there was a clear tendency to form clusters based on the concordance of predictions. The first and most populated cluster, colored in green, comprised the equivalent SMILES whose predictions are provided with full concordance. The second cluster, colored in yellow, grouped chemicals with four out of five concordant predictions across multiple equivalent SMILES. These were characterized by intermediate score variability values (standard deviation ranging from 0.08 to 0.15). The third cluster, colored in red, collected chemicals for which the prediction was highly discordant, i.e., 3 to 2. Of course, this happens most often for chemicals whose prediction scores are in proximity of the decision boundary (average prediction ∼ 0.5) and whose score dispersion is highest. In this respect, it is worth noting that a large majority of chemicals returned a full prediction concordance, with only 101 returning a 4/5 concordance and 74 returning 3/5 concordance. As a whole, 825 out of the 1000 chemicals of the external data set, exhibited fully concordant predictions, thereby underscoring the reliability of our model.

3.

3

Scatter plot of prediction consistency across canonical and randomized multiple equivalent SMILES (n = 1000). The x-axis represents the mean predicted scores across the five equivalent SMILES, and the y-axis the corresponding standard deviation. Green points depict molecules with fully concordant predictions (5/5, n = 825), yellow points indicate molecules with a single discordant equivalent SMILES (4/5 concordance, n = 101), and red points highlight cases with two discordant equivalent SMILES (3/5 concordance, n = 74). Overall, 17% of the data set exhibited discordant predictions.

For completeness, we finally focused on the top 100 chemicals exhibiting the highest prediction uncertainty, determined by their absolute distance from the decision threshold, as follows:

|0.5−p−|

where p− is the average predicted values. The 100 chemicals with scores closest to 0.5 were selected as the most uncertain. The goal was to evaluate whether the canonical SMILES was an effective reference for model decision-making compared to its multiple equivalent SMILES representations. As shown in Figure , the canonical SMILES is sometimes placed at the boundaries of the distribution, thereby showing the highest variability.

4.

4

Prediction scores of the top 100 most uncertain chemicals across canonical and randomized equivalent SMILES. Orange and blue points indicate canonical and the four randomized equivalent SMILES, connected by blue lines for clarity.

This trend indicated that canonical SMILES did not always provide a stable reference point and, in the worst cases, can even negatively influence the final classification outcome. This result suggests that model predictions are not inherently tied to the canonical form of a given chemical, and that data augmentation through noncanonical SMILES effectively enhanced model robustness by mitigating representation-dependent bias.

3.3. Explainability and Consistency Analysis on AR and ER on Expanded Cross-Validation Split

Following the curation phase, a dictionary containing 614 and 369 substructures were generated for the AR and ER end points. Based on the distribution and frequency analysis, a number of 20 substructures were prioritized and visualized in violin plots, thereby illustrating the values for toxic (red) and nontoxic (green) local contributions. As shown in Figure a, the violin plot reports the top 20 chemical substructures within the AR expanded cross-validation split. Interestingly, only four SMARTS are provided with nearly all negative attributions, thereby providing nontoxic contributions. Overall, the distributions appear consistent, suggesting that the model assigns homogeneous contributions to recurring substructures.

5.

5

(a) Violin plot of the 20 most informative SMARTS, ranked by distance from half of the expanded cross-validation split of AR end point. Positive attributions (red) correspond to local toxic contributions, whereas negative attributions (green) indicate local nontoxic contributions; (b) Representative chemicals containing the “N” substructure, highlighted with color maps, whose intensity is proportional to the attribution scores. The “N” substructure contributes to local toxicity, contoured in red (maximum attribution = 0.35), when it is part of the nitro group while does not contribute to toxicity, contoured in green (minimum attribution = −0.31), when it is part of the amine, cyan and amide groups. Moving clockwise from the upper left corner: 4-nitro-m-xylene, the veterinary antimicrobial agent ronidazole, the triazole fungicide myclobutanil, N,N,N′,N′-tetrabutylhexane-1,6-diamine, the skin penetration enhancer laurocapram and the systemic fungicide carboxin (6-methyl-N-phenyl-2,3-dihydro-1,4-oxathiine-5-carboxamide).

The consistency analysis revealed that the contribution of individual substructures to the overall toxicity prediction is highly dependent on their specific molecular environment. As illustrated in Figure b, the nitrogen (“N”) substructure exemplifies this context dependency. When it is part of a nitro group, it plays as a strong local flag of toxicity, consistent with its well-known electrophilic properties making the nitro group an acknowledged toxic structural alert. In contrast, when the same nitrogen (“N”) substructure is part of the amine, amide, or cyano groups, it is more frequently associated with nontoxic chemicals. This apparent divergence highlights instead the importance of chemical context, emphasizing that identical substructures may exert opposing influences on toxicity depending on their bonding partners and molecular surroundings. The violin plot in Figure a reports the top 20 chemical substructures within the ER expanded cross-validation split. Four SMARTS show predominantly negative attributions (green, associated with predictions of nontoxicity). The distributions are globally consistent, suggesting that the model assigns homogeneous contributions to recurring fragments across the data set. The overall consistency of the distributions indicates that each chemical substructure receives comparable attribution values across its occurrences in different SMILES representations, reflecting coherent substructure level contributions also for ER end point.

6.

6

(a) Violin plot of the 20 most informative SMARTS, ranked by distance from half of the expanded cross-validation split for ER end point. Positive attributions (red) correspond to local toxic contributions, whereas negative attributions (green) indicate local nontoxic contributions; b) Representative chemicals containing the “c1c­(ccc*1)­C” substructure, highlighted with color maps, whose intensity is proportional to the attribution scores. The “c1c­(ccc*1)­C” substructure contributes always to the local toxicity (0.21 < attributions < 0.97). Moving clockwise from the upper left corner: ethyl2-[[7-[[2-(3-chlorophenyl)-2-hydroxyethyl]­amino]-5,6,7,8-tetrahydronaphthalen-2-yl]­oxy]­acetate, 2-[2,6-di­(propan-2-yl)­phenyl]-5-hydroxyisoindole-1,3-dione, 2-(1,3-dioxo-2,3-dihydro-1h-inden-2-yl)­quinoline-6,8-disulfonic acid, the pyrethroid insecticide tefluthrin, the diphenyl ether herbicide acifluorfen and the plasticizer metabolite mono­(2-ethylhexyl)­phthalate.

As shown in Figure b, we extracted six representative chemicals in which the substituted tolyl fragment “c1c­(ccc*1)­C” was detected and whose attribution scores, as reported in the violin plot, are always positive, thereby indicating a local increase of toxicity. Interestingly, these chemicals belong to the class of environmental contaminants such as pyrethroid insecticide tefluthrin, the diphenyl ether herbicide acifluorfen and the plasticizer metabolite mono­(2-ethylhexyl)­phthalate. Although these chemicals belong to different chemical classes (i.e., pesticides, herbicides, plasticizers, dyes and industrial additives), they are all recognized as environmental pollutants with potential endocrine disrupting properties. Attribution maps consistently highlighted the monoalkylated phenyl ring (c1c­(ccc1)­C*) as a toxicophoric driver of binding predictions. This observation is fully consistent with the structural pharmacology of the ER, where an aromatic hydrophobic core provides anchoring interactions with residues of the ligand-binding pocket. Overall, the combination of the developed model and explainability analysis revealed substructures consistent with the toxicophores reported in the literature for ER disruption, reinforcing the mechanistic plausibility of the predictions and demonstrating the ability of our descriptor-free model to recover meaningful toxicophore models.

3.4. External Case Studies

To further validate the performance of the proposed BiLSTM models, several external case studies were examined. Both the reliability of the model predictions and their interpretability were evaluated using local attribution values. In particular, we selected four uncovered chemicals, for which the experimental activity toward nuclear receptors was previously reported. − In particular, nandrolone and zidovudine were experimentally flagged as toxic , and nontoxic, respectively, based on their binding to the AR. Similarly, kaempferol and phenobarbital were reported as toxic and nontoxic based on their binding to the ER. ,,,

As shown in Figure , our descriptor-free BiLSTM model correctly predicted nandrolone and kaempferol as toxic with classification scores equal to 0.87 and 0.83, respectively; as well as azidovudine and phenobarbital were predicted as nontoxic with classification scores equal to 0.30 and 0.23. Importantly, the models provided substructure level explanations that were consistent with chemical knowledge: the steroidal core of nandrolone was highlighted as the main determinant of toxicity, whereas in zidovudine, the azido and pyrimidine fragments received negative contributions, coherently explaining their safety. For kaempferol, the phenol rings, known to be critical for ER toxicity, were emphasized as positive contributors, while the barbituric scaffold of phenobarbital was associated with negative attributions.

7.

7

External case studies for (a) AR and (b) ER end points. Toxic and nontoxic substructures are highlighted with red and green color maps, whose intensity is proportional to the attribution scores.

Overall, the case studies illustrate that our approach can correctly discriminate toxic from nontoxic chemicals, while at the same time capturing chemically meaningful substructures that explain the outcome. This alignment between prediction and chemical intuition supports the reliability of the model and illustrates the potential of explainable chemical language models in predictive toxicology.

4. Conclusion

In this study, we developed a descriptor free deep learning strategy that directly learns from SMILES strings to predict toxic and safe chemicals with a focus on AR and ER end points. By training directly on SMILES strings, the proposed BiLSTM architecture learned chemical patterns without relying on predefined descriptors, demonstrating that molecular information can be effectively captured from simple textual representations. The switching from linear string notation to molecular substructures underscore the potential of identifying structural alerts. In this perspective, our strategy provides transparent information on the decision-making process, addressing one of the main challenges in applying deep learning to predictive toxicology, namely the “black box” nature of neural networks.

Overall, our results confirm that deep learning approaches can extract molecular knowledge directly from string representations, providing a valuable tool for predictive toxicology research. By combining predictive performance with interpretability, descriptor free strategies based on SMILES strings offer not only a practical tool for screening and prioritizing chemicals, but also a conceptual framework for understanding structure–activity relationships independently.

Supplementary Material

ao6c00663_si_001.pdf (223.6KB, pdf)
ao6c00663_si_002.zip (186.8KB, zip)

Glossary

Abbreviations

ACC

accuracy

AR

androgen receptor

AUC-ROC

area under the receiver operating characteristic curve

bACC

balanced accuracy

BiLSTM

bidirectional long short-term memory

CLP

chemical language processing

EDCs

endocrine disrupting chemicals

ER

estrogen receptor

FN

false negatives

FP

false positives

LSTM

long short-term memory

MCC

Matthews correlation coefficient

NAM

new approach methods

NLP

natural language processing

NPV

negative predictive value

PPV

positive predictive value

REACH

registration, evaluation, authorization and restriction of chemicals

ReLU

rectified linear unit

RNNs

recurrent neural networks

sens

sensitivity

SMILES

simplified molecular-input line-entry system

spec

specificity

TN

true negatives

TP

true positives

XAI

eXplainable Artificial Intelligence

The data and code adopted to conduct experiments are available at GitHub URL https://sourceforge.net/projects/toxfromsmiles.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.6c00663.

  • Performance of the AR and ER models available on the VEGA platform (Table S1), performance of the publicly available AR and ER consensus models (Table S2), performances of the BiLSTM model based on cross-validation analysis, based on 20 and 50% randomly stratified resampling and iterated for 10 timesmedian and standard deviation values are reported (Table S3), scaffold splits analysis for AR by computing Bemis–Murcko scaffolds (Table S4), scaffold splits analysis for ER by computing Bemis–Murcko scaffolds (Table S5); training and validation loss curves (mean across 10 folds) for the BiLSTM models trained on the VEGA end points: (a) AR and (b) ERshaded areas represent the standard deviation across folds (Figure S1) and mean ROC curves obtained from 10-fold cross-validation for the BiLSTM models trained on the VEGA end points: (a) AR and (b) ERmean AUC-ROC and its standard deviation are reported (Figure S2) (PDF)

  • Molecular indices, SMILES strings, and binary label for theAR data set without data augmentation (File_S1); molecular indices,SMILES strings, and binary label for the AR data set with 5-fold dataaugmentation (File_S2); molecular indices, SMILES strings, andbinary label for the ER data set without data augmentation­(File_S3); molecular indices, SMILES strings, and binarylabel for the ER data set with 5-fold data augmentation (File_S4) (ZIP)

#.

N.A. and O.N. are equal last author contribution.

F.C., N.A., and O.N conceived and designed the study. F.Cu., F.M., and F.C. developed and implemented the BiLSTM framework and conducted the model training and evaluation. F.Cu., M.V.T., V.A., V.B., A.T., L.S., and D.T. curated and preprocessed the data sets and contributed to performance benchmarking. F.M., N.G., and F.C. developed the explainable AI (XAI) module and generated fragment-level attribution maps. F.Cu., N.G. M.V.T., V.A., V.B., A.T., L.S., and D.T. analyzed the results. N.A. and O.N wrote the manuscript. All authors reviewed, edited, and approved the final version of the manuscript.

The authors declare no competing financial interest.

References

  1. Canipa S. J., Chilton M. L., Hemingway R., Macmillan D. S., Myden A., Plante J. P., Tennant R. E., Vessey J. D., Steger-Hartmann T., Gould J., Hillegass J., Etter S., Smith B. P. C., White A., Sterchele P., De Smedt A., O’Brien D., Parakhia R.. A quantitative in silico model for predicting skin sensitization using a nearest neighbours approach within expert-derived structure-activity alert spaces. J. Appl. Toxicol. 2017;37:985–995. doi: 10.1002/jat.3448. [DOI] [PubMed] [Google Scholar]
  2. Mastrolorito F., Gambacorta N., Ciriaco F., Cutropia F., Togo M. V., Belgiovine V., Tondo A. R., Trisciuzzi D., Monaco A., Bellotti R., Altomare C. D., Nicolotti O., Amoroso N.. Chemical Space Networks Enhance Toxicity Recognition via Graph Embedding. J. Chem. Inf. Model. 2025;65:1850. doi: 10.1021/acs.jcim.4c02140. [DOI] [PubMed] [Google Scholar]
  3. Maunz A., Gütlein M., Rautenberg M., Vorgrimmler D., Gebele D., Helma C.. lazar: a modular predictive toxicology framework. Front Pharmacol. 2013;4:38. doi: 10.3389/fphar.2013.00038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Noga M., Michalska A., Jurowski K.. Application of toxicology in silico methods for prediction of acute toxicity (LD50) for Novichoks. Arch. Toxicol. 2023;97:1691–1700. doi: 10.1007/s00204-023-03507-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Pinto C. L., Mansouri K., Judson R., Browne P.. Prediction of Estrogenic Bioactivity of Environmental Chemical Metabolites. Chem. Res. Toxicol. 2016;29:1410–1427. doi: 10.1021/acs.chemrestox.6b00079. [DOI] [PubMed] [Google Scholar]
  6. Zhang H., Shen C., Liu R.-Z., Mao J., Liu C.-T., Mu B.. Developing novel in silico prediction models for assessing chemical reproductive toxicity using the naïve Bayes classifier method. J. Appl. Toxicol. 2020;40:1198–1209. doi: 10.1002/jat.3975. [DOI] [PubMed] [Google Scholar]
  7. Zhang H., Ren J.-X., Kang Y.-L., Bo P., Liang J.-Y., Ding L., Kong W.-B., Zhang J.. Development of novel in silico model for developmental toxicity assessment by using naïve Bayes classifier method. Reproductive Toxicology. 2017;71:8–15. doi: 10.1016/j.reprotox.2017.04.005. [DOI] [PubMed] [Google Scholar]
  8. Ball N., Bars R., Botham P. A., Cuciureanu A., Cronin M. T. D., Doe J. E., Dudzina T., Gant T. W., Leist M., van Ravenzwaay B.. A framework for chemical safety assessment incorporating new approach methodologies within REACH. Arch. Toxicol. 2022;96:743–766. doi: 10.1007/s00204-021-03215-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Isaacs K. K., Egeghy P., Dionisio K. L., Phillips K. A., Zidek A., Ring C., Sobus J. R., Ulrich E. M., Wetmore B. A., Williams A. J., Wambaugh J. F.. The chemical landscape of high-throughput new approach methodologies for exposure. J. Expo Sci. Environ. Epidemiol. 2022;32:820–832. doi: 10.1038/s41370-022-00496-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Madden J. C., Enoch S. J., Paini A., Cronin M.T. D. A Review of In Silico Tools as Alternatives to Animal Testing, Principles Resources and Applications. Alternatives to Laboratory Animals. 2020;48:146–172. doi: 10.1177/0261192920965977. [DOI] [PubMed] [Google Scholar]
  11. Parish S. T., Aschner M., Casey W., Corvaro M., Embry M. R., Fitzpatrick S., Kidd D., Kleinstreuer N. C., Lima B. S., Settivari R. S., Wolf D. C., Yamazaki D., Boobis A.. An evaluation framework for new approach methodologies (NAMs) for human health safety assessment. Regul. Toxicol. Pharmacol. 2020;112:104592. doi: 10.1016/j.yrtph.2020.104592. [DOI] [PubMed] [Google Scholar]
  12. Schmeisser S., Miccoli A., von Bergen M., Berggren E., Braeuning A., Busch W., Desaintes C., Gourmelon A., Grafström R., Harrill J., Hartung T., Herzler M., Kass G. E. N., Kleinstreuer N., Leist M., Luijten M., Marx-Stoelting P., Poetz O., van Ravenzwaay B., Roggeband R., Rogiers V., Roth A., Sanders P., Thomas R. S., Vinggaard A. M., Vinken M., van de Water B., Luch A., Tralau T.. New approach methodologies in human regulatory toxicology – Not if, but how and when! Environ. Int. 2023;178:108082. doi: 10.1016/j.envint.2023.108082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Trisciuzzi D., Alberga D., Leonetti F., Novellino E., Nicolotti O., Mangiatordi G. F.. Molecular Docking for Predictive Toxicology. Methods Mol. Biol. 1800;2018:181–197. doi: 10.1007/978-1-4939-7899-1_8. [DOI] [PubMed] [Google Scholar]
  14. Born J., Markert G., Janakarajan N., Kimber T. B., Volkamer A., Martínez M. R., Manica M.. Chemical representation learning for toxicity prediction. Digital Discovery. 2023;2:674–691. doi: 10.1039/D2DD00099G. [DOI] [Google Scholar]
  15. Landrum, G. RDKit Documentation, (n.d.).
  16. Lilienblum W., Dekant W., Foth H., Gebel T., Hengstler J. G., Kahl R., Kramer P.-J., Schweinfurth H., Wollin K.-M.. Alternative methods to safety studies in experimental animals: role in the risk assessment of chemicals under the new European Chemicals Legislation (REACH) Arch. Toxicol. 2008;82:211–236. doi: 10.1007/s00204-008-0279-9. [DOI] [PubMed] [Google Scholar]
  17. Nicolotti O., Benfenati E., Carotti A., Gadaleta D., Gissi A., Mangiatordi G. F., Novellino E.. REACH and in silico methods: an attractive opportunity for medicinal chemists. Drug Discovery Today. 2014;19:1757–1768. doi: 10.1016/j.drudis.2014.06.027. [DOI] [PubMed] [Google Scholar]
  18. Punt A., Bouwmeester H., Blaauboer B. J., Coecke S., Hakkert B., Hendriks D. F. G., Jennings P., Kramer N. I., Neuhoff S., Masereeuw R., Paini A., Peijnenburg A. A. C. M., Rooseboom M., Shuler M. L., Sorrell I., Spee B., Strikwold M., Van der Meer A. D., Van der Zande M., Vinken M., Yang H., Bos P. M. J., Heringa M. B.. New approach methodologies (NAMs) for human-relevant biokinetics predictions. Meeting the paradigm shift in toxicology towards an animal-free chemical risk assessment, ALTEX. 2020;37:607–622. doi: 10.14573/altex.2003242. [DOI] [PubMed] [Google Scholar]
  19. Schoeters G.. The REACH perspective: toward a new concept of toxicity testing. J. Toxicol Environ. Health B Crit Rev. 2010;13:232–241. doi: 10.1080/10937404.2010.483938. [DOI] [PubMed] [Google Scholar]
  20. Öztürk H., Özgür A., Schwaller P., Laino T., Ozkirimli E.. Exploring chemical space using natural language processing methodologies for drug discovery. Drug Discovery Today. 2020;25:689–705. doi: 10.1016/j.drudis.2020.01.020. [DOI] [PubMed] [Google Scholar]
  21. Bijral R. K., Singh I., Manhas J., Sharma V.. Exploring Artificial Intelligence in Drug Discovery: A Comprehensive Review. Arch Computat Methods Eng. 2022;29:2513–2529. doi: 10.1007/s11831-021-09661-z. [DOI] [Google Scholar]
  22. Weininger D.. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988;28:31–36. doi: 10.1021/ci00057a005. [DOI] [Google Scholar]
  23. Bolton, E.E. ; Wang, Y. ; Thiessen, P.A. ; Bryant, S.H. , Chapter 12PubChem: Integrated Platform of Small Molecules and Biological Activities, in: Wheeler, R.A. ; Spellmeyer, D.C. (Eds.), Annual Reports in Computational Chemistry; Elsevier, 2008: pp 217–241. 10.1016/S1574-1400(08)00012-1. [DOI] [Google Scholar]
  24. Gaulton A., Bellis L. J., Bento A. P., Chambers J., Davies M., Hersey A., Light Y., McGlinchey S., Michalovich D., Al-Lazikani B., Overington J. P.. ChEMBL: a large-scale bioactivity database for drug discovery. Nucleic Acids Res. 2012;40:D1100–D1107. doi: 10.1093/nar/gkr777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Apweiler R., Bairoch A., Wu C. H., Barker W. C., Boeckmann B., Ferro S., Gasteiger E., Huang H., Lopez R., Magrane M., Martin M. J., Natale D. A., O’Donovan C., Redaschi N., Yeh L.-S. L.. UniProt: the Universal Protein knowledgebase. Nucleic Acids Res. 2004;32:D115–119. doi: 10.1093/nar/gkh131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Wishart D. S., Knox C., Guo A. C., Shrivastava S., Hassanali M., Stothard P., Chang Z., Woolsey J.. DrugBank: a comprehensive resource for in silico drug discovery and exploration. Nucleic Acids Res. 2006;34:D668–672. doi: 10.1093/nar/gkj067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Mastrolorito F., Ciriaco F., Togo M. V., Gambacorta N., Trisciuzzi D., Altomare C. D., Amoroso N., Grisoni F., Nicolotti O.. fragSMILES as a chemical string notation for advanced fragment and chirality representation. Commun. Chem. 2025;8:26. doi: 10.1038/s42004-025-01423-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kutchukian P. S., Shakhnovich E. I.. De novo design: balancing novelty and confined chemical space. Expert Opin Drug Discov. 2010;5:789–812. doi: 10.1517/17460441.2010.497534. [DOI] [PubMed] [Google Scholar]
  29. Alberga D., Gambacorta N., Trisciuzzi D., Ciriaco F., Amoroso N., Nicolotti O.. De Novo Drug Design of Targeted Chemical Libraries Based on Artificial Intelligence and Pair-Based Multiobjective Optimization. J. Chem. Inf. Model. 2020;60:4582–4593. doi: 10.1021/acs.jcim.0c00517. [DOI] [PubMed] [Google Scholar]
  30. Wang C., Kumar G. A., Rajapakse J. C.. Drug discovery and mechanism prediction with explainable graph neural networks. Sci. Rep. 2025;15:179. doi: 10.1038/s41598-024-83090-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Michels J., Bandarupalli R., Ahangar Akbari A., Le T., Xiao H., Li J., Hom E. F. Y.. Natural Language Processing Methods for the Study of Protein–Ligand Interactions. J. Chem. Inf. Model. 2025;65:2191–2213. doi: 10.1021/acs.jcim.4c01907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Özçelik R., Grisoni F.. A hitchhiker’s guide to deep chemical language processing for bioactivity prediction. Digital Discovery. 2025;4:316–325. doi: 10.1039/D4DD00311J. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Mansouri K., Kleinstreuer N., Abdelaziz A. M., Alberga D., Alves V. M., Andersson P. L., Andrade C. H., Bai F., Balabin I., Ballabio D., Benfenati E., Bhhatarai B., Boyer S., Chen J., Consonni V., Farag S., Fourches D., García-Sosa A. T., Gramatica P., Grisoni F., Grulke C. M., Hong H., Horvath D., Hu X., Huang R., Jeliazkova N., Li J., Li X., Liu H., Manganelli S., Mangiatordi G. F., Maran U., Marcou G., Martin T., Muratov E., Nguyen D.-T., Nicolotti O., Nikolov N. G., Norinder U., Papa E., Petitjean M., Piir G., Pogodin P., Poroikov V., Qiao X., Richard A. M., Roncaglioni A., Ruiz P., Rupakheti C., Sakkiah S., Sangion A., Schramm K.-W., Selvaraj C., Shah I., Sild S., Sun L., Taboureau O., Tang Y., Tetko I. V., Todeschini R., Tong W., Trisciuzzi D., Tropsha A., Van Den Driessche G., Varnek A., Wang Z., Wedebye E. B., Williams A. J., Xie H., Zakharov A. V., Zheng Z., Judson R. S.. CoMPARA: Collaborative Modeling Project for Androgen Receptor Activity. Environ. Health Perspect. 2020;128:027002. doi: 10.1289/EHP5580. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Mansouri K., Abdelaziz A., Rybacka A., Roncaglioni A., Tropsha A., Varnek A., Zakharov A., Worth A., Richard A. M., Grulke C. M., Trisciuzzi D., Fourches D., Horvath D., Benfenati E., Muratov E., Wedebye E. B., Grisoni F., Mangiatordi G. F., Incisivo G. M., Hong H., Ng H. W., Tetko I. V., Balabin I., Kancherla J., Shen J., Burton J., Nicklaus M., Cassotti M., Nikolov N. G., Nicolotti O., Andersson P. L., Zang Q., Politi R., Beger R. D., Todeschini R., Huang R., Farag S., Rosenberg S. A., Slavov S., Hu X., Judson R. S.. CERAPP: Collaborative Estrogen Receptor Activity Prediction Project. Environ. Health Perspect. 2016;124:1023. doi: 10.1289/ehp.1510267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Bjerrum, E. J. ; Threlfall, R. , Molecular Generation with Recurrent Neural Networks (RNNs), 2017. 10.48550/arXiv.1705.04612. [DOI]
  36. Segler M. H. S., Kogej T., Tyrchan C., Waller M. P.. Generating Focused Molecule Libraries for Drug Discovery with Recurrent Neural Networks. ACS Cent. Sci. 2018;4:120–131. doi: 10.1021/acscentsci.7b00512. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Rao, K. V. ; Rao, K. N. ; Ratnam, G. S. , Accelerating Drug Safety Assessment using Bidirectional-LSTM for SMILES Data, 2024. 10.48550/arXiv.2407.18919. [DOI]
  38. Benfenati, E. ; Manganaro, A. ; Gini, G. . VEGA-QSAR: AI inside a platform for predictive toxicology, (n.d.).
  39. Lundberg, S. ; Lee, S.-I. . A Unified Approach to Interpreting Model Predictions, 2017. 10.48550/arXiv.1705.07874. [DOI]
  40. Ribeiro, M. T. ; Singh, S. ; Guestrin, C. . “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, 2016. 10.48550/arXiv.1602.04938. [DOI]
  41. Kokhlikyan, N. ; Miglani, V. ; Martin, M. ; Wang, E. ; Alsallakh, B. ; Reynolds, J. ; Melnikov, A. ; Kliushkina, N. ; Araya, C. ; Yan, S. ; Reblitz-Richardson, O. ;. Captum: A unified and generic model interpretability library for PyTorch, 2020. 10.48550/arXiv.2009.07896. [DOI]
  42. Togo M. V., Mastrolorito F., Ciriaco F., Trisciuzzi D., Tondo A. R., Gambacorta N., Bellantuono L., Monaco A., Leonetti F., Bellotti R., Altomare C. D., Amoroso N., Nicolotti O.. TIRESIA: An eXplainable Artificial Intelligence Platform for Predicting Developmental Toxicity. J. Chem. Inf Model. 2023;63:56–66. doi: 10.1021/acs.jcim.2c01126. [DOI] [PubMed] [Google Scholar]
  43. Mastrolorito F., Togo M. V., Gambacorta N., Trisciuzzi D., Giannuzzi V., Bonifazi F., Liantonio A., Imbrici P., De Luca A., Altomare C. D., Ciriaco F., Amoroso N., Nicolotti O.. TISBE: A Public Web Platform for the Consensus-Based Explainable Prediction of Developmental Toxicity. Chem. Res. Toxicol. 2024;37:323. doi: 10.1021/acs.chemrestox.3c00310. [DOI] [PubMed] [Google Scholar]
  44. Gambacorta N., Ciriaco F., Amoroso N., Altomare C. D., Bajorath J., Nicolotti O.. CIRCE: Web-Based Platform for the Prediction of Cannabinoid Receptor Ligands Using Explainable Machine Learning. J. Chem. Inf. Model. 2023;63:5916–5926. doi: 10.1021/acs.jcim.3c00914. [DOI] [PubMed] [Google Scholar]
  45. Kabier M., Gambacorta N., Ciriaco F., Mastrolorito F., Kumar S., Mathew B., Nicolotti O.. PoseidonQ: A Free Machine Learning Platform for the Development, Analysis, and Validation of Efficient and Portable QSAR Models for Drug Discovery. J. Chem. Inf. Model. 2025;65:3944–3954. doi: 10.1021/acs.jcim.4c02372. [DOI] [PubMed] [Google Scholar]
  46. Gambacorta N., Mastrolorito F., Togo M. V., Amenduni V., Mele M., Liantonio A., Mele A., De Luca A., Altomare C. D., Belgiovine V., Tondo A. R., Cutropia F., Siragusa L., Amoroso N., Ciriaco F., Imbrici P., Trisciuzzi D., Nicolotti O.. CUPID: A free drug discovery platform for the explainable multi-ion channel assessment of cardiotoxicity. Eur. J. Med. Chem. 2025;290:117575. doi: 10.1016/j.ejmech.2025.117575. [DOI] [PubMed] [Google Scholar]
  47. Vittoria Togo M., Mastrolorito F., Orfino A., Graps E. A., Tondo A. R., Altomare C. D., Ciriaco F., Trisciuzzi D., Nicolotti O., Amoroso N.. Where developmental toxicity meets explainable artificial intelligence: state-of-the-art and perspectives. Expert Opin Drug Metab Toxicol. 2024;20:561–577. doi: 10.1080/17425255.2023.2298827. [DOI] [PubMed] [Google Scholar]
  48. Amoroso N., Gambacorta N., Mastrolorito F., Togo M. V., Trisciuzzi D., Monaco A., Pantaleo E., Altomare C. D., Ciriaco F., Nicolotti O.. Making sense of chemical space network shows signs of criticality. Sci. Rep. 2023;13:21335. doi: 10.1038/s41598-023-48107-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Roncaglioni A., Lombardo A., Benfenati E.. The VEGAHUB Platform: The Philosophy and the Tools. Alternatives to Laboratory Animals. 2022;50:121. doi: 10.1177/02611929221090530. [DOI] [PubMed] [Google Scholar]
  50. Danieli A., Colombo E., Raitano G., Lombardo A., Roncaglioni A., Manganaro A., Sommovigo A., Carnesecchi E., Dorne J.-L.C.M., Benfenati E.. The VEGA Tool to Check the Applicability Domain Gives Greater Confidence in the Prediction of In Silico Models. IJMS. 2023;24:9894. doi: 10.3390/ijms24129894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Diamanti-Kandarakis E., Bourguignon J.-P., Giudice L. C., Hauser R., Prins G. S., Soto A. M., Zoeller R. T., Gore A. C.. Endocrine-disrupting chemicals: an Endocrine Society scientific statement. Endocr Rev. 2009;30:293–342. doi: 10.1210/er.2009-0002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Arús-Pous J., Johansson S. V., Prykhodko O., Bjerrum E. J., Tyrchan C., Reymond J.-L., Chen H., Engkvist O.. Randomized SMILES strings improve the quality of molecular generative models. J. Cheminform. 2019;11:71. doi: 10.1186/s13321-019-0393-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Li X., Fourches D.. SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning. J. Chem. Inf. Model. 2021;61:1560–1569. doi: 10.1021/acs.jcim.0c01127. [DOI] [PubMed] [Google Scholar]
  54. Kimber T. B., Gagnebin M., Volkamer A.. Maxsmi: Maximizing molecular property prediction performance with confidence estimation using SMILES augmentation and deep learning. Artificial Intelligence in the Life Sciences. 2021;1:100014. doi: 10.1016/j.ailsci.2021.100014. [DOI] [Google Scholar]
  55. Tai, K.S. ; Socher, R. ; Manning, C.D. . Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks, 2015. 10.48550/arXiv.1503.00075. [DOI]
  56. Bjerrum, E. J. SMILES Enumeration as Data Augmentation for Neural Network Modeling of Molecules, 2017. 10.48550/arXiv.1703.07076. [DOI]
  57. Yang K., Swanson K., Jin W., Coley C., Eiden P., Gao H., Guzman-Perez A., Hopper T., Kelley B., Mathea M., Palmer A., Settels V., Jaakkola T., Jensen K., Barzilay R.. Analyzing Learned Molecular Representations for Property Prediction. J. Chem. Inf. Model. 2019;59:3370–3388. doi: 10.1021/acs.jcim.9b00237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Bemis G. W., Murcko M. A.. The Properties of Known Drugs. 1. Molecular Frameworks. J. Med. Chem. 1996;39:2887–2893. doi: 10.1021/jm9602928. [DOI] [PubMed] [Google Scholar]
  59. Ali S., Abuhmed T., El-Sappagh S., Muhammad K., Alonso J., Confalonieri R., Guidotti R., Del Ser J., Díaz-Rodríguez N., Herrera F.. Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence. Inf. Fus. 2023;99:101805. doi: 10.1016/j.inffus.2023.101805. [DOI] [Google Scholar]
  60. Galati S., Di Stefano M., Martinelli E., Macchia M., Martinelli A., Poli G., Tuccinardi T.. VenomPred: A Machine Learning Based Platform for Molecular Toxicity Predictions. Int. J. Mol. Sci. 2022;23:2105. doi: 10.3390/ijms23042105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Manganelli S., Roncaglioni A., Mansouri K., Judson R. S., Benfenati E., Manganaro A., Ruiz P.. Development, validation and integration of in silico models to identify androgen active chemicals. Chemosphere. 2019;220:204–215. doi: 10.1016/j.chemosphere.2018.12.131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Sagawa, T. ; Kojima, R. . How Well Do Large-Scale Chemical Language Models Transfer to Downstream Tasks?, 2026. 10.48550/arXiv.2602.11618. [DOI]
  63. Nepali K., Lee H.-Y., Liou J.-P.. Nitro-Group-Containing Drugs. J. Med. Chem. 2019;62:2851–2893. doi: 10.1021/acs.jmedchem.8b00147. [DOI] [PubMed] [Google Scholar]
  64. Sipes N. S., Martin M. T., Kothiya P., Reif D. M., Judson R. S., Richard A. M., Houck K. A., Dix D. J., Kavlock R. J., Knudsen T. B.. Profiling 976 ToxCast Chemicals across 331 Enzymatic and Receptor Signaling Assays. Chem. Res. Toxicol. 2013;26:878–895. doi: 10.1021/tx400021f. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Arévalo-Salina E. L., Nishigaki T., Olvera L., González-Andrade M., Xolalpa-Villanueva W., López-Romero E. O., Soberón X., Saab-Rincón G.. Change in selectivity of estrogen receptor alpha ligand-binding domain by mutations at residues H524/L525. Biochim. Biophys. Acta, Gen. Subj. 2025;1869:130775. doi: 10.1016/j.bbagen.2025.130775. [DOI] [PubMed] [Google Scholar]
  66. Trisciuzzi D., Alberga D., Mansouri K., Judson R., Novellino E., Mangiatordi G. F., Nicolotti O.. Predictive Structure-Based Toxicology Approaches To Assess the Androgenic Potential of Chemicals. J. Chem. Inf. Model. 2017;57:2874–2884. doi: 10.1021/acs.jcim.7b00420. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Trisciuzzi D., Alberga D., Mansouri K., Judson R., Cellamare S., Catto M., Carotti A., Benfenati E., Novellino E., Mangiatordi G. F., Nicolotti O.. Docking-Based Classification Models for Exploratory Toxicology Studies on High-Quality Estrogenic Experimental Data, Future. Med. Chem. 2015;7:1921–1936. doi: 10.4155/fmc.15.103. [DOI] [PubMed] [Google Scholar]
  68. Piir G., Sild S., Maran U.. Binary and multi-class classification for androgen receptor agonists, antagonists and binders. Chemosphere. 2021;262:128313. doi: 10.1016/j.chemosphere.2020.128313. [DOI] [PubMed] [Google Scholar]
  69. Kleinstreuer N. C., Ceger P., Watt E. D., Martin M., Houck K., Browne P., Thomas R. S., Casey W. M., Dix D. J., Allen D., Sakamuru S., Xia M., Huang R., Judson R.. Development and Validation of a Computational Model for Androgen Receptor Activity. Chem. Res. Toxicol. 2017;30:946–964. doi: 10.1021/acs.chemrestox.6b00347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Judson R. S., Magpantay F. M., Chickarmane V., Haskell C., Tania N., Taylor J., Xia M., Huang R., Rotroff D. M., Filer D. L., Houck K. A., Martin M. T., Sipes N., Richard A. M., Mansouri K., Setzer R. W., Knudsen T. B., Crofton K. M., Thomas R. S.. Integrated Model of Chemical Perturbations of a Biological Pathway Using 18 In Vitro High-Throughput Screening Assays for the Estrogen Receptor. Toxicol. Sci. 2015;148:137–154. doi: 10.1093/toxsci/kfv168. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ao6c00663_si_001.pdf (223.6KB, pdf)
ao6c00663_si_002.zip (186.8KB, zip)

Data Availability Statement

The data and code adopted to conduct experiments are available at GitHub URL https://sourceforge.net/projects/toxfromsmiles.


Articles from ACS Omega are provided here courtesy of American Chemical Society

RESOURCES