Abstract
The mutagenicity of chemical compounds is a key consideration in toxicology, drug development, and environmental safety. Traditional methods such as the Ames test, while reliable, are time-intensive and costly. With advances in imaging and machine learning (ML), high-content assays like cell painting offer new opportunities for predictive toxicology. Cell painting captures extensive morphological features of cells, which can correlate with chemical bioactivity. In this study, we leveraged cell painting data to develop ML models for predicting mutagenicity and compared their performance with structure-based models. We used two datasets: a Broad Institute dataset containing profiles of over 30 000 molecules and a U.S.-Environmental Protection Agency dataset with images of 1200 chemicals tested at multiple concentrations. By integrating these datasets, we aimed to improve the robustness of our models. Among three algorithms tested—Random Forest, Support Vector Machine, and Extreme Gradient Boosting—the third showed the best performance for both datasets. Notably, selecting the most relevant concentration per compound, the phenotypic altering concentration, significantly improved prediction accuracy. Our models outperformed traditional quantitative structure activity relationship (QSAR) tools such as the Virtual models for property Evaluation of chemicals within a Global Architecture (VEGA) and the CompTox Dashboard for the majority of compounds, demonstrating the utility of cell painting features. The cell painting-based models revealed morphological changes related to DNA and RNA perturbation, especially in mitochondria, endoplasmic reticulum and nuclei, aligning with mutagenicity mechanisms. Despite this, certain compounds remained challenging to predict due to inherent dataset limitations and inter-laboratory variability in cell painting technology. The findings highlight the potential of cell painting in mutagenicity prediction, offering a complementary perspective to chemical structure-based models. Future work could involve harmonizing cell painting methodologies across datasets and exploring deep learning techniques to enhance predictive accuracy. Ultimately, integrating cell painting data with QSAR descriptors in hybrid models may unlock novel insights into chemical mutagenicity.
Keywords: mutagenicity, cell painting, QSAR, chemical risk assessment
Introduction
Assessing the potential mutagenicity of chemical compounds is a crucial step not only in drug development but also in chemical risk assessment, environmental safety, and regulatory compliance [1, 2]. Traditional methods for assessing mutagenicity often rely on in vivo and in vitro tests, such as the Ames test [3], which can be time-consuming, laborious, and expensive. To limit the need for experimental lab investigations, in silico models have been developed to predict mutagenicity [4–6]. With the advent of high-throughput screening technologies and advanced computational methods, there is growing interest in cell-based assays to predict toxicity endpoints [7–10]. A promising and increasingly used technique in cell biology is cell painting, a high-content imaging assay that captures a wide range of morphological features of cells [11]. By staining cells with a combination of fluorescent dyes, cell painting makes it possible to visualize and quantify various cellular components, including the nucleus, cytoskeleton, and organelles. This technique can be applied to large sets of chemical compounds in a high-throughput manner to detect cell morphology properties in relation to pharmaceutical and/or toxicological effects [12].
Cell painting is a high-throughput phenotypic profiling (HTPP) technique, which can be leveraged in various innovative and impactful ways across numerous scientific studies. For example, it can be combined with high-throughput transcriptomics to quantify Per and polyfluoroalkyl substances (PFAS) -induced hepatosteatosis and mitochondrial damage, facilitating the calculation of point of departure values for phenotypic and transcriptomic responses using the Bayesian benchmark dose modeling approach [13]. Other studies integrating cell painting with transcriptomics (L1000) and/or chemical structure of compounds have suggested the mechanism of action and bioactivity of a molecule [14, 15]. In a more extensive way, an unbiased morphology-based atlas of genome-wide perturbations by combining the cell painting dataset with CRISPR-Cas9-based knockout experiments has been developed. It allows for precise linking of genetic perturbations to phenotypic morphological changes [16].
Recently, a chemoinformatics approach mining large-scale HTPP data to identify compounds with novel mechanisms of action has been reported, expanding the scope for drug development [17]. Cell painting assays can also be used to profile morphological changes in human osteosarcoma cells exposed to chemicals from plastic materials, providing insights that are comparable with cytotoxicity, osteogenesis, and in silico toxicity assays [18]. In another study, the morphological features from cell painting were combined with bioactivity readouts, revealing deeper phenotypic and cellular process insights [19]. Finally, some recent studies have used machine and deep learning modeling applied to cell painting data. As an example, compound activity across diverse assays can be predicted, enhancing drug discovery and toxicology studies [20]. The estrogenic exposure in breast cancer microtissues can be very precisely classified [21] and a convolution neural network (CNN) has been developed with cell painting data to accurately model the phenotypic effects of treatments on cells [15]. These diverse applications underscore the versatile utility of cell painting data in advancing biological and pharmaceutical research.
In our study, the objective was to develop machine learning (ML) models using, to our knowledge for the first time, cell painting data to predict the mutagenicity of various molecules and to compare their accuracy with more classical quantitative structure activity relationship (QSAR) models [22]. To do this, we used two independent cell painting datasets. The first dataset, generated by the Broad Institute, provided cellular phenotypic profiles in response to over 30 000 chemical molecules. The second dataset, from the US-EPA, provided a diverse range of cellular images corresponding to exposures to 1200 chemical molecules tested at different concentrations. By integrating these datasets, we sought to improve the robustness and generalizability of our mutagenicity prediction models. By extracting morphological features from cell paint images and then applying ML algorithms, we assumed that the morphological changes captured in the cell painting images would correlate with the mutagenic potential of the molecules, enabling accurate prediction of mutagenicity. The obtained results were promising and reinforced the growing amount of evidence supporting the integration of high-content imaging and advanced computational methods in toxicology and risk assessment.
Materials and methods
Cell painting datasets
The first cell painting dataset employed in this research was sourced from the Broad Institute’s Cell Profiling platform [11]. This dataset covers a vast collection of high-content imaging data capturing cellular morphology and subcellular organization of cells. The dataset includes 30 616 chemicals evaluated on U2OS human osteosarcoma cells and is available at http://gigadb.org/dataset/100351 (accessed 25 April 2024). The raw dataset contains a total of 406 384-well plates, with each plate containing 64 Dimethyl sulfoxide (DMSO) controls. Each compound was replicated 4−12 times across various plates. The images and the morphological features were identified using CellProfiler (http://cellprofiler.org/) [23]. The 1783 features that are identified encompassed diverse parameters, including intensity, size, area shape, texture, entropy, correlation, granularity, and the angle between neighboring cells [24]. Details about the features nomenclature can be found at the following link: https://github.com/carpenter-singh-lab/2023_Cimini_NatureProtocols/wiki/What-do-Cell-Painting-features-mean%3F (accessed 4 March 2025).
The second cell painting dataset used was acquired from the US-EPA’s Center for Computational Toxicology and Exposure [9]. It contains 1201 unique chemicals extracted from the ToxCast chemical library [25], tested for 24 hours at eight concentrations each, with one technical replicate and four biological replicates (i.e. independent cell cultures) on the U2OS human osteosarcoma cells. The original dataset comprises 29 384-well plates. Each dose plate holds a maximum of 42 test chemicals and three reference chemicals [dexamethasone, etoposide, all-trans retinoic acid (ATRA)] at eight concentrations with half-log-10 spacing. Additionally, each dose plate included three wells of a reference chemical (trichostatin A) and three wells of a cell viability positive control chemical (staurosporine) at a single concentration, along with 18 wells of vehicle control (0.5% DMSO). These dose plates were utilized to treat four assay plates across four biological replicates, representing independent cultures. The raw dataset contains 1300 features extracted using an Opera Phenix High Content Screening System and Harmony® software (v4.8).
Different levels of data were available for the two cell painting datasets. Numbered from 1 to 5, they represent the transition from raw data to fully processed data.
We have chosen to use Level 4, which corresponds to normalized morphological profile data per plate, with control and replicates z-scores. By aggregating single-cell morphological features at the well level, it reduces biological noise and technical variability. This aggregation also ensures proper correspondence between feature profiles and experimental annotations, which are defined at the well or treatment level. Moreover, working at this level significantly reduces data dimensionality and computational load, facilitating model training and interpretation. Finally, Level 4 profiles enhance reproducibility and generalizability across datasets by minimizing stochastic variation present at the single-cell level. Details of all levels are available at this address: https://clue.io/connectopedia/pdf/cell_painting_data_levels (accessed 24 September 2024).
Preprocessing stages
The data preprocessing stages performed included normalization (for the Broad Institute dataset only), feature selection, and spherization.
The normalization step was performed using the “normalize” function from the pycytominer package [26]. Profiles were normalized plate-wise, using DMSO-treated wells as the reference condition. The procedure leveraged predefined lists of morphological features and metadata to ensure consistent processing across samples. We applied the mad_robustize method, which scales each feature by subtracting the median and dividing by the median absolute deviation (MAD) calculated from the reference samples, thus providing robust normalization while mitigating the influence of outliers. As mentioned above, this method has only been applied to the Broad Institute dataset, as the US-EPA dataset is made available already normalized. They used the MAD method that is quite similar to the “mad_robustize” method.
For the basic feature selection, we used the “feature_select” function (also provided by the pycytominer package). We selected the operations “drop_na_columns,” “variance_threshold,” and “correlation_threshold,” which allowed us to respectively remove columns containing excessive missing data, discard columns with low variability and eliminate correlated columns to enhance data quality and ensure model robustness. We also set the cutoff of NA values to three, the correlation threshold to 0.95 and, for the Broad Institute dataset only, we provided the list of 87 blocklist features that are known to be noisy and generally unreliable [27]. A second feature selection step was carried out by applying a statistical test (a Wilcoxon-Mann-Whitney test) to the features to determine which of them were significant (P-value threshold at .05) in discriminating between mutagenic and nonmutagenic molecules.
To reduce potential biases caused by differences in image intensity across plates or experimental batches, we applied a normalization technique known as spherization. This method transforms the data so that all features contribute equally and are distributed in a way that is more comparable across conditions. It enhances the reproducibility of results and allows for more reliable comparisons between treatments. We implemented this process using the “normalize” function from the pycytominer package [26], applied to the previously filtered set of morphological features. DMSO-treated wells served as the reference samples, and the normalization method used was “spherize,” which standardizes the data by centering and projecting it onto a hypersphere.
Dose selection
Since each compound may act differently at different doses, and when the dataset allowed, we selected one dose for each single compound in our study. For the US-EPA dataset, all the compounds were tested at eight different dose level (min dose: 0.006 μM and max dose: 105 μM, see the distribution of concentrations in Supplementary Fig. SS1).
For each compound, a value for the most relevant dose is provided under the name “PAC,” for phenotypic altering concentration (PAC). It is defined in a previous study as the lower value between the global BMC and the category-level BMC [9], also defined as the concentration at which at least 30% of elements show a phenotypic effect of the compound while still being below the toxic dose. The file with the PAC values is available at the following link: https://clowder.edap-cluster.com/files/632b651ae4b04f6bb1354411?dataset=61147fefe4b0856fdc65639b&space=&folder=632b650be4b04f6bb135440d (accessed 22 July 2024). If the information is available, the dose selected for each of the compounds is the one indicated in the “cp_loec_dose_level” column (compounds-dependent). Otherwise, we identified the average selected dose level (column “cp_loec_dose_level”) among the compounds with a PAC, which corresponds to Level 7, and used it as the default concentration for compounds with no determined PAC. This strategy ensured a consistent and biologically relevant dose selection across the full set of compounds. Among the 1201 unique compounds, only 384 have a usable (calculated) PAC value. The doses of the compounds in this dataset range from 0.07 to 104 uM, with an average dose of 38.95 uM (±26.45).
An alternative dataset was set up with the same compounds, but all at dose-level 7, to compare the benefits of selecting the optimum dose (PAC) versus an imposed dose (dose-level 7). For this alternative dataset, the dose ranges from 6 uM to 31 uM with an average dose of 29.61 uM (±2.53).
To differentiate the two datasets, the one with the optimum-selected doses will be named “US-EPA_d7/PAC” and the one with the imposed dose “US-EPA_d7.”
Mutagenicity datasets
To determine whether or not a compound is mutagenic, we used a consensus of different sources. First, we used the ISSTOX Chemical Toxicity Databases (https://www.iss.it/isstox, accessed 25 April 2024) and in particular the ISSSTY results concerning the in vitro mutagenicity in Salmonella typhimurium (Ames test). It gathers experimental results for 7367 compounds, retrieved from the Chemical Carcinogenesis Research Information System (CCRIS) database, in the Toxnet databases cluster. To date, the TOXNET resource is not available anymore, but the information can be retrieved by using the filter “CCRIS” as “source name” in the Pubchem database. The following codes were used: 1 = negative; 2 = equivocal; 3 = positive; Inc = inconclusive and ND = no data.
Then, the EURL ECVAM Genotoxicity and Carcinogenicity Consolidated Database, which compiles data on genotoxicity and carcinogenicity of substances tested in the Ames test was considered. A curated collection of 211 substances with negative results and 726 substances with positive results in the Ames test has been added, complementing the previously published database. This database serves as a reference for various activities in genotoxicity testing across different sectors, aiding regulatory initiatives, project development, and the validation of new tests. Detailed information on the database construction is available in Madia et al. [28].
Third, we used a mutagenicity dataset from a study Hansen et al. [29]. This dataset contains ~6500 nonconfidential compounds along with Ames mutagenicity annotation. The report compares the performance of three commercial tools with four noncommercial ML implementations using this benchmark dataset.
A conservative consensus for mutagenicity determination has been automated: if at least one of the sources indicates that the compound is mutagenic, then a compound is classified as mutagenic. When the different sources disagree between “nonmutagenic” and “equivocal,” then the “equivocal” status is left and the compound is not considered mutagenic or nonmutagenic. Other notations such as “Inconclusive” or “No data” are considered as an absence of data and have no effect on the classification of other sources. At the end, the Broad dataset used for modeling comprises 111 compounds (15 mutagenic, 96 nonmutagenic). The “US-EPA_d7/PAC” and “US-EPA_d7” datasets include 307 (66 mutagenic, 241 nonmutagenic) and 294 compounds (55 mutagenic, 239 nonmutagenic), respectively. The list of mutagenicity consensus for each compound is available in Supplementary Table SS1.
GRIT score calculation
The GRIT (Global Robustness and Informativeness Test) score serves as a metric for evaluating the cellular activity patterns of compounds using cell painting datasets. It assesses both the robustness and informativeness of a compound’s impact on cells. Robustness measures the consistency of a compound’s effect across different biological replicates, calculated by assessing the standard deviation of the compound’s activity profile across replicates. Informativeness, gauges how distinct a compound’s effect is compared with a control condition, determined using distance metrics to compare the compound’s activity profile with that of controls. The final GRIT score is derived by integrating the robustness and informativeness measures, which is obtained by averaging the two components. The GRIT score was calculated using the cytominer_eval package (available at https://github.com/broadinstitute/grit-benchmark, accessed 25 April 2024) on Level 4 data. We compared all the feature values of all replicates of a particular compound with those of the DMSO controls from the plates where the compound was tested.
Feature selection
As the feature selection is important and may impact the performance of the model [30], an additional step of selecting the most significant features to distinguish mutagenic from nonmutagenic compounds was applied for Level 4 data. This involves calculating the P-value of a Mann–Whitney U test for each feature. If this value is below the .05 threshold, then it is considered discriminative for mutagenicity. The list of retained features for each dataset is available in Supplementary Table SS2.
Mapping visualization (Clustermap & t-SNE)
During our study, we visualized the cell painting datasets while clustering similar compounds according to feature values and clustering features together as well. The aim was to preview whether the feature values are useful for clustering mutagenic and nonmutagenic compounds.
We used the clustermap function (default parameters and “correlation” metric), from the Seaborn python library to generate hierarchical clustering heatmaps based on the cell painting dataset. The resulting dendrogram was visualized as a heatmap, where rows and columns correspond to samples, and the intensity of the colors represented the feature values of each sample.
To further explore the high-dimensional structure of the dataset and visualize the relationships between samples, we employed t-distributed stochastic neighbor embedding (t-SNE) [31]. We used the Python library “sklearn.manifold” to reduce the dimensionality of the dataset while emphasizing local relationships between samples. t-SNE was applied with default parameters (perplexity of 30, learning rate automatically calculated, and a maximum of 1000 iterations) to embed the samples into a 2D space. The resulting t-SNE plot provided a low-dimensional representation of the dataset, where each point represented a sample, and the spatial distribution reflected local neighborhood structures rather than preserving global distances.
Modelization/prediction
To assess whether the cell painting datasets are effective in predicting the mutagenicity of compounds, we tested different ML models and evaluated their performances. For each of the models tested, in addition to the initial models, we used the nested cross-validation (CV) protocol shown in Fig. 1 and detailed in each model section.
Figure 1.
Visual representation of nested CV applied on the cell painting datasets.
Random Forest
In this study, we used the Random Forest (RF) algorithm to predict the mutagenicity of compounds using cell painting data. Specifically, we used the Python programming language and the scikit-learn package (v1.3.0) [32] to implement the RF classifier. As shown in Fig. 1, the dataset was split into training and testing sets in an 80/20 or 70/30 ratio depending on the size of the initial dataset. As the mutagenic and nonmutagenic classes are unbalanced, we used the SMOTE tool (package imblearn.over_sampling v0.12.0, available at https://github.com/scikit-learn-contrib/imbalanced-learn/blob/master/imblearn/over_sampling/_smote/base.py, accessed 21 May 2024) that can handle highly skewed datasets with large class imbalances by oversampling the dataset, which aims to create a more robust and fair decision-making boundary.
To optimize the initial model and prevent overfitting, we used a nested CV protocol. The inner loop performed a grid search with CV to fine-tune the hyperparameters, including the number of trees, the maximum depth of each tree, and the minimum samples split. The outer loop provided an unbiased evaluation of the model’s performance. The CV procedure was repeated several times to improve the estimated performances with the “repeated stratified k-fold” function (2 x 5-fold using “n_splits = 5” and “n_repeats = 2” arguments).
Feature importance scores based on mean decrease in impurity were calculated to identify the most influential features in the classification task, using the “feature_importances_” function from scikit-learn package. The model’s performance was evaluated using metrics (also from scikit-learn package) such as balanced accuracy (average of the recall obtained for each class), precision (ability of positive predictions), recall (completeness of positive predictions), and the F1 score (harmonic mean of the precision and recall). Additionally, confusion matrices were generated to visualize the classification results and identify any misclassification patterns.
XGBoost
In addition to RF, we also tested another method for modeling the mutagenicity of compounds using Extreme Gradient Boosting (XGBoost) models. It is a powerful and efficient implementation of gradient boosting for supervised learning problems, which is widely used in various ML competitions and real-world applications due to its high predictive power and computational efficiency. We used the XGBoost Python module (v2.0.3) [33]. As for the RF modelization, the dataset was split into training and testing sets in an 80/20 or 70/30 ratio, the SMOTE tool was used for the training dataset resampling in combination with the nested CV procedure.
Support Vector Machine
As a third prediction method modeling, we considered Support Vector Machines (SVM) to predict the mutagenicity of chemicals based on their cell painting features. The SVM model was implemented using the Scikit-Learn library in Python. A radial basis function kernel was chosen due to its ability to handle nonlinear relationships in the data. Hyperparameters, including the regularization parameter (C) and the kernel coefficient (gamma), were optimized using a grid search with nested CV to avoid overfitting. The dataset was split into training and testing sets in an 80/20 or 70/30 ratio and resampling using SMOTE to ensure equal distribution of mutagenic and nonmutagenic compounds in the training subset.
Structural data
The performance of our prediction models was compared with the performance of models based on structural descriptors. Two types of descriptors were used and are described in this section.
Mordred descriptors
Mordred is a Python module (v 1.2.0) [34] that calculates more than 1800 2D and 3D molecular descriptors, and is freely available at https://github.com/mordred-descriptor/mordred (accessed 21 May 2024). Mordred has the advantage of automatically testing all the descriptors to check whether it can calculate accurate results using the reference values of the molecular descriptors (for more details in the original publication, see Motivation and Mordred concepts).
To calculate the molecular descriptors, we used the “MolFromSmiles” function from the RDKit tool (rdkit.Chem 2023.09.6) and the “Calculator” function from the Mordred python module.
The output descriptors were then subjected to a preprocessing step, where descriptors with low variance (<0.05) values, missing data, correlated descriptors (>0.95) or which are not significant for discriminating mutagenicity (Student’s t-test with a P-value threshold at .05) were excluded from subsequent analysis, ensuring the quality and relevance of the descriptor dataset used in this study. From the original 1826 descriptors provided, only 82 for the Broad and 73 were kept for the EPA dataset, respectively.
Morgan FingerPrint
We also used another powerful method to describe the structure of molecules as vectors: the Morgan Molecular Fingerprint [35]. It is a widely used and one of the most popular fingerprints in research [36]. As for Mordred, we used molecules in SMILE, the “MolFromSmiles” & “standardize_smiles” functions from the MolVS python package (v 0.1.1) freely available at https://github.com/mcs07/MolVS (accessed 21 May 2024) and the “GetMorganFingerprintAsBitVect” function from rdkit.Chem module by setting the radius to two and the number of bits to 2048.
Results
Datasets description
Broad institute dataset
The final Broad Institute dataset contains 230 compounds (including DMSO) of which 66 are mutagenic and 164 nonmutagenic. A set of 1105 features remains after the basic feature selection step, 361 are related to the “cells” compartment, 327 to the “cytoplasm” compartment, and 417 to the “nuclei” compartment.
US-EPA dataset
The US-EPA dataset remaining after all the processing steps is composed of 579 compounds (including the DMSO), of which 122 are mutagenic and 457 nonmutagenic. They are described by 694 features related to the following cellular compartments: AGP (actin, Golgi, and plasma membrane, 125 features), DNA (155 features), endoplasmic reticulum (141 features), mitochondria (134 features), and RNA (131 features). An exception is made for eight features that do not concern these compartments, but rather various shape and position measurements: “Position_both_Centroid_Dis,” “Position_both_Overlap_(%),” “Position_Cells_Contact_Area_with_Neighbors_(%),” “Shape_Cells_Area,” “Shape_Cells_Ratio_Width_to_Length,” “Shape_Cells_Roundness,” “Shape_Nuclei_Area,” and “Shape_Nuclei_Ratio_Width_to_Length.”
It has to be noted that both cell painting datasets used two different technologies of cell morphology measurement (Cell Profiler for the Broad Institute and Harmony® for US-EPA), therefore, it is not possible to concatenate the cell morphology read out from both datasets in a unique ML model.
Chemical space
By computing the Mordred descriptors for all the compounds selected after the GRIT score and for the three datasets, we can compare the chemical space of all the compounds. Figure 2 shows the overlap of compounds across datasets. Although only 24 compounds are in the three datasets, we checked the chemical space of all the compounds using the Mordred molecular descriptors. We could observe that the structures of the compounds occupy the same chemical space (Fig. 3). The distribution of mutagenicity also appears to occupy most of the chemical space. This could explain the lower performance of models using structure rather than cellular phenotypic data from cell painting.
Figure 2.
Venn diagram representing the number of compounds of each of the three datasets (broad and both of US-EPA).
Figure 3.
Visualization of the chemical space of molecules from the two datasets (broad & US-EPA).
A comparison of the chemical and phenotypic spaces for each of the two datasets separated is given in Supplementary Fig. SS2. This allows us to see whether, for each dataset separately, there is a visual separation between mutagenic and nonmutagenic compounds, whether using chemical or phenotypic descriptors. The difference in feature names between the two cell painting datasets makes it impossible to compare their phenotypic space jointly.
GRIT score
This metric is particularly valuable for distinguishing compounds with strong, reproducible effects from those with variable or weak responses. High GRIT scores indicate a strong and reproducible phenotypic effect, while low scores suggest variability or lack of a distinctive phenotypic signature. This analysis enabled us to prioritize compounds with robust phenotypic effects in order to create relevant and reliable training and test sets for the prediction models.
The number of compounds available according to the GRIT score threshold is detailed in Table 1. A threshold of one allows the selection of compounds that show significantly different cell morphology read out compared with DMSO controls, while theses read out are rather homogeneous between replicates for the same compounds. For both of the datasets however, this drastically reduces the number of compounds. So the GRIT score threshold was lowered to zero for nonmutagenic compounds, assuming that nonmutagenic compounds have a relatively low impact on cell morphology, but was set at one for mutagenic compounds to have more quality data for model training and testing.
Table 1.
Number of compounds captured in relation to the GRIT score.
| Broad | US-EPA_d7/PAC | US-EPA_d7 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Cpd | M | NM | Cpd | M | NM | Cpd | M | NM | |
| −2.0 | 230 | 66 | 164 | 506 | 133 | 373 | 506 | 133 | 373 |
| −1.0 | 230 | 66 | 164 | 497 | 133 | 364 | 498 | 133 | 365 |
| 0.0 | 137 | 41 | 96 | 346 | 105 | 241 | 345 | 106 | 239 |
| 0.5 | 39 | 18 | 21 | 239 | 85 | 154 | 213 | 74 | 139 |
| 1.0 | 30 | 15 | 15 | 176 | 66 | 110 | 144 | 55 | 89 |
| 1.5 | 26 | 14 | 12 | 149 | 56 | 93 | 110 | 43 | 67 |
| 2.0 | 22 | 11 | 11 | 140 | 53 | 87 | 95 | 39 | 56 |
| 2.5 | 19 | 11 | 8 | 126 | 47 | 79 | 85 | 36 | 49 |
| 3.0 | 17 | 10 | 7 | 113 | 41 | 72 | 74 | 30 | 44 |
| 3.5 | 15 | 9 | 6 | 96 | 33 | 63 | 67 | 26 | 41 |
| 4.0 | 12 | 7 | 5 | 85 | 26 | 59 | 59 | 21 | 38 |
| 4.5 | 10 | 5 | 5 | 77 | 24 | 53 | 56 | 21 | 35 |
| 5.0 | 10 | 5 | 5 | 67 | 22 | 45 | 52 | 21 | 31 |
For a GRIT score ranked between -2.0 and 5.0, the total number of compounds collected (Cpd) and the distribution between Mutagenic (M) and Non Mutagenic (NM) compounds are reported for the 3 datasets (Broad, US-EPA_d7/PAC and US-EPA_d7).
To summarize the details of the compounds by dataset: the Broad data subset used for modeling contains 111 compounds, 15 of which are mutagenic and 96 nonmutagenic; the “US-EPA_d7/PAC” subdataset is made up of 307 compounds (66 mutagenic and 241 nonmutagenic) and the “US-EPA_d7” subdataset is composed of 294 compounds (55 mutagenic and 239 nonmutagenic).
Modelization/prediction
Significant features
For the Broad Institute dataset, a set of 194 significant features have been selected (see Feature selection in section Methods) that discriminate mutagenic from nonmutagenic compounds. Among them, 60 are related to the “cells” compartment, 55 to the “cytoplasm” compartment, and 79 to the “nuclei” compartment.
For the US-EPA datasets, different significant features have been selected depending on whether the dose is chosen according to the compound or not. For the “US-EPA_d7/PAC” dataset, 114 features are significant to distinguish mutagenic and nonmutagenic compounds. They are related to the following cellular compartments: AGP (actin, Golgi, and plasma membrane, 11 features), DNA (37 features), endoplasmic reticulum (18 features), mitochondria (27 features), and RNA (18 features) and three features that do not concern these compartments, but rather various shape and position measurements: “Position_both_Overlap_(%),” “Position_Cells_Contact_Area_with_Neighbors_(%),” and the “Shape_Nuclei_Area.”
Similarly, the “US-EPA_d7” dataset has 136 significant features that distinguish between mutagenic and nonmutagenic compounds. They are related to the following cellular compartments: AGP (actin, Golgi, and plasma membrane, 23 features), DNA (31 features), endoplasmic reticulum (20 features), mitochondria (28 features), and RNA (30 features) and four features that do not concern these compartments, but rather various shape and position measurements: “Position_both_Centroid_Dis,” “Position_both_Overlap_(%),” “Position_Cells_Contact_Area_with_Neighbors_(%),” and the “Shape_Nuclei_Area.”
Therefore, for all three datasets, features related to the nucleus or DNA compartment have the highest counts. Given that mutagenicity directly concerns DNA alterations, it is logical that descriptors associated with DNA and the nucleus are the most relevant.
ClusterMaps
With the previously selected compounds and the significant features mentioned above for the three datasets, we performed the visualizations using the clustermaps. Since clustermaps are unsupervised clustering methods, these visualizations give us an idea of the ability of the selected features in our datasets to distinguish mutagenic from nonmutagenic compounds. This is a preliminary step to ML methods. On this clustermap, clusters of only mutagenic or only nonmutagenic compounds can be observed (see Fig. 4, color bar in blue and green), suggesting that distinct profiles may be associated with mutagenic compounds. The clustermap for the US-EPA dataset can be found in Supplementary Fig. S3.
Figure 4.
Clustermap of compounds based on cell painting features (broad dataset).
Prediction using cell painting dataset
All the models were trained with the oversampled training sets, and the performances evaluated on the test sets (see protocol Fig. 1). We computed initial models with the default parameters to compare with the final models where the parameter optimization was done with nested CV. Table 2 details the model performances evaluation for the three cell painting datasets and the three ML algorithms.
Table 2.
Performances of the predictive models.
| Data type | Model used | Perf. evaluation | Broad dataset | US-EPA_d7/PAC | US-EPA_d7 | |||
|---|---|---|---|---|---|---|---|---|
| Initial | Tuned | Initial | Tuned | Initial | Tuned | |||
| Cell painting | RF | Bal. acc | 0.86 | 0.81 | 0.78 | 0.72 | 0.66 | 0.6 |
| F1 | 0.73 | 0.57 | 0.69 | 0.59 | 0.53 | 0.42 | ||
| Precision | 0.67 | 0.44 | 0.77 | 0.73 | 0.71 | 0.57 | ||
| Recall | 0.8 | 0.8 | 0.62 | 0.5 | 0.42 | 0.33 | ||
| Mcc | 0.68 | 0.5 | 0.6 | 0.5 | 0.38 | 0.22 | ||
| XGB | Bal. acc | 0.85 | 0.91 | 0.76 | 0.79 | 0.52 | 0.52 | |
| F1 | 0.67 | 0.67 | 0.65 | 0.69 | 0.36 | 0.4 | ||
| Precision | 0.57 | 0.5 | 0.67 | 0.69 | 0.4 | 0.38 | ||
| Recall | 0.8 | 1 | 0.62 | 0.69 | 0.33 | 0.42 | ||
| Mcc | 0.61 | 0.64 | 0.53 | 0.58 | 0.05 | 0.04 | ||
| SVM | Bal. acc | 0.86 | 0.76 | 0.81 | 0.67 | 0.58 | 0.58 | |
| F1 | 0.73 | 0.47 | 0.73 | 0.52 | 0.35 | 0.35 | ||
| Precision | 0.67 | 0.33 | 0.79 | 0.53 | 0.6 | 0.6 | ||
| Recall | 0.8 | 0.8 | 0.69 | 0.5 | 0.25 | 0.25 | ||
| Mcc | 0.68 | 0.38 | 0.65 | 0.36 | 0.21 | 0.21 | ||
| Mordred | RF | Bal. acc | 0.67 | 0.68 | 0.71 | 0.73 | 0.76 | 0.76 |
| F1 | 0.5 | 0.46 | 0.57 | 0.62 | 0.55 | 0.55 | ||
| Precision | 1 | 0.43 | 0.67 | 0.8 | 0.46 | 0.46 | ||
| Recall | 0.33 | 0.5 | 0.5 | 0.5 | 0.67 | 0.67 | ||
| Mcc | 0.54 | 0.33 | 0.46 | 0.54 | 0.46 | 0.46 | ||
| XGB | Bal. acc | 0.66 | 0.69 | 0.71 | 0.68 | 0.73 | 0.76 | |
| F1 | 0.43 | 0.5 | 0.58 | 0.52 | 0.53 | 0.55 | ||
| Precision | 0.38 | 0.5 | 0.88 | 0.64 | 0.5 | 0.46 | ||
| Recall | 0.5 | 0.5 | 0.44 | 0.44 | 0.56 | 0.67 | ||
| Mcc | 0.28 | 0.39 | 0.54 | 0.4 | 0.44 | 0.46 | ||
| SVM | Bal. acc | 0.42 | 0.62 | 0.6 | 0.73 | 0.73 | ||
| F1 | 0.12 | 0.47 | 0.38 | 0.4 | 0.53 | |||
| Precision | 0.1 | 0.33 | 0.5 | 0.25 | 0.5 | |||
| Recall | 0.17 | 0.88 | 0.31 | 1 | 0.56 | |||
| Mcc | −0.14 | 0.23 | 0.24 | 0.34 | 0.44 | |||
| Morgan fingerprint | RF | Bal. acc | 0.41 | 0.47 | 0.75 | 0.77 | 0.7 | 0.65 |
| F1 | 0 | 0.15 | 0.62 | 0.65 | 0.45 | 0.37 | ||
| Precision | 0 | 0.14 | 0.62 | 0.61 | 0.38 | 0.28 | ||
| Recall | 0 | 0.17 | 0.62 | 0.69 | 0.56 | 0.56 | ||
| Mcc | −0.19 | −0.05 | 0.49 | 0.52 | 0.34 | 0.23 | ||
| XGB | Bal. acc | 0.41 | 0.45 | 0.7 | 0.58 | 0.64 | 0.61 | |
| F1 | 0 | 0.14 | 0.56 | 0.38 | 0.38 | 0.33 | ||
| Precision | 0 | 0.13 | 0.78 | 0.38 | 0.33 | 0.27 | ||
| Recall | 0 | 0.17 | 0.44 | 0.38 | 0.44 | 0.44 | ||
| Mcc | −0.19 | −0.08 | 0.49 | 0.16 | 0.25 | 0.19 | ||
| SVM | Bal. acc | 0.48 | 0.39 | 0.67 | 0.79 | 0.52 | 0.67 | |
| F1 | 0 | 0 | 0.5 | 0.69 | 0.14 | 0.4 | ||
| Precision | 0 | 0 | 0.75 | 0.69 | 0.2 | 0.31 | ||
| Recall | 0 | 0 | 0.38 | 0.69 | 0.11 | 0.56 | ||
| Mcc | −0.08 | −0.22 | 0.43 | 0.58 | 0.04 | 0.27 | ||
Using 3 data types (Cell painting, Mordred and Morgan fingerprint) and 3 machine learning algorithms (RF, XGB and SVM), 5 performances evaluation were computed i.e., Balanced accuracy (Bal. acc), F1, Precision, Recall and Matthews coefficient correlation (Mcc). The models performance are presented for 3 datasets (Broad, US-EPA_d7/PAC, US-EPA_d7) using the default parameters (initial) and with parameter optimization (Tuned). The closer to 1 are the performance values, the higher is the accuracy.
For the Broad dataset, all three models provide excellent performances (balanced accuracy between 0.76 and 0.91 on test set). Parameter optimization brings better performance for one model (XGB), while it does not improve models for the other two algorithms (RF and SVM). For this dataset, the optimized XGB model perform very well, having the highest balanced accuracy, and a recall of one which indicates the ability to classify nonmutagenic compounds correctly but with a precision of only 0.5, meaning that some nonmutagenic compounds are predicted as positive but the highly unbalanced number of positive/negative compounds reduces the precision. The second best model (RF initial), has a balanced accuracy of 0.86 but the other parameters (recall, precision, and F1) with values above 0.67, notably the recall at 0.8. By comparing the predictions of the two best models for the 33 compounds in the test set, the optimized XGB model misclassifies five compounds, only nonmutagenic as mutagenic: the pimozide, the clioquinol, the tamoxifen, the colchicine, and the paclitaxel. The RF model, with a higher precision only misclassifies two nonmutagenic compounds, the clioquinol and the paclitaxel while misclassifying only one mutagenic compound: the 1,4-dimethyl-6-nitro-9H-carbazole. It is interesting to note that the clioquinol and the paclitaxel are poorly predicted in both models. Although paclitaxel is classified as nonmutagenic in two of the three reference databases used in this study (not tested in Ecvam), it has been reported to induce mutations in the Ames test [37]. While this evidence does not necessarily contradict regulatory classifications, it indicates that under certain experimental conditions, paclitaxel may exhibit mutagenic potential. We compared the most important features identified by the two models. In the case of the optimized XGB model, the feature “Nuclei_Correlation_Correlation_Mito_RNA” emerges as the most dominant, exhibiting a markedly higher importance score (0.58) compared with the second-ranked feature (0.17). Additional features of significance include “Nuclei_Neighbors_SecondClosestDistance_1,” “Cells_Texture_SumEntropy_DNA_10_0,” and “Cytoplasm_Correlation_Costes_DNA_Mito,” the latter of which may reflect underlying biological processes such as DNA fragmentation and mitochondrial heterogeneity [38]. Also noteworthy are “Cytoplasm_RadialDistribution_MeanFrac_AGP_4of4” and “Cells_Correlation_K_DNA_AGP.” With the exception of the fourth feature, all are also ranked among the top 20 most important features (with relatively closer importance values) in the initial RF model, which also demonstrated strong performance (link to the features’ meaning in Materials and methods section).
For the US-EPA datasets, the performances are also higher when a “phenotypic” dose has been selected for some of the compounds (“US-EPA_d7/PAC” compared with “US-EPA_d7”). Among the evaluated models, the SVM initially achieved the highest balanced accuracy of 0.81 on the testing set, outperforming other approaches. However, hyperparameter optimization led to a substantial decrease in its performance, suggesting potential overfitting or a high sensitivity to parameter adjustments. In contrast, the XGBoost model showed a balanced accuracy of 0.76 before optimization, which improved to 0.79 following optimization. This enhancement highlights the model’s ability to generalize more effectively to unseen data. Additionally, the interpretability of XGBoost, combined with its stability after optimization, makes it a more reliable and informative approach for predicting compound mutagenicity based on cell painting data. Consequently, the optimized XGBoost model was selected as the most suitable candidate for this task. As mentioned earlier, the model chosen has a balanced accuracy value of 0.79 and the three scores for F1, precision, and recall are 0.69 and the MCC is 0.58. The five (out of 62) nonmutagenic compounds that are not well classified are: nordihydroguaiaretic acid, 17alpha-Ethinylestradiol, benomyl, curcumin, and triphenyltin hydroxide. Interestingly, according to the ECVAM source, the last two compounds were determined as positive for genotoxicity in the in vitro micronucleus (MN) assay and showed DNA damage in mammalian cells [38–40]. On the other hand, five mutagenic compounds are wrongly classified: 4,4′-methylenebis (N,N-dimethylaniline), 5-azacytidine, chloranil, diuron, and linuron. Most of these compounds are also incorrectly predicted by other (and less efficient) models we have developed.
Concerning the most important features from this model, the importance values of the top 20 features are very close (from 0.016 to 0.037), indicating that there is not one or a few features that are really more important than the others. From the top 20 most important features, one concerns the AGP compartment, two concern the mitochondria, seven are about the endoplasmic reticulum, four concern the RNA (especially located in the nuclei), and six the DNA. A barplot of the top 20 most important features for each model selected as the “best” for each dataset is available as Supplementary Fig. SS4.
The US-EPA_d7 model, on the other hand, has a lower performance. The best model, the XGBoost without optimization, gives a balanced accuracy value of 0.69, but a fairly low recall value (0.44), resulting in a low F1 value too (0.5). This shows that the model is not very efficient at predicting compounds that are actually mutagenic.
These values therefore suggest that selecting the optimal dose for each compound enables the model to better predict the (non)mutagenicity of the compounds. The dose of a compound directly influences the extent of the biological effects observed in the cells, which is fundamental for techniques such as cell painting. If the dose is too low, the signal is likely to be drowned out by noise, as the phenotypic changes will be minimal or even absent. Conversely, if the dose is too high, nonspecific toxic effects may dominate, leading to widespread cellular alterations that mask the more subtle and specific effects associated with mutagenicity. Choosing an optimal dose ensures that (i) compounds generate biological effects that are strong enough to be detected, but not so strong as to induce nonspecific perturbations, and (ii) the effects observed are directly linked to the mutagenic activity of the compound, which reinforces the accuracy of predictive models.
Comparison with structural data
Other prediction models using compound’s structure information from the three previous datasets were set up. This time, molecular descriptors from Mordred or Morgan Fingerprints were used (see Materials and Methods section). Details of performance are given in Table 2. None of the models (Mordred or Morgan Fingerprints) for any of the datasets (Broad or US-EPA_d7/PAC) performs better as the cell painting models. The only exception is the EPA_d7 dataset, for which the models using the Mordred descriptors perform better than those using the cell painting data (balanced accuracy of 0.76 compared with 0.69). The best model for the compounds in the Broad dataset is given by an optimized XGB algorithm (balanced accuracy = 0.69). The Morgan Fingerprint models perform very poorly. The best model for US-EPA_d7/PAC compounds is given by the optimized SVM algorithm on Morgan Fingerprint data (balanced accuracy = 0.79) or an RF optimized with molecular descriptors from Mordred (balanced accuracy = 0.73).
This highlights the potential of phenotypic cell data rather than structural information for predictive models. Nevertheless, since we were able to get good results with the US-EPA dataset, we could consider combining both cell painting data and structural data into one dataset and see if model performance can be improved.
A possible course of action could be to compute 3D Mordred descriptors for the compounds to check for any complementary information.
Discussion
We have developed in silico models that make it possible to use cell painting data to predict the mutagenicity of chemicals fairly reliably. We compared three ML algorithms, RF, SVM, and XGB, and used two cell painting datasets produced by two different technologies and image analysis software (the open-source software CellProfiler for the Broad Institute and Harmony® software at US-EPA). For the US-EPA dataset, chemicals were tested at several doses. A PAC value, indicating the most relevant dose for a compound, is also sometimes available in the dataset. This is a dose high enough for a phenotypic effect to be observed, but below the toxic dose. This enabled us to compare the performance of the models depending on whether we selected the PAC dose or the same dose level for all the compounds.
The best-performing model for the two cell painting sets (Broad and the US-EPA_d7/PAC set for which the PAC was chosen) is produced by the XGB algorithm (with parameter optimization in CV). The most effective model for the US-EPA_d7 dataset, in the absence of PAC dose selection, is also generated by an optimized XGB algorithm using cell painting data but is outperformed by models built with structural data. However, its performance is inferior to those that can be obtained using PAC doses, highlighting the critical importance of selecting the relevant dose for each compound and its substantial influence on the predictive accuracy of the models.
One possible explanation for XGB being the most effective algorithm for the Broad and US-EPA_d7/PAC datasets is that our datasets are quite unbalanced in terms of mutagenic and nonmutagenic compounds, and XGB may handle this aspect more effectively due to its boosting mechanism. A potential direction for future research is to evaluate alternative and more efficient algorithms when confronted with unbalanced datasets, such as the CSBBoost algorithm [41].
Various parameters are taken into account in the development of our models, and these can affect the final performance. For example, increasing the GRIT score threshold to >1 drastically reduces the number of mutagenic compounds used to develop predictive models, limiting the chemical space learnt by the model, and therefore reduces performance (results not shown). Decreasing the threshold, on the other hand, increases the number of chemicals that can be exploited for the model development but the data quality is less consistent and will reduce robustness of the model. The impact of the dose chosen (for the US-EPA dataset, apart from the PAC dose) also affects model performance. So, when possible, it is preferable to select cell painting data for which a phenotypic alteration is observable on a specific concentration. The same applies to the preprocessing parameters, the parameters of the different models, and the number of folds in the CV process.
Since the choice of dose impacts model performance, it is also crucial to understand which morphological features contribute the most to the prediction of mutagenicity. By analyzing these features, we can gain insights into the underlying biological mechanisms captured by our models. About the morphological features of predicted mutagens, we observe that some features contributing the most to the best-performing model’s performance are related to DNA and/or RNA perturbation. Furthermore, these features are primarily located in mitochondria and nucleus. As mutagenicity involves DNA intercalation and DNA transcription products, it gives plausibility to the output of the predictive model.
As stated in the Results section, the most important feature for the Broad dataset is “Nuclei_Correlation_Correlation_Mito_RNA” that we interpret as a measure of the co-localization between mitochondrial and RNA-associated fluorescence within nuclei. Its dominant importance (0.58) may indicate that mitochondrial–nuclear spatial relationships are highly predictive of mutagenic responses. Typically, if a change in co-localization of DNA and mitochondrial is induced by compounds, we could hypothesize that this translates into DNA fragmentation and hence DNA damage. The second important feature is “Nuclei_Neighbors_SecondClosestDistance_1” quantifying the distance between each nucleus and its second-closest neighboring nucleus, reflecting local cell density or spatial organization.
The most important feature of the best-performing model of the US-EPA_d7/PAC dataset is “ER_Ring_Texture_SER_Hole_2_px.” As 7 of the 20 most important features concern the ER, this could mean that the mutagenicity mechanisms of the compounds involve ER stress.
ER stress has been shown to promote oxidative stress through the production of reactive oxygen species (ROS), linking these two cellular responses [42]. Accumulation of ROS is a common feature in many cancer cells, leading to direct DNA damage by increasing mutation rates and acting as secondary messengers in signaling pathways that support oncogenesis [43]. These findings suggest that modification of the ER morphology may play a role in mediating the mutagenic effects of compounds.
Beyond ER-related features, some of the most important features (6 out of 20) are linked to DNA morphology, highlighting another potential mechanism involved in mutagenicity (features that begin with “DNA_Cells_Morph” or “DNA_Nuclei_Morph”). The exact meaning of these features is unclear, but they may reflect intracellular heterogeneity in DNA fluorescence distribution. These features might indicate cell cycle alterations or abnormalities (e.g. cell cycle arrest), as DNA is consistently distributed in G1 and G2 phase, and aggregated in S phase [44]. During transition to G0 phase or even when cells are undergoing apoptosis, DNA distribution shows heterogeneity, which can be visualized by imaging the nucleus [45]. Otherwise, changes in cellular regions in DNA staining channels may also suggest that spots that are distinct from the main nucleus, for example, micronuclei, may be recognized. MN is a well-known marker of genotoxicity and is also closely related to mutagenicity, although the mechanisms are generally different [46]. These features are generated by automated image analysis, so their biological interpretation is indirect, but often reveals phenotypic changes linked to mutagenicity: DNA fragmentation, cell reorganization, altered cell contacts, etc. Their diversity (nucleus, cytoplasm, membrane) suggests that mutagenicity affects several cell compartments. Their diversity (nucleus, cytoplasm, membrane) suggests that mutagenicity affects several cellular compartments.
While these comparisons highlight the strengths of our model, certain limitations remain, particularly regarding the dependency on image analysis software used in cell painting. Cell morphology readouts from Broad Institute and US-EPA are not comparable and integration of chemicals is not feasible at this moment. It could be interesting to harmonize the cell painting technology in order to combine readouts from a larger set of chemicals. It would be beneficial to the accuracy of the model and also to predict new chemicals independently to the technology used.
Although we are not aware of any other mutagenicity prediction model using cell painting, other “mutagenic” models do exist and use other types of data. For example, a ML model using over 4000 chemical’s structure features was created, with a balanced accuracy of 0.65 on a test set [6]. Another deep learning model, DeepAmes, trained on more than 10 k compounds using chemical–physical 1D/2D descriptors and compared with five ML models gives good performance with a balanced accuracy of 0.69 on their test set [47]. Molecular fingerprints are also in the spotlight with a graph CNN (MutagenPred-GCNNs), which predicts quite efficiently the mutagenicity of compounds with an accuracy about 76% for the validation set [48]. The advantage of this tool is that it identifies chemical features of relevance “toxicophores” that are often present in mutagenic compounds. Finally, another team developed a model (named Mol2vec) based on the natural language processing technique, using the Morgan algorithm-derived substructures of 6511 compounds (from the ZINC and ChEMBL databases) as words and the entire compounds as sentences [49]. They obtained good performances since both sensitivity and specificity are equal to 0.8.
We tested the prediction of the compounds in our test set with two free, easy-to-use tools: VEGA (a QSAR-based prediction tool) [22] and the CompTox Chemicals Dashboard v2.5.2 (https://comptox.epa.gov/dashboard/predictions). The QSAR VEGA tool can predict all the compounds in the Broad dataset. Three nonmutagenic compounds are falsely predicted to be mutagenic (nomifensine, resmethrin, pimozide), by VEGA, but the two first are correctly predicted by our model. For mutagenic compounds, two are predicted as nonmutagenic by VEGA (bergapten and etoposide), also correctly predicted by our XGB model. The CompTox Dashboard can predict 30 out of 33 compounds of the Broad dataset. Among them, six are falsely predicted as mutagenic by the tool: nomifensine, pimozide, phosalone, praziquantel, 1,10-phenanthroline, and fenbendazole. Two compounds, the bergapten and the 1,4-dimethyl-6-nitro-9H-carbazole are falsely predicted as nonmutagenic. All of these compounds (except the last mentioned before) are correctly predicted by our two best models. CompTox is able to predict in a correct way three nonmutagenic compounds that our best model has difficulty to predict: colchicine, clioquinol, tamoxifen, and paclitaxel. The bergapten, the nomifensine, and the pimozide seem to be harder to predict by both QSAR tools.
For the US-EPA dataset, the QSAR VEGA tool was able to predict most of the compounds in the US-EPA_d7/PAC dataset (three returned errors). Triphenyltin hydroxide, a nonmutagenic compound, is poorly predicted by both VEGA and our model On the other hand, the linuron, a mutagen, is poorly predicted by both our model and VEGA. The other poorly predicted or nonpredicted compounds (five for VEGA and eight for our model) differ according to the tool used to predict them (VEGA or XGB). VEGA mispredicts the following components: methyl violet, acetochlor, tetramethylthiuram monosulfide, methylene blue (all four mutagenics), and methidathion (nonmutagenic) whereas our XGB model mispredicts: 4,4′-methylenebis(N,N-dimethylaniline), 5-azacytidine, chloranil, diuron (all four mutagenics), curcumin, nordihydroguaiaretic acid, 17alpha-Ethinylestradiol, and benomyl (all four nonmutagenics).
The CompTox dashboard was able to predict all of the compounds from the US-EPA_d7/PAC dataset. But all of the 16 mutagenic compounds are mispredicted by the tool (5 of them are also wrongly predicted by our best model). The only compound predicted as mutagenic by CompTox (methidathion) is actually nonmutagenic. This demonstrates that regardless of the descriptors used, the mutagenicity of particular molecules is more difficult to predict than for others, although our models using cell painting appear to be very effective and promising.
It would appear that the use of cell painting to predict mutagenicity performs well, possibly reflecting a better understanding of the biological activity or toxic effects of molecules. Combining both models (cell painting and QSAR) as a consensus model would be interesting to investigate as the chemical space considered by both technologies are different. QSAR models are more chemical centric instead of cell painting, which is biological centric.
Nevertheless, since we were able to get good results with the US-EPA dataset, we could also consider combining both cell painting data and structural data into one dataset and see if model performance can be improved. A possible course of action could be to compute 3D Mordred descriptors for the compounds to check for any complementary information.
Another perspective to consider is the experimental conditions of cell painting and the annotations of compounds as potential factors contributing to the mispredictions. The Ames test, with its extensive history and comprehensive databases, often exhibits variability in results across different sources [50]. Because mutagenicity is generally regarded as having no threshold, implying that even minimal quantities of a compound can present a carcinogenicity risk, we employed a conservative approach in this study. Three databases were referenced in this study, and each compound was annotated as mutagenic if it yielded a positive result in any test, irrespective of species of strains, exposure doses, criteria for positive judgment, presence or absence of metabolic activation conditions, or whether the test was conducted under good laboratory practice conditions or not. In this study, our annotated mutagens included those that tested positive in the Ames test under metabolic activation conditions and negative under nonmetabolic activation conditions, such as chloranil, diuron, which were poorly predicted as negative. A plausible explanation for these mispredictions is that the U-2 OS cell line, derived from osteosarcoma, did not metabolize the tested compounds into mutagens when cell painting assay was performed, which led to an absence of morphological alterations and a negative test result. Although our current model has already achieved high accuracy, annotating the data with reflecting the experimental conditions of cell painting may lead to the development of an even more accurate model.
From the perspective of in vitro biological assays other than the Ames test, the in vitro MN test, which detects chromosomal damage, is commonly used as one of the standard test batteries for genotoxicity. However, the MN test cannot interpret the mechanism of action of genotoxicity, making it difficult to distinguish between clastogens and aneugens without combination of fluorescence in situ hybridization [51]. Therefore, compounds that test positive in this assay usually require follow-up testing to avoid misleading results, which makes the in vitro MN test alone insufficient for comprehensive genotoxicity assessment [52]. In the cell painting assay, Fig. 4 reveals a cluster of nonmutagenic compounds situated between two clusters of mutagenic compounds. These nonmutagenic compounds are suspected to be aneugens, as they affect microtubules: rotenone [53], colchicine [54], fenbendazole , [55], and podofilox [56]. A recent study has shown that cell painting is able to predict tubulin-targeting compounds [57], suggesting that our dataset here differentiates between potential aneugens and other compounds. Interestingly, our model predicted paclitaxel as mutagenic, whereas it is annotated as nonmutagenic in the reference mutagenicity database used for training. This discrepancy invites closer examination. Notably, a study published in 2013 reported mutagenic effects of paclitaxel in an Ames assay [37]. This suggests that paclitaxel’s mutagenic potential may be context-dependent and possibly underestimated in standard databases, which often consolidate data from heterogeneous protocols or consider only consensus outcomes. From a model perspective, the cell painting feature could reflect phenotypic alterations captured that are consistent with a prediction of mutagenicity.
Therefore, the large group that includes the two mutagenic clusters plus paclitaxel and the intermediate nonmutagenic cluster is thought to represent genotoxicity detectable by both Ames test and the in vitro MN test. In other words, while this study developed a prediction model for mutagenicity, it also suggested the potential to develop a model to predict genotoxicity that can distinguish between clastogens and aneugens. This provides additional weight of evidence to the prediction results for comprehensive chemical risk assessment.
On the subject of clastogens, some of the compounds used the test set of the Broad or the EPA dataset are detected as such by the QSAR Toolbox V4.6 [58] in the “worst case scenario, i.e. when it has at least one positive in vitro result, using all the available databases and ‘carcinogenicity’ & ‘genetic toxicity’” as endpoints. Among those 31 compounds, 13 are also mutagenic with only three wrongly predicted as nonmutagenic (5-azacytidine, chloranil, diuron). Three compounds are also nonmutagenic but wrongly predicted as mutagenic (curcumin, 17alpha-ethinylestradiol, triphenyltin hydroxide).
Finally, it could be interesting to study more complex models such as Deep Learning models (deep neural networks or graph neural networks for example) in order to try to better predict compounds that are poorly predicted by ML models. Another improvement would be to learn the models on more massive datasets, as other tools have done, and the expansion of the use of HTTP techniques should help with this. These future directions could further enhance the predictive power of our approach. Nevertheless, our current findings already demonstrate the potential of cell painting-based ML models for mutagenicity prediction.
Overall, based on the cell morphology readouts, features selection, and parameters optimization, we developed ML models that can contribute to the prediction of mutagenicity of chemicals. We found that XGBoost was the most suitable ML algorithm and provided good predictive performance. Such an approach is complementary to a predictive model based on a chemical’s structure. This study advances the use of high-content imaging approaches for toxicological assessment. It provides evidence that morphological cellular changes captured by cell painting correlate with mutagenic activity, particularly in DNA/RNA structures and mitochondria. Furthermore, it proposes a framework for integrating phenotypic data with traditional structure-based models to enhance predictive accuracy. Importantly, by developing alternatives to traditional in vivo mutagenicity testing, this work supports the 3Rs principle (replacement, reduction, and refinement), potentially reducing the need for animal experiments in toxicological assessments. The next step could be to combine both cell painting and chemical structure embedded features in a deep learning model for enhanced predictive capabilities.
Supplementary Material
Contributor Information
Natacha Cerisier, Université Paris Cité, INSERM U1133, Functional and Adaptive Biology Unit, CNRS UMR 8251, 75013 Paris, France.
Emily Truong, Université Paris Cité, INSERM U1133, Functional and Adaptive Biology Unit, CNRS UMR 8251, 75013 Paris, France.
Taku Watanabe, Scientific Product Assessment Center, Japan Tobacco Inc., 6-2, Umegaoka, Aoba-Ku, Yokohama, Kanagawa 227-8512, Japan.
Taro Oshiro, Scientific Product Assessment Center, Japan Tobacco Inc., 6-2, Umegaoka, Aoba-Ku, Yokohama, Kanagawa 227-8512, Japan.
Tomohiro Takahashi, Scientific Product Assessment Center, Japan Tobacco Inc., 6-2, Umegaoka, Aoba-Ku, Yokohama, Kanagawa 227-8512, Japan.
Shigeaki Ito, Scientific Product Assessment Center, Japan Tobacco Inc., 6-2, Umegaoka, Aoba-Ku, Yokohama, Kanagawa 227-8512, Japan.
Olivier Taboureau, Université Paris Cité, INSERM U1133, Functional and Adaptive Biology Unit, CNRS UMR 8251, 75013 Paris, France.
Conflict of interest statement: T.W., T.O., T.T., and S.I. are employees of Japan Tobacco Inc.
Funding
This work was supported by the project RISK-HUNT3R: RISK assessment of chemicals integrating HUman centric Next generation Testing strategies promoting the 3Rs. RISK-HUNT3R has received funding from the European Union’s Horizon 2020 research and innovation program under grant agreement No 964537 and is part of the ASPIS cluster.
This work reflects only the authors’ views, and the European Commission is not responsible for any use that may be made of the information it contains.
Data availability
The data underlying this article are available in the article and in its online supplementary material.
References
- 1. Benigni R, Bossa C. Mechanisms of chemical carcinogenicity and mutagenicity: a review with implications for predictive toxicology. Chem Rev 2011;111:2507–36. 10.1021/cr100222q [DOI] [PubMed] [Google Scholar]
- 2.Committee on Mutagenicity of Chemicals in Food, Consumer Products and the Environment. Guidelines for the testing of chemicals for carcinogenicity. Committee on carcinogenicity of chemicals in food, consumer products and the environment. Rep Health Soc Subj (Lond) 1982;25:1–30. [PubMed] [Google Scholar]
- 3. Maron DM, Ames BN. Revised methods for the Salmonella mutagenicity test. Mutat Res Mutagen Relat Subj 1983;113:173–215. 10.1016/0165-1161(83)90010-9 [DOI] [PubMed] [Google Scholar]
- 4. Cherkasov A, Muratov EN, Fourches D et al. QSAR modeling: where have you been? Where are you going to? J Med Chem 2014;57:4977–5010. 10.1021/jm4004285 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Muratov EN, Varlamova EV, Artemenko AG et al. Existing and developing approaches for QSAR analysis of mixtures. Mol Inform 2012;31:202–21. 10.1002/minf.201100129 [DOI] [PubMed] [Google Scholar]
- 6. Chu CSM, Simpson JD, O’Neill PM et al. Machine learning – predicting Ames mutagenicity of small molecules. J Mol Graph Model 2021;109:108011. 10.1016/j.jmgm.2021.108011 [DOI] [PubMed] [Google Scholar]
- 7. Chavan S, Scherbak N, Engwall M et al. Predicting chemical-induced liver toxicity using high-content imaging phenotypes and chemical descriptors: a random forest approach. Chem Res Toxicol 2020;33:2261–75. 10.1021/acs.chemrestox.9b00459 [DOI] [PubMed] [Google Scholar]
- 8. Lejal V, Cerisier N, Rouquié D et al. Assessment of drug-induced liver injury through cell morphology and gene expression analysis. Chem Res Toxicol 2023;36:1456–70. 10.1021/acs.chemrestox.2c00381 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Nyffeler J, Willis C, Harris FR et al. Application of cell painting for chemical hazard evaluation in support of screening-level chemical assessments. Toxicol Appl Pharmacol 2023;468:116513. 10.1016/j.taap.2023.116513 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Seal S, Trapotsi MA, Spjuth O et al. Cell painting: a decade of discovery and innovation in cellular imaging. Nat Methods 2025;22:254–68. 10.1038/s41592-024-02528-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Bray MA, Singh S, Han H et al. Cell painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes. Nat Protoc 2016;11:1757–74. 10.1038/nprot.2016.105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Cimini BA, Chandrasekaran SN, Kost-Alimova M et al. Optimizing the cell painting assay for image-based profiling. Nat Protoc 2023;18:1981–2013. 10.1038/s41596-023-00840-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Luo YS, Ying RY, Chen XT et al. Integrating high-throughput phenotypic profiling and transcriptomic analyses to predict the hepatosteatosis effects induced by per- and polyfluoroalkyl substances. J Hazard Mater 2024;469:133891. 10.1016/j.jhazmat.2024.133891 [DOI] [PubMed] [Google Scholar]
- 14. Cerisier N, Dafniet B, Badel A et al. Linking chemicals, genes and morphological perturbations to diseases. Toxicol Appl Pharmacol 2023;461:116407. 10.1016/j.taap.2023.116407 [DOI] [PubMed] [Google Scholar]
- 15. Moshkov N, Becker T, Yang K et al. Predicting compound activity from phenotypic profiles and chemical structures. Nat Commun 2023;14:1967. 10.1038/s41467-023-37570-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Ramezani M, Weisbart E, Bauman J et al. A genome-wide atlas of human cell morphology. Nat Methods 2025;22:621–33. 10.1038/s41592-024-02537-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Thomas JR, Shelton C, Murphy J et al. Enhancing the small-scale screenable biological space beyond known chemogenomics libraries with gray chemical matter—compounds with novel mechanisms from high-throughput screening profiles. ACS Chem Biol 2024;19:938–52. 10.1021/acschembio.3c00737 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Pahl I, Pahl A, Hauk A et al. Assessing biologic/toxicologic effects of extractables from plastic contact materials for advanced therapy manufacturing using cell painting assay and cytotoxicity screening. Sci Rep 2024;14:5933. 10.1038/s41598-024-55952-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Seal S, Carreras-Puigvert J, Singh S et al. From pixels to phenotypes: integrating image-based profiling with cell health data as BioMorph features improves interpretability. Mol Biol Cell 2024;35:mr2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Fredin Haslum J, Lardeau CH, Karlsson J et al. Cell painting-based bioactivity prediction boosts high-throughput screening hit-rates and compound diversity. Nat Commun 2024;15:3470. 10.1038/s41467-024-47171-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Li H, Seada H, Madnick S et al. Machine learning-assisted high-content imaging analysis of 3D MCF7 microtissues for estrogenic effect prediction. Sci Rep 2024;14:2999. 10.1038/s41598-024-53323-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Benfenati E, Manganaro A, Gini G. VEGA-QSAR: AI inside a platform for predictive toxicology. CEUR Workshop Proc 2013;1107. http://ceur-ws.org/Vol-1107/paper8.pdf [Google Scholar]
- 23. Stirling DR, Swain-Bowden MJ, Lucas AM et al. CellProfiler 4: improvements in speed, utility and usability. BMC Bioinformatics 2021;22:433. 10.1186/s12859-021-04344-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Gustafsdottir SM, Ljosa V, Sokolnicki KL et al. Multiplex cytological profiling assay to measure diverse cellular states. PLoS One 2013;8:e80999. 10.1371/journal.pone.0080999 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Richard AM, Judson RS, Houck KA et al. ToxCast chemical landscape: paving the road to 21st century toxicology. Chem Res Toxicol 2016;29:1225–51. 10.1021/acs.chemrestox.6b00135 [DOI] [PubMed] [Google Scholar]
- 26. Serrano E, Chandrasekaran SN, Bunten D et al. Reproducible image-based profiling with Pycytominer. Nat Methods 2025;22:677–80. 10.1038/s41592-025-02611-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Natoli T, Way G, Lu X et al. Broadinstitute/Lincs-Cell-Painting: Full Release of LINCS Cell Painting Dataset. Zenodo, 2021. 10.5281/zenodo.5008187 [Google Scholar]
- 28. Madia F, Corvi R. EURL ECVAM Genotoxicity and Carcinogenicity Consolidated Database of Ames Negative Chemicals [Internet]. Joint Research Centre (JRC): European Commission, 2020. http://data.europa.eu/89h/38701804-bc00-43c1-8af1-fe2d5265e8d7. [Google Scholar]
- 29. Hansen K, Mika S, Schroeter T et al. Benchmark data set for in silico prediction of Ames mutagenicity. J Chem Inf Model 2009;49:2077–81. 10.1021/ci900161g [DOI] [PubMed] [Google Scholar]
- 30. Shinada NK, Koyama N, Ikemori M et al. Optimizing machine-learning models for mutagenicity prediction through better feature selection. Mutagenesis. 2022;37:191–202. 10.1093/mutage/geac010 [DOI] [PubMed] [Google Scholar]
- 31. Maaten L v d, Hinton G. Visualizing data using t-SNE. J Mach Learn Res 2008;9:2579–605. [Google Scholar]
- 32. Pedregosa F, Varoquaux G et al. Scikit-learn: machine learning in python. J. Mach. Learn Res. 2011;12:2825–30. [Google Scholar]
- 33. Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining [Internet]. San Francisco California USA: ACM; 2016, 785–94. 10.1145/2939672.2939785 [DOI] [Google Scholar]
- 34. Moriwaki H, Tian YS, Kawashita N et al. Mordred: a molecular descriptor calculator. J Cheminformatics 2018;10:4. 10.1186/s13321-018-0258-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Morgan HL. The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. J Chem Doc 1965;5:107–13. 10.1021/c160017a018 [DOI] [Google Scholar]
- 36. Capecchi A, Probst D, Reymond JL. One molecular fingerprint to rule them all: drugs, biomolecules, and the metabolome. J Cheminformatics 2020;12:43. 10.1186/s13321-020-00445-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Vieira IL, de Souza DC, da Silva CL et al. In vitro mutagenicity and blood compatibility of paclitaxel and curcumin in poly (DL-lactide-co-glicolide) films. Toxicol in Vitro 2013;27:198–203. 10.1016/j.tiv.2012.10.013 [DOI] [PubMed] [Google Scholar]
- 38. Seal S, Carreras-Puigvert J, Trapotsi MA et al. Integrating cell morphology with gene expression and chemical structure to aid mitochondrial toxicity detection. Commun Biol 2022;5:1–15. 10.1038/s42003-022-03763-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Genetic Toxicology - Mammalian Cell Mutagenicity. Chemical effects in biological systems (CEBS). NTP study ID 174971, Triphenyltin hydroxide. https://cebs.niehs.nih.gov/cebs/genetox/002-02978-0008-0000-6/
- 40. Cao J, Jia L, Zhou HM et al. Mitochondrial and nuclear DNA damage induced by curcumin in human hepatoma G2 cells. Toxicol Sci 2006;91:476–83. 10.1093/toxsci/kfj153 [DOI] [PubMed] [Google Scholar]
- 41. Salehi AR, Khedmati M. A cluster-based SMOTE both-sampling (CSBBoost) ensemble algorithm for classifying imbalanced data. Sci Rep 2024;14:5152. 10.1038/s41598-024-55598-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Chen X, Shi C, He M et al. Endoplasmic reticulum stress: molecular mechanism and therapeutic targets. Signal Transduct Target Ther 2023;8:352. 10.1038/s41392-023-01570-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Ziech D, Franco R, Pappa A et al. Reactive oxygen species (ROS)––induced genetic and epigenetic alterations in human carcinogenesis. Mutat Res Mol Mech Mutagen. 2011;711:167–73. 10.1016/j.mrfmmm.2011.02.015 [DOI] [PubMed] [Google Scholar]
- 44. Ma Y, Kanakousaki K, Buttitta L. How the cell cycle impacts chromatin architecture and influences cell fate. Front Genet 2015;6:19. 10.3389/fgene.2015.00019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Duan DC, Pan G, Liu J et al. Cellular and intravital nucleus imaging by a D-π-A type of red-emitting two-photon fluorescent probe. Anal Chem 2024;96:20425–34. 10.1021/acs.analchem.4c04103 [DOI] [PubMed] [Google Scholar]
- 46. Hayashi M. The micronucleus test—most widely used in vivo genotoxicity test—. Genes Environ 2016;38:18. 10.1186/s41021-016-0044-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Li T, Liu Z, Thakkar S et al. DeepAmes: a deep learning-powered Ames test predictive model with potential for regulatory application. Regul Toxicol Pharmacol 2023;144:105486. 10.1016/j.yrtph.2023.105486 [DOI] [PubMed] [Google Scholar]
- 48. Li S, Zhang L, Feng H et al. MutagenPred-GCNNs: a graph convolutional neural network-based classification model for mutagenicity prediction with data-driven molecular fingerprints. Interdiscip Sci Comput Life Sci 2021;13:25–33. 10.1007/s12539-020-00407-2 [DOI] [PubMed] [Google Scholar]
- 49. Jaeger S, Fulle S, Turk S. Mol2vec: unsupervised machine learning approach with chemical intuition. J Chem Inf Model 2018;58:27–35. 10.1021/acs.jcim.7b00616 [DOI] [PubMed] [Google Scholar]
- 50. Furuhama A, Kitazawa A, Yao J et al. Evaluation of QSAR models for predicting mutagenicity: outcome of the second Ames/QSAR international challenge project. SAR QSAR Environ Res 2023;34:983–1001. 10.1080/1062936X.2023.2284902 [DOI] [PubMed] [Google Scholar]
- 51. Decordier I, Kirsch-Volders M. Fluorescence in situ hybridization (FISH) technique for the micronucleus test. In: Dhawan A, Bajpayee M (eds), Genotoxicity Assessment: Methods and Protocols [Internet]. Totowa, NJ: Humana Press; 2013, 237–44. 10.1007/978-1-62703-529-3_12 [DOI] [PubMed] [Google Scholar]
- 52. Kirkland D, Pfuhler S, Tweats D et al. How to reduce false positive results when undertaking in vitro genotoxicity testing and thus avoid unnecessary follow-up animal tests: report of an ECVAM workshop. Mutat Res Toxicol Environ Mutagen 2007;628:31–55. 10.1016/j.mrgentox.2006.11.008 [DOI] [PubMed] [Google Scholar]
- 53. Johnson GE, Parry EM. Mechanistic investigations of low dose exposures to the genotoxic compounds bisphenol-A and rotenone. Mutat Res Toxicol Environ Mutagen 2008;651:56–63. 10.1016/j.mrgentox.2007.10.019 [DOI] [PubMed] [Google Scholar]
- 54. Parry JM, Sors A. The detection and assessment of the aneugenic potential environmental chemicals: the European Community Aneuploidy Project. Mutat Res Mol Mech Mutagen 1993;287:3–15. 10.1016/0027-5107(93)90140-B [DOI] [PubMed] [Google Scholar]
- 55. Dogra N, Kumar A, Mukhopadhyay T. Fenbendazole acts as a moderate microtubule destabilizing agent and causes cancer cell death by modulating multiple cellular pathways. Sci Rep 2018;8:11926. 10.1038/s41598-018-30158-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Tateno H, Kamiguchi Y, Shimada M et al. Induction of aneuploidy in Chinese hamster oocytes following in vivo treatments with trimethoxybenzoic compounds and their analogues. Mutat Res Mol Mech Mutagen 1995;327:237–46. 10.1016/0027-5107(94)00192-8 [DOI] [PubMed] [Google Scholar]
- 57. Akbarzadeh M, Deipenwisch I, Schoelermann B et al. Morphological profiling by means of the cell painting assay enables identification of tubulin-targeting compounds. Cell Chem Biol 2022;29:1053–1064.e3. 10.1016/j.chembiol.2021.12.009 [DOI] [PubMed] [Google Scholar]
- 58. Dimitrov SD, Diderich R, Sobanski T et al. QSAR toolbox - workflow and major functionalities. SAR QSAR Environ Res 2016;27:203–19. 10.1080/1062936X.2015.1136680 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying this article are available in the article and in its online supplementary material.




