Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Nov 18.
Published in final edited form as: ACS Pharmacol Transl Sci. 2025 Sep 25;8(10):3568–3584. doi: 10.1021/acsptsci.5c00395

Integrated approach of Machine Learning and High Throughput Screening to Identify Chemical Probe Candidates Targeting Aldehyde Dehydrogenases

Adam Yasgar 1,, Sankalp Jain 1,, Marissa Davies 1, Carina Danchik 1, Taylor Niehoff 1, Jing Ran 1, Ganesha Rai 1, Shyh-Ming Yang 1, Anton Simeonov 1,*, Alexey V Zakharov 1,*, Natalia J Martinez 1,*
PMCID: PMC12519302  NIHMSID: NIHMS2114808  PMID: 41098570

Abstract

Selective chemical probes are essential for dissecting biological pathways and advancing drug discovery, yet developing high-quality probes for targets such as the Aldehyde Dehydrogenase (ALDH) family remains challenging. Here, we present a novel integrated approach combining experimental quantitative high-throughput screening (qHTS) with advanced machine learning (ML) and pharmacophore (PH4) modeling to rapidly identify selective inhibitors across multiple ALDH isoforms. We screened ~13K annotated compounds against biochemical and cellular assays. We then utilized the dataset to build ML and PH4 models to virtually screen a larger set of 174,000 compounds to enhance chemical diversity of hits. This approach led to the expansion of chemically diverse isoform-selective inhibitors that are potent in both biochemical and cell-based assays. Validation through cellular target engagement assays further confirmed the selective activity of these compounds leading to the discovery of ALDH1A2, ALDH1A3, ALDH2, and ALDH3A1 chemical probe candidates. Remarkably, this was achieved by employing just a single iteration of quantitative structure activity relationship (QSAR) and PH4 modeling for virtual screening. This combined in vitro and in silico strategy not only enhances the discovery of biologically relevant chemical probe candidates but also significantly expands the chemical diversity accessible for probe development, establishing a new platform for the rapid and resource-efficient identification of chemical probes against the ALDH enzyme family. The dataset generated, including hundreds of compounds thoroughly characterized across a spectrum of assays, is publicly available and can serve as a high-quality training set for future research initiatives and probe development efforts.

Keywords: chemical probes, quantitative high-throughput screening (qHTS), Aldehyde Dehydrogenase (ALDH), Machine Learning (ML), quantitative structure activity relationship (QSAR), Cellular Thermal Shift Assay (CETSA)

Graphical Abstract

graphic file with name nihms-2114808-f0001.jpg


Aldehyde detoxification is essential for maintaining cellular homeostasis. The accumulation of aldehydes, formed via uncatalyzed and catalyzed oxidation of metabolites, signaling molecules, and other species, can inflict significant damage upon cells, including DNA damage, impaired cellular homeostasis, and ultimately cell death1. The aldehyde dehydrogenase (ALDH) superfamily comprises 19 distinct isozymes that play a crucial role in mitigating this threat. ALDHs catalyze the irreversible oxidation of a wide range of endogenous and exogenous aldehydes to their respective carboxylic acid, using NAD(P)+ as a cofactor2, 3. Besides mitochondrial ALDH2, which is involved in alcohol metabolism, the cytosolic ALDH1A (1A) subfamily, composed of the ALDH1A1 (1A1), ALDH1A2 (1A2), and ALDH1A3 (1A3) isoforms, is arguably one of the best studied ALDH subfamilies to date. 1A enzymes are expressed in adult and embryonal tissues and, in addition to aldehyde detoxification, they also catalyze the conversion of the vitamin A metabolite retinal into retinoic acid, a key regulator of gene expression and cell differentiation4, 5. Moreover, increasing evidence suggests that high activity levels of 1A enzymes are linked to cancer pathogenesis, immunomudulation57 and chemoresistance, as these enzymes can metabolize chemotherapeutics8, 9. In addition to the 1A subfamily, other less studied ALDH isoforms may contribute to the pathology of cancer10. For instance, ALDH1B1, ALDH3A1 (3A1), ALDH7A1, and ALDH18A1 have also been shown overexpressed in different tumor types1114. Collectively, these findings position ALDH isozymes as plausible targets for drug development as inhibiting these enzymes may provide a novel avenue to overcome cancer growth and drug resistance15.

However, it remains challenging to dissect the functional role and pathological contribution of individual ALDH isozymes as different tumor types express different levels of each isozyme. One strategy to provide unambiguous insight into the effect of ALDH inhibition is the development of isozyme-selective high-quality chemical probes. Chemical probes are crucial tools that enable the generation of robust and reproducible insights into the cellular function of proteins of interest. Notably, while some chemical probes are also used as drugs, not all possess the requisite properties to confidently validate a biological target16. Certain criteria have previously been established to define high-quality probes, which include: i) potent target activity, ii) potent cell activity, iii) target engagement, and iv) selectivity (>30-fold within a target family)17. Unfortunately, many reported probes in the literature lack one or more of these attributes and thus, cannot be considered as high-quality18. Given the broad distribution of ALDHs in normal tissues, the development of selective inhibitors that can discriminate between isozymes is pivotal for reducing toxicity and enhancing therapeutic efficacy.

The traditional target-based drug discovery approach primarily utilizes high-throughput screening (HTS) of ALDH activity assays complemented by lead optimization and rational design of inhibitors based on available crystal structures or homology models19. However, challenges such as high sequence homology among isozymes, lack of available crystal structures for most ALDHs, low substrates specificity, and the limited availability of purified enzymes for assay development, severely constrain the development of potent and selective ALDH probes15. To date, only a few selective inhibitors such as NCT-505/506 for ALDH1A120 have been validated to meet all probe criteria, underscoring the need for innovative approaches in probe development21.

The recent advances in Artificial Intelligence (AI)/Machine Learning (ML) have enabled the development of QSAR (Quantitative Structure-Activity Relationships) models capable of exploring vast chemical spaces beyond the reach of traditional HTS22. Leveraging these modern in silico techniques, we previously identified novel inhibitors for SARS-CoV-223 and Zika virus proteases24, demonstrating the advantage of combining screening data with predictive modeling. Inspired by these successes, we expanded our methodology to develop a comprehensive drug discovery platform capable of identifying high-quality chemical probe candidates for target families.

In this study, we implemented quantitative high-throughput screening (qHTS) of biochemical and ALDEFLUOR cellular assays and identified a set of cell active compounds that selectively inhibit 1A2, 1A3, ALDH2, and 3A1 isozymes from a relatively small compound library (~13K). To expand our search and harness greater chemical diversity, we integrated in silico screening techniques to examine more extensive libraries (~170K), thereby enabling the swift identification of novel selective chemotypes while optimizing resource use. Finally, to further evaluate compound activity in cells, we employed high-throughput cellular target engagement assays for each isozyme using the SplitLuc system25. Our comprehensive platform not only identifies compounds that selectively engage specific ALDH isozymes in cells but also sets the stage for further medicinal chemistry efforts aimed at developing these into potent, isozyme-specific ALDH probes for proof-of-concept studies. This strategy, depicted in Fig. 1, exemplifies our streamlined approach to drug discovery, significantly reducing the time and resources typically required while maintaining a focus on high-impact therapeutic targets.

Figure 1. Schematic representation of screening strategy.

Figure 1.

Combined approach of qHTS (traditional) and virtual screening leading to the identification of chemical probe candidates. Under each respective step lists the isozymes, ALDH-expressing cell line, or SplitLuc CETSA number of compounds tested.

Results

Enzymatic activity assay to identify selective ALDH inhibitors

Expanding on our previous efforts at developing isozyme-selective inhibitors for the ALDH family20, 26, we sought to identify novel chemotypes that selectively inhibit 1A2, 1A3, ALDH2, and 3A1 isozymes. All three 1A isozymes share >70% sequence homology with each other and are closely related to ALDH2 (68% homology) but less so to the 3A1 (28.5% homology) subfamilies (Fig. S1A). We selected these isozymes based on the commercial availability of enzymatically active recombinant protein for biochemical assay development. In addition to the above panel, we included a biochemical assay of the 1A1 isozyme to further assess selectivity of candidate inhibitors. The biochemical assays utilized either propionaldehyde (1A1, 1A2, 1A3, and ALDH2) or benzaldehyde (3A1) as substrate, and NAD(P)+ as cofactor, following previously described methods20, 27. Assay readouts included monitoring NAD(P) + reduction in a coupled reaction with resorufin (1A1, 1A2, ALDH2 and 3A1) or pro-luciferin (1A3) (Fig. S1B). We utilized a qHTS approach, enabling the profiling and classification of compound activity through detailed concentration-response curves (CRC)28. Hence, the biochemical assays were miniaturized to a 4 μL volume across 1,536-well plates to maximize the use of resources and enhance throughput. In general, substrates and co-factors were tested at or above Km, and reactions run at < 20% conversion (see Methods). A set of 27 reference compounds, including biochemical selective and multi ALDH isozyme inhibitors from the literature (Table S1A), indicated that our assay conditions can recapitulate compound activity and selectivity. Our results confirmed the selectivity profiles of several compounds, particularly validating multiple inhibitors of 1A1 and the specific inhibition of 3A1 by CB2929. For 11 compounds, our assays identified activity against additional isozymes besides the reported one. For example, the reported 1A3-specific Compounds 68 and 6930 also inhibited other 1A family members. Finally, for 6 compounds, we did not detect inhibitory activity for their reported isozyme. For example, the ALDH2 inhibitor CVT-1021627 showed preferential inhibition towards 1A2 rather than ALDH2 in these conditions (Table S1A).

Biochemical assays for 1A2, 1A3, ALDH2, and 3A1 were each interrogated in a traditional screening approach against a relatively small set of 13,419 compounds from the LOPAC1280, NPACT, and NCATS Medicinal Chemistry collections in qHTS format (see Supporting Information and Methods for the description of the library). These libraries contain approved, bioactive, annotated, and structurally diverse compounds, which can facilitate hit identification and in silico screening (see below). Screening outcomes, including assay performance statistics, compound selection, and criteria cutoffs based on curve class, potency, and efficacy are described in detail in the Supporting Information. Briefly, our primary screen identified 2,132 compounds that exhibited inhibition against one or more isozymes (Fig. S1C).

From the initial screening hits, we selected and sourced 2,080 compounds for confirmatory testing in the same biochemical qHTS assays but at a higher dose response density (see Methods and Supporting Information; Table S2). As previously mentioned, we incorporated the 1A1 biochemical assay at this triage point to further assess compound selectivity (Fig. 1). Compounds were also counterscreened to remove any potential artifacts that interfere directly with the readouts rather than inhibiting ALDH activity (see Supporting Information for details). A total of 1,766 (85%) compounds confirmed their primary screening activity, with 1,151 (55%), 1,102 (54%), 726 (35%), 1,031 (50%), and 1,237 (59%) compounds demonstrating inhibition against 1A2, 1A3, ALDH2, 3A1, and 1A1, respectively (Table S2; Fig. 2A). Inspecting further, each isoform had between 12% and 24% of its respective hits exhibiting IC50’s < 10 μM. Then, a hit was deemed selective if it was active towards one isozyme and was either inactive against the other isozymes and/or exhibited a >5-fold IC50 ratio if activity was observed. Applying our selectivity criteria, 133 (6%), 114 (6%), 28 (1%) and 142 (7%) compounds were selective for 1A2, 1A3, ALDH2, and 3A1, respectively (Fig. 2B; Fig. S1D). The remaining compounds had multi-isozyme activity (Table S2). This data suggested that obtaining compounds exclusively selective for ALDH2 was more challenging compared to other isozymes. While our focus was on isoform-selective inhibitors, we also uncovered compounds with broader inhibition profiles. For example, there were candidates for pan-inhibition (Table S2), including ZM-39923 (NCGC00016107)31, Bedaquiline (NCGC00348215), Phenserine (NCGC00163250), NCGC00284018, NCGC00488260, and NCGC00385147, all found to be relatively potent (IC50 < 10 μM) across all isoforms.

Figure 2. Isozyme selectivity landscape of biochemical and ALDEFLUOR hits identified by traditional High-throughput screening.

Figure 2.

A) Number of biochemical (white) and ALDELFLUOR (grey) hits for ALDH1A1/OV90 (●), ALDH1A2/AN3CA (), ALDH1A3/PEO1 (), ALDH2/SKBR3 (), ALDH3A1/OE19 () and combinations. B) Overlap of selective hits in both biochemical (B) and ALEFLUOR (A) assays for each isozyme.

ALDEFLUOR assay to identify cell-active ALDH inhibitors

To explore the cellular activity of ALDH inhibitors and assess the translatability of biochemical hits to a cellular context, we employed our previously developed high-throughput, high-content ALDEFLUOR imaging assay27, 32. Given the ALDFLUOR substrate BAAA is utilized by multiple ALDHs10, 33, 34 we identified cell lines known to predominantly express particular ALDH isozymes: AN3CA (1A2), PEO1 (1A3), SKBR3 (ALDH2), OE19 (3A1), and OV90 (1A1)20, 35. Each cell line was validated for expression and substrate utilization, setting the stage for a high-throughput, high-content imaging assay to assess inhibitors’ cellular efficacy (Figs. S2AS2D and Supporting Information).

We first evaluated the ALDEFLUOR assay against the same reference set of 27 compounds characterized in the biochemical assays. Only 7 compounds, primarily 1A1-specific inhibitors, corroborated their biochemical activity profiles. Interestingly, the compounds designated as 1A3-specific inhibitors, namely Compound 68 and 69, demonstrated broader inhibitory effects, impacting other members of the 1A family despite their purported specificity30. For 9 compounds, our assays partially replicated the reported activity profile, for instance, by identifying activity against additional isozymes besides the reported one. Finally, for 11 compounds, we either did not detect any activity in cell assays or detected activity against a different isozyme (Table S3A).

Next, we applied multiple filters to the 1,766 hits confirmed in the biochemical assays, including potency and structural filters, selecting 810 for testing in the ALDEFLUOR assay against all five cell lines (Fig. 1, see Methods and Supporting Information for detailed selection criteria). This triage step yielded 353 compounds (44%; Table S3C and Fig. 2A) with inhibitory activity against at least one ALDH-expressing cell line: 42 (8%) for AN3CA, 280 (57%) for PEO1, 143 (29%) for SKBR3, 178 (36%) for OE19, and 43 (9%) for OV90. Among these, the potencies were moderate, with IC50 values under 10 μM observed in 9 (2%), 72 (15%), 52 (11%), 67 (14%) and 5 (1%) compounds for AN3CA, PEO1, SKBR3, OE19 and OV90, respectively. In contrast to the biochemical outcome, these results indicated that PEO1(1A3) had the highest hit rate. Applying stringent selectivity filters revealed 8, 103, 21, 45, and 1 compounds with the characteristic to selectively inhibit AN3CA, PEO1, SKBR3, OE19 and OV90 cell lines, respectively, highlighting the efficacy of our approach in bypassing ALDH1A1-selective inhibitors (Figs. 2A and 2B). Altogether, our HTS identified a limited number of compounds (totaling 24 or 0.2% of compounds screened) with consistent selectivity across both biochemical and cellular platforms: 2 for 1A2, 8 for 1A3, 7 for ALDH2, and 7 for 3A1 (Table S4 and Fig. 2B). These findings were consistent with our observations from the reference set, illustrating the challenges of identifying ALDH inhibitors that exhibit congruent selectivity profiles in both biochemical and cellular assays using conventional HTS methods and underscores the prevalence of polypharmacology in the beginning stages of drug discovery36.

Ligand-Based In silico screening to enhance ALDH inhibitor discovery

The low percentage of cell-active (44%) and selective compounds (0.2%) after the first two triage steps in our traditional screening paradigm highlights the need to broaden the chemical diversity of the screening libraries to increase the chance of identifying selective hits. This cannot be efficiently done via traditional screening due to the cost, time- and labor-intensive characteristics of cell assays like ALDEFLUOR. To address this limitation, we implemented a ligand-based virtual screening campaign utilizing Random Forest (RF), Deep Learning Consensus Architecture (DLCA), and pharmacophore (PH4) modeling, applied to an in-house library of ~174,000 compounds. (Fig. 1). These techniques were validated in previous studies for their capacity to efficiently navigate large chemical libraries and identify compounds with potential cellular activity23, 24, 37, 38 39. The novelty of the virtual screening performed here was due to strategically using the cell activity data from the ALDEFLUOR assays, aiming to prioritize compounds that exhibit potent ALDH inhibition in cellular environments, whereas previously the virtual screens were performed on biochemical data sets.

To optimize the training and test datasets for our ML models, we performed an additional confirmatory ALDEFLUOR screen to refine our list of active and inactive compounds, thus enhancing data quality by excluding entries with low-quality CRCs and weak potency (see Supporting Information for details). Over 95% of the selected compounds demonstrated consistent activity across their respective cell lines, affirming the robustness of our experimental design and resulting in a curated list of 452 compounds (Table S5A). Subsequently, using this refined dataset, we developed QSAR models, which typically demonstrate high precision (>50%)40, as detailed in Table S5B. However, the modeling for the AN3CA cell line (1A2) presented challenges due to the limited number of confirmed active compounds (19), restricting the robustness of our model training.

In parallel, we enriched the diversity and precision of our in silico compound selection through PH4 modeling. We selected the top active compounds for different ALDH isozymes: 9 for ALDH1A2, 20 for ALDH1A3, 26 for ALDH3A1, and 29 for ALDH2, and clustered them based on PH4 similarity with distances of 0.4, 0.6, 0.7, and 0.8. This approach generated ligand-based hypotheses, resulting in 8 PH4 hypotheses for ALDH1A2, 22 for ALDH1A3, 28 for ALDH3A1, and 30 for ALDH2, including both merged-features and shared-features PH4’s. Due to computational constraints, we further refined our selection to 2, 10, 7, and 23 PH4 models that targeted the majority (>20%) of active versus inactive compounds for each respective isozyme. The key PH4 features included hydrogen bond acceptors and donors, aromatic rings, and hydrophobic sites. Details on the PH4 models and specific PH4 hypotheses used for virtual screening are available in Table S5C of the Supporting Information, along with supporting files (PH4_Models.zip).

We then applied the above QSAR (RF and DLCA) and PH4 models to virtually screen our in-house library of ~174,000 compounds (see Methods section for details on the library composition). To minimize resources, we limited the number of selected hits to ~240 per cell line, representing two 1,536-well plates in qHTS format for experimental validation (Table S5D). That number also allowed us a statistical chance of observing a hit, even if the hit rate was as low as 1% (>2 compounds/cell line).

Experimental validation of in silico-identified hits

In our experimental validation, we selected 240 compounds for testing against AN3CA and 242 compounds each for the PEO1, SKBR3, and OE19 ALDEFLUOR assays (Fig. 1) (see Supporting Information for details). To minimize resource use, compounds were tested only in the cell line they were predicted to be active. From this group, 258 compounds (27%) exhibited inhibitory activity against their respective ALDH-expressing cell line, with 27 (11%), 105 (43%), 67 (28%), and 59 (24%), for AN3CA, PEO1, SKBR3, and OE19, respectively (Table S5D). These hit rates reflect the ability of our in silico models to mimic the cell-based assay performance observed in the traditional screening approaches.

Accounting for the number of compounds tested using a specific model, the best performing method based on the average hit rate was PH4 combined with DLCA (PH4+DLCA) followed by PH4 combined with RF (PH4+RF), RF, DLCA, and PH4 alone. Reviewing the hit rates for each model against the tested cell line, the models performed best for PEO1, followed by SKBR3, OE19, and AN3CA (Table S5D and Fig. 3A). The combination model PH4+DLCA yielded the best results for PEO1 (1A3) and SKBR3 (ALDH2), with hit rates of 56% (33 actives) and 41% (15 actives) and IC50 ranges of 0.1 – 35.5 μM and 0.4 – 31.6 μM, respectively. For the AN3CA (1A2) cell line, the combination model of PH4+RF yielded the highest hit rate of 26% with an IC50 range of 0.4 – 10 μM, while for the OE19 (3A1) cell line, the single model method RF proved best with a 33% hit rate (13 actives) and IC50’s ranging from 1.3 – 31.6 μM.

Figure 3. In silico screening performance and selectivity landscape of hits.

Figure 3.

Figure 3.

A) Machine Learning (ML) and Pharmacophore (PH4) hit rates against cell lines expressing ALDH isozymes. B) Chemical space distribution of AN3CA (ALDH1A2) actives identified through virtual screening (VS) and traditional high-throughput screening (HTS). A principal component analysis (PCA) plot illustrating the distribution of library compounds (), HTS actives (), and VS actives () in the chemical space of ALDH1A2 inhibitors. The placement of HTS hit NCGC00690141 and VS hit NCGC00601190, highlights the structural diversity among identified actives. The HTS-derived compound is located near previously explored regions, whereas the VS-identified compound occupies a more distant region, suggesting the discovery of novel scaffolds. C) Tanimoto Similarity between VS vs. HTS actives. D) Number of biochemical (white) and ALDELFLUOR (grey) hits for ALDH1A1/OV90 (●), ALDH1A2/AN3CA (), ALDH1A3/PEO1 (), ALDH2/SKBR3 (), ALDH3A1/OE19 () and combinations. E) Overlap of selective hits in both biochemical (B) and ALEFLUOR (A) assays for each isozyme.

Next, we sought to assess whether the QSAR and PH4 models increased the diversity of our cell-based actives. To evaluate this, we analyzed the chemical diversity of the compounds identified as actives in traditional high-throughput screening (HTS) compared to those in the virtual screening (VS) sets (Figs. 3B and Figs. S3AS3C). This analysis involved calculating Tanimoto coefficients among all compound pairs to quantify the variance in molecular structures, where 0.85 is a common cutoff used to determine whether compounds are similar. The comparison of Tanimoto similarity scores between VS and HTS hits highlights a significant structural diversity among the identified compounds. Across all ALDH isozymes, the majority of VS hits fall within the 0.2–0.4 similarity range, indicating only moderate resemblance to the training set (Fig. 3C). Notably, in PEO1(1A3) and SKBR3 (ALDH2), some compounds exhibit very low similarity scores (as low as 0.13–0.20), suggesting that VS identified novel scaffolds that were not well-represented in the training set. Meanwhile, only a small number of VS hits show high similarity (≥0.80) to HTS hits, indicating that while some structurally related compounds were retrieved, the majority of VS hits remain chemically diverse compared to HTS actives (Fig. 3C).

Given the number of cell actives identified and the increased diversity of chemical matter, we restricted our modeling to one iteration to minimize resource use. Mirroring the triage process from the traditional screening campaign, we tested the cell-active hits for enzymatic characterization (see Supporting Information for details). We selected a subset of 160 compounds: 10 from AN3CA, 58 from PEO1, 51 from SKBR3, and 41 from OE19; reflecting the distribution of actives across these cell lines (Fig. 1, Table S5E). This selection spanned all in silico models, with individual contributions from DLCA (30 compounds), RF (35), and PH4 (38), and from their combinations, PH4+DLCA (47) and PH4+RF (10) (Table S5E). Overall, 94 of 160 compounds (59%) exhibited inhibitory activity against at least one ALDH isozyme (Table S5E and Fig. S3D). Applying our selectivity criteria, the hit rates for isoform-selectivity combined with corresponding cell activity were 3 (30%), 5 (9%), 1 (5%), and 1 (2%), for 1A2, 1A3, ALDH2, and 3A1, respectively, for a total of 10 compounds (Table S5E and Fig. 3D; see Supporting Information for details). Profiling these 10 biochemically selective compounds across all cell lines using our ALDEFLUOR assay, only 4 compounds remained selective: 0/5 (0%), 3/5 (60%), 0/3 (0%), and 1/3 (33%) for 1A2, 1A3, ALDH2, and 3A1 isozymes (Table S5F and Fig. 3E; see Supporting Information for details). The remaining compounds were either inactive or displayed activity against multiple isozymes in the biochemical or cellular assays, indicating partial selectivity.

Our efforts to expand the diversity of cell-active chemotypes through in silico methods were successful, achieving a comparable proportion of active compounds to the traditional qHTS approach while utilizing fewer resources. These results demonstrate that after a single iteration the integrated application of QSAR and PH4 modeling on a high-quality, cell-active training set can effectively identify promising candidate inhibitors. Notably, this approach also increased the chemical diversity of identified hits, addressing a key limitation often encountered in traditional qHTS screenings. Finally, these findings highlight the inherent challenges in discovering ALDH inhibitors with both biochemical and cellular selectivity, underscoring the value of combining computational and experimental methodologies for drug discovery.

Target Engagement Assays to confirm binding of ALDH inhibitors

Cellular thermal shift assays (CETSA) can detect compound-target interactions in a cellular environment by quantifying the changes in the thermal stability of target proteins upon compound binding in cells41. We have previously developed a high-throughput CETSA platform that utilizes the split Nano Luciferase (NLuc) approach (SplitLuc CETSA) to characterize target engagement of 1A1 inhibitors in isothermal conditions25. Due to its wide applicability towards the ALDH family, we developed SplitLuc CETSA assays for the ALDH isozymes in this study and established isothermal conditions to test the identified inhibitors in dose-response (Figs. S4AE and Supporting Information).

We first tested the reference compounds in the above system (Table S6A). A systematic test of all compounds against all isozymes was performed for 21 reference compounds. We found that 57% of compounds tested did not show any activity in this assay. Nevertheless, we recapitulated the 1A1 selectivity of NCT505/506 inhibitors. Notably, the reported ALDH2 inhibitor CVT-10216, while broadly active in biochemical assays and inactive in ALDEFLUOR assays, showed selective stabilization of ALDH2 in CETSA. Similarly, the reported 1A1/ALDH2/3A1 inhibitor Isatin, showed selective activity against 3A1 in biochemical assays, was inactive in ALDEFLUOR assays but exhibited selective stabilization of 3A1 in CETSA. These results indicate that target engagement is not detectable for most ALDH inhibitors in the reference set and that once again, it is difficult to identify ALDH inhibitors with fully matching biochemical, ALDEFLUOR, and target engagement selectivity profiles.

In our subsequent target engagement studies, in addition to the 24 (23 sourced) and 4 selective compounds derived from traditional and in silico screens, respectively, we included additional compounds that encompass a variety of activity and selectivity profiles by incorporating less potent, non-selective, and pan inhibitors. This strategic inclusion was designed to validate the robustness of our target characterization approach, providing a clearer understanding of how broadly acting compounds interact with the ALDH isozymes under study in the context of this assay. We evaluated 87 compounds from traditional screening and 25 from virtual screening, along with two mixed inhibitors, Stattic and ZM-39923, totaling 114 compounds tested against all 5 ALDH isozymes (Fig. 4A, Tables S4, S5D, and S6B). Overall, 28 of 114 compounds (25%) exhibited thermal stabilization of at least one isozyme: 14 (12%), 7 (6%), 5 (4%), 14 (12%), and 4 (4%) compounds for 1A2, 1A3, ALDH2, 3A1, and 1A1, respectively (Table 6B; see Supporting Information for further details). Applying selectivity criteria, 6 (5%), 4 (4%), 2 (2%), 8 (7%), and 1 (1%) compounds were selective for 1A2, 1A3, ALDH2, 3A1, and 1A1, respectively (Fig. 4B). Of the 23 selective compounds identified in the traditional screening, 3 were also selective in the corresponding CETSA assays. Similarly, from the 4 selective hits identified in the virtual screen, 2 were also selective in the corresponding CETSA assays. As observed with the reference set, the target engagement assay shows a low hit rate, reflecting the inherent challenges in identifying compounds that stabilize isozymes without prior optimization through medicinal chemistry25, 42.

Figure 4. Target Engagement Assay Performance.

Figure 4.

A) Number of hits in the SplitLuc target engagement assay for ALDH1A1 (●), ALDH1A2 (), ALDH1A3 (), ALDH2 (), ALDH3A1 () and combinations; 86 inactives not shown. B) Clustering hits based on SplitLuc CETSA activity leads to the identification of selective compounds. C) Summary table of ALDH chemical probe candidates indicating potencies in various assays. D) Summary SAR table derived from analogs of the selective 3A1 inhibitor NCGC00274265.

Selective inhibitors as candidates for probe development

Our study highlighted several isoform-selective compounds that stand out as potential chemical probe candidates. These include NCGC00351858 for 1A2, NCGC00373989 for 1A3, NCGC00343742 for ALDH2, NCGC00480746 (MI-192) for 3A1, all demonstrating selectivity in enzymatic, ALDEFLUOR, and CETSA assays (Figs. 4B and 4C; Table S6B). NCGC00351858 is cataloged in PubChem (CID 1491276) with over 634 biological test entries, only 83 (13%) of which demonstrated bioactivity. It has been previously patented for use as anti-cancer treatment43 and as part of a collection of inhibitors against Thioredoxin reductase 1 (TrxR1)44, where it exhibited an IC50 value of 0.75 μM in the biochemical assay. Without any medicinal chemistry optimization, the compound showed an IC50 of 0.4 μM against 1A2. The 1A3 selective hit NCGC00373989 had only 13 biological test results documented in PubChem (CID 43840270), all deemed it inactive. Of note is the substructure of this molecule and its similarity to benzaldehyde, the substrate used in biochemical assays, and thus N,N-diethylamino benzaldehyde (DEAB), a pan-ALDH inhibitor45. Further characterization of this molecule to rule it out simply as a substrate mimetic is needed but outside the scope of this work. For ALDH2, we identified NCGC00343742, with only 8 biological test results documented in PubChem (CID 72710612), only exhibiting weak activity against PIP4Kα. Lastly, for 3A1, we identified the candidate probe NCGC00480746 (CID 56965342; MI-192), which has been reported as a potent and selective HDAC 2/3 inhibitor and potential anti-stroke therapy46.

Limited analog expansion for SAR determination

Structure-Activity Relationship (SAR) studies are critical to further develop screening hits into chemical probe candidates. Moreover, the availability of inactive analogs to serve as controls in proof-of-concept studies is desired for high-quality probes. We performed a limited analog expansion of the 4 selective chemotypes for 1A2, 1A3, ALDH2 and 3A1 shown in Figure 4B and discussed above, by identifying available analogs from our in-house compound collection.

We sourced 10 analogs of the selective 1A2 inhibitor NCGC00351858 (Fig. S5A and Table S7). In view of the enzymatic potency against 1A2, the SAR indicated the nitro group in the R1 position is vital to bring the activity down to low micromolar range. However, selectivity towards the other isozymes was reduced (Table S7). We identified inactive analogs, such as NCGC00039048, with H groups in the R1 position displaying no activity in enzymatic, ALDEFLUOR or CETSA assays (Table S7), which meets the criteria as an inactive control.

The 1A3-selective compound NCGC00373989 and NCGC00373963, both identified in virtual screening (Table S5F), are structurally related (Fig. S5B). We sourced 6 additional analogs around this chemotype (Table S7). Three of these analogs indicated that the position of aldehyde functionality is critical in terms of inhibitory activity in ALDH1A3 enzymatic assay. For instance, moving the aldehyde group from para-position of 2,4-dimethylphenyl ring (NCGC00373989, ALDH1A3 IC50: 0.11 μM) to meta- (NCGC00373987, ALDH1A3 IC50: 14.9 μM) or ortho-position (NCGC00373985, ALDH1A3 IC50: 43.5 μM) significantly decreased potency (Fig. S5B). This latter analog also lacked activity in cellular assays and has the potential to be used as an inactive control (Table S7).

Ten analogs of the ALDH2-selective inhibitor NCGC00343742 were sourced for profiling (Fig. S5C). The additional F group on the phenol ring was well tolerated (NCGC00343738, ALDH2 IC50: 0.9 μM). In contrast, the Cl analog showed a marked loss in potency (NCGC00343746, ALDH2 IC₅₀: 12.7 μM), likely due to the methyl substitution disrupting a key hydrogen bond between the NH of the indolinone core and a nearby cysteine residue Lastly, the more planar α,β-unsaturated double bond motif was found to be vital as the saturated cyclopropyl functionality losing potential π-π interactions with phenylaniline residues further decreased potency (NCGC00343744, ALDH2 IC50: >50 μM). The analog NCGC00485151 had minimal activity in biochemical assays and was inactive in cellular assay, thus representing a good candidate for inactive control (Fig. S5C and Table S7).

For the 3A1 selective compound NCGC00480746, we were only able to identify one additional analog, preventing any meaningful SAR exploration. Since a large portion of target engagement activity was observed for 3A1 (Fig. 4B), we searched for compounds that exhibited selective stabilization of 3A1, selective inhibition of 3A1 enzyme, but partial selectivity towards OE19 in ALDEFLUOR assays (Table S4). This pattern was observed for several compounds, including NCGC00274265, NCGC00371130 (Vilazodone), and NCGC00371121 (SKLB610). NCGC00274265 represents a novel chemotype for which we were able to source 27 compounds with Tanimoto similarities ranging from 0.7 to 0.96. The inhibitory activity of selected compounds in 3A1 enzymatic assay were compiled in Fig. 4D. The SAR revealed that adding 3-OH substitution (e.g., NCGC00019067) was fairly tolerated with 2-OH (NCGC00344854) or 4-OH (NCGC00344479) diminished potency significantly. Moreover, the R1 methyl substitution is also critical since NCGC00319044 (R1= H) weakly inhibited 3A1 activity (IC50= 45.8 μM). The H-bonding donor capability of 3-OH was found to be important as the analog with 3-OMe (e.g., NCGC00274245) was completely inactive. In addition, the R3 substitution seemed in favor of electron donating groups (e.g., OMe) than electron withdrawing groups, such as F, CF3, and OCF3. This observation was consistent with the reduced potency produced by the replacement of the phenyl group with a slightly electron deficient pyridine ring (NCGC00344482, IC50= 4.6 μM). Finally, the methyl ketone was found to be vital as missing the carbonyl group (NCGC00343528) or bearing phenyl ketone (NCGC00344775) both showed significant loss of potency. Upon testing analogs in target engagement and ALDEFLUOR assays, active analogs showed the same selectivity pattern as NCGC00274265, suggesting a medicinal chemistry effort could focus on improving the selectivity in ALDEFLUOR assays (Table S7). Several inactive analogs, such as NCGC00344494, were also identified (Table S7).

Favorable physicochemical properties, such as solubility, permeability, and stability, are desired in high-quality chemical probes. To understand the potential for further development, we assessed the in vitro ADME properties of the selective probe candidates and analogs sourced for limited SAR exploration (Table S8). Specifically, we assessed rat liver microsomal stability (RLM), kinetic solubility, and permeability (Parallel Artificial Membrane Permeability Assay or PAMPA). Most selective probe candidates exhibited good to moderate solubility and PAMPA permeability. However, their RLM stability was poor, indicating a focus area of further optimization for medicinal chemistry programs.

Binding hypothesis of identified inhibitors on ALDH isozymes

To elucidate the binding interactions of inhibitors with ALDH isozymes, we performed molecular docking studies using MOE 2019.147. We analyzed the four selective inhibitors identified: NCGC00351858 for 1A2, NCGC00373989for 1A3, NCGC00343742 for ALDH2, and NCGC00480746 for 3A1. The crystal structures of these isozymes were retrieved from the Protein Data Bank (PDB) with the following IDs: 1A2 (PDB ID: 6ALJ), 1A3 (PDB ID: 7A6Q), 3A1 (PDB ID: 4L2O), and ALDH2 (PDB ID: 4KWG). The substrate binding pockets of these isozymes exhibit significant overlap and share high sequence similarity (Fig. 5A). Since these selected PDB structures contained inhibitors bound within their catalytic (substrate) binding sites, the binding sites were defined based on the regions surrounding the crystal ligand. To assess the reliability of the docking protocol, we first re-docked the crystallographic ligands for each isoform. The Root Mean Square Deviation (RMSD) between crystal and re-docked poses was 2.9012 Å (1A2), 1.9843 Å (1A3), 0.5123 Å (3A1), and 1.0945 Å (ALDH2), confirming acceptable agreement across isoforms (see Fig. S6).

Figure 5. Docking.

Figure 5.

A) Protein structure and binding pocket alignment among the four ALDH isozymes, color-coded as follows: orange for ALDH1A2, blue for ALDH1A3, purple for ALDH3A1, and pink for ALDH2. This alignment highlights the overlapping and unique features of each isozyme’s binding pocket, offering insights into potential isozyme-specific interactions. (B-E) Docking poses of inhibitors with their respective ligands, showing blue for ALDH1A2 with NCGC00351858, purple for ALDH1A3 with NCGC00373949, pink for ALDH3A1 with NCGC00480746 and orange for ALDH2 with NCGC00343742.

For 1A2, the docking of NCGC00351858 revealed interactions with hydrophobic residues Leu477, Met495, Trp195, Ala479, Val138, Leu191, Phe314, and Phe188, with Phe314 and Phe188 engaging in aromatic hydrogen bond (h-bond) and π-π interactions (Fig. 5B). In 1A3, NCGC00373989 binds with hydrophobic residues such as Cys313, Cys314, Phe182, Ile132, Leu471, Ala473, Phe131, and Phe308, while polar residues like Arg139, Tyr472, Glu135, Asn469, and Gly136 indicate potential hydrogen bonding, with Asn469 also participating in aromatic h-bond interactions (Fig. 5C). For 3A1, NCGC00480746’s binding involves hydrophobic residues Val244, Ile391, Trp233, Ile394, Val392, and Met237, and further interactions are provided by aromatic residues Tyr115 and Trp233 through aromatic h-bond interactions, along with Phe401 showing π-π interactions (Fig. 5D). In ALDH2, NCGC00343742 interacts with hydrophobic residues Met174, Phe170, Phe401, Phe459, Phe296, Ala304, Phe465, Pro167, Cys301, Cys302, and Cys303, with Thr244 and Asn169 suggesting potential hydrogen bonding and Cys302 forming hydrogen bonds (Fig. 5E). These molecular docking studies highlight the critical role of hydrophobic interactions in binding the inhibitors within their respective ALDH isozymes’ binding pockets. The presence of aromatic and polar residues enhances binding through additional interactions such as hydrogen bonds and π-π interactions.

Discussion

Chemical probes are essential for exploring the role of target proteins in biological systems, diseases, and therapeutic interventions. Over the past two decades, our center, in collaboration with academic, governmental, and private entities, has been instrumental in advancing the accessibility of chemical probes for early-stage drug discovery4853. While significant progress has been achieved, the ambitious goals set by initiatives like Target 2035 indicate the current pace of probe development is insufficient54. It is evident that innovative approaches are necessary to expedite the discovery of probes for each protein.

In this study, we have implemented a multi-faceted platform that integrates qHTS of biochemical and cell-based activity assays, augmented by virtual screening methods and SplitLuc CETSA to target a panel of ALDH enzymes for probe development. This integrated approach facilitated the identification of four isozyme-selective chemotypes, which serve as foundational workflow for the development of selective chemical probes. Remarkably, this was achieved by physically screening fewer than 15,000 compounds and employing just a single iteration of QSAR and PH4 modeling for virtual screening. Unlike HTS, which is often biased towards known chemotypes, virtual screening enables the identification of novel and structurally distinct compounds. The broader distribution of in silico hits across lower Tanimoto similarity scores compared to HTS hits reinforces the idea that virtual screening is capturing more diverse chemical space, expanding the scope of potential ALDH inhibitors. These results highlight the advantage of computational approaches in discovering previously unexplored chemical matter, which is critical for early-stage drug discovery and scaffold hopping strategies.

While the single iteration of virtual screening managed to increase the chemical diversity of cell active ALDH inhibitors, we did not observe a significant increase in the number of hits when compared to the traditional screening method. However, it is important to note that the virtual screen experimentally interrogated a very small set of compounds. The number of hits was also low due to our application of stringent potency and efficacy criteria, emphasizing the trade-off between sensitivity and specificity in virtual screening. Overall, our platform identified a significant number of chemically diverse cell active compounds, ultimately leading to selective chemical probe candidates.

A key takeaway from this study, including the evaluation of reference compounds, is the challenge of aligning the compound’s activity profiles across biochemical, cellular, and target engagement assays. This could be due to numerous factors, including the choice of substrates used in biochemical and ALDEFLUOR assays, differences between isolated enzyme vs. enzyme within cellular context, off-target effects, active efflux and intracellular metabolization of compounds, differential assay sensitivity, strict activity and selectivity cutoffs applied, compound binding not leading to thermal stabilization, among others. Moreover, some of these factors could be variable across the different cell lines used. Of note, not all target-compound interactions are amenable to thermal shift characterization, and other target engagement strategies might be necessary to determine compound binding to ALDHs25, 55. Nevertheless, these challenges underscore the necessity of transparently sharing all available data to enable a comprehensive assessment during probe development. These results further highlight the need to develop truly selective inhibitors. While the primary goal was to identify selective probe candidates, the emergence of several non-selective or partially selective compounds throughout the triage process can provide alternative avenues into the modulation of biological pathways, enriching our understanding on the ALDH isoforms studied. The generated dataset not only offers extensive opportunities for further exploration of the ALDH isoforms in our study but also serves as a robust resource for advancing ML applications. To our knowledge, this collection represents one of the most expansive/comprehensive publicly available datasets for an enzyme family, encompassing not only biochemical data but also the challenging-to-acquire cell-active data. Our work here contributes hundreds of compounds, thoroughly characterized across a spectrum of assays that can serve as a high-quality training set for future research initiatives and probe development efforts.

While the identified chemotypes are selective towards the interrogated isozymes, their activity against other ALDHs not studied here, remains a potential. Future efforts should focus on developing assays for the remaining isozymes. In addition, identification of cell lines with unique isozyme expression, either endogenous or engineered, will be crucial. Regarding in silico approaches, future efforts will contain multiple iterations and the interrogation of larger libraries (i.e. Enamine22). While a detailed molecular-level understanding of binding mechanisms is beyond the scope of this study, leveraging insights into specific interactions within each binding pocket—such as hydrophobic contacts, hydrogen bonds, and π-π interactions—can aid in designing inhibitors that target unique structural features of each isozyme.

In conclusion, our innovative approach leverages both high-throughput screening and advanced computational models to streamline the probe discovery process, enhancing the efficiency of identifying potential therapeutics and robust probe candidates. This platform is not only applicable to our current library but is also designed to be adaptable to other chemical libraries and potentially to additional ALDH isozymes as suitable reagents become accessible. In addition, we also provide extensive publicly available datasets on an enzyme family, combining high-quality biochemical and cell-based data. This resource is poised to significantly advance the application of artificial intelligence in drug discovery, facilitating the development of more sophisticated models and supporting the broader scientific community in their research endeavors.

Materials and Methods

Compounds.

Libraries screened include the commercially available LOPAC collection (1,280 compounds) and two NCATS libraries, the NPACT library (5,099 compounds), and the NCATS Chemistry collection (7,448 compounds). The NCATS Pharmacologically Active Chemical Toolbox (NPACT) is a library of annotated compounds that provide information on phenotypes, biological pathways and cellular processes54. The NCATS Chemistry collection is a library consisting of internally synthesized compounds. For a full list of compound structures, see PubChem Assay Identifier (AID) 1508620. Removing overlapping structures, 13,419 unique compounds were tested. For confirmatory assays, all compounds were initially sourced from the National Center for Advancing Translational Sciences (NCATS)/National Institutes of Health (NIH). Select compounds were subjected to quality control by LC/UV, LC/MS, or High-resolution MS, with all compounds exhibiting >90% purity by peak area or m/z unless otherwise noted. NCATS in-house libraries virtually screened include: 1) the PubChem collection comprising drug-like compounds from a retired Pharma screening collection that contains a diversity of novel small molecules, with an emphasis on medicinal chemistry-tractable54;2) the Artificial Intelligence Driven (AID) Library (6,995 compounds) containing a subset of compounds selected from the Enamine Targeted Libraries collection; 3) the Genesis collection consisting of novel chemotypes including diverse compounds and targeted scaffolds56.

Enzymes.

Human ALDH1A1, ALDH1A2 and ALDH3A1 recombinant enzymes were purchased from R&D Systems (Minneapolis, MN). Human ALDH1A3 recombinant enzyme was purchased from Sino Biological (Wayne, PA). Human ALDH2 recombinant enzyme was purchased from Abcam Cambridge, MA).

Biochemical activity assays.

All enzyme concentrations are nominal. When possible, substrate concentrations were kept above Km to bias the assay for noncompetitive inhibitors. The inhibitory activity of compounds against ALDHs was measured according to protocols described previously and below20, 26, 27. For primary screens, we used the NCATS online robotic system in which compound libraries were formatted to qHTS and transferred via pintool (23 nL; Wako)57 for a final concentration range of 91 nM to 50 μM. For confirmatory screens, compounds were transferred either via pintool (23 nL; Wako) or acoustic droplet ejection (20 nL; Echo 655 Labcyte, Beckman Coulter). Compounds were formatted as either 10 mM, 1:3, 11-point interplate dilutions (final concentration range of 0.8 nM to 50 μM) or in a modified qHTS format, where each sample underwent interplate serial dilutions (1:2, 12-points) on 384-well plates, totaling 12 concentrations (final concentration range of 49 nM to 50 μM). The samples were then stamped into different quadrants of three 1536-well plates, with the first plate containing the top four concentrations, the second plate the next four concentrations, and the third plate the last four concentrations. Briefly, 3 μL of enzyme (final concentrations were [nM] 40, 25, 30, 50, and 5 nM for 1A1, 1A2, 1A3, ALDH2, and 3A1, respectively) or assay buffer (100 mM HEPES pH 7.5 with 0.01% Tween 20) were dispensed (BioRAPTR, Beckman Coulter) into a 1,536-well solid-bottom black (or white for ALDH1A3) plate (Greiner Bio One). Twenty or 40 nL of compounds (final concentration range 842 pM to 100 μM) or control (DEAB, WIN 18446 or ZM-39923 final concentration range 1.52 nM to 50 μM) were transferred via acoustic droplet ejection. Samples were incubated (room temperature, protected from light) for 15 minutes followed by addition of 1 μL of cofactor and substrate mix. The cofactor NAD+ was used at 1 mM for ALDH1A1, ALDH1A2, ALDH1A3, and ALDH2 (Km’s of 100 μM58, 20 μM59, 52.8 μM60, and 7.4 mM61, respectively) and 1 mM NADP+ for ALDH3A1 (Km of 260 μM62). The substrate propionaldehyde was used at 100 μM for ALDH1A1, ALDH1A2, ALDH1A3, and ALDH2 (Km’s of 21 μM63, ~50 – 61 μM3, 4, 64, 324 μM65, and 2.4 μM63, 64, 66, 67) and benzaldehyde at 200 μM for ALDH3A1 (Km of 280 μM45, 68, 69). ALDH1A3 substrate mixture was prepared with the Promega NADH-Glo Kit Reductase and Reductase Substrate (final concentrations of 0.5X, each) prepared in Promega Luciferase Detection Reagent. For ALDH1A1, ALDH1A2, ALDH2 and ALDH3A1, we used the coupling reagents Resazurin and Diaphorase (final concentrations of 100 μM and 100 μg/mL (or 0.7 U/mL)) to right-shift the detection by monitoring the production of resorufin.70 Plates were centrifuged at 1,000 rpm (164g) for 15 seconds, then read on a ViewLux High-throughput CCD imager (PerkinElmer) equipped with standard Rhodamine optics (525 nm excitation, 598 nm emission) for 30 minutes (< 20% conversion), 15 minutes (< 20% conversion), 30 or 60 minutes (< 50% conversion), and 10 minutes (< 20% conversion) for ALDH1A1, ALDH1A2, ALDH2, and ALDH3A1, respectively. For ALDH1A3, we used kinetic mode on a ViewLux High-throughput CCD imager equipped with standard Luminescence optics (clear filter, 20 second exposure) for 15 minutes (< 20% conversion). The change in fluorescence or luminescence intensity over the respective reaction period was normalized against no-inhibitor or no-enzyme for neutral (DMSO) or positive controls, respectively.

Enzymatic activity counterscreens.

To account for artifacts due to intrinsic compound fluorescence (in the above spectrum) or luminescence (due to inhibitors interacting with the NADH-Glo kit), compounds were also tested in a resorufin and luciferase counterscreen assay. Briefly, 3 μL of assay buffer were dispensed into a 1,536-well solid-bottom black (resorufin assay) or white (NADH-Glo assay) plate. Twenty or 40 nL of compounds (in mixed qHTS format; final concentration range 842 pM to 100 μM) were transferred via acoustic droplet ejection. For the resorufin assay, samples were incubated (room temperature, protected from light) for 15 minutes followed by a 1 μL substrate mixture of NAD+, NADH, and propionaldehyde (final concentrations of 800 μM, 200 μM and 100 μM, respectively) or mixture of NAD+ and propionaldehyde (final concentrations of 1 mM and 100 μM, respectively), in addition to the coupling reagents described previously. For the NADH-Glo counterscreen assay, samples were incubated (room temperature, protected from light) for 15 minutes followed by a 1 μL substrate mixture of NAD+, NADH, and propionaldehyde (final concentrations of 800 μM, 200 μM and 100 μM, respectively) or NAD+ (final concentrations of 1 mM), along with the Promega NADH-Glo Kit Reductase and Reductase Substrate (final concentrations of 0.5X, each) prepared in Promega Luciferase Detection Reagent. Plates were centrifuged at 1,000 rpm (164g) for 15 seconds, then read on a ViewLux imager equipped with standard Rhodamine optics (525 nm excitation, 598 nm emission) for the resorufin counterscreen or Luminescence optics (clear filter, 20 second exposure) for the luminescence counterscreen. The change in fluorescence or luminescence intensity over the respective reaction period was normalized against no-inhibitor and no-enzyme controls. Fluorescence or luminescence intensity was normalized against NADH or no-NADH for neutral (DMSO) or positive controls, respectively.

Cell lines and culture conditions.

HEK293, SKBR3, AN3CA, and OV90 cells were obtained from America Type Culture Collection, (ATCC, Manassas, VA; #CRL-1573, #HTB-30, #HTB-111, #CRL-11732, respectively), PEO1, PEO4, and OE19 cells were obtained from Sigma Aldrich (MilliporeSigma, Rockville, MD; #10032308–1VL, #10032309–1VL and #96071721–1VL, respectively). PEO1, PEO4, and OE19 were cultured in RPMI 1640 (Life Technologies), supplemented with 2mM L-Glutamine (Life Technologies), 10% HyClone fetal bovine serum (FBS, GE Healthcare) and 100 U/mL penicillin and 100 μg/mL streptomycin (referred to as 1% Pen/Strep, Life Technologies). SKBR3 were cultured in McCoy’s 5A (GIBCO) supplemented with 10% FBS and 1% Pen/Step. AN3CA were cultured in EMEM (ATCC) supplemented with 10% FBS and 1% Pen/Strep. OV90 were cultured in MCDB 105 (Cell Applications) and Medium 199 (ThermoFisher) supplemented with 15% FBS, 1.85 g/L sodium bicarbonate (ThermoFisher), and 1% Pen/Strep. HEK293 were grown in DMEM (ThermoFisher), supplemented with 10% FBS and 1% Pen/Strep. All cell lines were maintained at 37°C, 5% CO2, 85% RH, and routinely tested for mycoplasma contamination.

ALDH expression determination via Western Blot.

Whole cell lysates were prepared using RIPA buffer (Cell Signaling Technologies, Danvers, MA) and 1X Halt protease inhibitor cocktail (ThermoFisher). Cell lysates (10 μL) were quantified using the Pierce BCA Protein Assay Kit and read for absorbance (562 nm) on an Tecan infinite M1000Pro. Cell lysate solutions ranging from 15 to 30 μg of protein, or 25 ng of indicated recombinant ALDH, were prepared for gel-electrophoresis and western blot analysis. Samples (25 μL) were mixed with 25 μL of 2X sample loading buffer (NuPAGE® LDS Sample Buffer and Sample Reducing Agent) and transferred into 0.2 mL PCR tubes, heated for 10 min at 90°C (Veriti, Applied Biosystems), followed by centrifugation at 2,000 × g (4°C) for 4 minutes. Samples were then separated on a 15-lane 1.5 mm NuPAGE Novex 4–12% Bis-Tris gel in 1X MOPS SDS running buffer with 1X Antioxidant (ThermoFisher) for 60 minutes at 150V, then transferred to a Nitrocellulose membrane using an iBlot 2 Dry Blotting System (voltage and time settings of 20V for 2 min, 23V for 4 min, and 25V for the remaining 2 min; Life Technologies). Membranes were blocked for 1 hour with 5% milk (nonfat powdered) in PBST (phosphate-buffered saline, pH 7.4, with 0.05% Tween 20) before being incubated at 4°C overnight with rabbit monoclonal anti-ALDH1A1 (Cell Signaling Technologies catalog #12035S), rabbit monoclonal anti-ALDH1A2 (Abcam catalog #ab156019), rabbit polyclonal anti-ALDH1A3 (Abcam catalog #ab129815), rabbit polyclonal anti-ALDH2 (Cell Signaling Technologies catalog#18818S) or rabbit monoclonal anti-ALDH3A1 (Abcam catalog #ab129022) at 1:1,000 dilution, and 1:20,000 of mouse monoclonal anti-GAPDH (Sigma catalog#G8796) in 5% milk PBST. Blots were washed three times in 5% milk PBST and incubated with 1:10,000 anti-mouse or anti-rabbit HRP linked IgG (Cell Signaling Technologies) for 1 hour at room temperature. Blots were washed 3X with PBST and imaged with SuperSignal West Dura Extended Duration Substrate (Thermofisher) on a ChemiDoc Imaging System (Bio-Rad) for luminescence. Blot quantification was performed using ImageQuant TL v8.1 (GE Healthcare) or Image Lab v5.2.1 (Bio-Rad).

Cellular ALDEFLUOR activity assays.

ALDEFLUOR assays were performed as before27. Briefly, cells (5 μL) were filtered (40 μm cell strainer) then dispensed into black, optical quality, clear bottom, TC treated 1,536-well plates (Aurora Microplates) at a density of 1,000 (OV90, PEO1, SKBR3), 1,500 (OE19), or 2,000 (AN3CA) cells/well using a Multidrop Combi dispenser (ThermoFisher) and incubated overnight (37°C, 5% CO2, 85% RH). Media was subsequently removed by centrifuging plates upside down using a plate adaptor to collect media. Five μL/well of a solution of 500 nM (OV90, AN3CA, PEO1, SKBR3) or 1,000 nM (OE19) BAAA substrate (STEMCELL Technologies) and 0.5 nM Hoechst 33342 (ThermoFisher) in ALDEFLUOR buffer (STEMCELL Technologies) was dispensed onto cells using a Multidrop Combi followed by immediate acoustic transfer of 20 nL compound or control solutions using an Echo 655 acoustic dispenser. The neutral and positive controls were DMSO (0.5% final) and DEAB (final concentration 40 μM). Unless otherwise noted, all compounds were assayed as 10- or 12-point dilutions (the latter in mixed qHTS format) spanning a final concentration range of 19.5 nM to 50 μM. Cells were incubated for 1 hour at 37°C to allow the conversion of BAAA into BAA. Supernatant was subsequently removed by centrifugation as described above, then ALDEFLUOR buffer (3 or 5 μL/well) was dispensed by Multidrop Combi before imaging. Images were captured on an Opera Phenix (Perkin Elmer) widefield automated microscope using standard DAPI (390/18x, 432/48m) and FITC (475/28x, 525/48m) filter sets. Images were analyzed using Columbus image analysis system (Perkin Elmer).

Generation of 86b-tagged constructs.

The ALDH2, 1A2, 1A3, and 3A1 open reading frames were cloned as in frame fusions to either the N-term or C-term 86b tag (GSVSGWRLFKKISGS) by PCR amplifying the coding region with InFusion compatible oligonucleotides and ligating them into the acceptor pcDNA3.1–86b N-terminal or C-terminal backbone using BamHI/EcoRI or NheI/BamHI sites, respectively, as described before25. Primers used and source of DNA template are described in Table S9. All fusion constructs were verified by sequencing. The ALDH1A1 86b fusion constructs were previously described25.

Tmelt determination.

Determination of Tmelt was performed as previously described with some modifications. Briefly, HEK293 cells were transfected with indicated 86b-tagged construct using a reverse transfection procedure and plated onto 6-well dishes. Cells were harvested after 24 h and resuspended at 5 × 105 cells/mL in CETSA buffer (DPBS with CaCl2 and MgCl2 containing 1g/L glucose and 1X Halt Protease inhibitor cocktail, ThermoFisher), and 0.5% DMSO. Cells were aliquoted to PCR strips at 30 μL per tube and incubated at 37 °C for 1 h. Samples were then heated at the indicated temperature for 3.5 min using a pre-heated thermal cycler equipped with temperature gradient and then allowed to equilibrate to room temperature. 5 μL/well in triplicate were subsequently transferred onto white, 384-well solid bottom plates (Greiner Bio One) using a multichannel pipette. 2.5 μL/well of Nano-Glo HiBiT Lytic reagent (Promega) were added using a multichannel pipette. Samples were incubated for 15 minutes, centrifuged at 1,000 rpm for 30 seconds and analyzed for luminescence intensity using a ViewLux plate reader.

Isothermal SplitLuc CETSA assay.

SplitLuc CETSA was performed in isothermal conditions as described with some variations25. Briefly, HEK293 cells were transfected using a reverse transfection procedure. After 24 hours, cells were harvested and resuspended in CETSA buffer at a density of 5 × 105 cells/mL. Cells were dispensed (5 μL cell/well) into 1,536-well white plates (Aurora, cyclic olefin polymer) using a Multidrop Combi (Thermo Scientific). Twenty nL/well of compounds or controls (neutral control DMSO or positive control WIN1846 for1A2 and 1A3, and 2, NCT-505 for 1A1, or ZM-39923 for 3A1, at final concentration of 20 μM) were transferred using an Echo 550 acoustic dispenser. Compounds were tested at 11-point doses ranging from 39.8 μM to 0.67 nM. Plates were incubated for 1 hour at 37 °C and subsequently heated at 58°C (1A1, 1A2, and 1A3) or 46°C (ALDH2 and 3A1) for 9 minutes using a custom heating block. Plates were cooled down to room temperature and 3 μL/well of Nano-Glo HiBiT Lytic reagent was dispensed using a Multidrop Combi. Plates were incubated for 15 minutes, centrifuged at 1,000 rpm for 30 seconds and analyzed for luminescence intensity using a ViewLux plate reader.

Ligand-Based Virtual Screening (Machine Learning and Pharmacophore Modeling.

To expand the search space beyond traditional biochemical screening, we implemented a virtual screening approach relying exclusively on ligand-based methods. We applied Random Forest (RF) and Deep Learning Consensus Architecture (DLCA) QSAR modeling, combined with pharmacophore (PH4) modeling, to an internal library of ~174,000 compounds. These approaches were used to select candidates for experimental testing, thereby increasing the chemical diversity of the screened set. Below, we explain the steps involved in this process.

Data curation.

The dataset was curated according to a protocol established by Fourches et al7173. The following steps were undertaken: (i) inorganic compounds were removed based on their chemical formulas using MOE 2019.01 software (Chemical Computing Group); (ii) salts and compounds containing metals and/or rare or special atoms were excluded; (iii) chemical structures were standardized using Francis Atkinson’s standardizer (https://github.com/flatkinson/standardiser); and (iv) duplicates and permanently charged compounds were eliminated using MOE 2019.01.

Compound labeling.

To accurately classify compounds in our dataset, we employed a rigorous set of criteria based on established cellular parameters. Compounds were deemed active if they met the following conditions: a negative Curve Class Value (CCV), an AC50<30 μM, and efficacy<-50%74. Compounds were classified as inactive if they had a CCV of −4. Applying these criteria, we obtained the following active/inactive ratios: 19/70 for 1A2, 79/60 for 1A3, 98/72 for ALDH2, and 154/45 for 3A1. The detailed breakdown of these classifications is provided in Table S5A.

Descriptor calculation.

For all datasets, three distinct sets of descriptors were computed using RDKit (https://www.rdkit.org/):

  1. RDKit Descriptors: A comprehensive set of 119 descriptors based on the two-dimensional structure of the molecules.

  2. Morgan Fingerprints: Circular fingerprints generated with a length of 1024 bits, capturing structural features and their connectivity.

  3. Avalon Fingerprints: Another set of 1024-bit fingerprints, designed to represent molecular structure and properties efficiently.

Training and test set selection.

To ensure robust model training and validation, we implemented a stratified random sampling approach for splitting the data. From each class (active and inactive), 70% of the compounds were randomly selected to form the training set. The remaining 30% of compounds were designated as the test set. This methodology ensures that both the training and test sets are representative of the original data distribution, thereby enhancing the reliability and generalizability of our predictive models.

Virtual screening libraries.

We performed virtual screening using internal libraries, comprising approximately 174,000 compounds (see above). Compound libraries were curated following the same protocol outlined in the Data Curation section. During the screening process, the curated compounds were evaluated against the predictive model. Each compound was assigned a predicted activity score, which is indicative of its likelihood of being active against the corresponding isozyme. The compounds were then rank-ordered based on these scores, enabling us to prioritize those with the highest predicted activity for further investigation.

Machine Learning approaches.

For model development, we employed a combination of Random Forest (RF) and Deep Learning Consensus Architecture (DLCA). The DLCA integrates consensus and multitask deep learning methodologies to generate large-scale Quantitative Structure-Activity Relationship (QSAR) models75, 76.

Random Forest.

Random Forest (RF) is an ensemble learning method comprising multiple decision trees77. In this study, we used the knife implementation of RF. The number of trees was set to 100, a value within the optimal range of 64–128 trees, as suggested by prior research78, 79. Increasing the number of trees beyond this range typically does not enhance model performance. Due to the inherent robustness of the RF algorithm, no additional parameter optimization was performed80. This approach leverages the strength of RF in handling various datasets without the need for extensive tuning, thus ensuring reliable and consistent performance across different tasks.

Deep Learning Consensus Architecture (DLCA).

In this study, we employed a DLCA model that integrates both descriptor-based and descriptor-free models. The descriptor-based models utilized three distinct types of fingerprints—Morgan, Avalon, and AtomPair—as well as RDKit descriptors. The descriptor-free model was developed using SMILES notation and a convolutional neural network (CNN) architecture. This CNN architecture incorporated 1D convolutional layers and GlobalMax pooling layers, followed by hidden dense and output layers81, 82. Detailed architecture and training parameters are provided in Table S10. This hybrid approach leverages the strengths of both descriptor-based and descriptor-free methods, ensuring a robust and comprehensive modeling strategy that capitalizes on diverse molecular representations. The combination of these models within the DLCA framework aims to maximize predictive accuracy and reliability in the analysis of chemical data75.

Pharmacophore-Based screening.

In this study, LBP model generation, refinement, and VS were performed using LigandScout 4.4 Advanced (Inte:Ligand GmbH). Conformer libraries for both LBP modeling and VS were generated with i:Con (LigandScout)83, limiting conformers to a maximum of 200 per compound.

To design LBP models, actives from the training set were clustered based on pharmacophore similarity with cluster distances of 0.4, 0.6, 0.7, and 0.8. Merged-features pharmacophore (MFP) and shared-features pharmacophore (SFP) models were generated for each cluster, incorporating features from selected cluster members23. Effective pharmacophore models not only predict the activity of known actives but also identify active molecules from a vast pool of inactives. To select optimal screening models, we evaluated them against the complete dataset (training and test sets combined), calculating the percentage of actives and inactives that satisfied the pharmacophore features. Models that enriched for active compounds by 20% compared to inactives were chosen for final VS. Screening employed the iscreen module with default settings and a maximum of two omitted features.

Consensus models.

In this study, we used a consensus approach that utilizes the consensus of the predictions between the three different QSAR methods to identify active compounds against various ALDH isozymes. This strategy enhances the accuracy and reliability of our predictions by integrating multiple model outputs to achieve more robust and consistent results.

Following the identification of candidate compounds through QSAR and pharmacophore modeling, we next performed molecular docking studies to investigate their binding modes within ALDH isoforms.

Molecular docking.

To further elucidate the binding interactions of the identified inhibitors with ALDH isoforms, molecular docking studies were performed using MOE 2019.1. The inhibitors studied were NCGC00351858 for ALDH1A2, NCGC00373989 for ALDH1A3, NCGC00343742 for ALDH2, and NCGC00480746 for ALDH3A1. The crystal structures of the respective isozymes were retrieved from the Protein Data Bank (PDB) and prepared for docking using the Structure Preparation module of MOE 2019.1. The PDB IDs for the structures were: ALDH1A2 (PDB ID: 6ALJ)59, ALDH1A3 (PDB ID: 7A6Q)84, ALDH3A1 (PDB ID:4L2O)85, and ALDH2 (PDB ID: 4KWG)86. During preparation, crystallographic water molecules were removed, and protonation states of the amino acid residues were assigned at pH 7.0 using the Protonate 3D tool. Missing atoms and side chains were added, and the entire protein structure was subjected to energy minimization using the Amber10 force field to relieve any steric clashes and optimize the geometry. This ensured that the protein structures were suitable for subsequent docking studies.. The selected PDB structures contained inhibitors bound within their catalytic (substrate) binding sites, which provided a reliable basis for defining the active sites for our docking. Thus, the active sites were defined based on the regions surrounding the crystal ligand. As a quality control measure, the docking protocol was first validated by re-docking the crystal ligands for each isoform, yielding RMSD values of approximately 0.5–2.9 Å across the four ALDH isoforms, supporting the reliability of the docking procedure for subsequent binding mode analyses. Ligand structures were prepared using MOE 2019.1 prior to docking. The initial 2D structures of the ligands were converted to 3D using the Builder tool. Protonation states were assigned at physiological pH (7.0) using the Wash function. Conformational analysis was conducted to generate low-energy conformers, and energy minimization was performed using the MMFF94x force field to optimize the geometries. The prepared ligands were then saved in a suitable format for docking simulations.

Docking was performed using the triangle matcher as the placement algorithm, with the London dG scoring function and GBVI/WSA dG for refinement. This protocol ensured accurate and reliable interaction predictions between the inhibitors and the target proteins.

qHTS data analysis and statistics.

Data from each assay were normalized plate-wise to corresponding intra-plate controls as noted above. Controls were used for the calculation of the Z’ factor, a measure of assay quality control87. Concentration-response curves (CRCs) were fitted and classified as previously described, categorized into four classes: complete response curves (class 1), partial curves (class 2), single point actives (class 3), and inactives (class 4)28, 74, 8890. All CRCs were fitted as previously described and IC50 values were calculated using in-house software or GraphPad Prism (sigmoidal dose-response variable slope). Minimum significant ratio (MSR)91, a statistical parameter that characterizes the reproducibility of potency estimates from in vitro concentration-response (CRC) assays, was used to assess the performance of our intraplate controls. The chemical structures were standardized using the LyChI (Layered Chemical Identifier) program (version 20141028, https://github.com/ncats/lychi)92 using normalized charge. Hit selection criteria were aggregated for duplicate structures using LyChi-3 provided by the NCATS Resolver. This was all done within the Palantir Technologies Foundry Platform (Washington, DC), which is configured to ingest all HTS results generated at NCATS and harmonized this data with other sources such as ChEMBL and OrthoMCL. We used Spotfire (TIBCO) to perform clustering analysis. For activity clustering, we again used the UPGMA clustering method, Euclidean for distance measure, an ordering weight of average value, with a pruning line set to 1.07. All qHTS screening results are publicly available at PubChem (https://pubchem.ncbi.nlm.nih.gov/source/NCGC).

Supplementary Material

PH4_models
Supporting Information
Supporting Tables

Enzymatic assay, ALDH expression in various cell lines, In silico Screen Principal Component Analysis plots, SplitLuc CETSA Assay Optimization, and SAR tractability (PDF)

Enzymatic assay screening results, Cell activity assay screening results, Virtual screening results for both biochemical and cell-based assays, Target Engagement screening results, and Assay performance (XLSX)

SYNOPSIS.

Our study highlighted several isoform-selective compounds that stand out as potential chemical probe candidates. These include NCGC00351858 for 1A2, NCGC00373989 for 1A3, NCGC00343742 for ALDH2, and NCGC00480746 (MI-192) for 3A1, all demonstrating selectivity in enzymatic, ALDEFLUOR, and CETSA assays.

Acknowledgement

We thank NCATS’ compound management, analytical chemistry, ADME, automation, building and support teams. We thank Dvir Blivis and Ty Voss for assistance with automated ALDEFLUOR imaging analysis. This research was supported by the Intramural Research Program of the National Institutes of Health (NIH). The contributions of the NIH authors were made as part of their official duties as NIH federal employees, are in compliance with agency policy requirements, and are considered Works of the United States Government. However, the findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.

Abbreviations

AI

Artificial Intelligence

ALDH

Aldehyde Dehydrogenase

CETSA

Cellular Thermal Shift Assay

CRC

Concentration Response Curve

DCLA

Deep Learning Consensus Architecture

HTS

High-Throughput Screening

LOPAC

Library of Pharmacologically Active Compounds

ML

Machine Learning

NAD(P)+

Nicotinamide Adenine Dinucleotide Phosphate

NLuc

NanoLuciferase

NPACT

NCATS Pharmacologically Active Chemical Toolbox

PAMPA

Parallel Artificial Membrane Permeability

PH4

Pharmacophore

QSAR

Quantitative Structure-Activity Relationship

qHTS

quantitative High-Throughput Screening

RLM

Rat Liver Microsomal

RF

Random Forest

RMSD

Root Mean Square Deviation

SplitLuc

Split Luciferase

VS

Virtual Screening

References

  • 1.LoPachin RM, Gavin T. Molecular mechanisms of aldehyde toxicity: a chemical perspective. Chem Res Toxicol. Jul 21 2014;27(7):1081–91. doi: 10.1021/tx5001046 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Jackson B, Brocker C, Thompson DC, Black W, Vasiliou K, Nebert DW, Vasiliou V. Update on the aldehyde dehydrogenase gene (ALDH) superfamily. Hum Genomics. May 2011;5(4):283–303. doi: 10.1186/1479-7364-5-4-283 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Marchitti SA, Brocker C, Stagos D, Vasiliou V. Non-P450 aldehyde oxidizing enzymes: the aldehyde dehydrogenase superfamily. Expert Opin Drug Metab Toxicol. Jun 2008;4(6):697–720. doi: 10.1517/17425255.4.6.697 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Bchini R, Vasiliou V, Branlant G, Talfournier F, Rahuel-Clermont S. Retinoic acid biosynthesis catalyzed by retinal dehydrogenases relies on a rate-limiting conformational transition associated with substrate recognition. Chem Biol Interact. Feb 25 2013;202(1–3):78–84. doi: 10.1016/j.cbi.2012.11.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Tomita H, Tanaka K, Tanaka T, Hara A. Aldehyde dehydrogenase 1A1 in stem cells and cancer. Oncotarget. Mar 8 2016;7(10):11018–32. doi: 10.18632/oncotarget.6920 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Toledo-Guzman ME, Hernandez MI, Gomez-Gallegos AA, Ortiz-Sanchez E. ALDH as a Stem Cell Marker in Solid Tumors. Curr Stem Cell Res Ther. 2019;14(5):375–388. doi: 10.2174/1574888X13666180810120012 [DOI] [PubMed] [Google Scholar]
  • 7.Yu P, Cao S, Yang SM, Rai G, Martinez NJ, Yasgar A, Zakharov AV, Simeonov A, Molina Arocho WA, Lobel GP, Mohei H, Scott AL, Zhai L, Furth EE, Simon MC, Haldar M. RALDH1 Inhibition Shows Immunotherapeutic Efficacy in Hepatocellular Carcinoma. Cancer Immunol Res. Feb 2 2024;12(2):180–194. doi: 10.1158/2326-6066.CIR-22-1023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Januchowski R, Wojtowicz K, Zabel M. The role of aldehyde dehydrogenase (ALDH) in cancer drug resistance. Biomed Pharmacother. Sep 2013;67(7):669–80. doi: 10.1016/j.biopha.2013.04.005 [DOI] [PubMed] [Google Scholar]
  • 9.Huang CP, Tsai MF, Chang TH, Tang WC, Chen SY, Lai HH, Lin TY, Yang JC, Yang PC, Shih JY, Lin SB. ALDH-positive lung cancer stem cells confer resistance to epidermal growth factor receptor tyrosine kinase inhibitors. Cancer Lett. Jan 1 2013;328(1):144–51. doi: 10.1016/j.canlet.2012.08.021 [DOI] [PubMed] [Google Scholar]
  • 10.Zhou L, Sheng D, Wang D, Ma W, Deng Q, Deng L, Liu S. Identification of cancer-type specific expression patterns for active aldehyde dehydrogenase (ALDH) isoforms in ALDEFLUOR assay. Cell Biol Toxicol. Apr 2019;35(2):161–177. doi: 10.1007/s10565-018-9444-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.van den Hoogen C, van der Horst G, Cheung H, Buijs JT, Pelger RC, van der Pluijm G. The aldehyde dehydrogenase enzyme 7A1 is functionally involved in prostate cancer bone metastasis. Clin Exp Metastasis. Oct 2011;28(7):615–25. doi: 10.1007/s10585-011-9395-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Singh S, Arcaroli J, Chen Y, Thompson DC, Messersmith W, Jimeno A, Vasiliou V. ALDH1B1 Is Crucial for Colon Tumorigenesis by Modulating Wnt/beta-Catenin, Notch and PI3K/Akt Signaling Pathways. PLoS One. 2015;10(5):e0121648. doi: 10.1371/journal.pone.0121648 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Terzuoli E, Bellan C, Aversa S, Ciccone V, Morbidelli L, Giachetti A, Donnini S, Ziche M. ALDH3A1 Overexpression in Melanoma and Lung Tumors Drives Cancer Stem Cell Expansion, Impairing Immune Surveillance through Enhanced PD-L1 Output. Cancers (Basel). Dec 6 2019;11(12)doi: 10.3390/cancers11121963 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kardos GR, Wastyk HC, Robertson GP. Disruption of Proline Synthesis in Melanoma Inhibits Protein Production Mediated by the GCN2 Pathway. Mol Cancer Res. Oct 2015;13(10):1408–20. doi: 10.1158/1541-7786.MCR-15-0048 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dinavahi SS, Bazewicz CG, Gowda R, Robertson GP. Aldehyde Dehydrogenase Inhibitors for Cancer Therapeutics. Trends Pharmacol Sci. Oct 2019;40(10):774–789. doi: 10.1016/j.tips.2019.08.002 [DOI] [PubMed] [Google Scholar]
  • 16.Hartung IV, Rudolph J, Mader MM, Mulder MPC, Workman P. Expanding Chemical Probe Space: Quality Criteria for Covalent and Degrader Probes. J Med Chem. Jul 27 2023;66(14):9297–9312. doi: 10.1021/acs.jmedchem.3c00550 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Muller S, Ackloo S, Arrowsmith CH, Bauser M, Baryza JL, Blagg J, Bottcher J, Bountra C, Brown PJ, Bunnage ME, Carter AJ, Damerell D, Dotsch V, Drewry DH, Edwards AM, Edwards J, Elkins JM, Fischer C, Frye SV, Gollner A, Grimshaw CE, A IJ, Hanke T, Hartung IV, Hitchcock S, Howe T, Hughes TV, Laufer S, Li VM, Liras S, Marsden BD, Matsui H, Mathias J, O’Hagan RC, Owen DR, Pande V, Rauh D, Rosenberg SH, Roth BL, Schneider NS, Scholten C, Singh Saikatendu K, Simeonov A, Takizawa M, Tse C, Thompson PR, Treiber DK, Viana AY, Wells CI, Willson TM, Zuercher WJ, Knapp S, Mueller-Fahrnow A. Donated chemical probes for open science. Elife. Apr 20 2018;7doi: 10.7554/eLife.34311 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Knapp S, Muller S. Improving data quality in chemical biology. Nat Chem Biol. Nov 2023;19(11):1301–1302. doi: 10.1038/s41589-023-01449-5 [DOI] [PubMed] [Google Scholar]
  • 19.Gorgulla C, Boeszoermenyi A, Wang ZF, Fischer PD, Coote PW, Padmanabha Das KM, Malets YS, Radchenko DS, Moroz YS, Scott DA, Fackeldey K, Hoffmann M, Iavniuk I, Wagner G, Arthanari H. An open-source drug discovery platform enables ultra-large virtual screens. Nature. Apr 2020;580(7805):663–668. doi: 10.1038/s41586-020-2117-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yang SM, Martinez NJ, Yasgar A, Danchik C, Johansson C, Wang Y, Baljinnyam B, Wang AQ, Xu X, Shah P, Cheff D, Wang XS, Roth J, Lal-Nag M, Dunford JE, Oppermann U, Vasiliou V, Simeonov A, Jadhav A, Maloney DJ. Discovery of Orally Bioavailable, Quinoline-Based Aldehyde Dehydrogenase 1A1 (ALDH1A1) Inhibitors with Potent Cellular Activity. J Med Chem. Jun 14 2018;61(11):4883–4903. doi: 10.1021/acs.jmedchem.8b00270 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Sanfelice D, Antolin AA, Crisp A, Chen Y, Bellenie B, Brennan PE, Edwards A, Muller S, Al-Lazikani B, Workman P. The Chemical Probes Portal - 2024: update on this public resource to support best-practice selection and use of small molecules in biomedical research. Nucleic Acids Res. Jan 6 2025;53(D1):D1663–D1669. doi: 10.1093/nar/gkae1062 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Bedart C, Shimokura G, West FG, Wood TE, Batey RA, Irwin JJ, Schapira M. The Pan-Canadian Chemical Library: A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening. Sci Data. Jun 6 2024;11(1):597. doi: 10.1038/s41597-024-03443-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Jain S, Talley DC, Baljinnyam B, Choe J, Hanson Q, Zhu W, Xu M, Chen CZ, Zheng W, Hu X, Shen M, Rai G, Hall MD, Simeonov A, Zakharov AV. Hybrid In Silico Approach Reveals Novel Inhibitors of Multiple SARS-CoV-2 Variants. ACS Pharmacol Transl Sci. Oct 8 2021;4(5):1675–1688. doi: 10.1021/acsptsci.1c00176 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Abrams RPM, Yasgar A, Teramoto T, Lee MH, Dorjsuren D, Eastman RT, Malik N, Zakharov AV, Li W, Bachani M, Brimacombe K, Steiner JP, Hall MD, Balasubramanian A, Jadhav A, Padmanabhan R, Simeonov A, Nath A. Therapeutic candidates for the Zika virus identified by a high-throughput screen for Zika protease inhibitors. Proc Natl Acad Sci U S A. Dec 8 2020;117(49):31365–31375. doi: 10.1073/pnas.2005463117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Martinez NJ, Asawa RR, Cyr MG, Zakharov A, Urban DJ, Roth JS, Wallgren E, Klumpp-Thomas C, Coussens NP, Rai G, Yang SM, Hall MD, Marugan JJ, Simeonov A, Henderson MJ. A widely-applicable high-throughput cellular thermal shift assay (CETSA) using split Nano Luciferase. Sci Rep. Jun 21 2018;8(1):9472. doi: 10.1038/s41598-018-27834-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Yang SM, Yasgar A, Miller B, Lal-Nag M, Brimacombe K, Hu X, Sun H, Wang A, Xu X, Nguyen K, Oppermann U, Ferrer M, Vasiliou V, Simeonov A, Jadhav A, Maloney DJ. Discovery of NCT-501, a Potent and Selective Theophylline-Based Inhibitor of Aldehyde Dehydrogenase 1A1 (ALDH1A1). J Med Chem. Aug 13 2015;58(15):5967–78. doi: 10.1021/acs.jmedchem.5b00577 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yasgar A, Titus SA, Wang Y, Danchik C, Yang SM, Vasiliou V, Jadhav A, Maloney DJ, Simeonov A, Martinez NJ. A High-Content Assay Enables the Automated Screening and Identification of Small Molecules with Specific ALDH1A1-Inhibitory Activity. PLoS One. 2017;12(1):e0170937. doi: 10.1371/journal.pone.0170937 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Inglese J, Auld DS, Jadhav A, Johnson RL, Simeonov A, Yasgar A, Zheng W, Austin CP. Quantitative high-throughput screening: a titration-based approach that efficiently identifies biological activities in large chemical libraries. Proc Natl Acad Sci U S A. Aug 1 2006;103(31):11473–8. doi: 10.1073/pnas.0604348103 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Parajuli B, Georgiadis TM, Fishel ML, Hurley TD. Development of selective inhibitors for human aldehyde dehydrogenase 3A1 (ALDH3A1) for the enhancement of cyclophosphamide cytotoxicity. Chembiochem. Mar 21 2014;15(5):701–12. doi: 10.1002/cbic.201300625 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Huddle BC, Grimley E, Chtcherbinine M, Buchman CD, Takahashi C, Debnath B, McGonigal SC, Mao S, Li S, Felton J, Pan S, Wen B, Sun D, Neamati N, Buckanovich RJ, Hurley TD, Larsen SD. Development of 2,5-dihydro-4H-pyrazolo[3,4-d]pyrimidin-4-one inhibitors of aldehyde dehydrogenase 1A (ALDH1A) as potential adjuncts to ovarian cancer chemotherapy. Eur J Med Chem. Feb 5 2021;211:113060. doi: 10.1016/j.ejmech.2020.113060 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Chen Y, Sakamuru S, Huang R, Reese DH, Xia M. Identification of compounds that modulate retinol signaling using a cell-based qHTS assay. Toxicol In Vitro. Apr 2016;32:287–96. doi: 10.1016/j.tiv.2016.01.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Storms RW, Trujillo AP, Springer JB, Shah L, Colvin OM, Ludeman SM, Smith C. Isolation of primitive human hematopoietic progenitors on the basis of aldehyde dehydrogenase activity. Proc Natl Acad Sci U S A. Aug 3 1999;96(16):9118–23. doi: 10.1073/pnas.96.16.9118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Marcato P, Dean CA, Pan D, Araslanova R, Gillis M, Joshi M, Helyer L, Pan L, Leidal A, Gujar S, Giacomantonio CA, Lee PW. Aldehyde dehydrogenase activity of breast cancer stem cells is primarily due to isoform ALDH1A3 and its expression is predictive of metastasis. Stem Cells. Jan 2011;29(1):32–45. doi: 10.1002/stem.563 [DOI] [PubMed] [Google Scholar]
  • 34.Moreb JS, Ucar D, Han S, Amory JK, Goldstein AS, Ostmark B, Chang LJ. The enzymatic activity of human aldehyde dehydrogenases 1A2 and 2 (ALDH1A2 and ALDH2) is detected by Aldefluor, inhibited by diethylaminobenzaldehyde and has significant effects on cell proliferation and drug resistance. Chem Biol Interact. Jan 5 2012;195(1):52–60. doi: 10.1016/j.cbi.2011.10.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Uhlen M, Oksvold P, Fagerberg L, Lundberg E, Jonasson K, Forsberg M, Zwahlen M, Kampf C, Wester K, Hober S, Wernerus H, Bjorling L, Ponten F. Towards a knowledge-based Human Protein Atlas. Nat Biotechnol. Dec 2010;28(12):1248–50. doi: 10.1038/nbt1210-1248 [DOI] [PubMed] [Google Scholar]
  • 36.Manen-Freixa L, Antolin AA. Polypharmacology prediction: the long road toward comprehensively anticipating small-molecule selectivity to de-risk drug discovery. Expert Opin Drug Discov. Sep 2024;19(9):1043–1069. doi: 10.1080/17460441.2024.2376643 [DOI] [PubMed] [Google Scholar]
  • 37.Jain S, Yasgar A, Dalal A, Davies M, Nilova A, Martinez N, Simeonov A, Rai G, Zakharov A. AI-driven drug discovery: identification and optimization of ALDH3A1 selective inhibitors with nanomolar activity. 2024;
  • 38.Pourmousa M, Jain S, Barnaeva E, Jin W, Hochuli J, Itkin Z, Maxfield T, Melo-Filho C, Thieme A, Wilson K. AI-driven discovery of synergistic drug combinations against pancreatic cancer. 2024; [DOI] [PMC free article] [PubMed]
  • 39.Hochuli JE, Jain S, Melo C, Sessions ZL, Bobrowski T, Choe J, Zheng J, Eastman R, Talley DC, Rai G, Simeonov A, Tropsha A, Muratov EN, Baljinnyam B, Zakharov AV. Allosteric Binders of ACE2 Are Promising Anti-SARS-CoV-2 Agents. Acs Pharmacology & Translational Science. Jun 22 2022;doi: 10.1021/acsptsci.2c00049 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Wellnitz J, Jain S, Hochuli JE, Maxfield T, Muratov EN, Tropsha A, Zakharov AV. One size does not fit all: revising traditional paradigms for assessing accuracy of QSAR models used for virtual screening. J Cheminform. Jan 16 2025;17(1):7. doi: 10.1186/s13321-025-00948-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Martinez Molina D, Jafari R, Ignatushchenko M, Seki T, Larsson EA, Dan C, Sreekumar L, Cao Y, Nordlund P. Monitoring drug target engagement in cells and tissues using the cellular thermal shift assay. Science. Jul 5 2013;341(6141):84–7. doi: 10.1126/science.1233606 [DOI] [PubMed] [Google Scholar]
  • 42.Caballero IM, Lundgren S. A Shift in Thinking: Cellular Thermal Shift Assay-Enabled Drug Discovery. ACS Med Chem Lett. Apr 13 2023;14(4):369–375. doi: 10.1021/acsmedchemlett.2c00545 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Pelcman B, Suna E, Stafford W, Priede M. Hydrocarbylsulfonyl-substituted pyridines and their use in the treatment of cancer. Google Patents; 2022. [Google Scholar]
  • 44.Orwar O, Davidson M. Thioredoxin reductase inhibitors for use in the treatment of cancer. Google Patents; 2021. [Google Scholar]
  • 45.Morgan CA, Parajuli B, Buchman CD, Dria K, Hurley TD. N,N-diethylaminobenzaldehyde (DEAB) as a substrate and mechanism-based inhibitor for human ALDH isoenzymes. Chem Biol Interact. Jun 5 2015;234:18–28. doi: 10.1016/j.cbi.2014.12.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Demyanenko SV, Nikul VV, Uzdensky AB. The Neuroprotective Effect of the HDAC2/3 Inhibitor MI192 on the Penumbra After Photothrombotic Stroke in the Mouse Brain. Mol Neurobiol. Jan 2020;57(1):239–248. doi: 10.1007/s12035-019-01773-9 [DOI] [PubMed] [Google Scholar]
  • 47.MOE. Version 2019.1. Chemical Computing Group ULC; 2019. [Google Scholar]
  • 48.Antolin AA, Sanfelice D, Crisp A, Villasclaras Fernandez E, Mica IL, Chen Y, Collins I, Edwards A, Muller S, Al-Lazikani B, Workman P. The Chemical Probes Portal: an expert review-based public resource to empower chemical probe assessment, selection and use. Nucleic Acids Res. Jan 6 2023;51(D1):D1492–D1502. doi: 10.1093/nar/gkac909 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Skuta C, Southan C, Bartunek P. Will the chemical probes please stand up? RSC Med Chem. Aug 18 2021;12(8):1428–1441. doi: 10.1039/d1md00138h [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Antolin AA, Workman P, Al-Lazikani B. Public resources for chemical probes: the journey so far and the road ahead. Future Med Chem. Apr 2021;13(8):731–747. doi: 10.4155/fmc-2019-0231 [DOI] [PubMed] [Google Scholar]
  • 51.Wassermann AM, Camargo LM, Auld DS. Composition and applications of focus libraries to phenotypic assays. Front Pharmacol. 2014;5:164. doi: 10.3389/fphar.2014.00164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Schreiber SL, Kotz JD, Li M, Aube J, Austin CP, Reed JC, Rosen H, White EL, Sklar LA, Lindsley CW, Alexander BR, Bittker JA, Clemons PA, de Souza A, Foley MA, Palmer M, Shamji AF, Wawer MJ, McManus O, Wu M, Zou B, Yu H, Golden JE, Schoenen FJ, Simeonov A, Jadhav A, Jackson MR, Pinkerton AB, Chung TD, Griffin PR, Cravatt BF, Hodder PS, Roush WR, Roberts E, Chung DH, Jonsson CB, Noah JW, Severson WE, Ananthan S, Edwards B, Oprea TI, Conn PJ, Hopkins CR, Wood MR, Stauffer SR, Emmitte KA, Team NIHMLP. Advancing Biological Understanding and Therapeutics Discovery with Small-Molecule Probes. Cell. Jun 4 2015;161(6):1252–65. doi: 10.1016/j.cell.2015.05.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Roy A Early Probe and Drug Discovery in Academia: A Minireview. High Throughput. Feb 9 2018;7(1)doi: 10.3390/ht7010004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.https://ncats.nih.gov/preclinical/core/compound/npact
  • 55.Henderson MJ, Holbert MA, Simeonov A, Kallal LA. High-Throughput Cellular Thermal Shift Assays in Research and Drug Discovery. SLAS Discov. Feb 2020;25(2):137–147. doi: 10.1177/2472555219877183 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Clausse V, Fang Y, Tao D, Tagad HD, Sun H, Wang Y, Karavadhi S, Lane K, Shi ZD, Vasalatiy O, LeClair CA, Eells R, Shen M, Patnaik S, Appella E, Coussens NP, Hall MD, Appella DH. Discovery of Novel Small-Molecule Scaffolds for the Inhibition and Activation of WIP1 Phosphatase from a RapidFire Mass Spectrometry High-Throughput Screen. ACS Pharmacol Transl Sci. Oct 14 2022;5(10):993–1006. doi: 10.1021/acsptsci.2c00147 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Michael S, Auld D, Klumpp C, Jadhav A, Zheng W, Thorne N, Austin CP, Inglese J, Simeonov A. A robotic platform for quantitative high-throughput screening. Assay Drug Dev Technol. Oct 2008;6(5):637–57. doi: 10.1089/adt.2008.150 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Chtcherbinine M Kinetic and Structural Characterization of Isoenzyme-Selective Aldehyde Dehydrogenase 1A Inhibitors. 2016.
  • 59.Chen Y, Zhu JY, Hong KH, Mikles DC, Georg GI, Goldstein AS, Amory JK, Schonbrunn E. Structural Basis of ALDH1A2 Inhibition by Irreversible and Reversible Small Molecule Inhibitors. ACS Chem Biol. Mar 16 2018;13(3):582–590. doi: 10.1021/acschembio.7b00685 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Graham CE, Brocklehurst K, Pickersgill RW, Warren MJ. Characterization of retinaldehyde dehydrogenase 3. Biochem J. Feb 15 2006;394(Pt 1):67–75. doi: 10.1042/BJ20050918 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Perez-Miller S, Younus H, Vanam R, Chen CH, Mochly-Rosen D, Hurley TD. Alda-1 is an agonist and chemical chaperone for the common human aldehyde dehydrogenase 2 variant. Nat Struct Mol Biol. Feb 2010;17(2):159–64. doi: 10.1038/nsmb.1737 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Parajuli B Identification, kinetic and structural characterization of small molecule inhibitors of aldehyde dehydrogenase 3a1 (ALDH3A1) as an adjuvant therapy for reversing cancer chemoresistance. Indiana University; 2013. [Google Scholar]
  • 63.Jackson BC, Reigan P, Miller B, Thompson DC, Vasiliou V. Human ALDH1B1 polymorphisms may affect the metabolism of acetaldehyde and all-trans retinaldehyde--in vitro studies and computational modeling. Pharm Res. May 2015;32(5):1648–62. doi: 10.1007/s11095-014-1564-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Buchman CD, Hurley TD. Inhibition of the Aldehyde Dehydrogenase 1/2 Family by Psoralen and Coumarin Derivatives. J Med Chem. Mar 23 2017;60(6):2439–2455. doi: 10.1021/acs.jmedchem.6b01825 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Chadwick DJ, Goode JA. Acetaldehyde-related pathology: bridging the trans-disciplinary divide. John Wiley & Sons; 2007. [Google Scholar]
  • 66.Buchman CD, Mahalingan KK, Hurley TD. Discovery of a series of aromatic lactones as ALDH1/2-directed inhibitors. Chem Biol Interact. Jun 5 2015;234:38–44. doi: 10.1016/j.cbi.2014.12.038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Perez-Miller S, Younus H, Vanam R, Chen CH, Mochly-Rosen D, Hurley TD. Alda-1 is an agonist and chemical chaperone for the common human aldehyde dehydrogenase 2 variant. Nature Structural & Molecular Biology. Feb 2010;17(2):159–U4. doi: 10.1038/nsmb.1737 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Morgan CA, Hurley TD. Characterization of two distinct structural classes of selective aldehyde dehydrogenase 1A1 inhibitors. J Med Chem. Feb 26 2015;58(4):1964–75. doi: 10.1021/jm501900s [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Deitrich RA, Petersen D, Vasiliou V. Removal of acetaldehyde from the body. Wiley Online Library; 2006:23–51. [DOI] [PubMed] [Google Scholar]
  • 70.Davis MI, Shen M, Simeonov A, Hall MD. Diaphorase Coupling Protocols for Red-Shifting Dehydrogenase Assays. Assay Drug Dev Technol. Apr 2016;14(3):207–12. doi: 10.1089/adt.2016.706 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Fourches D, Muratov E, Tropsha A. Trust, but Verify II: A Practical Guide to Chemogenomics Data Curation. Journal of Chemical Information and Modeling. Jul 2016;56(7):1243–1252. doi:DOI 10.1021/acs.jcim.6b00129 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Fourches D, Muratov E, Tropsha A. Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling research. J Chem Inf Model. Jul 26 2010;50(7):1189–204. doi: 10.1021/ci100176x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Fourches D, Muratov E, Tropsha A. Curation of chemogenomics data. Nat Chem Biol. Aug 2015;11(8):535. doi: 10.1038/nchembio.1881 [DOI] [PubMed] [Google Scholar]
  • 74.Wang Y, Jadhav A, Southal N, Huang R, Nguyen DT. A grid algorithm for high throughput fitting of dose-response curve data. Curr Chem Genomics. Oct 21 2010;4:57–66. doi: 10.2174/1875397301004010057 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Jain S, Siramshetty VB, Alves VM, Muratov EN, Kleinstreuer N, Tropsha A, Nicklaus MC, Simeonov A, Zakharov AV. Large-Scale Modeling of Multispecies Acute Toxicity End Points Using Consensus of Multitask Deep Learning Methods. J Chem Inf Model. Feb 22 2021;61(2):653–663. doi: 10.1021/acs.jcim.0c01164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Zakharov AV, Zhao T, Nguyen DT, Peryea T, Sheils T, Yasgar A, Huang R, Southall N, Simeonov A. Novel Consensus Architecture To Improve Performance of Large-Scale Multitask Deep Learning QSAR Models. J Chem Inf Model. Nov 25 2019;59(11):4613–4624. doi: 10.1021/acs.jcim.9b00526 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Breiman L Random forests. Machine Learning. Oct 2001;45(1):5–32. doi:Doi 10.1023/A:1010933404324 [DOI] [Google Scholar]
  • 78.Oshiro TM, Perez PS, Baranauskas JA. How Many Trees in a Random Forest? Springer Berlin Heidelberg; 2012:154–168. [Google Scholar]
  • 79.Jain S, Kotsampasakou E, Ecker GF. Comparing the performance of meta-classifiers-a case study on selected imbalanced data sets relevant for prediction of liver toxicity. Journal of Computer-Aided Molecular Design. May 2018;32(5):583–590. doi: 10.1007/s10822-018-0116-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Huang BF, Boutros PC. The parameter sensitivity of random forests. BMC Bioinformatics. Sep 1 2016;17(1):331. doi: 10.1186/s12859-016-1228-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation. Nov 15 1997;9(8):1735–1780. doi:DOI 10.1162/neco.1997.9.8.1735 [DOI] [PubMed] [Google Scholar]
  • 82.Sherstinsky A Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network. Physica D-Nonlinear Phenomena. Mar 2020;404doi:ARTN 132306 10.1016/j.physd.2019.132306 [DOI] [Google Scholar]
  • 83.Friedrich NO, Kops CD, Flachsenberg F, Sommer K, Rarey M, Kirchmair J. Benchmarking Commercial Conformer Ensemble Generators. Journal of Chemical Information and Modeling. Nov 2017;57(11):2719–2728. doi: 10.1021/acs.jcim.7b00505 [DOI] [PubMed] [Google Scholar]
  • 84.Gelardi ELM, Colombo G, Picarazzi F, Ferraris DM, Mangione A, Petrarolo G, Aronica E, Rizzi M, Mori M, La Motta C, Garavaglia S. A Selective Competitive Inhibitor of Aldehyde Dehydrogenase 1A3 Hinders Cancer Cell Growth, Invasiveness and Stemness In Vitro. Cancers (Basel). Jan 19 2021;13(2)doi: 10.3390/cancers13020356 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Parajuli B, Fishel ML, Hurley TD. Selective ALDH3A1 inhibition by benzimidazole analogues increase mafosfamide sensitivity in cancer cells. J Med Chem. Jan 23 2014;57(2):449–61. doi: 10.1021/jm401508p [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Kimble-Hill AC, Parajuli B, Chen CH, Mochly-Rosen D, Hurley TD. Development of selective inhibitors for aldehyde dehydrogenases based on substituted indole-2,3-diones. J Med Chem. Feb 13 2014;57(3):714–22. doi: 10.1021/jm401377v [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Zhang JH, Chung TD, Oldenburg KR. A Simple Statistical Parameter for Use in Evaluation and Validation of High Throughput Screening Assays. J Biomol Screen. 1999;4(2):67–73. doi: 10.1177/108705719900400206 [DOI] [PubMed] [Google Scholar]
  • 88.Auld DS, Thorne N, Boxer MB, Southal N, Shen M, Thomas CJ, Inglese J. Understanding Enzymes as Reporters or Targets in Assays Using Quantitative High-throughput Screening (qHTS). Proceedings of the Beilstein Experimental Standard Conditions of Enzyme Characterizations Symposium. 2010:21–43. [Google Scholar]
  • 89.Huang RL. A Quantitative High-Throughput Screening Data Analysis Pipeline for Activity Profiling. High-Throughput Screening Assays in Toxicology. 2016;1473:111–122. doi: 10.1007/978-1-4939-6346-1_12 [DOI] [PubMed] [Google Scholar]
  • 90.Seethala R Handbook of drug screening. (No Title). 2009;13:489. [Google Scholar]
  • 91.Haas JV, Eastwood BJ, Iversen PW, Devanarayan V, Weidner JR. Minimum Significant Ratio - A Statistic to Assess Assay Variability. Assay Guidance Manual. 2004. [Google Scholar]
  • 92.Stefaniak F Prediction of Compounds Activity in Nuclear Receptor Signaling and Stress Pathway Assays Using Machine Learning Algorithms and Low-Dimensional Molecular Descriptors. Front Environ Sci. 2015;3:77. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

PH4_models
Supporting Information
Supporting Tables

RESOURCES