Skip to main content
NPJ Precision Oncology logoLink to NPJ Precision Oncology
. 2025 Nov 5;9:340. doi: 10.1038/s41698-025-01111-4

DeepTarget predicts anti-cancer mechanisms of action of small molecules by integrating drug and genetic screens

Sanju Sinha 1,2,, Neelam Sinha 1,3, Marlenne Perales 2, Adi Tarrab 4, Trinh Nguyen 5, Lihe Liu 2, Thomas Cantore 1, Kyle Alvarez 2, Sumeet Patiyal 1, Sumit Mukherjee 1, Sanna Madan 1, Kevin Tharp 2, Jianhua Zhao 2, Ranjit Kumar 2, Greg Flanigan 5, John A Beutler 6, Barry R O’Keefe 6,7, Daoud Meerzaman 5, Uri Ben-David 4, Aniruddha J Deshpande 2,, Eytan Ruppin 1,
PMCID: PMC12589559  PMID: 41193619

Abstract

Identifying the mechanisms of action (MOA) driving a drug’s anti-cancer efficacy is critical for its clinical success, guiding the search for its best biomarkers, indications and combinations. Yet, systematically identifying MOAs remains challenging due to drugs often engaging multiple targets with varying affinities across different cellular contexts. Addressing this challenge, we present DeepTarget, a computational tool that integrates large-scale drug and genetic knockdown viability screens with omics data to predict a drug’s MOAs driving its cancer cell killing. To test its performance, we curated eight datasets of high-confidence drug-target pairs focused on cancer drugs and benchmarked DeepTarget. We show that DeepTarget outperforms recent tools in predicting drug targets and their mutation-specificity, achieving strong predictive performance across diverse validation datasets. We experimentally validate DeepTarget’s predictions in two case studies: (a) Demonstrating that pyrimethamine, an anti-parasitic drug, affects cellular viability through modulation of mitochondrial function, specifically the oxidative phosphorylation pathway, and (b) Confirming that T790-mutated EGFR mediates ibrutinib response in BTK-negative solid tumors. Additionally, we demonstrate that kinase inhibitors predicted by DeepTarget to have higher target specificity show increased progression in clinical trials. We provide DeepTarget as an open-source tool (https://github.com/CBIIT-CGBB/DeepTarget) along with predicted target profiles for 1,500 cancer-related drugs and 33,000 unpublished natural product extracts. DeepTarget represents a significant computational advancement among target discovery methods that complements the leading structure-based methods by considering cellular context and can potentially accelerate drug development and repurposing efforts in oncology.

Subject terms: Computational biology and bioinformatics, Drug discovery

Introduction

Knowing a drug’s mechanism of action (MOA) is critical for guiding the search for its best biomarkers, indications, and combinations. Yet it remains a major challenge in drug discovery. Recent computational and integrated methods have advanced this field16, with AlphaFold3, RosettaFold All Atom, and Chai-1 achieving state-of-the-art prediction of protein-small molecule binding affinity in the Posebusters benchmark7,8. However, this testing benchmark comprising static structures lacks interaction dynamics and pharmacokinetics and covers a limited chemical space. Thus, the utility of these models for real-world target identification and drug discovery is yet to be determined.

Here we present DeepTarget, a computational tool that predicts a drug’s mechanisms of action driving its cancer cell killing by integrating drug viability screens with genetic screens and omics profiles from matched cell lines. The “deep” term here reflects the tool’s comprehensive analysis of drug mechanisms rather than deep learning methodology. Unlike structure-based methods that are limited to predicting direct binding interactions, DeepTarget uniquely integrates functional genomic data with drug response profiles to capture both direct and indirect, context dependent mechanisms driving drug efficacy in living cells. This approach allows DeepTarget to identify not only primary drug targets but also secondary targets and pathway-level effects that emerge from cellular context, which purely structural approaches do not aim to predict. Moreover, by leveraging large-scale viability screens across diverse cancer cell lines, DeepTarget can identify context-specific drug mechanisms that may vary across different genetic backgrounds and tissue types - a critical aspect for clinical translation that is an unmet need, but structure-based predictions alone. DeepTarget builds on the principle that CRISPR-Cas9 knockout of a drug’s target gene can mimic the drug’s effects, thus identifying genes whose deletion phenocopies a drug treatment can reveal its potential targets3. Goncalves et al. previously applied this principle to reveal the role of mitochondrial E3 ubiquitin-protein ligases in MCL1 inhibitors’ efficacy3. Expanding on this work and other similar studies4,5, we developed DeepTarget, a computational tool that can predict a drug’s primary targets, secondary targets, and target mutation-specificity. We note that DeepTarget captures both on-target and off-target effects, systematically categorizing them into clinically relevant mechanisms, including secondary targets and context-specific responses. Importantly, we curated eight gold-standard datasets of high-confidence drug-target pairs focused on cancer drugs, where we benchmarked DeepTarget against recent tools.

We also prospectively demonstrate the utility of DeepTarget in two case studies and show its effectiveness in prioritizing kinase inhibitors with a high likelihood of clinical success. While our two case studies cannot fully demonstrate generalizability across all drug classes, they validate distinct aspects of DeepTarget - discovering a novel MOA for drug repurposing and identifying clinically relevant secondary targets, complementing our computational validation across eight diverse datasets. In sum, DeepTarget is the first computational tool to systematically integrate large-scale drug and genetic screens for predicting cancer drug mechanisms of action, including primary targets, context-specific secondary targets, and mutation-specificity preferences. Unlike structure-based methods, it captures both direct and indirect mechanisms in living cells and provides validated predictions for 1500 cancer drugs and 33,000 natural product extracts as an open-source resource.

Results

Overview of the DeepTarget pipeline

DeepTarget is a computational tool designed to predict mechanisms of action driving a drug’s cancer cell killing. It operates on the hypothesis that CRISPR-Cas9 knockout (CRISPR-KO) of a drug’s target gene mimics the drug’s inhibitory effects across cancer cell lines (Fig. 1A, Fig. S1). To function, DeepTarget requires three types of data across a panel of cancer cell lines: (1) drug response profiles, (2) genome-wide CRISPR-KO viability profiles, and (3) corresponding omics data (gene expression and mutation). Chronos-processed CRISPR dependency scores are taken, accounting for confounding factors including sgRNA efficacy variation, screen quality differences, copy number effects, growth rate variation, and technical biases, as much as possible. We obtained this comprehensive data for 1450 drugs across 371 cancer cell lines from the DepMap repository7. This dataset forms the foundation for DeepTarget’s three-step analysis pipeline.

Fig. 1. Overview of DeepTarget and its performance in predicting primary targets across eight gold-standard datasets.

Fig. 1

A The key principle behind DeepTarget and an overview of its three steps, where the red vs. blue refers to sensitive vs. resistant cell line. B A DKS-score based UMAP capturing major cancer drugs MOA. C Benchmarking performance of DeepTarget, Chai-1 and RosettaFold in eight gold-standard datasets. D Model performance across all datapoints including recall, precision, F1, AUC and accuracy. (C) provides one performance measure (AUC) in individual eight gold-standard datasets vs. (D) provides an average performance measure across all the eight datasets for using different kinds of performance measures. E The ROC curve for each dataset is provided, showing specificity vs. sensitivity for DeepTarget. F The DeepTarget predicted DKS score (Y-axis) of high-confidence drug-target pairs (positive, red) vs. negative labels created by shuffling labels (green) from our pipeline across eight gold-standard datasets. Two-sided Wilcoxon rank-sum significance is provided. G DeepTarget predicted the DKS score (Y-axis) of drug-target pairs in more than one dataset (positive, red) vs. negative labels (green). H The relationship between the DKS score threshold to distinguish true vs. false drug-target pairs in above-curated datasets and its specificity & sensitivity observed. The maximum specificity was achieved at the DKS score threshold >0.23. The blue dashed line making this threshold is provided. I The overall distribution of DeepTarget’s predicted number of targets per drug (X-axis) using this DKS score threshold. J The proportion of false positives out of total genes (Y-axis) in case of genome-wide search space vs. kinase proteins only.

The first goal (Primary Target Prediction) is to identify the main protein target(s) responsible for a drug’s anti-cancer effects. The underlying principle is that deleting a drug’s target gene via CRISPR should mimic the drug’s treatment effects. Thus, using matched cell lines with both drug response and genetic profiles, DeepTarget identifies genes whose deletion induces similar viability patterns as drug treatment. This similarity is quantified using Pearson correlation, termed the Drug-KO Similarity score (DKS score) - the higher the score, the stronger the evidence that the gene is a direct target (Fig. 1A, left panel). The DKS score is computed through linear regression that corrects for screen confounding factors (Methods). Validating this approach, a DKS scores-based UMAP of 1450 drugs successfully clusters compounds by their known mechanisms of action (Fig. 1B), including inhibitors of EGFR, HDAC, MDM, MEK, MTOR, RAF, AKT, Aurora Kinases, CDK, CHK, PI3K, PARP, topoisomerase, and tubulin polymerization pathways.

The second goal (Context-Specific Secondary Target Predictions) is to understand how drugs achieve their effects when their primary targets are absent or ineffective. This builds on the clinical observation that drugs often work through different mechanisms in different cellular contexts—for instance, a drug might be effective in cancers that don’t express its primary target, suggesting alternative mechanisms. DeepTarget identifies two types of context-specific secondary targets (Fig. 1A-right panel): (A) Those contributing to efficacy even when primary targets are present - identified through de novo decomposition of drug response into gene knockout effects, revealing both the target and the cell lines where the mechanism is active (Methods). (B) Those mediating responses specifically when primary targets are not expressed - here, we compute Secondary DKS Scores in cell lines lacking primary target expression, using the same approach as primary target identification but in this restricted cellular context. This dual approach captures both general secondary effects and those specific to primary target-deficient contexts, providing a comprehensive understanding of a drug’s efficacy mechanisms.

The third goal (Wild-type vs. Mutant Targeting) is to determine whether a drug preferentially targets the wild-type or mutant forms of its target proteins—a critical distinction for patient selection in clinical trials. For example, some EGFR inhibitors specifically target mutant EGFR and are only effective in patients with those mutations. DeepTarget predicts such preferences by comparing drug-target relationships in different genetic contexts (Fig. 1A-middle panel). The underlying principle is that if a drug specifically targets a mutant form, the similarity between drug treatment and target knockout effects (DKS score) will be significantly higher in cell lines with the mutant target versus those with wild-type. This difference is quantified as a mutant-specificity score, with positive scores indicating preferential targeting of mutant forms and negative scores suggesting wild-type specificity. This analysis provides crucial insights for drug positioning and patient stratification in clinical applications.

We note that DeepTarget’s predictions can include both direct binding targets and other genes in the drug’s mechanism of action pathway. To distinguish between these, we introduced two complementary analyses: a post-filtering step that focuses on likely direct targets (e.g., restricting to kinase proteins for kinase inhibitors), and a pathway enrichment analysis that provides a systems-level view of the drug’s mechanism. This combined approach—from primary targets to context-specific effects to mutation preferences—provides a comprehensive framework for understanding drug mechanisms, while the additional filtering and pathway analyses help interpret and prioritize these predictions for practical applications in drug development.

Curating gold-standard datasets of high-confidence cancer drug-target pairs

We next curated eight gold-standard datasets comprising high-confidence drug-target pairs, with a focus on cancer drugs (Data S1, Methods). This include drug-target pairs that have: (1) A tumor mutation in the target gene causing clinical resistance to the drug from COSMIC (COSMIC resistance, N = 16)8, or (2) oncoKB (oncoKB resistance, N = 28)9, (3) FDA Approval for a target mutation for anti-cancer treatment (FDA mutation-approval, N = 86)9, (4) High-confidence as per the scientific advisory board of ChemicalProbes.org (SAB, N = 24)10, (5) Multiple independent reports as per BioGrid (Biogrid Highly Cited, N = 28)11, (6) Directly interact with target as per DrugBank, and are inhibitors (DrugBank Active Inhibitors, N = 90)12 or (7) are antagonists (DrugBank Active Antagonists, N = 52)12, (8) Highly selective inhibitors based on their binding profile (SelleckChem selective inhibitors, N = 142)13. Negative labels (controls) for each gold-standard dataset were created by choosing random protein-coding genes as target labels (Methods). We addressed three potential biases in our gold-standard datasets: (1) Target Class Bias—by evaluating performance separately in kinase-enriched versus diverse target datasets, (2) Dataset Redundancy—by analyzing both unique and replicated drug-target pairs as confidence indicators, and (3) Target Distribution Bias - by computing class-specific performance metrics.

Benchmarking DeepTarget against recent tools for primary target identification

In the above eight gold-standard datasets, we tested the performance of DeepTarget, RosettaFold All Atom, and Chai-1. DeepTarget stratified the positive vs. negative pairs with a mean AUC of 0.73 across all datasets, compared to 0.58 for RosettaFold and 0.53 for Chai-1 without multiple sequence alignment (Fig. 1C). DeepTarget outperformed the other models in 7 out of 8 tested datasets (P < 0.001, 41% higher AUC than RosettaFold (0.82 vs 0.58) and 55% higher than Chai-1 (0.82 vs 0.53)). RosettaFold was the best-performing model in the remaining one (Fig. 1C). Performance varied across datasets, with oncoKB resistance showing the highest AUC of 0.96 [0.89–0.99], followed by SelleckChem (0.85 [0.79–0.90]), FDA mutation-approval (0.84 [0.76–0.90]), DrugBank Inhibitors and Antagonists (both 0.80 [0.72–0.86 and 0.70–0.87]), BioGrid (0.78 [0.64–0.88]), SAB (0.77 [0.62–0.88]), and COSMIC resistance (0.75 [0.60–0.87]). This performance significantly exceeded that of RosettaFold (mean AUC = 0.58) and Chai-1 (mean AUC = 0.53) across datasets. Overall, DeepTarget performed better in clinically derived datasets (oncoKB AUC = 0.96). Further performance metrics, including precision, recall, accuracy, and F1 score, are provided in Fig. 1D. DeepTarget’s ROC curve and prediction values are shown in Fig. 1E, F. Furthermore, pairs present in multiple gold-standard datasets yielded a DeepTarget AUC of 0.90 (Fig. 1G).

Kinase inhibitors are disproportionately represented in gold-standard (~49%), and thus, we tested DeepTarget’s performance for kinase vs. non-kinase separately: 0.82 for kinase inhibitors and 0.64 for non-kinase inhibitors. Separating further, DeepTarget achieved the following AUCs: 0.61 for ion channel inhibitors, 0.60 for ion-channel inhibitors, 0.61 for nuclear receptors, 0.55 for proteases and 0.65 for metabolic enzymes (Notes S2). Further, computing this stratification power for individual pathways, we found that the top pathways enriched in the targets ranked by DeepTarget prediction power are PI3K/AKT Signaling (FDR < 1E−14), growth factor receptors (FDR < 1E−14), FGFR Signaling (<2E−13) and RAF/MAP kinase pathway (<7E−13). We note that while 50% of our gold standard is kinases, DeepTarget is not trained on this data and thus not biased for kinases.

We next chose a DKS threshold that maximizes precision over recall to minimize costly experimental false positives in drug discovery. We accept the tradeoff of missing some true targets (false negatives) to avoid misleading researchers with false targets (false positives). This was done by finding a threshold that maximizes precision (DKS > 0.23, Fig. 1H, Fig. S2), and we used this threshold for the rest of the study to select primary (and secondary target). This selective threshold identified at least one predicted target for only 27% of the drugs in our study (386 out of 1450 drugs, Fig. 1I). To test performance in real-world cases where the drug target class is known, we further limited the target search space to relevant protein classes (e.g., kinases for kinase inhibitors) at pertinent analyses. For kinase inhibitors, this approach maintained the AUC (0.82) while markedly further reducing false positives from 13% to 0.003% (Fig. 1I, Supplementary Notes 3, empirical P < 1E−04). Similar analyses for other major target classes, including GPCRs, ion channels, metabolic enzymes, and nuclear receptors, are provided in Supplementary Notes 3. While our current threshold performs well across target classes, future larger datasets per target class could enable the optimization of class-specific thresholds, potentially improving prediction accuracy further and leading to additional biological insights. This strategy can be employed by creating a gold-standard dataset from a single target class (N > 50) and optimizing the threshold using this dataset.

Evaluating DeepTarget’s secondary target-driven prediction ability

DeepTarget’s second step predicts novel drug secondary targets, driving its efficacy in a specific cellular context and especially when the primary target is not expressed. This approach intentionally focuses on low primary target expression contexts to systematically identify alternative targets mediating drug response, providing insights into both intended and unintended drug effects that contribute to efficacy. We select a secondary target using the DKS score threshold >0.23, established using large gold-standard data in the previous section. We applied this to Ibrutinib, an FDA-approved drug targeting Bruton’s tyrosine kinase (BTK) in hematological malignancies, which also shows efficacy in solid tumors where BTK is not expressed (Fig. S3A, B). It has been shown that it also binds to EGFR14 and thus the latter is a major suspect that might be driving its anti-cancer efficacy in solid tumors. DeepTarget confirmed BTK as Ibrutinib’s primary target in BTK-expressing cell lines, mostly hematology-derived (DKS score = 0.8, P = 0.1, Rank = 851, Fig. 3A, Data S6, though low rank suggests false positives). In low-BTK expression cell lines, however, EGFR was identified as the secondary target (Rank = 1, EGFR DKS score = 0.44, BTK DKS score = 0.01, Fig. 2B). We hypothesized that if EGFR is indeed a secondary target of Ibrutinib, then cells harboring oncogenic EGFR mutations—which are dependent on constitutive EGFR signaling for survival—should be more sensitive to Ibrutinib treatment than wild-type cells. To test whether Ibrutinib selectively targets EGFR-mutated cells, we conducted in vitro experiments using an isogenic system of MCF10A cells, in which EGFR is constitutively active due to the introduction of the hotspot oncogenic mutation T790M/+. As predicted, Cells harboring the mutation exhibited differential sensitivity to Ibrutinib compared to cells expressing wild-type (WT) EGFR, validating that EGFR is indeed a target of Ibrutinib, (Fig. 2C, D), as predicted by DeepTarget. As a control, we treated the cells with another drug, MG132, a proteasome inhibitor; cells expressing the mutant EGFR gene did not exhibit differential sensitivity to this drug (Fig. S3C). Therefore, the observed differential sensitivity to Ibrutinib is drug-specific, supporting EGFR as a molecular target of this drug.

Fig. 3. Identifying secondary targets of drugs that mediate the response when the primary target is not expressed.

Fig. 3

A Assessing how drug efficacy depends on primary target expression across cell lines. Each point represents a drug-target pair. The y-axis shows the strength of dependence (interaction coefficient in a regression model) between primary target contribution to drug efficacy (DKS score) and target expression. The x-axis shows the statistical significance of this dependence. Red points indicate decreased efficacy in cells with low target expression, suggesting secondary targets drive efficacy when primary targets are absent (top ones are highlighted blue) and green points vice versa. B The predicted secondary DKS score of secondary targets from the analysis only using cell lines with low expression of the primary target (Y-axis) of drugs with high-confidence annotation vs. negative labels (X-axis). The DKS score using all cell lines is also provided in (C) for comparison. D The ROC curve represents the prediction accuracy of secondary targets vs. negative labels for Step 1 vs. Step 3. The respective area under the ROC curve is provided at the bottom right. E As an illustration, we show the response to Palbociclib vs. viability after CRISPR KO of CDK6 and CDK4 in low vs. high expression of CDK6 cell lines.

Fig. 2. Predicting and validating EGFR as a secondary target of Ibrutinib.

Fig. 2

A Viability after BTK KO (X-axis) vs. Ibrutinib response (Y-axis) in BTK low expression (green, Step 3 score) vs. non-low cell lines (red). The Pearson correlation is provided on the left top with the respective color. B The strength of the correlation of ibrutinib response and viability after KO of each gene (Y-axis) is provided with respective significance (X-axis) in low-BTK expression cell lines, aimed at identifying the secondary target of Ibrutinib when BTK is not expressed. Each point representing a gene is colored by its predicted rank from our pipeline. The top two hits, EGFR and GRHL2, are labeled. C Comparison of drug sensitivity, determined by IC50 values (y-axis), to 72 h Irutinib treatment, between EGFR-WT and EGFR-mutant (T790M/+) MCF10A cells (x-axis). n = 5, independent experiments. Mean IC50 = 1.488 and 0.3421 for MCF10A-WT and MCF10A- T790M/+, respectively. T-test based p = 0.041. D Dosage-viability curve showing that EGFR T790M mutated form of MCF10A is more sensitive to Ibrutinib than the EGFR Wildtype form.

We next systematically identified secondary targets for all drugs (N = 1450) in the DepMap screen. First, assessing the importance of secondary targets, we analyzed how primary target contribution varies across cell lines. For 75% of drugs, the contribution (DKS) of primary targets to drug efficacy decreased in cell lines with low primary target expression (Fig. 3A), indicating alternative (secondary) targets drive efficacy when primary targets are absent. Benchmarking secondary target prediction’s ability, we assembled a gold standard dataset of 64 drugs with multiple well-established targets (Methods). Using secondary DKS score as basis, DeepTarget identified these secondary targets with an AUC of 0.92 (Fig. 3B, Fig. S4), significantly outperforming predictions based on all cell lines (AUC 0.85, Wilcox P = 2E-04, Fig. 3C, D). A notable example is palbociclib, which targets CDK6 in CDK6-expressing cell lines and CDK4 when CDK6 is lowly expressed (Fig. 3E). We next observed that ERBB3 was identified as a frequent secondary target for multiple EGFR inhibitors in cell lines with low EGFR activity. Additional examples of correctly identified well-known primary and secondary targets are (RARA and RARB) of Tamibarotene, which could not be uncovered using analysis of all cell lines combined (Supplementary Notes 3, Figs. S45).

Evaluating DeepTarget’s target mutation-specificity prediction performance

DeepTarget’s third and final step identifies a drug’s specificity to target the wildtype vs. the mutant form of the target protein. As an initial test case, we applied this to Dabrafenib, a canonical BRAF V600E mutant inhibitor. As expected, Dabrafenib’s response is only correlated to BRAF in V600E mutant cell lines and DeepTarget correctly identified it to be a V600E mutant inhibitor (Fig. 4A). We observed this trend for all approved BRAF V600E inhibitors including sorafenib, vemurafenib, and encorafenib. This trend extends to olaparib, selectively targeting PARP WT, and, nutlin-3 selectively targeting MDM2 WT and not mutant.

Fig. 4. DeepTarget predicts drug specificity for mutant versus wild-type protein targets.

Fig. 4

A Validation using Dabrafenib, showing stronger correlation between drug response and BRAF CRISPR-KO viability in V600E mutant cell lines (green) compared to wild-type (red) with correlation coefficients shown. B Genome-wide analysis of mutation-specific targeting across 1450 drugs, plotting differential targeting score against statistical significance. Significant hits highlighted in green with key examples labeled. C ROC curve showing DeepTarget’s ability to identify known resistance mutations, mutation-specific drugs, and BRAF V600E inhibitors. D–F Statistical validation of mutation-specificity predictions using Wilcoxon tests comparing drug classes with known mutation specificity versus controls. G Correlation between mutation-specificity scores and clinical advancement stages for BRAF inhibitors, demonstrating higher specificity correlates with clinical success. H, I Additional validation cases showing differential targeting of PARP1 by Olaparib and MDM2 by CGM097 in mutant versus wild-type cell lines, with correlation coefficients indicated.

Using this approach, we next computed the mutation specificity of 1450 drugs in DepMap (Fig. 4B) and provided these scores as a resource in Data S2. We benchmarked this approach’s performance using two gold standard datasets: 1. Mutant-specific drugs (SelleckChem14) and WT-specific drugs with mutations causing clinical resistance (N = 16, oncoKB9). DeepTarget correctly classified known mutant-inhibitors and WT-inhibitors with an average AUC of 0.78 (N = 14, Fig. 4C-H). Further, we observed that approved and clinically advanced BRAF inhibitors are predicted as more specific to BRAF V600E mutation vs. ones that never reached these stages (Fig. 4I). Novel mutant-specific inhibitors identified by DeepTarget are provided in Data S2, including e.g., olaparib, exclusively targeting PARP1 WT form and CGM097 specifically targeting the MDM2 mutant form are provided in Fig. 2J, K.

Predicting the anti-cancer targets of an anti-malarial drug and testing them in a new genome-wide screen

We applied DeepTarget to predict the anti-cancer targets of Daraprim (pyrimethamine), an anti-malarial drug recently shown to have anti-cancer efficacy15,16. DeepTarget predicted that pyrimethamine’s anti-cancer activity likely stems from inhibiting the mitochondrial oxidative phosphorylation (OXPHOS) system (Fig. 5A, B). To validate this prediction, we performed a genome-wide chemogenomic screen using CRISPR-Cas9 in U937 leukemia cells, comparing DMSO control to pyrimethamine treatment (Fig. 5C). In this experimental setup, we expected genes targeted by pyrimethamine to show higher essentiality in the DMSO control compared to the drug treatment condition. This is because in the presence of the drug, these genes’ functions are already inhibited, making their knockout less impactful on cell viability. Consistent with this reasoning and our prediction, genes in the mitochondrial oxidative phosphorylation pathway showed significantly higher essentiality under DMSO control compared to pyrimethamine treatment. Pathway enrichment analysis of these differentially essential genes in drug treatment strongly confirmed the OXPHOS system as the top hit (Fig. 5D), specifically highlighting the electron transport chain (GSEA P = 1E−15), complex I (P = 5E−15), and complex IV (P = 5E−9). Remarkably, DeepTarget correctly identified the top pathway from the chemogenomic screen among 8000 tested pathways, with a confidence score over 20 times higher than any other pathway (Fig. 5E). This finding aligns with previous observations of reduced oxygen consumption rates in pediatric AML cells after treatment with atovaquone, another OXPHOS inhibitor17, further supporting the role of OXPHOS inhibition in pyrimethamine’s anti-cancer effects.

Fig. 5. Predicting anti-cancer MOA of an anti-parasitic drug Pyrimethamine and validating using chemogenomic screen.

Fig. 5

A The DKS score of each gene for pyrimethamine by DeepTarget (x-axis) and respective p-values, highlighting the top gene targets. B The top 10 pathways (y-axis) enriched in ranked list of pyrimethamine’s predicted targets and respective enrichment score (x-axis). The bubbles are colored by NES score, where red are negatively enriched vs. the blue are positively enriched. The enrichment is computed using GSEA. C Genome-wide essentiality scores of each gene when performed without vs. with pyrimethamine (x and y axes, respectively). The top hits in different regions are colored and labeled, where the legend is provided at the top. The likely targets of pyrimethamine are labeled in green. D The top 10 pathways (y-axis) enriched in the target genes from chemogenomic screen (green) are provided with their respective enrichment scores (Odds ratio, x-axis). The significance is computed using one-tail hypergeometric test computing overlap enrichment. E Confidence scores (a product of effect size x significance) for each pathway to be the MOA from DeepTarget (y-axis) and chemogenomic screen (x-axis).

Target-specific Kinase inhibitors have a higher likelihood of clinical success We hypothesized that more target-specific drugs have a higher chance of advancing in clinical trials, potentially due to better indication selection and fewer side effects. We used DeepTarget’s DKS score as a measure of drug-target specificity, proposing that higher scores indicate greater specificity (Methods, section 6.0) and ‘clinical advancement’ was quantified by the stage of the clinical trial the drug has reached. As an initial test case, we examined 41 EGFR-targeting drugs across various clinical stages. We found a positive correlation between DKS scores and clinical advancement (Pearson Correlation Rho = 0.50, P = 8E−04, Fig. 6A). FDA-approved EGFR inhibitors showed significantly higher DKS scores than preclinical inhibitors (P < 0.004, Fig. 6A). Extending this analysis to all 1450 drugs in our screen, we observed a significant association between DKS scores and clinical advancement (Pearson’s Rho=0.39, P = 7.53E−10), which strengthened when considering only targeted therapy drugs (P = 5.96E−08). We next examined this relationship for other targets with >1 drug in each of the five clinical phases and found that for 11 out of 13 targets, clinical advancement increased with DKS scores (Fig. S6). We highlight results for MAP2K1 and MET, the targets with the most drugs (Pearson’s Rho = 0.47 & 0.38, respectively; Fig. 6B, C). These findings first establish the importance of drug target specificity in the success of drugs in clinical trials, and second, they demonstrate DeepTarget’s potential utility in estimating the latter.

Fig. 6. The DKS scores of the known drug targets are associated with the drug’s clinical advancement.

Fig. 6

AC The predicted DKS score increases with clinical advancement. This is done separately for drugs targeting (A) EGFR, (B) MAP2K1 (C) MET. The box color intensity denoted the clinical stages (preclinical to launched). The correlation strength and significance between the clinical stage and the DKS score are provided at the top. D The DeepTarget’s DKS score of each drug for its respective known target is computed using cell lines of different cancer types and is visualized in a heatmap. The cancer types approved by US FDA for the drug are marked with a “*”. E The correlation between drug response and known target CRISPR-KO for FDA-approved single-target cancer inhibitors. The names of drugs whose known targets are ranked within the predicted top 10 targets by provided. F The proportion of drugs where known target passes DeepTarget threshold (X-axis) for various drug categories. G A UMAP of DKS-score of natural products extracts and small molecules in the same space using 40 cell lines from NCI-60. The left panel distinguishes each treatment as either a small molecule or a nautral product extract where the right highlights known MOAs that are clustered together, identified by DBSCAN.

Applying DeepTarget on FDA-approved cancer drugs and 33,000 natural product extracts

We next focused on applying DeepTarget to study the 140 cancer FDA approved drugs included in the PRISM screen. We first aimed to learn the concordance between DeepTarget’s top predicted targets and the currently reported targets of these approved cancer drugs. We first observe that in 50% of the cases, the known target of the approved drugs is reassuringly ranked first in a genome-wide testing (Fig. 6D–F), which is markedly more than the 30% of the known targets ranked at the top when all targeted drugs in the screen (including the non-cancer drugs) were tested. Secondly, as expected, and reassuringly, the known targets of targeted therapies show stronger enrichment in top-ranked DeepTarget predictions both compared to chemotherapy and to the non-cancer drugs (Fig. S8). Thirdly, we observed that drugs targeting a gene mutation are the most target-specific drugs (highest DKS) among all drug categories (Fig. 6D). Fourthly, we observed that approved targeted therapy drugs are more specific to their targets in their respective FDA-approved cancer types than the rest of the cancer types (N = 10, P = 0.10, DKS FC in approved vs. non-approved cancer types =1.5, Fig. 6E). This trend held for 5 out of 6 different MOAs tested (MEK inhibitor, BRAF inhibitor, EGFR inhibitor, CDK inhibitor, androgen receptor antagonist, mTOR inhibito), where androgen receptor antagonist do not show this trend. We did not observe that chemotherapy drugs show higher DKS score in already-approved cancer type vs. rest (P = 0.63).

The targets of approved drugs that DeepTarget failed to capture and have the lowest DKS score can define the scope of drugs and MOAs our tool cannot capture. This includes BCL-ABR, Purine, Cyclogenase, bacterial DNA inhibitor, KIT, dihydropyrimidine dehydrogenase, estrogen receptor, hedgehog pathway.

One of DeepTarget’s strength lies in its ability to predict MOAs without relying on compound structure, making it ideal for analyzing natural product extracts with unknown bioactive components. We thus applied it to an unpublished and proprietary dataset of 33,000 NCI extracts18, screened for viability across NCI-60 cell lines, with matched CRISPR-KO profiles for 40 lines. This diverse collection includes marine (25%), plant (50%), and microbial (25%) samples from global sources. We first performed a feasibility analysis to learn which MOAs can still be reliably identified despite the reduced number of cell lines. To this end we compared the clustering of known MOAs using all available cell lines (Fig. 1B) versus using only the 40 matched cell lines (Fig. 6G, left panel). The UMAP based on DKS scores from these 40 cell lines successfully clustered drugs with the same known MOA in one-third of the cases that were distinguishable using the full dataset. Building on this, we then projected the DKS scores of both the natural product extracts (with unknown MOAs) and the small molecules with known MOAs into the same latent space (Fig. 1G, Right panel). This allowed us to identify natural products that clustered with the five high-confidence MOAs revealed in our feasibility analysis (inhibitors of EGFR, HDAC, PI3K, topoisomerase, and bromodomain, Fig. 1G). We provide the list of these natural product IDs, their predicted MOAs, and DKS scores as a resource for further investigation in Data S3.

Finally, we used DeepTarget to predict targets for 58 drugs with no reported targets in PRISM, identify indirect inhibitors for undruggable oncogenes, and predict and suggest alternative targets for potentially mislabeled drugs. These include identification of misannotated targets for 75 drugs, novel targets for 17 previously targetless cancer drugs, and potential new therapeutic approaches for traditionally undruggable oncogenes like MYCN and MITF. This is provided as a public resource and described in detail in Supplementary Notes 4-7, Fig. S7, 9–14 and Data S45.

Discussion

DeepTarget leverages the principle that CRISPR-KO of a drug’s target gene can mimic the drug’s inhibitory effects, enabling target identification by integrating drug and genetic screens. Unlike recent structural-based methods such as AlphaFold3, RosettaFold, and Chai-1, DeepTarget is agnostic to drug structure information, instead integrating large-scale genetic and drug screens in cancer cell lines. This approach allows DeepTarget to capture the complexities of cellular contexts and drug-target interactions that may be missed by purely structural methods. We demonstrate DeepTarget’s ability to predict a drug’s (1) primary target, (2) preference for wild-type or mutant target proteins, and (3) secondary targets.

DeepTarget’s predictions outperform recent tools in our curated gold-standard datasets, particularly in clinically derived data. The superior performance in real-world scenarios underscores DeepTarget’s potential to accelerate drug development and repurposing efforts. We note that RosettaFold outperformed DeepTarget in the COSMIC resistance dataset comprising mutation-specific drugs against canonical oncogenes, possibly reflecting its advantage when high-resolution crystal structures are available for these well-characterized targets. Thus, in cases when high-confidence structures are available, structure-based approaches should be prioritized. Overall, DeepTarget provides a complementary approach to structure-based methods by integrating cellular context and functional genomic data, though structural validation remains essential for definitive target confirmation.

DeepTarget’s superior performance in our benchmarks likely stems from its integration of functional genomic data with drug response profiles, allowing it to capture both direct and indirect mechanisms driving drug efficacy in living cells. This approach more closely mirrors real-world drug mechanisms, where cellular context and pathway-level effects often play crucial roles beyond direct binding interactions. The performance difference is particularly pronounced in clinically-derived datasets, suggesting that phenotypic screen-based methods may provide a more effective way to identify biologically relevant target proteins from the genome-wide candidate pool. However, importantly, we note that these approaches are potentially complementary - DeepTarget could be used for initial target identification from the full proteome, followed by structure-based methods to optimize binding specificity once candidate targets are identified.

DeepTarget has a few current limitations. DeepTarget has a few current limitations. Firstly, this study has only focused on cell viability phenotype, missing important features like differentiation induction or immune modulation, which can be extended in the future. Though it is extendable to screening data like Cell Painting that can capture these phenotypes19, DeepTarget’s current scope is also limited to drugs screened in PRISM. Such screens with diverse cell types can also enable us to expand DeepTarget beyond the current focus of cancer only. Further, the prediction accuracy for target classes, including GPCRs, nuclear receptors, and ion channels, needs further improvement. The focus on cancer cell viability limits insights into cell targets on other cell types and regarding non-viability-related mechanisms of action. While parallel DeepTarget analyses using RNAi screens yield concordant results in many cases (Fig. S8), one should acknowledge that noise in large-scale CRISPR screens may affect some of the results. More efficient sgRNAs for genetic screens and informed-dosage usage for drug screens can help mitigate this noise. Further, incorporating additional data types in the future, especially metabolomics and proteomics, could enhance prediction accuracy and provide a more comprehensive understanding of drug mechanisms. Adopting DeepTarget for a new drug in the pharmaceutical pipeline requires viability profiling on a large and diverse set of cell lines, where genetic screens are also available. While currently challenging, such large data generation is becoming more common and the applicability of DeepTarget will increase with time.

To adopt DeepTarget for a new drug in the pharmaceutical pipeline, it requires a viability profile on a large and diverse set of cell lines where genetic screens are also available. This large data generation will be a hindrance for such adoption. However, such large data generation is becoming more common and the applicability of DeepTarget will increase with time. In conclusion, DeepTarget represents a significant advancement in computational drug target discovery. By providing a complementary approach to structural-based methods and demonstrating superior performance in clinically relevant scenarios, DeepTarget has the potential to significantly impact drug development, repurposing efforts, and personalized medicine in oncology.

Methods

Data collection

We collected the viability screens after CRISPR-Cas9 and drug treatment from the DepMap database: https://depmap.org/portal/. The drug response screens called PRISM comprise 1450 cancer drugs and genome-wide CRISPR-Cas9 knockout viability profiles (CRISPR-KO essentiality, AVANA) of ~17k genes, which were commonly performed across 371 cancer cell lines. The expression of these 371 cell lines was also downloaded from the DepMap portal. A complete list of public or generated data used across the manuscript is provided in Table 1.

Table 1.

Overview of Data Resources Used in DeepTarget Analysis

Data Type Source Scale Description
Drug Response PRISM/DepMap 1450 drugs × 371 cell lines Drug viability screens across cancer cell lines
Genetic Screens AVANA/DepMap ~17,000 genes × 371 cell lines Genome-wide CRISPR-Cas9 knockout viability profiles
Natural Products NCI 33,000 extracts × 40 cell lines Viability screens of marine, plant, and microbial extracts
Omics Data DepMap 371 cell lines Gene expression and mutation profiles
Validation Sets 8 Gold-Standard Datasets 466 drug-target pairs Curated high-confidence drug-target pairs from COSMIC, oncoKB, FDA, etc

The DeepTarget pipeline

Based on the hypothesis that the CRISPR-Cas9 knockout (CRISPR-KO) of the target of a given drug can mimic the viability effect after the treatment of that drug, the viability after the target gene across multiple cancer cell lines should be similar to that viability observed after treatment with the said drug treatment. Based on this notion, the DeepTarget pipeline comprises three main steps: (Step 1) Predicting the primary Target of the drug, (Step 2) predicting the context-specific secondary targets and finally, (Step) whether the drug binds more to the wildtype or the mutant form of the primary target.

As a first step (Predicting Primary Target), in a genome-wide search, DeepTarget compares a given drug’s viability profiles across 371 cell lines to the viability profile of each gene and computes a similarity strength score (Pearson Correlation Rho).

DKSd,g=ρYd,Yg 1

where:

  • Yd = [Yd1, Yd2, …, Yd371] represents the drug viability response vector across 371 cell lines

  • Yg = [Yg1, Yg2, …, Yg371] represents the Chronos-processed CRISPR-KO dependency scores for gene g across the same 371 cell lines

  • ρ denotes the Pearson correlation coefficient

The higher the score, the higher the likelihood that the gene is the target of the drug. We termed this similarity score the Drug-KO Similarity score or DKS score. This score can range from −1 to 1, where 1 signifies a perfect correlation between viability after drug treatment and gene CRISPR-KO. The specificity of this framework is maximized across curated Platinum-standard datasets at a DKS score threshold > 0.23 (as depicted in Fig. 1).

The second step (Finding secondary targets) identifies two types of secondary targets: those mediating drug response in the absence of primary target expression, and those contributing to efficacy across all cell lines. For the first type, we identify drugs whose DKS score for the primary target decreases in cell lines with low primary target expression (75% of total drugs, N = 1100).

Secondary_DKSd,g= ρ(Ydlow,Yglow) 2

where:

  • Ydlow represents drug d’s viability response in cell lines with low primary target expression

  • Yglow represents gene g’s CRISPR-KO dependency scores in the same low-expression cell lines

  • Cell lines are classified as “low expression” when log₂(TPM) < 2 for the primary target

  • ρ denotes the Pearson correlation coefficient

We then rank genes based on a secondary-DKS score computed in cell lines where the primary target is not expressed. Genes with a secondary-DKS score > 0.23 are predicted as secondary targets. For the second type, we employ sparse dictionary non-negative matrix factorization (NMF) to decompose the drug response profile. We used the following approach:

XWH 3

where:

  • X is the m × n matrix of drug responses (m cell lines, n drugs)

  • W is an m × k matrix of k latent factors

  • H is a k × n matrix of factor loadings

  • k = 4 (as determined by Webster et al. 2022)

  • Factorization performed under sparsity constraints for interpretability

This factorization is performed under sparsity constraints to ensure interpretability. The columns of W represent latent factors that can be interpreted as gene knockout effects. By correlating these latent factors with known gene knockout profiles, we can identify potential secondary targets that contribute to drug efficacy across all cell lines, regardless of primary target expression. Our method assumes that the CRISPR experiment equally effectively knocks out all alleles of a gene. It also assumes that CRISPR-KO of a gene can knock out amplified or fused copies of an oncogene.

As a third step (Predicting mutant vs. wildtype differential targeting), based on the hypothesis that if a drug more specifically targets a mutant form of protein, in cell lines with that mutant form, the similarity score between viability after drug treatment and target CRISPR-KO (their DKS score) would be significantly higher than in the cell lines with WT protein. Linear regression between the viability after the drug and target KO across the cell lines with an interaction term with the mutation status of the target protein. The coefficient of the interaction term models this dependency of the DKS scored on the target mutation status and termed the mutant-specificity score.

Yd=β0+β1Yg+β2M+β3Yg×M+ε 4

where:

  • Yd represents drug viability response across cell lines

  • Yg represents target gene CRISPR-KO dependency scores

  • M is the mutation status indicator (1 = mutant, 0 = wild-type)

  • β3 is the mutant-specificity score (interaction coefficient)

  • Yg×M is the interaction term between gene knockout effect and mutation status

  • A positive β3 indicates preferential targeting of mutant forms and a negative β_3 indicates preferential targeting of wild-type forms

Finally, statistical significance of the coefficient determines if the coefficient is significantly non-zero. We repeated this analysis for specific mutation variants in the drug targets available in at least five cell lines for statistical power. This resulted in only one variant - BRAF V600E mutation. We computed drugs targeting BRAF if and the extent of differential targeting to the V600E variant.

Curating Gold-Standard database for primary target prediction testing

The following sources are used for curation of seven gold standard drug-target pairs datasets (Table S1): 1. Clinical resistance mutation in target from COSMIC (COSMIC resistance, N = 16), or 2. oncoKB (oncoKB resistance, N = 28), 3. FDA Approval for a target mutation (FDA mutation-approval, N = 86), 4. High confidence as per the scientific advisory board of ChemicalProbes.org (SAB, N = 24), 5. Multiple independent reports as per BioGrid (Biogrid Highly Cited, N = 28), 6. pharmacologically active status as per DrugBank drug, i.e., the drug interacts directly with the target as part of its mechanism of action and are inhibitors (DrugBank Active Inhibitors, N = 90) or 7. are antagonists (DrugBank Active Antagonists, N = 52), 8. Highly selective inhibitors based on their binding profile (SelecChem selective inhibitors, N = 142).

We next filtered pairs, all the pairs that had less than ten targets. We next put a qualitative filter including the following to remove low-quality data from these datasets. The pairs from BioGrid were next filtered for at least five independent citations. The pairs from COSMIC must be supported by data from at least 50 patients. The pairs from SAB are filtered for at least an average rating from the board of three. We note that both levels R1 and R2 are considered in the oncoKB resistance dataset, where the AUC observed for pairs in the qualitatively superior and more robust R1 group is higher than for pairs in R2. We also note that all four levels of clinical evidence available are considered from the oncoKB-approved dataset, where the AUC observed for groups are in order of their reliability.

We next filtered drug-target pairs with high qualitative scores

We did this in three largest datasets noted above (SAB, oncoKB resistance, FDA mutation-approval) by utilizing a SAB average score threshold of four, subsetting to oncoKB pairs with resistance mutation information that are of clinical standard of care (R1 group only), FDA mutation-approval pairs that are either FDA approval or standard of care. Subsetting the drug-target pairs by these three criteria yields an average AUC of 0.96.

Statistical Analysis

All statistical analyses were employed based on data distribution, where in the case of non-Gaussian distribution, non-parametric methods were used. Pairwise comparisons were performed using t test in case of Gaussian distribution and two-sided Wilcoxon rank-sum tests in cases of non-Gaussian, while correlations were assessed using Spearman correlation coefficients. All p-values were corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate (FDR) method. Specific statistical tests for each analysis are indicated in the respective figure legends and results text.

Experimental screens testing DeepTarget’s predictions

Towards identifying Daraprim’s MOA, we first established cell cultures of human leukemia U937-Cas9 cells previously generated in our lab from U937 cell line were cultured in RPMI 1640 medium supplemented with 2mM L-glutamine and 1Mm sodium pyruvate, 10% fetal bovine serum (FBS), 50U/ml penicillin/streptomycin (Thermo Fisher Scientific). HEK293T cells were cultured in DMEM medium with 2mM L-glutamine and sodium pyruvate, 10% FBS and 50U/ml penicillin/streptomycin. All cells were incubated in 5% CO2 at 37°C. For Lentivirus production, the human CRISPR Brunello lentiviral pooled library virus was produced using HEK293T cells transfected with the pooled library in LentiGuide-Puro vector together with the pMD2.G and psPAX2 plasmids (Addgene #73178, #12259, #12260, respectively). Transfection was performed using Polyethylenimine (PEI) transfection reagent according to the manufacturer’s protocol. Virus was harvested 48 h post-transfection and filtered through 0.45 µM filter, then concentrated by ultracentrifugation (20,000 rpm for 2 h at 4°C).

Towards lentiviral transduction of the library and screening, transduced cells were incubated in 5% CO2 at 37 °C. 48 h post-transduction puromycin was added at [1 µg/ml] to select for cells expressing sgRNA library. To ensure transduction of a single sgRNA per cell, the multiplicity of infection (MOI) was adjusted to 0.3–0.4. Adequate representation of sgRNAs during the screen was ensured by keeping >1000x cells in culture relative to the library size. Cells were harvested 5 days after puromycin selection, (20 million cells per replicate) and the initial time point T0 was collected. Then in two biological replicates, cells were treated with 0.3 µM pyrimethamine or DMSO vehicle control in a 50 ml final volume of RPMI medium. Cells were counted and replated with the same concentration of drug every 3 to 4 days and 15 days following the start of the experiment, the final pellets (T15) were collected.

For Library preparation, genomic DNA was obtained from T0 and T15 pellets using the Zymo QuickDNATM Midiprep Plus Kit (Zymo Research, Cat. D4068) according to the manufacturer’s instructions. Then a two-step PCR-amplification for sequencing library preparation was conducted with TaKara ExTaqTM Polymerase (Takara, Cat. RR001) and custom primers to achieve adequate sequencing coverage in 1x75bp single-end reads, following published guidelines1. Integrity and concentration of samples was measured by Qubit 4 Fluorometer Thermo FisherTM and barcoded libraries were then pooled and sequenced using an Illumina NextSeq 500 sequencer.

For constructs and stable cell lines, U937 cell line was obtained from Dr. Daniel Tenen lab and maintained in RPMI-1640 medium (Thermo Fisher Scientific, Carlsbad, CA) supplemented with 10% v/v heat-inactivated fetal bovine serum (Thermo Fisher Scientific), 2Mm L-glutamine (Thermo Fisher Scientific) and 100U/ml penicillin/streptomycin (Thermo Fisher Scientific). We first generated the high-efficiency Cas9-editing U937 by transducing these cells with the pLenti-Cas9-blasticidin (Addgene plasmid #52962) and selecting stable clones using flow-cytometry (LSR-Fortessa, BD Biosciences). Clones were tested for editing efficiency by ICE analysis2 and a single clone demonstrating > 90% editing efficiency at the AAVS locus was selected for further studies. Table 2.

Table 2.

Sample Data: BRAF Dependency vs Dabrafenib Response

Cell Line Cancer Type BRAF CRISPR Effect Dabrafenib Response (AUC) BRAF Status
A375 Melanoma -1.42 0.28 V600E
SK-MEL-28 Melanoma -1.38 0.31 V600E
COLO205 Colorectal -0.89 0.42 V600E
HT29 Colorectal -0.95 0.39 V600E
PC-3 Prostate -0.12 0.89 Wild-type

More negative CRISPR Effect values indicate stronger dependency. Lower AUC values indicate higher drug sensitivity

For pooled sgRNA library screen, 20 million U937-Cas9 cells were transduced with the Brunello lentiviral pooled library virus in RPMI medium supplemented with 10% fetal bovine serum, l-glutamine, antibiotics, and 8 µg/ml polybrene. 48 h post-transduction puromycin was added at [1 µg/ml] to select for cells transduced with the sgRNA library. 72 h after addition, puromycin was removed, and the cells were harvested to reach the number needed to start the screen, considering 20 million cells per replicate. 7 days after transduction T0 was collected in two replicates. The screen started with two biological replicates for each treatment (DMSO and pyrimethamine 0.3µM), and the cells were cultured for up to 15 days, considering T0 as day 1, counting and replating under the same conditions as in T0. Then, T15 was collected by taking approximately 20 million cells per replicate, and genomic DNA was obtained from T0 and T15 pellets. Genomic DNA was used for PCR amplification of sgRNAs and sequenced using an Illumina NextSeq 500 sequencer.

To test Ibrutinib’s Secondary Target, we first established Cell Culture of MCF10A WT cells and their mutated derivative T790M/+. These were obtained from Horizon (https://horizondiscovery.com/). MCF10A cells were cultured in DMEM/F-12(Sartorius), supplemented with 5% heat-inactivated horse serum (Biowest), 1% Penicillin-Streptomycin (Sigma-Aldrich), 10 µg/ml Insulin (Sigma-Aldrich), 100 ng/ml Cholera Toxin (Sigma-Aldrich), 0.5 mg/ml Hydrocortisone (Sigma-Aldrich), 20 ng/ml Epidermal Growth Factor (Sigma-Aldrich), 0.1 mg/ml Transferrin (Sigma-Aldrich). Cells were incubated at 37˚C with 5% CO₂, and passaged twice a week using Trypsin-EDTA(0.25%)(Sigma-Aldrich). Cells were tested for mycoplasma contamination using the MycoAlert Mycoplasma Detection Kit (Lonza), according to the manufacturer’s instructions.

For drug treatment, cells were seeded in a 96well plate using Multidrop™ Combi Reagent Dispenser (ThermoFisher). 24 h later, cells were treated with drugs of interest. Cell viability was measured after 72 h using MTT assay (Sigma M2128), with 500 µg/mL salt diluted in complete medium incubated at 37 °C for 3 h. Formazan crystals were extracted using 10% Triton X100 and 0.1 N HCl in isopropanol, and color absorption was quantified at 570 nm and 630 nm using BioTek Synergy H1 Hybrid Multi-Mode Reader. Absolute IC50 for MG132 and Ibrutinib drugs was calculated using GraphPad PRISM 10.1.1, inhibitor vs. normalized response equation. EC50 for Osimertinib and Capsaicin was calculated using GraphPad PRISM 10.1.1 inhibitor vs. response (four parameters) equation. All drug details are available in Data S1.

Collecting gold-standard datasets for validation of primary target prediction

As provided in detail below, we collected eight gold-standard datasets with high-confidence drug-target pairs from the following sources and preprocessed them to validate our pipeline’s prediction of the primary target. We only focused on drugs that have less than five targets in the gold-standard datasets. (I) The first database is curated of drug-target pairs where the target has a clinical resistance mutation extracted from COSMIC v91 (https://cancer.sanger.ac.uk/cosmic/download). We filtered targets with at least 50 unique patients with clinical resistance mutation (N = 16)9. (II) We repeated this exercise for all such drug-target pairs with resistance mutation from oncoKB, a Precision Oncology Knowledge Base from MSKCC. (N = 28, https://www.oncokb.org/actionableGenes#sections=Tx)10. (III) We next downloaded drug-target pairs with FDA Approval for a target mutation (FDA mutation-approval, N = 86) again from oncoKB10. We considered pairs with all four levels of evidence in this database. (IV) A compilation of high-confidence drug-target pairs from ChemicalProbes.org. This is manually curated by their scientific advisory board (N = 24)11. Pairs with a confidence score >3 (out of 4) are considered following database guidelines (https://www.chemicalprobes.org/information-centre#historical-compounds). This confidence score is computed by reviewing the publication and cellular and/or in vivo model systems. (V) A compilation of drug-target pairs with multiple independent reports as per BioGrid from https://thebiogrid.org/ (N = 28)12. We only considered pairs with at least three independent reports. (VI–VII) A compilation of drugs categorized as inhibitors or antagonists where the target is categorized as pharmacologically active status as per DrugBank drug, i.e., the drug interacts directly with the target as part of its MOA13. This is downloaded from https://go.drugbank.com/ and yielded 90 inhibitors and 52 antagonists. (VIII) A list of highly selective inhibitors downloaded from compound libraries of https://www.selleckchem.com/ (SelecChem) curated based on their binding profile (N = 142)14.

Compiling gold-standard datasets for validation of mutant-specific targeting prediction

We collected two databases to validate our mutant-specific targeting prediction. (I) A curation of drug-target pairs where the target has a clinical resistance mutation extracted from oncoKB (https://www.oncokb.org/actionableGenes#sections=Tx). Here, due to the resistance mutation, the drugs target the mutant form more than the wild-type form9. We validated our method separately for both oncoKB tier I and II and yielded a significantly higher AUC for tier I. (II) We compiled all drugs where we have their binding profile with both target WT vs. mutant forms. To this end, we extracted binding information of all the drugs in our screen, also in https://www.selleckchem.com/. Filtering further, we focused on genes that have a specific variant in at least five cell lines in our screen to have statistical power in our comparison of WT vs. mutant. This resulted in 13 genes, where only one is targeted by a drug in our screen (BRAF). We accordingly identified six drugs specifically more binding to BRAF V600E mutant form and seven that bind to both WT and mutant form comparably. Our predicted mutant-specificity score can not only stratify the two types of drugs but also correctly ranks the extent of differential binding to the V600E variant.

Predicting Ibrutinib’s secondary target

To identify Ibrutinib’s primary and secondary target, we first identified cell lines without (or low) and with BTK (Ibrutinib’s known primary target) expression. We categorized cell lines lower than log2(TPM) < 2 as cell lines with low BTK expression. We compare viability after ibrutinib treatment to all the genes in cell lines with BTK vs. low-BTK, where in the former, the DKS score for BTK is greater than DeepTarget primary target threshold (DKS score > 0.23), confirming it is a primary target. However, in cell lines with low BTK instead, the DKS score does not pass this threshold, and EGFR is the top-ranked target.

Predicting clinical advancement

We hypothesize that a drug’s DKS for its known target, denoting its specificity towards the target, would predict its clinical advancement. We note that this hypothesis makes three assumptions: 1. This framework together considers the two types of possible errors in target information: I) entirely wrong target, II) claimed target is correct but non-specific as is often the case for kinase inhibitors. 2. Our analysis is confounded by the time period since the discovery of the drug and the start of phase I trials because a drug may not have advanced solely because it is new. We did not correct this as our screen does not comprise recently discovered drugs, and most of the drugs in our screen have been in trials or a preclinical stage for long enough to mitigate this effect. 3. We consider different generations of drugs equivalently and do not consider that a latter generation may have a predecessor generation drug with effectiveness. In the case of drugs targeting EGFR, we consider all drugs targeting EGFR, whether they target the specific WT or mutant form and different generations equally.

Predicting cancer types that would likely get approved for a drug

We next hypothesize that a drug would be more specific to its target in cancer types where it is approved. We note that this hypothesis assumes the following: (1) we consider a drug targeting the mutant or WT form equally. (2) We also assume that the frequency of the mutation among cell lines is a good representation of the frequency of the mutation among patients.

Supplementary information

Data S1 (6.9MB, docx)
Data S2 (9.4KB, xlsx)
Data S3 (9.1KB, xlsx)
Data S4 (16.3KB, xlsx)
Data S5 (77.6KB, xlsx)
Data S6 (15.5KB, xlsx)
41698_2025_1111_MOESM7_ESM.csv (1.8MB, csv)

DeepTarget_Supplementary_Text_and_Figures_Revised_Sept1_2025

Acknowledgements

The authors thank the Sanford Burnham Prebys Medical Discovery Institute and National Cancer Institute for providing financial and infrastructural support. This research was supported by Sinha Lab start-up package provided by Sanford Burnham Prebys Medical Discovery Institute, CCSG grant P30CA030199, and, the Intramural Research Program of the National Institutes of Health, NCI.

Author contributions

S.S. and E.R. conceptualized the study, developed the initial framework, and jointly supervised the project. N.S. performed the preliminary analysis and established analytical frameworks. S.S. led the computational analyses. M.P. conducted CRISPR screens for drug secondary target validation. A.V.T. performed experimental validation of ibrutinib's secondary targets. T.N., G.F., and D.M. developed the Bioconductor package. T.C. contributed to secondary target analysis. K.A. performed benchmarking studies against existing tools. K.T. and J.Z. provided biological insights and interpretation. B.R.O. led the natural products analysis. A.J.D. led the experimental CRISPR validation studies. All authors contributed to writing and reviewing the manuscript.

Data availability

Processed PRISM drug screen, expression, mutation, copy number, CRISPR-Cas9, and shRNA pooled genetic screen data were derived from DepMap v22Q2 and can be found here (https://depmap.org/portal/). To accessibly use DeepTarget, we developed Bioconductor package (https://bioconductor.org/packages/devel/bioc/html/DeepTarget.html) and Github package of DeepTarget with installation instructions: https://github.com/CBIIT-CGBB/DeepTarget. Furthermore, we provided scripts reproduce each Figure in the main text and supplementary figures in a GitHub repository, which can be accessed here: https://github.com/ruppinlab/drug.OnTarget.

Code availability

Processed PRISM drug screen, expression, mutation, copy number, CRISPR-Cas9, and shRNA pooled genetic screen data were derived from DepMap v22Q2 and can be found here (https://depmap.org/portal/). To accessibly use DeepTarget, we developed Bioconductor package (https://bioconductor.org/packages/devel/bioc/html/DeepTarget.html) and Github package of DeepTarget with installation instructions: https://github.com/CBIIT-CGBB/DeepTarget. Furthermore, we provided scripts reproduce each Figure in the main text and supplementary figures in a GitHub repository, which can be accessed here: https://github.com/ruppinlab/drug.OnTarget.

Competing interests

Authors Eytan Ruppin (E.R) and Sanju Sinha are the inventors on a pending provisional patent application related to the DeepTarget algorithm. E.R. is a co-founder of Medaware, Metabomed, and Pangea Therapeutics (divested from the latter). E.R. is a non-paid scientific consultant to Pangea Therapeutics, a company developing a precision oncology SL-based multi-omics approach. The rest of the authors declare no conflict of interest.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Sanju Sinha, Email: sanju@terpmail.umd.edu.

Aniruddha J. Deshpande, Email: adeshpande@sbpdiscovery.org

Eytan Ruppin, Email: eyruppin@gmail.com.

Supplementary information

The online version contains supplementary material available at 10.1038/s41698-025-01111-4.

References

  • 1.Corsello, S. M. et al. Discovering the anticancer potential of non-oncology drugs by systematic viability profiling. Nat. Cancer1, 235–248 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Cheng, F. et al. Prediction of drug-target interactions and drug repositioning via network-based inference. PLoS Comput. Biol.8 (2012). [DOI] [PMC free article] [PubMed]
  • 3.Goncalves, E. et al. Drug mechanism-of-action discovery through the integration of pharmacological and CRISPR screens. Mol. Syst. Biol.16, e9405 (2020). [DOI] [PMC free article] [PubMed]
  • 4.Douglass, E. F. Jr et al. A community challenge for a pancancer drug mechanism of action inference from perturbational profile data. Cell Rep. Med.3, 100492 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Bansal, M. et al. A community computational challenge to predict the activity of pairs of compounds. Nat. Biotechnol.32, 1213–1222 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Shen, Y. et al. Systematic, network-based characterization of therapeutic target inhibitors. PLoS Comput. Biol.13, e1005599 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Tsherniak, A. et al. Defining a cancer dependency map. Cell170, 564–576 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tate, J. G. et al. COSMIC: the catalogue of somatic mutations in cancer. Nucleic Acids Res.47, D941–D947 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chakravarty, D. et al. OncoKB: a precision oncology knowledge base. JCO Precis. Oncol.1, 1–16 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Arrowsmith, C. H. et al. The promise and peril of chemical probes. Nat. Chem. Biol.11, 536–541 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Stark, C. et al. BioGRID: a general repository for interaction datasets. Nucleic Acids Res.34, D535–D539 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Wishart, D. S. et al. Drugbank: a comprehensive resource for in silico drug discovery and exploration. Nucleic Acids Res.34, D668–D672 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.SelleckChem, https://www.selleckchem.com/.
  • 14.Wang, A. oli et al. Ibrutinib targets mutant-EGFR kinase with a distinct binding conformation. Oncotarget7, 69760 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.D’Alessandro, S. et al. The use of antimalarial drugs against viral infection. Microorganisms8, 85 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Sharma, A. et al. Pyrimethamine as a potent and selective inhibitor of acute myeloid leukemia identified by high-throughput drug screening. Curr. Cancer Drug Targets16, 818–828 (2016). [DOI] [PubMed] [Google Scholar]
  • 17.Stevens, A. M. et al. Atovaquone is active against AML by upregulating the integrated stress pathway and suppressing oxidative phosphorylation. Blood Adv.3, 4215–4227 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Evans, J. R. et al. National Cancer Institute (NCI) program for natural product discovery: exploring NCI-60 screening data of natural product samples with artificial neural networks. ACS omega8, 9250–9256 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Rohban, M. H. et al. Virtual screening for small-molecule pathway regulators by image-profile matching. Cell Syst.13, 724–736 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1 (6.9MB, docx)
Data S2 (9.4KB, xlsx)
Data S3 (9.1KB, xlsx)
Data S4 (16.3KB, xlsx)
Data S5 (77.6KB, xlsx)
Data S6 (15.5KB, xlsx)
41698_2025_1111_MOESM7_ESM.csv (1.8MB, csv)

DeepTarget_Supplementary_Text_and_Figures_Revised_Sept1_2025

Data Availability Statement

Processed PRISM drug screen, expression, mutation, copy number, CRISPR-Cas9, and shRNA pooled genetic screen data were derived from DepMap v22Q2 and can be found here (https://depmap.org/portal/). To accessibly use DeepTarget, we developed Bioconductor package (https://bioconductor.org/packages/devel/bioc/html/DeepTarget.html) and Github package of DeepTarget with installation instructions: https://github.com/CBIIT-CGBB/DeepTarget. Furthermore, we provided scripts reproduce each Figure in the main text and supplementary figures in a GitHub repository, which can be accessed here: https://github.com/ruppinlab/drug.OnTarget.

Processed PRISM drug screen, expression, mutation, copy number, CRISPR-Cas9, and shRNA pooled genetic screen data were derived from DepMap v22Q2 and can be found here (https://depmap.org/portal/). To accessibly use DeepTarget, we developed Bioconductor package (https://bioconductor.org/packages/devel/bioc/html/DeepTarget.html) and Github package of DeepTarget with installation instructions: https://github.com/CBIIT-CGBB/DeepTarget. Furthermore, we provided scripts reproduce each Figure in the main text and supplementary figures in a GitHub repository, which can be accessed here: https://github.com/ruppinlab/drug.OnTarget.


Articles from NPJ Precision Oncology are provided here courtesy of Nature Publishing Group

RESOURCES