Abstract
In computational-aided drug discovery, structure-based drug design models are computationally intensive and rely on protein structures, limiting their scalability and generalization. Additionally, many existing models suffer from inflated false-positive rates due to the scarcity of negative binding data for training. To overcome these challenges, we present ProMol_Func, a structure-free deep learning framework that integrates graph-based encodings of small molecules with protein function embeddings derived solely from amino acid sequences. By augmenting the training data set with both experimentally validated inactives and randomly selected decoys, ProMol_Func improves screening power and generalization. The model achieves state-of-the-art performance on the challenging LIT-PCBA (Library of Integrated Targeted-Panel of Cell-Based Assays) benchmark, with an enrichment factor (EF1%) of 10.9, demonstrating robust screening power in realistic assay settings. Furthermore, in a zero-shot prospective application to E. coli DnaK, a protein chaperone without actives in the training set, ProMol_Func successfully identified compounds that inhibit its ATPase activity or alter the protein’s thermal stability, validating the potential of ProMol_Func for discovering binders toward novel targets. These results position ProMol_Func as an efficient and scalable alternative to traditional structure-dependent approaches in early stage hit discovery.
Keywords: virtual screening, structure-free, deep learning, protein−ligand binding, zero-shot learning, E. coli DnaK


Virtual screening (VS) is a computational technique in drug discovery used to search large libraries of small molecules and identify candidates likely to bind a specific biological target for experimental validation. − Docking, a key method in structure-based drug design (SBDD), simulates ligand–receptor binding by evaluating geometric fit and binding energy, often achieving higher hit rates than high-throughput screening (HTS). ,− Its success relies on sampling diverse ligand conformers to find optimal binding orientations and accurately estimating binding affinity using scoring functions. ,− The extensive conformational sampling leads to significant computational demands, making docking tools less practical for large-scale VS, particularly in the early phases of drug discovery, where speed and efficiency is crucial. Meanwhile, SBDD models face more challenges when applied to novel targets lacking high-quality crystal structures or targets without known or well-resolved binding sites. ,,
In recent decades, machine learning (ML) has significantly advanced the field of drug discovery. ,− However, data-driven deep learning models trained on PDBbind structures for regression tasks such as binding affinity prediction, lack negative data to enforce the training, thus their enrichments of true hits are limited or biased. ,, Many ML-driven hit discovery methods with experimental validation − remain target-specific and struggle to generalize beyond known binders. Recently, BIND introduced a structure-free deep learning framework that represents compounds as molecular graphs derived from SMILES and encodes proteins using ESM-2, a state-of-the-art protein language model. Although BIND achieved superior performance on well-established virtual screening benchmark test sets, it has not been prospectively validated experimentally, and thus the real-word applicability of structure-free deep learning approaches has not been demonstrated.
In realistic early stage hit discovery scenarios, researchers often begin without experimentally confirmed binders for a novel protein target while several binders may have already been found for homologues, proteins with similar molecular functions (MF). , Considering that small molecules often exert their effects by modulating MF, either directly or allosterically, − here we were motivated to explore a function-level modeling strategy in which protein molecular function embeddings serve as input features in binder classification. This enabled structure-free virtual screening by bypassing the need for 3D structural information altogether. To represent proteins, we used DeepFRI to predict MF from sequence alone, enabling applicability to proteins without structural data. These function scores were linearly transformed and added to trainable small molecule graph embeddings generated by KANO, as outlined in the ProMol_Func framework shown in Figure . Later we evaluated ProMol_Func’s zero-shot learning ability on a target chaperone protein, E. coli DnaK, and experimentally validated two ligands and one destabilizer.
1.
ProMol_Func incorporates a pretrained sequence-based protein function prediction model and a 2D graph encoder for small molecules to train a general classification model to prioritize binders over nonbinders. a) DeepCNN takes protein sequences as input (the structure shown is an AlphaFold2 prediction of DnaK based on the GenBank CAA41306 sequence), and 489 classes of protein molecular function prediction scores are transferred as protein embeddings and combined with b) KANO small-molecule representations, followed by feed-forward layers for binary classification. In KANO architecture, functional prompts were added to the Chemprop graph encoder. The KANO model schematic is adapted from ref , licensed under CC BY 4.0. All other aspects of the figure are original.
Model Development and Evaluation
Results on a pilot data set showed that adding protein function embeddings during training slightly outperforms concatenation-merging two vectors (Table S1). To expand the training data, we extracted 2.55 M active protein–ligand pairs covering 7445 proteins from BindingDB. Each pair was annotated with at least one of the following binding affinity metrics, K I , IC50, EC50 or K D , with values below 50 μM (Scheme S1). Guided by large-scale analyses of virtual screening campaigns and the potency distribution of LIT-PCBA actives, which show that most hits are reported in the low- to midmicromolar range, we adopted 50 μM as a practical and reproducible cutoff to define actives and inactives (see Supporting Information). Accordingly, 0.19 M compounds with any of the metrics higher than 50 μM were labeled as inactives. In addition, 1.30 M inactives (2000 inactive compounds per protein for around 650 proteins) from Pubchem bioassays were retrieved. To address data imbalance, we additionally sampled 2.55 M decoys, referring to assumed negatives, from BindingDB (details in SI). In total, the curated ProMol_Func data set includes 7817 proteins and 6.60 M compounds. To enhance robustness and generalization, we trained an ensemble of three models with different scaffold-balanced splitting seeds. This ensemble, referred to as ProMol_Func, with predictions averaged, was used for downstream benchmarking evaluations.
LIT-PCBA (Library of Integrated Targeted-Panel of Cell-Based Assays) is an unbiased subset of the broader PubChem BioAssay data consisting of 15 protein targets and is widely used for virtual screening and drug discovery. ProMol_Func achieved a top average 1% enrichment factor (EF1%) of 10.90 across all LIT-PCBA targets (Figure ), closely matching the best-performing model, BIND (EF1% = 10.93). This demonstrates ProMol_Func’s superior screening power in real-world scenarios. Per-target EF values are reported in Table S2, which also illustrates that certain targets (e.g., ESR1_ago and MTORC1) remain more challenging and provides a detailed view beyond the averaged results in Figure . For overlap analysis, we found highly similar protein–compound pairs between the training data and LIT-PCBA (Table ), with per-target counts summarized in Table S3. However, after removing the 209 identical compounds from LIT-PCBA, the ProMol_Func EF1% remained 10.90. Furthermore, after excluding samples whose protein-small-molecule similarity scoring (methods in SI) exceeded thresholds of 0.9, 0.7, or 0.5 from LIT-PCBA, ProMol_Func EF1% remain consistently strong for “dissimilar” samples as shown in Table , with per-target EF1% results provided in Tables S4–S6.
2.

Top 1% enrichment factor (EF1%) for the ProMol_Func general model (ensemble predictions of 3 models) on LIT-PCBA targets. Performance of other models is reported based on data from BIND and previous work. , Figure adapted from ref available under a Creative Commons CC BY license.
1. ProMol_Func EF1% on the LIT-PCBA Test Library after Removing Identical Compounds and Highly Similar Protein–Compound Pairs (Similarity Thresholds of 0.9, 0.7, and 0.5) from LIT-PCBA .
| Sim_Threshold | Samples Removed | EF1% |
|---|---|---|
| Identical | 209 | 10.90 |
| 0.9 | 15,036 | 10.89 |
| 0.7 | 31,580 | 10.76 |
| 0.5 | 135,369 | 8.91 |
“Sim_Threshold” refers to protein sequence identities × small molecule Morgan-fingerprints Tanimoto similarity. “Samples Removed” indicates the number of protein–compound pairs excluded for exceeding the corresponding similarity threshold.
Unlike LIT-PCBA, the DUD-E and DEKOIS2.0 data sets incorporate decoys. On DUD-E, ProMol_Func achieved the highest EF1% (48.67), surpassing BIND (46.35), while ranking second on DEKOIS2.0 with an EF1% of 22.94 (Figure S2). ProMol_Func also demonstrated consistently strong EF1% on the “dissimilar” subsets of both DUD-E and DEKOIS2.0 (Tables S7 and S8). These results confirm the robustness and generalizability of ProMol_Func in virtual screening tasks for various proteins.
Negatives data are essential for building robust classification models. , As shown in Table , the inclusion of decoys significantly enhances our model’s predictive performance. Model 1, trained on 1.02 M actives and 1.20 M inactives, achieved an EF1% of only 3.24 on the LIT-PCBA benchmark. Introducing 1.02 M decoys significantly improved the EF1% to 6.83. This improvement likely stems from the imbalance between inactives and actives for individual proteins (data distribution plotted in Figure S1), which decoys help to balance
2. Training Data Size and Top 1% Enrichment Factor (EF1%) on the LIT-PCBA Benchmark for Different Promol_Func Models.
| Model | Actives | Inactives | Decoys | EF1% |
|---|---|---|---|---|
| Model 1 | 1.02 M | 1.20 M | 0 | 3.24 |
| Model 2 | 1.02 M | 1.20 M | 1.02 M | 6.83 |
| Model 3 | 2.04 M | 1.20 M | 2.04 M | 8.22 |
| 3*Model 3 | – | – | – | 10.90 |
3*Model 3 means an ensemble of three Model 3 with averaged predictions and is referred to as ProMol_Func.
Zero Shot Learning Application
Heat shock protein families, namely HSP90 and HSP70, are ATP-dependent chaperones that maintain proteostasis in cells by folding and stabilizing client proteins or assisting in targeting proteins for degradation. Eukaryotic HSP70s and HSP90s have emerged as synergistic targets in anticancer therapy. − The training data included many actives and inactives for human HSP90 (Uniprot ID: P07900). Despite a highly imbalanced human HSP90 external test set (∼30 times more inactives; Table ), ProMol_Func achieved strong accuracy with an F1 score of 0.858 and precision of 0.941 at a 0.5 threshold (Table S9).
3. Overview of Active and Inactive Data for human HSP90 and E. coli DnaK .
| Training
Data set |
Test
Data set |
|||
|---|---|---|---|---|
| Target | Actives | Inactives | Actives | Inactives |
| HSP90 | 2658 | 9411 | 1579 | 284049 |
| ecDnaK | 0 | 24 | 24 | 3634 |
External data for human HSP90 was collected from Pubchem AID 1789, 1946, 1947, 1912, 1913 and chemBL, D3R AbbVie-CSAR. External E. coli DnaK (ecDnaK) data was collected from Pubchem AID 1033 and previous works. − Overlapping data were removed for a model test.
E. coli DnaK, a member of the HSP70 family − that shares 49.60% sequence identity (global alignment with Clustal Omega via UniProt Align) with a major human HSP70, HSPA1A, has been used as a model for human HSP70 inhibitor identification. Further, bacterial DnaKs are putative infectious disease targets. , However, E. coli DnaK has few known binders, limiting target-specific model training. Additionally, its available crystal structures are also scarce and often incomplete, typically representing only individual domains, the nucleotide-binding domain (NBD) or the substrate-binding domain (SBD). − E. coli DnaK contains multiple allosteric sites and exhibits distinct conformations in ATP- and ADP-bound states, , complicating structure-based drug design (SBDD) efforts. The hit rate for HSP70 is known to be lower than for HSP90 in HTS, likely due to HSP70s' tighter nucleotide binding, which makes competitive inhibition more difficult. , Additionally, chaperone inhibitors often show only micromolar potency and incomplete inhibition. ,
The training data of ProMol_Func did not include any active compounds for E. coli DnaK (Table ), making it an ideal target of zero-shot learning. Twenty-four active ligands and 3634 inactives from Pubchem and previous works − were used as the test set for evaluation (Table ). ProMol_Func achieved a precision of 0.381 at a 0.45 threshold and 0.143 at 0.5 on the E. coli DnaK test set (Table S9), highlighting the model’s ability to transfer across functionally similar proteins while not relying on high sequence identity or structural homology. Despite the low sequence identity between Human HSP90 and E. coli DnaK (20.86%, global alignment with Clustal Omega via UniProt Align), their functional similarity may still offer transferable signals. As shown in Figure , increasing the threshold reduces false positives in DnaK inhibitor selection. We applied ProMol_Func to screen the commercially available 575302 compounds from the Asinex library for potential inhibitors of E. coli DnaK (Uniprot ID: C3TRK2) and identified 185 compounds with predicted binding probabilities above 0.5. From these, 43 compounds were selected and purchased based on their predicted binding probability as well as favorable solubility profiles (Scheme S2).
3.

Count of true positives (TP) and false positives (FP) predicted by ProMol_Func on human HSP90 and E. coli DnaK (ecDnaK) test data sets across different thresholds.
DnaJ and GrpE are essential cofactors that modulate the ATPase and folding activities of DnaK, as illustrated in Figure a. DnaJ recognizes and binds unfolded proteins, subsequently delivering them to DnaK while stimulating its ATPase function. GrpE acts as a nucleotide exchange factor, promoting the exchange of ADP for ATP and facilitating substrate release. Together, DnaJ and GrpE tightly regulate DnaK’s ATP hydrolysis, which is critical for DnaK’s chaperone functions. , Among the 43 ProMol_Func-predicted binders (EG) of DnaK, EG36 emerged as a promising hit, showing 20–60% inhibition of DnaK ATPase activity stimulated by DnaJ-GrpE. A structural analog, EG35, also showed inhibition, but with weaker a potency (Figures b and S3). Co-sedimentation experiments confirmed that the compounds do not aggregate under ATPase assay conditions (Figure S4). Both exhibited modest potency relative to Telaprevir (TP), an HSP70 SBD inhibitor that allosterically reduces ATP turnover (Figure c). Notably, EG36 showed stronger inhibition of ADP production with increasing [DnaJ], suggesting it may disrupt functional DnaK–DnaJ interactions (Figure d).
4.
a) Scheme of cofactor-mediated (DnaJ and GrpE) activation of E. coli DnaK’s ATPase activity. b) Of the 43 compounds (100 μM) predicted by ProMol_Func to bind E. coli DnaK, two (EG35 and EG36) showed consistent inhibition of ATPase activity over two trials; chemical structures of hit compounds on the right (100 μM compound, 4 μM E. coli DnaK + 0.4 μM each of DnaJ2 and GrpE). c) Titration of EG35 and EG36 to DnaK-cofactor ATPase reactions (4 μM E. coli DnaK and 0.4 μM DnaJ2 and GrpE) indicates inhibition. Known inhibitor telaprevir (TP) serves as a positive control. EG20 did not show inhibition, thus acting as a negative control. d) Compounds’ inhibition of DnaK (4 μM) and GrpE (0.4 μM) reactions with increasing DnaJ. e) Morgan fingerprints Tanimoto similarity of EG36 to EG35 and other known E. coli DnaK inhibitors. f) Chemical structure of known DnaK inhibitors: Telaprevir (TP), JG-98, Quercetin and 2Cl-IB-MECA.
Notably, EG36 exhibits high structural similarity (0.85) to EG35 yet displays low similarity to known DnaK inhibitors (lower than 0.2) such as TP, JG-98, Quercetin and 2Cl-IB-MECA (Figures e,f and S5), underscoring their structural novelty. Fluorescence polarization (FP) assays showed that neither EG36 nor EG35 displaced the fluorescent peptide bound to the SBD, unlike TP, indicating they do not target the peptide-binding cleft (Figure S7). AlphaFold3 predicted EG36 binding at an allosteric site on the NBD, while Boltz-2 suggested interaction at the NBD-SBD linker domain (Figure S9). Both compounds inhibited the ATPase activity of the isolated NBD, supporting direct binding to this domain (Figure S10). Collectively, these results suggest that EG36 and EG35 represent a novel class of NBD-targeting inhibitors. Furthermore, they do not appear to be general Hsp inhibitors, as neither affected the thermal stability nor ATPase activity of a bacterial HSP90 called HtpG (Table S12, Figure S11). Nonetheless, their moderate inhibition of DnaK–cofactor ATPase activity indicates room for further optimization.
We next assessed the impact of ProMol_Func-predicted candidates on E. coli DnaK thermal stability using nanoDSF, a tagless thermal denaturation method (Figure S6). , While ligands like TP stabilized DnaK and increased its melting temperature (T m) by ∼1.2 °C, known inhibitors quercetin and 2Cl-IB-MECA reduced the T m by <1.0 °C, suggesting minimal destabilization (Table S10). Similarly, EG35 and EG36 did not cause a noticeable change in T m (Table S11, Figure S6), likely because they interact with an allosteric and shallow cleft and do not substantially perturb the thermal stability of the protein. − In contrast, EG31 (500 μM) reduced the T m by ∼1.5 °C (Figures S6 and S8), possibly by disrupting DnaK oligomerization. However, EG31 did not affect cofactor-stimulated ATPase activity, suggesting its binding is either disrupted by cofactors or occurs outside the ATPase-regulating site.
In summary, we present ProMol_Func, a generalizable framework by integrating sequence-derived protein function embeddings as global conditioning into trainable small-molecule graph encoders to predict protein–compound binding probabilities. ProMol_Func demonstrates strong screening power across multiple benchmarks, outperforming advanced SBDD models and rivaling the state-of-the-art structure-free model BIND. While we utilized 2.60 M trainable parameters, BIND operated with 9.26 M, so ProMol_Func has lower model complexity. Incredibly, ProMol_Func finished training on 5.3 M protein–ligand pairs for 9–10 epochs during 3–4 days on a single Nvidia A100 GPU. Importantly, ProMol_Func demonstrates zero-shot learning capability by successfully identifying novel binders for E. coli DnaK. Experimental validation confirmed this performance: two predicted compounds inhibited bacterial DnaK ATPase activity, and one compound destabilized DnaK, supporting the model’s reliability in prospective applications. We predict that the two inhibitors bind an allosteric NBD site, showing micromolar affinity and partial inhibition, consistent with typical chaperone inhibitor potencies. , In the future, structural characterization will be pursued following hit-to-lead optimization. These findings highlight ProMol_Func as a scalable, structure-free tool for early stage hit discovery in virtual screening, especially for targets lacking structural or biochemical data. However, since zero-shot learning is not universally guaranteed to perform well across all targets, users should carefully assess ProMol_Func’s performance on specific systems before relying on its zero-shot predictions.
Supplementary Material
Acknowledgments
This work was supported by the funding from U.S. National Institutes of Health (R35-GM127040) to Y.Z. and an NSF Career Award to T.L. Z.F. acknowledges partial support from a graduate fellowship from the Simons Center for Computational Physical Chemistry (SCCPC) at NYU. We thank NYU-ITS for providing computational resources. A.R. acknowledges partial support from the NYU Ted Keusseff Dissertation Fellowship.Z.F. acknowledges Mia Sheshova for the nanoDSF technology protocal and Weijun Zhou for compound water solubility predictions. We also thank Gideon Yawson for the synthesis of a fluorescent peptide for these studies.
Corresponding source code is available at https://github.com/ZixuanFeng-NYU/ProMol_Func. Data and saved model checkpoints are available at 10.5281/zenodo.16825387.
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jacsau.6c00173.
Additional figures and tables for extended benchmark results and experimental supporting assays (PDF)
Zixuan Feng – conceptualization, model development, biochemical experiments, writing – original draft. Max Kim – biochemical experiments, data analysis, writing – experimental methods. Aweon Richards – experimental design, mentorship, protein expression and purification, writing – experimental methods. Tania Lupoli – supervision, experimental design, funding acquisition, writing – review and editing. Yingkai Zhang – conceptualization, methodology, supervision, funding acquisition, writing – review and editing. All authors reviewed and approved the final manuscript.
The authors declare no competing financial interest.
References
- Shoichet B. K.. Virtual screening of chemical libraries. Nature. 2004;432:862–865. doi: 10.1038/nature03197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sadybekov A. V., Katritch V.. Computational approaches streamlining drug discovery. Nature. 2023;616:673–685. doi: 10.1038/s41586-023-05905-z. [DOI] [PubMed] [Google Scholar]
- Lavecchia A., Di Giovanni C.. Virtual screening strategies in drug discovery: a critical review. Curr. Med. Chem. 2013;20:2839–2860. doi: 10.2174/09298673113209990001. [DOI] [PubMed] [Google Scholar]
- Kitchen D. B., Decornez H., Furr J. R., Bajorath J.. Docking and scoring in virtual screening for drug discovery: methods and applications. Nat. Rev. Drug Discovery. 2004;3:935–949. doi: 10.1038/nrd1549. [DOI] [PubMed] [Google Scholar]
- Bender B. J., Gahbauer S., Luttens A., Lyu J., Webb C. M., Stein R. M., Fink E. A., Balius T. E., Carlsson J., Irwin J. J.. et al. A practical guide to large-scale docking. Nature protocols. 2021;16:4799–4832. doi: 10.1038/s41596-021-00597-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Doman T. N., McGovern S. L., Witherbee B. J., Kasten T. P., Kurumbail R., Stallings W. C., Connolly D. T., Shoichet B. K.. Molecular docking and high-throughput screening for novel inhibitors of protein tyrosine phosphatase-1B. Journal of medicinal chemistry. 2002;45:2213–2221. doi: 10.1021/jm010548w. [DOI] [PubMed] [Google Scholar]
- Jenkins J. L., Kao R. Y., Shapiro R.. Virtual screening to enrich hit lists from high-throughput screening: A case study on small-molecule inhibitors of angiogenin. Proteins: Struct., Funct., Bioinf. 2003;50:81–93. doi: 10.1002/prot.10270. [DOI] [PubMed] [Google Scholar]
- Coleman R. G., Carchia M., Sterling T., Irwin J. J., Shoichet B. K.. Ligand pose and orientational sampling in molecular docking. PloS one. 2013;8:e75992. doi: 10.1371/journal.pone.0075992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morris G. M., Huey R., Lindstrom W., Sanner M. F., Belew R. K., Goodsell D. S., Olson A. J.. AutoDock4 and AutoDockTools4: Automated docking with selective receptor flexibility. Journal of computational chemistry. 2009;30:2785–2791. doi: 10.1002/jcc.21256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Trott O., Olson A. J.. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of computational chemistry. 2010;31:455–461. doi: 10.1002/jcc.21334. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rarey M., Kramer B., Lengauer T., Klebe G.. A fast flexible docking method using an incremental construction algorithm. Journal of molecular biology. 1996;261:470–489. doi: 10.1006/jmbi.1996.0477. [DOI] [PubMed] [Google Scholar]
- Friesner R. A., Banks J. L., Murphy R. B., Halgren T. A., Klicic J. J., Mainz D. T., Repasky M. P., Knoll E. H., Shelley M., Perry J. K.. et al. Glide: a new approach for rapid, accurate docking and scoring. 1. Method and assessment of docking accuracy. Journal of medicinal chemistry. 2004;47:1739–1749. doi: 10.1021/jm0306430. [DOI] [PubMed] [Google Scholar]
- Yang C., Zhang Y.. Lin_F9: a linear empirical scoring function for protein–ligand docking. J. Chem. Inf. Model. 2021;61:4630–4644. doi: 10.1021/acs.jcim.1c00737. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tingle B. I., Irwin J. J.. Large-scale docking in the cloud. J. Chem. Inf. Model. 2023;63:2735–2741. doi: 10.1021/acs.jcim.3c00031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu F., Wu C.-G., Tu C.-L., Glenn I., Meyerowitz J., Kaplan A. L., Lyu J., Cheng Z., Tarkhanova O. O., Moroz Y. S.. et al. Large library docking identifies positive allosteric modulators of the calcium-sensing receptor. Science. 2024;385:eado1868. doi: 10.1126/science.ado1868. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shen C., Zhang X., Deng Y., Gao J., Wang D., Xu L., Pan P., Hou T., Kang Y.. Boosting protein–ligand binding pose prediction and virtual screening based on residue–atom distance likelihood potential and graph transformer. J. Med. Chem. 2022;65:10691–10706. doi: 10.1021/acs.jmedchem.2c00991. [DOI] [PubMed] [Google Scholar]
- Pan X., Wang H., Zhang Y., Wang X., Li C., Ji C., Zhang J. Z.. AA-score: a new scoring function based on amino acid-specific interaction for molecular docking. J. Chem. Inf. Model. 2022;62:2499–2509. doi: 10.1021/acs.jcim.1c01537. [DOI] [PubMed] [Google Scholar]
- Lionta E., Spyrou G., Vassilatis D., Cournia Z.. Structure-based virtual screening for drug discovery: principles, applications and recent advances. Current topics in medicinal chemistry. 2014;14:1923–1938. doi: 10.2174/1568026614666140929124445. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gimeno A., Ojeda-Montes M. J., Tomás-Hernández S., Cereto-Massagué A., Beltrán-Debón R., Mulero M., Pujadas G., Garcia-Vallvé S.. The light and dark sides of virtual screening: what is there to know? International journal of molecular sciences. 2019;20:1375. doi: 10.3390/ijms20061375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gligorijević V., Renfrew P. D., Kosciolek T., Leman J. K., Berenberg D., Vatanen T., Chandler C., Taylor B. C., Fisk I. M., Vlamakis H.. et al. Structure-based protein function prediction using graph convolutional networks. Nat. Commun. 2021;12:3168. doi: 10.1038/s41467-021-23303-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A.. et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fang Y., Zhang Q., Zhang N., Chen Z., Zhuang X., Shao X., Fan X., Chen H.. Knowledge graph-enhanced molecular contrastive learning with functional prompt. Nature Machine Intelligence. 2023;5:542–553. doi: 10.1038/s42256-023-00654-0. [DOI] [Google Scholar]
- Yang K., Swanson K., Jin W., Coley C., Eiden P., Gao H., Guzman-Perez A., Hopper T., Kelley B., Mathea M.. et al. Analyzing learned molecular representations for property prediction. J. Chem. Inf. Model. 2019;59:3370–3388. doi: 10.1021/acs.jcim.9b00237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gupta R., Srivastava D., Sahu M., Tiwari S., Ambasta R. K., Kumar P.. Artificial intelligence to deep learning: machine intelligence approach for drug discovery. Molecular diversity. 2021;25:1315–1360. doi: 10.1007/s11030-021-10217-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Babine R. E., Bender S. L.. Molecular recognition of protein- ligand complexes: Applications to drug design. Chem. Rev. 1997;97:1359–1472. doi: 10.1021/cr960370z. [DOI] [PubMed] [Google Scholar]
- Hagg A., Kirschner K. N.. Open-Source Machine Learning in Computational Chemistry. J. Chem. Inf. Model. 2023;63:4505. doi: 10.1021/acs.jcim.3c00643. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vatansever S., Schlessinger A., Wacker D., Kaniskan H. Ü., Jin J., Zhou M.-M., Zhang B.. Artificial intelligence and machine learning-aided drug discovery in central nervous system diseases: State-of-the-arts and future directions. Medicinal research reviews. 2021;41:1427–1473. doi: 10.1002/med.21764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Catacutan D. B., Alexander J., Arnold A., Stokes J. M.. Machine learning in preclinical drug discovery. Nat. Chem. Biol. 2024;20:960–973. doi: 10.1038/s41589-024-01679-1. [DOI] [PubMed] [Google Scholar]
- Vamathevan J., Clark D., Czodrowski P., Dunham I., Ferran E., Lee G., Li B., Madabhushi A., Shah P., Spitzer M.. et al. Applications of machine learning in drug discovery and development. Nat. Rev. Drug Discovery. 2019;18:463–477. doi: 10.1038/s41573-019-0024-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gentile F., Yaacoub J. C., Gleave J., Fernandez M., Ton A.-T., Ban F., Stern A., Cherkasov A.. Artificial intelligence–enabled virtual screening of ultra-large chemical libraries with deep docking. Nat. Protoc. 2022;17:672–697. doi: 10.1038/s41596-021-00659-2. [DOI] [PubMed] [Google Scholar]
- Volkov M., Turk J.-A., Drizard N., Martin N., Hoffmann B., Gaston-Mathé Y., Rognan D.. On the frustration to predict binding affinities from protein–ligand structures with deep neural networks. Journal of medicinal chemistry. 2022;65:7946–7958. doi: 10.1021/acs.jmedchem.2c00487. [DOI] [PubMed] [Google Scholar]
- Sundar V., Colwell L.. The effect of debiasing protein–ligand binding data on generalization. J. Chem. Inf. Model. 2020;60:56–62. doi: 10.1021/acs.jcim.9b00415. [DOI] [PubMed] [Google Scholar]
- Vignaux P. A., Minerali E., Foil D. H., Puhl A. C., Ekins S.. Machine learning for discovery of GSK3β inhibitors. ACS omega. 2020;5:26551–26561. doi: 10.1021/acsomega.0c03302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xing J., Lu W., Liu R., Wang Y., Xie Y., Zhang H., Shi Z., Jiang H., Liu Y.-C., Chen K.. et al. Machine-learning-assisted approach for discovering novel inhibitors targeting bromodomain-containing protein 4. J. Chem. Inf. Model. 2017;57:1677–1690. doi: 10.1021/acs.jcim.7b00098. [DOI] [PubMed] [Google Scholar]
- Yang S., Li S., Chang J.. Discovery of Cobimetinib as a novel A-FABP inhibitor using machine learning and molecular docking-based virtual screening. RSC Adv. 2022;12:13500–13510. doi: 10.1039/D2RA01057G. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lam H. Y. I., Guan J. S., Ong X. E., Pincket R., Mu Y.. Protein language models are performant in structure-free virtual screening. Briefings in Bioinformatics. 2024;25:bbae480. doi: 10.1093/bib/bbae480. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., Smetanin N., Verkuil R., Kabeli O., Shmueli Y.. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
- Hosfelt J., Richards A., Zheng M., Adura C., Nelson B., Yang A., Fay A., Resager W., Ueberheide B., Glickman J. F.. et al. An allosteric inhibitor of bacterial Hsp70 chaperone potentiates antibiotics and mitigates resistance. Cell Chemical Biology. 2022;29:854–869. doi: 10.1016/j.chembiol.2021.11.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Soper N., Yardumian I., Chen E., Yang C., Ciervo S., Oom A. L., Desvignes L., Mulligan M. J., Zhang Y., Lupoli T. J.. A Repurposed Drug Interferes with Nucleic Acid to Inhibit the Dual Activities of Coronavirus Nsp13. ACS Chem. Biol. 2024;19:1593–1603. doi: 10.1021/acschembio.4c00244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Buskirk A. R., Liu D. R.. Creating small-molecule-dependent switches to modulate biological functions. Chemistry & biology. 2005;12:151–161. doi: 10.1016/j.chembiol.2004.11.012. [DOI] [PubMed] [Google Scholar]
- Roskoski R. Jr. A historical overview of protein kinases and their targeted small molecule inhibitors. Pharmacol. Res. 2015;100:1–23. doi: 10.1016/j.phrs.2015.07.010. [DOI] [PubMed] [Google Scholar]
- Chène P.. ATPases as drug targets: learning from their structure. Nat. Rev. Drug Discovery. 2002;1:665–673. doi: 10.1038/nrd894. [DOI] [PubMed] [Google Scholar]
- Gilson M. K., Liu T., Baitaluk M., Nicola G., Hwang L., Chong J.. BindingDB in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic acids research. 2016;44:D1045–D1053. doi: 10.1093/nar/gkv1072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhu T., Cao S., Su P.-C., Patel R., Shah D., Chokshi H. B., Szukala R., Johnson M. E., Hevener K. E.. Hit identification and optimization in virtual screening: Practical recommendations based on a critical literature analysis: Miniperspective. Journal of medicinal chemistry. 2013;56:6560–6572. doi: 10.1021/jm301916b. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tran-Nguyen V.-K., Jacquemard C., Rognan D.. LIT-PCBA: an unbiased data set for machine learning and virtual screening. J. Chem. Inf. Model. 2020;60:4263–4273. doi: 10.1021/acs.jcim.0c00155. [DOI] [PubMed] [Google Scholar]
- Kim S., Chen J., Cheng T., Gindulyte A., He J., He S., Li Q., Shoemaker B. A., Thiessen P. A., Yu B.. et al. PubChem 2019 update: improved access to chemical data. Nucleic acids research. 2019;47:D1102–D1109. doi: 10.1093/nar/gky1033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gu Y., Xia S., Ouyang Q., Zhang Y.. Bioactivity Deep Learning for Complex Structure-Free Compound-Protein Interaction Prediction. J. Chem. Inf. Model. 2025;65:9910. doi: 10.1021/acs.jcim.5c00741. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang C., Zhang Y.. Delta Machine Learning to Improve Scoring-Ranking-Screening Performances of Protein–Ligand Scoring Functions. J. Chem. Inf. Model. 2022;62:2696–2712. doi: 10.1021/acs.jcim.2c00485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mysinger M. M., Carchia M., Irwin J. J., Shoichet B. K.. Directory of useful decoys, enhanced (DUD-E): better ligands and decoys for better benchmarking. Journal of medicinal chemistry. 2012;55:6582–6594. doi: 10.1021/jm300687e. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bauer M. R., Ibrahim T. M., Vogel S. M., Boeckler F. M.. Evaluation and optimization of virtual screening workflows with DEKOIS 2.0–a public library of challenging docking benchmark sets. J. Chem. Inf. Model. 2013;53:1447–1462. doi: 10.1021/ci400115b. [DOI] [PubMed] [Google Scholar]
- Wójcikowski M., Ballester P. J., Siedlecki P.. Performance of machine-learning scoring functions in structure-based virtual screening. Sci. Rep. 2017;7:46710. doi: 10.1038/srep46710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Genest O., Wickner S., Doyle S. M.. Hsp90 and Hsp70 chaperones: Collaborators in protein remodeling. J. Biol. Chem. 2019;294:2109–2120. doi: 10.1074/jbc.REV118.002806. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schopf F. H., Biebl M. M., Buchner J.. The HSP90 chaperone machinery. Nat. Rev. Mol. Cell Biol. 2017;18:345–360. doi: 10.1038/nrm.2017.20. [DOI] [PubMed] [Google Scholar]
- Radli M., Rüdiger S. G.. Dancing with the diva: Hsp90–client interactions. Journal of molecular biology. 2018;430:3029–3040. doi: 10.1016/j.jmb.2018.05.026. [DOI] [PubMed] [Google Scholar]
- Whitesell L., Lindquist S. L.. HSP90 and the chaperoning of cancer. Nature Reviews Cancer. 2005;5:761–772. doi: 10.1038/nrc1716. [DOI] [PubMed] [Google Scholar]
- Chang L., Miyata Y., Ung P. M., Bertelsen E. B., McQuade T. J., Carlson H. A., Zuiderweg E. R., Gestwicki J. E.. Chemical screens against a reconstituted multiprotein complex: myricetin blocks DnaJ regulation of DnaK through an allosteric mechanism. Chemistry & biology. 2011;18:210–221. doi: 10.1016/j.chembiol.2010.12.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang L., Bertelsen E. B., Wisén S., Larsen E. M., Zuiderweg E. R., Gestwicki J. E.. High-throughput screen for small molecules that modulate the ATPase activity of the molecular chaperone DnaK. Analytical biochemistry. 2008;372:167–176. doi: 10.1016/j.ab.2007.08.020. [DOI] [PubMed] [Google Scholar]
- Cesa L. C., Patury S., Komiyama T., Ahmad A., Zuiderweg E. R., Gestwicki J. E.. Inhibitors of difficult protein–protein interactions identified by high-throughput screening of multiprotein complexes. ACS Chem. Biol. 2013;8:1988–1997. doi: 10.1021/cb400356m. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Evans C. G., Chang L., Gestwicki J. E.. Heat shock protein 70 (hsp70) as an emerging drug target. Journal of medicinal chemistry. 2010;53:4585–4602. doi: 10.1021/jm100054f. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Murphy M. E.. The HSP70 family and cancer. Carcinogenesis. 2013;34:1181–1188. doi: 10.1093/carcin/bgt111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sherman M. Y., Gabai V. L.. Hsp70 in cancer: back to the future. Oncogene. 2015;34:4153–4161. doi: 10.1038/onc.2014.349. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lupoli T. J., Vaubourgeix J., Burns-Huang K., Gold B.. Targeting the proteostasis network for mycobacterial drug discovery. ACS infectious diseases. 2018;4:478–498. doi: 10.1021/acsinfecdis.7b00231. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu C.-C., Naveen V., Chien C.-H., Chang Y.-W., Hsiao C.-D.. Crystal structure of DnaK protein complexed with nucleotide exchange factor GrpE in DnaK chaperone system: insight into intermolecular communication. J. Biol. Chem. 2012;287:21461–21470. doi: 10.1074/jbc.M112.344358. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bertelsen E. B., Chang L., Gestwicki J. E., Zuiderweg E. R.. Solution conformation of wild-type E. coli Hsp70 (DnaK) chaperone complexed with ADP and substrate. Proc. Natl. Acad. Sci. U. S. A. 2009;106:8471–8476. doi: 10.1073/pnas.0903503106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rossi M.-A., Pozhidaeva A. K., Clerico E. M., Petridis C., Gierasch L. M.. New insights into the structure and function of the complex between the Escherichia coli Hsp70, DnaK, and its nucleotide-exchange factor, GrpE. J. Biol. Chem. 2024;300:105574. doi: 10.1016/j.jbc.2023.105574. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kityk R., Kopp J., Sinning I., Mayer M. P.. Structure and dynamics of the ATP-bound open conformation of Hsp70 chaperones. Molecular cell. 2012;48:863–874. doi: 10.1016/j.molcel.2012.09.023. [DOI] [PubMed] [Google Scholar]
- Massey A. J.. ATPases as drug targets: insights from heat shock proteins 70 and 90. Journal of medicinal chemistry. 2010;53:7280–7286. doi: 10.1021/jm100342z. [DOI] [PubMed] [Google Scholar]
- Chen I.-J., Hubbard R. E.. Lessons for fragment library design: analysis of output from multiple screening campaigns. Journal of computer-aided molecular design. 2009;23:603–620. doi: 10.1007/s10822-009-9280-5. [DOI] [PubMed] [Google Scholar]
- Gestwicki J. E., Shao H.. Inhibitors and chemical probes for molecular chaperone networks. J. Biol. Chem. 2019;294:2151–2161. doi: 10.1074/jbc.TM118.002813. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tetko I. V., Novotarskyi S., Sushko I., Ivanov V., Petrenko A. E., Dieden R., Lebon F., Mathieu B.. Development of dimethyl sulfoxide solubility models using 163 000 molecules: using a domain applicability metric to select more reliable predictions. J. Chem. Inf. Model. 2013;53:1990–2000. doi: 10.1021/ci400213d. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hartl F. U., Hayer-Hartl M.. Molecular chaperones in the cytosol: from nascent chain to folded protein. Science. 2002;295:1852–1858. doi: 10.1126/science.1068408. [DOI] [PubMed] [Google Scholar]
- Shao H., Gestwicki J. E.. Neutral analogs of the heat shock protein 70 (Hsp70) inhibitor, JG-98. Bioorganic & medicinal chemistry letters. 2020;30:126954. doi: 10.1016/j.bmcl.2020.126954. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Qiu X. B., Shao Y. M., Miao S., Wang L.. The diversity of the DnaJ/Hsp40 family, the crucial partners for Hsp70 chaperones. Cellular and Molecular Life Sciences CMLS. 2006;63:2560–2570. doi: 10.1007/s00018-006-6192-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Harrison C.. GrpE, a nucleotide exchange factor for DnaK. Cell stress & chaperones. 2003;8:218. doi: 10.1379/1466-1268(2003)008<0218:GANEFF>2.0.CO;2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang L., Thompson A. D., Ung P., Carlson H. A., Gestwicki J. E.. Mutagenesis reveals the complex relationships between ATPase rate and the chaperone activities of Escherichia coli heat shock protein 70 (Hsp70/DnaK) J. Biol. Chem. 2010;285:21282–21291. doi: 10.1074/jbc.M110.124149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liberek K., Marszalek J., Ang D., Georgopoulos C., Zylicz M.. Escherichia coli DnaJ and GrpE heat shock proteins jointly stimulate ATPase activity of DnaK. Proc. Natl. Acad. Sci. U. S. A. 1991;88:2874–2878. doi: 10.1073/pnas.88.7.2874. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McGovern S. L., Helfand B. T., Feng B., Shoichet B. K.. A specific mechanism of nonspecific inhibition. Journal of medicinal chemistry. 2003;46:4265–4272. doi: 10.1021/jm030266r. [DOI] [PubMed] [Google Scholar]
- Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A. J., Bambrick J.. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Passaro S., Corso G., Wohlwend J., Reveiz M., Thaler S., Ram Somnath V., Getz N., Portnoi T., Roy J., Stark H.. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. BioRxiv. 2025 doi: 10.1101/2025.06.14.659707. [DOI] [Google Scholar]
- Berisio R., Barra G., Napolitano V., Privitera M., Romano M., Squeglia F., Ruggiero A.. HtpGA Major Virulence Factor and a Promising Vaccine Antigen against Mycobacterium tuberculosis. Biomolecules. 2024;14:471. doi: 10.3390/biom14040471. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu T., Hornsby M., Zhu L., Yu J. C., Shokat K. M., Gestwicki J. E.. Protocol for performing and optimizing differential scanning fluorimetry experiments. STAR protocols. 2023;4:102688. doi: 10.1016/j.xpro.2023.102688. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Magnusson A. O., Szekrenyi A., Joosten H.-J., Finnigan J., Charnock S., Fessner W.-D.. nanoDSF as screening tool for enzyme libraries and biotechnology development. FEBS journal. 2019;286:184–204. doi: 10.1111/febs.14696. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Taylor I. R., Assimon V. A., Kuo S. Y., Rinaldi S., Li X., Young Z. T., Morra G., Green K., Nguyen D., Shao H.. et al. Tryptophan scanning mutagenesis as a way to mimic the compound-bound state and probe the selectivity of allosteric inhibitors in cells. Chemical Science. 2020;11:1892–1904. doi: 10.1039/C9SC04284A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vivoli M., Novak H. R., Littlechild J. A., Harmer N. J.. Determination of protein-ligand interactions using differential scanning fluorimetry. Journal of visualized experiments: JoVE. 2014:51809. doi: 10.3791/51809-v. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Real-Hohn A., Groznica M., Löffler N., Blaas D., Kowalski H.. nanoDSF: in vitro label-free method to monitor picornavirus uncoating and test compounds affecting particle stability. Frontiers in Microbiology. 2020;11:1442. doi: 10.3389/fmicb.2020.01442. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shao H., Oltion K., Wu T., Gestwicki J. E.. Differential scanning fluorimetry (DSF) screen to identify inhibitors of Hsp60 protein–protein interactions. Organic & Biomolecular Chemistry. 2020;18:4157–4163. doi: 10.1039/D0OB00928H. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Corresponding source code is available at https://github.com/ZixuanFeng-NYU/ProMol_Func. Data and saved model checkpoints are available at 10.5281/zenodo.16825387.


