Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Mar 20;148(12):13043–13054. doi: 10.1021/jacs.5c22222

Discovery of Covalent Ligands with AlphaFold3

Yoav Shamir , Ronen Gabizon , Adi Rogel , David Yin-wei Lin , Amy H Andreotti , Nir London †,*
PMCID: PMC13047693  PMID: 41857796

Abstract

Covalent inhibitors are a prominent modality for research and therapeutic tools. However, a scarcity of computational methods for their discovery slows progress in this field. AI models such as AlphaFold3 (AF3) have shown accuracy in ligand pose prediction, but their applicability for virtual screening campaigns was not assessed. We show that AF3 cofolding predictions and an associated predicted confidence metric ranks true covalent binders with near-optimal classification over property-matched decoys, significantly outperforming state-of-the-art covalent docking tools for a set of protein kinases. In a prospective virtual screening campaign against the model kinase BTK, we discovered a chemically distinct, novel, covalent small molecule that displays potent inhibition in vitro and in cells while maintaining marked kinome and proteomic selectivity. Co-crystallography validated the subangstrom accuracy of the predicted AF3 binding mode. These results demonstrate that AF3 can be practically used to discover novel chemical matter for kinases, one of the most prolific families of drug targets.


graphic file with name ja5c22222_0006.jpg


graphic file with name ja5c22222_0005.jpg

Introduction

Covalent acting small molecules represent new opportunities in chemical biology and drug discovery. Such ligands typically display increased potency, selectivity and prolonged target engagement compared with their noncovalent counterparts. Advances in chemical biology have mitigated early concerns regarding off-target reactivity and promiscuous irreversible binding and enabled rational and successful design processes. This has led to the discovery of drugs for decades-long challenging targets such as K-Ras. Currently more than 50 covalent drugs have been approved by the FDA. Recent examples include Nirmatrelvir for COVID-19, Ritlectinib for alopecia, and Adagrasib for cancer, with many others undergoing Phase III clinical trials.

Covalent binding is a two-step process, including the formation of a reversible, noncovalent complex (dominated by molecular recognition), followed by the formation of a covalent, typically irreversible bond, with kinetics that are determined by the intrinsic reactivities and relative orientation of the electrophile-nucleophile pair.

Several computational tools were developed for virtual screening of covalent libraries. Given a protein structure or model as input, these docking tools employ physics-based scoring functions to predict and energetically score the protein-bound pose of each ligand. However, covalent docking software typically ranks ligands in their bound (adduct) state and neglects to model the kinetics of covalent binding, ignoring both the reactivities of the reactants, as well as the orientation of the prereacted, or intermediate step of the covalent reaction. Nevertheless, multiple covalent binders targeting a range of proteins have been developed and experimentally validated based on the results of such in silico screening. ,− Several covalent data sets of experimental structures have been curated, enabling the evaluation of binding-pose prediction accuracy of docking tools. However, for virtual screening and practical applications, the ability of docking tools to rank a compound library such that active compounds are at the top is arguably more important than accurate pose recapitulation. While data sets designed to evaluate such enrichment are prevalent for noncovalent docking, we are unaware of any for the covalent domain.

AlphaFold3 (AF3), the latest AI-based all-atom structure prediction model from Google DeepMind, was released for free use in November 2024. AF3 facilitates high-accuracy prediction of biomolecular complexes, including the prediction of covalent protein–ligand complexes. To probe the enrichment performance of AF3 in the covalent domain, we constructed COValid, a first-of-its-kind benchmark set for the enrichment analysis of covalent virtual screening. We show that ranking the AF3-predicted compounds using a physics-based scoring function significantly outperforms the ranking produced by classical covalent docking tools. Strikingly, we find an AF3-predicted confidence metric that significantly outperforms all physics-based methods tested across all protein targets in COValid, yielding exceptionally high success rates in identifying active compounds. To mitigate concerns about training data leakage, we performed what is to our knowledge the first prospective covalent virtual screen with AF3, and identified novel, potent and selective BTK covalent inhibitors, that are active in cells.

Results

COValidA Novel Benchmark for Covalent Virtual Screening

Since cysteines are the nucleophile most commonly targeted by covalent binders, we focused our analysis on the most common electrophile that targets cysteinesacrylamide (Figure S1). As mentioned, almost all current docking methods neglect to account for the intrinsic reactivity of the docked electrophile, deeming the ranking of multiwarhead covalent libraries infeasible. Therefore, we avoided including additional cysteine-targeting electrophiles in our benchmark.

We curated compounds with experimental activity annotations from two databases – BindingDB and ChEMBL. The majority of targets we identified were protein kinases, likely due to the intense interest in covalent kinase inhibitors, ,, as they provide superior selectivity and potency for this widespread protein family. We included eight of the most populated kinases in COValid (Figure A) as well as K-RasG12C, a prime target for covalent inhibitor design. , For each target-ligand pair we also annotated the target cysteine position. We made sure that for each target the benchmark included structures with both a covalent and noncovalent binder (selected from the PDB; Data set S1).

1.

1

COValid enables optimization of covalent docking algorithms. (A) Composition of the COValid set. (B) Decoy design schemethe electrophile (E) is removed from the covalent “active”, next, property-matching, topologically dissimilar noncovalent decoys are collected for all actives. Finally, 36–50 decoy protomers are matched to each active protomer by attaching an acrylamide to free amines of the noncovalent decoys and ensuring property-matching to the active. (C) Property distributions of the active and decoy protomers (colored cyan and beige, respectively) for the six properties matched during decoy design. (D) Example of using COValid for parameter optimization. The angle range parameter of DOCKovalent defines a range Δ around a user-defined angle (109.5°), such that two bonds angles (angles a and b, defined by Cysteine Cβ, Sγ, and the acrylamide Cβ atoms, and by the Cysteine Sγ, acrylamide Cβ and Cα atoms) are sampled ± Δ during docking. Sampling with ±5° and further with ±10° leads to higher averaged adjusted LogAUC with statistical significance (p-values are 0.0002 and 0.004, respectively), but further increasing the exhaustiveness to Δ = 20° yields no significant improvement (p-value: 0.1053), while requiring significantly longer runtime. (E) Comparison of enrichment for the best identified configuration of the three covalent docking tools, docked to the ten covalent PDB structures.

Following the curation of active ligands, we generated 36–50 property-matched, acrylamide-bearing decoys for each ligand (see Supplementary Methods), to represent chemically feasible but likely inactive compounds. This computational design of decoys is necessary since annotations of “true” inactive compounds are scarce. The decoy generation protocol (Figure B) followed two design principles borrowed from the popular DUD-E benchmark set for noncovalent enrichment analysis. Firstthe physicochemical properties of the designed decoys should match those of the active. This mitigates the risk that different property distributions for the actives and decoys would yield artificially high enrichment owing to property mismatch (Figure C, Data set S2). The second principle - the chemical topology of the decoys should be dissimilar to that of the actives. We therefore scrambled the connectivity of the compounds to generate decoys that are compositionally matched to actives but unlikely to conserve favorable interactions with the target. Finally, the decoy scaffolds are selected from ZINC20 representing commercially available compounds to circumvent the risk of generating nonrealistic molecules. In total, COValid includes 874 active protomers and 37,919 decoy protomers, across ten cysteine attachment sites from nine protein targets (Data set S3).

COValid Enables Comparison and Optimization of Physics-Based Covalent Docking Tools

We evaluated three physics-based covalent docking tools using COValid: the DOCKovalent method based on DOCK3.7, the attach-and-grow method based on DOCK6.12, and the flexible side chain method of AutoDock. To quantify the enrichment performance, we used the adjusted LogAUC metric, which emphasizes early enrichment in ranking hit lists. In a real-life scenario it is more important to have active molecules at the top fraction of the hit list rather than overall better ranking across the entire library. An adjusted LogAUC of 0% equals random performance, whereas 85.5% is optimal performance, ranking all actives higher than all inactives.

Each docking tool offers many adjustable parameters and configurations that can be nontrivial to optimize and can result in significant increases in run-time. Since evaluating all different combinations of parameters is an intractable combinatorial problem, we chose several representative features for each software and assessed their effect on accuracy and runtime (see Figure D for one example, and the Supplementary Results for the full analysis).

The performance of leading configurations of DOCK6 and DOCKovalent were similar (Figure E; Table ; see docking parameters in Table S1). AutoDock, with a similar run-time showed an average enrichment slightly lower than that of DOCKovalent and DOCK6 (p-values of 0.05 and 0.06, respectively). For the three docking tools, the differences in the enrichment across the covalent-bound structures, compared to the noncovalent structures, were insignificant (p-values >0.1), indicating that noncovalent holo structures could be useful for covalent virtual screening when a covalent complex structure is not available.

1. Performance over the COValid Set.

  DOCKovalent DOCK6 AutoDock AF3 (Rosetta) AF3 (mPAE)
Average Adj. LocAUC (%) 9.9 ± 4.8 12.1 ± 9.9 5.7 ± 7.9 28.5 ± 16.1 71.8 ± 5.9
Average run-time per ligand (s) 8.4 ± 2.4 7.0 ± 2.0 6.7 ± 0.4 266 ± 114
a

The exact docking configurations corresponding to these runs are in Table S1.

b

Run-time does not include Rosetta rescoringonly AF3 modeling per ligand, run-time varied based on GPU used. The initial step (multiple sequence alignment generation and curation of template structures, using 20 allocated CPUs) took on average 22 ± 8 min per protein.

Physics-Based Scoring of AlphaFold3 Models Outperforms Docking Tools for Multiple Targets

AF3 is the latest AI model from DeepMind for all-atom prediction of biomolecular structures. It enables prediction of noncovalent as well as covalent protein–ligand complexes. On a pose-prediction benchmark set of covalent complexes curated by the AF3 authors, the model achieved a success rate (defined as pocket-aligned ligand RMSD < 2 Å) of 78.5%. To our knowledge AF3 performance in the context of covalent virtual screening (or in fact for virtual screening in general) has not been reported. We set out to use COValid to assess its capacity for screening. We used the prediction pipeline of AF3 to predict covalent complexes for all compounds in COValid using the protein sequence and geometrically optimized 3D ligand conformer as input. The input also includes specification of the ligand atom and the protein side-chain atom to be covalently bonded.

We sought to rank the predicted complex structures using a physics-based scoring function. We used Rosetta to minimize and energetically score all the predicted AF3 models and used this score to calculate the enrichment performance across COValid. Rescoring of AF3 models yielded significantly better results for five out of the ten sites (Figure S2) and an overall better average adjusted LogAUC across the ten sites (Table ). The fact that this solely based on the models produced by AF3 (and not by any AI-based metric) suggests the models are sufficiently realistic to differentiate true binding modes from artificial ones. Indeed, in the few cases where a crystal structure was available for a target-ligand pair from the benchmark, the AF3 models proved accurate (Figure S3). While most of these examples were likely included in the AF3 training set, one structure was determined after the training set cutoff, and in this case as well AF3 predicted it accurately (0.45Å pocket-aligned RMSD; PDB: 7O70, Figure A).

2.

2

Near-optimal enrichment of true inhibitors using AF3 mPAE. (A) Example of an accurate AF3 prediction of a covalent complex (K-RasG12C with Mg2+ and GDP, PDB 7O70, pocket-aligned RMSD 0.45 Å) not included in the training set (the experimental structure and the predicted model are colored in gray and purple, respectively). (B) Comparison of the two best physics-based covalent docking tools’ enrichment to AF3 structural predictions ranked by mPAE. (C) The distribution of the proportion of active COValid protomers, according to their AF3 mPAE element. Each column indicates the share of active protomers out of all protomers within the indicated mPAE value range. This could be interpreted as the probability for being “active” given a mPAE value. (D) AF3 mPAE enrichment performance on DUDE-Z targets. Across all 43 DUDE-Z targets, we compare the adjusted LogAUC values kindly provided by Balius et al. for DOCK3.7 and DOCK6.9 to values derived from our AF3 predictions, ranked by mPAE (actives are sorted prior to enrichment calculations, using only the best-scoring active protomer of each compound). Protein sequence input for the data step of AF3 was extracted from the DUDE-Z PDB files, and the noncovalent ligands were provided as SMILES strings from DUDE-Z.

AF3 Minimal Predicted Aligned Error Is a Superior Metric for Virtual Screening Enrichment

Along with the 3D coordinates of the predicted molecular structure, the AF3 model outputs predicted confidence metrics for the accuracy of the pose recapitulation. We were interested in the ability of such metrics to rank compound libraries and enrich actives, even though these metrics were not trained to predict activity or affinity. We evaluated the adjusted LogAUC values over the COValid set for various AF3-predicted confidence metrics (see definitions in Supplementary Methods and results in Figure S4). When attempting to rank COValid AF3 models by these metrics, we noticed that most metrics lack the resolution required for a decisive ranking of the lists. In several cases, thousands of protomers were ranked with the same exact predicted value (Figure S5). Nevertheless, to probe their enrichment performance, we evaluated the adjusted LogAUC value for each cysteine site as a range spanning two extremes: we assessed a best-case scenario, where we rank actives higher than decoys for the same metric value, and a worst-case scenario, where decoys are ranked higher than actives of the same value.

Ranking by a global metric (pTM) or protein-based metric (protein chain pTM) yielded large ranges between the best- and worst-case enrichment values (Figure S4). This lack of resolution is perhaps in line with the fact that the protein is constant for any specific target. Taking into account the protein–ligand interface (ipTM and “ranking score”), yielded smaller, more informative enrichment ranges for some targets, while other targets still displayed a wide range of enrichment values. The metric that yielded the most decisive adjusted LogAUC ranges of this group was the ligand chain pTM. Even when using the worst-case ranking, AF3 predictions sorted by the ligand chain pTM metrics significantly outperformed the physics-based docking tools across nine of the ten COValid Cys sites (Figure S6).

Finally, we studied the relevance of the predicted aligned error matrix (PAE), a structural confidence metric provided by AF3. In the internal representation used by the AF3 model, each protein residue and each ligand atom are represented by a single token. Each (i,j) element of the PAE matrix contains a prediction of the error in the positions of the AF3 token j when aligned to the ground-truth based on token i. We define mPAE as the minimal predicted error in the positions of the protein residues, when aligned to the ground-truth complex with respect to the ligand atoms (see details in Supplementary Methods; this value is provided as output by AF3).

Compared with other AF3-predicted confidence metrics, mPAE yields much better resolution between compounds, enabling conclusive ranking of all ten COValid compound lists. Remarkably, AF3 modeling followed by ranking by mPAE, significantly outperforms other metrics including all covalent docking tools across all ten cases (Figure B; Figure S4). The worst-case adjusted LogAUC values ranged from 56.4% for BMX to 83.8% for Cys477 of FGFR4, approaching optimal classification. Comparing the enrichment performance to that of the physics-based covalent docking tools, the differences are striking (Figure B). If we disregard early enrichment and report the more typical “AUC” (area under the receiver operating curve), the values range from 0.9082 to 0.9996 (Table S2).

We also studied the relevance of the absolute mPAE values for prospective covalent screening. Within the applicability domain of COValid, mPAE enables assignment of confidence to the classification of a compound as active (Figure C). Across all COValid protomers, 96.6% of the protomers with the lowest mPAE values (<0.85 Å) correspond to actives. As mPAE values increase (indicating a decrease in structural confidence) the probability that the protomer corresponds to an active binder drops monotonically, decreasing to less than 13% at mPAE values higher than 1.05Å.

We considered various confounding factors, such as the possibility of an overlap between the AF3 training set and the COValid compounds. If the model was trained on some COValid “actives”, it could potentially predict higher structural confidence for these compared to the designed decoys. We conducted ablation studies by removing active compounds with high topological similarity to any compound in the PDB (along with their corresponding decoys; Figure S7) and recalculated the adjusted LogAUC values in the absence of the removed protomers (Figure S8). Even at strict similarity cutoffs (Tanimoto coefficient = 0.4, Morgan fingerprint, radius = 2, 2048 bits) that remove most of the active and decoy protomers from the ranked lists, enrichment remains high, with little effect due to ablation. This suggests that the performance of mPAE is not due to overlap with the training data. In addition, analysis of the mPAE element of all the COValid “actives” against their maximal Tanimoto similarity to any PDB yielded no correlation between the two metrics (Pearson correlation coefficient −0.056, p-value 0.21; Figure S9).

Another potential confounding factor is that our decoys, despite matching in physicochemical properties and being based on commercially available scaffolds, differ in some rudimentary way from the “actives”. To address this, we searched ChEMBL 33 for experimental decoysacrylamide-bearing structures that were tested against benchmark targets but showed worse than 10 μM activity (Figure S10). We found 64 experimental decoys across six COValid kinases. We used AF3 to predict their covalent complexes and evaluated their mPAE values. The average mPAE was 1.5 ± 0.5 Å, with 86% of the decoys yielding a mPAE higher than 0.95 Å.

Recent work suggested that cofolding models do not learn the “physics” of ligand binding, and demonstrated their insensitivity to local mutations in the binding site. We found a similar trend. First, there is no apparent correlation between mPAE and affinity (Figure S11) suggesting it is a “coarse” metric useful primarily for classification. Looking further into selectivity, we conducted cross-docking experiments and calculated the enrichment for each kinase active and decoy set against noncognate kinases. (). While generally the highest enrichment is still achieved against the cognate kinase, clear clusters appear, in which kinases with a similarly positioned cysteine residue perform well in enriching noncognate ligand sets. To some extent this could be a result of nonselectivity of the active ligands, which are known to bind kinase off-targets with analogous cysteines. Overall it likely suggests that AF3 and mPAE are currently not suitable for selectivity prediction.

mPAE Performs Well Also for Noncovalent Virtual Screening

We tested the relevance of AF3 predictions with mPAE ranking for the noncovalent case using the DUDE-Z benchmark set. Across 42 out of 43 targets, AF3-based enrichment significantly outperformed the results reported for either DOCK3.7 or DOCK6.9 (Figure D; the exception is CXCR4, which yields near-random enrichment across all three screening methods). Across this diverse set, we observe a larger variability in enrichment values than in COValid, with an average enrichment of 50.9% ± 20.0%, suggesting some targets are more challenging than others. Indeed, kinases exhibited higher mPAE-based enrichment on average, when compared to other targets (11 kinases yielded an average adjusted LogAUC of 60.7% ± 11.1%, compared to 47.5% ± 21.3% on average for the other 32 protein targets). Across this more diverse set we do see a moderate correlation between performance and the average maximal chemical similarity of the active compounds to PDB ligands (Pearson correlation coefficient = 0.425, p-value = 0.00451, R2 = 0.18; Figure S13). Boltz-2, another AI based cofolding model showed similar performance to AF3 on this set (see Supplementary Results; Figure S14).

Covalent Prospective Screening Identifies Novel Kinase Inhibitors

To evaluate AF3-based covalent screening and probe its relevance for the discovery of novel covalent binders, we conducted a prospective campaign against Bruton’s tyrosine kinase (BTK), a well-studied kinase target, whose dysregulation is implicated in B-cell malignancies. BTK has multiple FDA-approved covalent inhibitors targeting Cys481 at the ATP binding pocket. We constructed a diverse virtual library of ∼906K acrylamide-bearing compounds and used AF3 to predict their covalent complexes with Cys481 of BTK (see Supplementary Methods; Data set S4). We ranked the predictions by their mPAE scores and filtered out compounds with mPAE greater than 0.9Å (resulting in 440 compounds; ∼0.05% of the library). In search of novel binders, we filtered out 50 potential BTK inhibitors showing even remote chemical similarity to known BTK binders (Tc > 0.35, Morgan fingerprint, radius = 2, 2048 bits). The remaining 390 compounds were clustered and manually inspected in search of compounds with diverse binding poses. Of these, we synthesized 13 compounds for experimental evaluation (Figure S15).

Intact protein LC/MS experiments indicated that three of the 13 compounds reached near-100% covalent labeling of BTK within 2 h (Figure A,B; Figures S16 and S17). We further characterized the three main hitsYS1, YS2, and YS3, and used ibrutinib, an FDA-approved drug for BTK, as a control (Figure C). We note that these compounds bear no chemical resemblance to any ligands across all of ChEMBL (Figure S18). Dose response and time course experiments indicated that YS1 covalently labels BTK rapidly, even in low concentrations (Figure D,E; in a 2 h dose response assay, ibrutinib and YS1 yield >90% labeling at 1 μM). To mitigate concerns that labeling arises from intrinsic reactivity of the molecules rather than target recognition, we conducted GSH reactivity assays, which indicated that binding is not driven by intrinsic electrophilicityall three compounds are less reactive than ibrutinib (YS1 and YS2 are significantly less reactive, Figure F). Differential scanning fluorimetry (DSF) experiments showed that all three compounds stabilize BTK, with YS1 leading to the largest shift in melting point (∼9 °C; Figure G). In vitro kinase activity assays identified YS1 as a potent inhibitor of BTK, (IC50 = 30 nM; Figure H; IC50 values for YS2 and YS3 are 77 μM and 8.4 μM, respectively). YS1 showed no activity with the C481S mutant of BTK (Figure S19), indicating that its binding is driven by the covalent interaction. We used a spectral shift time-course assay to kinetically profile YS1-YS3 as well as tirabrutinib (Figure S44). For YS1, we were able to determine both the specificity constant (k inact/KI = 915.85 M–1s–1; ) as well as individual components (k inact = 0.004887 s–1; KI = 5.34 μM). YS2 showed slow binding kinetics, and we were only able to extrapolate its k inact/KI ratio (229.8M–1s–1; Figure S44C). YS3 displayed autofluorescence which precluded its kinetic profiling using the spectral shift-based assay.

3.

3

Experimental characterization of novel BTK inhibitors discovered by prospective screening. (A) Intact protein LCMS covalent labeling percentage for 13 synthesized compounds (1 μM BTK, 200 μM compound, pH 7.5, RT, 2 h). n = 2, data represented as mean with SD as error bars. (B) Deconvoluted LC/MS spectrum of BTK incubated with YS1. The mass difference corresponds to the ligand adduct, validating covalent binding. (C) Three main binding hits. (D) LC/MS dose response labeling experiment (1 μM BTK, pH 7.5, RT, 2 h) for main hits and ibrutinib. (E) LC/MS time-course experiment (1 μM BTK, 5 μM compound, RT, pH 7.5). n = 2, data represented as mean with SD as error bars. (F) Reduced glutathione (GSH) assay for reactivity assessment (5 mM GSH, 0.2 mM compound, pH 8, 25 °C). (G) Differential scanning fluorimetry (DSF) analyzes the shift in BTK melting point after treatment with the compounds (5 μM BTK kinase domain, 0.02 mM compound, pH 7.5, overnight incubation at 25 °C). The derivative reporter, normalized by the absolute value of the minimum in each data set, was used to determine the melting point. Ibrutinib, YS1, YS2 and YS3 stabilize BTK by ∼13 °C, 9 °C, 5 °C, and 1 °C, respectively. BTK baseline curve is plotted in black. (H) Kinase inhibition analysis by a radiometric filter binding assay measuring phosphorylated substrates products (20 μM ATP, 3 nM BTK, 0.2 mg/mL pEY substrate, pH 7.5, RT). IC50 of YS1, YS2, and YS3 is 30 nM, 77 μM and 8.4 μM, respectively. (I) Kinome phylogenetic tree depicting inhibition data from an assay conducted with YS1 against 362 kinases. Each kinase tested is represented by a dot; dot size is relative to the mean percentage of inhibition with respect to DMSO across two replicates at a dose of 300 nM YS1. Dots are colored red if mean inhibition is greater than 40%, and green otherwise. Illustration reproduced courtesy of Cell Signaling Technology, Inc. (www.cellsignal.com). (J) Dose-dependent BTK activity assay in Mino cells as measured by autophosphorylation of BTK. The cells were incubated for 1 h with either DMSO or various concentrations of YS1. The cells were activated with anti-IgM, and BTK autophosphorylation was quantified by Western blot and normalized with respect to total BTK. IC50 (170 nM) was calculated by fitting the data to a dose–response curve using Prism software (n = 3). (K) YS1 selectivity quantification via competitive pull-down experiments with Probe 4, an alkynylated probe analog of ibrutinib (left) and XO44, a generic kinase probe (right), respectively. Mino cells were treated with either DMSO or 1 μM of YS1 for 1 h, followed by 45 min treatment with either 1 μM Probe 4 or 2 μM XO44 (n = 4). Proteins were quantified using label-free quantification. Proteins in the upper right segment show a significant change (fold change >2; p-value <0.01). In the Probe 4 experiment, 3,416 proteins are plotted (159 kinases and 3,257 nonkinases), out of which 11 exhibit significant competition (2 kinases and 9 nonkinases). In the XO44 experiment, 3,569 proteins are plotted (235 kinases and 3,334 nonkinases), out of which 5 exhibit significant competition (3 kinases and 2 nonkinases). Kinases were determined by a list of human kinase Uniprot entries from KinHub.

We conducted a kinome-wide inhibition assay for YS1 (300 nM) against 362 recombinantly expressed kinases (Figure I; Data set S5). Overall, YS1 exhibits marked selectivity across kinases, with only six off-target kinases inhibited by 40% or more. Three of the off-targets: JAK3, BMX, and TEC (85.1%, 95.3%, and 45.6% mean inhibition, respectively) are members of the TK kinase group, along with BTK, and all contain a Cys residue at an analogous position, making them susceptible to off-target covalent binding. BMX in particular is a challenging off-target since ten clinical BTK inhibitors are known to inhibit it in vitro. Of note, other kinases with an equivalent cysteine: ITK, BLK, TXK, HER2, HER4, EGFR and MKK7, were not significantly inhibited (Data set S5). The remaining off-targets were LIMK1, PKD3 and MNK2.

To measure the efficacy of the compounds in cells, we conducted dose-dependent cellular BTK activity assays measuring the autophosphorylation of BTK. Mino cells were preincubated with YS1, followed by BTK activation with an anti-IgM antibody and activity was measured by Western blotting. YS1 potently inhibits BTK activity in cells (Figure J, Figure S20, IC50 = 107 nM). To evaluate proteomic selectivity, we conducted two cellular proteomic competition assays of YS1 against “Probe 4”, an alkynylated probe analog of ibrutinib, and XO44, a generic kinase probe , (Figure K; Figure S21; Data set S6). In both assays, YS1 is shown to be extremely selective for BTK in cells. In a competition with Probe 4, only a single kinase off-targetTECis significantly competed (log fold-change >2; p-value < 0.01). In a competition experiment with XO44, two kinase off-targets are significantly competedTEC and to a much lesser extent MARK4. We should note that despite XO44 pulling down LIMK1, PKD3 and MNK2, none were competed in cells by YS1 ().

Finally, we crystallized and acquired X-ray diffraction data for YS1, YS2 and YS3 in complex with the BTK kinase domain. The YS1 and YS2 complex structures were determined to resolutions of 1.6 Å and 1.27 Å (PDB: 9ZLJ, 9ZLM respectively; Figure A,B; see Table S4 and Supplementary Methods). Unambiguous electron density is observed for the entirety of the ligands and the covalent attachment to Cys481 (). The structures closely match the AF3 predictions (Figure , ligand heavy atom RMSD of 0.50 Å and 0.41 Å, for YS1 and YS2 respectively, when aligned by the protein backbone). The YS1-BTK structure validates YS1 as the trans-isomer of the cyclohexyl moiety, as predicted by AF3; we were able to isolate a second isomer, presumably the cis-isomer, that produced similar, but slightly less potent results (Figure S23).

4.

4

AF3 structural predictions are accurate. Overlay representations of AF3 predictions (blue) and X-ray crystallography experimental results (gray) for YS1 (A) and YS2 (B) covalently attached to Cys481 of WT-BTK kinase domain. Overlay of the predicted ligand pose and the experimentally determined ligand structure are also shown separately for clarity. Heavy atom ligand RMSD values are 0.50 Å and 0.41 Å, respectively. Close up view of each ligand in the BTK active site are shown with interacting residues labeled. Putative hydrogen bonds are depicted as dashed lines.

Both YS1 and YS2 form hydrogen bonds with the backbone of M477; in YS1 the azaindole group forms two hydrogen bonds with the backbone of Met477 while in YS2 the atypical amide containing linker between the azetidine and benzyl groups forms the two hydrogen bonds to M477 (Figure A,B). Additional hydrogen bonds in the YS1 complex include the carbonyl of the acrylamide warhead with the backbone amide of Cys481 and a hydrogen bond between the Lys430 side chain and the other amide carbonyl in the ligand (Figure A).

Similar to many well characterized BTK inhibitors, YS2 extends into the “back pocket” of the BTK active site, , which is located between the gatekeeper and the C-helix (Figure B). YS1 is not a “back pocket” binder nor can it be characterized as a front pocket binder. YS1 protrudes under the kinase P-loop (or glycine rich loop) and extends its chloro-phenyl group into a subpocket formed by the side chains of M437 (next to the C-helix in the N-lobe) and L542/Y545/V546, which are located at the N-terminal end of the activation loop (Figure A). This subpocket sits above the pocket normally occupied by inhibitors that bind to the front pocket. The YS1 bound structure shows that the P-loop is repositioned into an open conformation to accommodate the ligand and the side chain of F413 adopts an unusual orientation where it points toward the ATP binding cleft rather than toward the activation loop (Figure A). In fact, structures of front pocket ligands bound to BTK show that the side chain of F413 typically occupies the pocket filled by the chloro-phenyl group of YS1 (Figure S24). In addition, front pocket binders can “sequester” the activation loop tyrosine, Y551; but the conformation of Y551 in the YS1 complex structure does not conform to the sequestered state. Thus, AF3 predicted a unique BTK inhibitor that expands our understanding of the druggable pockets within the BTK active site.

The X-ray data for the YS3-BTK kinase domain complex only diffracted to ∼3.5 Å. This resolution does not allow for detailed analysis of ligand binding as described above for YS1 and YS2, but is nevertheless sufficient to determine general ligand binding orientation in the crystal. Interestingly, the binding mode differs from the AF3 predicted binding pose of YS3 (Figure S25). The AF3 prediction places the hydrophobic spiro(3.5)­nonane portion of the YS3 ligand toward the solvent while the X-ray diffraction data show clear electron density extending toward the back pocket of the kinase active site. The predicted pose for YS3 and the pose suggested by the ligand electron density are related by rotation around one bond. It is worth noting that the predicted pose for YS3 is quite similar to the structure of CC-292 bound to the BTK kinase domain (Figure S26, PDB ID 5P9L). Solution studies of CC-292 bound to the BTK active site reveal multiple conformations and binding of CC-292 does not stabilize the BTK kinase domain. The observation that YS3 stabilizes BTK by only 1 °C (Figure G) and the fact that the AF3 prediction differs from the X-ray diffraction data for this ligand may be consistent with an inhibitor that adopts more than one conformation when covalently attached to BTK C481. Nevertheless, YS3 effectively inhibits BTK in a range of biochemical assays (Figure ).

Discussion

Virtual screening of large ligand libraries is a cornerstone of modern drug discovery. Despite the growing prominence of covalent drug discovery, prospective covalent virtual screening remains underdeveloped. We sought to test whether AF3 can address this gap. Inspired by the importance of benchmarks of noncovalent molecules for the development of virtual screening in that domain we developed COValid. Our key finding was that AF3 vastly outperformed traditional methods. This is especially encouraging as it only requires the sequence of the protein and the target ligand rather than a crystallographic structure of the protein as do conventional virtual screening approaches.

Previous studies have shown that AI models trained on ligand binding data could not surpass AutoDock Vina on the DUD-E enrichment test set. Moreover, an analysis conducted prior to the development of AF3 revealed that deep learning–based methods lagged behind classical docking tools, particularly in generating physically valid poses. Most such analyses focused, however, on pose reproduction. In terms of screening, AI models were only evaluated for their capacity to generate structural protein models to support virtual screens using traditional docking tools. In contrast, we show for the first time that AF3 outperformed classical docking in both covalent and noncovalent screening scenarios (Figure ).

Several factors may contribute to the exceptional performance. First, specifying the covalent attachment point may improve accuracy; prior work demonstrated that indicating the ligand binding pocket increased pose recapitulation accuracy from 81% to 93% in AF3. Second, by incorporating backbone flexibility, AF3 can resolve minor clashes that typically confound classical fixed-backbone docking methods or expose cryptic pockets. While some classical docking tools enable side chain flexibility for a subset of the residues near the ligand, these methods rarely allow for conformational changes of the protein backbone. The fact that AF3 can sample various backbone conformations may be key for success in the covalent domain, where the covalent bond itself might be to a residue found on a flexible loop, or when the covalently bound ligands induce conformational changes in flexible regions near the pocket (for example FGFR4 Cys477 and BTK Cys481, respectively, Figure S45). , This point is also exemplified by our prospective screen. When we used DOCKovalent to screen the same library against a static structure of BTK (covalent complex with ibrutinib; PDB: 5P9J), YS1 was ranked only at the 15th percentile of the list (as opposed to the top 0.05% via mPAE-ranked AF3) and its pose did not resemble the crystal structure, likely due to a clash with the side-chain of F413 (Figure S27).

A more cautious explanation, however, is that the structural data used to train AF3 might overlap with our test set. Kinasescommon drug targets that share similar featuresare heavily represented in both the PDB and in drug discovery generally. , Our data set is heavily biased toward kinases (nine out of ten targets), highlighting the potential for data leakage. Yet, this bias also underscores the importance of kinases in the field, given that only this class of proteins provided a critical mass of annotated covalent binders. It is possible, and even likely, that the impressive performance we report is limited to kinases; even so, the impact on accelerating ligand discovery for this crucial class would still be immense. Encouragingly, two observations suggest broader generalization. First, AF3 achieved high enrichment against K-Rasa GTPase unrelated to kinaseseven when considering only actives dissimilar to any PDB ligand (Figure S8). Second, AF3 demonstrated strong noncovalent enrichment across dozens of nonkinase targets, albeit not as high, on average, as the noncovalent enrichment for kinases (Figure D). This work also focused on cysteine residues. While the low abundance of Cys residues is beneficial for proteome-wide selectivity, it also limits the scope of proteins that could be targeted covalently. Efforts to expand covalent applicability has led to the development of novel warheads which can covalently bind to multiple residues, including Ser, Lys, Tyr, and His residues. Even within the domain of active sites across the kinome, experimentally validated, covalent AF3-based screening campaigns that target additional residues beyond Cys could increase the pool of relevant targets for covalent cofolding.

AF3 confidence metrics were trained to quantify modeling accuracy. Most predicted metrics lack the resolution to rank large libraries of models (Figure S5), particularly those focusing on the protein side. A recent adversarial experiment suggested that AF3 does not rely on specific physical interactions for ligand placement, and indeed, there is no apparent correlation between the mPAE score and the binding affinity of experimental actives (Figure S11). Nevertheless, the fact that rescoring AF3 models with Rosettaan established physics-based scoring methodsignificantly outperformed covalent docking tools in terms of enrichment for five of the ten sites suggests that the models generated by AF3 are physically sensible. In that respect, while AF3′s covalent cofolding tends to produce malformed poses with respect to the bond length and angle geometry (Figure S28), it compensates so that the overall ligand pose remains accurate.

The AF3 mis-prediction of the YS3-BTK complex in which an hydrophobic moiety was exposed to solvent may point to a weakness in interpreting physical interactions, already observed in a recent study that reported a tendency of AF3 to predict unfavorable interactions with regards to hydrophobicity. Another adversarial study, pointed to gaps in AF3’s representation of physical interactions. Together, these highlight an area of improvement for cofolding methods that could benefit from incorporation of physics-based knowledge.

An additional advantage of mPAE is its apparent insensitivity to the protein target, allowing comparisons across complexes and enabling the assignment of binding probabilities based on mPAE values (Figure C). Cross-docking experiments (Figure S12) demonstrated that mPAE can discriminate when covalent cofolding to different cysteine locations, but not so well within the same location, suggesting it should not be used currently for selectivity predictions. In an in silico analysis based on the noncovalent enrichment benchmark DEKOIS2.0, Shen et al. also report mPAE as the superior metric for enrichment performance for both AF3 as well as Boltz-2 across all evaluated confidence metrics. ,

In practical terms, despite AF3′s superior enrichment, its long run-time may render it impractical for screening the extremely large virtual libraries that have become standard. A feasible solution is to use a fast classical docking method to screen very large virtual librariespotentially comprising billions of compoundsfollowed by cofolding and rescoring the top million or so predicted compounds with AF3 to more effectively enrich true binders.

Another limitation to be considered is the heuristics we (and other covalent docking software) use in modeling covalent binding. Whereas, we model the covalent bound, adduct state, similarly to classical docking tools, AF3 does not explicitly account for the energetics of covalent bond formation, making it challenging to rank hits with different electrophiles, and potentially overemphasizing molecular recognition over covalent bond formation. Future investigation may examine the modeling of the covalent intermediate using AF3 to see if this better captures the kinetics of covalent binding.

To mitigate the aforementioned confounding factors, the ultimate validation of real-world applicability is experimental validation. The fact that we were able to discover novel chemical matter in a prospective manner, clearly indicates that at least for well-studied protein targets with well-defined pockets, such as kinases, AF3-based prospective screening can go beyond any examples it has observed during training. Such screening opens the door for incorporation of protein flexibility in large scale virtual screening for the first time, and may unlock more challenging, flexible targets that are recalcitrant to state-of-the-art docking methods.

Supplementary Material

ja5c22222_si_001.pdf (5.4MB, pdf)
ja5c22222_si_002.xlsx (11.1KB, xlsx)
ja5c22222_si_003.xlsx (9.3KB, xlsx)
ja5c22222_si_004.xlsx (2.1MB, xlsx)
ja5c22222_si_005.xlsx (33.5MB, xlsx)
ja5c22222_si_006.xlsx (30.6KB, xlsx)
ja5c22222_si_007.xlsx (575.5KB, xlsx)

Acknowledgments

We thank Sarel Fleishman for critical reading of the manuscript, the Irwin lab in UCSF for access to their cluster for DOCK6 and DOCKovalent ligand generation and in particular Dr. Khanh Tang for technical assistance. We thank Dr. Trent Balius for assistance with DOCK6.12, as well as sharing DUDE-Z enrichment calculation data, and Dr. Alexey Orlov for sharing his ligand generation pipeline. We acknowledge Crelux, a WuXi AppTec company, for performing the kinetic analysis of the covalent hits. Y.S. is funded by the CHE fellowship for data sciences. This research was generously supported by the Knell Family Institute of Artificial Intelligence. Research in the London lab is funded by the Abisch-Frenkel foundation, European Research Council (ERC_CoG 101125683), the Israel Science Foundation (1869/24), the Honey and Dr. Barry Sherman Lab, the Dr. Barry Sherman Institute for Medicinal Chemistry, the Abisch-Frenkel RNA Therapeutics Center, the Moross Integrated Cancer Center, the Goldhirsh-Yellin Foundation and Celia Zwillenberg-Fridman. D.Y.L. and A.H.A. thank the Roy J. Carver Charitable Trust for financial support. This work also utilized the resources at the NE-CAT beamlines (GM124165), a Pilatus detector (RR029205), and an Eiger detector (OD021527), all of which are funded by the NIH.

The crystal structures of YS1 and YS2 were deposited to the Protein Data Bank with PDB IDs 9ZLJ and 9ZLM, respectively. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the data set identifier PXD072258.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jacs.5c22222.

  • Benchmark set development, enrichment analysis, covalent docking methodology, AF3 cofolding, experimental validation, optimization of physics-based covalent docking tools, and Boltz-2 analysis for DUDE-Z (PDF)

  • Experimental structures curated for COValid (XLSX)

  • Physiochemical properties per protein in COValid (XLSX)

  • Active and decoy protomers in COValid (XLSX)

  • Prospective virtual library (XLSX)

  • YS1 kinome panel (XLSX)

  • YS1 proteomics (XLSX)

The authors declare no competing financial interest.

References

  1. Baillie T. A.. Targeted covalent inhibitors for drug design. Angew. Chem., Int. Ed. Engl. 2016;55:13408–13421. doi: 10.1002/anie.201601091. [DOI] [PubMed] [Google Scholar]
  2. Boike L., Henning N. J., Nomura D. K.. Advances in covalent drug discovery. Nat. Rev. Drug Discovery. 2022;21:881–898. doi: 10.1038/s41573-022-00542-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. London N.. Covalent proximity inducers. Chem. Rev. 2025;125:326–368. doi: 10.1021/acs.chemrev.4c00570. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Sutanto F., Konstantinidou M., Dömling A.. Covalent inhibitors: a rational approach to drug discovery. RSC Med. Chem. 2020;11:876–884. doi: 10.1039/D0MD00154F. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Ostrem J. M. L., Shokat K. M.. Targeting KRAS G12C with covalent inhibitors. Annu. Rev. Cancer Biol. 2022;6:49–64. doi: 10.1146/annurev-cancerbio-041621-012549. [DOI] [Google Scholar]
  6. Hammond J., Leister-Tebbe H., Gardner A., Abreu P., Bao W., Wisemandle W., Baniecki M., Hendrick V. M., Damle B., Simón-Campos A., Pypstra R., Rusnak J. M.. Oral nirmatrelvir for high-risk, nonhospitalized adults with Covid-19. N. Engl. J. Med. 2022;386:1397–1408. doi: 10.1056/NEJMoa2118542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Blair H. A.. Ritlecitinib: First approval. Drugs. 2023;83:1315–1321. doi: 10.1007/s40265-023-01928-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Jänne P. A., Riely G. J., Gadgeel S. M., Heist R. S., Ou S.-H. I., Pacheco J. M., Johnson M. L., Sabari J. K., Leventakos K., Yau E., Bazhenova L., Negrao M. V., Pennell N. A., Zhang J., Anderes K., Der-Torossian H., Kheoh T., Velastegui K., Yan X., Christensen J. G., Chao R. C., Spira A. I.. Adagrasib in non-small-cell lung cancer harboring a KRASG12C mutation. N. Engl. J. Med. 2022;387:120–131. doi: 10.1056/NEJMoa2204619. [DOI] [PubMed] [Google Scholar]
  9. London N., Miller R. M., Krishnan S., Uchida K., Irwin J. J., Eidam O., Gibold L., Cimermančič P., Bonnet R., Shoichet B. K., Taunton J.. Covalent docking of large libraries for the discovery of chemical probes. Nat. Chem. Biol. 2014;10:1066–1072. doi: 10.1038/nchembio.1666. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Bianco G., Forli S., Goodsell D. S., Olson A. J.. Covalent docking using autodock: Two-point attractor and flexible side chain methods: Covalent Docking with AutoDock. Protein Sci. 2016;25:295–301. doi: 10.1002/pro.2733. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Wu Y., Brooks C. L. Iii. Covalent docking in CDOCKER. J. Comput. Aided Mol. Des. 2022;36:563–574. doi: 10.1007/s10822-022-00472-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Toledo Warshaviak D., Golan G., Borrelli K. W., Zhu K., Kalid O.. Structure-based virtual screening approach for discovery of covalently bound ligands. J. Chem. Inf. Model. 2014;54:1941–1950. doi: 10.1021/ci500175r. [DOI] [PubMed] [Google Scholar]
  13. De Cesco S., Deslandes S., Therrien E., Levan D., Cueto M., Schmidt R., Cantin L.-D., Mittermaier A., Juillerat-Jeanneret L., Moitessier N.. Virtual screening and computational optimization for the discovery of covalent prolyl oligopeptidase inhibitors with activity in human cells. J. Med. Chem. 2012;55:6306–6315. doi: 10.1021/jm3002839. [DOI] [PubMed] [Google Scholar]
  14. Katritch V., Byrd C. M., Tseitin V., Dai D., Raush E., Totrov M., Abagyan R., Jordan R., Hruby D. E.. Discovery of small molecule inhibitors of ubiquitin-like poxvirus proteinase I7L using homology modeling and covalent docking approaches. J. Comput. Aided Mol. Des. 2007;21:549–558. doi: 10.1007/s10822-007-9138-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Tan Y. S., Chakrabarti M., Stein R. M., Prentis L. E., Rizzo R. C., Kurtzman T., Fischer M., Balius T. E.. Development of receptor desolvation scoring and covalent sampling in DOCK 6: Methods evaluated on a RAS test set. J. Chem. Inf. Model. 2025;65:722. doi: 10.1021/acs.jcim.4c01623. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Rachman M., Scarpino A., Bajusz D., Pálfy G., Vida I., Perczel A., Barril X., Keserű G. M.. DUckCov: A Dynamic Undocking-based virtual screening protocol for covalent binders. ChemMedChem. 2019;14:1011–1021. doi: 10.1002/cmdc.201900078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Zhang S., Tan J., Lai Z., Li Y., Pang J., Xiao J., Huang Z., Zhang Y., Ji H., Lai Y.. Effective virtual screening strategy toward covalent ligands: identification of novel NEDD8-activating enzyme inhibitors. J. Chem. Inf. Model. 2014;54:1785–1797. doi: 10.1021/ci5002058. [DOI] [PubMed] [Google Scholar]
  18. Shraga A., Olshvang E., Davidzohn N., Khoshkenar P., Germain N., Shurrush K., Carvalho S., Avram L., Albeck S., Unger T., Lefker B., Subramanyam C., Hudkins R. L., Mitchell A., Shulman Z., Kinoshita T., London N.. Covalent docking identifies a potent and selective MKK7 inhibitor. Cell Chem. Biol. 2019;26:98–108.e5. doi: 10.1016/j.chembiol.2018.10.011. [DOI] [PubMed] [Google Scholar]
  19. Nnadi C. I., Jenkins M. L., Gentile D. R., Bateman L. A., Zaidman D., Balius T. E., Nomura D. K., Burke J. E., Shokat K. M., London N.. Novel K-Ras G12C switch-II covalent binders destabilize Ras and accelerate nucleotide exchange. J. Chem. Inf. Model. 2018;58:464–471. doi: 10.1021/acs.jcim.7b00399. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Fink E. A., Bardine C., Gahbauer S., Singh I., Detomasi T. C., White K., Gu S., Wan X., Chen J., Ary B., Glenn I., O’Connell J., O’Donnell H., Fajtová P., Lyu J., Vigneron S., Young N. J., Kondratov I. S., Alisoltani A., Simons L. M., Lorenzo-Redondo R., Ozer E. A., Hultquist J. F., O’Donoghue A. J., Moroz Y. S., Taunton J., Renslo A. R., Irwin J. J., García-Sastre A., Shoichet B. K., Craik C. S.. Large library docking for novel SARS-CoV-2 main protease non-covalent and covalent inhibitors. Protein Sci. 2023;32:e4712. doi: 10.1002/pro.4712. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Scarpino A., Ferenczy G. G., Keserű G. M.. Comparative evaluation of covalent docking tools. J. Chem. Inf. Model. 2018;58:1441–1458. doi: 10.1021/acs.jcim.8b00228. [DOI] [PubMed] [Google Scholar]
  22. Gao M., Moumbock A. F. A., Qaseem A., Xu Q., Günther S.. CovPDB: a high-resolution coverage of the covalent protein-ligand interactome. Nucleic Acids Res. 2022;50:D445–D450. doi: 10.1093/nar/gkab868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Guo X.-K., Zhang Y.. CovBinderInPDB: A structure-based covalent binder database. J. Chem. Inf. Model. 2022;62:6057–6068. doi: 10.1021/acs.jcim.2c01216. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Du H., Zhang X., Wu Z., Zhang O., Gu S., Wang M., Zhu F., Li D., Hou T., Pan P.. CovalentInDB 2.0: an updated comprehensive database for structure-based and ligand-based covalent inhibitor design and screening. Nucleic Acids Res. 2025;53:D1322–D1327. doi: 10.1093/nar/gkae946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Wen C., Yan X., Gu Q., Du J., Wu D., Lu Y., Zhou H., Xu J.. Systematic studies on the protocol and criteria for selecting a covalent docking tool. Molecules. 2019;24:2183. doi: 10.3390/molecules24112183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Huang N., Shoichet B. K., Irwin J. J.. Benchmarking sets for molecular docking. J. Med. Chem. 2006;49:6789–6801. doi: 10.1021/jm0608356. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Mysinger M. M., Carchia M., Irwin J. J., Shoichet B. K.. Directory of useful decoys, enhanced (DUD-E): better ligands and decoys for better benchmarking. J. Med. Chem. 2012;55:6582–6594. doi: 10.1021/jm300687e. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Stein R. M., Yang Y., Balius T. E., O’Meara M. J., Lyu J., Young J., Tang K., Shoichet B. K., Irwin J. J.. Property-unmatched decoys in docking benchmarks. J. Chem. Inf. Model. 2021;61:699–714. doi: 10.1021/acs.jcim.0c00598. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Bauer M. R., Ibrahim T. M., Vogel S. M., Boeckler F. M.. Evaluation and optimization of virtual screening workflows with DEKOIS 2.0--a public library of challenging docking benchmark sets. J. Chem. Inf. Model. 2013;53:1447–1462. doi: 10.1021/ci400115b. [DOI] [PubMed] [Google Scholar]
  30. Rohrer S. G., Baumann K.. Maximum unbiased validation (MUV) data sets for virtual screening based on PubChem bioactivity data. J. Chem. Inf. Model. 2009;49:169–184. doi: 10.1021/ci8002649. [DOI] [PubMed] [Google Scholar]
  31. Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A. J., Bambrick J., Bodenstein S. W., Evans D. A., Hung C.-C., O’Neill M., Reiman D., Tunyasuvunakool K., Wu Z., Žemgulytė A., Arvaniti E., Beattie C., Bertolli O., Bridgland A., Cherepanov A., Congreve M., Cowen-Rivers A. I., Cowie A., Figurnov M., Fuchs F. B., Gladman H., Jain R., Khan Y. A., Low C. M. R., Perlin K., Potapenko A., Savy P., Singh S., Stecula A., Thillaisundaram A., Tong C., Yakneen S., Zhong E. D., Zielinski M., Žídek A., Bapst V., Kohli P., Jaderberg M., Hassabis D., Jumper J. M.. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Stecula A., Paul R., Litchfield K., Dalton S. E., Low C. M. R., Reis C. R., Congreve M.. The rise of AlphaFold in drug design. Prog. Med. Chem. 2025;64:99–147. doi: 10.1016/bs.pmch.2025.09.002. [DOI] [PubMed] [Google Scholar]
  33. Abdeldayem A., Raouf Y. S., Constantinescu S. N., Moriggl R., Gunning P. T.. Advances in covalent kinase inhibitors. Chem. Soc. Rev. 2020;49:2617–2687. doi: 10.1039/C9CS00720B. [DOI] [PubMed] [Google Scholar]
  34. Sotriffer C.. Docking of covalent ligands: Challenges and approaches. Mol. Inform. 2018;37:e1800062. doi: 10.1002/minf.201800062. [DOI] [PubMed] [Google Scholar]
  35. Gilson M. K., Liu T., Baitaluk M., Nicola G., Hwang L., Chong J.. BindingDB in 2015: A public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Res. 2016;44:D1045–53. doi: 10.1093/nar/gkv1072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Zdrazil B., Felix E., Hunter F., Manners E. J., Blackshaw J., Corbett S., Veij M. de, Ioannidis H., Lopez D. M., Mosquera J. F., Magarinos M. P., Bosc N., Arcila R., Kizilören T., Gaulton A., Bento A. P., Adasme M. F., Monecke P., Landrum G. A., Leach A. R.. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 2024;52:D1180–D1192. doi: 10.1093/nar/gkad1004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Liu Q., Sabnis Y., Zhao Z., Zhang T., Buhrlage S. J., Jones L. H., Gray N. S.. Developing irreversible inhibitors of the protein kinase cysteinome. Chem. Biol. 2013;20:146–159. doi: 10.1016/j.chembiol.2012.12.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Zhao Z., Bourne P. E.. Progress with covalent small-molecule kinase inhibitors. Drug Discovery Today. 2018;23:727–735. doi: 10.1016/j.drudis.2018.01.035. [DOI] [PubMed] [Google Scholar]
  39. Rathod L. S., Dabhade P. S., Mokale S. N.. Recent progress in targeting KRAS mutant cancers with covalent G12C-specific inhibitors. Drug Discovery Today. 2023;28:103557. doi: 10.1016/j.drudis.2023.103557. [DOI] [PubMed] [Google Scholar]
  40. Verdonk M. L., Berdini V., Hartshorn M. J., Mooij W. T. M., Murray C. W., Taylor R. D., Watson P.. Virtual screening using protein-ligand docking: avoiding artificial enrichment. J. Chem. Inf. Comput. Sci. 2004;44:793–806. doi: 10.1021/ci034289q. [DOI] [PubMed] [Google Scholar]
  41. Irwin J. J., Tang K. G., Young J., Dandarchuluun C., Wong B. R., Khurelbaatar M., Moroz Y. S., Mayfield J., Sayle R. A.. ZINC20-A free ultralarge-scale chemical database for ligand discovery. J. Chem. Inf. Model. 2020;60:6065–6073. doi: 10.1021/acs.jcim.0c00675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Bender B. J., Gahbauer S., Luttens A., Lyu J., Webb C. M., Stein R. M., Fink E. A., Balius T. E., Carlsson J., Irwin J. J., Shoichet B. K.. A practical guide to large-scale docking. Nat. Protoc. 2021;16:4799–4832. doi: 10.1038/s41596-021-00597-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Knight, I. S. , Naprienko, S. , Irwin, J. J. , Enrichment Score: a better quantitative metric for evaluating the enrichment capacity of molecular docking models. arXiv (2022). http://arxiv.org/abs/2210.10905 (accessed 2025-12-01).
  44. Mysinger M. M., Shoichet B. K.. Rapid context-dependent ligand desolvation in molecular docking. J. Chem. Inf. Model. 2010;50:1561–1573. doi: 10.1021/ci100214a. [DOI] [PubMed] [Google Scholar]
  45. Park H., Bradley P., Greisen P. Jr, Liu Y., Mulligan V. K., Kim D. E., Baker D., DiMaio F.. Simultaneous optimization of biomolecular energy functions on features from small molecules and macromolecules. J. Chem. Theory Comput. 2016;12:6201–6212. doi: 10.1021/acs.jctc.6b00819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Balius T. E., Tan Y. S., Chakrabarti M.. DOCK 6: Incorporating hierarchical traversal through precomputed ligand conformations to enable large-scale docking. J. Comput. Chem. 2024;45:47–63. doi: 10.1002/jcc.27218. [DOI] [PubMed] [Google Scholar]
  47. Masters M. R., Mahmoud A. H., Lill M. A.. Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions. Nat. Commun. 2025;16:8854. doi: 10.1038/s41467-025-63947-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Honigberg L. A., Smith A. M., Sirisawad M., Verner E., Loury D., Chang B., Li S., Pan Z., Thamm D. H., Miller R. A., Buggy J. J.. The Bruton tyrosine kinase inhibitor PCI-32765 blocks B-cell activation and is efficacious in models of autoimmune disease and B-cell malignancy. Proc. Natl. Acad. Sci. U. S. A. 2010;107:13075–13080. doi: 10.1073/pnas.1004594107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Estupiñán H. Y., Berglöf A., Zain R., Smith C. I. E.. Comparative analysis of BTK inhibitors and mechanisms underlying adverse effects. Front. Cell Dev. Biol. 2021;9:630942. doi: 10.3389/fcell.2021.630942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Barf T., Covey T., Izumi R., van de Kar B., Gulrajani M., van Lith B., van Hoek M., de Zwart E., Mittag D., Demont D., Verkaik S., Krantz F., Pearson P. G., Ulrich R., Kaptein A.. Acalabrutinib (ACP-196): A covalent Bruton tyrosine kinase inhibitor with a differentiated selectivity and in vivo potency profile. J. Pharmacol. Exp. Ther. 2017;363:240–252. doi: 10.1124/jpet.117.242909. [DOI] [PubMed] [Google Scholar]
  51. Passaro, S. , Corso, G. , Wohlwend, J. , Reveiz, M. , Thaler, S. , Somnath, V. R. , Getz, N. , Portnoi, T. , Roy, J. , Stark, H. , Kwabi-Addo, D. , Beaini, D. , Jaakkola, T. , Barzilay, R. , Boltz-2: Towards accurate and efficient binding affinity prediction. BioRxiv (2025). 10.1101/2025.06.14.659707 (accessed 2025-12-01). [DOI]
  52. Stanchina M. D., Montoya S., Danilov A. V., Castillo J. J., Alencar A. J., Chavez J. C., Cheah C. Y., Chiattone C., Wang Y., Thompson M., Ghia P., Taylor J., Alderuccio J. P.. Navigating the changing landscape of BTK-targeted therapies for B cell lymphomas and chronic lymphocytic leukaemia. Nat. Rev. Clin. Oncol. 2024;21:867–887. doi: 10.1038/s41571-024-00956-1. [DOI] [PubMed] [Google Scholar]
  53. Cameron F., Sanford M.. Ibrutinib: first global approval. Drugs. 2014;74:263–271. doi: 10.1007/s40265-014-0178-8. [DOI] [PubMed] [Google Scholar]
  54. Syed Y. Y.. Zanubrutinib: First approval. Drugs. 2020;80:91–97. doi: 10.1007/s40265-019-01252-4. [DOI] [PubMed] [Google Scholar]
  55. Markham A., Dhillon S.. Acalabrutinib: First global approval. Drugs. 2018;78:139–145. doi: 10.1007/s40265-017-0852-8. [DOI] [PubMed] [Google Scholar]
  56. Labanca C., Martino E. A., Vigna E., Bruzzese A., Mendicino F., Caridà G., Lucia E., Olivito V., Manicardi V., Amodio N., Neri A., Morabito F., Gentile M.. Rilzabrutinib for the treatment of immune thrombocytopenia. Eur. J. Haematol. 2025;115:4–15. doi: 10.1111/ejh.14425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Corrionero A., Zhang X., Alfonso P., Morris P. J., Klumpp-Thomas C., Melani C., McKnight C., Phelan J. D., Holland D., Wilson K., Hoyt S. B., Roschewski M., Tonge P. J., Wilson W., Ceribelli M., Staudt L. M., Thomas C. J.. An assessment of kinase selectivity, enzyme inhibition kinetics and in vitro activity for several bruton tyrosine kinase (BTK) inhibitors. ACS Pharmacol. Transl. Sci. 2025;8:4312. doi: 10.1021/acsptsci.5c00412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Lanning B. R., Whitby L. R., Dix M. M., Douhan J., Gilbert A. M., Hett E. C., Johnson T. O., Joslyn C., Kath J. C., Niessen S., Roberts L. R., Schnute M. E., Wang C., Hulce J. J., Wei B., Whiteley L. O., Hayward M. M., Cravatt B. F.. A road map to evaluate the proteome-wide selectivity of covalent kinase inhibitors. Nat. Chem. Biol. 2014;10:760–767. doi: 10.1038/nchembio.1582. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Zhao Q., Ouyang X., Wan X., Gajiwala K. S., Kath J. C., Jones L. H., Burlingame A. L., Taunton J.. Broad-spectrum kinase profiling in live cells with lysine-targeted sulfonyl fluoride probes. J. Am. Chem. Soc. 2017;139:680–685. doi: 10.1021/jacs.6b08536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Dhillon S.. Tirabrutinib: First approval. Drugs. 2020;80:835–840. doi: 10.1007/s40265-020-01318-8. [DOI] [PubMed] [Google Scholar]
  61. Eid S., Turk S., Volkamer A., Rippmann F., Fulle S.. KinMap: a web-based tool for interactive navigation through human kinome data. BMC Bioinformatics. 2017;18:16. doi: 10.1186/s12859-016-1433-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Möbitz H.. The ABC of protein kinase conformations. Biochim. Biophys. Acta. 2015;1854:1555–1566. doi: 10.1016/j.bbapap.2015.03.009. [DOI] [PubMed] [Google Scholar]
  63. Sydow D., Schmiel P., Mortier J., Volkamer A.. KinFragLib: Exploring the kinase inhibitor space using subpocket-focused fragmentation and recombination. J. Chem. Inf. Model. 2020;60:6081–6094. doi: 10.1021/acs.jcim.0c00839. [DOI] [PubMed] [Google Scholar]
  64. Bender A. T., Gardberg A., Pereira A., Johnson T., Wu Y., Grenningloh R., Head J., Morandi F., Haselmayer P., Liu-Bujalski L.. Ability of bruton’s tyrosine kinase inhibitors to sequester Y551 and prevent phosphorylation determines potency for inhibition of Fc receptor but not B-cell receptor signaling. Mol. Pharmacol. 2017;91:208–219. doi: 10.1124/mol.116.107037. [DOI] [PubMed] [Google Scholar]
  65. Joseph R. E., Amatya N., Fulton D. B., Engen J. R., Wales T. E., Andreotti A.. Differential impact of BTK active site inhibitors on the conformational state of full-length BTK. Elife. 2020;9:e60470. doi: 10.7554/eLife.60470. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Sadybekov A. V., Katritch V.. Computational approaches streamlining drug discovery. Nature. 2023;616:673–685. doi: 10.1038/s41586-023-05905-z. [DOI] [PubMed] [Google Scholar]
  67. Attwood M. M., Fabbro D., Sokolov A. V., Knapp S., Schiöth H. B.. Trends in kinase drug discovery: targets, indications and inhibitor design. Nat. Rev. Drug Discovery. 2021;20:839–861. doi: 10.1038/s41573-021-00252-y. [DOI] [PubMed] [Google Scholar]
  68. Cohen P., Cross D., Jänne P. A.. Kinase drug discovery 20 years after imatinib: progress and future directions. Nat. Rev. Drug Discovery. 2021;20:551–569. doi: 10.1038/s41573-021-00195-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Škrinjar, P. , Eberhardt, J. , Durairaj, J. , Schwede, T. , Have protein-ligand co-folding methods moved beyond memorisation? BioRxiv (2025). 10.1101/2025.02.03.636309 (accessed 2025-12-01). [DOI]
  70. Lyu J., Kapolka N., Gumpper R., Alon A., Wang L., Jain M. K., Barros-Álvarez X., Sakamoto K., Kim Y., DiBerto J., Kim K., Glenn I. S., Tummino T. A., Huang S., Irwin J. J., Tarkhanova O. O., Moroz Y., Skiniotis G., Kruse A. C., Shoichet B. K., Roth B. L.. AlphaFold2 structures guide prospective ligand discovery. Science. 2024;384:eadn6354. doi: 10.1126/science.adn6354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Holcomb M., Chang Y.-T., Goodsell D. S., Forli S.. Evaluation of AlphaFold2 structures as docking targets. Protein Sci. 2023;32:e4530. doi: 10.1002/pro.4530. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Karelina M., Noh J. J., Dror R. O.. How accurately can one predict drug binding modes using AlphaFold models? Elife. 2023;12:RP89386. doi: 10.7554/eLife.89386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Tan L., Wang J., Tanizaki J., Huang Z., Aref A. R., Rusan M., Zhu S.-J., Zhang Y., Ercan D., Liao R. G., Capelletti M., Zhou W., Hur W., Kim N., Sim T., Gaudet S., Barbie D. A., Yeh J.-R. J., Yun C.-H., Hammerman P. S., Mohammadi M., Jänne P. A., Gray N. S.. Development of covalent inhibitors that can overcome resistance to first-generation FGFR kinase inhibitors. Proc. Natl. Acad. Sci. U. S. A. 2014;111:E4869. doi: 10.1073/pnas.1403438111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Du G., Rao S., Gurbani D., Henning N. J., Jiang J., Che J., Yang A., Ficarro S. B., Marto J. A., Aguirre A. J., Sorger P. K., Westover K. D., Zhang T., Gray N. S.. Structure-based design of a potent and selective covalent inhibitor for SRC kinase that targets a P-loop cysteine. J. Med. Chem. 2020;63:1624–1641. doi: 10.1021/acs.jmedchem.9b01502. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Mukherjee H., Grimster N. P.. Beyond cysteine: recent developments in the area of targeted covalent inhibition. Curr. Opin. Chem. Biol. 2018;44:30–38. doi: 10.1016/j.cbpa.2018.05.011. [DOI] [PubMed] [Google Scholar]
  76. Zhang Z., Guiley K. Z., Shokat K. M.. Chemical acylation of an acquired serine suppresses oncogenic signaling of K-Ras­(G12S) Nat. Chem. Biol. 2022;18:1177–1183. doi: 10.1038/s41589-022-01065-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Kawano M., Murakawa S., Higashiguchi K., Matsuda K., Tamura T., Hamachi I.. Lysine-reactive N-acyl-N-aryl sulfonamide warheads: Improved reaction properties and application in the covalent inhibition of an ibrutinib-resistant BTK mutant. J. Am. Chem. Soc. 2023;145:26202–26212. doi: 10.1021/jacs.3c08740. [DOI] [PubMed] [Google Scholar]
  78. Hahm H. S., Toroitich E. K., Borne A. L., Brulet J. W., Libby A. H., Yuan K., Ware T. B., McCloud R. L., Ciancone A. M., Hsu K.-L.. Global targeting of functional tyrosines using sulfur-triazole exchange chemistry. Nat. Chem. Biol. 2020;16:150–159. doi: 10.1038/s41589-019-0404-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Jia S., He D., Chang C. J.. Bioinspired thiophosphorodichloridate reagents for chemoselective histidine bioconjugation. J. Am. Chem. Soc. 2019;141:7294–7301. doi: 10.1021/jacs.8b11912. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Childs, H. , Zhou, P. , Donald, B. R. , Has AlphaFold 3 solved the protein folding problem for D-peptides? BioRxiv (2025). 10.1101/2025.03.14.643307 (accessed 2025-12-01). [DOI]
  81. Shen C., Zhang X., Gu S., Zhang O., Wang Q., Du G., Zhao Y., Jiang L., Pan P., Kang Y., Zhao Q., Hsieh C.-Y., Hou T.. Unlocking the application potential of AlphaFold3-like approaches in virtual screening. Chem. Sci. 2026;17:2858. doi: 10.1039/D5SC06481C. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Lyu J., Wang S., Balius T. E., Singh I., Levit A., Moroz Y. S., O’Meara M. J., Che T., Algaa E., Tolmachova K., Tolmachev A. A., Shoichet B. K., Roth B. L., Irwin J. J.. Ultra-large library docking for discovering new chemotypes. Nature. 2019;566:224–229. doi: 10.1038/s41586-019-0917-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Gentile F., Yaacoub J. C., Gleave J., Fernandez M., Ton A.-T., Ban F., Stern A., Cherkasov A.. Artificial intelligence-enabled virtual screening of ultra-large chemical libraries with deep docking. Nat. Protoc. 2022;17:672–697. doi: 10.1038/s41596-021-00659-2. [DOI] [PubMed] [Google Scholar]
  84. Sadybekov A. A., Sadybekov A. V., Liu Y., Iliopoulos-Tsoutsouvas C., Huang X.-P., Pickett J., Houser B., Patel N., Tran N. K., Tong F., Zvonok N., Jain M. K., Savych O., Radchenko D. S., Nikas S. P., Petasis N. A., Moroz Y. S., Roth B. L., Makriyannis A., Katritch V.. Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature. 2022;601:452–459. doi: 10.1038/s41586-021-04220-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Perez-Riverol Y., Bandla C., Kundu D. J., Kamatchinathan S., Bai J., Hewapathirana S., John N. S., Prakash A., Walzer M., Wang S., Vizcaíno J. A.. The PRIDE database at 20 years: 2025 update. Nucleic Acids Res. 2025;53:D543–D553. doi: 10.1093/nar/gkae1011. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ja5c22222_si_001.pdf (5.4MB, pdf)
ja5c22222_si_002.xlsx (11.1KB, xlsx)
ja5c22222_si_003.xlsx (9.3KB, xlsx)
ja5c22222_si_004.xlsx (2.1MB, xlsx)
ja5c22222_si_005.xlsx (33.5MB, xlsx)
ja5c22222_si_006.xlsx (30.6KB, xlsx)
ja5c22222_si_007.xlsx (575.5KB, xlsx)

Data Availability Statement

The crystal structures of YS1 and YS2 were deposited to the Protein Data Bank with PDB IDs 9ZLJ and 9ZLM, respectively. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the data set identifier PXD072258.


Articles from Journal of the American Chemical Society are provided here courtesy of American Chemical Society

RESOURCES