Abstract
Computational screening of giga-scale chemical spaces opens a cost-effective path to high-quality hit identification, providing entry points for drug discovery. As these on-demand spaces grow and successful applications multiply, rigorous blind benchmarks like CACHE Challenges provide important performance metrics for computational tools. Here, we report the first application of the V-SYNTHES2 synthon-based screening approach to the 173-billion-compound Enamine xREAL Space, a 16-fold expansion beyond its previous benchmarks, demonstrating near-linear computational scaling with only a 10–15% increase in cost relative to the 11-billion-compound REAL Space. We applied this workflow in CACHE Challenge #2, targeting the RNA-binding site of NSP13 (SARS-CoV-2), and CACHE Challenge #4, targeting the tyrosine kinase-binding domain of CBLB, both pockets lacking established pharmacology and representing extreme hit-finding challenges. Under blinded, independently validated conditions, V-SYNTHES2 ranked among the top-performing submissions: the 8% hit rate for NSP13 exceeded the field average of 2.3% and placed the approach among the top three workflows, while for CBLB, one compound meeting predefined hit criteria was identified. These results demonstrate that V-SYNTHES2 maintains robust performance at giga-scale on ligand-depleted targets, precisely the conditions where data-driven approaches would face fundamental limitations, and establish a quantitative performance baseline for synthon-based screening of hundred-billion-compound chemical spaces.
Subject terms: Chemistry, Computational biology and bioinformatics, Drug discovery, Mathematics and computing
Introduction
Structure-based virtual screening (SBVS) has proved to be a successful alternative to high-throughput screening for hit-finding in the early drug discovery process, demonstrating high hit rates and the identification of high-quality hits for a number of clinically relevant targets1–6. In the last few years, SBVS has expanded to giga-scale chemical spaces containing billions of compounds7, such as Enamine REAL8, Chemspace Freedom Space9, WuXi AppTec GalaXi Space, Otava CHEMriya, and eMolecules eXplore10. The first two spaces are based on highly synthesizable drug-like molecules constructed from commercially available building blocks, experimentally qualified in the case of Enamine REAL Space, and ML-filtered in the case of Chemspace Freedom Space, using validated synthetic procedures. Compounds generated under the REAL concept typically result in synthesis success rates above 80% within 3–4 weeks11. Such on-demand chemical spaces offer several advantages for drug discovery: (i) large size, diversity, and drug-like properties increase the likelihood of identifying more potent and tractable hits12,13, (ii) a high synthesis success rate and competitive pricing enable reliable project budgeting, (iii) parallel chemistry allows for rapid hit expansion and SAR studies leading to fast development of lead series14. These benefits of on-demand spaces are highly attractive for both industry and academic efforts in hit and lead discovery. They have also enabled a series of CACHE (Critical Assessment of Computational Hit-Finding Experiments) Challenges15 by providing a reliable, uniform platform for fast, affordable on-demand synthesis. In the last few years, CACHE established itself in the community as a major open competition platform designed to accelerate early-stage drug discovery in areas of market failure by providing unbiased, high-quality experimental feedback on computational predictions16–18.
With the rapid expansion of virtual chemical spaces into billions and even trillions of molecules7,19, brute-force molecular docking methods are reaching their practical limits due to the major computational cost and time required20,21. Commonly used adaptations of molecular docking algorithms to multicore CPUs clusters or GPUs22,23 reduce computational time requirements for SBVS, but do not fully resolve the scaling problem due to the high cost of such computations. Several adaptive screening approaches have been suggested to reduce the cost by exploring synthon-based chemical spaces such as Enamine REAL, while still detecting the majority of high-scoring hits. This includes multi-stage filtering methods implemented in VirtualFlow and VirtualFlow25,24 and active learning (AL)-based approaches25–27.
Many of these approaches, however, are influenced by factors such as the acquisition function, batch size, model architecture, and library diversity, which require customization of the workflow for each screening campaign. Adoption of deep learning/AI approaches like DiffDock28 and co-folding (AlphaFold-329, Boltz-230, etc.) is an attractive alternative to physics-based SBVS. But these methods are even more computationally demanding, and generalization issues limit their applications to pharmacologically well-characterized targets. Moreover, these methods still require explicit enumeration of the whole library and brute-force molecular docking of the training set, typically around 1% of the entire library, which limits scalability to tens of billions of molecules31.
A conceptually different solution for ensuring the scalability of SBVS was introduced by synthon-based approaches like V-SYNTHES14. It takes full advantage of the modular organization of the combinatorial spaces based on their building blocks. The initial step involves docking of the so-called Minimal Enumeration Library (MEL) of fragments that represent all unique synthon-scaffold combinations and prioritizing the best MEL fragments for subsequent enumeration and molecular docking of larger derivatives of fragment hits. The V-SYNTHES approach was validated in SBVS of an 11-billion-size version of Enamine REAL Space, identifying qualified nanomolar hits and leads for the CB2 receptor and ROCK1 kinase. Efficiency and accuracy of the V-SYNTHES approach and its updated version V-SYNTHES232 were further validated across multiple diverse biological targets33,34. A similar synthon-based concept was reported for the Chemical Space Docking framework35. The latter, however, employs synthon fragments directly and differs methodologically from the V-SYNTHES approach. The growing interest in synthon-based approaches is further reflected in the recent introduction of exaScreen and exaDock, novel ligand-based and structure-based VS methodologies, respectively, that rely on the 3D conformation of a reference compound to prioritize synthons in ultra-large chemical spaces36.
Building on the V-SYNTHES2 framework, the present study addresses three questions not previously explored at this scale: whether V-SYNTHES2 maintains computational efficiency when the chemical space expands 16-fold to 173 billion compounds; whether the approach identifies confirmed hits on targets with no established pharmacology under blinded, independently validated conditions; and whether the synthon modularity of xREAL Space supports systematic SAR exploration around confirmed hits, and where its current limitations lie. In this study, we used the V-SYNTHES2 algorithm, expanded to screen 173 billion molecules from the Enamine xREAL Space against two targets, SARS-CoV-2 NSP13 helicase and CBLB, selected for CACHE Challenges #2 and #4, respectively. Both targets lack well-defined ligand binding pockets and existing ligand pharmacology, representing precisely the conditions where data-driven, ML-based approaches would face fundamental limitations due to insufficient training data. These pockets thus pose extreme hit-finding challenges for physics-based methods as well, as reflected in the low average hit rates across CACHE participants. V-SYNTHES2 screening of xREAL Space successfully identified several novel hit molecules for both targets. The hits were rigorously and independently validated in vitro as a part of the CACHE assessments. We further explored the xREAL Space for hit expansion and structure-activity relationships (SAR) studies.
Results
Scale of the V-SYNTHES approach
The version of Enamine xREAL Space used in this study contains 173 billion compounds generated through a combination of 103,879 unique building blocks and 165 chemical reactions. This includes 2.39 billion molecules from 2-component reactions and 171 billion molecules from 3-component reactions.
Compared with the 11-billion-size version of Enamine REAL Space, the total chemical space increased 16-fold. This rapid expansion was achieved by increasing the number of building blocks and the number of chemical reactions eligible for parallel synthesis. As expected, the majority of that growth is attributed to the 3-component space, which increased from 10 billion to over 171 billion molecules (Fig. 1c, d, Supplementary Table 1). The space of 2-component reactions became a much smaller percentage of the total chemical space, down from 9% in the 11 billion-size version of Enamine REAL Space to 1% in the xREAL Space, as shown in Supplementary Table 1. Even though the number of building blocks of 2-component space grew much more than that of 3-component space and now accounts for 99% of the total number of building blocks, up from 82%, this growth did not translate into a proportional increase in molecular diversity, as the vast majority of the chemical space expansion is driven by 3-component reactions.
Fig. 1. Chemical space distributions and library composition of the REAL and xREAL chemical spaces.

Distributions are shown for 1 million randomly enumerated compounds sampled from the 11-billion-compound REAL and 173-billion-compound xREAL chemical spaces, visualized using Uniform Manifold Approximation and Projection (UMAP). a Comparison of 2-component reaction spaces. Compounds from the REAL Space are shown as blue circles; compounds from the xREAL Space are shown as orange circles. Both libraries populate similar regions of chemical space; the xREAL Space shows denser clustering, reflecting the increased number of building blocks in this subset. b Comparison of 3-component reaction spaces. Compounds from the REAL Space are shown as blue circles; compounds from the xREAL Space are shown as orange circles. The xREAL Space occupies a broader area and exhibits substantially greater structural diversity, consistent with the dominant contribution of 3-component reactions to the overall xREAL expansion. c Size comparison of 2-component subspaces of the 11-billion-compound REAL and 173-billion-compound xREAL chemical spaces, showing the number of enumerated compounds, unique building blocks, and MEL fragments in each library. d Size comparison of 3-component subspaces of the 11-billion-compound REAL and 173-billion-compound xREAL chemical spaces, showing the number of enumerated compounds, unique building blocks, and MEL fragments in each library. UMAP Uniform Manifold Approximation and Projection, REAL Space REadily AccessibLe Space, xREAL Space expanded REadily AccessibLe Space, MEL Minimal Enumeration Library.
In the xREAL Space, 2-component molecules have largely retained properties similar to those from the 11 billion-size version of Enamine REAL Space (Supplementary Fig. 1). At the same time, 3-component xREAL Space was expanded towards larger molecules with a median molecular weight of ~550 Da and 90% of molecules within 420–690 Da. The V-SYNTHES2 approach opens the door to screening larger molecules in this xREAL Space as needed when targeting larger, more open binding pockets, considered highly challenging in drug discovery. To compare the chemical diversity of Enamine REAL Space (11 billion compounds, 2019) with the expanded Enamine xREAL Space (173 billion compounds, 2022), we visualized the distribution of compounds in chemical space using Uniform Manifold Approximation and Projection (UMAP). Fig. 1 displays 1 million randomly enumerated compounds from 2-component (panel a) and 3-component (panel b) reaction spaces.
In the 2-component space, compounds from both the REAL and xREAL libraries populate similar regions of chemical space. However, the xREAL Space shows denser clustering, reflecting the increased size of this subset.
In contrast, the 3-component reaction space reveals a striking difference. Compounds from xREAL occupy a broader area and exhibit substantially greater structural diversity compared to those from REAL. Since 3-component reactions account for a large share of the xREAL Space, this expansion significantly enhances its chemical diversity, making xREAL particularly valuable for hit discovery in giga-scale virtual screening.
As shown in Table 1, the computational cost required to virtually screen the full 173 billion space using the V-SYNTHES2 algorithm increased by only 10–15% compared to screening of 11 billion REAL Space. The increased cost comes from the growing number of building blocks, which translates into a larger MEL. In this study, we maintained the same number of fully enumerated molecules selected for final docking (1 million compounds). Even though the size of MEL increased from 659,000 compounds to 2 million molecules, the total computational cost required to screen 2 million MEL is still low in comparison to the fully enumerated library, as the MEL candidates are small fragment-like molecules with a low number of rotatable bonds that take 8 s on average to dock using the standard V-SYNTHES2 settings (docking effort 2).
Table 1.
V-SYNTHES2 computational cost based on benchmark calculationsa
| 11B | 173B | ||||
|---|---|---|---|---|---|
| MEL (+2nd enum for 3-comp) | Full molecules | MEL (+2nd enum for 3-comp) | Full molecules | ||
| 2-comp | hours | 1.2–1.3 | 5.4–8.4 | 4.5–4.6 | 5.4–8.4 |
| % of the total cost | 4–5 | 22–26 | 14–19 | 22–26 | |
| 3-comp | hours | 6.2–7.2 | 8–12 | 6.1–7.2 | 8–12 |
| % of the total cost | 22–26 | 33–37 | 23–26 | 33–37 | |
| Total time | hours | 21–29 | 24–32 | ||
aAs performed on a Linux cluster or computing cloud with 1000 CPU 3GHz cores.
Selection of SARS-CoV-2 NSP13 structural model
As part of CACHE Challenge #2, the V-SYNTHES2 approach was employed to screen 173 billion xREAL Space compounds against the SARS-CoV-2 helicase NSP13. NSP13 is a helicase of the superfamily 1B that utilizes nucleotide triphosphate hydrolysis to drive the 5′→3′ separation of double-stranded DNA or RNA. NSP13 was selected as a target in this challenge due to its critical role in SARS-CoV replication. Notably, the RNA-binding pocket of NSP13 is among the most highly conserved regions across coronaviruses, making it an attractive and potentially broad-spectrum antiviral target37,38.
The four structures solved as part of crystallographic screening for fragments targeting NSP13 (PDB IDs: 5RLH, 5RLZ, 5RML, 5RMM)38 were suggested by the CACHE Challenge organizer for virtual screening. While 5RLH, 5RLZ, and 5RMM all feature fragments bound in the RNA-binding pocket, 5RML contains a fragment positioned far outside this region and was therefore excluded from further analysis. The co-crystallized fragments in 5RLH (K2D), 5RLZ (VWM), and 5RMM (VXG) all contain carboxylic acid groups. Notably, in 5RLZ (VWM) and 5RMM (VXG), the carboxyl groups occupy the same binding site and form a conserved hydrogen bond with N516, emphasizing the importance of this residue in ligand recognition. In contrast, although 5RLH (K2D) also contains a carboxylic acid, it adopts a distinct spatial orientation, forming hydrogen bonds with Y515 and T552. Based on this conserved binding pose and interaction, 5RLZ and 5RMM were selected for further analysis.
Due to the lack of known active compounds, it was not feasible to perform rigorous model validation or ligand-guided optimization39. Therefore, pose reproducibility for two co-crystallized fragments (PDB IDs 5RMM and 5RLZ) sharing the same carboxylic acid binding position was used as the primary criterion for evaluating the docking models. Both models showed good agreement between docked and experimental poses. However, the model based on the 5RLZ structure achieved lower RMSD values (approximately 2.0 and 1.6 Å; Fig. 2) and produced more favorable docking scores. In addition, the carboxylic acid group of the docked fragment preserved a conserved hydrogen bond with residue N516, consistent with the crystal structure. Based on this high level of pose fidelity, the 5RLZ-derived model was selected for subsequent virtual screening.
Fig. 2. Validation of NSP13 docking models based on X-ray structures 5RMM and 5RLZ.

Experimental crystallographic binding poses of co-crystallized fragments are compared with computationally docked poses in the RNA-binding pocket of SARS-CoV-2 NSP13 helicase. Experimental poses for fragments VXG (PDB ID: 5RMM) and VWM (PDB ID: 5RLZ) are shown as gray sticks; docked poses are shown as orange sticks. The model based on structure 5RLZ achieved root-mean-square deviation (RMSD) values of approximately 2.0 Å and 0.7 Å for the two fragments respectively, and was selected for subsequent virtual screening based on superior pose fidelity and more favorable docking scores. The conserved hydrogen bond between the carboxylic acid group of the docked fragment and residue N516 is indicated by black dashed lines. NSP13 non-structural protein 13, SARS-CoV-2 severe acute respiratory syndrome coronavirus 2, PDB Protein Data Bank, RMSD root-mean-square deviation, RNA ribonucleic acid.
V-SYNTHES2 screening for candidate binders to SARS-CoV-2 NSP13
The compound screening with V-SYNTHES2 consists of several key steps: docking of MEL, iterative enumeration and docking of full molecules, semi-automated postprocessing, including filtering for drug-likeness and novelty, re-scoring of top hits, and synthon-based clustering, each performed using dedicated in-house scripts, and final expert visual inspection of the top-ranked poses (see Fig. 3).
Fig. 3. Overview of the V-SYNTHES2 virtual screening workflow used in the CACHE #2 and CACHE #4 challenges.

The figure illustrates the sequential stages of the V-SYNTHES2 structure-based virtual screening pipeline as applied in this study. Stage 1: generation of the Minimal Enumeration Library (MEL) of fragment-like compounds representing all unique synthon-scaffold combinations in the Enamine xREAL Space. Stage 2a: docking of MEL fragments into the prepared receptor binding pocket using ICM-Pro v3.9. Stage 2b: filtering of docked MEL fragments based on docking score and RTCNN score thresholds. Stage 3: iterative enumeration of selected MEL fragments into approximately 1-2 million fully enumerated compounds per reaction subset by replacing capped R-groups with real synthons from the xREAL Space. Stage 4: final docking of fully enumerated compounds followed by automated post-processing including PAINS filtering, Lipinski Rule of Five filtering, toxicity and reactivity filters, synthon-based clustering, and expert visual inspection to select approximately 100 compounds for synthesis and experimental testing.
As described before14,32 the 2-component and 3-component reaction libraries were screened separately. 1.88 million MEL fragments for 2-component reactions and 120 thousand MEL for 3-component reactions were docked into the NSP13 structural model. After iterative enumeration and docking of full molecules, 90,000 top-scoring compounds were selected for postprocessing, 30,000 for 2-component and 60,000 for 3-component reactions.
Initially, compounds were filtered based on docking score (cutoff: −35.0) in combination with Radial and Topological Convolutional Neural Network (RTCNN) score40 (cutoff: −15.0), as implemented in ICM-Pro (Molsoft). The prefiltered set was further refined by eliminating PAINS, undesirable functional groups, and compounds violating Lipinski’s Rule of Five (Ro5), as required by the CACHE Challenge guidelines. The impact of each filtering step on the number of retained compounds is summarized in Supplementary Table 2. The combined application of docking score, RTCNN, physicochemical, and toxicity/reactivity filters reduced the initial 90,000 top-scored compounds to approximately 3000 candidates prior to clustering and expert visual inspection.
The resulting subset was then re-docked using increased thoroughness settings (set to 5) to generate more accurate binding poses and binding scores. Ligand–protein interactions were subsequently analyzed to identify diverse binding hypotheses for how the compounds might occupy the target pocket (Supplementary Fig. 2). Four binding hypotheses were considered in the selection strategy. The first hypothesis proposed that compounds could mimic RNA interactions within the NSP13 RNA-binding pocket. The second focused on compounds that could occlude the entrance to the RNA-binding channel, thereby potentially blocking RNA access. The third hypothesis targeted compounds capable of engaging key residues near the RNA-binding surface, specifically N516 and S486, which were selected based on their conserved hydrogen bonding interactions with the carboxylic acid functional groups of co-crystallized fragments. Finally, the fourth hypothesis involved an unbiased selection of the best-scoring compounds, prioritized solely by docking scores, without constraints on binding mode.
To increase diversity of the final selection, the compounds in each hypothesis were clustered using a custom synthon-based algorithm that leverages the combinatorial nature of Enamine xREAL Space and provides improved grouping compared to traditional fingerprint-based clustering.
Finally, docking poses of 1000 compounds were reviewed by medicinal chemist experts, and a final list of compounds (100 compounds) was selected for synthesis and in vitro testing against NSP13.
Synthesis was successful for 77 compounds (translating to 77% synthesis success) for experimental assessment (Supplementary Table 3). Overall, 11 different synthetic procedures (see Methods) were used to assemble the molecules from the selection set. All selected compounds were synthesized by Enamine with >95% purity (with the exception of 3 compounds >90% purity). Analytical characterization data for all synthesized compounds are provided in the Supplementary Information.
Characterization of new SARS-CoV-2 NSP13 ligands
Primary screening of 77 synthesized compounds was performed using surface plasmon resonance (SPR) and ATPase inhibition assay at a concentration of 50 μM. The detailed protocols of biological characterization of the synthesized molecules were published earlier by CACHE organizers41. Overall, based on relative binding in the SPR assay and inhibition in the ATPase assay, 37 (48%) and 13 (17%) compounds, respectively, were advanced to dose-response experiments. In the SPR assay, 6 (8%) compounds demonstrated Kd values of 43 μM or below, with the best compound showing an affinity of 7 μM for NSP13. In the ATPase assay, the other 6 (8%) compounds demonstrated an IC50 less than 50 μM. As described in the CACHE #2 results paper17, a very poor correlation was observed between hit compounds identified by SPR and those detected in the ATPase assay across all potential hit selections. Consequently, the CACHE organizers decided to focus on SPR binding affinity data and exclude ATPase data from subsequent analysis for all participating groups.
All six hits in Fig. 4 did not show any significant binding to the unrelated protein WDR9, nor aggregation or solubility issues, and were designated as confirmed hits. It is also worth mentioning that comparison with SARS-CoV-2 helicase NSP13 ligands listed in ChEMBL confirmed no close analogs among the confirmed hits (Tanimoto similarity < 0.4 for all six hits; Supplementary Table 3). To further substantiate novelty, a comparison against the complete ChEMBL database was performed. For five of the six confirmed hits, no analogs with Tanimoto similarity >0.6 were found in ChEMBL against any biological target. For compound CACHE2_1418_3, two analogs (Tanimoto > 0.6) with reported activity against unrelated proteins, tyrosyl-DNA phosphodiesterase and histone-lysine N-methyltransferase, were identified, while no close analogs against NSP13 or related coronavirus targets were found.
Fig. 4. Hit compounds identified in round 1 of the CACHE #2 challenge targeting SARS-CoV-2 NSP13.

a Chemical structures and predicted docking poses of the six confirmed hit compounds in the RNA-binding groove of the SARS-CoV-2 NSP13 helicase. The protein surface is shown in green; hit compounds are shown as colored sticks. Hydrogen bond interactions with key binding site residues are indicated by black dashed lines. b Dose-response surface plasmon resonance (SPR) binding curves for the six confirmed hit compounds. Each curve represents binding response as a function of compound concentration, yielding dissociation constants (Kd) ranging from 7 µM to 43 µM. Compound identifiers follow the CACHE2_1418_XX naming convention assigned by the CACHE Challenge organizers. NSP13 non-structural protein 13, SARS-CoV-2 severe acute respiratory syndrome coronavirus 2, SPR surface plasmon resonance, Kd equilibrium dissociation constant, RNA ribonucleic acid, CACHE Critical Assessment of Computational Hit-Finding Experiments.
This corresponds to an 8% hit rate for the V-SYNTHES2 approach, placing it among the top three workflows in the hit identification stage of the CACHE Challenge #2. Most of the other 23 workflows yielded hit rates between 0% and 3%, with an overall average of 2.3%. Moreover, the V-SYNTHES2 approach was among the top10 performing methods based on the aggregated score of virtual screening hits, assigned by the CACHE Committee17.
Physicochemical properties and ligand efficiency of the six confirmed hits are summarized in Supplementary Table 3. The confirmed hits span a molecular weight range of 391–493 Da with calculated logP values between 1.1 and 4.0, consistent with drug-like chemical space. Ligand efficiency (LE) values range from 0.17 to 0.24 kcal/mol per heavy atom, reflecting the challenging nature of the NSP13 RNA-binding site. All six hits demonstrated selectivity over the unrelated protein WDR9, with no significant binding observed, confirming target-specific engagement.
Analysis of the predicted binding modes of the confirmed hits in relation to the initial binding hypotheses (Fig. 5a) revealed that four compounds likely act by occluding the entrance to the RNA-binding channel—CACHE2_1418_37, CACHE2_1418_52, CACHE2_1418_3, CACHE2_1418_39—while other two appear to mimic RNA interactions within the binding pocket—CACHE2_1418_6, CACHE2_1418_68.
Fig. 5. Round 2 hit expansion results for confirmed NSP13 hits from the CACHE #2 challenge.

Active analogs from two parent hit series and SPR binding data for all 32 round 2 compounds are shown. a Chemical structures and predicted docking poses of active round 2 analogs generated from parent compound CACHE2_1418_52, shown in the RNA-binding pocket of NSP13. b Chemical structure and predicted docking pose of the active round 2 analog generated from parent compound CACHE2_1418_68. c Surface plasmon resonance (SPR) binding data for all 32 round 2 synthesized compounds and 4 parent compounds. Each series is plotted in a distinct color: CACHE2_1418_3 series in green circles, CACHE2_1418_37 series in gray circles, CACHE2_1418_52 series in blue circles, and CACHE2_1418_68 series in red circles. Within each series, the parent compound is represented by a larger shape with more intense color shading; analogs are represented by smaller shapes with lighter color shading. Although the limited number of analogs was insufficient to establish a predictive QSAR model, structural analysis of active compounds reveals interpretable SAR trends identifying attachment points for future lead optimization. SPR surface plasmon resonance, Kd equilibrium dissociation constant, SAR structure-activity relationship, QSAR quantitative structure-activity relationship, NSP13 non-structural protein 13, xREAL Space expanded REadily AccessibLe Space, HBD hydrogen bond donor, CACHE Critical Assessment of Computational Hit-Finding Experiments.
Limited hit expansion of initial V-SYNTHES2 hits for SARS-CoV-2 NSP13
To generate 4 follow-up series (CACHE2_1418_3, CACHE2_1418_37, CACHE2_1418_52, and CACHE2_1418_68) for the four confirmed hits prioritized for further study, we employed several complementary approaches using the Enamine xREAL Space. First, close analogs were identified using fragment-space search methods, including FTrees, SpaceMACS, and SpaceLight42–44. Second, we exploited the synthon modularity of the xREAL Space by iteratively anchoring one synthon (for 2-component reactions) or two synthons (for 3-component reactions) while varying the remaining synthon(s), enabling systematic exploration of structure-activity relationships (SAR). The generated analogs were docked using the setup described above in ‘Model preparation and screening for SARS-CoV-2 NSP13’.
Of the 75 analogs advanced for synthesis, 32 (43%) were successfully synthesized across 4 follow-up series (Supplementary Table 4). The relatively low synthesis success rate was primarily due to the limited availability of one of the building blocks, which prevented the completion of all intended analogs. The strict project timeline further precluded resynthesis attempts. Analytical characterization data for all synthesized compounds are provided in the Supplementary Information.
To assess the novelty of the synthesized analogs, Tanimoto similarity was calculated against all ChEMBL compounds with reported activity against SARS-CoV-2 targets. No analog exceeded a Tanimoto similarity of 0.6, confirming that the synthesized compounds represent novel chemical matter relative to known SARS-CoV-2 ligands (Supplementary Table 5).
The 32 synthesized analogs were retested in a workflow similar to round 1, for which the detailed protocols were published earlier17. In series CACHE2_1418_3 (green) and CACHE2_1418_37 (gray), the follow-up molecules did not show binding activity (Fig. 5). In contrast, 4 from series CACHE2_1418_52 (blue) and 1 compound from series CACHE2_1418_68 (red) showed binding activity in a range of 48 μM to 189 μM (Supplementary Table 4).
In series CACHE2_1418_3, ten analogs were synthesized with the goal of replacing the original acyl group with structurally diverse alternatives, including bicyclic lactams, trimethoxyphenyl, saccharin derivatives, indazoles, and chlorinated heterocycles. None of the ten analogs showed binding activity, indicating that the tested acyl group replacements were not tolerated. The precise structural requirements at this position could not be resolved from the current dataset.
In series CACHE2_1418_37, five analogs modifying the pyrimidine ether and aminophenyl regions were all inactive. The limited number of analogs and absence of any activity preclude SAR interpretation for this series.
In the CACHE2_1418_52 series, four of twelve analogs showed binding activity (Kd = 48–189 µM) and three showed a linear SPR response without saturation. Two observations are consistent across the active analogs. First, the two most potent analogs: CACHE2-HO_1418_26 (Kd = 48 µM, 4-chlorothiazole acyl group, HBD = 2) and CACHE2-HO_1418_28 (Kd = 56 µM, bicyclic thiazoloindolizinone acyl group, HBD = 2), both carry compact sulfur-containing heterocycles as direct acyl groups without a methylene spacer, with only 1.5- to 1.75-fold potency loss relative to the parent. This suggests that the methylene spacer present in the parent is not required, and that compact aromatic sulfur heterocycles are tolerated at this position. Second, analogs bearing free primary amines on the acyl group: CACHE2-HO_1418_23 (4-aminophenyl-urea acyl group, HBD = 4, Kd = 167 µM) and CACHE2-HO_1418_24 (2-amino-5-bromopyridine acyl group, HBD = 3, Kd = 189 µM), show 5- to 6-fold weaker activity than the parent despite retaining the core scaffold. Both have higher HBD counts than the active analogs, consistent with a binding environment that may not readily accommodate additional hydrogen bond donors, though the dataset is too small to establish a firm trend. CACHE2-HO_1418_27, which modifies both core scaffold and linker geometry, and CACHE2-HO_1418_30, which alters core connectivity, are inactive, consistent with the core being essential for activity.
In the CACHE2_1418_68 series, one of five analogs showed binding activity. CACHE2-HO_1418_34 (Kd = 78 µM, HBD = 1) replaces the 2-amino-4-chloro-5-hydroxyphenyl acyl group of the parent with an indazolyl group, retaining a directly attached aromatic acyl group while removing all polar substituents, with a modest 3-fold potency loss. In contrast, HO_1418_32, which retains the NH₂ group but modifies the ring system, and CACHE2-HO_1418_33 and CACHE2-HO_1418_35, which introduce a methylene linker to non-aromatic acyl groups, are all inactive. CACHE2-HO_1418_36, with a directly attached fluorobenzochromanol acyl group, is also inactive. These observations are consistent with a requirement for direct aromatic acyl group attachment in this series, though the small number of analogs prevents firm conclusions about the contribution of individual substituents. The absence of potency improvement over the parent compounds in any of the four series reflects the combined effect of incomplete analog synthesis due to building block availability constraints, the limited number of analogs tested, and the inherent uncertainty of SPR measurements at these affinity levels — factors consistent with an early-stage SAR exploration rather than a complete optimization campaign. SAR findings across all four series are summarized in Supplementary Table 6.
Applying the V-SYNTHES2 workflow to CBLB
To further test applicability of our method, we employed the V-SYNTHES2 algorithm to screen for inhibitors of the tyrosine kinase binding (TKB) domain of E3 ubiquitin-protein ligase CBL-B (CBLB) in the CACHE Challenge #4. CBLB is an E3 ubiquitin ligase and an important target in cancer immunotherapy, as it plays a key role in the negative regulation of T cell activation45,46. CBLB consists of several domains, a short linker helix region (LHR), a substrate recognition tyrosine kinase-binding domain (TKBD), and a RING finger domain, which associates with E2-ubiquitin complexes. In the inactive (closed) conformation, residues from the TKBD domain and LHR interact with each other, resulting in the RING finger domain being positioned opposite the substrate binding site. After phosphorylation, the protein converts to the active (open) conformation, undergoing significant conformational changes that enable more efficient interaction between the RING finger domain and E2 proteins47. Thus, it is important to maintain the closed conformation of CBLB through small-molecule inhibition.
The crystal structure of CBLB in complex with inhibitor C7683 (PDB ID: 8GCY47, resolution 1.81 Å) was suggested for virtual screening by the CACHE Challenge. To validate the docking model based on the crystal structure with accession number 8GCY, we assembled a dataset comprising 29 active compounds and 1,450 decoy molecules. The actives were provided by the CACHE Challenge #4 organizers, while the decoys were generated using the DUD-E webserver.
Docking was performed using this benchmark set, and the model’s performance was assessed by generating a ROC curve and calculating the area under the curve (AUC). The model achieved an AUC of 96%. The exceptionally good docking results should be considered with the caveat that the active compounds share a similar scaffold, which reduces chemical diversity. Additionally, for the selected model redocking study was performed showing reproducibility with the co-crystallized bound ligand (Fig. 6).
Fig. 6. Validation of the CBLB docking model by binding pose comparison and active-decoy enrichment analysis.

a Comparison of the experimental crystallographic binding pose of inhibitor C7683 with the computationally docked pose in the tyrosine kinase-binding (TKB) domain of CBLB (PDB ID: 8GCY, resolution 1.81 Å). The binding pose from the crystal structure is shown as gray sticks; the pose obtained after molecular docking into the prepared docking model is shown as orange sticks. b Receiver operating characteristic (ROC) curve generated from docking of 29 active compounds and 1,450 DUD-E decoy molecules into the validated CBLB docking model. The solid curve represents model performance; the diagonal dashed line represents random performance (AUC = 50%). The model achieved an area under the curve (AUC) of 96%, indicating strong enrichment of active compounds, with the caveat that active compounds share a similar scaffold. CBLB Casitas B-cell lymphoma-b, TKB tyrosine kinase-binding domain, PDB Protein Data Bank, ROC receiver operating characteristic, AUC area under the curve, DUD-E Directory of Useful Decoys Enhanced.
Selection and synthesis of candidate hits for CBLB
We performed V-SYNTHES2 screening using a protocol similar to that above, with only minor changes to the CBLB procedure. We docked 1.88 million fragments from 2-component reactions and 120 thousand fragments from 3-component reactions into the prepared protein model. The best-scoring MEL fragments were enumerated into 1.5 million full molecules from each sub-space, and the resulting full compounds were docked. We obtained 30,000 top-scoring 2-component compounds and 60,000 3-component compounds for further processing.
The candidate selection process for synthesis involved several steps. The combined application of docking score, RTCNN, physicochemical, and toxicity/reactivity filters reduced the initial 90,000 top-scored compounds to approximately 2000 candidates prior to clustering and expert visual inspection (Supplementary Table 2). First, compounds were filtered using a docking score (cutoff –40.0) and an RTCNN score (cutoff –20.0). After this, we applied common filters such as PAINS, Rapid Elimination of Swill (REOS), Lipinski’s Rule of Five, and BadGroups (ICM-Pro) to remove compounds with unsuitable physicochemical properties, unwanted functional groups, and false-positive results.
Next, the top-scoring compounds were divided into three groups based on interactions with key residues Y363, which is crucial for the regulation of CBLB activity, and F263, which is located in the TKB domain. The first group comprises molecules that form hydrogen bonds with both F263 and Y363, which may contribute to the stability of the TKB domain and the LHR. The second group includes ligands that form hydrogen bonds only with Y363 (Supplementary Table 7). The third group consists of compounds that do not form hydrogen bonds with the key residues, but that show good docking scores and occupy a deeper region of the binding site. To ensure ligand diversity, we applied a synthon-based approach to cluster the compounds. The docking poses of the remaining 1000 molecules were then reviewed by experts.
As a result, 89 compounds were prioritized for synthesis, of which 69 were synthesized by Enamine (77.5% synthetic success rate) and tested in vitro against CBLB (Supplementary Table 7). Overall, 6 different synthetic procedures were used to assemble the molecules from the selection set (Methods). The selected compounds were synthesized by Enamine with >95% purity. Analytical characterization data for all synthesized compounds are provided in the Supplementary Information.
Characterization of new CBLB ligands
Fluorescence polarization (FP) assay was used for primary screening of 69 selected compounds at a concentration of 50 μM (Fig. 7a). Three compounds displaced the CBLB ligand, reducing the assay response to less than 85% of control, and these initial hits were advanced to a dose-response test. The single compound CACHE4HI_1955_60 (Z8710104284) demonstrated a dose-response curve with Kd of 111.9 μM that passed hit selection criteria (Kd < 150 μM) with a ligand efficiency of 0.21 kcal/mol per heavy atom. The compound was then counterscreened in the same dose-response concentration range against unrelated protein GID4 with no activity observed (Fig. 7c, d). After aggregability and solubility measurement with DLS, CACHE4HI_1955_60 met the predefined hit criteria and was advanced to hit expansion. Neither hit candidate, nor other tested compounds were reported for CBLB in ChEMBL (Supplementary Table 7). Notably, only 10 compounds out of 1,688 screened across 23 different CACHE participants were assigned as hit compounds and advanced to the hit expansion stage, each originating from a distinct group.
Fig. 7. Identification and characterization of the hit candidate for CBLB in the CACHE #4 challenge.

a Fluorescence polarization (FP) primary screening results for 69 synthesized compounds tested against CBLB at 50 µM. The horizontal dashed line indicates the 85% response threshold; filled circles below this threshold represent compounds advanced to dose-response experiments. b Chemical structure of the hit candidate CACHE4HI_1955_60 and its predicted binding pose in the TKB domain of CBLB. The nitrogen atom at the para position of the phenyl ring forms a hydrogen bond with residue S368, indicated by a green dashed line. c FP dose-response curve of CACHE4HI_1955_60 against CBLB, yielding a dissociation constant (Kd) of 111.9 µM (n = 2, technical replicates). The compound met the predefined hit selection criterion of Kd < 150 µM. d FP counter-screen dose-response curve of CACHE4HI_1955_60 against the unrelated protein GID4, showing no significant displacement of the substrate across the tested concentration range (n = 2, technical replicates), confirming selectivity of binding. FP fluorescence polarization, Kd equilibrium dissociation constant, CBLB Casitas B-cell lymphoma-b, TKB tyrosine kinase-binding domain, GID4 GID complex subunit 4, n number of replicates, para position 4 of the phenyl ring, CACHE Critical Assessment of Computational Hit-Finding Experiments.
Limited hit expansion of initial V-SYNTHES2 hits
The discovered hit CACHE4HI_1955_60 was obtained from a 2-component reaction. Follow-up molecules were generated by varying the synthons present in the hit molecule. In addition, we utilized methods such as FTrees, SpaceMACS, and SpaceLight to search for similar compounds in the REAL Space. However, the scaffold of the hit compound is poorly presented in Enamine xREAL Space, resulting in a limited number of analogs. After molecular docking, 26 analogs were prioritized for synthesis, of which 13 (50%) were successfully synthesized (Supplementary Table 8). As with the NSP13 series, the limited availability of one of the building blocks prevented completion of all intended analogs, and the strict project timeline did not permit resynthesis. The synthetic protocols are described in Methods, and QC data are provided in the Supplementary Information. All compounds were confirmed to have no close analogs (Tanimoto similarity > 0.6) in ChEMBL (Supplementary Table 8).
The synthesized compounds did not displace the substrate in the FP assay (Fig. 8c). A detailed analysis of these results revealed that all tested analogs differed from the initial hit in the position of the nitrogen atom (see Supplementary Information). In the initial hit (CACHE4HI_1955_60), the nitrogen atom is located at the para position of the phenyl ring, which, according to docking results, enables the formation of a hydrogen bond with S368 (Fig. 8a). The importance of this hydrogen bond was confirmed by testing the closest analog, CACHE4HO_1955_5 (Z9151908629), in which the nitrogen atom is located at the meta position. This positional change led to the loss of the hydrogen bond and, consequently, reduced activity in the FP assay (Fig. 8b), an important result for further optimization of this hit.
Fig. 8. Structure-activity relationship analysis of round 2 analogs for the CBLB hit candidate.

a Chemical structures of the round 1 hit candidate CACHE4HI_1955_60 (left) and its closest round 2 analog CACHE4HO_1955_5, highlighting the key structural difference: the nitrogen atom occupies the para position of the phenyl ring in CACHE4HI_1955_60 and the meta position in CACHE4HO_1955_5. b Superimposed predicted binding poses of CACHE4HI_1955_60 (shown as green sticks) and CACHE4HO_1955_5 (shown as gray sticks) in the TKB domain of CBLB. The para-nitrogen of CACHE4HI_1955_60 forms a hydrogen bond with residue S368, indicated by a green dashed line; the meta-nitrogen of CACHE4HO_1955_5 is positioned away from S368, abolishing this interaction and confirming the critical role of nitrogen position for target engagement. c Fluorescence polarization (FP) screening results for all 13 round 2 analogs tested at three concentration points, shown as filled circles. None of the round 2 analogs displaced the CBLB substrate above the activity threshold, consistent with loss of the key hydrogen bond with S368 observed in the docking analysis. FP fluorescence polarization, CBLB Casitas B-cell lymphoma-b, TKB tyrosine kinase-binding domain, SAR structure-activity relationship, para position 4 of the phenyl ring, meta position 3 of the phenyl ring, CACHE Critical Assessment of Computational Hit-Finding Experiments.
Discussion
This work establishes three findings that extend beyond the previously published V-SYNTHES2 framework. First, the method scales to 173 billion compounds, a 16-fold increase in chemical space, with only a 10–15% increase in computational cost, a direct consequence of the MEL architecture that insulates total docking effort from library size growth (Table 1). Second, a confirmed hit rate of 8% for NSP13 substantially exceeds the CACHE Challenge #2 average of 2.3%, placing V-SYNTHES2 among the top three workflows. For CBLB, one compound meeting predefined FP-based criteria was identified (Kd = 111.9 µM, FP assay) and advanced to hit expansion, demonstrating reliable performance on targets with no prior pharmacology, precisely the conditions under which data-driven methods face the greatest limitations. Third, independent experimental validation through the CACHE infrastructure provides a high-confidence performance benchmark, free from the post-hoc selection bias that can affect self-validated computational studies.
The hit-finding performance reported here is noteworthy in the context of how the virtual screening was conducted. Co-crystallized fragments (for NSP13) and known ligands (for CBLB) were used to validate the docking model and confirm the binding pocket geometry prior to screening, a standard and appropriate step when cognate ligands are available, which increases confidence in the structural model but may bias the search toward that binding site. Beyond this structural validation step, however, no ligand similarity information was used to guide compound selection or scoring45–48. The confirmed hits were therefore identified through physics-based scoring alone, without explicit optimization toward known chemotypes. This distinction is relevant because physics-based methods of this type are applicable even when no ligand data exist, as demonstrated in recent prospective campaigns where hits were identified for targets entirely lacking known binders29. ML-based approaches, by contrast, require sufficiently large and well-curated training datasets for the target or close homologs, which fundamentally limits their applicability to novel, ligand-depleted targets.
Hit identification represents only the initial phase of the drug discovery pipeline. Subsequent stages require the ability to efficiently explore vast combinatorial chemical spaces to enable hit expansion and optimization. Several strategies can support this exploration in giga-scale combinatorial spaces. Notably, FTrees, SpaceMACS, and SpaceLight are successfully utilized for hit expansion in the synthon spaces, including Enamine xREAL Space. Importantly, the V-SYNTHES2 extension synthon-based space search algorithm can be utilized for in-depth exploration of derivatives from identified scaffolds. The modularity of the REAL Space framework offers unique advantages in this regard, enabling the generation of tightly focused SAR series around validated hits. This level of precise SAR elaboration is more difficult to achieve using conventional fingerprint-based similarity approaches such as Morgan fingerprints. The hit expansion results from this study illustrate both the opportunities and the current limitations of this approach. For NSP13, analogs of four confirmed hits were generated and tested, with activity confirmed in two series across a range of 48–189 µM. While no improvement in potency over the parent compounds was observed, likely reflecting the limited number of analogs synthesized and the constraints imposed by the availability of building blocks and project timelines, the data nonetheless identify substitution-tolerant positions and provide actionable vectors for future optimization. For CBLB, the scaffold of the hit candidate was poorly represented in xREAL Space, restricting analog generation to a single close structural neighbor. Critically, even this minimal structural change, a shift in nitrogen position from para to meta, abolished activity entirely, demonstrating that the hit occupies a narrow SAR optimum and highlighting the sensitivity of this chemotype to substitution pattern. Together, these results reflect a broader finding: while xREAL Space provides strong coverage for initial hit identification, the representation of certain scaffolds remains uneven, and the availability of building blocks can constrain SAR exploration for chemically unique hits. Addressing this limitation is an active area of development. Make-on-demand building blocks (MADE), numbering in the billions, offer both synthetic tractability and cost efficiency, enabling construction of target-focused libraries that substantially extend the chemical space accessible around validated hits. As xREAL Space continues to expand toward the trillion-compound scale, combining the synthon-based efficiency of V-SYNTHES2 with active learning strategies represents a natural next-generation direction.
Methods
Preparation of structure-based docking models
Crystal structures of the two targets, SARS-CoV-2 NSP13 and CBLB (PDB IDs: 5RLZ and 8GCY, respectively), were prepared for virtual screening using ICM-Pro v3.9 (Molsoft). Each structure was converted from PDB coordinates to internal ICM objects via the ICM-Pro conversion tool, which included restoration of missing heavy atoms and hydrogens, local minimization of polar hydrogens, and optimization of the protonation states and rotamers of His, Asn, and Gln side chains.
Docking boxes were defined based on the positions of co-crystallized ligands, with box dimensions of 24 × 21 × 21 Å for 5RLZ and 30 × 24 × 25 Å for 8GCY.
Model validation and virtual screening
The docking model for SARS-CoV-2 NSP13 was validated by minimizing the root-mean-square deviation (RMSD) between the co-crystallized fragments and their redocked poses. In addition, the docking model for CBLB was validated by maximizing the area under the curve (AUC) of the ROC plot and minimizing the RMSD between the co-crystallized ligand and its predicted pose.
Virtual screening was performed using the V-SYNTHES2 algorithm as described here14,32. The docking studies were performed in ICM-Pro v3.9. To ensure comprehensive sampling, the top 30,000 compounds identified during the initial V-SYNTHES2 screening with docking thoroughness set to 2, were re-docked with increased thoroughness (set to 5) prior to final hit selection for experimental testing.
Biological assessment for SARS-CoV-2 NSP13
Biological procedures for the assessment of compounds targeted at SARS-CoV-2 NSP13 were published in detail in the CACHE assessment paper17.
Chemical synthesis
All chemicals for the synthesis of the studied compounds were provided by Enamine Ltd. (www.enamine.net). All solvents were treated according to standard methods. LC/MS analysis was performed utilizing an Agilent 1200 Series LC/MSD system with DAD/ELSD (column Zorbax SB-C18 1.8 µm 4.6 × 15 mm; solvent A (water, 0.1% formic acid) and solvent B (acetonitrile, 0.1% formic acid); gradient 0–100% solvent B, run time, 1.8 min; flow rate, 3 mL/min) and Agilent LC/MSD SL (G6130A) or SL (G6140A) mass spectrometer (APCI mode). All the LC/MS data were obtained using positive/negative mode switching.
Method 1. An aryl halide (100 mg), boronic derivative (1 mol equiv. to the aryl halide), sodium carbonate (1.5 mol equiv. to the aryl halide), 1,1′- [bis(diphenylphosphino)ferrocene]-dichloropalladium(II) (0.05 mol equiv. to the aryl halide), and a water-dioxane mixture 1:3 (1 mL) were placed in a vial. The vial was then placed into a thermostat and stirred for 24 h at 95 °C. The solvent was further evaporated. The crude sample was dissolved in 0.5 mL of DMSO and purified utilizing preparative HPLC.
Method 2. A diketone (100 mg), an aldehyde (1 mol equiv. to the ketone), and a urea (1 mol equiv. to the ketone) were dissolved in 0.6 mL of DMF and thoroughly mixed. Then, trimethylchlorosilane (1.1 mol equiv. to the ketone) was added dropwise, and the reaction mixture was kept shaken for 48 h at room temperature. The solvent was evaporated under reduced pressure, and the crude was dissolved in 0.5 mL of DMSO. The sample was further purified by preparative HPLC.
Method 3. An amino acid ester (100 mg), DIPEA (1.2 mol equiv. to the ester), and DMSO (0.5 mL) were placed into a 4 mL capped glass vial and stirred for 30 min. After the addition of an alkyl halide (1 mol equiv. to the ester), the vial was kept for 12 h at 100 °C. Then, CHCl3 (2 mL) was added, and the organic phase was washed with water, dried over Na2SO4, and evaporated under reduced pressure. The sample was dissolved in 0.5 mL of DMSO and hydrolyzed with 0.2 mL of 4 M KOH for 12 h at room temperature. Then, 50 μL of formic acid was added, and the reaction vial was shaken for 30 min at room temperature. The solvent was evaporated under reduced pressure, and the crude was dissolved in 0.5 mL of DMSO. The sample was further purified by preparative HPLC.
Method 4. An amine (100 mg), an acid (1.1 mol equiv. to the amine), and 0.5 mL of DMSO were placed into a 4 mL capped glass vial, and the mixture was stirred for 30 min. Then EDC (1.2 mol equiv. to the amine) was added, and the mixture was stirred for 1 h. If the solution was transparent, the mixture was left overnight at room temperature as is; otherwise, the vial was placed in the ultrasonic bath and left overnight. The solution was filtered, and the solvent and volatile components were evaporated under reduced pressure to give the crude product. The product was further purified by HPLC.
Method 5. An alkyl halide carboxylic acid ester (100 mg), DIPEA (1.2 mol equiv. to the ester), and DMSO (0.5 mL) were placed into a 4 mL capped glass vial and stirred for 30 min. After the addition of an amine (1 mol equiv. to the ester), the vial was kept for 8 h at 100 °C. Then, CHCl3 (2 mL) was added, and the organic phase was washed with water, dried over Na2SO4, and evaporated under reduced pressure. The sample was dissolved in 0.5 mL of DMSO and subjected to hydrolysis with 0.2 mL of 4 M KOH for 12 h at room temperature. Then, 50 μL of formic acid was added, and the reaction vial was shaken for 30 min at room temperature. The solvent was evaporated under reduced pressure, and the crude was dissolved in 0.5 mL of DMSO. The sample was further purified by preparative HPLC.
Method 6. An amine (100 mg) was dissolved in 0.5 mL of DMF and shaken for 20 min at room temperature. Then, an aldehyde (1.1 mol equiv. to the amine) was added. The resulting mixture was heated for 5 h at 100 °C. After cooling down to room temperature, methanol (1 mL) was added, followed by NaBH4 (2 mol equiv. to the amine). The reaction vial was left shaken overnight. The solvent was removed under reduced pressure, and the crude sample was purified by preparative HPLC.
Method 7. A methylene-active compound (100 mg) and an aldehyde (1 mol equiv. to the methylene-active compound)were dissolved in 0.6 mL of DMF and thoroughly mixed. Then, trimethylchlorosilane (1.1 mol equiv. to the methylene-active compound) was added dropwise, and the reaction mixture was kept in a sealed vial for 48 h at 100 °C. After cooling down to room temperature, 0.2 mL of DIPEA and 2 mL of methanol were added, and the vial was shaken at room temperature for 30 min. The solvent was further evaporated. The crude sample was dissolved in 0.5 mL of DMSO and purified utilizing preparative HPLC.
Method 8. The synthesis was performed according to the previously published procedure (DOI: https://doi.org/10.1002/ejoc.201501579).
Method 9. An amino acid (100 mg) was dissolved in 0.6 mL of DMSO and 0.1 mL of 1 M KOH and shaken for 20 min at room temperature. Then, an aldehyde (1.1 mol equiv. to the amino acid) was added. The resulting mixture was heated for 5 h at 100 °C. After cooling down to room temperature, methanol (1 mL) was added, followed by NaBH4 (2 mol equiv. to the amino acid). The reaction vial was left shaken overnight. Then, acetic acid was added to the neutral pH; the solvent was further removed under reduced pressure, and the crude sample was purified by preparative HPLC.
Method 10. An amine (100 mg) and an isothiocyanate (1 mol equiv. to the amine) were dissolved in 0.6 mL of acetonitrile and kept in a sealed vial in an oven for 30 min at 100 °C. Then, a maleimide (1 mol equiv. to the amine) was added, and the reaction mixture was kept in the oven for 6 h. After cooling down to room temperature, 0.2 mL of DIPEA and 2 mL of methanol were added, and the vial was shaken at room temperature for 30 min. The solvent was further evaporated. The crude sample was dissolved in 0.5 mL of DMSO and purified utilizing preparative HPLC.
Method 11. An amine (100 mg) and DIPEA (1.1 mol equiv. to the amine; additional equivalents were added when the amine used was in salt form) were dissolved in 0.5 mL DMSO. The vial was then shaken for 20 min at rt. An alkyl chloride (1 mol equiv. to the amine) was then added. The vial was sealed and stirred for 1 h. Then, the solution was heated for 8 h at 100 °C. After cooling down, the mixture was filtered; the solvent and volatile components were evaporated under reduced pressure to give the crude product. The product was further purified by HPLC.
Method 12. The synthesis was performed according to the previously published procedure (DOI: https://doi.org/10.1016/j.tet.2011.03.101 and https://doi.org/10.1021/co500025f).
Method 13. An amine (100 mg), DIPEA (1.2 mol equiv. to the amine), and DMSO (0.5 mL) were placed into a 4 mL capped glass vial and stirred for 30 min. After the addition of an aryl halide (1.2 mol equiv. to the amine), the mixture was stirred for 1 h at rt. Then, the vial was placed into a thermostat (set to 100 °C) for 9 h. After cooling down, the mixture was filtered; the solvent and volatile components were evaporated under reduced pressure to give the crude product. The product was further purified by HPLC.
Method 14. An amine (100 mg), DIPEA (1.2 mol equiv. to the amine), and DMSO (0.5 mL) were placed into a 4 mL capped glass vial and stirred for 30 min. After the addition of an alkyl halide (1.2 mol equiv. to the amine), the vial was stirred for 12 h at rt. Then the vial was placed into a thermostat (set to 100 °C) for 9 h. After cooling down, the mixture was filtered; the solvent and volatile components were evaporated under reduced pressure to give the crude product. The product was further purified by HPLC.
Biological assessment for CBLB
Biological assays were performed at the Structural Genomics Consortium, University of Toronto, and University Health Network. The experimental protocol was provided by the Structural Genomics Consortium and is described below. The compounds were dissolved in 100% DMSO. Fluorescence polarization (FP) assay was performed by addition of the compounds in DMSO to buffer (20 mM HEPES pH 7.4, 150 mM NaCl, 0.01% Triton X-100, 0.5 mM TCEP, 2% DMSO) that contained 625 nM of the CBLB protein with C-terminal Avi and N-terminal His tags (CBLB_3ZNI: 34-427—CBLB: MVC032-A08: P222444) and 10 nM of the BODIPY-labeled substrate C2787102 and incubation for 30 min. For the dose-response experiment, the compounds were titrated in pure DMSO (20 points, 200-0.0004 μM, double-dilution) and tested at a single concentration. FP counter screen was performed using an analogous procedure: serial dilution of compounds was added to buffer (50 mM Tris pH 7.5, 0.01% Triton X-100) that contained 5 μM of the GID4 protein (GID4_116-300—GID4: MVC011-A03: P222043) and 40 nM of the substrate, C-terminally FITC-labeled PGLWKS peptide. DLS assay was performed by adding the compounds in buffer (20 mM HEPES, pH 7.5, 100 mM NaCl, 0.5 mM TCEP, 4% DMSO) using the same protocol and equipment used for SARS-CoV-2 NSP13.
For round 2, the same setup was used for the FP assay. For surface plasmon resonance (SPR) assay, biotinylated CBLB (CBLB_3ZNI: 34-427—CBLB: MVC032-A08: P222444) was immobilized onto the flow cell of a streptavidin-conjugated chip at 4000-5000 RU. Compounds were added to the buffer (20 mM HEPES, pH 7.5, 100 mM NaCl, 0.5 mM TCEP, 0.01% Triton X-100, 4% DMSO) that was then injected over the surface at a flow rate of 40 μL/min. The multicycle kinetics mode with 60 s contact time and 120 s dissociation time was used to generate sensorgrams. An analogous procedure was performed for counter screening with immobilized at 2500 RU GID4 (GID4_116-300—GID4: MVC011-A03: P222043).
Supplementary information
Acknowledgements
No funding was received for this research. The authors thank the organizers of the CACHE (Critical Assessment of Computational Hit-finding Experiments) initiative for providing the framework for Challenges #2 and #4. Compounds selected in this study were tested experimentally at the Structural Genomics Consortium, University of Toronto and University Health Network. For CACHE Challenge #2, contributions to biological assessment from Cheryl Arrowsmith, Albina Bolotokova, Peter J. Brown, Irene Chau, Kristina Edfeldt, Pegah Ghiabi, Elisa Gibson, Rachel Harding, Zarah Hejazi, Oleksandra Herasymenko, Scott Houliston, Ashley Hutchinson, Carla Kharadjian, Helen Li, Peter Loppnau, Sumera Perveen, Matthieu Schapira, Almagul Seitova, and Madhushika Silva are acknowledged. For CACHE Challenge #4, contributions from Cheryl Arrowsmith, Albina Bolotokova, Irene Chau, Kristina Edfeldt, Elisa Gibson, Levon Halabelian, Rachel Harding, Oleksandra Herasymenko, Scott Houliston, Ashley Hutchinson, Peter Loppnau, Matthieu Schapira, Almagul Seitova, Madhushika Silva, Dakota Treleaven, and Hong Wu are gratefully acknowledged.
Author contributions
O.T. and V.K. conceived and designed the study. M.P., O.S., M.V., A.V.S., A.A.S., K.H., O.H., A.K., and A.N. performed the computational work, including virtual screening, molecular docking, structural model preparation and validation, postprocessing, hit selection and expansion. D.R. supervised the parallel chemical synthesis and quality control of the tested compounds. M.P., O.S., A.V.S., and K.H. analyzed the data and prepared the figures and tables. O.T., V.K., M.P., A.V.S., O.S., and K.H. wrote the manuscript. O.T. and V.K. supervised the project. All authors reviewed and approved the final manuscript.
Data availability
The datasets generated and/or analyzed during the current study are not publicly available in full because the virtual screening was performed against the Enamine xREAL Space, a proprietary chemical space that cannot be redistributed. Key data necessary to interpret the findings, including hit compound structures, physicochemical properties, ligand efficiency values, and biological screening data, is provided in the Supplementary Information. The xREAL Space input data used for screening is not publicly available, and was provided exclusively to Chmepsace by Enamine. Docking models were based on publicly available crystal structures (PDB IDs: 5RLZ and 8GCY). Independent experimental validation data from CACHE Challenges #2 and #4 are publicly available at https://cache-challenge.org. Any additional data is available from the corresponding author on reasonable request.
Code availability
The V-SYNTHES2 algorithm used in this study is freely available at https://github.com/katritchlab/V-SYNTHES2 under the GNU GPL open-source license. Virtual screening was performed using ICM-Pro v3.9 (Molsoft LLC). Custom post-processing scripts for filtering, synthon-based clustering, and candidate selection used in this study are available from the corresponding author on reasonable request.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Vsevolod Katritch, Email: katritch@usc.edu.
Olga Tarkhanova, Email: o.tarkhanova@chem-space.com.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s44386-026-00067-0.
References
- 1.Luttens, A. et al. Ultralarge virtual screening identifies SARS-CoV-2 main protease inhibitors with broad-spectrum activity against coronaviruses. J. Am. Chem. Soc.144, 2905–2920 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Lyu, J. et al. Ultra-large library docking for discovering new chemotypes. Nature566, 224–229 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Carlsson, J. & Luttens, A. Structure-based virtual screening of vast chemical space as a starting point for drug discovery. Curr. Opin. Struct. Biol.87, 102829 (2024). [DOI] [PubMed] [Google Scholar]
- 4.Stein, R. M. et al. Virtual discovery of melatonin receptor ligands to modulate circadian rhythms. Nature579, 609–614 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Gorgulla, C. et al. An open-source drug discovery platform enables ultra-large virtual screens. Nature580, 663–668 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.McGann, M. OpenEye Scientific. GigaDocking—Structure Based Virtual Screening of Over 1 Billion Molecules Webinarhttps://www.eyesopen.com/webinars/giga-docking-structure-based-virtual-screening.
- 7.Sadybekov, A. V. & Katritch, V. Computational approaches streamlining drug discovery. Nature616, 673–685 (2023). [DOI] [PubMed] [Google Scholar]
- 8.Enamine REAL Space. https://enamine.net/compound-collections/real-compounds/real-space-navigator.
- 9.Kapeliukha, A. et al. Freedom Space 3.0: ML-assisted selection of synthetically accessible small molecules. J. Chem. Inf. Model.65, 10338–10347 (2025). [DOI] [PubMed] [Google Scholar]
- 10.Korn, M., Ehrt, C., Ruggiu, F., Gastreich, M. & Rarey, M. Navigating large chemical spaces in early-phase drug discovery. Curr. Opin. Struct. Biol.80, 102578 (2023). [DOI] [PubMed] [Google Scholar]
- 11.Grygorenko, O. O. et al. Generating multibillion chemical space of readily accessible screening compounds. iScience23, 101681 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Gloriam, D. E. Bigger is better in virtual drug screens. Nature566, 193–194 (2019). [DOI] [PubMed] [Google Scholar]
- 13.Liu, F. et al. Large library docking identifies positive allosteric modulators of the calcium-sensing receptor. Science385, eado1868 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Sadybekov, A. A. et al. Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature601, 452–459 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.CACHE Challenge—Critical Assessment of Computational Hit-Finding Experiments. https://cache-challenge.org/. [DOI] [PMC free article] [PubMed]
- 16.Li, F. et al. CACHE Challenge #1: targeting the WDR domain of LRRK2, a Parkinson’s disease associated protein. J. Chem. Inf. Model.64, 8521–8536 (2024). [DOI] [PubMed] [Google Scholar]
- 17.Herasymenko, O. et al. CACHE Challenge #2: targeting the RNA site of the SARS-CoV-2 helicase Nsp13. J. Chem. Inf. Model.65, 6884–6898 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Herasymenko, O. et al. CACHE Challenge #3: targeting the Nsp3 macrodomain of SARS-CoV-2. J. Chem. Inf. Model. 66, 1566–1581 (2026). [DOI] [PMC free article] [PubMed]
- 19.Lyu, J., Irwin, J. J. & Shoichet, B. K. Modeling the expansion of virtual screening libraries. Nat. Chem. Biol.19, 712–718 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Grebner, C. et al. Virtual screening in the cloud: how big is big enough? J. Chem. Inf. Model.60, 4274–4282 (2020). [DOI] [PubMed] [Google Scholar]
- 21.Liu, F. et al. The impact of library size and scale of testing on virtual screening. Nat. Chem. Biol.21, 1039–1045 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Santos-Martins, D. et al. Accelerating AutoDock4 with GPUs and gradient-based local search. J. Chem. Theory Comput.17, 1060–1073 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Molsoft. GigaScreen—Deep Learning Approach for Screening Ultra Large Databases. https://www.molsoft.com/GigaScreen.html.
- 24.Cecchini, D. et al. AI-enhanced adaptive virtual screening platform enabling exploration of 69 billion molecules discovers structurally validated FSP1 inhibitors. Preprint at 10.1101/2023.04.25.537981 (2026). [DOI]
- 25.Luttens, A. et al. Rapid traversal of vast chemical space using machine learning-guided docking screens. Nat. Comput Sci.5, 301–312 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Zhou, G. et al. An artificial intelligence accelerated virtual screening platform for drug discovery. Nat. Commun.15, 7761 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Zhou, G. OpenVS for an artificial intelligence accelerated virtual screening platform for drug discovery. https://github.com/gfzhou/OpenVS (2024). [DOI] [PMC free article] [PubMed]
- 28.Corso, G., Stärk, H., Jing, B., Barzilay, R. & Jaakkola, T. DiffDock: diffusion steps, twists, and turns for molecular docking. Preprint at 10.48550/ARXIV.2210.01776 (2022). [DOI]
- 29.Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. Preprint at 10.1101/2025.06.14.659707 (2025). [DOI]
- 31.Wang, L. et al. The present state and challenges of active learning in drug discovery. Drug Discov. Today29, 103985 (2024). [DOI] [PubMed] [Google Scholar]
- 32.Nazarova, A. L. et al. V-SYNTHES2 - the next generation tool for structure-based virtual screening of gigascale chemical spaces. npj Drug Discov.3, 21 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sadybekov, A. V. et al. Development of potent, selective cPLA2 inhibitors for targeting neuroinflammation in Alzheimer’s disease and other neurodegenerative disorders. npj Drug Discov.3, 2 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Souza-Silva, I. M. et al. Development of an automated, high-throughput assay to detect angiotensin AT2-receptor agonistic compounds by nitric oxide measurements in vitro. Peptides172, 171137 (2024). [DOI] [PubMed] [Google Scholar]
- 35.Beroza, P. et al. Chemical space docking enables large-scale structure-based virtual screening to discover ROCK1 kinase inhibitors. Nat. Commun.13, 6447 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Medel-Lacruz, B. et al. Synthon-based strategies exploiting molecular similarity and protein–ligand interactions for efficient screening of ultra-large chemical libraries. J. Chem. Inf. Model.65, 7569–7583 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Adedeji, A. O. et al. Mechanism of nucleic acid unwinding by SARS-CoV helicase. PLoS ONE7, e36521 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Newman, J. A. et al. Structure, mechanism and crystallographic fragment screening of the SARS-CoV-2 NSP13 helicase. Nat. Commun.12, 4848 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Katritch, V., Rueda, M. & Abagyan, R. Ligand-guided receptor optimization. in Homology Modeling Vol. 857 (eds Orry, A. J. W. & Abagyan, R.) 189–205 (Humana Press, 2011). [DOI] [PubMed]
- 40.Raush, E., Abagyan, R. & Totrov, M. Graph-convolutional neural net model of the statistical torsion profiles for small organic molecules. J. Chem. Inf. Model.62, 5896–5906 (2022). [DOI] [PubMed] [Google Scholar]
- 41.Corona, A. et al. Natural compounds inhibit SARS-CoV-2 nsp13 unwinding and ATPase enzyme activities. ACS Pharmacol. Transl. Sci.5, 226–239 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Bellmann, L., Penner, P. & Rarey, M. Topological Similarity Search in Large Combinatorial Fragment Spaces. J. Chem. Inf. Model.61, 238–251 (2021). [DOI] [PubMed]
- 43.Schmidt, R., Klein, R. & Rarey, M. Maximum Common Substructure Searching in Combinatorial Make-on-Demand Compound Spaces. J. Chem. Inf. Model.62, 2133–2150 (2022). [DOI] [PubMed]
- 44.Lessel, U., Wellenzohn, B., Lilienthal, M. & Claussen, H. Searching Fragment Spaces with Feature Trees. J. Chem. Inf. Model.49, 270–279 (2009). [DOI] [PubMed]
- 45.Chiang, Y. J. et al. Cbl-b regulates the CD28 dependence of T-cell activation. Nature403, 216–220 (2000). [DOI] [PubMed] [Google Scholar]
- 46.Zhou, L. et al. Rising star in immunotherapy: development and therapeutic potential of small-molecule inhibitors targeting casitas B cell lymphoma-b (Cbl-b). J. Med. Chem.67, 816–837 (2024). [DOI] [PubMed] [Google Scholar]
- 47.Kimani, S. W. et al. The co-crystal structure of Cbl-b and a small-molecule inhibitor reveals the mechanism of Cbl-b inhibition. Commun. Biol.6, 1272 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Patel, N. et al. Structure-based discovery of potent and selective melatonin receptor agonists. eLife9, e53779 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets generated and/or analyzed during the current study are not publicly available in full because the virtual screening was performed against the Enamine xREAL Space, a proprietary chemical space that cannot be redistributed. Key data necessary to interpret the findings, including hit compound structures, physicochemical properties, ligand efficiency values, and biological screening data, is provided in the Supplementary Information. The xREAL Space input data used for screening is not publicly available, and was provided exclusively to Chmepsace by Enamine. Docking models were based on publicly available crystal structures (PDB IDs: 5RLZ and 8GCY). Independent experimental validation data from CACHE Challenges #2 and #4 are publicly available at https://cache-challenge.org. Any additional data is available from the corresponding author on reasonable request.
The V-SYNTHES2 algorithm used in this study is freely available at https://github.com/katritchlab/V-SYNTHES2 under the GNU GPL open-source license. Virtual screening was performed using ICM-Pro v3.9 (Molsoft LLC). Custom post-processing scripts for filtering, synthon-based clustering, and candidate selection used in this study are available from the corresponding author on reasonable request.
