Significance
Generative modeling has enabled novel approaches to the design of functional molecules. However, its practical utility has been impeded by its tendency to propose molecules that are difficult or impossible to synthesize. To address this issue, we present SynFormer, a generative framework that ensures every generated molecule has a viable synthetic pathway, enabling the design of analogs and optimization of molecular properties while maintaining synthetic feasibility. By providing effective and controllable navigation within synthesizable chemical space, SynFormer can accelerate the discovery of small organic molecules across a range of fields, including drug development and materials science.
Keywords: generative AI, molecular design, synthetic accessibility, deep learning
Abstract
We introduce SynFormer, a generative modeling framework designed to efficiently explore and navigate synthesizable chemical space. Unlike traditional molecular generation approaches, we generate synthetic pathways for molecules to ensure that designs are synthetically tractable. By incorporating a scalable transformer architecture and a diffusion module for building block selection, SynFormer surpasses existing models in synthesizable molecular design. We demonstrate SynFormer’s effectiveness in two key applications: 1) local chemical space exploration, where the model generates synthesizable analogs of a query molecule, and 2) global chemical space exploration, where the model aims to identify optimal molecules according to a black-box property prediction oracle. Additionally, we demonstrate the scalability of our approach via the improvement in performance as more computational resources become available. With our code and trained models openly available, we hope that SynFormer will find use across applications in drug discovery and materials science.
The discovery of novel functional molecules is a central challenge in chemical science and engineering and is crucial for addressing key societal and technological challenges, including those related to healthcare (1–3), energy (4, 5), and sustainability (6–8). However, the process of discovery is often risky, complex, time-consuming, and resource-intensive (9–11). Recent advances in AI (12), particularly in generative modeling (13–15), have opened up new avenues for generative molecular design (16–23). The expressivity of deep learning models, coupled with automatic differentiation (24) and increasingly affordable computational resources, have made it possible to model the complex distribution of molecular structures directly. This capability allows for efficient exploration of broader virtual chemical space than traditional virtual screening approaches (25–27), with more fine-grained and controllable navigation compared to conventional combinatorial fragment-based methods (28, 29).
Despite the promise of generative design methods, their adoption has remained somewhat limited. One major challenge is that most generative models often produce synthetically intractable molecular structures when targeting specific design goals, such as optimizing property scores (30–33). When designed molecules cannot be synthesized and validated in the lab at a reasonable cost, their practical value is limited. In addition, the actionability of computer-aided molecular design is crucial for achieving rapid design cycles in early-stage drug discovery (34, 35) and for enabling closed-loop autonomous discovery (36–40). This issue, though long-standing and observed in earlier combinatorial algorithms, has become more pronounced with the advent of deep generative AI models.
There has been considerable interest in incorporating synthetic accessibility into molecular design methods (30, 33). The most straightforward strategy is to incorporate a heuristic synthetic accessibility score as a design criterion (30, 41). However, quantitatively estimating synthetic accessibility must account for factors such as regioselectivity, functional group compatibility, and building block availability, all of which contribute to a rugged structure-synthesizability landscape that makes the design of such scores an ongoing challenge. While notable efforts have been made to quantify synthetic accessibility (42–44), achieving reliable quantification through a simple heuristic remains a distant goal. Advances in the sample efficiency of molecular optimization (45, 46) have made it feasible to conduct explicit retrosynthesis analysis for each designed molecule as a means of evaluating synthetic accessibility (47). However, the computational overhead remains significant, making this approach impractical for most existing generative design models. Furthermore, unsynthesizable molecules provide no learning signal, and when a model continues to propose such molecules, the sparse feedback hinders effective learning.
A more ideal and effective approach to synthesizable molecular design, in our opinion, involves constraining the design process to focus exclusively on synthesizable molecules by designing synthetic pathways rather than simply designing structures. With the growing role of massive make-on-demand libraries—whose sizes exceed what can be fully enumerated (27, 48–53)—this class of methods has gained increasing interest, spanning both traditional Monte Carlo techniques (54–60) and more recent generative approaches (61–69). Despite this progress, current synthesis-centric methods still fall short in terms of controllability and efficiency in navigating the chemical space, and as a result have not achieved widespread adoption. These methods that rely on property scores to guide the design process often exhibit lower efficiency in optimizing expensive oracles compared to structure-centric methods (45), as they model the synthetic action sequence-property landscape, which is inherently more complex than the structure-property landscape. Some recent models offer a more controlled exploration of local chemical spaces by using a query molecule as input (61, 64, 65, 67). However, the low reconstruction rates for theoretically feasible molecules suggest a risk of partial collapse during decoding, suggesting that the practical chemical space accessible to these models is smaller than the theoretically accessible space. This behavior leads to certain regions of the chemical space becoming inaccessible, regardless of the input, ultimately hindering the overall design process.
Herein, we build upon our previous contributions (65, 67) to introduce SynFormer, a generative AI framework designed for efficient and controllable navigation within a synthesizable chemical space. Like earlier synthesizable design methods, SynFormer is a synthesis-centric approach, generating synthetic pathways by readily available building blocks through robust chemical transformations, ensuring synthetic traceability subject to the limitations of those transformation rules. SynFormer distinguishes itself from previous models with its compute-efficient and scalable transformer architecture (14), which empirically shows improvements in performance as the training dataset grows. To select suitable molecular building blocks from the large, discrete, and multimodal space of commercially available options, SynFormer incorporates a denoising diffusion model as a token head module (15). The framework is fully end-to-end differentiable, enabling effective training and optimization.
We demonstrate the utility and versatility of the SynFormer framework through two instantiations: 1) SynFormer-ED, an encoder–decoder model that generates synthetic pathways corresponding to a given input molecule for exact or approximate reconstruction of that input, and 2) SynFormer-D, a decoder-only model for generating synthetic pathways that is amenable to fine-tuning toward specific property goals. Both models are trained on a simulated chemical space derived from a curated set of 115 reaction templates and 223,244 commercially available building blocks, extending beyond Enamine’s REAL Space (50). We validate SynFormer’s capabilities by demonstrating success on a) reconstructing molecules within both the Enamine REAL and ChEMBL (70) chemical spaces; b) local synthesizable chemical space exploration given a query molecule; and c) global synthesizable chemical space exploration guided by black-box property prediction model. These results not only highlight SynFormer’s ability to navigate synthesizable chemical space using varied control strategies but also demonstrate its efficiency and practical applicability in real-world molecular design use cases. The scalability of the SynFormer framework with respect to both training data and model size suggests considerable potential for further performance enhancement. In conclusion, these findings underscore SynFormer’s flexibility and potential to impact molecular design across various domains, including drug development and materials science.
SynFormer: A Generative Framework for Synthesizable Molecular Design
SynFormer is a generative framework designed for modeling synthesizable chemical spaces, as schematically illustrated in Fig. 1. Although it is challenging to predict which molecules can be synthesized without experimental validation, the progress of make-on-demand libraries such as Enamine REAL Space (50), GalaXi (51), and eXplore (52) has demonstrated that molecules constructed by connecting purchasable molecular building blocks through a series of curated, reliable reactions have a high likelihood of being synthesizable. Therefore, within the scope of this work, we define synthesizable chemical space as encompassing all molecules that can be formed within the reaction network by linking purchasable building blocks through up to five steps of known chemical transformations, though in principle, there is no limit to the length of pathways that can be generated, and building blocks and reactions beyond our selected list can also be utilized.
Fig. 1.

Schematic illustration of the SynFormer framework and architecture. (A) The SynFormer-ED architecture is an encoder–decoder that takes a molecule as input and outputs a synthetic route to the same or an analogous molecule. (B) SynFormer-D is a decoder-only framework designed to generate synthetic routes. (C) Synthetic routes are tokenized using a postfix notation to make them amenable to autoregressive generation. Routes are constructed by applying 115 reactions to a set of 223,244 molecular building blocks, covering a synthesizable chemical space estimated as 1060 molecules. (D) During generation, a token generated by the transformer is first classified by token type. If the token represents a reaction or reactant (building block), it undergoes an additional classification process to select the appropriate reactions or a denoising diffusion process followed by a nearest neighbor search to select the appropriate building block(s).
To effectively model chemically relevant and practically useful spaces, we adapted the reaction set used to construct REAL Space, which primarily focuses on bi- and trimolecular reactions that couple multiple building blocks together. We further augment this set with additional common reactions, such as functional group interconversions, to better simulate a general organic synthesis endeavor, resulting in a total of 115 reaction templates. As the set of purchasable molecular building blocks, we use Enamine’s U.S. stock catalog to ensure high availability, which also approximates those used to generate REAL Space. Although the proportion of synthesizable molecules may vary due to the broader focus of the chemical space, our model theoretically covers a chemical space broader than then tens of billions currently represented in the Enamine REAL Space, making it suitable for extensive exploration during the molecular design stage. Both the choice of templates and the choice of building blocks can be straightforwardly modified prior to retraining and are not inherent to SynFormer’s approach.
Generating synthetic pathways requires careful consideration of their data structure. Bottom–up synthesis planning involves sequential decision-making with branches and merges, encompassing data modalities that span reaction templates and molecular building blocks, which represent small and very large discrete design spaces, respectively. To manage this complexity, we adopt the postfix notation to represent synthetic pathways linearly, as described in Luo et al. (67) (Fig. 1C). Specifically, we use four types of tokens to represent synthetic paths: a start token [START], an end token [END], reaction tokens [RXN], and building block tokens [BB]. Similar to the postfix notation of mathematical formulae, the reactions are placed after the reagents, allowing the sequence to be processed in a step-by-step manner without loss of generality in the synthetic pathways (e.g., the ability to accommodate any linear or convergent sequence). This linear notation also enables the use of autoregressive decoding via a transformer architecture (14), which is widely recognized as a scalable model backbone. A stack of standard transformer layers processes the sequence as it is decoded, taking the previous tokens and outputting the embedding for the next token. Each token embedding then passes through a multilayer perceptron (MLP) classification head to determine the token type. If the token is [END], the generation process terminates. If the token is [RXN] or [BB], the embedding is used to select which building blocks to add and which reaction type to perform.
The selection of reactions can be relatively straightforward using an additional classification head; however, the number of purchasable building blocks can easily exceed hundreds of thousands, millions, or perhaps billions, and will continue to grow over time. Instead of relying on a static classification network, our past work (65, 67) has involved generating Morgan fingerprints (71) and then retrieving the nearest building block from a list of candidates to enable generalization to unseen building blocks. To further enhance this building block selection process, we adopt a denoising diffusion probabilistic module (15) to predict the posterior distribution of molecular fingerprints conditioned on the token embedding. Specifically, we model the molecular fingerprints using an -dimensional joint Bernoulli distribution (where is the length of the fingerprint) and employ a multinomial diffusion framework (72) to handle the discrete nature of the data. Using the denoising diffusion module as the building block token head offers several advantages: While maintaining the end-to-end differentiable nature, it allows for scaling to longer fingerprints (2,048 in this work compared to 256 in previous studies), thereby avoiding bit collisions that can occur with shorter fingerprints and enabling generation at a finer resolution. Additionally, this approach more effectively handles multimodal distributions, where multiple “true” answers exist for a single decision–a particularly common scenario when modeling synthetic pathways.
Based on this generative modeling framework, we implemented two specific models: SynFormer-D, which straightforwardly implements a transformer decoder with the aforementioned reaction classification heads and building block diffusion heads (Fig. 1B), and SynFormer-ED, which adds a standard transformer encoder to condition generation on an input SMILES string (73) (Fig. 1A). To train the models, we uniformly sample synthetic pathways from the synthesizable chemical space simulated from the reaction templates and building block list as training data. Detailed descriptions of the model architectures, training protocols, and hyperparameter settings are provided in Materials and Methods and SI Appendix.
Results
Molecule Reconstruction and Chemical Space Coverage.
An important factor that affects a model’s performance in molecular design applications is its coverage of chemical space (74). Even though the theoretical synthesizable chemical space constructed from reaction templates and building blocks is vast, previous models have shown that a trained model can struggle to generate molecules that are known to be synthesizable (57, 64, 65, 67). This limitation can hinder the model’s effectiveness in molecular design if it cannot access the chemical space where the optimal design resides. To assess the coverage of chemical space, we tested the SynFormer-ED model on molecule reconstruction: how often SynFormer-ED can reconstruct a viable synthetic path given a large set of potential molecules of interest.
We first validated that the model successfully reconstructed 92.5% of 1,000 molecules enumerated from the reactions and building blocks used. We then evaluated its reconstruction performance on a set of molecules more relevant to real-world applications: a randomly selected subset of 1,000 molecules from the ChEMBL Database and Enamine’s REAL Diversity Set (Fig. 2 A and B). SynFormer-ED successfully recovers 66% of molecules from REAL, notably higher than that achieved by ChemProjector (67) and SynNet (30). ChEMBL, by comparison, is a curated database of bioactive molecules with drug-like properties, which includes molecules whose syntheses may require building blocks or reactions beyond SynFormer’s training set. This results in a reconstruction rate of 20%—lower than in REAL Space but still an improvement over previous models. A t-SNE visualization of the different distributions of molecules in REAL Space, ChEMBL, and the training data for our model is shown in Fig. 2C.
Fig. 2.
Model performance on molecular reconstruction. (A and B) Comparison of the reconstruction rate and average structural (Tanimoto) similarity between input and output molecules for SynFormer-ED, ChemProjector (67), and SynNet (65) on 1,000 randomly selected molecules from (A) REAL Diversity Set and (B) ChEMBL Database. Error bars represent the SD across three independent runs. (C) A t-SNE visualization of molecules in REAL Space, ChEMBL, and the training data for SynFormer, representing the theoretical synthesizable chemical space covered by the model. ChEMBL and REAL Space cover qualitatively different areas of chemical space. (D) Scaling of model performance is measured by the training loss (binary cross-entropy (BCE) of the molecular fingerprint (FP) prediction) as model size and training data size increase.
We further demonstrate that the model’s performance, as measured by the binary cross-entropy (BCE) of the molecular fingerprints (FPs) (71), improves with increasing model size and training data (Fig. 2D). Notably, we also observed that the performance of the largest model did not outperform the second-largest model when insufficient data was available. This finding suggests that estimating training performance based on early epochs may not be an effective strategy when the model size surpasses a certain threshold (75). These results emphasize the scalability of the architecture (76) and indicate that further performance improvements can be achieved by increasing both the amount of data and computational resources. They also underscore the importance of scaling model size and training data in tandem to fully realize performance gains. As SynFormer is trained using simulated pathways according to its building blocks and reaction rules, there is virtually no limit to the size of training data available.
Local Chemical Space Exploration with SynFormer.
In this section, we present the application of SynFormer-ED in exploring the local chemical space around a given query molecule. We demonstrate two use cases of our model in molecular design: 1) generating synthesizable analogs for unsynthesizable designs and 2) hit expansion within synthesizable chemical space.
Synthesizable analog generation.
In molecule reconstruction experiments, we observed that SynFormer-ED is able to generate structurally similar outputs even when provided with unsynthesizable (according to its building blocks and reactions) molecules. This capability allows SynFormer-ED to be applied to generate synthesizable structural analogs of unsynthesizable designs, as illustrated in Fig. 3A. The implicit goal is to preserve as much of the overall structure and key pharmacophores as possible, effectively transforming nonfeasible molecules into synthesizable compounds, akin to “projecting” the input molecule onto the synthesizable space. It is important to note that this “molecular projection” is not a rigorous mathematical concept but rather a conceptual description of the model’s behavior.
Fig. 3.
Application of SynFormer in projecting unsynthesizable design into synthesizable chemical space. (A) Schematic illustration of SynFormer-ED generating synthesizable structural analogs for unsynthesizable molecules. (B) Normalized distribution of SA scores for the originally designed molecules and their corresponding synthesizable analogs. Note that the distributions are normalized to a peak value of 1 for a clearer comparison. (C) Scatter plot comparing objective scores of originally designed molecules versus their generated analogs. Points are colored by the structural similarity between them, showing that structurally similar analogs tend to possess close properties. (D) Examples of originally designed unsynthesizable molecules and their SynFormer-generated synthesizable analogs. Objective scores are shown beneath each molecule, demonstrating comparable activities for the generated analogs. The modified fragments or atoms are highlighted in light blue. (E) The workflow shows SynFormer-ED generating synthesizable analogs for ligands generated by structure-based drug design. (F) Scatter plot comparing the Vina docking scores of originally designed ligands and their generated analogs. Points are colored based on the structural (Tanimoto) similarity between the input and output, showing strong agreement in general and an ability to generate analogs with comparable scores. (G) The original design and generated analog for Estrogen receptor alpha, with their SA Score and Vina score below. The modified fragments or atoms are highlighted in light blue.
To evaluate the impact of molecular projection, we first followed the experimental setup of Luo et al. (67) and used SynFormer-ED to generate synthesizable analogs for molecules identified as unsynthesizable by ASKCOS (77) in ref. 30. The original molecules were optimized using de novo design methods targeting ten multiobjective scores, each ranging from 0 to 1, where higher values are preferable. These scores encompass structural features, physicochemical properties, and metrics related to similarity or dissimilarity to known drugs, as well as the presence of specific substructures (78). After the generation of analogs, we evaluated them using the same objective scores that the original designs were optimized for.
As shown in Fig. 3B, the synthesizable projection effectively eliminated the peak in synthetic accessibility (SA) scores (42) at around 6 to 7 observed in the original designs, shifting the distribution toward more easily synthesizable molecules. Since SynFormer explicitly generates synthetic routes, all generated molecules are inherently synthesizable. The SA Scores are included as a quick, quantitative measure of structural complexity but are neither a necessary nor a sufficient criterion for assessing improved synthesizability. In addition, a significant fraction of the generated analogs exhibited objective scores comparable to those of the original molecules (Fig. 3C and SI Appendix, Fig. S3). However, we also observed that a notable fraction of molecules exhibited a significant drop in design objectives. Analyzing the relationship between changes in objective scores, the similarity between query and output molecules, and the SA score of the query molecules revealed that the observed drop in design objectives can be attributed to a failure in decoding query molecules into meaningfully similar analogs (SI Appendix, Figs. S2 and S3). This failure primarily occurs when the query molecules are too structurally complex and distant from synthesizable chemical space. These findings highlight the importance of starting with query molecules that are closer to synthesizable space to ensure the molecular projection generates meaningful analogs. A more detailed discussion, along with a fragment analysis of molecules that can and cannot be decoded into similar analogs, is provided in SI Appendix, section 2E and Figs. S4–S7.
We show two examples in Fig. 3D. Starting with high-scoring but seemingly unsynthesizable molecules, SynFormer-ED generated analogs that preserved key structural features while correcting the unsynthesizable fragments. The generated analogs retained objective scores comparable to those of the original molecules and are inherently synthesizable (SI Appendix, Fig. S10), demonstrating SynFormer-ED’s ability to transform an unsynthesizable molecule into a synthesizable analog while maintaining structural integrity.
To further validate the utility of SynFormer-ED in realistic drug discovery scenarios, we explored its application in structure-based drug design (SBDD) (79–81). While numerous SBDD algorithms can generate ligands with promising predicted binding affinities (e.g., via 3D pocket-conditioned generation), these methods are often criticized for overlooking synthetic accessibility (82, 83). SynFormer-ED addresses this gap by projecting the outputs of such models into synthesizable chemical space (Fig. 3E). We demonstrate this approach using Pocket2Mol (84) as an exemplary SBDD generative model.
Targeting 15 protein targets from the LIT-PCBA dataset (85), we used SynFormer-ED to generate synthesizable analogs of molecules initially designed by Pocket2Mol. SynFormer-ED successfully produced analogs that preserve similar Vina docking scores (86) while ensuring that each design has a synthetic pathway (Fig. 3F). Two examples are shown in Fig. 3G: In the first example, the original design had a Vina score of 11.2 kcal/mol but a high SA score of 5.84, indicating poor synthesizability. SynFormer-ED generated a structurally similar analog with an improved SA score of 2.99 while exhibiting a Vina score of 9.8 kcal/mol. In the second example, although the original design was likely synthesizable with access to the appropriate building blocks, SynFormer-ED further reduced the SA score from 2.44 to 1.99 with minor improvements in Vina score. These results demonstrate SynFormer-ED’s potential to complement existing SBDD algorithms by generating more practical and synthesizable drug candidates, suitable for experimental evaluation.
Hit expansion within synthesizable chemical space.
Building on SynFormer-ED’s ability to explore local synthesizable chemical space around query molecules, we applied it to hit expansion within synthesizable chemical space. By allowing the model to select suboptimal intermediate choices of building blocks or reactions via a beam search, SynFormer-ED can generate sets of close analogs around known hits, as illustrated in Fig. 4A. This approach aims to produce additional molecules that preserve the core motifs and overall structural integrity, potentially enhancing design objectives and expanding the pool of candidates for downstream processes, such as hit-to-lead optimization.
Fig. 4.

Application of SynFormer in hit expansion. (A) Schematic illustrating SynFormer-ED expanding a known hit compound into structurally similar, synthesizable analogs. (B) Normalized distribution of predicted JNK3 inhibition scores for screening ZINC250k, nearest neighbors in Enamine REAL of the hits, and SynFormer-generated analogs of the hits, highlighting the enrichment of high-scoring ligands among SynFormer analogs. Note that the distributions are normalized to a peak value of 1 for a clearer comparison. (C) Examples of hits from screening ZINC250k and corresponding best-scored generated analogs, alongside the best-scored molecules identified in the nearest neighbor search within Enamine REAL, demonstrate SynFormer-ED’s ability to generate high scoring, synthesizable compounds. The motifs retained from the hits are colored in red. (D and E) Expanding experimentally validated ligands for the PKM2 target (PDB: 3ME3) (D) and the KAT2A target (PDB: 5MLJ) (E). Three representative analogs are shown for each target, along with their synthetic pathways. The structural motifs retained from the hits are colored in red.
Following the setup in ref. 87, we screened ZINC250k (88), a subset of ZINC database (89), using a predictive model for inhibition activity against c-Jun NH2-terminal kinase 3 (JNK3) (90–92). The top 10 scoring molecules, with scores ranging from 0.49 to 0.68 (the higher, the better), were selected as hits and input into SynFormer-ED to generate analogs, which were then evaluated with the JNK3 inhibition predictor. As a baseline, a nearest neighbor search within Enamine REAL (approximately 7 billion molecules) was performed for the hits, retrieving approximately 100 of the most similar molecules for each hit and evaluating them for JNK3 inhibition. While both SynFormer-ED and the nearest neighbor search identified additional molecules with scores higher than those from a simple screening of ZINC250k, SynFormer-generated analogs exhibited a larger enrichment in the high-scoring region (Fig. 4B). We present two sets of examples in Fig. 4C. Though both methods produced analogs retaining the overall structure with higher predicted inhibition than the original hits, the SynFormer-ED analogs consistently achieved higher scores compared to those from the nearest neighbor search.
To further validate SynFormer-ED’s capability in a more realistic hit expansion scenario, we used experimentally verified ligands for Human Pyruvate Kinase M2 (PKM2, PDB ID: 3ME3) (93) and Lysine Acetyltransferase 2A (KAT2A, PDB ID: 5MLJ) (94) as query molecules and evaluated their Vina docking scores. SynFormer-ED was used to expand the reported ligands, generating 191 analogs for PKM2, with 179 of them with Tanimoto similarities greater than 0.5 and Vina score changes lower than 1 kcal/mol. For KAT2A, SynFormer-ED generated 200 analogs, 111 of which met the criteria of Tanimoto similarity above 0.5 and Vina score changes lower than 1 kcal/mol. Fig. 4 D and E present three generated analogs for ligands targeting PKM2 and KAT2A, respectively, showing competitive Vina scores while preserving the overall structure, along with their synthetic pathways (see SI Appendix, Figs. S17 and S18 for more examples). Notably, in the case of KAT2A, the synthetic pathways of the generated analogs are diverse despite the structural similarity of the final analogs, highlighting that the SynFormer-based approach is not limited by the synthetic pathways of the input hits. This flexibility enables generation of analogs with distinct synthetic routes, which traditional path-based enumeration strategies (87, 95) cannot achieve. Combined with these results, SynFormer-ED demonstrates significant potential in hit expansion, enabling the generation of practical, synthesizable analogs suitable for further drug development stages.
Global Chemical Space Navigation with SynFormer.
In this section, we demonstrate the application of SynFormer to global chemical space navigation, enabling the optimization of molecular properties treated as a black-box function across broad chemical spaces (45). This process, also known as de novo molecular optimization, is key to exploring novel molecular designs. We first show that SynFormer-D can be fine-tuned using reinforcement learning to generate high-scoring molecules (Fig. 5A). Furthermore, we demonstrate how SynFormer-ED can be integrated as a mutation step within an evolutionary algorithm (Fig. 5C), achieving state-of-the-art sample efficiency while constraining the design space to synthesizable chemical space.
Fig. 5.

Application of SynFormer in global chemical space exploration. (A) Illustration of fine-tuning SynFormer-D with reinforcement learning. (B) Performance comparison of SynFormer fine-tuning with reinforcement learning against other popular methods, showing the average top-10 molecule scores versus the number of oracle calls (45). The plots represent the mean performance curves of 5 independent runs, with the shaded region for SynFormer-D indicating the range across these runs. (C) Illustration of a genetic algorithm with SynFormer-ED used for mutation steps. (D) AUC Top-10 performance comparison across different molecular design methods [GraphGA-SF, GraphGA (29), AugMem (96), SynNet (65), DoG-Gen (64), and SyntheMol (59)] for four tasks from GuacaMol (78). (E) Distribution of SA Scores (42) for the top 25 molecules at various optimization steps, with colors representing objective scores, demonstrating how SynFormer effectively constrains its design space to synthesizable space exclusively. (F) The best molecules generated by GraphGA (Top) and GraphGA-SF (Bottom) show that GraphGA-SF identifies a more synthetically tractable candidate with a reduced SA score, albeit with a minor sacrifice in the objective score.
SynFormer-D with reinforcement learning.
We fine-tuned SynFormer-D using reinforcement learning (RL) to guide its generation toward high-scoring molecules. Specifically, we adopted a variant of the REINFORCE algorithm (97), which iteratively generates a batch of molecules, evaluates their properties, and fine-tunes the model parameters with the goal of maximizing the desired properties (detailed methods can be found in Materials and Methods). In our experiments, we optimized predicted binding affinity against the dopamine receptor D2 (DRD2) (98) as an exemplary oracle function. As shown in Fig. 5B, our method combining SynFormer-D with RL successfully biases generation toward high-scoring molecules and outperforms several popular optimization methods (29, 99). Although its sample efficiency is lower than the most efficient algorithms like GraphGA (29) and REINVENT (19), these results confirm the feasibility and effectiveness of fine-tuning SynFormer-D for de novo molecular optimization. Given that much of the algorithm’s design space of RL is yet to be explored, the potential for additional algorithmic improvements in RL-based fine-tuning remains substantial.
SynFormer-ED as a module in a genetic algorithm.
Last, we demonstrate that SynFormer-ED can be incorporated as a module to confine unconstrained design methods onto synthesizable space. Specifically, we present a genetic algorithm framework incorporating SynFormer-ED projection as a mutation step. After the standard processes of crossover and mutation to diversify the candidate pool, SynFormer-ED projects all candidates into synthesizable chemical space. This step not only corrects unsynthesizable fragments and linkages, ensuring synthetic accessibility but also introduces additional modifications to candidates compared to the original candidate pool. Generated molecules are then scored and ranked, with the top-performing ones selected for further iterations (Fig. 5C). For crossover and mutation rules, we follow the GraphGA algorithm (29), a well-established molecular optimization method that regularly performs well in molecular optimization benchmarks (45, 100) and real-world applications (41).
We evaluated the performance of this genetic algorithm, denoted as GraphGA-SF, across four tasks from the GuacaMol benchmark (78), including four Multi-Property Optimization (MPO) tasks: Sitagliptin MPO, Scaffold Hop, Perindopril MPO, and Ranolazine MPO. As shown in Fig. 5D, GraphGA-SF demonstrated optimization efficiency comparable to the original GraphGA and the state-of-the-art RL method, Augmented Memory (denoted as AugMem) (96), while significantly outperforming previous methods for synthesizable molecular optimization. Notably, in the Sitagliptin MPO task, where the primary objective score often favored unsynthesizable molecules, conventional methods without synthesizability constraints tended to generate exclusively unsynthesizable compounds (See SA Score in Fig. 5E and structure in Fig. 5F). On the other hand, synthesis-centric design methods, such as SynNet (65), DoG-Gen (64), and SyntheMol (59), struggled to find meaningful optimization signals due to the sparse signal and their limited coverage of chemical space. In contrast, GraphGA-SF matched the performance of GraphGA while ensuring each design has a synthetic pathway like other synthesis-centric methods. We highlight the best designs from each genetic algorithm in Fig. 5F, showing that the molecules generated by GraphGA-SF were significantly more synthesizable, with only a modest sacrifice in the main objective score.
Notably, applying synthesizable projection after molecular design can lead to structures that are too dissimilar from the original designs and exhibit lower desired properties. This occurs because synthesizable projection is ineffective for molecules that deviate too far from the synthesizable chemical space. To illustrate this, we applied synthesizable projection to the top-ranked design from GraphGA in the Sitagliptin MPO task. As shown in SI Appendix, Fig. S23, the projection resulted in a molecule with a Tanimoto similarity of only 0.186 to the original design, leading to a significant loss in the objective score. This outcome underscores the importance of incorporating SynFormer as a modular component to refine intermediate steps in the design process, preventing designs from drifting too far from the manifold of synthesizable structures.
Overall, our results demonstrate that by combining the optimization capability of GraphGA with SynFormer-ED’s ability to restrict the design space to synthesizable compounds, GraphGA-SF can effectively and efficiently optimize molecular properties while maintaining synthetic feasibility. Furthermore, SynFormer-ED’s integration as a mutation operator exemplifies its potential as a generalizable component in molecular design algorithms, allowing these frameworks to confine their search to synthesizable chemical spaces. This adaptability enables SynFormer-ED to leverage ongoing advances in de novo molecular design algorithms, offering a versatile and robust solution to the problem of synthesizable molecular design.
Discussion
In this paper, we introduced SynFormer, a generative modeling framework designed to efficiently explore and navigate synthesizable chemical space. We curated a set of reaction rules that allow the model to construct a synthesizable chemical space beyond Enamine REAL Space, ensuring that the designed molecules are feasible for synthesis subject to the reliability of such rules. By integrating transformer architectures with a diffusion module, SynFormer enhances the ability to make beneficial choices of building blocks and reactions during autoregressive bottom–up pathway generation, surpassing previous models for synthesizable molecular design. We demonstrated two key applications of SynFormer: local chemical space exploration and global optimization within synthesizable chemical space, both of which yielded successful results. Additionally, we showed that SynFormer is a scalable architecture, with the potential for further performance improvements with increased resources. While our demonstration focuses on applications in drug discovery based on topological structure, the local and global chemical space exploration capability is task-agnostic, and thus, the models are readily applicable to other domains of small organic molecule design, with or without the addition of conformation generation as a postprocessing step (4, 7).
There are several areas where we expect further improvements can be made. First, the reconstruction rate of SynFormer-ED is not perfect, suggesting that certain regions of synthesizable chemical space may remain inaccessible. Additionally, RL-based fine-tuning in SynFormer-D currently requires a relatively large number of oracle calls to achieve competitive results. Another limitation lies in the quality and coverage of reaction templates and building blocks. These templates do not account for stereochemistry, meaning that chiral centers can only be inherited from building blocks, not installed during synthesis. Furthermore, the Enamine building blocks we choose are designed for combinatorial library construction and do not comprehensively cover commercially available materials. These constraints may limit the model’s applicability to somewhat flat, linearly constructed sp2-rich structures and restrict its use for more complex molecules like natural products (101). Expanding the SynFormer framework to include more sophisticated reactions and a broader range of building blocks is possible but will require further validation.
Materials and Methods
Synthesizable Chemical Space Construction.
A synthesizable chemical space is defined as a set of molecules that can be accessed from purchasable starting materials (building blocks) through a series of reliable chemical transformations.
Reaction templates.
Reaction rules are encoded as strings using the SMARTS grammar (102), known as reaction templates. We created a set of 86 reaction templates, which directly implement the 169 reactions used in the January 2022 version of the REAL Space. This set includes both bimolecular and trimolecular reactions, with deprotection reactions treated as separate entities. Reactions that can be represented by the same transformation rule are encoded using a single SMARTS string. In addition to linking building blocks, we also considered more general chemical synthesis processes. We manually selected 29 additional reaction templates from Hartenfeller et al.’s (103) and Button et al.’s (104) template collection, which were not covered by the initial 86 reactions. This expanded set of reactions complements the REAL reactions by providing a closer approximation to general laboratory organic synthesis.
Building blocks.
The building block list comprises reagents or intermediates that are commercially available and can be directly purchased from the market. For this study, we utilized Enamine’s U.S. in-stock collection of building blocks, which was available for delivery as of October 1st, 2023. The building block list contains a total of 223,244 molecules. This selection was made to ensure the accessibility and immediate availability of the building blocks for synthesis.
Representing synthetic pathways using postfix notation.
Each synthesis pathway is encoded as a sequence of tokens, where each token represents either a reaction step or a building block, represented using the postfix notation of synthesis (67). In this format, the operators (chemical reactions) follow their operands (molecular building blocks), providing a structured and concise way to represent reaction sequences.
SynFormer.
SynFormer is based on a transformer architecture with a diffusion module to select molecular building blocks. We designed two variants of the model to handle distinct molecular design tasks: SynFormer-D, a decoder-only model designed for generating new molecules or optimizing molecules based on feedback from a black-box objective function; SynFormer-ED, an encoder–decoder model that generates synthetic pathways based on a given reference molecule, enabling local chemical space exploration.
Transformer decoder.
SynFormer adopts the transformer decoder architecture as its backbone (14). The decoder generates postfix notations of synthesis in an autoregressive manner, which is, predicting the next token according to previous tokens. We assign embedding vectors to each reaction template and the [START] token, and convert building blocks fingerprint into continuous-valued vectors with a MLP in the beginning. To indicate the position of each token within the sequence, we add positional encodings (14) to the embedding vectors.
Following the transformer decoder, we use a classifier based on an MLP to predict the next token’s type. If it is predicted as an [END] token, the decoding process will be terminated. If it is predicted as a reaction, another classifier is used to predict the probability of each reaction template being the next token. When the next token is predicted as a building block, a diffusion module is used to generate molecular fingerprints which are subsequently used to retrieve building blocks from the database.
Transformer encoder.
The transformer decoder can be conditioned by a transformer encoder, i.e., given reference molecule represented by a SMILES string. The transformer encoder produces vector representations for each token in the SMILES string, which the decoder can attend to via the cross attention mechanism (14).
Diffusion module for building block fingerprints.
We build a denoising diffusion model to learn the distribution of molecular fingerprints, represented as -dimensional binary vectors, where denotes the number of bits in the fingerprint. This distribution is modeled as a joint Bernoulli binary distribution of dimensions.
During training, the forward diffusion process perturbs the fingerprint vector by randomly flipping each bit according to the following noise distribution:
| [1] |
where represents the -th bit of the ground truth fingerprint vector, and represents the perturbed bit. Here, is a value that controls the noise level, which we define using a formulation similar to multinomial diffusion (72). starts at 1 and monotonically decreases to 0. When , the distribution becomes the uniform Bernoulli distribution, where the noise level reaches the maximum. The denoiser network learns the bitwise Bernoulli distribution of ground truth fingerprints conditioned on the perturb fingerprint and the embeddings from the transformer decoder. It is trained by minimizing the binary cross entropy loss that measures the discrepancy between the distribution and the ground truth fingerprint.
The inference process is known as the reverse diffusion process, where we start with a random bit vector drawn from the uniform Bernoulli distribution, and denoise it iteratively with the denoiser network. Finally, the denoised bit vector is used as a fingerprint to retrieve building block molecules from the database.
Reinforcement learning.
For global optimization tasks, SynFormer-D was fine-tuned using a variant of the REINVENT algorithm (19) based on REINFORCE (97), a type of RL. Our RL setup was designed to minimize the following loss function, which incentivizes the model to bring the pseudo-log-likelihood (PLL) of generating a synthetic pathway closer to the objective score of the product molecules:
| [2] |
where represents the trainable network parameters, is the synthetic pathway, is the product molecule, is the objective score of molecule , typically scaled between 0 and 1 (the higher, the better), is a constant scaling factor, is the pseudo-log-likelihood of generating the pathway leading to molecule . The pseudo-log-likelihood, , is computed as the sum of three terms: the log-likelihood of selecting each token (reaction step or building block), the log-likelihood of selecting each reaction type, and the empirical log-likelihood value of the building block fingerprints:
| [3] |
where represents each component of a synthetic path (token, reaction). The empirical likelihood of the building block fingerprints is approximated by
| [4] |
where is the denoised fingerprint vector, is the ground truth fingerprint.
Genetic algorithm with SynFormer-ED as a mutation operator.
We integrated SynFormer-ED as a mutation step into an evolutionary framework for global chemical space exploration and optimization. The genetic algorithm follows a standard workflow: Starting with a pool of candidates randomly sampled from the ZINC database, pairs of molecules are selected to undergo crossover, producing an offspring pool. Mutation is then applied to each offspring with a specific probability, followed by an additional mutation step where SynFormer-ED projects all offspring into synthesizable space. This step ensures that unsynthesizable fragments are corrected and adds additional diversity to the candidate pool. The offspring pool will be evaluated, and the highest-scored ones will form the next generation for the next iteration. The crossover and mutation rules were adopted from the GraphGA (29) algorithm; specifically, crossover involves graph matching and exchanging molecular halves, while mutation employs a set of hand-coded rules including both atom- and fragment-level modifications. We adopted the hyperparameters from ref. 100.
Evaluation Details.
This section describes the precise settings used in the evaluation of SynFormer-ED and SynFormer-D.
Reconstruction and local chemical space exploration.
SynFormer-ED was employed for reconstruction, analog generation, and hit expansion. The workflows for these tasks are similar: SynFormer-ED encodes the input molecule and decodes it into synthetic pathways, using a search width of 24 and an exhaustiveness of 64 during decoding. For reconstruction and analog generation, Tanimoto (Jaccard) similarity (105) to the input molecule based on Morgan fingerprints (106) was evaluated for each generated molecule, with the most similar molecule selected as the output. In hit expansion, generated molecules were evaluated based on both the design objective score and Tanimoto similarity. We tested the reconstruction performance on a random sample of 1,000 molecules from Enamine’s REAL Space (50) and the ChEMBL database (107). Molecules were sampled from the REAL Diversity Set, which includes 48.2 million molecules and maximizes diversity to represent the broader REAL Space. These molecules comply with the Rule of 5 (Ro5) (108) and Veber criteria (109), making them suitable for drug-like properties. ChEMBL samples were taken from version 29 of the database, released in 2021.
The SA Score (42), a heuristic measure evaluating structural complexity based on fragment frequency in PubChem (110), was used to assess the synthesizability of the generated molecules. For synthesizable analog generation, 10 objective scores from the GuacaMol benchmark (78) were used, which involve optimizing various physicochemical properties, similarities, or dissimilarities to known drugs, and specific structural motifs. In hit expansion, we also used a predictive model (92) to assess inhibition against JNK3, a member of the mitogen-activated protein kinase family. This model is a random forest classifier utilizing ECFP6 fingerprints trained on the ExCAPE-DB dataset (91). AutoDock Vina (86) was employed to evaluate binding affinities against protein targets in both structure-based drug design and hit expansion tasks. Protein structures and pocket information were sourced from the LIT-PCBA dataset (85), and an exhaustiveness setting of 16 was used for most calculations.
For baseline comparisons in analog finding, we conducted a nearest neighbor search within Enamine REAL. This was done by accessing the Enamine website (https://new.enaminestore.com/draw-search) and downloading approximately 100 nearest neighbors by adjusting the similarity threshold value.
Global chemical space exploration.
For both SF-RL and GraphGA-SF, we employed these models to solve molecular optimization problems without predefined starting points (45). In each experiment, we assumed the presence of an oracle, i.e., a function that evaluates the desired property of a molecule, providing the ground truth value. Each experiment was limited to 10,000 oracle calls, meaning up to 10,000 molecules could be evaluated. During optimization, all evaluated molecules were recorded, and the primary performance metric was the area under the curve (AUC) of the top- average property value versus the number of oracle calls (AUC top-). The reported AUC values were min-max scaled to the range [0, 1].
We evaluated the global chemical space optimization capabilities of SF-RL and GraphGA-SF on optimizing DRD2 inhibition (111) and four multiobjective properties from the GuacaMol benchmark (78): Sitagliptin MPO, Scaffold Hop, Perindopril MPO, and Ranolazine MPO. DRD2 inhibition was assessed using a classifier trained on the ExCAPE-DB dataset (91), employing a support vector machine with a Gaussian kernel and ECFP6 fingerprints to distinguish active from inactive compounds for the DRD2 (111). The multiobjective tasks from the GuacaMol benchmark are described above (78). All oracle functions were accessed through the Therapeutic Data Commons (112, 113).
Baseline methods in this study include: GraphGA (29) is a popular heuristic algorithm inspired by natural evolutionary processes, featuring crossover rules derived from graph matching and mutations applied at both the atom- and fragment-level. Graph MCTS (29) uses Monte Carlo Tree Search to explore molecular structures by locally searching each branch of the current state (molecule or partial molecule) and selecting the most promising candidates based on property scores for subsequent iterations. REINVENT (19) utilizes a policy-based RL approach where agents take actions in an environment to maximize cumulative rewards, tuning recurrent neural networks (RNNs) to generate SMILES strings. Augmented Memory (AugMem) (96) integrates SMILES-based reinforcement learning with augmented training through experience replay and selective memory purge to prevent model collapse. This approach achieves state-of-the-art sample efficiency, demonstrating the highest reported performance in the PMO benchmark. GFlowNet (99) is a generative AI model that treats the generative process as a flow network and trains it with a temporal difference-like loss function, aligning the generation process with the target property distribution by matching the property of interest to the flow volume. DoG-Gen (64) models synthetic pathways as Directed Acyclic Graphs (DAGs) and uses an RNN-based generator, optimizing via an iterative learning method that incorporates high-scoring molecules into the training data for successive fine-tuning of the generative model. SynNet (65) is a synthesis-based method that applies a genetic algorithm to molecular fingerprints, decoding them into synthetic pathways using trained neural networks. SyntheMol (59) is a synthesis-focused method leveraging the Monte Carlo tree search algorithm to explore chemical space. The implementations of these methods were adapted from their original publications with minor modifications, and hyperparameters were set according to the specifications in ref. 45, except AugMem which is from the original publication.
Supplementary Material
Appendix 01 (PDF)
Acknowledgments
This research was supported by the Office of Naval Research under grant number N00014-21-1-2195 and the AI2050 program at Schmidt Futures under grant number G-22-64475. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Office of Naval Research. W.G. received additional funding from the Google Ph.D. fellowship. We thank Kaiming He and Yurii Moroz for their helpful discussions and comments on the manuscript.
Author contributions
W.G., S.L., and C.W.C. designed research; W.G. and S.L. performed research; W.G., S.L., and C.W.C. analyzed data; and W.G., S.L., and C.W.C. wrote the paper.
Competing interests
The authors declare no competing interest.
Footnotes
This article is a PNAS Direct Submission.
Data, Materials, and Software Availability
The code and reaction templates are available at: https://github.com/wenhao-gao/synformer (114), while the model weights can be accessed on https://huggingface.co/whgao/synformer (115).
Supporting Information
References
- 1.Walters W. P., Green J., Weiss J. R., Murcko M. A., What do medicinal chemists actually make? A 50-year retrospective J. Med. Chem. 54, 6405–6416 (2011). [DOI] [PubMed] [Google Scholar]
- 2.Wu P., Nielsen T. E., Clausen M. H., Small-molecule kinase inhibitors: An analysis of FDA-approved drugs. Drug Discov. Today 21, 5–10 (2016). [DOI] [PubMed] [Google Scholar]
- 3.Lyu J., et al. , Ultra-large library docking for discovering new chemotypes. Nature 566, 224–229 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Yu Z., et al. , Molecular design for electrolyte solvents enabling energy-dense and long-cycling lithium metal batteries. Nat. Energy 5, 526–533 (2020). [Google Scholar]
- 5.Hachmann J., et al. , The harvard clean energy project: Large-scale computational screening and design of organic photovoltaics on the world community grid. J. Phys. Chem. Lett. 2, 2241–2251 (2011). [Google Scholar]
- 6.Rai P., Mehrotra S., Priya S., Gnansounou E., Sharma S. K., Recent advances in the sustainable design and applications of biodegradable polymers. Bioresour. Technol. 325, 124739 (2021). [DOI] [PubMed] [Google Scholar]
- 7.Diederichsen K. M., et al. , Electrochemical methods for carbon dioxide separations. Nat. Rev. Methods Primers 2, 68 (2022). [Google Scholar]
- 8.Peng J., et al. , Human-and machine-centred designs of molecules and materials for sustainability and decarbonization. Nat. Rev. Mater. 7, 991–1009 (2022). [Google Scholar]
- 9.Pyzer-Knapp E. O., Suh C., Gómez-Bombarelli R., Aguilera-Iparraguirre J., Aspuru-Guzik A., What is high-throughput virtual screening? A perspective from organic materials discovery Annu. Rev. Mater. Res. 45, 195–216 (2015). [Google Scholar]
- 10.Smietana K., Siatkowski M., Møller M., Trends in clinical success rates. Nat. Rev. Drug Discov. 15, 379–380 (2016). [DOI] [PubMed] [Google Scholar]
- 11.Pushpakom S., et al. , Drug repurposing: Progress, challenges and recommendations. Nat. Rev. Drug Discov. 18, 41–58 (2019). [DOI] [PubMed] [Google Scholar]
- 12.LeCun Y., Bengio Y., Hinton G., Deep learning. Nature 521, 436–444 (2015). [DOI] [PubMed] [Google Scholar]
- 13.Wang Y., Blei D., Variational bayes under model misspecification. Adv. Neural Inf. Process. Syst. 32, 10859 (2019). [Google Scholar]
- 14.A. Vaswani et al. , Attention is all you need. arXiv [Preprints] (2017). 10.48550/arXiv.1706.03762 (Accessed 20 September 2024). [DOI]
- 15.Ho J., Jain A., Abbeel P., Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 33, 6840–6851 (2020). [Google Scholar]
- 16.Sanchez-Lengeling B., Aspuru-Guzik A., Inverse molecular design using machine learning: Generative models for matter engineering. Science 361, 360–365 (2018). [DOI] [PubMed] [Google Scholar]
- 17.Gómez-Bombarelli R., et al. , Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent. Sci. 4, 268–276 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.W. Jin, R. Barzilay, T. Jaakkola, “Junction tree variational autoencoder for molecular graph generation” in International conference on machine learning, J. Dy, A. Krause, Eds. (PMLR, 2018), pp. 2323–2332.
- 19.Blaschke T., et al. , Reinvent 2.0: An AI tool for de novo drug design. J. Chem. Inf. Model. 60, 5918–5922 (2020). [DOI] [PubMed] [Google Scholar]
- 20.T. Fu et al. , Differentiable scaffolding tree for molecular optimization. arXiv [Preprints] (2021). 10.48550/arXiv.2109.10469 (Accessed 20 September 2024). [DOI]
- 21.Meyers J., Fabian B., Brown N., De novo molecular design and generative models. Drug Discov. Today 26, 2707–2715 (2021). [DOI] [PubMed] [Google Scholar]
- 22.Anstine D. M., Isayev O., Generative models as an emerging paradigm in the chemical sciences. J. Am. Chem. Soc. 145, 8736–8750 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.A. Subramanian et al., Closing the execution gap in generative AI for chemicals and materials: Freeways or safeguards. An MIT Exploration of Generative AI (2024). https://mit-genai.pubpub.org/pub/681kpeoa. Accessed 20 September 2024.
- 24.A. Paszke et al. , Automatic differentiation in Pytorch. Conf. Neural Inf. Process. Syst. (2017). https://openreview.net/forum?id=BJJsrmfCZ. Accessed 20 September 2024.
- 25.Walters W. P., Stahl M. T., Murcko M. A., Virtual screening–an overview. Drug Discov. Today 3, 160–178 (1998). [Google Scholar]
- 26.Shoichet B. K., Virtual screening of chemical libraries. Nature 432, 862–865 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Walters W. P., Virtual chemical libraries: Miniperspective. J. Med. Chem. 62, 1116–1124 (2018). [DOI] [PubMed] [Google Scholar]
- 28.Venkatasubramanian V., Chan K., Caruthers J. M., Evolutionary design of molecules with desired properties using the genetic algorithm. J. Chem. Inf. Comput. Sci. 35, 188–195 (1995). [Google Scholar]
- 29.Jensen J. H., A graph-based genetic algorithm and generative model/Monte Carlo tree search for the exploration of chemical space. Chem. Sci. 10, 3567–3572 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Gao W., Coley C. W., The synthesizability of molecules proposed by generative models. J. Chem. Inf. Model. 60, 5714–5723 (2020). [DOI] [PubMed] [Google Scholar]
- 31.Renz P., Van Rompaey D., Wegner J. K., Hochreiter S., Klambauer G., On failure modes in molecule generation and optimization. Drug Discov. Today Technol. 32, 55–63 (2019). [DOI] [PubMed] [Google Scholar]
- 32.Walters W. P., Barzilay R., Critical assessment of AI in drug discovery. Expert Opin. Drug Discov. 16, 937–947 (2021). [DOI] [PubMed] [Google Scholar]
- 33.Stanley M., Segler M., Fake it until you make it? Generative de novo design and virtual screening of synthesizable molecules Curr. Opin. Struct. Biol. 82, 102658 (2023). [DOI] [PubMed] [Google Scholar]
- 34.Nicolaou C. A., et al. , Idea2Data: Toward a new paradigm for drug discovery. ACS Med. Chem. Lett. 10, 278–286 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Brocklehurst C. E., et al. , Microcycle: An integrated and automated platform to accelerate drug discovery. J. Med. Chem. 67, 2118–2128 (2024). [DOI] [PubMed] [Google Scholar]
- 36.Schneider G., Automating drug discovery. Nat. Rev. Drug Discov. 17, 97–113 (2018). [DOI] [PubMed] [Google Scholar]
- 37.Coley C. W., Eyke N. S., Jensen K. F., Autonomous discovery in the chemical sciences Part I: Progress. Angew. Chem. Int. Ed. 59, 22858–22893 (2020). [DOI] [PubMed] [Google Scholar]
- 38.Coley C. W., Eyke N. S., Jensen K. F., Autonomous discovery in the chemical sciences Part II: Progress. Angew. Chem. Int. Ed. 59, 23414–23436 (2020). [DOI] [PubMed] [Google Scholar]
- 39.Koscher B. A., et al. , Autonomous, multiproperty-driven molecular discovery: From predictions to measurements and back. Science 382, eadi1407 (2023). [DOI] [PubMed] [Google Scholar]
- 40.Strieth-Kalthoff F., et al. , Delocalized, asynchronous, closed-loop discovery of organic laser emitters. Science 384, eadk9227 (2024). [DOI] [PubMed] [Google Scholar]
- 41.Seumer J., Kirschner Solberg Hansen J., Brøndsted Nielsen M., Jensen J. H., Computational evolution of new catalysts for the Morita-Baylis-Hillman reaction. Angew. Chem. Int. Ed. 62, e202218565 (2023). [DOI] [PubMed] [Google Scholar]
- 42.Ertl P., Schuffenhauer A., Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. J. Cheminform. 1, 1–11 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Thakkar A., Chadimová V., Bjerrum E. J., Engkvist O., Reymond J. L., Retrosynthetic accessibility score (rascore)-rapid machine learned synthesizability classification from AI driven retrosynthetic planning. Chem. Sci. 12, 3339–3349 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Liu C. H., et al. , Retrognn: Fast estimation of synthesizability for virtual screening and de novo design by learning from slow retrosynthesis software. J. Chem. Inf. Model. 62, 2293–2300 (2022). [DOI] [PubMed] [Google Scholar]
- 45.Gao W., Fu T., Sun J., Coley C., Sample efficiency matters: A benchmark for practical molecular optimization. Adv. Neural Inf. Process. Syst. 35, 21342–21357 (2022). [Google Scholar]
- 46.J. Guo, P. Schwaller, Saturn: Sample-efficient generative molecular design using memory manipulation. arXiv [Preprint] (2024). http://arxiv.org/abs/2405.17066 (Accessed 5 September 2024).
- 47.J. Guo, P. Schwaller, Directly optimizing for synthesizability in generative molecular design using retrosynthesis models. arXiv [Preprint] (2024). http://arxiv.org/abs/2407.12186 (Accessed 5 September 2024). [DOI] [PMC free article] [PubMed]
- 48.Hoffmann T., Gastreich M., The next level in chemical space navigation: Going far beyond enumerable compound libraries. Drug Discov. today 24, 1148–1156 (2019). [DOI] [PubMed] [Google Scholar]
- 49.Patel H., et al. , Savi, in silico generation of billions of easily synthesizable compounds through expert-system type rules. Sci. Data 7, 384 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Grygorenko O. O., et al. , Generating multibillion chemical space of readily accessible screening compounds. Iscience 23, 101681 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.WuXi AppTec, Virtual screening - hit finding and screening services (2024). https://wuxibiology.com/drug-discovery-services/hit-finding-and-screening-services/. Accessed 8 December 2022.
- 52.Neumann A., Marrison L., Klein R., Relevance of the trillion-sized chemical space “explore’’ as a source for drug discovery. ACS Med. Chem. Lett. 14, 466–472 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Bedart C., et al. , The pan-canadian chemical library: A mechanism to open academic chemistry to high-throughput virtual screening. Sci. Data 11, 597 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Vinkers H. M., et al. , Synopsis: Synthesize and optimize system in silico. J. Med. Chem. 46, 2765–2773 (2003). [DOI] [PubMed] [Google Scholar]
- 55.Hartenfeller M., et al. , DOGs: Reaction-driven de novo design of bioactive compounds. PLoS Comput. Biol. 8, e1002380 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.K. Korovina et al. , “Bayesian optimization of small organic molecules with synthesizable recommendations” in International Conference on Artificial Intelligence and Statistics, S. Chiappa, R. Calandra, Eds. (PMLR, 2020), pp. 3393–3403.
- 57.D. H. Nguyen, K. Tsuda, A generative model for molecule generation based on chemical reaction trees. arXiv [Preprint] (2021). http://arxiv.org/abs/2106.03394 (Accessed 5 September 2024).
- 58.Sadybekov A. A., et al. , Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature 601, 452–459 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Swanson K., et al. , Generative AI for designing and validating easily synthesizable and structurally novel antibiotics. Nat. Mach. Intell. 6, 338–353 (2024). [Google Scholar]
- 60.M. Sun et al., Syntax-guided procedural synthesis of molecules. arXiv [Preprint] (2024). http://arxiv.org/abs/2409.05873 (Accessed 5 September 2024).
- 61.Bradshaw J., Paige B., Kusner M. J., Segler M., Hernández-Lobato J. M., A model to search for synthesizable molecules. Adv. Neural Inf. Process. Syst. 32, 05221 (2019). [Google Scholar]
- 62.S. K. Gottipati et al. , “Learning to navigate the synthetically accessible chemical space using reinforcement learning” in International Conference on Machine Learning, H. Daumé III, A. Singh, Eds. (PMLR, 2020), pp. 3668–3679.
- 63.Horwood J., Noutahi E., Molecular design in synthetically accessible chemical space via deep reinforcement learning. ACS Omega 5, 32984–32994 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Bradshaw J., Paige B., Kusner M. J., Segler M., Hernández-Lobato J. M., Barking up the right tree: An approach to search over molecule synthesis DAGs. Adv. Neural Inf. Process. Syst. 33, 6852–6866 (2020). [Google Scholar]
- 65.W. Gao, R. Mercado, C. W. Coley, Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design. arXiv [Preprints] (2021). 10.48550/arXiv.2110.06389 (Accessed 20 September 2024). [DOI]
- 66.M. Koziarski et al., Rgfn: Synthesizable molecular generation using gflownets. arXiv [Preprint] (2024). http://arxiv.org/abs/2406.08506 (Accessed 5 September 2024).
- 67.S. Luo et al. , Projecting molecules into synthesizable chemical spaces. arXiv [Preprints] (2024). 10.48550/arXiv.2406.04628 (Accessed 20 September 2024). [DOI]
- 68.Wang M., et al. , Clickgen: Directed exploration of synthesizable chemical space via modular reactions and reinforcement learning. Nat. Commun. 15, 10127 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.S. Seo et al., Generative flows on synthetic pathway for drug design. arXiv [Preprint] (2024). http://arxiv.org/abs/2410.04542 (Accessed 5 September 2024).
- 70.Mendez D., et al. , Chembl: Towards direct deposition of bioassay data. Nucleic Acids Res. 47, D930–D940 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Morgan H. L., The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. J. Chem. Doc. 5, 107–113 (1965). [Google Scholar]
- 72.Hoogeboom E., Nielsen D., Jaini P., Forré P., Welling M., Argmax flows and multinomial diffusion: Learning categorical distributions. Adv. Neural Inf. Process. Syst. 34, 12454–12465 (2021). [Google Scholar]
- 73.Weininger D., Smiles, A chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 28, 31–36 (1988). [Google Scholar]
- 74.Zhang J., Mercado R., Engkvist O., Chen H., Comparative study of deep generative models on chemical space coverage. J. Chem. Inf. Model. 61, 2572–2581 (2021). [DOI] [PubMed] [Google Scholar]
- 75.Frey N. C., et al. , Neural scaling of deep chemical models. Nat. Mach. Intell. 5, 1297–1305 (2023). [Google Scholar]
- 76.J. Kaplan et al., Scaling laws for neural language models. arXiv [Preprint] (2020). http://arxiv.org/abs/2001.08361 (Accessed 5 September 2024).
- 77.Z. Tu et al., Askcos: An open source software suite for synthesis planning. arXiv [Preprint] (2025). http://arxiv.org/abs/2501.01835 (Accessed 5 September 2024).
- 78.Brown N., Fiscato M., Segler M. H., Vaucher A. C., Guacamol: Benchmarking models for de novo molecular design. J. Chem. Inf. Model. 59, 1096–1108 (2019). [DOI] [PubMed] [Google Scholar]
- 79.Gillet V. J., et al. , Sprout: Recent developments in the de novo design of molecules. J. Chem. Inf. Comput. Sci. 34, 207–217 (1994). [DOI] [PubMed] [Google Scholar]
- 80.Bohacek R. S., McMartin C., Guida W. C., The art and practice of structure-based drug design: A molecular modeling perspective. Med. Res. Rev. 16, 3–50 (1996). [DOI] [PubMed] [Google Scholar]
- 81.Wang R., Gao Y., Lai L., Ligbuilder: A multi-purpose program for structure-based drug design. Mol. Model. Annu. 6, 498–516 (2000). [Google Scholar]
- 82.Gillet V. J., Myatt G., Zsoldos Z., Johnson A. P., Sprout, hippo and caesa: Tools for de novo structure generation and estimation of synthetic accessibility. Perspect. Drug Discov. Des. 3, 34–50 (1995). [Google Scholar]
- 83.C. Harris et al., Benchmarking generated poses: How rational is structure-based drug design with generative models? arXiv [Preprint] (2023). http://arxiv.org/abs/2308.07413 (Accessed 5 September 2024).
- 84.Luo S., Guan J., Ma J., Peng J., A 3d generative model for structure-based drug design. Adv. Neural Inf. Process. Syst. 34, 6229–6239 (2021). [Google Scholar]
- 85.Tran-Nguyen V. K., Jacquemard C., Rognan D., LIT-PCBA: An unbiased data set for machine learning and virtual screening. J. Chem. Inf. Model. 60, 4263–4273 (2020). [DOI] [PubMed] [Google Scholar]
- 86.Trott O., Olson A. J., Autodock vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J. Comput. Chem. 31, 455–461 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Levin I., Fortunato M. E., Tan K. L., Coley C. W., Computer-aided evaluation and exploration of chemical spaces constrained by reaction pathways. AIChE J. 69, e18234 (2023). [Google Scholar]
- 88.M. J. Kusner, B. Paige, J. M. Hernández-Lobato, “Grammar variational autoencoder” in International Conference on Machine Learning, D. Precup, Y. W. Teh, Eds. (PMLR, 2017), pp. 1945–1954.
- 89.Sterling T., Irwin J. J., Zinc 15-ligand discovery for everyone. J. Chem. Inf. Model. 55, 2324–2337 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Kuan C. Y., et al. , A critical role of neural-specific JNK3 for ischemic apoptosis. Proc. Natl. Acad. Sci. U.S.A. 100, 15184–15189 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Sun J., et al. , Excape-db: An integrated large scale dataset facilitating big data analysis in chemogenomics. J. Cheminf. 9, 1–9 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Li Y., Zhang L., Liu Z., Multi-objective de novo drug design with conditional graph generative model. J. Cheminf. 10, 1–24 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Anastasiou D., et al. , Pyruvate kinase M2 activators promote tetramer formation and suppress tumorigenesis. Nat. Chem. Biol. 8, 839–847 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Humphreys P. G., et al. , Discovery of a potent, cell penetrant, and selective p300/CBP-associated factor (PCAF)/general control nonderepressible 5 (GCN5) bromodomain chemical probe. J. Med. Chem. 60, 695–709 (2017). [DOI] [PubMed] [Google Scholar]
- 95.Dolfus U., Briem H., Rarey M., Synthesis-aware generation of structural analogues. J. Chem. Inf. Model. 62, 3565–3576 (2022). [DOI] [PubMed] [Google Scholar]
- 96.Guo J., Schwaller P., Augmented memory: Sample-efficient generative molecular design with reinforcement learning. Jacs Au 4, 2160–2172 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Williams R. J., Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn. 8, 229–256 (1992). [Google Scholar]
- 98.Vallar L., Meldolesi J., Mechanisms of signal transduction at the dopamine D2 receptor. Trends Pharmacol. Sci. 10, 74–77 (1989). [DOI] [PubMed] [Google Scholar]
- 99.Bengio E., Jain M., Korablyov M., Precup D., Bengio Y., Flow network based generative models for non-iterative diverse candidate generation. Adv. Neural Inf. Process. Syst. 34, 27381–27394 (2021). [Google Scholar]
- 100.A. Tripp, J. M. Hernández-Lobato, Genetic algorithms are strong baselines for molecule generation. arXiv [Preprint] (2023). http://arxiv.org/abs/2310.09267 (Accessed 5 September 2024).
- 101.Mullowney M. W., et al. , Artificial intelligence for natural product drug discovery. Nat. Rev. Drug Discov. 22, 895–916 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Daylight Chemical Information Systems, Daylight theory: Smarts (2024). https://www.daylight.com/dayhtml/doc/theory/theory.smarts.html. Accessed 5 September 2024.
- 103.Hartenfeller M., et al. , A collection of robust organic synthesis reactions for in silico molecule design. J. Chem. Inf. Model. 51, 3093–3098 (2011). [DOI] [PubMed] [Google Scholar]
- 104.Button A., Merk D., Hiss J. A., Schneider G., Automated de novo molecular design by hybrid machine intelligence and rule-driven chemical synthesis. Nat. Mach. Intell. 1, 307–315 (2019). [Google Scholar]
- 105.Jaccard P., Étude comparative de la distribution florale dans une portion des alpes et des jura. Bull. Soc. Vaudoise Sci. Nat. 37, 547–579 (1901). [Google Scholar]
- 106.Rogers D., Hahn M., Extended-connectivity fingerprints. J. Chem. Inf. Model. 50, 742–754 (2010). [DOI] [PubMed] [Google Scholar]
- 107.Zdrazil B., et al. , The ChEMBL database in 2023: A drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 52, D1180–D1192 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Lipinski C. A., Lombardo F., Dominy B. W., Feeney P. J., Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. Drug Deliv. Rev. 23, 3–25 (1997). [DOI] [PubMed] [Google Scholar]
- 109.Veber D. F., et al. , Molecular properties that influence the oral bioavailability of drug candidates. J. Med. Chem. 45, 2615–2623 (2002). [DOI] [PubMed] [Google Scholar]
- 110.Kim S., et al. , Pubchem substance and compound databases. Nucleic Acids Res. 44, D1202–D1213 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Olivecrona M., Blaschke T., Engkvist O., Chen H., Molecular de-novo design through deep reinforcement learning. J. Cheminf. 9, 1–14 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.K. Huang et al. , Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. arXiv [Preprints] (2021). 10.48550/arXiv.2102.09548 (Accessed 20 September 2024). [DOI]
- 113.Huang K., et al. , Artificial intelligence foundation for therapeutic science. Nat. Chem. Biol. 18, 1033–1036 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.W. Gao, S. Luo, C. W. Coley, SynFormer code and reaction templates. GitHub. https://github.com/wenhao-gao/synformer. Accessed 20 September 2024.
- 115.W. Gao, S. Luo, C. W. Coley, SynFormer pre-trained models and datasets. Hugging Face. https://huggingface.co/whgao/synformer. Accessed 20 September 2024.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Data Availability Statement
The code and reaction templates are available at: https://github.com/wenhao-gao/synformer (114), while the model weights can be accessed on https://huggingface.co/whgao/synformer (115).


