Abstract
Efficient exploration of chemical space is an essential component of modern generative drug design. Herein, we introduce ChemBang, a computational engine that grows small molecules based on chemical transformations extracted by matched molecular pair analysis of all structures available in catalogues of synthesized molecules. Each chemical transformation is mapped onto its associated atomic environment defined as the substructure within a three‐atom radius from the transformation site. Unsupervised chemical evolution is then performed in cycles by systematically applying chemical transformations to all exposed atomic environments present in a seed structure. Multiple physicochemical properties and substructural alerts are incorporated to effectively guide the generation of drug‐like synthetically accessible molecules. As a use case, the generation of the Erdafitinib structure from any of its three ring systems (pyrazole, benzene and quinoxaline), and the evolution of the property distributions from all molecules generated in each cycle, are discussed in detail. The ability to explore the chemical space of pharmaceutical relevance is shown by successfully generating the exact chemical structure of 95.3% of all 2,809 small‐molecule ATC drugs from their constituting fragments.
Keywords: chemical space exploration, computational chemistry, in‐silico library design, lead optimization, matched molecular pair, molecular generator, scaffold decoration
Schematic diagram of the underlying concepts behind ChemBang and the molecular generation process from a seed molecular structure by applying the three main categories of chemical transformations, namely, interchange of small substituents, replacement of ring systems, and scaffold enlargement by additional ring systems.

Abbreviations
- A
Alimentary tract and metabolism
- additionalRingSystem
the addition of new ring systems
- ATC
Anatomical Therapeutic Chemical
- B
Blood and blood forming organs
- Bio space
bioactive compounds
- C
Cardiovascular system
- Change1
original substituent
- Change2
single modification
- changeRingSystem
scaffold replacements in which one ring system is exchanged for another while maintaining the same molecular graph
- Chem space
synthetic compound collections
- CTs
mapped chemical transformations
- D
Dermatologicals
- EAS+change1
initial condition
- EAS+change2
final condition
- ER diversity set
Enamine REAL diversity set
- fCTS
filtered chemical transformations occurring in more than 1% of ATC drugs and ChEMBL33 therapeutics
- FGFR
fibroblast growth factor receptor
- G
Genito‐urinary system and sex hormones
- H
Systemic hormonal preparations
- excl.
sex hormones and insulins
- J
Anti‐infectives for systemic use
- L
Antineoplastic and immunomodulating agents
- M
Musculo‐skeletal system
- maxSubs5 Atm
small substituent modifications involving up to five atoms
- MCS
maximum common substructure
- MLSMR
NIH Molecular Libraries Small Molecule Repository
- MMPs
Matched molecular pairs
- MW
molecular weight
- N
Nervous system
- P
Antiparasitic products, insecticides and repellents
- PAINS
Pan‐Assay Interference Compounds
- QED
quantitative estimate of drug‐likeness
- R
Respiratory system
- S
Sensory organs
- SA
synthetic accessibility
- SAR
structure‐activity relationships
- TPSA
topological polar surface area
- USPTO‐50K
USPTO‐50K reaction dataset
- V
Various
1. Introduction
Drug discovery fundamentally implies exploring the vast landscape of chemical space to identify regions enriched with molecules having optimal physicochemical, pharmacological, and toxicological parameters [1]. Recent AI‐based generative models hold promise in efficiently creating new chemical entities with desired properties without the need to evaluate ultra‐large numbers of enumerated molecules [2, 3]. Methods based on recurrent neural networks, variational autoencoders, graph neural networks and transformer‐based architectures have demonstrated the ability to generate structurally diverse molecules [4, 5, 6, 7, 8, 9, 10]. However, despite significant progress in the field, synthetic feasibility and chemical coverage remain a challenge [11]. Common limitations are the frequent generation of compounds structurally similar to the training data and of compounds with impossible chemical structures [12, 13]. In this respect, it has been shown that integrating empirical transformation knowledge helps connecting generative approaches to more realistic chemistry, thereby improving the relevance and viability of the proposed compounds for laboratory synthesis [14].
Matched molecular pairs (MMPs) are chemical substructures that vary at a single site as a result of a specific, well‐defined transformation [15]. MMP analysis has become a pivotal approach for systematic molecular transformation and property optimization [16, 17]. Since MMPs differ by a single chemical transformation, they enable insights into structure–activity relationships (SAR) and facilitate hypothesis‐driven design [18, 19]. Leveraging large datasets of experimentally derived MMPs may promote the generation of transformations to expand chemical coverage and diversity. This strategy not only enhances the expansion of chemical libraries but also aligns with medicinal chemistry goals of compound synthesizability and progression.
Recent studies have demonstrated that combining data‐driven and rule‐based approaches can enhance the exploration of more synthetically‐tractable areas of chemical space [20]. For instance, the integration of machine‐learned retrosynthetic analysis with template‐based transformations has the potential to accelerate the identification of viable synthetic routes, with advantages in lead optimization campaigns [21, 22, 23]. Moreover, large‐scale compilations of MMPs from commercial catalogues and public databases have enabled the construction of extensive libraries of transformation rules, capturing the diversity and properties inherent in medicinal chemistry datasets [24]. By leveraging this knowledge, computational tools can efficiently generate chemical libraries that better explore and cover the chemical space of pharmaceutical relevance [25, 26, 27, 28, 29].
In this study, we introduce ChemBang, a novel computational engine designed to expand the chemical space around small molecules with novel drug‐like synthetically accessible chemical matter. Using a seed chemical structure as a starting point, it applies MMP‐based transformations to generate plausible diverse molecules. By incorporating curated transformation rules, property filters, and structural constraints, it addresses the need for efficient exploration, while ensuring the chemical validity of the generated molecules. As proof of concept, we validated the performance of our approach through case studies in drug discovery, showcasing its ability to generate molecules that are not only structurally diverse but also relevant in medicinal chemistry. By combining state‐of‐the‐art MMP methodologies with scalable computational pipelines, we achieved a significant step forward in chemical space navigation.
2. Methods
ChemBang has been developed utilizing MMP‐based methodologies to facilitate the systematic expansion of molecular seeds. This approach employs thoroughly defined transformation rules, derived from extensive chemical datasets, which focus on single‐point structural modifications to induce desirable changes in molecular properties, while aligning with the knowledge of chemists [30].
2.1. Data Sources and Structural Preparation
A diverse dataset of unique molecular structures was assembled from eight in‐stock chemical catalogues—ChemDiv, ChemSpace, LifeChemicals, Mcule, Vitas, WixilabNetwork, Akos, and Zinc20 [31]—as well as the curated bioactivity database ChEMBL33 [32]. These sources (Table S1) provided a robust foundation for the identification of MMPs that underpin the ChemBang platform. The number of original identifiers, unique standardized molecules, and scaffolds contributed by each source is reported in Table S2.
All molecules were converted to SMILES notation, and their corresponding InChIKeys were generated. Compounds were then standardized and processed to extract core scaffolds and capture relevant molecular transformations. An internal molecular standardization pipeline was applied to eliminate inconsistencies and ensure compatibility with the algorithms used for MMPs identification and transformation generation. This process involved the use of RDKit libraries [33], including rdmolops, SaltRemover, and rdMolStandardize, to perform sanitization, remove salts based on a custom list while retaining the largest structure, normalize acid forms, adjust ionizable charges when a chemically valid neutral form exists, canonicalize tautomers, and filter reactive groups. These steps ensured a consistent and standardized molecular dataset, ready for further computational analysis.
2.2. Matched Molecular Pair Identification
MMPs are identified through a two‐step process. The first step involves a pre‐screening phase designed to efficiently reduce the combinatorial space of possible molecular pairs, thereby narrowing down the set of candidate pairs for further evaluation. This is followed by the detection of the maximum common substructure (MCS) [34, 35, 36, 37], which is then used to define both the anchor point and its surrounding chemical environment, referred to as the Extended Anchor Substructure (EAS). This includes all atoms within a three‐bond radius of the modification site. This step is critical because MMP rules are strongly influenced by the local structural context surrounding transformations [28, 38]. The structural differences between the two molecules in each MMP are classified as change1 (derived from the first molecule) and change2 (derived from the second molecule).
Transformations are classified into three categories: small substituent modifications involving up to five atoms (maxSubs5Atm), scaffold replacements in which one ring system is exchanged for another while maintaining the same molecular graph (changeRingSystem), and the addition of new ring systems (additionalRingSystem). This classification ensures a systematic and controlled approach to generating molecular diversity and facilitates the recognition of MMPs.
As part of the prescreening process, a system of nested dictionaries was developed to facilitate the comparison of molecules. This approach focuses on pairs of compounds that differ by small substituent modifications, specifically allowing for differences of up to five additional atoms per molecule while maintaining the same Murcko scaffold. Two distinct scenarios are considered: (a) pairs where one molecule is a substructure of the other, differing by a single substituent (e.g., benzene vs. bromobenzene), and (b) pairs in which both molecules contain a single substituent at the same position, each differing by up to five atoms (e.g., chlorobenzene vs. bromobenzene). In both cases, the comparison restricts changes to only one additional or distinct substituent per molecule, with all modifications limited to a maximum of five atoms.
For the addition of new ring systems, molecules are compared to identify pairs in which one molecule is a substructure of the other but with a missing ring system. This comparison highlights compounds with an additional ring system that can be part of a longer substituent, where the smaller one is a substructure of the larger one. Only one additional ring system is allowed.
Finally, the pre‐screening phase for scaffold replacements compares molecules with the same molecular graph and substituents but different scaffolds. In this case, only one distinct ring system is allowed to differ between both molecules, while preserving the same molecular graph.
Once the pre‐screening phase is complete, the prefiltered molecular pairs are submitted to the MMP confirmation process via the rdFMCS.FindMCS function of RDKit [33]. Here, the MCS is identified, and further analysis is performed to describe the transformation environment. This involves detecting the surrounding atoms of the change, starting from the anchor point, and detecting the EAS in a radius of 3 atoms from the transformation point, as well as calculating other structural properties related to the change.
To improve execution time, a custom MMPs detection algorithm was developed for specific maxSubs5Atm cases, avoiding the use of RDKit's native function. The algorithm begins by ensuring that both molecules share the same scaffold. The substituents are then extracted using the RDKit RGroupDecompose module, and the number of substituents in each molecule is calculated. If both molecules have the same number of substituents, one substituent must differ between the two. The position of the differing substituent is identified, ensuring that only one substituent differs at that position in each molecule. The algorithm also identifies identical substituents that are in the same position in both molecules. If common substituents are found, they are merged into the scaffold to update the preliminary version of the MCS. The algorithm proceeds by identifying common elements to generate the MCS, including the largest shared subgraph found in the differing substituents. The distinct substituents (change1 and change2) are then determined, with a maximum of one distinct substituent allowed for each molecule and identical scaffolds required.
The resulting MMPs were used to construct an internal library of chemical transformations containing detailed information, including the InChIKeys of both molecules involved, and their SMILES representations, the InChIKey of the MCS, the MCS with dummy atoms, and the structural changes observed in each molecule. Additional data stored in the library includes the number of atoms in change1 and change2, the number of atoms in the MCS, the number of aromatic and aliphatic rings in the MCS, the bond types at the anchor points in both molecules, the anchor atom, and the extended anchor substructure (within a three‐bond radius), among other features.
2.3. Generation of the Internal Library of Chemical Transformations
Chemical transformations are generated through a customized Python script, utilizing RDKit and based on MMPs information. Initially, stereochemical information is removed from the input molecule to simplify the pattern recognition task. SMARTS patterns are then derived for both the reactants and products, incorporating detailed atom indexing to accurately define the transformation sites. Specifically, SMARTS patterns are generated for EAS + change1 (representing the SMARTS of the main reactant) and EAS + change2 (representing the SMARTS of the main product). The number of hydrogens attached to the terminal atom of these substructures is also included to enhance specificity in identifying the initial condition (match) required for the transformation.
Next, the SMARTS pattern for the second product is generated, with atom numbering starting from the last index of the main product, based on change1 (the substructure to be replaced). Similarly, the SMARTS pattern for the second reactant is generated, with atom numbering starting from the last index of the main reactant, based on change2 (the substructure to be added). Reaction strings are then constructed, with all SMARTS patterns containing indices that specify the relative positions of atoms within the substructures.
In cases where the replaced substructure is a hydrogen atom, it is unnecessary to specify it as part of the second product. If hydrogen replaces another substituent, it is also unnecessary to include it as part of the second reactant. Otherwise, the substructure to be added is incorporated as the second reactant, and the substructure to be replaced is incorporated as the second product.
Finally, each transformation is validated by applying it to the initial molecule and confirming that the resulting product matches the intended MMP counterpart. Transformations are further refined using rule‐based filters that prioritize synthetically feasible and drug‐like molecules.
2.3.1. Filters Applied to Chemical Transformations
A literature review was carried out to obtain a list of structural alerts for drugs [39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56]. This was supplemented with custom alerts, and all structural alerts were expressed as SMARTS. Then, their frequency in molecules from the Anatomical Therapeutic Chemical (ATC) classification system and ChEMBL33 therapeutics with molecular weight (MW) below 1000 was calculated based on the number of drugs in these databases that matched each structural alert. Those with a frequency of 1% or less in both databases were used to filter out chemical transformations whose change2 matched these alerts.
In addition, atoms in molecules from ATC and ChEMBL33 therapeutics were characterized using RDKit native functions to determine the symbol, degree, hybridization, explicit valence, total number of hydrogens, number of implicit hydrogens, number of radical electrons, formal charge, and whether the atom is part of an aromatic system. Atoms present in less than 1% of the ATC, ChEMBL therapeutics, and ChEMBL active sets were marked as blacklisted atoms if all parameters for that atom type coincided. Transformations containing blacklisted atoms in change2 were filtered out.
Change2 was characterized by determining the number of unlisted atoms different from H, C, N, O, F, P, S, Cl, Br, I, and B, as well as MW, number of heavy atoms, number of spiro atoms, number of headbridge atoms, number of rotatable and rigid bonds, number of ring systems, maximum size of the ring systems in change2, number of rings in change2, and formal charge. Pre‐defined thresholds of these properties were used to filter transformations (Table S3), including a maximum of 11 heavy atoms per iteration in change2 [36, 57]. These properties were also monitored in ATC and ChEMBL therapeutics to avoid an abrupt decrease in ChemBang's ability to generate real drugs.
2.3.2. Most Frequent Chemical Transformations
To identify the most common chemical transformations, we conducted a statistical analysis of the MMP‐based transformation dataset available within the ChemBang platform. The dataset was queried across three categories of structural changes—maxSubs5 Atm, additionalRingSystem, and changeRingSystem—to extract and quantify individual modifications defined by pairs of chemical changes (Change1 → Change2). Here, Change1 represents the original substituent, and Change2 the modified group, as defined by transformation rules. The frequency of each transformation was assessed using the EAS count, which reflects the number of transformations that have been observed for a specific modification. Additionally, we recorded the total number of MMPs supporting these transformations, providing further insight into their observed frequency and relevance in medicinal chemistry applications.
2.3.3. Mapping Local Atomic Environments to Chemical Transformations
This study employed the USPTO‐50K reaction dataset (USPTO‐50K) curated by Coley et al. [58]., which comprises around 50 thousand chemically diverse reactions derived from patent literature [58, 59]. Each reaction is categorized into one of ten high‐level transformation classes, including heteroatom alkylation and arylation, acylation, carbon–carbon bond formation, heterocycle formation, protection and deprotection strategies, redox processes, and functional group interconversions or additions [58]. These chemical transformations collectively represent a broad swath of reactivity central to organic synthesis and reaction prediction research.
We mapped MMP‐based transformations of ChemBang to explore how local atomic environments relate to USPTO‐50K reactions. ChemBang encodes transformations as a combination of an original substituent (Change1) associated with a single modification (Change2), and an EAS that captures the local molecular environment. These elements are combined into a STRING representation, allowing transformation patterns to be matched across structurally diverse molecules.
Only a subset of the USPTO‐50K reactions fell within the applicability domain of the current version of ChemBang, due to inherent constraints of the MMP‐based transformation model. ChemBang is designed to capture single‐point transformations, consistent with the MMP formalism, which limits its ability to represent reactions involving multiple simultaneous structural changes. Additionally, in the case of acyclic substituents, the atom‐mapped transformation is restricted to five atoms or fewer. For transformations involving ring systems, at present ChemBang permits either the addition of a single ring system or a ring system replacement, but only when the graph topologies of the replaced rings are isomorphic. Reactions involving more extensive skeletal rearrangements, multiple ring additions or removals, or transformations affecting larger substructures fall outside these constraints and are thus not captured by the ChemBang representation.
The mapping process aligned reactant–product pairs from USPTO‐50K reactions with the transformation library of ChemBang, retaining only those that could be fully represented within the scope of ChemBang transformation rules. These were then analyzed to assess the diversity of local environments contributing to each reaction class.
2.4. Molecular Generation
The molecular generation process uses validated transformations to iteratively grow and diversify seed structures (Figure 1). Transformations are applied based on user‐defined parameters, guiding molecular modifications (growth, shrinkage, or size retention) through substituent changes, scaffold replacements, or ring additions. The algorithm begins by removing stereochemistry from the input structure to simplify its structural representation. It then scans the seed structure to identify the presence of atomic environments associated with chemical transformations in the library. If a match is found, the transformation is applied. Finally, only RDKit‐validated SMILES strings are preserved to ensure that the generated molecules are chemically plausible [60].
FIGURE 1.

ChemBang molecular generator pipeline.
Transformation accuracy is evaluated using multiple metrics. These include the number of MMPs supporting a transformation, as well as the number of distinct molecules exhibiting the initial condition (EAS + change1) and the number of distinct molecules exhibiting the final condition (EAS + change2). This approach helps to avoid the generation of unusual structures. Additionally, a frequency score is used to assess the likelihood of a specific transformation occurring in a chemical environment. This frequency score is calculated as the ratio of molecules exhibiting the transformation to those possessing the corresponding initial condition. This metric ensures the reliability of the generated transformations and facilitates the creation of structurally diverse and synthetically accessible molecular libraries. In this calculation the numerator represents the number of molecules with a specific environment that undergo the transformation. This value is derived by grouping the database table by the reaction string “rxn_string” and the EAS + change1 smart. The denominator counts the number of molecules containing a specific EAS linked to a particular substituent ("change1"), and this value is obtained by grouping the database table by the EAS + change1 smart.
2.4.1. User Control of the Generation Process
On the other hand, ChemBang provides a comprehensive suite of parameters that allow users to fine‐tune the molecular transformation process, offering flexibility and control over the generated molecules. Key parameters include the aforementioned frequency_score, which sets a minimum frequency threshold for the transformations, and mmps_count, which filters transformations based on the minimum number of MMPs supporting a transformation required for execution. Additional options enhance the specificity of transformations, the change parameter allows for filtering based on the type of structural modification, including options like ‘total’, ‘subsmax5atm’, ‘changerings’, and ‘additional_ringsystem’. The initial parameter controls the minimum number of occurrences of both the initial and final conditions, ensuring relevant transformations are applied. Here, initial and final conditions refer to the number of molecules exhibiting EAS + Change1 and EAS + Change2, respectively (e.g., if set to 10, at least 10 molecules must exhibit each condition), helping to ensure that both the chemical environment and the substructures involved in the transformation are sufficiently common. The require_substructure option specifies whether the generated molecules should retain the seed structure, offering further control over the transformation process. Further customization is enabled by the direction parameter, which filters transformations based on the atom count gradient (e.g., ’grow’, 'shrink’, ’equal’, or ’all’), which is designed to specify the transformation direction and other parameters. The accepted_charges and ncharged_atoms parameters allow users to define specific charge states or limits on the number of charged atoms in the transformation, while accepted_atoms permits the specification of allowed atom types, such as carbon, nitrogen and oxygen, among others.
Lastly, the config_json option provides a way to define advanced transformation parameters in a single JSON document, streamlining the configuration process. Together, these parameters offer a powerful toolkit for tailoring molecular transformations to meet precise requirements in medicinal chemistry. This will allow users to define different options for gradients, change2, or the properties of the resulting molecules. Among them, the user can decide on the number or gradient of heavy atoms, the molecular weight of heavy atoms, or the hydrogen bond donor and hydrogen bond acceptor counts. All these options work by pre‐filtering transformations that will be used to obtain results with the desired characteristics, if any. In addition, specific positions can be blocked to avoid changes in certain atoms of the molecule. To perform multiple changes the user can utilize the n‐generation script that allows n consecutive iterations of single changes based on MMP transformations.
2.4.2. Molecular Generation Capabilities
To evaluate ChemBang's capabilities for molecular generation, we systematically examined its capacity to produce structurally diverse, pharmaceutically relevant compounds. Using seed structures comprising ring systems, scaffolds, and terminal fragments (defined by a 3‐atom radius counted from the terminal atom for acyclic molecules), we applied iterative transformations across up to 30 generations while monitoring the emergence of target molecules from the ATC. This approach enables a comprehensive assessment of ChemBang coverage across therapeutic categories, ensuring broad sampling of pharmacologically relevant chemical space.
We implemented two complementary strategies to evaluate whether ChemBang transformations are sufficient to efficiently reconstruct ATC drugs from simple fragments. The decomposition‐based approach first dissected each target drug into a core scaffold/ring system with precisely annotated substituents and their EAS. Reconstruction proceeds deterministically by applying ChemBang transformations sequentially, starting from the unsubstituted seeds (EAS with change1 = H) and progressing through intermediate substitution states (EAS with change2 = substituent) until the full target structure is recovered. Parallelly, the fragment‐mapping approach identified single step substitutions (change1 = H and change2 = substituent) by mapping fragments of the target molecule. These transformations are then applied sequentially and aggregated to reconstruct the target. Both strategies independently assess ChemBang's ability to reconstruct molecules from minimal building blocks under complementary perspectives, providing robust validation of the transformation set.
For each target molecule reconstructed, we collected (i) the number of transformation generations required, (ii) the complete molecular paths from seed to target, and (iii) the success rates across ATC categories. This quantitative analysis verified the ability of ChemBang to recover bioactive compounds from minimal, chemically intuitive seeds. The dual‐method validation not only demonstrates robust navigation of drug‐like chemical space but also underscores the potential of ChemBang for scaffold‐hopping and structure‐based drug discovery applications.
Building on these reconstruction experiments, a stratified analysis was conducted to characterize the performance of ChemBang across therapeutic categories. Drug molecules were grouped according to their ATC classification, and reconstruction success rates were calculated independently for each category and seed type. This allowed assessment of the ability of ChemBang to uniformly generate molecules from diverse pharmacological classes from minimal input structures. In parallel, we analyzed the number of transformation steps required to regenerate each compound, with each generation (step) corresponding to a single MMP‐based structural modification. Generation distributions were computed cumulatively to determine how rapidly ChemBang could recover target molecules across ATC. These analyses enabled a comprehensive characterization of the generative process of ChemBang across drug‐like chemical space. In a complementary set of experiments, independent molecular generation was initiated from two simple seed fragments independently, benzene and indole, to evaluate the ability of ChemBang to produce pharmaceutically relevant compounds spanning multiple ATC therapeutic classes. Starting from these structurally minimal seeds, iterative transformations were applied to systematically explore and expand chemical space, yielding molecules across a wide range of pharmacological categories.
In a third experiment aimed at examining the diversity of reconstruction paths and convergence behavior, we performed targeted regeneration of Erdafitinib, a kinase inhibitor with multiple fused aromatic systems. Independent simulations were initiated from each of the three ring systems present in the Erdafitinib structure: quinoxaline, benzene, and pyrazole. For each starting fragment, ChemBang was used to iteratively generate molecules up to 10 generations. Molecules were filtered at each generation based on substructure inclusion criteria, and only candidates matching intermediate or final Erdafitinib substructures were retained for subsequent transformations. We documented the number of successful paths leading to full reconstruction of the target molecule, the number of generations, and SMILES of intermediates. A comparative analysis of these paths across seeds was performed to evaluate the ability of ChemBang to navigate convergent and divergent molecular routes. Besides, molecules generated at each cycle were also screened against compound databases such as chemical catalogues, namely ChEMBL, SureChEMBL, and Enamine REAL, to assess novelty and coverage of existing chemical space.
2.4.3. Chemical Diversity and Scaffold Novelty
The structural diversity of molecules generated at each ChemBang iteration was quantified using the instant similarity (iSIM) framework, which enables the efficient calculation of the average pairwise similarity of large molecular sets with linear computational scaling. For each generation, molecules were represented using 2048‐bit ECFP4 fingerprints, and similarity was evaluated using the Jaccard–Tanimoto (JT) index. Fingerprints were aggregated column‐wise, and the resulting bit‐count vectors were used to compute the iSIM Tanimoto value, which corresponds to the mean of all distinct pairwise Tanimoto similarities within the set. This approach allows direct comparison of diversity trends across generations without explicit enumeration of all molecular pairs. All calculations were performed using the iSIM Python implementation, following a methodology reported previously [61, 62].
Additionally, to assess structural novelty beyond internal diversity during the reconstruction of Erdafitinib, a scaffold‐based analysis was performed for molecules generated at each iteration. For each generation and starting seed, Bemis–Murcko scaffolds were extracted from the generated molecules and compared against the set of scaffolds present in ChEMBL33. The number and proportion of scaffolds not observed in ChEMBL33 were quantified per generation, providing an operational measure of scaffold‐level novelty relative to known bioactive chemical space. This scaffold‐based analysis complements the iSIM diversity evaluation by assessing exploration of core structural frameworks not represented in curated bioactivity datasets, rather than relying on molecular‐level uniqueness. For comparative benchmarking, the same diversity and scaffold analyses were applied to molecules generated with REINVENT4 (LibINVENT4), using Erdafitinib‐derived substructures as seeds with high‐throughput GPU sampling (105 molecules) and attachment points constrained to those present in the target drug.
2.4.4. Molecular Properties, Drug‐Likeness and Synthetic Accessibility
The evolution of molecular properties along the reconstruction trajectories of Erdafitinib was examined by analyzing the distributions of MW, topological polar surface area (TPSA), calculated logP, and quantitative estimate of drug‐likeness (QED) across eight transformation generations [9, 13, 60]. This analysis was conducted separately for each of the three molecular seeds used to initiate the reconstruction. For each generation, all produced molecules were evaluated to characterize property shifts associated with progressive structural modifications.
To evaluate the synthetic accessibility of the generated molecules, we computed two established metrics, SA [63] and SYBA [64] scores, for all structures produced across generations during the reconstruction of Erdafitinib from the three selected seeds. SA scores were calculated using the RDKit implementation [33], and molecules with SA scores below 4.5 were considered synthetically tractable under stringent criteria, with 6.0 being the threshold originally suggested by the authors [63]. SYBA scores were also computed, with positive values indicating predicted ease of synthesis [64]. Distributions of both metrics were analyzed per generation and compared against reference datasets representing: (i) Bio space, consisting of bioactive molecules from ChEMBL and SureChEMBL (∼25 million compounds); (ii) Chem space, comprising commercially available compounds from the chemical catalogues included in this study (∼18 million compounds); and (iii) the ER diversity set, a chemically diverse virtual library from Enamine (∼48 million compounds). These comparisons enabled benchmarking of ChemBang outputs against established chemical spaces and provided a robust assessment of the ability of this platform to generate synthetically accessible molecules.
2.4.5. Comparison against a Reference Platform
To contextualize ChemBang performance against a representative generative AI framework, a parallel molecular generation benchmark was conducted using REINVENT4 (LibINVENT component). For each of the three Erdafitinib‐derived seeds, REINVENT4 benchmarks were executed on an NVIDIA RTX A6000 GPU (device cuda:0) using both standard multinomial sampling and a score‐directed Staged Learning protocol (DAP strategy, sigma = 128, learning rate = 10−4). The global seed was fixed to 42 and SMILES randomization was disabled. For the baseline benchmarking, all possible attachment sites were considered, whereas for high‐volume sampling scales (104, 105 and 106 molecules) and reinforcement learning, substitution points were specifically constrained to the positions present in Erdafitinib to evaluate recovery under guided conditions. For each REINVENT4 run, execution time, number of generated molecules, throughput, and time‐per‐molecule metrics were recorded analogously to the ChemBang analysis. In addition, the number of unique Bemis–Murcko scaffolds, scaffold counts per generation, and iSIM‐based diversity metrics were computed for the REINVENT‐generated molecules using the same fingerprinting and analysis protocols applied to ChemBang outputs. This standardized benchmarking framework enables a direct, methodologically consistent comparison of iterative rule‐based generation and deep generative sampling under matched seed and computational conditions.
2.4.6. Data Visualization and Computational Cost
Molecular and data visualizations were generated using various Python libraries to facilitate result interpretation and presentation. RDKit [33] was employed for rendering molecular structure images. Quantitative analyses, including bar charts, stacked bar plots, and density distributions, were created using Matplotlib. To depict hierarchical relationships and proportions within the generation of ATC drugs by ChemBang, a Sunburst diagram was generated with Plotly Express. Key molecular properties such as MW, TPSA, logP, and QED, were visualized using violin plots produced with Matplotlib, enabling detailed assessment of their distributions and variability across generations.
To address the computational efficiency and scalability of ChemBang, execution‐time benchmarks were performed during the targeted reconstruction of Erdafitinib from three distinct starting seeds (quinoxaline, benzene, and pyrazole). For each reconstruction route, execution time was recorded independently for every transformation generation. All computations were performed on a single workstation equipped with an AMD Ryzen 7 1700 processor (8 cores, 16 threads, up to 3.0 GHz), 64 GB of RAM, running Ubuntu 22.04.5 LTS (Linux kernel 6.8.0‐90‐generic). From these measurements, multiple performance metrics were derived, including: (i) runtime per generation, (ii) cumulative runtime as a function of generation depth, (iii) throughput expressed as the number of generated structures per second per seed, (iv) average time per seed, and (v) average time per molecule. To evaluate computational scaling with molecular complexity, the average molecular weight of compounds generated at each generation was correlated with the corresponding per‐molecule execution time. These metrics provide a quantitative characterization of how ChemBang performance evolves with increasing generation depth and molecular size.
3. Results
ChemBang has been designed as a evolutionary molecular generator aiming at producing chemical structures with good synthetic feasibility and efficiently covering the chemical space of therapeutic relevance. The platform leverages MMP‐based transformations extracted from multiple chemical repositories of synthesized compounds to systematically evolve seed fragments into diverse drug‐like small molecules. Although outside the scope of this work, the generated structures can be further processed with other complementary tools, such as structure‐based docking methods, to score them on their fitness for protein targets.
3.1. Generation of a Chemical Transformation Database
A dataset comprising 19,185,826 unique molecules was used to identify MMPs that encode the internal catalogue of chemical transformations. MMPs were extracted from eight in‐stock chemical repositories, namely, ChemDiv, ChemSpace, LifeChemicals, Mcule, Vitas, WixilabNetwork, Zinc20, and Akos (Table S1), as well as ChEMBL33.
The number of original identifiers, unique standardized molecules and scaffolds per database are shown in Figure 2 (exact numbers provided in Table S2). These diverse sources provide a robust reference framework of synthesized molecules from which frequently observed chemical transformations can be identified by MMP analysis, a feature that distinguishes our approach from other existing molecular generators that rely mostly on a single repository (generally ChEMBL) [6, 7, 9, 10, 13, 60, 65].
FIGURE 2.

Data sources of chemical structures used for MMP analysis.
3.1.1. Matched Molecular Pair Analysis
Using a two‐step MMP detection workflow, a total of 100,424,135 MMPs were identified across all chemical catalogues. An initial pre‐screening phase was performed to significantly reduce the combinatorial search space (from 3.7 × 1014 possibilities to a subset in the order of 108), enabling scalable MMP identification across millions of candidate structures. This pre‐screening step applied rule‐based filters to select molecular pairs differing by a single, well‐defined transformation—such as small substituent changes, scaffold replacements, or the addition of ring systems—while preserving core structural features. In a second phase, an MCS analysis was used to define the transformation sites and map them with its local atomic environment (see Methods section). For every MMP detected, structural differences were consistently encoded as change1 and change2, corresponding to the respective variable fragments in each pair of molecules [66]. This standardized representation allowed for constructing an internal MMP library and supported systematic grouping of transformation patterns (Table 1), uncovering a wide range of chemically meaningful modifications across diverse scaffolds and compound classes.
TABLE 1.
Count summary of all matched molecular pairs (MMPs), originally mapped chemical transformations (CTs) and final set of filtered chemical transformations occurring in more than 1% of ATC drugs and ChEMBL33 therapeutics (fCTS).
| Structural change | Description | MMPs | CTs | fCTs |
|---|---|---|---|---|
| maxSubs5Atm | Small substituent modifications involving up to five atoms. | 75,645,663 | 3,822,986 | 1,546,230 |
| additionalRingSystem | Addition of new ring systems to the molecular structure. | 22,317,200 | 1,778624 | 995,038 |
| changeRingSystem | Replacement of one ring system with another while retaining core topology. | 2,461,272 | 306,768 | 216,019 |
| Total | 100,424,135 | 5,908,378 | 2,757,287 |
A comprehensive catalogue of 5,908,378 chemical transformations was constructed by translating structural changes observed in the MMP library into SMARTS‐based reaction rules. These transformations, while not representing experimentally validated chemical reactions, are derived directly from MMPs and effectively reflect local structural modifications along with their surrounding atomic environments. As such, they offer valuable information for computational applications in drug discovery, particularly in tasks involving the generation of novel molecules and scaffold hopping.
The resulting catalogue represents a diverse and chemically meaningful set of structural changes. To ensure accuracy, each transformation was applied back to its originating molecule to verify the regeneration of its matched molecular pair counterpart. This validation ensures that all transformations in the library are syntactically correct and chemically consistent with the behavior observed in the original dataset of MMPs. In line with the classification scheme used for MMPs, each transformation was assigned to one of three categories based on the nature of the structural change (Table 1).
3.1.2. Final Filtered Set of Chemical Transformations
The initial chemical transformation catalogue was subsequently refined with property and structural filters to improve the likelihood of generating drug‐like synthetically accessible small molecules of therapeutic relevance. A literature review identified 2,092 structural alerts with valid SMARTS representations [39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56], covering undesirable substructures, acute oral toxicity, highly reactive motifs, bioactivation‐prone fragments, Pan‐Assay Interference Compounds (PAINS), genotoxicity, carcinogenicity, mutagenicity, frequent hitters, and other problematic functionalities. However, to ensure relevance to drug discovery, we applied a filtering step in which only alerts matching less than 1% of drugs—within both the ATC classification and ChEMBL33 therapeutics datasets—were retained. This filtering step yielded a final collection of 1,775 structural alerts, which were employed in our workflow to remove transformations leading to substructures (Change2) containing any of these alerts.
On the other hand, an analysis of uncommon atom types – defined by atomic symbol, degree, hybridization state, explicit valence, hydrogen count (explicit and implicit), radical electron count, and formal charge – was conducted. The identification of atom types occurring in less than 1% of ATC and ChEMBL33 drug structures resulted in 386 atoms present in change2 that are not common in these datasets. Therefore, transformations containing them were removed.
Finally, some property filters (Table S3) were applied to further refine the transformation set, enhancing the likelihood that the generated molecules would exhibit drug‐like properties. After the whole filtering process, 2,757,287 transformations were retained (46.7% of all chemical transformations originally extracted). The distribution of these transformations across the three major structural modification categories is summarized in Table 1.
The molecular generation framework implemented in the workflow identifies the atomic environments of seed molecules and intermediates to apply MMP‐based transformations in a stepwise fashion, across single or multiple generations with user‐customized settings. Each generation introduces a single structural modification to the preceding molecule, based on three transformation types that control changes: maxSubs5Atm, additionalRingSystem, changeRingSystem (Table 1; see Methods section for details). This controlled process enables targeted diversification guided by the various options available in the molecular generator engine, allowing for flexible exploration of chemical space.
3.1.3. Most Frequent Chemical Transformations
The most common MMP‐based transformations identified in the analyzed chemical catalogues include acyclic transformations, the addition of aromatic systems, and ring replacements involving topologically equivalent scaffolds.
As shown in Table 2, the most frequent acyclic transformations derived from MMPs primarily involve simple substituent modifications. These changes are consistent with well‐established bioisosteric transformations [67, 68] and they have been reported to be highly prevalent in the NIH Molecular Libraries Small Molecule Repository (MLSMR) [19].
TABLE 2.
Top 10 most frequent small substituent transformations (maxSubs5Atm) in in‐stock catalogues.
| Change1 | Change2 | Number of transformations (EAS count) | Number of MMPs |
|---|---|---|---|
|
|
6,118 | 3,096,714 |
|
|
2,205 | 223,819 |
|
|
2,093 | 429,332 |
|
|
1,753 | 149,217 |
|
|
1,443 | 45,768 |
|
|
1,438 | 74,940 |
|
|
1,419 | 105,810 |
|
|
1,325 | 1,095,096 |
|
|
1,304 | 82,407 |
|
|
1,281 | 742,281 |
The single most common transformation is the replacement of a hydrogen atom with a methyl group (–CH3), observed in 6,118 transformations and represented in over 3 million MMPs. This prevalence emphasizes both the ubiquity and synthetic accessibility of this modification. Notably, the methyl group ranks among the most common functional groups in bioactive compounds, appearing in more than 67% of top‐selling pharmaceutical drugs [69]. The transformation of a C–H bond to a C–Me group (methylation) has shown remarkable potential to enhance drug potency [70], with some cases reporting over a 100‐fold improvement in activity [69, 71].
Other highly frequent transformations include the substitution of hydrogen with hydroxyl (–OH, 2,205 transformations), ethyl (–CH2CH3, 2,093), and carboxylic acid (–COOH, 1,753). Insertions of polar functional groups, such as –CH2OH (1,443), –COOCH3 (1,438), and –NH2 (1,419), also rank among the most frequent modifications. Notably, most of these groups have been reported among the top 20 most frequently used R‐groups in ChEMBL bioactive compounds [72].
Electronegative substituents appear prominently as well, with chlorine (–Cl) showing particularly high frequency (1,325 transformations; >1 million MMP instances). Chlorine is present in over 250 drugs approved by the United States Food and Drug Administration (FDA) [73]. Moreover, the replacement of –H with –Cl can lead to substantial changes in pharmacokinetic properties and biological potency, with reported enhancements of up to 100,000‐fold [74]. Interestingly, fluorine (–F) also presented a high frequency of transformations (1,281). Insertion of halogen atoms is a commonly employed strategy to fine‐tune a range of molecular properties, including metabolic stability, membrane permeability, steric profile, and electronic distribution [75, 76]. Additionally, they play a critical role in optimizing target binding affinity and selectivity [76, 77].
The molecular scaffold, as the core structural framework of a compound, defines its shape, rigidity, and substituent positioning, while also influencing key properties such as hydrophobicity, polarity, reactivity, metabolic stability, and overall behavior in biological systems [78]. The most frequent ring transformations are summarized in Table S4. Such modifications represent some of the most common medicinal chemistry strategies employed in lead optimization and the discovery of novel core substructures [79]. The addition of ring systems allows for extending the core structure to a new scaffold that explores a whole new region of increased complexity in chemical space, enabling the modulation of pharmacologically relevant features, such as molecular size and shape, rigidity, and electrostatics, potentially leading to improved binding affinity when optimizing protein–ligand interactions [80].
The most prevalent transformation in this category is the substitution of hydrogen with a phenyl group (–H → –Ph), observed in 2,589 transformations and associated with more than 433,000 MMPs. Other frequent transformations include additions of benzyl groups (–CH2Ph, 1,340 transformations), acyl‐substituted phenyl groups (e.g., phenyl ketone derivatives), and saturated or small aliphatic rings such as cyclopropyl (734 transformations), cyclohexyl (592), and piperidine (531). These ring insertions are widely utilized in medicinal chemistry to support the design of novel drug candidates and refine their interaction profiles with biological targets [81, 82, 83].
The addition of heteroaromatic rings, including pyridines, pyrazoles, and thiazoles, also feature prominently, consistent with their established roles as privileged scaffolds in drug design [84, 85, 86, 87, 88]. In addition, modifications incorporating morpholine, piperazine, and other nitrogen‐ or oxygen‐containing heterocycles occur frequently. These heterocycles have been reported to enhance pharmacokinetics profiles, metabolic properties and biological activity while maintaining synthetic accessibility [89, 90].
The most common scaffold‐level transformations involving ring replacements—where one ring is substituted for another within an otherwise preserved molecular framework—are summarized in Table S5. As a fundamental scaffold‐hopping approach, changeRingSystem transformations allow for the efficient exploration of chemical space, metabolic stability and optimization of biological activity while preserving key pharmacophoric elements [78, 91].
The most frequent ring system replacement is benzene to pyridine (3‐ and 2‐substituted isomers), with 845 and 843 transformations, respectively. Other prominent modifications include benzene to pyrazine (714 transformations), thiophene to furan (469) and morpholine to piperidine (422). Many of the ring systems involved in this kind of transformations are among the top 100 most commonly used rings in FDA‐approved drugs, as listed in the Orange Book [83] and are also highlighted in medicinal chemistry resources such as the ring replacement recommender, as common ring replacements [92]. These include benzene, pyridine, thiophene, furan, morpholine, piperidine, cyclohexane and pyrazine [83].
3.1.4. Mapping Local Atomic Environments to Chemical Transformations
Chemical transformations within acyclic systems include single modifications affecting no more than five atoms. For ring systems, only the addition of a single ring is permitted. Ring substitutions are allowed solely when the topological graphs of the original and resulting rings are isomorphic. Consequently, complex reactions—such as multi‐step processes, scaffold rearrangements, or transformations involving multiple simultaneous bond changes—fall outside the applicability domain of single transformations and were excluded from the analysis.
Mapping the USPTO‐50K dataset to our catalogue of chemical transformations resulted in 15,325 successfully matched reactions. These matched reactions span 9 of the 10 reaction classes defined in the USPTO‐50K classification scheme (Table S6). Particularly, the heterocycle formation class was not represented, as its structural complexity is beyond the scope of the MMP‐based transformation model.
While this limitation reflects the intrinsic design of our framework, it is important to note that chemical transformations are intended to be applied iteratively, allowing the expansion of a molecular seed over multiple transformation generations. Therefore, the single‐change feature does not significantly limit the utility of our approach, but rather supports a gradual, modular exploration of chemical space. It should be emphasized that transformations in ChemBang do not correspond to real chemical reactions, and the generated steps or paths are not intended as practical laboratory synthesis routes. ChemBang is not designed as a synthetic guidance tool; consequently, discrepancies with the USPTO‐50K dataset do not undermine the validity of the transformations for systematic chemical space exploration. Despite these constraints, the mapped subset provides meaningful insight into how local atomic environments encode broader chemical transformations (Table 3), reinforcing their value for systematic, context‐aware reaction modelling.
TABLE 3.
Top 10 most frequent USPTO‐50K reactions mapped to ChemBang transformations, showing the associated changes and number of unique local environments (EAS count).
| ChemBang change based on MMPs |
Number of ChemBang transformations (EAS count) |
Number of USPTO‐50K reactions mapped to ChemBang transformations |
|---|---|---|
| *C → *H | 12 | 2202 |
| *CC → *H | 2 | 1390 |
| *H → *C | 168 | 789 |
| *Cc1ccccc1 → *H | 44 | 732 |
| *=O → *O | 84 | 532 |
| *=O → *H | 94 | 476 |
| *O → *=O | 69 | 436 |
| *H → *C(C)=O | 79 | 435 |
| *C(C)=O → *H | 39 | 355 |
| *C(C)(C)C → *H | 2 | 346 |
The mapped reactions reveal important insights into the relationship between local atomic environments and globally defined chemical transformations. A key observation is the presence of moderate local environment diversity within certain transformation classes, where a small number of chemical transformations, defined by a specific structural change and EAS, account for a large number of USPTO‐50K reactions. This indicates that some chemical transformations, despite varying in substrate structure, can be abstracted into a small set of generalized, MMP‐based templates. For example, the demethylation transformation (*C → *H) is captured by just 12 chemical transformations, yet these map to over 2,000 reactions in the USPTO‐50K dataset.
The abstraction of high‐level transformations into context‐aware fragment representations is conceptually aligned with methodologies like AutoTemplate [93], which enhance reaction datasets by addressing the limitations of whole‐molecule representations and emphasizing the reactive core. While our implementation achieves this through its defined EAS, AutoTemplate automates the extraction of generalized reaction templates. This shared paradigm of isolating the essential chemical event from the molecular scaffold enables more robust and interpretable modelling of reaction patterns.
Figure 3 illustrates this concept by presenting an example transformation involving the replacement of a hydrogen atom with a methyl, mapped to multiple local atomic environments as defined by EAS. Each EAS represents a unique context in which the same high‐level transformation occurs. Importantly, these local environments correspond to well‐documented pharmaceutical reactions described in U.S. patents. Panel A shows a transformation used in the synthesis of doravirine (US20130296382A1), where an H → C substitution takes place within a constrained heteroaromatic system. Panel B presents a similar transformation in the production of aminoquinazoline‐based protein tyrosine kinase modulators (US20140228361A1) but occurring within a distinct heterocyclic framework. Panel C shows the transformation within the context of substituted tricyclic FGFR inhibitors (US20130338134A1). These examples underscore how a single formal transformation, when abstracted from an MMP, can be embedded within highly diverse structural motifs and synthetic routes.
FIGURE 3.

Illustrative example of MMP‐based chemical transformations (H → C), mapped to multiple atomic environments represented as Extended Anchor Substructures (EAS). These transformations are associated with known pharmaceutical reactions described in US patents (right). (A) Synthesis of Doravirine (US20130296382A1). (B) Production of aminoquinazoline‐based protein tyrosine kinase modulators (US20140228361A1). (C) Formation of substituted tricyclic FGFR inhibitors (US20130338134A1).
This figure highlights the key point that even when transformations are grouped under a common category, such as H → C substitution, the specific atomic environments in which they occur can differ substantially. By capturing local chemical context through EAS representations, we can enable a more nuanced and chemically meaningful characterization of transformations. This localized framework provides a clear advantage in representing transformation patterns relevant to medicinal chemistry and drug development.
3.2. Use Case: Generation of Erdafitinib from Its Ring Systems
To illustrate the ability of our molecular generator to delineate a path of chemical transformations that allow a simple ring or fragment to evolve into a drug‐like small molecule, we will use the case of generating the exact chemical structure of Erdafitinib, a fibroblast growth factor receptor (FGFR) inhibitor, from its ring systems (pyrazole, benzene, and quinoxaline). This experiment uses a target‐guided reconstruction in which only molecules that are substructures of Erdafitinib advance, allowing analysis of transformation coverage, path diversity, and synthetic accessibility while property trends should be interpreted within this guided context. This constrained reconstruction represents an analogue of practical iterative generation workflows, where molecules are typically selected at each cycle based on target affinity, predicted activity, synthetic accessibility, or other scoring criteria. The results are summarized in Table 4.
TABLE 4.
Number of molecular reconstruction paths generated by ChemBang to obtain Erdafitinib from different molecular seeds, along with the corresponding number of generations (transformation steps) and the count of distinct molecules produced in each generation. PYR: pyrazole; BEN: benzene; QUI: quinoxaline; ISC + CB33: in‐stock catalogues and ChEMBL33; E‐REAL: Enamine REAL.
| Seed | Paths | Source | Count of distinct molecules produced by generation | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | |||
| PYR | 448 | ChemBang | 1,492 | 21,097 | 78,488 | 289,294 | 594,929 | 819,375 | 730,905 | 275,915 |
| ISC + CB33 | 283 | 413 | 17 | 2 | 8 | 66 | 131 | 42 | ||
| E‐REAL | 232 | 187 | 3 | 0 | 0 | 0 | 1 | 1 | ||
| Substructures of Erdafitinib | 3 | 5 | 12 | 21 | 27 | 24 | 9 | 1 | ||
| BEN | 528 | ChemBang | 3,798 | 66,319 | 386,349 | 1,000,824 | 1,538,469 | 1,459,253 | 871,828 | 358,947 |
| ISC + CB33 | 2,127 | 5,039 | 1,795 | 200 | 17 | 25 | 126 | 57 | ||
| E‐REAL | 548 | 1809 | 526 | 73 | 9 | 0 | 1 | 1 | ||
| Substructures of Erdafitinib | 5 | 20 | 42 | 57 | 49 | 29 | 12 | 1 | ||
| QUI | 1344 | ChemBang | 8,154 | 73,734 | 300,067 | 644,985 | 1,039,248 | 1,088,180 | 730,905 | 275,915 |
| ISC + CB33 | 210 | 62 | 2 | 2 | 8 | 66 | 131 | 42 | ||
| E‐REAL | 76 | 18 | 0 | 0 | 0 | 0 | 1 | 1 | ||
| Substructures of Erdafitinib | 4 | 13 | 24 | 36 | 36 | 24 | 9 | 1 | ||
Starting from any of the three ring systems, ChemBang is able to reconstruct the exact chemical structure of Erdafitinib in 7–8 generations. Interestingly, a strong signal of convergent evolution is observed and the Erdafitinib structure is not reproduced once but 2,320 times, 448 of those originating from pyrazole, 528 from benzene, and 1,344 from quinoxaline (Table 4). This means that not only different seed ring systems can lead to the same chemical structure, but the same ring system is capable of evolving to the Erdafitinib structure from multiple unique paths that apply different chemical transformations in a different order. This exceptional chemical coverage is reassuring of the diversity of atomic environments mapped to chemical transformations implemented in ChemBang. A representative 7‐generation path leading to Erdafitinib from each of the three ring systems is provided in Figure S1.
The structural diversity of molecules generated at each ChemBang iteration was systematically evaluated using the iSIM framework, providing an efficient, scalable metric of mean pairwise similarity across large molecular sets. Figure 4 summarizes the generative process starting from minimal seeds (pyrazole, benzene and quinoxaline), with subsequent generations derived from molecules containing substructures of Erdafitinib identified in the previous iteration. The figure illustrates both the total count of molecules generated per iteration and the corresponding iSIM values calculated stepwise. Notably, the overall iSIM remained at levels comparable to large chemical databases such as PubChem [61], highlighting that ChemBang maintains a broad chemical space coverage. Within each generation, iSIM consistently decreased when comparing the initial seed pool to the newly generated molecules, demonstrating a clear expansion of structural diversity at every iteration and reflecting the capacity of ChemBang to explore novel regions of chemical space efficiently.
FIGURE 4.

Count of structures generated by ChemBang at each iteration during the reconstruction of Erdafitinib from different molecular seeds, along with the corresponding iSIM values. Panels correspond to paths starting from pyrazole (A), benzene (B), and quinoxaline (C). For each step, molecules that are substructures of Erdafitinib and used as seeds for the next generation are shown as gray dots, while newly generated molecules are shown as blue dots. Dashed lines represent the expansion from seeds to generated molecules (left) and the reduction of similarity—increase of diversity—between seeds and generated molecules (right). Generation steps correspond to successive transformation steps.
In addition to overall molecular diversity, scaffold‐level novelty was assessed to capture structural innovation beyond internal similarity metrics. For each generation, Bemis–Murcko scaffolds were extracted from the generated molecules and compared against scaffolds cataloged in ChEMBL33. Figure 5 shows the number and proportion of scaffolds absent from ChEMBL33, providing an operational measure of scaffold‐level novelty relative to known bioactive chemical space. This analysis revealed that successive ChemBang iterations progressively sample previously unrepresented scaffolds, complementing the iSIM‐based assessment by demonstrating the exploration of core structural frameworks rather than merely molecular‐level uniqueness. Together, these analyses indicate that ChemBang simultaneously expands both the diversity and the novelty of the generated molecular sets, achieving a balance between extensive coverage of chemical space and the generation of previously unseen scaffolds.
FIGURE 5.

Number of unique Murcko scaffolds generated by ChemBang during the reconstruction of Erdafitinib from different molecular seeds, along with the number of scaffolds not present in ChEMBL33. Panels correspond to paths starting from (A) pyrazole, (B) benzene, and (C) quinoxaline. Generations represent successive transformation steps.
As can be observed (Table 4), the count of unique molecules produced at each generation grows rapidly during the early to mid‐stages (generations 3–6), peaking around generation 5 or 6 depending on the seed. For example, chemical transformations applied to quinoxaline and its transformed derivatives produce almost 1.1 million unique molecules at generation 6, illustrating the degree of chemical diversity generated by the process. Then, the number of unique molecules produced decays rapidly at generations 7 and 8. This is an expected outcome as there should be a turning point (generation 7 for this case) in which the size and complexity of the molecules start reducing the possibility of being a substructure of the target molecule that one is trying to reconstruct. Indeed, the analysis of how several molecular properties (MW, TPSA, and logP) evolve over generations (Figure 6) displays a gradual increase in all three properties, with logP reaching a plateau at around generation 6. This trend is mirrored by a gradual decrease of QED in all molecules generated from the three ring systems.
FIGURE 6.

Distribution of molecular weight (MW), Topological Polar Surface Area (TPSA), calculated logP and quantitative estimation of drug‐likeness (QED) for each generation in the reconstruction paths of Erdafitinib generated by ChemBang, starting from three different molecular seeds: (A) pyrazole, (B) benzene, and (C) quinoxaline.
The analysis of how synthetic accessibility (SA) scores evolve over generations reveals that most molecules produced by ChemBang have SA scores below 4.5 (Figure 7), a more stringent cutoff than the commonly used threshold of 6, which typically distinguishes between compounds considered easy or difficult to synthesize [63]. This lower threshold still retains the majority of molecules included in three reference sets, namely, ChEMBL bioactive compounds (Bio space), synthetic compound collections (Chem space), and the Enamine REAL diversity set (ER diversity set). Notably, the average SA scores obtained for molecules generated by ChemBang are lower than those reported for other machine learning–based generative models [6, 10] or evolutionary algorithms [94]. On the other hand, most molecules generated by ChemBang across all generations exhibit SYBA scores above 0 (Figure S2), indicating a predicted ease of synthesis. Furthermore, the SYBA score distributions of all molecules produced follow closely those observed in the three reference spaces, providing further support for the validity of the chemical structures generated.
FIGURE 7.

Synthetic Accessibility scores (SA scores) for each generation in the reconstruction paths of Erdafitinib generated by ChemBang, starting from three different molecular seeds: (A) pyrazole, (B) benzene, and (C) quinoxaline. The dashed gray line indicates a stricter SA score reference of 4.5, while the commonly suggested threshold by the original SA score authors is 6. “Bio space” refers to molecules from ChEMBL33 and SureChEMBL; “Chem space” includes in‐stock compound catalogues used in this study; and the “ER diversity set” corresponds to the Enamine REAL diversity set.
Finally, we checked the percentage of molecules generated at each cycle that were present in all in‐stock chemical catalogues considered (directly purchasable) and/or Enamine REAL (readily accessible). As captured in Table 4, 56.0%, 19.0%, and 2.6% of all molecules generated in the first generation from benzene, pyrazole, and quinoxaline, respectively, could be directly purchased from some chemical provider catalogue. The number of generated purchasable molecules decay for the following 3–4 generations until it increases again to peak at the turning point generation (generation 7 in this case). At generation 7, a total of 126, 131, and 131 molecules evolved from benzene, pyrazole, and quinoxaline, respectively, are directly purchasable. Since generation 7 is the generation at which the structure of Erdafitinib is produced from most paths, it is likely that this result reflects the presence of Erdafitinib analogues in catalogues and that they are produced by ChemBang.
3.3. Computational Performance and Benchmarking
To quantify the computational efficiency of the ChemBang reconstruction process, execution‐time metrics were collected for each generation during the targeted regeneration of Erdafitinib from the three molecular seeds (pyrazole, benzene, and quinoxaline). Figure 8 summarizes the runtime behavior across generations, including per‐generation runtime, cumulative runtime, throughput, and time normalized per seed and per molecule. Across all three starting points, runtime increased with generation depth, reflecting the rapid growth in the number and complexity of generated molecules during intermediate stages of the reconstruction. Despite this increase, ChemBang maintained high throughput in the early and mid‐generations, producing hundreds of structures per second per seed, before a gradual reduction in throughput at later stages as molecular size and structural constraints increased. Analysis of time per molecule as a function of average molecular weight shows a clear scaling relationship, indicating that computational cost increases smoothly with molecular complexity rather than exhibiting abrupt degradation. These results demonstrate that ChemBang enables large‐scale exploration of chemically relevant space with predictable and tractable computational scaling during iterative, target‐guided generation.
FIGURE 8.

Computational performance of ChemBang during the reconstruction of Erdafitinib from three molecular seeds. (A) Runtime per generation; (B) cumulative runtime as a function of generation depth; (C) throughput expressed as generated structures per second per seed; (D) time per seed; (E) time per molecule as a function of generation; and (F) time per molecule as a function of average molecular weight. Results are shown for paths initiated from pyrazole (PYR), benzene (BEN), and quinoxaline (QUI). Generation steps correspond to successive transformation iterations. Parentheses in A show seeds per generation.
To contextualize these results, analogous molecular generation experiments were conducted using REINVENT4 (LibINVENT4) across expanded sampling scales and score‐directed optimization. In contrast to ChemBang, REINVENT4 did not produce the exact structure of Erdafitinib from any of the three seeds via score‐directed optimization or at standard sampling scales. Recovery of the target was only achieved from the quinoxaline seed when sampling was increased to 105 molecules under constrained substitution conditions (limited to those positions present in Erdafitinib). Notably, Erdafitinib was not recovered from benzene or pyrazole seeds under the same constrained substitution conditions even with an increased sampling size of 106 molecules, which resulted in prohibitively expensive execution times. Moreover, no substructures of Erdafitinib were generated in the first sampling iteration starting from the benzene or pyrazole seeds. For the quinoxaline seed, sampled molecules in later iterations did not introduce novel Erdafitinib substructures beyond those already observed in earlier outputs, indicating limited progressive coverage of the target structure. These results highlight a fundamental methodological difference between iterative, transformation‐based generation and single‐pass deep generative sampling when applied to constrained target reconstruction tasks.
The generative outputs of REINVENT4 were further analyzed using the same diversity and scaffold‐based metrics applied to ChemBang. As shown in Figure 9 (quinoxaline seed), REINVENT4 sampling produced a comparatively limited number of distinct scaffolds, with most generated structures remaining close to the initial seed frameworks and exhibiting relatively stable iSIM values across sampling steps. In contrast, ChemBang shows a clear generational expansion in both molecular diversity and scaffold novelty, while still retaining chemically meaningful trajectories toward the target structure. Notably, ChemBang simultaneously achieves increasing scaffold‐level novelty and convergent reconstruction of Erdafitinib. In contrast, the REINVENT4 (LibINVENT4) sampling protocol, even when scaled to 105 molecules with guided substitution points, favors local exploration around the input scaffolds; while it recovers Erdafitinib from the quinoxaline seed (Figure 10), it fails to reconstruct the target molecule from the benzene or pyrazole seeds. Together, these analyses illustrate how iterative, chemically interpretable transformations enable ChemBang to balance diversity, novelty, and target coverage in a manner that is not readily achieved by default generative AI sampling workflows.
FIGURE 9.

Diversity and scaffold analysis of molecules generated using REINVENT4 with quinoxaline to produce Erdafitinib (achieved in three iterations). (A) Number of generated molecules across high‐volume sampling steps (105 molecules with GPU); (B) corresponding iSIM values; and (C) number of distinct Bemis–Murcko scaffolds, including the fraction not present in ChEMBL33. Results indicate limited scaffold expansion and stable similarity profiles across sampling iterations even under these expanded sampling conditions.
FIGURE 10.

Computational performance of REINVENT4 during 105 sampling from quinoxaline to produce Erdafitinib (achieved in three iterations) on an NVIDIA RTX A6000 GPU. (A) Runtime per sampling step; (B) cumulative runtime; (C) throughput per seed; (D) time per seed; (E) time per molecule as a function of sampling step; and (F) time per molecule as a function of average molecular weight. At this scale, Erdafitinib was only recovered from the quinoxaline seed. Parentheses show seeds per step.
3.4. Chemical Coverage of ATC Drugs
ChemBang's ability to generate small molecules of therapeutical interest from chemical fragments was evaluated on the list of 2,809 drugs from the ATC. ChemBang is able to reconstruct 95.3% of all ATC drugs by using ring systems (79.6%), scaffolds (9.3%; non‐ring‐reconstructible), and terminal acyclic fragments (3‐atom radius; 5.6%) as initial molecular seeds (Figure 11). We employed both generation algorithms described in Section 2.4.2 in parallel for the reconstruction of ATC drugs. The decomposition‐based approach applied ChemBang transformations sequentially to unsubstituted seeds, progressing through intermediate substitution states until the full target structure was recovered. Simultaneously, the fragment‐mapping approach identified single‐step substitutions by mapping fragments of the target molecules and sequentially applied these transformations to reconstruct the compounds. This dual strategy ensured robust reconstruction from minimal building blocks, while providing independent validation of ChemBang's ability to generate structurally diverse, pharmaceutically relevant molecules. Unlike the constrained regeneration experiments performed for Erdafitinib, ATC drug reconstruction did not involve iterative growth or substructure‐based pruning. Instead, the experiment verified the existence of valid reconstruction paths within the ChemBang transformation space that enable the production of the majority of ATC drugs from minimal building blocks, confirming that these compounds can be reconstructed whether substructure‐based constraints are applied or unconstrained runs are carried out. This near‐complete exploration of pharmaceutical chemical space from minimal structural inputs reflects a good coverage of the chemical transformations extracted from current catalogues of synthesized molecules. Comparatively, generative models, like those implemented in REINVENT, have been reported to exhibit low recovery rates of middle‐ to late‐stage compounds from real‐world drug discovery campaigns, highlighting the inherent difficulties to reproduce complex, pharmaceutically relevant structures [65].
FIGURE 11.

Coverage of ATC drugs generated by ChemBang from ring system, scaffold, and acyclic seeds (terminal fragments, radius = 3 atoms).
A close analysis at Murcko scaffolds of drugs revealed robust decoration capabilities. We assessed the ability of ChemBang to reconstruct drugs through stepwise addition of open‐chain substituents (≤5 atoms) or elongation of existing chains. Of the initial set of 2,809 drugs, 204 compounds lacked scaffold/ring systems, while 81 possessed scaffolds identical to the parent drug. From the remaining 2,524 molecules, ChemBang successfully reconstructed 2,353 drugs (93.23%) from their core scaffolds, demonstrating robust decoration capabilities. This performance was consistent across all ATC categories, with each group achieving >70% coverage.
We further probed the ability of ChemBang to generate bioactive compounds from minimal structural elements by using either ring systems or terminal fragments (3‐atom radius for acyclic drugs) as seeds. Within the test set, 17 acyclic drugs were identical to their terminal fragments (radius: 3 atoms), and 7 compounds were composed exclusively of a single ring system without substituents. From the remaining 2,785 molecules, ChemBang reconstructed 2,392 drugs (85.89%), with 2,236 derived from ring system seeds and 156 from acyclic terminal fragments. These results demonstrate the versatility of ChemBang in exploring drug‐like chemical space through systematic construction of bioactive molecules from scaffolds, ring systems or acyclic fragments (Table S7).
The percentage of drug molecules successfully generated by ChemBang exceeds 85% across all ATC categories when considering all studied seed types—ring systems, Murcko scaffolds (non‐ring‐reconstructible), and terminal acyclic fragments (Figure 12). This uniform performance across therapeutic classes underscores the broad applicability and robustness of ChemBang in reconstructing structurally diverse bioactive compounds. Together, these results validate the potential of ChemBang as a useful molecular generator for navigating pharmaceutical chemical space from minimal starting units.
FIGURE 12.

Percentage of drug molecules in each ATC category successfully generated by ChemBang from ring systems (cyclic drugs), Murcko scaffolds (cyclic drugs), and acyclic fragments. (A: Alimentary tract and metabolism, B: Blood and blood forming organs, C: Cardiovascular system, D: Dermatologicals, G: Genito‐urinary system and sex hormones, H: Systemic hormonal preparations, excl. sex hormones and insulins, J: Anti‐infectives for systemic use, L: Antineoplastic and immunomodulating agents, M: Musculo‐skeletal system, N: Nervous system, P: Antiparasitic products, insecticides and repellents, R: Respiratory system, S: Sensory organs, V: Various).
We employed a hierarchical strategy with ChemBang for molecular reconstruction. For each drug, generation was first attempted from ring systems, followed by Murcko scaffolds if ring‐based reconstruction failed, and finally terminal acyclic fragments (radius = 3 atoms) for molecules lacking rings. Using this approach, 2,236 drugs were successfully reconstructed from ring seeds (mean generations: 5.39, median: 5), 260 from scaffold seeds (mean generations: 3.97, median: 3), and 156 from acyclic fragments (mean generations: 3.55, median: 3). Reconstruction from ring systems generally requires more generations, reflecting the greater challenge of building complete molecules from minimal cyclic motifs, whereas scaffold and acyclic seeds facilitated faster recovery. These results demonstrate that ChemBang efficiently generates bioactive molecules from diverse starting fragments while highlighting the impact of initial seed type on reconstruction speed and complexity.
The analysis of the number of generations (steps) associated with the molecular reconstruction process by ChemBang reveals that 50% of ATC‐classified drugs are successfully generated within just four transformation steps, 75% within six steps, and over 90% by generation nine (Figure 13). Since each generation corresponds to a single structural change derived from MMP transformations, these results highlight the efficiency of ChemBang in traversing drug‐like chemical space through minimal, chemically grounded modifications. Most molecules are reconstructed between generations 2 and 8, indicating that relatively few iterations are needed to recover most compounds. Only a small fraction (<5%) of drugs require more than 12 steps, underscoring the ability of ChemBang to reach a broad range of bioactive structures through concise, stepwise transformations. This performance reinforces the applicability of ChemBang to fragment‐based design strategies.
FIGURE 13.

Number and cumulative distribution of ATC‐classified drugs reconstructed by ChemBang across generations. A, Number of drugs successfully reconstructed at each generation. B. Cumulative distribution of reconstructed drugs over successive generations.
To further illustrate the capacity of ChemBang to explore drug‐like chemical space from minimal inputs, Figure 14 showcases representative generated structures across all ATC categories, starting from a simple benzene seed. Despite the constrained molecular origin, ChemBang produces chemically diverse and therapeutically relevant compounds, spanning alimentary tract modulators (A), antineoplastics (L), nervous system drugs (N), and other major classes.
FIGURE 14.

Chemical structures of representative ATC‐class drugs generated from benzene as a seed using ChemBang. Numbers in parentheses indicate the minimum generations required for derivation from the benzene seed. (A: Alimentary tract and metabolism, B: Blood and blood forming organs, C: Cardiovascular system, D: Dermatologicals, G: Genito‐urinary system and sex hormones, H: Systemic hormonal preparations, excl. sex hormones and insulins, J: Anti‐infectives for systemic use, L: Antineoplastic and immunomodulating agents, M: Musculo‐skeletal system, N: Nervous system, P: Antiparasitic products, insecticides and repellents, R: Respiratory system, S: Sensory organs, V: Various).
Comparable results were obtained with indole as the starting seed, generating more than 25 ATC‐classified drugs spanning multiple therapeutic categories (Table S8 and Figure S3). This demonstrates the ability of ChemBang to efficiently convert small, well‐defined fragments (e.g., ring systems, scaffolds, or short chains) into structurally varied, bioactive molecules, reinforcing its utility in scaffold hopping, lead optimization, and drug design.
4. Discussion
One of the key strengths of ChemBang lies in its knowledge‐based chemistry‐driven design. By leveraging MMPs extracted from several curated chemical catalogues, ChemBang ensures that its transformations reflect real‐world, experimentally observed chemical modifications. This contrasts with many deep generative models, which often produce structures that are synthetically implausible or chemically unstable. The rule‐based framework of ChemBang prioritizes frequently‐observed synthetically‐reasonable chemical transformations applied to chemically‐functionalizable atomic environments. In terms of chemical space exploration, ChemBang demonstrates remarkable coverage of drug‐like molecules. It can reconstruct 95.3% of ATC‐classified drugs from minimal molecular seeds, a capability achieved primarily from simple ring systems, but also with significant contributions from core scaffolds and acyclic fragments, thereby highlighting its versatility in navigating chemical space from diverse starting points. This efficiency is further highlighted by the observation that half of the known drugs are recovered in four steps or fewer, and 90% in nine steps or less, underscoring the applicability of ChemBang in fragment‐based drug design and scaffold amplification. Another important advantage of ChemBang is its controlled and interpretable generation process. Each transformation is validated for chemical correctness using RDKit [33], and users can adjust multiple parameters—including atom type restrictions, frequency scores, and charge constraints—to direct the generative process. This level of user control and mechanistic transparency is rarely available in black‐box AI models. ChemBang also excels in terms of synthetic feasibility. Molecules generated tend to have low SA scores [63], typically below 4.5, suggesting a high likelihood of laboratory synthesis. Furthermore, the calculated SYBA scores [64] of generated compounds align well with those of known drug‐like molecules, reinforcing the practical utility of the platform in real‐world drug discovery settings. Finally, ChemBang promotes molecular novelty and diversity, generating millions of unique molecules not present in major commercial or public databases, such as commercially available chemical catalogues, Enamine REAL or ChEMBL [32]. This capability allows researchers to navigate underexplored regions of chemical space, potentially uncovering new bioactive scaffolds or chemotypes.
One of the limitations of ChemBang's stepwise, single‐transformation approach is that it can be slower than the one step full molecule generation performed by diffusion models, which may limit rapid hypothesis generation. Additionally, there is a potential for property drift in extended generation chains. While early generations tend to maintain favorable drug‐like properties (e.g., logP, QED), later generations may accumulate undesirable features, such as excessive molecular weight or suboptimal polarity, as a result of increasing chemical complexity. To mitigate this, property‐based constraints could be incorporated dynamically at each generation step to guide molecule selection and ensure retention of drug‐like characteristics throughout the expansion process. This could be coupled with complementary tools that enable intermediate filtering and a more streamlined workflow. Integrating ChemBang into automated pipelines for prioritization, scoring, or post‐processing could mitigate these limitations and enhance its applicability in high‐throughput workflows.
The performance benchmarking and comparative experiments with REINVENT4 provide critical insight into the distinct operational profiles of iterative rule‐based generation versus deep generative sampling. For the constrained task of reconstructing Erdafitinib from minimal ring‐system seeds, ChemBang achieved complete, convergent regeneration in minutes. In stark contrast, REINVENT4 sampling failed to recover the Erdafitinib structure from simpler seeds, despite substantially longer execution times; high‐volume 105 sampling runs required approximately 2 h of GPU compute time without achieving reconstruction from benzene or pyrazole; recovery was only achieved from the most structurally similar seed (quinoxaline) under constrained substitution conditions. Notably, score‐directed optimization (reinforcement learning) also failed to yield the exact target structure across all seeds within the allocated 500 optimization steps. This outcome highlights a fundamental methodological divergence: ChemBang is engineered to exhaustively enumerate chemically validated transformation paths under structural constraints, ensuring coverage and convergence. Deep generative models like REINVENT4, optimized for rapid exploration of the latent space around an input scaffold, do not inherently guarantee such path completeness or target reachability. Moreover, the model's dependence on user‐marked substitution sites further constrains the search space. As demonstrated in our benchmarks, even when these sites were specifically guided to match the target drug, the deep generative approach lacked the path completeness necessary to reconstruct the target from more distant starting points.
Therefore, ChemBang's computational cost is not merely “higher per molecule"; rather, it is intrinsically devoted to a different objective—guaranteeing the systematic exploration of chemically plausible routes to generate molecules, including therapeutically relevant and bioactive compounds, through an exhaustive, path‐complete enumeration. This stands in contrast to the benchmarked generative approach, which, in this focused reconstruction scenario from a minimal seed, did not demonstrate an ability to reach the specified target under the same iterative, substructure‐filtered workflow.
This makes ChemBang's paradigm particularly suited for target‐driven, fragment‐based, and scaffold‐hopping workflows where path interpretability and transformation completeness are paramount. Crucially, this systematic approach is designed to be coupled with user‐defined filters, scoring functions, or active learning models at each generation cycle. Such integration provides the necessary control to manage the combinatorial expansion, prioritize promising chemical subspaces, and focus computational resources on the most relevant branches, effectively transforming the exhaustive enumeration into a guided and tractable exploration process for multi‐cycle hit expansion and optimization.
Moreover, ChemBang is dependent on the coverage and diversity of its underlying MMPs dataset. While extensive, the transformation rules are constrained by the content of public databases, and rare or proprietary transformations may be underrepresented. This could limit the ability of the platform to model highly specific or unconventional chemical modifications, such as the ones leading to complex cyclic drugs and long aliphatic drugs that mainly compose the 4.7% of ATC drugs that ChemBang is not able to generate at present. Besides, unlike retrosynthesis tools, ChemBang does not explicitly validate whether the sequence of transformations corresponds to feasible chemical reactions, as these are merely based on MMPs. Although structures remain chemically valid, synthetic routes are not provided. These limitations are also evident in diffusion‐based molecular generative models, where generalizability to novel or chemically distinct scaffolds remains a significant challenge, particularly when such regions of chemical space are underrepresented in the training data. Moreover, these generative approaches often demonstrate an inherent disconnect between computational metrics and real‐world feasibility, especially in terms of synthetic accessibility and biological activity, underscoring the need for complementary validation tools to support downstream applications more effectively [95].
Finally, all stereochemical information is removed during the molecular generation process by ChemBang, which may pose challenges for applications involving chirality‐sensitive targets such as GPCRs or metabolic enzymes. This simplification may lead to the omission of crucial stereoisomers that differ significantly in activity, selectivity, or pharmacokinetics. One possible way to address this limitation is to reconstruct 3D molecular models with plausible stereoisomeric configurations from the structures generated by ChemBang, prior to further in silico evaluation using techniques such as molecular docking or pharmacophore modelling. This post‐processing step could restore stereochemical detail and improve relevance for structure‐based design workflows.
Beyond demonstrating the ability of ChemBang to reconstruct known drug structures, we quantitatively assessed whether the observed growth in library size corresponds to a genuine increase in chemical diversity, rather than a simple expansion in molecular cardinality. To this end, structural diversity across generations was evaluated using the iSIM methodology originally proposed to disentangle diversity from set size effects in large chemical libraries. Using this framework, the average pairwise similarity of ChemBang‐generated molecules remained consistently low across generations, with mean iSIM values below 0.3 and comparable to those reported for large, heterogeneous databases such as PubChem. Moreover, within each generation, iSIM values decreased systematically from the seed sets to the generated molecules, indicating that ChemBang actively expands the structural diversity encoded in the initial fragments. These observations directly address the well‐established notion that an increase in the number of molecules does not necessarily translate into increased chemical diversity, and place ChemBang's generative behavior within the same quantitative framework previously used to analyze diversity trends in public compound repositories.
Scaffold‐level analysis further supports the observed expansion of structural diversity while clarifying the nature of the novelty introduced by ChemBang. Although the platform can reproduce known scaffolds along reconstruction paths, a substantial fraction of the generated Bemis–Murcko scaffolds are not present in curated bioactivity datasets such as ChEMBL33, with 81.9–99.9% of scaffolds from the second generation onward being absent from this reference set. Importantly, ChemBang does not introduce new chemical rules; instead, it recombines experimentally observed MMPs transformations derived from in‐stock compounds to generate novel, synthetically accessible molecules. As a consequence, exploration of chemical space is inherently bounded by the coverage of the underlying MMPs dataset, and novelty arises primarily through the recombination and extension of existing ring systems and scaffolds rather than de novo chemistry. This constraint has been explicitly acknowledged as a limitation, but it also ensures that the generated molecules remain chemically realistic and experimentally plausible. Taken together, these results indicate that ChemBang expands drug‐like chemical space in a controlled and interpretable manner, balancing scaffold novelty with synthetic feasibility rather than aiming for unconstrained or speculative chemical exploration.
5. Conclusions
We have demonstrated that ChemBang is a robust and versatile tool for the in silico generation of diverse molecular libraries with high relevance to medicinal chemistry. By leveraging a range of validated transformations, including the addition and removal of short substituents, the inclusion and interchange of ring systems, and scaffold modifications, ChemBang facilitates the generation of structurally diverse compounds and supports the exploration of a broad chemical space. Notably, ChemBang is capable of reconstructing the vast majority of ATC drugs from their ring systems, scaffolds or acyclic fragments demonstrating its potential in drug discovery and lead optimization. This functionality enables scaffold hopping and molecular decoration for yielding viable drug candidates from both small and larger molecular seeds. To ensure the quality and developability of the generated compounds, ChemBang incorporates medicinal chemistry filters that promote synthetic accessibility, safety, and drug‐likeness. These features make ChemBang a valuable tool for building ultra‐large virtual libraries of novel synthesizable molecules that can be computationally prioritized for synthesis and biological evaluation, accelerating the path from virtual molecules to real hits, leads, and drug candidates.
Supporting Information
Additional supporting information can be found online in the Supporting Information section.
Author Contributions
DMG, RGS, and JM conceived the study, designed ChemBang and planned the experiments. DMG implemented ChemBang. DMG, LM, and RGS ran experiments. DMG and JM wrote the manuscript.
Conflicts of Interest
DMG, LM, RGS, and JM are currently employees of the company Chemotargets, of which JM is the founder and co‐owner.
Supporting information
Supplementary Material
Acknowledgments
The authors gratefully acknowledge all researchers at Chemotargets for their input and feedback during the construction of the platform.
Contributor Information
Diana Montes‐Grajales, Email: diana.montes@chemotargets.com.
Jordi Mestres, Email: jordi.mestres@chemotargets.com.
Data Availability Statement
A description of the chemical catalogues used to extract all chemical transformations is provided in the Supplementary Information. ChemBang is a proprietary software developed by Chemotargets S.L. and it is not publicly available. Access to the code and implementation details is therefore restricted to collaborative work.
References
- 1. Mullard A., “The Drug‐Maker's Guide to the Galaxy,” Nature 549 (2017): 445–447, 10.1038/549445a. [DOI] [PubMed] [Google Scholar]
- 2. Pirnay J., Rittig J. G., Wolf A. B., et al., “GraphXForm: Graph Transformer for Computer‐Aided Molecular Design,” Digital Discovery 4 (2025): 1052–1065, 10.1039/D4DD00339J. [DOI] [Google Scholar]
- 3. Tong X., Liu X., Tan X., et al., “Generative Models for De Novo Drug Design,” Journal of Medicinal Chemistry 64 (2021): 14011–14027, 10.1021/acs.jmedchem.1c00927. [DOI] [PubMed] [Google Scholar]
- 4. Ishitani R., Kataoka T., and Rikimaru K., “Molecular Design Method Using a Reversible Tree Representation of Chemical Compounds and Deep Reinforcement Learning,” Journal of Chemical Information and Modeling 62 (2022): 4032–4048, 10.1021/acs.jcim.2c00366. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Nigam A., Pollice R., Krenn M., et al., “Beyond Generative Models: Superfast Traversal, Optimization, Novelty, Exploration and Discovery (STONED) Algorithm for Molecules Using SELFIES,” Chemical Science 12 (2021): 7079–7090, 10.1039/D1SC00231G. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Tan Y., Dai L., Huang W., et al., “DRlinker: Deep Reinforcement Learning for Optimization in Fragment Linking Design,” Journal of Chemical Information and Modeling 62 (2022): 5907–5917, 10.1021/acs.jcim.2c00982. [DOI] [PubMed] [Google Scholar]
- 7. Loeffler H. H., He J., Tibo A., et al., “Reinvent 4: Modern AI–driven Generative Molecule Design,” Journal of Cheminformatics 16 (2024): 20, 10.1186/s13321-024-00812-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Mercado R., Rastemo T., Lindelöf E., et al., “Graph Networks for Molecular Design,” Machine Learning: Science and Technology 2 (2021): 025023, 10.1088/2632-2153/abcf91. [DOI] [Google Scholar]
- 9. Kotsias P.‐C., Arús‐Pous J., Chen H., et al., “Direct Steering of De Novo Molecular Generation with Descriptor Conditional Recurrent Neural Networks,” Nature Machine Intelligence 2 (2020): 254–265. [Google Scholar]
- 10. Ai C., Yang H., Liu X., et al., “MTMol‐GPT: De Novo Multi‐Target Molecular Generation with Transformer‐Based Generative Adversarial Imitation Learning,” PLOS Computational Biology 20 (2024): e1012229, 10.1371/journal.pcbi.1012229. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Bilodeau C., Jin W., Jaakkola T., et al., “Generative Models for Molecular Discovery: Recent Advances and Challenges,” WIREs Computational Molecular Science 12 (2022): e1608. [Google Scholar]
- 12. Blanchard A. E., Stanley C., and Bhowmik D., “Using GANs with Adaptive Training Data to Search for New Molecules,” Journal of Cheminformatics 13 (2021): 14, 10.1186/s13321-021-00494-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Arús‐Pous J., Patronov A., Bjerrum E. J., et al., “SMILES‐Based Deep Generative Scaffold Decorator for De‐Novo Drug Design,” Journal of Cheminformatics 12 (2020): 38, 10.1186/s13321-020-00441-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. de la Vega de León A., and Bajorath J., “Matched Molecular Pairs Derived by Retrosynthetic Fragmentation,” Medicinal Chemistry Communications 5 (2014): 64–67, 10.1039/C3MD00259D. [DOI] [Google Scholar]
- 15. Tyrchan C. and Evertsson E., “Matched Molecular Pair Analysis in Short: Algorithms, Applications and Limitations,” Computational and Structural Biotechnology Journal 15 (2017): 86–90, 10.1016/j.csbj.2016.12.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Landry M. L., “Deriving Insights for Molecular Design with MMP Analysis,” Trends in Chemistry 6 (2024): 346–348, 10.1016/j.trechm.2024.04.001. [DOI] [Google Scholar]
- 17. Dossetter A. G., Griffen E. J., and Leach A. G., “Matched Molecular Pair Analysis in Drug Discovery,” Drug Discovery Today 18 (2013): 724–731, 10.1016/j.drudis.2013.03.003. [DOI] [PubMed] [Google Scholar]
- 18. Kramer C., Fuchs J. E., Whitebread S., Gedeck P., and Liedl K. R., “Matched Molecular Pair Analysis: Significance and the Impact of Experimental Uncertainty,” Journal of Medicinal Chemistry 57 (2014): 3786–3802, 10.1021/jm500317a. [DOI] [PubMed] [Google Scholar]
- 19. Hussain J. and Rea C., “Computationally Efficient Algorithm to Identify Matched Molecular Pairs (MMPs) in Large Data Sets,” Journal of Chemical Information and Modeling 50 (2010): 339–348, 10.1021/ci900450m. [DOI] [PubMed] [Google Scholar]
- 20. Kattuparambil A. A., Chaurasia D. K., Shekhar S., et al., “Exploring Chemical Space for “Druglike” Small Molecules in the Age of AI,” Frontiers in Molecular Biosciences 12 (2025), 10.3389/fmolb.2025.1553667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Genheden S., Thakkar A., Chadimová V., et al., “AiZynthFinder: A Fast, Robust and Flexible Open‐Source Software for Retrosynthetic Planning,” Journal of Cheminformatics 12 (2020): 70, 10.1186/s13321-020-00472-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Luttens A., Cabeza de Vaca I., Sparring L., et al., “Rapid Traversal of Vast Chemical Space Using Machine Learning‐Guided Docking Screens,” Nature Computational Science 5 (2025): 301–312, 10.1038/s43588-025-00777-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Chen S., Noh J., Jang J., Kim S., Gu G. H., and Jung Y., “Reaction Templates: Bridging Synthesis Knowledge and Artificial Intelligence,” Accounts of Chemical Research 57 (2024): 1964–1972, 10.1021/acs.accounts.4c00261. [DOI] [PubMed] [Google Scholar]
- 24. Kramer C., Ting A., Zheng H., et al., “Learning Medicinal Chemistry Absorption, Distribution, Metabolism, Excretion, and Toxicity (ADMET) Rules from Cross‐Company Matched Molecular Pairs Analysis (MMPA,” Journal of Medicinal Chemistry 61 (2018): 3277–3292, 10.1021/acs.jmedchem.7b00935. [DOI] [PubMed] [Google Scholar]
- 25. O’Boyle N. M., Boström J., Sayle R. A., and Gill A., “Using Matched Molecular Series as a Predictive Tool To Optimize Biological Activity,” Journal of Medicinal Chemistry 57 (2014): 2704–2713, 10.1021/jm500022q. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Raut J. A. and Dixit V. A., “A Context‐Based Matched Molecular Pair Analysis Identifies Structural Transformations that Reduce CYP1A2 Inhibition,” RSC Medicinal Chemistry 16 (2025): 3281–3290, 10.1039/D4MD01012D. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Lester C. C. and Yan G., “A Matched Molecular Pair (MMP) Approach for Selecting Analogs Suitable for Structure Activity Relationship (SAR)‐Based Read across,” Regulatory Toxicology and Pharmacology 124 (2021): 104966, 10.1016/j.yrtph.2021.104966. [DOI] [PubMed] [Google Scholar]
- 28. Dalke A., Hert J., and Kramer C., “Mmpdb: An Open‐Source Matched Molecular Pair Platform for Large Multiproperty Data Sets,” Journal of Chemical Information and Modeling 58 (2018): 902–910, 10.1021/acs.jcim.8b00173. [DOI] [PubMed] [Google Scholar]
- 29. Sushko Y., Novotarskyi S., Körner R., Vogt J., Abdelaziz A., and Tetko I. V., “Prediction‐Driven Matched Molecular Pairs to Interpret QSARs and Aid the Molecular Optimization Process,” Journal of Cheminformatics 6 (2014): 48, 10.1186/s13321-014-0048-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. He J., You H., Sandström E., et al., “Molecular Optimization by Capturing Chemist's Intuition Using Deep Neural Networks,” Journal of Cheminformatics 13 (2021): 26, 10.1186/s13321-021-00497-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Irwin J. J., Tang K. G., Young J., et al., “ZINC20—A Free Ultralarge‐Scale Chemical Database for Ligand Discovery,” Journal of Chemical Information and Modeling 60 (2020): 6065–6073, 10.1021/acs.jcim.0c00675. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Zdrazil B., Felix E., Hunter F., et al., “The ChEMBL Database in 2023: A Drug Discovery Platform Spanning Multiple Bioactivity Data Types and Time Periods,” Nucleic Acids Research 52 (2024): D1180–D1192, 10.1093/nar/gkad1004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. RDKit , accessed August 4, 2025, https://www.rdkit.org/.
- 34. Raymond J. W., Watson I. A., and Mahoui A., “Rationalizing Lead Optimization by Associating Quantitative Relevance with Molecular Structure Modification,” Journal of Chemical Information and Modeling 49 (2009): 1952–1962, 10.1021/ci9000426. [DOI] [PubMed] [Google Scholar]
- 35. Sheridan R. P., “The Most Common Chemical Replacements in Drug‐Like Compounds,” Journal of Chemical Information and Computer Sciences 42 (2002): 103–108, 10.1021/ci0100806. [DOI] [PubMed] [Google Scholar]
- 36. Leach A. G., Jones H. D., Cosgrove D. A., et al., “Matched Molecular Pairs as a Guide in the Optimization of Pharmaceutical Properties; a Study of Aqueous Solubility, Plasma Protein Binding and Oral Exposure,” Journal of Medicinal Chemistry 49 (2006): 6672–6682, 10.1021/jm0605233. [DOI] [PubMed] [Google Scholar]
- 37. Sheridan R. P., Hunt P., and Culberson J. C., “Molecular Transformations as a Way of Finding and Exploiting Consistent Local QSAR,” Journal of Chemical Information and Modeling 46 (2006): 180–192, 10.1021/ci0503208. [DOI] [PubMed] [Google Scholar]
- 38. Papadatos G., Alkarouri M., Gillet V. J., et al., “Lead Optimization Using Matched Molecular Pairs: Inclusion of Contextual Information for Enhanced Prediction of hERG Inhibition, Solubility, and Lipophilicity,” Journal of Chemical Information and Modeling 50 (2010): 1872–1886, 10.1021/ci100258p. [DOI] [PubMed] [Google Scholar]
- 39. Sirois S., Hatzakis G., Wei D., Du Q., and Chou K.‐C., “Assessment of Chemical Libraries for Their Druggability,” Computational Biology and Chemistry 29 (2005): 55–67, 10.1016/j.compbiolchem.2004.11.003. [DOI] [PubMed] [Google Scholar]
- 40. Ashby J. and Tennant R. W., “Chemical Structure, Salmonella Mutagenicity and Extent of Carcinogenicity as Indicators of Genotoxic Carcinogenesis among 222 Chemicals Tested in Rodents by the U.S. NCI/NTP,” Computational Biology and Chemistry 204 (1988): 17–115, 10.1016/0165-1218(88)90114-0. [DOI] [PubMed] [Google Scholar]
- 41. Bailey A. B., Chanderbhan R., Collazo‐Braier N., et al., “The use of Structure–activity Relationship Analysis in the Food Contact Notification Program,” Regulatory Toxicology and Pharmacology 42 (2005): 225–235, 10.1016/j.yrtph.2005.04.006. [DOI] [PubMed] [Google Scholar]
- 42. Barratt M. D., Basketter D. A., Chamberlain M., Admans G. D., and Langowski J. J., “An Expert System Rulebase for Identifying Contact Allergens,“ Toxicology in Vitro 8 (1994): 1053–1060, 10.1016/0887-2333(94)90244-5. [DOI] [PubMed] [Google Scholar]
- 43. Benigni R. and Bossa C., “Structure Alerts for Carcinogenicity, and the Salmonella Assay System: A Novel Insight through the Chemical Relational Databases Technology,” Mutation Research/Reviews in Mutation Research 659 (2008): 248–261, 10.1016/j.mrrev.2008.05.003. [DOI] [PubMed] [Google Scholar]
- 44. Brenk R., Schipani A., James D., et al., “Lessons Learnt from Assembling Screening Libraries for Drug Discovery for Neglected Diseases,” ChemMedChem 3 (2008): 435–444, 10.1002/cmdc.200700139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Claesson A. and Minidis A., “Systematic Approach to Organizing Structural Alerts for Reactive Metabolite Formation from Potential Drugs,” Chemical Research in Toxicology 31 (2018): 389–411, 10.1021/acs.chemrestox.8b00046. [DOI] [PubMed] [Google Scholar]
- 46. Ehmki E. S. R., Schmidt R., Ohm F., and Rarey M., “Comparing Molecular Patterns Using the Example of SMARTS: Applications and Filter Collection Analysis,” Journal of Chemical Information and Modeling 59 (2019): 2572–2586, 10.1021/acs.jcim.9b00249. [DOI] [PubMed] [Google Scholar]
- 47. Hann M., Hudson B., Lewell X., et al., “Strategic Pooling of Compounds for High‐Throughput Screening,” Journal of Chemical Information and Computer Sciences 39 (1999): 897–902, 10.1021/ci990423o. [DOI] [PubMed] [Google Scholar]
- 48. Kalgutkar A. S. and Soglia J. R., “Minimising the Potential for Metabolic Activation in Drug Discovery,“ Expert Opinion on Drug Metabolism and Toxicology 1 (2005): 91–142, 10.1517/17425255.1.1.91. [DOI] [PubMed] [Google Scholar]
- 49. Kazius J., McGuire R., and Bursi R., “Derivation and Validation of Toxicophores for Mutagenicity Prediction,” Journal of Medicinal Chemistry 48 (2005): 312–320, 10.1021/jm040835a. [DOI] [PubMed] [Google Scholar]
- 50. Lagorce D., Sperandio O., Baell J. B., Miteva M. A., and Villoutreix B. O., “FAF‐Drugs3: A Web Server for Compound Property Calculation and Chemical Library Design,” Nucleic Acids Research 43 (2015): W200–W207, 10.1093/nar/gkv353. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Castro Nascimento C. M., Guimarães Moura P., and Silva Pimentel A., “Generating Structural Alerts from Toxicology Datasets Using the Local Interpretable Model‐Agnostic Explanations Method,“ Digital Discovery (2023), 10.1039/D2DD00136E. [DOI] [Google Scholar]
- 52. Olsen L., Montefiori M., Tran K. P., and Jørgensen F. S., “SMARTCyp 3.0: Enhanced Cytochrome P450 Site‐of‐Metabolism Prediction Server,” Bioinformatics 35 (2019): 3174–3175, 10.1093/bioinformatics/btz037. [DOI] [PubMed] [Google Scholar]
- 53. Pearce B. C., Sofia M. J., Good A. C., Drexler D. M., and Stock D. A., “An Empirical Process for the Design of High‐Throughput Screening Deck Filters,” Journal of Chemical Information and Modeling 46 (2006): 1060–1068, 10.1021/ci050504m. [DOI] [PubMed] [Google Scholar]
- 54. Pizzo F., Gadaleta D., Lombardo A., Nicolotti O., and Benfenati E., “Identification of Structural Alerts for Liver and Kidney Toxicity Using Repeated Dose Toxicity Data,” Chemistry Central Journal 9 (2015): 62, 10.1186/s13065-015-0139-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Tinkov O. V., Grigorev V. Y., Polishchuk P. G., Yarkov A. V., and Raevsky O. A., “QSAR Investigation of Acute Toxicity of Organic Compounds during Oral Administration to Mice,“ Biomeditsinskaya Khimiya 65 (2019): 123–132, 10.18097/PBMC20196502123. [DOI] [PubMed] [Google Scholar]
- 56. Wijeyesakere S. J., Wilson D. M., Settivari R., Auernhammer T. R., Parks A. K., and Marty M. S., “Development of a Profiler for Facile Chemical Reactivity Using the Open‐Source Konstanz Information Miner,“ Applied In Vitro Toxicology 4 (2018): 202–213, 10.1089/aivt.2017.0040. [DOI] [Google Scholar]
- 57. Fink T. and Reymond J.‐L., “Virtual Exploration of the Chemical Universe up to 11 Atoms of C, N, O, F: Assembly of 26.4 Million Structures (110.9 Million Stereoisomers) and Analysis for New Ring Systems, Stereochemistry, Physicochemical Properties, Compound Classes, and Drug Discovery,” Journal of Chemical Information and Modeling 47 (2007): 342–353, 10.1021/ci600423u. [DOI] [PubMed] [Google Scholar]
- 58. Coley C. W., Rogers L., Green W. H., and Jensen K. F., “Computer‐Assisted Retrosynthesis Based on Molecular Similarity,” ACS Central Science 3 (2017): 1237–1245, 10.1021/acscentsci.7b00355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Schneider N., Stiefl N., and Landrum G. A., “What's What: The (Nearly) Definitive Guide to Reaction Role Assignment,” Journal of Chemical Information and Modeling 56 (2016): 2336–2346, 10.1021/acs.jcim.6b00564. [DOI] [PubMed] [Google Scholar]
- 60. Berenger F. and Tsuda K., “Molecular Generation by Fast Assembly of (Deep)SMILES Fragments,” Journal of Cheminformatics 13 (2021): 88, 10.1186/s13321-021-00566-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Lopez Perez K., López‐López E., Soulage F., Felix E., Medina‐Franco J. L., and Miranda‐Quintana R. A., “Growth Vs Diversity: A Time‐Evolution Analysis of the Chemical Space,” Journal of Chemical Information and Modeling 65 (2025): 6788–6796, 10.1021/acs.jcim.5c00347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. López‐Pérez K., Kim T. D., and Miranda‐Quintana R. A., “iSIM: instant similarity,” Digital Discovery 3 (2024): 1160–1171, 10.1039/D4DD00041B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Ertl P. and Schuffenhauer A., “Estimation of Synthetic Accessibility Score of Drug‐Like Molecules Based on Molecular Complexity and Fragment Contributions,” Journal of Cheminformatics 1 (2009): 8, 10.1186/1758-2946-1-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Voršilák M., Kolář M., Čmelo I., and Svozil D., “SYBA: Bayesian Estimation of Synthetic Accessibility of Organic Compounds,” Journal of Cheminformatics 12 (2020): 35, 10.1186/s13321-020-00439-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Handa K., Thomas M. C., Kageyama M., Iijima T., and Bender A., “On the Difficulty of Validating Molecular Generative Models Realistically: A Case Study on Public and Proprietary Data,” Journal of Cheminformatics 15 (2023): 112, 10.1186/s13321-023-00781-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Lukac I., Zarnecka J., Griffen E. J., et al., “Turbocharging Matched Molecular Pair Analysis: Optimizing the Identification and Analysis of Pairs,” Journal of Chemical Information and Modeling 57 (2017): 2424–2436, 10.1021/acs.jcim.7b00335. [DOI] [PubMed] [Google Scholar]
- 67. Lima L. M. and Barreiro E. J., “Bioisosterism: A Useful Strategy for Molecular Modification and Drug Design,“ Current Medicinal Chemistry 12 (2005): 23–49, 10.2174/0929867053363540. [DOI] [PubMed] [Google Scholar]
- 68. Ali G., Subhan F., and Islam N. U. et al., Input of Isosteric and Bioisosteric Approach in Drug Design.
- 69. Schönherr H. and Cernak T., “Profound Methyl Effects in Drug Discovery and a Call for New C‐H Methylation Reactions,” Angewandte Chemie International Edition 52 (2013): 12256–12267, 10.1002/anie.201303207. [DOI] [PubMed] [Google Scholar]
- 70. Barreiro E. J., Kümmerle A. E., and Fraga C. A. M., “The Methylation Effect in Medicinal Chemistry,” Chemical Reviews 111 (2011): 5215–5246, 10.1021/cr200060g. [DOI] [PubMed] [Google Scholar]
- 71. Pinheiro P. S. M., Franco L. S., and Fraga C. A. M., “The Magic Methyl and Its Tricks in Drug Discovery and Development,” Pharmaceuticals 16 (2023): 1157, 10.3390/ph16081157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Takeuchi K., Kunimoto R., and Bajorath J., “R‐group Replacement Database for Medicinal Chemistry,” Future Science OA 7 (2021): FSO742, 10.2144/fsoa-2021-0062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Joshi S. and Srivastava R., “Effect of “Magic Chlorine” in Drug Discovery: An in Silico Approach,” RSC Advances 13 (2023): 34922–34934, 10.1039/D3RA06638J. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Chiodi D. and Ishihara Y., ““Magic Chloro”: Profound Effects of the Chlorine Atom in Drug Discovery,” Journal of Medicinal Chemistry 66 (2023): 5305–5331, 10.1021/acs.jmedchem.2c02015. [DOI] [PubMed] [Google Scholar]
- 75. Meanwell N. A., “Fluorine and Fluorinated Motifs in the Design and Application of Bioisosteres for Drug Design,” Journal of Medicinal Chemistry 61 (2018): 5822–5880, 10.1021/acs.jmedchem.7b01788. [DOI] [PubMed] [Google Scholar]
- 76. Hernandes M., Cavalcanti S. M., Moreira D. R., de Azevedo Junior W., and Leite A. C., “Halogen Atoms in the Modern Medicinal Chemistry: Hints for the Drug Design,“ Current Drug Targets 11 (2010): 303–314, 10.2174/138945010790711996. [DOI] [PubMed] [Google Scholar]
- 77. Lu Y., Liu Y., Xu Z., Li H., Liu H., and Zhu W., “Halogen Bonding for Rational Drug Design and New Drug Discovery,“ Expert Opinion on Drug Discovery 7 (2012): 375–383, 10.1517/17460441.2012.678829. [DOI] [PubMed] [Google Scholar]
- 78. Ertl P., “Database of Bioactive ring Systems with Calculated Properties and Its use in Bioisosteric Design and Scaffold Hopping,” Bioorganic and Medicinal Chemistry 20 (2012): 5436–5442, 10.1016/j.bmc.2012.02.058. [DOI] [PubMed] [Google Scholar]
- 79. Wang S., Zhang R., Li X., et al., “Recent Advances in Molecular Representation Methods and Their Applications in Scaffold Hopping,” npj Drug Discovery 2 (2025): 14, 10.1038/s44386-025-00017-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80. Shearer J., Castro J. L., Lawson A. D. G., MacCoss M., and Taylor R. D., “Rings in Clinical Trials and Drugs: Present and Future,” Journal of Medicinal Chemistry 65 (2022): 8699–8712, 10.1021/acs.jmedchem.2c00473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Bauer M. R., Di Fruscia P., Lucas S. C. C., et al., “Put a ring on It: Application of Small Aliphatic Rings in Medicinal Chemistry,” RSC Medicinal Chemistry 12 (2021): 448–471, 10.1039/d0md00370k. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82. Talele T. T., “The “Cyclopropyl Fragment” Is a Versatile Player that Frequently Appears in Preclinical/Clinical Drug Molecules,” Journal of Medicinal Chemistry 59 (2016): 8712–8756, 10.1021/acs.jmedchem.6b00472. [DOI] [PubMed] [Google Scholar]
- 83. Taylor R. D., MacCoss M., and Lawson A. D. G., “Rings in Drugs,” Journal of Medicinal Chemistry 57 (2014): 5845–5859, 10.1021/jm4017625. [DOI] [PubMed] [Google Scholar]
- 84. Ling Y., Hao Z.‐Y., Liang D., Zhang C.‐L., Liu Y.‐F., and Wang Y., “The Expanding Role of Pyridine and Dihydropyridine Scaffolds in Drug Design,“ Drug Design, Development and Therapy 15 (2021): 4289–4338, 10.2147/DDDT.S329547. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85. De S., A. K. S. K., Kumar Shah S., et al., “Pyridine: The Scaffolds with Significant Clinical Diversity,” RSC Advances 12 (2022): 15385–15406, 10.1039/D2RA01571D. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86. Alam M. A., “Pyrazole: an Emerging Privileged Scaffold in Drug Discovery,” Future Medicinal Chemistry 15 (2023): 2011–2023, 10.4155/fmc-2023-0207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87. Sarwar T., Mustafa G., Zafar W., Hassan A. U., Sumrra S. H., and Asif A., “Pyrazoles a Privileged Scaffold in Drug Discovery – Synthetic Strategies & Exploration of Pharmacological Potential,” Journal of Molecular Structure 1348 (2025): 143523, 10.1016/j.molstruc.2025.143523. [DOI] [Google Scholar]
- 88. Alam M. A., “Chapter 1 ‐ Thiazole, a Privileged Scaffold in Drug Discovery,” in Privileged Scaffolds in Drug Discovery, eds. Yu B., Li N. and Fu C. (Academic Press, 2023), 1–19. [Google Scholar]
- 89. Kourounakis A. P., Xanthopoulos D., and Tzara A., “Morpholine as a Privileged Structure: A Review on the Medicinal Chemistry and Pharmacological Activity of Morpholine Containing Bioactive Molecules,” Medicinal Research Reviews 40 (2020): 709–752, 10.1002/med.21634. [DOI] [PubMed] [Google Scholar]
- 90. Kant R. and Maji S., “Recent Advances in the Synthesis of Piperazine Based Ligands and Metal Complexes and Their Applications,” Dalton Transactions 50 (2021): 785–800, 10.1039/D0DT03569F. [DOI] [PubMed] [Google Scholar]
- 91. Sun H., Tawa G., and Wallqvist A., “Classification of Scaffold Hopping Approaches,“ Drug Discovery Today 17 (2012): 310–324, 10.1016/j.drudis.2011.10.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92. Ertl P., Altmann E., Racine S., and Lewis R., “Ring Replacement Recommender: Ring Modifications for Improving Biological Activity,” European Journal of Medicinal Chemistry 238 (2022): 114483, 10.1016/j.ejmech.2022.114483. [DOI] [PubMed] [Google Scholar]
- 93. Chen L.‐Y. and Li Y.‐P., “AutoTemplate: Enhancing Chemical Reaction Datasets for Machine Learning Applications in Organic Chemistry,” Journal of Cheminformatics 16 (2024): 74, 10.1186/s13321-024-00869-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94. Kerstjens A. and De Winter H., “LEADD: Lamarckian Evolutionary Algorithm for De Novo Drug Design,” Journal of Cheminformatics 14 (2022): 3, 10.1186/s13321-022-00582-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95. Alakhdar A., Poczos B., and Washburn N., “Diffusion Models in De Novo Drug Design,” Journal of Chemical Information and Modeling 64 (2024): 7238–7256, 10.1021/acs.jcim.4c01107. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Material
Data Availability Statement
A description of the chemical catalogues used to extract all chemical transformations is provided in the Supplementary Information. ChemBang is a proprietary software developed by Chemotargets S.L. and it is not publicly available. Access to the code and implementation details is therefore restricted to collaborative work.
