Abstract
Designing target-relevant molecules is particularly difficult when only limited bioactivity data are available. We present MMP2Mol, a ligand-based generative framework that combines matched molecular pair (MMP) analysis with a pretrained chemical language model (CLM). MMP2Mol extracts target-specific structure–activity transformations from available ligands, applies the prioritized transformations to construct a virtual focused library, and fine-tunes the CLM toward target-relevant chemical space. The framework was evaluated across ten therapeutic targets and compared with direct CLM fine-tuning, Seq2Seq, and Reinvent 4. In repeated experiments on F2, BRD4, and PARP1, MMP2Mol achieved 92.8%–95.3% validity and 99.5%–99.7% uniqueness. When both methods were evaluated against the same original active-compound reference, molecular novelty increased from 57.8%–64.0% for the CLM baseline to 79.8%–82.0% for MMP2Mol, while internal diversity remained broadly comparable. The gains were most evident for F2 and BRD4, whereas performance varied across the smaller target datasets. These findings indicate that MMP-derived chemical knowledge can improve the focus and reproducibility of ligand-based molecular generation under limited-data conditions. MMP2Mol therefore provides a practical strategy for computational candidate generation and prioritization in early-stage drug discovery.
Keywords: matched molecular pair (MMP), ligand-based drug design, chemical language models (CLMs), data-scarce drug discovery, structure–activity relationship (SAR)
Introduction
Drug development aims to identify compounds with favourable bioactivity, safety, and pharmacokinetics [1–3]. However, the drug-like chemical space is tiny compared to the whole chemical universe [4], and traditional drug discovery methods are slow, costly, and often unpredictable [5]. Computational approaches have therefore become indispensable complements to experimental screening [6]. More recently, artificial intelligence (AI), including machine learning and deep learning, has accelerated target identification, molecular design, and property prediction, giving rise to AI-driven drug discovery (AIDD) [7–11]. AIDD can improve efficiency and reduce costs while supporting structural optimization, efficacy prediction, and toxicity assessment [12–14].
Among AI methods, deep generative models have become important tools for de novo drug design. They generate virtual libraries from one-dimensional SMILES strings [4, 15–17], two-dimensional molecular graphs [18–20], or three-dimensional structures [21–24], often relying on large ligand-receptor bioactivity datasets. Such data are often limited for transmembrane or highly flexible proteins, protein-nucleic-acid complexes, and glycoproteins, which can reduce model reliability. Chemical language models (CLMs) can produce target-focused libraries [16, 25–28], but small fine-tuning sets may inadequately define the relevant chemical space. Drug-likeness and synthetic feasibility also remain important constraints [29].
Data scarcity provides a strong rationale for integrating explicit chemical knowledge with data-driven molecular generation. Matched molecular pair (MMP) analysis, introduced in 2005 [30], compares molecule pairs that differ by a single well-defined structural change and extracts local structure–activity relationship (SAR) information [31–34]. MMP analysis can provide interpretable local quantitative structure–activity relationship (QSAR) guidance [35] and has been applied to ADMET (absorption, distribution, metabolism, excretion, and toxicity) optimization, bioisosteric replacement, and activity-cliff analysis [36–42]. Recent generative approaches have incorporated MMP-derived knowledge more directly. He et al. [43] framed molecular optimization as a translation task and trained Seq2Seq and Transformer models on MMPs to generate optimized analogues conditioned on a starting molecule and desired property changes. More recently, Pan et al. [44] introduced a variable-to-variable foundation model for MMP transformations, together with a retrieval-augmented framework that uses reference analogues to improve transformation diversity and controllability.
MMP2Mol differs from these approaches in how MMP knowledge is incorporated into molecular generation. It does not formulate generation as direct translation from a starting molecule to an optimized product, nor does it retrieve or generate individual transformations during inference. Instead, MMP2Mol extracts target-specific transformations from curated ligand data, prioritizes them using observed activity changes, rule support, and physicochemical-property criteria, applies retained transformations to known active compounds, and computationally prioritizes the resulting virtual focused library. The selected full molecules are then used to fine-tune a pretrained CLM for stochastic de novo generation. Thus, MMP knowledge is used primarily for target-specific data augmentation and distribution shaping rather than as an inference-time edit generator. The rule-ranking stage is treated as a prioritization procedure rather than confirmatory inference because the original rule-level P-values were not adjusted for the large number of candidate transformations.
This staged workflow preserves the interpretability of local MMP transformations while allowing the fine-tuned CLM to generate complete molecules beyond direct one-step analogues. Building on prior work in low-data generative molecular design [45], we therefore evaluate whether target-specific MMP-guided augmentation improves reproducible generation validity, structural novelty, chemical-space focus, and computational prioritization relative to direct CLM fine-tuning, and under which target and data conditions those gains remain credible.
Results and discussion
Overview of MMP2Mol
MMP2Mol comprises four stages (Fig. 1): target-ligand curation and activity modelling; MMP transformation prioritization by activity change, positive-pair fraction, support, quantitative estimate of drug-likeness (QED) [46], and synthetic accessibility score (SA) [47]; application of retained rules to known actives followed by computational prioritization of the virtual focused library; and fine-tuning and stochastic sampling of a pretrained CLM to generate complete molecules.
Figure 1.

Overall MMP2Mol architecture. (a) Curation of target-specific active and inactive compounds and construction of activity-prediction QSAR models. A scaffold-disjoint development/locked-test workflow was used to construct Locked-QSAR models for retrospective reliability assessment, whereas the selected descriptor-classifier combination was refitted on the complete labelled dataset to construct a Full-data QSAR model for application-stage screening. (b) MMPA-driven rule extraction, filtering, and virtual focused library generation. (c) Molecules retained by QSAR screening were used to fine-tune a ChEMBL 24-pretrained CLM for target-specific molecular generation.
To assess QSAR reliability [48], standardized and deduplicated compounds were divided by non-overlapping Bemis–Murcko scaffolds into a development set (~80%) and a locked test set (~20%). Four prespecified molecular representations were evaluated: 167-bit Molecular ACCess System (MACCS) keys, 2048-bit extended connectivity fingerprint 6 (ECFP6), 2048-bit functional connectivity fingerprint with diameter 6 (FCFP6) [49–51], and a concatenated representation comprising MACCS, ECFP6, and FCFP6. Each representation was combined with six classifiers: random forest (RF) [52], extreme gradient boosting (XGBoost) [53], support-vector machine (SVM) [54, 55], Naive Bayes (NB) [56], logistic regression [57], and multilayer perceptron (MLP) [58], giving 24 candidate descriptor-classifier combinations per target. Model selection was performed exclusively within the development set using five-fold stratified scaffold-grouped cross-validation (CV) [59]. The selected combination was refitted on the complete development set and evaluated on the locked test set using a fixed probability threshold of 0.5. Development-set results for all 240 combinations are provided in Table S1 and Fig. S1; complete locked-test results are provided in Table S3 and Fig. S2. Full-data QSAR model-selection results are provided in Table S2 and Fig. S3.
To assess QSAR generalization independently of application-stage screening, we established target-specific scaffold-disjoint evaluation models, hereafter termed locked-QSAR. Models were selected by five-fold scaffold-grouped CV within development partitions and evaluated once on untouched locked test sets (Table 1). Performance was strongest for PDE5A and PARP1 and remained useful for AR, BRD4, and F2. Estimates for smaller targets require caution; HMGCR’s ROC-AUC of 1.000 was based on 36 compounds, whereas PKCeta showed no discrimination (ROC-AUC = 0.479; MCC = −0.138). Locked-QSAR models were subsequently used for reduced-data and ablation analyses, while separate full-data QSAR models supported application-stage screening.
Table 1.
Target-specific locked-QSAR model selection and performance on the scaffold-disjoint locked test set.
| a Target | Selected model | Development CV ROC-AUC | Locked-test n (active %) | Locked-test ROC-AUC (95% CI) | PR -AUC | MCC | AD coverage |
|---|---|---|---|---|---|---|---|
| AR | Combined- XGBoost | 0.822 ± 0.036 | 444 (61.5%) | 0.849 (0.797–0.889) | 0.885 | 0.575 | 73.6% |
| BRD4 | Combined-RF | 0.914 ± 0.017 | 871 (55.6%) | 0.888 (0.842–0.928) | 0.902 | 0.644 | 85.0% |
| F2 | Combined-RF | 0.884 ± 0.026 | 951 (54.3%) | 0.868 (0.825–0.905) | 0.888 | 0.581 | 74.7% |
| PARP1 | Combined-RF | 0.916 ± 0.035 | 644 (80.0%) | 0.905 (0.855–0.945) | 0.974 | 0.596 | 81.1% |
| PDE5A | ECFP6- SVM | 0.918 ± 0.032 | 383 (65.8%) | 0.931 (0.888–0.962) | 0.942 | 0.704 | 90.6% |
| ERR1 | ECFP6-NB | 0.870 ± 0.130 | 29 (65.5%) | 0.989 (0.899–1.000) | 0.995 | 0.847 | 82.8% |
| HMGCR | FCFP6- SVM | 0.804 ± 0.076 | 36 (66.7%) | 1.000 (1.000–1.000) | 1.000 | 0.938 | 91.7% |
| KAT6A | FCFP6-RF | 0.875 ± 0.088 | 39 (59.0%) | 0.921 (0.878–1.000) | 0.957 | 0.624 | 64.1% |
| PKCeta | ECFP6- XGBoost | 0.642 ± 0.105 | 44 (84.1%) | 0.479 (0.163–0.735) | 0.841 | −0.138 | 93.2% |
| VDR | Combined-RF | 0.908 ± 0.101 | 75 (46.7%) | 0.969 (0.888–1.000) | 0.950 | 0.948 | 84.0% |
aCombined denotes concatenated MACCS, ECFP6, and FCFP6 descriptors. Development values are mean ± SD from five-fold scaffold-grouped CV. Confidence intervals were obtained from 1000 scaffold-grouped bootstrap replicates. AD denotes applicability-domain coverage. Complete confidence intervals, balanced accuracy, F1 score, Brier score, sensitivity, specificity, and in-domain/out-of-domain performance are reported in Table S3.
For application-stage screening, all 24 prespecified descriptor-classifier combinations were evaluated for each target using five-fold stratified scaffold-grouped CV. The best-performing combination was subsequently refitted on the complete curated labeled dataset to establish the target-specific full-data QSAR model. Generated molecules with predicted probabilities ≥0.5 were classified as active. These full-data QSAR models were used for virtual-library and generated-molecule screening. The five-fold CV ROC curves and complete model-selection results are presented in Fig. S3 and Table S2, respectively.
Performance of MMP2Mol in bioactive-molecule generation
We evaluated MMP2Mol on ten therapeutic targets from ChEMBL, divided into data-scarce (<500 active compounds: ERR1 [60, 61], HMGCR [62, 63], KAT6A [64], VDR [65], PKCeta [66]) and data-dense (>1000 active compounds: PARP1 [67–69], BRD4 [70], AR [71], PDE5A [72, 73], F2 [74, 75]). The original application-stage analysis used 10 000 generated molecules per target for docking evaluation. Independent stochastic repetitions were subsequently performed for all four generators using fixed method-specific configurations. For the principal repeated full-data QSAR, QED, and SA analyses, CLM and MMP2Mol outputs generated after epochs 1–30 were pooled within each independent run, standardized, and deduplicated by canonical SMILES; each pooled run was treated as one statistical replicate. Reinvent 4 and the Seq2Seq baseline, implemented using the Mol2Mol module in Reinvent 4, were evaluated using independently sampled run-level libraries under fixed official configurations. In contrast, the separate scale-matched CLM-MMP2Mol comparison used epoch-30 outputs only to evaluate validity, uniqueness, molecular novelty relative to the original target-specific active set, ECFP4 internal diversity, and Bemis–Murcko scaffold novelty (Table S4).
Across the five data-dense targets, MMP2Mol increased mean validity by 1.8–7.3 percentage points and molecular novelty by 15.8–22.6 percentage points relative to CLM, while maintaining 98.6%–99.7% uniqueness and broadly comparable internal diversity (Table S4). Its mean full-data QSAR-predicted active rate exceeded those of all three baselines for every target, ranging from 77.2% to 93.6% in the data-dense tasks (Table 2) and from 51.8% to 89.9% in the data-scarce tasks (Table 3). QED, SA, and docking performance remained competitive. The predicted-active rate denotes the proportion of valid, unique molecules assigned a probability ≥0.5 by the full-data QSAR model; model-selection results are provided in Table S2 and Fig. S3, with detailed data-scarce results in Table S5.
Table 2.
Performance on data-dense targets.
| a Target | Method | Best docking score↓ (kcal/mol) | Docking score, mean ± SD↓ (kcal/mol) | Full-data QSAR-predicted active, mean (95% CI)↑ | QED, mean (95% CI)↑ | SA, mean (95% CI)↓ |
|---|---|---|---|---|---|---|
| F2 | MMP2Mol | −15.77 | −12.75 ± 0.80 | 77.2% (75.4–78.4) | 0.39 (0.39–0.40) | 3.16 (3.15–3.17) |
| CLM | −15.76 | −12.84 ± 0.88 | 65.2% (64.3–66.5) | 0.38 (0.38–0.39) | 3.10 (3.09–3.10) | |
| Seq2Seq | −13.02 | −8.04 ± 1.01 | 72.6% (71.1–74.1) | 0.32 (0.32–0.33) | 3.35 (3.33–3.37) | |
| Reinvent 4 | −13.22 | −10.94 ± 0.58 | 60.8% (59.4–62.2) | 0.41 (0.40–0.42) | 3.08 (3.04–3.12) | |
| AR | MMP2Mol | −12.63 | −10.91 ± 0.31 | 86.3% (83.3–88.4) | 0.66 (0.66–0.67) | 3.19 (3.18–3.21) |
| CLM | −12.15 | −10.82 ± 0.30 | 78.3% (77.0–79.6) | 0.67 (0.67–0.68) | 3.24 (3.24–3.24) | |
| Seq2Seq | −12.25 | −9.39 ± 0.60 | 68.4% (66.8–70.0) | 0.72 (0.71–0.72) | 3.87 (3.84–3.91) | |
| Reinvent 4 | −12.83 | −10.73 ± 0.29 | 73.1% (71.0–75.1) | 0.70 (0.69–0.71) | 3.42 (3.40–3.44) | |
| BRD4 | MMP2Mol | −9.78 | −7.68 ± 0.40 | 82.7% (79.7–84.4) | 0.50 (0.50–0.51) | 3.10 (3.09–3.11) |
| CLM | −9.71 | −7.53 ± 0.43 | 68.5% (68.0–68.9) | 0.53 (0.52–0.53) | 3.05 (3.04–3.06) | |
| Seq2Seq | −10.37 | −6.54 ± 1.74 | 63.7% (62.1–65.4) | 0.50 (0.49–0.5) | 3.33 (3.30–3.36) | |
| Reinvent 4 | −10.44 | −8.07 ± 0.60 | 53.5% (52.2–54.8) | 0.54 (0.53–0.55) | 2.98 (2.95–3.02) | |
| PARP1 | MMP2Mol | −15.17 | −12.74 ± 0.62 | 93.6% (92.4–94.4) | 0.66 (0.65–0.67) | 2.86 (2.85–2.88) |
| CLM | −15.05 | −12.67 ± 0.57 | 90.1% (89.9–90.5) | 0.64 (0.64–0.65) | 2.72 (2.71–2.73) | |
| Seq2Seq | −14.17 | −10.42 ± 1.17 | 80.1% (79.4–80.7) | 0.57 (0.56–0.58) | 2.93 (2.90–2.96) | |
| Reinvent 4 | −14.27 | −9.57 ± 1.40 | 64.2% (63.6–64.8) | 0.64 (0.63–0.64) | 2.70 (2.66–2.74) | |
| PDE5A | MMP2Mol | −14.26 | −12.38 ± 0.47 | 90.2% (80.3–96.0) | 0.53 (0.52–0.53) | 3.01 (2.99–3.02) |
| CLM | −14.33 | −12.09 ± 0.58 | 73.3% (69.0–76.1) | 0.51 (0.50–0.52) | 3.00 (2.99–3.01) | |
| Seq2Seq | −13.04 | −8.57 ± 1.38 | 77.2% (74.2–78.9) | 0.48 (0.47–0.48) | 3.22 (3.20–3.40) | |
| Reinvent 4 | −10.38 | −7.11 ± 0.66 | 60.3% (58.1–62.5) | 0.55 (0.54–0.56) | 2.88 (2.83–2.92) |
aFull-data QSAR-predicted active rates, QED, and SA are reported as run-level means with 95% confidence intervals. For CLM and MMP2Mol, outputs generated after each of epochs 1–30 were pooled within the corresponding independent run before standardization and canonical-SMILES deduplication, and each run was weighted equally irrespective of the resulting library size. Docking statistics were obtained from the original application-stage libraries and represent molecule-level summaries rather than between-run confidence intervals.
Table 3.
Comparison of MMP2Mol and other methods on data-scarce targets.
| a Target | Full-data QSAR-predicted active, mean (95% CI)↑ | Docking Score↓ (kcal/mol) | ||||||
|---|---|---|---|---|---|---|---|---|
| MMP2Mol | CLM | Seq2Seq | Reinvent 4 | MMP2Mol | CLM | Seq2Seq | Reinvent 4 | |
| VDR | 69.3% (53.6–72.7) | 24.9% (23.7–25.8) | 43.3% (42.2–44.5) | 45.2% (44.9–45.5) | −10.56 | −10.42 | −10.25 | −9.86 |
| HMGCR | 63.0% (58.1–66.1) | 33.0% (30.8–34.5) | 41.7% (40.1–42.9) | 48.1% (47.3–48.9) | −10.91 | −10.08 | −7.59 | −7.58 |
| ERR1 | 51.8% (44.7–55.4) | 16.3% (15.8–16.7) | 37.2% (36.6–37.8) | 43.8% (42.2–44.5) | −8.89 | −8.50 | −7.75 | −8.40 |
| PKCeta | 89.9% (89.1–90.5) | 87.4% (86.9–88.5) | 82.1% (80.9–83.3) | 60.6% (59.8–61.4) | −9.22 | −9.41 | −9.06 | −7.96 |
| KAT6A | 71.7% (65.9–75.2) | 57.3% (56.8–57.5) | 65.2% (63.8–66.3) | 52.1% (50.8–53.5) | −10.27 | −9.36 | −8.13 | −8.80 |
aFull-data QSAR-predicted active rates are reported as run-level means with 95% confidence intervals. For CLM and MMP2Mol, outputs generated after epochs 1–30 were pooled within each independent run before standardization and canonical-SMILES deduplication. Reinvent 4 and Seq2Seq were evaluated using independently sampled run-level libraries. Docking statistics are descriptive molecule-level summaries from the original application-stage libraries.
We further compared the docking-score distributions of prioritized MMP2Mol-generated molecules with those of known actives (Fig. 2a). t-Distributed stochastic neighbour embedding (t-SNE) [76] analysis of equal-size MMP2Mol and CLM samples (Fig. 2b) revealed partially overlapping yet complementary chemical spaces, consistent with MMP-guided enrichment of target-relevant chemotypes beyond direct CLM fine-tuning.
Figure 2.

Docking scores and chemical space distribution of MMP2Mol-generated molecules. (a) Docking-score distributions of the 2000 valid and unique MMP2Mol-generated molecules with the most favourable AutoDock Vina scores for each target. For each target, the known bioactive set comprises the top 500 compounds with the highest pChEMBL values, or all available compounds when fewer than 500 were present. The boxes represent the interquartile range, the central lines indicate the medians, and the whiskers extend to 1.5 times the interquartile range. (b) Chemical-space comparison of MMP2Mol- and CLM-generated molecules using equal-size samples of 5000 valid and unique molecules from each method for each target. Molecules were sampled proportionally across ten ECFP6-based structural clusters to reduce overplotting while preserving chemical-space coverage and were visualized using t-SNE (perplexity = 30).
Overall, MMP2Mol balanced validity, novelty, computational activity enrichment, and drug-like properties across data-dense and data-scarce tasks. For some small datasets, stronger focus reduced internal or scaffold diversity and increased QSAR uncertainty. Accordingly, QSAR classifications and docking scores are interpreted as computational prioritization endpoints.
Performance under reduced-data scenarios
To examine generation under reduced target-specific data conditions, we constructed case studies for F2, AR, PARP1, and BRD4 using 10% subsets selected by ECFP6 structure-cluster-stratified sampling. Only these selected subsets were used for MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning. For these targets, canonical-SMILES matches to the complete target-specific sets were additionally removed from the ChEMBL 24 pretraining corpus before pretraining the corresponding CLMs. Generated molecules were evaluated using the locked-QSAR and target-specific re-scoring models as in the main application workflow. Thus, the reduced-data constraint applies to target-specific rule extraction and generator adaptation, whereas the reported classifier outputs are practical post-generation screening endpoints. The re-scoring models used AutoDock Vina-derived residue-interaction descriptors within 5 Å of the docked poses and random-forest classifiers at a probability threshold of 0.5; their CV performance is reported in Figs S4–S5 and Tables S7–S8. As shown in Table 4, MMP2Mol achieved higher locked-QSAR-predicted active rates than CLM for F2 (39.5% versus 24.2%), AR (60.9% versus 47.2%), PARP1 (85.6% versus 76.2%), and BRD4 (51.5% versus 17.2%). The comparatively high PARP1 rates are consistent with target-specific enrichment relative to the locked evaluator’s decision boundary. Because Table 2 used full-data QSAR models, absolute values between the two settings are not directly comparable. The docking distributions in Fig. 3 further showed a smaller change for MMP2Mol than for CLM after target-specific data reduction. Detailed AD coverage, out-of-domain fractions, domain-restricted predicted-active rates, and 95% confidence intervals are shown (Table S6).
Table 4.
Performance of MMP2Mol in the data-reduced scenarios.
| Target | Compounds | Applicable rules | Locked-QSAR-predicted active rate (%)↑ | Re-scoring-predicted active rate (%)↑ | ||||
|---|---|---|---|---|---|---|---|---|
| 100% | 10% | 100% | 10% | MMP2Mol | CLM | MMP2Mol | CLM | |
| F2 | 4822 | 473 | 6659 | 78 | 39.5 | 24.2 | 61.6 | 54.9 |
| AR | 3179 | 216 | 6813 | 61 | 60.9 | 47.2 | 41.8 | 31.4 |
| PARP1 | 3223 | 318 | 21,235 | 173 | 85.6 | 76.2 | 89.4 | 82.4 |
| BRD4 | 4359 | 433 | 11,848 | 154 | 51.5 | 17.2 | 59.6 | 37.6 |
aThe 10% subsets were obtained by proportional sampling across ECFP6 structural clusters and were fixed before model training. Compounds outside the subsets were excluded from MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning. QSAR-predicted active rates were calculated after generation using the locked-QSAR at a probability threshold of 0.5; re-scoring-predicted active rates were calculated using the target-specific re-scoring models at the same threshold.
Figure 3.

Docking-score distributions under complete-data and 10% target-specific data settings. (a) Distributions for MMP2Mol and CLM using the complete target-specific datasets. (b) Corresponding distributions using the 10% cluster-stratified subsets. (c) Comparison of CLM-generated F2 molecules under the complete-data and 10% settings. (d) Corresponding comparison for MMP2Mol.
Before CLM fine-tuning, we evaluated the MMP-derived virtual libraries (Table 5). From 181 HMGCR compounds, 205 rules generated 5783 molecules, of which 5579 (96.5%) were full-data QSAR predicted active. Across five targets, the libraries preserved useful QED and SA distributions (Fig. 4), while subsequent CLM fine-tuning added structures beyond direct transformations. In the size-matched comparison (Table 6), MMP2Mol exceeded CLM in QSAR- and re-scoring-predicted active rates for all targets; corrected PKCeta values were 92.5% versus 91.7% and 80.4% versus 71.0%, respectively. SA was comparable, and QED was similar or higher for MMP2Mol in four targets. Together, these analyses support reusable MMP-guided expansion under limited data; activity assignments remain computational prioritization endpoints.
Table 5.
Performance of the MMP-derived virtual libraries within the MMP2Mol framework.
| a Target | Compounds | Transformations | Applicable rules | MMPA- derived molecules | Full-data QSAR- predicted active molecules | Avg QED ↑ | Avg SA↓ |
|---|---|---|---|---|---|---|---|
| VDR | 394 | 35,417 | 1289 | 1527 | 1412 | 0.3898 (0.3649) | 4.291 (4.207) |
| HMGCR | 181 | 5825 | 205 | 5783 | 5579 | 0.4677 (0.3675) | 3.758 (3.784) |
| ERR1 | 165 | 8811 | 404 | 1579 | 1546 | 0.4290 (0.3847) | 2.947 (3.171) |
| PKCeta | 225 | 10,890 | 424 | 8056 | 7931 | 0.4675 (0.3720) | 3.132 (3.521) |
| KAT6A | 246 | 52,471 | 2412 | 23,796 | 21,383 | 0.7022 (0.6783) | 2.245 (2.221) |
aValues in parentheses in the QED and SA columns are the corresponding means for the original target-specific active-compound set.
Figure 4.

Molecular-property distributions of the virtual libraries constructed by MMP2Mol and the corresponding known bioactive compounds. Kernel-density estimates show the distributions of (a) QED and (b) SA across five targets. Higher QED values indicate more favourable drug-like properties, whereas lower SA scores indicate greater predicted synthetic accessibility. The dashed lines indicate the predefined SA threshold of 6.0.
Table 6.
Scale-matched comparison of MMP2Mol and CLM across data-scarce targets.
| Target | Full-data QSAR-predicted active rate↑ | Re-scoring-predicted active rate↑ | Avg SA↓ | Avg QED↑ | ||||
|---|---|---|---|---|---|---|---|---|
| MMP2Mol | CLM | MMP2Mol | CLM | MMP2Mol | CLM | MMP2Mol | CLM | |
| VDR | 69.3% | 24.9% | 41.1% | 25.5% | 3.95 | 3.23 | 0.43 | 0.42 |
| HMGCR | 63.0% | 33.0% | 62.1% | 45.8% | 3.56 | 3.31 | 0.58 | 0.42 |
| ERR1 | 51.8% | 16.3% | 71.8% | 67.4% | 2.36 | 2.25 | 0.67 | 0.73 |
| PKCeta | 92.5% | 91.7% | 80.4% | 71.0% | 2.46 | 2.46 | 0.74 | 0.72 |
| KAT6A | 91.1% | 71.0% | 42.1% | 22.3% | 2.27 | 2.30 | 0.68 | 0.67 |
Application and comparative evaluation in ligand-based drug design
To characterize the structural proximity and novelty of the generated molecules, we calculated Morgan-fingerprint-based Tanimoto similarities between the valid and unique molecules generated by each method and the corresponding target-specific original active-compound set. Table 7 summarizes nearest-neighbour Morgan-fingerprint Tanimoto similarity between valid, unique generated molecules and original target actives (range, 0.186–0.671). MMP2Mol generally remained near known-active chemical space while achieving higher full-data QSAR-predicted active rates. For AR, the full-data QSAR-predicted active rates were 86.3%, 78.3%, 68.4%, and 73.1% for MMP2Mol, CLM, Seq2Seq, and Reinvent 4, respectively, whereas the corresponding mean nearest-reference similarities were 0.445, 0.367, 0.416, and 0.412. Thus, enrichment was not explained solely by similarity (Figs S6–S7).
Table 7.
Mean nearest-reference ECFP4 Tanimoto similarities of molecules generated by the four methods.
| Target | MMP2Mol | CLM | Reinvent 4 | Seq2Seq |
|---|---|---|---|---|
| F2 | 0.522 | 0.451 | 0.487 | 0.605 |
| AR | 0.445 | 0.367 | 0.412 | 0.416 |
| PARP1 | 0.549 | 0.517 | 0.533 | 0.644 |
| BRD4 | 0.392 | 0.446 | 0.465 | 0.603 |
| PDE5A | 0.544 | 0.464 | 0.444 | 0.671 |
| VDR | 0.463 | 0.310 | 0.346 | 0.671 |
| HMGCR | 0.373 | 0.269 | 0.267 | 0.600 |
| ERR1 | 0.344 | 0.186 | 0.227 | 0.578 |
| PKCeta | 0.330 | 0.264 | 0.251 | 0.550 |
| KAT6A | 0.518 | 0.356 | 0.413 | 0.597 |
In a retrospective PARP1 low-data case study, ECFP6 cluster-stratified sampling assigned 318 of the 3219 unique curated compounds to the reduced-data subset, leaving 2901 compounds as a reference holdout. Nearest-neighbour ECFP6 analysis yielded a median maximum Tanimoto similarity of 0.600 (interquartile range, 0.443–0.700); 25.1% of the holdout compounds had a maximum similarity of ≥0.7, only 0.5% had a similarity of ≥0.9, and 33.5% shared a Bemis–Murcko scaffold with the subset (Table S9). Under identical model and sampling settings, MMP2Mol and CLM each generated 40 000 molecules, yielding 37 368 and 36 741 valid and unique structures, respectively. Among molecules with nearest-reference Tanimoto similarity >0.9, 65.4% and 67.2%, respectively, met a QSAR probability ≥0.5 and docking score ≤ −8.0 kcal/mol; MMP2Mol recovered five exact holdout structures versus one for CLM (Fig. 5). At similarities of 0.2–0.5, joint pass rates were 24.8% and 13.3% (Fig. 6), indicating improved recovery and computational prioritization. Compounds 4 and 8 from Fig. 6, selected for complementary predicted activity, QED, SA, and novelty, were subjected to retrosynthetic analysis (Fig. S9).
Figure 5.

Exact reference-holdout recovery by MMP2Mol and the CLM baseline in the PARP1 reduced-data case study. Both models were trained on the same 10% target-specific subset (318 molecules) and each generated 40 000 molecules. MMP2Mol recovered five exact structures from the reference holdout, compared with one for CLM.
Figure 6.

Representative MMP2Mol-generated molecules with nearest-reference Tanimoto similarities of 0.2–0.5 in the PARP1 reduced-data experiment. QED and SA are computational estimates; the closest reference compounds are shown in Fig. S8.
To test whether diversity compromised beneficial MMP transformations, we measured rule-containing generated molecules. For F2, 86% incorporated at least two rule-defined substructure modifications, whereas 14% contributed scaffold-level diversity. Thus, diversity and rule adherence coexisted (Fig. S10), supporting target-focused exploration while preserving MMP-derived information.
Ablation analysis of MMP-derived augmentation and QSAR filtering
To examine the contributions of the main MMP2Mol components, a configuration-based ablation study was conducted on PARP1 and BRD4. Four configurations were compared: A0, direct CLM fine-tuning on the original active compounds; A1, fine-tuning on molecules generated using all retained MMP transformations, including activity-increasing and activity-decreasing directions; A2, fine-tuning on the original actives plus molecules generated using activity-increasing transformations, without final QSAR filtering; and A3, the complete MMP2Mol workflow using the QSAR-filtered expanded library. Each used three prespecified independent runs under identical settings; extra runs were excluded. Size-matched outputs were assessed for validity, uniqueness, novelty, internal and scaffold diversity, and active & AD under the scaffold-disjoint oracle, with novelty referenced to the same original active set.
The unfiltered all-rule variant A1 produced the broadest exploration for both targets (Table 8). A1 produced the greatest chemical-space expansion, with molecular novelty of 97.99% for PARP1 and 96.76% for BRD4, but its active & AD rates decreased to 62.07% and 44.87%, respectively. Restricting augmentation to activity-increasing transformations in A2 increased the corresponding active & AD rates to 74.48% and 74.13%. The complete A3 configuration provided the most favourable overall balance. For PARP1, A3 yielded 96.67% validity, 79.25% novelty, and 80.28% active & AD, compared with 92.00%, 57.71%, and 79.53% for A0. For BRD4, the corresponding A3 values were 95.33%, 82.01%, and 79.54%, compared with 91.67%, 58.57%, and 74.86% for A0. Uniqueness remained above 98% across all configurations, whereas the lower internal diversity of A3 relative to A2 indicated greater concentration within the model-supported region.
Table 8.
Configuration-based ablation results for PARP1 and BRD4. Values are reported as mean ± standard deviation across independent runs.
| a Target | Configuration | Validity (%) ↑ | Uniqueness (%) ↑ | Molecular novelty (%) ↑ | Internal diversity ↑ | Scaffold novelty (%) ↑ | Active & AD (%) ↑ |
|---|---|---|---|---|---|---|---|
| PARP1 | A0 | 92.00 ± 1.00 | 99.28 ± 1.26 | 57.71 ± 7.75 | 0.8126 ± 0.0180 | 35.74 ± 6.95 | 79.53 ± 4.59 |
| A1 | 84.67 ± 6.51 | 100.00 ± 0.00 | 97.99 ± 1.93 | 0.8349 ± 0.0027 | 80.60 ± 6.82 | 62.07 ± 4.12 | |
| A2 | 89.67 ± 3.06 | 98.88 ± 1.95 | 83.51 ± 3.20 | 0.8249 ± 0.0074 | 60.78 ± 3.76 | 74.48 ± 1.98 | |
| A3 | 96.67 ± 0.58 | 99.66 ± 0.60 | 79.25 ± 3.65 | 0.8048 ± 0.0062 | 58.80 ± 3.44 | 80.28 ± 5.21 | |
| BRD4 | A0 | 91.67 ± 2.52 | 98.55 ± 0.59 | 58.57 ± 8.04 | 0.8464 ± 0.0048 | 39.99 ± 6.53 | 74.86 ± 4.87 |
| A1 | 83.00 ± 3.46 | 100.00 ± 0.00 | 96.76 ± 1.48 | 0.8587 ± 0.0037 | 79.94 ± 2.14 | 44.87 ± 3.99 | |
| A2 | 86.33 ± 2.89 | 98.46 ± 0.64 | 75.69 ± 3.19 | 0.8525 ± 0.0022 | 53.49 ± 1.77 | 74.13 ± 2.78 | |
| A3 | 95.33 ± 0.58 | 99.30 ± 1.22 | 82.01 ± 4.09 | 0.8156 ± 0.0210 | 60.68 ± 9.85 | 79.54 ± 3.75 |
aValues are mean ± sample standard deviation across three independent runs per configuration. All molecular endpoints were calculated from a fixed sample of 100 raw SMILES generated from the epoch-30 checkpoint per independent run. Molecular and scaffold novelty were calculated against the same target-specific original active set. Active & AD represents the proportion simultaneously predicted active and located within the AD of the independent scaffold-disjoint oracle. Bold values indicate the highest result for each target and metric.
MMP augmentation therefore expanded chemical space, activity-increasing rule selection limited inconsistent expansion, and QSAR filtering enriched oracle-supported in-domain molecules. Because fine-tuning-set compositions differed, these are configuration-level rather than strictly additive module effects.
Discussion
MMP2Mol integrates MMP analysis with chemical language modelling to address limited target-specific bioactivity data in ligand-based molecular generation. Its principal contribution is the use of nominal Wilcoxon-prioritized [77], activity-enhancing molecular transformations as chemical priors. These transformations expand the target-focused fine-tuning space and guide the generative model toward structurally meaningful regions that may be difficult to learn directly from a limited number of active compounds. The repeated-generation analysis further showed that the improvements in molecular novelty were most consistent for data-dense targets, indicating that MMP-derived augmentation can enrich the information available to the language model without causing a general loss of generation quality.
The structural analyses provide additional insight into how MMP2Mol balances exploitation and exploration. Molecules generated by MMP2Mol generally remained within a moderate similarity range to known active compounds while achieving increased computational activity enrichment. Together with the molecular novelty, internal diversity, and Bemis–Murcko scaffold novelty results, this pattern is consistent with focused exploration around target-relevant chemical space rather than simple reproduction of the fine-tuning compounds. The reduced-data case studies further showed that MMP-derived augmentation can preserve useful target-specific structural information when only a fraction of the available compounds is used for model adaptation. In the PARP1 analysis, MMP2Mol recovered more structures held out from target-specific fine-tuning and provided greater enrichment than CLM among molecules with low-to-moderate similarity to the reference set.
The generators evaluated in this study were constructed using transformations prioritized by a nominal one-sided Wilcoxon criterion (P < .05), together with the prespecified effect-size and positive-pair filters. A complementary multiplicity-controlled audit using the BH-FDR procedure showed that, under the primary criterion (support ≥5, q ≤ 0.05, median ΔpChEMBL ≥0.30, and positive-pair fraction ≥0.50), no transformations were retained for AR, VDR, HMGCR, ERR1, or KAT6A, and only one was retained for PKCeta (Table S10). Thus, statistically robust rule support was weak or absent for several targets, indicating that the reliability of MMP-derived guidance depends on the quantity, diversity, and consistency of the available target-specific data. Settings with fewer than ten active compounds were not evaluated and may yield too few robust transformations. QSAR and re-scoring reliability also decreases with scarce inactives or out-of-domain molecules. Although MMP2Mol generation is fully ligand based, docking and interaction-based re-scoring require a reliable target structure. These analyses were used only for post-generation ranking, not as training signals; when structures are unavailable, scaffold-disjoint QSAR with applicability-domain assessment or ligand-based pharmacophore models can be used instead. Future work will emphasize uncertainty-aware prediction, domain-guided selection, temporal or prospective validation, and stronger rule-reliability estimates. Incorporating 3D pharmacophore or structure-based constraints and multi-objective optimization of potency, selectivity, ADMET, and synthesis may improve lead optimization, while alternative backbones, cross-target rule transfer, and active learning could extend MMP2Mol to more extreme low-data settings. To support experimental translation, candidates could be ranked by predicted activity, applicability-domain confidence, QED, SA, novelty, and structural alerts, diversity-selected for synthesis, and assessed in target-appropriate biochemical assays before dose–response and cellular confirmation.
Conclusion
We present MMP2Mol, a ligand-based framework that combines MMP-derived transformations with CLM for de novo molecular design. The method addresses limited target-specific data by extracting SAR information through a nominal Wilcoxon-based prioritization workflow, constructing focused virtual libraries, and fine-tuning a pretrained CLM. Across diverse targets, MMP2Mol improved predicted activity enrichment and molecular novelty relative to direct CLM fine-tuning while maintaining competitive drug-likeness and synthetic-accessibility profiles. Its modular workflow provides a practical strategy for early-stage computational candidate generation, particularly for ligand-defined targets with limited activity data.
Methods
Data collection and pretreatment
The ChEMBL 24 collection, containing more than 1.1 million molecular records, was used for pretraining the CLM. Target-specific datasets were restricted to Homo sapiens single-protein targets and comprised five data-scarce targets with fewer than 500 active compounds (ERR1, HMGCR, KAT6A, VDR, and PKCeta) and five data-dense targets with more than 1000 active compounds (PARP1, BRD4, AR, PDE5A, and F2). Molecules were processed with RDKit (https://www.rdkit.org) and represented as canonical SMILES. Salts were removed by retaining the largest organic fragment, invalid structures and records without usable activity values were discarded, and SMILES longer than 90 characters were excluded. Stereochemical annotations present in the source SMILES were preserved. Redundant records were collapsed to a single entry, and when duplicate records had discordant activity values, the record with the higher pChEMBL value was retained. Activity labels were assigned using a 500 nM threshold: records with reported activity ≤500 nM were labelled active and those with activity >500 nM were labelled inactive (equivalent to a pChEMBL boundary of ~6.30) [78, 79]. Both classes were retained for classifier development, whereas active compounds were used as the starting structures for MMP transformation and target-specific CLM fine-tuning.
Application-stage full-data QSAR models
All classifier hyperparameters were fixed a priori and were not optimized separately for individual targets. For each target, the descriptor-classifier combination with the best CV ROC-AUC in the original full-data search was refitted on all curated labelled compounds and used to prioritize valid and unique generated molecules at a probability threshold of 0.5. The full-data QSAR-predicted active rate was defined as the proportion assigned a probability of at least 0.5. Five-fold CV ROC curves for these application models are shown in Fig. S3.
Workflow of molecular generation
-
(1)
MMPA-guided focused molecular library generation: Rule databases were generated with mmpdb 3.1 using default parameters. For each directed transformation, paired activity differences were defined as ΔpChEMBL = pChEMBL(υB) − pChEMBL(υA) and evaluated using a one-sided Wilcoxon signed-rank test against a non-positive median change. The application-stage workflow used the original nominal Wilcoxon screening criterion (P < .05), followed by a median ΔpChEMBL ≥0.30 and a positive-pair fraction ≥0.50. These nominal Wilcoxon-prioritized rules were used to construct the virtual libraries and train the generators reported in this study. To assess the robustness of rule selection under multiple testing, the complete rule sets were additionally evaluated using Benjamini–Hochberg false-discovery-rate (BH-FDR) [80] correction within each support family, with a primary audit tier requiring support ≥5, q ≤ 0.05, median ΔpChEMBL ≥0.30, and positive-pair fraction ≥0.50; 95% bootstrap confidence intervals used 2000 resamples. The robustness-audit results are reported in Table S10. This analysis was used to characterize the statistical reliability of the extracted transformations, whereas generator training and the reported generation experiments were based on the nominal-Wilcoxon-prioritized rule sets described above.
-
(2)
Products generated by the retained MMP transformations were sanitized and represented as canonical SMILES using RDKit. Salts were reduced to the largest organic fragment, SMILES longer than 90 characters were discarded, and duplicate products were collapsed by canonical SMILES. Stereochemical annotations present in the source structures were preserved.
-
(3)
We pretrained the chemical language model (CLM) on more than 1.1 million molecules from ChEMBL 24. The model was implemented in Python (v 3.6.5) using Keras (v 2.2.0, https://keras.io/) with a TensorFlow GPU backend (v 1.9.0, https://www.tensorflow.org). It comprised two batch-normalization layers and two long short-term memory (LSTM) layers with 1024 and 256 units, respectively, totalling 5,820,515 parameters. One-hot-encoded SMILES were used as input, and the model was trained for 30 epochs using the Adam optimizer (learning rate = 10−3) and categorical cross-entropy loss. During target-specific fine-tuning, both LSTM layers remained trainable, and all models were trained for a prespecified 30 epochs with an initial learning rate of 1 × 10−4 [16, 45]. No early stopping was applied; instead, the learning rate was reduced by a factor of 0.5 after three epochs without improvement, to a minimum of 5 × 10−5. Epoch 30 was prespecified as the terminal training checkpoint and was not selected post hoc using generation performance. It was used for the convergence assessment and the epoch-30-specific scale-matched, training-set-overlap, medicinal-chemistry-alert, and ablation analyses. The principal repeated full-data QSAR, QED, and SA analyses instead used outputs pooled across epochs 1–30 within each independent run, as described below. Training and validation losses were recorded at each epoch. Across the 60 audited training histories, validation loss decreased between epochs 25 and 30 in every run; epoch 30 yielded the minimum validation loss in 53 runs, whereas the remaining seven reached comparable minima at epochs 28 or 29. Potential memorization was assessed using exact canonical-SMILES overlap and nearest-neighbour ECFP4 similarity to the corresponding fine-tuning library. Complete convergence and overlap diagnostics are provided in Table S11.
Molecules were sampled using a temperature of 1.0, top-k of 0, and top-p of 0.9. For the principal repeated analysis, SMILES strings generated from the checkpoints saved at epochs 1–30 were pooled within each independent CLM or MMP2Mol run, standardized, and deduplicated using canonical SMILES. Full-data QSAR-predicted active rates, QED values, and SA scores were calculated separately for each pooled run, with each run contributing equally to the summary statistics. By contrast, the separate scale-matched analyses of molecular quality, training-set overlap, and medicinal-chemistry alerts used 100 raw SMILES strings generated from the epoch-30 checkpoint per run. Reinvent 4 and Seq2Seq were evaluated using three independently sampled run-level libraries because these official pretrained implementations did not involve target-specific 30-epoch fine-tuning. Docking statistics were obtained from the original 10 000-molecule application-stage libraries and were treated as descriptive molecule-level summaries. Valid and canonical-unique molecules from the epoch-30 scale-matched samples were retrospectively screened for medicinal-chemistry alerts. Pan-assay interference compounds and Brenk alerts were identified using the corresponding RDKit FilterCatalog definitions [81, 82], supplemented with prespecified SMARTS patterns for potentially reactive functional groups and strained or unstable ring systems. Alerts were treated as advisory flags rather than definitive evidence of chemical invalidity. The alert definitions and target-specific results are reported in Table S12.
Baseline generative models
The CLM baseline used the same pretrained architecture and sampling settings as MMP2Mol but was fine-tuned on the corresponding original active-compound set without MMP-derived augmentation. The Reinvent 4 de novo generator and the Seq2Seq comparator, which was implemented using the conditional Mol2Mol module provided in the Reinvent 4 package (https://github.com/MolecularAI/REINVENT4), were run using the corresponding official pretrained models and software-default configurations. Only the target-specific input or generation mode and the required output size were changed, and no target-specific hyperparameter tuning was performed.
Optimization of transformation rules
Optimization of application-stage transformation rules: Candidate rules passing the nominal Wilcoxon screen were oriented toward increased activity and further prioritized using their median ΔpChEMBL, positive-pair fraction, support, and associated QED and SA changes. Rules with a median ΔpChEMBL <0.30 or a positive-pair fraction <0.50 were excluded. The retained transformations were converted to SMARTS, applied to compounds labelled active, and filtered for adverse physicochemical-property changes before construction of the focused library. The separate BH-FDR analysis described above was a retrospective sensitivity audit and did not alter these application-stage rule sets.
Reduced-data case studies and pretraining-overlap control
Reduced-data experiments were conducted for F2, AR, PARP1, and BRD4. ECFP6 fingerprints were used to construct structural clusters, and ~10% of the target compounds were selected by proportional sampling across clusters; the fixed subsets were reused for all method comparisons. MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning were restricted to these subsets. For each of the four targets (F2, AR, PARP1, and BRD4), molecules matching the complete corresponding target-specific dataset by canonical SMILES were removed from the ChEMBL 24 corpus before pretraining the CLMs used in the reduced-data experiments. This exclusion covered both the selected 10% subset and the remaining target-specific compounds. Generated molecules were evaluated using the corresponding scaffold-disjoint QSAR model selected within the development partition and fitted without the locked-test compounds. Each model was fixed before scoring the generated libraries and applied at a probability threshold of 0.5. The locked test partition was used only to report the reliability results in Table 1. Target-specific re-scoring models were retained as secondary post-generation screening endpoints.
Molecular docking and target-specific re-scoring
Docking was performed with AutoDock Vina (https://github.com/ccsb-scripps/AutoDock-Vina) using the following experimentally determined target structures: AR (PDB 2PNU), ERR1 (3K6P), BRD4 (4HXM), PKCeta (8FP1), F2 (1SL3), PARP1 (7KK4), HMGCR (2Q1L), KAT6A (8DD5), PDE5A (4XOZ), and VDR (5 V39). Receptor and ligand PDBQT files were prepared using standard AutoDockTools routines, including addition of polar hydrogen atoms and assignment of Gasteiger charges. A fixed target-specific binding pocket was used for all molecules and all generation methods for a given target. Vina exhaustiveness and number-of-modes settings were left at their software defaults. Docking was used only as a post-generation ranking procedure and was not included in CLM training.
For target-specific re-scoring, known active and inactive compounds were docked with the same protocol. Residue-level interaction terms within 5 Å of each docked ligand, together with the Vina docking output, were used as descriptors for random-forest classifiers. Activity labels used the same 500 nM threshold as the re-scoring QSAR models. Re-scoring models were evaluated by five-fold CV and applied at a probability threshold of 0.5. The re-scoring-predicted active rate was the proportion of successfully docked generated molecules assigned a probability of at least 0.5 by the corresponding target-specific re-scoring classifier.
Configuration-based ablation analysis
A configuration-based ablation study was performed for PARP1 and BRD4. A0 used direct CLM fine-tuning on the original active compounds. A1 used molecules produced using all retained MMP transformations, including activity-increasing and activity-decreasing directions. A2 used the original actives together with molecules produced by activity-increasing transformations but omitted final deployment-QSAR filtering. A3 represented the complete MMP2Mol workflow with the QSAR-filtered expanded library. All configurations used the same CLM architecture, 30-epoch fine-tuning schedule, checkpoint epoch, and sampling parameters. Three stochastic runs per target were included for every configuration; for A0 and A3, the first three prespecified run identifiers were used and any fourth completed run was excluded from the ablation summary. The evaluation was output-size matched at 100 raw epoch-30 sequences per run; larger outputs were deterministically subsampled using a fixed analysis seed (20260818) combined with the run identifier.
Ablation endpoints comprised validity, uniqueness, molecular novelty, ECFP4 internal diversity, Bemis–Murcko scaffold novelty, and the proportion of valid and unique molecules simultaneously assigned active and within the AD (active & AD) by the prespecified scaffold-based evaluation model. The activity threshold was 0.5, and the AD used an ECFP4 nearest-neighbour criterion whose target-specific threshold was the fifth percentile of leave-one-out nearest-neighbour similarities within the development partition. Values were summarized as mean and sample standard deviation across runs. Fine-tuning-library sizes and compositions were audited but were not forced to be equal; consequently, the analysis estimates configuration-level effects rather than strictly additive effects of isolated modules.
Evaluation
For the target-level structural comparison, ECFP4 fingerprints (radius 2, 2048 bits) were calculated for valid and unique generated molecules and the corresponding original active-compound set. The maximum Tanimoto similarity of each generated molecule to any reference active was determined, and Table 7 reports the mean of these nearest-reference similarities. Drug-likeness was assessed using QED, calculated with RDKit. QED ranges from 0 to 1, with higher values indicating a more favourable drug-like property profile. Synthetic accessibility was evaluated using the SA score, for which lower values indicate easier synthesis. Rule retention was assessed by matching the SMARTS patterns derived from the retained MMP transformations to generated structures.
For chemical-space visualization, equal-sized samples of 5000 valid and unique molecules from MMP2Mol and CLM were encoded as Morgan fingerprints and jointly embedded by t-SNE. For the docking-distribution visualization, all application experiments began from 10 000 generated molecules per target, and the 2000 molecules with the most favourable docking scores were displayed for each target; the known-compound comparison set comprised the 500 compounds with the highest pChEMBL values, or all available compounds when fewer than 500 were present. These subsets were used for visualization only and did not replace the full-library molecular-quality calculations.
Key Points
MMP2Mol integrates MMP analysis with chemical language models to enable ligand-based de novo drug design, effectively extracting and optimizing transformation rules from limited bioactivity data to guide generative models.
The framework constructs target-focused virtual libraries that improve predicted activity enrichment relative to direct CLM fine-tuning while maintaining competitive docking, drug-likeness, and synthetic-accessibility (SA) profiles.
MMP2Mol maintains substantial molecular diversity while generating structurally novel compounds with favourable computational estimates of synthetic accessibility and can recover known active molecules absent from target-specific fine-tuning.
Because generation does not require target protein structures, MMP2Mol is applicable across diverse ligand-defined targets and provides a practical framework for early-stage computational molecular design under data-scarce conditions.
Supplementary Material
Contributor Information
Dekun Chen, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.
Hui Li, Xiangya School of Pharmaceutical Sciences, Central South University, No. 172 Tongzipo Road, Yuelu District, Changsha, Hunan 410013, China.
Jianwang Liu, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.
Yueping Jiang, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.
Dongsheng Cao, Xiangya School of Pharmaceutical Sciences, Central South University, No. 172 Tongzipo Road, Yuelu District, Changsha, Hunan 410013, China.
Shao Liu, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.
Author contributions
Dekun Chen collected the data, developed the models, analyzed the data, and wrote the manuscript; Hui Li and Jianwang Liu helped interpret the results with constructive discussions; Yueping Jiang, Dongsheng Cao, and Shao Liu conceived and supervised the project, interpreted the results, and wrote the manuscript.
Conflicts of interest
There are no conflicts to declare.
Funding
None declared.
Data availability
The codes used for this research work are made publicly available in the GitHub repository [https://github.com/DKCCCChen/MMP2Mol].
References
- 1. Hughes JP, Rees S, Kalindjian SB et al. Principles of early drug discovery. Br J Pharmacol 2011;162:1239–49. 10.1111/j.1476-5381.2010.01127.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Xiong G, Wu Z, Yi J et al. ADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties. Nucleic Acids Res 2021;49:W5–w14. 10.1093/nar/gkab255 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Cumming JG, Davis AM, Muresan S et al. Chemical predictive modelling to improve compound quality. Nat Rev Drug Discov 2013;12:948–62. 10.1038/nrd4128 [DOI] [PubMed] [Google Scholar]
- 4. Gómez-Bombarelli R, Wei JN, Duvenaud D et al. Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent Sci 2018;4:268–76. 10.1021/acscentsci.7b00572 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Goodnow RA. Hit and lead identification: integrated technology-based approaches. Drug Discov Today Technol 2006;3:367–75. 10.1016/j.ddtec.2006.12.009 [DOI] [Google Scholar]
- 6. Neves BJ, Braga RC, Melo-Filho CC et al. QSAR-based virtual screening: advances and applications in drug discovery. Front Pharmacol 2018;9:1275. 10.3389/fphar.2018.01275 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Singh S, Kaur N, Gehlot A. Application of artificial intelligence in drug design: a review. Comput Biol Med 2024;179:108810. 10.1016/j.compbiomed.2024.108810 [DOI] [PubMed] [Google Scholar]
- 8. Gholap AD, Uddin MJ, Faiyazuddin M et al. Advances in artificial intelligence for drug delivery and development: a comprehensive review. Comput Biol Med 2024;178:108702. 10.1016/j.compbiomed.2024.108702 [DOI] [PubMed] [Google Scholar]
- 9. Yang X, Wang Y, Byrne R et al. Concepts of artificial intelligence for computer-assisted drug discovery. Chem Rev 2019;119:10520–94. 10.1021/acs.chemrev.8b00728 [DOI] [PubMed] [Google Scholar]
- 10. You Y, Lai X, Pan Y et al. Artificial intelligence in cancer target identification and drug discovery. Signal Transduct Target Ther 2022;7:156. 10.1038/s41392-022-00994-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Moingeon P, Kuenemann M, Guedj M. Artificial intelligence-enhanced drug design and development: toward a computational precision medicine. Drug Discov Today 2022;27:215–22. 10.1016/j.drudis.2021.09.006 [DOI] [PubMed] [Google Scholar]
- 12. Carini C, Seyhan AA. Tribulations and future opportunities for artificial intelligence in precision medicine. J Transl Med 2024;22:411. 10.1186/s12967-024-05067-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Vamathevan J, Clark D, Czodrowski P et al. Applications of machine learning in drug discovery and development. Nat Rev Drug Discov 2019;18:463–77. 10.1038/s41573-019-0024-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Dara S, Dhamercherla S, Jadav SS et al. Machine learning in drug discovery: a review. Artif Intell Rev 2022;55:1947–99. 10.1007/s10462-021-10058-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Lim J, Ryu S, Kim JW et al. Molecular generative model based on conditional variational autoencoder for de novo molecular design. J Cheminform 2018;10:31. 10.1186/s13321-018-0286-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Moret M, Pachon Angona I, Cotos L et al. Leveraging molecular structure and bioactivity with chemical language models for de novo drug design. Nat Commun 2023;14:114. 10.1038/s41467-022-35692-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Kong W, Hu Y, Zhang J et al. Application of SMILES-based molecular generative model in new drug design. Front Pharmacol 2022;13:1046524. 10.3389/fphar.2022.1046524 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Guo X, Zhao X, Lu X et al. A deep learning-driven discovery of berberine derivatives as novel antibacterial against multidrug-resistant helicobacter pylori. Signal Transduct Target Ther 2024;9:183. 10.1038/s41392-024-01895-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Li Y, Zhang L, Liu Z. Multi-objective de novo drug design with conditional graph generative model. J Cheminform 2018;10:33. 10.1186/s13321-018-0287-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Lim J, Hwang SY, Moon S et al. Scaffold-based molecular design with a graph generative model. Chem Sci 2019;11:1153–64. 10.1039/c9sc04503a [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Zhu H, Zhou R, Cao D et al. A pharmacophore-guided deep learning approach for bioactive molecular generation. Nat Commun 2023;14:6234. 10.1038/s41467-023-41454-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Imrie F, Hadfield TE, Bradley AR et al. Deep generative design with 3D pharmacophoric constraints. Chem Sci 2021;12:14577–89. 10.1039/D1SC02436A [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Li Y, Pei J, Lai L. Structure-based de novo drug design using 3D deep generative models. Chem Sci 2021;12:13664–75. 10.1039/D1SC04444C [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Bai Q, Xu T, Huang J et al. Geometric deep learning methods and applications in 3D structure-based drug design. Drug Discov Today 2024;29:104024. 10.1016/j.drudis.2024.104024 [DOI] [PubMed] [Google Scholar]
- 25. Wu K, Xia Y, Deng P et al. TamGen: drug design with target-aware molecule generation through a chemical language model. Nat Commun 2024;15:9360. 10.1038/s41467-024-53632-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Yue J, Peng B, Chen Y et al. Unlocking comprehensive molecular design across all scenarios with large language model and unordered chemical language. Chem Sci 2024;15:13727–40. 10.1039/D4SC03744H [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Isigkeit L, Hörmann T, Schallmayer E et al. Automated design of multi-target ligands by generative deep learning. Nat Commun 2024;15:7946. 10.1038/s41467-024-52060-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Grisoni F. Chemical language models for de novo drug design: challenges and opportunities. Curr Opin Struct Biol 2023;79:102527. 10.1016/j.sbi.2023.102527 [DOI] [PubMed] [Google Scholar]
- 29. Gangwal A, Ansari A, Ahmad I et al. Current strategies to address data scarcity in artificial intelligence-based drug discovery: a comprehensive review. Comput Biol Med 2024;179:108734. 10.1016/j.compbiomed.2024.108734 [DOI] [PubMed] [Google Scholar]
- 30. Kenny PW, Sadowski J. Structure modification in chemical databases. Chemoinform Drug Discov 2005;23:271–85. 10.1002/3527603743.ch11 [DOI] [Google Scholar]
- 31. Leach AG, Jones HD, Cosgrove DA et al. Matched molecular pairs as a guide in the optimization of pharmaceutical properties; a study of aqueous solubility, plasma protein binding and oral exposure. J Med Chem 2006;49:6672–82. 10.1021/jm0605233 [DOI] [PubMed] [Google Scholar]
- 32. Sheridan RP, Hunt P, Culberson JC. Molecular transformations as a way of finding and exploiting consistent local QSAR. J Chem Inf Model 2006;46:180–92. 10.1021/ci0503208 [DOI] [PubMed] [Google Scholar]
- 33. Griffen E, Leach AG, Robb GR et al. Matched molecular pairs as a medicinal chemistry tool. J Med Chem 2011;54:7739–50. 10.1021/jm200452d [DOI] [PubMed] [Google Scholar]
- 34. Yang Z, Shi S, Fu L et al. Matched molecular pair analysis in drug discovery: methods and recent applications. J Med Chem 2023;66:4361–77. 10.1021/acs.jmedchem.2c01787 [DOI] [PubMed] [Google Scholar]
- 35. Griffen EJ, Dossetter AG, Leach AG et al. Can we accelerate medicinal chemistry by augmenting the chemist with big data and artificial intelligence? Drug Discov Today 2018;23:1373–84. 10.1016/j.drudis.2018.03.011 [DOI] [PubMed] [Google Scholar]
- 36. Kramer C, Fuchs JE, Whitebread S et al. Matched molecular pair analysis: significance and the impact of experimental uncertainty. J Med Chem 2014;57:3786–802. 10.1021/jm500317a [DOI] [PubMed] [Google Scholar]
- 37. Kramer C, Ting A, Zheng H et al. Learning medicinal chemistry absorption, distribution, metabolism, excretion, and toxicity (ADMET) rules from cross-company matched molecular pairs analysis (MMPA). J Med Chem 2018;61:3277–92. 10.1021/acs.jmedchem.7b00935 [DOI] [PubMed] [Google Scholar]
- 38. Fu L, Yang ZY, Yang ZJ et al. QSAR-assisted-MMPA to expand chemical transformation space for lead optimization. Brief Bioinform 2021;22:bbaa374. 10.1093/bib/bbaa374 [DOI] [PubMed] [Google Scholar]
- 39. Wang D, Jin J, Shi G et al. ADMET evaluation in drug discovery: 21. Application and industrial validation of machine learning algorithms for Caco-2 permeability prediction. J Cheminform 2025;17:3. 10.1186/s13321-025-00947-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Glyn RJ, Pattison G. Effects of replacing oxygenated functionality with fluorine on lipophilicity. J Med Chem 2021;64:10246–59. 10.1021/acs.jmedchem.1c00668 [DOI] [PubMed] [Google Scholar]
- 41. Hu X, Hu Y, Vogt M et al. MMP-cliffs: systematic identification of activity cliffs on the basis of matched molecular pairs. J Chem Inf Model 2012;52:1138–45. 10.1021/ci3001138 [DOI] [PubMed] [Google Scholar]
- 42. Dimova D, Hu Y, Bajorath J. Matched molecular pair analysis of small molecule microarray data identifies promiscuity cliffs and reveals molecular origins of extreme compound promiscuity. J Med Chem 2012;55:10220–8. 10.1021/jm301292a [DOI] [PubMed] [Google Scholar]
- 43. He J, You H, Sandström E et al. Molecular optimization by capturing chemist’s intuition using deep neural networks. J Cheminform 2021;13:26. 10.1186/s13321-021-00497-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Pan B, Zhang PZ, Pang HW et al. Retrieval-augmented foundation models for matched molecular pair transformations to recapitulate medicinal chemistry intuition. arXiv [Preprint]. 2026. arXiv:2602.16684. 10.48550/arXiv.2602.16684 [DOI]
- 45. Moret M, Friedrich L, Grisoni F et al. Generative molecular design in low data regimes. Nat Mach Intell 2020;2:171–80. 10.1038/s42256-020-0160-y [DOI] [Google Scholar]
- 46. Bickerton GR, Paolini GV, Besnard J et al. Quantifying the chemical beauty of drugs. Nat Chem 2012;4:90–8. 10.1038/nchem.1243 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Ertl P, Schuffenhauer A. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. J Cheminform 2009;1:8. 10.1186/1758-2946-1-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Beck JM, Springer C. Quantitative structure-activity relationship models of chemical transformations from matched pairs analyses. J Chem Inf Model 2014;54:1226–34. 10.1021/ci500012n [DOI] [PubMed] [Google Scholar]
- 49. Landrum G. Rdkit documentation. Release 2013;1:4. [Google Scholar]
- 50. Rogers D, Hahn M. Extended-connectivity fingerprints. J Chem Inf Model 2010;50:742–54. 10.1021/ci100050t [DOI] [PubMed] [Google Scholar]
- 51. Cereto-Massagué A, Ojeda MJ, Valls C et al. Molecular fingerprint similarity search in virtual screening. Methods 2015;71:58–63. 10.1016/j.ymeth.2014.08.005 [DOI] [PubMed] [Google Scholar]
- 52. Breiman L. Random forests. Mach Learn 2001;45:5–32. 10.1023/A:1010933404324 [DOI] [Google Scholar]
- 53. Chen T, Guestrin C. Xgboost: A scalable tree boosting system. In: Krishnapuram B, Shah M, Smola AJ, Aggarwal CC, Shen D, Rastogi R (eds). Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining. New York, NY, USA: Association for Computing Machinery, 2016, 785–794. 10.1145/2939672.2939785 [DOI]
- 54. Huang X, Cao D-S, Xu Q-S et al. A novel tree kernel support vector machine classifier for modeling the relationship between bioactivity and molecular descriptors. Chemom Intell Lab Syst 2013;120:71–6. 10.1016/j.chemolab.2012.11.008 [DOI] [Google Scholar]
- 55. Hearst MA, Dumais ST, Osuna E et al. Support vector machines. IEEE Intelligent Systems and their applications 1998;13:18–28. 10.1109/5254.708428 [DOI] [Google Scholar]
- 56. Webb GI, Keogh E, Miikkulainen R. Naïve Bayes. Encycl Mach Learn 2010;15:713–4. 10.1007/978-0-387-30164-8_576 [DOI] [Google Scholar]
- 57. Cox DR. The regression analysis of binary sequences. J R Stat Soc Ser B: Stat Methodol 1958;20:215–32. 10.1111/j.2517-6161.1958.tb00292.x [DOI] [Google Scholar]
- 58. Rumelhart DE, Hinton GE, Williams RJ. Learning representations by back-propagating errors. Nature 1986;323:533–6. 10.1038/323533a0 [DOI] [Google Scholar]
- 59. Wang JB, Cao DS, Zhu MF et al. In silico evaluation of logD7. 4 and comparison with other prediction methods. J Chemom 2015;29:389–98. 10.1002/cem.2718 [DOI] [Google Scholar]
- 60. Ranhotra HS. The estrogen-related receptors in metabolism and cancer: newer insights. J Recept Signal Transduct Res 2018;38:95–100. 10.1080/10799893.2018.1456552 [DOI] [PubMed] [Google Scholar]
- 61. Crevet L, Vanacker JM. Regulation of the expression of the estrogen related receptors (ERRs). Cell Mol Life Sci 2020;77:4573–9. 10.1007/s00018-020-03549-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Sharpe LJ, Brown AJ. Controlling cholesterol synthesis beyond 3-hydroxy-3-methylglutaryl-CoA reductase (HMGCR). J Biol Chem 2013;288:18707–15. 10.1074/jbc.R113.479808 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Jiang SY, Li H, Tang JJ et al. Discovery of a potent HMG-CoA reductase degrader that eliminates statin-induced reductase accumulation and lowers cholesterol. Nat Commun 2018;9:5138. 10.1038/s41467-018-07590-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Wiesel-Motiuk N, Assaraf YG. The key roles of the lysine acetyltransferases KAT6A and KAT6B in physiology and pathology. Drug Resist Updat 2020;53:100729. 10.1016/j.drup.2020.100729 [DOI] [PubMed] [Google Scholar]
- 65. De Martinis M, Allegra A, Sirufo MM et al. Vitamin D deficiency, osteoporosis and effect on autoimmune diseases and hematopoiesis: a review. Int J Mol Sci 2021;22:8855. 10.3390/ijms22168855 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Basu A. The enigmatic protein kinase C-eta. Cancers (Basel) 2019;11:214. 10.3390/cancers11020214 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Noordermeer SM, van Attikum H. PARP inhibitor resistance: a tug-of-war in BRCA-mutated cells. Trends Cell Biol 2019;29:820–34. 10.1016/j.tcb.2019.07.008 [DOI] [PubMed] [Google Scholar]
- 68. Ray Chaudhuri A, Nussenzweig A. The multifaceted roles of PARP1 in DNA repair and chromatin remodelling. Nat Rev Mol Cell Biol 2017;18:610–21. 10.1038/nrm.2017.53 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Wei X, Zhou F, Zhang L. PARP1-DNA co-condensation: the driver of broken DNA repair. Signal Transduct Target Ther 2024;9:135. 10.1038/s41392-024-01832-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Donati B, Lorenzini E, Ciarrocchi A. BRD4 and cancer: going beyond transcriptional regulation. Mol Cancer 2018;17:164. 10.1186/s12943-018-0915-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Jamroze A, Chatta G, Tang DG. Androgen receptor (AR) heterogeneity in prostate cancer and therapy resistance. Cancer Lett 2021;518:1–9. 10.1016/j.canlet.2021.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Andersson KE. PDE5 inhibitors - pharmacology and clinical applications 20 years after sildenafil discovery. Br J Pharmacol 2018;175:2554–65. 10.1111/bph.14205 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Das A, Durrant D, Salloum FN et al. PDE5 inhibitors as therapeutics for heart disease, diabetes and cancer. Pharmacol Ther 2015;147:12–21. 10.1016/j.pharmthera.2014.10.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Mann KG, Brummel K, Butenas S. What is all that thrombin for? J Thromb Haemost 2003;1:1504–14. 10.1046/j.1538-7836.2003.00298.x [DOI] [PubMed] [Google Scholar]
- 75. Negrier C, Shima M, Hoffman M. The central role of thrombin in bleeding disorders. Blood Rev 2019;38:100582. 10.1016/j.blre.2019.05.006 [DOI] [PubMed] [Google Scholar]
- 76. Van der Maaten L, Hinton G. Visualizing data using t-SNE. J Mach Learn Res 2008;9:2579–2605. [Google Scholar]
- 77. Wilcoxon F. Individual Comparisons by Ranking Methods. Breakthroughs in Statistics: Methodology and Distribution. New York, NY, USA: Springer, 1992. 196–202. 10.1007/978-1-4612-4380-9_16 [DOI] [Google Scholar]
- 78. Mervin LH, Bulusu KC, Leen K et al. Orthologue chemical space and its influence on target prediction. Bioinformatics 2017;34:72–9. 10.1093/bioinformatics/btx525 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Bosc N, Atkinson F, Felix E et al. Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery. J Cheminform 2019;11:4. 10.1186/s13321-018-0325-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol 1995;57:289–300. 10.1111/j.2517-6161.1995.tb02031.x [DOI] [Google Scholar]
- 81. Baell JB, Holloway GA. New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. J Med Chem 2010;53:2719–40. 10.1021/jm901137j [DOI] [PubMed] [Google Scholar]
- 82. Brenk R, Schipani A, James D et al. Lessons learnt from assembling screening libraries for drug discovery for neglected diseases. ChemMedChem 2008;3:435–44. 10.1002/cmdc.200700139 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The codes used for this research work are made publicly available in the GitHub repository [https://github.com/DKCCCChen/MMP2Mol].
