Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 Sep 30;27(5):bbag548. doi: 10.1093/bib/bbag548

MMP2Mol: a matched molecular pairs-based framework for ligand-based de novo drug design

Dekun Chen 1,2,3, Hui Li 4, Jianwang Liu 5,6,7, Yueping Jiang 8,9,10, Dongsheng Cao 11,✉, Shao Liu 12,13,14,✉
PMCID: PMC13625653  PMID: 42814605

Abstract

Designing target-relevant molecules is particularly difficult when only limited bioactivity data are available. We present MMP2Mol, a ligand-based generative framework that combines matched molecular pair (MMP) analysis with a pretrained chemical language model (CLM). MMP2Mol extracts target-specific structure–activity transformations from available ligands, applies the prioritized transformations to construct a virtual focused library, and fine-tunes the CLM toward target-relevant chemical space. The framework was evaluated across ten therapeutic targets and compared with direct CLM fine-tuning, Seq2Seq, and Reinvent 4. In repeated experiments on F2, BRD4, and PARP1, MMP2Mol achieved 92.8%–95.3% validity and 99.5%–99.7% uniqueness. When both methods were evaluated against the same original active-compound reference, molecular novelty increased from 57.8%–64.0% for the CLM baseline to 79.8%–82.0% for MMP2Mol, while internal diversity remained broadly comparable. The gains were most evident for F2 and BRD4, whereas performance varied across the smaller target datasets. These findings indicate that MMP-derived chemical knowledge can improve the focus and reproducibility of ligand-based molecular generation under limited-data conditions. MMP2Mol therefore provides a practical strategy for computational candidate generation and prioritization in early-stage drug discovery.

Keywords: matched molecular pair (MMP), ligand-based drug design, chemical language models (CLMs), data-scarce drug discovery, structure–activity relationship (SAR)

Introduction

Drug development aims to identify compounds with favourable bioactivity, safety, and pharmacokinetics [1–3]. However, the drug-like chemical space is tiny compared to the whole chemical universe [4], and traditional drug discovery methods are slow, costly, and often unpredictable [5]. Computational approaches have therefore become indispensable complements to experimental screening [6]. More recently, artificial intelligence (AI), including machine learning and deep learning, has accelerated target identification, molecular design, and property prediction, giving rise to AI-driven drug discovery (AIDD) [7–11]. AIDD can improve efficiency and reduce costs while supporting structural optimization, efficacy prediction, and toxicity assessment [12–14].

Among AI methods, deep generative models have become important tools for de novo drug design. They generate virtual libraries from one-dimensional SMILES strings [4, 15–17], two-dimensional molecular graphs [18–20], or three-dimensional structures [21–24], often relying on large ligand-receptor bioactivity datasets. Such data are often limited for transmembrane or highly flexible proteins, protein-nucleic-acid complexes, and glycoproteins, which can reduce model reliability. Chemical language models (CLMs) can produce target-focused libraries [16, 25–28], but small fine-tuning sets may inadequately define the relevant chemical space. Drug-likeness and synthetic feasibility also remain important constraints [29].

Data scarcity provides a strong rationale for integrating explicit chemical knowledge with data-driven molecular generation. Matched molecular pair (MMP) analysis, introduced in 2005 [30], compares molecule pairs that differ by a single well-defined structural change and extracts local structure–activity relationship (SAR) information [31–34]. MMP analysis can provide interpretable local quantitative structure–activity relationship (QSAR) guidance [35] and has been applied to ADMET (absorption, distribution, metabolism, excretion, and toxicity) optimization, bioisosteric replacement, and activity-cliff analysis [36–42]. Recent generative approaches have incorporated MMP-derived knowledge more directly. He et al. [43] framed molecular optimization as a translation task and trained Seq2Seq and Transformer models on MMPs to generate optimized analogues conditioned on a starting molecule and desired property changes. More recently, Pan et al. [44] introduced a variable-to-variable foundation model for MMP transformations, together with a retrieval-augmented framework that uses reference analogues to improve transformation diversity and controllability.

MMP2Mol differs from these approaches in how MMP knowledge is incorporated into molecular generation. It does not formulate generation as direct translation from a starting molecule to an optimized product, nor does it retrieve or generate individual transformations during inference. Instead, MMP2Mol extracts target-specific transformations from curated ligand data, prioritizes them using observed activity changes, rule support, and physicochemical-property criteria, applies retained transformations to known active compounds, and computationally prioritizes the resulting virtual focused library. The selected full molecules are then used to fine-tune a pretrained CLM for stochastic de novo generation. Thus, MMP knowledge is used primarily for target-specific data augmentation and distribution shaping rather than as an inference-time edit generator. The rule-ranking stage is treated as a prioritization procedure rather than confirmatory inference because the original rule-level P-values were not adjusted for the large number of candidate transformations.

This staged workflow preserves the interpretability of local MMP transformations while allowing the fine-tuned CLM to generate complete molecules beyond direct one-step analogues. Building on prior work in low-data generative molecular design [45], we therefore evaluate whether target-specific MMP-guided augmentation improves reproducible generation validity, structural novelty, chemical-space focus, and computational prioritization relative to direct CLM fine-tuning, and under which target and data conditions those gains remain credible.

Results and discussion

Overview of MMP2Mol

MMP2Mol comprises four stages (Fig. 1): target-ligand curation and activity modelling; MMP transformation prioritization by activity change, positive-pair fraction, support, quantitative estimate of drug-likeness (QED) [46], and synthetic accessibility score (SA) [47]; application of retained rules to known actives followed by computational prioritization of the virtual focused library; and fine-tuning and stochastic sampling of a pretrained CLM to generate complete molecules.

Figure 1.

MMP2Mol workflow integrating matched molecular pair rules, QSAR screening, and chemical language model fine-tuning. Separate QSAR branches provide retrospective reliability assessment and full-data screening. The screened virtual library guides target-specific molecular generation.

Overall MMP2Mol architecture. (a) Curation of target-specific active and inactive compounds and construction of activity-prediction QSAR models. A scaffold-disjoint development/locked-test workflow was used to construct Locked-QSAR models for retrospective reliability assessment, whereas the selected descriptor-classifier combination was refitted on the complete labelled dataset to construct a Full-data QSAR model for application-stage screening. (b) MMPA-driven rule extraction, filtering, and virtual focused library generation. (c) Molecules retained by QSAR screening were used to fine-tune a ChEMBL 24-pretrained CLM for target-specific molecular generation.

To assess QSAR reliability [48], standardized and deduplicated compounds were divided by non-overlapping Bemis–Murcko scaffolds into a development set (~80%) and a locked test set (~20%). Four prespecified molecular representations were evaluated: 167-bit Molecular ACCess System (MACCS) keys, 2048-bit extended connectivity fingerprint 6 (ECFP6), 2048-bit functional connectivity fingerprint with diameter 6 (FCFP6) [49–51], and a concatenated representation comprising MACCS, ECFP6, and FCFP6. Each representation was combined with six classifiers: random forest (RF) [52], extreme gradient boosting (XGBoost) [53], support-vector machine (SVM) [54, 55], Naive Bayes (NB) [56], logistic regression [57], and multilayer perceptron (MLP) [58], giving 24 candidate descriptor-classifier combinations per target. Model selection was performed exclusively within the development set using five-fold stratified scaffold-grouped cross-validation (CV) [59]. The selected combination was refitted on the complete development set and evaluated on the locked test set using a fixed probability threshold of 0.5. Development-set results for all 240 combinations are provided in Table S1 and Fig. S1; complete locked-test results are provided in Table S3 and Fig. S2. Full-data QSAR model-selection results are provided in Table S2 and Fig. S3.

To assess QSAR generalization independently of application-stage screening, we established target-specific scaffold-disjoint evaluation models, hereafter termed locked-QSAR. Models were selected by five-fold scaffold-grouped CV within development partitions and evaluated once on untouched locked test sets (Table 1). Performance was strongest for PDE5A and PARP1 and remained useful for AR, BRD4, and F2. Estimates for smaller targets require caution; HMGCR’s ROC-AUC of 1.000 was based on 36 compounds, whereas PKCeta showed no discrimination (ROC-AUC = 0.479; MCC = −0.138). Locked-QSAR models were subsequently used for reduced-data and ablation analyses, while separate full-data QSAR models supported application-stage screening.

Table 1.

Target-specific locked-QSAR model selection and performance on the scaffold-disjoint locked test set.

a Target Selected model Development CV ROC-AUC Locked-test n (active %) Locked-test ROC-AUC (95% CI) PR -AUC MCC AD coverage
AR Combined- XGBoost 0.822 ± 0.036 444 (61.5%) 0.849 (0.797–0.889) 0.885 0.575 73.6%
BRD4 Combined-RF 0.914 ± 0.017 871 (55.6%) 0.888 (0.842–0.928) 0.902 0.644 85.0%
F2 Combined-RF 0.884 ± 0.026 951 (54.3%) 0.868 (0.825–0.905) 0.888 0.581 74.7%
PARP1 Combined-RF 0.916 ± 0.035 644 (80.0%) 0.905 (0.855–0.945) 0.974 0.596 81.1%
PDE5A ECFP6- SVM 0.918 ± 0.032 383 (65.8%) 0.931 (0.888–0.962) 0.942 0.704 90.6%
ERR1 ECFP6-NB 0.870 ± 0.130 29 (65.5%) 0.989 (0.899–1.000) 0.995 0.847 82.8%
HMGCR FCFP6- SVM 0.804 ± 0.076 36 (66.7%) 1.000 (1.000–1.000) 1.000 0.938 91.7%
KAT6A FCFP6-RF 0.875 ± 0.088 39 (59.0%) 0.921 (0.878–1.000) 0.957 0.624 64.1%
PKCeta ECFP6- XGBoost 0.642 ± 0.105 44 (84.1%) 0.479 (0.163–0.735) 0.841 −0.138 93.2%
VDR Combined-RF 0.908 ± 0.101 75 (46.7%) 0.969 (0.888–1.000) 0.950 0.948 84.0%

aCombined denotes concatenated MACCS, ECFP6, and FCFP6 descriptors. Development values are mean ± SD from five-fold scaffold-grouped CV. Confidence intervals were obtained from 1000 scaffold-grouped bootstrap replicates. AD denotes applicability-domain coverage. Complete confidence intervals, balanced accuracy, F1 score, Brier score, sensitivity, specificity, and in-domain/out-of-domain performance are reported in Table S3.

For application-stage screening, all 24 prespecified descriptor-classifier combinations were evaluated for each target using five-fold stratified scaffold-grouped CV. The best-performing combination was subsequently refitted on the complete curated labeled dataset to establish the target-specific full-data QSAR model. Generated molecules with predicted probabilities ≥0.5 were classified as active. These full-data QSAR models were used for virtual-library and generated-molecule screening. The five-fold CV ROC curves and complete model-selection results are presented in Fig. S3 and Table S2, respectively.

Performance of MMP2Mol in bioactive-molecule generation

We evaluated MMP2Mol on ten therapeutic targets from ChEMBL, divided into data-scarce (<500 active compounds: ERR1 [60, 61], HMGCR [62, 63], KAT6A [64], VDR [65], PKCeta [66]) and data-dense (>1000 active compounds: PARP1 [67–69], BRD4 [70], AR [71], PDE5A [72, 73], F2 [74, 75]). The original application-stage analysis used 10 000 generated molecules per target for docking evaluation. Independent stochastic repetitions were subsequently performed for all four generators using fixed method-specific configurations. For the principal repeated full-data QSAR, QED, and SA analyses, CLM and MMP2Mol outputs generated after epochs 1–30 were pooled within each independent run, standardized, and deduplicated by canonical SMILES; each pooled run was treated as one statistical replicate. Reinvent 4 and the Seq2Seq baseline, implemented using the Mol2Mol module in Reinvent 4, were evaluated using independently sampled run-level libraries under fixed official configurations. In contrast, the separate scale-matched CLM-MMP2Mol comparison used epoch-30 outputs only to evaluate validity, uniqueness, molecular novelty relative to the original target-specific active set, ECFP4 internal diversity, and Bemis–Murcko scaffold novelty (Table S4).

Across the five data-dense targets, MMP2Mol increased mean validity by 1.8–7.3 percentage points and molecular novelty by 15.8–22.6 percentage points relative to CLM, while maintaining 98.6%–99.7% uniqueness and broadly comparable internal diversity (Table S4). Its mean full-data QSAR-predicted active rate exceeded those of all three baselines for every target, ranging from 77.2% to 93.6% in the data-dense tasks (Table 2) and from 51.8% to 89.9% in the data-scarce tasks (Table 3). QED, SA, and docking performance remained competitive. The predicted-active rate denotes the proportion of valid, unique molecules assigned a probability ≥0.5 by the full-data QSAR model; model-selection results are provided in Table S2 and Fig. S3, with detailed data-scarce results in Table S5.

Table 2.

Performance on data-dense targets.

a Target Method Best docking score↓ (kcal/mol) Docking score, mean ± SD↓ (kcal/mol) Full-data QSAR-predicted active, mean (95% CI)↑ QED, mean (95% CI)↑ SA, mean (95% CI)↓
F2 MMP2Mol −15.77 −12.75 ± 0.80 77.2% (75.4–78.4) 0.39 (0.39–0.40) 3.16 (3.15–3.17)
CLM −15.76 −12.84 ± 0.88 65.2% (64.3–66.5) 0.38 (0.38–0.39) 3.10 (3.09–3.10)
Seq2Seq −13.02 −8.04 ± 1.01 72.6% (71.1–74.1) 0.32 (0.32–0.33) 3.35 (3.33–3.37)
Reinvent 4 −13.22 −10.94 ± 0.58 60.8% (59.4–62.2) 0.41 (0.40–0.42) 3.08 (3.04–3.12)
AR MMP2Mol −12.63 −10.91 ± 0.31 86.3% (83.3–88.4) 0.66 (0.66–0.67) 3.19 (3.18–3.21)
CLM −12.15 −10.82 ± 0.30 78.3% (77.0–79.6) 0.67 (0.67–0.68) 3.24 (3.24–3.24)
Seq2Seq −12.25 −9.39 ± 0.60 68.4% (66.8–70.0) 0.72 (0.71–0.72) 3.87 (3.84–3.91)
Reinvent 4 −12.83 −10.73 ± 0.29 73.1% (71.0–75.1) 0.70 (0.69–0.71) 3.42 (3.40–3.44)
BRD4 MMP2Mol −9.78 −7.68 ± 0.40 82.7% (79.7–84.4) 0.50 (0.50–0.51) 3.10 (3.09–3.11)
CLM −9.71 −7.53 ± 0.43 68.5% (68.0–68.9) 0.53 (0.52–0.53) 3.05 (3.04–3.06)
Seq2Seq −10.37 −6.54 ± 1.74 63.7% (62.1–65.4) 0.50 (0.49–0.5) 3.33 (3.30–3.36)
Reinvent 4 −10.44 −8.07 ± 0.60 53.5% (52.2–54.8) 0.54 (0.53–0.55) 2.98 (2.95–3.02)
PARP1 MMP2Mol −15.17 −12.74 ± 0.62 93.6% (92.4–94.4) 0.66 (0.65–0.67) 2.86 (2.85–2.88)
CLM −15.05 −12.67 ± 0.57 90.1% (89.9–90.5) 0.64 (0.64–0.65) 2.72 (2.71–2.73)
Seq2Seq −14.17 −10.42 ± 1.17 80.1% (79.4–80.7) 0.57 (0.56–0.58) 2.93 (2.90–2.96)
Reinvent 4 −14.27 −9.57 ± 1.40 64.2% (63.6–64.8) 0.64 (0.63–0.64) 2.70 (2.66–2.74)
PDE5A MMP2Mol −14.26 −12.38 ± 0.47 90.2% (80.3–96.0) 0.53 (0.52–0.53) 3.01 (2.99–3.02)
CLM −14.33 −12.09 ± 0.58 73.3% (69.0–76.1) 0.51 (0.50–0.52) 3.00 (2.99–3.01)
Seq2Seq −13.04 −8.57 ± 1.38 77.2% (74.2–78.9) 0.48 (0.47–0.48) 3.22 (3.20–3.40)
Reinvent 4 −10.38 −7.11 ± 0.66 60.3% (58.1–62.5) 0.55 (0.54–0.56) 2.88 (2.83–2.92)

aFull-data QSAR-predicted active rates, QED, and SA are reported as run-level means with 95% confidence intervals. For CLM and MMP2Mol, outputs generated after each of epochs 1–30 were pooled within the corresponding independent run before standardization and canonical-SMILES deduplication, and each run was weighted equally irrespective of the resulting library size. Docking statistics were obtained from the original application-stage libraries and represent molecule-level summaries rather than between-run confidence intervals.

Table 3.

Comparison of MMP2Mol and other methods on data-scarce targets.

a Target Full-data QSAR-predicted active, mean (95% CI)↑ Docking Score↓ (kcal/mol)
MMP2Mol CLM Seq2Seq Reinvent 4 MMP2Mol CLM Seq2Seq Reinvent 4
VDR 69.3% (53.6–72.7) 24.9% (23.7–25.8) 43.3% (42.2–44.5) 45.2% (44.9–45.5) −10.56 −10.42 −10.25 −9.86
HMGCR 63.0% (58.1–66.1) 33.0% (30.8–34.5) 41.7% (40.1–42.9) 48.1% (47.3–48.9) −10.91 −10.08 −7.59 −7.58
ERR1 51.8% (44.7–55.4) 16.3% (15.8–16.7) 37.2% (36.6–37.8) 43.8% (42.2–44.5) −8.89 −8.50 −7.75 −8.40
PKCeta 89.9% (89.1–90.5) 87.4% (86.9–88.5) 82.1% (80.9–83.3) 60.6% (59.8–61.4) −9.22 −9.41 −9.06 −7.96
KAT6A 71.7% (65.9–75.2) 57.3% (56.8–57.5) 65.2% (63.8–66.3) 52.1% (50.8–53.5) −10.27 −9.36 −8.13 −8.80

aFull-data QSAR-predicted active rates are reported as run-level means with 95% confidence intervals. For CLM and MMP2Mol, outputs generated after epochs 1–30 were pooled within each independent run before standardization and canonical-SMILES deduplication. Reinvent 4 and Seq2Seq were evaluated using independently sampled run-level libraries. Docking statistics are descriptive molecule-level summaries from the original application-stage libraries.

We further compared the docking-score distributions of prioritized MMP2Mol-generated molecules with those of known actives (Fig. 2a). t-Distributed stochastic neighbour embedding (t-SNE) [76] analysis of equal-size MMP2Mol and CLM samples (Fig. 2b) revealed partially overlapping yet complementary chemical spaces, consistent with MMP-guided enrichment of target-relevant chemotypes beyond direct CLM fine-tuning.

Figure 2.

Box plots compare docking scores of selected MMP2Mol-generated molecules and known actives across ten targets. Chemical-space maps show overlapping and distinct regions occupied by MMP2Mol- and CLM-generated molecules.

Docking scores and chemical space distribution of MMP2Mol-generated molecules. (a) Docking-score distributions of the 2000 valid and unique MMP2Mol-generated molecules with the most favourable AutoDock Vina scores for each target. For each target, the known bioactive set comprises the top 500 compounds with the highest pChEMBL values, or all available compounds when fewer than 500 were present. The boxes represent the interquartile range, the central lines indicate the medians, and the whiskers extend to 1.5 times the interquartile range. (b) Chemical-space comparison of MMP2Mol- and CLM-generated molecules using equal-size samples of 5000 valid and unique molecules from each method for each target. Molecules were sampled proportionally across ten ECFP6-based structural clusters to reduce overplotting while preserving chemical-space coverage and were visualized using t-SNE (perplexity = 30).

Overall, MMP2Mol balanced validity, novelty, computational activity enrichment, and drug-like properties across data-dense and data-scarce tasks. For some small datasets, stronger focus reduced internal or scaffold diversity and increased QSAR uncertainty. Accordingly, QSAR classifications and docking scores are interpreted as computational prioritization endpoints.

Performance under reduced-data scenarios

To examine generation under reduced target-specific data conditions, we constructed case studies for F2, AR, PARP1, and BRD4 using 10% subsets selected by ECFP6 structure-cluster-stratified sampling. Only these selected subsets were used for MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning. For these targets, canonical-SMILES matches to the complete target-specific sets were additionally removed from the ChEMBL 24 pretraining corpus before pretraining the corresponding CLMs. Generated molecules were evaluated using the locked-QSAR and target-specific re-scoring models as in the main application workflow. Thus, the reduced-data constraint applies to target-specific rule extraction and generator adaptation, whereas the reported classifier outputs are practical post-generation screening endpoints. The re-scoring models used AutoDock Vina-derived residue-interaction descriptors within 5 Å of the docked poses and random-forest classifiers at a probability threshold of 0.5; their CV performance is reported in Figs S4–S5 and Tables S7–S8. As shown in Table 4, MMP2Mol achieved higher locked-QSAR-predicted active rates than CLM for F2 (39.5% versus 24.2%), AR (60.9% versus 47.2%), PARP1 (85.6% versus 76.2%), and BRD4 (51.5% versus 17.2%). The comparatively high PARP1 rates are consistent with target-specific enrichment relative to the locked evaluator’s decision boundary. Because Table 2 used full-data QSAR models, absolute values between the two settings are not directly comparable. The docking distributions in Fig. 3 further showed a smaller change for MMP2Mol than for CLM after target-specific data reduction. Detailed AD coverage, out-of-domain fractions, domain-restricted predicted-active rates, and 95% confidence intervals are shown (Table S6).

Table 4.

Performance of MMP2Mol in the data-reduced scenarios.

Target Compounds Applicable rules Locked-QSAR-predicted active rate (%)↑ Re-scoring-predicted active rate (%)↑
100% 10% 100% 10% MMP2Mol CLM MMP2Mol CLM
F2 4822 473 6659 78 39.5 24.2 61.6 54.9
AR 3179 216 6813 61 60.9 47.2 41.8 31.4
PARP1 3223 318 21,235 173 85.6 76.2 89.4 82.4
BRD4 4359 433 11,848 154 51.5 17.2 59.6 37.6

aThe 10% subsets were obtained by proportional sampling across ECFP6 structural clusters and were fixed before model training. Compounds outside the subsets were excluded from MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning. QSAR-predicted active rates were calculated after generation using the locked-QSAR at a probability threshold of 0.5; re-scoring-predicted active rates were calculated using the target-specific re-scoring models at the same threshold.

Figure 3.

Density plots compare docking scores under complete-data and 10% data conditions. Comparisons include MMP2Mol versus CLM and within-method changes for F2, with MMP2Mol showing a smaller distributional shift after data reduction.

Docking-score distributions under complete-data and 10% target-specific data settings. (a) Distributions for MMP2Mol and CLM using the complete target-specific datasets. (b) Corresponding distributions using the 10% cluster-stratified subsets. (c) Comparison of CLM-generated F2 molecules under the complete-data and 10% settings. (d) Corresponding comparison for MMP2Mol.

Before CLM fine-tuning, we evaluated the MMP-derived virtual libraries (Table 5). From 181 HMGCR compounds, 205 rules generated 5783 molecules, of which 5579 (96.5%) were full-data QSAR predicted active. Across five targets, the libraries preserved useful QED and SA distributions (Fig. 4), while subsequent CLM fine-tuning added structures beyond direct transformations. In the size-matched comparison (Table 6), MMP2Mol exceeded CLM in QSAR- and re-scoring-predicted active rates for all targets; corrected PKCeta values were 92.5% versus 91.7% and 80.4% versus 71.0%, respectively. SA was comparable, and QED was similar or higher for MMP2Mol in four targets. Together, these analyses support reusable MMP-guided expansion under limited data; activity assignments remain computational prioritization endpoints.

Table 5.

Performance of the MMP-derived virtual libraries within the MMP2Mol framework.

a Target Compounds Transformations Applicable rules MMPA- derived molecules Full-data QSAR- predicted active molecules Avg QED ↑ Avg SA↓
VDR 394 35,417 1289 1527 1412 0.3898 (0.3649) 4.291 (4.207)
HMGCR 181 5825 205 5783 5579 0.4677 (0.3675) 3.758 (3.784)
ERR1 165 8811 404 1579 1546 0.4290 (0.3847) 2.947 (3.171)
PKCeta 225 10,890 424 8056 7931 0.4675 (0.3720) 3.132 (3.521)
KAT6A 246 52,471 2412 23,796 21,383 0.7022 (0.6783) 2.245 (2.221)

aValues in parentheses in the QED and SA columns are the corresponding means for the original target-specific active-compound set.

Figure 4.

Density plots compare drug-likeness and synthetic accessibility between virtual libraries and known actives for ERR1, PKCeta, VDR, HMGCR, and KAT6A. Higher QED and lower SA scores are preferred; dashed lines mark SA = 6.

Molecular-property distributions of the virtual libraries constructed by MMP2Mol and the corresponding known bioactive compounds. Kernel-density estimates show the distributions of (a) QED and (b) SA across five targets. Higher QED values indicate more favourable drug-like properties, whereas lower SA scores indicate greater predicted synthetic accessibility. The dashed lines indicate the predefined SA threshold of 6.0.

Table 6.

Scale-matched comparison of MMP2Mol and CLM across data-scarce targets.

Target Full-data QSAR-predicted active rate↑ Re-scoring-predicted active rate↑ Avg SA↓ Avg QED↑
MMP2Mol CLM MMP2Mol CLM MMP2Mol CLM MMP2Mol CLM
VDR 69.3% 24.9% 41.1% 25.5% 3.95 3.23 0.43 0.42
HMGCR 63.0% 33.0% 62.1% 45.8% 3.56 3.31 0.58 0.42
ERR1 51.8% 16.3% 71.8% 67.4% 2.36 2.25 0.67 0.73
PKCeta 92.5% 91.7% 80.4% 71.0% 2.46 2.46 0.74 0.72
KAT6A 91.1% 71.0% 42.1% 22.3% 2.27 2.30 0.68 0.67

Application and comparative evaluation in ligand-based drug design

To characterize the structural proximity and novelty of the generated molecules, we calculated Morgan-fingerprint-based Tanimoto similarities between the valid and unique molecules generated by each method and the corresponding target-specific original active-compound set. Table 7 summarizes nearest-neighbour Morgan-fingerprint Tanimoto similarity between valid, unique generated molecules and original target actives (range, 0.186–0.671). MMP2Mol generally remained near known-active chemical space while achieving higher full-data QSAR-predicted active rates. For AR, the full-data QSAR-predicted active rates were 86.3%, 78.3%, 68.4%, and 73.1% for MMP2Mol, CLM, Seq2Seq, and Reinvent 4, respectively, whereas the corresponding mean nearest-reference similarities were 0.445, 0.367, 0.416, and 0.412. Thus, enrichment was not explained solely by similarity (Figs S6–S7).

Table 7.

Mean nearest-reference ECFP4 Tanimoto similarities of molecules generated by the four methods.

Target MMP2Mol CLM Reinvent 4 Seq2Seq
F2 0.522 0.451 0.487 0.605
AR 0.445 0.367 0.412 0.416
PARP1 0.549 0.517 0.533 0.644
BRD4 0.392 0.446 0.465 0.603
PDE5A 0.544 0.464 0.444 0.671
VDR 0.463 0.310 0.346 0.671
HMGCR 0.373 0.269 0.267 0.600
ERR1 0.344 0.186 0.227 0.578
PKCeta 0.330 0.264 0.251 0.550
KAT6A 0.518 0.356 0.413 0.597

In a retrospective PARP1 low-data case study, ECFP6 cluster-stratified sampling assigned 318 of the 3219 unique curated compounds to the reduced-data subset, leaving 2901 compounds as a reference holdout. Nearest-neighbour ECFP6 analysis yielded a median maximum Tanimoto similarity of 0.600 (interquartile range, 0.443–0.700); 25.1% of the holdout compounds had a maximum similarity of ≥0.7, only 0.5% had a similarity of ≥0.9, and 33.5% shared a Bemis–Murcko scaffold with the subset (Table S9). Under identical model and sampling settings, MMP2Mol and CLM each generated 40 000 molecules, yielding 37 368 and 36 741 valid and unique structures, respectively. Among molecules with nearest-reference Tanimoto similarity >0.9, 65.4% and 67.2%, respectively, met a QSAR probability ≥0.5 and docking score ≤ −8.0 kcal/mol; MMP2Mol recovered five exact holdout structures versus one for CLM (Fig. 5). At similarities of 0.2–0.5, joint pass rates were 24.8% and 13.3% (Fig. 6), indicating improved recovery and computational prioritization. Compounds 4 and 8 from Fig. 6, selected for complementary predicted activity, QED, SA, and novelty, were subjected to retrosynthetic analysis (Fig. S9).

Figure 5.

Structures of five reference-holdout PARP1 compounds recovered by MMP2Mol, including niraparib, and one recovered by CLM. Annotations identify each compound and its reported activity value.

Exact reference-holdout recovery by MMP2Mol and the CLM baseline in the PARP1 reduced-data case study. Both models were trained on the same 10% target-specific subset (318 molecules) and each generated 40 000 molecules. MMP2Mol recovered five exact structures from the reference holdout, compared with one for CLM.

Figure 6.

Nine representative MMP2Mol-generated PARP1 candidates, annotated with drug-likeness, synthetic accessibility, docking scores, and nearest-reference similarities of ~0.2–0.5. Reference compounds are identified by ChEMBL identifiers.

Representative MMP2Mol-generated molecules with nearest-reference Tanimoto similarities of 0.2–0.5 in the PARP1 reduced-data experiment. QED and SA are computational estimates; the closest reference compounds are shown in Fig. S8.

To test whether diversity compromised beneficial MMP transformations, we measured rule-containing generated molecules. For F2, 86% incorporated at least two rule-defined substructure modifications, whereas 14% contributed scaffold-level diversity. Thus, diversity and rule adherence coexisted (Fig. S10), supporting target-focused exploration while preserving MMP-derived information.

Ablation analysis of MMP-derived augmentation and QSAR filtering

To examine the contributions of the main MMP2Mol components, a configuration-based ablation study was conducted on PARP1 and BRD4. Four configurations were compared: A0, direct CLM fine-tuning on the original active compounds; A1, fine-tuning on molecules generated using all retained MMP transformations, including activity-increasing and activity-decreasing directions; A2, fine-tuning on the original actives plus molecules generated using activity-increasing transformations, without final QSAR filtering; and A3, the complete MMP2Mol workflow using the QSAR-filtered expanded library. Each used three prespecified independent runs under identical settings; extra runs were excluded. Size-matched outputs were assessed for validity, uniqueness, novelty, internal and scaffold diversity, and active & AD under the scaffold-disjoint oracle, with novelty referenced to the same original active set.

The unfiltered all-rule variant A1 produced the broadest exploration for both targets (Table 8). A1 produced the greatest chemical-space expansion, with molecular novelty of 97.99% for PARP1 and 96.76% for BRD4, but its active & AD rates decreased to 62.07% and 44.87%, respectively. Restricting augmentation to activity-increasing transformations in A2 increased the corresponding active & AD rates to 74.48% and 74.13%. The complete A3 configuration provided the most favourable overall balance. For PARP1, A3 yielded 96.67% validity, 79.25% novelty, and 80.28% active & AD, compared with 92.00%, 57.71%, and 79.53% for A0. For BRD4, the corresponding A3 values were 95.33%, 82.01%, and 79.54%, compared with 91.67%, 58.57%, and 74.86% for A0. Uniqueness remained above 98% across all configurations, whereas the lower internal diversity of A3 relative to A2 indicated greater concentration within the model-supported region.

Table 8.

Configuration-based ablation results for PARP1 and BRD4. Values are reported as mean ± standard deviation across independent runs.

a Target Configuration Validity (%) ↑ Uniqueness (%) ↑ Molecular novelty (%) ↑ Internal diversity ↑ Scaffold novelty (%) ↑ Active & AD (%) ↑
PARP1 A0 92.00 ± 1.00 99.28 ± 1.26 57.71 ± 7.75 0.8126 ± 0.0180 35.74 ± 6.95 79.53 ± 4.59
A1 84.67 ± 6.51 100.00 ± 0.00 97.99 ± 1.93 0.8349 ± 0.0027 80.60 ± 6.82 62.07 ± 4.12
A2 89.67 ± 3.06 98.88 ± 1.95 83.51 ± 3.20 0.8249 ± 0.0074 60.78 ± 3.76 74.48 ± 1.98
A3 96.67 ± 0.58 99.66 ± 0.60 79.25 ± 3.65 0.8048 ± 0.0062 58.80 ± 3.44 80.28 ± 5.21
BRD4 A0 91.67 ± 2.52 98.55 ± 0.59 58.57 ± 8.04 0.8464 ± 0.0048 39.99 ± 6.53 74.86 ± 4.87
A1 83.00 ± 3.46 100.00 ± 0.00 96.76 ± 1.48 0.8587 ± 0.0037 79.94 ± 2.14 44.87 ± 3.99
A2 86.33 ± 2.89 98.46 ± 0.64 75.69 ± 3.19 0.8525 ± 0.0022 53.49 ± 1.77 74.13 ± 2.78
A3 95.33 ± 0.58 99.30 ± 1.22 82.01 ± 4.09 0.8156 ± 0.0210 60.68 ± 9.85 79.54 ± 3.75

aValues are mean ± sample standard deviation across three independent runs per configuration. All molecular endpoints were calculated from a fixed sample of 100 raw SMILES generated from the epoch-30 checkpoint per independent run. Molecular and scaffold novelty were calculated against the same target-specific original active set. Active & AD represents the proportion simultaneously predicted active and located within the AD of the independent scaffold-disjoint oracle. Bold values indicate the highest result for each target and metric.

MMP augmentation therefore expanded chemical space, activity-increasing rule selection limited inconsistent expansion, and QSAR filtering enriched oracle-supported in-domain molecules. Because fine-tuning-set compositions differed, these are configuration-level rather than strictly additive module effects.

Discussion

MMP2Mol integrates MMP analysis with chemical language modelling to address limited target-specific bioactivity data in ligand-based molecular generation. Its principal contribution is the use of nominal Wilcoxon-prioritized [77], activity-enhancing molecular transformations as chemical priors. These transformations expand the target-focused fine-tuning space and guide the generative model toward structurally meaningful regions that may be difficult to learn directly from a limited number of active compounds. The repeated-generation analysis further showed that the improvements in molecular novelty were most consistent for data-dense targets, indicating that MMP-derived augmentation can enrich the information available to the language model without causing a general loss of generation quality.

The structural analyses provide additional insight into how MMP2Mol balances exploitation and exploration. Molecules generated by MMP2Mol generally remained within a moderate similarity range to known active compounds while achieving increased computational activity enrichment. Together with the molecular novelty, internal diversity, and Bemis–Murcko scaffold novelty results, this pattern is consistent with focused exploration around target-relevant chemical space rather than simple reproduction of the fine-tuning compounds. The reduced-data case studies further showed that MMP-derived augmentation can preserve useful target-specific structural information when only a fraction of the available compounds is used for model adaptation. In the PARP1 analysis, MMP2Mol recovered more structures held out from target-specific fine-tuning and provided greater enrichment than CLM among molecules with low-to-moderate similarity to the reference set.

The generators evaluated in this study were constructed using transformations prioritized by a nominal one-sided Wilcoxon criterion (P < .05), together with the prespecified effect-size and positive-pair filters. A complementary multiplicity-controlled audit using the BH-FDR procedure showed that, under the primary criterion (support ≥5, q ≤ 0.05, median ΔpChEMBL ≥0.30, and positive-pair fraction ≥0.50), no transformations were retained for AR, VDR, HMGCR, ERR1, or KAT6A, and only one was retained for PKCeta (Table S10). Thus, statistically robust rule support was weak or absent for several targets, indicating that the reliability of MMP-derived guidance depends on the quantity, diversity, and consistency of the available target-specific data. Settings with fewer than ten active compounds were not evaluated and may yield too few robust transformations. QSAR and re-scoring reliability also decreases with scarce inactives or out-of-domain molecules. Although MMP2Mol generation is fully ligand based, docking and interaction-based re-scoring require a reliable target structure. These analyses were used only for post-generation ranking, not as training signals; when structures are unavailable, scaffold-disjoint QSAR with applicability-domain assessment or ligand-based pharmacophore models can be used instead. Future work will emphasize uncertainty-aware prediction, domain-guided selection, temporal or prospective validation, and stronger rule-reliability estimates. Incorporating 3D pharmacophore or structure-based constraints and multi-objective optimization of potency, selectivity, ADMET, and synthesis may improve lead optimization, while alternative backbones, cross-target rule transfer, and active learning could extend MMP2Mol to more extreme low-data settings. To support experimental translation, candidates could be ranked by predicted activity, applicability-domain confidence, QED, SA, novelty, and structural alerts, diversity-selected for synthesis, and assessed in target-appropriate biochemical assays before dose–response and cellular confirmation.

Conclusion

We present MMP2Mol, a ligand-based framework that combines MMP-derived transformations with CLM for de novo molecular design. The method addresses limited target-specific data by extracting SAR information through a nominal Wilcoxon-based prioritization workflow, constructing focused virtual libraries, and fine-tuning a pretrained CLM. Across diverse targets, MMP2Mol improved predicted activity enrichment and molecular novelty relative to direct CLM fine-tuning while maintaining competitive drug-likeness and synthetic-accessibility profiles. Its modular workflow provides a practical strategy for early-stage computational candidate generation, particularly for ligand-defined targets with limited activity data.

Methods

Data collection and pretreatment

The ChEMBL 24 collection, containing more than 1.1 million molecular records, was used for pretraining the CLM. Target-specific datasets were restricted to Homo sapiens single-protein targets and comprised five data-scarce targets with fewer than 500 active compounds (ERR1, HMGCR, KAT6A, VDR, and PKCeta) and five data-dense targets with more than 1000 active compounds (PARP1, BRD4, AR, PDE5A, and F2). Molecules were processed with RDKit (https://www.rdkit.org) and represented as canonical SMILES. Salts were removed by retaining the largest organic fragment, invalid structures and records without usable activity values were discarded, and SMILES longer than 90 characters were excluded. Stereochemical annotations present in the source SMILES were preserved. Redundant records were collapsed to a single entry, and when duplicate records had discordant activity values, the record with the higher pChEMBL value was retained. Activity labels were assigned using a 500 nM threshold: records with reported activity ≤500 nM were labelled active and those with activity >500 nM were labelled inactive (equivalent to a pChEMBL boundary of ~6.30) [78, 79]. Both classes were retained for classifier development, whereas active compounds were used as the starting structures for MMP transformation and target-specific CLM fine-tuning.

Application-stage full-data QSAR models

All classifier hyperparameters were fixed a priori and were not optimized separately for individual targets. For each target, the descriptor-classifier combination with the best CV ROC-AUC in the original full-data search was refitted on all curated labelled compounds and used to prioritize valid and unique generated molecules at a probability threshold of 0.5. The full-data QSAR-predicted active rate was defined as the proportion assigned a probability of at least 0.5. Five-fold CV ROC curves for these application models are shown in Fig. S3.

Workflow of molecular generation

  • (1)

    MMPA-guided focused molecular library generation: Rule databases were generated with mmpdb 3.1 using default parameters. For each directed transformation, paired activity differences were defined as ΔpChEMBL = pChEMBL(υB) − pChEMBL(υA) and evaluated using a one-sided Wilcoxon signed-rank test against a non-positive median change. The application-stage workflow used the original nominal Wilcoxon screening criterion (P < .05), followed by a median ΔpChEMBL ≥0.30 and a positive-pair fraction ≥0.50. These nominal Wilcoxon-prioritized rules were used to construct the virtual libraries and train the generators reported in this study. To assess the robustness of rule selection under multiple testing, the complete rule sets were additionally evaluated using Benjamini–Hochberg false-discovery-rate (BH-FDR) [80] correction within each support family, with a primary audit tier requiring support ≥5, q ≤ 0.05, median ΔpChEMBL ≥0.30, and positive-pair fraction ≥0.50; 95% bootstrap confidence intervals used 2000 resamples. The robustness-audit results are reported in Table S10. This analysis was used to characterize the statistical reliability of the extracted transformations, whereas generator training and the reported generation experiments were based on the nominal-Wilcoxon-prioritized rule sets described above.

  • (2)

    Products generated by the retained MMP transformations were sanitized and represented as canonical SMILES using RDKit. Salts were reduced to the largest organic fragment, SMILES longer than 90 characters were discarded, and duplicate products were collapsed by canonical SMILES. Stereochemical annotations present in the source structures were preserved.

  • (3)

    We pretrained the chemical language model (CLM) on more than 1.1 million molecules from ChEMBL 24. The model was implemented in Python (v 3.6.5) using Keras (v 2.2.0, https://keras.io/) with a TensorFlow GPU backend (v 1.9.0, https://www.tensorflow.org). It comprised two batch-normalization layers and two long short-term memory (LSTM) layers with 1024 and 256 units, respectively, totalling 5,820,515 parameters. One-hot-encoded SMILES were used as input, and the model was trained for 30 epochs using the Adam optimizer (learning rate = 10−3) and categorical cross-entropy loss. During target-specific fine-tuning, both LSTM layers remained trainable, and all models were trained for a prespecified 30 epochs with an initial learning rate of 1 × 10−4 [16, 45]. No early stopping was applied; instead, the learning rate was reduced by a factor of 0.5 after three epochs without improvement, to a minimum of 5 × 10−5. Epoch 30 was prespecified as the terminal training checkpoint and was not selected post hoc using generation performance. It was used for the convergence assessment and the epoch-30-specific scale-matched, training-set-overlap, medicinal-chemistry-alert, and ablation analyses. The principal repeated full-data QSAR, QED, and SA analyses instead used outputs pooled across epochs 1–30 within each independent run, as described below. Training and validation losses were recorded at each epoch. Across the 60 audited training histories, validation loss decreased between epochs 25 and 30 in every run; epoch 30 yielded the minimum validation loss in 53 runs, whereas the remaining seven reached comparable minima at epochs 28 or 29. Potential memorization was assessed using exact canonical-SMILES overlap and nearest-neighbour ECFP4 similarity to the corresponding fine-tuning library. Complete convergence and overlap diagnostics are provided in Table S11.

Molecules were sampled using a temperature of 1.0, top-k of 0, and top-p of 0.9. For the principal repeated analysis, SMILES strings generated from the checkpoints saved at epochs 1–30 were pooled within each independent CLM or MMP2Mol run, standardized, and deduplicated using canonical SMILES. Full-data QSAR-predicted active rates, QED values, and SA scores were calculated separately for each pooled run, with each run contributing equally to the summary statistics. By contrast, the separate scale-matched analyses of molecular quality, training-set overlap, and medicinal-chemistry alerts used 100 raw SMILES strings generated from the epoch-30 checkpoint per run. Reinvent 4 and Seq2Seq were evaluated using three independently sampled run-level libraries because these official pretrained implementations did not involve target-specific 30-epoch fine-tuning. Docking statistics were obtained from the original 10 000-molecule application-stage libraries and were treated as descriptive molecule-level summaries. Valid and canonical-unique molecules from the epoch-30 scale-matched samples were retrospectively screened for medicinal-chemistry alerts. Pan-assay interference compounds and Brenk alerts were identified using the corresponding RDKit FilterCatalog definitions [81, 82], supplemented with prespecified SMARTS patterns for potentially reactive functional groups and strained or unstable ring systems. Alerts were treated as advisory flags rather than definitive evidence of chemical invalidity. The alert definitions and target-specific results are reported in Table S12.

Baseline generative models

The CLM baseline used the same pretrained architecture and sampling settings as MMP2Mol but was fine-tuned on the corresponding original active-compound set without MMP-derived augmentation. The Reinvent 4 de novo generator and the Seq2Seq comparator, which was implemented using the conditional Mol2Mol module provided in the Reinvent 4 package (https://github.com/MolecularAI/REINVENT4), were run using the corresponding official pretrained models and software-default configurations. Only the target-specific input or generation mode and the required output size were changed, and no target-specific hyperparameter tuning was performed.

Optimization of transformation rules

Optimization of application-stage transformation rules: Candidate rules passing the nominal Wilcoxon screen were oriented toward increased activity and further prioritized using their median ΔpChEMBL, positive-pair fraction, support, and associated QED and SA changes. Rules with a median ΔpChEMBL <0.30 or a positive-pair fraction <0.50 were excluded. The retained transformations were converted to SMARTS, applied to compounds labelled active, and filtered for adverse physicochemical-property changes before construction of the focused library. The separate BH-FDR analysis described above was a retrospective sensitivity audit and did not alter these application-stage rule sets.

Reduced-data case studies and pretraining-overlap control

Reduced-data experiments were conducted for F2, AR, PARP1, and BRD4. ECFP6 fingerprints were used to construct structural clusters, and ~10% of the target compounds were selected by proportional sampling across clusters; the fixed subsets were reused for all method comparisons. MMP rule extraction, virtual-library construction, and target-specific CLM fine-tuning were restricted to these subsets. For each of the four targets (F2, AR, PARP1, and BRD4), molecules matching the complete corresponding target-specific dataset by canonical SMILES were removed from the ChEMBL 24 corpus before pretraining the CLMs used in the reduced-data experiments. This exclusion covered both the selected 10% subset and the remaining target-specific compounds. Generated molecules were evaluated using the corresponding scaffold-disjoint QSAR model selected within the development partition and fitted without the locked-test compounds. Each model was fixed before scoring the generated libraries and applied at a probability threshold of 0.5. The locked test partition was used only to report the reliability results in Table 1. Target-specific re-scoring models were retained as secondary post-generation screening endpoints.

Molecular docking and target-specific re-scoring

Docking was performed with AutoDock Vina (https://github.com/ccsb-scripps/AutoDock-Vina) using the following experimentally determined target structures: AR (PDB 2PNU), ERR1 (3K6P), BRD4 (4HXM), PKCeta (8FP1), F2 (1SL3), PARP1 (7KK4), HMGCR (2Q1L), KAT6A (8DD5), PDE5A (4XOZ), and VDR (5 V39). Receptor and ligand PDBQT files were prepared using standard AutoDockTools routines, including addition of polar hydrogen atoms and assignment of Gasteiger charges. A fixed target-specific binding pocket was used for all molecules and all generation methods for a given target. Vina exhaustiveness and number-of-modes settings were left at their software defaults. Docking was used only as a post-generation ranking procedure and was not included in CLM training.

For target-specific re-scoring, known active and inactive compounds were docked with the same protocol. Residue-level interaction terms within 5 Å of each docked ligand, together with the Vina docking output, were used as descriptors for random-forest classifiers. Activity labels used the same 500 nM threshold as the re-scoring QSAR models. Re-scoring models were evaluated by five-fold CV and applied at a probability threshold of 0.5. The re-scoring-predicted active rate was the proportion of successfully docked generated molecules assigned a probability of at least 0.5 by the corresponding target-specific re-scoring classifier.

Configuration-based ablation analysis

A configuration-based ablation study was performed for PARP1 and BRD4. A0 used direct CLM fine-tuning on the original active compounds. A1 used molecules produced using all retained MMP transformations, including activity-increasing and activity-decreasing directions. A2 used the original actives together with molecules produced by activity-increasing transformations but omitted final deployment-QSAR filtering. A3 represented the complete MMP2Mol workflow with the QSAR-filtered expanded library. All configurations used the same CLM architecture, 30-epoch fine-tuning schedule, checkpoint epoch, and sampling parameters. Three stochastic runs per target were included for every configuration; for A0 and A3, the first three prespecified run identifiers were used and any fourth completed run was excluded from the ablation summary. The evaluation was output-size matched at 100 raw epoch-30 sequences per run; larger outputs were deterministically subsampled using a fixed analysis seed (20260818) combined with the run identifier.

Ablation endpoints comprised validity, uniqueness, molecular novelty, ECFP4 internal diversity, Bemis–Murcko scaffold novelty, and the proportion of valid and unique molecules simultaneously assigned active and within the AD (active & AD) by the prespecified scaffold-based evaluation model. The activity threshold was 0.5, and the AD used an ECFP4 nearest-neighbour criterion whose target-specific threshold was the fifth percentile of leave-one-out nearest-neighbour similarities within the development partition. Values were summarized as mean and sample standard deviation across runs. Fine-tuning-library sizes and compositions were audited but were not forced to be equal; consequently, the analysis estimates configuration-level effects rather than strictly additive effects of isolated modules.

Evaluation

For the target-level structural comparison, ECFP4 fingerprints (radius 2, 2048 bits) were calculated for valid and unique generated molecules and the corresponding original active-compound set. The maximum Tanimoto similarity of each generated molecule to any reference active was determined, and Table 7 reports the mean of these nearest-reference similarities. Drug-likeness was assessed using QED, calculated with RDKit. QED ranges from 0 to 1, with higher values indicating a more favourable drug-like property profile. Synthetic accessibility was evaluated using the SA score, for which lower values indicate easier synthesis. Rule retention was assessed by matching the SMARTS patterns derived from the retained MMP transformations to generated structures.

For chemical-space visualization, equal-sized samples of 5000 valid and unique molecules from MMP2Mol and CLM were encoded as Morgan fingerprints and jointly embedded by t-SNE. For the docking-distribution visualization, all application experiments began from 10 000 generated molecules per target, and the 2000 molecules with the most favourable docking scores were displayed for each target; the known-compound comparison set comprised the 500 compounds with the highest pChEMBL values, or all available compounds when fewer than 500 were present. These subsets were used for visualization only and did not replace the full-library molecular-quality calculations.

Key Points

  • MMP2Mol integrates MMP analysis with chemical language models to enable ligand-based de novo drug design, effectively extracting and optimizing transformation rules from limited bioactivity data to guide generative models.

  • The framework constructs target-focused virtual libraries that improve predicted activity enrichment relative to direct CLM fine-tuning while maintaining competitive docking, drug-likeness, and synthetic-accessibility (SA) profiles.

  • MMP2Mol maintains substantial molecular diversity while generating structurally novel compounds with favourable computational estimates of synthetic accessibility and can recover known active molecules absent from target-specific fine-tuning.

  • Because generation does not require target protein structures, MMP2Mol is applicable across diverse ligand-defined targets and provides a practical framework for early-stage computational molecular design under data-scarce conditions.

Supplementary Material

MMP2Mol_SI_Final_bbag548

Contributor Information

Dekun Chen, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.

Hui Li, Xiangya School of Pharmaceutical Sciences, Central South University, No. 172 Tongzipo Road, Yuelu District, Changsha, Hunan 410013, China.

Jianwang Liu, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.

Yueping Jiang, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.

Dongsheng Cao, Xiangya School of Pharmaceutical Sciences, Central South University, No. 172 Tongzipo Road, Yuelu District, Changsha, Hunan 410013, China.

Shao Liu, Department of Pharmacy, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; Clinical Pharmacology Research Center of Hunan, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China; National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, No. 87 Xiangya Road, Kaifu District, Changsha, Hunan 410008, China.

Author contributions

Dekun Chen collected the data, developed the models, analyzed the data, and wrote the manuscript; Hui Li and Jianwang Liu helped interpret the results with constructive discussions; Yueping Jiang, Dongsheng Cao, and Shao Liu conceived and supervised the project, interpreted the results, and wrote the manuscript.

Conflicts of interest

There are no conflicts to declare.

Funding

None declared.

Data availability

The codes used for this research work are made publicly available in the GitHub repository [https://github.com/DKCCCChen/MMP2Mol].

References

  • 1. Hughes  JP, Rees  S, Kalindjian  SB  et al. Principles of early drug discovery. Br J Pharmacol  2011;162:1239–49. 10.1111/j.1476-5381.2010.01127.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Xiong  G, Wu  Z, Yi  J  et al. ADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties. Nucleic Acids Res  2021;49:W5–w14. 10.1093/nar/gkab255 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Cumming  JG, Davis  AM, Muresan  S  et al. Chemical predictive modelling to improve compound quality. Nat Rev Drug Discov  2013;12:948–62. 10.1038/nrd4128 [DOI] [PubMed] [Google Scholar]
  • 4. Gómez-Bombarelli  R, Wei  JN, Duvenaud  D  et al. Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent Sci  2018;4:268–76. 10.1021/acscentsci.7b00572 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Goodnow  RA. Hit and lead identification: integrated technology-based approaches. Drug Discov Today Technol  2006;3:367–75. 10.1016/j.ddtec.2006.12.009 [DOI] [Google Scholar]
  • 6. Neves  BJ, Braga  RC, Melo-Filho  CC  et al. QSAR-based virtual screening: advances and applications in drug discovery. Front Pharmacol  2018;9:1275. 10.3389/fphar.2018.01275 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Singh  S, Kaur  N, Gehlot  A. Application of artificial intelligence in drug design: a review. Comput Biol Med  2024;179:108810. 10.1016/j.compbiomed.2024.108810 [DOI] [PubMed] [Google Scholar]
  • 8. Gholap  AD, Uddin  MJ, Faiyazuddin  M  et al. Advances in artificial intelligence for drug delivery and development: a comprehensive review. Comput Biol Med  2024;178:108702. 10.1016/j.compbiomed.2024.108702 [DOI] [PubMed] [Google Scholar]
  • 9. Yang  X, Wang  Y, Byrne  R  et al. Concepts of artificial intelligence for computer-assisted drug discovery. Chem Rev  2019;119:10520–94. 10.1021/acs.chemrev.8b00728 [DOI] [PubMed] [Google Scholar]
  • 10. You  Y, Lai  X, Pan  Y  et al. Artificial intelligence in cancer target identification and drug discovery. Signal Transduct Target Ther  2022;7:156. 10.1038/s41392-022-00994-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Moingeon  P, Kuenemann  M, Guedj  M. Artificial intelligence-enhanced drug design and development: toward a computational precision medicine. Drug Discov Today  2022;27:215–22. 10.1016/j.drudis.2021.09.006 [DOI] [PubMed] [Google Scholar]
  • 12. Carini  C, Seyhan  AA. Tribulations and future opportunities for artificial intelligence in precision medicine. J Transl Med  2024;22:411. 10.1186/s12967-024-05067-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Vamathevan  J, Clark  D, Czodrowski  P  et al. Applications of machine learning in drug discovery and development. Nat Rev Drug Discov  2019;18:463–77. 10.1038/s41573-019-0024-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Dara  S, Dhamercherla  S, Jadav  SS  et al. Machine learning in drug discovery: a review. Artif Intell Rev  2022;55:1947–99. 10.1007/s10462-021-10058-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Lim  J, Ryu  S, Kim  JW  et al. Molecular generative model based on conditional variational autoencoder for de novo molecular design. J Cheminform  2018;10:31. 10.1186/s13321-018-0286-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Moret  M, Pachon Angona  I, Cotos  L  et al. Leveraging molecular structure and bioactivity with chemical language models for de novo drug design. Nat Commun  2023;14:114. 10.1038/s41467-022-35692-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Kong  W, Hu  Y, Zhang  J  et al. Application of SMILES-based molecular generative model in new drug design. Front Pharmacol  2022;13:1046524. 10.3389/fphar.2022.1046524 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Guo  X, Zhao  X, Lu  X  et al. A deep learning-driven discovery of berberine derivatives as novel antibacterial against multidrug-resistant helicobacter pylori. Signal Transduct Target Ther  2024;9:183. 10.1038/s41392-024-01895-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Li  Y, Zhang  L, Liu  Z. Multi-objective de novo drug design with conditional graph generative model. J Cheminform  2018;10:33. 10.1186/s13321-018-0287-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Lim  J, Hwang  SY, Moon  S  et al. Scaffold-based molecular design with a graph generative model. Chem Sci  2019;11:1153–64. 10.1039/c9sc04503a [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Zhu  H, Zhou  R, Cao  D  et al. A pharmacophore-guided deep learning approach for bioactive molecular generation. Nat Commun  2023;14:6234. 10.1038/s41467-023-41454-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Imrie  F, Hadfield  TE, Bradley  AR  et al. Deep generative design with 3D pharmacophoric constraints. Chem Sci  2021;12:14577–89. 10.1039/D1SC02436A [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Li  Y, Pei  J, Lai  L. Structure-based de novo drug design using 3D deep generative models. Chem Sci  2021;12:13664–75. 10.1039/D1SC04444C [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Bai  Q, Xu  T, Huang  J  et al. Geometric deep learning methods and applications in 3D structure-based drug design. Drug Discov Today  2024;29:104024. 10.1016/j.drudis.2024.104024 [DOI] [PubMed] [Google Scholar]
  • 25. Wu  K, Xia  Y, Deng  P  et al. TamGen: drug design with target-aware molecule generation through a chemical language model. Nat Commun  2024;15:9360. 10.1038/s41467-024-53632-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Yue  J, Peng  B, Chen  Y  et al. Unlocking comprehensive molecular design across all scenarios with large language model and unordered chemical language. Chem Sci  2024;15:13727–40. 10.1039/D4SC03744H [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Isigkeit  L, Hörmann  T, Schallmayer  E  et al. Automated design of multi-target ligands by generative deep learning. Nat Commun  2024;15:7946. 10.1038/s41467-024-52060-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Grisoni  F. Chemical language models for de novo drug design: challenges and opportunities. Curr Opin Struct Biol  2023;79:102527. 10.1016/j.sbi.2023.102527 [DOI] [PubMed] [Google Scholar]
  • 29. Gangwal  A, Ansari  A, Ahmad  I  et al. Current strategies to address data scarcity in artificial intelligence-based drug discovery: a comprehensive review. Comput Biol Med  2024;179:108734. 10.1016/j.compbiomed.2024.108734 [DOI] [PubMed] [Google Scholar]
  • 30. Kenny  PW, Sadowski  J. Structure modification in chemical databases. Chemoinform Drug Discov  2005;23:271–85. 10.1002/3527603743.ch11 [DOI] [Google Scholar]
  • 31. Leach  AG, Jones  HD, Cosgrove  DA  et al. Matched molecular pairs as a guide in the optimization of pharmaceutical properties; a study of aqueous solubility, plasma protein binding and oral exposure. J Med Chem  2006;49:6672–82. 10.1021/jm0605233 [DOI] [PubMed] [Google Scholar]
  • 32. Sheridan  RP, Hunt  P, Culberson  JC. Molecular transformations as a way of finding and exploiting consistent local QSAR. J Chem Inf Model  2006;46:180–92. 10.1021/ci0503208 [DOI] [PubMed] [Google Scholar]
  • 33. Griffen  E, Leach  AG, Robb  GR  et al. Matched molecular pairs as a medicinal chemistry tool. J Med Chem  2011;54:7739–50. 10.1021/jm200452d [DOI] [PubMed] [Google Scholar]
  • 34. Yang  Z, Shi  S, Fu  L  et al. Matched molecular pair analysis in drug discovery: methods and recent applications. J Med Chem  2023;66:4361–77. 10.1021/acs.jmedchem.2c01787 [DOI] [PubMed] [Google Scholar]
  • 35. Griffen  EJ, Dossetter  AG, Leach  AG  et al. Can we accelerate medicinal chemistry by augmenting the chemist with big data and artificial intelligence?  Drug Discov Today  2018;23:1373–84. 10.1016/j.drudis.2018.03.011 [DOI] [PubMed] [Google Scholar]
  • 36. Kramer  C, Fuchs  JE, Whitebread  S  et al. Matched molecular pair analysis: significance and the impact of experimental uncertainty. J Med Chem  2014;57:3786–802. 10.1021/jm500317a [DOI] [PubMed] [Google Scholar]
  • 37. Kramer  C, Ting  A, Zheng  H  et al. Learning medicinal chemistry absorption, distribution, metabolism, excretion, and toxicity (ADMET) rules from cross-company matched molecular pairs analysis (MMPA). J Med Chem  2018;61:3277–92. 10.1021/acs.jmedchem.7b00935 [DOI] [PubMed] [Google Scholar]
  • 38. Fu  L, Yang  ZY, Yang  ZJ  et al. QSAR-assisted-MMPA to expand chemical transformation space for lead optimization. Brief Bioinform  2021;22:bbaa374. 10.1093/bib/bbaa374 [DOI] [PubMed] [Google Scholar]
  • 39. Wang  D, Jin  J, Shi  G  et al. ADMET evaluation in drug discovery: 21. Application and industrial validation of machine learning algorithms for Caco-2 permeability prediction. J Cheminform  2025;17:3. 10.1186/s13321-025-00947-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Glyn  RJ, Pattison  G. Effects of replacing oxygenated functionality with fluorine on lipophilicity. J Med Chem  2021;64:10246–59. 10.1021/acs.jmedchem.1c00668 [DOI] [PubMed] [Google Scholar]
  • 41. Hu  X, Hu  Y, Vogt  M  et al. MMP-cliffs: systematic identification of activity cliffs on the basis of matched molecular pairs. J Chem Inf Model  2012;52:1138–45. 10.1021/ci3001138 [DOI] [PubMed] [Google Scholar]
  • 42. Dimova  D, Hu  Y, Bajorath  J. Matched molecular pair analysis of small molecule microarray data identifies promiscuity cliffs and reveals molecular origins of extreme compound promiscuity. J Med Chem  2012;55:10220–8. 10.1021/jm301292a [DOI] [PubMed] [Google Scholar]
  • 43. He  J, You  H, Sandström  E  et al. Molecular optimization by capturing chemist’s intuition using deep neural networks. J Cheminform  2021;13:26. 10.1186/s13321-021-00497-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Pan  B, Zhang  PZ, Pang  HW  et al. Retrieval-augmented foundation models for matched molecular pair transformations to recapitulate medicinal chemistry intuition. arXiv [Preprint]. 2026. arXiv:2602.16684. 10.48550/arXiv.2602.16684 [DOI]
  • 45. Moret  M, Friedrich  L, Grisoni  F  et al. Generative molecular design in low data regimes. Nat Mach Intell  2020;2:171–80. 10.1038/s42256-020-0160-y [DOI] [Google Scholar]
  • 46. Bickerton  GR, Paolini  GV, Besnard  J  et al. Quantifying the chemical beauty of drugs. Nat Chem  2012;4:90–8. 10.1038/nchem.1243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Ertl  P, Schuffenhauer  A. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. J Cheminform  2009;1:8. 10.1186/1758-2946-1-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Beck  JM, Springer  C. Quantitative structure-activity relationship models of chemical transformations from matched pairs analyses. J Chem Inf Model  2014;54:1226–34. 10.1021/ci500012n [DOI] [PubMed] [Google Scholar]
  • 49. Landrum  G. Rdkit documentation. Release  2013;1:4. [Google Scholar]
  • 50. Rogers  D, Hahn  M. Extended-connectivity fingerprints. J Chem Inf Model  2010;50:742–54. 10.1021/ci100050t [DOI] [PubMed] [Google Scholar]
  • 51. Cereto-Massagué  A, Ojeda  MJ, Valls  C  et al. Molecular fingerprint similarity search in virtual screening. Methods  2015;71:58–63. 10.1016/j.ymeth.2014.08.005 [DOI] [PubMed] [Google Scholar]
  • 52. Breiman  L. Random forests. Mach Learn  2001;45:5–32. 10.1023/A:1010933404324 [DOI] [Google Scholar]
  • 53. Chen  T, Guestrin  C. Xgboost: A scalable tree boosting system. In: Krishnapuram B, Shah M, Smola AJ, Aggarwal CC, Shen D, Rastogi R (eds). Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining. New York, NY, USA: Association for Computing Machinery, 2016, 785–794. 10.1145/2939672.2939785 [DOI]
  • 54. Huang  X, Cao  D-S, Xu  Q-S  et al. A novel tree kernel support vector machine classifier for modeling the relationship between bioactivity and molecular descriptors. Chemom Intell Lab Syst  2013;120:71–6. 10.1016/j.chemolab.2012.11.008 [DOI] [Google Scholar]
  • 55. Hearst  MA, Dumais  ST, Osuna  E  et al. Support vector machines. IEEE Intelligent Systems and their applications  1998;13:18–28. 10.1109/5254.708428 [DOI] [Google Scholar]
  • 56. Webb  GI, Keogh  E, Miikkulainen  R. Naïve Bayes. Encycl Mach Learn  2010;15:713–4. 10.1007/978-0-387-30164-8_576 [DOI] [Google Scholar]
  • 57. Cox  DR. The regression analysis of binary sequences. J R Stat Soc Ser B: Stat Methodol  1958;20:215–32. 10.1111/j.2517-6161.1958.tb00292.x [DOI] [Google Scholar]
  • 58. Rumelhart  DE, Hinton  GE, Williams  RJ. Learning representations by back-propagating errors. Nature  1986;323:533–6. 10.1038/323533a0 [DOI] [Google Scholar]
  • 59. Wang  JB, Cao  DS, Zhu  MF  et al. In silico evaluation of logD7. 4 and comparison with other prediction methods. J Chemom  2015;29:389–98. 10.1002/cem.2718 [DOI] [Google Scholar]
  • 60. Ranhotra  HS. The estrogen-related receptors in metabolism and cancer: newer insights. J Recept Signal Transduct Res  2018;38:95–100. 10.1080/10799893.2018.1456552 [DOI] [PubMed] [Google Scholar]
  • 61. Crevet  L, Vanacker  JM. Regulation of the expression of the estrogen related receptors (ERRs). Cell Mol Life Sci  2020;77:4573–9. 10.1007/s00018-020-03549-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Sharpe  LJ, Brown  AJ. Controlling cholesterol synthesis beyond 3-hydroxy-3-methylglutaryl-CoA reductase (HMGCR). J Biol Chem  2013;288:18707–15. 10.1074/jbc.R113.479808 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Jiang  SY, Li  H, Tang  JJ  et al. Discovery of a potent HMG-CoA reductase degrader that eliminates statin-induced reductase accumulation and lowers cholesterol. Nat Commun  2018;9:5138. 10.1038/s41467-018-07590-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Wiesel-Motiuk  N, Assaraf  YG. The key roles of the lysine acetyltransferases KAT6A and KAT6B in physiology and pathology. Drug Resist Updat  2020;53:100729. 10.1016/j.drup.2020.100729 [DOI] [PubMed] [Google Scholar]
  • 65. De Martinis  M, Allegra  A, Sirufo  MM  et al. Vitamin D deficiency, osteoporosis and effect on autoimmune diseases and hematopoiesis: a review. Int J Mol Sci  2021;22:8855. 10.3390/ijms22168855 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Basu  A. The enigmatic protein kinase C-eta. Cancers (Basel)  2019;11:214. 10.3390/cancers11020214 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Noordermeer  SM, van  Attikum  H. PARP inhibitor resistance: a tug-of-war in BRCA-mutated cells. Trends Cell Biol  2019;29:820–34. 10.1016/j.tcb.2019.07.008 [DOI] [PubMed] [Google Scholar]
  • 68. Ray Chaudhuri  A, Nussenzweig  A. The multifaceted roles of PARP1 in DNA repair and chromatin remodelling. Nat Rev Mol Cell Biol  2017;18:610–21. 10.1038/nrm.2017.53 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Wei  X, Zhou  F, Zhang  L. PARP1-DNA co-condensation: the driver of broken DNA repair. Signal Transduct Target Ther  2024;9:135. 10.1038/s41392-024-01832-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Donati  B, Lorenzini  E, Ciarrocchi  A. BRD4 and cancer: going beyond transcriptional regulation. Mol Cancer  2018;17:164. 10.1186/s12943-018-0915-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71. Jamroze  A, Chatta  G, Tang  DG. Androgen receptor (AR) heterogeneity in prostate cancer and therapy resistance. Cancer Lett  2021;518:1–9. 10.1016/j.canlet.2021.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Andersson  KE. PDE5 inhibitors - pharmacology and clinical applications 20 years after sildenafil discovery. Br J Pharmacol  2018;175:2554–65. 10.1111/bph.14205 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Das  A, Durrant  D, Salloum  FN  et al. PDE5 inhibitors as therapeutics for heart disease, diabetes and cancer. Pharmacol Ther  2015;147:12–21. 10.1016/j.pharmthera.2014.10.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74. Mann  KG, Brummel  K, Butenas  S. What is all that thrombin for?  J Thromb Haemost  2003;1:1504–14. 10.1046/j.1538-7836.2003.00298.x [DOI] [PubMed] [Google Scholar]
  • 75. Negrier  C, Shima  M, Hoffman  M. The central role of thrombin in bleeding disorders. Blood Rev  2019;38:100582. 10.1016/j.blre.2019.05.006 [DOI] [PubMed] [Google Scholar]
  • 76. Van der Maaten  L, Hinton  G. Visualizing data using t-SNE. J Mach Learn Res  2008;9:2579–2605. [Google Scholar]
  • 77. Wilcoxon  F. Individual Comparisons by Ranking Methods. Breakthroughs in Statistics: Methodology and Distribution. New York, NY, USA: Springer, 1992. 196–202. 10.1007/978-1-4612-4380-9_16 [DOI] [Google Scholar]
  • 78. Mervin  LH, Bulusu  KC, Leen  K  et al. Orthologue chemical space and its influence on target prediction. Bioinformatics  2017;34:72–9. 10.1093/bioinformatics/btx525 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79. Bosc  N, Atkinson  F, Felix  E  et al. Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery. J Cheminform  2019;11:4. 10.1186/s13321-018-0325-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80. Benjamini  Y, Hochberg  Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol  1995;57:289–300. 10.1111/j.2517-6161.1995.tb02031.x [DOI] [Google Scholar]
  • 81. Baell  JB, Holloway  GA. New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. J Med Chem  2010;53:2719–40. 10.1021/jm901137j [DOI] [PubMed] [Google Scholar]
  • 82. Brenk  R, Schipani  A, James  D  et al. Lessons learnt from assembling screening libraries for drug discovery for neglected diseases. ChemMedChem  2008;3:435–44. 10.1002/cmdc.200700139 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

MMP2Mol_SI_Final_bbag548

Data Availability Statement

The codes used for this research work are made publicly available in the GitHub repository [https://github.com/DKCCCChen/MMP2Mol].


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES