Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Feb 18;16:9772. doi: 10.1038/s41598-026-39960-z

Stacked ensemble learning and in-silico profiling reveal dual DPP-IV and SGLT2 inhibitors from Moringa oleifera metabolites

M K Letuku 1, M G Mohlala 1, P Appiah-Kubi 2, A Singh 2, Y B Nuapia 1,
PMCID: PMC13013813  PMID: 41708792

Abstract

Diabetes mellitus (DM) is a growing global health challenge, particularly in low-resource settings where access to effective therapies remains limited. Dual inhibition of dipeptidyl peptidase IV (DPP-IV) and sodium–glucose co-transporter 2 (SGLT2) offers a synergistic therapeutic strategy by enhancing insulin secretion and promoting glucose excretion. This study developed an integrated in silico framework combining stacked ensemble machine learning, molecular docking, MD Simulations, and ADMET profiling to identify dual DPP-IV/SGLT2 inhibitors from M. oleifera metabolites. Baseline models were generated using 110 algorithm–descriptor combinations per target, and stacking significantly improved predictive accuracy, achieving Matthews Correlation Coefficients of 0.968 (training) and 0.937 (testing) for DPP-IV and 0.968 (training) and 0.861 (testing) for SGLT2. Validation against FDA-approved inhibitors confirmed the models’ reliability and generalisability. LC–MS/MS profiling of M. oleifera revealed several metabolites with high predicted activity, among which vitexin, homoorientin, lariciresinol 4-O-β-D-glucopyranoside, and N,α-L-rhamnopyranosyl vincosamide showed strong binding affinities and favourable pharmacokinetic properties. MD Simulations analyses further position N,α-L-rhamnopyranosyl vincosamide as a potential dual DPP-IV/SGLT2 hit inhibitor. The findings highlight M. oleifera as a promising source of multitarget antidiabetic compounds and demonstrate the potential of stacked ensemble learning in accelerating natural product-based drug discovery.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-39960-z.

Keywords: M. oleifera, Stacked ensemble learning, DPP-IV, SGLT2, MD simulations, Molecular docking, MM/GBSA, Natural product drug discovery

Subject terms: Computational biology and bioinformatics, Drug discovery

Introduction

Diabetes mellitus (DM) is a major global health challenge, affecting over 537 million people worldwide, including approximately 24 million in Africa and 4.2 million in South Africa1,2. The disease is associated with significant morbidity and mortality, contributing to 6.7 million deaths in 2021, and its prevalence is projected to increase by 46% globally and by 129% in Africa by 20453,4. These trends underscore the urgent need for affordable, effective, and innovative therapeutic strategies, particularly in low-resource settings.

DM is characterised by chronic hyperglycaemia due to insufficient insulin production, impaired insulin action, or both5. Type 1 DM results from autoimmune destruction of pancreatic β cells, while Type 2 DM involves progressive β cell dysfunction and reduced insulin sensitivity1,6. Two key molecular targets involved in glucose homeostasis are dipeptidyl peptidase IV (DPP-IV) and sodium glucose co-transporter 2 (SGLT2). DPP-IV rapidly degrades incretin hormones such as glucagon-like peptide 1 (GLP-1), thereby diminishing insulin secretion, whereas SGLT2 mediates about 90% of renal glucose reabsorption7. Their inhibition has complementary effects: DPP-IV inhibition prolongs incretin action and enhances insulin secretion, while SGLT2 inhibition lowers the renal glucose threshold, promoting glucose excretion8,9. Consequently, dual inhibition of DPP-IV and SGLT2 represents a rational strategy to improve glycaemic control through synergistic mechanisms.

Several synthetic inhibitors targeting these proteins, such as sitagliptin, saxagliptin, and empagliflozin, have been approved globally and are clinically effective10,11. However, these drugs remain expensive, are not widely accessible in many African regions, and may be associated with adverse effects12. This has renewed interest in medicinal plants as affordable sources of bioactive compounds with antihyperglycemic potential. M. oleifera, widely used in African traditional medicine, contains diverse phytochemicals including flavonoids, alkaloids, and lignans that exhibit antidiabetic activity1316. Nevertheless, conventional screening of plant extracts is labour intensive and costly, and most studies on M. oleifera have focused narrowly on docking or ADMET profiling against single targets17,18.

Recent advances in machine learning (ML) provide powerful tools to accelerate virtual screening and activity prediction for both synthetic and natural compounds1922. Stacked ensemble learning, which integrates multiple base learners into a meta model, offers superior accuracy, calibration, and generalisability compared to single algorithms2325. Despite its success in other pharmacological domains, its application to systematic dual-target prediction for M. oleifera metabolites remains largely unexplored.

This study addresses this gap by developing a stacked ensemble ML framework for dual DPP-IV and SGLT2 inhibition, integrated with LC–MS/MS metabolite profiling, molecular docking, Molecular Dynamics Simulation, and ADMET analysis. By bridging traditional ethnopharmacology with modern computational methods, the work aims to prioritise M. oleifera metabolites with dual antidiabetic potential, thereby supporting the discovery of affordable, locally relevant therapeutic scaffolds.

Methodology

Dataset preparation for DPP-IV and SGLT2 inhibitors

Datasets for modelling DPP-IV and SGLT2 inhibitors were compiled from ChEMBL, BindingDB, and PubChem BioAssay. For DPP-IV, data were primarily retrieved from CHEMBL284, while SGLT2 bioactivity records were obtained from CHEMBL2977 and its associated entries (BindingDB and PubChem BioAssay). A uniform curation pipeline was applied to ensure data quality and chemical diversity. Only compounds with exact IC50 values, complete chemical structures, and unambiguous activity annotations were included. Records containing inequality qualifiers, missing SMILES, or incomplete metadata were excluded. Duplicate entries were removed using canonical SMILES, and chemical structures were standardised using RDKit, including salt stripping, tautomer canonicalisation, and charge normalisation. IC50 values were converted to pIC50 to harmonise activity representation.

Following curation, the final DPP-IV dataset comprised 4918 compounds, of which approximately 65% were synthetic molecules derived from medicinal-chemistry databases (e.g., ChEMBL and BindingDB), while the remaining 35% corresponded to phytochemicals curated from natural-product databases and peer-reviewed literature. Using a pIC50 threshold of 5.00; 2318 compounds were classified as active and 2600 as inactive. For SGLT2, the final curated dataset comprised 1584 compounds, of which approximately 70% were synthetic and 30% were phytochemicals. Synthetic molecules were predominantly sourced from ChEMBL, whereas phytochemical structures were compiled from natural product surveys and ethnopharmacological reports. Based on the same activity threshold, 823 compounds were labelled as active and 761 as inactive.

For each target, the curated datasets were randomly stratified into training (80%) and independent test (20%) sets while preserving the active–inactive class distribution. All feature selection, model training, and stratified ten-fold cross-validation procedures were performed exclusively on the training sets, whereas the independent test sets were reserved for a single final evaluation of the baseline and stacked ensemble models.

Chemical space analysis

Chemical space analysis was performed to characterise the physicochemical landscape of the SGLT2 and DPP-IV datasets and to identify features distinguishing active from inactive compounds. Canonical smiles were converted into molecular objects using RDKit (Python 3.13.7), and eight key descriptors related to Lipinski’s Rule of Five26 and molecular complexity were calculated: molecular weight, ALogP, number of hydrogen bond acceptors and donors, aromatic atom ratio, number of rings, rotatable bonds, and benzene-like rings. Descriptor values were standardised using z-score normalisation, and Lipinski’s thresholds were applied to assess drug likeness. Descriptor distributions were summarised using descriptive statistics and visualised through scatter plots and kernel density estimates in Seaborn and Matplotlib. MW ALogP plots illustrate the distribution of compounds within drug-like space. The nonparametric Mann–Whitney U test (p < 0.001) was used to assess differences between active and inactive compounds for each descriptor, analysed separately for the two targets.

Fingerprint generation

Molecular fingerprints were generated to numerically encode the chemical structures27. Ten types of descriptors, such as AtomPair2D (AP2D), Electrotopological State (EState), MACCS keys (MACCS), PubChem, Substructure (FP4), SubstructureCount (FP4C), Klekota-Roth (KR), Klekota-Roth Count (KRC), CDK Extended (CDKExt), and CDK Graph-Only (CDKGR) were selected to capture complementary structural and physicochemical information. Table S1 summarises the characteristics and key features of the molecular fingerprint descriptors employed in this study. All descriptor calculations were performed within a Python-based computational environment, following the preprocessing steps of salt removal, tautomer normalisation, and smiles verification. The resulting descriptor sets provided diverse chemical coverage, enabling robust model training.

Baseline model development

To develop predictive frameworks for DPP-IV and SGLT2 inhibition, baseline models were built using established machine learning algorithms and then integrated into a stacked ensemble. Before modeling, feature selection was performed to eliminate redundant descriptors using a Genetic Algorithm Subset Attribute Reduction (GA SAR) strategy (population = 10, generations = 10, crossover = 0.5, mutation = 0.2, random seed = 42), with twofold cross-validation to identify the most informative features. A comprehensive baseline modelling framework was implemented to evaluate a wide range of machine-learning algorithms and molecular descriptor types before constructing the stacked ensemble. For each target, 110 baseline models were generated by systematically combining algorithm classes and descriptor categories. Molecular descriptors included extended-connectivity fingerprints (ECFP), MACCS keys, PubChem fingerprints, topological fingerprints, and RDKit-derived physicochemical property descriptors. Multiple supervised machine-learning algorithms were explored, including AdaBoost (Ada), Bagging (Bag), Decision Trees (DT), Gradient Boosting (GB), k-Nearest Neighbours (KNN), Logistic Regression (LR), Naïve Bayes (NB), Multi-Layer Perceptron (MLP), Random Forests (RF), Support Vector Machines (SVM), X Gradient Boosting (XGBoost).

Full pairwise combinations of each descriptor type with each algorithm were constructed, yielding 110 algorithm–descriptor configurations; these are listed in Supplementary Table S3 to ensure transparency and reproducibility. Hyperparameter tuning was performed using stratified tenfold cross-validation on the training set. Baseline models were ranked by AUC, accuracy, sensitivity, and specificity, and the top-performing classifiers were selected as members of the final stacked ensemble.

Stacked ensemble development

Outputs from all baseline models were integrated into a stacked ensemble framework to enhance predictive accuracy and robustness. Each compound CCC was represented by a meta-feature vector constructed from the probability outputs of all baseline classifiers. Specifically, probabilities from 110 baseline models (10 molecular fingerprints × 11 machine learning algorithms) were concatenated to form a 110-dimensional meta-feature vector (FV(C)) representing the predictive behaviour of all base learners.

These meta-features were then used as input for the meta-classifier. A logistic regression (LR) model was employed as the meta-learner because of its strong generalisation capacity and interpretability in probabilistic settings. To further improve feature quality and reduce redundancy, a Genetic Algorithm–Structure–Activity Relationship (GA-SAR) feature optimisation step was applied within the stacking workflow, retaining the most informative subset of the original 110 meta-features.

graphic file with name d33e363.gif 1

where Inline graphic denotes the probability factor derived from the baseline model trained with the ith machine learning algorithm and the jth molecular encoding. As a result, FV(C) was converted into a 110-dimensional (D) feature vector. In this study, we employed a logistic regression (LR) meta-model to construct the stacked ensemble (named StackD) due to its inherent capability for feature design and extraction28. StackD was developed by training the LR meta-model on the 110-D meta-feature vectors to predict compound activity.

Hyperparameters of the LR meta-model were optimised via grid search using tenfold stratified cross-validation, with the Matthews Correlation Coefficient (MCC) as the selection criterion. The hyperparameter grid included regularisation strength C ∈ {0.01, 0.1, 1, 10, 100} and penalty types ∈ {ℓ2, ℓ1}, with appropriate solvers (liblinear or saga). The final model, designated StackD, was retrained on the complete training dataset using the best configuration and subsequently evaluated on an independent test set for accuracy, calibration, and generalisability.

Model evaluation metrics

To rigorously assess the predictive performance of both the baseline and the stacked ensemble (StackD) models, a comprehensive suite of evaluation metrics was used. These included accuracy (ACC), precision (positive predictive value), recall (sensitivity), specificity (true negative rate), F1-score, Matthew’s correlation coefficient (MCC), balanced accuracy (BACC), Cohen’s kappa (κ), and the area under the receiver operating characteristic curve (ROC AUC). Each metric provides complementary insights into different aspects of classification performance, allowing a multidimensional evaluation of model generalisation across diverse datasets. All metrics were computed using Python’s scikit-learn and imbalanced-learn libraries within a tenfold stratified cross-validation framework, as well as on an independent test set to ensure the robustness, fairness, and generalisability of the results. These metrics are described as follows:

graphic file with name d33e382.gif 2
graphic file with name d33e386.gif 3
graphic file with name d33e390.gif 4
graphic file with name d33e394.gif 5
graphic file with name d33e398.gif 6
graphic file with name d33e402.gif 7
graphic file with name d33e406.gif 8
graphic file with name d33e410.gif 9
graphic file with name d33e414.gif 10
graphic file with name d33e418.gif 11

External validation using approved antidiabetic drugs

External validation was used to assess the reliability and generalisability of the models. A benchmark set of eight FDA approved antidiabetic drugs with known mechanisms was compiled, including four DPP-IV inhibitors (sitagliptin, saxagliptin, linagliptin, alogliptin) and four SGLT2 inhibitors (dapagliflozin, empagliflozin, canagliflozin, ertugliflozin). These clinically relevant and structurally diverse compounds provided a rigorous test of model performance. To avoid bias, all benchmark compounds were excluded from training and cross-validation before model development, preventing data leakage and ensuring a genuine external test. The molecules were retrieved in smiles format and processed using the same standardisation pipeline as the training data, including tautomer normalisation, salt removal, and molecular fingerprint generation. The processed fingerprints were then evaluated using the final stackD classifiers for DPP-IV and SGLT2 inhibition, producing both binary activity predictions and probability scores for each compound.

Extraction of metabolites from M. oleifera leaves

The extraction of M. oleifera leaf phytochemicals was performed following the methodology described by Myoli et al.29, with minor modifications. Dried leaves were finely ground using a laboratory grinder and stored in airtight containers until analysis. A 2 g portion of the powdered sample was extracted with 20 mL of 80% methanol in 50 mL Falcon tubes. The samples were placed on a digital tube rotisserie rotator (Thermo Fisher Scientific, Johannesburg, South Africa) and agitated overnight at 150 rpm at room temperature to facilitate metabolite extraction. After extraction, the crude mixtures were centrifuged at 5678 g for 15 min using a benchtop fixed-angle centrifuge (Thermo Fisher Scientific, Johannesburg, South Africa). The resulting supernatants were collected, filtered through 0.22 µm syringe filters, and stored at 20 °C until LC MS/MS analysis.

LC–MS/MS profiling and metabolite identification

Methanolic extracts of M. oleifera leaves were analysed using LC–qTOF–MS on a Shimadzu LCMS-9030 platform following the method developed by Myoli et al.29. 3 μL aliquot of the extract was injected onto a Shim-pack C18 column (100 mm × 2.1 mm, 2.7 μm) maintained at 50 °C. Separation was achieved using a binary solvent system comprising 0.1% formic acid in Milli-Q water (A) and methanol with 0.1% formic acid (B) at a flow rate of 0.4 mL/min. The gradient started at 5% B (0–3 min), increased to 40% B (3–5 min), then to 95% B (5–12 min), was held until 18 min, and returned to 5% B (18–20 min), followed by a 3 min equilibration. Mass spectrometric detection was performed using electrospray ionisation in negative mode. Instrument settings included an interface voltage of 4.5 kV, an interface temperature of 300 °C, nebulising and drying gas flows of 3 L/min, a desolvation line temperature of 250 °C, a heat block temperature of 400 °C, and a detector voltage of 1.8 kV. Sodium iodide (NaI) was used for mass calibration. Data were acquired over an m/z range of 100–1200 in MS1 and MS2 modes using data-dependent acquisition with a 5000-count threshold. Fragmentation was induced by collision-induced dissociation with argon at 30 eV ± 5 eV.

Metabolite annotation and identification

Metabolite annotation was performed by integrating experimental spectral matching within in silico structural prediction. Raw LC-q-TOF–MS data were converted to mzML using MSConvert (ProteoWizard) and processed in MS-DIAL for peak picking, deconvolution, and alignment under the following conditions: MS1 tolerance 0.015 Da, MS1/MS2 tolerances 0.25/0.1 Da, minimum peak height 2000, and retention time tolerance 0.1 min. The processed files were exported in GNPS-compatible format and submitted to GNPS2 for classification-based molecular networking, using precursor and fragment tolerances of 0.25 Da, a cosine score ≥ 0.65, and ≥ 4 matched fragments.

Spectral library searches were performed against MassBank, ReSpect, and NIST, and the resulting molecular networks were visualised in Cytoscape (v3.8.2). Unmatched nodes were manually dereplicated using exact mass, isotopic distribution, and fragmentation patterns. To expand annotation coverage, SIRIUS (v6) with CSI:FingerID was used to predict molecular formulas and elucidate putative structures. Results were cross-referenced with PubChem, ChEBI, SUPNAT, DrugBank, and the Dictionary of Natural Products. All metabolite annotations followed the Metabolomics Standards Initiative (MSI) guidelines and were reported at Level 2 (putatively identified compounds) or Level 3 (putative compound classes).

Machine learning-QSAR virtual screening

Following the LC–MS/MS identification and annotation of metabolites in M. oleifera leaves, the resulting structures were converted to standardised SMILES strings and processed in PaDEL to compute the same panel of 10 molecular fingerprints used during model development. These descriptors were submitted to the optimised stackD classification model to predict DPP-IV and SGLT2 inhibitory activity. The meta-classifier generated probability scores indicating the likelihood of each phytochemical acting as an inhibitor, and compounds with predicted probabilities ≥ 0.5 were flagged as potential hits. The top-ranked phytochemicals were then subjected to structure-based docking to evaluate binding modes and affinities.

Molecular docking evaluation against DPP-IV and SGLT2

Protein and ligand structure preparations

The crystal structures of DPP-IV (PDB ID: 4N8D; Resolution: 1.65 Å) and SGLT2 (PDB ID: 8HEZ; Resolution: 2.8 Å) were sourced from the RCSB Protein Data Bank (https://www.rcsb.org/)30. The protein crystal structures were prepared using the Protein Preparation Wizard in Schrödinger Maestro, version 2023–2. This methodology encompassed assigning bond orders, incorporating hydrogen atoms, and reconstructing missing residues or atoms using the Prime module with the OPLS4 force field31. Co-crystallized water molecules were excluded from the analysis, and the protonation states of ionizable residues were established at a pH of 7.0, as determined by PROPKA. After optimizing the hydrogen bond network, the entire system was minimized until the root-mean-square deviation (RMSD) of heavy atoms reached a convergence threshold of 0.3 Å.

The three-dimensional (3D) chemical structures of N,α-L-rhamnopyranosyl vincosamide (PubChem ID: 71,717,770), Sitagliptin (PubChem ID: 4,369,359), Vitexin (PubChem ID: 5,280,441), Homoorientin (PubChem ID: 114,776), Lariciresinol 4-O-β-D-glucopyranoside (PubChem ID: 73,157,776) and Empagliflozin (PubChem ID: 11,949,646) were retrieved from the PubChem chemical database (https://pubchem.ncbi.nlm.nih.gov/)31. Ligand preparation was performed using the LigPrep module in Schrödinger Suite version 2023–2 (Schrödinger, LLC, New York, NY, USA). This process included generating low-energy 3D conformations, assigning protonation states at physiological pH (7.2 ± 0.2) using the Epik classic tool, and energy minimization using the OPLS4 force field. The protocol was configured to provide a single conformer for each ligand, and the generation of alternative tautomers was disabled.

Receptor Grid generation

The binding sites of DPP-IV (PDB ID: 4N8D) and SGLT2 (PDB ID: 8HEZ) were defined using the co-crystalized ligand coordinates of the retrieved proteins. The docking grid box was centred on the co-crystal ligand centroid, with inner dimensions of approximately 20 × 20 × 20 Å and an outer box measuring 40 Å. All critical active-site residues were covered, and no constraints were enforced throughout the docking process.

Molecular docking with induced-fit docking (IFD)

The Induced Fit Docking (IFD) protocol implemented in the Schrödinger Suite was used to account for receptor flexibility during ligand binding32. The multiple-step IFD workflow began with a preliminary soft-docking phase, in which the van der Waals radii for both receptor and ligand atoms were reduced by a factor of 0.5. For each resulting ligand conformation, the side chains of residues situated within a 5 Å proximity were remodeled and subsequently subjected to minimization using the Prime module. The top 20 scoring protein–ligand complexes derived from this refinement phase were subsequently progressed to a final, rigorous redocking step. This final stage involved docking each ligand into its corresponding induced-fit receptor conformation using Glide in Extra Precision (XP) mode.

Validation of docking protocol

The docking protocol was validated by redocking the original crystallized inhibitors into their respective binding pockets. Specifically, the co-crystallized ligands, dapagliflozin in SGLT2 (PDB ID: 8HEZ) and syn-7aa in DPP-IV (PDB ID: 4N8D), were removed from their binding sites, prepared, and subsequently redocked into the catalytic pocket using the above docking protocol.

Molecular dynamics simulations

All-atom molecular dynamics (MD) simulations were performed using AMBER20 to evaluate the conformational stability of the DPP-IV and SGLT2 protein systems. Protein parameters were assigned using the ff14SB force field33, and each system was solvated in explicit TIP3P water34 within a truncated octahedral box extending at least 14 Å from the solute. System neutrality and physiological ionic strength were achieved by adding Na⁺ and Cl⁻ ions to a final concentration of 0.1 M using a 1 Å placement grid. Ligand parameters were generated with GAFF235, with partial atomic charges derived using the AM1-BCC method36. Any missing force-field parameters were resolved using the parmchk2 module, and hydrogen atoms were added using LEaP module37.

All bonds involving hydrogen atoms were constrained using the SHAKE algorithm38, enabling a 2 fs integration time step. Long-range electrostatic interactions were treated using the particle-mesh Ewald (PME) method, with an initial nonbonded cutoff of 8 Å. Energy minimization was carried out in two stages: 5000 steps of steepest descent followed by 10,000 steps of conjugate gradient minimization. The systems were first equilibrated under constant volume (NVT) conditions at 100 K for 1 ns, with positional restraints of 100 kcal mol−1 Å−2 applied to all solute atoms. Temperature regulation was achieved using a Langevin thermostat39 with a collision frequency of 1 ps−1.

Subsequently, the systems were gradually heated to 310 K under constant pressure (NPT) conditions over 200 ps, followed by a 1 ns equilibration phase during which positional restraints were progressively released. Production MD simulations were then conducted for 120 ns under NPT conditions at 300 K without restraints. The Langevin thermostat and SHAKE constraints were maintained throughout, and a 12 Å nonbonded cutoff was used. Trajectory snapshots were saved every 20 ps for downstream analyses.

MM/GBSA Binding free energy calculations

End-point binding free energies (ΔGbind) of the SGLT2– and DPP-IV–ligand complexes were estimated by post-processing molecular dynamics trajectories using the molecular mechanics/generalized Born surface area (MM/GBSA) method, as implemented in the AMBER20 suite40,41. A total of 1200 snapshots were extracted at 100-ps intervals over the entire molecular dynamics trajectory. MM/GBSA calculations were carried out using the MMPBSA.py module with the mbondi3 atomic radii set, and a solvent probe radius of 2 Å was used for the estimation of the solvent-accessible surface area. Entropic contributions (TΔS) were not included because of computational cost. All energy values are reported in kcal/mol.

ADMET prediction

The pharmacokinetic characteristics of the bioactive compounds active against the SGLT2 and DPP-IV receptors were assessed using the SwissADME (http://www.swissadme.ch) and ADMETlab 2.0 (https://admetmesh.scbdd.com) web servers. The compounds included Vitexin (PubChem CID: 5,280,441), Homoorientin (PubChem CID: 114,776), Lariciresinol 4-O-β-D-glucopyranoside (PubChem CID: 11,972,395), and N,α-L-rhamnopyranosyl-vincosamide (PubChem CID: 71,717,770). Physicochemical properties such as molecular weight (MW), number of hydrogen bond donors (nHD) and acceptors (nHA), rotatable bonds (nRot), lipophilicity (cLog P), solubility, and topological polar surface (TPSA) area were analyzed. Drug-likeness was evaluated using Lipinski’s Rule of Five. Absorption-related parameters, including Caco-2 cell permeability (C2P) and human intestinal absorption (HIA), were examined, while excretion potential was estimated based on predicted clearance and biological half-life (T1/2).

Data analysis

All plots and data were generated and statistically analyzed using Origin data analysis software42. Results are presented as either averages or as mean ± standard deviation (SD). Docking complexes were visualized using DS Visualizer43, UCSF Chimera44, and Schrödinger Maestro version 2023–2.

Results and discussion

Chemical space analysis

In this section, we performed chemical space analysis to characterise the structural and physicochemical patterns that distinguish active and inactive compounds in the DPP-IV and SGLT2 datasets. Initially, the general chemical space was visualised as a function of molecular weight (MW) and lipophilicity (ALogP). Additionally, Ro5 descriptors were used to compare active and inactive compounds. Ro5 determines the drug likeness of compounds based on molecular properties including MW (< 500 Da), ALogP (< 5), nHAcc (< 10), and nHDon (< 5).

For the DPP-IV dataset, the MW and ALogP distribution is shown in Additional file 1: Fig. S1. Most compounds were clustered within the MW range of 200–500 Da and an ALogP between 1 and 5. Additional file 1: Fig. S2 presents the distribution of active and inactive compounds based on the Ro5 descriptors. Statistical analysis using the Mann–Whitney U test revealed significant differences (p < 0.001) between active and inactive compounds. Actives displayed higher MW (419.31 ± 89.58 Da) than inactives (376.55 ± 104.86 Da), lower ALogP (2.17 ± 1.42 vs. 2.31 ± 1.50; p < 0.001), more hydrogen bond acceptors (5.71 ± 1.97 vs. 4.72 ± 2.00; p < 0.01), and greater structural complexity with higher nCIC (3.52 ± 1.09 vs. 3.13 ± 1.26; p < 0.001) and nBnz (1.98 ± 1.14 vs. 1.75 ± 1.03; p = 2.18 × 10⁻12). No significant differences were observed for nHDon, ARR, or RBN.

Furthermore, molecular complexity descriptors ARR, nCIC, RBN, and nBnz were analysed to assess differences between active and inactive compounds. Additional file 1: Fig. S3 shows that active compounds generally exhibit lower ARR ratios, fewer rotatable bonds, and more benzene-like rings than inactive compounds, and these differences are statistically significant (p < 0.001). These results indicate that active DPP-IV inhibitors occupy chemical regions characterised by larger molecular frameworks, moderate lipophilicity, increased polarity, and enriched aromatic content. This structural profile aligns with previously reported structure–activity relationships, which show that increased molecular size and hydrogen bonding capacity are critical determinants of DPP-IV inhibitory potency. Polar heteroaromatic scaffolds enable multiple interactions within the catalytic pocket45,46, while rigid heteroaromatic cores support π-π stacking and shape complementarity47,48.

For the SGLT2 dataset, the MW and ALogP plots are shown in Additional file 1: Fig. S4. Most molecules fall within the MW range of 200–500 Da and ALogP between 1 and 5. Additional file 1: Fig. S5 presents the distributions of Ro5 descriptors for active and inactive compounds. Statistical analyses revealed significant differences between active and inactive groups for MW (p = 3.66 × 10−2), ALogP (p = 2.72 × 10−3), nHAcc (p = 6.50 × 10−14), ARR (p = 2.14 × 10−18), nCIC (p = 2.97 × 10−5), RBN (p = 1.88 × 10−5), and nBnz (p = 6.93 × 10−15). Molecular complexity analysis (Additional file 1: Fig. S6) showed that active SGLT2 compounds are characterised by moderate molecular size, balanced lipophilicity, reduced polarity, increased aromatic content, and greater structural rigidity compared to inactive molecules. These properties correspond to the C aryl glucoside scaffold architecture of marketed SGLT2 inhibitors. Moderate MW and controlled polarity are consistent with medicinal chemistry optimisation strategies to achieve adequate solubility and permeability while maintaining hydrogen bonding via conserved hydroxyl groups49,50. The lower number of rotatable bonds and higher aromatic ratios among actives indicate scaffold rigidity, which reduces entropic penalties upon binding and enhances affinity51,52. The presence of para- or meta-substituted aryl rings and heteroaryl glucoside frameworks reflects motifs found in clinically approved gliflozins, such as dapigliflozin and canagliflozin53.

Although structural preferences differ between the datasets, both occupy overlapping drug-like chemical space defined by Ro5. Natural product scaffolds such as flavonoids, lignans, and indole alkaloids often occupy this space, combining aromatic frameworks with tunable polarity. This shared landscape supports the application of stacked ensemble machine learning to prioritise novel multitarget DPP-IV SGLT2 inhibitors.

Baseline model development and evaluation

In this section, a comprehensive comparative analysis of 220 baseline machine-learning (ML) classifiers was performed using multiple algorithms and molecular descriptor sets to establish a robust foundation for dual-target prediction. For DPP-IV inhibition, 110 classifiers were constructed by combining 11 ML algorithms with 10 molecular fingerprint/descriptor types, while an equivalent 110 classifiers were developed for SGLT2 inhibition using the same algorithm–descriptor combinations. A complete overview of all evaluated ML algorithms, descriptor types, and their pairwise combinations is provided in Supplementary Table S3a for DPP-IV training dataset and Table 3b for DPP-IV testing dataset, and Table S4a for SGLT2 training dataset and Table 4b for SGLT2 testing dataset.

Table 3.

Comparison of baseline learners and stacked ensemble models for DPP-IV and SGLT2 (test set performance).

Target Model ACC (%) Balanced accuracy (%) ROC-AUC PR-AUC MCC
DPP-IV RF 90.5 89.8 0.912 0.884 0.785
XGBoost 91.2 90.4 0.918 0.889 0.792
SVM 89.8 88.6 0.904 0.871 0.760
StackD 93.9 93.9 0.946 0.935 0.937
SGLT2 RF 92.1 90.7 0.913 0.902 0.771
XGBoost 92.5 91.1 0.917 0.906 0.781
MLP 91.6 90.3 0.908 0.894 0.756
StackeD 95.9 91.5 0.935 0.926 0.861

Table 4.

Predicted dual-target activity of known antidiabetic drugs against DPP-IV and SGLT2.

Compounds DPP-IV SGLT2
Activity Prob Activity Prob
Sitagliptin 1 0.874 0 0.089
Saxagliptin 1 0.625 0 0.103
Linagliptin 1 0.924 0 0.110
Alogliptin 1 0.889 0 0.133
Dapagliflozin 0 0.108 1 0.923
Empagliflozin 0 0.126 1 0.886
Canagliflozin 0 0.084 1 0.895
Ertugliflozin 0 0.117 1 0.934

All classifiers were evaluated using stratified tenfold cross-validation followed by assessment on an independent test set. The classifier achieving the highest Matthews correlation coefficient (MCC) during cross-validation was selected as the best-performing baseline model.

For DPP-IV, the overall performance trends of all 110 classifiers are visualised in Fig. 1, with detailed quantitative metrics including accuracy, sensitivity, specificity, AUC, and MCC. The top five baseline classifiers were RF–KR, SVM–CDKEXT, Bag–CDKEXT, XGB–CDK, and XGB–KRC, achieving cross-validation MCC values of 0.97, 0.96, 0.94, 0.95, and 0.95, respectively, with corresponding AUC values consistently exceeding 0.96.

Fig. 1.

Fig. 1

MCC and AUC values of all 110 DPP-IV baseline models under cross-validation and independent test evaluations.

Among these, RF–KR emerged as the strongest single-descriptor baseline model, achieving a cross-validation MCC of 0.970, an accuracy of 97.2%, a sensitivity of 96.8%, a specificity of 97.6%, and an AUC of 0.970. When evaluated on the independent test set, RF–KR retained robust predictive performance with an MCC of 0.830, accuracy of 85.1%, sensitivity of 83.9%, specificity of 86.2%, and AUC of 0.850, confirming good generalisation beyond the training data.

For SGLT2 inhibition, the overall performance of the 110 baseline classifiers is illustrated in Fig. 2, with detailed evaluation metrics summarising the associations between ML algorithms and molecular descriptor sets. The top-performing baseline classifiers for SGLT2 were XGB–CDK, RF–CDK, SVM–CDK, Bag–CDKEXT, and RF–CDKEXT, achieving cross-validation MCC values of 0.95, 0.92, 0.95, 0.94, and 0.92, respectively. All five models consistently yielded AUC values exceeding 0.96, indicating strong discriminative ability. Among these, XGB–CDK emerged as the best single-descriptor baseline model, with a cross-validation MCC of 0.950, accuracy of 96.0%, sensitivity of 94.8%, specificity of 95.2%, and AUC of 0.970.

Fig. 2.

Fig. 2

MCC and AUC values of all 110 SGLT2 baseline models under cross-validation and independent test evaluations.

When evaluated on the independent test set, the XGB–CDK model maintained robust predictive performance, achieving an MCC of 0.830, accuracy of 84.2%, sensitivity of 85.1%, specificity of 83.0%, and an AUC of 0.860, confirming good generalisation capability for SGLT2 inhibition prediction.

These results indicate that both datasets are well structured and chemically separable, enabling reliable classification across multiple descriptor–algorithm combinations. Tree-based ensemble methods (RF, XGB, Bag) and SVM with CDK/CDKEXT descriptors consistently ranked among the top performers, reflecting their capacity to model complex non-linear structure–activity relationships in high-dimensional molecular spaces. This observation aligns with previous studies that highlighted the effectiveness of ensemble and kernel-based methods for cheminformatics applications51,54,55.

Furthermore, the interaction patterns between descriptors and algorithms indicate that KR, CDK, and CDKEXT fingerprints encode richer structural information relevant to DPP-IV and SGLT2 inhibition than simpler keys such as MACCS or PubChem. This is consistent with previous reports showing that extended circular and KR fingerprints outperform simpler substructure descriptors for complex pharmacological targets because they capture both local and global molecular environments51,56.

Taken together, these findings confirm the predictive strength of multiple single-algorithm classifiers and provide a strong foundation for integrating them into a stacked ensemble framework to enhance robustness and generalisability for dual-target activity prediction.

Stacked ensemble model development and evaluation

In this section, the performance of stacked ensemble learning models developed for DPP-IV and SGLT2 inhibition was systematically evaluated using stratified tenfold cross-validation and independent test sets. The evaluation focused on key classification and calibration metrics, enabling comparison with individual baseline models and assessment of generalisation capacity.

For the DPP-IV dataset, the stacked ensemble model achieved outstanding performance during cross-validation, with an accuracy of 98.4%, a balanced accuracy of 98.4%, an ROC-AUC of 0.998, an AUC of 0.901, and an MCC of 0.968 (Table 1). Probability calibration was excellent, as indicated by a Brier score of 0.099 and an expected calibration error (ECE) of 0.022. On the independent test set, the model maintained strong predictive performance, with an accuracy of 93.9%, balanced accuracy of 93.9%, ROC-AUC of 0.946, AUC of 0.935, and MCC of 0.937.

Table 1.

Performance metrics of the stacked ensemble model on the DPP-IV training and test datasets.

Performance metric CV-train Test
Acc (%) 0.984 0.939
Precision (%) 0.981 0.874
Recall/sensitivity (%) 0.983 0.850
Specificity (%) 0.983 0.850
F1-score (%) 0.982 0.862
Balanced accuracy (%) 0.984 0.939
ROC-AUC 0.998 0.946
AUC 0.901 0.935
MCC 0.968 0.937
Cohen’s κ 0.958 0.927
Brier score 0.099 0.092
ECE (10 bins) 0.022 0.033

Calibration remained stable (Brier score 0.092; ECE 0.033), confirming the reliability of probability estimates on unseen data. The confusion matrix (Fig. 3) shows balanced classification performance: 430 true positives, 484 true negatives, 62 false positives, and 76 false negatives, yielding sensitivity and specificity of 85.0% each, a precision of 87.4%, and an F1 score of 86.2%.

Fig. 3.

Fig. 3

Confusion matrices for the stacked DPP-IV classifier at a fixed decision threshold of 0.50.

Similarly, the SGLT2 stacked ensemble model showed excellent predictive power (Table 2). During cross-validation, it achieved an accuracy of 98.6%, a balanced accuracy of 99.7%, an ROC-AUC of 0.901, an AUC of 0.956, and an MCC of 0.968, with well-calibrated probabilities (Brier score 0.092; ECE 0.040). Independent test results further confirmed its robustness, with 95.9% accuracy, 91.5% balanced accuracy, ROC-AUC of 0.935, AUC of 0.926, and MCC of 0.861.

Table 2.

Performance metrics of the stacked ensemble model on the Sglt2 training and test datasets.

Performance metric CV-train Test
ACC (%) 98.60 95.90
Precision (%) 96.70 94.65
Recall/sensitivity (%) 98.68 91.73
Specificity (%) 92.42 89.53
F1-score (%) 97.68 96.17
Balanced accuracy (%) 0.997 0.915
ROC-AUC 0.901 0.935
PR-AUC (AP) 0.956 0.926
MCC 96.78 86.14
Cohen’s κ 0,979 0,928
Brier score 0.092 0.089
ECE (10 bins) 0.040 0.061

The confusion matrix (Fig. 4) shows accurate classification: 195 true positives, 82 true negatives, 20 false positives, and 10 false negatives, yielding a sensitivity of 91.7%, specificity of 89.5%, precision of 94.6%, and an F1 score of 96.2%.

Fig. 4.

Fig. 4

Confusion matrices for the stacked SGLT-2 classifier at a fixed decision threshold of 0.50.

As shown in Table 3, the stacked ensembles consistently outperformed the strongest baseline models, including Random Forest, XGB, and SVM. While RF and XGB were the top single learners, their MCC values were generally below 0.80 for both targets, compared with 0.937 and 0.861 for the DPP-IV and SGLT2 stacked ensembles, respectively. The improvement was particularly notable in PR-AUC, indicating superior enrichment of true actives compared with single algorithms.

Beyond discrimination, calibration behaviour was markedly better for the stacked ensembles than for individual models. This is especially relevant for virtual screening, where well-calibrated probabilities are essential for prioritising candidate compounds and setting rational thresholds. These findings reflect the theoretical advantages of stacking, which integrates complementary decision boundaries from diverse learners into a meta-model, reducing both variance and bias.

The results align with recent studies demonstrating the superiority of stacking in cheminformatics applications. Charoenkwan et al.57 and Zhang et al.58 reported that stacking improves both predictive accuracy and calibration compared with single learners.

Similarly, Wang et al.59 found that stacked architectures perform more robustly across chemically diverse datasets. Malik et al.60 showed that ensembles can achieve state-of-the-art MCC values with lower computational cost than deep learning models. Collectively, these findings confirm that stacked ensemble models provide a powerful, well-calibrated, and generalisable predictive framework for DPP-IV and SGLT2 inhibition, supporting their use in virtual screening and early-stage antidiabetic drug discovery.

External validation with approved drugs

A benchmark set of eight FDA-approved antidiabetic drugs was used to validate the predictive performance and specificity of the stacked ensemble models. This set included four DPP-IV inhibitors (sitagliptin, saxagliptin, linagliptin, and alogliptin) and four SGLT2 inhibitors (dapagliflozin, empagliflozin, canagliflozin, and ertugliflozin). The predictions are summarised in Table 4.

The DPP-IV stacked model correctly classified all four DPP-IV inhibitors as active, with prediction probabilities of 0.924 for linagliptin, 0.889 for alogliptin, 0.874 for sitagliptin and 0.625 for saxagliptin. All SGLT2 inhibitors were correctly labelled as inactive by the DPP-IV model, with probabilities below the activity threshold. For example, dapagliflozin scored 0.108 and canagliflozin 0.084. Conversely, the SGLT2 stacked model identified all four SGLT2 inhibitors as active with probabilities of 0.934 for ertugliflozin, 0.923 for dapagliflozin, 0.895 for canagliflozin and 0.886 for empagliflozin. It also correctly assigned low probability scores to the DPP-IV inhibitors, recording 0.089 for sitagliptin, 0.103 for saxagliptin, 0.110 for linagliptin and 0.133 for alogliptin, indicating inactivity on the SGLT2 target.

External validation on the benchmark dataset demonstrated the robustness and predictive selectivity of the stacked ensemble models. Each model achieved 100% sensitivity and specificity for its respective drug class, accurately distinguishing structurally distinct inhibitors without cross-target misclassification. For example, sitagliptin and linagliptin were correctly rejected by the SGLT2 model, while dapagliflozin and canagliflozin were accurately classified only by the SGLT2 model. This reflects precise target recognition based on substructural and pharmacophoric features unique to each enzyme. These findings align with previous reports showing that stacked ensembles outperform single classifiers by reducing variance, improving generalisation, and enhancing sensitivity and precision in chemogenomic tasks55,56,60,61. The perfect alignment with known drug activities and absence of cross-target confusion confirm the reliability and translational potential of the models for virtual screening and drug discovery.

LC–MS/MS metabolite profiling

Comprehensive phytochemical profiling of the methanolic extract of M. oleifera leaves was carried out using LC–ESI–QTOF–MS/MS. Analysis of the base peak chromatograms enabled the tentative identification of 44 secondary metabolites through spectral library matching (GNPS2) and in silico fragmentation (SIRIUS), supported by dereplication against MassBank, ReSpect, PubChem, the Dictionary of Natural Products, and literature sources. Table S5 presents the retention time, accurate mass, molecular formula, and key fragment ions for the annotated compounds. The annotated compounds, classified at Metabolomics Standards Initiative level 2, comprised amino acids, organic and phenolic acids, flavonoids, glucosinolates, alkaloids, lignans, and terpenoid glycosides. All screened phytochemicals have previously been reported in M. oleifera leaves in peer-reviewed studies59,62,63. All structures were converted into SMILES for cheminformatics analyses. Kaempferol-3-O-glucoside (astragalin) illustrates the workflow: its MS/MS spectrum showed a [M–H] ion at m/z 447.09 with a − 162 Da neutral loss, generating aglycone fragments at m/z 285.040 and 284.032.

The fragmentation behaviour of representative compounds provided diagnostic evidence for structural annotation. Kaempferol-3-O-glucoside showed characteristic C-ring decarbonylation ions, hydroxycinnamoylquinic acids produced quinic and phenyl-rich fragments, and (R)-linalyl β-vicianoside generated sugar-derived ions typical of terpenoid glycosides.

Untargeted LC–MS profiling revealed a chemically diverse metabolite composition, dominated by flavonoid glycosides, including vitexin, homoorientin, rutin, and isorhamnetin derivatives. These results align with recent LC–MS/MS studies that consistently report abundant flavonoid C- and O-glycosides in M. oleifera leaves6265. Complex acylated glycosides, including kaempferol and quercetin derivatives, were also detected, reflecting typical moringa phytochemistry66,67.

Beyond flavonoids, hydroxycinnamic acid derivatives (coumaroyl-, caffeoyl-, feruloylquinic acids), lignans (lariciresinol glycosides), and alkaloids (N-α-L-rhamnopyranosyl vincosamide) were identified, consistent with previous metabolomic surveys65. Overall, these findings reinforce the characteristic secondary metabolite profile of M. oleifera and provide a robust chemical basis for subsequent virtual screening and docking analyses.

Stacked machine learning virtual screening

A stacked ensemble approach was applied to evaluate the antidiabetic potential of 44 compounds tentatively annotated from the methanolic extract of M. oleifera leaves. Each compound, represented as a SMILES string, was processed through the stacked classification models to predict inhibitory activity against SGLT2 and DPP-IV.

The SGLT2 classification model identified five compounds as putative active inhibitors, with predicted probabilities ranging from 0.70 to 0.82 (Fig. 5). These included lignan glycosides such as lariciresinol 4-O-β-D-glucopyranoside (p = 0.78) and isolariciresinol 9′-O-β-D-glucoside, together with glycosylated flavonoids including vitexin (p = 0.79) and homoorientin (p = 0.74), as well as the indole alkaloid glycoside N,α-L-rhamnopyranosyl vincosamide (p = 0.79).

Fig. 5.

Fig. 5

Predicted active compounds (probability ≥ 0.5) from M. oleifera methanolic extract against SGLT2 and DPP-IV using the stacked ensemble model.

For DPP-IV, only N,α-L-rhamnopyranosyl-vincosamide was predicted to be active, with a probability of 0.75 (Fig. 5), making it the sole dual-target candidate within the screened library.

The stacked ensemble screening highlights the power of integrating machine learning with natural product diversity. Prioritised compounds encompassed both classical antidiabetic scaffolds and novel chemotypes. Predicted SGLT2 inhibitors included flavonoid glycosides such as vitexin and homoorientin, whose antidiabetic effects improve insulin sensitivity, lower fasting glucose, and protect β cells are well68,69. This confirms that the model effectively recognised established flavonoid scaffolds.

The identification of lignan glycosides such as lariciresinol-4-O-β-D-glucopyranoside broadens the predicted chemical space. Although transporter inhibition data are limited, their aglycones exhibit α-glucosidase inhibition and glucose-lowering effects70, suggesting their potential as pro scaffolds with improved solubility and stability. Notably, N,α-L-rhamnopyranosyl vincosamide emerged as a dual target candidate. Its indole rhamnoside structure is well suited for interactions within both DPP-IV and SGLT2 sites, indicating an underexplored multitarget potential for alkaloid glycosides.

The convergence between computational predictions and experimental evidence for flavonoid and lignan glycosides validates the modelling approach, while the discovery of N,α-L-rhamnopyranosyl-vincosamide’s dual target potential illustrates its exploratory power. These findings support prioritising flavonoid and lignan glycosides for transporter inhibition assays and evaluating vincosamide derivatives for dual activity opening new avenues for natural product-based antidiabetic drug discovery.

Molecular docking analysis of the predicted lead compounds

Molecular docking was conducted to elucidate the binding modes and interactions between the predicted M. oleifera metabolites and the dual antidiabetic targets SGLT2 and DPP-IV, thereby supporting the predictive reliability of the stacked ensemble machine-learning models. To validate the docking protocol, the co-crystalized ligands syn-7aa and dapagliflozin were redocked into the binding site of DPP-IV and SGLT2, respectively. The redocked complexes successfully reproduced their native poses, with RMSD values of 1.43 Å and 0.44 Å for syn-7aa and dapagliflozin, respectively (Figure S7). These deviations fall well below the commonly accepted 2 Å threshold, confirming the reliability of the docking protocol71.

The docking scores, poses, and interaction profiles of the phytochemicals vitexin, homoorientin, lariciresinol-4-O-β-D-glucopyranoside, and N,α-L-rhamnopyranosyl-vincosamide against the SGLT2, using Empagliflozin as the reference inhibitor, are presented in Table 5 and Figure S8. All compounds demonstrated strong predicted affinities for SGLT2, with IFD scores ranging from − 1141.21 to − 1146.18 kcal/mol, comparable to or exceeding that of Empagliflozin (− 1143.24 kcal/mol) (Table 5). Among the identified phytochemicals, the docking scores revealed that lariciresinol-4-O-β-D-glucopyranoside (IFD: -1145.79 kcal/mol) and N,α-L-rhamnopyranosyl-vincosamide (IFD: -1146.18 kcal/mol) exhibited binding energies surpassing the reference compound Empagliflozin (IFD: -1143.24 kcal/mol), indicating potentially superior inhibitory potential.

Table 5.

Molecular docking (IFD) scores and detailed interaction analysis for compounds against SGLT2.

Compound IFD score (kcal/mol) H-bond π–π stacked Hydrophobic interactions π–sigma
π–Alkyl
Sodium–glucose co-transporter 2 (SGLT2)
Empagliflozin -1143.24

Phe98, Glu99

Asn75, Lys320 Trp291, Ser287

Gln457

His80, Phe98 Leu84, Phe98, Ala102, Trp291, Tyr290, Val286, Leu283, Val95, Val157 Val95, Leu84
Vitexin -1141.21

Gly79, Glu99

Asn75, Lys321

Phe98, Trp291 Ser287, Glu457

Ser74, Lys154

His80 Leu84, Val157, Leu283, Val286, Trp291, Phe98, Ala102, Phe453 Val157
Homoorientin -1143.97 Glu99, Asp158, Tyr290,Thr153 Asp454 Tyr526

Phe45

Phe98

Val157, Ile76, Ile397, Phe98, Leu84, Val95, Leu274, Phe453, Ile456 Leu84, Val95
Lariciresinol-4-O-β-D-glucopyranoside -1145.79

Phe98,Trp291

Asn75,Gln457 Ser287

His80

Ala107, Trp291, Phe98,Leu274, Leu84,Leu283, Val157,Val286,

Phe453,

Tyr526

Val95

Leu84

Tyr290

N,α-L-rhamnopyranosyl-vincosamide -1146.18

Gln457, Asn75

Ser287, Trp291

Phe98, Tyr526

Asp454,

Asp273

Tyr290

Val195, Ala96

Phe98, Trp291

Ala102, Leu274

Ile456, Phe453

Lys154

Val157

His80

Leu84

The binding of Empagliflozin is stabilized within the SGLT2 binding pocket through multiple hydrogen bonds with key residues (Phe98, Glu99, Asn75, Lys320, Gln457, Trp291, Ser287), π–π stacking with His80 and Phe98, and extensive hydrophobic interactions (Leu84, Ala102, Trp291, Tyr290, Val286, Leu283, Val95, Val157). Additional π-alkyl contacts with Val95 and Leu84 enhance ligand stabilization. Lariciresinol-4-O-β-D-glucopyranoside engages in hydrogen bond formations with Phe98, Trp291, Asn75, Gln457, and Ser287, while π–π stacking with His80 enhances ligand stability. Extensive hydrophobic interactions (Ala107, Trp291, Phe98, Leu274, Leu84, Leu283, Val157, Val286, Phe453) and π-sigma/π-alkyl contacts (Tyr526, Val95) contributed to robust binding, positioning it as a potential SGLT2 inhibitor. On the other hand, N,α-L-rhamnopyranosyl-vincosamide makes extensive hydrogen bonds with Gln457, Asn75, Ser287, Trp291, Phe98, Tyr526, Asp454, and Asp273, indicating potential strong polar engagement. Aromatic π–π stacking with Tyr290, hydrophobic contacts with Val195, Ala96, Phe98, Trp291, Ala102, Leu274, Ile456, and Phe453, and π-alkyl interactions (Lys154, Val157, His80, and Leu84) further reinforced binding. This combination of polar, aromatic, and hydrophobic interactions highlights its potential as a hit compound towards SGLT2 inhibition.

Vitexin and homoorientin showed comparable binding affinities of (IFD: -1141.21 kcal/mol) and (IFD: -1143.97 kcal/mol), respectively. Vitexin was predicted to form hydrogen bonds Gly79, Glu99, Asn75, Lys321, Phe98, Trp291, Ser287, Glu457, Ser74, and Lys154, indicating strong polar interactions. Aromatic stabilization via π–π stacking with His80 and hydrophobic contacts with Leu84, Val157, Leu283, Val286, Trp291, Phe98, Ala102, and Phe453 were observed. Polar interactions of Homoorientin involved Glu99, Asp158, Tyr290, Thr153, Asp454, and Tyr526, while aromatic π–π stacking with Phe45 and Phe98 and hydrophobic contacts with Val157, Ile76, Ile397, Phe98, Leu84, Val95, Leu274, Phe453, and Ile456 contributed to overall stabilization. Although vitexin and homoorientin displayed slightly lower binding affinity, they remain promising hits for further structural optimization.

For DPP-IV, the docking analysis of the reference inhibitor sitagliptin showed IFD score of –1607.67 kcal/mol (Table 6, Figure S9). Sitagliptin is predicted to form a single stabilizing hydrogen bond with Arg669 and to engage in π–π stacking with Phe357, supported by hydrophobic interactions with Tyr666, Tyr547, and Tyr662. In contrast, N,α-L-rhamnopyranosyl vincosamide was predicted to bind the DPP-IV active site, with a slightly favourable IFD score of –1611.09 kcal/mol (Table 6, Figure S9). N,α-L-rhamnopyranosyl vincosamide engages in multiple hydrogen bonds with Phe357, Arg358, Glu361, Glu205, and Glu206, indicating deeper anchoring within the catalytic pocket. It also establishes dual π–π stacking interactions with Phe357 and Tyr547 and engages an expanded hydrophobic environment, including Cys551, Tyr585, Tyr547, Tyr662, and Pro359. Additional π–sigma interactions with Phe357 further reinforce the stability of the complex. Taken together, the induced fit docking findings suggest that N,α-L-rhamnopyranosyl vincosamide is a structurally compatible and energetically favorable binder of DPP-IV, with predicted binding affinity and interaction patterns comparable to those of sitagliptin, supporting its potential as a promising hit compound.

Table 6.

Molecular docking (IFD) scores and detailed interaction analysis for the compounds against DPP-IV.

Compound IFD score (kcal/mol) H-bond π–π stacked Hydrophobic interactions π–sigma
π–Alkyl
Dipeptidyl peptidase IV (DPP-IV)
Sitagliptin -1607.67 Arg669 Phe357 Phe357, Tyr666, Tyr547, Tyr662, Tyr634, Trp659, Tyr631, Val656, Pro655, Ala654, Val 711, Tyr666
N,α-L-rhamnopyranosyl-vincosamide -1611.09 Phe357, Arg358, Glu361, Glu205, Glu206 Tyr547, Phe357 Phe357,Cys551, Tyr585, Tyr547, Tyr662, Tyr666, Pro359 Phe357

Molecular dynamics (MD) simulation

Following molecular docking, molecular dynamics (MD) simulations were performed to assess the structural stability and conformational behavior of the SGLT2– and DPP-IV ligand complexes. Structural and dynamic descriptors, including root mean square deviation (RMSD), root mean square fluctuation (RMSF), solvent-accessible surface area (SASA), and radius of gyration (RoG), were systematically evaluated. These analyses were used to monitor system equilibration, residue-level flexibility, global compactness, and the overall structural integrity and binding stability of the protein–ligand complexes throughout the simulations.

Structural stability of SGLT2- and DPP-IV-ligand complexes

Backbone RMSD trajectory analysis of SGLT2 indicated that all systems reached equilibrium within 20–40 ns and remained stable throughout the simulation (Fig. 6A, Table S6). The SGLT2–empagliflozin complex exhibited one of the lowest average RMSD values (1.72 ± 0.25 Å), confirming its well-established stabilizing effect on the protein. The investigated ligand-bound systems had slightly higher RMSD values (1.88–1.90 Å) than those of the reference complex, except for Vitexin, which had a lower RMSD (1.70 ± 0.26 Å). These results indicate that the tested ligands maintain structural stability comparable to empagliflozin without inducing major conformational perturbations. Furthermore, the apo SGLT2 system exhibited an average RMSD of 1.72 ± 0.16 Å, suggesting that ligand binding does not cause large-scale structural deviations in SGLT2.

Fig. 6.

Fig. 6

Post–molecular dynamics simulation analysis of SGLT2 Apo and ligand-bound systems over 120 ns. (A) Backbone RMSD as a function of simulation time, (B) Radius of gyration (RoG) profiles showing global compactness, (C) Residue-wise RMSF indicating local flexibility, and (D) Solvent-accessible surface area (SASA) variations during the simulation.

In the DPP-IV systems, backbone RMSD profiles (Fig. 7A, Table S7) indicate that all systems reached equilibrium within the first ~ 30 ns and remained stable thereafter. The DPP-IV–sitagliptin complex exhibited the lowest average RMSD (1.60 ± 0.24 Å), reflecting a highly stable protein–ligand complex. The DPP-IV–N,α-L-rhamnopyranosyl vincosamide complex showed a moderately higher RMSD (2.04 ± 0.36 Å), suggesting increased conformational flexibility relative to the reference inhibitor while remaining within an acceptable stability range for a well-equilibrated system. Notably, the Apo DPP-IV system showed RMSD values comparable to those of the sitagliptin-bound complex (1.60 ± 0.16 Å), indicating that ligand binding does not induce large-scale destabilization of the DPP-IV backbone.

Fig. 7.

Fig. 7

Molecular dynamics simulation analyses of DPP-IV in the Apo state and in complex with sitagliptin (SITAG) and the test ligand (NL) over 120 ns. (A) Backbone RMSD versus time, (B) Radius of gyration (RoG) profiles, (C) Per-residue RMSF, and (D) Solvent-accessible surface area (SASA) as a function of simulation time.

Structural compactness analysis of SGLT2- and DPP-IV-ligand complexes

The compactness of the SGLT2 protein upon ligand binding was evaluated using radius of gyration (RoG) analysis (Fig. 6B, Table S6). All SGLT2 systems maintained stable compactness throughout the simulation, with average RoG values ranging from 23.7 to 23.9 Å, indicating preservation of the global fold. Most ligand-bound complexes exhibited slightly reduced RoG fluctuations relative to the Apo protein (23.83 Å), suggesting ligand-induced structural compactness. An exception was observed for SGLT2–Vitexin (23.85 Å), which showed a marginally higher RoG value. Notably, SGLT2–N,α-L-rhamnopyranosyl vincosamide (23.72 Å) and SGLT2–lariciresinol 4-O-β-D-glucopyranoside (23.76 Å) displayed RoG profiles comparable to the reference inhibitor empagliflozin (23.77 Å), reflecting similar global folding behavior and compactness.

RoG profiles for the DPP-IV systems (Fig. 7B, Table S7) also revealed minimal fluctuations, particularly after the first 50 ns of simulation. Among these, the DPP-IV–sitagliptin complex maintained the lowest average RoG (26.94 Å), indicating a more compact protein conformation compared to DPP-IV-N,α-L-rhamnopyranosyl vincosamide (27.07 Å) and the Apo protein (27.14 Å). These results suggest that ligand binding enhances structural compactness in DPP-IV.

Collectively, the RoG analyses demonstrate that ligand binding preserves the global compactness of both SGLT2 and DPP-IV proteins. The observed ligand-induced stabilization is consistent with the RMSD findings and further supports the formation of stable protein–ligand complexes during the simulations.

Residue-level flexibility analysis of SGLT2- and DPP-IV-ligand complexes

To evaluate protein flexibility upon ligand binding, residue-wise root mean square fluctuation (RMSF) analysis was performed for both SGLT2 and DPP-IV systems. For SGLT2, RMSF analysis (Fig. 6C, Table S6) revealed that most residues exhibited fluctuations below 1.5 Å, indicating overall structural rigidity during the simulation. Higher flexibility was primarily confined to loop and terminal regions, including Gly235–Ile251, His555–Ser581, Leu611–Gly621, and Glu631–Glu641. Importantly, residues lining the SGLT2 binding pocket—Asn75–Glu99, Tyr150–Phe160, Leu280–Met325, and Gly450–Ser460 showed reduced fluctuations in all ligand-bound systems compared to the apo protein. This reduction highlights the stabilizing influence of ligand binding on functionally relevant regions of SGLT2.

For DPP-IV, RMSF profiles (Fig. 7C, Table S7) indicated that most residues fluctuated below 1.5 Å, reflecting preserved structural integrity. Elevated flexibility was mainly observed in loop regions, particularly around Asp230–Gly260. In contrast, residues comprising the active site remained comparatively stable in both ligand-bound systems. Notably, the DPP-IV–N,α-L-rhamnopyranosyl vincosamide complex exhibited RMSF behavior comparable to the reference DPP-IV–sitagliptin complex, indicating that the test ligand effectively stabilizes the active-site architecture.

Overall, RMSF analyses demonstrate that ligand binding reduces local flexibility within key functional regions of both SGLT2 and DPP-IV while preserving global protein dynamics, further supporting the formation of stable protein–ligand complexes.

Solvent exposure and surface area analysis of SGLT2- and DPP-IV-ligand complexes

The solvent-accessible surface area (SASA) was used to assess protein exposure to the surrounding solvent and to infer ligand-induced conformational effects. Variations in SASA provide insight into changes in protein folding, structural stability, and solvent interactions that accompany complex formation72. SASA analysis for SGLT2 (Fig. 6D, Table S6) demonstrated stable solvent exposure across all systems, with ligand-bound complexes showing slightly lower SASA values (24,482.08–24,925.69 Å2) relative to the Apo protein (25,157.35 Å2). This reduction suggests efficient burial of hydrophobic surfaces upon ligand binding, consistent with tight ligand accommodation within the binding cavity.

For DPPIV, SASA analysis (Fig. 7D, Table S7) showed that DPP-IV–Sitagliptin had the lowest solvent exposure (29,227.55 Å2), followed by DPP-IV, N,α-L-rhamnopyranosyl vincosamide (29,366.83 Å2) and the Apo form (29,746.69 Å2). This indicates increased surface burial upon ligand binding, with a slightly greater effect for Sitagliptin.

Number of hydrogen-bond interaction analysis of SGLT2- and DPP-IV-ligand complexes

Hydrogen-bond analysis (Table S6) revealed that SGLT2–Empagliflozin maintained an average number of 4.46 ± 1.00 intermolecular hydrogen bonds, serving as a benchmark for interaction stability. Among the investigated ligands, particularly SGLT2–Vitexin (4.25 ± 1.18) exhibited comparable number of hydrogen bonds to Empagliflozin, while lower intermolecular hydrogen bonds observed in SGLT2–Homoorientin (3.86 ± 1.14) and SGLT2–Lariciresinol 4-O-β-D-glucopyranoside (2.81 ± 1.00). However, SGLT2–N,α-L-rhamnopyranosyl vincosamide displayed the highest hydrogen-bond persistence (5.13 ± 1.38), indicating a stable hydrogen-bonding network among the potential hits and contributing to enhanced binding stability.

For DPP-IV, hydrogen bond analysis (Table S7) revealed that DPP-IV–N,α-L-rhamnopyranosyl vincosamide formed a higher average number of hydrogen bonds (3.33 ± 2.19) compared to the reference DPP-IV–Sitagliptin complex (1.90 ± 0.38). This increase suggests that N,α-L-rhamnopyranosyl vincosamide engages in more frequent or transient polar interactions with the binding site. However, despite the higher hydrogen-bond count, the increased RMSD and SASA values suggest that these interactions may be less optimally oriented or more dynamic than those formed by sitagliptin.

In summary, the post-MD analyses collectively demonstrate that ligand binding does not induce deleterious structural perturbations in either protein. Instead, most ligand-bound systems preserved the targets’ native fold and dynamic equilibrium, underscoring the reliability of the docking-derived binding modes and the suitability of these ligands for further investigation.

Among the evaluated compounds, N,α-L-rhamnopyranosyl vincosamide emerged as a particularly compelling candidate due to its consistent and favorable dynamic behavior across both targets. In the SGLT2 system, this ligand exhibited backbone stability, global compactness, and residue-level rigidity comparable to those of the reference inhibitor, empagliflozin. Notably, it displayed the highest persistence of number intermolecular hydrogen bonds among the screened ligands, highlighting the formation of a robust and well-maintained interaction network within the SGLT2 binding pocket. The reduction in RMSF values for key binding-site residues, together with slightly decreased SASA and stable RoG profiles, further indicates efficient ligand accommodation and sustained stabilization of functionally relevant regions.

In the DPP-IV system, N,α-L-rhamnopyranosyl vincosamide also formed a stable complex, maintaining acceptable backbone RMSD values and preserving the protein’s overall compactness. Although its conformational fluctuations and solvent exposure were marginally higher than those observed for the reference inhibitor sitagliptin, these parameters remained within the range of a well-equilibrated system. Importantly, the ligand had a higher average number of hydrogen bonds than sitagliptin, suggesting greater polar engagement with the DPP-IV active site. The RMSF analysis further revealed effective stabilization of catalytic residues, indicating that, despite increased dynamism, the ligand can maintain the integrity of the active-site architecture.

MM/GBSA binding free energy estimations

Binding free energy estimation methods derived from molecular dynamics trajectories, such as MM/PBSA and MM/GBSA, are widely used in computational drug discovery workflows due to their robust predictive accuracy, relatively low energy consumption, and high computational efficiency. Thus, the integration of molecular docking with MM/GBSA calculations enhances the identification of promising lead compounds7375. To further elucidate the binding affinity and energetic determinants governing ligand–protein interactions, molecular docking results were complemented with MM/GBSA binding free energy calculations for SGLT2 (Table 7) and DPP-IV complexes (Table 8).

Table 7.

Predicted MM/GBSA free energies (kcal/mol) analysis of the identified compounds and the referenced drug in complex with SGLT2.

Energy components Complexes
114,776 11,949,646 5,280,441 71,717,770 73,157,776
ΔvDWaals -45.27 ± 5.39 -58.71 ± 3.64 -45.64 ± 4.32 -70.53 ± 5.94 -63.74 ± 3.62
ΔELECT -66.34 ± 15.32 -57.18 ± 7.24 -54.87 ± 9.08 -69.05 ± 13.08 -39.65 ± 5.06
ΔEGB 65.41 ± 7.93 49.04 ± 4.50 53.96 ± 4.98 63.65 ± 8.32 43.72 ± 3.56
ΔESURF -6.74 ± 0.42 -8.40 ± 0.15 -7.05 ± 0.33 -9.99 ± 0.40 -9.02 ± 0.19
ΔGGAS -111.61 ± 13.03 -115.88 ± 6.41 -100.50 ± 9.58 -139.58 ± 12.63 103.40 ± 4.54
ΔGSOLV 58.67 ± 7.92 40.64 ± 4.53 46.91 ± 5.01 53.65 ± 8.17 34.70 ± 3.56
ΔGTOTAL -52.95 ± 6.91 -75.24 ± 3.78 -53.59 ± 8.01 -85.92 ± 6.87 -68.69 ± 3.84

ΔvDWaals = Van der Waal energy; ΔELECT = Electrostatic energy; ΔEGB = Polar solvation energy; ΔESURF = Non-polar solvation energy; ΔGGAS = Net gas phase energy; ΔGSOLV = Net solvation energy; Homoorientin (114,776) ; Empagliflozin (11,949,646) ; N, α-L-rhamnopyranosyl vincosamide (71,717,770) ; Vitexin (5,280,441) ; Lariciresinol 4-O-β-D-glucopyranoside (73,157,776).

Table 8.

Predicted MM/GBSA free energies (kcal/mol) analysis of the two selected ligands in complex with DPP-IV.

Energy components Complexes
4,369,359 71,717,770
ΔvDWaals -34.78 ± 3.58 -36.01 ± 4.85
ΔELECT -366.33 ± 16.77 -76.13 ± 24.61
ΔEGB 370.58 ± 15.60 85.85 ± 14.06
ΔESURF -5.48 ± 0.27 -5.45 ± 0.76
ΔGGAS -401.11 ± 17.33 -112.15 ± 22.79
ΔGSOLV 365.10 ± 15.52 80.41 ± 13.47
ΔGTOTAL -36.01 ± 4.70 -31.74 ± 10.74

ΔvDWaals = Van der Waal energy; ΔELECT = Electrostatic energy; ΔEGB = Polar solvation energy; ΔESURF = Non-polar solvation energy; ΔGGAS = Net gas phase energy; ΔGSOLV = Net solvation energy; N, α-L-rhamnopyranosyl vincosamide (71,717,770); Sitagliptin (4,369,359).

The estimated total binding free energies (ΔGtotal) in SGLT2-Ligand complexes indicate that all ligands bind favorably to SGLT2, with N,α-L-rhamnopyranosyl vincosamide exhibiting the strongest affinity (− 85.92 ± 6.87 kcal/mol), followed by Empagliflozin (− 75.24 ± 3.78 kcal/mol), Lariciresinol 4-O-β-D-glucopyranoside (− 68.69 ± 3.84 kcal/mol), Vitexin (− 53.59 ± 8.01 kcal/mol), and Homoorientin (− 52.95 ± 6.91 kcal/mol) (Table 7). Across all SGLT2 complexes, the binding process is dominated by highly favorable van der Waals (ΔvDWaals) and electrostatic contributions (ΔELECT), which together account for the strongly negative gas-phase interaction energies (ΔGGAS). In particular, N,α-L-rhamnopyranosyl vincosamide displayed the most favorable van der Waals contribution (− 70.53 ± 5.94 kcal/mol). Electrostatic interactions also played a significant role, especially for ligands N,α-L-rhamnopyranosyl vincosamide and Homoorientin, which showed markedly favorable electrostatic energy values (− 69.05 ± 13.08 and − 66.34 ± 15.32 kcal/mol, respectively). As expected, the unfavorable solvation energies (ΔGSOLV) in all the systems were countered by the favorable gas-phase energy interactions (ΔGGAS).

For DPPIV complexes, the total binding free energies (ΔGtotal) reveal that both compounds bind favorably to the target protein, with Sitagliptin exhibiting a marginally stronger binding affinity (− 36.01 ± 4.70 kcal/mol) compared to N,α-L-rhamnopyranosyl vincosamide (− 31.74 ± 10.74 kcal/mol) (Table 8). Binding of Sitagliptin was dominated by a highly favorable electrostatic contribution, which was largely offset by a substantial polar solvation penalty, suggesting extensive charge-mediated interactions. In contrast, N, α-L-rhamnopyranosyl vincosamide showed more moderate electrostatic interactions and lower desolvation costs, reflecting a more balanced energetic profile. Comparable van der Waals and favorable non-polar solvation energies for both ligands highlight the contribution of hydrophobic interactions to DPP-IV binding.

Residue-wise binding energy decomposition and hydrogen bond occupancy analysis of SGLT2– and DPP-IV–ligand complexes

Hydrogen-bond occupancy analysis across the MD trajectories showed that SGLT2–ligand complexes are stabilized by persistent interaction networks within the substrate-binding pocket (Table S8). The Empagliflozin complex exhibited the highest hydrogen-bond persistence, with interactions dominated by Glu99 (98.78%), Phe98 (94.18%), and Gln457 (88.51%), indicative of a highly stable binding mode. Similarly, Vitexin similarly formed persistent hydrogen bonds with Glu99 (88.12%), Phe98 (87.18%), Ser287 (78.55%), Ser393 (41.04%), Trp291 (40.74) and Lys321 (10.33%). In contrast, N-α-L-rhamnopyranosyl vincosamide exhibited distributed hydrogen-bond occupancies across multiple residues such as Glu99 (72.35%), Phe453 (47.26%), His80 (39.71%), Ser287 (32.99%), Trp291 (29.21%), Asn75 (22.53%), Gln457 (19.46%), Tyr526 (11.95%) and Phe98 (10.18%). The Homoorientin complex displayed moderately high-occupancy hydrogen bonds, mainly with Asp158 (64.13%), Asn75 (64.06%), Asp454 (57.13%), Tyr290 (18.52%), and Lys154 (13.43%). Similarly, Lariciresinol 4-O-β-D-glucopyranoside formed limited but notable interactions with Glu99 (81.95%), Ser287 (61.31%), Gln457 (55.70%), and Ser460 (10.20%).

The per-residue binding energy decomposition for SGLT2 (Table S9) further highlight Phe98 and Glu99 contributing the most favorable energies across complexes, with maximal stabilization observed in the Empagliflozin system (− 4.41 ± 0.65 and − 4.11 ± 1.52 kJ/mol, respectively). In the empagliflozin-bound complex, strong contributions from Gln457 (− 4.83 ± 1.03 kJ/mol), together with Phe98 and Glu99, reflected a cooperative interaction network. Additional residues, including His80, Tyr290, Trp291, Phe453, and Ser287, provided moderate but consistent stabilization, particularly in the Vitexin and N,α-L-rhamnopyranosyl vincosamide complexes. Overall, the combined hydrogen-bond and energy decomposition analyses identify a conserved interaction core centered on Phe98, Glu99, His80, Tyr290, Trp291, and Gln457 that governs ligand stabilization in SGLT2. Ligands that engage this core with high persistence exhibit enhanced energetic favorability, providing a mechanistic basis for the rational optimization of SGLT2 inhibitors.

In the DPP-IV ligand systems, MD simulations revealed stable hydrogen-bond occupancy networks within the catalytic pocket (Table S10). Sitagliptin binding was dominated by persistent interactions with the catalytic residues Glu205 and Glu206 (57.70% and 37.28%, respectively), indicating a well-defined and stable binding mode. In contrast, N-α-L-rhamnopyranosyl vincosamide exhibited a more distributed hydrogen-bonding pattern involving Glu361 (38.49%), Glu206 (25.47%), Cys301 (19.33%), Tyr585 (18.20%), Pro359 (17.03%), and Arg358 (14.21) each with moderate occupancies. Residue-wise MM/GBSA energy decomposition corroborated these trends (Table S11). Sitagliptin binding was primarily stabilized by Glu205 (− 3.52 ± 1.14 kJ/mol) and Glu206 (− 3.36 ± 1.52 kJ/mol), with additional contributions from Tyr547 and Tyr666. In contrast, the natural product ligand N-α-L-rhamnopyranosyl vincosamide showed distributed energetic contributions from Phe357, Glu361, Gln553, and Ser552, consistent with its diffuse hydrogen-bond profile. Overall, persistent engagement of the catalytic Glu205/Glu206 pair underlies stronger and more focused stabilization in the sitagliptin complex, whereas N-α-L-rhamnopyranosyl vincosamide relies on a broader hydrogen bond interaction network, providing a structural basis for differences in binding efficiency.

Among the evaluated compounds, N,α-L-rhamnopyranosyl vincosamide consistently exhibited the most favorable profile across both targets. In SGLT2, it maintained global compactness comparable to empagliflozin while inducing reduced flexibility within the binding pocket and forming persistent number of hydrogen-bond networks. End-point binding free energy calculations further revealed that this ligand exhibits a stronger binding affinity for SGLT2, driven by highly favorable van der Waals and electrostatic interactions. Residue-wise energy decomposition identified a conserved interaction core involving Phe98, Glu99, His80, Tyr290, Trp291, and Gln457 as the primary contributors to stabilization. Although N,α-L-rhamnopyranosyl vincosamide formed moderately persistent hydrogen bonds compared with empagliflozin, its distributed interaction network enabled substantial cumulative stabilization, which explains its higher binding free energy. A recent substrate-recognition study identified His80, Asn75, Glu99, Lys321, Trp291, Gln457, and Ser460 as key residues that form an extensive hydrogen-bonding network within the SGLT2 substrate-binding site, which is critical for substrate recognition76. Consistent with this, the hydrogen-bond occupancy analysis revealed that N,α-L-rhamnopyranosyl vincosamide engages six (Asn75, His80, Glu99, Trp291, Lys321, and Gln457) of these seven residues, whereas the reference inhibitor empagliflozin interacts with only four (His80, Glu99, Trp291, and Gln457). Furthermore, small-molecule inhibitor identification studies targeting SGLT2 have independently identified Asn75, Phe98, Glu99, Ser287, Gln457, and Trp291 as critical interaction hotspots for potent inhibitors77. Together, these observations indicate that N,α-L-rhamnopyranosyl vincosamide more engages the native substrate-recognition interaction network of SGLT2, providing a structural basis for its enhanced binding stability and affinity.

In DPP-IV, N,α-L-rhamnopyranosyl vincosamide formed a stable yet more dynamic complex than sitagliptin, characterized by distributed hydrogen-bonding interactions and a balanced energetic profile with reduced desolvation penalties. The end-point binding energy results revealed favorable binding for both sitagliptin and N,α-L-rhamnopyranosyl vincosamide, with sitagliptin showing a marginally stronger affinity due to focused electrostatic interactions with the catalytic Glu205/Glu206 pair critical for substrate and ligand recognition (Gnoth et al., 2024; Kim et al., 2006). In contrast, N,α-L-rhamnopyranosyl vincosamide exhibited a more balanced energetic profile, characterized by moderate electrostatics (moderate interaction with Glu206), lower desolvation costs, and distributed stabilizing contributions across several residues within the active site. This distributed interaction pattern, as evidenced by hydrogen-bond occupancy and residue-wise energy decomposition analyses, indicates enhanced conformational adaptability rather than diminished binding strength. Overall, the binding free energy and per-residue analyses underscore the strong, mechanistically distinct binding of N,α-L-rhamnopyranosyl vincosamide to both SGLT2 and DPP-IV, reinforcing its potential as a dual-target hit and providing a robust energetic basis for further optimization and experimental validation. Dual inhibition of DPP-IV and SGLT2 represents a complementary therapeutic strategy that targets glucose homeostasis through distinct yet synergistic mechanisms. While DPP-IV inhibition enhances endogenous incretin signaling, leading to glucose-dependent insulin secretion and reduced glucagon release, SGLT2 inhibition promotes urinary glucose excretion independently of insulin action. The combination of these mechanisms has the potential to improve glycaemic control while mitigating adverse effects such as hypoglycaemia and weight gain commonly associated with insulin-centric therapies. Therefore, identifying natural compounds with concurrent DPP-IV and SGLT2 inhibitory potential is clinically relevant and supports the rationale for exploring dual-target scaffolds as next-generation antidiabetic agents.

ADMET profiling

The absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties of selected M.oleifera metabolites were evaluated in silico to assess their pharmacokinetic behaviour and safety profiles. The ADMET is summarized in Table 9 below. Vitexin displayed a molecular weight of 432.11 g/mol, 10 hydrogen bond acceptors, 7 hydrogen bond donors, 3 rotatable bonds, a cLogP of 1.232, and a topological polar surface area (TPSA) of 181.05 Å2. It satisfied Lipinski’s rule but showed low predicted human intestinal absorption, a clearance of 4.091 mL/min/kg, a short half-life of 0.765 h, and was classified as soluble. Homoorientin exhibited similar characteristics, with a molecular weight of 448.10 g/mol, high polarity (TPSA = 201.28 Å2), and multiple hydrogen-bonding sites, resulting in low predicted absorption and a short half-life of 0.823 h. Lariciresinol 4-O-β-D-glucopyranoside had a higher molecular weight (522.21 g/mol) and 11 hydrogen-bond acceptors, resulting in low intestinal absorption and rapid clearance (4.909 mL/min/kg), with a half-life of 0.390 h.

Table 9.

ADMET properties of prioritized compounds.

Compounds Physicochemical Medicinal Absorption range Excretion Solubility
MW nHA nHD nRot cLog P TPSA Lipinski C2P HIA Clear-ance T1/2 ESOL log Solubility class
Vitexin 432.11 10 7 3 1.232 181.05 Accepted -6.22 Low 4.091 0.765 -284 Soluble
Homoorientin 448.10 11 8 3 0.708 201.28 Rejected -6.25 Low 4.067 0.823 -2.70 Soluble
Lariciresinol 4-O-beta-D-glucopyranoside 522.21 11 6 9 -0.365 167.53 Rejected -5.99 Low 4.909 0.390 -3.10 Soluble
N, α-L-rhamnopyranosyl vincosamide 178.08 5 3 1 -2.26 79.15 Accepted -5.25 High 3.853 0.716 0.10 Highly soluble

N,α-L-rhamnopyranosyl-vincosamide displayed a distinct pharmacokinetic profile compared to glycosylated flavonoids and lignans. It showed low molecular weight, moderate polarity, high predicted intestinal absorption, good solubility, and compliance with Lipinski’s rules, indicating favourable systemic bioavailability. In contrast, C glycosyl flavonoids such as vitexin and homoorientin exhibited high polarity and extensive hydrogen bonding, leading to poor gastrointestinal absorption and significant first-pass metabolism, consistent with reported low oral bioavailability in vivo47. Lignan glycosides showed moderate polarity and rapid elimination, matching literature reports of low bioavailability for dietary lignans78.

Plasma protein binding varied by scaffold, with moderate binding for lariciresinol glycosides and high binding for isolariciresinol derivatives. None of the compounds were predicted to be P glycoprotein substrates, and all showed minimal CYP inhibition, aligning with previous reports on M. oleifera phytochemicals12. Toxicity predictions indicated low risk for hERG inhibition, mutagenicity, and DILI, with vincosamide showing the lowest DILI score, consistent with its reported lack of hepatotoxicity79.

Overall, scaffold-dependent structure property relationships emerged. Glycosylated flavonoids and lignan glycosides are pharmacologically safe but poorly absorbed, suggesting local gastrointestinal activity. In contrast, vincosamide combines favourable absorption, physicochemical properties, and low toxicity, making it a promising scaffold for developing dual-target antidiabetic drugs. Enhancing the bioavailability of glycosylated compounds may require formulation strategies such as nanoparticle encapsulation, co-crystal formation, or enzymatic deglycosylation.

Conclusion

This study presents an integrated in silico framework combining stacked ensemble machine learning, metabolomics, molecular docking, molecular dynamics, and ADMET profiling to identify potential dual DPP-IV and SGLT2 inhibitors from M. oleifera. The ensemble models demonstrated strong predictive performance and reliability for both targets, successfully recognising known clinical inhibitors and prioritising bioactive phytochemicals with favourable drug-like properties.

Virtual screening highlighted flavonoid and lignan glycosides, including vitexin, homoorientin, and lariciresinol-4-O-β-D-glucopyranoside, as promising SGLT2 inhibitors, while N,α-L-rhamnopyranosyl vincosamide emerged as a potential dual-target scaffold. Docking and molecular dynamics analyses revealed stable binding conformations and key interactions within the active sites of both enzymes, comparable to those of reference inhibitors. These findings support the therapeutic rationale of dual DPP-IV/SGLT2 modulation through complementary incretin enhancement and glucosuria-mediated glucose control.

Although the results provide strong computational and structural support, they remain predictive. Experimental validation through in vitro and in vivo studies is required to confirm dual inhibitory activity and pharmacological efficacy. Overall, this work demonstrates the utility of ensemble learning for natural product–based multitarget drug discovery and positions M. oleifera as a promising source of antidiabetic lead scaffolds.

Supplementary Information

Acknowledgements

The authors gratefully acknowledge the NRF for funding and institutional support provided through the University of Limpopo. Computational analyses were conducted using resources and infrastructure supported by the Department of Pharmacy, University of Limpopo. We also acknowledge the Centre for High Performance Computing (CHPC), Cape Town, South Africa, for providing access to the Schrödinger software suite and for running the AMBER molecular dynamics simulation package on the CHPC computing cluster.

Author contributions

Letuku MK collected the plant material, performed the phytochemical analysis, developed the Python scripts, and drafted the manuscript. Appiah-Kubi P. performed the docking experiments and MD simulations and drafted the relevant manuscript sections. Singh A. supervised the MD simulations and revised them. Mohlala MG critically revised the manuscript. Nuapia YB conceived and supervised the study and contributed to the revision of both the manuscript and the Python scripts. All authors read and approved the final version of the manuscript.

Funding

This research was supported by the National Research Foundation (NRF) of South Africa under grant number CSUR240322210391.

Data availability

The Python scripts used for dataset preprocessing, model training, and stacked ensemble construction are available on GitHub at: https://github.com/Yanbelo/Dual-inhibition-Dppiv-and-Sglt2-

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Gairing, S. J., Schleicher, E. M. & Labenz, C. Diabetes Mellitus – Risk Factor and Potential Future Target for Hepatic Encephalopathy in Patients with Liver Cirrhosis? (Springer, 2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Cole, J. B. & Florez, J. C. Genetics of Diabetes Mellitus and Diabetes Complications (Nature Research, 2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Jiang, S., Yu, T., Di, D., Wang, Y. & Li, W. Worldwide burden and trends of diabetes among people aged 70 years and older, 1990–2019: A systematic analysis for the Global Burden of Disease Study 2019. Diabetes Metab. Res. Rev.10.1002/dmrr.3745 (2024). [DOI] [PubMed] [Google Scholar]
  • 4.Zhang, F.-S. et al. Global trends and hotspots of type 2 diabetes in children and adolescents: A bibliometric study and visualization analysis.. World J. Diabetes16(1), 96032 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Li, J. et al. Epidemiological status, development trends, and risk factors of disability-adjusted life years due to diabetic kidney disease: A systematic analysis of Global Burden of Disease Study 2021. Chin. Med. J.10.1097/cm9.0000000000003428 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Diawara, Abdoulaye, Coulibaly, Djibril Mamadou, Hussain, Talib Yusuf Abbas, Cisse, Cheickna, Li, Jian, Wele, Mamadou, Diakite, Mahamadou, Traore, Kassim, Doumbia, Seydou O., and Shaffer, Jeffrey G., Type 2 diabetes prevalence, awareness, and risk factors in rural Mali: a cross-sectional study, Sci. Rep.13 no. 1 (2023). [DOI] [PMC free article] [PubMed]
  • 7.Hu, Tairan, and Fang, Zhaohui, Explore potential immune-related targets of leeches in the treatment of type 2 diabetes based on network pharmacology and machine learning, Front. Genet.16 (2025). [DOI] [PMC free article] [PubMed]
  • 8.Hasan, I. et al. SGLT2 inhibitors: Beyond glycemic control. J. Clin. Transl. Endocrinol.10.1016/j.jcte.2024.100335 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Soliman, A. R. et al. Dual-faced guardians: SGLT2 inhibitors’ kidney protection and health challenges: A position statement by Kasralainy nephrology group (KANG). Diabetol. Metab. Syndr.10.1186/s13098-025-01790-w (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Liang, Q. et al. Exploration of ACE inhibitory peptides from sea cucumber viscera with DPP-IV inhibitory activity: Virtual screening, characterization, and potential mechanism investigation. Int. J. Biol. Macromol.10.1016/j.ijbiomac.2025.143843 (2025). [DOI] [PubMed] [Google Scholar]
  • 11.McLean, Patrick, Bennett, Josiah, “Trey” Woods, Edward, Chandrasekhar, Sanjay, Newman, Noah, Mohammad, Yusuf, Khawaja, Muzamil, Rizwan, Affan, Siddiqui, Riyan, Birnbaum, Yochai, Lavie, Carl J., Virani, Salim, Hachem, Karim El, Wilson Tang, W. H., Ahuja, Tania, Isaacs, Scott, and Krittanawong, Chayakrit, SGLT2 inhibitors across various patient populations in the era of precision medicine: The multidisciplinary team approach, Springer Nature, 2025.
  • 12.Chen, D. et al. A screening strategy for bioactive peptides from enzymolysis extracts of Lentinula edodes based on molecular docking and molecular dynamics simulation. J. Future Foods5(4), 388–397 (2025). [Google Scholar]
  • 13.Tshabalala, T., Ndhlala, A. R., Ncube, B., Abdelgadir, H. A. & Staden, J. V. Potential substitution of the root with the leaf in the use of Moringa oleifera for antimicrobial, antidiabetic and antioxidant properties. S. Afr. J. Bot.129, 106–112 (2020). [Google Scholar]
  • 14.Asafo-Agyei, T., Appau, Y., Barimah, K. B. & Asase, A. Medicinal plants used for management of diabetes and hypertension in Ghana. Heliyon10.1016/j.heliyon.2023.e22977 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Yedjou, C. G. et al. The Management of Diabetes Mellitus Using Medicinal Plants and Vitamins (Multidisciplinary Digital Publishing Institute (MDPI), 2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Aljazzaf, Badriyah, Regeai, Sassia, Elghmasi, Sana, Alghazir, Nadia, Balgasim, Amal, Hdud Ismail, Ismail M., Eskandrani, Areej A., Shamlan, Ghalia, Alansari, Wafa S., Al-Farga, Ammar, and Alghazeer, Rabia, Evaluation of antidiabetic effect of combined leaf and seed extracts of Moringa oleifera (Moringaceae) on alloxan-induced diabetes in mice: A biochemical and histological study, Oxidati. Med. Cell. Long.2023 (2023). [DOI] [PMC free article] [PubMed]
  • 17.Hossain, Alomgir, Rahman, Md Ekhtiar, Faruqe, Md Omar, Saif, Ahmed, Suhi, Suzzada, Zaman, Rashed, Hirad, Abdurahman Hajinur, Matin, Mohammad Nurul, Rabbee, Muhammad Fazle, and Baek, Kwang Hyun, Characterization of plant-derived natural inhibitors of dipeptidyl peptidase-4 as potential antidiabetic agents: A computational study, Pharmaceutics 16 no. 4 (2024). [DOI] [PMC free article] [PubMed]
  • 18.Li, M. et al. Novel computational approaches in the discovery and identification of bioactive peptides: A bioinformatics perspective. J. Agric. Food Chem.73(22), 13212–13228 (2025). [DOI] [PubMed] [Google Scholar]
  • 19.Srisongkram, T., Waithong, S., Thitimetharoch, T. & Weerapreeyakul, N. Machine learning and in vitro chemical screening of potential α-amylase and α-glucosidase inhibitors from Thai indigenous plants. Nutrients10.3390/nu14020267 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Tufail, S., Riggs, H., Tariq, M. & Sarwat, A. I. Advancements and Challenges in Machine Learning: A Comprehensive Review of Models, Libraries, Applications, and Algorithms (MDPI, 2023). [Google Scholar]
  • 21.Obaido, G. et al. Supervised machine learning in drug discovery and development: Algorithms, applications, challenges, and prospects. Mach. Learn. Appl.17, 100576 (2024). [Google Scholar]
  • 22.Oselusi, Samson O., Dube, Phumuzile, Odugbemi, Adeshina I., Akinyede, Kolajo A., Ilori, Tosin L., Egieyeh, Elizabeth, Sibuyi, Nicole RS, Meyer, Mervin, Madiehe, Abram M., Wyckoff, Gerald J., and Egieyeh, Samuel A., The role and potential of computer-aided drug discovery strategies in the discovery of novel antimicrobials, Elsevier Ltd, 2024. [DOI] [PubMed]
  • 23.Giordano, Deborah, Biancaniello, Carmen, Argenio, Maria Antonia, and Facchiano, Angelo, Drug Design by Pharmacophore and Virtual Screening Approach, MDPI, 2022. [DOI] [PMC free article] [PubMed]
  • 24.Jamir, Lemnaro, and Hariprasad, P., Employing machine learning models to predict potential α-glucosidase inhibitory plant secondary metabolites targeting type-2 diabetes and their in vitro validation, J. Chem. Inform. Model. (2024). [DOI] [PubMed]
  • 25.Liu, X. W. et al. iPADD: A computational tool for predicting potential antidiabetic drugs using machine learning algorithms. J. Chem. Inf. Model.63(15), 4960–4969 (2023). [DOI] [PubMed] [Google Scholar]
  • 26.Lipinski, Christopher A., Lead- and drug-like compounds: The rule-of-five revolution, 2004. [DOI] [PubMed]
  • 27.Boldini, Davide, Ballabio, Davide, Consonni, Viviana, Todeschini, Roberto, Grisoni, Francesca, and Sieber, Stephan A., Effectiveness of molecular fingerprints for exploring the chemical space of natural products, J. Cheminform.16 no. 1 (2024). [DOI] [PMC free article] [PubMed]
  • 28.Ambhore, Jaya P., Laddha, Purushottam R., Nandedkar, Anjali, Ajmire, Prashant V., Chumbhale, Deshraj S., Navghare, Ashish B., Kuchake, Vitthal G., Chaudhari, Prashant J., and Adhao, Vaibhav S., Medicinal chemistry of non-peptidomimetic dipeptidyl peptidase IV (DPP IV) inhibitors for treatment of Type-2 diabetes mellitus: Insights on recent development, Elsevier B.V., 2023.
  • 29.Myoli, Akhona, Choene, Mpho, Kappo, Abidemi Paul, Madala, Ntakadzeni Edwin, Hooft, Justin J.J. van der, and Tugizimana, Fidele, Charting the Cannabis plant chemical space with computational metabolomics, Metabolomics 20 no. 3 (2024). [DOI] [PMC free article] [PubMed]
  • 30.Burley, S. K. et al. RCSB protein data bank: Biological macromolecular structures enabling research and education in fundamental biology, biomedicine, biotechnology and energy. Nucleic Acids Res.10.1093/nar/gky1004 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Kim, S. et al. PubChem 2023 update. Nucleic Acids Res.51(D1), D1373–D1380 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Miller, E. B. et al. Reliable and accurate solution to the induced fit docking problem for protein-ligand binding. J. Chem. Theory Comput.17(4), 2630–2639 (2021). [DOI] [PubMed] [Google Scholar]
  • 33.Maier, J. A. et al. ff14SB: Improving the accuracy of protein side chain and backbone parameters from ff99SB. J. Chem. Theory Comput.11(8), 3696–3713 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Mark, P. & Nilsson, L. Structure and dynamics of the TIP3P, SPC, and SPC/E water models at 298 K. J. Phys. Chem. A105(43), 9954–9960 (2001). [Google Scholar]
  • 35.He, X., Man, V. H., Yang, W., Lee, T. S. & Wang, J. A fast and high-quality charge model for the next generation general AMBER force field. J. Chem. Phys.10.1063/5.0019056 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Gallegos, M. & Martín Pendás, Á. Developing a user-friendly code for the fast estimation of well-behaved real-space partial charges. J. Chem. Inf. Model.63(13), 4100–4114 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Schneider, Jakob, Ribeiro, Rui, Alfonso-Prieto, Mercedes, Carloni, Paolo, and Giorgetti, Alejandro, Hybrid MM/CG webserver: Automatic set up of molecular mechanics/coarse-grained simulations for human g protein-coupled receptor/ligand complexes, Front. Mol. Biosci.7 (2020). [DOI] [PMC free article] [PubMed]
  • 38.López-Villellas, L. et al. Accurate and efficient constrained molecular dynamics of polymers using Newton’s method and special purpose code. Comput. Phys. Commun.10.1016/j.cpc.2023.108742 (2023). [Google Scholar]
  • 39.Agarwal, Shubham, Sukhomlinov, Sergey V., Honecker, Marc, and Müser, Martin H., Advanced Langevin thermostats: Properties, extensions to rheology, and a lean momentum-conserving approach, (2025). [DOI] [PubMed]
  • 40.King, E., Aitchison, E., Li, H. & Luo, R. Recent Developments in Free Energy Calculations for Drug Discovery (Frontiers Media S.A., 2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Wang, J., Hou, T. & Xu, X. Recent advances in free energy calculations with a combination of molecular mechanics and continuum models. Current Comput Aid Drug Des2(3), 287–306 (2006). [Google Scholar]
  • 42.Moberly, James G., Bernards, Matthew T., and Waynant, Kristopher V., Key features and updates for origin 2018, BioMed Central Ltd., 2018. [DOI] [PMC free article] [PubMed]
  • 43.Kemmish, H., Fasnacht, M. & Yan, L. Fully automated antibody structure prediction using BIOVIA tools: Validation study. PLoS ONE10.1371/journal.pone.0177923 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Pettersen, E. F. et al. UCSF Chimera - A visualization system for exploratory research and analysis. J. Comput. Chem.25(13), 1605–1612 (2004). [DOI] [PubMed] [Google Scholar]
  • 45.Mathur, V. et al. Insight into structure activity relationship of DPP-4 inhibitors for development of antidiabetic agents. Molecules10.3390/molecules28155860 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Gao, F. et al. Effects of different dietary flavonoids on dipeptidyl peptidase-IV activity and expression: Insights into structure-activity relationship. J. Agric. Food Chem.68(43), 12141–12151 (2020). [DOI] [PubMed] [Google Scholar]
  • 47.Istrate, D. & Crisan, L. Natural compounds as DPP-4 inhibitors: 3D-similarity search, ADME toxicity, and molecular docking approaches. Symmetry10.3390/sym14091842 (2022). [Google Scholar]
  • 48.Wakchaure, P. D. & Ganguly, B. Unraveling the role of π-stacking interactions in ligand binding to the thiamine pyrophosphate riboswitch with high-level quantum chemical calculations and docking study. J. Phys. Chem. B126(5), 1076–1084 (2022). [DOI] [PubMed] [Google Scholar]
  • 49.Souza, MMd. et al. Prodrug approach as a strategy to enhance drug permeability. Pharmaceuticals10.3390/ph18030297 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Prajapati, P., Pandey, J., Tandon, P., Sinha, K. & Shimpi, M. R. Molecular structural, hydrogen bonding interactions, and chemical reactivity studies of Ezetimibe-L-Proline cocrystal using spectroscopic and quantum chemical approach. Front. Chem.10.3389/fchem.2022.848014 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Yu, Y., Xia, Y. & Liang, G. Exploring novel lead scaffolds for SGLT2 inhibitors: Insights from machine learning and molecular dynamics simulations. Int. J. Biol. Macromol.10.1016/j.ijbiomac.2024.130375 (2024). [DOI] [PubMed] [Google Scholar]
  • 52.Moinul, M. et al. Exploring sodium glucose cotransporter (SGLT2) inhibitors with machine learning approach: A novel hope in anti-diabetes drug discovery. J. Mol. Graph. Model.10.1016/j.jmgm.2021.108106 (2022). [DOI] [PubMed] [Google Scholar]
  • 53.Yang, L. et al. Identifying patients at risk of acute kidney injury among Medicare beneficiaries with type 2 diabetes initiating SGLT2 inhibitors: A machine learning approach. Front. Pharmacol.10.3389/fphar.2022.834743 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Bao, X. et al. DPP-IV inhibitory peptides from highland barley via machine learning and multi-scale validation. Food Chem.10.1016/j.foodchem.2025.144979 (2025). [DOI] [PubMed] [Google Scholar]
  • 55.Parigi, A. D., Tang, W., Liu, D., Lee, C. & Pratley, R. Machine learning to identify predictors of glycemic control in type 2 diabetes: An analysis of target HbA1c reduction using Empagliflozin/Linagliptin data. Pharm. Med.33(3), 209–217 (2019). [DOI] [PubMed] [Google Scholar]
  • 56.Shoombuatong, W., Charoenkwan, P., Kanthawong, S., Nantasenamat, C. & Hasan, M. M. IDPPIV-SCM: A sequence-based predictor for identifying and analyzing dipeptidyl peptidase IV (DPP-IV) inhibitory peptides using a scoring card method. J. Proteome Res.19(10), 4125–4136 (2020). [DOI] [PubMed] [Google Scholar]
  • 57.Charoenkwan, P. et al. StackIL6: A stacking ensemble model for improving the prediction of IL-6 inducing peptides. Brief. Bioinform.22(6), 172 (2021). [DOI] [PubMed] [Google Scholar]
  • 58.Zhang, Y. et al. Mining bovine milk proteins for DPP-4 inhibitory peptides using machine learning and virtual proteolysis. Research7, 0391 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wang, D. et al. StructuralDPPIV: a novel deep learning model based on atom structure for predicting dipeptidyl peptidase-IV inhibitory peptides. Bioinformatics40(2), btae057 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Malik, A. A. et al. StackHCV: A web-based integrative machine-learning framework for large-scale identification of hepatitis C virus NS5B inhibitors. J. Comput. Aided Mol. Des.35(10), 1037–1053 (2021). [DOI] [PubMed] [Google Scholar]
  • 61.Sampath, P. et al. Robust diabetic prediction using ensemble machine learning models with synthetic minority over-sampling technique. Sci. Rep.14(1), 28984 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.El-Sherbiny, G. M., Alluqmani, A. J., Elsehemy, I. A. & Kalaba, M. H. Antibacterial, antioxidant, cytotoxicity, and phytochemical screening of Moringa oleifera leaves. Sci. Rep.10.1038/s41598-024-80700-y (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Seghir, A. et al. Comprehensive chemical profiling of Moringa oleifera leaves extracts by LC–MS/MS followed by in silico ADMET prediction using SwissADME. Biomed. Chromatogr.10.1002/bmc.70110 (2025). [DOI] [PubMed] [Google Scholar]
  • 64.Ma, Z. F., Ahmad, J., Zhang, H., Khan, I. & Muhammad, S. Evaluation of phytochemical and medicinal properties of Moringa (Moringa oleifera) as a potential functional food. S. Afr. J. Bot.129, 40–46 (2020). [Google Scholar]
  • 65.Wang, J. et al. LC-MS/MS-based chemical profiling of water extracts of Moringa oleifera leaves and pharmacokinetics of their major constituents in rat plasma. Food Chem. X10.1016/j.fochx.2024.101585 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Pop, Oana Lelia, Kerezsi, Andreea Diana, and Ciont, Călina, A Comprehensive review of Moringa Oleifera bioactive compounds—cytotoxicity evaluation and their encapsulation, MDPI, 2022. [DOI] [PMC free article] [PubMed]
  • 67.Chokoe, T. A. et al. Phenolic metabolite profiling of Moringa oleifera Lam waste under different drying methods and extraction solvents. S. Afr. J. Bot.184, 182–192 (2025). [Google Scholar]
  • 68.Ansari, Prawej, Khan, Joyeeta T., Chowdhury, Suraiya, Reberio, Alexa D., Kumar, Sandeep, Seidel, Veronique, Abdel-Wahab, Yasser H.A., and Flatt, Peter R., Plant-based diets and phytochemicals in the management of diabetes mellitus and prevention of its complications: A review, Multidisciplinary Digital Publishing Institute (MDPI), 2024. [DOI] [PMC free article] [PubMed]
  • 69.Abdulai, Ibrahim Luru, Kwofie, Samuel Kojo, Gbewonyo, Winfred Seth, Boison, Daniel, Puplampu, Joshua Buer, and Adinortey, Michael Buenor, Multitargeted effects of vitexin and isovitexin on diabetes mellitus and its complications, Hindawi Limited, 2021. [DOI] [PMC free article] [PubMed]
  • 70.Han, X. et al. α-Glucosidase inhibition mechanism and anti-hyperglycemic effects of flavonoids from Astragali Radix and their mixture effects. Pharmaceuticals10.3390/ph18050744 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Gong, Q. et al. New techniques and strategies in drug discovery (2020–2024 update). Chin. Chem. Lett.10.1016/j.cclet.2024.110456 (2025). [Google Scholar]
  • 72.Dorosh, Lyudmyla, Wille, Holger, and Stepanova, Maria, Multiscale simulations of folded and intrinsically disordered region-containing protein condensates, Biophys. J. (2025). [DOI] [PMC free article] [PubMed]
  • 73.Elsaman, T., Awadalla, M. K. A., Mohamed, M. S., Eltayib, E. M. & Mohamed, M. A. Identification of microbial-based natural products as potential CYP51 inhibitors for Eumycetoma treatment: Insights from molecular docking, MM-GBSA calculations, ADMET analysis, and molecular dynamics simulations. Pharmaceuticals10.3390/ph18040598 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Shafiq, N. et al. Integrated computational modeling and in-silico validation of flavonoids-Alliuocide G and Alliuocide A as therapeutic agents for their multi-target potential: Combination of molecular docking, MM-GBSA, ADMET and DFT analysis. S. Afr. J. Bot.169, 276–300 (2024). [Google Scholar]
  • 75.Alzain, A. A. et al. Integrating computational methods guided the discovery of phytochemicals as potential Pin1 inhibitors for cancer: Pharmacophore modeling, molecular docking, MM-GBSA calculations and molecular dynamics studies. Front. Chem.10.3389/fchem.2024.1339891 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Cui, W. et al. Mechanism of substrate recognition and release of human SGLT2. Nat. Commun.10.1038/s41467-025-62421-6 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Haque, A., Alenezi, K. M., Alam, M. J. & Tazyeen, S. Identification of new small-molecule SGLT2 inhibitors: In-silico approach to therapeutics for diabetes-related complications. ACS Omega10.1021/acsomega.5c05715 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Yang, Y. et al. Identification of potential dipeptidyl peptidase (DPP)-IV inhibitors among Moringa oleifera phytochemicals by virtual screening, molecular docking analysis, ADME/T-based prediction, and in vitro analyses. Molecules10.3390/molecules25010189 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Chigurupati, S. et al. Molecular docking of phenolic compounds and screening of antioxidant and antidiabetic potential of Moringa oleifera ethanolic leaves extract from Qassim region, Saudi Arabia. Saudi J. Biol. Sci.29(2), 854–859 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The Python scripts used for dataset preprocessing, model training, and stacked ensemble construction are available on GitHub at: https://github.com/Yanbelo/Dual-inhibition-Dppiv-and-Sglt2-


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES