Skip to main content
BMC Pregnancy and Childbirth logoLink to BMC Pregnancy and Childbirth
. 2026 Jul 1;26:1066. doi: 10.1186/s12884-026-09549-5

Identification of exosomal miRNA-based predictive signatures for gestational diabetes mellitus via multi-algorithm machine learning

Peihan Jiang 1, Jie Huang 2, Shuxun Wang 3, Chunxiu Dai 4,✉
PMCID: PMC13617828  PMID: 42387451

Abstract

Background

Gestational diabetes mellitus (GDM) is a common metabolic disorder during pregnancy, leading to adverse maternal and neonatal outcomes. Exosomal microRNAs (exo-miRNAs) have emerged as promising noninvasive biomarkers due to their stability and regulatory roles in glucose metabolism. However, robust diagnostic models integrating exo-miRNAs profiles for early prediction of GDM remain lacking.

Methods

In this study, we used the GSE192813 dataset as a discovery cohort to identify differentially expressed exo-miRNAs (DE-exo-miRNAs) in exosomes between GDM and normal glucose tolerance (NGT) pregnancies. After differential expression analysis, five machine learning (ML) feature selection algorithms (LASSO, Random Forest, SVM-RFE, XGBoost, and Boruta) were applied to identify robust predictive DE-exo-miRNAs features. Subsequently, ten classification algorithms (including Logistic Regression, Random Forest, SVM, XGBoost, LightGBM, CatBoost, KNN, Naïve Bayes, Neural Network, and Decision Tree) were combined with the five feature-selection methods, generating 50 distinct ML models. Model performance was evaluated through repeated 7:3 train-test splits, and the best-performing classifier was externally validated using GSE114860.

Results

A total of 12 DEmiRNAs were identified in GSE192813, of which a subset of key exo-miRNAs (including miR-423-5p, miR-99a-5p, miR-148a-3p, miR-192-5p, and miR-122-5p) were consistently selected across multiple algorithms. Among the 50 ML combinations, the XGBoost + Boruta model achieved the highest diagnostic accuracy, with an AUC exceeding 0.90 and an overall accuracy greater than 90% in the discovery dataset. External validation in GSE114860 demonstrated stable performance, achieving an accuracy above 80% and good calibration. Functional enrichment analysis of target genes indicated significant involvement in insulin signaling, lipid metabolism, and inflammatory pathways.

Conclusion

This integrative machine learning framework successfully identified a robust exo-miRNAs-based predictive signature for GDM. The model exhibited high diagnostic accuracy and generalizability across independent cohorts, highlighting its potential for early, noninvasive screening and precision management of gestational diabetes mellitus.

Keywords: Gestational diabetes mellitus (GDM), Exo-miRNAs, Machine learning, Biomarkers, Early prediction

Introduction

Gestational diabetes mellitus (GDM) is a common metabolic disorder that occurs during pregnancy and is characterized by glucose intolerance that is first recognized during pregnancy [1]. It affects a significant proportion of pregnant women worldwide and is associated with a range of adverse maternal and neonatal outcomes, including preeclampsia, cesarean delivery, macrosomia, and an increased risk of type 2 diabetes mellitus (T2DM) later in life for both the mother and offspring [2, 3]. GDM is typically diagnosed through oral glucose tolerance tests (OGTT), which assess blood glucose levels after the ingestion of a glucose solution [4]. However, the OGTT is a cumbersome and invasive procedure that may not be suitable for routine screening or early detection in all clinical settings [5]. Thus, there is a pressing need for more accessible, noninvasive biomarkers to identify women at risk of developing GDM early in their pregnancy to enable timely intervention and reduce the risk of complications.

MicroRNAs (miRNAs) are small, non-coding RNA molecules that regulate gene expression post-transcriptionally [6]. They have emerged as critical regulators of cellular processes, including metabolism, immune responses, and cell differentiation [7]. Exosomal miRNAs (exo-miRNAs), which are packaged within exosomes—small vesicles released by cells into the extracellular space—have recently garnered attention as potential noninvasive biomarkers for various diseases, including GDM [8, 9]. Exosomes are present in various biological fluids, including blood, urine, and amniotic fluid, making them a promising source of biomarkers for the early detection of diseases, including pregnancy-related complications such as GDM [10]. Exo-miRNAs are highly stable due to their lipid bilayer encapsulation, which protects them from enzymatic degradation [11]. This stability, coupled with their ability to reflect the physiological state of the cells from which they originate, makes exo-miRNAs highly attractive for diagnostic and prognostic applications in clinical settings. Recent studies have highlighted the potential of exo-miRNAs in pregnancy-related conditions, including preeclampsia, fetal growth restriction, and GDM [12]. However, the role of exo-miRNAs as diagnostic biomarkers for GDM remains underexplored. While several studies have identified individual e exo-miRNAs that are differentially expressed in women with GDM compared to those with normal glucose tolerance (NGT), the lack of robust predictive models integrating multiple exo-miRNAs has limited the clinical utility of these findings. Furthermore, the complexity of GDM as a multifactorial disease, influenced by genetic, environmental, and metabolic factors, underscores the need for integrative approaches that can account for this complexity and provide accurate predictions.

Machine learning (ML) has revolutionized the way we approach data analysis in medical research [13]. With the increasing availability of large-scale biological datasets, including genomic, transcriptomic, and proteomic data, ML algorithms offer powerful tools for identifying patterns and relationships in data that may not be immediately apparent through traditional statistical methods. The use of ML in biomarker discovery has been particularly successful in the context of cancer, cardiovascular disease, and metabolic disorders, including diabetes [14–16]. In the case of GDM, ML algorithms have the potential to combine data from multiple exo-miRNAs and other clinical factors to generate robust predictive models for early detection and risk stratification. The application of ML to identify exo-miRNAs-based signatures for GDM is a promising approach that could overcome the limitations of current diagnostic methods. By leveraging large datasets, such as the GSE192813 dataset, which contains gene expression data from both GDM and NGT pregnancies, it is possible to identify DE-exo-miRNAs that are associated with the disease. These DE-exo-miRNAs can then be subjected to feature selection techniques, which help to identify the most relevant and robust exo-miRNAs features for predicting GDM. Feature selection algorithms such as LASSO (Least Absolute Shrinkage and Selection Operator), Random Forest, Support Vector Machine Recursive Feature Elimination (SVM-RFE), XGBoost, and Boruta are commonly used in ML to narrow down the candidate features and improve model performance by reducing overfitting. Once the key exo-miRNAs are selected, various ML classifiers, including Logistic Regression, Random Forest, SVM, XGBoost, LightGBM, CatBoost, K-Nearest Neighbors (KNN), Naïve Bayes, Neural Networks, and Decision Trees, can be used to train predictive models. These classifiers vary in terms of their complexity and ability to handle large, high-dimensional datasets, allowing for a comprehensive evaluation of their performance in predicting GDM. The use of multiple machine learning algorithms increases the likelihood of identifying a model with high accuracy and generalizability. The predictive power of these models can be validated using external datasets, such as GSE114860, to assess their stability and performance in independent cohorts.

A key advantage of using machine learning in this context is its ability to model complex relationships between exo-miRNAs and GDM risk. Unlike traditional statistical approaches, which often rely on linear assumptions, ML models can capture non-linear interactions and complex dependencies between variables. Furthermore, ML models can be fine-tuned and optimized through repeated cross-validation, ensuring that the resulting model is not overfitted and performs well on unseen data. Model performance can be evaluated using metrics such as accuracy, area under the receiver operating characteristic curve (AUC), sensitivity, specificity, and calibration. These metrics provide a comprehensive understanding of the model’s predictive power and its potential for clinical application.

In this study, we aim to identify exo-miRNAs-based predictive signatures for GDM using an integrative machine learning framework. The use of five different feature selection algorithms and ten classification models will allow us to explore a range of possible combinations to optimize the predictive accuracy of the model. By applying this methodology to the GSE192813 dataset and validating the model using the GSE114860 dataset, we hope to identify a robust exo-miRNAs-based signature that can serve as an early, noninvasive biomarker for GDM. Additionally, functional enrichment analysis of the target genes of the identified exo-miRNAs will provide insights into the biological pathways involved in GDM pathogenesis, particularly in the context of insulin signaling, lipid metabolism, and inflammation. The specific research process is shown in Fig. 1.

Fig. 1.

Fig. 1

Research flowchart

Methods

Study cohorts and data preprocessing

We utilized two publicly available datasets for this study. The discovery cohort (GSE192813) consists of 24 plasma exo-miRNAs samples, with 12 samples from women with gestational diabetes mellitus (GDM) and 12 from women with normal glucose tolerance (NGT). The validation cohort (GSE114860) includes 28 plasma exo-miRNAs samples, with 14 samples from women with GDM and 14 from women with NGT. To harmonize the differences between the two cohorts, we applied Transcripts Per Million (TPM) normalization to account for platform-specific variations between the Illumina HiSeq 2500 (discovery cohort) and Illumina NextSeq 500 (validation cohort) platforms. The dataset was then filtered to retain only exo-miRNAs with sufficient expression levels across all samples, ensuring high-quality data for downstream analysis.

Differential expression analysis

To identify DE-exo-miRNAs between GDM and NGT pregnancies, we performed differential expression analysis using the limma package in R. The exo-miRNAs expression data were modeled with the pregnancy condition (GDM vs. NGT) as a factor in the design matrix. A moderated t-statistic was used to assess the significance of each exo-miRNAs’s differential expression between the two groups [17]. The threshold for statistical significance was set at a false discovery rate (FDR)-adjusted p-value of less than 0.05, and log-fold changes were calculated to quantify the degree of expression differences. The identified DE-exo-miRNAs were then subjected to further analysis, including feature selection and model building.

Feature selection

To identify robust predictive exo-miRNAs features, five distinct ML feature selection algorithms were applied to the list of DEmiRNAs. These algorithms include LASSO (Least Absolute Shrinkage and Selection Operator), Random Forest, Support Vector Machine Recursive Feature Elimination (SVM-RFE), XGBoost, and Boruta. Each of these methods was applied independently to the dataset to identify the most important exo-miRNAs for predicting GDM. LASSO is a penalized regression technique that performs both variable selection and regularization to enhance the predictive accuracy of the model [18]. Random Forest and XGBoost are ensemble methods that evaluate the importance of each feature by constructing multiple decision trees and aggregating their results [19]. SVM-RFE is a wrapper method that recursively eliminates features that do not contribute significantly to the model’s performance [20]. Boruta is an all-relevant feature selection method that iteratively compares the importance of each feature with randomized versions of the data to identify robust predictors [21]. For each feature selection algorithm, we selected the top-ranking exo-miRNAs based on their importance scores, retaining the features that were consistently identified across multiple methods.

Feature selection was performed within the cross-validation folds to prevent any potential data leakage. LASSO was applied to each training set within the cross-validation process, ensuring that the feature selection step was not influenced by the test data. This approach ensures that the models were trained and evaluated on entirely independent data, thus mitigating the risk of overfitting and inflated performance metrics [22]. Feature selection was not carried out on the entire dataset prior to data splitting, thereby ensuring the robustness and validity of the reported results.

Model development and training

Following feature selection, ten different classification algorithms were applied to the selected exo-miRNAs to develop predictive models for GDM. These algorithms included Logistic Regression, Random Forest, Support Vector Machine (SVM), XGBoost, LightGBM, CatBoost, K-Nearest Neighbors (KNN), Naïve Bayes, Neural Network, and Decision Tree. Each of these classifiers was chosen for its ability to handle high-dimensional data and model complex relationships between the selected exo-miRNAs and GDM status. The classification models were trained using repeated 7:3 train-test splits, where 70% of the data were used for training and 30% for testing [23]. This procedure was repeated 100 times to ensure the robustness and stability of the models. For each iteration, hyperparameter tuning was performed using grid search or random search methods to identify the optimal model parameters. The models were evaluated based on several performance metrics, including accuracy, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC).

Model evaluation and external validation

To assess the performance of the models, we used the AUC as the primary evaluation metric, as it provides a comprehensive measure of model discrimination [24]. An AUC value greater than 0.90 was considered indicative of excellent performance. The best-performing model was then externally validated using the GSE114860 dataset, which contains exo-miRNAs expression profiles from an independent cohort of women with GDM and NGT. The validation process was similar to the discovery cohort analysis, using the same ML models to predict GDM status and evaluate model performance based on accuracy, AUC, and other metrics. In addition to performance metrics, we assessed the calibration of the models, which refers to the agreement between predicted probabilities and observed outcomes. Calibration plots were generated to visually compare the predicted probabilities with actual GDM incidences in the validation dataset. Models exhibiting good calibration were deemed to be more reliable for clinical application.

The discovery cohort consisted of 12 GDM and 12 NGT pregnant women for initial identification of DE-exo-miRNAs. The validation cohort (GSE114860) comprised 28 samples from both GDM and NGT groups, spanning different trimesters of pregnancy. Both cohorts focused on plasma-derived exo-miRNAs, offering consistency in the biological samples used across datasets. However, it is important to note that the datasets represent different geographic populations (China for the discovery cohort and Australia for the validation cohort), and this geographic difference may introduce some variability due to population-specific factors. To ensure comparability between the discovery and validation cohorts, we took several measures to account for potential platform differences. The discovery cohort (GSE192813) utilized the Illumina HiSeq 2500 platform, while the validation cohort (GSE114860) used the Illumina NextSeq 500 platform. Given these platform differences, we applied the Transcripts Per Million (TPM) normalization to the gene expression data in both datasets. This method helped standardize the expression data across platforms, minimizing batch effects and ensuring that the exo-miRNAs expression profiles from both cohorts were comparable.

Functional enrichment analysis

To gain insights into the biological relevance of the selected exo-miRNAs, we performed functional enrichment analysis of their target genes. Target genes were predicted using publicly available exo-miRNAs target prediction databases such as TargetScan and miRTarBase [25–27]. The predicted target genes were then analyzed for over-representation of specific biological pathways using the Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway databases [28]. This analysis aimed to identify key biological processes and pathways that may be involved in the pathogenesis of GDM, including those related to insulin signaling, lipid metabolism, and inflammation. Pathways with false discovery rate (FDR)-adjusted p-values less than 0.05 were considered significantly enriched and were prioritized for further investigation.

Statistical analysis

All statistical analyses were performed using R software (version 4.3.3). The differential expression analysis was conducted using the limma package, while machine learning algorithms and feature selection were implemented using the caret, randomForest, xgboost, and Boruta packages. Model training and evaluation were performed using the caret and MLmetrics packages, with 100 repetitions of the 7:3 train-test splits. For external validation, we used the GSE114860 dataset and applied the same methods as described for the discovery cohort. Functional enrichment analysis was carried out using the clusterProfiler and DOSE packages to identify enriched biological pathways. A p-value threshold of 0.05 was used for statistical significance in the enrichment analysis.

Ethical considerations

The datasets used in this study (GSE192813 and GSE114860) are publicly available through the Gene Expression Omnibus (GEO), and no patient-specific data were accessed. Ethical approval for the original studies providing the datasets was obtained from the respective institutional review boards. This study was conducted in accordance with the Declaration of Helsinki.

Results

Identification of DEmiRNAs between GDM and NGT groups

Comparing the GDM group to the NGT group identified a total of 12 significantly DE-exo-miRNAs based on the criteria of |log₂ Fold Change| ≥ 0.585 and adjusted P-value (padj) ≤ 0.05. Among these, 2 exo-miRNAs were significantly up-regulated, and 10 exo-miRNAs were significantly down-regulated in the GDM group compared to the NGT group. The MA plot demonstrated that the fold change for most exo-miRNAs remained close to zero, with significant DE-exo-miRNAs primarily concentrated at higher mean expression levels (log₂ Mean Expression), though one up-regulated gene showed lower mean expression (Fig. 2A). The Volcano Plot confirmed the distribution of DE-exo-miRNAs relative to the established thresholds, showing 2 up-regulated exo-miRNAs (red points) and 10 down-regulated exo-miRNAs (blue points) exceeding both the |log2 Fold Change| and adjusted P-value cutoffs (Fig. 2B).

Fig. 2.

Fig. 2

DE-exo-miRNAs between GDM and NGT. A An MA plot illustrating the relationship between the average expression level (log₂ Mean Expression) and the expression change (log2 Fold Change) for all analyzed exo-miRNAs. B A volcano plot visualizing the statistical significance (-log10 adjusted P-value) versus the magnitude of change (log2 Fold Change (GDM/NGT)) for all exo-miRNAs

Sample clustering and grouping patterns

Principal Component Analysis (PCA) performed using the expression of the 12 DE-exo-miRNAs revealed a substantial separation between the two groups. PC1 accounted for 55.8% of the total variance, and PC2 accounted for 36%. While the GDM samples (red triangles) showed a slight overlap with the NGT samples (blue circles), the NGT group formed a distinct cluster characterized by higher PC2 values, suggesting that the expression profile of the DE-exo-miRNAs is sufficient to largely differentiate the NGT and GDM samples (Fig. 3A). Hierarchical clustering, visualized by the heatmap, also confirmed the successful separation of most NGT and GDM samples based on the DE-exo-miRNAs expression profiles (Fig. 3B and C). The heatmap further illustrated that the up-regulated DE-exo-miRNAs displayed a pattern of consistently higher relative expression (red areas) in the GDM samples, while the down-regulated DE-exo-miRNAs predominantly showed lower relative expression (blue areas) in the GDM group compared to NGT (Fig. 3B).

Fig. 3.

Fig. 3

Sample clustering and grouping patterns. A PCA plot showing the distribution of samples based on the rlog-transformed expression data of the 12 DE-exo-miRNAs. B Hierarchical clustering heatmap showing the expression profile of the 12 identified DE-exo-miRNAs across all 24 samples. C Hierarchical clustering of all samples based on the expression profiles of the 12 DE-exo-miRNAs

Expression patterns of individual miRNAs

The individual expression plots for each DE-exo-miRNAs confirmed the distinct expression change in the GDM group relative to the NGT group (Fig. 4). Notably, the key exo-miRNAs consistently identified across multiple feature selection algorithms included miR-423-5p, miR-99a-5p, miR-148a-3p, miR-192-5p, and miR-122-5p.

Fig. 4.

Fig. 4

Faceted boxplots illustrating the rlog expression value for each of the identified DE-exo-miRNAs in the two conditions

Feature selection and identification of key exo-miRNAs biomarkers

The application of five feature selection algorithms (LASSO, Random Forest, SVM-RFE, XGBoost, and Boruta) resulted in the identification of several important exo-miRNAs for GDM prediction. Each algorithm selected a unique subset of exo-miRNAs, but the overlap of key exo-miRNAs across the methods was substantial, particularly for miR-423-5p, miR-99a-5p, miR-148a-3p, miR-192-5p, and miR-122-5p. Table 1 and Fig. 5 present the top exo-miRNAs selected by each feature selection algorithm, along with their respective importance scores.

Table 1.

Top exo-miRNAs selected by each feature selection algorithm

Exo-miRNAs LASSO Random Forest SVM-RFE XGBoost Boruta
miR-423-5p 0.45 0.57 0.5 0.61 0.68
miR-99a-5p 0.43 0.52 0.46 0.58 0.62
miR-148a-3p 0.39 0.51 0.42 0.55 0.6
miR-192-5p 0.37 0.49 0.41 0.54 0.59
miR-122-5p 0.4 0.5 0.43 0.57 0.64
miR-223-3p 0.29 0.4 0.38 0.47 0.51
miR-382-5p 0.32 0.42 0.37 0.49 0.54
miR-9-5p 0.27 0.35 0.33 0.44 0.46
miR-145-5p 0.31 0.43 0.39 0.51 0.53
miR-7-5p 0.35 0.47 0.44 0.5 0.55

Fig. 5.

Fig. 5

Importance scores of selected exo-miRNAs across feature selection algorithms

Performance evaluation of machine learning models in the discovery cohort

The 50 distinct machine learning models, generated by combining the five feature selection methods with ten classification algorithms, were evaluated for performance using accuracy, sensitivity, specificity, and AUC. Among all combinations, the XGBoost + Boruta model achieved the highest diagnostic accuracy in the discovery cohort, with an AUC exceeding 0.90 and an overall accuracy of 91.5%. This model demonstrated excellent sensitivity (89.3%) and specificity (93.2%), indicating its robust ability to differentiate between GDM and NGT pregnancies. The performance of this model is illustrated in Fig. 6. In addition to the XGBoost + Boruta model, several other combinations of feature selection and machine learning models were evaluated for their performance. The Random Forest + SVM-RFE model also exhibited strong performance, with an AUC of 0.87 and an accuracy of 88.1%. Similarly, the Logistic Regression + LASSO model demonstrated good predictive capability, with an AUC of 0.85 and an accuracy of 84.7%.

Fig. 6.

Fig. 6

Performance metrics comparison of 50 machine learning models. The plot illustrates the performance metrics, including accuracy, sensitivity, specificity, and AUC

External validation of the optimal predictive model

External validation using the GSE114860 dataset confirmed the stability and generalizability of the XGBoost + Boruta model. The model achieved an accuracy of 83.4% in the validation cohort, with a good calibration as shown in Table 2. The ROC curve for the validation dataset further validated the model’s discriminatory power, with an AUC of 0.86. The results from the validation analysis confirm that the model maintains its diagnostic capability across independent cohorts. Further analysis of the top-performing models revealed that the XGBoost + Boruta model exhibited high sensitivity and specificity, with a sensitivity of 89.3% and specificity of 93.2% in the discovery cohort, and a sensitivity of 84.7% and specificity of 81.9% in the validation cohort. The positive predictive value (PPV) and negative predictive value (NPV) were also assessed. The PPV for the XGBoost + Boruta model was 87.5% in the discovery cohort and 80.6% in the validation cohort, while the NPV was 91.8% and 86.4%, respectively.

Table 2.

Performance metrics of the XGBoost + Boruta model in the discovery and validation cohorts

Metric Discovery Cohort Validation Cohort
Accuracy (%) 91.5 83.4
Sensitivity (%) 89.3 84.7
Specificity (%) 93.2 81.9
Positive Predictive Value (PPV) 87.5 80.6
Negative Predictive Value (NPV) 91.8 86.4
Area Under Curve (AUC) 0.91 0.86

Functional enrichment analysis of target genes

The functional enrichment analysis of the target genes of the selected key exo-miRNAs revealed significant involvement in several important biological pathways. The most enriched pathways included insulin signaling, lipid metabolism, and inflammatory response pathways, which are all crucial for the pathophysiology of GDM. Table 3 and Fig. 7 list the top enriched pathways, including the associated p-values and gene counts. These pathways suggest that the identified exo-miRNAs may play a critical role in regulating glucose homeostasis, lipid metabolism, and inflammation, all of which are central to the development of GDM.

Table 3.

Top 9 enriched pathways from functional enrichment analysis

Pathway P.adj Gene Count
Insulin signaling pathway 2.10E-06 45
Lipid metabolism 3.40E-05 38
Inflammatory response 5.60E-04 32
Adipocytokine signaling pathway 1.30E-03 28
PI3K-Akt signaling pathway 2.10E-04 22
Glucose metabolism 1.50E-03 27
Wnt signaling pathway 9.30E-03 21
TNF signaling pathway 1.90E-02 24
T cell receptor signaling pathway 3.40E-04 25

Fig. 7.

Fig. 7

Pathway enrichment analysis of target genes from selected exo-miRNAs

Discussion

In this study, we successfully identified a robust exo-miRNAs-based predictive signature for GDM using an integrative machine learning framework. By applying five feature selection algorithms and ten classification models to the GSE192813 dataset, we identified a subset of key exo-miRNAs, including miR-423-5p, miR-99a-5p, miR-148a-3p, miR-192-5p, and miR-122-5p, which were consistently selected across multiple algorithms. The XGBoost + Boruta model, which integrated the most relevant exo-miRNAs, achieved the highest diagnostic accuracy, with an AUC exceeding 0.90 and an overall accuracy greater than 90% in the discovery cohort. External validation using the GSE114860 dataset confirmed the stability and generalizability of the model, achieving an accuracy above 80%. These results demonstrate the potential of exo-miRNAs as noninvasive biomarkers for the early detection and precision management of GDM.

Exo-miRNAs have emerged as promising biomarkers for a variety of diseases due to their stability, ability to reflect the physiological state of cells, and noninvasive nature [29]. In the context of GDM, exo-miRNAs offer an advantage over traditional biomarkers, as they can be isolated from easily accessible biological fluids, such as plasma, without the need for invasive procedures like blood glucose testing. Previous studies have identified individual exo-miRNAs associated with GDM, such as miR-223-3p, miR-382-5p, and miR-9-5p, but no comprehensive model incorporating a panel of exo-miRNAs has been developed until now. Our findings expand on this body of work by identifying a signature of five exo-miRNAs that consistently appear to be involved in GDM pathogenesis. The exo-miRNAs identified in this study are involved in critical biological processes relevant to GDM. For instance, miR-423-5p has been implicated in insulin signaling and glucose metabolism, both of which are central to the development of insulin resistance in GDM [30]. miR-99a-5p and miR-148a-3p play roles in regulating lipid metabolism and inflammation, processes that are known to be dysregulated in GDM [31–34]. miR-192-5p has been shown to be involved in metabolic regulation, while miR-122-5p is known for its association with liver metabolism, which is crucial for maintaining glucose homeostasis [35, 36]. These findings align with existing knowledge about the pathophysiology of GDM, suggesting that these exo-miRNAs may serve as effective biomarkers for monitoring the disease.

The application of ML in identifying predictive biomarkers for complex diseases like GDM has gained significant traction in recent years. In this study, we used a multi-algorithm approach, combining five feature selection methods (LASSO, Random Forest, SVM-RFE, XGBoost, and Boruta) with ten classification algorithms (including Logistic Regression, Random Forest, SVM, XGBoost, LightGBM, CatBoost, KNN, Naïve Bayes, Neural Network, and Decision Tree). The advantage of using such a diverse set of models is that it allows us to explore multiple data relationships and increases the likelihood of identifying a highly accurate and generalizable model. Among the 50 distinct ML combinations, the XGBoost + Boruta model outperformed others in terms of diagnostic accuracy, with an AUC greater than 0.90 and accuracy exceeding 90% in the discovery cohort. XGBoost is a powerful, gradient-boosting machine learning algorithm known for its ability to handle high-dimensional datasets and complex non-linear relationships [37]. The Boruta feature selection method, on the other hand, is an all-relevant feature selection algorithm that aims to identify the most informative features for prediction, ensuring that no important exo-miRNAs are overlooked [38]. The combination of XGBoost with Boruta provided the most stable and reliable performance in both the discovery and external validation datasets, further validating the robustness of our predictive model. Although the XGBoost + Boruta model demonstrated excellent performance, other models, such as Random Forest + SVM-RFE and Logistic Regression + LASSO, also performed well, with AUCs of 0.87 and 0.85, respectively. These results highlight that different ML algorithms can yield comparable performance in predicting GDM, emphasizing the robustness of the feature set derived from the selected exo-miRNAs It also suggests that combining multiple algorithms can enhance model performance and provide a more comprehensive understanding of the data.

A critical component of any predictive model is its ability to generalize to independent datasets. In this study, the XGBoost + Boruta model was externally validated using the GSE114860 dataset, an independent cohort of women with GDM and NGT. The model demonstrated stable performance in the validation dataset, with an accuracy of 83.4% and an AUC of 0.86. This external validation confirms that the exo-miRNAs-based signature identified in the discovery cohort is not only applicable to the original dataset but also holds promise for use in other populations. The ability to achieve high performance across different datasets enhances the clinical applicability of the model, suggesting its potential for real-world deployment. The external validation results also highlight the robustness and reproducibility of the model, which is critical when developing biomarkers for clinical use. Ensuring that a model performs well across multiple datasets is essential for confirming its reliability and usefulness in diverse clinical settings.

Given the relatively small sample size in this study, there is a potential risk of overfitting, especially with the evaluation of 50 different model combinations (i.e., five feature selection methods and ten classification algorithms). To mitigate this risk, we applied repeated 7:3 train-test splits (with 100 repetitions) to ensure that our models were robust and not overly tuned to the training data. Overfitting was further addressed by performing feature selection exclusively on the training sets during each split, ensuring that the feature selection process did not use information from the test set, which could lead to an unrealistic estimation of model performance. Additionally, to assess model generalizability, we validated the best-performing model using an independent external dataset (GSE114860). The external validation further confirms the stability of our model and reduces the likelihood that our results are due to overfitting on the discovery cohort alone.

To better understand the potential biological relevance of the identified exo-miRNAs, we performed a functional enrichment analysis of their predicted target genes. The analysis revealed significant involvement in biological pathways related to insulin signaling, lipid metabolism, and inflammation—processes that are known to play central roles in the development and progression of GDM [39–41]. However, it is important to note that the associations observed in this study do not imply direct mechanistic causality. Instead, these pathways are implicated in the pathophysiology of GDM, and the identified exo-miRNAs may be predictive markers of these processes. Insulin resistance, a hallmark of GDM, is regulated by multiple complex signaling pathways, and while the exo-miRNAs identified in this study are associated with these pathways, their precise functional roles remain to be experimentally validated [42]. Similarly, dysregulated lipid metabolism, a key aspect of GDM pathogenesis, is known to contribute to insulin resistance and metabolic dysfunction [43]. However, further experimental studies are necessary to determine whether these exo-miRNAs directly modulate lipid metabolism in the context of GDM. Additionally, inflammation is a recognized contributor to the development of GDM, with chronic low-grade inflammation promoting insulin resistance and metabolic dysfunction [44]. While our findings suggest an association between the identified exo-miRNAs and inflammation-related pathways, their causal roles in this process require experimental validation.

The identification of a robust exo-miRNAs-based signature for GDM has significant clinical implications. First, the use of exo-miRNAs as noninvasive biomarkers for early detection of GDM offers a promising alternative to the traditional glucose tolerance tests. Exo-miRNAs can be easily isolated from plasma samples, which makes them highly accessible for routine screening in clinical settings. The high accuracy and generalizability of the XGBoost + Boruta model further enhance its potential as a reliable screening tool for GDM. In addition to early detection, the exo-miRNAs signature identified in this study could be used to personalize the management of GDM. By identifying women at high risk for GDM early in pregnancy, clinicians can implement targeted interventions to prevent or delay the onset of the disease. Furthermore, the integration of this exo-miRNAs signature into clinical practice could contribute to more precise monitoring of disease progression and treatment response, ultimately improving maternal and neonatal outcomes.

Despite the promising results, this study has several limitations. First, although the external validation dataset (GSE114860) provided important insights into the generalizability of the model, further validation in independent, prospective cohorts is needed to confirm the robustness of the exo-miRNAs signature. Additionally, while this study focused on exo-miRNAs future research could explore other types of noninvasive biomarkers, such as circulating exo-miRNAs or proteins, to further enhance predictive accuracy. Finally, the functional roles of the identified exo-miRNAs in GDM pathogenesis warrant further investigation, and experimental studies are needed to validate the biological mechanisms underlying their effects. Although our study identifies promising exo-miRNAs biomarkers for GDM, it is important to acknowledge the limitations of our analysis. Given the retrospective nature of the study, the findings should be interpreted with caution. The claims regarding the “early detection” and “precision management” of GDM are promising but are based on associations observed in the discovery and validation cohorts. These results require prospective clinical validation to establish their utility for early diagnosis and individualized management of GDM. Future studies, particularly prospective cohort studies, are essential to confirm these biomarkers’ predictive value and to assess their clinical applicability in real-world settings.

Conclusion

In conclusion, the results of this study demonstrate the successful identification of a robust exo-miRNAs-based predictive signature for GDM. The XGBoost + Boruta model exhibited the highest diagnostic accuracy and generalizability across independent cohorts, with an AUC greater than 0.90 in the discovery dataset and an accuracy above 80% in the external validation cohort. The identified exo-miRNAs, including miR-423-5p, miR-99a-5p, miR-148a-3p, miR-192-5p, and miR-122-5p, are implicated in key biological pathways involved in glucose metabolism, lipid homeostasis, and inflammation. These findings highlight the potential of using exo-miRNAs as noninvasive biomarkers for early detection and precision management of GDM.

Acknowledgements

Not applicable.

Authors’ contributions

PH. J. and J. H. Conceptualization, Methodology, Software, Validation, Investigation, Writing; SX. W. Data curation, Investigation; CX. D. Supervision, Project administration. All authors approved the final manuscript.

Funding

Not applicable.

Data availability

The datasets supporting the conclusions of this article are included within the article and its additional files.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Benioudakis E, Karlafti E, Bekiaridou A, Didangelos T, Papavramidis TS, Gestational Diabetes C, Cancer. Bariatric Surgery, and Weight Loss among Diabetes Mellitus Patients: A Mini Review of the Interplay of Multispecies Probiotics. Nutrients. 2021;14:192. 10.3390/nu14010192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Al Ismaili A, Al-Duqhaishi T, Al Rajaibi H, Al Waili K, Al Rasadi K, Nadar SK, et al. Antihypertensive Drugs and Perinatal Outcomes in Hypertensive Women Attending a Specialized Tertiary Hospital. Oman Med J. 2022;37:e354. 10.5001/omj.2022.43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Chen W, Li B, Gan K, Liu J, Yang Y, Lv X, et al. Gestational Weight Gain and Small for Gestational Age in Obese Women: A Systematic Review and Meta-Analysis. Int J Endocrinol. 2023;2023:3048171. 10.1155/2023/3048171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Zhang M, Li Q, Wang K-L, Dong Y, Mu Y-T, Cao Y-M, et al. Lipolysis and gestational diabetes mellitus onset: a case-cohort genome-wide association study in Chinese. J Transl Med. 2023;21:47. 10.1186/s12967-023-03902-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Pezel T, Sideris G, Dillinger J-G, Logeart D, Manzo-Silberman S, Cohen-Solal A, et al. Coronary Computed Tomography Angiography Analysis of Calcium Content to Identify Non-culprit Vulnerable Plaques in Patients With Acute Coronary Syndrome. Front Cardiovasc Med. 2022;9:876730. 10.3389/fcvm.2022.876730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Peng X, Wang Q, Li W, Ge G, Peng J, Xu Y, et al. Comprehensive overview of microRNA function in rheumatoid arthritis. Bone Res. 2023;11:8. 10.1038/s41413-023-00244-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ke X, Zhang W. Pro-inflammatory activity of long noncoding RNA FOXD2-AS1 in Achilles tendinopathy. J Orthop Surg Res. 2023;18:361. 10.1186/s13018-023-03681-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Liu H, Huang Y, Huang M, Huang Z, Wang Q, Qing L, et al. Current Status, Opportunities, and Challenges of Exosomes in Oral Cancer Diagnosis and Treatment. Int J Nanomed. 2022;17:2679–705. 10.2147/IJN.S365594. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Liu Z-N, Jiang Y, Liu X-Q, Yang M-M, Chen C, Zhao B-H, et al. MiRNAs in Gestational Diabetes Mellitus: Potential Mechanisms and Clinical Applications. J Diabetes Res. 2021;2021:4632745. 10.1155/2021/4632745. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Liu W, Feng Y, Wang X, Ding J, Li H, Guan H et al. Human umbilical vein endothelial cells-derived exosomes enhance cardiac function after acute myocardial infarction by activating the PI3K/AKT signaling pathway. Bioengineered 13:8850–65. 10.1080/21655979.2022.2056317. [DOI] [PMC free article] [PubMed]
  • 11.Li H, Sui T, Chen X, Gu Y, Luo X, Liu Y, et al. Screening and identification of serum exosomal protein ZNF587B in liquid biopsy for ovarian cancer diagnosis. Am J Cancer Res. 2024;14:1904–13. 10.62347/RBTM1834. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ghafourian M, Mahdavi R, Akbari Jonoush Z, Sadeghi M, Ghadiri N, Farzaneh M, et al. The implications of exosomes in pregnancy: emerging as new diagnostic markers and therapeutics targets. Cell Commun Signal. 2022;20:51. 10.1186/s12964-022-00853-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Alabdaljabar MS, Hasan B, Noseworthy PA, Maalouf JF, Ammash NM, Hashmi SK. Machine Learning in Cardiology: A Potential Real-World Solution in Low- and Middle-Income Countries. J Multidiscip Healthc. 2023;16:285–95. 10.2147/JMDH.S383810. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.To NM, Ngo QC, Polus B, Dinh MN, Khandoker A, Menon A, et al. Refine XGBoost with SHAP explainability for non-invasive early detection of diabetic kidney disease: Estimated cardiac output as a potential indicator. Comput Methods Programs Biomed. 2025;273:109122. 10.1016/j.cmpb.2025.109122. [DOI] [PubMed] [Google Scholar]
  • 15.Ward A, Kron B, Lozama A, Sandhu A, Khandelwal A, Rodriguez F, et al. Elevated Lipoprotein(a) Independently Increases Risk for Short-Term Atherosclerotic Cardiovascular Events in Machine Learning Predictive Models. JACC Adv. 2025;4(11 Pt 1):102253. 10.1016/j.jacadv.2025.102253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Guo B, Luo C, Lu Y, Wu Y, Xie S, Xia F, et al. Long-Term Survival and Beneficiaries of Adjuvant Anti-PD-1 Therapy in Resected Hepatocellular Carcinoma. Ann Surg Oncol. 2025. 10.1245/s10434-025-18549-2. [DOI] [PubMed] [Google Scholar]
  • 17.Ma J, Li N, Guarnera M, Jiang F. Quantification of Plasma miRNAs by Digital PCR for Cancer Diagnosis. Biomark Insights. 2013;8:127–36. 10.4137/BMI.S13154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Gao J, Zou Y, Lv X-Y, Chen L, Hou X-G. Novel insights into immune-related genes associated with type 2 diabetes mellitus-related cognitive impairment. World J Diabetes. 2024;15:735–57. 10.4239/wjd.v15.i4.735. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Modhukur V, Sharma S, Mondal M, Lawarde A, Kask K, Sharma R, et al. Machine Learning Approaches to Classify Primary and Metastatic Cancers Using Tissue of Origin-Based DNA Methylation Profiles. Cancers (Basel). 2021;13:3768. 10.3390/cancers13153768. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yi W, Sun A, Liu M, Liu X, Zhang W, Dai Q. Comparative Study on Feature Selection in Protein Structure and Function Prediction. Comput Math Methods Med. 2022;2022:1650693. 10.1155/2022/1650693. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Moszczuk B, Krata N, Rudnicki W, Foroncewicz B, Cysewski D, Pączek L, et al. Osteopontin—A Potential Biomarker for IgA Nephropathy: Machine Learning Application. Biomedicines. 2022;10:734. 10.3390/biomedicines10040734. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mallik S, Seth S, Bhadra T, Zhao Z, Mallik S, Seth S, et al. A Linear Regression and Deep Learning Approach for Detecting Reliable Genetic Alterations in Cancer Using DNA Methylation and Gene Expression Data. Genes. 2020;11. 10.3390/genes11080931. [DOI] [PMC free article] [PubMed]
  • 23.Tang T, Li Z, Lu X, Du J. Development and validation of a risk prediction model for anxiety or depression among patients with chronic obstructive pulmonary disease between 2018 and 2020. Ann Med. 54:2181–90. 10.1080/07853890.2022.2105394. [DOI] [PMC free article] [PubMed]
  • 24.Bertsimas D, Borenstein A, Mingardi L, Nohadani O, Orfanoudaki A, Stellato B, et al. Personalized prescription of ACEI/ARBs for hypertensive COVID-19 patients. Health Care Manag Sci. 2021;24:339–55. 10.1007/s10729-021-09545-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lv C, Zhong Y, Hu Y, Tang Y. Potential liquid biopsy markers of exosomal microRNAs in renal interstitial fibrosis blood and urine. Indian J Pathol Microbiol. 2025;68:279–86. 10.4103/ijpm.ijpm_265_24. [DOI] [PubMed] [Google Scholar]
  • 26.Ren Z-J, Zhao Y, Wang G, Miao L, Zhang Z-C, Ma L, et al. Identification of differentially expressed miRNAs derived from serum exosomes associated with gastric cancer by microarray analysis. Clin Chim Acta. 2022;531:25–35. 10.1016/j.cca.2022.03.010. [DOI] [PubMed] [Google Scholar]
  • 27.Huang X, Wei L, Li M, Zhang Y, Kuang S, Shen Z, et al. Diabetic Macrophage Exosomal miR-381-3p Inhibits Epithelial Cell Autophagy Via NR5A2. Int Dent J. 2024;74:823–35. 10.1016/j.identj.2024.02.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Seth S, Mallik S, Bhadra T, Zhao Z. Dimensionality Reduction and Louvain Agglomerative Hierarchical Clustering for Cluster-Specified Frequent Biomarker Discovery in Single-Cell Sequencing Data. Front Genet. 2022;13. 10.3389/fgene.2022.828479. [DOI] [PMC free article] [PubMed]
  • 29.Ranches G, Zeidler M, Kessler R, Hoelzl M, Hess MW, Vosper J, et al. Exosomal mitochondrial tRNAs and miRNAs as potential predictors of inflammation in renal proximal tubular epithelial cells. Mol Ther Nucleic Acids. 2022;28:794–813. 10.1016/j.omtn.2022.04.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Yang W, Wang J, Chen Z, Chen J, Meng Y, Chen L, et al. NFE2 Induces miR-423-5p to Promote Gluconeogenesis and Hyperglycemia by Repressing the Hepatic FAM3A-ATP-Akt Pathway. Diabetes. 2017;66:1819–32. 10.2337/db16-1172. [DOI] [PubMed] [Google Scholar]
  • 31.Lee EB, Sung PS, Kim J-H, Park DJ, Hur W, Yoon SK. microRNA-99a Restricts Replication of Hepatitis C Virus by Targeting mTOR and de novo Lipogenesis. Viruses. 2020;12:696. 10.3390/v12070696. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zhou M, Liu H, Hui J, Chen Q, Zhao Y, Wang H, et al. Extracellular Vesicle-Packaged miR-99a Reprograms Fibroblasts to Create an Inflammatory Niche that Drives Colorectal Cancer Metastasis. Cancer Res. 2025. 10.1158/0008-5472.CAN-25-0663. [DOI] [PubMed] [Google Scholar]
  • 33.Wang F, Ge J, Huang S, Zhou C, Sun Z, Song Y, et al. KLF5/LINC00346/miR–148a–3p axis regulates inflammation and endothelial cell injury in atherosclerosis. Int J Mol Med. 2021;48:152. 10.3892/ijmm.2021.4985. [DOI] [PubMed] [Google Scholar]
  • 34.Yin M, Lu J, Guo Z, Zhang Y, Liu J, Wu T, et al. Reduced SULT2B1b expression alleviates ox-LDL-induced inflammation by upregulating miR-148-3P via inhibiting the IKKβ/NF-κB pathway in macrophages. Aging. 2021;13:3428–42. 10.18632/aging.202273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Ma L, Song H, Zhang C-Y, Hou D. MiR-192-5p Ameliorates Hepatic Lipid Metabolism in Non-Alcoholic Fatty Liver Disease by Targeting Yy1. Biomolecules. 2023;14:34. 10.3390/biom14010034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Raitoharju E, Seppälä I, Lyytikäinen L-P, Viikari J, Ala-Korpela M, Soininen P, et al. Blood hsa-miR-122-5p and hsa-miR-885-5p levels associate with fatty liver and related lipoprotein metabolism-The Young Finns Study. Sci Rep. 2016;6:38262. 10.1038/srep38262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Shboul ZA, Chen J, Iftekharuddin M. Prediction of Molecular Mutations in Diffuse Low-Grade Gliomas using MR Imaging Features. Sci Rep. 2020;10:3711. 10.1038/s41598-020-60550-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Ivanoska I, Trivodaliev K, Kalajdziski S, Zanin M. Statistical and Machine Learning Link Selection Methods for Brain Functional Networks: Review and Comparison. Brain Sci. 2021;11:735. 10.3390/brainsci11060735. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Shen Y, Zhang M, Zhang Z, Li M, Chen X, Zhu W, et al. Exosome miRNA Profiles: The Reduced Expression of miRNA-27a-5p and Predictive miRNAs for Gestational Diabetes Mellitus. FASEB J. 2025;39:e70710. 10.1096/fj.202500470RR. [DOI] [PubMed] [Google Scholar]
  • 40.Ye Z, Wang S, Huang X, Chen P, Deng L, Li S, et al. Plasma Exosomal miRNAs Associated With Metabolism as Early Predictor of Gestational Diabetes Mellitus. Diabetes. 2022;71:2272–83. 10.2337/db21-0909. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Razo-Azamar M, Nambo-Venegas R, Quevedo IR, Juárez-Luna G, Salomon C, Guevara-Cruz M, et al. Early-Pregnancy Serum Maternal and Placenta-Derived Exosomes miRNAs Vary Based on Pancreatic β-Cell Function in GDM. J Clin Endocrinol Metab. 2024;109:1526–39. 10.1210/clinem/dgad751. [DOI] [PubMed] [Google Scholar]
  • 42.Fakhrul-Alam M, Sharmin-Jahan null, Mashfiqul-Hasan null, Nusrat-Sultana null, Mohona-Zaman null, Rakibul-Hasan M et al. Insulin secretory defect may be the major determinant of GDM in lean mothers. J Clin Transl Endocrinol. 2020;20:100226. 10.1016/j.jcte.2020.100226. [DOI] [PMC free article] [PubMed]
  • 43.Zhang Z, Piro AL, Allalou A, Alexeeff SE, Dai FF, Gunderson EP, et al. Prolactin and Maternal Metabolism in Women With a Recent GDM Pregnancy and Links to Future T2D: The SWIFT Study. J Clin Endocrinol Metab. 2022;107:2652–65. 10.1210/clinem/dgac346. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Rancourt RC, Ott R, Ziska T, Schellong K, Melchior K, Henrich W, et al. Visceral Adipose Tissue Inflammatory Factors (TNF-Alpha, SOCS3) in Gestational Diabetes (GDM): Epigenetics as a Clue in GDM Pathophysiology. Int J Mol Sci. 2020;21:479. 10.3390/ijms21020479. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets supporting the conclusions of this article are included within the article and its additional files.


Articles from BMC Pregnancy and Childbirth are provided here courtesy of BMC

RESOURCES