Abstract
Background
Gestational diabetes mellitus (GDM) is a prevalent pregnancy complication that can pose numerous adverse health effects on both mothers and newborns. Accurate prediction of the risk of GDM serves as a valuable supplement to prenatal education and clinical decision-making. Compared with traditional prediction models, artificial intelligence (AI) algorithms have demonstrated higher predictive accuracy and stronger individualization capabilities. However, the application of AI models in GDM prediction is still in a developmental stage, and their performance and clinical utility have not been thoroughly evaluated. Therefore, this study aims to systematically review and critically appraise the published predictive performance of AI models for GDM prediction and to offer insights for future research and practical application.
Methods
A systematic literature search will be performed across six databases (PubMed, Web of Science, Cochrane Library, Scopus, EMBASE, and OVID). Screening of titles and abstracts, full-text review, and data extraction will be independently completed by two authors. Qualitative data on the characteristics of the included studies, methodological quality, and the applicability of models will be summarized through narrative descriptions and tabulated formats. For models with predictive performance data from multiple studies, a random-effects meta-analysis or meta-regression will be employed to synthesize the findings, considering potential heterogeneity.
Ethics and dissemination
Ethical approval is deemed not applicable for this systematic review and meta-analysis. The findings will be based on published literature, disseminated through publication in a peer-reviewed journal, and presented at major conferences focused on clinical healthcare.
Systematic review registration
PROSPERO registration number CRD42025645913
Supplementary Information
The online version contains supplementary material available at 10.1186/s13643-026-03167-0.
Keywords: Artificial intelligence, Gestational diabetes mellitus, Prediction model, Meta-analysis, Protocols
Introduction
Gestational diabetes mellitus (GDM) is defined as glucose metabolism abnormalities of varying severity that first emerge during pregnancy, representing a prevalent metabolic condition associated with pregnancy [1]. According to the 10th edition of the “Diabetes Atlas” published by the International Diabetes Federation, 16.7% of women of reproductive age experience elevated blood glucose levels during pregnancy, with approximately 80.3% of these cases attributed to GDM [2]. GDM significantly increases the risk of adverse pregnancy and delivery outcomes, including macrosomia, preterm birth, preeclampsia, and the future development of type 2 diabetes [3, 4]. Additionally, offspring of women with GDM are more susceptible to neonatal hypoglycemia, respiratory distress syndrome, and long-term metabolic disorders [5–7]. Therefore, early identification and management of GDM are critical for improving maternal and neonatal health outcomes.
Accurate prediction of GDM risk is crucial for timely intervention and effective management. Traditional risk assessment methods, which are usually based on clinical risk factors such as age, body mass index, and family history of diabetes, have limitations in predictive accuracy and accounting for individual variability [8–11]. However, the advent of artificial intelligence (AI) algorithms has markedly enhanced the precision of GDM prediction by autonomously detecting patterns within vast datasets and integrating multiple data sources—spanning clinical, biochemical, and lifestyle factors—to provide a holistic assessment of GDM risk [12–14].
Currently, the application of AI algorithms in GDM risk prediction extends across two major areas: machine learning and deep learning [15, 16]. Machine learning algorithms, including Gaussian Naive Bayes, decision trees, random forest (RF), gradient boosting machines (GBM), and support vector machines (SVM), have proven effective in predicting GDM during early pregnancy, achieving remarkable predictive performance [12, 17–19]. However, the predictive performance of these algorithms varies significantly across different studies. For example, a machine learning model developed by Gallardo et al. based on routine early-pregnancy examination data demonstrated high predictive accuracy in a specific population but may perform poorly in other GDM populations due to differences in data characteristics [20]. Meanwhile, deep learning algorithms, like artificial neural networks, also exhibit noticeable fluctuations in performance across different populations [21–23]. This discrepancy not only reflects the intrinsic constraints of the algorithms but also highlights the profound impact of dataset characteristics and population diversity on model performance [24, 25]. This underscores the need for further validation of the applicability and generalizability of AI algorithms across diverse populations, despite their promising potential in GDM prediction [14].
Despite the growing interest in AI-driven predictive models for GDM, the field remains in its developmental stage [18, 26, 27]. The performance and clinical applicability of these models vary widely across studies, and there is a need for a systematic evaluation of predictive accuracy, generalizability, and potential impact on clinical practice [28, 29]. To address this gap, we propose a systematic review and meta-analysis of published studies on AI algorithms for GDM prediction. This protocol outlines the methodology for identifying, appraising, and synthesizing evidence on AI prediction models, aiming to provide a comprehensive overview of their status and potential for future application in a clinical setting.
Research aims
This study aims to systematically evaluate available evidence on AI algorithms for predicting GDM. It will identify existing diagnostic prediction models and establish the most effective ones to inform clinical decision-making. The specific objectives of this systematic review and meta-analysis are:
To identify and catalogue existing AI-based diagnostic prediction models for GDM and qualitatively describe their features in the included studies.
To summarize and compare the predictive performance of current AI-based diagnostic prediction models for GDM.
To critically appraise the methodological quality and reporting standards of included studies.
To identify the most effective and highest-performing diagnostic prediction model for GDM and provide evidence for clinical decision-making.
Methods and design
Protocol registration and reporting
The present study protocol has been registered within PROSPERO under registration number (No. CRD42025645913) and is being reported in accordance with the reporting guidance provided in the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols (PRISMA-P) guidelines (see checklist in Additional file 1). The methodology for data extraction and management will be guided by the Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHRAMS) checklist and the recommendations reported by Debray et al. [30, 31].
Eligibility criteria
The study selection will be based on predefined eligibility criteria established using the PICOTS system [31], as detailed in Table 1. PICOTS is a modification of the established PICO system, tailored to meet the specific requirements of systematic reviews of prediction models, and additionally considers time (prediction period and timing for model use) and clinical setting [31].
Table 1.
Eligibility criteria for the systematic review framed using the PICOTS system
| Inclusion criteria | Exclusion criteria | |
|---|---|---|
| Population | Pregnant women from around the globe who are in the first or second trimester and have not been diagnosed with GDM |
Women with pre-gestational diabetes (type 1 and type 2 diabetes) |
| Women with diagnosis of GDM | ||
| Index |
Development or external validation of a diagnostic prediction model for GDM (e.g. prediction model for the diagnosis of GDM) |
Prognostic prediction model for women (e.g., prediction model for women with GDM to predict pregnancy complications) |
| Comparator | Not applicable | |
| Outcome | GDM (any diagnostic criteria) | |
| Timing | The period of gestation before being diagnosed with GDM | The period of gestation following a confirmed diagnosis of GDM |
| Setting | Diagnostic prediction models intended for use by healthcare professionals in prenatal clinical settings prior to the diagnosis of GDM during pregnancy, aimed at informing clinical decision-making for physicians | Diagnostic prediction models, which are intended to be used after the diagnosis of GDM |
GDM gestational diabetes mellitus
Population
Studies proposing diagnostic prediction models for pregnant women from around the globe who are in the first or second trimester and have not been diagnosed with GDM will be considered for inclusion. Studies proposing models for women with pre-gestational diabetes (type 1 and type 2 diabetes) will be excluded. Additionally, studies proposing models for women who have been diagnosed with GDM will also be excluded.
Intervention
Studies involving the development of prediction models (with or without external validation) and external model validation studies (with or without model updating) that aim to clinical decision-making regarding the diagnosis of GDM or prediction of risk will be considered for inclusion.
Outcomes
The primary outcome to be predicted in the included studies will be the diagnosis of GDM (based on any diagnostic criteria). Secondary outcomes and effect measures will be defined by the study authors. Predictions for all these outcomes will be conducted before the diagnosis of GDM.
Timing
Included studies need to report on prediction models for timing before pregnancy or during the period before being diagnosed with GDM. Prediction models for the period of gestation following a confirmed diagnosis of GDM will be excluded.
Setting
Diagnostic prediction models intended for use by healthcare professionals in prenatal clinical settings prior to the diagnosis of GDM during pregnancy, aimed at informing clinical decisions for physicians will be considered for inclusion. Models intended to be used after the diagnosis of GDM will be excluded.
Types of studies
All quantitative studies (including cohort studies, case-control studies, and clinical trials) that investigate the application of AI algorithms to predict the occurrence of GDM in pregnant women and involve the development and/or validation of at least one prediction model will be eligible for this study, without any restrictions on publication date. Provided that the full text of the literature is accessible, a comprehensive review and analysis of eligible studies from any region of the world will be conducted, with pregnant women as the subjects. The initial literature search will cover the period from the inception of the database to 1 June 2025. The literature search process will remain ongoing until the study is completed, with supplementary searches conducted periodically to incorporate the latest developments.
Condition/domain being studied
This review will focus on GDM as the condition of interest. Given the variations in GDM diagnostic criteria across different regions and studies, we will standardize all GDM diagnostic data from all included studies to align with the International Association of Diabetes and Pregnancy Study Groups (IADPSG) criteria [32]. According to the IADPSG criteria, GDM is diagnosed using the 75-g oral glucose tolerance test (OGTT) conducted during 24–28 weeks of gestation, with the diagnostic criteria being: fasting blood level ≥ 5.1 mmol/L, a 1-h glucose level ≥ 10.0 mmol/L, or a 2-h glucose level ≥ 8.5 mmol/L [32]. For original studies using diagnostic thresholds that differ from the IADPSG criteria, conversions will be made by leveraging the correlation between blood glucose levels and the risk of GDM. Additionally, for studies that cannot be standardized to the IADPSG criteria, a sensitivity analysis will be performed to evaluate the impact of the diverse diagnostic criteria on the pooled effect size.
Search strategy
A comprehensive search will be conducted across six databases, including PubMed, Web of Science, Cochrane Library, Scopus, EMBASE, and OVID. To enhance the accuracy of the search results and minimize the possibility of missing relevant studies, the author team has devised a rigorous search strategy by combining Medical Subject Headings and keywords. Additionally, the bibliometric networks have been visualized using the VOSviewer software tool for analysis (see Fig. 1). The tailored search strategies for each database are outlined in Table 2, summarizing the key search terms related to the target population and AI algorithms. Detailed descriptions of the search strategies for all databases are described in Additional file 2.
Fig. 1.
Cluster analysis of keywords from databases
Table 2.
Key terms for developing search strategy
| Population (P) | Intervention (I) |
|---|---|
|
‘Pregnancy induced diabetes’ ‘Diabetes in pregnancy’ ‘Gestational diabetes mellitus’ ‘Maternal diabetes’ ‘Pregnancy diabetes mellitus’ ‘GDM’ |
‘Artificial Intelligence’ ‘Deep Learning’ ‘Machine Learning’ ‘Knowledge Acquisition (Computer)’ ‘Hierarchical Learning’ ‘Ensemble Learning’ ‘Transfer Learning’ ‘Knowledge Representation (Computer)’ ‘Diagnostic prediction model’ ‘Risk prediction’ ‘Risk assessment’ |
To evaluate the effectiveness and accuracy of the keywords, search strategies, and search scope in identifying relevant literature, the preliminary literature search will be conducted in the PubMed database. Subsequently, the search strategies will be adjusted accordingly based on the results obtained, and comprehensive searches will be carried out in the remaining electronic databases. In addition, we will also review the reference lists of relevant literature, particularly systematic reviews related to the topic of this study, and conduct additional searches in the electronic databases to minimized the omission of the key literature as much as possible. All searches will be conducted under the supervision of an academic librarian.
Screening and selection procedure
Following the completion of the initial search based on the search strategies, all retrieved records from the databases will be exported into the Endnote 21 reference management software, and then the “Find Duplicates” function of the software will be used to remove duplicates. Each literature will undergo a dual screening by two authors based on the title and abstract to initially identify studies that are relevant. Following this initial screening, the same process will be applied to the full-text screening, with a rigorous evaluation conducted according to the exclusion. The rationale for excluding each literature will be recorded. Any disagreements between the two authors will initially be resolved through discussion. If consensus cannot be achieved, a third author will be consulted. The PRISMA flow diagram (see Fig. 2) will be used to depict the literature screening process.
Fig. 2.
Preferred reporting items for systematic reviews and meta-analysis flow diagram of the identification, screening, and eligibility of included articles
Data extraction and management
Data extraction from selected studies will be guided by CHARMS checklist [30]. Two authors will independently perform data extraction using a pre-formatted extraction template in Microsoft Excel (Microsoft, WA, USA). The following data details will be extracted: (1) title of the study; (2) author of the study; (3) time of publication; (4) country of the study; (5) type of study; (6) dataset size (including sample size and any information on missing data); (7) source of participants; (8) predicted outcomes; (9) diagnostic criteria for GDM; (10) potential predictors; (11) incidence of GDM; (12) odds ratio or risk ratio for predictor; (13) type of AI model (and its algorithms); (14) model performance (properties of discrimination with confidence intervals, calibration, classification, and overall performance); (15) model validation method. If multiple models are used in a study, we will extract performance data from all models to enable comparative analysis. In cases where the information presented in the literature is ambiguous, the researchers will proactively contact the corresponding author to acquire the relevant information. If the required data cannot be obtained from the author, the relevant literature will be excluded from further consideration. After the data extraction is completed, the two authors will cross-check each other’s extracted information. Any discrepancies will be resolved by a third-party arbitrator.
In addition, the following data related to AI models will be extracted:
Type of AI model. The AI models are classified into machine learning and deep learning. Machine learning models encompass algorithms, including SVM, decision trees, RF, k-nearest neighbors, GBM, and other relevant algorithms. Deep learning models mainly consist of algorithms based on neural networks.
Model validation method. The approaches for validating the predictive performance of AI models mainly consist of cross-validation, K-fold cross-validation, external validation, and additional specific validation methods.
Sample size and missing data. The sample size of each study and any reported information on missing data will be recorded, including the extent of missingness and how the authors addressed these missing data.
Critical appraisal
The prediction model risk of bias assessment tool (PROBAST) will be employed to evaluate the methodological quality (risk of bias) and relevance to the review question (applicability) of the included studies [33]. PROBAST is composed of four key domains: participants, predictors, outcomes, and analysis. The risk of bias for each domain is categorized as “high”, “low”, or “unclear”. The evaluation of applicability is employed to determine whether the model development/validation studies align with our systematic review question regarding the target population, predictors, or outcomes of interest. The above process will be conducted independently by two authors. Any discrepancies will be resolved through discussion and consultation with a third author.
Data synthesis and analysis
Data synthesis
After the completion of data extraction, a narrative synthesis method will be utilized to systematically elucidate the quantitative data derived from the included studies, focusing on the characteristics and features of the data. This will cover key aspects such as predictive factors, performance metrics, classification indicators, and descriptive analyses of critical items. To facilitate comparison, the findings of each included study will ultimately be presented in tabular form. For studies reporting multiple AI models, all models will be included in the narrative synthesis to provide comparative insights.
Meta-analysis and investigation of heterogeneity
If the AI models identified in the included studies demonstrate sufficient homogeneity, data synthesis will be conducted through meta-analysis stratified by the type of AI modelling study. The following conditions will be deemed to satisfy homogeneity:
AI model development studies where the target population, predicted outcomes, and anticipated time frames for model application are comparable or
Several validation studies targeting the identical AI model.
During the meta-analysis, a random-effects model will be applied to pool and analyze the performance metrics of AI models (such as discrimination and calibration), thereby estimating the overall average performance of the models included in the studies. The performance of the diagnostic prediction model will be based on the following metrics [34–36], detailed in Table 3. To further conduct a comprehensive assessment of the models’ overall performance, multivariate meta-analysis will be employed for the joint analysis of discriminative ability and calibration, while considering their correlation. Besides, the restricted maximum likelihood estimation and the Hartung-Knapp-Sidik-Jonkman methods will be applied to assess between-study heterogeneity and the 95% confidence intervals for average model performance.
Table 3.
Summary of measuring performance of diagnostic prediction models
| Items | Explanation | Performance measures/statistics | Values for better performance | Visualization |
|---|---|---|---|---|
| Discrimination | To measure the models’ ability to distinguish between cases and non-cases | AUROC | Higher | ROC curve |
| AUPRC | Higher | PRC curve | ||
| ACC | Higher | |||
| BER | Lower | |||
| D statistics | Lower | |||
| MCC | Higher | |||
| F1 score | Higher | |||
| Log-rank | Lower, p > 0.05 | |||
| Calibration | To evaluate consistency between predicted probabilities of the model and the actual observed results | Calibration curve, slope, intercept | Slope closer to 1 and intercept closer to 0 | Calibration plot |
| Reliability-deviation | Lower | |||
| Reliability-within-bin variation | Lower | |||
| Reliability-within-bin covariance | Higher | |||
| Resolution | Higher | |||
| Predictive range | Higher | |||
| Hosmer–Lemeshow test | Higher | |||
| Total O:E ratios | Higher | |||
| Classification | To measure the models’ ability to correctly classify individuals as cases or non-cases | Sensitivity | Higher | Reclassification scatter plot |
| Specificity | Higher | |||
| PPV | Higher | |||
| NPV | Higher | |||
| RMSE via neighborhood estimate | Lower | |||
| Overall performance | To comprehensively evaluate the overall performance of the models, including its accuracy and practicability | Brier score | Lower | |
| Brier skill score | Higher | |||
| Prediction squared error | Lower | |||
| Decision curve analysis | / | DCA plot |
AUROC area under receiver operating characteristic curve, AUPRC area under precision-recall curve, ACC accuracy, BER balanced error rate, MCC Matthews correlation coefficient, PPV positive predictive value, NPV negative predictive value, RMSE root-mean-squared error
Heterogeneity in performance metrics is anticipated due to differences in study design and population. The potential range of model performance across different populations can be estimated by calculating an approximate 95% prediction interval. This will help assess the generalizability of AI models across diverse settings, addressing concerns related to population diversity and language bias raised in the limitations. Variations in case-mix within each study will be quantified by estimating the standard deviation of the linear predictor [31]. When performance metrics or uncertainty measures are unreported, we will estimate these values based on sample size, event rate, or confusion matrices using normal approximation [31]. If estimation is not feasible, the missing data will be recorded and included only in the narrative synthesis. Statistical heterogeneity will be assessed using the I2 test, with I2 values above 50% indicating moderate to high heterogeneity. Potential sources of heterogeneity will be explored through meta-regression analysis (P < 0.05).
Sensitivity analysis
To assess the robustness of our primary findings, we will conduct a series of sensitivity analyses. First, to address potential unit-of-analysis issues arising from studies that report multiple AI models, our primary meta-analysis will include only the primary model (as specified by study authors) or the model with the highest accuracy. We will then perform a sensitivity analysis by including all models in a separate meta-analysis, where appropriate, to evaluate the impact of our model selection strategy on the pooled estimates.
If enough studies are included, subgroup analyses will be performed based on
Region—categorized according to the Organization for Economic Co-operation and Development classification into low/middle-income countries and high-income countries.
Population characteristics—including ethnicity, medical history, and pregnancy trimester—to assess differences in model performance across various subgroups.
Type of AI modelling study—development or validation.
Type of AI models—categorized into machine learning (e.g., RF, SVM, gradient boosting machine) and deep learning (e.g., artificial neural networks), with further comparisons across specific algorithm families if data permit.
Predictors—type or number of predictors to assess their impact on model performance.
Study quality—risk of bias.
This overall process will be conducted using Stata 17.0 software, in accordance with the Meta-analysis of Observational Studies in Epidemiology (MOOSE) guidelines [37].
Reporting and presentation of findings
The findings of this study will be presented in accordance with the transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD) statement and the PRISMA guideline [38, 39]. The GRADE approach (grading of recommendations, assessment, development, and evaluation) will also be applied to assess the certainty of the evidence [40]. The certainty of the evidence will be categorized into “high”, “moderate”, “low”, or “very low”.
Discussion
This systematic review and meta-analysis aims to synthesize the current evidence on AI algorithms for predicting GDM and to explore the potential of AI-driven models in clinical practice. By systematically evaluating available studies, this review seeks to identify the most effective AI algorithms based on their predictive accuracy, which may ultimately facilitate the early detection of high-risk populations for GDM. The findings are expected to advance the application of AI algorithms in GDM prediction and contribute valuable insights to the broader field of medical AI.
Moreover, for these algorithms to achieve meaningful clinical impact, several critical factors beyond predictive accuracy should be considered. First, clinical integration requires that AI models be seamlessly embedded into antenatal care workflows, with clear guidance on when and how clinicians should act on predictions. Second, model interpretability is essential for clinician trust. Many high-performing algorithms, particularly deep learning models, operate as “black boxes,” obscuring the rationale behind risk assessments. Emerging explainable AI techniques, such as SHapley Additive exPlanations and Local Interpretable Model-agnostic Explanations, can help address this challenge by quantifying the contribution of individual risk factors to each prediction, thereby enabling clinicians to understand why a specific patient is classified as high-risk.
However, this protocol also has several limitations. First, by restricting our search to studies published in English, we may introduce language bias, potentially excluding relevant studies published in other languages. Consequently, our findings may not fully capture the global distribution of AI-based GDM prediction research, particularly regarding research progress in non-English speaking regions. Second, as this review will primarily focus on model development and validation studies, evidence on clinical implementation, interpretability, and cost-effectiveness may remain limited. Third, included studies may rarely report on ethical considerations such as data privacy, fairness, or bias mitigation, which could limit our ability to assess their readiness for clinical deployment. Nevertheless, this review is expected to systematically synthesize the current state of research on AI algorithms in GDM prediction and provide a reference for future studies and clinical practice.
Supplementary Information
Additional file 1: PRISMA 2020 Checklist.
Additional file 2: Search strategy used for the electronic databases.
Acknowledgements
Authors would like to acknowledge all clinicians, staff, and patient stakeholders who participated in the co-design of this protocol.
Abbreviations
- GDM
Gestational diabetes mellitus
- AI
Artificial intelligence
- RF
Random Forest
- GBM
Gradient boosting machines
- SVM
Support vector machines
- PRISMA-P guidelines
Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols guidelines
- CHRAMS checklist
Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies checklist
- IADPSG criteria
International Association of Diabetes and Pregnancy Study Groups criteria
- OGTT
Oral Glucose Tolerance Test
- PROBAST
Prediction Model Risk of Bias Assessment Tool
- MOOSE guidelines
Meta-analysis of Observational Studies in Epidemiology guidelines
- TRIPOD statement
Transparent Reporting of A Multivariable Prediction Model for Individual Prognosis Or Diagnosis statement
Authors’ contributions
YN.L and MY.L designed and developed the research question. YN.L and JY.S developed the search strategy. YN.L registered the protocol and wrote the first draft of the manuscript. YH.S and ZY.L are the guarantors. YP.Y and JY.S developed the risk of bias assessment strategy. AR.D and ZL.Z revised the manuscript. All authors read and approved the final manuscript.
Funding
This work was supported by the National Natural Science Foundation of China (No. 32070189), Hunan Province Health Commission Scientific Research Project (No. D202314038701), and Postgraduate Scientific Research Innovation Project of Hunan Province (No. CX20251466).
Data availability
The datasets during and/or analyzed during the current study are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
All authors declare that they have no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yingni Liang and Meiyan Luo contributed equally to this work.
Contributor Information
Yinhua Su, Email: 382373646@qq.com.
Zhongyu Li, Email: lzhy1023@hotmail.com.
References
- 1.Cate JJM, Bloom E, Chu A, Bauer ST, Kuller JA, Dotters-Katz SK. Suboptimally controlled diabetes in pregnancy: a review to guide antepartum and delivery management. Obstet Gynecol Surv. 2024;79(6):348–65. [DOI] [PubMed] [Google Scholar]
- 2.Wang H, Li N, Chivese T, Werfalli M, Sun H, Yuen L, et al. IDF diabetes atlas: estimation of global and regional gestational diabetes mellitus prevalence for 2021 by International Association of Diabetes in Pregnancy Study Group’s criteria. Diabetes Res Clin Pract. 2022;183:109050. [DOI] [PubMed] [Google Scholar]
- 3.Sweeting A, Wong J, Murphy HR, Ross GP. A clinical update on gestational diabetes mellitus. Endocr Rev. 2022;43(5):763–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Cho NH, Shaw JE, Karuranga S, Huang Y, da Rocha Fernandes JD, Ohlrogge AW, et al. IDF diabetes atlas: global estimates of diabetes prevalence for 2017 and projections for 2045. Diabetes Res Clin Pract. 2018;138:271–81. [DOI] [PubMed] [Google Scholar]
- 5.Yang F, Liu H, Ding C. Gestational diabetes mellitus and risk of neonatal respiratory distress syndrome: a systematic review and meta-analysis. Diabetol Metab Syndr. 2024;16(1):294. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Zhang Y, Chen L, Ouyang Y, Wang X, Fu T, Yan G, et al. A new classification method for gestational diabetes mellitus: a study on the relationship between abnormal blood glucose values at different time points in oral glucose tolerance test and adverse maternal and neonatal outcomes in pregnant women with gestational diabetes mellitus. AJOG Glob Rep. 2024;4(4):100390. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Rizzo HE, Escaname EN, Alana NB, Lavender E, Gelfond J, Fernandez R, et al. Maternal diabetes and obesity influence the fetal epigenome in a largely Hispanic population. Clin Epigenetics. 2020;12(1):34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Darzi AJ, Busse JW, Torabiardakani K, Phillips M, Thabane L, Bhandari M, et al. Risk assessment models: considerations prior to use in clinical practice. Eye (Lond). 2024. 10.1038/s41433-024-03557-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Xing J, Dong K, Liu X, Ma J, Yuan E, Zhang L, et al. Enhancing gestational diabetes mellitus risk assessment and treatment through GDMPredictor: a machine learning approach. J Endocrinol Invest. 2024;47(9):2351–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Quotah OF, Andreeva D, Nowak KG, Dalrymple KV, Almubarak A, Patel A, et al. Interventions in preconception and pregnant women at risk of gestational diabetes; a systematic review and meta-analysis of randomised controlled trials. Diabetol Metab Syndr. 2024;16(1):8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Zhang X, Zhao X, Huo L, Yuan N, Sun J, Du J, et al. Risk prediction model of gestational diabetes mellitus based on nomogram in a Chinese population cohort study. Sci Rep. 2020;10(1):21223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T, Yavuz AA. The early prediction of gestational diabetes mellitus by machine learning models. BMC Pregnancy Childbirth. 2024;24(1):574. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Hah H, Goldin DS. How clinicians perceive artificial intelligence-assisted technologies in diagnostic decision making: mixed methods approach. J Med Internet Res. 2021;23(12):e33540. [DOI] [PMC free article] [PubMed]
- 14.Kokori E, Olatunji G, Aderinto N, Muogbo I, Ogieuhi IJ, Isarinade D, et al. The role of machine learning algorithms in detection of gestational diabetes; a narrative review of current evidence. Clin Diabetes Endocrinol. 2024;10(1):18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Bigdeli SK, Ghazisaedi M, Ayyoubzadeh SM, Hantoushzadeh S, Ahmadi M. Predicting gestational diabetes mellitus in the first trimester using machine learning algorithms: a cross-sectional study at a hospital fertility health center in Iran. BMC Med Inform Decis Mak. 2025;25(1):3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Kurt B, Gürlek B, Keskin S, Özdemir S, Karadeniz Ö, Kırkbir İB, et al. Prediction of gestational diabetes using deep learning and Bayesian optimization and traditional machine learning techniques. Med Biol Eng Comput. 2023;61(7):1649–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Li YX, Liu YC, Wang M, Huang YL. Prediction of gestational diabetes mellitus at the first trimester: machine-learning algorithms. Arch Gynecol Obstet. 2024;309(6):2557–66. [DOI] [PubMed] [Google Scholar]
- 18.Zhou F, Ran X, Song F, Wu Q, Jia Y, Liang Y, et al. A stepwise prediction and interpretation of gestational diabetes mellitus: foster the practical application of machine learning in clinical decision. Heliyon. 2024;10(12):e32709. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Watanabe M, Eguchi A, Sakurai K, Yamamoto M, Mori C. Prediction of gestational diabetes mellitus using machine learning from birth cohort data of the Japan Environment and Children’s Study. Sci Rep. 2023;13(1):17419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Cubillos G, Monckeberg M, Plaza A, Morgan M, Estevez PA, Choolani M, et al. Development of machine learning models to predict gestational diabetes risk in the first half of pregnancy. BMC Pregnancy Childbirth. 2023;23(1):469. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Gallardo-Rincón H, Ríos-Blancas MJ, Ortega-Montiel J, Montoya A, Martinez-Juarez LA, Lomelín-Gascón J, et al. MIDO GDM: an innovative artificial intelligence-based prediction model for the development of gestational diabetes in Mexican women. Sci Rep. 2023;13(1):6992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wang Y, Sun P, Zhao Z, Yan Y, Yue W, Yang K, Liu R, Huang H, Wang Y, Chen Y et al. Identify gestational diabetes mellitus by deep learning model from cell-free DNA at the early gestation stage. Brief Bioinform. 2023;25(1):bbad492. [DOI] [PMC free article] [PubMed]
- 23.Wu YT, Zhang CJ, Mol BW, Kawai A, Li C, Chen L, et al. Early prediction of gestational diabetes mellitus in the Chinese population via advanced machine learning. J Clin Endocrinol Metab. 2021;106(3):e1191–205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Spyropoulos CD. Ai planning and scheduling in the medical hospital environment. Artif Intell Med. 2000;20(2):101–11. [DOI] [PubMed] [Google Scholar]
- 25.Kufel J, Bargieł-Łączek K, Kocot S, Koźlik M, Bartnikowska W, Janik M, Czogalik Ł, Dudek P, Magiera M, Lis A et al. What is machine learning, artificial neural networks and deep learning?-Examples of practical applications in medicine. Diagnostics. 2023;13(15):2582. [DOI] [PMC free article] [PubMed]
- 26.Zhao M, Yao Z, Zhang Y, Ma L, Pang W, Ma S, et al. Predictive value of machine learning for the progression of gestational diabetes mellitus to type 2 diabetes: a systematic review and meta-analysis. BMC Med Inform Decis Mak. 2025;25(1):18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Theodorou B, Danek B, Tummala V, Kumar SP, Malin B, Sun J. Improving medical machine learning models with generative balancing for equity and excellence. NPJ Digit Med. 2025;8(1):100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Chen M, Xu W, Guo Y, Yan J. Predicting recurrent gestational diabetes mellitus using artificial intelligence models: a retrospective cohort study. Arch Gynecol Obstet. 2024;310(3):1621–30. [DOI] [PubMed] [Google Scholar]
- 29.Magrabi F, Ammenwerth E, McNair JB, De Keizer NF, Hyppönen H, Nykänen P, et al. Artificial intelligence in clinical decision support: challenges for evaluating AI and practical implications. Yearb Med Inform. 2019;28(1):128–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Moons KG, de Groot JA, Bouwmeester W, Vergouwe Y, Mallett S, Altman DG, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Debray TP, Damen JA, Snell KI, Ensor J, Hooft L, Reitsma JB, et al. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460. [DOI] [PubMed] [Google Scholar]
- 32.Metzger BE, Gabbe SG, Persson B, Buchanan TA, Catalano PA, Damm P, et al. International association of diabetes and pregnancy study groups recommendations on the diagnosis and classification of hyperglycemia in pregnancy. Diabetes Care. 2010;33(3):676–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Moons KGM, Wolff RF, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: explanation and elaboration. Ann Intern Med. 2019;170(1):W1-w33. [DOI] [PubMed] [Google Scholar]
- 34.Cabot JH, Ross EG. Evaluating prediction model performance. Surgery. 2023;174(3):723–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Huang C, Li SX, Caraballo C, Masoudi FA, Rumsfeld JS, Spertus JA, et al. Performance metrics for the comparative analysis of clinical risk prediction models employing machine learning. Circ Cardiovasc Qual Outcomes. 2021;14(10):e007526. [DOI] [PubMed] [Google Scholar]
- 36.Obuchowski NA, Bullen JA. Receiver operating characteristic (ROC) curves: review of methods with applications in diagnostic medicine. Phys Med Biol. 2018;63(7):07TR01. [DOI] [PubMed]
- 37.Stroup DF, Berlin JA, Morton SC, Olkin I, Williamson GD, Rennie D, Moher D, Becker BJ, Sipe TA, Thacker SB. Meta-analysis of observational studies in epidemiology: a proposal for reporting. Meta-analysis Of Observational Studies in Epidemiology (MOOSE) group. Jama. 2000;283(15):2008–2012. [DOI] [PubMed]
- 38.Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Iorio A, Spencer FA, Falavigna M, Alba C, Lang E, Burnand B, et al. Use of GRADE for assessment of evidence about prognosis: rating confidence in estimates of event rates in broad categories of patients. BMJ. 2015;350:h870. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Additional file 1: PRISMA 2020 Checklist.
Additional file 2: Search strategy used for the electronic databases.
Data Availability Statement
The datasets during and/or analyzed during the current study are available from the corresponding author on reasonable request.


