Skip to main content
PLOS One logoLink to PLOS One
. 2026 Sep 25;21(9):e0358557. doi: 10.1371/journal.pone.0358557

Artificial intelligence for women’s reproductive health: A scoping review of global diagnostic trends, methodological gaps, and a translational research agenda for low-resource settings

Mohammad Mehedi Hasan Munna 1,*,#, Toriqul Islam 1,#, Omar Faruk 2, Reefka Fabliha Tulona 3, Afsin Sultana 2,‡, Fairuj Saima 2,‡, Acramul Haque Kabir 4, Tania Sultana 1, Raihan Ul Islam 1
Editor: Kwang-Sig Lee5
PMCID: PMC13614794  PMID: 42789829

Abstract

Women’s reproductive and endocrine disorders including Polycystic Ovary Syndrome (PCOS), endometriosis, thyroid disorders, infertility, and pregnancy-related complications remain a major global health burden. These conditions are especially difficult to manage in low- and middle-income countries (LMICs), where diagnostic facilities and specialist care are limited. We conducted a PRISMA-ScR–guided scoping review of artificial intelligence (AI) and machine learning (ML) applications in women’s reproductive and hormonal health. Our search covered five databases PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar and included peer-reviewed studies published between 2021 and 2025. A total of 116 eligible studies were identified and classified into six categories: clinical prediction and risk modeling, medical imaging and AI, biomarker discovery and multi-omics, reproductive and endocrine health applications, clinical decision support systems, and AI/ML methodology development. Ensemble methods such as Random Forest and Gradient Boosting showed strong and consistent diagnostic performance. Convolutional neural networks performed well in single-centre settings but showed reduced performance upon external validation, consistent with optimism bias. Overall, 83.6% of studies were classified as high risk of bias and 75.5% lacked external validation. After adjusting for population size, high-income countries produced 23 times more studies per million women of reproductive age than LMICs. Sub-Saharan Africa contributed fewer than 0.1 studies per million women, and no studies were identified from Nepal or Sri Lanka. Current evidence, which is heavily concentrated in high-income and East Asian settings, is insufficient to support population-level deployment. We introduce a four-tier LMIC feasibility classification system that links data modality, infrastructure requirements, and personnel needs. We also propose a phased research agenda that separates near-term priorities (0–2 years) from medium-term goals that depend on infrastructure investment.

1. Introduction

Women’s reproductive and hormonal disorders constitute a significant yet persistently underaddressed global health challenge, with profound implications for fertility, metabolic function, mental health, and quality of life [1–3]. Reproductive and hormonal disorders are complex conditions resulting from a mix of genetic, environmental and lifestyle factors [1,4,5]. Although these are very common, they are highly under-diagnosed and under-treated, resulting in considerable health burden, sub-fertility, disability-adjusted life years (DALYs), and negative health outcomes in later life, such as cardiovascular and metabolic diseases and psychological disturbances [2,5]. PCOS affects 5–18% of women of reproductive age, and is closely associated with a number of other important comorbidities, such as insulin resistance, obesity, dyslipidaemia and depression [1,4,6]. Endometriosis affects 5–10% of women, with diagnostic delays often exceeding 8–12 years [7,8]. Women’s health is a systemic domain where women’s symptoms are often normalised or dismissed [7–9], and hence inadequate care is a common issue affecting these disorders.

A huge burden of reproductive and hormonal disorders also exists in low- and middle-income countries (LMICs) in South Asia, Sub-Saharan Africa, and South East Asia, where there is a coexistence of socio-economic burden, absence of specialist healthcare services, limited diagnostic capacity, and societal stigma [10]. In Bangladesh, the Maternal Mortality Ratio (MMR) has been reported to decrease from 323 deaths/100,000 live births in 2001–116 deaths/100,000 live births in 2017, but non-acute reproductive and hormonal disorders are not optimally managed. The prevalence of PCOS in Bangladesh has been estimated to be 7.8–20% as per the Rotterdam criteria, and diagnostic delays of over two years have been reported in rural Bangladesh [10,11]. Rapid expansion in mobile phone ownership and internet connectivity across LMICs provides a viable foundation for AI-enabled health solutions, yet pilot studies including work on early gestational diabetes prediction using wearable devices and explainable AI in South Africa highlight the gap between technological readiness and clinical integration [10,11].

Artificial intelligence (AI) and machine learning (ML) have shown strong potential in healthcare. They are particularly useful for pattern recognition, risk prediction, and clinical decision support. These methods have been applied to many types of data, including clinical records, hormonal profiles, medical imaging, and lifestyle indicators. Studies have reported high diagnostic performance for conditions such as PCOS, endometriosis, and gestational diabetes [12–16]. However, most existing models have important limitations. They were developed in high-income countries using small or single-centre datasets [12,14]. Many lack external validation or interpretability assessment. This raises serious concerns about whether these models can work reliably in low- and middle-income country (LMIC) settings [10]. No existing review has systematically examined methodological trends and translational readiness from an LMIC perspective. Consequently, the findings of this review reflect trends predominantly from East Asian and high-income settings and should not be interpreted as globally generalizable trends. Translational recommendations for LMICs remain theoretical until validated in local, resource-constrained populations. Some recent reviews have addressed related topics. Mengistu et al. [15] conducted a broad scoping review of AI in sexual and reproductive health. Ghaderzadeh et al. [12] systematically reviewed AI applications specifically in PCOS. However, neither review examined methodological rigor, resource-stratified deployment feasibility, and geographic equity together. These three factors are essential for assessing translational readiness in LMIC settings. The present review addresses this gap through three key contributions. First, it proposes a six-category methodological taxonomy covering all major reproductive and endocrine conditions. Second, it introduces a four-tier LMIC feasibility classification system that links data modality, infrastructure, and personnel requirements. Third, it presents a population-adjusted geographic equity analysis. This analysis reveals significant research voids in Nepal, Sri Lanka, and most of Sub-Saharan Africa [10,11]. This paper presents a PRISMA-ScR–guided scoping review of AI and ML applications in women’s reproductive and hormonal health (Fig 1). It covers studies published between 2021 and 2025. The review focuses explicitly on translational applicability in low-resource settings. Its goal is to support the development of robust, equitable, and context-aware AI solutions for LMICs [10,12].

Fig 1. PRISMA flow diagram.

Fig 1

Illustrating the study identification, screening, eligibility assessment, and inclusion process.

2. Materials and methods

This scoping review followed the PRISMA-ScR guidelines (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) [17]. PRISMA-ScR includes 20 essential items and 2 optional items. The completed checklist is provided as S1 File. These guidelines were chosen to make the review transparent, reproducible, and complete. This was especially important because the review aimed to explore and map existing literature. A predefined protocol guided the entire review process. It covered all major stages: literature search, screening, eligibility assessment, data extraction, study classification, and synthesis. Since this is a scoping review rather than an intervention-based systematic review, formal registration in PROSPERO was not required. However, the protocol was developed before data extraction began and was deposited in the Mendeley Data repository [18]. The repository includes screening metadata for all 273 identified records. It contains the reasons for exclusion at both the first and second screening stages. It also includes final inclusion decisions for the 116 eligible studies and their taxonomy classifications. This was done in accordance with the PRISMA 2020 data sharing guidelines.

2.1 Research questions

The research questions were formulated using the PICO (Population, Intervention, Comparison, Outcome) framework [19]. The population comprised women affected by reproductive and hormonal disorders; the intervention included AI and ML–based diagnostic or predictive approaches; comparisons involved conventional or non–AI-based strategies where applicable; and outcomes focused on diagnostic accuracy, risk prediction, clinical decision support, and translational feasibility in low-resource healthcare settings. The review was guided by the following research questions:

  • Q1: What are the clinical applications of AI and ML in women’s reproductive and hormonal health?

  • Q2: Which AI/ML tasks and models are most frequently used across different reproductive and endocrine disorders?

  • Q3: What data modalities are employed to develop AI-based diagnostic and predictive systems?

  • Q4: How are existing studies distributed across disease-specific and methodological taxonomies, and what translational gaps exist for low-resource healthcare settings?

2.2 Search strategy

A comprehensive literature search was conducted across five major academic databases: PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar. All searches were performed between December 2025 and February 2026, and results were date-stamped to ensure reproducibility. The search strategy was designed to identify peer-reviewed studies applying artificial intelligence, machine learning, or deep learning techniques to women’s reproductive and hormonal health. Search terms were organized into two conceptual domains and combined using Boolean operators (AND/OR). The complete database-specific search queries are summarized in Table 1. The 2021 start date was selected because it marks the inflection point where transformer-based architectures and the post-COVID acceleration of digital health research substantially reshaped the AI-in-healthcare literature, rendering pre-2021 methodological surveys insufficient for characterizing the current state of the field. Earlier work is well covered by previous reviews [12,15,16]. To ensure full reproducibility, we explicitly restricted the search to English-language publications. Grey literature, including preprints, unpublished theses, and conference abstracts without full peer-reviewed papers, were explicitly excluded to ensure the inclusion of only methodologically complete, validated studies. The complete, database-specific search strings, including all Boolean operators and field tags, are provided in Table 1.

Table 1. Database-specific search parameters.

Database AI/ML Terms Health Domain Terms Field Tags Date Range
PubMed “artificial intelligence”, “machine learning”, “deep learning” “women’s health”, “PCOS”, “endometriosis”, “thyroid”, “reproductive disorder”, “hormonal disorder”, “menstrual disorder” [TIAB] 2021–2025
Scopus Same as above Same as above TITLE-ABS-KEY 2021–2025
Web of Science Same as above Same as above TS= 2021–2025
IEEE Xplore Same as above Same as above All Metadata 2021–2025
Google Scholar Same as above “women’s reproductive health”, “PCOS”, “endometriosis”, “infertility”, “menstrual disorder”, “thyroid” Full-text 2021–2025

All term groups combined using Boolean OR within each concept block and AND between blocks. Searches restricted to English-language peer-reviewed articles and conference proceedings. Google Scholar: first 200 results screened per query. All searches were conducted between December 2025 and February 2026.

2.3 Eligibility criteria

Studies were considered eligible if they satisfied all of the following criteria:

  • Original peer-reviewed research articles;

  • Application of AI, ML, or deep learning techniques;

  • Focus on women’s reproductive or hormonal disorders;

  • Use of human clinical, imaging, biochemical, lifestyle, or multimodal data;

  • Publication in English between January 1, 2021 and December 31, 2025.

Studies were excluded if they met any of the following conditions: review articles, meta-analyses, editorials, commentaries, or book chapters; conference abstracts without full papers; animal or laboratory-based studies; studies unrelated to reproductive or endocrine health; absence of an AI/ML methodological component; or inaccessible full text or insufficient methodological detail.

2.4 Study selection process

All retrieved records were exported into Microsoft Excel, and duplicate entries were removed using automated deduplication followed by manual verification. Study selection was conducted using a predefined two-stage screening process managed through Rayyan AI [20,21] screening software. In the first stage, titles and abstracts were independently screened by two reviewers (MMHM and TI) to assess relevance to the research questions. Studies deemed irrelevant by both reviewers were excluded, while discrepancies were resolved through discussion with a third reviewer (TS). In the second stage, full-text articles of potentially eligible studies were independently assessed against the predefined eligibility criteria by two reviewers (AFS and FS). Reasons for exclusion at this stage were documented using a standardized classification system. Initial database searches identified 402 records. After removing 129 duplicates, 273 records underwent title/abstract screening. Of these, 167 proceeded to full-text assessment after excluding 73 records at title screening and 23 at abstract screening. Full-text review of 167 articles resulted in a further 51 exclusions, yielding 116 studies that satisfied all inclusion criteria for final synthesis (Fig 1). Of the 116 final studies, 94 (81.0%) were retrieved from PubMed, Scopus, IEEE Xplore, or Web of Science, and 22 (19.0%) were unique to Google Scholar. The Google Scholar 200-result limit therefore affected approximately one-fifth of the final corpus, and sensitivity to this cap was assessed by cross-checking citation overlap across databases.

2.5 Data extraction and charting

Data extraction was performed using a standardized Excel-based data-charting form, which was iteratively refined during a pilot phase involving 10 randomly selected studies. For each included study, data were extracted across six specific domains:

  1. Bibliographic information (database source, title, authors, publication year, journal/conference);

  2. Study characteristics (research objective, target disease/condition, study design, sample size, geographic region, funding source);

  3. Data modalities (clinical records, hormonal/biochemical markers, medical imaging, lifestyle/demographic data, multimodal integration);

  4. AI/ML components (tasks, algorithms employed, software/implementation details);

  5. Performance metrics (accuracy, sensitivity, specificity, AUC, F1-score, validation methods); and

  6. Translational indicators (data accessibility, computational requirements, cost considerations, ethical considerations, deployment feasibility in LMICs).

A complete data extraction dictionary detailing the operational definitions for all extracted variables is provided in the supplementary repository.

To ensure high fidelity and minimize extraction errors, the data underwent a rigorous, multi-tier independent verification process. Initial data extraction was conducted by one reviewer (TI). In the second stage, the extracted data were independently verified by three reviewers (OF, AFS, and FS). Subsequently, a third verification was conducted by MMHM to ensure comprehensive accuracy across the entire dataset.

A 20% random sample of the fully verified data was cross-verified, with inter-rater agreement calculated using Cohen’s κ (κ=0.92). Any remaining discrepancies at any stage of the verification process were resolved through discussion and consensus with the senior authors (RUI and TS).

2.6 Study classification and taxonomy

The classification taxonomy was developed through an iterative three-step process: (1) initial framework based on prior AI-in-healthcare reviews [12,15], (2) pilot testing and refinement using 20 randomly selected studies, and (3) finalization through team consensus meetings involving all authors. Each study was assigned to a single primary taxonomy category based on its principal research objective, data modalities, and AI/ML tasks. In cases where studies addressed multiple domains, classification was based on the primary contribution of the work. The final taxonomy comprised six categories: (1) Clinical Prediction and Risk Modeling; (2) Medical Imaging and Artificial Intelligence; (3) Biomarker Discovery and Multi-Omics Analysis; (4) Reproductive and Endocrine Health Applications; (5) Clinical Decision Support and Digital Health Tools; and (6) AI/ML Methodology and Algorithmic Development. To resolve boundary cases between Category 2 and Category 4, the following decision rule was applied prospectively: Category 2 was assigned when the primary methodological contribution was an imaging architecture or segmentation method; Category 4 was assigned when the primary contribution was clinical integration of AI with reproductive health data, regardless of whether imaging was involved. Studies combining ultrasound features with clinical parameters without proposing a novel imaging methodology were classified as Category 4.

2.7 Data synthesis and reporting

To ensure transparency in all descriptive summaries, the total number of included studies (N = 116) serves as the primary denominator for all reported counts, proportions, and percentages throughout the results, unless a specific subgroup denominator (e.g., studies specific to a single disease domain or algorithm) is explicitly stated in the corresponding table or text.

Given the heterogeneity in datasets, disease domains, and evaluation metrics across included studies, quantitative meta-analysis was not performed. Three structural violations precluded valid pooling: first, substantial clinical heterogeneity across disease domains (PCOS, endometriosis, thyroid, GDM) means that a pooled AUC estimate would constitute an ecological fallacy; second, 31 of 116 studies report multiple algorithms on identical datasets, violating the independence assumption required for meta-analytic pooling; third, overlapping training cohorts from shared public repositories (Kaggle PCOS, UCI thyroid) further compound statistical dependence. Formal heterogeneity testing (I²) was therefore not performed, as the independence violations render such statistics uninterpretable. This decision is consistent with established scoping review methodology, which prioritises breadth of evidence mapping over quantitative synthesis. All performance summaries are presented as descriptive ranges to characterise the reviewed literature, in alignment with scoping review methodology. Although descriptive AUC ranges are reported to characterise central tendency within algorithm categories, several structural threats to validity must be acknowledged: (1) ecological fallacy, as study-level AUC estimates aggregate outcomes across different diseases, populations, and imaging modalities; (2) statistical dependence, as 31 of 116 studies report multiple algorithms on identical datasets, violating independence assumptions; and (3) overlapping training cohorts from shared public repositories (e.g., UCI thyroid, Kaggle PCOS datasets). Given these violations, no inferential statistical testing (p-values, t-tests, or significance claims) is reported, and all quantitative summaries in Tables 4 and 5 are presented as descriptive approximations; no causal or superiority claim should be inferred.

Extracted data were summarised descriptively and thematically in relation to the four research questions: (1) clinical applications by disease domain, (2) AI/ML task and model distribution, (3) data modality trends, and (4) taxonomy-based patterns with translational implications. Descriptive AUC ranges were compiled from author-reported metrics for each algorithm category.

2.8 Risk-of-bias and quality assessment

The included corpus spans four methodologically distinct study types — diagnostic imaging, clinical prediction modeling, biomarker discovery, and algorithm development — each governed by a different validated risk-of-bias instrument (QUADAS-AI for imaging, PROBAST-AI for prediction, and TRIPOD-AI reporting standards for algorithm development). Item-by-item application of any single instrument across this heterogeneous corpus would have produced systematic misclassification, because items relevant to one study type are inapplicable to another. To preserve cross-corpus comparability while retaining domain-specific validity, we adopted a harmonized five-domain framework synthesizing the shared core items of PROBAST-AI, QUADAS-AI, and TRIPOD-AI. Each study was evaluated across five domains: (1) participant selection and data representativeness; (2) predictor measurement; (3) outcome assessment and ground truth reliability; (4) model development and validation strategy; and (5) reporting transparency. Studies were categorized as having low, moderate, or high risk of bias based on cumulative assessment across domains. The complete scoring rubric, including explicit criteria for low- and high-risk classification in each domain, is provided in S1 Table. Studies meeting high-risk criteria in 1–2 domains were classified as moderate risk; studies meeting high-risk criteria in 3 or more domains were classified as high risk. The number of studies at each risk level per domain is reported in Table 10.

Table 10. Risk of bias assessment using adapted PROBAST-AI framework (N = 116).

Domain Low Risk n (%) High Risk n (%) Most Common Deficiencies
Participant Selection 37 (31.9%) 79 (68.1%) Single-center recruitment (78%); convenience sampling (64%); inadequate sample size justification (89%)
Predictor Measurement 51 (43.7%) 65 (56.3%) Retrospective data extraction without blinding (72%); inconsistent feature engineering across validation sets (41%)
Outcome Assessment 73 (62.9%) 43 (37.1%) Reference standard not blinded to AI prediction (58%); surgical confirmation unavailable for negative cases (endometriosis studies)
Model Development 40 (34.5%) 76 (65.5%) No hyperparameter tuning documentation (67%); class imbalance not addressed (53%); feature selection leakage (31%)
Model Validation 26 (22.7%) 90 (77.3%) No external validation (75.5%); temporal validation lacking (89%); inadequate sample size for validation set (<20% of total) (44%)

2.9 LMIC applicability assessment

Each included study was evaluated for deployment feasibility in LMIC settings using a four-tier qualitative classification scheme, defined prospectively based on infrastructure requirements, personnel needs, cost per test, and appropriate care level. The four categories are: (1) High Feasibility — requires only basic clinical variables, no specialist personnel, cost < $10/test, deployable in primary care; (2) Moderate Feasibility — requires routine laboratory tests or basic ultrasound, nurse- or midwife-level personnel, cost $10–25/test, deployable in secondary care; (3) Low Feasibility — requires specialist imaging or advanced laboratory panels, specialist personnel, cost $25–50/test, requires district hospital; and (4) Very Low Feasibility — requires MRI or advanced imaging or multi-omics, highly specialized personnel, cost > $50/test, requires tertiary centre only. The full categorical criteria are provided in S2 Table. This qualitative scheme was adopted in preference to a continuous numeric score to avoid conveying unwarranted precision and to ensure reproducibility of classification across reviewers.

2.10 Limitations of the methodology

  • Language restriction: Inclusion of only English-language publications may exclude relevant studies from non-English speaking LMICs.

  • Database coverage: While five major databases were searched, some relevant studies in regional or specialized databases may have been missed.

  • Grey literature exclusion: Preprints, theses, and conference proceedings without full papers were excluded, potentially missing recent innovations.

  • Rapid field evolution: The AI/ML field evolves quickly, and some advances from late 2025 may not be fully captured.

  • Scoping nature: As a scoping review, detailed quality assessment of individual studies was not the primary focus, though general methodological trends were assessed.

2.11 Sensitivity analysis for overlapping public datasets

Several included studies draw on identical publicly available repositories, including the Kaggle PCOS dataset and the UCI thyroid dataset. To assess sensitivity of descriptive AUC ranges to dataset duplication, a collapsed dataset-level analysis was performed. Studies sharing a common training source were treated as a single observation for descriptive pooling purposes. After collapsing, effective independent dataset counts were:

  • PCOS (Kaggle-derived): 20 studies collapsed to 11 independent datasets;

  • Thyroid (UCI-derived): 24 studies collapsed to 14 independent datasets.

Collapsed descriptive AUC ranges shifted by less than 0.02 units across all algorithm families, suggesting descriptive patterns are not materially driven by dataset duplication. However, algorithm frequency rankings (Table 3) should be interpreted with caution, as Random Forest’s apparent dominance in PCOS studies partly reflects repeated application to the same Kaggle benchmark rather than independent clinical validation.

Table 3. Distribution of studies across six taxonomy categories with validation quality assessment.

Taxonomy Category n (%) Primary Algorithms Mean Sample Ext. Val. Multi-center Risk of Bias
Clinical Prediction & Risk Modeling 41 (35.3%) RF (65.7%), XGBoost (62.8%) 7,700.5 6 (16.7%) 15 (41.7%) Moderate-High
Medical Imaging + AI 22 (18.9%) CNN variants (86%) 920 2 (9.0%) 5 (22.7%) High
Biomarker Discovery (Multi-Omics) 13 (11.2%) LASSO (77%), SVM-RFE (69%) 282 5 (38.5%) 7 (53.8%) Moderate
Reproductive & Endocrine Health 19 (16.4%) Ensemble methods (73.7%) 2,259 2 (10.5%) 10 (52.6%) Moderate
Clinical Decision Support & Digital Health 10 (8.6%) RF (60%), Logistic Regression (60%) 1,072 0 (0.0%) 3 (30.0%) Moderate-High
AI/ML Methodology Development 11 (9.5%) Hybrid architectures (92%) 3,004 0 (0.0%) 1 (8.3%) High

Risk of bias assessed using adapted PROBAST-AI framework; “High” indicates single-center design + no external validation + small sample size (<300). Ext. Val. = External Validation.

2.12 Use of artificial intelligence tools

During the preparation of this manuscript, the authors used Grammarly and Quillbot for language editing and manuscript refinement. The authors reviewed and edited the AI-generated content and take full responsibility for the accuracy, validity, and originality of the final text.

2.13 Ethics statement

Ethical approval was not required for this study, as it is a scoping review of published literature and involved no human participants or primary data collection.

3. Results

3.1 Overview of included studies

Data from 116 studies were synthesised. An increase in the number of publications was observed from 2021 to 2022, and a larger surge was seen from 2023 onwards. Studies were reported from 27 countries across all regions (Table 2). Although the number of studies in LMICs was marginally higher in absolute terms (51.7%), when adjusted for the number of women of reproductive age per region, the number of studies per million women in HICs was 23 times that in LMICs, reflecting disparities in institutional and funding capacity [13,22–136].

Table 2. Structural characteristics of included studies (n = 116).

Characteristic Category n (%) Key Observations Characteristic Category n (%)
Publication Year 2021 13 (11.2%) Early exploratory phase Study Design Retrospective 86 (74.1%)
2022 10 (8.6%) Imaging focus Prospective 15 (12.9%)
2023 24 (20.7%) Multimodal emergence Mixed Methods 15 (12.9%)
2024 32 (27.6%) XAI growth Sample Size <100 11 (9.5%)
2025 37 (31.9%) Clinical implementation 100–500 33 (28.4%)
Origin High-Income 56 (48.3%) US, S. Korea, EU 501–1,000 13 (11.2%)
LMICs 60 (51.7%) China, India, Bangladesh >1,000 59 (50.9%)

HIC = High-Income Countries; LMIC = Low- and Middle-Income Countries; XAI = Explainable Artificial Intelligence. Population-adjusted study density in HICs was 23 times that of LMICs.

3.2 Findings by taxonomy category

The distribution of included studies across the six taxonomy categories is reported in Table 3.

Clinical Prediction & Risk Modeling (41 studies) included work on PCOS, gestational diabetes mellitus (GDM), and preeclampsia, with Random Forest and XGBoost as dominant algorithms.

Medical Imaging + AI (22 studies) applied convolutional neural networks (CNNs) to endometriosis, uterine fibroid, and ovarian abnormality diagnosis, with accuracies in the range of 85–99%; almost all studies used single-centre datasets.

Biomarker Discovery (13 studies) applied multi-omics approaches to identify potential novel biomarkers for endometriosis and PCOS, using models such as PLS-DA and LASSO regression.

Reproductive & Endocrine Health (19 studies) predominantly addressed PCOS diagnosis, often combining ultrasound features with clinical parameters (accuracy range 89–100%).

Clinical Decision Support (10 studies) reported mobile health and digital triage applications, though very few included validation in LMIC populations.

AI/ML Methodology (11 studies) described novel hybrid CNN-ML approaches and interpretability approaches such as SHAP [137] (Fig 2).

Fig 2. Classification of women’s reproductive health disorders.

Fig 2

Illustrating major categories including menstrual dysfunction, pelvic structural disorders, androgen and metabolic abnormalities, endocrine gland dysfunction, fertility and pregnancy-related complications, and ovarian aging.

3.3 Disease-specific analysis

As summarized in Table 4, endometriosis (n = 27, 23.3%) and PCOS (n = 20, 17.2%) represented the most extensively researched domains. PCOS studies reported accuracy in the range of 88–99% and AUC in the range of 0.89–0.98, supported by the integration of clinical (52%) and ultrasound (31%) data. Thyroid disorder studies (n = 24, 20.7%) similarly reported strong discrimination (AUC range 0.89–0.99), reflecting standardized biomarker thresholds in hormonal data. Endometriosis studies reported the widest range of performance metrics (AUC range 0.72–0.94) and the lowest proportion of externally validated models (11.1%). Infertility/IVF studies (n = 4, 3.4%) reported the narrowest range of performance (AUC range 0.74–0.89) and none had external validation.

Table 4. Disease domain distribution, data modalities, and performance metrics.

Disease Domain n (%) Primary Data Modalities Top Algorithms Accuracy Range Ext. Val.
PCOS 20 (17.2%) Clinical (52%), US (31%), Hormonal (17%) RF (38%), XGBoost (28%) 0.89−0.98 10.0%
Endometriosis 27 (23.3%) Imaging (48%), Clinical (32%), Biomarkers (20%) CNN (52%), RF (26%) 0.72−0.94 11.1%
Thyroid Disorders 24 (20.7%) Hormonal (65%), Clinical (35%) RF (45%), SVM (25%) 0.89−0.99 12.5%
GDM 12 (10.3%) Clinical (60%), First-trimester labs (40%) XGBoost (50%), RF (32%) 0.81−0.94 16.7%
Preeclampsia 11 (9.5%) Clinical (45%), Doppler (30%), Biomarkers (25%) XGBoost (75%), RF (67%) 0.79−0.94 18.2%
Infertility/IVF 4 (3.4%) Embryo imaging (75%), Hormonal (25%) CNN (75%), RF (50%) 0.74−0.89 0%
Uterine Disorders 18 (15.5%) Ultrasound (62%), MRI (38%) CNN (79%), U-Net (36%) 0.84−0.97 11.1%

US = Ultrasound; RF = Random Forest; GDM = Gestational Diabetes Mellitus; Ext. Val. = External Validation. *Includes fibroids, adenomyosis, endometrial hyperplasia/cancer.

When performance is stratified by validation rigor (Table 5), a directional pattern of AUC reduction is observed as validation moves from single-centre to external testing. This pattern is directionally consistent with optimism bias across all algorithm categories and should not be interpreted as a formal statistical comparison, given the small matched subsets, heterogeneity in disease contexts, and violations of independence. Fig 3 illustrates the directional reduction in the reported AUC ranges for five algorithm families moving from single-centre to external validation settings. Calibration reporting was assessed as an additional aspect of study quality: calibration metrics (Brier score, calibration curve, or Hosmer–Lemeshow test) were reported in 8 of 116 studies (6.9%). Decision curve analysis or net benefit were reported in 3 studies (2.6%), all focused on GDM or preeclampsia. Clinical decision thresholds were clearly mentioned in the text of only 1 study (0.9%). A model may have high AUC but be poorly calibrated, resulting in a systematic bias toward over- or under-estimation of individual risk. Calibration metrics, in addition to discrimination metrics, should be included in minimum reporting standards.

Table 5. Descriptive AUC ranges reported across validation settings.

Algorithm Studies Single-Center (AUC) Multi-Center (AUC) Ext. Validation (AUC)
Random Forest 51 0.89−0.96 0.86−0.92 0.84−0.92
XGBoost 38 0.88−0.96 0.85−0.92 0.83−0.92
CNN Variants 33 0.91−0.96 0.87−0.93 0.85−0.92
SVM 28 0.84−0.93 0.82−0.90 0.80−0.90
Logistic Regression 25 0.79−0.89 0.77−0.88 0.76−0.87
Hybrid/Ensemble 22 0.93−0.97 0.89−0.94 0.87−0.93

AUC values are descriptive approximations computed from author-reported metrics within each algorithm category. All differences across validation tiers should be interpreted directionally as exploratory benchmarks only; no inferential comparison is intended. Total algorithm counts exceed 116 because individual studies may report multiple algorithms for comparative purposes; each algorithm-study pair is counted independently.

Fig 3. Descriptive AUC by validation stringency (exploratory).

Fig 3

Bar chart shows mean AUC values reported across single-center, multi-center, and external validation settings for five algorithm families. Differences across validation tiers are directionally consistent with optimism bias; formal significance testing was not performed given violations of independence and small matched subsets (minimum n = 5 per algorithm). Values should not be interpreted as comparative efficacy estimates.

3.3.1 Sensitivity analysis of performance patterns.

To assess robustness of the descriptive AUC decay patterns, three supplementary exploratory checks were performed.

  1. Leave-one-out analysis: the largest single study contributing to each algorithm category was sequentially excluded; for Random Forest (largest study n = 12,000), exclusion changed the descriptive single-centre AUC range by less than 0.01, suggesting results are not driven by outlier studies.

  2. Disease-stratified sub-analysis: decay patterns were re-examined separately for PCOS, endometriosis, and thyroid disorders. Thyroid models showed the smallest decay (mean ΔAUC≈0.01 across validation tiers), consistent with their well-standardised biomarker reference ranges. Endometriosis models showed the largest decay (ΔAUC≈0.09), consistent with heterogeneous imaging protocols and the absence of consensus reference standards.

  3. Modality stratification: imaging-based studies (ultrasound + MRI) showed larger decay than clinical-data-only models (ΔAUC≈0.08 vs. 0.04), consistent with known scanner-level and protocol-level domain shift in medical imaging.

These exploratory sensitivity checks reinforce the narrative interpretation that directional patterns are robust to outlier and stratification perturbations, and support prioritising clinical-data models for near-term LMIC deployment. These should not be interpreted as confirmatory analyses.

3.4 Data modality and resource requirements

Basic clinical variables were used in 98 (84.5%) of the studies, and routine laboratory tests in 65 (56.0%), indicating reliance on low-cost and readily available data. Tier 1 modalities were used in all the studies concerning PCOS, GDM, and Preeclampsia, supporting feasibility in primary and secondary care settings. Refer to Table 6.

Table 6. Data modality utilization patterns across disease domains.

Data Modality Overall Usage PCOS Endo. GDM PE Resource Tier LMIC Feasibility
Basic Clinical (Age, BMI) 98 (84.5%) 20 (100%) 22 (81.5%) 12 (100%) 11 (100%) Tier 1 (Low) Highest
Routine Labs (CBC, Glu) 65 (56.0%) 12 (60.0%) 8 (29.6%) 12 (100%) 11 (100%) Tier 1 (Low) High
Hormonal Panel 42 (36.2%) 18 (90.0%) 5 (18.5%) 6 (50.0%) 2 (18.2%) Tier 2 (Med) Moderate
Ultrasound Imaging 38 (32.8%) 16 (80.0%) 12 (44.4%) 2 (16.7%) 4 (36.4%) Tier 2 (Med) Moderate
MRI/Advanced Imaging 22 (19.0%) 0 (0%) 15 (55.6%) 1 (8.3%) 2 (18.2%) Tier 3 (High) Low
Multi-omics 18 (15.5%) 4 (20.0%) 10 (37.0%) 2 (16.7%) 1 (9.1%) Tier 3 (High) Very Low

Resource tiers defined by equipment requirements, specialist interpretation needs, and cost per test in LMIC contexts. PE = Preeclampsia; Endo. = Endometriosis.

Use of hormonal panels (36.2%) and ultrasound (32.8%) suggests a moderate level of resource dependency, implying a need for some level of laboratory and clinical skill. The low frequency of use of MRI/Advanced Imaging (19.0%) and multi-omics platforms (15.5%) suggests these modalities are primarily being used in tertiary care in HICs and in specialized centres, and are therefore not widely applicable in resource-constrained settings. As shown in Table 6, models requiring Tier 3 resources have the lowest LMIC feasibility. As shown in Table 7, studies that do not require specialist personnel (e.g., community hypertension screening, symptom-based endometriosis classifiers) fall into High Feasibility, while ultrasound-based fibroid and PCOS diagnostics fall into Low Feasibility.

Table 7. Resource-constraint analysis: clinical applicability scoring.

Study Disease Focus Required Infrastructure Specialist Dependency Cost per Test (USD) LMIC Score Setting
[22] Osteopenia BIA device Minimal (nurse) 15–25 Moderate Primary care
[25] Hypertension BP cuff, WC tape Minimal (nurse) <5 High Community health
[27] Endometriosis Clinical history only Moderate (GP) 0 High Primary care
[32] Uterine Fibroids Ultrasound High (radiologist) 25–40 Low District hospital
[45] Endometriosis Urine test + spectrometer Moderate (lab tech) 8–12 Moderate Secondary care
[66] Endometriosis Electrochemical sensor Minimal (nurse) 3–5 High Primary care
[75] GDM First-trimester labs Minimal (midwife) 12–18 Moderate Antenatal clinic
[119] PCOS Ultrasound High (radiologist) 25–40 Low Tertiary center
[36] PCOS Clinical + USG Moderate (GP + tech) 15–25 Moderate Secondary care

Resource tiers defined by equipment requirements, specialist interpretation needs, and cost per test in LMIC contexts. PE = Preeclampsia; Endo. = Endometriosis.

The resource-applicability patterns illustrated in Fig 4 highlight that low-resource models tend to fall within the primary-care-compatible quadrants, further emphasising the importance of simple infrastructure for feasibility and scalability.

Fig 4. Resource-applicability matrix for AI diagnostic tools.

Fig 4

The scatter plot visualizes the trade-off between infrastructure requirements and specialist dependency, with the LMIC feasibility threshold indicated by the dashed line.

3.5 Geographic and equity analysis

The geographic distribution showed notable discrepancies. As shown in Table 8, East Asia contributed the largest proportion of studies (39.6%), followed by North America (17.2%) and Europe (16.4%). More than one-third of the studies in each of these regions were multicentre (37–45%), with external validation rates ranging from 24% to 35%, and open-access rates reaching 60% in North America.

Table 8. Geographic distribution with methodological quality indicators.

Region n (%) Median Sample Multi-center Ext. Validation Open Code/Data Mean AUC Key Limitations
North America 20 (17.2%) 541 9 (45.0%) 7 (35.0%) 12 (60.0%) 0.87 Limited LMIC validation
East Asia 46 (39.6%) 1,114 18 (39.1%) 11 (23.9%) 8 (17.4%) 0.89 Single-country focus
South Asia 18 (15.5%) 309 4 (22.2%) 3 (16.7%) 2 (11.1%) 0.85 Small samples; urban bias
Europe 19 (16.4%) 419 7 (36.8%) 5 (26.3%) 10 (52.6%) 0.88 High-resource bias
Middle East 6 (5.2%) 694 1 (16.7%) 1 (16.7%) 0 (0%) 0.83 Limited generalizability
Latin America 5 (4.3%) 305 0 (0%) 0 (0%) 1 (20.0%) 0.81 Severe validation gap
Sub-Saharan Africa 2 (1.7%) 104 0 (0%) 0 (0%) 0 (0%) 0.79 Critical research gap

The lower validation rates observed in South Asia (15.5%) and Sub-Saharan Africa (1.7%) were accompanied by smaller median sample sizes (309 and 104, respectively) and limited data-sharing practices (only 11.1% of South Asian studies shared code or data). The country-level analysis (Table 9) further highlights under representation of South Asian countries: there were only 4 studies from Bangladesh (3.4%), all from a single tertiary care hospital and using only internal cross-validation, and none from Nepal or Sri Lanka. Taken together, the geographic distribution data in Table 8 and the inequality metrics illustrated in Fig 5 highlight the severe under representation of Sub-Saharan Africa and rural South Asia, the exclusion of marginalized populations, and a high risk of bias arising from the absence of diverse training data.

Table 9. LMIC representation analysis: Bangladesh and regional context.

Country n (%) Disease Focus Data Source Validation Approach Key Strengths Critical Gaps
Bangladesh 4 (3.4%) PCOS (2), Thyroid (2) Single tertiary centers Internal CV only Context-specific patterns No rural/external validation
India 18 (15.5%) PCOS (8), Thyroid (4), Other (6) Multi-center urban 2 with external validation Large diverse samples Urban bias
Pakistan 3 (2.6%) PCOS (2), Thyroid (1) Single centers Internal CV only Low-cost feature selection Small samples (<350)
Nepal 0 (0%) — — — — Complete research gap
Sri Lanka 0 (0%) — — — — Complete research gap
SAARC Avg.* 17 (14.3%) PCOS dominant (65%) Urban tertiary centers 11.8% external validation Growing capacity Severe rural gap

SAARC Avg.* = South Asian Association for Regional Cooperation Average, comprising Afghanistan, Bangladesh, Bhutan, India, Maldives, Nepal, Pakistan, and Sri Lanka. Percentages are calculated based on the total number of LMIC studies (n = 119).

Fig 5. Global research equity gap analysis.

Fig 5

(A) Geographic distribution of included studies across seven regions, highlighting the concentration of research in East Asia and high-income countries. (B) Key inequality metrics: 75.5% of studies lack external validation, 83.6% are classified as high risk of bias per PROBAST-AI criteria, and high-income countries outproduce LMICs by a 23:1 ratio after DALY-burden adjustment.

3.6 Methodological quality and bias assessment

Methodological appraisal using the adapted PROBAST-AI framework and the scoring rubric in S1 Table revealed pervasive limitations across multiple domains. As shown in Table 10, high risk of bias was observed in participant selection (68.1%), predictor measurement (56.3%), and model development (65.5%). These deficiencies were primarily driven by convenience sampling, retrospective feature extraction, inadequate documentation of preprocessing pipelines, and insufficient handling of class imbalance.

Validation represented the most critical weakness, with 77.3% of studies classified as high risk and 75.5% lacking any form of external testing. Temporal validation was absent in 89% of studies, substantially increasing susceptibility to optimism bias and overestimation of real-world performance.

The multidimensional quality patterns illustrated in Fig 6 further demonstrate systematic gaps in validation rigor, data transparency, and LMIC applicability. Overall, 83.6% of studies were classified as high risk of bias, substantially restricting the reliability of current evidence for clinical and policy-level decision-making. Explainability and clinical integration features summarized in Table 11 indicate limited adoption of transparency mechanisms. SHAP or LIME was implemented in only 6.9% of studies, while the majority (26.7%) employed no interpretability tools. Although attention-based visualizations were used in 4.3% of imaging studies, their clinical actionability remained limited.

Fig 6. Methodological quality disparities between HIC and LMIC studies.

Fig 6

Dumbbell chart comparing mean quality scores (1–5 scale) across six dimensions. Blue dots represent HIC-originating studies; red dots represent LMIC-originating studies. Delta values indicate the HIC–LMIC score difference. Validation rigor (highlighted) shows the largest gap (+2.9), driven by differences in external validation rates (35.0% vs 16.7%) and multi-center participation (45.0% vs 22.2%).

Table 11. Explainability and clinical integration features.

Feature n (%) Domains Utility
SHAP/LIME 8 (6.9%) PCOS, GDM, PE High
Attention/Grad-CAM 5 (4.3%) Imaging Moderate
Decision Thresholds 1 (0.9%) GDM, PE High
Web/Mobile 2 (1.7%) PCOS, thyroid Moderate
EHR Integration 1 (0.9%) US, China High
No Explainability 31 (26.7%) Imaging, biomarker Low

PE = Preeclampsia; GDM = Gestational Diabetes Mellitus; EHR = Electronic Health Record.

3.7 Temporal trends and innovation trajectory

The temporal analysis reported in Fig 7 shows that AI research methodologies have evolved substantially over 2021–2025. Multimodal integration increased from 15.4% in 2021 to 51.4% in 2025. Explainable AI features evolved from 0% in 2021 to 34.3% in 2025. The average sample size of datasets used in the studies increased from approximately 920 in 2021–4,500 in 2025. Open science practices also increased markedly, from 23.1% in 2021 to 65.7% in 2025. Based on the timeline in Fig 7, after 2023, multimodal modeling, validation frameworks, and transparency mechanisms have been adopted more rapidly. However, as of 2025, the percentage of AI models with some form of external validation is only 37.1%, and prospective studies remain below 20%.

Fig 7. Innovation trajectory timeline.

Fig 7

Gantt-style timeline showing emergence and adoption rates of key methodological innovations: multimodal integration, external validation, XAI, prospective design, and open science practices. Color intensity represents adoption percentage per year.

4. Discussion

4.1 Synthesis of key findings

AI and ML applications in women’s reproductive health have advanced substantially from 2021 to 2025, yet a critical translational gap persists between algorithmic performance and real-world applicability, particularly in LMICs. Three cross-cutting patterns define the field. First, ensemble methods — Random Forest and XGBoost — dominate clinical prediction tasks, reporting consistently strong performance across studies and outperforming imaging-based approaches in resource efficiency. Second, CNN-based imaging models show a directional pattern of performance reduction moving from single-centre evaluations to external validation settings, consistent with the well-documented literature on optimism bias in internally validated clinical AI. This pattern is observed across all five major algorithm families and replicated in disease-stratified sensitivity analyses. We emphasise, however, that the magnitude of this reduction carries considerable uncertainty: estimates are derived from small matched subsets (n = 5–51 per algorithm), aggregate heterogeneous disease contexts, and assume independence that is partially violated by overlapping training cohorts. The observed pattern of lower reported performance in externally validated versus internally validated studies should therefore be interpreted as descriptive evidence consistent with established literature on optimism bias in clinical AI, rather than a precise measurement of any single model family’s generalisation gap. Third, despite multimodal integration rising from 15.4% in 2021 to 51.4% in 2025 and XAI adoption growing from 0% to 34.3%, external validation remained below 37.1% even in 2025, indicating that methodological sophistication has outpaced clinical validation infrastructure. Table 12 summarises domain-level findings alongside LMIC-specific recommendations.

Table 12. Summary of domain-level findings and LMIC recommendations.

Domain Key Finding LMIC Recommendation
Clinical Prediction (PCOS, GDM, Thyroid) Ensemble methods (RF, XGBoost) dominant; AUC 0.87–0.95 on Tier 1 data. Prioritise for primary care deployment; low-cost inputs feasible.
Medical Imaging (Endometriosis, Fibroids) CNN AUC 0.91 (single-centre) drops to 0.85 externally; Tier 3 infrastructure. Not viable in primary care; invest in non-invasive biosensor alternatives.
Explainability & Validation SHAP/LIME in 6.9% only; 75.5% lack external validation; 83.6% high bias risk. Mandate XAI and multi-centre validation before deployment.
Geographic Equity 23:1 HIC-to-LMIC ratio; East Asia 39.6%, Africa 1.7%; Bangladesh 4 studies only. Fund LMIC-led validation consortia; build rural data infrastructure.
Decision Support & Digital Health Promising for LMICs but 0% external validation; no LMIC population testing. Pilot mobile-health triage tools with community health workers.

4.2 Disease-specific interpretation

Predictive performance across disease domains is shaped by biological complexity, diagnostic standardisation, and data modality requirements. Thyroid disorder studies reported the highest and most stable performance range, underpinned by well-validated biomarker thresholds (TSH, T3, T4) and Tier 1–2 resource inputs, making them the strongest candidates for near-term LMIC deployment. PCOS studies similarly reported strong performance, yet 80% relied on ultrasound — a Tier 2 resource classified as Low Feasibility in our LMIC scheme — limiting scalability beyond district hospital settings. The heterogeneity introduced by competing diagnostic criteria (Rotterdam, NIH, AES) further complicates cross-study comparisons and model replication in populations with different phenotypic distributions. Endometriosis presented the most challenging diagnostic landscape, with the widest range of reported performance and only 11.1% external validation, reflecting dependence on laparoscopic confirmation as the reference standard. Emerging non-invasive approaches — including electrochemical biosensors, urine infrared spectrometry, and symptom-based ML classifiers classified as High Feasibility — offer more viable LMIC pathways and directly address the 8–12 year diagnostic delay that causes preventable morbidity. GDM and preeclampsia studies demonstrated the most credible external validation evidence (16.7–18.2%), with first-trimester clinical and laboratory inputs costing $12–18 per assessment supporting feasibility for antenatal clinic deployment by midwifery-level personnel. Infertility and IVF studies (0% external validation) remain the least translatable domain given their dependence on specialised embryo imaging technology.

4.3 Methodological quality, systemic bias, and explainability

The adapted PROBAST-AI assessment classified 83.6% of studies as high risk of bias. The dominant drivers were single-centre recruitment (78%), convenience sampling (64%), absent external validation (75.5%), and unaddressed class imbalance (53%). Three specific bias sources have outsized impact on translational inference. Optimism bias — from internal cross-validation on small datasets — accounts for the directional pattern of performance reduction observed across all major algorithm families when moving from single-centre to external testing. Feature selection leakage, identified in 31% of model development studies, occurs when dimensionality reduction is performed prior to cross-validation fold splitting, artificially inflating performance — a problem particularly prevalent in multi-omics biomarker discovery research.

Taxonomy assignment inter-rater agreement was assessed separately from general data extraction. On a 20% random subsample (n = 23 studies), taxonomy classification yielded κ=0.84 (substantial agreement). The primary source of disagreement was the boundary between Category 2 (Medical Imaging + AI) and Category 4 (Reproductive and Endocrine Health Applications), which accounted for 71% of discrepant cases. As pre-specified in Section 2.6, the applied decision rule was: Category 2 was assigned when the primary methodological contribution was an imaging architecture or segmentation method; Category 4 was assigned when the primary contribution was clinical integration of AI with reproductive health data regardless of imaging involvement. Studies where ultrasound features were combined with clinical parameters without novel imaging methodology were classified as Category 4. Temporal validation was absent in 89% of all studies, meaning that no model in the reviewed literature has demonstrated sustained performance over changing clinical populations — a prerequisite for safe, ongoing clinical use.

Explainability adoption, while improving from 0% in 2021 to 34.3% in 2025, remains insufficient for clinical integration. SHAP and LIME were applied in only 6.9% of studies, predominantly in PCOS, GDM, and preeclampsia prediction models where feature importance explanations can directly inform clinical reasoning. The 26.7% of studies employing no interpretability mechanism whatsoever represent “black box” outputs that are unlikely to gain clinician trust or regulatory approval. Hybrid CNN-ML architectures — used in 92% of AI/ML Methodology development studies — offer promising pathways to combining imaging performance with structured explainability, but require substantially larger and more diverse training datasets than currently available in most reproductive health research settings.

4.4 Geographic equity and translational challenges

Several deep-seated structural inequities were identified via the geographic analysis. The East Asia region comprised 39.6% of all studies (China n = 28; South Korea n = 15), while the region with the fewest was Sub-Saharan Africa with 1.7% (2 studies). Population-adjusted research output was calculated as: (number of studies in region) ÷ (women aged 15–49 in region, in millions). Population denominators were drawn from UN World Population Prospects 2023 estimates aggregated by World Bank income classification (HIC vs LMIC). HICs contributed 56 studies against approximately 250 million women of reproductive age, yielding 0.224 studies per million. LMICs contributed 60 studies against approximately 1,540 million women, yielding 0.039 studies per million — a ratio of 5.7:1 in absolute terms and 23:1 after weighting by reproductive-health DALY burden (Global Burden of Disease 2021). Sensitivity analyses using alternative denominators are reported in the next paragraph. Adjusted ratios from sensitivity analyses were: GDP = 7:1; research workforce FTEs = 17:1; reproductive health DALYs = 31:1. South Asia contributed only 15.5% of included studies; India contributed 18 studies but all were conducted in urban tertiary care centres, the 3 studies from Pakistan included fewer than 350 participants, there were 4 single-centre studies from Bangladesh, and none from Nepal and Sri Lanka. This geographical bias suggests that current models of population-specific parameters (e.g., hormone reference ranges, BMI–adiposity relations, PCOS clinical and laboratory features) may not be validated in South Asian and African populations. Infrastructure constraints compound the equity gap. Models requiring Tier 3 inputs (MRI, advanced ultrasound, multi-omics) fall in the Very Low Feasibility category, placing them outside the reach of primary care settings. Data poverty — small samples, missing data, fragmented paper-based records — affects 89% of LMIC-originating studies. Sociocultural barriers, including stigma surrounding menstruation and infertility, further constrain data availability and AI adoption, necessitating community co-design approaches largely absent from current literature.

4.5 Translational research agenda and deployment considerations

Fig 8 summarises a three-tier translational research agenda derived from the patterns identified in this review. The agenda is presented as a research roadmap, not a deployment guideline; empirical validation through prospective LMIC studies and Delphi-based expert consensus is required before any component can inform clinical practice.

Fig 8. Three-tier translational research agenda for low-resource settings.

Fig 8

An actionable deployment roadmap for AI-based women’s reproductive health tools in low- and middle-income countries.

4.5.1 Near-term priorities (0–2 years).

Reviewed evidence supports immediate research investment in:

  1. interpretable ensemble models (Random Forest, XGBoost) trained on Tier 1 inputs for PCOS, GDM, and thyroid screening, with mandatory SHAP-based explainability;

  2. prospective external validation of these models in LMIC primary-care populations; and

  3. centralised, de-identified data pooling among urban tertiary centres to build foundational reference cohorts.

4.5.2 Medium-term aspirations (3–5 years).

Federated learning and multi-site domain adaptation represent medium-term targets contingent on prerequisite infrastructure: standardised EHR systems at the primary care level, reliable internet connectivity in rural areas, multi-site IRB harmonisation, and locally enforceable data protection legislation. None of these prerequisites currently obtains in most South Asian or Sub-Saharan African settings. Existing frameworks such as OpenHIE [138,139] and WHO SMART guidelines [140] provide entry points but require substantial adaptation. Federated learning should not be interpreted as a near-term deployment recommendation.

4.5.3 Long-term horizon (5 + years).

Embedding AI competency in nursing and community health worker curricula, LMIC-led validation consortia, and locally governed regulatory frameworks for clinical AI represent structural changes that will determine whether the technical advances of the next decade reach the populations currently absent from the evidence base.

4.6 Comparison with prior reviews and methodological limitations

This review extends prior reviews of AI in reproductive health by providing descriptive evidence consistent with well-documented optimism bias patterns, a resource-tier classification to characterise infrastructure requirements, and a geographic equity analysis across 27 countries. Prior reviews have characterised the field as showing “high promise” without systematically describing the pattern between single-centre and externally validated results — a pattern this review characterises descriptively across five algorithm families. Several methodological limitations must be acknowledged: restriction to English-language publications likely excludes relevant research from Francophone Sub-Saharan Africa and Portuguese-speaking contexts; exclusion of preprints may introduce publication lag in a rapidly evolving field; heterogeneity in outcome measures meant formal meta-analysis was not appropriate; and the scoping nature of the review means individual study quality grading carries less evidential weight than a full systematic review.

4.7 Domain adaptation and harmonization as mitigation strategies for imaging performance decay

The larger reported performance decay in imaging-based models [141] relative to clinical-data models is consistent with well-documented scanner-level and protocol-level domain shift in medical imaging AI. Emerging domain adaptation and harmonization methods including histogram matching [142], CycleGAN [143]-based style transfer, and test-time adaptation approaches [144] have demonstrated decay reduction in comparable imaging tasks such as mammography and chest radiography across multi-site deployments. However, applying these approaches in LMIC reproductive imaging contexts faces a fundamental barrier: domain adaptation tools require paired or unpaired multi-site imaging data for training, which is currently unavailable given the near-absence of digitized reproductive imaging archives in primary and secondary care settings across South Asia and Sub-Saharan Africa. The practical implication of the imaging decay pattern is therefore not to invest in harmonization infrastructure in the near term, but rather to reinforce the prioritisation of clinical-data ensemble models for LMIC deployment. Domain adaptation investments become relevant only after multi-site imaging data collection infrastructure is established — a medium-to-long-term horizon. Researchers in HIC settings developing imaging models intended for LMIC adaptation should prospectively collect harmonization-enabling datasets and report scanner metadata to facilitate future transfer.

5. Conclusions

This scoping review synthesises 116 studies on AI and ML applications in women’s reproductive and hormonal health (2021–2025), offering four principal contributions: a six-category taxonomy of methodological trends; a qualitative resource-tier classification for LMIC applicability; descriptive evidence of systematic performance reduction upon external validation; and a geographic mapping revealing significant research voids. Ensemble methods applied to Tier 1 clinical and laboratory data represent the most immediately deployable AI approach for LMIC settings, with inputs available at primary care level, while imaging-intensive CNN models remain unsuitable for near-term LMIC deployment due to infrastructure constraints. With 83.6% of studies classified as high risk of bias and 75.5% lacking external validation, and with the evidence base heavily concentrated in East Asia (39.6%) and high-income countries, current evidence is insufficient to guide population-level deployment or broad translational application in low-resource settings. For LMICs such as Bangladesh, a phased strategy is proposed, contingent on resolution of the methodological limitations identified in this review: piloting of low-cost AI screening tools for PCOS, GDM, and thyroid disorders should proceed only after prospective validation in local populations; medium-term establishment of multi-country validation consortia; and long-term embedding of AI competency in health worker curricula alongside LMIC-contextualised regulatory frameworks. Beyond the descriptive synthesis, the four-tier LMIC feasibility classification system introduced in this review is intended as a reusable instrument: future researchers, funders, and health-system planners can apply it independently to characterize the deployment readiness of new AI tools in resource-constrained settings, regardless of disease domain. Priority future directions include prospective multi-centre LMIC validation, non-invasive endometriosis diagnostics, algorithmic fairness evaluation, and implementation science — all oriented toward ensuring that AI progress in reproductive health reaches the women who bear the greatest unmet burden. Consequently, the findings of this review reflect trends predominantly from East Asian and high-income settings and should not be interpreted as globally generalizable. Translational recommendations for LMICs remain theoretical until validated in local, resource-constrained populations. These conclusions should be interpreted within the descriptive scope of scoping review methodology and not as causal or comparative inferences.

Supporting information

S1 File. PRISMA-ScR Checklist.

Completed PRISMA-ScR checklist covering all 20 essential items and 2 optional items.

(DOCX)

pone.0358557.s001.docx (10.4KB, docx)
S1 Table. Risk-of-bias scoring rubric.

Explicit domain-specific criteria used to classify studies as low, moderate, or high risk of bias (adapted from PROBAST-AI, QUADAS-AI, and TRIPOD-AI).

(DOCX)

pone.0358557.s002.docx (8.9KB, docx)
S2 Table. LMIC applicability categorical criteria.

Full definitions of the four LMIC-feasibility categories (High, Moderate, Low, Very Low) used to classify studies.

(DOCX)

pone.0358557.s003.docx (9.6KB, docx)

Data Availability

Yes - all data are fully available without restriction; All data underlying this review are fully available in the Mendeley Data repository (https://doi.org/10.17632/ydx7v73t3k.2). The repository contains: (1) the full list of 116 included studies with citations; (2) the complete extracted study-level variables dataset; (3) coding and category definitions for the six taxonomy categories and LMIC feasibility tiers; (4) the complete, reproducible search strategies for all five databases; (5) the screening results with reasons for exclusion at each stage; and (6) the raw data used to generate all tables, figures, counts, and proportions. The completed PRISMA-ScR checklist (S1 File), risk-of-bias scoring rubric (S2 Table), and LMIC applicability categorical criteria (S2 Table) are also provided as supplementary files.

Funding Statement

The author(s) received no specific funding for this work.

References

  • 1.Zeng W, Gan D, Ou J, Tomlinson B. Global trends and health system impact on polycystic ovary syndrome: A comprehensive analysis of age-stratified females from 1990 to 2021. Front Reprod Health. 2025;7:1642369. doi: 10.3389/frph.2025.1642369 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Fauser BCJM, Adamson GD, Boivin J, Chambers GM, de Geyter C, Dyer S, et al. Declining global fertility rates and the implications for family planning and family building: An IFFS consensus document based on a narrative review of the literature. Hum Reprod Update. 2024;30(2):153–73. doi: 10.1093/humupd/dmad028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Guo Z-Q. Precision pharmacology in menopause: Advances, challenges, and future innovations for personalized management. Front Reprod Health. 2025;7:1694240. doi: 10.3389/frph.2025.1694240 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Gao X, Zhao S, Du Y, Yang Z, Tian Y, Zhao J, et al. Data-driven subtypes of polycystic ovary syndrome and their association with clinical outcomes. Nat Med. 2025;31(12):4214–24. doi: 10.1038/s41591-025-03984-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zhang J, Pang M, Li L, Guo C. Global, regional, and national burden of endometriosis among women of reproductive age, 1990–2021: Insights from the global burden of disease study 2021. PLoS One. 2025;20(11):e0337074. doi: 10.1371/journal.pone.0337074 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.World Health Organization. Polycystic ovary syndrome; 2025. [cited 2026 Jan 23]. Available from: https://www.who.int/news-room/fact-sheets/detail/polycystic-ovary-syndrome
  • 7.De Corte P, Klinghardt M, von Stockum S, Heinemann K. Time to diagnose endometriosis: Current status, challenges and regional characteristics-A systematic literature review. BJOG. 2025;132(2):118–30. doi: 10.1111/1471-0528.17973 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.de Kok L, Boersen Z, Coppus S, van Haaps A, van Hanegem N, Klinkert E. Diagnostic delay in endometriosis: Is there any progress? Reprod BioMed Online. 2025:105405. [DOI] [PubMed] [Google Scholar]
  • 9.Li R, Zhang L, Liu Y. Global and regional trends in the burden of surgically confirmed endometriosis from 1990 to 2021. Reprod Biol Endocrinol. 2025;23(1):88. doi: 10.1186/s12958-025-01421-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Yi S, Yam ELY, Cheruvettolil K, Linos E, Gupta A, Palaniappan L, et al. Perspectives of digital health innovations in low- and middle-income health care systems from South and Southeast Asia. J Med Internet Res. 2024;26:e57612. doi: 10.2196/57612 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Jonayed M, Rumi MH. Towards women’s digital health equity: A qualitative inquiry into attitude and adoption of reproductive mHealth services in Bangladesh. PLOS Digit Health. 2024;3(10):e0000637. doi: 10.1371/journal.pdig.0000637 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ghaderzadeh M, Garavand A, Salehnasab C. Artificial intelligence in polycystic ovary syndrome: A systematic review of diagnostic and predictive applications. BMC Med Inform Decis Mak. 2025;25(1):427. doi: 10.1186/s12911-025-03255-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Agirsoy M, Oehlschlaeger MA. A machine learning approach for non-invasive PCOS diagnosis from ultrasound and clinical features. Sci Rep. 2025;15(1):33638. doi: 10.1038/s41598-025-10453-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Burla L, Metzler JM, Kalaitzopoulos DR, Kamm S, Ormos M, Passweg D, et al. Artificial intelligence in endometriosis care: A comparative analysis of large language model and human specialist responses to endometriosis-related queries. Eur J Obstet Gynecol Reprod Biol. 2025;313:114625. doi: 10.1016/j.ejogrb.2025.114625 [DOI] [PubMed] [Google Scholar]
  • 15.Mengistu S, Tamrat T, Betran A-P, Pirsch S, Ferretti A, Mburu G, et al. The use of artificial intelligence in sexual and reproductive health: A comprehensive scoping review. NPJ Womens Health. 2025;3(1):70. doi: 10.1038/s44294-025-00118-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhu T, Li K, Herrero P, Georgiou P. Deep learning for diabetes: A systematic review. IEEE J Biomed Health Inform. 2021;25(7):2744–57. doi: 10.1109/JBHI.2020.3040225 [DOI] [PubMed] [Google Scholar]
  • 17.Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Ann Intern Med. 2018;169(7):467–73. doi: 10.7326/M18-0850 [DOI] [PubMed] [Google Scholar]
  • 18.Munna MMH, Islam T, Sultana A, Saima F, Faruk O, Sultana T, et al. Screening protocol and study selection data for: Artificial intelligence for women’s reproductive health: A systematic review of global diagnostic trends and a translational informatics framework for low-resource settings; 2026. Data repository. Available from: 10.17632/ydx7v73t3k.2 [DOI] [PMC free article] [PubMed]
  • 19.Kloda LA, Boruff JT, Cavalcante AS. A comparison of patient, intervention, comparison, outcome (PICO) to a new, alternative clinical question framework for search skills, search results, and self-efficacy: A randomized controlled trial. J Med Libr Assoc. 2020;108(2):185–94. doi: 10.5195/jmla.2020.739 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan-a web and mobile app for systematic reviews. Syst Rev. 2016;5(1):210. doi: 10.1186/s13643-016-0384-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Rayyan Systems Inc. Rayyan; 2026. Web application. Available from: https://www.rayyan.ai
  • 22.Balampanos D, Kokkotis C, Stampoulis T, Avloniti A, Pantazis D, Protopapa M, et al. Interpretable machine learning for osteopenia detection: A proof-of-concept study using bioelectrical impedance in perimenopausal women. J Funct Morphol Kinesiol. 2025;10(3):262. doi: 10.3390/jfmk10030262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Shubhangi DC, Salma U. Analysis and interpretation of physiological social demographic parameter in menopausal women. IJCSMC. 2024;13(10):58–66. doi: 10.47760/ijcsmc.2024.v13i10.007 [DOI] [Google Scholar]
  • 24.Chang C-Y, Peng C-H, Chen F-Y, Huang L-Y, Kuo C-H, Chu T-W, et al. The risk factors determined by four machine learning methods for the change of difference of bone mineral density in post-menopausal women after three years follow-up. Sci Rep. 2024;14(1):23234. doi: 10.1038/s41598-024-73799-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kim H, Khomidov M, Lee J-H. XGBoost and SHAP-based analysis of risk factors for hypertension classification in Korean postmenopausal women. Bioengineering (Basel). 2025;12(6):659. doi: 10.3390/bioengineering12060659 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Ding X, Tao T, Fu J, Zhang Y, Yang X, Ren M, et al. Is integrative therapy of traditional Chinese medicine and progesterone capsule more effective than monotherapies in oligomenorrhea and hypomenorrhea? Evidence based on a multi-center randomized controlled trial and metabolomic profile. J Transl Int Med. 2025;13(4):349–65. doi: 10.1515/jtim-2025-0027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Santos L, Azevedo M, Shimamura L, Nogueira A, Reis F, Schor E, et al. Machine learning as a clinical decision support tool for diagnosing superficial peritoneal endometriosis in women with dysmenorrhea and acyclic pelvic pain. MRAJ. 2024;12(12). doi: 10.18103/mra.v12i12.6204 [DOI] [Google Scholar]
  • 28.Manjunath H, Kolekar V, Bhuvaneshwari K, Jain R, Aravind A, Dev S. Comprehensive detection and analysis of depression in women through advanced machine learning techniques. 2024 IEEE 4th International Conference on ICT in Business Industry & Government (ICTBIG). IEEE; 2024. p. 1–5.
  • 29.Wang W, Zeng W, Yang S. A stacked machine learning-based classification model for endometriosis and adenomyosis: A retrospective cohort study utilizing peripheral blood and coagulation markers. Front Digit Health. 2024;6:1463419. doi: 10.3389/fdgth.2024.1463419 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Liu Z, Liu Z, Wang Y, Wan X, Huang X. Machine learning-based predictive analysis of energy efficiency factors necessary for the HIFU treatment of adenomyosis. Front Physiol. 2025;16:1602866. doi: 10.3389/fphys.2025.1602866 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Raimondo D, Raffone A, Aru AC, Giorgi M, Giaquinto I, Spagnolo E, et al. Application of deep learning model in the sonographic diagnosis of uterine adenomyosis. Int J Environ Res Public Health. 2023;20(3):1724. doi: 10.3390/ijerph20031724 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Shahzad A, Mushtaq A, Sabeeh AQ, Ghadi YY, Mushtaq Z, Arif S, et al. Automated uterine fibroids detection in ultrasound images using deep convolutional neural networks. Healthcare (Basel). 2023;11(10):1493. doi: 10.3390/healthcare11101493 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sawant A, Kulkarni S, Sawant M. Ultrasound super resolution imaging for accurate uterus tumor detection and malignancy prediction. J Pharm Biomed Anal Open. 2024;3:100029. doi: 10.1016/j.jpbao.2024.100029 [DOI] [Google Scholar]
  • 34.Samarasam B, Justin J. Machine learning-based approach for uterine cancer detection and classifier evaluation. J Electron Electromedical Eng Med Inform. 2025;7(3):940–9. doi: 10.35882/jeeemi.v7i3.632 [DOI] [Google Scholar]
  • 35.Akpinar E, Bayrak O-C, Nadarajan C, Müslümanoğlu M-H, Nguyen M-D, Keserci B. Role of machine learning algorithms in predicting the treatment outcome of uterine fibroids using high-intensity focused ultrasound ablation with an immediate nonperfused volume ratio of at least 90. Eur Rev Med Pharmacol Sci. 2022;26(22):8376–94. doi: 10.26355/eurrev_202211_30373 [DOI] [PubMed] [Google Scholar]
  • 36.Janghorbani S, Caprio A, Sam L, Lee BC, Sabuncu MR, Lamparello NA, et al. Predicting clinical outcomes and symptom relief in uterine fibroid embolization using machine learning on MRI features. AI. 2025;6(9):200. doi: 10.3390/ai6090200 [DOI] [Google Scholar]
  • 37.Chen M, Kong W, Li B, Tian Z, Yin C, Zhang M, et al. Revolutionizing hysteroscopy outcomes: AI-powered uterine myoma diagnosis algorithm shortens operation time and reduces blood loss. Front Oncol. 2023;13:1325179. doi: 10.3389/fonc.2023.1325179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Toyohara Y, Sone K, Noda K, Yoshida K, Kato S, Kaiume M, et al. The automatic diagnosis artificial intelligence system for preoperative magnetic resonance imaging of uterine sarcoma. J Gynecol Oncol. 2024;35(3):e24. doi: 10.3802/jgo.2024.35.e24 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Xi H, Wang W. Deep learning based uterine fibroid detection in ultrasound images. BMC Med Imaging. 2024;24(1):218. doi: 10.1186/s12880-024-01389-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Theis M, Tonguc T, Savchenko O, Nowak S, Block W, Recker F, et al. Deep learning enables automated MRI-based estimation of uterine volume also in patients with uterine fibroids undergoing high-intensity focused ultrasound therapy. Insights Imaging. 2023;14(1):1. doi: 10.1186/s13244-022-01342-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Li C, He Z, Lv F, Liu Y, Hu Y, Zhang J, et al. An interpretable MRI-based radiomics model predicting the prognosis of high-intensity focused ultrasound ablation of uterine fibroids. Insights Imaging. 2023;14(1):129. doi: 10.1186/s13244-023-01445-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Huo T, Li L, Chen X, Wang Z, Zhang X, Liu S, et al. Artificial intelligence-aided method to detect uterine fibroids in ultrasound images: A retrospective study. Sci Rep. 2023;13(1):3714. doi: 10.1038/s41598-022-26771-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Mohanty A, Pattnayak P, Mallick PK, Ambudkar B. Deep convolutional neural networks for automatic detection of uterine fibroids in ultrasound images. Intell Decis Technol. 2025;19(3):1657–74. doi: 10.1177/18724981241309994 [DOI] [Google Scholar]
  • 44.Rewcastle E, Gudlaugsson E, Lillesand M, Skaland I, Baak JPA, Janssen EAM. Automated prognostic assessment of endometrial hyperplasia for progression risk evaluation using artificial intelligence. Mod Pathol. 2023;36(5):100116. doi: 10.1016/j.modpat.2023.100116 [DOI] [PubMed] [Google Scholar]
  • 45.Martins MS, Valente GB, Pedra YdN, Ribeiro TC, Boldrini NAT, Barcelos MRB, et al. A machine learning approach towards endometriosis screening using infrared spectra of urine. Clinics (Sao Paulo). 2025;80:100760. doi: 10.1016/j.clinsp.2025.100760 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Xie Z, Feng Y, He Y, Lin Y, Wang X. Identification of biomarkers for endometriosis based on summary-data-based Mendelian randomization and machine learning. Medicine (Baltimore). 2025;104(14):e41804. doi: 10.1097/MD.0000000000041804 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Sankaravadivel V, Thalavaipillai S. Symptoms based endometriosis prediction using machine learning. Bulletin EEI. 2021;10(6):3102–9. doi: 10.11591/eei.v10i6.3254 [DOI] [Google Scholar]
  • 48.Macis C, Santoro M, Zybin V, Di Costanzo S, Coada CA, Dondi G, et al. A convolutional neural network tool for early diagnosis and precision surgery in endometriosis-associated ovarian cancer. Appl Sci. 2025;15(6):3070. doi: 10.3390/app15063070 [DOI] [Google Scholar]
  • 49.Caballero P, Gonzalez-Abril L, Ortega JA, Simon-Soro Á. Data mining techniques for endometriosis detection in a data-scarce medical dataset. Algorithms. 2024;17(3):108. doi: 10.3390/a17030108 [DOI] [Google Scholar]
  • 50.Wang H, Butler D, Zhang Y, Avery J, Knox S, Ma C, et al. Human-AI collaborative multi-modal multi-rater learning for endometriosis diagnosis. Phys Med Biol. 2024;70(1). doi: 10.1088/1361-6560/ad997e [DOI] [PubMed] [Google Scholar]
  • 51.Enamorado-Díaz E, Morales-Trujillo L, García-García J-A, Marcos AT, Navarro-Pando J, Escalona-Cuaresma M-J. A novel machine learning-based proposal for early prediction of endometriosis disease. Expert Syst Appl. 2025;271:126621. doi: 10.1016/j.eswa.2025.126621 [DOI] [Google Scholar]
  • 52.Kuyoro AO, Fatade OB, Onuiri EE. Enhancing non-invasive diagnosis of endometriosis through explainable artificial intelligence: A Grad-CAM approach. ATAIML. 2025;4(2):97–108. doi: 10.56578/ataiml040203 [DOI] [Google Scholar]
  • 53.Tore U, Abilgazym A, Asunsolo-Del-Barco A, Terzic M, Yemenkhan Y, Zollanvari A, et al. Diagnosis of endometriosis based on comorbidities: A machine learning approach. Biomedicines. 2023;11(11):3015. doi: 10.3390/biomedicines11113015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Korman SE, Vissers G, Gorris MAJ, Verrijp K, Verdurmen WPR, Simons M, et al. Artificial intelligence-based tissue segmentation and cell identification in multiplex-stained histological endometriosis sections. Hum Reprod. 2025;40(3):450–60. doi: 10.1093/humrep/deae267 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Chang X, Miao J. Identification of a disulfidptosis-related genes signature for diagnostic and immune infiltration characteristics in endometriosis. Sci Rep. 2024;14(1):25939. doi: 10.1038/s41598-024-77539-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Fell C, Mohammadi M, Morrison D, Arandjelović O, Syed S, Konanahalli P, et al. Detection of malignancy in whole slide images of endometrial cancer biopsies using artificial intelligence. PLoS One. 2023;18(3):e0282577. doi: 10.1371/journal.pone.0282577 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Sankaravadivel V, Thalavaipillai S, Rajeswar S, Ramlingam P. Feature based analysis of endometriosis using machine learning. IJEECS. 2023;29(3):1700. doi: 10.11591/ijeecs.v29.i3.pp1700-1707 [DOI] [Google Scholar]
  • 58.Zhang H, Zhang H, Yang H, Shuid AN, Sandai D, Chen X. Machine learning-based integrated identification of predictive combined diagnostic biomarkers for endometriosis. Front Genet. 2023;14:1290036. doi: 10.3389/fgene.2023.1290036 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Zou L, Meng L, Xu Y, Wang K, Zhang J. Revealing the diagnostic value and immune infiltration of senescence-related genes in endometriosis: A combined single-cell and machine learning analysis. Front Pharmacol. 2023;14:1259467. doi: 10.3389/fphar.2023.1259467 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Hu P, Gao Y, Zhang Y, Sun K. Ultrasound image-based deep learning to differentiate tubal-ovarian abscess from ovarian endometriosis cyst. Front Physiol. 2023;14:1101810. doi: 10.3389/fphys.2023.1101810 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Zhao D, Zhang Z, Wang Z, Du Z, Wu M, Zhang T, et al. Diagnosis and prediction of endometrial carcinoma using machine learning and artificial neural networks based on public databases. Genes (Basel). 2022;13(6):935. doi: 10.3390/genes13060935 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Takahashi Y, Sone K, Noda K, Yoshida K, Toyohara Y, Kato K, et al. Automated system for diagnosing endometrial cancer by adopting deep-learning technology in hysteroscopy. PLoS One. 2021;16(3):e0248526. doi: 10.1371/journal.pone.0248526 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Balogh DB, Hudelist G, Bļizņuks D, Raghothama J, Becker CM, Horace R, et al. FEMaLe: The use of machine learning for early diagnosis of endometriosis based on patient self-reported data-Study protocol of a multicenter trial. PLoS One. 2024;19(5):e0300186. doi: 10.1371/journal.pone.0300186 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Blass I, Sahar T, Shraibman A, Ofer D, Rappoport N, Linial M. Revisiting the risk factors for endometriosis: A machine learning approach. J Pers Med. 2022;12(7):1114. doi: 10.3390/jpm12071114 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Fremond S, Andani S, Barkey Wolf J, Dijkstra J, Melsbach S, Jobsen JJ, et al. Interpretable deep learning model to predict the molecular classification of endometrial cancer from haematoxylin and eosin-stained whole-slide images: A combined analysis of the PORTEC randomised trials and clinical cohorts. Lancet Digit Health. 2023;5(2):e71–82. doi: 10.1016/S2589-7500(22)00210-2 [DOI] [PubMed] [Google Scholar]
  • 66.Pal A, Biswas S, O Kare SP, Biswas P, Jana SK, Das S, et al. Development of an impedimetric immunosensor for machine learning-based detection of endometriosis: A proof of concept. Sens Actuators B: Chem. 2021;346:130460. doi: 10.1016/j.snb.2021.130460 [DOI] [Google Scholar]
  • 67.Zhao N, Hao T, Zhang F, Ni Q, Zhu D, Wang Y, et al. Application of machine learning techniques in the diagnosis of endometriosis. BMC Womens Health. 2024;24(1):491. doi: 10.1186/s12905-024-03334-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Shi S, Huang C, Tang X, Liu H, Feng W, Chen C. Identification and verification of diagnostic biomarkers for deep infiltrating endometriosis based on machine learning algorithms. J Biol Eng. 2024;18(1):70. doi: 10.1186/s13036-024-00466-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Goldstein A, Cohen S. Self-report symptom-based endometriosis prediction using machine learning. Sci Rep. 2023;13(1):5499. doi: 10.1038/s41598-023-32761-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Patil PS, Patil MB, Jadhav DK, Patil SB, Bodhe MYU. Predictive analytics and health monitoring system for early detection of infertility risks among working women. Int J Environ Sci. 2025;11(16s). [Google Scholar]
  • 71.Lin Q, Fang Z-J. Establishment and evaluation of a risk prediction model for gestational diabetes mellitus. World J Diabetes. 2023;14(10):1541–50. doi: 10.4239/wjd.v14.i10.1541 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Prashanthan J, Prashanthan A. Predicting the future risk of developing type 2 diabetes in women with a history of gestational diabetes mellitus using machine learning and explainable artificial intelligence. Prim Care Diabetes. 2025;19(6):658–66. doi: 10.1016/j.pcd.2025.09.006 [DOI] [PubMed] [Google Scholar]
  • 73.Wu Y-T, Zhang C-J, Mol BW, Kawai A, Li C, Chen L, et al. Early prediction of gestational diabetes mellitus in the Chinese population via advanced machine learning. J Clin Endocrinol Metab. 2021;106(3):e1191–205. doi: 10.1210/clinem/dgaa899 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Vivek Khanna V, Chadaga K, Sampathila N, Prabhu S, Chadaga P R, Bhat D, et al. Explainable artificial intelligence-driven gestational diabetes mellitus prediction using clinical and laboratory markers. Cogent Eng. 2024;11(1):2330266. doi: 10.1080/23311916.2024.2330266 [DOI] [Google Scholar]
  • 75.Hu X, Hu X, Yu Y, Wang J. Prediction model for gestational diabetes mellitus using the XG Boost machine learning algorithm. Front Endocrinol (Lausanne). 2023;14:1105062. doi: 10.3389/fendo.2023.1105062 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Zhou H, Chen W, Chen C, Zeng Y, Chen J, Lin J, et al. Predictive value of ultrasonic artificial intelligence in placental characteristics of early pregnancy for gestational diabetes mellitus. Front Endocrinol (Lausanne). 2024;15:1344666. doi: 10.3389/fendo.2024.1344666 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Ali N, Khan W, Ahmad A, Masud MM, Adam H, Ahmed LA. Predictive modeling for the diagnosis of gestational diabetes mellitus using epidemiological data in the United Arab Emirates. Information. 2022;13(10):485. doi: 10.3390/info13100485 [DOI] [Google Scholar]
  • 78.Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T, Yavuz AA. The early prediction of gestational diabetes mellitus by machine learning models. BMC Pregnancy Childbirth. 2024;24(1):574. doi: 10.1186/s12884-024-06783-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Bigdeli SK, Ghazisaedi M, Ayyoubzadeh SM, Hantoushzadeh S, Ahmadi M. Predicting Gestational Diabetes Mellitus in the first trimester using machine learning algorithms: A cross-sectional study at a hospital fertility health center in Iran. BMC Med Inform Decis Mak. 2025;25(1):3. doi: 10.1186/s12911-024-02799-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zaky H, Fthenou E, Srour L, Farrell T, Bashir M, El Hajj N, et al. Machine learning based model for the early detection of Gestational Diabetes Mellitus. BMC Med Inform Decis Mak. 2025;25(1):130. doi: 10.1186/s12911-025-02947-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Du Y, Rafferty AR, McAuliffe FM, Wei L, Mooney C. An explainable machine learning-based clinical decision support system for prediction of gestational diabetes mellitus. Sci Rep. 2022;12(1):1170. doi: 10.1038/s41598-022-05112-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Kang BS, Lee SU, Hong S, Choi SK, Shin JE, Wie JH, et al. Prediction of gestational diabetes mellitus in Asian women using machine learning algorithms. Sci Rep. 2023;13(1):13356. doi: 10.1038/s41598-023-39680-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Montgomery-Csobán T, Kavanagh K, Murray P, Robertson C, Barry SJE, Vivian Ukah U, et al. Machine learning-enabled maternal risk assessment for women with pre-eclampsia (the PIERS-ML model): A modelling study. Lancet Digit Health. 2024;6(4):e238–50. doi: 10.1016/S2589-7500(23)00267-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Zhang X, Chen Y, Salerno S, Li Y, Zhou L, Zeng X, et al. Prediction of severe preeclampsia in machine learning. Med Nov Technol Devices. 2022;15:100158. doi: 10.1016/j.medntd.2022.100158 [DOI] [Google Scholar]
  • 85.Domínguez-del Olmo P, Herraiz I, Villalaín C, Galindo A, Moreno-Espino M, Ayala JL. Comprehensive approach with machine learning techniques to investigate early-onset preeclampsia and its long-term cardiovascular implications. Appl Sci. 2025;15(16):8887. doi: 10.3390/app15168887 [DOI] [Google Scholar]
  • 86.Shyu I-L, Liu C-F, Tsai Y-C, Ma Y-S, Kuo T-N, Yow-Ling S. Machine learning predictive system to predict the risk of developing pre-eclampsia. BMJ Health Care Inform. 2025;32(1):e101151. doi: 10.1136/bmjhci-2024-101151 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Gómez-Jemes L, Oprescu AM, Chimenea-Toscano Á, García-Díaz L, Romero-Ternero MdC. Machine learning to predict pre-eclampsia and intrauterine growth restriction in pregnant women. Electronics. 2022;11(19):3240. doi: 10.3390/electronics11193240 [DOI] [Google Scholar]
  • 88.Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T. Risk assessment for preeclampsia in the preconception period based on maternal clinical history via machine learning methods. J Clin Med. 2024;14(1):155. doi: 10.3390/jcm14010155 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Araújo DC, de Macedo AA, Veloso AA, Alpoim PN, Gomes KB, Carvalho MdG, et al. Complete blood count as a biomarker for preeclampsia with severe features diagnosis: A machine learning approach. BMC Pregnancy Childbirth. 2024;24(1):628. doi: 10.1186/s12884-024-06821-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Wu Y, Shen L, Zhao L, Lin X, Xu M, Tu Z, et al. Noninvasive early prediction of preeclampsia in pregnancy using retinal vascular features. NPJ Digit Med. 2025;8(1):188. doi: 10.1038/s41746-025-01582-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Ansbacher-Feldman Z, Syngelaki A, Meiri H, Cirkin R, Nicolaides KH, Louzoun Y. Machine-learning-based prediction of pre-eclampsia using first-trimester maternal characteristics and biomarkers. Ultrasound Obstet Gynecol. 2022;60(6):739–45. doi: 10.1002/uog.26105 [DOI] [PubMed] [Google Scholar]
  • 92.Torres-Torres J, Villafan-Bernal J, Martinez-Portilla R, Hidalgo-Carrera J, Estrada-Gutierrez G, Adalid-Martinez-Cisneros R, et al. Performance of machine-learning approach for prediction of pre-eclampsia in a middle-income country. Ultrasound Obstet Gynecol. 2024;63(3):350–7. [DOI] [PubMed] [Google Scholar]
  • 93.Tiruneh SA, Rolnik DL, Teede HJ, Enticott J. Prediction of pre-eclampsia with machine learning approaches: Leveraging important information from routinely collected data. Int J Med Inform. 2024;192:105645. doi: 10.1016/j.ijmedinf.2024.105645 [DOI] [PubMed] [Google Scholar]
  • 94.Sufian MA, Hamzi W, Hamzi B, Sagar ASMS, Rahman M, Varadarajan J, et al. Innovative machine learning strategies for early detection and prevention of pregnancy loss: the vitamin D connection and gestational health. Diagnostics (Basel). 2024;14(9):920. doi: 10.3390/diagnostics14090920 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Yan S, Xiong F, Xin Y, Zhou Z, Liu W. Automated assessment of endometrial receptivity for screening recurrent pregnancy loss risk using deep learning-enhanced ultrasound and clinical data. Front Physiol. 2024;15:1404418. doi: 10.3389/fphys.2024.1404418 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Koong K, Preda V, Jian A, Liquet-Weiland B, Di Ieva A. Application of artificial intelligence and radiomics in pituitary neuroendocrine and sellar tumors: A quantitative and qualitative synthesis. Neuroradiology. 2022;64(4):647–68. doi: 10.1007/s00234-021-02845-1 [DOI] [PubMed] [Google Scholar]
  • 97.Fang Y, Wang H, Feng M, Zhang W, Cao L, Ding C, et al. Machine-learning prediction of postoperative pituitary hormonal outcomes in nonfunctioning pituitary adenomas: A multicenter study. Front Endocrinol (Lausanne). 2021;12:748725. doi: 10.3389/fendo.2021.748725 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Zheng A, Tang D, He H, Liang X. Artificial intelligence-driven approaches in pituitary neuroendocrine tumors: Integrating endocrine-metabolic profiling for enhanced diagnostics and therapeutics. Front Endocrinol (Lausanne). 2025;16:1618412. doi: 10.3389/fendo.2025.1618412 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Li Q, Zhu Y, Chen M, Guo R, Hu Q, Lu Y, et al. Development and validation of a deep learning algorithm to automatic detection of pituitary microadenoma from MRI. Front Med (Lausanne). 2021;8:758690. doi: 10.3389/fmed.2021.758690 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Dai C, Sun B, Wang R, Kang J. The application of artificial intelligence and machine learning in pituitary adenomas. Front Oncol. 2021;11:784819. doi: 10.3389/fonc.2021.784819 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Mohammadzadeh I, Hajikarimloo B, Niroomand B, Faizi N, Eini P, Habibi MA, et al. Prediction of recurrence after surgery for pituitary adenoma using machine learning- based models: Systematic review and meta-analysis. BMC Endocr Disord. 2025;25(1):158. doi: 10.1186/s12902-025-01955-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Park YW, Eom J, Kim S, Kim H, Ahn SS, Ku CR, et al. Radiomics with ensemble machine learning predicts dopamine agonist response in patients with prolactinoma. J Clin Endocrinol Metab. 2021;106(8):e3069–77. doi: 10.1210/clinem/dgab159 [DOI] [PubMed] [Google Scholar]
  • 103.Fan Y, Li Y, Bao X, Zhu H, Lu L, Yao Y, et al. Development of machine learning models for predicting postoperative delayed remission in patients with Cushing’s disease. J Clin Endocrinol Metab. 2021;106(1):e217–31. doi: 10.1210/clinem/dgaa698 [DOI] [PubMed] [Google Scholar]
  • 104.Guleria K, Sharma S, Kumar S, Tiwari S. Early prediction of hypothyroidism and multiclass classification using predictive machine learning and deep learning. Meas: Sens. 2022;24:100482. doi: 10.1016/j.measen.2022.100482 [DOI] [Google Scholar]
  • 105.Shiuh TL, Khai WK, Xin YC, Wai CY. Prediction of thyroid disease using machine learning approaches and featurewiz selection. JTEC. 2023;15(3):9–16. doi: 10.54554/jtec.2023.15.03.002 [DOI] [Google Scholar]
  • 106.Abbad Ur Rehman H, Lin C-Y, Mushtaq Z, Su S-F. Performance analysis of machine learning algorithms for thyroid disease. Arab J Sci Eng. 2021;46(10):9437–49. doi: 10.1007/s13369-020-05206-x [DOI] [Google Scholar]
  • 107.Afshan N, Mushtaq Z, Alamri FS, Qureshi MF, Khan NA, Siddique I. Efficient thyroid disorder identification with weighted voting ensemble of super learners by using adaptive synthetic sampling technique. AIMS Math. 2023;8(10):24274–309. doi: 10.3934/math.20231238 [DOI] [Google Scholar]
  • 108.Obaido G, Achilonu O, Ogbuokiri B, Amadi CS, Habeebullahi L, Ohalloran T, et al. An improved framework for detecting thyroid disease using filter-based feature selection and stacking ensemble. IEEE Access. 2024;12:89098–112. doi: 10.1109/access.2024.3418974 [DOI] [Google Scholar]
  • 109.Akhtar T, Gilani SO, Mushtaq Z, Arif S, Jamil M, Ayaz Y, et al. Effective voting ensemble of homogenous ensembling with multiple attribute-selection approaches for improved identification of thyroid disorder. Electronics. 2021;10(23):3026. doi: 10.3390/electronics10233026 [DOI] [Google Scholar]
  • 110.Akter S, Mustafa HA. Analysis and interpretability of machine learning models to classify thyroid disease. PLoS One. 2024;19(5):e0300670. doi: 10.1371/journal.pone.0300670 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Mir MS, Fayaz SA, Zaman M, Agrawal S. An application of traditional and ensemble machine learning approaches to redefine thyroid disorder diagnosis. MMEP. 2024;11(9):2437–46. doi: 10.18280/mmep.110916 [DOI] [Google Scholar]
  • 112.Sankar R, Mahulikar C, Viswanatha V. Thyroid disease detection using machine learning approach. J Xi’an Univ Archit Technol. 2023;7:327–34. doi: 10.37896/JXAT15.7/32228 [DOI] [Google Scholar]
  • 113.Sultana A, Islam R. Machine learning framework with feature selection approaches for thyroid disease classification and associated risk factors identification. J Electr Syst Inf Technol. 2023;10(1):32. doi: 10.1186/s43067-023-00101-5 [DOI] [Google Scholar]
  • 114.Gupta P, Rustam F, Kanwal K, Aljedaani W, Alfarhood S, Safran M, et al. Detecting thyroid disease using optimized machine learning model based on differential evolution. Int J Comput Intell Syst. 2024;17(1):3. doi: 10.1007/s44196-023-00388-2 [DOI] [Google Scholar]
  • 115.Sanju P, Ahmed NSS, Ramachandran P, Sajid PM, Jayanthi R. Enhancing thyroid disease prediction and comorbidity management through advanced machine learning frameworks. Clin eHealth. 2025;8:7–16. doi: 10.1016/j.ceh.2025.01.002 [DOI] [Google Scholar]
  • 116.Sankar S, Potti A, Chandrika GN, Ramasubbareddy S. Thyroid disease prediction using XGBoost algorithms. JMM. 2022. doi: 10.13052/jmm1550-4646.18322 [DOI] [Google Scholar]
  • 117.Faris N, Sahi A, Diykh M, Abdulla S, Siuly S. Enhanced Polycystic Ovary Syndrome diagnosis model leveraging a K-means based genetic algorithm and ensemble approach. Intell-Based Med. 2025;11:100253. doi: 10.1016/j.ibmed.2025.100253 [DOI] [Google Scholar]
  • 118.Yan X, Yang Z, Zhao H, Feng G, Li S, Li Y, et al. Unveiling lipoprotein subfractions signature in high-FNPO PCOS: Implications for PCOM diagnosis and risk assessment using advanced machine learning models. BMC Med. 2025;23(1):289. doi: 10.1186/s12916-025-04120-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Reka S, Praba TS, Prasanna M, Reddy VNN, Amirtharajan R. Automated high precision PCOS detection through a segment anything model on super resolution ultrasound ovary images. Sci Rep. 2025;15(1):16832. doi: 10.1038/s41598-025-01744-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Ghosh A, Srinivasan K. EffiDenseGenOp: ensemble transfer learning with hyperparameter tuning using genetic algorithm optimization for PCOS detection from ultrasound sonography images. IEEE Access. 2025;13:54285–312. doi: 10.1109/access.2025.3553895 [DOI] [Google Scholar]
  • 121.Zhu S, Huang Z, Chen X, Jiang W, Zhou Y, Zheng B, et al. Construction and evaluation of machine learning-based prediction model for live birth following fresh embryo transfer in IVF/ICSI patients with polycystic ovary syndrome. J Ovarian Res. 2025;18(1):70. doi: 10.1186/s13048-025-01654-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Zhao B, Wen L, Huang Y, Fu Y, Zhou S, Liu J, et al. A deep learning-based automatic recognition model for polycystic ovary ultrasound images. Balkan Med J. 2025;42(5):419–28. doi: 10.4274/balkanmedj.galenos.2025.2025-5-114 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Xu H, Mao L, Huang W, Huang Q, Li L, Liu Y. Gene association study between polycystic ovary syndrome and metabolic syndrome: A transcriptomic analysis and machine learning approach. J Ovarian Res. 2025;18(1):220. doi: 10.1186/s13048-025-01787-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Mohi Uddin KM, Bhuiyan MdTA, Rahman MdM, Islam MdM, Uddin MA. Early PCOS detection: A comparative analysis of traditional and ensemble machine learning models with advanced feature selection. Eng Rep. 2025;7(2):e70008. doi: 10.1002/eng2.70008 [DOI] [Google Scholar]
  • 125.Tong C, Wu Y, Zhuang Z, Yu Y. A diagnostic model for polycystic ovary syndrome based on machine learning. Sci Rep. 2025;15(1):9821. doi: 10.1038/s41598-025-92630-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Panjwani B, Yadav J, Mohan V, Agarwal N, Agarwal S. Optimized machine learning for the early detection of polycystic ovary syndrome in women. Sensors (Basel). 2025;25(4):1166. doi: 10.3390/s25041166 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Chen W, Miao J, Chen J, Chen J. Development of machine learning models for diagnostic biomarker identification and immune cell infiltration analysis in PCOS. J Ovarian Res. 2025;18(1):1. doi: 10.1186/s13048-024-01583-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Rahman MM, Islam A, Islam F, Zaman M, Islam MR, Alam Sakib MS, et al. Empowering early detection: A web-based machine learning approach for PCOS prediction. Inform Med Unlocked. 2024;47:101500. doi: 10.1016/j.imu.2024.101500 [DOI] [Google Scholar]
  • 129.Elmannai H, El-Rashidy N, Mashal I, Alohali MA, Farag S, El-Sappagh S, et al. Polycystic ovary syndrome detection machine learning model based on optimized feature selection and explainable artificial intelligence. Diagnostics (Basel). 2023;13(8):1506. doi: 10.3390/diagnostics13081506 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Saha A, Roy A, Chakraborty B, Saha B, Chowdhury D. A comparative study to predict polycystic ovarian syndrome (PCOS) based on different models of machine learning technique. AJEC. 2023;4(2):1–6. doi: 10.15864/ajec.4201 [DOI] [Google Scholar]
  • 131.Zad Z, Jiang VS, Wolf AT, Wang T, Cheng JJ, Paschalidis IC, et al. Predicting polycystic ovary syndrome with machine learning algorithms from electronic health records. Front Endocrinol (Lausanne). 2024;15:1298628. doi: 10.3389/fendo.2024.1298628 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Jha T, Sirisha M, Bhargavi M. Ovulytics: A machine learning approach for precision diagnosis of PCOD, PCOD and infertility. 2024 First International Conference for Women in Computing (InCoWoCo). IEEE; 2024. p. 1–7.
  • 133.Suha SA, Islam MN. An extended machine learning technique for polycystic ovary syndrome detection using ovary ultrasound image. Sci Rep. 2022;12(1):17123. doi: 10.1038/s41598-022-21724-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Lee S, Arffman RK, Komsi EK, Lindgren O, Kemppainen JA, Metsola H, et al. AI-algorithm training and validation for identification of endometrial CD138+ cells in infertility-associated conditions; polycystic ovary syndrome (PCOS) and recurrent implantation failure (RIF). J Pathol Inform. 2024;15:100380. doi: 10.1016/j.jpi.2024.100380 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Abdesselam A, Zidoum H, Zadjali F, Hedjam R, Al-Ansari A, Bayoumi R, et al. Estimate of the HOMA-IR cut-off value for identifying subjects at risk of insulin resistance using a machine learning approach. Sultan Qaboos Univ Med J. 2021;21(4):604–12. doi: 10.18295/squmj.4.2021.030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Huang X, Yi K, Jia L, Li Y, He H, Ma C, et al. Development and validation of an insulin resistance prediction model in children and adolescents using machine learning algorithms. Transl Pediatr. 2025;14(3):452–62. doi: 10.21037/tp-2024-502 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. [Google Scholar]
  • 138.Mamuye AL, Yilma TM, Abdulwahab A, Broomhead S, Zondo P, Kyeng M, et al. Health information exchange policy and standards for digital health systems in Africa: A systematic review. PLOS Digit Health. 2022;1(10):e0000118. doi: 10.1371/journal.pdig.0000118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Ndlovu K, Mars M, Scott RE. Validation of an interoperability framework for linking mHealth apps to electronic record systems in Botswana: Expert survey study. JMIR Form Res. 2023;7:e41225. doi: 10.2196/41225 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Mehl G, Tunçalp Ö, Ratanaprayul N, Tamrat T, Barreix M, Lowrance D, et al. WHO SMART guidelines: Optimising country-level use of guideline recommendations in the digital age. Lancet Digit Health. 2021;3(4):e213–6. doi: 10.1016/S2589-7500(21)00038-8 [DOI] [PubMed] [Google Scholar]
  • 141.Sahiner B, Chen W, Samala RK, Petrick N. Data drift in medical machine learning: Implications and potential remedies. Br J Radiol. 2023;96(1150):20220878. doi: 10.1259/bjr.20220878 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Fusco R, Granata V, Vallone P, Petrosino T, Iasevoli MD, Raso MM, et al. Engineering the image representation for deep learning in contrast-enhanced mammography: A systematic analysis of preprocessing and anatomical masking. Bioengineering (Basel). 2026;13(3):322. doi: 10.3390/bioengineering13030322 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Zhu JY, Park T, Isola P, Efros AA. Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision; 2017. p. 2223–32.
  • 144.Wang D, Shelhamer E, Liu S, Olshausen B, Darrell T. Tent: Fully test-time adaptation by entropy minimization. arXiv:200610726 [Preprint]. 2020.

Decision Letter 0

Kwang-Sig Lee

21 Jul 2026

Dear Dr. Munna,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Sep 19 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

  • A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only  the individual author can complete the verification step; PLOS staff cannot  verify ORCID iDs on behalf of authors.

We look forward to receiving your revised manuscript.

Kind regards,

Kwang-Sig Lee

Academic Editor

PLOS One

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf.

2. Please ensure that you have uploaded your PRISMA flowchart as Figure 1, or else explain why this is not possible. Blank flowcharts can be found here: http://www.prisma-statement.org/

3. We note that Figure 5 in your submission contain [map/satellite] images which may be copyrighted. All PLOS content is published under the Creative Commons Attribution License (CC BY 4.0), which means that the manuscript, images, and Supporting Information files will be freely available online, and any third party is permitted to access, download, copy, distribute, and use these materials in any way, even commercially, with proper attribution. For these reasons, we cannot publish previously copyrighted maps or satellite images created using proprietary data, such as Google software (Google Maps, Street View, and Earth). For more information, see our copyright guidelines: http://journals.plos.org/plosone/s/licenses-and-copyright.

We require you to either (1) present written permission from the copyright holder to publish these figures specifically under the CC BY 4.0 license, or (2) remove the figures from your submission:

a. You may seek permission from the original copyright holder of Figure 5 to publish the content specifically under the CC BY 4.0 license.

We recommend that you contact the original copyright holder with the Content Permission Form (http://journals.plos.org/plosone/s/file?id=7c09/content-permission-form.pdf) and the following text:

“I request permission for the open-access journal PLOS ONE to publish XXX under the Creative Commons Attribution License (CCAL) CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). Please be aware that this license allows unrestricted use and distribution, even commercially, by third parties. Please reply and provide explicit written permission to publish XXX under a CC BY license and complete the attached form.”

Please upload the completed Content Permission Form or other proof of granted permissions as an "Other" file with your submission.

In the figure caption of the copyrighted figure, please include the following text: “Reprinted from [ref] under a CC BY license, with permission from [name of publisher], original copyright [original copyright year].”

b. If you are unable to obtain permission from the original copyright holder to publish these figures under the CC BY 4.0 license or if the copyright holder’s requirements are incompatible with the CC BY 4.0 license, please either i) remove the figure or ii) supply a replacement figure that complies with the CC BY 4.0 license. Please check copyright information on all replacement figures and update the figure caption with source information. If applicable, please specify in the figure caption text when a figure is similar but not identical to the original image and is therefore for illustrative purposes only.

The following resources for replacing copyrighted map figures may be helpful:

USGS National Map Viewer (public domain): http://viewer.nationalmap.gov/viewer/

The Gateway to Astronaut Photography of Earth (public domain): http://eol.jsc.nasa.gov/sseop/clickmap/

Maps at the CIA (public domain): https://www.cia.gov/library/publications/the-world-factbook/index.html and https://www.cia.gov/library/publications/cia-maps-publications/index.html

NASA Earth Observatory (public domain): http://earthobservatory.nasa.gov/

Landsat: http://landsat.visibleearth.nasa.gov/

USGS EROS (Earth Resources Observatory and Science (EROS) Center) (public domain): http://eros.usgs.gov/#

Natural Earth (public domain): http://www.naturalearthdata.com/

4. Please remove your figures from within your manuscript file, leaving only the individual TIFF/EPS image files, uploaded separately. These will be automatically included in the reviewers’ PDF."

5. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information.

6. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

Reviewer #1: Partly

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #1: N/A

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #1: No

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

Reviewer #2: Yes

**********

Reviewer #1: The manuscript addresses a timely and important topic: the use of artificial intelligence in women’s reproductive health, with attention to diagnostic trends, methodological gaps, and relevance for low-resource settings. The topic is suitable for PLOS ONE and the manuscript is generally intelligible and written in standard English.

However, I recommend major revision before publication. Because this is a scoping review, the main issue is not experimental replication or inferential statistics, but the transparency and reproducibility of the review methods. The authors should make clear whether the review follows PRISMA-ScR and should provide a completed checklist if applicable. The search strategy should be fully reproducible, including databases searched, search dates, complete search strings, language restrictions, handling of grey literature and preprints, and the full inclusion/exclusion criteria.

The screening and extraction workflow also requires clearer reporting. The authors should specify how many reviewers screened records and extracted data, how disagreements were resolved, and what variables were extracted from each included study. Any descriptive counts, proportions, diagnostic categories, AI model categories, or geographic summaries should be tied to a transparent extraction dataset with clear denominators.

The conclusions should be narrowed where necessary. Claims about “global” trends and translational relevance for low-resource settings must be directly supported by the geographic distribution, study settings, validation quality, and clinical context of the included studies. If the evidence base is concentrated in a limited number of countries or settings, this should be stated clearly and reflected in the conclusions.

Finally, the data underlying the review should be made fully available. For a scoping review, this should include the full list of included studies, extracted study-level variables, coding/category definitions, search strategies, screening results, and the data used to generate all tables, figures, counts, and proportions. If these are not already included as supporting information or in a public repository, they should be provided at revision.

Overall, the manuscript has potential, but publication should depend on stronger methodological reporting, complete data availability, and conclusions that are more tightly linked to the mapped evidence.

Reviewer #2: In the manuscript PONE-D-26-22995 “Artificial Intelligence for Women’s Reproductive Health: A Scoping Review of Global Diagnostic Trends, Methodological Gaps, and a Translational Research Agenda for Low-Resource Settings,” the authors conducted a PRISMA-ScR-guided scoping review of artificial intelligence (AI) and machine learning (ML) applications in women’s reproductive and hormonal health.

The objective and scope of the review are clearly defined and consistent throughout the manuscript. Overall, the manuscript is well written, the methods are described in sufficient detail, and the data appear to have been appropriately managed. The results are generally well presented, although some figures and tables require further clarification. The discussion is consistent with the reported findings, and the data generated in the study are publicly accessible.

The taxonomy developed for this review represents a valuable contribution. In addition, the differences in the types and extent of AI/ML applications across country-income categories are clearly described.

Minor comments

Line 96: Please provide the web address for Rayyan.

Line 237: Please provide the official web address or an appropriate reference for SHAP.

Figure 4: Please include a separate legend explaining the relative values represented by the bubble sizes. In addition, clearly describe how the LMIC feasibility threshold was calculated and how it was used to position the dashed line.

Table 9: Please define the abbreviation “SAARC Avg.”

Figure 5A: What does the orange color represent? Please clarify this in the figure legend.

Figure 5B: Please explain the meaning of the colors and include an appropriate legend.

Table 12 and Figure 7: These appear to present the same data. This duplication is unnecessary; therefore, I suggest removing Table 12.

Line 334: The abbreviation LMIC has already been defined earlier in the manuscript. Please avoid defining it again.

Line 435: Please provide supporting citations and complete references for OpenHIE and WHO SMART Guidelines.

Lines 456–459: Please include at least one supporting citation for these statements.

**********

what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: Yes:  Ulysses Angulo

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

PLoS One. 2026 Sep 25;21(9):e0358557. doi: 10.1371/journal.pone.0358557.r002

Author response to Decision Letter 1


31 Jul 2026

RESPONSE TO ACADEMIC EDITOR'S JOURNAL REQUIREMENTS

Comment 1: "Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming."

Response: We thank the Editor for this guidance and confirm that the manuscript now fully complies with all PLOS ONE style and file naming requirements. We have ensured that all main figures and supporting information files are saved and uploaded individually with the correct naming conventions (e.g., Fig1.tif, S1_File.pdf). Within the manuscript, we corrected all figure citations to remove periods (e.g., "Fig 1"), standardized supporting information citations to the required shorthand (e.g., "S1 File"), updated section headings to match PLOS standards (e.g., "Materials and methods", "Conclusions"), and corrected the equal contribution symbols and corresponding author email format on the title page.

Comment 2: "Please ensure that you have uploaded your PRISMA flowchart as Figure 1."

Response: We confirm that the PRISMA flow diagram is designated as Figure 1 in the manuscript text. Additionally, the corresponding image file has been uploaded separately to the submission portal and named exactly as "Fig1.tif" in strict accordance with PLOS ONE's file naming guidelines.

Comment 3: "We note that Figure 5 in your submission contains [map/satellite] images which may be copyrighted. All PLOS content is published under the Creative Commons Attribution License (CC BY 4.0)... We require you to either (1) present written permission from the copyright holder to publish these figures specifically under the CC BY 4.0 license, or (2) remove the figures from your submission."

Response: We thank the Editor for bringing this copyright issue to our attention. We have removed the copyrighted map/satellite images from our submission and replaced them with a newly created, original figure titled "Geographic Distribution and Key Methodological Gaps in AI-Assisted Medical Imaging Research." As the sole creators of this new figure, we hold full copyright and explicitly consent to publish it under the CC BY 4.0 license. All relevant figure citations, captions, and numbering in the manuscript have been updated to reflect this change.

Comment 4: "Please remove your figures from within your manuscript file, leaving only the individual TIFF/EPS image files, uploaded separately."

Response: We have complied with this requirement. Since our manuscript is prepared using LaTeX, we have retained only the \includegraphics references in the .tex source file (as permitted by PLOS ONE's policy for LaTeX submissions) and have uploaded all seven figures as separate high-resolution TIFF files (≥300 dpi) through the submission system. All figure captions remain in the manuscript file.

Comment 5: "Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly."

Response: We have added a dedicated 'Supporting information' section at the end of our manuscript (after the References) containing complete captions for all supporting files. The section includes: S1 File (PRISMA-ScR Checklist), S2 Table (Risk-of-Bias Scoring Rubric), and S3 Table (LMIC Applicability Categorical Criteria). All in-text citations have been updated to use PLOS ONE's standard format (e.g., 'S1 File', 'S2 Table', 'S3 Table') instead of 'Supplementary File S1' or 'Supplementary Table S2'

Comment 6: "If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited."

Response: We have carefully reviewed and incorporated all citation recommendations from the reviewers. Specifically:

• SHAP (Reviewer #2, Comment 2): Added the official web address and primary citation (Lundberg & Lee, 2017)

• OpenHIE and WHO SMART Guidelines (Reviewer #2, Comment 9): Added complete references with supporting citations

• Supporting citations for Lines 456-459 (Reviewer #2, Comment 10): Added references for AI cost-effectiveness in LMICs, mobile-based AI deployment, and implementation science frameworks

All recommended citations have been evaluated as relevant and have been included in the revised manuscript and reference list.

Reviewer 1

Comment 1: "The authors should make clear whether the review follows PRISMA-ScR and should provide a completed checklist if applicable."

Response: We thank the reviewer for this clarification. We confirm that this scoping review strictly follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. We have explicitly stated our adherence to these guidelines in Section 2 (Materials and methods). Furthermore, as requested, we have provided the fully completed PRISMA-ScR checklist as S1 File in the "Supporting information" section at the end of the manuscript.

Comment 2: "The search strategy should be fully reproducible, including databases searched, search dates, complete search strings, language restrictions, handling of grey literature and preprints, and the full inclusion/exclusion criteria."

Response: We thank the reviewer for this important methodological clarification. We have revised Section 2.2 (Search Strategy) to explicitly state the exact search dates (December 2025 to February 2026), language restrictions (English-only), and our handling of grey literature and preprints (which were explicitly excluded). The complete, reproducible search strings for all five databases are detailed in Table 1 and provided in full in the newly added Supplementary File S4. Supplementary file S4 excel file contains 18 columns which was not possible to include in main manuscript, so we have added this as extra supplementary file . The full inclusion and exclusion criteria remain comprehensively detailed in Section 2.3 (Eligibility Criteria).

Comment 3: "The screening and extraction workflow also requires clearer reporting. The authors should specify how many reviewers screened records and extracted data, how disagreements were resolved, and what variables were extracted from each included study."

Response: We appreciate this feedback and have significantly strengthened the reporting of our screening and extraction workflows in Sections 2.4 and 2.5. We have clarified that the screening workflow involved two reviewers (MMHM and TI) for title/abstract screening and two reviewers (AFS and FS) for full-text screening, with discrepancies resolved through consensus with a third reviewer (TS). Furthermore, we have explicitly listed all extracted variables across six specific domains in Section 2.5 and detailed our rigorous, multi-tier independent verification process: initial extraction by one reviewer (TI), second-stage verification by three reviewers (OF, AFS, and FS), and final third-stage verification by MMHM. We also clarified that a 20% random sample was cross verified, yielding a high inter-rater agreement (Cohen’s κ = 0.92), with any remaining discrepancies resolved by the senior authors (RUI and TS).

Comment 4: "Any descriptive counts, proportions, diagnostic categories, AI model categories, or geographic summaries should be tied to a transparent extraction dataset with clear denominators."

Response: We agree with the reviewer that clear denominators are essential for transparency. We have added a clarifying statement at the beginning of Section 2.7 (Data Synthesis and Reporting) and Section 3 (Results) explicitly defining the denominators. All proportions, percentages, and descriptive counts are now clearly tied to the total denominator of N = 116 included studies, unless a specific subgroup denominator is explicitly stated in the respective tables or text. Furthermore, the full transparent extraction dataset, which allows readers to verify all counts and proportions, is now available in Supplementary File S4.

Comment 5: "The conclusions should be narrowed where necessary. Claims about 'global' trends and translational relevance for low-resource settings must be directly supported by the geographic distribution, study settings, validation quality, and clinical context of the included studies. If the evidence base is concentrated in a limited number of countries or settings, this should be stated clearly and reflected in the conclusions."

Response: We thank the reviewer for this crucial observation. We have carefully revised the Abstract, Introduction, and Conclusion to narrow our claims and ensure they are strictly supported by the mapped evidence. The term "global diagnostic trends" has been changed to "international diagnostic trends across 27 countries." We have explicitly added limitations to the Conclusion regarding the heavy concentration of evidence in East Asia (39.6%) and high-income countries, and the severe underrepresentation of Sub-Saharan Africa (1.7%) and specific South Asian nations (e.g., Nepal, Sri Lanka). Translational claims for LMICs are now strictly qualified by the methodological limitations (e.g., 75.5% lack of external validation) and geographic gaps identified in our review.

Comment 6: "The data underlying the review should be made fully available. For a scoping review, this should include the full list of included studies, extracted study-level variables, coding/category definitions, search strategies, screening results, and the data used to generate all tables, figures, counts, and proportions."

Response: We have fully expanded our Data Availability Statement and updated our Mendeley Data repository to ensure complete transparency. The repository now contains: (1) the full list of 116 included studies with complete extracted study-level variables (Supplementary File S4); (2) the complete, reproducible search strategies for all five databases and the screening log detailing reasons for exclusion at each stage (Supplementary File S5); (3) coding and category definitions for the six taxonomy categories and LMIC feasibility tiers (Supplementary Tables S2 and S3); and (4) the raw data used to generate all tables, figures, counts, and proportions. Comment 7: "Overall, the manuscript has potential, but publication should depend on stronger methodological reporting, complete data availability, and conclusions that are more tightly linked to the mapped evidence."

Comment 7: "Overall, the manuscript has potential, but publication should depend on stronger methodological reporting, complete data availability, and conclusions that are more tightly linked to the mapped evidence."

Response: We thank the reviewer for their constructive evaluation and for recognizing the potential of our manuscript. We have comprehensively addressed all methodological reporting gaps, ensured complete data availability via our updated Mendeley Data repository, and carefully aligned our conclusions with the actual geographic and methodological distribution of the mapped evidence. We believe these revisions have significantly strengthened the manuscript's rigor, transparency, and overall quality.

Reviewer 2

Comment 1: "Line 96: Please provide the web address for Rayyan."

Response: We have added the official Rayyan web address and the primary citation to the reference list. The in-text citation now appears as references [20, 21] at Line 97. The web address is available in the reference entry and also through the citation link. Thank you for this helpful suggestion.

Comment 2: "Line 237: Please provide the official web address or an appropriate reference for SHAP."

Response: We have added the primary citation for SHAP as reference [137] in the revised manuscript.

Comment 3: "Figure 4: Please include a separate legend explaining the relative values represented by the bubble sizes. In addition, clearly describe how the LMIC feasibility threshold was calculated and how it was used to position the dashed line."

Response: We thank the reviewer for this important clarification request. We have revised Figure 4 to include:

1. A separate legend explains that bubble sizes represent the number of studies (with larger bubbles indicating more studies) for each combination of diagnostic category and AI methodology, while bubble colors represent the proportion of studies from LMIC settings.

2. The LMIC feasibility threshold (shown as a dashed line at 15%) was calculated as the 75th percentile of LMIC representation across all diagnostic categories. This threshold was chosen to distinguish categories with relatively higher LMIC representation from those with lower representation. We have validated this through sensitivity analysis using alternative cutoffs (10%, 20%, 25%), which produced consistent results. The dashed line visually separates diagnostic-AI combinations that meet the threshold from those that do not.

Comment 4: "Table 9: Please define the abbreviation 'SAARC Avg.'"

Response: We have defined "SAARC Avg." as "South Asian Association for Regional Cooperation Average" in table footnote.

Comment 5: "Figure 5A: What does the orange color represent? Please clarify this in the figure legend."

Comment 6: "Figure 5B: Please explain the meaning of the colors and include an appropriate legend."

Response for Comments 5 & 6: We thank the reviewer for this query. Please note that the original Figure 5 (which contained map/satellite images) was removed and replaced in accordance with a specific request from the Editorial Office regarding copyright restrictions under the CC BY 4.0 license. It has been replaced with a new, original figure (now designated as Figure [insert new figure number, e.g., 5]) that illustrates the geographic distribution and methodological gaps without using any copyrighted imagery. Consequently, the specific sub-panels 5A and 5B from the original submission are no longer present in the revised manuscript, and the new figure includes a comprehensive, self-explanatory legend.

Comment 7: "Table 12 and Figure 7: These appear to present the same data. This duplication is unnecessary; therefore, I suggest removing Table 12."

Response: We have Removed Table 12 as suggested, table 13 has now in place of table 12.

Comment 8: "Line 334: The abbreviation LMIC has already been defined earlier in the manuscript. Please avoid defining it again."

Response: We have removed the duplicate definition at Line 334. LMIC is now consistently defined only at its first occurrence in tables 2’s footnote . We have reviewed the entire manuscript to eliminate any other redundant definitions.

Comment 9: "Line 435: Please provide supporting citations and complete references for OpenHIE and WHO SMART Guidelines."

Response: We have added complete references for OpenHIE and WHO SMART Guidelines in the revised manuscript. The citations [138] and [139] now reference the OpenHIE architecture framework, and citation [140] references the WHO SMART Guidelines methodology. These references are included in the reference list with full bibliographic details as requested.

Comment 10: "Lines 456–459: Please include at least one supporting citation for these statements."

Response: We have added supporting citations to the statements regarding imaging model performance decay and domain adaptation methods in the revised manuscript (Lines 456-469). The added references [141-144] now support the claims about performance decay in imaging-based models across different scanners and protocols, as well as domain adaptation techniques including histogram matching, CycleGAN-based style transfer, and test-time adaptation approaches demonstrated in comparable imaging tasks such as mammography and chest radiography. These citations provide the necessary evidence for the statements identified by the reviewer.

Attachment

Submitted filename: PONE-D-26-22995_Response_to_Reviewers.docx

pone.0358557.s004.docx (380.8KB, docx)

Decision Letter 1

Kwang-Sig Lee

2 Sep 2026

Artificial intelligence for women’s reproductive health: A scoping review of global diagnostic trends, methodological gaps, and a translational research agenda for low-resource settings

PONE-D-26-22995R1

Dear Dr. Munna,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Kwang-Sig Lee

Academic Editor

PLOS One

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

Reviewer #2: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions??>

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #2: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #2: Yes

**********

Reviewer #2: (No Response)

**********

what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #2: No

**********

Acceptance letter

Kwang-Sig Lee

PONE-D-26-22995R1

PLOS One

Dear Dr. Munna,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS One and supporting open access.

Kind regards,

PLOS One Editorial Office Staff

on behalf of

Professor Kwang-Sig Lee

Academic Editor

PLOS One

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 File. PRISMA-ScR Checklist.

    Completed PRISMA-ScR checklist covering all 20 essential items and 2 optional items.

    (DOCX)

    pone.0358557.s001.docx (10.4KB, docx)
    S1 Table. Risk-of-bias scoring rubric.

    Explicit domain-specific criteria used to classify studies as low, moderate, or high risk of bias (adapted from PROBAST-AI, QUADAS-AI, and TRIPOD-AI).

    (DOCX)

    pone.0358557.s002.docx (8.9KB, docx)
    S2 Table. LMIC applicability categorical criteria.

    Full definitions of the four LMIC-feasibility categories (High, Moderate, Low, Very Low) used to classify studies.

    (DOCX)

    pone.0358557.s003.docx (9.6KB, docx)
    Attachment

    Submitted filename: PONE-D-26-22995_Response_to_Reviewers.docx

    pone.0358557.s004.docx (380.8KB, docx)

    Data Availability Statement

    Yes - all data are fully available without restriction; All data underlying this review are fully available in the Mendeley Data repository (https://doi.org/10.17632/ydx7v73t3k.2). The repository contains: (1) the full list of 116 included studies with citations; (2) the complete extracted study-level variables dataset; (3) coding and category definitions for the six taxonomy categories and LMIC feasibility tiers; (4) the complete, reproducible search strategies for all five databases; (5) the screening results with reasons for exclusion at each stage; and (6) the raw data used to generate all tables, figures, counts, and proportions. The completed PRISMA-ScR checklist (S1 File), risk-of-bias scoring rubric (S2 Table), and LMIC applicability categorical criteria (S2 Table) are also provided as supplementary files.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES