Abstract
Objective
Polycystic ovary syndrome (PCOS) is a leading cause of female infertility, and early pregnancy outcomes in affected women are affected by multiple interactive factors; however, no well-established predictive tools are available. This study aimed to develop and validate machine learning (ML) models for predicting early pregnancy in women with PCOS and compare the predictive performance of different ML algorithms.
Methods
This was a secondary analysis of a multicenter randomized controlled trial including 994 PCOS patients recruited from 21 research centers. Candidate predictors included reproductive history, ultrasound parameters, and serum hormone levels. Missing data were processed using multiple imputation. Variables were initially screened by LASSO regression, and final predictors were determined by univariate and multivariate logistic regression. Three ML models—random forest (RF), support vector machine (SVM), and decision tree (DT)—were constructed. The dataset was randomly split into training (70%) and validation (30%) cohorts. Model performance was evaluated using receiver operating characteristic (ROC) curves, area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1 score.
Results
LASSO regression and logistic regression analyses identified four key predictors: body mass index (BMI), body weight, endometrial thickness, and sex hormone-binding globulin (SHBG). In the validation cohort, the RF model showed the best overall performance, with an AUC of 0.56 (95% CI: 0.493–0.628), sensitivity of 65.1%, and specificity of 47.0%. All models exhibited moderate predictive efficacy with an AUC range of 0.505–0.56. The RF model yielded a PPV of 24.8%, NPV of 83.3%, accuracy of 50.8%, and F1 score of 0.360.
Conclusion
ML-based predictive models for early pregnancy in PCOS patients were successfully established. All models demonstrated moderate discriminative ability, with the RF model showing relatively superior performance. These models may serve as adjunctive tools for clinical decision-making, while the moderate predictive power underscores the multifactorial complexity of early pregnancy in PCOS.
Keywords: Polycystic ovary syndrome, Early pregnancy, Predictive models, Machine learning
Introduction
Polycystic ovary syndrome (PCOS) is one of the most prevalent endocrine and metabolic disorders affecting women of reproductive age, with a global prevalence of 6% to 20% and an increasing trend year by year [1, 2]. Core phenotypic features include clinical or biochemical hyperandrogenism, ovulatory dysfunction, and polycystic ovarian morphology. PCOS is not only a major cause of anovulatory infertility but also predisposes affected women to adverse pregnancy complications such as early miscarriage, gestational diabetes mellitus, and preeclampsia [3, 4]. Even with assisted reproductive technology (ART), pregnancy outcomes vary substantially among PCOS individuals, presenting substantial challenges to personalized reproductive management [5]. Therefore, accurate prediction of early pregnancy probability in PCOS patients is clinically imperative to optimize treatment strategies and improve reproductive outcomes.
Current clinical prediction of pregnancy in PCOS mainly relies on conventional indicators such as age, baseline follicle-stimulating hormone (FSH), and anti-Müllerian hormone (AMH), often combined with traditional statistical methods such as logistic regression to construct predictive models [6, 7]. Nevertheless, these approaches have critical limitations. First, the included predictors are relatively limited and cannot comprehensively capture the heterogeneous endocrine and metabolic characteristics of PCOS. Second, traditional statistical methods are less capable of capturing non-linear relationships and high-dimensional interactions, leading to suboptimal model calibration and limited generalizability in external validation, which fail to meet the requirements of precision reproductive medicine [8, 9].
In recent years, machine learning (ML) techniques have been increasingly applied in reproductive medicine owing to their superior capacity in processing complex high-dimensional data and detecting non-linear associations [10–12]. At the methodological level of model construction, interdisciplinary algorithmic integration provides valuable reference. In one study on lung adenocarcinoma (LUAD), a low-density lipoprotein-associated signature (LAS) was constructed by fusing the CoxBoost and Ridge regression algorithms, which demonstrated substantially superior performance in C-index compared with conventional models [13]. Accumulating evidence indicates that ML algorithms, including random forest and support vector machine, outperform conventional statistical models in ovarian reserve assessment, embryo quality evaluation, and pregnancy outcome prediction, providing a novel technical strategy for precise prediction of pregnancy in PCOS patients [14, 15]. However, studies specifically employing ML to predict early pregnancy in PCOS remain scarce, are mostly limited to single-center small-sample designs, and their generalizability and clinical applicability require further validation [16].
Accordingly, this study integrated multicenter clinical and ART data from PCOS patients. After standardized preprocessing, we developed and validated ML-based prediction models for early pregnancy, with the goal of establishing a quantitative and generalizable predictive tool to support individualized clinical decision-making and provide a methodological reference for predictive modeling in reproductive medicine [17].
Materials and methods
Study design and ethics
This study was a retrospective secondary analysis of a multicenter randomized controlled trial (RCT), designed to develop and validate ML models for predicting early pregnancy in women with PCOS. Data were derived from PCOS patients enrolled at 21 clinical centers between July 2012 and November 2014. The study protocol was approved by the Ethics Committee of the First Affiliated Hospital of Heilongjiang University of Chinese Medicine (Approval No.: 2010 HZYLL-010). The clinical trial was registered at ClinicalTrials.gov (NCT01573858) and the Chinese Clinical Trial Registry (ChiCTRTRC-12002081; registered on April 10, 2012).Inclusion criteria: ①Met the Rotterdam diagnostic criteria for PCOS; ② Aged 20–40 years with infertility and desire for pregnancy; ③Male partner with sperm concentration ≥ 15 × 10⁶/mL, total motility ≥ 40%, or total motile sperm count ≥ 10 million; ④ Agreed to regular intercourse (2–3 times weekly during the intervention period). Exclusion criteria: ①Other endocrine disorders;②Use of hormones or other medications within the previous 3 months;③Pregnancy detected within the previous 6 weeks;④Postpartum period or within 6 weeks after delivery;⑤ Breastfeeding within the previous 6 months; ⑥ Unclear partner smoking status; ⑦ Incomplete laboratory data; ⑧ Loss to follow-up or unavailable pregnancy outcome; ⑨ Refusal to provide written informed consent.
Study population and clinical variables
A total of approximately 1000 PCOS patients were initially enrolled. Associations of anthropometric indices, sex hormones, gynecological ultrasound parameters, and reproductive history with early pregnancy were systematically analyzed. Early pregnancy was defined as the detection of a gestational sac by ultrasound within 12 weeks of gestation. All anthropometric measurements and blood samples were obtained after an overnight fast, and biochemical analyses were conducted under standardized laboratory conditions. Predictor variables comprised demographic characteristics, anthropometric parameters, pregnancy-related history, gynaecological ultrasound indicators, and serum hormone levels. Demographic variables included age. Anthropometric measurements included height and weight, from which body mass index (BMI, weight [kg]/height [m]2) was calculated. Medical history encompassed duration of infertility and prior pregnancy history.Gynaecological ultrasound parameters included ovarian volume, antral follicle count (AFC), endometrial thickness, uterine longitudinal diameter, uterine transverse diameter, uterine anteroposterior diameter, uterine morphology, and the presence of concomitant endometriosis or uterine fibroids. Blood hormone assays measured anti-Müllerian hormone (AMH), follicle-stimulating hormone (FSH), luteinising hormone (LH), testosterone, total testosterone (TT), free testosterone (FT), sex hormone-binding globulin (SHBG), oestradiol (E2), and progesterone (P).Research assistants received standardized training during the project initiation phase to ensure uniformity in anthropometric measurements, ultrasound procedures, and hormone testing protocols, thereby ensuring data consistency.
Data analysis
Prior to data analysis, descriptive statistics were performed for all variables. The normality of continuous variables was assessed using the Shapiro–Wilk test. Normally distributed continuous variables were expressed as mean ± standard deviation, whereas non-normally distributed variables were presented as median (interquartile range, Q1–Q3). Categorical variables were reported as frequencies (percentages). Comparisons between the training and test sets were conducted using the independent samples t-test or the Mann–Whitney U test, as appropriate. A P value < 0.05 was considered statistically significant. The overall proportion of missing data in the study population was approximately 8.6%. Certain variables (e.g., ultrasound-related indicators) exhibited higher rates of missingness, ranging from 16 to 25%. Cases with more than 30% missing data were excluded; ultimately, 6 cases were removed, yielding a final sample size of 994 cases for analysis. The dataset was randomly divided into a training set (70%, n = 697) and a test set (30%, n = 297), with strict separation maintained throughout the analysis to prevent data leakage. Missing data handling was performed exclusively within the training set. For the training set, multiple imputation by chained equations (MICE) was applied to address missing data, with the number of imputations set to m = 5. Continuous variables were imputed using the predictive mean matching (PMM) method. Missing values in the test set were handled using statistical characteristics derived from the training set: continuous variables were imputed using the training set mean, and categorical variables were imputed using the training set mode.Variable selection was performed exclusively in the training set. LASSO logistic regression with tenfold cross-validation was used for preliminary screening to identify candidate predictors with non-zero coefficients. Univariate logistic regression was then applied, and variables with statistical significance or clear clinical relevance were entered into multivariate logistic regression to determine the final predictive factors.Three ML models—random forest (RF), support vector machine (SVM), and decision tree (DT)—were constructed. Model performance was assessed using ROC curves and AUC in both cohorts. In the validation set, sensitivity, specificity, PPV, NPV, accuracy, and F1 score were calculated. Confusion matrices were generated to visualize classification performance. The optimal model was determined by comprehensive comparison of all metrics.Statistical analyses were performed using SPSS 27.0 (IBM Corp., Armonk, NY, USA) and R 4.3.2 (R Foundation for Statistical Computing, Vienna, Austria). A two-sided P < 0.05 was considered statistically significant. The study workflow is summarized in Figs. 1 and 2.
Fig. 1.

Flowchart
Fig. 2.

Variable selection process in Lasso regression. A Illustrates the LASSO coefficient path. The horizontal axis represents?log(?), and the vertical axis shows the regression coefficients of each candidate variable. As the penalty parameter? decreases (i.e.,?log(?) increases), variables progressively enter the model, and the absolute values of their coefficients increase. The numbers above the plot indicate the number of non-zero coefficients (i.e., selected variables) at each corresponding? value. B Presents the results of tenfold cross-validation for optimal? selection. The horizontal axis represents?log(?), and the vertical axis denotes the binomial deviance. Red dots indicate the mean cross-validation error at each? value, while the grey error bars represent ± 1 standard error. The two dashed lines correspond to?_min and?_1se, respectively. Based on the?_1se criterion, the most parsimonious model was selected, ultimately retaining 19 variables for subsequent multivariable analysis
Results
Baseline characteristics of the study participants
The present study enrolled 994 patients with PCOS, of whom 697 were randomly allocated to the training set and 297 to the validation set. Baseline characteristics—including age, anthropometric parameters, medical and obstetric histories, gynecological ultrasound findings, and serum hormone levels—are summarized in Table 1. No significant differences were observed between the training and validation sets in key demographic or clinical variables (all P > 0.05), confirming adequate balance between the groups.
Table 1.
Baseline characteristics of patients in the training and validation cohorts
| Total (N = 994) | Training dataset(N = 697) | Validation dataset (N = 297) | P value | |
|---|---|---|---|---|
| Anthropometric | ||||
| Age (years) | 28(26,30) | 28(26,30) | 27(25.5,30) | 0.1591 |
| Height (cm) | 161 (158, 165) | 161(158, 165) | 160(158, 165) | 0.1359 |
| Weight (kg) | 61(54, 70.1) | 61(54.1, 70.35) | 62(53.75, 70.00) | 0.5036 |
| bmi (kg/m2) | 23.7(21,26.7) | 23.75(20.97,27) | 23.59(21.02,26.54) | 0.7495 |
| Ultrasound imaging parameters | ||||
| Uterine length(cm) | 4.6(4.2,5) | 4.6(4.1,5.06) | 4.5(4.19,4.92) | 0.2328 |
| Uterine width(cm) | 3.7(3.3,4.2) | 3.7(3.3,4.2) | 3.7(3.3,4.2) | 0.8042 |
| Anteroposterior diameter of the uterus(cm) | 3.7(3.3,4.2) | 3.7(3.3,4.2) | 3.7(3.3,4.2) | 0.9605 |
| Endometrial thickness(mm) | 6.1(5,7.725) | 6.3(5,8) | 6.0(5,7.4) | 0.326 |
| Uterine morphology | Normal(980,98.5%) | Normal(688,98.7%) | Normal(292,98.3%) | 0.8522 |
| Abnormal(14,1.4%) | Abnormal(9,1.3%) | Abnormal(5,1.7%) | ||
| Do you have endometriosis? | No(991,99.7%) | No(694,99.6%) | No(297,100%) | 0.5586 |
| Yes(3,0.3%) | Yes(3,0.4%) | Yes(0,0%) | ||
| Do you have uterine fibroids? | No(973,97.9%) | No(683,98.0%) | No(290,97.6%) | 0.9135 |
| Yes(21,2.1%) | Yes(14,2.0%) | Yes(7,2.4%) | ||
| Left ovary volume | 10.1(7.11,13.65) | 10.1(7.245,13.61) | 10.1(6.94,13.915) | 0.7234 |
| Number of follicles in the left ovary | 12(12,12) | 12(12,12) | 12(12,12) | 0.5172 |
| Right ovarian volume | 10.9(8.05,14.5) | 10.84(8.01,14.49) | 11.00(8.11,14.815) | 0.9635 |
| Number of follicles in the right ovary | 12(12,12) | 12(12,12) | 12(12,12) | 0.4846 |
| Left ovary morphology | Normal(819,82.4%) | Normal(575,82.5%) | Normal(232,78.1%) | 0.6142 |
| Abnormal(175,17.6%) | Abnormal(110,15.8%) | Abnormal(65,21.9%) | ||
| Right ovary morphology | Normal(862,86.7%) | Normal(608,87.2%) | Normal(254,85.5%) | 0.5091 |
| Abnormal(132,13.3%) | Abnormal(89,12.7%) | Abnormal(43,14.5%) | ||
| Is there a cyst on the left ovary? | No(979,98.5%) | No(688,98.7%) | No(291,98.0%) | 0.5628 |
| Yes(15,1.5%) | Yes(9,1.3%) | Yes(6,2.0%) | ||
| Is there a cyst on the right ovary? | No(978,98.4%) | No(687,98.6%) | No(291,98.0%) | 0.6921 |
| Yes(16,1.6%) | Yes(10,1.4%) | Yes(6,2.0%) | ||
| Sex hormones | ||||
| LH (mIU/mL) | 9.19(6.088,13.823) | 9.12(5.85,13.65) | 9.54(6.255,14.635) | 0.172 |
| FSH (mIU/mL) | 6.00(5.030,7.052) | 5.97(4.95,7.000) | 6.08(5.13,7.195) | 0.1619 |
| LH/FSH ratio | 1.57(1.040,2.300) | 1.56(1.030,2.280) | 1.57(1.065,2.330) | 0.7841 |
| Estradiol (pmol/L) | 199.30(157.70,265.475) | 199.90(158.30,268.25) | 198.40(157.60,264.40) | 0.746 |
| Progesterone (nmol/L) | 1.730(1.21,2.40) | 1.710(1.22,2.395) | 1.750(1.195,2.43) | 0.8378 |
| Free testosterone (pmol/L) | 2.20(1.669,2.827) | 2.185(1.662,2.800) | 2.233(1.679,2.859) | 0.4802 |
| SHBG (nmol/L) | 33.70(22.000,54.875) | 33.70(21.800,53.500) | 33.90(21.400,56.950) | 0.9244 |
| Total testosterone | 1.590(1.190,2.030) | 1.580(1.190,2.030) | 1.630(1.190,2.055) | 0.6789 |
| AMH | 11.334(7.133,15.683) | 11.206(7.152,15.435) | 11.401(6.916,16.799) | 0.3721 |
| Medical history | ||||
| How long have you been trying to get pregnant?(Month) | 23(12,30) | 22(12,30) | 24(12,30) | 0.5358 |
| Have you ever been diagnosed with infertility? | No(32, 3.2%) | No(19, 2.7%) | No(13, 4.4%) | 0.2487 |
| Yes(962, 96.8%) | Yes(678, 97.3%) | Yes(284, 95.6%) | ||
| Causes of infertility |
PCOS (991, 99.7%) Tubal factors(2, 0.2%) Unknown factors(1, 0.1%) |
PCOS (696, 99.9%) Unknown factors(1, 0.1%) | 1-pcos(297, 100%) | 0.5153 |
| Have you ever received medication for infertility? | No(439, 44.2%) | No(304, 43.6%) | No(135, 45.5%) | 0.6422 |
| Yes(555, 55.8%) | Yes(393, 56.4%) | Yes(162, 54.5%) | ||
| Have you ever received medication treatment for infertility? | No(940, 94.6%) | No(657, 94.3%) | No(283, 95.3%) | 0.6172 |
| Yes(54, 5.4%) | Yes(40, 5.7%) | Yes(14, 4.7%) | ||
| Have you ever had surgery on your abdomen, uterus, cervix, or ovaries? | No(882, 88.7%) | No(608, 87.2%) | No(274, 92.3%) | 0.4657 |
| Yes(112, 11.3%) | Yes(89, 12.8%) | Yes(23, 7.7%) | ||
| Have you used any contraceptive methods in the past? | No(716, 72%) | No(497, 71.3%) | No(219, 73.7%) | 0.481 |
| Yes(278, 28%) | Yes(200, 28.7%) | Yes(78, 26.3%) | ||
| Number of previous pregnancies | 0(630, 63.4%) | 0(433, 62.1%) | 0(197, 66.3%) | 0.0933 |
| 1(269, 27.1%) | 1(186, 26.8%) | 1(82, 27.6%) | ||
| 2(67, 6.7%) | 2(52, 7.5%) | 2(15, 5.1%) | ||
| 3(18, 1.8%) | 3(17, 2.4%) | 3(1, 0.3%) | ||
| 4(7, 0.7%) | 4(6, 0.9%) | 4(1, 0.3%) | ||
| 5(3, 0.3%) | 5(2, 0.3%) | 5(1, 0.3%) | ||
| Number of previous deliveries | 0(939, 94.5%) | 0(939, 94.5%) | 0(283, 95.3%) | 0.456 |
| 1(53, 5.3%) | 1(53, 5.3%) | 1(14, 4.7%) | ||
| 2(2, 0.2%) | 2(2, 0.2%) | |||
| Number of previous spontaneous miscarriages | 0(878, 88.3%) | 0(616, 88.4%) | 0(262, 88.2%) | 0.953 |
| 1(101, 10.2%) | 1(70, 10.0%) | 1(31, 10.4%) | ||
| 2(14, 1.4%) | 2(10, 1.4%) | 2(4, 1.3%) | ||
| 3(1, 0.1%) | 3(1, 0.1%) | |||
| Number of previous induced abortions | 0(781, 78.6%) | 0(537, 77.0%) | 0(244, 82.2%) | 0.0552 |
| 1(166, 16.7%) | 1(121, 17.4%) | 1(45, 15.2%) | ||
| 2(36, 3.6%) | 2(30, 4.3%) | 2(6, 2.0%) | ||
| 3(10, 1.0%) | 3(9, 1.3%) | 3(1, 0.3%) | ||
| 4(1, 0.1%) | 4(1, 0.3%) | |||
| Number of previous induced abortions | 0(976, 98.2%) | 0(682, 97.8%) | 0(294, 99.0%) | 0.3009 |
| 1(18, 1.8%) | 1(15, 2.2%) | 1(3, 1.0%) | ||
| Number of previous preterm births | 0(988, 99.4%) | 0(693, 99.4%) | 0(295, 99.3%) | 1 |
| 1(6, 0.6%) | 1(4, 0.6%) | 1(2, 0.7%) | ||
| Number of previous full-term deliveries | 0(956, 96.2%) | 0(669, 96.0%) | 0(287, 96.6%) | 0.6225 |
| 1(37, 3.7%) | 1(27, 3.9%) | 1(10, 3.4%) | ||
| 2(1, 0.1%) | 2(1, 0.1%) | |||
| Number of previous ectopic pregnancies | 0(976, 98.2%) | 0(684, 98.1%) | 0(292, 98.3%) | 1 |
| 1(18, 1.8%) | 1(13, 1.9%) | 1(5, 1.7%) | ||
Variable screening results
LASSO logistic regression was first applied to screen candidate predictors in the training set. The optimal penalty parameter (λ) was selected by tenfold cross-validation, as illustrated in the coefficient path plot and cross-validation error curve (Fig. 1). As λ increased, the regression coefficients of most variables progressively shrank toward zero. At the optimal λ, a subset of predictors associated with early pregnancy in PCOS patients was retained. The 19 variables selected by LASSO regression were subsequently evaluated by univariate logistic regression in the training set. This analysis revealed that body mass index (BMI), weight, endometrial thickness, and sex hormone-binding globulin (SHBG) were significantly associated with early pregnancy (all P < 0.05). Specifically, BMI, weight, and endometrial thickness showed negative associations, whereas SHBG exhibited a positive association. Age approached statistical significance (P = 0.059). Regression coefficients, odds ratios (ORs), and 95% confidence intervals for these variables are presented in Table 2.
Table 2.
Univariate and multivariate logistic analyses of early pregnancy in the training set
| Variables | Univariate logistic analysis | Multivariate logistic analysis | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| β | S.E | Z | P | OR (95%CI) | β | S.E | Z | P | OR (95%CI) | |
| BMI | −0.073 | 0.023 | −3.119 | 0.002 | 0.929(0.887–0.972) | −0.024 | 0.071 | −0.34 | 0.734 | 0.976(0.848–1.122) |
| Weight | −0.024 | 0.008 | −2.945 | 0.003 | 0.976(0.961–0.992) | −0.014 | 0.025 | −0.586 | 0.558 | 0.986(0.939–1.034) |
| Endometrial thickness | −0.124 | 0.046 | −2.679 | 0.007 | 0.884(0.805–0.965) | −0.122 | 0.049 | −2.507 | 0.012 | 0.885(0.802–0.972) |
| SHBG | 0.007 | 0.003 | 2.483 | 0.013 | 1.007(1.001–1.013) | 0.004 | 0.003 | 1.202 | 0.23 | 1.004(0.997–1.01) |
| Age | −0.053 | 0.028 | −1.888 | 0.059 | 0.948(0.896–1.002) | −0.06 | 0.031 | −1.952 | 0.051 | 0.942(0.887–1) |
| T | −0.215 | 0.147 | −1.456 | 0.145 | 0.807(0.601–1.072) | −0.106 | 0.161 | −0.662 | 0.508 | 0.899(0.653–1.226) |
These variables were then entered into a multivariate logistic regression model. Multivariate analysis identified endometrial thickness as the sole independent predictor of early pregnancy (β = − 0.122, P = 0.012, OR = 0.885, 95% CI: 0.802–0.972). In contrast, BMI, weight, SHBG, age, and testosterone (T) were not statistically significant after adjustment for the other covariates (all P > 0.05). Ultimately, considering both the multivariate findings and clinical relevance, the four variables—BMI, weight, SHBG, and endometrial thickness—were selected for subsequent machine learning model construction.
Construction and performance comparison of machine learning models
Based on the aforementioned variable selection results, four predictors were incorporated to develop three machine learning models: random forest (RF), support vector machine (SVM), and decision tree. These models were used to predict the study outcome. Model performance was evaluated in both the training and validation datasets. The discriminative ability and stability of the models were systematically assessed using receiver operating characteristic (ROC) curves, confusion matrices, and multiple classification performance metrics.
Within the training set, the random forest model demonstrated the best discriminative performance, with an AUC of 0.841 (95% CI: 0.821–0.860), substantially outperforming both the SVM model [AUC = 0.517 (95% CI: 0.488–0.546)] and the decision tree model [AUC = 0.547 (95% CI: 0.523–0.572)]. The random forest model also achieved perfect sensitivity (1.000) in the training set, albeit with relatively lower specificity (0.682), indicating a strong ability to identify positive cases.In contrast, the SVM and decision tree models exhibited high specificity (0.914 and 0.995, respectively) but markedly low sensitivity (0.120 and 0.100, respectively), with overall discriminative performance approaching that of random classification (Table 3).
Table 3.
Summary of model performance metrics
| Training group | |||||||
|---|---|---|---|---|---|---|---|
| Sensitivity | Specificity | AUC | |||||
| Random_Forest | 1 | 0.682 | 0.841(0.821–0.86) | ||||
| SVM | 0.12 | 0.914 | 0.517(0.488–0.546) | ||||
| Decision_Tree | 0.1 | 0.995 | 0.547(0.523–0.572) | ||||
| Validation group | |||||||
|---|---|---|---|---|---|---|---|
| Sensitivity | Specificity | AUC | PPV | NPV | Accuracy | F1_score | |
| Random Forest | 0.651 | 0.47 | 0.56(0.493–0.628) | 0.248 | 0.833 | 0.508 | 0.36 |
| SVM | 0.111 | 0.936 | 0.524(0.481–0.566) | 0.318 | 0.796 | 0.761 | 0.165 |
| Decision Tree | 0.032 | 0.979 | 0.505(0.481–0.529) | 0.286 | 0.79 | 0.778 | 0.057 |
In the validation set, the overall predictive performance of all three models declined substantially. The random forest model achieved an AUC of 0.560 (95% CI: 0.493–0.628), which remained slightly higher than that of the support vector machine model (AUC = 0.524, 95% CI: 0.481–0.566) and the decision tree model (AUC = 0.505, 95% CI: 0.481–0.529); however, this advantage was markedly attenuated compared with that observed in the training set. In the validation set, the random forest model showed a sensitivity of 0.651 and a specificity of 0.470, indicating a moderate ability to identify positive cases but with a relatively high false-positive rate. In contrast, both the SVM and decision tree models continued to exhibit a pattern of high specificity and low sensitivity. Specifically, the SVM model achieved the highest specificity (0.936) but had a sensitivity of only 0.111, whereas the decision tree model showed the lowest sensitivity (0.032).
Further analysis of the confusion matrices and comprehensive performance metrics in the validation set (Fig. 3) showed that, although the random forest model had a relatively modest F1 score (0.360), both its F1 score and negative predictive value (NPV = 0.833) were higher than those of the support vector machine and decision tree models. These findings suggest that the random forest model provided comparatively better overall discrimination of positive cases in the presence of class imbalance.By contrast, although the support vector machine and decision tree models achieved higher accuracies (0.761 and 0.778, respectively), these values were primarily driven by correct classification of negative cases. Their F1 scores were only 0.165 and 0.057, respectively, indicating limited clinical utility.
Fig. 3.

Confusion matrix heatmap for the three models on the test set. The horizontal axis represents the actual outcome (0 = not pregnant, 1 = pregnant), while the vertical axis denotes the predicted outcome (0 = predicted not pregnant, 1 = predicted pregnant). The color intensity reflects the sample count in each cell. The left, middle, and right panels present the confusion matrices for the decision tree, random forest, and support vector machine (SVM) models, respectively.Each cell shows the number of samples corresponding to a specific prediction-outcome combination: the top-left cell indicates true positives (TP), the top-right false positives (FP), the bottom-left false negatives (FN), and the bottom-right true negatives (TN)
The ROC curve comparison (Fig. 4) showed that, in the training set, the random forest model curve was markedly higher than those of the other two models. In the validation set, all three ROC curves approached the diagonal line, indicating limited overall generalizability. Nevertheless, the random forest model retained relatively superior discriminative performance.
Fig. 4.

ROC curves in training group (A) and validation group (B)
Considering the combined results of AUC, sensitivity, F1 score, and ROC curves across both the training and validation sets, the random forest model demonstrated the most favorable overall performance among the three machine learning approaches. Accordingly, it was selected as the final predictive model and was subsequently used for variable importance analysis.
On this basis, variable importance analysis was performed for the random forest model (Fig. 5; Table 4). The results showed that SHBG contributed most substantially to the model, with a mean decrease in Gini index of 40.793, accounting for 30.7% of the total importance. This was followed by body mass index (BMI) (mean decrease in Gini = 33.474, 25.2%) and body weight (30.877, 23.2%). Endometrial thickness also demonstrated notable importance (27.906, 21.0%).These findings suggest that sex hormone–related indicators and BMI play pivotal roles in predicting early pregnancy outcomes in patients with PCOS within the random forest model.
Fig. 5.

Random forest feature importance ranking chart
Table 4.
Random forest feature importance scores
| Variable | Mean Decrease Gini | Relative Importance Percent |
|---|---|---|
| SHBG | 40.793 | 30.7 |
| BMI | 33.474 | 25.2 |
| Weight | 30.877 | 23.2 |
| Endometrial thickness | 27.906 | 21 |
Discussion
The establishment and maintenance of early pregnancy represent not only a critical stage for women with polycystic ovary syndrome (PCOS) to achieve conception, but also a pivotal window that influences subsequent pregnancy outcomes. Therefore, identifying key determinants of pregnancy in PCOS and developing reliable predictive tools are of substantial clinical importance for optimizing individualized reproductive management.
Although the random forest model achieved satisfactory fit in the training set (AUC = 0.841, 95% confidence interval [CI]: 0.821–0.860), the predictive performance of all three machine learning algorithms was notably limited when applied to the independent validation cohort. Specifically, the random forest model yielded an AUC of only 0.560 (95% CI: 0.493–0.628), while the corresponding values for the support vector machine and decision tree models were 0.524 (95% CI: 0.481–0.566) and 0.505 (95% CI: 0.481–0.529), respectively—all approaching the threshold of random classification (AUC = 0.5). Moreover, the positive predictive value (PPV) of the random forest model was merely 24.8%, and its F1 score stood at 0.360. Despite exhibiting relatively high sensitivity (0.651), the model suffered from low specificity (0.470), which generated a substantial burden of false-positive predictions and consequently diminished its clinical applicability. The support vector machine and decision tree models produced even lower F1 scores (0.165 and 0.057, respectively), further underscoring the inherent limitations of these approaches in forecasting early pregnancy outcomes among patients with polycystic ovary syndrome (PCOS).
Obesity and metabolic dysfunction are well-established determinants of reproductive outcomes in polycystic ovary syndrome (PCOS). Previous retrospective cohort analyses have demonstrated that an elevated body mass index (BMI) markedly reduces the likelihood of conception (hazard ratio [HR]: 0.37, 95% confidence interval [CI]: 0.31–0.44), whereas weight reduction confers a substantial improvement in pregnancy outcomes (HR: 1.68) [18]. Mendelian randomisation studies further indicate a causal relationship between genetically predicted body mass index (BMI) and visceral adiposity and the risk of developing polycystic ovary syndrome (PCOS) and associated reproductive disorders (odds ratio [OR] range: 1.150–3.080) [19]. Moreover, anthropometric measures exhibit modest discriminatory capacity in forecasting reproductive outcomes among women with PCOS; among these parameters, BMI demonstrates the greatest utility in predicting ovulation [20]. Although BMI did not retain independent statistical significance in the multivariate model, it exhibited a high ranking in the variable importance assessment derived from the random forest analysis. This observation lends further credence to the prognostic relevance of obesity-related anthropometric parameters in determining pregnancy outcomes.
Beyond metabolic factors, this study further reveals that SHBG and endometrial thickness are strongly linked to pregnancy outcomes, highlighting the central importance of the uterine environment in PCOS pregnancies. Hyperandrogenism may compromise pregnancy maintenance by disrupting endometrial receptivity and metabolic pathways. Prior research demonstrates that hyperandrogenism downregulates the expression of endometrial receptivity markers such as IGFBP-1 and LIF, thereby impairing embryo implantation [21]; Women with polycystic ovary syndrome exhibit persistently elevated androgen levels both prior to conception and throughout gestation [22]. Further research has demonstrated that testosterone, fasting insulin, and blood glucose levels constitute independent predictors of threatened miscarriage [23]. Although T did not attain statistical significance in this study, it exhibited a robust association with early pregnancy outcomes in PCOS patients during Lasso variable selection, implying that it may still possess predictive value within more complex, non-linear modeling frameworks.
Endometrial dysfunction also represents an important mechanism underlying adverse pregnancy outcomes in PCOS. Systematic reviews have demonstrated the presence of endometrial insulin resistance and alterations in the inflammatory microenvironment in patients with PCOS [24]. In this study, endometrial thickness was included as a multifactorial independent variable, reflecting endometrial receptivity to some extent. Its incorporation into the predictive model underscores its importance in the early implantation and development of fertilised ova. Furthermore, optimized endometrial preparation protocols have been shown to improve live birth rates (63.1% vs 56.8%) [25], further emphasizing the critical role of the endometrial environment in achieving early pregnancy in patients with polycystic ovary syndrome.
In addition, although this study focused on clinical phenotype–based prediction, the potential mechanisms underlying the high-contribution features identified by the model (e.g., sex hormone–binding globulin) warrant further investigation. In the tumor microenvironment, low-density lipoprotein (LDL)–related genes have been shown to influence prognosis by reshaping intercellular communication and modulating immune surveillance [13]. These findings suggest that future PCOS prediction models may benefit from investigating intrinsic differences in the “metabolic microenvironment,” which may, in turn, inform the individualized optimization of assisted reproduction strategies to maximize clinical benefit. In the field of machine learning, previous studies have reported that XGBoost or random forest models can achieve AUC values ranging from 0.782 to 0.901 for predicting live birth or adverse pregnancy outcomes [16, 26]. Additionally, endometriosis and PCOS were included as risk factors for cardiovascular disease in the Cox proportional hazards model [27]. Although endometriosis was recorded only as part of the medical history in this study and was not included in the predictive model, this indirectly supports the reliability of the underlying data. Consequently, these findings further highlight the advantages of machine learning in assessing pregnancy risk in PCOS.The random forest model demonstrated satisfactory performance in the training set but showed a decline in AUC in the validation set, suggesting a potential risk of overfitting, possibly related to sample size and variable dimensionality. Nevertheless, its high sensitivity indicates potential utility for screening high-risk populations.
Additionally, pregnancy outcomes in PCOS are influenced by peripregnancy metabolic and inflammatory states. Studies have shown that women with PCOS exhibit metabolic abnormalities and elevated inflammatory levels during pregnancy [20, 28], leading to an increased overall risk of adverse pregnancy outcomes [23]. Lifestyle interventions and weight loss can improve insulin resistance and increase pregnancy rates [29]. The exclusion of inflammation and treatment variables from the study model may explain the moderate AUC in the validation set; however, BMI and body weight remained the strongest predictors, further underscoring the central role of metabolic status in PCOS pregnancy outcomes.
The suboptimal predictive performance of the models in this study may be attributable to multiple factors. Although a multicenter retrospective cohort design was employed, the relatively low incidence of the positive outcome (early pregnancy) resulted in pronounced class imbalance, which weakened the models’ ability to identify positive cases in the independent validation set, as reflected by generally low PPV values. From the machine learning algorithms’ perspective, although the random forest, support vector machine, and decision tree models achieved good fitting performance in the training set, their performance declined markedly in the validation set, suggesting substantial overfitting risk with these conventional approaches. Moreover, these algorithms have limited capacity to capture complex nonlinear interactions among features. In this study, the number of predictive variables was relatively small, and feature importance analysis highlighted only four major contributors—SHBG, BMI, body weight, and endometrial thickness—thus failing to fully characterize the multifactorial and complex mechanisms underlying early pregnancy outcomes in patients with PCOS. In addition, only three traditional machine learning models (random forest, support vector machine, and decision tree) were developed, without exploration of more advanced algorithmic architectures, thereby further constraining model potential. Potential issues in implementation details, including data preprocessing, feature engineering, and hyperparameter tuning, may also have introduced subtle biases or inconsistencies, ultimately affecting model generalizability. Furthermore, the inherent heterogeneity of multicenter retrospective data, along with potential missing information, may have further compromised overall model robustness. Although feature importance analysis identified these four variables as the primary contributors, neither their individual use nor their combination was sufficient to construct a high-accuracy clinical prediction tool.
In summary, the machine learning models developed in this study are not yet suitable for clinical translation, and their limited predictive performance restricts their practical application in the management of early pregnancy in patients with PCOS. Nevertheless, the models provide some reference value, as the findings highlight body composition and the intrauterine environment as key determinants of pregnancy in PCOS. Moreover, this study was based on a secondary analysis of multicenter randomized controlled trial (RCT) data, which ensured relatively high data quality and consistency. Standardized procedures for missing-value handling and variable screening minimized bias. The three machine learning models were systematically compared and evaluated in an independent validation set, and key factors associated with early pregnancy in PCOS were ultimately identified through converging evidence from variable importance analysis and regression modeling.Compared with previous studies, which were predominantly single-center or single-model investigations [16, 26], this research demonstrates a degree of innovation in data sources and model comparisons.
However, as a secondary analysis, the scope of study variables was inherently constrained by the original research design, and certain treatment regimens were excluded from the model analysis. In addition, the absence of independent external cohort validation requires further confirmation of the model’s generalizability. Furthermore, the moderate AUC in the validation set suggests that pregnancy outcomes in PCOS are jointly regulated by multidimensional factors, limiting the predictive capability of baseline indicators alone. Finally, the relatively limited interpretability of the random forest model represents a common challenge in the clinical application of machine learning models.
Future studies aimed at predicting pregnancy outcomes in PCOS may incorporate indicators of inflammatory and metabolic pathways to enhance the biological interpretability of predictive models. These models could subsequently be integrated with treatment protocols and variables pertaining to assisted reproductive strategies, thereby enabling the construction of dynamic prediction frameworks. Furthermore, the inclusion of additional machine learning approaches for comparative evaluation—such as neural networks, Cox proportional hazards models, LightGBM, and Naive Bayes models—may facilitate the clinical translation of model findings. Coupled with explainable AI methodologies to delineate variable interaction pathways, this integrated strategy holds considerable promise for advancing the clinical utility of predictive models in this population.
Authors’ contributions
Y.L. and H.Y. conceptualized the methodology. Y.L. and J.Z. wrote the original draft of the manuscript. J.P. created Fig. 1. Z.G. and J.N. reviewed and edited the manuscript. X.W. provided funding acquisition and project administration, and was responsible for supervision and validation.
Funding
This work is supported by (1) The National key R&D Program of China (2019YFC1709500); (2) The National Collaboration Project of Critical Illness by Integrating Chinese Medicine and Western Medicine; (3) Project of Heilongjiang Province Innovation Team “TouYan” (LH2019H046); (4) Heilongjiang Provincial Clinical Research Centre for Ovary Diseases (LC2020R009); (5) Traditional Chinese Medicine Research Project of Heilongjiang Administration of Traditional Chinese Medicine (ZHY2022-124); (6) The project of Evidence-based capacity in Traditional Chinese Medicine (TCM Sci-Tech Internal Letter [2023] No. 24). (7) New Round of Provincial "Double First-Class" Construction—Traditional Chinese Medicine (15041250006).
Data availability
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yuxiao Liu and Hong Yu contributed equally to this work.
References
- 1.Neven ACH, Forslund M, Ranasinha S, Sethi P, Dhungana RR, Mousa A, et al. Prevalence of polycystic ovary syndrome: a global and regional systematic review and meta-analysis[J]. Hum Reprod Update. 2026. [DOI] [PubMed]
- 2.Hoeger KM, Dokras A, Piltonen T. Update on PCOS: consequences, challenges, and guiding treatment. J Clin Endocrinol Metab. 2021;106:e1071–83. [DOI] [PubMed] [Google Scholar]
- 3.Xie N, Zhao W. Adverse pregnancy and perinatal outcomes in women with polycystic ovary syndrome undergoing assisted reproductive technology: a systematic review and meta-analysis. Front Med Lausanne. 2025;12:1656389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Bahri Khomami M, Shorakae S, Hashemi S, Harrison CL, Piltonen TT, Romualdi D, et al. Systematic review and meta-analysis of pregnancy outcomes in women with polycystic ovary syndrome[J]. Nat Commun. 2024;15:5591. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Helvaci N, Yildiz BO. The impact of ageing and menopause in women with polycystic ovary syndrome. Clin Endocrinol. 2022;97:371–82. [DOI] [PubMed] [Google Scholar]
- 6.Guixue G, Yifu P, Yuan G, Xialei L, Fan S, Qian S, et al. Progress of the application clinical prediction model in polycystic ovary syndrome[J]. J Ovarian Res. 2023;16:230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Liu Y, Gao J, Ge H, Feng J, Wang Y, Wu X. Development and validation of a LASSO logistic regression based nomogram for predicting live births in women with polycystic ovary syndrome: a retrospective cohort study[J]. Front Endocrinol Lausanne. 2025;16:1525823. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Wang J, Zhou W, Song Z, Ni T, Zhang Q, Chen ZJ, et al. Does the risk of embryo abnormality increase in PCOS women? A secondary analysis of a multicenter, randomized controlled trial [J]. J Clin Endocrinol Metabolism. 2023;108:e249–57. [DOI] [PubMed] [Google Scholar]
- 9.Sufriyana H, Husnayain A, Chen YL, Kuo CY, Singh O, Yeh TY, et al. Comparison of multivariable logistic regression and other machine learning algorithms for prognostic prediction studies in pregnancy care: systematic review and meta-analysis. JMIR Med Inform. 2020;8:e16503. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xin X, Wu S, Xu H, Ma Y, Bao N, Gao M, et al. Non-invasive prediction of human embryonic ploidy using artificial intelligence: a systematic review and meta-analysis[J]. EClinicalMedicine. 2024;77:102897. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Curchoe CL, Letterie GS, Quaas AM. Unlocking the potential of artificial intelligence (AI) in reproductive medicine: the JARG collection on assisted reproductive technology (ART) and machine learning. J Assist Reprod Genet. 2023;40:2079–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Ferrand T, Boulant J, He C, Chambost J, Jacques C, Pena CA, et al. Predicting the number of oocytes retrieved from controlled ovarian hyperstimulation with machine learning. Hum Reprod. 2023;38:1918–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zhang P, Wu X, Wang D, Zhang M, Zhang B, Zhang Z. Unraveling the role of low-density lipoprotein-related genes in lung adenocarcinoma: insights into tumor microenvironment and clinical prognosis. Environ Toxicol. 2024;39:4479–95. [DOI] [PubMed] [Google Scholar]
- 14.Li L, Cui X, Yang J, Wu X, Zhao G. Using feature optimization and LightGBM algorithm to predict the clinical pregnancy outcomes after in vitro fertilization. Front Endocrinol. 2023;14:1305473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Wang CW, Kuo CY, Chen CH, Hsieh YH, Su EC. Predicting clinical pregnancy using clinical features and machine learning algorithms in in vitro fertilization. PLoS One. 2022;17:e0267554. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Mogos R, Gheorghe L, Carauleanu A, Vasilache IA, Munteanu IV, Mogos S, et al. Predicting unfavorable Pregnancy Outcomes in Polycystic Ovary Syndrome (PCOS) patients using machine learning algorithms [J]. Medicina (Kaunas, Lithuania). 2024;60. [DOI] [PMC free article] [PubMed]
- 17.Zad Z, Jiang VS, Wolf AT, Wang T, Cheng JJ, Paschalidis IC, et al. Predicting polycystic ovary syndrome with machine learning algorithms from electronic health records. Front Endocrinol. 2024;15:1298628. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Haase CL, Varbo A, Laursen PN, Schnecke V, Balen AH. Association between body mass index, weight loss and the chance of pregnancy in women with polycystic ovary syndrome and overweight or obesity: a retrospective cohort study in the UK. Hum Reprod (Oxford, England). 2023;38:471–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Venkatesh SS, Ferreira T, Benonisdottir S, Rahmioglu N, Becker CM, Granne I, et al. Obesity and risk of female reproductive conditions: a Mendelian randomisation study [J]. PLoS Med. 2022;19:e1003679. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Xia Q, Wu Q, Feng J, He H, Cai W, Li J, et al. The discriminatory capability of anthropometric measures in predicting reproductive outcomes in Chinese women with PCOS[J]. J Ovarian Res. 2024;17:186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wei SY, Zhang JL, Guan HQ, Cai JJ, Jiang XF, Wang H, et al. High androgen level during controlled ovarian stimulation cycle impairs endometrial receptivity in PCOS patients[J]. Sci Rep. 2024;14:23100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Peigné M, Simon V, Pigny P, Mimouni NEH, Martin C, Dewailly D, et al. Changes in circulating forms of anti-Muüllerian hormone and androgens in women with and without PCOS: a systematic longitudinal study throughout pregnancy[J]. Hum Reprod (Oxford, England). 2023;38:938–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ma L, Zhou Q. Predictors of miscarriage in polycystic ovary syndrome patients with threatened abortion: development and validation of a nomogram model[J]. Front Endocrinol. 2025;16:1689878. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Palomba S, Piltonen TT, Giudice LC. Endometrial function in women with polycystic ovary syndrome: a comprehensive review[J]. Hum Reprod Update. 2021;27:584–618. [DOI] [PubMed] [Google Scholar]
- 25.Wang Y, Hu WH, Wan Q, Li T, Qian Y, Chen MX, et al. Effect of artificial cycle with or without GnRH-a pretreatment on pregnancy and neonatal outcomes in women with PCOS after frozen embryo transfer: a propensity score matching study[J]. Reprod Biol Endocrinol. 2022;20:56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Zhu S, Huang Z, Chen X, Jiang W, Zhou Y, Zheng B, et al. Construction and evaluation of machine learning-based prediction model for live birth following fresh embryo transfer in IVF/ICSI patients with polycystic ovary syndrome[J]. J Ovarian Res. 2025;18:70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Powell MJ, Fuller S, Gunderson EP, Benz CC. Reduced cardiovascular risks in women with endometriosis or polycystic ovary syndrome carrying a common functional IGF1R variant[J]. Hum Reprod (Oxford, England). 2022;37:1083–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Suha SA, Islam MN. An extended machine learning technique for polycystic ovary syndrome detection using ovary ultrasound image[J]. Sci Rep. 2022;12:17123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Moro F, Giudice MT, Ciancia M, Zace D, Baldassari G, Vagni M, et al. Application of artificial intelligence to ultrasound imaging for benign gynecological disorders: systematic review[J]. Ultrasound Obstet Gynecol. 2025;65:295–302. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
