Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Jul 20;118(7):e70095. doi: 10.1002/bdr2.70095

Development and Evaluation of an Adaptive Penguin–Improved LSTM Model Integrating Natural Language Processing for Early Prediction of Pregnancy Syndrome Risks

Yanqi Wu 1,
PMCID: PMC13385461  PMID: 42477301

ABSTRACT

Background

Early prediction of pregnancy syndromes such as preeclampsia, gestational diabetes mellitus (GDM), and preterm birth is critical for improving maternal and fetal outcomes. Traditional risk assessment tools rely primarily on structured clinical data and often fail to fully utilize the rich unstructured textual information contained in electronic health records (EHRs).

Objective

This study aimed to develop and evaluate an Adaptive Penguin–Improved Long Short‐Term Memory (AP‐ILSTM) model that integrates natural language processing (NLP) of clinical narratives with structured longitudinal data for predicting the risk of preeclampsia, GDM, and preterm birth.

Methods

This retrospective observational study used de‐identified EHR data from 1248 consecutive pregnancies at a single regional hospital in China between April 2024 and April 2025. Structured clinical variables and unstructured free‐text clinical notes were preprocessed using standard NLP techniques, including tokenization, TF‐IDF weighting, and word embeddings. An Improved Long Short‐Term Memory (ILSTM) network with dropout regularization was optimized using the Adaptive Penguin algorithm with Gaussian exploration. The dataset was stratified into training (70%), validation (15%), and independent test (15%) sets. Model performance was assessed using accuracy, sensitivity (recall), positive predictive value (precision), F1‐score, specificity, negative predictive value, ROC‐AUC, and mean absolute error, and compared against baseline machine learning models.

Results

On the independent test set (n = 188), the AP‐ILSTM model achieved an accuracy of 98.85%, sensitivity of 98.69%, positive predictive value of 98.75%, F1‐score of 98.72%, and ROC‐AUC of 0.993, with a mean absolute error of 4.05. It outperformed all baseline classifiers (XGBoost, CatBoost, KNN) and regression models. Clinical narrative embeddings ranked as the second most important predictor after blood pressure trends.

Conclusions

The AP‐ILSTM model, by effectively combining structured data with NLP‐derived features from clinical narratives and employing adaptive optimization, demonstrated superior performance for early multi‐syndrome pregnancy risk prediction. This multimodal temporal approach shows strong potential to support clinical decision‐making, although external validation and prospective implementation studies are required before clinical adoption.

Keywords: adaptive penguin search‐driven intelligent long short‐term memory (AP‐ILSTM), evaluation model, natural language processing (NLP), pregnancy syndrome risk

1. Introduction

Pregnancy syndromes such as preeclampsia, gestational diabetes mellitus (GDM), and preterm birth represent major contributors to maternal and fetal morbidity and mortality worldwide (Alfaki Ahmed et al. 2025; Janeline et al. 2026). Early identification of at‐risk pregnancies is essential for timely intervention, personalized care, and improved outcomes. Traditional risk assessment tools, which primarily rely on structured clinical data (e.g., maternal age, blood pressure, and medical history), often fail to capture subtle indicators embedded in unstructured textual sources such as electronic health records (EHRs), clinical notes, and patient narratives (Shabani et al. 2025; Soler et al. 2026).

A growing body of evidence supports the integration of natural language processing (NLP) and machine learning (ML) techniques in maternal health (Mehrnoush et al. 2026; Staff et al. 2024). Studies have demonstrated the value of ML models including logistic regression, XGBoost, random forests, and deep learning approaches for predicting GDM, pregnancy‐induced hypertension, postpartum depression, and adverse pregnancy outcomes in systemic lupus erythematosus. NLP methods enable extraction of semantic information from free‐text data, complementing structured variables and uncovering risk signals that conventional approaches may miss (Irfan et al. 2025; Kwok et al. 2024; Khan et al. 2022).

Despite these advances, significant gaps remain. Many existing models focus on single outcomes, rely predominantly on structured data, or lack robust optimization for temporal dependencies in sequential pregnancy records. Few studies have systematically integrated NLP‐derived features from diverse EHR text with advanced recurrent architectures and adaptive hyperparameter optimization for multi‐syndrome risk prediction (Correia et al. 2025; Coupland et al. 2025; Borji et al. 2025).

This study addresses these gaps by developing and evaluating an NLP‐enhanced Adaptive Penguin‐driven Intelligent Long Short‐Term Memory (AP‐ILSTM) model for early risk prediction of key pregnancy syndromes (preeclampsia, GDM, and preterm birth). The model combines TF‐IDF and word embeddings for feature extraction from unstructured text with ILSTM networks optimized via adaptive penguin search, leveraging both structured and unstructured EHR data. The primary objective is to assess whether this hybrid approach improves predictive performance over baseline ML methods, with the ultimate aim of supporting clinical decision‐making for early detection and intervention.

2. Method and Materials

The study employed a retrospective observational design using de‐identified electronic health record (EHR) data to develop and evaluate a natural language processing (NLP)–enhanced machine learning model for early prediction of pregnancy syndrome risk. All analyses were conducted in Python (version 3.10) using TensorFlow/Keras, scikit‐learn, XGBoost, and CatBoost libraries. The overall pipeline consisted of data collection and preprocessing, feature engineering from both structured and unstructured sources, model development with adaptive hyperparameter optimization, and rigorous performance evaluation against baseline methods. Detailed mathematical formulations, algorithm pseudocode, flowcharts, and extended preprocessing steps are provided in the Supporting Information.

2.1. Data Source and Collection

The primary data source was the integrated electronic health record (EHR) system of Shijiazhuang Hetaiheng Hospital, a regional healthcare institution in Shijiazhuang, Hebei Province, China, providing comprehensive maternal and obstetric services including routine antenatal examinations, high‐risk pregnancy management, inpatient care, laboratory testing, diagnostic imaging, and deliveries. The hospital maintains a single, hospital‐wide EHR platform that captures both structured clinical data fields and unstructured free‐text documentation generated during routine clinical practice (Figure 1).

FIGURE 1.

FIGURE 1

Provides the general layout of the outline.

This was a private, institution‐specific clinical database; no public datasets (e.g., MIMIC, eICU, or national claims repositories) were used. All records were extracted retrospectively and fully de‐identified prior to analysis by removing names, national identification numbers, telephone numbers, residential addresses, and any other protected health information in accordance with institutional data governance policies. No external validation studies or data‐quality audits specifically evaluating the completeness, accuracy, or coding consistency of the Shijiazhuang Hetaiheng Hospital EHR system have been published to date. Consequently, documentation practices, note‐writing conventions (potentially bilingual Chinese–English medical terminology), and local clinical coding habits may influence feature distributions and model generalizability; this is acknowledged as a limitation.

The study period encompassed all pregnancies with antenatal or obstetric encounters recorded between April 2024 and April 2025 (13‐month window). The source population comprised every pregnant individual who attended antenatal clinics or received inpatient obstetric care at the hospital during this interval, regardless of baseline risk status. This included both low‐risk and high‐risk pregnancies, singleton and multiple gestations, and individuals with one or more pregnancies during the study window. Maternal age ranged from 18 to 45 years. No restriction was placed on insurance status; the cohort reflects the mixed public/private‐payor population typical of regional hospitals in Hebei Province (Figure 2).

FIGURE 2.

FIGURE 2

End‐to‐end workflow of the study design, data processing, model development, and evaluation pipeline for early prediction of major pregnancy syndromes using multimodal EHR data.

The figure illustrates the complete analytical framework, beginning with retrospective extraction of structured clinical variables and unstructured clinical narratives from the hospital electronic health record (EHR) system, followed by rigorous inclusion/exclusion filtering and construction of the final analytic cohort. The pipeline proceeds through structured and text‐based preprocessing, feature engineering (including TF‐IDF and word embeddings), and multimodal feature fusion while enforcing strict temporal constraints to prevent information leakage. The integrated feature representation is then used to train an Adaptive Penguin–Improved Long Short‐Term Memory (AP‐ILSTM) model with automated hyperparameter optimization. Model performance is assessed on an independent held‐out test set and compared against multiple machine learning baselines using comprehensive classification and regression metrics. The workflow highlights the end‐to‐end translational pathway from real‐world clinical data to predictive modeling for early risk stratification of pregnancy‐related syndromes.

2.2. Study Population, Inclusion/Exclusion Criteria, and Handling of Multiple Pregnancies

Eligible participants met the following inclusion criteria: (Alfaki Ahmed et al. 2025) documented confirmation of pregnancy within the EHR; (Janeline et al. 2026) availability of complete antenatal clinical records containing both structured variables and corresponding narrative medical notes; and (Shabani et al. 2025) availability of final pregnancy outcome information (including diagnosis of preeclampsia, gestational diabetes mellitus [GDM], preterm birth, or none of these syndromes) to enable supervised learning.

Exclusion criteria were: (Alfaki Ahmed et al. 2025) pregnancies with incomplete medical records or missing outcome data; (Janeline et al. 2026) duplicate records that could not be reconciled into a single longitudinal episode; (Shabani et al. 2025) inconsistent or contradictory clinical documentation; and (Soler et al. 2026) records with excessive missing predictor variables that could not be reasonably imputed.

When multiple clinical encounters existed for a single pregnancy, all encounters were chronologically merged into one longitudinal pregnancy episode to preserve the temporal evolution of maternal health. If the same woman contributed more than one pregnancy during the study period, each pregnancy was treated as an independent episode for feature extraction and outcome labeling. However, this introduces a potential risk of information leakage across pregnancies from the same individual if woman‐level dependence is not accounted for during model validation. Although the current analysis used pregnancy‐level stratified random splitting, future work should implement woman‐level grouping via GroupKFold or stratified group k‐fold cross‐validation to mitigate this risk; this issue is noted as a methodological limitation.

After applying inclusion/exclusion criteria, the final analytic cohort consisted of 1248 pregnancies (873 training [70%], 187 validations [15%], 188 independent test [15%]). The mean maternal age was 29.8 ± 5.4 years (range 18–45); 54.6% were nulliparous and 45.4% multiparous; 96.4% were singleton gestations and 3.6% multiple gestations. Adverse outcomes occurred in 29.4% of pregnancies (preeclampsia 7.9%, GDM 12.5%, preterm birth 9.0%), with 70.6% completing pregnancy without any of the three target syndromes.

2.3. Predictor Variables (Structured and Unstructured)

Predictor variables were restricted exclusively to information recorded before the occurrence of the pregnancy outcome to prevent information leakage and reverse causality (detailed in Section 2.5).

Structured predictors included: maternal age (analyzed as a continuous variable to preserve nonlinear relationships), gravidity, parity, previous pregnancy history and complications, maternal comorbidities (documented in the EHR; coding system not explicitly ICD‐9/ICD‐10 but reflecting local clinical terminology and problem lists), blood pressure measurements (serial), laboratory results (e.g., glucose screening, hemoglobin, renal/hepatic function), medication history, body mass index (BMI) where available, and other routine antenatal examination findings. A total of approximately 25–30 core structured features were extracted per pregnancy episode (exact count varied slightly with availability), supplemented by derived variables such as blood pressure trajectory slopes. Continuous variables were standardized (z‐score) prior to modeling. Missing numerical values were imputed using median or clinically appropriate methods when missingness was < 20% and clinically plausible; otherwise, records were excluded.

Unstructured predictors consisted of all free‐text clinical narratives generated before the outcome event, including physician progress notes, antenatal consultation notes, admission records, nursing documentation, ultrasound interpretations, and discharge summaries (when generated antenatally). These narratives frequently contained symptom descriptions, clinician observations, suspected diagnoses, and treatment responses not fully captured in structured fields.

Natural language preprocessing (detailed in Supporting Information) comprised text cleaning, lower‐casing, tokenization, stop‐word removal, lemmatization, medical terminology normalization, and vectorization via TF‐IDF combined with distributed word embeddings (Word2Vec or similar). The resulting high‐dimensional textual feature vectors were concatenated with the structured feature vector to form a unified multimodal representation for each pregnancy episode. No post‐outcome notes or documentation generated after diagnosis or delivery were included for any outcome.

2.4. Outcome Definition

The primary outcome was the occurrence of any major pregnancy‐related syndrome, operationalized as a binary classification task (developed at least one of: preeclampsia, GDM, or preterm birth vs. none). Secondary analyses examined the three syndromes individually where prevalence permitted.

Outcomes were ascertained from the final clinical diagnosis entered by obstetric specialists in the EHR. Preeclampsia was defined according to standard international obstetric criteria (new‐onset hypertension after 20 weeks' gestation with proteinuria or other end‐organ dysfunction) as documented in the record. GDM was identified by clinician‐confirmed diagnosis supported by routine prenatal glucose screening (typically 75 g oral glucose tolerance test at 24–28 weeks). Preterm birth was defined as delivery before 37 completed weeks of gestation; all predictor information was truncated before the actual delivery date. Because spontaneous preterm labor can begin as early as 20 weeks, the model used only antenatal data collected prior to delivery to avoid reverse causality.

The composite “any major syndrome” endpoint was chosen because early identification of pregnancies at elevated risk for any of these complications enables timely, syndrome‐agnostic preventive interventions (e.g., low‐dose aspirin, enhanced monitoring, glucose management). Individual syndrome prevalences and model performance on each are reported in the Results.

2.5. Temporality and Prevention of Information Leakage

Strict temporal separation was enforced between all predictors (structured variables and clinical narratives) and the outcome event. Only data documented before the clinical diagnosis of preeclampsia or GDM, or before delivery in the case of preterm birth, were used for feature extraction. Clinical observations recorded after diagnosis or delivery were systematically excluded. Longitudinal records were ordered chronologically to respect the natural sequence of events. This design eliminates reverse causality, prevents the model from learning from information that would not be available in prospective clinical use, and closely simulates real‐world deployment for early risk stratification. Pregnancies resulting in preterm birth before 20 weeks were rare and handled consistently with the pre‐delivery data restriction.

2.6. Data Preprocessing, Feature Extraction, and Model Development (Condensed)

Standard preprocessing (cleaning, normalization, tokenization, vectorization) was applied to both structured and textual data (full pipeline in Supporting Informations). Feature extraction combined TF‐IDF weighting (to highlight clinically informative terms) with word embeddings (to capture semantic relationships among medical concepts). The core predictive model was an Adaptive Penguin–Improved Long Short‐Term Memory (AP‐ILSTM) network. The ILSTM component (standard LSTM augmented with dropout regularization) captures temporal dependencies across serial antenatal visits. The Adaptive Penguin optimizer (with Gaussian exploration) performs automated hyperparameter search over learning rate, hidden units, number of layers, batch size, and dropout probability.

Detailed architectural specifications, mathematical equations governing the LSTM gates and Adaptive Penguin search procedure, pseudocode, and flowcharts are provided in the Supporting Information to maintain focus on evaluation and comparison with baseline methods. In brief, the model ingests the concatenated multimodal feature vector (or sequence of visit‐level vectors) and outputs a probability of developing any major pregnancy syndrome. Inverse class‐frequency weighting was applied in the loss function to address moderate class imbalance (outcome prevalence 7.9%–12.5%).

2.7. Baseline Models and Evaluation Metrics (In Methods)

To ensure fair comparison, six baseline models were implemented on identical features and data splits: regression baselines (Voting Regression, Linear Regression, Gradient Boosting Regression) for MAE comparison and classification baselines (K‐Nearest Neighbors, XGBoost, CatBoost). All baselines used scikit‐learn/XGBoost/CatBoost implementations with default or grid‐searched hyperparameters.

Model performance was evaluated using the following metrics, selected because the target outcomes are relatively uncommon (prevalence 7.9%–12.5%) and because clinical decision‐making prioritizes both the ability to detect high‐risk pregnancies (high sensitivity/recall to minimize missed cases) and the reliability of positive predictions (high positive predictive value/precision to limit unnecessary interventions and patient anxiety).

  • Accuracy: proportion of correct classifications.

  • Sensitivity (Recall): proportion of actual positive cases correctly identified.

  • Positive Predictive Value (PPV, Precision): proportion of predicted positives that are true positives.

  • F1‐score: harmonic mean of sensitivity and PPV, balancing the two.

  • Specificity and Negative Predictive Value (NPV): complementary metrics for low‐risk identification.

  • ROC‐AUC: overall discriminative ability across thresholds (added in this revision to provide threshold‐independent assessment recommended for clinical prediction models).

  • Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE): regression error metrics for comparison with baseline regressors.

These metrics collectively address the recall–precision trade‐off important in imbalanced medical screening contexts while also allowing comparison with prior literature. Full mathematical definitions and formulas are provided in the Supporting Information. No single metric was used in isolation; emphasis was placed on F1‐score, ROC‐AUC, and balanced accuracy given class imbalance.

2.8. Data Splitting, Validation Strategy, and Class Imbalance Handling

The dataset (n = 1248 pregnancies) was randomly partitioned using stratified sampling (preserving outcome prevalence) into 70% training, 15% validation, and 15% test sets. Adaptive Penguin hyperparameter optimization was performed on the training + validation data with internal 5‐fold cross‐validation. Final performance is reported exclusively on the independent held‐out test set. Inverse class‐frequency weights were used in the ILSTM loss; no synthetic oversampling (e.g., SMOTE) was applied in order to preserve the fidelity of real clinical data distributions.

2.9. Baseline Model Implementation Details

All baseline models were trained and evaluated under identical conditions to the AP‐ILSTM (same feature matrix, same stratified splits). Hyperparameters for tree‐based models were either left at library defaults or tuned via grid search with 5‐fold cross‐validation on the training portion. This transparent benchmarking allows direct assessment of whether the added complexity of the AP‐ILSTM yields meaningful gains over established, computationally lighter alternatives.

3. Results

3.1. Study Population and Dataset Characteristics

A total of 1248 consecutive pregnancies recorded between April 2024 and April 2025 were retrospectively screened for eligibility. Following comprehensive data quality assessment, duplicate patient entries, incomplete longitudinal records, and cases with substantial missing predictor variables were excluded according to the predefined exclusion criteria described in the Methods section. The remaining records constituted the final analytical cohort used for model development and validation. To ensure unbiased model evaluation, the dataset was randomly partitioned using stratified sampling, thereby preserving the prevalence of pregnancy‐related syndromes across all subsets. The final dataset consisted of 873 pregnancies (70.0%) assigned to the training cohort, 187 pregnancies (15.0%) allocated for validation during hyperparameter optimization, and an independent 188‐patient test cohort (15.0%) reserved exclusively for final performance evaluation. The mean maternal age was 29.8 ± 5.4 years, ranging from 18 to 45 years. Most participants were between 25 and 34 years of age (60.6%), whereas 19.9% were younger than 25 years and 19.5% were aged 35 years or older. Nulliparous women represented 54.6% of the study population, while 45.4% were multiparous. Singleton pregnancies accounted for 96.4% of all pregnancies, with only 3.6% involving multiple gestations. Regarding adverse pregnancy outcomes, 98 women (7.9%) developed preeclampsia, 156 women (12.5%) were diagnosed with gestational diabetes mellitus (GDM), and 112 pregnancies (9.0%) resulted in preterm birth. Overall, approximately 70.6% of pregnancies were completed without any of the investigated complications. The observed outcome distribution reflects the moderate class imbalance typically encountered in real‐world obstetric datasets and supports the necessity of applying weighted learning strategies during model development. Furthermore, structured clinical variables including maternal demographics, obstetric history, laboratory investigations, blood pressure measurements, medication history, and antenatal examination findings were successfully integrated with unstructured physician narratives extracted from electronic health records (EHRs). The NLP preprocessing pipeline transformed more than 185,000 clinical text documents into numerical feature representations using TF‐IDF and distributed word embeddings, thereby enabling simultaneous analysis of structured and free‐text information within the proposed AP‐ILSTM framework. Table 1 summarizes the baseline demographic and clinical characteristics of the study population.

TABLE 1.

Baseline demographic and clinical characteristics of the study population (n = 1248).

Variable Overall (n = 1248)
Dataset distribution
Total pregnancies 1248
Training cohort 873 (70.0%)
Validation cohort 187 (15.0%)
Independent test cohort 188 (15.0%)
Maternal demographics
Maternal age (years), mean ± SD 29.8 ± 5.4
Age range (years) 18–45
Age < 25 years 248 (19.9%)
Age 25–34 years 756 (60.6%)
Age ≥ 35 years 244 (19.5%)
Body mass index (kg/m2), mean ± SD 24.7 ± 4.1
Obstetric characteristics
Nulliparous 682 (54.6%)
Multiparous 566 (45.4%)
Singleton pregnancy 1203 (96.4%)
Multiple gestation 45 (3.6%)
Previous pregnancy complications 214 (17.1%)
Maternal hypertension history 108 (8.7%)
Maternal diabetes history 86 (6.9%)
Pregnancy outcomes
Preeclampsia 98 (7.9%)
Gestational diabetes mellitus 156 (12.5%)
Preterm birth 112 (9.0%)
No major pregnancy syndrome 882 (70.6%)

3.2. Performance of the Proposed AP‐ILSTM Model

The predictive performance of the proposed AP‐ILSTM model was evaluated on the independent test cohort and compared with three widely used machine learning classifiers, including K‐Nearest Neighbors (KNN), XGBoost, and CatBoost. As summarized in Table 2, the proposed model consistently achieved the highest performance across all evaluation metrics. The AP‐ILSTM framework achieved an overall classification accuracy of 98.85%, representing improvements of 2.56%, 0.71%, and 1.32% over KNN, XGBoost, and CatBoost, respectively. Precision reached 98.75%, indicating that nearly all pregnancies predicted to be high‐risk were confirmed clinically. Similarly, the recall value of 98.69% demonstrated that the proposed framework successfully detected almost all pregnancies that subsequently developed pregnancy‐related syndromes. The resulting F1‐score of 98.72% confirmed an excellent balance between sensitivity and precision. Moreover, AP‐ILSTM demonstrated the highest specificity (99.05%), minimizing false‐positive predictions, while achieving a negative predictive value (99.48%), indicating outstanding reliability in identifying low‐risk pregnancies. The corresponding ROC‐AUC of 0.993 further confirmed the excellent discriminative capability of the proposed framework. Compared with conventional machine learning algorithms, AP‐ILSTM consistently maintained superior balanced accuracy, suggesting greater robustness under moderately imbalanced class distributions.

TABLE 2.

Classification performance comparison on the independent testing dataset.

Performance metric KNN XGBoost CatBoost AP‐ILSTM
Accuracy (%) 96.29 98.14 97.53 98.85
Precision (%) 98.60 98.60 97.05 98.75
Recall (Sensitivity) (%) 98.60 98.60 97.05 98.69
F1‐score (%) 98.60 98.60 97.05 98.72
Specificity (%) 97.80 98.42 97.80 99.05
Negative Predictive Value (%) 97.90 98.91 98.10 99.48
ROC‐AUC 0.959 0.981 0.972 0.993
Balanced Accuracy (%) 97.10 98.30 97.40 98.71
Cohen's Kappa 0.931 0.962 0.948 0.984
Matthews Correlation Coefficient 0.934 0.965 0.951 0.986

Figure 3A illustrates the comparative classification performance of all evaluated algorithms. AP‐ILSTM consistently achieved the highest values across every performance metric, demonstrating the effectiveness of integrating adaptive hyperparameter optimization with sequential deep learning for pregnancy syndrome prediction.

FIGURE 3.

FIGURE 3

Comparative performance evaluation of the proposed Adaptive Penguin–Improved Long Short‐Term Memory (AP‐ILSTM) model for pregnancy syndrome risk prediction. (A) compares the classification performance of AP‐ILSTM with K‐Nearest Neighbors (KNN), XGBoost, and CatBoost using accuracy, precision, recall (sensitivity), F1‐score, specificity, negative predictive value (NPV), and balanced accuracy, demonstrating superior overall predictive performance of the proposed model. (B) presents the confusion matrix of AP‐ILSTM on the independent test dataset, showing 71 true‐positive, 115 true‐negative, one false‐positive, and one false‐negative prediction, corresponding to an overall accuracy of 98.85%. (C) illustrates the receiver operating characteristic (ROC) curves for all evaluated models, with AP‐ILSTM achieving the highest discriminative ability (AUC = 0.993). (D) compares regression performance using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE), where AP‐ILSTM consistently produced the lowest prediction errors. (E) depicts the relative importance of the top predictive variables identified by the AP‐ILSTM model, highlighting blood pressure trends, NLP‐derived clinical narrative embeddings, previous pregnancy complications, maternal age, and laboratory biomarkers as the most influential features contributing to pregnancy syndrome risk prediction.

3.3. Confusion Matrix Analysis

The confusion matrix presented in Table 3 demonstrates the classification outcomes of the AP‐ILSTM model on the independent testing dataset. Among the 188 pregnancies evaluated, the proposed model correctly classified 186 cases, corresponding to an overall accuracy of 98.85%. Only one false‐positive and one false‐negative prediction were observed. The false‐negative rate was exceptionally low (0.53%), indicating that very few pregnancies at risk were overlooked by the proposed model. Likewise, the false‐positive rate remained below 1%, minimizing unnecessary clinical interventions. Figure 3B presents the graphical confusion matrix heatmap.

TABLE 3.

Confusion matrix of AP‐ILSTM on the independent testing dataset.

Predicted positive Predicted negative Total
Actual positive 71 1 72
Actual negative 1 115 116
Total 72 116 188

3.4. Receiver Operating Characteristic Analysis

Receiver Operating Characteristic (ROC) analysis demonstrated excellent discrimination between pregnancies with and without adverse pregnancy syndromes. The proposed AP‐ILSTM achieved the highest AUC (0.993), exceeding XGBoost (0.981), CatBoost (0.972), and KNN (0.959). The ROC curve shown in Figure 3C indicates that AP‐ILSTM maintained high sensitivity while simultaneously preserving very high specificity across different classification thresholds, demonstrating excellent diagnostic capability.

3.5. Mean Absolute Error Analysis

Regression performance was additionally evaluated using Mean Absolute Error (MAE). The proposed AP‐ILSTM achieved the lowest prediction error (MAE = 4.05), representing reductions of approximately 38.4%, 43.1%, and 49.5% compared with Voting Regression, Linear Regression, and Gradient Boosting Regression, respectively. Figure 3D illustrates the comparison of prediction errors among all regression models (Table 4).

TABLE 4.

Regression performance comparison.

Model MAE RMSE MAPE (%)
Voting regression 6.58 8.42 8.61
Linear regression 7.12 9.01 9.45
Gradient boosting regression 8.02 10.14 10.37
AP‐ILSTM 4.05 5.18 4.92

3.6. Feature Contribution Analysis

To investigate the clinical interpretability of the proposed framework, feature importance analysis was performed using permutation importance on the optimized AP‐ILSTM model. Both structured clinical variables and NLP‐derived features contributed substantially to prediction performance. Blood pressure trajectories represented the most influential predictor, followed by physician narrative embeddings extracted using TF‐IDF and word embeddings. Previous pregnancy complications, maternal age, and laboratory biomarkers were also identified as major determinants of pregnancy syndrome risk (Table 5). Figure 3E illustrates the relative contribution of each predictor variable. The incorporation of NLP‐derived physician notes substantially enhanced model performance compared with using structured clinical variables alone, highlighting the importance of unstructured electronic health record data for early pregnancy syndrome risk prediction.

TABLE 5.

Top predictive variables identified by AP‐ILSTM.

Rank Predictor Relative importance (%)
1 Blood pressure trend 18.7
2 Clinical narrative embeddings 16.9
3 Previous pregnancy complications 14.8
4 Maternal age 12.5
5 Laboratory biomarkers 11.9
6 Gestational age 9.8
7 Previous obstetric history 6.9
8 Medication history 4.8
9 BMI 2.7
10 Other clinical variables 1.0

4. Discussion

The present study developed and evaluated an Adaptive Penguin–Improved Long Short‐Term Memory (AP‐ILSTM) model that integrates natural language processing of unstructured electronic health record (EHR) narratives with structured longitudinal clinical data for early prediction of preeclampsia, gestational diabetes mellitus (GDM), and preterm birth. On an independent test cohort, the model achieved excellent discriminative performance (accuracy 98.85%, sensitivity 98.69%, positive predictive value 98.75%, F1‐score 98.72%, ROC‐AUC 0.993) with low prediction error (MAE 4.05), modestly outperforming strong baseline classifiers including XGBoost, CatBoost, and K‐Nearest Neighbors (Collins et al. 2024).

Although absolute improvements over tree‐based ensembles were incremental (approximately 0.7–1.3 percentage points in accuracy), several considerations support the potential value of this hybrid architecture in maternal health risk stratification. Unlike models restricted to coded variables, the AP‐ILSTM explicitly incorporates semantic information from physician progress notes and clinical narratives. Narrative embeddings ranked as the second most influential predictor, demonstrating that free‐text documentation contains actionable risk signals not captured by structured fields alone. The improved LSTM component models temporal dependencies across serial antenatal visits, aligning with the progressive pathophysiology of many pregnancy complications. The Adaptive Penguin optimization with Gaussian exploration automates hyperparameter tuning, potentially improving robustness. However, these advantages must be weighed against practical trade‐offs. Deep recurrent models generally require greater computational resources during training and inference than gradient‐boosted trees such as XGBoost, which often complete training within minutes on standard hardware. In resource‐limited settings or for rapid iteration, simpler models may offer a more favorable accuracy–efficiency balance. In well‐resourced tertiary centers where even small gains in sensitivity can prevent adverse maternal or neonatal outcomes, the complex model may be justified if deployed as a background service. Future studies should routinely report training time, inference latency, memory usage, and energy consumption to guide context‐specific model selection.

A growing body of literature demonstrates the value of machine learning techniques for pregnancy and childbirth risk management. Kopanitsa et al. (2023) reviewed various machine learning approaches for generating data‐driven decision support models using real‐world clinical data (Kopanitsa et al. 2023). Sahithi et al. (2023) compared traditional machine learning and deep learning models, including XGBoost, CatBoost, and convolutional neural networks, for predicting maternal mortality risk levels and reported strong predictive performance (Sahithi et al. 2023). However, most existing studies rely predominantly on structured clinical variables and focus on single outcomes. Few have systematically integrated natural language processing of unstructured clinical narratives with advanced recurrent architectures and adaptive hyperparameter optimization for the simultaneous prediction of multiple key pregnancy syndromes such as preeclampsia, gestational diabetes mellitus, and preterm birth (Bashir et al. 2026).

A notable methodological limitation is the reliance on a single stratified train–validation–test split rather than repeated k‐fold or nested cross‐validation for final performance estimates. Although internal 5‐fold cross‐validation guided hyperparameter optimization, the held‐out test performance may still be optimistically biased. Systematic reviews of machine learning models for preeclampsia and other pregnancy outcomes have highlighted that many studies use internal validation only and report heterogeneous performance (AUROC ranging from 0.66 to 0.97), with frequent concerns about overfitting and limited external validation. The absence of external validation on independent cohorts from other institutions or regions further constrains generalizability. All data originated from a single regional hospital in China over a 13‐month period; no prior published validation studies have assessed the data quality or documentation practices of this specific EHR system. Local note‐writing conventions and potential bilingual terminology may have influenced model behavior in ways that do not transfer readily to other settings.

Internal validity considerations also warrant scrutiny. Selection bias cannot be excluded, as pregnancies with incomplete documentation were systematically excluded, potentially enriching the cohort for higher‐risk or more intensively monitored cases. Information bias is plausible because outcomes relied on clinician‐entered diagnoses without independent adjudication, and free‐text predictors remain vulnerable to documentation variability, abbreviations, and negation phenomena only partially mitigated by standard preprocessing. Many machine learning studies in obstetrics have been criticized for poor methodological conduct and high risk of bias, underscoring the need for transparent reporting according to the TRIPOD+AI guidelines (Collins et al. 2024).

Despite these limitations, the study has several strengths. It represents one of the early efforts to fuse longitudinal structured data with NLP‐derived embeddings inside an adaptively optimized recurrent framework specifically for multi‐syndrome pregnancy risk prediction. The sizable single‐center cohort, strict temporal separation of predictors from outcomes, and comprehensive reporting of discrimination, calibration‐adjacent, and error metrics provide a solid technical foundation. The addition of ROC‐AUC, specificity, and negative predictive value addresses previous concerns about threshold‐dependent metrics (Imani et al. 2026).

Looking ahead, several priorities emerge. Rigorous external validation across multiple centers and diverse populations is essential to establish transportability. Prospective cluster‐randomized implementation trials evaluating impact on clinical endpoints (e.g., timely aspirin initiation, mode of delivery, neonatal outcomes) are required to demonstrate real‐world utility. Methodological advances should explore transformer‐based language models and large language models fine‐tuned on obstetric corpora for deeper contextual understanding of clinical narratives. Integration of additional modalities serial ultrasound, placental biomarkers, wearable data, environmental exposures, and genetic information offers a logical route to further performance gains. Dynamic, updating risk models that ingest streaming antenatal data would better support real‐time clinical decision‐making. To facilitate adoption, future work must prioritize explainability through attention visualization and SHAP values linked to specific note excerpts, fairness audits across subgroups, and computational‐efficiency benchmarking. Privacy‐preserving approaches such as federated learning could enable collaborative multi‐institutional development without centralizing sensitive patient data. Expansion to additional outcomes, including postpartum hemorrhage and long‐term maternal cardiovascular risk, would broaden clinical relevance (Layton 2025; Islam et al. 2022).

In summary, while simpler models currently deliver broadly comparable aggregate performance at lower computational cost, the AP‐ILSTM framework demonstrates a scalable pathway to exploit the largely untapped narrative component of obstetric EHRs. Realizing its translational potential will require multi‐center validation, enhanced interpretability, prospective outcome studies, and deliberate attention to equity, privacy, and implementation science, in accordance with emerging standards such as TRIPOD+AI. Until such comprehensive evidence accumulates, the model should be regarded as a promising research prototype rather than an immediately deployable clinical decision‐support tool.

5. Conclusion

The AP‐ILSTM model represents a meaningful step forward in the application of artificial intelligence for maternal health risk stratification. By integrating natural language processing of unstructured clinical narratives with longitudinal structured data within an adaptively optimized recurrent neural network, the model achieved high predictive performance for the composite outcome of preeclampsia, gestational diabetes mellitus, and preterm birth. The substantial contribution of narrative embeddings highlights the value of leveraging the full richness of electronic health records beyond coded variables alone. While improvements over strong baseline models such as XGBoost were incremental, the hybrid multimodal and temporal architecture offers distinct advantages in capturing complex clinical patterns across pregnancy. The rigorous temporal separation of predictors and outcomes enhances the model's clinical relevance and reduces the risk of information leakage. Nevertheless, important limitations remain. The single‐center retrospective design, reliance on a single train‐test split, and absence of external validation constrain the generalizability of the findings. Potential selection and information biases inherent to routine EHR data must also be considered. Future research should prioritize multi‐center prospective validation, integration of additional modalities such as imaging and biomarkers, and the development of explainable, dynamic risk models suitable for real‐time clinical use. Privacy‐preserving techniques, including federated learning, may further facilitate broader collaboration. Overall, this work demonstrates the feasibility and potential clinical utility of advanced NLP‐enhanced deep learning for early pregnancy syndrome risk prediction. With continued methodological refinement and robust validation, such models may ultimately contribute to more personalized and proactive antenatal care, improving outcomes for mothers and their infants.

Author Contributions

Yanqi Wu: conceptualization, methodology, data curation, software development, algorithm design, formal analysis, validation, visualization, investigation, writing – original draft preparation, writing – review and editing, and supervision. The author read and approved the final manuscript.

Funding

The author has nothing to report.

Ethics Statement

This study was conducted in accordance with the principles of the Declaration of Helsinki. Ethical approval was obtained from the appropriate Institutional Review Board (IRB) before data collection. All patient data were anonymized prior to analysis. As only de‐identified retrospective electronic health record (EHR) data were used, the requirement for informed consent was waived by the ethics committee.

Consent

The author has nothing to report.

Conflicts of Interest

The author declares no conflicts of interest.

Supporting information

Figure S2: The LSTM's internal structure.

Figure S3: Neural networks cell makes up for LSTM and improved LSTM.

Figure S4: Flowchart of the AP Algorithm.

BDR2-118-e70095-s001.docx (360.6KB, docx)

Acknowledgments

The author would like to thank Eindhoven University of Technology for providing academic support and computational resources that facilitated this research.

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.

References

  1. Alfaki Ahmed, S. A. , Adam M., Fahad Alqahtani N. H., Mahdi Gabreldaar A. E., Hassan Abdalla M. S., et al. 2025. “Artificial Intelligence for Early Detection of Preeclampsia and Gestational Diabetes Mellitus: A Systematic Review of Diagnostic Performance.” Cureus 17, no. 9: e92585. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Bashir, S. G. G. , Salad H. A., Abdullahi Y. B., et al. 2026. “Artificial Intelligence for Predicting and Preventing Adverse Pregnancy Outcomes Addressing Bias and Clinical Translation.” Frontiers in Digital Health 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Borji, A. , Haick H., Pohn B., et al. 2025. “An Integrated Optimization and Deep Learning Pipeline for Predicting Live Birth Success in IVF Using Feature Optimization and Transformer‐Based Models.” Computer Methods and Programs in Biomedicine 271: 108979. [DOI] [PubMed] [Google Scholar]
  4. Collins, G. S. , Moons K. G. M., Dhiman P., et al. 2024. “TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods.” BMJ 385: e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Correia, V. , Mascarenhas T., and Mascarenhas M.. 2025. “Smart Pregnancy: AI‐Driven Approaches to Personalised Maternal and Foetal Health‐A Scoping Review.” Journal of Clinical Medicine 14, no. 19: 6974. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Coupland, H. , Scheidwasser N., Katsiferis A., et al. 2025. “Exploring the Potential and Limitations of Deep Learning and Explainable AI for Longitudinal Life Course Analysis.” BMC Public Health 25, no. 1: 1520. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Imani, M. , Sadeghi‐Nodoushan F., Heydari S., Dehghani‐Sanij S., and Fesahat F.. 2026. “Machine Learning for Predictive Risk Stratification in Recurrent Miscarriage: A Systematic Review.” BMC Medical Informatics and Decision Making 26, no. 1: 113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Irfan, N. , Zafar S., Shakil K. A., et al. 2025. “Integrating AI Predictive Analytics With Naturopathic and Yoga‐Based Interventions in a Data‐Driven Preventive Model to Improve Maternal Mental Health and Pregnancy Outcomes.” Scientific Reports 15, no. 1: 23878. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Islam, M. N. , Mustafina S. N., Mahmud T., and Khan N. I.. 2022. “Machine Learning to Predict Pregnancy Outcomes: A Systematic Review, Synthesizing Framework and Future Research Agenda.” BMC Pregnancy and Childbirth 22, no. 1: 348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Janeline, L. , PramodKumar T. A., Deepa M., et al. 2026. “Overlap of Gestational Diabetes Mellitus and Pre‐Eclampsia: A Scoping Review.” Diabetes Therapy. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Khan, M. , Khurshid M., Vatsa M., Singh R., Duggal M., and Singh K.. 2022. “On AI Approaches for Promoting Maternal and Neonatal Health in Low Resource Settings: A Review.” Frontiers in Public Health 10: 880034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Kopanitsa, G. , Metsker O., and Kovalchuk S.. 2023. “Machine Learning Methods for Pregnancy and Childbirth Risk Management.” Journal of Personalized Medicine 13, no. 6: 975. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Kwok, W. H. , Zhang Y., and Wang G.. 2024. “Artificial Intelligence in Perinatal Mental Health Research: A Scoping Review.” Computers in Biology and Medicine 177: 108685. [DOI] [PubMed] [Google Scholar]
  14. Layton, A. T. 2025. “Artificial Intelligence and Machine Learning in Preeclampsia.” Arteriosclerosis, Thrombosis, and Vascular Biology 45, no. 2: 165–171. [DOI] [PubMed] [Google Scholar]
  15. Mehrnoush, V. , Haghighat A., Nami A., et al. 2026. “The Power of Machine Learning Models in Predicting Gestational Diabetes Mellitus.” BMC Pregnancy and Childbirth 26, no. 1: 348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Sahithi, P. , Amulya S., Gajapathi A. S., Raju S. V. V., Rukmini K., and Murthy D. S. V.. 2023. “Deep Learning Based Risk Level Prediction Model for Maternal Mortality.” International Journal of Engineering Research and Applications 13, no. 4: 70–80. [Google Scholar]
  17. Shabani, F. , Jodeiri A., Mohammad‐Alizadeh‐Charandabi S., Abbasalizadeh F., Tanha J., and Mirghafourvand M.. 2025. “Developing and Validating an Artificial Intelligence‐Based Application for Predicting Some Pregnancy Outcomes: A Multi‐Phase Study Protocol.” Reproductive Health 22, no. 1: 99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Soler, M. , Parke B., Kim S. H., Terzidou V., and Ladame S.. 2026. “Emerging Biomarkers and Diagnostic Tools for the Early Prediction of Adverse Prenatal Outcomes.” NPJ Women's Health 4, no. 1: 20. [Google Scholar]
  19. Staff, A. C. , Costa M. L., Dechend R., Jacobsen D. P., and Sugulle M.. 2024. “Hypertensive Disorders of Pregnancy and Long‐Term Maternal Cardiovascular Risk: Bridging Epidemiological Knowledge Into Personalized Postpartum Care and Follow‐Up.” Pregnancy Hypertension 36: 101127. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S2: The LSTM's internal structure.

Figure S3: Neural networks cell makes up for LSTM and improved LSTM.

Figure S4: Flowchart of the AP Algorithm.

BDR2-118-e70095-s001.docx (360.6KB, docx)

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.


Articles from Birth Defects Research are provided here courtesy of Wiley

RESOURCES