Abstract
Background
Fetal hypoxia is a leading cause of neonatal morbidity and mortality. Cardiotocography (CTG) is widely used to predict fetal hypoxia during labor, but its interpretation remains suboptimal. Artificial intelligence (AI) models have been developed for CTG interpretation, but their clinical utility is limited by two major challenges: demonstrating superiority over human experts and ensuring explainability in real-world settings.
Methods
A large dataset containing CTG traces from three tertiary hospitals between January 2014 and May 2022 was built for model development. Deep learning architectures, named Cardiotocography Artificial-intelligence Predictors (CAPs), were trained to predict fetal hypoxia from CTG traces based on CNN (CAP-C), Transformer (CAP-T), LSTM (CAP-L), and CfC (CAP-CfC) algorithms. The outcome was fetal hypoxia, determined by either low Apgar score (≤ 7 at 1 or 5 min) or umbilical artery acidemia (grade 1: pH of umbilical artery (pHa) < 7.20; grade 2: pHa < 7.15; grade 3: pHa < 7.10). Model performance was determined by area under the receiver operating characteristic curve (AUROC), evaluated through nationwide AI-human comparison and validated on the CTU-UHB dataset. Gradient-weighted class activation mapping (Grad-CAM) was applied to highlight the CTG regions that contributed most to the model’s predictions.
Results
A total of 20,780 CTG traces were obtained for model development, and 467 cases were held out for the nationwide AI-human comparison. Among all models, CAP-L achieved highest AUROC in predicting fetal hypoxia (grade 1: 0.758, 95% CI: 0.754–0.761; grade 2: 0.770, 95% CI: 0.764–0.776; grade 3: 0.716, 95% CI: 0.700–0.732). In comparison with 10,571 expert responses, all CAP models achieved higher AUROC (0.757–0.789 vs. 0.715, P values in Delong test < 0.05). On the public CTU-UHB dataset, CAP-L achieved AUROC of 0.709, 0.727, and 0.730 in predicting fetal hypoxia with grade 1, 2, and 3 acidemia. Grad-CAM analysis showed that the CAP models leveraged variable and prolonged decelerations to predict fetal hypoxia, verified by perturbation-based faithfulness test.
Conclusions
The CAP algorithms developed in this study showed superior performance in detecting fetal hypoxia from CTG traces compared to human experts, and demonstrated promising explainability, supporting clinical CTG interpretation.
Trial registration
Clinical trial registration number: ChiCTR2100045316, ChiCTR2100052695, ChiCTR2400085338.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12916-026-04794-z.
Keywords: Fetal hypoxia, Artificial intelligence, Deep learning, AI-human comparison, Explainability
Background
Fetal hypoxia is a leading cause of fetal and neonatal mortality and morbidity, often resulting in cerebral palsy, organ dysfunction, and even intrauterine or neonatal death [1–7]. Cardiotocography (CTG) is a major tool to detect fetal hypoxia [8–15], but the interpretation of CTG is highly subjective, leading to variability in diagnosis and either delayed or unnecessary intervention [16–18]. Although computer programs are used in CTG interpretation, recent study showed that existing software did not improve predictive accuracy of fetal hypoxia [19, 20].
Artificial intelligence (AI) has been increasingly applied in medical field [21–25], including CTG interpretation [26–32]. Petrozziello et al. applied convolutional neural network (CNN) to predict fetal hypoxia [33], and further reported a multimodal convolutional neural network (MCNN) to detect fetal compromise [34]. Cömert and Kocamaz applied short-time Fourier transform (STFT) followed by CNN to predict fetal hypoxia [35]. Recently, McCoy et al. compared six deep learning algorithms in predicting fetal acidemia from CTG trace [36].
Despite recent advances, AI-assisted CTG interpretation still faces several significant challenges. First, most deep learning models were trained on small, highly selected datasets, which often fail to capture the wide variability present in real-world CTG recordings. To date, only a few studies have trained models on cohorts exceeding 10,000 cases [33, 34, 36, 37]. Second, it remains unclear whether deep learning algorithms outperform trained human interpreters in real clinical settings. Prior studies typically compared AI models to routine clinical practice [34, 38] or assessments by limited expert panels [39], thus restricting the generalizability of their findings. Third, model explainability is unknown for most AI models applied to CTG interpretation, which become a major challenge in their clinical application.
To address this gap, we developed a set of deep learning models to interpretate CTG using data from a large multicenter cohort and applied gradient-weighted class activation mapping (Grad-CAM) to explore the decision-making process of the models. Importantly, we validated the model performances against trained clinicians and midwives in a nationwide AI-human comparison, aiming to explore the clinical utility of AI-assisted CTG interpretation for decision support in detecting fetal hypoxia and improving neonatal outcomes.
Methods
The overall study design contains four phases: deep learning architectures development, nationwide AI-human comparison, external validation, and evaluation of model explainability (Fig. 1).
Fig. 1.
An overview of the study design. A total of 20,780 CTG traces were collected for the model development. Four deep learning architectures (CAP-C, CAP-T, CAP-L, CAP-CfC) were built to predict fetal hypoxia. A dataset containing 467 CTG traces was held out for the nationwide AI-human comparison. Human interpreters can get access to the nationwide AI-human comparison by scanning the quick-response code. The generalizability of CAP models was examined in the CTU-UHB dataset. Gradient-weighted class activation mapping (Grad-CAM) were employed to evaluate the interpretation of CAP models
CTG traces for deep learning architectures development
The CTG traces and clinical information were obtained from three hospitals, the First Affiliated Hospital of Sun Yat-sen University, the Guangzhou Women and Children’s Medical Center in Guangdong province, and the Sanming First Hospital in Fujian province. Singleton pregnant women aged ≥ 18 years and had vaginal delivery with cephalic presentation between January 1, 2014, and May 31, 2022, were recruited and matched with the intrapartum CTG records. The exclusion criteria were (1) gestational age < 34 weeks, (2) fetal growth restriction found before delivery, (3) no neonatal record in detail, (4) CTG trace lasting < 40 min, and (5) CTG traces contain missing values > 20% (Additional file 1: Fig. S1). Clinical data were collected from medical records, including demographic features, delivery related information, and neonatal information.
The study was approved by the ethical committees of the First Affiliated Hospital of Sun Yat-sen University ([2020]278–1, [2021]431–1, [2022]451) and registered at chictr.org.cn on April 12, 2021 (ChiCTR2100045316), November 3, 2021 (ChiCTR2100052695), and June 5, 2024 (ChiCTR2400085338) and followed the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis for Artificial Intelligence (TRIPOD + AI) reporting guideline [40]. For the retrospective analysis of anonymized CTG and clinical data, the requirement for informed consent was waived by the ethical committees. For the online survey, informed consent was obtained from all participants.
Data labeling and grouping
The primary outcome of the present study is fetal hypoxia [41, 42], determined by either low Apgar scores (≤ 7 at 1 min or 5 min) or umbilical artery acidemia (grade 1: pH of umbilical artery (pHa) < 7.20; grade 2: pHa < 7.15; grade 3: pHa < 7.10), aligning with the diagnostic criteria for neonatal asphyxia from the Chinese Society of Perinatal Medicine [43]. The graded acidemia thresholds were adopted to reflect severity strata consistent with international guidelines [44–46] and prior evidence [36, 41].
Among the dataset, 20,780 cases were used for model development, and 467 cases were held out for the nationwide AI-human comparison (Additional file 1: Fig. S1). The development set was randomly split into train-validation set and test set with a ratio of 10:1, and the train-validation set were further randomly divided into ten parts with balanced hypoxia distribution for the tenfold cross-validation.
Data pre-processing
The raw CTG data was 1 Hz signal with fetal heart rate (beat per min) and pressure of uterine contraction (kPa). This data was pre-processed before it can be input into the deep learning models as follows: (1) signal range 2400 s before delivery was selected, (2) outliers (FHR > 200 bpm or < 50 bpm) were removed, (3) missing values were replaced by Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) spline interpolation [47] using the PchipInterpolator class from the scipy.interpolate module of the SciPy Python library [48]. For each FHR tracing, PCHIP interpolation generated a cubic function between adjacent observed data points, while preserving local monotonicity, thereby preventing the introduction of non-physiological oscillations or spurious peaks, and (4) uterine contraction signal was smoothed by applying the average value of 13 signals centered on it.
Deep learning algorithms and training
A series of deep learning models, named Cardiotocography Artificial-intelligence Predictors (CAPs), were developed to interpretate CTG traces. In CAP-C, CTG signal was input to one-dimensional convolutional neural networks (1D CNNs). In CAP-T, CAP-L, and CAP-CfC architectures, CTG signal was input to 1D CNNs followed by the Transformer [49], long-term short-term memory (LSTM), or closed-form continuous-time neural networks (CfC) layers [50] (Additional file 2: Fig. S2). Model parameters were shown in Additional file 3: Tables S1–S4. The primary models for each architecture were built by using full 40-min signal of fetal heart rate (FHR) and uterine contraction (UC). Additionally, separate models were also trained using only the FHR signal, and signals of the last 20 min CTG segments to accommodate different clinical scenarios.
Models were trained using tenfold cross-validation with a 9:1 train-validation split in each fold. The best-performing algorithm on the validation set was applied to the test set. Training was performed with the Adam optimizer (learning rate = 0.001) for 200 epochs, and weighted binary cross-entropy was used to address class imbalance [51]. The deep learning code was written in Python (3.8 version) using the Keras library (2.7.0 version), and the training process was completed in the Tianhe-2 National Supercomputer Center in Guangzhou, China.
Evaluation of model performance
To compare the performance of different algorithms, the area under the receiver operating characteristic curves (AUROC), F1 score, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were assessed in the test set for each model. A standard probability threshold of 0.5 was applied to binarize model predictions for the calculation of F1 score, sensitivity, specificity, PPV, and NPV.
Nationwide AI-human comparison on CTG interpretation
To investigate whether the CAP deep learning architectures outperform trained human interpreters, we conducted an online nationwide AI-human comparison.
CTG images for the nationwide AI-human comparison
A holdout dataset containing 467 CTG trace images were used for this AI-human comparison. The 467 cases were selected from the initial phase of the data collection to facilitate model development and the online survey in parallel. To avoid few abnormal CTG in each test, we included all available cases with fetal hypoxia from the initial cohort and randomly selected a comparable number of normal cases to create a balanced set. In this dataset, 172 cases with pHa < 7.20 were labeled as fetal hypoxia, and the remaining 295 cases were taken as normal. The CTG tracings were converted into digital images following a fixed ratio corresponding to 1 cm/min. The CTG trace images were uploaded to the online database with the website https://www.wjx.cn/, and participants can get access to the online survey by scanning the quick-response code (shown in Fig. 1) with their phones or pads.
The procedure of the online survey
The online CTG survey were broadcasted among obstetricians, gynecologists, reproductive specialists, and midwives through academic conferences, lectures, and researchers’ social networks.
Before the test, participants were required to provide basal information, including the city they were working in, the specialties, and years of practice. To ensure respondent professionalism, three threshold questions were established: “What is the normal range for fetal heart rate baseline?,” “Which of the following is a contraindication for vaginal delivery?,” and “What is the definition of the first stage of labor?.” Participants who correctly answer all three threshold questions are deemed valid respondents.
The system randomly assigned 10 from the 467 CTG images without clinical information to the participants with two questions: (1) What is the category of the CTG image? I, II, or III. (According to the Expert Consensus on Application of Electronic Fetal Heart Monitoring [52]) (2) According to this CTG image, please predict the outcome of the case. Fetal hypoxia (pHa < 7.20) or normal (pHa ≥ 7.20). After confirming each prediction, participants received immediate feedback on the correctness of their judgment on the web interface. Once confirmed, the answer was locked and could not be changed, and that specific CTG tracing would not be presented again for the remainder of that test session. Upon completing all assessments, participants were shown their total score. The questionnaire format and the procedure of the online survey are illustrated in Additional file 4: Fig. S3, and the questions in the online survey are listed in Additional file 5.
Comparison of deep learning algorithms and human interpreters
To investigate whether the CAP architectures outperform human in predicting intrauterine hypoxia, we used trained deep learning models to interpret the CTG traces for the nationwide online survey. Model predicted probabilities of fetal hypoxia (using 40 min FHR and UC signal) of each architecture were compared with the average probability from human interpreters. The AUROC, as well as sensitivity, specificity, and F1 score (calculated on binary outputs using a standard threshold of 0.5), were compared between the deep learning algorithms and human interpreters.
Evaluation of model generalization in the CTU-UHB dataset
To evaluate the generalization of the deep learning models developed in the present study, we investigated the models’ performance in the CTU-UHB dataset [53]. The CTU-UHB dataset is a public dataset containing 552 CTG traces from Czech Technical University (CTU) and University Hospital in Brno (UHB). The raw data in the CTU-UHB dataset was 4 Hz. The signal was smoothed and down sampled to 1 Hz and then pre-processed with the same approaches described above for the model developing data. Then, all models developed in the present study was applied to this external validating set, to investigate the performance of the deep learning models in new CTG traces.
Evaluation of model explainability
To achieve model explainability, we applied gradient-weighted class activation mapping (Grad-CAM) [54] to highlight CTG regions most influential to model predictions. Saliency maps were generated by extracting activation values from key layers (the final convolutional layer of CAP-C model, the first LSTM layer of CAP-L model, the last multi-head self-attention layer of CAP-T model, and the CfC layer of CAP-CfC model) and were superimposed on the original CTG traces.
Then, a perturbation-based evaluation was conducted to quantitatively assess the faithfulness of the Grad-CAM explanations. For each test sample, the Grad-CAM attention weights were computed from the target layer, and the top and bottom 10% of time points ranked by attention importance were identified. Critical regions (top 10%) were ablated by replacing them with local baseline values calculated from a 300-s preceding window, with a global mean fallback. Control regions (bottom 10%) were perturbed identically. Faithfulness was quantified by comparing the prediction probability drop after ablating important versus unimportant features.
Results
Baseline characteristics
A total of 34,438 cases that met the inclusion criteria were recruited in the multicenter cohort. Among them, 450 (1.3%) with gestational weeks < 34, 450 (1.3%) with a diagnosis of fetal growth restriction found before delivery, and 412 (1.2%) without neonatal record in detail were excluded due to clinical factors. In addition, 4709 (13.7%) with CTG length less than 40 min and 7170 (20.8%) with missing value more than 20% were excluded due to CTG reasons. Ultimately, a total of 21,247 cases with CTG and clinical information were included in the final analysis.
In the final cohort, a dataset of 20,780 cases was used to train, validate, and test the deep learning algorithms. Among them, 2219 (10.7%) cases with pHa < 7.20 or Apgar ≤ 7, 1177 (5.7%) cases with pHa < 7.15 or Apgar ≤ 7, and 647 (3.1%) cases with pHa < 7.10 or Apgar ≤ 7. A holdout dataset containing 467 CTG traces was used to compare the performance of deep learning models with trained human interpreters. The clinical characteristics of the study population were shown in Table 1.
Table 1.
Clinical characteristics of the research population
| Training, validating, and testing set (mean ± SD) | Online testing set (mean ± SD) | |||||||
|---|---|---|---|---|---|---|---|---|
| Overall | Normal | Hypoxia* | P | Overall | Normal | Hypoxia* | P | |
| N | 20,780 | 18,561 | 2219 | 467 | 295 | 172 | ||
| Age | 29.949 ± 4.343 | 29.895 ± 4.376 | 30.400 ± 4.031 | < 0.001 | 30.310 ± 4.215 | 30.451 ± 4.119 | 30.070 ± 4.375 | 0.354 |
| Gestational age (week) | 38.796 ± 1.386 | 38.779 ± 1.395 | 38.940 ± 1.300 | < 0.001 | 38.931 ± 1.140 | 38.885 ± 1.110 | 39.012 ± 1.190 | 0.255 |
| Parity, N (%) | ||||||||
| Nulliparity | 12,339 (59.597) | 10,583 (57.252) | 1756 (79.135) | < 0.001 | 392 (83.940) | 245 (83.051) | 147 (85.465) | 0.579 |
| Multiparity | 8365 (40.403) | 7902 (42.748) | 463 (20.865) | 75 (16.060) | 50 (16.949) | 25 (14.535) | ||
| Delivery mode, N (%) | ||||||||
| Natural | 19,085 (91.843) | 17,447 (93.998) | 1638 (73.817) | < 0.001 | 358 (76.660) | 238 (80.678) | 120 (69.767) | 0.010 |
| Instrumental | 1695 (8.157) | 1114 (6.002) | 581 (26.183) | 109 (23.340) | 57 (19.322) | 52 (30.233) | ||
| Anesthesia, N (%) | 5866 (28.229) | 5104 (27.499) | 762 (34.340) | < 0.001 | 23 (4.925) | 15 (5.085) | 8 (4.651) | 1.000 |
| Stages of labor (min) | ||||||||
| First | 477.745 ± 259.775 | 468.422 ± 255.548 | 555.727 ± 280.945 | < 0.001 | 488.418 ± 292.072 | 481.237 ± 305.813 | 500.807 ± 267.126 | 0.471 |
| Second | 42.461 ± 39.092 | 40.491 ± 37.764 | 58.939 ± 45.588 | < 0.001 | 44.651 ± 42.868 | 42.817 ± 44.904 | 47.797 ± 39.058 | 0.210 |
| Third | 7.080 ± 6.861 | 7.004 ± 6.779 | 7.719 ± 7.483 | < 0.001 | 10.276 ± 7.566 | 10.292 ± 7.345 | 10.250 ± 7.952 | 0.955 |
| Total | 527.274 ± 275.936 | 515.924 ± 271.060 | 622.225 ± 297.324 | < 0.001 | 543.378 ± 309.477 | 534.346 ± 324.229 | 558.959 ± 282.466 | 0.391 |
| Birthweight (g) | 3179.8 ± 395.9 | 3178.7 ± 395.7 | 3188.5 ± 398.4 | 0.276 | 3215.1 ± 354.4 | 3201.8 ± 343.9 | 3237.9 ± 371.9 | 0.299 |
| Gender of the neonate, N (%) | ||||||||
| Female | 9793 (47.127) | 8754 (47.163) | 1039 (46.823) | 0.779 | 229 (49.036) | 144 (48.814) | 85 (49.419) | 0.976 |
| Male | 10,987 (52.873) | 9807 (52.837) | 1180 (53.177) | 238 (50.964) | 151 (51.186) | 87 (50.581) | ||
| Apgar scores | ||||||||
| 1 min | 9.499 ± 0.679 | 9.567 ± 0.534 | 8.925 ± 1.252 | < 0.001 | 9.647 ± 0.894 | 9.895 ± 0.358 | 9.221 ± 1.292 | < 0.001 |
| 5 min | 9.938 ± 0.286 | 9.956 ± 0.212 | 9.781 ± 0.601 | < 0.001 | 9.955 ± 0.236 | 9.993 ± 0.082 | 9.890 ± 0.366 | < 0.001 |
| 10 min | 9.964 ± 0.212 | 9.972 ± 0.175 | 9.899 ± 0.400 | < 0.001 | 9.989 ± 0.103 | 10.000 ± 0.000 | 9.971 ± 0.168 | 0.025 |
| pH in umbilical artery | 7.232 ± 0.074 | 7.268 ± 0.046 | 7.146 ± 0.056 | < 0.001 | 7.215 ± 0.079 | 7.264 ± 0.038 | 7.133 ± 0.060 | < 0.001 |
| PROM, N (%) | 5097 (24.528) | 4559 (24.562) | 538 (24.245) | 0.763 | 114 (24.411) | 67 (22.712) | 47 (27.326) | 0.314 |
| Cord entanglement, N (%) | 5055 (24.326) | 4305 (23.194) | 750 (33.799) | < 0.001 | 152 (32.548) | 88 (29.831) | 64 (37.209) | 0.124 |
| Cord torsion, N (%) | 452 (2.175) | 368 (1.983) | 84 (3.785) | < 0.001 | 30 (6.424) | 17 (5.763) | 13 (7.558) | 0.570 |
| Macrosomia, N (%) | 378 (1.819) | 334 (1.799) | 44 (1.983) | 0.598 | 5 (1.071) | 3 (1.017) | 2 (1.163) | 1.000 |
| Congenital heart disease, N (%) | 87 (0.431) | 75 (0.414) | 12 (0.574) | 0.378 | 5 (1.071) | 4 (1.356) | 1 (0.581) | 0.656 |
| Hypertension, N (%) | 928 (4.466) | 800 (4.310) | 128 (5.768) | 0.002 | 16 (3.426) | 7 (2.373) | 9 (5.233) | 0.169 |
| Gestational diabetes, N (%) | 3391 (16.319) | 3000 (16.163) | 391 (17.621) | 0.084 | 76 (16.274) | 43 (14.576) | 33 (19.186) | 0.241 |
| Obesity, N (%) | 393 (1.891) | 303 (1.632) | 90 (4.056) | < 0.001 | 22 (4.711) | 13 (4.407) | 9 (5.233) | 0.857 |
| Prior cesarean, N (%) | 422 (2.031) | 369 (1.988) | 53 (2.388) | 0.236 | 11 (2.355) | 7 (2.373) | 4 (2.326) | 1.000 |
| Intrauterine infection, N (%) | 413 (1.987) | 292 (1.573) | 121 (5.453) | < 0.001 | 24 (5.139) | 11 (3.729) | 13 (7.558) | 0.112 |
| Assisted reproductive, N (%) | 811 (3.903) | 686 (3.696) | 125 (5.633) | < 0.001 | 34 (7.281) | 18 (6.102) | 16 (9.302) | 0.272 |
| Anemia, N (%) | 1496 (19.515) | 1208 (19.114) | 288 (21.397) | 0.060 | 128 (27.409) | 87 (29.492) | 41 (23.837) | 0.225 |
| ICP, N (%) | 218 (1.049) | 196 (1.056) | 22 (0.991) | 0.864 | 2 (0.428) | 1 (0.339) | 1 (0.581) | 1.000 |
| Maternal heart diseases, N (%) | 617 (2.969) | 528 (2.845) | 89 (4.011) | 0.003 | 23 (4.925) | 17 (5.763) | 6 (3.488) | 0.382 |
*Fetal hypoxia cases in this table defined as pHa < 7.20 or Apgar score ≤ 7. PROM, premature rupture of membranes; ICP, intrahepatic cholestasis of pregnancy
Performance of the deep learning algorithms
Four deep learning architectures were developed to predict fetal hypoxia from CTG traces, based on CNN (CAP-C), Transformer (CAP-T), LSTM (CAP-L), and CfC (CAP-CfC) algorithms.
The performance of CAPs in the test set was shown in Table 2. In predicting fetal hypoxia with grade 1 acidemia (pHa < 7.20), the CAP-C and CAC-L both achieved highest AUROC (CAP-C: 0.758, 95% CI: 0.755–0.761; CAP-L: 0.758, 95% CI: 0.754–0.761). In predicting fetal hypoxia with grade 2 acidemia (pHa < 7.15), the CAP-L achieved highest AUROC (0.770, 95% CI: 0.764–0.776). In predicting fetal hypoxia with grade 3 acidemia (pHa < 7.10), the CAP-T and CAP-L both achieved highest AUROC (CAP-T: 0.716, 95% CI: 0.709–0.723; CAP-L: 0.716, 95% CI: 0.700–0.732).
Table 2.
Performance of Cardiotocography Artificial-intelligence Predictors (CAPs) in the testing set (mean, 95% CI)
| Architecture | AUROC | F1 scores | Sensitivity | Specificity | PPV | NPV |
|---|---|---|---|---|---|---|
| Prediction of pHa < 7.20 or Apgar ≤ 7 | ||||||
| CAP-C | 0.758 (0.755, 0.761) | 0.325 (0.308, 0.342) | 0.628 (0.532, 0.724) | 0.722 (0.640, 0.805) | 0.232 (0.204, 0.259) | 0.944 (0.936, 0.953) |
| CAP-T | 0.753 (0.748, 0.757) | 0.308 (0.293, 0.323) | 0.738 (0.685, 0.790) | 0.626 (0.562, 0.691) | 0.197 (0.181, 0.213) | 0.953 (0.949, 0.958) |
| CAP-L | 0.758 (0.754, 0.761) | 0.320 (0.308, 0.331) | 0.681 (0.633, 0.729) | 0.687 (0.640, 0.734) | 0.211 (0.198, 0.225) | 0.948 (0.944, 0.952) |
| CAP-CfC | 0.748 (0.744, 0.753) | 0.314 (0.304, 0.325) | 0.670 (0.628, 0.711) | 0.686 (0.642, 0.730) | 0.207 (0.196, 0.218) | 0.946 (0.943, 0.950) |
| Prediction of pHa < 7.15 or Apgar ≤ 7 | ||||||
| CAP-C | 0.742 (0.736, 0.748) | 0.184 (0.164, 0.204) | 0.642 (0.585, 0.699) | 0.721 (0.656, 0.786) | 0.110 (0.093, 0.126) | 0.977 (0.975, 0.978) |
| CAP-T | 0.765 (0.761, 0.769) | 0.159 (0.139, 0.179) | 0.776 (0.740, 0.812) | 0.584 (0.503, 0.664) | 0.089 (0.076, 0.103) | 0.982 (0.981, 0.983) |
| CAP-L | 0.770 (0.764, 0.776) | 0.197 (0.175, 0.219) | 0.663 (0.593, 0.732) | 0.732 (0.660, 0.804) | 0.118 (0.102, 0.135) | 0.978 (0.976, 0.981) |
| CAP-CfC | 0.757 (0.753, 0.761) | 0.215 (0.190, 0.241) | 0.599 (0.536, 0.662) | 0.783 (0.701, 0.864) | 0.135 (0.114, 0.157) | 0.976 (0.974, 0.978) |
| Prediction of pHa < 7.10 or Apgar ≤ 7 | ||||||
| CAP-C | 0.709 (0.705, 0.713) | 0.105 (0.094, 0.115) | 0.628 (0.601, 0.655) | 0.664 (0.612, 0.716) | 0.057 (0.051, 0.064) | 0.983 (0.982, 0.983) |
| CAP-T | 0.716 (0.709, 0.723) | 0.106 (0.095, 0.117) | 0.619 (0.565, 0.674) | 0.667 (0.586, 0.747) | 0.059 (0.051, 0.066) | 0.983 (0.982, 0.983) |
| CAP-L | 0.716 (0.700, 0.732) | 0.120 (0.105, 0.136) | 0.593 (0.539, 0.647) | 0.720 (0.650, 0.791) | 0.068 (0.057, 0.078) | 0.983 (0.982, 0.984) |
| CAP-CfC | 0.714 (0.704, 0.723) | 0.161 (0.131, 0.191) | 0.523 (0.489, 0.557) | 0.821 (0.771, 0.871) | 0.098 (0.076, 0.121) | 0.982 (0.982, 0.983) |
AUROC area under the receiver operating characteristic curve, PPV positive predictive value, NPV negative predictive value, CAP-C architecture with convolutional neural network, CAP-T architecture with Transformer, CAP-L architecture with long short-term memory, CAP-CfC architecture with closed-form continuous-time neural networks
For models trained using only the fetal heart rate signal, the highest AUROC in predicting fetal hypoxia with different grades of acidemia were CAP-C for grade 1 (0.760, 95% CI: 0.758–0.762), CAP-T for grade 2 (0.760, 95% CI: 0.754–0.767), and CAP-CfC for grade 3 (0.716, 95% CI: 0.710–0.722) (Additional file 3: Table S5).
In models trained on the last 20 min CTG signal, CAP-C model (using only the FHR signal) achieved the highest AUROC in predicting fetal hypoxia with grade 1 acidemia (0.764, 95% CI: 0.762–0.767). CAP-L model (using both FHR and UC signals) achieved the highest AUROC in predicting fetal hypoxia with grade 2 (0.770, 95% CI: 0.764–0.776) and grade 3 acidemia (0.729, 95% CI: 0.723–0.735) (Additional file 3: Table S6).
Performance of the human interpreters
In the nationwide AI-human comparison, 11,427 responses were provided by 7540 human interpreters from 34 provinces and districts in China. Among them, 11,064 were from obstetrician, gynecologists, reproductive specialists, and midwives. A total of 493 responses who provided wrong answers to the threshold questions were excluded. Ultimately, 10,571 responses were included in the study. Among them, 7435 were from obstetricians, 812 were from gynecologists or reproductive specialists, and 2324 were from midwives (Fig. 2a). A number of 4255 responses were provided by participants that had less than 10 years of experience, and the other 6316 had equal or more than 10 years of experience. The average number of ratings per CTG image is 274.0 (95% CI: 244.0–303.9), and the proportion of pathological CTGs in each test is 31.6% (95% CI: 31.5%–31.9%).
Fig. 2.
Performance of human interpreters in the nationwide AI-human comparison in CTG interpretation. a Components of the human interpreters in the nationwide survey. The sunburst chart showed participants different in specialty (inner ring) and year of clinical practice (outer ring). In the outer ring, colors from light to dark indicates years of clinical practice, from 1 year to up to 10 years. b The prediction performance of human interpreters with different specialties. Red: obstetricians. Blue: gynecologists or reproductive specialists. Teal: midwives. c The AUROC of human interpreters with different specialties. d The AUROC of human interpreters with different years of practice. Senior: more than 10 years of practice; Junior: less or equal to 10 years of practice. GYN & RS: gynecologists or reproductive specialists
The overall accuracy of human interpreters was 69.2%. Interestingly, midwives had higher AUROC, accuracy, sensitivity, sensitivity, PPV, NPV, and F1 score than obstetricians and gynecologists/reproductive specialists (Fig. 2b, c). Participants that had up to 10 years of experience did not show much higher AUROC than junior practitioner (Fig. 2d).
AI vs. human to predict fetal hypoxia
The performance of the CAP models on the 467 CTG cases in the AI-human comparison was analyzed (Table 3). CAP-L achieved the highest AUROC (0.789, 95% CI: 0.783–0.794) in predicting fetal hypoxia with grade 1 acidemia. In predicting fetal hypoxia with grade 2 and 3 acidemia, CAP-CfC achieved best AUROC (grade 2: 0.801, 95% CI: 0.794–0.807; grade 3: 0.829, 95% CI: 0.823–0.834), followed by CAP-L (grade 2: 0.800, 95% CI: 0.791–0.808; grade 3: 0.825, 95% CI: 0.819–0.831).
Table 3.
Performance of Cardiotocography Artificial-intelligence Predictors (CAPs) in AI-human comparison (mean, 95% CI)
| Architecture | AUROC | F1 scores | Sensitivity | Specificity | PPV | NPV |
|---|---|---|---|---|---|---|
| Prediction of pHa < 7.20 or Apgar ≤ 7 | ||||||
| CAP-C | 0.779 (0.774, 0.784) | 0.619 (0.597, 0.641) | 0.753 (0.657, 0.849) | 0.608 (0.489, 0.727) | 0.556 (0.505, 0.606) | 0.828 (0.795, 0.861) |
| CAP-T | 0.779 (0.776, 0.783) | 0.632 (0.621, 0.643) | 0.829 (0.776, 0.881) | 0.534 (0.446, 0.622) | 0.519 (0.488, 0.550) | 0.855 (0.829, 0.881) |
| CAP-L | 0.789 (0.783, 0.794) | 0.651 (0.646, 0.656) | 0.807 (0.768, 0.846) | 0.608 (0.557, 0.659) | 0.550 (0.530, 0.570) | 0.848 (0.830, 0.867) |
| CAP-CfC | 0.757 (0.750, 0.765) | 0.618 (0.604, 0.633) | 0.774 (0.739, 0.809) | 0.571 (0.505, 0.637) | 0.520 (0.492, 0.547) | 0.813 (0.800, 0.826) |
| Prediction of pHa < 7.15 or Apgar ≤ 7 | ||||||
| CAP-C | 0.771 (0.765, 0.778) | 0.460 (0.444, 0.477) | 0.751 (0.677, 0.825) | 0.640 (0.554, 0.726) | 0.343 (0.310, 0.376) | 0.923 (0.909, 0.937) |
| CAP-T | 0.792 (0.788, 0.796) | 0.430 (0.407, 0.454) | 0.884 (0.832, 0.936) | 0.468 (0.373, 0.564) | 0.288 (0.262, 0.315) | 0.953 (0.940, 0.967) |
| CAP-L | 0.800 (0.791, 0.808) | 0.487 (0.464, 0.510) | 0.776 (0.702, 0.850) | 0.662 (0.574, 0.749) | 0.366 (0.330, 0.402) | 0.933 (0.917, 0.948) |
| CAP-CfC | 0.801 (0.794, 0.807) | 0.497 (0.471, 0.523) | 0.725 (0.656, 0.794) | 0.711 (0.618, 0.804) | 0.390 (0.354, 0.427) | 0.922 (0.909, 0.935) |
| Prediction of pHa < 7.10 or Apgar ≤ 7 | ||||||
| CAP-C | 0.790 (0.785, 0.795) | 0.319 (0.302, 0.335) | 0.845 (0.807, 0.883) | 0.568 (0.511, 0.624) | 0.198 (0.183, 0.212) | 0.969 (0.963, 0.974) |
| CAP-T | 0.808 (0.799, 0.817) | 0.328 (0.300, 0.356) | 0.867 (0.827, 0.907) | 0.561 (0.471, 0.652) | 0.205 (0.182, 0.227) | 0.974 (0.967, 0.981) |
| CAP-L | 0.825 (0.819, 0.831) | 0.361 (0.329, 0.393) | 0.837 (0.784, 0.891) | 0.639 (0.562, 0.716) | 0.234 (0.205, 0.264) | 0.972 (0.964, 0.979) |
| CAP-CfC | 0.829 (0.823, 0.834) | 0.401 (0.373, 0.430) | 0.735 (0.658, 0.813) | 0.751 (0.685, 0.817) | 0.290 (0.247, 0.332) | 0.960 (0.953, 0.968) |
AUROC area under the receiver operating characteristic curve, PPV positive predictive value, NPV negative predictive value, CAP-C architecture with convolutional neural network, CAP-T architecture with Transformer, CAP-L architecture with long short-term memory, CAP-CfC architecture with closed-form continuous-time neural networks
In the AI-human comparison, the CAP models predicting fetal hypoxia with grade 1 acidemia (pHa < 7.20) were compared with human experts. All CAP models achieved higher AUROC (CAP-C: 0.779, CAP-T: 0.779, CAP-L: 0.789, CAP-CfC: 0.757) than human interpreters (0.715), examining by Delong test (P values: CAP-C: 0.002, CAP-T: < 0.001, CAP-L: < 0.001, CAP-CfC: 0.026). Figure 3a–d shows the difference of AUROC between CAPs and human interpreters. CAPs also achieved higher F1 scores (CAPs: 0.598–0.651 vs. human: 0.550) and sensitivity (CAPs: 0.753–0.872 vs. human: 0.512) than human interpreters (Fig. 3e, f). On the other hand, human interpreters had better specificity than deep learning algorithms (CAPs: 0.384–0.608 vs. human: 0.797, Fig. 3g).
Fig. 3.
Comparison of AUROC, F1 scores, sensitivity, and specificity between CAP models and human experts in nationwide AI-human comparison in CTG interpretation. The performance of all the CAP models was compared with human experts. a–d All CAP models achieve higher AUROC than human interpreters in the AI-human comparison and CAP-L achieve highest AUROC among all CAPs. a CAP-C vs. human; b CAP-T vs. human; c CAP-L vs. human; d CAP-CfC vs. human. CAP models achieve higher F1 scores (e) and sensitivity (f) but lower specificity (g) than human experts
The deep learning models analyzing only the fetal heart rate signal, or CTG traces of the last 20 min also had higher AUROC, F1 scores, sensitivity, but lower specificity than human interpreters (Additional file 3: Tables S7 and S8).
External validation of the deep learning algorithms
The CAP models were applied in the public CTU-UHB dataset for external validation. In predicting fetal hypoxia with grade 1 acidemia, CAP-T achieved highest AUROC (0.715, 95% CI: 0.710–0.721), followed by CAP-L (0.709, 95% CI: 0.703–0.714). CAP-L achieved highest AUROC in the other two grades (grade 2: 0.727, 95% CI: 0.718–0.736; grade 3: 0.730, 95% CI: 0.723–0.738) (Table 4). In most architectures, models using only FHR signal achieved even higher AUROC than the primary model, partly due to the lack of UC signal in some cases in the CTU-UHB dataset (Additional file 3: Table S9). The performances of CAPs analyzing the last 20 min CTG traces of the CTU-UHB dataset were shown in Additional file 3: Table S10.
Table 4.
Performance of Cardiotocography Artificial-intelligence Predictors (CAPs) in analyzing CTU-UHB dataset (mean, 95% CI)
| Model | AUROC | F1 score | Sensitivity | Specificity | PPV | NPV |
|---|---|---|---|---|---|---|
| Prediction of pHa < 7.20 or Apgar ≤ 7 | ||||||
| CAP-C | 0.634 (0.623, 0.646) | 0.441 (0.382, 0.500) | 0.438 (0.313, 0.563) | 0.720 (0.602, 0.837) | 0.532 (0.489, 0.575) | 0.678 (0.662, 0.695) |
| CAP-T | 0.715 (0.710, 0.721) | 0.579 (0.558, 0.600) | 0.661 (0.593, 0.730) | 0.616 (0.552, 0.681) | 0.525 (0.509, 0.540) | 0.750 (0.729, 0.772) |
| CAP-L | 0.709 (0.703, 0.714) | 0.595 (0.579, 0.611) | 0.666 (0.612, 0.720) | 0.644 (0.592, 0.697) | 0.546 (0.527, 0.564) | 0.758 (0.742, 0.775) |
| CAP-CfC | 0.671 (0.658, 0.684) | 0.546 (0.529, 0.563) | 0.576 (0.526, 0.626) | 0.671 (0.620, 0.723) | 0.530 (0.508, 0.552) | 0.717 (0.707, 0.728) |
| Prediction of pHa < 7.15 or Apgar ≤ 7 | ||||||
| CAP-C | 0.712 (0.704, 0.719) | 0.522 (0.510, 0.535) | 0.685 (0.609, 0.760) | 0.623 (0.543, 0.704) | 0.438 (0.406, 0.469) | 0.836 (0.821, 0.851) |
| CAP-T | 0.725 (0.717, 0.734) | 0.519 (0.507, 0.531) | 0.782 (0.725, 0.839) | 0.498 (0.411, 0.585) | 0.395 (0.369, 0.421) | 0.855 (0.841, 0.868) |
| CAP-L | 0.727 (0.718, 0.736) | 0.520 (0.507, 0.533) | 0.701 (0.614, 0.789) | 0.596 (0.491, 0.701) | 0.432 (0.395, 0.469) | 0.841 (0.821, 0.861) |
| CAP-CfC | 0.712 (0.706, 0.718) | 0.521 (0.502, 0.540) | 0.602 (0.527, 0.678) | 0.717 (0.640, 0.794) | 0.480 (0.445, 0.516) | 0.821 (0.806, 0.836) |
| Prediction of pHa < 7.10 or Apgar ≤ 7 | ||||||
| CAP-C | 0.726 (0.722, 0.730) | 0.469 (0.456, 0.482) | 0.761 (0.715, 0.808) | 0.542 (0.471, 0.613) | 0.343 (0.322, 0.364) | 0.884 (0.875, 0.893) |
| CAP-T | 0.724 (0.712, 0.736) | 0.462 (0.444, 0.481) | 0.747 (0.682, 0.812) | 0.538 (0.444, 0.633) | 0.342 (0.313, 0.371) | 0.880 (0.866, 0.894) |
| CAP-L | 0.730 (0.723, 0.738) | 0.474 (0.457, 0.490) | 0.737 (0.665, 0.809) | 0.570 (0.469, 0.671) | 0.360 (0.325, 0.395) | 0.883 (0.870, 0.896) |
| CAP-CfC | 0.724 (0.716, 0.731) | 0.492 (0.480, 0.503) | 0.615 (0.558, 0.673) | 0.727 (0.667, 0.788) | 0.423 (0.386, 0.460) | 0.863 (0.854, 0.871) |
AUROC area under the receiver operating characteristic curve. PPV positive predictive value, NPV negative predictive value, CAP-C architecture with convolutional neural network, CAP-T architecture with Transformer, CAP-L architecture with long short-term memory, CAP-CfC architecture with closed-form continuous-time neural networks
Using the CTU-UHB dataset where all cases have both pHa and Apgar scores, we compared model performance when the outcome was defined by acidemia alone (pHa < 7.20) versus Apgar alone (Apgar ≤ 7). The pre-trained CAP models (using 40-min signals) were applied to this dataset. Overall, models achieved higher AUROC for acidemia-based prediction. The best-performing models were CAP-T (FHR only) for acidemia (AUROC: 0.728, 95% CI: 0.725–0.732) and CAP-L (FHR only) for Apgar-based prediction (AUROC: 0.718, 95% CI: 0.711–0.726) (Additional file 3: Table S11).
Further, we have conducted a comparison between CAP models and expert annotations [55] in the CTU-UHB dataset. According to the original study design, the outcome was defined as pHa ≤ 7.15. Thus, we employed the pre‑trained CAP models (which was developed to predict pHa < 7.15 or Apgar ≤ 7) to analyze this outcome and compared with the nine experts. As shown in Additional file 3: Table S12, all CAP models achieved higher F1 score (mean range: 0.402–0.488 vs. 0.278), sensitivity (mean range: 0.595–0.799 vs. 0.441), PPV (mean range: 0.288–0.410 vs. 0.236), and NPV (mean range: 0.862–0.904 vs. 0.855) than the experts in the interpretations of CTU‑UHB data. CAP-C (0.714), CAP-T (0.734), and CAP-CfC (0.748) models using only FHR data, and CAP-CfC models using FHR and UC (0.693) even had higher specificity than the experts (0.668).
Explainability of deep learning algorithms
To achieve model explainability, we employed gradient-weighted class activation mapping (Grad-CAM) to show the focusing period of the models. The saliency map showed that, in the CAP-L model, the LSTM layer primarily highlight prolonged decelerations to predict fetal hypoxia (Fig. 4a). In the CAP-CfC model, the CfC layer focused on both variable and prolonged decelerations (Fig. 4b). However, the saliency maps of the CAP-C and CAP-T models did not highlight clinically interpretable features within the CTG traces (Fig. 4c, d).
Fig. 4.
Gradient-weighted class activation mapping (Grad-CAM) of the CAP models. Saliency map of the regions of CTG signal contributed to the prediction of fetal hypoxia. The red region shows the gradient of prediction of fetal hypoxia with respect to the CTG signal, where darker regions representing greater salience. a In the LSTM layer 1 of the CAP-L model, the region of prolonged deceleration showed greater salience. b In the CfC layer of the CAP-CfC model, the region of prolonged deceleration and variable decelerations showed greater salience. c The CAP-C model did not highlight clinically interpretable features within the CTG traces. d The CAP-T modes did not highlight clinically interpretable features within the CTG traces
To quantitatively validate the faithfulness of the Grad-CAM results, we performed a perturbation-based analysis. The principle of this faithfulness test is that removing truly important features should significantly degrade model performance. The mean probability drop was significantly greater when perturbing the top 10% critical regions, compared with perturbing the bottom 10% non-critical regions (CAP-L: 0.053 ± 0.071 vs. 0.021 ± 0.038, P < 0.001; CAP-CfC: 0.111 ± 0.093 vs. 0.004 ± 0.023, P < 0.001). These results strongly support the faithfulness of the Grad-CAM and the explainability of the models.
Web application of the Cardiotocography Artificial-intelligence Predictors
To bridge the gap between algorithm development and practical utility, an interactive web application (https://www.ctgaipredictor.online) was developed. The CAP-L variant was selected for deployment based on its superior and balanced performance across internal test, AI-human comparison, external validation, and model explainability. It uses 40 min of FHR and uterine contraction signals to predict fetal hypoxia with grade 1 acidemia. Users can upload CTG recordings (1 Hz sampling rate) in the specified format and the application will display the uploaded tracing alongside the model’s predicted probability of fetal hypoxia (see instructions in Additional file 6).
Discussion
In this study, we developed a series of deep learning models to predict fetal hypoxia based on CTG traces using a large dataset from a multicenter cohort and conducted a nationwide AI-human comparison to evaluate the performances of deep learning models. The models developed in this study demonstrated superior performance in identifying fetal hypoxia compared with trained human interpreters and showed promising generalizability in external validation. Notably, model interpretability analyses revealed that the CAP architectures mainly leveraged variable and prolonged decelerations of the CTG traces to predict fetal hypoxia.
Interpretation of CTG remains challenging in clinical practice [56, 57], and several deep learning algorithms have been developed to analyze CTG [33, 34, 36–38, 58]. Trained in the Oxford dataset of over 35,000 CTG traces, the CNN algorithms predict fetal acidemia with an AUC of 0.68 [33], which was later improved to 0.77 by the MCNN algorithms [34]. Another study showed conventional algorithms with AUC of 0.73 [37]. McCoy et al. trained six algorithms on 10,182 traces and achieved AUC of 0.75 at pHa threshold of 7.20 [36]. More recently, DeepCTG® 2.0 was reported to achieved AUC of 0.72 to 0.74 [38]. The CAP models developed in this study achieve promising accuracy in CTG interpretation in comparison with previous models and the experts in the AI-human comparison.
The primary challenge in clinical application of AI technologies is to determine whether AI models outperform human experts in real-world clinical settings, and how they can be integrated into decision-making. While prior models have been compared with documented decisions [34, 38] or small expert panels [39], the strength of CAP algorithms was identified by comparing with a large cohort of trained clinicians and midwives through the nationwide AI-human comparison. Moreover, result of the AI-human comparison provided a strategy for AI-human cooperation, that AI may serve as a sensitive screening tool, and human interpreters can provide final clinical judgment based on the AI-generated outputs.
Another challenge in application of AI is the lack of model explainability [59, 60]. Clinical practice may judge an AI-based predictor not only by its precision, but also by how the model make decision. In this view, the Grad-CAM method and the saliency maps provide the opportunity to clarify the underlining mechanism how AI models make the prediction. Grad-CAM method analyzes the gradient of the target layer of a deep learning model and reflect which neurons in the model have greater contribution in prediction [54]. Saliency map shows the Grad-CAM result by highlighting the temporal regions of the time series signal that most influenced the prediction, enabling human experts to determine whether the highlighting region represent clinically significant features. Grad-CAM and saliency map has been used in medical image analysis [61]. In signal processing, Grad-CAM and saliency map has been used in the study of electrocardiogram [62].
In this study, we employed Grad-CAM to elucidate how CAP architectures interpret CTG signals. Our findings revealed that the CAP-L model highlighted prolonged decelerations, and CAP-CfC models highlighted variable and prolonged decelerations, features that are pathophysiologically meaningful and clinically relevant. In CTG guidelines [45, 46, 52], variable and prolonged decelerations are critical features indicating risk of fetal hypoxia. Therefore, this finding enhances model transparency and supports their potential integration into clinical decision-making processes.
Algorithm selection and input configuration are critical in AI-based CTG interpretation [63]. Prior studies mainly employed CNN architecture [33, 34, 37] and varied in whether uterine contraction (UC) signals were included [36, 58]. In this study, we systematically evaluated how different algorithms (CNN, Transformer [49], LSTM, and CfC [50]), input strategies (FHR versus FHR and UC), and temporal windows (20 vs. 40 min) impact the prediction results.
For algorithms, CNN was chosen to extract multidimensional features from the overall CTG data, Transformer was applied to capture temporal patterns of CTG traces with multi-head attention mechanisms [49], LSTM served as a classical time series analysis model with long- and short-term memory capabilities, and closed-form continuous-time neural networks (CfC) is novel liquid neural network-based model for time series data [50]. In the present study, the CAP-L model showed solid and robust performance among internal test, AI-human comparison, and the external validation, and demonstrated superior performance of explainability. Further research is needed to evaluate the performance of CAP algorithms in a prospective, real-time clinical settings, and to establish whether and how cumulative or trend-based predictions from the model can be integrated into a standardized clinical workflow to improve outcomes.
This study makes several distinct contributions to AI-based CTG interpretation. First, it provides a direct and prospective validation of deep learning algorithms by benchmarking them against interpretations from a broad, nationwide cohort of clinicians, offering a rigorous clinical performance assessment. Second, to our knowledge, it is the first application of Grad-CAM technology with faithfulness test to achieve model explainability in CTG interpretation. Furthermore, we introduced the novel closed-form continuous-time neural network (CfC) architecture into fetal monitoring. Ultimately, by implementing the CAP model as a publicly available web application with built-in input data assessment, we translated deep learning algorithm into a resource for clinical interaction. In the context of previous studies, our work substantially broadens both the methodological depth and clinical relevance of AI-based CTG interpretation.
This study has several strengths. First, the large dataset derived from multicenter real-world clinical cases facilitates the training of deep learning models with improved robustness. Second, the comprehensive exploration of different algorithms and input combinations enabled us to select the model with precision and achieve optimal results. Third, the nationwide AI-human comparison identified the performance of deep learning algorithms in real clinical settings. By evaluating newly developed AI models against large cohort of experts, this study introduces a novel and scalable framework for evaluating AI performance in clinical practice. Fourth, the application of Grad-CAM enhanced the explainability of the deep learning models.
The present study also has some limitations. First, the primary outcome combines Apgar scores with pHa, which, though clinically aligned, may reduce specificity. Second, a fundamental asymmetry existed in the decision-making processes between AI and human experts. The AI makes a direct data-driven prediction, whereas clinicians rely on visual pattern recognition based on established grading criteria. Third, the online survey design, despite its nationwide scale, is susceptible to biases from random dissemination and participant self-selection. Finally, the enrichment strategy used to select CTG cases for the human-AI benchmark, while necessary for a balanced comparison, may affect the model generalizability in an unselected population.
Conclusions
The CAP deep learning algorithms developed in this study effectively leverage variable and prolonged decelerations of CTG trace to identify fetal hypoxia and demonstrated superior performance compared to human interpreters in the nationwide AI-human comparison. These findings highlight the clinical application of AI-based CTG interpretation.
Supplementary Information
Additional file 1: Fig. S1 The flow chart of the dataset generation and nationwide AI-human comparison.
Additional file 2: Fig. S2 The structure of the Cardiotocography Artificial-intelligence Predictors (CAPs).
Additional file 3: Tables S1–S12. Tables S1–S4 Summary of the CAP-C, CAP-T, CAP-L, and CAP-CfC architectures. Table S5 Performance of CAPs in analyzing the fetal heart rate signal in the testing set. Table S6 Performance of CAPs in analyzing CTG of the last 20 min in the testing set. Table S7 Performance of CAPs in analyzing the fetal heart rate signal in the AI-human comparison. Table S8 Performance of CAPs in analyzing CTG of the last 20 min in the AI-human comparison. Table S9 Performance of CAPs in analyzing the fetal heart rate signal of the CTU-UHB dataset. Table S10 Performance of CAPs in analyzing CTG of the last 20 min in the CTU-UHB dataset. Table S11 Performance of CAPs in analyzing CTU-UHB dataset in subgroups. Table S12 Comparison of CAPs with expert annotations in the CTU-UHB dataset.
Additional file 4: Fig. S3 The questionnaire format and the procedure of the online survey. (a) Interface for the three-category CTG classification; (b) interface for hypoxia prediction; (c) interface displaying the judgment result and feedback; (d) interface displaying the final test score.
Additional file 5: Questions in the online survey.
Additional file 6: Instructions for using the web application.
Acknowledgements
We appreciate the National Supercomputer Center in Guangzhou, China to provide supercomputer. We thanked Professor KK Cheng from the University of Birmingham, and Dr. Guangyao Cai from Sun Yat-sen University Cancer Center for valuable comments to the manuscript.
Abbreviations
- CTG
Cardiotocography
- AI
Artificial intelligence
- CNN
Convolutional neural network
- Grad-CAM
Gradient-weighted class activation mapping
- pHa
PH of umbilical artery
- FHR
Fetal heart rate
- UC
Uterine contraction
- CAP
Cardiotocography Artificial-intelligence Predictor
- LSTM
Long-term short-term memory
- CfC
Closed-form continuous-time neural networks
- AUROC
Area under the receiver operating characteristic curves
- PPV
Positive predictive value
- NPV
Negative predictive value
- CTU-UHB
Czech Technical University and University Hospital in Brno
Authors’ contributions
SL and BL designed the study. SL collected and pre-processed CTG records and built the online database. XD and MY provided clinical data in their institutes. HW and BL developed and trained the deep learning algorithms. ZS and RH assisted with the data analysis. NW and DW provided support during the training process with supercomputer. SL and BL analyzed the data and draft the manuscript. YX and ZW revised the manuscript. All authors had full access to all the data in the study. All authors read and approved the final manuscript.
Funding
This study was supported by the National Natural Science Foundation of China (No. 82371689 and 81771602) and Noncommunicable Chronic Diseases-National Science and Technology Major Project (No. 2024ZD0532100).
Data Availability
The CTG records and clinical data supporting this study are not publicly available due to data use agreement restrictions. De-identified CTG records and clinical data will be made available on reasonable request to the corresponding author.
Declarations
Ethics approval and consent to participate
The study was approved by the ethical committees of the First Affiliated Hospital of Sun Yat-sen University ([2020]278–1, [2021]431–1, [2022]451). For the retrospective analysis of anonymized CTG and clinical data, the requirement for informed consent was waived by the ethical committees. For the online survey, informed consent was obtained from all participants.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.McClure EM, Saleem S, Goudar SS, Tikmani SS, Dhaded SM, Hwang K, et al. The causes of stillbirths in south Asia: results from a prospective study in India and Pakistan (PURPOSe). Lancet Glob Health. 2022;10(7):e970–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Driscoll DJO, Felice VD, Kenny LC, Boylan GB, O’Keeffe GW. Mild prenatal hypoxia-ischemia leads to social deficits and central and peripheral inflammation in exposed offspring. Brain Behav Immun. 2018;69:418–27. [DOI] [PubMed] [Google Scholar]
- 3.Giussani DA. Breath of life: heart disease link to developmental hypoxia. Circulation. 2021;144(17):1429–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.O’Sullivan MP, Looney AM, Moloney GM, Finder M, Hallberg B, Clarke G, et al. Validation of altered umbilical cord blood microRNA expression in neonatal hypoxic-ischemic encephalopathy. JAMA Neurol. 2019;76(3):333–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Palanisamy A, Giri T, Jiang J, Bice A, Quirk JD, Conyers SB, et al. In utero exposure to transient ischemia-hypoxemia promotes long-term neurodevelopmental abnormalities in male rat offspring. JCI Insight. 2020;5(10):e133172. [DOI] [PMC free article] [PubMed]
- 6.Wilkinson LJ, Neal CS, Singh RR, Sparrow DB, Kurniawan ND, Ju A, et al. Renal developmental defects resulting from in utero hypoxia are associated with suppression of ureteric beta-catenin signaling. Kidney Int. 2015;87(5):975–83. [DOI] [PubMed] [Google Scholar]
- 7.Andersson CB, Klingenberg C, Thellesen L, Johnsen SP, Kesmodel US, Petersen JP. Umbilical cord pH levels and neonatal morbidity and mortality. JAMA Netw Open. 2024;7(8):e2427604. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Lovers A, Daumer M, Frasch MG, Ugwumadu A, Warrick P, Vullings R, et al. Advancements in fetal heart rate monitoring: a report on opportunities and strategic initiatives for better intrapartum care. BJOG. 2025;132(7):853–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Alfirevic Z, Devane D, Gyte GM, Cuthbert A. Continuous cardiotocography (CTG) as a form of electronic fetal monitoring (EFM) for fetal assessment during labour. Cochrane Database Syst Rev. 2017;2(2):CD006066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Devane D, Lalor JG, Daly S, McGuire W, Cuthbert A, Smith V. Cardiotocography versus intermittent auscultation of fetal heart on admission to labour ward for assessment of fetal wellbeing. Cochrane Database Syst Rev. 2017;1(1):CD005122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Evans MI, Britt DW, Evans SM, Devoe LD. Improving the interpretation of electronic fetal monitoring: the fetal reserve index. Am J Obstet Gynecol. 2023;228(5S):S1129–43. [DOI] [PubMed] [Google Scholar]
- 12.Antepartum fetal surveillance. ACOG practice bulletin, number 229. Obstet Gynecol. 2021;137(6):e116–27. [DOI] [PubMed] [Google Scholar]
- 13.Ayres-de-Campos D, Spong CY, Chandraharan E, Panel FIFMEC. FIGO consensus guidelines on intrapartum fetal monitoring: cardiotocography. Int J Gynaecol Obstet. 2015;131(1):13–24. [DOI] [PubMed] [Google Scholar]
- 14.National Institute for Health and Care Excellence: Guidelines. Intrapartum care for healthy women and babies. London: National Institute for Health and Care Excellence (NICE) Copyright © NICE 2023.; 2022. [Google Scholar]
- 15.Macones GA, Hankins GD, Spong CY, Hauth J, Moore T. The 2008 National Institute of Child Health and Human Development workshop report on electronic fetal monitoring: update on definitions, interpretation, and research guidelines. Obstet Gynecol. 2008;112(3):661–6. [DOI] [PubMed] [Google Scholar]
- 16.Sabiani L, Le Du R, Loundou A, d’Ercole C, Bretelle F, Boubli L, et al. Intra- and interobserver agreement among obstetric experts in court regarding the review of abnormal fetal heart rate tracings and obstetrical management. Am J Obstet Gynecol. 2015;213(6):856-e1-8. [DOI] [PubMed] [Google Scholar]
- 17.Blackwell SC, Grobman WA, Antoniewicz L, Hutchinson M, Gyamfi Bannerman C. Interobserver and intraobserver reliability of the NICHD 3-tier fetal heart rate interpretation system. Am J Obstet Gynecol. 2011;205(4):378-e1-5. [DOI] [PubMed] [Google Scholar]
- 18.Hirsch E. Electronic fetal monitoring to prevent fetal brain injury: a ubiquitous yet flawed tool. JAMA. 2019;322(7):611–2. [DOI] [PubMed] [Google Scholar]
- 19.Group IC. Computerised interpretation of fetal heart rate during labour (INFANT): a randomised controlled trial. Lancet. 2017;389(10080):1719–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Nunes I, Ayres-de-Campos D, Ugwumadu A, Amin P, Banfield P, Nicoll A, et al. Central fetal monitoring with and without computer analysis: a randomized controlled trial. Obstet Gynecol. 2017;129(1):83–90. [DOI] [PubMed] [Google Scholar]
- 21.Cid YD, Macpherson M, Gervais-Andre L, Zhu Y, Franco G, Santeramo R, et al. Development and validation of open-source deep neural networks for comprehensive chest X-ray reading: a retrospective, multicentre study. Lancet Digit Health. 2024;6(1):e44–57. [DOI] [PubMed] [Google Scholar]
- 22.Ouyang D, Theurer J, Stein NR, Hughes JW, Elias P, He B, et al. Electrocardiographic deep learning for predicting post-procedural mortality: a model development and validation study. Lancet Digit Health. 2024;6(1):e70–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Beam AL, Drazen JM, Kohane IS, Leong TY, Manrai AK, Rubin EJ. Artificial intelligence in medicine. N Engl J Med. 2023;388(13):1220–1. [DOI] [PubMed] [Google Scholar]
- 24.Fung R, Villar J, Dashti A, Ismail LC, Staines-Urias E, Ohuma EO, et al. Achieving accurate estimates of fetal gestational age and personalised predictions of fetal growth based on data from an international prospective cohort study: a population-based machine learning study. Lancet Digit Health. 2020;2(7):e368–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Hu Y, Xu J, Huang L, Zheng Z, Zhao J, Chen T, et al. Artificial intelligence-assisted endoscopic diagnosis system for diagnosing Helicobacter pylori infection: a multicenter study. BMC Med. 2025;23(1):540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Zhou Z, Zhao Z, Zhang X, Zhang X, Jiao P, Ye X. Identifying fetal status with fetal heart rate: deep learning approach based on long convolution. Comput Biol Med. 2023;159:106970. [DOI] [PubMed] [Google Scholar]
- 27.Clapp MA, Li S, James KE, Reiff ES, Little SE, McCoy TH, et al. Development of a practical prediction model for adverse neonatal outcomes at the start of the second stage of labor. Obstet Gynecol. 2025;145(1):73–81. [DOI] [PubMed] [Google Scholar]
- 28.Spairani E, Daniele B, Signorini MG, Magenes G. A deep learning mixed-data type approach for the classification of FHR signals. Front Bioeng Biotechnol. 2022;10:887549. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Melaet R, de Vries IR, Kok RD, Guid Oei S, Huijben IAM, van Sloun RJG, et al. Artificial intelligence based cardiotocogram assessment during labor. Eur J Obstet Gynecol Reprod Biol. 2024;295:75–85. [DOI] [PubMed] [Google Scholar]
- 30.Mendis L, Palaniswami M, Brownfoot F, Keenan E. Computerised cardiotocography analysis for the automated detection of fetal compromise during labour: a review. Bioengineering. 2023;10(9):1007. [DOI] [PMC free article] [PubMed]
- 31.Ben M’Barek I, Jauvion G, Ceccaldi PF. Computerized cardiotocography analysis during labor - a state-of-the-art review. Acta Obstet Gynecol Scand. 2023;102(2):130–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Jones GD, Cooke WR, Vatish M, Redman CWG. Computerized analysis of antepartum cardiotocography: a review. Matern Fetal Med. 2022;4(2):130–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Petrozziello A, Jordanov I, Aris Papageorghiou T, Christopher Redman WG, Georgieva A. Deep learning for continuous electronic fetal monitoring in labor. Annu Int Conf IEEE Eng Med Biol Soc. 2018;2018:5866–9. [DOI] [PubMed] [Google Scholar]
- 34.Petrozziello A, Redman CWG, Papageorghiou AT, Jordanov I, Georgieva A. Multimodal convolutional neural networks to detect fetal compromise during labor and delivery. IEEE Access. 2019;7:112026–36. [Google Scholar]
- 35.Cömert Z, Kocamaz AF, editors. Fetal hypoxia detection based on deep convolutional neural network with transfer learning approach. Software engineering and algorithms in intelligent systems; 2019 2019//; Cham: Springer International Publishing.
- 36.McCoy JA, Levine LD, Wan G, Chivers C, Teel J, La Cava WG. Intrapartum electronic fetal heart rate monitoring to predict acidemia at birth with the use of deep learning. Am J Obstet Gynecol. 2024. [DOI] [PMC free article] [PubMed]
- 37.Ogasawara J, Ikenoue S, Yamamoto H, Sato M, Kasuga Y, Mitsukura Y, et al. Deep neural network-based classification of cardiotocograms outperformed conventional algorithms. Sci Rep. 2021;11(1):13367. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Ben M’Barek I, Jauvion G, Merrer J, Koskas M, Sibony O, Ceccaldi PF, et al. DeepCTG(R) 2.0: development and validation of a deep learning model to detect neonatal acidemia from cardiotocography during labor. Comput Biol Med. 2025;184:109448. [DOI] [PubMed]
- 39.Ben M’Barek I, Jauvion G, Vitrou J, Holmstrom E, Koskas M, Ceccaldi PF. DeepCTG(R) 1.0: an interpretable model to detect fetal hypoxia from cardiotocography data during labor and delivery. Front Pediatr. 2023;11:1190441. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Malin GL, Morris RK, Khan KS. Strength of association between umbilical cord pH and perinatal and long term outcomes: systematic review and meta-analysis. BMJ. 2010;340:c1471. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Sabol BA, Caughey AB. Acidemia in neonates with a 5-minute Apgar score of 7 or greater - what are the outcomes? Am J Obstet Gynecol. 2016;215(4):486.e1-6. [DOI] [PubMed] [Google Scholar]
- 43.Neonatal Resuscitation Group of Chinese Society of Perinatal Medicine. Expert consensus on the diagnosis of neonatal asphyxia. Chin J Perinat Med. 2016;9(1):3–6.
- 44.Ayres-de-Campos D, Arulkumaran S, Panel FIFMEC. FIGO consensus guidelines on intrapartum fetal monitoring: physiology of fetal oxygenation and the main goals of intrapartum fetal monitoring. Int J Gynaecol Obstet. 2015;131(1):5–8. [DOI] [PubMed]
- 45.ACOG clinical practice guideline no. 10: intrapartum fetal heart rate monitoring: interpretation and management. Obstet Gynecol. 2025;146(4):583–99. [DOI] [PubMed] [Google Scholar]
- 46.Dore S, Ehman W. No. 396-fetal health surveillance: intrapartum consensus guideline. J Obstet Gynaecol Can. 2020;42(3):316-48 e9. [DOI] [PubMed] [Google Scholar]
- 47.Fritsch FN, Carlson RE. Monotone piecewise cubic interpolation. SIAM J Numer Anal. 1980;17(2):238–46. [Google Scholar]
- 48.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. arxiv:170603762[csCL,csLG].
- 50.Hasani R, Lechner M, Amini A, Liebenwein L, Ray A, Tschaikowski M, et al. Closed-form continuous-time neural networks. Nature Machine Intelligence. 2022;4(11):992–1003. [Google Scholar]
- 51.Johnson JM, Khoshgoftaar TM. Survey on deep learning with class imbalance. J Big Data. 2019;6(1):27. [Google Scholar]
- 52.Chinese Society of Perinatal Medicine. Expert consensus on application of electronic fetal heart monitoring (in Chinese). Chin J Perinat Med. 2015;18(7):486–90. [Google Scholar]
- 53.Chudacek V, Spilka J, Bursa M, Janku P, Hruban L, Huptych M, et al. Open access intrapartum CTG database. BMC Pregnancy Childbirth. 2014;14:16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D, editors. Grad-CAM: visual explanations from deep networks via gradient-based localization. 2017 IEEE International Conference on Computer Vision (ICCV); 2017 22–29 Oct. 2017.
- 55.Hruban L, Spilka J, Chudacek V, Janku P, Huptych M, Bursa M, et al. Agreement on intrapartum cardiotocogram recordings between expert obstetricians. J Eval Clin Pract. 2015;21(4):694–702. [DOI] [PubMed] [Google Scholar]
- 56.Caughey AB. Electronic fetal monitoring-imperfect but opportunities for improvement. JAMA Netw Open. 2020;3(2):e1921352. [DOI] [PubMed] [Google Scholar]
- 57.Beermann SE, Watkins VY, Frolova AI, Raghuraman N, Cahill AG. The relationship between maternal anemia and electronic fetal monitoring patterns. Am J Obstet Gynecol. 2023;229(4):449 e1-e6. [DOI] [PubMed] [Google Scholar]
- 58.Asfaw D, Jordanov I, Impey L, Namburete A, Lee R, Georgieva A. Multimodal deep learning for predicting adverse birth outcomes based on early labour data. Bioengineering (Basel). 2023;10(6):730. [DOI] [PMC free article] [PubMed]
- 59.Mesinovic M, Watkinson P, Zhu T. Explainability in the age of large language models for healthcare. Commun Eng. 2025;4(1):128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Ennab M, McHeick H. Enhancing interpretability and accuracy of AI models in healthcare: a comprehensive review on challenges and future directions. Front Robot AI. 2024;11:1444763. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Mala DJ, Chattopadhyay M, Mukhopadhyay P, Sinha R. Visualizing UNet decisions: an explainable AI perspective for brain MRI segmentation. IEEE Access. 2025;13:133869–81. [Google Scholar]
- 62.Khurshid S, Friedman S, Reeder C, Di Achille P, Diamant N, Singh P, et al. ECG-based deep learning and clinical risk factors to predict atrial fibrillation. Circulation. 2022;145(2):122–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Ahmed SF, Alam MSB, Hassan M, Rozbu MR, Ishtiak T, Rafa N, et al. Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artif Intell Rev. 2023;56(11):13521–617. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Additional file 1: Fig. S1 The flow chart of the dataset generation and nationwide AI-human comparison.
Additional file 2: Fig. S2 The structure of the Cardiotocography Artificial-intelligence Predictors (CAPs).
Additional file 3: Tables S1–S12. Tables S1–S4 Summary of the CAP-C, CAP-T, CAP-L, and CAP-CfC architectures. Table S5 Performance of CAPs in analyzing the fetal heart rate signal in the testing set. Table S6 Performance of CAPs in analyzing CTG of the last 20 min in the testing set. Table S7 Performance of CAPs in analyzing the fetal heart rate signal in the AI-human comparison. Table S8 Performance of CAPs in analyzing CTG of the last 20 min in the AI-human comparison. Table S9 Performance of CAPs in analyzing the fetal heart rate signal of the CTU-UHB dataset. Table S10 Performance of CAPs in analyzing CTG of the last 20 min in the CTU-UHB dataset. Table S11 Performance of CAPs in analyzing CTU-UHB dataset in subgroups. Table S12 Comparison of CAPs with expert annotations in the CTU-UHB dataset.
Additional file 4: Fig. S3 The questionnaire format and the procedure of the online survey. (a) Interface for the three-category CTG classification; (b) interface for hypoxia prediction; (c) interface displaying the judgment result and feedback; (d) interface displaying the final test score.
Additional file 5: Questions in the online survey.
Additional file 6: Instructions for using the web application.
Data Availability Statement
The CTG records and clinical data supporting this study are not publicly available due to data use agreement restrictions. De-identified CTG records and clinical data will be made available on reasonable request to the corresponding author.




