Skip to main content
BMC Medical Imaging logoLink to BMC Medical Imaging
. 2026 Jan 29;26:107. doi: 10.1186/s12880-026-02192-8

Multimodal deep learning using preoperative CT and ultrasound for recurrence risk prediction in high-grade serous ovarian carcinoma

Silin Nie 1,#, Yumin Jiang 2,#, Yanan Duan 1, Yulu Han 3, Danni Jiang 4, Aiping Chen 1,✉, Huijun Chu 1,✉
PMCID: PMC12924501  PMID: 41612257

Abstract

Background

High-grade serous ovarian carcinoma (HGSOC) is associated with a high risk of postoperative recurrence, and the timing of recurrence is closely related to patient survival outcomes and subsequent treatment strategies. However, current postoperative surveillance primarily relies on clinical indicators and routine imaging examinations, which offer limited predictive accuracy. Although deep learning–based image analysis has shown promise in capturing tumor heterogeneity and improving prognostic assessment, studies integrating preoperative contrast-enhanced computed tomography (CE-CT) and ultrasound imaging for recurrence risk prediction in HGSOC remain limited.

Methods

This single-center retrospective study enrolled 293 patients with pathologically confirmed HGSOC, who were randomly assigned to a training cohort and an internal validation cohort. Separate two-dimensional deep learning survival models were developed using preoperative CE-CT and ultrasound images. A Cox partial likelihood–based time-to-event loss function was applied to generate modality-specific deep learning scores (DL-scores). Independent clinical predictors were identified through univariate and multivariate Cox regression analyses and were integrated with CT-DL and US-DL scores to construct a multimodal Cox prognostic model, which was visualized using a nomogram. Model performance was assessed using the concordance index (C-index), time-dependent receiver operating characteristic (ROC) curves, Kaplan–Meier survival analysis, calibration curves, and decision curve analysis. Bootstrap resampling was performed to evaluate model robustness, and Gradient-weighted Class Activation Mapping (Grad-CAM) was used to visualize model attention.

Results

The multimodal integrated model demonstrated superior prognostic performance, achieving C-index values of 0.840 in the training cohort and 0.722 in the internal validation cohort. Time-dependent ROC analysis showed that the combined model achieved areas under the curve (AUCs) of 0.880, 0.890, and 0.892 for predicting 1-, 2-, and 3-year recurrence-free survival (RFS) in the training cohort, and 0.864, 0.781, and 0.772 in the validation cohort, respectively. Kaplan–Meier survival analysis revealed a significant separation between high- and low-risk groups (log-rank test, all P < 0.001). Calibration and decision curve analyses indicated good agreement between predicted and observed outcomes and a higher net clinical benefit. Bootstrap resampling further confirmed the robustness of the multimodal model. Grad-CAM visualizations suggested that the model primarily focused on tumor and peritumoral regions.

Conclusions

This study proposes a multimodal prognostic framework integrating deep learning features from CE-CT and ultrasound with clinical variables for preoperative prediction of postoperative recurrence risk in HGSOC. Using routinely acquired imaging data, the model shows consistent interpretability and performance, supporting its use as an exploratory decision-support and risk-stratification tool. Prospective, multicenter validation is required before clinical implementation.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12880-026-02192-8.

Keywords: CE-CT, US, HGSOC, Deep learning, Multimodal imaging, Recurrence-free survival, Prognostic model

Background

Ovarian cancer (OC) remains the most lethal gynecologic malignancy worldwide [1]. Among its histological subtypes, high-grade serous ovarian carcinoma (HGSOC) is the most prevalent and biologically aggressive form [2, 3]. Standard treatment for ovarian cancer consists of comprehensive staging surgery or primary cytoreductive surgery followed by platinum-based chemotherapy. For patients who are unlikely to achieve optimal cytoreduction, neoadjuvant chemotherapy followed by interval debulking surgery represents an alternative therapeutic strategy. Based on individual clinical and molecular characteristics, eligible patients are recommended to receive maintenance therapy with bevacizumab and/or poly(ADP-ribose) polymerase (PARP) inhibitors [4, 5]. Although standardized first-line treatment strategies have been widely implemented and approximately 80% of patients with HGSOC achieve clinical remission after initial therapy [6], the majority ultimately experience disease recurrence. The recurrence rate is approximately 25% in early-stage ovarian cancer (FIGO stages I–II) and exceeds 80% in advanced-stage disease (FIGO stages III–IV) [7]. Recurrence represents a major cause of mortality in ovarian cancer [8, 9], and early recurrence is strongly associated with unfavorable survival outcomes. Importantly, the timing of recurrence plays a critical role in guiding subsequent treatment decisions [7, 10, 11], including the reintroduction of platinum-based chemotherapy and adjustments to maintenance strategies and surveillance intensity. Collectively, these observations underscore a pressing need to develop reliable methods for accurately predicting recurrence risk in patients with HGSOC.

At present, recurrence surveillance primarily relies on clinical symptoms, serum CA-125 levels, and routine imaging examinations; however, these approaches are limited by suboptimal sensitivity and specificity and the absence of objective, quantitative biomarkers [12, 13]. In this context, deep learning (DL) has emerged as a transformative technology in medical imaging, facilitating a shift from qualitative image interpretation toward high-dimensional quantitative analysis and predictive modeling. By automatically extracting hierarchical, task-specific features, DL can capture intratumoral heterogeneity, tumor microenvironment–related information, and subtle morphological patterns beyond human visual perception, thereby demonstrating superior representational capacity and generalizability compared with traditional handcrafted radiomics approaches [14–17]. In parallel, accumulating evidence highlights the potential of artificial intelligence–driven multimodal prognostic models to integrate imaging data with clinical variables and molecular information, ultimately improving outcome prediction [18, 19]. In particular, Sadeghi et al. systematically reviewed deep learning–based multimodal imaging approaches for enhancing diagnostic accuracy in ovarian cancer, demonstrating the methodological feasibility and potential advantages of multimodal image integration and providing a valuable foundation for extending these strategies to prognostic modeling [20].

Among available imaging modalities, contrast-enhanced computed tomography (CE-CT) and ultrasound (US) remain central to the evaluation of ovarian cancer [21]. CE-CT enables comprehensive assessment of tumor burden, disease dissemination, and surgical resectability [22], whereas US provides noninvasive, real-time visualization of pelvic soft tissues and vascular characteristics [23]. Owing to their distinct imaging mechanisms, these modalities offer complementary macroscopic and microscopic insights into tumor biology. Recent studies have demonstrated the feasibility of CE-CT– or US-based DL models for predicting survival and recurrence outcomes in epithelial ovarian cancer, as well as for accurately classifying adnexal masses in multicenter cohorts [24–27]. Despite the growing interest in multimodal deep learning approaches, no previous study has integrated preoperative CE-CT and US within a unified DL framework to predict postoperative recurrence risk in patients with HGSOC undergoing primary cytoreductive surgery.

Accordingly, the present study aimed to develop and validate a multimodal deep learning model that integrates preoperative CE-CT and US imaging features with clinical parameters to predict postoperative recurrence risk in patients with HGSOC.

Method

Patients

This retrospective study was approved by the Ethics Committee of the Affiliated Hospital of Qingdao University and was conducted in accordance with the principles of the Declaration of Helsinki. Given the retrospective nature of the study, the requirement for written informed consent was waived, and patient confidentiality was strictly protected throughout the study.

Patients treated at our institution between January 2014 and September 2024 with pathologically confirmed HGSOC were retrospectively enrolled. The inclusion criteria were as follows: (1) histopathological confirmation of HGSOC and receipt of standard surgical staging and cytoreductive surgery; (2) administration of postoperative platinum-based chemotherapy; (3) completion of both CE-CT and US examinations within one month prior to initial therapy; and (4) availability of complete follow-up data. The exclusion criteria were as follows: (1) coexistence of other malignant tumors; (2) a follow-up duration of less than 6 months or incomplete follow-up information; and (3) poor-quality CE-CT or US images, or inadequate delineation of the region of interest (ROI). Ultimately, 293 eligible patients were included and randomly allocated to a training cohort (n = 204) and an internal validation cohort (n = 89) at a 7:3 ratio. The overall study workflow is illustrated in Fig. 1, and the detailed inclusion and exclusion process is presented in Fig. 2.

Fig. 1.

Fig. 1

Overall study framework and analysis workflow

Fig. 2.

Fig. 2

Patient selection flowchart

Follow-up and clinical endpoints

The follow-up cutoff date for this study was April 30, 2025. Follow-up was conducted in accordance with the National Comprehensive Cancer Network (NCCN) guidelines. After completion of initial therapy, patients were followed at intervals of 2–4 months during the first two years and 3–6 months during years 3–5. Follow-up was primarily performed through outpatient visits, with telephone contact used as a supplementary approach when necessary. At each visit, symptom assessment, comprehensive physical examination, and serial monitoring of tumor markers were routinely performed. In the presence of suspicious symptoms, abnormal clinical findings, or persistent elevation of tumor markers, CE-CT of the chest, abdomen, and pelvis was performed. Additional imaging examinations, including magnetic resonance imaging (MRI) or positron emission tomography–computed tomography (PET-CT), were conducted when clinically indicated to further assess disease recurrence.

Recurrence-free survival (RFS) was defined as the interval from the date of surgery to the first documented tumor recurrence or the end of follow-up, whichever occurred first. To ensure consistency of the analytical time window, a postoperative period of 3 years was predefined as the observation endpoint. For patients who experienced recurrence within 3 years after surgery, the actual RFS was recorded. For patients without recurrence, censoring was applied as follows: patients with a follow-up duration exceeding 3 years were censored at 3 years postoperatively, whereas those with a follow-up duration of less than 3 years were censored at the time of their last follow-up.

Ultrasound image acquisition and preprocessing

All patients underwent preoperative transvaginal gynecologic ultrasound examinations. For patients who were not suitable for transvaginal examination, transabdominal ultrasound was additionally performed. All examinations were conducted by experienced sonographers using ultrasound systems, including ACUSON Sequoia (Siemens Healthineers, USA). During image acquisition, grayscale B-mode images of the lesions were routinely obtained in both transverse and longitudinal planes, and the plane demonstrating the maximum tumor diameter was selected as the representative image. For each patient, representative images were manually reviewed by a physician with extensive experience in gynecologic ultrasound diagnosis, with preference given to images showing well-defined lesion boundaries, minimal artifacts, and complete visualization of tumor morphology. All selected images were exported in JPG format for subsequent analysis.

Prior to model development, a unified and standardized preprocessing pipeline was applied to all ultrasound images to minimize variability introduced by different operators and examination conditions. First, images were cropped to remove irrelevant background regions while preserving complete visualization of the lesion, thereby reducing the influence of surrounding normal tissues and noise on model learning. Subsequently, intensity normalization was applied to harmonize grayscale distribution differences caused by variations in imaging devices and acquisition parameters. To accommodate deep learning models pretrained on ImageNet, all ultrasound images were further standardized prior to network input using channel-wise means and standard deviations derived from the ImageNet dataset (means: [0.485, 0.456, 0.406]; standard deviations: [0.229, 0.224, 0.225]). For grayscale ultrasound images, single-channel images were replicated into three channels before normalization to meet model input requirements. This preprocessing pipeline was consistently applied during both the training and validation stages.

CT image acquisition, preprocessing, and tumor segmentation

All patients underwent preoperative pelvic contrast-enhanced CT examinations using multi-detector spiral CT scanners, including Siemens Somatom Definition Flash (Germany), Philips iCT 256 (the Netherlands), and GE Optima CT670 (USA). To ensure consistency of contrast enhancement and emphasize tumor vascular characteristics while minimizing heterogeneity introduced by multiphase imaging, only arterial-phase images were selected for subsequent analysis. All images were stored in DICOM format. Standardized acquisition parameters were applied across scanners to reduce inter-scanner variability.

All CT images were resampled to isotropic voxels with a spatial resolution of 1 mm × 1 mm × 1 mm using cubic B-spline interpolation. Fixed window width and window level settings were applied to all images to minimize manual intervention and ensure consistency of image intensity distributions across patients and scanners. Tumor segmentation was performed using ITK-SNAP software (version 3.8.0).

In this study, three-dimensional tumor segmentation was not intended for direct construction of three-dimensional deep learning models. Instead, it served as an auxiliary tool to objectively and standardly identify the axial slice containing the maximum tumor cross-sectional area. Specifically, two radiologists with more than five years of experience in gynecologic imaging manually delineated tumor boundaries on all axial slices to generate a three-dimensional volume of interest (VOI), with careful exclusion of adjacent blood vessels, bowel loops, and pelvic organs. In cases of disagreement, a senior radiologist with over twenty years of experience in gynecologic imaging reviewed the segmentations and reached a consensus. Based on the finalized VOI, the system automatically identified the axial slice with the largest tumor cross-sectional area and exported this slice as the representative two-dimensional CT image for subsequent two-dimensional deep learning analysis. All exported images were subsequently re-evaluated by the two radiologists to ensure complete lesion coverage and the absence of significant artifacts.

Deep learning model development and training

To predict postoperative recurrence risk in patients with ovarian cancer, two-dimensional deep learning–based survival models were developed. Multiple architectures based on convolutional neural networks (CNNs) and Transformer models were evaluated, including AlexNet, VGG, ResNet, DenseNet, GoogLeNet, and Vision Transformer (ViT). All models were initialized with ImageNet-pretrained weights. Input images were resized to the required input dimensions of each network and normalized using the ImageNet dataset mean and standard deviation. To improve model robustness, data augmentation strategies—including random flipping, rotation, and cropping—were applied to the training set, whereas only deterministic preprocessing was applied to the internal validation set.

Model training was performed on a single NVIDIA GPU using the stochastic gradient descent (SGD) optimizer, with an initial learning rate of 0.01, a batch size of 32, and 64 training epochs. The models were optimized using a Cox regression–based time-to-event loss function, generating continuous deep learning–derived risk scores (DL-scores) for each patient. During validation, the model achieving the highest concordance index (C-index) was selected as the optimal network. For each imaging modality, the architecture with the best performance in the internal validation cohort was used to generate modality-specific DL-scores. Specifically, the ultrasound DL-score was derived from representative preoperative grayscale ultrasound images, whereas the CT DL-score was derived from the axial slice corresponding to the maximum tumor cross-sectional area. Higher DL-scores indicated a higher predicted risk of postoperative recurrence. The overall workflow of deep learning model development is illustrated in Fig. 3.

Fig. 3.

Fig. 3

Construction and analysis pipeline of deep learning models based on ultrasound and contrast-enhanced CT

Clinical data collection and clinical model construction

Using the collected clinical variables, univariate Cox proportional hazards regression analysis was first performed to identify potential prognostic factors significantly associated with RFS. Variables demonstrating statistical significance in the univariate analysis (P < 0.05) were subsequently entered into a multivariate Cox proportional hazards regression model to determine independent predictors of recurrence. Only variables that remained statistically significant (P < 0.05) in the multivariate analysis were retained for construction of the final clinical model. For each independent predictor, hazard ratios (HRs) and corresponding 95% confidence intervals (CIs) were calculated to quantify their associations with recurrence risk.

Deep learning–clinical model development and validation

Based on the identified independent clinical predictors from multivariate analysis, together with DL-scores derived from CE-CT and US (CT-DL score and US-DL score), an integrated prognostic model was constructed to assess postoperative recurrence risk. These variables were jointly incorporated into a multivariate Cox proportional hazards model. A nomogram was subsequently developed based on the Cox regression coefficients to visualize the relative contribution of each variable and to facilitate individualized recurrence risk prediction.

Model performance was evaluated using multiple complementary metrics. Discriminative performance was quantified using the C-index and time-dependent receiver operating characteristic (ROC) analysis, with areas under the curve (AUCs) calculated at 1-, 2-, and 3-year postoperative time points. Kaplan–Meier (KM) survival analysis combined with the log-rank test was used to compare survival outcomes between predicted high- and low-risk groups, thereby assessing the model’s risk stratification ability. Calibration curves were generated to evaluate the agreement between predicted and observed RFS probabilities, and decision curve analysis (DCA) was performed to assess the clinical utility and net benefit of the model across a range of threshold probabilities.

Bootstrap resampling analysis

To compare predictive performance between the combined model and other models, a nonparametric bootstrap resampling approach was employed to evaluate differences in model performance. Using the combined model as the reference, differences in C-index and time-dependent AUC were calculated between the combined model and the clinical model, single-modality deep learning models (CT-DL and US-DL), as well as partially fused models (Δ = Combined − comparator).

In each bootstrap iteration, random sampling with replacement was performed from the original dataset, with the resampled sample size identical to that of the original cohort. Model performance was recalculated on the resampled data to obtain the corresponding ΔC-index and ΔAUC values. Time-dependent AUCs were computed at postoperative 1-, 2-, and 3-year time points. If only a single outcome category was present at a given time point in a specific resampling iteration, that iteration was excluded from the AUC calculation at the corresponding time point. This resampling procedure was repeated 1,000 times. The distributions of performance differences across all iterations were summarized to calculate the mean values and 95% percentile intervals (2.5th–97.5th percentiles). In addition, the proportion of iterations with positive performance differences (P(Δ > 0)) was calculated to quantify the frequency with which the combined model outperformed the comparator models during resampling. All bootstrap analyses were conducted separately in the training cohort and the internal validation cohort.

Model interpretability using grad-CAM

To enhance the interpretability of deep learning–based predictions, gradient-weighted class activation mapping (Grad-CAM) was applied to visualize the trained models. Grad-CAM generates heatmaps in the input image space by computing gradient-weighted activations from the final convolutional feature maps with respect to the model output. These visualizations were used to qualitatively illustrate the image regions contributing to recurrence risk prediction and to assess whether model attention was predominantly focused on tumor-related areas.

Grad-CAM analysis was performed on the best-performing deep learning model for each imaging modality and was computed using feature maps from the final convolutional layer. For each input image, gradients of the predicted risk score with respect to the selected convolutional feature maps were calculated and subsequently averaged across spatial dimensions via global average pooling to obtain channel-wise importance weights. The weighted feature maps were then linearly combined and passed through a rectified linear unit (ReLU) function to retain only regions with positive contributions to the prediction. The resulting heatmaps were upsampled using bilinear interpolation to match the spatial resolution of the original input images and overlaid on the corresponding original images for visualization.

Statistical analysis

The normality of continuous variables was assessed using the Shapiro–Wilk test. Depending on data distribution characteristics, normally distributed variables were compared using the independent-samples t test, whereas non-normally distributed variables were analyzed using the Mann–Whitney U test. Categorical variables were compared using either the χ² test or Fisher’s exact test, as appropriate. All statistical tests were two-sided, and a P value < 0.05 was considered statistically significant. Patients were stratified into high and low-risk groups according to the median value of the corresponding prognostic score. All statistical analyses were performed using Python (version 3.7.12) on the OnekeyAI platform (version 3.1.8).

Results

Patient characteristics

A total of 293 patients with pathologically confirmed HGSOC were included, among whom 135 (46.1%) experienced recurrence within the predefined 3-year follow-up period. The mean age was 55.0 ± 9.6 years, and 68.3% of patients were postmenopausal. Median serum CA-125 and HE4 levels were 593.7 U/mL and 314.0 pmol/L, respectively. Most patients (71.3%) underwent primary debulking surgery (PDS) followed by platinum-based chemotherapy, whereas 28.7% received neoadjuvant chemotherapy (NACT) followed by interval debulking surgery (IDS). With respect to disease stage, 57.3% and 14.3% of patients were classified as FIGO stage III and stage IV, respectively. Baseline clinical characteristics did not differ significantly between the training and internal validation cohorts (all P > 0.05). Detailed demographic and clinical characteristics are summarized in Table 1.

Table 1.

Baseline demographic and clinical characteristics of the study population

Characteristic All cohort
(N = 293)
Training cohort
(N = 204)
Validation cohort
(N = 89)
P value
Age (years) 55.00 ± 9.63 55.65 ± 8.62 57.37 ± 9.27 0.139
CA125 (U/mL) 593.70 (161.00, 1316.00) 595.74 (160.55, 1300.00) 591.90 (162.40, 1380.50) 0.732
HE4 (pmol/L) 314.00 (142.10, 683.00) 295.00 (134.15, 707.80) 344.30 (172.60, 625.00) 0.436
AFP (ng/mL) 2.90 (2.10, 3.95) 2.88 (2.06, 3.87) 2.94 (2.14, 4.01) 0.563
CEA (ng/mL) 1.21 (0.69, 1.91) 1.22 (0.70, 1.92) 1.19 (0.68, 1.89) 0.746
CA19-9 (U/mL) 11.22 (5.91, 18.80) 10.62 (6.03, 18.15) 12.74 (5.91, 19.90) 0.396
Platelets (×10⁹/L) 292.00 (242.00, 361.00) 290.00 (240.00, 372.50) 292.00 (248.00, 358.00) 0.903
T lymphocytes (×10⁹/L) 1.55 (1.26, 1.87) 1.58 (1.25, 1.93) 1.54 (1.31, 1.76) 0.609
Monocytes (×10⁹/L) 0.46 (0.35, 0.58) 0.46 (0.35, 0.58) 0.47 (0.35, 0.60) 0.492
Neutrophils (×10⁹/L) 4.33 (3.34, 5.84) 4.31 (3.27, 5.75) 4.44 (3.62, 5.94) 0.276
WBC (×10⁹/L) 6.60 (5.44, 8.05) 6.66 (5.43, 8.05) 6.39 (5.45, 8.04) 0.824
PLR 186.97 (135.95, 253.23) 185.97 (132.66, 254.75) 189.81 (146.67, 248.83) 0.529
NLR 2.80 (2.03, 4.03) 2.63 (1.91, 4.09) 3.04 (2.24, 3.86) 0.123
SII 801.42 (483.84, 1370.29) 751.64 (465.99, 1417.37) 879.63 (571.10, 1342.12) 0.198
LMR 3.53 (2.57, 4.85) 3.58 (2.62, 5.00) 3.29 (2.56, 4.38) 0.203
dNLR 2.05 (1.53, 2.68) 2.00 (1.46, 2.76) 2.14 (1.65, 2.56) 0.498
Therapeutic strategy, n(%) 1.0
NACT + IDS 84 (28.7%) 58 (28.4%) 26 (29.2%)
PDS + Chemotherapy 209 (71.3%) 146 (71.6%) 63 (70.8%)
Adjuvant therapy, n(%) 0.442
No 153 (52.2%) 103 (50.5%) 50 (56.2%)
Yes 140 (47.8%) 101 (49.5%) 39 (43.8%)
Omentectomy, n(%) 1.0
No 9 (3.1%) 6 (2.9%) 3 (3.4%)
Yes 284 (96.9%) 198 (97.1%) 86 (96.6%)
Residual disease status, n(%) 0.955
R0 183 (62.5%) 128 (62.7%) 55 (61.8%)
R1 106 (36.2%) 73 (35.8%) 33 (37.1%)
R2 4 (1.4%) 3 (1.5%) 1 (1.1%)
Chemotherapy cycle(s), n(%) 0.87
≤ 4 32 (10.9%) 21 (10.3%) 11 (12.4%)
5–6 173 (59.0%) 121 (59.3%) 52 (58.4%)
≥ 7 88 (30.0%) 62 (30.4%) 26 (29.2%)
FIGO stage, n(%) 0.42
I 32 (10.9%) 26 (12.7%) 6 (6.7%)
II 51 (17.4%) 34 (16.7%) 17 (19.1%)
III 168 (57.3%) 117 (57.4%) 51 (57.3%)
IV 42 (14.3%) 27 (13.2%) 15 (16.9%)
Menopausal status, n(%) 0.803
Premenopausal 93 (31.7%) 65 (31.8%) 28 (31.5%)
Postmenopausal 200 (68.3%) 139 (68.2%) 61 (68.5%)

Clinical prognostic factors and performance of the clinical model

Univariate Cox proportional hazards regression analysis identified several clinical variables significantly associated with postoperative RFS, including age, CA-125, HE4, inflammatory markers, therapeutic strategy, adjuvant therapy, omentectomy status, and FIGO stage. Multivariate Cox regression analysis confirmed six variables as independent predictors of recurrence. Therapeutic strategy, adjuvant therapy, and omentectomy were significantly associated with a lower risk of recurrence, whereas higher FIGO stage and elevated CA-125 and HE4 levels were associated with an increased risk of recurrence. Detailed results of the univariate and multivariate Cox regression analyses are presented in Table 2.

Table 2.

Univariate and multivariate Cox regression analyses of clinical factors associated with recurrence in patients with HGSOC

Characteristic Univariate HR (95% CI) P value Multivariate HR (95% CI) P value
Age (years) 1.03 (1.01–1.05) 0.010
CA125 (U/mL) 1.01 (1.01–1.01) 0.007 1.01 (1.01–1.01) 0.013
HE4 (pmol/L) 1.01 (1.01–1.01) < 0.001 1.01 (1.01–1.01) < 0.001
AFP (ng/mL) 1.01 (1.01–1.01) 0.002
CEA (ng/mL) 0.95 (0.89–1.01) 0.129
CA199 (U/mL) 1.00 (0.99–1.00) 0.295
Platelet count (×10⁹/L) 1.01 (1.01–1.01) < 0.001
T cells (×10⁹/L) 0.96 (0.87–1.06) 0.433
Monocytes (×10⁹/L) 0.97 (0.79–1.20) 0.791
Neutrophils (×10⁹/L) 1.00 (0.98–1.02) 0.920
WBC (×10⁹/L) 1.00 (0.99–1.00) 0.674
PLR 1.01 (1.01–1.01) < 0.001
NLR 1.01 (0.99–1.03) 0.301
SII 1.00 (1.00–1.00) 0.079
LMR 0.87 (0.79–0.97) 0.012
dNLR 1.15 (1.01–1.32) 0.040
Therapeutic strategy
NACT + IDS 1.00 (Reference) 1.00 (Reference)
PDS + Chemotherapy 0.31 (0.22–0.44) < 0.001 0.48 (0.32–0.71) < 0.001
Adjuvant therapy
No 1.00 (Reference) 1.00 (Reference)
Yes 0.31 (0.21 ~ 0.46) < 0.001 0.30 (0.20 ~ 0.46) < 0.001
Omentectomy
No 1.00 (Reference) 1.00 (Reference)
Yes 0.35 (0.17–0.76) 0.008 0.32 (0.15–0.72) 0.005
Chemotherapy cycle(s)
≤ 4 1.00 (Reference)
5–6 0.63 (0.37–1.08) 0.094
≥ 7 1.53 (0.89–2.64) 0.124
FIGO stage
I 1.00 (Reference) 1.00 (Reference)
II 2.13 (0.85–5.38) 0.108
III 4.64 (2.03–10.63) < 0.001 2.66 (1.81–3.90) < 0.001
IV 4.39 (1.76–10.97) 0.002 2.42 (1.46–4.02) < 0.001
Menopausal status
No 1.00 (Reference)
Yes 1.21 (0.83–1.77) 0.313

Based on these six independent predictors, a clinical Cox prognostic model was constructed. The model achieved C-index values of 0.779 in the training cohort and 0.684 in the internal validation cohort (Fig. 4a and b). Kaplan–Meier survival analysis demonstrated significant differences in RFS between the high- and low-risk groups stratified by the clinical model in both cohorts (log-rank test, P < 0.05). Time-dependent ROC analysis showed that the clinical model yielded AUCs of 0.831, 0.830, and 0.829 for predicting 1-, 2-, and 3-year postoperative RFS in the training cohort, and 0.814, 0.730, and 0.750 in the internal validation cohort, respectively.

Fig. 4.

Fig. 4

C-index values of the clinical and single-modality deep learning models for recurrence-free survival prediction

Performance of deep learning models based on ultrasound and CT

Two-dimensional deep learning survival models were independently developed using preoperative US and CE-CT images. Among the evaluated architectures, ResNet-18 achieved the best performance for US images, whereas DenseNet-161 demonstrated optimal performance for CT images.

For the CT-based deep learning model, the C-index was 0.729 in the training cohort and 0.638 in the internal validation cohort (Fig. 4c and d). The AUCs for predicting postoperative RFS at 1, 2, and 3 years were 0.701, 0.774, and 0.760 in the training cohort, and 0.753, 0.673, and 0.632 in the internal validation cohort, respectively. For the US-based deep learning model, the C-index was 0.732 in the training cohort and 0.640 in the internal validation cohort (Fig. 4e and f). The corresponding AUCs for predicting 1-, 2-, and 3-year postoperative RFS were 0.807, 0.759, and 0.773 in the training cohort, and 0.709, 0.723, and 0.663 in the internal validation cohort.

Kaplan–Meier survival analysis showed significant differences in RFS between the high- and low-risk groups stratified by either the US- or CT-based deep learning scores in both the training and internal validation cohorts (log-rank test, all P < 0.001).

Performance of the multimodal combined model

To systematically evaluate the contribution of different data modalities and their combinations to recurrence prediction, multiple predictive models were constructed and compared, including a clinical model, single-modality deep learning models (CT-DL and US-DL), and multimodal models integrating imaging and clinical information. Given the widespread clinical use and interpretability of CE-CT and US in gynecologic oncology, analyses focused primarily on the CT-DL model, the US-DL model, and their combinations with clinical variables. Comparative predictive performance of all models at different time points is summarized in Table 3.

Table 3.

Predictive performance of different models in the training and validation cohorts

Survival Cohort Signature Accuracy AUC 95% CI Sensitivity Specificity
3 year Training Clinical 0.761 0.829 0.7612–0.8974 0.792 0.741
3 year Training DL_CT 0.696 0.760 0.6792–0.8405 0.736 0.671
3 year Training DL_US 0.746 0.773 0.6929–0.8525 0.623 0.824
3 year Training Clinic_CT 0.819 0.876 0.8157–0.9370 0.736 0.871
3 year Training Clinic_US 0.768 0.870 0.8133–0.9270 0.962 0.647
3 year Training CT_US 0.746 0.821 0.7495–0.8922 0.755 0.741
3 year Training Combined 0.790 0.892 0.8377–0.9461 0.925 0.706
3 year Validation Clinical 0.741 0.750 0.6211–0.8789 0.708 0.765
3 year Validation DL_CT 0.672 0.632 0.4797–0.7850 0.708 0.647
3 year Validation DL_US 0.707 0.663 0.5112–0.8148 0.417 0.912
3 year Validation Clinic_CT 0.724 0.718 0.5837–0.8526 0.708 0.735
3 year Validation Clinic_US 0.724 0.749 0.6205–0.8771 0.667 0.765
3 year Validation CT_US 0.707 0.697 0.5539–0.8407 0.667 0.735
3 year Validation Combined 0.707 0.772 0.6488–0.8953 0.875 0.588
2 year Training Clinical 0.759 0.830 0.7689–0.8911 0.735 0.794
2 year Training DL_CT 0.706 0.774 0.7038–0.8449 0.618 0.838
2 year Training DL_US 0.682 0.759 0.6856–0.8317 0.588 0.824
2 year Training Clinic_CT 0.794 0.869 0.8170–0.9218 0.765 0.838
2 year Training Clinic_US 0.829 0.865 0.8093–0.9214 0.882 0.750
2 year Training CT_US 0.788 0.830 0.7687–0.8922 0.784 0.794
2 year Training Combined 0.806 0.890 0.8418–0.9376 0.765 0.868
2 year Validation Clinical 0.731 0.730 0.6045–0.8553 0.744 0.714
2 year Validation DL_CT 0.672 0.673 0.5397–0.8064 0.667 0.679
2 year Validation DL_US 0.672 0.723 0.5991–0.8478 0.462 0.964
2 year Validation Clinic_CT 0.716 0.723 0.5972–0.8497 0.692 0.750
2 year Validation Clinic_US 0.716 0.735 0.6125–0.8582 0.718 0.714
2 year Validation CT_US 0.731 0.762 0.6461–0.8777 0.692 0.786
2 year Validation Combined 0.746 0.781 0.6696–0.8927 0.795 0.679
1 year Training Clinical 0.688 0.831 0.7616–0.9005 0.637 0.937
1 year Training DL_CT 0.630 0.701 0.6004–0.8011 0.619 0.687
1 year Training DL_US 0.786 0.807 0.7217–0.8923 0.794 0.750
1 year Training Clinic_CT 0.693 0.829 0.7596–0.8982 0.656 0.875
1 year Training Clinic_US 0.812 0.885 0.8334–0.9357 0.794 0.906
1 year Training CT_US 0.828 0.822 0.7447–0.8998 0.856 0.687
1 year Training Combined 0.792 0.880 0.8275–0.9319 0.775 0.875
1 year Validation Clinical 0.812 0.814 0.6891–0.9386 0.818 0.786
1 year Validation DL_CT 0.713 0.753 05956-0.9109 0.682 0.857
1 year Validation DL_US 0.500 0.709 0.5843–0.8334 0.394 1.000
1 year Validation Clinic_CT 0.775 0.848 0.7438–0.9532 0.742 0.929
1 year Validation Clinic_US 0.750 0.801 0.6825–0.9193 0.742 0.788
1 year Validation CT_US 0.787 0.813 0.7030–0.9226 0.788 0.945
1 year Validation Combined 0.800 0.864 0.7749–0.9524 0.773 0.929

CT-DL scores, US-DL scores, and independent clinical predictors identified by multivariate Cox regression were subsequently integrated to construct the final multimodal prognostic model, which was visualized using a nomogram (Fig. 5C). Among all evaluated models, the combined model achieved the best predictive performance, with a C-index of 0.840 in the training cohort and 0.722 in the internal validation cohort (Fig. 5a and b). Time-dependent ROC analysis showed that the combined model yielded AUCs of 0.880, 0.890, and 0.892 for predicting 1-, 2-, and 3-year postoperative RFS in the training cohort, and 0.864, 0.781, and 0.772 in the internal validation cohort, respectively. Comparisons of AUCs across models at each time point are presented in Fig. 6.

Fig. 5.

Fig. 5

Discriminative performance and nomogram of the multimodal combined model for recurrence-free survival

Fig. 6.

Fig. 6

Time-dependent AUC comparison of different models for recurrence-free survival prediction

Kaplan–Meier survival analysis demonstrated that the combined model effectively stratified patients into high- and low-risk groups in both the training and internal validation cohorts (log-rank test, all P < 0.001). Calibration curves showed good agreement between predicted and observed RFS probabilities at each time point (Supplementary Fig. 1). Decision curve analysis indicated that, across a range of threshold probabilities, the combined model provided a higher net benefit than the clinical model and the single-modality deep learning models (Supplementary Fig. 2).

Analysis of model performance and interpretability

To assess the robustness of differences in predictive performance across models, bootstrap resampling analysis was conducted. Bootstrap estimates of C-index differences between the combined model and comparator models in the training and internal validation cohorts are shown in Fig. 7. Overall, the combined model exhibited superior discriminative performance relative to the single-modality models. In most comparisons, C-index differences favored the combined model. This advantage was more pronounced in the training cohort, whereas in the internal validation cohort, the 95% confidence intervals for some comparisons included zero, indicating variability in performance gains across datasets. Bootstrap results for time-dependent AUCs at 1-, 2-, and 3-year time points, together with complete analyses for both cohorts, are provided in the Supplementary Materials (Supplementary Fig. 3).

Fig. 7.

Fig. 7

Bootstrap comparison of C-index between the combined model and alternative models

Model interpretability was further examined using Grad-CAM to visualize image regions contributing to recurrence risk prediction (Fig. 8). The Grad-CAM heatmaps showed that model attention was primarily concentrated on the tumor and adjacent regions, offering qualitative insight into the imaging features underlying the deep learning–based predictions.

Fig. 8.

Fig. 8

Nomogram-based individualized recurrence risk prediction with representative cases and Grad-CAM visualizations

Discussion

Consistent with recent deep learning–based survival and recurrence prediction studies, the present work further demonstrates that joint modeling of imaging features and clinical information can improve the stability and discriminative performance of recurrence risk stratification. The OvarXNet framework proposed by Sadeghi et al. integrates longitudinal PET/CT imaging with clinical variables to predict recurrence risk in HGSOC, underscoring the potential value of combining imaging phenotypes with clinical data in survival modeling [28]. Moreover, accumulating evidence indicates that multimodal deep learning survival models can achieve improved prognostic stratification across a range of solid tumors [29, 30]. Nevertheless, systematic reviews in ovarian cancer have noted that most existing survival prediction studies still rely on single data modalities or traditional machine learning approaches, and that multimodal deep learning frameworks remain underexplored despite being recognized as a critical direction for enhancing predictive accuracy and model interpretability [31]. Against this background, the present study integrates routine CE-CT, US imaging, and clinical variables to develop a multimodal deep survival learning framework for predicting postoperative recurrence risk in patients with HGSOC.

To clarify the prognostic contribution of each imaging modality, independent deep learning survival models were first developed using CE-CT and US. Although three-dimensional deep learning models may theoretically provide more comprehensive characterization of volumetric tumor heterogeneity, a two-dimensional slice-based strategy was adopted to balance robustness, reproducibility, and clinical feasibility. In a single-center setting with a limited sample size, two-dimensional models offer improved training stability, reduced overfitting risk, and substantially lower annotation and computational demands [32–34]. Accordingly, arterial-phase axial slices corresponding to the maximum tumor cross-sectional area were selected as representative inputs for CE-CT. For US—an inherently two-dimensional, real-time imaging modality—modeling based on representative grayscale planes is consistent with its imaging characteristics and routine clinical use.

Among the evaluated architectures, ResNet-18 exhibited the most stable convergence and generalization performance for US-based modeling. Trained end-to-end using a Cox partial likelihood–based loss function, the model captures both local echotextural features and global morphological patterns. For CE-CT–based modeling, DenseNet-161 achieved optimal performance when arterial-phase axial slices corresponding to the maximum tumor cross-sectional area were used as input, facilitating the identification of enhancement and texture features associated with tumor vascularity and aggressiveness. Despite being based on two-dimensional inputs, both models demonstrated meaningful prognostic discrimination. Notably, US primarily reflects tumor microstructure and echogenic characteristics, whereas CE-CT emphasizes vascular supply, enhancement patterns, and invasive extent; thus, the two modalities differ fundamentally in imaging physics and the biological information they encode. Integrating deep learning representations from US and CE-CT therefore enables complementary characterization of tumor heterogeneity across multiple imaging dimensions, providing more comprehensive information for recurrence risk assessment.

Building on these imaging-based models, deep learning–derived scores were integrated with independent clinical predictors identified through multivariate Cox regression to construct a multimodal prognostic model. The included clinical variables—CA-125, HE4, FIGO stage, therapeutic strategy, adjuvant therapy, and omentectomy status—have been well established as prognostic factors in ovarian cancer. Although the hazard ratios of CA-125 and HE4 are close to unity on a per-unit basis, their wide dynamic ranges as continuous variables can yield cumulative effects that materially influence recurrence risk, supporting their identification as independent predictors in multivariate analysis [35]. Importantly, surgery- and treatment-related variables in retrospective studies are more likely to reflect baseline tumor burden and resectability rather than direct causal treatment effects [36, 37]. Accordingly, in the present study, these variables were incorporated primarily for risk stratification rather than for inferring treatment efficacy.

The robustness of performance differences among models was further evaluated using bootstrap resampling. The results indicate that the combined model demonstrated superior discriminative performance in the majority of resampling iterations across both the training and internal validation cohorts. For certain metrics, bootstrap confidence intervals crossed zero, which may reflect the high correlation and shared features among the compared models, as well as the conservative nature of resampling-based inference in nested model comparisons. Therefore, the bootstrap findings should be interpreted as evidence supporting the stability and non-inferiority of the combined model, rather than proof of strict statistical superiority. In terms of model interpretability, Grad-CAM visualizations showed that model attention was predominantly concentrated on the tumor parenchyma and adjacent regions, which are commonly associated with tumor proliferation, angiogenesis, and local invasion. These observations suggest imaging plausibility of the model’s predictions. However, such visualizations reflect attention patterns during prediction and do not constitute direct validation of underlying biological mechanisms or histopathological correlates, which warrant further investigation in future studies incorporating pathological data.

From a clinical perspective, this study proposes an exploratory, hypothesis-generating framework for postoperative risk assessment based on routinely acquired preoperative imaging and clinical information, with potential utility for identifying patients at elevated risk of recurrence. Nevertheless, several limitations should be acknowledged. First, the CT-based deep learning model relied on a single two-dimensional axial slice corresponding to the maximum tumor cross-sectional area, which may limit capture of full three-dimensional spatial heterogeneity. Although such slices can reflect key imaging features, they cannot fully substitute for volumetric three-dimensional modeling in spatial information integration. However, under the current sample size and single-center design, two-dimensional modeling offers clear advantages in training stability, annotation consistency, and computational efficiency. Future studies in larger, multicenter cohorts should further explore three-dimensional deep learning models to determine whether they provide incremental value for recurrence risk prediction using CE-CT imaging. Second, this study employed a single-center retrospective design with internal validation only; consequently, the model is not yet suitable for immediate clinical decision-making. Its clinical applicability requires confirmation through prospective, multicenter external validation. In addition, despite efforts to standardize imaging acquisition and preprocessing, variability in imaging protocols across institutions may still limit cross-center generalizability.

Future research will focus on prospective validation in large, multicenter cohorts and on integrating three-dimensional imaging features with multi-omics data to further enhance model robustness, generalizability, and potential clinical utility.

Conclusion

In conclusion, this study developed and internally validated a multimodal deep survival learning framework that integrates preoperative CE-CT and ultrasound imaging features with clinical variables to predict postoperative recurrence risk in patients with HGSOC. By leveraging complementary information from different imaging modalities and established clinical predictors, the combined model demonstrated improved discriminative performance and stable risk stratification compared with clinical or single-modality deep learning models.

The proposed framework relies exclusively on routinely acquired preoperative data and incorporates interpretable modeling strategies, including Cox-based survival learning and Grad-CAM visualization, providing qualitative insight into imaging regions contributing to recurrence risk prediction. Although the model showed promising performance in internal validation, its current role should be regarded as exploratory and hypothesis-generating.

Future work will focus on prospective validation in multicenter cohorts and on extending the framework through the integration of three-dimensional imaging features and multi-omics data, with the goal of further improving model robustness, generalizability, and potential clinical utility in postoperative risk stratification for HGSOC.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

The authors would like to thank all participating centers and radiology departments for their support in data collection and validation.

Abbreviations

AUC

Area under the receiver operating characteristic curve

CE-CT

Contrast-enhanced computed tomography

CI

Confidence interval

C-index

Concordance index

DCA

Decision curve analysis

DICOM

Digital Imaging and communications in medicine

dNLR

Derived neutrophil-to-lymphocyte ratio

DL

Deep learning

FIGO

International federation of gynecology and obstetrics

Grad-CAM

Gradient-weighted class activation mapping

HR

Hazard ratio

HGSOC

High-grade serous ovarian carcinoma

IDS

Interval debulking surgery

KM

Kaplan–Meier

LMR

Lymphocyte-to-monocyte ratio

NACT

Neoadjuvant chemotherapy

NCCN

National comprehensive cancer network

NLR

Neutrophil-to-lymphocyte ratio

PLR

Platelet-to-lymphocyte ratio

RFS

Recurrence-free survival

ROC

Receiver operating characteristic

US

Ultrasound

SGD

Stochastic gradient descent

ViT

Vision Transformer

VOI

Volume of interest

Author contributions

SN contributed to data and image collection, image annotation, figure preparation, manuscript drafting, and proposed the initial research concept. YJ was responsible for statistical analysis, model construction, figure preparation, and manuscript drafting. YD assisted in the literature review and general data searching. YH contributed to data acquisition. DJ provided technical support and equipment for image acquisition. AC and HC jointly supervised the study. HC served as the corresponding author, while AC acted as the co-corresponding author. All authors read and approved the final manuscript.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Data availability

The numerical datacohort analyzed in this study is available upon reasonable request from the corresponding author. Due to privacy restrictions, JPG files cannot be provided freely.

Declarations

Ethics approval and consent to participate

This retrospective study was reviewed and approved by the Ethics Committees of the Affiliated Hospital of Qingdao University. Written informed consent was waived due to the retrospective nature of the study, and patient data confidentiality was strictly ensured. All methods were carried out in accordance with relevant guidelines and regulations, including the ethical principles of the Declaration of Helsinki.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Silin Nie and Yumin Jiang contributed equally to this work.

Contributor Information

Aiping Chen, Email: chenaiping@qdu.edu.cn.

Huijun Chu, Email: chuhuijun@qdu.edu.cn.

References

  • 1.Ittner E, et al. Diagnostic and prognostic biomarkers associated with histotype in advanced epithelial ovarian cancer. Sci Rep. Oct 23 2025;15(1):37171. 10.1038/s41598-025-24938-0 [DOI] [PMC free article] [PubMed]
  • 2.Momenimovahed Z, Tiznobaik A, Taheri S, Salehiniya H. Ovarian cancer in the world: epidemiology and risk factors. Int J Womens Health. 2019;11:287–99. 10.2147/ijwh.S197604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lisio M-A, Fu L, Goyeneche A, Gao ZH, Telleria C. High-grade serous ovarian cancer: basic sciences, clinical and therapeutic standpoints. Int J Mol Sci. 2019;20(4):952. https://www.mdpi.com/1422-0067/20/4/952 [DOI] [PMC free article] [PubMed]
  • 4.O’Malley DM, Krivak TC, Kabil N, Munley J, Moore KN. PARP inhibitors in ovarian cancer: a review. Target Oncol. 2023 Jul;18(4):471–503. 10.1007/s11523-023-00970-w [DOI] [PMC free article] [PubMed]
  • 5.Ledermann JA, et al. ESGO-ESMO-ESP consensus conference recommendations on ovarian cancer: pathology and molecular biology and early, advanced and recurrent disease. Ann Oncol. 2024 Mar;35(3):248–266. 10.1016/j.annonc.2023.11.015 [DOI] [PubMed]
  • 6.Tewari KS, et al. Final overall survival of a randomized trial of bevacizumab for primary treatment of ovarian Cancer. J Clin Oncol. Sep 10 2019;37:2317–28. 10.1200/jco.19.01009. [DOI] [PMC free article] [PubMed]
  • 7.Garzon S, et al. Secondary and tertiary ovarian cancer recurrence: what is the best management? Gland Surg. 2020 Aug;9(4):1118–1129. 10.21037/gs-20-325 [DOI] [PMC free article] [PubMed]
  • 8.Shah JB, et al. Analysis of matched primary and recurrent BRCA1/2 mutation-associated tumors identifies recurrence-specific drivers. Nat Commun. 2022 Nov 7;13(1):6728. 10.1038/s41467-022-34523-y [DOI] [PMC free article] [PubMed]
  • 9.Luvero D, Milani A, Ledermann JA. Treatment options in recurrent ovarian cancer: latest evidence and clinical potential. Ther Adv Med Oncol. 2014 Sep;6(5):229–239. 10.1177/1758834014544121 [DOI] [PMC free article] [PubMed]
  • 10.Ushijima K. Treatment for recurrent ovarian cancer at first relapse. J Oncol. 2010;2010:497429. 10.1155/2010/497429 [DOI] [PMC free article] [PubMed]
  • 11.Nag S, et al. Pan-Asia adapted ESMO clinical practice guideline for the management of patients with newly diagnosed and relapsed epithelial ovarian cancer. ESMO Open. 2025 Jun;10(6):105125. 10.1016/j.esmoop.2025.105125. [DOI] [PMC free article] [PubMed]
  • 12.Giampaolino P, Foreste V, Della Corte L, Di Filippo C, Iorio G, Bifulco G. Role of biomarkers for early detection of ovarian cancer recurrence. Gland Surg. 2020 Aug;9(4):1102–1111. 10.21037/gs-20-544 [DOI] [PMC free article] [PubMed]
  • 13.Du K, et al. An increase of serum CA-125 to two times of nadir level strongly predicts the image-identified relapse of serous ovarian cancer. Sci Rep. Jul 1 2024;14(1):14986. 10.1038/s41598-024-65760-4. [DOI] [PMC free article] [PubMed]
  • 14.Zhou SK, et al. A review of deep learning in medical imaging: imaging traits, technology trends, case studies with progress highlights, and future promises. Proc IEEE Inst Electr Electron Eng. 2021 May;109(5):820–838. 10.1109/jproc.2021.3054390 [DOI] [PMC free article] [PubMed]
  • 15.Ardila D, et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nat Med. 2019 Jun;25(6):954–961. 10.1038/s41591-019-0447-x [DOI] [PubMed]
  • 16.Osny A, Parmar C, Quackenbush J, Schwartz LH, Aerts H. Artificial intelligence in radiology. Nat Rev Cancer. 2018 Aug;18(8):500–510. 10.1038/s41568-018-0016-5 [DOI] [PMC free article] [PubMed]
  • 17.Bi WL, et al. Artificial intelligence in cancer imaging: clinical challenges and applications. CA Cancer J Clin. 2019 Mar;69(2):127–157. 10.3322/caac.21552 [DOI] [PMC free article] [PubMed]
  • 18.Schouten D, et al. Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications. Med Image Anal. 2025 Oct;105:103621. 10.1016/j.media.2025.103621 [DOI] [PubMed]
  • 19.Yang H, Yang M, Chen J, Yao G, Zou Q, Jia L. Multimodal deep learning approaches for precision oncology: a comprehensive review. Brief Bioinform. 2024 Nov 22;26(1). 10.1093/bib/bbae699 [DOI] [PMC free article] [PubMed]
  • 20.Sadeghi MH, Sina S, Omidi H, Farshchitabrizi AH, Alavi M. Deep learning in ovarian cancer diagnosis: a comprehensive review of various imaging modalities. Pol J Radiol. 2024;89(48):e30–e. 10.5114/pjr.2024.134817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Rizzo S, et al. Ovarian cancer staging and follow-up: updated guidelines from the European Society of Urogenital Radiology female pelvic imaging working group. Eur Radiol. 2025 Jul;35(7):4029–4039. 10.1007/s00330-024-11300-7 [DOI] [PMC free article] [PubMed]
  • 22.Sahdev A. CT in ovarian cancer staging: how to review and report with emphasis on abdominal and pelvic disease for surgical planning. Cancer Imaging. 2016 Aug 2;16(1):19. 10.1186/s40644-016-0076-2 [DOI] [PMC free article] [PubMed]
  • 23.Tantawy HFA, Ebrahim SAM, Kamal MRA, Hassan RM. The diagnostic performance of ultrasound in the diagnosis of indeterminate adnexal masses based on the O-RADS US scoring system. Egypt J Radiol Nucl Med. 2024 Jan 8;55(1):11. 10.1186/s43055-024-01184-4
  • 24.Xie W, et al. Developing a deep learning model for predicting ovarian cancer in Ovarian-Adnexal Reporting and Data System Ultrasound (O-RADS US) Category 4 lesions: a multicenter study. J Cancer Res Clin Oncol. 2024 Jul 9;150(7):346. 10.1007/s00432-024-05872-6. [DOI] [PMC free article] [PubMed]
  • 25.Zheng Y, et al. Preoperative CT-based deep learning model for predicting overall survival in patients with high-grade serous ovarian cancer. Front Oncol. 2022;12:986089. 10.3389/fonc.2022.986089 [DOI] [PMC free article] [PubMed]
  • 26.Wang S, et al. Deep learning provides a new computed tomography-based prognostic biomarker for recurrence prediction in high-grade serous ovarian cancer. Radiother Oncol. 2019 Mar;132:171–177. 10.1016/j.radonc.2018.10.019 [DOI] [PubMed]
  • 27.Su C, Miao K, Zhang L, Dong X. Deep learning based on ultrasound images to predict platinum resistance in patients with epithelial ovarian cancer. Biomed Eng Online. 2025 May 13;24(1):58. 10.1186/s12938-025-01391-8 [DOI] [PMC free article] [PubMed]
  • 28.Sadeghi MH, et al. Longitudinal deep learning models for tracking disease progression in ovarian cancer using PET/CT imaging and clinical reports. Phys Eng Sci Med. 2025 Nov 10. 10.1007/s13246-025-01669-0. [DOI] [PubMed]
  • 29.Zhang L, et al. Deep learning predicts overall survival of patients with unresectable hepatocellular carcinoma treated by transarterial chemoembolization plus Sorafenib. Front Oncol. 2020;10:593292. 10.3389/fonc.2020.593292. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Mahootiha M, Qadir HA, Bergsland J, Balasingham I. Multimodal deep learning for personalized renal cell carcinoma prognosis: integrating CT imaging and clinical data. Comput Methods Programs Biomed. 2024 Feb;244:107978. 10.1016/j.cmpb.2023.107978. [DOI] [PubMed]
  • 31.Asadi F, Rahimi M, Ramezanghorbani N, Almasi S. Comparing the effectiveness of artificial intelligence models in predicting ovarian cancer survival: a systematic review. Cancer Rep (Hoboken). 2025 Mar;8(3):e70138. 10.1002/cnr2.70138 [DOI] [PMC free article] [PubMed]
  • 32.Zhang Y, Liao Q, Ding L, Zhang J. Bridging 2D and 3D segmentation networks for computation-efficient volumetric medical image segmentation: an empirical study of 2.5D solutions. Comput Med Imaging Graph. 2022 Jul;99:102088. 10.1016/j.compmedimag.2022.102088. [DOI] [PubMed]
  • 33.Singh SP, Wang L, Gupta S, Goli H, Padmanabhan P, Gulyás B. 3D deep learning on medical images: a review. Sensors (Basel). 2020 Sep 7;20(18). 10.3390/s20185097 [DOI] [PMC free article] [PubMed]
  • 34.Tajbakhsh N, Jeyaseelan L, Li Q, Chiang JN, Wu Z, Ding X. Embracing imperfect datasets: a review of deep learning solutions for medical image segmentation. Med Image Anal. 2020 Jul;63:101693. 10.1016/j.media.2020.101693. [DOI] [PubMed]
  • 35.Anastasi E, et al. Recent insight about HE4 role in ovarian cancer oncogenesis. Int J Mol Sci. 2023 Jun 22;24(13). 10.3390/ijms241310479. [DOI] [PMC free article] [PubMed]
  • 36.Bryant A, et al. Impact of residual disease as a prognostic factor for survival in women with advanced epithelial ovarian cancer after primary surgery. Cochrane Database Syst Rev. 2022 Sep 26;9(9):CD015048. 10.1002/14651858.CD015048.pub2 [DOI] [PMC free article] [PubMed]
  • 37.Eisenkop SM, Spirtos NM, Friedman RL, Lin WC, Pisani AL, Perticucci S. Relative influences of tumor volume before surgery and the cytoreductive outcome on survival for patients with advanced ovarian cancer: a prospective study, (in eng). Gynecol Oncol. Aug 2003;90(2):390–6. 10.1016/s0090-8258(03)00278-6. [DOI] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The numerical datacohort analyzed in this study is available upon reasonable request from the corresponding author. Due to privacy restrictions, JPG files cannot be provided freely.


Articles from BMC Medical Imaging are provided here courtesy of BMC

RESOURCES