Skip to main content
Radiation Oncology (London, England) logoLink to Radiation Oncology (London, England)
. 2025 Aug 13;20:127. doi: 10.1186/s13014-025-02695-8

A stacking ensemble framework integrating radiomics and deep learning for prognostic prediction in head and neck cancer

Bingzhen Wang 1,2, Jinghua Liu 3,4, Xiaolei Zhang 2, Jianpeng Lin 2, Shuyan Li 5, Zhongxiao Wang 2, Zhendong Cao 6, Dong Wen 7, Tiange Liu 7, Hafiz Rashidi Harun Ramli 1, Hazreen Haizi Harith 1, Wan Zuha Wan Hasan 1,, Xianling Dong 2,8,9,
PMCID: PMC12351975  PMID: 40804402

Abstract

Background

Radiomics models frequently face challenges related to reproducibility and robustness. To address these issues, we propose a multimodal, multi-model fusion framework utilizing stacking ensemble learning for prognostic prediction in head and neck cancer (HNC). This approach seeks to improve the accuracy and reliability of survival predictions.

Methods

A total of 806 cases from nine centers were collected; 143 cases from two centers were assigned as the external validation cohort, while the remaining 663 were stratified and randomly split into training (n = 530) and internal validation (n = 133) sets. Radiomics features were extracted according to IBSI standards, and deep learning features were obtained using a 3D DenseNet-121 model. Following feature selection, the selected features were input into Cox, SVM, RSF, DeepCox, and DeepSurv models. A stacking fusion strategy was employed to develop the prognostic model. Model performance was evaluated using Kaplan-Meier survival curves and time-dependent ROC curves.

Results

On the external validation set, the model using combined PET and CT radiomics features achieved superior performance compared to single-modality models, with the RSF model obtaining the highest concordance index (C-index) of 0.7302. When using deep features extracted by 3D DenseNet-121, the PET + CT-based models demonstrated significantly improved prognostic accuracy, with Deepsurv and DeepCox achieving C-indices of 0.9217 and 0.9208, respectively. In stacking models, the PET + CT model using only radiomics features reached a C-index of 0.7324, while the deep feature-based stacking model achieved 0.9319. The best performance was obtained by the multi-feature fusion model, which integrated both radiomics and deep learning features from PET and CT, yielding a C-index of 0.9345. Kaplan–Meier survival analysis further confirmed the fusion model’s ability to distinguish between high-risk and low-risk groups.

Conclusion

The stacking-based ensemble model demonstrates superior performance compared to individual machine learning models, markedly improving the robustness of prognostic predictions.

Keywords: Head and neck Cancer, Radiomics, DenseNet, Stacking, Ensemble, Prognostic

Introduction

Head and neck cancer (HNC) refers to malignancies occurring in the head and neck region, including oropharyngeal, nasopharyngeal, hypopharyngeal cancers, and other related anatomical sites [1]. Among them, head and neck squamous cell carcinoma (HNSCC) is the predominant subtype, accounting for approximately 90% of cases [2]. In 2018, there were an estimated 890,000 new cases of HNC worldwide, with nearly 450,000 associated deaths [3, 4]. Major risk factors include tobacco use, alcohol consumption, and human papillomavirus (HPV) infection [5]. HNC is characterized by high aggressiveness and a strong propensity for lymphatic metastasis [6]. In recent years, radiomics has demonstrated significant prognostic value in HNC [7, 8]. Numerous studies have shown that texture features extracted from medical images can aid in tumor diagnosis and prognosis [9]. Moreover, advances in image standardization, multimodal image fusion, and deep learning techniques have substantially improved the accuracy of tumor segmentation and feature extraction, thereby enhancing prognostic performance in HNC [1012].

Despite the significant progress of radiomics in tumor diagnosis and prognosis, it still faces considerable challenges, including limited model robustness, poor reproducibility, and pronounced batch effects in multi-center data [13, 14]. A major contributing factor is the lack of standardized feature extraction protocols. Although Zwanenburg et al. emphasized this issue in the Image Biomarker Standardisation Initiative (IBSI), the framework has not been widely adopted with the rise of deep learning [14, 15]. Moreover, limited training data and the absence of independent validation cohorts hinder model performance; Grau et al. reported a marked decline in accuracy when tested on external datasets [16, 17]. Imaging and equipment heterogeneity across centers further compromises model training, as highlighted by Chen et al., who noted its significant impact on model stability [18]. In addition, different models often yield inconsistent predictions on the same dataset [19], and single-model approaches are highly sensitive to data distribution and parameter settings, limiting their generalizability [20]. Collectively, these limitations hinder the clinical translation of radiomics.

Ensemble learning, particularly stacking, is widely recognized for enhancing model robustness and reproducibility [21]. By integrating predictions from multiple base learners through a meta-model, stacking effectively reduces overfitting and mitigates noise [22]. Saad et al. demonstrated that ensemble approaches outperform single models in predicting radiotherapy outcomes for lung cancer [23]. S Similarly, Zhang et al. reported superior diagnostic accuracy of stacking models over individual classifiers in detecting sacroiliitis associated with ankylosing spondylitis [24]. Despite these advantages, most existing studies rely on single-center, single-modality data, leaving key issues in radiomics unresolved. Moreover, optimal weight allocation between traditional models (e.g., Cox regression) and deep learning models (e.g., DeepSurv) within ensemble frameworks remains unclear. There is an urgent need to develop a multi-center, multi-modality ensemble framework that integrates diverse model architectures and systematically evaluates their suitability for survival prediction, aiming to advance the clinical translation of radiomics.

Building on the existing background, this study collected PET and CT imaging data from patients with HNC across nine centers. Both deep learning features based on a 3D DenseNet-121 and traditional radiomics features were extracted. Subsequently, a stacking ensemble learning framework was constructed using Cox regression, support vector machine (SVM), random survival forest (RSF), DeepCox, and DeepSurv as base models, with DeepSurv serving as the meta-learner. The framework was validated on multi-center, multi-modality data, providing insights for the further optimization and clinical translation of radiomics.

Methods

Data collection

This study utilized a multicenter dataset of 806 patients with histologically confirmed HNC, sourced from 5 datasets encompassing 9 centers, all of whom underwent pre-treatment 18F-FDG PET/CT imaging. Dataset 1 included 296 patients from four Canadian centers: Centre Hospitalier de l’Université de Montréal (CHUM, n = 65), Centre Hospitalier Universitaire de Sherbrooke (CHUS, n = 100), Jewish General Hospital (HGJ, n = 90), and Hôpital Maisonneuve-Rosemont (HMR, n = 41). Dataset 2 contained 158 patients from the M.D. Anderson Cancer Center (MDACC, n = 158). Dataset 3 involved 74 patients from Maastricht Radiation Oncology (MAASTRO, n = 74) in the Netherlands. Dataset 4 consisted of 29 patients from the University of Pittsburgh, available via The Cancer Genome Atlas (TCGA, n = 29). Dataset 5 included 243 patients from the University of Iowa: Quantitative Imaging Network – Site 1 (QIN1, n = 130) and Quantitative Imaging Network – Site 2 (QIN2, n = 113). The median follow-up across all cohorts was 45 months (range: 1–133 months). This dataset was used to develop a prognostic model for overall survival in HNC patients.

Study design

Figure 1 illustrates the overall experimental design. Initially, a total of 663 cases from seven centers—CHUS, CHUM, HGJ, HMR, MDACC, MAASTRO and QIN 1—were stratified and randomly allocated into a training set and an internal validation set at an 80%:20% ratio. The training set comprised 530 cases, while the internal validation set included 133 cases. Additionally, 114 cases from QIN2 and 29 cases from TCGA were utilized as an external validation set to further evaluate the model’s performance. The detailed dataset division and baseline characteristics of the cases are provided in Table 1. For the internal validation set, a ten-fold cross-validation approach was applied for model training and evaluation, with the final model performance being determined by averaging the results from all ten validation rounds.

Fig. 1.

Fig. 1

The entire experimental workflow of this study, including image acquisition, radiomics and deep learning feature extraction, feature selection, construction of the ensemble learning model, and model evaluation

Table 1.

Baseline characteristics and overview of dataset partition (training, internal validation, and external validation sets)

Train (n = 530) Internal Validation (n = 133) External Validation (n = 143) p-value
Age (mean (SD)) a 60.61 (10.14) 59.71 (10.08) 57.68 (10.50) 0.009
Sex (%) b 0.103
Female 99 (18.7) 33 (24.8) 29 (20.3)
Male 431 (81.3) 100 (75.2) 114 (79.7)
T-stage (%) c 0.154
T0 5 (0.9) 0 (0.0) 0 (0.0)
T1 85 (16.0) 18 (13.5) 25 (17.5)
T2 176 (33.2) 40 (30.1) 55 (38.5)
T3 148 (27.9) 45 (33.8) 24 (16.8)
T4 105 (19.8) 27 (20.3) 36 (25.2)
Tx 11 (2.1) 3 (2.3) 3 (2.1)
N-stage (%) c 0.819
N0 115 (21.7) 24 (18.0) 35 (24.5)
N1 65 (12.3) 19 (14.3) 17 (11.9)
N2 316 (59.6) 83 (62.4) 85 (59.4)
N3 34 (6.4) 7 (5.3) 6 (4.2)
M-stage (%) c 0.014
M0 526 (99.2) 133 (100.0) 140 (97.9)
Mx 4 (0.8) 0 (0.0) 2 (1.4)
Unknown 0 (0.0) 0 (0.0) 1 (0.7)
AJCC_Stage (%) c 0.329
Stage I 18 (3.4) 3 (2.3) 7 (4.9)
Stage II 38 (7.2) 7 (5.3) 15 (10.5)
Stage III 92 (17.4) 32 (24.1) 24 (16.8)
Stage IV 381 (71.9) 90 (67.7) 97 (67.8)
unknown 1 (0.2) 1 (0.8) 0 (0.0)
Primary Site (%) c < 0.001
Oropharynx 303 (57.2) 75 (56.4) 28 (19.6)
Larynx 76 (14.3) 11 (8.3) 9 (6.3)
Nasopharynx 27 (5.1) 9 (6.8) 1 (0.7)
Base of Tongue 22 (4.2) 5 (3.8) 25 (17.5)
Tonsil 24 (4.5) 9 (6.8) 52 (36.4)
Glottis 16 (3.0) 7 (5.3) 2 (1.4)
Hypopharynx 16 (3.0) 7 (5.3) 2 (1.4)
Supraglottis 11 (2.1) 6 (4.5) 6 (4.2)
Others 22 (5.8) 2 (1.6) 13(9.1)
Unknown 16 (3.0) 2 (1.5) 2 (1.4)
Death (%) b 0.785
Yes 164 (30.9) 41 (30.8) 40 (28.0)
No 366 (69.1) 92 (69.2) 103 (72.0)

SD Standard deviation

a: Kruskal–Wallis test; b: Pearson Chi-square test; c: Fisher’s exact test

Tumor segmentation

In this study, the primary tumor and lymph nodes (GTV and GTVnd) were used as mask files. For patients with available DICOM segmentation files (SEG) or radiation therapy structure sets (RTSTRUCT), mask files were generated directly. For patients lacking such files, a radiologist manually delineated the regions using the semi-automatic tool in 3D Slicer version 4.10.2.

Image preprocessing and feature extraction

In this dataset, the CT image pixel spacing ranged from 0.68 to 1.95 mm, with slice thicknesses between 1.5 and 5 mm, and matrix sizes of 512. PET image pixel spacing ranged from 1.95 to 5.47 mm, with slice thicknesses between 2.02 and 5 mm, and matrix sizes ranging from 128 to 256. Prior to feature extraction, all images were resampled to a 1 × 1 × 1 mm³ resolution using the nearest neighbor interpolation method. PET images were converted to standardized uptake values (SUVs) before feature extraction [25].

For radiomics feature extraction, the open-source package PyRadiomics was utilized [26]. Three types of images were processed: the original image (Original), Laplacian of Gaussian (LoG) filtered image, and wavelet filtered image. The original images were used to extract unprocessed basic features, while LoG filtering was applied to capture multi-scale texture variations and edge information using Gaussian Laplacian filters at different scales (σ = 1.0, 3.0, and 5.0). Wavelet filtering decomposed the image into multiple frequency bands, enabling the extraction of both high- and low-frequency features. From both PET and CT images, 1,130 radiomics features were extracted for each modality per case, yielding a total of 2,260 features.

For deep learning feature extraction, a 3D convolutional neural network based on the 3D DenseNet-121 architecture was developed [27]. Tumor regions were cropped from the PET and CT scans using a 96 × 96 × 64 mm³ bounding box and input into the model. The network comprised an initial convolutional layer, followed by four Dense Blocks (with 6, 12, 24, and 16 convolutional layers, respectively, and a growth rate of 32), and three transition layers. Regularization was implemented with a dropout probability of 0.2, and ReLU activation functions were applied throughout. The extracted features were mapped to the risk space via adaptive average pooling and a linear classifier. Model parameters were optimized using the Cox Proportional Hazards loss function (CoxPHLoss) [28], with the Adam optimizer and an initial learning rate and weight decay of 1 × 10⁻⁵. The training process incorporated an adaptive learning rate scheduler and an early stopping mechanism, with a maximum of 300 epochs and a batch size of 4. After training, the extracted features were saved and combined with subsequent models for further prognostic analysis. The network architecture of 3D DenseNet-121 is depicted in Fig. 2. Deep learning features (1,024 per modality) were extracted from both PET and CT images for each case, resulting in a total of 2,048 features.

Fig. 2.

Fig. 2

Network architecture of 3D DenseNet-121

Multicenter feature harmonization and feature selection

This study utilized imaging data from 5 datasets, representing a total of 9 distinct centers. Given the substantial inter-center variations in the imaging data, feature harmonization was performed using the NeuroCombat method to ensure data consistency and comparability [29]. This method is based on the ComBat model, which assumes that biases introduced by different centers can be modeled through additive and multiplicative effects. The specific model is as follows:

graphic file with name d33e1044.gif

xij denotes the original observed value of the i-th sample on the j-th feature;

αj is the overall intercept for feature j;

Zi is the covariate vector for sample i (including sex, age, and survival status), and β represents the corresponding regression coefficients;

γb[i],j and δb[i],j represent the additive and multiplicative batch effects for feature j in batch (or center) b[i], respectively.

This method has been widely applied for harmonizing multicenter radiomic features, effectively removing inter-center biases while preserving biological covariates and enhancing model reproducibility across centers.

For feature selection, univariate Cox regression analysis was initially applied to features extracted from various modalities and methods, retaining those with a p-value < 0.05. Redundant features were then removed by calculating Spearman’s correlation coefficients and excluding those exhibiting high correlation. Subsequently, the Least Absolute Shrinkage and Selection Operator (LASSO) regression was employed to further refine the selection of important features. These selected features were used in subsequent analyses and model development, enhancing the reliability and robustness of the results.

Development of a prognostic model using stacking-based ensemble learning

In ensemble learning, Stacking is a commonly used nonlinear fusion strategy that enhances model generalization by combining the outputs of multiple base learners and training a meta-learner on these predictions. Unlike Bagging and Boosting, Stacking allows the integration of heterogeneous models, making it well-suited for capturing complex relationships. Specifically, given a training set Inline graphic, one first constructs M machine learning models f1, f2, …fM, , and obtains the prediction results for each sample as:

Inline graphic

These predictions are then used as new features to train a meta-learner g(⋅), which produces the final prediction as:

Inline graphic

The outputs of base learners are typically obtained through cross-validation to prevent information leakage and ensure reliable model performance evaluation. This approach offers high structural flexibility and demonstrates strong generalization capability across diverse tasks.

The stacking ensemble learning model used in this study incorporates base models in the first layer, including Cox, RSF, SVM, DeepSurv, and DeepCox. These models generate corresponding survival risk scores (e.g., Rad-Cox-Score and DL-Cox-Score) based on radiomics and deep learning features. The second layer is the meta-model, where DeepSurv serves as the integrative model. This layer combines the risk scores from all base models in the first layer as input features, effectively consolidating predictive information across multiple models to generate the final survival prediction [30]. The network architectures of the DeepCox and DeepSurv are depicted in Fig. 3 [31, 32].

Fig. 3.

Fig. 3

Network architecture of DeepCox and DeepSurv

DeepSurv employs a neural network with 2 hidden layers, each comprising 8 neurons, using the ReLU activation function. To mitigate overfitting, the model incorporates L2 regularization (with a strength of 16.094) and Dropout (with a probability of 0.147). It is optimized using a negative log-likelihood loss function, with a learning rate of 0.067, a batch size of 256, and a maximum of 200 training epochs, including early stopping. DeepCox, on the other hand, features 2 hidden layers with 256 neurons each, combined with Batch Normalization, LeakyReLU activation, and Dropout (probability of 0.3) to enhance model generalization. The training process uses the AdamW optimizer (with a learning rate of 0.001 and weight decay of 1e-4) and a negative partial log-likelihood loss function. The model also trains for up to 200 epochs with early stopping.

Model training and testing

In this study, the selected imaging features underwent a comprehensive performance evaluation. PET radiomics features, PET deep learning features, CT radiomics features, CT deep learning features, and combined PET + CT features were sequentially input into five distinct base models, including Cox, RSF, SVM, DeepSurv, and DeepCox, to assess their prognostic prediction performance. Subsequently, a Stacking-based ensemble learning approach was employed to further evaluate the integration of multimodal and multi-feature data. The final evaluation result was determined by averaging the ten-fold cross-validation outcomes.

Statistical analysis

Statistical analyses were conducted using R software (version 4.2.1). Kaplan–Meier survival curves were generated to evaluate survival outcomes, with intergroup differences assessed using the log-rank test. Multivariate Cox proportional hazards regression was analyzed using Harrell’s C-index, and time-dependent receiver operating characteristic (ROC) curves with corresponding area under the curve (AUC) values were plotted using the timeROC package. The NeuroCombat method, implemented via the neuroCombat package, was employed to correct features across multicenter datasets. Statistical significance was defined as a two-sided p value < 0.05.

Results

Comparison of predictive model discrimination ability of each base model

Table 2 presents the concordance index (C-index) values for radiomics features extracted from three imaging modalities (PET, CT, PET + CT) across various models (Cox, SVM, RSF, DeepCox, DeepSurv). For the PET modality, the C-index for radiomics models on the External Validation Set ranged from 0.6392 (SVM) to 0.6715 (RSF), while deep learning models exhibited C-index values between 0.8803 (SVM) and 0.904 (Cox). In the CT modality, the radiomics models’ C-index on the External Validation Set varied from 0.6642 (Cox) to 0.7109 (RSF), with deep learning models ranging from 0.8119 (SVM) to 0.8252 (RSF). For the combined PET + CT modality, radiomics models showed C-index values between 0.6671 (DeepSurv) and 0.7302 (RSF), while deep learning models demonstrated C-index values ranging from 0.9193 (SVM) to 0.9217 (DeepSurv).

Table 2.

C-index of radiomics and deep learning models using PET and CT features on the external validation set

Radiomics Model Deepling Learning Model
Model Training Validation Test Model Training Validation Test
PET Cox 0.6963 0.679 0.6711 Cox 0.8727 0.8718 0.904
SVM 0.6757 0.6658 0.6392 SVM 0.8656 0.8675 0.8803
RSF 0.7059 0.6979 0.6715 RSF 0.9007 0.874 0.8917
DeepSurv 0.6937 0.6782 0.6603 DeepSurv 0.8733 0.8735 0.9025
DeepCox 0.7096 0.6953 0.6665 DeepCox 0.8736 0.8778 0.9009
CT Cox 0.7325 0.6953 0.6642 Cox 0.7993 0.7973 0.8233
SVM 0.6923 0.6682 0.6872 SVM 0.7764 0.773 0.8119
RSF 0.7336 0.6981 0.7109 RSF 0.8374 0.7871 0.8252
DeepSurv 0.7372 0.7041 0.6677 DeepSurv 0.7968 0.7978 0.82
DeepCox 0.7814 0.7198 0.6778 DeepCox 0.8031 0.8103 0.8198
PET + CT Cox 0.757 0.7075 0.6817 Cox 0.8958 0.8906 0.9201
SVM 0.7064 0.6817 0.693 SVM 0.8817 0.88 0.9193
RSF 0.7451 0.7155 0.7302 RSF 0.9121 0.8917 0.9184
DeepSurv 0.7492 0.6975 0.6671 DeepSurv 0.8953 0.8927 0.9217
DeepCox 0.8065 0.7286 0.7017 DeepCox 0.8945 0.9007 0.9208

Figure 4 compares the performance of different features and models on the external validation set. As shown in the figure, the PET + CT model consistently outperforms single-modality models in most cases, followed by deep learning models based on PET and CT. Prognostic models utilizing radiomics features generally exhibit slightly lower performance, although the combination of PET and CT still surpasses the performance of single-modality models. Among models using radiomics features from CT and PET + CT, the RSF model achieves the highest performance. In deep learning models based on PET, the Cox model demonstrates relatively strong performance, while in models incorporating both PET and CT, DeepSurv yields the best results. Among the five models evaluated, the SVM model shows comparatively lower overall performance.

Fig. 4.

Fig. 4

Performance comparison of multi-modality radiomics and deep learning features across models on the external validation set (A: Performance of multi-modality radiomics features across different models on the external validation set, B: Performance of multi-modality deep learning features across different models on the external validation set)

Results of multimodal radiomics features in multi-model prognostic evaluation

Figure 5 presents Kaplan-Meier survival curves based on radiomics features derived from three imaging modalities (PET, CT, PET + CT) and evaluated across various models (Cox, SVM, RSF, DeepCox, DeepSurv). The curves illustrate the survival probabilities of high-risk and low-risk groups over time. Statistical significance between these groups is assessed using p-values. DeepSurv exhibits superior discriminative power in the PET and PET + CT modalities, with a p-value of 0.0003 in the PET + CT modality, surpassing traditional models such as Cox and SVM. Furthermore, the integration of multi-modal data (PET + CT) enhances model discrimination in Cox, RSF, and DeepSurv, significantly outperforming single-modality data. However, the SVM model experiences a notable decline in performance when combining PET and CT features, failing to distinguish high-risk from low-risk groups effectively (p-value = 0.1126). In contrast, the fusion of PET + CT in the Cox and DeepCox models results in a slight performance decrease compared to single-modality models.

Fig. 5.

Fig. 5

Performance of radiomics features in multi-modality multi-prognostic models on the external validation set

Results of multimodal deep learning features in multi-model prognostic evaluation

Figure 6 presents Kaplan-Meier survival curves based on deep learning features extracted from three imaging modalities (PET, CT, PET + CT) and evaluated using various models (Cox, SVM, RSF, DeepCox, DeepSurv). The survival curves for the “high-risk” and “low-risk” groups show clear separation across all modalities, demonstrating the models’ effectiveness in distinguishing between risk levels. However, subtle variations in the degree of separation are observed among the models. Notably, the combination of PET + CT outperforms the single-modality models. DeepCox and DeepSurv exhibit the strongest separation across all modalities, with DeepSurv particularly excelling in the PET + CT modality, where it displays a narrower confidence interval, indicating superior stability and reliability.

Fig. 6.

Fig. 6

Performance of deep learning features in multi-modality multi-prognostic models on the external validation set

Prognostic evaluation using stacking ensemble learning models

Table 3 presents the C-index values for stacking-based ensemble models employing different modalities and feature combinations on the External Validation Set. Notably, the radiomics and deep learning model combining PET and CT achieved a C-index of 0.9345. Compared to the models in Table 2, most models in Table 3 demonstrated superior performance with the same modality and feature combination, underscoring their robustness on the External Validation Set. The only exception was the stacking ensemble model based on PET-derived deep learning features, which exhibited slightly lower performance than the base model using PET-based deep learning features.

Table 3.

C-index of ensemble learning models with different modalities and feature combinations on the external validation set

Radiomics Model Deepling Learning Model Deepling Learning Model+
Radiomics
PET 0.6862 0.895 0.9233
CT 0.7131 0.8232 0.9219
PET + CT 0.7324 0.9319 0.9345

Figure 7 presents Kaplan-Meier survival curves for the PET + CT-based ensemble learning model, incorporating radiomics features, deep learning features, and their combination. The analysis reveals statistically significant survival differences (p < 0.05) between high- and low-risk groups for both radiomics and deep learning feature-based models. Importantly, the inclusion of deep learning features substantially improved the model’s discriminatory power in predicting survival outcomes. Moreover, the ensemble model, which integrates both radiomics and deep learning features, further enhances the distinction between risk groups, showcasing superior predictive accuracy. When considered alongside the results in Figs. 5 and 6, it is clear that the ensemble learning model outperforms single-modality models in survival prediction.

Fig. 7.

Fig. 7

Performance of radiomics and deep learning features in the PET + CT ensemble learning model on the external validation set

Figure 8 illustrates a comparative analysis of time-dependent ROC curves and their corresponding AUC values for the combination of different feature sets and models in PET + CT, evaluated on the external validation set at three distinct time points: the first quartile (959.5 days), the second quartile (1534.0 days), and the third quartile (2283.5 days). At the first time point, the AUC of the radiomics model ranges from 0.68 (SVM) to 0.76 (Stacking), while the AUC of the deep learning model increases to 0.96 (Cox) and 0.97 (Stacking). The fusion of radiomics and deep learning features further enhances model performance, with the AUC of the stacking model across all time points ranging from 0.97 to 0.98. Notably, the stacking ensemble model consistently outperforms other individual models. These results highlight the superior predictive performance of multi-modal feature fusion and stacking models in survival prediction.

Fig. 8.

Fig. 8

ROC curves of prediction accuracy for radiomics and deep learning features in different prognostic models at 1/4, 2/4, and 3/4 time points on the external validation set

Discussion

The robustness and reproducibility of radiomics models are critical prerequisites for their broad clinical application [3335]. In this study, we investigate a prognostic evaluation method for HNC patients, leveraging radiomics and deep learning features derived from cases across nine different centers. Our findings demonstrate that integrating multiple machine learning and deep learning models within a stacked ensemble learning framework markedly improves both the accuracy and robustness of predictions.

The performance of ensemble learning models is largely determined by the quality of the base models and the training data [36]. In this study, we integrate traditional models such as Cox, SVM, and RSF with DeepCox and DeepSurv, employing the widely validated DeepSurv as the meta-model [3740] to further enhance the performance and stability of the ensemble. Our results indicate that the performance of individual models is highly dependent on feature type and image modality. Specifically, significant discrepancies were observed both across different models using the same features and within the same model using different feature sets. Specifically, for radiomics-based features, deep learning models (e.g., DeepCox and DeepSurv) outperform Cox and SVM in predictive accuracy, though they still lag behind RSF. Interestingly, RSF, when utilizing traditional radiomics features, exceeds the performance of other models; however, its efficacy diminishes when deep learning features are incorporated, falling behind both DeepCox and DeepSurv. A similar trend is observed for Cox and SVM models.

Further analysis reveals that the performance of different models varies depending on the imaging modality and feature type used. Taking the SVM model as an example, as shown in Fig. 5F, when combined with traditional radiomics features, SVM was unable to effectively distinguish between high- and low-risk groups in the PET + CT model. However, as shown in Fig. 6F, when using deep learning features, the performance of the SVM in distinguishing between high-risk and low-risk groups is significantly superior to that of other models. The results in Table 2; Fig. 4 also highlight that SVM’s performance across different modalities and feature combinations is less stable than that of other models, particularly when deep learning features are included, where it performs notably worse. Notably, when only traditional radiomics features from CT images are used, SVM outperforms Cox, DeepSurv, and DeepCox models on the external validation set.

Furthermore, a “specific case” presented in Figs. 5 and 6 provides additional support for the above findings. This case resulted in death approximately 6300 days after diagnosis. As shown in Fig. 5, models based on radiomics features classified the case as high risk, with the exception of those in Fig. 5E and F. In contrast, models utilizing deep learning features in Fig. 6 classified the case as low risk. Since the patient’s death occurred 6,300 days after tumor treatment and the exact cause remains unclear, classifying the case as either high-risk or low-risk is considered reasonable. This result also indicates that variations in features and models can lead to differences in predictive outcomes.

The results of this study suggest that in certain cases, when base models are too similar or exhibit poor performance, the stacking model may underperform relative to the individual models, aligning with the findings of Kunapuli et al. [41]. As shown in Table 3, the ensemble model incorporating PET deep learning features (C-index = 0.895) demonstrates comparable predictive performance to the base models, such as the Cox model (C-index = 0.904), DeepSurv model (C-index = 0.9025), and DeepCox model (C-index = 0.9009). Although the C-index of the ensemble model is slightly lower, the overall performance differences are minimal.

Overall, the stacking model outperformed individual models in C-index across multimodal and multi-feature settings, and most clearly distinguished high- and low-risk groups in Kaplan–Meier analysis. The integration of radiomic and deep learning features further improved its performance. Although the improvements were modest (as shown in Tables 2 and 3), they highlight the complementary nature of radiomic and deep learning features in enhancing the model’s ability to capture tumor heterogeneity and texture.

This study has several limitations. First, although PET and CT data from nine centers were included, the sample size per center was limited, and variations in imaging equipment, acquisition protocols, and preprocessing introduced bias. Despite incorporating both traditional radiomics and deep learning features, standardized imaging and feature extraction procedures were lacking. Second, feature harmonization remains a challenge [42]. Although the NeuroCombat method was applied, it is a linear harmonization technique based on an empirical Bayes framework and fails to remove nonlinear multi-center batch effects [29]. Finally, the stacked framework requires substantial computational resources and is highly dependent on the performance of base models, which may affect its generalizability [4345].

In summary, the stacked ensemble model, integrating multimodal imaging and diverse features, effectively improves the accuracy and robustness of HNC prediction. Future research should focus on harmonizing multi-center data, developing more advanced deep learning feature extractors, and incorporating clinical variables to enhance predictive performance and generalizability, thereby supporting personalized treatment and risk stratification in HNC.

Abbreviations

HNC

Head and Neck Cancer

TCIA

The Cancer Imaging Archive

HPV

Human Papillomavirus

PET

Positron Emission Tomography

CT

Computed Tomography

MRI

Magnetic Resonance Imaging

Cox

Cox Proportional Hazards Model

RSF

Random Survival Forest

SVM

Support Vector Machine

DeepCox

Deep Learning-based Cox Proportional Hazards Model

DeepSurv

Deep Learning-based Survival Model

ROC

Receiver Operating Characteristic Curve

AUC

Area Under the Curve

CHUS

Centre Hospitalier Universitaire de Sherbrooke

CHUM

Centre Hospitalier de l’Université de Montréal

HGJ

Jewish General Hospital

HMR

Hôpital Maisonneuve-Rosemont

MDACC

MD Anderson Cancer Center

MAASTRO

Maastricht Radiation Oncology

QIN

Quantitative Imaging Network

TCGA

The Cancer Genome Atlas

HNSCC

Head and Neck Squamous Cell Carcinoma

GTV

Gross Tumor Volume

RTSTRUCT

Radiotherapy Structure Set

SEG

Segmentation

SUV

Standardized Uptake Value

Author contributions

B.W. and J.L. conceptualized the study and designed the methodology. X.Z., J.L., and S.L. performed the data collection and analysis. Z.W. and Z.C. contributed to the validation and interpretation of the results. D.W. and T.L. were responsible for software development and data visualization. H.R.H.R., H.H.H., and W.Z.W.H. supervised the study and provided critical revisions. X.D. contributed to manuscript drafting and review. All authors participated in manuscript writing, revision, and approval of the final version.

Funding

This work was supported by Hebei Province Introduced Returned Overseas Chinese Scholars Funding Project (C20220107), Chengde Biomedicine Industry Research Institute Funding project (202205B086), Chengde Medical university Project (202307,202404), Science and Technology Project of Hebei Education Department (BJK2023060).

Data availability

No datasets were generated or analysed during the current study.

Declarations

Ethics approval and consent to participate

All data used in this study is sourced from the publicly available TCIA (The Cancer Imaging Archive) dataset (https://www.cancerimagingarchive.net/). This dataset complies with the data usage agreement, and all participants’ privacy has been protected through de-identification, ensuring that no personal identification information is disclosed. The use of this data does not involve obtaining informed consent from participants, and the study adheres to relevant ethical guidelines and international standards for data protection laws. The data is used solely for academic research and in accordance with the terms of use provided by the dataset provider.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Wan Zuha Wan Hasan, Email: wanzuha@upm.edu.my.

Xianling Dong, Email: dongxl@cdmc.edu.cn.

References

  • 1.Samara P, Athanasopoulos M, Mastronikolis S, Kyrodimos E, Athanasopoulos I, Mastronikolis NS. The role of oncogenic viruses in head and neck cancers: epidemiology, pathogenesis, and advancements in detection methods. Microorganisms. 2024;12(7). [DOI] [PMC free article] [PubMed]
  • 2.Chow LQM. Head and neck Cancer. N Engl J Med. 2020;382(1):60–72. [DOI] [PubMed] [Google Scholar]
  • 3.Yan K, Agrawal N, Gooi Z. Head and neck masses. Med Clin North Am. 2018;102(6):1013–25. [DOI] [PubMed] [Google Scholar]
  • 4.Zhuang RY, Xu HG. Head and neck Cancer. N Engl J Med. 2020;382(20):e57. [DOI] [PubMed] [Google Scholar]
  • 5.Huang SH, O’Sullivan B. Overview of the 8th edition TNM classification for head and neck Cancer. Curr Treat Options Oncol. 2017;18(7):40. [DOI] [PubMed] [Google Scholar]
  • 6.Bhatia A, Burtness B. Treating head and neck Cancer in the age of immunotherapy: A 2023 update. Drugs. 2023;83(3):217–48. [DOI] [PubMed] [Google Scholar]
  • 7.Aerts HJ, Velazquez ER, Leijenaar RT, Parmar C, Grossmann P, Carvalho S, et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nat Commun. 2014;5:4006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lambin P, Rios-Velazquez E, Leijenaar R, Carvalho S, van Stiphout RG, Granton P, et al. Radiomics: extracting more information from medical images using advanced feature analysis. Eur J Cancer. 2012;48(4):441–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zwanenburg A, Vallières M, Abdalah MA, Aerts H, Andrearczyk V, Apte A, et al. The image biomarker standardization initiative: standardized quantitative radiomics for High-Throughput image-based phenotyping. Radiology. 2020;295(2):328–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ai Y, Liu J, Li Y, Wang F, Du X, Jain RK et al. SAMA: A Self-and-Mutual attention network for accurate recurrence prediction of Non-Small cell lung Cancer using genetic and CT data. IEEE J Biomed Health Inf. 2024;Pp. [DOI] [PubMed]
  • 11.Chang B, Geng Z, Mei J, Wang Z, Chen P, Jiang Y, et al. Application of multimodal deep learning and multi-instance learning fusion techniques in predicting STN-DBS outcomes for parkinson’s disease patients. Neurotherapeutics. 2024;21(6):e00471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Liao H, Yuan J, Liu C, Zhang J, Yang Y, Liang H, et al. One novel transfer learning-based CLIP model combined with self-attention mechanism for differentiating the tumor-stroma ratio in pancreatic ductal adenocarcinoma. Radiol Med. 2024;129(11):1559–74. [DOI] [PubMed] [Google Scholar]
  • 13.Regodić M, Bardosi Z, Freysinger W. Automated fiducial marker detection and localization in volumetric computed tomography images: a three-step hybrid approach with deep learning. J Med Imaging (Bellingham). 2021;8(2):025002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Shiri I, Rahmim A, Ghaffarian P, Geramifar P, Abdollahi H, Bitarafan-Rajabi A. The impact of image reconstruction settings on 18F-FDG PET radiomic features: multi-scanner Phantom and patient studies. Eur Radiol. 2017;27(11):4498–509. [DOI] [PubMed] [Google Scholar]
  • 15.Spaanderman DJ, Hakkesteegt SN, Hanff DF, Schut ARW, Schiphouwer LM, Vos M, et al. Multi-center external validation of an automated method segmenting and differentiating atypical lipomatous tumors from lipomas using radiomics and deep-learning on MRI. EClinicalMedicine. 2024;76:102802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Grau ET, Charles M, Féménia M, Rebours E, Vaiman A, Rocha D. Survey of mitochondrial sequences integrated into the bovine nuclear genome. Sci Rep. 2020;10(1):2077. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Orellana B, Monclús E, Brunet P, Navazo I, Bendezú Á, Azpiroz F. A scalable approach to T2-MRI colon segmentation. Med Image Anal. 2020;63:101697. [DOI] [PubMed] [Google Scholar]
  • 18.Chen C, Mat Isa NA, Liu X. A review of convolutional neural network based methods for medical image classification. Comput Biol Med. 2024;185:109507. [DOI] [PubMed] [Google Scholar]
  • 19.Chen J, Wee L, Dekker A, Bermejo I. Using 3D deep features from CT scans for cancer prognosis based on a video classification model: A multi-dataset feasibility study. Med Phys. 2023;50(7):4220–33. [DOI] [PubMed] [Google Scholar]
  • 20.Santinha J, Pinto Dos Santos D, Laqua F, Visser JJ, Groot Lipman KBW, Dietzel M, et al. ESR essentials: radiomics-practice recommendations by the European society of medical imaging informatics. Eur Radiol. 2025;35(3):1122–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Mahajan P, Uddin S, Hajati F, Moni MA. Ensemble learning for disease prediction: A review. Healthc (Basel). 2023;11(12). [DOI] [PMC free article] [PubMed]
  • 22.Parvez A, Ali SD, Tayara H, Chong KT. Stacking based ensemble learning framework for identification of Nitrotyrosine sites. Comput Biol Med. 2024;183:109200. [DOI] [PubMed] [Google Scholar]
  • 23.Saad MB, Hong L, Aminu M, Vokes NI, Chen P, Salehjahromi M, et al. Predicting benefit from immune checkpoint inhibitors in patients with non-small-cell lung cancer by CT-based ensemble deep learning: a retrospective study. Lancet Digit Health. 2023;5(7):e404–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhang K, Liu C, Pan J, Zhu Y, Li X, Zheng J, et al. Use of MRI-based deep learning radiomics to diagnose sacroiliitis related to axial spondyloarthritis. Eur J Radiol. 2024;172:111347. [DOI] [PubMed] [Google Scholar]
  • 25.Albano D, Gatta R, Marini M, Rodella C, Camoni L, Dondi F et al. Role of (18)F-FDG PET/CT radiomics features in the differential diagnosis of solitary pulmonary nodules: diagnostic accuracy and comparison between two different PET/CT scanners. J Clin Med. 2021;10(21). [DOI] [PMC free article] [PubMed]
  • 26.van Griethuysen JJM, Fedorov A, Parmar C, Hosny A, Aucoin N, Narayan V, et al. Computational radiomics system to Decode the radiographic phenotype. Cancer Res. 2017;77(21):e104–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Huang G, Liu Z, Van Der Maaten L, Weinberger KQ, editors. Densely connected convolutional networks. Proceedings of the IEEE conference on computer vision and pattern recognition; 2017.
  • 28.Van Ness M, Udell M. Interpretable prediction and feature selection for survival analysis. ArXiv Preprint arXiv:240414689. 2024.
  • 29.Richter S, Winzeck S, Correia MM, Kornaropoulos EN, Manktelow A, Outtrim J, et al. Validation of cross-sectional and longitudinal combat harmonization methods for magnetic resonance imaging data on a travelling subject cohort. Neuroimage: Rep. 2022;2(4):100136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Pavlyshenko B, editor. Using stacking approaches for machine learning models. 2018 IEEE second international conference on data stream mining & processing (DSMP); 2018: IEEE.
  • 31.Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. 2018;18:1–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Nagpal C, Yadlowsky S, Rostamzadeh N, Heller K, editors. Deep cox mixtures for survival regression. Machine Learning for Healthcare Conference; 2021: PMLR.
  • 33.Duman A, Sun X, Thomas S, Powell JR, Spezi E. Reproducible and interpretable machine Learning-Based radiomic analysis for overall survival prediction in glioblastoma multiforme. Cancers (Basel). 2024;16:19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kallos-Balogh P, Vas NF, Toth Z, Szakall S, Szabo P, Garai I, et al. Multicentric study on the reproducibility and robustness of PET-based radiomics features with a realistic activity painting Phantom. PLoS ONE. 2024;19(10):e0309540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Santinha J, Pinto Dos Santos D, Laqua F, Visser JJ, Groot Lipman KBW, Dietzel M et al. ESR essentials: radiomics-practice recommendations by the European society of medical imaging informatics. Eur Radiol. 2024. [DOI] [PMC free article] [PubMed]
  • 36.Mohammed A, Kora R. A comprehensive review on ensemble deep learning: opportunities and challenges. J King Saud University-Computer Inform Sci. 2023;35(2):757–74. [Google Scholar]
  • 37.Adam N, Wieder R. AI survival prediction modeling: the importance of considering treatments and changes in health status over time. Cancers (Basel). 2024;16(20). [DOI] [PMC free article] [PubMed]
  • 38.Germer S, Rudolph C, Labohm L, Katalinic A, Rath N, Rausch K, et al. Survival analysis for lung cancer patients: A comparison of Cox regression and machine learning models. Int J Med Inf. 2024;191:105607. [DOI] [PubMed] [Google Scholar]
  • 39.Jiao Y, Ye J, Zhao W, Fan Z, Kou Y, Guo S, et al. Development and validation of a deep learning-based survival prediction model for pediatric glioma patients: A retrospective study using the SEER database and Chinese data. Comput Biol Med. 2024;182:109185. [DOI] [PubMed] [Google Scholar]
  • 40.Teshale AB, Htun HL, Vered M, Owen AJ, Ryan J, Tonkin A et al. Artificial intelligence improves risk prediction in cardiovascular disease. Geroscience. 2024. [DOI] [PMC free article] [PubMed]
  • 41.Kunapuli G. Ensemble methods for machine learning. Simon and Schuster; 2023.
  • 42.Atasever S, Azginoglu N, Terzi DS, Terzi R. A comprehensive survey of deep learning research on medical image analysis with focus on transfer learning. Clin Imaging. 2023;94:18–41. [DOI] [PubMed] [Google Scholar]
  • 43.Aboneh T, Rorissa A, Srinivasagan R. Stacking-based ensemble learning method for multi-spectral image classification. Technologies. 2022;10(1):17. [Google Scholar]
  • 44.Dey R, Mathur R, editors. Ensemble learning method using stacking with base learner, a comparison. International Conference on Data Analytics and Insights; 2023: Springer.
  • 45.Dong X, Yu Z, Cao W, Shi Y, Ma Q. A survey on ensemble learning. Front Comput Sci. 2020;14:241–58. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from Radiation Oncology (London, England) are provided here courtesy of BMC

RESOURCES