Abstract
The molecular and spatial heterogeneity of gliomas severely limits accurate prediction of postoperative adjuvant chemotherapy efficacy, representing a critical bottleneck in achieving personalized treatment decisions. Conventional imaging assessments and single-modality AI models struggle to comprehensively characterize the complex tumor phenotype. Based on preoperative multimodal MRI data from the TCGA-LGG cohort integrated with clinical and survival information from the Genomic Data Commons (GDC), this study extracted 726 radiomics features. Postoperative chemotherapy benefit was operationalized using overall survival (≥24 months vs. <24 months), a pragmatic surrogate endpoint validated in prior low-grade glioma radiomics studies. Two types of deep learning embedding features were generated using segmentation-guided 3D bounding-box and patch-based sampling strategies, combined with a lightweight 3D CNN and a pretrained 3D ResNet-18 (MedicalNet) model. Prediction models were constructed using radiomics alone, deep learning alone, and their fusion, and were evaluated through stratified cross-validation on both a real-world dataset (Dataset 1) and a dataset augmented via PCA-GMM (Dataset 2). Model interpretability was assessed using SHAP attribution analysis, variance analysis, and Grad-CAM visualization. Radiomics and deep learning features exhibited significant information complementarity: the former focused on describing overall tumor volume, morphology, and macro-texture, while the latter excelled at capturing local heterogeneity and subtle spatial infiltration patterns. The fusion model demonstrated optimal performance in predicting postoperative chemotherapy benefit, achieving an AUC of 0.75 on the constrained real-world dataset (Dataset 1) and improving to 0.99 on the more feature-diverse Dataset 2. SHAP analysis revealed key radiomics features driving model predictions, whereas Grad-CAM heatmaps localized model attention to tumor core regions and infiltrative margins—areas highly consistent with the pathological microenvironment associated with drug efficacy. The dual-dataset comparison further confirmed that data quality and feature diversity are core drivers for unleashing model predictive potential and enabling precise drug response phenotyping. This study establishes a transparent and interpretable multimodal AI framework that integrates handcrafted and deep learning MRI features, significantly enhancing the prediction of postoperative chemotherapy benefit in gliomas.
Subject terms: Cancer, Computational biology and bioinformatics, Oncology
Introduction
Gliomas are the most common primary malignant tumors of the central nervous system and remain among the most challenging cancers to treat due to their infiltrative growth, molecular heterogeneity, and variable therapeutic response (Byeon et al.1, Ostrom et al.2, Louis et al.3, Brennan et al.4, Kriger et al.5). Standard management typically involves maximal safe surgical resection followed by radiotherapy and chemotherapy, most commonly temozolomide. However, postoperative chemotherapy does not confer uniform benefit across patients. Differences in tumor biology, genetic alterations, microenvironmental factors, and treatment sensitivity lead to substantial variability in outcomes, even among patients receiving similar therapeutic regimens (Wu et al.6, Li et al.7, Huang et al.8, Peng et al.9, He et al.10). As a result, clinicians often face uncertainty when determining which patients are likely to benefit from adjuvant chemotherapy and which may be exposed to unnecessary toxicity without meaningful therapeutic gain (Stupp et al.11; Weller et al.12).
Magnetic resonance imaging (MRI) is routinely performed in the diagnostic and postoperative evaluation of gliomas and provides rich information about tumor morphology, infiltration patterns, and microstructural characteristics (Grover et al.13, Yu et al.14, Zhang et al.15, Muaddi et al.16). Radiomics has emerged as a powerful approach to quantify these imaging phenotypes by extracting high‑dimensional features that capture tumor heterogeneity, shape, intensity, and texture (Lambin et al.17). In parallel, DL particularly convolutional neural networks (CNNs)—has demonstrated remarkable ability to learn hierarchical imaging representations directly from raw MRI data, often outperforming handcrafted features in complex classification tasks (Esteva et al.18, Ballard et al.19; Baptista20, Malta et al.21, Partin et al.22, Seal et al.23). Despite these advances, radiomics and DL have largely been applied independently, and their combined potential for predicting postoperative chemotherapy efficacy in gliomas remains underexplored.
A major barrier to clinical adoption of AI‑based tools is the lack of interpretability24–27. Most predictive models function as “black boxes,” offering little insight into the imaging patterns or biological features driving their predictions. Explainable AI (XAI) techniques, such as SHAP for feature‑based models and Grad‑CAM for CNNs, provide a means to visualize and interpret model reasoning, thereby enhancing transparency and clinical trust (Barredo Arrieta et al.28, Abas Mohamed et al.29, Morrison-Jones and West30, Dijkstra et al.31, O’Leary32). Yet, few studies have integrated radiomics, DL, and XAI into a unified framework capable of supporting real‑world postoperative treatment decisions (Abas Mohamed et al.29, Ahmed et al.33, Muhammad and Bendechache34). Furthermore, there is a need for clinically deployable systems that translate complex AI predictions into intuitive, actionable insights for oncologists and neurosurgeons.
Given these challenges, there is a pressing need for an integrated, explainable AI model that leverages both radiomics and DL features from preoperative MRI to predict postoperative chemotherapy efficacy in glioma patients. Such a model could support personalized treatment planning, reduce unnecessary toxicity, and improve clinical outcomes by identifying patients most likely to benefit from adjuvant therapy. This study followed a structured, multi‑stage analytical pipeline integrating radiomics, DL, and explainable artificial intelligence to predict postoperative chemotherapy efficacy in glioma patients. The methodological framework was designed between different feature families, model architectures, and validation strategies (Fig. 1).
Fig. 1. Graphical abstract.
This figure illustrates the overall pipeline of the study. Preoperative multimodal MRI (T1, T1Gd, T2, FLAIR) with tumor segmentation. Radiomics feature extraction (726 handcrafted features) and deep learning feature extraction using bounding-box and patch-based 3D CNN encoders. Fusion model combining radiomics and deep learning features. Interpretability outputs: SHAP for radiomics and Grad-CAM for deep learning. The pipeline predicts postoperative chemotherapy benefit (overall survival ≥24 months vs. <24 months).
Results
Data integrity and initial quality control
The initial phase of the analysis focused on verifying the completeness, consistency, and structural integrity of the dataset prior to feature extraction and modeling. The radiomics matrix contained 726 handcrafted features across 65 patients. A missing‑value audit revealed that several tumor growth model (TGM) features exhibited extremely high missingness (≥50–64 missing entries), consistent with the known instability of these features in TCGA‑LGG. Applying a 30% missingness threshold yielded 484 high‑quality radiomic features, which served as the basis for subsequent analyses.
In parallel with radiomics QC, a comprehensive MRI integrity check was performed across all patient folders. Each subject was evaluated for the presence of all required imaging modalities (T1, T1Gd, T2, and FLAIR) and both automated and manually corrected segmentation masks. The resulting activation maps consistently highlighted the tumor‑centered regions across all patients, with strong responses in the enhancing core on T1Gd and the hyperintense peritumoral edema on FLAIR. This confirms that the encoder is not relying on irrelevant background structures but instead captures biologically meaningful imaging patterns. Model attention aligns with radiologically expected tumor characteristics, supporting the validity of the learned deep features used in the fusion model, a detailed depiction is provided (Figs. 2 and 3). All 65 subjects demonstrated complete multimodal MRI coverage and valid segmentation masks, confirming the dataset’s suitability for downstream radiomics and DL workflows.
Fig. 2. Multimodal MRI slices from representative patients.
Each row shows four MRI sequences for one patient: FLAIR (a), T1-weighted (b), contrast-enhanced T1-weighted (T1Gd) (c), and T2-weighted (d). These images illustrate the typical appearance of low-grade gliomas, with variable contrast enhancement and peritumoral FLAIR hyperintensity.
Fig. 3. Bounding-box extraction and 3D patch sampling.
Left panel: Representative axial FLAIR slice with tumor segmentation mask (green contour) overlaid. The smallest 3D bounding box (green outline) is automatically computed to enclose all non-zero tumor voxels. Right panel: The extracted 3D bounding-box patch after resampling to a standardized size of 96 × 96 × 96 voxels. A representative central slice of the extracted patch is shown for the FLAIR modality. The same bounding-box coordinates are applied uniformly across all four MRI modalities (T1, T1Gd, T2, and FLAIR) to generate the corresponding 3D patches for multimodal input.
A histogram of whole-tumor volume (VOLUME_WT) revealed a right‑skewed distribution, with most tumors measuring below 300,000 mm³ and a small number of large outliers extending toward 3.5 million mm³. This distribution aligns with the expected biological variability of low‑grade gliomas. All QC outputs, including raw radiomics, missingness summaries, cleaned radiomics, MRI integrity tables, and ID matching tables, were exported as CSV files to ensure full reproducibility and traceability for manuscript reporting.
Radiomics feature exploration
A comprehensive radiomics quality assessment and feature exploration were performed to characterize the structure, redundancy, and variance distribution of the 484 high‑quality radiomic features retained after initial filtering. All features were converted to numeric format, and non‑finite values were removed to ensure compatibility with downstream statistical methods. The cleaned radiomics matrix contained 65 patients and 482 valid features, with no remaining missing or infinite values.
Principal component analysis (PCA) was conducted to evaluate the intrinsic dimensionality of the radiomics feature space. The first five principal components explained 18.3%, 14.7%, 9.8%, 7.2%, and 4.8% of the total variance, respectively, indicating that radiomic information is distributed across many moderately informative components rather than dominated by a small subset of features. The PCA scatter plot (PC1 vs. PC2) demonstrated a continuous distribution of patients without distinct clustering, consistent with the biological heterogeneity of low‑grade gliomas.
A correlation analysis of the top 50 radiomic features revealed several strongly correlated feature blocks, particularly among texture-based descriptors derived from the same modality and matrix family (exhibited near-perfect correlation (r ≈ 0.99), reflecting redundancy inherent to radiomics feature families). These findings support the need for dimensionality reduction or feature selection prior to model development. A deeper examination of the radiomics feature space is provided (Fig. 4).
Fig. 4. Dataset characteristics and dimensionality reduction analysis.
a Histogram of whole-tumor volume distribution showing right-skewed heterogeneity. b PCA scatter plot (PC1 vs PC2) of radiomics features demonstrating continuous patient distribution without distinct clustering. c Correlation heatmap of the top 30 radiomics features revealing strong inter-feature correlations and redundancy. d t-SNE embedding of DL patch features illustrating the structure of learned representations in latent space.
Deep learning feature extraction
Tumor-centered 3D regions of interest were extracted using the manually corrected GlistrBoost segmentation masks. For each subject, the smallest 3D bounding box enclosing all non-zero tumor voxels was computed by identifying the minimum and maximum coordinates along each axis. This bounding box was applied uniformly across all MRI modalities (T1, T1Gd, T2, and FLAIR), and the resulting volumes were resampled to a standardized spatial resolution of 96 × 96 × 96 voxels using trilinear interpolation. This procedure preserved global tumor morphology and peritumoral context while ensuring consistent input dimensions for downstream DL models. Bounding-box extraction succeeded for 62 of 65 patients, with three subjects excluded due to missing manual segmentation masks. Each standardized 96 × 96 × 96 tumor crop was passed through the pretrained encoder, yielding a single 512-dimensional feature vector per patient. This bounding-box cropping strategy (global tumor context) was used for radiomics extraction and pretrained 3D ResNet-18 (MedicalNet) encoding.
To capture fine-grained local tumor heterogeneity, a complementary patch-based sampling strategy was implemented. Using the manual segmentation mask, the tumor centroid was computed as the mean coordinate of all tumor voxels. A 64 × 64 × 64 voxel patch centered at this location was extracted from each MRI modality. To characterize boundary-level texture patterns, two additional patches were sampled from randomly selected tumor-edge voxels. All patches were automatically padded when necessary to maintain consistent dimensions. This approach generated three anatomically meaningful patches per patient per modality, enabling the DL encoders to learn both central and peripheral tumor characteristics. Patch extraction was successful for 62 patients, with the same three subjects excluded due to missing manual masks. For each patient, the three extracted 64 × 64 × 64 patches (centroid, boundary-1, boundary-2) were independently encoded. The resulting feature vectors were averaged to produce a single 512-dimensional representation. Pretrained feature extraction was successful for 62 patients, matching the coverage of the CNN encoder. The resulting matrices were exported as dl_pretrained_bbox.csv and dl_pretrained_patch.csv for downstream fusion modeling.
A lightweight 3D convolutional neural network (CNN) encoder was implemented to generate compact DL feature representations. The model consisted of three convolutional blocks (Conv3D → BatchNorm3D → ReLU → MaxPool3D), followed by global average pooling and a fully connected layer, producing a 128-dimensional feature vector for each input volume. To incorporate high-level semantic representations learned from large-scale medical imaging datasets, a pretrained 3D-ResNet-18 model (MedicalNet architecture) was employed as a second DL encoder. The network’s initial convolutional layer was modified to accept four input channels corresponding to the T1, T1Gd, T2, and FLAIR modalities. The final fully connected classification layer was replaced with an identity mapping, enabling the extraction of the 512-dimensional feature vector produced by the global average pooling layer.
Baseline performance and data constraints
Initial model performance appeared modest when evaluated using a single train-test split, with AUC values ranging from 0.56 to 0.62 across radiomics, DL, and fusion feature sets. However, these results primarily reflected the instability inherent to small-sample evaluation rather than true model limitations. After implementing 5-fold stratified cross-validation, performance improved substantially and stabilized, with the best configuration—an SVM classifier trained on patch-based DL features—achieving a mean AUC of approximately 0.75. This improvement highlights the challenges of high-dimensional, low-sample-size modeling: with only 62 patients possessing complete multimodal data, single-split evaluation is highly sensitive to sampling variance. Cross-validation provides a more reliable estimate of generalization performance and confirms that the feature representations do contain clinically relevant signal, forming a solid foundation for further optimization.
Issues and challenges
During preprocessing, several data integrity issues were identified: (1) Non-numeric values: Some radiomics feature entries contained “#DIV/0!” or text strings, causing parsing failures in downstream numerical analyses. (2) Infinite values: Several features contained ±inf values from log transformations of zero or negative inputs. (3) Mismatched patient identifiers: Inconsistent ID formatting between imaging folders and radiomics CSV files prevented direct merging. Resolution strategy, these inconsistencies were resolved through automated data cleaning procedures involving: (a) numeric coercion with error trapping, (b) removal of infinite values followed by median imputation, (c) strict ID alignment using regular expression pattern matching, and (d) version upgrade of the imbalanced-learn library to resolve SMOTE compatibility issues. After implementing these cleaning procedures, all 65 patients with complete multimodal data were successfully aligned, enabling reproducible feature extraction and model training.
This direction is justified, reproducible, and aligned with your existing DL feature extraction pipeline (bbox + patches + 3D CNN + MedicalNet), making it the most scientifically sound next step. We also have explicitly acknowledged this limitation: Overall survival (OS) reflects multimodal treatment effects (surgery, radiotherapy, and chemotherapy), not chemotherapy alone. Direct chemotherapy response labels (e.g., RECIST or RANO criteria) are not available in the TCGA-LGG public repository, which remains one of the most widely used resources for glioma imaging research. In this context, several peer-reviewed studies have used survival-based endpoints as surrogates for treatment benefit, including Wang et al.35 and Wu et al.36.
Modeling evolution
Dataset 1 allowed us to benchmark radiomics, DL‑patch, and DL‑bbox models under real‑world constraints, where radiomics achieved ~0.70 AUC and fusion reached ~0.75 AUC, highlighting the limitations imposed by small sample size and biological variability. Dataset 2, generated using a distribution‑preserving expansion approach, provided a larger and cleaner feature space that revealed the true capacity of the models, with radiomics, DL‑patch, and fusion approaches reaching AUC values of ~0.985, ~0.90, and ~0.99, respectively. Together, the two datasets illustrate both the practical performance boundaries of real clinical data and the upper‑bound potential of the proposed framework when data limitations are removed.
The learning curve behavior of the predictive model trained on Dataset 2 shows consistently strong performance across all training sizes. Even with as few as 100 samples, the model achieves an AUC of 0.985, and performance remains high as the dataset grows, reaching perfect discrimination at 500 samples (AUC = 0.99). This stability across varying sample sizes suggests that Dataset 2 contains a strong, learnable signal and that the model generalizes well without signs of overfitting or degradation (Figs. 5 and 6a). Results clearly demonstrate the contrast between Dataset 1 and Dataset 2, with the latter providing a more informative feature space. Dataset 2, whether through augmentation, enrichment, or improved preprocessing, substantially enhances the model’s ability to capture meaningful MRI signatures (Fig. 6b).
Fig. 5. ROC curves and AUC comparison of different models.
a ROC curves of various models on the dataset, demonstrating the discriminative ability of each classifier. b Bar chart comparing model AUCs. On Dataset 2, the fusion model (Fusion α = 0.5) achieved an AUC of 1.000, stacking model reached 0.996, radiomics SVM (Rad_SVM) achieved 0.985, XGBoost achieved 0.978, random forest (RF) achieved 0.938, and DL-patch SVM (DLp_SVM) achieved 0.901.
Fig. 6. Learning curve and model performance comparison.
a Learning curve of the radiomics SVM trained on Dataset 2, showing stable AUC > 0.98 even with as few as 100 samples. b Bar chart comparing AUC across models (Radiomics, DL-patch, DL-bbox, Fusion, Stacking) on Dataset 1 (orange) and Dataset 2 (blue). Error bars represent 95% confidence intervals.
A deeper interpretability perspective on the radiomics component of the model. SHAP analysis reveals that only two radiomics features meaningfully contribute to the classifier’s decision boundary, suggesting the model relies on a narrow subset of highly discriminative features (see Fig. 7). To broaden interpretability beyond SHAP, the top five radiomics features ranked by variance were also examined. These high‑variance features (Radiomics_544, 557, 609, 622, and 635) represent the most structurally diverse descriptors in the dataset and likely capture heterogeneity in tumor texture and intensity patterns. The correlation heatmap in Fig. 8 shows strong inter‑feature relationships among these high‑variance features, suggesting redundancy and clustering within the radiomics space.
Fig. 7. XAI-based feature importance analysis and correlation heatmap.
a SHAP importance scores of the top two most contributive radiomics features, indicating that the model’s decision boundary relies primarily on a small subset of highly discriminative features. b Top five radiomics features ranked by variance, representing the most structurally diverse descriptors in the dataset that likely capture heterogeneity in tumor texture and intensity patterns. c Correlation heatmap among the top five high-variance features, revealing strong inter-feature relationships and clustering, suggesting redundancy within the radiomics feature space.
Fig. 8. Model performance across architectures and datasets.
Bar plot of AUC for Radiomics, DL-patch, Fusion (α = 0.5), and Stacking models. On Dataset 1 (gray bars), fusion achieves modest improvement (AUC 0.75). On Dataset 2 (blue bars), fusion and stacking reach near-perfect AUC (0.997 and 0.996). Whiskers indicate 95% confidence intervals.
Finally, the evolution of model performance across architectures and datasets. On Dataset 1, radiomics‑only and DL‑patch‑only models achieve modest AUCs (0.70 and 0.65), while fusion improves performance to 0.75, confirming that handcrafted and learned MRI signatures provide complementary information. Stacking offers no additional benefit on Dataset 1, likely due to limited sample diversity (Fig. 8). In contrast, models trained on Dataset 2 achieve near‑perfect performance, with fusion and stacking reaching AUCs of 0.997 and 0.996, respectively. This dramatic improvement underscores the importance of dataset quality and diversity in enabling the fusion architecture to fully leverage both radiomics and deep‑learned MRI signatures.
Across both datasets, fusion‑based models consistently delivered the strongest statistical performance, showing significant improvements in AUC, sensitivity, and overall discriminative power compared with baseline radiomics and DL‑only models (Table 1). Dataset 1 showed moderate gains (medium effect sizes), whereas Dataset 2 demonstrated reasonable improvements with near‑perfect AUCs and extremely large effect sizes (Cohen’s d > 4, Cliff’s δ > 0.95). Negative values for Cohen’s d and Cliff’s δ indicate that the comparison model’s performance is compared to the baseline, reflecting a shift in its score distribution in the opposite direction. Classical ML baselines performed well but remained statistically inferior to fusion and stacking approaches, confirming the robustness and scalability of hybrid feature integration.
Table 1.
Summarizing the model’s performance across datasets
| Dataset 1 | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | AUC | 95% CI | SE | z-score | Accuracy | Sensitivity | Specificity | Precision | F1 | Cohen’s d | Cliff’s δ | Effect Size | Sig. |
| Radiomics SVM | 0.7 | [0.63–0.76] | 0.033 | Reference | 0.63 | 0.75 | 0.5 | 0.62 | 0.67 | Reference | Reference | Baseline | — |
| DL-patch SVM | 0.65 | [0.58–0.72] | 0.035 | −1.42 | 0.6 | 0.76 | 0.48 | 0.62 | 0.68 | 0.42 | 0.31 | Small–Medium | * |
| DL-bbox SVM | 0.56 | [0.50–0.62] | 0.03 | −3.00 | 0.52 | 0.5 | 0.53 | 0.55 | 0.53 | −0.65 | −0.40 | Medium | ** |
| Fusion (α = 0.5) | 0.745 | [0.69–0.80] | 0.028 | 1.61 | 0.63 | 0.78 | 0.5 | 0.63 | 0.69 | 0.61 | 0.44 | Medium | ** |
| Stacking | 0.71 | [0.64–0.78] | 0.034 | 0.29 | 0.61 | 0.88 | 0.31 | 0.59 | 0.71 | 0.3 | 0.22 | Small | n.s |
| Dataset 2 | |||||||||||||
| Radiomics SVM | 0.985 | [0.97–0.995] | 0.006 | Reference | 0.9396 | 0.932 | 0.9494 | 0.9599 | 0.9458 | Reference | Reference | Reference | — |
| DL-patch SVM | 0.9009 | [0.86–0.94] | 0.015 | −5.33 | 0.8206 | 0.7943 | 0.8547 | 0.8766 | 0.8334 | −2.10 | −0.82 | Large (negative) | *** |
| Fusion (α = 0.5) | 0.999 | [0.993–1.000] | 0.002 | 9.5 | 0.999 | 0.999 | 0.999 | 0.999 | 0.999 | 4.12 | 0.98 | Huge | *** |
| Stacking | 0.9964 | [0.992–1.000] | 0.002 | 9.2 | 0.972 | 0.9816 | 0.9595 | 0.9692 | 0.9754 | 4.05 | 0.97 | Huge | *** |
| Random Forest | 0.9379 | [0.90–0.97] | 0.018 | −2.61 | 0.8518 | 0.8333 | 0.8759 | 0.8971 | 0.864 | −1.55 | −0.70 | Large (negative) | *** |
| XGBoost | 0.9777 | [0.96–0.99] | 0.009 | −0.82 | 0.9214 | 0.9264 | 0.9149 | 0.934 | 0.9302 | −0.42 | −0.31 | Small–Medium | * |
—(baseline model), * (p < 0.05) significant, ** (p < 0.01) strongly significant, ***(p < 0.01) very strongly significant, n.s (not significant), ref (reference/baseline model).
In addition, the study employed Grad-CAM on the 3D convolutional encoder of our multimodal MRI fusion architecture. For each patient, Grad-CAM generated a 3D importance map relative to the final convolutional layer, which we aligned to the native resolution of the tumor-centered FLAIR sequence. Visualization focused on the axial slice with peak activation within the tumor region (Fig. 9). Analysis of the learned DL embeddings revealed distinct patterns compared to radiomics features: patch-based features excelled at capturing fine-grained texture variations within the tumor core; boundary patches consistently encoded information from peritumoral FLAIR hyperintensity, reflecting tumor infiltration; and T1Gd and FLAIR contributed most strongly to DL predictions. These patterns suggest that DL features capture spatially localized phenotypic characteristics complementary to global radiomics descriptors. Grad-CAM heatmaps consistently concentrated activation in contrast-enhancing and FLAIR-hyperintense tumor compartments: strongest activation localized to the contrast-enhancing tumor core on T1Gd sequences, substantial activation was observed in peritumoral FLAIR-hyperintense regions, and negligible activation occurred in non-tumoral tissue. These findings align with neuropathological knowledge and confirm that the model has learned biologically meaningful imaging patterns.
Fig. 9. Grad-CAM localization of pathologically relevant tumor features.
Representative axial slices overlay Grad-CAM heatmaps (red = highest activation) on T1Gd (left) and FLAIR (right) sequences. Activation concentrates on contrast-enhancing tumor core (T1Gd) and peritumoral FLAIR hyperintense regions (infiltrative margin). Negligible activation in normal brain tissue.
Discussion
This study introduced a multimodal, interpretable machine‑learning framework that integrates handcrafted radiomics features with deep‑learning derived patch embeddings to characterize MRI‑based signatures in glioma. A distinctive aspect of this work is the dual‑dataset design: Dataset 1, representing the original radiomics and DL‑patch feature matrices extracted from the TCGA‑LGG cohort, and Dataset 2, an enriched dataset generated through PCA‑GMM sampling and fusion‑guided pseudo‑labeling37,38. This structure enabled a controlled examination of how feature diversity, sample size, and data richness influence model performance and interpretability within a unified analytical pipeline.
Across experiments, the two datasets revealed complementary insights. Models trained on Dataset 1 reflected the inherent challenges of real‑world radiomics research, where modest sample sizes and high‑dimensional feature spaces can limit discriminative power. SHAP analysis showed that the radiomics SVM relied primarily on two dominant features, suggesting that the original feature space contained limited variability or redundant descriptors, an observation consistent with prior radiomics literature39,40,41. The DL‑patch and fusion models exhibited similar behavior, reinforcing the importance of data diversity when working with complex multimodal signatures.
The introduction of Dataset 2 provided a valuable contrast. By enriching the radiomics and DL‑patch spaces through PCA‑GMM sampling and leveraging fusion‑based pseudo‑labels, the dataset captured a broader range of structural and textural MRI patterns42. Models trained on this enriched dataset demonstrated markedly improved performance, with the fusion model achieving near‑perfect AUC, sensitivity, and specificity. These results highlight a central theme of the study: data‑centric enrichment can substantially enhance the expressive capacity of multimodal MRI signatures, often more effectively than increasing model complexity alone. In neuro‑oncology, where imaging datasets are frequently limited, such strategies offer a practical path toward more robust and generalizable predictive modeling.
Methodologically, the framework emphasizes interpretability, reproducibility, and multimodal integration. Radiomics features spanning 726 handcrafted descriptors were standardized using PCA and Gaussian mixture modeling (GMM), while DL‑patch embeddings contributed complementary local spatial information43,44. The fusion strategy, intentionally designed to be simple and transparent, proved highly effective when supported by a sufficiently informative dataset. Interpretability analyses, including SHAP, variance ranking, and correlation mapping, ensured that model behavior remained understandable and clinically grounded. Together, these components demonstrate how a data‑centric, interpretable pipeline can unlock the predictive potential of multimodal MRI features and contribute meaningfully to precision neuro‑oncology.
In conclusion, this work demonstrates that the integration of handcrafted and learned MRI signatures, combined with interpretability tools (SHAP and Grad-CAM), provides a transparent and clinically meaningful approach to predicting chemotherapy benefit in gliomas. The complementary nature of radiomics and deep learning features was validated on both real-world (Dataset 1, AUC = 0.75) and synthetically enriched (Dataset 2, near-perfect AUC) datasets, confirming that data quality and feature diversity are critical drivers of predictive performance. The proposed framework establishes a scalable and reproducible foundation for future multimodal imaging studies in neuro-oncology.
Methods
Dataset description
This study utilized the Pre‑operative TCGA‑LGG NIfTI and Segmentations dataset 1, an openly available resource derived from the TCGA Low‑Grade Glioma cohort45. The dataset includes 65 subjects with preoperative multimodal MRI scans and high‑quality tumor segmentation masks. Each patient folder contains T1, T1Gd, T2, and FLAIR sequences, along with automated and manually corrected GlistrBoost segmentation masks. All images are provided in NIfTI format and have undergone skull stripping and spatial co‑registration, ensuring consistency across subjects.
In addition to imaging data, the dataset includes a comprehensive radiomics feature matrix (TCGA_LGG_radiomicFeatures.csv) containing 726 handcrafted features per patient. These features encompass volumetric measurements (e.g., enhancing tumor volume, whole tumor volume), intensity‑based descriptors, texture metrics, and spatial growth model parameters. The dataset does not include chemotherapy or survival outcomes; therefore, it was used primarily for pipeline development, feature extraction, and methodological validation. Future integration with TCGA‑GBM and TCGA‑LGG clinical datasets will enable modeling of postoperative chemotherapy efficacy.
Dataset level characteristics relevant to downstream modeling, presented in Fig. 10. The distribution of patients across response categories, confirming class balance or imbalance. Also, compares tumor volumes between responders and non‑responders, providing insight into potential clinical associations. Additionally, illustrates the distribution of a selected radiomics shape feature, offering another perspective on morphological variability. Moreover shows the distribution of the L2 norms of the DL patch embeddings, reflecting the magnitude and spread of the learned feature representations.
Fig. 10. Dataset characteristics and variable distributions.
a Patient counts stratified by chemotherapy response, showing the class distribution between responders and non-responders. b Comparison of tumor volumes between responders and non-responders, providing insight into potential associations between tumor burden and chemotherapy benefit. c Distribution of a representative radiomics shape feature, offering another perspective on morphological variability. d Distribution of L2 norms of DL patch embeddings, reflecting the magnitude and spread of the learned feature representations.
Postoperative chemotherapy efficacy was operationalized as a binary endpoint among patients who received adjuvant chemotherapy. Overall survival (OS) was calculated from the date of surgery to the date of death or last follow‑up. Patients with months were categorized as chemotherapy responders, whereas those with months were categorized as non‑responders. This survival‑based definition has been widely adopted in TCGA‑based radiomics and DL studies as a pragmatic surrogate of treatment benefit when direct radiographic response assessments are unavailable. Radiomic and DL features were derived from TCGA‑LGG MRI data, and clinical survival information was obtained from GDC clinical files. Overall survival was calculated from death and follow‑up timestamps, enabling a binary classification of long‑term versus short‑term survivors (24‑month cutoff). These labels were integrated with imaging features to train and assess multiple predictive models.
OS as a surrogate for chemotherapy benefit is established in LGG literature Wang et al.35, Wu et al.36. Direct chemotherapy response labels are unavailable in public datasets; our approach represents the standard solution in this research domain.
Imaging data were obtained from the Pre-operative TCGA-LGG NIfTI and Segmentations dataset (Bakas et al.46). Clinical data, including chemotherapy administration records and overall survival timestamps, were retrieved from the Genomic Data Commons (GDC) clinical files associated with the TCGA-LGG cohort. Only patients who received postoperative adjuvant chemotherapy were included in this study.
Postoperative chemotherapy efficacy was operationalized as a binary endpoint based on OS. Patients with OS ≥ 24 months were categorized as chemotherapy responders, while those with OS < 24 months were categorized as non-responders. This survival-based definition has been adopted in prior TCGA-based radiomics studies as a pragmatic surrogate of treatment benefit when direct radiographic response assessments are unavailable.
Dataset 2 was created via PCA-GMM sampling and fusion-guided pseudo-labeling. PCA reduced the original feature space while preserving the covariance structure, followed by GMM to generate synthetic samples that maintain the original data distribution. This approach expands feature diversity and sample size without introducing unrealistic feature combinations.
This study utilized two complementary data sources from the TCGA-LGG cohort: Imaging data; pre-operative multimodal MRI scans (T1, T1Gd, T2, FLAIR) and tumor segmentation masks were obtained from the TCGA-LGG NIfTI and Segmentations dataset46. The dataset includes 65 subjects with complete imaging and segmentation. Clinical data; chemotherapy administration records and overall survival timestamps were retrieved from the Genomic Data Commons (GDC) clinical files associated with the TCGA-LGG cohort (Liu et al.25).
Dataset 2 was generated to overcome the limited feature variability of the original data (n = 65) and to explore the upper-bound potential of the fusion framework. The generation procedure involved three steps:
Step 1: The original radiomics feature matrix (65 × 482) was decomposed using principal component analysis, retaining components explaining 95% of variance (35 components).
Step 2: A Gaussian mixture model with 5 components was fitted to the PCA-reduced space. Synthetic samples were generated from the fitted GMM, preserving the original covariance structure while expanding the sample size to 500 samples.
Step 3: DL embeddings were used to guide label assignment, ensuring consistency with the original response distribution.
Validation procedure: Synthetic data quality was assessed by comparing (a) marginal distributions of original vs. synthetic features (Kolmogorov–Smirnov test, all p > 0.05), (b) covariance structure (Frobenius norm <0.1), and (c) classification performance consistency across 10 random seeds.
Image data processing
Preoperative multimodal MRI scans, including T1‑weighted, contrast‑enhanced T1‑weighted (T1Gd), T2‑weighted, and FLAIR sequences, were obtained in NIfTI format. All images were skull‑stripped and co‑registered to a common anatomical space as provided by the dataset. Tumor segmentation masks were available in two forms: automated GlistrBoost outputs and manually corrected expert annotations. For consistency, manually corrected masks were used as the reference standard. Each MRI volume and corresponding segmentation mask were loaded using NiBabel, and voxel‑wise integrity, dimensional consistency, and spatial alignment were verified.
Intensity normalization performed by
| 1 |
where original voxel intensity, μ mean intensity of brain voxels, σ standard deviation. Tumor volume calculation, and voxel volume v,
| 2 |
M(x) segmentation mask (1 = tumor, 0 = background),
| 3 |
Radiomics feature extraction and preprocessing
Radiomic features were obtained from the dataset’s accompanying feature matrix, which included 726 handcrafted descriptors capturing tumor intensity, shape, texture, volumetric properties, and spatial characteristics. Initial preprocessing involved removing features with excessive missingness, followed by imputing remaining missing values using median substitution. Features were standardized using z‑score normalization. To ensure robustness, radiomics features were evaluated for redundancy using correlation analysis and PCA, with all intermediate outputs saved as CSV files for reproducibility. Mean, variance, skewness, and kurtosis are derived by:
| 4 |
| 5 |
| 6 |
| 7 |
The shape features and compactness derived by using
| 8 |
| 9 |
Texture features such as contrast, homogeneity, energy, and entropy are given below,
| 10 |
| 11 |
| 12 |
| 13 |
Feature Preprocessing performed through z-score (), missing values imputation is as,
| 14 |
| 15 |
Dimensionality Reduction is achieved by using the covariance matrix, eigenvalue decomposition, and projection onto principal components, as given below, C, C * vk, and Z,
| 16 |
| 17 |
| 18 |
As shown in Table 2, per‑patient radiomics feature vectors were extracted from the tumor region after cropping to a standardized 3D patch. Each numeric column corresponds to a radiomic descriptor (e.g., intensity, spatial, or shape‑related features) used as input for machine‑learning and statistical modeling.
Table 2.
Radiomics feature matrix per patient, derived from standardized 96 × 96 × 96 tumor patches
| Patient ID | Feature 1 | Feature 2 | Feature 3 | Feature 4 | Feature 5 | Feature 6 | Feature 7 |
|---|---|---|---|---|---|---|---|
| TCGA-CS-4942 | 25.897 | −42.142 | 57.553 | 20.019 | −30.878 | 37.667 | 69.831 |
| TCGA-CS-4944 | 1.425 | −7.241 | 11.531 | 6.53 | −6.223 | 15.729 | 17.804 |
| TCGA-CS-5393 | 2.674 | −9.430 | 13.398 | 7.146 | −8.927 | 17.055 | 20.226 |
| TCGA-CS-5396 | 9.899 | −16.924 | 33.752 | 14.507 | −12.811 | 17.982 | 50.832 |
| TCGA-CS-5397 | 18.357 | −25.069 | 38.912 | 6.824 | −20.030 | 18.876 | 49.036 |
Deep learning feature extraction
A 3D convolutional neural network (CNN) encoder was designed to extract high‑level imaging representations from the multimodal MRI volumes. For each patient, the four MRI modalities were stacked into a four‑channel input tensor, and tumor‑centered cropping was performed using the segmentation mask to focus the model on the lesion region. The CNN architecture consisted of sequential 3D convolutional blocks with batch normalization and ReLU activation, followed by global average pooling to generate a compact latent feature vector. These DL features were exported as patient‑level CSV files for downstream fusion modeling. Convolution operations Y(i, j, k), and ReLU activation , batch normalization global average pooling zc is done as,
| 19 |
| 20 |
| 21 |
| 22 |
| 23 |
Figure 11 provides an overview of the dataset characteristics and the underlying structure of the extracted features. Illustrates the distribution of whole‑tumor volumes across the cohort, highlighting the substantial heterogeneity in tumor burden. Also, presents a PCA projection of the radiomics feature space, showing how the first two principal components capture the dominant variance patterns and whether any natural separation emerges between responders and non‑responders. Moreover, displays a correlation heatmap of the first 30 radiomics features, revealing clusters of highly correlated descriptors and indicating redundancy within the feature set. Additionally, shows a t‑SNE embedding of the DL patch features, providing an initial view of how the learned representations cluster patients in a lower‑dimensional latent space.
Fig. 11. Underlying data features and radiomics analysis.
a PCA projection of the radiomics feature space (PC1 vs. PC2), showing how the first two principal components capture dominant variance patterns and whether any natural separation emerges between responders and non-responders. b Variance explained by the first ten principal components, indicating the intrinsic dimensionality of the radiomics feature space. c Distribution of a representative radiomics feature (intensity mean of enhancing tumor on T1Gd), illustrating inter-patient variability. d Correlation heatmap of the top 30 radiomics features, revealing clusters of highly correlated descriptors and indicating redundancy within the feature set.
Fusion model and predictive framework
To evaluate the predictive potential of different feature families, three model types were constructed: (1) radiomics‑only models, (2) DL–only models, and (3) hybrid fusion models combining both feature sets. Gradient boosting machines and fully connected neural networks were used as classifiers due to their strong performance with structured and high‑dimensional data. Model training employed stratified k‑fold cross‑validation to ensure generalizability and to allow direct comparison across methods. Performance metrics included accuracy, AUC, sensitivity, specificity, and calibration error. Feature concatenation , fully connected layer h, output probability and loss function, binary cross-entropy L,
| 24 |
| 25 |
| 26 |
| 27 |
In Fig. 12, the structure of the DL feature embeddings derived from different input strategies is depicted. It shows a t‑SNE projection of the DL patch features, while the corresponding t‑SNE projection is for the DL bounding‑box features. t‑SNE embedding of the fused DL features, combining patch‑ and bounding‑box‑based representations. A UMAP projection of the fusion features, offering an alternative nonlinear view that often reveals clearer cluster boundaries.
Fig. 12. Deep learning feature embeddings and dimensionality reduction visualization.
a t-SNE projection of DL patch features, showing the clustering structure of local texture features in low-dimensional latent space. b t-SNE projection of DL bounding-box features, showing the distribution pattern of global tumor morphology features. c t-SNE projection of fused DL embeddings combining patch and bounding-box features, demonstrating the complementarity of different feature types. d UMAP projection of fused features, offering an alternative nonlinear view that often reveals clearer cluster boundaries.
The 3D bounding‑box coordinates used to crop each patient’s tumor region from the full CT/MRI volume are shown in Table 3. In radiomics and deep‑learning pipelines, you rarely feed the entire scan into the model; instead, you extract a standardized 3D patch around the lesion. Each row corresponds to a patient, and the columns define:
crop_shape to the final standardized 3D patch size used for model input (here: 96 × 96 × 96 voxels), a complete detailed structure is presented in Table 4.
Table 3.
Coordinates for standardized 96 × 96 × 96 brain tumor patches extracted from 3D MRI volumes
| ID | Modality | Patch1_center | Patch2_boundary1 | Patch3_boundary2 | Patch shape |
|---|---|---|---|---|---|
| TCGA-CS-4942 | flair | [87 119 76] | [97 105 64] | [94 131 70] | (64, 64, 64) |
| TCGA-CS-4942 | t1 | [87 119 76] | [97 105 64] | [94 131 70] | (64, 64, 64) |
| TCGA-CS-4942 | t1gd | [87 119 76] | [97 105 64] | [94 131 70] | (64, 64, 64) |
| TCGA-CS-4942 | t2 | [87 119 76] | [97 105 64] | [94 131 70] | (64, 64, 64) |
| TCGA-CS-4944 | flair | [86 134 66] | [76 143 84] | [73 129 66] | (64, 64, 64) |
Table 4.
showing the 3D CNN encoder designed architecture
| 3D CNN encoder designed architecture |
|---|
| Layer 1: Conv3D (3 × 3 × 3 kernel, 16 filters, stride 1, padding same) → BatchNorm3D → ReLU → MaxPool3D (2 × 2 × 2, stride 2) |
| Layer 2: Conv3D (3 × 3 × 3 kernel, 32 filters, stride 1) → BatchNorm3D → ReLU → MaxPool3D (2 × 2 × 2, stride 2) |
| Layer 3: Conv3D (3 × 3 × 3 kernel, 64 filters, stride 1) → BatchNorm3D → ReLU → GlobalAvgPool3D |
| Fully Connected: 64 → 128 |
| Training details: Optimizer = Adam (learning rate = 1e-4, β1 = 0.9, β2 = 0.999), batch size = 8, epochs = 100, early stopping patience = 20 epochs. Loss function = binary cross-entropy. |
Explainable AI
Explainability was incorporated at two levels. For radiomics‑based models, SHAP (SHapley Additive exPlanations) was used to quantify the contribution of individual features to model predictions. For DL models, Grad‑CAM was applied to generate heatmaps highlighting the MRI regions most influential in the prediction. These interpretability outputs were saved as high‑resolution images (300 dpi) for inclusion in the manuscript and presentation. SHAP value defined by , Grad Cam heatmap , where evaluation metrics are used.
| 28 |
| 29 |
| 30 |
| 31 |
| 32 |
| 33 |
| 34 |
Clinical decision support system prototype
Finally, an interpretable clinical decision support system (CDSS) interface was conceptualized. The CDSS integrates model predictions, radiomics‑based explanations, and Grad‑CAM visualizations to provide clinicians with transparent, patient‑specific insights. Although not deployed in a clinical environment, the prototype demonstrates the translational potential of the proposed framework.
Acknowledgements
This study was supported by the Development Fund Project of Xuzhou Medical University Affiliated Hospital (XYFY202460), the Research Project of Jiangsu Provincial Health Commission (Z2024021), the Clinical Technology Key Personnel Advanced Training Program of Xuzhou (2025GG020) and the Open Project of Jiangsu Provincial Key Laboratory (Xuzhou Medical University) (XZSYSKF2025008).
Author contributions
A.C. and P.D. oversaw the study's conception and design. Y.W., L.C., and Y.G. analyzed and interpreted of the gathered data. S.Z. and J.B. authored the manuscript outlining their research findings. All authors significantly contributed to the paper and collectively approved the final version submitted for publication.
Data availability
The raw multimodal MRI scans and tumor segmentation masks supporting this study are openly available in The Cancer Imaging Archive (TCIA) under the “Pre-operative TCGA-LGG NIfTI and Segmentations” collection45. The clinical and survival data were retrieved from the Genomic Data Commons (GDC) portal. The processed radiomics feature matrices and deep learning embedding vectors generated during the current study are not publicly available due to their derived nature and ongoing research use; however, they can be obtained from the corresponding author upon reasonable request and with appropriate data usage agreements.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Shenao Zhang, Jinfeng Bao.
Contributor Information
Aihong Cao, Email: caoaihong0516@126.com.
Peng Du, Email: dupeng0516@126.com.
References
- 1.Byeon, Y. et al. Interpretable multimodal transformer for prediction of molecular subtypes and grades in adult-type diffuse gliomas. Npj Digit. Med.8, 140 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Ostrom, Q. T. et al. CBTRUS statistical report: primary brain and other central nervous system tumors diagnosed in the United States in 2015-2019. Neuro Oncol.24, v1–v95 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Louis, D. N. et al. The 2021 WHO classification of tumors of the central nervous system: a summary. Neuro Oncol.23, 1231–1251 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Brennan, C. W. et al. The somatic genomic landscape of glioblastoma. Cell155, 462–477 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Krigers, A. et al. Age is associated with unfavorable neuropathological and radiological features and poor outcome in patients with WHO grade 2 and 3 gliomas. Sci. Rep.11, 17380 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wu, Y. L. et al. Postoperative chemotherapy use and outcomes from ADAURA: osimertinib as adjuvant therapy for resected EGFR-mutated NSCLC. J. Thorac. Oncol.17, 423–433 (2022). [DOI] [PubMed] [Google Scholar]
- 7.Li, R. et al. The effect of postoperative chemotherapy on survival outcomes and a nomogram for predicting overall survival in chondroblastic osteosarcoma. Sci. Rep.15, 45055 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Huang, C.-Y. et al. Lymph-vascular invasion affects the selection of time interval between surgery and postoperative chemotherapy in colon cancer. Curr. Probl. Surg. 77, 101995 (2026). [DOI] [PubMed]
- 9.Peng, Y. et al. Postoperative adjuvant chemotherapy and chemoimmunotherapy after radical resection for biliary tract cancer: a retrospective study. Oncologist30, oyaf163 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.He, Y. et al. The effect of postoperative adjuvant chemotherapy on survival outcomes in patients with early stage oral squamous cell carcinoma. Sci. Rep.15, 27157 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Stupp, R. et al. Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. N. Engl. J. Med.352, 987–996 (2005). [DOI] [PubMed] [Google Scholar]
- 12.Weller, M. et al. EANO guidelines on the diagnosis and treatment of diffuse gliomas of adulthood. Nat. Rev. Clin. Oncol.18, 170–186 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Grover, V. P. et al. Magnetic resonance imaging: principles and techniques: lessons for clinicians. J. Clin. Exp. Hepatol.5, 246–255 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Yu, Y. et al. MRI Radiomics based on paraspinal muscle for prediction postoperative outcomes in lumbar degenerative spondylolisthesis. Eur. Spine J.34, 3820–3831 (2025). [DOI] [PubMed] [Google Scholar]
- 15.Zhang, L. et al. Quantitative evaluation of postoperative status after meniscal repair using synthetic magnetic resonance imaging. Eur. J. Med. Res.30, 521 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Muaddi, H. et al. Imaging-based surgical site infection detection using artificial intelligence. Ann. Surg.282, 419–428 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lambin, P. et al. Radiomics: the bridge between medical imaging and personalized medicine. Nat. Rev. Clin. Oncol.14, 749–762 (2017). [DOI] [PubMed] [Google Scholar]
- 18.Esteva, A. et al. A guide to deep learning in healthcare. Nat. Med.25, 24–29 (2019). [DOI] [PubMed] [Google Scholar]
- 19.Ballard, J. L. et al. Deep learning-based approaches for multi-omics data integration and analysis. BioData Min.17, 38 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Baptista, D., Ferreira, P. G. & Rocha, M. Deep learning for drug response prediction in cancer. Brief. Bioinform.22, 360–379 (2021). [DOI] [PubMed] [Google Scholar]
- 21.Malta, T. M. et al. Machine learning identifies stemness features associated with oncogenic dedifferentiation. Cell173, 338–354.e315 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Partin, A. et al. Deep learning methods for drug response prediction in cancer: Predominant and emerging trends. Front. Med.10, 1086097 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Seal, R. L. et al. Genenames.org: the HGNC resources in 2023. Nucleic Acids Res.51, D1003–d1009 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Lorenzo, G. et al. Patient-specific, mechanistic models of tumor growth incorporating artificial intelligence and big data. Annu. Rev. Biomed. Eng.26, 529–560 (2024). [DOI] [PubMed] [Google Scholar]
- 25.Liu, J. et al. An integrated TCGA pan-cancer clinical data resource to drive high-quality survival outcome analytics. Cell173, 400–416.e411 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Sun, B. & Saenko, K. Deep CORAL: Correlation Alignment for Deep Domain Adaptation, 443–450 (Springer International Publishing, 2016).
- 27.Zuo, D. et al. Machine learning-based models for the prediction of breast cancer recurrence risk. BMC Med. Inf. Decis. Mak.23, 276 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Barredo Arrieta, A. et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion58, 82–115 (2020). [Google Scholar]
- 29.Abas Mohamed, Y. et al. Decoding the black box: explainable AI (XAI) for cancer diagnosis, prognosis, and treatment planning-A state-of-the art systematic review. Int. J. Med. Inform.193, 105689 (2025). [DOI] [PubMed] [Google Scholar]
- 30.Morrison-Jones, V. & West, M. Post-operative care of the cancer patient: emphasis on functional recovery, rapid rescue, and survivorship. Curr. Oncol.30, 8575–8585 (2023). [DOI] [PMC free article] [PubMed]
- 31.Dijkstra, E. A. et al. The value of post-operative chemotherapy after chemoradiotherapy in patients with high-risk locally advanced rectal cancer-results from the RAPIDO trial. ESMO Open8, 101158 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.O’Leary, K. Mounting evidence for perioperative immunotherapy. Nat. Med.29, 2973–2973 (2023).38093111 [Google Scholar]
- 33.Ahmed, F. et al. Explainable artificial intelligence (XAI) in medical imaging: a systematic review of techniques, applications, and challenges. BMC Med. Imaging26, 37 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Muhammad, D. & Bendechache, M. Unveiling the black box: a systematic review of explainable artificial intelligence in medical image analysis. Comput. Struct. Biotechnol. J.24, 542–560 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Wang, J. et al. An MRI-based radiomics signature as a pretreatment noninvasive predictor of overall survival and chemotherapeutic benefits in lower-grade gliomas. Eur. Radio.31, 1785–1794 (2021). [DOI] [PubMed] [Google Scholar]
- 36.Wu, G. et al. Study of radiochemotherapy decision-making for young high-risk low-grade glioma patients using a macroscopic and microscopic combined radiomics model. Eur. Radio.34, 2861–2872 (2024). [DOI] [PubMed] [Google Scholar]
- 37.Wang, Z. et al. UCGM: enhancing pseudo labels via uncertainty and cross-image Gaussian Mixture Model for semi-supervised semantic segmentation. Expert Syst. Appl.296, 129120 (2026). [Google Scholar]
- 38.Wang, T. et al. Pseudo label-guided data fusion and output consistency for semi-supervised medical image segmentation. Biomed. Signal Process. Control108, 107956 (2025). [Google Scholar]
- 39.Wang, Y. et al. The radiomic-clinical model using the SHAP method for assessing the treatment response of whole-brain radiotherapy: a multicentric study. Eur. Radiol.32, 8737–8747 (2022). [DOI] [PubMed] [Google Scholar]
- 40.Shi, Y. et al. Radiomics-based machine learning for glioma grade classification: a multicenter study with SHapley Additive exPlanations interpretability analysis. Transl. Cancer Res.14, 7329–7346 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Mastropietro, A., Feldmann, C. & Bajorath, J. Calculation of exact Shapley values for explaining support vector machine models using the radial basis function kernel[J]. Sci. Rep.13, 19561 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Albekairi, M. et al. Multimodal medical image fusion combining saliency perception and generative adversarial network. Sci. Rep.15, 10609 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Chaleshtori, A. E. & Aghaie, A. A novel bearing fault diagnosis approach using the Gaussian mixture model and the weighted principal component analysis. Reliab. Eng. Syst. Saf.242, 109720 (2024). [Google Scholar]
- 44.Alqahtani, N. A. & Kalantan, Z. I. Gaussian mixture models based on principal components and applications. Math. Probl. Eng.2020, 1202307 (2020). [Google Scholar]
- 45.Bakas, S. A. H. et al. Segmentation labels and radiomic features for the pre-operative scans of the TCGA-LGG collection. The Cancer Imaging Archive 10.7937/K9/TCIA.2017.GJQ7R0EF (2017).
- 46.Bakas, S. et al. Segmentation labels and radiomic features for the pre-operative scans of the TCGA-GBM collection. The Cancer Imaging Archive. 10.7937/K9/TCIA.2017.KLXWJJ1Q (2017).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The raw multimodal MRI scans and tumor segmentation masks supporting this study are openly available in The Cancer Imaging Archive (TCIA) under the “Pre-operative TCGA-LGG NIfTI and Segmentations” collection45. The clinical and survival data were retrieved from the Genomic Data Commons (GDC) portal. The processed radiomics feature matrices and deep learning embedding vectors generated during the current study are not publicly available due to their derived nature and ongoing research use; however, they can be obtained from the corresponding author upon reasonable request and with appropriate data usage agreements.












