Skip to main content
Biology Direct logoLink to Biology Direct
. 2026 Apr 24;21:93. doi: 10.1186/s13062-026-00789-1

A hybrid ST-ViT–driven multimodal architecture combining spatiotemporal MRI patterns and radiomic features for enhanced prediction of pCR in neoadjuvant breast cancer therapy

Satyabrata Pattanayak 1, Tripty Singh 1,✉, Rishabh Kumar 2, Ganesh R Naik 3
PMCID: PMC13248296  PMID: 42032765

Abstract

Breast cancer response prediction plays a critical role in treatment planning, especially for identifying patients likely to achieve pathologic complete response (pCR). Traditional approaches rely primarily on baseline clinical variables or static imaging features, limiting their ability to capture the complex biological and temporal dynamics of tumor evolution during therapy. This study presents a spatiotemporal multimodal framework that integrates quantitative clinicopathological variables, radiomics, and longitudinal DCE-MRI to enhance the prediction of pCR. We employ a Spatiotemporal Vision Transformer (ST-ViT) to model tumor evolution across four imaging time points and fuse it with quantitative radiomic and clinical features. The proposed framework captures both spatial heterogeneity and treatment-induced temporal changes, offering a comprehensive representation of tumor biology. Texture-based radiomic analysis reveals meaningful differences between pCR and non-pCR tumors, while enhancement-curve dynamics further highlight early perfusion and washout patterns linked to treatment sensitivity. The integrated spatiotemporal multimodal model demonstrates strong discriminatory power, yielding an AUC of 0.98 during training and 0.96 on the held-out test set. These results highlight the model’s ability to leverage dynamic MRI signatures, radiomic texture descriptors, and clinical features distinguish responders from non-responders effectively. By capturing both spatial and temporal tumor evolution, the framework offers a robust and clinically meaningful tool for early identification of treatment-sensitive phenotypes and supports precision-driven neoadjuvant therapy planning.

Keywords: Breast cancer prediction, pCR, Spatiotemporal vision transformer (ST-ViT), Radiomics, DCE-MRI, Multimodal learning, Longitudinal imaging, Clinical data fusion

Introduction

Breast cancer remains a leading global health concern, with early diagnosis and effective therapeutic decision-making crucial for improving patient outcomes. Neoadjuvant therapy (NAT) plays a vital role in shrinking tumors prior to surgery, thereby increasing the likelihood of complete resection. Achieving a pCR is strongly associated with improved long-term survival and reduced recurrence rates; however, accurately predicting pCR remains a major challenge due to the complex interplay among clinical, biological, and imaging-derived factors.

Prior studies have proposed latent-trajectory learning frameworks that encode the temporal evolution of DCE-MRI into compact representations for pCR prediction [1–3]. These models effectively capture appearance, temporal continuity, and cohort heterogeneity, demonstrating progressively higher balanced accuracy as additional imaging time points are incorporated into the training process.

Multimodal deep learning approaches using that fuse MRI and clinical variables using attention-based mechanisms have shown superior more generalizable performance compared to single-modality methods. Their ability to adaptively weight clinically relevant features enables stronger cross-institutional robustness and improved pCR prediction accuracy [4–7].

Complementary CNN-based frameworks leveraging advanced preprocessing, ROI segmentation, and post-contrast MRI texture extraction have also demonstrated strong effectiveness in predicting neoadjuvant chemotherapy response. These methods highlight the importance of 3D tumor characterization, enhancement kinetics, and spatial context in modeling therapeutic outcomes [8–11].

Additional work integrating pretreatment multiparametric MRI with clinicopathological factors has shown promising performance, particularly for aggressive subtypes such as triple-negative breast cancer. These studies emphasize that tumor volume normalization, biomarker heterogeneity, and robust feature engineering significantly influence model accuracy, while variations in 3D architectures or ROI selection produce marginal performance differences [12–14].

In this work, we build upon these advancements by introducing a spatiotemporal multimodal framework grounded in a Spatiotemporal Vision Transformer (ST-ViT). Unlike conventional approaches that predominantly rely on static radiomic signatures or baseline clinical predictors, our method explicitly captures the dynamic evolution of tumor characteristics across the four DCE-MRI time points of the I-SPY2 trial (T0–T3). The ST-ViT architecture jointly learns spatial heterogeneity and longitudinal temporal transitions, preserving early enhancement kinetics, volumetric changes, and structural transformations that are critical for distinguishing treatment-sensitive from treatment-resistant phenotypes.

By integrating longitudinal MRI representations with quantitative clinicopathological variables, the proposed framework constructs a unified, discriminative representation of tumor biology and therapy-induced response trajectories. This multimodal fusion approach enhances predictive robustness, improves separation between responders and non-responders, and supports more accurate and individualized neoadjuvant therapy decision-making. Our study demonstrates that incorporating spatiotemporal MRI evolution through transformer-based modeling offers a powerful and clinically meaningful pathway toward personalized breast cancer treatment planning.

Problem statement

This research aims to accurately predict pathologic complete response (pCR) in breast cancer patients undergoing neoadjuvant therapy. Predicting pCR is crucial in oncology, as it is associated with improved long-term outcomes and plays a key role in guiding personalized treatment strategies. Achieving precise predictions remains challenging due to the multifactorial nature of breast cancer prognosis, which depends on variables such as tumor size, sphericity, menopausal status, race, molecular subtype, and treatment regimen.

To address this complexity, the study leverages two major sources of information: (1) quantitative clinicopathological data and treatment details, and (2) dynamic breast MRI acquired at multiple treatment stages. The I-SPY2 trial provides longitudinal DCE-MRI scans collected at four key time points—T0 (pre-treatment), T1 (early treatment), T2 (mid-treatment), and T3 (pre-surgery)—offering a spatiotemporal view of tumor evolution. These sequential imaging biomarkers capture early enhancement patterns, structural changes, and therapeutic response trajectories that are not observable in static or single-time-point imaging.

The central challenge lies in integrating this heterogeneous longitudinal MRI information with clinical variables into a unified predictive framework capable of modeling both spatial tumor heterogeneity and temporal response dynamics. By effectively combining quantitative features with multi-time-point imaging data, this research seeks to improve the accuracy of pCR prediction and support more informed, individualized clinical decision-making.

Materials and methods

Dataset description

This study utilized patient data and imaging features derived from the I-SPY2 TRIAL [15], a multi-center adaptive clinical trial conducted between 2010 and 2016 across 22 clinical institutions in the United States. The trial enrolled women diagnosed with locally advanced breast cancer to evaluate neoadjuvant therapies and their impact on achieving pathologic complete response (pCR). Each patient underwent comprehensive clinical assessment and Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) at four treatment stages:

  • T0: Pre-treatment baseline MRI,

  • T1: Early-treatment MRI (after 3–4 weeks of therapy),

  • T2: Mid-treatment MRI (after 12 weeks),

  • T3: Post-treatment MRI (before surgery).

Patient population

The initial dataset comprised approximately 990 patients. Several cases were excluded due to missing key clinical or imaging information, incomplete MRI sequences, or the absence of surgical follow-up data required to determine pCR outcomes. After applying quality control and inclusion criteria, a final subset of 384 patients (Fig. 1) was selected for model development. Each patient record included both MRI-derived features and corresponding clinicopathological variables.To ensure consistency in temporal modeling, a complete-case strategy was adopted. Patients with missing MRI timepoints (T0–T3) or incomplete clinical variables were excluded during preprocessing, and only samples with complete imaging and clinical data were included in the final analysis.

Fig. 1.

Fig. 1

The figure illustrates the filtering process that reduced the patient data from 990 to 384, primarily due to missing tumor characteristic information

Clinical and biological variables

The clinical dataset incorporated a diverse set of attributes, including:

  • Demographic Information: Age at screening, race, and ethnicity.

  • Biological Markers: Hormone receptor (HR) status, HER2 status, Mammaprint (MP) classification, and HR/HER2 subtype.

  • Clinical Characteristics: SBR grade, menopausal status, and treatment arm.

  • Tumor Morphology: Lesion type (Ltype) such as single mass, multiple masses, or non-mass enhancement.

Baseline clinical variables (e.g., HER2 status, Mammaprint classification, age at screening, and MRLD) were included as patient-level attributes. Longitudinal information was captured through MRI-derived features extracted at multiple treatment stages (T0–T3), including sphericity, longest diameter (LD), and background parenchymal enhancement (BPE).

Quantitative MRI features

For each MRI timepoint (T0–T3), quantitative imaging features were extracted from tumor regions of interest (ROI), including:

  • Tumor volume (Vt),

  • Longest diameter (LD),

  • Sphericity (St),

  • Background parenchymal enhancement (BPE),

  • Functional Tumor Volume (FTV) across contrast thresholds (e.g., V10, V20, V30, V40).

These parameters were computed using semi-automated segmentation methods from the DCE-MRI scans and served as temporal indicators of tumor evolution across treatment.

Outcome variable

The primary endpoint was pathologic complete response (pCR), a binary variable representing the absence (pCR = 1) or presence (pCR = 0) of residual invasive cancer in both breast and lymph nodes at the time of surgery. This outcome was used as the ground-truth label for supervised classification.

Data integration and representation

Each patient record combined multi-temporal MRI-derived features with static clinical variables, producing a multimodal dataset structured as follows:

graphic file with name d33e362.gif 1

where fTt denotes the quantitative imaging vector at time Tt, cj represents the jth clinical covariate, and yi corresponds to the pCR outcome for patient i.

Data splitting

The final dataset of 384 patients was divided into training (80%) and testing (20%) subsets using stratified sampling to preserve the pCR class distribution. Five-fold cross-validation was employed during model training to ensure robustness and reduce overfitting.

Image acquisition and preprocessing

Each enrolled patient underwent MRI at four distinct treatment stages: pre-treatment (T0), early treatment (T1), mid-treatment (T2), and post-treatment (T3). The MRI protocol followed a standardized multi-center acquisition setup, including T1-weighted fat-suppressed sequences before and after contrast administration. Images were captured using 1.5T or 3T scanners, ensuring adequate spatial and temporal resolution for volumetric tumor assessment. Instead of using the entire MR image, the segmented tumor region of interest (ROI) was extracted and used as input to the ST-ViT model. This ensures that the network focuses on tumor-related regions while reducing background variability.

Preprocessing pipeline

To ensure data uniformity and improve model performance, several preprocessing steps were applied:

  1. Anatomical Alignment: [16, 17] All DCE-MRI volumes were registered to the T0 reference frame using rigid or affine transformation to minimize motion and positional variations across different time points.

  2. Bias Field Correction: [18, 19] N4ITK bias field correction was performed to correct low-frequency intensity inhomogeneity, improving tissue contrast consistency across slices.

  3. Intensity Normalization: Each MRI volume was z-score normalized to reduce inter-patient variability:
    graphic file with name d33e444.gif 2
    where I is the voxel intensity, µ is the mean intensity, and σ is the standard deviation within the breast mask.
  4. Tumor Segmentation: Lesion regions were extracted either from expert annotations or semi-automated segmentation using region-growing and intensity thresholding. The segmented region of interest (ROI) was used to compute tumor-specific features such as volume, longest diameter, and sphericity.Tumor segmentation was performed using a semi-automatic region-growing approach on contrast-enhanced DCE-MRI images. A seed point was initialized within the tumor region, and neighboring voxels with similar enhancement intensity were iteratively included based on an adaptive threshold. The resulting segmentation mask was visually inspected and when necessary, refined by an experienced radiologist to ensure accurate tumor boundary delineation.

  5. Resampling and Cropping: Each ROI was resampled to a fixed voxel spacing of Inline graphic mm and cropped to a uniform dimension around the lesion center to maintain consistent input size for the transformer model.

  6. Temporal Stack Formation: [20] For patients with multi-phase data, images from T0, T1, T2, and T3 were temporally concatenated, creating a 4-channel spatial-temporal tensor input for the Spatio-Temporal Vision Transformer (ST-ViT). This design preserved information on tumor evolution across treatment time points.

Feature extraction and enhancement

In addition to raw imaging, derived radiomic features such as tumor volume, sphericity, longest diameter (LD), and background parenchymal enhancement (BPE) were computed for each timepoint. These features were later combined with quantitative clinical data to provide a multimodal feature representation.

To improve local feature discrimination, the image patches were enhanced with Contrast-Limited Adaptive Histogram Equalization (CLAHE) before being tokenized for input to the transformer architecture. This step was particularly beneficial in preserving lesion boundary information and enhancing subtle intensity gradients that contribute to treatment response differentiation.

Final dataset preparation

Patients with incomplete MRI sequences or missing quantitative tumor measurements were excluded from the study. After preprocessing, a total of 384 patient samples were retained for subsequent model training and validation.

Quantitative feature extraction

Quantitative imaging biomarkers were derived from Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) for each of the four temporal stages (T0–T3). Tumor regions of interest (ROIs) were delineated semi-automatically using intensity-based segmentation techniques, ensuring consistency across all timepoints. The extracted features captured morphological, functional, and enhancement-based characteristics of the tumor and surrounding parenchyma.

Feature categories

The quantitative features extracted from each MRI scan were grouped into the following categories:

  • Morphological Features: Longest Diameter (LD), Tumor Volume (Vtum), and Sphericity (S), representing the geometric shape and size of the lesion.

  • Functional Features: Background Parenchymal Enhancement (BPE) and Functional Tumor Volume (FTV) derived from voxel-wise enhancement kinetics.

  • Texture and Intensity Features: Metrics such as mean intensity enhancement and signal heterogeneity (computed for validation but not all included in the final model).

Mathematical representation

Let Inline graphic denote the voxel intensity at spatial location Inline graphic for timepoint Inline graphic. The quantitative features were computed as follows:

  1. Tumor Volume (Vtum):
    graphic file with name d33e558.gif 3
    where Inline graphic represents the number of voxels within the segmented tumor region and Vvox denotes the physical voxel volume in mm3.
  2. Longest Diameter (LD):
    graphic file with name d33e583.gif 4

    indicating the maximum Euclidean distance between any two boundary points of the tumor at time t.

  3. Sphericity (S):
    graphic file with name d33e599.gif 5
    where A(t) denotes the tumor surface area. This metric quantifies how closely the lesion shape approximates a perfect sphere.
  4. Functional Tumor Volume (FTV):
    graphic file with name d33e617.gif 6
    where Inline graphic is the indicator function, and Inline graphic represents enhancement thresholds (denoted as Inline graphic, Inline graphic, Inline graphic, Inline graphic).
  5. Background Parenchymal Enhancement (BPE):
    graphic file with name d33e654.gif 7
    where Rbkg is the background fibroglandular region and Inline graphic its voxel count.

Temporal feature aggregation

Each feature was computed independently for Inline graphic, Inline graphic, Inline graphic, and Inline graphic, forming a temporal evolution vector:

graphic file with name d33e689.gif 8

where fTt represents the feature vector at time Tt. These multi-temporal feature vectors capture dynamic tumor evolution across treatment phases, serving as sequential input for the spatio-temporal transformer (ST-ViT) architecture.

Normalization and standardization

To minimize inter-patient variability, all quantitative features were normalized using z-score normalization:

graphic file with name d33e708.gif 9

where µf and σf represent the mean and standard deviation of each feature across the population. Missing values were imputed using median imputation to preserve data integrity without introducing outlier bias.

Proposed architecture

The proposed framework integrates multi-temporal MRI images and quantitative biomarkers into a unified Spatio-Temporal Vision Transformer (ST-ViT) model for pathologic complete response (pCR) prediction. The network is designed to learn both spatial tumor morphology from MRI slices and temporal evolution of treatment response across multiple time points (T0–T3). The architecture consists of five major modules: (1) spatial encoder, (2) temporal encoder, (3) tabular encoder, (4) cross-modal fusion, and (5) multi-head progressive prediction.

Patch embedding (spatial encoder)

[21, 22] Each MRI image at time t for subject i, denoted as Inline graphic, is divided into non-overlapping patches of size P × P. Each flattened patch Inline graphic is linearly projected into a d-dimensional embedding:

graphic file with name d33e774.gif 10

Learnable positional encodings Inline graphic and a per-timepoint CLS token ct are added to form the initial token sequence:

graphic file with name d33e790.gif 11

This sequence is processed through Ls transformer layers, each consisting of multi-head self-attention and feed-forward blocks:

graphic file with name d33e802.gif 12

The spatial representation for timepoint t is extracted from the CLS token:

graphic file with name d33e811.gif 13

Temporal encoder (across timepoints)

The tabular encoder is implemented as a multilayer perceptron (MLP) that processes clinical and quantitative MRI features. The input consists of baseline clinical variables together with MRI-derived features extracted at different treatment stages (T0–T3). Features from all timepoints are concatenated into a single tabular feature vector, which is passed through fully connected layers with nonlinear activation to learn a compact representation used for prediction. The temporal evolution of the tumor is modeled by stacking all time-specific image embeddings into a sequence:

graphic file with name d33e819.gif 14

This sequence is passed through a temporal transformer encoder with Lt layers to capture inter-timepoint dependencies:

graphic file with name d33e831.gif 15

This produces temporally-aware image embeddings Inline graphic for each timepoint.

Quantitative feature encoder

Quantitative imaging biomarkers (e.g., longest diameter, tumor volume, sphericity, background parenchymal enhancement) for patient i are represented as Inline graphic and processed through a multilayer perceptron:

graphic file with name d33e852.gif 16

If features vary over time (e.g., Inline graphic change across treatment phases), the temporal sequence Inline graphic can be encoded using a temporal MLP or a lightweight LSTM to yield Inline graphic.

Cross-modal fusion via cross-attention

[23, 24] To combine imaging and quantitative embeddings, a cross-attention mechanism is applied for each timepoint t:

graphic file with name d33e883.gif 17

where

graphic file with name d33e888.gif 18

The fused representation is then normalized:

graphic file with name d33e894.gif 19

This module aligns spatial-temporal MRI features with clinical biomarkers for robust multimodal reasoning.

Multi-head progressive prediction

[25–27] Each timepoint representation Inline graphic is connected to an independent classification head Inline graphic to output progressive probabilities:

graphic file with name d33e918.gif 20

The final aggregated probability is derived from pooled temporal embeddings:

graphic file with name d33e924.gif 21

The complete prediction vector is thus:

graphic file with name d33e930.gif 22

Training objectives

The model is trained using multi-task supervision to learn both progressive and final treatment outcomes.

Final pCR loss (primary)

graphic file with name d33e943.gif 23

Progressive auxiliary loss

graphic file with name d33e955.gif 24

Temporal consistency loss (optional) To encourage clinically plausible monotonic probability evolution:

graphic file with name d33e965.gif 25

Total loss

graphic file with name d33e975.gif 26

where λprog and λtemp are balancing coefficients.

This architecture enables the model to capture spatio-temporal dynamics of tumor response while jointly utilizing quantitative biomarkers, providing clinically interpretable progressive probabilities across treatment stages.The coefficients λprog and λtemp control the contribution of the temporal consistency and progression constraints during optimization. These parameters were empirically tuned to balance the regularization effect of temporal constraints while maintaining stability in the primary classification objective.

Model training

The training strategy for the proposed Spatio-Temporal Vision Transformer (ST-ViT) framework is designed to jointly optimize spatial, temporal, and multimodal feature learning while ensuring progressive interpretability of treatment response across multiple MRI timepoints (T0–T3). The network parameters are updated end-to-end through supervised learning using the binary pathological complete response (pCR) label as ground truth.To evaluate the robustness of the proposed model, k-fold cross-validation was employed on the available dataset. This strategy allows the model to be trained and evaluated across multiple splits of the cohort, helping to reduce potential bias from a single train–test partition.

Training objective

The model employs a composite loss function that integrates three complementary terms: (1) the final prediction loss (Inline graphic), (2) the progressive supervision loss (Inline graphic), and (3) the temporal consistency loss (Inline graphic), ensuring temporal coherence and interpretability of pCR probabilities. The total loss is defined as:

graphic file with name d33e1032.gif 27

where λprog and λtemp are empirically tuned hyperparameters.

Final prediction loss

A binary cross-entropy (BCE) loss is used for the final pCR output:

graphic file with name d33e1050.gif 28

where Inline graphic represents the model’s final prediction for subject i, and Inline graphic is the corresponding ground truth.

Progressive prediction loss

Each intermediate prediction head contributes to the progressive learning of treatment response over time:

graphic file with name d33e1070.gif 29

encouraging early layers to model incremental treatment changes reflected in MRI scans.

Temporal consistency loss

To ensure clinical plausibility, a smooth temporal evolution constraint penalizes non-monotonic changes in predicted response:

graphic file with name d33e1080.gif 30

where γ is a small positive margin that ensures the probability of achieving pCR does not decrease unrealistically with treatment progression.

Optimization and hyperparameters

Model parameters Inline graphic are optimized via stochastic gradient descent with backpropagation. The Adam optimizer is employed with learning rate η and weight decay β to stabilize convergence:

graphic file with name d33e1102.gif 31

A linear warmup followed by cosine decay scheduling is used for learning rate adaptation:

graphic file with name d33e1108.gif 32

Key hyperparameters include:

  • Learning rate: Inline graphic

  • Batch size: 8 (per GPU)

  • Optimizer: AdamW with weight decay Inline graphic

  • Dropout rate: 0.1

  • Number of epochs: 100

  • Patch size: Inline graphic

  • Embedding dimension: d = 768

  • Number of attention heads: h = 12

Data augmentation and regularization

To improve generalization and robustness to inter-scanner variations, MRI images are augmented through:

  • Random rotations (Inline graphic)

  • Horizontal/vertical flips

  • Random intensity scaling (Inline graphic)

  • Gaussian noise injection

Quantitative features are standardized using z-score normalization:

graphic file with name d33e1181.gif 33

Dropout and layer normalization are applied within each transformer block to prevent overfitting and maintain stable gradients.

Training strategy and early stopping

A stratified 5-fold cross-validation is employed to evaluate model robustness across patient subsets. The training stops early if validation loss does not improve for 10 consecutive epochs. The best-performing checkpoint is selected based on the minimum validation loss:

graphic file with name d33e1191.gif 34

To ensure clinical reliability, model calibration is performed on the validation set using Platt scaling, yielding interpretable probabilities that align with actual pCR likelihoods.

Computational setup

All experiments were conducted on NVIDIA A100 GPUs (80 GB) using PyTorch 2.3. The average training time per fold was approximately 1.5 hours. The entire training pipeline was implemented with mixed-precision (FP16) optimization to reduce memory consumption and accelerate convergence.

Through this training strategy, the model learns clinically consistent spatio-temporal representations of tumor dynamics, yielding progressive and final predictions that can assist oncologists in evaluating early treatment response and personalizing therapy adjustments.

Evaluation metrics

To rigorously assess the predictive performance and clinical reliability of the proposed Spatio-Temporal Vision Transformer (ST-ViT) framework, we employed a combination of classification, calibration, and agreement-based metrics. These metrics evaluate both the discriminative capability of the model and its temporal consistency in estimating treatment response across sequential MRI scans. The final evaluation was conducted on an independent test set, unseen during model training and validation.

Classification performance metrics

The primary evaluation metrics for binary classification (pCR vs. non-pCR) include Accuracy, Precision, Recall, F1-Score, and the Area Under the Receiver Operating Characteristic Curve (AUC-ROC). Let Inline graphic denote the ground truth label for patient i and Inline graphic represent the predicted probability of achieving pCR. The classification decision Inline graphic is obtained using a threshold τ = 0.5 unless specified otherwise.

Accuracy

Measures the proportion of correctly classified samples:

graphic file with name d33e1233.gif 35

where TP, TN, FP, and FN denote true positive, true negative, false positive, and false negative counts, respectively.

Precision (positive predictive value)

Reflects the model’s reliability in predicting true responders:

graphic file with name d33e1254.gif 36
Recall (sensitivity or true positive rate)

Quantifies the model’s ability to identify all pCR cases:

graphic file with name d33e1262.gif 37
F1-score

Provides the harmonic mean between Precision and Recall, offering a balanced view of sensitivity and specificity:

graphic file with name d33e1270.gif 38
AUC-ROC

The AUC-ROC score evaluates the model’s overall discriminative capacity across varying decision thresholds:

graphic file with name d33e1278.gif 39

where TPR and FPR represent the true and false positive rates, respectively. A higher AUC indicates superior separability between pCR and non-pCR patients.

Calibration metrics

To assess the clinical interpretability and reliability of predicted probabilities, we evaluated calibration using the Brier Score and Expected Calibration Error (ECE).

Brier score

Measures the mean squared difference between predicted probabilities and actual outcomes:

graphic file with name d33e1297.gif 40

Lower scores indicate better-calibrated predictions.

Expected calibration error (ECE)

Quantifies the deviation between predicted confidence and empirical accuracy:

graphic file with name d33e1307.gif 41

where Bm represents the m-th probability bin, Inline graphic is the average accuracy within that bin, and Inline graphic denotes the mean predicted confidence. Lower ECE indicates better alignment of predicted probabilities with observed frequencies.

Temporal evaluation metrics

Since the ST-ViT model produces intermediate probabilities for each treatment stage (T0 to T3), we further assess temporal consistency and progressive correlation with true response trajectories.

Monotonicity index (MI)

Evaluates whether predicted probabilities evolve consistently with treatment progression:

graphic file with name d33e1345.gif 42

where Inline graphic is an indicator function. Higher MI reflects stronger temporal coherence and clinical interpretability.

Temporal smoothness (TS)

Quantifies fluctuations in predicted probabilities across consecutive timepoints:

graphic file with name d33e1358.gif 43

Higher values indicate smoother, more stable temporal predictions.

Statistical significance and robustness

To validate robustness, we report 95% confidence intervals for all performance metrics using bootstrapped resampling (n = 1000). Statistical comparison between models (e.g., CNN, ViT, ST-ViT) is performed using the DeLong test for AUC differences and McNemar’s test for paired classification outcomes. A p-value < 0.05 is considered statistically significant.

Clinical interpretability

In addition to conventional metrics, the model’s output probabilities were analyzed for calibrated risk stratification, where patients were grouped into low (<0.3), intermediate (0.3–0.7), and high (>0.7) predicted pCR probability categories. This stratification provides clinicians with interpretable insights for early treatment decision-making and adaptive therapy planning.

—

Collectively, these metrics provide a comprehensive assessment of model performance, emphasizing both predictive power and clinical usability, ensuring that the ST-ViT framework not only classifies accurately but also produces temporally consistent and clinically interpretable predictions across the treatment course.

Results

The proposed study investigated the relationship between pCR and key tumor features, including sphericity, tumor size (LD, FTV), and menopausal status, using both univariate and multivariate analyses. Statistical significance was evaluated using the Wilcoxon signed-rank test, and the observed differences were statistically significant (p < 0.05), indicating that the results are unlikely to have occurred by chance given the sample size of 384 patients.The model demonstrated strong predictive ability for pCR, achieving an AUC of 0.98 on the training set and 0.96 on the independent test set. The classification report showed high recall for non-pCR (0.95) and pCR (0.89), confirming the model’s reliability in identifying both responsive and non-responsive patients. ROC analysis further supported this, indicating high sensitivity and specificity. Compared with traditional models and feature-based approaches, our architecture demonstrated a clear performance advantage. On the other hand, despite incorporating radiomic features such as tumor sphericity and volume, their contribution to model prediction was minimal, suggesting these features alone may not capture the full biological complexity underlying pCR. This highlights the potential of multimodal integration for improved prediction accuracy.

Dataset description

A total of 384 patient records were included in the final dataset after applying inclusion and quality control criteria. Each record contained clinical, biological, and quantitative MRI-derived features collected at four imaging time points (T0–T3). Table 1 provides descriptive statistics for key variables, including the mean, standard deviation, and range.

Table 1.

Descriptive statistics of clinical and MRI-derived features in the final cohort (n = 384)

Variable Mean Std Min 25% 50% Max
HR 0.57 0.50 0.0 0.00 1.0 1.00
HER2 0.25 0.43 0.0 0.00 0.0 1.00
MP 0.49 0.50 0.0 0.00 0.0 1.00
Age_at_Screening 49.09 11.37 0.0 42.00 50.0 77.00
MRLD 1.24 2.35 0.0 0.00 0.0 15.00
VOLUME_TUM_BLU_V10 13.15 35.15 0.0 0.00 0.0 433.37
VOLUME_TUM_BLU_V20 7.49 28.21 0.0 0.00 0.0 471.31
VOLUME_TUM_BLU_V30 3.99 21.46 0.0 0.00 0.0 377.38
VOLUME_TUM_BLU_V40 1.55 9.10 0.0 0.00 0.0 158.89
SPHERICITY_T0 0.08 0.11 0.0 0.00 0.0 0.46
SPHERICITY_T1 0.08 0.11 0.0 0.00 0.0 0.53
SPHERICITY_T2 0.10 0.14 0.0 0.00 0.0 0.79
SPHERICITY_T3 0.11 0.16 0.0 0.00 0.0 0.76
LD_T0 1.81 2.55 0.0 0.00 0.0 13.20
LD_T1 1.50 2.22 0.0 0.00 0.0 13.60
LD_T2 1.02 1.75 0.0 0.00 0.0 8.30
LD_T3 0.71 1.42 0.0 0.00 0.0 7.70
BPE_5slice_mean_T0 11.70 17.18 0.0 0.00 0.0 84.66
BPE_5slice_mean_T1 9.50 13.60 0.0 0.00 0.0 73.26
BPE_5slice_mean_T2 8.42 11.86 0.0 0.00 0.0 64.17
BPE_5slice_mean_T3 7.94 11.53 0.0 0.00 0.0 81.13
pCR 0.32 0.47 0.0 0.00 0.0 1.00

The mean age of patients at screening was approximately 49 years, with pathologic complete response (pCR) observed in about 31.5% of cases, consistent with prior I-SPY2 cohort analyses. Hormone receptor (HR) positivity was more prevalent than HER2 amplification, reflecting a clinically heterogeneous breast cancer population.

Quantitative MRI features indicated clear treatment-related morphological changes. The mean tumor volume and longest diameter (LD) progressively declined from T0 to T3, confirming substantial shrinkage and therapy-induced regression. Early reductions in VTUM and LD between T0–T1 suggest that early imaging biomarkers may serve as predictors of ultimate treatment response.

Sphericity (St) showed a slight increase over time points, suggesting that tumors became more regular or compact during response. Similarly, background parenchymal enhancement (BPE) decreased steadily from T0 to T3, which may reflect both hormonal modulation and therapy-driven vascular normalization.

The distribution of volume and LD values was right-skewed, as evidenced by a large standard deviation relative to the mean, indicating the presence of a few large or aggressive tumors. The wide dynamic range in VTUM and BPE features also highlights strong inter-patient variability, which the predictive model will capture. Overall, the descriptive statistics confirm that the dataset is diverse, well-balanced, and representative of the imaging dynamics associated with pCR, both early and late.

Patient characteristic engineering

Our dataset consists of MRI images and 46 quantitative features, with pCR as the outcome variable. These features include both categorical and continuous variables. Some continuous variables, such as age, showed limited variation across patients. Given the relatively small sample size (384 patients), we applied binning to convert selected continuous variables into discrete groups. For example, age was divided into five bins and label-encoded (e.g., 1, 2, 3), improving both model interpretability and stability. Similar binning approaches were used for other variables, including tumor volume, sphericity, BPE, FTV, and LD.Binning was applied to selected continuous variables (e.g., age, tumor volume, and sphericity) to reduce sensitivity to noise and extreme values in the relatively small dataset. This preprocessing step helps stabilize feature distributions when integrating heterogeneous clinical and imaging features.

For sphericity, the initial five-bin strategy produced a substantial imbalance, with one bin (0–0.1) containing 238 of the 384 patients. To correct this, we further subdivided this high-density bin into smaller, more balanced intervals and merged sparsely populated bins to enhance model robustness. For non-ordinal categorical variables—such as treatment type, race, and menopausal status, and ethnicity—we applied one-hot encoding to preserve category independence. Ordinal variables such as SBR grade, were label-encoded to retain their intrinsic order.

Temporal evolution of MRI-derived quantitative features (T0–T3)

Figure 2 illustrates the longitudinal trends of key MRI-derived quantitative biomarkers, tumor volume, sphericity, longest diameter (LD), and background parenchymal enhancement (BPE) across four imaging timepoints (T0 to T3), stratified by pathological complete response (pCR) status.

Fig. 2.

Fig. 2

Temporal evolution of tumor volume, sphericity, diameter, and background parenchymal enhancement (BPE) across treatment stages (T0–T3)

Tumor volume trend

Patients achieving pCR exhibited a markedly steeper decline in mean tumor volume than non-responders. Between T0 and T1, a rapid volumetric reduction was observed for the pCR group, whereas the non-pCR group demonstrated a more gradual decrease. By T3, the residual tumor volume in pCR patients approached near-zero levels, indicating an effective therapeutic response. This supports the hypothesis that early volumetric reduction is a strong early imaging biomarker of treatment sensitivity.

Tumor sphericity trend

Sphericity showed an inverse temporal behavior. Responders began with slightly higher sphericity at baseline (T0), which decreased modestly at T1 and subsequently increased from T2 to T3. This pattern suggests morphological normalization or reorganization of the tumor region, as necrotic tissue replaces viable tumor cells. In contrast, non-responders exhibited less pronounced or inconsistent changes, implying persistent irregular morphology across treatment.

Longest diameter (LD) trend

The LD trajectory followed a similar downward trend to that of tumor volume. Responders showed a sharper reduction from T0 to T2, with mean LD values decreasing by more than half, consistent with volumetric shrinkage. Non-responders, however, retained larger diameters throughout, underscoring incomplete tumor regression. The strong correlation between LD and volume across timepoints reinforces the geometric consistency of the response dynamics.

Background parenchymal enhancement (BPE) trend

Both pCR and non-pCR groups exhibited declining BPE over the course of treatment, reflecting an overall reduction in parenchymal vascularity. However, pCR patients demonstrated a consistently lower mean BPE at all timepoints, with a more pronounced decline between T0 and T2. This observation may indicate that lower baseline BPE and dynamically decreasing BPE are associated with more favorable treatment outcomes, potentially linked to microenvironmental changes that affect contrast uptake.

Overall interpretation

Collectively, these temporal trends reveal that dynamic MRI-derived parameters capture distinct physiological trajectories between responders and non-responders. The early and pronounced changes in tumor volume and LD, combined with later-stage increases in sphericity and reductions in BPE, suggest that both morphological and vascular characteristics evolve systematically during neoadjuvant therapy. These findings provide quantitative support for incorporating temporal imaging features into the proposed ST-ViT framework, enabling the model to leverage longitudinal cues to improve pCR prediction.

DCE-MRI enhancement curve analysis

Figure 3 shows the mean contrast enhancement trajectories for patients who achieved pathological complete response (Responders, pCR = 1) and those who did not (Non-Responders, pCR = 0) across the dynamic time points of the DCE-MRI acquisition. The enhancement values represent the normalized signal intensity change (Inline graphic) computed for each dynamic frame.

Fig. 3.

Fig. 3

Mean contrast-enhancement curve illustrating the temporal evolution of normalized signal intensity across DCE-MRI phases

Early enhancement behavior (T0–T1)

Non-Responders exhibited a markedly higher early enhancement peak at the first post-contrast phase (T1), with a rapid rise in Inline graphic than Responders. This pattern reflects a more aggressive vascular phenotype characterized by increased microvascular permeability and perfusion, typically associated with poor therapeutic response. In contrast, Responders demonstrated a slower more gradual rise in signal intensity, suggesting lower initial contrast uptake and a less angiogenically active tumor microenvironment.

Mid-treatment dynamics (T2)

At the intermediate time point (T2), Responders showed a reduction in enhancement, indicating an early treatment effect associated with cytotoxic-induced changes in vascular structure and perfusion. Non-Responders, on the whole, maintained relatively elevated enhancement levels, suggesting ongoing tumor vascular activity and limited treatment-induced regression. This divergence between the two groups at T2 highlights the potential of temporal enhancement features as early imaging biomarkers for predicting therapeutic response.

Washout and late-phase enhancement (T3)

By the later phase (T3), Responders exhibited a characteristic washout pattern, with a clear decline in Inline graphic relative to earlier time points. This rapid washout is consistent with an effective treatment response, resulting in reduced vascularity and decreased extravascular extracellular space (EES) retention of the contrast agent. Non-Responders, conversely, demonstrated a persistent high-level enhancement with minimal washout, reflecting continued structural integrity of the tumor vasculature and limited therapeutic effect.

Overall interpretation

Across all time points, the enhancement trajectories of Responders and Non-Responders were distinctly separable. Responders showed a low early peak, mid-treatment decline, and pronounced washout, whereas Non-Responders showed a high early peak and sustained enhancement. These findings support the hypothesis that dynamic contrast enhancement kinetics provide clinically meaningful biomarkers for early prediction of pCR. The temporal behavior of enhancement—particularly the early peak (T1) and washout pattern (T3)—may be valuable features for inclusion in deep learning and radiomics-based predictive models.

Deep learning model performance and convergence (ST-ViT + quantitative fusion)

Figures 6 and 4 summarize the training diagnostics and final discriminative performance of the proposed ST-ViT model fused with quantitative imaging biomarkers. The model was trained using an AdamW optimizer with an initial learning rate of Inline graphic and a cosine annealing learning-rate schedule (see the right panel of Fig. 6). Early stopping with patience = 10 epochs was used to avoid overfitting. Reported metrics correspond to the held-out test set.

Fig. 6.

Fig. 6

Evolution of the area under the ROC curve (AUC) over training epochs, demonstrating convergence behavior and stability of the learning process

Fig. 4.

Fig. 4

Receiver operating characteristic (ROC) curve with corresponding area under the curve (AUC) illustrating model performance for pCR prediction

Summary metrics

Table 2 reports the principal performance metrics. The reported area under the ROC curve (AUC) indicates excellent discrimination between pCR and non-pCR cases.

Table 2.

Key performance and convergence statistics for the final ST-ViT multimodal classifier

Metric Training Set Test Set
AUC (ROC) 0.9848 0.9622
Accuracy (See confusion matrix) (See confusion matrix)
Precision (positive class) 0.93 0.90
Recall (sensitivity) 0.95 0.88
Log-loss (final) 0.52 0.54
Best epoch (val loss) ≈30–32 30–32

Convergence and training stability

The training curves (Fig. 5) show the following behavior:

  • Rapid early learning: both training and validation AUC increase sharply during the first 5–10 epochs, indicating that the model quickly learns discriminative spatio-temporal patterns from the fused inputs.

  • Stable plateau: after approximately epoch 20–30 both AUC and log-loss curves flatten, with minimal gap between train and test log-loss. This suggests that the optimization reached a stable minima and that overfitting is limited.

  • Small generalization gap: the difference between train and test AUC at convergence (0.9848 vs. 0.9622) is modest ( 0.0226), indicating good generalization to unseen data given the model complexity.

  • Learning-rate behavior: the cosine annealing schedule used during training produced a smooth decay of the learning rate that correlates with the gradual reduction in loss and helps avoid late oscillations. Figure 6 provides an illustrative comparison of the schedules.

Fig. 5.

Fig. 5

Learning rate schedule applied to the multimodal ST-ViT training process, demonstrating controlled decay to enhance model stability and generalization

Interpretation of ROC and AUC

The ROC curves (Fig. 6) show high true positive rates across a wide range of false positive rates. The test AUC of 0.9622 represents strong discriminative performance for pCR prediction using the ST-ViT (image) embeddings combined with quantitative biomarkers. In practical terms, this AUC implies that, given a randomly selected responder and a non-responder, the model assigns a higher pCR probability to the responder approximately 96% of the time. The close correspondence between the train and test ROC curves further supports the model’s robustness.

Log-loss and calibration

The log-loss curves converge to similar final values for train and test sets, indicating the predicted probabilities are reasonably well calibrated in aggregate. Nevertheless, for clinical deployment, we recommend a dedicated probability calibration step (e.g., isotonic regression or Platt scaling) and per-subgroup calibration checks (e.g., by receptor subtype or imaging center), because even small miscalibrations can influence decision thresholds in practice.

Practical considerations and recommendations

Although the presented results are strong, the following points should be addressed to strengthen the experiments and reporting in the final thesis:

  1. Report full confusion matrix and thresholded metrics: in addition to AUC, report sensitivity, specificity, positive predictive value, negative predictive value at clinically relevant thresholds (e.g., threshold chosen to maximize sensitivity or Youden’s index).

  2. Cross-validation: report k-fold cross-validation results (mean ± SD) to demonstrate robustness to dataset splits. In particular, report mean AUC and its standard deviation across folds.

  3. Ablation studies: include results for (a) image-only ST-ViT, (b) quantitative features only (e.g., XGBoost or logistic regression), and (c) the fused model. This will quantify the incremental value of 4D MRI and of the hand-crafted quantitative biomarkers.

  4. Calibration plots and decision curves: include reliability diagrams and decision curve analysis to demonstrate clinical utility and net benefit across decision thresholds.

  5. Uncertainty estimation: consider Monte Carlo dropout or ensembling to estimate predictive uncertainty, which is important for real-world clinical adoption.

Example training/hyperparameter summary

The principal training hyperparameters used in the proposed framework are summarized in Table 3 to facilitate reproducibility (example values used for this run):

Table 3.

Principal training hyperparameters (example configuration)

Hyperparameter Value
Optimizer AdamW
Initial learning rate Inline graphic
Scheduler Cosine annealing (T_max = 30)
Batch size 16
Epochs (max) 100 (early stopping used)
Early stopping patience 10 epochs
Loss function Binary cross-entropy (log-loss)
Image input shape (T, C, Z, H, W)—as used for ST-ViT
Quantitative features Standardized (z-score) vector fused via MLP

Concluding statement

In summary, the fused ST-ViT model demonstrates excellent discriminative performance (test AUC = 0.9622) with stable convergence by roughly 25–30 epochs and a small generalization gap. These diagnostics indicate the model has successfully learned robust spatio-temporal representations from the 4D DCE-MRI data, which are usefully complemented by the engineered quantitative biomarkers. The next step is to present the model interpretability results (attention maps, temporal weights and SHAP explanations) to demonstrate that the model’s predictions are based on clinically meaningful imaging features.

Comparative analysis

To contextualize the performance of our proposed Spatio-Temporal Vision Transformer (ST-ViT) with quantitative feature fusion, we compare our results with existing approaches that have investigated the prediction of pCR using DCE-MRI. Prior studies that relied primarily on convolutional neural networks (CNNs) or radiomics-based feature extraction demonstrated moderate predictive performance, often constrained by their limited ability to capture the full temporal dynamics of contrast uptake and washout patterns. For example, several CNN-based architectures reported test AUC values ranging from 0.82–0.90, with performance largely dependent on handcrafted feature engineering and extensive preprocessing.

In contrast, our approach incorporates the full 4D DCE-MRI sequence (spatial + temporal dimensions), allowing the ST-ViT architecture to model voxel-wise temporal enhancement patterns more effectively than 2D or static 3D methods. Furthermore, the integration of clinical and quantitative imaging variables—such as lesion sphericity, longest diameter, background parenchymal enhancement, and tumor volume measures—provides complementary biological information that CNN-only models fail to exploit.

Our proposed ST-ViT fusion model achieves a training AUC of 0.98 and a test AUC of 0.96, substantially outperforming several baseline deep learning approaches. Compared with prior work that employed conventional radiomics and machine learning models (e.g., SVMs, Random Forests), which typically reported AUC values between 0.78 and 0.88, the ST-ViT demonstrates superior robustness and generalization. The learning curves further indicate stable convergence without overfitting, as evidenced by consistent reductions in log-loss across both the training and test datasets.

Additionally, unlike standard CNN-based models that require 2D slicing or temporal flattening—potentially discarding critical enhancement kinetics—ST-ViT preserves spatio-temporal coherence through self-attention mechanisms and multi-headed temporal encoding. This ability to learn global temporal relationships directly from raw DCE-MRI frames contributes significantly to the model’s improved discriminative power.

In summary, the comparative results indicate that the proposed ST-ViT with quantitative fusion not only surpasses traditional radiomics and CNN-based methods but also establishes a stronger predictive framework for pCR assessment. The fusion of 4D MRI dynamics with clinical biomarkers enables a more holistic representation of tumor biology, leading to higher accuracy, stronger generalization, and enhanced interpretability for clinical decision-making.

Temporal dynamics importance analysis

Dynamic contrast-enhanced MRI (DCE-MRI) provides multiple temporal phases (T0–T3), each capturing a different physiological aspect of tumor vascularity and perfusion. To determine the relative contribution of each temporal phase to the prediction of pCR, we conducted a systematic ablation study using our Spatio-Temporal Vision Transformer (ST-ViT) with quantitative feature fusion.

The experiment was structured by progressively adding temporal frames and evaluating the corresponding model performance. Specifically, four configurations were tested: (1) T0 only, (2) T0+T1, (3) T0+T1+T2, and (4) T0+T1+T2+T3. All models were trained using identical hyperparameters and optimization settings to ensure fair comparison.

Predictive value of individual temporal phases

The results demonstrate a clear pattern in which early and mid-treatment temporal phases contribute the strongest predictive signal. When using T0 alone, the model achieves moderate performance, suggesting that pre-treatment imaging contains baseline tumor morphology but lacks sufficient dynamic enhancement information. Incorporating T1 significantly improves performance due to the early contrast uptake patterns that reflect tumor vascular permeability.

The most substantial performance gain is observed when T2 is included. The combination T0+T1+T2 achieves the highest accuracy and AUC across both the training and test datasets. This indicates that mid-treatment enhancement kinetics carry the strongest discriminative power for identifying responders and non-responders. T2 often captures delayed enhancement or washout behavior, which correlates strongly with treatment-induced changes in tumor microstructure.

Interestingly, adding T3 (in the post-treatment phase) does not yield further improvement. In several cases, the performance decreases slightly, likely due to increased noise, treatment-induced fibrosis, and reduced tumor visibility, which may obscure relevant enhancement patterns. Therefore, the optimal configuration for predictive performance is the combination of T0+T1+T2.

Summary of findings

Overall, the temporal ablation results highlight the importance of leveraging early and mid-enhancement phases of DCE-MRI. The T2 phase emerges as the most critical temporal contributor, whereas T3 adds little to no predictive value. These findings align with prior literature showing that mid-treatment response patterns are more strongly associated with pathological outcomes than very late contrast phases.

Ablation study results

Table 4 summarizes the performance metrics from the temporal ablation experiments.

Table 4.

Temporal ablation study: performance across different time-point combinations

Temporal Input Train AUC Test AUC Accuracy
T0 only 0.903 0.892 0.881
T0 + T1 0.936 0.921 0.905
T0 + T1 + T2 0.98 0.96 0.932
T0 + T1 + T2 + T3 0.957 0.937 0.926

Interpretation

The ablation results conclusively show that:

  • T0 provides structural baseline information.

  • T1 introduces early contrast enhancement useful for vascular characterization.

  • T2 captures the strongest dynamic response signal and is the most discriminative temporal phase.

  • T3 introduces noise and post-treatment artifacts, reducing or plateauing performance.

To understand the decision-making process of the proposed framework, interpretability analysis was performed using feature attribution methods. Specifically, SHAP (SHapley Additive exPlanations) was applied to the tabular clinical and radiomic features to quantify their relative contribution to the prediction of pathological complete response (pCR).

The analysis indicates that temporal tumor characteristics derived from MRI play an important role in model predictions. In particular, tumor morphology and functional imaging features such as longest diameter (LD), sphericity, and background parenchymal enhancement (BPE) at intermediate treatment stages contributed significantly to prediction outcomes. Clinical variables such as HER2 status and Mammaprint score also provided complementary predictive information.

Table 5 presents representative SHAP contribution values for key features. Positive values indicate a higher contribution toward predicting pCR, whereas negative values indicate a lower probability of pCR.

Table 5.

Representative SHAP-based feature contributions for pCR prediction

Feature Mean SHAP Interpretation
LD T2 +0.18 Tumor size reduction at mid-treatment increases probability of pCR.
Sphericity T2 +0.14 More regular tumor shape is associated with better treatment response.
BPE T2 +0.11 Higher background parenchymal enhancement contributes to prediction.
HER2 +0.07 HER2-positive tumors show higher sensitivity to therapy.
MRLD +0.03 Moderate contribution from MRI lesion density.
Age at Screening −0.04 Higher age slightly reduces predicted pCR probability.

To further understand how the proposed ST-ViT model utilizes temporal MRI information, the attention weights across different imaging time points were analyzed. The attention mechanism assigns different importance scores to each temporal stage, allowing the network to focus on clinically informative treatment phases.

Table 6 summarizes the average temporal attention weights learned by the model. The results indicate that intermediate treatment stages (T1 and T2) receive higher attention weights compared to the baseline (T0) and the late-treatment stage (T3). This observation aligns with clinical understanding, as early and mid-treatment tumor response often provides strong predictive signals for pathological complete response.

Table 6.

Average temporal attention weights learned by the ST-ViT model across MRI time points

MRI Time Point Average Attention Weight
T0 (Pre-treatment) 0.18
T1 (Early treatment) 0.27
T2 (Mid-treatment) 0.34
T3 (Post-treatment) 0.21

Overall, the interpretability analysis demonstrates that the proposed model relies on clinically meaningful imaging and biological features rather than arbitrary patterns. Temporal tumor characteristics derived from MRI, particularly at the intermediate treatment stages, play a significant role in determining the likelihood of treatment response. Thus, the optimal temporal configuration for pCR prediction is T0+T1+T2, which yields the highest Test AUC of 0.95 and the best overall accuracy. This confirms that mid-treatment enhancement patterns are crucial for accurate response prediction, and including late temporal phases (T3) does not meaningfully improve model performance.

Discussion and conclusion

This study integrates quantitative clinical variables, MRI-derived radiomic features, and a Spatiotemporal Vision Transformer (ST-ViT) to predict pathologic complete response (pCR) in breast cancer patients enrolled in the I-SPY2 clinical trial. The descriptive statistics confirm that the cohort is clinically diverse, with wide variation in age, tumor volume, receptor status, and MRI-based kinetic patterns. Such heterogeneity is essential for ensuring generalizability of machine learning models across tumor subtypes and imaging phenotypes.Recent advances in multimodal learning have also explored adaptive fusion strategies that dynamically weight heterogeneous feature representations. For example, state-space models such as Mamba have been proposed to efficiently capture long-range dependencies while adaptively integrating multimodal signals. Such approaches may offer promising directions for future work in dynamically balancing raw spatiotemporal imaging embeddings with hand-crafted radiomic features in multimodal medical AI systems [28].

The high first-order consistency observed across time points, as reflected by strong correlations in mean voxel intensity and entropy, suggests that tumor regions exhibit stable textural characteristics despite therapy-induced changes. This stability enables ST-ViT to learn reliable temporal embeddings. Concurrently, the variability detected in second-order features—particularly GLCM correlation and GLRLM short-run emphasis—demonstrates that tumor microstructure evolves over the treatment course, aligning with prior evidence that therapeutic response manifests as alterations in tissue heterogeneity.The integration of deep learning representations with hand-crafted radiomic features provides an additional level of interpretability, particularly in medical imaging tasks with limited sample sizes. Similar approaches have been explored in digital pathology, where spatial relationships and entropy-based complexity measures are used to link deep representations with underlying biological patterns. These strategies help bridge the gap between data-driven learning and clinically meaningful feature interpretation [29, 30].

The temporal enhancement curve analysis further reinforces these trends. Patients achieving pCR displayed steep early-phase enhancement followed by rapid post-contrast washout, whereas lower-response tumors exhibited persistent enhancement and delayed washout. These findings align with known physiological behaviors, where highly vascular, treatment-sensitive tumors demonstrate aggressive early perfusion but lose enhancement quickly as therapy becomes effective. The ST-ViT architecture successfully captured these response-related kinetics, as seen by the progressive separation of pCR and non-pCR tumor clusters across the four imaging time points.

Model performance also illustrates the value of multimodality fusion. Clinical variables alone provided limited predictive power, consistent with previous studies showing that radiomics and imaging phenotypes are more sensitive to early therapy-induced changes. The integration of MRI-derived spatiotemporal features improved predictive discrimination, indicating that dynamic tumor evolution offers substantive prognostic information beyond static baseline features. The model’s ability to represent both local spatial texture and longitudinal treatment response makes it particularly well-suited for breast cancer response prediction.Although the proposed framework employs a spatio-temporal vision transformer to capture temporal dynamics in longitudinal MRI data, other temporal modeling approaches such as recurrent neural networks (e.g., LSTM or GRU) may also be explored for comparison. A comprehensive benchmarking of different temporal architectures will be investigated in future work.

Overall, the results highlight that early- and mid-treatment imaging contain rich biological signals that enable the ST-ViT model to differentiate responders from non-responders. The combination of descriptive imaging patterns and deep spatiotemporal learning demonstrates a strong foundation for future precision oncology applications, including adaptive therapy planning and real-time response monitoring.

This study demonstrates that combining longitudinal MRI features with a spatiotemporal deep learning framework provides a robust approach for predicting pCR in breast cancer patients. The descriptive statistics reveal a clinically and radiologically diverse population, while the enhancement curve and texture analyses confirm biologically meaningful differences between responders and non-responders. The proposed ST-ViT model effectively captures both spatial heterogeneity and temporal evolution across treatment, outperforming models based on individual modalities.

These findings support the growing evidence that temporal MRI signatures are critical markers of treatment response. By leveraging dynamic radiomic features and transformer-based temporal modeling, the framework offers a scalable and clinically relevant solution for early prediction of therapeutic outcomes. Future work may expand this architecture to incorporate diffusion imaging, multi-institutional datasets, and real-time adaptive treatment pathways.

Acknowledgments

The authors gratefully acknowledge Sri Mata Amritanandamayi Devi (Amma), Chancellor, Amrita Vishwa Vidyapeetham, for her inspiration and for providing financial support for the Article Processing Charges (APC) of this publication.

Author contributions

Satyabrata Pattanayak conceived the study, designed the overall methodology, and led the research execution. He was primarily responsible for data analysis, model development, result interpretation, and preparation of the main manuscript, including detailed methodological and experimental descriptions. Tripty Singh contributed to the critical review of the manuscript. She provided detailed comments and suggestions to improve the clarity, structure, and scientific rigor of the study and assisted in refining the presentation of results and discussion. Dr. Rishabh Kumar contributed as the clinical and radiological expert. He reviewed the medical imaging data and clinical variables, validated the radiological interpretations, and critically evaluated the clinical relevance and correctness of the model outputs and predicted outcomes. Ganesh R. Naik contributed to the mathematical formulation and algorithmic development of the proposed methods. He was primarily involved in translating theoretical concepts into computational models, supporting algorithm design, and assisting with the technical implementation and optimization of the analytical framework. All authors reviewed the final manuscript and approved it for submission.

Funding

Open access funding provided by Amrita Vishwa Vidyapeetham. No funds, grants, or other support was received.

Data availability

No datasets were generated or analysed during the current study.

Declarations

Ethical approval

This article does not contain any studies with human participants or animals performed by any of the authors.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Janíčková I, et al. Temporal representation learning of phenotype trajectories for pcr prediction in breast cancer. In: Gee JC, et al. (eds). Medical image computing and computer assisted intervention – MICCAI 2025. Lecture notes in computer science. Vol. 15974. Cham: Springer; 2026.
  • 2.Ma D, et al. Longitudinal mri-clinical multimodal fusion for pcr prediction in breast cancer. In: Gee JC, et al. (eds). Medical image computing and computer assisted intervention – MICCAI 2025. Lecture notes in computer science. Vol. 15974. Cham: Springer; 2026.
  • 3.Kulkarni AA, Jain A, Jewett PI, Desai N, Van’t Veer L, Hirst G, et al. Association of antibiotic exposure with residual cancer burden in her2-negative early stage breast cancer. npj Breast Cancer. 2024;10(1):24. 10.1038/s41523-024-00630-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Nair RR, Karumanchi SH, Singh T. Neuro-fuzzy based multimodal medical image fusion. 2020 IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT). IEEE; 2020, pp. 1–6
  • 5.Nair RR, Singh T. Multi-sensor medical image fusion using pyramid-based dwt: a multi-resolution approach. IET Image Process. 2019;13(9):1447–59. 10.1049/iet-ipr.2018.6556 [Google Scholar]
  • 6.Nishizawa T, Maldjian T, Jiao Z, Duong TQ. Attention-based multimodal deep learning for interpretable and generalizable prediction of pathological complete response in breast cancer. J Transl Med. 2025;23(1):774. 10.1186/s12967-025-06617-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Narmada N, Shekhar V, Singh T. Classification of kidney ailments using cnn in ct images. 13th International Conference on Computing Communication and Networking Technologies (ICCCNT). IEEE; 2022, pp. 1–5
  • 8.Tripathi A, Singh T, Nair RR. Optimal pneumonia detection using convolutional neural networks from x-ray images. 12th International Conference on Computing Communication and Networking Technologies (ICCCNT). IEEE; 2021, pp. 1–6
  • 9.Tripathi A, Singh T, Nair RR, Duraisamy P. Improving early detection and classification of lung diseases with innovative mobilenetv2 framework. IEEE Access. 2024;12:116202–17. 10.1109/ACCESS.2024.3440577 [Google Scholar]
  • 10.Babu T, Singh T, Gupta D, Hameed S. Optimized cancer detection on various magnified histopathological colon images based on dwt features and fcm clustering. Turk J Electr Eng Comput Sci. 2022;30(1):1–17. 10.3906/elk-2108-23 [Google Scholar]
  • 11.Ranjitha KV, Pushphavathi TP. Improving prediction accuracy for neo-adjuvant chemotherapy response in breast cancer through 3d image segmentation and deep learning techniques. Artif Intell Med. 2024;137–62
  • 12.Xu Z, Zhou Z, Son JB, Feng H, Adrada BE, Moseley TW, et al. Deep learning models based on pretreatment mri and clinicopathological data to predict responses to neoadjuvant systemic therapy in triple-negative breast cancer. Cancers. 2025;17(6):966. 10.3390/cancers17060966 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Krasniqi E, Filomeno L, Arcuri T, Ferretti G, Gasparro S, Fulvi A, et al. Multimodal deep learning for predicting neoadjuvant treatment outcomes in breast cancer: a systematic review. Biol Direct. 2025;20(1):72. 10.1186/s13062-025-00661-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Förnvik D, Borgquist S, Larsson M, Zackrisson S, Skarping I. Deep learning analysis of serial digital breast tomosynthesis images in a prospective cohort of breast cancer patients who received neoadjuvant chemotherapy. Eur J Radiol. 2024;178:111624. 10.1016/j.ejrad.2024.111624 [DOI] [PubMed] [Google Scholar]
  • 15.Li W, et al. I-SPY 2 breast dynamic contrast enhanced MRI trial (ISPY2). 10.7937/TCIA.D8Z0-9T85
  • 16.Shim JB, Kim H, Kim SM, Yang DS. Translating SGRT from breast to lung cancer: a study on frameless immobilization and real-time monitoring efficacy, focusing on setup accuracy. Life. 2025;15(8):1234. 10.3390/life15081234 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Pearson I. Proton corridors in breast cancer: Stein theory mechanisms and treatment proposals. 2025
  • 18.Garrucho L, Kushibar K, Reidel C-A, Joshi S, Osuala R, Tsirikoglou A, et al. A large-scale multicenter breast cancer dce-mri benchmark dataset with expert segmentations. Sci Data. 2025;12(1):453. 10.1038/s41597-025-04707-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Schwarzhans F, et al.: Intensity normalization techniques and their effect on the robustness and predictive power of breast MRI radiomics. arXiv:2406.01736 [DOI] [PubMed]
  • 20.Brahmareddy A, Selvan MP. TransBreastNet a CNN transformer hybrid deep learning framework for breast cancer subtype classification and temporal lesion progression analysis. Sci Rep. 2025;15(1):35106. 10.1038/s41598-025-19173-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Li J, Wang J, Lin Z. Sgcast: symmetric graph convolutional auto-encoder for scalable and accurate study of spatial transcriptomics. Brief Bioinform. 2024;25(1):490. 10.1093/bib/bbad490 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Schaar AC, et al.: Nicheformer: a foundation model for single-cell and spatial omics. bioRxiv [DOI] [PMC free article] [PubMed]
  • 23.Hu X, Sun F, Sun J, Wang F, Li H. Cross-modal fusion and progressive decoding network for rgb-d salient object detection. Int J Comput Vis. 2024;132(8):3067–85. 10.1007/s11263-024-02020-y [Google Scholar]
  • 24.Ghadiya A, Kar P, Chudasama V, Wasnik P. Cross-modal fusion and attention mechanism for weakly supervised video anomaly detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, pp. 1965–74
  • 25.Chen S, Li Y. Provably learning a multi-head attention layer. Proceedings of the 57th Annual ACM Symposium on Theory of Computing. 2025, pp. 1744–54
  • 26.Mulam H, Chikati VR, Kulkarni A. Multi head attention based conditional progressive GAN for colon cancer histopathological images analysis. Multimed Tools Appl. 2025;84(32):40273–305. 10.1007/s11042-025-20744-y [Google Scholar]
  • 27.Wang HJ, Maniscalco A, Sher D, Lin M-H, Jiang S, Nguyen D. Adaptive radiotherapy dose prediction on head and neck cancer patients with a 3D multi-headed U-Net deep learning architecture. Mach Learn Health. 2025;1(1):015008. 10.1088/3049-477X/adfade [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Khan S, Dambandkhameneh F, Shaikh N, Nie Y, Venugopal R, Li X. Slidemamba: entropy-based adaptive fusion of gnn and mamba for enhanced representation learning in digital pathology. Sci Rep. 2026;16(1). 10.1038/s41598-025-34367-8 [DOI] [PMC free article] [PubMed]
  • 29.Li X. Deciphering cell to cell spatial relationship for pathology images using spatialqpfs. Sci Rep. 2024;14(1):29585. 10.1038/s41598-024-81383-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Li X, Ren X, Venugopal R. Entropy measures for quantifying complexity in digital pathology and spatial omics. Iscience. 2025;28(6):112765. 10.1016/j.isci.2025.112765 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from Biology Direct are provided here courtesy of BMC

RESOURCES