Abstract
Precise prediction of cancer therapy response remains challenging because conventional biomarkers capture molecular features but overlook the physical state of tumors. We developed a multimodal deep-learning framework integrating ultrasound shear wave elastography (SWE) images with quantitative stiffness measurements (elastic modulus, kPa) to predict treatment outcomes in preclinical murine tumors. Each image–stiffness pair is tokenized within a transformer that learns interactions between local elastographic texture and global rigidity. A lightweight convolutional encoder extracts image features, the modulus is embedded as a numeric token, and self-attention fuses both modalities for classification. Trained on 1578 baseline SWE images from five syngeneic tumor models, the model classified tumors as responders, stable, or non-responders. Across five random seeds, it achieved 92.4% ± 1.3% accuracy, macro-F1 0.92, and ROC-AUC 0.99 on a held-out test set, with well-calibrated probabilities. Matched-split ablations—image-only, stiffness-token-removed, stiffness-shuffled, late-fusion, and stiffness-only—showed that performance reflected genuine cross-modal learning rather than scalar stiffness alone; shuffling image–stiffness pairings significantly reduced accuracy (all corrected p < 0.01). Leave-one-tumor-model-out analysis demonstrated generalization to unseen tumor types, with 95.5% ± 1.5% accuracy (range 93.9–97.9%) across five held-out models. Lower baseline stiffness correlated with better response, supporting the hypothesis that mechanically normalized tumors respond more effectively. These preclinical proof-of-concept findings establish tumor mechanics as candidate predictive biomarkers and transformer-based multimodal learning as a scalable approach to biomechanically informed response prediction, while requiring validation in human cohorts.
Keywords: Shear wave elastography, Elastic modulus, Multimodal learning, Vision transformer, Tumor microenvironment, Cancer therapy response, Medical image classification
Graphical abstract

Highlights
-
•
A multimodal transformer predicts cancer therapy response from tumor imaging.
-
•
Fusing elastography images with tissue stiffness improves response prediction.
-
•
Shuffling the image-stiffness pairing lowers accuracy, showing cross-modal learning.
-
•
The model generalizes to tumor types entirely absent from its training data.
-
•
Softer tumors at baseline respond better to treatment in preclinical mouse models.
1. Introduction
In the fight against cancer, tumor heterogeneity [1] – both between patients and within an individual tumor over time – remains a major challenge, often leading to divergent therapeutic outcomes. Some patients respond well to treatment, whereas others show minimal benefit or develop resistance [2]. A central goal of oncology is thus to predict, for each patient, which therapies are most likely to succeed, moving closer to personalized medicine. The tumor microenvironment (TME) is defined by biomechanical abnormalities, such as elevated mechanical stresses, increased stiffness, and impaired perfusion, that collectively drive therapeutic resistance [[3], [4], [5]]. As tumors expand within the confined space of the host tissue, they experience mechanical compression that collapses intratumoral blood and lymphatic vessels, drastically reducing perfusion and causing hypoxia and acidosis [6,7] (Fig. 1A). These conditions compromise oxygen-dependent treatments such as radiotherapy, inhibit delivery of blood-borne chemotherapies, and foster genetic instability, angiogenesis, invasion, and metastasis [8]. Hypoxia also remodels the immune landscape by upregulating checkpoint molecules, polarizing macrophages toward immunosuppressive phenotypes, and impairing T cell effector function [9].
Fig. 1.

Tumor mechanopathology and multimodal transformer-based prediction of therapy response. (A) Abnormal versus normalized tumor microenvironment (TME). Solid tumors generate mechanical forces that compress blood vessels, impair perfusion, and create hypoxic regions, limiting drug and immune cell delivery and driving therapy resistance. Mechanotherapeutic agents remodel the stroma, reduce mechanical stress, decompress vessels, and restore perfusion, enhancing therapeutic delivery and efficacy. (B) Experimental workflow in murine tumor models. Cancer cells were implanted in mice, producing tumors with elevated stiffness and abnormal mechanics. Mechanotherapeutic pretreatment preceded chemo-immunotherapy. Tumor stiffness was quantified non-invasively by shear wave elastography (SWE), yielding both elastography images and bulk stiffness values (kPa). Response was assessed by RECIST and classified as response, stable disease, or non-response. (C) Multimodal transformer model. SWE images pass through a convolutional neural network (CNN) backbone that extracts deep texture features as image tokens, while the stiffness value (elastic modulus, kPa) is embedded as a separate numeric token. Both token types enter a transformer encoder, whose self-attention captures relationships between local elastographic patterns and global stiffness. The classifier outputs the predicted response category (non-responder, stable disease, responder). Created with BioRender.com.
These barriers are most evident in desmoplastic tumors. Cancer-associated fibroblasts deposit dense networks of collagen and hyaluronan, producing a highly crosslinked extracellular matrix [10,11]. Collagen confers tensile strength, while hyaluronan absorbs water and swells, elevating interstitial fluid pressure and further compressing vessels. These matrix alterations markedly increase tissue stiffness, frequently measured by the elastic modulus – the ratio of applied stress to induced strain and a direct quantitative index of rigidity [11]. Tumors often exhibit elastic modulus values an order of magnitude greater than adjacent normal tissues, creating formidable barriers to perfusion, immune infiltration, and drug delivery [12,13].
While predictive biomarkers have traditionally focused on molecular and genomic features, their clinical translation has been slow, with only a handful of assays gaining regulatory approval [14,15]. In this context, the mechanopathology of solid tumors offers a compelling but underexplored source of biomarkers for treatment response, as evidenced by recent work on mechanical barrier characterization [3], immune microenvironment regulation by mechanical forces [4], drug delivery optimization through microenvironment reengineering [5], and vascular function restoration via mechanical modulation [7]. Strategies that normalize or remodel the TME (Fig. 1A) have shown therapeutic benefit: carefully dosed anti-angiogenic agents can restore perfusion and alleviate hypoxia, while stromal re-engineering – through enzymatic degradation of hyaluronan, inhibition of collagen crosslinking, or antifibrotic drugs such as losartan or ketotifen – reduces solid stress, decompresses vessels, and enhances drug penetration and immunostimulation. These approaches exemplify mechanotherapeutics: interventions that modulate tumor mechanics to overcome biophysical barriers and synergize with conventional treatments [7,16]. Increasing evidence supports incorporating biomechanical markers – such as stiffness (quantified by the elastic modulus) and perfusion – into predictive models for therapy response, underscoring the dual role of tumor mechanics as both barriers and actionable biomarkers [17,18]. Tumor stiffness can now be measured in patients via imaging. Ultrasound shear wave elastography (SWE) is a non-invasive technique that quantifies tissue elasticity by tracking the propagation of shear waves through tissue [19]. SWE has broad clinical utility – for instance, characterizing breast lesions, staging rectal tumors, and distinguishing malignant from benign thyroid nodules with high accuracy [20,21]. Beyond diagnosis, SWE is being explored for monitoring treatment response [22,23]. Early studies in breast cancer have shown that reductions in SWE-derived stiffness during neoadjuvant chemotherapy correlate strongly with pathological complete response, whereas resistant tumors often remain stiff [13]. Such elastography-based metrics could therefore serve as early-response biomarkers, enabling timely treatment adaptation.
Transformers – deep learning models built around self-attention – have revolutionized sequence and language modeling and are increasingly applied to biomedical imaging and multimodal tasks [24]. Unlike convolutional neural networks (CNNs), which extract features hierarchically from local neighborhoods, transformers learn pairwise relationships between all input tokens – whether image patches or embedded scalar features [25]. This captures both local texture (e.g., fine elastographic patterns) and global dependencies (e.g., how regional stiffness variations relate to the overall elastic modulus). Importantly, the transformer's modality-agnostic token design is well-suited to heterogeneous data streams, enabling the elastic modulus to be treated as an equal partner to image tokens in the self-attention process [26]. The model can thus learn nuanced biomechanical interactions that conventional late-fusion CNNs often miss. Building on these advances, we developed a multimodal framework tailored for SWE. Each elastography image was paired with its elastic modulus (kPa), ensuring that both spatial texture and quantitative stiffness were represented. A CNN backbone encodes fine elastographic features, while the scalar stiffness value is embedded into a parallel numeric token. These embeddings are assembled into a joint sequence and processed by a transformer encoder (Fig. 1C), whose self-attention learns how local image heterogeneities relate to global stiffness magnitude. The challenge of fusing heterogeneous modalities — where images must be jointly modeled with scalar or tabular features — remains an active area of research. Most existing multimodal frameworks in medical image classification address the fusion of two image-based modalities (e.g., MRI and PET) or of images with free-text reports. Self-adaptive image–text fusion approaches have recently been developed, highlighting the persistent difficulty of bridging semantic gaps between modalities with fundamentally different representations [27]. Transformer-based architectures capture inter-modal dependencies that convolutional approaches overlook [28], and attention-based fusion strategies, including mask-attention mechanisms, consistently outperform fixed-weight alternatives [29]. However, comparatively little attention has been given to the fusion of images with scalar numeric features — a setting that arises naturally in clinical imaging, where quantitative biomarkers (e.g., tissue stiffness, perfusion indices, or blood-based measurements) accompany each scan.
Our work addresses this gap by treating the scalar elastic modulus as a dedicated token within the transformer's attention, enabling direct cross-modal interaction between image patches and a numeric biomarker — a fusion paradigm that extends beyond the image–image and image–text settings dominant in current literature. Trained on 1578 baseline SWE images from five distinct tumor models (Fig. 1B), this multimodal transformer achieved strong predictive performance, with an overall accuracy of 0.92 in classifying treatment outcomes (responder, stable disease, or non-responder). To establish that these gains reflect genuine cross-modal learning rather than the scalar stiffness value alone — and that the model generalizes beyond the specific tumor types seen during training — we further perform systematic fusion ablations (image-only, stiffness-token-removed, stiffness-shuffled, and late-fusion baselines) and a leave-one-tumor-model-out generalization analysis, with all comparisons reported using confidence intervals and multiple-comparison correction. To our knowledge, this represents the first systematic application of a transformer-based architecture to SWE data for response prediction, and the first fusion of elastographic imaging with quantitative biomechanical indices in a deep learning framework. Beyond technical novelty, the approach illustrates how tumor mechanics could serve as a non-invasive source of predictive information. We emphasize that the present study is a preclinical proof-of-concept in murine tumor models; the learned stiffness–response relationships are not yet validated in humans. Pending such validation, a biomechanically informed framework of this kind could ultimately help identify tumors unlikely to benefit from standard chemo- or immunotherapy alone and that may require mechanotherapeutic intervention, supporting more effective, personalized treatment strategies.
2. Methods
2.1. Cell culture and tumor models
All in vivo experiments have been previously conducted and reported in detail elsewhere [23]. In brief, our team and collaborators cultured a panel of murine cancer cell lines under established, cell-line–specific conditions. The breast adenocarcinoma lines 4T1 and E0771 were maintained in RPMI-1640 supplemented with fetal bovine serum (FBS) and antibiotics; B16F10 melanoma cells were cultured in DMEM with identical supplements; the osteosarcoma K7M2 and fibrosarcoma MCA205 lines were expanded in RPMI-1640 containing L-glutamine, sodium pyruvate, non-essential amino acids, β-mercaptoethanol, FBS, and antibiotics. All cell lines were incubated at 37 °C in a humidified 5% CO2 environment, with specific conditions determined by prior protocols. Syngeneic tumor models were established by implanting defined numbers of cancer cells into immunocompetent hosts. Orthotopic mammary carcinomas were generated by inoculating 4T1 or E0771 cells into the mammary fat pad of female mice, while K7M2, MCA205, and B16F10 cells were implanted subcutaneously into the flanks of male or female mice to model osteosarcoma, fibrosarcoma, and melanoma, respectively.
All animal procedures complied with the animal welfare regulations of the Republic of Cyprus and the European Union. A comprehensive breakdown of the dataset by tumor model and treatment regimen is provided in Fig. 2F and G, with a cross-tabulation of SWE images by cell line and treatment in Supplementary Table S1.
Fig. 2.

Therapy response is shaped by tumor stiffness and modulated by mechanotherapeutic interventions. (A) Class distribution of treatment outcomes across the full dataset (n = 1578): 36.3% responders (n = 573), 31.2% stable disease (n = 492), and 32.5% non-responders (n = 513). (B) Elastic modulus distributions by response category. Ridge density plots show clear separation in stiffness (kPa) between groups: responders had the lowest mean (26.3 ± 5.4 kPa), stable disease intermediate values (38.5 ± 6.8 kPa), and non-responders the highest (52.0 ± 8.3 kPa). The dataset mean (38.4 kPa) is indicated by a red dashed line. (C) Statistical comparison of stiffness across groups. Boxplots confirm significant differences between response categories (***p < 0.001, ANOVA with post-hoc testing); lower elastic modulus was associated with response, while elevated stiffness predicted stable disease or non-response. (D) Treatment modality and outcome. Heatmap of response distributions stratified by therapy type; mechanotherapeutic pretreatment combined with chemo- and/or immunotherapy shifted outcomes toward higher response rates than chemotherapy or immunotherapy alone. (E) Cell-line-specific response patterns across five murine tumor models (4T1, E0771, MCA205, B16F10, K7M2), highlighting heterogeneity in therapeutic sensitivity. (F) Number of tumors and SWE images per treatment regimen. (G) Number of tumors and SWE images per tumor model. Color gradient reflects sample counts, with warmer colors indicating higher numbers.
2.2. Treatment protocol
Mice received mechanotherapy consisting of tranilast 200 mg/kg (Rizaben, Kissei Pharmaceutical, Japan), bosentan 1 mg/kg (Selleckchem, USA) or ketotifen 10 mg/kg (Sigma-Aldrich, USA). When mean tumor volume reached ∼150 mm3, mechanotherapy was initiated. Upon growth to ∼350 mm3, we introduced chemotherapy (Doxil, 3 mg/kg, Epirubicin micelles, 6 mg/kg, intravenous, administered daily), immunotherapy (anti-PD-L1, 10 mg/kg, intraperitoneal, every three days for three total doses), or the combination. Tumor dimensions were monitored throughout using digital calipers. Treatment response was evaluated at a predefined endpoint: the day after the third immunotherapy dose, or seven days after chemotherapy initiation.
2.3. Response classification criteria
Tumor dimensions were monitored using digital calipers. Baseline measurements were recorded immediately prior to chemo-immunotherapy initiation (at approximately 350 mm3), and endpoint measurements were obtained at the predefined timepoint described above. Treatment response was classified according to adapted RECIST 1.1 criteria [30], applied to caliper-derived tumor measurements: Response was defined as a ≥30% decrease in the sum of tumor diameters relative to baseline; Non-Response as a ≥20% increase relative to the smallest recorded sum; and Stable Disease as neither sufficient shrinkage nor sufficient increase to meet either threshold. Labels were assigned at a single predefined endpoint rather than from longitudinal trajectories. The Stable Disease category is inherently the most ambiguous, as tumors near either threshold boundary may be subject to classification uncertainty — a factor consistent with our model performance, where most misclassifications occurred in this class.
2.4. Ultrasound imaging
Shear-wave elastography (SWE) was performed on a Philips EPIQ Elite system equipped with an eL18-4 linear array transducer. The transducer generated and tracked two-dimensional shear waves, and the resulting stiffness maps (kPa) were superimposed on B-mode images using a blue-to-red color scale (soft-to-stiff). A confidence display ensured adequate shear-wave quality within the region of interest, and a ∼1.5 cm gel pad was used to mitigate boundary artifacts. Acquisition parameters were held constant throughout (frequency 10 MHz, acoustic power 52%, B-mode gain 22 dB, dynamic range 62 dB), with tumors positioned so that the focused radiation force impulse targeted the approximate midplane. All imaging for the prognostic task was performed when tumors were approximately 350 mm3 and prior to chemo-immunotherapy. Representative SWE images corresponding to each response category are provided in Supplementary Fig. S1–S3.
2.5. Image preparation and segmentation
Tumor segmentation was carried out using our previously developed Auto-Prognose-CNNattention model (SegforClass) [23]. Because elastograms are overlaid on image anatomy, SWE and B-mode frames were geometrically registered before analysis. Tumor regions were manually annotated to generate ground-truth masks, and a key-point detection algorithm was first applied to crop elastography and B-mode inputs, which were then processed by a U-Net backbone within the SegforClass framework.
To enhance robustness and reduce overfitting, we used training-only augmentations reflecting realistic variability in probe orientation and acoustic conditions. Augmentation parameters were deliberately conservative to preserve the mechanobiological fidelity of SWE images. Unlike natural photographs, elastograms encode quantitative stiffness information through a calibrated color map, where pixel intensity directly corresponds to tissue rigidity in kilopascals. Aggressive spatial transformations or color perturbations could distort this quantitative mapping, introducing artificial stiffness patterns that do not reflect genuine tissue mechanics. We therefore restricted augmentations to horizontal flips (which preserve stiffness values), small rotations (±5%, corresponding to approximately ±18°), translations (±5% of image dimensions), and modest contrast adjustments (±5%). No geometric warping, elastic deformations, or color-space randomization was applied. Validation and test data underwent deterministic preprocessing only. Unless stated otherwise, classification images were standardized to a uniform size, cast to float32, and scaled to the [0, 1] range.
2.6. Image – numeric pairing and dataset assembly
Each image used for classification was paired one-to-one with its corresponding scalar elastic modulus (kPa). Pairing was performed by normalizing filenames to basenames, recording the data-loader order, and joining to a curated annotation table. We verified alignment with spot checks and an assertion on equal counts across images, labels, and numeric features. Samples missing either component were excluded. When multiple SWE images were acquired from the same tumor (typically two to three per animal), animal-level grouping was enforced during data partitioning to prevent information leakage between training, validation, and test sets (see Data splits and cross-validation). The kPa feature was standardized within each training fold only and the same transformation was applied to the corresponding validation and test partitions to prevent information leakage.
2.7. Data splits and cross-validation
Unless otherwise specified, data were partitioned at the animal level into training (70%), validation (15%), and test (15%) sets, with stratification across the three outcome classes. Each animal was assigned to exactly one partition, and all SWE images acquired from that animal (typically two to three, reflecting natural probe-positioning variability) were assigned together with it. No animal therefore contributed images to more than one partition, eliminating information leakage. Because grouping is enforced at the animal level, the validation and test partitions each contain 237 images from 98 unique animals (2.4 images per animal), comparable to the training set (1104 images from 514 animals, 2.2 images per animal). This constraint was enforced identically across all cross-validation folds and across the ablation and leave-one-tumor-model-out analyses. For model selection, we additionally conducted five-fold stratified cross-validation on the 85% development pool (training + validation), keeping the 15% test set untouched for a single final evaluation. Random seeds were fixed across NumPy, TensorFlow, and Python to support reproducibility. The composition of each partition is summarized in Table S2.
To probe the source of predictive performance and its generalizability, two additional evaluation protocols were used. First, a fusion ablation trained five model variants under identical data splits and seeds: the full multimodal model, an image-only model, a model with the stiffness token removed, a model with stiffness values randomly permuted across samples (breaking the image–stiffness pairing while preserving the marginal stiffness distribution), and a late-fusion baseline; a stiffness-only logistic-regression model provided a scalar reference. Second, a leave-one-tumor-model-out (LOTO) analysis trained the model on four tumor models and evaluated it on the entirely held-out fifth, repeated for each of the five tumor models. This analysis requires the tumor model of origin to be known for every image and was therefore restricted to the 1365 images (86.5% of the cohort) for which the cell line could be confirmed against the original acquisition records. The remaining 213 images were retained in the main training, validation, and test analyses — which require only the image, its paired elastic modulus, and its outcome label — but were excluded from the LOTO analysis, since assigning them to a cell line on an inferred basis would compromise the held-out partition. The main and LOTO analyses therefore operate on different subsets of the cohort (1578 and 1365 images, respectively), and their performance figures are not directly comparable; the LOTO results should be read as a generalization test across tumor histologies rather than as a re-estimate of main-analysis performance.
2.8. Deep learning models
We evaluated six multimodal architectures to predict treatment response from SWE images combined with a per-image elastic modulus value. All models accept an RGB image and a standardized scalar (kPa) and output probabilities for response, stable disease, or non-response. Extended layer-by-layer definitions, hyperparameters, and seeds are provided in the repository (see Code Availability).
2.8.1. Baseline multimodal CNN (MModal-Baseline)
The baseline model uses a compact convolutional encoder to extract image features, progressively increasing representational depth while controlling capacity with batch normalization, non-linear activations, spatial down-sampling, and light dropout. Features are aggregated by global pooling into a fixed-length embedding. In parallel, the kPa value is mapped to a small dense embedding. We then fuse image and numeric embeddings by late concatenation and apply a regularized classifier to produce three-way probabilities [23]. Training used momentum-based stochastic gradient descent with a cosine learning-rate schedule and early stopping on validation loss.
2.8.2. CNN with data augmentation (MM-augmentation)
This model retains the baseline architecture but is trained with an augmented image pipeline designed specifically for ultrasound elastography. Augmentations include mirrored views, small rotations and translations, modest contrast and color-space jitter, and low-amplitude noise, all applied on-the-fly to the training split only. The numeric pathway and late-fusion classifier are unchanged. The optimization protocol mirrors MModal-Baseline while allowing for longer training under the cosine schedule, with safeguards to reduce the learning rate and restore the best validation weights if overfitting is detected [31].
2.8.3. CNN with trainable soft attention layer (MM-attention)
To emphasize tumor-relevant regions and suppress background, we integrated two trainable soft-attention modules at intermediate depths of the convolutional encoder. Each module derives a spatial weighting map by comparing localized image descriptors to a global context summary learned by the network. The attention maps are normalized and used to compute attended feature summaries, which are subsequently combined with the numeric kPa embedding. The fused representation passes through a regularized classifier [32,33]. The network is trained end-to-end so that attention weights align with discriminative tumor patterns.
2.8.4. CNN–transformer (MM-transformer)
This model couples a conventional convolutional front-end with a lightweight transformer encoder [34]. Convolutional features are reshaped into a short sequence of tokens that capture spatial context; in parallel, the scalar kPa is projected to a token with matching dimensionality. Positional information is added, and the sequence is processed by multi-head self-attention layers to model longer-range relationships and modality interactions. The resulting representation is summarized and classified. This design allows the model to retain strong local image descriptors from the CNN while leveraging attention to integrate broader context and numeric conditioning.
2.8.5. CNN with trainable soft attention layer and transformer (MM-transformer_attention)
We combined (i) a CNN branch with global-context, compatibility-based soft attention applied at three feature scales, producing three 16-dimensional attended vectors and a 16-dimensional global descriptor; (ii) a Vision Transformer–style branch that tokenizes images into non-overlapping 25×25 patches, adds positional embeddings and a CLS token, and processes the sequence with four pre-normalized multi-head self-attention blocks, yielding a 16-dimensional global representation; and (iii) a numeric branch that embeds the elastic modulus (kPa) via a two-layer MLP to 16 dimensions [35]. The three streams are concatenated and passed through a regularized classifier, culminating in a softmax over three classes. Models were trained with AdamW (initial learning rate 2 × 10−4) using cosine decay with warm restarts, with early stopping and learning-rate reduction on plateau; the loss was sparse categorical cross-entropy.
2.8.6. Vision Transformer (MM-ViT16)
Finally, we further explored a pure transformer approach. Images were partitioned into non-overlapping patches and projected into a sequence of patch tokens [34]. The standardized kPa value was embedded as a dedicated numeric token and prepended to the sequence so that cross-attention could explicitly relate tumor stiffness to image regions. Learnable positional encodings were used throughout. After several pre-norm transformer blocks, the model read out a compact representation—centered on the numeric token—through a small classification head with dropout and weight decay [36].
The same multimodal architecture underlies the ablation variants described in Data splits; individual arms differ only in whether and how the stiffness token is incorporated, so that any performance difference is attributable to the fusion mechanism rather than to architectural or training changes.
2.9. Stiffness-only logistic regression baseline
To quantify the independent predictive contribution of the scalar elastic modulus, we trained a multinomial logistic regression model using only the standardized kPa value as input. The model was implemented in scikit-learn with an L-BFGS solver, maximum 1000 iterations, and default regularization (C = 1.0). The same 70/15/15 stratified split and random seed [37] used for the deep learning models were applied to ensure a fair comparison. Five-fold stratified cross-validation was performed on the development set. This single-feature baseline serves as an ablation control, isolating the contribution of tissue stiffness alone and establishing a lower bound against which the added value of SWE image features can be measured.
2.10. Training procedures and reproducibility
All models were implemented in TensorFlow/Keras. Unless memory required otherwise, we trained with mini-batches of 32, used sparse categorical cross-entropy, and monitored validation loss for early stopping with best-weight restoration. Learning-rate schedules were cosine-decayed; weight decay and dropout were used for regularization. Sessions were cleared between runs, and all operations involving randomization (data shuffling, splitting, initialization) were seeded [37]. Complete training scripts, environment files, and exact hyperparameters are provided in the Supplementary Information and repository. Hyperparameter selection followed an informed, iterative approach rather than exhaustive grid or random search. Initial configurations — including learning rate, batch size, embedding dimensions, dropout rates, and weight decay — were guided by our previously validated CNN-attention framework for the same SWE classification task [23], as well as established defaults from the Keras library for convolutional and transformer architectures. From this starting point, we conducted manual tuning, systematically varying one parameter at a time while monitoring validation loss and accuracy across five-fold cross-validation. Learning rates were explored in the range 1 × 10−4 to 5 × 10−3, batch sizes of 16 and 32 were compared, and transformer depth was varied between one and four encoder blocks. The number of attention heads (2 and 4) and embedding dimensions (16, 32, and 64) were also evaluated. In total, approximately 15–20 configurations per architecture were assessed.
2.11. Evaluation, statistics, and interpretability
Performance was quantified on the held-out 15% test set that remained untouched during development and cross-validation. We report overall accuracy and, to account for class balance, per-class precision, recall, and F1-score, alongside macro one-vs-rest ROC–AUC and macro average precision from precision–recall curves. Model calibration was assessed using reliability (calibration) curves and Brier scores. Where appropriate, pairwise differences in ROC–AUC between architectures were evaluated using DeLong's test (two-sided). To account for training variability, each model configuration was trained over five random seeds, and accuracy, macro-F1, macro one-vs-rest ROC–AUC, and log-loss are reported as the mean with 95% confidence intervals across seeds. Model calibration was additionally quantified by the expected calibration error (ECE). For the fusion ablation, the full model was compared against each ablated variant using paired t-tests across seeds, with Holm–Bonferroni correction for multiple comparisons.
To enable side-by-side appraisal, we generated a set of harmonized visual summaries. First, macro-averaged ROC curves were plotted for all six models on the common test fold, annotated with AUC and 95% CIs. Second, we produced a performance map in which macro AUC (x-axis) and macro F1 (y-axis) are shown per model with bubble size proportional to parameter count, providing an at-a-glance view of the performance–complexity trade-off. To examine class-specific behavior, radar plots depict precision, recall, and F1 for Response, Stable, and Non-Response classes separately. We summarized multi-criteria performance with a heatmap of normalized metrics (AUC, F1, and an inverted parameter proxy for simplicity) and a Pareto-style efficiency plot (AUC per million parameters versus AUC), highlighting models that are simultaneously performant and parameter-efficient.
2.12. Software and computing
Training and evaluation were performed on NVIDIA GPUs (e.g., A100 or V100). Data pipelines used tf. data for shuffling, preprocessing, batching, and prefetching. Experiment tracking captured learning-rate schedules, training/validation curves, predictions, and intermediate embeddings to support downstream visualization and exact reproducibility.
2.13. Data and code availability
SWE/B-mode images, segmentation masks, and the per-image elastic modulus (kPa) table are available from the corresponding author on reasonable request. All source code — including full layer-level model specifications, hyperparameters, random seeds, preprocessing details, fixed data-split files, and training scripts — is openly available at https://github.com/anabiosi-data/Prognostic_AnaBioSi-Data.
3. Results
3.1. Elastic modulus stratifies treatment response groups
To investigate the relationship between tissue stiffness and treatment response, we analyzed the distribution of elastic modulus (kPa) across response categories in a cohort of 1578 samples (Response, n = 573; Stable Disease, n = 492; Non-Response, n = 513). Box plot analysis revealed distinct distributions between groups, with significantly lower median stiffness in the Response group compared with both Stable and Non-Response categories (**p < 0.001, Fig. 2C). Ridge density estimates further highlighted these differences, showing that Response samples clustered at lower elastic modulus values (Ε = 26.3 ± 5.4 kPa), Stable samples at intermediate levels (Ε = 38.5 ± 6.8 kPa), and Non-Response samples at higher levels (Ε = 52.0 ± 8.3 kPa) (Fig. 2B). The overall cohort mean was 38.4 kPa. Importantly, group sizes were well balanced (36.3% Response, 31.2% Stable, 32.5% Non-Response), minimizing class imbalance and supporting robust downstream modeling (Fig. 2A). The dataset comprised 710 animals across five syngeneic tumor models, yielding 1578 SWE images (approximately 2.2 images per tumor). Among the 1365 images with a confirmed cell-line annotation (Methods 2.7), the largest contributors were 4T1 and E0771 breast adenocarcinoma models (346 and 333 images, respectively), followed by MCA205 fibrosarcoma (300), K7M2 osteosarcoma (272), and B16F10 melanoma (114) (Fig. 2G). Among treatment regimens, the combined mechanotherapeutics + immunotherapy + chemotherapy arm contributed the most images (490), while immunotherapy alone contributed the fewest (158) (Fig. 2F). Collectively, these findings indicate that higher elastic modulus is associated with poorer treatment response, while lower stiffness is a hallmark of favorable response [37,38].
3.2. Multimodal transformer and ViT16 architectures integrate imaging and quantitative stiffness features
To test whether combining quantitative tumor stiffness with elastography images enhances prediction of therapeutic response, we designed two multimodal architectures: a CNN–Transformer and a Vision Transformer (ViT16)-based framework (Fig. 3). In both cases, shear wave elastography (SWE) frames (300 × 400 pixels) were paired one-to-one with their corresponding elastic modulus values (kPa) and stratified into training, validation, and independent test sets. Elastic modulus values were standardized within training folds only, ensuring no information leakage. To increase generalizability, the training images were augmented with random flips, small-angle rotations, translations, and contrast adjustments, while all images underwent resizing, normalization, and contrast scaling (Fig. 3A). In the MM-transformer, images were first processed through a compact convolutional backbone with progressively deeper layers (32, 64, and 128 filters) and global average pooling to generate a compact image embedding. In parallel, the scalar stiffness measurement was projected into the same latent space via a dense layer. Both embeddings were reshaped into tokens and concatenated with positional encoding before being passed through a two-layer transformer encoder. Leveraging self-attention, the encoder captured cross-modal relationships between image-derived texture and spatial heterogeneity and the absolute stiffness magnitude, producing a fused representation for classification (Fig. 3B). The ViT16 multimodal variant adopted a pure transformer design. Images were partitioned into non-overlapping 16 × 16 patches and projected into a sequence of patch tokens. The standardized stiffness value was embedded as a dedicated numeric token and prepended to the sequence, allowing cross-attention to explicitly model interactions between global image regions and tumor stiffness. Learnable positional encodings preserved spatial context, and stacked transformer blocks refined the multimodal representation. A classification head centered on the numeric token provided final predictions. These designs enabled modeling of both local SWE texture and global biomechanical context, making stiffness not merely an auxiliary feature but an integrated conditioning factor in the attention process. Across experiments, both multimodal models consistently outperformed single-modality baselines, with the ViT16 framework yielding the highest predictive accuracy and calibration. This represents a methodological advance over conventional CNNs, allowing the models to learn nuanced biomechanical–imaging interactions with direct clinical relevance.
Fig. 3.

Multimodal transformer and Vision Transformer (ViT16) architectures for treatment response prediction from elastography data. (A) Data preprocessing pipeline. SWE images (300 × 400 × 3) were paired with per-image elastic modulus values (kPa) and stratified into training (70%), validation (15%), and test (15%) sets. Training images underwent augmentation (horizontal flips, random rotations ±0.05, translations ±0.05, contrast adjustments ±0.05) to improve robustness. All data were standardized, normalized, and resized before modeling. (B) MM–Transformer and ViT16 multimodal architectures. In the MM–Transformer, the image branch comprises three convolutional blocks (32, 64, 128 filters) with batch normalization, ReLU activation, max pooling, and dropout, followed by global average pooling to produce a 128-dimensional image token; in parallel, the numeric branch projects stiffness values into a 128-dimensional embedding. In the ViT16 variant, SWE images are partitioned into 16 × 16 non-overlapping patches, yielding 475 patch tokens (128 dimensions each); a dedicated numeric stiffness token is prepended to the sequence, enabling cross-attention between global image regions and stiffness, while learnable positional embeddings preserve spatial context. (C) Transformer encoder and classification. In the MM–Transformer, image and numeric tokens are concatenated with positional encodings and processed through two transformer encoder blocks, each with multi-head attention (4 heads × 32 dimensions), a feed-forward network (128 hidden units, ReLU), residual connections, dropout, and layer normalization. The fused representation is flattened and passed through fully connected layers to softmax probabilities across three categories (response, stable disease, non-response). In the ViT16 variant, stacked transformer blocks refine the full patch–numeric token sequence, and a classification head on the numeric token generates final predictions. Created with BioRender.com.
3.3. Stiffness-only baseline confirms the added value of SWE image features for response prediction
To isolate the contribution of each modality, we performed an ablation study comparing a stiffness-only baseline against the full multimodal framework. A multinomial logistic regression model trained exclusively on elastic modulus (kPa) values achieved a test accuracy of 76.4% and a macro AUC of 0.900 (Response AUC = 0.953, Stable AUC = 0.828, Non-Response AUC = 0.918) (Fig. 4B). The model demonstrated strong discrimination for Response cases (precision = 0.84, recall = 0.90) and Non-Response cases (precision = 0.79, recall = 0.75), but performed notably worse on the Stable class (precision = 0.64, recall = 0.62), consistent with the boundary ambiguity inherent to this intermediate category (Fig. 4F). The decision boundary analysis revealed that the logistic regression learned kPa thresholds at approximately 30 and 45 kPa, below which tumors were predominantly classified as responsive and above which as non-responsive (Fig. 4D and E). While these results confirm that elastic modulus alone carries substantial predictive information, the full multimodal ViT16 achieved considerably higher accuracy of 92.4% (macro-F1 0.922, macro AUC 0.987); five seeds, held-out test set demonstrating that SWE image features capture spatial and textural patterns beyond what a single scalar stiffness measurement can encode.
Fig. 4.

Stiffness-only logistic regression ablation. A multinomial logistic regression was trained on elastic modulus (kPa) values alone to predict treatment response. (A) Confusion matrix on the held-out test set (n = 237); overall accuracy 76.4%. (B) One-vs-rest ROC curves per class, macro AUC = 0.900. Response showed the highest discriminability (AUC = 0.953), while Stable Disease was the most challenging (AUC = 0.828). (C) Distribution of prediction confidence scores across classes (mean 0.744). (D) Calibration curves of predicted probabilities against observed frequencies for each class. (E) Decision boundary showing predicted class probabilities as a function of elastic modulus; vertical dashed lines mark transition thresholds at approximately 30 and 45 kPa. (F) Radar plot of per-class precision, recall, and F1-score, highlighting the relative weakness in Stable Disease classification. This single-feature baseline establishes that tumor stiffness alone is informative but insufficient compared with the full multimodal framework. Values shown are from a single-fit reference model; the corresponding five-seed value used for comparison against the other ablation arms is 78.1% ± 1.9% accuracy (Section 3.4). Abbreviations: R, Response; SD, Stable Disease; NR, Non-response.
3.4. Architecture ablations confirm that image–stiffness fusion drives accuracy
To dissect how the multimodal ViT16 achieves its performance, we ablated the model along its two defining design choices — the fusion of the stiffness token and its integration by self-attention — training each variant on the identical stratified split across five random seeds (mean ± 95% CI reported). The full multimodal model, which prepends the standardized elastic modulus as a dedicated token to the image patches, achieved the strongest performance (accuracy = 0.924 ± 0.016, macro-F1 = 0.922 ± 0.016, macro AUC = 0.987 ± 0.003; expected calibration error, ECE = 0.011). Removing the stiffness token entirely (image-only) reduced accuracy to 0.850 ± 0.038 and macro AUC to 0.962 ± 0.015, a significant drop (Δaccuracy = 0.074, p = 0.0008; Δmacro-F1 = 0.075, p = 0.0006, paired t-test with Holm correction), confirming that quantitative stiffness adds information beyond image texture alone. Critically, when the image–stiffness pairing was broken by randomly shuffling kPa values across samples, performance fell to 0.841 ± 0.037 (Δaccuracy = 0.083, p = 0.0011) — no better than image-only — demonstrating that it is the correct biomechanical pairing, not merely the presence of a numeric input, that the model exploits. A late-fusion variant, which concatenates image and stiffness features only at the classifier rather than allowing them to interact through self-attention, likewise underperformed the full model (0.853 ± 0.031; Δaccuracy = 0.071, p = 0.0003), indicating that early, attention-based cross-modal conditioning is superior to post-hoc feature concatenation. All learned variants substantially exceeded the stiffness-only logistic baseline (0.781 ± 0.019; Δaccuracy = 0.143, p = 5 × 10−5). The full model was also the best calibrated (ECE = 0.011 vs 0.018–0.077 for the ablated arms). Together, these ablations establish that both the inclusion of the elastic modulus and its attention-based fusion with correctly paired imaging are necessary for the ViT16's predictive gains (Fig. 5).
Fig. 5.

Architecture ablation of the multimodal ViT16. Five model variants - Full (image + kPa), Image-only (stiffness token removed), Shuffled kPa (image–stiffness pairing randomly permuted), Late fusion (image and stiffness concatenated only at the classifier), and the Stiffness-only logistic-regression baseline - were each trained on the identical stratified split across five random seeds and evaluated on the common held-out test set (n = 237). (A) Normalized multi-metric overview (macro-F1, accuracy, calibration, and macro ROC–AUC) for all arms. (B) Per-arm accuracy distribution across the five seeds (violins; points = individual seeds, diamond = mean). (C) Macro-averaged one-vs-rest ROC curves (predictions pooled over seeds) with AUC per arm. (D) Accuracy drop of each ablation relative to the Full model, with Holm-corrected significance (paired t-test across seeds; **p < 0.01, ***p < 0.001). (E) Validation-loss convergence over training (mean ± SD across seeds). (F) Confusion matrices per arm (mean counts per test set, averaged over five seeds; color encodes row proportion). The Full model outperforms every ablation, and the collapse of the Shuffled-kPa arm to Image-only performance confirms that the model exploits the genuine image–stiffness relationship rather than the mere presence of a numeric feature. Abbreviations: R, Response; SD, Stable Disease; NR, Non-response.
3.5. Leave-one-tumor-out validation confirms cross-tumor generalization
To test whether the multimodal ViT16 generalizes to tumor types unseen during training — rather than memorizing model-specific features — we performed a leave-one-tumor-out (LOTO) analysis across the five syngeneic tumor models. In each fold, all images from one tumor model were withheld entirely from training and validation and used only for testing, with the model retrained from scratch (five random seeds per fold; mean ± SD). Generalization was strong and consistent across all held-out tumors: accuracy ranged from 0.939 ± 0.007 (K7M2 osteosarcoma) to 0.979 ± 0.010 (B16F10 melanoma), with macro AUC ≥ 0.993 in every fold (4T1 = 0.994 ± 0.001; E0771 = 0.993 ± 0.000; MCA205 = 0.997 ± 0.000; K7M2 = 0.993 ± 0.002; B16F10 = 0.999 ± 0.000). Averaged over all five held-out tumors (n = 1365 test images), the model achieved 0.955 ± 0.015 accuracy, 0.954 ± 0.015 macro-F1, and 0.995 ± 0.002 macro AUC - comparable to, and in several folds exceeding, its performance under the standard stratified split. That predictive accuracy is retained when an entire tumor histology is excluded from training indicates that the model learns a transferable biomechanical signature of therapy response rather than tumor-type-specific texture, supporting the generalizability of the SWE-based framework across cancer types (Fig. 6). We note that this analysis is performed on the 1365-image confirmed-annotation subset rather than the full 1578-image cohort (Methods 2.7). The higher accuracy observed here relative to the stratified split should therefore not be interpreted as an improvement over the main model, as the two are evaluated on different data and under different partitioning schemes.
Fig. 6.

Leave-one-tumor-model-out generalization of the multimodal ViT16. In each fold, all images from one syngeneic tumor model (4T1, E0771, MCA205, K7M2, or B16F10) were held out entirely and used only for testing, while training and validation used the remaining four models; each fold was retrained from scratch across five random seeds. (A) Accuracy (solid) and macro-F1 (hatched) on the unseen tumor per fold (mean ± 95% CI); the dashed line marks the mean accuracy across folds (0.955). (B) Per-seed accuracy distribution for each held-out tumor (violins; points = individual seeds, diamond = mean). (C) Macro-averaged one-vs-rest ROC curves per held-out tumor (predictions pooled over seeds), with AUC. (D) Data partition per fold — train/validation/test image counts (and % of dataset) with the corresponding test accuracy; the test set is the entire held-out cell line (1365 unseen images across all folds). (E) Confusion matrices on each unseen tumor (mean counts per test set, averaged over five seeds). Performance is preserved for every held-out tumor, demonstrating that the model generalizes to tumor types absent from training. Abbreviations: R, Response; SD, Stable Disease; NR, Non-response.
3.6. Understanding prediction dynamics in baseline, transformer, and ViT16 frameworks
We next compared the prediction behaviors of the baseline multimodal CNN, the Transformer-based hybrid, and the ViT16 architecture (Fig. 7). The baseline CNN displayed moderate discrimination, with ROC analysis showing strong but less decisive separability between classes (macro-AUC = 0.985) (Fig. 7B). Misclassifications were concentrated in the Stable group, reflecting the intermediate stiffness and heterogeneous features that blur the distinction between clear responders and non-responders. Moreover, confidence analyses revealed a broad spread, with many predictions clustered in the mid-confidence range, underscoring limited decisiveness (Fig. 7D). The Transformer model improved on this by reducing cross-class confusion, particularly for Stable tumors (Fig. 7E), and shifting the overall confidence distribution upward (mean confidence = 0.91), with sharper separation of classes (Fig. 7H). The ViT16 model further consolidated these gains, achieving the cleanest confusion matrix with minimal misclassification, the highest discriminative power (macro-AUC = 0.991) (Fig. 7J), and the most decisive predictions (mean confidence = 0.92). Importantly, confidence histograms confirmed that the ViT16 model not only produced highly accurate classifications but also well-calibrated and reliable probabilities across all three outcome categories (Fig. 7K). Incorporating attention mechanisms and transformer-based architectures improves both discriminative performance and predictive reliability, with ViT16 offering the strongest balance for clinical translation. Class-specific histograms showed that Non-response predictions were the most confidently assigned by the Transformer and ViT16 architectures, aligning with their improved confusion patterns.
Fig. 7.

Comparative prediction behaviors across multimodal architectures. (A – D) Baseline multimodal CNN. Misclassifications concentrate in the Stable group, while Response and Non-response are generally distinguished (A). ROC curves show strong but less decisive separability (macro-AUC = 0.985) (B). Confidence histograms reveal broad variability and many mid-confidence predictions (mean 0.76) (C). Class-specific distributions show less reliable discrimination, particularly between Stable and Non-response (D). (E – H) Transformer-based multimodal model. Cross-class errors are reduced, especially in the Stable category (E). ROC curves indicate improved discrimination (macro-AUC = 0.987) (F). Predictions are more polarized, with most cases assigned high probabilities (mean 0.91) (G). Class-wise distributions show clearer stratification, particularly in Non-response (H). (I – L) Vision Transformer (ViT16) multimodal model. The confusion matrix is cleanest, minimizing misclassifications across all categories (I). ROC curves confirm discriminative ability (macro-AUC = 0.991) (J). Predictions are strongly decisive (mean 0.92) with reduced variability (K). Class-specific distributions confirm high confidence across all categories, identifying ViT16 as the most reliable architecture (L). Values in this figure derive from the initial single-run comparison; seed-averaged benchmark metrics are reported in Section 3.8.
3.7. Reliability and calibration characteristics of the MM-Transformer and MM-ViT16 models
To probe the reliability and calibration of our Transformer-based models, we examined training dynamics, confidence–uncertainty relationships, and error stratification (Fig. 8). Learning curves revealed clear differences in stability: while the Transformer exhibited mild fluctuations across epochs (Fig. 8A and B), the ViT16 model converged more smoothly (Fig. 8J and K), with no evidence of overfitting, underscoring its robustness even on modest datasets. Radar plots summarizing per-class metrics confirmed balanced precision, recall, and F1-scores across Response, Stable, and Non-response tumors, consistent with biologically meaningful stratification of treatment outcomes (Fig. 8C and L). Confidence versus uncertainty analyses revealed the expected inverse relationship, with higher predicted probabilities corresponding to lower entropy (Fig. 8D and M). Crucially, misclassifications were disproportionately concentrated in the lowest confidence quartile, whereas high-confidence predictions were almost uniformly correct (Fig. 8E and N). This effect was especially pronounced in ViT16, which displayed a steeper decline in error with increasing confidence, reflecting clear “decisiveness”. Distribution analyses further underscored these trends: incorrect predictions clustered at low confidence and high entropy (Fig. 8F–G, Fig. 8O and P), while correct predictions formed narrow peaks at high confidence and low entropy. Boxplots confirmed consistently high confidence across all classes, with Non-response showing the tightest spread in both models—an important feature for clinical application, since failing to detect resistant tumors could have significant therapeutic consequences (Fig. 8I and R). Calibration plots demonstrated that both architectures produced well-aligned probability estimates, closely matching the ideal diagonal. The ViT16 model adhered most closely, particularly in intermediate confidence bins, suggesting that its probability outputs can be directly trusted for clinical decision-making (Fig. 8Q). The transformer-based models not only achieve high discriminative performance but also yield calibrated, uncertainty-aware predictions. Importantly, the ViT16 model provided the most stable training, decisive confidence distributions, and clinically interpretable reliability profiles—critical qualities for deploying AI to guide mechanotherapeutic treatment strategies.
Fig. 8.

Confidence and calibration analysis of transformer-based multimodal models. (A – I) MM-Transformer model. Training and validation curves show mild fluctuations during training (A, B). Radar plots confirm balanced precision, recall, and F1-scores across Response, Stable, and Non-response categories (C). Confidence–uncertainty plots show an inverse correlation, with most errors occurring at low confidence (D). Error rate stratification shows that high-confidence predictions are overwhelmingly correct (E). Confidence and entropy distributions reveal systematic separation between correct (high confidence, low entropy) and incorrect (low confidence, high entropy) predictions (F – G). The calibration curve aligns closely with the ideal diagonal, with minor deviations in mid-confidence bins (H). (J – R) Vision Transformer (ViT16) model. Training curves show smoother convergence and greater stability than the MM-Transformer (J – K). Radar plots again show balanced performance across all classes (L). Confidence–uncertainty plots and quartile error analysis confirm the most decisive behavior, with sharp error reduction as confidence increases (M − N). Distributions show narrow peaks of correct predictions at high confidence and low entropy (O – P). Calibration plots show near-perfect alignment with the ideal diagonal, confirming the reliability of predicted probabilities (Q). Class-wise boxplots show consistently high confidence across categories, with Non-response predictions most tightly constrained (R).
3.8. Comparative evaluation of six multimodal architectures reveals superior calibration and generalization in transformer-based models
To benchmark multimodal learning performance, we compared six architectures differing in their integration of convolutional, attention-based, and transformer components (Fig. 9A–F). All models demonstrated strong discriminative capacity, with macro-AUC values exceeding 0.97 (Fig. 9D). ViT16 achieved the highest performance (macro-AUC = 0.988 ± 0.002), followed by the MM-Transformer (macro-AUC = 0.982), underscoring the advantage of global attention mechanisms in modeling multimodal dependencies. Cross-validation analyses confirmed consistent gains in transformer-based designs: mean validation accuracy increased from 0.877 in the baseline CNN to 0.927 in ViT16, while validation loss decreased from 0.498 to 0.201 (Fig. 9Aand B). Test set evaluation yielded concordant trends – ViT16 attained the lowest test loss (0.189 ± 0.024) and highest accuracy (0.925 ± 0.021), outperforming both the MM-Transformer (0.917 ± 0.024) and the MM-attention (0.841 ± 0.029) (Fig. 9C). Calibration analyses revealed that all architectures assigned higher confidence to correct than incorrect predictions, but this separation was most pronounced in transformer-based models (Fig. 9E). Confidence–accuracy plots further indicated that CNN-based models underestimated true accuracy, whereas transformer-based networks approached the line of perfect calibration, with ViT16 exhibiting the most reliable probability estimates (Fig. 9F). Incorporating attention mechanisms improves both discrimination and uncertainty estimation, with fully transformer-based architectures—particularly ViT16—delivering the most generalizable and clinically reliable predictions. Pairwise DeLong tests confirmed that these differences were statistically significant where they mattered most: ViT16's macro-AUC significantly exceeded the attention-augmented CNN (ΔAUC = +0.026, Holm-adjusted p < 0.001) and the baseline CNN (p < 0.001), whereas the three transformer-based models (ViT16, MM-Transformer, MM-Transformer-Attention) were statistically indistinguishable from one another (p > 0.05), indicating a comparable top-performing tier from which ViT16 emerges as the most parameter-efficient (Fig. 9G).
Fig. 9.

Comparative benchmarking of multimodal architectures. (A) Cross-validation accuracy (mean ± s.d.) across folds. Transformer-based architectures achieved consistently higher validation performance, with ViT16 reaching the highest mean accuracy (0.927). (B) Cross-validation loss (mean ± s.d.), showing improved optimization stability in transformer-based models and the lowest loss in ViT16 (0.201). (C) Held-out test performance (accuracy and loss per model). ViT16 achieved the best generalization (accuracy = 0.925 ± 0.021; loss = 0.189 ± 0.024), followed by the MM-Transformer (accuracy = 0.917 ± 0.024). (D) Receiver operating characteristic (ROC) curves (macro-average). All architectures showed strong discriminative ability (macro-AUC > 0.97), with ViT16 performing best (AUC = 0.988 ± 0.002). (E) Mean predicted confidence for correct (green) and incorrect (red) predictions, showing greater separation in attention- and transformer-based architectures. (F) Confidence–accuracy plots comparing predicted probabilities with empirical accuracy; ViT16 lies closest to the diagonal, indicating near-ideal calibration and reliable probability estimates. (G) Pairwise DeLong comparison of macro-averaged ROC–AUC between all six architectures. Cell color encodes the AUC difference (row − column model; red = row higher, blue = column higher, grey diagonal = self-comparison), and stars denote Holm-adjusted significance (*p < 0.05, **p < 0.01, ***p < 0.001).
3.9. Transformer-based architectures outperform simple CNNs
To comprehensively assess model behavior, we compared six multimodal architectures across class-specific, global, and efficiency-adjusted performance dimensions (Fig. 10A–E). At the class level, all models maintained balanced precision, recall, and F1-scores across the three outcome categories—Response, Stable, and Non-Response (Fig. 10A). Transformer-based architectures were more stable in recall for the Stable class, the most error-prone category, while ViT16 achieved the most uniform performance across all metrics. In the AUC–F1 landscape (Fig. 10B), ViT16 occupied the top-right quadrant, combining high discriminative power with robust balance. The attention-based CNN reached comparable F1 but lower AUC, while the multimodal transformer sat at an intermediate balance. Notably, models with extensive parameter counts—MM-transformer_attention and MM-attention—did not yield commensurate gains, suggesting diminishing returns with increasing complexity. Heatmap analyses integrated efficiency and reliability (Fig. 10C–E). Parameter–performance mapping (Fig. 10C) showed ViT16 achieving the best trade-off, combining top-tier accuracy with a compact parameter footprint (∼0.32 M), versus attention-based models exceeding 15 M parameters with no substantive gain. The aggregated performance matrix (Fig. 10D) showed ViT16 ranking highest or near-highest across accuracy, mean confidence, and AUC. A final composite ranking (Fig. 10E) placed ViT16 first, followed by the MM-transformer and MM-attention. While all models achieved strong discrimination (macro-AUC > 0.97), attention-based mechanisms systematically improved calibration and stability. Together, these results identify ViT16 as the most efficient and generalizable architecture for multimodal SWE-based response prediction, balancing predictive strength, probabilistic reliability, and computational economy.
Fig. 10.

Integrated benchmarking of multimodal architectures for treatment response prediction. (A) Class-level radar plots of precision, recall, and F1-score across the three outcome categories (Response, Stable, Non-Response). Transformer-based models achieved more consistent recall in the Stable class, and ViT16 showed the most balanced performance across all metrics. (B) AUC–F1 scatter plot (bubble size proportional to parameter count). ViT16 occupies the top-right quadrant, combining high discriminative power with balanced classification and moderate complexity. Larger architectures, such as the transformer–attention hybrid and attention-only CNN, carried substantial parameter loads (>15 M) without proportional performance gains. (C) Composite heatmaps of AUC, F1-score, and parameter counts, summarizing absolute performance across models. ViT16 achieved the best trade-off between accuracy and efficiency, while CNN-only baselines lagged. (D) Aggregated performance matrix integrating accuracy, mean confidence, and AUC. Transformer-based architectures consistently outperformed CNN variants, with ViT16 showing the highest overall stability and reliability. (E) Relative performance ranking heatmap (1 = best, 6 = worst) consolidating all metrics. ViT16 ranked first overall, followed by the multimodal transformer and attention-enhanced CNNs, highlighting the progressive benefit of incorporating attention into multimodal learning. Panel B is reproduced from our initial single-run benchmark; the corresponding three-seed values are reported in Section 3.8, and the ranking of architectures is identical under both analyses.
4. Discussion
4.1. Integrating biomechanical and imaging biomarkers
The findings of this study highlight the value of combining quantitative tissue stiffness with imaging features for therapy response prediction. In our preclinical models, baseline tumor stiffness (elastic modulus) emerged as a strong biomarker: tumors that ultimately responded to treatment were significantly softer at baseline than those with stable or progressive disease [12]. This finding was valid across several tumor types, which aligns with the mechanistic understanding that higher mechanical stress and matrix rigidity in tumors can impair drug delivery and foster resistance [23,39,40]. While the scalar stiffness value alone stratified outcomes to an extent (with responders averaging lower kPa than non-responders), the multimodal approach substantially improved predictive power. By fusing the shear-wave elastography (SWE) images with the corresponding stiffness measurements, our model captured both global and local mechanical cues that a single-modality model might miss. Deep learning analysis of the full elastography map allows the model to “scan” the spatial distribution of stiffness across the entire tumor, rather than relying on an average value [23]. Indeed, intratumoral heterogeneity in stiffness contains critical information – for example, pockets of high stiffness may indicate hypoxic, therapy-resistant regions, whereas uniformly softer tumors might be more perfused and drug-accessible. The superior performance of our multimodal model over a single-input alternative underscores that these imaging and numeric features provide complementary information. Integrating the two enabled more nuanced discrimination between response categories, reducing confusion especially between partial responders and non-responders. This result supports a broader conclusion that mechanopathological markers can enrich outcome prediction beyond what traditional anatomical or genomic markers offer. By leveraging a mechanical biomarker (elastic modulus) alongside imaging, we effectively created a composite predictor of how well a tumor's microenvironment can support therapy efficacy. In practical terms, this means we can potentially identify, before treatment begins, which tumors are predisposed to respond to chemo-immunotherapy and which are likely to resist. Early stratification could guide treatment decisions – for instance, patients with very stiff tumors (predictive of poor response) might be triaged for adjunct mechanomodulatory therapy or alternative regimens upfront.
4.2. Advances through a transformer-based architecture
To our knowledge, this is the first transformer-based architecture applied to shear-wave elastography for the prediction of cancer therapy response. Unlike conventional multimodal pipelines that combine imaging and numeric features by late concatenation, our framework introduces the quantitative elastic modulus as a dedicated token within the transformer's self-attention, enabling direct, learnable cross-modal interaction between spatial stiffness texture and the global stiffness magnitude. Our ablation experiments confirm that this token-level fusion — not the scalar feature alone — drives the performance gain, since randomly breaking the image–stiffness pairing collapses accuracy to image-only levels. Concretely, the ViT16 model delivered a marked performance boost over the MModal-Baseline, achieving ∼92% accuracy with an AUC ≈0.99 in classifying responders vs. stable vs. non-responders, compared to about 87% accuracy and AUC 0.96 with our previous best CNN-attention model [23,41]. This improvement is especially significant given that earlier work had already approached human-expert performance with deep CNNs [23]. The transformer's self-attention mechanism appears to capture subtle patterns that CNNs with or without soft attention only partially addressed. By treating the CNN-extracted image features and the stiffness value as a set of “tokens,” the model can attend to how global stiffness context interacts with specific image regions. For example, the network can learn whether small stiff areas in an otherwise soft tumor carry more prognostic weight than uniformly intermediate stiffness – a level of conditional reasoning that a simple late-fusion of features might miss. We also observed that the ViT16 architecture produced better-calibrated confidence estimates and more decisive predictions [23]. The attention-based models were particularly adept at confidently identifying the truly non-responsive tumors, a critical class to get right so that ineffective therapies can be promptly adjusted. Interestingly, the best-performing ViT16 was a relatively lightweight model (∼0.3 million parameters), far smaller than some MM-attention networks we tested (>15 million parameters). Its strong performance despite the compact size suggests that architectural efficiency trumped brute force complexity. In data-constrained domains like ours, this is an encouraging result – it illustrates that incorporating domain-specific design (e.g. including a stiffness token, using appropriate patch sizes or convolutional embedding) can enable transformers to excel without requiring exorbitant data or pre-training. In fact, our results resonate with emerging evidence that transformers can equal or surpass CNNs on modest medical datasets when carefully optimized. Beyond raw performance metrics, the transformer's multi-head attention offers a degree of interpretability; by inspecting attention maps, we could confirm that the model learns to focus on meaningful regions (e.g. high-stiffness foci or tumor periphery) and cross-reference them with the global stiffness level – essentially learning the same “high-and-low stiffness” complementary focus that our prior attention CNN model demonstrated [23]. This ability to capture both localized texture and holistic context is pushing the envelope of multimodal AI in oncology. Deploying a transformer in our setting advanced the state-of-the-art for SWE-based prognosis, and it showcases how self-attention can unlock complex cross-modal interactions that improve predictive modeling of cancer outcomes.
4.3. Mechanotherapeutic implications for tumor response
Our study has direct implications for mechanotherapies – treatments that modify the tumor's physical microenvironment. All tumors in our experiments were pre-treated with mechanotherapeutic agents, including the anti-fibrotic drug tranilast, the endothelin receptor antagonist bosentan, or the antihistamine ketotifen, prior to standard chemo-immunotherapy. This combinatorial strategy was intended to “prime” tumors by reducing stiffness and mechanical stress, thereby improving therapeutic efficacy. The prognostic significance of stiffness we observed strongly reinforces the biological rationale for this approach. We found that even with mechanotherapeutic pre-treatment, the residual stiffness of the tumor at therapy onset was predictive of outcome: tumors that remained mechanically rigid tended not to respond well, whereas those that had softened were more likely to regress. This aligns with prior reports that alleviating mechanical stress in tumors can markedly improve treatment efficacy by decompressing blood vessels and enhancing drug delivery. For example, Papageorgis et al. showed that tranilast-induced stress relaxation in breast tumors boosted the effect of chemotherapy and nanotherapy [23]. Our results extend this concept by suggesting that measuring the post-mechanotherapy stiffness provides a quantitative gauge of how well the microenvironment has been normalized. In cases where a tumor remains highly stiff (perhaps due to dense collagen or incomplete response to the mechanotherapeutic), our model would flag a likely treatment failure. Clinically, this could prompt consideration of alternative strategies – such as intensifying the mechanotherapeutic regimen, switching therapies, or enrolling the patient in trials for drugs targeting the extracellular matrix. Conversely, identification of a tumor as mechanically “responsive” (i.e. significantly softened) could justify continuing with the planned chemo-immunotherapy regimen with greater confidence. It is also notable that our predictive framework performed consistently across different treatment modalities (chemotherapy alone, immunotherapy alone, and the combination) in our dataset, correctly identifying responders in each group at similar rates. This suggests that tumor stiffness is a generalizable biomarker of therapeutic vulnerability, not tied to one specific drug mechanism or tumor type [39]. Biologically, a less fibrotic, better-perfused tumor microenvironment should aid many forms of therapy – from small-molecule drugs to immune cell infiltration – so it makes sense that elastic modulus correlates with outcome across diverse treatments [41]. Taken together, these findings highlight tumor mechanopathology as a key piece of the precision oncology puzzle: they underscore that how a tumor feels (its physical properties) can be as prognostically important as its genomic or proteomic profile. By embracing this principle, future therapies could combine molecular targeting with stroma-normalizing treatments to tackle both the seeds and the soil of cancer. Our work provides a quantitative foundation for that approach, supporting the notion that measuring and modulating tumor mechanics can significantly impact patient outcomes.
4.4. Translational outlook and clinical relevance
An exciting aspect of using SWE-based biomarkers is the immediate path to clinical translation. Ultrasound shear wave elastography is already FDA-approved and widely used for diagnostic applications in breast, liver, thyroid, and other tissues. This means the leap from our preclinical findings to testing in humans is far smaller than if a novel imaging modality were required. Indeed, early clinical studies have hinted at the prognostic value of elastography – for instance, in breast cancer patients receiving neoadjuvant chemotherapy, tumors that ultimately achieve pathological complete response tend to exhibit a greater drop in stiffness after just a couple of treatment cycles [42]. Similarly, baseline SWE measurements in patients have been correlated with eventual therapy outcomes, reinforcing that mechanical phenotype is relevant in the clinical setting [43,44]. Building on these insights, our model could be deployed during routine ultrasound exams to provide an on-the-spot risk stratification. For example, when a breast or sarcoma tumor is first imaged, one could acquire an elastography scan and input it to our AI system, which would then output a prediction. This information could support oncologists in crafting a personalized treatment plan before starting therapy – potentially sparing patients with predicted non-responsive tumors from months of ineffective therapy and toxicity, and directing them sooner to alternate interventions or clinical trials. We are actively working to validate this approach in the clinical arena. Ongoing collaborative trials at our institutions are enrolling sarcoma and breast cancer patients to test whether SWE-based transformer predictions align with actual treatment responses. These studies will also address practical considerations for clinical deployment, such as inter-observer consistency in acquiring elastograms and the integration of this tool into clinical workflow. A challenge we anticipate is the variability in stiffness ranges between murine models and humans: for example, murine breast tumors in our study averaged ∼30–50 kPa in stiffness, whereas human breast cancers can reach 100+ kPa. Despite these differences, the underlying biology of how fibrosis and pressure affect therapy is conserved across species. Early results from our patient scans indicate a broad distribution of stiffness values across tumor types and even within a given cancer type, underscoring the need for individualized analysis – precisely what our model is designed to do. Ultimately, successful clinical translation of this work could make SWE-based response prediction an inexpensive, non-invasive adjunct to precision medicine.
4.5. Limitations and future directions
While our preclinical results are encouraging, several challenges remain for clinical translation. First, training labels were defined using the RECIST 1.1 criterion; although standard and reproducible, RECIST is a size-based anatomical measure that does not capture functional, metabolic, or survival endpoints, and future work should incorporate volumetric, perfusion-based, or survival outcomes. Importantly, all imaging was performed at therapy onset — after mechanotherapeutic priming but before response was known — so the model uses the residual post-priming stiffness to predict the subsequently RECIST-defined outcome; longitudinal stiffness change was not used as a label, and treatment-induced extracellular-matrix softening is therefore part of the predictive response phenotype rather than a confounding factor. Second, shear-wave elastography is subject to necrosis-related artifacts: coagulative necrosis can elevate apparent stiffness, whereas liquefactive necrosis prevents shear-wave propagation and yields invalid kPa measurements. Baseline imaging of small tumors (∼350 mm3) and enforcement of the scanner's shear-wave confidence map within the region of interest minimize these effects in the present cohort, but in advanced or heavily necrotic tumors they could confound stiffness interpretation, and pairing SWE with B-mode or necrosis-aware assessment is a valuable extension. Third, human tumors vary more in depth, size, and viscoelastic properties than subcutaneous murine models, and clinical SWE acquisitions are subject to motion artifacts from breathing and cardiac pulsation — factors largely absent in anesthetized mice — while deeper tumors yield lower signal-to-noise ratios as shear waves attenuate through intervening tissue. Importantly, the same ultrasound platform (Philips EPIQ Elite) and acquisition protocol used in our preclinical experiments are directly applicable to human imaging, generating elastograms on an identical color-mapped stiffness scale. This hardware consistency means image-level features — spatial stiffness patterns, texture heterogeneity, and color distributions — are encoded in the same representational space across species. We therefore envision transfer learning as the most practical deployment path: pretrained weights would serve as initialization, followed by fine-tuning on a smaller human cohort to recalibrate stiffness thresholds and adapt to clinical variability, with domain adaptation bridging the gap if systematic differences in image statistics emerge. On the technical side, we distilled each tumor's mechanics into a single mean elastic modulus, which captures overall stiffness but loses distributional subtleties; although the SWE image input implicitly encodes spatial variation, explicitly adding summary statistics or multi-channel elastography inputs could further improve performance. The model also focuses on a single therapeutic context, and different therapies may interact with tumor mechanics in distinct ways, warranting evaluation across additional treatment scenarios; an ambitious extension would be a model that recommends the therapy most likely to succeed for a given tumor profile. Finally, improving interpretability — by highlighting image regions driving predictions or quantifying the stiffness token's influence — will be important for clinician trust. Taken together, and given that all experiments were conducted in preclinical murine models, the present findings should be regarded as proof-of-concept; clinical readiness will require prospective validation in independent patient cohorts.
4.6. Conclusions
This study demonstrates that tumor mechanical properties — captured jointly through shear wave elastography (SWE) images and scalar elastic modulus values — serve as powerful, non-invasive predictors of treatment response in preclinical cancer models. By developing a multimodal deep learning framework that integrates spatial elastographic texture with quantitative tissue stiffness within a transformer architecture, we achieved 92.4% classification accuracy (macro-F1 0.922, macro AUC 0.987; five seeds, held-out test set) in distinguishing responders, stable disease, and non-responders across five syngeneic tumor models, and 95.5% accuracy on tumor models entirely absent from training in a leave-one-tumor-model-out analysis. These results represent a substantial improvement over both single-modality baselines (78.1% accuracy with stiffness alone, under the same five-seed protocol) and conventional convolutional approaches, establishing that the combination of local image heterogeneity and global mechanical context provides complementary predictive information that neither modality captures independently.
A principal strength of this work lies in the architectural design of the multimodal fusion. Unlike late-fusion CNN approaches, the transformer's self-attention mechanism enables direct interaction between image-derived tokens and the numeric stiffness embedding, allowing the model to learn conditional relationships — for example, whether focal regions of high stiffness within an otherwise compliant tumor carry different prognostic weight than uniformly intermediate stiffness. The compact Vision Transformer (ViT16), with ∼0.32 million parameters, achieved the strongest performance among six evaluated architectures, outperforming models exceeding 15 million parameters. This favorable accuracy-to-complexity ratio is particularly relevant for data-constrained biomedical domains, where parameter-efficient models are less prone to overfitting and more amenable to deployment on resource-limited hardware. Additionally, the ViT16 exhibited well-calibrated probability estimates, with predicted confidence closely aligning with empirical accuracy — a critical property for translating model outputs into actionable clinical decisions. From a pattern recognition perspective, this work contributes to the growing evidence that transformer-based multimodal fusion can outperform established CNN architectures when heterogeneous data types — images and scalar features — must be jointly modeled. The systematic comparison of six architectures, evaluated not only by discriminative metrics but also by calibration, confidence reliability, and parameter efficiency, provides a benchmarking framework that may inform future multimodal model selection in medical imaging tasks beyond oncology.
Several limitations must be acknowledged. All experiments were conducted in murine tumor models; while biologically relevant, the direct applicability of learned stiffness thresholds and texture features to human tumors remains to be validated, given the greater heterogeneity in depth, size, viscoelastic properties, and clinical acquisition variability present in patient imaging. The framework was trained within a specific therapeutic context — chemo-immunotherapy following mechanotherapeutic priming — and its generalizability to other treatment modalities has not been assessed. The Stable Disease category proved most challenging for all models, reflecting genuine ambiguity at classification boundaries. Finally, more systematic explainability analyses would strengthen the case for clinical adoption.
Looking ahead, clinical validation through prospective trials in sarcoma and breast cancer patients is underway at our institutions, employing the same Philips EPIQ Elite SWE platform, which ensures representational consistency across species and enables a transfer learning strategy — initializing with preclinical weights and fine-tuning on human cohorts. Expanding the feature set to include additional biomechanical descriptors, perfusion indices, or molecular biomarkers within a multi-omic framework represents a logical extension, leveraging the transformer's inherent capacity for flexible token-based data fusion. Developing uncertainty-aware prediction pipelines and clinically interpretable attention visualizations will be essential for bedside deployment. Ultimately, this work lays a foundation for integrating biomechanical imaging into the precision oncology toolkit alongside genomic and molecular profiling, ensuring that treatment is optimized not only for a tumor's biological identity but also for its physical microenvironment.
Data and Code Availability: The datasets generated and/or analyzed during the current study were obtained from previous research conducted by our laboratory [23]. The complete code used for model development, training, and evaluation is available on GitHub.
CRediT authorship contribution statement
Styliana Georgiou: Data curation, Formal analysis, Methodology, Validation, Visualization, Writing – original draft, Writing – review & editing. Triantafyllos Stylianopoulos: Funding acquisition, Supervision, Writing – original draft, Writing – review & editing. Chrysovalantis Voutouri: Funding acquisition, Investigation, Methodology, Supervision, Visualization, Writing – original draft, Writing – review & editing.
Ethics statement
Mice were purchased from the Cyprus Institute of Neurology and Genetics and all experiments were conducted in accordance with the animal welfare regulations and guidelines of the European Union (European Directive 2010/63/EU and Cyprus Legislation for the protection and welfare of animals, Laws 1994–2013) under a license acquired and approved (CY/EXP/PR.L01/2024) by the Cyprus Veterinary Services committee, the Cyprus national authority for monitoring the welfare of animals in research. The mice were housed at the Transgenic Mouse Facility (TMF) of the Cyprus Institute of Neurology and Genetics. Animals were kept in specific pathogen free conditions, according to regulations contained in the Cyprus Law N.55 (I)/2013, which is fully harmonized to the EU Directive 2010/63/EU. All the bedding and water for the mice was sterilized by autoclaving. Special care was taken to minimize potential pain, suffering or distress, and enhance animal welfare for the animals still used. Mice were anesthetized during tumor implantation with Avertin (250 mg/kg), and placed on a heating pad to maintain body temperature at 37 °C. At the endpoint of the study, animals were euthanized via CO2 asphyxiation followed by cervical dislocation. The authors have adhered to the ARRIVE guidelines.
Funding
This work is part of the Research and Innovation Foundation of Cyprus through the projects PROGNOSTIC-CODEVELOP-AG-SH-HE/0823/0202 and OncoPredict-POST-DOC/0524/0068. The project is implemented under the programme of social cohesion “THALIA 2021- 2027” co-funded by the European Union, through Research and Innovation Foundation. Ιt has also received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No. 101141357). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Author contributions: S. Georgiou: Data Curation, Formal Analysis, Investigation, Methodology, Resources, Software, Validation, Writing – original draft, Writing – review and editing. T. Stylianopoulos: Conceptualization, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Writing – review and editing. C. Voutouri: Conceptualization, Data Curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Writing – original draft, Writing – review and editing. Competing interests: All authors declare no financial or non-financial competing interests.
Declaration of competing interests
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.array.2026.101168.
Appendix A. Supplementary data
The following is the Supplementary data to this article.
Data availability
Data will be made available on request.
References
- 1.Jacquemin V., Antoine M., Dom G., Detours V., Maenhaut C., Dumont J.E. Dynamic cancer cell heterogeneity: diagnostic and therapeutic implications. Cancers (Basel) 2022;14 doi: 10.3390/cancers14020280. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Dagogo-Jack I., Shaw A.T. Tumour heterogeneity and resistance to cancer therapies. Nat Rev Clin Oncol. 2018;15:81–94. doi: 10.1038/nrclinonc.2017.166. [DOI] [PubMed] [Google Scholar]
- 3.Angeli S., Neophytou C., Kalli M., Stylianopoulos T., Mpekris F. The mechanopathology of the tumor microenvironment: detection techniques, molecular mechanisms and therapeutic opportunities. Front Cell Dev Biol. 2025;13:2025. doi: 10.3389/fcell.2025.1564626. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Kalli M., Stylianopoulos T. Toward innovative approaches for exploring the mechanically regulated tumor-immune microenvironment. APL Bioeng. 2024;8 doi: 10.1063/5.0183302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Stylianopoulos T., Munn L.L., Jain R.K. Reengineering the physical microenvironment of tumors to improve drug delivery and efficacy: from mathematical modeling to bench to bedside. Trends Cancer. 2018;4:292–319. doi: 10.1016/j.trecan.2018.02.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Matuszewska K., Pereira M., Petrik D., Lawler J., Petrik J. Normalizing tumor vasculature to reduce hypoxia, enhance perfusion, and optimize therapy uptake. Cancers (Basel) 2021;13 doi: 10.3390/cancers13174444. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Mpekris F., Panagi M., Charalambous A., Voutouri C., Stylianopoulos T. Modulating cancer mechanopathology to restore vascular function and enhance immunotherapy. Cell Rep Med. 2024;5 doi: 10.1016/j.xcrm.2024.101626. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Jain R.K., Martin J.D., Stylianopoulos T. The role of mechanical forces in tumor growth and therapy. Annu Rev Biomed Eng. 2014;16:321–346. doi: 10.1146/annurev-bioeng-071813-105259. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Byrnes J.R., Weeks A.M., Shifrut E., Carnevale J., Kirkemo L., Ashworth A., Marson A., Wells J.A. Hypoxia is a dominant remodeler of the effector T cell surface proteome relative to activation and regulatory T cell suppression. Mol Cell Proteomics. 2022;21 doi: 10.1016/j.mcpro.2022.100217. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Feng X., Cao F., Wu X., Xie W., Wang P., Jiang H. Targeting extracellular matrix stiffness for cancer therapy. Front Immunol. 2024;15:2024. doi: 10.3389/fimmu.2024.1467602. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Xu Q., Torres J.E., Hakim M., Babiak P.M., Pal P., Battistoni C.M., Nguyen M., Panitch A., Solorio L., Liu J.C. Collagen- and hyaluronic acid-based hydrogels and their biomedical applications. Mater Sci Eng R Rep. 2021;146 doi: 10.1016/j.mser.2021.100641. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Kontomaris S.V., Malamou A., Stylianou A. The young's modulus as a mechanical biomarker in AFM experiments: a tool for cancer diagnosis and treatment monitoring. Sensors. 2025;25 doi: 10.3390/s25113510. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Voutouri C., Mpekris F., Panagi M., Krolak C., Michael C., Martin J.D., Averkiou M.A., Stylianopoulos T. Ultrasound stiffness and perfusion markers correlate with tumor volume responses to immunotherapy. Acta Biomater. 2023;167:121–134. doi: 10.1016/j.actbio.2023.06.007. [DOI] [PubMed] [Google Scholar]
- 14.Zhou Y., Tao L., Qiu J., Xu J., Yang X., Zhang Y., Tian X., Guan X., Cen X., Zhao Y. Tumor biomarkers for diagnosis, prognosis and targeted therapy. Signal Transduct Targeted Ther. 2024;9:132. doi: 10.1038/s41392-024-01823-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Vargas A.J., Harris C.C. Biomarker development in the precision medicine era: lung cancer as a case study. Nat Rev Cancer. 2016;16:525–537. doi: 10.1038/nrc.2016.56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Kalli M., Mpekris F., Charalambous A., Michael C., Stylianou C., Voutouri C., Hadjigeorgiou A.G., Papoui A., Martin J.D., Stylianopoulos T. Mechanical forces inducing oxaliplatin resistance in pancreatic cancer can be targeted by autophagy inhibition. Commun Biol. 2024;7:1581. doi: 10.1038/s42003-024-07268-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Neophytou C., Stylianopoulos T., Mpekris F. The synergistic potential of mechanotherapy and sonopermeation to enhance cancer treatment effectiveness. npj Biological Physics and Mechanics. 2025;2:13. doi: 10.1038/s44341-025-00017-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Koutsi M., Stylianopoulos T., Mpekris F. Optimizing therapeutic outcomes with mechanotherapy and ultrasound sonopermeation in solid tumors. PLoS Comput Biol. 2025;21 doi: 10.1371/journal.pcbi.1012676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Englezos D., Voutouri C., Stylianopoulos T. Machine learning analysis reveals tumor stiffness and hypoperfusion as biomarkers predictive of cancer treatment efficacy. Transl Oncol. 2024;44 doi: 10.1016/j.tranon.2024.101944. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Marturano-Kruik A., Villasante A., Yaeger K., Ambati S.R., Chramiec A., Raimondi M.T., Vunjak-Novakovic G. Biomechanical regulation of drug sensitivity in an engineered model of human tumor. Biomaterials. 2018;150:150–161. doi: 10.1016/j.biomaterials.2017.10.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Chen Z., Chambara N., Lo X., Liu S.Y.W., Gunda S.T., Han X., Ying M.T.C. Improving the diagnostic strategy for thyroid nodules: a combination of artificial intelligence-based computer-aided diagnosis system and shear wave elastography. Endocrine. 2025;87:744–757. doi: 10.1007/s12020-024-04053-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wang G., Ouyang Y., Ruan L. Clinical value of quantitative analysis of Sonazoid-contrast enhanced ultrasound combined with shear wave elastography in discriminating and diagnosing breast tumor characteristics. Front Oncol. 2025;15:2025. doi: 10.3389/fonc.2025.1485671. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Voutouri C., Englezos D., Zamboglou C., Strouthos I., Papanastasiou G., Stylianopoulos T. A convolutional attention model for predicting response to chemo-immunotherapy from ultrasound elastography in mouse tumor models. Commun Med. 2024;4:203. doi: 10.1038/s43856-024-00634-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Evans A., Armstrong S., Whelehan P., Thomson K., Rauchhaus P., Purdie C., Jordan L., Jones L., Thompson A., Vinnicombe S. Can shear-wave elastography predict response to neoadjuvant chemotherapy in women with invasive breast cancer? Br J Cancer. 2013;109:2798–2802. doi: 10.1038/bjc.2013.660. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser Ł., Polosukhin I. Proceedings of the 31st international conference on neural information processing systems. Curran Associates Inc.; Long Beach, California, USA: 2017. Attention is all you need; pp. 6000–6010. [Google Scholar]
- 26.Madan S., Lentzen M., Brandt J., Rueckert D., Hofmann-Apitius M., Fröhlich H. Transformer models in biomedicine. BMC Med Inf Decis Making. 2024;24:214. doi: 10.1186/s12911-024-02600-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Zhang J., He K., Feng Z., Sun S., Sun X., Yuan Z., Yu J. Self-adaptive image-text fusion for medical image classification. Pattern Recogn. 2025;167 [Google Scholar]
- 28.Dai Y., Gao Y., Liu F. TransMed: transformers advance multi-modal medical image classification. Diagnostics. 2021;11 doi: 10.3390/diagnostics11081384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Wang X., Fang L., Zhao J., Pan Z., Li H., Li Y. MMAE: a universal image fusion method via mask attention mechanism. Pattern Recogn. 2025;158 [Google Scholar]
- 30.Eisenhauer E.A., Therasse P., Bogaerts J., Schwartz L.H., Sargent D., Ford R., Dancey J., Arbuck S., Gwyther S., Mooney M., Rubinstein L., Shankar L., Dodd L., Kaplan R., Lacombe D., Verweij J. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1) Eur J Cancer. 2009;45:228–247. doi: 10.1016/j.ejca.2008.10.026. [DOI] [PubMed] [Google Scholar]
- 31.Kolluri J., Kotte V.K., Phridviraj M.S.B., Razia S. 2020 4th international conference on trends in electronics and informatics (ICOEI)(48184) 2020. Reducing overfitting problem in machine learning using novel L1/4 regularization method; pp. 934–938. [Google Scholar]
- 32.Zhao X., Wang L., Zhang Y., Han X., Deveci M., Parmar M. A review of convolutional neural networks in computer vision. Artif Intell Rev. 2024;57:99. [Google Scholar]
- 33.Islam N., Hasib K.M., Mridha M.F., Alfarhood S., Safran M., Bhuyan M.K. Fusing global context with multiscale context for enhanced breast cancer classification. Sci Rep. 2024;14 doi: 10.1038/s41598-024-78363-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Zhang Y., Wang J., Gorriz J.M., Wang S. Deep learning and vision transformer for medical image analysis. Journal of Imaging. 2023;9:147. doi: 10.3390/jimaging9070147. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Azad R., Kazerouni A., Heidari M., Aghdam E.K., Molaei A., Jia Y., Jose A., Roy R., Merhof D. Advances in medical image analysis with vision transformers: a comprehensive review. Med Image Anal. 2024;91 doi: 10.1016/j.media.2023.103000. [DOI] [PubMed] [Google Scholar]
- 36.Aburass S., Dorgham O., Al Shaqsi J., Abu Rumman M., Al-Kadi O. Vision transformers in medical imaging: a comprehensive review of advancements and applications across multiple diseases. J Imaging Inform Med. 2025;38:3928–3971. doi: 10.1007/s10278-025-01481-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Ku B., Liu J., Na S., Yokota H., Hong C., Lim H. Effect of elastic modulus of tumour and non-tumour cells on vibration-induced behaviours. Sci Rep. 2025;15 doi: 10.1038/s41598-025-97837-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Voutouri C., Stylianopoulos T. Accumulation of mechanical forces in tumors is related to hyaluronan content and tissue stiffness. PLoS One. 2018;13 doi: 10.1371/journal.pone.0193801. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Ali M., Benfante V., Basirinia G., Alongi P., Sperandeo A., Quattrocchi A., Giannone A.G., Cabibi D., Yezzi A., Di Raimondo D., Tuttolomondo A., Comelli A. Applications of artificial intelligence, deep learning, and machine learning to support the analysis of microscopic images of cells and tissues. J Imaging. 2025;11 doi: 10.3390/jimaging11020059. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Neophytou C., Charalambous A., Voutouri C., Angeli S., Panagi M., Stylianopoulos T., Mpekris F. Sonopermeation combined with stroma normalization enables complete cure using nano-immunotherapy in murine breast tumors. J Contr Release. 2025;382 doi: 10.1016/j.jconrel.2025.113722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., Unterthiner T., Dehghani M., Minderer M., Heigold G., Gelly S., Uszkoreit J., Houlsby N. 2021. An image is worth 16x16 words: transformers for image recognition at scale. [Google Scholar]
- 42.Papageorgis P., Polydorou C., Mpekris F., Voutouri C., Agathokleous E., Kapnissi-Christodoulou C.P., Stylianopoulos T. Tranilast-induced stress alleviation in solid tumors improves the efficacy of chemo- and nanotherapeutics in a size-independent manner. Sci Rep. 2017;7 doi: 10.1038/srep46140. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Chen X.M., Cui L.G., He P., Shen W.W., Qian Y.J., Wang J.R. Shear wave elastographic characterization of normal and torn achilles tendons: a pilot study. J Ultrasound Med. 2013;32:449–455. doi: 10.7863/jum.2013.32.3.449. [DOI] [PubMed] [Google Scholar]
- 44.Hong E.K., Choi Y.H., Cheon J.E., Kim W.S., Kim I.O., Kang S.Y. Accurate measurements of liver stiffness using shear wave elastography in children and young adults and the role of the stability index. Ultrasonography. 2018;37:226–232. doi: 10.14366/usg.17025. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
SWE/B-mode images, segmentation masks, and the per-image elastic modulus (kPa) table are available from the corresponding author on reasonable request. All source code — including full layer-level model specifications, hyperparameters, random seeds, preprocessing details, fixed data-split files, and training scripts — is openly available at https://github.com/anabiosi-data/Prognostic_AnaBioSi-Data.
Data will be made available on request.
