Skip to main content
Medicine logoLink to Medicine
. 2026 Jun 26;105(26):e49226. doi: 10.1097/MD.0000000000049226

Advancing breast cancer diagnosis through ultrasound imaging: A cross-attention multi-scale vision transformer approach with comparative analysis against ResNet architectures

Haewon Byeon a,*
PMCID: PMC13313694  PMID: 42363450

Abstract

In recent years, the field of medical imaging has witnessed substantial progress due to the integration of advanced machine learning techniques, particularly in the diagnosis of critical conditions such as breast cancer. This study aims to improve the predictive accuracy of breast cancer diagnosis using ultrasound images by employing the cross-attention multi-scale vision transformer (CrossViT). The proposed methodology involves a dual-branch architecture in which each branch processes image patches of different sizes, thereby capturing both fine-grained and coarse-grained features. The model incorporates a cross-attention mechanism that efficiently fuses these multi-scale features, enhancing its ability to discern complex patterns in medical images. A public ultrasound dataset was partitioned at the patient level using a stratified 80/10/10 train/validation/test split. The development data used for model optimization included 4074 benign and 4042 malignant training images, along with 500 benign and 400 malignant validation images after preprocessing and augmentation, and final model performance was assessed on a held-out test set. Hyperparameters were fine-tuned using a grid search strategy to optimize performance, and training was conducted with stochastic gradient descent and regularization techniques to support stable convergence. Results from this single-dataset experiment showed that CrossViT yielded higher observed performance metrics than the evaluated ResNet architectures across accuracy, precision, recall, F1-score, and AUC. However, these findings should be interpreted as exploratory comparative observations rather than statistically confirmed evidence of model superiority. In conclusion, CrossViT represents a promising technical approach in medical imaging and may have potential utility for automated breast cancer diagnosis in future clinically validated settings.

Keywords: breast cancer, cross-attention, machine learning, medical imaging, vision transformer

1. Introduction

Breast cancer is one of the most common cancers among women globally, and early detection significantly improves treatment outcomes. Ultrasound imaging, in particular, is a crucial tool for early breast cancer diagnosis due to its noninvasive nature and wide accessibility.[1] However, ultrasound images can be challenging to interpret,[2] necessitating the development of accurate automated diagnostic systems. Predicting breast cancer based on ultrasound images is a critical topic in medical artificial intelligence (AI), largely because of the images’ inherent complexity and the subjectivity of their interpretation.

To address these challenges, deep learning techniques have recently been utilized, with a notable focus on models that combine the strengths of Vision Transformers (ViT) and convolutional neural network (CNN)-based architectures.[3,4] Traditional CNN models excel at capturing local spatial information but are limited in their ability to model long-range dependencies across the entire image.[4] In contrast, vision transformers can better capture such global dependencies through self-attention mechanisms.[5] Against this backdrop, the development of architectures that efficiently integrate local and global representations is highly significant.[3,4] In particular, multi-scale approaches based on cross-attention can effectively combine features from different spatial resolutions, leading to richer image representations.[6]

Recent advancements in vision transformers have shown promising results in image classification tasks.[5] In parallel, the importance of multi-scale feature representation has been emphasized in newer transformer architectures.[6,7] In this study, we developed a breast cancer prediction model based on the cross-attention multi-scale vision transformer (CrossViT). This model processes image patches of different sizes through 2 distinct branches and integrates these multi-scale features using a cross-attention module. The objective of this research is to perform breast cancer prediction using breast ultrasound images and to compare the observed performance of CrossViT with that of traditional deep learning architectures such as ResNet-18, ResNet-34, and ResNet-50. This paper is organized as follows: first, we introduce the theoretical background and related works, followed by a detailed description of the proposed model’s design and implementation. We then evaluate the model’s performance through various experiments and finally discuss the results and present our conclusion.

2. Related work

In computer vision research, the incorporation of attention modules into CNNs has driven remarkable progress in recent years. Approaches such as squeeze-and-excitation network, convolutional block attention module, and efficient channel attention network utilize channel-based and/or spatial attention to strengthen feature representation.[8–10] These developments helped establish the conceptual foundation for vision transformers, which replace conventional convolutional operations with self-attention mechanisms.

The introduction of the vision transformer (ViT)[5] marked a turning point, demonstrating that transformer-based models can achieve competitive performance in image classification benchmarks. ViT reinterprets an image as a sequence of smaller patches, similar to how text is represented as a sequence of words, enabling the model to capture global context across the entire image.[5] Nevertheless, its dependence on large-scale datasets and substantial computational resources has encouraged further research into more efficient transformer designs.

One such development is the cross-attention multi-scale vision transformer (CrossViT),[6] which employs a two-branch design to process patches at different resolutions for multi-scale feature extraction. Cross-attention layers then merge information between branches, improving representational richness while controlling computational cost. This design has enabled CrossViT to remain competitive with other contemporary transformer models, including DeiT and swin transformer,[7,11] highlighting the value of cross-attention for robust feature learning.

Within medical imaging – particularly in ultrasound-based breast cancer diagnosis – the potential of ViT and its variants, including CrossViT, has drawn growing interest. While CNN architectures such as residual network (ResNet) remain widely used in medical image classification,[12] their reliance on local receptive fields can limit their ability to capture long-range dependencies across the entire image.[4,12] By contrast, vision transformers can leverage global attention to extract richer and more comprehensive features.

Even so, certain obstacles persist. The heavy computational requirements of transformer architectures highlight the need for efficient designs, especially in medical settings where computing resources may be limited. Moreover, the scarcity of large, well-annotated datasets in healthcare remains a major barrier to the broader deployment of transformer-based systems.

3. Methodology

The proposed model leverages the CrossViT to enhance the predictive accuracy of breast cancer diagnosis using ultrasound images (https://www.kaggle.com/datasets/vuppalaadithyasairam/ultrasound-breast-images-for-breast-cancer). This approach integrates fine-grained and coarse-grained feature representations through a dual-branch architecture. The dataset consists of ultrasound images categorized as benign and malignant and was divided at the patient level using a stratified 80/10/10 train/validation/test scheme. The development set included 4074 benign and 4042 malignant training images, along with 500 benign and 400 malignant validation images. To mitigate overfitting and improve generalization on this medical imaging dataset, data augmentation was performed, with rotational transformation serving as a primary augmentation strategy. The detailed composition of the dataset is presented in Table 1, representative examples of the augmented images are shown in Figure 1, and the overall research workflow is presented in Figure 2.

Table 1.

Dataset composition and augmentation.

Aspect Description
Model used Cross-attention multi-scale vision transformer (CrossViT)
Objective Enhance predictive accuracy of breast cancer diagnosis using ultrasound images
Architecture Dual-branch architecture integrating fine-grained and coarse-grained feature representations
Dataset composition
 Training set 4074 benign images, 4042 malignant images
 Validation set 500 benign images, 400 malignant images
Data augmentation Rotation applied to increase diversity and robustness

Figure 1.

Figure 1.

Representative examples of augmented ultrasound breast images for cancer diagnosis.

Figure 2.

Figure 2.

Overall research methodology pipeline for breast cancer diagnosis using CrossViT. CrossViT = cross-attention multi-scale vision transformer.

3.1. Data preprocessing and augmentation

Each ultrasound image was resized to a uniform input resolution (224 × 224 pixels), converted to grayscale if needed, z-score normalized per image, and lightly denoised (median filter, 3 × 3) prior to augmentation. Augmentation included random rotations (±15°), horizontal flips (P = .5), and limited random translations (±5%). Data were split at the patient level using a stratified 80/10/10 train/validation/test scheme with a fixed random seed (seed = 42), preserving the benign/malignant class ratio across splits; no subject overlapped across partitions. The held-out test set was used only for final evaluation. Models were trained on Ubuntu 22.04/ CUDA 12.1 using NVIDIA RTX 3090 (24GB VRAM) and AMD Ryzen 9 5950X (128GB RAM). Source codes are provided in Supplementary Material S1, Supplemental Digital Content 1.

3.2. Model architecture

The CrossViT[6] is an architecture designed to improve image classification tasks by effectively capturing multi-scale features through a dual-branch transformer framework (Fig. 3). This architecture was selected to address the challenge of extracting both local and global features from ultrasound images, a task for which conventional CNNs may be less effective when long-range dependencies are important.

Figure 3.

Figure 3.

The overall architecture of the CrossViT model. CrossViT = cross-attention multi-scale vision transformer.

Architecture of CrossViT. The CrossViT architecture is fundamentally composed of multiple multi-scale transformer encoders (Fig. 4), each incorporating 2 distinct branches: the large (L) branch (Fig. 4) and the small (S) branch. As shown in Figure 3, cross-attention fuses S-branch patch tokens with the L-branch classification token to retain local detail while preserving global context. These branches operate on different scales of the input image, leveraging the strengths of varying patch sizes to capture a comprehensive range of features.

Figure 4.

Figure 4.

Feature fusion mechanisms in multi-scale vision transformers (A) All-attention fusion, in which every token is combined without taking into account any of its distinguishing features. (B) Fusion of class tokens, in which only CLS tokens are combined since they serve as a representation of a single branch globally. (C) Pairwise fusion, in which CLS tokens are fused independently and tokens at the corresponding spatial locations are fused together. (D) Cross-attention is the process of fusing patch tokens from one branch with CLS tokens from another. CLS = classification.

Large branch (L-Branch): This branch focuses on capturing coarse-grained features by processing larger patch sizes (Fig. 5). It can effectively handle broad, low-resolution features throughout the image because it uses more transformer encoders and wider embedding dimensions. This branch thus serves as the backbone for capturing the overall structure and significant patterns within the data.

Figure 5.

Figure 5.

Cross-attention mechanism within the CrossViT architecture. CrossViT = cross-attention multi-scale vision transformer.

Small branch (S-Branch): Complementing the L-Branch, the S-Branch is tasked with handling smaller patch sizes. This approach facilitates the extraction of fine-grained details, enabling the model to focus on local patterns and textures that might be critical for distinguishing subtle differences, particularly in medical imaging scenarios like ultrasound breast cancer detection.

3.3. Cross-attention mechanism

The core of CrossViT’s design is its innovative cross-attention mechanism, which efficiently fuses features from its dual-branch architecture (Fig. 3). This mechanism operates by allowing the classification token of one branch to interact with the patch tokens of the other branch, effectively merging the diverse feature sets. This fusion is critical as it combines the detailed insights from the S-Branch with the structural context from the L-Branch, enhancing the model’s ability to interpret complex image data. The specific configurations of these branches, including their patch sizes and dimensions, are detailed in Table 2.

Table 2.

CrossViT architecture configurations.

Model Patch size (S) Patch size (L) Dim (S) Dim (L) # Heads M L
CrossViT-small 12 16 192 384 6 4 1
CrossViT-base 12 16 384 768 12 4 1
CrossViT-custom 12 16 224 448 7 6 3

CrossViT = cross-attention multi-scale vision transformer.

3.4. Efficiency and performance

The design of CrossViT balances predictive performance with computational efficiency. By fusing features through cross-attention, the model reduces some of the computational burden associated with broader all-token attention schemes. Empirical studies of the original architecture have shown that CrossViT can maintain competitive accuracy with a moderate increase in floating-point operations and parameters, making it a relevant option for applications requiring both representational richness and efficiency.

The CrossViT architecture, with its dual-scale processing and cross-attention fusion, provides a practical framework for medical image classification tasks that require the integration of detailed local information and broader contextual structure. In the context of ultrasound-based breast cancer prediction, this design offers a promising basis for improved automated diagnostic support.

Training and Optimization. The CrossViT model was fine-tuned on the augmented ultrasound dataset. A comprehensive grid search was conducted to optimize key hyperparameters for both CrossViT and the benchmark ResNet architectures (ResNet-18, -34, -50), ensuring a fair comparison. The search ranges were as follows:

Learning rate: 1 × 10−4 to 1 × 10−2

Batch size: 16, 32, 64

Epochs: 50, 100, 150

All models were trained using the stochastic gradient descent optimizer. To facilitate stable and efficient convergence, a StepLR scheduler was applied, reducing the learning rate by a factor of 0.1 every 30 epochs. To mitigate overfitting, a dropout rate of 0.3 and a weight decay of 1 × 10−4 were implemented for all models.

3.5. Evaluation metric

The performance of the model was assessed using a comprehensive set of metrics: accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), and model loss. These metrics provide a holistic view of the model’s ability to correctly classify benign and malignant cases. The performance of CrossViT was compared with that of traditional CNN architectures such as ResNet-18, ResNet-34, and ResNet-50 in order to characterize relative model behavior on this dataset.

4. Results

The evaluation of CrossViT for breast cancer diagnosis using ultrasound images was conducted using multiple performance metrics and compared with traditional CNN architectures such as ResNet-18, ResNet-34, and ResNet-50. The evaluation metrics included accuracy, precision, recall, F1-score, AUC, and model loss. Together, these metrics provide a detailed view of the model’s ability to distinguish between benign and malignant cases.

4.1. Model performance metrics

The results presented in Table 3 and Figures 6–8 show that, within this dataset, CrossViT yielded the highest observed performance metrics among the evaluated models for breast cancer ultrasound image diagnosis.

Table 3.

Performance metrics for each model.

Model Accuracy Precision Recall F1 Sensitivity Specificity AUC
CrossViT 0.935 0.928 0.942 0.935 0.942 0.928 0.96
ResNet-18 0.912 0.905 0.910 0.907 0.910 0.914 0.94
ResNet-34 0.921 0.917 0.920 0.918 0.920 0.922 0.95
ResNet-50 0.928 0.920 0.925 0.922 0.925 0.931 0.95

AUC = area under the receiver operating characteristic curve, CrossViT = cross-attention multi-scale vision transformer, ResNet = residual network.

Figure 6.

Figure 6.

Comparison of F-1 scores for each model.

Figure 8.

Figure 8.

Comparison of model loss for each model.

Figure 7.

Figure 7.

Comparison of ROC curves for each model. ROC = receiver operating characteristic.

CrossViT achieved the highest observed values across the key evaluation metrics. Specifically, accuracy was 0.935, compared with 0.928 for ResNet-50, 0.921 for ResNet-34, and 0.912 for ResNet-18. The AUC reached 0.96, whereas the ResNet models showed values ranging from 0.94 to 0.95. Furthermore, the F1-score was 0.935, indicating balanced overall predictive performance. However, because confidence intervals and formal statistical significance tests were not performed, these differences should be interpreted as comparative observations rather than statistically confirmed improvements.

CrossViT also showed a favorable balance between precision and recall, which are clinically relevant metrics. Precision was 0.928 and recall was 0.942, indicating that the model maintained a balance between false-positive reduction and true-positive detection. Sensitivity was 0.942, which was numerically higher than that of the ResNet models, including ResNet-50 (0.925). Specificity was 0.928, compared with 0.914 for ResNet-18 and 0.922 for ResNet-34. These results suggest that CrossViT performed well across both types of diagnostic error, although the observed differences require further statistical validation.

The observed performance pattern may be related to the ability of CrossViT’s dual-branch architecture to capture multi-scale features through its cross-attention mechanism. In addition, CrossViT recorded a relatively low model loss of 0.24, suggesting stable optimization and consistent predictive behavior compared with the evaluated ResNet models. Collectively, these results indicate that CrossViT is a promising model for further investigation in automated breast cancer diagnosis, particularly in the context of exploratory single-dataset evaluation.

5. Discussion

The results of this study highlight the potential of the cross-attention multi-scale vision transformer (CrossViT) for breast cancer diagnosis using ultrasound images. In this dataset, CrossViT showed higher observed accuracy, precision, recall, F1-score, and AUC values than the evaluated ResNet architectures. These findings suggest that the model’s dual-branch design may be advantageous for integrating both fine-grained image details and broader contextual information. For example, the cross-attention mechanism can combine texture information from small image patches with morphological context from larger patches, which is relevant when differentiating malignant from benign tumors.[6] Nevertheless, these findings should be interpreted as preliminary because they were obtained from a single public dataset without formal statistical comparison.

Our findings are broadly consistent with recent studies reporting promising applications of transformer-based models in breast ultrasound and related medical imaging tasks.[13–15] While CNNs such as ResNet remain strong baseline models in medical image analysis,[12] their convolution-based structure can be less effective at modeling long-range dependencies across the full image.[4,12] By contrast, transformer-based approaches can incorporate global contextual relationships more directly. In breast ultrasound analysis, this capability may be relevant because lesion shape, margin characteristics, and internal echo patterns all contribute to diagnostic interpretation.

Despite these promising results, this study has several limitations that warrant future research. First, the model was not externally validated on an independent clinical cohort, and all results were derived from a single public dataset, which limits generalizability. Second, the public dataset does not provide per-case patient demographics or lesion characteristics such as lesion size, BI-RADS category, or tissue density, all of which may influence model performance; future work should incorporate datasets with richer metadata to assess subgroup effects. Future studies could also explore self-supervised or transfer learning approaches[16] to reduce data demands. Third, although we report point estimates for clinically relevant metrics such as sensitivity and specificity, we did not report 95% confidence intervals or conduct formal statistical testing such as bootstrap-based confidence intervals, McNemar tests for paired proportions, or DeLong tests for AUC. Therefore, the comparative differences reported in this study should not be interpreted as statistically confirmed superiority. Future work should include uncertainty estimation and hypothesis testing, with appropriate correction for multiple comparisons, to strengthen the robustness of the conclusions. Fourth, although the cross-attention mechanism is relatively efficient for a transformer, the overall model complexity may still present challenges in resource-constrained environments. Optimization strategies such as pruning or quantization should therefore be investigated for mobile or edge deployment. Finally, interpretability remains a major challenge for AI systems intended for clinical use. Medical professionals require a clear understanding of model decision-making to support trust and patient safety. Our current study focused on predictive performance, but an important next step will be to implement explainable AI techniques such as attention maps or Grad-CAM to visualize the image regions that most strongly influence model predictions. Such approaches may help bridge the gap between technical performance and clinical validation.[17]

6. Conclusion

In conclusion, this exploratory single-dataset study suggests that CrossViT is a promising approach for automated breast cancer diagnosis using ultrasound images. Its multi-scale feature processing and cross-attention mechanism enabled higher observed performance metrics than those of the evaluated ResNet architectures in this dataset. However, because external validation, confidence intervals, and formal statistical significance testing were not included, these results should be interpreted cautiously as preliminary comparative findings rather than definitive evidence of model superiority. With further validation using multicenter data and rigorous statistical evaluation, CrossViT may become a useful tool for clinical decision support in early breast cancer detection.

Author contributions

Conceptualization: Haewon Byeon.

Data curation: Haewon Byeon.

Formal analysis: Haewon Byeon.

Funding acquisition: Haewon Byeon.

Investigation: Haewon Byeon.

Methodology: Haewon Byeon.

Project administration: Haewon Byeon.

Resources: Haewon Byeon.

Software: Haewon Byeon.

Supervision: Haewon Byeon.

Validation: Haewon Byeon.

Visualization: Haewon Byeon.

Writing – original draft: Haewon Byeon.

Writing – review & editing: Haewon Byeon.

medi-105-e49226-s001.pdf (63.8KB, pdf)

Abbreviations:

AI
artificial intelligence
AUC
area under the receiver operating characteristic curve
CNN
convolutional neural network
CrossViT
cross-attention multi-scale vision transformer
L-Branch
large branch
RAM
random access memory
ResNet
residual network
S-Branch
small branch
ViT
vision transformer
VRAM
video random access memory

This research was supported by the Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (NRF- RS-2023-00237287).

The author has no conflicts of interest to declare.

The datasets generated during and/or analyzed during the current study are publicly available.

Supplemental Digital Content is available in the online version of this article (http://dx.doi.org/10.1097/MD.0000000000049226).

How to cite this article: Byeon H. Advancing breast cancer diagnosis through ultrasound imaging: A cross-attention multi-scale vision transformer approach with comparative analysis against ResNet architectures. Medicine 2026;105:26(e49226).

References

  • [1].Iacob R, Stoicescu ER, Ghenciu DM, et al. Evaluating the role of breast ultrasound in early detection of breast cancer in low- and middle-income countries: a comprehensive narrative review. Bioengineering. 2024;11:262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [2].Sood R, Rositch AF, Shakoor D, et al. Ultrasound for breast cancer detection globally: a systematic review and meta-analysis. J Glob Oncol. 2019;5:1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Takahashi S, Sakaguchi Y, Kouno N, et al. Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review. J Med Syst. 2024;48:84. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Li J, Chen J, Tang Y, et al. Transforming medical imaging with Transformers? A comparative review of key properties, current progresses, and future perspectives. Front Med. 2023;10:1086097. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [5].Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16 × 16 words: transformers for image recognition at scale. In: International Conference on Learning Representations (ICLR). 2021. [Google Scholar]
  • [6].Chen CF, Fan Q, Panda R. CrossViT: cross-attention multi-scale vision transformer for image classification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021:347–56. [Google Scholar]
  • [7].Liu Z, Lin Y, Cao Y, et al. Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021:9992–10002. [Google Scholar]
  • [8].Hu J, Shen L, Sun G. Squeeze-and-excitation networks. IEEE Trans Pattern Anal Mach Intell. 2020;42:2011–23. [DOI] [PubMed] [Google Scholar]
  • [9].Woo S, Park J, Lee JY, Kweon IS. CBAM: convolutional block attention module. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018:3–19. [Google Scholar]
  • [10].Wang Q, Wu B, Zhu P, Li P, Zuo W, Hu Q. ECA-Net: efficient channel attention for deep convolutional neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020:11531–9. [Google Scholar]
  • [11].Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H. Training data-efficient image transformers & distillation through attention. In: Proceedings of the 38th International Conference on Machine Learning (ICML). 2021;139:10347–57. [Google Scholar]
  • [12].He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016:770–8. [Google Scholar]
  • [13].Ayana G, Choe S-W. Vision transformers-based transfer learning for breast mass classification from multiple diagnostic modalities. J Electr Eng Technol. 2024;19:3391–410. [Google Scholar]
  • [14].Ayana G, Choe S-W. BUViTNet: breast ultrasound detection via vision transformers. Diagnostics (Basel). 2022;12:2654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Li M, Yin F, Shao J, et al. Vision transformer-based multimodal fusion network for classification of tumor malignancy on breast ultrasound: a retrospective multicenter study. Int J Med Inform. 2025;196:105793. [DOI] [PubMed] [Google Scholar]
  • [16].Zhou SK, Greenspan H, Davatzikos C, et al. A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises. Proc IEEE Inst Electr Electron. 2021;109:820–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Bashir Z, Lin M, Feragen A, et al. Clinical validation of explainable AI for fetal growth scans through multi-level, cross-institutional prospective end-user evaluation. Sci Rep. 2025;15:2074. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

medi-105-e49226-s001.pdf (63.8KB, pdf)

Articles from Medicine are provided here courtesy of Wolters Kluwer Health

RESOURCES