Skip to main content
RSNA Journals logoLink to RSNA Journals
. 2025 Jun 25;7(5):e240550. doi: 10.1148/ryai.240550

High-Performance Open-Source AI for Breast Cancer Detection and Localization in MRI

Lukas Hirsch 1, Elizabeth J Sutton 2, Yu Huang 1, Beliz Kayis 1, Mary Hughes 2, Danny Martinez 2, Hernan A Makse 3, Lucas C Parra 1,✉
PMCID: PMC12464713  PMID: 40560044

Abstract

Purpose

To develop and evaluate an open-source deep learning model for detection and localization of breast cancer on MRI scans.

Materials and Methods

In this retrospective study, a deep learning model for breast cancer detection and localization was trained on the largest breast MRI dataset to date. Data included all breast MRI examinations conducted at a tertiary cancer center in the United States between 2002 and 2019. The model was validated on sagittal MRI scans from the primary site (n = 6615 breasts). Generalizability was assessed by evaluating model performance on axial data from the primary site (n = 7058 breasts) and a second clinical site (n = 1840 breasts).

Results

The primary site dataset included 30 672 sagittal MRI examinations (52 598 breasts) from 9986 female patients (mean age, 52.1 years ± 11.2 [SD]). The model achieved an area under the receiver operating characteristic curve of 0.95 for detecting cancer in the primary site. At 90% specificity (5717 of 6353), model sensitivity was 83% (217 of 262), which was comparable to historical performance data for radiologists. The model generalized well to axial examinations, achieving an area under the receiver operating characteristic curve of 0.92 on data from the same clinical site and 0.92 on data from a secondary site. The model accurately located the tumor in 88.5% (232 of 262) of sagittal images, 92.8% (272 of 293) of axial images from the primary site, and 87.7% (807 of 920) of secondary site axial images.

Conclusion

The model demonstrated state-of-the-art performance on breast cancer detection. Code and weights are openly available to stimulate further development and validation.

Keywords: Computer-aided Diagnosis (CAD), MRI, Neural Networks, Breast

Supplemental material is available for this article.

See also commentary by Moassefi and Xiao in this issue.

© RSNA, 2025

Keywords: Computer-aided Diagnosis (CAD), MRI, Neural Networks, Breast


graphic file with name ryai.240550.VA.jpg


Summary

An open-source deep learning model developed and trained on the largest breast MRI dataset to date achieved state-of-the-art performance in breast cancer detection and localization.

Key Points

  • ■ A two-dimensional convolutional neural network trained on a large breast MRI dataset achieved state-of-the-art performance (area under the receiver operating characteristic curve, 0.95), which is comparable to radiologist performance.

  • ■ The model generalized well across acquisition orientations and clinical sites.

  • ■ Open-source code and pretrained model weights are made publicly available.

Introduction

Breast cancer remains a leading cause of cancer-related deaths among women in the United States (1). Early detection is crucial for successful treatment and improved patient outcomes (2). For women at high risk of developing breast cancer, annual MRI screening is recommended in addition to mammography (3). Breast MRI is also used diagnostically when a tumor is suspected based on clinical findings, mammography, or US. It is effective for early breast cancer detection (4), including in women with dense breast tissue for whom mammography may be less reliable. Supplemental breast MRI use is expected to rise after recent recommendations to screen women with extremely dense breasts (5).

Breast MRI interpretation is time-consuming and requires specialized training, as radiologists must review multiple sections in each volume. Automated reading has the potential to assist radiologists by identifying MRI sections most likely to contain a tumor (6) or triaging low-probability images that do not require reading (7). Given that three of four biopsies are negative in routine clinical practice (8,9), reliable prediction of negative outcomes through automation could also help reduce the biopsy burden.

Substantial progress has been made with deep learning models in radiology (10–18), with the ability to detect subtle patterns and abnormalities by analyzing large datasets. This progress is most evident in mammography, in which population-wide screening programs have generated large datasets containing hundreds of thousands of images (19–22), enabling the development and validation of accurate models and fueling a growing interest in automation. In contrast, breast MRI screening reaches a smaller population, resulting in more limited datasets. Therefore, model performance in breast MRI has yet to consistently match that of radiologists and often excludes complex cases, such as those involving implants or postsurgical changes (23,24). Moreover, most existing models provide only a probability of malignancy (10) without localizing the area of concern on MRI scan, limiting their practical utility for radiologists.

Modern deep learning models contain a large number of parameters, making them prone to overfitting when trained on small, single-site datasets, which may limit generalizability to other sites (25–27). This problem is particularly relevant for MRI, for which image acquisition parameters can vary substantially between institutions. Even within a single institution, temporal changes, such as a shift from sagittal- to axial-plane acquisition, can introduce variability (28). To date, only two published studies have included multisite validation for breast cancer detection in MRI (10,29). Cross-site validation remains challenging, largely because most trained models have not been publicly released, with one recent exception (10). Additionally, performance may be constrained by the relatively small size of available datasets, typically limited to a few hundred examinations (30–33).

This study aimed to develop an open-source deep learning model for MRI-based breast cancer detection trained on a large dataset comprising tens of thousands of examinations from Memorial Sloan-Kettering Cancer Center in New York. The model was validated across different imaging planes and clinical sites, including an external dataset from Duke University (34). The model was designed to both detect and localize the cancer, thus aiding radiologists during interpretation. By making the model and its parameters openly available, we aim to foster further research and development in this field.

Materials and Methods

Study Sample and Data Partition

The use of this retrospective data was approved by the City College of New York institutional review board with a waiver of informed consent, and all procedures were compliant with the Health Insurance Portability and Accountability Act. Identifiable patient information was removed, and MRI scans were saved with anonymized identifiers before analysis.

We used three distinct datasets for training and testing from a primary and secondary site (Fig 1).

Figure 1:

Overview of three datasets (sagittal and axial) used for model training and testing, with patient counts, malignancy labels, and annotation availability shown for each partition.

Data overview for breast cancer detection model development and evaluation. Three datasets were used: sagittal data from the primary clinical site (training, validation, testing), axial data from the primary site (testing), and axial data from a secondary clinical site (testing). The sagittal dataset was partitioned at the patient level (indicated as percentage for each partition). No patient overlap existed between the primary site's sagittal and axial data. Each box displays the number of patients, examinations, individual breasts (accounting for bilateral and unilateral examinations), malignant breasts, and malignant breasts with radiologist annotations indicating cancer location.

Primary site.

Data included all breast MRI examinations conducted in women at Memorial Sloan-Kettering Cancer Center between January 2002 and December 2019. The inclusion criterion was a complete sequence of dynamic contrast-enhanced MRI and available pathology or clinical follow-up of 2 years. Data included both screening and diagnostic imaging with multiple examinations for each screening patient. Data were excluded for benign breast images from 332 screening patients who eventually developed cancer to avoid potential false-negative results. Primary site data were then separated into sagittal and axial examinations (details in Table 1 and below). To assign labels for each breast, bilateral examinations were separated into right and left breasts. MRI scans were labeled as “malignant” when there was biopsy-proven cancer and “benign” if there was no cancer diagnosis within 2 years of clinical follow-up. Two years is the standard follow-up period for treatment studies (35,36) and is more stringent than previous deep learning studies when labeling healthy breasts (10). Model performance was evaluated on individual breasts as well as examinations.

Table 1:

Summary of Patient Demographics, Examination Counts, Imaging Protocol Distribution, Cancer Labels, and BI-RADS Categories for the Sagittal Training, Validation, and Test Sets and the External Axial Test Set

Variable Sagittal Training Set Sagittal Validation Set Sagittal Test Set Sagittal Total Axial Total Overall Total
No. of patients 9986 (81) 1110 (9) 1233 (10) 12 329 3219 15 548
Race
 Asian Far East or Indian subcontinental 376 (3.77) 30 (2.7) 52 (4.22) 458 60 (1.86) 518
 Black or African American 533 (5.34) 67 (6.04) 74 (6) 674 62 (1.93) 736
 Native American or Alaska Native 1 (0.01) 0 0 1 0 1
 Native Hawaiian or Pacific Islander 3 (0.03) 0 0 3 0 3
 White 6744 (67.53) 753 (67.84) 832 (67.48) 8329 732 (22.74) 9061
 Unknown or missing 2329 (23.32) 260 (23.42) 275 (22.3) 2864 107 (3.32) 2971
Ethnicity
 Hispanic 459 (4.6) 55 (4.95) 76 (6.16) 590 56 (1.74) 646
 Not Hispanic 7252 (72.62) 797 (71.8) 886 (71.86) 8935 823 (25.57) 9758
 Other or unknown 2275 (22.78) 258 (23.24) 271 (21.98) 2804 81 (2.52) 2885
No. of examinations 30 672 3463 3870 38 005 3873 41 878
Age (y) 52.1 ± 11.2 51.9 ± 11.1 52.3 ± 10.8 … 54.4 ± 10.6 …
Examination label
 Malignant 2114 (6.89) 238 (6.87) 255 (6.59) 2607 688 (17.76) 3295
 Benign or negative 28 558 (93.11) 3225 (93.13) 3615 (93.41) 35 398 3178 (82.06) 38 576
BI-RADS category
 0 347 (1.13) 40 (1.16) 44 (1.14) 431 4 (0.1) 435
 1 3949 (12.87) 545 (15.74) 534 (13.8) 5028 870 (22.46) 5898
 2 17 712 (57.75) 1913 (55.24) 2182 (56.38) 21 807 2054 (53.03) 23 861
 3 3792 (12.36) 409 (11.81) 489 (12.64) 4690 248 (6.4) 4938
 4 2455 (8) 287 (8.29) 321 (8.29) 3063 195 (5.03) 3258
 5 446 (1.45) 46 (1.33) 72 (1.86) 564 10 (0.26) 574
 6 1943 (6.33) 221 (6.38) 224 (5.79) 2388 492 (12.7) 2880
Lesion size (cm) 1.8 ± 1.2 1.8 ± 1.2 1.8 ± 1.2 … 2.7 ± 1.2 …

Note.—Data are presented as numbers with percentages in parentheses or means ± SDs. BI-RADS = Breast Imaging Reporting and Data System.

Primary-site sagittal-plane examinations (training, validation, and testing sets).

Sagittal data included 38 005 examinations from 2002 to 2014 (31 564 screening, 6015 diagnostic, 426 unknown or not applicable) from 12 329 patients. Counting each breast individually yielded 65 105 sagittal breast images with 2690 malignant images. This dataset was drawn from the same patient cohort reported in previous work (37), which was used to develop a lesion segmenter and is otherwise independent from this study. The data were randomly divided by patient into training, validation, and test sets (90 for training and 10 for test and subsequently 90 for training and 10 for validation patients). This dataset included radiologist segmentation in two dimensions (2D) for the section containing the largest (index) cancer for all 2690 malignant breast images (termed the index section). Radiologists selected a single index section per breast showing the largest tumor extent.

Primary-site axial-plane examinations (testing set).

To evaluate the model's performance on a different imaging protocol, axial MRI scans, excluding patients from the sagittal cohort, were used. This dataset comprised 3873 examinations from 2013 to 2019 (3069 screening, 720 diagnostic, 84 unknown or not applicable) from 3219 patients and 7058 breasts. The dataset contained 688 malignant and 6370 benign breast images. Volumetric segmentation generated by radiologists for the index lesions was available for 293 of the malignant breasts. Cancer localization could be evaluated only on this subset.

Secondary-site axial-plane examinations (testing set).

Performance was also evaluated on axial MRI scans from a secondary clinical site, using a public dataset released by Duke University (34). This dataset included 922 axial examinations from patients with confirmed breast cancer and excluded two examinations due to issues identifying pre- and postcontrast images. This dataset included pathology information and radiologist annotations on the extent of the index lesion in one breast. This information was provided as a three-dimensional (3D) bounding box, only for one lesion, even if there was multicentric or bilateral breast cancer present. Breast images with malignant pathology were labeled as “malignant” (n = 948, as some examinations had malignancy in both breasts). Contralateral breast images without malignant pathology were labeled as “benign” (n = 892) (see Appendix S1).

Model Architecture

We used a conventional 2D convolutional neural network (CNN) designed to detect breast cancer in the 3D volume by assigning a probability of containing cancer to each 2D sagittal section of a breast (Fig 2). The maximum probability across all sections was used as the prediction for the whole breast. Localization performance was evaluated in terms of the distance of the maximum probability section from the reference standard provided by radiologists on the location of the cancer. The inputs to the network were three input channels capturing dynamic contrast enhancement (T1-weighted postcontrast, dynamic contrast-enhanced in, dynamic contrast-enhanced out) for the corresponding 2D section. The model consisted of 13 2D convolutional layers (3 × 3 kernel size), each followed by batch normalization and ReLU activation. A max pooling layer was added after every two convolutional layers for a total of five downsampling steps. This layering mapped the input (of dimensions 512,512,3) to image features (of dimensions 16,16,252), which were flattened, processed by one dense layer, and then concatenated with clinical features, ending in two final dense layers and a Softmax activation. In total, this network had 4 126 294 trainable parameters. Code and model weights can be accessed at https://github.com/lkshrsch/BreastCancerDiagnosisMRI.

Figure 2:

Diagram showing model input (sagittal MRI sections with three channels), CNN-based processing, and output used for cancer detection (max probability) and localization (argmax index) compared to reference standards.

Cancer detection and localization using artificial intelligence. The input of the model (purple box) consists of a sagittal section from the full three-dimensional MRI with three channels (T1-weighted postcontrast [T1post], dynamic contrast enhancement in [DCE-in], and dynamic contrast enhancement out [DCE-out]). The model (blue box) is a two-dimensional (2D) deep convolutional neural network (CNN) that outputs a probability of cancer for each section. Model output is evaluated in two tasks: detection and localization (yellow box). The maximum probability (max) is used as the prediction for the whole volume and evaluated in the detection task against the reference standard (pathology from the biopsy of clinical follow-up). Similarly, the section index with maximum probability (argmax) serves as the estimate of tumor location and is evaluated against the reference standard provided as annotations on cancer location made by radiologists. BatchNorm = batch normalization, Conv2D = 2D convolutional layer, Hx = history, x5 = times five.

Data Preprocessing and Harmonization

Preprocessing followed our previous work in segmentation (37). Briefly, pre- and postcontrast T1-weighted images were coregistered using NiftyReg (38), and dynamic contrast enhancement was summarized into images capturing initial contrast material uptake (dynamic contrast-enhanced in) and washout (dynamic contrast-enhanced out), alongside the first postcontrast T1-weighted image (T1-weighted postcontrast). These three channels were normalized by dividing by the 95th percentile of the precontrast T1-weighted image in each examination. To adjust for interchannel differences, each channel was divided by its 95th percentile across the training set. The sagittal MRI data had varying in-plane resolutions (0.4–0.8 mm). Low-resolution images were upsampled by a factor of two for harmonization. Axial images were resampled to match a sagittal in-plane resolution of 0.4 mm and separated into left and right breasts. All images were cropped to 512 × 512 pixels, ensuring the breast was centered. For the axial examinations, the network operated on these resampled sagittal images.

Demographic Data

Demographic information included age and 11 categorical variables with one-hot encoding: family history of breast cancer (yes or no); ethnicity (Hispanic or Latino, not Hispanic, unknown); and race (Asian Far East or Indian subcontinental, Black or African American, Native American or American Islander, Native Hawaiian or Pacific Islander, White, and unknown). All information was self-reported, with missing ethnicity and race imputed as “unknown.”

Training

The model was trained using index sections from malignant images as positive examples. As negative examples, we selected the center section and one randomly selected section from benign images. All models were trained using a focal loss (39) with α = 5, using the Adam optimizer (40) with learning rate of 1e-5. All models were trained for 100 epochs with early stopping (Fig S1), and the network weights with the lowest validation loss were saved for evaluation. Unless otherwise specified, all models were trained with data augmentation, consisting of random rotation within 60 degrees, random shear of scale 0.1, random horizontal and vertical flips, and random intensity scaling of 0.8–1.2, all implemented using the TensorFlow preprocessing ImageDataGenerator library (version 2.0).

Validation and Model Selection

Various model architectures and hyperparameters were compared based on their area under the receiver operating characteristic curve (AUC) performance on the sagittal validation set (Fig S2). Comparisons included different loss functions (binary cross-entropy vs focal loss; Fig S2A), the effect of data augmentation (Fig S2B), and training data sizes (10%, 50%, 100% of the whole data) (Fig S2C). The addition of the contralateral breast as input (Fig S2D) and comparison of the architecture to ResNet-50 (with and without ImageNet pretrained weights) were also tested (Fig S2A). Fine-tuning the pretrained ResNet-50 outperformed training it from scratch, but the CNN, trained from scratch with demographic information, achieved the best validation set performance. Adding the contralateral breast did not significantly improve performance and was excluded from the final model. Data augmentation substantially boosted validation set performance and was included in the final model training.

Categorization of Image Quality in the Sagittal Test Set

The sagittal test set was visually inspected (blinded to labels and predictions by L.H., with 6 years of experience with breast MRI) to categorize image quality and determine its effect on performance. Categories that could occur simultaneously included biopsy clips, implants, large postsurgery changes, and poor image quality (blurriness, movement or fat-saturation artifacts, enhancing nipple tissue).

Statistical Analysis

Performance was evaluated using AUC. Bootstrapped CIs for AUC were obtained by resampling with replacement subjects and the predicted probabilities 1000 times. CIs were then computed from the 2.5 and 97.5 percentiles of the AUC values derived from each data drawn during bootstrapping.

AUC differences between examination categories were assessed using a bootstrap sample. All images were combined, ignoring categories, and randomly drawn with replacements to match benign or malignant numbers in categories. AUC differences from this bootstrap sample were used to compute P values in a one-sided test, assuming conservatively that the presence of these clinical and imaging abnormalities worsened model performance (Fig S3). Wilcoxon signed rank tests and Pearson correlation tests were performed with the scipy.stats package in Python (version 3.1; Python Software Foundation). P < .05 was considered statistically significant.

Results

We trained deep networks with various configurations using the sagittal examinations from the primary site. The primary site dataset included 30 672 sagittal MRI examinations (52 598 breasts) from 9986 female patients (mean age, 52.1 years ± 11.2 [SD]; range, 13–93 years) (Fig 2; Table 1). The AUC measured on a validation set showed the benefit of the large data size, 2D data augmentation, and superiority of the CNN over a ResNet-50 (Fig S2). Based on this, we selected for final testing a 2D deep CNN (Fig 2) trained with data augmentation, focal loss function, and demographic information.

Detection Performance on Sagittal Images

Evaluation on a random subset of patients not included in the training set (6615 breast sagittal images, 262 with biopsy-confirmed cancers) demonstrated an AUC of 0.95 (95% CI: 0.93, 0.96) (Fig 3A). Evaluating results by the outcome of the examination instead of each individual breast achieved an AUC of 0.94 (Fig S4A). We also evaluated results using only the first examination for screening patients to rule out repeated measures and obtain the same AUC of 0.94 (Fig S5A). All examinations underwent routine clinical Breast Imaging Reporting and Data System assessment by radiologists, and the estimated cancer probability generally increased with the Breast Imaging Reporting and Data System score (Fig S6), supporting internal validity of the model's prediction. A network trained on demographic information alone achieved an AUC of 0.60 (95% CI: 0.56, 0.63) (Fig S2A). Of the demographic variables, only family history and unknown race and ethnicity were individually associated with outcome (Table S2).

Figure 3:

Histograms of cancer probability predictions and ROC curves across three test sets: (A) sagittal primary site, (B) axial primary site, and (C) axial secondary site.

Test set performance on sagittal data and generalization to axial data across sites. Top: Histograms of predicted probability of cancer for all breasts, color coded by breast outcome. Bottom: Receiver operating characteristic curves and area under the receiver operating characteristic curves. The 95th percentile bootstrap CIs are shown in blue. (A) Sagittal test set from the primary site. (B) Axial test set from the primary site. (C) Axial test set from the secondary site.

Generalization of Detection Performance to Axial Images

MRI volumes have a higher in-plane resolution. Although the training dataset primarily consisted of sagittal images (higher resolution in depth and lateral directions), current clinical practice often uses axial acquisition. To assess the model's generalization to axial images (higher lateral and vertical resolution), we resampled axial images in the sagittal plane to match the resolution of the training data and processed them with the same trained model. Without fine-tuning, the model achieved an AUC of 0.92 (95% CI: 0.91, 0.93) on the primary site axial data (7058 breasts, 688 cancers) (Fig 3B) and an AUC of 0.92 (95% CI: 0.91, 0.93) on axial scans from a secondary site (948 malignant, 892 benign) (Fig 3C). Because the axial data from the primary site included screening data, multiple samples were from the same patient. After removing this correlation and evaluating results only considering the first examination, the model achieved a performance of 0.90 (Fig S4B). Similarly, when only evaluating performance based on examination outcomes instead of individual breasts, the model achieved an AUC of 0.92 (Fig S5B). Data from the secondary site already consisted of single patients, and the outcome of all examinations was positive; therefore, these evaluations could not be applied.

Cancer Localization

The model estimated cancer probability for each 2D section in the MRI volume (Fig 2). The section with the maximum probability localized the tumor. To determine the accuracy of this localization for the sagittal data from the primary site, we used the 2D segmentation provided by radiologists for the index lesion and extended these to 3D using volumetric automatic segmentation (37). The model's maximum probability section intersected the lesion volume in 88.5% (232 of 262) of breasts (Fig 4A). Additionally, the maximum probability section correlated with the index section provided by the radiologist (n = 262; Pearson correlation, r = 0.87; P < .001) (Fig S7).

Figure 4:

Predicted tumor locations (dots) compared to radiologist-defined cancer extents (horizontal lines) across three datasets: (A) sagittal primary, (B) axial secondary, and (C) axial primary, with hits and misses color-coded.

Predicted location and lesion lateral extent. Comparison of the location predicted by the network (dots) relative to the lateral extent of the cancer provided by radiologists (horizontal lines and shading). The lateral extent indicates the range of sagittal sections that contain cancer. The vertical axis indicates different breasts sorted by extent. A “hit” indicates that the location predicted by the network falls within the lesion (blue); a “miss” indicates no overlap (orange). The percentage of hits is shown in blue text. Breasts are sorted by decreasing lesion size for both hits and misses. (A) Sagittal test data of the primary site (n = 262). Lateral extent for each cancer is based on a semiautomatic segmentation. (B) Axial data from the secondary site (n = 920). Lateral extent is determined from bounding boxes provided by radiologists. (C) Axial data from the primary site (n = 293). Lateral extent is based on radiologists’ volumetric segmentations. Most of the large cancers missed in the axial primary site correspond to cancers found in the left breast, which explains the rightward shift of the orange dots (this effect, however, is not statistically significant [binomial test: cancers, 21; left breast, 14; population frequency of left cancers, 0.53 (156 of 293); P = .15]).

For axial data, the maximum probability section intersected the 3D segmentation provided by radiologists in 92.8% (272 of 293) of breasts from the primary site (Fig 4C) and the bounding box around the cancers in 87.7% (807 of 920) of breasts from the secondary site (Fig 4B).

Post Hoc Exploratory Analyses of Test Case Subsets

We conducted exploratory analyses on subsets of the sagittal test data to further characterize the model's performance. To assess its ability to predict biopsy outcomes, we evaluated the performance on the subset that received biopsies (Breast Imaging Reporting and Data System 4 and 5: n = 578; 94 cancers), achieving an AUC of 0.86 (Fig S8). Given that this subset consisted only of suspicious cases requiring biopsies, a lower performance was expected.

To determine the robustness of the model, we analyzed performance on cases typically excluded from previous studies, such as implants and imaging artifacts (10,41). All images were categorized into one of four categories (Table 2). For categories with both benign and malignant images, we evaluated AUC and assessed if performance differed significantly between breasts with or without the category (eg, presence or absence of biopsy clips). We found that performance was comparable across all these categories.

Table 2:

Model Performance with Challenging Cases

Attribute Absence of Each Attribute Presence of Each Attribute AUC Difference
Benign Malignant AUC Benign Malignant AUC
Implant 5913 260 0.95 393 2 >0.99 P = .79
Biopsy clip 4347 173 0.94 1940 89 0.95 P = .73
Postsurgery change 6184 258 0.96 122 4 0.91 P = .20
Poor image quality 5091 207 0.94 1215 55 0.96 P = .75

Note.—Unless otherwise indicated, data are numbers of breasts. Model performance did not significantly change for cases considered challenging that were excluded in previous studies (10,23,24). “Poor image quality” encompasses various issues (see Methods section). AUC = area under the receiver operating characteristic curve.

Discussion

We demonstrated that training a 2D CNN from scratch with a uniquely large MRI dataset can lead to a new state-of-the-art performance in breast cancer detection (Table 3). Performance also benefited from efficient implementations of 2D data augmentation methods. Notably, we leveraged information from radiologists regarding the location of the cancer along one dimension (section number with the index lesion). This information was considerably more informative than a single overall diagnostic label for the entire volume and allowed us to design a network that highlighted the image most likely to contain a tumor. In doing so, the network provided interpretable results that may assist radiologists in their diagnostic workflow.

Table 3:

Summary of Studies to Date Involved in Detection of Breast Cancer in MRI Using Deep Neural Networks

Publication Pretrained Model AUC No. in Training Set No. in Test Set Comments Published Model Use Contralateral Multisite MRI Protocol
Herent et al 2019 (30) ResNet-50 (ImageNet) 0.82 335 168 Evaluation on single section. Generates attention heatmaps through initial segmentation. No Yes (axial) No Axial
Truhn et al 2019 (43) ResNet-18 (ImageNet) 0.88 1294 647 Manual cropping of lesions. Radiologist performance of 0.98. Postcontrast sequence as RGB channels. No Unclear (no) No Axial
Amit et al 2017 (23) CIFAR-10, VGGNet 0.91 1256 1256 Requires manual selection of ROIs on lesions. Only BI-RADS 2 benign and BI-RADS 5 malignant. Only single lesions. No asymmetric BPE. Pretrained models performed worse than model trained from scratch. No Not specified No Not specified
Hu et al 2021 (41) VGG16 (ImageNet) 0.93 1455 535 Dynamic contrast-enhanced time sequence as RGB channels. Manual cropping of lesions before input to CNN. No No No Sagittal
Zhou et al 2019 (24) No 0.86 1073 307 Only single lesions. No asymmetric BPE. Only evident BPE. No Yes (axial) No Axial
Verburg et al 2022 (29) No 0.83 9162 4581 Triages 40% of normal breasts at NPV of 100% in women with extremely dense breasts (DENSE trial). No Yes (axial) Yes Axial
Dalmış et al 2017 (42) No 0.81 201 160 No spatial registration of contralateral. No use of postcontrast images. No Yes No Transverse and coronal
Li et al 2017 (32) No 0.84 80 43 Manual cropping of lesions. No No No Axial
Witowski et al 2022 (10) 3D ResNet-18 (Kinetics-400) 0.92 14198 3936 No interpretability of results. Under request case by case Yes (axial) Yes Sagittal and axial
Zhang et al 2023 (33) Mask R-CNN and ResNet-50 Sensitivity, 96%; specificity, 70% 241 176 Outputs bounding box on lesion. No Yes (axial) No Axial
This study No 0.95 (axial 0.92) 52 598 breasts; 30 672 examinations 6615 Identifies relevant section and tumor location in section. Yes Yes Yes Sagittal and axial

Note.—AUC = area under the receiver operating characteristic curve, BI-RADS = Breast Imaging Reporting and Data System, BPE = background parenchymal enhancement, CNN = convolutional neural network, DENSE = Dense Tissue and Early Breast Neoplasm Screening, NPV = negative predictive value, RGB = red, green, blue, ROI = region of interest.

This work addresses several limitations of previous studies: data size, case exclusions, out-of-plane and cross-site validation, interpretability, comparison with radiologist performance, and public release of the trained model. We will discuss each aspect in turn while providing an overview of published results on deep learning–based breast cancer detection in MRI (Table 3).

A key distinguishing factor of this study was the size of the dataset used. Recently, curated MRI datasets have grown from a few hundred (30,32,33,42) to thousands of images (10,29,41). Larger datasets offer greater diversity, enabling the training of large models from scratch, as we have shown here (Fig 5).

Figure 5:

Scatterplot comparing training set size and internal test AUC across related studies, with color coding for model/code availability status.

Comparison of related studies to date. Each study is shown in terms of the size of the training data (number of examinations) versus performance on the internal test set of each study (area under the receiver operating characteristic curve [AUC]). The colors indicate if the study makes code and trained weights openly available (red: not available, orange: upon request, green: no restriction).

In contrast, previous studies (23,30,32,33,41,43) with smaller datasets relied on fine-tuning pretrained models from unrelated tasks, like ImageNet or Kinetics-400 (44). However, features extracted by such models may differ substantially from those needed for analyzing 3D MR images. Some studies have shown that models trained on just over a thousand examinations can outperform those pretrained on ImageNet (23). Others have shown that fine-tuning a pretrained model can improve performance, as compared with training from scratch, even when a larger MRI dataset is available (10). Our validation set analysis found that fine-tuning ResNet-50 trained on ImageNet numerically outperformed training it from scratch, supporting the view that pretrained models do help, even if the imaging domains are quite distinct. However, the simpler CNN model, trained from scratch, outperformed a ResNet-50, showing that with enough data, the model architecture becomes less important.

Unlike many studies that exclude difficult cases like breast implants, postoperative changes, or imaging artifacts (10,23,24), our training included such examples. We hypothesized that the larger training set would encompass enough of these anomalies to avoid exclusions. Indeed, our overall performance (AUC, 0.95; 95% CI: 0.93, 0.96) exceeded the top-performing artificial intelligence study to date (10) (AUC, 0.92; 95% CI: 0.92, 0.93), despite including previously excluded cases. Beyond the obvious benefit of providing detection for all cases, avoiding exclusions is essential for handling tens of thousands of examinations, as manual visual inspection is impractical. Full automation also eliminates the need for manual region selection, as required in previous studies (23,43).

This study trained a model using sagittal MRI scans because this was the largest dataset available to us at the time. Since MRI resolution is higher in plane, axial test set images were upsampled vertically. Nonetheless, the model performed well on both primary- and secondary-site axial scans without fine-tuning, demonstrating its ability to generalize beyond the training data. Notably, it performed well in both cancer classification and localization on axial examinations.

Trained on individual sections without global position knowledge, the model accurately selected sections containing index lesions. This accuracy was evidenced by the high hit rates and correlation with the index section from radiologists. Although some previous models estimated the location of the tumor (29,30,33,43), many did not (10,23,32,41). Providing such information is crucial for integrating artificial intelligence into clinical workflows, building confidence in its detection, and potentially guiding abbreviated radiologist reevaluations.

The network's detection performance appears to be comparable to historical data on radiologist performance. Among the four studies with such measures, Witowski et al (10) reported an average AUC of 0.89 (95% CI: 0.85, 0.95) for five readers on 100 examinations, numerically lower than our model's AUC of 0.95 (95% CI: 0.93, 0.96). Other studies reported varying radiologist sensitivity and specificity, with our model demonstrating numerically superior specificity at those sensitivities. Dalmiş et al (31) reported a radiologist sensitivity of 98% (specificity, 28%) and Zhou et al (24) a sensitivity of 59% (specificity, 86%). At these sensitivities, our model achieved numerically superior specificities (62% and 99%, respectively). One outlier was the study by Truhn et al (43) that reported a higher radiologist AUC of 0.98 (95% CI: 0.96, 0.99), likely due to an easier curated dataset (eg, with high prevalence, including only large and enhancing lesions, excluding high-risk benign lesions). However, direct statistical comparison is impossible as datasets are unavailable and due to potential heterogeneity. Although our model outperformed the existing state-of-the-art study (10) when tested on sagittal data, the performance matched that on axial test data using a much simpler model. We expect performance gains when fine-tuning the model on such axial MRI scans.

This study had limitations. First, although our dataset was uniquely large, it primarily comprised sagittal MRI scans from a single institution. Although the model generalized well to external axial scans, broader multi-institutional training would likely improve performance further. Second, the training approach relied on section-level annotations, which may not be routinely available. Third, although we did include demographic information, we did not explore its added value in detail. Finally, although model performance was comparable to radiologists on retrospective datasets, a prospective reader study on the same data is needed for a direct comparison.

Training a 2D CNN from scratch on a large and diverse breast MRI dataset enabled state-of-the-art cancer detection, even in challenging clinical cases. Incorporating section-level lesion annotations during training improved both classification performance and interpretability by highlighting relevant sections. The model demonstrated generalizability across different MRI protocols and institutions. Future studies should explore prospective clinical validation. To this end, we are openly releasing source code and trained weights. We hope this will enhance reproducibility in artificial intelligence for radiology and encourage further technical development.

Funding: This work was supported by National Institutes of Health grant R01CA247910 with additional support from National Institutes of Health National Cancer Institute grant P30 CA008748 and National Institutes of Health BRAIN Initiative grant R01 EB028157.

Disclosures of conflicts of interest: L.H. No relevant relationships. E.J.S. No relevant relationships. Y.H. No relevant relationships. B.K. No relevant relationships. M.H. No relevant relationships. D.M. No relevant relationships. H.A.M. No relevant relationships. L.C.P. No relevant relationships.

Abbreviations:

AUC
area under the receiver operating characteristic curve
CNN
convolutional neural network
3D
three-dimensional
2D
two-dimensional

References

  • 1. Siegel RL , Miller KD , Fuchs HE , Jemal A . Cancer statistics, 2022 . CA . CA A Cancer J Clinicians 2022. ; 72 ( 1 ): 7 – 33 . [DOI] [PubMed] [Google Scholar]
  • 2. Pinsky PF . Principles of Cancer Screening . Surg Clin North Am 2015. ; 95 ( 5 ): 953 – 966 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Mann RM , Kuhl CK , Moy L . Contrast-enhanced MRI for breast cancer screening . J Magn Reson Imaging 2019. ; 50 ( 2 ): 377 – 390 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Hylton NM , Gatsonis CA , Rosen MA , et al. ; ACRIN 6657 Trial Team and I-SPY 1 TRIAL Investigators . Neoadjuvant Chemotherapy for Breast Cancer: Functional Tumor Volume by MR Imaging Predicts Recurrence-free Survival-Results from the ACRIN 6657/CALGB 150007 I-SPY 1 TRIAL . Radiology 2016. ; 279 ( 1 ): 44 – 55 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Mann RM , Athanasiou A , Baltzer PAT , et al. ; European Society of Breast Imaging (EUSOBI) . Breast cancer screening in women with extremely dense breasts recommendations of the European Society of Breast Imaging (EUSOBI) . Eur Radiol 2022. ; 32 ( 6 ): 4036 – 4045 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Sheth D , Giger ML . Artificial intelligence in the interpretation of breast cancer on MRI . Eur J Magn Reson Imaging 2020. ; 51 ( 5 ): 1310 – 1324 . [DOI] [PubMed] [Google Scholar]
  • 7. Bhowmik A , Eskreis-Winkler S . Deep learning in breast imaging . BJR Open 2022. ; 4 ( 1 ): 20210060 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Gavenonis SC , Lee JM , Halpern EF , Rafferty EA . Positive predictive value of breast MRI in cancer detection . Cancer Res 2009. ; 69 ( 2_Supplement ): 4007 . [Google Scholar]
  • 9. Han BK , Schnall MD , Orel SG , Rosen M . Outcome of MRI-Guided Breast Biopsy . AJR Am J Roentgenol 2008. ; 191 ( 6 ): 1798 – 1804 . [DOI] [PubMed] [Google Scholar]
  • 10. Witowski J , Heacock L , Reig B , et al . Improving Breast Cancer Diagnostics with Artificial Intelligence for MRI . Science Translational Medicine 2022. ; 14 ( 664 ): eabo4802 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Wang X , Yang W , Weinreb J , et al . Searching for prostate cancer by fully automated magnetic resonance imaging classification: deep learning vs non-deep learning . Sci Rep 2017. ; 7 ( 1 ): 15415 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Kumar A , Kim J , Lyndon D , Fulham M , Feng D . An Ensemble of Fine-Tuned Convolutional Neural Networks for Medical Image Classification . IEEE J Biomed Health Inform 2017. ; 21 ( 1 ): 31 – 40 . [DOI] [PubMed] [Google Scholar]
  • 13. Zhu Z , Albadawy E , Saha A , et al . Deep learning for identifying radiogenomic associations in breast cancer . Comput Biol Med 2019. ; 109 : 85 – 90 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Suk HI , Lee SW , Shen D ; Alzheimer's Disease Neuroimaging Initiative . Deep ensemble learning of sparse regression models for brain disease diagnosis . Med Image Anal 2017. ; 37 : 101 – 113 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Kim DH , MacKinnon T . Artificial intelligence in fracture detection: transfer learning from deep convolutional neural networks . Clin Radiol 2018. ; 73 ( 5 ): 439 – 445 . [DOI] [PubMed] [Google Scholar]
  • 16. Yoo Y , Tang LYW , Brosch T , et al . Deep learning of joint myelin and T1w MRI features in normal-appearing brain tissue to distinguish between multiple sclerosis patients and healthy controls . Neuroimage Clin 2018. ; 17 : 169 – 178 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Li N , Haopeng L , Bin Q , et al . Detection and Attention: Diagnosing Pulmonary Lung Cancer from CT by Imitating Physicians . arXiv 2017 . Preprint posted online December 14, 2017; https://arxiv.org/abs/1712.05114 .
  • 18. Saiz F , Barandiaran I . COVID-19 Detection in Chest X-ray Images using a Deep Learning Approach . 10.9781/ijimai.2020.04.003. Published June 2020. Accessed January 2025. [DOI]
  • 19. Halling-Brown MD , Warren LM , Ward D , et al . OPTIMAM Mammography Image Database: A Large-Scale Resource of Mammography Images and Clinical Data . Radiol Artif Intell 2021. ; 3 ( 1 ): e200103 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Cai H , Wang J , Dan T , et al . An Online Mammography Database with Biopsy Confirmed Types . Sci Data 2023. ; 10 ( 1 ). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Lopez MG, Posada N, Moura DC, et al. BCDR: a breast cancer digital repository. In: 15th International conference on experimental mechanics, volume 1215: 113–120. https://www.researchgate.net/profile/Jose-Franco-Valiente/publication/258243150_BCDR_A_BREAST_CANCER_DIGITAL_REPOSITORY/links/59afe98a0f7e9bf3c72930e5/BCDR-A-BREAST-CANCER-DIGITAL-REPOSITORY.pdf. [Google Scholar]
  • 22. Heath M , et al . Current Status of the Digital Database for Screening Mammography . In: Karssemeijer N , Thijssen M , Hendriks J , Van Erning L , eds. Digital Mammography . Springer Netherlands; , Dordrecht: , 1998. : 457 – 460 . [Google Scholar]
  • 23. Amit G , Ben-Ari R , Hadad O , Monovich E , Granot N , Hashoul S . Classification of breast MRI lesions using small-size training sets: comparison of deep learning approaches . In: Medical Imaging, Computer-Aided Diagnosis; SPIE , 2017. : 374 – 379 . [Google Scholar]
  • 24. Zhou J , Luo LY , Dou Q , et al . Weakly supervised 3D deep learning for breast cancer classification and localization of the lesions in MR images . J Magn Reson Imaging 2019. ; 50 ( 4 ): 1144 – 1151 . [DOI] [PubMed] [Google Scholar]
  • 25. McDermott MBA , Wang S , Marinsek N , et al . Reproducibility in machine learning for health research: Still a ways to go . Sci Transl Med 2021. ; 13 ( 586 ): eabb1655 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Burns ML , Kheterpal S . Machine Learning Comes of Age: Local Impact vs National Generalizability . Anesthesiology 2020. ; 132 ( 5 ): 939 – 941 . [DOI] [PubMed] [Google Scholar]
  • 27. Barak-Corren Y , Chaudhari P , Perniciaro J , et al . Prediction across healthcare settings: a case study in predicting emergency department disposition . npj Digit Med 2021. ; 4 ( 1 ): 1 – 7 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Mann RM , Cho N , Moy L . Breast MRI: State of the Art . Radiology 2019. ; 292 ( 3 ): 520 – 536 . [DOI] [PubMed] [Google Scholar]
  • 29. Verburg E , van Gils CH , van der Velden BHM , et al . Deep Learning for Automated Triaging of 4581 Breast MRI Examinations from the DENSE Trial . Radiology 2022. ; 302 ( 1 ): 29 – 36 . [DOI] [PubMed] [Google Scholar]
  • 30. Herent P , Schmauch B , Jehanno P , et al . Detection and characterization of MRI breast lesions using deep learning . Diagn Interv Imaging 2019. ; 100 ( 4 ): 219 – 225 . [DOI] [PubMed] [Google Scholar]
  • 31. Dalmiş MU , Gubern-Mérida A , Vreemann S , et al . Artificial Intelligence–Based Classification of Breast Lesions Imaged With a Multiparametric Breast MRI Protocol With Ultrafast DCE-MRI, T2, and DWI . Invest Radiol 2019. ; 54 ( 6 ): 325 – 332 . [DOI] [PubMed] [Google Scholar]
  • 32. Li J , Fan M , Zhang J , Li L . Discriminating between benign and malignant breast tumors using 3D convolutional neural network in dynamic contrast enhanced-MR images . In: Cook TS , Zhang J , eds. Proceedings of the SPIE 2017. , vol 10138 ; 1013808 . [Google Scholar]
  • 33. Zhang Y , Liu YL , Nie K , et al . Deep Learning-based Automatic Diagnosis of Breast Cancer on MRI Using Mask R-CNN for Detection Followed by ResNet50 for Classification . Acad Radiol 2023. ; 30 Suppl 2(Suppl 2) : S161 – S171 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Saha A , Harowicz MR , Grimm LJ , et al . A machine learning approach to radiogenomics of breast cancer: a study of 922 subjects and 529 DCE-MRI features . Br J Cancer 2018. ; 119 ( 4 ): 508 – 516 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Smith I , Procter M , Gelber RD , et al. ; HERA study team . 2-year follow-up of trastuzumab after adjuvant chemotherapy in HER2-positive breast cancer: a randomised controlled trial . Lancet 2007. ; 369 ( 9555 ): 29 – 36 . [DOI] [PubMed] [Google Scholar]
  • 36. Emens LA , Davidson NE . The follow-up of breast cancer . Semin Oncol 2003. ; 30 ( 3 ): 338 – 348 . [DOI] [PubMed] [Google Scholar]
  • 37. Hirsch L , Huang Y , Luo S , et al . Radiologist-Level Performance by Using Deep Learning for Segmentation of Breast Cancers on MRI Scans . Radiol Artif Intell 2022. ; 4 ( 1 ): e200231 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Modat M , Ridgway GR , Taylor ZA , et al . Fast free-form deformation using graphics processing units . Comput Methods Programs Biomed 2010. ; 98 ( 3 ): 278 – 284 . [DOI] [PubMed] [Google Scholar]
  • 39. Lin TY , Goyal P , Girshick R , He K , Dollar P . Focal Loss for Dense Object Detection . 2017 IEEE International Conference on Computer Vision (ICCV) , Venice, Italy : 2999 – 3007 . [Google Scholar]
  • 40. Kingma DP , Ba J . Adam: A Method for Stochastic Optimization . arXiv 2014 . Preprint posted online December 22, 2014; https://arxiv.org/abs/1412.6980 .
  • 41. Hu Q , Whitney HM , Li H , et al . Improved Classification of Benign and Malignant Breast Lesions Using Deep Feature Maximum Intensity Projection MRI in Breast Cancer Diagnosis Using Dynamic Contrast-enhanced MRI . Radiol Artif Intell 2021. ; 3 ( 3 ): e200159 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Dalmış MU , Litjens G , Holland K , et al . Using deep learning to segment breast and fibroglandular tissue in MRI volumes . Med Phys 2017. ; 44 ( 2 ): 533 – 546 . [DOI] [PubMed] [Google Scholar]
  • 43. Truhn D , Schrading S , Haarburger C , et al . Radiomic vs Convolutional Neural Networks Analysis for Classification of Contrast-enhancing Lesions at Multiparametric Breast MRI . Radiology 2019. ; 290 ( 2 ): 290 – 297 . [DOI] [PubMed] [Google Scholar]
  • 44. Kay W , Carreira J , Simonyan K , et al . The Kinetics Human Action Video Dataset . arXiv 2017 . Preprint posted online May 19, 2017; https://arxiv.org/abs/1705.06950 .

Articles from Radiology: Artificial Intelligence are provided here courtesy of Radiological Society of North America

RESOURCES