Abstract
Background
Accurately differentiating early-stage breast cancer from benign lesions on MRI is essential to reduce unnecessary biopsies. However, the limited interpretability of current deep learning models hinders their clinical trustworthiness and adoption. This study aimed to develop a clinically interpretable concept bottleneck model (CBM) that integrates radiologist-specific knowledge and automatically generates structured reports, thereby improving diagnostic accuracy and consistency in breast MRI interpretation.
Methods
Preoperative breast MR images and radiological reports were retrospectively collected from five institutions (January 2016–July 2025) and allocated to internal, external and multi-reader cohorts. Lesion-related descriptors from free-text MRI reports were standardized into BI-RADS-compliant concepts. These concepts, alongside multiparametric MR sequences, were input into the CBM for classification and structured reporting of the lesions annotated by radiologists using bounding boxes. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and compared against a black-box deep learning model. The accuracy of CBM-generated concepts was evaluated. A two-phase multi-reader study was further conducted to assess clinical utility.
Results
A total of 1,695 pathology-confirmed breast lesions (857 malignant and 838 benign) from 1,634 patients (median age 46 years, IQR 39–53) were included. The CBM achieved an AUC of 0.92 (95%CI 0.90–0.93) on the test set, comparable to the black-box model (AUC: 0.93, 95%CI 0.92–0.94). Concept accuracy ranged from 0.64 to 1.00. In the multi-reader study, the CBM matched the diagnostic accuracy of one radiologist and exceeded that of seven others (all P < 0.05). With CBM assistance, radiologists correctly downgraded 22.1% of lesions to benign. Diagnostic accuracy improved for three radiologists (from 0.71 to 0.72 to 0.82–0.91, all P < 0.05), and inter-reader agreement increased for both concept recognition and BI‑RADS category (Gwet’s AC1: 0.27-1.00 to 0.46-1.00).
Conclusions
The CBM provides a versatile framework for classifying early breast cancer and benign lesions. By employing an image-concept alignment strategy, it enhances intrinsic interpretability and offers radiologists clinically relevant, intelligible decision support that serves both diagnostic and educational needs. Moreover, this retrospective study demonstrates its potential to reduce unnecessary biopsies for benign breast lesions and to improve reporting consistency in breast MRI.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12916-026-04889-7.
Keywords: Breast cancer, Biopsy, Interpretable AI, Magnetic resonance imaging, Structured reporting, Breast imaging reporting and data system
Background
MRI demonstrates exceptional sensitivity for breast cancer (BC) diagnosis 1, 2. Yet, its clinical application faces two major challenges. First, the relatively low positive predictive value (PPV), ranging from 35% to 64% in high-risk women and from 19.6% to 35.7% in intermediate- or average-risk populations, implies that its diagnostic advantage may be offset by a considerable number of false-positives, potentially leading to unnecessary biopsies 1, 3, 4, 5, 6. Second, breast MRI interpretation remains labor-intensive, expertise-dependent, and prone to interobserver variability 7, 8. These constraints underscore the need for reliable computer-assisted approaches to increase the accuracy and efficiency of breast MRI analysis and to standardize reporting.
Despite the promise of deep learning (DL) for classifying benign and malignant breast lesions on MRI, its “black-box” nature limits interpretability and hinders clinical adoption 6, 9. Specifically, most current DL models take images as inputs and directly output predictions without providing insight into which image features drive the decision 10. In clinical practice, radiologists do not merely render a benign or malignant diagnosis; rather, they describe specific lesion characteristics (e.g., “spiculated margin”, “irregular shape”) and then synthesize these features to form a comprehensive assessment 11. An ideal lesion classification model should emulate this structured reasoning to earn clinicians’ trust. Therefore, developing interpretable models that elucidate the decision-making rationale and are tailored to radiologists’ needs is imperative for successful clinical translation 10, 12.
Previous studies have often employed post-hoc interpretability methods, e.g., class activation mapping (CAM) and Shapley Additive exPlanations (SHAP), to generate heatmaps or feature importance scores after a model makes a prediction, highlighting the image regions or features most influential on the output 9, 12. However, these methods cannot reveal the model’s internal decision process or translate findings into clinically meaningful patterns 9, 12, 13. Bridging this gap necessitates intrinsically interpretable strategies, such as the concept bottleneck model (CBM), which incorporates radiologist-specific knowledge during training by aligning image features with textual concepts derived from radiological reports, thereby allowing the model to learn radiologists’ reasoning processes 12, 14, 15.
Meanwhile, multimodal large language models (MLLMs) have garnered significant attention in medicine for their ability to efficiently process multimodal inputs, extract and interpret textual and visual information, demonstrate powerful associative reasoning, and translate complex machine learning outputs into comprehensible narratives 9, 16, 17, 18, 19, 20, 21. These strengths position MLLMs as powerful tools to facilitate standardized medical image interpretation and optimize reporting workflow 9, 21, 22, 23. Nevertheless, their full potential in assisting breast lesion classification and presenting model predictions in a format readily understandable to radiologists remains largely untapped.
Therefore, we aimed to develop a CBM to classify early-stage BC and benign lesions. Given a user‑provided bounding box that encloses the lesion, the model first predicted a set of clinically meaningful concepts from the image. It then performed the final classification on the basis of the predicted concepts, thereby providing diagnostic rationales understandable to radiologists. With the assistance of an embedded GPT-5, it also generated structured reports. We evaluated the CBM on several key aspects: classification performance, concept prediction accuracy, and the ability to reduce unnecessary biopsies among benign lesions categorized as BI‑RADS category 4. A two-phase multi-reader diagnostic study further assessed its clinical utility.
Methods
Patient Cohorts
This study was approved by the local institutional ethics committee (approval No. 2025 − 949) and adhered to the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines specific to artificial intelligence 24. Informed consent was waived due to the retrospective design of the study. Between January 2016 and July 2025, 1,634 patients with 1,695 breast lesions were enrolled from five hospitals: West China Hospital (WCH), Chengdu Second People’s Hospital (CDSPH), Guang’an People’s Hospital (GAPH), The First People’s Hospital of Guangyuan (GYFPH), and Pidu District People’s Hospital (PDDPH). Inclusion criteria were: (1) pathologically confirmed early-stage breast cancer (TNM stages T1-2, N0-1, M0) or benign lesions of any pathological type, with preoperative multiparametric breast MRI performed for screening or surgical planning 25, 26. Non-mass enhancement (NME) lesions were excluded due to the frequent difficulty in distinguishing them from physiological background parenchymal enhancement (BPE), leading to potential mismatches between pathological lesion areas and MRI enhancement regions and thus unreliable annotation 5. Further exclusion criteria included: (1) severe artifacts affecting MRI interpretation, incomplete sequences, or lesions not visible on MRI; (2) ambiguous pathology or lesions not definitively classifiable as benign or malignant (e.g., borderline phyllodes tumor); (3) concurrent ipsilateral benign and malignant lesions; (4) missing clinicopathologic data or MRI reports; and (5) prior treatment before MRI. Details on sample size calculation are provided in Additional file 1-Supplementary methods 27, 28. Clinical characteristics recorded included lesion location, size, amount of fibroglandular tissue, background parenchymal enhancement level, and patient age.
The data were partitioned into three independent cohorts: (1) Internal cohort: Comprising 1,425 lesions (712 malignant, 713 benign) from WCH (January 2016 to April 2024), randomly divided into training and test sets in an 8:2 ratio. Five-fold cross-validation was performed on the training set. For each fold, the best-performing model was retained and evaluated on the held-out test set. (2) External cohort: Consisting of 205 lesions (114 malignant, 91 benign) collected between April 2021 and May 2025 from four external institutions (CDSPH, GAPH, GYFPH, and PDDPH), used to assess model generalizability. (3) Multi-reader cohort: Including 65 lesions (31 malignant and 34 benign) from WCH (May to July, 2025), utilized for clinical utility assessment. No data overlap across the datasets. All the data were de-identified before analysis.
MRI Acquisition and Processing
MR images were acquired using 3.0 T or 1.5 T scanners from four manufacturers, all of which were equipped with dedicated breast coils. The following imaging sequences were utilized: (1) fat-suppressed T2-weighted imaging (T2WI); (2) diffusion-weighted imaging (DWI) with b values of 800–1000 s/mm2 and corresponding apparent diffusion coefficient (ADC) maps; (3) early-phase contrast-enhanced T1-weighted imaging (CE); and (4) time‒intensity curves (TICs). Key imaging parameters and details of image preprocessing are summarized in the Additional file 1-Supplementary methods and Additional file 2-Table S16, 29. Regions of interest (ROIs) were manually delineated on the CE series by a board-certified radiologist using ITK-SNAP (version 3.8.0, https://www.itksnap.org/), referencing pathology‑confirmed lesion locations. Three-dimensional rectangular bounding boxes encapsulating each lesion were generated on the basis of the maximum dimensions of the individual lesions. These bounding boxes were then propagated to the other registered volumes.
Concept Bank Establishment
Given the potential variability in descriptors used by radiologists for identical MRI findings, we standardized the relevant terminology. GPT-4o (OpenAI) was prompted to extract lesion-relevant concepts from unstructured free-text reports. These concepts were subsequently standardized and semantically categorized according to the BI-RADS lexicon 11. A radiologist (J.Q., with 5 years of experience) reviewed the corresponding images and original reports, and used Visual Studio Code (version 1.100.3; https://code.visualstudio.com/) to verify and correct the GPT-extracted concepts saved in JSON format. All the concepts were then reviewed by a senior breast-dedicated radiologist (J.H., with more than 20 years of experience). This process resulted in a standardized concept bank of 47 items, which comprehensively covers the BI-RADS terms used to describe lesions and associated features (Additional file 2-Table S2) and has been widely applied in clinical practice. This rigorous two-step review ensured semantic accuracy and facilitated standardization across multiple centers. Further details are available in Additional file 1-Supplementary methods and Additional file 2-Figs. S1-S4).
CBM Construction, Evaluation and Comparison
An image-text alignment model named CBM was developed. The CBM offered a unified and extensible framework for integrating multiparametric MR images and textual concepts. Within this framework, images were processed by sequence-specific encoders (EfficientNet backbones for T2WI, DWI, ADC, and CE) to generate intermediate low-dimensional representations. The kinetic information describing lesion enhancement patterns in the initial and delayed phases was extracted from TICs using GPT-4o (Additional file 1-Supplementary methods and Additional file 2-Fig. S5). These features were converted into one-hot labels and subsequently processed by a multilayer perceptron. The outputs from all the image encoders and the multilayer perceptron were then integrated using an attention pooling module to merge information across the five modalities.
For textual data, multiple linear layers were trained to perform binary classification for each concept. Given image-concept pairs, images corresponding to the target concept were designated as positive samples, and all others as negatives. This process yielded a set of well-trained linear layers capable of projecting image representations into the concept space. The resulting projections were employed to compute alignment scores, which quantified the affinity between the input image and each concept. These alignment scores were subsequently fed into a linear classifier for the estimation of malignant probability.
Essentially, the CBM aligned intermediate image representations with radiologist-interpretable concepts, allowing intuitive interpretation of model predictions through the activation and linear combination of these concept representations. The activated concepts function as evidence, mirroring image features observed by radiologists during MRI interpretation 15. Ultimately, the CBM leverages these activated concepts, along with its malignancy prediction, to generate structured reports containing standardized lesion descriptions, malignancy probabilities, BI-RADS categories, and management recommendations. In this stage, GPT-5 was positioned after the linear layer, seamlessly transforming the output of the CBM into coherent clinical narratives while adhering to the required structured formats. Further details on CBM construction are provided in Additional file 1-Supplementary methods and Additional file 2-Fig. S6.
The CBM was evaluated across three aspects: (1) Classification performance: Metrics included the area under the receiver operating characteristic curve (AUC), accuracy, recall, precision, and F1-score, using a probability threshold of ≥ 0.5 to define malignancy 30. All the metrics were presented with 95% CIs, calculated using t-distribution statistics. (2) Concept accuracy: Assessed through directly comparison of CBM-generated concepts against radiologist-determined ground-truth 18. (3) Ability to reduce unnecessary biopsy: The CBM was specifically evaluated for its potential to reduce unnecessary biopsies in benign lesions classified as BI-RADS category 4 across both the test and external datasets.
In concept-based interpretable models, a trade-off between accuracy and interpretability is commonly observed. Imposing prior constraints on the CBM—requiring input images to align with predefined concepts—introduces additional regularization into the hypothesis space, potentially compromising classification performance 14, 15. For reference, we additionally trained and evaluated an end-to-end black-box model without conceptual components on the same datasets (see Additional file 1-Supplementary methods and Additional file 2-Fig. S7). The performance of this model served as a benchmark against that of the CBM.
Multi-reader Study
For a preliminary evaluation of the clinical utility of CBM, 65 lesions from the multi-reader cohort were used. The assessments involved (1) comparing the diagnostic accuracy of CBM with that of radiologists and (2) examining the impact of CBM assistance on radiologists’ diagnostic decisions and interpretation time. Eight radiologists from different tertiary hospitals, with 3–11 years of experience, participated in this two-phase multi-reader study via an online survey platform (www.wjx.cn). They were grouped by experience: three with less than 5 years (less-experienced) and five with 5 years or more (experienced) 31. Each reader assessed all the cases twice, with a two-week interval between assessments, yielding 130 independent evaluations per reader. Prior to the study, all the readers received training on the study protocol and the use of the DICOM viewer (RadiAnt DICOM Viewer, Version 2020.2.3).
During MRI interpretation, readers were provided with only brief clinical complaints and multiparametric MR images. Pathology localization details were supplied exclusively for cases with multifocal lesions. A structured reporting template was employed to standardize the documentation (Additional file 2-Fig. S2). Based on the imaging findings, readers selected relevant descriptors from the template and assigned a BI-RADS category to each lesion. For comparison with CBM, the reader-assigned BI-RADS categories were dichotomized as benign (categories 2 and 3) or malignant (categories 4 and 5) 32. After a two-week washout period, the cases were reshuffled and reinterpreted. In this phase, readers were additionally provided with structured reports generated by the CBM. Changes in concept identification, BI-RADS reclassification, and MRI interpretation time were recorded for each reader.
Assessment of Multiparametric MRI Superiority and Quantification of Concept Importance
The superiority of multiparametric MRI over single-sequence strategies in distinguishing benign from malignant lesions remains unclear; t-SNE dimensionality reduction was thus applied to map high-dimensional features into a low-dimensional space, facilitating the visualization of complex feature distributions across classes 33. Additionally, the classification performance of the models based on individual MRI sequences was separately calculated to provide supplementary insight. Capitalizing on the intrinsic interpretability of the CBM, concept importance was quantified through weight-based analysis to identify the most discriminative concepts for classifying breast lesions. Specifically, by examining the absolute values of the classifier weights—which reflect the contribution of each concept to predicting malignancy versus benignity—concepts critical for lesion discrimination were pinpointed 34.
Subsequently, for the key concepts identified, a correlation analysis using the chi-square test along with Cramér’s V coefficient was performed to quantitatively evaluate the agreement between the concepts predicted by the CBM from the input images and those annotated by radiologists 35, 36. This analysis assesses the consistency between the visual patterns recognized by the CBM and their corresponding concept representations.
Statistical analysis
Statistical analysis was conducted using R (version 3.6.1; http://www.R-project.org/) and Python (version 3.11; https://www.python.org/). Differences in continuous variables were assessed using the independent t‒test or Mann‒Whitney U test, as appropriate, whereas categorical variables were compared using the chi-square test or Fisher’s exact test. Comparisons of AUCs were performed using DeLong’s test. The impact of the CBM on clinical decision-making was evaluated via decision curve analysis. McNemar’s test was employed to compare the diagnostic accuracy of radiologists with and without CBM assistance. Interpretation times were analyzed using the Wilcoxon signed-rank test. Inter-radiologist agreement regarding concepts and BI-RADS categories was assessed using Gwet’s agreement coefficient (AC1 value) 20. A P-value of less that 0.05 was considered statistically significant.
Results
Patient and Lesion Characteristics
The final analysis cohort comprised 1,695 lesions (median size 16 mm, IQR 11–22) from 1,634 patients (median age 46 years, IQR 39–53) (Fig. 1). Of these, 857 lesions (median size 20 mm, IQR 15–25) were pathologically confirmed as early-stage BC, diagnosed between January 2022 and July 2025 in patients with a median age of 48 years (IQR 43–56). The remaining 838 lesions (median size 12 mm, IQR 10–16) were benign, diagnosed between January 2016 and July 2025 in patients with a median age of 42 years (IQR 35–49). Patient and lesion characteristics are summarized in Table 1. A schematic overview of the study design and workflow is presented in Fig. 2.
Fig. 1.
Flowchart of patient inclusion. Initially, a total of 2,478 patients with breast lesions were consecutively enrolled, all of whom underwent biopsy or surgery. After applying the exclusion criteria, 1,634 patients were finally included. Abbreviations: WCH: West China Hospital; CDSPH: Chengdu Second People’s Hospital; GAPH: Guang’an People’s Hospital; GYFPH: The First People’s Hospital of Guangyuan; PDDPH: Pidu District People’s Hospital
Table 1.
Patient and lesion characteristics
| Characteristic | Internal Cohort (n = 1,425) | External Cohort (n = 205) | Multi-reader Cohort (n = 65) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BC | Benign | P value | BC | Benign | P value | BC | Benign | P value | |||
| No. lesions | 712 | 713 | NA | 114 | 91 | NA | 31 | 34 | NA | ||
| Location | (χ² = 1.18, P = 0.28) | (χ² = 0.00, P > 0.99) | (χ² = 0.00, P > 0.99) | ||||||||
| left | 340 (47.8%) | 362 (50.8%) | 61 (53.5%) | 48 (52.7%) | 17 (54.8%) | 19 (55.9%) | |||||
| right | 372 (52.2%) | 351 (49.2%) | 53 (46.5%) | 43 (47.3%) | 14 (45.2%) | 15 (44.1%) | |||||
| Lesion size, median (IQR), cm | 2.0 (1.5–2.5) | 1.2 (0.9–1.6) | (Z = 18.12, P < 0.001) *** | 2.23 ± 0.67 # | 1.4 (1.0-1.8) | (Z = 6.90, P < 0.001) *** | 1.94 ± 0.53 # | 1.0 (1.0-1.3) | (Z = 5.74, P < 0.001) *** | ||
| Amount of FGT | (χ² = 87.27, P < 0.001) *** | (χ² = 18.83, P < 0.001) *** | (χ² = 2.28, P = 0.53) | ||||||||
| Almost entirely fat | 9 (1.3%) | 2 (0.3%) | (χ² = 3.31, P = 0.28) | 1 (0.9%) | 0 (0.00%) | (χ² = 0.00, P > 0.99) | 1 (3.2%) | 0 (0.00%) | |||
| Scattered fibroglandular tissue | 111 (15.6%) | 57 (8.0%) | (χ² = 19.04, P < 0.001) *** | 43 (37.7%) | 15 (16.5%) | (χ² = 10.23, P = 0.01) * | 2 (6.5%) | 1 (2.9%) | |||
| Heterogeneous fibroglandular tissue | 464 (65.2%) | 375 (52.6%) | (χ² = 22.75, P < 0.001) *** | 44 (38.6%) | 62 (68.1%) | (χ² = 16.52, P < 0.001) *** | 25 (80.6%) | 27 (79.4%) | |||
| Extreme fibroglandular tissue | 128 (17.9%) | 279 (39.1%) | (χ² = 77.09, P < 0.001) *** | 26 (22.8%) | 14 (15.4%) | (χ² = 1.33, P = 0.99) | 3 (9.7%) | 6 (17.7%) | |||
| BPE level | (χ² = 62.63, P < 0.001) *** | (χ² = 15.93, P = 0.001) ** | (χ² = 3.11, P = 0.37) | ||||||||
| Minimal | 285 (40.0%) | 151 (21.2%) | (χ² = 58.73, P < 0.001) *** | 34 (29.8%) | 17 (18.7%) | (χ² = 2.79, P = 0.38) | 13 (42.0%) | 16 (47.0%) | |||
| Mild | 314 (44.1%) | 386 (54.1%) | (χ² = 13.96, P < 0.001) *** | 54 (47.4%) | 29 (31.9%) | (χ² = 4.42, P = 0.14) | 16 (51.6%) | 14 (41.2%) | |||
| Moderate | 84 (11.8%) | 136 (19.1%) | (χ² = 13.90, P < 0.001) *** | 18 (15.8%) | 32 (35.1%) | (χ² = 9.28, P = 0.009) ** | 1 (3.2%) | 4 (11.8%) | |||
| Marked | 29 (4.1%) | 40 (5.6%) | (χ² = 1.51, P = 0.88) | 8 (7.0%) | 13 (14.3%) | (χ² = 2.17, P = 0.56) | 1 (3.2%) | 0 (0.0%) | |||
| No. patients | 707 | 668 | 113 | 82 | 31 | 33 | |||||
| Age, median (IQR), yrs | 48 (42–56) | 42 (35–49) | (Z = 11.63, P < 0.001) *** | 51 (47–59) | 45.64 ± 10.34 # | (Z = 4.32, P < 0.001) *** | 47.10 ± 8.75 # | 45.18 ± 10.0 # | (t = 0.81, P = 0.42) | ||
| ≤ 40 | 140 (19.8%) | 289 (43.3%) | (χ² = 86.99, P < 0.001) *** | 10 (8.8%) | 23 (28.0%) | (χ² = 11.13, P = 0.001) ** | 6 (19.4%) | 10 (30.3%) | (χ² = 0.52, P = 0.47) | ||
| > 40 | 567 (80.2%) | 379 (56.7%) | 103 (91.2%) | 59 (72.0%) | 25 (80.6%) | 23 (69.7%) | |||||
Note: A total of 1,634 patients (1,695 lesions) were included in this study, among whom six were diagnosed with bilateral breast cancer and fifty-five with bilateral benign breast lesions. For the 61 patients with bilateral breast lesions, each breast was analyzed separately. Except where indicated, data are numbers of lesions or patients, with percentages in parentheses. Abbreviations: IQR, interquartile range; yrs, years; BPE, background parenchymal enhancement; FGT, amount of fibroglandular tissue. # Data are expressed as means ± SDs, with normality confirmed by Shapiro‒Wilk tests. Significance levels: * P < 0.05, ** P < 0.01, *** P < 0.001
Fig. 2.
Study design and workflow. Some icons are from the open-source libraries Bioicons (https://bioicons.com/) and PinClipart (https://www.pinclipart.com/)
The malignancies included invasive breast carcinoma (94.5%, 810/857, including invasive breast carcinomas of no special type, invasive lobular, mucinous, and metaplastic carcinomas); ductal carcinoma in situ (4.0%, 34/857); papillary neoplasms (1.0%, 9/857; comprising invasive and solid papillary carcinomas); and other types (0.5%, 4/857, including adenoid cystic carcinoma, malignant adenomyoepithelioma, malignant phyllodes tumor, and medullary carcinoma). The benign lesions included adenosis and benign sclerosing lesions (50.7%, 425/838); fibroadenoma (37.7%, 316/838); intraductal papilloma (5.6%, 47/838); and other benign conditions (6.0%, 50/838; including benign epithelial proliferation and precursors, benign phyllodes tumor, adenomyoepithelioma, adiponecrosis, desmoid fibromatosis, granulomatous mastitis, granulosa cell tumor, hemangioma, hamartoma, myofibroblastoma, neurofibroma, schwannoma, and tubular adenoma). The classification of subtypes followed the World Health Organization (WHO) guidelines 37.
Model Performance and Comparison
The CBM achieved an AUC of 0.92 ± 0.01 in the test set for classifying early-stage BC and benign lesions, performing comparably to the black-box model (AUC: 0.93 ± 0.01; DeLong’s test P = 0.10; Table 2 and Additional file 2-Fig. S8A). P‑values from DeLong’s test for each fold of the 5‑fold cross‑validation are provided in Additional file 2-Fig. S9. The 95% CI for the difference in the AUC between the CBM and the black-box model was (-0.036, 0.002), with a t-statistic of -2.50 (P = 0.07). Across a wide range of threshold probabilities, the CBM exhibited a high net benefit in the decision curve analysis (Additional File 2-Fig. S8B). In the external cohort, the CBM maintained robust diagnostic performance, attaining an AUC of 0.93, along with an accuracy of 0.86, a recall of 0.84, a precision of 0.89, and an F1-score of 0.87. Detailed performance metrics for each participating center in the external cohort are presented in Additional file 2-Fig. S10. For conceptual prediction, the CBM achieved accuracies ranging from 0.64 to 1.00 in the test set, 0.58 to 1.00 in the external cohort, and 0.53 to 1.00 in the multi-reader cohort (Additional file 2-Table S2).
Table 2.
Performance of the models in differentiating early-stage breast cancer from benign lesions in the test set
| Model | Sequence | AUC | 95% CI | Accuracy | 95% CI | Recall | 95% CI | Precision | 95% CI | F1-score | 95% CI |
|---|---|---|---|---|---|---|---|---|---|---|---|
| CBM | ADC | 0.85 ± 0.02 | (0.82, 0.88) | 0.77 ± 0.02 | (0.75, 0.80) | 0.77 ± 0.07 | (0.68, 0.85) | 0.78 ± 0.03 | (0.74, 0.81) | 0.77 ± 0.03 | (0.73, 0.81) |
| CE | 0.90 ± 0.01 | (0.89, 0.92) | 0.83 ± 0.03 | (0.79, 0.86) | 0.81 ± 0.05 | (0.75, 0.88) | 0.84 ± 0.05 | (0.77, 0.90) | 0.82 ± 0.03 | (0.79, 0.86) | |
| DWI | 0.84 ± 0.01 | (0.82, 0.86) | 0.75 ± 0.02 | (0.73, 0.78) | 0.71 ± 0.07 | (0.63, 0.80) | 0.78 ± 0.05 | (0.72, 0.84) | 0.74 ± 0.03 | (0.71, 0.77) | |
| T2WI | 0.82 ± 0.02 | (0.80, 0.85) | 0.76 ± 0.01 | (0.75, 0.77) | 0.71 ± 0.04 | (0.66, 0.75) | 0.79 ± 0.03 | (0.76, 0.82) | 0.75 ± 0.01 | (0.73, 0.76) | |
| TIC | 0.75 ± 0.00 | (0.75, 0.75) | 0.69 ± 0.01 | (0.68, 0.70) | 0.80 ± 0.01 | (0.79, 0.81) | 0.65 ± 0.01 | (0.64, 0.67) | 0.72 ± 0.00 | (0.72, 0.72) | |
| Multi-parametric | 0.92 ± 0.01 | (0.90, 0.93) | 0.85 ± 0.01 | (0.83, 0.86) | 0.84 ± 0.02 | (0.82, 0.86) | 0.85 ± 0.02 | (0.82, 0.88) | 0.84 ± 0.01 | (0.83, 0.86) | |
| Black-box | ADC | 0.85 ± 0.01 | (0.84, 0.86) | 0.77 ± 0.01 | (0.75, 0.78) | 0.80 ± 0.05 | (0.74, 0.87) | 0.75 ± 0.03 | (0.71, 0.79) | 0.77 ± 0.01 | (0.76, 0.79) |
| CE | 0.90 ± 0.01 | (0.89, 0.92) | 0.81 ± 0.03 | (0.77, 0.84) | 0.78 ± 0.10 | (0.65, 0.91) | 0.83 ± 0.04 | (0.78, 0.88) | 0.80 ± 0.05 | (0.74, 0.86) | |
| DWI | 0.83 ± 0.02 | (0.80, 0.86) | 0.74 ± 0.02 | (0.72, 0.76) | 0.78 ± 0.05 | (0.72, 0.84) | 0.72 ± 0.03 | (0.69, 0.76) | 0.75 ± 0.02 | (0.73, 0.77) | |
| T2WI | 0.83 ± 0.02 | (0.80, 0.86) | 0.75 ± 0.03 | (0.72, 0.78) | 0.76 ± 0.04 | (0.71, 0.81) | 0.75 ± 0.05 | (0.69, 0.81) | 0.75 ± 0.02 | (0.73, 0.77) | |
| TIC | 0.75 ± 0.01 | (0.75, 0.76) | 0.69 ± 0.01 | (0.68, 0.70) | 0.79 ± 0.01 | (0.78, 0.80) | 0.66 ± 0.01 | (0.65, 0.67) | 0.72 ± 0.00 | (0.72, 0.72) | |
| Multi-parametric | 0.93 ± 0.01 | (0.92, 0.94) | 0.85 ± 0.01 | (0.83, 0.86) | 0.84 ± 0.04 | (0.79, 0.88) | 0.86 ± 0.02 | (0.83, 0.89) | 0.85 ± 0.01 | (0.83, 0.86) |
Abbreviation: CBM, Concept Bottleneck Model. Classification performance metrics for the multiparametric strategy are highlighted in bold. Recall (also known as sensitivity) is defined as the proportion of positive cases that are correctly identified, calculated as TP / (TP + FN), where TP and FN denote true positives and false negatives, respectively. Precision, also referred to as the positive predictive value (PPV), is defined as TP / (TP + FP), with FP representing false positives 38. The F1-score, which represents the harmonic mean of precision and recall (F1-score = 2 × precision × recall / [precision + recall]), penalizes imbalances between these two metrics 17
Among the 143 benign lesions in the test set, 68 (47.6%) were classified as BI-RADS category 4 by radiologists. Of these, the CBM correctly identified 50 (73.5%) as benign. A similar trend was observed in the external cohort: among the 54 (59.3%) benign lesions categorized as BI-RADS 4, the CBM accurately predicted 45 (83.3%) as benign.
The Impact of CBM on Radiologists’ Diagnostic Performance
In the multi-reader study, the CBM again exhibited high diagnostic accuracy (0.89 [58/65]) in distinguishing benign lesions from BC, performing comparably to one radiologist and outperforming the other seven (all P < 0.05; Fig. 3A and Additional file 2-Table S3). With CBM assistance, the diagnostic accuracy of the eight readers improved from a range of 0.71–0.79 to 0.77–0.91 (Fig. 3B), while their reading time significantly decreased (143 ± 107 vs. 104 ± 66 s, P < 0.001; Fig. 3C, D).
Fig. 3.
A Comparison of classification performance between the CBM and eight radiologists. B Diagnostic performance of radiologists without versus with CBM assistance, evaluated using McNemar’s test (dotted border: less-experienced radiologists; solid border: experienced radiologists). C Average reading time across all readers with and without CBM assistance; the difference was analyzed using the Wilcoxon signed‑rank test. D Box plot showing the reading time for each radiologist under both conditions. Gray lines connect the reading times for the same case without and with CBM assistance, and the paired differences were statistically assessed using the Wilcoxon signed‑rank test
The principal advantage of the CBM lies in improving the PPV (precision) of less‑experienced radiologists, as evidenced by the following findings: with CBM assistance, diagnostic sensitivity (recall) remained comparable in both groups, whereas overall accuracy increased (Fig. 3B). Notably, the less‑experienced group showed more pronounced improvements—accuracy increased from 0.74 to 0.88, and PPV rose from 0.65 to 0.81 (χ² = 5.82, P = 0.02). In contrast, the experienced group showed moderate gains, with accuracy improving from 0.77 to 0.82 and PPV from 0.69 to 0.74 (χ² = 0.57, P = 0.45).
The use of CBM led to adjustments in BI-RADS categories across all readers, changing management recommendations in 7 to 15 of the 65 reports (Table 3). Importantly, with CBM assistance, readers correctly reclassified 4 to 13 cases that had been initially misclassified, including 3 to 12 benign lesions that had been inappropriately assigned as BI-RADS 4 (see example in Fig. 4).
Table 3.
Changes in Radiologists’ BI-RADS Category Assignments Without vs. With CBM Assistance
| Outcome | Reader 1 | Reader 2 | Reader 3 | Reader 4 | Reader 5 | Reader 6 | Reader 7 | Reader 8 |
|---|---|---|---|---|---|---|---|---|
| Gwet’s AC1 (95% CI) * | 0.63 (0.47, 0.77) | 0.44 (0.28, 0.58) | 0.63 (0.47, 0.76) | 0.72 (0.58, 0.83) | 0.59 (0.43, 0.74) | 0.74 (0.60, 0.86) | 0.49 (0.34, 0.63) | 0.66 (0.52, 0.79) |
| No change in BIRADS category | 72.3% (47/65) | 55.4% (36/65) | 69.2% (45/65) | 76.9% (50/65) | 67.7% (44/65) | 78.5% (51/65) | 60.0% (39/65) | 72.3% (47/65) |
| Any change in BIRADS category | 27.7% (18/65) | 44.6% (29/65) | 30.8% (20/65) | 23.1% (15/65) | 32.3% (21/65) | 21.5% (14/65) | 40.0% (26/65) | 27.7% (18/65) |
| BIRADS category Downgraded | 94.4% (17/18) | 82.8% (24/29) | 35.0% (7/20) | 73.3% (11/15) | 42.9% (9/21) | 35.7% (5/14) | 61.5% (16/26) | 66.7% (12/18) |
| From BIRADS 4 to BIRADS 2/3 | 76.5% (13/17) | 58.3% (14/24) | 42.9% (3/7) | 90.9% (10/11) | 66.7% (6/9) | 80.0% (4/5) | 62.5% (10/16) | 58.3% (7/12) |
| From BIRADS 5 to BIRADS 2/3 | 0.0% (0/17) | 0.0% (0/24) | 0.0% (0/7) | 0.0% (0/11) | 0.0% (0/9) | 0.0% (0/5) | 0.0% (0/16) | 0.0% (0/12) |
| From BIRADS 3 to BIRADS 2 | 0.0% (0/17) | 4.2% (1/24) | 57.1% (4/7) | 9.1% (1/11) | 0.0% (0/9) | 20.0% (1/5) | 0.0% (0/16) | 0.0% (0/12) |
| From BIRADS 5 to BIRADS 4 | 23.5% (4/17) | 37.5% (9/24) | 0.0% (0/7) | 0.0% (0/11) | 33.3% (3/9) | 0.0% (0/5) | 37.5% (6/16) | 41.7% (5/12) |
| BIRADS category Upgraded | 5.6% (1/18) | 17.2% (5/29) | 65.0% (13/20) | 26.7% (4/15) | 57.1% (12/21) | 64.3% (9/14) | 38.5% (10/26) | 33.3% (6/18) |
| From BIRADS 2/3 to BIRADS 4 | 100.0% (1/1) | 20.0% (1/5) | 61.5% (8/13) | 75.0% (3/4) | 16.7% (2/12) | 33.3% (3/9) | 10.0% (1/10) | 16.7% (1/6) |
| From BIRADS 2/3 to BIRADS 5 | 0.0% (0/1) | 0.0% (0/5) | 0.0% (0/13) | 0.0% (0/4) | 0.0% (0/12) | 0.0% (0/9) | 0.0% (0/10) | 0.0% (0/6) |
| From BIRADS 2 to BIRADS 3 | 0.0% (0/1) | 40.0% (2/5) | 23.1% (3/13) | 0.0% (0/4) | 66.7% (8/12) | 55.6% (5/9) | 70.0% (7/10) | 16.7% (1/6) |
| From BIRADS 4 to BIRADS 5 | 0.0% (0/1) | 40.0% (2/5) | 15.4% (2/13) | 25.0% (1/4) | 16.7% (2/12) | 11.1% (1/9) | 20.0% (2/10) | 66.7% (4/6) |
| Comparison with Pathology | ||||||||
| Correct downgrading from BI-RADS 4 to 2/3 | 92.3% (12/13) | 85.7% (12/14) | 100.0% (3/3) | 100.0% (10/10) | 83.3% (5/6) | 75.0% (3/4) | 90.0% (9/10) | 85.7% (6/7) |
| Incorrect downgrading from BI-RADS 4 to 2/3 | 7.7% (1/13) | 14.3% (2/14) | 0.0% (0/3) | 0.0% (0/10) | 16.7% (1/6) | 25.0% (1/4) | 10.0% (1/10) | 14.3% (1/7) |
| Correct upgrading from BI-RADS 2/3 to 4 | 100.0% (1/1) | 100.0% (1/1) | 25.0% (2/8) | 33.3% (1/3) | 50.0% (1/2) | 33.3% (1/3) | 0.0% (0/1) | 100.0% (1/1) |
| Incorrect upgrading from BI-RADS 2/3 to 4 | 0.0% (0/1) | 0.0% (0/1) | 75.0% (6/8) | 66.7% (2/3) | 50.0% (1/2) | 67.7% (2/3) | 100.0% (1/1) | 0.0% (0/1) |
Abbreviation: BI-RADS, Breast Imaging Reporting and Data System. *, The inter-rater agreement among eight radiologists in assigning BI-RADS categories with and without CBM assistance was assessed by calculating Gwet’s Agreement Coefficient (AC1) and its 95% confidence interval (95% CI) 20. The results were interpreted according to the Landis and Koch benchmark scale: 0-0.20, slight agreement; 0.21–0.40, fair agreement; 0.41–0.60, moderate agreement; 0.61–0.80, substantial agreement; 0.81-1.00, almost perfect agreement 39
Fig. 4.
An example demonstrating the impact of CBM assistance on radiologists’ diagnostic decisions. A mass was identified in the upper inner quadrant of the left breast of a 51-year-old female. The lesion appeared isointense on T2WI (A), showed no restricted diffusion (B, C), and displayed an irregular shape with uneven margins along with heterogeneous internal enhancement (D). The time-intensity curve (TIC) exhibited a rapid initial rise followed by a persistent pattern (E). Without CBM assistance, all eight radiologists classified the lesion as BI-RADS category 4. With the aid of structured report generated by CBM (F), six radiologists downgraded the assessment to BI-RADS category 3, while two maintained a BI-RADS category 4 rating. Pathological examination confirmed a fibroadenoma with adenosis
In parallel, with CBM support, the eight readers each misclassified 1 to 6 cases. An analysis of error patterns revealed that: among the three cases that radiologists incorrectly downgraded from BI-RADS 4 to BI-RADS 2/3, two were also misclassified by CBM (Fig. 5A); among the ten cases that radiologists erroneously upgraded from BI-RADS 2/3 to BI-RADS 4, only one was misclassified by CBM (Fig. 5B). Representative examples are illustrated in Fig. 5C, D.
Fig. 5.
Analysis of radiologists’ error patterns with CBM assistance. A Incorrect downgrading from BI-RADS category 4 to category 2/3. B Incorrect upgrading from BI-RADS category 2/3 to category 4. C Two cases initially assessed as BI-RADS 4 were erroneously downgraded to BI-RADS 3 with CBM support. MRI revealed irregular masses in the left breast (maximum diameters: 1.1 cm and 0.8 cm) with ill‑defined margins and heterogeneous enhancement. DWI showed no significant restriction, and TIC analysis displayed a plateau pattern. Pathology confirmed invasive carcinoma in both cases. D In a 39‑year‑old woman with left breast fibroadenoma, MRI identified an oval mass in the outer quadrant with circumscribed margins, high signal intensity on T2-weighted imaging, and restricted diffusion on DWI. Despite a CBM‑estimated malignancy probability of 8.9%, Reader 4 incorrectly upgraded the assessment from BI‑RADS 3 to 4 under CBM guidance
Furthermore, inter-reader agreement in concept recognition and BI-RADS classification, measured by Gwet’s AC1, improved from a range of 0.27–1.00 to 0.46–1.00 (Additional file 2-Table S4). The accuracy of concept recognition also increased, from 0.61 to 1.00 to 0.74–1.00 (Additional file 2-Figs. S11, S12).
Comparison of Multi- and Single-Sequence Strategies and Concept Importance Analysis
t-SNE visualization of feature representations extracted by the sequence-specific encoder revealed partial overlap in the distributions of benign and malignant lesions under single-sequence conditions. In contrast, the multiparametric strategy improved class separation, yielding tighter and less dispersed clusters (Fig. 6A). Moreover, both the CBM and the black-box model exhibited inferior performance in classifying benign versus malignant lesions using a single MRI sequence, with AUC values ranging from 0.75 ± 0.00 to 0.90 ± 0.01, significantly lower than those achieved with the multiparametric strategy (all P < 0.05; Table 2; Fig. 6B).
Fig. 6.
A Comparison between the multiparametric MRI strategy and the single-sequence approaches. The t-SNE visualization reveals that the multiparametric method (denoted as MM) achieves a more distinct separation between benign and malignant breast lesions than any of the single-sequence methods. B DeLong’s test heatmap for pairwise comparisons of AUC values. The comparisons involve both the single-sequence and multiparametric strategies, each implemented with either the CBM or the black-box model. C Heatmap of concept weights, indicating the contribution of each concept to the discrimination between benign and malignant breast lesions
Concept importance analysis identified “restricted diffusion”, “spiculated margin”, and the “wash-out” pattern on TIC as the most discriminative semantic features for diagnosing malignancies, whereas homogeneous enhancement, persistent pattern on the delayed phase of TIC, and dark internal septations were the top three features most indicative of benign lesions (Fig. 6C). Correlation analysis based on these six concepts demonstrated statistically significant moderate-to-large associations between the visual patterns identified by the CBM and the corresponding radiological semantic features across all validation datasets (test set, external cohort, and multi-reader cohort), with Cramer’s V values ranging from 0.31 to 0.60 (Additional file 2-Fig. S13). An example of Grad-CAM visualization is provided to illustrate the spatial localization of the concepts recognized by the CBM within the images (Additional file 2-Fig. S14).
Discussion
This multicenter study developed an MRI-based framework (CBM) for classifying early-stage breast cancer and benign lesions while concurrently generating structured reports. Its core innovation lies in aligning multiparametric image features with textual concepts derived from radiological reports, thereby enhancing clinical interpretability. By incorporating expert-identified semantic features as prior knowledge, augmented with GPT-5 embedding, the CBM outputs radiologist-intelligible structured reports that include BI-RADS-compliant lesion descriptors, malignancy probabilities, and management recommendations. The model achieved promising classification performance (AUC: 0.92 ± 0.01, PPV: 0.85 ± 0.02, F1-score: 0.84 ± 0.01), comparable to that of state-of-the-art end-to-end black-box model. In our retrospective cohort, the CBM also demonstrated potential to reduce unnecessary biopsies for benign lesions initially overclassified as BI-RADS category 4 by less experienced radiologists, while improving both inter-reader consistency and interpretation efficiency in breast MRI analysis. Given the increasing adoption of structured reporting systems such as LI-RADS and PI-RADS in radiology, our approach provides a versatile framework that can be extended to other organs.
In terms of classification performance, the CBM achieves results comparable to those of earlier approaches, as evidenced by previous meta-analyses reporting AUC values ranging from 0.87 to 0.90 for MRI-based DL in breast cancer diagnosis 40, 41. While prior studies predominantly relied on single post-contrast sequence, multiparametric MRI better captures the diverse biological processes of the tumor microenvironment 13, 42, 43. Supporting this notion, t-SNE visualization revealed that a multiparametric strategy improved the distinction between benign and malignant lesions compared with single‑sequence approaches. This finding is further corroborated by earlier evidence and our subsequent analysis, indicating that multi-sequence fusion outperformed CE-MRI alone in breast lesion classification 44. The key semantic feature for distinguishing benign from malignant breast lesions, as identified by our concept importance analysis, reinforces established pathological knowledge and aligns with previous findings 1, 2, 43.
Understanding the biological meaning of high-level features derived from DL remains challenging 45. By integrating structured concepts from MRI reports with multiparametric images and incorporating radiologists’ domain knowledge to guide image representation learning, the CBM constrains feature vectors within a semantically relevant space, establishing a lesion-level mapping between computer vision representations and clinically relevant textual concepts. Ingeniously, the CBM first predicts clinically meaningful concepts (e.g., mass margin and internal enhancement characteristics) directly from input images, and then linearly combines these concepts to infer the probability of malignancy. This effectively simulates the reasoning process that radiologists follow when interpreting MRI scans. The entire decision-making process is fully transparent: the generated report explicitly lists the model’s intermediate predictions (i.e., concepts), enabling radiologists to intuitively discern which concepts are activated in the input image. Without compromising performance compared to uninterpretable black-box models, this framework partially bridges the gap between sophisticated model behavior and human-interpretable decision-making 9. To our knowledge, this approach represents the first application of its kind for classifying benign and malignant breast lesions using MRI.
Radiology reports function as an essential communication medium between radiologists and referring clinicians, documenting critical observations and evidence that support medical decisions 46. Compared with traditional narrative reports, structured reports are preferred by clinicians for their clarity, ability to facilitate clinical decisions, and potential benefits in workflow optimization 22, 47, 48. By leveraging GPT-5, a model known for its reduced tendency to hallucinate, to convert standardized BI-RADS-aligned concepts predicted by the CBM into narrative content that follows established formats and is comprehensible to radiologists, the CBM generates evidence-based structured reports 49. In this process, GPT-5 serves solely as an auxiliary tool to convert the highly structured concept outputs (Additional file 2-Fig. S15) into a templated report format, thereby minimizing the risk of hallucination. This was further supported by a factual consistency evaluation of GPT-5 described in Additional file 1-Supplementary methods and Additional file 2-Fig. S16, in which the participating radiologists rated the resulting structured reports highly 49, 50. Given the high accuracy of most concepts predicted by the CBM, along with their demonstrated potential to improve inter-reader agreement and accuracy in concept recognition and BI-RADS categories, as shown in our multi-reader study, the elements of these structured reports may serve a dual purpose: as rationales for breast lesion classification decisions and as valuable educational resources for radiology trainees.
Spanning a broad range of malignancy probabilities, BI-RADS category 4 lesions account for the majority of false-positive findings on breast MRI 51. Current guidelines recommend tissue diagnosis for these lesions, yet many prove pathologically benign. Frequent false-positives not only impose a psychological burden on patients but also reflect inefficient use of healthcare resources from a health-economics perspective. A well-trained CAD system capable of effectively reducing unnecessary biopsies of benign lesions would thus constitute a major clinical advancement. Focusing on benign lesions that were overclassified as BI-RADS 4 by radiologists in both the test set and the external cohort, we observed that the CBM correctly identified a large number of these lesions as benign. This finding was further corroborated in the multi-reader study: with CBM assistance, radiologists appropriately reclassified an average of 7.5 out of 34 benign lesions (22.1%) from BI-RADS category 4 to category 2 or 3, highlighting its potential to reduce unnecessary biopsies in patients with benign breast lesions. Meanwhile, it should be noted that for very small lesions lacking typical malignant features (as shown in Fig. 5C), the CBM may erroneously lead radiologists to classify them as benign, which represents an important direction for future optimization of the CBM.
The proportion of DCIS in our cohort was lower than the epidemiologically reported range of 15%–25%, a finding consistent with a previous study 52. MRI is primarily used in cases with clinical suspicion of more extensive disease, whereas most DCIS are typically detected by characteristic microcalcifications on screening mammography and subsequently confirmed by biopsy 2, 53. Furthermore, the exclusion of lesions presenting as NME, which is a common manifestation of DCIS but is difficult to correlate precisely between pathology and MRI, also contributed to the observed lower proportion 2. Additionally, the median age of malignant cases in this study was 48 years, which is lower than that reported in some studies but aligns with the regional epidemiological profile, where breast cancer incidence peaks between 45 and 54 years of age 54.
Our study has several limitations. First, the limited size of the external dataset necessitates further validation in a larger, preferably prospective, cohort. To this end, we have registered a clinical trial at the Chinese Clinical Trial Registry (Registration No. ChiCTR2500103426) to collect additional data for more extensive validation. Second, while the CBM was developed on a histopathology-confirmed cohort, clinical practice involves negative MRIs and a wider spectrum of benign lesions—some of which are confirmed through follow-up rather than biopsy. Thus, the real-world performance of the CBM, particularly its false-positive rate, requires validation in a more representative dataset covering all relevant clinical scenarios. Additionally, given that the primary task of this study was breast lesion classification, the current model relies on lesion location bounding boxes provided by radiologists; consequently, it cannot accurately estimate false positive detections. This makes the model more suitable for diagnostic settings where suspicious lesions have already been identified, rather than for screening scenarios. Third, this study did not encompass all elements specified in the BI-RADS lexicon (e.g., the amount of fibroglandular tissue), as their extraction would involve segmentation tasks beyond the core objectives of this study. Fourth, image-concept alignment was achieved at the whole-lesion level rather than through pixel-wise spatial mapping; future work will focus on exploring pixel-level alignment methods. Moreover, the image‑to‑concept mapping relies on a deep visual encoder whose internal computations are not fully transparent at the pixel level—a challenge that is widely shared in current DL-based medical imaging research. Therefore, the interpretability contribution of the CBM lies in rendering the decision-making stage transparent and amenable to clinical review, a stage that bears the most direct relevance to clinical trust and radiologist intervention. Last but not least, although the multi-reader study provides preliminary evidence for the clinical utility of the CBM, it was conducted in a simulated setting where radiologists only saw the structured reports generated by the CBM, without being informed of which concepts drove the malignancy prediction or the contribution weights of individual concepts. Future work will involve deploying the model into clinical workflows and developing a user-friendly interface that clearly displays the Top-k activated concepts along with their confidence scores, while also allowing radiologists to interactively adjust the model’s output during inference by manually correcting concept predictions (Additional file 1-Supplementary methods and Additional file 2-Fig. S17). Through this “radiologist-in-the-loop” strategy, we aim to progressively improve and validate the model’s effectiveness and reliability in real-world scenarios.
Conclusions
In conclusion, the CBM enhances clinical interpretability through image-concept alignment and effectively classifies early breast cancer and benign lesions. It improves radiologists’ diagnostic performance, particularly among less-experienced readers, increases reporting consistency, and shows potential to reduce unnecessary biopsies for benign breast lesions.
Supplementary Information
Below is the link to the electronic supplementary material.
Supplementary Material 1: Additional file 1: Supplementary methods
Supplementary Material 2: Additional file 2: Tables S1-S4 and Figures S1-S17
Acknowledgements
Not applicable.
Abbreviations
- ADC
apparent diffusion coefficient
- AUC
area under the receiver operating characteristic curve
- BC
breast cancer
- BI-RADS
Breast Imaging Reporting and Data System
- CAM
class activation mapping
- CBM
concept bottleneck model
- CE
contrast-enhanced T1-weighted imaging
- DL
deep learning
- DWI
diffusion-weighted imaging
- MLLM
multimodal large language model
- PPV
positive predictive value
- ROI
region of interest
- SHAP
Shapley Additive exPlanations
- T2WI
T2-weighted imaging
- TIC
time‒intensity curve
Author contributions
Conceptualization: JQ, YL1, HS, JJ2, SG, SL; Literature research: JQ, YL1; Data acquisition: JQ, MZ, HX, JY, LZ, JZ1; Methodology: JQ, YL1; Statistical analysis: JQ, YL1; Data analysis and interpretation: JQ, YL1, MZ, HX, JY, MX, LY, WQ, YL2, JJ1, JH; Algorithm implementation: YL1, DT, JZ2, JS; Funding acquisition: JQ, JJ2, SL; Manuscript drafting: JQ, YL1; All authors read and approved the final manuscript.
Funding
This study was supported by (1) Noncommunicable Chronic Diseases-National Science and Technology Major Project (No. 2023ZD0502300 [2023ZD0502304]); (2) National Guidance Fund on Developing Local Science and Technology for Sichuan Province (No. 2023ZYD0167); (3) National Natural Science Foundation of China (Project Nos. 82441007, 82120108014); (4) 1.3.5 project for disciplines of excellence, West China Hospital, Sichuan University (Project No. ZYGD23003); (5) Sichuan Science and Technology Program (Grant Nos. 2026NSFSC1775); (6) Chengdu Science and Technology Office, major technology application demonstration project (Project No. 2022-GH03-00017-HZ); and (7) the Postdoctor Research Fund of West China Hospital, Sichuan University (Project No. 2025HXBH040).
Data availability
The data supporting the findings of this study cannot be made publicly available due to restrictions in the data use agreement and concerns regarding patient privacy. The original data are available from the corresponding author upon reasonable request, subject to data sharing agreements that comply with institutional and national regulations. The code used in this work is publicly available at ( *https://github.com/ly1998117/MammoCBM* ).
Declarations
Ethics approval and consent to participate
This study was approved by the Ethics Committee on Biomedical Research, West China Hospital of Sichuan University (approval No. 2025 − 949). Informed consent was waived given the retrospective design of the study.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Jiao Qu and Yang Liu contributed equally to this work.
Contributor Information
Jing Jing, Email: jingjing@wchscu.edu.cn.
Hubing Shi, Email: shihb@scu.edu.cn.
Shi Gu, Email: gus@zju.edu.cn.
Su Lui, Email: lusuwcums@hotmail.com.
References
- 1.Kuhl C. The current status of breast MR imaging. Part I. Choice of technique, image interpretation, diagnostic accuracy, and transfer to clinical practice. Radiology. 2007;244(2):356–78. [DOI] [PubMed] [Google Scholar]
- 2.Mann RM, Cho N, Moy L, Breast MRI. State of the Art. Radiology. 2019;292(3):520–36. [DOI] [PubMed] [Google Scholar]
- 3.Saadatmand S, Geuzinge HA, Rutgers EJT, et al. MRI versus mammography for breast cancer screening in women with familial risk (FaMRIsc): a multicentre, randomised, controlled trial. Lancet Oncol. 2019;20(8):1136–47. [DOI] [PubMed] [Google Scholar]
- 4.Verburg E, van Gils CH, van der Velden BHM, et al. Validation of Combined Deep Learning Triaging and Computer-Aided Diagnosis in 2901 Breast MRI Examinations From the Second Screening Round of the Dense Tissue and Early Breast Neoplasm Screening Trial. Invest Radiol. 2023;58(4):293–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Oviedo F, Kazerouni AS, Liznerski P, et al. Cancer Detection in Breast MRI Screening via Explainable AI Anomaly Detection. Radiology. 2025;316(1):e241629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Witowski J, Heacock L, Reig B, et al. Improving breast cancer diagnostics with deep learning for MRI. Sci Transl Med. 2022;14(664):eabo4802. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Dalmiş MU, Gubern-Mérida A, Vreemann S, et al. Artificial Intelligence–Based Classification of Breast Lesions Imaged With a Multiparametric Breast MRI Protocol With Ultrafast DCE-MRI, T2, and DWI. Invest Radiol. 2019;54(6):325–32. [DOI] [PubMed] [Google Scholar]
- 8.Luo L, Wu M, Li M, et al. A large model for non-invasive and personalized management of breast cancer from multiparametric MRI. Nat Commun. 2025;16(1):3647. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bilal A, Ebert D, Lin B. LLMs for Explainable AI: A Comprehensive Survey. arXiv:250400125. 2025.
- 10.Gastounioti A, Kontos D. Is It Time to Get Rid of Black Boxes and Cultivate Trust in AI? Radiol Artif Intell. 2020;2(3):e200088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.D'Orsi CJ, Sickles EA, Mendelson EB, Morris EA. ACR BI-RADS Atlas, Breast Imaging Reporting and Data System. 5th ed. Reston, Va: American College of Radiology, 2013.
- 12.An S, Teo K, McConnell MV, Marshall J, Galloway C, Squirrell D. AI explainability in oculomics: How it works, its role in establishing trust, and what still needs to be addressed. Prog Retin Eye Res. 2025;106:101352. [DOI] [PubMed] [Google Scholar]
- 13.Adam R, Dell’Aquila K, Hodges L, Maldjian T, Duong TQ. Deep learning applications to breast cancer detection by magnetic resonance imaging: a literature review. Breast Cancer Res. 2023;25(1):87. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Koh PW, Nguyen T, Tang YS, et al. Concept Bottleneck Models. 2020. 10.48550/arXiv.2007.04612. [Google Scholar]
- 15.Wu Y, Liu Y, Yang Y, et al. A concept-based interpretable model for the diagnosis of choroid neoplasias using multimodal data. Nat Commun. 2025;16(1):3504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Adams LC, Truhn D, Busch F, et al. Leveraging GPT-4 for Post Hoc Transformation of Free-text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology. 2023;307(4):e230725. [DOI] [PubMed] [Google Scholar]
- 17.Benary M, Wang XD, Schmidt M, et al. Leveraging Large Language Models for Decision Support in Personalized Oncology. JAMA Netw Open. 2023;6(11):e2343689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Bhayana R, Nanda B, Dehkharghanian T, et al. Large Language Models for Automated Synoptic Reports and Resectability Categorization in Pancreatic Cancer. Radiology. 2024;311(3):e233117. [DOI] [PubMed] [Google Scholar]
- 19.Schiaffino S, Zhang T, Mann RM, Pinker K. The Role of Large Language Models (LLMs) in Breast Imaging Today and in the Near Future. J Magn Reson Imaging. 2025;62(5):1296–304. [DOI] [PubMed] [Google Scholar]
- 20.Cozzi A, Pinker K, Hidber A, et al. BI-RADS Category Assignments by GPT-3.5, GPT-4, and Google Bard: A Multilanguage Study. Radiology. 2024;311(1):e232133. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wu SH, Tong WJ, Li MD, et al. Collaborative Enhancement of Consistency and Accuracy in US Diagnosis of Thyroid Nodules Using Large Language Models. Radiology. 2024;310(3):e232255. [DOI] [PubMed] [Google Scholar]
- 22.Busch F, Hoffmann L, Dos Santos DP, et al. Large language models for structured reporting in radiology: past, present, and future. Eur Radiol. 2025;35(5):2589–602. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Das A, Talati IA, Chaves JMZ, Rubin D, Banerjee I. Weakly supervised language models for automated extraction of critical findings from radiology reports. NPJ Digit Med. 2025;8(1):257. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Pescia C, Guerini-Rocco E, Viale G, Fusco N. Advances in Early Breast Cancer Risk Profiling: From Histopathology to Molecular Technologies. Cancers (Basel). 2023;15(22):5430. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kim GR, Choi JS, Han B-K, et al. Preoperative Axillary US in Early-Stage Breast Cancer: Potential to Prevent Unnecessary Axillary Lymph Node Dissection. Radiology. 2018;288(1):55–63. [DOI] [PubMed] [Google Scholar]
- 27.Truhn D, Schrading S, Haarburger C, Schneider H, Merhof D, Kuhl C. Radiomic versus Convolutional Neural Networks Analysis for Classification of Contrast-enhancing Lesions at Multiparametric Breast MRI. Radiology. 2019;290(2):290–7. [DOI] [PubMed] [Google Scholar]
- 28.Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29–36. [DOI] [PubMed] [Google Scholar]
- 29.Hosny A, Parmar C, Quackenbush J, Schwartz LH, Aerts HJWL. Artificial intelligence in radiology. Nat Rev Cancer. 2018;18(8):500–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhou J, Zhang Y, Chang KT, et al. Diagnosis of Benign and Malignant Breast Lesions on DCE-MRI by Using Radiomics and Deep Learning With Consideration of Peritumor Tissue. J Magn Reson Imaging. 2019;51(3):798–809. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Choi SY, Min JH, Kim JH, et al. Interobserver Variability and Diagnostic Performance in Predicting Malignancy of Pancreatic Intraductal Papillary Mucinous Neoplasm with MRI. Radiology. 2023;308(1):e222463. [DOI] [PubMed] [Google Scholar]
- 32.Romeo V, Clauser P, Rasul S, et al. AI-enhanced simultaneous multiparametric 18F-FDG PET/MRI for accurate breast cancer diagnosis. Eur J Nucl Med Mol Imaging. 2021;49(2):596–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.van der Maaten L, Hinton G. Visualizing Data using t-SNE. J Mach Learn Res. 2008;9:2579–605. [Google Scholar]
- 34.Liu Y, Zhang T, Gu S, Hybrid Concept Bottleneck M. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2025.
- 35.McGillivray B, De Ranieri E. Uptake and outcome of manuscripts in Nature journals by review model and author characteristics. Res Integr Peer Rev. 2018;3:5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Cramér’s V and Measures of Association for Nominal Variables. R Handbook. https://rcompanion.org/handbook/H_10.html
- 37.WHO Classification of Tumours. https://publicationsiarcwhoint/Book-And-Report-Series/Who-Classification-Of-Tumours/Breast-Tumours-2019
- 38.Hirsch L, Huang Y, Makse HA, et al. Early Detection of Breast Cancer in MRI Using AI. Acad Radiol. 2025;32(3):1218–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–74. [PubMed] [Google Scholar]
- 40.Aggarwal R, Sounderajah V, Martin G, et al. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. npj Digit Med. 2021;4(1):65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Abdullah KA, Marziali S, Nanaa M, Escudero Sanchez L, Payne NR, Gilbert FJ. Deep learning-based breast cancer diagnosis in breast MRI: systematic review and meta-analysis. Eur Radiol. 2025;35(8):4474–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Chen Y, Wang L, Luo R, et al. A deep learning model based on dynamic contrast-enhanced magnetic resonance imaging enables accurate prediction of benign and malignant breast lessons. Front Oncol. 2022;12:943415. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Hoffmann E, Masthoff M, Kunz WG, et al. Multiparametric MRI for characterization of the tumour microenvironment. Nat Rev Clin Oncol. 2024;21(6):428–48. [DOI] [PubMed] [Google Scholar]
- 44.Hu Q, Whitney HM, Giger ML. A deep learning methodology for improved breast cancer diagnosis using multiparametric MRI. Sci Rep. 2020;10(1):10536. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Howard FM, Hieromnimon HM, Ramesh S, et al. Generative adversarial networks accurately reconstruct pan-cancer histology from pathologic, genomic, and radiographic latent features. Sci Adv. 2024;10(46):eadq0856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Doshi R, Amin KS, Khosla P, Bajaj S, Chheang S, Forman HP. Quantitative Evaluation of Large Language Models to Streamline Radiology Report Impressions: A Multimodal Retrospective Analysis. Radiology. 2024;310(3):e231593. [DOI] [PubMed] [Google Scholar]
- 47.Brook OR, Brook A, Vollmer CM, Kent TS, Sanchez N, Pedrosa I. Structured reporting of multiphasic CT for pancreatic cancer: potential effect on staging and surgical planning. Radiology. 2015;274(2):464–72. [DOI] [PubMed] [Google Scholar]
- 48.Brady A, Brink J, Slavotinek J. Radiology and Value-Based Health Care. JAMA. 2020;324(13):1286–7. [DOI] [PubMed] [Google Scholar]
- 49.Handler R, Sharma S, Hernandez-Boussard T. The fragile intelligence of GPT-5 in medicine. Nat Med. 2025;31(12):3968–70. [DOI] [PubMed] [Google Scholar]
- 50.Zheng S, Zhao N, Wang J, et al. Comparison of a Specialized Large Language Model with GPT-4o for CT and MRI Radiology Report Summarization. Radiology. 2025;316(2):e243774. [DOI] [PubMed] [Google Scholar]
- 51.Daimiel Naranjo I, Gibbs P, Reiner JS, et al. Radiomics and Machine Learning with Multiparametric Breast MRI for Improved Diagnostic Accuracy in Breast Cancer Diagnosis. Diagnostics (Basel). 2021;11(6):919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Hu Q, Whitney HM, Li H, Ji Y, Liu P, Giger ML. Improved Classification of Benign and Malignant Breast Lesions Using Deep Feature Maximum Intensity Projection MRI in Breast Cancer Diagnosis Using Dynamic Contrast-enhanced MRI. Radiol Artif Intell. 2021;3(3):e200159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Delaloge S, Khan SA, Wesseling J, Whelan T. Ductal carcinoma in situ of the breast: finding the balance between overtreatment and undertreatment. Lancet. 2024;403(10445):2734–46. [DOI] [PubMed] [Google Scholar]
- 54.The Society of Breast Cancer China A-CA. Guidelines for breast cancer diagnosis and treatment by China Anticancer Association (2026 edition). CHINA Oncol. 2025;35(12):9–104. Breast Oncology Group of the Oncology Branch of the Chinese Medical Association. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Material 1: Additional file 1: Supplementary methods
Supplementary Material 2: Additional file 2: Tables S1-S4 and Figures S1-S17
Data Availability Statement
The data supporting the findings of this study cannot be made publicly available due to restrictions in the data use agreement and concerns regarding patient privacy. The original data are available from the corresponding author upon reasonable request, subject to data sharing agreements that comply with institutional and national regulations. The code used in this work is publicly available at ( *https://github.com/ly1998117/MammoCBM* ).






