Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 15.
Published in final edited form as: J Am Coll Radiol. 2026 May 2;23(8):1579–1583. doi: 10.1016/j.jacr.2026.04.023

Using Artificial Intelligence to Improve Timeliness of Follow-Up in Breast Cancer Screening

Diana L Miglioretti a,b, Matt Ponzini a, Ojas A Ramwala c,d, Evan de Bie a, Shadi Aminololama-Shakeri e, Elizabeth Morris e, Christoph I Lee c
PMCID: PMC13265059  NIHMSID: NIHMS2182827  PMID: 42082064

Brief description of the problem

Regular screening mammography reduces breast cancer mortality through early detection and diagnosis; however, in the United States, more than 3.8 million individuals are recalled each year for additional evaluation of suspicious screening findings.1 Because diagnostic evaluation typically requires at least one additional visit, recalls can lead to delays in diagnostic resolution, anxiety, financial and opportunity costs, and loss to follow-up.2 These challenges disproportionately affect under-resourced and rural populations, who are more likely to experience longer diagnostic delays or not return for diagnostic work-up.3

One approach to reduce diagnostic delays is immediate interpretation of screening examinations while patients wait, enabling same-day diagnostic imaging when needed.4 Immediate interpretation of all screening examinations is typically infeasible due to workflow constraints; therefore, strategies that prioritize examinations most likely to be recalled for immediate interpretation with same-day diagnostic evaluation when indicated may reduce the number of patients who must return for a second visit. Artificial intelligence (AI) algorithms applied to screening mammography show promise for identifying examinations with suspicious findings and could support targeted immediate interpretation workflows.5,6

University of California (UC), Davis recently launched a mobile mammography program to increase access to screening in underserved communities. In this program, we observed substantial delays in diagnostic evaluation following positive screening results, underscoring the need for approaches that enable more timely diagnostic work-up. Implementing immediate interpretation on a mobile unit is particularly challenging because of limitations in transmitting images to interpreting radiologists and in coordinating real-time diagnostic mammography with remote radiologist support. A potential solution is to use AI to flag examinations with high malignancy suspicion scores for priority upload and immediate remote interpretation. When indicated, diagnostic mammography could be completed on the van with virtual radiologist support (e.g., using Philips Lumify portable ultrasound) or by providing free transportation to the main imaging facility.

What the authors did

To evaluate the potential of AI for identifying screening examinations requiring additional diagnostic work-up and those that result in a breast cancer diagnosis, we identified 3,535 consecutive screening digital breast tomosynthesis examinations performed at UC Davis between 10/17/2022 and 12/31/2022, allowing at least 1 year of complete cancer capture via linkage with the California Cancer Registry. Of these, 55 examinations were excluded because participants opted out of research. For the remaining 3,480 examinations, we applied two FDA-cleared AI algorithms (iCAD Profound version 0.2.0.2 and Lunit INSIGHT DBT v1.0.0.2) using the ClinValAI framework7 to generate examination-level malignancy suspicion scores. Algorithm identities were masked as AI1 and AI2 due to agreements with the AI vendors who supplied their algorithms for research use. Of the 3,480 examinations, 12 were excluded because neither algorithm could generate a score, leaving 3,468 examinations for analysis. Among these 3,468 examinations, 4 were scored by AI2 but not AI1, and 3 were scored by AI1 but not AI2.

Examination-level malignancy suspicion scores (0–100 raw scores) were standardized to percentile ranks within the study sample to support vendor anonymity and comparability and were linked to patient characteristics and outcomes from the Sacramento Area Breast Imaging Registry (SABIR). SABIR links to the California Cancer Registry for complete capture of breast cancer outcomes, ensuring accurate ground truth for examination-level AI algorithm performance.

We evaluated distributions of AI examination-level malignancy suspicion scores by recall and screen-detected cancer status and computed areas under the ROC curve (AUC) for identifying recalls, screen-detected cancers, and cancers diagnosed within 1 year.

For each of the two AI scores, we calculated cumulative percentages of recalls and screen-detected cancers captured at thresholds corresponding to the top 2%, 5%, 10%, 15%, 20%, 25%, 50%, and 75% of examinations after ranking examinations from highest to lowest score; 100% represented all examinations. For each threshold, we calculated efficiency ratios as the percentage of recalls or screen-detected cancers captured above the threshold divided by the percentage of examinations above the threshold, with a value of 1 indicating performance equivalent to random selection. The thresholds evaluated represent exploratory workflow strategies and may not correspond directly to FDA-cleared operating thresholds or intended use labeling. We conducted all analyses using R version 4.5.2.

Outcomes and limitations

Patient characteristics reflected a typical screening population (Table 1).1 Of 3,468 screening mammograms, 213 (6.1%) were recalled. Within 1 year of screening, 25 breast cancers were diagnosed (7.2/1,000): 20 were detected following a positive screening mammogram (cancer detection rate 5.8/1,000), and 5 were diagnosed within 1 year of a negative screening mammogram (false-negative rate 1.4/1,000).

Table 1:

Characteristics of study sample of 3,468 screening mammograms.

Overall Recall Rate, % Cancer Detection Rate, per 1000 Cancer Rate, per 1000
Characteristic N (Column %)

N (row %) 3,468 6.1% 5.8 7.2
Age, years
 <40 30 (1%) 13.3% 0.0 0.0
 40–49 639 (18%) 10.6% 4.7 7.8
 50–59 856 (25%) 4.9% 4.7 5.8
 60–69 1,086 (31%) 4.3% 4.6 6.4
 70+ 857 (25%) 6.1% 9.3 9.3
Race and ethnicity*
 Hispanic/Latina 340 (10%) 7.6% 2.9 2.9
 Asian/Hawaiian/Pacific Islander 400 (12%) 4.5% 7.5 10.0
 Black 180 (5%) 5.6% 5.6 5.6
 White 2,133 (62%) 6.2% 6.1 7.5
 Multiracial/Other/Unknown 415 (12%) 6.5% 4.8 7.2
BI-RADS breast density
 Almost entirely fatty 680 (20%) 4.6% 5.9 5.9
 Scattered fibroglandular densities 1,265 (37%) 5.5% 5.5 6.3
 Heterogeneously dense 1,303 (38%) 6.7% 6.1 9.2
 Extremely dense 220 (6%) 11.4% 4.5 4.5
Personal history of breast cancer
 Yes 262 (8%) 5.0% 11.5 19.1
 No/None recorded 3,206 (92%) 6.2% 5.3 6.2
First degree family history of breast cancer
 Yes 580 (17%) 6.7% 10.3 15.5
 No/None recorded 2,888 (83%) 6.0% 4.8 5.5
Time since prior mammogram
 No previous mammogram 334 (10%) 14.1% 6.0 6.0
 <11 months 60 (2%) 8.3% 0.0 0.0
 1 year (11–18 months) 2,049 (59%) 4.1% 3.4 4.9
 2 years (19–30 months) 556 (16%) 5.4% 7.2 9.0
 3+ years (31+ months) 469 (14%) 9.8% 14.9 17.1
*

Racial groups shown for those who self-identified as non-Hispanic

The two AI algorithms performed similarly, with modest discrimination for identifying mammograms recalled for additional imaging (AUC AI1, AI2: 0.676, 0.670) and excellent discrimination for identifying mammograms with screen-detected cancer (AUC AI1, AI2: 0.952, 0.908) and cancers diagnosed within 1 year (AUC AI1, AI2: 0.852, 0.907) (Figure 1). Mammograms in the highest 2% of AI scores captured 7.5%–9.0% of all recalls (efficiency ratio, 3.9–4.5) and 50% of all screen-detected cancers (efficiency ratio, 25.1–25.9) (Table 2). Within this top 2% of mammograms, the recall rate was 23.9%–27.5% and the cancer detection rate was 14.5%–14.9%, compared with an overall recall rate of 6.1% and an overall cancer detection rate of 0.6% in the full sample. Decreasing the AI-score threshold for immediate interpretation captured a larger share of recalls and screen-detected cancers, but efficiency ratios decreased (Table 2). For example, the highest 5% of AI scores accounted for 14.6%–16.5% of recalls (efficiency ratio 2.9–3.3) and 60% of screen-detected cancers (efficiency ratio 12.0–12.2). The highest 10% accounted for 25.8%–28.8% of recalls (efficiency ratio 2.7–2.9) and 75%–90% of screen-detected cancers (efficiency ratio 7.8–9.0). Capturing all screen-detected cancers would require lowering the threshold to include 36%–58% of mammograms, which would also capture 57–77% of recalls.

Figure 1.

Figure 1.

Box plots of AI score distributions by AI model, recall status, screen-detected cancer status, and cancer status; and ROC curves assessing the discriminatory accuracy of the AI score for predicting recalls, detected cancers, and cancers within 1 year.

Table 2.

Percentage of recalls for additional imaging and screen-detected cancers, efficiency ratios, and recall and cancer detection rates for different percentages of exams ranked by highest to lowest AI score. Number pairs are AI1, AI2.

Recalls for Additional Imaging Screen-detected Cancers
Target % of Exams Cumulative Number of Exams Recalls (%) Efficiency Ratio Recall Rate (%) Detected cancers (%) Efficiency Ratio Cancer Detection Rate (%)

2% 69, 67 9.0, 7.5 4.5, 3.9 27.5, 23.9 50, 50 25.1, 25.9 14.5, 14.9
5% 173, 171 16.5, 14.6 3.3, 2.9 20.2, 18.1 60, 60 12.0, 12.2 6.9, 7.0
10% 346, 335 28.8, 25.8 2.9, 2.7 17.6, 16.4 90, 75 9.0, 7.8 5.2, 4.5
15% 517, 519 34.4, 33.3 2.3, 2.2 14.1, 13.7 95, 80 6.4, 5.3 3.7, 3.1
20% 692, 674 40.6, 38.5 2.0, 2.0 12.4, 12.2 95, 80 4.8, 4.1 2.7, 2.4
25% 866, 862 45.3, 44.6 1.8, 1.8 11.1, 11.0 95, 85 3.8, 3.4 2.2, 2.0
50% 1732, 1715 72.6, 74.2 1.5, 1.5 8.9, 9.2 100, 95 2.0, 1.9 1.2, 1.1
75% 2585, 2586 90.6, 88.3 1.2, 1.2 7.4, 7.3 100, 100 1.3, 1.3 0.8, 0.8
100% 3464, 3465 100.0, 100.0 1.0, 1.0 6.1, 6.1 100, 100 1.0, 1.0 0.6, 0.6
*

Efficiency ratios are based on the actual percentage of exams above the AI-score threshold, which may differ slightly from the target percentage of exams.

These findings suggest that while AI may not reliably identify all patients who will be recalled, it can concentrate a large fraction of screen-detected cancers into a small subset of mammograms with high malignancy suspicion scores. This pattern is not surprising, because these mammography AI algorithms were optimized to identify cancer (or cancer-associated imaging features), rather than to predict radiologist recall decisions, which are influenced by additional factors (e.g., cautionary practice patterns and benign but suspicious-appearing findings). Accordingly, AI-based triage may be more effective for prioritizing examinations most likely to benefit from expedited diagnostic evaluation than for capturing all recalls.

Our findings support a clinically feasible strategy in which AI is used to triage a limited number of screening mammograms for immediate remote interpretation, enabling same-day diagnostic mammography for patients most likely to benefit (i.e., those with cancer visible at screening) and potentially decreasing the time to diagnostic resolution. A prior strategy based on time since prior mammography, targeting 9.2% of patients undergoing either a first mammogram or a mammogram ≥5 years after their previous mammogram, captured 19.2% of patients with abnormal findings and could potentially reduce the percentage requiring a second visit from 8.8% to 7.1%.8 In contrast, an AI-based approach selecting the top 10% of AI scores captured a larger share of recalls (26%–29%) and, importantly, most screen-detected cancers (75%–90%). Immediate interpretation based on time since prior mammogram may be useful if women need to be pre-scheduled for same-day interpretation appointments. In contrast, an AI-based implementation requires near-real-time processing while the patient waits and the capacity for add-on diagnostic mammography on mobile vans or diagnostic imaging at a nearby facility when indicated. However, women may differ systematically in their ability to use these same-day opportunities because of employment, childcare, transportation, or other barriers, which could contribute to delays in diagnostic work-up.

Study limitations include the evaluation of only two commercially available, FDA-cleared AI algorithms at a single institution with only 25 breast cancers. Due to the limited sample size, we did not evaluate the impact of patient characteristics such as race or ethnicity on AI-triage performance. This study was performed retrospectively and provides a proof-of-concept for the prospective implementation of an AI-driven workflow. AI was not used during the interpretation of the mammograms, which could have influenced recall decisions. AI-driven strategies for more rapid diagnostic evaluation in a subset of women will require substantial operational resources. Implementing same-day diagnostic imaging after a positive screening mammogram requires radiologist availability for rapid (potentially remote) interpretation, workflow and IT support for expedited image transfer from mobile units, and accounting for same-day services in daily clinical appointment grids.8 Reserving same-day diagnostic slots may reduce overall volume if slots are not consistently filled. In our study sample, the proportion of examinations exceeding a given AI malignancy score threshold varied substantially from day to day. Thus, threshold-based triage, whether based on AI scores or another rule, must account for daily variation in the proportion of examinations exceeding the selected threshold.6

Sources of Support:

This work was supported by the UC Davis Comprehensive Cancer Center’s Women’s Cancer Care & Research (WeCARE) program; by residual class settlement funds in the matter of April Krueger v. Wyeth Inc., case No. 03-cv-2496 (US District Court, SD of Calif); and a gift from the Safeway Foundation.

Leadership Roles:

Dr. Miglioretti is Division Chief of Biostatistics for the Department of Public Health Sciences and Program Leader for the Comprehensive Cancer Center for the University of California, Davis. Dr. Morris is Chair of the Department of Radiology at the University of California, Davis. Dr. Aminololama-Shakeri is the Division Chief of Breast Radiology, Program Director of the Breast Imaging Fellowship, and Vice Chair of Academic Affairs for the University of California. Dr. Lee is Vice Chair of Research for the Department of Radiology at the University of Wisconsin-Madison School of Medicine and Public Health; Dr. Lee is also Deputy Editor of JACR. All authors report being employed at non-profit institutions.

Footnotes

Disclaimers: The views expressed in this article are the coauthors and do not necessarily represent views of the institutions or funders.

Artificial Intelligence: During the preparation of this work, the first author used ChatGPT 5.3 to improve language and readability. After using this tool, the first author carefully reviewed and edited the content and takes full responsibility for the content of the publication.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • 1.Lee CI, Abraham L, Miglioretti DL, et al. National Performance Benchmarks for Screening Digital Breast Tomosynthesis: Update from the Breast Cancer Surveillance Consortium. Radiology. May 2023;307(4):e222499. doi: 10.1148/radiol.222499 [DOI] [Google Scholar]
  • 2.Nelson HD, Pappas M, Cantor A, Griffin J, Daeges M, Humphrey L. Harms of Breast Cancer Screening: Systematic Review to Update the 2009 U.S. Preventive Services Task Force Recommendation. Ann Intern Med. Feb 16 2016;164(4):256–67. doi: 10.7326/M15-0970 [DOI] [PubMed] [Google Scholar]
  • 3.Lawson MB, Bissell MCS, Miglioretti DL, et al. Multilevel Factors Associated With Time to Biopsy After Abnormal Screening Mammography Results by Race and Ethnicity. JAMA Oncol. Aug 1 2022;8(8):1115–1126. doi: 10.1001/jamaoncol.2022.1990 [DOI] [Google Scholar]
  • 4.Dontchos BN, Achibiri J, Mercaldo SF, et al. Disparities in Same-Day Diagnostic Imaging in Breast Cancer Screening: Impact of an Immediate-Read Screening Mammography Program Implemented During the COVID-19 Pandemic. AJR Am J Roentgenol. Feb 2022;218(2):270–278. doi: 10.2214/AJR.21.26597 [DOI] [PubMed] [Google Scholar]
  • 5.Yoon JH, Strand F, Baltzer PAT, et al. Standalone AI for Breast Cancer Detection at Screening Digital Mammography and Digital Breast Tomosynthesis: A Systematic Review and Meta-Analysis. Radiology. Jun 2023;307(5):e222639. doi: 10.1148/radiol.222639 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lin Y, Hoyt AC, Manuel VG, et al. Risk-Stratified Screening: A Simulation Study of Scheduling Templates on Daily Mammography Recalls. Journal of the American College of Radiology : JACR. Mar 2025;22(3):297–306. doi: 10.1016/j.jacr.2024.12.010 [DOI] [Google Scholar]
  • 7.Ramwala OA, Lowry KP, Hippe DS, et al. ClinValAI: A framework for developing Cloud-based infrastructures for the External Clinical Validation of AI in Medical Imaging. Pac Symp Biocomput. 2025;30:215–228. doi: 10.1142/9789819807024_0016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ho TH, Bissell MCS, Lee CI, et al. Prioritizing Screening Mammograms for Immediate Interpretation and Diagnostic Evaluation on the Basis of Risk for Recall. Journal of the American College of Radiology : JACR. Mar 2023;20(3):299–310. doi: 10.1016/j.jacr.2022.09.030 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES