Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 23.
Published in final edited form as: Cancer Epidemiol Biomarkers Prev. 2026 Aug 3;35(8):1375–1385. doi: 10.1158/1055-9965.EPI-25-2052

Performance of Statistical and Machine Learning Risk Prediction Models for Advanced Breast Cancers

Shuai Chen 1, Karla Kerlikowske 2,3,4, Yu-Ru Su 5,6, Rebecca A Hubbard 7, Charlotte C Gard 8, Jeffrey A Tice 9, Brian L Sprague 10,11, Thomas P Ahern 10,12, Diana L Miglioretti 1,5
PMCID: PMC13286512  NIHMSID: NIHMS2182002  PMID: 42132486

Abstract

Background:

Machine learning enables complex risk prediction models, but comparative performance with statistical approaches remains context-dependent. We compared statistical and machine learning models for predicting advanced breast cancer risk.

Methods:

Using data from 968,178 women (40–74 years) undergoing 2,796,459 annual or 812,126 biennial screening mammograms (2005-2019) in the Breast Cancer Surveillance Consortium, we cross-validated models predicting advanced breast cancer within 12 months (annual) or 24 months (biennial) following screening. Models included conventional logistic regression, regularized regressions (LASSO, Elastic net), and machine learning methods (random forests, gradient boosting), considering a modest number of clinical and demographic predictors. Performance was assessed using calibration and area under the receiver operating characteristic curve (AUC).

Results:

Discrimination was similar across models (AUC 0.677–0.690). Calibration differences were more pronounced. Regularized regressions achieved the most favorable calibration overall and across racial and ethnic groups, with AUC 0.689 (95%CI = 0.676–0.701). Gradient boosting showed comparable AUC but suboptimal calibration (calibration slope 1.12; 95%CI = 1.04–1.20). Conventional logistic regression had slightly lower AUC (0.683; 95%CI = 0.671–0.696) and calibration slope of 0.90 (95%CI = 0.83–0.96). Regression-based approaches were generally well calibrated across racial and ethnic groups (E/O ratio 0.96–1.03; calibration intercept −0.03 to 0.04), with some subgroup deviations in calibration slopes (<1).

Conclusions:

For predicting advanced breast cancers, regularized regression demonstrated similar discrimination and generally more favorable calibration than other approaches.

Impact:

In settings with rare outcomes and low dimensional features, regularized regression may offer a practical balance between performance and interpretability.

Introduction

Numerous prediction modeling methods have been applied across various healthcare settings to estimate clinical risk (1-3). Machine learning methods can flexibly incorporate nonlinearities and interactions. However, in many real-world healthcare settings with hundreds to a few thousand outcome events and tens (rather than hundreds or thousands) of candidate predictors, regression-based approaches often achieve similar or better performance across multiple clinical domains. Prior studies reported small differences in discrimination measured by area under the receiver operating characteristic curve (AUC) (absolute differences <0.01–0.03), and no consistent improvement in calibration with machine learning methods (4-8). Regression-based approaches also provide analytic expressions for the prediction model, facilitating interpretation and clinical dissemination. Importantly, no single prediction modeling approach consistently outperforms others (7, 9), underscoring the relative utility of different modeling approaches in different contexts.

Beyond overall predictive performance, evaluation for model fairness is needed to ensure that models perform equitably across racial and ethnic groups and do not perpetuate or exacerbate racial and ethnic disparities (10-16). This concern is especially salient for machine learning methods, which generally require large numbers of events per variable to achieve good performance (17) and may therefore perform poorly in underrepresented populations.

Screening mammography reduces breast cancer mortality by lowering the incidence of advanced breast cancer (18-20). Advanced breast cancer, defined as American Joint Committee on Cancer (21) prognostic pathologic stage II or higher, occurs in 22%-24% of routinely screened women diagnosed with invasive breast cancer (22-24) and is associated with worse survival than early-stage disease (19, 23-25). Since advanced breast cancer represents clinically meaningful disease burden, accurate prediction of advanced disease risk is clinically relevant for guiding screening strategies, including identifying high-risk women who might benefit from shorter screening intervals or supplemental imaging, and low-risk women who might safely undergo less intensive screening (thereby reducing harms from false-positives associated with annual screening). The Breast Cancer Surveillance Consortium (BCSC) developed the first actionable advanced breast cancer risk prediction model to guide decisions on screening interval and supplemental imaging (https://tools.bcsc-scc.ucdavis.edu/AdvBC6yearRisk/#/) (26), using conventional logistic regression and routinely collected clinical predictors (menopausal status, age, race and ethnicity, family history, benign biopsy history and result, body mass index, and breast density). While this model provides a transparent and clinically interpretable framework, it relies on prespecified functional forms and interactions. More advanced statistical and machine learning approaches could enhance predictive performance, but they also carry tradeoffs in model transparency, interpretability and fairness.

We aimed to develop and compare various statistical and machine learning models for predicting advanced breast cancer risk. We evaluate relative performance of modeling approaches, by assessing calibration and discrimination, overall and across racial and ethnic groups, to guide method selection for similar clinical contexts characterized by rare outcomes (despite large sample sizes) and tens of candidate predictors (i.e., low- to moderate-dimensional settings with events-per-variable ratios well above commonly cited minimum thresholds of 10–20 events per variable), particularly in the domain of cancer risk prediction.

Materials and Methods

Study Setting and Data Sources

Data were obtained from the BCSC mammography registries (https://www.bcsc-research.org/), a large multi-regional U.S. screening cohort whose demographic characteristics are broadly similar to those of U.S. women undergoing screening mammography, with slight differences in rurality and racial/ethnic composition (27-29). Women’s demographics and mammography from radiology facilities were prospectively collected. Breast cancer diagnoses and tumor characteristics were obtained by linking women to pathology databases, regional Surveillance, Epidemiology, and End Results (SEER) programs, and state or regional tumor registries. Deaths were ascertained by linking to state death records. Registries and the Statistical Coordinating Center (SCC) received approvals to perform the studies by the Institutional Review Boards (IRBs) of the University of North Carolina at Chapel Hill, Kaiser Permanente, Dartmouth-Hitchcock Medical Center, the University of California San Francisco, the University of Vermont and the UVM Medical Center, and the University of Illinois at Chicago. Depending on registry site, facility, and study year, written informed consent was obtained from participants, passive permission procedures were used, or informed consent requirements were waived under IRB-approved protocols. All procedures were conducted in accordance with the U.S. Common Rule and were Health Insurance Portability and Accountability Act compliant, and registries and the SCC received a Federal Certificate of Confidentiality and other protections for the identities of women, physicians, and facilities.

Participants

We applied the same inclusion/exclusion criteria and used the data from the same registry sites as the original BCSC advanced cancer risk model (26), with an additional 1.5 years of new data. To reflect women routinely screened, screening mammograms from January 2005 through June 2019 among women aged 40-74 years were included if a prior mammogram occurred 11-30 months earlier. Annual and biennial screening were defined as 11-18 and 19-30 months since the prior mammogram, respectively. We excluded screens from women with a history of breast cancer, mastectomy, or lobular carcinoma in situ, and excluded mammograms that were either unilateral or preceded by mammography within 9 months. To align with the intended use for women undergoing routine screening mammography without supplemental imaging, we excluded screening examinations with same-day screening ultrasound or screening MRI in the preceding or following 12 months. These criteria are consistent with the original BCSC advanced cancer risk model.

Measures, Definitions, Outcomes

Risk factors were defined based on data available at or prior to the screening mammogram. Demographic and breast health history variables were primarily obtained from self-administered questionnaires at the time of screening, supplemented by electronic health record data when questionnaire data were unavailable or incomplete. Because data were integrated from multiple sources, individual variables may reflect combined sources. Race and ethnicity were self-reported and categorized as Hispanic, non-Hispanic African American/Black, Asian/Pacific Islander, non-Hispanic White, Other/multiracial (non-Hispanic Native American/Alaskan Native, or with two or more reported races, or other). Radiologists categorized breast density during clinical interpretation using Breast Imaging Reporting and Data System (BI-RADS) density categories (30): almost entirely fat, scattered fibroglandular densities, heterogeneously dense, and extremely dense. Menopausal status was based on self-reported questionnaire at the time of screening and categorized using the standard BCSC hierarchical algorithm (31-33). Postmenopausal women were those with both ovaries removed, self-reporting natural menopause, ≥365 days since last menstrual period, current postmenopausal hormone therapy users, or age ≥60. Premenopausal women were those self-reporting a period within the last 180 days or birth control hormone users. Perimenopausal women were those reporting uncertainty about menopausal status or a last menstrual period 180-364 days prior. Women reporting surgical menopause without bilateral oophorectomy, cessation of menses for other reasons, or insufficient information were classified as unknown. Body mass index (BMI) was categorized into <18.5kg/m2=underweight, 18.5–24.9kg/m2=normal weight, 25.0–29.9kg/m2=overweight, 30.0–34.9kg/m2=obese I, and ≥35.0kg/m2=obese II/III (34) when needed in analysis. Mammography type was classified as film, digital (2-D only), and digital breast tomosynthesis. Parity and age at first live birth were combined and categorized as nulliparous, age <30, or ≥30 years. Prior benign breast diagnoses from clinical pathology reports were grouped based on the highest grade as proliferative with atypia > proliferative without atypia > non-proliferative, according to published taxonomy (35-37).

The outcome was advanced breast cancer diagnosed within 12 months (annual) or 24 months (biennial) after mammography. Advanced cancer was defined as prognostic pathologic stage II or higher using American Joint Committee on Cancer, 8th edition (21), based on anatomic staging elements (TNM) and tumor grade, and estrogen, progesterone, and human epidermal growth factor receptor status as in the original BCSC advanced cancer risk model (26). If prognostic stage could not be calculated (16%), we used anatomic stage IIb or higher (14%) or other combined information based on high differentiation grade, high SEER summary stage, large tumor size, positive lymph node status, and treatment details (2%). Thirteen mammograms (<0.001%) had missing outcome due to insufficient staging information.

Statistical Approach

Overview

The screening mammogram was the unit of analysis, and the target estimand was the probability of advanced breast cancer following one screening round (12 months for annual; 24 months for biennial). Because the primary goal was prediction rather than inference on regression coefficients, within-woman correlation from repeated mammograms was not explicitly modeled. Logistic models under a working independence correlation yield consistent predicted risk estimates (38, 39), and performance metrics depend only on these predicted probabilities. In sensitivity analysis using generalized estimating equations with clustering at the woman level, coefficient estimates were unchanged and robust standard errors differed minimally (relative differences <8%), indicating negligible impact. Women entered the analysis each time they received a screening mammogram at a participating facility, with the demographic and clinical predictors ascertained at or prior to the time of the mammogram, reflecting information available at the screening encounter. We restricted analyses to all screens with complete outcome ascertainment, defined as availability of cancer registry follow-up for the full risk window (12 or 24 months for annual or biennial screens, respectively).

Models were constructed to compare modeling strategies in a staged manner. First, we evaluated progressive refinements of the expert-driven logistic regression framework, including alternative functional forms, predictor expansion, and regularization. Second, we compared different modeling structures using comparable candidate predictors, with differences arising from stratification versus interaction parameterization and functional form handling (splines in regression-based models versus original continuous variables in machine learning methods). Model predictors are documented in Supplemental Table S1.

The workflow of analysis is illustrated in Supplemental Fig. S1. Within each imputed dataset, we first evaluated the prediction performance through 5-fold cross-validation, and then fit the final models using all data within this imputed dataset to obtain variable importance and coefficient estimates. Data were analyzed using R version 4.4.2 (RRID: SCR_001905), SAS version 9.4 (RRID: SCR_008567), and Python version 3.13 (RRID: SCR_008394). Relevant code is available at: https://codeocean.com/capsule/5083809/tree/v1.

Missing data

Missing values in variables involved in analyses were imputed using multiple imputation by chained equations (MICE) (40, 41), separately for annual and biennial screens. All model building, predictions and assessments were based on 5 imputed datasets unless stated otherwise. Spline terms of age were used to impute other variables to allow nonlinear age effects. Both perimenopausal and unknown menopausal status were considered missing and imputed into either premenopausal or postmenopausal status. There were 13 missing advanced cancer outcomes, and the imputation model for the outcome was built within the subset of screens associated with cancers only. Digital breast tomosynthesis (DBT) was approved by the FDA in 2011 and hence the type of mammography could be DBT only in 2011 and later. Thus, the imputation for type of mammography was performed separately within the two subsets (screens prior to 2011; screens in 2011 and later). Such imputation model built within subsets was implemented during the multiple imputation procedure using MNAR statement in SAS procedure PROC MI. Other variables were imputed using all annual or biennial screens. To better handle missing BMI values and its possibly nonlinear association with other variables, we imputed the missing data in two steps: We first obtained 5 imputations using categorical BMI in the imputation model, and Supplemental Table S2 summarizes the variables in the imputation model performed in SAS. For each of the 5 imputed datasets from the first step, missing continuous BMI was imputed with a single imputation based on the predictive mean matching method, including categorical BMI and other variables as predictors in the imputation model.

We then performed imputation diagnostics for each imputed variable, by comparing summary statistics (mean, standard deviation, range for continuous variables; proportion for categorical variables) and graphical plots (density plot for continuous variables; barplot for categorical variables) between observed and imputed values. The results displayed similar distributions between observed and imputed for all imputed variables.

We also evaluated fraction of missing information (42) and the impact of increasing the number of imputations, focusing on the precision of point estimates for performance metrics. Supplemental Table S3 demonstrated that increasing number of imputations beyond the current 5 would reduce the standard errors of performance metrics by <6% compared with an infinite number of imputations, indicating that 5 imputations are adequate for this study. We also performed sensitivity analyses for two alternative ways to handle missing data (excluding the 13 screens with missing outcomes and increasing number of imputations to 10). The absolute differences are ≤0.001, ≤0.001, ≤0.001, and ≤0.021 across all modeling approaches for AUC, expected-to-observed event ratio (E/O ratio), calibration intercept and slope, respectively, indicating that the results are robust (Supplemental Table S4).

Regression techniques

We considered an “Expert” model based on the original BCSC advanced cancer model structure (26), with separate logistic regressions stratified by menopausal status (premenopausal and postmenopausal) and screening interval. To avoid bias from overlap with the original training data used by BCSC advanced cancer model, models were refitted within the cross-validation framework. Predictors included age (linear and quadratic), race and ethnicity, first-degree family history of breast cancer, history of benign biopsy, categorical BMI, and breast density. Variants included spline-based modeling (43) of continuous age and BMI (Expert model with splines), and an Expansion model that further added three new predictors to the Expert model with splines: prior false-positive results within 5 years, parity and age at first live birth, and mammography type.

We considered two types of regularized regressions, LASSO (44) and Elastic net (45), which perform simultaneous variable selection and coefficient shrinkage using penalized optimization. We considered two modeling strategies: stratified models by menopausal status and screening interval, including all Expansion model predictors; and a single unstratified model including all stratification variables (menopausal status, interval), Expansion model predictors, and full interactions between stratification variables and predictors (details in Supplemental Table S1). For unpenalized logistic regressions, stratified models are equivalent to an unstratified model with full interactions between stratification variables and predictors. This equivalence does not hold under penalization, where penalties are applied jointly across all terms, making penalized regression a data-adaptive alternative for modeling interactions within an unstratified model. We also evaluated logistic regressions with model terms selected by LASSO or Elastic net but refitted without a penalty term (Refitted LASSO, Refitted elastic net), since regularized parameter estimates are biased (46). Regularized regressions were fitted using the R package “glmnet” (version 4.1.8).

Machine learning algorithms

We considered two ensemble tree-based machine learning approaches: random forests (47) and gradient boosting (48). Models were trained on the full dataset without stratification, including menopausal status, interval, and all Expansion model predictors. These approaches capture nonlinearities and interactions adaptively without explicit specification. Random forests average predictions across trees built on subsamples with feature subsetting, while gradient boosting aggregate trees sequentially built to improve prediction error. We used histogram-based gradient boosting for computational efficiency (49, 50).

Additional analyses evaluated performance of the two machine learning methods handling missing predictors directly without imputation or creation of an “unknown” category. The algorithms allowed the tree grower to learn whether samples with missing values should go to the left or right child at each split point based on the potential gain. This is particularly advantageous when handling missing continuous or ordinal predictors when comparing to creation of an “unknown” category. The 13 screens with missing advanced cancer outcome were excluded for these approaches.

The hyperparameters for machine learning methods include the number of randomly selected candidate features at each split for random forests, maximum tree depth, the learning rate, and the maximum number of leaves for each tree. We built 600 decision trees in each random forest, and a maximum of 10000 boosted trees, allowing for early stopping when the fitting performance criteria did not improve. The performance criteria used the AUC to select hyperparameters and determine whether to stop early.

As advanced cancer is rare, we considered two alternative strategies to handle class imbalance for machine learning methods, including using a weighting approach (i.e., applying more weights to oversample the advanced cancers) to achieve better balance, and using area under the precision-recall curve as the criteria to select hyperparameters. However, the performance of the machine learning approaches was not improved noticeably when using these alternative strategies. Thus, we only reported the results using original unweighted approaches based on criteria of AUC.

Models were fitted using “RandomForestClassifier” and “HistGradientBoostingClassifier” in Python module “scikit-learn” (version 1.6.1) (51).

Cross-validation

Cross-validation for performance evaluation:

The 5-fold cross-validation was chosen to provide fair, internally validated comparisons of predictive performance across all modeling approaches. We used clustered randomization to ensure that all screens of a woman were assigned to the same fold and that the same assignment was applied to all imputed datasets, preventing information leakage. Within each training set (data from 4 folds), we further conducted 5-fold cross-validation to select hyperparameters, which were used to build the model using the entire training set. The built model was used to generate predicted risks for the validation set (data from the remaining fold). The predicted risks from 5 validation folds were pooled together to evaluate performance.

Cross-validation for final model building:

For the final model building, we used 5-fold cross-validation to select hyperparameter values within each imputed dataset. The selected hyperparameters were used to build the final model using the entire imputed dataset, leading to one final model within each imputed dataset.

Performance evaluation

Model performance was assessed using calibration (52) and discrimination. Calibration was assessed using the E/O ratio, Cox calibration intercept and slope, reflecting systematic bias and over/underfitting. Here, overfitting and underfitting refer to cases in which predicted risks in high- and low-risk groups are more or less extreme, respectively, than observed risks. Rather than relying on hypothesis testing, we interpreted the magnitude and precision of these estimates relative to their ideal values, which are E/O ratio = 1, calibration intercept = 0, calibration slope = 1. The 95% CI for E/O ratio was calculated by assuming that the observed advanced breast cancer events follow a Poisson distribution; thus, we calculated the CI using (E/O ratio)*exp (± 1.96/sqrt [O]), which is a standard method to construct CI for E/O ratio (53-55). For the Cox calibration intercept and slope, their point estimates and CIs were obtained using logistic calibration model, where predicted risks are treated as fixed covariates as is standard in validation settings. We also performed decile-based and graphical assessment via calibration plots with Hosmer-Lemeshow (HL) test p-values. Because HL tests may detect statistically significant deviations from perfect calibration even when miscalibration is small in magnitude in very large samples, we emphasize the magnitude and graphical assessment. Calibration was also assessed within racial and ethnic subgroups to evaluate model fairness.

The discriminatory accuracy was evaluated using the AUC. Sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were also calculated using a threshold of top 25% of predicted risk (corresponding to intermediate/high risk), defined separately for annual and biennial screens due to different follow-up periods.

Performance metrics were estimated within each imputed dataset and combined to obtain the final estimates and confidence intervals using Rubin’s rule (56). For the E/O ratio, Rubin’s rule was applied to log scale of E/O ratio, and the combined estimates and confidence interval limits were then exponentiated to obtain final estimates and intervals for the E/O ratio. For stratified models, predicted risks from all strata were pooled to obtain overall performance metrics.

Variable importance in final models

Final models were refit using all observations within each imputed dataset. Variable importance for regularized regressions was assessed by selection frequency, and for machine learning models by mean decrease in AUC (57). Final variable importance was obtained by averaging the variable importance estimates across imputations.

Data availability

The data underlying this article will be shared upon email request to the Corresponding Author or BCSC Statistical Coordinating Center (kpwa.scc@kp.org) with appropriate regulatory approvals.

Results

Descriptive statistical analysis

The cohort included 968,178 women contributing 2,796,459 annual (from 808,376 women) and 812,126 biennial (from 491,540 women) screening mammograms, with 1,128 and 791 advanced cancers, respectively. An average of 3.7 screens per woman were contributed (median=3, IQR=1–5). Compared with women screened biennially, women screened annually tended to be older, non-Hispanic White, have a family history of breast cancer, and have a history of breast biopsy (Table 1). The averaged time since prior screening mammogram is 13.3 or 23.7 months for annual or biennial screens, respectively.

Table 1.

Exam-level characteristics of women undergoing annual and biennial screening

Characteristics Annual (N = 2,796,459) Biennial (N = 812,126)
No Advanced
cancera
Advanced breast cancerb No Advanced
cancera
Advanced breast cancerb
No. Column % No. Column % Row % No. Column % No. Column % Row %
Screening examinations 2,795,324 99.96 1,128 c 0.04 811,329 99.90 791 c 0.10
Age, years
 40-49 681,301 24.4 214 19.0 0.03 212,912 26.2 170 21.5 0.08
 50-59 981,651 35.1 394 34.9 0.04 300,937 37.1 294 37.2 0.10
 60-69 846,932 30.3 392 34.8 0.05 230,614 28.4 246 31.1 0.11
 70-74 285,440 10.2 128 11.3 0.04 66,866 8.2 81 10.2 0.12
Race/ethnicity
 Asian/Pacific Islander 258,938 9.6 68 6.2 0.03 113,065 14.3 89 11.5 0.08
 Black, non-Hispanic 275,057 10.2 207 19.0 0.08 75,396 9.5 108 13.9 0.14
 Hispanic 132,070 4.9 49 4.5 0.04 53,162 6.7 48 6.2 0.09
 White, non-Hispanic 1,987,107 73.6 744 68.1 0.04 527,787 66.8 511 65.9 0.10
 Other/multiracial 48,022 1.8 24 2.2 0.05 20,352 2.6 19 2.5 0.09
 Unknown 94,130 3.4 36 3.2 0.04 21,567 2.7 16 2.0 0.07
Menopausal
 No 698,701 30.2 247 25.7 0.04 215,275 31.9 203 30.2 0.09
 Yes 1,616,589 69.8 713 74.3 0.04 458,588 68.1 470 69.8 0.10
 Unknownd 480,034 17.2 168 14.9 0.03 137,466 16.9 118 14.9 0.09
1st degree family history of breast cancere
 No 2,210,635 81.8 815 74.4 0.04 686,628 86.8 645 83.3 0.09
 Yes 491,449 18.2 280 25.6 0.06 104,598 13.2 129 16.7 0.12
 Unknown 93,240 3.3 33 2.9 0.04 20,103 2.5 17 2.1 0.08
History of breast biopsy
 None (no prior biopsy) 2,148,546 76.9 741 65.7 0.03 679,594 83.8 594 75.1 0.09
 Prior biopsy, benign diagnosis unknown 415,614 14.9 257 22.8 0.06 95,553 11.8 150 19.0 0.16
 Non-proliferative 163,601 5.9 91 8.1 0.06 26,693 3.3 33 4.2 0.12
 Proliferative without atypia 56,804 2.0 30 2.7 0.05 8,363 1.0 12 1.5 0.14
 Proliferative with atypia 10,759 0.4 9 0.8 0.08 1,126 0.1 2 0.3 0.18
BI-RADS breast density
 Almost entirely fat 263,060 9.5 46 4.2 0.02 79,258 10.0 30 4.0 0.04
 Scattered fibroglandular densities 1,191,962 43.2 391 35.7 0.03 325,463 41.2 274 36.3 0.08
 Heterogeneously dense 1,079,622 39.1 547 50.0 0.05 318,123 40.3 372 49.3 0.12
 Extremely dense 224,631 8.1 110 10.1 0.05 66,807 8.5 78 10.3 0.12
 Unknown 36,049 1.3 34 3.0 0.09 21,678 2.7 37 4.7 0.17
Body mass index, kg/m2
 Underweight (<18.5) 26,739 1.5 10 1.6 0.04 8,769 1.6 6 1.1 0.07
 Normal (18.5-24.9) 736,413 41.5 216 33.6 0.03 228,676 40.7 189 34.2 0.08
 Overweight (25.0-29.9) 519,905 29.3 182 28.3 0.03 161,486 28.7 179 32.4 0.11
 Obese I (30.0-34.9) 281,363 15.8 120 18.7 0.04 89,244 15.9 121 21.9 0.14
 Obese II/III (≥35.0) 211,356 11.9 115 17.9 0.05 73,986 13.2 58 10.5 0.08
 Unknown 1,019,548 36.5 485 43.0 0.05 249,168 30.7 238 30.1 0.10
Parity and age at first live birth
 Nulliparous 433,594 21.0 182 22.4 0.04 139,065 22.3 149 24.1 0.11
 Age < 30 years 1,235,657 59.9 487 60.0 0.04 359,410 57.7 331 53.6 0.09
 Age ≥ 30 years 394,484 19.1 142 17.5 0.04 124,637 20.0 137 22.2 0.11
 Unknown 731,589 26.2 317 28.1 0.04 188,217 23.2 174 22.0 0.09
Type of mammogram
 Film 606,903 21.7 303 26.9 0.05 227,848 28.1 272 34.4 0.12
 Digital (2-D only) 1,900,985 68.0 719 63.7 0.04 538,955 66.4 475 60.1 0.09
 Tomosynthesis 257,314 9.2 87 7.7 0.03 37,731 4.7 36 4.6 0.10
 Not otherwise specified 30,122 1.1 19 1.7 0.06 6,795 0.8 8 1.0 0.12
Prior false positive result within 5 years
 No 2,181,151 78.0 805 71.4 0.04 696,048 85.8 626 79.1 0.09
 Yes 614,173 22.0 323 28.6 0.05 115,281 14.2 165 20.9 0.14
a

Includes non-advanced breast cancers. American Joint Committee on Cancer (AJCC), Breast Imaging Reporting and Data System (BI-RADS).

b

Invasive cancer AJCC 8th edition prognostic pathologic stage II or higher within 12 or 24 months of screening mammography.

c

Not applicable.

d

Includes perimenopausal, surgical menopause, periods stopped for other (not natural or surgical) reason, and other unknown menopausal status due to not available information.

e

Defined as first-degree relative (mother, sister, or daughter) with breast cancer.

Calibration

All approaches demonstrated E/O ratios of 1.00 and calibration intercepts of 0.00, indicating no systematic biases (Fig. 1). Calibration slope differentiated methods. The calibration slopes of unstratified LASSO (0.99; 95%CI = 0.93–1.05) and Elastic net (0.99; 95%CI = 0.93–1.06) were closest to ideal 1. The Expert model had a calibration slope of 0.90 (95%CI = 0.83–0.96), suggesting overfitting with approximately 10% attenuation relative to the ideal slope of 1. Other regression-based approaches had calibration slopes in the range 0.85–0.92, suggesting some overfitting. Refitted regularized models suffered more overfitting than their penalized counterparts. Machine learning approaches showed underfitting, with calibration slopes >1, most notably for random forests (1.65 and 1.68 for random forests using unimputed and imputed data, respectively). Gradient boosting showed moderate underfitting (1.13 and 1.12 for gradient boosting using unimputed and imputed data, respectively) (Fig. 1). Calibration plots (Fig. 2) confirmed these patterns: Unstratified regularized models demonstrated the best calibration across all risk deciles; Gradient boosting underpredicted risks slightly in the highest risk decile; Random forests severely underpredicted the risks in high-risk deciles while overpredicting the risks in low-risk deciles, which is consistent with its known behavior due to high variance resulted from feature subsetting (58), especially in rare-outcome setting.

Figure 1.

Figure 1.

Model calibration measured by three metrics, including the ratio of expected to observed events (E/O ratio), calibration intercept, and calibration slope for each prediction modeling approach. Points and error bars represent point estimates and 95% CIs for each metric. Vertical reference lines show the ideal value for each metric (1 for E/O ratio, 0 for calibration intercept, and 1 for calibration slope). “Stratified” indicates separate logistic regressions stratified by menopausal status and interval; “unstratified” indicates single unstratified model; “imputed” indicates machine learning models applied to multiply imputed data; “unimputed” indicates machine learning models handling missing predictors directly without imputation.

Figure 2.

Figure 2.

Weak calibration for selected risk prediction models for advanced cancer. Each panel demonstrates the weak calibration of an individual modeling approach by comparing the mean predicted risk (x-axis) to the observed risk of advanced cancer (y-axis) in 10 deciles determined by the predicted risk. The vertical error bars show the 95% CI of the observed risk of advanced cancer in individual deciles. “Stratified” indicates separate logistic regressions stratified by menopausal status and interval; “unstratified” indicates single unstratified model; “imputed” indicates machine learning models applied to multiply imputed data. The machine learning approaches using unimputed data directly had similar patterns to the approaches using imputed data and thus were omitted. HL test: Hosmer-Lemeshow test.

Calibration stratified by race and ethnicity

Regression-based models were robustly unbiased across all racial and ethnic groups regarding E/O ratios (range: 0.96 – 1.03) and calibration intercepts (range: -0.03 – 0.04), while evidence of overfitting was observed in non-white subgroups (calibration slope <1), most notably in the extremely small Other/multiracial subgroup (Fig. 3). Unstratified regularized models were calibrated best across all racial and ethnic groups, and demonstrated relatively small between-subgroup variation, whereas machine learning approaches exhibited greater variability between subgroups (Supplemental Tables S5, S6). The gradient boosting approaches showed slight underfitting in Non-Hispanic White women as suggested by its high calibration slopes (1.19 and 1.16 for the approaches using unimputed and imputed data, respectively), but no lack of calibration in other groups or according to other calibration criteria. Random forests approaches demonstrated more noticeable lack of calibration in several groups (Fig. 3).

Figure 3.

Figure 3.

Race and ethnicity-stratified model calibration for advanced cancer by three metrics: the ratio of expected and observed events (E/O ratio), the calibration intercept, and the calibration slope within racial and ethnic subgroups. Error bars depict 95% CIs for each metric. Vertical lines show the ideal value for each metric (i.e., 1 for E/O ratio, 0 for calibration intercept, and 1 for calibration slope). “Stratified” indicates separate logistic regressions stratified by menopausal status and interval; “unstratified” indicates single unstratified model; “imputed” indicates machine learning models applied to multiply imputed data; “unimputed” indicates machine learning models handling missing predictors directly without imputation. PI, Pacific Islander.

Discrimination

Discriminatory accuracy was similar across all approaches (AUC 0.677–0.690; Table 2). Unstratified LASSO, Elastic net, and gradient boosting achieved the highest AUCs (0.689–0.690), representing only modest improvement over the Expert model (0.683). Confidence intervals overlapped substantially, indicating small differences.

Table 2.

Assessment of discriminatory accuracy using area under the receiver operating characteristic curve (AUC) for various risk modeling approaches, along with the corresponding 95% confidence interval (CI)

Modeling Strategy Approach AUC (95% CI)
Stratified Models by Menopausal and Interval Expert (original) 0.683 (0.671, 0.696)
Expert (splines) 0.683 (0.671, 0.696)
Expansion 0.686 (0.673, 0.698)
Refitted LASSO 0.684 (0.672, 0.697)
LASSO 0.684 (0.672, 0.697)
Refitted elastic net 0.685 (0.672, 0.698)
Elastic net 0.684 (0.671, 0.697)
Single Unstratified Model Refitted LASSO 0.687 (0.675, 0.700)
LASSO 0.689 (0.676, 0.701)
Refitted elastic net 0.687 (0.675, 0.700)
Elastic net 0.689 (0.676, 0.701)
Gradient boosting (imputed) 0.690 (0.677, 0.702)
Gradient boosting (unimputed) 0.686 (0.674, 0.698)
Random forests (imputed) 0.681 (0.668, 0.695)
Random forests (unimputed) 0.677 (0.665, 0.689)

Note: “Expert (original)” indicates the stratified logistic models using the same set of predictors of the original BCSC advanced cancer model; “Expert (splines)” is similar to “Expert (original)” except that splines were used to model continuous age and BMI; “Expansion” is similar to “Expert (splines)” but further added three new predictors; LASSO and Elastic net indicate logistic regressions including Expansion model predictors as candidate predictors, with model terms selected by LASSO or Elastic net; “Refitted” LASSO and Elastic net indicate logistic regressions with model terms selected by LASSO or Elastic net but re-estimated without a penalty term; “imputed” for machine learning methods indicates models fitted using imputed data; “unimputed” indicates the machine learning methods handling missing predictors directly without imputation. All logistic regression models and regularized models (LASSO and Elastic net) were fitted using imputed data.

Performance stratified by menopausal status and screening interval

The comparative performance patterns of the approaches were consistent across menopausal status and screening interval, including similar AUCs and similar comparative calibration metrics within each subgroup (Supplemental Table S7, Supplemental Fig. S2, S3). At thresholds corresponding to the top 25% of predicted risk, specificity was approximately 0.75, NPV >0.999, and PPV was uniformly low with minimal variation between methods, reflecting the rare outcome. Sensitivity varied modestly (0.44–0.47 for annual and 0.41–0.44 for biennial screens), consistent with similar AUCs across methods (Supplemental Table S8).

Variable importance

The frequency of selected predictors in unstratified LASSO and elastic-net models are summarized in Supplemental Fig. S4. The estimated effect sizes for the selected features were similar between LASSO and Elastic net (Supplemental Table S9). Breast density and BMI ranked highest for both machine learning approaches using imputed data, followed by history of breast biopsy and age at screening with moderate importance (Supplemental Fig. S5).

Discussion

We compared statistical and machine learning approaches for predicting advanced breast cancer in a large screening cohort. Discrimination was similar across methods, whereas calibration differences were more pronounced. Unstratified regularized regression provided most favorable calibration while maintaining comparable discrimination. For informing patients and medical decision-making, good calibration is a necessary criterion for model performance and fairness (59-63). Our findings highlight the importance of evaluating calibration in addition to discrimination, especially in similarly large datasets with rare outcomes and modest predictor dimensionality.

Although post-selection refitting may reduce coefficient bias, it removes the regularization that can improve predictive performance. In prediction settings, retaining shrinkage often yields better performance than refitting (64-66), consistent with our findings.

Machine learning methods did not outperform regression-type approaches despite greater flexibility. Random forests showed clear underfitting, while gradient boosting achieved similar discrimination as unstratified regularized regressions but less favorable calibration. In this clinical context, regularized regression offers a practical balance between flexibility, interpretability, and predictive performance. Our use of a large and diverse sample of women undergoing routine screening increases the applicability of these results.

This screen-level framework provides a basis for estimating longer-term cumulative risk. In future work, we will aggregate predicted risks from the validated screen-level advanced cancer and competing event models (e.g., early-stage cancer and death) within a discrete-time survival framework to estimate 6-year cumulative risk (67, 68). This extension is a deterministic aggregation and does not involve refitting or model selection using the same data.

Although internal cross-validation enabled fair comparison across approaches, external validation is needed to assess generalizability. Future work will evaluate these models in other sites of BCSC data or independent populations as our next priority.

Our results reinforce the conclusion that clinical risk models in settings with modest numbers of predictors and rare events may not benefit from machine learning approaches, aligning with prior studies showing no systematic benefit of machine learning in clinical risk prediction (7-9). By contrast, studies showing substantial gains from machine learning (69) typically involve high-dimensional data and higher-incidence outcomes. Beyond predictive performance, interpretability and ease of implementation are critical considerations in clinical risk modeling (70). Regression-based models can be disseminated via analytic equations and directly interpreted, whereas machine learning models generally require software deployment and lack transparency, posing barriers to implementation and clinical uptake for patients, providers, and health systems (13, 70).

Fairness in predictions

Fairness evaluation depend on intended use of the model (11) and different criteria cannot generally be satisfied simultaneously (71). We prioritized subgroup calibration because the models are intended for risk-based screening decisions, where accurate estimation of absolute risk within each group is critical. Systematic overprediction of risk in some groups and underprediction in others could lead to inequitable resource allocation (11, 72), translating into differences in access to more frequent screening or supplemental imaging. While discrimination was similar across models, calibration varied more across racial and ethnic groups for machine learning approaches than for regression approaches. Regularized regression demonstrated more stable calibration across subgroups, suggesting lower risk of introducing additional disparities.

All models were trained using observed diagnoses of advanced breast cancer, and therefore reflect underlying disparities in access to screening, diagnostic evaluation, and cancer ascertainment. As a result, predictive models may reproduce these structural inequities. However, prior study suggests that racial and ethnic differences in screening and diagnostic utilization contribute limited detection bias, with relative risks of disease onset broadly similar to those based on diagnosed cases (73). Addressing these disparities will likely require system-level interventions beyond statistical model refinement.

Limitations

This study has several limitations. First, machine learning performance depends on hyperparameter tuning; although optimized via cross-validation, results might vary under alternative specifications. Second, findings are specific to the study context and are likely generalizable to other similar settings with a large sample size (e.g., hundreds of thousands of observations or more), rare outcome (e.g., <1–5%) with hundreds to a few thousand outcome events, tens (rather than hundreds or thousands) of predictors with a similar mix of continuous and categorical predictors. This may include settings of risk modeling for other rare diseases or cancer types in a large screening population. Third, multiple imputation was performed in a supervised manner (including outcomes when imputing predictors, and vice versa) prior to cross-validation, which may introduce some optimism (arXiv:2010.00718). Fourth, exclusion of mammograms with screening MRI within 12 months relied on information not available at prediction time; however, such cases were rare (<0.2%) and unlikely to affect results. Fifth, menopausal status was based on a standardized BCSC algorithm using self-reported data and may be misclassified, particularly among perimenopausal women or those using hormone therapies; pregnancy status was not captured. Such misclassification may attenuate associations but is likely to have limited impact on risk estimates (74). Finally, smaller number of outcomes in some racial and ethnic subgroups could limit the reliability of subgroup-specific calibration assessment, and external validation in more diverse cohorts is warranted.

Conclusions

Regularized regression predicted advanced cancer risks that were well-calibrated with comparable discriminatory to alternative approaches. Regularized regression balances flexibility, interpretability, and ease of implementation, and holds promise for supporting risk-guided clinical decision-making in the context of breast cancer screening and other similar settings with rare outcomes and modest predictor dimensionality.

Supplementary Material

Supplementary Tables & Figures

Acknowledgements

Research reported in this work was funded by a grant from the National Cancer Institute (P01CA154292), which supported Drs. S. Chen, K. Kerlikowske, Y. Su, R.A. Hubbard, J.A. Tice, B.L. Sprague, D.L. Miglioretti. Data collection was additionally supported by the National Institute of General Medical Sciences (U54GM115516), which supported Dr. B.L. Sprague. Cancer and vital status data collection was supported by several state public health departments and cancer registries throughout the U.S. (http://www.bcsc-research.org/work/acknowledgement.html). Dr. B.L. Sprague was also supported by a grant from the National Cancer Institute (R01CA248068). Dr. T.P. Ahern was supported by a grant from the National Cancer Institute (R01CA286069). All statements in this report, including its findings and conclusions, are solely those of the authors and do not necessarily represent the views of the National Institutes of Health. We thank the participating women, mammography facilities, and radiologists for the data they have provided for this study.

Footnotes

Conflict of interest disclosure statement: The authors declare no potential conflicts of interest.

References

  • 1.Parikh RB, Manz C, Chivers C, Regli SH, Braun J, Draugelis ME, et al. Machine Learning Approaches to Predict 6-Month Mortality Among Patients With Cancer. JAMA Netw Open, 2019. 2(10): e1915997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Goldstein BA, Navar AM, Carter RE. Moving beyond regression techniques in cardiovascular risk prediction: applying machine learning to address analytic challenges. Eur Heart J, 2017. 38(23): 1805–1814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Ming C, Viassolo V, Probst-Hensch N, Dinov ID, Chappuis PO, Katapodi MC. Machine learning-based lifetime breast cancer risk reclassification compared with the BOADICEA model: impact on screening recommendations. Br J Cancer, 2020. 123(5): 860–867. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Gravesteijn BY, Nieboer D, Ercole A, Lingsma HF, Nelson D, van Calster B, et al. Machine learning algorithms performed no better than regression models for prognostication in traumatic brain injury. J Clin Epidemiol, 2020. 122: 95–107. [DOI] [PubMed] [Google Scholar]
  • 5.Nusinovici S, Tham YC, Chak Yan MY, Wei Ting DS, Li J, Sabanayagam C, et al. Logistic regression was as good as machine learning for predicting major chronic diseases. J Clin Epidemiol, 2020. 122: 56–69. [DOI] [PubMed] [Google Scholar]
  • 6.Witteveen A, Nane GF, Vliegen IMH, Siesling S, MJ IJ. Comparison of Logistic Regression and Bayesian Networks for Risk Prediction of Breast Cancer Recurrence. Med Decis Making, 2018. 38(7): 822–833. [DOI] [PubMed] [Google Scholar]
  • 7.Christodoulou E, Ma J, Collins GS, Steyerberg EW, Verbakel JY, Van Calster B. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol, 2019. 110: 12–22. [DOI] [PubMed] [Google Scholar]
  • 8.Su YR, Buist DSM, Lee JM, Ichikawa L, Miglioretti DL, Bowles EJA, et al. Performance of Statistical and Machine Learning Risk Prediction Models for Surveillance Benefits and Failures in Breast Cancer Survivors. Cancer Epidemiol Biomarkers Prev, 2023. 32(4): 561–571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Sun Z, Dong W, Shi H, Ma H, Cheng L, Huang Z. Comparing Machine Learning Models and Statistical Models for Predicting Heart Failure Events: A Systematic Review and Meta-Analysis. Front Cardiovasc Med, 2022. 9: 812276. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Mittelstadt BD, Allo P, Taddeo M, Wachter S, Floridi L. The ethics of algorithms: Mapping the debate. Big Data & Society, 2016. 3(2): 2053951716679679. [Google Scholar]
  • 11.Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann Intern Med, 2018. 169(12): 866–872. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Vyas DA, Eisenstein LG, Jones DS. Hidden in Plain Sight - Reconsidering the Use of Race Correction in Clinical Algorithms. N Engl J Med, 2020. 383(9): 874–882. [DOI] [PubMed] [Google Scholar]
  • 13.Paulus JK, Kent DM. Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities. NPJ Digit Med, 2020. 3: 99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Oni-Orisan A, Mavura Y, Banda Y, Thornton TA, Sebro R. Embracing Genetic Diversity to Improve Black Health. N Engl J Med, 2021. 384(12): 1163–1167. [DOI] [PubMed] [Google Scholar]
  • 15.Waters EA, Colditz GA, Davis KL. Essentialism and Exclusion: Racism in Cancer Risk Prediction Models. J Natl Cancer Inst, 2021. 113(12): 1620–1624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Wang Y, Wang L, Zhou Z, Laurentiev J, Lakin JR, Zhou L, et al. Assessing fairness in machine learning models: A study of racial bias using matched counterparts in mortality prediction for patients with chronic diseases. J Biomed Inform, 2024. 156: 104677. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.van der Ploeg T, Austin PC, Steyerberg EW. Modern modelling techniques are data hungry: a simulation study for predicting dichotomous endpoints. BMC Med Res Methodol, 2014. 14: 137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Nelson HD, Cantor A, Humphrey L, Fu R, Pappas M, Daeges M, et al. U.S. Preventive Services Task Force Evidence Syntheses, formerly Systematic Evidence Reviews. Screening for Breast Cancer: A Systematic Review to Update the 2009 U.S Preventive Services Task Force Recommendation. 2016, Rockville (MD): Agency for Healthcare Research and Quality (US). [Google Scholar]
  • 19.Autier P, Hery C, Haukka J, Boniol M, Byrnes G. Advanced breast cancer and breast cancer mortality in randomized controlled trials on mammography screening. J Clin Oncol, 2009. 27(35): 5919–23. [DOI] [PubMed] [Google Scholar]
  • 20.Duffy SW, Tabar L, Yen AM, Dean PB, Smith RA, Jonsson H, et al. Mammography screening reduces rates of advanced and fatal breast cancers: Results in 549,091 women. Cancer, 2020. 126(13): 2971–2979. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Hortobagyi G, Connolly J, Edge S, et al. Breast. In: Amin MB ES, Greene F, et al, eds. AJCC Cancer Staging Manual. 8th ed. 2016, New York, NY: Springer International Publishing. [Google Scholar]
  • 22.Henderson LM, Miglioretti DL, Kerlikowske K, Wernli KJ, Sprague BL, Lehman CM. Breast cancer characteristics associated with digital versus screen-film mammography for screen-detected and interval cancers. AJR Am J Roentgenol, 2015. 205(3): 676–684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Kerlikowske K, Sprague BL, Tosteson ANA, Wernli KJ, Rauscher GH, Johnson D, et al. Strategies to Identify Women at High Risk of Advanced Breast Cancer During Routine Screening for Discussion of Supplemental Imaging. JAMA Intern Med, 2019. 179(9): 1230–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Kerlikowske K, Bissell MCS, Sprague BL, Buist DSM, Henderson LM, Lee JM, et al. Advanced Breast Cancer Definitions by Staging System Examined in the Breast Cancer Surveillance Consortium. J Natl Cancer Inst, 2021. 113(7): 909–916. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Weiss A, Chavez-MacGregor M, Lichtensztajn DY, Yi M, Tadros A, Hortobagyi GN, et al. Validation Study of the American Joint Committee on Cancer Eighth Edition Prognostic Stage Compared With the Anatomic Stage in Breast Cancer. JAMA Oncol, 2018. 4(2): 203–209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kerlikowske K, Chen S, Golmakani MK, Sprague BL, Tice JA, Tosteson ANA, et al. Cumulative Advanced Breast Cancer Risk Prediction Model Developed in a Screening Mammography Population. J Natl Cancer Inst, 2022. 114(5): 676–685. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ballard-Barbash R, Taplin SH, Yankaskas BC, Ernster VL, Rosenberg RD, Carney PA, et al. Breast Cancer Surveillance Consortium: a national mammography screening and outcomes database. AJR Am J Roentgenol, 1997. 169(4): 1001–8. [DOI] [PubMed] [Google Scholar]
  • 28.Sickles EA, Miglioretti DL, Ballard-Barbash R, Geller BM, Leung JW, Rosenberg RD, et al. Performance benchmarks for diagnostic mammography. Radiology, 2005. 235(3): 775–90. [DOI] [PubMed] [Google Scholar]
  • 29.Lehman CD, Arao RF, Sprague BL, Lee JM, Buist DS, Kerlikowske K, et al. National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium. Radiology, 2017. 283(1): 49–58. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.American College of Radiology. American College of Radiology Breast Imaging Reporting and Data System Atlas (BI-RADS® Atlas). Vol. 5. 2013, Reston, VA: American College of Radiology. [Google Scholar]
  • 31.Ahn J, Schatzkin A, Lacey JV Jr., Albanes D, Ballard-Barbash R, Adams KF, et al. Adiposity, adult weight change, and postmenopausal breast cancer risk. Arch Intern Med, 2007. 167(19): 2091–102. [DOI] [PubMed] [Google Scholar]
  • 32.Kerlikowske K, Miglioretti DL, Ballard-Barbash R, Weaver DL, Buist DS, Barlow WE, et al. Prognostic characteristics of breast cancer among postmenopausal hormone users in a screened population. J Clin Oncol, 2003. 21(23): 4314–21. [DOI] [PubMed] [Google Scholar]
  • 33.Kerlikowske K, Miglioretti D, Buist D, Walker R, Carney P. Declines in invasive breast cancer and use of postmenopausal hormone therapy in a screening mammography population. J Natl Cancer Inst, 2007. 99(17): 1335–1339. [DOI] [PubMed] [Google Scholar]
  • 34.Executive summary of the clinical guidelines on the identification, evaluation, and treatment of overweight and obesity in adults. Arch Intern Med, 1998. 158(17): 1855–1867. [DOI] [PubMed] [Google Scholar]
  • 35.Dupont WD, Page DL. Risk factors for breast cancer in women with proliferative breast disease. N Engl J Med, 1985. 312(3): 146–51. [DOI] [PubMed] [Google Scholar]
  • 36.Page DL, Schuyler PA, Dupont WD, Jensen RA, Plummer WD Jr, Simpson JF. Atypical lobular hyperplasia as a unilateral predictor of breast cancer risk: a retrospective cohort study. Lancet, 2003. 361: 125–129. [Google Scholar]
  • 37.Tice JA, Miglioretti DL, Li CS, Vachon CM, Gard CC, Kerlikowske K. Breast Density and Benign Breast Disease: Risk Assessment to Identify Women at High Risk of Breast Cancer. J Clin Oncol, 2015. 33(28): 3137–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Kerlikowske K, Walker R, Miglioretti DL, Desai A, Ballard-Barbash R, Buist DS. Obesity, mammography use and accuracy, and advanced breast cancer risk. J Natl Cancer Inst, 2008. 100(23): 1724–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Fitzmaurice GM, Laird NM, Ware JH. Applied longitudinal analysis. 2012: John Wiley & Sons. [Google Scholar]
  • 40.van Buuren S Multiple imputation of discrete and continuous data by fully conditional specification. Stat Methods Med Res, 2007. 16(3): 219–42. [DOI] [PubMed] [Google Scholar]
  • 41.White IR, Royston P, Wood AM. Multiple imputation using chained equations: Issues and guidance for practice. Stat Med, 2011. 30(4): 377–99. [DOI] [PubMed] [Google Scholar]
  • 42.von Hippel PT. How Many Imputations Do You Need? A Two-stage Calculation Using a Quadratic Rule. Sociol Methods Res, 2020. 49(3): 699–718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Wood SN. Generalized additive models: an introduction with R. 2006, Boca Raton, Florida: Chapman and Hall/CRC. [Google Scholar]
  • 44.Tibshirani R. Regression shrinkage and selection via the Lasso. Journal of the Royal Statistical Society (B), 1996. 58(1): 267–288. [Google Scholar]
  • 45.Zou H, Hastie T. Regularization and Variable Selection Via the Elastic Net. Journal of the Royal Statistical Society Series B: Statistical Methodology, 2005. 67(2): 301–320. [Google Scholar]
  • 46.Javanmard A, Montanari A. Confidence Intervals and Hypothesis Testing for High-Dimensional Regression. Journal of Machine Learning Research, 2014. 15(82): 2869–2909. [Google Scholar]
  • 47.Breiman L Random forests. Machine Learning, 2001. 45: 5–32. [Google Scholar]
  • 48.Friedman JH. Greedy function approximation: a gradient boosting machine. The Annals of Statistics, 2001. 29(5): 1189–1232. [Google Scholar]
  • 49.Jin R, Agrawal G. Communication and Memory Efficient Parallel Decision Tree Construction. Proceedings of the 2003 SIAM International Conference on Data Mining (SDM), 2003: 119–129. [Google Scholar]
  • 50.Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, et al. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. Advances in Neural Information Processing Systems 30 (NIPS 2017), 2017: 3149–3157. [Google Scholar]
  • 51.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 2011. 12: 2825–2830. [Google Scholar]
  • 52.Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology, 2010. 21(1): 128–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Tice JA, Cummings SR, Smith-Bindman R, Ichikawa L, Barlow WE, Kerlikowske K. Using clinical factors and mammographic breast density to estimate breast cancer risk: development and validation of a new predictive model. Ann Intern Med, 2008. 148(5): 337–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Viallon V, Ragusa S, Clavel-Chapelon F, Benichou J. How to evaluate the calibration of a disease risk prediction tool. Stat Med, 2009. 28(6): 901–16. [DOI] [PubMed] [Google Scholar]
  • 55.Breslow NE, Day NE. Statistical methods in cancer research. Volume II--The design and analysis of cohort studies. IARC Sci Publ, 1987(82): 1–406. [Google Scholar]
  • 56.Marshall A, Altman DG, Holder RL, Royston P. Combining estimates of interest in prognostic modelling studies after multiple imputation: current practice and guidelines. BMC Med Res Methodol, 2009. 9: 57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Han H, Guo X, Yu H. Variable selection using Mean Decrease Accuracy and Mean Decrease Gini based on Random Forest. in 2016 7th IEEE International Conference on Software Engineering and Service Science (ICSESS). 2016: 219–224. IEEE. [Google Scholar]
  • 58.Niculescu-Mizil A, Caruana R. Predicting good probabilities with supervised learning. in Proceedings of the 22nd international conference on Machine learning. 2005, Association for Computing Machinery: Bonn, Germany. 625–632. [Google Scholar]
  • 59.Hedden B. On statistical criteria of algorithmic fairness. . Phil. & Pub. Aff, 2021. 49: 209. [Google Scholar]
  • 60.Steyerberg E. Clinical prediction models: a practical approach to development, validation, and updating. 2009, New York: Springer-Verlag. [Google Scholar]
  • 61.Kim KI, Simon R. Probabilistic classifiers with high-dimensional data. Biostatistics, 2011. 12(3): 399–412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Pepe M, Janes H, Methods for evaluating prediction performance of biomarkers and tests, in Risk assessment and evaluation of predictions. 2013, Springer. 107–142. [Google Scholar]
  • 63.Van Calster B, Nieboer D, Vergouwe Y, De Cock B, Pencina MJ, Steyerberg EW. A calibration hierarchy for risk models was defined: from utopia to empirical data. J Clin Epidemiol, 2016. 74: 167–76. [Google Scholar]
  • 64.Chzhen E, Hebiri M, Salmon J. On lasso refitting strategies. Bernoulli, 2019. 25(4A): 3175–3200. [Google Scholar]
  • 65.Van Houwelingen JC. Shrinkage and Penalized Likelihood as Methods to Improve Predictive Accuracy. Statistica Neerlandica, 2001. 55(1): 17–34. [Google Scholar]
  • 66.Musoro JZ, Zwinderman AH, Puhan MA, ter Riet G, Geskus RB. Validation of prediction models based on lasso regression with multiply imputed data. BMC Med Res Methodol, 2014. 14: 116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Hubbard RA, Miglioretti DL. A semiparametric censoring bias model for estimating the cumulative risk of a false-positive screening test under dependent censoring. Biometrics, 2013. 69(1): 245–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Hubbard RA, Ripping TM, Chubak J, Broeders MJ, Miglioretti DL. Statistical Methods for Estimating the Cumulative Risk of Screening Mammography Outcomes. Cancer Epidemiol Biomarkers Prev, 2016. 25(3): 513–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Couronne R, Probst P, Boulesteix AL. Random forest versus logistic regression: a large-scale benchmark experiment. BMC Bioinformatics, 2018. 19(1): 270. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Buist DSM. Factors to Consider in Developing Breast Cancer Risk Models to Implement into Clinical Care. Curr Epidemiol Rep, 2020. 7(2): 113–116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Corbett-Davies S, Gaebler JD, Nilforoshan H, Shroff R, Goel S. The Measure and Mismeasure of Fairness. Journal of Machine Learning Research, 2023. 24(312): 1–117. [Google Scholar]
  • 72.Kerlikowske K, Chen S, Sprague BL, Tice JA, Miglioretti D, Hubbard R. Effect of race and ethnicity on advanced breast cancer risk prediction model performance. npj Digital Medicine, 2025. 8: 771. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Gard CC, Lange J, Miglioretti DL, O'Meara ES, Lee CI, Etzioni R. Risk of cancer versus risk of cancer diagnosis? Accounting for diagnostic bias in predictions of breast cancer risk by race and ethnicity. J Med Screen, 2023. 30(4): 209–216. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Phipps AI, Ichikawa L, Bowles EJ, Carney PA, Kerlikowske K, Miglioretti DL, et al. Defining menopausal status in epidemiologic studies: A comparison of multiple approaches and their effects on breast cancer rates. Maturitas, 2010. 67(1): 60–6. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Tables & Figures

Data Availability Statement

The data underlying this article will be shared upon email request to the Corresponding Author or BCSC Statistical Coordinating Center (kpwa.scc@kp.org) with appropriate regulatory approvals.

RESOURCES