Abstract
The ELDER-ICU model, a machine learning tool for predicting in-hospital mortality in critically ill older adults ( ≥ 65 years), was externally validated across 12 international centers in the US, Austria, South Korea, and China, where we assessed three model updating strategies: recalibration, incremental training, and retraining. While maintaining robust performance in US and Austrian cohorts (AUROC 0.804–0.864), significant drops occurred in Asian sites (South Korea: 0.753; China: 0.698). Incremental training enhanced performance in most centers, while retraining significantly improved AUROC by 0.066 and 0.076 in the two Asian sites (South Korea and China, respectively). Isotonic regression and Platt scaling improved calibration performance globally. This study demonstrates the varying robustness of the ELDER-ICU model and the differential effectiveness of model updating strategies across temporal shifts, populations, and clinical practice environments. Rigorous validation and proactive model adaptation are essential before clinical deployment in settings with heterogeneous populations and clinical practice.
Subject terms: Diseases, Health care, Medical research, Risk factors
Introduction
The escalating global aging population has led to a disproportionate rise in older adults admitted to intensive care units (ICUs), amplifying demands for tailored risk stratification in this vulnerable cohort1. Older patients, characterized by geriatric-specific factors such as cognitive impairment, multimorbidity, and frailty, face significantly higher mortality rates compared to younger populations2,3. Timely, personalized assessment of illness severity is therefore critical to optimizing clinical outcomes and resource allocation in ICU settings.
We developed the ELDER-ICU model, an interpretable machine learning tool designed to predict in-hospital mortality in critically ill older adults (≥65 years) using clinical variables recorded during the first 24 h of ICU admission. The model was trained and externally validated using multicenter data from the United States and Amsterdam, the Netherlands, demonstrating robust performance across diverse clinical settings4. However, its generalizability during public health emergencies (e.g., COVID-19) and across geographically diverse populations (Austria, South Korea and China) remains untested. A model’s performance on its development dataset alone is insufficient to ensure clinical reliability; true generalizability requires validation in external populations distinct from the training cohort5. Thus, pre-deployment validation in new clinical environments is essential to confirm performance stability6–8. When performance degrades in new settings, updating the existing model—rather than developing a new one—offers a pragmatic solution. Common adaptation strategies include retrain, recalibration, and feature extension9–11. Previous studies have demonstrated the effectiveness of retraining and recalibration in addressing performance degradation12–14. However, the application of incremental training—a method that fine-tunes models using newly acquired data while preserving the core architecture—remains limited in critical care settings.
In this study, we aim to evaluate the robustness of the ELDER-ICU model across two dimensions of real-world deployment: temporal shifts (during the pandemic) and geographic heterogeneity (across centers in the United States, Austria, South Korea, and China). And we aim to quantify the impact of model updating on diverse clinical environments. This study serves as a use case illustrating how the generalizability of ML models may be validated and addressed by model updating.
Results
Baseline characteristics
We conducted an international multicenter cohort study utilizing data from 12 ICU sites across the U.S., Austria, South Korea, and China to systematically validate the generalizability of the ELDER-ICU model. We assessed the effectiveness of three model updating strategies—recalibration, incremental training, and retraining. Finally, we performed subgroup and distributional shift analyses to investigate model performance across different clinical subgroups and explore variations in the distributions of key predictors across diverse clinical settings. The study design is illustrated in Fig. 1. The flowchart of patient selection for each dataset is presented in Supplementary Fig. 1. Table 1 summarizes the baseline characteristics of patients from eight representative sites, with data from the remaining four US sites detailed in Supplementary Table 1. Site-specific mortality rates exhibited notable geographic variation, ranging from approximately 6% to 22%, with the lowest rates observed in South Korea and the highest in China; rates for all other sites fell between 10% and 21%. Comparative analyses between non-survivors and survivors within each site are provided in Supplementary Tables 2–13. Details regarding the randomized partitioning schema (training/calibration, validation, and test sets) for each dataset are provided in Supplementary Table 14.
Fig. 1. Study overview.
This figure shows the design of this study. a Multicenter cohort construction: we established an international retrospective cohort of critically ill adults (≥65 years) data from 12 centers across four US regions (Midwest, Northeast, South, West), Austria, South Korea, and China. b Model Validation & Updating: The ELDER-ICU model was tested in US, European, and Asian cohorts, with performance assessed via discrimination, calibration, and clinical utility metrics. Population-specific recalibration (isotonic regression and Platt scaling), incremental learning and retraining were applied, followed by re-evaluation in respective test sets. c Subsequent analysis: subgroup assessments stratified by ICU type, ethnicity, age, gender, and Charlson Comorbidity Index (CCI), and SHAP analysis quantified feature contributions and distributional shifts. The icons utilized in the figure are sourced from iconfont.cn, and they are permitted for non-commercial use under the MIT license (https://pub.dev/packages/iconfont/license).
Table 1.
Baseline characteristics of eight validation cohorts
| US (BIDMC) (N = 4175) | US (MidWest) A (N = 2070) | US (NorthEast) A (N = 1769) | US (South) A (N = 1563) | US (West) A (N = 837) | Austria (N = 6371) | South Korea (N = 1871) | China (N = 1409) | |
|---|---|---|---|---|---|---|---|---|
| Mortality | 637 (15.3) | 237 (11.5) | 297 (16.8) | 250 (16.0) | 170 (20.3) | 686 (10.8) | 120 (6.4) | 313 (22.2) |
| Setting information | ||||||||
| Location | USA_Boston | USA_MidWest | USA_Northeast | USA_South | USA_West | Austria_Salzburg | South Korea_Soul | China_Zhejiang |
| Teaching hospital | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes |
| Number of beds | >500 | >500 | >500 | >500 | 250–500 | >500 | >500 | >500 |
| Basic information | ||||||||
| Age | 75.5 (70.1,81.9) | 74.0 (69.0,80.8) | 76.0 (70.0,83.0) | 77.0 (70.0,82.0) | 76.0 (71.0,82.0) | 75.0 (70.0,80.0) | 70.0 (65.0,75.0) | 76.0 (70.0,83.0) |
| Gender | 2425 (58.1) | 1151 (55.6) | 979 (55.3) | 867 (55.5) | 469 (56.0) | 3922 (61.6) | 1141 (61.0) | 871 (61.8) |
| BMI | 27.0 (23.3,31.3) | 28.1 (24.3,32.1) | 26.1 (22.7,30.1) | 26.9 (23.4,31.4) | 28.3 (24.3,32.9) | 26.0 (23.2,29.4) | 22.5 (20.2,25.4) | 23.7 (30.0,27.6) |
| Care unit | ||||||||
| CCU | 454 (10.9) | – | 274 (15.5) | – | – | – | – | – |
| CSRU | 932 (22.3) | 1017 (49.1) | 314 (17.8) | 433 (27.7) | 325 (38.8) | – | – | – |
| MICU | 473 (11.3) | 275 (13.3) | 333 (18.8) | 550 (35.2) | 317 (37.9) | – | – | – |
| Med-SurgICU | 458 (11.0) | 5 (0.2) | 118 (6.7) | 14 (0.9) | 195 (23.3) | – | – | – |
| NICU | 1074 (25.7) | 452 (21.8) | 225 (12.7) | – | – | – | – | – |
| SICU | 386 (9.3) | 321 (15.5) | 505 (28.6) | 566 (36.2) | – | – | 1871 (100.0) | – |
| TSICU | 398 (9.5) | – | – | – | – | – | – | – |
| CCI score | 6.0 (4.0,8.0) | 4.0 (3.0,4.0) | 4.0 (4.0,4.0) | 4.00 (4.0,4.0) | 4.0 (4.0,4.0) | 4.0 (3.0,4.0) | 5.0 (4.0,6.0) | 5.0 (4.0,6.0) |
| Physical frailty | ||||||||
| Activity stand | 916 (21.9) | – | – | 46 (2.9) | – | – | – | 1 (0.1) |
| Activity sit | 1098 (26.3) | 1 (0.1) | – | 26 (1.7) | – | – | – | 10 (0.7) |
| Activity bed | 2158 (51.7) | – | 753 (42.6) | 344 (22.0) | – | – | – | 1102 (78.2) |
| GNRI | 97.7 (89.0,107.4) | 96.8 (88.4,106.1) | 97.4 (86.7,107.4) | 94.7 (85.0,105.1) | 102.6(92.0,114.7) | 94.7 (88.3,102.1) | 84.9 (77.1,92.3) | 89.7 (82.2,98.2) |
| Treatment | ||||||||
| Ventilation | 1470 (35.2) | 805 (38.9) | 704 (39.8) | 582 (37.2) | 416 (49.7) | 4539 (71.2) | 454 (24.3) | 1113 (79.0) |
| Partial code | 61 (1.5) | 322 (15.6) | 104 (5.8) | 258 (16.5) | 60 (7.1%) | – | – | – |
| Clinical score | ||||||||
| OASIS | 32.0 (26.0,38.0) | 28.0 (21.0,34.0) | 31.0 (24.0,40.0) | 26.0 (19.0,33.0) | 30.0 (23.0,38.0) | 34.0 (27.0,41.0) | 30.0 (26.0,37.0) | 31.0 (26.0,36.0) |
| SAPS II | 38.0 (31.0,47.5) | 40.0 (31.0,53.0) | 44.00 (34.0,59.0) | 41.0 (32.0,52.0) | 43.0 (33.0,57.0) | 35.0 (29.0,41.5) | 42.0 (35.0,50.0) | 39.0 (31.0,47.0) |
| SOFA | 4.0 (2.0,6.0) | 5.0 (4.0,8.0) | 6.0 (3.0,9.0) | 5.0 (4.0,8.0) | 6.0 (4.0,9.0) | 6.0 (3.0,7.0) | 7.0 (5.0,10.0) | 8.0 (6.0,10.0) |
| Outcome | ||||||||
| Days of ICU admission | 2.9 (1.7,5.4) | 2.3 (1.6,4.0) | 2.9 (1.8,5.1) | 2.7 (1.8,4.7) | 2.2 (1.5,4.2) | 2.9 (1.8,5.8) | 1.9 (1.1,3.6) | 6.1 (2.8,14.1) |
| Days of hospital admission | 8.8 (5.6,15.7) | 7.7 (4.2,13.1) | 9.8 (5.3,17.7) | 7.3 (4.6,11.6) | 6.3 (3.8,11.8) | 3.1 (2.0,6.0) | 15.0 (10.0,24.0) | 17.6 (9.7,28.8) |
| Days before ICU admission | 0.1 (0.0,0.9) | 0.1 (0.0,1.6) | 0.25 (0.1,1.9) | 0.1 (0.1,0.9) | 0.2 (0.1,1.1) | 0.1 (0.0,0.2) | 2.9 (1.8,5.7) | 0.1 (0.0,0.5) |
Data are n (%) or median (IQR). For each US region, Hospital A indicates the first hospital analyzed, with Hospital B presented in the Supplementary information. CCU coronary care unit, CSRU cardiac surgery recovery unit, MICU medical intensive care unit, ICU intensive care unit, NICU neuro intensive care unit, SICU surgical intensive care unit, TSICU trauma surgical intensive care unit, CCI Charlson comorbidity index, GNRI geriatric nutritional risk index, OASIS Oxford acute severity of illness score, SAPS II simplified acute physiology score, SOFA sequential organ failure assessment.
Discrimination performance
Figure 2 illustrates the AUROC comparison across the original model, updated model, retrained model and conventional clinical scores on eight representative individual sites. The performance of additional sites was fully reported in Supplementary Fig. 2. External validation revealed geographically heterogeneous performance patterns. The US sites and the Austrian sites generally maintained robust discrimination; In contrast, there was markedly degraded performance in both Asian sites. Incremental training improved model performance across all sites, yielding modest gains in US cohorts and substantial improvements in Austria and Asian centers (ΔAUROC: Austria: +0.022; South Korea: +0.048; China: +0.062). Conversely, full retraining yielded limited or even slightly negative effects in most US sites and only a modest gain in Austria. However, retraining markedly outperformed incremental training in Asia, achieving larger AUROC gains in both South Korea (ΔAUROC: +0.066) and China (ΔAUROC: +0.076). Detailed performance was provided in Supplementary Table 15. Statistical comparisons using the DeLong test indicated that the updated models showed significant improvements in select sites, including the US (BIDMC), Austria, and both Asian cohorts, whereas no significant differences were observed in the remaining US sites. Notably, in South Korea and China, the retrained models demonstrated significant superiority over the original model (p < 0.001). Detailed statistical results are provided in Supplementary Table 16.
Fig. 2. The discrimination performance of eight sites.
The figure shows the AUROC of the original model, updated model, retrained model and three conventional clinical scores, including SAPS II, OASIS, and SOFA. Performance metrics reflect mean ± SD over 10 random test splits; statistical comparisons (p-values) are based on the most representative split. a US (BIDMC); b US (Midwest) A; c US (Northeast) A; d US (South) A; e US (West) A; f Austria; g South Korea; h China. OASIS Oxford Acute Severity of Illness Score, SAPS II simplified acute physiology score, SOFA sequential organ failure assessment. ***P < 0.001; **P < 0.05; *P > 0.05.
Calibration performance
Analysis of Brier score comparisons demonstrated statistically significant performance improvements with both isotonic regression and Platt scaling, affirming the efficacy of model recalibration strategies (P < 0.001, Supplementary Tables 17 and 18). Notably, isotonic regression exhibited superior calibration performance in 12 out of the 14 evaluated datasets when compared to Platt scaling. Conversely, Platt scaling outperformed isotonic regression in the South Korean and Chinese cohorts. In the South Korean cohort, Platt scaling reduced the Brier score from 0.057 (SD = 0.002) to 0.055 (SD = 0.002), compared with 0.056 (SD = 0.002) achieved by isotonic regression. In the Chinese cohort, both calibration methods yielded identical point estimates (Brier score: 0.158, down from the baseline of 0.171); however, Platt scaling exhibited a narrower standard deviation (±0.002) than isotonic regression (±0.003), suggesting greater calibration stability despite equivalent central performance. Calibration curves for eight representative sites (Fig. 3) and additional sites (Supplementary Fig. 3) are provided.
Fig. 3. The Calibration plots are based on eight sites.
The dots represent quintiles of the predictions according to the predicted values, with the mean prediction of each quintile on the x-axis and the observed outcome proportion of that group on y-axis, including 95% CIs obtained by bootstrapping using 1000 bootstrap samples. Error bars indicate 95% CIs. a US (BIDMC); b US (Midwest) A; c US (Northeast) A; d US (South) A; e US (West) A; f Austria; g South Korea; h China.
Decision curve analysis
Figure 4 demonstrates the decision curve analysis (DCA) for original, recalibrated, updated and retrained ELDER-ICU models across eight sites, with additional datasets shown in Supplementary Fig. 4. In the DCA plots, the x-axis represents the threshold probability, which indicates the minimum predicted risk of in-hospital mortality at which a clinician would opt for intervention or escalated care. This threshold reflects a trade-off between the benefit of correctly identifying high-risk patients (e.g., initiating appropriate life-saving measures) and the harm of over-treating low-risk patients (e.g., allocating scarce ICU resources to those who would otherwise survive or causing unnecessary anxiety). The y-axis represents net benefit; a higher value implies that the model yields greater clinical utility at that threshold than alternative strategies. Overall, the recalibrated and updated models consistently demonstrated superior net benefit compared to the original model. Notably, the retrained model provided the greatest clinical utility in the Chinese cohorts compared to other models.
Fig. 4. The decision curve analysis on eight sites.
The x-axis represents the threshold probability, defined as the minimum predicted risk of in-hospital mortality at which a clinician would opt for intervention. The y-axis represents the net benefit; a higher net benefit indicates a better clinical utility. a US (BIDMC); b US (Midwest) A; c US (Northeast) A; d US (South) A; e US (West) A; f Austria; g South Korea; h China.
Subgroup analysis
Detailed subgroup analyses revealed context-dependent performance patterns across demographic and clinical strata (Supplementary Table 19). Sex-based differences were negligible in the US and Austrian cohorts but pronounced in the Asian sites, where AUROCs were consistently higher among females than males. Age stratification generally favored younger patients (<80 years) across most sites, with two exceptions: the US (West) A and South Korea (Supplementary Fig. 5). Ethnicity-based analyses were limited to US sites due to data availability. Among White patients, the largest and most consistently represented group, model performance remained stable. In contrast, estimates for Black, Asian, and Hispanic subgroups were highly unstable in several sites, reflecting high uncertainty due to limited sample sizes (Supplementary Fig. 6). Finally, subgroup analysis by ICU type (assessed in US cohorts) showed that surgical ICUs (CSRU, SICU, NICU) consistently demonstrated higher AUROCs than medical ICUs (MICU) (Supplementary Fig. 7). Performance based on CCI scores revealed heterogeneous patterns, with substantial differences between CCI strata observed in most sites but not in US (South) A or South Korea.
Feature importance and distribution shift analysis
Analysis of SHAP feature importance rankings based on updated models (Supplementary Figs. 8–19) demonstrated consistent patterns across 12 sites, with GCS score, total urine output, respiratory rate, mechanical ventilation use, activity status (bed), CCI score, and GNRI maintaining top-tier contributions aligned with the original model. Figure 5 demonstrates marked disparities in data and Shapley values distributions for critical predictors—GCS scores, respiratory rates, urine output, and mechanical ventilation—across geographic cohorts. While US sites exhibited relative homogeneity in GCS score, urine output and respiratory rate patterns, Asian cohorts showed significant divergence. SHAP value distributions mirrored these contrasts. Mechanical ventilation utilization further illustrated geographic contrasts: South Korea dataset had the lowest rate (24.27%), whereas China (78.99%) and Austria dataset (71.24%) significantly exceeded US averages (40.17%), a pattern mirrored in SHAP value heterogeneity. Supplementary Fig. 20 provides distribution comparisons for the additional four US sites.
Fig. 5. The distributional shifts of key features in eight sites.
a Data distribution of GCS score; b data distribution of respiratory rate; c data distribution of urine output; d data distribution of ventilation; e Shapley values distribution of GCS score; f Shapley values distribution of respiratory rate; g Shapley values distribution of urine output; h Shapley values distribution of ventilation.
Discussion
In this multinational validation across 12 sites, the ELDER-ICU model demonstrated significant performance heterogeneity-highlighting the necessity of context-aware deployment strategies. Discrimination remained robust in the US and Austrian cohorts but degraded in the Asian populations. The effectiveness of updating strategies varied by context: incremental training yielded consistent, modest gains in US sites, whereas full retraining was essential for South Korea and China but offered minimal or even slightly negative effects elsewhere. Both Platt scaling and isotonic regression significantly improved calibration. The SHAP analysis revealed consistent feature importance rankings across sites, identifying GCS, urine output, respiratory rate, mechanical ventilation use, activity status (bed), CCI score, and GNRI as key predictors.
The original model exhibited divergent performance across US, Austrian, and Asian datasets, with potential causes identifiable from our analysis. This divergence likely stems from two key factors: outcome prevalence disparities and distributional shifts in critical predictors. First, mortality rates varied substantially across regions—China cohort exhibited the highest mortality (22.21%), while South Korea reported the lowest (6.41%), contrasting sharply with US and European ranges (10–20%). Such outcome heterogeneity destabilizes calibration by altering the baseline risk landscape, which gradient-boosted models like ELDER-ICU struggle to adapt to without localized recalibration. Second, distributional shift analysis revealed marked differences in feature distributions between Asian and US populations, with key predictors such as GCS scores, total urine output, respiratory rates, and mechanical ventilation utilization exhibiting significant regional variations. Moreover, researchers have previously identified broader factors contributing to model performance degradation. Threats to generalizability extend beyond covariate shift and include: (1) temporal changes in clinical practice; (2) differences in practice between health systems; (3) patient-level biological and phenotypic heterogeneity; (4) technical variation in data capture infrastructure; and (5) broader contextual health factors, including social, environmental, political, and cultural dimensions. The observed disparities in our data can be systematically categorized under these five established factors15.
Our findings indicate that the optimal updating strategy depends on the divergence between the original and target populations and clinical environments. We introduce a pragmatic and qualitative implementation process involving a pre-deployment gap analysis of data quality (e.g., missingness in key variables), clinical practice, and population characteristics. For settings similar to the original cohort, recalibration is likely sufficient. For moderate divergence with adequate data, incremental training is recommended; For substantial clinical or demographic shifts, full retraining is optimal. Regardless of the strategy, regular reassessment using new local data is essential. Future prospective studies are needed to establish evidence-based, quantitative decision thresholds for model updating. Furthermore, we encourage adopting implementation lifecycle frameworks like TEHAI to address technical, clinical, and system adoption challenges, bridging the translational gap in medical AI16.
This study has several strengths. First, this work represents the inaugural multinational validation of the ELDER-ICU model, encompassing 12 sites across four nations through both pandemic and non-pandemic clinical contexts. This multifaceted evaluation—assessing discrimination, calibration, and clinical utility—serves as a valuable example for external validation of clinical prediction models. Second, this study introduces methodological advancements through comparative recalibration, incremental learning, and retraining approaches to facilitate continuous performance enhancement. There are also limitations to acknowledge. First, the predictive efficacy of the model exhibits heterogeneity across clinical/demographic subgroups. Future iterations will prioritize targeted refinement protocols such as stratified representation learning. Second, potential sources of bias related to insurance type and escalation limits were not systematically explored. Most critically, the ELDER-ICU model maintained robust performance across the US and Austrian sets, with performance restored in Asian sites through model updating. However, the potential for harmful self-fulfilling prophecies inherent in the clinical utilization implies that a high-performing model may inflict patient harm17. For older critically ill patients, a high predicted mortality may lead clinicians or families to forgo beneficial interventions or withdraw care prematurely, misinterpreting a probabilistic risk as a deterministic verdict. Similarly, in resource-constrained settings, models could inadvertently reinforce age-based triage biases by systematically deprioritizing older adults for ICU admission or escalation. Therefore, future efforts should focus on prospectively evaluating the model’s impact on clinical decision-making and patient outcomes. A cluster randomized trial represents the most suitable study design for this purpose, comparing ICUs using the model as a decision-support tool versus usual care18. The study would then prospectively compare key patient-centered outcomes, such as ICU length of stay, levels of decisional conflict, and the alignment of treatment with patient goals. When such RCTs are impractical, alternatives-including interviews, case-based surveys, and comparison of decisions-may provide valuable insights into how the model influences clinician behavior and patient experiences19.
In summary, the ELDER-ICU model generally performed well in US and Austrian cohorts but showed significant degradation in South Korean and Chinese populations. While incremental training is applicable generally, retraining proved superior for populations with substantial divergence. This study highlights the importance of rigorous validation and contextual model adaptation as critical considerations for deploying machine learning models across heterogeneous clinical populations.
Methods
Study design and data sources
We gathered an international multicenter retrospective cohort from critically ill patients aged ≥65 years admitted to 12 centers across four US regions (Midwest, Northeast, South, West), Austria, South Korea, and China. To evaluate the ELDER-ICU model’s generalizability, we first validated the original model in independent test sets from pandemic-era US sites, European and Asian cohorts, assessing discrimination, calibration, and clinical utility. Following external validation, we performed recalibration via isotonic regression and Platt scaling, incremental training by fine-tuning the original model and retraining on population-specific training subsets, after which updated models were re-evaluated in corresponding test sets to quantify performance improvements. Concurrently, we benchmarked model performance against conventional severity scores (SAPS II, OASIS, SOFA). Then we conducted subgroup analysis across ICU types, age groups, gender, ethnicity, and Charlson Comorbidity Index (CCI) score, while SHapley Additive exPlanations (SHAP) analysis was employed to characterize feature contributions and distributional shifts driving predictive process and outcome. The reporting of this study follows the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis using regression or machine learning methods (TRIPOD + AI) Statement20. And we applied the PROBAST + AI tool to assess the risk of bias and applicability21. These checklists are available in the Supplementary Materials.
We utilized five de-identified, publicly available datasets spanning four countries: Medical Information Mart for Intensive Care Database v3.0 (MIMIC-IV) collected from the Beth Israel Deaconess Medical Center from 2020 to 202222, eICU Collaborative Research Database (eICU-CRD-II) from 2020 to 202123, Salzburg Intensive Care database (SICdb v1.0.8) from the University Hospital Salzburg from 2013 to 202124, INformative Surgical Patient dataset for Innovative Research Environment (INSPIRE) from the Seoul National University Hospital from 2011 to 2020 and Chinese critical care database from Zhejiang Provincial People’s Hospital from 2012 to 202225,26. The MIMIC-IV v3.0 contains 364,627 unique patients with 546,028 hospital admissions and 94,458 ICU admissions. The SICdb database currently comprises 27,386 intensive care admissions from 21,583 patients. The INSPIRE database includes approximately 130,000 cases (50% of all surgical cases) that underwent anesthesia for surgery between 2011 and 2020. The Chinese critical care database comprises 8180 unique hospital admissions for 7638 individual patients from January 2012 to May 2022. To evaluate the ELDER-ICU model’s generalizability across U.S. regions during public health emergencies, we leveraged the multicenter eICU-CRD-II, which encompasses ICU data from 182 hospitals across four US regions (Midwest, Northeast, South, West). We selected the two hospitals with the largest sample sizes from each of the four US regions and denoted them as A and B, respectively, to ensure adequate statistical power while capturing regional practice variations for the primary analysis. Additionally, we created two combined datasets to provide a comprehensive assessment. First, all external validation cohorts (from the US, China, South Korea, and Austria) were combined to form a pooled dataset, representing a diverse international patient pool. Second, to reflect the relative homogeneity of healthcare systems within the country, we combined the eight sites from eICU-CRD (Midwest, Northeast, South, West) into a Pooled US Validation Cohort. The inclusion and exclusion criteria were implemented in accordance with the methodology employed in the prior publication4.
Ethics approval and patient consent
This study utilized de-identified, publicly available clinical databases and was conducted in accordance with the ethical standards of the institutional review boards (IRBs) of the corresponding data providers. All data were accessed following appropriate approval and in compliance with relevant data usage agreements. As this was a retrospective cohort study involving anonymized patient information, the requirement for individual informed consent was waived by the respective ethics committees. The study was performed in accordance with the principles outlined in the Declaration of Helsinki. The MIMIC, INSPIRE and SICdb datasets are publicly available under credentialed access on PhysioNet. The Chinese critical care database is available from the National Genomic Data Center database under accession number PRJCA006118.
ELDER-ICU model
The ELDER-ICU model is a machine learning model that leverages data from the first ICU day to provide an early assessment of disease severity by predicting in-hospital mortality for elderly patients (≥65 years). Developed and evaluated on electronic health records from multicenter cohorts, it serves as an early-warning tool for mortality risk, enabling timely initiation of goals-of-care discussions and personalized care planning for older adults at high risk of in-hospital death. It integrates sixty routinely collected clinical variables spanning six domains, including basic information, laboratory tests, vital signs, treatment, and urine output and geriatric-specific factors such as nutritional status, Charlson Comorbidity Index (CCI), physical frailty markers, and code status documented during the initial 24-h period of intensive care unit admission. The AUROCs of the model were 0.866 (95% CI: 0.851–0.880) for the internal validation set (2001–2016), 0.838 (0.829–0.847) for the external validation (US) set, 0.833 (0.812–0.853) for the external validation (Europe) set, and 0.884 (0.869–0.897) for the temporal validation set (2017–2019). An in-depth description of the model was reported in the original publication4.
Model updating and validation design
Each dataset was randomly partitioned into training/recalibration (40%), validation (10%), and test (50%) subsets for each cycle of analysis. The training/recalibration set was independently and exclusively used to (1) fit post-hoc calibration functions (Platt scaling and isotonic regression), (2) perform incremental training of the original ELDER-ICU model, or (3) train a new ELDER-ICU model from scratch in the case of full retraining. The validation set was employed during incremental training to monitor loss and prevent overfitting via early stopping. The test set remained completely held out throughout all model development and updating steps. It was used at the evaluation stage to compare the performance of the original ELDER-ICU model, the recalibrated model, the incrementally updated model, and the fully retrained model. This process was repeated 10 times using distinct random seeds. In each repetition, performance metrics were computed independently on the held-out test set. The final results are reported as the mean ± standard deviation of the 10 independently calculated values, thereby accounting for variability due to data partitioning. Outliers beyond the valid data range were excluded. Missing data were handled using multiple imputation by chained equations (MICE)-consistent with the ELDER-ICU model’s native protocol. To balance methodological rigor and practical deployment needs, five imputed datasets were generated, and a final consolidated dataset was created by averaging the imputed values across the five replicates.
The original ELDER-ICU model was first evaluated on test sets derived from pandemic-era data collected across US hospitals spanning the Midwest, Northeast, South, and West regions, as well as international cohorts from Austria, South Korea, and China. Model performance was benchmarked against conventional severity scores (OASIS, SAPS II, SOFA). To address observed performance gaps, three updating strategies were implemented: recalibration, incremental training and retraining. We employed two widely used post-hoc calibration methods, Platt scaling and isotonic regression, to recalibrate model outputs without modifying their internal parameters, ensuring predicted probabilities align with empirical outcome distributions in target populations. Platt scaling fits a logistic regression model to the original predicted probabilities (or logit scores) to produce calibrated outputs, which is suitable for small datasets. As a non-parametric, monotonicity-preserving technique, isotonic regression adjusts raw model outputs while maintaining rank-order relationships, thereby correcting calibration drift without compromising discriminative performance27,28. Specifically, we used the CalibratedClassifierCV function from scikit-learn (with cv = “prefit”), treating the original ELDER-ICU as a fixed predictor. The calibration functions were fitted on the 40% recalibration subset, and the resulting calibrated models were evaluated on the held-out 50% test set. Leveraging the inherent incremental learning capability of XGBoost, we updated the ELDER-ICU model by sequentially integrating new training data into the existing architecture, eliminating the need for resource-intensive retraining from scratch. This approach preserves learned representations from the original training phase while adapting to novel data patterns, achieving dual advantages of computational efficiency and enhanced generalizability. To prevent overfitting, we implemented early stopping through callback functions that halted training when validation performance failed to improve for 10 consecutive boosting rounds. Full retraining involved training a new ELDER-ICU model from scratch using local data, without leveraging the original model’s parameters or structure. The resulting model is thus a site-specific, independently trained instance of the ELDER-ICU architecture, reflecting the local population’s outcome distribution and feature-outcome relationships. Moreover, to ensure a fair comparison across updating strategies, neither the incrementally updated model nor the fully retrained model was subjected to additional calibration.
Model performance was evaluated across three key dimensions-discrimination (AUROC), calibration (Brier score), and clinical utility (decision curve analysis)-with both the original and updated models assessed on identical test sets to isolate the effect of each adaptation strategy. To assess and report the statistical significance of performance differences, we selected the most representative data partition, defined as the repetition in which the AUROC of the original ELDER-ICU model was closest to the mean across all 10 runs, and performed all inferential analyses on this fixed test set. These analyses include: 95% confidence intervals for performance metrics, pairwise comparisons of AUROC using the DeLong test, significance testing of Brier score differences, as well as the generation of calibration plots and decision curve analyses.
Subgroup analysis
The discriminative performance of the refined model was rigorously evaluated across clinically relevant subgroups, including: intensive care unit type (cardiac care unit [CCU], cardiac surgery recovery unit [CSRU], surgical ICU [SICU], medical-surgical ICU [Med-Surg_ICU], medical ICU [MICU], and neurosurgical ICU [NICU]); ethnicity (White, Black, Asian, Hispanic); age stratification ( < 80 years versus ≥80 years); gender (male versus female); and multimorbidity burden categorized by CCI scores (3–4 versus ≥5). Notably, subgroup analyses pertaining to ICU type and ethnicity were exclusively conducted utilizing data from U.S.-based institutions, as these demographic variables were uniquely available within this geographical cohort.
Model explanation and data distribution analysis
We employed the SHAP algorithm to quantify feature contributions to mortality predictions. To identify drivers of performance variability across sites, we analyzed distributions of four critical variables: Glasgow Coma Scale (GCS) score, respiratory rate, urine output, and mechanical ventilation. Continuous variables (GCS, respiratory rate, urine output) were normalized to [0,1)] using Min–Max scaling to eliminate unit-based discrepancies, while binary variables (mechanical ventilation) were analyzed categorically. This approach revealed significant distributional shifts across sites, correlating with observed performance degradation.
Statistical analysis
Categorical and continuous variables were reported as count (percentage) and median (IQR), respectively. Mann–Whitney U test and Fisher’s exact test were used to test the statistical significance of continuous variables and categorical variables, respectively. DeLong’s test was employed to assess the statistical significance of differences in AUROC. The Brier score was calculated as the mean squared difference between predicted probabilities and observed binary outcomes (in-hospital mortality: 0 or 1) across all patients in the test set. Because the original and recalibrated models were evaluated on the identical test set, their Brier scores form paired observations. To assess the statistical significance, we applied the Wilcoxon signed-rank test—a non-parametric test appropriate for paired, non-normally distributed data. Bootstrapping with 1000 resamples to compute the 95% Confidence interval. A two-sided p-value of less than 0.05 was considered statistically significant. All analyses were performed using Python 3.10.0 (released by the Python Software Foundation).
Supplementary information
Acknowledgements
The study was supported by the Beijing Natural Science Foundation (7252298), National Natural Science Foundation of China (82502525 and 62571550), Beijing Municipal Science and Technology Project (Z241100007724003), Project of Drug Clinical Evaluate Research of Chinese Pharmaceutical Association (NO.CPA-Z06-ZC-2021-004), a collaborative scientific project co-established by the Science and Technology Department of the National Administration of Traditional Chinese Medicine and the Zhejiang Provincial Administration of Traditional Chinese Medicine (GZY-ZJ-KJ-24082), Project of Zhejiang University Longquan Innovation Center (ZJDXLQCXZCJBGS2024016), National Institute of Health (R01 EB017205), DS-I Africa U54 TW012043-01 and Bridge2AI OT2OD032701, and National Science Foundation (ITEST #2148451).
Author contributions
M.D., X.L., W.Y., and S.H. collected data, validated and updated models, and drafted the paper. J.R. and T.L. edited the paper and checked the results. P.H., C.L., L.C., Z.L., and Z.G. provided the expertise for the validation study design. D.C., F.Z., Zho. Z., Zhe. Z., and L.A.C. designed the study and critically reviewed the core content of the paper. All authors contributed to the methodology, results analysis, and discussions of the modeling process. All authors had full access to the datasets used in this study and confirmed the fidelity of the results. All authors had final responsibility for the decision to submit for publication.
Data availability
The MIMIC and INSPIRE datasets are publicly available under credentialed access on PhysioNet (https://physionet.org/content/mimiciv/, https://physionet.org/content/inspire/1.3/). The SICdb datasets are available for research use following submission of a data access request via the PhysioNet (https://physionet.org/content/sicdb/1.0.8/). The Chinese critical care database is available from the National Genomic Data Center database under accession number PRJCA006118. The eICU-CRD-II dataset is not yet publicly available; please contact the corresponding author for access.
Code availability
Code is available at https://github.com/dmj163/ELDER-ICU-Ext-Val.
Competing interests
Zhongheng Zhang is an Editorial Board Member of NPJ Digital Medicine; he was not involved in the peer-review process or in any editorial decisions related to this paper. The remaining authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Minjie Duan, Xiaoli Liu, Wesley Yeung, Sicheng Hao.
These authors jointly supervised this work: Desen Cao, Feihu Zhou, Zhongheng Zhang, Zhengbo Zhang.
Contributor Information
Desen Cao, Email: caodesen2012@126.com.
Feihu Zhou, Email: feihuzhou301@126.com.
Zhongheng Zhang, Email: zh_zhang1984@zju.edu.cn.
Zhengbo Zhang, Email: zhangzhengbo@301hospital.com.cn.
Supplementary information
The online version contains supplementary material available at 10.1038/s41746-026-02472-1.
References
- 1.Ho, L. et al. Performance of models for predicting 1-year to 3-year mortality in older adults: a systematic review of externally validated models. Lancet Healthy Longev.5, e227–e235 (2024). [DOI] [PubMed] [Google Scholar]
- 2.Wang, S. M. et al. Association of multimorbidity patterns and order of physical frailty and cognitive impairment occurrence: a prospective cohort study. Age and Ageing54, 10 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Guidet, B. et al. The trajectory of very old critically ill patients. Intensive Care Med.50, 181–194 (2024). [DOI] [PubMed] [Google Scholar]
- 4.Liu, X. et al. Illness severity assessment of older adults in critical illness using machine learning (ELDER-ICU): an international multicentre study with subgroup bias evaluation. Lancet Digit Health5, e657–e667 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Efthimiou, O. et al. Developing clinical prediction models: a step-by-step guide. Br. Med. J.386, e078276 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Subbaswamy, A. et al. A data-driven framework for identifying patient subgroups on which an AI/machine learning model may underperform. NPJ Digit. Med.7, 334 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Davis, S. E., Dorn, C., Park, D. J. & Matheny, M. E. Emerging algorithmic bias: fairness drift as the next dimension of model maintenance and sustainability. J. Am. Med Inf. Assoc.32, 845–854 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Windecker, D. et al. Generalizability of FDA-approved AI-enabled medical devices for clinical use. JAMA Netw. Open8, e258052 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lee, C. S. & Lee, A. Y. Clinical applications of continual learning machine learning. Lancet Digit. Health2, e279–e281 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Finlayson, S. G. et al. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med.385, 283–286 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Otokiti, A. U. et al. The need to prioritize model-updating processes in clinical artificial intelligence (AI) models: protocol for a scoping review. JMIR Res. Protoc.12, e37685 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liou, L. et al. Assessing calibration and bias of a deployed machine learning malnutrition prediction model within a large healthcare system. NPJ Digit Med.7, 149 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.de Hond, A. A. H. et al. Predicting readmission or death after discharge from the ICU: external validation and retraining of a machine learning model. Crit. Care Med.51, 291–300 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Barak-Corren, Y. et al. Prediction across healthcare settings: a case study in predicting emergency department disposition. NPJ Digit. Med.4, 169 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Futoma, J., Simons, M., Panch, T., Doshi-Velez, F. & Celi, L. A. The myth of generalisability in clinical research and machine learning in health care. Lancet Digit. Health2, e489–e492 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Reddy, S. et al. Evaluation framework to guide implementation of AI systems into healthcare settings. BMJ Health & Care Informatics28, e100444 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.van Amsterdam, W. A. C., van Geloven, N., Krijthe, J. H., Ranganath, R. & Cinà, G. When accurate prediction models yield harmful self-fulfilling prophecies. Patterns6, 101229 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Leuchter, R. K., Turner, W. B. & Ouyang, D. Evaluating translational AI: a two-way moving target problem. NEJM AI2, AIp2500705 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Janse, R. J. et al. When impact trials are not feasible: alternatives to study the impact of prediction models on clinical practice. Nephrol. Dial. Transpl.40, 27–33 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. Br. Med. J.385, e078378 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Moons, K. G. M. et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. Br. Med. J.388, e082505 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data10, 1 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Pollard, T. J. et al. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Sci. Data5, 180178 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Rodemund, N., Wernly, B., Jung, C., Cozowicz, C. & Koköfer, A. Harnessing big data in critical care: exploring a new European Dataset. Sci. Data11, 320 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Lim, L. et al. INSPIRE, a publicly available research dataset for perioperative medicine. Sci. Data11, 655 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Jin, S. et al. Establishment of a Chinese critical care database from electronic healthcare records in a tertiary care medical center. Sci. Data10, 49 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.van der Laan, L., Ulloa-Pérez, E., Carone, M. & Luedtke, A. Causal isotonic calibration for heterogeneous treatment effects. Proc. Mach. Learn Res.202, 34831–34854 (2023). [PMC free article] [PubMed] [Google Scholar]
- 28.Jiang, X., Osl, M., Kim, J. & Ohno-Machado, L. Smooth isotonic regression: a new method to calibrate predictive models. AMIA Jt Summits Transl. Sci. Proc.2011, 16–20 (2011). [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The MIMIC and INSPIRE datasets are publicly available under credentialed access on PhysioNet (https://physionet.org/content/mimiciv/, https://physionet.org/content/inspire/1.3/). The SICdb datasets are available for research use following submission of a data access request via the PhysioNet (https://physionet.org/content/sicdb/1.0.8/). The Chinese critical care database is available from the National Genomic Data Center database under accession number PRJCA006118. The eICU-CRD-II dataset is not yet publicly available; please contact the corresponding author for access.
Code is available at https://github.com/dmj163/ELDER-ICU-Ext-Val.





