Abstracts
Alert fatigue remains a major barrier to the effective deployment of predictive models in emergency care, particularly in the context of rare but critical outcomes such as in-hospital mortality (IHM), which often occurs in less than 5.0% of patients admitted from the emergency department (ED). Severe class imbalance leads to low positive predictive value (PPV), undermining the clinical utility of even high-performance predictive models. To address this issue, we propose AI-TEW (Artificial Intelligence-powered Tiered Early Warning), a novel two-stage early warning framework designed to reduce false alarms and improve clinical interpretability. In Stage 1, a robust machine learning model was developed and validated using data from 174,292 ED visits across three hospitals in China and the United States. The model demonstrated strong discriminative ability for IHM prediction, achieving AUROCs ranging from 0.84 (95% CI, 0.81–0.86) to 0.91 (95% CI, 0.90–0.91) in internal and external validation cohorts. In Stage 2, AI-TEW implements a tiered risk stratification strategy by optimizing decision thresholds to prioritize high-risk patients, thereby increasing PPV from baseline levels of 9.8–18.8% to 32.5–40.5% across sites, while maintaining a high negative predictive value (NPV) of over 98% for low-risk individuals. To further refine alert precision, a knowledge-based filtering layer is introduced, leveraging large language models (LLM) to interpret patient-specific risk factors derived from SHAP (Shapley Additive exPlanations) method. Integrating explainable AI with clinical reasoning enhances contextual understanding and reduces spurious alerts, leading to an 11.53% increase in PPV in external validation (p = 0.0092 for MedGemma). By integrating improved predictive efficiency with interpretable, knowledge-informed filtering, AI-TEW reduces alert burden while supporting timely clinical intervention, demonstrating a promising approach to mitigating the impact of class imbalance in emergency risk prediction.
Subject terms: Computational biology and bioinformatics, Diseases, Health care, Mathematics and computing, Medical research, Risk factors
Introduction
Early identification of patients at risk of in-hospital mortality (IHM) is critical for timely intervention and improved outcomes in emergency department (ED) settings1,2. Delayed recognition of clinical deterioration can lead to preventable harm, making accurate risk stratification a cornerstone of emergency care3–5. In recent years, artificial intelligence (AI) and machine learning (ML) models have shown promise in predicting adverse events by analyzing large-scale electronic health record data, often achieving high discrimination as measured by the area under the receiver operating characteristic curve (AUROC)6–8.
However, high discrimination alone does not guarantee clinical utility. A major challenge of AI-driven early warning systems is their susceptibility to generate excessive false alarms, particularly when predicting rare outcomes such as IHM, which occurs in fewer than 5% of ED admissions9. This extreme class imbalance results in low positive predictive value (PPV), severely limiting the practical utility of even high-performing models. For example, Hsieh et al.10 developed an IHM prediction model with a robust AUROC of 0.81 but a PPV of only 10.55%, meaning nearly 90% of alerts were false positives. PPV is a critical metric reflecting the correctness of alerts, as false alarms can prompt unnecessary urgent responses, contributing to alert fatigue, clinician desensitization, and diminished trust in decision support systems11.
Efforts to improve PPV by prioritizing high-precision alerts often come at the cost of sensitivity, risking missed true events. Moreover, traditional machine learning algorithms tend to bias predictions toward the majority class in imbalanced datasets, overlooking rare but clinically important cases12. This issue arises not from poor model accuracy but from the inherent mathematical challenges associated with rare event prediction. Low PPV stemming from class imbalance exacerbates alert fatigue, a well-established phenomenon linked to delayed responses and missed critical events13–15.
Various strategies, including resampling and cost-sensitive learning, have been proposed to address dataset imbalance. However, these approaches frequently fail to align with the nuanced priorities of clinical decision-making or lack transparency. For instance, Goorbergh et al16. showed that over-sampling methods like SMOTE (Synthetic Minority Over-sampling TEchnique) often cause overestimation of minority class probabilities, resulting in mis-calibrated risk predictions that undermine clinical confidence. Despite an increasing number of IHM prediction models, few have systematically evaluated how class imbalance impacts clinical usability through measures like PPV, nor explored thresholding strategies that balance sensitivity and specificity under these imbalanced conditions17–19. Furthermore, most existing early warning systems operate as monolithic predictors without tiered prioritization, producing alerts that lack actionable clarity and do not align well with clinical workflows.
Explainability is increasingly recognized as vital for clinician trust and successful AI adoption. SHAP (SHapley Additive exPlanations) facilitates transparent, quantitative attribution of feature importance at the individual patient level7,20,21. Nonetheless, SHAP lacks the clinical context necessary to interpret the significance of certain risk factors—for example, why a mildly elevated creatinine may warrant heightened concern in patients with underlying renal disease. A recent systematic review by Porto et al.22 noted that while combining ML with traditional natural language processing can enhance predictive performance, it remains limited in explainability, lacking clinical context and integration of complex medical knowledge. Large language models (LLMs), trained on extensive medical corpora, offer the potential to fill this gap by generating natural language explanations that embed predictive risks within clinical narratives, differential diagnoses, and physiological reasoning23,24. Combining SHAP’s precise local interpretability with LLMs’ semantic and domain expertise enables multi-layered, tier-specific explanations that range from concise alerts for high-risk patients to detailed summaries for medium-risk cases. However, current applications rarely integrate SHAP and LLMs within a cohesive, hierarchical framework for staged clinical decision support in emergency risk prediction.
To address these challenges, we propose AI-TEW (AI-powered Tiered Early Warning), a novel two-stage framework designed for settings with severe outcome imbalance. AI-TEW introduces three key components: (1) a multitiered risk stratification strategy that progressively focuses alerts on the highest-risk patients, substantially increasing PPV; (2) SHAP-based explainability at each decision tier to enhance transparency and support clinician interpretation; and (3) a knowledge-driven filtering layer powered by LLMs that reduces false alarms by contextualizing predictions with medical reasoning. Developed and validated using data from three hospitals in China and the United States, AI-TEW illustrates that improving clinical utility in imbalanced settings requires not only predictive accurate but also thoughtful design choices—such as interpretable outputs, risk-adapted alerting, and integration of domain knowledge—that align AI outputs with clinical decision-making needs.
Results
Study cohorts and baseline characteristics
The cohorts used for model development and internal validation comprised 6364 patients from hospital H1 with an IHM of 3.65% (232/6364), 8097 patients from hospital H2 with an IHM of 3.24% (262/8097), and 159,831 patients from H3 with an IHM of 2.45% (3923/159,831). Specifically, the temporal external validation cohort consisted of 1723 patients from hospital H2 with an IHM of 5.22% (90/1723). The elevated mortality in the H2 external cohort is largely attributable to the post–COVID-19 policy transition, during which a surge of severe infections resulted in a higher-acuity patient population and, consequently, a shorter average Emergency Department Length of Stay (EDLOS). The cohort screening procedure is illustrated in Supplementary Fig. S1. Baseline characteristics of these cohorts are summarized in Table 1. The median age was 63.0 years (interquartile range [IQR], 51.0–72.0) in H1, 55.0 years (IQR, 41.0–69.0) in H2, 64.0 years (IQR, 50.0–77.0) in H3, and 58.0 years (IQR, 45.0–73.0) in the temporal external validation cohort. A more detailed comparison of patient characteristics between survivors and non-survivors across hospital sites is provided in Supplementary Table S1. For example, median systolic blood pressure was significantly lower in non-survivors than in survivors at all sites: in H1, 119.5 mmHg (IQR: 101.0–150.5) vs. 133.0 mmHg (IQR: 114.0–155.0), respectively (p < 0.001); in H2, 127.0 mmHg (IQR: 106.8–151.0) vs. 136.0 mmHg (IQR: 119.0–156.0) (p < 0.001); and in H3, 119.0 mmHg (IQR: 103.0–138.0) vs. 132.0 mmHg (IQR: 118.0–149.0) (p < 0.001). Figure 1 outlines the study design, depicting sample selection (Panel A), and AI-TEW framework to enhance clinical utility of prediction models (Panel B).
Table 1.
Baseline characteristics across three internal development cohorts and one temporal external validation cohort
| Characteristics | Five-fold internal cross validation cohorts | External validation cohort | ||
|---|---|---|---|---|
| H1 (n = 6364) in Jul 2020–Oct 2022 |
H2 (n = 8097) in Jan 2019–Oct 2022 |
H3 (n = 159,831) in 2011–2019 |
H2Ext (n = 1723) in Nov 2022–Apr 2023 |
|
| Age, years, median (IQR) | 63.0 (51.0–72.0) | 55.0 (41.0–69.0) | 64.0 (50.0–77.0) | 58.0 (45.0–73.0) |
| Older population (age > =65), No. (%) | 2931 (46.1) | 2690 (33.2) | 78,581 (49.2) | 699 (40.6) |
| Male, No. (%) | 4305 (67.6) | 5132 (63.4) | 77,809 (48.7) | 1084 (62.9) |
| Mode of arrival transport, No. (%) | ||||
| Unknown | 244 (3.8) | 141 (1.7) | 6276 (3.9) | 10 (0.6) |
| Walk-in | 1699 (26.7) | 3859 (47.7) | 70,593 (44.2) | 740 (42.9) |
| Ambulance | 2433 (38.2) | 2662 (32.9) | 82,548 (51.6) | 618 (35.9) |
| Other | 1988 (31.2) | 1435 (17.7) | 414 (0.3) | 355 (20.6) |
| Triage and acuity score, No. (%) | ||||
| Triage Level 1 (Resuscitation) | 466 (7.3) | 287 (3.5) | 17,879 (11.2) | 102 (5.9) |
| Triage Level 2 (Emergent) | 4638 (72.9) | 1255 (15.5) | 78,380 (49.0) | 241 (14.0) |
| Triage Level 3 (Urgent) | 1236 (19.4) | 6420 (79.3) | 62,984 (39.4) | 1322 (76.7) |
| Triage Level 4 (Less urgent) | 24 (0.4) | 135 (1.7) | 588 (0.4) | 58 (3.4) |
| Vital signs, median (IQR) | ||||
| Diastolic blood pressure, mmHg | 77.0 (68.0–87.0) | 81.0 (71.0–92.0) | 75.0 (64.0–85.0) | 82.0 (72.0–91.0) |
| Heart rate, bpm | 86.0 (73.0–100.0) | 85.0 (75.0–99.0) | 85.0 (73.0–99.0) | 120.0 (110.5–120.5) |
| Respiratory rate, bpm | 20.0 (20.0–20.0) | 20.0 (20.0–20.0) | 18.0 (16.0–18.0) | 20.0 (20.0–20.0) |
| Systolic blood pressure, mmHg | 133.0 (113.0–155.0) | 135.0 (119.0–156.0) | 132.0 (117.0–149.0) | 136.5 (119.0–155.0) |
| SpO2, % | 97.0 (96.0–99.0) | 99.0 (98.0–99.0) | 98.0 (97.0–100.0) | 99.0 (98.0–100.0) |
| Temperature, °C | 36.5 (36.3–36.7) | 36.5 (36.3–36.8) | 36.7 (36.4–37.0) | 36.6 (36.5–36.9) |
| Lab tests, median (IQR) | ||||
| Alkaline phosphatase, U/L | 96.5 (74.5–117.5) | 83.0 (66.8–106.0) | 90.0 (68.0–130.0) | 84.0 (66.0–106.2) |
| Anion gap, mEq/L | 5.9 (0.9–10.0) | 12.0 (8.0–16.0) | 16.0 (14.0–18.0) | 11.0 (7.0–16.0) |
| Bicarbonate, mEq/L | 22.4 (20.6–25.8) | 24.3 (21.6–26.7) | 24.0 (22.0–26.0) | 22.7 (20.3–24.6) |
| Creatinine, mg/dL | 0.9 (0.7–1.3) | 0.9 (0.7–1.1) | 0.9 (0.7–1.3) | 0.9 (0.7–1.3) |
| Glucose, mmol/L | 7.3 (6.0–9.5) | 7.4 (6.2–9.6) | 6.2 (5.3–7.9) | 7.8 (6.5–10.1) |
| Hematocrit, % | 38.3 (32.0–42.6) | 39.1 (33.8–43.2) | 37.0 (32.3–40.9) | 40.6 (35.0–45.2) |
| Hemoglobin, g/dL | 12.7 (10.5–14.3) | 13.2 (11.4–14.8) | 12.1 (10.4–13.6) | 13.4 (11.3–15.0) |
| Neutrophils, % | 75.2 (67.0–83.1) | 74.2 (63.8–83.8) | 72.8 (62.9–81.6) | 75.8 (64.7–83.9) |
| PCO2, mmHg | 29.6 (27.7–31.6) | 35.0 (29.0–42.0) | 42.0 (35.0–50.0) | 33.0 (26.2–38.8) |
| Partial thromboplastin time, sec | 36.3 (33.2–40.6) | 26.6 (23.9–29.6) | 30.4 (27.5–34.5) | 27.8 (25.5–30.6) |
| Platelet count, 109/L | 217.0 (169.5–271.0) | 219.0 (178.0–267.0) | 228.0 (174.0–291.0) | 217.0 (167.0–269.5) |
| Potassium, mmol/L | 3.8 (3.5–4.1) | 3.8 (3.5–4.2) | 4.2 (3.9–4.7) | 3.7 (3.4–4.1) |
| Prothrombin time, sec | 13.6 (13.0–14.6) | 11.6 (10.9–12.5) | 12.3 (11.2–14.9) | 12.0 (11.3–13.3) |
| Sodium, mmol/L | 136.9 (134.8–139.0) | 139.0 (136.0–141.1) | 138.0 (136.0–141.0) | 139.0 (136.0–141.0) |
| Total CO2, mEq/L | 29.0 (27.4–30.7) | 23.3 (18.9–27.3) | 26.0 (22.0–29.0) | 22.7 (16.9–26.5) |
| White blood cells, 109/L | 9.2 (7.0–11.9) | 8.9 (6.8–12.3) | 8.6 (6.5–11.7) | 8.9 (6.4–12.1) |
| pH Blood (no unit) | 7.4 (7.4–7.4) | 7.4 (7.3–7.4) | 7.4 (7.3–7.4) | 7.4 (7.3–7.5) |
| Specific disease groups, No. (%) | ||||
| Cerebrovascular disease patients | 1317 (20.7) | 958 (11.8) | 4707 (2.9) | 252 (14.6) |
| Diabetes mellitus patients | 717 (11.3) | 508 (6.3) | 9506 (5.9) | 195 (11.3) |
| Ischemic heart disease patients | 1693 (26.6) | 544 (6.7) | 4572 (2.9) | 163 (9.5) |
| Liver disease patients | 290 (4.6) | 131 (1.6) | 3081 (1.9) | 35 (2.0) |
| Pneumonia patients | 167 (2.6) | 1887 (23.3) | 8311 (5.2) | 230 (13.3) |
| Renal failure patients | 339 (5.3) | 273 (3.4) | 7195 (4.5) | 93 (5.4) |
| Respiratory failure patients | 323 (5.1) | 169 (2.1) | 8291 (5.2) | 55 (3.2) |
| Systemic disease groups, No. (%) | ||||
| Certain infectious and parasitic diseases | 168 (2.6) | 36 (0.4) | 3793 (2.4) | 19 (1.1) |
| Diseases of the genitourinary system | 458 (7.2) | 442 (5.5) | 16,813 (10.5) | 138 (8.0) |
| Diseases of the nervous system | 183 (2.9) | 170 (2.1) | 5367 (3.4) | 62 (3.6) |
| Diseases of the circulatory system | 4828 (75.9) | 2360 (29.1) | 34,883 (21.8) | 678 (39.3) |
| Diseases of the respiratory system | 1005 (15.8) | 2348 (29.0) | 14,947 (9.4) | 359 (20.8) |
| EDLOS, hours, median (IQR) | 24.5 (12.8–45.3) | 5.8 (3.4–8.9) | 6.3 (4.4–8.8) | 2.8 (1.6–4.4) |
| Boarding time, minutes, median (IQR) | 80.5 (30.4–187.9) | - | 76.0 (4.9–106.0) | - |
| Non-ICU admitted population, No. (%) | 4897 (76.9) | 7522 (92.9) | 130,634 (81.7) | 1513 (87.8) |
| ICU admitted population, No. (%) | 1467 (23.1) | 575 (7.1) | 29,197 (18.3) | 210 (12.2) |
| In-hospital survival, No. (%) | 6132 (96.4) | 7835 (96.8) | 155,908 (97.5) | 1633 (94.8) |
| In-hospital death, No. (%) | 232 (3.6) | 262 (3.2) | 3923 (2.5) | 90 (5.2) |
Data are presented as median (interquartile range) for continuous variables and frequency (percentage) for categorical variables. Missing data are reported where applicable. Cohorts include: Guangdong Provincial People’s Hospital (H1), Shenzhen Baoan District People’s Hospital (H2 and H2Ext), and Beth Israel Deaconess Medical Center (H3). The external validation cohort (H2Ext) consists of prospectively collected, temporally distinct ED visits from November 2022 to April 2023.
IQR interquartile range between the 25th percentile (Q1) and the 75th percentile (Q3), No., number, EDLOS emergency department length of stay, ICU intensive care unit, SpO2, peripheral capillary oxygen saturation, PCO2, partial pressure of carbon dioxide, Total CO2 total carbon dioxide, pH Blood potential of hydrogen in blood.
Fig. 1. Study design and the AI-Powered Tiered Early Warning (AI-TEW) framework.
A Shows the patient selection process across three multicenter cohorts, resulting in internal (H1, H2, H3) and temporal external (H2Ext) validation sets, where the internal validation sets employ a five-fold cross-validation approach for model training and testing. B Outlines the two-stage design of the AI-TEW system. Stage 1 develops artificial intelligence models to estimate mortality risk, illustrating how class imbalance in rare-event settings leads to high false alert rates and alert fatigue in current systems. Stage 2 integrates machine learning-based risk stratification (multi-layer decision-making) with LLM-driven interpretability and alert filtering to reduce false positives and enhance clinical decision support in ED workflows. (ED emergency department, IHM in-hospital mortality, AUROC area under the receiver operating characteristic, PPV positive predictive value, LLM large language model, AI artificial intelligence).
Performance of the machine learning models
Seven distinct machine learning algorithms were evaluated for predicting IHM among ED patients. Figure 2A compares the performance of these models across the three internal cohorts by five-fold cross-validation method (H1, H2, and H3). Among all methods, LightGBM consistently achieved the highest AUROC: 0.84 (95% CI, 0.81–0.86) in H1, 0.87 (95% CI, 0.85–0.90) in H2, and 0.91 (95% CI, 0.90–0.91) in H3. CatBoost and XGBoost, also demonstrated strong performance, with XGBoost achieving AUROC values of 0.81 (95% CI, 0.78–0.84) in H1, 0.87 (95% CI, 0.85–0.89) in H2, and 0.90 (95% CI, 0.90–0.91) in H3; CatBoost achieved 0.81 (95% CI, 0.79–0.84), 0.87 (95% CI, 0.84–0.89), and 0.89 (95% CI, 0.89–0.90) in H1, H2, and H3, respectively. LightGBM also attained the highest area under the precision-recall curve (AUPRC) values: 0.24 (95% CI, 0.19–0.29) in H1, 0.33 (95% CI, 0.27–0.39) in H2, and 0.31 (95% CI, 0.29–0.32) in H3. These results suggest that LightGBM exhibits superior discrimination performance in indicating IHM cases.
Fig. 2. Internal and external validation of machine learning–based in-hospital mortality prediction models.
A Compares the performance of seven distinct machine learning models across three internal validation cohorts (H1, H2, and H3) using five-fold cross-validation method. B Represents the performance in the internal cross validation and external temporal validation cohorts for the LightGBM model which had the best predictive capabilities. Evaluation metrics include the area under the receiver operating characteristic curve (AUROC), the area under the precision-recall curve (AUPRC), and decision curve analysis (DCA). (Catboost, categorical boosting, LightGBM light gradient boosting machine, XGBoost extreme gradient boosting, DTC decision tree classifier, KNN k-nearest neighbor, RF random forest, LR logistic regression, ROC receiver operating characteristic, PR precision-recall).
Due to its consistently strong performance, LightGBM was selected for temporal external validation (i.e., verifying on an independent dataset from a later period), as shown in Fig. 2B. In the external cohort, the model maintained robust discrimination with an AUROC of 0.87 (95% CI, 0.83–0.90) and an AUPRC of 0.40 (95% CI, 0.31–0.50), closely mirroring its internal validation results. Furthermore, decision curve analysis (DCA), presented in Fig. 2B (Panel b3), demonstrated that the LightGBM model provides a clinically meaningful net benefit across a wide range of threshold probabilities (0.1–0.4), significantly outperforming both the “treat all” and “treat none” strategies. This suggests strong potential for clinical utility in guiding risk-informed decision-making.
Supplementary Fig. S2 presents pairwise model comparisons using DeLong’s test with z-scores, confirming the statistical significance of performance differences. Supplementary Fig. S3 shows the evolution of SHAP feature importance and model performance with increasing feature count in H1, H2, and H3, providing insight into the feature selection process. importance and model performance change with more features. For instance, in H3, the top 50 features contributed most substantially to predictive performance, with diminishing returns observed upon inclusion of additional variables, suggesting that a compact feature set may be sufficient for robust prediction. Collectively, these findings underscore the robustness, generalizability, and clinical relevance of the LightGBM model for IHM prediction in ED settings.
Impact of class imbalance on predictive performance
Table 2 presents key clinical evaluation metrics—sensitivity, specificity, PPV and negative predictive value (NPV)—across varying decision probability thresholds for the final LightGBM model. The model was trained using a down-sampled dataset with a 1:3 positive-to-negative ratio to mitigate class imbalance during learning, but it was evaluated on real-world, highly imbalanced test datasets reflecting natural event rates: 1:26 in H1, 1:29 in H2, and 1:39 in H3. At decision thresholds determined by the Youden index (0.200 for H1, 0.175 for H2, and 0.225 for H3), the model achieved high NPVs (98.7% in H1, 99.1% in H2, and 99.5% in H3), indicating strong ability to identify low-risk patients and support safe clinical de-escalation. However, PPVs remained low (13.9% in H1, 11.0% in H2, and 9.8% in H3), revealing a high false positive burden under realistic deployment conditions.
Table 2.
Clinical performance of the in-hospital mortality prediction model at varying probability thresholds across three internal validation cohorts
| Probability cutoff |
TP | FP | TN | FN | PPV, % |
NPV, % |
Sensitivity, % |
Specificity, % |
F1 score, % |
Youden, % |
|---|---|---|---|---|---|---|---|---|---|---|
| H1 (Five-fold cross-validation): n = 6364; Jul 2020–Oct 2022 | ||||||||||
| 0.010 | 224 | 4396 | 1736 | 8 | 4.8 | 99.5 | 96.6 | 28.3 | 9.2 | 24.9 |
| 0.050 | 206 | 2724 | 3408 | 26 | 7.0 | 99.2 | 88.8 | 55.6 | 13.0 | 44.4 |
| 0.100 | 185 | 1776 | 4356 | 47 | 9.4 | 98.9 | 79.7 | 71.0 | 16.9 | 50.8 |
| 0.150 | 173 | 1306 | 4826 | 59 | 11.7 | 98.8 | 74.6 | 78.7 | 20.2 | 53.3 |
| 0.200 | 163 | 1011 | 5121 | 69 | 13.9 | 98.7 | 70.3 | 83.5 | 23.2 | 53.8 * |
| 0.250 | 149 | 834 | 5298 | 83 | 15.2 | 98.5 | 64.2 | 86.4 | 24.5 | 50.6 |
| 0.300 | 140 | 697 | 5435 | 92 | 16.7 | 98.3 | 60.3 | 88.6 | 26.2 | 49.0 |
| 0.400 | 119 | 496 | 5636 | 113 | 19.3 | 98.0 | 51.3 | 91.9 | 28.1 | 43.2 |
| 0.450 | 111 | 422 | 5710 | 121 | 20.8 | 97.9 | 47.8 | 93.1 | 29.0 | 41.0 |
| 0.500 | 102 | 357 | 5775 | 130 | 22.2 | 97.8 | 44.0 | 94.2 | 29.5 | 38.1 |
| H2 (Five-fold cross-validation): n = 8097; Jan 2019–Oct 2022 | ||||||||||
| 0.010 | 255 | 6509 | 1326 | 7 | 3.8 | 99.5 | 97.3 | 16.9 | 7.3 | 14.3 |
| 0.050 | 240 | 3866 | 3969 | 22 | 5.8 | 99.4 | 91.6 | 50.7 | 11.0 | 42.3 |
| 0.110 | 222 | 2330 | 5505 | 40 | 8.7 | 99.3 | 84.7 | 70.3 | 15.8 | 55.0 |
| 0.150 | 215 | 1866 | 5969 | 47 | 10.3 | 99.2 | 82.1 | 76.2 | 18.4 | 58.2 |
| 0.175 | 209 | 1685 | 6150 | 53 | 11.0 | 99.1 | 79.8 | 78.5 | 19.4 | 58.3 * |
| 0.250 | 190 | 1212 | 6623 | 72 | 13.6 | 98.9 | 72.5 | 84.5 | 22.8 | 57.1 |
| 0.300 | 179 | 992 | 6843 | 83 | 15.3 | 98.8 | 68.3 | 87.3 | 25.0 | 55.7 |
| 0.400 | 160 | 674 | 7161 | 102 | 19.2 | 98.6 | 61.1 | 91.4 | 29.2 | 52.5 |
| 0.450 | 148 | 541 | 7294 | 114 | 21.5 | 98.5 | 56.5 | 93.1 | 31.1 | 49.6 |
| 0.500 | 140 | 445 | 7390 | 122 | 23.9 | 98.4 | 53.4 | 94.3 | 33.1 | 47.8 |
| H3 (Five-fold cross-validation): n = 159,831; 2011–2019 | ||||||||||
| 0.010 | 3921 | 134,113 | 21,795 | 2 | 2.8 | 100.0 | 99.9 | 14.0 | 5.5 | 13.9 |
| 0.050 | 3829 | 80,080 | 75,828 | 94 | 4.6 | 99.9 | 97.6 | 48.6 | 8.7 | 46.2 |
| 0.120 | 3587 | 49,219 | 106,689 | 336 | 6.8 | 99.7 | 91.4 | 68.4 | 12.6 | 59.9 |
| 0.180 | 3402 | 36,419 | 119,489 | 521 | 8.5 | 99.6 | 86.7 | 76.6 | 15.6 | 63.4 |
| 0.225 | 3276 | 30,111 | 125,797 | 647 | 9.8 | 99.5 | 83.5 | 80.7 | 17.6 | 64.2 * |
| 0.260 | 3155 | 26,116 | 129,792 | 768 | 10.8 | 99.4 | 80.4 | 83.2 | 19.0 | 63.7 |
| 0.300 | 3037 | 22,450 | 133,458 | 886 | 11.9 | 99.3 | 77.4 | 85.6 | 20.7 | 63.0 |
| 0.400 | 2747 | 15,508 | 140,400 | 1176 | 15.0 | 99.2 | 70.0 | 90.1 | 24.8 | 60.1 |
| 0.450 | 2587 | 12,892 | 143,016 | 1336 | 16.7 | 99.1 | 65.9 | 91.7 | 26.7 | 57.7 |
| 0.500 | 2424 | 10,695 | 145,213 | 1499 | 18.5 | 99.0 | 61.8 | 93.1 | 28.4 | 54.9 |
TP true positive, FP false positive, TN true negative, FN false negative, PPV positive predictive value, NPV negative predictive value, H1 Guangdong Provincial People’s Hospital (GDPH), H2 Shenzhen Baoan District People’s Hospital (SZBA), H3 Beth Israel Deaconess Medical Center (BIDMC). Youden index is a statistical metric for evaluating the overall diagnostic ability of a binary classification model. The bold values and “*” both indicate the threshold with the highest Youden index.
Figure 3 helps explain this limitation by illustrating how class imbalance in the training data affects PPV. As the proportion of positive (i.e., IHM) cases decreases, the training PPV declines substantially. For example, in H1, the original training imbalance ratio of 1:26 (1 death per 27 patients) resulted in a training PPV of only 10.1%. When negative samples were randomly down-sampled to achieve a more balanced 1:3 ratio, the training PPV increased dramatically to 49.3%. Notably, while the test PPV remained low and largely unchanged across different training ratios when evaluated on the original imbalanced test sets (e.g., 1:26 in H1), the AUROC demonstrated robust stability and maintained a high level (e.g., >0.84), underscoring the model’s consistent discriminative performance even under severe class imbalance. This pattern was consistent across all three hospital sites (H1, H2, and H3), demonstrating that while down-sampling improves apparent performance during training, it does not translate into meaningful gains in PPV under real-world conditions where event rates remain low. This underscores the need for complementary strategies (e.g., risk stratification) to reduce false alarms and enhance clinical utility.
Fig. 3. Exploration of the impact of down-sampling ratio (samples size of positive predictive vs.
negative predictive) on training and testing stage. Panel A represents the results for H1, B represents H2, and C represents H3. Specifically, subgraphs a1 (H1), b1 (H2), and c1 (H3) represent validation during training using datasets with different down-sampling ratios, revealing that low PPV is related to data imbalance (e.g., rare events such as low IHM). Subgraphs a2 (H1), b2 (H2), and c2 (H3) represent testing on real-scale (not down-sampled) data, demonstrating the model’s low PPV challenge in real-world application. (PPV positive predictive value, NPV negative predictive value, AUROC area under the receiver operating characteristic curve).
Effectiveness of the tiered early warning approach
Figure 4A illustrates the standardized net benefit (relative to the “treat all” and “treat none” strategy) across varying threshold combinations for risk stratification: the low/medium-risk cutoff () and the medium/high-risk cutoff (). As these thresholds are adjusted, the model’s clinical utility—measured by net benefit—varies accordingly. In H2, when (, ) falls within the region bounded by those coordinates— (0.01, 0.20), (0.19, 0.20), (0.19, 0.80), (0.16, 0.84), (0.15, 0.85), and (0.01, 0.97)—the standardized net benefit exceeds zero, indicating that the stratified approach provides added value over default strategies. The corresponding net benefit curves for H1 and H3 are provided in Supplementary Fig. S4.
Fig. 4. Optimization of risk stratification thresholds (k1, k2) and clinical benefit analysis.
A Displays the standardized net benefit derived from a three-tier risk stratification strategy-low (Prisk ≤ k1), intermediate (), and high-risk ()—using decision curve analysis principles, compared to treat-all and treat-none strategies. The yellow region (a1) indicates threshold combinations yielding maximal net benefit, while the red region (a2) represents a clinically acceptable range where net benefit remains positive, supporting robust model deployment under uncertainty. B Shows the corresponding negative predictive value (NPV) for the low-risk group (b1) and positive predictive value (PPV) for the high-risk group (b2) across varying and values. This trade-off visualization enables calibration of the AI-Powered Tiered Early Warning (AI-TEW) system to simultaneously improve patient safety (via high NPV) and reduce alert burden (via high PPV), addressing key limitations of conventional single-threshold models in emergency care.
Figure 4B demonstrates the trade-offs among PPV, NPV, and patient coverage as the stratification thresholds are varied. As increases, the low-risk group expands and includes more patients, but its NPV gradually declines. Conversely, as increases, the high-risk group becomes more selective: fewer patients are classified as high-risk, but the PPV of this group improves significantly. For example, in H3, setting a low = 0.22 yields a high-risk PPV of 9.8% and captures 21.28% of all patients. In contrast, raising to 0.79 increases the PPV to 35.7%, albeit at the cost of reduced coverage (2.15% of patients). By carefully calibrating and , it is possible to simultaneously achieve an NPV > 99.0% in the low-risk group, a PPV > 40.0% in the high-risk group, and a positive standardized net benefit—proving a practical framework for clinical implementation.
Figure 5 compares the performance of three risk stratification schemes under optimized decision thresholds: a traditional two-tier model (high vs. low risk), a three-tier model (high/medium/low risk), and a refined multitiered decision-making approach that further stratifies the medium-risk group to enable graded clinical interventions. In the internal validation cohort (H2; Fig. 5B), transitioning from the two-tier to the three-tier framework increased the PPV of the high-risk group from 11.03 to 40.49%, substantially improving case enrichment and reducing the number of false positives among high-alert patients. This allows for more efficient targeting of intensive monitoring and early interventions.
Fig. 5. Threshold selection for risk stratification, and evaluation of stratification effectiveness.
A Shows the threshold division (dashed lines) for traditional binary classification (a1), risk stratification (a2) and multitiered decision approach (a3). B (H2, internal validation) and C (, external validation) show PPV variation (IHM proportions within each risk category) for traditional binary classification (b1 and c1), risk stratification (b2 and c2) and multitiered decision approach (b3 and c3). (PPV, positive predictive value; NPV, negative predictive value).
Patients classified as medium risk exhibit considerable heterogeneity and may benefit from additional clinical assessment or closer observation. To support more nuanced decision-marking and reduce physician workload, this group was subdivided into three subcategories: lower-intermediate (M-), intermediate (M), and upper-intermediate (M + ). Notably, the M+ subgroup had a markedly higher mortality rate (20.14%) compared to M (7.64%) and M- (3.22%), indicating meaningful risk stratification within the intermediate zone. This hierarchical differentiation enables a tiered clinical response—ranging from routine monitoring for M− to prompt evaluation and escalation for M + —thereby aligning prediction output with actionable care pathways.
As shown in Fig. 5C, in the external validation cohort, the PPV of the high-risk group similarly improved from 18.81 (two-tier scheme) to 37.17% after applying the three-tier stratification, demonstrating consistent performance gains across independent populations. Stratification results for H1 and H3 are provided in Supplementary Fig. S5. These findings illustrate that moving beyond binary risk prediction toward a multitiered warning system enhances both the precision and clinical utility of in-hospital mortality forecasting in emergency settings.
Enhanced precision with LLM-based alert filtering
Figure 6 evaluates the performance of six LLMs in filtering false alerts generated by the multitiered decision model for patients classified as high-risk in both the internal (H2) and external validation cohorts. The LLMs were tasked with reviewing patient-level risk alerts and classifying them as “Yes” (confirm alert), “No” (reject as false alert), or “Uncertain” (requires further clinical assessment), thereby acting as a secondary validation layer to improve alert precision. The models exhibited notable variation in their alerting tendencies. In the external validation cohort, DeepSeek-V3 issued a “Yes” response in 78.8% of cases, “No” in 0.9%, and “Uncertain” in 20.3%, indicating a conservative alerting strategy. In contrast, GPT-4.1 showed a higher propensity to affirm alerts, responding “Yes” in 90.3% of cases, “No” in 3.5%, and “Uncertain” in 6.2%, reflecting a more aggressive confirmation bias.
Fig. 6. Large language model (LLM)-based filtering of risk alerts and its impact on false alarm reduction and clinical performance.
A Shows the proportional distribution of LLM responses to AI-generated high-risk alerts: “Yes” (proceed with alert), “Uncertain” (recommend clinical review), and “No” (suppress alert as likely false positive). This demonstrates the LLM’s role in contextual interpretation and alert triaging. B Compares the positive predictive value (PPV) and sensitivity across different LLM models. The red dashed line indicates the baseline PPV after risk stratification but before LLM filtering; red bars show PPV after LLM integration—highlighting substantial improvements in alert precision. The purple line represents sensitivity in the high-risk group, with all models maintaining 90% sensitivity, ensuring that critical cases are not missed. These results illustrate how LLMs can enhance the clinical utility of AI early warning systems by reducing false alerts while preserving patient safety.
To characterize the nature of uncertain outputs from the LLM-based filter, we categorized all cases flagged as “Uncertain” during validation (see Fig. 7A). In the external validation set (), uncertainty most commonly arose from missing critical clinical features (66.7%), followed by conflicting clinical features (26.2%) and insufficient contextual information (7.1%). These categories consistently reflect scenarios where the LLM identified either incomplete or conflicting clinical data (e.g., missing laboratory results, contradictory vital signs) or deemed available evidence insufficient to corroborate the machine learning model’s high-risk prediction. The LLM’s accompanying explanatory narratives aligned closely with the conservative, information-seeking decision-making employed by clinicians in situations of genuine diagnostic uncertainty, suggesting that the model’s ‘uncertainty’ constitutes a reasoned, rather than arbitrary output. As further illustrated in Fig. 7B, under the same clinical information, different LLMs may emphasize different risk factors based on their underlying clinical knowledge, thereby arriving at divergent classifications for the same patient. For example, GPT-4.1 highlighted the absence of essential laboratory indicators (e.g., lactate or organ-function markers) and therefore issued an “Uncertain” judgment; MedGemma placed greater weight on the patient’s stable vital signs and normal creatinine, leading to a “No” classification; whereas DeepSeek-V3 focused on high-risk features such as fever, underlying malignancy, and elevated inflammatory markers (e.g., white blood cell count [WBC] and lactate dehydrogenase [LDH]), resulting in a “Yes” decision.
Fig. 7. Distribution of reasons for “Uncertain” responses from the LLM alert filter layer in the external validation cohort (H2Ext).
A Shows the reason for uncertainty most commonly arises from (a) missing critical clinical features (66.7%); (b) conflicting clinical features (26.2%); and (c) insufficient clinical context (7.1%). B Illustrates the divergent classification outcomes (Yes/No/Uncertain) across different LLMs for representative cases, highlighting model-specific reasoning patterns. (LLM Large language model).
In terms of clinical performance, all LLMs achieved high sensitivity (90%) in identifying true IHM events, indicating a low rate of missed critical cases. Importantly, the PPV value of every LLM surpassed that of the risk stratification model alone. MedGemma achieved the highest PPV in both internal (45.8%) and external (48.7%) validation, with sensitivities of 90.9% and 90.5%, respectively. This represents a significant reduction in false alerts without compromising case detection, outperforming both traditional two-tier model (PPV: 11.03% [internal], 18.81% [external]) and three-tier risk stratification (PPV: 40.49% [internal], 37.17% [external]). As shown in Supplementary Table S2, in the external validation cohort (), the MedGemma-based false alarm filtering approach significantly reduced the alert rate compared to ML-based risk stratification (z-test statistic = 2.606, p = 0.0092). These results demonstrate that integrating LLMs as a post-hoc refinement step for risk alerts can significantly improve the precision of AI-driven early warning systems, thereby reduce alert fatigue and enhance clinical trust and adoption. Supplementary Fig. S6 presents the standardized prompt template used for the question-answering task and illustrates how different LLMs provide divergent judgments and interpretations when evaluating the same IHM risk alert for an emergency-admitted patient, highlighting model-specific reasoning patterns and variability in clinical inference. A direct comparison of the ML-based risk stratification with an LLM-based direct risk assessment (Supplementary Fig. S7) revealed that LLMs used as standalone classifiers exhibit lower PPV and weaker risk separation, reinforcing their optimal role as complementary reasoning tools rather than independent prognostic models.
Interpretability and clinical alignment
Figure 8 illustrates the risk factors for patients at different risk levels (low, medium, high) identified using the SHAP method, providing personalized explanations for each patient’s risk classification. Figure 8A shows the differences in feature importance and contributions to model predictions across the three risk groups, with risk factors more pronounced in the high-risk group. As shown in Supplementary Fig. S8, elevated urea nitrogen (>20 mg/dL), elevated lactate (>2 mmol/L), low albumin (<3.5 g/dL), and low lymphocyte percentage (<20%) are all associated with increased risk. A SHAP-based analysis of missing data (Supplementary Fig. S9) confirmed that missingness patterns themselves carry clinically meaningful signals. For instance, the absence of lactate or triage SpO₂ measurements was associated with lower predicted mortality, aligning with selective testing practices in the ED. Figure 8B, C depict the predicted risk probabilities for representative patients from each risk group—low-risk patient A (4.97%), medium-risk patient B (49.5%), and high-risk patient C (85.05%)—and provide detailed personalized explanations, emphasizing the key contributing features to their respective risk classifications. These results underscore that the contribution and ranking of risk factors differ markedly by risk group, supporting more precise and personalized risk stratification. Furthermore, a web-based IHM prediction calculator is available online at https://emergency-death-risk-prediction.streamlit.app/ for real-time patient risk assessment, offering tailored risk predictions.
Fig. 8. Personalized and interpretable risk prediction across low-, medium-, and high-risk patient groups.
A Displays global feature importance using SHAP, with each dot representing a patient in the cohort. Features are ranked by mean absolute SHAP value (bar length), indicating their overall impact on IHM prediction. Feature values are color-coded: deep orange (high) to deep blue (low). B Compares individual predicted risk (star marker) against the global risk distribution (gray density curve), with colors indicating risk category (green: low; yellow: medium; red: high). C Presents SHAP waterfall plots for representative patients in each risk group, showing how individual features cumulatively shift the prediction from baseline (expected value) toward survival (blue contributions) or non-survival (red contributions). D Demonstrates how the LLM enhances interpretability by generating concise, clinically meaningful natural language explanations from structured SHAP outputs, enabling intuitive understanding for frontline clinicians. (SHAP SHapley Additive exPlanations, IHM in-hospital mortality, Vitals vital signs, SpO₂ peripheral oxygen saturation, DBP diastolic blood pressure, SBP systolic blood pressure, RDW red cell distribution width. Multiple vital sign measurements are labeled with “last” indicating the most recent value.).
In addition to SHAP-based interpretability, LLMs also offer enhanced explainability through natural language generation (see Fig. 8D). Specifically, LLMs can articulate complex model outputs in plain language, providing clinicians with intuitive and actionable insights. For instance, when evaluating a high-risk alert, an LLM might generate a narrative explanation such as: “This patient presents has multiple critical abnormalities, including severe lactic acidosis (lactic acid 18.5 mmol/L), profound hypoalbuminemia (albumin 1.79 g/dL), markedly low bicarbonate (6.8 mmol/L), acute kidney injury (creatinine 3.35 mg/dL), these findings indicate significant organ dysfunction, placing the patient in high-risk category.” Such narratives facilitate better understanding and trust in AI-driven predictions, ultimately improving clinical decision-making.
Discussion
Faced with the challenge of alert fatigue in EDs—a consequence of low PPV when predicting rare but critical outcomes such as IHM—this study introduces AI-TEW (AI-powered tiered early warning), a two-stage framework designed to improve the clinical utility of risk prediction under severe class imbalance. By combining optimized risk stratification with an interpretable knowledge-informed filtering layer, AI-TEW generates more precise alerts than conventional risk scoring approaches, supporting timely clinical review in resource-constrained ED settings.
First, our model demonstrates robust generalizability across heterogeneous healthcare systems and patient populations. Trained and validated using data from three distinct EDs in China and the United States, it achieved consistently high discrimination, with AUROC values of 0.84 (95% CI, 0.81–0.86), 0.87 (95% CI, 0.85–0.90), and 0.91 (95% CI, 0.90–0.91) in internal validation for H1, H2, and H3, respectively (Fig. 2A). Notably, the model maintained strong performance in temporal external validation (AUROC: 0.87; 95% CI, 0.83–0.90) which verifying on an independent dataset from a later period, underscoring its stability over time and adaptability to evolving clinical. This cross-institutional consistency highlights the model’s resilience to variations in patient demographics, data collection protocols, and care delivery models, supporting its potential for broad implementation in global emergency care settings. Leveraging a comprehensive set of routinely collected variables from electronic medical records (such as demographics, vital signs, and laboratory results), the model integrates seamlessly into existing workflows without additional data entry or specialized infrastructure. To bridge the gap between complex AI and frontline care, we developed a web-based IHM prediction calculator using Streamlit, providing real-time, personalized risk assessment with intuitive visualizations (e.g., SHAP plots) and tiered alerts to prioritize high-risk patients.
Second, a central innovation of AI-TEW lies in its explicit optimization for clinical utility—particularly the mitigation of low PPV, a long-standing limitation in rare-event prediction within ED settings. While prior studies25 have reported strong discrimination (AUROC 0.87–0.93) for IHM prediction, few have addressed the consequent high false-positive rates that undermine clinical trust and exacerbate alert fatigue. In our internal cohort (H2), the baseline model yielded a PPV of only 11.0% (Table 2), implying that nearly 9 out of 10 high-risk alerts were false alarms—a scenario that exacerbates alert fatigue and erodes clinician trust. To mitigate this, we introduced a multitiered risk stratification strategy, categorizing patients into low-, medium-, and high-risk groups under optimized thresholds (Fig. 4). This approach substantially improved PPV for high-risk patients (from 11.0 to 40.5% in H2) while enhancing NPV (~99.37%) for low-risk individuals (Fig. 5), enabling safer de-escalation of care and more efficient allocation of limited resources. The reported PPV of 40.5% reflects a deliberate balance between alert precision and operational feasibility, rather than a ceiling of the model’s discriminative ability. The selected threshold (= 0.790) identifies a clinically manageable high-risk cohort (2.01% of admissions)—thereby striking a deliberate balance between precision and operational feasibility in a real-world ED setting.
Further refinement was achieved through the integration of LLMs as post-hoc alert filters. By leveraging their contextual reasoning and medical knowledge, LLMs evaluated high-risk alerts and classified them as “Yes”, “No” or “Uncertain”. All tested LLMs preserved high sensitivity (90%), ensuring critical cases were not missed, while significantly boosting PPV. Notably, MedGemma achieved the highest performance, with PPVs reaching 45.8% (internal) and 48.7% (external)—surpassing both traditional models and risk stratification alone (Fig. 6). While only MedGemma demonstrated statistically significant performance (p = 0.0092) among the six LLMs tested (Supplementary Table S2), the overall trend suggests that LLMs may hold untapped potential for improving clinical alert systems. This dual-layer architecture—statistical risk stratification followed by knowledge-guided LLM filtering—offers a promising pathway to reduce false alerts without compromising case detection.
Third, beyond predictive performance, AI-TEW prioritizes interpretability and personalization to support informed clinical action. Rather than reporting only global feature importance26–30, our approach uses SHAP analysis to identify patient-specific risk drivers and clinically meaningful thresholds (e.g., lactate >2 mmol/L, blood urea nitrogen >20 mg/dL), aligning with established medical knowledge. Risk factors outside normal ranges were markedly concentrated in the high-risk stratum, validating the biological plausibility of our stratification. At the individual level, the system generates tailored explanations illustrating how combinations of features shape each patient’s risk profile—providing actionable insight that complements traditional population-level metrics like odds ratios, which are ill-suited for nonlinear models31. Furthermore, LLMs enhance interpretability through natural language generation. For example, an alert might be accompanied by: “This patient presents with severe lactic acidosis (7.4 mmol/L), acute kidney injury (creatinine 1.62 mg/dL), hypoalbuminemia (2.96 g/dL), tachypnea (respiratory rate 25 breaths/min), and low bicarbonate (18 mmol/L), indicating multi-organ dysfunction consistent with a high-risk mortality alert.” Such narrative outputs translate complex model logic into clinician-friendly reasoning, fostering transparency, trust, and safer adoption.
Despite these advances, several limitations warrant acknowledgement. First, the model is trained exclusively on patient-level clinical data and does not incorporate system-level factors such as admission policies, staffing levels, or bed availability. Although its strong AUROC across sites suggests generalizable risk patterns, PPV—which depends on local event prevalence—will vary by institution. Therefore, the decision thresholds (k1, k2) must be locally calibrated prior to deployment. Second, while IHM is a validated proxy for overall severity among undifferentiated ED patients, it is a distal and multifactorial endpoint. Future iterations should target more proximal, intervention-linked outcomes (such as cardiac arrest or septic shock) to better align predictions with specific clinical protocols. Third, although LLMs improve alert precision, they introduce risks of automation bias32. Our design mitigates this by restricting the LLM to an auxiliary role—never as an autonomous decision-maker—and by including an “Uncertain” category that mandates human review. However, current LLMs lack well-calibrated confidence estimates. Future work should explore uncertainty quantification and domain-specific fine-tuning to better align LLM reasoning with clinical safety standards. Fourth, the current Streamlit-based calculator represents an initial proof-of-concept implementation and does not yet incorporate LLM-based refinement. This reflects a deliberate, phased development strategy that prioritizes retrospective validation before prospective clinical integration; LLM-enhanced functionality is planned pending institutional review and API stability. Finally, prospective, multi-center implementation studies will be necessary to assess AI-TEW’s real-world impact—particularly its effectiveness in reducing alert fatigue among frontline clinicians. While PPV and NPV provide useful alert precision, they do not reflect key human factors such as cognitive load, workflow disruption, or perceived trustworthiness, which are critical to successful clinical adoption.
In conclusion, we present AI-TEW, a novel two-stage early warning framework designed to mitigate alert fatigue in emergency care—a challenge driven by the severe class imbalance inherent in predicting rare outcomes such as IHM. By integrating optimized risk stratification with a knowledge-informed filtering layer powered by LLM and explainable AI, AI-TEW achieves substantial improvements in PPV while maintaining high sensitivity, as demonstrated across three hospitals in China and the United States. Our findings suggest that combining machine learning with interpretable, clinically grounded post-processing can reduce false alarms and enhance the contextual relevance of risk alerts. With consistent performance in both internal and external validation cohorts, AI-TEW offers a practical and adaptable approach to improving early warning systems in resource-constrained emergency departments, particularly where outcome rarity limits the utility of conventional predictive models.
Methods
Ethical approval
This study was approved by the Institutional Review Board (IRB) and the Data Review Board of Guangdong Provincial People’s Hospital (IRB No. KY-N-2022-105-01) and Shenzhen Baoan District People’s Hospital (IRB No. 2023041712495656) in the People’s Republic of China. Due to the retrospective study design and the use of fully deidentified patient information, the requirement for individual informed consent was waived by the IRB. Additionally, the study utilized the publicly available Medical Information Mart for Intensive Care (MIMIC-IV) database (version 2.2). The use of MIMIC-IV was approved by the IRB of Beth Israel Deaconess Medical Center (BIDMC), which granted a waiver of informed consent for research use of the dataset. Access to MIMIC-IV was obtained under credentialed access via PhysioNet33 (Project Certification Number: 34034170).
Study population and design
This multicenter, retrospective cohort study included 1,388,586 ED visits from three institutions: Guangdong Provincial People’s Hospital (H1; =163,563 visits, July 2020 to October 2022) and Shenzhen Baoan District People’s Hospital (H2; =777,311 visits, January 2019 to April 2023) in China, and Beth Israel Deaconess Medical Center (H3; =447,712 visits, 2011–2019) in the United States. The study population was restricted to adult patients (aged 18 years) who were transferred from the ED to inpatient wards. Exclusion criteria included: (a) age < 18 years; (b) absence of essential emergency records (e.g., diagnostic information); (c) missing transfer admission data; and (d) illogical temporal sequences (e.g., discharge time preceding admission time). After applying these criteria (see Supplementary Fig. S1), a total of 174,292 ED visits were included in the final analysis. The primary objective was to develop and externally validate a machine learning-based multitiered risk stratification model for IHM prediction in ED patients. Model development and internal validation were performed using data from H1, H2 (prior to November 2022), and H3. Internal validation was further supplemented with five-fold cross-validation to evaluate model stability. Temporal external validation was conducted on an independent, prospective-seeming cohort from Shenzhen Baoan District People’s Hospital (H2Ext; = 1723 visits, November 2022 to April 2023), which was temporally distinct from the model training period and used to assess real-world generalizability (see Fig. 1A).
Data collection and preprocessing
Patient-level data were extracted from the electronic medical record system and included demographics (age, gender, mode of arrival: unknown, walk-in, ambulance, or other), triage acuity scores (level-1: resuscitation, level-2: emergency, level-3: urgent, and level-4: less urgent), vital signs (systolic and diastolic blood pressure, respiratory rate, heart rate, temperature, and peripheral oxygen saturation), laboratory results (e.g., creatinine, glucose, hematocrit, hemoglobin, alkaline phosphatase, anion gap, and bicarbonate), clinical disease diagnoses defined by the International Classification of Diseases with 9th and 10th Revisions (ICD-9 and ICD-10) 34,35, timestamps for key clinical events (e.g., triage time, admission time, and discharge time), and patient severity and outcomes (inpatient unit: ICU or general ward; patient outcome at discharge: survival or in-hospital death).
To ensure data quality and analytical consistency, a standardized preprocessing pipeline was implemented. Categorical variables with missing values were assigned a distinct “unknown” category to preserve information about data completeness. For continuous variables, missing data were retained during modeling with tree-based ensemble methods, which natively handle missingness through mechanisms such as surrogate splits or gradient-aware optimization. For models requiring complete data (e.g., Logistic Regression), missing values were imputed using the median (for skewed distributions) or mean (for approximately normal distributions). Outliers in continuous variables, such as vital signs and laboratory data, were identified using a method based on the interquartile range (values < Q1 − 1.5× [Q3 − Q1] or > Q3 + 1.5× [Q3 − Q1]) and further evaluated against clinically plausible thresholds derived from domain knowledge36. Extreme but biologically implausible values were replaced with the nearest clinically acceptable boundary value (a method known as winsorization). Categorical variables (e.g., mode of arrival) were one-hot encoded to facilitate numerical processing. Temporal variables were standardized to compute clinically relevant intervals (e.g., time from triage to admission). ICD-9 codes were systematically mapped to ICD-10 to ensure diagnostic consistency across datasets34,35. Records with critical missing data (e.g., absent triage timestamps) or logical inconsistencies (e.g., discharge before admission) were excluded prior to analysis.
AI-TEW framework development
The Artificial Intelligence-powered Tiered Early Warning (AI-TEW) system is a structured decision-support framework designed to enhance the clinical utility of IHM prediction in emergency care. Rather than relying solely on raw risk scores, AI-TEW integrates a sequential pipeline to enhance both precision and interpretability. Initially, a machine learning model generates risk estimate using routinely collected clinical data. Next, SHAP-based explanations provide personalized, feature-level attribution to support transparent decision-making. A tiered alert mechanism then applies risk-stratified thresholds to prioritize high-concern cases and reduce low-value notifications. Importantly, LLMs are incorporated to generate natural language rationales and function as a knowledge-based filter, assessing the clinical plausibility of alerts and further minimizing false positives. This hierarchical approach ensures that alerts are not only statistically robust but also clinically relevant. The following sections detail each component of the AI-TEW framework.
Machine learning for initial risk prediction
For the initial risk prediction component of the AI-TEW framework, we evaluated multiple machine learning algorithms to identify the optimal model, including Categorical Boosting (CatBoost)37, Light Gradient Boosting Machine (LightGBM)38, eXtreme Gradient Boosting (XGBoost)39, Decision Tree Classifier (DTC)40, Random Forest (RF)41, K-Nearest Neighbors (KNN)42, and Logistic Regression (LR)43. Model selection was guided by performance across key metrics—particularly AUROC, sensitivity, and PPV—with emphasis on high-risk case detection. Hyperparameter tuning was performed using Bayesian optimization, and to ensure robustness and minimize variability from data partitioning, internal validation was conducted using five-fold cross-validation44. The final model was selected based on its superior generalizability, stability across validation folds, and computational efficiency, ensuring reliable risk estimation for downstream tiered alerting and interpretation components.
SHAP-based personalized risk interpretation
To enhance model interpretability and support clinical transparency, SHAP45 were applied to quantify the contribution of each input feature to individual risk predictions. By computing locally accurate, game-theoretic attributions, SHAP enables personalized interpretation of the model’s output—highlighting which clinical variables (e.g., elevated lactate, low systolic blood pressure) most strongly drive high-risk alerts for a given patient. This post-hoc explainability method provides both global model insights and instance-level explanations, facilitating trust and informed clinical review. For practical deployment and real-time decision support, the model was implemented as an interactive web application using the Streamlit framework (https://streamlit.io/). The interface not only delivers real-time IHM risk estimates but also visualizes key risk-influencing factors for each patient, enabling clinicians to rapidly assess the rationale behind AI-generated alerts.
Tiered alert system for mitigating alert fatigue
To address class imbalance and reduce false alarm rates, we implemented a resampling strategy46 during model development, systematically varying the positive-to-negative sample ratio () to evaluate its impact on predictive performance. This enabled optimization of model calibration under realistic prevalence conditions. For clinical deployment, we designed a tiered alert system to support risk-stratified decision-making and mitigate alert fatigue. Three risk categories were defined using two clinically interpretable thresholds: a low/medium-risk cutoff () and a medium/high-risk cutoff (), resulting in: low-risk group ( ) characterized by high NPV, supporting safe expedited discharge or observation; medium-risk group ( ) representing uncertain cases requiring further evaluation; and high-risk group ( ) with high PPV, prioritized for immediate intervention. Thresholds and were calibrated using precision-recall trade-offs to maximize clinical utility—minimizing unnecessary alerts in low-risk patients while ensuring high sensitivity for critical cases. To enhance nuance in intermediate-risk management, an optional refinement subdivides this group into lower-intermediate, intermediate, and upper-intermediate subgroups, enabling graded monitoring intensity and improving PPV in higher subcategories.
a) Utility function definition. To evaluate the clinical net benefit of the multitiered risk stratification framework, we extended standard decision curve analysis (DCA)47,48 to incorporate three distinct clinical pathways: immediate intervention for high-risk patients, continued observation for medium-risk patients, and discharge or minimal monitoring for low-risk patients. The utility function assigns values based on the clinical consequences of correct and incorrect classifications, with representing the maximum benefit of correctly identifying and intervening on a high-risk case ( and ). The cost of a false positive in the high-risk group is defined as , reflecting the burden of unnecessary intervention, while the cost of a false negative in the low-risk group is , representing the harm of prematurely discharging a patient who later deteriorates. For medium-risk cases, delayed identification yields a reduced benefit where accounts for the time-sensitive nature of the condition, and over-observation incurs a lower cost with reflecting less resource utilization than full intervention. In this study, and were used as weighting coefficients and were set to 0.6 and 0.1, respectively. The complete utility function for each patient with true outcome and predicted risk is therefore defined as:
| 1 |
This formulation enables a clinically meaningful comparison of net benefit across different risk-based decision strategies.
b) Total net benefit calculation. The total net benefit () for the multitiered model can be calculated as follows:
| 2 |
where is the total number of patients.
c) Comparison with extreme strategies. We compare our multitiered model against two extreme strategies: “treat all” and “treat none”. Under the “treat all” strategy, all samples are classified as high-risk (i.e., intervention). If = 1, it is a TP with a benefit of ; if = 0, it is a FP with a loss of . Thus:
| 3 |
Under the “treat none” strategy, all patients are classified as low-risk (i.e., no treatment). If = 1, it is a FN with a loss of ; if = 0, it is a TN with a utility of 0. Thus:
| 4 |
d) Standardized net benefit. To quantify the clinical utility of our multitiered model relative to these extreme strategies, we calculate the standardized net benefit ():
| 5 |
A positive SNB indicates that our model provides greater clinical value than either extreme strategy. By enabling rapid discharge or observation of low-risk patients and focusing resources on high-risk cases, context-specific alerts were designed to trigger targeted clinical actions.
LLM for knowledge-based filtering and interpretability
To enhance the specificity of high-risk alerts, multiple LLMs—including open-source models (MedGemma49, DeepSeek-V350,, Baichuan-M151, LLaMA-3.352, Qwen-353) and the closed-source GPT-4.154—were integrated for their clinical language understanding. Detailed parameter configurations and selection rationale are provided in Supplementary Table S3. Each model processed structured inputs containing patient demographics, vital signs, laboratory results, and clinical history through a consistent prompting framework (see Supplementary Fig. S6), and was tasked with answering the structured question: “Based on your clinical knowledge, please assess the model’s risk alert level: if you agree with the alert level, respond with ‘Yes’; if you believe the alert is inaccurate, respond with ‘No’; if you’re uncertain, respond with ‘Uncertain’.” Models generated responses in a uniform schema, providing a categorical decision (“Yes”, “No”, or “Uncertain”) along with a concise, natural language rationale. This dialog-based approach leverages the LLM’s capacity to integrate multisource clinical data and apply contextual medical reasoning, effectively serving as a knowledge-driven filter to distinguish clinically meaningful alerts from spurious ones. Crucially, the generated rationales provide transparent, interpretable justifications for each decision, reducing the cognitive burden of manual validation for clinicians and supporting more efficient, trustworthy alert management.
Statistical analysis
Categorical variables were compared using the Chi-square test55 and reported as frequencies and percentages. Continuous variables were analyzed using the t-test56 for normally distributed data or nonparametric tests (e.g., Kruskal-Wallis test57) for non-normally distributed data, with results presented as mean (standard deviation) or median (interquartile range, IQR). Model discrimination was evaluated by constructing receiver operating characteristic (ROC) and precision-recall (PR) curves, and quantified by the area under the ROC curve (AUROC) and the area under the PR curve (AUPRC). AUPRC was included because it provides a more informative assessment of classifier performance in imbalanced datasets, where the number of negative outcomes substantially exceeds that of positive cases. To account for uncertainty, 95% confidence intervals (CIs) for AUROC and AUPRC were estimated using bootstrapping with 1000 resampling iterations on the test set. Sensitivity, specificity, PPV, NPV, and F1 score were computed across a range of risk thresholds to support clinically relevant interpretation and facilitate threshold selection in practice. Decision curve analysis (DCA) was performed to evaluate the clinical utility of the model by quantifying the net benefit across varying threshold probabilities. Differences in alert burden were statistically assessed by comparing alert rates between methods using a two-proportion z-test. A two-tailed p-value < 0.05 was considered statistically significant.
Supplementary information
npj_IHM_Supplementary Materials_20260204
Acknowledgements
This research was supported by the National Key R&D Program of China-Intergovernmental Key Projects (No. 2023YFE0114300), Guangdong Natural Science Foundation General Project (No. 2024A1515012112), Guangdong Medical Research Fund Project (No. A2024044), and National Natural Science Fund of China (No.82302462), Guang Dong Basic and Applied Basic Research Foundation (No. 2022A1515111206). The funders of the study had no role in study design, data collection, data analysis, data interpretation, or writing of the report. We thank the collaborating hospitals for providing access to the clinical datasets that made this study possible. We gratefully acknowledge the editors and reviewers for their helpful and perceptive feedback.
Author contributions
L.Wu. had full access to all the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis.Concept and design: A.B. and L.Wu. Acquisition, analysis, or interpretation of data: L.Wu., L.M., H.W., J.H., X.H., and X.Z. Drafting of the manuscript: L.Wu., L.M., and H.W. Critical revision of the manuscript for important intellectual content: A.B., L.Wu., X.L., J.L., and H.L.Revision of the manuscript: A.K., V.A.K., J.P., D.T., S.H.I., S.W.L., G.S., N.R., K.T., A.S., M.C., M.Q., H.H., S.Gro., B.Hu., H.Wa., B.He., P.L., B.O., S.Gem., L.K., E.L., J.L., H.L., and X.L. Statistical analysis: L.Wu., L.M., and J.H. Administrative, technical, or material support: L.Wu., L.M., and J.H. Supervision: A.B., L.Wu., X.L., H.L., and J.L. All authors have read and approved the manuscript.
Data availability
Data sources include the public MIMIC dataset and restricted Chinese clinical data (not publicly available due to privacy regulations). A lightweight, standalone version of the risk prediction calculator (implementing the core machine-learning tier without LLM filtering) is available for interactive use at: https://emergency-death-risk-prediction.streamlit.app/. The core code has been openly shared on GitHub (https://github.com/AI-Medicine-team/AI-TEW-Framework).
Code availability
Data sources include the public MIMIC dataset and restricted Chinese clinical data (not publicly available due to privacy regulations). A lightweight, standalone version of the risk prediction calculator (implementing the core machine-learning tier without LLM filtering) is available for interactive use at: https://emergency-death-risk-prediction.streamlit.app/. The core code has been openly shared on GitHub (https://github.com/AI-Medicine-team/AI-TEW-Framework).
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Lijuan Wu, Liyi Mai, Hongnian Wang.
Contributor Information
Jinle Lin, Email: lljjll2010@126.com.
Huiying Liang, Email: lianghuiying@gdph.org.cn.
Xin Li, Email: sylixin@scut.edu.cn.
Abdelouahab Bellou, Email: abellou402@gmail.com.
Supplementary information
The online version contains supplementary material available at 10.1038/s41746-026-02522-8.
References
- 1.Tamara, M. O. et al. Mortality after emergency department discharge: an analysis of 453599 cases. Emergencias36, 168–178 (2024). [DOI] [PubMed] [Google Scholar]
- 2.Yuan, S., Yang, Z., Li, J., Wu, C. & Liu, S. AI-Powered early warning systems for clinical deterioration significantly improve patient outcomes: a meta-analysis. BMC Med. Inform. Decis. Mak.25, 203 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lin, M. P. et al. Potential diagnostic error for emergency conditions, mortality, and healthy days at home. JAMA Netw. Open8, e2516400 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Honarmand, K. et al. Society of critical care medicine guidelines on recognizing and responding to clinical deterioration outside the ICU: 2023. Critic. Care Med. 52, 314–330 (2024). [DOI] [PubMed]
- 5.Sax, D. R., Warton, E. M., Mark, D. G. & Reed, M. E. Emergency department triage accuracy and delays in care for high-risk conditions. JAMA Netw. Open8, e258498 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Son, B. et al. Improved patient mortality predictions in emergency departments with deep learning data-synthesis and ensemble models. Sci. Rep.13, 15031 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Boulitsakis Logothetis, S., Green, D., Holland, M. & Al Moubayed, N. Predicting acute clinical deterioration with interpretable machine learning to support emergency care decision making. Sci. Rep.13, 13563 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Hsu, C. C. et al. Using artificial intelligence to predict adverse outcomes in emergency department patients with hyperglycemic crises in real time. BMC Endocr. Disord.23, 234 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Goldstein, B. A. & Bedoya, A. D. Guiding clinical decisions through predictive risk rules. JAMA Netw. Open3, e2013101 (2020). [DOI] [PubMed] [Google Scholar]
- 10.Hsieh, M. J. et al. Developing and validating a model for predicting 7-day mortality of patients admitted from the emergency department: an initial alarm score by a prospective prediction model study. BMJ Open11, e040837 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Bedoya, A. D. et al. Minimal impact of implemented early warning score and best practice alert for patient deterioration. Critic. Care Med. 47, 49–55 (2019). [DOI] [PMC free article] [PubMed]
- 12.Salmi, M., Atif, D., Oliva, D., Abraham, A. & Ventura, S. Handling imbalanced medical datasets: review of a decade of research. Artif. Intell. Rev.57, 273 (2024). [Google Scholar]
- 13.Cartus, A. R., Samuels, E. A., Cerdá, M. & Marshall, B. D. L. Outcome class imbalance and rare events: an underappreciated complication for overdose risk prediction modeling. Addiction118, 1167–1176 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ruppel, H., Dougherty, M., Bonafide, C. P. & Lasater, K. B. Alarm burden and the nursing care environment: a 213-hospital cross-sectional study. BMJ Open Qual.12, e002342 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Anderson, H. R. et al. Stats on the desats: alarm fatigue and the implications for patient safety. BMJ Open Qual.12, e002262 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.van den Goorbergh, R., van Smeden, M., Timmerman, D. & Van Calster, B. The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression. J. Am. Med. Inform. Assoc.29, 1525–1534 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Au-Yeung, W. T. M., Sahani, A. K., Isselbacher, E. M. & Armoundas, A. A. Reduction of false alarms in the intensive care unit using an optimized machine learning based approach. NPJ Digital Med.2, 86 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Segal, G. et al. Utilizing risk-controlling prediction calibration to reduce false alarm rates in epileptic seizure prediction. Front. Neurosci. 17, 1184990 (2023). [DOI] [PMC free article] [PubMed]
- 19.Xu, D. et al. Exploring ICU nurses’ response to alarm management and strategies for alleviating alarm fatigue: a meta-synthesis and systematic review. BMC Nurs.24, 412 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Takefuji, Y. Limitations of XGBoost-SHAP integration for interpretable machine learning in antimicrobial resistance prediction. J. Infect.91, 106575 (2025). [DOI] [PubMed] [Google Scholar]
- 21.Ponce-Bobadilla, A. V., Schmitt, V., Maier, C. S., Mensing, S. & Stodtmann, S. Practical guide to SHAP analysis: explaining supervised machine learning model predictions in drug development. Clin. Transl. Sci.17, e70056 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Porto, B. M. Improving triage performance in emergency departments using machine learning and natural language processing: a systematic review. BMC Emerg. Med.24, 219 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhou, S. et al. Explainable differential diagnosis with dual-inference large language models. NPJ Health Syst.2, 12 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Williams, C. Y. K., Miao, B. Y., Kornblith, A. E. & Butte, A. J. Evaluating the use of large language models to provide clinical recommendations in the Emergency Department. Nat. Commun.15, 8236 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Kuo, K. M. & Chang, C. S. A meta-analysis of the diagnostic test accuracy of artificial intelligence predicting emergency department dispositions. BMC Med. Inform. Decis. Mak.25, 187 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Tschoellitsch, T. et al. Using emergency department triage for machine learning-based admission and mortality prediction. Eur. J. Emergency Med. 30, 408–416 (2023). [DOI] [PubMed]
- 27.Klug, M. et al. A gradient boosting machine learning model for predicting early mortality in the emergency department triage: devising a nine-point triage score. J. Gen. Intern. Med.35, 220–227 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Ye, C. et al. A real-time early warning system for monitoring inpatient mortality risk: prospective study using electronic medical record data. J. Med. Internet Res.21, e13719 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Brajer, N. et al. Prospective and external evaluation of a machine learning model to predict in-hospital mortality of adults at time of admission. JAMA Netw. Open3, e1920733 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Naemi, A. et al. Machine learning techniques for mortality prediction in emergency departments: a systematic review. BMJ Open11, e052663 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Lundberg, S. M. et al. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nat. Biomed. Eng.2, 749–760 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Steyvers, M. et al. What large language models know and what people think they know. Nat. Mach. Intell.7, 221–231 (2025). [Google Scholar]
- 33.Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data10, 1 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Gosselin, A. et al. Detection of serious adverse drug reactions using diagnostic codes in the International Statistical Classification of Diseases and Related Health Problems. J. Popul. Ther. Clin. Pharmacol.27, e35–e48 (2020). [DOI] [PubMed] [Google Scholar]
- 35.Theodosakis, N. et al. Validation of case identification for melasma using international statistical classification of diseases and related health problems, tenth revision codes. JAMA Dermatol.158, 1453–1454 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Henny, J. et al. Recommendation for the review of biological reference intervals in medical laboratories. Clin. Chem. Lab. Med.54, 1893–1900 (2016). [DOI] [PubMed] [Google Scholar]
- 37.Prasanna Venkatesh, N., Pradeep Kumar, R., Chakravarthy Neelapu, B., Pal, K. & Sivaraman, J. CatBoost-based improved detection of P-wave changes in sinus rhythm and tachycardia conditions: a lead selection study. Phys. Eng. Sci. Med.46, 925–944 (2023). [DOI] [PubMed] [Google Scholar]
- 38.Wang, Y., Liu, J. xing, Wang, J., Shang, J. & Gao, Y. lian A graph representation approach based on light gradient boosting machine for predicting drug–disease associations. J. Comput. Biol.30, 937–947 (2023). [DOI] [PubMed] [Google Scholar]
- 39.Mahapatra, S., Gupta, V. R., Sahu, S. S. & Panda, G. Deep neural network and extreme gradient boosting based hybrid classifier for improved prediction of protein-protein interaction. IEEE/ACM Trans. Comput. Biol. Bioinform.19, 155–165 (2022). [DOI] [PubMed] [Google Scholar]
- 40.Botha, D. & Steyn, M. The use of decision tree analysis for improving age estimation standards from the acetabulum. Forensic Sci. Int.341, 111514 (2022). [DOI] [PubMed] [Google Scholar]
- 41.Wallace, M. L. et al. Use and misuse of random forest variable importance metrics in medicine: demonstrations through incident stroke prediction. BMC Med. Res. Methodol.23, 144 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Wang, J. & Geng, X. Large margin weighted k-nearest neighbors label distribution learning for classification. IEEE Trans. Neural. Netw. Learn. Syst.35, 16720–16732 (2024). [DOI] [PubMed] [Google Scholar]
- 43.Song, X., Liu, X., Liu, F. & Wang, C. Comparison of machine learning and logistic regression models in predicting acute kidney injury: a systematic review and meta-analysis. Int. J. Med. Inform.151, 104484 (2021). [DOI] [PubMed] [Google Scholar]
- 44.Wong, T. T. & Yang, N. Y. Dependency analysis of accuracy estimates in k-fold cross validation. IEEE Trans. Knowl. Data Eng.29, 2417–2427 (2017). [Google Scholar]
- 45.Nohara, Y., Matsumoto, K., Soejima, H. & Nakashima, N. Explanation of machine learning models using shapley additive explanation and application for real data in hospital. Comput. Methods Prog. Biomed.214, 106584 (2022). [DOI] [PubMed] [Google Scholar]
- 46.Ma, Y. et al. Advancing preeclampsia prediction: a tailored machine learning pipeline integratingresampling and ensemble models for handling imbalanced medical data. BioData Min.18, 25 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Vickers, A. J. & Holland, F. Decision curve analysis to evaluate the clinical benefit of prediction models. Spine J.21, 1643–1648 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Van Calster, B. et al. Reporting and interpreting decision curve analysis: a guide for investigators. Eur. Urol.74, 796–804 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Sellergren, A. et al. MedGemma Technical Report. Google Research and Google DeepMind. https://arxiv.org/abs/2507.05201 (2025).
- 50.Sandmann, S. et al. Benchmark evaluation of DeepSeek large language models in clinical decision-making. Nat. Med.31, 2546–2549 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Ray, P. P. Toward transparent AI-enabled patient selection in cosmetic surgery by integrating reasoning and medical LLMs. Aesthetic Plastic Surg.49, 5641–5642 (2025). [DOI] [PubMed] [Google Scholar]
- 52.Lee, Y., Chang, C. H. & Yang, C. C. Enhancing patient-physician communication: simulating african american vernacular english in medical diagnostics with large language models. J. Healthc. Inform. Res.9, 119–153 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Huang, Y. et al. Evaluation of large language models for providing educational information in orthokeratology care. Contact Lens Anterior Eye. 48, 102384 (2025). [DOI] [PubMed]
- 54.Hou, W. & Ji, Z. Assessing GPT-4 for cell type annotation in single-cell RNA-seq analysis. Nat. Methods21, 1462–1465 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Aslam, M. Chi-square test under indeterminacy: an application using pulse count data. BMC Med. Res. Methodol.21, 201 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Francis, G. & Jakicic, V. Equivalent statistics for a one-sample t-test. Behav. Res. Methods55, 77–84 (2023). [DOI] [PubMed] [Google Scholar]
- 57.Clark, J. S. C. et al. Empirical investigations into Kruskal-Wallis power studies utilizing Bernstein fits, simulations and medical study datasets. Sci. Rep.13, 2352 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
npj_IHM_Supplementary Materials_20260204
Data Availability Statement
Data sources include the public MIMIC dataset and restricted Chinese clinical data (not publicly available due to privacy regulations). A lightweight, standalone version of the risk prediction calculator (implementing the core machine-learning tier without LLM filtering) is available for interactive use at: https://emergency-death-risk-prediction.streamlit.app/. The core code has been openly shared on GitHub (https://github.com/AI-Medicine-team/AI-TEW-Framework).
Data sources include the public MIMIC dataset and restricted Chinese clinical data (not publicly available due to privacy regulations). A lightweight, standalone version of the risk prediction calculator (implementing the core machine-learning tier without LLM filtering) is available for interactive use at: https://emergency-death-risk-prediction.streamlit.app/. The core code has been openly shared on GitHub (https://github.com/AI-Medicine-team/AI-TEW-Framework).








