Skip to main content
Respiratory Research logoLink to Respiratory Research
. 2026 Feb 23;27:126. doi: 10.1186/s12931-026-03581-x

A clinically interpretable prediction model for acute mortality in patients with pneumonia requiring mechanical ventilation

Haiming Hu 1,2,#, Yao Zu 1,#, Lijuan Zhao 2,3, Qing Zhang 1, Geer Zhou 1,2, Pan Xu 1, Anqi Zhao 1,2, Fei Yin 1,2, Lokesh Sharma 4,✉, De Chang 1,✉
PMCID: PMC12980919  PMID: 41731516

Abstract

Background

Patients with pneumonia admitted to the intensive care unit (ICU) requiring early mechanical ventilation are at high risk for short-term mortality. Early risk assessment is crucial for timely intervention, optimal resource allocation, and improved outcomes. This study aimed to develop and validate a clinically interpretable prediction model for predicting 7-day mortality in this high-risk population.

Methods

Data from the MIMIC-IV database were used for model development and internal validation, while the SCRIPT and PLAGH cohorts served as external validation cohorts to assess generalizability across geographical regions and ethnicities. Feature selection was conducted using Lasso regression. We compared nine machine learning algorithms and selected the optimal model based on performance metrics including the area under the curve (AUC), calibration, decision curve analysis (DCA) and precision-recall (PR) curves. Model interpretability was enhanced through a nomogram for individualized risk visualization, supplemented by SHAP analysis to rank feature importance. A freely accessible web-based calculator was developed to facilitate individualized risk assessment in clinical practice.

Results

This study included 6,720 patients from MIMIC-IV, 492 from SCRIPT, and 136 from PLAGH. Eight predictors were selected: age, SOFA score, heart rate, respiratory rate, urine output, hemoglobin, lactate, and tracheostomy. The logistic regression model demonstrated the best performance, with an AUC of 0.79 (95%CI: 0.73–0.85) in the SCRIPT cohort and 0.90 (95% CI: 0.83–0.96) in the PLAGH cohort, outperforming the SOFA score (DeLong test, P < 0.01). The model exhibited good calibration, and DCA confirmed its clinical net benefit. A user-friendly web-based calculator was developed to facilitate clinician use.

Conclusions

In this study, we developed and multicentrically validated an interpretable prediction model to predict short-term mortality in patients with pneumonia requiring early mechanical ventilation. The model demonstrated robust and consistent performance across multiple independent cohorts. This tool may assist clinicians in early risk stratification, facilitating timely intervention for high-risk patients while avoiding unnecessary treatments in those at low risk. Implementation as a freely available web-based calculator may further enhance care efficiency by aligning interventions with patient risk.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12931-026-03581-x.

Keywords: Pneumonia, Mechanical ventilation, Machine learning, Predictive model, Short-term mortality

Introduction

Pneumonia, a common respiratory infection, remains a leading cause of global mortality [1]. It is particularly prone to progressing to severe pneumonia in older adults and individuals with comorbidities, often requiring intensive care unit (ICU) admission. The management of severe pneumonia in the ICU is increasingly challenged by a growing aging population and the emergence of drug-resistant pathogens [2]. Despite advances in supportive care, incidence and mortality rates remain high, particularly among patients who require mechanical ventilation.

Patients with pneumonia requiring early mechanical ventilation in the ICU are at high risk for organ failure and death, which imposes a significant healthcare burden [3]. The initial days after ICU admission represent a critical window of highest risk for this population, often driven by refractory hypoxemia and rapid clinical deterioration. Although conventional prognostic scoring systems such as APACHE II, SOFA, and CURB-65 scores are widely used in clinical practice, their accuracy in this specific population remains limited [4]. Furthermore, they fail to provide a clinician-friendly visualization of mortality risk. These scores remain abstract numbers that are difficult to translate into tangible, individualized prognoses for clinical decision-making.

Machine learning (ML) has demonstrated growing potential for prognostic prediction in critical care [5, 6]. Its applications range from unsupervised clustering, which has identified robust prognostic phenotypes in hospital-acquired pneumonia [7], to various supervised models for mortality prediction. Several supervised learning models for pneumonia-related mortality have also been developed and externally validated. Examples include an XGBoost model (9 variables) derived from public databases, which showed stable external performance [8]; a logistic regression model for ICU all-cause mortality (8 variables) achieving an AUC of 0.76, sensitivity of 0.76, and specificity of 0.64 [9]; and another XGBoost model (7 variables) for severe pneumonia mortality that outperformed traditional scores like SOFA with an AUC of 0.85 [10]. Notably, another logistic regression model developed in a cohort of 875 patients with severe pneumonia achieved an AUC of 0.88, surpassing APACHE II and SOFA scores [11]. Collectively, these studies highlight the promise of ML for risk stratification.

Despite these advances, critical gaps remain. First, the opaque nature of many complex models obscures their decision-making process, which limits their clinical application. More importantly, a specific predictive tool for short-term mortality is lacking for pneumonia patients requiring early mechanical ventilation, a high-risk subgroup in whom rapid disease progression makes timely and interpretable risk stratification essential to guide interventions and optimize resource allocation.

To address these gaps, we aimed to develop and externally validate a clinically interpretable prediction model for short-term mortality in this high-risk population. By leveraging a multicenter study design, we sought to create a practical tool to aid clinicians in early, interpretable risk stratification, with the goal of guiding timely interventions and optimizing resource allocation in critical care settings.

Methodology

Study population

This retrospective study utilized data from three independent cohorts: the Medical Information Mart for Intensive Care IV (MIMIC-IV, version 3.0), the Successful Clinical Response in Pneumonia Therapy (SCRIPT, version 1.1.0) dataset [12], and a cohort from the People's Liberation Army General Hospital (PLAGH). The MIMIC-IV database contains comprehensive health records from 546,028 critical care admissions at Beth Israel Deaconess Medical Center [13], while the SCRIPT dataset comprises 585 ICU patients with suspected severe pneumonia who received mechanical ventilation [12]. Access to the MIMIC-IV and SCRIPT databases was authorized under Record ID 13485805, which waived the requirement for additional ethical approval.

To independently validate the model's generalizability, we identified ICU patients with pneumonia who required mechanical ventilation within the first 7 days of admission from the PLAGH cohort, spanning March 2020 to August 2025. This part of the study was approved by the Ethics Committee of the Seventh Medical Center of PLA General Hospital (No. S2025-113–01). All patient records across cohorts were de-identified, and the need for informed consent was waived.

Population selection and definitions

This study included adult patients diagnosed with pneumonia in ICU. The inclusion criteria were: (1) admission to the ICU; (2) diagnosis of pneumonia (Table S1); (3) initiation of invasive mechanical ventilation within the first 7 days following ICU admission. The exclusion criteria were: (1) age < 18 years; (2) implausible values for clinical data (e.g., negative RR); (3) discharge from the ICU or death within 24 h of admission. For patients with multiple ICU admissions during a single hospitalization, only the first ICU stay was considered for analysis. The study flowchart and model development process are presented in Fig. 1.

Fig. 1.

Fig. 1

Study flowchart and model development process. A workflow of model development, external validation, and clinical application. B detailed flowchart of cohort selection, model training, and performance evaluation

The diagnosis of pneumonia was established based on International Classification of Diseases-9/10 (ICD-9/10) codes, supplemented by clinical assessment [14]. Patients in the MIMIC-IV and PLAGH cohorts were identified using these codes (Table S1). In the SCRIPT cohort, the pneumonia diagnosis for all patients was additionally confirmed by at least one bronchoalveolar lavage (BAL) procedure, ensuring a high diagnostic standard. Early mechanical ventilation was defined as the initiation of invasive mechanical ventilation within 7 days of ICU admission, with a minimum duration exceeding one day. The urine output was defined as the earliest recorded 24-h urine volume following ICU admission. The primary outcome was short-term mortality (7-day all-cause mortality) after ICU admission. Patients who were discharged alive from the hospital before day 7 were considered survivors. Mortality events included both deaths during the ICU stay and discharge to hospice care.

Data collection and preprocessing

Candidate predictors for this study were selected based on their availability in the MIMIC-IV and SCRIPT databases. The selected variables represented the earliest recorded clinical data following the initiation of mechanical ventilation. After initial screening, a total of 33 features were retained for analysis, categorized as follows: (1) Demographic information: age, sex, and race; (2) Medical history: smoking status, ICU history, admission source (non-medical unit vs. medical unit), and COVID-19 infection; (3) Vital signs: temperature, heart rate (HR), systolic blood pressure (SBP), diastolic blood pressure (DBP), respiratory rate (RR), urine output, pH, PaCO2, and PaO2; (4) Laboratory data: white blood cell (WBC), lymphocyte, neutrophil, hemoglobin (Hb), blood platelet count (PLT), bicarbonate, creatinine, albumin, bilirubin, and lactate; (5) Clinical scores: the Glasgow Coma Scale score (GCS score) and sequential organ failure assessment score (SOFA score); (6) Interventions: tracheostomy, vasoactive agents, and continuous renal replacement therapy (CRRT). Additionally, the length of ICU stay and 7-day mortality were also collected.

Variables with missing data exceeding 20% were excluded from the analysis. For the remaining variables, multiple imputation was performed using a random forest algorithm. This non-parametric method effectively captures complex, nonlinear relationships and interactions among variables of mixed types without assuming a specific data distribution [15]. The imputation process was conducted separately for each cohort to preserve their unique data structures and prevent information leakage. Subsequently, Rubin's rules were applied to pool the imputed datasets for downstream modeling, ensuring unbiased parameter estimates and valid standard errors.

Model development and tuning

For model development, the MIMIC-IV dataset was randomly split into a training set (70%) and an internal validation set (30%). The SCRIPT cohort was used for external validation, and the PLAGH cohort was used to further validate the model in a distinct regional and racial population.

Feature selection was performed using Lasso regression, which simultaneously shrinks coefficients and performs automated variable selection [16]. To identify the optimal predictive model, we compared the performance of nine mainstream machine learning algorithms: Decision Tree, ElasticNet (Enet), K-Nearest Neighbors (KNN), light gradient boosting machine (LightGBM), Logistic Regression, Multi-layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM) and XGBoost [17]. Model hyperparameters were optimized and overfitting was controlled via fivefold cross-validation on the training set [18]. The optimal threshold was determined using Youden’s J statistic. Thereafter, the top four models ranked by their performance on the internal validation set were further evaluated on the SCRIPT cohort. Subsequently, the best-performing model was finally assessed on the PLAGH cohort to test its generalizability and predictive power.

Model evaluation and interpretation

Model performance was comprehensively evaluated using a range of metrics, including the area under the ROC curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). The 95% confidence intervals (CIs) for AUC estimates were calculated using methods appropriate to each dataset: bootstrap resampling (R = 1,000) was applied to the MIMIC-IV cohort (after its split into training and internal validation sets) to assess stability [19], while DeLong’s method was used for the external validation cohorts (SCRIPT and PLAGH) [20]. Calibration curves were plotted to visualize the agreement between predicted probabilities and observed outcomes, and decision curve analysis (DCA) was applied to assess the clinical net benefit across different probability thresholds [21, 22]. To account for class imbalance in the outcome, the area under the precision-recall curve (AUPRC) was also calculated as a complementary performance measure, which reflects the average positive predictive value across all sensitivity thresholds [23]. To evaluate the incremental predictive value of our model, we compared its performance with that of the SOFA score across both the training and all validation cohorts. The statistical significance of differences in AUC was assessed using DeLong’s test [24].

To enhance model interpretability and clinical utility, we employed visualization tools tailored for clinical practice. A nomogram was constructed to translate model parameters into an individualized short-term mortality risk prediction, visually illustrating the contribution of each predictor. Furthermore, to provide a data-driven ranking of feature importance, Shapley Additive Explanation (SHAP) analysis was performed on the final model. Mean absolute SHAP values were used to rank predictors by their overall contribution to model predictions [25]. This analysis offers an alternative perspective on variable importance, complementing the nomogram and standard regression outputs.

Finally, a freely accessible, investigational web-based calculator was developed to facilitate risk assessment and model dissemination. This tool allows clinicians to input patient-specific variables to obtain an individualized estimate of short-term mortality risk and is intended for exploratory use [26].

Statistical methods

Statistical analyses were performed using R (version 4.4.1). Continuous variables are presented as mean ± standard deviation (SD) or median [interquartile range (IQR)], based on their distribution. Categorical variables are expressed as numbers (percentages). Group comparisons employed comparisons employed Student’s t-test or the Mann–Whitney U test for continuous variables and the chi-square test for categorical variables. A two-sided p-value of < 0.05 was considered statistically significant.

Results

Basic characteristics

Following screening, a total of 6,720 eligible participants from the MIMIC-IV database were included in the study. They were randomly allocated into a training set (N = 4,704, 70%) and an internal validation set (N = 2,016, 30%). Additionally, 492 patients from the SCRIPT cohort and 136 patients from the PLAGH cohort were included as independent external validation cohorts.

Missing data patterns varied across cohorts. In both MIMIC-IV and SCRIPT cohorts, missing values were primarily observed in laboratory variables. Oxygen saturation was excluded from the analysis due to a high rate of missing data (34.55% in MIMIC-IV). In contrast, the PLAGH cohort (enrolled within the past five years) had no missing values, reflecting more complete data collection. The detailed proportions of missing values for each variable are provided in Table S2.

Notably, three cohorts showed distinct distributions in terms of demographic and clinical characteristics. First, patients in the PLAGH cohort exhibited greater overall illness severity, characterized by older age, increased utilization of vasoactive agents, higher mortality rates, elevated SOFA score, and longer length of ICU stay compared to the other two cohorts. Second, due to differences in the periods of data collection, the prevalence of COVID-19 infection in MIMIC- IV cohort (2.5%) was significantly lower than SCRIPT cohort (38.6%) and PLAGH cohort (21.3%). Finally, racial composition also differed across the cohorts: the PLAGH cohort was exclusively of Asian descent, whereas the majority of patients in the MIMIC-IV and SCRIPT cohorts were White. The baseline characteristics of the MIMIC- IV, SCRIPT, and PLAGH cohort are presented in Table 1.

Table 1.

Baseline characteristics of patients in four cohorts: training, internal validation, SCRIPT, and PLAGH

Patients Training cohort
(N = 4704)
Internal validation cohort
(N = 2016)
SCRIPT cohort
(N = 492)
PLAGH cohort
(N = 136)
Age (years) 66.70 [55.58, 76.53] 66.15 [55.27, 76.40] 63.00 [52.00, 72.00] 79.50 [70.75, 86.00]
Gender
 Male [n (%)] 2782 (59.1) 1225 (60.8) 299 (60.8) 95 (69.9)
 Female [n (%)] 1922 (40.9) 791 (39.2) 193 (39.2) 41 (30.1)
Race
 White [n (%)] 2763 (58.7) 1196 (59.3) 285 (57.9) 0 (0)
 Others [n (%)] 1941 (40.3) 820 (40.7) 207 (42.1) 136 (100)
Smoking
 Yes [n (%)] 740 (15.7) 300 (14.9) 302 (61.4) 36 (26.5)
 No [n (%)] 3964 (84.3) 1716 (85.1) 190 (38.6) 100 (73.5)
ICU history
 Yes [n (%)] 273 (5.8) 113 (5.6) 105 (21.3) 58 (42.6)
 No [n (%)] 4431 (94.2) 1903 (94.4) 387 (78.7) 78 (57.4)
Admission source
 NMU [n (%)] 232 (4.9) 96 (4.8) 280 (56.9) 4 (2.9)
 MU [n (%)] 4472 (95.1) 1920 (95.2) 212 (43.1) 132 (97.1)
COVID-19 infected
 Yes [n (%)] 116 (2.5) 49 (2.4) 190 (38.6) 29 (21.3)
 No [n (%)] 4588 (97.5) 1967 (97.6) 302 (61.4) 107 (78.7)
Tracheostomy
 Yes [n (%)] 251 (5.3) 106 (5.3) 136 (27.6) 3 (2.2)
 No [n (%)] 4453 (94.7) 1910 (94.7) 356 (72.4) 133 (97.8)
Vasoactive agent
 Yes [n (%)] 2865 (60.9) 1261 (62.5) 225 (45.7) 111 (81.6)
 No [n (%)] 1839 (39.1) 755 (37.5) 267 (54.3) 25 (18.4)
CRRT
 Yes [n (%)] 619 (13.2) 271 (13.4) 16 (3.3) 34 (25.0)
 No [n (%)] 4085 (86.8) 1745 (86.6) 476 (96.7) 102 (75.0)
Clinically scores
 SOFA score 8.00 [5.00, 11.00] 8.00 [5.00, 11.00] 9.00 [5.00, 12.00] 9.00 [7.00, 13.00]
 GCS scores 9.00 [5.00, 13.00] 9.00 [4.00, 13.00] 8.00 [3.00, 14.25] 10.00 [7.00, 12.25]
Vital signs
 Temperature (℃) 37.01 [36.70, 37.40] 37.00 [36.68, 37.37] 36.95 [36.56, 37.53] 37.40 [36.80, 38.20]
 Heart rate (/min) 86.29 [75.20, 98.95] 86.27 [75.32, 98.11] 91.87 [78.75, 105.32] 90.00 [80.00, 100.50]
 SBP (mmHg) 113.52 [105.16, 124.92] 113.12 [104.85, 124.01] 116.59 [106.08, 129.10] 128.50 [111.75, 140.00]
 DBP (mmHg) 61.61 [55.53, 68.23] 61.64 [55.30, 68.60] 62.82 [56.81, 70.34] 65.50 [58.00, 72.00]
 Respiratory rate (/min) 20.06 [17.62, 23.15] 20.33 [17.71, 23.32] 23.17 [20.00, 27.34] 20.00 [16.00, 21.25]
 Urine output (mL/day) 1374.00 [826.50, 2155.00] 1345.00 [775.00, 2142.50] 622.00 [274.00, 1130.00] 1860.00 [1130.00, 2500.00]
 PH 7.36 [7.29, 7.42] 7.36 [7.28, 7.42] 7.37 [7.31, 7.43] 7.43 [7.38, 7.48]
 PaCO2 (mmHg) 43.00 [37.00, 52.00] 43.00 [36.00, 52.00] 39.00 [33.00, 45.67] 37.15 [31.20, 43.00]
 PaO2 (mmHg) 91.00 [57.00, 162.00] 89.00 [56.00, 166.00] 98.00 [78.41, 120.54] 88.80 [74.60, 117.25]
Laboratory data
 WBC (× 10^9/L) 11.30 [7.90, 15.90] 11.30 [8.00, 15.90] 10.75 [7.27, 15.60] 12.04 [8.27, 14.83]
 Lymphocyte (× 10^9/L) 0.99 [0.57, 1.55] 0.99 [0.59, 1.59] 0.80 [0.50, 1.30] 0.55 [0.35, 0.95]
 Neutrophil (× 10^9/L) 8.91 [5.95, 13.04] 9.00 [6.04, 12.87] 8.90 [5.29, 13.00] 10.77 [7.12, 13.58]
 Hemoglobin (g/dL) 10.80 [9.10, 12.60] 10.80 [9.20, 12.60] 10.93 [8.80, 12.90] 9.90 [8.28,11.85]
 Platelet (× 10^9/L) 200.00 [145.00, 270.00] 197.00 [142.75, 271.00] 197.00 [125.00, 275.50] 139.50 [94.25, 207.25]
 Bicarbonate (mmol/L) 23.00 [20.00, 26.00] 23.00 [20.00, 26.00] 23.00 [20.00, 26.31] 24.20 [19.80, 28.20]
 Creatinine (mg/dL) 1.00 [0.80, 1.70] 1.10 [0.80, 1.60] 1.14 [0.79, 1.90] 1.16 [0.76, 2.36]
 Albumin (g/dL) 3.00 [2.60, 3.40] 3.00 [2.50, 3.40] 3.10 [2.70, 3.54] 2.90 [2.70, 3.18]
 Bilirubin (mg/dL) 0.50 [0.30, 1.00] 0.60 [0.40, 1.00] 0.70 [0.50, 1.10] 0.68 [0.47, 1.13]
 Lactate (mmol/L) 1.60 [1.10, 2.40] 1.60 [1.10, 2.50] 1.70 [1.20, 2.38] 1.80 [1.37, 2.73]
Clinical outcomes
 ICU stay length (days) 9.72 [5.11, 16.85] 9.37 [4.92, 16.79] 15.00 [7.00, 28.00] 18.00 [8.75, 36.25]
Short mortality
 Alive [n (%)] 4264 (90.6) 1826 (90.6) 453 (92.1) 104 (76.5)
 Died [n (%)] 440 (9.4) 190 (9.4) 39 (7.9) 32 (23.5)

CRRT Continuous renal replacement therapy, GCS Glasgow Coma Scale score, SOFA sequential organ failure assessment score, SBP Systolic blood pressure, DBP Diastolic blood pressure, WBC White blood cell count

Feature selection

To balance model performance with parsimony, Lasso regression with fivefold cross-validation was applied to the training set to select features. Figure 2A illustrates the shrinkage of coefficient estimates as the log-transformed regularization parameter (log(λ)) increases. Figure 2B presents the corresponding cross-validated deviance, which was used to identify the optimal λ value. To prioritize a simpler and more clinically interpretable model, λ.1se was chosen as the final regularization parameter. This procedure ultimately retained eight features with non-zero coefficients: age, SOFA score, heart rate, respiratory rate, urine output, hemoglobin, lactate, and tracheostomy status.

Fig. 2.

Fig. 2

Lasso regression for feature selection: coefficient shrinkage (A) and cross-validated deviance (B)

Model development

Among the nine machine learning methods constructed, the logistic regression model exhibited the strongest discrimination capability, achieving an AUC of 0.77 (95% CI: 0.73–0.80) on the internal validation set. The remaining models, ranked in descending order of AUC, were: Enet (AUC = 0.77, 95%CI: 0.73–0.80), XGBoost (AUC = 0.75, 95%CI: 0.71–0.79), LightGBM (AUC = 0.75, 95%CI: 0.71–0.78), MLP (AUC = 0.74, 95%CI: 0.71–0.78), SVM (AUC = 0.73, 95%CI: 0.69–0.77), RF (AUC = 0.70, 95%CI: 0.66–0.74), DT (AUC = 0.63, 95%CI: 0.59–0.66) and KNN (AUC = 0.62, 95%CI: 0.58–0.67). A comprehensive summary of performance metrics for all models is provided in Fig. 3 and Table S3.

Fig. 3.

Fig. 3

Column chart comparison of multi-metric classification performance across nine machine learning models: training and validation sets

Overall, the logistic regression model demonstrated balanced and stable performance across both the training set and the internal validation set. On the internal validation set, it achieved an AUC of 0.77 (95%CI: 0.73–0.80), an accuracy of 0.66 (95%CI: 0.64–0.68), a sensitivity of 0.71 (95%CI: 0.64–0.77), and a specificity of 0.65 (95%CI: 0.63–0.67). Clinically, its high NPV (0.96, 95%CI: 0.94–0.97) is valuable, as it helps clinicians to exclude low-risk patients, avoid unnecessary interventions, and thereby direct critical medical resources toward those at higher risk. Although some ensemble algorithms such as RF and KNN showed nearly perfect performance on the training set, their performance declined substantially on the internal validation set, suggesting potential risks of overfitting.

The ROC, calibration, and decision curve analysis (DCA) based the internal validation cohorts are presented in Fig. 4A-C, respectively. Additional ROC evaluations focusing on the logistic regression model are provided in Figure S1A and S1B: Figure S1A compares performance between the training and internal validation sets, while Figure S1B displays the ROC curves obtained from fivefold cross-validation on the training set.

Fig. 4.

Fig. 4

The ROC (A), calibration (B), and decision curve analysis (C) curves in internal validation

External validation

To further evaluate model robustness and generalizability, the top four internally validated models (logistic regression, Enet, XGboost and LightGBM) were tested on the independent SCRIPT cohort. Notably, the logistic model maintained strong performance in the SCRIPT cohort (AUC = 0.79, 95%CI: 0.73–0.85). The ROC curves for all four models in the SCRIPT cohort are provided in Figure S2. Given its optimal balance of simplicity, competitive performance, and clinical interpretability, we ultimately selected the logistic regression model as the final model.

We then compared the discrimination performance of our model with that of the SOFA score across the training and validation cohorts. The logistic model consistently outperformed SOFA (Fig. 5), with DeLong's test confirming statistically significant differences in discrimination (P < 0.01).

Fig. 5.

Fig. 5

Comparison of predictive performance between the logistic regression model and SOFA score across different cohorts

Subsequently, to assess the model’s generalizability across distinct populations, we evaluated it on the geographically and racially different PLAGH cohort. In this cohort, the logistic regression model attained an AUC of 0.90 (95% CI: 0.83–0.96). The corresponding ROC, calibration, decision and precision-recall curves are presented in Fig. 6A-D, respectively. The calibration curve demonstrated good concordance between predicted probabilities and observed outcomes, which was supported by quantitative metrics: the Hosmer–Lemeshow test indicated adequate fit (χ2 = 9.66, P = 0.209), the Brier score was 0.119, and the calibration parameters showed minimal deviation from ideal (calibration slope = 0.979; calibration intercept = −0.000044). Decision curve analysis revealed that the model provided greater net benefits than both treat-all and treat-none strategies across a wide range of threshold probabilities.

Fig. 6.

Fig. 6

The ROC (A), calibration (B), decision curve analysis (C) and precision-recall (D) curves in PLAGH cohort

Interpretability analysis

To enhance clinical applicability and interpretability of the final model, a nomogram was constructed as the primary visualization tool (Fig. 7A). Each variable was assigned a specific point based on its regression coefficient, and the total points were summed to determine the probability of short-term mortality.

Fig. 7.

Fig. 7

Nomogram (A) and forest plot (B) for the logistic regression model. Note: Urine output: 24-h urine output (ml)

To provide a statistical summary of the predictor-outcome associations, a forest plot of the adjusted odds ratios (ORs) from the multivariable logistic regression was generated (Fig. 7B). This plot quantifies the direction and magnitude of each predictor’s association with short-term mortality. Specific details are reported in Table S4.

As a supplementary, data-driven assessment of feature importance, we performed a SHAP analysis. The results confirmed age, respiratory rate, and SOFA score as the top contributors to model predictions. The SHAP summary swarm plot and feature importance ranking based on mean SHAP values are presented in Figure S3.

Development of an investigational web-based calculator

To translate model predictions into a clinically usable format, an interactive, web-based calculator was developed. This tool offers a dynamic interface for individualized risk estimation (available at: https://pneumoniaml.shinyapps.io/EVTPPM_model/). To ensure full transparency and reproducibility, the regression coefficients, intercept, and key parameters of the final logistic regression model are listed in Table S4.

Discussion

Hospital mortality for severe pneumonia ranges from 12.3% to 48%, with approximately 14% of patients requiring mechanical ventilation [27–30]. In our study, short-term mortality among pneumonia patients receiving early mechanical ventilation ranged from 7.9% to 23.5%, which underscores the substantial risk in this population. These findings highlight the critical need for early identification of risk factors contributing to the mortality and appropriate intervention to improve clinical outcomes.

Machine learning has shown considerable potential for improving prognostic prediction in critical care. Various prediction models for pneumonia-related mortality, ranging from interpretable logistic regression to complex ensemble methods, have been developed and externally validated, often outperforming traditional scores [7–9]. Research on mortality prediction for mechanically ventilated pneumonia patients spans decades, identifying factors such as advanced age and organ failure as key determinants [31]. Despite these advances, high-quality studies with large sample sizes, rigorous external validation, and clear demonstration of clinical utility remain scarce [27]. Moreover, most existing prognostic models have focused on longer-term endpoints, such as 28-day or in-hospital mortality. While important, these outcomes are increasingly influenced by secondary complications (e.g., severe nosocomial infection, drug resistance, or prolonged organ dysfunction) that may not be directly attributable to the initial severity of pneumonia. In contrast, mortality within the first week is more closely linked to the acute, life-threatening sequelae of the disease itself, such as refractory hypoxemia or shock. A model that can identify patients at highest risk during this critical early window may therefore support more timely and tailored interventions, potentially altering the initial clinical trajectory.

To address these gaps, we developed and validated machine learning models using three independent databases to predict 7-day mortality in critically ill pneumonia patients treated with early mechanical ventilation. In the internal validation cohort, the logistic regression model demonstrated strong overall performance (AUC = 0.77, 95%CI: 0.73–0.80), with excellent calibration and a substantial net benefit in clinical decision-making scenarios. Importantly, it achieved a consistently high NPV of 0.96 in both training and validation sets, which may help clinicians confidently identify low-risk patients and prioritize resources for higher-risk individuals [32]. Although the PPV was relatively low, which is typical for low-incidence outcomes, the model ultimately provides a critical tool for aiding clinical decision-making [33]. Validation in the SCRIPT cohort yielded an AUC of 0.79 (95% CI: 0.73–0.85), confirming its robustness. Notably, its superior performance in the PLAGH cohort (AUC = 0.90, 95% CI: 0.83–0.96), indicates its particular appropriateness for elderly patients with severe pneumonia.

To enhance clinical interpretability, we employed visualization tools tailored to clinical practice. A nomogram was constructed based on the multivariate logistic regression, providing a visual representation of how each variable contributes to the estimated risk. The model identified several predictors of short-term mortality. For instance, older age was associated with increased mortality risk (OR = 1.03, 95%CI: 1.02–1.04, P < 0.001), consistent with evidence that reduced physiological reserve worsens outcomes in critical illness [31, 34]. Elevated lactate (OR = 1.08, 95%CI:1.04–1.12, P < 0.001) and higher SOFA score (OR = 1.07, 95%CI:1.04–1.10, P < 0.001) were also strong predictors, reflecting tissue hypoxia and multi-organ dysfunction, respectively [35, 36].

Tracheostomy was associated with a lower predicted risk of short-term mortality (OR = 0.19, 95%CI: 0.08–0.41, P < 0.001). However, this association requires cautious interpretation due to potential selection and time-dependent biases. Baseline comparisons between patients who did and did not receive tracheostomy in the MIMIC-IV cohort (Table S5) showed similar SOFA score (8.0 vs. 8.0, P = 0.083) but significant differences in other characteristics: the proportion of smokers was lower in the tracheostomy group (2.2% vs. 16.2%, P < 0.001), and their length of ICU stay was longer (22.9 days vs. 9.3 days, P < 0.01). These findings suggest that tracheostomy was typically performed in patients who had already survived the initial acute phase of illness. A sensitivity analysis removing tracheostomy from the model resulted in a statistically significant decline in predictive performance across all datasets (DeLong test P < 0.05, Figure S4), indicating that tracheostomy contributed non-negligibly to the model's discriminative ability. However, this finding does not imply that tracheostomy directly reduces mortality. Instead, it likely reflects survival bias (patients must live long enough to undergo the procedure) and clinical indication bias (clinicians selectively perform tracheostomy in patients with more favorable baseline status).

Several limitations of this study should be acknowledged. First, its retrospective design using public datasets carries inherent risks of information bias and unmeasured confounding factors. Although we performed comparative and sensitivity analyses, residual confounding, particularly from clinicians’ subjective assessments, cannot be fully ruled out, especially for time-dependent interventions like tracheostomy. Second, variability in clinical protocols across institutions, particularly in intubation criteria and outcome classification, may limit generalizability. The higher AUC observed in the PLAGH cohort may be attributable to its stricter patient selection and definitive outcome measures, where only patients discharged due to imminent death were classified as receiving hospice care. In contrast, in the MIMIC-IV and SCRIPT cohorts included cases with discontinued treatment who survived beyond 7 days. However, the relatively small sample size of the PLAGH cohort (N = 136) calls for cautious interpretation, and future prospective studies with standardized protocols are warranted. Third, potentially relevant predictors may have been omitted due to missing data or lack of recording in the MIMIC-IV and SCRIPT, which could affect model performance. Finally, treatment details, such as specific antibiotic, glucocorticoid, and ventilator parameter adjustments, were not incorporated and may influence the prognosis.

Despite these limitations, we have developed and externally validated a clinically interpretable tool to predict short-term mortality in critically ill pneumonia patients requiring early mechanical ventilation. This model may assist clinicians in early risk stratification, optimizing resource allocation, and facilitating timely decisions. To enhance accessibility, the model has been implemented as a freely available web-based calculator.

Supplementary Information

Supplementary Material 1. (47.7KB, docx)

Acknowledgements

We acknowledge the support of all investigators of the MIMIC-Ⅳ and SCRIPT group.

Abbreviations

ML

Machine learning

MIMIC-Ⅳ

Medical Information Mart for Intensive Care Ⅳ

SCRIPT

Successful Clinical Response in Pneumonia Therapy

PLAGH

People's Liberation Army General Hospital

ICU

Intensive care unit

NMU

Non-medical unit

MU

Medical unit

AUC

Area under the curve

AUPRC

Area under the precision-recall curve

Enet

ElasticNet

KNN

K-Nearest Neighbors

MLP

Multi-layer Perceptron

SVM

Support Vector Machine

SHAP

Shapley Additive Explanation

ICD

International Classification of Diseases

DCA

Decision curve analysis

SD

Standard deviation

IQR

Interquartile range

PPV

Positive predictive value

NPV

Negative predictive value

Authors’ contributions

Dr Chang and Mr Hu had full access to all of the data in the study and take responsibility for the integrity of the data. Concept and design: Chang, Sharma L, Hu. Acquisition, analysis, or interpretation of data: All authors. Drafting of the manuscript: Hu, Lijuan Zhao, AQ Zhao, Zhou, Qing Zhang and Pan Xu. Revision of the manuscript: Hu, Y.Z, Sharma L, Chang. Statistical analysis: Hu, Y.Z, Yin. Obtained funding: Chang. Administrative, technical, or material support: Chang. Supervision: Hu, Chang.

Funding

This study was supported by funding from National Science and Technology Major Special Projects (DC; No.2025ZD01903600), The National Key Research and Development Program of China (DC; No. 2021YFC2302300), Talent Project (DC; No. 02-SWKJYCJJ21, No.2021–439, 2022QN07351, 20230315).

Data availability

The PLAGH dataset can be made available upon reasonable request to the first author and corresponding authors on reasonable request. Datasets from MIMIC- IV and SCRIPT are available from below link: [] (https://physionet.org/content/mimiciv/3.1). Please note that access to the two databases requires completing relevant training and obtaining proper permission. Further inquiries can be directed to Haiming Hu.

Declarations

Ethics approval and consent to participate

The analysis of the MIMIC-IV and SCRIPT databases was conducted using de-identified, publicly available data previously authorized by relevant institutional review boards (IRB), which waived the need for additional ethical approval. The PLAGH cohort study was approved by the Ethics Committee of the Seventh Medical Center of PLA General Hospital (No. S2025-113–01). The study was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines. All data were handled in compliance with applicable privacy regulations and institutional policies to protect patient confidentiality.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Haiming Hu and Yao Zu contributed equally to this work.

Contributor Information

Lokesh Sharma, Email: sharmalk2@upmc.edu.

De Chang, Email: changde@301hospital.com.cn.

References

  • 1.Reyes LF, Conway Morris A, Serrano-Mayorga C, et al. Community-acquired pneumonia. Lancet. 2025;406(10517):2371–88. [DOI] [PubMed] [Google Scholar]
  • 2.Torres A, Chalmers JD, Dela Cruz CS, et al. Challenges in severe community-acquired pneumonia: a point-of-view review. Intensive Care Med. 2019;45(2):159–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Núñez SA, Roveda G, Zárate MS, et al. Ventilator-associated pneumonia in patients on prolonged mechanical ventilation: description, risk factors for mortality, and performance of the SOFA score. J Bras Pneumol. 2021;47(3):e20200569. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Sun Y, Li H, Pei Z, et al. Incidence of community-acquired pneumonia in urban China: A national population-based study. Vaccine. 2020;38(52):8362–70. [DOI] [PubMed] [Google Scholar]
  • 5.Greener JG, Kandathil SM, Moffat L, et al. A guide to machine learning for biologists. Nat Rev Mol Cell Biol. 2021;23(1):40–55. [DOI] [PubMed] [Google Scholar]
  • 6.Rubulotta F, Bahrami S, Marshall DC, et al. Machine Learning Tools for Acute Respiratory Distress Syndrome Detection and Prediction. Crit Care Med. 2024;52(11):1768–80. [DOI] [PubMed] [Google Scholar]
  • 7.Martin FP, Poulain C, Mulier JH, et al. Identification and validation of robust hospital-acquired pneumonia subphenotypes associated with all-cause mortality: a multi-cohort derivation and validation. Intensive Care Med. 2025;51(4):692–707. [DOI] [PubMed] [Google Scholar]
  • 8.Chen J, Hou DYS. Development and multi-database validation of interpretable machine learning models for predicting In-Hospital mortality in pneumonia patients: a comprehensive analysis across four healthcare systems. Respir Res. 2025;26(1):279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Niu J, Lv X, Gao L, et al. Development and validation of a machine learning-based prediction model for in-ICU mortality in severe pneumonia: a dual-center retrospective study. Int J Med Inform. 2025;204:106075. [DOI] [PubMed]
  • 10.Xie K, Huang X, Li Z, et al. Development and validation of a clinical prediction model for in-hospital mortality of severe pneumonia based on machine learning. Front Pharmacol. 2025;16:1660893 Published 2025 Nov 26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhao W, Li X, Gao L, et al. Machine learning-based model for predicting all-cause mortality in severe pneumonia. BMJ Open Respir Res. 2025;12(1):e001983. [DOI] [PMC free article] [PubMed]
  • 12.Markov, Nikolay, al. e. SCRIPT CarpeDiem Dataset: demographics, outcomes, and per-day clinical parameters for critically ill patients with suspected pneumonia (version 1.1.0). PhysioNet (2023), RRID:SCR_007345. 10.13026/5phr-4r89.
  • 13.Johnson AEW, Bulgarelli L, Shen L, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10(1):1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Metlay JP, Waterer GW, Long AC, et al. Diagnosis and treatment of adults with community-acquired pneumonia. An official clinical practice guideline of the American Thoracic Society and Infectious Diseases Society of America. Am J Respir Crit Care Med. 2019;200(7):e45–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Shah AD, Bartlett JW, Carpenter J, et al. Comparison of random forest and parametric imputation models for imputing missing data using MICE: a CALIBER study. Am J Epidemiol. 2014;179(6):764–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Bainter SA, et al. Comparing Bayesian variable selection to Lasso approaches for applications in psychology. Psychometrika. 2023;88(3):1032–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Binson VA, Thomas S, Subramoniam M, et al. A review of machine learning algorithms for biomedical applications. Ann Biomed Eng. 2024;52(5):1159–83. [DOI] [PubMed] [Google Scholar]
  • 18.Bates S, Hastie T, Tibshirani R. Cross-validation: what does it estimate and how well does it do it? J Am Stat Assoc. 2024;119(546):1434–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Xie W, Liu M, Okoli CTC, et al. Construction and evaluation of a predictive model for compassion fatigue among emergency department nurses: a cross-sectional study. Int J Nurs Stud. 2023;148:104613. [DOI] [PubMed]
  • 20.Koga D, Kaneda R, Komiya C, et al. Artificial intelligence identifies individuals with prediabetes using single-lead electrocardiograms. Cardiovasc Diabetol. 2025;24(1):415. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Van Calster B, Nieboer D, Vergouwe Y, et al. A calibration hierarchy for risk models was defined: from utopia to empirical data. J Clin Epidemiol. 2016;74:167–76. [DOI] [PubMed] [Google Scholar]
  • 22.Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Lawler PR, Goligher EC, Berger JS, et al. Therapeutic Anticoagulation with Heparin in Noncritically Ill Patients with Covid-19. N Engl J Med. 2021;385(9):790–802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Lu R, Lumish HS, Hasegawa K, et al. Prediction of new-onset atrial fibrillation in patients with hypertrophic cardiomyopathy using machine learning. Eur J Heart Fail. 2024;27(2):275–84. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lundberg SM, Erion G, Chen H, et al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell. 2020;2(1):56–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Jalali A, Alvarez-Iglesias A, Roshan D, et al. Visualising statistical models using dynamic nomograms. PLoS ONE. 2019;14(11):e0225253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Kohno S, Seki M, Takehara K, et al. Prediction of requirement for mechanical ventilation in community-acquired pneumonia with acute respiratory failure: a multicenter prospective study. Respiration. 2013;85(1):27–35. [DOI] [PubMed] [Google Scholar]
  • 28.Phua J, et al. Severe community-acquired pneumonia: timely management measures in the first 24 hours. Crit Care. 2016;20(1):237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Angus DC, Marrie TJ, Obrosky DS, et al. Severe community-acquired pneumonia: use of intensive care services and evaluation of American and British Thoracic Society diagnostic criteria. Am J Respir Crit Care Med. 2002;166(5):717–23. [DOI] [PubMed] [Google Scholar]
  • 30.Woodhead M, Welch CA, Harrison DA, et al. Community-acquired pneumonia on the intensive care unit: secondary analysis of 17,869 cases in the ICNARC Case Mix Programme Database. Crit Care. 2006;10 Suppl 2(Suppl 2):S1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Pascual FE, Matthay MA, Bacchetti P, et al. Assessment of prognosis in patients with community-acquired pneumonia who require mechanical ventilation. Chest. 2000;117(2):503–12. [DOI] [PubMed] [Google Scholar]
  • 32.Palmqvist S, Tideman P, Mattsson-Carlgren N, et al. Blood biomarkers to detect Alzheimer disease in primary care and secondary care. JAMA. 2024;332(15):1245–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hansen L, Bernstorff M, Enevoldsen K, et al. Predicting diagnostic progression to schizophrenia or bipolar disorder via machine learning. JAMA Psychiat. 2025;82(5):459–69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Rosenthal GE, Kaboli PJ, Barnett MJ, et al. Age and the risk of in-hospital death: insights from a multihospital study of intensive care patients. J Am Geriatr Soc. 2002;50(7):1205–12. [DOI] [PubMed] [Google Scholar]
  • 35.Yuzefpolskaya M, Schwartz S, Ladanyi A, et al. The role of lactate metabolism in heart failure and cardiogenic shock: clinical insights and therapeutic implications. J Card Fail. 2026;32(1):115–25. [DOI] [PubMed]
  • 36.Do SN, Dao CX, Nguyen TA, et al. Sequential organ failure assessment (SOFA) score for predicting mortality in patients with sepsis in Vietnamese intensive care units: a multicentre, cross-sectional study. BMJ Open. 2023;13(3):e064870. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (47.7KB, docx)

Data Availability Statement

The PLAGH dataset can be made available upon reasonable request to the first author and corresponding authors on reasonable request. Datasets from MIMIC- IV and SCRIPT are available from below link: [] (https://physionet.org/content/mimiciv/3.1). Please note that access to the two databases requires completing relevant training and obtaining proper permission. Further inquiries can be directed to Haiming Hu.


Articles from Respiratory Research are provided here courtesy of BMC

RESOURCES