Skip to main content
European Journal of Medical Research logoLink to European Journal of Medical Research
. 2025 Apr 24;30:328. doi: 10.1186/s40001-025-02588-2

Construction and validation of prognostic model for ICU mortality in cardiac arrest patients: an interpretable machine learning modeling approach

Yong Li 1, Ying Liu 1, Qing Zhang 1, Hongwei Zhu 1, Chengli Wen 2,, Xian Jiang 1,
PMCID: PMC12020013  PMID: 40275415

Abstract

Background

The incidence and mortality of cardiac arrest (CA) is high. We developed interpretable machine learning models for early prediction of ICU mortality risk in patients diagnosed with CA.

Methods

Data from the Medical Information Mart for Intensive Care (MIMIC-IV, version 2.2) was randomized to training set (0.7) and internal validation set (0.3), and data from eICU(version 2.0.1) was used as external validation set. Five models including Logistic Regression (LR), Random Forest (RF), K Nearest Neighbor (KNN), Decision Tree (DT), and Extreme Gradient Boost (XGBoost) were developed. The model with the largest area under the Receiver Operating Characteristic (ROC) curve (AUC) and good performance in other features was defined as the best model, and Shapley Additive Explanations (SHAP) was used to improve the interpretability of the optimal model.

Results

A total of 1088 patients from MIMIC-IV, and 3542 patients from eICU were included. Seven variables were selected to construct models by Least Absolute Shrinkage and Selection Operator (LASSO) regression. The RF model was the best predictive model with AUC and 95% CI at 0.83 (0.78–0.88) in internal validation set, and 0.71(0.68–0.74) in external validation set. SHAP analysis found that the variables that had a high impact on the risk of ICU death were minimal Glasgow Coma Scale (GCS), base excess, anion gap, and urine output.

Conclusion

RF is the optimal model for predicting the risk of ICU death in CA patients. The development of this model is important for early identification and intervention of CA patients who are at risk of dying in the ICU.

Supplementary Information

The online version contains supplementary material available at 10.1186/s40001-025-02588-2.

Keywords: Cardiac arrest, Random forest, ICU mortality risk, Shapley additive explanations

Introduction

Cardiac arrest (CA) refers to the sudden cessation of the heart's ejection function, leading to disappearance of large artery pulsations and heart sounds, as well as severe ischemia and hypoxia of vital organs such as the heart, liver, kidneys, and brain. It is characterized by the interruption of body circulation, respiratory arrest, and loss of consciousness [1, 2]. The incidence of CA is high and continues to rise annually [3]. With a high mortality rate, CA is one of the leading causes of death. According to the 2023 heart disease and stroke statistics, the incidence of in-hospital cardiac arrest (IHCA) among hospitalized patients in the United States is 292,000 cases per year. Additionally, the incidence of out-of-hospital cardiac arrest (OHCA) treated by Emergency Medical Services (EMS) is 92.3 per 100,000 population across all ages [4]. Despite increased focus on rescue CA from healthcare system, the survival rate of CA patients remained low. The annual survival rate for CA is less than 20% [5]. Early assessment of the prognosis of CA patients is crucial for reducing mortality rates [6].

Patients with CA are initially admitted to intensive care unit (ICU) for advanced life support treatment. Clinical information is importance in predicting their prognosis. Machine learning has been extensively studied in prognostic prediction of ICU patients, in the prediction of risk of ICU death in sepsis [7], trauma [8], and Covid 19 [9], and sometimes with automated machine-learning [9]. Several studies have already reported on the prediction of in-hospital mortality rates for CA patients [6, 10]. Some of these studies focused on predicting the in-hospital mortality rates of CA patients using individual single indicators [1115]. Their predictive performance is limited and no one has focused on predicting the risk of ICU death in CA patients. To better categorize patients based on their risk profiles and facilitate stratified diagnosis and treatment, thereby reducing mortality rates among CA patients, we have developed and validated machine learning-based prediction model for ICU mortality risk. This study holds significant implications for intensive care, subsequent decision-making, and treatment of CA patients.

Methods

Data source and study design

This study is a retrospective cohort study based on an open, large-scale critical care database called the Medical Information Mart for Intensive Care IV (MIMIC-IV, version 2.2), and multicenter critical care database called eICU (version 2.0.1), collected from Beth Israel Deaconess Medical Center in Boston between 2008 and 2019 and multiple U.S. hospitals between 2014 and 2015, respectively. By obtaining a certificate (Certificate ID: 11718300) from the Collaborative Institutional Training Initiative (CITI), we were granted permission to extract data from the database.

We extracted data from the MIMIC-IV and eICU database that met the inclusion criteria. The data from MIMIC-IV was randomly divided into a training set and internal validation set in a 7:3 ratio. Firstly, we extracted the clinical information of CA patients within 24 h of ICU admission. Secondly, variables with more than 20% missing values were removed. Then, Least Absolute Shrinkage and Selection Operator (LASSO) regression was used to select variables with predictive value for ICU mortality in CA patients. The training set was used for model training, while the internal and external validation sets were used for model validation. The best model obtained was compared with traditional disease severity scores, such as Sequential Organ Failure Assessment (SOFA), Oxford Acute Severity of Illness Score (OASIS), Simplified Acute Physiology Score II (SAPS II), Acute Physiology Score III (APS III), and Logistic Organ Dysfunction Score (LODS) to assess which is better at assessing the severity of CA patients. Finally, the Shapley Additive Explanations (SHAP) method was used to interpret the optimal model.

Study patients

Patients first time diagnosed with CA who were over 18 years old and treatment in ICU for more than 24 h were included in the study.

Data extraction

We used Structured Query Language (SQL, version 15.1) to extract data from MIMIC-IV (version 2.2) and eICU (version 2.0.1). The predictive models included clinical indicators and laboratory data from patients'first day in the ICU. If multiple measurements of vital signs and laboratory parameters were recorded on the first day, the average values were used for subsequent analysis. We extracted six types of data: ① demographic features, including age, gender, weight, height, and Body Mass Index(BMI); ② comorbidities, such as hypertension, diabetes, congestive heart failure, myocardial infarction, peptic ulcer, cerebrovascular disease, chronic obstructive pulmonary disease, and the Charlson Comorbidity Index; ③ vital signs, including heart rate (HR), systolic blood pressure, diastolic blood pressure, mean arterial pressure, respiratory rate (RR), body temperature, and peripheral oxygen saturation; ④ laboratory results, such as white blood cell count, red blood cell distribution width, platelet count, hematocrit, hemoglobin, prothrombin time, international normalized ratio (INR), partial thromboplastin time, fibrinogen, alanine aminotransferase (ALT), alkaline phosphatase, aspartate aminotransferase (AST), total bilirubin, lactate dehydrogenase, albumin, blood urea nitrogen (BUN), creatinine (Cr), troponin T(TnT), creatine kinase MB(CK-MB), lactate, pH, PO2, PCO2, PO2/FIO2, base excess, anion gap, bicarbonate, serum calcium, serum chloride, serum sodium, serum potassium, and blood glucose; ⑤ other indicators, including duration of ICU stay and first-day urine output; ⑥ traditional severity scores, such as the Glasgow Coma Scale (GCS), Sequential Organ Failure Assessment (SOFA), Oxford acute severity of illness score (OASIS), Simplified Acute Physiology Score II(SAPS II), Acute Physiology and Chronic Health Evaluation III (APS III), and Logistic Organ Dysfunction Score(LODS). The clinical outcome of this study is ICU mortality rate. Table S1 illustrate a detailed list of the included variables.

Ethics statements

The database received approval from the Massachusetts Institute of Technology and Beth Israel Deaconess Medical Center. This study is a retrospective study and does not affect clinical treatment and care. Therefore, ethical approval statements and informed consent from the subjects were given up [16]. This study is consistent with the Transparent Reporting of Multivariate Predictive Models for Individual Prognosis or Diagnosis (TRIPOD): TRIPOD statement [17], showed in Table S2.

Data preprocessing

Data processing was performed using R and Python, in which Python was only used to model interpretation. Firstly, the MIMIC-IV cohort was divided into survival and death groups based on clinical outcomes, and the clinical variables of the two groups were compared for differences. Secondly, missing values analysis was conducted, as presented in Table S3, and data with more than 20% missing values were excluded. Subsequently, multiple imputation was performed to fill missing values in data with less than 20% missing. Fig S1 shows the density plot after imputation. Fig S2 illustrates the data distribution before and after imputation, showing that interpolation has less impact on the data. The Glasgow Coma Scale (GCS) score was included in the modeling data, since it is only an assessment of the nervous system, while other disease severity scores were excluded. LASSO regression was employed for variables selection. Fig S3a illustrates the cross-validation plot for LASSO regression. Fig S3b displays the coefficient profile plot against the Log Lambda sequence for predicting mortality risk in the ICU. Table S4 presents the selected variables, their coefficient not equal to 0 by LASSO regression, for constructing the mortality risk prediction model for CA patients. Selecting these features in eICU database that met the inclusion criteria.

Model development and validation

The data from MIMIC-IV was divided into a training set and internal validation set in a 7:3 ratio, and data from eICU was used for external validation. We trained the models, including Logistic Regression (LR), Random Forest (RF), K-nearest Neighbor (KNN), Extreme Gradient Boosting (XGBoost), and Decision Tree (DT), on the training set and assessed their performance on the validation sets. Performance features such as the area under the Receiver Operating Characteristic (ROC) curve (AUC), accuracy, precision, recall, specificity, and F1 score were used to evaluate the models'performance and select the best-performing model with the highest AUC and other good performance. This model was then compared with traditional severity scores (SOFA, OASIS, SAPS II, APS III, and LODS) to identify a clinical prediction model that better predicts the risk of death in ICU for CA patients.

Model explainability

SHAP was used to improve the interpretability of the best-performing model. SHAP is a machine learning interpretation method that explains the importance of variables in the model's prediction results [18].It is based on the concept of SHAP values in cooperative game theory and uses an additive method to calculate the contribution of each feature to the model's prediction results [19]. Each sample is considered a participant in a game, and the combination of feature values is regarded as a cooperative strategy of the participants. SHAP can rank the importance of individual features by order of magnitude, making it easier to identify features that have a greater impact on the outcome.

Statistical analysis

All data processing, statistical analysis, model construction, and model interpretation were performed in R 4.2.1 and Python 3.7 software. Continuous variables were expressed as mean ± standard deviation (SD) for those that conform to a positive distribution, median (interquartile spacing) for those that are not, and number (percentage) for categorical variables. Continuous variables were tested using t-tests and rank-sum tests, and categorical variables were tested using chi-square tests. After data preprocessing, selection predictive models (LR, RF, KNN, XGBoost, DT) were constructed to predict ICU risk of death in CA patients.

Results

Baseline characteristics

There was a total of 299,712 patients’ data in MIMIC-IV, and 1,088 patients were included in this study, of which 761 patients were included in training set and 327 patients were included in internal validation set, and 3542 patients from eICU included in external validation set. Figure 1 shows the flowchart of this study. Table 1 shows baseline characteristics of the entire MIMIC-IV cohort, the ICU death and survival cohort. There’s no statistically significant difference between ICU survivors and the death groups in age, height, weight, and BMI. The incidence of CA patients was higher in males (674 (61.9%)) compared to females (414 (38%)), and proportion of males in the death group was also higher than that of females. Patients with cerebrovascular disease comorbidities were at higher risk in ICU mortality. Indicators related to liver, kidney, and myocardial injury were significantly higher in the ICU death group than survival group. ALT [109.3 (39.7, 298.0) vs. 51.5 (27.0, 143.5)], AST [166.0 (70.0, 457.4) vs. 80.0 (40.5, 203.2)], and ALP [95.0 (66.5, 140.0) vs. 79.9 (57.0, 109.8)] values for the ICU death group and survival group were used to assess the liver function. Meanwhile, BUN [26.6 (18.3, 44.2) vs. 22.0 (15.5, 34.0)] and creatinine [1.5 (1.0, 2.3) vs. 1.1 (0.8, 1.8)] were used to assess renal function. The indicators of myocardial injury that we are most concerned about in CA patients were TnT [0.4 (0.1, 1.6) vs 0.2 (0.1, 0.8)] and CKMB [18.0 (7.0, 52.6) vs 9.0 (4.0, 27.0)], and they were also significantly higher in the ICU death group. Some indicators of circulatory monitoring, such as lactate [3.7 (2.2, 5.7) vs 2.4 (1.6, 3.8)] and urine [1072.5 (401.5, 1914.3) vs 1455.0 (904.5, 2317.5)] were worse in the ICU death group than survival group. Similarly, the scores related to traditional disease severity [SOFA (11 (8, 13) vs 8 (4, 11)], OASIS [44 (38, 50) vs 37 (30, 45], SAPSII [51 (40, 63) vs 41 (32, 52)], APSIII [85 (65, 102) vs 56 (49, 82)], and LODS [10 (7, 12) vs 7 (4, 10)] showed the same trend.

Fig. 1.

Fig. 1

Flowchart of the study

Table 1.

Baseline characteristics of the MIMIC-IV cohort, and ICU death and survival groups

Characteristics All (N = 1088) Survival (N = 702) Non-survival(N = 386) P value
Demographic
 Age, year 67.1 (55.8,79.6) 67.1 (56.5,79.7) 67.2 (54.7,79.5) 0.709
Gender 0.035
 Male, n (%) 674 (61.9) 451 (64.2) 223(57.8)
 Female, n (%) 414 (38.1) 251 (35.8) 163 (42.2)
Weight, kg 81.0 (69.0,96.9) 82 (69.9,92.3) 80.0 (67.4,95.1) 0.028
Height, cm 170(163,178) 170(163,178) 170(163,178) 0.386
BMI 27.8(24.4,32.3) 28.1(24.6,32.4) 27.2(24.1,32.3) 0.146
Comorbidities
 Charlson Comorbidity Index 6(4,8) 6(4,8) 6(4,8) 0.961
 Diabetes, n (%) 419(38.5) 307(43.7) 112(29.0)  < 0.001
 Hypertension, n (%) 696(64.0) 461(65.7) 235(60.9) 0.116
All (N = 1088) Survival (N = 702) Non-survival (N = 386) P value
Congestive heart failure, n (%) 419 (38.5) 307 (43.7) 112 (29.0)  < 0.001
Myocardial infarction, n (%) 326 (30.0) 222 (31.6) 104 (28.0) 0.107
Peptic ulcer, n (%) 27 (2.5) 22 (3.1) 5 (1.3) 0.062
Cerebrovascular disease, n (%) 182 (16.7) 103 (14.7) 79 (20.5) 0.014
Chronic pulmonary disease, n (%) 295 (27.1) 188 (26.8) 107 (22.7) 0.739
Renal disease, n (%) 271 (24.9) 186 (26.5) 85 (22.0) 0.103
Vital signs on day 1
 Heart rate, bpm 82 (71,95) 81 (70,93) 86 (72,100) 0.002
 Systolic blood pressure, mmHg 113 (105,123) 113 (106,124) 112 (104,122) 0.054
 Diastolic blood pressure, mmHg 63 (56,70) 62 (56,69) 64 (55,71) 0.515
 Mean arterial pressure, mmHg 78 (72,85) 78 (72,84) 78 (71,85) 0.508
 Respiratory rate 20 (18,23) 19 (17,22) 22 (19,25)  < 0.001
 Body temperature, ℃ 36.7 (36.3,37.1) 36.8 (36.5,37.1) 36.6 (35.8,37.0)  < 0.001
SpO2, % 98 (96,99) 98 (96,99) 98 (96,99) 0.079
Laboratory findings on day 1
 White blood cell, × 103/µL 12.9 (9.3,16.9) 12.3 (9.0,16.2) 14.2 (10.1,18.7)  < 0.001
 Red blood cell distribution width 14.6 (13.6,16.0) 14.4 (13.5,15.6) 14.8 (13.8,16.4)  < 0.001
 Platelets, × 103/µL 191 (142,253) 191 (144,249) 192 (139,255) 0.941
 Hematocrit, % 34.3 (29.1,39.7) 34.3 (29.0,39.4) 34.4 (29.6,40.1) 0.275
 Hemoglobin, g/dL 11.2 (9.6,13.1) 11.2 (9.6,13.1) 11.2 (9.46,13.1) 0.721
 Prothrombin Time, s 14.3 (12.6,17.6) 13.9 (12.4,16.5) 14.9 (13.0,19.3)  < 0.001
 International normalized ratio 1.3 (1.1,1.6) 1.3 (1.1,1.5) 1.4 (1.2,1.8)  < 0.001
 Partial thromboplastin time, s 35.6 (29.0,53.0) 34.1 (28.7,49.4) 38.6 (29.8,56.4) 0.006
 Alkaline phosphatase, U/L 84.0 (61.0,122.7) 79.9 (57.0,109.8) 95.0 (66.5,140.0)  < 0.001
All (N = 1088) Survival (N = 702) Non-survival (N = 386) P value
Lactate dehydrogenase, U/L 429.0 (286.0,795.3) 363.5 (269.0,591.0) 573.5 (351.1,1200.8  < 0.001
Albumin, g/L 3.3 (2.7,3.7) 3.3 (2.9,3.8) 3.2 (2.6,3.6) 0.002
Blood urea nitrogen, mg/dL 23.5 (16.0,38.0) 22.0 (15.5,34.0) 26.6 (18.3,44.2)  < 0.001
Serum creatinine, mg/dL 1.2 (0.9,2.0) 1.1(0.8,1.8) 1.5 (1.0,2.3)  < 0.001
Troponin T, ug/L 0.3 (0.1,1.0) 0.2 (0.1,0.8) 0.4 (0.1,1.6)  < 0.001
CK_MB, U/L 12.0 (5.0,36.0) 9.0 (4.0,27.0) 18.0 (7.0,52.6)  < 0.001
Lactate, mmol/L 2.7 (1.7,4.6) 2.4 (1.6,3.8) 3.7 (2.2,5.7)  < 0.001
pH 7.3 (7.3,7.4) 7.4 (7.3,7.4) 7.3 (7.3,7.4)  < 0.001
pO2, mmHg 142.5 (105.4,195.6) 148.5 (111.6,201.1) 133.8 (97.4,186.3)  < 0.001
pCO2, mmHg 39.0 (34.6,43.9) 39.0 (35.5,43.9) 39.0 (33.7,44.0) 0.113
PaO2/FiO2 ratio 244.3 (176.0,336.2) 249.7 (185.9,336.5) 232.1 (159.6,334.3) 0.050
Base excess − 3.4(− 7.0, − 0.3) − 2.5(− 5.6,0) − 5.2 (− 8.8, − 1.3)  < 0.001
Anion gap 16.0 (13.6,19.0) 15.3 (13.0,17.6) 17.5 (14.5,21.0)  < 0.001
Bicarbonate, mmol/L 19.5 (16.3,22.0) 20.0 (17.0,23.2) 18.0 (15.5,20.0) 0.483
Serum calcium, mmol/L 1.1 (1.1,1.2) 1.1 (1.1,1.2) 1.1 (1.1,1.2) 0.214
Serum chloride, mmol/L 105.9 (5.9) 105.6 (5.1) 106.5 (6.0) 0.359
Serum sodium, mmol/L 137.0 (134.4,140.0) 136.4 (134.0,139.0) 138.0 (135.0,140.5) 0.005
Serum potassium, mmol/L 4.1 (3.6,4.7) 4.1 (3.7,4.7) 4.1 (3.6,4.8) 0.869
Blood glucose, mg/dL 169.5 (134.0,222.3) 165.0 (131.0,210.6) 180.7 (140.4,251.1) 0.012
Others
 Duration of ICU stay this time, day 4.0 (2.2,8.1) 4.2 (2.3,8.6) 3.7 (1.9,6.8)  < 0.001
Severity of illness scores
 GCS 12.0 (5,15) 13 (8,14) 9 (3,15)  < 0.001
 SOFA 9 (5,12) 8 (4,11) 11 (8,13)  < 0.001
 OASIS 40 (33,47) 37 (30,45) 44 (38,50)  < 0.001
 SAPSII 44 (34,57) 41 (32,52) 51 (40,63)  < 0.001
 APSIII 68 (45,93) 56 (49,82) 85 (65,102)  < 0.001
 LODS 8 (5,11) 7 (4,10) 10 (7,12)  < 0.001

Model development and validation

We used LASSO regression to select variables to develop models. In the prediction model of ICU mortality risk in CA patients, RF showed the best performance as compared to the other models (LR, KNN, XGBoost, and DT). Their AUC values and 95% CI were 0.83 (0.78–0.88), 0.79 (0.74–0.84), 0.76 (0.70–0.82), 0.72 (0.66–0.78), and 0.74 (0.69–0.80) in internal validation set, and 0.71(0.68–0.74), 0.70(0.68–0.74), 0.65(0.62–0.68), 0.68(0.65–0.71), and 0.61(0.58–0.64) in external validation set, respectively. The ROC curve, Precision-Recall (PR) Curve, and Decision Curve Analysis (DCA) curve of RF in internal validation set were shown to have optimal performance, as can be seen in Fig. 2a, 2c, and 2e, respectively. The confusion matrix visualization of the models in internal validation set is shown in Fig S4. Fig S4c shows the confusion matrix for RF, which can be seen to have the highest number of correct predictions for the outcome (the largest sum of the number of correct predictions for survival and dead patients) compared to the other models. In addition, we compared the RF with the traditional disease severity scores, and RF still showed the best predictive performance. The ROC curves, PR curves, and DCA curves comparing RF to traditional disease severity scores in internal validation set are shown in Fig. 2b, d, f respectively.

Fig. 2.

Fig. 2

Comparison of the predictive performance of multiple models for ICU mortality risk in internal validation set. a Comparing the ROC curves of multiple models. b Comparison of ROC curves for RF versus traditional disease severity scores. c Comparing the PR curves of multiple models. d Comparison of PR curves for RF versus traditional disease severity scores. e Comparing the DCA curves of multiple models. f Comparison of DCA curves for RF versus traditional disease severity scores

Confusion matrix visualizations of traditional disease severity scores in internal validation set are showed in Fig S5, RF still has the best predictive performance. Table 2 shows the AUC, Accuracy, Precision, Recall, Specificity, and F1, and their 95%CI of multiple models in internal validation set. After a comprehensive comparison, RF was selected as the best model. We compared RF with traditional disease severity score in internal validation set in Table S5 and found RF also showed best performance. In external validation set also showed the similarity result that RF was the best performance model (table S6 and S7).

Table 2.

Comparison of multi-model related indicators in predicting ICU mortality risk in CA patients from internal validation set

Classifiers
95%CI
AUC ACC (%) Precision (%) Recall Specificity (%) F1
RF

0.83

0.75–0.91

75.23

70.55–79.91

76.63

68.93–84.33

0.89

0.82–0.96

50.86

44.59–57.13

0.82

0.79–0.86

LR

0.79

0.74–0.84

74.70

69.99–79.41

74.31

66.36–82.26

0.91

0.84–0.98

43.10

37.05–49.15

0.82

0.78–0.56

KNN

0.76

0.70–0.82

72.17

67.31–77.03

75.64

67.83–83.45

0.84

0.77–0.91

50.86

44.45–57.27

0.80

0.75–0.84

XGBoost

0.74

0.69–0.80

69.42

64.43–74.41

74.24

66.28–82.20

0.81

0.73–0.89

49.14

42.66–55.62

0.77

0.73–0.82

DT

0.72

0.66–0.78

71.56

66.67–76.45

73.98

66.00–81.96

0.86

0.78–0.94

44.83

38.62–51.04

0.80

0.75–0.84

Model explainability

After the construction and comparison of the models, RF was found to be the optimal prediction model. RF impurity/permutation importance for testing the strength of association between the dependent variable was performed, and showed in Fig S6. GCS_min, los_icu, aniongap, base excess, urine output were the top five important features for RF. We used SHAP values to interpret the model, and Fig. 3a show the SHAP summary plots for each variable of the model. The color represents the level of the eigenvalue, red for high eigenvalues and blue for low eigenvalues. The line with a SHAP value of 0 is the baseline. Each dot on the way represents an eigenvalue, and the farther it is from the baseline, the stronger its influence on the results. If it is located on the side with a positive SHAP value, it has a positive effect on the results, and if it is located on the side with a negative SHAP value, it has a negative effect on the results. In Fig. 3a, we found that GCS_min and the eigenvalue of expelled urine positively affected the prediction of ICU death in CA patients when the eigenvalues of GCS_min and expelled urine were low, while anion gap positively affected the prediction of ICU death when the eigenvalue of anion gap was high.

Fig. 3.

Fig. 3

SHAP summary and Mean absolute SHAP values for each clinical variable of the RF. a SHAP summary plot of the RF prediction model for ICU mortality risk in CA patients. b Mean absolute SHAP values for each clinical feature in RF for prediction of ICU mortality risk

We also plotted the mean of the absolute values of SHAP for each of the clinical features that were included in the model construction, as presented in Fig. 3b, which illustrate the extent to which variables affect the predictive performance of the model. The longer the bar in the graph, the greater the impact on model performance. The variables that had a greater impact on ICU death prediction performance were GCS_min, base excess, anion gap, and urine output.

Based on these results, we plotted the SHAP dependence of the variables that had significant impact on RF to explain the specific impact of clinical variables on the risk of ICU death. Figure 4 show the SHAP dependency plots of clinical variables on ICU mortality risk in CA patients. The vertical axis of the SHAP dependency plot is the SHAP value for the clinical characteristic, and the horizontal axis is the range of variation, with SHAP values greater than 0 indicating an increased risk of death. The value of GCS_min had the greatest effect on RF model (Fig. 4a), with a SHAP value of 0 for a GCS_min of 12. The smaller the value of GCS_min, the greater the value of SHAP with a higher risk of patient death. The condition of the internal environment has a large impact on death in the ICU, and the base excess and anion gap are important indicators of the internal environment, which may reflect the patient's disturbed internal environment (Fig. 4b, c). The critical value of urine was about 1000–1200 ml, and the smaller the urine output, the higher the risk of ICU death for the patient (Fig. 4d).

Fig. 4.

Fig. 4

SHAP dependency plots for the top 4 clinical variables that have the greatest impact on the outcome of the ICU mortality model for CA patients. a SHAP dependency plot of GCS_min. b SHAP dependency plot of Base excess. c SHAP dependency plot of Anion gap. d SHAP dependency plot of Urine output

Discussion

When a patient diagnosed with CA, there is a stagnation of blood circulation and a lack of nutrients and oxygen, resulting in a lack of supply to his organs and brain, and damage to the organs and the nervous system, which lead to post-cardiac arrest syndrome [20, 21]. The treatment of CA patients begins with first aid, followed by advanced life support, with a goal to restore organ function, especially neurological function. Most of the first aid is done at the scene or in emergency room, before admitting the patient into the ICU for advanced life support for recovery of organ and nerve function [2225]. Recovery of neurologic function affects patient prognosis [2628]. ICU is the second important barrier that CA patients need to go through, and it is also the key stage of the CA patient's rescue treatment.

Shock status on admission was associated with ICU death in CA patients [29]. Corrected serum calcium and lactate dehydrogenase to albumin ratio may be associated with prognosis in CA patients [30, 31]. Difficulty in accurately predicting the prognosis of a disease by a single indicator. Some studies have explored the prediction model of in-hospital mortality risk in patients after cardiac arrest [6, 10]. They missed the importance of predicting the risk of ICU death in CA patients. For building a valid, stable, and interpretable model to predict the risk of ICU mortality in CA patients, we developed predictive models using a large data from MIMIC-IV and eICU database. Feature selection is crucial in developing predicting models. We used LASSO regression to select features, and finally, six features were selected to construct the ICU mortality risk prediction models for CA patients. RF was the most reliable and stable model, which had the best predictive performance. We compared RF with traditional disease severity scores (SOFA, OASIS, SAPSII, APSIII, and LODS) and found that RF remained the most optimal.

Traditional machine learning algorithms are often disfavored for their lack of transparency and interpretability. Another strength of this study is the use of SHAP values to explain these machine learning models and reveal their black-box problems. RF, as the best performing model, was the one we focused on explaining. We calculated the SHAP values for each of the feature variables to assess their contribution to the prediction results. The overall SHAP summary plot helps us understand which features positively or negatively influence the predicted outcomes, while the importance feature plot provides an average assessment of the importance of the features incorporated into the modeling for the entire dataset. In addition, the SHAP dependency plot helps to observe how these features affect the output of the predictive model at different levels.

We found that GCS_min is very important in RF model we constructed. GCS is an objective rating scale used to assess states of consciousness and the degree of damage to the nervous system. It is the sum of the scores for the eye-opening, speech, and motor components. The total score is 15, with 13–15 being classified as mildly impaired consciousness, 9–12 as moderately impaired consciousness, and 8 or less as coma; the lower the score, the poorer the state of consciousness [32, 33]. We took the minimum GCS score within 24 h after admission to ICU, which can reflect the worst conscious state of the patient's and the severity of neurological function impairment, for model construction. The plot of SHAP dependence showed that the critical value of GCS for the prediction of ICU mortality risk was 12, below which patients have an increased risk of ICU death. Therefore, death of CA patients is closely related to the severity of their brain injury. A lower GCS score indicates worsen brain function, which may lead to a higher risk of death [34].

The anion gap is the difference between unmeasured anions and unmeasured cations in plasma, which is derived from the concentrations of three commonly measured ions, such as Na+, Cl−, HCO3[35]. It is one of the indicators for evaluating the acid–base balance of the biological internal environment [36]. Jun Chen et al. explored the relationship between anion gap and in-hospital mortality in patients with CA, and found that patients with higher anion gap values had significantly higher in-hospital mortality rates than patients with lower anion gap values and concluded that anion gap value was a predictor of mortality in patients with CA [13]. Beiping Hu et al. found that elevated albumin corrected anion gap is associated with poor in-hospital prognosis in CA patients [14]. We found that negative anion gap was a predictor of the risk of ICU death in CA patients, and the greater the anion gap value, the greater the SHAP value, with a higher their risk of ICU death (Fig. 3a and Fig. 4c). Base excess is also one of the indicators for detecting acid–base balance within organisms [37, 38]. Figure 4b illustrate that too low and too high a base excess, SHAP values are elevated and patients are at increased risk of death. Severe disturbances in the internal environment exacerbate damage to the patient's organs, and the patient's response to catecholamine rescue medications is markedly diminished, leading to failure of rescue CA patients [39]. Urine output is related to the patient's kidney function and the recovery of blood circulation [40]. The severity of the patient's impaired consciousness, the degree of disturbance of the internal environment, and the amount of urine output within 24 h of ICU admission were important features in predicting the risk of ICU death in patients with CA.

The successful construction of a prediction model for the risk of ICU death in patients with CA allows early warning of patients at high-risk of ICU death. The model can help physicians to focus on the condition of high-risk CA patients and take effective measures to rescue them, which may improve the success of rescuing CA patient and reduce the mortality of high-risk CA patients.

Machine learning modeling is a branch of artificial intelligence (AI) that is widely used in several areas of medicine, affecting clinical treatment decisions as well as patient prognosis. Therefore, the safety of AI should not be underestimated. Europe has issued an AI security approach that emphasizes the challenges posed by AI technology in 2024 [41]. Data security and privacy protection are important for AI [42]. The MIMIC-IV and eICU database are de-identified to ensure that patient privacy is protected and access to data is only granted to users who have signed the DUA credentials, signed the PhysioNet Authenticated Health Data Usage Agreement 1.5.0, and have been trained in CITI data or sample-only studies. Therefore, the data from this study is secure and guarantees patient privacy.

Study limitation

There are some limitations in our study. This study was a retrospective study. In addition, although certain methods were taken to ensure the reliability and generalizability of the model, the results still need to be verified prospective clinical studies. Furthermore, we focused only on the clinical indicators within 24 h after ICU admission and did not assess the impact of changes in the clinical features on the outcomes during the ICU stay. Therefore, further design of multicenter prospective studies is needed to validate our findings.

Conclusion

In summary, we successfully used machine learning methods to predict the ICU risk of death in CA patients. RF was the best performing model in our study. We employed SHAP to explain the optimal model, help determine the importance of the variables incorporated into the model, and show how each variable affects the model. This is important for early identification and intervention of CA patients who are at risk of dying at ICU.

Supplementary Information

Acknowledgements

The authors sincerely thank the bioinformatics team of Southwest Medical University for supporting this study. This research fund by Luzhou Medical Association (grant number: 2024-YXXM-040, Yong Li).

Abbreviations

AUC

Area of ROC curve

APS III

Acute physiology score III

ALT

Alanine aminotransferase

AST

Alkaline phosphatase

BMI

Body mass index

BUN

Blood urea nitrogen

CA

Cardiac arrest

CITI

Collaborative Institutional Training Initiative

Cr

Creatinine

CK-MB

Creatine kinase MB

DT

Decision Tree

DCA

Decision curve analysis

EMS

Emergency medical services

GCS

Glasgow coma scale

HR

Heart rate

IHCA

In-hospital cardiac arrest

ICU

Intensive care unit

INR

International normalized ratio

KNN

K nearest neighbor

LR

Logistic regression

LASSO

Least absolute Shrinkage and selection operator

LODS

Logistic organ dysfunction score

MIMIC-IV

Medical information mart for intensive care

OHCA

Out-of-hospital cardiac arrest

OASIS

Oxford acute severity of illness score

PR

Precision-recall

RF

Random forest

ROC

Receiver operating characteristic

RR

Respiratory rate

SOFA

Sequential organ failure assessment

SAPS II

Simplified acute physiology score II

SQL

Structured query language

SHAP

Shapley additive explanations

SD

Standard deviation

TRIPOD

Transparent reporting of multivariate predictive models for individual

TnT

Troponin T

XGBoost

Extreme gradient boost

Author contributions

Y. L. developed inclusion and exclusion criteria, extracted the data, processed the data, and prepared the initial draft of the manuscript. Y. L. developed inclusion and exclusion criteria and extracted the data. Q. Z. and H. Z. performed a literature search. C.W. and X. J. revised the paper and reviewed the language of the final version of the manuscript. All authors have read and agreed to the published version of the manuscript.

Availability of data and materials

No datasets were generated or analysed during the current study.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Chengli Wen, Email: wenchengli076@swmu.edu.cn.

Xian Jiang, Email: 441315510@qq.com.

References

  • 1.Jacobs I, Nadkarni V, Bahr J, Berg RA, Billi JE, Bossaert L, Cassan P, Coovadia A, D’Este K, Finn J, Halperin H, Handley A, Herlitz J, Hickey R, Idris A, Kloeck W, Larkin GL, Mancini ME, Mason P, Mears G, Monsieurs K, Montgomery W, Morley P, Nichol G, Nolan J, Okada K, Perlman J, Shuster M, Steen PA, Sterz F, Tibballs J, Timerman S, Truitt T, Zideman D. Cardiac arrest and cardiopulmonary resuscitation outcome reports: update and simplification of the Utstein templates for resuscitation registries: a statement for healthcare professionals from a task force of the International Liaison Committee on Resuscitation (American Heart Association, European Resuscitation Council, Australian Resuscitation Council, New Zealand Resuscitation Council, Heart and Stroke Foundation of Canada, InterAmerican Heart Foundation, Resuscitation Councils of Southern Africa). Circulation. 2004;110(21):3385–97. 10.1161/01.Cir.0000147236.85306.15. [DOI] [PubMed] [Google Scholar]
  • 2.Elfassy MD, Randhawa VK, Allan KS, Dorian P. Understanding etiologies of cardiac arrest: seeking definitional clarity. Can J Cardiol. 2022;38(11):1715–8. 10.1016/j.cjca.2022.08.005. [DOI] [PubMed] [Google Scholar]
  • 3.Ravindran R, Kwok CS, Wong CW, Siller-Matula JM, Parwani P, Velagapudi P, Fischman DL, Alraies C, Michos ED, Mamas MA. Cardiac arrest and related mortality in emergency departments in the United States: analysis of the nationwide emergency department sample. Resuscitation. 2020;157:166–73. 10.1016/j.resuscitation.2020.10.005. [DOI] [PubMed] [Google Scholar]
  • 4.Tsao CW, Aday AW, Almarzooq ZI, Anderson CAM, Arora P, Avery CL, Baker-Smith CM, Beaton AZ, Boehme AK, Buxton AE, Commodore-Mensah Y, Elkind MSV, Evenson KR, Eze-Nliam C, Fugar S, Generoso G, Heard DG, Hiremath S, Ho JE, Kalani R, Kazi DS, Ko D, Levine DA, Liu J, Ma J, Magnani JW, Michos ED, Mussolino ME, Navaneethan SD, Parikh NI, Poudel R, Rezk-Hanna M, Roth GA, Shah NS, St-Onge MP, Thacker EL, Virani SS, Voeks JH, Wang NY, Wong ND, Wong SS, Yaffe K, Martin SS. Heart disease and stroke statistics-2023 update: a report from the American heart association. Circulation. 2023;147(8):e93–621. 10.1161/cir.0000000000001123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Andersen LW, Holmberg MJ, Løfgren B, Kirkegaard H, Granfeldt A. Adult in-hospital cardiac arrest in Denmark. Resuscitation. 2019;140:31–6. 10.1016/j.resuscitation.2019.04.046. [DOI] [PubMed] [Google Scholar]
  • 6.Chen J, Mei Z, Wang Y, Shou X, Zeng R, Chen Y, Liu Q. A nomogram to predict in-hospital mortality in post-cardiac arrest patients: a retrospective cohort study. Pol Arch Intern Med. 2023. 10.2452/pamw.16325. [DOI] [PubMed] [Google Scholar]
  • 7.Zheng R, Qian S, Shi Y, Lou C, Xu H, Pan J. Association between triglyceride-glucose index and in-hospital mortality in critically ill patients with sepsis: analysis of the MIMIC-IV database. Cardiovasc Diabetol. 2023;22(1):307. 10.1186/s12933-023-02041-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Qi L, Geng X, Feng R, Wu S, Fu T, Li N, Ji H, Cheng R, Wu H, Wu D, Huang L, Long Q, Wang X. Association of glycemic variability and prognosis in patients with traumatic brain injury: a retrospective study from the MIMIC-IV database. Diabetes Res Clin Pract. 2024;217:111869. 10.1016/j.diabres.2024.111869. [DOI] [PubMed] [Google Scholar]
  • 9.Sakagianni A, Koufopoulou C, Kalles D, Loupelis E, Verykios VS, Feretzakis G. Automated ML techniques for predicting COVID-19 mortality in the ICU. Stud Health Technol Inform. 2023;305:517–20. 10.3233/shti230547. [DOI] [PubMed] [Google Scholar]
  • 10.Sun Y, He Z, Ren J, Wu Y. Prediction model of in-hospital mortality in intensive care unit patients with cardiac arrest: a retrospective analysis of MIMIC -IV database based on machine learning. BMC Anesthesiol. 2023;23(1):178. 10.1186/s12871-023-02138-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhang N, Lin Q, Jiang H, Zhu H. Age-adjusted Charlson comorbidity index as effective predictor for in-hospital mortality of patients with cardiac arrest: a retrospective study. BMC Emerg Med. 2023;23(1):7. 10.1186/s12873-022-00769-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lin Q, Zhang N, Zhu H. The relationship between the level of NMLR on admission and the prognosis of patients after cardiopulmonary resuscitation: a retrospective observational study. Eur J Med Res. 2023;28(1):424. 10.1186/s40001-023-01407-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chen J, Dai C, Yang Y, Wang Y, Zeng R, Li B, Liu Q. The association between anion gap and in-hospital mortality of post-cardiac arrest patients: a retrospective study. Sci Rep. 2022;12(1):7405. 10.1038/s41598-022-11081-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Hu B, Zhong L, Yuan M, Min J, Ye L, Lu J, Ji X. Elevated albumin corrected anion gap is associated with poor in-hospital prognosis in patients with cardiac arrest: a retrospective study based on MIMIC-IV database. Front Cardiovasc Med. 2023;10:1099003. 10.3389/fcvm.2023.1099003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tang Y, Sun J, Yu Z, Liang B, Peng B, Ma J, Zeng X, Feng Y, Chen Q, Zha L. Association between prothrombin time-international normalized ratio and prognosis of post-cardiac arrest patients: a retrospective cohort study. Front Public Health. 2023;11:1112623. 10.3389/fpubh.2023.1112623. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Johnson AE, Pollard TJ, Shen L, Lehman LW, Feng M, Ghassemi M, Moody B, Szolovits P, Celi LA, Mark RG. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3:160035. 10.1038/sdata.2016.35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Heus P, Reitsma JB, Collins GS, Damen J, Scholten R, Altman DG, Moons KGM, Hooft L. Transparent reporting of multivariable prediction models in journal and conference abstracts: TRIPOD for abstracts. Ann Intern Med. 2020. 10.7326/m20-0193. [DOI] [PubMed] [Google Scholar]
  • 18.Nordin N, Zainol Z, Mohd Noor MH, Chan LF. An explainable predictive model for suicide attempt risk using an ensemble learning and shapley additive explanations (SHAP) approach. Asian J Psychiatr. 2023;79:103316. 10.1016/j.ajp.2022.103316. [DOI] [PubMed] [Google Scholar]
  • 19.Ning Y, Ong MEH, Chakraborty B, Goldstein BA, Ting DSW, Vaughan R, Liu N. Shapley variable importance cloud for interpretable machine learning. Patterns. 2022;3(4):100452. 10.1016/j.patter.2022.100452. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Binks A, Nolan JP. Post-cardiac arrest syndrome. Minerva Anestesiol. 2010;76(5):362–8. [PubMed] [Google Scholar]
  • 21.Medicherla CB, Lewis A. The critically ill brain after cardiac arrest. Ann N Y Acad Sci. 2022;1507(1):12–22. 10.1111/nyas.14423. [DOI] [PubMed] [Google Scholar]
  • 22.Lavonas EJ, Akpunonu PD, Arens AM, Babu KM, Cao D, Hoffman RS, Hoyte CO, Mazer-Amirshahi ME, Stolbach A, St-Onge M, Thompson TM, Wang GS, Hoover AV, Drennan IR. American heart association focused update on the management of patients with cardiac arrest or life-threatening toxicity due to poisoning: an update to the american heart association guidelines for cardiopulmonary resuscitation and emergency cardiovascular care. Circulation. 2023;148(16):e149–84. 10.1161/cir.0000000000001161. [DOI] [PubMed] [Google Scholar]
  • 23.Nielsen N, Skrifvars MB. Oxygenation and blood-pressure targets in the ICU after cardiac arrest - one step forward. N Engl J Med. 2022;387(16):1517–8. 10.1056/NEJMe2211024. [DOI] [PubMed] [Google Scholar]
  • 24.Peltan ID, Poll JB, Guidry D, Brown SM, Beninati W. Acceptability and perceived utility of telemedical consultation during cardiac arrest resuscitation a multicenter survey. Ann Am Thorac Soc. 2020;17(3):321–8. 10.1513/AnnalsATS.201906-485OC. [DOI] [PubMed] [Google Scholar]
  • 25.Nolan JP, Berg RA, Bernard S, Bobrow BJ, Callaway CW, Cronberg T, Koster RW, Kudenchuk PJ, Nichol G, Perkins GD, Rea TD, Sandroni C, Soar J, Sunde K, Cariou A. Intensive care medicine research agenda on cardiac arrest. Intensive Care Med. 2017;43(9):1282–93. 10.1007/s00134-017-4739-7. [DOI] [PubMed] [Google Scholar]
  • 26.Benghanem S, Pruvost-Robieux E, Bouchereau E, Gavaret M, Cariou A. Prognostication after cardiac arrest: how EEG and evoked potentials may improve the challenge. Ann Intensive Care. 2022;12(1):111. 10.1186/s13613-022-01083-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Dragancea I, Wise MP, Al-Subaie N, Cranshaw J, Friberg H, Glover G, Pellis T, Rylance R, Walden A, Nielsen N, Cronberg T. Protocol-driven neurological prognostication and withdrawal of life-sustaining therapy after cardiac arrest and targeted temperature management. Resuscitation. 2017;117:50–7. 10.1016/j.resuscitation.2017.05.014. [DOI] [PubMed] [Google Scholar]
  • 28.Paul M, Bougouin W, Geri G, Dumas F, Champigneulle B, Legriel S, Charpentier J, Mira JP, Sandroni C, Cariou A. Delayed awakening after cardiac arrest: prevalence and risk factors in the parisian registry. Intensive Care Med. 2016;42(7):1128–36. 10.1007/s00134-016-4349-9. [DOI] [PubMed] [Google Scholar]
  • 29.Ni J, Liu Y, Wu M, Wang J, Sha D, Xu B. Shock on admission as a potential marker for ICU mortality of cardiac arrest patients. Int Heart J. 2020;61(4):795–8. 10.1536/ihj.20-040. [DOI] [PubMed] [Google Scholar]
  • 30.Zhong L, Lu J, Sun X, Sun Y. The association between albumin-corrected calcium and prognosis in patients with cardiac arrest: a retrospective study based on the MIMIC-IV database. Eur J Med Res. 2024;29(1):251. 10.1186/s40001-024-01841-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Ye L, Lu J, Yuan M, Min J, Zhong L, Xu J. Correlation between lactate dehydrogenase to albumin ratio and the prognosis of patients with cardiac arrest. Rev Cardiovasc Med. 2024;25(2):65. 10.31083/j.rcm2502065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Mehta R, Chinthapalli K. Glasgow coma scale explained. Bmj. 2019;365:l1296. 10.1136/bmj.l1296. [DOI] [PubMed] [Google Scholar]
  • 33.Reith FC, Van den Brande R, Synnot A, Gruen R, Maas AI. The reliability of the glasgow coma scale: a systematic review. Intensive Care Med. 2016;42(1):3–15. 10.1007/s00134-015-4124-3. [DOI] [PubMed] [Google Scholar]
  • 34.Seymour CW, Kahn JM, Cooke CR, Watkins TR, Heckbert SR, Rea TD. Prediction of critical illness during out-of-hospital emergency care. Jama. 2010;304(7):747–54. 10.1001/jama.2010.1140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Moe OW, Fuster D. Clinical acid-base pathophysiology: disorders of plasma anion gap. Best Pract Res Clin Endocrinol Metab. 2003;17(4):559–74. 10.1016/s1521-690x(03)00054-x. [DOI] [PubMed] [Google Scholar]
  • 36.Achanti A, Szerlip HM. Acid-base disorders in the critically Ill patient. Clin J Am Soc Nephrol. 2023;18(1):102–12. 10.2215/cjn.04500422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Berend K. Diagnostic use of base excess in acid-base disorders. N Engl J Med. 2018;378(15):1419–28. 10.1056/NEJMra1711860. [DOI] [PubMed] [Google Scholar]
  • 38.Langer T, Brusatori S, Gattinoni L. Understanding base excess (BE): merits and pitfalls. Intensive Care Med. 2022;48(8):1080–3. 10.1007/s00134-022-06748-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Yagi K, Fujii T. Management of acute metabolic acidosis in the ICU: sodium bicarbonate and renal replacement therapy. Crit Care. 2021;25(1):314. 10.1186/s13054-021-03677-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Verma S, Kellum JA. Defining acute kidney injury. Crit Care Clin. 2021;37(2):251–66. 10.1016/j.ccc.2020.11.001. [DOI] [PubMed] [Google Scholar]
  • 41.Kalodanis K, Rizomiliotis P, Anagnostopoulos D. European artificial intelligence act: an AI security approach. Inform Computer Secur. 2024;32(3):265–81. [Google Scholar]
  • 42.Feretzakis G, Verykios VS. Trustworthy AI: securing sensitive data in large language models. AI. 2024;5(4):2773–800. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from European Journal of Medical Research are provided here courtesy of BMC

RESOURCES