Abstract
BACKGROUND Hyperglycemic crisis is associated with substantial morbidity and mortality. Existing prognostic tools largely rely on static baseline measurements and do not exploit the longitudinal data generated during acute care. We aimed to develop a real-time dynamic prediction model for in-hospital mortality within the next 24 h in adults with hyperglycemic crisis.
METHODS We performed a multicenter retrospective study using eICU-CRD and MIMIC-IV databases. We developed HCNet, a temporal deep learning architecture that integrates longitudinal trajectories with static features to generate a continuously updated 24-hour mortality risk. The eICU-CRD was used for development and internal testing; MIMIC-IV served as an external test cohort. Discrimination, calibration, and decision curve analysis were compared with those of eight machine learning models, including logistic regression (LR), random forest (RF), eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), light gradient boosting machine (LightGBM), multilayer perceptron (MLP), tabular data network (TabNet), and long short-term memory (LSTM).
RESULTS The study included 3,436 eICU-CRD training admissions, 1,473 eICU-CRD internal test admissions, and 1,182 MIMIC-IV external test admissions. In-hospital mortality was 1.8%, 1.9%, and 3.1%, respectively. HCNet achieved an AUC of 0.941 (95% CI: 0.939 to 0.943) in internal testing and 0.913 (95% CI: 0.911 to 0.916) in external testing, outperforming all comparators (P<0.05). The strongest baseline models according to the AUC were CatBoost (internal 0.916; external 0.871) and LSTM (internal 0.911; external 0.890). Temporal analyses showed improved discrimination with increasing length of stay and as death approached, as well as a reduced false alarm burden.
CONCLUSION HCNet provides real-time 24-hour mortality risk estimates for adults with hyperglycemic crisis using routinely collected data. Its generalizable performance and interpretable temporal patterns support its potential as a clinical decision support tool for early warning and escalation of care during hyperglycemic crisis.
Keywords: Hyperglycemic crisis, Mortality, Deep learning
INTRODUCTION
Hyperglycemic crisis is among the most serious acute complications of diabetes mellitus and encompasses diabetic ketoacidosis (DKA), a hyperosmolar hyperglycemic state (HHS), and mixed DKA-HHS.[1] Despite advances in diabetes care and improved recognition of acute metabolic decompensation, the incidence of DKA and HHS continues to increase, particularly among younger adults. Epidemiological data suggest that approximately 10% of deaths in individuals with diabetes are attributable to confirmed or suspected DKA or hyperglycemic coma, and reported in-hospital mortality for HHS ranges from 5% to 16%, approximately an order of magnitude higher than for DKA.[1-4] These findings underscore both the substantial clinical burden of hyperglycemic crisis and the marked heterogeneity in prognosis among patients who appear to have similar metabolic derangements.
Standard management of hyperglycemic crisis, including prompt fluid resuscitation, intravenous insulin, and careful correction of electrolyte and acid-base disturbances, has been well established in international guidelines.[5,6] Nonetheless, outcomes remain highly variable. Some patients recover uneventfully with routine care, whereas others experience rapid clinical deterioration, multiorgan failure, and death despite guideline-concordant treatment.[5,7,8] Early recognition of those at highest risk is therefore crucial to guide timely escalation of care, such as transfer to a higher-acuity unit, initiation of organ support, or more intensive monitoring.
Most prior prognostic studies in hyperglycemic crisis have relied on conventional statistical models or simple scoring systems derived from baseline clinical and biochemical variables.[9-13] While such tools are easy to implement, they inherently treat risk as static and assume that a limited set of early measurements can fully characterize the patient’s trajectory.[14-16] In clinical practice, however, the condition of patients with hyperglycemic crisis evolves rapidly over hours in response to treatment, intercurrent infection, hemodynamic changes, and the development or resolution of organ dysfunction.[17] Repeated measurements of vital signs, mental status, and laboratory indices, such as lactate, electrolyte, and acid-base parameters, provide rich longitudinal information that is largely ignored when only the first set of results is used for risk estimation.[8,18] As a result, many models are not well suited to real-time decision making in acute care.
Temporal machine learning models can use longitudinal data to update risk as new information becomes available.[19-21] Prior prognostic work in hyperglycemic crisis has typically relied on real-time or early aggregated variables, has often been developed in single-center cohorts, and has offered limited insight into how risk evolves during admission.[9,11 -13,22 -25]
To address these gaps, we developed hierarchical context network (HCNet), a real-time dynamic prediction model for in-hospital mortality in adults with hyperglycemic crisis using multicenter data. HCNet integrates time-varying clinical trajectories with static patient characteristics to generate continuously updated 24-hour mortality risk estimates throughout the hospital stay. We conducted external validation, assessed the model’s temporal performance and false alarm profile, and applied interpretable modeling techniques to characterize the evolving contribution of key clinical variables.
METHODS
Study population
We conducted a retrospective cohort study of adult patients who presented with hyperglycemic crises, identified using International Classification of Diseases codes (ICD-9-CM: 250.1 or 250.2; ICD-10: E11.0 or E11.1). Data were obtained from two large, independent critical care databases: the eICU Collaborative Research Database (eICU-CRD v2.0),[26] and the Medical Information Mart for Intensive Care Database (MIMIC-IV v2.2).[27] The MIMIC-IV contains 431,231 admissions to the ED or ICU at Beth Israel Deaconess Medical Center (BIDMC) between 2008 and 2019. From the MIMIC-IV, we included encounters for which clinical data were available from the index ED visit and, when applicable, from any subsequent inpatient stay. The eICU-CRD includes data from more than 200,000 adult ICU stays across 208 hospitals in the United States between 2014 and 2015. Unlike the MIMIC-IV, the eICU-CRD does not provide a directly linked ED dataset. Therefore, in the eICU-CRD we identified all the ICU stays with a diagnosis of hyperglycemic crisis and extracted all available relevant clinical variables recorded during the ICU stay and any linked ward stay, when available. For both databases, we excluded admissions if the patient was younger than 18 years, if the total hospital length of stay exceeded 60 days, or if all variables required for model development were missing. Details of the cohort selection process are summarized in supplementary Figure 1.
Ethical approval
The MIMIC-IV database is publicly available to qualified researchers following Institutional Review Board (IRB) approval obtained from Beth Israel Deaconess Medical Center (Boston, USA) and the Massachusetts Institute of Technology (Cambridge, USA). The eICU-CRD is similarly accessible for public research use under appropriate IRB oversight from the 208 participating hospitals across the United States. All datasets consisted of deidentified patient records; therefore, individual informed consent was not needed. The principal investigator (PGX) obtained access to both databases after completing the mandated National Institutes of Health (NIH) training and successfully passing the relevant certification examinations (Record ID: 51524821).
Data extraction and preprocessing
The overall study workflow, including data extraction, preprocessing, model development, validation, and interpretation, is presented in Figure 1. Guided by prior literature and specialist clinical expertise, we selected 38 routinely available variables, including demographic characteristics (e.g., age and sex), vital signs, and commonly obtained laboratory measurements. For each patient, we generated a sample at every time point when at least one of these variables was recorded. Values lying outside predefined clinically plausible ranges were considered invalid and were treated as missing. The predefined ranges for all variables are listed in supplementary Table 1.
Figure 1. Overview of the study framework. A: data extraction and preprocessing; B: architecture of the proposed HCNet model. eICU-CRD: eICU Collaborative Research Database; MIMIC: Medical Information Mart for Intensive Care; HR: hazard ratio; SBP: systolic blood pressure; LSTM: long short-term memory.

To ensure a fair and unbiased comparison between HCNet and all comparator models, we applied an identical preprocessing pipeline across all cohorts and methods. This pipeline included standardized variable extraction, filtering based on clinically plausible ranges, normalization, and construction of two temporal features: (i) time elapsed since the last available vital sign measurement and (ii) time elapsed since the last available laboratory test. Missing values were handled using a forward-imputation strategy.[28] For the first occurrence of a missing value without any prior observation, we imputed the cohort median. For subsequent missing values, we carried forward the most recent available observation.
To assess robustness, we additionally performed sensitivity analyses using two alternative imputation strategies: median imputation only and K-nearest neighbors (KNN) imputation. All downstream model development and evaluation procedures were repeated under these alternative preprocessing schemes using the same training-test partitions and evaluation framework.
Prediction target and sample construction
The target of this study was not generic in-hospital mortality at an arbitrary time point, but rather dynamic prediction of death within the next 24 h. Accordingly, for patients who died during their hospital stay, we constructed samples from measurements obtained within the 24 h preceding death. For patients who survived to discharge, we constructed samples from measurements collected throughout the entire hospital stay. In all the cases, data from the time of presentation to the ED through to the end of the index hospital stay (ED discharge or inpatient discharge) were included.
Model development
We used the multicenter eICU-CRD cohort for model development and reserved the MIMIC-IV cohort exclusively for external validation. The eICU-CRD dataset was randomly partitioned into a training set and an internal test set using stratified sampling by outcome status, with a 7:3 split at the admission level. The MIMIC-IV cohort was not involved in model training or hyperparameter selection and served solely as an external test cohort to assess out-of-sample performance and generalizability. All comparator models underwent hyperparameter tuning using the same training cohort and five-fold cross-validation, with harmonized preprocessing and the same evaluation framework.
All the experiments were conducted on a workstation equipped with a 3.50 GHz 13th Generation Intel Core i5-13600KF central processing unit and an NVIDIA RTX A6000 graphics processing unit with 48 GB of memory. We optimized HCNet using the Adam algorithm with a weight decay of 0.0001, β1 of 0.9, and β2 of 0.999. The initial learning rate was set to 0.001, the mini-batch size was set to 64, and training proceeded for up to 50 epochs. Model parameters corresponding to the best validation performance during cross-validation were retained for subsequent testing.
HCNet is a temporal deep learning architecture designed to integrate time-varying clinical trajectories with static patient characteristics. At each time step, the input vector comprises 40 features, including 2 static variables (age and sex) and 38 longitudinal clinical measurements. The time-varying component is first processed by a two-layer unidirectional long short-term memory (LSTM) network with a hidden dimension of 128 units.
The LSTM outputs are passed through a linear projection layer with a residual connection and layer normalization to stabilize training and enrich the learned temporal representation. On top of these contextualized temporal features, we apply a multi-head self-attention module along the time dimension. This temporal attention mechanism employs a causal attention mask to ensure that each time point can attend only to its current and preceding time steps, thereby preventing information leakage from future observations. A padding mask is applied simultaneously to exclude positions beyond the true sequence length from the attention computation. The attention outputs are combined with the LSTM representations via residual connections and layer normalization, and a global temporal representation is obtained by mean pooling across the time dimension.
Static patient features are modeled in a parallel branch. These variables are passed through a two-layer multilayer perceptron with rectified linear unit activation and dropout, yielding a dense static representation in the same hidden space. This static embedding is concatenated with the global temporal representation to form a fused patient-level feature vector that captures both baseline risk and the dynamic evolution of clinical status over time.
For risk estimation, the attention-enhanced temporal features are further processed by a fully connected prediction head to generate time-step-specific mortality risk scores, yielding a continuous risk trajectory throughout hospitalization. This design allows the model to update the estimated risk as new measurements become available while strictly preserving the temporal ordering of the data.
Comparative models
We compared the performance of HCNet with eight widely used baseline methods. These included a conventional linear model, logistic regression (LR), and several tree-based machine learning algorithms: random forest (RF),[29] eXtreme Gradient Boosting (XGBoost),[30] Categorical Boosting (CatBoost),[31] and light gradient boosting machine (LightGBM).[32] For deep learning comparators, we implemented a multilayer perceptron (MLP), TabNet, and a standard LSTM network, the latter being a commonly used architecture for time-series prediction tasks. The hyperparameters for all the comparative models were tuned using five-fold cross-validation on the training cohort to ensure fair and robust performance estimates.
Model evaluation
Model discrimination was primarily assessed using the area under the receiver operating characteristic curve (AUC). We also reported the balanced accuracy, sensitivity, specificity, F1 score, positive predictive value (PPV), negative predictive value (NPV), and false alarms per 100 patient-days (FAC/100 patient-days). To quantify uncertainty, we calculated 95% confidence intervals using 1,000 bootstrap resamples.
Because the task is dynamic prediction, risk scores were generated repeatedly across time within each admission as new measurements became available. For this reason, operational metrics such as FAC/100 patient-days depend on two choices: the probability threshold used to trigger an alert and the frequency at which the model is evaluated over time. To keep comparisons consistent across methods, we selected a single threshold for each model using the Youden index during cross-validation on the training cohort. We then applied this fixed threshold to the internal and external test cohorts. This approach was used to standardize evaluation across models; in clinical implementation, the threshold and alerting policy would need to be selected based on local workflow, staffing, and tolerance for alarm burden.
Calibration was assessed using calibration plots, expected calibration error (ECE), and the Brier score. To explore potential clinical usefulness across a range of decision thresholds, we performed decision curve analysis (DCA).
Statistical analysis
Continuous variables are summarized as the mean ± standard deviation (SD), and categorical variables are reported as counts and percentages. Comparative differences in predictive performance between the models were evaluated using the DeLong test for correlated AUCs. A two-sided P-value < 0.05 was considered to indicate statistical significance.
Model explanation
To interpret model predictions, we employed the Shapley Additive Explanations (SHAP) framework, which enables both global and instance-level assessments of feature contributions.[33] To reduce computational burden while maintaining class balance, we constructed explanation cohorts for both the internal and external test sets by randomly sampling a subset of survivors equal in number to the deceased admissions in each set. Within these balanced subsets, we quantified overall feature importance and characterized how the contribution of each predictor evolved over time.
RESULTS
Study cohorts
After applying all inclusion and exclusion criteria, 3,436 admissions (323,769 time-stamped samples) from the eICU-CRD were included in the training cohort. An additional 1,473 admissions (132,596 samples) from eICU-CRD and 1,182 admissions (103,670 samples) from MIMIC-IV were used as the internal and external test cohorts, respectively. Overall, 62 patients (1.8%) in the training cohort, 27 patients (1.9%) in the internal test cohort, and 35 patients (3.1%) in the external test cohort died. Baseline demographic and clinical characteristics, as well as the means and standard deviations of all input variables, are summarized by cohort and outcome group in supplementary Table 2.
Predictive performance
HCNet significantly outperformed all eight comparator models across both internal and external test cohorts (P< 0.05). Specifically, the model achieved an AUC of 0.941 (95% CI: 0.939-0.943) in the internal test dataset, 0.913 (95% CI: 0.911-0.916) in the external test dataset as illustrated in Figure 2 A and B. Additional evaluation metrics, including balanced accuracy, sensitivity, specificity, and F1 score, PPV, NPV, FAC/100 patient-days, ECE, and Brier score are detailed in supplementary Tables 3 and 4. The results of the calibration analysis demonstrated excellent agreement between predicted and observed risks. The calibration curves for HCNet closely followed the ideal 45° line and were associated with the lowest ECE among all models, with an ECE of 0.002 (95% CI: 0.002-0.002) in the internal test cohort and 0.004 (95% CI: 0.004-0.004) in the external test cohort (Figure 2 C and D). DCA further indicated that HCNet provided a higher net benefit than the comparators across a broad range of clinically relevant decision thresholds in both test cohorts, supporting its potential clinical utility (Figure 2 E and F). To evaluate the robustness of HCNet to missing-data handling, we compared the primary median + forward fill strategy with median imputation only and KNN imputation (supplementary Table 5). The performance of this model was similar across the three methods, indicating the reliability of the main research findings.
Figure 2. Performance of the models in internal and external test cohorts. A and B display the receiver operating characteristic curves in the internal and external test cohorts, respectively; C and D present the calibration curves for the same cohorts; E and F show the decision curve analysis results for the internal and external cohorts, respectively. LR: logistic regression; RF: random forest; XGBoost: eXtreme Gradient Boosting; LightGBM: light gradient boosting machine; CatBoost: Categorical Boosting; MLP: multilayer perceptron; TabNet: tabular data network; LSTM: long short-term memory; HCNet: hierarchical context network; AUC: area under the receiver operating characteristic curve; ECE: expected calibration error.

Temporal analysis
We next examined how the discriminative performance of HCNet varied over time, both as a function of time since ICU admission and as a function of time preceding death. This dual temporal perspective provided a comprehensive view of the model’s dynamic behavior (supplementary Figure 2). Model discrimination improved as the length of stay increased and clearly tended to increase as the terminal event was approached. These findings suggest that HCNet is capable of both early detection and increasing sensitivity to physiological deterioration as death nears. Importantly, this pattern underscores the value of leveraging longitudinal time-series data rather than relying solely on baseline measurements at admission.
To further evaluate clinical applicability, we analyzed the temporal profile of actionable lead time together with FAC/100 patient-days (supplementary Figure 3). Specifically, for each time window prior to death, we quantified the proportion of high-risk patients correctly identified. A patient was considered “identified” if their predicted probability of death exceeded a predefined risk threshold at least once before the specified time horizon. In parallel, we assessed how FAC/100 patient-days evolved over time. As the time of death approached, the proportion of correctly identified high-risk patients increased, while FAC/100 patient-days generally declined. This combination indicates that the model becomes both more sensitive and more efficient (i.e., fewer false alerts per unit patient-time) closer to the critical event, which is desirable for real-time clinical monitoring.
Feature importance
Global feature importance, quantified by SHAP values, is presented in supplementary Figure 4 for the 20 most influential predictors. Peripheral oxygen saturation (SpO2), systolic blood pressure (SBP), Glasgow Coma Scale (GCS) score, respiratory rate, age, and lactate emerged as dominant contributors to the model’s mortality risk estimates. These features reflect core aspects of respiratory, hemodynamic, and neurological function, which is consistent with the established clinical understanding of acute decompensation in hyperglycemic crises.
To characterize dynamic effects, we further investigated how feature importance changed over time (supplementary Figure 5). This temporal interpretation was based on the absolute sum of SHAPley values for the top 10 features across different time points. In non-survivors, lactate levels and GCS scores progressively decreased as death progressed, indicating the role of these parameters in signaling imminent clinical deterioration. In contrast, among survivors, most features tended to drive down the predicted risk over time, with the GCS score and respiratory rate playing particularly important protective roles. The same feature could have opposite effects on predicted mortality risk at different time points. This bidirectional influence emphasizes the need for real-time patient characteristics instead of relying on static, single-time-point assessments.
DISCUSSION
In this multicenter retrospective study, we developed and externally validated HCNet, a real-time dynamic prediction framework for short-horizon in-hospital mortality risk within the next 24 h among patients who presented with hyperglycemic crisis. By continuously integrating routinely collected longitudinal data from vital signs and laboratory tests with static patient characteristics, HCNet provides a patient-specific 24-hour mortality risk trajectory throughout admission, rather than a single risk estimate at admission. Across both internal and external test cohorts, the model demonstrated high discrimination, good calibration, and favorable decision-analytic performance, suggesting potential utility as a clinical early warning tool.
Most prior prognostic work in hyperglycemic crisis has used baseline or early aggregated measurements and has typically been developed within a single center, with limited external testing and limited ability to describe how risk drivers evolve over time.[9,11 -13,22 -25] These design choices make bedside use challenging in settings where physiology and treatment response can change rapidly over hours. In contrast, HCNet was built for time-updated prediction and evaluated across two independent databases using a harmonized preprocessing and evaluation framework. In our experiments, HCNet maintained high AUC in both cohorts, and the best-performing baseline methods by AUC (such as CatBoost, RF, and a standard LSTM) consistently performed below HCNet, particularly in the external test cohort.
A key contribution of this work is the explicit modeling and evaluation of temporal dynamics. Clinically, the course of hyperglycemic crisis is highly labile: patients may improve rapidly with appropriate therapy or deteriorate quickly due to infection, hemodynamic collapse, or evolving organ failure.[34,35] Our temporal analyses showed that discrimination improved as more data became available during the admission and increased as death approached. This pattern is clinically intuitive and suggests that compared with a single baseline snapshot, repeated reassessment using longitudinal information can better reflect the evolving state of an acutely ill patient.
False alarms are an equally important part of real-world feasibility. In this study, HCNet achieved strong discrimination but did not always minimize FAC/100 patient-days at the fixed operating point used for standardized comparisons. This is not contradictory. The AUC describes how well a model ranks risk across all thresholds, while FAC depends on the chosen threshold and on how often predictions are generated. When risk is recalculated many times during one admission, a threshold that favors sensitivity can trigger more alerts, especially in a rare-outcome setting where most alerts will necessarily occur in survivors. For bedside use, the threshold should not be treated as a universal constant. It should be selected to match local escalation capacity and tolerance for alert burden. In addition, alerting policies can substantially reduce alarm fatigue without changing the underlying model, for example by requiring persistent elevation over multiple updates, using tiered alerts (watch vs. urgent), or limiting repeat notifications within a defined time window.
From an implementation perspective, HCNet uses routinely recorded variables and is designed to be updated as new measurements appear in the electronic health records. A practical deployment would therefore involve automated data ingestion, consistent real-time preprocessing, and a display that supports action rather than simply reporting a probability. Equally important is governance: prospective “silent mode” evaluation, selection of an operating threshold aligned with workflow, and ongoing monitoring for calibration drift. These steps are particularly relevant because mortality in hyperglycemic crisis is uncommon, and alert policies need to be designed to avoid excessive alarms.[36]
An important consideration in this study is the very low mortality rate across cohorts. In rare-event prediction settings, the PPV and F1 score are strongly constrained by event prevalence, even when discrimination is high. Therefore, the relatively low PPV observed here should not be interpreted in isolation as evidence that the model lacks value. Instead, sensitivity, NPV, calibration, decision curves, and the intended clinical use case should be considered. At the same time, false alarms are a major practical concern for any early warning model. In our study, the threshold was chosen using the Youden index to standardize model comparison, not to define a clinically optimized alert policy. If HCNet were to be implemented prospectively, threshold selection would need to be tailored to local workflow, staffing, escalation capacity, and tolerance for alert burden. Additional strategies, such as tiered alerts, persistence requirements, or integration with clinician review, may be needed to reduce alarm fatigue while preserving clinical usefulness. Accordingly, the present findings should be interpreted within the scope of warning of near-term deterioration and should not be generalized to arbitrary-time, whole-stay mortality prediction.
A further point that merits clarification is the asymmetric sampling strategy between event and non-event admissions. This design follows directly from the target of the study: prediction of death within the next 24 h. For admissions ending in death, the final 24 h before death represent the clinically actionable window of interest. Including large amounts of much earlier data from event admissions would dilute the short-horizon prediction target and shift the task toward long-term prognosis, which was not the purpose of this model. Accordingly, the present findings should be interpreted within the scope of near-term deterioration warning and not generalized to arbitrary-time, whole-stay mortality prediction.
Several limitations should be noted. First, the study is retrospective and relies on historical data. The eICU-CRD and MIMIC-IV were collected during different time periods, and management practices, documentation, and monitoring patterns may have changed over time. This temporal shift can affect transportability in current practice. While external testing across two independent databases provides useful evidence of robustness under heterogeneity, it does not replace prospective validation in contemporary cohorts. Second, the outcome was rare, which constrains PPV and F1 score and increases the importance of carefully choosing thresholds and warning logic for any real-world use. Third, although we provided model explanations using SHAP, these explanations describe associations within the model rather than causal effects. Finally, subgroup performance assessments (e.g., by age, sex, and ethnicity), including both discrimination and calibration, may be explored in future work. This could be informative given potential differences in ethnicity documentation completeness and coding practices across datasets and institutions.
CONCLUSION
In summary, HCNet provides time-updated estimates of next-24-hour in-hospital mortality risk in adults with hyperglycemic crisis using routinely collected longitudinal data. The model showed strong discrimination and calibration in internal testing and external validation across two independent critical care databases. Given the retrospective design and the historical, non-concurrent nature of the datasets, prospective validation in contemporary cohorts and workflow-based evaluation of alert thresholds and policies are needed before clinical deployment.
Funding: The present study was supported by Chongqing Municipal Health Commission Medical Research Project (2025WSJK086). Chongqing Science and Health Joint Medical Research Project (2025QNXM045).
Ethical approval: The MIMIC-IV is publicly available after Institutional Review Board (IRB) approval by the Beth Israel Deaconess Medical Center in Boston, MA, USA and the Massachusetts Institute of Technology, MA, USA. The eICU-CRD is publicly available with appropriate IRB approval from 208 hospitals in the USA. All databases contain anonymized patient information, eliminating the need for individual informed consent.
Conflicts of interest: The authors declare no competing interests.
Contributors: YT, PGX, JX contributed equally to this article. Scientific guarantor: YM, GYY; study concepts or data acquisition or data analysis, all authors; manuscript drafting or manuscript revision for important intellectual content, all authors; approval of final version of submitted manuscript, all authors; agrees to ensure any questions related to the work are appropriately resolved, all authors; literature research, YT, PGX, JX; experimental studies, YT, PGX, JX, JXX, FYJ, HPF, JHY, BL; and manuscript editing, YM, GYY; all authors have accessed and verified the study data.
All the supplementary files in this paper are available at http://wjem.com.cn.
Contributor Information
Yu Ma, Email: 81846846@qq.com.
Gangyi Yang, Email: gangyiyang@hospital.cqmu.edu.cn.
REFERENCES
- [1]. Fayfman M, Pasquel FJ, Umpierrez GE. . Management of hyperglycemic crises: diabetic ketoacidosis and hyperglycemic hyperosmolar state. Med Clin North Am. 2017; 101(3): 587-606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2]. Bragg F, Holmes MV, Iona A, Guo Y, Du HD, Chen YP, et al. . Association between diabetes and cause-specific mortality in rural and urban areas of China. JAMA. 2017; 317(3): 280. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3]. Benoit SR, Hora I, Pasquel FJ, Gregg EW, Albright AL, Imperatore G. . Trends in emergency department visits and inpatient admissions for hyperglycemic crises in adults with diabetes in the U.S., 2006-2015. Diabetes Care. 2020; 43(5): 1057-64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4]. Umpierrez G, Korytkowski M. . Diabetic emergencies-ketoacidosis, hyperglycaemic hyperosmolar state and hypoglycaemia. Nat Rev Endocrinol. 2016; 12(4): 222-32. [DOI] [PubMed] [Google Scholar]
- [5]. Kitabchi AE, Umpierrez GE, Murphy MB, Barrett EJ, Kreisberg RA, Malone JI, et al. . Management of hyperglycemic crises in patients with diabetes. Diabetes Care. 2001; 24(1): 131-53. [DOI] [PubMed] [Google Scholar]
- [6]. Umpierrez GE, Davis GM, ElSayed NA, Fadini GP, Galindo RJ, Hirsch IB, et al. . Hyperglycaemic crises in adults with diabetes: a consensus report. Diabetologia. 2024; 67(8): 1455-79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7]. Galindo RJ, Ali MK, Funni SA, Dodge AB, Kurani SS, Shah ND, et al. . Hypoglycemic and hyperglycemic crises among U.S. adults with diabetes and end-stage kidney disease: population-based study, 2013-2017. Diabetes Care. 2022; 45(1): 100-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8]. González-Vidal T, Lambert C, García AV, Villa-Fernández E, Pujante P, Ares-Blanco J, et al. . Hypoglycemia during hyperosmolar hyperglycemic crises is associated with long-term mortality. Diabetol Metab Syndr. 2024; 16(1): 83. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9]. Xie PG, Yang C, Yang GY, Jiang YZ, He M, Jiang XY, et al. . Mortality prediction in patients with hyperglycaemic crisis using explainable machine learning: a prospective, multicentre study based on tertiary hospitals. Diabetol Metab Syndr. 2023; 15(1): 44. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10]. Adegbosin OT, Olamoyegun MA, Olarewaju SO. . Determinants and predictors of early re-admission of patients with hyperglycemic crises: a machine learning-based analysis. J Diabetes Metab Disord. 2025; 24(1): 85. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11]. He R, Zhang KB, Li H, Gu MP. . Development and validation of inpatient mortality prediction models for patients with hyperglycemic crisis using machine learning approaches. BMC Endocr Disord. 2025; 25(1): 86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12]. Hsu CC, Kao Y, Hsu CC, Chen CJ, Hsu SL, Liu TL, et al. . Using artificial intelligence to predict adverse outcomes in emergency department patients with hyperglycemic crises in real time. BMC Endocr Disord. 2023; 23(1): 234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13]. Huang CC, Kuo SC, Chien TW, Lin HJ, Guo HR, Chen WL, et al. . Predicting the hyperglycemic crisis death (PHD) score: a new decision rule for emergency and critical care. Am J Emerg Med. 2013; 31(5): 830-4. [DOI] [PubMed] [Google Scholar]
- [14]. Xie JY, Gao JD, Yang MT, Zhang T, Liu YC, Chen YT, et al. . Prediction of sepsis within 24 hours at the triage stage in emergency departments using machine learning. World J Emerg Med. 2024; 15(5): 379-85. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15]. Wu WM, Li M, Jiang HL, Sun M, Zhu YC, Zhu GX, et al. . Development of an emergency department length-of-stay prediction model based on machine learning. World J Emerg Med. 2025; 16(3): 220. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16]. Liu QY, Zhang YX, Sun J, Wang KP, Wang YG, Wang YL, et al. . Early identification of high-risk patients admitted to emergency departments using vital signs and machine learning. World J Emerg Med. 2025, 16(2): 113-20 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17]. Steenkamp DW, Alexanian SM, McDonnell ME. . Adult hyperglycemic crisis: a review and perspective. Curr Diabetes Rep. 2013; 13(1): 130-7. [DOI] [PubMed] [Google Scholar]
- [18]. He R, Zhang KB, Li H, Fu SM, Chen Z, Gu MP. . Impact of Charlson Comorbidity Index on in-hospital mortality of patients with hyperglycemic crises: a propensity score matching analysis. Evaluation Clinical Practice. 2024; 30(6): 977-88. [DOI] [PubMed] [Google Scholar]
- [19]. Xie PG, Hu Y, Li J, Ma Y, Xiao JJ. . Unlocking the potential of real-time ICU mortality prediction: redefining risk assessment with continuous data recovery. NPJ Digit Med. 2025; 8: 733. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20]. Xie PG, Wang H, Xiao J, Xu F, Guo R, Xie BB, et al. . Generative AI-powered explainable prediction model: enhancing early in-hospital mortality alert for patients with acute myocardial infarction. Int J Cardiol. 2025; 440: 133649. [DOI] [PubMed] [Google Scholar]
- [21]. Rhee TM, Ko YK, Kim HK, Lee SB, Kim BS, Choi HM, et al. . Machine learning-based discrimination of cardiovascular outcomes in patients with hypertrophic cardiomyopathy. JACC Asia. 2024; 4(5):375-86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22]. Ding Y, Qu L, Zhang J, Li T, Wang Y, Xu W, et al. . Parental health at preconception and gestational age at birth: evidence from a population-based cohort using double machine learning. World J Pediatr Surg. 2025; 8(4):e001078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23]. Kaewkrasaesin C, Kositanurit W, Chotwanvirat P, Laichuthai N. . Enhancing outcome prediction by applying the 2019 WHO DM classification to adults with hyperglycemic crises: a single-center cohort in Thailand. Diabetes Metab Syndr Clin Res Rev. 2024; 18(4): 103012. [DOI] [PubMed] [Google Scholar]
- [24]. Deng LL, Xie PG, Chen Y, Rui SL, Yang C, Deng B, et al. . Impact of acute hyperglycemic crisis episode on survival in individuals with diabetic foot ulcer using a machine learning approach. Front Endocrinol. 2022; 13: 974063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25]. Liao WT, Lee CC, Kuo CL, Lin KC. . Predicting readmission due to severe hyperglycemia after a hyperglycemic crisis episode. Diabetes Res Clin Pract. 2022; 192: 110115. [DOI] [PubMed] [Google Scholar]
- [26]. Pollard TJ, Johnson AEW, Raffa JD, Celi LA, Mark RG, Badawi O. . The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Sci Data. 2018; 5: 180178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27]. Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. . MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023; 10: 1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28]. Liu JH, Huang WC, Hu J, Hong N, Rhee Y, Li Q, et al. . Validating machine learning models against the saline test gold standard for primary aldosteronism diagnosis. JACC Asia. 2024; 4(12):972-84. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29]. Breiman L. . Random forests. Mach Learn. 2001;45:5-32. [Google Scholar]
- [30]. Chen TQ, Guestrin C. . XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery; San Francisco California, USA; 2016: 785-794. [Google Scholar]
- [31]. Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. . CatBoost: unbiased boosting with categorical features. Available at: https://arxiv.org/pdf/1706.09516 https://arxiv.org/pdf/1706.09516 [Google Scholar]
- [32]. Ke GL, Meng Q, Finley T, Wang TF, Chen W, Ma WD, et al. . LightGBM: a highly efficient gradient boosting decision tree. Neural Information Processing Systems. 2017. [Google Scholar]
- [33]. Nohara Y, Matsumoto K, Soejima H, Nakashima N. . Explanation of machine learning models using shapley additive explanation and application for real data in hospital. Comput Meth Programs Biomed. 2022; 214: 106584. [DOI] [PubMed] [Google Scholar]
- [34]. Gunst J, Derese I, Aertgeerts A, Ververs EJ, Wauters A, Van den Berghe G, et al. . Insufficient autophagy contributes to mitochondrial dysfunction, organ failure, and adverse outcome in an animal model of critical illness. Crit Care Med. 2013; 41(1): 182-94. [DOI] [PubMed] [Google Scholar]
- [35]. Richards JE, Scalea TM, Mazzeffi MA, Rock P, Galvagno SM. . Does lactate affect the association of early hyperglycemia and multiple organ failure in severely injured blunt trauma patients? Anesth Analg. 2018; 126(3): 904-10. [DOI] [PubMed] [Google Scholar]
- [36]. Doig GS, Heighes PT, Simpson F, Sweetman EA, Davies AR. . Early enteral nutrition, provided within 24h of injury or intensive care unit admission, significantly reduces mortality in critically ill patients: a meta-analysis of randomised controlled trials. Intensive Care Med. 2009; 35(12): 2018-27 [DOI] [PubMed] [Google Scholar]
