Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2026 Jun 11;9:728. doi: 10.1038/s41746-026-02874-1

Development and prospective evaluation of a real-time deep learning model for inpatient hypoglycemia prediction

Amanda Momenzadeh 1,✉, Caleb Cranney 1, Dennis H Chen 2, Elizabeth Vi Nguyen 3, Roma Gianchandani 2, Jesse G Meyer 1,✉
PMCID: PMC13612521  PMID: 42277312

Abstract

Inpatient hypoglycemia is associated with increased morbidity, mortality, length of stay, and healthcare costs, yet current management remains reactive due to the lack of real-time prediction tools. We developed, validated, and prospectively evaluated a real-time long short-term memory (LSTM) model to predict hypoglycemia within 24 h using electronic health record (EHR) data from 143,124 adult inpatient admissions across three hospitals between 2014 and 2025. Eligible patients were ≥ 18 years old, hospitalized for ≥ 24 h, and received at least one antihyperglycemic medication. Time-series predictors included medications, laboratory values, diet orders, and percentage of meals consumed over a 5-day lookback window segmented into 4-hour intervals, alongside static demographic variables. The primary outcome was blood glucose (BG) < 70 mg/dL within 24 hours following each prediction timepoint. Model hyperparameters were optimized using Bayesian optimization, and performance was compared with logistic regression, dense neural network, and XGBoost baselines using F1 score, precision, recall, area under the precision-recall curve (AUPRC), and calibration metrics. The best-performing LSTM model achieved an F1 score of 0.30 (95% CI 0.296–0.305), precision of 0.23, recall of 0.44, and AUPRC of 0.23 at a decision threshold of 0.7, outperforming all baseline models. Performance remained stable during prospective daily validation using live EHR extracts. SHapley Additive exPlanations (SHAP) identified clinically meaningful temporal predictors, including recent insulin administration and prior hypoglycemia. Model performance remained consistent across most demographic subgroups. This real-time deployable LSTM model provides a clinically interpretable prediction of inpatient hypoglycemia and may support proactive glycemic stewardship workflows in hospitalized patients.

Subject terms: Diseases, Health care, Medical research, Risk factors

Introduction

Inpatient hypoglycemia, defined as a blood glucose (BG) level below 70 mg/dL1,2, is one of the most common adverse events during diabetes treatment in hospitalized patients3. In the U.S., inpatient hypoglycemia affects individuals with and without diabetes4, occurring in approximately 20% of patients receiving insulin5, 10% of intensive care unit (ICU) patients6, and 3.5% of non-ICU patients6. Hypoglycemia can cause troublesome, acute symptoms such as confusion, impaired vision, and seizures7–10. Inpatient hypoglycemia is also associated with increased risk of long-term cerebrovascular and cardiovascular complications7–12, prolonged length of stay, and higher in-hospital mortality13,14. Despite its clinical significance, most inpatient hypoglycemia management remains reactive, with treatment adjustments commonly made only after an inpatient hypoglycemia event has already occurred3,15,16.

Predicting inpatient hypoglycemia is inherently complex. Risk factors include a history of prior hypoglycemia3 (e.g., 84% of inpatients with a BG < 40 mg/dL had a prior BG < 70 mg/dL in the same stay2), intensive insulin therapy, end-stage renal disease, cognitive impairment, female sex, and age ≥ 75 years2. Many risk factors are poorly understood1; for example, unexpected inpatient medications can have significant associations with BG4. These risks fluctuate during hospitalization due to changes in acute illness, nutritional intake, and evolving insulin needs5, and standard inpatient insulin dosing protocols often do not account for these factors6. Other barriers, such as provider confidence, knowledge gaps, complex workflows, and clinical inertia, may impede effective inpatient glycemic management7. Current care is usually reactive, supporting the adjustment of treatment plans only after a hypoglycemic event has occurred3,4,16. Although the 2024 American Diabetes Association (ADA) guidelines recommend individualized strategies for inpatient hypoglycemia prevention4, there remains no widely adopted tool to support the proactive prediction of inpatient hypoglycemia in real time other than provider experience and adjustment.

Machine Learning (ML) offers promising capabilities for transforming inpatient diabetes management by leveraging the wealth of electronic health records (EHR), drawing from past clinician decisions and patient outcomes to inform individualized care8. However, 77% of published hypoglycemia prediction models have been developed using continuous glucose monitoring (CGM) device data from outpatient populations with Type 1 diabetes mellitus (T1DM)17–19. In the hospital setting, patients rarely wear CGM devices; instead, BG is monitored via point-of-care (POC) fingerstick testing at intervals (e.g., before meals or every 4–6 h). Only 5% of hypoglycemia prediction studies use EHR data that includes POC BG from inpatients with T1DM, Type 2 diabetes mellitus (T2DM), and stress hyperglycemia19. Among these few models that use inpatient EHR data, most predict outcomes at the admission level16,20–23. Models that predict a single outcome for an entire admission substantially oversimplify the prediction task, limiting their clinical utility. In contrast, real-time prediction requires that models learn from all BG measurements, as patients at risk for hypoglycemia may experience both high and low values across time. Additionally, current ML studies using POC BG data from inpatient EHRs apply standard ML methods such as random forest (RF), which do not use time-series data and are trained instead on aggregated historical features (e.g., mean, median, and standard deviation [SD] of medications and lab values). Lastly, these studies typically report metrics (e.g., area under receiver operating characteristic [AUROC]) that can be misleading for rare event prediction problems like inpatient hypoglycemia. We report F1, the harmonic mean of precision and recall, to balance false positives and false negatives and support deployment-oriented threshold selection, and area under the precision-recall curve (AUPRC), which summarizes performance without selecting a single operating threshold24,25.

To our knowledge, only two studies have utilized inpatient EHR data that incorporates all POC BG values, not CGM, for real-time hypoglycemia prediction. Neither used time-series models, and both relied on descriptive statistics as inputs. Mathioudakis et al. (2021)26 used a stochastic gradient boosting (SGB) model on 1.6 million POC BG values from 55,000 admissions to predict BG < 70 mg/dL within a rolling 24-hour window. The model achieved precision (positive predictive value) 0.09, recall (sensitivity) 0.82, and F1 score 0.16. Similarly, Zale et al. (2022)27 used a RF model on 4.5 million POC BG values from 185,000 admissions to perform multiclass classification (BG < 70, 70–180, > 180 mg/dL), achieving precision 0.12, recall 0.72, and F1 score 0.21. Specificity, although high in the study by Zale et al., was largely driven by correctly predicting true negatives27. The low precision ( ~ 0.1) in both studies indicates high number of false alerts, contributing to alert fatigue. Translation of a clinical decision support tool into practice requires careful threshold selection to optimize the trade-off between sensitivity and alert burden. For example, lowering the threshold increases sensitivity and identifies more patients who will experience hypoglycemia, but raises the false positive rate, contributing to alert fatigue. Conversely, raising the threshold reduces alert burden and improves precision, but at the cost of missed events. Thus, probability thresholds should be selected not only based on model performance metrics, but also on the practical workflow and tolerance for alert burden in the clinical setting.

Recurrent neural networks (RNNs) capture relationships between sequential datapoints, such as POC BGs over time, and can prioritize relevant features out of millions of variables28–30. Long short-term memory (LSTM) models are a type of RNN that have an intermediate memory cell that can retain information over extended periods, making them ideal for modeling complex temporal dynamics.

LSTMs have been widely used with CGM data in the outpatient management of T1DM8,11,31. However, existing studies have not demonstrated that CGM-based hypoglycemia prediction models generalize directly to patients monitored using inpatient POC BG testing. CGM-based models rely on high-frequency BG trajectories, using features such as the difference between the current and 10 minutes prior collected CGM value32,33, which are not observable with the sparse and irregular sampling of POC BG readings. These current CGM-based models also do not incorporate clinical variables critical for hypoglycemia risk in the hospital, such as medications or indicators of acute illness. The time horizon of prediction further distinguishes these approaches; models using CGM focus on prediction of imminent risk of hypoglycemia (e.g., 30–60 min ahead)20,21, whereas longer prediction horizons (e.g.,12–24 h ahead) are more clinically actionable in the hospital setting as they allow for proactive insulin and nutritional management.

To address these critical gaps, we developed a LSTM-based model for predicting inpatient hypoglycemia within a 24 h prediction horizon using POC BG data and structured clinical variables from the EHR that advances the field in several ways. First, we emphasize interpretability by using SHapley Additive exPlanations (SHAP) to generate patient-level, temporal explanations of model predictions34, helping clinicians understand not only the estimated risk of hypoglycemia but also the clinical factors contributing to that risk and when they occurred. Second, we perform longitudinal external validation using daily EHR extracts from a live clinical environment to ensure the model maintains generalizability over time. Third, we assess model performance across demographic subgroups, recognizing that even well-performing models can exacerbate healthcare disparities if subgroup performance is not evaluated35–37. By integrating real-time prediction, patient-level interpretability, ongoing external validation, and fairness evaluation, our model provides a robust foundation for clinically actionable decision support that aims to improve inpatient glycemic safety.

Results

Model training and optimization

Models were trained and tested on 143,124 inpatient admissions collected between May 14, 2014 and March 26, 2025. There were 14,999 medications, 5,808 labs, 104 diet orders, 34,713 ICD-10 codes, and 13 socio-demographic features available in the dataset. Across the inpatient admissions, there were 4.8 million POC BG measurements, averaging ~30 POC BGs per admission. Hypoglycemia (BG < 70 mg/dL) occurred in 19% of admissions, with 2.8 events per admission. There was an average of 23 hypoglycemic events experienced by 8 patients per day. An analysis of temporal patterns of POC BG values indicate the frequency of hypoglycemia peaks between 6 and 9 AM (SFig. 1A), the annual distribution of POC BG values across four BG categories indicates stable proportions of hypoglycemia and hyperglycemia over time with a slight increase in hyperglycemia and decrease in hypoglycemia (SFig. 1B), and most POC BG readings fall between 80 and 200 mg/dL(SFig. 1C). The proportion of BGs < 70 mg/dL out of all recorded BG values per year remains consistently low, ranging from 1.3% to 1.9% annually.

Among all admissions, 10% had T1DM, 78% T2DM, and the remainder received antihyperglycemic agents without a formal diabetes diagnosis. The average age was 66.5 years, body mass index (BMI) 27.8 kg/m2, 43% were female, 59% Caucasian, 18% Black or African American (AA), and 9% Asian. Additionally, 58% had chronic kidney disease (CKD), 10% had a prior hypoglycemia diagnosis code, 50% had Medicare insurance, and 44% had commercial insurance. 10% were admitted to the ICU, and 88% of admissions occurred at CSMC, with 8% at HH and 4% at MDRH. STable 1 presents patient characteristics across the train, validation, test, and daily external validation cohorts. Demographic and clinical features are well-balanced across splits. Most patients were admitted to the academic medical center, Cedars-Sinai Medical Center (CSMC), with smaller proportions from two affiliated community hospitals, Marina Del Rey Hospital (MDRH) and Huntington Health (HH).

A total of 6,426,254 prediction windows were generated across all admissions. Consistent with the rarity of inpatient hypoglycemia, only 4.4% of prediction windows were labeled positive, reflecting significant class imbalance and the low incidence of inpatient hypoglycemia in the dataset (Table 1). The class distribution remained stable across splits.

Table 1.

Distribution of hypoglycemia and control labels across the training, validation, and test sets

# (%) Hypoglycemic # (%) Control
Training output 181,705 (4.43%) 3,915,364 (95.57%)
Validation output 44,866 (4.32%) 993,671 (95.68%)
Test output 57,181 (4.43%) 1,233,467 (95.57%)

Labels were generated at 4-hour intervals based on whether a hypoglycemic event (BG < 70 mg/dL) occurred within the subsequent 24 h.

To identify the optimal temporal context for hypoglycemia prediction, we evaluated model performance across lookback (LB) windows of 2, 3, 5, and 7 days, each segmented into 4 h intervals (STable 2). For each window length, models were trained using both class weighting (CW) and weighted binary cross-entropy (WBCE) to address class imbalance, with a fixed batch size of 4096. The best model (5-day LB, CW, 1 LSTM layer with 112 units and dropout 0.5, learning rate 0.0003) achieved the highest mean F1 score of 0.302 (95% CI: 0.298–0.306), with mean precision of 0.23 and mean recall of 0.44 when evaluated on a held-out test set at a decision threshold of 0.7, offering the best overall balance between sensitivity and specificity. Mean AUPRC, which summarizes discrimination across all decision thresholds (DTs) and is therefore threshold-independent, was 0.23. Figure 1 illustrates model performance across DTs using bootstrapped evaluation on the held-out test set. The LSTM consistently achieved the highest F1 values, with peak performance at 0.7 (F1 0.30; Fig. 1a). XGBoost demonstrated intermediate performance, while the dense neural network and LR showed substantially lower F1 scores across thresholds. As expected, precision increased (Fig. 1b) and recall decreased as the threshold increased (Fig. 1c), reflecting the expected tradeoff between sensitivity and specificity. The dense neural network, for example, had low overall F1 scores as the reduction in recall offset its precision gains. In contrast, the LSTM maintained a more favorable balance between precision and recall.

Fig. 1. Bootstrapped model performance across decision thresholds.

Fig. 1

(a) F1 score, (b) Precision, and (c) Recall for the LSTM, XGBoost, dense neural network, and logistic regression (LR) models across decision thresholds from 0.5 to 0.9 on the held-out test set. Points represent the mean performance across 100 bootstrap resamples at each threshold, and error bars denote 95% confidence intervals. This figure was generated in Python using Matplotlib.

This equates to 15 alerts/day, with 3.4 false positives per true alert, and the opportunity to prevent 3.5 hypoglycemia events per day (44% of the 8 anticipated hypoglycemic patients/day). We also evaluated performance under a patient-level (rather than encounter-level) split, finding that the best LSTM achieved a mean F1 score of 0.295 (95% CI: 0.291–0.298) at a threshold of 0.7, which was comparable to that observed under the admission-level split. Both shorter and longer LB windows (e.g., 2, 3, or 7 days) resulted in lower F1 scores, likely due to insufficient context or dilution of predictive signal from older data. These results support the use of a 5-day LB window with CW as the optimal configuration for real-time prediction of hypoglycemia.

Model calibration was evaluated on the held-out test set, which was divided into equal-sized validation and final test subsets. As shown in SFig. 2, uncalibrated predictions consistently underestimated the true outcome rate, particularly at higher predicted probabilities. In contrast, the isotonic calibrations improved alignment with the ideal diagonal line, indicating improved reliability. This improvement is quantitatively supported by the Brier score, which decreased from 0.127 (uncalibrated) to 0.0378 (isotonic-calibrated).

Comparator model performance was initially evaluated using temporally flattened 2D representations of the full longitudinal inputs, which discard explicit temporal structure but preserve all observed values. As shown in Table 2, the dense neural network, logistic regression (LR), and XGBoost models achieved mean F1 scores of 0.18, 0.15, and 0.26, respectively, under this representation. In contrast, the LSTM model, which explicitly modeled temporal dependencies, achieved a significantly higher F1 score of 0.30, as assessed by Mann-Whitney U testing of bootstrapped F1 distributions.

Table 2.

Model performance comparison between use of sequential inputs (LSTM) and temporally flattened inputs (Dense, LR and XGBoost)

Model Optimized HPs Best DT F1 (95% CI) Precision (95% CI) Recall (95% CI) AUPRC (95% CI) P-value
Best LSTM Dense units: 88, Learning rate: 0.000308, # LSTM layers: 1, LSTM units layer 1: 112, Dropout layer 1: 0.5 0.7 0.301 (0.296–0.305) 0.227 (0.223–0.231) 0.444 (0.439–0.452) 0.232 (0.228–0.237) —
Dense Learning rate: 0.000939, # LSTM layers: 2, LSTM units layer 1: 128, Dropout layer 1: 0.6, LSTM units layer 2: 104, Dropout layer 2: 0.5 0.5 0.175 (0.173–0.178) 0.107 (0.105–0.108) 0.490 (0.484–0.496) 0.111 (0.108–0.113) < 0.01
LR max_iter = 10, class_weight = ‘balanced’, solver = ‘saga' 0.5 0.151 (0.149–0.152) 0.090 (0.089–0.091) 0.464 (0.459–0.466) 0.088 (0.087–0.089) < 0.01
XGBoost n_estimators=100, max_depth=6, Learning rate=0.05 0.7 0.264 (0.260–0.267) 0.236 (0.233–0.238) 0.300 (0.197–0.303) 0.198 (0.195–0.201) < 0.01

The LSTM model was trained on full longitudinal time-series data, whereas baseline models (Dense neural network, logistic regression (LR), and XGBoost) used temporally flattened two-dimensional inputs that ignore sequential structure. For each model, the best decision threshold (DT) was selected based on F1 score. Reported values represent the mean performance on the held-out test set with 95% confidence intervals (CIs) estimated via bootstrap resampling. Model performance is summarized using F1 score, precision, recall, and area under the precision-recall curve (AUPRC). Statistical comparisons between the LSTM and each baseline model were performed using the Mann-Whitney U test applied to the bootstrapped F1 score distributions at the selected DT, with resulting p-values reported in the final column.

The three baseline models were then trained using summaries of the time-varying features over the 5-day LB window (i.e., mean, minimum, maximum, most recent value, SD, range, and interquartile range [IQR]) combined with the static features. When trained on these summary-statistic representations, LR, XGBoost, and dense neural network models demonstrated consistently lower performance compared with their counterparts trained on flattened longitudinal inputs (STable 3), indicating loss of predictive information when short-term temporal patterns were collapsed into aggregate summaries.

To assess the relative contribution of each input modality, we trained separate models using only a single input source using a 5-day LB window and 4-hour intervals, as shown in Table 3. All models used a batch size of 4096 and CW. Among individual inputs, laboratory values and medications yielded the highest F1 scores (0.21 and 0.22, respectively). In contrast, models using only % meals consumed (F1 0.069), diet orders (F1 0.093), or static history (F1 0.15) performed poorly, highlighting their limited predictive value in isolation. Notably, the best-performing configuration was achieved when all input modalities were combined (F1 0.302; Table 2), underscoring the value of multiple structured EHR data inputs to capture complementary signals for hypoglycemia prediction.

Table 3.

Model performance using single input types compared to all inputs

Single Input Best DT Optimized HPs F1 (95% CI) Precision (95% CI) Recall (95% CI)
Medications 0.7 Dense units: 80, Learning rate: 0.000248, # LSTM layers: 2, LSTM units layer 1: 152, Dropout layer 1: 0.5, LSTM units layer 2: 112, Dropout layer 2: 0.5 0.217 (0.213– 0.220) 0.188 (0.185–0.191) 0.256 (0.252–0.260)
Laboratory 0.7 Dense units: 88, Learning rate: 1.518e-05, # LSTM layers: 1, LSTM units layer 1: 240, Dropout layer 1: 0.7 0.214 (0.211–0.218) 0.190 (0.186–0.194) 0.246 (0.241–0.250)
% Meals Consumed 0.5 Dense units: 64, Learning rate: 4.693e-05, # LSTM layers: 2, LSTM units layer 1: 32, Dropout layer 1: 0.8, LSTM units layer 2: 24, Dropout layer 2: 0.9 0.0688 (0.0671–0.0701) 0.0420 (0.0409–0.0429) 0.190 (0.186–0.194)
Diet Orders 0.5 Dense units: 64, Learning rate: 5.892e-05, # LSTM layers: 2, LSTM units layer 1: 120, Dropout layer 1: 0.8, LSTM units layer 2: 96, Dropout layer 2: 0.5 0.0926 (0.0912–0.0939) 0.0523 (0.0515–0.0531) 0.401 (0.395– 0.406)
Static History 0.6 Dense units: 72, Learning rate: 0.000451, # LSTM layers: 1, LSTM units layer 1: 56, Dropout layer 1: 0.6 0.153 (0.150–0.156) 0.112 (0.110–0.115) 0.239 (0.235–0.244)

Performance of LSTM models trained on individual input modalities. Input types included medications, laboratory values, % meals consumed, diet orders, and static patient history. For each model, the best-performing decision threshold (DT), optimized hyperparameters (HPs), and resulting mean F1 score, precision, and recall with 95% confidence intervals are shown.

We compared an expanded input dataset that included the top 140 most administered inpatient medications, 102 most ordered inpatient labs, 61 diet orders, percent meal consumed, and 65 static features to the curated one used up to now that is based on clinical relevance and prior studies (33 medications, 13 labs, 4 diet types, % meal consumed and 65 static features). Both were evaluated using CW or WBCE, with a 5-day LB window and were trained with batch sizes of either 2048 or 4096. The expanded input set, despite including more variables, performed worse overall (STable 4). The best F1 achieved was 0.268 using WBCE and batch size 2048, with a higher recall of 0.459 but lower precision of 0.189 compared to the optimal model trained with the curated input dataset. These findings suggest that expanding the feature set may introduce noise rather than signal, and support the use of a clinically-informed curated input set for predictive modeling of inpatient hypoglycemia.

We illustrate the behavior of the highest F1 achieving LSTM model using a representative test-set patient example (Fig. 2). In this case, the patient experienced one hypoglycemic event on 01/02/2015 (9:07 AM). To capture risk before these events, the model is trained using binary outcome labels that switch to 1 beginning 24 h before each event (Fig. 2a). Model predictions were generated every 4 h starting from 4 h after the first POC BG measurement. As shown in Fig. 2b, the actual POC BG measurements confirm the hypoglycemic event on 01/02/2015. Figure 2c displays the predicted probability of hypoglycemia over time. Using the decision threshold of 0.7 that led to the best F1, the model issues a correct alert starting on 01/01/2015, one day before the first event, flags the event day, and also produces a false alert on 01/06/2015 when the BG decreases rapidly and approaches 70 mg/dL but does not fall below it.

Fig. 2. Model predictions for example test patient.

Fig. 2

a Binary outcome label used for training, where label = 1 is assigned starting 24 h prior to each BG < 70 mg/dL event. This patient experienced 1 hypoglycemic event on 01/02/2015 at 9:07AM, marked by the vertical red line. b Actual POC BG values for this test patient. The horizontal red line marks BG 70 mg/dL. c Model’s predicted probability of BG < 70 mg/dL within the next 24 h. A threshold of 0.7 (horizontal red line) triggers an alert beginning on 01/01/2015, one day before the first event. This figure was generated in Python using Matplotlib.

Model Interpretation

Understanding the clinically relevant factors that contribute to the predicted hypoglycemia risk enhances confidence in the model’s accuracy and its resilience to confounding influences. To enhance clinical interpretability, we used SHapley Additive exPlanations (SHAP)38 with GradientExplainer39, specifically the positional SHAP (PoSHAP) concept to interpret time series34. PoSHAP computes each feature’s contribution to the model prediction at each timestep, quantifying how much each feature increases or decreases hypoglycemia risk relative to a baseline40. Positive SHAP values indicate increased risk, while negative values indicate decreased risk of hypoglycemia. For each 4-hour prediction time, we recorded the top 10 contributing features across all input types. This enabled identification of consistently influential predictors and assessment of their modifiability in clinical workflows. We used the same representative test patient as used in Fig. 2 to examine the most influential predictors of the first hypoglycemia prediction generated on the day preceding the event (Fig. 3). In this test patient, insulin glargine given at timesteps -2, -5, and -8 (where timestep 0 denotes the most recent 4 h interval prior to prediction) were the most influential modifiable contributors to this patient’s increased (positive SHAP values) hypoglycemia risk (Fig. 3a). The test patient’s insulin glargine doses can be visualized in Fig. 3b, showing the doses of 60, 85, and 65 units given at -2, -5, and -8 timesteps prior to the prediction made at the end of timestep 0.

Fig. 3. Top 10 SHAP predictors and insulin dosing pattern preceding hypoglycemia for example test patient.

Fig. 3

a Top 10 predictors for an individual test case derived using SHAP values, which quantify each feature’s contribution to the model’s predicted risk of hypoglycemia. Features are labeled with the timestep relative to prediction time (e.g., “GLUCOSE-POC (-6)” denotes a point-of-care glucose measurement 6 timesteps ago). b Time series plot of insulin glargine doses over the 30-timestep lookback window, where timestep 0 is the most recent prior to the prediction. Red dashed vertical lines indicate the three SHAP-selected timesteps where glargine contributed most to predicted hypoglycemia risk. This figure was generated in Python using Matplotlib.

To evaluate whether the drivers of hypoglycemia risk differed across key demographic subgroups, we performed subgroup-specific SHAP analyses restricted to prediction windows associated with a hypoglycemic event (y = 1). SFig. 3 shows the mean absolute SHAP values at the most recent 4-hour timestep (t0) for the top features across subgroup comparisons defined by sex, race, and age. Across all subgroup comparisons, feature importance profiles were highly concordant, with Spearman rank correlations between subgroup-specific mean |SHAP| vectors ranging from 0.97 to 0.99. This indicates the relative ranking of influential predictors was largely preserved across demographic strata. Core drivers of hypoglycemia risk, including recent BMI, age, comorbid CKD, prior hypoglycemia, and recent POC BG were consistently among the top contributors in all panels. Across all three demographic subgroup analyses, BMI emerged as a consistently influential predictor of hypoglycemia risk. In the sex-stratified analysis, BMI exhibited a significantly larger mean |SHAP| value in males, indicating a stronger contribution to hypoglycemia risk among male patients. Similarly, in the age-stratified analysis, BMI showed a significantly higher attribution in patients older than 75 years. The number of features with significant subgroup differences varied by comparison, ranging from 24 (sex) to 38 (race). Importantly, these differences occurred in the context of highly correlated global attribution patterns, indicating that subgroup-specific variation reflects differences in the strength of individual risk factors.

Daily prospective hypoglycemia prediction evaluation

The daily cohort validation included 1190 unique admissions and a total of 5299 admissions who met inclusion criteria (i.e., adult, receiving an antihyperglycemic in the hospital, and a length of stay at least 24 h) between 6/26/25 and 7/13/25. While the training/validation/test sets (5/14/2014-3/26/2025) were sampled retrospectively to support model development and evaluation, the daily validation cohort aligns with how the model would be deployed operationally in real-time. As shown in STable 1, compared to the model development cohort, the daily validation cohort included a higher proportion of CSMC patients (96.5% vs. 88.1% in the test set), a greater percentage of males (70.8% vs. 58.0%), and higher ICU representation (18.5% vs. 10.3%). It also contained a higher prevalence of individuals with T2DM (83.2% vs. 77.1%) and CKD (65.2% vs. 58.3%) but lower rates of malignancy and liver failure. We evaluated the daily performance of the best-trained LSTM hypoglycemia prediction model across multiple decision thresholds (0.5, 0.6, 0.7, 0.8, 0.9; Fig. 4, SFig. 4). Overall global precision, recall and F1 scores were computed by summing true positives, false positives, and false negatives across all evaluation days and then calculating the corresponding metrics. When the best LSTM model was evaluated daily during this period, it maintained comparable performance to that achieved during training at a 0.7 threshold (F1 score of 0.30), supporting its ability to generalize to daily cohorts and its potential for EHR-integrated deployment. A threshold of 0.5 consistently achieved the highest recall (0.59), but also resulted in the most false positives (479). Conversely, higher thresholds (e.g., 0.8 and 0.9) improved precision (0.36 and 0.38, respectively) but reduced recall and the number of true positives (26 and 3, respectively). True negatives had higher counts at stricter thresholds due to reduced alerting. False negatives increased at higher thresholds, with a steep rise at 0.9 (124), indicating more missed events. SFig. 5 shows daily performance metrics using the calibrated model over the same time period. Notably, the maximum predicted probability after isotonic calibration was 0.6; thus, at thresholds > 0.6, the model made no positive predictions. At lower thresholds (e.g., 0.1-0.3), the model captured more true positives (e.g., 60 at threshold 0.1), resulting in higher recall but lower precision. The highest F1 score (0.29) was achieved at the 0.2 threshold. Calibrated probabilities ensure that predicted probabilities more accurately reflect observed risk, which may support better threshold selection in clinical practice. For example, knowing that a model-predicted risk of 0.5 corresponds to an actual 50% event rate empowers clinicians to act appropriately based on clinical context and competing priorities.

Fig. 4. Daily performance of the best-trained LSTM hypoglycemia prediction model across decision thresholds in real-time simulation.

Fig. 4

Each panel shows model performance metrics calculated daily from 6/26/2025 to 7/14/2025 across five decision thresholds (0.5–0.9), applied to adult inpatients receiving antihyperglycemic medications with a length of stay > 24 h. Cumulative performance values for each threshold are included in the legends. This figure was generated in Python using Matplotlib.

Subgroup fairness evaluation

STable 5 summarizes pairwise comparisons of model F1 scores across demographic and clinical subgroups. No significant performance differences were observed between males and females, Asian and White patients, or Hispanic and Non-Hispanic patients. A small but statistically significant difference was noted between age groups, with the model performing slightly better in patients ≤ 75 years (F1 0.305) compared to those > 75 years (F1 0.294). The model performed better in patients with CKD vs non-CKD (F1 0.310 vs. 0.278), T1DM vs T2DM (F1 0.353 vs. 0.304), and those not in the ICU (F1 0.321 vs. 0.298). Hypoglycemia rates were higher in these populations (i.e., 10% T1DM vs 4.8% T2DM, 5.8% ICU vs. 4.3% non-ICU, and 5.3% CKD vs 3.2% non-CKD; STable 6), suggesting that higher incidence enables the model to learn more informative patterns and generate more true positives, improving F1 score. Similarly, Black patients showed modestly higher F1 scores (0.320 vs. 0.295 for White), which aligns with their higher baseline hypoglycemia rate (5.5% vs. 4.2%; STable 6). Overall, the model demonstrates stronger performance in groups with greater hypoglycemia incidence, which may reflect improved detection performance in higher-risk groups.

Discussion

This work addresses longstanding challenges in inpatient glycemic management, where care remains largely reactive; current practices typically adjust treatment only after a hypoglycemic episode occurs13,15,16, despite clear evidence that many events are predictable. Our model shifts this paradigm by enabling real-time, proactive prediction using EHR data that includes structured clinical variables and POC BG. By continuously synthesizing heterogeneous EHR data streams, the model can surface patients whose risk may not be immediately apparent, enabling earlier, proactive intervention.

The LSTM model that was developed, evaluated, and validated in this study has the potential to prevent an average of 3.5 inpatient hypoglycemia events per day (44% of expected daily cases). At the 0.7 decision threshold, this benefit is achieved with an alert burden of ~15 alerts per day and a positive predictive value (PPV; precision) of 23%, reflecting 3.4 false positives per true alert. A prior ICU hypoglycemia alert study reported a PPV of 14%, yet was still considered clinically actionable given the severity of hypoglycemia and low risk associated with recommendations41.

EHR-integrated hypoglycemia risk tools have the potential to influence clinician behavior, such as increased modification of insulin doses42. Rather than presenting interruptive alerts to all ordering clinicians, we envision integration of this model into a diabetes stewardship workflow, such as a prioritized list reviewed by a pharmacist. Prior studies have used a similar approach for real-world implementation of hypoglycemia prediction tools, including delivering daily emails with high-risk patient lists to nurse practitioners43. The alert threshold can also be tuned for each institution, allowing the hospital to trade sensitivity for alert volume. The broader clinical and financial implications of this are substantial; scaled nationally to ~900,000 hospital beds in the U.S., this would equate to prevention of nearly 4000 inpatient hypoglycemia events daily. These events are associated with increased morbidity, length of stay, and cost7–14, making even partial prevention a meaningful advancement for patient safety and hospital performance metrics.

The model’s LSTM architecture was designed to explicitly capture temporal dependencies across heterogeneous inputs, including medications, lab values, diet orders, meal intake, and static clinical features, over a rolling 5-day window discretized into 4-hour intervals. By preserving the ordering and spacing of events, the model learns short-term trajectories, dose timing patterns, and evolving physiologic trends rather than relying on collapsed representations. In contrast, prior inpatient EHR hypoglycemia studies have relied on non-sequential models. For example, Mathioudakis et al. used SGB26, and Zale et al. used a RF model27. In those studies, precision was 0.09 with recall 0.82 (F1 = 0.16)26 and precision 0.12 with recall 0.72 (F1 = 0.21)27. While recall was high, precision near 0.1 implies a substantial false-alert burden that may limit clinical deployment. In contrast, the LSTM in this study achieved precision 0.23 and recall 0.44 (F1 = 0.30), more than doubling precision while maintaining clinically meaningful sensitivity, corresponding to approximately 3.4 false positives per true alert. This shift toward higher precision directly addresses alert fatigue, a major barrier to clinical adoption44. Importantly, when we implemented representative non-temporal approaches within our dataset including LR and gradient-boosted trees trained on flattened inputs or summary statistics (mean, minimum, maximum, most recent value, SD, IQR), performance consistently declined relative to the LSTM. Reducing longitudinal inputs to descriptive summaries further degraded performance, indicating temporal structure contains predictive information not captured by aggregate statistics. Together, these findings demonstrate that preserving temporal dynamics provides measurable predictive benefit.

Our daily validation cohort, drawn from real-time EHR extracts, demonstrated consistent performance over time, confirming the model’s robustness under operational conditions. Daily evaluation showed that model sensitivity and specificity could be tuned via threshold selection: lower thresholds yielded higher recall with more false positives, while higher thresholds improved precision but missed more events.

We also emphasize interpretability, which is essential for trust and adoption in clinical settings. Using PoSHAP34, we identified the most influential features driving individual predictions at specific timepoints. This allows clinicians to understand why patients are at risk, and whether those risks stem from modifiable factors such as recent insulin administration. Visualizing these contributors supports safer, patient-specific interventions and opens the door to integrating model explanations into clinical workflows. While SHAP-based explanations can help clinicians understand whether the model uses clinically plausible signals to make predictions, evidence that explanations alone change clinician decisions or improve outcomes remains limited45. Prospective studies are needed as the next step in understanding whether explanations alter treatment decisions and patient outcomes45. Additionally, a fairness evaluation, an essential component of equitable healthcare especially for real-world deployment46 but lacking in the currently published ML-based rolling hypoglycemia prediction studies, further demonstrated that model performance was stable across sex and ethnicity. Performance was modestly higher in groups with higher baseline hypoglycemia risk, likely due to more positive training examples. This suggests that differences in performance may reflect both the distribution of events across subgroups and potential limitations in model generalizability. These findings highlight the need for ongoing fairness evaluations and possibly subgroup-specific calibration or training strategies to ensure equitable performance47.

Our study has several limitations. First, despite outperforming prior inpatient EHR-based models and non-sequential baselines, performance remains modest and likely reflects sparse and irregular POC BG sampling, as well as the partially stochastic nature of inpatient hypoglycemia precipitants. Meaningful gains will likely require complementary data such as clinical notes, vital sign trajectories, or nursing observations. Second, development and prospective evaluation cohorts were drawn from a single health system in Los Angeles, with the academic medical center accounting for most of the development data and daily prospective cohort, which limits external generalizability. Third, prospective evaluation spanned only 2.5 weeks; longer surveillance is needed to detect performance drift. Fourth, the model uses intermittent POC BG and may miss transient hypoglycemia between fingersticks; inpatient CGM data, when available routinely in the inpatient setting, may further improve sensitivity. Fifth, time-series missingness was encoded as zero to mirror EHR information availability, but this conflates “not measured” with “not occurring”.

Several directions follow from these limitations. The most important next step is a prospective, ideally randomized, interventional study embedding the model in clinician workflow and measuring clinically meaningful endpoints, including overall and severe (BG < 54 mg/dL) hypoglycemia rates, time in target glycemic range, length of stay, ICU transfers, and 30-day readmissions, alongside operational metrics such as alert response rates, time-to-action, and clinician workload. External multi-site validation across health systems is required before broad deployment and will determine whether site-specific recalibration or fine-tuning is needed.

We developed and validated a real-time, deployable LSTM model that predicts inpatient hypoglycemia within a 24-hour horizon using routinely collected EHR data and sparse POC glucose measurements. By explicitly modeling temporal dynamics across medications, labs, diet, meal intake, and static clinical factors, the model achieved improved discrimination over strong non-sequential baselines while sustaining equivalent performance in daily prospective validation, supporting robustness under operational conditions. Patient-level, time-resolved explanations using PoSHAP further increase clinical transparency by highlighting modifiable drivers such as recent insulin dosing. Together, these results provide a practical foundation for EHR-integrated decision support, ideally embedded within stewardship workflows, to enable proactive interventions, reduce preventable hypoglycemia, and improve inpatient glycemic safety.

Methods

Data sources

We analyzed EHR data from Cedars-Sinai Health System hospitals, including Cedars-Sinai Medical Center (CSMC; 886 beds) and two affiliate community hospitals, Marina Del Rey Hospital (MDRH; 133 beds) and Huntington Health (HH; 619 beds), spanning May 14, 2014 to March 26, 2025. Eligible admissions met the following inclusion criteria: (i) age ≥ 18 years, (ii) length of stay ≥ 24 h, and (iii) receipt of at least one antihyperglycemic medication during the admission. The study was approved by the Cedars-Sinai Institutional Review Board (STUDY00002306, MOD00011744) and the Huntington Hospital Clinical Research Committee (MOD00012907). The requirement for informed consent and HIPAA authorization was waived by the approving ethics committees because the study involved retrospective EHR review with no participant interaction or intervention and was determined to pose no greater than minimal risk to participants.

Model output definition and prediction window construction

Ground truth labels were generated by assigning a label every 4 h starting from 4 h after the first recorded POC BG measurement and continuing until the last POC BG for that admission. Each prediction window was assigned a binary label indicating whether hypoglycemia occurred within the subsequent 24 h (Fig. 5). Specifically, the label was set to 1 if any BG < 70 mg/dL was observed in the 24 h following the prediction time; otherwise the label was set to 0. Because predictions occur repeatedly over the course of an admission, each admission contributes multiple prediction windows, and a single hypoglycemic event can correspond to multiple positive windows. The 24 h prediction window was selected based on Cedars-Sinai endocrinologist consensus as a clinically actionable timeframe for intervention.

Fig. 5. Rolling prediction framework for inpatient hypoglycemia prediction.

Fig. 5

Predictions are generated every 4 h using a sliding window. At each prediction time, the model uses a lookback (LB) window of time-series and static data to assess risk of BG < 70 mg/dL within a prediction horizon (PH). The PH extends 24 h beyond each prediction point to identify whether BG < 70 mg/dL occurs. In this example, the red BG of 60 mg/dL falls within the 4th PH, which is assigned a positive output label. For patients hospitalized fewer than 5 days, the LB window is zero-filled to maintain uniform input dimensions. This figure was created in PowerPoint.

Model input representation and feature engineering

The model uses four time series modalities (33 medications, 13 laboratory values, 4 diet order types, and percent meal consumption) alongside 65 static variables comprised of social, demographic, and historical clinical data (STable 7). Features were selected based on clinical relevance and prior studies identifying inpatient medications and labs associated with BG variation3,48–52.

All time series inputs were constructed over a configurable lookback window (LB) preceding each prediction time and discretized into 4 h bins. Within each 4 h timestep, features were aggregated as follows: medication doses were summed, lab values were averaged, diet type was encoded, and percent meals consumed were totaled. If an input was absent during a given 4 h bin, a value of 0 was inserted for that interval. LB window length was treated as a modeling choice and evaluated empirically.

To assess whether including a broader set of structured features improved performance, we constructed an expanded input set containing the curated features plus additional high-frequency variables (top 140 administered medications, 102 most ordered labs, 61 diet orders, percent meal consumed, and 65 static features). The curated and expanded feature sets were evaluated under matched training settings (lookback window, timestep resolution, and imbalance strategy).

Missing data handling

Missingness was handled to reflect how information naturally occurs in the EHR, with separate approaches for time-series and static features. Time-series inputs were aggregated into 4 h intervals, and if a feature was not recorded within a given interval, a value of zero was used to indicate absence of observation during that window. No model-based imputation was performed for time-series features. Imputation was applied only to static variables, all of which had < 20% missingness. Continuous static features (e.g., BMI) were imputed using scikit-learn’s IterativeImputer, fitted exclusively on the training set and applied unchanged to validation, test, and prospective cohorts to prevent data leakage. Categorical static variables were imputed by sampling from the empirical distribution of observed categories in the training data.

Model architecture, training and optimization

Data was split by hospital admission to reflect real-world deployment, in which models are trained on historical admissions and applied to future admissions, including readmissions of patients previously seen by the system. Data was split into training (64%), validation (16%), and test (20%) sets such that no admission appeared in more than one split. For each prediction timepoint, features were constructed using only data available up to that timepoint. As a sensitivity analysis, we additionally evaluated performance under a patient-level split. Under this setting, all admissions from a patient were assigned to a single split.

A multi-branch deep learning (DL) architecture was implemented in TensorFlow/Keras to predict inpatient hypoglycemia within a 24-hour prediction horizon. Each time series input was passed through a dedicated branch composed of up to three bidirectional LSTM layers with dropout and batch normalization applied to each recurrent layer. The number of LSTM layers, units per layer, and dropout rates were all treated as tunable hyperparameters. Static inputs were processed through a dense ReLU-activated layer with dropout. Outputs from all branches were concatenated, followed by an additional dense layer with ReLU activation and dropout, and a final sigmoid-activated output layer for binary classification.

Model training was conducted on a high-performance computing cluster equipped with NVIDIA A100 80GB GPUs. Predictions simulate a real-time system that updates every 4 h as new clinical data becomes available. Once fit, the model generates each prediction in 0.07 s using only CPU resources, making frequent updates computationally feasible.

To address class imbalance given the low incidence of inpatient hypoglycemia, we applied two strategies during training: (1) class weighting (CW) to increase the penalty for misclassified positive cases53, and (2) a custom weighted binary cross-entropy (WBCE) loss function to assign different penalties to positive versus negative classes during backpropagation. Class weights were empirically set to 11.3 for positive cases and 0.5 for negative cases based on the inverse frequency in the training set.

Model hyperparameters, including the number of LSTM layers, LSTM units per layer, dropout rates, dense layer size, and learning rate, were optimized using Keras Tuner’s Bayesian Optimization54 to maximize F1 score on the validation set. All models were trained for up to 10 epochs with early stopping (patience = 3), using 10 hyperparameter tuning trials, and evaluated on the held-out test set. The final model’s output probabilities were thresholded at multiple decision thresholds (DT, 0.5–0.9) to explore trade-offs between precision and recall. For each model, the DT that maximized F1 on the validation set was selected and the test set metrics at that DT were reported. Bootstrap resampling of test prediction windows (100 iterations with replacement, each comprising 50% of the test set) were used to estimate 95% CIs.

To compare the performance of the best LSTM model to that of baseline models, we trained a dense neural network, LR and XGBoost on temporally flattened inputs. Time-series features (medications, lab values, diet orders, and meal intake) were reshaped from 3D (samples × timesteps × features) to 2D representations by concatenating all timepoints along the feature axis, then combining them with static features to create a single fixed-length input per prediction window. For the dense neural network, this input was passed through a feedforward network with two ReLU-activated dense layers. Hyperparameters, including layer size, dropout rate, and learning rate were optimized using Bayesian Optimization. LR and XGBoost models were trained using CW to address class imbalance, and their respective hyperparameters were tuned independently. Performance was evaluated on the held-out test set using 100 bootstrap iterations, with 95% CIs calculated for precision, recall, and F1. In addition to thresholded metrics, AUPRC was computed on the held-out test set to quantify threshold-free discrimination under class imbalance. To compare the LSTM model performance with each baseline, we applied a two-sided Mann-Whitney U test to the bootstrapped F1 distributions at the selected DT. To evaluate whether performance of the three baseline models depended on how data were represented, we additionally computed a set of summary statistics, including mean, minimum, maximum, most recent value, standard deviation (SD), range, and interquartile range (IQR), for each time-series variables over the 5-day LB window. These summaries were then combined with the static features to form a single fixed-length input and used to train and optimize LR, XGBoost, and two-layer dense neural network models, which were evaluated on the same held-out test set using the identical bootstrap procedure.

Calibration

To evaluate the calibration of predicted probabilities from the LSTM model, we first divided the held-out 20% test set into equal-sized validation and final test subsets using stratified random sampling to preserve outcome distribution. On the validation subset, predicted probabilities were generated using the uncalibrated model and then used to fit an isotonic regression model. This mapping was trained to align predicted probabilities with actual outcome frequencies and was then applied to the final test set predictions to produce calibrated probabilities.

Model interpretation

We used SHapley Additive exPlanations (SHAP)38 to interpret model outputs at the patient level. GradientExplainer was initialized with a reference background of 5000 samples drawn from the validation set. For each test set prediction window, we passed individualized input features (time series of labs, meds, diets, meals, and static characteristics) to the explainer to compute SHAP values. These values were flattened across time and ranked to identify the top clinical features and the timestep at which they occurred driving each prediction, allowing clinicians to understand and act on the factors and timing of these factors contributing to a patient’s hypoglycemia risk.

To evaluate whether model feature importances differed across demographic subgroups, including sex, race, and age, we performed subgroup SHAP analyses using a previously constructed SHAP explainer among the previously held-out test samples with a hypoglycemic event within a 24 h prediction horizon. SHAP values were computed at timestep 0, the most recent 4 hour interval prior to the hypoglycemic event, corresponding to the clinical state at the time of prediction. For each pairwise subgroup comparison, balanced cohorts were then constructed by randomly sampling an equal number of samples from each group, with the sample size per group determined by the smaller available subgroup. For each prediction window, SHAP values at timestep 0 were extracted for all input features, yielding one SHAP value per feature per sample. For each subgroup, feature importance was summarized as the mean absolute SHAP value (mean ∣SHAP∣) per feature across all sampled windows in that subgroup. Overall concordance between subgroup attribution profiles was quantified using the Spearman rank correlation between the subgroup-specific mean ∣SHAP∣ vectors across all features.

To identify features with statistically different attribution magnitudes between subgroups, we compared the per-feature ∣SHAP∣ distributions between groups using two-sided Mann-Whitney U tests. P-values were adjusted for multiple comparisons across features using the Benjamini-Hochberg false discovery rate (FDR) procedure, with FDR < 0.05 considered statistically significant. For visualization, each panel displayed the top 12 features ranked by pooled importance (sum of subgroup mean ∣SHAP∣ values) and annotated features that were FDR-significant with significance stars (*FDR < 0.05; **FDR < 0.01; ***FDR < 0.001).

Daily prospective hypoglycemia prediction evaluation

To evaluate model performance prospectively, we applied our best-performing LSTM model to daily EHR extracts including patients admitted to the hospital that day who met the same inclusion criteria that was used for model development (i.e., all inpatients ≥ 18 years old with a length of stay ≥ 24 h who received at least one anti-hyperglycemic medication) from 6/26/2025 to 7/13/2025. For each date, labs, medication administrations, diet orders, meal intake, and static history were extracted. The model was loaded from a saved H5 file and applied to data from each day using a consistent 5-day lookback window to generate risk probabilities for hypoglycemia within the subsequent 24 h window.

Predictions were matched to POC BG values collected over the 24 h period post-prediction for each patient. Each day’s predictions were evaluated using multiple thresholds (0.5–0.9) to compute F1 score, precision, recall, and confusion matrix components for each day. Overall performance at each threshold was calculated by aggregating true positives, false positives, and false negatives across all evaluation days before computing the corresponding metrics. Model inputs were preprocessed using trained imputers and encoders from the model training phase.

Subgroup fairness evaluation

To evaluate model fairness across demographic and clinical subgroups, we defined binary masks for each subgroup using the test static input matrix. These included age ( ≤ 75 vs. > 75), sex, comorbidities (CKD, T1DM, T2DM), insurance type, race, ethnicity, and ICU status. Using these masks, we computed F1 scores separately for each subgroup by comparing the model’s predicted labels with the ground truth labels. To assess whether differences in model performance between subgroups were statistically significant, we implemented a bootstrapping procedure. For each predefined subgroup pair, we resampled the F1 scores 1000 times with replacement from the respective subgroup indices. We calculated the mean F1 difference and corresponding 95% CI from the bootstrapped distribution. Statistical significance was determined by whether the 95% CI excluded zero.

Supplementary information

Acknowledgements

We thank the Enterprise Data Intelligence team at Cedars-Sinai Medical Center, including Kevin Japardi, for assistance with data extraction. We also thank Alexandre Hutton for assistance with the use of high-performance computing resources, and Dr. Omar Ibrahim for assistance in obtaining Huntington Health EHR data for this study. This work was supported by the NIGMS R35GM142502 and NIH National Center for Advancing Translational Science (NCATS) UCLA CTSI Grant Number UL1TR001881.

Author contributions

A.M. conceived the study, performed the data analysis, generated the figures, interpreted the results, and drafted the manuscript. C.C. assisted with data processing. D.C., E.V.N., and R.G. provided clinical expertise and contributed to interpretation of the results. J.G.M. supervised the project, conceived the study, contributed to study design and methodology, and interpreted the results. All authors reviewed and approved the manuscript.

Data availability

The data used in this study were derived from the EHR systems of Cedars-Sinai Medical Center and affiliated hospitals (Marina Del Rey Hospital and Huntington Health) and include protected health information. Due to institutional policies and patient privacy regulations, the individual-level EHR data cannot be publicly shared. This study was reported in accordance with the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis + Artificial Intelligence (TRIPOD + AI) statement55; the completed checklist is provided in STable 8.

Code availability

The model development, training, evaluation, and interpretation code is publicly available at: https://github.com/xomicsdatascience/Hypoglycemia-DL. All protected health information and institution-specific infrastructure components have been removed.

Competing interests

Cedars-Sinai Medical Center has filed patent applications related to the hypoglycemia prediction model described in this manuscript. Jesse Meyer, Amanda Momenzadeh, and Caleb Cranney are listed as inventors on U.S. Provisional Application No. 63/675,784, U.S. Provisional Application No. 63/688,658, and PCT Application No. PCT/US2025/039346. The patent status is pending.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Amanda Momenzadeh, Email: Amanda.Momenzadeh@cshs.org.

Jesse G. Meyer, Email: Jesse.Meyer@cshs.org

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s41746-026-02874-1.

References

  • 1.American Diabetes Association 15. Diabetes care in the hospital: Standards of Medical Care in Diabetes—2019. Diab. Care 42, S173–S181 (2019). [DOI] [PubMed]
  • 2.Cryer, P. E. et al. Evaluation and management of adult hypoglycemic disorders: an endocrine society clinical practice guideline. J. Clin. Endocrinol. Metab.94, 709–728 (2009). [DOI] [PubMed] [Google Scholar]
  • 3.Dhatariya, K., Corsino, L. & Umpierrez, G. E. Management of Diabetes and Hyperglycemia in Hospitalized Patients. in Endotext (eds Feingold, K. R. et al.) (MDText.com, Inc., South Dartmouth (MA), 2000).
  • 4.American, D. et al. 16. Diabetes care in the hospital. Stand. Care Diab.—2024. Diab. Care47, S295–S306 (2024). [Google Scholar]
  • 5.Brodovicz, K. G. et al. Association between hypoglycemia and inpatient mortality and length of hospital stay in hospitalized, insulin-treated patients. Curr. Med. Res. Opin.29, 101–107 (2013). [DOI] [PubMed] [Google Scholar]
  • 6.Cook, C. B. et al. Inpatient glucose control: a glycemic survey of 126 U.S. hospitals. J. Hosp. Med. 4, E7-E14 (2009). [DOI] [PubMed]
  • 7.Gold, A. E. & Marshall, S. M. Cortical blindness and cerebral infarction associated with severe hypoglycemia. Diab. Care19, 1001–1003 (1996). [DOI] [PubMed] [Google Scholar]
  • 8.Desouza, C., Salazar, H., Cheong, B., Murgo, J. & Fonseca, V. Association of hypoglycemia and cardiac ischemia. Diab. Care26, 1485–1489 (2003). [DOI] [PubMed] [Google Scholar]
  • 9.Wei, M. et al. Low fasting plasma glucose level as a predictor of cardiovascular disease and all-cause mortality. Circulation101, 2047–2052 (2000). [DOI] [PubMed] [Google Scholar]
  • 10.National Diabetes In-patient Audit-Harms. NHS Digitalhttps://digital.nhs.uk/data-and-information/clinical-audits-and-registries/national-diabetes-in-patient-audit-nadia-harms.
  • 11.Amiel, S. A. et al. Hypoglycaemia, cardiovascular disease, and mortality in diabetes: epidemiology, pathogenesis, and management. Lancet Diab. Endocrinol.7, 385–396 (2019). [DOI] [PubMed] [Google Scholar]
  • 12.Petersen, K.-G., Schlüter, K. J. & Kerp, L. Regulation of serum potassium during insulin-induced hypoglycemia. Diabetes31, 615–617 (1982). [DOI] [PubMed] [Google Scholar]
  • 13.Lake, A. et al. The effect of hypoglycaemia during hospital admission on health-related outcomes for people with diabetes: a systematic review and meta-analysis. Diabet. Med. J. Br. Diabet. Assoc.36, 1349–1359 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Investigators, The NICE-. Hypoglycemia and risk of death in critically ill patients. N. Engl. J. Med.367, 1108–1118 (2012). [DOI] [PubMed] [Google Scholar]
  • 15.American Diabetes Association Professional Practice Committee 16. Diabetes Care in the Hospital: Standards of Medical Care in Diabetes—2022. Diab. Care 45, S244–S253 (2022).. [DOI] [PubMed]
  • 16.Witte, H., Nakas, C., Bally, L. & Leichtle, A. B. Machine learning prediction of hypoglycemia and hyperglycemia from electronic health records: algorithm development and validation. JMIR Form. Res.6, e36176 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.De La Cruz, M., Garnica, O., Cervigon, C., Velasco, J. M. & Hidalgo, J. I. Explainable hypoglycemia prediction models through dynamic structured grammatical evolution. Sci. Rep.14, 12591 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Duckworth, C. et al. Explainable machine learning for real-time hypoglycemia and hyperglycemia prediction and personalized control recommendations. J. Diab. Sci. Technol.18, 113–123 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Mujahid, O., Contreras, I. & Vehi, J. Machine learning techniques for hypoglycemia prediction: trends and challenges. Sensors21, 546 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Agraz, M., Deng, Y., Karniadakis, G. E. & Mantzoros, C. S. Enhancing severe hypoglycemia prediction in type 2 diabetes mellitus through multi-view co-training machine learning model for imbalanced dataset. Sci. Rep.14, 22741 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Shi, M. et al. A novel electronic health record-based, machine-learning model to predict severe hypoglycemia leading to hospitalizations in older adults with diabetes: a territory-wide cohort and modeling study. PLOS Med21, e1004369 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mantena, S. et al. Predicting hypoglycemia in critically Ill patients using machine learning and electronic health records. J. Clin. Monit. Comput.36, 1297–1303 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Ruan, Y. et al. Predicting the risk of inpatient hypoglycemia with machine learning using electronic health records. Diab. Care43, 1504–1511 (2020). [DOI] [PubMed] [Google Scholar]
  • 24.Saito, T. & Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE10, e0118432 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Davis, J. & Goadrich, M. The relationship between Precision-Recall and ROC curves. in Proceedings of the 23rd international conference on Machine learning - ICML ’06 233–240 (ACM Press, Pittsburgh, Pennsylvania, 2006). 10.1145/1143844.1143874. [DOI]
  • 26.Mathioudakis, N. N. et al. Development and validation of a machine learning model to predict near-term risk of iatrogenic hypoglycemia in hospitalized patients. JAMA Netw. Open4, e2030913 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Zale, A. D., Abusamaan, M. S., McGready, J. & Mathioudakis, N. Development and validation of a machine learning model for classification of next glucose measurement in hospitalized patients. EClinicalMedicine44, 101290 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Rumelhart, D. E., Hinton, G. E. & Williams, R. J. Learning Internal Representations by Error Propagation. in Readings in Cognitive Science 399–421 (Elsevier, 1988). 10.1016/B978-1-4832-1446-7.50035-2. [DOI]
  • 29.What is a Neural Network? - Artificial Neural Network Explained - AWS. Amazon Web Services, Inc. https://aws.amazon.com/what-is/neural-network/ (2023).
  • 30.Lipton, Z. C., Berkowitz, J. & Elkan, C. A critical review of recurrent neural networks for sequence learning. Preprint at http://arxiv.org/abs/1506.00019 (2015).
  • 31.Bian, Q., As’arry, A., Cong, X., Rezali, K. A. B. M. & Raja Ahmad, R. M. K. B. A hybrid Transformer-LSTM model apply to glucose prediction. PLOS ONE19, e0310084 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Dave, D. et al. Feature-based machine learning model for real-time hypoglycemia prediction. J. Diab. Sci. Technol.15, 842–855 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Dassau, E. et al. Real-time hypoglycemia prediction suite using continuous glucose monitoring. Diab. Care33, 1249–1254 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Dickinson, Q. & Meyer, J. G. Positional SHAP (PoSHAP) for Interpretation of machine learning models trained from biological sequences. PLOS Comput. Biol.18, e1009736 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hort, M., Chen, Z., Zhang, J. M., Harman, M. & Sarro, F. Bias mitigation for machine learning classifiers: a comprehensive survey. ACM J. Responsible Comput.1, 1–52 (2024). [Google Scholar]
  • 36.Kamiran, F., Karim, A. & Zhang, X. Decision Theory for Discrimination-Aware Classification. in 2012 IEEE 12th International Conference on Data Mining 924–929 (IEEE, Brussels, Belgium, 2012). 10.1109/ICDM.2012.45. [DOI]
  • 37.Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. & Galstyan, A. A survey on bias and fairness in machine learning. ACM Comput. Surv.54, 1–35 (2022). [Google Scholar]
  • 38.Ribeiro, M. T., Singh, S. & Guestrin, C. ‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 1135–1144 (ACM, San Francisco California USA, 2016). 10.1145/2939672.2939778. [DOI]
  • 39.shap.DeepExplainer — SHAP latest documentation. https://shap-lrjball.readthedocs.io/en/latest/generated/shap.DeepExplainer.html.
  • 40.Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. in Proceedings of the 31st International Conference on Neural Information Processing Systems 4768–4777 (Curran Associates Inc., Red Hook, NY, USA, 2017).
  • 41.Horton, W. B. et al. Accuracy of a risk alert threshold for ICU hypoglycemia: retrospective analysis of alert performance and association with clinical deterioration events. Crit. Care Med.51, 136–140 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Fenske, M., Brant, A., Willner, M., Eastman, K. & Ketz, J. Use of an electronic medical record integrated tool to identify patients at highest risk of hypoglycemia. AACE Endocrinol. Diab.12, 355–361 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Fralick, M. et al. Using real-time machine learning to prevent in-hospital hypoglycemia: a prospective study. Intern. Emerg. Med.18, 325–328 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Khairat, S., Marc, D., Crosby, W. & Al Sanousi, A. Reasons for physicians not adopting clinical decision support systems: critical analysis. JMIR Med. Inform.6, e24 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Abbas, Q., Jeong, W. & Lee, S. W. Explainable AI in clinical decision support systems: a meta-analysis of methods, applications, and usability challenges. Healthc. Basel Switz.13, 2154 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science366, 447–453 (2019). [DOI] [PubMed] [Google Scholar]
  • 47.Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G. & Chin, M. H. Ensuring fairness in machine learning to advance health equity. Ann. Intern. Med.169, 866–872 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Pratiwi, C., Mokoagow, M. I., Made Kshanti, I. A. & Soewondo, P. The risk factors of inpatient hypoglycemia: a systematic review. Heliyon6, e03913 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Carey, M., Boucai, L. & Zonszein, J. Impact of hypoglycemia in hospitalized patients. Curr. Diab Rep.13, 107–113 (2013). [DOI] [PubMed] [Google Scholar]
  • 50.Rubin, D. J. & Golden, S. H. Hypoglycemia in non-critically ill, hospitalized patients with diabetes: evaluation, prevention, and management. Hosp. Pract.41, 109–116 (2013). [DOI] [PubMed] [Google Scholar]
  • 51.Maynard, G. A., Huynh, M. P. & Renvall, M. Iatrogenic inpatient hypoglycemia: risk factors, treatment, and prevention. Diab. Spectr.21, 241–247 (2008). [Google Scholar]
  • 52.Hulkower, R. D., Pollack, R. M. & Zonszein, J. Understanding hypoglycemia in hospitalized patients. Diab. Manag. Lond. Engl.4, 165–176 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Sivarajkumar, S., Huang, Y. & Wang, Y. Fair patient model: mitigating bias in the patient representation learned from the electronic health records. J. Biomed. Inform.148, 104544 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Keras: Deep learning for humans. https://keras.io/.
  • 55.Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ385, e078378 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The data used in this study were derived from the EHR systems of Cedars-Sinai Medical Center and affiliated hospitals (Marina Del Rey Hospital and Huntington Health) and include protected health information. Due to institutional policies and patient privacy regulations, the individual-level EHR data cannot be publicly shared. This study was reported in accordance with the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis + Artificial Intelligence (TRIPOD + AI) statement55; the completed checklist is provided in STable 8.

The model development, training, evaluation, and interpretation code is publicly available at: https://github.com/xomicsdatascience/Hypoglycemia-DL. All protected health information and institution-specific infrastructure components have been removed.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES