Abstract
Objective:
Determine the efficacy of commonly used approaches to handling missing and/or imbalanced Electronic Health Record (EHR) data on the performance of predictive models targeting risk of admission, intensive care unit (ICU) use, or prolonged length of stay (PLOS) among presenting febrile pediatric emergency department (ED) patients.
Materials and methods:
Historical ED EHR data was used to train a series of XGBoost (XGB) and logistic regression (LR) classifiers. Data handling strategies included imputation methods (multiple imputation (MI), median imputation, complete case (CC) analysis), and imbalanced data corrections (minority oversampling, stratified sub-group analysis). Model performance was evaluated using discriminative (AUC, AUPRC) and calibration metrics (Brier score, Z-scores, p-values).
Results:
Among the study population, 34 % were admitted, 2 % utilized the ICU, and 7 % had a PLOS. Significant data missingness was observed and determined to be not at random (MNAR). In predicting admissions using data recorded within the first two hours of presentation, LR trained using full cohort with median imputation was comparable to MI yielding well-calibrated admissions models with an AUC/AUPRC of 0.82/0.73 while CC analysis yielded an AUC/AUPRC of 0.76/0.78. XGB, trained with unimputed data, produced a well-calibrated admissions classifier with an AUC/AUPRC of 0.85/0.78. In contrast, imbalanced data correction techniques, including synthetic minority oversampling (SMOTE), risk stratification, or the use of XGB did not significantly improve the poor AUPRC and calibration performance of LR models predicting ICU and PLOS.
Conclusion:
Both XGB and LR with median imputation demonstrated robust performance in predicting admissions in the presence of missing data. However, deriving clinically useful models for rare outcomes, such as ICU use or PLOS, remains a challenge due to poor precision/recall and calibration performance. Further research is needed to improve the prediction of rare outcomes in this population.
Keywords: Pediatric emergency medicine, Febrile, Machine learning, Imputation, Imbalanced data
1. Background and significance
Supervised machine learning (ML) is increasingly being applied in the field of pediatric emergency medicine to derive classifiers that predict admission, intensive care unit (ICU) use, and/or prolonged length of stay (PLOS) [1,2]. In addition to the potential utility of such models in improving emergency department (ED) workflow, such predictions may help identify patients that may have been under-triaged [3] or not recognized as having a serious illness in need of time-sensitive care [4]. Children with fever, specifically, account for up to 20 % of pediatric ED visits [5] with underlying disorders ranging from mild, self-limited viral infections [6] to life-threatening bacterial illnesses. Timely recognition of febrile patients at risk for a rapidly progressive, serious bacterial infection in need of hospitalization from other non-seriously ill patients remains a challenge in pediatric emergency medicine [7–9].
In supervised ML a classifier model commonly used in emergency medicine is trained to predict a binary outcome variable commonly called a “label” (e.g. admission, mortality, risk of deterioration) as a function of covariates (e.g. age, patient history, presenting vitals, labs, administered treatments) that can be either ordinal, continuous or binary, and are also referred to as “features”. Classifier performance is evaluated using metrics such as area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPRC) and calibration measures (e.g. Brier score). The quality of the training data significantly affects classifier performance with factors such as non-overlapping class distributions, the inclusion of relevant features to distinguish target classes, and manageable levels of missing data playing key roles.
Logistic regression (LR) is a favored statistical machine learning (ML) tool in clinical applications due to its transparency, ease of use, and ability for clinicians to directly interpret its regression coefficients [10,11]. However, LR requires complete data and thus relies heavily on appropriate imputation or exclusion strategies to maintain model performance [12]. While the literature on handling missing data is well-developed including approaches such as multiple imputation with chained equations (MICE) [13], univariate imputation (mean, median, mode) [14], reduced datasets using complete case (CC) analysis [15], or the use of advanced ML algorithms like extreme gradient boosting (XGB) that inherently handle missing data [16], the relative impact of these techniques on predictive classifier performance is unclear [17–19]. Additionally XGB has demonstrated superior predictive performance compared to LR in medical studies [20], crucial for making accurate diagnoses or risk assessments in precision medicine.
Another major challenge in ML is addressing data imbalance, which arises when predicting relatively rare high-risk conditions in a study population, resulting in a disproportionate number of negative to positive labels in the training data. Classifiers trained on imbalanced data often prioritize maximizing overall accuracy, resulting in poor recall (sensitivity) and precision (also referred to as positive predictive value PPV) for the minority class [21]. Developing predictors of rare events with high recall and precision using imbalance training data is a critical ML challenge. While synthetic minority over-sampling techniques (SMOTE) have been widely used to address class imbalance, recent reports have questioned their clinical validity [22]. Moreover, although XGB, either alone [23], or combined with SMOTE [24] has been reported to effectively handle class imbalance, results have been mixed [25]. Fig. 1 (flowchart) provides a summary of our patient selection criteria, pathway analysis, and key findings.
Fig. 1.

Patient Selection, Pathway Analysis, and Key Findings Flowchart.
2. Objective
The primary objective of this study is to compare methods for addressing missing/imbalanced EHR data in machine learning designed to predict pediatric ED outcomes.
3. Materials and methods
3.1. Data Sources
ML predictive models were developed using anonymized Emergency Department (ED) electronic health record (EHR) data from Children’s National Hospital (CNH), encompassing patient visits to the main hospital and a satellite facility between January 2017 and December 2020. Data were de-identified according to the “Safe Harbor” standard of the HIPAA Privacy Rule (Section 164.514 [a]). Under this standard, beyond removal of all direct identifiers, all date/time tags (admission, discharge, etc.) were randomly shifted so that the relative timing between events (e.g. observations recorded within 2 h of admission) was retained without disclosure of any actual time tag. The project was acknowledged by the CNH Institutional Review Board as not constituting human subjects’ research.
3.2. Study population and feature Generation
From an initial database of 427,173 encounters, patients < 19 years of age were included in the febrile cohort (Cohort-1) based on the following criteria:
a chief complaint or provisional triage diagnoses that contained the term “fever” or a recorded presenting temperature >= 38C
length-of-stay (LOS) in the ED was >= 2 h
lab tests were ordered, and key vital signs recorded within the first two hours of arrival
an ICD-coded problem list enabling the identification of complex patients with chronic conditions
an encounter disposition as either an out-patient, in-patient (admitted), or critical care (ICU) patient
a recorded discharge date enabling the calculation of PLOS defined as > 95th percentile of LOS of all febrile patients that were admitted.
Features recorded within the first two hours of arrival were categorized as follows:
Demographics and contextual variables: age, gender and race (white, black, Latino modeled as binary variables), emergency severity index (ESI) assigned at triage, and complex patient status score (CPS) based on the Pediatric Complex Chronic Condition System [26]
Empiric treatments: Binary indicators of IV antibiotics, ≥2 fluid boluses (20 mL/kg of isotonic crystalloid fluid, delivered over 5–20 min), culture tests, and administration of critical care drugs (vasopressors, hydrocortisone (stress dose), 3 % NaCl, sodium bicarbonate).
Vitals and labs: Continuous values recorded within the first two hours of arrival
Missing data flags: binary indicators for missing values [27]
Outcomes: Binary indicators of admissions, ICU utilization and PLOS
Beyond Cohort-1 we identified two CC reduced data cohorts: Cohort-2: patients with non-missing data when highly missing vitals and labs are dropped as features; Cohort-3: patients with non-missing data when highly missing vitals are dropped.
3.3. Machine learning
For ML we employed LR with least absolute shrinkage and selection operator (LASSO) regularization [28] and the XGBoost (XGB) [29] algorithm for gradient boosting modeling. Both algorithms can generate feature importance plots and related quantification: coefficients/odds ratios in the case of LR, and Shapely Additive Explanations (SHAP) in the case of XGB. The quantification of feature importance available from both LR and XGB algorithms is particularly beneficial in scenarios where interpretability is as important as prediction accuracy, as it helps in identifying the most significant predictors.
3.4. Feature engineering
It is well known that highly correlated features can affect ML algorithms [30]. While it has been reported that both regularized regression [31] and gradient boosting [32] predictive algorithms can effectively deal with correlated features from an accuracy perspective, multicollinearity may hinder model interpretability [32]. To address this issue, once early encounter available features were identified, we performed a covariance analysis to identify highly correlated features and used Principal Component Analysis (PCA) to replace highly correlated features with the minimum number of components required to explain 95 % of the variance associated with correlated features for predictive modeling [33].
In addition to addressing multicollinearity through covariance analysis and PCA, feature engineering included harmonizing semantically equivalent features. For example, lab test or vitals values recorded with different units (e.g., mmol/L vs. mg/dL) or different names (e.g. “Sodium” vs “Whole Blood Sodium”, “Oxygen Sat” vs “SpO2”) were standardized to ensure consistency across the dataset. Outlier filtering was applied to identify and remove physiologically implausible values (e.g., negative vital sign readings, oxygen saturation > 100 %, or blood pressures outside human survival thresholds). Clinician input and published guidelines [34] were used to define acceptable ranges for each feature.
3.5. Statistics and predictive modeling
In the analysis of group comparative statistics, the mean, median and quartiles for continuous features were calculated, and the Mann-Whitney U test was employed to determine the statistical significance of feature distribution differences between outcome groups, providing p-values for each comparison.
We then compared the discriminative and calibration performance:
LR models to predict admission outcome using full cohort with median vs MICE imputation vs CC subsets with no imputation
XGB models to predict admission outcome using full cohort with missingness vs CC subsets [35]
LR and XGB models using SMOTE to predict severely imbalanced outcomes (ICU, PLOS)
LR and XGB models predicting all outcomes using risk stratification to improve data balance
3.6. Parsimonious models
It is well established that complex models risk fitting noise in the data and that parsimonious models tend to be more precise and generalizable [36]. Guided by feature importance plots, regression coefficients and SHAP [37] values 5-,6-, and 7-element feature sets were used to train parsimonious admission prediction models.
3.7. Evaluation and Optimized model performance
Ten-fold cross-validation was used to evaluate model performance, a method commonly employed to ensure robust generalization across datasets [38]. Metrics such as the area under the receiver operating characteristic curve (AUC) and area under the precision-recall curve (AUPRC) were averaged across the folds. To quantify uncertainty in these metrics, 95 % confidence intervals for these metrics were computed using bootstrapping with 1,000 resamples.
Receiver operating characteristic (ROC) curve analysis to identify optimal probability thresholds for comparative model analysis [39]. In this study, predictive models were developed to identify high-risk clinical outcomes, including hospital admission, ICU use, and extended length of stay (LOS). These outcomes are associated with highly imbalanced datasets, where the minority class (e.g., ICU use) represents only a small fraction of cases. To account for this imbalance and align with clinical priorities, we used a sensitivity threshold of 0.8 to evaluate model performance for all models, given that all outcomes are associated with high-risk. Prioritizing sensitivity ensures that most high-risk patients are identified, minimizing the risk of false negatives, which could lead to delayed or inadequate care. While this approach may increase false positives, it aligns with the clinical principle of erring on the side of caution when managing critical outcomes such as ICU admission or extended LOS. A sensitivity threshold of 0.8 was selected as a pragmatic balance between clinical utility and model feasibility, ensuring high-risk patients are identified while maintaining acceptable specificity to avoid overburdening healthcare resources [40,41]. Precision-recall curves were used to illustrate the sensitivity/PPV tradeoffs, particularly useful in class imbalance situations [42].
The Brier score, also known as the Brier loss or the quadratic scoring rule, was used as a measure of calibration performance [43]. Beyond its value in model comparisons, reporting on calibration performance is recommended by the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis) guidelines for prediction modeling studies [35]. However, while the Brier score measures the accuracy of probabilistic predictions, with lower values indicating better accuracy, it has limitations in the context of imbalanced data. If a classifier consistently predicts the majority class with high probability, it can achieve a low Brier score even if the model is poorly calibrated for the minority class. Class probability estimates obtained via supervised learning in imbalanced scenarios systematically underestimate the probabilities for minority class [44] resulting in poor calibration.
Calibration performance was further assessed [45] using the Spiegelhalter Z-test, which evaluates the agreement between predicted probabilities and observed outcomes. This test calculates a Z-score based on the difference between observed outcomes and predicted probabilities, normalized by the variance of the predicted probabilities. The corresponding p-value indicates whether deviations from perfect calibration are statistically significant, with a low p-value signifying a significant deviation. Together, the Brier score and Spiegelhalter Z-test, with associated p-values, provided complementary insights into the reliability of probabilistic predictions, ensuring a robust evaluation of model calibration.
To address calibration challenges and ensure that the predicted probabilities reflected the true likelihood of the outcome, we explored the use of two well-known recalibration techniques: Platt scaling and isotonic regression to adjust the predicted probabilities of a model so that they better reflect the observed outcomes in the data. Recalibration was performed by randomly splitting the classifier training data into two subsets: a main training set (80 %) for initial model fitting and a calibration set (20 %) for model recalibration. The choice between Platt scaling and isotonic regression was determined based on calibration performance metrics and visual inspection of calibration curves on the calibration set.
3.8. Imbalanced data oversampling
To address the inherent imbalance in the datasets used in this study, where minority class outcomes (e.g., admissions, ICU use, extended LOS) constitute only a fraction of all presenting ED cases, the Synthetic Minority Oversampling Technique (SMOTE) [46,47] was explored as a method to enhance minority class representation in the training set. The impact of SMOTE on model performance was evaluated by comparing AUC, AUPRC, and calibration metrics between models trained with and without oversampling.
3.9. Risk stratification for improved data balance
Beyond examining the potential benefits of SMOTE and XGB in dealing with this difficult problem, we also examined the utility of risk-stratifying the study population towards improving the balance of training data. Specifically, we explored the impact of developing models using febrile patients with complex chronic conditions and suspected infection (defined as having received an IV antibiotic and/or culture test within the two-hour window) as well as replacing a very rare outcome such as prolonged length of stay defined as exceeding the 95th percentile with cases where LOS was 24 h or more as the study population towards a potentially clinically useful classifier with improved AUPRC performance resulting from improved data balance.
3.10. Computing environment
All ML model development and statistical analysis was performed using a Jupyter 3.2.1 SPARK notebook environment that accessed the SKLEARN, LIBLINEAR and XGBoost [35] libraries.
4. Results
4.1. Study cohort and clinical features
Fig. 1 shows the overall process flow used to identify the 12,649 patients that met the febrile patient inclusion criteria, and the early encounter associated features described above and shown in Table 1. As detailed in Supplement 1 patient problem list diagnosis codes were matched with CCC codes to establish the complex patient score (CP score) feature. Supplement 2 describes the derivation of empiric treatment features (culture test, antibiotics, bolus, critical care drugs).
Table 1.
Feature sets for machine learning (ML).
| Feature Type | FS-1 (55 features) | FS-2 (34 features, 5445 complete cases) | FS-3 (48 features, 1406 complete cases) |
|---|---|---|---|
|
| |||
| Contextual/Demographics (8) | Age, gender (M/F), race (white/black/Latino), CP score, acuity score (ESI) | Age, gender (M/F), race (white/black/Latino), CP score, acuity score (ESI) | Age, gender (M/F), race (white/black/Latino), CP score, acuity score (ESI) |
| Treatment (4) | Culture test, Antibiotics, Bolus, Critical care drug | Culture test, Antibiotics, Bolus, Critical care drug | Culture est, Antibiotics, Bolus, Critical care drug |
| Vitals (6) | HR z-score, Max RR, Min MAP, Min SpO2, Min SFR, Max Temp | HR z-score, Max RR, Max Temp | HR z-score, Max RR, Max Temp |
| CBC labs (19) | AAL, AAM, ABP, AEP, AGP, AIGP, ALP, AMP, NRBC, HCT, Hb, MCH, MCHC, MCV, MPV, PLT, RBC, RDW SD, WBC | AAL, AAM, ABP, AEP, AGP, AIGP, ALP, AMP, NRBC, HCT, Hb, MCH, MCHC, MCV, MPV, PLT, RBC, RDW SD, WBC | AAL, AAM, ABP, AEP, AGP, AIGP, ALP, AMP, NRBC, HCT, Hb, MCH, MCHC, MCV, MPV, PLT, RBC, RDW SD, WBC |
| CMP labs (14) | ALK, AST, ALT, ALB, BUN, BR, CO2, Ca, Cl, Cr, BGL, k, Na, TP | ALK, AST, ALT, ALB, BUN, BR, CO2, Ca, Cl, Cr, BGL, k, Na, TP | |
| Missing data flags (4) | Missing CMP labs, missing MAP, SpO2 and SFR | ||
Abbreviations: CP Score: score representing the number of complex chronic conditions associated with a patient, Antibiotics: empiric IV antibiotic administered, Bolus: 2 or more fluid boluses administered, HR heart rate, RR respiratory rate, MAP mean arterial pressure, SpO2 blood oxygen saturation, SFR SpO2/FiO2 ratio where FiO2 is the % of inspired O2, ALK Alkaline Phosphatase, ALP Auto Lymphocytes Percent, AST Aspartate Aminotransferase, ALT Alanine Aminotransferase, ALB Albumin, AAL Auto Abs Lymph, AAM Auto Abs Monos, ABP Auto Basophils Percent, AEP Auto EOS Percent, AGP Auto Gran Percent, AIGP Auto Immature Gran Percent, ALP Auto Lymph Percent, AMP Auto Mono Percent, NRBC2 Automated NRBC, BUN Blood Urea Nitrogen, BR Bilirubin, CO2 Carbon dioxide, Ca calcium, Cl Chloride, Cr creatinine, BGL Blood Glucose, HCT Hematocrit, Hb Hemoglobin, MCH Mean corpuscular hemoglobin, MCHC Mean corpuscular hemoglobin concentration, MCV mean corpuscular volume, MPV mean platelets volume, PLT Platelet count, RBC red blood cell count, RDW SD Red cell distribution width standard deviation, Na Sodium, TP Total Protein, WBC white blood cell count.
For the included patients, within the two-hour window, the most common recorded vitals were HR, RR and temperature and the most common labs ordered were CMP, CBC panels and urinalysis. Urine tests were excluded due to the frequent use of free text to describe results. Within the CMP and CBC panels there was significant variability in non-null results. Tables S3a and S3b in Supplement 3 show the count and percent of non-missingness associated with febrile patient CBC and CMP tests recorded within the two-hour window. We used all available CMP labs and CBC labs with at least a 15 % prevalence in the febrile cohort as lab features Table 1).
The provided data included (randomized) ED arrival date, ED LOS and patient disposition that identified both admitted patents as well as patients that received “critical care” (ICU utilization). For admitted patients the data also included the discharge date which was used to identify the total LOS of admitted patients. The mean ED LOS in the study cohort was 4.9 h with a median of 4.3 h (IQR 3.1–9.7 h), suggesting that a useful admission predictive model should be based on observations captured no later than within the first two hours following arrival at the ED. Among admitted febrile patients, the median total hospitalized LOS was 64.7 h (IQR 41.4–114.1 h) and the 95th percentile was 127.1 h, used to define PLOS. Table 2 shows the relationship of cohorts to outcomes showing the severe imbalance associated with the ICU and LOS training data and the improved balance associated with the risk stratified cohort.
Table 2.
The relationship of cohorts to outcomes showing the severe imbalance associated with the ICU and LOS training data.
| Cohorts used for training | Admitted, N (%) | ICU, N (%) | PLOS, N (%) |
|---|---|---|---|
|
| |||
| Cohort 1 (All) N = 12,649 | 4293 (34 %) | 215 (1.6 %) | 917 (7.2 %) |
| Cohort 2 (CC reduced features) N = 5445 | 2285 (42 %) | 125 (2 %) | 475 (8.7 %) |
| Cohort 3 (CC reduced cases) N = 1406 | 711 (51 %) | 34 (2 %) | 199 (14.1 %) |
| Febrile patients with complex chronic condition and suspected infection N = 1436 | 731 (51 %) | 29 (2 %) | 209 (14.4 %) |
Abbreviations: ICU: intensive care unit, PLOS: prolonged length-of-stay, CC: complete case.
Tables S4a and S4b in Supplement 4 show the feature missingness and statistics (mean, median/IQR) distributed across the 3 binary outcomes (admission, ICU, PLOS) with statistical significance (p-values) indicated. Importantly, these tables show that the distribution of missing feature flags against the admission, ICU and PLOS outcomes are statistically significant (p < 0.001). These tables also disclose disproportionate race and gender associations with outcomes (e.g. white males more likely to be admitted than female minorities).
Fig. S1 and Table S5a in Supplement 5 show the 12 highly correlated FS-1 features. Three principal components identified by PCA analysis that explain at least 95 % of the variance associated with the original features as shown in Table S5b were used for FS-1-based predictive modeling.
4.2. Predicting admission
Using the Riley [48] sample size formula with an expected R2 of 0.145 based on previous literature [1], an admission prevalence of 0.51 and 34 features, we estimated the required sample size required for admissions predictive model training to be 1100 suggesting either Cohort-3 with FS-3 features, or risk stratified complex patients with suspected infection to be sufficiently powered at N = 1406 and N = 1436, respectively.
Similarly, using a rule-of-thumb metric like Events per Variable (EPV)>= 10 criteria [49] to determine the minimum sample size and/or maximum number of predictors that can be examined, 711 or 731 events were sufficient to examine 34 features.
Fig. 2 shows the ROC curves and performance characteristics associated with LR and XGB models trained to predict admissions with varying cohorts/feature sets.
Fig. 2.

Comparative Admission Model Performance Characteristics. a: Comparing median, MICE imputation vs CC Datasetsb: Comparing XBG with no imputation vs CC Datasets. c: LR Calibration Curves. d: XGB Calibration Curves. e: LR vs XGB AUC/AUPRC CP score > 0, infection suspect. f: LR vs XGB calibration CP score > 0, infection suspect.
Fig. 2A ROC curves and performance statistics associated with LR show that Cohort 1 using MICE (red) and CC techniques (Cohort 2 orange, Cohort 3 green) do not outperform Cohort1 using median imputation (blue). The MICE and median techniques generate nearly identical curves with AUC of 0.82 and AUPRC of 0.73–074. This is expected for several reasons:
Most of the data in Cohort 1 is from non-critically ill patients, leading both imputation methods to impute similar values (central tendency) for missing data.
Class imbalance causes the performance to be driven largely by how well the model classifies the non-severe majority class, reducing the influence of differences in imputation for severe cases.
The AUC, AUPRC and calibration performance insensitivity to minor differences in imputation, particularly when the missingness predominantly affects a small subset of critically ill patients, contributes to the similar classifier performance.
While from an AUC perspective Cohort-1 with either MICE and median imputation significantly outperform Cohort-2 or −3, from a AUPRC perspective Cohort-3 with a AUPRC of 0.78 outperformed Cohort-1 with imputation, likely due to the class imbalance of Cohort-1 at 34 % vs the balanced Cohort-3 at 51 % (Table 2).
Fig. 2B shows the performance of XGB processing original Cohort 1 data (with no imputation) achieving an AUC of 0.85, AUPRC 0.78 vs Cohort-2 and Cohort-3 with AUC 0.74 respectively showing, consistent with LR findings, that there is no benefit of CCs in dealing with missing data [50]. These results also suggest that compared to LR, XGB does improve class imbalance (AUPRC) performance. From a calibration perspective, Fig. 2C and D show the calibration curves for LG and XGB models indicating that, while the LR models are better aligned with the diagonal, both LG and XGB models can be considered well calibrated.
4.3. Predicting admissions in a risk stratified cohort
Fig. 2E shows improved AUPRC/PPV performance when an improved balance study cohort (N = 2691 risk stratified febrile patients with a CP-score > 0 and suspected infection; 1702 (63 %) admissions) combined with Cohort-1 is used for training with XGB marginally outperforming LR. With Z scores close to zero and non-significant p-values, Fig. 2F shows that following isotonic regression adjustment the LR model was well calibrated with a Z-score of −1.15 and p = 0.2.
In an effort to fine tune these calibration results using recalibration to address deviations (especially for XGB), Table 3 shows the key calibration metrics (Brier, Z score and p value) associated with the LR and XGB models, comparing the performance of models with no recalibration adjustments against Platt and isotonic recalibration. From a Brier perspective, these tables show that in general models based on Cohort −1 were better calibrated (Brier 0.16 vs 0.20), likely due to the retention of more information and reduced bias associated with CC subsets that inherently sacrifice sample size and variability of the original data. These tables also show that in general Platt and isotonic recalibration methods did not improve calibration performance (as indicated by Z-scores and p-values).
Table 3.
LR/XGB calibration metrics.
| Cohort | Met | LR |
XGB |
||||||
|---|---|---|---|---|---|---|---|---|---|
| IMP | No Recal | Platt | Iso | IMP | No Recal | Platt | Iso | ||
|
| |||||||||
| 1 | Brier | Median | 0.16 | 0.16 | 0.16 | None | 0.159 | 0.16 | 0.16 |
| Z | 2.2 | 2.85 | 3.0 | 1.93 | 2.85 | 3.0 | |||
| P | 0.05 | 0.00 | 0.003 | 5.4e-02 | 4.3e-03 | 2.7e-03 | |||
| 1 | Brier | MICE | 0.16 | 0.16 | 0.16 | ||||
| Z | 1.77 | 2.69 | 2.61 | ||||||
| P | 0.07 | 0.00 | 0.009 | ||||||
| 2 | Brier | None | 0.2 | 0.2 | 0.2 | None | 0.197 | 0.197 | 0.198 |
| Z | −0.63 | 0.55 | 0.81 | −0.63 | 0.55 | 0.81 | |||
| P | 0.528 | 0.582 | 0.416 | 5.2e-01 | 5.8e-01 | 4.2e-01 | |||
| 3 | Brier | None | 0.21 | 0.21 | 0.21 | None | 0.206 | 0.201 | 0.206 |
| Z | 0.76 | 0.74 | 0.86 | 0.76 | 0.74 | 0.86 | |||
| P | 0.45 | 0.46 | 0.39 | 4.4e-01 | 4.5e-01 | 3.8e-01 | |||
Abbreviations: Met: Calibration Metric, IMP: imputation method, No Recal: no recalibration, Iso: isotonic regression.
Tables S6a and S6b in Supplement 6 show the most significant (top 15) LR coefficients and corresponding XGB SHAP values associated with the LR and XGB models predicting admission using Cohort-1 (Fig. 2A and B). Table S6c shows that out of the top 15 features identified as most important by LR and XGB, only 6 features were common between the two models. This discrepancy highlights differences in the way these models handle data and evaluate feature importance. Nonlinear feature/outcome interactions and class imbalance affect the models differently: LR adjusts for class imbalance primarily through class weighting and assumes linear relationships between features and admissions whereas tree-based XGB handles imbalance by learning complex/nonlinear interactions, which may highlight different features as important. SHAP values in XGB consider the marginal contribution of features across all trees, whereas LR coefficients reflect the global weight of features in a linear decision boundary.
Given the class imbalance and differences in model structures, combining insights from both models provides a more comprehensive understanding of the predictors driving patient admission [51]. Figs. S6a and S6b in Supplement 6 show additional feature importance metrics: bar plots of the absolute value of LR coefficients and, for XGB, the gain defined as the average improvement in model performance each time the feature is used in a split. Notably these results show that the “CMP” feature was ranked high by both LR and XGB suggesting that the missingness of this key feature is not likely at random.
Table 4 shows the performance of 3 parsimonious LR models predicting admission: 1) Cohort-1 using the 6 features identified in the feature importance analysis shown in Supplement 6 (CP score, RBC, CMP, acuity, HCT and Hb) with median imputation; 2) Cohort-2 with no imputation using the 5 features (CP score, RBC, HCT, Hb, and acuity) and 3) Cohort-3 with no imputation using 7 features (CP score, bolus, acuity, RBC, HCT, ALB and Hb). The AUC performance of these potentially clinically useful models was found to be comparable to the full feature set models (AUC 0.72–0.80). Additionally, the Z-scores and p-values indicate all models are well calibrated. Among parsimonious LR models the cohort 3-based CC model appears to have a moderate precision (AUPRC) performance benefit likely due to the improved balance of data [52–54].
Table 4.
LR Parsimonious Model Performance.
| Cohort-1 6 features |
Cohort-2 5 features |
Cohort-3 7 features |
|
|---|---|---|---|
|
| |||
| AUC | 0.80 0.79, 0.82) | 0.72(0.70, 0.75) | 0.79 (0.74,0.83) |
| Sens | 0.80 | 0.80 | 0.80 |
| Spec | 0.65 | 0.48 | 0.58 |
| Acc | 0.70 | 0.61 | 0.69 |
| PPV | 0.54 | 0.52 | 0.66 |
| AUPRC | 0.70 | 0.68 | 0.81 |
| Brier | 0.16 | 0.20 | 0.18 |
| z-score | Z: −0.10 p: 0.91 | Z: −0.60 p:0.5 | Z:0.40 p:0.8 |
Abbreviations: AUC: area under ROC curve, Sens: sensitivity, Spec: specificity, Acc: accuracy, PPV: positive predictive value, AUPRC: area under the precision-recall curve.
4.4. Predicting ICU utilization
Given the severity of class imbalance as shown in Table 2 we did not have sufficient sample size to power a predictive modeling analysis using 46 uncorrelated features (replacing 12 correlated features in FS-1 with 3 PCs). For example, the number of Cohort-1 ICU events was 215 for an EPV of 4.7. Consequently, our ICU analysis started with the development of a parsimonious feature set that was based on the 6 clinical features identified for admission prediction (CP Score, RBC, CMP, acuity, HCT and Hb) augmented by 8 vitals/labs routinely recorded in the ED that are commonly used in critical care (e.g. WBC, Cr, BUN, HR, Temp, SpO2, SFR and RR) [55] resulting in an EPV of 15.3.
Fig. 3A shows the discrimination performance characteristics of the LR and XGB ICU models with and without the SMOTE correction. From a discrimination perspective this figure shows that while high AUCs were achieved (0.86–0.91), AUPRCs remained low ((0.24–27) despite the use of SMOTE indicating challenges in achieving high precision and recall for the minority class in severe data imbalance situations. From a calibration perspective Fig. 3B shows the corresponding calibration curves without recalibration, indicating that LR No SMOTE with a Brier of 0.014, Z-score of −0.91, p of 0.3 and curve close to the diagonal is well calibrated while adding SMOTE significantly degrades performance (e.g. LR SMOTE (red)). Conversely the XGB model with a higher Z-score, p near zero suggests poor calibration. Adding SMOTE to the XGB model improved Z-score, p-value, but from a calibration perspective, without recalibration, No SMOTE LR was the best performer.
Fig. 3.

Comparing ICU Model Performance Characteristics. a: LR and XGB ROC curves with and without SMOTE. b: LR and XGB Cal curves with and without SMOTE. c: Platt Scaling Recalibrated Curves. d: Isotonic Regression Recalibrated Curves.
Fig. 3C shows the results of Platt scaling indicating that in general all models improved after this recalibration with the most dramatic improvements in SMOTE corrected models (e.g. the LR SMOTE model). Fig. 3D shows the results of recalibration via isotonic regression indicating that from a Z-score/p-value perspective, similar improvements as Platt rescaling were achieved, but resulted in truncated curves at probabilities > 0.6. likely due to the sparse distribution of high-probability predictions caused by class imbalance.
In this series of models LR without SMOTE was the clear best performer given that this model achieved the best AUC/AUPRC of 0.9/0.27 with Z-score of −0.9 and p of 0.3. These results suggest that while SMOTE helps balance the dataset, it may introduce synthetic samples that do not fully capture the underlying minority class distribution, leading to a tradeoff between improved probability range and calibration performance [56,57]. Overall, in cases of severe class imbalance, the combination of No SMOTE and no recalibration appears to provide the best calibration performance, minimizing both overconfidence and oscillatory miscalibration.
4.5. Predicting ICU use in febrile patients with complex chronic conditions (CCC)
In our efforts to identify a clinically useful ICU utilization predictor, we explored a model where we risk stratified the training population as febrile patients that had a cp score > 0 and had a culture test (2691 patients). Among these cases 135 utilized the ICU, reflecting a modest data imbalance improvement (from 1.5 % to 5 %).
Using the same 14 features identified for ICU prediction we developed parallel LR and XGB models exploring the impact of varying the recalibration approach. Fig. 4A and B show the ROC and calibration curves with metrics for LG and XGB models with no recalibration. From a discrimination perspective the Fig. 4A ROC curves show that the LR and XGB models based on stratified data achieved AUCs of 0.90 and 0.87 respectively and LR achieved a AUPRC of 0.43 with corresponding PPV of 0.22, representing a significant improvement from the LR No Smote AUPRC of 0.27 and PPV of 0.07 using the original data (Fig. 3A).
Fig. 4.

Predicting ICU among Febrile Patients with a Chronic Condition and Culture (14 features). a: Risk Stratified LR vs XGB ICU ROC Curves. b: LR vs XGB Calibration (No Recalibration). c: Calibration Curves with Platt Recalibration. d: Calibration Curves with Isotonic Regression.
From a calibration perspective, with no recalibration Fig. 4B, compared to XGB, the LG model had superior calibration performance with a Z-score of − 0.49 and a p-value of 0.63. Fig. 4C and D show that neither Platt rescaling or isotonic regression significantly improved the calibration performance of either the LR or XGB models. While LR with no recalibration exhibits no significant miscalibration (Z = −0.49, p = 0.63), the calibration curve deviates significantly from the diagonal at probabilities > 0.8. Graphically both Platt and isotonic recalibration show significant deviations at all probability ranges. The results demonstrate that in cases of severe imbalance “no recalibration” outperformed recalibration approaches—likely because recalibration methods struggled to handle the extreme skew in data distribution, leading to miscalibration or oscillatory behavior.
These findings underscore the importance of improving data balance through clinically meaningful stratification or similar techniques before applying recalibration. Notably, stratification, which maintains the integrity of the underlying data distribution, proved more effective than resampling methods like SMOTE, which can distort minority class distributions and degrade recalibration performance. Overall, these results highlight the critical interplay between data preprocessing, model calibration, and the choice of recalibration techniques in developing clinically meaningful predictive models. Table 5 shows the LR coefficients and XGB SHAP values of the ICU LR and XGB models with no recalibration.
Table 5.
Risk Stratified ICU LR Coefficients and XGB SHAP values.
| Feature | LR Coefficient (95 % CI) | Mean SHAP Value |
|---|---|---|
|
| ||
| cp_score | 1.249 (1.169, 1.328) | 0.836 |
| Cr | −0.636 (−0.773, −0.498) | 0.136 |
| Hb_z_score | −0.481 (−0.636, −0.326) | 0.298 |
| acuity | 0.330 (0.220, 0.440) | 0.123 |
| max_temp | −0.238 (−0.267, −0.210) | 0.503 |
| RBC | −0.209 (−0.322, −0.096) | 0.626 |
| HCT | 0.195 (0.166, 0.224) | 0.854 |
| max_RR | 0.037 (0.032, 0.042) | 0.583 |
| min_SpO2 | −0.021 (−0.031, −0.012) | 0.343 |
| CMP | 0.013 (−0.091, 0.116) | 0.083 |
| WBC | 0.012 (0.008, 0.016) | 0.513 |
| max_HR | 0.004 (0.002, 0.006) | 0.455 |
| min_sfr | −0.003 (−0.005, −0.002) | 0.332 |
| BUN | 0.000 (−0.011, 0.011) | 0.162 |
4.6. Predicting prolonged length of stay (PLOS) and 24 hr. LOS
As in the case of ICU utilization, predicting prolonged LOS (PLOS) is challenged by class imbalance. Fig. 5A and B show the ROC and calibration performance of the 14-feature LR and XGB models predicting PLOS (defined as >95th percentile of all hospital stays of admitted patients) with AUC of 0.80–0.83 and AUPRC (0.31–0.37) with the LR model showing good calibration with a Brier score of 0.06, Z-score of 0.02 with p-value of 0.99. Fig. 5C and D shows significantly improved performance when outcome is modified to define LOS as an admission > 24 hrs1. With this modified outcome, class imbalance improves significantly (e.g. 30 % had an admission with a LOS > 24 hrs. compared to.07 % (PLOS). Predicting 24 hr. LOS using the same 14 features used for ICU predictions generates significantly improved discrimination with AUPRCs of 0.67 and 0.71 with corresponding PPVs 0.51, 0.49 (Fig. 5C) compared to AUPRCs of 0.34 and 0.28 with PPVs of 0.18 and 0.15. From a calibration perspective, using no recalibration, the 24 Hr. LOS curves (Fig. 5D) show better overall alignment compared to the PLOS curves (Fig. 5B).
Fig. 5.

LR vs XGB Predicting PLOS (>95th percentile) vs 24hrs LOS. a: LR vs XGB ROC Curves: PLOS. b: LR vs XGB ROC Calibration Curves (No recalibration). c: LR vs XGB ROC Curves: 24hrs LOS. d: LR vs XGB Cal Curves: 24hrs LOS.
Table S7 in Supplement 7 presents a performance summary of all derived models.
5. Discussion
In this study, we aim to address key questions in data imputation and predictive modeling to determine the most effective strategies for handling missingness and class imbalance. Specifically, we investigate which imputation method yields the best performance to identify the most robust solution. Additionally, we assess how XGB compares to LR in predictive performance. Finally, we explore the optimal approach for managing class imbalance, a critical factor in ensuring fair and reliable model outcomes. By answering these questions, we seek to provide data-driven insights that can enhance predictive modeling in various applications. It is important noting that we found disproportionately high rates of hospitalization, ICU utilization and PLOS for White children, in concert with studies that have shown Black and Hispanic children are less likely than White to be admitted [58–60]. Additionally, our results are consistent with studies that have shown a significant gender difference in pediatric hospital admissions [61]. Our findings must be interpreted with caution given severe sepsis causing frequent hospitalization or ICU admission in White, male children could be driven by provider diagnosis patterns (e.g., greater likelihood to identify severe sepsis if a patient is a White male, lesser likelihood to identify less severe cases of sepsis if a patient is a White male), complex healthcare system incentives around ICU hospitalization that could disproportionately apply to patients of a specific demographic, or other complex interactions between diagnostic patterns, patient factors, and billing incentives that are not well understood.
5.1. Optimizing missing data imputation: the effectiveness of median imputation
Consistent with a recent analysis of EHR data imputation techniques [62] median imputation was noninferior to MICE, likely due, in part, to non-missingness of most important predictors [19], as well as our finding that missingness was predictive [63–66] and not at random (MNAR) [2,67,68]. The MICE algorithm considers the assumption that missing values are missing at random (MAR), which means that its use in a dataset where the missing values are not MAR could generate biased imputations [69,70]. While MNAR subjects all methods that infer missing values from observed data to problems of bias, articulation of missingness mechanisms is seldom reported in predictive model studies (e.g. the MOFICHE [1] study) employing imputed EHR data [71].
Since most pediatric patients presenting to the ED are not critically ill, and clinicians are less likely to order tests for children who appear healthy, the median imputation technique reflecting the central tendency that is also robust to data noise (e.g. outliers/implausible values [72]), common in EHR data [73], was effective. Moreover, in contrast to studies that support the use of reduced data CC in dealing with missing data [74], in our study, median imputation outperformed CC likely due to additional statistical power resulting from the addition of cases with missing data. We propose that, in general, median imputation can be an effective and widely accepted method in dealing with missingness observed in research studies involving pediatric ED EHR data [75].
5.2. Comparing XGB and LR: Modest Gains vs. Explainability Trade-offs
While there have been numerous studies that have compared the predictive performance of traditional LR with XGB, results have been mixed [8,20,62,76,77]. While our results are consistent with studies that show in balanced, or moderately imbalanced data situations, XGB outperforms LR [20,76], in our case improvements were modest, and may not outweigh the well-established explainability benefits of LR [62,78,79].
In a recent study [80], LR and XGB algorithms were compared in predicting admission among all pediatric patients presenting to the ED. Using all ED visitors resulted in a highly imbalanced (2 % prevalence) training dataset. The study examined 83 potential predictors that included demographics, acuity, physiological parameters and symptoms extracted from free text notes associated with 88,282 patients and concluded that XGB was the best performing algorithm with an AUC of 0.86 and AUPRC of 0.26.
Using ED triage data, the MOFICHE [1] study developed an LR model using MICE with an AUC of 0.84 to predict admissions or ICU in children presenting to EDs with fever. Given the data imbalance associated with ICU use among febrile ED patients, at any meaningful sensitivity (e.g. > 0.5), the PPV for the MOFICHE ICU prediction model was < 0.2, limiting clinical utility.
5.3. Risk stratification: the best approach for handling moderate class imbalance in pediatric prediction models
Risk stratifying the febrile population improved the moderate data imbalance associated with admission and yielded a potentially useful classifier with acceptable [81] AUC/AUPRC performance. Due to the severity of imbalance, this technique was ineffective in addressing ICU and PLOS outcomes. Neither XGB75 nor SMOTE [21,22,82–85] were effective in cases of severe data imbalance. Beyond discrimination performance, calibration often degrades due to systematic underestimation of probabilities for the minority class [86] resulting from severe class imbalance. In cases of severe imbalance resampling techniques such as SMOTE and recalibration methods such as Platt scaling or isotonic regression were not found to be effective in improving calibration performance and, in general, for rare events such as ICU and PLOS, both LR and XGB demonstrated poor AUPRC and calibration, especially at higher predicted probabilities (e.g. >0.6).
Future studies should explore how more advanced risk scoring tools [80,87–90] might be useful towards the development of better-balanced datasets that may yield high AUPRC well-calibrated classifiers of rare conditions like pediatric sepsis. Additionally, methods that combine probabilistic adjustments with more robust calibration techniques, such as Bayesian recalibration or stratified isotonic regression, may offer more effective solutions. Currently, severely imbalanced data used for training classifiers and/or deriving scoring tools (e.g. predicting mortality) is common in pediatric medicine and lead to poorly calibrated, low precision models [56] despite high AUC performance levels [91].
5.4. Limitations
This study employed randomized 10-fold cross-validation to evaluate model performance across various techniques for handling missing and imbalanced data. While this approach fulfills the primary goal of the study, temporal data splitting would enhance the assessment of real-world model applicability. By training models on earlier data and testing them on later data, temporal splitting could reveal shifts in data distributions reflecting shifts in clinical practices (e.g. new treatment protocols) or patient populations (e.g. influenced by pandemics) and inform model recalibration needs [92]. Unfortunately, while potentially useful, such temporal analysis is exacerbated by commonly used data de-identification strategies that randomly anonymize time tags associated with encounter data [93].
Additionally, as a single center study, we would not expect our specific predictive models to be generalizable to other EDs with differing treatment protocols and patient population characteristics. Differences in clinical practice between a facility like CNH, a large, urban, tertiary/academic care center and a major referral hospital with highly specialized protocols, advanced diagnostic capabilities, and subspecialty expertise, combined with CNH’s highly diverse racial and ethnic patient population may result in ML features that differ and/or are not available or used as frequently in smaller hospitals. Urban hospitals like CNH may see a higher burden of severe conditions (e.g., septic shock, trauma, complex chronic conditions) compared to smaller hospitals, which might serve more patients with mild or moderate illnesses. Consequently, CNH’s data may overrepresent high-risk patients, skewing predictive models toward severe cases. These differences in treatment protocols, population characteristics, and healthcare infrastructure could limit the generalizability of models developed at CNH. Future work should include multicenter validation to ensure that predictive models are robust across varied clinical environments.
6. Conclusion
Using simple median imputation, highly interpretable LR models with strong predictive performance can be developed to identify febrile pediatric ED patients at risk of hospitalization. However, addressing severe data imbalances remains a significant challenge in developing predictive models for rare outcomes, such as ICU utilization and PLOS. Future research should focus on advanced calibration techniques and multicenter validation to improve the utility and generalizability of these models.
Supplementary Material
Acknowledgments
The study was supported by a National Institute of Allergy and Infectious Diseases R41 award #1R41AI167224-01A1 (IK, TV).
Appendix A. Supplementary material
Supplementary data to this article can be found online at https://doi.org/10.1016/j.ijmedinf.2025.105905.
Footnotes
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
CRediT authorship contribution statement
Tom Velez: Writing – original draft, Visualization, Software, Methodology, Investigation, Funding acquisition, Formal analysis, Data curation, Conceptualization. Zara Ibrahim: Writing – original draft. Kanayo Duru: Writing – review & editing. Dante Velez: Writing – review & editing. Maria Triantafyllou: Writing – review & editing. Kenneth McKinley: Writing – review & editing. Pasha Saif: Writing – review & editing. Panagiotis Kratimenos: Writing – review & editing. Andy Clark: Writing – review & editing, Software, Data curation. Ioannis Koutroulis: Writing – review & editing, Supervision, Investigation, Funding acquisition, Conceptualization.
References
- [1].Borensztajn DM, et al. , A NICE combination for predicting hospitalisation at the Emergency Department: a European multicentre observational study of febrile children, Lancet Reg. Health Eur. 8 (2021) 100173, 10.1016/j.lanepe.2021.100173. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Fernandes M, et al. , Predicting Intensive Care Unit admission among patients presenting to the emergency department using machine learning and natural language processing, PLoS One 15 (2020) e0229331, 10.1371/journal.pone.0229331. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Berkowitz D, Cohen JS, McCollum N, Rojas CR, Chamberlain JM, Delays in treatment and disposition attributable to undertriage of pediatric emergency medicine patients, Am. J. Emerg. Med. 74 (2023) 130–134, 10.1016/j.ajem.2023.09.054. [DOI] [PubMed] [Google Scholar]
- [4].Goel NN, Durst MS, Vargas-Torres C, Richardson LD, Mathews KS, Predictors of delayed recognition of critical illness in emergency department patients and its effect on morbidity and mortality, J. Intensive Care Med. 37 (2022) 52–59, 10.1177/0885066620967901. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Alpern ER, et al. , Epidemiology of a pediatric emergency medicine research network: the PECARN Core Data Project, Pediatr. Emerg. Care 22 (2006) 689–699, 10.1097/01.pec.0000236830.39194.c0. [DOI] [PubMed] [Google Scholar]
- [6].Muth M, Statler J, Gentile DL, Hagle ME, Frequency of fever in pediatric patients presenting to the emergency department with non-illness-related conditions, J. Emerg. Nurs. 39 (2013) 389–392, 10.1016/j.jen.2012.09.018. [DOI] [PubMed] [Google Scholar]
- [7].Guo BC, et al. , Predictors of bacteremia in febrile infants under 3 months old in the pediatric emergency department, BMC Pediatr. 23 (2023) 444, 10.1186/s12887-023-04271-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Tsai CM, et al. , Use of machine learning to differentiate children with Kawasaki disease from other febrile children in a pediatric emergency department, JAMA Netw. Open 6 (2023) e237489, 10.1001/jamanetworkopen.2023.7489. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Mercurio L, Pou S, Duffy S, Eickhoff C, Risk factors for pediatric sepsis in the emergency department: a machine learning pilot study, Pediatr. Emerg. Care 39 (2023) e48–e56, 10.1097/PEC.0000000000002893. [DOI] [PubMed] [Google Scholar]
- [10].Larburu N, Azkue L, Kerexeta J, Predicting hospital ward admission from the emergency department: a systematic review, J. Pers. Med. 13 (2023), 10.3390/jpm13050849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Anderson RP, Jin R, Grunkemeier GL, Understanding logistic regression analysis in clinical reports: an introduction, Ann. Thorac. Surg. 75 (2003) 753–757, 10.1016/s0003-4975(02)04683-0. [DOI] [PubMed] [Google Scholar]
- [12].Mashoufi M, Ayatollahi H, Khorasani-Zavareh D, Talebi Azad Boni T, Data quality assessment in emergency medical services: an objective approach, BMC Emerg. Med. 23 (10) (2023), 10.1186/s12873-023-00781-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Azur MJ, Stuart EA, Frangakis C, Leaf PJ, Multiple imputation by chained equations: what is it and how does it work? Int. J. Methods Psychiatr. Res. 20 (2011) 40–49, 10.1002/mpr.329. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Jazayeri A, Liang OS, Yang CC, Imputation of missing data in electronic health records based on patients’ similarities, J. Healthc. Inform. Res. 4 (2020) 295–307, 10.1007/s41666-020-00073-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Ross RK, Breskin A, Westreich D, When is a complete-case approach to missing data valid? The importance of effect-measure modification, Am. J. Epidemiol. 189 (2020) 1583–1589, 10.1093/aje/kwaa124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Zhang X, Yan C, Gao C, Malin BA, Chen Y, Predicting missing values in medical data via XGBoost regression, J. Healthc. Inform. Res. 4 (2020) 383–394, 10.1007/s41666-020-00077-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Hurst M, O’Neill M, Pagalan L, Diemert LM, Rosella LC, The impact of different imputation methods on estimates and model performance: an example using a risk prediction model for premature mortality, Popul. Health Metr. 22 (2024) 13, 10.1186/s12963-024-00331-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Gorelick MH, Bias arising from missing data in predictive models, J. Clin. Epidemiol. 59 (2006) 1115–1123, 10.1016/j.jclinepi.2004.11.029. [DOI] [PubMed] [Google Scholar]
- [19].Berkelmans GFN, et al. , Population median imputation was noninferior to complex approaches for imputing missing values in cardiovascular prediction models in clinical practice, J. Clin. Epidemiol. 145 (2022) 70–80, 10.1016/j.jclinepi.2022.01.011. [DOI] [PubMed] [Google Scholar]
- [20].Liu P, Li XJ, Zhang T, Huang YH, Comparison between XGboost model and logistic regression model for predicting sepsis after extremely severe burns, J. Int. Med. Res. 52 (2024) 3000605241247696, 10.1177/03000605241247696. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Zhao Y, Wong ZS, Tsui KL, A framework of rebalancing imbalanced healthcare data for rare events’ classification: a case of look-alike sound-alike mix-up incident detection, J. Healthc. Eng. 2018 (2018) 6275435, 10.1155/2018/6275435. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Gholampour S, Impact of nature of medical data on machine and deep learning for imbalanced datasets: clinical validity of SMOTE is questionable, Mach. Learn. Knowl. Extr. 6 (2024) 827–841. [Google Scholar]
- [23].Yan Z, et al. , XGBoost algorithm and logistic regression to predict the postoperative 5-year outcome in patients with glioma, Ann. Transl. Med. 10 (2022) 860, 10.21037/atm-22-3384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Zhang S, Yao Z, Wang X, Lei Z, Improvement of the performance of models for predicting coronary artery disease based on XGBoost algorithm and feature processing technology, Electronics 11 (2022) 315. [Google Scholar]
- [25].Rahmatinejad Z, et al. , A comparative study of explainable ensemble learning and logistic regression for predicting in-hospital mortality in the emergency department, Sci. Rep. 14 (2024) 3406, 10.1038/s41598-024-54038-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Feinstein JA, Hall M, Davidson A, Feudtner C, Pediatric complex chronic condition system version 3, JAMA Netw. Open 7 (2024) e2420579, 10.1001/jamanetworkopen.2024.20579. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Singh J, Sato M, Ohkuma T, On missingness features in machine learning models for critical care: observational study, JMIR Med. Inform. 9 (2021) e25022, 10.2196/25022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Kim SM, Kim Y, Jeong K, Jeong H, Kim J, Logistic LASSO regression for the diagnosis of breast cancer using clinical demographic data and the BI-RADS lexicon for ultrasonography, Ultrasonography 37 (2018) 36–42, 10.14366/usg.16045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Zhang Z, et al. , Predictive analytics with gradient boosting in clinical medicine, Ann. Transl. Med. 7 (2019) 152, 10.21037/atm.2019.03.29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Nicodemus KK, Malley JD, Predictor correlation impacts machine learning algorithms: implications for genomic studies, Bioinformatics 25 (2009) 1884–1890, 10.1093/bioinformatics/btp331. [DOI] [PubMed] [Google Scholar]
- [31].Schreiber-Gregory DN, Bader K, Regulation Techniques for Multicollinearity: Lasso, Ridge, and Elastic Nets, 2018. [Google Scholar]
- [32].Chan J-Y-L, Bea KT, et al. , Mitigating the multicollinearity problem and its machine learning approach: a review, Mathematics 10 (2022) 1283. [Google Scholar]
- [33].Gregorich M, Strohmaier S, Dunkler D, Heinze G, Regression with highly correlated predictors: variable omission is not the solution, Int. J. Environ. Res. Public Health 18 (2021), 10.3390/ijerph18084259. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Williams JB, Ghosh D, Wetzel RC, Applying machine learning to pediatric critical care data, Pediatr. Crit. Care Med. 19 (2018) 227–233, 10.1097/PCC.0000000000001567. [DOI] [PubMed] [Google Scholar]
- [35].Moons KG, et al. , Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): explanation and elaboration, Ann. Intern. Med. 162 (2015) W1–W, 10.7326/M14-0698. [DOI] [PubMed] [Google Scholar]
- [36].Sandokji I, Yamamoto Y, Biswas A, Arora T, Ugwuowo U, Simonov M, Saran I, Martin M, Testani JM, Mansour S, Moledina DG, Greenberg JH, Wilson FP, A time-updated, parsimonious model to predict AKI in hospitalized children, J. Am. Soc. Nephrol. 31 (2020) 654–664, 10.1681/ASN.2019070745. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].Yi F, et al. , XGBoost-SHAP-based interpretable diagnostic framework for Alzheimer’s disease, BMC Med. Inf. Decis. Making 23 (2023) 137, 10.1186/s12911-023-02238-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Kohavi R, A study of cross-validation and bootstrap for accuracy estimation and model selection, in: Proc. 14th Int. Joint Conf. Artif. Intell. (IJCAI), 1995, pp. 1137–1143. [Google Scholar]
- [39].Shreffler J, Huecker MR, StatPearls, 2024. [Google Scholar]
- [40].Baratloo A, Hosseini M, Negida A, El Ashal G, Part 1: simple definition and calculation of accuracy, sensitivity and specificity, Emerg (Tehran) 3 (2015) 48–49. [PMC free article] [PubMed] [Google Scholar]
- [41].Scarffe A, Coates A, Brand K, Michalowski W, Decision threshold models in medical decision making: a scoping literature review, BMC Med. Inf. Decis. Making 24 (2024) 12, 10.1186/s12911-024-02681-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Saito T, Rehmsmeier M, The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets, PLoS One 10 (2015) e0118432, 10.1371/journal.pone.0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [43].Steyerberg EW, et al. , Assessing the performance of prediction models: a framework for traditional and novel measures, Epidemiology 21 (2010) 128–138, 10.1097/EDE.0b013e3181c30fb2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Wallace BCD, Improving class probability estimates for imbalanced data, Knowl. Inf. Syst. 41 (2014) 33–52. [Google Scholar]
- [45].Niculescu-Mizil A, Caruana R, Predicting good probabilities with supervised learning, in: Proc. 22nd Int. Conf. Mach. Learn. (ICML), 2005, pp. 625–632, doi: 10.1145/1102351.1102430. [DOI] [Google Scholar]
- [46].Alkhawaldeh IM, Albalkhi I, Naswhan AJ, Challenges and limitations of synthetic minority oversampling techniques in machine learning, World J. Methodol. 13 (2023) 373–378, 10.5662/wjm.v13.i5.373. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [47].Rudd K, et al. , Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study, Lancet 394 (10200) (2019) 1555–1568, 10.1016/S0140-6736(19)32989-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [48].Riley RD, et al. , Calculating the sample size required for developing a clinical prediction model, BMJ 368 (2020) m441, 10.1136/bmj.m441. [DOI] [PubMed] [Google Scholar]
- [49].Dhiman P, et al. , Sample size requirements are not being considered in studies developing prediction models for binary outcomes: a systematic review, BMC Med. Res. Method. 23 (2023) 188, 10.1186/s12874-023-02008-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [50].Al-Hadad M, Care of the critically ill begins in the emergency medicine setting, Eur. J. Emerg. Med. 31 (2024) 165–168, 10.1097/MEJ.0000000000001134. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [51].Chen Z, Li T, Guo S, Zeng D, Wang K, Machine learning-based in-hospital mortality risk prediction tool for intensive care unit patients with heart failure, Front. Cardiovasc. Med. 10 (2023) 1119699, 10.3389/fcvm.2023.1119699. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [52].Marsaglia GT, Wang J, Evaluating Kolmogorov’s distribution J Stat. Softw. 8 (2003) 1–4. [Google Scholar]
- [53].Eisenberg MA, Balamuth F, Pediatric sepsis screening in US hospitals, Pediatr. Res. 91 (2022) 351–358, 10.1038/s41390-021-01708-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [54].Dodge Y, The Concise Encyclopedia of Statistics 283–287, Springer, New York, 2008. [Google Scholar]
- [55].Potes C, et al. , A clinical prediction model to identify patients at high risk of hemodynamic instability in the pediatric intensive care unit, Crit. Care 21 (2017) 282, 10.1186/s13054-017-1874-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [56].van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B, The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression, J. Am. Med. Inform. Assoc. 29 (2022) 1525–1534, 10.1093/jamia/ocac093. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [57].Carriero A, Luijken K, de Hond A, Moons KGM, van Calster B, van Smeden M, The harms of class imbalance corrections for machine learning-based prediction models: a simulation study, arXiv 2404.19494 (2024), arXiv:2404.19494. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Zhang X, Carabello M, Hill T, He K, Friese CR, Mahajan P, Racial and ethnic disparities in emergency department care and health outcomes among children in the United States, Front. Pediatr. (2019), 10.3389/fped.2019.00525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Mackey K, et al. , Racial and ethnic disparities in COVID-19-related infections, hospitalizations, and deaths: a systematic review, Ann. Intern. Med. 174 (2021) 362–373, 10.7326/M20-6306. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [60].Price-Haywood EG, Burton J, Fort D, Seoane L, Hospitalization and mortality among black patients and white patients with Covid-19, N. Engl. J. Med. 382 (2020) 2534–2543, 10.1056/NEJMsa2011686. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [61].Mehdi AH, Riaz K, Ghazal N, Kamran NS, Saboohi E, Malick AHH, et al. , Disparities among pediatric hospital admissions according to gender, Int. J. Sci. Rep. 6 (2020), 10.18203/issn.2454-2156.IntJSciRep20202642. [DOI] [Google Scholar]
- [62].de Hond AAH, et al. , Machine learning did not beat logistic regression in time series prediction for severe asthma exacerbations, Sci. Rep. 12 (2022) 20363, 10.1038/s41598-022-24909-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [63].Lipton ZC, ICML Workshop on Human Interpretability in Machine Learning (WHI), 2016. [Google Scholar]
- [64].Che Z, Purushotham S, Cho K, Sontag D, Liu Y, Recurrent neural networks for multivariate time series with missing values, Sci. Rep. 8 (2018) 6085, 10.1038/s41598-018-24271-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [65].Mikalsen K-Ø-S-R, Bianchi FM, Revhaug A, Jenssen R, Time series cluster kernels to exploit informative missingness and incomplete label information, Pattern Recognition 120 (2021). [Google Scholar]
- [66].Singh AO, Lee Y, Wang Y, Utilizing Informative Missingness for Early Detection of Sepsis, 2020. [Google Scholar]
- [67].Austin PC, White IR, Lee DS, van Buuren S, Missing data in clinical research: a tutorial on multiple imputation, Can. J. Cardiol. 37 (2021) 1322–1331, 10.1016/j.cjca.2020.11.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [68].Lin JH, Haug PJ, Exploiting missing clinical data in Bayesian network modeling for predicting medical problems, J. Biomed. Inform. 41 (2008) 1–14, 10.1016/j.jbi.2007.06.001. [DOI] [PubMed] [Google Scholar]
- [69].Sterne JA, et al. , Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls, BMJ 338 (2009) b2393, 10.1136/bmj.b2393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [70].Mera-Gaona M, Neumann U, Vargas-Canas R, Lopez DM, Evaluating the impact ´ of multivariate imputation by MICE in feature selection, PLoS One 16 (2021) e0254720, 10.1371/journal.pone.0254720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [71].Haneuse S, Arterburn D, Daniels MJ, Assessing missing data assumptions in EHR-based studies: a complex and underappreciated task, JAMA Netw. Open 4 (2021) e210184, 10.1001/jamanetworkopen.2021.0184. [DOI] [PubMed] [Google Scholar]
- [72].Hartwig FP, et al. , The median and the mode as robust meta-analysis estimators in the presence of small-study effects and outliers, Res. Synth. Methods 11 (2020) 397–412, 10.1002/jrsm.1402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [73].Estiri H, Klann JG, Murphy SN, A clustering approach for detecting implausible observation values in electronic health records data, BMC Med. Inf. Decis. Making 19 (2019) 142, 10.1186/s12911-019-0852-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [74].Mukaka M, et al. , Is using multiple imputation better than complete case analysis for estimating a prevalence (risk) difference in randomized controlled trials when binary outcome observations are missing? Trials 17 (2016) 341, 10.1186/s13063-016-1473-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [75].Chen X, et al. , Dealing with missing, imbalanced, and sparse features during the development of a prediction model for sudden death using emergency medicine data: machine learning approach, JMIR Med. Inform. 11 (2023) e38590, 10.2196/38590. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [76].Moore A, Bell M, XGBoost, a novel explainable AI technique, in the prediction of myocardial infarction: a UK Biobank Cohort Study, Clin. Med. Insights Cardiol. 16 (2022) 11795468221133611, 10.1177/11795468221133611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [77].Ravi LG, Kauppi K, Big data analytics in logistics and supply chain management: a review of literature, Universal J. Manag. 9 (2021) 29–42, 10.13189/ujm.2021.090201. [DOI] [Google Scholar]
- [78].Gravesteijn BY, et al. , Machine learning algorithms performed no better than regression models for prognostication in traumatic brain injury, J. Clin. Epidemiol. 122 (2020) 95–107, 10.1016/j.jclinepi.2020.03.005. [DOI] [PubMed] [Google Scholar]
- [79].Muehlematter UJ, Daniore P, Vokinger KN, Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015–20): a comparative analysis, Lancet Digit Health 3 (2021) e195–e203, 10.1016/S2589-7500(20)30292-2. [DOI] [PubMed] [Google Scholar]
- [80].Hatachi T, et al. , Machine learning-based prediction of hospital admission among children in an emergency care center, Pediatr. Emerg. Care 39 (2023) 80–86, 10.1097/PEC.0000000000002648. [DOI] [PubMed] [Google Scholar]
- [81].Mandrekar JN, Receiver operating characteristic curve in diagnostic test assessment, J. Thorac. Oncol. 5 (2010) 1315–1316, 10.1097/JTO.0b013e3181ec173d. [DOI] [PubMed] [Google Scholar]
- [82].Alghamdi M, et al. , Predicting diabetes mellitus using SMOTE and ensemble machine learning approach: The Henry Ford ExercIse Testing (FIT) project, PLoS One 12 (2017) e0179805, 10.1371/journal.pone.0179805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [83].Technique S-S, Chawla NVB, Hall LO, Kegelmeyer WP, SMOTE: synthetic minority over-sampling technique, J. Artif. Intell. Res. 16 (2002) 321–357. [Google Scholar]
- [84].Hassanzadeh R, Farhadian M, Rafieemehr H, Hospital mortality prediction in traumatic injuries patients: comparing different SMOTE-based machine learning algorithms, BMC Med. Res. Method. 23 (2023) 101, 10.1186/s12874-023-01920-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [85].Sowjanya AM, Mrudula O, Effective treatment of imbalanced datasets in health care using modified SMOTE coupled with stacked deep learning algorithms, Appl. Nanosci. 13 (2023) 1829–1840, 10.1007/s13204-021-02063-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [86].Challen R, Denny J, Pitt M, Gompels L, Edwards T, Tsaneva-Atanasova K, Artificial intelligence, bias and clinical safety, BMJ Qual. Saf. 28 (2019) 231–237, 10.1136/bmjqs-2018-008370. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [87].Barak-Corren Y, Fine AM, Reis BY, Early prediction model of patient hospitalization from the pediatric emergency department, Pediatrics 139 (2017), 10.1542/peds.2016-2785. [DOI] [PubMed] [Google Scholar]
- [88].Balamuth F, et al. , Validation of the pediatric sequential organ failure assessment score and evaluation of third international consensus definitions for sepsis and septic shock definitions in the pediatric emergency department, JAMA Pediatr. 176 (2022) 672–678, 10.1001/jamapediatrics.2022.1301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [89].Romaine ST, et al. , Accuracy of a modified qSOFA score for predicting critical care admission in febrile children, Pediatrics 146 (2020), 10.1542/peds.2020-0782. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [90].Gomez J-M, Fugar S, Sarau A, Simmons JA, Clark B, Sanghani RM, Aggarwal NT, Williams KA, Doukky R, Volgman AS, Sex differences in COVID-19 hospitalization and mortality, J. Womens Health (2021), 10.1089/jwh.2020.8948. [DOI] [PubMed] [Google Scholar]
- [91].Shadbahr T, et al. , The impact of imputation quality on machine learning classifiers for datasets with missing values, Commun. Med. (Lond.) 3 (2023) 139, 10.1038/s43856-023-00356-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [92].Morita K, Mizuno T, Kusuhara H, Investigation of a data split strategy involving the time axis in adverse event prediction using machine learning, J. Chem. Inf. Model. 62 (2022), 10.1021/acs.jcim.2c00765. [DOI] [PubMed] [Google Scholar]
- [93].Hripcsak G, Mirhaji P, Low AF, Malin BA, Preserving temporal relations in clinical data while maintaining privacy, J. Am. Med. Inform. Assoc. 23 (2016) 1040–1045, 10.1093/jamia/ocw001. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
