Abstract
We developed and externally validated an interpretable machine-learning model for observed continuous renal replacement therapy (CRRT) initiation in sepsis-associated acute kidney injury (SA-AKI) using a strict 24-hour landmark framework. This multicenter retrospective study used MIMIC-IV and eICU-CRD data. Predictors were restricted to the period from sepsis diagnosis to 24 h, and the outcome was first observed CRRT initiation after the landmark and within 7 days. Patients who died, initiated CRRT, or were no longer under observation by 24 h were excluded. Eight algorithms were evaluated, and a seven-predictor Gradient Boosting model was selected. Discrimination, calibration, incremental value, SHAP-based interpretation, and exploratory prediction-subgroup outcomes were assessed. The final MIMIC-IV cohort included 5,238 patients, with 239 CRRT events; 3,666 were assigned to training and 1,572 to internal validation. The eICU-CRD cohort included 4,683 patients and 157 events. The model retained serum creatinine, AKI stage, SOFA score, urine output, lactate, red blood cell distribution width, and peripheral oxygen saturation. AUCs were 0.905 internally and 0.816 externally, with an external calibration slope of 0.567. In an exploratory age- and sex-matched eICU analysis, the high predicted-risk/no observed CRRT subgroup had higher ICU, 7-day, and 28-day mortality than the true-negative subgroup. The model provided interpretable risk estimates for observed CRRT initiation, but external calibration and incremental value were limited. Findings should not be interpreted as evidence of CRRT indication, undertreatment, or treatment benefit. Prospective validation and local recalibration are required.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1007/s10238-026-02305-1.
Keywords: Sepsis, Acute kidney injury, CRRT, Machine learning, Landmark analysis
Introduction
Sepsis-associated acute kidney injury (SA-AKI) is a frequent and serious complication among critically ill patients and is associated with substantial short- and long-term mortality [1–3]. Despite advances in critical care and organ support, outcomes remain poor, particularly among patients with severe kidney dysfunction, some of whom subsequently undergo renal replacement therapy during critical illness [4–6]. Continuous renal replacement therapy (CRRT) is commonly used to manage fluid overload, severe metabolic and electrolyte disturbances, and uremic complications in hemodynamically unstable patients [5]. However, CRRT initiation represents a real-world treatment decision rather than a direct or objective measure of physiological treatment necessity.
Clinical assessment preceding CRRT initiation commonly incorporates serum creatinine, urine output, electrolyte and acid–base disturbances, fluid status, hemodynamic condition, and overall illness severity, including the Sequential Organ Failure Assessment (SOFA) score [6–8]. Although these indicators are indispensable in routine care, individual renal or severity measures may not fully capture the evolving and multidimensional clinical status of patients with SA-AKI [9, 10]. Moreover, the initiation and selection of kidney replacement modalities are influenced by clinician judgment, contraindications, goals of care, institutional practice, affordability, and resource availability [5, 6, 11–13]. Consequently, patients with similar physiological profiles may follow different treatment trajectories across institutions. These considerations underscore the need for multivariable approaches that estimate the probability of subsequent observed CRRT initiation without interpreting that probability as an objective indication for treatment [6].
Machine-learning methods can integrate multidomain information from electronic health records and are increasingly used in biomedical and critical care research [14, 15]. These methods have been applied to the prediction of acute kidney injury, mortality, renal recovery, and CRRT-related outcomes in critically ill populations [16–21]. However, the interpretation of complex models may be difficult, particularly when prediction is driven by high-dimensional and correlated clinical variables [22]. Recent work by Zhuang et al. applied machine learning to mortality prediction among patients with SA-AKI who had already received CRRT and examined the prognostic relevance of CRRT initiation timing [23]. However, evidence specifically addressing subsequent observed CRRT initiation in a broad SA-AKI population using a temporally separated prediction framework and independent external validation remains limited.
An interpretable machine-learning framework based exclusively on information available before a prespecified landmark may help reduce temporal leakage and provide transparent estimates of subsequent CRRT-initiation risk. Beyond conventional measures of discrimination, such a model should be evaluated in terms of calibration, external transportability, and incremental value over established severity and renal indicators. Model interpretation and probability-based subgroup analyses may additionally characterize heterogeneity in physiological risk, but they should not be used to infer CRRT indication, treatment benefit, or undertreatment.
Accordingly, we conducted a multicenter retrospective study using the MIMIC-IV and eICU-CRD databases to develop and externally validate an interpretable machine-learning model for predicting first observed CRRT initiation after a strict 24-hour landmark and within 7 days after sepsis diagnosis among patients with SA-AKI. We compared multiple candidate algorithms, developed a compact predictor set, assessed calibration and external validation performance, and evaluated incremental predictive value relative to SOFA-only, AKI stage-only and renal-only models. We further used model-interpretation methods and exploratory prediction-subgroup analyses to characterize model behavior and risk heterogeneity rather than to infer treatment necessity or effect.
Methods
Study design and data sources
We conducted a retrospective multicenter study using the Medical Information Mart for Intensive Care IV (MIMIC-IV, version 2.2) and the eICU Collaborative Research Database (eICU-CRD). MIMIC-IV contains 73,181 ICU stays involving 50,920 unique ICU patients admitted to Beth Israel Deaconess Medical Center between 2008 and 2019 [24]. The eICU-CRD contains 200,859 patient-unit encounters involving 139,367 patients admitted to 335 units across 208 U.S. hospitals between 2014 and 2015 [25]. Both databases provide deidentified, time-stamped electronic health record data, including demographics, vital signs, laboratory measurements, treatments, and clinical outcomes.
Study population
Adult patients aged ≥ 18 years with sepsis-associated acute kidney injury (SA-AKI) were eligible for inclusion. For patients with repeated ICU admissions, only the first eligible ICU admission was analyzed. Patients with chronic kidney disease (baseline estimated glomerular filtration rate < 60 mL/min/1.73 m² for > 3 months), end-stage kidney disease requiring maintenance dialysis, or advanced malignancy with an expected survival of < 6 months were excluded. Additional exclusions required by the strict 24-hour landmark design are described in the Outcome Definition subsection and summarized in Fig. 1.
Fig. 1.

Cohort construction and strict 24-hour landmark framework. The left pathway shows the sequential exclusions used to construct the MIMIC-IV development and internal validation cohort, and the right pathway shows application of the harmonized eligibility criteria to the eICU-CRD external validation cohort. Predictors were ascertained from sepsis diagnosis (T0) to the 24-hour landmark. Only patients who were alive, remained under observation, and had not initiated CRRT at or before the landmark entered the prediction risk set. The outcome was first observed CRRT initiation after the landmark and by day 7 after sepsis diagnosis. AKI, acute kidney injury; CRRT, continuous renal replacement therapy; eICU-CRD, eICU Collaborative Research Database; ICU, intensive care unit; MIMIC-IV, Medical Information Mart for Intensive Care IV; SA-AKI, sepsis-associated acute kidney injury; T0, sepsis diagnosis time
Definitions of Sepsis and AKI
Sepsis was identified according to the Sepsis-3 criteria as suspected or documented infection accompanied by an acute increase in the Sequential Organ Failure Assessment (SOFA) score of ≥ 2 points [26]. Acute kidney injury (AKI) was identified using the serum creatinine components of the Kidney Disease: Improving Global Outcomes (KDIGO) criteria when either of the following conditions was met: (1) an increase in serum creatinine of ≥ 0.3 mg/dL (26.5 µmol/L) within 48 h; or (2) an increase in serum creatinine to ≥ 1.5 times the baseline value, known or presumed to have occurred within the prior 7 days [7].AKI severity was classified as KDIGO stages 1–3. Patients who met both the Sepsis-3 and KDIGO AKI criteria were classified as having SA-AKI.
Outcome definition
The primary outcome was first observed continuous renal replacement therapy (CRRT) initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. The first documented CRRT start time was identified from the electronic health records. The landmark risk set comprised patients who were alive, remained under ICU observation, and had not initiated CRRT at 24 h after sepsis diagnosis. Patients who died, were discharged from the ICU, or initiated CRRT at or before the landmark were excluded from the risk set. Only predictors available from sepsis diagnosis to the 24-hour landmark were used. The modeling endpoint was landmark_crrt_24h_to_7d; the variable indicating any CRRT record was not used as the modeling outcome.
Data collection and preprocessing
Missingness was first assessed among all extracted candidate variables in the MIMIC-IV landmark cohort and the eICU-CRD external validation cohort. Variables with more than 30% missingness were excluded before downstream modeling. Spearman correlation analyses were then performed among candidate variables retained after missingness filtering to evaluate collinearity. Highly correlated variables were screened using an absolute Spearman correlation threshold of |ρ| > 0.90, and clinically redundant variables were removed. Missing data in the remaining variables were imputed using median imputation fitted in the training set. Continuous predictors were standardized using z-score normalization based on the training set, and the same scaling parameters were applied to the validation cohorts. In the MIMIC-IV cohort, 64 candidate variables were initially extracted, 42 remained after missingness filtering, and 39 remained after correlation screening for downstream model development. In the eICU-CRD external validation cohort, 47 candidate variables were initially assessed, 44 remained after missingness filtering, and 41 remained after correlation screening.
External validation
An independent external validation cohort was constructed from the eICU-CRD database using a harmonized strict 24-hour landmark framework. Patients were selected according to the same clinical target population and landmark eligibility principles as the MIMIC-IV cohort. Predictor variables were ascertained within the first 24 h after sepsis diagnosis, and the external validation outcome was defined as first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. Patients who did not remain eligible for risk prediction at the 24-hour landmark or who had initiated CRRT at or before the landmark were not included in the external validation risk set. The final eICU-CRD external validation cohort included 4,683 patients, including 157 observed CRRT initiation events.
Machine learning model development and comparison
For the MIMIC landmark analysis, the final landmark cohort was partitioned into a training set (70%) and an internal validation set (30%) using stratified random sampling with a fixed random seed. All imputation, standardization, and model fitting steps were fit within the training set and then applied to the internal validation set.
Using the same 39 screened candidate variables, we systematically evaluated eight machine-learning algorithms: Logistic Regression, Linear Discriminant Analysis, Random Forest, Gradient Boosting, Extreme Gradient Boosting, Light Gradient Boosting Machine, Adaptive Boosting, and Extra Trees. Candidate algorithms were compared in the held-out MIMIC-IV internal validation cohort. The primary model was selected by considering discrimination together with overfitting risk, parsimony, interpretability, and the feasibility of subsequent variable reduction, rather than by selecting the algorithm with the highest numerical AUC alone [22, 27].
A full 39-variable Gradient Boosting model was subsequently used to derive the model-based variable-importance ranking. Nested reduced Gradient Boosting models were then constructed by sequentially adding predictors according to this ranking to evaluate the relationship between predictor count and model performance. The final seven-predictor model was developed by integrating model-based importance, parsimony, clinical interpretability, cross-database availability, and variable harmonization. The retained predictors were serum creatinine, AKI stage, SOFA score, urine output, lactate, red blood cell distribution width, and peripheral oxygen saturation. All models predicted first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis.
To quantify incremental predictive value beyond routinely available severity and renal indicators, the final seven-predictor Gradient Boosting model was compared with three.
simpler Gradient Boosting models: a SOFA-only model; an AKI stage-only model; a renal-only model including serum creatinine, AKI stage, and urine output. All comparator models were trained in the MIMIC-IV training cohort using the same preprocessing strategy and were evaluated without refitting in the MIMIC-IV internal validation and eICU-CRD external validation cohorts. AUCs were compared using paired DeLong tests, and AUPRC and Brier scores were also reported.
Feature selection and model explanation
Given the complexity of sepsis-associated AKI, model interpretability was evaluated using model-based feature importance for the primary Gradient Boosting model. Feature contributions were interpreted as associations with model predictions for subsequent observed CRRT initiation, not as evidence of a true physiological indication for CRRT or a treatment-effect relationship. After the final model and predictor set had been fixed, SHapley Additive exPlanations (SHAP) were calculated to characterize the contribution of each retained predictor to model predictions [28]. Global model behavior was summarized using a SHAP beeswarm plot, with predictors ranked according to their mean absolute SHAP values. Local waterfall plots were generated for representative patients with high and low predicted risks to illustrate how individual predictor values shifted the model output from the expected value to the final prediction.
Statistical analysis
All analyses were performed using R with standard statistical and machine learning packages. Continuous variables were expressed as medians with interquartile ranges (IQR) due to non-normal distributions, with between-group comparisons conducted using non-parametric Mann-Whitney U tests (two groups) or Kruskal-Wallis tests (multiple groups). Categorical data were presented as frequencies and percentages, analyzed using chi-square tests or Fisher’s exact tests for small sample sizes. Discriminatory ability was quantified using the area under the receiver operating characteristic curve (AUC-ROC), with 95% confidence intervals. Probability thresholds were determined by maximizing Youden’s index in the training set. Model calibration was evaluated using Brier scores and calibration plots. Differences between AUCs were tested using DeLong’s method for correlated ROC curves [29].
The Youden-index cutoff was determined in the MIMIC-IV training cohort and locked at 0.0295. The final seven-predictor Gradient Boosting model and the locked cutoff were applied to the eICU-CRD external validation cohort. Patients were classified as true negative, false negative, false positive, or true positive according to predicted risk and observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. To explore prognosis among patients with high predicted risk but no observed CRRT initiation, the false-positive subgroup was compared with the true-negative subgroup. A propensity score based on age and sex was estimated. Exact matching on sex and 1:1 nearest-neighbor matching without replacement were performed using a caliper of 0.10 on the logit of the propensity score. Covariate balance was assessed using standardized mean differences. ICU, 7-day, and 28-day mortality were compared in the matched cohort, and odds ratios with 95% confidence intervals and two-sided P values were reported.
Ethical considerations
The study adhered to STROBE and TRIPOD guidelines. The study was exempt from institutional review board approval due to the use of de-identified data, and all authors completed required CITI training (certification 54747348). It should be noted that the initiation of continuous renal replacement therapy (CRRT) reflects a complex real-world clinical decision influenced by disease severity, clinician judgment, institutional practice patterns, and resource availability, rather than a uniform physiological indication.
Results
Study population and cohort construction
The cohort selection process and strict 24-hour landmark framework are summarized in Fig. 1. In the MIMIC-IV cohort, an operational Sepsis-3 source cohort was first assembled and then restricted to an at-risk landmark cohort by excluding patients younger than 18 years, those with chronic kidney disease or advanced malignancy when applicable, those with insufficient ICU observation before the 24-hour landmark or an ICU stay shorter than 24 h, patients who died before the landmark, patients who had CRRT initiated before or at the landmark, patients with severe missingness, and patients with AKI stage 0 or no AKI when applicable. This yielded a strict 24-hour landmark SA-AKI cohort of 5,238 patients, including 239 patients with first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis and 4,999 patients without observed CRRT initiation. The MIMIC-IV cohort was randomly split into a training cohort (n = 3,666) and an internal validation cohort (n = 1,572).
The same strict 24-hour landmark framework and outcome definition were then applied to the eICU-CRD cohort to support external validation. The final eICU strict 24-hour landmark external validation cohort included 4,683 patients, of whom 157 had first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis and 4,526 did not. Out-of-hospital mortality was not available in eICU-CRD; therefore, subsequent external outcome analyses used available ICU or hospital outcomes. Missingness patterns and correlation-based preprocessing are shown in Supplementary Figs. S1-S4.
Baseline characteristics by CRRT status
Baseline characteristics of the MIMIC-IV strict 24-hour landmark cohort are summarized in Table 1. Patients were grouped according to first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis.
Table 1.
Baseline characteristics of the MIMIC-IV strict 24-hour landmark SA-AKI cohort stratified by subsequent observed CRRT initiation
| Characteristic | Overall | No observed CRRT initiation | Observed CRRT initiation | P value |
|---|---|---|---|---|
| Age (years) | 66.06 [54.99, 76.72] | 66.38 [55.32, 76.97] | 60.66 [49.22, 70.62] | < 0.001 |
| Sex, Female/Male |
Female: 3087 (58.9%) Male: 2151 (41.1%) |
Female: 2955 (59.1%) Male: 2044 (40.9%) |
Female: 132 (55.2%) Male: 107 (44.8%) |
0.233 |
| Height (cm) | 170.00 [163.00, 178.00] | 170.00 [163.00, 178.00] | 170.00 [163.00, 178.00] | 0.205 |
| Weight (kg) | 85.80 [71.88, 101.15] | 85.80 [71.71, 101.00] | 86.50 [73.02, 106.80] | 0.075 |
| Body mass index (kg/m^2) | 29.39 [25.42, 34.33] | 29.30 [25.40, 34.26] | 30.52 [26.13, 37.00] | 0.009 |
| SOFA score | 6.00 [4.00, 9.00] | 6.00 [4.00, 8.00] | 12.00 [10.00, 14.00] | < 0.001 |
| LODS score | 5.00 [4.00, 8.00] | 5.00 [4.00, 7.00] | 10.00 [8.00, 11.00] | < 0.001 |
| SAPS II | 39.00 [32.00, 49.00] | 39.00 [31.00, 48.00] | 53.00 [45.00, 63.00] | < 0.001 |
| SIRS score | 3.00 [2.00, 4.00] | 3.00 [2.00, 4.00] | 3.00 [3.00, 4.00] | < 0.001 |
| AKI stage |
Stage 1 1591 (30.4%) Stage 2 3043 (58.1%) Stage 3 604 (11.5%) |
Stage 1 1574 (31.5%) Stage 2 2930 (58.6%) Stage 3 495 (9.9%) |
Stage 1 17 (7.1%) Stage 2 113 (47.3%) Stage 3 109 (45.6%) |
< 0.001 |
| Maximum heart rate (beats/min) | 102.00 [90.00, 116.00] | 102.00 [89.00, 115.00] | 114.00 [98.00, 128.00] | < 0.001 |
| Minimum mean arterial pressure (mmHg) | 60.00 [55.00, 65.00] | 60.00 [55.00, 65.00] | 58.00 [52.50, 64.50] | 0.007 |
| Maximum respiratory rate (breaths/min) | 26.00 [23.00, 30.00] | 26.00 [23.00, 30.00] | 28.50 [24.00, 32.50] | < 0.001 |
| Minimum SpO2 (%) | 93.00 [91.00, 95.00] | 93.00 [91.00, 96.00] | 92.00 [89.00, 94.00] | < 0.001 |
| Minimum temperature | 36.50 [36.06, 36.89] | 36.50 [36.06, 36.89] | 36.50 [35.94, 36.89] | 0.971 |
| Maximum lactate (mmol/L) | 2.00 [1.30, 3.30] | 2.00 [1.30, 3.10] | 4.40 [2.40, 7.30] | < 0.001 |
| Minimum pH | 7.34 [7.29, 7.39] | 7.35 [7.29, 7.39] | 7.26 [7.17, 7.34] | < 0.001 |
| Maximum pH | 7.41 [7.37, 7.45] | 7.41 [7.37, 7.45] | 7.38 [7.32, 7.44] | < 0.001 |
| Minimum PaO2/FiO2 ratio | 195.00 [130.00, 274.82] | 197.50 [134.00, 277.50] | 122.50 [84.00, 186.00] | < 0.001 |
| Minimum RBC count | 3.19 [2.74, 3.72] | 3.20 [2.76, 3.73] | 2.92 [2.38, 3.45] | < 0.001 |
| Maximum RDW (%) | 14.60 [13.60, 15.90] | 14.50 [13.60, 15.90] | 15.90 [14.60, 17.80] | < 0.001 |
| Maximum white blood cell count | 15.60 [11.80, 20.20] | 15.40 [11.80, 20.05] | 17.70 [12.45, 24.90] | < 0.001 |
| Minimum platelet count | 150.00 [107.00, 207.00] | 153.00 [109.00, 209.00] | 99.00 [59.50, 156.00] | < 0.001 |
| Maximum blood urea nitrogen (mg/dL) | 20.00 [15.00, 30.00] | 20.00 [15.00, 29.00] | 36.00 [26.00, 56.00] | < 0.001 |
| Maximum creatinine (mg/dL) | 1.05 [0.80, 1.40] | 1.00 [0.80, 1.40] | 2.50 [1.70, 3.45] | < 0.001 |
| Maximum potassium (mmol/L) | 4.60 [4.20, 5.00] | 4.50 [4.20, 5.00] | 4.90 [4.40, 5.70] | < 0.001 |
| Maximum sodium (mmol/L) | 141.00 [138.00, 143.00] | 141.00 [138.00, 143.00] | 140.00 [136.00, 144.50] | 0.378 |
| Urine output (mL) | 1615.00 [1015.00, 2299.00] | 1655.00 [1075.00, 2330.00] | 535.00 [247.00, 1155.00] | < 0.001 |
| Minimum hemoglobin (g/dL) | 9.60 [8.20, 11.10] | 9.70 [8.30, 11.10] | 8.70 [7.30, 10.60] | < 0.001 |
Continuous variables are reported as median [IQR] and compared using the Wilcoxon rank-sum test. Categorical variables are reported as n (%) and compared using the chi-square test or Fisher exact test, as appropriate. SMDs are rounded to 3 decimals. AKI stage uses a generalized multi-category
Patients with subsequent observed CRRT initiation had higher illness severity, more advanced AKI stage, higher serum creatinine and lactate concentrations, lower urine output, and lower oxygen saturation than patients without observed CRRT initiation in the landmark prediction window.
Using the MIMIC-IV strict 24-hour landmark cohort, we first restricted the prediction task to clinical variables available before the landmark and then applied leakage-safe preprocessing. After excluding post-landmark length-of-stay variables, 64 candidate variables were initially considered, 42 remained after missingness filtering, and 39 screened candidate variables remained after correlation screening. These 39 screened candidate variables were then used to train and compare eight machine learning algorithms in the MIMIC-IV internal validation cohort for the outcome of first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. Figure 2A shows the ROC curves for the eight algorithms, and Supplementary Table S1 provides the corresponding performance metrics.
Fig. 2.

Model development workflow in the MIMIC-IV strict 24-hour landmark cohort. (a) Receiver operating characteristic curves of eight machine-learning algorithms evaluated in the held-out MIMIC-IV internal validation cohort. All candidate algorithms were trained using the same 39 screened predictor variables. (b) Internal validation AUCs of nested reduced Gradient Boosting models constructed by sequentially adding predictors according to the variable-importance ranking derived from the full 39-variable Gradient Boosting model. Points represent AUC estimates and the shaded area represents the corresponding 95% confidence intervals. The red dashed line at seven predictors indicates an approximate performance plateau and does not represent the final clinically retained seven-predictor model. (c) Accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1 score, Brier score, and calibration slope of the nested Gradient Boosting models across different predictor counts. The red dashed line again indicates the approximate performance plateau around seven predictors. (d) Normalized variable-importance ranking derived from the full 39-variable Gradient Boosting model. The 20 highest-ranked predictors are displayed; red bars and right-side annotations identify variables retained in the final clinically selected model, whereas blue bars represent other screened candidate variables. The final seven predictors—serum creatinine, AKI stage, SOFA score, urine output, lactate, red blood cell distribution width, and peripheral oxygen saturation—were selected by integrating model-based importance, parsimony, clinical interpretability, cross-database availability, and variable harmonization rather than by simply retaining the seven highest-ranked variables. The predicted outcome was first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. Abbreviations: AKI, acute kidney injury; AUC, area under the receiver operating characteristic curve; CRRT, continuous renal replacement therapy; LDA, linear discriminant analysis; NPV, negative predictive value; PPV, positive predictive value; RDW, red blood cell distribution width; ROC, receiver operating characteristic; SOFA, Sequential Organ Failure Assessment; SpO₂, peripheral oxygen saturation
On the basis of this comparison, several algorithms showed broadly comparable discrimination, and the final model choice was not based on AUC alone. Instead, the primary model was carried forward as a Gradient Boosting model because it offered a practical balance of discrimination, flexibility for predictor ranking, parsimony, and interpretability. The primary Gradient Boosting model was then used to generate a full 39-variable importance ranking. Figure 2D presents this ranking and highlights the clinically retained predictors. Importantly, the final seven clinically retained predictors were not identical to the top seven importance-ranked variables. Final predictor selection therefore integrated model-based importance, parsimony, clinical interpretability, cross-database availability, and variable harmonization rather than relying on rank order alone.
To examine how model complexity affected discrimination and other performance characteristics, we next constructed nested reduced Gradient Boosting models by sequentially adding predictors according to the full 39-variable importance ranking. Figure 2B and C, together with Supplementary Table S2, summarize these nested models. The red vertical line at seven predictors marks an approximate performance plateau and should not be interpreted as indicating that the final clinically retained seven predictors were simply the top seven importance-ranked variables. In the nested reduction analysis, the row with “n_predictors = 7” reflects the performance achieved when the first seven predictors in the full 39-variable importance ranking were included, not the final clinically retained seven-predictor model.
The final clinically retained predictors were serum creatinine, AKI stage, SOFA score, urine output, lactate, RDW, and SpO2.
Internal stability of the final seven-predictor Gradient Boosting model was assessed using 5-fold and 10-fold cross-validation in the MIMIC-IV strict 24-hour landmark training set. The model showed broadly consistent discriminative performance, with pooled out-of-fold AUCs of 0.914 and 0.905 in the 5-fold and 10-fold analyses, respectively (Supplementary Figs. S5–S6).
External validation and model robustness
The final seven-predictor Gradient Boosting model was further evaluated in the MIMIC-IV internal validation cohort and in an independent eICU-CRD strict 24-hour landmark external validation cohort. The eICU-CRD cohort included 4,683 patients, of whom 157 patients experienced first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis, corresponding to an event rate of 3.35%. In the external validation cohort, the model showed acceptable discrimination, with an AUC of 0.816 (95% CI, 0.790–0.842) (Fig. 3A). The calibration intercept was − 2.496 and the calibration slope was 0.567(Fig. 3B). Details are summarized in Supplementary Table S3.
Fig. 3.

External validation performance of the final seven-predictor Gradient Boosting model in the eICU-CRD cohort. (a) Receiver operating characteristic curve of the primary Gradient Boosting model in the eICU-CRD strict 24-hour landmark external validation cohort. The model showed acceptable external discrimination, with an AUC of 0.816 (95% CI, 0.790–0.842). (b) Calibration plot showing observed CRRT initiation rates across quintiles of predicted risk in the eICU-CRD cohort. The calibration intercept was − 2.496 and the calibration slope was 0.567, indicating poor external calibration with substantial overprediction in the external validation cohort. The outcome was first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis
Incremental value over conventional severity and renal indicators
To assess whether the final seven-predictor Gradient Boosting model provided incremental predictive information beyond conventional severity and renal indicators, we compared it with SOFA-only, AKI stage-only and renal-only models using the same strict 24-hour landmark framework. In the MIMIC-IV internal validation cohort, the final model achieved an AUC of 0.905, an AUPRC of 0.367, and a Brier score of 0.036. It showed higher discrimination than the SOFA-only model (AUC, 0.841; ΔAUC, 0.064; DeLong P < 0.001) and the AKI stage-only model (AUC, 0.707; ΔAUC, 0.198; DeLong P < 0.001). The difference from the renal-only model was not statistically significant (AUC, 0.877; ΔAUC, 0.028; DeLong P = 0.099).
In the eICU-CRD external validation cohort, the final model achieved an AUC of 0.816 and an AUPRC of 0.124. It outperformed the AKI stage-only model (AUC, 0.578; ΔAUC, 0.238; DeLong P < 0.001) and the renal-only model (AUC, 0.773; ΔAUC, 0.043; DeLong P < 0.001), but did not significantly outperform the SOFA-only model (AUC, 0.806; ΔAUC, 0.010; DeLong P = 0.520). (Supplementary Table S4).
To enhance model transparency and support clinical interpretation of model behavior, SHAP analysis was performed for the final seven-predictor Gradient Boosting model in the MIMIC-IV strict 24-hour landmark internal validation cohort. As shown in Fig. 4a, serum creatinine, urine output, SOFA score, and lactate showed the largest average contributions to the model-predicted probability of first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis. SpO₂, AKI stage, and RDW provided additional explanatory contributions.
Fig. 4.

SHAP-based interpretation of the final seven-predictor Gradient Boosting model. (a) SHAP summary beeswarm plot showing the distribution of SHAP values for the seven retained predictors in the MIMIC-IV strict 24-hour landmark internal validation cohort. Points are colored according to feature values, and predictors are ordered by mean absolute SHAP value. (b) Local SHAP waterfall-style explanation for a representative high predicted-risk patient from the internal validation cohort, showing how individual predictors shifted the model-predicted probability of observed CRRT initiation from the baseline expectation E[f(X)] to the final prediction f(x). (c) Local SHAP waterfall-style explanation for a representative low predicted-risk patient from the internal validation cohort. SHAP values indicate contributions to the model-predicted probability only and should not be interpreted as causal effects or treatment guidance
Representative local SHAP waterfall plots further illustrated how the retained predictors contributed to individual-level predictions. In the high predicted-risk patient, elevated creatinine, higher SOFA score, reduced urine output, advanced AKI stage, lower SpO₂, higher RDW, and lactate collectively shifted the predicted probability upward from the baseline expectation to the final model prediction (Fig. 4b). In contrast, in the low predicted-risk patient, the retained predictors shifted the model output toward a lower predicted probability (Fig. 4c).
In the eICU-CRD external validation cohort, prediction subgroups were defined using the locked Youden cutoff of 0.0295 derived from the MIMIC-IV training cohort. The final seven-predictor Gradient Boosting model and cutoff were applied without re-optimization. Of the 4,683 patients, 1,761 were classified as predicted low risk and 2,922 as predicted high risk, corresponding to 1,758 true negatives, 3 false negatives, 2,768 false positives, and 154 true positives (Supplementary Table S5).
As an exploratory sensitivity analysis, age- and sex-based propensity score matching was performed between the false-positive and true-negative subgroups. Exact matching on sex and 1:1 nearest-neighbor matching without replacement using a caliper of 0.10 on the logit of the propensity score yielded 1,740 matched pairs. Covariate balance improved after matching, with the standardized mean difference decreasing from 0.120 to 0.061 for age and from 0.035 to 0.000 for sex (Supplementary Table S6).
In the matched cohort, the false-positive subgroup had higher post-landmark mortality than the true-negative subgroup, including ICU mortality (21.6% vs. 7.1%; OR, 3.58; 95% CI, 2.86–4.48; P < 0.001), 7-day mortality (21.8% vs. 7.9%; OR, 3.21; 95% CI, 2.59–3.97; P < 0.001), and 28-day mortality (31.8% vs. 13.4%; OR, 3.05; 95% CI, 2.55–3.66; P < 0.001) (Supplementary Table S7).
Discussion
In this strict 24-hour landmark analysis of MIMIC-IV with external validation in eICU-CRD, we developed an interpretable machine-learning model for predicting first observed CRRT initiation after the 24-hour landmark and by day 7 after sepsis diagnosis among patients with sepsis-associated acute kidney injury. The final seven-predictor Gradient Boosting model showed strong discrimination in internal validation and retained acceptable discriminatory performance in the external cohort. Model interpretation and exploratory subgroup analyses were used to characterize prediction patterns and risk heterogeneity rather than to infer CRRT indication, treatment necessity, or treatment effect.
The landmark framework was a central methodological feature of the study. Predictors were restricted to information available from sepsis diagnosis to the 24-hour landmark, whereas the outcome was defined only after the landmark. Patients who died, underwent CRRT initiation, or were no longer under observation by the landmark were excluded from the corresponding at-risk cohort. This design reduced the likelihood that information occurring concurrently with or after CRRT initiation contributed to prediction and therefore addressed an important source of temporal leakage in the original analysis. Nevertheless, the study remained retrospective, and the landmark approach cannot eliminate residual bias arising from unmeasured clinical decisions or incomplete data capture.
The comparison of candidate algorithms also illustrates why discrimination alone may be insufficient for selecting a primary prediction model in a low-event-rate setting, because models with similar population-level performance may still generate inconsistent risk estimates for individual patients [27]. Although XGBoost achieved the highest numerical AUC among the eight algorithms evaluated using the same 39 screened variables, the differences among several leading models were relatively small. Highly flexible tree-based methods also showed very high apparent training performance, raising concern about optimism when model selection is based only on the highest validation AUC. Gradient Boosting was therefore carried forward after considering discrimination together with parsimony, interpretability, cross-database variable availability, and the feasibility of constructing a compact model. The final model retained serum creatinine, AKI stage, SOFA score, urine output, lactate, RDW, and SpO₂, representing complementary domains of renal dysfunction, systemic organ failure, metabolic stress, hematologic-inflammatory disturbance, and oxygenation impairment.
The present study addresses a different clinical question from previous CRRT-related prediction research. Zhuang et al. [23], for example, evaluated mortality risk among patients with SA-AKI who had already received CRRT and examined the prognostic relevance of CRRT initiation timing. In contrast, the present analysis included the broader population of patients with SA-AKI and modeled first observed CRRT initiation after a prespecified 24-hour landmark. Other studies have evaluated CRRT initiation timing, mortality, or renal recovery among patients with AKI who had already been selected for or received CRRT [18–19, 23, 30–32], rather than predicting subsequent CRRT initiation across the broader SA-AKI population.
Accordingly, the outcome in the present study should be interpreted as an observed treatment-utilization event influenced by physiological severity, clinician judgment, institutional practice, goals of care, and resource availability, rather than as an objective measure of CRRT indication or evidence regarding the benefit of earlier treatment [5, 6, 11, 13].
The incremental analyses further clarify the model’s added value. In MIMIC-IV, the final model improved discrimination over the SOFA-only and AKI stage-only models, whereas the improvement over the renal-only model was modest and not statistically significant. In eICU-CRD, the final model outperformed the AKI stage-only and renal-only models but did not significantly outperform the SOFA-only model. These findings indicate that the model provides additional predictive information, particularly beyond AKI stage alone. However, its added value beyond routine renal indicators was not consistent across cohorts, and its incremental value beyond established clinical severity assessment remains modest.
SHAP analysis was used to characterize how the final Gradient Boosting model assigned predicted probabilities. Serum creatinine and urine output were among the largest contributors, followed by SOFA score and lactate, with additional contributions from AKI stage, RDW, and SpO2. Serum creatinine, AKI stage, and urine output reflect the severity of renal dysfunction, whereas the SOFA score reflects overall organ failure severity [1, 6, 7, 13, 32]. Lactate reflects metabolic and circulatory stress and has been associated with adverse outcomes in critically ill patients with sepsis or AKI requiring CRRT [33–34]. RDW has also been associated with mortality in patients with sepsis and SA-AKI [35–37]. SpO₂ represents oxygenation status and was also identified as an important prognostic feature in the CRRT-treated SA-AKI population studied by Zhuang et al. [23]. Nevertheless, the biological interpretation of individual feature contributions should remain cautious. SHAP values explain model behavior within the fitted dataset and do not establish causal effects, specific pathophysiological mechanisms, CRRT indication, treatment necessity, or recommendations for CRRT initiation [22, 28].
Risk-stratification analyses identified discordance between predicted probability and observed CRRT initiation. Within the eICU-CRD external validation cohort, patients with high predicted risk but no observed CRRT initiation had more severe physiological abnormalities and poorer outcomes than true-negative patients. After exact matching on sex and 1:1 nearest-neighbor propensity-score matching, the false-positive subgroup had approximately threefold higher odds of ICU, 7-day, and 28-day mortality. These findings suggest that some apparent false-positive classifications may identify a prognostically adverse phenotype rather than random classification error.
This interpretation remains descriptive. Matching accounted only for age and sex and did not balance physiological severity, treatment limitations, contraindications, goals of care, competing clinical priorities, institutional practice, or resource availability [5, 11, 13]. In addition, some patients may have died after the landmark before CRRT could be initiated, making early death a competing event for subsequent observed CRRT initiation. We therefore cannot determine whether CRRT was clinically indicated, withheld, delayed, or capable of improving outcomes in this subgroup. This subgroup should be interpreted as a high predicted-risk/no observed CRRT initiation phenotype with adverse prognosis rather than as an undertreated population. Prospective studies are needed to determine whether this type of probability-based stratification can support phenotypic characterization, monitoring strategies, or trial enrichment.
At the model-performance level, external validation demonstrated acceptable discrimination but attenuated calibration and limited threshold-based classification performance. The external AUC of 0.816 indicated that the model retained discriminatory ability in a separate critical care system. However, the calibration slope of 0.567, together with the low PPV and F1 score at the locked MIMIC-derived cutoff, showed that predicted probabilities and threshold-based classifications were less well aligned with the external event distribution. Differences in case mix, variable measurement, database structure, missingness patterns, CRRT practices, and institutional resources may have contributed to this attenuation. These results support moderate rather than robust transportability and indicate that local validation and recalibration would be required before prospective clinical evaluation.
Several limitations should be acknowledged. First, this was a retrospective analysis of two critical care databases and was therefore susceptible to missing data, residual confounding, and measurement error. Second, observed CRRT initiation is a treatment-utilization outcome influenced by clinician judgment, institutional practice, contraindications, goals of care, resource availability, and factors not fully represented in structured databases; it is not equivalent to an objective physiological indication for CRRT. Third, external calibration was attenuated, and the locked MIMIC-derived cutoff showed limited specificity and positive predictive value in eICU-CRD. Fourth, the age- and sex-matched sensitivity analysis was exploratory and did not account for broader clinical or institutional differences between the false-positive and true-negative subgroups. Fifth, the small number of false-negative patients limited the reliability of comparisons involving that subgroup. Sixth, the 24-hour landmark summarized only early information and did not capture subsequent changes in physiology or treatment. Finally, out-of-hospital mortality information was not consistently available in eICU-CRD, limiting some external outcome assessments.
Conclusion
In patients with sepsis-associated acute kidney injury, the final seven-predictor Gradient Boosting model provided interpretable early risk estimates for observed CRRT initiation after a strict 24-hour landmark and within 7 days after sepsis diagnosis. The model demonstrated strong discrimination in the MIMIC-IV internal validation cohort and acceptable discrimination in the eICU-CRD external validation cohort, although external calibration and threshold-based classification performance were attenuated. Its incremental value over conventional severity and renal indicators was modest, and the model should therefore be interpreted as a framework for risk characterization and hypothesis generation rather than as evidence of CRRT indication or as a directive for treatment initiation. Prospective validation, local recalibration, and formal clinical-impact assessment are required before clinical implementation.
Supplementary Information
Below is the link to the electronic supplementary material.
Supplementary Material 1: Checklist. TRIPOD checklist for prediction model development and validation.
Supplementary Material 2: Checklist. STROBE checklist for observational studies.
Acknowledgements
This work was supported by the Hebei Provincial Medical Applicable Technology Tracking Project (2025), entitled “Application and Promotion of Driving Pressure and Mechanical Power in Guiding ARDS Therapy and Big Data Predictive Modeling” (Grant No. GZ20250010), and the Hebei Provincial Medical Scientific Research Project Plan (Guiding Project) (2024), entitled “The Effect and Impact of Early Programmed Critical Care Rehabilitation Nursing on Respiratory Function in ARDS” (Grant No. 20241884). Both projects were funded by the Hebei Provincial Health Commission. The funders had no role in the study design, data collection, analysis, interpretation of data, manuscript preparation, or the decision to submit the manuscript for publication.The authors acknowledge the contributors of the MIMIC-IV and eICU Collaborative Research Database for providing access to the data used in this study.
Author contributions
Y.T. and S.Z. contributed equally to this work. Y.T. and S.Z. were responsible for study conception and design, data extraction, and statistical analysis. Q.W. and Q.Z. contributed to data interpretation and manuscript drafting. L.J. and R.L. assisted with data preprocessing, model development, and figure preparation. M.F. supervised the study, provided critical revisions for important intellectual content, and approved the final manuscript. All authors reviewed and approved the final version of the manuscript.
Data availability
The data analyzed in this study are available from the Medical Information Mart for Intensive Care IV (MIMIC–IV) and the eICU Collaborative Research Database (eICU–CRD).The processed datasets and analysis code supporting the conclusions of this article are available from the corresponding author upon reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Ethics Statement
The study was conducted in accordance with the Declaration of Helsinki. This study used data from the Medical Information Mart for Intensive Care IV (MIMIC–IV) and the eICU Collaborative Research Database (eICU–CRD), both of which contain fully de-identified patient information. As all data were anonymized and publicly available, ethical approval and informed consent were waived according to institutional and national guidelines. Access to the databases was granted after completion of the required data use agreements and training.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yaoyao Tang and Shuai Zhang contributed equally to this work and share first authorship.
References
- 1.Ostermann M, Lumlertgul N, Jeong R, See E, Joannidis M, James M. Acute kidney injury. Lancet. 2025;405(10474):241–56. 10.1016/S0140-6736(24)02385-7. [DOI] [PubMed] [Google Scholar]
- 2.Manrique-Caballero CL, Del Rio-Pertuz G, Gomez H. Sepsis-Associated Acute Kidney Injury. Crit Care Clin. 2021;37(2):279–301. 10.1016/j.ccc.2020.11.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Aguilar MG, AlHussen HA, Gandhi PD, et al. Sepsis-associated acute kidney injury: pathophysiology and treatment modalities. Cureus. 2024;16(12):e75992. 10.7759/cureus.75992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Bottari G, Ranieri VM, Ince C, et al. Use of extracorporeal blood purification therapies in sepsis: the current paradigm, available evidence, and future perspectives. Crit Care. 2024;28(1):432. 10.1186/s13054-024-05220-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Galindo P, Neyra JA. Continuous Renal Replacement Therapy: What have we learned and what are key milestones for the years to come? Rev Invest Clin. 2023;75(6):348–58. 10.24875/RIC.23000221. [DOI] [PubMed] [Google Scholar]
- 6.Zarbock A, Nadim MK, Pickkers P, et al. Sepsis-associated acute kidney injury: consensus report of the 28th Acute Disease Quality Initiative workgroup. Nat Rev Nephrol. 2023;19(6):401–17. 10.1038/s41581-023-00683-3. [DOI] [PubMed] [Google Scholar]
- 7.Palevsky PM, Liu KD, Brophy PD, et al. KDOQI US commentary on the 2012 KDIGO clinical practice guideline for acute kidney injury. Am J Kidney Dis. 2013;61(5):649–72. 10.1053/j.ajkd.2013.02.349. [DOI] [PubMed] [Google Scholar]
- 8.Hu H, Li L, Zhang Y, et al. A prediction model for assessing prognosis in critically ill patients with sepsis-associated acute kidney injury. Shock. 2021;56(4):564–72. 10.1097/SHK.0000000000001768. [DOI] [PubMed] [Google Scholar]
- 9.Ostermann M, Zarbock A, Goldstein S, et al. Recommendations on acute kidney injury biomarkers from the acute disease quality initiative consensus conference: a consensus statement. JAMA Netw Open. 2020;3(10):e2019209. 10.1001/jamanetworkopen.2020.19209. [DOI] [PubMed] [Google Scholar]
- 10.Xu H, Mo R, Liu Y, Niu H, Cai X, He P. L-shaped association between triglyceride-glucose body mass index and short-term mortality in ICU patients with sepsis-associated acute kidney injury. Front Med (Lausanne). 2024;11:1500995. 10.3389/fmed.2024.1500995. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Neyra JA, Yessayan L, Thompson Bastin ML, Wille KM, Tolwani AJ. How to prescribe and troubleshoot continuous renal replacement therapy: a case-based review. Kidney360. 2020;2(2):371–84. 10.34067/KID.0004912020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Taha AKA, Shigidi MMT, Abdulfatah NM, Alsayed RK. The use of sustained low-efficiency dialysis in the treatment of sepsis-associated acute kidney injury in a low-income country: a Prospective Cohort Study. Indian J Crit Care Med. 2024;28(1):30–5. 10.5005/jp-journals-10071-24595. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zarbock A, Koyner JL, Gomez H, Pickkers P, Forni L, Acute Disease Quality Initiative Group. Sepsis-associated acute kidney injury—treatment standard. Nephrol Dial Transpl. 2023;39(1):26–35. 10.1093/ndt/gfad142. [DOI] [PubMed] [Google Scholar]
- 14.Goecks J, Jalili V, Heiser LM, Gray JW. How machine learning will transform biomedicine. Cell. 2020;181(1):92–101. 10.1016/j.cell.2020.03.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Topaloğlu D, Polat O. Trends and methods in intensive care unit (ICU) research using machine learning: latent dirichlet allocation (LDA)-based thematic literature review. BMC Med Inf Decis Mak. 2025;25(1):282. 10.1186/s12911-025-03132-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Barboi C, Tzavelis A, Muhammad LN. Comparison of severity of illness scores and artificial intelligence models that are predictive of intensive care unit mortality: meta-analysis and review of the literature. JMIR Med Inf. 2022;10(5):e35293. 10.2196/35293. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Tseng PY, Chen YT, Wang CH, et al. Prediction of the development of acute kidney injury following cardiac surgery by machine learning. Crit Care. 2020;24(1):478. 10.1186/s13054-020-03179-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Wang Y, Sun X, Lu J, Zhong L, Yang Z. Construction and evaluation of a mortality prediction model for patients with acute kidney injury undergoing continuous renal replacement therapy based on machine learning algorithms. Ann Med. 2024;56(1):2388709. 10.1080/07853890.2024.2388709. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhong L, Liang Y, Xu L, Ji XW, Xie B, Yang XH. Construction and evaluation of prediction model for renal function recovery in acute kidney injury patients undergoing continuous renal replacement therapy based on machine learning algorithms. Ann Med. 2025;57(1):2561794. 10.1080/07853890.2025.2561794. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Luo XQ, Yan P, Duan SB, et al. Development and validation of machine learning models for real-time mortality prediction in critically ill patients with sepsis-associated acute kidney injury. Front Med (Lausanne). 2022;9:853102. 10.3389/fmed.2022.853102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Fan Z, Jiang J, Xiao C, et al. Construction and validation of prognostic models in critically ill patients with sepsis-associated acute kidney injury: interpretable machine learning approach. J Transl Med. 2023;21(1):406. 10.1186/s12967-023-04205-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206–15. 10.1038/s42256-019-0048-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhuang C, Hu R, Li K, et al. Machine learning prediction models for mortality risk in sepsis-associated acute kidney injury: evaluating early versus late CRRT initiation. Front Med (Lausanne). 2025;11:1483710. 10.3389/fmed.2024.1483710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Johnson AEW, Bulgarelli L, Shen L, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10(1):1. 10.1038/s41597-022-01899-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Pollard TJ, Johnson AEW, Raffa JD, Celi LA, Mark RG, Badawi O. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Sci Data. 2018;5:180178. 10.1038/sdata.2018.178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Singer M, Deutschman CS, Seymour CW, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801–10. 10.1001/jama.2016.0287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Li Y, Sperrin M, Ashcroft DM, van Staa TP. Consistency of variety of machine learning and statistical models in predicting clinical risks of individual patients: longitudinal cohort study using cardiovascular disease as exemplar. BMJ. 2020;371:m3919. 10.1136/bmj.m3919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Samadi ME, Nikulina K, Fritsch SJ, Schuppert A. GPT-4o and the quest for machine learning interpretability in ICU risk of death prediction. BMC Med Inf Decis Mak. 2025;25(1):373. 10.1186/s12911-025-03224-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Bantis LE, Nakas CT, Reiser B. Statistical inference for the difference between two maximized Youden indices obtained from correlated biomarkers. Biom J. 2021;63(6):1241–53. 10.1002/bimj.202000128. [DOI] [PubMed] [Google Scholar]
- 30.Nuermaimaiti M, Wang M, Lou R, et al. The impact of initiation timing of continuous renal replacement therapy on outcomes in critically ill patients with acute kidney injury: a retrospective study from the MIMIC-IV database. Sci Rep. 2025;15(1):10922. 10.1038/s41598-024-84435-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Murugan R, Kerti SJ, Chang CH, et al. Association between net ultrafiltration rate and renal recovery among critically ill adults with acute kidney injury receiving continuous renal replacement therapy: an observational cohort study. Blood Purif. 2022;51(5):397–409. 10.1159/000517281. [DOI] [PubMed] [Google Scholar]
- 32.Hansrivijit P, Yarlagadda K, Puthenpura MM, et al. A meta-analysis of clinical predictors for renal recovery and overall mortality in acute kidney injury requiring continuous renal replacement therapy. J Crit Care. 2020;60:13–22. 10.1016/j.jcrc.2020.07.012. [DOI] [PubMed] [Google Scholar]
- 33.Kim SG, Lee J, Yun D, et al. Hyperlactatemia is a predictor of mortality in patients undergoing continuous renal replacement therapy for acute kidney injury. BMC Nephrol. 2023;24(1):11. 10.1186/s12882-023-03063-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Chertoff J, Chisum M, Garcia B, Lascano J. Lactate kinetics in sepsis and septic shock: a review of the literature and rationale for further research. J Intensive Care. 2015;3:39. 10.1186/s40560-015-0105-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Wu H, Liao B, Cao T, Ji T, Huang J, Ma K. Diagnostic value of RDW for the prediction of mortality in adult sepsis patients: a systematic review and meta-analysis. Front Immunol. 2022;13:997853. 10.3389/fimmu.2022.997853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Lai H, Wu G, Zhong Y, et al. Red blood cell distribution width improves the prediction of 28-day mortality for patients with sepsis-induced acute kidney injury: A retrospective analysis from MIMIC-IV database using propensity score matching. J Intensive Med. 2023;3(3):275–82. 10.1016/j.jointm.2023.02.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Mo M, Huang Z, Huo D, et al. Influence of red blood cell distribution width on all-cause death in critical diabetic patients with acute kidney injury. Diabetes Metab Syndr Obes. 2022;15:2301–9. 10.2147/DMSO.S377650. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Material 1: Checklist. TRIPOD checklist for prediction model development and validation.
Supplementary Material 2: Checklist. STROBE checklist for observational studies.
Data Availability Statement
The data analyzed in this study are available from the Medical Information Mart for Intensive Care IV (MIMIC–IV) and the eICU Collaborative Research Database (eICU–CRD).The processed datasets and analysis code supporting the conclusions of this article are available from the corresponding author upon reasonable request.
