Abstract
Acute Kidney Injury (AKI) is a common complication in hospitalized patients and is associated with increased in-hospital mortality, readmission, and chronic kidney disease. Early identification of patients at risk of AKI non-recovery can improve discharge planning and follow-up. Using a retrospective cohort of 7,667 patient encounters diagnosed with AKI from the University of California San Diego Health, we compared traditional machine learning (ML) and temporal deep learning (DL) models to predict three AKI recovery outcomes: Recovery, Partial Recovery, and Non-Recovery. The ML models evaluated were Logistic Regression, Random Forest, and XGBoost; while, the DL models were Gated Recurrent Unit (GRU) and Long-Short Term Memory. On the test set, DL models consistently outperformed traditional ML approaches. The GRU model achieved the highest Macro-Area Under the Curve (AUC) (0.822) with strong discrimination for the Non-Recovery class (AUC 0.932). This work demonstrates that temporal modeling of clinical trajectories can enhance AKI recovery prediction.
Introduction
Acute Kidney Injury (AKI) is an abrupt reduction in renal function, diagnosed by a rapid rise in serum creatinine (SCr) and/or a decrease in urine output1,2. It occurs in up to one in five hospitalized patients3 and affects up to half of those in intensive care settings4. Beyond its acute phase, incomplete or delayed recovery of kidney function is strongly associated with adverse long-term outcomes including the progression to chronic kidney disease (CKD), kidney replacement therapy (KRT), end stage renal disease (ESRD), and higher mortality5. Multiple determinants have been shown to influence renal recovery including baseline kidney function, fluid balance, the severity and duration of the AKI episode, and comorbidities such as diabetes mellitus, hypertension, and obesity6. In addition, exposure to nephrotoxic medications can further reduce the likelihood of recovery7. Older adults and certain racial minorities have also been shown to be less likely to achieve full renal recovery8,9. As the trajectory of AKI unfolds over the duration of a patient’s hospital stay, timely intervention through the reduction of nephrotoxic drug exposure, hemodynamic stabilization, and coordinated follow-up are essential to mitigate long-term kidney damage and downstream complications10.
Previous machine learning (ML) models have shown promising performance in predicting AKI recovery. However, these studies are often focused on a binary prediction of recovery versus non-recovery and have lacked detailed assessment of exposure to nephrotoxic drugs11–13. In addition, existing models may not capture the complex time-dependent relationship between clinical predictors and AKI recovery. Temporal recurrent neural networks (RNN) have shown strong capabilities in the early detection of AKI onset14,15; however, its application to AKI recovery remains underexplored.
In this study, we investigate whether temporal deep learning (DL) models can improve the prediction of AKI recovery status at hospital discharge compared to traditional ML approaches. Specifically, we aim to predict three recovery outcomes: Recovery, Partial Recovery, and Non-Recovery. This granularity enables improved differentiation of clinically meaningful renal recovery states and identification of patients who may require closer monitoring and individualized care planning during both hospitalization and post-discharge.
Methods
Dataset
This retrospective study utilized routinely collected clinical data from the University of California San Diego Health System. The dataset included patient demographics, comorbidities, laboratory measurements, vital signs, hospital utilization, and medication records for over 100,000 hospital encounters between June 26, 2009 and February 28, 2021. AKI cases were identified using ICD-9 and ICD-10 diagnostic codes. Approval for anonymous analysis of this de-identified data was granted by the Institutional Review Board of the University of California San Diego, with a waiver of informed consent.
Cohort Selection
We included adult patients who develop AKI during their hospital stay as defined by Kidney Disease: Improving Global Outcomes (KDIGO) criteria2. We excluded patients without documented baseline SCr. Additionally, we excluded patients who were recipients of kidney transplants and those with ESRD. To ensure adequate observation of AKI trajectory, we excluded encounters with a hospital length of stay less than two days. Lastly, patients who died during hospitalization were excluded, as recovery status and attribution to AKI-related death could not be reliably determined.
Recovery Outcome Definitions and AKI Staging
The primary prediction goal of this study was to determine a patient’s AKI recovery status at the time of hospital discharge using information available within a prediction window. To ensure predictions provided clinically actionable timing, the prediction window spanned from hospital admission to the day of peak SCr, representing kidney injury progression to its highest severity.
Discharge recovery outcomes were categorized into three groups based on the ratio of discharge SCr to baseline SCr. Recovery was defined as discharge SCr ≤ 110% of baseline, indicating a return to near pre-admission kidney function. Partial Recovery was defined as discharge SCr between 110% and 200% of baseline. Non-Recovery was defined as discharge SCr > 200% of baseline or on dialysis by discharge15.
AKI severity was classified according to KDIGO criteria, which define stages based on the magnitude of increase in SCr relative to baseline. Baseline SCr was determined using the most recent minimum creatinine value on record prior to admission. Urine output criteria were not applied due to data sparsity.
Feature Engineering
For the traditional ML models, we constructed a comprehensive set of window-based aggregated features and domain-specific clinical features to summarize each patient’s AKI trajectory. All predictors were aggregated across the defined prediction window to provide a structured representation for model input. SCr related features included information on the minimum and maximum SCr relative to baseline as well as SCr at admission. Medication related features were derived directly from each patient’s full medication administration record and included cumulative nephrotoxic exposure, stratified by low, medium, and high nephrotoxicity, and by pharmacologic drug class17. Additional clinical descriptors included ICU duration, blood pressure measures, and comorbidities. In total, 82 features were prepared for the ML models. Following feature construction, recursive feature selection was applied to identify the most informative predictors and reduce dimensionality for the ML models.
For the DL models, we preserved the data’s temporal structure by representing patient trajectories at a 24-hour resolution across the same admission to peak SCr window. Each daily time step included raw clinical measurements (i.e., SCr, systolic blood pressure, nephrotoxic medication usage), relative daily changes (i.e., percent change in SCr from previous day and blood pressure variability across three days), and rolling aggregates (i.e., maximum SCr since admission and cumulative nephrotoxic exposure). These temporal representations resulted in 204 input features and enabled the RNNs to capture both abrupt changes and sustained longitudinal trends in AKI progression.
For both ML and DL models, continuous variables were standardized using z-score scaling. Additionally, missing values were imputed using the patient-level median when available, and the population-level median otherwise. Figure 2 illustrates the feature engineering frameworks across the prediction window for both the ML and DL approaches.
Figure 2.
Prediction window for a patient trajectory across both machine learning and deep learning frameworks. Abbreviations: AKI = acute kidney injury; DL = deep learning; ML = machine learning; SCr = serum creatinine.
Model Architectures
Three traditional ML models were selected based on their demonstrated performance in prior AKI recovery prediction studies: LASSO Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost). LASSO Logistic Regression applies an L1 regularization to the conventional regression model. Random Forest is an ensemble learning method that makes classifications based on the majority vote of a large collection of decision trees. XGBoost builds decision trees sequentially, with each new tree trained to correct the residual errors from all previous trees.
For our DL models, we implemented two RNN architectures: Gated Recurrent Units (GRU) and Long-Short Term Memory (LSTM). RNNs are designed to process ordered sequences of data, making them suitable for time-series analysis. The GRU architecture utilizes a two gate mechanism – a reset gate to control how much past information is discarded and an update gate to regulate the information carried forward to the next time steps. The LSTM architecture is more complex with three gates, a forget gate, input gate, and output gate, providing more control over how information is retained and propagated. These gating mechanisms enable the RNNs to preserve long range temporal dependencies while amplifying days that carry informative signals for AKI recovery. The GRU model contained 457,923 total parameters, while the LSTM model contained 610,371 total parameters. Both DL models were trained using an AdamW optimizer and a Focal Loss function. By amplifying the loss of misclassified, low-confidence predictions, Focal Loss concentrates on improving its discrimination of clinically challenging recovery trajectories where temporal patterns may overlap across outcome classes18,19.
To address the class imbalance of the three recovery groups, class weights were applied to each model’s loss function. These weights were computed using the inverse frequency of each recovery category to ensure the minority class contributed proportionally to parameter updates.
Model Training and Hyperparameter Optimization
The dataset was split into 80% training, 10% validation, and 10% testing using stratified sampling to preserve class distributions of the three recovery outcomes. All models were trained exclusively on the training set, with the validation set used for hyperparameter selection. The test set was reserved for final performance evaluation.
Hyperparameters of all five prediction models were tuned through Bayesian optimization20. This strategy utilizes a probabilistic surrogate model to efficiently explore the parameter space and identify configurations that maximize a predefined objective function. For the LASSO Logistic Regression, the primary hyperparameter tuned was the inverse regularization strength. For Random Forest, the number of estimators, maximum depth, maximum number of features, and minimum number of samples required to split a node were optimized. For XGBoost, the learning rate, number of estimators, maximum depth, subsampling ratio, and regularization penalties were optimized. For both DL models, the hyperparameters included the hidden size, number of layers, dropout rate, and learning rate. In order to maximize discriminative performance of the classification models while minimizing overfitting, we formulated the optimization score as:
Evaluation and Interpretability of Prediction Models
Model performance was assessed on a 10% held-out test set. Metrics used to evaluate each models’ multiclass performance include overall accuracy, macro-averaged precision, macro recall, macro F1 score, and macro AUC. Macro-averaged metrics were used to ensure that each recovery outcome contributed equally to evaluation, regardless of its prevalence.
To provide interpretability for each of the classification models, we examined the features that had the greatest impact on model predictions. For the traditional ML models, the top 10 predictors were extracted using each model’s built-in feature importance measures. LASSO Logistic Regression ranks features based on their absolute beta coefficient magnitude. Random Forest calculates feature importance based on average impurity reduction, while XGBoost utilizes the cumulative information gain from each feature split to compute importance. For the DL models, feature contributions were quantified using Shapley Additive exPlanation (SHAP) values21. SHAP values are a model-agnostic method to derive marginal feature contributions across all possible combinations of features. Since the DL models contain a larger feature set compared to the ML models, we report on the top 20 predictors.
To further characterize model behavior across the three recovery outcomes, we analyzed per-class performance using both confusion matrices and multiclass receiver operating characteristic (ROC) curves. Confusion matrices provided a systematic identification of misclassification patterns. In these matrices, diagonal elements represent the recall for each recovery class. Additionally, multiclass ROC curves plot the true positive rate against the false positive rate across all possible classification thresholds. In a multiclass setting, each curve represents the model’s One-vs-Rest discriminative ability to separate each specific recovery class from all other classes. Utilizing these class specific analyses illuminated the model’s strongest and weakest decision boundaries.
Descriptive Analyses and Software
Descriptive statistics were used to summarize the patient characteristics across the Recovery, Partial Recovery, Non-Recovery, and overall cohorts. Continuous variables were reported as mean (standard deviation), and categorical variables as counts (%). All analyses were done using Python 3.10. Data preprocessing was conducted with NumPy 1.26.4. Deep learning model development and training were implemented using TensorFlow 2.20.0.
Results
Distribution of Patient Characteristics Across Recovery Groups
After applying our exclusion criteria, the final cohort consisted of 7,667 AKI hospital encounters. Table 1 summarizes patient characteristics across the AKI recovery groups. The average age was 61.62 years, 53% were white patients, and 60.8% were male. The average length of stay (LOS) was 4.99 days. Stratifying this cohort by recovery status at discharge yielded 3,521 Recovery patients, 3,409 Partial Recovery patients, and 737 Non-Recovery patients.
Table 1.
Distribution of features across AKI Recovery, Partial Recovery, and Non Recovery groups. Continuous variables are presented as mean (standard deviation), while categorical variables as counts (%). SCr min/max 24hr refers to the minimum and maximum serum creatinine values per day. Abbreviations: AKI = acute kidney injury; BMI = body mass index; CAD = coronary artery disease; CKD = chronic kidney disease; eGFR1 EPI = estimated glomerular filtration rate calculated using the 2021 CKD-EPI equation and 24hr minimum serum creatinine; eGFR2 EPI = estimated glomerular filtration rate calculated using the 2021 CKD-EPI equation and 24hr maximum serum creatinine; ICU = intensive care unit; Misc. = miscellaneous; SBP = systolic blood pressure; SCr = serum creatinine.
| Recovery (n=3521) | Partial Recovery (n=3409) | Non-Recovery (n=737) | Total (n=7667) | |
|---|---|---|---|---|
| Age (years) | 62.31 (16.36) | 61.94 (16.08) | 56.91 (15.93) | 61.62 (16.26) |
| Male Gender | 2201 (62.5%) | 2029 (59.5%) | 430 (58.3%) | 4660 (60.8%) |
| Race | ||||
| White | 1905 (54.1%) | 1779 (52.2%) | 381 (51.7%) | 4065 (53.0%) |
| Black | 440 (12.5%) | 406 (11.9%) | 67 (9.1%) | 913 (11.9%) |
| Asian | 232 (6.6%) | 251 (7.4%) | 51 (6.9%) | 534 (7.0%) |
| Other | 944 (26.6%) | 973 (28.5%) | 238 (32.3%) | 2155 (28.1%) |
| Weight (kg) | 80.21 (23.91) | 79.32 (23.61) | 79.42 (23.27) | 79.74 (23.71) |
| Comorbidities | ||||
| Anemia | 1418 (40.3%) | 1380 (40.5%) | 274 (37.2%) | 3072 (40.1%) |
| BMI | 52 (1.5%) | 40 (1.2%) | 3 (0.4%) | 95 (1.2%) |
| CAD | 759 (21.6%) | 800 (23.5%) | 96 (13.0%) | 1655 (21.6%) |
| Cancer | 1217 (34.6%) | 1121 (32.9%) | 260 (35.3%) | 2598 (33.9%) |
| CKD | 1140 (32.4%) | 1066 (31.3%) | 189 (25.6%) | 2395 (31.2%) |
| Diabetes | 1277 (36.3%) | 1285 (37.7%) | 200 (27.1%) | 2762 (36.0%) |
| Liver | 658 (18.7%) | 711 (20.9%) | 181 (24.6%) | 1550 (20.2%) |
| Sepsis | 883 (25.1%) | 761 (22.3%) | 160 (21.7%) | 1804 (23.5%) |
| Vital signs (mmHg) | ||||
| SBP min 24hr | 93.64 (22.00) | 94.09 (21.58) | 89.86 (22.39) | 93.47 (21.89) |
| SBP max 24hr | 149.82 (30.83) | 152.71 (30.64) | 155.68 (30.04) | 151.67 (30.72) |
| Utilization (days) | ||||
| Length of Stay | 4.09 (7.70) | 5.21 (7.12) | 8.32 (13.65) | 4.99 (8.31) |
| ICU length | 1.23 (3.65) | 1.16 (3.20) | 2.56 (11.89) | 1.33 (4.94) |
| Ventilator Support length | 0.57 (2.64) | 0.46 (2.27) | 1.79 (11.48) | 0.64 (4.28) |
| Dialysis length | 0.43 (3.02) | 0.40 (2.75) | 2.68 (10.90) | 0.63 (4.40) |
| Serum Labs | ||||
| Baseline SCr (mg/dL) | 1.87 (1.64) | 1.27 (0.81) | 1.24 (1.08) | 1.54 (1.31) |
| SCr min 24hr (mg/dL) | 2.45 (1.84) | 2.05 (1.39) | 2.89 (2.41) | 2.31 (1.74) |
| SCr max 24hr (mg/dL) | 2.70 (2.13) | 2.19 (1.60) | 3.22 (2.77) | 2.53 (2.02) |
| SCr max at Admission (mg/dL) | 2.49 (2.09) | 1.97 (1.56) | 2.58 (2.81) | 2.27 (1.98) |
| eGFR1 EPI 24hr (mL/min/1.73m2) | 40.74 (24.24) | 47.68 (25.82) | 44.63 (29.20) | 44.20 (25.67) |
| eGFR2 EPI 24hr (mL/min/1.73m2) | 36.57 (21.46) | 44.73 (24.63) | 41.61 (28.91) | 40.68 (24.01) |
| Albumin (g/dL) | 3.37 (0.67) | 3.30 (0.62) | 3.14 (0.60) | 3.32 (0.65) |
| Max AKI Stage before peak | ||||
| Stage I | 2245 (63.8%) | 1848 (54.2%) | 70 (9.5%) | 4163 (54.3%) |
| Stage II | 735 (20.8%) | 971 (28.5%) | 212 (28.8 %) | 1918 (25.0%) |
| Stage III | 541 (15.4%) | 590 (17.3 %) | 455 (61.7%) | 1586 (20.7%) |
| Nephrotoxic Medication Frequency | ||||
| Antivirals | 1.51 (8.08) | 1.64 (10.88) | 1.22 (6.51) | 1.54 (9.31) |
| Diuretics | 6.88 (15.46) | 7.38 (15.46) | 5.60 (11.40) | 6.98 (15.12) |
| Misc. anti-infectives | 1.91 (13.46) | 1.61 (9.93) | 1.57 (10.59) | 1.75 (11.74) |
| Penicillins | 0.73 (3.64) | 0.74 (4.35) | 0.85 (3.41) | 0.75 (3.95) |
| Pressors | 0.06 (0.44) | 0.07 (0.43) | 0.09 (0.40) | 0.07 (0.43) |
| High nephrotoxicity drug | 6.55 (23.50) | 6.65 (19.55) | 8.95 (22.20) | 6.83 (21.71) |
| Medium nephrotoxicity drug | 6.85 (15.00) | 7.79 (19.00) | 5.72 (13.05) | 7.16 (16.75) |
| Low nephrotoxicity drug | 19.27 (41.31) | 19.22 (39.86) | 23.94 (45.10) | 19.70 (41.08) |
Non-Recovery patients were on average younger (56.91 years old) with a slightly higher proportion of female patients (41.7%). They had greater systolic blood pressure (SBP) variability with both the lowest daily minimum and highest maximum SBP measures. In addition, Non-Recovery patients experienced longer LOS (8.32 days), longer ICU stays (2.56 days), and longer durations of dialysis (2.68 days). Nephrotoxic medication exposure also reveals that the Non-Recovery group had higher frequencies of both high and low nephrotoxic usage.
Looking at serum laboratory trends, patients in the Recovery group had the highest baseline SCr (1.87 mg/dL). On average, the Partial Recovery group had the lowest daily SCr minimum and lowest daily SCr maximum. Lastly, maximum AKI severity displayed the strongest gradient. Recovery patients were most likely to remain in AKI Stage I, whereas the majority of Non-Recovery patients had progressed to Stage III.
Feature Importance
Across the three traditional ML models, maximum SCr to baseline ratio, LOS, maximum SCr at admission, and maximum AKI Stage before peak were consistently among the top five predictors. Additionally, nephrotoxic exposure, particularly diuretic usage, and albumin levels also appeared among the top 10 predictors across the ML models. The Logistic Regression model uniquely selected BMI and duration in ICU, while the Random Forest and XGBoost models selected cumulative nephrotoxic drug frequency and SBP.
Figure 3 displays the top 20 predictors for the GRU and LSTM models computed through SHAP values. For both DL architectures, SCr related features are among the top 4 contributors to model predictions. Similarly to the ML models, daily SCr minimum (denoted SCr1) to baseline ratio, maximum AKI stage to date, and diuretic usage were within the top 10 predictors for both DL models. Among the top 20, both GRU and LSTM models selected comorbidity of cancer, SCr1 percent change from the previous day, and medium to low-risk nephrotoxic exposure. The GRU model additionally emphasized coronary artery disease, anemia, race, as well as SBP standard deviation across 3 days, which captures hemodynamic instability. The LSTM approach alternatively selected in-ICU status, high-risk nephrotoxic exposure, and minimum albumin to date.
Figure 3.
Top 20 predictors for the deep learning models computed using SHAP values. Abbreviations: AKI = acute kidney injury; CAD = coronary artery disease; dx = diagnosis; eGFR1 EPI = estimated glomerular filtration rate calculated using the 2021 CKD-EPI equation and 24hr minimum serum creatinine; eGFR2 EPI = estimated glomerular filtration rate calculated using the 2021 CKD-EPI equation and 24hr maximum serum creatinine; ICU = intensive care unit; Nephro = nephrotoxic; SBP = systolic blood pressure; SCr = serum creatinine; SCr1 = 24hr minimum serum creatinine; SCr2 = 24hr maximum serum creatinine; STD = standard deviation.
Held-out Test Performance
Performance metrics on the held-out test set (n=767) are shown in Table 2. Across overall accuracy, macro-averaged precision, recall, F1 score, and AUC, the DL architectures outperformed the traditional ML models for AKI recovery prediction. Among the traditional ML models, XGBoost is the best performing with an accuracy of 0.613, a recall of 0.665, and an AUC of 0.793. This performance is followed by the Random Forest model and then the Logistic Regression model. GRU achieved the strongest performance across the 5 evaluation metrics with an accuracy of 0.640, macro precision of 0.609, macro recall of 0.694, macro F1 score of 0.629, and macro-AUC of 0.822. The LSTM model performed comparably with an accuracy of 0.625, recall of 0.672, and AUC of 0.795. This performance is a slight improvement from XGBoost and exceeds the Logistic Regression and Random Forest approaches.
Table 2.
Classification performance on the held-out test set. The table contains overall accuracy and all other performance metrics are macro-averaged across the three recovery outcomes. Abbreviations: AUC = area under receiver operating characteristic curve.
| Accuracy | Precision | Recall | F1 Score | AUC | |
|---|---|---|---|---|---|
| Logistic Regression | 0.555 | 0.520 | 0.596 | 0.522 | 0.744 |
| Random Forest | 0.593 | 0.555 | 0.643 | 0.570 | 0.776 |
| XGBoost | 0.613 | 0.574 | 0.664 | 0.584 | 0.793 |
| GRU | 0.640 | 0.609 | 0.694 | 0.629 | 0.822 |
| LSTM | 0.625 | 0.588 | 0.672 | 0.601 | 0.795 |
We further evaluated the per-class performance of the top three prediction models (XGBoost, GRU, and LSTM) through multiclass confusion matrices shown in Figure 4 and multiclass ROC curves shown in Figure 5. All models demonstrated strong discrimination for the Non-Recovery class. XGBoost and LSTM both achieved a class-specific recall of 0.81, while the GRU model achieved 0.85. This performance is reflected in the corresponding AUCs for the Non-Recovery class ranging from 0.91 to 0.93 across the three models. For the Recovery group, LSTM achieved the highest recall of 0.74, followed by GRU (0.72) and XGBoost (0.68). AUCs for this class ranged from 0.76 to 0.79. Partial Recovery remained the most challenging class to predict. LSTM reflected the weakest performance for this class with a recall of 0.47 and an AUC of 0.69. XGBoost achieved a recall of 0.50 and AUC of 0.71, and GRU achieved the highest relative performance with a recall of 0.51 and AUC of 0.74.
Figure 4.
Confusion matrices of the top three performing models on the held-out test set.
Figure 5.
Multiclass ROC curves for the top three performing models on the held-out test set. Abbreviations: ROC = receiver operating characteristic; AUC = area under receiver operating characteristic curve.
Discussion
In this study, we developed and evaluated traditional ML and temporal DL frameworks to predict AKI recovery status by time of discharge. Utilizing routinely collected clinical and medication data, we demonstrated that temporal modeling of AKI trajectories improved recovery prediction compared to window-based ML approaches. By integrating nephrotoxic exposures directly from each patient’s complete medication record, our framework captured a rich quantification of therapeutic burden across both pharmacologic classes and toxicity levels. SHAP analyses and multiclass performance assessments further enhanced interpretability by characterizing feature contributions and elucidating class-specific model behavior. Notably, temporal instability emerged as an important signal uniquely highlighted by the deep learning models. Among all five prediction models, the GRU model demonstrated the strongest overall performance, with particularly high discrimination of both the Recovery and Non-Recovery classes and modest gains for the Partial Recovery class. GRU’s superior performance compared to the LSTM model may be attributable to its simpler gating mechanism. The overall reduced model complexity mitigates overfitting and provides a better bias-variance trade off, allowing the GRU model to better generalize to unseen AKI recovery trajectories.
Prior studies in AKI recovery prediction have primarily used Logistic Regression and tree-based decision models to predict binary recovery outcomes. For example, Liu et al. employed an XGBoost model (AUC = 0.807) to predict non-recovery, defined as SCr > 1.5x baseline within 3 days of discharge. Consistent with our findings, their top predictors included baseline SCr, base albumin, and cancer comorbidity11. For an ICU population of 12,321 AKI patients, Zhao et al. applied a Random Forest model (AUC = 0.830), defining non-recovery as a discharge SCr > 30% above baseline12. On a general hospital population of 350,345 encounters, Cho et al. developed a Categorical Boosting model, an alternative decision-tree based method, that achieved an AUC of 0.780. Recovery was defined as a 33% decrease in SCr levels from AKI onset22. These studies demonstrated that AKI recovery can be predicted using structured EHR data; however short prediction horizons limited clinical actionability for predicting recovery at hospital discharge23,24. In contrast, our work introduces a trajectory-based framework that leverages data from hospital admission through peak kidney injury, allowing for earlier risk stratification while the patient’s clinical course is still modifiable. Additionally, our models demonstrate comparable AUC performance relative to these approaches, while extending the prediction task from binary to multiclass recovery outcomes. This extension provides more granular insight into the spectrum of renal recovery, distinguishing patients who may otherwise be overlooked in binary models.
While deep learning methods have been increasingly applied to AKI detection, few have applied these neural network architectures to post-AKI recovery. In a study by Tomašev et al., they developed an RNN-based model for prediction of any AKI within 48 hours, achieving an AUC of 0.921 and outperforming XGBoost (0.889), Random Forest (0.871), and Logistic Regression (0.863)25. Furthermore, Rank et al. demonstrated that an RNN model significantly outperformed clinicians in predicting AKI within the first seven postoperative days following cardiothoracic surgery (AUC = 0.901 vs 0.745)15. Wu et al. further showed that these temporal deep learning models could be utilized to predict downstream post-AKI outcomes. In predicting 7-day death and dialysis outcomes, Wu et al.’s LSTM-based model achieved an AUC of 0.931 and 0.780, respectively26. However, these studies mainly addressed AKI onset or severe downstream complications. Our work addresses this gap by applying temporal RNN architectures to model post-AKI recovery trajectories, ultimately revealing informative, time-dependent predictors and challenges distinct to each recovery class.
Despite these advances, Partial Recovery remains the most elusive outcome to predict. Defined by a discharge SCr between 110% and 200% of baseline, this outcome category is inherently heterogeneous and exhibits overlapping clinical signatures with both Recovery and Non-Recovery. In our cohort, the Partial Recovery group also deviated from expected SCr trajectory patterns, displaying the lowest daily SCr minimum and lowest daily SCr maximum values, despite being an intermediate recovery outcome. This atypical pattern suggests that Partial Recovery lacks a clear linear severity gradient and may reflect a mixture of underlying dysfunction that SCr cannot fully capture alone. Improving discrimination for this group may require the incorporation of additional biomarkers such as serum cystatin C and urine output27.
An additional observation that requires careful interpretation is that patients in the Recovery group had higher baseline SCr and a greater prevalence of preexisting CKD compared to the other two recovery groups. Since the AKI recovery classes were defined relative to baseline SCr, patients with chronically elevated baseline creatinine may meet recovery criteria by returning to their pre-injury renal state, even if overall kidney function remains impaired. Importantly, individuals with higher baseline SCr have a wider absolute margin to meet recovery criteria compared to patients with lower baseline SCr. Therefore, these recovery definitions reflect a return to an individualized baseline reference state rather than restoration of normal, healthy kidney function.
Several limitations should also be acknowledged. First, our study was evaluated exclusively on an internal dataset from a single health system, which may limit generalizability across institutions whose patient population and practice differ. Second, older adults with severe AKI may be underrepresented in our study population. As individuals who died during hospitalization were excluded from analysis, older adults – who experience higher mortality rates during severe AKI episodes8 – were less likely to be captured in the final cohort. Third, our models relied heavily on SCr related features. Although central to AKI staging, SCr may provide an incomplete picture of kidney injury and recovery. Notably, SCr is a delayed marker of kidney deterioration and can be influenced by factors not related to renal function such as muscle mass and catabolic states28. Future work should assess model robustness when SCr measurements are missing or sparsely recorded. Third, nephrotoxic exposure was derived from the frequency and duration of medication administration. Greater specificity through drug dosage measures may provide a more accurate measure of nephrotoxic burden and improve prognostic detail. Lastly, in order for clinical utility and real-time integration to be fully assessed, future work will require prospective validation across diverse hospital settings. In addition, our dataset was derived from a clinical registry which may limit richness of available predictors. Incorporation of the patient’s full EHR and clinical notes via natural language processing may provide greater clinical context to AKI trajectories and further improve recovery prediction performance.
This study underscores the clinical potential of temporal modeling to advance individualized AKI care and management. By identifying patients at risk of incomplete recovery before discharge, such models can inform early consultation, targeted medication adjustments, and post-discharge follow-up. Future work should prioritize external multi-institution validation, integration of complementary biomarkers, and implementation in real-time workflows.
Conclusion
In this work, we demonstrate that temporal deep learning architectures can improve prediction of AKI recovery at the time of hospital discharge compared with traditional machine learning approaches. Importantly, this study extends prior work by moving beyond binary outcomes to characterize a more holistic spectrum of recovery. In addition, our framework systematically incorporates detailed nephrotoxic medication exposures from complete medication records. Although further validation is needed, the approaches presented demonstrate the promising potential for temporal modeling to enhance kidney care.
Figures & Tables
Figure 1.
Study Design. Abbreviations: AKI = acute kidney injury; AUC = area under the receiver operating characteristic curve; DL = deep learning; ML = machine learning; ROC = receiver operating characteristic.
References
- 1.Kellum JA, Romagnani P, Ashuntantang G, Ronco C, Zarbock A, Anders HJ. Acute kidney injury. Nat Rev Dis Primers. 2021 July 15;7(1):52. doi: 10.1038/s41572-021-00284-z. [DOI] [PubMed] [Google Scholar]
- 2.Khwaja A. KDIGO clinical practice guidelines for acute kidney injury. Nephron Clin Pract. 2012 Aug 7;120(4):c179–84. doi: 10.1159/000339789. [DOI] [PubMed] [Google Scholar]
- 3.Susantitaphong P, Cruz DN, Cerda J, et al. World incidence of AKI: a meta-analysis. Clinical Journal of the American Society of Nephrology. 2013 Sept;8(9):1482–93. doi: 10.2215/CJN.00710113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hoste EAJ, Bagshaw SM, Bellomo R, et al. Epidemiology of acute kidney injury in critically ill patients: the multinational AKI-EPI study. Intensive Care Med. 2015 Aug;41(8):1411–23. doi: 10.1007/s00134-015-3934-7. [DOI] [PubMed] [Google Scholar]
- 5.Gameiro J, Marques F, Lopes JA. Long-term consequences of acute kidney injury: a narrative review. Clinical Kidney Journal. 2021 Mar 23;14(3):789–804. doi: 10.1093/ckj/sfaa177. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Forni LG, Darmon M, Ostermann M, et al. Renal recovery after acute kidney injury. Intensive Care Med. 2017 June;43(6):855–66. doi: 10.1007/s00134-017-4809-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Griffin BR, Wendt L, Vaughan-Sarrazin M, et al. Nephrotoxin exposure and acute kidney injury in adults. CJASN. 2023 Feb;18(2):163–72. doi: 10.2215/CJN.0000000000000044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Abdel-Rahman EM, Turgut F, Gautam JK, Gautam SC. Determinants of outcomes of acute kidney injury: Clinical Predictors and Beyond. JCM. 2021 Mar 11;10(6):1175. doi: 10.3390/jcm10061175. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zhao X, Li C, Lu Y, et al. Characteristics and risk factors for renal recovery after acute kidney injury in critically ill patients in cohorts of elderly and non-elderly: a multicenter retrospective cohort study. Renal Failure. 2023 Dec 31;45(1):2166531. doi: 10.1080/0886022X.2023.2166531. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Gameiro J, Fonseca JA, Outerelo C, Lopes JA. Acute kidney injury: from diagnosis to prevention and treatment strategies. JCM. 2020 June 2;9(6):1704. doi: 10.3390/jcm9061704. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Liu CL, Tain YL, Lin YC, Hsu CN. Prediction and clinically important factors of acute kidney injury non-recovery. Front Med. 2022 Jan 17;8:789874. doi: 10.3389/fmed.2021.789874. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zhao X, Lu Y, Li S, et al. Predicting renal function recovery and short-term reversibility among acute kidney injury patients in the ICU: comparison of machine learning methods and conventional regression. Renal Failure. 2022 Dec 31;44(1):1327–38. doi: 10.1080/0886022X.2022.2107542. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Neyra JA, Ortiz-Soriano V, Liu LJ, et al. Prediction of mortality and major adverse kidney events in critically ill patients with acute kidney injury. American Journal of Kidney Diseases. 2023 Jan;81(1):36–47. doi: 10.1053/j.ajkd.2022.06.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Heo S, Kang EA, Yu JY, et al. Time series AI model for acute kidney injury detection based on a multicenter distributed research network: development and verification study. JMIR Med Inform. 2024 July 5;12:e47693–e47693. doi: 10.2196/47693. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Rank N, Pfahringer B, Kempfert J, et al. Deep-learning-based real-time prediction of acute kidney injury outperforms human predictive performance. npj Digit Med. 2020 Oct 26;3(1):139. doi: 10.1038/s41746-020-00346-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Xu J, Xu X, Shen B, et al. Evaluation of five different renal recovery definitions for estimation of long-term outcomes of cardiac surgery associated acute kidney injury. BMC Nephrol. 2019 Dec;20(1):427. doi: 10.1186/s12882-019-1613-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Stottlemyer BA, Abebe KZ, Palevsky PM, et al. Expert consensus on the nephrotoxic potential of 195 medications in the non-intensive care setting: a modified delphi method. Drug Saf. 2023 July;46(7):677–87. doi: 10.1007/s40264-023-01312-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lin TY, Goyal P, Girshick R, He K, Dollár P. Focal loss for dense object detection [Internet] arXiv. 2017 [cited 2025 Nov 30]. Available from: https://arxiv.org/abs/1708.02002 . [DOI] [PubMed]
- 19.Loshchilov I, Hutter F. Decoupled weight decay regularization [Internet] arXiv. 2019 [cited 2025 Nov 30]. Available from: http://arxiv.org/abs/1711.05101 .
- 20.Wu J, Chen XY, Zhang H, Xiong LD, Lei H, Deng SH. Hyperparameter optimization for machine learning models based on Bayesian optimization
- 21.Lundberg S, Lee SI. A unified approach to interpreting model predictions [Internet] arXiv. 2017 [cited 2025 Nov 30]. Available from: https://arxiv.org/abs/1705.07874 .
- 22.Cho NJ, Jeong I, Kim Y, et al. A machine learning-based approach for predicting renal function recovery in general ward patients with acute kidney injury. Kidney Res Clin Pract. 2024 July 31;43(4):538–47. doi: 10.23876/j.krcp.23.330. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Huang CY, Güiza F, Wouters P, et al. Development and validation of the creatinine clearance predictor machine learning models in critically ill adults. Crit Care. 2023 July 6;27(1):272. doi: 10.1186/s13054-023-04553-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Ghanbari G, Lam JY, Shashikumar SP, et al. Development and validation of a deep learning algorithm for the prediction of serum creatinine in critically ill patients. JAMIA Open. 2024 July 1;7(3):ooae097. doi: 10.1093/jamiaopen/ooae097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Tomašev N, Glorot X, Rae JW, et al. A clinically applicable approach to continuous prediction of future acute kidney injury. Nature. 2019 Aug 1;572(7767):116–9. doi: 10.1038/s41586-019-1390-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Wu C, Zhang Y, Nie S, et al. Predicting in-hospital outcomes of patients with acute kidney injury. Nat Commun. 2023 June 22;14(1):3739. doi: 10.1038/s41467-023-39474-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Kashani K, Rosner MH, Ostermann M. Creatinine: from physiology to clinical application. European Journal of Internal Medicine. 2020 Feb;72:9–14. doi: 10.1016/j.ejim.2019.10.025. [DOI] [PubMed] [Google Scholar]
- 28.Fu Z, Deng Y. Assessment, molecular mechanisms and therapeutic targets for renal functional reserve. Renal Failure. 2025 Dec 31;47(1):2526686. doi: 10.1080/0886022X.2025.2526686. [DOI] [PMC free article] [PubMed] [Google Scholar]





