Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Nov 28;15:42631. doi: 10.1038/s41598-025-26774-8

Development and validation of an interpretable predictive machine learning model for successful weaning of continuous renal replacement therapy

Benjamin Popoff 1,8,✉, Boris Delange 1,3, Jonathan Nicolas 4, Arthur Le Gall 5, Badisse Dahamna 7, Marc Cuggia 1, Thomas Clavier 2,6, Guillaume Bouzillé 1
PMCID: PMC12663152  PMID: 41315416

Abstract

Continuous renal replacement therapy (CRRT) is a vital intervention for critically ill patients with severe acute kidney injury, yet no standardized criteria exist to determine the optimal time for its discontinuation. We developed and validated machine learning models to predict successful CRRT weaning, defined as survival without any form of renal replacement therapy for at least seven days after discontinuation. This retrospective multicenter study used data from two French university hospitals and the publicly available MIMIC-IV critical care database. Predictive variables were selected from routinely collected clinical and biological data to ensure real-world applicability. Models were trained on the Rouen cohort and externally validated on the Rennes and MIMIC-IV cohorts. Among the tested algorithms, the random forest model achieved the best performance, with an area under the receiver operating characteristic curve (AUROC) of 0.86 (95% CI, 0.82–0.91) in the training cohort, 0.81 (95% CI, 0.71–0.90) in the Rennes cohort, and 0.72 (95% CI, 0.65–0.78) in the MIMIC cohort. These results demonstrate the feasibility of a robust and interpretable prediction model that relies solely on routinely available data and has potential for integration into clinical workflows to support CRRT weaning decisions.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-26774-8.

Keywords: Renal replacement therapy, Acute kidney injury, Machine learning, Critical care, Clinical decision support system

Subject terms: Computational biology and bioinformatics, Diseases, Health care, Medical research, Nephrology

Introduction

Acute kidney injury (AKI) is a severe condition affecting over half of patients admitted to the intensive care unit (ICU)1. In severe cases, renal replacement therapy (RRT) is a critical and life-saving therapeutic intervention. Among the various modalities of RRT, continuous renal replacement therapy (CRRT) is the most frequently used in critically ill patients. Unlike intermittent hemodialysis, CRRT provides continuous 24-h support, making it more suitable for hemodynamically unstable patients2. Despite this, mortality among AKI patients treated with CRRT remains high, ranging between 40 and 60%3,4, and the dependence on RRT at 90 days is estimated to be between 16 and 29%5.

To date, most studies have focused on identifying the optimal timing for initiating CRRT in AKI patients6–8. However, the equally critical question of when and how to wean patients from CRRT remains unresolved9. Determining the appropriate timing for CRRT discontinuation is crucial. Premature weaning can lead to the recurrence of renal dysfunction and associated complications, such as fluid overload and electrolyte imbalances, often requiring the reintroduction of renal support10,11. Conversely, prolonged CRRT increases the risks of catheter-related complications, infections, unnecessary resource utilization, extended ICU stays and delayed renal recovery12. Addressing these challenges by identifying reliable criteria and developing decision-support tools for successful weaning from CRRT is an essential and ongoing area of research.

Given the complexity of CRRT weaning and its impact on patient outcomes, developing a clinical decision support system (CDSS) to assist intensivists in this critical decision is a promising approach. Advances in artificial intelligence (AI), particularly machine learning (ML), have enabled the creation of numerous CDSS tools in the ICU, targeting areas such as sepsis prediction, early detection of patient deterioration, and management of mechanical ventilation13–16. However, the deployment of AI-based CDSS in critical care remains limited by a lack of rigorous evaluation, with most articles focusing on model development rather than external validation or prospective evaluation14. A review of 521 devices authorized by the U.S. Food and Drug Administration (FDA) for AI/ML applications identified only a handful applicable to ICU patients, with most relying on minimal clinical evidence, limited safety assessments, and no evaluation of performance bias. This underscores the need for robust validation and implementation strategies to ensure the reliability and safety of AI tools in high-stakes settings like the ICU.

The development of mature AI models for routine ICU care is hindered at various stages of the CDSS development cycle. Recent frameworks have been proposed to ensure a more rigorous development methodology, emphasizing the importance of focusing on the clinical question the model aims to address and align with clinicians’ expectations for AI-based CDSS in practice17,18. Involving healthcare providers at every stage of development and implementation through a user-centered design approach has been identified as a key strategy to overcome these barriers and enhance the applicability of CDSS in critical care19,20.

To address these challenges, we adopted a stepwise approach for the development and evaluation of a CDSS for CRRT weaning. In a pre-development phase, we conducted a survey among ICU physicians to identify their needs, expectations, and concerns regarding AI-based CDSS, as well as the clinical variables they consider most relevant for decision-making in the context of CRRT weaning21. Based on these insights, we developed a predictive machine learning model to assist with CRRT weaning decisions, focusing on routinely collected clinical and biological variables to ensure practicality and usability in real-world settings. In this study, we present the development and validation of this model, emphasizing its robust evaluation across three distinct clinical contexts.

The objective of this study was to develop and validate a user-centered predictive model for successful CRRT weaning, defined as RRT-free survival within seven days of discontinuation. This approach aims to integrate key clinical and biological variables, providing clinicians with a robust, AI-based tool to optimize decision-making while addressing the practical challenges of CDSS development in critical care.

Results

Patient characteristics

A total of 1,133 (Rouen), 389 (Rennes) and 1,246 (BIDMC) patients were admitted to the ICU and received CRRT. After applying eligibility criteria, 941 patients remained for analysis: 382 in the Rouen cohort, 166 in the Rennes cohort, and 393 in the MIMIC cohort (Fig. 1). The included patients had a median CRRT duration of 4.0 (IQR, 3.0–6.0) days in the Rouen cohort, 5.0 (IQR, 3.0–9.0) in the Rennes cohort and 5.0 (IQR, 3.0–8.0) in the MIMIC cohort, with a CRRT weaning success rate of 48.2% (184/382), 35.5% (59/166) et 44.0% (173/393) respectively. Table 1 summarizes baseline characteristics of included patients. Detailed information in each cohort patients stratified by CRRT weaning status are presented in Tables S1, S2 and S3.

Fig. 1.

Fig. 1

Study flowchart. Abbreviations: CRRT, continuous renal replacement therapy; MIMIC-IV, Medical Information Mart for Intensive Care-IV.

Table 1.

Patients baseline characteristics.

Rouen
(n = 382)
Rennes
(n = 166)
MIMIC
(n = 393)
Successful weaning 184 (48.2%) 59 (35.5%) 173 (44.0%)
Age, years 64 (65–71) 65 (54 – 72) 61 (51, 70)
Gender
- Male 252 (66.0%) 110 (66.3%) 230 (58.5%)
- Female 130 (34.0%) 56 (34.7%) 163 (41.5%)
BMI, kg/m2 28.0 (25.2–32.6) 28.0 (24.1–33.4) 31.1 (26.3–37.4)
SAPS II 57.0 (43.0–69.0) 62.0 (48.0–73.0) 56.0 (46.0–65.0)
Comorbidities
- Myocardial infarction 12 (3.1%) 5 (3.0%) 62 (15.8%)
- Congestive heart failure 87 (22.8%) 26 (15.7%) 125 (31.8%)
- Peripheral vascular disease 32 (8.4%) 20 (12.0%) 36 (9.2%)
- Cerebrovascular disease 20 (5.2%) 14 (8.4%) 39 (9.9%)
- Dementia 2 (0.5%) 1 (0.6%) 2 (0.5%)
- Chronic pulmonary disease 32 (8.4%) 15 (9.0%) 106 (27.0%)
- Rheumatic disease 7 (1.8%) 3 (1.8%) 16 (4.1%)
- Peptic ulcer disease 19 (5.0%) 10 (6.0%) 17 (4.3%)
- Liver disease 47 (12.3%) 36 (21.7%) 193 (49.1%)
- Diabetes mellitus 127 (33.2%) 53 (31.9%) 155 (39.4%)
- Hemi or paraplegia 18 (4.7%) 20 (12.0%) 6 (1.5%)
- Chronic kidney disease 53 (13.9%) 26 (15.7%) 139 (35.4%)
- Malignant cancer 63 (16.5%) 24 (14.5%) 71 (18.1%)
- Hypertension 223 (58.4%) 82 (49.4%) 221 (56.2%)
Admission type
- Medical 183 (47.9%) 92 (55.4%) 201 (51.1%)
- Surgical 199 (52.1%) 74 (44.6%) 92 (23.4%)
- Medical/Surgical 0 (0%) 0 (0%) 100 (25.4%)
Baseline serum creatinine, μmol/L 86.3 (69.0–88.6) 87.3 (71.5–111.0) 70.7 (44.2–97.2)
Admission serum creatinine, μmol/L 206.5 (122.0–332.5) 210.0 (104.0–358.0) 252.0 (159.1–389.0)
Hospital LOS, days 38.4 (14.8–82.1) 36.0 (20.0–56.0) 22.0 (14.0–35.0)
ICU LOS, days 14.0 (7.0–27.0) 24.0 (13.0–38.0) 11.0 (7.0–17.0)
CRRT duration, days 4.0 (3.0–6.0) 5.0 (3.0–9.0) 5.0 (3.0–8.0)
ICU mortality 95 (24.9%) 42 (25.3%) 87 (22.1%)

Data presented as n (%) or median (Q1–Q3).

Abbreviations: BMI, body mass index; CRRT, continuous renal replacement therapy; ICU, intensive care unit; LOS, length of stay; MIMIC, Medical Information Mart for Intensive Care; SAPS II, Simplified Acute Physiology Score II.

Feature selection

A total of 111 features were explored for the development of the models. These variables are routinely collected in the ICUs electronic health records, including demographic data such as age, gender, SAPS II score, Charlson comorbidity index, as well as weight, height, and admission type. Biological variables include creatinine, urea, sodium, potassium, and other parameters measured at different time points. Clinical variables cover hourly urine output, body temperature, daily weight, among others. Finally, therapeutic variables consist of information on mechanical ventilation, as well as the administration of catecholamines and diuretics. These variables were selected due to their systematic availability in clinical practice and their relevance in predicting the success of CRRT weaning.

After performing LASSO feature selection on the training cohort, the variables retained for the final model included urine output at t0 (in mL/kg/h), invasive mechanical ventilation at t0, serum phosphate level at t-1, daily median body temperature at t0, white blood cell (WBC) count at t-1, urine output at t-1 (in mL/kg/h), WBC count at t0, prothrombin time (PT) at t-1, norepinephrine infusion at t0, and serum bicarbonate level at t0. t0 represents the day of CRRT discontinuation, and t-1 refers to the day prior.

Performance of the models

After training and optimizing the hyperparameters, the models achieved the following AUROC performances on the Rouen cohort using tenfold cross validation: 0.86 (95% CI, 0.82–0.91) for RF, 0.86 (95% CI, 0.82–0.90) for XGB, 0.84 (95% CI, 0.80–0.88) for LR, 0.84 (95% CI, 0.80–0.88) for SVM, and 0.81 (95% CI, 0.78–0.84) for KNN, using an optimized classification thresholds of 0.476, determined by maximizing Youden’s index.

The models demonstrated robust performance on both external validation cohorts. On the Rennes cohort AUROC values were 0.82 (95% CI, 0.74–0.91) for RF, 0.81 (95% CI, 0.71–0.90) for XGB, 0.78 (95% CI, 0.68–0.88) for KNN, 0.74 (95% CI, 0.66–0.84) for LR, and 0.68 (95% CI, 0.58–0.79) for SVM. Similarly, on the MIMIC cohort, models achieved AUROC values of 0.73 (95% CI, 0.66–0.79) for RF, 0.72 (95% CI, 0.65–0.78) for XGB, 0.68 (95% CI, 0.61–0.73) for KNN, 0.71 (95% CI, 0.64–0.77) for LR, and 0.70 (95% CI, 0.63–0.77) for SVM, using the same classification thresholds. Detailed results are presented in Table 2. Results with the default classification threshold of 0.5 are presented in Table S4.

Table 2.

Models performance metrics for prediction of successful continuous renal replacement therapy weaning with optimized classification thresholds.

Models AUROC F1 score Accuracy Precision Recall Specificity MCC

Rouen

(tenfold cross-validation)

RF

0.86

(0.82–0.91)

0.81

(0.68–0.88)

0.79

(0.66–0.87)

0.77

(0.66–0.89)

0.85

(0.70–0.95)

0.79

(0.69–0.94)

0.58

(0.39–0.77)

XGB

0.86

(0.82–0.90)

0.81

(0.71–0.88)

0.81

(0.71–0.87)

0.83

(0.75–0.93)

0.80

(0.71–0.88)

0.81

(0.69–0.94)

0.62

(0.44–0.74)

LR

0.84

(0.80–0.88)

0.77

(0.65–0.86)

0.76

(0.62–0.86)

0.77

(0.63–0.88)

0.78

(0.66–0.89)

0.74

(0.56–0.88)

0.53

(0.24–0.72)

SVM

0.84

(0.80–0.88)

0.77

(0.73–0.81)

0.75

(0.71–0.79)

0.75

(0.70–0.80)

0.80

(0.76–0.83)

0.70

(0.62–0.78)

0.50

(0.42–0.59)

KNN

0.81

(0.78–0.84)

0.75

(0.61–0.82)

0.73

(0.60–0.81)

0.72

(0.62–0.83)

0.78

(0.60–0.89)

0.68

(0.59–0.82)

0.46

(0.20–0.63)

Rennes RF

0.82

(0.74–0.91)

0.83

(0.75–0.90)

0.78

(0.69–0.85)

0.81

(0.73–0.93)

0.85

(0.76–0.95)

0.65

(0.45–0.83)

0.51

(0.32–0.70)

XGB

0.81

(0.71–0.90)

0.84

(0.77–0.9)

0.79

(0.71–0.88)

0.83

(0.75–0.93)

0.85

(0.76–0.95)

0.68

(0.54–0.85)

0.54

(0.34–0.73)

LR

0.74

(0.66–0.84)

0.78

(0.70–0.84)

0.69

(0.61–0.77)

0.72

(0.61–0.80)

0.85

(0.77–0.94)

0.40

(0.24–0.57)

0.27

(0.09–0.48)

SVM

0.68

(0.58–0.79)

0.41

(0.27–0.51)

0.51

(0.41–0.62)

0.88

(0.72–1.0)

0.27

(0.17–0.36)

0.93

(0.83–1.00)

0.25

(0.07–0.36)

KNN

0.78

(0.68–0.88)

0.81

(0.75–0.88)

0.74

(0.66–0.82)

0.76

(0.67–0.87)

0.87

(0.78–0.95)

0.51

(0.34–0.68)

0.41

(0.20–0.60)

MIMIC RF

0.73

(0.66–0.79)

0.64

(0.56–0.73)

0.64

(0.58–0.70)

0.71

(0.61–0.80)

0.59

(0.51–0.69)

0.70

(0.61–0.79)

0.29

(0.18–0.41)

XGB

0.72

(0.65–0.78)

0.68

(0.61–0.74)

0.67

(0.61–0.73)

0.72

(0.64–0.81)

0.65

(0.57–0.72)

0.70

(0.61–0.78)

0.34

(0.23–0.45)

LR

0.71

(0.64–0.77)

0.61

(0.52–0.68)

0.63

(0.57–0.69)

0.72

(0.62–0.82)

0.53

(0.45–0.61)

0.75

(0.66–0.85)

0.28

(0.16–0.40)

SVM

0.70

(0.63–0.77)

0.14

(0.07–0.21)

0.48

(0.42–0.55)

0.89

(0.67–1.00)

0.07

(0.04–0.12)

0.99

(0.97–1.00)

0.15

(0.05–0.22)

KNN

0.68

(0.61–0.73)

0.6

(0.54–0.66)

0.61

(0.55–0.66)

0.69

(0.59–0.77)

0.53

(0.46–0.61)

0.71

(0.61–0.78)

0.24

(0.12–0.33)

Models performance across the three cohorts are presented in Fig. 2 with ROC curves (2A) highlighting their discriminative ability and calibration curves (2B) assessing the agreement between predicted and observed outcomes. Figure 3 displays the confusion matrices for the test and external validation cohorts.

Fig. 2.

Fig. 2

Models performance for prediction of successful continuous renal replacement therapy weaning: receiver operating characteristic (ROC) curves (A) and calibration curves (B). Abbreviations: KNN, k-nearest neighbors; LR, logistic regression; MIMIC, Medical Information Mart for Intensive Care; RF, random forest; SVM, support vector machine; XGB, eXtreme gradient boosting.

Fig. 3.

Fig. 3

Confusion matrices of the models with optimized classification thresholds on the Rennes (A) and MIMIC (B) cohorts. Values in parentheses represent 95% confidence intervals, calculated using normal approximation for cross-validation performance assessment on the training cohort, or using bootstrapping otherwise. Abbreviations: AUROC, area under the receiver operating characteristic curve; KNN, k-nearest neighbors LR, logistic regression; MCC, Matthews correlation coefficient; MIMIC, Medical Information Mart for Intensive Care; RF, random forest; SVM, support vector machine; XGB, eXtreme gradient boosting.

Model interpretations

With the XGB and RF models achieving the highest AUROC values among the tested models, their explainability was further explored using the SHAP method (Fig. 4). A positive SHAP value indicates that the feature increases the probability of CRRT weaning, while a negative value suggests weaning failure. To further illustrate individual level explainability, Fig. 5 presents SHAP force plots for two representative patients, one correctly predicted as successfully weaned and another as not successfully weaned, highlighting the contributions of key features to the model’s predictions.

Fig. 4.

Fig. 4

SHapley Additive exPlanation (SHAP) value of the Random Forest (A) and XGBoost (B) models on the testing cohort (Rennes).

Fig. 5.

Fig. 5

Individual SHapley Additive exPlanation (SHAP) force plot for random forest model predictions of continuous renal replacement therapy (CRRT) weaning success: (A) a patient correctly predicted as successfully weaned from CRRT, and (B) a patient correctly predicted as not successfully weaned. Positive SHAP values represent features contributing to the prediction of successful weaning, while negative values reflect features contributing to the prediction of weaning failure. These plots demonstrate the impact of key variables on the model’s decision-making for each patient. Abbreviations: WBC, white blood cells.

The variables with the greatest influence on the models were urine output on the day and the day before weaning, the presence of invasive mechanical ventilation, and the median body temperature for XGB model. Figure S1 presents partial dependent plots which illustrate the marginal effect of selected features on the predicted probability of successful CRRT weaning, providing insights into how changes in individual variables influence model predictions. The interpretation of SHAP values indicates that higher urine output, discontinuation of mechanical ventilation and norepinephrine support, and lower WBC counts are associated with an increased likelihood of successful CRRT weaning.

The consistency of feature contributions across centers was assessed in the Random Forest model (Supplementary Figure S2), which showed similar patterns of feature importance, with urine output remaining the strongest predictor and PT contributing minimally in all cohorts. Supplementary Table S5 details the distribution of the ten selected features across the three cohorts, showing overall comparable values, although fewer patients were mechanically ventilated in the MIMIC cohort compared with the French centers.

Variables are ranked from the most important at the top to the less important at the bottom. Variable influence on continuous renal replacement therapy (CRRT) weaning prediction is represented from the right (positive) to left (negative) and the value of the observation is colored from blue (lowest value for continuous variables and "no” for binary variables) to green (highest value for continuous variables and "yes” for binary variables). SHAP values for the RF model are computed in probabilities, whereas for the XGB model, they are computed in log-odds. Abbreviations: WBC, white blood cells.

Discussion

Summary of findings

This study presents the development and evaluation of machine learning models designed to predict the success of weaning from CRRT using routinely collected clinical data from the 48 h preceding its interruption. The model underwent external validation on two distinct cohorts: one from another French hospital and another from the publicly available MIMIC-IV database, which includes patients from a U.S. hospital. This dual validation allowed for the assessment of the model’s robustness in two different contexts. The French cohort provided a validation dataset from a center within the same country as the training cohort, with a more similar clinical environment and patient population. In contrast, the MIMIC cohort represented a significantly different clinical context, with variations in practices, data structure, and patient characteristics. The model demonstrated strong performance on the training cohort and maintained good robustness on the external validation cohorts, despite minor performance declines. This highlights the potential of the model for generalization across varying clinical settings.

Ten routinely available features were selected, insuring interpretability and practicality of the final models. These variables were identified based on their relevance in clinical practice, as highlighted by clinicians in a prior survey, and the total number of selected features aligns with the number they suggested as optimal for usability. The selected features include commonly available parameters such as urine output, invasive ventilation, and serum phosphate levels21. By focusing on routine, easily accessible variables, the model simplifies its integration into clinical workflows and enhances its potential for adoption in real-world settings.

The involvement of clinicians throughout the CDSS development cycle is essential to avoid common pitfalls of developing models that are technically robust but poorly suited to clinical needs or workflows22. By incorporating clinician feedback during feature selection and model design, the resulting tool aligns more closely with real-world requirements. This aligns with the principles of a Learning Health System (LHS), where continuous feedback from end-users is leveraged to refine and adapt the system over time23. An LHS approach not only enhances the practical value of the CDSS but also fosters clinician engagement and trust in the tool, further supporting its implementation in diverse clinical environments.

Furthermore, this study demonstrates the feasibility of developing and validating a predictive model using data from multiple centers with varying data structures, thanks to the adoption of the OMOP common data model. The structural and semantic mapping of data to OMOP facilitated the harmonization of datasets from different sources, ensuring consistency in data preparation and enabling robust analysis24. This approach underscores the value of standardized data models in the development of CDSS. While OMOP was used in this study, the Fast Healthcare Interoperability Resources (FHIR) standard could also represent a viable alternative, offering additional interoperability features for real-time clinical integration25.

Comparison with prior works

Several predictive models for CRRT weaning have been recently developed and published26–31. Definitions of successful CRRT weaning vary widely across studies, ranging from no CRRT resumption within 72 h30 to CRRT-free survival at hospital discharge27. The most widely accepted definition of successful CRRT weaning is CRRT-free survival for at least seven days, as used in our study28,29.

Our RF model’s performance, with an AUC of 0.86 (95% CI, 0.82–0.91), 0.81 (95% CI, 0.71–0.90) and 0.72 (95% CI, 0.65–0.78) on the three cohorts respectively, aligns with previously reported models, such as Sheng et al. (AUC 0.87, 95% CI 0.85–0.89, KNN) and Zhong et al. (AUC 0.80, 95% CI 0.75–0.85, XGBoost). However, unlike many studies that rely solely on MIMIC for model development and validation, we used MIMIC exclusively for external validation, alongside a French cohort, making this the only study to validate across two distinct clinical contexts.

A key feature across previously published models is the consistent identification of urine output as one of the most critical predictors of successful CRRT weaning. This aligns with the findings of our study, where urine output at both t0 and t-1 were among the most influential predictors in our model. Unlike previous studies, our approach was informed by a prior survey of ICU physicians, allowing us to identify both the most relevant predictive variables and the maximum number of features they deemed acceptable for practical use. By incorporating this user-centered design, we ensured that our model aligns with clinical needs and workflow constraints, making it more likely to be adopted in real-world practice. A major strength of our approach is the use of only ten variables that can be automatically extracted, ensuring practicality and scalability. In contrast, some models require extensive features (e.g., 90 variables in Wang et al.27) or biomarkers and specialized tests, such as the furosemide stress test31, which limit their applicability.

Limitations

This study has several limitations. First, the retrospective nature of the data collection introduces inherent biases and limits the ability to control for unmeasured confounders. Additionally, the harmonization of datasets required extensive efforts to align terminologies and structures, particularly between the French cohorts and the MIMIC database, which may have introduced variability.

Second, differences in CRRT weaning practices across centers were evident, such as the lower weaning success rate observed in Rennes. This discrepancy is likely due to a local preference for transitioning patients to intermittent hemodialysis following CRRT discontinuation, which was considered a weaning failure according to our definition. Despite these differences, the model demonstrated robust performance and adaptability, suggesting its potential applicability across diverse clinical environments.

Third, the heterogeneity between the French datasets and MIMIC posed specific challenges that likely contributed to the observed performance drop in the U.S. cohort. Differences in clinical context, coding practices, and patient case-mix may have affected model generalizability. For instance, the MIMIC cohort included a higher proportion of patients with chronic kidney and liver disease, as determined from administrative billing codes, which may not accurately reflect the true clinical burden and could modify the recovery dynamics of acute kidney injury, even though baseline serum creatinine levels were similar across cohorts. Moreover, fewer patients were mechanically ventilated in the MIMIC cohort compared with the French centers, suggesting potential differences in case severity or local management strategies. Another key challenge was the difference in PT calculation between French centers and MIMIC, which may have affected the predictive contribution of this variable, although SHAP analyses indicated that PT was among the least influential predictors in all cohorts. Therapeutic limitations, explicitly documented in the French datasets due to national regulations, could not be extracted from MIMIC as this information is unstructured. Despite these challenges, the adoption of the OMOP common data model enabled substantial standardization, allowing for robust multicenter and international validation.

Furthermore, the sample sizes, particularly in the Rennes cohort, were relatively modest for machine learning applications, this may have contributed to the wider confidence intervals observed in model performance estimates. To limit overfitting and enhance reliability, we used parsimonious algorithms, a small number of predictive variables, and robust resampling techniques, including tenfold cross-validation. The modest sample size also underscores the need for further evaluation of model stability and generalizability in larger, prospective real-world settings. A prospective observational study is currently planned to assess the model’s clinical performance in routine practice.

In addition, while several predictors identified by the model, such as urine output, discontinuation of mechanical ventilation, or norepinephrine use, are potentially modifiable, this study was purely predictive and does not establish causality. Future work using causal inference frameworks, such as target trial emulation, could assess whether interventions on these variables directly improve CRRT weaning outcomes, thereby bridging the gap between prediction and causal understanding32.

Finally, patients requiring CRRT constitute a heterogeneous population with diverse etiologies, comorbidities, and recovery trajectories. This study specifically included patients treated with CRRT for more than 48 h who survived long enough to undergo a first weaning attempt, representing a well-defined subset of all CRRT-treated individuals. While this selection was necessary to address the study objective, it limited the possibility of identifying distinct patient subgroups with different recovery kinetics. Future studies involving larger populations could explore clustering or subphenotyping approaches to better characterize these subgroups and refine predictive performance.

Future directions

This study is the second step of a broader initiative aimed at developing a clinically relevant CDSS for CRRT weaning. It builds on a first step pre-development phase conducted by our research group, which assessed the needs and expectations of ICU physicians, ensuring that the model addresses practical clinical challenges. The next and third step, which is currently underway, consists of an observational evaluation of the model within the Rouen center with clinicians blinded to the results, allowing for real-world validation without influencing clinical decisions. If this third step is conclusive, the fourth step will be a randomized study exploring the potential benefits of our CDSS for CRRT weaning (reduction of CRRT length and ICU length of stay, improvement of the renal function after ICU discharge, etc.).

While the current model focuses on short-term CRRT weaning success, future research could extend this approach to predict long-term renal recovery after ICU discharge. Integrating longitudinal outcomes would provide complementary insights into post-ICU kidney trajectories and strengthen the model’s clinical relevance across the continuum of care. Similarly, the potential value of using earlier time windows to anticipate weaning readiness could be explored in future research, supporting more proactive clinical decision-making.

In a future implementation, the model will be enhanced with graphical tools like the SHAP-based force plots presented in Fig. 5. At each step, we will collect users’ feedback to optimize its interpretability and ease of use. This process appears essential in the development of CDSS, which goes far beyond the simple creation of a statistically satisfying algorithm.

This global approach aligns with modern guidelines for the development of CDSS, emphasizing iterative design, user-centered development, and rigorous evaluation17,34. By adhering to these principles, this tool has the potential to become a valuable resource for intensivists managing CRRT weaning in diverse clinical settings.

In conclusion, this study demonstrates the development and validation of a robust machine learning model for predicting successful CRRT weaning, defined as RRT-free survival for at least seven days following discontinuation. The model was validated across two distinct clinical contexts (highlighting its generalizability and robustness despite variations in clinical practices. Its reliance on only 10 commonly available variables enhances its practicality for real-world integration and is in concordance with clinicians’ expectations. Building on this work, a prospective evaluation phase is planned to assess its clinical utility, aligning with modern guidelines for CDSS development and ensuring its relevance.

Methods

Study design

This retrospective, multicenter study used data from tertiary care ICUs of two French university hospitals (Rouen and Rennes) and the publicly available Medical Information Mart for Intensive Care-IV (MIMIC-IV) database version 2.2, containing ICU data from the Beth Israel Deaconess Medical Center (BIDMC) in Boston, Massachusetts, U.S.A.35. This article follows the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis Or Diagnosis – Artificial Intelligence (TRIPOD + AI) statement36 (Supplementary materials).

Ethics

Ethical approval was granted by the scientific and ethical boards of Rennes (registration 24.18) and Rouen (CSE2024_DA054) hospitals. In accordance with French law and GDPR regulations, patient consent was not required, rather patients were informed of their right to oppose to the use of their data37. MIMIC patient records were de-identified, with data collection approved by the BIDMC and Massachusetts Institute of Technology Institutional Review Boards, waiving the requirement for individual patient informed consent.

Participants

All adult patients ICU who underwent CRRT assessed for eligibility. Inclusion criteria were CRRT for more than 48 h and a first discontinuation attempt. Exclusion criteria were: CRRT duration of less than 48 h, history of chronic dialysis, perioperative context of cardiac surgery, withdrawal of life-sustaining treatments regarding dialysis, death during the observation period, or missing data on urine output during the observation period.

Primary outcome

The primary outcome was the successful CRRT weaning. There is no universally accepted definition of weaning, and the international Kidney Disease Improving Global Outcome (KDIGO) guidelines define it as the “discontinuation of RRT when it is no longer required, either because intrinsic kidney function has sufficiently recovered to meet patient needs or because RRT is no longer aligned with the goals of care”38. Given the lack of explicit detail in this definition, the primary outcome in this study was defined as patients being alive and free from any form of RRT for at least 7 days following discontinuation. This definition was chosen for its widespread use and practical interpretability28,29,39,40.

Data extraction

The training dataset (Rouen University Hospital) was collected from June 2016 to July 2024 using the EDSaN (Entrepôt de Données de Santé Normand) clinical data warehouse (CDW) and ICU software ICCA© (IntelliSpace Critical Care and Anesthesia©, Koninklijke Philips N.V.)41.

Two external validation cohorts were used to assess generalizability of the models. The first validation cohort consisted of data from another French center, Rennes University Hospital, collected between 2020 and December 2023 from the eHOP CDW and the Metavision© software (iMDsoft Ltd.)42. The second validation cohort was extracted from the publicly available MIMIC-IV database, containing ICU data from the BIDMC in Boston, USA, spanning 2008 to 2019. The MIMIC data, stored in PostgreSQL format, were extracted using Structured Query Language (SQL) queries from the MIMIC repository and then processed and standardized to align with the other datasets43.

The final datasets included clinical, therapeutic, and biological data. To ensure interoperability between the databases, the data were transformed into the OMOP (Observational Medical Outcomes Partnership) version 5.4 format, and medical concepts were standardized using international standard terminologies24. All extracted data were deidentified. Ultimately, three datasets in the same format were created, enabling distributed learning and evaluation. We did not undertake any prior power calculations and used all the data available.

Data preprocessing

The distribution of continuous variables was assessed graphically and extreme values beyond physiological plausibility were removed. Continuous variables with high dispersion were normalized, if required. Variables with more than 20% missing data were excluded. Patients with missing data on urine output during the observation period were excluded from the analysis, as this variable was consistently identified as the most important predictor in the literature and by surveyed clinicians. For variables measured at multiple timepoints, missing values were imputed using the last observation carried forward (LOCF) method. For remaining variables, missing values were imputed using the median values. Non-binary categorical variables were one-hot encoded, and static variables were repeated across all timepoints.

The perioperative context of cardiac surgery was determined using billing data, with CCAM (Classification Commune des Actes Médicaux) codes for the French cohorts and DRG (Diagnosis Related Group) and ICD (International Classification of Diseases) procedure codes for the MIMIC cohort.

Feature selection

Predictors were selected based on a prior clinical survey21 and relevant literature10,44,45. Key variables included urine output, age, CRRT duration, invasive mechanical ventilation, fluid overload status and serum creatinine (SCr) levels. Patient comorbidities were extracted using the Charlson Comorbidity Index. Baseline SCr was defined as the last available measurement between the 365th and the 7th day before ICU admission. If unavailable, SCr was estimated assuming a renal clearance of 75 mL/min/1.73m2, using the CKD-EPI formula according to the KDIGO guidelines38,46, without race adjustment due to its absence in French datasets47.

Feature selection was performed using the Least Absolute Shrinkage and Selection Operator (LASSO) method on the training dataset, which allowed us to retain the most predictive variables while minimizing overfitting48.

Model development and evaluation

The models evaluated included: penalized logistic regression (LR), random forests (RF), extreme gradient boosting (XGB), k-nearest neighbors (KNN), and support vector machine (SVM). These models were selected for their suitability with relatively small datasets and prioritization of interpretability over complex “black box” models, following recommendations from the literature49.

Models were trained and fine-tuned on the Rouen cohort and validated externally on two independent datasets: one from another French center (Rennes) and the other from the MIMIC-IV database, representing distinct clinical and geographical contexts. This dual validation assessed the generalizability and robustness of the models.

Training and hyperparameter tuning employed tenfold cross-validation, where the training dataset was split into ten groups. For each iteration, one group was excluded from training and used as the validation set. This procedure was repeated ten times, with a different group serving as the validation set in each iteration. Aggregated results from all iterations optimized hyperparameters while minimizing overfitting (a common issue in machine learning where models perform well on training data but poorly on unseen data). Final models were retrained on the entire training dataset.

Predictive variables included immutable (e.g., age, comorbidities) and time-varying variables (e.g., urine output, laboratory results). Time-varying variables were considered over 24-h intervals, focusing on the two intervals immediately preceding the weaning attempt: t0 (discontinuation day) and t−1 (previous day) (Fig. 6). Classification threshold was optimized by maximizing Youden’s J index (J = Sensitivity + Specificity – 1)50.

Fig. 6.

Fig. 6

Study design. t0 represents the first CRRT weaning attempt, t-1 is the day before, and the observation period (t-1 to t0) is when the selected features are analyzed by the models for prediction. t7 marks the evaluation of the primary outcome, defined as CRRT-free survival for at least 7 days. Abbreviations: CRRT, continuous renal replacement therapy; ICU, intensive care unit.

Performance metrics included Area Under the Receiver Operating Characteristic curve (AUROC), accuracy, precision, recall, specificity, F1-score (combining precision and recall), and Matthews Correlation Coefficient (MCC). The 95% confidence intervals were calculated using tenfold cross-validation with the assumption of normality for the training cohort, and by bootstrapping with 1000 samples with replacement for the Rennes and MIMIC cohorts. The MCC is a robust quality measure for classification models, often described as a discretization of the Pearson correlation coefficient for binary variables51. It ranges from 1 (perfectly correct predictions for all examples) to − 1 (perfectly incorrect predictions for all examples)52.

Model interpretability was assessed using SHapley Additive exPlanations (SHAP), which quantified the contribution of each variable to predictions53. SHAP analysis identified key global predictors and provided insights into individual predictions, enhancing transparency and clinical relevance.

All analyses were performed using R version 3.6 (R Foundation for Statistical Computing, Vienna, Austria) and Python version 3.9.2 (Python Software Foundation).

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

The authors would like to thank the engineering teams from the clinical data warehouses of Rennes and Rouen for their invaluable assistance in the realization of this study. The authors also extend their gratitude to the team responsible for the MIMIC-IV database and the contributors of the MIMIC code repository for their essential support and resources.

Abbreviations

AI

Artificial intelligence

AKI

Acute kidney injury

AUROC

Area under the receiver operating characteristic curve

CDSS

Clinical decision support system

CDW

Clinical data warehouse

CRRT

Continuous renal replacement therapy

EDSaN

Entrepôt de Données de Santé Normand

FDA

Food and Drug Administration

ICCA©

IntelliSpace Critical Care and Anesthesia©

ICU

Intensive care unit

IRB

Institutional Review Board

KDIGO

Kidney disease improving global outcome

KNN

K-nearest neighbor

LASSO

Least absolute shrinkage and selection operator

LOCF

Last observation carried forward

LR

Logistic regression

MCC

Matthews correlation coefficient

ML

Machine learning

PT

Prothrombin time

RF

Random forest

SCr

Serum creatinine

SQL

Structured query language

SVM

Support vector machine

XGB

EXtreme gradient boosting

Author contributions

Benjamin Popoff: Conceptualization, Data curation, Formal Analysis, Methodology, Software, Visualization, Writing – original draft. Boris Delange: Writing – review & editing. Jonathan Nicolas: Conceptualization, Writing – review & editing. Arthur Le Gall: Writing – review & editing. Badisse Dahamna: Software, Resources, Writing – review & editing. Marc Cuggia: Writing – review & editing. Thomas Clavier: Conceptualization, Supervision, Writing – review & editing. Guillaume Bouzillé: Conceptualization, Methodology, Resources, Supervision, Writing – review & editing.

Funding

Support was provided solely from departmental sources.

Data availability

Data from the MIMIC-IV database are publicly available upon request at the PhysioNet website (https://physionet.org). The code used in this study, as well as the data from the French centers, are available from the corresponding author (B.P.) upon reasonable request.

Competing interests

The authors declare that they have no competing interest.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Hoste, E. A. J. et al. Global epidemiology and outcomes of acute kidney injury. Nat. Rev. Nephrol.14, 607–625 (2018). [DOI] [PubMed] [Google Scholar]
  • 2.Tandukar, S. & Palevsky, P. M. Continuous renal replacement therapy: Who, when, why, and how. Chest155, 626–638 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Rewa, O. G. et al. Epidemiology and outcomes of AKI treated with Continuous Kidney Replacement therapy: The Multicenter CRRTnet Study. Kidney Med.5, 100641 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Uchino, S. et al. Acute renal failure in critically ill patients: A multinational, multicenter study. JAMA294, 813–818 (2005). [DOI] [PubMed] [Google Scholar]
  • 5.Cerdá, J. et al. Promoting kidney function recovery in patients with AKI Requiring RRT. Clin J Am Soc Nephrol10, 1859–1867 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.STARRT-AKI Investigators et al. Timing of initiation of renal-replacement therapy in acute kidney injury. N. Engl. J Med.383, 240–251 (2020). [DOI] [PubMed]
  • 7.Gaudry, S. et al. Initiation strategies for renal-replacement therapy in the Intensive Care Unit. N Engl J Med375, 122–133 (2016). [DOI] [PubMed] [Google Scholar]
  • 8.Gaudry, S. et al. Comparison of two delayed strategies for renal replacement therapy initiation for severe acute kidney injury (AKIKI 2): A multicentre, open-label, randomised, controlled trial. Lancet397, 1293–1300 (2021). [DOI] [PubMed] [Google Scholar]
  • 9.Katulka, R. J. et al. Determining the optimal time for liberation from renal replacement therapy in critically ill patients: A systematic review and meta-analysis (DOnE RRT). Crit. Care24, 50 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Uchino, S. et al. Discontinuation of continuous renal replacement therapy: A post hoc analysis of a prospective multicenter observational study. Crit. Care Med.37, 2576–2582 (2009). [DOI] [PubMed] [Google Scholar]
  • 11.Schiffl, H. & Lang, S. M. Current approach to successful liberation from renal replacement therapy in critically ill patients with severe acute kidney injury: The Quest for Biomarkers Continues. Mol. Diagn. Ther.25, 1–8 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Clark, E. G. & Bagshaw, S. M. Unnecessary renal replacement therapy for acute kidney injury is harmful for renal recovery. Semin. Dial.28, 6–11 (2015). [DOI] [PubMed] [Google Scholar]
  • 13.Hong, N. et al. State of the art of machine learning-enabled clinical decision support in intensive care units: Literature review. JMIR Med. Inform.10, e28781 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.van de Sande, D., van Genderen, M. E., Huiskens, J., Gommers, D. & van Bommel, J. Moving from bytes to bedside: A systematic review on the use of artificial intelligence in the intensive care unit. Intensive Care Med.47, 750–760 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Xu, H. et al. Machine learning-based risk prediction model construction of difficult weaning in ICU patients with mechanical ventilation. Sci. Rep.14, 20875 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C. & Faisal, A. A. The Artificial intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nat. Med.10.1038/s41591-018-0213-5 (2018). [DOI] [PubMed] [Google Scholar]
  • 17.de Hond, A. A. H. et al. Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: A scoping review. npj Digit. Med.5, 1–13 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.van de Sande, D. et al. Developing, implementing and governing artificial intelligence in medicine: A step-by-step approach to prevent an artificial intelligence winter. BMJ Health Care Inform.29, e100495 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Seneviratne, M. G. et al. User-centred design for machine learning in health care: A case study from care management. BMJ Health Care Inform.29, e100656 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Borges do Nascimento, I. J. et al. Barriers and facilitators to utilizing digital health technologies by healthcare professionals. npj Digit. Med.6, 1–28 (2023). [DOI] [PMC free article] [PubMed]
  • 21.Popoff, B., Cabon, S., Cuggia, M., Bouzillé, G. & Clavier, T. Expectations of intensive care physicians regarding an artificial intelligence-based decision support system for weaning from continuous renal replacement therapy: A pre-development survey study. JMIR Med .Inform. 13, e63709. 10.2196/63709 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. AI in health and medicine. Nat Med28, 31–38 (2022). [DOI] [PubMed] [Google Scholar]
  • 23.Friedman, C. P. et al. The science of Learning Health Systems: Foundations for a new journal. Learning Health Syst.1, e10020 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Popoff, B. et al. How to accurately detect renal replacement therapy weaning in intensive care: Data quality and standardization considerations for the OMOP common data model. Stud. Health Technol. Inform.316, 1584–1588 (2024). [DOI] [PubMed] [Google Scholar]
  • 25.Index - FHIR v5.0.0. https://www.hl7.org/fhir/.
  • 26.Liu, C. et al. Predicting successful continuous renal replacement therapy liberation in critically ill patients with acute kidney injury. J. Crit. Care66, 6–13 (2021). [DOI] [PubMed] [Google Scholar]
  • 27.Wang, T.-J. et al. Predictive approach for liberation from acute dialysis in ICU patients using interpretable machine learning. Sci. Rep.14, 13142 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Jeon, J. et al. Validation of prediction model for successful discontinuation of continuous renal replacement therapy: A multicenter cohort study. Kidney Res. Clin. Pract.10.23876/j.krcp.23.308 (2024) [DOI] [PMC free article] [PubMed]
  • 29.Zhong, L., Min, J., Zhang, J., Hu, B. & Qian, C. Risk prediction models for successful discontinuation in acute kidney injury undergoing continuous renal replacement therapy. iScience27, (2024). [DOI] [PMC free article] [PubMed]
  • 30.Sheng, S. et al. Factors and machine learning models for predicting successful discontinuation of continuous renal replacement therapy in critically ill patients with acute kidney injury: A retrospective cohort study based on MIMIC-IV database. BMC Nephrol.25, 407 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Liang, Q., Xu, X., Ding, S., Wu, J. & Huang, M. Prediction of successful weaning from renal replacement therapy in critically ill patients based on machine learning. Ren. Fail.46, 2319329 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Yang, J. et al. A comprehensive step-by-step approach for the implementation of target trial emulation: Evaluating fluid resuscitation strategies in post-laparoscopic septic shock as an example. Laparosc. Endosc. Robot. Surg.8, 28–44 (2025). [Google Scholar]
  • 33.Yang, J. et al. Identification of clinical subphenotypes of sepsis after laparoscopic surgery. Laparosc. Endosc. Robot. Surg.7, 16–26 (2024). [Google Scholar]
  • 34.FDA. Clinical Decision Support Software. Guidance for Industry and Food and Drug Administration Staff (2022).
  • 35.Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data10, 1 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ385, e078378 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Toulouse, E., Lafont, B., Granier, S., Mcgurk, G. & Bazin, J.-E. French legal approach to patient consent in clinical research. Anaesth. Crit. Care Pain Med.39, 883–885 (2020). [DOI] [PubMed] [Google Scholar]
  • 38.Khwaja, A. KDIGO clinical practice guidelines for acute kidney injury. Nephron Clin. Pract.120, c179-184 (2012). [DOI] [PubMed] [Google Scholar]
  • 39.Katayama, S. et al. Factors predicting successful discontinuation of continuous renal replacement therapy. Anaesth. Intensive Care44, 453–457 (2016). [DOI] [PubMed] [Google Scholar]
  • 40.Fröhlich, S., Donnelly, A., Solymos, O. & Conlon, N. Use of 2-hour creatinine clearance to guide cessation of continuous renal replacement therapy. J. Crit. Care27(744), e1-5 (2012). [DOI] [PubMed] [Google Scholar]
  • 41.Pressat-Laffouilhère, T. et al. Evaluation of Doc’EDS: A French semantic search tool to query health documents from a clinical data warehouse. BMC Med. Inform. Decis. Mak.22, 34 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Madec, J. et al. eHOP Clinical Data Warehouse: From a Prototype to the Creation of an Inter-Regional Clinical Data Centers Network. Stud. Health Technol. Inform.264, 1536–1537 (2019). [DOI] [PubMed] [Google Scholar]
  • 43.Johnson, A. E., Stone, D. J., Celi, L. A. & Pollard, T. J. The MIMIC Code Repository: enabling reproducibility in critical care research. J. Am. Med. Inform. Assoc.25, 32–39 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Raurich, J. M. et al. Successful weaning from continuous renal replacement therapy. Associated risk factors. J. Crit.l Care45, 144–148 (2018). [DOI] [PubMed]
  • 45.Mendu, M. L. et al. A decision-making algorithm for initiation and discontinuation of RRT in severe AKI. Clin. J. Am. Soc. Nephrol.12, 228–236 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Levey, A. S. et al. A new equation to estimate glomerular filtration rate. Ann. Intern. Med.150, 604–612 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Yap, E. et al. The implication of dropping race from the MDRD equation to estimate GFR in an African American-Only Cohort. Int J Nephrol2021, 1880499 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Tibshirani, R. Regression shrinkage and selection via the Lasso. J. Roy. Stat. Soc.: Ser. B (Methodol.)58, 267–288 (1996). [Google Scholar]
  • 49.Herm, L.-V., Heinrich, K., Wanner, J. & Janiesch, C. Stop ordering machine learning algorithms by their explainability! A user-centered investigation of performance and explainability. Int. J. Inf. Manage.69, 102538 (2023). [Google Scholar]
  • 50.Youden, W. J. Index for rating diagnostic tests. Cancer3, 32–35 (1950). [DOI] [PubMed] [Google Scholar]
  • 51.Matthews, B. W. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochimica et Biophysica Acta (BBA) - Protein Structure405, 442–451 (1975). [DOI] [PubMed]
  • 52.Boughorbel, S., Jarray, F. & El-Anbari, M. Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric. PLoS ONE12, e0177678 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Molnar, C. 9.6 SHAP (SHapley Additive exPlanations) | Interpretable Machine Learning.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

Data from the MIMIC-IV database are publicly available upon request at the PhysioNet website (https://physionet.org). The code used in this study, as well as the data from the French centers, are available from the corresponding author (B.P.) upon reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES