Abstract
Long-term outcomes of kidney allografts vary significantly among deceased donor kidney transplant recipients, and current prediction tools struggle to integrate comprehensive pre- and post-transplant factors. Extended longitudinal follow-up data beyond five years remains particularly scarce in kidney transplantation research despite being crucial for understanding true long-term outcomes. To address this, we developed and validated machine learning models to predict 5-year allograft survival using a distinctive cohort of 940 adult deceased donor kidney transplantation recipients with extended follow-up exceeding 5 years. Two predictive models were developed: a pre-transplant model (Kidney Allograft Prediction of Transplant Outcome Risk, KAPTOR-pre) using pre-transplant donor-recipient matching data, and a 1-year landmark conditional prediction model (KAPTOR-full) incorporating both pre- and post-transplant parameters, pathological data, and laboratory markers from the first year. KAPTOR-full achieved excellent discrimination with area under the receiver operating characteristic of 0.904, while KAPTOR-pre performed well at 0.813. In internal validation, both models showed higher C-index and improved risk stratification compared with established prognostic tools including KDPI. The extended follow-up period allowed internal assessment of model performance for 5-year outcomes. Ultimately, our models integrating routine clinical variables demonstrated excellent predictive performance for long-term graft survival. While the pre-transplant model achieved good discrimination, the addition of first-year post-transplant data significantly enhanced predictive accuracy. Both models outperformed existing tools in internal validation and may support personalized risk assessment, pending independent multicenter validation.
Keywords: Kidney transplantation, deceased donor, machine learning, prognosis prediction, clinical decision support
Introduction
Kidney transplantation (KT) represents the most effective treatment for end-stage renal disease, not only significantly improving patients’ quality of life but also proving more cost-effective than dialysis in the long term [1]. However, more than 20% of patients experience graft failure within five years [2], with failure rates reaching 20–30% at ten years [3]. This issue presents unique clinical challenges in China: since the implementation of the citizen organ donation system in 2015, deceased donor (DD) transplants have accounted for over 80% [4] of all kidney transplants. However, the long-term graft outcomes of DD kidneys are influenced by multiple factors [5]. The various forms of injury include damage from underlying diseases such as hypertension [6,7] and diabetes [8], warm ischemic injury during cardiopulmonary resuscitation [9], nosocomial infections from invasive treatments during ICU stays [10], systemic inflammatory responses induced by brain death [11], extended cold ischemia time during organ allocation [12–14], and reperfusion injury.
Existing predictive tools such as Kidney Donor Profile Index (KDPI) [15] were primarily developed based on European and American population data and lack proficiency in fully reflecting the characteristics of Chinese donors and recipients, while also showing limitations in predictive accuracy [16,17]. Notably, these tools often overlook the dynamic changes in early post-transplant clinical indicators (within the first year), which could provide crucial early warning information for timely intervention. Identifying high-risk patients promptly and implementing targeted measures is crucial for improving outcomes.
With the advancing application of machine learning techniques in KT [18], integrating multidimensional clinical data for precise prediction has become feasible. This study aims to develop a prediction model system for DDKT recipients, providing quantitative evidence for clinical decision-making. The establishment of this early warning system will provide a scientific basis for individualized treatment planning, contributing to improved long-term survival rates among Chinese kidney transplant recipients. Additionally, this research offers valuable methodological insights for other developing countries facing similar challenges.
To our knowledge, this is the first study to develop a machine learning-based prediction system for long-term graft survival specifically in Chinese deceased donor kidney transplant recipients. Furthermore, unlike most existing models that provide only a single time-point prediction, our two-stage framework (pre-transplant and conditional 1-year post-transplant) enables dynamic risk assessment aligned with clinical decision nodes.
Materials and methods
Data sources
This retrospective study utilized data collected from the Kidney Disease Center of the First Affiliated Hospital of Zhejiang University. The study cohort comprised adult patients (≥18 years) who underwent KT between 2010 and 2020. The inclusion criteria were as follows: (1) deceased donor kidney transplantation; (2) recipients of a primary kidney transplant; (3) single kidney transplantation; (4) a minimum follow-up period of 5 years. Based on these selection criteria, a total of 940 patients were included in the study.
Outcomes, predictors, and model definitions
The primary endpoint was 5-year graft survival. Graft failure was defined as a composite endpoint of return to dialysis, retransplantation, or all-cause mortality, whichever occurred first. Two models were developed according to predictor availability. KAPTOR-pre used only variables available before or at the time of transplantation and was intended for baseline prediction. KAPTOR-full additionally included first-year post-transplant variables and was designed as a 1-year landmark conditional prediction model among recipients with a functioning graft at 1 year.
Baseline demographic data and laboratory results were extracted from the hospital information system.
Donor-related variables included demographic characteristics (age, sex, and body mass index (BMI)), Donation after circulatory death (DCD) status, history of hypertension and diabetes, length of ICU stay, final serum creatinine (SCr) level before organ procurement, cold ischemia time (CIT), and warm ischemia time. The cause of death was classified into four categories: traumatic brain injury, intracerebral hemorrhage, hypoxic-ischemic encephalopathy, and other causes (including drug intoxication, hepatorenal syndrome, aortic dissection involving renal artery ostium, etc.).
Recipient-related variables comprised demographic information (age, sex, and BMI), pre-transplant dialysis duration, panel reactive antibody (PRA), and Human leukocyte antigen (HLA) mismatches (A, B, and DR).
Postreperfusion graft biopsy Remuzzi score: postreperfusion biopsies were obtained using a 16-gauge Tru-Cut needle immediately after clamp release following anastomosis completion. All tissue samples were processed following standardized protocols. The pathological assessment followed the established Remuzzi scoring system [19], which evaluates four key parameters: glomerulosclerosis, interstitial fibrosis, tubular atrophy, and vascular disease. Each parameter was scored on a scale of 0–3, yielding a total Remuzzi score ranging from 0 to 12. All biopsy specimens were evaluated by dedicated renal pathologists from our center.
Ultrasonographic graft arterial resistance index (RI): RI was evaluated using Doppler ultrasonography by an experienced dedicated ultrasonographer within the first week after transplantation. Measurements were obtained from the renal, segmental, interlobar, and arcuate arteries. The peak systolic velocity (PSV) and the end diastolic velocity (EDV) were measured to calculate RI, using the following formula: RI = (PSV - EDV)/PSV.
Post-transplantation laboratory parameters: SCr levels measured at 1 week, 1 month, 3 months, 6 months, and 12 months after transplantation; Urinary albumin-to-creatinine ratio (UACR), blood urea nitrogen (BUN), and serum albumin (ALB) levels at 12 months post-transplantation. Estimated glomerular filtration rate (eGFR) was calculated using the EPI 2021 equation [20].
Biopsy indications and AR diagnostic criteria: Renal allograft biopsy was performed when any of the following occurred: (1) unexplained persistent elevation of serum creatinine (≥20% increase from baseline); (2) new-onset proteinuria (urinary albumin-to-creatinine ratio > 300 mg/g) or a significant increase from baseline; (3) delayed graft function persisting beyond 2 weeks without an identifiable cause; or (4) protocol biopsy for specific clinical research patients. Acute rejection was diagnosed by biopsy in all cases, using the Banff 2017 classification system for pathological diagnosis and subtyping. AR was classified as T cell-mediated rejection (TCMR) or antibody-mediated rejection (ABMR), with respective Banff grading. Subclinical rejection (biopsy findings without concomitant graft dysfunction) was excluded from the AR definition.
Data processing
The dataset was randomly divided into training and testing sets at an 80%–20% ratio, with the split performed at the patient level. The models were trained on the training set using a range of hyperparameters and subsequently validated on the testing set to avoid overfitting and ensure robust performance.
To ensure data quality, variables with more than 25% missing data were excluded. Additionally, three new variables were generated: Donor-to-recipient weight ratio was created by dividing the donor’s weight by the recipient’s weight. The product of CIT and donor age was used to represent the interaction between these two factors. %Change in eGFR was defined as the relative change in eGFR over time, calculated as (eGFR at 6 months – eGFR at 12 months)/eGFR at 6 months post-transplantation. Missing data were addressed using simple imputation, which involved mean imputation for continuous or quantitative features and mode imputation for categorical or qualitative features. Importantly, the imputation process was performed exclusively on the training set, and the calculated mean and mode values were subsequently applied to impute missing values in the testing set.
The initial cohort of 940 patients was initially used as the pre-KT cohort, comprising 18 variables that were available prior to transplantation. For the KAPTOR-full model, the post-KT cohort was restricted to 879 patients who had a functioning graft and remained under follow-up at 12 months after transplantation. This design ensured that predictors were measured before prediction and that KAPTOR-full was interpreted as a conditional model among 1-year graft survivors. After generating dummy variables for categorical data, the postoperative cohort expanded to include 34 variables.
Prior to feature selection, pairwise correlations among candidate predictors were examined using a correlation matrix and substantial correlations were observed among several clinically related variables (Supplementary Figure S1). Given the limited number of graft failure events, feature prescreening was performed to reduce redundancy, improve model stability, and mitigate the risk of overfitting.
After dummy-variable encoding where appropriate, candidate predictors were ranked separately according to the absolute values of their Least Absolute Shrinkage and Selection Operator (LASSO) coefficients and according to Gradient Boosting Machine (GBM) variable importance. For each method, predictors ranked within the top two-thirds were retained and then combined to form an initial candidate feature pool. This initial pool was further selected according to LASSO coefficient magnitude, GBM importance ranking, clinical interpretability, missingness, routine clinical availability, and redundancy with other predictors. After preprocessing, the preoperative cohort was reduced to 6 key features, and the postoperative cohort was narrowed down to 12 key features. These selected features were subsequently utilized for model development and analysis.
Model training
In this study, we used 5 linear and non-linear machine learning algorithms to develop prediction models including Logistic Regression (LR), random forest (RF), Support Vector Machine (SVM), extreme gradient boosting decision tree (XGB), and Light Gradient Boosting Machine (LightGBM). For each algorithm, a grid search with cross-validation was conducted on the training set to identify the optimal hyperparameter combinations. The final models, trained with the optimal hyperparameters, were then evaluated on the held-out test dataset to assess generalization performance. Model performance was assessed using three key metrics: (1) Area under the Receiver Operating Characteristic (AUROC), which evaluates overall risk stratification capability by considering the tradeoff between sensitivity and specificity.; (2) Area under the Precision-Recall curve (AUPRC), which assesses clinical utility by balancing sensitivity and positive predictive value; (3) and Brier Score, which measures the overall accuracy of predicted probabilities. The model achieving the highest performance across the metrics was selected for further evaluation on the test dataset. To quantify the robustness of the results, 95% confidence intervals (CI) were calculated through 1,000 bootstrap iterations of the test split. The final selected models, referred to as KAPTOR-pre for the pre-KT cohort and KAPTOR-full for the post-KT cohort, underwent additional evaluation using the C-index metric [21] and Kaplan-Meier survival analysis to further validate their performance.
Performance evaluation
We compared the predictive performance of the KAPTOR models with eight established prognostic tools identified through a comprehensive literature review, including four pre-transplant scoring systems (Pessione 2003 [22], DDS 2005 [23], DRS 2005 [24], and KDPI), three post-transplant models (Foucher 2010 [25], Moore 2011 [26], Shabir 2014 [27]), and the widely used clinical marker, 12-month post-transplant eGFR [28,29]. All risk scores were calculated using formulas from their original publications. Comparative analyses were conducted using C-index and survival analyses to assess discriminative ability across all models. KDPI is derived from the Kidney Donor Risk Index (KDRI), calculated for each donor using a 2020 scaling factor [30], then mapped to a cumulative percentage scale that ranges from 0% to 100% [31]. The KDRI is calculated using the following 10 donor-specific clinical characteristics: age, height, weight, ethnicity, history of hypertension, history of diabetes, cause of death, serum creatinine, hepatitis C virus status, and donation after cardiac death status.
To assess the independent prognostic value of external scoring systems in the presence of KAPTOR risk scores, we conducted multivariable Cox proportional hazards regression analyses. Each external score was first evaluated individually, then jointly analyzed with KAPTOR-pre or KAPTOR-full scores to observe any remaining predictive value.
Explainability
Interpretable machine learning models are essential for clinical tasks to ensure transparency, trust, and correct decision-making. Beyond evaluating technical performance, we utilized Shapley additive explanations (SHAP) [32] to visualize feature importance and assess how the KAPTOR-full model derives its predictions. SHAP values were analyzed for four patient subsets – true positives (TPs), true negatives (TNs), false positives (FPs), and false negatives (FNs) – to visualize the contribution of each feature for correctly and incorrectly classified cases, providing insight into potential model biases or limitations. Next, SHAP summary plots, with LOESS smoothing applied, were generated to examine individual feature contributions to the predicted probability of long-term graft survival. This analysis was conducted for all 12 features of the KAPTOR-full model.
To further assess the clinical utility of KAPTOR-pre and KAPTOR-full, we performed decision curve analysis (DCA) to quantify the net benefit of each model in predicting 5-year graft survival.
Statistical analysis
All statistical analyses were performed using R 4.3.2. Baseline characteristics of recipients and donors, along with pre-operative details, were presented as means and standard deviations for continuous variables, and as frequencies and percentages for categorical variables. Group differences were evaluated using independent t tests (for continuous variables) and chi-square (χ2) tests (for categorical variables). Multiple groups were compared using Kruskal-Wallis test. All statistical tests were two-sided and evaluated at a significance level of p < 0.05, with the Bonferroni correction for multiple testing.
Results
Study population
A total of 940 patients with a median follow-up of 6.9 years (95%CI 6.66–7.18) who underwent DDKT at the Kidney Disease Center of the First Affiliated Hospital of Zhejiang University between 2010 and 2020 were included in this study (Figure 1). For outcome assessment, a binary outcome variable was constructed to indicate whether a graft survived for at least 5 years. The pre-KT cohort consisted of all 940 patients, utilizing only preoperative parameters to predict graft outcomes. Among them, 879 patients whose grafts survived for more than one year formed the post-KT cohort, which incorporated both preoperative parameters and first-year post-transplant variables.
Figure 1.
Flow chart of patient enrollment.
Baseline demographic and clinical characteristics of the study population are summarized in Table 1. In the pre-KT cohort, the mean age of recipients at the time of KT was 44.37 years (SD: 10.47), with 59.6% of the patients being male. Key preoperative variables included recipient and donor demographics, HLA mismatches, dialysis duration prior to KT, and other relevant clinical parameters. For the post-KT cohort, additional first-year post-transplant variables were incorporated (Tab S1).
Table 1.
Recipient and donor characteristics stratified by graft survival status.
| pre-KT cohort |
post-KT cohort |
|||||||
|---|---|---|---|---|---|---|---|---|
| Overall | No graft failure | Graft failure | P value | Overall | No graft failure | Graft failure | P value | |
| n (%) | 940 (100) | 769 (81.8) | 171 (18.2) | 879 (100) | 769 (87.5) | 110 (12.5) | ||
| Recipient demographics | ||||||||
| Age (years), mean (SD) | 44.37 ± 10.47 | 44.10 ± 10.12 | 45.56 ± 11.88 | 0.100 | 44.44 ± 10.22 | 44.10 ± 10.12 | 46.79 ± 10.64 | 0.010 |
| Sex, male, n (%) | 560 (59.6) | 446 (58.0) | 114 (66.7) | 0.041 | 517 (58.8) | 446 (58.0) | 71 (64.5) | 0.010 |
| Body mass index (kg/m²), mean (SD) | 21.38 ± 3.11 | 21.31 ± 2.96 | 21.72 ± 3.72 | 0.114 | 21.37 ± 3.05 | 21.31 ± 2.96 | 21.81 ± 3.65 | 0.109 |
| Dialysis duration before transplant (month), mean (SD) | 49.82 ± 30.27 | 48.28 ± 29.01 | 56.73 ± 34.67 | 0.001 | 49.77 ± 30.23 | 48.28 ± 29.01 | 60.20 ± 36.16 | <0.001 |
| PRA positive, n (%) | 28 (3.0) | 20 (2.6) | 8 (4.6) | 0.068 | 24 (2.7) | 20 (2.6) | 4 (3.6) | 0.091 |
| HLA-A/B/DR mismatch, mean (SD) | 2.93 ± 1.28 | 2.97 ± 1.23 | 2.78 ± 1.45 | 0.079 | 2.93 ± 1.25 | 2.97 ± 1.23 | 2.67 ± 1.35 | 0.021 |
| D/R weight ratio, mean (SD) | 1.10 ± 0.28 | 1.11 ± 0.27 | 1.03 ± 0.33 | 0.001 | 1.10 ± 0.28 | 1.11 ± 0.27 | 1.05 ± 0.32 | 0.037 |
| Donor demographics | ||||||||
| Age (years), mean (SD) | 40.19 ± 15.19 | 39.68 ± 14.60 | 42.50 ± 17.48 | 0.028 | 40.19 ± 14.88 | 39.68 ± 14.60 | 43.74 ± 16.34 | 0.007 |
| Sex, male, n (%) | 758 (80.6) | 639 (83.1) | 119 (69.6) | 0.004 | 718 (81.7) | 639 (83.1) | 79 (71.8) | 0.009 |
| Donor final Scr (μmol/L), mean (SD) | 106.34 ± 81.22 | 102.35 ± 76.41 | 124.30 ± 98.31 | 0.001 | 104.67 ± 79.77 | 102.35 ± 76.41 | 120.91 ± 99.05 | 0.022 |
| Body mass index (kg/m²), mean (SD) | 22.67 ± 3.14 | 22.68 ± 3.09 | 22.60 ± 3.35 | 0.104 | 22.67 ± 3.14 | 22.68 ± 3.09 | 22.61 ±3.26 | 0.824 |
| Donor ICU stay (days), mean (SD) | 6.85 ± 6.49 | 6.66 ± 5.64 | 7.70 ± 9.38 | 0.057 | 6.81 ± 6.35 | 6.66 ± 5.64 | 7.87 ± 9.98 | 0.061 |
| Donor cause of death, n (%) | <0.001 | <0.001 | ||||||
| Traumatic Brain Injury | 540 (57.4) | 474 (61.6) | 66 (38.6) | 520 (59.2) | 474 (61.6) | 46 (41.8) | ||
| Intracerebral Hemorrhage | 288 (30.6) | 211 (27.4) | 77 (45.0) | 258 (29.4) | 211 (27.4) | 47 (42.7) | ||
| Hypoxic-Ischemic Encephalopathy | 82 (8.7) | 60 (7.8) | 22 (12.9) | 76 (8.6) | 60 (7.8) | 16 (14.5) | ||
| Others | 30 (3.2) | 24 (3.1) | 6 (3.5) | 25 (2.8) | 24 (3.1) | 1 (0.9) | ||
| Donor DCD status, n (%) | 851 (90.5) | 702 (91.3) | 149 (87.1) | 0.099 | 796 (90.6) | 702 (91.3) | 94 (85.5) | 0.075 |
| Cold ischemic time (hours), mean (SD) | 7.05 ± 3.50 | 7.00 ± 3.41 | 7.28 ± 3.84 | 0.348 | 7.00 ± 3.39 | 7.00 ± 3.41 | 7.01 ± 3.19 | 0.970 |
| Warm ischemic time (mins), mean (SD) | 10.20 ± 9.56 | 9.96 ± 6.65 | 11.25 ± 17.45 | 0.113 | 9.89 ± 6.77 | 9.96 ± 6.65 | 9.40 ± 7.61 | 0.416 |
Scr, serum creatinine; HLA, Human leukocyte antigen; D/R, Donor-to-Recipient; DCD, Donation after circulatory death; ICU, Intensive care unit.
Evaluation of model performance
The correlation matrix heatmap (Figure S1) illustrates the Pearson correlation coefficients between model variables. Substantial collinearity existed among several candidate predictors, particularly among serial SCr measurements at different post-transplant time points, between eGFR and SCr, and among arterial resistance indices from different renal artery branches. This multicollinearity justified the need for feature pre-selection to ensure model stability and interpretability.
The features identified by LASSO regression and GBM as key predictors of long-term graft survival in the pre-KT and post-KT cohorts are summarized in Tab S2-S5. These selected features formed the basis for model training across both cohorts and algorithms.
The predictive performance of the KAPTOR-pre and KAPTOR-full models was evaluated across five machine learning algorithms – LR, RF, SVM, XGB, and LightGBM. These algorithms were trained and tested on the pre-KT and post-KT cohorts, respectively (Tab S6). Among these, XGB consistently outperformed the other models in both cohorts, achieving the highest AUROC and AUPRC values. In the post-KT cohort, KAPTOR-full achieved an AUROC of 0.904 (95% CI: 0.820–0.973), and an AUPRC of 0.771 (95% CI: 0.601–0.903), showcasing excellent discrimination, strong calibration, and practical clinical utility. By comparison, KAPTOR-pre, trained only on the preoperative dataset, achieved an AUROC of 0.813 (95% CI: 0.698–0.912), and an AUPRC of 0.710 (95% CI: 0.558–0.843). Calibration curves for both models are presented in Figure 2. KAPTOR-full demonstrated good calibration, with predicted probabilities closely aligning with observed graft failure rates across the entire risk spectrum. KAPTOR-pre also showed acceptable calibration. While both models demonstrated reliable predictive capabilities, the post-KT data enhanced the performance of KAPTOR- full, emphasizing the significant value of incorporating first-year post-transplant data into the predictive model (Figure 2).
Figure 2.
Model performance of KAPTOR-pre and KAPTOR-full models in predicting 5-year allograft outcomes.
Orange lines represent KAPTOR-pre model (pre-KT Cohort, n = 940) trained on preoperative data only, while purple lines represent KAPTOR-full model (post-KT cohort, n = 879) incorporating both preoperative and first-year post-transplant data.
(a) Receiver operating characteristic curves for the prediction of allograft outcome. Summarizing metric is the receiver operating characteristic area under the curve (AUROC).
(b) Precision curves for the prediction of allograft outcome. Summarizing metric is the precision recall area under the curve (AUPRC).
(c) Calibration curves for the prediction of allograft outcome. The x-axis represents the mean predicted probability of graft failure, and the y-axis represents the observed fraction of positives. The diagonal dashed line indicates perfect calibration. Brier scores for KAPTOR-pre and KAPTOR-full were 0.124 and 0.059, respectively, indicating good overall calibration.
Comparison with current prognostic scores
To evaluate prognostic discrimination, Kaplan-Meier survival analyses were conducted for all scores. In the post-KT cohort, based on the pooled out-of-fold predictions from the five-fold cross-validation, KAPTOR-full demonstrated superior stratification of risk groups, with 5-year graft survival probabilities of 95.0% in the low-risk group (n = 794) and 17.6% in the high-risk group (n = 85), surpassing all existing scores through five-fold cross-validation (Tab S7, S8>). Kaplan-Meier analysis demonstrated that KAPTOR-pre and KAPTOR-full effectively stratified patients into distinct risk groups (log-rank p < 0.001 for both models; Figure 3a and b, Figure S2), illustrated the risk stratification ability of the KAPTOR models.
Figure 3.
Performance of KAPTOR.
(a, b) The graft surviving probability analysis using the Kaplan–Meier method by KAPTOR risk groups in the pre-operative cohort and log rank test P value.
(c) Comparison of concordance index between KAPTOR models and traditional scoring systems. Box plots demonstrate the C-index distributions across five-fold cross-validation. The horizontal dashed line indicates the C-index of KAPTOR models. Box plots show median, interquartile range, and whiskers extend to 1.5 times the interquartile range.
(d) Residual prognostic value of all established clinical risk scores when using KAPTOR, predicted risk scores in a multivariable analysis on pre-operative cohort (n = 940) and post-operative cohort (n = 879). Data are presented as the HRs and 95% CIs.
The discriminative performance of KAPTOR-pre and KAPTOR-full was evaluated using C-index and compared with existing scoring systems. Specifically, five-fold cross-validation revealed that KAPTOR models outperformed traditional scores that used similar input parameters (Figure 3c, Table S9). The Kruskal-Wallis test showed significant differences between KAPTOR-pre and traditional pre-KT scores (p < 0.001) as well as between KAPTOR-full and post-KT scores (p < 0.001). Subsequent Dunn’s test with Bonferroni correction confirmed that KAPTOR-pre achieved significantly higher C-index compared to pre-KT scores (all p < 0.001), while KAPTOR-full also outperformed post-KT scores (all p < 0.001) (Tab S10, S11).
Multivariable analysis further validated the independent prognostic value of the KAPTOR models. In the pre-KT cohort, KAPTOR-pre maintained its prognostic significance, whereas traditional pre-KT scores were not independently prognostic (red dots). In the post-KT cohort, KAPTOR-full demonstrated strong independent prognostic value (HR = 1.038; 95% CI: 1.031–1.044; p = 8.60 × 10–30), with only Foucher2010 retaining statistical significance (HR = 1.035; 95% CI: 1.005–1.066; p = 0.02) (Figure 3d, Tables S12–S14).
Model and feature interpretation
The final model configurations used for interpretation are summarized in Tables S15 and S16. We analyzed the SHAP values on the post-KT cohort to understand feature contributions to model predictions (Figure 4a). The SHAP analysis showed the relative importance and direction of the top clinical and demographic variables, with positive SHAP values indicating increased predicted risk and negative values indicating decreased predicted risk. Recipient SCr at 12 mo emerged as the strongest predictor, followed by recipient serum Alb at 12 mo, donor/recipient weight ratio, recipient UACR at 12 mo, and recipient age. We further performed a descriptive subgroup SHAP analysis according to prediction outcomes (Figure 4b). This analysis showed that several dominant predictors, particularly recipient SCr and serum Alb at 12 months, remained highly ranked across multiple prediction-outcome groups. TP and FN cases showed similar rankings for the leading post-transplant markers, whereas FP predictions were more influenced by pre-transplant factors. TN cases showed a slightly different pattern, with postperfusion biopsy Remuzzi score gaining greater relative prominence.
Figure 4.
Model explanations for the model trained and evaluated on post-operative cohort.
Orange indicates that a higher feature value has the corresponding impact, as indicated by the x-axis, on model output. Purple indicates the impact of lower feature values on model output.
(a) Ordered ranking of the ten most important features by average magnitude of SHAP values and direction of influence on output predictions.
(b) Ordered ranking of the ten most important features by average magnitude of SHAP values when isolating subpopulations of patients in the test set that were classified as true positives, true negatives, false positives, and false negatives. We reported the corresponding rank of each feature in the ordered ranking of feature importance using the entire dataset for each feature (for the top ten features) (n = 879).
The overlap in SHAP importance patterns indicates that similar predictors contributed to risk estimation across several groups; however, it does not fully explain the mechanisms underlying misclassification. Misclassified cases may reflect unmeasured clinical factors, residual confounding, threshold-dependent classification effects, post-transplant management changes, or complex interactions not completely captured by the model or by SHAP visualization. These findings suggest post-transplant laboratory markers, donor-related characteristics, and pre-transplant parameters provide clinically interpretable information together for post-transplant risk stratification.
To further understand the relationship between individual features and model predictions, we generated SHAP interaction plots for all 12 variables in the KAPTOR-full model (Figure 5). These plots visualized how changes in each feature value affect the predicted risk while accounting for feature interactions. The plots reveal several non-linear relationships, particularly for laboratory markers at 12 months (SCr, %Change in eGFR, UACR, and Alb). The donor-related features (D/R weight ratio, donor final SCr) and histological parameters (postperfusion biopsy Remuzzi score) also showed distinct patterns of influence on risk prediction.
Figure 5.
SHAP value plots showing the impact of key features on model predictions.
The plots illustrate the relationship between feature values (x-axis) and their SHAP values (y-axis) for twelve important predictors. Each point represents a single patient, with the red line showing the general trend. Higher SHAP values indicate stronger positive impact on model predictions. Blue dots represent individual observations; darker areas indicate higher density of points. The dashed horizontal line at y = 0 represents neutral impact on graft failure risk.
%Change in eGFR shows escalating risk above 50% decline, SCr at 12 mo demonstrates positive association with graft failure risk, UACR at 12 mo shows initial steep rise in risk before stabilizing; Alb at 12 months reveals protective effect above 40 g/L; Recipient age shows U-shaped relationship with graft failure risk; D/R weight ratio indicates optimal range around 1.0 for minimal risk; Dialysis duration before KT demonstrates increasing graft failure risk with longer exposure; Graft renal artery RI shows exponential risk increase; CIT × Donor age interaction shows positive impact on risk; Postreperfusion Remuzzi score indicates a 3-mark threshold effect for graft failure; Donor final SCr also shows U-shaped risk pattern, where both extremes – high (kidney dysfunction) and low (ICU-related muscle wasting) – correlate with increased graft failure risk. Donor cause of death categories demonstrate varying risk impacts. Traumatic brain injury was associated with the lowest risk, while other causes (including drug intoxication, hepatorenal syndrome, and aortic dissection) showed the highest risk.
Clinical considerations of an applied model
Decision curve analysis demonstrated the clinical utility of KAPTOR models (Figure S3). The KAPTOR-full model exhibited consistently higher net benefit across most threshold probabilities, indicating its superior clinical value. As threshold probabilities increased, the net benefit of both models gradually decreased, reflecting the tradeoff FP and FN.
For clinical implementation, threshold selection requires careful consideration of potential consequences. In the pre-KT cohort, a moderate threshold is recommended for the KAPTOR-pre model since FPs may lead to inappropriate candidate exclusion while FNs could deprive suitable candidates of transplantation opportunities. Similarly, for the KAPTOR-full model, a balanced threshold helps avoid both unnecessary interventions (FPs) and delayed treatment (FNs) that could compromise graft outcomes. While the KAPTOR-full model is preferred when post-transplant data is available due to its consistently higher net benefit, the KAPTOR-pre model remains clinically valuable, particularly at low to moderate thresholds, when only pre-transplant information is accessible. This demonstrates that both models can effectively support clinical decision-making at different stages of KT, with threshold selection carefully balanced against the clinical consequences of misclassification.
Discussion
In this study, we developed KAPTOR, a machine learning-based prognostic model for deceased donor kidney transplantation, achieving accurate prediction of 5-year graft survival through analysis of 940 patients. The XGB-based KAPTOR-full model, incorporating one-year post-transplant data, achieved superior performance (AUC 0.904) compared to existing predictive tools [33], while KAPTOR-pre demonstrated reliable pre-transplant prediction capability (AUC 0.813).
Under internal validation, KAPTOR showed higher discrimination than traditional scoring systems in our cohort. Compared to traditional scoring systems, KAPTOR showed significant superiority in C-index comparisons (p < 0.001), maintaining this advantage in multivariate analysis. From a clinical perspective, KAPTOR offers dual-timepoint risk assessment capabilities. KAPTOR-pre may provide supplementary pre-transplant risk information, while the post-transplant model provides more accurate long-term prognosis prediction, guides immunosuppression management and complication prevention. For patients with expected long-term survival, clinical teams can emphasize preventive healthcare measures, including cancer screening, cardiovascular event prevention, and quality of life improvement. For identified high-risk patients, early consideration can be given to adjusting immunosuppression regimens or evaluating re-transplantation possibilities [34].
Compared with existing donor- or pre-transplant-based prediction tools, KAPTOR demonstrated improved discrimination in our cohort. Conventional indices such as KDRI and KDPI were designed mainly to characterize donor quality and may not fully capture recipient-specific risk, post-transplant clinical evolution, or population-specific features of Chinese DDKT. This may partly explain their modest performance in our dataset. KAPTOR-pre provides an individualized pre-transplant risk estimate using donor-recipient matching information, whereas KAPTOR-full incorporates early post-transplant variables and functions as a conditional model among 1-year graft survivors. The superior performance of KAPTOR-full likely reflects its ability to capture early graft function, recipient response, and evolving post-transplant risk, which is consistent with previous evidence that dynamic or landmark models using early follow-up data outperform static baseline models. Thus, KAPTOR-pre and KAPTOR-full should be considered complementary models for different clinical time points rather than directly interchangeable prediction tools.
Key features incorporated in KAPTOR-full provided valuable insights. Postreperfusion graft biopsy Remuzzi score [19] is widely used in assessing marginal donors or elderly donor kidneys. A Remuzzi score greater than 3 typically suggests poor donor kidney quality [35], where the cumulative effects of pathological changes may lead to poor graft function recovery or long-term outcomes. Our analysis revealed that when the total Remuzzi score exceeds 2–3, the SHAP value becomes positive and shows a continuous upward trend, validating the important role of high Remuzzi scores in predicting poor graft outcomes, consistent with previous studies [36].
Previous studies have reported potential interaction effects between CIT and donor age on transplant outcomes [37].
We also observed a significant joint effect between CIT and donor age in a Generalized Additive Model (GAM) analysis (Figure S4). The SHAP dependency plot demonstrated that moderate to high values of CIT × Donor Age values (300–900 h-years) showed increased risk.
Increased renal artery RI typically reflects increased vascular resistance, potentially associated with delayed graft function, acute rejection, chronic allograft dysfunction, or other hemodynamic abnormalities such as renal vein thrombosis [38]. SHAP analysis showed graft renal artery RI exceeding 0.8 was associated with adverse outcomes, consistent with previous findings linking elevated resistance to graft dysfunction [39,40].
Unexpectedly, AR and DGF showed limited contribution to KAPTOR-full. We offer several explanations for this finding. First and most importantly, KAPTOR-full was developed on a conditional cohort restricted to patients with graft survival beyond 1 year, while AR and DGF predominantly occur early post-transplant and are strong predictors of early graft failure. Second, the influence of AR and DGF on long-term outcomes may be partially mediated through subsequent markers of graft function. After adjusting for these direct measures, their independent contribution is diluted. Third, the relatively low incidence of AR (9.9%) and DGF (21.8%) in our cohort may limit statistical power for variable selection.
Several limitations exist in our study. First, as a single-center study, the model’s generalizability requires multi-center validation. While our models demonstrated better performance compared with existing scoring systems under internal validation, these comparisons are subject to optimism bias. We acknowledge that NRI/IDI analyses were not performed in the current study; these should be included in future external validation studies to quantify the added predictive value of KAPTOR. Second, although the sample size of 940 cases is relatively substantial in similar studies, it may still limit predictive capability for rare scenarios, suggesting room for improvement in model performance. Third, due to unavailability of continuous post-transplant monitoring data and protocol biopsy results, our performance comparison was restricted to conventional risk scores rather than more sophisticated scoring systems [41,42]. Fourth, while our median follow-up of years exceeds the 5-year primary outcome window, future studies with extended follow-up are still needed to validate the model’s predictive performance over longer time horizons. Fifth, as a single-center study, the model’s generalizability is constrained by center-specific practice patterns, including immunosuppressive protocols, biopsy indications, and follow-up schedules. Multi-center external validation remains a priority for our ongoing work. Finally, while SHAP analysis provides feature importance interpretation, the ‘black box’ nature of machine learning models may affect clinicians’ intuitive understanding and application of prediction results.
The translational implications of KAPTOR should be interpreted within the framework of responsible AI use in kidney care [43]. In this context, KAPTOR should be viewed as a clinical decision support tool rather than an autonomous decision-making system. Its intended role is to assist transplant clinicians in risk stratification and to help identify recipients who may benefit from closer surveillance or individualized management, while final decisions regarding immunosuppression, biopsy, follow-up intensity, and other clinical actions should remain under physician supervision and be informed by comprehensive clinical judgment.
Before clinical deployment, KAPTOR should be prospectively validated across different regions, donor profiles, recipient populations, and clinical practice settings, integrated into electronic health record workflows without increasing clinician burden, aligned with regulatory requirements for clinical decision support software, and evaluated for clinical utility and net benefit. Future studies should determine whether KAPTOR-guided risk stratification improves surveillance efficiency, supports timely intervention, reduces unnecessary testing, and ultimately improves long-term graft and patient outcomes.
In conclusion, we developed predictive models for long-term graft survival using a Chinese deceased donor kidney transplant cohort. Our results demonstrate that equitable prediction tools can be constructed using readily available clinical data. By highlighting the potential for predicting long-term survival, this work shifts focus to long-term outcomes and provides a valuable prognostic tool for the rapidly growing deceased donor kidney transplant population. Future studies should focus on prospective validation and optimal implementation of these models in clinical practice.
Supplementary Material
Funding Statement
The work was supported by grants from the National Nature Science Foundation of China (82300852 and U21A20350).
Ethics approval and consent to participate
This study received ethical approval from First Affiliated Hospital of Zhejiang University’s Institutional Review Board (IIT20250221) and was conducted in accordance with the Declaration of Helsinki. All participants provided written informed consent. We explicitly declare that no organs were obtained from condemned executed prisoners.
Disclosure statement
No potential conflict of interest was reported by the author(s).
Data availability and materials statement
The authors confirm that the data supporting the findings of this study are available within the article or its supplementary materials. More details of the data of this study are available from the corresponding authors on request.
References
- 1.Axelrod DA, Schnitzler MA, Xiao H, et al. An economic assessment of contemporary kidney transplant practice. Am J Transplant. 2018;18(5):1168–1176. doi: 10.1111/ajt.14702. [DOI] [PubMed] [Google Scholar]
- 2.United States Renal Data System . USRDS Annual Data Report: epidemiology of kidney disease in the United States. 2024. Bethesda, MD: National Institutes of Health, National Institute of Diabetes and Digestive and Kidney Diseases; 2024. [Google Scholar]
- 3.Hariharan S, Israni AK, Danovitch G.. Long-term survival after kidney transplantation. N Engl J Med. 2021;385(8):729–743. doi: 10.1056/NEJMra2014530. [DOI] [PubMed] [Google Scholar]
- 4.Zhang Z, Liu Z, Shi B.. Global perspective on kidney transplantation: China. Kidney360. 2022;3(2):364–367. doi: 10.34067/KID.0003302021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Yaffe HC, Von Ahrens D, Urioste A, et al. Impact of deceased-donor acute kidney injury on kidney transplantation. Transplantation. 2024;108(6):1283–1295. Published online November 22. doi: 10.1097/TP.0000000000004848. [DOI] [PubMed] [Google Scholar]
- 6.Pippias M, Stel VS, Arnol M, et al. Temporal trends in the quality of deceased donor kidneys and kidney transplant outcomes in Europe: an analysis by the ERA-EDTA Registry. Nephrol Dial Transplant. 2021;37(1):175–186. doi: 10.1093/ndt/gfab156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Fang X, Wang Y, Liu R, et al. Long-term outcomes of kidney transplantation from expanded criteria donors with Chinese novel donation policy: donation after citizens’ death. BMC Nephrol. 2022;23(1):325. doi: 10.1186/s12882-022-02944-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Cohen JB, Bloom RD, Reese PP, et al. National outcomes of kidney transplantation from deceased diabetic donors. Kidney Int. 2015;89(3):636–647. doi: 10.1038/ki.2015.325. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Tennankore KK, Kim SJ, Alwayn IPJ, et al. Prolonged warm ischemia time is associated with graft failure and mortality after kidney transplantation. Kidney Int. 2016;89(3):648–658. doi: 10.1016/j.kint.2015.09.002. [DOI] [PubMed] [Google Scholar]
- 10.Zhang X, Shan H, Zhang M, et al. Donor-derived infection’s prevention and control in kidney transplantation. Transplant Proc. 2023;55(1):22–29. doi: 10.1016/j.transproceed.2022.12.009. [DOI] [PubMed] [Google Scholar]
- 11.De Vries DK, Lindeman JHN, Ringers J, et al. Donor brain death predisposes human kidney grafts to a proinflammatory reaction after transplantation. Am J Transplant. 2011;11(5):1064–1070. doi: 10.1111/j.1600-6143.2011.03466.x. [DOI] [PubMed] [Google Scholar]
- 12.Kasimsetty SG, McKay DB.. Ischemia as a factor affecting innate immune responses in kidney transplantation. Curr Opin Nephrol Hypertens. 2016;25(1):3–11. doi: 10.1097/MNH.0000000000000190. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Del Río F, Andrés A, Padilla M, et al. Kidney transplantation from donors after uncontrolled circulatory death: the Spanish experience. Kidney Int. 2019;95(2):420–428. doi: 10.1016/j.kint.2018.09.014. [DOI] [PubMed] [Google Scholar]
- 14.Peters-Sengers H, Houtzager JHE, Idu MM, et al. Impact of cold ischemia time on outcomes of deceased donor kidney transplantation: an analysis of a national registry. Transplant Direct. 2019;5(5):e448. doi: 10.1097/TXD.0000000000000888. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Rao PS, Schaubel DE, Guidinger MK, et al. A comprehensive risk quantification score for deceased donor kidneys: the kidney donor risk index. Transplantation. 2009;88(2):231–236. doi: 10.1097/TP.0b013e3181ac620b. [DOI] [PubMed] [Google Scholar]
- 16.Bachmann Q, Haberfellner F, Büttner-Herold M, et al. The kidney donor profile index (KDPI) correlates with histopathologic findings in post-reperfusion baseline biopsies and predicts kidney transplant outcome. Front Med (Lausanne). 2022;9:875206. doi: 10.3389/fmed.2022.875206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Sexton DJ, O’Kelly P, Kennedy C, et al. Assessing the discrimination of the Kidney Donor Risk Index/Kidney Donor Profile Index scores for allograft failure and estimated glomerular filtration rate in Ireland’s National Kidney Transplant Programme. Clin Kidney J. 2019;12(4):569–573. doi: 10.1093/ckj/sfy130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ravindhran B, Chandak P, Schafer N, et al. Machine learning models in predicting graft survival in kidney transplantation: meta-analysis. BJS Open. 2023;7(2):zrad011. doi: 10.1093/bjsopen/zrad011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Remuzzi G, Grinyò J, Ruggenenti P, et al. Early experience with dual kidney transplantation in adults using expanded donor criteria. J Am Soc Nephrol. 1999;10(12):2591–2598. doi: 10.1681/ASN.V10122591. [DOI] [PubMed] [Google Scholar]
- 20.National Kidney Foundation . CKD-EPI creatinine equation (2021); 2021. https://www.kidney.org/.
- 21.Uno H, Cai T, Pencina MJ, et al. On the C‐statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Stat Med. 2011;30(10):1105–1117. doi: 10.1002/sim.4154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Pessione F, Cohen S, Durand D, et al. Multivariate analysis of donor risk factors for graft survival in kidney transplantation. Transplantation. 2003;75(3):361–367. doi: 10.1097/01.TP.0000044171.97375.61. [DOI] [PubMed] [Google Scholar]
- 23.Nyberg SL, Baskin-Bey ES, Kremers W, et al. Improving the prediction of donor kidney quality: deceased donor score and resistive indices. Transplantation. 2005;80(7):925–929. doi: 10.1097/01.TP.0000173798.04043.AF. [DOI] [PubMed] [Google Scholar]
- 24.Schold JD, Kaplan B, Baliga RS, et al. The broad spectrum of quality in deceased donor kidneys. Am J Transplant. 2005;5(4 Pt 1):757–765. doi: 10.1111/j.1600-6143.2005.00770.x. [DOI] [PubMed] [Google Scholar]
- 25.Foucher Y, Daguin P, Akl A, et al. A clinical scoring system highly predictive of long-term kidney graft survival. Kidney Int. 2010;78(12):1288–1294. doi: 10.1038/ki.2010.232. [DOI] [PubMed] [Google Scholar]
- 26.Moore J, He X, Shabir S, et al. Development and evaluation of a composite risk score to predict kidney transplant failure. Am J Kidney Dis. 2011;57(5):744–751. doi: 10.1053/j.ajkd.2010.12.017. [DOI] [PubMed] [Google Scholar]
- 27.Shabir S, Halimi JM, Cherukuri A, et al. Predicting 5-year risk of kidney transplant failure: a prediction instrument using data available at 1 year posttransplantation. Am J Kidney Dis. 2014;63(4):643–651. doi: 10.1053/j.ajkd.2013.10.059. [DOI] [PubMed] [Google Scholar]
- 28.Clayton PA, Lim WH, Wong G, et al. Relationship between eGFR decline and hard outcomes after kidney transplants. J Am Soc Nephrol. 2016;27(11):3440–3446. doi: 10.1681/ASN.2015050524. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Pruett TL, Vece GR, Carrico RJ, et al. US deceased kidney transplantation: estimated GFR, donor age and KDPI association with graft survival. EClinicalMedicine. 2021;37:100980. doi: 10.1016/j.eclinm.2021.100980. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.OPTN: Organ Procurement and Transplantation Network . A guide to calculating and interpreting the kidney donor profle index (KDPI). 2020. https://optn.transplant.hrsa.gov/media/1512/guide_to_calculating_interpreting_kdpi.pdf.
- 31.OPTN: Organ Procurement and Transplantation Network . KDRI to KDPI mapping table; 2024. https://optn.transplant.hrsa.gov/media/wnmnxxzu/kdpi_mapping_table.pdf.
- 32.Lundberg SM, Lee SI.. A unified approach to interpreting model predictions. 2017;30:4765–4774. [Google Scholar]
- 33.Kim JM, Jung H, Kwon HE, et al. Predicting prognostic factors in kidney transplantation using a machine learning approach to enhance outcome predictions: a retrospective cohort study. Int J Surg. 2024;110(11):7159–7168. doi: 10.1097/JS9.0000000000002028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Davis S, Mohan S.. Managing patients with failing kidney allograft: many questions remain. Clin J Am Soc Nephrol. 2022;17(3):444–451. doi: 10.2215/CJN.14620920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Wang M, Lv J, Zhao J, et al. Postreperfusion renal allograft biopsy predicts outcome of single-kidney transplantation: a 10-year observational study in China. Kidney Int Rep. 2024;9(1):96–107. doi: 10.1016/j.ekir.2023.10.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Naesens M. Zero-time renal transplant biopsies: A comprehensive review. Transplantation. 2016;100(7):1425–1439. doi: 10.1097/tp.0000000000001018. [DOI] [PubMed] [Google Scholar]
- 37.Helanterä I, Ibrahim HN, Lempinen M, et al. Donor Age, cold ischemia time, and delayed graft function. Clin J Am Soc Nephrol. 2020;15(6):813–821. doi: 10.2215/CJN.13711119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Hidaka Y, Yamanaga S, Kawabata C, et al. The resistive index by doppler ultrasonography as a predictor of the long-term outcomes after kidney transplantation. Transplant Proc. 2023;55(4):777–781. doi: 10.1016/j.transproceed.2023.04.006. [DOI] [PubMed] [Google Scholar]
- 39.Loock MT, Bamoulid J, Courivaud C, et al. Significant increase in 1-year posttransplant renal arterial index predicts graft loss. Clin J Am Soc Nephrol. 2010;5(10):1867–1872. doi: 10.2215/CJN.01210210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Lee M, Hsu B.. Inverse association of free oxygen radicals defence (ford) with peripheral arterial stiffness in kidney transplantation patients.: Abstract# B892. Transplantation. 2014;98:518. doi: 10.1097/00007890-201407151-01736. [DOI] [Google Scholar]
- 41.Raynaud M, Aubert O, Divard G, et al. Dynamic prediction of renal survival among deeply phenotyped kidney transplant recipients using artificial intelligence: an observational, international, multicohort study. Lancet Digit Health. 2021;3(12):e795–e805. doi: 10.1016/S2589-7500(21)00209-0. [DOI] [PubMed] [Google Scholar]
- 42.Loupy A, Aubert O, Orandi BJ, et al. Prediction system for risk of allograft loss in patients receiving kidney transplants: international derivation and validation study. BMJ. 2019;366:l4923. Published online September 17. doi: 10.1136/bmj.l4923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Tangri N, Cheungpasitporn W, Crittenden SD, et al. Responsible use of artificial intelligence to improve kidney care: a statement from the American society of nephrology. J Am Soc Nephrol. 2026;37(4):881–890. doi: 10.1681/ASN.0000000929. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The authors confirm that the data supporting the findings of this study are available within the article or its supplementary materials. More details of the data of this study are available from the corresponding authors on request.





