Abstract
Artificial intelligence (AI) is reshaping kidney care by integrating electronic health records (EHRs), imaging, digital pathology, genomics, wearable biosensors, and longitudinal physiologic data into dynamic risk-prediction systems. Traditional nephrology risk stratification has relied on static variables and regression-based models such as estimated glomerular filtration rate (eGFR), albuminuria, and composite clinical scores, which often inadequately capture nonlinear disease trajectories and phenotypic heterogeneity. Advances in machine learning, deep learning, and multimodal foundation models are accelerating the shift toward predictive, preventive, and precision nephrology. This narrative review evaluated AI applications across acute kidney injury (AKI), chronic kidney disease (CKD), dialysis, and transplantation. PubMed, Embase, Scopus, Web of Science, and IEEE Xplore were searched for English-language studies published from January 2015 through July 2026, prioritizing those reporting discrimination, calibration, external validation, or implementation outcomes. Across AKI, CKD, dialysis, and kidney transplantation, AI-based models have frequently demonstrated improved discrimination relative to conventional approaches, although the magnitude of improvement varies substantially across populations, prediction targets, comparators, and validation settings. In transplantation, multimodal systems integrating histopathology, donor-derived biomarkers, and clinical variables have improved prediction of rejection and graft failure. Emerging bioprognostic frameworks incorporating wearables, dialysis telemetry, and molecular biomarkers may support dynamic or near–real-time risk estimation of hyperkalemia, intradialytic hypotension, and cardiovascular instability. Prospective validation, assessment of calibration drift, and formal fairness analyses remain limited across the published literature, while dataset shift, restricted generalizability, algorithmic opacity, and workflow integration continue to constrain clinical adoption. AI-driven prediction and bioprognostics have the potential to support a transition from predominantly reactive kidney care toward more anticipatory and continuously informed precision care, although improved predictive performance has not yet been consistently shown to translate into improved patient outcomes. Realizing this potential will require rigorous external validation, prospective clinical-impact evaluation, equitable deployment, interoperable infrastructure, and human-in-the-loop oversight.
keywords: artificial intelligence, machine learning, acute kidney injury, chronic kidney disease, risk prediction, super Intelligence
Introduction
Kidney disease represents a growing global public health challenge, affecting hundreds of millions of individuals worldwide and contributing substantially to cardiovascular morbidity, mortality, healthcare utilization, and economic burden.1,2 Chronic kidney disease (CKD) affects approximately 10% of the global population,1,2 while acute kidney injury (AKI) complicates up to 20% of hospitalizations and remains strongly associated with progression to kidney failure, cardiovascular disease, and death.3,4 Simultaneously, increasing numbers of patients require maintenance dialysis or kidney transplantation, further intensifying demands on healthcare systems.1,5 Despite advances in therapeutics and supportive care, nephrology remains largely reactive, with interventions frequently initiated only after substantial physiologic deterioration or irreversible nephron loss has occurred.2,6
Traditional kidney risk stratification has relied predominantly on static clinical variables and regression-based prediction tools, including estimated glomerular filtration rate (eGFR), albuminuria, and composite models such as the Kidney Failure Risk Equation (KFRE).7,8 Although these approaches have improved prognostication at the population level, they incompletely capture the dynamic, nonlinear, and heterogeneous trajectories characteristic of kidney disease. Conventional models are typically derived from episodic measurements obtained at isolated clinical encounters and often fail to incorporate evolving physiologic trends, multimorbidity interactions, temporal variability, or high-dimensional biologic data.9,10 As a result, clinically meaningful deterioration may remain undetected until advanced disease manifestations emerge.9
Concurrently, the rapid expansion of digital health ecosystems has transformed the quantity and granularity of data available for kidney care.11 Modern nephrology increasingly generates longitudinal multimodal datasets through electronic health records (EHRs), wearable physiologic biosensors, dialysis telemetry, digital pathology, imaging platforms, and multi-omics technologies including genomics, transcriptomics, proteomics, and metabolomics.12–14 These data streams provide an opportunity to move beyond static risk estimation toward continuously adaptive models capable of dynamic clinical surveillance and individualized prognostication.
Advances in artificial intelligence (AI), including machine learning, deep learning, transformer architectures, and multimodal foundation models, are expanding opportunities for predictive and precision approaches in nephrology.11,15 These technologies increasingly enable the integration of longitudinal clinical, physiologic, imaging, pathologic, and molecular data and create opportunities to move beyond conventional static risk prediction. Emerging AI-driven predictive models have demonstrated promising discrimination and risk-stratification performance for AKI detection, CKD progression forecasting, dialysis instability prediction, and transplant outcome assessment compared with traditional approaches.11,15 However, improved predictive performance should not be equated with clinical effectiveness. Technical gains in discrimination, earlier prediction, or multimodal integration do not by themselves establish transportability, actionability, safety, equity, or improvement in patient outcomes. Randomized and prospective implementation studies remain limited, and available evidence demonstrates that accurate or earlier risk identification may fail to improve outcomes when predictions are not coupled to effective, timely, and consistently implemented clinical interventions. Accordingly, this review adopts a deliberately critical translational perspective, evaluating not only where AI models perform well but also where validation is incomplete, implementation has not improved outcomes, or evidence remains insufficient for routine clinical use.
This narrative review focuses on AI-driven predictive modeling and the proposed AI-bioprognostics framework across major domains of kidney care, including AKI, CKD, dialysis, and kidney transplantation. We emphasize patient-level prognostication, longitudinal risk prediction, multimodal data integration, and clinically relevant applications incorporating EHR data, physiologic monitoring, imaging and digital pathology, molecular biomarkers, wearable technologies, and related longitudinal data streams. We also consider emerging applications in pediatric nephrology, glomerular and rare kidney diseases, and home dialysis, where evidence remains comparatively limited but clinically relevant. The review is not intended to provide a comprehensive overview of all applications of AI in nephrology. We therefore exclude stand-alone imaging segmentation or classification studies without a prognostic or clinically relevant risk-prediction component, AI applications primarily directed toward drug discovery or molecular target identification, and predominantly administrative or operational uses such as coding, scheduling, billing, or resource allocation. Imaging and pathology applications are included when they contribute directly to disease characterization, risk prediction, or outcome prognostication.
This review complements, but is distinct from, prior work on the broader clinical integration of AI in nephrology and on workflow-integrated clinical AI agents.15,16 Whereas those reviews emphasized the overall roadmap for AI implementation and the transition from prediction to agent-enabled clinical action, respectively, the present narrative review focuses specifically on AI-driven prognostication and longitudinal multimodal risk assessment. We examine how clinical, physiologic, imaging, pathologic, molecular, and wearable-derived data may be integrated across time to support dynamic individualized prediction in AKI, CKD, dialysis, and kidney transplantation. Within this narrower focus, we propose the concept of “AI-bioprognostics” as an organizing framework for continuously updated, multimodal prognostic intelligence. Thus, the primary contribution of this review is to synthesize current predictive applications and define a forward-looking framework that links conventional risk prediction with longitudinal biologic and physiologic data integration, rather than to provide another general overview of AI implementation or agentic clinical workflows.
In this review, we propose the term “AI-bioprognostics” to describe a conceptual framework in which AI-enabled systems integrate biologic, molecular, physiologic, imaging, and longitudinal clinical information to generate dynamic, individualized prognostic estimates and potentially actionable therapeutic insights. We use this term as a proposed organizing framework rather than as established terminology in nephrology or clinical informatics. The framework builds upon existing developments in AI-enabled clinical prediction, multimodal data integration, digital biomarkers, and precision kidney care.11,15,17 Rather than assuming that such systems will improve care, we use this framework to examine the conditions under which longitudinal multimodal prediction may become clinically useful, including external validation, calibration, fairness, workflow integration, human oversight, and prospective demonstration of benefit. Together, these approaches may support a transition from predominantly reactive kidney care toward more anticipatory, continuously informed, and precision-guided clinical management (Figure 1).
Figure 1.

Evolution from Traditional Risk Prediction to AI-Driven Predictive Kidney Care. The schematic contrasts traditional episodic risk assessment based on limited clinical variables with an AI-bioprognostic framework integrating longitudinal clinical data, laboratory trajectories, imaging, digital pathology, molecular and multi-omics information, physiologic monitoring, and wearable or device-derived data to generate dynamically updated individualized risk estimates and clinically relevant prognostic outputs. Horizontal arrows indicate the conceptual progression from traditional reactive kidney care, through expansion of the digital kidney ecosystem, toward continuously adaptive predictive kidney care. Arrows within the predictive-care panel indicate the flow from multimodal data integration to continuous physiologic surveillance, dynamic risk recalibration, individualized prognostication, and precision intervention; they do not denote an increase or decrease in a quantitative clinical variable. The lower horizontal continuum illustrates the intended transition from delayed recognition of disease toward earlier prediction and more anticipatory, targeted intervention.
Foundations of AI-Driven Predictive Modeling in Nephrology
AI-driven predictive modeling builds upon established statistical approaches traditionally used in nephrology while expanding the ability to analyze high-dimensional, nonlinear, and longitudinal data.17 Classical prediction models, including logistic regression and Cox proportional hazards models, typically specify functional relationships between predictors and outcomes and rely on modeling assumptions that may include linearity on the chosen scale and, for Cox models, proportional hazards over time.18 These approaches can incorporate nonlinear terms, interactions, time-varying covariates, and other extensions, but their performance depends on appropriate model specification. In contrast, many machine-learning approaches can flexibly learn complex nonlinear relationships and higher-order interactions from multidimensional data with less reliance on prespecified functional forms. Widely used nephrology tools such as the KFRE have substantially improved individual-level risk estimation and clinical prognostication;7,8 however, the KFRE and many other commonly used clinical prediction models are based on a limited set of variables obtained at defined clinical time points and therefore may not fully capture evolving longitudinal patterns or the dynamic and heterogeneous trajectories characteristic of kidney disease progression.9,10
In contrast, machine learning systems are capable of identifying complex nonlinear relationships, high-order interactions, and temporal dependencies across large multidimensional datasets without requiring explicit specification of functional relationships between variables.17 Conventional statistical models are typically “fixed” after derivation, whereas many AI systems can be periodically recalibrated or updated as new data become available.11,15 This distinction is particularly relevant in nephrology, where physiologic states frequently evolve over time in response to comorbid illness, medications, dialysis-related factors, and acute clinical events.
Several major AI methodologies have emerged within kidney care. Supervised learning approaches, including gradient boosting and random forests, are trained using labeled outcomes such as AKI, kidney failure, or mortality and currently represent the most widely deployed predictive systems in nephrology.15,17 Unsupervised learning techniques identify latent phenotypes and hidden patient clusters without predefined outcomes, enabling recognition of biologically and clinically distinct kidney disease subgroups.19–21 Reinforcement learning models further extend predictive analytics by optimizing sequential clinical decisions through iterative feedback and reward-based learning, with emerging applications in dialysis management and hemodynamic optimization.22,23 Deep learning architectures, particularly convolutional neural networks (CNNs), have demonstrated strong performance in imaging and digital pathology interpretation, including fibrosis quantification and glomerular classification.24,25 More recently, transformer architectures and multimodal foundation models have enabled integration of longitudinal clinical data, free-text notes, imaging, physiologic telemetry, and molecular information within unified predictive frameworks.26,27
The rapid expansion of digital health ecosystems has accelerated development of kidney AI models.11,15 Contemporary systems increasingly integrate longitudinal EHR data, laboratory trajectories, imaging, whole-slide histopathology, wearable physiologic monitoring, dialysis telemetry, and multi-omics platforms including genomics, proteomics, metabolomics, and transcriptomics.11,15,28–30 These multimodal datasets support the transition from static risk estimation toward continuously adaptive prediction systems capable of near–real-time clinical surveillance.11,15,31,32
Evaluation of predictive performance requires rigorous statistical assessment beyond discrimination alone.33,34 Common metrics include the area under the receiver-operating-characteristic curve (AUROC) and C-statistics, which quantify discriminative ability, as well as calibration, which assesses agreement between predicted and observed outcomes.33,34 Precision-recall curves may provide additional insight in imbalanced clinical datasets.33–35 External validation across geographically and demographically distinct populations remains essential to establish generalizability and transportability.36,37 Increasing attention has also focused on calibration drift, dataset shift, and temporal degradation in model performance following clinical deployment, emphasizing the need for ongoing monitoring and recalibration of AI systems in real-world nephrology practice.38–40 The major AI methodologies currently applied in nephrology, along with their strengths, limitations, and clinical applications, are summarized in Table 1.
Table 1.
Common AI Methodologies in Kidney Predictive Modeling
| AI Methodology | Core Principle | Common Kidney Applications | Major Strengths | Key Limitations |
|---|---|---|---|---|
| Logistic Regression | Regression-based probability modeling | CKD progression, AKI risk scores41 | Interpretable, clinically familiar | Limited nonlinear modeling |
| Cox Proportional Hazards | Time-to-event survival analysis | Kidney failure prediction, graft survival42 | Handles censored outcomes | Assumes proportional hazards |
| Random Forests | Ensemble decision trees | AKI prediction, dialysis outcomes43 | Captures nonlinear interactions | Reduced interpretability |
| Gradient Boosting | Sequential error-correcting trees | CKD progression, hospitalization prediction44 | High predictive accuracy | Risk of overfitting |
| Unsupervised Clustering | Identification of latent subgroups | AKI phenotyping, CKD subtypes21,45 | Detects hidden phenotypes | Clinical interpretability challenges |
| Reinforcement Learning | Reward-based sequential optimization | Dialysis management, hemodynamic optimization46 | Adaptive decision support | Limited prospective validation |
| Convolutional Neural Networks (CNNs) | Deep image feature extraction | Digital pathology, imaging interpretation47 | Strong imaging performance | High computational requirements |
| Transformer Architectures | Longitudinal sequence modeling | EHR prediction, multimodal integration48 | Handles temporal complexity | Large data requirements |
| Foundation Models | Multimodal large-scale representation learning | Clinical copilots, integrated prediction49 | Broad contextual integration | Operational and governance challenges |
Notes: Methodology names are presented without study-specific citations because these represent general statistical or machine-learning approaches rather than methods originating from individual nephrology studies. References accompanying the clinical applications identify representative studies in which the corresponding methods have been evaluated in kidney care. General methodological principles and evaluation considerations are discussed in the main text.
Abbreviations: AI, Artificial Intelligence; AKI, Acute Kidney Injury; CKD, Chronic Kidney Disease; CNNs, Convolutional Neural Networks; EHR, Electronic Health Record.
AI Predictive Modeling in AKI
AKI represents one of the most extensively studied applications of AI in nephrology. AKI frequently develops rapidly in hospitalized and critically ill patients, often preceding detectable rises in serum creatinine by several hours to days. Because delayed recognition is strongly associated with increased mortality, dialysis dependence, prolonged hospitalization, and progression to CKD, AI-driven prediction systems have emerged as promising tools for earlier detection and proactive intervention.17,50
Early AKI Prediction
AI-based AKI prediction models increasingly leverage high-frequency EHR data to enable continuous physiologic surveillance within inpatient and intensive care unit (ICU) settings. Contemporary systems integrate longitudinal laboratory trajectories, vital signs, medication exposures, hemodynamic variables, fluid balance, vasopressor requirements, and clinical documentation to identify evolving patterns preceding overt kidney injury. Machine learning approaches including gradient boosting, random forests, recurrent neural networks, and transformer-based architectures have demonstrated the ability to detect AKI risk substantially earlier than conventional rule-based systems.51
Several EHR-triggered surveillance platforms have reported strong discriminative performance, with AUROC values commonly ranging from approximately 0.85 to 0.93 for prediction of moderate-to-severe AKI within 12 to 48 hours prior to creatinine-based diagnosis.51–54 Importantly, many models maintain high negative predictive values, supporting potential utility in identifying low-risk patients while prioritizing high-risk individuals for earlier nephrology consultation or preventive interventions.54 Real-time deployment strategies increasingly emphasize dynamic longitudinal monitoring rather than isolated snapshot prediction, allowing continuous recalibration of risk as new clinical data become available.55,56
A prominent example is the deep-learning model developed by Tomašev et al using longitudinal EHR data from 703,782 adults within the US Department of Veterans Affairs, which demonstrated the ability to predict a substantial proportion of AKI events up to 48 hours in advance.50 However, the source population was predominantly male, reflecting the demographic composition of the Veterans Affairs population, and this limits assumptions regarding direct transportability to more sex-balanced or demographically distinct populations.
Prediction of Severe Outcomes
Beyond early AKI detection, AI systems have been developed to predict clinically significant downstream outcomes, including dialysis requirement, persistent AKI, ICU transfer, prolonged hospitalization, and mortality. Models incorporating temporal physiologic trends and multiorgan dysfunction patterns frequently outperform conventional severity scores in identifying patients likely to experience progressive kidney injury or hemodynamic deterioration.57 Deep learning systems trained on large ICU datasets have demonstrated improved discrimination for prediction of renal replacement therapy initiation and mortality compared with traditional risk models, particularly among critically ill populations with complex comorbidity profiles.58,59
Emerging studies further suggest that AI-based trajectory modeling may distinguish transient from persistent AKI phenotypes, potentially enabling earlier identification of patients at elevated risk for maladaptive renal recovery or progression to CKD. These approaches are particularly relevant because serum creatinine kinetics alone often inadequately characterize biologic heterogeneity across AKI syndromes.60,61
Comparative Performance and Limitations
AI-driven prediction systems have frequently demonstrated improved predictive performance compared with conventional nephrology and critical care scoring systems, including logistic regression–based models, Sequential Organ Failure Assessment (SOFA) scores, and KDIGO stage–based approaches.57,62 However, the magnitude of improvement varies substantially across studies because of differences in patient populations, prediction targets, time horizons, comparator models, and performance metrics. Accordingly, comparative performance is better interpreted using study-specific AUROC, C-statistic, precision-recall, calibration, and other reported metrics rather than a single summary percentage improvement. Precision-recall analyses have additionally demonstrated improved performance in imbalanced ICU datasets where severe AKI events are relatively infrequent.62,63
Importantly, improved discrimination or earlier detection should not be assumed to translate into improved patient outcomes. Randomized implementation evidence provides an important counterpoint to retrospective studies reporting strong predictive performance. In the randomized controlled trial by Wilson et al,64 automated electronic AKI alerts did not improve clinical outcomes despite providing earlier notification of kidney injury, and signals of possible harm were observed in some patient strata. Although this intervention was an electronic AKI alert rather than a contemporary machine-learning prediction system, the study illustrates a central translational principle that remains directly relevant to AI-enabled prediction: identifying risk earlier is clinically useful only when the information leads to an effective, timely, and appropriately targeted change in management. Related implementation evidence further demonstrates that the clinical value of predictive systems depends not only on model accuracy but also on alert design, workflow integration, clinician response, actionability, and the availability of interventions capable of modifying the predicted outcome.57,64
Accordingly, the principal challenge for AI-enabled AKI prediction is not simply to maximize AUROC or extend the prediction horizon, but to demonstrate incremental clinical utility beyond existing care. Prospective evaluation should therefore assess whether deployment changes clinician behavior, reduces preventable nephrotoxic exposures or hemodynamic insults, improves patient-centered outcomes, and avoids unintended consequences such as alert fatigue, unnecessary testing, or inappropriate intervention. These considerations underscore the distinction between model performance and clinical effectiveness and support the need for prospective, preferably randomized, implementation studies before assuming that more accurate prediction will improve kidney outcomes.
Despite these advances, important translational limitations remain. Many published models rely on retrospective single-center datasets and may demonstrate reduced transportability across healthcare systems with differing patient populations, laboratory workflows, or EHR structures. Alert fatigue and excessive false-positive notifications remain important operational concerns, particularly when predictive systems are deployed without clearly actionable response pathways.56,65 Additionally, relatively few studies have undergone rigorous prospective validation or assessed long-term calibration drift following real-world deployment. These limitations underscore the need for cautious implementation, continuous monitoring, and human-in-the-loop oversight as AI-enabled AKI prediction systems move from experimental development toward clinical integration.66,67 A representative workflow illustrating real-time AI-enabled AKI prediction and dynamic risk recalibration within the EHR is shown in Figure 2.
Figure 2.

AI-Enabled AKI Prediction Workflow within the EHR. Conceptual workflow illustrating real-time AI-based AKI prediction using longitudinal EHR data. Clinical variables, including laboratory values, vital signs, medications, clinical context, and hemodynamic parameters, are integrated through preprocessing and machine-learning algorithms to generate continuously updated risk estimates. Dynamic risk stratification may support earlier nephrology evaluation and targeted preventive measures through human-in-the-loop oversight. Horizontal arrows indicate the direction of data and workflow progression, whereas upward and downward arrows in the outcomes panel indicate the intended direction of potentially favorable outcomes (increased renal recovery and decreased severe AKI, dialysis requirement, and length of stay) and do not represent observed effect sizes or established clinical benefit.
Predictive Modeling in CKD and Dialysis
AI-driven predictive modeling is increasingly transforming the management of CKD and dialysis by enabling dynamic longitudinal risk assessment, continuous physiologic surveillance, and individualized therapeutic optimization. Unlike AKI, which often develops over hours to days, CKD progression typically unfolds over prolonged and heterogeneous trajectories influenced by diabetes, hypertension, cardiovascular disease, inflammation, medication exposure, and socioeconomic determinants. Conventional nephrology models frequently rely on static measurements obtained at isolated clinical encounters and may inadequately characterize temporal variability and nonlinear disease progression. AI systems have therefore emerged as promising tools for continuously adaptive prediction across the CKD continuum.
CKD Progression Prediction
Machine learning models have demonstrated improved performance for prediction of eGFR decline, kidney failure, hospitalization, and cardiovascular events compared with conventional regression-based approaches. Contemporary AI systems increasingly integrate longitudinal laboratory trajectories, medication exposures, comorbidity patterns, blood pressure variability, imaging findings, and clinical documentation derived from EHRs.10 Gradient boosting models,44 recurrent neural networks,68 and transformer-based architectures15,48 have shown particular promise in capturing complex temporal interactions underlying CKD progression.
Several studies have reported improved discrimination for prediction of kidney failure and rapid eGFR decline relative to traditional risk equations, with gains particularly evident among patients with irregular disease trajectories or multiple interacting comorbidities.9,42,44 AI models have additionally demonstrated utility in forecasting hospitalization risk, cardiovascular instability, and healthcare utilization, supporting broader operational applications within population kidney health management.69–71 More recently, large-scale external validation of the Klinrisk machine-learning model in approximately 4.8 million commercially insured, Medicare, and Medicaid beneficiaries demonstrated robust prediction of 2-year CKD progression across diverse US populations and superior discrimination compared with the Kidney Disease: Improving Global Outcomes (KDIGO) heatmap-based staging system.9 Klinrisk is a commercially developed AI-enabled prognostic platform. Although this study evaluated the model in external US datasets distinct from the original Canadian development setting, several study authors were affiliated with Klinrisk Inc.; accordingly, the findings represent external validation in independent data but should not be interpreted as fully investigator-independent validation. This distinction is relevant when assessing the strength and generalizability of evidence supporting commercially developed predictive systems.
A notable example of clinically implemented AI-enabled kidney prognostication is KidneyIntelX.dkd,72 a machine-learning–derived prognostic assay that integrates circulating biomarkers with clinical variables to stratify the risk of progressive kidney-function decline in patients with diabetic kidney disease. The assay incorporates plasma biomarkers, including soluble tumor necrosis factor receptors 1 and 2 and kidney injury molecule-1, together with clinical information to generate an individualized risk score. KidneyIntelX.dkd72 received FDA De Novo marketing authorization in 2023, representing an important transition from retrospective predictive modeling toward regulated clinical deployment of AI-enabled kidney risk stratification. Clinical and real-world implementation studies have further evaluated its ability to refine risk assessment and influence disease-management pathways in early diabetic kidney disease.
AI Versus the KFRE
The KFRE remains one of the most widely validated tools for prediction of progression to kidney failure and has substantially improved standardized CKD prognostication. However, KFRE primarily relies on static variables including age, sex, eGFR, and albuminuria obtained at single time points.7 Many contemporary AI models extend beyond conventional static risk estimation by incorporating longitudinal trajectory modeling and continuously evolving clinical data streams.
Transformer architectures and temporal deep learning systems are particularly well suited for modeling irregularly sampled longitudinal clinical data, enabling dynamic recalibration of individualized kidney risk over time. Rather than generating a fixed prediction at a single encounter, these systems continuously reinterpret evolving laboratory trends, medication changes, hemodynamic patterns, and intercurrent clinical events.26,27 Such approaches may better capture biologic heterogeneity and temporal fluctuations in CKD progression risk compared with traditional regression-based models.
Dialysis Applications
AI-driven prediction systems are also expanding across dialysis care, where physiologic instability and high-frequency data generation create opportunities for continuous monitoring and precision-guided intervention. Machine learning models have demonstrated utility in prediction of intradialytic hypotension,73 hyperkalemia,74 volume overload,75 hospitalization,76 and mortality77 among maintenance dialysis populations. Dialysis telemetry systems incorporating ultrafiltration parameters, blood pressure trends, conductivity measurements, and treatment-session dynamics enable near–real-time physiologic surveillance during dialysis therapy.78,79 Such telemetry systems may be vendor- or device-specific, and models developed using proprietary device data or platform-specific measurements may not necessarily generalize across dialysis technologies, manufacturers, or healthcare environments. Accordingly, external validation across independent devices and clinical settings should be considered when evaluating the transportability of telemetry-based prediction systems.
Emerging reinforcement learning and adaptive prediction approaches further support individualized dialysis prescription optimization, including ultrafiltration adjustment and hemodynamic management.46 These systems may ultimately facilitate “precision dialysis”, in which treatment parameters are dynamically individualized according to continuously evolving physiologic responses rather than standardized population-based protocols.
Remote Monitoring and Wearables
Rapid expansion of remote physiologic monitoring technologies has further accelerated development of AI-enabled longitudinal kidney surveillance systems. Home blood pressure monitoring,80 smartwatch-derived physiologic data,74 wearable biosensors,81 and remote dialysis telemetry82 increasingly provide continuous streams of real-world patient data outside traditional clinical settings. Emerging biosensor platforms capable of monitoring potassium, hydration status, physical activity, and cardiovascular parameters may further extend AI-based risk prediction into ambulatory and home dialysis environments.
Home dialysis represents a particularly relevant setting in which longitudinal and patient-generated data may enable forms of prediction that differ from those used in conventional in-center dialysis.83 Peritoneal dialysis and home hemodialysis can generate repeated treatment, physiologic, device, adherence, and symptom data outside the traditional clinical encounter.84 Integration of these data with EHR information may support dynamic prediction of treatment instability, technique failure, hospitalization, volume-related complications, infection risk, and the need for clinical intervention. However, the evidence base for AI-enabled prediction in home dialysis remains less mature than for many hospital-based applications. Models developed for this setting must also account for missing or irregularly captured data, heterogeneity across device platforms, variation in patient engagement and digital literacy, connectivity limitations, and the possibility that unequal access to digital technologies could exacerbate disparities in AI-enabled monitoring and care.
Collectively, these advances support a transition from episodic CKD management toward continuously adaptive predictive nephrology centered on dynamic longitudinal monitoring, earlier detection of physiologic deterioration, and proactive individualized intervention. Current clinical applications of AI-driven predictive modeling across CKD and dialysis are summarized in Table 2.
Table 2.
Current Clinical Applications of AI Predictive Modeling Across CKD and Dialysis
| Clinical Domain | AI Applications | Common Data Sources | Potential Clinical Impact | Current Limitations |
|---|---|---|---|---|
| CKD Progression42 | eGFR decline prediction, kidney failure forecasting | EHRs, laboratory trajectories, medications | Earlier identification of rapid progressors | Limited prospective validation |
| Cardiovascular Risk71 | Heart failure and cardiovascular event prediction | Vital signs, imaging, comorbidity data | Improved risk stratification | Dataset heterogeneity |
| Hospitalization Prediction76 | Admission and readmission forecasting | EHRs, utilization patterns | Population health management | Calibration drift |
| Intradialytic Hypotension73 | Hemodynamic instability prediction | Dialysis telemetry, BP trends | Improved dialysis safety | Workflow integration challenges |
| Hyperkalemia Prediction74 | Potassium instability forecasting | Laboratory trends, telemetry, wearables | Earlier intervention | Limited wearable validation |
| Volume Overload75 | Fluid status prediction | Bioimpedance, dialysis data, wearables | Precision ultrafiltration management | Sensor variability |
| Mortality Prediction77 | Survival and complication modeling | Longitudinal multimodal datasets | Individualized prognostication | Generalizability concerns |
| Remote Monitoring79 | Home surveillance and physiologic monitoring | Smartwatches, biosensors, home BP | Continuous outpatient monitoring | Data interoperability limitations |
Abbreviations: AI, artificial intelligence; BP, blood pressure; CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate; EHRs, electronic health records.
Pediatric, Glomerular, and Rare Kidney Diseases
Pediatric nephrology and glomerular and rare kidney diseases represent important but comparatively underdeveloped areas for AI-enabled prognostication. Pediatric prediction models must account for age-dependent physiology, growth and developmental changes, differences in disease distribution, and smaller available datasets, all of which may limit direct transfer of models developed in adult populations.85 Glomerular and rare kidney diseases similarly encompass heterogeneous phenotypes, relatively small patient cohorts, and infrequent clinical outcomes, creating challenges for conventional model development, external validation, and reliable estimation of subgroup performance. At the same time, these conditions may be particularly well suited to multimodal approaches that integrate longitudinal clinical trajectories with kidney histopathology, molecular biomarkers, genomics, transcriptomics, proteomics, and other disease-specific data.86,87 Such integration may support more refined disease phenotyping, risk stratification, and prediction of treatment response or disease progression. Future progress will likely require multicenter and international data sharing, harmonized phenotyping and outcome definitions, collaborative or federated model development, and rigorous validation across age groups, disease subtypes, healthcare settings, and geographic populations.
Across these kidney-care domains, reported predictive performance is frequently strong, but the maturity of supporting validation evidence varies substantially.9,52,73,88–90 Figure 3 provides a descriptive comparison of representative AI-enabled prediction models across AKI, CKD, in-center dialysis, home dialysis, glomerular disease, and kidney transplantation, displaying reported AUROC or C-statistic values together with validation status.
Figure 3.

Reported predictive performance and validation maturity of representative AI-enabled models across kidney care. Reported AUROC or C-statistic values are shown for representative models across acute kidney injury, chronic kidney disease, in-center dialysis, home dialysis, glomerular disease, and kidney transplantation. Filled circles indicate externally or temporally validated point estimates, whereas the open circle denotes internal/test-set validation. Horizontal whiskers indicate 95% confidence intervals where reported, and thicker horizontal bars indicate the reported range of performance across multiple validation cohorts rather than confidence intervals. Validation status is categorized as internal/test-set, temporal, external, or external multicenter validation. The figure is descriptive, does not represent a pooled analysis, and is not intended to support direct comparison or ranking across heterogeneous studies with different populations, prediction targets, time horizons, and validation strategies.
AI-Bioprognostics and Multimodal Kidney Intelligence
The emergence of multimodal AI systems is expanding nephrology beyond conventional predictive analytics toward integrated biologic and physiologic intelligence platforms. Traditional kidney risk prediction has generally relied on individual clinical variables or single-domain biomarkers;91 however, kidney diseases are biologically heterogeneous, temporally dynamic, and influenced by complex interactions across molecular, physiologic, imaging, and environmental domains. Increasingly, AI systems are being designed to integrate these diverse data streams into unified prognostic ecosystems capable of continuously adaptive and individualized risk prediction.29
Proposed Framework of AI-Bioprognostics
Building on the definition introduced in the Introduction, the proposed AI-bioprognostics framework extends conventional prediction by emphasizing longitudinal multimodal integration, repeated updating of individualized risk, and characterization of evolving disease states and physiologic trajectories rather than reliance on isolated variables or static clinical snapshots.92
The conceptual basis of this proposed framework reflects the increasing convergence of several previously distinct areas of kidney medicine, including longitudinal EHR-based prediction, digital biomarkers and wearable sensing, computational imaging and pathology, molecular and multi-omics profiling, and multimodal AI. Although these individual domains are supported by an expanding literature, their integration into a unified longitudinal prognostic framework remains an evolving area. We therefore use AI-bioprognostics as an organizing construct for considering how these complementary data streams might be integrated across AKI, CKD, dialysis, and kidney transplantation.
This transition from isolated biomarkers and episodic risk estimation toward integrated longitudinal prognostic systems may be particularly relevant in nephrology, where conventional markers such as serum creatinine incompletely reflect concurrent tissue injury, fibrosis, inflammation, hemodynamic instability, or immunologic activation.8 Multimodal AI systems may enable changes across these domains to be synthesized over time, potentially allowing clinically important deterioration to be recognized before it becomes apparent through conventional measures alone.
Digital Pathology and Imaging
Digital pathology represents one of the most rapidly advancing applications of multimodal AI in kidney care. Whole-slide imaging platforms combined with convolutional neural networks (CNNs) enable automated characterization of glomeruli, tubulointerstitial fibrosis, vascular injury, and inflammatory infiltrates with increasing precision. AI-based image analysis systems have demonstrated strong performance in glomerular segmentation, lesion classification, fibrosis quantification, and prediction of kidney outcomes from biopsy specimens.14,24,25,47
Similarly, AI-assisted imaging analysis using ultrasound,93 computed tomography,94 and magnetic resonance imaging95 increasingly supports noninvasive assessment of kidney structure and disease progression. Radiomics approaches may further extract high-dimensional imaging features not readily detectable through conventional visual interpretation, potentially enabling earlier identification of fibrosis, vascular remodeling, or allograft dysfunction.
Molecular and Omics Integration
Rapid advances in genomics, transcriptomics, proteomics, and metabolomics have generated increasingly complex multidimensional datasets within nephrology. AI systems are particularly well suited for integration of these high-dimensional molecular data because they can identify nonlinear interactions and latent biologic patterns beyond human interpretability. Machine learning approaches have been used to identify molecular signatures associated with CKD progression, AKI recovery, fibrosis, and transplant rejection.30,91,92
In kidney transplantation, donor-derived cell-free DNA, transcriptomic signatures, and proteomic biomarkers increasingly complement conventional laboratory monitoring. Integration of these molecular signals with longitudinal clinical data may support earlier detection of immune activation and subclinical allograft injury.96,97
Kidney Transplantation Applications
Kidney transplantation represents one of the most data-rich environments in nephrology and is therefore particularly well suited for multimodal AI integration. AI systems have been developed for prediction of delayed graft function,98 acute rejection,96,99 graft survival,88 infection risk,100,101 and post-transplant complications.102 Multimodal frameworks integrating histopathology, donor characteristics, immunologic profiles, molecular biomarkers, and longitudinal laboratory trajectories have demonstrated improved predictive performance compared with conventional isolated-variable approaches.97,103 More recently, integrative platforms such as the iBox prognostic model have demonstrated the value of combining longitudinal clinical, functional, immunologic, and histopathologic data to provide dynamic prediction of long-term allograft survival. iBox was developed through an international academic collaboration and externally validated in geographically distinct European and US cohorts, supporting transportability across multiple transplant populations.104 The algorithm has subsequently been subject to commercial licensing and commercialization arrangements. Because major validation studies included investigators involved in development of the model, the use of external validation cohorts should be distinguished from fully investigator-independent validation. These considerations do not negate the reported predictive performance but are relevant to interpretation of the maturity and independence of the supporting evidence. However, recent randomized evidence also demonstrates that accurate risk prediction alone may be insufficient to improve clinical care.105 This gap between predictive performance and clinical benefit is considered in greater detail below together with analogous evidence from AKI alert systems. Dynamic graft-state modeling may ultimately enable continuously adaptive surveillance of transplant recipients, supporting earlier intervention before irreversible allograft injury develops.
Prediction without Benefit: from Risk Estimation to Clinical Utility
An increasingly important lesson from kidney prediction research is that accurate or earlier prediction does not necessarily improve clinical outcomes. Electronic AKI alert systems provide a well-established example. A systematic review and meta-analysis of electronic AKI alert systems demonstrated the difficulty of translating earlier recognition into consistent clinical benefit.56,57 Similarly, in the randomized controlled trial by Wilson et al,64 automated electronic AKI alerts did not improve clinical outcomes despite providing earlier notification of kidney injury, with signals of possible harm observed in some patient strata. Although these systems differ from contemporary machine-learning prediction models, they illustrate a central translational principle that remains directly relevant to AI-enabled prediction: identifying risk earlier is clinically useful only when that information leads to an effective, timely, and appropriately targeted change in management. More recently, this principle was tested directly in the ESTOP-AKI randomized clinical trial,106 in which a machine-learning AKI risk score was used to identify hospitalized patients at elevated risk for stage 2 AKI and trigger structured early nephrology consultation. Among 180 randomized patients, early nephrology consultation substantially increased nephrology involvement and generated more management recommendations but did not significantly improve the primary outcome of 7-day change in serum creatinine compared with usual care (0.04 versus −0.03 mg/dL; P =0.30). The intervention also did not significantly reduce development of stage 1 or higher AKI (42% versus 36%; P =0.47), stage 2 or higher AKI (19% versus 13%; P =0.28), 90-day readmission (34.1% versus 44.4%; P =0.21), or 90-day mortality (14.8% versus 18.7%; P =0.62).106 Notably, several medication, fluid, diuretic, and vasopressor recommendations were followed less frequently in the intervention group, further illustrating that model-generated risk identification alone does not ensure effective downstream implementation.
A similar principle has now been demonstrated in kidney transplantation. In a randomized trial of EHR-integrated AI risk prediction, the system accurately identified transplant recipients at increased risk of graft loss, yet passive presentation of these risk estimates did not improve shared decision-making or clinical outcomes.105,107 This finding highlights the distinction between identifying risk and changing care: a prediction may be technically valid but clinically ineffective when it is not accompanied by actionable recommendations, appropriate workflow integration, or a clearly defined response pathway.
Taken together, these studies suggest that evaluation of predictive AI in kidney care should extend beyond discrimination and calibration to demonstration of clinical utility. Relevant outcomes include whether the system changes clinician behavior, enables timely modification of nephrotoxic exposure or hemodynamic risk, improves shared decision-making, reduces preventable complications, and ultimately improves patient-centered outcomes. Implementation factors such as alert burden, timing, clinician trust, actionability, workflow integration, and the availability of effective downstream interventions may determine whether technically accurate prediction produces meaningful benefit or simply adds information without improving care. Thus, the most meaningful benchmark for AI-driven kidney prediction is not whether a model can predict an adverse outcome earlier, but whether that prediction enables an intervention that meaningfully alters the patient’s clinical trajectory.
Foundation Models and Multimodal AI
Recent advances in transformer architectures and multimodal foundation models are further accelerating development of integrated kidney intelligence systems. Unlike conventional single-task prediction models, foundation models can process heterogeneous data modalities simultaneously, including clinical notes, laboratory trajectories, imaging, pathology, telemetry, and molecular information. These architectures may function as longitudinal synthesis engines capable of contextual interpretation across complex clinical environments.48,49,108
Future applications may include AI-assisted clinical copilots that support longitudinal kidney disease surveillance, risk prioritization, diagnostic synthesis, and workflow coordination. However, foundation models and large language models introduce failure modes that differ qualitatively from those of conventional tabular prediction models. Hallucination may lead to fluent but unsupported clinical statements, recommendations, or references; prompt sensitivity may cause materially different outputs in response to relatively small changes in wording, context, or information ordering; and sycophancy may cause a model to align inappropriately with a user’s stated belief or preferred conclusion rather than maintaining evidence-based independence. These characteristics are particularly important in clinical settings because linguistic fluency and apparent reasoning coherence may create an impression of reliability even when the underlying output is incorrect. Accordingly, evaluation of foundation-model applications in kidney care should extend beyond conventional predictive performance to include factual accuracy, consistency across prompts, robustness to conflicting or incomplete context, uncertainty communication, source attribution, and resistance to user-induced bias.
Mitigation strategies may include grounding outputs in validated clinical data and authoritative sources, retrieval-augmented generation, structured prompting, explicit citation and source verification, constrained output formats, and human review of clinically consequential recommendations. However, these safeguards do not eliminate risk and should themselves be prospectively evaluated under realistic clinical conditions. Substantial challenges also remain, including explainability,107 computational requirements,109 interoperability,110 governance,111 and the need for rigorous prospective validation before widespread clinical implementation. The conceptual architecture of multimodal AI-bioprognostic systems integrating clinical, molecular, imaging, pathology, and wearable data streams is illustrated in Figure 4.
Figure 4.

Multimodal AI-Bioprognostic Architecture in Precision Kidney Care. Schematic representation of a proposed AI-bioprognostic platform integrating longitudinal clinical data, wearable and remote-monitoring data, imaging, digital pathology, molecular and multi-omics information, and transplant biomarkers through multimodal fusion, temporal modeling, latent phenotype identification, and dynamic recalibration. The rightward arrow indicates the conceptual flow from multimodal data integration and adaptive modeling toward potential clinical outputs, including earlier disease detection, individualized prognostication, precision therapeutics, operational intelligence, and AI-assisted clinical synthesis. These outputs represent proposed applications of the AI-bioprognostic framework and should not be interpreted as established clinical benefits or effect sizes.
Challenges, Bias, and Translational Barriers
Despite rapid advances in AI across nephrology and kidney transplantation, substantial translational barriers continue to limit safe and scalable clinical deployment. Many currently published AI models remain proof-of-concept systems developed under highly controlled retrospective conditions, with limited demonstration of sustained real-world performance across diverse healthcare environments.
Critical Appraisal and Reporting Frameworks for Kidney AI
Interpretation of AI prediction studies in kidney care should be guided by established methodological and reporting frameworks rather than by discrimination metrics alone. TRIPOD+AI provides reporting guidance for studies developing or evaluating prediction models using machine-learning or other AI methods, including transparent description of data sources, predictors, outcomes, model development, validation, performance, and intended use.112,113 For studies involving large language models, TRIPOD-LLM extends TRIPOD+AI to address the specific reporting requirements of LLM development, tuning, prompt engineering, evaluation, and clinical use, with particular emphasis on transparency, human oversight, model and task specification, and task-specific performance assessment.114 PROBAST and PROBAST+AI provide complementary frameworks for assessing risk of bias and applicability, including concerns related to participant selection, predictor and outcome definition, analytic methods, overfitting, validation, and model evaluation.112 For reviews of prediction models, CHARMS provides a structured framework for study selection, data extraction, and appraisal of model-development and validation studies.115,116 Together, these instruments provide a more rigorous basis for determining whether apparently strong predictive performance is likely to be reproducible, transportable, and clinically meaningful.
Framework selection should also reflect the stage and modality of AI evaluation. DECIDE-AI is particularly relevant to early-stage prospective evaluation of AI-based clinical decision-support systems, emphasizing human–AI interaction, workflow integration, safety, and real-world performance.117 For interventional studies, SPIRIT-AI and CONSORT-AI extend established protocol and randomized-trial reporting standards to AI interventions and emphasize issues such as algorithm specification, input data, human interaction, error handling, and implementation context.118,119 For imaging and digital pathology applications, CLAIM provides domain-specific guidance for transparent reporting of AI studies involving medical imaging.120 Broader frameworks such as FUTURE-AI further emphasize trustworthiness across dimensions including fairness, universality, traceability, usability, robustness, and explainability.121
Because the present article is a narrative review, we did not formally score individual studies using PROBAST or PROBAST+AI or conduct a systematic risk-of-bias assessment using these instruments.116 Rather, we use these frameworks to contextualize recurrent methodological limitations in the literature and to define expectations for future research. In practical terms, high-quality kidney AI evidence should demonstrate transparent reporting, appropriate internal and external validation, calibration assessment, evaluation of subgroup performance and fairness, prospective workflow-based testing when clinically deployed, and, where the AI system is intended to alter care, assessment of whether its use improves clinician decisions or patient outcomes. Application of these frameworks may help shift the field from demonstrations of technical performance toward reproducible, clinically useful, and trustworthy AI-enabled kidney care.
External Validation and Dataset Shift
One of the most significant limitations of current AI-bioprognostic systems is inadequate external validation. Many kidney AI models are trained using single-center datasets derived from tertiary academic institutions with relatively homogeneous patient populations, institutional workflows, and laboratory practices. Consequently, models often demonstrate performance degradation when applied to external populations, a phenomenon commonly referred to as dataset shift or transportability failure.66,122
Variations in EHR structure, laboratory calibration, imaging protocols, biopsy interpretation practices, and patient demographics can substantially alter model behavior. For example, AKI prediction algorithms trained in ICU populations may underperform in outpatient nephrology settings or community hospitals. Similarly, transplant rejection models developed in predominantly White recipient cohorts may not generalize to more diverse populations with different immunologic or socioeconomic profiles. The distinction between dataset size and dataset representativeness is also particularly important. For example, the widely cited AKI prediction study by Tomašev et al was developed using a very large longitudinal dataset from the US Department of Veterans Affairs, comprising 703,782 adults across 172 inpatient and 1,062 outpatient sites.50 Despite this scale and multisite design, the cohort was predominantly male, reflecting the demographic characteristics of the Veterans Affairs population. Consequently, strong performance within this setting should not be assumed to translate directly to more sex-balanced populations, pediatric populations, community health systems, or healthcare environments outside the United States. This example illustrates that large sample size alone does not ensure transportability and that external validation should explicitly examine whether model performance remains stable across demographic groups, institutions, and clinical settings that differ materially from the development cohort.
A prominent cautionary example comes from the independent external validation of the proprietary Epic Sepsis Model.123 Wong et al123 evaluated the model in 38,455 hospitalizations at an Academic Medical Center after it had already been implemented across hundreds of US hospitals and found that its real-world performance was substantially less favorable than might have been inferred from prior vendor-reported development results. The evaluation also demonstrated that threshold selection could generate a substantial alert burden, highlighting that discrimination alone does not capture the operational consequences of deployment. Although this model was developed for sepsis rather than kidney disease, the findings are directly relevant to nephrology because increasingly complex predictive systems may be deployed through proprietary EHR platforms or commercial clinical decision-support tools. This example illustrates that performance achieved during model development or vendor validation may not transport reliably across institutions, patient populations, workflows, or clinical environments and reinforces the need for independent external validation before broad clinical adoption.
For kidney AI, external validation should therefore assess not only discrimination but also calibration, threshold-specific sensitivity and positive predictive value, subgroup performance, alert burden, and clinically meaningful utility within the intended deployment environment. These considerations are especially important for proprietary models, in which limited transparency regarding training data, model architecture, feature engineering, and updating procedures may constrain independent scrutiny. Prospective local validation and ongoing post-deployment monitoring are therefore essential to determine whether performance remains stable after implementation and whether model outputs improve rather than disrupt clinical care. Continuous recalibration, multicenter validation, temporal validation, and prospective implementation studies are therefore essential before widespread adoption.36,66
Bias and Fairness
AI systems may inadvertently amplify existing healthcare disparities if trained on incomplete or unrepresentative datasets. Structural inequities embedded within healthcare systems can become encoded into predictive models through differential access to care, delayed referral patterns, missing laboratory data, or inconsistent follow-up.49,124
A landmark example of algorithmic bias outside nephrology was reported by Obermeyer et al,125 who evaluated a widely used commercial population-health algorithm and identified substantial racial bias in the selection of patients for additional care. Although race was not explicitly used as a predictor, the algorithm relied on healthcare expenditures as a proxy for health need. Because healthcare spending reflected pre-existing disparities in access to and utilization of care, Black patients were assigned lower predicted need than White patients with comparable levels of illness, resulting in systematic under-identification for enhanced care management.125 This example illustrates that removal of race from model inputs is not, by itself, sufficient to ensure fairness. Bias may also arise through proxy variables, outcome definitions, training labels, or optimization targets that encode structural inequities. In kidney care, these concerns are particularly relevant when predictive models influence referral, transplantation evaluation, dialysis planning, or allocation of disease-management resources.
A related nephrology-specific example is the adoption of the 2021 CKD-EPI creatinine- and creatinine–cystatin C-based equations that removed the race coefficient from eGFR estimation.126 Because eGFR is a foundational input used directly or indirectly in many models for CKD progression, kidney failure, transplantation, and other kidney outcomes, revision of the underlying estimating equation may alter patient classification, predicted risk, model calibration, and downstream clinical decision-making. This change illustrates how equity considerations can prompt revision of a core clinical variable and how such revisions may propagate through prediction systems that depend on that variable. Accordingly, fairness assessment in kidney AI should extend beyond evaluation of the predictive algorithm itself to include scrutiny of the variables, labels, proxy measures, and clinical assumptions embedded within model inputs and targets. When foundational inputs such as eGFR are revised, existing prediction models may require reassessment, external validation, and, where appropriate, recalibration to ensure that their performance remains reliable and equitable across patient populations.
Racial and socioeconomic bias remain particularly concerning in kidney disease because historically marginalized populations frequently experience higher CKD burden, delayed transplantation access, and disparities in dialysis care. Inadequate representation of rural populations, uninsured patients, pediatric recipients, and global populations may further compromise fairness. Importantly, algorithmic fairness cannot be assumed solely from high overall model accuracy, as subgroup calibration and error distribution may differ substantially across populations. Fairness evaluation should therefore include subgroup-specific calibration and error rates, assessment of potentially biased proxy variables and outcome definitions, transparent reporting of demographic composition, and inclusion of diverse training and validation cohorts. Bias auditing, transparent reporting of demographic composition, fairness metrics, and inclusion of diverse training cohorts are critical safeguards for equitable deployment.127,128
Interpretability and Explainability
The increasing complexity of deep learning and transformer-based architectures has intensified concerns regarding interpretability. Many multimodal AI systems function as “black-box” models in which the internal reasoning process remains difficult to understand even for developers. This opacity may reduce clinician trust, particularly in high-stakes nephrology decisions involving dialysis initiation, transplant rejection treatment, or immunosuppression adjustment.11,129,130
Explainable AI approaches aim to partially address these concerns through feature attribution maps, saliency visualization, Shapley additive explanations, counterfactual modeling, and uncertainty estimation. However, explainability tools themselves may not fully capture biologic causality and can occasionally provide misleading reassurance. Clinically useful interpretability therefore requires alignment between model outputs and established pathophysiologic reasoning rather than solely mathematical transparency.107,131
Operational Challenges
Operational integration remains a major determinant of clinical success. Even highly accurate AI systems may fail if they disrupt workflow efficiency or increase cognitive burden. Excessive notifications and poorly calibrated alerts can contribute to alert fatigue, reducing clinician responsiveness and potentially undermining patient safety.17,64
Interoperability challenges further complicate deployment, particularly across fragmented EHR systems and heterogeneous imaging or laboratory platforms. Regulatory oversight of adaptive and AI-enabled clinical systems is increasingly being operationalized through specific lifecycle-based mechanisms. In the United States, the FDA’s Predetermined Change Control Plan (PCCP) framework allows manufacturers of AI-enabled medical devices to prospectively describe specified modifications, the methods by which those modifications will be developed, validated, and implemented, and an assessment of their anticipated impact.132 When reviewed and authorized as part of a marketing submission, a PCCP can permit implementation of modifications within the authorized scope without requiring a separate marketing submission for each change. Importantly, a PCCP does not provide unrestricted authorization for autonomous or continuously changing algorithms; rather, it establishes predefined boundaries and validation requirements for anticipated modifications, thereby supporting controlled model updating while maintaining lifecycle oversight of safety and effectiveness.132
In the European Union, the EU Artificial Intelligence Act establishes an additional risk-based governance framework.133,134 Under Article 6, AI systems that are themselves products, or safety components of products, covered by specified Union harmonization legislation and subject to third-party conformity assessment are classified as high-risk, a category that encompasses many AI-enabled medical devices. High-risk systems are subject to requirements addressing lifecycle risk management, documentation, human oversight, performance, and post-market governance. These frameworks illustrate that regulation of adaptive clinical AI has progressed beyond general principles toward defined mechanisms for managing model modification and lifecycle risk. Nevertheless, important implementation questions remain regarding model drift, local adaptation, responsibility for post-deployment updates, and coordination between regulatory requirements and health-system governance.
Health-economic considerations are an additional determinant of whether AI-enabled prediction systems can be implemented and sustained in routine kidney care. Deployment costs extend beyond acquisition of the algorithm itself and may include EHR integration, data engineering, computing infrastructure, cybersecurity, model validation and monitoring, clinician training, workflow redesign, technical support, and governance. These requirements may be particularly burdensome for smaller or resource-constrained health systems and may contribute to unequal adoption even when a model demonstrates acceptable predictive performance. Accordingly, implementation feasibility should be evaluated together with technical performance and clinical effectiveness.
Cost-effectiveness should also be demonstrated rather than assumed. AI-enabled prediction may generate economic value if earlier identification of high-risk patients prevents hospitalization, delays kidney disease progression, reduces avoidable dialysis-related complications, improves transplant outcomes, or decreases unnecessary testing and clinical utilization. However, these potential savings must be weighed against implementation and maintenance costs, false-positive–driven downstream testing, clinician time, and opportunity costs. Future evaluations should therefore incorporate formal health-economic outcomes, including budget impact, incremental cost-effectiveness, cost per adverse event prevented, and, where appropriate, cost per quality-adjusted life-year, alongside conventional measures of discrimination, calibration, and clinical utility.
Reimbursement pathways represent a related but distinct challenge. Sustainable adoption may depend on whether AI-enabled prediction is incorporated into an existing reimbursed clinical service, supported through a diagnostic or care-management pathway, financed through institutional or value-based care models, or associated with a specific payment mechanism. Thus, reimbursement and business-model considerations should be evaluated early in implementation planning rather than after technical deployment. In practice, the likelihood that an AI system will be adopted at scale may depend as much on demonstrable clinical and economic value, implementation burden, and alignment with reimbursement structures as on predictive accuracy itself. Additional concerns include cybersecurity vulnerabilities, data governance, medico-legal accountability, reimbursement uncertainty, and institutional infrastructure limitations.11,12,135
Human-in-the-Loop Oversight
Current evidence suggests that AI is most effective when deployed as a cognitive augmentation tool rather than a replacement for nephrologists or transplant clinicians. Human-in-the-loop oversight remains essential to contextualize AI recommendations within individual patient circumstances, competing comorbidities, and evolving clinical trajectories.11,12,17,111,135–139
Physician supervision is particularly important when AI outputs conflict with bedside assessment or when predictions involve uncertainty, rare phenotypes, or ethically sensitive decisions. The most effective future models will likely combine computational scalability with expert clinical judgment, creating collaborative intelligence systems that enhance precision, efficiency, and consistency while preserving human accountability and patient-centered care.11,12,17,111,135–139 Major translational barriers to clinical implementation and potential mitigation strategies are summarized in Table 3.
Table 3.
Major Translational Challenges and Mitigation Strategies for AI Deployment in Kidney Care
| Challenge Domain | Key Barrier | Clinical Consequence | Potential Mitigation Strategy |
|---|---|---|---|
| External Validation9,36 | Single-center overfitting | Reduced generalizability | Multicenter and prospective validation |
| Dataset Shift12,40 | Differences in EHRs, workflows, populations | Performance degradation | Dynamic recalibration and monitoring |
| Bias and Fairness49,124 | Underrepresentation of vulnerable groups | Health disparities amplification | Diverse datasets and fairness auditing |
| Interpretability129 | Black-box decision making | Reduced clinician trust | Explainable AI and uncertainty reporting |
| Workflow Integration110 | Poor EHR integration | Alert fatigue and low adoption | Human-centered workflow design |
| Interoperability12 | Heterogeneous data systems | Fragmented deployment | Standardized data architectures and FHIR integration |
| Regulatory Oversight137 | Adaptive model evolution | Safety and liability uncertainty | Continuous post-deployment surveillance |
| Human Oversight11,17,136 | Overreliance on automation | Diagnostic or therapeutic error | Human-in-the-loop supervision |
| Data Governance111,135 | Privacy and cybersecurity risks | Loss of trust and legal exposure | Federated learning and secure infrastructures |
| Resource Constraints138,139 | Limited technical infrastructure | Unequal AI access | Scalable cloud-based implementation strategies |
Abbreviations: AI, artificial intelligence; EHR, electronic health record; FHIR, Fast Healthcare Interoperability Resources.
Future Directions
The future role of AI in nephrology should be considered conditional rather than inevitable. High-performing models may ultimately support earlier recognition, improved prognostication, or more efficient clinical workflows, but these benefits must be demonstrated rather than assumed. Evidence of strong discrimination should be followed by rigorous external validation, prospective evaluation in the intended clinical environment, assessment of calibration and subgroup performance, and, where the model is intended to alter care, controlled clinical-impact studies demonstrating that deployment leads to meaningful improvement in decision-making, safety, efficiency, or patient outcomes. Within this evidence-based framework, future kidney AI systems may evolve beyond static prediction models toward more adaptive, multimodal, and workflow-integrated approaches. Such systems may combine biologic, physiologic, imaging, wearable, and longitudinal clinical data into continuously updated platforms, but their clinical value will depend on whether these capabilities translate into reliable, equitable, and actionable improvements in care.
Continuously Learning Kidney Systems
Future kidney AI systems may incorporate adaptive architectures that are periodically recalibrated or retrained as new clinical, laboratory, imaging, and outcome data become available, rather than relying exclusively on fixed models. This capability may improve resilience against dataset shift and changing practice patterns while enabling more accurate longitudinal risk estimation.15,140 However, adaptation should not be interpreted as unconstrained continuous learning. Changes to a clinically deployed model may alter discrimination, calibration, subgroup performance, and the overall benefit-risk profile and therefore require predefined validation, monitoring, and governance processes.
Future nephrology platforms may incorporate real-time recalibration pipelines embedded within EHRs, allowing dynamic adjustment of predictions for AKI progression, dialysis instability, transplant rejection, and CKD progression. Regulatory frameworks are beginning to provide mechanisms for managing this evolution. In the United States, FDA-authorized PCCPs132 can prospectively define specified model modifications and the procedures used to develop, validate, and implement them, thereby permitting controlled updates within an authorized scope. In the European Union, AI-enabled medical devices meeting the criteria for high-risk classification under the EU AI Act are subject to lifecycle-oriented risk-management and governance requirements. Future adaptive kidney AI systems will therefore require not only technical mechanisms for recalibration or retraining but also traceable version control, prespecified change boundaries, ongoing performance monitoring, subgroup and calibration surveillance, and coordinated regulatory and institutional oversight throughout the model lifecycle. Importantly, the ability to update a model should not itself be considered evidence of improved clinical utility; each material modification should preserve or improve performance, safety, and equity in the intended population.
Federated Learning
Federated learning represents a promising strategy for multicenter kidney AI development while preserving patient privacy and institutional data sovereignty. Rather than transferring sensitive patient-level data to a centralized repository, decentralized training allows models to learn across multiple institutions while data remain locally stored.141,142
This approach may be particularly valuable in nephrology, where rare kidney diseases, transplant populations, and underrepresented demographic groups often require collaboration across geographically distributed centers. Federated infrastructures could improve model generalizability while reducing legal, ethical, and cybersecurity concerns associated with large-scale centralized data sharing. Integration with standardized data architectures and interoperable frameworks may further accelerate international kidney AI collaboration. However, federated learning does not automatically eliminate bias, heterogeneity, or site-specific performance differences; harmonized data definitions, transparent aggregation methods, and external evaluation across participating and nonparticipating centers remain necessary.
Agentic AI and Clinical Decision Support
The evolution of agentic AI introduces the possibility of semi-autonomous clinical support systems capable of persistent monitoring, contextual synthesis, and longitudinal reasoning. Rather than functioning solely as passive prediction tools, future nephrology AI agents may continuously analyze laboratory trends, medications, hemodynamics, pathology findings, and wearable telemetry to identify evolving kidney risk states.16,140
Such systems could support automated surveillance for nephrotoxin exposure, transplant rejection signals, dialysis access complications, or early fluid overload while generating prioritized recommendations for clinician review. AI-assisted longitudinal synthesis may also reduce cognitive fragmentation by integrating multimodal patient data into concise, clinically interpretable summaries. These potential advantages must be weighed against additional risks introduced by agentic systems, including hallucination, prompt sensitivity, sycophancy, automation bias, propagation of erroneous intermediate inferences, and unintended downstream actions. Clinical deployment will therefore require restricted tool permissions, prespecified action boundaries, source-grounded outputs, uncertainty disclosure, human confirmation before consequential actions, audit trails, version control, and prospective evaluation of clinician over-reliance and downstream error.
Digital Twins in Nephrology
Digital twin technologies represent an emerging frontier in precision nephrology. These systems aim to create individualized computational representations of patients by integrating biologic, physiologic, imaging, genomic, and treatment-response data into dynamic simulation models.15,109
In kidney care, digital twins could eventually enable personalized prediction of disease progression, dialysis responses, immunosuppression optimization, and fluid management strategies before clinical interventions are implemented. Although current applications remain largely investigational, digital twins may ultimately support individualized therapeutic simulation and scenario testing within complex kidney disease phenotypes. At present, however, evidence supporting clinically actionable kidney digital twins remains limited, and substantial methodological, validation, interoperability, and governance challenges must be addressed before simulated treatment responses can be relied upon in patient care.
Precision Predictive Kidney Care
Collectively, these advances support a broader transition from reactive nephrology toward precision predictive kidney care. Future kidney intelligence ecosystems may prioritize earlier disease detection, individualized risk forecasting, and proactive intervention before irreversible injury occurs.6,17,143 However, such a transition should not be assumed to follow automatically from technical advances. The relevant benchmark is whether predictive systems demonstrate reproducible performance across settings, remain calibrated and equitable over time, integrate effectively into clinical workflows, and improve clinically meaningful outcomes when acted upon. Rather than focusing predominantly on late-stage complications, nephrology may increasingly evolve toward preventive, continuously monitored, and precision-informed care pathways supported by adaptive multimodal AI systems (Figure 5).
Figure 5.

Future AI Ecosystem for Predictive Kidney Care. Proposed future-state architecture for precision nephrology in which multimodal clinical, physiologic, imaging, pathology, molecular, wearable, home-monitoring, dialysis-telemetry, and transplant biomarker data are integrated within a kidney intelligence platform to support early disease detection, individualized prognostication, precision therapeutics, automated surveillance, and AI-assisted clinical synthesis under human oversight. Directional arrows indicate the conceptual flow of data, model updating, clinical information, and care progression rather than observed causal effects. Upward and downward arrows in the “ultimate goal” panel indicate the intended direction of desirable outcomes, such as improved quality of life and health equity and reduced kidney failure, cardiovascular events, hospitalizations, and cost of care; they do not represent measured effect sizes or established clinical benefits. The lower continuum depicts the proposed evolution from reactive care toward predictive, precision, and preventive care, while the human-in-the-loop pathway emphasizes continued physician supervision, explainability and transparency, fairness auditing, safety monitoring, and governance throughout the system.
Conclusion
Although AI-based models have frequently demonstrated improved discrimination, earlier risk identification, and more granular prognostic stratification across kidney care, these technical gains have not yet been shown consistently to translate into improved patient outcomes. Indeed, the randomized clinical-impact studies discussed in this review, including electronic AKI alerting, machine-learning-triggered early nephrology consultation, and EHR-integrated transplant risk prediction, did not demonstrate improvement in the relevant clinical outcomes despite earlier or more accurate identification of risk. These findings underscore that predictive performance should not be interpreted as evidence of clinical benefit.
The next phase of kidney AI should prioritize evidence quality and clinical utility rather than incremental gains in model complexity alone. First, studies of AI-enabled prediction in nephrology should meet a minimum reporting standard consistent with TRIPOD+AI and related frameworks, including transparent specification of the intended use, target population, data sources, predictor and outcome definitions, missing-data handling, model-development procedures, model version, discrimination, calibration, external validation, subgroup performance, and clinically relevant decision thresholds. For large language model and foundation-model applications, additional reporting should include model version, prompting or tuning strategy, input context, human oversight, and task-specific evaluation.
Second, a nephrology model-validation registry could provide a standardized mechanism for tracking the lifecycle of clinically relevant AI systems. At minimum, such a registry should capture the model name and version, intended clinical use, development population, external validation populations, prediction target and horizon, discrimination and calibration, subgroup-specific performance, temporal validation, evidence of dataset shift or calibration drift, prospective clinical-impact evaluation, commercial or proprietary status, investigator independence of validation, and subsequent model updates or recalibration. Such infrastructure could help distinguish models supported only by retrospective performance from those with evidence of transportability, stability, and clinical utility.
Third, movement from prediction to demonstrated benefit will require prospective clinical-impact evaluation. Silent prospective validation may establish whether performance persists in the deployment environment, but models intended to alter care should subsequently be evaluated using pragmatic randomized, cluster-randomized, stepped-wedge, or other appropriately controlled implementation designs when feasible. These studies should assess not only discrimination and calibration but also whether AI changes clinician behavior, treatment decisions, workflow efficiency, alert burden, safety, patient-centered outcomes, healthcare utilization, and unintended consequences. The relevant benchmark is therefore not simply whether an adverse kidney outcome can be predicted earlier, but whether acting on that prediction improves care.
Fourth, fairness auditing should become a minimum requirement for kidney AI rather than an optional secondary analysis. At a minimum, studies should report subgroup-specific discrimination, calibration, false-positive and false-negative rates, and relevant uncertainty across clinically important demographic and socioeconomic groups; examine whether proxy variables, outcome labels, missingness patterns, or clinical assumptions encode structural inequities; and reassess model performance when foundational inputs, such as eGFR equations, are revised. Models intended for broad deployment should additionally undergo validation in populations that differ meaningfully from the development cohort with respect to sex, race and ethnicity, age, geography, socioeconomic context, and healthcare setting.
Taken together, these priorities provide a practical pathway for moving kidney AI from technically impressive prediction toward clinically useful and trustworthy implementation. Accordingly, progress in kidney AI should be judged not by the sophistication of the underlying algorithm or by predictive accuracy alone, but by transparent reporting, reproducible and independent validation, equitable performance, clinically appropriate workflow integration, and prospective demonstration of benefit. The appropriate objective is not greater use of AI per se, but selective implementation of AI systems that provide demonstrable value beyond existing clinical care.
Funding Statement
No specific funding was received for the preparation of this manuscript.
Data Sharing Statement
No new datasets were generated or analyzed during the preparation of this narrative review.
Author Contributions
All authors made a significant contribution to the work reported, whether that is in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; have agreed on the journal to which the article has been submitted; and agree to be accountable for all aspects of the work.
Disclosure
The authors declare no conflicts of interest and report no financial, advisory, consulting, equity, licensing, or other relevant commercial relationships with the developers or vendors of the proprietary AI systems discussed in this review
References
- 1.Mark PB, Stafford LK, Grams ME. Global, regional, and national burden of chronic kidney disease in adults, 1990-2023, and its attributable risk factors: a systematic analysis for the Global Burden of Disease Study 2023. Lancet. 2025;406(10518):2461–25. doi: 10.1016/S0140-6736(25)01853-7 [DOI] [PubMed] [Google Scholar]
- 2.KDIGO. Clinical practice guideline for the evaluation and management of chronic kidney disease. Kidney Int. 2024;105(4s):S117–s314. doi: 10.1016/j.kint.2023.10.018 [DOI] [PubMed] [Google Scholar]
- 3.Ostermann M, Lumlertgul N, Jeong R, See E, Joannidis M, James M. Acute kidney injury. Lancet. 2025;405(10474):241–256. doi: 10.1016/S0140-6736(24)02385-7 [DOI] [PubMed] [Google Scholar]
- 4.Chalikias G, Tziakas D. Cardiovascular consequences of acute kidney injury. N Engl J Med. 2020;383(11):1094. doi: 10.1056/NEJMc2023901 [DOI] [PubMed] [Google Scholar]
- 5.Rafferty Q, Stafford LK, Vos T. Global, regional, and national prevalence of kidney failure with replacement therapy and associated aetiologies, 1990-2023: a systematic analysis for the Global Burden of Disease Study 2023. Lancet Glob Health. 2025;13(8):e1378–e95. doi: 10.1016/S2214-109X(25)00198-6 [DOI] [PubMed] [Google Scholar]
- 6.Ortiz A, Yanagita M, Yokoi H, Torra R. Evolving strategies for early diagnosis, proactive prevention and treatment of CKD. Nephrol Dial Transplant. 2026;41(3):418–427. doi: 10.1093/ndt/gfaf151 [DOI] [PubMed] [Google Scholar]
- 7.Grams ME, Brunskill NJ, Ballew SH, et al. The kidney failure risk equation: evaluation of novel input variables including eGFR estimated using the CKD-EPI 2021 equation in 59 cohorts. J Am Soc Nephrol. 2023;34(3):482–494. doi: 10.1681/ASN.0000000000000050 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Levey AS, Grams ME, Inker LA. Uses of GFR and albuminuria level in acute and chronic kidney disease. N Engl J Med. 2022;386(22):2120–2128. doi: 10.1056/NEJMra2201153 [DOI] [PubMed] [Google Scholar]
- 9.Tangri N, Ferguson TW, Teng CC, et al. Validation of the Klinrisk machine learning model for CKD progression in a large representative US population. J Am Soc Nephrol. 2026;37(2):326–337. doi: 10.1681/ASN.0000000817 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Tangri N, Inker LA, Hiebert B, et al. A dynamic predictive model for progression of CKD. Am J Kidney Dis. 2017;69(4):514–520. doi: 10.1053/j.ajkd.2016.07.030 [DOI] [PubMed] [Google Scholar]
- 11.Tangri N, Cheungpasitporn W, Crittenden SD, et al. Responsible use of artificial intelligence to improve kidney care: a statement from the American Society of Nephrology. J Am Soc Nephrol. 2026;37(4):881–890. doi: 10.1681/ASN.0000000929 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Kaur N, Bhattacharya S, Butte AJ. Big Data in Nephrology. Nat Rev Nephrol. 2021;17(10):676–687. doi: 10.1038/s41581-021-00439-x [DOI] [PubMed] [Google Scholar]
- 13.Madhvapathy SR, Cho S, Gessaroli E, et al. Implantable bioelectronics and wearable sensors for kidney health and disease. Nat Rev Nephrol. 2025;21(7):443–463. doi: 10.1038/s41581-025-00961-2 [DOI] [PubMed] [Google Scholar]
- 14.Barisoni L, Lafata KJ, Hewitt SM, Madabhushi A, Balis UGJ. Digital pathology and computational image analysis in nephropathology. Nat Rev Nephrol. 2020;16(11):669–685. doi: 10.1038/s41581-020-0321-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Cheungpasitporn W, Athavale A, Ghazi L, et al. Transforming nephrology through artificial intelligence: a state-of-the-art roadmap for clinical integration. Clin Kidney J. 2026;19(2):sfag004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Collaco BG, Haider SA, Prabha S, et al. The role of agentic artificial intelligence in healthcare: a scoping review. NPJ Digit Med. 2026;9(1). doi: 10.1038/s41746-026-02517-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Loftus TJ, Shickel B, Ozrazgat-Baslanti T, et al. Artificial intelligence-enabled decision support in nephrology. Nat Rev Nephrol. 2022;18(7):452–465. doi: 10.1038/s41581-022-00562-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Hunter DJ, Holmes C. Where medical statistics meets artificial intelligence. N Engl J Med. 2023;389(13):1211–1219. doi: 10.1056/NEJMra2212850 [DOI] [PubMed] [Google Scholar]
- 19.Gulamali FF, Sawant AS, Nadkarni GN. Machine learning for risk stratification in kidney disease. Curr Opin Nephrol Hypertens. 2022;31(6):548–552. doi: 10.1097/MNH.0000000000000832 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wang Y, Zhao Y, Therneau TM, et al. Unsupervised machine learning for the discovery of latent disease clusters and patient subgroups using electronic health records. J Biomed Inform. 2020;102:103364. doi: 10.1016/j.jbi.2019.103364 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Zheng Z, Waikar SS, Schmidt IM, et al. Subtyping CKD patients by consensus clustering: the Chronic Renal Insufficiency Cohort (CRIC) study. J Am Soc Nephrol. 2021;32(3):639–653. doi: 10.1681/ASN.2020030239 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Carone M, Rotnitzky A. Reinforcement learning for finding optimal dynamic treatment regimes using observational data. Jama. 2026;335(3):267–268. doi: 10.1001/jama.2025.20541 [DOI] [PubMed] [Google Scholar]
- 23.Adiyeke E, Liu T, Dheeraj Naganaboin VS, et al. Reinforcement learning for intraoperative hypotension management with consideration to postoperative acute kidney injury. Kidney360. 2026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Becker JU, Mayerich D, Padmanabhan M, et al. Artificial intelligence and machine learning in nephropathology. Kidney Int. 2020;98(1):65–75. doi: 10.1016/j.kint.2020.02.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Hermsen M, de Bel T, den Boer M, et al. Deep Learning–based histopathologic assessment of kidney tissue. J Am Soc Nephrol. 2019;30(10):1968–1979. doi: 10.1681/ASN.2019020144 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Yang Z, Mitra A, Liu W, Berlowitz D, Yu H. TransformEHR: transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nat Commun. 2023;14(1):7857. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Nerella S, Bandyopadhyay S, Zhang J, et al. Transformers and large language models in healthcare: a review. Artif Intell Med. 2024;154:102900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Eddy S, Mariani LH, Kretzler M. Integrated multi-omics approaches to improve classification of chronic kidney disease. Nat Rev Nephrol. 2020;16(11):657–668. doi: 10.1038/s41581-020-0286-5 [DOI] [PubMed] [Google Scholar]
- 29.Lopes MB, Coletti R, Duranton F, et al. The omics-driven machine learning path to cost-effective precision medicine in chronic kidney disease. Proteomics. 2025;25(11–12):e202400108. doi: 10.1002/pmic.202400108 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Rroji M, Spasovski G. Omics studies in CKD: diagnostic opportunities and therapeutic potential. Proteomics. 2025;25(11–12):e202400151. doi: 10.1002/pmic.202400151 [DOI] [PubMed] [Google Scholar]
- 31.Dong Z, Wang X, Pan S, et al. A multimodal transformer system for noninvasive diabetic nephropathy diagnosis via retinal imaging. NPJ Digit Med. 2025;8(1):50. doi: 10.1038/s41746-024-01393-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Mistry NS, Koyner JL. Artificial intelligence in acute kidney injury: from static to dynamic models. Adv Chronic Kidney Dis. 2021;28(1):74–82. doi: 10.1053/j.ackd.2021.03.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Alba AC, Agoritsas T, Walsh M, et al. Discrimination and calibration of clinical prediction models: users’ guides to the medical literature. Jama. 2017;318(14):1377–1384. doi: 10.1001/jama.2017.12126 [DOI] [PubMed] [Google Scholar]
- 34.Khan SS, Greenland P, Hayman LL, et al. Criteria to assess the predictive and clinical utility of novel models, biomarkers, and tools for risk of cardiovascular disease: a scientific statement from the American Heart Association. Circulation. 2026;153(11):e953–e70. doi: 10.1161/CIR.0000000000001401 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi: 10.1371/journal.pone.0118432 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Ramspek CL, Jager KJ, Dekker FW, Zoccali C, van Diepen M. External validation of prognostic models: what, why, how, when and where? Clin Kidney J. 2021;14(1):49–58. doi: 10.1093/ckj/sfaa188 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Tangri N, Grams ME, Levey AS, et al. Multinational assessment of accuracy of equations for predicting risk of kidney failure: a meta-analysis. Jama. 2016;315(2):164–174. doi: 10.1001/jama.2015.18202 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Davis SE, Greevy RA Jr, Lasko TA, Walsh CG, Matheny ME. Detection of calibration drift in clinical prediction models to inform model updating. J Biomed Inform. 2020;112:103611. doi: 10.1016/j.jbi.2020.103611 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Subasri V, Krishnan A, Kore A, et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Network Open. 2025;8(6):e2513685. doi: 10.1001/jamanetworkopen.2025.13685 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Guo LL, Pfohl SR, Fries J, et al. Evaluation of domain generalization and adaptation on improving model robustness to temporal dataset shift in clinical medicine. Sci Rep. 2022;12(1):2726. doi: 10.1038/s41598-022-06484-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wan T, Chen Q, Gao Y, Luo R, Li N, Feng Y. Development and validation of early-stage and progression prediction models for chronic kidney disease: a retrospective study. PeerJ. 2026;14:e20931. doi: 10.7717/peerj.20931 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zacharias HU, Altenbuchinger M, Schultheiss UT, et al. A predictive model for progression of CKD to kidney failure based on routine laboratory tests. Am J Kidney Dis. 2022;79(2):217–30e1. doi: 10.1053/j.ajkd.2021.05.018 [DOI] [PubMed] [Google Scholar]
- 43.Chiofolo C, Chbat N, Ghosh E, Eshelman L, Kashani K. Automated continuous acute kidney injury prediction and surveillance: a random forest model. Mayo Clin Proc. 2019;94(5):783–792. doi: 10.1016/j.mayocp.2019.02.009 [DOI] [PubMed] [Google Scholar]
- 44.Li J, Du X, Zhang R, et al. Machine learning models for predicting short-term progression in patients with stage 4 chronic kidney disease: a multi-center validation study. Sci Rep. 2025;15(1):39285. doi: 10.1038/s41598-025-23037-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Bhatraju PK, Zelnick LR, Herting J, et al. Identification of acute kidney injury subphenotypes with differing molecular signatures and responses to vasopressin therapy. Am J Respir Crit Care Med. 2019;199(7):863–872. doi: 10.1164/rccm.201807-1346OC [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Yang Z, Tian Y, Zhou T, et al. Optimization of dry weight assessment in hemodialysis patients via reinforcement learning. IEEE J Biomed Health Inform. 2022;26(10):4880–4891. doi: 10.1109/JBHI.2022.3192021 [DOI] [PubMed] [Google Scholar]
- 47.Jacq A, Tarris G, Jaugey A, et al. Automated evaluation with deep learning of total interstitial inflammation and peritubular capillaritis on kidney biopsies. Nephrol Dial Transplant. 2023;38(12):2786–2798. [DOI] [PubMed] [Google Scholar]
- 48.Ratchatorn A, Ketdao N, Sonsilphong S, Triamwichanon D, Panitchote A. Deep learning approaches for time series prediction of renal recovery in medical critically Ill patients with acute kidney injury: LSTM, GRU, and transformer models. Crit Care. 2026;30(1). doi: 10.1186/s13054-026-05942-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Wang Z, Fang Y, Li Q. Foundation models in healthcare: a comprehensive review from technical advances to clinical translation. J Transl Med. 2026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Tomasev N, Glorot X, Rae JW, et al. A clinically applicable approach to continuous prediction of future acute kidney injury. Nature. 2019;572(7767):116–119. doi: 10.1038/s41586-019-1390-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Koyner JL, Carey KA, Edelson DP, Churpek MM. The development of a machine learning inpatient acute kidney injury prediction model. Crit Care Med. 2018;46(7):1070–1077. doi: 10.1097/CCM.0000000000003123 [DOI] [PubMed] [Google Scholar]
- 52.Churpek MM, Carey KA, Edelson DP, et al. Internal and external validation of a machine learning risk score for acute kidney injury. JAMA Network Open. 2020;3(8):e2012892. doi: 10.1001/jamanetworkopen.2020.12892 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Yu X, Wang W, Wu R, Gong X, Ji Y, Feng Z. Construction of a machine learning-based interpretable prediction model for acute kidney injury in hospitalized patients. Sci Rep. 2025;15(1):9313. doi: 10.1038/s41598-025-90459-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Ruinelli L, Cippa P, Sieber C, Di Serio C, Ferrari P, Bellasi A. Usability of machine learning algorithms based on electronic health records for the prediction of acute kidney injury and transition to acute kidney disease: a proof of concept study. PLoS One. 2025;20(7):e0326124. doi: 10.1371/journal.pone.0326124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Kotwal S, Herath S, Erlich J, et al. Electronic alerts and a care bundle for acute kidney injury-an Australian cohort study. Nephrol Dial Transplant. 2023;38(3):610–617. doi: 10.1093/ndt/gfac155 [DOI] [PubMed] [Google Scholar]
- 56.Chen JJ, Lee TH, Chan MJ, et al. Electronic alert systems for patients with acute kidney injury: a systematic review and meta-analysis. JAMA Network Open. 2024;7(8):e2430401. doi: 10.1001/jamanetworkopen.2024.30401 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Neyra JA, Ortiz-Soriano V, Liu LJ, et al. Prediction of mortality and major adverse kidney events in critically Ill patients with acute kidney injury. Am J Kidney Dis. 2023;81(1):36–47. doi: 10.1053/j.ajkd.2022.06.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Wu C, Zhang Y, Nie S, et al. Predicting in-hospital outcomes of patients with acute kidney injury. Nat Commun. 2023;14(1):3739. doi: 10.1038/s41467-023-39474-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Jiang X, Hu Y, Guo S, Du C, Cheng X. Prediction of persistent acute kidney injury in postoperative intensive care unit patients using integrated machine learning: a retrospective cohort study. Sci Rep. 2022;12(1):17134. doi: 10.1038/s41598-022-21428-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Takkavatakarn K, Oh W, Chan L, et al. Machine learning derived serum creatinine trajectories in acute kidney injury in critically ill patients with sepsis. Crit Care. 2024;28(1):156. doi: 10.1186/s13054-024-04935-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Luo XQ, Yan P, Zhang NY, et al. Machine learning for early discrimination between transient and persistent acute kidney injury in critically ill patients with sepsis. Sci Rep. 2021;11(1):20269. doi: 10.1038/s41598-021-99840-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Yue S, Li S, Huang X, et al. Machine learning for the prediction of acute kidney injury in patients with sepsis. J Transl Med. 2022;20(1):215. doi: 10.1186/s12967-022-03364-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Liu K, Zhang X, Chen W, et al. Development and validation of a personalized model with transfer learning for acute kidney injury risk estimation using electronic health records. JAMA Network Open. 2022;5(7):e2219776. doi: 10.1001/jamanetworkopen.2022.19776 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Wilson FP, Shashaty M, Testani J, et al. Automated, electronic alerts for acute kidney injury: a single-blind, parallel-group, randomised controlled trial. Lancet. 2015;385(9981):1966–1974. doi: 10.1016/S0140-6736(15)60266-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Song X, Asl Y, Kellum JA, et al. Cross-site transportability of an explainable artificial intelligence model for acute kidney injury prediction. Nat Commun. 2020;11(1):5668. doi: 10.1038/s41467-020-19551-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Rehman AU, Neyra JA, Chen J, Ghazi L. Machine learning models for acute kidney injury prediction and management: a scoping review of externally validated studies. Crit Rev Clin Lab Sci. 2025;62(6):454–476. doi: 10.1080/10408363.2025.2497843 [DOI] [PubMed] [Google Scholar]
- 67.Vagliano I, Chesnaye NC, Leopold JH, Jager KJ, Abu-Hanna A, Schut MC. Machine learning models for predicting acute kidney injury: a systematic review and critical appraisal. Clin Kidney J. 2022;15(12):2266–2280. doi: 10.1093/ckj/sfac181 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Zhu Y, Bi D, Saunders M, Ji Y. Prediction of chronic kidney disease progression using recurrent neural network and electronic health records. Sci Rep. 2023;13(1):22091. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Chung MC, Yu TM, Wang MS, et al. Machine-learning prediction of 30-day infection-related hospitalization in advanced CKD. BMC Infect Dis. 2026;26. doi: 10.1186/s12879-026-13308-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Ma J, Wang J, Ying J, et al. Long-term efficacy of an AI-based health coaching mobile app in slowing the progression of nondialysis-dependent chronic kidney disease: retrospective cohort study. J Med Internet Res. 2024;26:e54206. doi: 10.2196/54206 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Zhu H, Qiao S, Zhao D, et al. Machine learning model for cardiovascular disease prediction in patients with chronic kidney disease. Front Endocrinol. 2024;15:1390729. doi: 10.3389/fendo.2024.1390729 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Lam D, Nadkarni GN, Mosoyan G, et al. Clinical utility of kidney IntelX in early stages of diabetic kidney disease in the CANVAS trial. Am J Nephrol. 2022;53(1):21–31. doi: 10.1159/000519920 [DOI] [PubMed] [Google Scholar]
- 73.Zhang H, Wang LC, Chaudhuri S, et al. Real-time prediction of intradialytic hypotension using machine learning and cloud computing infrastructure. Nephrol Dial Transplant. 2023;38(7):1761–1769. doi: 10.1093/ndt/gfad070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Chiu IM, Wu PJ, Zhang H, et al. Serum potassium monitoring using ai-enabled smartwatch electrocardiograms. JACC Clin Electrophysiol. 2024;10(12):2644–2654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Huang LT, Zheng XY, Zhang ZH, et al. Exploring factors and models to predict post-dialysis volume overload status in maintenance hemodialysis patients based on pre-dialysis parameters. Clin Nephrol. 2026;105(3):151–159. doi: 10.5414/CN111762 [DOI] [PubMed] [Google Scholar]
- 76.Karpinski S, Sibbel S, Gray K, et al. Predicting hospitalizations for patients with chronic kidney disease. Am J Manag Care. 2023;29(9):e262–e6. [DOI] [PubMed] [Google Scholar]
- 77.Peng Z, Zhong S, Li X, et al. An artificial intelligence model to predict mortality among hemodialysis patients: a retrospective validated cohort study. Sci Rep. 2025;15(1):27699. doi: 10.1038/s41598-025-06576-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Mambelli E, Grandi F, Santoro A. Comparison of blood volume biofeedback hemodialysis and conventional hemodialysis on cardiovascular stability and blood pressure control in hemodialysis patients: a systematic review and meta-analysis of randomized controlled trials. J Nephrol. 2024;37(4):897–909. doi: 10.1007/s40620-023-01844-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Randhay A, Eldehni MT, Selby NM. Feedback control in hemodialysis. Semin Dial. 2025;38(1):62–70. doi: 10.1111/sdi.13185 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Bansal N, Artinian NT, Bakris G, et al. Hypertension in patients treated with in-center maintenance hemodialysis: current evidence and future opportunities: a scientific statement from the American Heart Association. Hypertension. 2023;80(6):e112–e22. [DOI] [PubMed] [Google Scholar]
- 81.Lindeboom L, Lee S, Wieringa F, et al. On the potential of wearable bioimpedance for longitudinal fluid monitoring in end-stage kidney disease. Nephrol Dial Transplant. 2022;37(11):2048–2054. [DOI] [PubMed] [Google Scholar]
- 82.Ali H, Mohamed MM, Fulop T, Hamer R. Outcomes of remote patient monitoring in peritoneal dialysis: a meta-analysis and review of practical implications for COVID-19 epidemics. Asaio j. 2023;69(4):e142–e8. doi: 10.1097/MAT.0000000000001891 [DOI] [PubMed] [Google Scholar]
- 83.Arya S, Zhang S, Fu X. Remote patient monitoring in home dialysis patients: enhancing care for home hemodialysis and peritoneal dialysis. Semin Dial. 2026;39(4):157–167. doi: 10.1111/sdi.70041 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Lew SQ, Manani SM, Ronco C, Rosner MH, Sloand JA. Effect of Remote and Virtual Technology on Home Dialysis. Clin J Am Soc Nephrol. 2024;19(10):1330–1337. doi: 10.2215/CJN.0000000000000405 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Pellegrino G, Strippoli GFM. Artificial intelligence in pediatric nephrology: current applications and emerging frameworks for evidence generation. Pediatr Nephrol. 2026. doi: 10.1007/s00467-026-07223-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Chen X, Burgun A, Boyer O, Knebelmann B, Garcelon N. Artificial intelligence and perspective for rare genetic kidney diseases. Kidney Int. 2025;108(3):388–393. [DOI] [PubMed] [Google Scholar]
- 87.Zhou XJ, Zhong XH, Duan LX. Integration of artificial intelligence and multi-omics in kidney diseases. Fundam Res. 2023;3(1):126–148. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Loupy A, Aubert O, Orandi BJ, et al. Prediction system for risk of allograft loss in patients receiving kidney transplants: international derivation and validation study. BMJ. 2019;366:l4923. doi: 10.1136/bmj.l4923 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Chen T, Chen T, Xu W, et al. Development and external validation of a multidimensional deep learning model to dynamically predict kidney outcomes in IgA nephropathy. Clin J Am Soc Nephrol. 2024;19(7):898–907. doi: 10.2215/CJN.0000000000000471 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Cheng C, Li B, Li J, et al. Multi-stain deep learning prediction model of treatment response in lupus nephritis based on renal histopathology. Kidney Int. 2025;107(4):714–727. doi: 10.1016/j.kint.2024.12.007 [DOI] [PubMed] [Google Scholar]
- 91.Chen K, Zheng Y, Huang S, Zhang L. From albuminuria to multi-omics signatures: emerging biomarkers and drug targets for early-stage chronic kidney disease. Front Pharmacol. 2026;17:1765974. doi: 10.3389/fphar.2026.1765974 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Qiao Y, Zhou H, Liu Y, et al. A multi-modal fusion model with enhanced feature representation for chronic kidney disease progression prediction. Brief Bioinform. 2024;26(1). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Chen Z, Ying MTC, Wang Y, et al. Ultrasound-based radiomics analysis in the assessment of renal fibrosis in patients with chronic kidney disease. Abdom Radiol. 2023;48(8):2649–2657. doi: 10.1007/s00261-023-03965-3 [DOI] [PubMed] [Google Scholar]
- 94.Gregory AV, Khalifa M, Im J, et al. Deep learning-based instance-level segmentation of kidney and liver cysts in computed tomography images of patients affected by polycystic kidney disease. Kidney360. 2026;7(1):117–130. doi: 10.34067/KID.0000000924 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Hua C, Qiu L, Zhou L, et al. Value of multiparametric magnetic resonance imaging for evaluating chronic kidney disease and renal fibrosis. Eur Radiol. 2023;33(8):5211–5221. doi: 10.1007/s00330-023-09674-1 [DOI] [PubMed] [Google Scholar]
- 96.Obrisca B, Butiu M, Sibulesky L, et al. Combining donor-derived cell-free DNA and donor specific antibody testing as non-invasive biomarkers for rejection in kidney transplantation. Sci Rep. 2022;12(1):15061. doi: 10.1038/s41598-022-19017-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Gupta G, Athreya A, Kataria A. Biomarkers in kidney transplantation: a rapidly evolving landscape. Transplantation. 2025;109(3):418–427. doi: 10.1097/TP.0000000000005122 [DOI] [PubMed] [Google Scholar]
- 98.Tirasattayapitak S, Ratanatharathorn C, Thotsiri S, et al. Integrating clinical and histopathological data to predict delayed graft function in kidney transplant recipients using machine learning techniques. J Clin Med. 2024;13(24):7502. doi: 10.3390/jcm13247502 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Wu J, Gan L, Shen X, et al. Multiple omics-based machine learning reveals peripheral blood immune cell landscape during acute rejection of kidney transplantation and constructs a precise non-invasive diagnostic strategy. Mamm Genome. 2025;36(4):1192–1214. doi: 10.1007/s00335-025-10149-5 [DOI] [PubMed] [Google Scholar]
- 100.Kherabi Y, Messika J, Peiffer-Smadja N. Machine learning, antimicrobial stewardship, and solid organ transplantation: is this the future? Transpl Infect Dis. 2022;24(5):e13957. doi: 10.1111/tid.13957 [DOI] [PubMed] [Google Scholar]
- 101.Maung Myint T, Chong CH, von Huben A, et al. Serum and urine nucleic acid screening tests for BK polyomavirus-associated nephropathy in kidney and kidney-pancreas transplant recipients. Cochrane Database Syst Rev. 2024;11(11):CD014839. doi: 10.1002/14651858.CD014839.pub2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Arjmandmazidi S, Heidari HR, Ghasemnejad T, et al. An In-depth overview of artificial intelligence (AI) tool utilization across diverse phases of organ transplantation. J Transl Med. 2025;23(1):678. doi: 10.1186/s12967-025-06488-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Labriffe M, Woillard JB, Gwinner W, et al. Machine learning-supported interpretation of kidney graft elementary lesions in combination with clinical data. Am J Transplant. 2022;22(12):2821–2833. doi: 10.1111/ajt.17192 [DOI] [PubMed] [Google Scholar]
- 104.Raynaud M, Aubert O, Divard G, et al. Dynamic prediction of renal survival among deeply phenotyped kidney transplant recipients using artificial intelligence: an observational, international, multicohort study. Lancet Digit Health. 2021;3(12):e795–e805. doi: 10.1016/S2589-7500(21)00209-0 [DOI] [PubMed] [Google Scholar]
- 105.Osmanodja B, Spencker JJ, Ö e Ö, et al. Randomized trial of electronic health record implemented AI risk prediction in kidney transplant care. NPJ Digit Med. 2026;9(1):373. doi: 10.1038/s41746-026-02757-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Churpek MM, Fatima A, Anjorin O, et al. Early nephrology consultation and acute kidney injury in hospitalized patients: a randomized clinical trial. JAMA Network Open. 2026;9(7):e2622554. doi: 10.1001/jamanetworkopen.2026.22554 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.van Royen FS, Weerts HJP, de Hond AAH, et al. In humble defense of unexplainable black box prediction models in healthcare. J Clin Epidemiol. 2026;189:112013. doi: 10.1016/j.jclinepi.2025.112013 [DOI] [PubMed] [Google Scholar]
- 108.Li Y, Mamouei M, Salimi-Khorshidi G, et al. Hi-BEHRT: hierarchical transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records. IEEE J Biomed Health Inform. 2023;27(2):1106–1117. doi: 10.1109/JBHI.2022.3224727 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Kashani M, Ninan J, Wei L, et al. International Delphi consensus on acute kidney injury: foundations for AI-driven digital twin development in critical care nephrology. PLoS One. 2026;21(3):e0344991. doi: 10.1371/journal.pone.0344991 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Deng T, Xue Y, Methakanjanasak N. Digital health integration in chronic kidney disease. Clin Chim Acta. 2026;582:120749. doi: 10.1016/j.cca.2025.120749 [DOI] [PubMed] [Google Scholar]
- 111.Wi CI, Overgaard S, Malik M, et al. A lifecycle governance and learning health system framework for trustworthy, generalizable, and sustainable human-ai partnership in clinical practice: lessons from the asthma-guidance and prediction system (A-GPS). J Natl Med Assoc. 2026;118:789–816. doi: 10.1016/j.jnma.2026.04.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Collins GS, Dhiman P, Andaur Navarro CL, et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ open. 2021;11(7):e048008. doi: 10.1136/bmjopen-2020-048008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Gallifant J, Afshar M, Ameen S, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nature Med. 2025;31(1):60–69. doi: 10.1038/s41591-024-03425-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Moons KG, de Groot JA, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744. doi: 10.1371/journal.pmed.1001744 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Wolff RF, Moons KGM, Riley RD et al, PROBAST Group. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann Intern Med. 2019;170(1):51–58. doi: 10.7326/M18-1376 [DOI] [PubMed] [Google Scholar]
- 117.Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Med. 2022;28(5):924–933. doi: 10.1038/s41591-022-01772-9 [DOI] [PubMed] [Google Scholar]
- 118.Rivera SC, Liu X, Chan A-W, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Lancet Digital Health. 2020;2(10):e549–e60. doi: 10.1016/S2589-7500(20)30219-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 119.Ibrahim H, Liu X, Rivera SC, et al. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines. Trials. 2021;22(1):11. doi: 10.1186/s13063-020-04951-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Mongan J, Moy L, Kahn CE Jr. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): A Guide for Authors and Reviewers. Radiol Artif Intell. 2020;2(2):e200029. doi: 10.1148/ryai.2020200029 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Lekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025:e081554. doi: 10.1136/bmj-2024-081554 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Milders J, Ramspek CL, Janse RJ, et al. Prognostic models in nephrology: where do we stand and where do we go from here? Mapping out the evidence in a scoping review. J Am Soc Nephrol. 2024;35(3):367–380. doi: 10.1681/ASN.0000000000000285 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med. 2021;181(8):1065–1070. doi: 10.1001/jamainternmed.2021.2626 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Ferryman K, Mackintosh M, Ghassemi M. Considering biased data as informative artifacts in AI-assisted health care. N Engl J Med. 2023;389(9):833–838. doi: 10.1056/NEJMra2214964 [DOI] [PubMed] [Google Scholar]
- 125.Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447–453. doi: 10.1126/science.aax2342 [DOI] [PubMed] [Google Scholar]
- 126.Inker LA, Eneanya ND, Coresh J, et al. New creatinine- and cystatin C-based equations to estimate GFR without race. N Engl J Med. 2021;385(19):1737–1749. doi: 10.1056/NEJMoa2102953 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Wilk AS, Cummings JR, Plantinga LC, Franch HA, Lea JP, Patzer RE. Racial and ethnic disparities in kidney replacement therapies among adults with kidney failure: an observational study of variation by patient age. Am J Kidney Dis. 2022;80(1):9–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Rizzolo K, Shen JI. Barriers to home dialysis and kidney transplantation for socially disadvantaged individuals. Curr Opin Nephrol Hypertens. 2024;33(1):26–33. doi: 10.1097/MNH.0000000000000939 [DOI] [PubMed] [Google Scholar]
- 129.Ferber D, Hilgers L, Hoper C, et al. Towards autonomous medical artificial intelligence agents. Nature. 2026;655:1282–1291. doi: 10.1038/s41586-026-10675-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Tun HM, Rahman HA, Naing L, Malik OA. Trust in artificial intelligence-based clinical decision support systems among health care workers: systematic review. J Med Internet Res. 2025;27:e69678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Jin D, Sergeeva E, Weng WH, Chauhan G, Szolovits P. Explainable deep learning in healthcare: a methodological survey from an attribution view. WIREs Mech Dis. 2022;14(3):e1548. doi: 10.1002/wsbm.1548 [DOI] [PubMed] [Google Scholar]
- 132.Food US, Drug A. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions: Guidance for Industry and Food and Drug Administration Staff. Final Guidance. Silver Spring, MD: US Food and Drug Administration; 2025. [Google Scholar]
- 133.Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act), (2024/06/13, 2024).
- 134.Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), (2024/06/13, 2024).
- 135.Ho CW, Caals K. A call for an ethics and governance action plan to harness the power of artificial intelligence and digitalization in nephrology. Semin Nephrol. 2021;41(3):282–293. [DOI] [PubMed] [Google Scholar]
- 136.Parsons CS, Zuiderwijk A, Orchard NA, Oosterhoff JHF, de Reuver M. Task-technology fit of artificial intelligence-based clinical decision support systems: a review of qualitative studies. BMC Med Inform Decis Mak. 2025;25(1):397. doi: 10.1186/s12911-025-03237-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 137.Potnis KC, Ross JS, Aneja S, Gross CP, Richman IB. Artificial intelligence in breast cancer screening: evaluation of FDA device regulation and future recommendations. JAMA Intern Med. 2022;182(12):1306–1312. doi: 10.1001/jamainternmed.2022.4969 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 138.Nesa L, Rony MKK, Chowdhury S, et al. Artificial intelligence in healthcare: a scoping review of medical professionals’ acceptance and institutional challenges in implementation. J Eval Clin Pract. 2025;31(4):e70170. doi: 10.1111/jep.70170 [DOI] [PubMed] [Google Scholar]
- 139.Li LT, Haley LC, Boyd AK, Bernstam EV. Technical/Algorithm, Stakeholder, and Society (TASS) barriers to the application of artificial intelligence in medicine: a systematic review. J Biomed Inform. 2023;147:104531. doi: 10.1016/j.jbi.2023.104531 [DOI] [PubMed] [Google Scholar]
- 140.Thongprayoon C, Pesce F, Cheungpasitporn W. Clinical artificial intelligence agents in nephrology: from prediction to action through workflow-native intelligence-a roadmap for workflow-integrated care. J Clin Med. 2026;15(7). doi: 10.3390/jcm15072576 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 141.Bakas S, Li X, Shah P, Roth HR. Federated learning in healthcare: from research to real-world deployment. Annu Rev Biomed Eng. 2026;28(1):163–186. doi: 10.1146/annurev-bioeng-080125-041414 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 142.Sheller MJ, Edwards B, Reina GA, et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci Rep. 2020;10(1):12598. doi: 10.1038/s41598-020-69250-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Wang Q, Luo Q, Ding Y, Wan S, Zhang Y, Xiong F Prediction of imminent peritoneal dialysis-associated peritonitis using time-updated electronic health records and machine learning: A temporal validation study. J Inflamm Res. 2026;19. doi: 10.2147/JIR.S595197:595197. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No new datasets were generated or analyzed during the preparation of this narrative review.
