ABSTRACT
Aim
Develop and evaluate a chronic kidney disease (CKD) electronic (e‐) phenotype to identify and risk‐stratify CKD patients at scale using electronic health record (EHR) data.
Methods
Patient encounter and laboratory data determining kidney function and proteinuria from a four‐year period (2021–2024) at a large Australian tertiary referral hospital was extracted from the hospital EHR. Following exclusion of kidney transplant recipients and patients on dialysis, eligible patients were classified as having CKD based on ICD‐10 codes, decreased estimated glomerular filtration rate (eGFR, < 60 mL/min per 1.73 m2) and/or albuminuria (urine albumin‐to‐creatinine ratio [uACR] ≥ 3 mg/mmol, A2/A3 equivalent) for a minimum of 3 months. Patients with sufficient laboratory data were also staged according to the Kidney Disease Improving Global Outcomes (KDIGO) A‐by‐G grid.
Results
From a total undifferentiated hospital cohort of n = 342 146, the CKD e‐phenotype algorithm identified n = 17 908 likely CKD cases (5.2%). Algorithm‐detected CKD cases were validated against blinded manual chart review (n = 200), revealing sensitivity and specificity for CKD by ICD‐10 codes (0.85 and 0.83), eGFR (0.62 and 0.98) and proteinuria criteria (0.33 and 0.75). The proportion of patients that were ICD‐10‐coded for CKD increased significantly (p < 0.001) with increasing G‐stage (odds ratio [OR] 3.05) and A‐stage (OR 2.31).
Conclusion
Incorporating eGFR and proteinuria data into the e‐phenotype algorithm appears to identify novel CKD cases with high specificity, particularly in the earlier stages of CKD progression. However, overall sensitivity is limited when compared to ICD‐10 CKD codes alone.
Keywords: chronic, electronic health records, glomerular filtration rate, International Classification of Diseases, phenotype, renal insufficiency
Summary at a Glance
Chronic kidney disease (CKD) can be readily identified via an electronic (e‐) phenotype established using data extracted from the electronic health record. This article describes the design, implementation and validation of a CKD e‐phenotype to detect and risk‐stratify patients across a diverse Australian tertiary hospital cohort.
1. Introduction
Chronic kidney disease (CKD) represents a spectrum of persistent deterioration in kidney function that has evolved into a rising global healthcare burden [1]. Advanced CKD is associated with significant healthcare utilisation, profound morbidity and high mortality [2]. However, a prolonged asymptomatic period between early‐ and late‐stage disease allows an indolent and often clinically silent progression towards irreversible kidney failure [3]. CKD is therefore frequently under‐diagnosed, under‐recognised and under‐investigated by healthcare providers [4]. Late disease detection also stalls timely implementation of kidney‐protective measures, such as renin‐angiotensin system (RAS) blockade and sodium‐glucose cotransporter‐2 (SGLT‐2) inhibitors, that are crucial for preserving kidney function in CKD [5, 6].
Electronic health records (EHRs) have emerged as promising avenues to bridge this gap in delayed CKD identification and allow targeted interventions to improve CKD care [7]. Given that CKD is typically detected and staged on the basis of objective laboratory criteria, including estimated glomerular filtration rate (eGFR) and urine albumin‐to‐creatinine ratio (uACR), it represents an ideal condition for an EHR‐derived electronic (e‐) phenotype [8]. Leveraging these significant clinical data sets towards a CKD e‐phenotype presents a novel opportunity to facilitate large‐scale screening, public health resourcing and clinical trial recruitment for patients with CKD [9]. In addition, a robust CKD e‐phenotype could offer rapid population‐level identification of patients meeting criteria for appropriate nephrology or surgical referral [8]. In practice, various iterations of a CKD e‐phenotype have therefore been developed and interrogated, with a broad range observed in algorithm performance [10].
Though simple, CKD diagnostic codes alone, which are often organised according to the International Classification of Diseases (ICD), have been repeatedly established as inadequate, demonstrating poor sensitivity and specificity for detecting CKD [11, 12, 13]. Even in recent studies, the sensitivity of clinical coding data for CKD has been reported as low as 0.45 [13, 14] and interpretation is often limited by varied practices across institutions [11]. On the opposing end of the complexity spectrum are sophisticated machine learning models that integrate laboratory values, clinical notes and patient demographic data to identify cases of CKD [15, 16, 17]. However, this e‐phenotype approach is often accurate at the expense of EHR interoperability and algorithm transparency and is limited by the need for advanced health information technology (IT) infrastructure [16].
More pragmatic methods have evaluated an e‐phenotype defined by Kidney Disease Improving Global Outcomes (KDIGO) criteria that incorporate eGFR, proteinuria and evidence of chronicity [3]. Notable efforts have reported e‐phenotype sensitivity and specificity as high as 0.99 [18], but many studies have been vulnerable to healthcare fragmentation, incomplete data and confounding due to intercurrent acute kidney injury (AKI) [19, 20, 21]. In addition, despite the emphasis placed by KDIGO on the importance of CKD risk stratification by both eGFR (G1–G5) and albuminuria (A1–A3) staging, e‐phenotype studies directly incorporating these elements are rare [21, 22]. Moreover, and importantly given the heterogeneity in both the CKD population and EHR practices across healthcare institutions, there has been limited exploration of these e‐phenotype elements in an Australian context [22, 23, 24].
This study therefore aimed to develop and validate a CKD e‐phenotype defined by a combination of ICD 10th Revision (ICD‐10) coding data and laboratory values (eGFR and proteinuria) across a diverse inpatient and outpatient hospital cohort. ICD‐10 codes were incorporated both to compare diagnostic performance against laboratory criteria and provide a surrogate for additional KDIGO markers of kidney damage such as structural abnormalities. Secondary objectives included demonstrating clinical applications of the CKD e‐phenotype, most notably its potential for large‐scale CKD risk‐stratification by KDIGO eGFR (G) and albuminuria (A) categories and evaluating the distribution, consistency and appropriateness of CKD‐related ICD‐10 diagnostic codes across the EHR. It was hypothesised that e‐phenotype sensitivity and specificity would consistently outperform ICD‐10 coding systems alone, in particular to identify CKD patients in low‐ and moderate‐risk categories.
2. Methods
2.1. Patient Cohort
The study population was a single‐centre cohort from The Royal Melbourne Hospital (RMH), a large inpatient tertiary referral hospital with integrated emergency department and outpatient services in Melbourne, Australia. Adult patients (≥ 18 years of age) with EHR evidence of an inpatient admission, emergency department presentation or an outpatient encounter within the time period from 1 January 2021 to 31 December 2024 were eligible for data extraction and subsequent e‐phenotype evaluation. Of note, the hospital EHR (Epic, Epic Systems Corporation) had only been introduced at RMH since August 2020.
Patient sub‐cohorts were defined as CKD e‐phenotype negative, CKD e‐phenotype positive, Dialysis, Transplant and Other (representing insufficient data). Summary statistics were computed for cohort age, sex, overseas nationality, First Nations status and inpatient admission data. Patient postcodes also facilitated cohort stratification by Modified Monash Model (MMM) category, which classifies Australian state regions according to geographical remoteness and town size [25]. MMM categories range from 1 (metropolitan areas) to 7 (very remote communities) and reflect patient access to healthcare workforce and services [25].
2.2. Data Extraction and Cleaning
Once the overall patient cohort was defined, various data points were derived from the hospital EHR system (Epic) for the purposes of demographic analysis and e‐phenotype algorithm implementation. The Epic Clarity database and hospital billing infrastructure were queried to tabulate all encounters within the study time period and their associated ICD‐10 diagnostic codes. Encounter‐level ICD‐10 data were then parsed to generate flags for AKI‐related encounters and aggregated to generate a patient‐level framework for CKD, dialysis and kidney transplant cohorts. A complete overview of relevant ICD‐10 codes incorporated into the e‐phenotype algorithm is presented in Table S1.
Laboratory data was also extracted across four different pathology services (including in‐house pathology) that integrate results within the RMH EHR. Laboratory variables of interest comprised eGFR as well as various proteinuria markers, such as ‘spot’ uACR and 24‐h uACR (A24), ‘spot’ and 24‐h urine protein‐to‐creatinine ratio (uPCR and P24 respectively), urine point‐of‐care (POC) dipstick (urinalysis) and urine microscopy culture and sensitivity (MC&S) results. Patient eGFR values were calculated using the creatinine‐based CKD Epidemiology Collaboration (CKD‐EPI) equation [26].
Data cleaning and all subsequent statistical analysis were performed using base R Statistical Software (version 4.4.2) [27] with the addition of the tidyverse set of packages for assistance with data transformation [28].
2.3. CKD e‐Phenotype Algorithm Development
Following the data extraction and cleaning phases, the final data set was reviewed and an e‐phenotype algorithm for CKD identification and classification was designed and implemented. The CKD e‐phenotype comprised sequential ICD‐10 code logic checks and a range of EHR‐derived laboratory values (eGFR and proteinuria markers). The initial patient triage process of the executed CKD e‐phenotype algorithm is provided in Figure 1.
FIGURE 1.

Initial triage of patient cohort with exclusion of kidney transplant and dialysis cases, followed by CKD e‐phenotype classification by ICD‐10 codes, eGFR and proteinuria criteria. CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate; ICD‐10, International Classification of Diseases, 10th Revision; KDIGO, Kidney Disease Improving Global Outcomes; MC&S, microscopy, culture and sensitivity; uACR, urine albumin‐to‐creatinine ratio; uPCR, urine protein‐to‐creatinine ratio; POC, point‐of‐care.
First, patients with an associated ICD‐10 code for kidney transplant, then dialysis, were excluded from the CKD e‐phenotype cohort. Dialysis maintenance therapy consistent with kidney failure was defined as evidence of at least two dialysis‐coded encounters. Then, in recognition of concurrent AKI as a potential confounder of creatinine steady state, eGFR laboratory values from patient encounters with ICD‐10 AKI codes were excluded from the e‐phenotype data set.
The next stage of the e‐phenotype algorithm created three overlapping KDIGO‐aligned CKD positive sub‐cohorts, as outlined in Figure 1. Patients were considered e‐phenotype positive if they had evidence of decreased eGFR (< 60 mL/min per 1.73 m2) or albuminuria (uACR ≥ 3 mg/mmol, A2/A3 equivalent) present for a minimum of 3 months (90 days) [3]. The presence of an ICD‐10 CKD code was used as a surrogate for alternative markers of kidney damage that can also constitute potential CKD diagnostic criteria [3].
Accessible EHR laboratory data further facilitated classification of patient CKD severity by eGFR (G‐staging, G1–G5) and/or persistent albuminuria values (A‐staging, A1–A3), which represent important dimensions for risk stratification and prognostication in CKD. To account for the CKD chronicity criterion, patients required at least two laboratory results for eGFR and/or proteinuria at least 90 days apart to be eligible for G‐ and A‐staging respectively, and were otherwise classified as incomplete cases (G0) for each dimension. Overviews of the G‐ and A‐staging algorithms are provided in Figures 2 and 3, respectively.
FIGURE 2.

CKD e‐phenotype eGFR‐based G‐staging algorithm to classify patients according to G1–G5 categories or G0 if incomplete. CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate; ICD‐10, International Classification of Diseases, 10th Revision.
FIGURE 3.

CKD e‐phenotype proteinuria‐based A‐staging algorithm to classify patients according to A1–A3 categories or A0 if incomplete. CKD, chronic kidney disease; MC&S, microscopy, culture and sensitivity; uACR, urine albumin‐to‐creatinine ratio; uPCR, urine protein‐to‐creatinine ratio; POC, point‐of‐care.
Given the anticipated paucity of gold‐standard albuminuria (e.g., uACR) data, a number of surrogate proteinuria markers were incorporated to improve e‐phenotype A‐staging sensitivity, including uPCR and urine dipstick results. The alignment process for consistent A‐staging across different proteinuria results is described alongside relevant test thresholds for A1–A3 stages in Figure 3. Urine PCR values were treated quantitatively, whereas urine dipstick proteinuria values were correlated to semi‐quantitative ordinal albuminuria categories. For proteinuria values, the e‐phenotype algorithm was also adjusted to emphasise clinical utility over recency. As such, proteinuria markers were prioritised in the order of uACR (spot or 24‐h) and uPCR (spot or 24‐h) followed by dipstick (POC) and MC&S results, which were only relied upon if a quantitative urine protein test was unavailable.
Of note for G‐staging, as the G1A1 and G2A1 CKD classifiers only represent CKD in the setting of kidney damage, these patients were therefore further divided by the presence (disease) or absence (no disease) of an associated ICD‐10 CKD code as a proxy marker of kidney damage. Finally, to improve coverage of the KDIGO A‐by‐G grid, patients that were initially eligible for staging in one dimension but not the other (e.g., G0A2 or G4A0), were staged in the missing dimension according to their most recent eGFR result or most recent (adjusted) urine protein test, if available.
2.4. Evaluation of ICD‐10 CKD Coding Trends
The proportion of patients in each A‐by‐G grid category that were classified as having an associated ICD‐10 CKD code was also computed. Binomial logistic regression models were used to evaluate trends in CKD coding behaviour across disease severity, testing independent effects of increasing A‐ and G‐stage after controlling for age and sex covariates. Odds ratios (OR) and 95% confidence intervals (CI) were calculated with statistical significance defined at a Bonferroni‐corrected threshold of α adjusted = 0.025.
2.5. CKD e‐Phenotype Validation
Following implementation of the e‐phenotype as described, a subset of the overall patient cohort was randomly selected for manual chart validation as a gold‐standard comparison. CKD e‐phenotype algorithm positive cases (total n = 150) were randomly drawn from each of the sub‐cohorts including CKD by ICD‐10 codes (n = 50), CKD by eGFR (n = 50) and CKD by proteinuria (n = 50). To reduce bias from missing data and sparsely populated EHR files, CKD e‐phenotype negative cases (n = 50) were derived from the G1A1 and G2A1 patients that lacked an ICD‐10 CKD code.
Each patient within the total validation cohort (n = 200) then underwent blinded manual chart review for classification as CKD (positive/negative) for comparison against algorithm‐derived diagnoses. The presence of a CKD‐related diagnosis within the patient's ‘Problem List’ or ‘Medical History’ EHR fields was also recorded. Chart review was performed with reference to KDIGO 2024 CKD diagnostic criteria and evaluation of any information accessible to the front‐end EHR user interface was permitted, including clinical notes, laboratory data, imaging and biopsy findings.
Algorithm‐identified dialysis and kidney transplant cohorts were validated with reference to a dedicated RMH clinical nephrology audit database, NephWorks, as gold standard. NephWorks is a bespoke, local database which records demographics and treatment details of all dialysis and transplant patients registered to the RMH nephrology network. This data set facilitated automated cross‐referencing, allowing validation of the entire patient cohort. In all cases, validation was limited to only information available prior to conclusion of the study time period on 31 December 2024.
Collected validation data was then un‐blinded and used to calculate true/false positives and true/false negatives alongside overall performance characteristics of the CKD e‐phenotype algorithm, such as sensitivity, specificity and accuracy. Corresponding 95% CI were calculated using binom.test in R version 4.4.2 [27].
2.6. Ethics Approval
The study obtained Melbourne Health ethics approval in December 2024 (Project number QA2024180).
3. Results
3.1. Population and Demographics
Initial EHR data extraction retrieved a total of 2 049 410 patient encounters within the specified timeframe, representing 342 146 unique patients. A total of 1677 kidney transplant recipients and 802 dialysis patients were identified and excluded, with the CKD e‐phenotype algorithm further classifying 17 908 CKD positive and 11 353 CKD negative cases. The remaining 310 406 patients (Other subgroup) either lacked a relevant ICD‐10 code or had insufficient laboratory data for full CKD e‐phenotype evaluation.
Descriptive and summary statistics of demographic and healthcare utilisation data for each patient subgroup as counts, proportions (%) or median (interquartile range [IQR]) are presented in Table 1. In general, patients identified by the e‐phenotype as CKD positive had a median age of 74 (60–84) years, 55.5% were male, 1.3% were born overseas, 0.9% identified as First Nations and were primarily from major metropolitan areas (MMM category 1). The CKD positive cohort had a median of 7 [3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16] encounters, 3 (0–9) eGFR values and 2 (0–4) proteinuria results.
TABLE 1.
CKD e‐phenotype cohort demographic and summary statistics, with values presented as counts, proportions (%) or median (IQR).
| CKD negative | CKD positive | Dialysis | Transplant | Other | |
|---|---|---|---|---|---|
| N | 11 353 | 17 908 | 802 | 1677 | 310 406 |
| Age | 52 (34–71) | 74 (60–84) | 69 (58–78) | 56 (45–66) | 42 (29–61) |
| Male | 42.1 | 55.5 | 65.0 | 62.3 | 50.5 |
| Overseas‐born | 0.6 | 1.3 | 2.8 | 0.7 | 1.7 |
| First Nations | 1.7 | 0.9 | 3.2 | 1.7 | 0.7 |
| MMM | 1 (1–1) | 1 (1–1) | 1 (1–3) | 1 (1–2) | 1 (1–1) |
| Encounters | 9 (4–21) | 7 (3–16) | 32 (12–169) | 27 (13–60) | 2 (1–5) |
| Ever inpatient | 94.8 | 86.1 | 97.5 | 79.8 | 43.6 |
| Ever ED | 98.0 | 73.5 | 63.2 | 50.3 | 53.6 |
| Ever outpatient | 77.4 | 80.4 | 86.4 | 98.0 | 65.7 |
| eGFR results | 6 (3–13) | 3 (0–9) | 12 (2–32) | 7 (0–23) | 0 (0–1) |
| Proteinuria results | 2 (1–4) | 2 (0–4) | 0 (0–2) | 3 (0–11) | 0 (0–0) |
Abbreviations: CKD, chronic kidney disease; ED, Emergency Department; eGFR, estimated glomerular filtration rate; MMM, Modified Monash Model.
3.2. CKD e‐Phenotype Implementation
Following exclusion of dialysis patients (n = 802) and kidney transplant recipients (n = 1677), there were 339 667 unique patients available for further CKD e‐phenotype consideration. Within this cohort, 15 849 patients were classified as CKD positive due to associated ICD‐10 codes, 4015 by eGFR criteria and 1955 by proteinuria criteria. The overlaps and corresponding patient counts, including relative proportions within these subgroups, following e‐phenotype implementation are represented in Figure 4. Although a majority of cases were identified by ICD‐10 codes alone (70.5%), a significant proportion of laboratory‐identified CKD cases (2059 out of 5283, 39%) did not have evidence of an ICD‐10 CKD code. Demographic data for each e‐phenotype subgroup is presented in Table S2, with accompanying histograms of most recent eGFR values provided in Figures S1–S4.
FIGURE 4.

Venn diagram of CKD e‐phenotype positive cases and overlaps across ICD‐10 code, eGFR and proteinuria subgroups. Kidney transplant recipients and dialysis patients are excluded from analysis. Values and overlaps are presented as counts with relative cohort proportions in brackets. CKD, chronic kidney disease; EHR, electronic health record; eGFR, estimated glomerular filtration rate; ICD‐10, International Classification of Diseases, 10th Revision.
A total of 19 865 patients were eligible for further evaluation with classification via the G‐ and A‐staging algorithm. The distribution of these cases across the dimensions of eGFR (G1–G5) and albuminuria (A1–A3) is shown in Table 2, alongside when a given dimension did not have adequate laboratory data for full staging (G0 and/or A0). Colours correspond to KDIGO‐aligned CKD risk categories, including low (green), moderate (yellow), high (orange) and very high (red) risk groups.
TABLE 2.
Counts of patients meeting CKD e‐phenotype criteria across KDIGO eGFR (G) and albuminuria (A) dimensions (with corresponding ICD‐10 CKD‐coded proportions).
| Albuminuria (A)‐staging (mg/mmol) | |||||||
|---|---|---|---|---|---|---|---|
| A0 | A1 | A2 | A3 | ||||
| Incomplete | Normal to mildly increased | Moderately increased | Severely increased | ||||
| < 3 | 3–30 | > 30 | |||||
| eGFR (G)‐staging (mL/min/1.73 m2) | G0 | Incomplete | 308 312 (0.034) | 464 (0.250) | 270 (0.363) | 84 (0.702) | |
| G1 | Normal or high | ≥ 90 | 7022 (0.009) | 7749 (0.046) | 1798 (0.119) | 241 (0.253) | |
| G2 | Mildly decreased | 60–89 | 2917 (0.055) | 4694 (0.153) | 1534 (0.300) | 268 (0.485) | |
| G3a | Mildly to moderately decreased | 45–59 | 319 (0.310) | 846 (0.461) | 499 (0.649) | 109 (0.817) | |
| G3b | Moderately to severely decreased | 30–44 | 270 (0.522) | 608 (0.664) | 465 (0.778) | 178 (0.916) | |
| G4 | Severely decreased | 15–29 | 105 (0.714) | 236 (0.814) | 289 (0.910) | 167 (0.970) | |
| G5 | Kidney failure | < 15 | 39 (0.795) | 34 (0.853) | 81 (0.951) | 69 (0.986) | |
Note: Low (green), moderate (yellow), high (orange) and very high (red) risk groups.
Abbreviations: CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate; ICD‐10, International Classification of Diseases, 10th Revision; KDIGO, Kidney Disease Improving Global Outcomes.
When stratified by KDIGO risk category, there were 12 443 patients determined to be low‐risk (1072 with CKD by ICD‐10 codes), 4178 patients at moderately increased risk, 1616 patients at high risk and 1628 patients at very high risk of CKD progression. Incomplete data (G0 and/or A0) was available for a total of 308 312 patients in both dimensions, 818 patients for G‐staging alone and 10 672 patients for A‐staging alone. The proportion of patients in each A‐by‐G grid cell that had evidence of an ICD‐10 CKD‐related code is also given in Table 2 in brackets beneath each count. There was a significant trend observed in the CKD‐coded proportion of each subgroup with advancing A‐stage (from A1–A3) and G‐stage (from G1–G5). This trend persisted after controlling for age and sex as potential confounders, with an OR (95% CI) of 3.05 (2.90, 3.20) for increasing G‐stage and 2.31 (2.16, 2.48) for increasing A‐stage (p < 0.001).
3.3. CKD e‐Phenotype Performance
Overall test performance characteristics of the CKD e‐phenotype, dialysis and kidney transplant algorithms, including relevant counts, sensitivity, specificity and accuracy, are described in Table 3. Diagnostic performance statistics are presented alongside computed 95% CI. Following manual chart review of the 150 algorithm‐positive and 50 algorithm‐negative cases, there were 113 true CKD positive patients and 87 true CKD negative cases identified using KDIGO 2024 CKD criteria as the reference gold standard.
TABLE 3.
Algorithm performance characteristics (95% CI) of the CKD e‐phenotype, EHR documentation and dialysis and kidney transplant detection following validation with manual chart review.
| TP | FP | TN | FN | Sensitivity | Specificity | Accuracy | ||
|---|---|---|---|---|---|---|---|---|
| CKD e‐phenotype positive | ICD‐10 | 96 | 15 | 72 | 17 | 0.85 (0.77, 0.91) | 0.83 (0.73, 0.90) | 0.84 (0.78, 0.89) |
| eGFR | 70 | 2 | 85 | 43 | 0.62 (0.52, 0.71) | 0.98 (0.92, 1.00) | 0.78 (0.71, 0.83) | |
| Proteinuria | 37 | 22 | 65 | 76 | 0.33 (0.24, 0.42) | 0.75 (0.64, 0.83) | 0.51 (0.44, 0.58) | |
| Any | 113 | 37 | 50 | 0 | 1.00 (0.97, 1.00) | 0.57 (0.46, 0.68) | 0.82 (0.75, 0.87) | |
| CKD in PL or MHx | 48 | 1 | 86 | 65 | 0.42 (0.33, 0.52) | 0.99 (0.94, 1.00) | 0.67 (0.60, 0.73) | |
| Dialysis | 733 | 69 | 340 618 | 726 | 0.50 (0.48, 0.53) | 1.00 (1.00, 1.00) | 1.00 (1.00, 1.00) | |
| Kidney transplant | 1580 | 97 | 340 221 | 248 | 0.86 (0.85, 0.88) | 1.00 (1.00, 1.00) | 1.00 (1.00, 1.00) | |
Abbreviations: CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate; EHR, electronic health record; FN, false negative; FP, false positive; ICD‐10, International Classification of Diseases, 10th Revision; MHx, medical history; PL, problem list; TN, true negative; TP, true positive.
The combined algorithm demonstrated a sensitivity of 1.00 (as expected given the sampling strategy), but a relatively poor specificity of 0.57. ICD‐10 codes had high sensitivity and specificity (0.85 and 0.83, respectively), whereas eGFR criteria had very high specificity (0.98) at the expense of sensitivity (0.62). Finally, proteinuria criteria had the poorest sensitivity (0.33) with only moderate specificity (0.75). Examples of discrepancies between algorithm‐derived ICD‐10 CKD diagnoses and manual chart review included transient hydronephrosis without persistent kidney dysfunction and a kidney donor code attributed to a patient later determined ineligible. CKD cases by proteinuria criteria were also impacted by low‐level intermittent positive urine dipstick results in patients with recurrent urinary tract infections or renal calculi that lacked a quantitative proteinuria test.
Manual chart review also involved assessment of appropriate CKD documentation in each patient's Problem List or Medical History EHR fields. After unblinding, CKD documentation was calculated to have poor sensitivity (0.42) but excellent specificity (0.99). In addition, kidney transplant and dialysis components of the algorithm were evaluated at a full cohort level with automated reference to the NephWorks registry database as gold standard. The e‐phenotype algorithm demonstrated limited sensitivity to identify dialysis patients (0.50), but improved sensitivity for kidney transplant recipients (0.86).
4. Discussion
A CKD e‐phenotype algorithm comprising ICD‐10 codes, eGFR and proteinuria values was implemented and validated across a diverse tertiary hospital population within a four‐year study period (n = 342 146). A total of 17 908 patients (5.2%) were classified by the e‐phenotype algorithm as CKD positive, despite latest Australia‐specific estimates reporting an overall CKD prevalence of 11% [29]. Incorporating eGFR and proteinuria data into the e‐phenotype algorithm appears to identify novel CKD cases with high specificity, particularly in the earlier stages of CKD progression. Overall sensitivity, however, is limited when compared to ICD‐10 CKD codes alone and likely still underestimates true cohort CKD prevalence.
Landmark CKD e‐phenotyping efforts by Norton et al. using similar eGFR and proteinuria criteria demonstrated an impressive sensitivity and specificity of 0.99 across five healthcare organisations [18]. However, the study validation cohort was drawn solely from the laboratory‐defined CKD subgroup, thereby possibly missing cases of CKD from diagnostic codes that could have otherwise limited its overall sensitivity [18]. In contrast to our study, this approach did not adapt its algorithm to exclude AKI‐related results, representing a potential confounder for eGFR criteria [18].
In an Australian setting, Chen et al. describe a CKD e‐phenotype algorithm with an overall sensitivity of 0.93 and specificity of 0.73, identifying patients with albuminuria or persistent eGFR < 60 mL/min/1.73 m2 [22]. Improved sensitivity in this study was likely facilitated by using a database of ‘active’ patients with pre‐identified CKD or associated CKD risk factors that integrated up to 24 years of longitudinal hospital and primary care data [22]. By comparison, in our study, many patients lacked sufficient long‐term laboratory data to satisfy or be adequately evaluated for persistent changes in eGFR or proteinuria.
Recently, Shang et al. have reported a CKD e‐phenotype algorithm that incorporates a range of diagnostic code terminology, eGFR and proteinuria tests, achieving 0.87 sensitivity and 0.97 specificity [21]. The study also stratified its CKD cohort according to the KDIGO A‐by‐G grid, although it employed more complex machine learning strategies to optimise urine dipstick and uPCR values across albuminuria categories [21]. Our approach, which prioritises test accuracy over recency with simple rule‐based ordinal cut‐offs, may therefore be preferable for healthcare institutions that lack this dedicated informatics infrastructure.
Strengths of our study include an implemented CKD e‐phenotype algorithm which is practical, modifiable and maintains high performance characteristics across a diverse and undifferentiated hospital inpatient and outpatient cohort. Additional strengths include the exclusion of AKI‐associated eGFR data to limit confounding from intercurrent fluctuations in serum creatinine steady state, improving the specificity of CKD by eGFR criteria. The described strategy in our study also manages to prioritise clinically relevant, quantitative proteinuria tests (uACR, uPCR), rather than attempting to infer albuminuria categories or persistent proteinuria from recent urine dipstick results alone.
Indeed, the algorithm in our study can therefore extend on a simple CKD case‐finding e‐phenotype approach, facilitating cohort risk stratification by both KDIGO G‐ and A‐staging dimensions. In particular, the CKD e‐phenotype algorithm can be reproduced and scaled with relatively few modifications. Future replication may therefore benefit from increased longitudinal data accumulation and improved interoperability between electronic health systems.
The accuracy and reliability of any EHR‐based retrospective data analysis however is inherently limited by the completeness of its underlying data set as entered during routine clinical care. Incomplete ICD‐10 coding data, sparse laboratory values and limited follow‐up therefore likely limited the overall sensitivity of the CKD e‐phenotype and subsequent KDIGO G‐ and A‐staging evaluation. As such, the described e‐phenotype approach remained vulnerable to healthcare fragmentation and variations in data formatting such as inaccessible scanned or faxed laboratory results. In addition, the use of CKD‐related ICD‐10 codes represented an imperfect surrogate for alternative markers of kidney damage. This information, including imaging or biopsy findings, often resides in free text documentation that was not captured by the data extraction process.
Regarding laboratory results, cohort‐level extraction and strict cut‐off criteria will lack the discretion of clinician‐adjudicated CKD, with the algorithmic approach swayed disproportionately by outlier or threshold values. For example, alternate influences on serum creatinine steady state that may not have been explicitly coded as episodes of AKI, such as sepsis or shock, may escape exclusion from the eGFR data set and impact specificity. Similarly, although the implemented algorithm prioritised more accurate quantitative proteinuria results (uACR, uPCR), supplemental urine dipstick values are more susceptible to confounding by infection, renal calculi or fluctuations in specific gravity.
Moreover, our study evaluated only a cross‐sectional application of the CKD e‐phenotype algorithm. Analysis of variables such as ‘time to CKD diagnosis’ for each set of criteria may have provided insight into the influence of observation duration, loss to follow‐up and documentation practices on e‐phenotype accuracy. Our study therefore did not explore temporal relationships, such as whether CKD positivity by eGFR criteria tended to precede or follow ICD‐10 coding and how this may have affected sensitivity estimates. In addition, the impact of individual clinical settings on e‐phenotype performance was not assessed in favour of integrating patient data across inpatient, outpatient and emergency department encounters.
Ongoing research in this area will likely involve CKD e‐phenotype algorithm refinement, further characterisation of the CKD positive cohort and evaluation of the e‐phenotype's clinical applications. Altering ICD‐10 code criteria and adjusting for potential confounders of proteinuria results, such as urinary tract infections or renal calculi, could further improve algorithm specificity. In addition, free text analysis of clinician documentation (notes, letters, discharge summaries) for CKD‐related terms could prove a useful e‐phenotype adjunct to capture CKD cases with insufficient or inaccessible laboratory values [23].
Future efforts may investigate additional characteristics of the CKD e‐phenotype cohort, such as prescribed medications and screen for evidence of CKD‐related complications (haemoglobin, parathyroid hormone) or comorbidities (diabetes mellitus, hypertension, ischaemic heart disease) [14, 19, 21]. Indeed, the ability to improve CKD detection at the moderately increased to high‐risk categories that correlate with earlier disease stages may enable prompt, targeted evaluation for initiation of kidney‐protective RAS blockade or SGLT‐2 inhibitor therapy [6, 14, 30]. For patients identified at later stages of CKD, e‐phenotype elements and laboratory data could be leveraged to determine eligibility for nephrologist referral and allow planning for kidney replacement therapy or conservative kidney care [31].
Finally, at a population level, the CKD e‐phenotype may have applications for providing a valuable database for public health resourcing, forecasting and clinical trial recruitment [7, 9]. EHR‐based clinical trials offer a platform to accelerate enrolment, allocation and data collection and will benefit from leveraging an e‐phenotyping algorithm that can reliably identify patients with CKD and/or AKI [32]. This may include integrated evaluation of interventions such as decision support systems to aid CKD diagnosis, identify risk factors and promote guideline‐directed therapy [33, 34].
In conclusion, a CKD e‐phenotype algorithm comprising ICD‐10 codes, eGFR results and proteinuria criteria was designed, implemented and validated across a diverse and undifferentiated hospital inpatient and outpatient cohort. The described algorithm is practical, modifiable and scalable for downstream longitudinal reproducibility. Our study suggests an important role for laboratory‐defined CKD e‐phenotype elements, particularly to identify CKD cases at earlier stages of disease.
Funding
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Table S1: ICD‐10 diagnostic codes used in the CKD e‐phenotype algorithm. Unless otherwise stated as excluded (excl.), all subcodes within each listed category were included.
Table S2: CKD e‐phenotype positive overlapping subgroup demographic data, with values presented as counts, proportions (%) or median (IQR).
Figure S1: Histogram of most recent eGFR values available for patients meeting any of the CKD e‐phenotype ICD‐10, eGFR or proteinuria criteria. eGFR values were not available for n = 5035 patients.
Figure S2: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype ICD‐10 criteria. eGFR values were not available for n = 5029 patients.
Figure S3: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype eGFR criteria.
Figure S4: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype proteinuria criteria. eGFR values were not available for n = 51 patients.
Acknowledgements
The authors have nothing to report. Open access publishing facilitated by The University of Melbourne, as part of the Wiley ‐ The University of Melbourne agreement via the Council of Australasian University Librarians.
Data Availability Statement
The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.
References
- 1. Xie Y., Bowe B., Mokdad A. H., et al., “Analysis of the Global Burden of Disease Study Highlights the Global, Regional, and National Trends of Chronic Kidney Disease Epidemiology From 1990 to 2016,” Kidney International 94, no. 3 (2018): 567–581, 10.1016/j.kint.2018.04.011. [DOI] [PubMed] [Google Scholar]
- 2. Evans M., Lewis R. D., Morgan A. R., et al., “A Narrative Review of Chronic Kidney Disease in Clinical Practice: Current Challenges and Future Perspectives,” Advances in Therapy 39, no. 1 (2022): 33–43, 10.1007/s12325-021-01927-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Stevens P. E., Ahmed S. B., Carrero J. J., et al., “KDIGO 2024 Clinical Practice Guideline for the Evaluation and Management of Chronic Kidney Disease,” Kidney International 105, no. 4 (2024): S117–S314, 10.1016/j.kint.2023.10.018. [DOI] [PubMed] [Google Scholar]
- 4. Bansal S., Mader M., and Pugh J. A., “Screening and Recognition of Chronic Kidney Disease in VA Health Care System Primary Care Clinics,” Kidney360 1, no. 9 (2020): 904–915, 10.34067/KID.0000532020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Shlipak M. G., Tummalapalli S. L., Boulware L. E., et al., “The Case for Early Identification and Intervention of Chronic Kidney Disease: Conclusions From a Kidney Disease: Improving Global Outcomes (KDIGO) Controversies Conference,” Kidney International 99, no. 1 (2021): 34–47, 10.1016/j.kint.2020.10.012. [DOI] [PubMed] [Google Scholar]
- 6. Zhuo M., Li J., Buckley L. F., et al., “Prescribing Patterns of Sodium‐Glucose Cotransporter‐2 Inhibitors in Patients With CKD: A Cross‐Sectional Registry Analysis,” Kidney360 3, no. 3 (2022): 455–464, 10.34067/KID.0007862021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Tummalapalli S. L. and Peralta C. A., “An Electronic CKD Phenotype: A Step Forward in Improving Kidney Care,” Clinical Journal of the American Society of Nephrology 14, no. 9 (2019): 1277–1279, 10.2215/CJN.08180719. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Drawz P. E., Archdeacon P., McDonald C. J., et al., “CKD as a Model for Improving Chronic Disease Care Through Electronic Health Records,” Clinical Journal of the American Society of Nephrology 10, no. 8 (2015): 1488–1499, 10.2215/CJN.00940115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Navaneethan S. D., Jolly S. E., Sharp J., et al., “Electronic Health Records: A New Tool to Combat Chronic Kidney Disease?,” Clinical Nephrology 79, no. 3 (2013): 175–183, 10.5414/CN107757. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Sparks C., Steinberg A. G., and Toussaint N. D., “Identifying and Characterising a Chronic Kidney Disease Electronic‐Phenotype Using Electronic Health Record‐Derived Data: A Narrative Review of Strategies and Applications,” Nephrology 30, no. 9 (2025): e70118, 10.1111/nep.70118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Grams M. E., Plantinga L. C., Hedgeman E., et al., “Validation of CKD and Related Conditions in Existing Data Sets: A Systematic Review,” American Journal of Kidney Diseases 57, no. 1 (2011): 44–54, 10.1053/j.ajkd.2010.05.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Norton J. M., Grunwald L., Banaag A., et al., “CKD Prevalence in the Military Health System: Coded Versus Uncoded CKD,” Kidney Medicine 3, no. 4 (2021): 586–595.e1, 10.1016/j.xkme.2021.03.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Sisk R., Cameron R., Tahir W., and Sammut‐Powell C., “Diagnosis Codes Underestimate Chronic Kidney Disease Incidence Compared With eGFR‐Based Evidence: A Retrospective Observational Study of Patients With Type 2 Diabetes in UK Primary Care,” BJGP Open 8, no. 1 (2024): BJGPO.2023.0079, 10.3399/BJGPO.2023.0079. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Molokhia M., Okoli G. N., Redmond P., et al., “Uncoded Chronic Kidney Disease in Primary Care: A Cross‐Sectional Study of Inequalities and Cardiovascular Disease Risk Management,” British Journal of General Practice 70, no. 700 (2020): e785–e792, 10.3399/bjgp20X713105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Mansour O., Paik J. M., Wyss R., et al., “A Novel Chronic Kidney Disease Phenotyping Algorithm Using Combined Electronic Health Record and Claims Data,” Clinical Epidemiology 15, no. 101531700 (2023): 299–307, 10.2147/CLEP.S397020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Rashidian S., Hajagos J., Moffitt R. A., et al., “Deep Learning on Electronic Health Records to Improve Disease Coding Accuracy,” AMIA Joint Summits on Translational Science proceedings 2019, no. 101539486 (2019): 620–629. [PMC free article] [PubMed] [Google Scholar]
- 17. Weber C., Roschke L., Modersohn L., et al., “Optimized Identification of Advanced Chronic Kidney Disease and Absence of Kidney Disease by Combining Different Electronic Health Data Resources and by Applying Machine Learning Strategies,” Journal of Clinical Medicine 9, no. 9 (2020): 2955, 10.3390/jcm9092955. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Norton J. M., Ali K., Jurkovitz C. T., et al., “Development and Validation of a Pragmatic Electronic Phenotype for CKD,” Clinical Journal of the American Society of Nephrology 14, no. 9 (2019): 1306–1314, 10.2215/CJN.00360119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Navaneethan S. D., Jolly S. E., Schold J. D., et al., “Development and Validation of an Electronic Health Record‐Based Chronic Kidney Disease Registry,” Clinical Journal of the American Society of Nephrology 6, no. 1 (2011): 40–49, 10.2215/CJN.04230510. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Ernecoff N. C., Wessell K. L., Hanson L. C., et al., “Electronic Health Record Phenotypes for Identifying Patients With Late‐Stage Disease: A Method for Research and Clinical Application,” Journal of General Internal Medicine 34, no. 12 (2019): 2818–2823, 10.1007/s11606-019-05219-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Shang N., Khan A., Polubriaginof F., et al., “Medical Records‐Based Chronic Kidney Disease Phenotype for Clinical Care and “Big Data” Observational and Genetic Studies,” npj Digital Medicine 4, no. 1 (2021): 70, 10.1038/s41746-021-00428-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Chen W., Abeyaratne A., Gorham G., et al., “Development and Validation of Algorithms to Identify Patients With Chronic Kidney Disease and Related Chronic Diseases Across the Northern Territory, Australia,” BMC Nephrology 23, no. 1 (2022): 320, 10.1186/s12882-022-02947-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Canaway R., Chidgey C., Hallinan C. M., Capurro D., and Boyle D. I., “Undercounting Diagnoses in Australian General Practice: A Data Quality Study With Implications for Population Health Reporting,” BMC Medical Informatics and Decision Making 24, no. 1 (2024): 155, 10.1186/s12911-024-02560-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Ko S., Venkatesan S., Nand K., Levidiotis V., Nelson C., and Janus E., “International Statistical Classification of Diseases and Related Health Problems Coding Underestimates the Incidence and Prevalence of Acute Kidney Injury and Chronic Kidney Disease in General Medical Patients,” Internal Medicine Journal 48, no. 3 (2018): 310–315, 10.1111/imj.13729. [DOI] [PubMed] [Google Scholar]
- 25. Australian Government Department of Health, Disability and Ageing , Modified Monash Model, 2025, https://www.health.gov.au/topics/rural‐health‐workforce/classifications/mmm.
- 26. Levey A. S., Stevens L. A., Schmid C. H., et al., “A New Equation to Estimate Glomerular Filtration Rate,” Annals of Internal Medicine 150 (2009): 604–612. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. R Core Team , R: A Language and Environment for Statistical Computing. Published Online 2024, https://www.R‐project.org/.
- 28. Wickham H., Averick M., Bryan J., et al., “Welcome to the Tidyverse,” Journal of Open Source Software 4, no. 43 (2019): 1686, 10.21105/joss.01686. [DOI] [Google Scholar]
- 29. Australian Institute of Health and Welfare , Chronic Kidney Disease: Australian Facts (Australian Institute of Health and Welfare, 2024), 2024, https://www.aihw.gov.au/reports/chronic‐kidney‐disease/chronic‐kidney‐disease. [Google Scholar]
- 30. Forbes A. K., Hinton W., Feher M. D., et al., “Implementation of Chronic Kidney Disease Guidelines for Sodium‐Glucose Co‐Transporter‐2 Inhibitor Use in Primary Care in the UK: A Cross‐Sectional Study,” EClinicalMedicine 68, no. 101733727 (2024): 102426, 10.1016/j.eclinm.2024.102426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Mendu M. L., Ahmed S., Maron J. K., et al., “Development of an Electronic Health Record‐Based Chronic Kidney Disease Registry to Promote Population Health Management,” BMC Nephrology 20, no. 1 (2019): 72, 10.1186/s12882-019-1260-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Abdel‐Kader K. and Jhamb M., “EHR‐Based Clinical Trials: The Next Generation of Evidence,” Clinical Journal of the American Society of Nephrology 15, no. 7 (2020): 1050–1052, 10.2215/CJN.11860919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Pefanis A., Botlero R., Langham R. G., and Nelson C. L., “eMAP:CKD: Electronic Diagnosis and Management Assistance to Primary Care in Chronic Kidney Disease,” Nephrology, Dialysis, Transplantation 33, no. 1 (2018): 121–128, 10.1093/ndt/gfw366. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Khoong E. C., Karliner L., Lo L., et al., “A Pragmatic Cluster Randomized Trial of an Electronic Clinical Decision Support System to Improve Chronic Kidney Disease Management in Primary Care: Design, Rationale, and Implementation Experience,” JMIR Research Protocols 8, no. 6 (2019): e14022, 10.2196/14022. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Table S1: ICD‐10 diagnostic codes used in the CKD e‐phenotype algorithm. Unless otherwise stated as excluded (excl.), all subcodes within each listed category were included.
Table S2: CKD e‐phenotype positive overlapping subgroup demographic data, with values presented as counts, proportions (%) or median (IQR).
Figure S1: Histogram of most recent eGFR values available for patients meeting any of the CKD e‐phenotype ICD‐10, eGFR or proteinuria criteria. eGFR values were not available for n = 5035 patients.
Figure S2: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype ICD‐10 criteria. eGFR values were not available for n = 5029 patients.
Figure S3: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype eGFR criteria.
Figure S4: Histogram of most recent eGFR values available for patients meeting CKD e‐phenotype proteinuria criteria. eGFR values were not available for n = 51 patients.
Data Availability Statement
The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.
