Skip to main content
Clinical Kidney Journal logoLink to Clinical Kidney Journal
. 2025 Mar 8;18(4):sfaf073. doi: 10.1093/ckj/sfaf073

Identifying a cohort of hospitalized chronic kidney disease patients using electronic health records: lessons learnt and implications for future research and clinical practice guidelines

Daniel Fernández-Llaneza 1,2,3,, Luuk B Hilbrands 4, Liffert Vogt 5, Rik H G Olde Engberink 6, Joanna E Klopotowska 7,8,9; the LEAPfROG Consortium
PMCID: PMC11986813  PMID: 40226368

ABSTRACT

Background

Safe medication prescribing for hospitalized chronic kidney disease (CKD) patients is challenging. Leveraging electronic health records (EHRs) offers potential for decision support. A first step is to capture the CKD cohort through so called electronic phenotypes (e-phenotypes). However, available e-phenotypes, defined by logical rules applied to EHR data, lack consensus and are often inconsistently aligned with the Kidney Disease – Improving Global Outcomes (KDIGO) guideline for CKD (KDIGO-CKD). Therefore, local analyses and formalization efforts are essential to derive logical rules for CKD cohort selection.

Methods

We analyzed routinely collected EHR data from adults hospitalized at Amsterdam University Medical Centre (2018–23). Six logical rules were investigated: four derived from KDIGO-CKD (reduced glomerular filtration rate, albuminuria, kidney replacement therapy, and other markers of kidney damage) and two from published studies (diagnosis codes and medications).

Results

The study included 108 854 hospitalized patients. Extensive efforts were needed to formalize the clinical CKD definition from KDIGO-CKD and adapt it to EHR data, including selecting appropriate CKD diagnosis codes, medications, and computable criteria. Pooling six logical rules resulted in identifying 17 805 hospitalized CKD patients (16.4%), showcasing varying CKD patient counts per rule (with proportions ranging from 2.1% to 8.4%). Nonetheless, baseline characteristics across cohorts were comparable. Over one-third of patients identified by decreased eGFR or albuminuria/proteinuria measurements lacked a corresponding diagnosis code.

Conclusions

Deriving and formalizing six logical rules required close collaboration between nephrologists, EHR data experts, and medical informaticians. Our study provides groundwork towards a computer-interpretable CKD definition to standardize cohort capture in EHR-based studies.

Keywords: chronic kidney disease, clinical practice guideline, electronic health record, epidemiology, nephrology

Graphical Abstract

Graphical Abstract.

Graphical Abstract


KEY LEARNING POINTS.

What was known:

  • Various logical rules have been proposed to identify chronic kidney disease (CKD) patients using electronic health record (EHR) data, resulting in highly variable CKD counts.

  • None of these rules has achieved widespread acceptance, and their alignment with the prevailing CKD definition is inconsistent.

  • Informed decisions on CKD cohort selection require local analyses and rule formalization efforts.

This study adds:

  • Formalizing the KDIGO-CKD definition into logical rules and selecting CKD diagnosis codes and medications required extensive efforts.

  • Substantial variability in CKD counts arises from the logical rules used, albeit resulting cohorts show comparable baseline characteristics.

  • Cohort overlap differences underscore the need to harness diverse data elements to identify CKD patients.

Potential impact:

  • Our findings highlight the need to formalize the CKD clinical definition into a computer-interpretable one, to avoid inconsistencies in EHR implementation across studies.

  • Such efforts require close collaboration between nephrologists, EHR data experts and medical informaticians.

  • The results provide insights for further improvements in EHR-based CKD cohort selection.

INTRODUCTION

Chronic kidney disease (CKD) is an emerging global public health issue [1, 2]. CKD patients often present complex and multi-morbid conditions [1, 3], but they are frequently excluded from clinical trials [4, 5], which limits the evidence base for safe prescribing [6, 7]. To bridge this gap, leveraging electronic health record (EHR) data is a promising strategy [8–10]. Accurate identification of CKD patients using EHR data is an essential first step towards this goal [11]. CKD cohort selection can be achieved by using the so-called electronic phenotypes or e-phenotypes [12], consisting of logical rules derived from clinical practice guidelines [12, 13]. In this vein, recent U.S. Food and Drug Administration guidance emphasizes the importance of precise operational definitions to capture the intended patient cohorts when using real-world data [13].

The Kidney Disease – Improving Global Outcomes (KDIGO) guideline for CKD (hereinafter referred to as KDIGO-CKD) has established a clinical definition for CKD [14]. Using KDIGO-CKD as a starting point, multiple CKD e-phenotypes have been developed and validated for hospital settings, each employing different operational definitions to capture a CKD cohort [15–18]. Indeed, the varied practical implementations in EHR data have brought to the fore the challenges and inconsistencies in such captures [19–21]. Adding to this challenge, there is currently no widely accepted or endorsed CKD e-phenotype to standardize such efforts. This variability has contributed to notable discrepancies in CKD patient population estimates and characterizations, even within similar healthcare environments [22–24].

Therefore, urged by the need to capture a CKD cohort to study medication safety using EHR data, we conducted in-depth analyses using EHR data from our organization. For this purpose, we used the KDIGO-CKD guideline [14] and previous CKD e-phenotypes [15–18, 25, 26] to derive and formalize a set of logical rules and criteria to be able to make better-informed decisions on EHR-based CKD cohort selection.

MATERIALS AND METHODS

This study is reported in alignment with the REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) (Supplementary Table S1).

Setting and study design

This is a cross-sectional descriptive study conducted at Amsterdam University Medical Centre (Amsterdam UMC) in the Netherlands.

Participants

All adult patients with at least one hospitalization of ≥24 h in Amsterdam UMC between 1 January 2018 and 31 December 2023 were included. Patients with multiple hospitalization episodes were assigned the same patient ID.

Data sources and data access

Pseudonymized routinely collected EHR data, extracted retrospectively by authorized research data management staff on 30 June 2023 from the Epic® EHR System were made available for this study. For the included patients, the whole hospital trajectory between the above-specified dates was fully traceable. In addition, EHR data from Amsterdam UMC outpatient specialist clinics included patient data between 1 January 2015 and 31 December 2023.

Study size

No formal power calculations are applicable.

Data elements and logical rules exploration

Since the KDIGO-CKD guideline [14] necessitates translating the clinical criteria into logical rules for defining CKD for computer frameworks, we collaborated closely with nephrologists and EHR data experts to adapt the clinical definition for EHR research. Additionally, we introduced logical rules commonly used for EHR-based CKD cohort selection [15, 16, 23, 25, 26]. After approval from three nephrologists (L.B.H., L.V., and R.H.G.O.E.), we formulated six logical rules as outlined in Table 1.

Table 1:

Definition of logical rules derived from KDIGO-CKD and studies on CKD phenotyping and formalization steps applied.

No. Name Definition Formalization steps applied Data elements used
1 KDIGO-CKD Decreased eGFR logical rule Based on the most recent eGFR <60 mL/min/1.73 m2 and another reduced eGFR measurement at least 90 days apart (i.e. chronicity criterion). eGFR measurements within the 90-day time window should remain decreased, that is, <60 mL/min/1.73 m2 (i.e. persistence criterion). Since non-steady-state eGFR is frequent in hospital settings, eGFR laboratory measurements ±31 days from registration of an acute episode (Supplementary Table S3) or dialysis encounters were discarded (i.e. acute disease episode unconfoundedness criterion). Given the lack of guidance in KDIGO-CKD on how to formalize the acute disease episode unconfoundedness criterion into a computer-interpretable one, we relied on previous studies for such formalization [18]. Furthermore, KDIGO-CKD does not provide guidance on which dialysis procedures and ICD-10 codes should be used to identify acute episodes or dialysis encounters. Therefore, we defined suitable ICD-10 codes. Additionally, conditions upon which relaxation of the 90-day time window for eGFR measurements should be considered are lacking, too. Therefore, we also analyzed the impact of threshold modifications (Table 2). Laboratory measurements, ICD-10 diagnosis codes, procedure codes, dialysis measurements
2 KDIGO-CKD Albuminuria logical rule Based on the most recent albuminuria/proteinuria measurement like UACR ≥3 mg/mmol, AER ≥30 mg/24 h or equivalent with another albuminuria/proteinuria qualifying measurement at least 90 days apart (i.e. chronicity criterion). The increased measurements should remain within the 90-day time window (i.e. persistence criterion). See Supplementary Table S4 for more details on KDIGO-CKD equivalences for albuminuria/proteinuria laboratory measurements. KDIGO-CKD does not provide conditions upon which relaxation of the 90-day time window for albuminuria/proteinuria measurements should be considered. Therefore, we also analyzed the impact of threshold modifications (Table 2). Laboratory measurements
3 KDIGO-CKD KRT logical rule Includes patients with history of kidney transplantation or dependent on dialysis. Given the lack of guidance in KDIGO-CKD on how to formalize dependence on dialysis (also known as maintenance dialysis) into a computer-interpretable criterion, we defined it as two dialysis encounters at least 90 days apart with at least two dialysis encounters per week (i.e. dialysis measurements, procedure codes or ICD-10 codes) or the equivalent total (i.e. 25 encounters in any 90-day time window) or a registration of an ICD-10 code pertaining to maintenance dialysis. KDIGO-CKD does not provide guidance on which procedure or ICD-10 codes should be used to detect chronic dialysis. Therefore, we defined suitable ICD-10 codes. ICD-10 diagnosis codes, procedure codes, dialysis measurements
4 KDIGO-CKD Other Markers of Kidney Damage logical rule Includes conditions specified in KDIGO-CKD that are indicative of CKD when present for at least 90 days, namely, urinary sediment abnormalities, renal tubular disorders, pathologic abnormalities detected by histology or inferred or structural abnormalities detected by imaging. For urine sediment abnormalities, two abnormal measurements had to be recorded at least 90 days apart. See Supplementary Table S5 for more details on the urine sediment abnormality laboratory measurements and their corresponding thresholds. KDIGO-CKD does not provide guidance on which procedure or ICD-10 codes should be used to detect other markers of kidney damage. Therefore, we defined suitable ICD-10 codes and validated them with: (i) an eGFR <60 mL/min/1.73 m2 measurement; (ii) UACR ≥30 mg/g, AER ≥30 mg/24 h or equivalent, or (iii) registration of the same ICD-10 code at least 90 days apart. ICD-10 codes, laboratory measurements
5 Certain CKD Diagnosis Codes logical rule Utilizes only ICD-10 codes that specifically mention presence of CKD or pertain to structural abnormalities or pathologic abnormalities that are congenital (i.e. polycystic kidneys and dysplastic kidneys) as per previous studies on CKD phenotyping [15–18, 23, 26]. KDIGO-CKD does not provide guidance on which ICD-10 codes should be used to detect CKD diagnoses and their certainty in terms of actual CKD. Therefore, we collected suitable ICD-10 codes from literature and assessed which can be viewed as certain CKD diagnosis codes. ICD-10 diagnosis codes
6 CKD Medications logical rule Utilizes the prescription of medications used in advanced stages of CKD combined with at least one measurement of eGFR <60 mL/min/1.73 m2 or UACR ≥30 mg/g, AER ≥30 mg/24 h or equivalent at least 90 days apart from the first prescription of one of the selected medications (Supplementary Table S6), as per previous studies on CKD phenotyping [26, 27]. KDIGO-CKD does not provide guidance on which medications indicative of CKD should be considered for e-phenotypes. Therefore, we identified suitable medications based on Dutch prescribing guidelines. ATC codes, laboratory measurements

AER, albumin extraction rate; ATC, anatomical therapeutic chemical.

These six logical rules were applied separately to assess the impact on CKD patient counts and patient characteristics and were finally pooled into one computational implementation to provide an overview of the potential CKD cohort in our institution. For the pooled CKD cohort, the index date corresponds with the earliest date of CKD diagnosis. See Fig. 1 for a diagram of the study design.

Figure 1:

Figure 1:

Multi-panel diagram illustrating descriptive study design choices. (A) Longitudinal illustration of CKD patient trajectories and identification. (B) Baseline data collection timeframes. BMI, body mass index.

Variables and definitions

  • Index date. The earliest date of CKD diagnosis as detected by one of the six logical rules (Supplementary Table S2). Since our logical rules are purposely not formally validated, the CKD diagnoses should be interpreted as potential CKD diagnoses.

  • CKD diagnosis codes. To identify CKD for logical rules 1, 3, 4, and 5, ICD-10 Dutch codes were used mapped from the literature [15, 16, 18, 28–30]. The codes deemed certain were reviewed by three nephrologists (L.B.H., L.V., and R.H.G.O.E.).

  • Patient characteristics at baseline. The demographic variables extracted were age, sex, and body mass index. Comorbidities were extracted using Dutch ICD-10 codes (Supplementary Table S7). The most recent Charlson comordibidity index (CCI) from index date was estimated. Healthcare utilization metrics were described as number of hospitalizations, number of healthcare contact days 1 year prior to index date, and the top five medical specialities at index date. The number of active prescriptions at index date was also reported. Electrolytes and biochemical measurements were extracted using internal laboratory measurement codes during the 3-year period prior to index date.

  • eGFR. eGFR was estimated using the Chronic Kidney Disease Epidemiology Collaboration creatinine (CKD-EPI) 2009 equation without correction for race [31].

  • Albuminuria/proteinuria. Albuminuria was identified using urine albumin–creatinine ratio (UACR), albumin excretion rate (AER), and albumin dipstick, whilst proteinuria was identified using urine protein–creatinine ratio (UPCR), protein excretion rate (PER), and protein dipstick. Albuminuria/proteinuria categories were assigned using the conversion table provided by KDIGO-CKD (Supplementary Table S4). Prioritization of albuminuria/proteinuria measurement type was done as per KDIGO-CKD [32].

  • Kidney replacement therapy identification. Kidney transplantation was identified via registered ICD-10 Dutch codes and procedure codes. Dialysis treatment was identified with ICD-10 Dutch codes, procedure codes or dialysis encounters.

  • Other markers of kidney damage. Identified via ICD-10 Dutch diagnosis codes that indicate ‘possible’ CKD, directly derived from the literature and further validated as specified under logical rule 4 (see Table 1).

  • Medication use. Prescription of phosphate-binding drugs, selected vitamin D receptor agonists, cation exchangers, or specific dose regimens of erythropoiesis-stimulating agents highly indicative of CKD according to Dutch prescribing guidelines (Supplementary Table S6).

  • CKD staging. G and A stages were assigned based on KDIGO-CKD categorization using the last qualifying eGFR or albuminuria/proteinuria measurements [19, 21]. However, eGFR measurements lying within ±31 days from an acute disease episode were discarded (i.e. as per the ‘acute disease episode unconfoundedness’ criterion). People on maintenance dialysis were excluded from G and A staging.

Statistical methods and sensitivity analyses

Analyses are mainly descriptive in nature. CKD counts are presented as absolute and relative frequencies out of the total number of included patients. Categorical data are summarized by counts and percentages and continuous variables by using mean and standard deviation or median and quartiles for non-normally distributed data. We also illustrate patient distribution across logical rules with a flowchart, along with an UpSet plot to assess their overlap and concordance patterns.

With regard to missing data, comorbidities identified in the medical history without an associated start date were assigned the registration date.

We assess the consequences of altering chronicity and persistence criteria on KDIGO-CKD Decreased eGFR and KDIGO-CKD albuminuria logical rules. In addition, we also explore the impact of altering the related acute disease episode unconfoundedness criterion in the KDIGO-CKD Decreased eGFR logical rule. These sensitivity analyses with criteria are described in more detail in Table 2.

Table 2:

Sensitivity analysis with laboratory measurements to explore consequences of altering existing thresholds in three criteria.

Criterion Definition Applicable logical rules Bounds under study
KDIGO-CKD chronicity criterion Time window between two laboratory qualifying measurements KDIGO-CKD Decreased eGFR;
KDIGO-CKD Albuminuria
30–120 days
KDIGO-CKD persistence criterion Measurements should be below a certain threshold between the two laboratory qualifying measurements (i.e. eGFR <60 mL/min/1.73 m2 and A stage >1) KDIGO-CKD Decreased eGFR;
KDIGO-CKD Albuminuria
0–100 eGFR spikes;
0–15 albuminuria troughs
Acute disease episode unconfoundedness criterion Timespan from ICD-10 code registration of an acute disease episode or short-term dialysis where eGFR measurements are discarded KDIGO-CKD Decreased eGFR ±7 to ±70 days

Bias

To minimize bias, we implemented several steps. Logical rules, including decreased eGFR and albuminuria, incorporated various KDIGO-CKD derived criteria. For maintenance dialysis, we required at least two dialysis encounters per week to avoid acute dialysis encounters. We also applied the chronicity criterion with eGFR or albuminuria measurements to validate the markers of kidney damage logical rule. In the medications logical rule, we tailored dosing regimens specifically for CKD patients and validated them with laboratory measurements. To prevent measurement bias in logical rules using laboratory measurements, we excluded measurements that were not physiologically plausible.

RESULTS

Formalization efforts

For all six logical rules extensive formalization efforts and assumptions were needed to enable leveraging EHR data (Table 1), which required close collaboration between medical informaticians, nephrologists, and EHR data experts. These efforts included identification of suitable ICD-10 codes (see Supplementary data, Excel file), medications, defining additional computable criteria (e.g. how maintenance dialysis is defined), additional time windows (e.g. number of days to discard around acute disease episodes), and validation criteria (e.g. for kidney tubular disorders or pathological abnormalities diagnosis codes).

Chronic kidney disease patient identification workflow

Out of 108 854 included patients, the implementation of the six logical rules and their pooling resulted in 17 805 CKD patients (16.4%). The highest proportion was obtained with the certain CKD diagnosis codes logical rule (n = 9121; 8.4%) and the lowest using medications (n = 2467; 2.3%), and kidney replacement therapy (KRT) (n = 2277; 2.1%) logical rules (Fig. 2).

Figure 2:

Figure 2:

Flow diagram to arrive at the final CKD patient counts using logical rules and criteria. Validation in the KDIGO-CKD Other Markers of Kidney Damage is colored with the data element (i.e. eGFR, albuminuria/proteinuria measurements or ICD-10 other markers of kidney damage diagnosis codes) from the corresponding logical rule.

After applying the chronicity criterion, CKD counts using KDIGO-CKD Decreased eGFR and KDIGO-CKD Albuminuria logical rules were reduced by factors of 1.9 and 3.8, respectively, while the KDIGO-CKD KRT logical rule led to a 4.4-fold reduction. For the KDIGO-CKD Other Markers of Kidney Damage logical rule, there was a 6.5- and 1.6-fold decrease in the urine sediments and ICD-10 branches, respectively. The persistence criterion also had a substantial impact on CKD counts in the KDIGO-CKD KRT logical rule, leading to a 2.5-fold decrease.

Patient characteristics at index date across cohorts

The proportion of females ranged from 39.4% to 43.8% and the median age between 58 and 73 years across all cohorts (Table 3). Concerning comorbidities, ≥75% of CKD patients had moderate to high CCI. The majority of CKD patients across cohorts had more than five active prescriptions. The Medications and KDIGO-CKD Albuminuria cohorts displayed the highest healthcare utilization, and KDIGO-CKD KRT and KDIGO-CKD Other Markers of Kidney Damage cohorts showed the lowest.

Table 3:

Demographic characteristics, comorbidities, healthcare utilization and medical speciality at baseline for CKD patients identified with six different logical rules.

  Logical rule
Patient characteristic KDIGO-CKD Decreased eGFR cohort KDIGO-CKD Albuminuria cohort KDIGO-CKD KRT cohort KDIGO-CKD Other Markers of Kidney Damage cohort Certain CKD Diagnosis Codes cohort Medications cohort All
Total, N (%) 8021 (45.0) 4465 (25.1) 2277 (12.8) 7620 (42.8) 9121 (51.2) 2467 (13.9) 17 805 (100.0)
Female, n (%) 3461 (43.1) 1846 (41.3) 926 (40.7) 3326 (43.6) 3874 (42.5) 973 (39.4) 7937 (44.6)
Age at index date, median (Q1, Q3) 73 (64, 80) 68 (53, 76) 58 (47, 68) 63 (51, 73) 69 (57, 77) 64 (50, 72) 68 (56, 77)
BMI, kg m−2,
median (Q1, Q3)
missing, n (%)

26.0 (23.0, 29.5)
3175 (39.6)

25.5 (22.2, 29.4)
1350 (30.2)

25.9 (22.8, 29.4)
126 (5.5)

26.0 (23.0, 29.7)
415 (5.5)

25.9 (23.0, 29.6)
1170 (12.8)

25.9 (22.8, 29.6)
136 (5.5)

25.9 (22.9, 29.6)
3123 (17.5)
CCI, n (%)
None (CCI = 0) 414 (5.2) 488 (10.9) 449 (19.7) 1144 (15.0) 1015 (11.1) 331 (13.4) 2251 (12.6)
Moderate (CCI = 1–2) 818 (10.2) 634 (14.2) 746 (32.8) 1895 (24.9) 1813 (19.9) 543 (22.0) 3542 (19.9)
Severe (CCI ≥ 3) 6789 (84.6) 3343 (74.9) 1082 (47.5) 4581 (60.1) 6293 (69.0) 1593 (64.6) 12 012 (67.5)
Comorbidities, n (%)
All-site site cancer
Atrial fibrillation/flutter
Depression
HF
HIV
Hypertension
Liver disease
MI
PAD
Pulmonary disease
Stroke
T1DM
T2DM

3109 (38.8)
2410 (30.0)
261 (3.3)
2427 (30.3)
119 (1.5)
5160 (64.3)
593 (7.4)
1702 (21.2)
1,427 (17.8)
1422 (17.7)
1121 (14.0)
131 (1.6)
2430 (30.3)

1721 (38.5)
1015 (22.7)
183 (4.1)
982 (22.0)
103 (2.3)
2613 (58.5)
415 (9.3)
716 (16.0)
765 (17.1)
784 (17.6)
668 (15.0)
116 (2.6)
1461 (32.7)

521 (22.9)
421 (18.5)
78 (3.4)
434 (19.1)
32 (1.4)
1661 (72.9)
206 (9.0)
339 (14.9)
493 (21.7)
260 (11.4)
244 (10.7)
65 (2.9)
766 (33.6)

3221 (42.3)
1628 (21.4)
283 (3.7)
1532 (20.1)
109 (1.4)
4079 (53.5)
674 (8.8)
1148 (15.1)
1119 (14.7)
1221 (16.0)
870 (11.4)
129 (1.7)
1984 (26.0)

3162 (34.7)
2468 (27.1)
291 (3.2)
2565 (28.1)
99 (1.1)
5479 (60.1)
661 (7.2)
1806 (19.8)
1528 (16.8)
1567 (17.2)
1145 (12.6)
141 (1.5)
2749 (30.1)

832 (33.7)
611 (24.8)
91 (3.7)
710 (28.8)
35 (1.4)
1734 (70.3)
291 (11.8)
448 (18.2)
536 (21.7)
386 (15.6)
303 (12.3)
67 (2.7)
879 (35.6)

6777 (38.1)
4374 (24.6)
584 (3.3)
4115 (23.1)
218 (1.2)
9521 (53.5)
1312 (7.4)
3042 (17.1)
2479 (13.9)
2981 (16.7)
2161 (12.1)
279 (1.6)
4594 (25.8)
Number of healthcare contact days
median (Q1, Q3)

18.0 (9.0, 37.0)

24.0 (11.0, 48.0)

5.0 (2.0, 15.0)

5.0 (2.0, 14.0)

8.0 (3.0, 21.0)

18.0 (4.0, 39.0)

7.0 (2.0, 19.0)
Number of hospitalizations, n (%)
0
1
2 or 3
>3

0 (0.0)
2034 (25.4)
1831 (22.8)
3001 (37.4)

0 (0.0)
771 (17.3)
784 (17.6)
2274 (50.9)

0 (0.0)
712 (31.3)
387 (17.0)
845 (37.1)

0 (0.0)
2088 (27.4)
1455 (19.1)
3052 (40.1)

0 (0.0)
3430 (37.6)
1726 (18.9)
2835 (31.1)

0 (0.0)
488 (19.8)
386 (15.6)
1258 (51.0)

0 (0.0)
6282 (35.3)
3654 (20.5)
5515 (31.0)
Number of active prescriptions, n (%)
No medications
One to five
More than five

1508 (18.8)
3100 (38.6)
3413 (42.6)

398 (8.9)
1273 (28.5)
2794 (62.6)

102 (4.5)
314 (13.8)
1861 (81.7)

867 (11.4)
2299 (30.2)
4454 (58.5)

1174 (12.9)
2249 (24.7)
5698 (62.5)

24 (1.0)
333 (13.5)
2110 (85.5)

2584 (14.5)
5431 (30.5)
9790 (55.0)
Medical speciality at index date, n (%)
Internal medicine, General
Nephrology
Surgery
Cardiology
Urology
Any other speciality
988 (12.3)
720 (9.0)
1234 (15.4)
1517 (18.9)
641 (8.0)
2921 (36.4)
801 (17.9)
409 (9.2)
550 (12.3)
379 (8.5)
419 (9.4)
2558 (42.7)
321 (14.1)
914 (40.1)
451 (19.8)
130 (5.7)
85 (3.7)
376 (16.5)
1075 (14.1)
742 (9.7)
908 (11.9)
707 (9.3)
1562 (20.5)
2626 (34.5)
1179 (12.9)
925 (10.1)
1539 (17.1)
1521 (16.9)
816 (9.0)
3041 (33.3)
409 (16.6)
637 (25.8)
408 (16.7)
201 (8.2)
122 (5.0)
690 (28.0)
2108 (11.8)
1141 (6.4)
2668 (15.0)
2554 (14.3)
2030 (11.4)
7304 (41.0)

HF, heart failure; HIV, Human Immunodeficiency Virus; MI, myocardial infarction; PAD, peripheral artery disease; Q, quartile; T1DM, type 1 diabetes mellitus; T2DM, type 2 diabetes mellitus.

In general, stages G3 and A3 were most frequent (Table 4). Serum electrolytes and biochemical measurement summary statistics were comparable across cohorts. Overall, the KDIGO-CKD KRT and Medications logical rules had the highest proportions of abnormal serum electrolytes and biochemical measurements.

Table 4:

Staging, serum electrolytes and biochemical measurements at baseline for chronic kidney disease patients identified with six different logical rules.

  Logical rule  
Patient characteristic KDIGO-CKD Decreased eGFR cohort KDIGO-CKD Albuminuria cohort KDIGO-CKD KRT cohorta KDIGO-CKD Other Markers of Kidney Damage cohort Certain CKD Diagnosis Codes cohort Medications cohort All
Total, N (%) 8021 (45.0) 4465 (25.1) 2277 (12.8) 7620 (42.8) 9121 (51.2) 2467 (13.9) 17 805 (100.0)
Stages,b  n (%)
G stage
Stage G1
Stage G2
Stage G3
Stage G4
Stage G5
Unstaged
A stage
Stage A1
Stage A2
Stage A3
Unstaged
On maintenance dialysis


N/A
N/A
5,714 (71.2)
1477 (18.4)
542 (6.8)
N/A

2296 (28.6)
1195 (14.9)
2684 (33.5)
1558 (19.4)
288 (3.6)


811 (18.2)
905 (20.3)
1,076 (24.1)
651 (14.6)
421 (9.4)
485 (10.9)

N/A
1636 (36.6)
2829 (63.4)
N/A
116 (2.6)


40 (1.8)
157 (6.9)
491 (21.6)
398 (17.5)
612 (26.9)
145 (6.4)

664 (29.2)
328 (14.4)
768 (33.7)
83 (3.6)
434 (19.1)


1,259 (16.5)
1,718 (22.5)
2029 (26.6)
925 (12.5)
663 (8.7)
823 (10.8)

2553 (33.5)
1,237 (16.2)
2919 (38.3)
708 (9.3)
203 (2.7)


647 (7.1)
1193 (13.1)
3115 (34.2)
1831 (20.1)
1112 (12.2)
825 (9.0)

2503 (27.4)
1,273 (14.0)
3296 (36.1)
1651 (18.1)
398 (4.4)


173 (7.0)
235 (9.5)
508 (20.6)
562 (22.8)
651 (26.4)
160 (6.5)

583 (23.6)
445 (18.0)
1260 (51.1)
1 (<0.1)
178 (7.2)


2363 (13.3)
3490 (19.6)
5959 (33.5)
2216 (12.4)
1233 (6.9)
2110 (11.9)

5045 (28.3)
2820 (15.8)
6233 (35.0)
3273 (18.4)
434 (2.4)
Serum electrolytes and biochemical measurements at baselinec
Bicarbonate, mmol L−1
Mean (SD)
<21, n (%)
Missingness, n (%)
Calcium, mmol L−1
Mean (SD)
<2.1, n (%)
Missingness (%)
Haemoglobin, mmol L−1
Mean (SD)
<7, n (%)
Missingness, n (%)
Phosphate, mmol L−1
Mean (SD)
≥1.5, n (%)
Missing, n (%)
Potassium, mmol L−1
Mean (SD)
≥5.5, n (%)
Missingness, n (%)
Cystatin C, mg L−1
Mean (SD)
Missingness, n (%)



23.4 (4.1)
1207 (15.0)
2843 (35.4)

2.4 (0.2)
410 (5.1)
2208 (27.5)

7.8 (1.3)
1870 (23.3)
249 (3.1)

1.2 (0.4)
700 (8.7)
3140 (39.1)

4.4 (0.6)
343 (4.3)
392 (4.9)

2.1 (1.1)
7723 (96.3)



23.4 (4.5)
831 (18.6)
1005 (22.5)

2.3 (0.2)
310 (6.9)
767 (17.2)

7.8 (1.4)
1125 (25.2)
41 (0.9)

1.1 (0.4)
432 (9.7)
1240 (27.8)

4.3 (0.6)
178 (4.0)
170 (3.8)

2.0 (1.2)
4222 (94.6)



23.0 (3.9)
407 (17.9)
699 (30.7)

2.3 (0.2)
160 (7.0)
609 (26.7)

7.3 (1.2)
632 (27.8)
531 (23.3)

1.3 (0.5)
449 (19.7)
636 (27.9)

4.6 (0.8)
213 (9.4)
554 (24.3)

2.6 (1.4)
2239 (98.3)



23.8 (4.2)
530 (7.0)
4922 (64.6)

2.3 (0.2)
282 (3.7)
3893 (51.1)

8.0 (1.3)
1124 (14.8)
2243 (29.4)

1.1 (0.4)
333 (4.4)
4806 (63.1)

4.3 (0.6)
181 (2.4)
2577 (33.8)

1.0 (1.1)
7553 (99.1)



23.2 (4.0)
274 (3.0)
7990 (87.6)

2.3 (0.2)
122 (1.3)
7799 (85.5)

7.6 (1.3)
484 (5.3)
7462 (81.8)

1.2 (0.5)
197 (2.2)
7929 (86.9)

4.5 (0.7)
122 (1.3)
7507 (82.3)

2.0 (1.4)
9079 (99.5)



22.8 (4.5)
559 (22.7)
557 (22.6)

2.3 (0.2)
208 (8.4)
456 (18.5)

7.3 (1.3)
810 (32.8)
328 (13.3)

1.3 (0.5)
392 (15.9)
552 (22.4)

4.5 (0.7)
218 (8.8)
344 (13.9)

2.0 (1.4)
2404 (97.4)



23.9 (4.2)
1155 (6.5)
11 759 (66.0)

2.3 (0.2)
582 (3.3)
9767 (54.9)

8.0 (1.3)
2421 (13.6)
5255 (29.5)

1.1 (0.4)
600 (3.4)
11 860 (66.6)

4.3 (0.6)
328 (1.8)
6021 (33.8)

1.1 (1.0)
17 643 (99.1)
a

Applies to post-transplant patients only.

b

G stage and A stage reported based on the last available measurement, discarding eGFR measurements ±31 days from an acute disease episode. Patients on maintenance dialysis were not staged. Patients who received a kidney transplant are staged after the transplant procedure.

c

Percentages for concentration cut-off subpopulations are specified out of each cohort. Note that when all CKD patients are pooled, the earliest CKD diagnosis date is taken as the index date and therefore missingness patterns for laboratory measurements would be affected.

N/A, not applicable; Q, quartile; PTH, parathyroid hormone.

Analysis of overlap between logical rules

The six logical rules resulted in cohorts with high overlap ranging from 67.0% to 93.7% (Fig. 3). The Medications logical rule and KDIGO-CKD KRT logical rules showed the highest overlap with others, both at 93.7%, whilst the Certain CKD Diagnosis Codes logical rule identified the largest proportion of non-overlapping CKD patients (36.0%). Notably, 2725 (34.0%) CKD patients identified by the KDIGO-CKD Decreased eGFR logical rule and 1585 (35.5%) identified by the KDIGO-CKD Albuminuria logical rule lacked a certain CKD diagnosis code. Additionally, 4387 (48.1%) CKD patients identified with the Certain CKD Diagnosis Codes logical rule were not captured by any laboratory-based logical rule. A total of 3392 (45.5%) CKD patients identified with the KDIGO-CKD Other Markers of Kidney Damage logical rule were also missed by laboratory-based ones. Additional analyses can be found in Supplementary Fig. S3. Lastly, as CKD is primarily diagnosed based on eGFR or albuminuria measurements as per KDIGO-CKD, we additionally analyzed the degree of non-overlap (i.e. missingness) of KDIGO-CKD eGFR and KDIGO-CKD Albuminuria logical rules with all logical rules (Fig. 3): (i) KDIGO-CKD Decreased eGFR rule: in 6020 (75.1%) patients missing albuminuria; (ii) KDIGO-CKD Albuminuria rule: in 2464 (55.2%) missing eGFR; (iii) KDIGO-CKD KRT rule (transplant only): in 689 (30.3%) patients missing eGFR and in 1165 (51.2%) patients missing albuminuria; (iv) KDIGO-CKD Other Markers of Kidney Damage rule: in 4288 (65.3%) patients missing eGFR and in 5366 (70.4%) patients missing albuminuria; (v) Certain CKD Diagnosis Codes rule: in 4870 (53.4%) patients missing eGFR and in 7177 (78.0%) patients missing albuminuria; (vi) Medication rule: in 764 (31.0%) patients missing eGFR and in 1241 (50.3%) patients missing albuminuria.

Figure 3:

Figure 3:

UpSet plot for assessing overlap and concordance (intersection) patterns between logical rules. The UpSet plot is a visual way of representing a multidimensional Venn diagram and consists of three parts: a bar plot (top), a dot and segments heat map (bottom), and a horizontal bar plot (left). In the bar plot (top), the x-axis shows the unique intersection patterns between logic rules and the y-axis represents the absolute frequency of intersection patterns. The bars and dots are colored according to the number of intersections (overlap), i.e. coral for one intersection (CKD patients exclusively identified by that one logic rule) and teal for more than one intersection (CKD patients who are identified by more than one logic rule). When interpreting the graph, the bar plot on the left-hand side provides information about the total counts, whilst the bar plot at the top provides information about the frequency of patterns. Beneath the graph, there is a table with statistics for each logical rule. The data elements used in each logical rule are indicated in brackets. The statistics shown are the total unique CKD patient counts, total unique CKD patient counts non-overlapping with other logic rules, total unique CKD patient counts overlapping with other logic rules, total unique CKD patient counts overlapping with logical rules using laboratory values only (i.e. KDIGO-CKD Decreased eGFR and KDIGO-CKD Albuminuria), total unique CKD patient counts overlapping with logical rules using ICD-10 codes only (i.e. KDIGO-CKD Other Markers of Kidney Damage, Certain CKD Diagnosis Codes) and the logical rule with which there is most overlap. Notably, the Certain CKD Diagnosis Codes and the KDIGO-CKD KRT logical rules are mutually exclusive, as patients were programmatically removed from the former logical rule if they were also identified by the latter logical rule. ATC, anatomical therapeutic chemical; N/A, not applicable.

Sensitivity analyses

Imposing the 90-day chronicity window resulted in 12.2% and 15.8% decrease in counts for the KDIGO-CKD Decreased eGFR and KDIGO-CKD Albuminuria logical rules, respectively, when compared with a 30-day window (Supplementary Fig. S4). For the persistence criterion, the greatest count rise was noted after allowing a single eGFR spike, leading to an increase from 8382 (i.e. no spikes within 90-day window) to 8944 CKD patients (+6.7%) (Supplementary Fig. S5). For albuminuria/proteinuria persistence criterion, a change from 4465 CKD patients with no A-stage troughs to 5090 CKD patients (+14.0%) was noted (Supplementary Fig. S5). For the acute disease episode unconfoundedness criterion, 8193 CKD patients were identified using a ±7-day window and 8021 CKD patients using a ±31-day window, representing a 2.1% decrease in counts (Supplementary Fig. S6).

Electronic health record data quality analysis

The application of logical rules is impacted by EHR data quality. Patient demographic data lacked race/ethnicity categories. The absence of internationally accepted clinical terminology, like LOINC, necessitated using multiple local codes to extract similar laboratory measurements. Additionally, lack of code granularity affected the discernment between CKD stages (e.g. G3a vs G3b) and identifying specific conditions like Alport syndrome, requiring manual verification in the problem list. Additionally, medical diagnoses can be registered in various places in an EHR system, leading to discrepancies that need to be resolved.

DISCUSSION

Extensive efforts and a substantial number of assumptions were needed to formalize the CKD definition as specified in KDIGO-CKD into a set of computer-interpretable logical rules. Furthermore, the rules derived from existing CKD e-phenotypes required additional formalization efforts to align with the Dutch healthcare setting and the assessment of certainty for a set of ICD-10 codes. Implementation of the derived six logical rules on EHR data resulted in a cohort of 17 805 CKD patients (16.4%). We observed substantial variability in identified CKD proportions across logical rules (2.1%–8.4%), although baseline characteristics were broadly similar. Overlap analysis revealed variations in the percentage of patients commonly identified by the logical rules (67.0%–93.7%). In addition, CKD counts were clearly impacted by criteria like chronicity or persistence, which to date are still applied inconsistently [19, 33, 34].

Previous studies have shown that CKD prevalence in hospital EHR data varies between 1.5% and 30.7% and our estimates fall within this range [22, 35, 36]. However, direct comparisons are challenging due to differences in settings, data elements, definitions, and logical rules. In our study, we validated ICD-10 codes and applied progressively stringent criteria to laboratory measurements to ensure alignment with KDIGO-CKD [14], revealing potential sources of variability between EHR-based studies. Consistent with a prior Danish study [23], we found notable variability in CKD counts depending on which criteria were applied to laboratory-based logical rules. Sensitivity analyses further emphasized the influence of different thresholds, underscoring the importance of harmonizing CKD definitions across epidemiological studies and implementing them consistently [19, 25, 37].

Beyond laboratory measurements, our study identified another overlooked source of variability, namely, the selection of diagnosis codes. For instance, the Certain CKD Diagnosis Codes logical rule, confirmed in this study by three nephrologists, and the KDIGO-CKD Other Markers of Kidney Damage logical rule, based on KDIGO-CKD and CKD phenotyping studies [15–18, 26], utilized distinct sets of ICD-10 codes. Comparing diagnosis codes across studies revealed limited consensus in diagnosis code selection [22, 38]. Therefore, we explored the impact of further validation for diagnosis codes indicative of structural abnormalities with laboratory measurements or an additional diagnosis code.

Ideally, ICD-10 codes should reflect CKD status based on eGFR or albuminuria/proteinuria findings. However, over one-third of patients identified via decreased eGFR or albuminuria/proteinuria measurements had no corresponding certain CKD ICD-10 code registered (Fig. 3), concordant with results from a Canadian study [28]. Issues with ICD-10 code registration, uncertainty around the parties responsible for maintaining the diagnosis list, and a limited perceived utility may contribute to this [39]. Conversely, over a half of CKD patients identified with the Certain CKD Diagnosis Codes logical rule lacked corresponding eGFR (53.4%) or albuminuria (78.0%) data (Fig. 3). This may be due to multiple factors, such as the academic nature of our hospital, where patients are often referred from other settings with a CKD diagnosis but may not have enough laboratory measurements within our institution to confirm it. Additionally, some patients may already have a diagnosis of kidney disease, but have not yet progressed to stage G3 or A2, such as those with polycystic kidney disease. Notably, relying exclusively on eGFR values to identify CKD patients would capture less than half of the actual CKD population in our hospital (n = 8021; 45.0%). This contrasts with previous reports, which suggest that ‘at least a fifth’ of the CKD population would be understudied, whereas our findings indicate a much higher proportion [19]. These observations further reinforce the idea that relying on a single data element can lead to an underestimation of CKD amongst hospitalized patients [35, 40].

Based on our findings, the Medication logical rule, which includes a decreased eGFR/albuminuria validation step, identifies relatively few additional CKD patients (Fig. 3; Supplementary  Fig. S3). Whilst the inclusion of the validation step ensures specificity by confirming CKD, its standalone utility may be limited due to the small number of new patients identified. However, this rule may still serve as a valuable additional layer of validation for CKD e-phenotypes, particularly in cases where medication use strongly indicates advanced stages of CKD [27, 41].

Key strengths of our study include a multidisciplinary approach to derive six logical rules, leveraging various EHR data elements, a comprehensive ICD-10 code selection informed by literature, and detailed cohort analyses, all serving as an opportunity to increase standardization for a computable CKD definition.

However, limitations exist, such as the absence of a formal clinical validation for the proposed logical rules, potentially leading to false positives. However, this was a deliberate decision given the objective of this study. Misclassification bias is also possible, for instance, in relation to the fact that we did not utilize cystatin C as a confirmatory test due to limited use [14, 42]. We also note that this was not a part of a systematic screening effort and therefore EHR registration of CKD may have been impacted by established practices in our setting. Finally, this study does not include general practitioners’ laboratory data as in other studies [23, 35, 36], given the lack of a suitable data infrastructure in the Netherlands, in contrast to a few other European countries [43]. Thus, it is hoped that initiatives like the European Health Data Space spearheaded by the European Commission will spur advances in data exchange [43].

Our findings carry relevant implications for clinicians, researchers, and guideline developers. Clinicians can use our findings to improve EHR data quality of the CKD patients, by resolving the discrepancies between laboratory findings and diagnosis code registration. For kidney (pharmaco)epidemiology researchers, our results indicate that relying solely on logical rules based on laboratory measurements and certain CKD diagnosis codes may not always adequately capture a CKD cohort [19, 23, 25], as outlined in the KDIGO-CKD recommendations [14]. Notably, KDIGO-CKD allows the derivation of additional logical rules [37], like KDIGO-CKD Other Markers of Kidney Damage, which considers structural and urine sediment abnormalities, and KDIGO-CKD KRT, both of which were introduced in this study, along with a medication-based rule. Whilst it is acknowledged that diagnosis codes may be better suited for detecting patients with structural abnormalities than laboratory measurements [19], efforts to standardize their identification through diagnosis codes have not yet reached maturity. Additionally, the underutilization of urine sediments as potential CKD indicators highlights a gap in current practices. Therefore, refining logical rules for CKD cohort capture using EHR data will require ongoing collaboration across fields to ensure accurate representation and optimal extraction [11]. Further discussions within the research and clinical communities are essential to achieve consistency and applicability in various research contexts.

Given the persisting discrepancies in CKD e-phenotypes, this study also highlights opportunities to enhance the CKD definition for computer frameworks. Some key steps may include a structured approach to designing crisp logical rules, a careful consideration for selecting the most appropriate diagnosis codes to identify CKD patients, and targeted efforts to emphasize the importance of ascertainment criteria like chronicity, persistence, how to handle acute disease episodes in inpatient settings, and any other computable validation requirements. In addition, it provides a basis for starting to engage in discussions on other aspects of CKD identification, such as an age-adjusted CKD definition and classification [44, 45] or biological fluctuations of laboratory value thresholds [46]. Indeed, these factors can impact the inclusion of groups at risk of progressing to CKD who are also candidates for medication safety studies. Ultimately, before designing and validating yet another CKD e-phenotype, achieving a broad consensus on a computer-interpretable CKD definition is essential to minimize redundancy, enhance comparability across studies, and ensure accurate identification and monitoring of hospitalized CKD patients using EHR data.

CONCLUSIONS

As more EHR-based (pharmaco)epidemiological analyses are being conducted, careful attention must be given to EHR-based CKD cohort capture. This study demonstrates proportions, overlap, and the sources of variability in CKD patient counts resulting from the application of six logical rules, whilst showing broadly comparable baseline characteristics across cohorts. Our study can help guide efforts in the direction of formalizing and standardizing a CKD definition for computer frameworks.

Supplementary Material

sfaf073_Supplemental_Files

ACKNOWLEDGEMENTS

The authors would like to thank Remko A. de Jong for his contribution to data curation and René Huijsman for his support in clarifying medical coding practices in our hospital. In addition, the authors would like to acknowledge the contribution from attendees at the Kidney Research Meeting in Amsterdam UMC held in October 2023 for giving their expert advice on the experimental design, and nephrologists Femke Waanders, MD PhD (Isala Hospital, Zwolle, the Netherlands) and Jacobien C. Verhave, MD PhD (Rijnstate Hospital, Arnhem, the Netherlands) for their suggestions on which drugs to consider for the Medication logical rule. LEAPfROG consortium: The LEAPfROG consortium which aims to leverage real-world data, machine learning and knowledge representation methods to improve pharmacotherapy outcomes in patients with multimorbidity, is a unique cross-sectional collaboration between patients, healthcare professionals, epidemiologists, medical informatics, and computer science researchers, as well as partners from the medical technology industry, regulatory bodies, and healthcare insurance. It aims to develop data-driven computerized decision support tools to support clinicians in data quality improvement, drug side-effect detection, and causal inference using data from electronic health records. The project focuses on clinically relevant use case of drug-induced kidney disease in people with chronic kidney disease. Principal investigator: Dr Joanna E. Klopotowska. Work package (WP) leaders: Dr Ronald Cornet (WP1), Prof. Dr Annette ten Teije (WP2), Dr Giovani Cina (WP3), and Dr Stephanie Medlock (WP4). Beneficiaries of the LEAPfROG consortium: Amsterdam University Medical Center, leadership and coordination; Vrije Universiteit Amsterdam, and Open University Heerlen. Ongoing LEAPfROG PhDs: Daniel Fernández-Llaneza, Romy M. P. Vos, and Joris E. Lieverse. The LEAPfROG Consortium members include: Ameen Abu-Hanna (Amsterdam University Medical Center), Annette ten Teije (Vrije Universiteit Amsterdam), Birgit A. Damoiseaux (Amsterdam University Medical Center), Cornelis Boersma (Open Universiteit), Dave A. Dongelmans (Nationale Intensive Care Evaluatie foundation), David H. de Koning (Amsterdam University Medical Center), Erol S. Hofmans (College ter Beoordeling van Geneesmiddelen), Evelien Tiggelaar (Z-Index), Frank van Harmelen (Vrije Universiteit Amsterdam), Gerty Holla (Amsterdam Economic Board), Heiralde Marck (Koninklijke Nederlandse Maatschappij ter bevordering der Pharmacie [KNMP] Geneesmiddelen Informatie Centrum), Iacopo Vagliano (Amsterdam University Medical Center), Jan Pander (AstraZeneca), Joanna E. Klopotowska (Amsterdam University Medical Center), Jurjen van der Schans (Open Universiteit), Joris E. Lieverse (Amsterdam University Medical Center), Kitty J. Jager (Amsterdam University Medical Center, European Renal Association-European Dialysis and Transplantation Association [ERA-EDTA] Registry), Linda Dusseljee-Peute (Amsterdam University Medical Center), Luuk B. Hilbrands (Radboud University Medical Centre), Marieke A. R. Bak (Amsterdam University Medical Center), Mariette van den Hoven (Amsterdam University Medical Centre), Martijn G. Kersloot (Castor), Menno Maris (Amsterdam University Medical Center), Nicolette F. de Keizer (Amsterdam University Medical Center), Otto R. Maarsingh (Amsterdam University Medical Center), Paul Blank (NWO), Piet Heingraaf (Pitts.AI), Renée de Wildt (Nierpatiënten Vereniging Nederland), Rosa M. P. Vos (Vrije Universiteit Amsterdam), Ron Herings (PHARMO Institute for Drug Outcomes Research & Amsterdam University Medical Center), Ron J. Keizer (InsightRX), Ronald Cornet (Amsterdam University Medical Centre), Ruben Boyd (IXA), Sebastiaan L. Knijnenburg (Castor), Sipke Visser (Digital Health Link), Stephanie Medlock (Amsterdam University Medical Centre), Stijn Gremmen (Dutch Kidney Foundation/Nierstichting), Teun van Gelder (Leiden University Medical Center), Tjerk S. Heijmens Visser (CZ Health Insurance, Zorgverzekeraars Nederland), and Vianda S. Stel (Amsterdam University Medical Center, ERA-EDTA Registry).

Contributor Information

Daniel Fernández-Llaneza, Department of Medical Informatics, Amsterdam University Medical Centre, University of Amsterdam, Amsterdam, the Netherlands; Amsterdam Public Health Institute, Digital Health, Amsterdam, the Netherlands; Amsterdam Public Health Institute, Methodology, Amsterdam, the Netherlands.

Luuk B Hilbrands, Department of Nephrology, Radboud University Medical Centre, Nijmegen, the Netherlands.

Liffert Vogt, Department of Internal Medicine and Nephrology, Amsterdam University Medical Centre, University of Amsterdam, Amsterdam, the Netherlands.

Rik H G Olde Engberink, Department of Internal Medicine and Nephrology, Amsterdam University Medical Centre, University of Amsterdam, Amsterdam, the Netherlands.

Joanna E Klopotowska, Department of Medical Informatics, Amsterdam University Medical Centre, University of Amsterdam, Amsterdam, the Netherlands; Amsterdam Public Health Institute, Digital Health, Amsterdam, the Netherlands; Amsterdam Public Health Institute, Quality of Care, Amsterdam, the Netherlands.

the LEAPfROG Consortium:

Ameen Abu-Hanna, Annette ten Teije, Birgit A Damoiseaux, Cornelis Boersma, Dave A Dongelmans, David H de Koning, Erol S Hofmans, Evelien Tiggelaar, Frank van Harmelen, Gerty Holla, Heiralde Marck, Iacopo Vagliano, Jan Pander, Joanna E Klopotowska, Jurjen van der Schans, Joris E Lieverse, Kitty J Jager, Linda Dusseljee-Peute, Luuk B Hilbrands, Marieke A R Bak, Mariette van den Hoven, Martijn G Kersloot, Menno Maris, Nicolette F de Keizer, Otto R Maarsingh, Paul Blank, Piet Heingraaf, Renée de Wildt, Rosa M P Vos, Ron Herings, Ron J Keizer, Ronald Cornet, Ruben Boyd, Sebastiaan L Knijnenburg, Sipke Visser, Stephanie Medlock, Stijn Gremmen, Teun van Gelder, Tjerk S Heijmens Visser, and Vianda S Stel

AUTHORS’ CONTRIBUTIONS (CRediT TAXONOMY)

D.F.L.: conceptualization, methodology, software, validation, formal analysis, investigation, data curation; writing—original draft, visualization; L.B.H., L.V., R.H.G.O.E.: methodology, validation, writing—review and editing; J.E.K.: conceptualization, data curation, formal analysis, funding acquisition, methodology, project administration, validation, writing—review and editing. All authors and LEAPfROG consortium members gave final approval of the submitted version.

FUNDING

This work was supported by Nederlandse Organisatie voor Wetenschappelijk Onderzoek (NWO; Dutch Research Council) (KICH1.ST01.20.011) and co-funded in cash by Dutch Kidney Foundation and National Intensive Care Evaluation (NICE) foundation, and in kind by PHARMO Institute for Drug Outcomes Research, Castor, InsightRX, Z-Index, Digital Health Link.

DATA AVAILABILITY STATEMENT

The EHR data reused and analyzed in this study are not publicly available due to Dutch privacy regulation and data sharing regulations at Amsterdam UMC. Access to the data might only be provided after reasonable request and explicit consent from Amsterdam UMC.

The integral LEAPfROG project protocol was reviewed by the Medical Ethics Committee of Amsterdam UMC (the Netherlands). This committee provided a waiver from formal approval (W22_340 # 22.412) since the LEAPfROG project, including all substudies, does not fall within the scope of the Dutch Medical Research Involving Human Subjects Act (WMO). Furthermore, due to the data size, approaching all patients for informed consent would require an unreasonable effort. Under Dutch law, this is a valid argument for research to be exempted from requesting informed consent. Patients that actively opted out from reuse of their EHR for research were excluded. Prior to obtaining the data, a Data Protection Impact Assessment (DPIA) was conducted and the data collection was registered in a national register.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflicts of interest.

REFERENCES

  • 1. Francis  A, Harhay  MN, Ong  ACM  et al.  Chronic kidney disease and the global public health agenda: an international consensus. Nat Rev Nephrol  2024;20:473–85. 10.1038/s41581-024-00820-6 [DOI] [PubMed] [Google Scholar]
  • 2. Romagnani  P, Remuzzi  G, Glassock  R  et al.  Chronic kidney disease. Nat Rev Dis Primers  2017;3:17088. 10.1038/nrdp.2017.88 [DOI] [PubMed] [Google Scholar]
  • 3. Fraser  SD, Taal  MW. Multimorbidity in people with chronic kidney disease: implications for outcomes and treatment. Curr Opin Nephrol Hypertens  2016;25:465–72. 10.1097/MNH.0000000000000270 [DOI] [PubMed] [Google Scholar]
  • 4. Colombijn  JMT, Idema  DL, Van Beem  S  et al.  Representation of patients with chronic kidney disease in clinical trials of cardiovascular disease medications: a systematic review. JAMA Netw Open  2024;7:e240427–. 10.1001/jamanetworkopen.2024.0427 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Kitchlu  A, Shapiro  J, Amir  E  et al.  Representation of patients with chronic kidney disease in trials of cancer therapy. JAMA  2018;319:2437–9. 10.1001/jama.2018.7260 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Baigent  C, Herrington  WG, Coresh  J  et al.  Challenges in conducting clinical trials in nephrology: conclusions from a Kidney Disease – Improving Global Outcomes (KDIGO) Controversies Conference. Kidney Int  2017;92:297–305. 10.1016/j.kint.2017.04.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Perkovic  V, Craig  JC, Chailimpamontree  W  et al.  Action plan for optimizing the design of clinical trials in chronic kidney disease. Kidney Int Suppl  2017;7:138–44. 10.1016/j.kisu.2017.07.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Navaneethan  SD, Jolly  SE, Sharp  J  et al.  Electronic health records: a new tool to combat chronic kidney disease?  Clin Nephrol  2013;79:175–83. 10.5414/CN107757 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Drawz  PE, Archdeacon  P, McDonald  CJ  et al.  CKD as a model for improving chronic disease care through electronic health records. Clin J Am Soc Nephrol  2015;10:1488–99. 10.2215/CJN.00940115 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Tummalapalli  SL, Peralta  CA. An electronic CKD phenotype: a step forward in improving kidney care. Clin J Am Soc Nephrol  2019;14:1277–9. 10.2215/CJN.08180719 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Wei  WQ, Denny  JC. Extracting research-quality phenotypes from electronic health records to support precision medicine. Genome Med  2015;7:41. 10.1186/s13073-015-0166-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Richesson  RL, Hammond  WE, Nahm  M  et al.  Electronic health records based phenotyping in next-generation clinical trials: a perspective from the NIH Health Care Systems Collaboratory. J Am Med Inform Assoc  2013;20:e226–31. 10.1136/amiajnl-2013-001926 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Use of Real-World Evidence to Support Regulatory Decision-Making for Medical Devices. Draft Guidance for Industry and Food and Drug Administration Staff . 2023.
  • 14. Stevens  PE, Ahmed  SB, Carrero  JJ  et al.  KDIGO 2024 clinical practice guideline for the evaluation and management of chronic kidney disease. Kidney Int  2024;105:S117–314. 10.1016/j.kint.2023.10.018 [DOI] [PubMed] [Google Scholar]
  • 15. Chen  W, Abeyaratne  A, Gorham  G  et al.  Development and validation of algorithms to identify patients with chronic kidney disease and related chronic diseases across the Northern Territory, Australia. BMC Nephrol  2022;23:320. 10.1186/s12882-022-02947-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Norton  JM, Ali  K, Jurkovitz  CT  et al.  Development and validation of a pragmatic electronic phenotype for CKD. Clin J Am Soc Nephrol  2019;14:1306–14. 10.2215/CJN.00360119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Ozrazgat-Baslanti  T, Ren  Y, Adiyeke  E  et al.  Development and validation of a race-agnostic computable phenotype for kidney health in adult hospitalized patients. PLoS One  2024;19:e0299332. 10.1371/journal.pone.0299332 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Shang  N, Khan  A, Polubriaginof  F  et al.  Medical records-based chronic kidney disease phenotype for clinical care and “big data” observational and genetic studies. NPJ Digit Med  2021;4:70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Carrero  JJ, Fu  EL, Vestergaard  SV  et al.  Defining measures of kidney function in observational studies using routine health care data: methodological and reporting considerations. Kidney Int  2023;103:53–69. 10.1016/j.kint.2022.09.020 [DOI] [PubMed] [Google Scholar]
  • 20. Kovesdy  CP, Epidemiology of chronic kidney disease: an update 2022. Kidney Int Suppl  2022;12:7–11. 10.1016/j.kisu.2021.11.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. McDonald  HI, Shaw  C, Thomas  SL  et al.  Methodological challenges when carrying out research on CKD and AKI using routine electronic health records. Kidney Int  2016;90:943–9. 10.1016/j.kint.2016.04.010 [DOI] [PubMed] [Google Scholar]
  • 22. Grams  ME, Plantinga  LC, Hedgeman  E  et al.  Validation of CKD and related conditions in existing data sets: a systematic review. Am J Kidney Dis  2011;57:44–54. 10.1053/j.ajkd.2010.05.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Vestergaard  SV, Christiansen  CF, Thomsen  RW  et al.  Identification of patients with CKD in medical databases: a comparison of different algorithms. Clin J Am Soc Nephrol  2021;16:543–51. 10.2215/CJN.15691020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Glassock  RJ, Warnock  DG, Delanaye  P. The global burden of chronic kidney disease: estimates, variability and pitfalls. Nat Rev Nephrol  2017;13:104–14. 10.1038/nrneph.2016.163 [DOI] [PubMed] [Google Scholar]
  • 25. Anderson  J, Glynn  LG. Definition of chronic kidney disease and measurement of kidney function in original research papers: a review of the literature. Nephrol Dial Transplant  2011;26:2793–8. 10.1093/ndt/gfq849 [DOI] [PubMed] [Google Scholar]
  • 26. Mansouri  I, Raffray  M, Lassalle  M  et al.  An algorithm for identifying chronic kidney disease in the French national health insurance claims database. Nephrol Ther  2022;18:255–62. 10.1016/j.nephro.2022.03.003 [DOI] [PubMed] [Google Scholar]
  • 27. Marino  C, Ferraro  PM, Bargagli  M  et al.  Prevalence of chronic kidney disease in the Lazio region, Italy: a classification algorithm based on health information systems. BMC Nephrol  2020;21:23. 10.1186/s12882-020-1689-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Fleet  JL, Dixon  SN, Shariff  SZ  et al.  Detecting chronic kidney disease in population-based administrative databases using an algorithm of hospital encounter and physician claim codes. BMC Nephrol  2013;14:81. 10.1186/1471-2369-14-81 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Norton  JM, Grunwald  L, Banaag  A  et al.  CKD prevalence in the military health system: coded versus uncoded CKD. Kidney Med  2021;3:586–595.e1.e1. 10.1016/j.xkme.2021.03.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. So  L, Evans  D, Quan  H. ICD-10 coding algorithms for defining comorbidities of acute myocardial infarction. BMC Health Serv Res  2006;6:161. 10.1186/1472-6963-6-161 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Levey  AS, Stevens  LA, Schmid  CH  et al.  A new equation to estimate glomerular filtration rate. Ann Intern Med  2009;150:604–12. 10.7326/0003-4819-150-9-200905050-00006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. KDIGO 2012 clinical practice guideline for the evaluation and management of chronic kidney disease. Kidney Int Suppl  2012;3:134–5. [DOI] [PubMed] [Google Scholar]
  • 33. Delanaye  P, Glassock  RJ, De Broe  ME. Epidemiology of chronic kidney disease: think (at least) twice!  Clin Kidney J  2017;10:370–4. 10.1093/ckj/sfw154 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. De Broe  ME, Gharbi  MB, Zamd  M  et al.  Why overestimate or underestimate chronic kidney disease when correct estimation is possible?  Nephrol Dial Transplant  2017;32:ii136–41. 10.1093/ndt/gfw267 [DOI] [PubMed] [Google Scholar]
  • 35. Gasparini  A, Evans  M, Coresh  J  et al.  Prevalence and recognition of chronic kidney disease in Stockholm healthcare. Nephrol Dial Transplant  2016;31:2086–94. 10.1093/ndt/gfw354 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Sundström  J, Bodegard  J, Bollmann  A  et al.  Prevalence, outcomes, and cost of chronic kidney disease in a contemporary population of 2·4 million patients from 11 countries: the CaReMe CKD study. Lancet Reg Health Eur  2022;20:100438. 10.1016/j.lanepe.2022.100438 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Perez-Gomez  MV, Bartsch  L-A, Castillo-Rodriguez  E  et al.  Clarifying the concept of chronic kidney disease for non-nephrologists. Clin Kidney J  2019;12:258–61. 10.1093/ckj/sfz007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Vlasschaert  MEO, Bejaimal  SAD, Hackam  DG  et al.  Validity of administrative database coding for kidney disease: a systematic review. Am J Kidney Dis  2011;57:29–43. 10.1053/j.ajkd.2010.08.031 [DOI] [PubMed] [Google Scholar]
  • 39. Klappe  ES, de Keizer  NF, Cornet  R. Factors influencing problem list use in electronic health records—application of the unified theory of acceptance and use of technology. Appl Clin Inform  2020;11:415–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Robertson  LM, Denadai  L, Black  C  et al.  Is routine hospital episode data sufficient for identifying individuals with chronic kidney disease? A comparison study with laboratory data. Health Informatics J  2016;22:383–96. 10.1177/1460458214562286 [DOI] [PubMed] [Google Scholar]
  • 41. Huber  CA, Szucs  TD, Rapold  R  et al.  Identifying patients with chronic conditions using pharmacy data in Switzerland: an updated mapping approach to the classification of medications. BMC Public Health  2013;13:1030. 10.1186/1471-2458-13-1030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Fan  Li, Inker  LA, Rossert  J  et al.  Glomerular filtration rate estimation using cystatin C alone or combined with creatinine as a confirmatory test. Nephrol Dial Transplant  2014;29:1195–203. 10.1093/ndt/gft509 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Raab  R, Küderle  A, Zakreuskaya  A  et al.  Federated electronic health records for the European Health Data Space. Lancet Digit Health  2023;5:e840–7. 10.1016/S2589-7500(23)00156-5 [DOI] [PubMed] [Google Scholar]
  • 44. Alfano  G, Perrone  R, Fontana  F  et al.  Rethinking chronic kidney disease in the aging population. Life  2022;12:1724. 10.3390/life12111724 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Delanaye  P, Jager  KJ, Bökenkamp  A  et al.  CKD: a call for an age-adapted definition. J Am Soc Nephrol  2019;30:1785–805. 10.1681/ASN.2019030238 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Waikar  SS, Rebholz  CM, Zheng  Z  et al.  Biological variability of estimated GFR and albuminuria in CKD. Am J Kidney Dis  2018;72:538–46. 10.1053/j.ajkd.2018.04.023 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

sfaf073_Supplemental_Files

Data Availability Statement

The EHR data reused and analyzed in this study are not publicly available due to Dutch privacy regulation and data sharing regulations at Amsterdam UMC. Access to the data might only be provided after reasonable request and explicit consent from Amsterdam UMC.

The integral LEAPfROG project protocol was reviewed by the Medical Ethics Committee of Amsterdam UMC (the Netherlands). This committee provided a waiver from formal approval (W22_340 # 22.412) since the LEAPfROG project, including all substudies, does not fall within the scope of the Dutch Medical Research Involving Human Subjects Act (WMO). Furthermore, due to the data size, approaching all patients for informed consent would require an unreasonable effort. Under Dutch law, this is a valid argument for research to be exempted from requesting informed consent. Patients that actively opted out from reuse of their EHR for research were excluded. Prior to obtaining the data, a Data Protection Impact Assessment (DPIA) was conducted and the data collection was registered in a national register.


Articles from Clinical Kidney Journal are provided here courtesy of Oxford University Press

RESOURCES