Abstract
Understanding how dietary patterns influence chronic kidney disease (CKD) development is crucial for effective prevention strategies. This study identified distinct dietary patterns among Korean adults and investigated their association with CKD development. This retrospective cohort study used data from the Korean Genome and Epidemiology Study health examinee study database of community-dwelling adults aged ≥ 40 years in South Korea (2004–2016). Then, dietary patterns were identified using K-means clustering analysis based on the quantity (weights) of 106 foods and intakes of energy and 22 nutrients. The dependent variable for Cox regression analyses was the development of new-onset CKD. A total of 57,213 participants were classified into three dietary clusters. Cluster C, characterized by lower overall food, energy, and nutrient intakes and higher carbohydrate intake, was independently associated with increased CKD risk (adjusted hazard ratio, 1.59; 95% confidence interval, 1.04–2.41; P = 0.031) compared to Cluster A, characterized by higher intake of vegetables and fish/shellfish. Subgroup analyses revealed that Cluster C still had a significantly high risk for CKD development in age ≥ 65 years, male sex, previous cardiovascular disease, systolic blood pressure ≥ 130 mm Hg, and body mass index ≥ 25 kg/m2. Both the quantity and quality of food intake may influence CKD development in Korean adults. Maintaining a balanced, nutrient-rich diet could be key for CKD prevention, especially in high-risk subgroups.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-025-18627-1.
Keywords: Chronic kidney disease, Cluster analysis, Cohort studies, Diet, Nutrition survey
Subject terms: Epidemiology, Kidney diseases, Nutrition
Introduction
Chronic kidney disease (CKD) is associated with significant morbidity, mortality, and healthcare costs1,2. In South Korea, CKD prevalence has steadily increased, driven by an aging population and the increased lifestyle-related risk factors such as hypertension, diabetes, and obesity3,4. Given the crucial role of dietary intake in modulating these risk factors5,6, understanding nutritional intake patterns specific to the Korean population is essential for developing targeted strategies to mitigate CKD risk.
Previous studies have primarily focused on nutritional management in patients already with CKD7. Some studies have examined the relationship between specific dietary components and CKD risk;8 however, few studies have addressed overall dietary patterns reflecting complex interactions between various nutrients and food groups9. Furthermore, understanding the complex interplay between overall dietary patterns and CKD risk can be essential for developing effective prevention and management strategies10.
Analyzing dietary patterns involving diverse nutrient interactions often requires managing complex, high-dimensional datasets. Traditional statistical methods may not be able to capture underlying patterns. Clustering, an unsupervised machine learning technique, provides a robust solution to this challenge11,12. By grouping individuals with similar dietary patterns, clustering reduces data dimensionality, identifies distinct dietary profiles, and reveals potential associations with health outcomes13. This approach facilitates a holistic understanding of dietary habits, offering insights beyond individual nutrient or food group analysis for assessing CKD-associated dietary risks.
This study’s objectives were twofold; it aimed to leverage clustering techniques to (1) identify distinct dietary patterns among Korean adults and (2) investigate their association with CKD development using a Korean urban cohort. By elucidating the relationship between dietary patterns and CKD risk, this study aimed to enhance the understanding of how dietary patterns contribute to CKD risk, ultimately providing evidence to inform dietary recommendations tailored to the Korean population to mitigate CKD risk.
Methods
Data source and participants
This study utilized the Korean Genome and Epidemiology Study (KoGES) Health Examinees (HEXA) Study database, a population-based cohort of community-dwelling adults aged ≥ 40 years at baseline recruited from the national health examinee registry14. Baseline data were collected from 2004 to 2013, with follow-up assessments conducted between 2012 and 2016. Detailed methodology and study design of the HEXA study have been described previously15. Participants included individuals who completed both baseline and follow-up assessments. Participants with missing food frequency questionnaire (FFQ) data, implausible total energy intake (< 500 or > 4,000 kcal/d), or missing serum creatinine or urine protein measurements at enrollment were excluded. Additionally, participant were excluded based on an underlying kidney disease, a history of kidney disease, an estimated glomerular filtration rate (eGFR) of < 60 mL/min/1.73 m², or a random urine protein dipstick result of ≥ 1 + at enrollment. Furthermore, participants with extreme hemoglobin (hemoglobin < 6 g/dL or > 18 g/dL) or triglyceride (triglyceride > 1,000 mg/dL) values and those with missing height or weight data. A total of 57,213 participants were included in the final analysis (Fig. 1).
Fig. 1.
Study flowchart. eGFR, estimated glomerular filtration rate; FFQ, food frequency questionnaire; HEXA, health examinees; KoGES, Korean Genome and Epidemiology Study.
This study was approved by the Institutional Review Board at Pusan National University Hospital (Approval No. 2111-005-108), and the Institutional Review Board waived the requirement for informed consent given the study’s retrospective design and the use of anonymized data. This study adhered to the principles of the Declaration of Helsinki.
Dietary assessment, food grouping, and food categories
Dietary intake was evaluated at enrollment using a validated semiquantitative FFQ comprising 106 food items. Participants were required to recall the frequency and portion sizes of food items consumed over the past 12 months. Intake frequency was recorded using a 9-point scale: (< 1 time/month or never, 1 time/month, 2–3 times/month, 1–2 times/week, 3–4 times/week, 5–6 times/week, 1 time/day, 2 times/day, or 3 times/day). The amount consumed was evaluated using a 3-point scale: less, standard, or more. Accordingly, data on the daily intake (g/d) of 106 food items were calculated using this method. The 106 food items were categorized into 21 groups based on similar types, adapted from a previous report: rice, other grains, noodles and dumplings, wheat flour and bread, potatoes, sweets, soybean pastes, bean, tofu and soymilk, nuts, vegetables, kimchi, mushroom, fruits, red meat and its products, white meat and its products, eggs, fish and shellfish, seaweeds, milk and dairy products, beverages, and coffee and tea (Table S1)16. These groups were further consolidated into seven food categories, following modifications to food exchange list guidelines from the Korean Diabetes Association17. Six categories from the food exchange list—grains, fish and meat, vegetables, fats, milk, and fruits—were used, with the addition of a sugary and caffeinated consumables category.
Daily energy and specific nutrient intake, including macronutrients (protein, fat, and carbohydrates), vitamins, and minerals, were also considered. We evaluated macronutrient intake and the energy contribution ratios of carbohydrates, protein, and fat. The energy contribution from each macronutrient was calculated based on the intake and caloric values of carbohydrates (4 kcal/g), protein (4 kcal/g), and fat (9 kcal/g). The energy contribution ratio for carbohydrates was determined as A × 100/(A + B + C), for protein as B × 100/(A + B + C), for fat as C × 100/(A + B + C), where A, B, and C represent the energy derived from carbohydrates, protein, and fat, respectively. These ratios were classified according to the Manual of KoGES in the following manner: for carbohydrates, a ratio of < 55% was considered insufficient, 55–70% as adequate, and ≥ 70% as excessive; for protein, < 7% was considered insufficient, 7–20% as adequate, and ≥ 20% as excessive; and for fat, < 15% was considered insufficient, 15–25% as adequate, and ≥ 25% as excessive.
Covariates
Given the potential variations in food intake based on age, sex, and body mass index (BMI), these factors were included as covariates. Additionally, underlying comorbidities, including diabetes mellitus (DM), hypertension (HTN), and cardiovascular disease (CVD), were considered. Lifestyle habits, including smoking (never, former smoker, or current smoker) and drinking status (non-, former, or current drinker), were also considered. We included education (low, middle, or high) and income (low, middle, or high) status as covariates. Educational levels were defined as follows: low—less than high school completion; middle—high school graduate; and high—post-secondary or graduate-level education. Income levels were categorized based on monthly family income in the following manner: low, <₩1,500,000 (approximately $1,130 USD); middle, ₩1,500,000 ₩4,000,000 (approximately $1,130–$3,000 USD); and high, >₩4,000,000 (approximately $3,000 USD). Serum hemoglobin, albumin, and total cholesterol levels, which serve as indicators of nutritional status, were also considered. Serum creatinine levels were measured using the Jaffe method throughout the study period. The serum creatinine values were reduced by a calibration factor of 5% to standardize them to the isotope dilution mass spectrometry reference method18,19. The eGFR was derived using the CKD-Epidemiology Collaboration Eq.20.
Determining outcome
Primary outcome was new-onset CKD, defined as an eGFR < 60 mL/min/1.73 m² during follow-up. Follow-up durations were calculated from the baseline date to either the date of the follow-up assessment or new-onset CKD development, whichever occurred first.
Clustering and statistical analysis
Data normality was assessed using histograms for visual inspection, with skewness and kurtosis used to quantify symmetry and peakedness, respectively. Normally distributed variables are expressed as mean ± standard deviation and non-normally distributed variables as median (interquartile range). Categorical variables were expressed as frequencies (proportion). Group comparisons were made using the analysis of variance or Kruskal–Wallis test for continuous variables and the chi-square test for categorical variables. Post-hoc analysis for multiple comparisons was conducted using the Bonferroni method to adjust for Type I error. Cox regression was employed to assess the adjusted hazard ratios (aHRs) and 95% confidence intervals (CIs) for CKD risk. Cox proportional hazards regression models were constructed with progressive increases in the adjustments for potential confounders: Model I was unadjusted; Model II was adjusted for age and sex; Model III was further adjusted for BMI, HTN, DM, preexisting CVD, and systolic blood pressure; and Model IV was adjusted for eGFR, hemoglobin, serum albumin, total cholesterol, and educational status.
Clustering analysis was performed on the quantity (weights) of 106 daily foods, energy, and 22 nutrient variables using the K-means method. We selected the K-means clustering algorithm given its computational efficiency, scalability to high-dimensional continuous data, and ease of interpretation. Compared to alternative methods, such as hierarchical clustering or latent class analysis, K-means offers a favorable balance between speed and interpretability, making it particularly suited for large-scale dietary datasets with numerous variables.
The optimal number of clusters was determined through an internal validation process based on multiple criteria: (1) the within-cluster sum of squares (Within SS) to evaluate cluster cohesion, (2) the statistical significance of differences in CKD incidence among clusters, and (3) the clinical interpretability of the dietary patterns identified within each cluster. Based on these metrics, we selected a three-cluster solution as the optimal choice. These results, including sensitivity analyses across varying cluster numbers and variable sets, are presented in Fig. S1. All clustering procedures were conducted using raw intake data, without adjusting for covariates, such as age, sex, or energy intake, to reflect actual consumption patterns. To facilitate the interpretation of dietary patterns following clustering, residuals (i.e., actual intake minus predicted intake) were calculated for each dietary variable using linear regression models adjusted for age, sex, and BMI. These residuals were not used in cluster formation, but instead employed post hoc to illustrate relative intake differences across clusters after accounting for potential confounders. P-values < 0.05 were considered statistically significant. All statistical analyses were performed using R version 4.4.1. (R Core Team, 2023) with additional packages (stats, tableone, ggplot2, factoextra, Rtsne, survival, and survminer).
Results
Clusters and baseline characteristics
Baseline characteristics of the study population are presented in Table 1, across three clusters. Figure 2 depicts the distribution of clustering participants. Cluster C had the highest mean age (54.13 ± 7.98 years), followed by Cluster B (52.89 ± 7.91 years) and Cluster A (52.55 ± 7.77 years) (P < 0.001). Sex distribution was similar across clusters, with male participants comprising approximately one-third of each group, the highest being in Cluster B (34.9%) (P < 0.001). Significant differences were observed in clinical factors, including systolic blood pressure (SBP); BMI; and DM, HTN, and CVD prevalences (P < 0.001 for all); Cluster C had the highest DM (9.5%), HTN (29.1%), and CVD (4.1%) prevalences. Smoking and drinking behaviors varied significantly; Cluster C had the lowest proportion of current smokers (9.8%) and the highest proportion of non-drinkers or former drinkers (58.1%). Socioeconomic factors, including educational and income status, showed substantial variation across the clusters (P < 0.001). Cluster C had the highest proportion of individuals with low-education levels (35.7%) and low-income status (20.3%), whereas Cluster A had the highest proportion with high-education levels (37.1%) and high-income status (12.7%). Laboratory parameters, including eGFR, hemoglobin, albumin levels, and total cholesterol, showed small but statistically significant differences (P < 0.001 for all comparisons). Detailed post-hoc analysis results are presented in Table S2.
Table 1.
Baseline characteristics.
| Characteristic | Overall (n = 57,213) | Cluster A (n = 4,152) | Cluster B (n = 22,127) | Cluster C (n = 30,934) | P-value |
|---|---|---|---|---|---|
| Male, n (%) | 19,227 (33.6) | 1,373 (33.1) | 7,729 (34.9) | 10,125 (32.7) | < 0.001 |
| Age, years | 53.54 ± 7.97 | 52.55 ± 7.77 | 52.89 ± 7.91 | 54.13 ± 7.98 | < 0.001 |
| SBP, mm Hg | 122.07 ± 14.88 | 121.54 ± 14.81 | 121.74 ± 14.63 | 122.38 ± 15.05 | < 0.001 |
| BMI, kg/m2 | 23.87 ± 2.85 | 24.08 ± 2.80 | 23.93 ± 2.87 | 23.80 ± 2.84 | < 0.001 |
| DM, n (%) | 5,065 (8.9) | 337 (8.1) | 1,791 (8.1) | 2,937 (9.5) | < 0.001 |
| HTN, n (%) | 15,900 (27.8) | 1,058 (25.5) | 5,835 (26.4) | 9,007 (29.1) | < 0.001 |
| CVD, n (%) | 2,124 (3.7) | 125 (3.0) | 741 (3.3) | 1,258 (4.1) | < 0.001 |
| Smoking status, n (%) | < 0.001 | ||||
| Never | 42,250 (74.1) | 3,083 (74.9) | 16,036 (72.8) | 23,131 (75.0) | |
| Former smoker | 8,853 (15.5) | 589 (14.3) | 3,572 (16.2) | 4,692 (15.2) | |
| Current smoker | 5,879 (10.3) | 445 (10.8) | 2,427 (11.0) | 3,007 (9.8) | |
| Drinking status, n (%) | < 0.001 | ||||
| Non-drinker or former drinker | 31,942 (56.1) | 2,223 (54.0) | 11,812 (53.6) | 17,907 (58.1) | |
| Current drinker | 25,021 (43.9) | 1,894 (46.0) | 10,218 (46.4) | 12,909 (41.9) | |
| Educational status*, n (%) | < 0.001 | ||||
| Low | 17,519 (30.6) | 928 (22.4) | 5,555 (25.1) | 11,036 (35.7) | |
| Middle | 21,338 (37.3) | 1,639 (39.5) | 8,609 (38.9) | 11,090 (35.9) | |
| High | 17,736 (31.0) | 1,541 (37.1) | 7,728 (34.9) | 8,467 (27.4%) | |
| Unknown | 620 (1.1) | 44 (1.1) | 235 (1.1) | 341 (1.1) | |
| Income status, n (%) | < 0.001 | ||||
| Low | 9,749 (17.0) | 500 (12.0) | 2,954 (13.4) | 6,295 (20.3) | |
| Middle | 28,440 (49.7) | 2,032 (48.9) | 11,550 (52.2) | 14,858 (48.0) | |
| High | 13,304 (23.3) | 1,093 (26.3) | 5,499 (24.9) | 6,712 (21.7) | |
| Unknown | 5,720 (10.0) | 527 (12.7) | 2,124 (9.6) | 3,069 (9.9) | |
| eGFR, mL/min/m2 | 92.04 ± 11.33 | 92.51 ± 11.23 | 92.31 ± 11.44 | 91.78 ± 11.26 | < 0.001 |
| Hemoglobin, g/dL | 13.84 ± 1.44 | 13.83 ± 1.46 | 13.88 ± 1.47 | 13.81 ± 1.42 | < 0.001 |
| Albumin, g/dL, | 4.63 ± 0.26 | 4.65 ± 0.26 | 4.63 ± 0.26 | 4.62 ± 0.26 | < 0.001 |
| Total cholesterol, median (Q1–Q3), mg/dL | 195 (173–219) | 195 (176–220) | 196 (173–220) | 194 (172–218) | < 0.001 |
SBP, systolic blood pressure; BMI, body mass index; DM, diabetes mellitus; HTN, hypertension; CVD, cardiovascular disease; eGFR, estimated glomerular filtration rate.
Fig. 2.
K-means clustering of participants, visualized using principal component analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE) methods. All variables were standardized to z-scores before clustering. Participants were grouped into three clusters (A–C) using the K-means algorithm (K = 3). (A) In the PCA plot, each point represents a participant projected into a two-dimensional PCA space, where Dim1 and Dim2 account for 19.9% and 4.2% of the total variance, respectively. Convex hulls outline the outer boundaries of each cluster. (B) In the t-SNE plot, the same individuals are projected, which emphasizes the local structure and preservation of neighborhoods.
Dietary features of each cluster
The unsupervised learning (K-means) method categorized participants into three clusters: Cluster A, B, and C with 4,152, 22,127, and 30,934 participants, respectively. Each cluster’s participant distribution is described in Table S3. Radar plots of normalized mean and standard deviations for each food group revealed that Cluster A had higher consumption levels across most food categories, particularly fish/shellfish and vegetables, indicating a diverse dietary intake compared to Clusters B and C. Cluster C showed relatively lower intake, even less than overall in all categories. Cluster B fell between the two clusters, with moderate intake levels for most food groups (Fig. 3A and Table S4). Energy and nutrient intake profiles normalized by mean and standard deviation demonstrated that Cluster A had highest levels of most nutrients, whereas Cluster C exhibited the lowest levels (Fig. 3B and Table S5).
Fig. 3.
Radar plots for (A) food and (B) nutrient consumption by dietary cluster groups. Cluster A shows higher consumption levels across most food groups, particularly fish/shellfish and vegetables, and higher nutrient consumption levels than other clusters. The range of the axes is the z-score value, normalized to a mean of 0 and a standard deviation of 1.
Among the seven food categories, Cluster A had the highest overall intake, particularly in fish and meat and vegetables, whereas Cluster C had the lowest intake (Fig. S2 and Table S6).
Cluster A had the highest daily intake of essential macronutrients—protein, fat, and carbohydrates—per unit body weight (P < 0.001 for all comparisons), whereas Cluster C had the lowest (Fig. 4). Macronutrient energy contribution ratios showed that the energy contribution from protein is the highest in Cluster A, although the energy contribution from carbohydrates was the highest in Cluster C. (Fig. 5A). Overall, Cluster A displayed balanced ratios of protein, fat, and carbohydrates, while Cluster B exhibited an excess of carbohydrates and fat deficiency. The imbalance was more severe in Cluster C, characterized by excessive carbohydrates and marked fat deficiency (Fig. 5B).
Fig. 4.
Daily essential nutrients intake: protein, fat, and carbohydrate per body weight according to each cluster. Panel (A) shows protein intake, Panel (B) shows fat intake, and Panel (C) shows carbohydrate intake per kilogram of body weight. Cluster A provides the highest daily intake of essential macronutrients, and the intake of each nutrient decreases from Cluster B to Cluster C.
Fig. 5.
Energy intake of each essential nutrient (protein, fat, and carbohydrates). (A) Cluster A shows the highest proportion of energy contribution of protein. Conversely, Cluster C shows the highest proportion of energy contribution of carbohydrate. (B) Cluster A has mostly adequate ratios of protein, fat, and carbohydrates. In contrast, Cluster B shows excess carbohydrates and a fat insufficiency. In Cluster C, this imbalance is more severe, given that excess carbohydrates and inadequate fats are the most prominent.
CKD development
The Kaplan–Meier curves demonstrate that Cluster C showed the highest cumulative incidence of CKD development, while Cluster A had the lowest (P < 0.001) (Fig. 6). Cox regression revealed that Cluster C was independently associated with CKD development in the fully adjusted model (Model IV: aHR, 1.59; 95% CI, 1.04–2.41; P = 0.031) (Table 2). A summary table including the number of CKD events and person-years of follow-up for each cluster is provided in Table S7. Moreover, we included the results of the Schoenfeld residuals test performed to verify the Cox proportional hazards assumption (Fig. S3); for the primary variable of interest, the cluster group, no statistically significant violations were identified (P = 0.099).
Fig. 6.

The cumulative incidence of chronic kidney disease according to dietary cluster groups. This Kaplan–Meier curves show that the cluster C group had a significantly higher cumulative incidence of chronic kidney disease (P < 0.001).
Table 2.
CKD development according to dietary cluster groups.
| Clusters | Model Ia | Model IIb | Model IIIc | Model IVd | ||||
|---|---|---|---|---|---|---|---|---|
| HR (95% CI) | P-value | aHR (95% CI) | P-value | aHR (95% CI) | P-value | aHR (95% CI) | P-value | |
| A | Reference | Reference | Reference | Reference | ||||
| B | 1.63 (1.06–2.49) | 0.025 | 1.40 (0.91–2.15) | 0.121 | 1.42 (0.93–2.18) | 0.106 | 1.41 (0.91–2.16) | 0.120 |
| C | 2.09 (1.38–3.17) | < 0.001 | 1.54 (1.02–2.33) | 0.042 | 1.54 (1.01–2.33) | 0.042 | 1.59 (1.04–2.41) | 0.031 |
aHR, adjusted hazard ratio; HR, hazard ratio; BMI, body mass index; DM, diabetes mellitus; HTN, hypertension; CVD, cardiovascular disease; eGFR, estimated glomerular filtration rate.
aModel I: Unadjusted; χ² = 19.10; df = 2; P < 0.001.
bModel II: Adjusted for age, sex; χ² = 19.10; df = 2; P < 0.001.
cModel III: Model 2 + BMI, HTN, DM, preexisting CVD, systolic blood pressure; χ² = 18.89; df = 2; P < 0.001.
dModel IV: Model 3 + eGFR, hemoglobin, serum albumin, total cholesterol, education status; χ² = 18.88; df = 2; P < 0.001.
Subgroup analysis revealed that Cluster C had a significantly higher risk for CKD development in participants aged ≥ 65 years (aHR, 3.58; 95% CI, 1.31–9.77; P = 0.013), male participants (aHR, 2.00; 95% CI, 1.08–3.71; P = 0.028), those with previous CVD (aHR, 4.62; 95% CI, 1.08–19.72; P = 0.039), SBP ≥ 130 mmHg (aHR, 1.96; 95% CI, 1.08–3.56; P = 0.026), and BMI ≥ 25 kg/m2 (aHR, 3.62; 95% CI, 1.48–8.84; P = 0.005). No differences were observed based on sex, the presence of DM, or the presence of HTN (Table S8).
Discussion
This study identified dietary patterns using a clustering algorithm and evaluated their association with CKD development in Korean adults. Our findings emphasized the importance of a more comprehensive assessment that complements traditional nutrient-specific analyses; by employing an unsupervised clustering approach, this research provided a detailed characterization of dietary patterns and their links to CKD, particularly emphasizing subgroup-specific variations in food intake. Our findings offer new insights into nutritional risk factors for CKD.
Previous research on diet and CKD has predominantly focused on specific nutrients or micronutrients and their relationship with disease risk. However, studies exploring comprehensive dietary patterns reflecting real-world eating habits remain limited. A prior cross-sectional study found that individuals in the lowest quartile of phosphorus, potassium, iron, and zinc levels had significantly higher odds of advanced CKD compared to those in the reference quartile21. Another study showed that insufficient dietary zinc intake may increase CKD risk in individuals with normal kidney function22. Unlike these studies, our clustering-based approach captures comprehensive, real-world dietary patterns by simultaneous integration of multiple foods and nutrients, reflecting the complex interactions inherent in habitual diets. This perspective may better represent actual dietary behaviors and their relationships with CKD risk. Meanwhile, previous studies on dietary patterns have utilized factor analysis to categorize these patterns16,23. In these studies, the highest loadings of each factor were used to define groups based on specific variables. Although this approach predicted CKD development based on the order of values across various factors, it lacked the capacity to integrate patterns cohesively for group classification. Additionally, these studies were limited by the low explanatory power of dietary patterns, ranging from 4.5 to 32.6%.
As mentioned before, one key strength of our study is the use of clustering algorithms, which provide a holistic understanding of dietary patterns by integrating multiple nutrients and food groups24. This approach categorizes individuals based on nutritional behaviors, capturing complex interactions between food items that traditional nutrient-specific analyses may overlook in large cohort samples. Using clustering, we identified three distinct dietary patterns with varying food and nutrient intake levels. Cluster A, characterized by higher overall food and nutrient intakes, demonstrated the importance of a healthy dietary pattern in CKD prevention, emphasizing sufficient essential nutrient intake. In contrast, Clusters B and C showed relatively lower intake levels overall. The higher educational and income levels in Cluster A likely contributed to better dietary patterns, which aligned with a systematic review that reported healthier diets and greater diversity among individuals with higher socioeconomic status25. Our findings suggest the importance of both dietary quality and quantity in CKD risk and highlight the influence of distinct dietary behaviors on health outcomes.
Cluster C, characterized by higher carbohydrate intake and lower fat intake, showed a significantly higher CKD risk than that of Cluster A, which had greater consumption of vegetables and fish/shellfish. These results align with prior studies26. A previous study conducted in both urban and rural areas of Korea reported that high-carbohydrate diets increased CKD risk in individuals without DM27. Previous studies have demonstrated a significant association between a high dietary glycemic index and CKD, particularly highlighting an elevated risk of incident CKD linked to nutrient-poor sources of carbohydrates28. Another study revealed diets with low fat-to-carbohydrate ratios were associated with faster renal function decline and increased CKD risk in the general population29. Previous research has demonstrated that partially replacing carbohydrates with protein or monounsaturated fats can significantly reduce blood pressure, improve lipid profiles, and decrease cardiovascular risk30. Additionally, another study found that a high carbohydrate intake was linked to increased overall mortality risk31. Although excessive carbohydrate consumption increases metabolic risk, high-fat intake may also have adverse effects, highlighting the need for further research. Current CKD nutrition guidelines lack specific recommendations for optimal carbohydrate and fat intake, reinforcing the necessity for additional studies in this area32.
Although traditional CKD risk factors such as HTN, DM, and CVD were more significant33, dietary patterns, as captured by clustering, emerged as potential contributors to CKD development. Furthermore, our subgroup analyses revealed that demographic and clinical conditions such as older age, higher BMI, and preexisting CVD amplified the associations between dietary patterns and CKD incidents. Therefore, these subgroups may have an increased susceptibility to dietary imbalances. Furthermore, individuals in these subgroups were more sensitive to nutritional factors because their preexisting conditions had increased the stress on their kidney functions. Therefore, there was a potential synergistic burden on CKD risk due to the combination of unbalanced dietary patterns with the physiological stress of aging, obesity-related metabolic disturbances, and CVD-associated vascular damage34,35.
This study has some limitations. As an observational study, it cannot eliminate potential biases or unmeasured confounders. However, adjustments were made for confounding factors, and various statistical methods were applied to predict CKD development. Additionally, FFQ, which is subject to recall bias, was used to determine the dietary intake. However, the FFQ is a widely accepted tool for estimating nutritional intake in observational research36. Notably, individual foods within the same food group may have distinct effects on kidney health. For example, foods classified under the same grain group can vary considerably in their sodium and saturated fat content, leading to diverse impacts on renal outcomes. Furthermore, dietary sources vary significantly across geographical regions, which could further influence their effects on kidney health. Moreover, several critical dietary confounders known to affect renal outcomes—such as sodium, phosphorus, and distinctions between refined and whole grains—were unavailable in our dataset and thus could not be accounted for. Future studies with detailed dietary data are warranted to better clarify these effects. Finally, although clustering methods offer valuable insights into complex dietary patterns, they have inherent limitations, including sensitivity to outliers, potential variability in solutions, and a risk of misclassification. Accordingly, we applied data preprocessing and validation analyses to enhance robustness and interpreted the findings within the context of these methodological constraints.
Conclusions
This study identified an association between dietary clusters and CKD risk in Korean adults, indicating that both the quality and quantity of food consumed may be related to CKD development. Dietary habits characterized by high overall intake, particularly vegetables and fish/shellfish, were associated with lower CKD risk, whereas a carbohydrate-rich diet with lower overall intake was linked to increased risk. These findings highlight the importance of healthy dietary behaviors in CKD prevention and can guide the development of targeted dietary intervention strategies. Prospective, multi-ethnic, and multinational studies are required to confirm these associations and investigate the underlying mechanisms.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
Data in this study were from the Korean Genome and Epidemiology Study (KoGES; 4851-302), National Institute of Health, Korea Disease Control and Prevention Agency, Republic of Korea.
Abbreviations
- aHR
Adjusted hazard ratio
- BMI
Body mass index
- CI
Confidence interval
- CKD
Chronic kidney disease
- CVD
Cardiovascular disease
- DM
Diabetes mellitus
- eGFR
Estimated glomerular filtration rate
- FFQ
Food frequency questionnaire
- HEXA
Health Examinees
- HTN
Hypertension
- KoGES
Korean Genome and Epidemiology Study
Author contributions
Conceptualization: DP and HJK; Data curation: JK, DWK, and DL; Formal analysis: JK, TK, DK, and YL; Funding acquisition: HJK; Investigation: DP, JK, DWK, and HJK; Methodology: DP, JK, TK, DK, and YL; Supervision: WHK; Validation: DP, WHK, and HJK; Visualization: JK, and HJK; Writing – original draft: DP, JK, and HJK; Writing– review & editing: DP, JK, and HJK. All authors have read and approved the final manuscript.
Funding
This research was supported by the Bio & Medical Technology Development Program of the National Research Foundation (NRF) funded by the Korean government (MSIT) (No. RS-2023-00223764).
Data availability
The data from this study are not publicly available due to privacy and ethical restrictions of the Korea Genome and Epidemiology Study (KoGES; 4851-302) but are available from the corresponding author on reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Ethical approval
This study was approved by the Institutional Review Board at Pusan National University Hospital (Approval No. 2111-005-108), and the Institutional Review Board waived an informed consent due to the study’s retrospective design and the use of anonymized data. This study adhered to the principle of the Declaration of Helsinki.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Dougho Park and Jinmi Kim contributed equally to this work.
References
- 1.Bikbov, B. et al. Global, regional, and National burden of chronic kidney disease, 1990–2017: a systematic analysis for the global burden of disease study 2017. Lancet395, 709–733. 10.1016/s0140-6736(20)30045-3 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Makmun, A. et al. The burden of chronic kidney disease in Asia region: a review of the evidence, current challenges, and future directions. Kidney Res. Clin. Pract.44, 411–433. 10.23876/j.krcp.23.194 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Oh, K. H., Park, S. K., Kim, J. & Ahn, C. The Korean cohort study for outcomes in patients with chronic kidney disease (KNOW-CKD): A Korean chronic kidney disease cohort. J. Prev. Med. Public. Health. 55, 313–320. 10.3961/jpmph.22.031 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Kuma, A. & Kato, A. Lifestyle-related risk factors for the incidence and progression of chronic kidney disease in the healthy young and middle-aged population. Nutrients14. 10.3390/nu14183787 (2022). [DOI] [PMC free article] [PubMed]
- 5.Thomas, G. et al. Metabolic syndrome and kidney disease. Clin. J. Am. Soc. Nephrol.6, 2364–2373. 10.2215/cjn.02180311 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Narasaki, Y., Siu, M. K., Nguyen, M., Kalantar-Zadeh, K. & Rhee, C. M. Personalized nutritional management in the transition from non-dialysis dependent chronic kidney disease to Dialysis. Kidney Res. Clin. Pract.43, 575–585. 10.23876/j.krcp.23.142 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Apetrii, M., Timofte, D., Voroneanu, L. & Covic, A. Nutrition in chronic kidney disease—The role of proteins and specific diets. Nutrients13. 10.3390/nu13030956 (2021). [DOI] [PMC free article] [PubMed]
- 8.He, L. Q., Wu, X. H., Huang, Y. Q., Zhang, X. Y. & Shu, L. Dietary patterns and chronic kidney disease risk: a systematic review and updated meta-analysis of observational studies. Nutr. J.20. 10.1186/s12937-020-00661-6 (2021). [DOI] [PMC free article] [PubMed]
- 9.Kramer, H. Diet and chronic kidney disease. Adv. Nutr.10, S367–S379. 10.1093/advances/nmz011 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Li, Z., Forester, S., Jennings-Dobbs, E., Heber, D. & Perspective A comprehensive evaluation of data quality in nutrient databases. Adv. Nutr.14, 379–391. 10.1016/j.advnut.2023.02.005 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Aletti, G. & Micheletti, A. A clustering algorithm for multivariate data streams with correlated components. J. Big Data4. 10.1186/s40537-017-0109-0 (2017).
- 12.Pan, W., Shen, X. & Liu, B. Cluster analysis: unsupervised learning via supervised learning with a non-convex penalty. J. Mach. Learn. Res.14, 1865 (2013). [PMC free article] [PubMed] [Google Scholar]
- 13.Jimenez-Lopez, E. et al. Clustering of mediterranean dietary patterns linked with health-related quality of life in adolescents: the EHDLA study. Eur. J. Pediatr.182, 4113–4121. 10.1007/s00431-023-05069-y (2023). [DOI] [PubMed] [Google Scholar]
- 14.Kim, Y. & Han, B. G. Cohort profile: the Korean genome and epidemiology study (KoGES) consortium. Int. J. Epidemiol.46, e20. 10.1093/ije/dyv316 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Study, H. E. The health examinees (HEXA) study: rationale, study design and baseline characteristics. Asian Pac. J. Cancer Prev.16, 1591–1597. 10.7314/apjcp.2015.16.4.1591 (2015). [DOI] [PubMed] [Google Scholar]
- 16.Fu, J. & Shin, S. The association of dietary patterns with incident chronic kidney disease and kidney function decline among middle-aged Korean adults: a cohort study. Epidemiol. Health. 45, e2023037. 10.4178/epih.e2023037 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lim, J. H. Guidelines for the use of 2023 food exchange lists for diabetes meal planning. J. Korean Diabetes. 25, 42–51. 10.4093/jkd.2024.25.1.42 (2024). [Google Scholar]
- 18.Levey, A. S. et al. Expressing the modification of diet in renal disease study equation for estimating glomerular filtration rate with standardized serum creatinine values. Clin. Chem.53, 766–772. 10.1373/clinchem.2006.077180 (2007). [DOI] [PubMed] [Google Scholar]
- 19.Matsushita, K. et al. Comparison of risk prediction using the CKD-EPI equation and the MDRD study equation for estimated glomerular filtration rate. JAMA307, 1941–1951. 10.1001/jama.2012.3954 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Inker, L. A. et al. Estimating glomerular filtration rate from serum creatinine and Cystatin C. N Engl. J. Med.367, 20–29. 10.1056/NEJMoa1114248 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kim, J. et al. Association between dietary mineral intake and chronic kidney disease: the health examinees (HEXA) study. Int. J. Environ. Res. Public Health. 15. 10.3390/ijerph15061070 (2018). [DOI] [PMC free article] [PubMed]
- 22.Joo, Y. S. et al. Dietary zinc intake and incident chronic kidney disease. Clin. Nutr.40, 1039–1045. 10.1016/j.clnu.2020.07.005 (2021). [DOI] [PubMed] [Google Scholar]
- 23.Lee, J. E. et al. Dietary pattern classifications with nutrient intake and health-risk factors in Korean men. Nutrition27, 26–33. 10.1016/j.nut.2009.10.011 (2011). [DOI] [PubMed] [Google Scholar]
- 24.Silva, V. C. et al. Clustering analysis and machine learning algorithms in the prediction of dietary patterns: Cross-sectional results of the Brazilian longitudinal study of adult health (ELSA-Brasil). J. Hum. Nutr. Diet.35, 883–894. 10.1111/jhn.12992 (2022). [DOI] [PubMed] [Google Scholar]
- 25.Mayén, A. L., Marques-Vidal, P., Paccaud, F., Bovet, P. & Stringhini, S. Socioeconomic determinants of dietary patterns in low- and middle-income countries: a systematic review. Am. J. Clin. Nutr.100, 1520–1531. 10.3945/ajcn.114.089029 (2014). [DOI] [PubMed] [Google Scholar]
- 26.Rhee, C. M. et al. Nutritional and dietary management of chronic kidney disease under Conservative and preservative kidney care without Dialysis. J. Ren. Nutr.33, S56–S66. 10.1053/j.jrn.2023.06.010 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Nam, K. H. et al. Carbohydrate-rich diet is associated with increased risk of incident chronic kidney disease in non-diabetic subjects. J. Clin. Med.8. 10.3390/jcm8060793 (2019). [DOI] [PMC free article] [PubMed]
- 28.Gopinath, B. et al. Carbohydrate nutrition is associated with the 5-year incidence of chronic kidney disease. J. Nutr.141, 433–439. 10.3945/jn.110.134304 (2011). [DOI] [PubMed] [Google Scholar]
- 29.Kim, H. et al. Relationship between carbohydrate-to-fat intake ratio and the development of chronic kidney disease: A community-based prospective cohort study. Clin. Nutr.40, 5346–5354. 10.1016/j.clnu.2021.09.001 (2021). [DOI] [PubMed] [Google Scholar]
- 30.Appel, L. J. et al. Effects of protein, monounsaturated fat, and carbohydrate intake on blood pressure and serum lipids. Jama294. 10.1001/jama.294.19.2455 (2005). [DOI] [PubMed]
- 31.Dehghan, M. et al. Associations of fats and carbohydrate intake with cardiovascular disease and mortality in 18 countries from five continents (PURE): a prospective cohort study. Lancet390, 2050–2062. 10.1016/s0140-6736(17)32252-3 (2017). [DOI] [PubMed] [Google Scholar]
- 32.Ikizler, T. A. et al. KDOQI clinical practice guideline for nutrition in CKD: 2020 update. Am. J. Kidney Dis.76, S1–S107. 10.1053/j.ajkd.2020.05.006 (2020). [DOI] [PubMed] [Google Scholar]
- 33.Lo, R., Narasaki, Y., Lei, S. & Rhee, C. M. Management of traditional risk factors for the development and progression of chronic kidney disease. Clin. Kidney J.16, 1737–1750. 10.1093/ckj/sfad101 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Nomura, I., Kato, J. & Kitamura, K. Association between body mass index and chronic kidney disease: a population-based, cross-sectional study of a Japanese community. Vasc Health Risk Manag. 5, 315–320. 10.2147/vhrm.s5522 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Bansal, N., Vittinghoff, E., Plantinga, L. & Hsu, C. Y. Does chronic kidney disease modify the association between body mass index and cardiovascular disease risk factors. J. Nephrol.25, 317–324. 10.5301/JN.2011.8454 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Naska, A., Lagiou, A. & Lagiou, P. Dietary assessment methods in epidemiological research: current state of the Art and future prospects. F1000Research6. 10.12688/f1000research.10703.1 (2017). [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data from this study are not publicly available due to privacy and ethical restrictions of the Korea Genome and Epidemiology Study (KoGES; 4851-302) but are available from the corresponding author on reasonable request.





