Skip to main content
Diabetology & Metabolic Syndrome logoLink to Diabetology & Metabolic Syndrome
. 2026 Jun 1;18:168. doi: 10.1186/s13098-026-02176-2

Development and temporal external validation of a high-specificity XGBoost rule-in model for diabetes in middle-aged and older Korean adults

Soo Myeong Kim 1, Jung Min Cho 2,✉
PMCID: PMC13440108  PMID: 42226262

Abstract

Early identification of diabetes in older adults is essential for preventing complications, yet many high‑risk individuals remain undetected in community settings. Using recent cycles of the nationally representative Korea National Health and Nutrition Examination Survey (KNHANES 2020–2023), we developed and temporally validated an Extreme Gradient Boosting (XGBoost) model to rule-in diabetes among Korean adults aged ≥ 50 years. Candidate predictors included sociodemographic factors, health behaviors, anthropometric indices, blood pressure, medical history, and simple laboratory markers. Data from 2020 to 2022 were used for model development, with the 2023 cycle reserved as a temporal external validation cohort. We prespecified a high‑specificity rule‑in threshold based on the development cohort and evaluated discrimination (area under the receiver operating characteristic curve (AUROC) and average precision), calibration, Brier score, classification metrics, decision‑curve net benefit, and Shapley additive explanation (SHAP) values. In temporal external validation, the XGBoost model demonstrated robust performance (AUROC 0.868; average precision 0.646; Brier score 0.101) and achieved high rule-in accuracy (0.866), specificity (97.3%), positive predictive value (76.7%), and F1-score (0.521) at the prespecified threshold. Compared with logistic regression and random forest, the model showed superior rule-in performance and performed comparably to Light Gradient Boosting Machine (LightGBM), a gradient boosting framework based on decision tree ensembles, in terms of specificity and positive predictive value, while intentionally accepting reduced sensitivity consistent with a high-specificity design. SHAP analyses identified urine creatinine, urine specific gravity, urine albumin, total cholesterol and other lipids, body mass index, waist circumference, and a history of hypertension and dyslipidemia as major contributors to model predictions. These findings indicate that an XGBoost-based rule-in model using routinely collected survey variables can efficiently identify older Korean adults with a high probability of diabetes and may serve as a practical decision-support tool for prioritizing confirmatory testing and targeted screening in community settings with limited resources.

Supplementary Information

The online version contains supplementary material available at 10.1186/s13098-026-02176-2.

Keywords: Diabetes mellitus, Type 2; Machine learning; Health surveys; Middle aged; Predictive value of tests

Introduction

Diabetes mellitus (DM) imposes a substantial and growing burden on global health systems, contributing to morbidity, mortality, and escalating healthcare expenditures. According to a recent Global Burden of Disease analysis, an estimated 529 million people were living with diabetes worldwide in 2021, a number projected to exceed 1.31 billion by 2050 [1]. A large fraction of individuals with diabetes remains undiagnosed, delaying therapeutic intervention and increasing the risk of complications: the International Diabetes Federation estimates that roughly 240 million adults—nearly one in two people with diabetes—are undiagnosed globally, with particularly high proportions in low- and middle-income countries [2]. Undiagnosed or late-diagnosed diabetes is associated with prolonged hyperglycemia, early onset of microvascular and macrovascular complications, and higher long-term healthcare costs [3].

Given this burden, early detection of diabetes or high-risk states is a major public health priority. Over the past decades, numerous diabetes risk prediction models have been developed using traditional statistical methods, primarily logistic regression or Cox proportional hazards regression, incorporating demographic, lifestyle, and biochemical predictors [4, 5]. Tools such as the Finnish Diabetes Risk Score (FINDRISC) and the Korean Diabetes Risk Score have informed population-level screening and prevention programs [6, 7]. However, many regression-based models are optimized for sensitivity—to identify as many at-risk individuals as possible—at the expense of specificity, leading to a high false-positive rate and relatively low positive predictive value (PPV). Such sensitivity-oriented models may be useful for initial broad screening, but they are less suited to rule-in clinical decision-making, where reliable confirmation of disease risk is required before initiating costly or invasive follow-up tests.

Several groups have applied machine learning (ML) to address these limitations. ML algorithms can capture complex, non-linear interactions among demographic, biochemical, and nutritional variables without assuming linearity or independence. Extreme Gradient Boosting (XGBoost) has shown strong performance for structured health data by sequentially building regularized decision trees with inherent handling of missing values [8]. Systematic reviews of ML-based diabetes prediction models suggest that gradient-boosting and related ensemble methods often outperform conventional logistic regression in terms of discrimination, while relying on overlapping sets of anthropometric and metabolic predictors [9]. At the same time, many ML models yield poorly calibrated probabilities and lack transparency, which restricts their clinical adoption. Classical calibration techniques such as Platt scaling and isotonic regression, together with more recent guidance on calibration assessment, can substantially improve the reliability of predicted probabilities [10, 11]. Complementary explainable AI methods, including SHapley Additive exPlanations (SHAP), enable visual and quantitative assessment of feature contributions at both global and individual levels and have been applied to a variety of diabetes-related prediction tasks [12, 13]. Recent bibliometric and scoping reviews further highlight rapid growth, methodological heterogeneity, and variable reporting quality among AI-based diabetes risk prediction models, underscoring the need for robust validation and clearly defined clinical use cases [14–16]. A recent scoping review of 40 studies on AI-based methods for type 2 diabetes risk prediction reported that most models used classical machine learning applied to electronic health records, with only five studies reporting external validation and only five providing any assessment of calibration [16, 17].

In this context, the present study was undertaken to address several critical gaps in current diabetes risk prediction research. Specifically, we aimed (1) to develop a high-specificity rule-in diabetes prediction model using Extreme Gradient Boosting (XGBoost) based on routinely collected sociodemographic, anthropometric, clinical, and simple laboratory variables from a nationally representative population of middle-aged and older Korean adults; (2) to conduct rigorous temporal external validation using an independent survey cycle, with comprehensive evaluation of discrimination, calibration, and overall predictive performance, thereby responding to persistent limitations in external validation and calibration reporting among existing AI-based diabetes prediction models [11, 18, 19]; and (3) to examine the clinical utility of a prespecified high-specificity rule-in threshold by prioritizing specificity and positive predictive value, supported by decision-curve analysis and explainable artificial intelligence approaches [20, 21], including SHapley Additive exPlanations, to enhance interpretability and practical relevance for targeted confirmatory testing in community settings.

Methods and materials

Data source and study design

The Korea National Health and Nutrition Examination Survey (KNHANES) is an ongoing, nationally representative, cross-sectional survey designed to assess the health and nutritional status of the non-institutionalized Korean population. Conducted annually by the Korea Disease Control and Prevention Agency, KNHANES employs a stratified, multistage probability sampling design based on geographic area, sex, and age, with primary sampling units selected from across all administrative regions of Korea. Each survey cycle typically includes approximately 8,000–10,000 participants, ensuring coverage of both urban and rural areas nationwide. The survey integrates interviewer-administered questionnaires on sociodemographic characteristics, health behaviors, and medical history with standardized physical examinations and centralized laboratory testing conducted under rigorous quality-control protocols. In the present study, we used publicly available KNHANES data from survey years 2020–2023 and restricted the analysis to participants aged 50 years or older. Individuals with missing or inconsistent information required to define diabetes status or lacking core examination variables essential for model development were excluded. Participants from 2020–2022 were used to construct the development cohort, while those from 2023 were reserved as an independent temporal external validation cohort. Figure 1 provides an overview of the study design, including cohort selection, model development, calibration, and temporal external validation. The detailed selection process of study participants is illustrated in Fig. 2.

Fig. 1.

Fig. 1

Study design and analytical framework for the XGBoost-based diabetes rule-in model. Schematic overview of the data source (KNHANES 2020–2023), eligibility criteria for adults aged ≥ 50 years, A development cohort (2020–2022) and an external validation cohort (2023)

Fig. 2.

Fig. 2

Schematic of the post-hoc calibration procedure for the primary XGBoost model. Raw predicted probabilities from the trained XGBoost model were calibrated using a held-out calibration subset within the development cohort. The post-hoc calibration procedure combined isotonic regression for lower-to-intermediate predicted probabilities, logistic recalibration for the high-risk region, and selective upper-tail temperature scaling for extreme predicted probabilities. The resulting calibration mapping was fixed before external validation and then applied unchanged to the external validation cohort

Diabetes outcome definition and ascertainment

Diabetes status was defined and ascertained according to contemporary diagnostic criteria and standardized KNHANES measurement protocols. Participants were classified as having diabetes if they met any of the following conditions: fasting plasma glucose ≥ 126 mg/dL measured after an overnight fast, glycated hemoglobin (HbA1c) ≥ 6.5% based on centralized laboratory assays, current use of glucose-lowering medication as reported in the health interview, or self-reported physician diagnosis of diabetes. Fasting blood samples were collected and analyzed using standardized procedures with rigorous internal and external quality control, ensuring comparability across survey cycles. Medication use and physician diagnosis were obtained through structured interviewer-administered questionnaires. Individuals with missing values across all diabetes-defining components or with clearly inconsistent information (e.g., discordant reporting incompatible with laboratory findings) were excluded during cohort construction.

Rationale for age restriction (≥ 50 years)

The study population was restricted to adults aged ≥ 50 years based on both epidemiological and physiological considerations relevant to the study objective. In Korea, the prevalence of diabetes increases markedly after midlife; national data from the KNHANES and Diabetes Fact Sheets indicate that prevalence rises from approximately 3.5%–7.4% in individuals in their 40 s to over 25–30% among those aged ≥ 50–60 years. From a physiological perspective, aging is associated with progressive β-cell dysfunction, increased insulin resistance, sarcopenia, and accumulation of visceral adiposity, all of which contribute to a higher baseline risk of diabetes. Given that the present study aimed to develop a high-specificity rule-in model prioritizing positive predictive value and clinical utility, focusing on a population with a higher baseline prevalence was methodologically appropriate. Restricting the cohort to middle aged and older adults also reduces population heterogeneity and improves model stability. Although diabetes onset is increasingly observed at younger ages, inclusion of individuals aged < 50 years would likely alter model operating characteristics. Accordingly, we determined that individuals aged ≥ 50 years represent an appropriate target population for the present high-specificity rule-in model.

Prespecified candidate predictors

Candidate predictors were prespecified a priori based on clinical relevance, evidence from prior epidemiologic and machine-learning studies on diabetes risk, and consistent availability across KNHANES survey cycles. The prespecification strategy was adopted to minimize data-driven variable selection and reduce overfitting, in line with contemporary guidance for prediction model development. Predictors encompassed multiple domains reflecting established and emerging determinants of diabetes. Demographic and socioeconomic variables included age, sex, educational attainment, and household income. Lifestyle factors comprised smoking status, alcohol consumption, physical activity, and selected dietary variables routinely assessed in KNHANES. Self-reported health indicators, including physician-diagnosed comorbidities and oral health measures, were incorporated to capture health status and care-seeking behavior. Anthropometric and clinical measurements included body mass index, waist and neck circumference, systolic and diastolic blood pressure, all obtained using standardized examination protocols. Biochemical predictors consisted of routinely measured serum markers relevant to metabolic and cardiometabolic risk, including lipid profiles, liver enzymes, serum creatinine, uric acid, and blood cell counts. In addition, urine-based markers—urine creatinine, albumin, specific gravity, pH, sodium, and protein—were included to reflect renal function, hydration status, and early metabolic alterations that may precede overt diabetes. Collectively, these prespecified predictors were selected to reflect variables plausibly available in community or screening settings, while avoiding inclusion of highly specialized or costly tests inconsistent with the intended rule-in application of the model.

Missing data and preprocessing

Variable coding was harmonized across survey years to ensure consistency in predictor definitions. Continuous predictors were examined for outliers and distributional skewness; when marked skewness was observed, simple transformations or winsorization were considered while preserving clinically interpretable units. For categorical predictors, sparse categories were collapsed a priori when clinically justified. Participants with missing outcome data were excluded during cohort construction. No participants were excluded from model development or validation because of missing predictor values; instead, missingness in predictors was addressed through the imputation strategy described below. Predictors exhibiting very high levels of missingness in the development cohort were not considered as candidate variables. For model fitting, a unified preprocessing pipeline was implemented and applied identically across all algorithms. Continuous predictors were imputed using the median and standardized to zero mean and unit variance, whereas categorical predictors were imputed using the most frequent category and one-hot encoded, with unseen categories in the validation cohort ignored. This common preprocessing approach ensured that all models—including tree-based methods and logistic regression—were trained on identical inputs and did not rely on native missing-value handling or complete-case analysis. The preprocessing pipeline was fit exclusively on the development cohort (2020–2022) and then applied unchanged to the temporal external validation cohort (2023).

Model development and hyperparameter tuning

Extreme Gradient Boosting (XGBoost) was selected as the primary modeling algorithm because of its strong empirical performance on structured tabular data and its capacity to model complex non-linear relationships and higher-order interactions among predictors. Within the development cohort, data were further partitioned into training and internal validation subsets for model tuning. Hyperparameter optimization was conducted using a limited random search over a prespecified parameter space, including the learning rate, maximum tree depth, number of trees, subsample fraction, column subsample fraction, and regularization parameters. Model selection was guided by joint consideration of discrimination and overall predictive accuracy, as assessed by cross-validated area under the receiver operating characteristic curve (AUROC) and Brier score. A single hyperparameter configuration that provided a favorable trade-off between discrimination and prediction error was selected and subsequently fixed. Using this locked configuration, the final XGBoost model was refitted on the full development cohort. No additional hyperparameter tuning or model modification was performed after this stage, ensuring that all subsequent evaluations reflected genuine model generalization rather than post hoc optimization.

Post-hoc probability calibration

To improve the reliability and clinical interpretability of predicted probabilities, we applied a prespecified multi-step post-hoc calibration strategy within the development cohort. For transparency, the post-hoc calibration procedure is summarized in Fig. 2. Calibration was performed using a held-out calibration subset that was not used for initial model fitting. First, isotonic regression was used to estimate a flexible monotonic mapping between the raw XGBoost-predicted probabilities and observed diabetes outcomes across lower-to-intermediate predicted probability ranges. Second, logistic recalibration was applied in the high-risk region by fitting a logistic regression model with observed diabetes status as the outcome and the logit-transformed raw predicted probabilities as the sole predictor. Third, selective upper-tail temperature scaling was applied to stabilize extreme predicted probabilities while leaving lower-risk predictions unchanged. All calibration parameters were estimated exclusively within the development cohort and fixed prior to validation. The resulting calibration mapping was then applied unchanged to the final XGBoost model refitted on the full development data and to the independent 2023 temporal external validation cohort.

Prespecified rule-in threshold

The model was explicitly designed as a rule-in tool, with primary emphasis on achieving high specificity and positive predictive value. Within the development cohort, model performance was systematically evaluated across a range of candidate probability thresholds. Based on this evaluation, a probability cut-off of 0.637 was selected as the prespecified rule-in threshold, corresponding to very high specificity while preserving a high positive predictive value. This threshold was determined a priori to external validation and was fixed thereafter; it was not modified or re-optimized based on results from the temporal external validation cohort.

Temporal external validation and performance assessment

The locked and calibrated Extreme Gradient Boosting (XGBoost) model, together with the prespecified rule-in threshold, was applied without modification to the independent temporal external validation cohort from 2023 (n = 3,635). Model discrimination was evaluated using the AUROC and average precision. Calibration was assessed using the Brier score, calibration slope and intercept, and the observed-to-expected event ratio. Calibration performance was further examined graphically by plotting observed versus predicted risks across deciles of predicted probability. At the prespecified rule-in threshold, classification performance was summarized using sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and the F1-score. To evaluate the robustness and generalizability of model performance, subgroup analyses stratified by sex, age tertiles, and household income tertiles were conducted, examining both discrimination and calibration across key population subgroups.

Comparator models and incremental value

To benchmark the performance of the primary XGBoost model, three additional supervised learning models—multivariable logistic regression, random forest, and Light Gradient-Boosting Machine (LightGBM)—were developed using the same prespecified set of candidate predictors and an identical preprocessing pipeline. For each comparator model, hyperparameter tuning was performed within the development cohort using a limited random search over prespecified parameter grids, with cross-validated AUROC and Brier score jointly guiding model selection. In addition, an uncalibrated XGBoost comparator model was fitted using the same hyperparameter configuration as the final calibrated XGBoost model but evaluated without application of the hybrid calibration strategy or upper-tail temperature scaling. This comparison allowed isolation of the incremental contribution of the calibration procedure beyond the underlying algorithmic performance. All comparator models were trained exclusively in the development cohort and subsequently applied, without further modification, to the 2023 temporal external validation cohort to generate predicted probabilities. To further quantify the incremental prognostic value of the XGBoost-based approach relative to conventional logistic regression, net reclassification improvement (NRI) and integrated discrimination improvement (IDI) were calculated in accordance with established definitions. Accordingly, the calibrated primary XGBoost model reported in Table 2 should not be interpreted as identical to the uncalibrated XGBoost comparator reported in Table 3. Because the aim of the present study was to develop a clinically interpretable rule-in model based on predicted probability, the calibrated XGBoost model was prespecified as the primary model. The uncalibrated XGBoost model was retained as a comparator to examine how post-hoc calibration affected predicted probabilities and threshold-based performance. Discrimination, overall predictive performance, and rule-in operating characteristics of the primary and comparator models in the validation cohort are summarized in Table 3 and Supplementary Figure S3.

Table 2.

Performance estimates are reported for the 2023 external validation cohort (95% CI)

Accuracy Sensitivity Specificity PPV NPV F1 AUC AP Brier

0.866

(0.856–0.877)

0.394

(0.357–0.431)

0.973

(0.967–0.979)

0.767

(0.724–0.811)

0.877

(0.865–0.888)

0.521

(0.484–0.557)

0.868

(0.854–0.882)

0.646

(0.612–0.680)

0.101

(0.095–0.107)

All metrics are presented with 95% bootstrap confidence intervals. The operating point corresponds to the prespecified rule-in probability threshold (0.637) selected in the main calibrated XGBoost analysis. PPV; positive predictive value, NPV; negative predictive value, AUC; area under the receiver operating characteristic curve, AP; average precision (area under the precision–recall curve)

Table 3.

Comparison of benchmark models at the prespecified rule-in threshold in the external validation cohort

Model Accuracy Sensitivity Specificity PPV NPV F1-score AUC AP Brier NRI vs LR (95% CI) IDI vs LR (95% CI)
Logistic 0.805 (0.793–0.818) 0.727 (0.696–0.760) 0.823 (0.809–0.837) 0.481 (0.451–0.513) 0.930 (0.921–0.940) 0.579 (0.551–0.607) 0.865 (0.852–0.879) 0.646 (0.610–0.679) 0.177 (0.170–0.184) - -
RandomForest 0.851 (0.839–0.862) 0.378 (0.341–0.414) 0.958 (0.950–0.965) 0.668 (0.620–0.714) 0.872 (0.860–0.884) 0.482 (0.446–0.519) 0.859 (0.845–0.873) 0.597 (0.560–0.638) 0.134 (0.129–0.138)

-0.359

(-0.410–-0.307)

-0.098

(-0.108–-0.088)

XGBoost

(uncalibrated comparator)

0.856 (0.844–0.868) 0.255 (0.223–0.289) 0.992 (0.988–0.995) 0.872 (0.827–0.918) 0.855 (0.843–0.867) 0.395 (0.354–0.436) 0.869 (0.856–0.883) 0.665 (0.631–0.699) 0.102 (0.095–0.109) 0.067 (0.039–0.098)

-0.071

(-0.088–-0.054)

LightGBM 0.850 (0.839–0.862) 0.524 (0.488–0.561) 0.924 (0.915–0.933) 0.609 (0.571–0.649) 0.896 (0.885–0.906) 0.563 (0.531–0.596) 0.857 (0.842–0.872) 0.645 (0.611–0.679) 0.118 (0.112–0.125) 0.454 (0.386–0.526) 0.041 (0.025–0.057)

Entries show mean performance in the 2023 external validation cohort (KNHANES 2023), including accuracy, sensitivity, specificity, PPV, NPV, F1-score, AUC, and AP, each with 95% bootstrap confidence intervals. All models were evaluated at the same prespecified rule-in probability threshold (0.637) derived from the main calibrated XGBoost pipeline. The row labeled “XGBoost (uncalibrated comparator)” refers to the uncalibrated comparator model, whereas Table 2 reports the final calibrated primary XGBoost model. PPV; positive predictive value, NPV; negative predictive value, AUC; area under the receiver operating characteristic curve, AP; average precision (area under the precision–recall curve), NRI; net reclassification improvement, LR; logistic regression, IDI; integrated discrimination improvement

Clinical utility (decision curve analysis)

Clinical utility was evaluated using decision-curve analysis in the 2023 temporal external validation cohort. For each model, net benefit was calculated across a clinically relevant range of threshold probabilities (0.01–0.80), reflecting risk levels above which clinicians might reasonably opt to initiate confirmatory diagnostic testing or more intensive clinical management. Net benefit curves for the primary XGBoost model and comparator models were plotted and compared with the default strategies of treating all individuals and treating none. This analysis was used to assess the relative clinical value of the proposed rule-in model across plausible decision thresholds.

Model explainability (SHAP-based analysis)

To enhance model transparency and facilitate clinical interpretability, we conducted a SHAP–based analysis for the final XGBoost model. SHAP values quantify the marginal contribution of each predictor to an individual’s predicted diabetes risk while ensuring consistency with the model’s overall output. Global feature importance was assessed using bar plots of mean absolute SHAP values, which rank predictors according to their average contribution to model predictions. In addition, summary beeswarm plots were generated to visualize the distribution and directionality of SHAP values for the top 15 predictors across the study population. A tabular summary of the leading predictors and their corresponding mean absolute SHAP values is provided in Supplementary Table S4.

Positioning relative to existing screening approaches

Conventional approaches for identifying individuals at risk of diabetes, including fasting glucose measurement, HbA1c testing, and self-monitoring of blood glucose or point-of-care testing in community settings, are widely used and remain effective for detecting current glycemic abnormalities. However, these methods are primarily based on single time-point biochemical measurements and may not fully capture underlying or preclinical risk profiles. In contrast, the present study adopts a multidimensional risk-based approach by integrating sociodemographic, anthropometric, clinical, and behavioral variables to estimate the probability of diabetes. Accordingly, the proposed model is not intended to replace existing screening methods, but rather to complement them by enabling risk stratification and supporting more targeted use of confirmatory biochemical testing.

Statistical analysis

All statistical analyses were performed using Python (version 3.11.7), with established Python packages used for gradient-boosting model development, probability calibration, and decision-curve analysis. Prior to model development, data integrity checks, variable consistency assessments, and descriptive inspections were conducted using SPSS (IBM SPSS Statistics) to verify coding accuracy, ranges, and distributions across survey cycles. Uncertainty in model performance was quantified using nonparametric bootstrap resampling with replacement in the temporal external validation cohort (B = 1,000). For each performance metric, 95% confidence intervals were estimated using the percentile method. A fixed random seed was applied throughout the analysis to ensure reproducibility. Where reported, two-sided p-values were interpreted descriptively rather than as strict thresholds for statistical significance, in keeping with contemporary recommendations for predictive modeling studies.

Data availability and ethics

KNHANES protocols are approved by the institutional review board of the Korea Disease Control and Prevention Agency, and all participants provide written informed consent. The present analysis used de-identified, publicly available data and was conducted in accordance with relevant guidelines and regulations.

Results

Study population

Of the 27,643 KNHANES participants surveyed between 2020 and 2023, we excluded those aged under 50, individuals with missing or indeterminate diabetes status, and those with incomplete core examination data (Fig. 3). The final analytical sample consisted of 13,579 adults aged over 50 years. This sample was divided into a development cohort of 9,944 participants (2020–2022) and a temporal external validation cohort of 3,635 participants (2023). Within the development cohort, 1,848 participants had diabetes and 8,096 did not; in the validation cohort, 670 participants had diabetes and 2,965 did not.

Fig. 3.

Fig. 3

Flowchart of participant selection process from the Korean National Health and Nutrition Examination Survey (KNHANES) 2020–2023. The development cohort comprised KNHANES 2020–2022, and KNHANES 2023 served as temporal external validation cohort

Participant characteristics by diabetes status

Participant characteristics stratified by diabetes status are summarized in Table 1. In both the development and validation cohorts, individuals with diabetes were older than those without diabetes (development cohort: 68.20 ± 8.59 vs 64.90 ± 9.24 years; validation cohort: 67.69 ± 8.30 vs 64.48 ± 8.93 years) and exhibited higher measures of adiposity, including body mass index, waist circumference, and neck circumference. The proportion of men was consistently higher among participants with diabetes across both cohorts. Other laboratory and urinalysis parameters stratified by diabetes status are presented in Supplementary Tables S5–S6, and additional baseline characteristics for the development cohort are provided in Supplementary Table S7.

Table 1.

Basic characteristics of the development and external validation cohorts by diabetes status

Variable Development cohort (n = 9,944) External validation cohort (n = 3,635)
Non-diabetes
(n = 8,096)
Diabetes
(n = 1,848)
Non-diabetes
(n = 2,965)
Diabetes
(n = 670)

Sex (n, %)

Male

Female

3428 (42.3%)

4668 (57.7%)

888 (48.1%)

960 (51.9%)

1218 (41.1%)

1747 (58.9%)

336 (50.1%)

334 (49.9%)

Age (years) 64.90 ± 9.24 68.20 ± 8.59 64.48 ± 8.93 67.69 ± 8.30

Education level

Elementary school or below

Middle school

High school

College or above

2,137 (28.8%)

1,110 (15.0%)

2,397 (32.3%)

1,770 (23.9%)

678 (41.0%)

253 (15.3%)

417 (25.2%)

304 (18.4%)

708 (24.7%)

428 (14.9%)

987 (34.4%)

745 (26.0%)

217 (33.9%)

127 (19.8%)

201 (31.4%)

95 (14.8%)

BMI(kg/m2) 24.09 ± 3.27 24.75 ± 3.32 23.92 ± 3.24 24.79 ± 3.53
Waist circumference (cm) 85.24 ± 9.51 89.24 ± 9.17 84.58 ± 9.45 89.32 ± 9.40
Neck circumference (cm) 34.66 ± 3.18 35.67 ± 3.16 34.53 ± 3.14 35.76 ± 3.23
Alcohol use (frequency) 10.65 ± 37.26 12.75 ± 23.91 9.41 ± 17.33 5.88 ± 6.40
Average monthly household income 401.26 ± 344.02 312.07 ± 289.00 465.73 ± 388.02 375.00 ± 336.41

Self-rated health status

Very good

Good

Average

Bad

Very bad

Don’t know / No response

355 (4.4%)

1,840 (22.7%)

3,631 (44.8%)

1,217 (15.0%)

343 (4.2%)

710 (8.8%)

43 (2.3%)

238 (12.9%)

762 (41.2%)

461 (24.9%)

160 (8.7%)

184 (10.0%)

169 (5.7%)

670 (22.6%)

1,347 (45.4%)

419 (14.1%)

109 (3.7%)

251 (8.5%)

18 (2.7%)

96 (14.3%)

263 (39.3%)

157 (23.4%)

52 (7.8%)

84 (12.5%)

Continuous variables are summarized as mean ± standard deviation and categorical variables as n (%). BMI; body mass index

External validation performance at the prespecified rule-in threshold

At the locked rule-in threshold of 0.637, evaluated in the 2023 temporal external validation cohort, the XGBoost model demonstrated strong rule-in performance, achieving an accuracy of 0.866, specificity of 0.973, positive predictive value of 0.767, and negative predictive value of 0.877, with an F1-score of 0.521. As expected under a high-specificity design, sensitivity was lower (0.394). Overall discrimination and calibration remained good, with an area under the receiver operating characteristic curve of 0.868, average precision of 0.646, and Brier score of 0.101 (Table 2 and Fig. 4).

Fig. 4.

Fig. 4

Confusion matrix at the prespecified rule-in threshold (External 2023). Confusion matrix for predictions in the 2023 external validation cohort at the prespecified rule-in threshold. The model yields few false positives and a high positive predictive value, at the cost of reduced sensitivity, consistent with a rule-in strategy

Discrimination in temporal external validation (AUROC and average precision)

In the temporal external validation cohort, the XGBoost model demonstrated good discrimination between participants with and without diabetes, with an AUROC of 0.868 and an average precision (area under the precision–recall curve) of 0.646 (Fig. 5 and Table 2). Both the receiver operating characteristic and precision–recall curves indicated robust discriminative performance across a broad range of candidate probability thresholds, including the prespecified rule-in region.

Fig. 5.

Fig. 5

Discrimination performance in external validation (ROC and precision–recall curves). (A) Receiver operating characteristic (ROC) curve in the 2023 external validation cohort, summarized by the area under the ROC curve (AUROC = 0.868). (B) Precision–recall (PR) curve in the 2023 external validation cohort, summarized by the area under the PR curve (AUPRC/average precision = 0.646). The no-skill baseline (outcome prevalence = 0.184) is shown to facilitate interpretation under outcome imbalance

Calibration in temporal external validation (Brier score, calibration slope/intercept, and E/O ratio)

Overall calibration in the 2023 temporal external validation cohort was acceptable (Fig. 6A). The Brier score was 0.101 (95% CI, 0.095–0.107), indicating good overall predictive accuracy. The calibration slope was 0.87 (95% CI, 0.81–0.94), suggesting that predicted risks were moderately overdispersed, while the calibration intercept was –0.41 (95% CI, –0.53 to –0.29), reflecting a modest tendency to overestimate average risk. Consistent with these findings, the expected-to-observed (E/O) ratio was 1.16, indicating slight overprediction of the overall event rate. Calibration plots demonstrated reasonable agreement between predicted and observed risks across deciles of predicted probability.

Fig. 6.

Fig. 6

Calibration and clinical utility (calibration curve + decision curve). (A) Calibration plot comparing mean predicted probabilities with observed event rates across deciles; estimates generally track the 45-degree reference line, suggesting acceptable agreement. (B) Decision curve analysis (DCA) comparing the model with treat-all and treat-none strategies. The model shows higher net benefit across clinically relevant threshold probabilities (e.g., 0.01–0.80), including the pre-specified rule-in threshold (0.637), supporting potential clinical utility

Clinical utility in temporal external validation (decision-curve analysis and net benefit)

Decision-curve analysis in the 2023 temporal external validation cohort (Fig. 6B; threshold probability range 0.01–0.80) demonstrated that the XGBoost model yielded non-negative net benefit across the evaluated range and consistently outperformed the treat-none strategy, indicating incremental clinical value. As expected, the treat-all strategy crossed zero net benefit at the observed event prevalence (0.184) and became increasingly unfavorable at higher threshold probabilities. In contrast, the model provided greater net benefit than the treat-all strategy for threshold probabilities exceeding this region. Notably, at the prespecified rule-in threshold of 0.637, the model maintained net benefit above treat-none and substantially exceeded treat-all, supporting its application as a high-specificity rule-in tool to prioritize confirmatory testing.

Subgroup performance and comparator models

Subgroup analyses stratified by sex, age tertiles, and income tertiles demonstrated that rule-in performance was broadly preserved across key population subgroups (Supplementary Tables S2–S3). Across all income strata, specificity consistently exceeded 0.96 and positive predictive value remained high, although sensitivity varied modestly between subgroups. External validation performance of the comparator models—multivariable logistic regression, random forest, LightGBM, and an uncalibrated XGBoost model with identical hyperparameters to the primary model—is summarized in Table 3. In that table, the row labeled “XGBoost (uncalibrated comparator)” refers to the same underlying XGBoost algorithm evaluated without the post-hoc calibration procedure used for the final primary model reported in Table 2. Among these, the uncalibrated XGBoost comparator achieved the highest area under the receiver operating characteristic curve and average precision, the lowest Brier score, and the greatest specificity and positive predictive value at the prespecified rule-in threshold. LightGBM exhibited a more balanced performance profile, characterized by higher sensitivity and F1-scores, whereas logistic regression and random forest demonstrated intermediate performance. Category-free net reclassification improvement indicated modestly improved risk stratification for the XGBoost comparator relative to logistic regression, while the integrated discrimination improvement was close to zero or slightly negative, suggesting limited gains in average risk separation; differences relative to LightGBM were small.

Model explainability (SHAP-based feature attribution)

SHAP-based analyses identified the predictors most influential for the XGBoost model’s risk predictions (Fig. 7; Supplementary Figure S2; Supplementary Table S4). Highly impactful features included urine creatinine, urine specific gravity, urine albumin, total cholesterol and other lipid measures, BMI, waist and neck circumferences, the presence and timing of hypertension and dyslipidemia, liver enzymes, uric acid, white blood cell count, and several dietary intake variables. Higher predicted diabetes risk was associated with patterns that are consistent with established pathophysiological mechanisms, including central adiposity, atherogenic lipid profiles, renal involvement, and systemic inflammation.

Fig. 7.

Fig. 7

SHAP beeswarm plot (Top features). Beeswarm plot of SHAP values indicating individual feature contributions. Urine creatinine, cholesterol, urine albumin/specific gravity, and waist circumference are among the most influential predictors. Red indicates high feature values contributing to risk, blue indicates low values. SHAP; Shapley Additive Explanations, AST; Aspartate aminotransferase, ALT; Alanine aminotransferase, WBC; White blood cell count, RBC; Red blood cell count, Hz; Hertz

Discussion

This study has several novel and distinguishing strengths that merit emphasis. Unlike many prior diabetes prediction studies that prioritize sensitivity for broad screening, we explicitly developed and evaluated a high-specificity rule-in model, addressing a clinically underrepresented but practically important objective—prioritizing confirmatory testing among individuals with a high probability of diabetes in settings where resources are constrained and false positives incur substantial downstream costs. The use of a large, nationally representative survey, together with temporal external validation, enhances the generalizability of our findings to middle-aged and older Korean adults in routine practice [22, 23]. Importantly, we followed contemporary guidance for prediction-model development, including prespecification of the rule-in threshold, comprehensive assessment of calibration and decision-curve net benefit, and transparent reporting, thereby reducing optimism and improving real-world credibility [9, 18]. Beyond discrimination, we demonstrated reliable probability estimates, net clinical benefit, and biologically plausible feature attributions through integrated calibration analyses, decision-curve analysis, and SHAP-based explainability. Finally, the model relies exclusively on variables that are typically available in health surveys and primary-care encounters, facilitating potential implementation without substantial additional costs. Collectively, these features address several gaps highlighted in recent work on AI-based diabetes risk prediction, particularly the scarcity of externally validated and well-calibrated models developed in representative populations with clearly defined clinical use cases [16, 17].

The focus on a high‑specificity rule‑in decision distinguishes this work from most previous KNHANES‑based prediction models, which have primarily optimized overall discrimination or undiagnosed diabetes detection in broader adult populations [24]. In settings where confirmatory laboratory testing, specialist referral slots, or patient time are constrained, prioritizing a small subset of very high‑risk individuals may be more actionable than maximizing sensitivity at the expense of false positives. Our decision‑curve analysis showed that the proposed rule‑in model provides greater net benefit than treating all or no individuals as high risk, and then a conventional risk‑score‑based comparator, across a clinically relevant range of threshold probabilities. In interpreting comparative performance, we treated net reclassification improvement and integrated discrimination improvement as supportive rather than decisive evidence, because these indices can be sensitive to calibration and may overstate incremental value in large samples [20, 21].

Our results are broadly consistent with large‑scale machine‑learning studies showing that gradient‑boosting models offer incremental improvements over traditional logistic regression for diabetes prediction, while also highlighting that much of the predictive signal resides in a relatively small set of anthropometric and metabolic features [9, 25, 26]. In the UK Biobank, Lugner et al. reported an AUROC of 0.90 for 10‑year incident diabetes using a full XGBoost model and 0.88 using only the top ten predictors, with glycated hemoglobin, body mass index, and waist circumference providing the strongest contributions [25]. An Iranian cohort analysis using XGBoost and SHAP achieved an AUC above 0.99 for prevalent type 2 diabetes, again emphasizing fasting blood sugar, markers of adiposity, and selected comorbidities [26]. Compared with these highly optimised models, our AUROC of 0.868 in temporal external validation is slightly lower, which likely reflects our reliance on a finite set of survey variables and the older age structure of the target population; however, the trade‑off for high specificity and rule‑in utility is intentional.

The pattern of influential predictors in our SHAP analyses largely aligns with current epidemiological understanding of diabetes in older Koreans. Waist‑to‑height ratio, body mass index, and blood pressure have been repeatedly associated with incident and prevalent diabetes in KNHANES analyses and related national reports [22, 27, 28]. Resting heart rate has previously been linked to undiagnosed diabetes in Korean adults and incorporated into risk‑score‑based models using the same survey platform [24, 29]. Our finding that lifestyle factors such as smoking, alcohol consumption, and physical activity contribute less strongly than anthropometric indices is also consistent with UK Biobank results and prior systematic reviews of prediction models [9, 25].

From a clinical perspective, the proposed model is best viewed as a triage tool that flags older adults who are most likely to have diabetes and who should therefore be prioritized for confirmatory testing and comprehensive evaluation. Rather than replacing existing diagnostic criteria, the model outputs a probability that can be mapped to action thresholds that reflect local resource constraints and clinical workflows. In settings where additional laboratory testing is feasible, clinicians may choose a lower probability cut‑off to increase sensitivity; conversely, when resources are limited, the high‑specificity rule‑in threshold we pre‑specified may help concentrate efforts on those with the highest likelihood of diabetes.

The relatively low sensitivity observed in our model should be interpreted in the context of its intended use as a high-specificity rule-in tool. Unlike screening-oriented models that aim to maximize sensitivity, our approach prioritizes positive predictive value and minimizes false positives, thereby identifying a smaller subset of individuals with a high likelihood of diabetes. This design is particularly relevant for clinical decision-making scenarios in which confirming a high-risk state is more critical than broadly identifying all potential cases. It is fully acknowledged that diabetes can be effectively screened using widely available and inexpensive biochemical tests such as fasting glucose and HbA1c, which remain the most practical and cost-effective methods for population-level screening. Importantly, the present model is not intended to replace these standard approaches. In fact, acquiring the full set of input variables required for the model may, in certain settings, be more resource-intensive than performing a single biochemical test.

Rather, the proposed model should be understood as a complementary risk stratification and triage tool. By integrating multidimensional information—including sociodemographic, anthropometric, clinical, and behavioral factors—the model may capture underlying or preclinical risk patterns that are not fully reflected by single-point biochemical measurements. This enables identification of individuals with a high underlying probability of diabetes, thereby supporting more targeted and timely confirmatory testing. From an operational perspective, this approach may be particularly valuable in settings where such multidimensional data are already routinely available, such as health surveys or electronic health record–based systems. In these contexts, the model can be applied without substantial additional cost and may enhance the efficiency of clinical decision-making by prioritizing high-risk individuals while avoiding unnecessary testing in low-risk groups.

This study also has limitations. KNHANES is cross‑sectional and does not capture incident diabetes, so our model predicts prevalent rather than future diabetes; prospective validation is needed before the model can be used for long‑term risk prediction. Misclassification of diabetes status is possible because we relied on single‑time‑point measurements and self‑reported diagnosis or treatment, and we could not distinguish between type 1 and type 2 diabetes. Some potentially informative predictors, such as detailed medication patterns or family history beyond first‑degree relatives, were not available or could not be harmonized across survey cycles. As with other machine‑learning models, performance may degrade if the underlying population or care patterns change over time, underscoring the need for ongoing monitoring and recalibration [9]. Finally, we handled missing predictor values using single imputation with median and most frequent values rather than multiple imputation; although simple imputation is widely used and often performs comparably to more complex methods in prediction studies, residual bias related to unmodelled missingness mechanisms cannot be excluded.

In addition, a common limitation of machine learning models, including gradient boosting algorithms such as XGBoost, is their “black box” nature, which can limit transparency and hinder clinical interpretability. Unlike traditional regression models, the complex and non-linear structure of these algorithms makes it difficult to directly trace how individual predictors contribute to model outputs. To mitigate this limitation, we applied SHAP, which provide both global and individual-level insights into feature contributions. The SHAP-based analyses in our study identified clinically plausible predictors, such as adiposity measures, lipid profiles, and renal markers, supporting the internal consistency and face validity of the model. However, it is important to note that SHAP and similar post hoc explainability methods do not fully resolve the inherent opacity of machine learning models. Therefore, caution is warranted in interpreting model outputs, and further work is needed to enhance transparency, validation, and clinical integration of such approaches.

Future research could extend this work by examining incident diabetes and downstream clinical outcomes in prospective cohorts, by assessing the impact of implementing the rule‑in model on testing patterns and time to diagnosis, and by exploring hybrid approaches that combine survey‑based predictions with electronic health record, imaging, or biosignal data. In addition, qualitative and implementation‑science studies are needed to understand how clinicians and patients perceive algorithm‑based risk communication in older adults and how such tools can be integrated into existing Korean screening programs.

From an implementation perspective, the practical utility of the proposed model depends on the availability and integration of multidimensional data within existing healthcare systems. In environments where such data are routinely collected—such as national health surveys, periodic health examinations, or electronic health record–based systems—the model can be deployed with minimal additional burden, functioning as an automated risk stratification layer. In contrast, in settings lacking structured data infrastructure, direct biochemical screening may remain more feasible. Therefore, the model should be viewed as a context-dependent tool whose value is maximized when integrated into data-rich healthcare environments, where it can support scalable, data-driven prioritization of high-risk individuals.

In conclusion, the present findings suggest that an XGBoost-based rule-in model trained on recent KNHANES data can identify middle-aged and older Korean adults with a very high probability of diabetes using routinely available information. When coupled with careful local calibration, prospective validation, and thoughtful implementation strategies, such models may support more targeted and resource-efficient approaches to diabetes identification and management in ageing populations.

Supplementary information

Additional file 1 . (1.1MB, docx)

Author contributions

Soo Myeong Kim and Jung Min Cho jointly conceived and designed the study, developed the analytical framework, and played leading roles in the statistical analyses and interpretation of the results. All authors critically reviewed the manuscript, provided intellectual input, and approved the final version.

Funding

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2025—2144296851582065930001) and was supported by the Daegu Haany University Regional Innovation System & Education (RISE) Glocal project program [Industry-Academia Collaborative R&D Projects in Specialized Fields] through the Gyeongbook RISE center, funded by the Ministry of Education (MOE) and the Gyeongsangbookdo, Republic of Korea. (2026-RISE—15—110).

Data availability

The data described in the manuscript, code book, and analytic code will be made available upon reasonable request [Korea Disease Control and Prevention Agency at https://chs.kdca.go.kr/].

Declarations

Ethics approval

This study was approved by the Institutional Review Board of the Korea Centers for Disease Control and Prevention (KCDC) in accordance with the Declaration of Helsinki. Written informed consent was obtained before Korea National Health and Nutrition Examination Survey participation, and all data analyses were conducted in accordance with the guidelines and regulations of the KCDC.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Ong KL, Stafford LK, McLaughlin SA, Boyko EJ, Vollset SE, Smith AE, et al. Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the global burden of disease study 2021. The Lancet. 2023;402(10397):203–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.International Diabetes Federation. IDF Diabetes Atlas, 10th edn. Brussels, Belgium: International Diabetes Federation; 2021.
  • 3.Beagley J, Guariguata L, Weil C, Motala AA. Global estimates of undiagnosed diabetes in adults. Diabetes Res Clin Pract. 2014;103(2):150–60. [DOI] [PubMed] [Google Scholar]
  • 4.Collins GS, Mallett S, Omar O, Yu LM. Developing risk prediction models for type 2 diabetes: a systematic review of methodology and reporting. BMC Med. 2011;9(1):103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Abbasi A, Peelen LM, Corpeleijn E, van der Schouw YT, Stolk RP, Spijkerman AM, et al. Prediction models for risk of developing type 2 diabetes: systematic literature search and independent external validation study. BMJ. 2012;345:e5900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lindstrom J, Tuomilehto J. The diabetes risk score: a practical tool to predict type 2 diabetes risk. Diabetes Care. 2003;26(3):725–31. [DOI] [PubMed] [Google Scholar]
  • 7.Ha KH, Lee YH, Song SO, Lee JW, Kim DW, Cho KH, et al. Development and validation of the Korean diabetes risk score: a 10-year national cohort study. Diabetes Metab J. 2018;42(5):402–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016.
  • 9.Silva KD, Lee WK, Forbes A, et al. Use and performance of machine learning models for type 2 diabetes prediction in community settings: a systematic review and meta‑analysis. Int J Med Inform. 2020;143:104268. [DOI] [PubMed] [Google Scholar]
  • 10.Niculescu-Mizil A, Caruana R. Predicting good probabilities with supervised learning. In: Proceedings of the 22nd International Conference on Machine Learning; 2005.
  • 11.Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17(1):230 (for the topic group ‘evaluating diagnostic tests and prediction models’ of the STRATOS initiative). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30.
  • 13.Hasan R, Dattana V, Mahmood S, Hussain S. Towards transparent diabetes prediction: combining AutoML and explainable AI for improved clinical insights. Information. 2024;16(1):7. [Google Scholar]
  • 14.Fazakis N, Kocsis O, Dritsas E, Alexiou S, Fakotakis N, Moustakas K. Machine learning tools for long-term type 2 diabetes risk prediction. IEEE Access. 2021;9:103737–57. [Google Scholar]
  • 15.Kiran M, Xie Y, Anjum N, Ball G, Pierscionek B, Russell D. Machine learning and artificial intelligence in type 2 diabetes prediction: a comprehensive 33-year bibliometric and literature analysis. Front Digit Health. 2025;7:1557467. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mohsen F, Al-Absi HR, Yousri NA, El Hajj N, Shah Z. A scoping review of artificial intelligence-based methods for diabetes risk prediction. NPJ Digit Med. 2023;6(1):197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang SCY, Nickel G, Venkatesh KP, Raza MM, Kvedar JC. AI-based diabetes care: risk prediction models and implementation concerns. npj Digit Med. 2024. 10.1038/s41746-024-01034-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Piovani D, Sokou R, Tsantes AG, Vitello AS, Bonovas S. Optimizing clinical decision making with decision curve analysis: insights for clinical investigators. Healthcare (Basel). 2023;11(14):2053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Pencina MJ, D’Agostino RB Sr, Steyerberg EW. Extensions of net reclassification improvement calculations to measure usefulness of new biomarkers. Stat Med. 2011;30(1):11–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Pencina MJ, D’Agostino RB Sr, D’Agostino RB Jr, Vasan RS. Evaluating the added predictive ability of a new marker: from area under the ROC curve to reclassification and beyond. Stat Med. 2008;27(2):157–72. [DOI] [PubMed] [Google Scholar]
  • 22.Kim BY, Kim H, Kim BS, et al. Diabetes mellitus in the elderly adults in Korea: based on data from the Korea national Health and nutrition examination survey. Diabetes Metab J. 2023. 10.4093/dmj.2023.0041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Kweon S, Kim Y, Jang MJ, et al. Korea national health and nutrition examination survey: 20th anniversary, accomplishments and future directions. Epidemiol Health. 2021;43:e2021025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Choi SG, Oh M, Park DH, et al. Comparisons of the prediction models for undiagnosed diabetes between machine learning versus traditional statistical methods. Sci Rep. 2023;13:13101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lugner M, Rawshani A, Helleryd E, Eliasson B. Identifying top ten predictors of type 2 diabetes through machine learning analysis of UK Biobank data. Sci Rep. 2024;14:2102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Rafie Z, Sedaghat Talab M, Ebrahim Zadeh Koor B, et al. Leveraging XGBoost and explainable AI for accurate prediction of type 2 diabetes. BMC Public Health. 2025;25:3688. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Rhee EJ, Kim JH, Lee EY, et al. Diabetes fact sheet in Korea 2021. Diabetes Metab J. 2022;46:417–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Jee YH, Bae J, Kim JH, et al. Diabetes fact sheets in Korea 2024. Diabetes Metab J. 2025. 10.4093/dmj.2024.0818. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Park DH, Cho W, Lee YH, Ji S, Jeon JJ. The predicting value of resting heart rate to identify undiagnosed diabetes in Korean adults: Korea national health and nutrition examination survey. Epidemiol Health. 2022;44:e2022009. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Additional file 1 . (1.1MB, docx)

Data Availability Statement

KNHANES protocols are approved by the institutional review board of the Korea Disease Control and Prevention Agency, and all participants provide written informed consent. The present analysis used de-identified, publicly available data and was conducted in accordance with relevant guidelines and regulations.

The data described in the manuscript, code book, and analytic code will be made available upon reasonable request [Korea Disease Control and Prevention Agency at https://chs.kdca.go.kr/].


Articles from Diabetology & Metabolic Syndrome are provided here courtesy of BMC

RESOURCES