Skip to main content
Frontiers in Endocrinology logoLink to Frontiers in Endocrinology
. 2026 Sep 3;17:1905144. doi: 10.3389/fendo.2026.1905144

Machine-learning-based prediction model of type 2 diabetes using liver enzymes: a cross-sectional study

Yaru Bi 1,2, Xiaojie Yuan 2, Yufeng Wang 3, Chenglin Sun 1,2,*, Suyan Tian 4,*
PMCID: PMC13581821  PMID: 42755850

Abstract

Introduction

Although observational studies have established associations between three major liver enzymes (alanine aminotransferase (ALT), aspartate aminotransferase (AST), and gamma-glutamyl transferase (GGT)) and type 2 diabetes (T2D), their ability to improve T2D identification using machine learning (ML) approaches remains underexplored. This study aimed to develop and validate ML-based diagnostic models for identifying T2D by incorporating these liver enzymes.

Methods

Data from two independent cohorts were analyzed: the US National Health and Nutrition Examination Survey (N = 15,528) and a Chinese health examination database (N = 4,952). Twelve demographic and biochemical features were combined with liver enzymes to train eight ML models, including logistic regression, support vector machine, random forest, K-nearest neighbors, classification and regression trees, gradient boosting decision tree, LightGBM, and XGBoost. Model performance was assessed using standard metrics including accuracy, precision, recall, F1-score, and area under the curve (AUC).

Results

Incorporating liver enzymes consistently improved model performance across the algorithms. The XGBoost model showed particularly strong performance, with baseline metrics (accuracy = 0.696, precision = 0.314, F1 = 0.450, AUC = 0.803) increasing to 0.743, 0.346, 0.468, and 0.814, respectively, after inclusion of liver enzymes. Comparable improvements were observed across algorithms and in both the US and Chinese cohorts.

Conclusions

ML models incorporating routinely measured liver enzymes improve T2D identification across US and Chinese datasets. These findings indicate the potential utility of liver enzymes as accessible adjunctive indicators for diabetes risk stratification and improved T2D identification.

Keywords: liver enzymes, type 2 diabetes, machine learning, prediction model, NHANES

Introduction

Diabetes mellitus is a chronic metabolic disorder characterized by insulin resistance and beta-cell dysfunction and represents a major global health challenge. Current estimates indicate that more over 10% of adults worldwide had diabetes in 2021, with projections suggesting an increase to 783 million cases by 2045 (1). Notably, type 2 diabetes (T2D) represents 90% of all diabetes cases. This condition markedly increases the risk of severe complications including cardiovascular disease, renal failure, vision impairment, lower-extremity amputations, and premature mortality (2, 3). The economic burden is also substantial, with global diabetes-related expenditures projected to reach $1.054 trillion by 2045 (1). These concerning trends underscore the urgent need for early identification of T2D.

Machine learning (ML), a core component of artificial intelligence, uses algorithmic approaches to extract meaningful patterns from complex datasets for predictive modeling (4). Recent advances in ML methods have transformed disease prediction, including T2D, by enabling early risk stratification and progression monitoring (5, 6). Existing T2D prediction models have incorporated diverse data types, including demographic characteristics such as sex, age, and ethnicity; lifestyle factors such as physical activity and smoking status; and basic clinical parameters (7). More sophisticated models have integrated biochemical markers such as fasting glucose and lipid profiles (8–11), genetic predisposition assessed using polygenic risk scores (12), and environmental exposure data such as chemical contaminants and heavy metals (13, 14).

Growing epidemiological evidence indicates a potential association between liver enzymes, particularly alanine aminotransferase (ALT), aspartate aminotransferase (AST), and gamma-glutamyl transferase (GGT), and T2D development (15–20). A recent cross-sectional analysis of the Azar cohort study (N = 14,865 participants aged 35–70 years) found significantly higher serum ALT, AST, and GGT levels in prediabetic and diabetic patients than in controls (P < 0.05). Multivariable logistic regression further revealed a dose-response relationship for all liver enzymes.

Although machine learning models achieve excellent T2D prediction, liver enzymes (ALT, AST, GGT) have not been fully leveraged despite their metabolic importance. We address this gap by developing and comparing multiple ML algorithms, including logistic regression (LR), support vector machines (SVM), K-nearest neighbors (KNN), classification and regression trees (CART), random forest (RF), gradient boosting decision trees (GBDT), LightGBM, and XGBoost, to quantify the incremental value of liver enzymes beyond established clinical features. The second aim of this study is to improve early screening approaches for T2D.

Materials and methods

Experimental data

This study analyzed two distinct datasets: (1) the US National Health and Nutrition Examination Survey (NHANES) spanning 2011–2018 and (2) a health examination study from Northeast China. The NHANES protocol was approved by the NCHS Ethics Review Board, and the Chinese protocol was approved by the Ethics Committee of the First Hospital of Jilin University (AF–IRB–032–06). All participants provided their written informed consents to participate in this study.

NHANES Cohort: We applied selection criteria consistent with our prior research (17). From the NHANES demographic, anthropometric, questionnaire, and biochemical data, we excluded participants aged <20 years; individuals with missing liver enzyme measurements; patients with hepatitis B/C; pregnant women; cancer patients; heavy alcohol consumers (>30 g/day [men] or >20 g/day [women]); and those diagnosed with diabetes before age 30 (to reduce potential misclassification of type 1 diabetes). The final analytical sample included 15,528 eligible participants.

In this cohort, non-alcoholic fatty liver disease (NAFLD) was defined using the Fatty Liver Index (FLI≥60), calculated from waist circumference (WC), triglycerides (TG), body mass index (BMI), and GGT. Diabetes was defined by meeting any of the following criteria: fasting plasma glucose (FPG) ≥7.0 mmol/L, 2-hour oral glucose tolerance test (OGTT) glucose ≥11.1 mmol/L, glycated hemoglobin (HbA1c) ≥6.5%, or current use of hypoglycemic medication or insulin.

Chinese Cohort: We recruited 5,300 adults (≥20 years) undergoing routine health examinations at the First Hospital of Jilin University (November 2022–March 2023). Demographic characteristics, anthropometric measurements, laboratory test results, and abdominal ultrasound data (for fatty liver diagnosis) were collected. Exclusion criteria included: non-T2D diagnoses; excessive alcohol intake (>30 g/day [men] or >20 g/day [women]); missing liver enzyme data; significant liver injury (ALT>250 IU/L, AST>200 IU/L, GGT>300 IU/L); and renal insufficiency (serum creatinine >333 μmol/L). After exclusions, 4,952 participants remained for analysis. In this cohort, T2D was defined based on FPG ≥7.0 mmol/L, HbA1c ≥6.5%, self-reported physician-diagnosed T2D, or current use of hypoglycemic medication or insulin.

Data pre-processing

Before model development, several pre-processing steps were performed. Features with >15% missing data, including smoking status and low-density lipoprotein cholesterol, were excluded from further analysis. Features with ≤15% missingness, such as blood pressure, high-density lipoprotein cholesterol (HDL-C), and drinking status, were imputed using multiple imputation by chained equations (MICE) with 10 iterations. In addition, binary features, including sex (female/male), drinking status (no/yes), and NAFLD (no/yes), were encoded as 0/1 numeric variables, whereas continuous variables were standardized using the Z-score normalization. After pre-processing, 12 features were included for model development: sex, age, BMI, WC, systolic blood pressure (SBP), diastolic blood pressure (DBP), drinking status, serum creatinine (Scr), total cholesterol (TC), TG, HDL-C, NAFLD, and the three liver enzymes (ALT, AST, GGT).

Model construction and evaluation

First, the NHANES dataset was divided into a training set and an internal validation set at a 7:3 ratio. A grid search strategy was combined with five-fold cross-validation within the training set to identify the optimal hyperparameters for each algorithm, while the Chinese dataset was used for external validation. The analytical strategy was then reversed: models were trained on the Chinese dataset and externally validated using NHANES data to assess model generalizability (Figure 1).

Figure 1.

Two flowchart panels labeled (a) and (b) show machine learning model development and validation workflows. In panel (a), the NHANES dataset is split into a 70% training set and a 30% internal validation set. Within the training set, five-fold cross-validation combined with grid search optimizes hyperparameters for eight classifiers: LR, CART, RF, GBDT, SVM, KNN, LightGBM, and XGBoost. Model performance is evaluated on both the NHANES internal validation set and the Chinese external dataset. Panel (b) mirrors this workflow, where the Chinese dataset is used as the training set and the NHANES dataset as the external validation set.

Flowchart of the machine learning model development and validation. (a) The NHANES dataset used as the training set, and the Chinese dataset as the external validation set; (b) the Chinese dataset used as the training set, and the NHANES dataset as the external validation set.

Eight ML algorithms were evaluated: logistic regression (LR), support vector machine (SVM), k-nearest neighbors (KNN), classification and regression trees (CART), random forest (RF), gradient-boosted decision trees (GBDT), LightGBM, and XGBoost. LR is a generalized linear model that uses the sigmoid function to convert linear risk scores into probabilities (ranging from 0 to 1) for event prediction. SVM determines an optimal hyperplane that maximizes the margin between distinct classes, allowing effective classification. KNN is a distance-based algorithm that classifies new samples based on the majority class among their k nearest neighbors. CART is a nonparametric decision tree method that recursively partitions data into binary branches to produce classification rules.

RF is as an ensemble model that constructs multiple decision trees trained on random subsets of features and samples, with final predictions obtained through majority voting. GBDT is a boosting method that sequentially trains weak learners by reweighting misclassified samples to enhance model performance. XGBoost and LightGBM are advanced, computationally efficient derivatives of GBDT. XGBoost improves computational speed and predictive accuracy, whereas LightGBM optimizes memory use and reduces training time. The hyperparameters explored are listed in Supplementary Table 1.

Model performance was evaluated using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC), with AUC prioritized as the primary indicator of predictive performance. Calibration was additionally evaluated using calibration curves, the calibration intercept and slope, and the Brier score (21). Decision curve analysis (DCA) was used to assess the clinical applicability and net benefit of the predictive models across different threshold probabilities (22).

To improve interpretability, we applied Shapley Additive exPlanations (SHAP) to the optimal model. SHAP quantifies the contribution of individual features to model predictions and reveals both the magnitude and direction (positive or negative) of their influence. Features with larger SHAP values were considered more influential in the model’s decision-making process.

Comparison with risk score

We compared our models with the Chinese Diabetes Risk Score (CDRS) (23), which included six variables: age, sex, BMI, WC, SBP, and family history of diabetes. Because family history of diabetes was unavailable in the Chinese cohort, we used a modified CDRS that excluded this variable. The optimal cutoff for the modified CDRS was identified using the Youden index.

Programming language

All ML models were implemented in Python (version 3.13.0). The XGBoost and LightGBM models were developed using the “xgboost” (version 3.3.0) and “lightgbm” (version 4.7.0) packages, respectively, while the other algorithms were built using “scikit-learn” (version 1.5.2). SHAP analysis was conducted with the “shap” (version 0.52.0) package. Fixed random seed (2026) was set for model training and cross-validation to ensure reproducibility of the results.

This study complies with the TRIPOD+AI reporting guideline (24), and the corresponding TRIPOD+AI checklist is provided as Supplementary Material.

Results

Baseline characteristics of the participants

The study included 15,528 participants from the NHANES dataset and 4,952 from the Chinese dataset, of whom 2,426 (15.62%) and 447 (9.03%) were identified as having T2D, respectively. Baseline characteristics of participants in both datasets are presented in Supplementary Table 2. Briefly, individuals with T2D had significantly higher values for age, BMI, WC, SBP, ALT, AST, and GGT than the non-diabetes group in both cohorts.

The NHANES dataset served as the training set for developing eight ML models based on 12 features: sex, age, BMI, WC, SBP, DBP, NAFLD status, Scr, TC, TG, HDL-C, and drinking status. The performance of these models on the Chinese external validation set is presented in Table 1. In addition, the inclusion of ALT, AST, and GGT improved the performance of T2D detection. Figure 2a presents the receiver operating characteristic (ROC) curves of these models. DeLong’s test indicated that the AUC improvements were statistically significant (P < 0.05) except for CART and KNN. The XGBoost model yielded the best performance. Without liver enzymes, the XGBoost model achieved the following metrics: accuracy = 0.709, precision = 0.207, recall = 0.785, F1-score = 0.328, and AUC = 0.802. After inclusion of liver enzymes, these metrics changed to: accuracy = 0.707, precision = 0.207, recall = 0.794, F1-score = 0.328, and AUC = 0.824. Supplementary Tables 3, 4 provide the confusion matrices, specificity, positive predictive value, and negative predictive value for the XGBoost models in the Chinese external validation set.

Table 1.

Performance metrics of the prediction models validated on the Chinese cohort.

Model Accuracy Precision Recall F1-score AUC (95% CI) DeLong’s test P
LR 0.790 0.231 0.570 0.329 0.792 (0.772, 0.811) 0.047
LR (+enzymes) 0.786 (−0.4%) 0.235 (+0.4%) 0.609 (+3.9%) 0.339 (+1%) 0.798 (0.779, 0.816)
CART 0.565 0.154 0.855 0.262 0.736 (0.714, 0.758) 0.459
CART (+enzymes) 0.610 (+4.5%) 0.162 (+0.8%) 0.794 (−6.1%) 0.269 (+0.7%) 0.739 (0.717, 0.761)
RF 0.751 0.220 0.689 0.333 0.800 (0.783, 0.819) 0.001
RF (+enzymes) 0.714 (−3.7%) 0.207 (−1.3%) 0.765 (+7.6%) 0.326 (−0.7%) 0.812 (0.795, 0.831)
GBDT 0.758 0.225 0.689 0.339 0.802 (0.784, 0.821) <0.001
GBDT (+enzymes) 0.741 (−1.7%) 0.220 (−0.5%) 0.736 (+4.7%) 0.339 (+0%) 0.820 (0.804, 0.839)
SVM 0.733 0.163 0.472 0.242 0.638 (0.608, 0.669) <0.001
SVM (+enzymes) 0.812 (+7.9%) 0.239 (+7.6%) 0.494 (+2.2%) 0.322 (+8%) 0.710 (0.683, 0.739)
KNN 0.719 0.186 0.624 0.286 0.760 (0.741, 0.781) 0.106
KNN (+enzymes) 0.711 (−0.8%) 0.190 (+0.4%) 0.678 (+5.4%) 0.297 (+1.1%) 0.773 (0.753, 0.795)
LightGBM 0.806 0.247 0.562 0.343 0.806 (0.788, 0.824) <0.001
LightGBM (+enzymes) 0.704 (−10.2%) 0.205 (−4.2%) 0.787 (+22.5%) 0.325 (−1.8%) 0.821 (0.803, 0.839)
XGBoost 0.709 0.207 0.785 0.328 0.802 (0.784, 0.820) <0.001
XGBoost (+enzymes) 0.707 (−0.2%) 0.207 (+0%) 0.794 (+0.9%) 0.328 (+0%) 0.824 (0.806, 0.842)

Figure 2.

Two side-by-side ROC curve panels compare external-validation performance for eight machine learning models. Panel (a) evaluates model generalizability from the NHANES to the Chinese cohort; panel (b) evaluates generalizability from the Chinese dataset to the NHANES cohort. Each panel contains multiple colored model curves, and the legend lists each algorithm together with its corresponding AUC value and confidence interval. The dashed diagonal reference line represents random prediction. The horizontal x-axis is 1-Specificity, and the vertical y-axis is Sensitivity.

Receiver operating characteristic (ROC) curves of machine learning models from the Chinese external validation (a) and the NHANES external validation (b).

For the XGBoost model, the calibration curve showed satisfactory calibration in the Chinese external validation set (Brier score: 0.07, calibration intercept: -0.21, calibration slope: 1.19) (Figure 3a). The DCA curve indicated that the XGBoost model achieved greater net benefit than the treat-all or treat-none strategies across a clinically meaningful range of threshold probabilities (0.1 to 0.35). Furthermore, the XGBoost model incorporating liver enzymes exhibited a slightly higher net benefit than the model without liver enzymes (threshold range: 0.2–0.35) (Figure 4a). SHAP analysis indicated that age had the largest predictive weight for T2D detection, followed by WC and GGT. Among the liver enzymes, GGT, AST, and ALT ranked 3rd, 5th, and 10th in predictive contribution, respectively. Within this model, liver enzymes showed a greater contribution to T2D detection than BMI, a well-established marker linked to T2D (Figure 5a).

Figure 3.

Two XGBoost calibration-plot panels display observed outcome probabilities against predicted probabilities. Panel (a), titled “NHANES to Chinese external validation”, shows the calibration curve in the Chinese external validation set. Panel (b), titled “Chinese to NHANES external validation”, presents the calibration curve in the NHANES external validation set. Each plot includes a solid blue calibration line, light-blue shaded 95% bootstrap confidence intervals, scatter points for 10 quantile‑group predictions, and a grey dashed perfect-calibration diagonal. The x‑axis shows mean predicted probability; the y‑axis shows observed probability.

Calibration curve of the XGBoost models from the Chinese external validation (a) and the NHANES external validation (b).

Figure 4.

Two side‑by‑side decision-curve-analysis plots show net benefit against threshold probability for cross‑cohort external‑validation analyses. Panel (a) corresponds to NHANES‑to‑Chinese external validation; panel (b) corresponds to Chinese‑to‑NHANES external validation. A solid blue line represents the prediction model incorporating AST, ALT, and GGT, and a dashed red line represents the model excluding these liver‑enzyme features. A solid black line denotes the “All‑intervention” strategy, and a dotted black line denotes the “None‑intervention” strategy. The horizontal x‑axis shows threshold probability, and the vertical y‑axis shows net benefit.

Decision curve analysis of the XGBoost models from the Chinese external validation (a) and the NHANES external validation (b).

Figure 5.

Panel (a) shows a beeswarm plot of SHAP values from an XGBoost model for the NHANES internal test dataset, ranking features such as age, waist circumference, and GGT by their contribution. Panel (b) displays a similar beeswarm plot for the Chinese internal test dataset, ranking age, GGT, and TG as top features. Both plots use a blue-to-red color gradient to indicate low to high feature values, with SHAP value ranges shown along the x-axis.

Shapley Additive exPlanations (SHAP) values for the XGBoost models features from the NHANES internal test (a) and Chinese internal test (b).

Subsequently, eight machine learning models were trained on the Chinese dataset using twelve baseline features. After ALT, AST, and GGT were incorporated into the models, performance improved in the NHANES external validation set except for CART, and the improvements were statistically significant (DeLong’s test P < 0.05) (Table 2, Figure 2b). The performance gains were particularly evident for ensemble methods. For the GBDT model, consistent gains were observed across all evaluation metrics. Accuracy increased from 0.709 to 0.717 (+0.008), and precision rose from 0.316 to 0.328 (+0.012). The recall rate increased from 0.737 to 0.773 (+0.036), accompanied by an F1-score improvement from 0.442 to 0.461(+0.019). Most notably, AUC increased from 0.792 to 0.808 (+0.016), indicating enhanced discriminative ability.

Table 2.

Performance metrics of the prediction models validated on the NHANES cohort.

Model Accuracy Precision Recall F1-score AUC (95% CI) DeLong’s test P
LR 0.743 0.326 0.602 0.423 0.767 (0.758, 0.777) <0.001
LR (+enzymes) 0.762 (+1.9%) 0.340 (+1.4%) 0.559 (−4.3%) 0.423 (0%) 0.775 (0.765, 0.784)
CART 0.562 0.244 0.859 0.380 0.709 (0.699, 0.719) 0.002
CART (+enzymes) 0.554 (−0.8%) 0.243 (−0.1%) 0.876 (+1.7%) 0.381 (+0.1%) 0.698 (0.688, 0.708)
RF 0.631 0.279 0.863 0.422 0.794 (0.785, 0.803) <0.001
RF (+enzymes) 0.667 (+3.6%) 0.300 (+2.1%) 0.849 (−1.4%) 0.443 (+2.1%) 0.810 (0.801, 0.818)
GBDT 0.709 0.316 0.737 0.442 0.792 (0.783, 0.800) <0.001
GBDT (+enzymes) 0.717 (+0.8%) 0.328 (+1.2%) 0.773 (+3.6%) 0.461 (+1.9%) 0.808 (0.799, 0.816)
SVM 0.641 0.237 0.584 0.337 0.651 (0.640, 0.662) 0.090
SVM (+enzymes) 0.649 (+0.8%) 0.242 (+0.5%) 0.586 (+0.2%) 0.343 (+0.6%) 0.660 (0.648, 0.673)
KNN 0.619 0.270 0.840 0.408 0.771 (0.763, 0.780) 0.030
KNN (+enzymes) 0.607 (−1.2%) 0.265 (−0.5%) 0.853 (+1.3%) 0.404 (−0.4%) 0.778 (0.769, 0.788)
LightGBM 0.678 0.300 0.799 0.437 0.798 (0.790, 0.806) <0.001
LightGBM (+enzymes) 0.745 (+6.7%) 0.344 (+4.4%) 0.698 (−10.1%) 0.461 (+2.4%) 0.810 (0.801, 0.818)
XGBoost 0.696 0.314 0.797 0.450 0.803 (0.794, 0.811) <0.001
XGBoost (+enzymes) 0.743 (+4.7%) 0.346 (+3.2%) 0.723 (−7.4%) 0.468 (+1.8%) 0.814 (0.805, 0.822)

The XGBoost model showed similar improvements. Accuracy increased from 0.696 to 0.743 (+0.047), while precision improved by 0.032 (0.314 to 0.346). A modest gain was observed for F1-score (0.450 to 0.468, + 0.018). The AUC increased from 0.803 to 0.814 (+0.011), and this improvement was statistically significance (DeLong’s test P < 0.001), maintaining XGBoost as the highest-performing model. Supplementary Table 3 provides the confusion matrix obtained from the NHANES external validation cohort.

Calibration curves indicated moderate agreement (Brier score: 0.12, calibration intercept: 0.53, calibration slope: 0.87) (Figure 3b). The DCA curve for the XGBoost model incorporating liver enzymes lies above that of the model without liver enzymes (threshold range: 0.25–0.45), indicating greater net clinical benefit (Figure 4b). SHAP analyses identified the leading features in descending order of contribution as age, GGT, and TG. Among the liver enzymes, GGT, AST, and ALT ranked 2nd, 8th, and 9th, respectively (Figure 5b). Within this fitted model, liver enzymes had larger mean absolute SHAP values than BMI, a well-established correlate of T2D.

We compared model performance with the modified CDRS. In the Chinese external validation cohort, the modified CDRS achieved an AUC of 0.776 (95% CI: 0.756–0.795) at the optimal cutoff of 28.5, with a sensitivity of 80.3% and specificity of 63.3%. In comparison, our XGBoost model showed superior performance with an AUC of 0.824 (95% CI: 0.806–0.842). The DeLong’s test indicated a statistically significant difference between the two AUCs (P < 0.05). Similar findings were observed in the NHANES external validation cohort, where the modified CDRS achieved an AUC of 0.780 (95%CI: 0.771–0.788) compared with 0.814 (95% CI: 0.805–0.822) for our XGBoost model (P < 0.05) (Supplementary Table 5).

Discussion

This study demonstrated that incorporating liver enzymes improved the performance of models for identifying T2D across different populations. The consistent improvements in model metrics and robust performance during external validation suggest that liver enzymes provide incremental predictive value beyond routine clinical features. This finding is particularly notable because liver enzymes maintained their predictive power even after adjustment for established clinical correlates of T2D, including BMI, WC, and lipid profiles.

Incorporating liver enzymes improved model predictive performance for T2D identification and conferred a favorable net benefit across clinically relevant threshold probabilities. Although inclusion of liver enzymes slightly increased model complexity, these routinely measured, low-cost laboratory tests impose no additional clinical or economic burden. Furthermore, liver enzymes are nonspecific markers influenced by obesity and hepatic steatosis. During model construction, conventional metabolic factors and fatty liver were already included in the models. The improved predictive capacity suggests that liver enzymes provide additional information beyond existing metabolic factors and hepatic steatosis.

The superior performance of XGBoost among the eight ML algorithms evaluated is consistent with previous studies showing the effectiveness of ensemble-based methods for medical prediction tasks (25–27). Our SHAP analysis identified several clinically meaningful patterns. Among the three liver enzymes, GGT made the most prominent contribution to model prediction, with its SHAP value exceeding those of conventional metabolic markers such as blood pressure. All three enzymes also showed substantial absolute SHAP values that were larger than those of BMI in the models. The predictive contribution of liver enzymes was consistent across both US and Chinese datasets, suggesting that these biomarkers may be generalizable across ethnicities.

The association of GGT with T2D is explained by its glutathione-related roles in oxidative stress-mediated beta-cell damage, inflammation-induced insulin resistance, and dysregulation of hepatic glucose metabolism (28–30). ALT elevations are associated with liver fat content and hepatic insulin resistance (31, 32). These enzymes collectively indicate progressive metabolic liver injury, with GGT reflecting redox imbalance, ALT indicating steatotic lipotoxicity, and AST signaling mitochondrial distress during NASH progression (28, 31, 33). Their complementary elevations capture distinct diabetes-pathogenic processes, from initial hepatic lipid accumulation (ALT) to subsequent inflammatory/oxidative damage (GGT/AST), explaining their combined incremental value beyond conventional risk factors.

Although the current study leveraged two large, well-characterized datasets, several limitations warrant consideration. First, inclusion of FLI-defined NAFLD status alongside its constituent indicators (BMI, WC, and TG) may introduce construct overlap, particularly because GGT, a component of FLI, is subsequently added to the model to evaluate its incremental value. Nevertheless, we used binary FLI-based NAFLD status rather than continuous FLI scores, meaning that the continuous, graded information remains available for the model to leverage. Second, the diagnostic criteria for NAFLD differed between the cohorts, potentially introducing minor bias. Third, the prevalence of NAFLD was 47.05% and 51.80% in the NHANES and Chinese cohorts, which was higher than the reported NAFLD prevalence of approximately 32% in the general adult population (34). The generalizability to low-risk community settings therefore requires further validation. Lastly, the inherently observational design precludes definitive causal inference regarding the observed association.

Conclusions

In conclusion, this study provides consistent evidence that incorporating liver enzymes could improve the performance of ML-based diagnostic prediction models for T2D across different populations and algorithms. These findings indicate that routinely available liver enzymes could improve the identification of individuals with T2D. Future longitudinal research should examine the temporal relationship between liver enzyme changes and diabetes development, elucidate the underlying mechanisms linking these changes to diabetes pathogenesis, and explore the potential for personalized prevention strategies based on these biomarkers.

Acknowledgments

We thank SageSci Editing for their professional English language editing services.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by the First Hospital of Jilin University (Doctor of Excellence Program (DEP): JDYY-DEP-2024001) and The Science & Technology Department of Jilin Province (20260203188SF).

Footnotes

Edited by: Tong Yue, University of Science and Technology of China, China

Reviewed by: Tien Van Nguyen, Thai Binh University of Medicine and Pharmacy, Vietnam

Carlos Fernando Martínez-Cabrera, Centro de Investigación y Gastroenterología, Mexico

Data availability statement

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding authors.

Ethics statement

The studies involving humans were approved by Ethics Committee of the First Hospital of Jilin University. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

YB: Data curation, Formal analysis, Writing – original draft, Funding acquisition. XY: Data curation, Writing – original draft. YW: Formal analysis, Writing – original draft. CS: Conceptualization, Writing – review & editing. ST: Conceptualization, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fendo.2026.1905144/full#supplementary-material

DataSheet1.pdf (213.7KB, pdf)
Supplementaryfile1.pdf (329.9KB, pdf)

References

  • 1. Sun H, Saeedi P, Karuranga S, Pinkepank M, Ogurtsova K, Duncan BB, et al. Idf diabetes atlas: global, regional and country-level diabetes prevalence estimates for 2021 and projections for 2045. Diabetes Res Clin Pract. (2022) 183:109119. doi:  10.1016/j.diabres.2021.109119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Baena-Díez JM, Peñafiel J, Subirana I, Ramos R, Elosua R, Marín-Ibañez A, et al. Risk of cause-specific death in individuals with diabetes: a competing risks analysis. Diabetes Care. (2016) 39:1987–95. doi:  10.2337/dc16-0614 [DOI] [PubMed] [Google Scholar]
  • 3. Zheng Y, Ley SH, Hu FB. Global aetiology and epidemiology of type 2 diabetes mellitus and its complications. Nat Rev Endocrinol. (2018) 14:88–98. doi:  10.1038/nrendo.2017.151 [DOI] [PubMed] [Google Scholar]
  • 4. Khan S. Artificial intelligence and machine learning in clinical medicine. N Engl J Med. (2023) 388:2398. doi:  10.1056/NEJMc2305287 [DOI] [PubMed] [Google Scholar]
  • 5. Fregoso-Aparicio L, Noguez J, Montesinos L, García-García JA. Machine learning and deep learning predictive models for type 2 diabetes: a systematic review. Diabetol Metab Syndrome. (2021) 13:148. doi:  10.1186/s13098-021-00767-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Kavakiotis I, Tsave O, Salifoglou A, Maglaveras N, Vlahavas I, Chouvarda I. Machine learning and data mining methods in diabetes research. Comput Struct Biotechnol J. (2017) 15:104–16. doi:  10.1016/j.csbj.2016.12.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Xie Z, Nikolayeva O, Luo J, Li D. Building risk prediction models for type 2 diabetes using machine learning techniques. Preventing Chronic Dis. (2019) 16:E130. doi:  10.5888/pcd16.190109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Deberneh HM, Kim I. Prediction of type 2 diabetes based on machine learning algorithm. Int J Environ Res Public Health. (2021) 18:3317. doi:  10.3390/ijerph18063317 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Joshi RD, Dhakal CK. Predicting type 2 diabetes using logistic regression and machine learning approaches. Int J Environ Res Public Health. (2021) 18:7346. doi:  10.3390/ijerph18147346 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Xiong XL, Zhang RX, Bi Y, Zhou WH, Yu Y, Zhu DL. Machine learning models in type 2 diabetes risk prediction: results from a cross-sectional retrospective study in Chinese adults. Curr Med Sci. (2019) 39:582–8. doi:  10.1007/s11596-019-2077-4 [DOI] [PubMed] [Google Scholar]
  • 11. Lee H, Hwang SH, Park S, Choi Y, Lee S, Park J, et al. Prediction model for type 2 diabetes mellitus and its association with mortality using machine learning in three independent cohorts from South Korea, Japan, and the UK: a model development and validation study. EClinicalMedicine. (2025) 80:103069. doi:  10.1016/j.eclinm.2025.103069 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Hahn SJ, Kim S, Choi YS, Lee J, Kang J. Prediction of type 2 diabetes using genome-wide polygenic risk score and metabolic profiles: a machine learning analysis of population-based 10-year prospective cohort study. EBioMedicine. (2022) 86:104383. doi:  10.1016/j.ebiom.2022.104383 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Wei H, Sun J, Shan W, Xiao W, Wang B, Ma X, et al. Environmental chemical exposure dynamics and machine learning-based prediction of diabetes mellitus. Sci Total Environ. (2022) 806:150674. doi:  10.1016/j.scitotenv.2021.150674 [DOI] [PubMed] [Google Scholar]
  • 14. Zhao M, Wan J, Qin W, Huang X, Chen G, Zhao X. A machine learning-based diagnosis modelling of type 2 diabetes mellitus with environmental metal exposure. Comput Methods Programs BioMed. (2023) 235:107537. doi:  10.1016/j.cmpb.2023.107537 [DOI] [PubMed] [Google Scholar]
  • 15. Wang YL, Koh WP, Yuan JM, Pan A. Association between liver enzymes and incident type 2 diabetes in Singapore Chinese men and women. BMJ Open Diabetes Res Care. (2016) 4:e000296. doi:  10.1136/bmjdrc-2016-000296 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Zhang J, Cheng N, Ma Y, Li H, Cheng Z, Yang Y, et al. Liver enzymes, fatty liver and type 2 diabetes mellitus in a Jinchang cohort: a prospective study in adults. Can J Diabetes. (2018) 42:652–8. doi:  10.1016/j.jcjd.2018.02.002 [DOI] [PubMed] [Google Scholar]
  • 17. Bi Y, Yang Y, Yuan X, Wang J, Wang T, Liu Z, et al. Association between liver enzymes and type 2 diabetes: a real-world study. Front Endocrinol. (2024) 15:1340604. doi:  10.3389/fendo.2024.1340604 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Faramarzi E, Mehrtabar S, Molani-Gol R, Dastgiri S. The relationship between hepatic enzymes, prediabetes, and diabetes in the Azar cohort population. BMC Endocr Disord. (2025) 25:41. doi:  10.1186/s12902-025-01871-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Ashoobi MT, Joukar F, Mojtahedi K, Maroufizadeh S, Javid M, Parvaneh A, et al. Elevated liver enzymes and diabetes in the Persian Guilan cohort study. Caspian J Internal Med. (2025) 16:73–82. doi:  10.22088/cjim.16.1.73 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Minato-Inokawa S, Tsuboi-Kaji A, Honda M, Takeuchi M, Kitaoka K, Kurata M, et al. The different associations of serum gamma-glutamyl transferase and alanine aminotransferase with insulin secretion, β-cell function, and insulin resistance in non-obese Japanese. Sci Rep. (2024) 14:19234. doi:  10.1038/s41598-024-70396-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Van Calster B, Nieboer D, Vergouwe Y, De Cock B, Pencina MJ, Steyerberg EW. A calibration hierarchy for risk models was defined: from utopia to empirical data. J Clin Epidemiol. (2016) 74:167–76. doi:  10.1016/j.jclinepi.2015.12.005 [DOI] [PubMed] [Google Scholar]
  • 22. Van Calster B, Wynants L, Verbeek JFM, Verbakel JY, Christodoulou E, Vickers AJ, et al. Reporting and interpreting decision curve analysis: a guide for investigators. Eur Urol. (2018) 74:796–804. doi:  10.1016/j.eururo.2018.08.038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Zhou X, Qiao Q, Ji L, Ning F, Yang W, Weng J, et al. Nonlaboratory-based risk assessment algorithm for undiagnosed type 2 diabetes developed on a nation-wide diabetes survey. Diabetes Care. (2013) 36:3944–52. doi:  10.2337/dc13-0593 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. Tripod+ai statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ (Clinical Res Ed). (2024) 385:e078378. doi:  10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Dinh A, Miertschin S, Young A, Mohanty SD. A data-driven approach to predicting diabetes and cardiovascular disease with machine learning. BMC Med Inf Decis Making. (2019) 19:211. doi:  10.1186/s12911-019-0918-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Li L, Cheng Y, Ji W, Liu M, Hu Z, Yang Y, et al. Machine learning for predicting diabetes risk in Western China adults. Diabetol Metab Syndrome. (2023) 15:165. doi:  10.1186/s13098-023-01112-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Liu Q, Zhang M, He Y, Zhang L, Zou J, Yan Y, et al. Predicting the risk of incident type 2 diabetes mellitus in Chinese elderly using machine learning techniques. J Personalized Med. (2022) 12:905. doi:  10.3390/jpm12060905 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Koenig G, Seneff S. Gamma-glutamyltransferase: a predictive biomarker of cellular antioxidant inadequacy and disease risk. Dis Markers. (2015) 2015:818570. doi:  10.1155/2015/818570 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Kunutsor SK. Gamma-glutamyltransferase-friend or foe within? Liver Int Off J Int Assoc For Study Liver. (2016) 36:1723–34. doi:  10.1111/liv.13221 [DOI] [PubMed] [Google Scholar]
  • 30. Singh A, Kukreti R, Saso L, Kukreti S. Mechanistic insight into oxidative stress-triggered signaling pathways and type 2 diabetes. Molecules (Basel Switzerland). (2022) 27:950. doi:  10.3390/molecules27030950 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Korenblat KM, Fabbrini E, Mohammed BS, Klein S. Liver, muscle, and adipose tissue insulin action is directly related to intrahepatic triglyceride content in obese subjects. Gastroenterology. (2008) 134:1369–75. doi:  10.1053/j.gastro.2008.01.075 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Vozarova B, Stefan N, Lindsay RS, Saremi A, Pratley RE, Bogardus C, et al. High alanine aminotransferase is associated with decreased hepatic insulin sensitivity and predicts the development of type 2 diabetes. Diabetes. (2002) 51:1889–95. doi:  10.2337/diabetes.51.6.1889 [DOI] [PubMed] [Google Scholar]
  • 33. Kojima H, Sakurai S, Uemura M, Fukui H, Morimoto H, Tamagawa Y. Mitochondrial abnormality and oxidative stress in nonalcoholic steatohepatitis. Alcoholism Clin Exp Res. (2007) 31:S61–6. doi:  10.1111/j.1530-0277.2006.00288.x [DOI] [PubMed] [Google Scholar]
  • 34. Teng ML, Ng CH, Huang DQ, Chan KE, Tan DJ, Lim WH, et al. Global incidence and prevalence of nonalcoholic fatty liver disease. Clin Mol Hepatol. (2023) 29:S32–42. doi:  10.3350/cmh.2022.0365 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

DataSheet1.pdf (213.7KB, pdf)
Supplementaryfile1.pdf (329.9KB, pdf)

Data Availability Statement

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding authors.


Articles from Frontiers in Endocrinology are provided here courtesy of Frontiers Media SA

RESOURCES