Abstract
Background
Metabolic dysfunction-associated steatotic liver disease (MASLD) is a leading chronic liver disorder closely linked to diabetes mellitus (DM) and its cardiovascular and renal complications. Early identification of diabetes risk in this population is essential for timely intervention.
Objective
To develop machine learning (ML) models to predict diabetes risk in individuals with MASLD and to identify key predictive factors using a nationally representative dataset.
Methods
Data from 6310 MASLD participants (2007–2018) were analysed and classified into DM and non-DM groups. Feature selection was performed using Random Forest, Least Absolute Shrinkage and Selection Operator and Support Vector Machine Recursive Feature Elimination. Based on selected features, nine ML models were developed. Model performance was evaluated using accuracy, sensitivity, area under the curve, F1 score, Rank Score and Brier Score. SHapley Additive exPlanations (SHAP) were used for interpretability.
Results
Eight key variables (age, urinary albumin (Ualb), total cholesterol (TC), lipid accumulation product (LAP), urinary creatinine, white blood cell count, uric acid and Visceral Adiposity Index) were identified and used for model construction. Among nine algorithms, the Light Gradient Boosting Machine (LightGBM) model showed superior predictive performance. SHAP analysis revealed that Ualb, age, TC and LAP were the most influential predictors.
Conclusion
Our ML-based model effectively identifies individuals with MASLD at high risk for developing DM. The LightGBM algorithm outperformed other models in both accuracy and interpretability. Key predictors such as Ualb and LAP highlight the importance of renal and metabolic markers in early diabetes risk prediction, offering a new approach for individualised intervention and clinical decision-making.
Keywords: Diabetes & endocrinology, Machine Learning, Risk Factors
STRENGTHS AND LIMITATIONS OF THIS STUDY.
This study integrates three complementary feature selection algorithms (Least Absolute Shrinkage and Selection Operator, Random Forest and Support Vector Machine Recursive Feature Elimination) to identify robust, biologically meaningful predictors of diabetes risk in metabolic dysfunction-associated steatotic liver disease (MASLD), enhancing model stability and interpretability.
Nine mainstream machine learning algorithms were systematically compared, and Light Gradient Boosting Machine demonstrated superior discrimination and calibration performance, supporting its potential use as an efficient tool for individualised diabetes risk stratification in MASLD.
As this study is based on cross-sectional National Health and Nutrition Examination Survey data, temporal directionality cannot be established, and the model predicts current diabetes status rather than future disease onset.
External validation using independent longitudinal cohorts was not feasible, and future studies are required to assess the generalisability and predictive stability of the model in real-world clinical settings.
Introduction
Metabolic dysfunction-associated steatotic liver disease (MASLD) is defined by the accumulation of excessive fat in the liver without significant alcohol intake. It is considered the hepatic manifestation of metabolic syndrome and is frequently associated with obesity, diabetes mellitus (DM) and hyperlipidaemia.1 Due to globalisation and lifestyle changes, MASLD has become one of the most common chronic liver diseases, affecting approximately 25% of adults worldwide.2 According to the American Gastroenterological Association, MASLD is expected to surpass all other causes and become the leading indication for liver transplantation in the USA by 2030.3
MASLD can progress insidiously to non-alcoholic steatohepatitis (NASH), liver fibrosis, cirrhosis and even hepatocellular carcinoma (HCC), markedly increasing morbidity and mortality and imposing a heavy burden on patients, families and the entire healthcare system.4 Beyond liver-related complications, MASLD is strongly associated with multiple metabolic disorders, especially a markedly elevated risk of DM.
Hepatic fat accumulation has been identified as one of the earliest pathological changes in the progression of diabetes. MASLD exacerbates systemic metabolic stress by inducing hepatic insulin resistance, enhancing hepatic glucose output and disturbing lipid metabolism.5 Furthermore, steatotic liver tissue secretes multiple pro-inflammatory cytokines (eg, tumor necrosis factor alpha (TNF-α) and interleukin-6 (IL-6)), which not only lead to chronic hepatic injury but also impair insulin signalling in skeletal muscle and adipose tissue, thereby further disturbing systemic glucose homeostasis.6 7 Meanwhile, dysregulated hepatokine secretion (eg, fetuin-A) and oxidative stress caused by mitochondrial dysfunction may also contribute to the onset of diabetes by impairing pancreatic β-cell function.8 9 Therefore, MASLD should be regarded not only as a hepatic manifestation of metabolic dysregulation but also as a critical contributor to diabetes pathogenesis. Early identification of high-risk individuals with MASLD, combined with timely and effective intervention and prevention strategies, is essential for both clinical management and public health.
Although numerous studies have achieved a strong association between MASLD and diabetes, accurately identifying individuals at genuinely high risk of developing diabetes within the MASLD population remains a significant challenge. Most existing risk assessment tools are based on traditional statistical models such as logistic regression (Logistic), which assume linear relationships among variables and are limited in capturing complex interactions among metabolic, biochemical and behavioural factors. Moreover, individuals with MASLD exhibit considerable heterogeneity driven by diverse factors, including genetic background, age, obesity, insulin sensitivity, dietary habits, physical activity, gut microbiota composition and systemic inflammation. This heterogeneity not only complicates risk prediction but also undermines the effectiveness of standardised intervention strategies. In recent years, machine learning (ML) techniques have shown remarkable accuracy and adaptability in predicting chronic diseases. Compared with traditional approaches, ML can process high-dimensional and non-linear data, making it more suitable for identifying hidden risk factors for diabetes in MASLD populations. Furthermore, the integration of interpretability algorithms such as SHapley Additive exPlanations (SHAP) has improved the transparency of model outputs, thereby enhancing the clinical credibility of ML-based applications. Early identification of high-risk individuals for diabetes among those with MASLD requires access to high-quality, representative, large-scale datasets. The National Health and Nutrition Examination Survey (NHANES), a comprehensive and authoritative database covering demographic, lifestyle, biochemical and disease-related information, provides a robust foundation for developing ML-based predictive models and performing feature selection.
This study aimed to identify key risk factors associated with diabetes in individuals with MASLD by employing multiple mainstream ML algorithms using data from NHANES. We integrated multidimensional variables—including demographic characteristics, biochemical markers and metabolic indices—and systematically compared the predictive performance of multiple modelling approaches. Ultimately, we identified age, urinary albumin (Ualb), total cholesterol (TC), lipid accumulation product (LAP), urinary creatinine (Ucr), white blood cell count (WBC), uric acid (UA) and Visceral Adiposity Index (VAI) as the most influential predictive features. Furthermore, we used the SHAP algorithm for model interpretability, revealing the relative contribution of each variable to diabetes risk prediction and offering novel insights into the potential mechanisms driving the progression from MASLD to diabetes. This study provides theoretical support for early diabetes screening and the development of personalised intervention strategies, laying a foundation for precise risk stratification in individuals with MASLD.
Methods
Study population
NHANES is a nationally representative, cross-sectional survey conducted biennially since 1999 to collect health and nutrition data from the non-institutionalised civilian population in the USA. It uses a stratified, multistage and complex probability sampling design. Each participant is assigned a sampling weight to account for unequal probabilities of selection and potential non-response.
This study used data collected from 2007 to 2018.10 Participants with MASLD were defined as individuals with a Hepatic Steatosis Index (HSI) ≥36.4 11–13 The HSI estimates the probability of MASLD based on body mass index (BMI), the alanine aminotransferase (ALT) to aspartate aminotransferase (AST) ratio, and the presence of diabetes or female sex. The HSI is calculated using the following formula:
Diabetes was defined as a fasting plasma glucose (FPG) level ≥126 mg/dL, glycated haemoglobin (HbA1c) ≥ 6.5% or a self-reported physician diagnosis of diabetes during the interview.14 All participants with MASLD were included, while individuals younger than 18 or older than 74 years were excluded. Additionally, individuals with alcohol-related or virus-associated liver diseases were excluded. Excessive alcohol consumption was defined as more than three drinks per day for men or more than two drinks per day for women.15 Viral liver disease was defined as the presence of hepatitis B surface antigen or detectable hepatitis C virus RNA.
To ensure data quality and temporal consistency, records with more than 10% missing values were excluded, while multiple imputation was performed for records with minor missingness. The outcome events (diabetes) accounted for ~20% of the sample. Given that the minority class proportion exceeded conventional thresholds for severe class imbalance, no dedicated oversampling, undersampling or synthetic minority oversampling methods were implemented. Instead, classification algorithms incorporating class-weight adjustment (where available) were used, and model calibration was assessed via calibration curves and Brier Score. No additional recalibration specifically for class imbalance was required. In addition, individuals without data on FPG, HbA1c or self-reported diabetes diagnosis were excluded. Ultimately, individuals with MASLD were classified into diabetes and non-diabetes groups according to their diabetes diagnosis status. A detailed flowchart of the screening process and study design is shown in figure 1. This study followed the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis–Artificial Intelligence (TRIPOD+AI) reporting guideline to ensure transparency in model development.16
Figure 1.

Flowchart of participant selection and ML model construction for diabetes prediction in MASLD patients. This flowchart illustrates the stepwise selection of participants from the NHANES 2007–2018 dataset and the overall analysis pipeline. From 59 842 individuals, 14 123 were identified with MASLD based on HSI >36. After excluding those with age >74 or <18 years, excessive alcohol intake or viral hepatitis, 7994 participants remained. Further exclusions were applied for missing data (>10%) and incomplete diabetes-related records, resulting in 6310 eligible individuals (1258 with diabetes and 5052 without). A total of 48 differential variables were screened using LASSO, RF and SVM, yielding eight common predictive features. These were used to construct prediction models using nine ML algorithms. The dataset was split into training (70%) and testing (30%) sets. Model performance was evaluated by AUC, DCA, calibration curves and SHAP analysis. AUC, area under the curve; DCA, decision curve analysis; DM, diabetes mellitus; FPG, fasting plasma glucose; HbA1c, glycated haemoglobin; HSI, Hepatic Steatosis Index; LASSO, Least Absolute Shrinkage and Selection Operator; MASLD, metabolic dysfunction-associated steatotic liver disease; NHANES, National Health and Nutrition Examination Survey; RF, Random Forest; SHAP, SHapley Additive exPlanations; SVM, Support Vector Machine.
Blinding of outcome assessment
Outcome assessments in NHANES were performed independently by trained laboratory and interview staff who were blinded to any modelling hypotheses or predictor information. Therefore, no additional blinding procedures were required.
Variable selection and measurement
This study systematically selected a range of variables to investigate factors associated with diabetes development in individuals with MASLD. Participant data were extracted, including demographic characteristics, examination findings, questionnaire responses and laboratory results. Demographic variables included sex, age and race/ethnicity. Educational attainment was classified into three categories: less than high school, high school graduate and more than high school. Marital status was categorised as married, widowed, divorced, separated, never married or cohabiting with a partner. The poverty income ratio (PIR) was calculated by dividing family income by the federal poverty threshold corresponding to the survey year and geographic location. Lifestyle variables included smoking status (smoker or non-smoker) and alcohol use (drinker or non-drinker). Physical activity levels were evaluated based on responses to the NHANES physical activity questionnaire. Physical activity was quantified as metabolic equivalent task minutes per week (MET·min/week), derived from self-reported frequency (days/week), duration (minutes/session) and intensity. According to NHANES guidelines, moderate-intensity activities (eg, brisk walking) were assigned a MET value of 4.0, and vigorous-intensity activities (eg, running) were assigned a MET value of 8.0. Total MET-minutes were calculated using the formula:
Based on total MET values, individuals were classified into three physical activity levels: low (<600 MET·min/week), moderate (600–1200 MET·min/week) and high (>1200 MET·min/week).17 This classification was based on international physical activity guidelines and is widely used in studies evaluating the impact of physical activity on health outcomes. Clinical variables included self-reported or clinically confirmed diagnoses of diabetes, hypertension, heart failure, angina pectoris and stroke.
Anthropometric and biochemical indicators were assessed during the NHANES survey cycles. Anthropometric measurements included height, BMI and waist circumference. The biochemical indicators analysed in this study encompassed markers of renal function, liver function, lipid metabolism and haematologic status. The biochemical indicators analysed in this study encompassed markers of renal function, liver function, lipid metabolism and haematologic status. Renal function markers included Ualb, Ucr, serum creatinine (Scr), UA and blood urea nitrogen (BUN), which were used to evaluate renal metabolic status. Liver function markers included albumin (Alb), ALT, AST, alkaline phosphatase, gamma-glutamyl transferase (GGT) and total bilirubin, which primarily reflect hepatic synthetic function and the degree of hepatocellular injury. Lipid metabolism markers included high-density lipoprotein (HDL), TC, triglycerides (TG), LAP and VAI, which were used to assess lipid metabolic status. In this study, LAP and VAI were included as key markers reflecting fat distribution and metabolic function. LAP is a sex-specific index derived from waist circumference (WC) and TG levels, used to estimate abdominal fat accumulation indirectly.18 The formula for LAP is as follows:
The VAI is a sex-specific composite index calculated from WC, BMI, TG and high-density lipoprotein cholesterol (HDL-C), used to evaluate visceral fat function and its related metabolic risk.19 The sex-specific formulas for VAI calculation are as follows:
Haematological parameters included WBC, red blood cell count (RBC), haemoglobin (Hb), platelet count, haematocrit, mean corpuscular volume, mean corpuscular haemoglobin (MCH), red cell distribution width (RDW), UA and Scr levels. These parameters were used to comprehensively evaluate systemic inflammation, immune function and anaemia status.
Additionally, this study collected data on participants’ nutrient intake, including total energy, protein, sugar, fat and beta-carotene, as well as vitamin-related indicators such as vitamins B12, C, D and K. Zinc (Zn), copper (Cu) and selenium (Se) are essential trace elements that play critical roles in glucose metabolism, oxidative stress defence and the regulation of insulin function. Previous studies have reported strong associations between these trace elements and the risk of diabetes. Nutrient intake data were obtained from the NHANES 24-hour dietary recall questionnaire, which assesses individuals’ daily dietary consumption. By analysing the associations between nutritional components and disease status, the potential contributions of dietary factors to health risks can be further elucidated.
Model building process
Key features screening by three different algorithms
To comprehensively evaluate the predictive value of each variable with respect to the study outcome, and to improve model stability and interpretability, multiple ML algorithms were applied for feature selection. These methods included the Random Forest (RF) algorithm, Least Absolute Shrinkage and Selection Operator (LASSO) regression and Support Vector Machine Recursive Feature Elimination (SVM-RFE).
First, the RF model was employed to assign importance scores to variables based on their contributions to model performance at DT split nodes. Variables were ranked based on these scores, and the top 20 were selected for subsequent feature screening. LASSO regression was subsequently applied to select candidate variables. By applying an L1 regularisation penalty, LASSO enables simultaneous feature selection and model complexity control, effectively removing redundant or irrelevant variables. The optimal regularisation parameter (λ) was determined via cross-validation, and variables with non-zero coefficients were retained as candidate features. SVM-RFE was employed to evaluate all variables. This method iteratively eliminates the least important variables by training a Support Vector Machine (SVM) model, ultimately identifying the optimal subset of features contributing most to classification performance.20 Finally, the sets of variables selected by all three methods were extracted, and their intersection was taken as the final feature set for downstream model construction and analysis. This integrative strategy leveraged the strengths of multiple feature selection algorithms to enhance the robustness and biological interpretability of the selected features.
Development of diagnostic models based on machine learning algorithms
After completing feature selection, nine widely used ML algorithms were implemented to comprehensively evaluate the predictive performance of the selected features for the study outcome. The algorithms used included Multilayer Perceptron (MLP), Light Gradient Boosting Machine (LightGBM), K-Nearest Neighbours (KNN), Logistic, RF, eXtreme Gradient Boosting (XGBoost), Radial Basis Function Support Vector Machine (RSVM), Elastic Net Regression (ENET) and Decision Tree (DT).
Prior to model development, the original dataset was randomly split into a training set and a testing set at a ratio of 7:3. The training set was used for model fitting and hyperparameter tuning, whereas the testing set was reserved for evaluating the model’s generalisation performance. To improve model robustness and accuracy, fivefold cross-validation was conducted on the training set for both model training and hyperparameter optimisation.
Optimal algorithm selection and model performance evaluation
Model performance was assessed using both standard and composite evaluation metrics. Standard metrics included the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity and the F1 score. Among these, the F1 score is a critical metric that balances precision and recall, making it particularly suitable for datasets with imbalanced class distributions.21 Composite metrics included the Rank Score and the Brier Score. The Rank Score evaluates the overall relative performance ranking of the model across multiple individual metrics. A lower Rank Score indicates superior overall performance, reflecting greater robustness and consistency at the global level.22
The Brier Score was employed to assess the accuracy of predicted probabilities generated by each model. The Brier Score is defined as the mean squared difference between predicted probabilities and actual binary outcomes; a lower score reflects better calibration and overall predictive performance.23 This metric is particularly well suited for binary classification tasks, as it captures both the discrimination and calibration components of model performance. All Brier Scores were computed using the testing set. The optimal algorithms were defined as those ranking within the top three across both the Rank Score and Brier Score metrics.
To further evaluate the clinical applicability and predictive performance of the constructed models, a comprehensive assessment was performed across three dimensions: interpretability, clinical utility and goodness-of-fit. Specifically, SHAP summary plots were used to interpret model predictions, decision curve analysis (DCA) to assess clinical utility, and calibration curves to evaluate the agreement between predicted probabilities and observed outcomes.
First, SHAP summary plots were generated based on the best-performing model to visualise the relative contribution of each predictor to the outcome, thereby enhancing model interpretability and supporting individualised risk assessment. In addition, SHAP values help identify key high-risk features, thereby informing the development of targeted clinical intervention strategies. Second, DCA was employed to assess the net benefit of each model across a range of threshold probabilities, providing insights into its potential value in real-world clinical decision-making.20 DCA considers both false-positive and false-negative outcomes, making it a valuable tool for comparing the clinical utility of predictive models. Finally, calibration curves were constructed to evaluate the agreement between predicted probabilities and observed outcomes. By examining the relationship between predicted probabilities and observed event rates, calibration curves indicate whether the model tends to overestimate or underestimate risk, thereby reflecting its calibration accuracy. Good model calibration is indicated by close alignment between predicted and observed values, with the calibration curve approaching the ideal 45° diagonal reference line.24
Statistical analysis
NHANES employs a complex, multistage probability sampling design to produce nationally representative estimates. In the descriptive analysis, weighted percentages and 95% CIs were calculated to account for the complex survey design. Categorical variables were summarised as frequencies (n) and weighted percentages (%), whereas continuous variables were reported as medians with IQRs.
For group comparisons, weighted χ2 tests were applied to categorical variables, and Mann-Whitney U tests were used for non-parametric analysis of continuous variables, following NHANES analytic guidelines. Model development and performance evaluation were performed using unweighted data. All statistical tests were two-sided, with a p value <0.05 considered statistically significant.
Results
Baseline characteristics
A total of 6310 participants were included in this study, comprising 5052 individuals without diabetes and 1258 individuals with diabetes. Compared with individuals without diabetes, those with diabetes were older (59.00 (50.00, 65.00) vs 45.00 (33.00, 58.00) years, p<0.001), had a higher proportion of females (p=0.003) and had a significantly greater proportion of non-Hispanic Black individuals (p<0.001). Significant differences were also observed between the two groups in marital status, educational attainment and physical activity level (p<0.001). Additionally, individuals with diabetes had higher proportions of smokers (p<0.001) and alcohol drinkers (p=0.003).
In terms of comorbidities, the prevalence of heart failure, angina, stroke and essential hypertension (EH) was significantly higher among individuals with diabetes compared with those without diabetes (p<0.001). Regarding anthropometric parameters, individuals with diabetes exhibited significantly higher values for BMI (33.32 (29.60, 38.20) vs 31.60 (29.00, 35.57)), WC, HSI, LAP and VAI (p<0.001).
Regarding biochemical indicators, individuals with diabetes exhibited significantly lower levels of ALB and HDL and significantly higher levels of BUN, GGT, TG, UA, Ualb and WBC (p<0.001). Regarding nutrient intake, individuals with diabetes reported significantly lower consumption of total energy (1844.50 (1373.00, 2461.75) vs 2027.00 (1527.75, 2661.00) kcal), protein, fat and sugar compared with those without diabetes (p<0.001).
With respect to micronutrients, individuals with diabetes exhibited significantly lower intake of several vitamins—including B1, B2, B6, B12, C, D and E—as well as reduced serum concentrations of Se, Zn and Cu, compared with those without diabetes (p<0.05). All key variables are summarised in online supplemental table S1.
bmjopen-16-1-s003.pdf (123.1KB, pdf)
Identification of key features
To explore linear associations among variables, a correlation heatmap was generated using data from the NHANES dataset (online supplemental figure S1). The upper triangle represents the direction and strength of correlations using colour intensity, while the lower triangle denotes statistical significance through circle size. The results revealed predominantly weak to moderate correlations among variables, with no evidence of substantial multicollinearity observed overall. Ualb was strongly positively correlated with Ucr, and BMI showed a high correlation with WC, both consistent with physiological expectations. TG, LAP and VAI were strongly positively correlated (r>0.7), indicating consistency in their assessment of fat metabolism and accumulation. HDL was moderately negatively correlated with TG, LAP and VAI, supporting its protective inverse association with metabolic risk. Among sociodemographic variables, the PIR was moderately positively correlated with education level and nutrient intake (eg, protein, carbohydrates, vitamins), and moderately negatively correlated with EH and estimated glomerular filtration rate (eGFR). Sex was significantly associated with haematological parameters such as Hb, MCH and RBC, suggesting potential sex-based differences. Additionally, variables such as smoking and physical activity level showed only weak correlations with most metabolic indicators, suggesting limited independent contributions. To assess the discriminatory power of individual variables in identifying diabetes among obese individuals, AUC values were calculated for each variable (online supplemental figure S2). Age, history of EH and WC exhibited the highest discriminatory ability, as reflected by top-ranking AUC values. Variables including LAP, Ualb, sugar intake, RDW and VAI also exhibited good predictive performance. Several metabolic indicators (eg, TG, TC, BUN, BMI, GGT) and nutrient intake variables (eg, energy, protein, vitamin B6) suggested moderate predictive ability. In contrast, sociodemographic factors—including marital status, education level, physical activity level and alcohol consumption—showed relatively weaker predictive power.
bmjopen-16-1-s001.pdf (341.7KB, pdf)
bmjopen-16-1-s002.pdf (271.7KB, pdf)
To identify key variables associated with diabetes onset in the obese population, an integrated feature selection approach was applied, incorporating LASSO regression, RF and SVM. As shown in figure 2, these methods evaluated variable importance from different perspectives, thereby enhancing the robustness of the feature selection process.
Figure 2.

Feature selection for diabetes prediction in MASLD patients using LASSO, RF and SVM algorithms. (A) 10-fold cross-validation for tuning the LASSO regression model. The optimal lambda value was selected based on the minimum mean cross-validated error. (B) LASSO coefficient profiles of all features across a sequence of lambda values. (C) Feature importance ranking by the RF model. (D) Feature importance ranking by the SVM model. (E) Venn diagram showing the overlap of selected features among LASSO, RF and SVM. Eight common variables were identified by all three methods. ALP, alkaline phosphatase; BMI, body mass index; BUN, blood urea nitrogen; EH, essential hypertension; HB, haemoglobin; HF, heart failure; LAP, lipid accumulation product; LASSO, Least Absolute Shrinkage and Selection Operator; MASLD, metabolic dysfunction-associated steatotic liver disease; MC, mean corpuscular haemoglobin concentration; PLT, platelet count; RDW, red cell distribution width; RF, Random Forest; SCR, serum creatinine; SVM, Support Vector Machine; TC, total cholesterol; TG, triglyceride; UA, uric acid; Ualb, urinary albumin; Ucr, urinary creatinine; VAI, Visceral Adiposity Index; WBC, white blood cell count.
In the LASSO regression, the optimal regularisation parameter was determined using 10-fold cross-validation, and variables with non-zero coefficients were selected, including age, Ualb, TC, Ucr, UA and WBC (figure 2A,B). The variable importance rankings derived from the RF model (figure 2C) indicated that LAP, VAI, age, TG and WC were among the top predictors. Results from the SVM model (figure 2D) suggested that Ualb, VAI, UA and age substantially contributed to the model’s predictive performance. To integrate the outputs of the three algorithms, common features were extracted from their intersection (figure 2E), ultimately yielding eight core variables: age, Ualb, TC, LAP, Ucr, WBC, UA and VAI. These variables span multiple domains, including demographic characteristics, lipid metabolism, renal function and inflammatory status. These variables supported clear biological relevance and were subsequently used as key input features for model development.
Construction of diagnostic models
To evaluate the predictive performance of various ML algorithms for diabetes risk in individuals with MASLD, nine classification models were constructed using the eight previously identified core variables (age, Ualb, TC, LAP, Ucr, WBC, UA and VAI). The models included Logistic, ENET, DT, RF, XGBoost, RSVM, MLP, LightGBM and KNN.
To ensure robust model development and evaluation, a total of 6310 patients with MASLD were randomly divided into a training set (70%) and a testing set (30%). The training set was used for model development and internal validation, while the testing set was reserved for independent performance evaluation.
A comparison of baseline characteristics, including diabetes prevalence, sex ratio and the eight core diagnostic variables (age, Ualb, TC, LAP, Ucr, WBC, UA and VAI), was performed between the two sets. No statistically significant differences were observed between the training and testing cohorts (all p>0.05), indicating that the two datasets were comparable and representative of the overall MASLD population.
These findings support the validity of the data partitioning strategy for subsequent model construction and evaluation (online supplemental table S2).
bmjopen-16-1-s004.pdf (88.5KB, pdf)
DCA revealed that most ML models outperformed both the ‘treat-all’ and ‘treat-none’ strategies across a wide range of threshold probabilities, indicating favourable clinical applicability (figure 3A,B). Notably, the RF, XGBoost, MLP and LightGBM models achieved higher net benefits, underscoring their potential utility in clinical decision-making. Receiver operating characteristic (ROC) curve analysis (figure 3C,D) further revealed the strong discriminative capabilities of these models, with most achieving an AUC exceeding 0.75 in both the training and testing sets. Among them, XGBoost (AUC=0.823), RF (AUC=0.826) and LightGBM (AUC=0.821) presented the highest predictive performance in the testing set. In addition, calibration curve analysis (figure 3E) showed that XGBoost, RF, MLP and LightGBM exhibited close alignment between predicted probabilities and observed outcomes, indicating excellent calibration and predictive reliability. By contrast, models such as KNN and RSVM displayed notable calibration deviation and reduced stability.
Figure 3.

Evaluation of ML models for predicting diabetes in MASLD patients. (A, B) DCA curves for training (A) and testing (B) sets, showing the net clinical benefit of each model across a range of threshold probabilities. (C, D) ROC curves for training (C) and testing (D) sets. The AUC is labelled for each model. Calibration curves for nine ML models, comparing predicted and observed probabilities in the testing set. The diagonal line represents perfect calibration. AUC, area under the curve; DCA, decision curve analysis; dt, Decision Tree; enet, Elastic Net Regression; KNN, K-Nearest Neighbours; lightgbm, Light Gradient Boosting Machine; logistic, logistic regression; MASLD, metabolic dysfunction-associated steatotic liver disease; ML, machine learning; mlp, Multilayer Perceptron; rf, Random Forest; ROC, receiver operating characteristic; rsvm, Radial Basis Function Support Vector Machine; xgboost, eXtreme Gradient Boosting.
To systematically compare the predictive performance of different ML algorithms for diabetes risk among individuals with MASLD, nine models were evaluated using five key classification metrics: accuracy, precision, sensitivity, F1 score and AUC. The MLP, LightGBM and KNN models indicated superior performance across multiple dimensions. Specifically, MLP exhibited a balanced profile with an F1 score of 0.805 and accuracy of 0.722, KNN achieved the highest sensitivity (0.754) and LightGBM outperformed others in precision (0.927). XGBoost achieved the highest AUC (0.801), suggesting strong overall discriminative power despite its relatively lower sensitivity (0.636) (figure 4A). In terms of probability calibration, assessed using the Brier Score (figure 4B), RSVM, Logistic and LightGBM showed the best calibration performance, with Brier Scores of 0.115, 0.116 and 0.117, respectively. In contrast, MLP—despite strong classification ability—showed notable calibration deviation (Brier Score=0.153). Based on a comprehensive weighted ranking across all performance metrics (figure 4C), MLP ranked highest (score=2.9), followed by LightGBM (3.7) and KNN (4.0), indicating these models achieved an optimal balance between classification performance and model stability. The DT model consistently underperformed across all indicators and received the lowest overall rank.
Figure 4.

Comprehensive performance comparison of ML models for diabetes prediction in MASLD patients. (A) Heatmap of five evaluation metrics (accuracy, precision, sensitivity, F1 score and AUC) for each ML model on the testing set. (B) Brier score of each model, with lower values indicating better calibration. (C) Overall ranking score integrating all evaluation metrics to determine comprehensive performance. Lower scores represent better overall performance. AUC, area under the curve; DT, Decision Tree; ENET, Elastic Net Regression; KNN, K-Nearest Neighbours; Lightgbm, Light Gradient Boosting Machine; Logistic, logistic regression; MASLD, metabolic dysfunction-associated steatotic liver disease; ML, machine learning; MLP, Multilayer Perceptron; RF, Random Forest; RSVM, Radial Basis Function Support Vector Machine; XGBoost, eXtreme Gradient Boosting.
These findings indicate that while various ML algorithms can effectively predict diabetes risk in individuals with MASLD, ensemble learning methods such as LightGBM and neural network-based approaches such as MLP exhibit superior discriminative performance and overall model stability. These characteristics highlight their potential for clinical application in individualised risk stratification. Considering both discrimination and calibration metrics, LightGBM was selected as the optimal modelling algorithm, as it was the only method consistently ranked among the top three in both Brier Score and Rank Score evaluations.
Model decision of SHAP
To enhance model interpretability and assess the influence of individual features on predictive outcomes, the SHAP was applied to the LightGBM model (figure 5). In the SHAP summary plot, each row corresponds to a specific feature, and each dot represents a single participant. The colour gradient reflects the magnitude of the feature value, with red indicating higher values and blue indicating lower values. The x-axis displays the SHAP value, which quantifies the contribution of each feature to the model’s prediction, thereby elucidating both the direction and strength of its influence.
Figure 5.
SHAP summary plot of the LightGBM model for diabetes prediction in MASLD patients. SHAP summary plot showing the impact of each feature on the LightGBM model’s output. Each point represents a SHAP value for an individual, with colour indicating the original feature value (red=high, blue=low). Features are ranked by importance from top to bottom. Age, Ualb and TC were the top contributors to the prediction model. LAP, lipid accumulation product; LightGBM, Light Gradient Boosting Machine; MASLD, metabolic dysfunction-associated steatotic liver disease; SHAP, SHapley Additive exPlanations; TC, total cholesterol; UA, uric acid; Ualb, urinary albumin; Ucr, urinary creatinine; VAI, Visceral Adiposity Index; WBC, white blood cell count.
In the LightGBM model, age, Ualb and TC emerged as the three most influential predictors, as indicated by their wide SHAP value distributions and high feature importance rankings. SHAP values for age were predominantly positive, demonstrating that increasing age is strongly associated with a higher predicted risk of diabetes. Likewise, Ualb consistently contributed positively to the model’s output, highlighting the potential role of renal dysfunction as a key indicator of diabetes risk in individuals with MASLD. TC also showed a meaningful impact on prediction, further underscoring the relevance of lipid metabolism in diabetes pathogenesis within this population.
Discussion
This study leveraged data from the NHANES database (2007–2018) to develop a predictive model for diabetes risk among individuals with MASLD. By integrating a wide range of variables—including demographic characteristics, biochemical markers, medical history and dietary intake—we systematically evaluated the performance of multiple ML algorithms. Through the combined application of three feature selection methods—LASSO, RF and SVM—we identified eight robust and biologically meaningful core predictors: age, Ualb, TC, LAP, Ucr, WBC, UA and VAI. Based on these features, nine ML models were constructed, all demonstrating favourable predictive performance, with LightGBM outperforming the others. SHAP analysis was employed to interpret model predictions, revealing Ualb, age, TC and LAP as the most influential contributors. These key variables are closely aligned with established mechanisms of diabetes development, including metabolic dysregulation, renal impairment and abnormal fat distribution, thereby reinforcing the biological plausibility and clinical relevance of the proposed predictive model.
Metabolic inter-relationship and biological interpretation between MASLD and diabetes
MASLD and diabetes share a closely intertwined metabolic and pathological foundation. As a hepatic manifestation of insulin resistance, MASLD is recognised not only as an independent risk factor for diabetes but also as a potential prediabetic condition.25 26 On the one hand, MASLD contributes to hepatic glucose overproduction, adipose tissue inflammation and chronic low-grade systemic inflammation, all of which disrupt insulin signalling and accelerate the development and progression of diabetes.27 28 On the other hand, diabetes may reciprocally exacerbate hepatic lipid accumulation, oxidative stress and endoplasmic reticulum stress, thereby driving the progression of MASLD towards NASH and hepatic fibrosis.29 30 Clinical evidence suggests that individuals with coexisting MASLD and diabetes exhibit more severe metabolic derangements, heightened cardiovascular risk and greater renal impairment compared with those with either condition alone, indicating a potential bidirectional and self-reinforcing pathological loop between the two diseases.31 32
Notably, individuals with MASLD often exhibit a constellation of metabolic abnormalities, including abdominal obesity, hyperuricaemia, chronic low-grade inflammation, dyslipidaemia and early renal dysfunction.33–35 These abnormalities constitute core pathophysiological drivers of diabetes and align closely with the key predictors identified in this study. Specifically, LAP and VAI reflect fat accumulation and visceral adiposity; Ualb and Ucr serve as markers of glomerular filtration and tubular function; WBC captures the presence of systemic inflammation; and UA and TC represent indicators of metabolic dysregulation. This mechanistic concordance not only reinforces the biological plausibility of the predictive model but also highlights the clinical necessity of implementing diabetes risk assessment strategies within the MASLD population. Early identification of high-risk individuals through such targeted models holds promise for guiding personalised interventions and mitigating the dual burden of MASLD and diabetes.
In the predictive model constructed in this study, renal function-related indicators—Ualb and Ucr—presented high importance in SHAP analysis, suggesting that renal stress may represent an early and meaningful precursor in the pathogenesis of diabetes among individuals with MASLD. This finding is consistent with prior research indicating that, under obese conditions, glomerular hyperfiltration, tubular reabsorption abnormalities and lipid-induced toxicity collectively contribute to early renal dysfunction and the emergence of microalbuminuria.36–38 By assigning high variable importance to Ualb and Ucr, our model quantitatively reinforces the emerging concept of ‘metabolic-associated kidney disease’ as a critical component in diabetes risk stratification frameworks. In addition to renal markers, several metabolic indicators—such as TG and TC—also emerged as consistent and influential predictors across models. These lipid parameters, strongly associated with hepatic lipid overproduction and impaired lipid clearance, reflect canonical features of insulin resistance and β-cell dysfunction, both of which are central to diabetes pathogenesis.39 40 Furthermore, the inclusion of two adiposity indices—VAI and LAP—extends the model’s capacity beyond traditional metabolic markers by capturing the effects of visceral fat distribution, inflammation and adipose tissue-driven insulin resistance. Interestingly, although some nutrition intake variables were retained during initial feature selection, they were not preserved in the final predictive model. This suggests that dietary behaviours may exert their influence indirectly, primarily through intermediary processes such as lipid metabolism and renal impairment, rather than serving as independent predictors. Collectively, these results underscore the importance of incorporating multidimensional biochemical and metabolic variables—rather than relying solely on behavioural data—when constructing predictive models for diabetes in metabolically at-risk populations such as those with MASLD, thereby enhancing both accuracy and biological interpretability.
Clinical implications and perspectives for diabetes prevention
This study represents a multidimensional advancement in diabetes risk prediction by focusing on individuals with MASLD, a metabolically heterogeneous and high-risk population. Unlike traditional models that rely on linear assumptions and limited feature interactions, we incorporated a suite of nonlinear ML algorithms—including LightGBM, MLP and SVM—and conducted a systematic evaluation of their predictive accuracy, robustness and interpretability. Among them, the LightGBM model showed the most consistent and reliable performance across key metrics such as accuracy, AUC and F1 score, with minimal disparity between training and testing sets, thereby underscoring its strong generalisability and practical applicability. To enhance transparency and overcome the ‘black box’ limitations often encountered in complex models, we used SHAP to deconstruct and visualise individual feature contributions. Notably, Ualb, age, TC and LAP consistently emerged as top-ranking variables across models, reinforcing the model’s biological plausibility and clinical relevance. Compared with general population-based models, our focus on the MASLD subpopulation—considered a prediabetic phenotype—enabled the early identification of subtle yet critical metabolic perturbations prior to the onset of overt diabetes, thereby supporting more precise and proactive risk stratification. This targeted approach not only improves the efficiency and precision of screening strategies but also provides a theoretical and methodological foundation for early intervention and stratified management of diabetes. At the broader public health level, such precision-oriented predictive tools tailored to specific high-risk groups may serve as powerful adjuncts in developing efficient and scalable metabolic disease prevention frameworks.
Limitations
This study has several limitations. First, the NHANES database is cross-sectional, which precludes the determination of temporal or causal relationships between MASLD and diabetes. The exact time of disease onset or duration could not be established; therefore, our model identifies individuals with a higher likelihood of diabetes at the time of survey rather than predicting future disease development. Second, although the model demonstrated good internal performance, external validation using independent datasets was not feasible due to data constraints. Future studies incorporating longitudinal cohort data or real-world clinical data are warranted to evaluate the model’s predictive ability over time and to enhance its clinical applicability in dynamic risk assessment.
Conclusion
This study developed a diabetes risk prediction model for MASLD patients using NHANES data and identified eight key predictors, including age, Ualb, TC and LAP. Among several ML algorithms, LightGBM showed the best performance with strong predictive power and interpretability. SHAP analysis highlighted the critical roles of renal function, lipid metabolism and inflammation. The model holds promise for early screening and personalised prevention of diabetes in MASLD populations.
bmjopen-16-1-s005.pdf (111.1KB, pdf)
Supplementary Material
Footnotes
SY, PZ and WH contributed equally.
Contributors: PZ designed the study. WW, RM and WH conducted data extraction and interpreted the study results. SY and WH wrote and revised the manuscript. All authors have read and agreed to the published version of the manuscript. SY is identified as the guarantor.
Funding: This study was supported by grants from the National Natural Science Foundation of China (82260448), Guangxi Natural Science Foundation of China (2024GXNSFDA010024), Guangxi Key Research and Development Plan (2021AB11027) and State Key Laboratory of Advanced Optical Communication Systems and Networks, China (2024GZKF13).
Competing interests: None declared.
Patient and public involvement: Patients and/or the public were not involved in the design, conduct, reporting or dissemination plans of this research.
Provenance and peer review: Not commissioned; externally peer reviewed.
Supplemental material: This content has been supplied by the author(s). It has not been vetted by BMJ Publishing Group Limited (BMJ) and may not have been peer-reviewed. Any opinions or recommendations discussed are solely those of the author(s) and are not endorsed by BMJ. BMJ disclaims all liability and responsibility arising from any reliance placed on the content. Where the content includes any translated material, BMJ does not warrant the accuracy and reliability of the translations (including but not limited to local regulations, clinical guidelines, terminology, drug names and drug dosages), and is not responsible for any error and/or omissions arising from translation and adaptation or otherwise.
Data availability statement
Data are available in a public, open access repository. Data are available upon reasonable request. Data may be obtained from a third party and are not publicly available. The data used in this study are publicly available from the National Health and Nutrition Examination Survey (NHANES) database, which is maintained by the Centers for Disease Control and Prevention (CDC). The datasets analysed from the 2007–2018 cycles can be accessed at: https://wwwn.cdc.gov/nchs/nhanes/default.aspx.
Ethics statements
Patient consent for publication
Not applicable.
Ethics approval
As NHANES is a publicly accessible dataset that has been approved by the Institutional Review Board (IRB) of the National Center for Health Statistics (NCHS), our institution confirmed that no additional ethical approval was necessary. Moreover, the IRB acknowledges that NCHS follows strict ethical protocols during data collection and processing, including obtaining informed consent from all participants and anonymising personal data. These procedures ensure full compliance with ethical standards for secondary data analysis.
References
- 1.Byrne CD, Targher G. NAFLD: A multisystem disease. J Hepatol 2015;62:S47–64. 10.1016/j.jhep.2014.12.012 [DOI] [PubMed] [Google Scholar]
- 2.Cotter TG, Rinella M. Nonalcoholic Fatty Liver Disease 2020: The State of the Disease. Gastroenterology 2020;158:1851–64. 10.1053/j.gastro.2020.01.052 [DOI] [PubMed] [Google Scholar]
- 3.Wattacheril JJ, Abdelmalek MF, Lim JK, et al. AGA Clinical Practice Update on the Role of Noninvasive Biomarkers in the Evaluation and Management of Nonalcoholic Fatty Liver Disease: Expert Review. Gastroenterology 2023;165:1080–8. 10.1053/j.gastro.2023.06.013 [DOI] [PubMed] [Google Scholar]
- 4.Wang J-L, Jiang S-W, Hu A-R, et al. Non-invasive diagnosis of non-alcoholic fatty liver disease: Current status and future perspective. Heliyon 2024;10:e27325. 10.1016/j.heliyon.2024.e27325 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Bhat N, Mani A. Dysregulation of Lipid and Glucose Metabolism in Nonalcoholic Fatty Liver Disease. Nutrients 2023;15:2323. 10.3390/nu15102323 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Rafaqat S, Gluscevic S, Mercantepe F, et al. Interleukins: Pathogenesis in Non-Alcoholic Fatty Liver Disease. Metabolites 2024;14:153. 10.3390/metabo14030153 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chao H-W, Chao S-W, Lin H, et al. Homeostasis of Glucose and Lipid in Non-Alcoholic Fatty Liver Disease. Int J Mol Sci 2019;20:298. 10.3390/ijms20020298 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Dinić S, Arambašić Jovanović J, Uskoković A, et al. Oxidative stress-mediated beta cell death and dysfunction as a target for diabetes management. Front Endocrinol (Lausanne) 2022;13:1006376. 10.3389/fendo.2022.1006376 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Eguchi N, Vaziri ND, Dafoe DC, et al. The Role of Oxidative Stress in Pancreatic β Cell Dysfunction in Diabetes. Int J Mol Sci 2021;22:1509. 10.3390/ijms22041509 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.National health and nutrition examination survey. n.d. Available: https:// www.cdc.gov/nchs/index.htm
- 11.Lee Y, Bang H, Park YM, et al. Non–Laboratory-Based Self-Assessment Screening Score for Non-Alcoholic Fatty Liver Disease: Development, Validation and Comparison with Other Scores. PLoS ONE 2014;9:e107584. 10.1371/journal.pone.0107584 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Chon YE, Jung KS, Kim SU, et al. Controlled attenuation parameter (CAP) for detection of hepatic steatosis in patients with chronic liver diseases: a prospective study of a native Korean population. Liver Int 2014;34:102–9. 10.1111/liv.12282 [DOI] [PubMed] [Google Scholar]
- 13.Shih K-L, Su W-W, Chang C-C, et al. Comparisons of parallel potential biomarkers of 1H-MRS-measured hepatic lipid content in patients with non-alcoholic fatty liver disease. Sci Rep 2016;6:24031. 10.1038/srep24031 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.ElSayed NA, Aleppo G, Aroda VR. 2. Classification and Diagnosis of Diabetes: Standards of Care in Diabetes—2023. Diabetes Care 2023;46:S19–40. 10.2337/dc23-ad08 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Yang B, Lu H, Ran Y. Advancing non-alcoholic fatty liver disease prediction: a comprehensive machine learning approach integrating SHAP interpretability and multi-cohort validation. Front Endocrinol 2024;15:1450317. 10.3389/fendo.2024.1450317 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. 10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Herrmann SD, Willis EA, Ainsworth BE. The 2024 Compendium of Physical Activities and its expansion. J Sport Health Sci 2024;13:1–2. 10.1016/j.jshs.2023.09.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Kahn HS. The “lipid accumulation product” performs better than the body mass index for recognizing cardiovascular risk: a population-based comparison. BMC Cardiovasc Disord 2005;5:26. 10.1186/1471-2261-5-26 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Amato MC, Giordano C, Galia M, et al. Visceral Adiposity Index: a reliable indicator of visceral fat function associated with cardiometabolic risk. Diabetes Care 2010;33:920–2. 10.2337/dc09-1825 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Chen Q, Meng Z, Liu X, et al. Decision Variants for the Automatic Determination of Optimal Feature Subset in RF-RFE. Genes (Basel) 2018;9:301. 10.3390/genes9060301 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Humphrey A, Kuberski W, Bialek J, et al. Machine-learning classification of astronomical sources: estimating F1-score in the absence of ground truth. Mon Not R Astron Soc 2022;517:L116–20. 10.1093/mnrasl/slac120 [DOI] [Google Scholar]
- 22.Zamo M, Naveau P. Estimation of the Continuous Ranked Probability Score with Limited Information and Applications to Ensemble Weather Forecasts. Math Geosci 2018;50:209–34. 10.1007/s11004-017-9709-7 [DOI] [Google Scholar]
- 23.Assel M, Sjoberg DD, Vickers AJ. The Brier score does not evaluate the clinical utility of diagnostic tests or prediction models. Diagn Progn Res 2017;1:19. 10.1186/s41512-017-0020-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Fu J, Wu Y, Feng H, et al. Development of a nomogram for predicting the outcome in patients with prolonged disorders of consciousness based on the multimodal evaluative information. BMC Neurol 2025;25:175. 10.1186/s12883-025-04189-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Jung I, Koo D-J, Lee W-Y. Insulin Resistance, Non-Alcoholic Fatty Liver Disease and Type 2 Diabetes Mellitus: Clinical and Experimental Perspective. Diabetes Metab J 2024;48:327–39. 10.4093/dmj.2023.0350 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Gariani K, Philippe J, Jornayvaz FR. Non-alcoholic fatty liver disease and insulin resistance: from bench to bedside. Diabetes Metab 2013;39:16–26. 10.1016/j.diabet.2012.11.002 [DOI] [PubMed] [Google Scholar]
- 27.Armandi A, Rosso C, Caviglia GP, et al. Insulin Resistance across the Spectrum of Nonalcoholic Fatty Liver Disease. Metabolites 2021;11:155. 10.3390/metabo11030155 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Dharmalingam M, Yamasandhi PG. Nonalcoholic fatty liver disease and Type 2 diabetes mellitus. Indian J Endocr Metab 2018;22:421. 10.4103/ijem.IJEM_585_17 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Martín-Fernández M, Arroyo V, Carnicero C, et al. Role of Oxidative Stress and Lipid Peroxidation in the Pathophysiology of NAFLD. Antioxidants (Basel) 2022;11:2217. 10.3390/antiox11112217 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Ma Y, Lee G, Heo S-Y, et al. Oxidative Stress Is a Key Modulator in the Development of Nonalcoholic Fatty Liver Disease. Antioxidants (Basel) 2021;11:91. 10.3390/antiox11010091 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Caussy C, Aubin A, Loomba R. The Relationship Between Type 2 Diabetes, NAFLD, and Cardiovascular Risk. Curr Diab Rep 2021;21:15. 10.1007/s11892-021-01383-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Targher G, Corey KE, Byrne CD. NAFLD, and cardiovascular and cardiac diseases: Factors influencing risk, prediction and treatment. Diabetes Metab 2021;47:101215. 10.1016/j.diabet.2020.101215 [DOI] [PubMed] [Google Scholar]
- 33.Paschos P, Paletas K. Non alcoholic fatty liver disease and metabolic syndrome. Hippokratia 2009;13:9–19. [PMC free article] [PubMed] [Google Scholar]
- 34.Carrillo-Larco RM, Guzman-Vilca WC, Castillo-Cara M, et al. Phenotypes of non-alcoholic fatty liver disease (NAFLD) and all-cause mortality: unsupervised machine learning analysis of NHANES III. BMJ Open 2022;12:e067203. 10.1136/bmjopen-2022-067203 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Huang Y, Jin T, Ni W, et al. Baseline and change in serum lipid and uric acid level over time and incident of nonalcoholic fatty liver disease (NAFLD) in Chinese adults. Sci Rep 2024;14:18547. 10.1038/s41598-024-69411-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Nawaz S, Chinnadurai R, Al-Chalabi S, et al. Obesity and chronic kidney disease: A current review. Obes Sci Pract 2023;9:61–74. 10.1002/osp4.629 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Ye M, Yang M, Dai W, et al. Targeting Renal Proximal Tubule Cells in Obesity-Related Glomerulopathy. Pharmaceuticals (Basel) 2023;16:1256. 10.3390/ph16091256 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Wei L, Li Y, Yu Y, et al. Obesity-Related Glomerulopathy: From Mechanism to Therapeutic Target. Diabetes Metab Syndr Obes 2021;14:4371–80. 10.2147/DMSO.S334199 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Chen Y, Jiang H, Zhan Z, et al. Restoration of lipid homeostasis between TG and PE by the LXRα-ATGL/EPT1 axis ameliorates hepatosteatosis. Cell Death Dis 2023;14:85. 10.1038/s41419-023-05613-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Alves-Bezerra M, Cohen DE. Triglyceride Metabolism in the Liver. Compr Physiol 2017;8:1–8. 10.1002/cphy.c170012 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
bmjopen-16-1-s003.pdf (123.1KB, pdf)
bmjopen-16-1-s001.pdf (341.7KB, pdf)
bmjopen-16-1-s002.pdf (271.7KB, pdf)
bmjopen-16-1-s004.pdf (88.5KB, pdf)
bmjopen-16-1-s005.pdf (111.1KB, pdf)
Data Availability Statement
Data are available in a public, open access repository. Data are available upon reasonable request. Data may be obtained from a third party and are not publicly available. The data used in this study are publicly available from the National Health and Nutrition Examination Survey (NHANES) database, which is maintained by the Centers for Disease Control and Prevention (CDC). The datasets analysed from the 2007–2018 cycles can be accessed at: https://wwwn.cdc.gov/nchs/nhanes/default.aspx.

