Abstract
Background
Previous studies have emphasized the critical role of diet and gut microbiome in Metabolic syndrome (MetS). The dietary index for gut microbiota (DI-GM) represents a novel dietary index that effectively reflects the diversity of gut microbiota; nevertheless, its applicability to MetS and its components remains unknown.
Methods
For this study, we enrolled 19,702 individuals from NHANES 2007–2020. DI-GM comprises dietary information of 14 dietary components, including 10 beneficial and 4 unfavorable ones. Weighted logistic regressions evaluated associations of DI-GM with MetS and its components, whereas weighted linear regression analyzed its association with 6 MetS-related biochemical indicators. Modified Poisson regression, sensitivity analyses after multiple imputation and subgroup analyses ensured robustness. Restricted cubic spline (RCS) analysis explored whether a non-linear relationship exists. Nine machine-learning models were developed for MetS prediction, and six discrimination characteristics selected the optimal model. SHapley Additive exPlanations (SHAP) was utilized to interpret the contributions of variables for model decision-making capacity.
Results
After fully adjusting for confounders, the DI-GM score exhibited a noticeable negative correlation with the prevalence of MetS (OR: 0.95, 95% CI: 0.93–0.97, P-value < 0.001), along with elevated waist circumference (OR: 0.91, 95% CI: 0.88–0.94, P-value < 0.001), elevated blood pressure (OR: 0.95, 95% CI: 0.93–0.98, P-value < 0.001), reduced high-density lipoprotein (OR: 0.95, 95% CI: 0.93–0.98, P-value < 0.001) and elevated fasting blood glucose (OR: 0.94, 95% CI: 0.90–0.98, P-value = 0.002). RCS exhibited a significant inverse association of DI-GM with MetS for non-linear relationship when met the score of 5. Subgroup analysis demonstrated that the association remained stable and consistent across the majority of the subgroups. XGboost presented superior performance and SHAP analysis revealed that higher DI-GM exhibited considerable inverse influence, ranking after “BMI ≥ 30”, “Age”, “Race-Non-Hispanic Black” and “CKD-Yes”.
Conclusions
Our study presents compelling evidence that higher scores of the DI-GM are associated with a lower prevalence of Mets and its components. Dietary strategies that incorporate the DI-GM score could contribute to the harmonious ecological state of the gut microbiome and be crucial in the prevention of MetS.
Supplementary Information
The online version contains supplementary material available at 10.1186/s41043-025-01137-1.
Keywords: DI-GM, Gut microbiota, Metabolic syndrome, Machine learning, NHANES
Introduction
Metabolic syndrome (MetS) comprises a combination of metabolic disturbances including high blood pressure, raised fasting glucose, dyslipidemia, and central obesity [1]. MetS has developed into a serious public health concern. From 2000 to 2020, the prevalence of MetS in the U.S. experienced significant increases from 26.7% to 45.9%, marking a substantial increase of 71.9% [2, 3]. The accumulation of the different components of MetS considerably elevates the risk of chronic diseases such as chronic kidney disease (CKD) and cardiovascular disease (CVD), ultimately leading to earlier mortality [4, 5]. MetS develops through a multifaceted interplay of various factors, encompassing adverse dietary patterns and poor lifestyles being major contributors [1, 6]. A broad consensus suggests that prevention rather than treatment should be the primary targeted [1, 7].
In recent years, gut microbiota research has emerged as a prominent focus [8, 9]. The gut microbiome assumes a vital role in nutrition, metabolism, immune function, and various other dimensions of health [10, 11]. Existing studies have observed alterations in gut microbiota composition and abundance between MetS patients and healthy individuals [12]. Dietary pattern presents one of the crucial determinants shaping the composition of the gut microbiome, thus targeted approaches for gut microbiota modulation have demonstrated promising potential [13]. The Dietary Index for Gut Microbiota (DI-GM), an innovative metric developed by Kase et al., identified 14 food categories or nutrients that could enhance or impair the gut health [14]. The DI-GM can capture variations in gut microbiota diversity, short-chain fatty acid (SCFA) production quantities, and changes within certain bacterial phyla, enabling its application across a broad spectrum of research beyond individual bacterial clusters [14, 15]. This dietary index offers individuals a multifaceted and dependable evaluation of their dietary quality in relation to ecological balance of the gut microbiome, facilitating suitable dietary interventions for enhanced health. At present, several studies have investigated the association between the DI-GM and other medical disorders [16–18]; nevertheless, its applicability to metabolic syndrome and its components remains unknown.
Traditional statistical methods for disease identification entail specific prerequisites for data preparation and a large quantity of structured data with great quality of distribution [19–21]. Fortunately, Machine Learning (ML) presents the possibility for us to identify interactions among multiple variables in large dataset [22]. ML applies intricate mathematical algorithms for detecting and classifying patterns in diverse, complex datasets, with the objective of assisting and optimizing decision-making [23, 24]. These sophisticated techniques are capable of processing extensive and sparse data matrices, thereby enabling the analysis of a greater quantity of information, promoting hazard detection, and guiding health-related decision-making [25]. For interpreting ML prediction results, SHapley Additive exPlanation (SHAP) provides a clear explanation for each feature’s contribution, serving as a reliable and promising tool for feature selection in medical diagnosis [26].
Our study aims to provide evidence-based dietary strategies for promoting gut microbiome homeostasis and preventing MetS through the application of DI-GM scoring. We performed a cross-sectional analysis based on the available data from National Health and Nutrition Examination Survey (NHANES) 2007–2020. A variety of statistical methods were employed to thoroughly assess the association between DI-GM and MetS.
Materials and methods
The participants in the study
The NHANES employed a multistage and intricately stratified sampling design throughout the country and has established itself as a large-scale updated database. Prior to being incorporated into the database, each individual had already provided written informed consent, and this consent had received approval from the National Center for Health Statistics Ethics Review Board (NCHS).
Our study performed the analysis based on the available data obtained from NHANES 2007–2020, during which all relevant indicators were consistently included in the survey questionnaires, avoiding partial availability due to variations across cycles. A total of 66,148 participants were initially recruited for our analysis. Firstly, participants under 20 years of age were excluded (N = 27,715), as metabolic syndrome diagnosis in this study followed the NCEP ATPIII-2005 criteria (National Cholesterol Education Program Adult Treatment Panel III)—widely used in clinical and epidemiological research—which focus on adults aged ≥ 20 years. As pregnancy greatly perturbs the human gut microbiome, pregnant individuals were excluded from participant selection (N = 298). Secondly, we excluded participants with missing data of blood pressure (N = 2,951), high-density lipoprotein(N = 2,686), total triglycerides (N = 131), waist circumference (N = 1,328) and our main exposure DI-GM (N = 8,160). Finally, participants lacking information on covariates were also excluded from our research (N = 3,177). After the aforementioned screening procedures, the study encompassed a complete cohort comprising 19,702 participants. (Fig. 1).
Fig. 1.
Flowchart of the selection of participants for this study from NHANES 2007–2020
Diagnosis of MetS
MetS was diagnosed following the NCEP ATPIII-2005 criteria, which have been widely used in clinical and epidemiological research [27, 28]. Specifically, the diagnosis was established when an individual satisfied three or more of the following criteria: (1) Elevated WC, characterized by attaining a value of 102 cm or more in males and 88 cm or more in females; (2) Elevated BP, characterized by blood pressure exceeding 130/85 mmHg, or under drug treatment; (3) Reduced HDL-C, established when levels measured less than 50 mg/dL in women or less than 40 mg/dL in men, or receiving lipid-lowering treatment; (4) Elevated TG, defined as TG concentration being equal to or exceeding 150 mg/dL, or the current use of triglyceride-lowering medications; (5) Elevated FG, defined as >100 mg/dL or current use of glucose-lowering medications.
Assessment of DI-GM
The dietary data was obtained using two 24-hour dietary recollection interviews [29]. The initial interview was performed at the Mobile Examination Center, requiring participants to report every food and beverage consumed in the preceding 24 h. The follow-up telephone interview occurred 3 to 10 days after the initial data collection. Data from the two assessment days were averaged to derive the value of each dietary component. The assessment of DI-GM involves 14 food categories and nutrients. Beneficial components encompass avocados, broccoli, chickpeas, coffee, cranberries, fermented dairy products, fiber, soybeans, whole grains, and green tea. Conversely, refined grains, processed meats, red meat, and high-fat diets (characterized by ≥ 40% of energy from fat) are considered detrimental. Given that the dietary data in NHANES lack explicit records of tea consumption, the DI-GM scoring system spans from 0 to 13. DI-GM scores are calculated using sex-specific medians with consideration of the marked differences between men and women in physiological intake and dietary preference. Participants were assigned a score of 1 if they consumed above the sex-specific median for beneficial components or below the sex-specific median for unfavorable components. Conversely, a score of 0 was given to those who consumed below the sex-specific median for beneficial components or above the sex-specific median for unfavorable components. More comprehensive and in-depth details regarding the composition and calculation methodology of the DI-GM was presented in Table S1.
Covariates
Drawing upon prior investigations, we identified several potential confounding factors in our analysis, including age, sex, ethnicity, education, poverty-to-income ratio (PIR), smoking status, alcohol consumption, body mass index (BMI), energy intake, chronic kidney disease and cardiovascular disease. Age, sex, race and socioeconomic factors are well-established demographic correlates of metabolic dysfunction [6]. BMI, cigarette use, and drinking status are lifestyle factors linked to metabolic health [30–32]. Energy intake is a key confounder in diet-health associations [33]. CKD and CVD are recognized comorbidities often clustered with MetS [4, 5]. The PIR, reflecting economic status, has been pre-calculated and can be obtained from the NHANES documentation. We classified smoking status according to the following criteria: (1) Current smoker, defined as smoking exceeding 100 cigarettes in their lifetime and are now involved in smoking behavior; (2) Former smoker, who have abstained from smoking until now; (3) Never smoker, individuals who had smoked fewer than 100 cigarettes throughout their entire lifetime. Our study divided alcohol consumption into two distinct groups, defined by whether or not consumed alcohol more than 12 times a year. BMI was categorized into two groups by the value of 30 kg/m2. CKD was diagnosed when albuminuria was present (characterized by a urinary albumin-to-creatinine ratio exceeding 30 mg/g) or when there was a reduced estimated glomerular filtration rate (specified as an eGFR of 60 mL/min/1.73 m² or lower). CVD was defined as a physician-diagnosed condition encompassing one or more of the following: angina, heart attack, congestive heart failure, and coronary artery disease.
Statistical analysis
All statistical methods were conducted using R software Version 4.4.2 in our study. To generate estimates that were representative of the entire national population, the sampling design integrated stratification, clustering, and differential selection probabilities. “SDMVSTRA” enabled stratification by grouping samples according to geographic and demographic factors to address population heterogeneity. “SDMVPSU” was utilized to address clustering by denoting the primary sampling unit (PSU), capturing the within-PSU observation correlation. Population representativeness was achieved by applying “WTINT2YR”, the interview sampling weights, to adjust for unequal selection probabilities. Initially, participants were stratified into two groups according to the diagnosis of MetS, and then the demographic characteristics between these groups were compared. The Wilcoxon rank-sum test was adopted for the analysis of continuous variables, with the chi-square test being used for categorical variables. A series of multivariate weighted logistic and modified Poisson regression were employed to assess the associations of DI-GM, both as a continuous variable and its quartile subgroups, with the prevalence of MetS and its components. A series of multivariate weighted linear regression to examine the association of DI-GM with six MetS related biochemical indicators (waist circumference, systolic blood pressure, diastolic blood pressure, high-density lipoprotein, total triglycerides and fasting blood glucose). Model 1 represented the crude analysis without adjustment for confounding factors. Model 2 was adjusted for age, sex, ethnicity, education, PIR, and BMI. Model 3 was additionally adjusted for smoking status, alcohol consumption, CKD, and CVD. Furthermore, to assess whether missing covariates affected the research results, we performed multiple imputation to generate new complete datasets and replicated the aforementioned regression models as the sensitivity analysis. Missing data were handled using a well-established method via the R “mice” package [34, 35]. The MICE algorithm employed predictive mean matching, utilizing a chained equation approach as the imputation algorithm with 5 imputed datasets, where missing entries were probabilistically replaced based on the actual distribution of the observed data. Given the characteristics of study variables and observed missing patterns, we assumed a missing-at-random (MAR) mechanism for variables with missing data in our research (blood pressure, HDL, TG, WC, education, PIR, BMI, smoking status, alcohol consumption, chronic kidney disease and cardiovascular diseases). Under this mechanism, the probability of missingness depends only on observed data, rendering multiple imputation via MICE methodologically appropriate. Subgroup analysis enables us to identify heterogeneity across these relationships and enhances the extent of the research findings among different groups. Restricted cubic spline (RCS) analysis was conducted to explore whether a non-linear relationship exists. For findings of all statistical methods, a two-tailed P-value < 0.05 were considered significant.
Model development for machine learning
Nine machine learning models were developed to identify the contributions of input relevant variables and assess their positive/negative associations across the MetS prediction. Our machine learning model adopted the following variable inclusion approach: demographic variables (age, sex, race, education, PIR), clinical indicators not used in outcome definition (BMI, energy intake, CKD and CHD), lifestyle variables (smoking status, alcohol consumption), and our primary novel indicator (DI-GM). The machine learning dataset was randomly divided into training and testing subsets at a 7:3 ratio, following a pre-specified sampling strategy. Multicollinearity among variables was evaluated using the variance inflation factor (VIF) method. Any variable with a VIF score surpassing 5 was excluded from the study’s variable set. Five-fold cross-validation was implemented on the training dataset to iteratively perform testing and optimize hyperparameters, thereby ensuring the model’s robustness and identifying the optimal parameter configuration. Nine algorithms were utilized to analyze the selected variables, including Decision Tree (DT), Elastic Net (ENET), K-nearest Neighbors (KNN), Light Gradient Boosting Machine (LightGBM), Logistic Regression (LR), Multilayer Perceptron (MLP), Random Forest (RF), Support Vector Machine (SVM) and Extreme Gradient Boosting (XGBoost). Model accuracy and clinical prediction performance were assessed using a set of evaluation measures: accuracy, the area under the receiver operating characteristic curve (AUC), precision, sensitivity/recall, F1 score and specificity. Calibration curves served as the method for judging the accuracy of absolute risk prediction. Decision curve analysis (DCA) was then employed to assess the clinical net benefit of the model over different threshold probabilities. The SHAP method was employed to evaluate the contribution of variables for model interpretability. The method produced a predicted value for each sample, depicted in the SHAP Summary plot where purple dots symbolize high feature values while yellow dots represent low feature values.
Results
Demographic features of the participants
The demographic features of 19,702 participants were presented in Table 1, consisting of 6,689 individuals with MetS and 13,013 individuals without MetS. Statistical significance was observed for the differences between the two groups in age, race, education, PIR, BMI, smoking status, alcohol consumption, energy intake, CKD and CVD. Moreover, a significant difference of DI-GM was also observed (P-value<0.05).
Table 1.
Basic characteristics of the study population
| Characteristic | Overall (N = 197021) |
Non-MetS (N = 130131) |
MetS (N = 66891) |
P-value |
|---|---|---|---|---|
| Age (years) | 47.34 (16.62) | 44.59 (16.54) | 53.53 (15.07) | < 0.001 |
| Age groups | < 0.001 | |||
| ≤ 39 | 6,631 (36%) | 5,428 (43%) | 1,203 (20%) | |
| 40–59 | 6,659 (39%) | 4,216 (37%) | 2,443 (43%) | |
| ≥ 60 | 6,412 (26%) | 3,369 (21%) | 3,043 (37%) | |
| Sex | 0.5 | |||
| Female | 9,813 (50%) | 6,309 (50%) | 3,504 (51%) | |
| Male | 9,889 (50%) | 6,704 (50%) | 3,185 (49%) | |
| Race | < 0.001 | |||
| Mexican American | 2,970 (8.0%) | 1,780 (7.6%) | 1,190 (8.9%) | |
| Other Hispanic | 2,020 (5.2%) | 1,276 (5.2%) | 744 (5.1%) | |
| Non-Hispanic White | 8,919 (70%) | 5,836 (70%) | 3,083 (71%) | |
| Non-Hispanic Black | 3,912 (10%) | 2,709 (10%) | 1,203 (9.2%) | |
| Other/multiracial | 1,881 (6.6%) | 1,412 (7.0%) | 469 (5.6%) | |
| Education | < 0.001 | |||
| Under high school | 4,596 (15%) | 2,677 (14%) | 1,919 (19%) | |
| High school | 4,490 (22%) | 2,838 (21%) | 1,652 (25%) | |
| More than high school | 10,616 (63%) | 7,498 (66%) | 3,118 (56%) | |
| PIR groups | < 0.001 | |||
| < 1.0 | 4,144 (14%) | 2,649 (14%) | 1,495 (15%) | |
| ≥ 1, < 3 | 8,223 (36%) | 5,206 (34%) | 3,017 (39%) | |
| ≥ 3 | 7,335 (50%) | 5,158 (52%) | 2,177 (46%) | |
| BMI (kg/m2 ) | 28.96 (6.65) | 27.08 (5.78) | 33.21 (6.53) | < 0.001 |
| BMI groups | < 0.001 | |||
| < 30 | 12,190 (63%) | 9,830 (77%) | 2,360 (33%) | |
| ≥ 30 | 7,512 (37%) | 3,183 (23%) | 4,329 (67%) | |
| Smoking Status | < 0.001 | |||
| Current smoker | 4,108 (20%) | 2,773 (20%) | 1,335 (20%) | |
| Former smoker | 4,850 (25%) | 2,888 (23%) | 1,962 (30%) | |
| Never smoker | 10,742 (55%) | 7,351 (57%) | 3,391 (50%) | |
| Alcohol consumption | < 0.001 | |||
| Yes | 14,406 (79%) | 9,840 (81%) | 4,566 (74%) | |
| No | 5,296 (21%) | 3,173 (19%) | 2,123 (26%) | |
| Elevated WC | < 0.001 | |||
| Yes | 11,304 (57%) | 5,219 (41%) | 6,085 (93%) | |
| No | 8,398 (43%) | 7,794 (59%) | 604 (7.3%) | |
| Elevated BP | < 0.001 | |||
| Yes | 9,675 (44%) | 4,356 (29%) | 5,319 (77%) | |
| No | 10,027 (56%) | 8,657 (71%) | 1,370 (23%) | |
| Reduced HDL-C | < 0.001 | |||
| Yes | 6,322 (31%) | 1,922 (14%) | 4,400 (68%) | |
| No | 13,380 (69%) | 11,091 (86%) | 2,289 (32%) | |
| Elevated TG | < 0.001 | |||
| Yes | 7,474 (37%) | 2,360 (19%) | 5,114 (80%) | |
| No | 12,228 (63%) | 10,653 (81%) | 1,575 (20%) | |
| Elevated FG | < 0.001 | |||
| Yes | 3,274 (12%) | 522 (2.8%) | 2,752 (34%) | |
| No | 16,428 (88%) | 12,491 (97%) | 3,937 (66%) | |
| Chronic kidney disease | < 0.001 | |||
| Yes | 3,447 (14%) | 1,587 (9.8%) | 1,860 (23%) | |
| No | 16,255 (86%) | 11,426 (90%) | 4,829 (77%) | |
| Cardiovascular disease | < 0.001 | |||
| Yes | 1,988 (8.0%) | 862 (5.2%) | 1,126 (14%) | |
| No | 17,714 (92%) | 12,151 (95%) | 5,563 (86%) | |
| Energy Intake (kcal/day) | 2,126.59 (845.66) | 2,150.24 (854.67) | 2,073.27 (822.57) | < 0.001 |
| DI-GM | 4.88 (1.68) | 4.95 (1.70) | 4.73 (1.64) | < 0.001 |
| DI-GM groups | < 0.001 | |||
| Quartile 1(0–4) | 8,873 (43%) | 5,704 (41%) | 3,169 (46%) | |
| Quartile 2(5) | 4,508 (23%) | 2,975 (23%) | 1,533 (23%) | |
| Quartile 3(6) | 3,344 (18%) | 2,237 (18%) | 1,107 (17%) | |
| Quartile 4(7–11) | 2,977 (17%) | 2,097 (18%) | 880 (14%) |
1Mean (SD); n (unweighted) (%). MetS, Metabolic syndrome; PIR, family income-to-poverty ratio; WC, Waist Circumference (cm); BP, blood pressure; HDL-C, high-density lipoprotein; TG, total triglycerides; FG, fasting blood glucose. The bold values denote statistically significant difference at p < 0.05 level
Association of DI-GM with MetS and its components by logistic and modified poisson regression
Table 2 provides the outcomes of weighted multivariate logistic regression. After controlling for a variety of confounders in Model 3, the DI-GM score continued to exhibit a noticeable negative association with the odds of MetS (OR: 0.95, 95% CI: 0.93–0.97, P-value < 0.001), along with elevated waist circumference (OR: 0.91, 95% CI: 0.88–0.94, P-value < 0.001), elevated blood pressure (OR: 0.95, 95% CI: 0.93–0.98, P-value < 0.001), reduced high-density lipoprotein (OR: 0.95, 95% CI: 0.93–0.98, P-value < 0.001) and elevated fasting blood glucose (OR: 0.94, 95% CI: 0.90–0.98, P-value = 0.002). Subsequently, the DI-GM was categorized into quartiles for a more detailed exploration (Table 2). In Model 3, DI-GM was correlated with a lower odds for the following subclasses: the MetS in the 3rd, and the 4th (OR: 0.76, 95% CI: 0.67–0.87, P-value < 0.001) quartile; the elevated WC in the 2nd, 3rd, and the 4th (OR: 0.62, 95% CI: 0.52–0.73, P-value < 0.001) quartile; the elevated BP in the 2nd, the 3rd, and the 4th (OR: 0.83, 95% CI: 0.73–0.95, P-value = 0.006) quartile; the reduced HDL-C in the 3rd, and the 4th (OR: 0.76, 95% CI: 0.66–0.87, P-value < 0.001) quartile; the elevated TG in the 4th (OR: 0.87, 95% CI: 0.78–0.98, P-value = 0.026) quartile; the elevated FG in the 2nd, the 3rd, and the 4th (OR: 0.82, 95% CI: 0.69–0.98, P-value = 0.033) quartile, comparing with those in the reference quartiles.
Table 2.
Association of DI-GM with MetS and its components via weighted multivariate logistic regression
| Characteristic | OR1 (95% CI1), P-value | |||
|---|---|---|---|---|
| Model1 | Model2 | Model3 | ||
| MetS | ||||
| DI-GM-Continuous | 0.92 (0.90, 0.94), < 0.001 | 0.95 (0.92, 0.97), < 0.001 | 0.95 (0.93, 0.97), < 0.001 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.88 (0.80, 0.97), 0.009 | 0.89 (0.79, 1.00), 0.058 | 0.90 (0.80, 1.01), 0.077 | |
| Quartile 3(6) | 0.83 (0.75, 0.92), < 0.001 | 0.85 (0.76, 0.94), 0.004 | 0.86 (0.77, 0.96), 0.008 | |
| Quartile 4(7–11) | 0.68 (0.60, 0.76), < 0.001 | 0.74 (0.65, 0.84), < 0.001 | 0.76 (0.67, 0.87), < 0.001 | |
| Elevated WC | ||||
| DI-GM-Continuous | 0.93 (0.90, 0.95), < 0.001 | 0.92 (0.88, 0.95), < 0.001 | 0.91 (0.88, 0.94), < 0.001 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.92 (0.83, 1.02), 0.099 | 0.88 (0.75, 0.99), 0.038 | 0.86 (0.74, 0.99), 0.033 | |
| Quartile 3(6) | 0.86 (0.77, 0.95), 0.005 | 0.84 (0.71, 1.00), 0.044 | 0.84 (0.71, 0.99), 0.041 | |
| Quartile 4(7–11) | 0.69 (0.62, 0.78), < 0.001 | 0.63 (0.53, 0.75), < 0.001 | 0.62 (0.52, 0.73), < 0.001 | |
| Elevated BP | ||||
| DI-GM-Continuous | 0.96 (0.94, 0.98), < 0.001 | 0.95 (0.92, 0.97), < 0.001 | 0.95 (0.93, 0.98), < 0.001 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.87 (0.79, 0.96), 0.005 | 0.86 (0.76, 0.97), 0.018 | 0.87 (0.77, 0.99), 0.032 | |
| Quartile 3(6) | 0.88 (0.79, 0.97), 0.015 | 0.80 (0.70, 0.92), 0.001 | 0.81 (0.71, 0.93), 0.002 | |
| Quartile 4(7–11) | 0.89 (0.80, 0.98), 0.023 | 0.81 (0.71, 0.91), 0.001 | 0.83 (0.73, 0.95), 0.006 | |
| Reduced HDL-C | ||||
| DI-GM-Continuous | 0.90 (0.88, 0.93), < 0.001 | 0.94 (0.92, 0.97), < 0.001 | 0.95 (0.93, 0.98), < 0.001 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.87 (0.78, 0.97), 0.013 | 0.90 (0.80, 1.01), 0.082 | 0.91 (0.80, 1.02), 0.104 | |
| Quartile 3(6) | 0.78 (0.70, 0.87), < 0.001 | 0.86 (0.77, 0.97), 0.016 | 0.88 (0.78, 0.99), 0.032 | |
| Quartile 4(7–11) | 0.60 (0.52, 0.69), < 0.001 | 0.73 (0.64, 0.84), < 0.001 | 0.76 (0.66, 0.87), < 0.001 | |
| Elevated TG | ||||
| DI-GM-Continuous | 0.96 (0.94, 0.98), < 0.001 | 0.98 (0.96, 1.00) 0.045 | 0.98 (0.96, 1.00), 0.067 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.99 (0.91, 1.07), 0.772 | 1.02 (0.93, 1.11), 0.712 | 1.02 (0.93, 1.11), 0.700 | |
| Quartile 3(6) | 0.91 (0.82, 1.00), 0.055 | 0.92 (0.83, 1.02), 0.119 | 0.93 (0.84, 1.03), 0.149 | |
| Quartile 4(7–11) | 0.79 (0.71, 0.89), < 0.001 | 0.87 (0.77, 0.98), 0.019 | 0.87 (0.78, 0.98), 0.026 | |
| Elevated FG | ||||
| DI-GM-Continuous | 0.91 (0.89, 0.95), < 0.001 | 0.93 (0.90, 0.97), < 0.001 | 0.94 (0.90, 0.98), 0.002 | |
| Quartile 1(0–4) | — | — | — | |
| Quartile 2(5) | 0.80 (0.69, 0.93), 0.003 | 0.81 (0.70, 0.95), 0.009 | 0.82 (0.70, 0.95), 0.012 | |
| Quartile 3(6) | 0.78 (0.66, 0.93), 0.005 | 0.89 (0.64, 0.96), 0.017 | 0.80 (0.65, 0.99), 0.037 | |
| Quartile 4(7–11) | 0.72 (0.62, 0.84), < 0.001 | 0.78 (0.65, 0.93), 0.005 | 0.82 (0.69, 0.98), 0.033 | |
1OR = Odds Ratio, CI = Confidence Interval. Model 1 was adjusted for no variables. Model 2 was adjusted for age, sex, race, education, family income-to-poverty ratio, and BMI. Model 3 was further adjusted for smoking status, alcohol consumption, energy intake, chronic kidney disease, and cardiovascular disease based on Model 2. Abbreviations: MetS, Metabolic syndrome; WC, Waist Circumference (cm); BP, blood pressure; HDL-C, high-density lipoprotein; TG, total triglycerides; FG, fasting blood glucose. The bold values denote statistically significant difference at P-value < 0.05 level
Modified Poisson regression (Table S2) was employed to directly estimate prevalence ratios (PR) for the association between exposures and outcomes. In the fully adjusted Model 3, the DI-GM score maintained a marked negative link to MetS prevalence (PR: 0.97, 95% CI: 0.96–0.99, P-value < 0.001), along with elevated waist circumference (PR: 0.98, 95% CI: 0.97–0.99, P-value < 0.001), elevated blood pressure (PR: 0.98, 95% CI: 0.97–0.99, P-value < 0.001), reduced high-density lipoprotein (PR: 0.97, 95% CI: 0.95–0.98, P-value < 0.001) and elevated fasting blood glucose (PR: 0.95, 95% CI: 0.93–0.98, P-value = 0.001). More specifically, DI-GM was correlated with a lower prevalence for the following subclasses: the MetS in the 3rd, and the 4th (PR: 0.86, 95% CI: 0.80–0.93, P-value < 0.001) quartile; the elevated WC in the 4th (PR: 0.90, 95% CI: 0.87–0.94, P-value < 0.001) quartile; the elevated BP in the 2nd, the 3rd, and the 4th (PR: 0.93, 95% CI: 0.88–0.98, P-value = 0.006) quartile; the reduced HDL-C in the 3rd, and the 4th (PR: 0.83, 95% CI: 0.76–0.91, P-value < 0.001) quartile; the elevated TG in the 4th (PR: 0.92, 95% CI: 0.86–0.99, P-value = 0.022) quartile; the elevated FG in the 2nd, the 3rd, and the 4th (PR: 0.86, 95% CI: 0.75–0.98, P-value = 0.028) quartile, comparing with those in the reference quartiles.
The RCS analysis demonstrated that DI-GM was nonlinearly associated with MetS (Fig. 2A) (P for non-linearity = 0.030), along with elevated WC (Fig. 2B) (P for non-linearity = 0.004), reduced HDL-C (Fig. 2D) (P for non-linearity = 0.021) and elevated FG (Fig. 2F) (P for non-linearity = 0.012) when met the DI-GM score of 5. Other components included elevated BP (Fig. 2C) and elevated TG (Fig. 2E) showed a linear relationship.
Fig. 2.
The RCS analysis of DI-GM with MetS and its components
Association of DI-GM with six MetS related biochemical indicators by linear regression
Table 3 presents the relationship between DI-GM and six MetS related biochemical indicators. In Model 3, after adjusting for various confounders, DI-GM was significantly negatively correlated with the levels of WC (β: −0.35, 95% CI: −0.45–0.25, P-value < 0.001), systolic blood pressure (β: −0.20, 95% CI: −0.33–0.06, P-value = 0.004), diastolic blood pressure (β: −0.13, 95% CI: −0.23–0.02, P-value = 0.017) and fasting blood glucose (β: −0.27, 95% CI: −0.49–0.05, P-value = 0.017, while positively correlated with HDL-C (β: 0.19, 95% CI: 0.07–0.31, P-value = 0.002). These findings are generally consistent with our prior logistic regression results. Notably, this analysis may suggest that total triglyceride levels are unlikely to represent the primary biological mediator in the DI-GM–MetS association.
Table 3.
Association between DI-GM and various MetS related biochemical indicators via weighted multivariate linear regression
| Characteristic | β (95% CI1), P-value | |||
|---|---|---|---|---|
| Model1 | Model2 | Model3 | ||
| DI-GM | ||||
| Waist circumference | −0.87 (−1.00, −0.73), < 0.001 | −0.33 (−0.43, −0.24), < 0.001 | −0.35 (−0.45, −0.25), < 0.001 | |
| Systolic blood pressure | −0.17 (−0.33, −0.01), 0.035 | −0.23 (−0.36, −0.10), < 0.001 | −0.20 (−0.33, −0.06), 0.004 | |
| Diastolic blood pressure | −0.18 (−0.29, −0.07), 0.002 | −0.11 (−0.22, −0.01), 0.036 | −0.13 (−0.23, −0.02), 0.017 | |
| High-density lipoprotein | 0.92 (0.77, 1.1), < 0.001 | 0.25 (0.13, 0.37), < 0.001 | 0.19 (0.07, 0.31), 0.002 | |
| Total triglyceride | −2.1 (−3.0, −1.2), < 0.001 | −0.67 (−1.60, 0.23) 0.144 | −0.52 (−1.4, 0.38), 0.253 | |
| Fasting blood glucose | −0.63 (−0.84, −0.43), < 0.001 | −0.34 (−0.55, −0.13), 0.002 | −0.27 (−0.49, −0.05), 0.017 | |
1 CI = Confidence Interval. Model 1 was adjusted for no variables. Model 2 was adjusted for age, sex, race, education, family income-to-poverty ratio, and BMI. Model 3 was further adjusted for smoking status, alcohol consumption, energy intake, chronic kidney disease, and cardiovascular disease based on Model 2. The bold values denote statistically significant difference at P-value < 0.05 level
Sensitivity analysis
With the aim of reducing the influence of missing variables on the results, the weighted multivariate regression models were repeated using new dataset derived from multiple imputation as a sensitivity analysis. Specifically, consistent with the primary analysis, we maintained the exclusion criteria of participants aged < 20 years and pregnant women. After performing multiple imputation for missing variables (as detailed in Fig. 1), the new analytic sample comprised 38,135 participants. As shown in Table S3, the association between DI-GM and Mets with its components remained significant for continuous DI-GM except Elevated TG in Model 3. Despite minor changes in significance across certain quartiles, all 4th quartile remained significant. Table S4 demonstrates consistency with our prior findings, further substantiating the robustness of our main conclusions. Collectively, these sensitivity analyses validate the associations estimated in our primary analysis, demonstrating they are not excessively affected by missing data.
Subgroups analysis
A subgroup analysis was carried out to investigate potential links between DI-GM and MetS across various subgroups defined by age, sex, race, PIR, BMI, cigarette use, drinking status, CKD and CVD. The findings of the research, as depicted in Fig. S1, demonstrated that the association between DI-GM and MetS remained consistent across the majority of the subgroups examined. Significant interactive effects were detected within the PIR groups and the smoking status (P for interaction < 0.001).
Construction and validation of nine different ML models
The ROC curves of nine machine learning for both the training set (Fig. 3A) and the testing set (Fig. 3B) models indicated that XGboost achieved the highest AUC in both the training set (AUC = 0.91) and testing set (AUC = 0.83). Besides, after removing the main exposure DI-GM as a predictor, most models had lower AUC values, with Enet being the exception as depicted in Fig. S2. Simultaneously, the Xgboost remained attaining the highest AUC for the prediction of MetS in our research. Moreover, Fig. 4; Table 4 provided detailed discriminative characteristics for nine machine-learning algorithms. Importantly, XGboost performed exceptionally well across several important metrics, attaining the highest values of Accuracy (0.76), Sensitivity/Recall (0.82), AUC (0.83), and F1 score (0.82), along with a great Precision (0.82).
Fig. 3.
The ROC curves of nine machine-learning models. (A) Training set. (B) Testing set
Fig. 4.
Discriminative characteristics of nine machine-learning models
Table 4.
Detailed discriminative characteristics of ROC curve for nine machine-learning algorithms
| Model | Accuracy | Precision | Sensitivity/Recall | AUC | F1 score | Specificity |
|---|---|---|---|---|---|---|
| Decision tree | 0.72 | 0.82 | 0.75 | 0.78 | 0.78 | 0.68 |
| Enet | 0.70 | 0.84 | 0.66 | 0.78 | 0.74 | 0.76 |
| KNN | 0.67 | 0.78 | 0.71 | 0.72 | 0.74 | 0.61 |
| LightGBM | 0.71 | 0.83 | 0.69 | 0.78 | 0.76 | 0.73 |
| Logistic | 0.71 | 0.84 | 0.69 | 0.78 | 0.76 | 0.74 |
| MLP | 0.72 | 0.83 | 0.73 | 0.78 | 0.77 | 0.70 |
| Random forest | 0.72 | 0.80 | 0.77 | 0.77 | 0.79 | 0.62 |
| RSVM | 0.70 | 0.84 | 0.68 | 0.77 | 0.75 | 0.74 |
| XGBoost | 0.76 | 0.82 | 0.82 | 0.83 | 0.82 | 0.65 |
The calibration curve of XGBoost presented in Fig. S3 closely matched the reference line, highlighting its superior predictive accuracy. Furthermore, as depicted in Fig. S4 for the decision curve analysis, XGboost consistently remains above other models across a broad span of threshold probabilities, demonstrating a positive net benefit and corroborating its exceptional capacity in clinical decision-making.
Visualization of feature importance
Figure 5 depicted the distribution of SHAP values based on XGBoost, providing insight into the contributions of various features to the predictions. The color scheme, where purple represented high values and yellow represented low values, effectively conveyed the different degrees of feature influence. Notably, the variable “BMI ≥ 30” exhibited the most substantial influence on the predicted outcome, with higher SHAP values on the right side corresponding to a positive relationship with the outcome (MetS). Moreover, higher DI-GM exhibited considerable inverse influence, ranking after “Age”, “Race-Non-Hispanic Black” and “CKD-Yes”.
Fig. 5.
SHAP Summary plot based on the XGboost
Discussion
To the best of our knowledge, our study stands as the first attempt to systematically evaluate the association of DI-GM with MetS and its components using NHANES 2007–2020 data. Moreover, this study also constitutes the first effort to develop nine machine learning models as a promising tool for feature selection related to DI-GM in medical prediction of MetS, utilizing the explainable SHAP methodology.
We presented novel and multiple evidence from diverse perspectives. Weighted logistic and modified Poisson regression revealed that higher DI-GM related to lower prevalence of MetS includes abdominal obesity, high blood pressure, dyslipidemia and glucose metabolism dysfunction. Weighted linear regression examined significant association of DI-GM with five MetS related biochemical indicators including WC, systolic blood pressure, diastolic blood pressure and fasting blood glucose and HDL-C. This analysis may suggest that total triglyceride levels are unlikely to represent the primary biological mediator in the DI-GM–MetS association. Sensitivity analyses after multiple imputation validate the associations estimated in our primary analysis, demonstrating they are not excessively affected by missing data. RCS exhibited a significant inverse association of DI-GM with MetS for non-linear relationship when met the score of 5. Subgroup analysis demonstrated that this inverse association remained stable and consistent across the majority of the subgroups examined. ML employed sophisticated mathematical algorithms to detect and classify patterns in complex, heterogeneous datasets, facilitating us to enhance decision-making processes [23, 24]. SHAP methodology incorporated the individual effects of various features as well as the effects arising from their interactions [26]. Based on SHAP summary plots derived from XGboost, we observed that higher DI-GM exhibited considerable inverse influence, ranking after “BMI ≥ 30”, “Age”, “Race-Non-Hispanic Black” and “CKD-Yes”. This work was also aimed at providing secondary validation complementary to our prior logistic regression analysis, further ensuring the robust of the main conclusion.
Gut microbiota disorders have gradually been well-established as a critical factor for MetS. For example, evidence from a large-scale gut microbiome epidemiological survey conducted in an Eastern-nation revealed that Bacteroidetes and Ruminococcaceae were negatively associated with MetS [12]. Another multi-ethnic cohort study involving 3926 participants from Netherlands diagnosed with MetS reported a higher abundance of Enterobacteriaceae and lower of peptostreptococcaceae [36]. A study exploring the microbiota patterns in Romanian discovered that patients with MetS possessed a microbiome enrichment of Enterobacteriaceae,and Clostridium leptum, while beneficial taxa like Butyricicoccus sp. exhibited a decrease [37].
The DI-GM score primarily reflects the balance of gut microbiome and the generation of beneficial metabolites such as SCFAs. Several mechanisms may account for the link between the DI-GM and MetS as well as its individual components. Evidences have suggested that the diversity of gut microbiota contributed to alleviating insulin resistance and ameliorating glucose metabolism disorders [38]. Furthermore, SCFAs could improve pancreatic β-cell proliferation and prevent the transdifferentiation into α-cells, ultimately leading to an elevation release of insulin [39]. Studies also revealed that SCFAs involved in the regulation of fatty acid metabolism, stimulating leptin production in adipocytes and decreasing accumulation of visceral fat [40]. Qin et al. demonstrated that altered gut microbiota composition in metabolic syndrome patients correlates with enhanced inflammatory responses, mediated by suppressed SCFA, thereby promoting metabolic syndrome related diseases [41]. Another investigation found that the microbiome of obese individuals exhibited an enhanced capacity to derive energy from dietary intake, causing disrupted nutrient distribution and the rapid progression of adiposity [42, 43]. Moreover, probiotic strains generated peptides with ACE-inhibitory activity, thereby exerting a blood-pressure lowering effect [44]. Other potential explanations encompass the modulation of hepatic gluconeogenesis, the regulation of circadian host rhythms, and the influence on insulin signaling by the microbiome [44].
Our study possesses several strengths and clinical significance. The findings suggest that maintaining a high DI-GM score helps reduce the prevalence of metabolic syndrome. This dietary index offers individuals a multifaceted and dependable evaluation of their dietary quality in relation to ecological balance of the gut microbiome, facilitating suitable dietary interventions for enhanced health. Furthermore, we conducted advanced machine learning techniques, along with a series of regression analysis to validate our findings from diverse perspectives in a large-scale database. Appropriate sampling weights were considered to reduce potential oversampling bias and enhance the generalizability of the conclusion to the U.S. population. Despite its strengths, our study has several limitations. Firstly, the cross-sectional dietary assessment, measured only once, cannot capture temporal changes, failing to accurately reflect long-term dietary patterns. Secondly, the unavailability of green tea consumption metrics in NHANES datasets might have introduced underestimation in DI-GM calculation, potentially compromising association estimates with metabolic outcomes. Thirdly, the findings from NHANES are primarily applicable to the U.S. population; therefore, the current results cannot be fully extrapolated to all racial groups worldwide. Despite our efforts to control for relevant confounders in the current model, the possibility of unaccounted covariates (e.g., antibiotic use) cannot be entirely excluded. Finally, we could not demonstrate the causal relationships. Future prospective cohorts with repeated and comprehensive measures, randomized controlled interventions, and replication studies in diverse geographic populations are required to confirm the long-term effects of dietary interventions on DI-GM and MetS, thereby enabling stronger causal inference and enhancing the generalizability of the findings.
Conclusion
Combining the outcomes, our study presents compelling evidence that higher scores of the DI-GM are associated with a lower prevalence of Mets and its components. Dietary strategies that incorporate the DI-GM score could contribute to the harmonious ecological state of the gut microbiome and be crucial in the prevention of MetS.
Supplementary Information
Acknowledgements
We thank the staff at the National Center for Health Statistics of the Centers for Disease Control for designing, collecting, and collating the NHANES data and creating the public database.
Author contributions
Y.C. contributed to conception, overall design, main data analysis, and paper writing. S.W. contributed to formal analysis and paper writing. Y.T. contributed to visualization and investigation. D.L., S.C. and H.L. contributed to supervision, funding acquisition, review and edit the manuscript.
Funding
This research was supported by the National Natural Science Foundation of China (grant number: 82274419), Guangdong Basic and Applied Basic Research Foundation (No: 2021A1515220177), and Sanming Project of Medicine in Shenzhen (No.SZZYSM202411016).
Data availability
This study used data from a free and open public database, which can be found here: www.cdc.gov/nchs/nhanes/.
Declarations
Ethics approval and consent to participate
Each participant provided a written informed agreement before inclusion in the NHANES database, which was examined and allowed by the National Center for Health Statistics Ethics Review Board.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yu Cai and Sheng-Jia Wang contributed equally to this work.
Contributor Information
Shu-Fang Chu, Email: chushufanggzhtcm@163.com.
Hui-Lin Li, Email: sztcmlhl@163.com.
References
- 1.Saklayen MG. The global epidemic of the metabolic syndrome. Curr Hypertens Rep. 2018. 10.1007/s11906-018-0812-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Kim YJ, Kim S, Seo JH, Cho SK. Prevalence and associations between metabolic syndrome components and hyperuricemia by race: findings from US population, 2011–2020. Arthritis Care Res (Hoboken). 2024;76:1195–202. 10.1002/acr.25338. [DOI] [PubMed] [Google Scholar]
- 3.Kang Q, Mei X, Guo C, Si Y, Wang N. Association between mediterranean diet and metabolic syndrome: analysis of NHANES 2007–2020. Int J Food Sci Nutr. 2025;76:209–22. 10.1080/09637486.2025.2450452. [DOI] [PubMed] [Google Scholar]
- 4.Huh JH, Yadav D, Kim JS, Son JW, Choi E, Kim SH, et al. An association of metabolic syndrome and chronic kidney disease from a 10-year prospective cohort study. Metabolism. 2017;67:54–61. 10.1016/j.metabol.2016.11.003. [DOI] [PubMed] [Google Scholar]
- 5.Silveira Rossi JL, Barbalho SM, de Araujo R, Bechara R, Sloan MD, K. P., Sloan LA. Metabolic syndrome and cardiovascular diseases: going beyond traditional risk factors. Diabetes Metab Res Rev. 2021;38. 10.1002/dmrr.3502. [DOI] [PubMed]
- 6.Neeland IJ, Lim S, Tchernof A, Gastaldelli A, Rangaswami J, Ndumele CE, et al. Metabolic syndrome. Nat Rev Dis Primers. 2024. 10.1038/s41572-024-00563-5. [DOI] [PubMed] [Google Scholar]
- 7.Grundy SM, Stone NJ, Bailey AL, Beam C, Birtcher KK, Blumenthal RS, AHA/ACC/AACVPR/AAPA/ABC/ACPM/ADA/AGS/APhA et al. /ASPC/NLA/PCNA Guideline on the Management of Blood Cholesterol: A Report of the American College of Cardiology/American Heart Association Task Force on Clinical Practice Guidelines. Circulation (2019) 139.10.1161/cir.0000000000000625 [DOI] [PMC free article] [PubMed]
- 8.Fan Y, Pedersen O. Gut microbiota in human metabolic health and disease. Nat Rev Microbiol. 2020;19:55–71. 10.1038/s41579-020-0433-9. [DOI] [PubMed] [Google Scholar]
- 9.Huang X, Hu L, Li J, Xie X, Meng C, Liu Y, et al. Dietary live microorganisms and depression-driven mortality in hypertensive patients: NHANES 2005–2018. J Health Popul Nutr. 2025;44117. 10.1186/s41043-025-00861-y. [DOI] [PMC free article] [PubMed]
- 10.Adak A, Khan MR. An insight into gut microbiota and its functionalities. Cell Mol Life Sci. 2019;76:473–93. 10.1007/s00018-018-2943-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Du YZ, Hu HJ, Dong QX, Guo B, Zhou Q, Guo J. The relationship between dietary live microbe intake and overactive bladder among American adults: a cross-sectional study from NHANES 2007–2018. J Health Popul Nutr. 2024;43120. 10.1186/s41043-024-00612-5. [DOI] [PMC free article] [PubMed]
- 12.He Y, Wu W, Wu S, Zheng H-M, Li P, Sheng H-F, et al. Linking gut microbiota, metabolic syndrome and economic status based on a population-level analysis. Microbiome. 2018. 10.1186/s40168-018-0557-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Thomas MS, Blesso CN, Calle MC, Chun OK, Puglisi M, Fernandez ML. Dietary influences on gut microbiota with a focus on metabolic syndrome. Metab Syndr Relat Disord. 2022;20:429–39. 10.1089/met.2021.0131. [DOI] [PubMed] [Google Scholar]
- 14.Kase BE, Liese AD, Zhang J, Murphy EA, Zhao L, Steck SE. The development and evaluation of a literature-based dietary index for gut microbiota. Nutrients. 2024. 10.3390/nu16071045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Wu Z, Gong C, Wang B. The relationship between dietary index for gut microbiota and diabetes. Sci Rep. 2025. 10.1038/s41598-025-90854-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Liu J, Huang S. Dietary index for gut microbiota is associated with stroke among US adults. Food Funct. 2025;16:1458–68. 10.1039/d4fo04649h. [DOI] [PubMed] [Google Scholar]
- 17.Xiao D, Sun X, Li W, Wen Z, Zhang WH, Yang L. Associations of dietary index for gut microbiota and flavonoid intake with female infertility in the united States. Food Sci Nutr. 2025;13:e70098. 10.1002/fsn3.70098. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Yan X, Shao X, Zeng T, Zhang Q, Deng J, Xie J. National analysis of the dietary index for gut microbiota and kidney stones: evidence from NHANES (2007–2018). Front Nutr. 2025;121540688. 10.3389/fnut.2025.1540688. [DOI] [PMC free article] [PubMed]
- 19.Li X, Zhao Y, Zhang D, Kuang L, Huang H, Chen W, et al. Development of an interpretable machine learning model associated with heavy metals’ exposure to identify coronary heart disease among US adults via SHAP: findings of the US NHANES from 2003 to 2018. Chemosphere. 2023. 10.1016/j.chemosphere.2022.137039. [DOI] [PubMed] [Google Scholar]
- 20.Liu J, Li X, Zhu P. Effects of various heavy metal exposures on insulin resistance in non-diabetic populations: interpretability analysis from machine learning modeling perspective. Biol Trace Elem Res. 2024;202:5438–52. 10.1007/s12011-024-04126-3. [DOI] [PubMed] [Google Scholar]
- 21.Dinh A, Miertschin S, Young A, Mohanty SD. A data-driven approach to predicting diabetes and cardiovascular disease with machine learning. BMC Med Inform Decis Mak. 2019. 10.1186/s12911-019-0918-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Rahnenführer J, De Bin R, Benner A, Ambrogi F, Lusa L, Boulesteix A-L, et al. Statistical analysis of high-dimensional biomedical data: a gentle introduction to analytical goals, common approaches and challenges. BMC Med. 2023. 10.1186/s12916-023-02858-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Akyea RK, Qureshi N, Kai J, Weng SF. Performance and clinical utility of supervised machine-learning approaches in detecting familial hypercholesterolaemia in primary care. Npj Digit Med. 2020. 10.1038/s41746-020-00349-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Stafford IS, Kellermann M, Mossotto E, Beattie RM, MacArthur BD, Ennis S. A systematic review of the applications of artificial intelligence and machine learning in autoimmune diseases. Npj Digit Med. 2020;3:30. 10.1038/s41746-020-0229-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Alber M, Buganza Tepole A, Cannon WR, De S, Dura-Bernal S, Garikipati K, et al. Integrating machine learning and multiscale modeling—perspectives, challenges, and opportunities in the biological, biomedical, and behavioral sciences. NPJ Digit Med. 2019. 10.1038/s41746-019-0193-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Ejiyi CJ, Qin Z, Ukwuoma CC, Nneji GU, Monday HN, Ejiyi MB, et al. Comparative performance analysis of Boruta, SHAP, and Borutashap for disease diagnosis: A study with multiple machine learning algorithms. NCNS. 2024;1–38. 10.1080/0954898x.2024.2331506. [DOI] [PubMed]
- 27.Grundy SM, Cleeman JI, Daniels SR, Donato KA, Eckel RH, Franklin BA, et al. Diagnosis and management of the metabolic syndrome. Circulation. 2005;112:2735–52. 10.1161/circulationaha.105.169404. [DOI] [PubMed] [Google Scholar]
- 28.Zhang Q, Wu Y, Luo B. Association of oxidative balance score with metabolic syndrome and its components in middle-aged and older individuals in the United States. Front Nutr. 2025. 10.3389/fnut.2025.1523791. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Johnson CL, Paulose-Ram R, Ogden CL, Carroll MD, Kruszon-Moran D, Dohrmann SM et al. National health and nutrition examination survey: analytic guidelines, 1999–2010. Vital and health statistics. Series 2, Data evaluation and methods research (2013): 1–24. [PubMed]
- 30.Guo T, Zheng S, Chen T, Chu C, Ren J, Sun Y, et al. The association of long-term trajectories of BMI, its variability, and metabolic syndrome: a 30-year prospective cohort study. EClinicalMedicine. 2024;69102486. 10.1016/j.eclinm.2024.102486. [DOI] [PMC free article] [PubMed]
- 31.Wang J, Bai Y, Zeng Z, Wang J, Wang P, Zhao Y, et al. Association between life-course cigarette smoking and metabolic syndrome: a discovery-replication strategy. Diabetol Metab Syndr. 2022;1411. 10.1186/s13098-022-00784-2. [DOI] [PMC free article] [PubMed]
- 32.Choi S, Kim K, Lee JK, Choi JY, Shin A, Park SK, et al. Association between change in alcohol consumption and metabolic syndrome: analysis from the health examinees study. Diabetes Metab J. 2019;43:615–26. 10.4093/dmj.2018.0128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Angelico F, Baratta F, Coronati M, Ferro D, Del Ben M. Diet and metabolic syndrome: a narrative review. Intern Emerg Med. 2023;18:1007–17. 10.1007/s11739-023-03226-7. [DOI] [PubMed] [Google Scholar]
- 34.Guo K, Ni W, Du L, Zhou Y, Cheng L, Zhou H. Environmental chemical exposures and a machine learning-based model for predicting hypertension in NHANES 2003–2016. BMC Cardiovasc Disord. 2024. 10.1186/s12872-024-04216-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Chevret S, Seaman S, Resche-Rigon M. Multiple imputation: a mature approach to dealing with missing data. Intensive Care Med. 2015;41:348–50. 10.1007/s00134-014-3624-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Deschasaux M, Bouter K, Prodan A, Levin E, Groen A, Herrema H, et al. Differences in gut microbiota composition in metabolic syndrome and type 2 diabetes subjects in a multi-ethnic population: the HELIUS study. Proc Nutr Soc. 2020. 10.1017/s0029665120001317. [Google Scholar]
- 37.Gradisteanu Pircalabioru G, Ilie I, Oprea L, Picu A, Petcu LM, Burlibasa L, et al. Microbiome, mycobiome and related metabolites alterations in patients with metabolic syndrome—a pilot study. Metabolites. 2022. 10.3390/metabo12030218. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Ranaivo H, Thirion F, Béra-Maillet C, Guilly S, Simon C, Sothier M, et al. Increasing the diversity of dietary fibers in a daily-consumed bread modifies gut microbiota and metabolic profile in subjects at cardiometabolic risk. Gut Microbes. 2022. 10.1080/19490976.2022.2044722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Teyani R, Moniri NH. Gut feelings in the islets: the role of the gut microbiome and the FFA2 and FFA3 receptors for short chain fatty acids on β-cell function and metabolic regulation. Br J Pharmacol. 2023;180:3113–29. 10.1111/bph.16225. [DOI] [PubMed] [Google Scholar]
- 40.Dong G, Zhang J, Yang Z, Feng X, Li J, Li D, et al. The association of gut microbiota with idiopathic central precocious puberty in girls. Front Endocrinol (Lausanne). 2020. 10.3389/fendo.2019.00941. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Qin Q, Yan S, Yang Y, Chen J, Li T, Gao X, et al. A metagenome-wide association study of the gut microbiome and metabolic syndrome. Front Microbiol. 2021. 10.3389/fmicb.2021.682721. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Turnbaugh PJ, Ley RE, Mahowald MA, Magrini V, Mardis ER, Gordon JI. An obesity-associated gut microbiome with increased capacity for energy harvest. Nature. 2006;444:1027–31. 10.1038/nature05414. [DOI] [PubMed] [Google Scholar]
- 43.Green M, Arora K, Prakash S. Microbial medicine: prebiotic and probiotic functional foods to target obesity and metabolic syndrome. Int J Mol Sci. 2020. 10.3390/ijms21082890. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Gonzalez-Gonzalez C, Gibson T, Jauregi P. Novel probiotic-fermented milk with angiotensin I-converting enzyme inhibitory peptides produced by bifidobacterium bifidum MF 20/5. Int J Food Microbiol. 2013;167:131–7. 10.1016/j.ijfoodmicro.2013.09.002. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
This study used data from a free and open public database, which can be found here: www.cdc.gov/nchs/nhanes/.





