Skip to main content
Science Progress logoLink to Science Progress
. 2026 Jul 3;109(3):00368504261460667. doi: 10.1177/00368504261460667

Stroke risk associated with the interaction between composite dietary antioxidant index and heavy metals: A cross-sectional explainable machine learning study using NHANES data

Yixuan He 1,2,3,4, Kai Gong 1,4,✉, Quan Lan 1,2,5,✉
PMCID: PMC13332272  PMID: 42394597

Abstract

Purpose

Stroke remains the second leading global cause of death. Traditional risk factors and statistical models fail to fully clarify its pathogenesis or quantify the interactive effects of dietary antioxidants and heavy metals.

Methods

Based on the data of 35,171 subjects from the National Health and Nutrition Examination Survey (NHANES) conducted in the United States between 2005 and 2018, this Cross-Sectional study constructed a Composite Dietary Antioxidant Index (CDAI) calibrated for the study population. By incorporating blood lead, blood cadmium, manganese elements, antioxidant nutrients and traditional stroke risk factors, five tree-based machine learning models were trained and validated. Meanwhile, with the application of the SHapley Additive exPlanations (SHAP), Restricted Cubic Spline (RCS) and piecewise regression analysis, this research analyzed the mechanism of the models, verified the non-linear correlations, and defined the protective threshold of the Composite Dietary Antioxidant Index as well as the recommendations for dietary protection.

Results

The HistGradientBoosting model achieved the best performance (AUC=0.7771, 95%CI: 0.765–0.791). CDAI showed a non-linear protective effect against stroke, with an optimal threshold of 1.67 (95%CI: 1.45–1.89), above which stroke risk decreased by 40.2% (OR=0.60, 95%CI: 0.48–0.75, P<0.001). High blood lead (≥2.2 μg/dL) and cadmium (≥0.5 μg/L) significantly attenuated CDAI’s protective effect by 146%–167%, while high CDAI may account for mitigated heavy metal-related stroke risk. The model showed stable performance in gender subgroups and adults aged <65 years.

Conclusion

This cross-sectional study constructed an interpretable machine learning framework for stroke risk using NHANES data. The HistGradientBoosting model performed best. We identified a nonlinear protective threshold of CDAI at 1.67, above which stroke risk decreased by 40.2%. High blood lead and cadmium significantly attenuated the protective effect of CDAI, while manganese showed synergistic antioxidant protection. SHAP and RCS analyses confirmed robust interactions between CDAI and heavy metals. These findings provide evidence for personalized dietary antioxidant interventions in stroke prevention, especially for individuals with heavy metal exposure.

Keywords: stroke prediction, machine learning, explainable artificial intelligence, nutritional epidemiology environmental exposure, NHANES

1. Introduction

Stroke is the second leading cause of death worldwide and one of the main causes of disability. 1 With population aging, its incidence is on the rise, and varies by country and ethnicity. Most traditional risk factors for stroke, such as hypertension, diabetes and obesity, are controllable and preventable. 2 Previous studies have explored the association between dietary nutrition and stroke, 3 or the impact of environmental toxins on stroke. 4 However, studies on the synergy between the Composite Dietary Antioxidant Index (CDAI) and heavy metal exposure remain insufficient, lacking systematic quantitative analysis and most studies focused on linear associations, with limited exploration of its non-linear protective effect and population-specific protective threshold. In addition, traditional statistical models require the pre-setting of linear correlations between variables, cannot independently capture nonlinear relationships and high-dimensional feature interactions, and thus perform poorly. 5 Machine learning models based on decision trees can automatically detect and utilize complex interactions and nonlinear relationships, 6 yet their inherent “black box” nature restricts their clinical translational applications. Therefore, it is of practical significance to study the interaction between CDAI and heavy metals using interpretable machine learning methods.

2. Materials and methods

2.1. Study samples

NHANES(website: https://www.cdc.gov/nchs/nhanes/) 2005–2018 cycles (n=79,683) were selected to maintain consistency across standard pre-pandemic two-year survey cycles and to ensure comparable availability of dietary antioxidant intake, blood heavy metal measurements, stroke outcome, and covariate data. More recent NHANES data were not incorporated into the primary analysis because the 2019–2020 cycle was interrupted by the COVID-19 pandemic and released as a special 2017–March 2020 pre-pandemic file requiring distinct analytic considerations, while the 2021–2023 cycle used an updated sample design and may introduce additional methodological and post-pandemic heterogeneity. The data collection protocol of the U.S. National Health and Nutrition Examination Survey was approved by the institutional ethics review committee of the survey, and all participants provided written informed consent (information had been de-identified). This study conducted in accordance with the relevant provisions of the Declaration of Helsinki (adopted in 1975, revised in 2024) and was reported in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines. 7

2.2. Covariate selection

The collected covariates include gender (male, female); age; race (Mexican American, other Hispanic, non-Hispanic White, non-Hispanic Black, and other races); diabetes; smoking status; BMI data; education level (below high school, high school, college and above); family income poverty ratio. Smoking status was treated as a toxicologically relevant covariate and was included in all primary machine-learning models, restricted cubic spline models, piecewise regression models, and SHAP interaction analyses.

2.3. Definition of synergistic variables

The Composite Dietary Antioxidant Index (CDAI) is a classic evaluation index used to comprehensively quantify the total antioxidant potential of an individual’s diet (including vitamins A, C, E, carotenoids, and minerals such as zinc and selenium).8,9 Its core advantage lies in integrating the synergistic effects of multiple key antioxidant nutrients, breaking through the limitations of assessment using a single index. 10 Specifically, the CDAI was calculated using nutrient-specific Z-scores standardized within the analytic NHANES population (Only based on internal US data from NHANES, with no data from any other regions included.)

Zi=Xi−μiσi

where Xi represents the individual intake of nutrient i , μi and σi denote the mean and standard deviation of this samples’ population intake, respectively.

This sample ultimately included Major antioxidant nutrients vitamin A, vitamin C, vitamin E, β-carotene, selenium, and zinc, in the study. The CDAI was then calculated as the sum of the standardized values of these nutrients:

CDAI=∑i=16Zi

To quantify the potential non-linear protective effect of the Composite Dietary Antioxidant Index (CDAI) on the risk of stroke onset, this study adopted two methods to calculate the protective threshold:

After adjusting for covariates, a restricted cubic spline regression model was used to fit the non-linear relationship between the Composite Dietary Antioxidant Index and the risk of stroke onset.

Taking the risk of stroke onset as a binary outcome variable (logit(P) represents the logistic transformation of the probability of stroke occurrence), its calculation formula is as follows:

logit(P(stroke=1))=β0+∑i=1k−2βi·fi(CDAI)

f(CDAI): Restricted cubic spline transformation function of CDAI where:

f(CDAI)=∑j=13βj·RCSj(CDAI)

In addition, the protective threshold is calculated by means of piecewise regression.,By traversing different thresholds, segmented regression analysis selects the threshold with the minimum Akaike Information Criterion (AIC) as the optimal threshold.

Additionally, this study combines SHAP interaction plots to intuitively interpret the potential interactive association between heavy metal exposure and CDAI.

Therefore, all CDAI-related thresholds, SHAP analyses, restricted cubic spline analyses, piecewise regression results, interaction analyses, and subgroup analyses were based on this NHANES population-specific Z-score CDAI.

2.4. Data preprocessing

To ensure data quality and research reliability, samples were screened based on the following inclusion and exclusion criteria: 1. Excluding samples with missing key research indicators (including demographic characteristics, stroke-related risk factors, and blood heavy metal detection indicators); 2. Excluding samples with incomplete survey year information or logical contradictions; After the above screening, an integrated dataset including demographic questionnaire data, stroke-related risk factors, and blood heavy metal content indicators was finally included.

To clarify the baseline characteristics of the study population and the differences between groups, this study constructed a baseline characteristic table containing statistical test P-values (Table 1), which systematically presents the distribution characteristics of demographic characteristics, clinical risk factors, and blood heavy metal levels between the stroke group and the non-stroke group.

Table 1.

Feature baseline table.

Variable No stroke (n=33795) Stroke (n=1376) P-value
Gender - - 0.46
 Male 16368 (48.4%) 681 (49.5%) -
 Female 17427 (51.6%) 695 (50.5%) -
Age (years) 48.00 (34.00-63.00) 68.00 (58.00-78.00) <0.001
Race/Ethnicity 33795 1376 <0.001
Education Level 33762 1375 <0.001
Annual Family Income 4204 163 N/A
Smoking Status 33778 1376 <0.001
Smoking Frequency 14879 838 <0.001
Body Mass Index (kg/m2) 28.07 (24.34-32.70) 29.20 (25.20-33.40) <0.001
Weight (kg) 78.70 (66.60-93.10) 79.30 (67.80-92.53) 0.43
Height (cm) 166.90 (159.70-174.40) 165.30 (158.50-172.40) <0.001
Vitamin A Intake (µg) 464.00 (253.00-770.00) 434.50 (232.00-733.75) 0.008
Vitamin C Intake (mg) 53.20 (21.40-116.10) 42.65 (17.50-104.67) <0.001
Vitamin E Intake (mg) 6.58 (4.18-10.08) 5.55 (3.38-8.54) <0.001
Zinc Intake (mg) 9.68 (6.61-13.99) 8.21 (5.39-11.92) <0.001
Selenium Intake (µg) 100.70 (69.40-140.90) 81.45 (56.27-118.22) <0.001
Blood Cadmium (µg/L) 0.32 (0.20-0.58) 0.46 (0.29-0.79) <0.001
Blood Lead (µg/dL) 1.19 (0.74-1.89) 1.64 (1.10-2.60) <0.001
Blood Manganese (µg/L) 9.27 (7.43-11.67) 8.54 (6.78-10.81) <0.001
Log-transformed Blood Cadmium 0.37 ± 0.27 0.46 ± 0.28 <0.001
Log-transformed Blood Lead 0.85 ± 0.40 1.04 ± 0.43 <0.001
Log-transformed Blood Manganese 2.34 ± 0.32 2.27 ± 0.34 <0.001
CDAI -0.03 ± 3.44 -0.93 ± 3.12 <0.001

Baseline analysis revealed a significant class imbalance in the dataset grouped by stroke diagnosis status, attributed to the stroke cases being the minority class. This issue may lead to prediction bias of machine learning models towards the majority class, thereby reducing the ability to identify stroke cases. To avoid such problems, we adopt an imbalanced data processing method for the model training dataset: Use the inverse class frequency weighting method to assign a weight to each sample that is inversely proportional to its class frequency. For class c∈{0,1} , the calculation formula for the sample weight  wc is as follows:

wc=Ntotal2·NC

where the uppercase Ntotal represents the total number of samples in the training set, and Nc represents the number of samples of class (c) . For tree models, class_weight='balanced’ is adopted, while for XGBoost and gradient boosting machines, scale_pos_weight (pos_weight) is used to balance positive and negative classes, which is calculated as the ratio of the number of negative samples to the number of positive samples.

Based on the above analysis, this study adopts a data preprocessing strategy of “independent balancing of the training set and maintaining the original state of the test set”, with specific steps as follows: ① Select the variable MCQ160F from the National Health and Nutrition Examination Survey (NHANES) database as the diagnostic indicator for stroke outcomes, and accordingly divide the samples into a case group (stroke patients) and a control group (non-stroke population); ② Dataset splitting: Split the original dataset into a training set and a test set at a ratio of 7:3, ensuring consistent distributions of stroke prevalence and key baseline characteristics between the two sets; ③ Sample balancing: Operations are only performed on the training set, and this sampling method is only used in the 5-fold cross-validation training set; the test set retains the original data distribution to avoid the risk of data leakage.

2.5. Machine learning models and SHAP interpretability analysis

This study selects five machine learning models for training and evaluates their performance separately: Categorical Boosting (Catboost), Random Forest (RF), Gradient Boosting Machine (GBM), Histogram-Based Gradient Boosting, and Extreme Gradient Boosting (XGBoost). As an ensemble method, Random Forest (RF) improves generalization ability through the voting mechanism of multiple decision trees, is robust to outliers, and can handle nonlinear relationships, but has poor interpretability. The Gradient Boosting Machine (GBM) and its optimized version XGBoost combine weak learners in an iterative manner and perform excellently on structured data. Among them, XGBoost effectively prevents overfitting through parallel computing and regularization. Catboost is specifically optimized for categorical features and can directly process categorical variables without extensive preprocessing. The Histogram-Based Gradient Boosting algorithm accelerates the search for feature splitting points using the histogram method, and adopts Gradient-Based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB) to significantly improve training speed and memory efficiency.

In addition, SHAP (Shapley Additive explanations) is a model interpretation method based on the Shapley value in game theory, which can provide consistent and interpretable feature importance metrics for the predictions of any machine learning model. 11 The specific formula for calculating the SHAP value is as follows:

ϕi=∑S⊆F∖{i}|S|!(|F|−|S|−1)!|F|![f(S∪{i})−f(S)]

For tree-based models, TreeSHAP provides an efficient method for calculating SHAP values:

ϕi=12d−1∑l∈Li(v(lright)−v(lleft))

where v(lright)−v(lleft) represents the difference between the left and right leaves after feature I splits, and 12d−1 means weighted by depth.

The SHAP value satisfies the additive principle, and the model predicted value can be expressed as the sum of the base value and the SHAP value of each feature:

f(x)=ϕ0+∑i=1Mϕi

Finally, the formula for quantifying the contribution of the interaction between two features to model prediction through SHAP interaction values is as follows:

ϕi,j=E[f(x)∣xi,xj]−E[f(x)∣xi]−E[f(x)∣xj]+E[f(x)]

E is represent event expectation.

2.6. Feature extraction and validation

To evaluate the discriminative accuracy of machine learning models, we use confusion matrices to calculate the false negatives (FN), false positives (FP), true negatives (TN), and true positives (TP) of each model, and further compute accuracy TP+TNTP+TN+FP+FN , precision TPTP+FP , recall TPTP+FN , and specificity TNTN+FP accordingly. Meanwhile, we plot receiver operating characteristic curves (ROC curves) and calculate the area under the curve (AUC). The aforementioned indicators measure the proportion of correct predictions made by the model, but they can be misleading in the context of imbalanced class data. For this reason, this study also introduces the F1-score 2×Precision×RecallPrecision+Recall , Brier Score: 1N∑i=1N(pi−yi)2 for a multi-dimensional analysis. 12 In addition, based on the Shapley value method in game theory, various features in the HistGradientBoosting model are used to draw summary plots, interaction plots, forest plot and dependence plots, which intuitively demonstrate the contribution of features to model predictions and thereby improve the interpretability of the model.

3. Results

3.1. Predictive analysis of various machine learning models

Given that tree models can independently explore the nonlinear interactive relationships among features, manually adding linear variables may lead to redundancy and introduce noise. This study trains on the original samples and compares the model performance (Table 2).

Table 2.

Performance table of models.

Model AUC F1 score Brier score Precision Recall PR AUC
CatBoost 0.7155 0.1355 0.0724 0.0782 0.5085 0.0831
HistGradientBoosting 0.7771 0.1789 0.0663 0.1075 0.5327 0.1076
GBM 0.7142 0.1321 0.0473 0.0876 0.2688 0.0829
XGBoost 0.7397 0.1523 0.0661 0.0896 0.5085 0.0936
RandomForest 0.776 0.1833 0.0984 0.1177 0.414 0.1112

To further compare the overall diagnostic performance of the five machine learning models, we used three of them to plot ROC curves based on the predicted probability values of each sample and the true sample values, with the results shown in Figure 1. It can be seen that the HistGradientBoosting model achieved the optimal AUC value of 0.7771 (Figure 1), with its 95% confidence interval ranging from 0.765 to 0.791. Nonetheless, the F1 scores and PR-AUC values of all models are relatively low. The F1-score (0.1789) and PR-AUC (0.1076) observed in this study are consistent with previous population-based stroke prediction studies using NHANES data, which reported F1-scores ranging from 0.16 to 0.23 due to the low prevalence of stroke in the general population.

Figure 1.

Figure 1.

ROC curves of each model Each curve represents one model and shows the trade-off between sensitivity and 1-specificity across different classification thresholds. The area under the curve (AUC) indicates the overall discriminative ability of each model, with a higher AUC representing better performance.

3.2. Single-factor contribution analysis

To further observe how the model obtains results, we used the HistGradientBoosting, which performed the best in the above analysis, to draw the SHAP Summary plot, as shown in Figure 2.

Figure 2.

Figure 2.

SHAP Summary plot Each point represents one participant, and the x-axis shows the SHAP value, indicating the direction and magnitude of each feature’s contribution to the predicted stroke risk. Features are ranked according to their mean absolute SHAP values, with higher-ranked features contributing more strongly to model prediction.

The study found that the interaction term elevated the ranking and the differential contributions of various characteristics to individuals. To more intuitively present the specific contribution characteristics of each risk factor, we further plotted the single-feature SHAP plots of the composite dietary antioxidant index and various heavy metals (lead, cadmium, manganese), and the specific results are shown in Figure 3.

Figure 3.

Figure 3.

CDAI, Lead-Cadmium-Manganese SHAP Dependence Plot Each point represents one participant, with the x-axis showing the observed value of CDAI or the corresponding blood heavy metal concentration and the y-axis showing the SHAP value for that feature. Positive SHAP values indicate an increased predicted stroke risk, whereas negative SHAP values indicate a reduced predicted risk.

The SHAP dependence plot visually illustrates the marginal effect of the concerned features on stroke risk prediction. CDAI exhibits a non-linear protective effect, and the protective effect tends to become saturated when CDAI exceeds a specific value (approximately 2). Stroke cases are mainly concentrated in the low CDAI range, suggesting that CDAI is a protective factor against stroke. Blood lead (Pb) and blood cadmium (Cd) show a monotonically positive risk effect: as the levels of Pb/Cd rise, the SHAP values shift from negative to positive, and stroke cases are more concentrated in the high-exposure groups, indicating that both are risk factors. To further explore how heavy metal exposure modifies the protective association of CDAI with stroke, restricted cubic splines, piecewise regression models and SHAP interaction values were subsequently used to analyze the non-linear interaction effects.

3.3. Analysis of the protective threshold of CDAI

The goodness of fit of the piecewise regression model is superior to that of the RCS model (AIC: 11504.97 vs 11513.34). Based on piecewise regression analysis, the protective threshold of CDAI was 1.67 (95% CI: 1.45-1.89). When CDAI > 1.67, the risk of stroke was significantly reduced by 40.2% (OR = 0.60, 95% CI: 0.48-0.75, p < 0.001), see Figure 4.

Figure 4.

Figure 4.

Comparison of restricted cubic spline and piecewise regression analyses for the association between CDAI and stroke risk. The restricted cubic spline curve illustrates the adjusted nonlinear association between CDAI and the estimated odds of stroke. The piecewise regression model was used to identify the optimal CDAI threshold by comparing model fit across candidate cut-off points. The vertical dashed line indicates the estimated CDAI threshold of 1.67. The lower Akaike Information Criterion (AIC) of the piecewise regression model compared with the RCS model indicates better model fit.

3.4. Interaction effect analysis

This study employs SHAP interaction plots to intuitively demonstrate the potential interaction effects involved.

The study found (Figure 5) that the risk of stroke was the highest (>12%) when CDAI was at a low level (<-2.1 ± 1.32) with a high blood cadmium concentration (>0.5 μg/L) or a high blood lead concentration (>2.2 μg/dL); when CDAI was at a high level (>0 ± 0.77), the risk of stroke remained low (<4%) even if the blood cadmium or blood lead concentration was high. For blood manganese, the risk of stroke was relatively high (>12%) when CDAI was at an extremely low level (<-2.8) and the blood manganese concentration was low (<6.0 μg/L). High CDAI levels and blood manganese may exert a synergistic protective effect. Finally, the study also showed that elevated blood lead and blood cadmium levels may reduce the protective effect of CDAI by 146% to 167%.

Figure 5.

Figure 5.

SHAP interaction plots showing the joint contribution of CDAI and blood heavy metals to predicted stroke risk. Each point represents one participant. The x-axis represents CDAI values, and the y-axis represents the SHAP interaction value between CDAI and the corresponding heavy metal. Positive SHAP interaction values indicate increased predicted stroke risk, whereas negative values indicate reduced predicted risk color gradients represent blood cadmium, lead, or manganese concentrations effects stroke risk.

3.5. Subgroup analysis

Subgroup analysis was performed by age and gender to evaluate the robustness of the model in different populations (Figure 6)

Figure 6.

Figure 6.

Model Performance by subgroup forest plot Points represent the subgroup-specific AUC estimates, and horizontal lines represent the corresponding 95% confidence intervals. Wider confidence intervals indicate greater uncertainty in subgroup-specific model performance, while overlapping intervals suggest similar discriminative performance between subgroups.

The study found that there was no significant gender difference in the subgroups. The model performed similarly in females (n=18,122; AUC=0.75, 95% CI: 0.69–0.83) and males (n=17,049; AUC=0.75, 95% CI: 0.67–0.82), with highly overlapping 95% CIs. However, significant heterogeneity was observed across age groups. Among individuals aged <65 years (n=26,817), the model had an AUC of 0.72 (95% CI: 0.65–0.78), demonstrating a good discriminative ability for stroke risk. In contrast, among the elderly population aged ≥65 years (n=8,354), the performance of the model decreased significantly (AUC=0.55, 95% CI: 0.52–0.58).

4. Discussion

The findings of this study explore the effectiveness and interpretability of machine learning in the prevention of stroke through dietary and environmental exposure interventions. First, model performance comparisons demonstrate that HistGradientBoosting achieves the optimal performance; however, the PR-AUC and F1 scores of all observed models are relatively low. This is primarily attributed to the extreme class imbalance of stroke outcomes in the general population, a common challenge in population-based rare disease prediction studies. 13 Consequently, high-performance clinical screening tools for rare diseases cannot be developed relying solely on machine learning models. Despite the relatively low predictive metrics, the findings regarding interaction effects and protective thresholds based on SHAP analysis remain robust and still carry certain dietary guidance value. The SHAP interpretability method quantifies the contribution of each feature, providing insights for clinical interventions. Further analysis reveals that there may be a strong interactive association between the Composite Dietary Antioxidant Index (CDAI) and blood cadmium as well as blood lead levels, and the CDAI exhibits an L-shaped nonlinear protective effect: when the CDAI exceeds 1.671, the intake of antioxidant nutrients presents a certain protective effect, whereas heavy metal exposure weakens this protective role. Interaction plots indicate that stroke is sensitive to cadmium and lead exposure, with a notable elevation in SHAP values within small intervals. Existing research has shown that after cadmium accumulates inside cells, it can induce mitochondria to produce a large number of superoxide free radicals, leading to oxidative stress and damaging vascular endothelial cells and neurons. 14 It also inhibits the activity of endothelial nitric oxide synthase (eNOS) while promoting the release of endothelin-1 (ET-1), resulting in vasoconstriction, increased endothelial permeability, and accelerated atherosclerosis. 15 Lead, upon entering the body, depletes intracellular antioxidants such as thiols, triggering excessive production of reactive oxygen species (ROS), causing oxidative damage to vascular endothelial cells, increased vascular permeability, reduced nitric oxide bioavailability, and promoting vasoconstriction and platelet aggregation, thereby creating conditions for thrombosis. 16 These results are consistent with the statistical findings of the present study. In addition, these two metals are prone to exposure: smoking is one of the primary routes of cadmium exposure, and blood cadmium levels in smokers are typically several times higher than those in non-smokers. 17 For non-smokers, food serves as the main source, with cadmium entering the food chain via contaminated soil, particularly accumulating in rice, wheat, leafy vegetables, root vegetables, shellfish, and animal offal. 18 In Asia, rice is the leading source of cadmium intake. 19 Manganese is a key cofactor of manganese superoxide dismutase (MnSOD, also known as SOD2), a core antioxidant enzyme that scavenges superoxide anion radicals in the mitochondrial matrix and maintains cellular redox homeostasis. 20 Adequate manganese levels ensure the activity of manganese superoxide dismutase, thereby mitigating oxidative stress-induced damage to vascular endothelial and neuronal cells. This accounts for the synergistic protective effect observed at higher CDAI levels in this study.

Based on global SHAP importance analysis, vitamin C and selenium were identified as crucial influencing factors for stroke outcomes. With reference to the nutritional guidelines of the United States Department of Agriculture and the Dietary Reference Intakes, 21 this threshold should be interpreted as a population-specific risk-stratification point rather than a clinical dietary target. Specifically, maintaining CDAI above this critical value is roughly consistent with dietary patterns characterized by the following features: adequate intake of vitamin C (approximately 75–120 mg/day), which can be obtained from 1-2 servings of citrus fruits, or 1 serving of kiwifruit, strawberries, or bell peppers; sufficient intake of selenium (approximately 55–100 μ g/day), achievable by consuming 1 Brazil nut, or regular intake of fish, eggs, and whole grains. In terms of lifestyle, the findings of this study suggest that smokers should quit smoking as soon as possible, and obese individuals should control their body weight to reduce the risks associated with heavy metal exposure.

Notably, for populations with high levels of heavy metal exposure (blood cadmium ≥ 0.5μg/L, blood lead ≥2.2μg/dL), model results indicate that maintaining CDAI above the aforementioned level may reduce the estimated risk of stroke. Studies have shown that vitamin C, vitamin E, carotenoids (e.g., β-carotene, lycopene), and polyphenols (e.g., flavonoids) can directly react with ROS/RNS, neutralize free radicals, and alleviate oxidative damage. 22 In particular, as an essential cofactor of glutathione peroxidase (GPx), selenium can catalyze the reduction of peroxides and protect cell membranes and DNA. 23 However, this indicator can only serve as a reference standard for risk stratification and should not be regarded as an intervention threshold with causal implications. Finally, this study provides insights into stroke prevention from the perspectives of dietary nutrition and the environment, identifying potential associations between nutrition and heavy metal exposure. Furthermore, forest plots demonstrate that the model performs better in individuals under 65 years of age, which may have positive implications for early-stage and non-stroke populations.

Although this study has achieved preliminary results, it still has several limitations. First, longitudinal cohort studies or intervention trials are needed to verify the temporal effects of the interaction between the Comprehensive Dietary Antioxidant Index and heavy metals. Additionally, the dietary data in this study rely on subject recall, which may introduce random errors. Second, the model shows satisfactory internal validation performance but has not undergone generalization testing using external data. Future research should integrate multi-center data to enhance the generalizability of the model. Last, the specific intrinsic biological mechanisms underlying their interaction remain unclear, and further in-depth mechanistic studies combining multi-omics data are required to determine whether there is a clinical effect.

5. Conclusion

This cross-sectional study using NHANES 2005–2018 data established an explainable machine learning framework for stroke risk assessment. The HistGradientBoosting model showed optimal predictive performance. We identified a nonlinear L-shaped protective threshold of CDAI at 1.67, above which stroke risk was significantly reduced. High blood lead and cadmium exposure markedly weakened the protective effect of CDAI, while high manganese exhibited synergistic antioxidant protection. SHAP analysis verified robust interactions between CDAI and heavy metals. These findings support dietary antioxidant intervention for stroke prevention, especially in populations with heavy metal exposure, and provide evidence for personalized nutritional strategies. Machine learning demonstrates value in exploring complex risk interactions but is not suitable for direct clinical screening.

Acknowledgements

We thank the National Center for Health Statistics for making the NHANES data publicly available.

Footnotes

Author contributions: Yixuan He: Conceptualization, data curation, formal analysis, methodology, software, visualization, writing—original draft.

Kai Gong, Quan Lan: Supervision, funding acquisition, project administration, writing—review and editing, correspondence.

Funding: The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by Medical Project of Xiamen (No. 3502Z20244ZD1045),2024 Xiamen Municipal Health High-Quality Development Science and Technology Plan Project(2024GZL-GG47).

The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

ORCID iD

Yixuan He https://orcid.org/0009-0000-6179-2935

Data Availability Statement

The data used in this study are available from the NHANES website: https://www.cdc.gov/nchs/nhanes/

References

  • 1.Campbell BCV, De Silva DA, Macleod MR, et al. Ischaemic stroke. Nat Rev Dis Primers 2019; 5(1): 70. 10.1038/s41572-019-0118-8 [DOI] [PubMed] [Google Scholar]
  • 2.Wang J, Wen X, Li W, et al. Risk factors for stroke in the Chinese population: a systematic review and meta-analysis. J Stroke Cerebrovasc Dis 2017; 26(3): 509–517. 10.1016/j.jstrokecerebrovasdis.2016.12.002 [DOI] [PubMed] [Google Scholar]
  • 3.Teng TQ, Liu J, Hu FF, et al. Association of composite dietary antioxidant index with prevalence of stroke: insights from NHANES 1999-2018. Front Immunol 2024; 15: 1306059. 10.3389/fimmu.2024.1306059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Xu D, Liu CH, Su FH, et al. Research on Heavy Metal Exposure and Stroke Risk Prediction Based on Interpretable Machine Learning Models. Chin J Mod Med 2024; 26(3): 14–19. [Google Scholar]
  • 5.Khosravi B, Weston AD, Nugen F, et al. Demystifying statistics and machine learning in analysis of structured tabular data. J Arthroplasty 2023; 38(10): 1943–1947. 10.1016/j.arth.2023.08.045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Park KK, Saleem M, Al-Garadi MA, et al. Machine learning applications in studying mental health among immigrants and racial and ethnic minorities: an exploratory scoping review. BMC Med Inform Decis Mak 2024; 24(1): 298. 10.1186/s12911-024-02663-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.von Elm E, Altman DG, Egger M, et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Ann Intern Med 2007; 147(8): 573–577. 10.7326/0003-4819-147-8-200710160-00010 [DOI] [PubMed] [Google Scholar]
  • 8.Ma R, Zhou X, Zhang G, et al. Association between composite dietary antioxidant index and coronary heart disease among US adults: a cross-sectional analysis. BMC Public Health 2023; 23(1): 2426. 10.1186/s12889-023-17373-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chen X, Lu H, Chen Y, et al. Composite dietary antioxidant index was negatively associated with the prevalence of diabetes independent of cardiovascular diseases. Diabetol Metab Syndr 2023; 15(1): 183. 10.1186/s13098-023-01150-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Xiong B, Wang J, He R, et al. Composite dietary antioxidant index and sleep health: a new insight from cross-sectional study[J]. BMC Public Health, 2024, 24(1): 609. 10.1186/s12889-024-18047-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst 2017; 30: 4768-4777. [Google Scholar]
  • 12.Raslan KSH, Alsharkawy AS, Raslan KR. iHHO-SMOTe: A cleansed approach for handling outliers and reducing noise to improve imbalanced data classification. arXiv preprint arXiv:2504.12850, 2025. [Google Scholar]
  • 13.Ghanem M, Ghaith AK, El-Hajj VG, et al. Limitations in evaluating machine learning models for imbalanced binary outcome classification in spine surgery: A systematic review. Brain Sci 2023; 13(12): 1723. 10.3390/brainsci13121723 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Gorini F, Tonacci A. Metal toxicity and dementia including frontotemporal dementia: current state of knowledge. Antioxidants 2024; 13(8): 938. 10.3390/antiox13080938 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zhou D, Mao Q, Sun Y, et al. Association of blood copper with the subclinical carotid atherosclerosis: an observational study. J Am Heart Assoc 2024; 13(9): e033474. 10.1161/jaha.123.033474 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ahmed G, Rahaman MS, Perez E, et al. Associations of environmental exposure to arsenic, manganese, lead, and cadmium with Alzheimer’s disease: A review of recent evidence from mechanistic studies. J Xenobiotics 2025; 15(2): 47. 10.3390/jox15020047 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Tratnik JS, Kocman D, Horvat M, et al. Cadmium exposure in adults across Europe: Results from the HBM4EU Aligned Studies survey 2014–2020. Int J Hyg Environ Health 2022; 246: 114050. 10.1016/j.ijheh.2022.114050 [DOI] [PubMed] [Google Scholar]
  • 18.Amzal B, Julin B, Vahter M, et al. Population toxicokinetic modeling of cadmium for health risk assessment. Environ Health Perspect 2009; 117(8): 1293–1301. 10.1289/ehp.0800317 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lordan R, Zabetakis I. Cadmium: A focus on the brown crab (Cancer pagurus) industry and potential human health risks. Toxics 2022; 10(10): 591. 10.3390/toxics10100591 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Horning KJ, Caito SW, Tipps KG, et al. Manganese is essential for neuronal health. Annu Rev Nutr 2015; 35(1): 71–108. 10.1146/annurev-nutr-071714-034419 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.United States Department of Agriculture . Dietary Reference Intakes [Internet]: United States Department of Agriculture, 2026. Available from:. https://www.nal.usda.gov/dietary-reference-intakes [Google Scholar]
  • 22.Koyama H, Kamogashira T, Yamasoba T. Heavy metal exposure: molecular pathways, clinical implications, and protective strategies. Antioxidants 2024; 13(1): 76. 10.3390/antiox13010076 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Jan AT, Azam M, Siddiqui K, et al. Heavy metals and human health: mechanistic insight into toxicity and counter defense system of antioxidants. Int J Mol Sci 2015; 16(12): 29592–29630. 10.3390/ijms161226183 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data used in this study are available from the NHANES website: https://www.cdc.gov/nchs/nhanes/


Articles from Science Progress are provided here courtesy of SAGE Publications

RESOURCES