Abstract
Patients with type 2 diabetes mellitus (T2DM) have a significantly higher risk of cardiovascular disease (CVD) compared to the general population. Accurately predicting this risk is crucial for developing personalized treatment plans and public health interventions. This study aims to develop and validate a model for predicting CVD risk in T2DM patients using the Boruta feature selection algorithm and machine learning methods. We analyzed data from the National Health and Nutrition Examination Survey (NHANES) from 1999 to 2018. Six machine learning (ML) models, including Multilayer Perceptron (MLP), Light Gradient Boosting Machine (LightGBM), Decision Tree (DT), Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), and k-Nearest Neighbors (KNN), were employed for model development and validation. Boruta was used for optimal feature selection. The performance of the machine learning models was comprehensively evaluated using ROC curves, accuracy, and other related metrics. Shapley Additive Explanation (SHAP) analysis was conducted for visual interpretation, and the Shinyapps.io platform was utilized to deploy the best-performing models as web-based applications. A total of 4,015 T2DM patients were included, among which 999 (24.9%) had CVD. Model evaluation revealed significant overfitting with the KNN algorithm, which showed perfect discrimination in the training set but performed poorly in the test set (AUC = 0.64). In contrast, XGBoost demonstrated more consistent performance between training and testing datasets (AUC = 0.75 and 0.72, respectively), indicating better generalization ability and making it more suitable for clinical application. Using SHAP analysis, the top 10 important influencing factors identified by the XGBoost model were utilized to construct a CVD risk prediction platform for T2DM patients. The prediction model based on Boruta feature selection and machine learning shows promising results in assessing the CVD risk among T2DM patients. This study provides a viable tool for clinical use, facilitating early intervention and precision treatment.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-025-18443-7.
Keywords: Boruta feature selection, Machine learning, Type 2 diabetes, Cardiovascular disease, T2DM, CVD
Subject terms: Cardiology, Cardiovascular diseases, Diabetes
Introduction
Cardiovascular disease (CVD) is the leading cause of morbidity and mortality in patients with type 2 diabetes mellitus (T2DM)1,2. Compared to non-diabetic populations, T2DM patients have a significantly increased risk of developing CVD, along with poorer prognoses3. Specifically, the risk of CVD events such as coronary heart disease, stroke, and heart failure in T2DM patients is 2 to 4 times higher than that in the general population, and the mortality rate is significantly elevated4–6. The pathophysiological mechanisms behind this increased CVD risk are complex and closely related to various metabolic abnormalities commonly seen in T2DM patients, including insulin resistance, hyperglycemia, dyslipidemia, hypertension, and chronic inflammatory states7–9. Furthermore, T2DM patients often experience microvascular complications and endothelial dysfunction, which further exacerbate the occurrence and progression of CVD10. Therefore, accurately predicting the CVD risk in T2DM patients is crucial for developing personalized treatment plans, improving patient outcomes, and optimizing public health interventions.
Traditional cardiovascular disease (CVD) risk assessment primarily relies on risk scoring systems, such as the Framingham Risk Score, UKPDS Risk Score, and Reynolds Risk Score11–13. However, the predictive accuracy of these scoring systems in patients with type 2 diabetes mellitus (T2DM) is subject to numerous limitations and may not fully capture the complexity and heterogeneity of CVD risk in T2DM patients. On one hand, traditional risk scores are often based on a limited number of clinical variables, such as age, sex, blood pressure, and cholesterol levels, making it difficult to comprehensively reflect the overall risk status of T2DM patients14–16. On the other hand, these scoring systems generally assume a linear relationship among variables, which does not align with the complex pathophysiological mechanisms of CVD risk in T2DM patients. In fact, the CVD risk in T2DM patients is influenced by multiple interacting factors, including genetic factors, environmental factors, and lifestyle17. In recent years, with the rapid advancement of big data technology and artificial intelligence, machine learning (ML) methods have shown tremendous potential in the field of disease risk prediction18–20, providing new insights for improving the accuracy and individualization of CVD risk prediction in T2DM patients.
In recent years, diverse machine learning (ML) algorithms have been increasingly used for disease prediction in clinical settings. Each model presents specific strengths and limitations depending on the nature of the classification task and clinical requirements. For instance, logistic regression is widely valued for its interpretability and simplicity, but may not capture intricate, non-linear relationships between risk factors and outcomes21. Tree-based ensemble models such as Random Forest and XGBoost are robust to outliers, accommodate variable interactions, and can process both categorical and continuous data; however, they are prone to overfitting, especially with small sample sizes. Neural network-based models, for example multilayer perceptrons, can leverage large datasets with complex relationships but often require considerable computational resources and are frequently viewed as “black boxes” regarding clinical interpretability22. K-Nearest Neighbors (KNN) is a straightforward approach suitable for small-scale problems, yet its performance diminishes with high-dimensional data and class imbalance23.
Recent studies have demonstrated the advantages of ML models in disease risk prediction. For example, Alaa et al.20 used automated machine learning to predict cardiovascular disease (CVD) risk in the UK Biobank cohort and achieved better performance than conventional risk scores.However, the application of ML in medical research faces substantial challenges, such as class imbalance, heterogeneous data distributions, missing values, the risk of overfitting, and the need for interpretability to enable clinical adoption24.
Feature selection is essential for enhancing both model performance and clinical interpretability by identifying the most informative variables while reducing redundancy and noise25. Among various feature selection techniques, we used the Boruta algorithm due to its high stability and effectiveness in high-dimensional, potentially correlated clinical datasets. Boruta is a random forest-based wrapper algorithm that iteratively compares feature importance with that of randomly permuted “shadow” features, thus identifying all relevant predictors rather than just a minimal subset26. This approach is especially advantageous in clinical research, where disease risk is typically influenced by multiple interacting factors rather than a single predictor.
Recent comparative studies and reviews have further validated the utility of Boruta in clinical prediction. For example, a 2022 study by Degenhardt, Seifert, and Szymczak systematically evaluated various feature selection methods on high-dimensional biomedical data and found that Boruta consistently performed well in terms of selection stability and the identification of relevant variables, which is crucial for reproducible clinical findings27. Furthermore, another study by Speiser et al., which compared multiple machine learning algorithms for clinical outcome prediction, demonstrated that combining tree-based models with wrapper feature selection methods like Boruta can effectively handle multicollinearity and interaction effects common in clinical datasets28. These recent studies provide strong evidence supporting our choice of Boruta as a scientifically rigorous and appropriate algorithm for identifying the complex multivariate patterns that characterize cardiovascular disease risk in patients with type 2 diabetes.
This study aims to develop and validate a novel predictive model for cardiovascular disease (CVD) risk in patients with type 2 diabetes mellitus (T2DM) using machine learning methods combined with the Boruta feature selection algorithm. We chose the Boruta algorithm for feature selection to overcome the limitations of traditional statistical methods in handling high-dimensional data, effectively identifying key variables associated with CVD risk. The Boruta algorithm is a feature selection method based on random forests that iteratively compares original features with randomly generated shadow features, allowing for the effective identification of key features related to the target variable while removing redundant and noisy features29. We utilized data from the National Health and Nutrition Examination Survey (NHANES) from 1999 to 2018, which has the advantages of a large sample size, strong representativeness, and rich variable information, providing a reliable data foundation for this research. We compared six commonly used machine learning models; additionally, an online predictive model was developed on the Shinyapps.io platform, aimed at providing a valuable tool for clinicians to facilitate early intervention and precise treatment. This online prediction tool features a user-friendly interface and convenient operation, helping clinicians quickly assess the CVD risk of T2DM patients and formulate individualized treatment plans.
Methods
Data source and study population
The National Health and Nutrition Examination Survey (NHANES) (https://www.cdc.gov/nchs/nhanes/index.html) is an ongoing study aimed at assessing the health status of U.S. residents. This survey has received approval from the National Center for Health Statistics Institutional Review Board, and all participants provided written informed consent. Therefore, the Ethics Committee of The Fourth Affiliated Hospital of Soochow University exempted this study from ethical review. The data used in this study comes from the NHANES database for the period of 1999–2018.
This study included 9,038 patients with type 2 diabetes aged 18 years and older from 1999 to 2018. Participants were excluded for lacking cardiovascular disease data (n = 26), demographic information (n = 1471),including marital status (n = 64), education (n = 16), smoking (n = 6), alcohol consumption (n = 1034), hypertension (n = 1), chronic kidney disease (CKD) (n = 349), and chronic obstructive pulmonary disease (COPD) (n = 1); as well as for lacking fasting blood glucose data (n= 3426). The missing values for other continuous variables were all below 10%,to address missing values within our dataset, we performed multiple imputations using the Multiple Imputation by Chained Equations (MICE) method. MICE is a flexible and widely used imputation technique that models each variable with missing data conditionally on the other variables in an iterative fashion. Specifically, MICE performs the following steps1: Initially, missing values are filled in with simple estimates (for example, mean or median)2. Then, for each variable with missing data, a regression model is built where the missing values are treated as the dependent variable and all other variables are used as predictors3. The missing values are imputed based on the model predictions, and this process is repeated for each variable with missing data in a chained manner4. The whole process is cycled several times, resulting in multiple complete datasets that reflect the uncertainty of the missing values.
We chose MICE over other imputation methods (such as k-nearest neighbors (kNN), random forest (RF), or PLS-based approaches) because MICE is particularly well-suited for clinical datasets containing different types of variables (continuous, categorical, binary) and missing patterns. Unlike simpler methods, MICE accounts for multivariate relationships among variables and produces multiple imputed datasets, which fully incorporates the uncertainty caused by missingness and reduces the risk of bias. Compared to kNN or RF, which are non-parametric and may be sensitive to parameter settings or the structure of missing data, MICE is more interpretable and extensible. PLS-based imputation is mainly applicable to continuous variables and usually less flexible for mixed data types.
Regarding the option of deleting entries (rows) with missing data (“complete case analysis”), we decided against this approach as it would significantly reduce sample size and statistical power, and may introduce selection bias if the data are not missing completely at random. MICE allows the analysis to benefit from the full available information and thus is preferable in our context.Ultimately, the eligible participants included in the analysis numbered 4,015, as shown in Figs. 1, 2, 3, 4.
Fig. 1.

Study flow chart.
Fig. 2.
Feature importance ranking from the Boruta feature selection algorithm. Features are ranked in descending order of importance score. Green bars represent confirmed important features that significantly contribute to cardiovascular disease prediction, while red bars represent features identified as unimportant by the algorithm. The x-axis shows the feature importance score derived from random forest importance measure, while the y-axis lists the features with their full descriptive names. Higher scores indicate greater variable importance in predicting cardiovascular disease.PIR: Poverty income ratio, BMI: Body mass index, HbA1c: Glycosylated hemoglobin, BUN: Blood urea nitrogen, CKD: Chronic kidney disease, BMI: Body mass index, TG: Triglyceride, TC: Total cholesterol, LDL: Low density lipoprotein, FPG: Fasting plasma glucose, eGFR: Estimated glomerular filtration rate, WC: Waist circumference, ALT: Alanine aminotransferase, AST: Aspartate aminotransferase, GGT: Gamma-glutamyl transferase, COPD: Chronic obstructive pulmonary disease.
Fig. 3.
Performance comparison of different machine learning models on the training (A) and test (B) sets across multiple evaluation metrics. The heatmaps display the performance of MLP: Multi-Layer Perceptron, LightGBM: Light Gradient Boosting Machine, DT: Decision Tree, XGBoost: Extreme Gradient Boosting, LR: Logistic Regression, KNN: K-Nearest Neighbors.Each cell represents the value of a specific evaluation metric, including accuracy, balanced accuracy, F1 score, J-index, kappa, Matthew’s correlation coefficient (MCC), positive predictive value (PPV), negative predictive value (NPV), precision, recall, ROC AUC, sensitivity (sens), and specificity (spec). Higher values are indicated by red, while lower values are represented by green, showing the model’s effectiveness in both training and test sets.
Fig. 4.
ROC Analysis (A, B), Decision Curve Analysis (C) and Calibration Curve (D) of Machine Learning Model; MLP: Multi-Layer Perceptron, LightGBM: Light Gradient Boosting Machine, DT: Decision Tree, XGBoost: Extreme Gradient Boosting, LR: Logistic Regression, KNN: K-Nearest Neighbors. (C). Decision curve analysis (DCA) for different machine learning models.The y-axis represents net benefit, and the x-axis represents threshold probability. The ‘Treat All’ line indicates the net benefit if all patients were to receive intervention; the ‘Treat None’ line shows the net benefit if no patients were to receive intervention. Curves above these two reference lines indicate clinical utility of the model. The XGBoost model (red line) shows the highest net benefit in the threshold range of 10%−40%, which aligns with clinically relevant decision thresholds for CVD risk in T2DM patients, indicating optimal clinical application value in real-world settings. (D). Calibration curve for the XGBoost model. The x-axis represents predicted probability of CVD, and the y-axis represents the actual observed proportion of CVD. The diagonal line (dashed line) represents the ideal calibration line where predicted probabilities perfectly match observed rates. The calibration curve of the XGBoost model (solid blue line) closely follows the ideal calibration line, particularly in the moderate risk range (20%−60%), indicating good agreement between predicted probabilities and actual risk. This good calibration performance is crucial for accurate risk stratification and clinical decision-making.
Diagnosis of type 2 diabetes(T2DM)
The diagnostic criteria for type 2 diabetes mellitus (T2DM) include the following indicators: fasting blood glucose ≥ 7.0 mmol/L, or a 2-hour oral glucose tolerance test result ≥ 11.1 mmol/L; random blood glucose ≥ 11.1 mmol/L; glycated hemoglobin (HbA1c) ≥ 6.5%; additionally, individuals using diabetes medications or insulin, or those with a diabetes diagnosis confirmed by a physician, are also considered to meet the criteria30.
Diagnosis of cardiovascular disease (CVD)
The diagnosis of cardiovascular disease (CVD) was conducted through standardized medical questionnaires administered during personal interviews, where participants confirmed their own health status through self-reporting. Investigators asked participants whether they had ever been diagnosed with conditions such as coronary artery disease, congestive heart failure, angina, myocardial infarction, or stroke. If participants responded “yes,” they were classified as having cardiovascular disease31.
Covariates
This study included various covariates such as age, gender, race, educational background, marital status, family income poverty ratio (PIR), hypertension, chronic obstructive pulmonary disease (COPD), chronic kidney disease (CKD), use of antidiabetic medications, body mass index (BMI), smoking and drinking habits. Race was categorized into non-Hispanic Black, non-Hispanic White, Mexican American, and others. The family income poverty ratio (PIR) was divided into three categories: <1.3, 1.3–3.5, and > 3.5. Marital status included married, divorced, unmarried, and other options. Smoking was classified into three groups: ① never smoked; ② formerly smoked; ③ currently smoking. Drinking status was categorized into five groups: never drank, previously drank but currently does not, heavy drinker, moderate drinker, and light drinker. Laboratory indicators included estimated glomerular filtration rate (eGFR), waist circumference (WC), lymphocytes, neutrophils, hemoglobin, platelets, creatinine, uric acid, blood urea nitrogen (BUN), glucose, glycosylated hemoglobin (HbA1c), alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin, albumin, gamma-glutamyl transferase (GGT), triglycerides (TG), total cholesterol (TC), high-density lipoprotein (HDL), and low-density lipoprotein (LDL).
Statistical analysis
Basic analysis and regression analysis
This study utilized R software (version 4.3.2, https://www.r-project.org) and SPSS 26.0 for data analysis. We followed the analysis and reporting standards of NHANES, taking into account the complex sampling design and its weights. In the weighted analysis, the MEC sample weights (a combination of WTMEC2YR/4 and WTMEC2YR/8) were used. A multicollinearity test was conducted on the continuous variables included in the analysis; the results showed that the variance inflation factor (VIF) for HDL exceeded 5, so it was excluded, while the VIFs for the other variables were all below 5, indicating that no significant multicollinearity existed.
In this study, after careful consideration of multiple feature selection techniques, we chose the Boruta algorithm for variable feature screening.Boruta offers distinct advantages over methods such as LASSO and embedded boosting techniques that made it particularly suitable for our research objectives.Boruta’s strengths include1: its ability to capture non-linear relationships through its Random Forest foundation, which is essential for the complex interactions in cardiovascular disease pathophysiology2; its statistically rigorous approach that establishes significance thresholds for feature importance rather than using arbitrary cutoffs; and3 its algorithm-agnostic nature, allowing consistent feature selection across the multiple machine learning models we compared in this study.While LASSO offers computational efficiency and embedded methods provide algorithm-specific optimization, Boruta’s all-relevant feature selection approach and ability to handle complex, non-linear relationships made it the most appropriate choice for our specific research context.In comparing the characteristics of T2DM patients with and without CVD, one-way ANOVA and Pearson’s Chi-square test were employed to assess the differences between continuous and categorical variables. A p-value of less than 0.05 (two-tailed) was considered statistically significant.
Construction of machine learning models
In this study, we randomly divided the dataset into a training set and a testing set with a ratio of 7:3. The training set was used for model parameter computation and training, while the testing set was used to evaluate model performance. We selected six machine learning models: Multilayer Perceptron (MLP), Light Gradient Boosting Machine (LightGBM), Decision Tree (DT), Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), and K-Nearest Neighbors (KNN) to predict the CVD risk in T2DM patients.Multilayer Perceptron (MLP): A feedforward artificial neural network capable of modeling complex non-linear relationships. Our implementation used a three-layer architecture (input, hidden, output) with ReLU activation functions for hidden layers and sigmoid activation for the output layer to model binary classification probabilities.Light Gradient Boosting Machine (LightGBM): A gradient boosting framework that uses tree-based algorithms and implements histogram-based algorithms for faster training. LightGBM is particularly advantageous for handling heterogeneous clinical data with both categorical and continuous variables.Decision Tree (DT): A non-parametric supervised learning method that creates a model predicting target values by learning simple decision rules from data features. We implemented the CART (Classification and Regression Trees) algorithm with Gini impurity as the splitting criterion. Extreme Gradient Boosting (XGBoost): An optimized implementation of gradient boosting that employs regularization techniques to prevent overfitting. XGBoost is especially suitable for structured clinical data with potential non-linear relationships among variables.Logistic Regression (LR): A linear model for binary classification that estimates the probability of class membership using a logistic function. Despite its simplicity, LR often serves as a strong baseline for classification tasks and offers high interpretability.K-Nearest Neighbors (KNN): A non-parametric method that classifies objects based on the majority class among their k nearest neighbors, making it useful for detecting local patterns in the feature space.To address the class imbalance between the CVD and non-CVD groups, we applied the Synthetic Minority Over-sampling Technique (SMOTE) only to the training set to generate synthetic samples for the minority class, optimizing model performance32. It is important to emphasize that SMOTE was applied exclusively to the training data and not to the entire dataset before splitting, ensuring that the test set remained completely independent, thus avoiding data leakage and biased estimates.Additionally, to prevent overfitting and enhance model generalization ability, we implemented rigorous 10-fold cross-validation for each model. Specifically, the training data was randomly divided into 10 equal subsets, with 9 subsets used for training and the remaining subset used for validation. This process was repeated 10 times, ensuring each subset was used once as a validation set. The average performance metrics across all 10 validation runs were used to evaluate the models. This repeated cross-validation strategy is particularly suitable for handling the heterogeneous data present in our study, as it ensures that the model performs well across different data distributions. Eventually, through multiple iterations and parameter adjustments, we selected the best-performing model.The optimal hyperparameters for all models are presented in Table S1.Our hyperparameter optimization process followed these steps:1. Initial hyperparameter range definition based on literature review and domain knowledge.2. Implementation of grid search with 10-fold cross-validation on the training set to systematically explore all possible combinations.3. Selection of optimal hyperparameters for each model based on AUC-ROC maximization.4. Final model training using the entire training set with the optimal hyperparameter configuration.5. Performance validation on the held-out test set to assess generalizability.We tested all parameter combinations through grid search to find the best performance parameters. Model performance was evaluated using metrics such as ROC curve, accuracy, sensitivity, specificity, F1 score, recall, and PR.
Explanability of machine learning models
We used the SHAP algorithm to calculate SHAP values for each variable to identify the best-performing model. SHAP (Shapley Additive Explanations) is a game theory-based machine learning prediction method that provides a unified framework for interpreting machine learning results, aimed at analyzing complex black-box models33.Our visualizations employ a consistent color scheme where blue represents higher feature values and red represents lower values. SHAP values indicate the magnitude and direction of a feature’s influence on prediction deviation from baseline, with positive values indicating increased CVD risk and negative values indicating decreased risk. Figure 5A displays features ranked by importance (based on mean |SHAP| values), Fig. 5B shows the relationship between feature values and their predictive impact, while Fig. 6 uses force plots to illustrate individual patient predictions and major contributing factors. This multi-level visualization approach allows us to understand the model’s predictive mechanisms from both aggregate and individual perspectives.We showcased the relationship between features and the prediction model using a macro feature importance graph and a beeswarm plot, and provided a detailed analysis of the top 10 important variables. Additionally, we developed a simple online prediction model for healthcare practitioners to access conveniently via Shinyapps.io.A comprehensive “Usage Guide” has been integrated into the Shiny application, providing instructions on input data, application use, and output interpretation. Additionally, a disclaimer appears on both the main page and the guide, emphasizing that predictions are for research and educational reference only and should not replace clinical judgment.
Fig. 5.
Feature importance analysis and SHAP value distribution in the XGBoost model.Feature-ranking plots (A) and summary plots (B). (A). Feature importance ranking plot showing the average contribution of each feature to CVD risk prediction (sorted by mean |SHAP value|). Bar length indicates feature importance, with longer bars showing greater impact on prediction outcomes. Age, estimated glomerular filtration rate (eGFR), and neutrophil count are the three most important predictors of CVD risk. (B). SHAP value summary plot showing how feature values influence predictions. The x-axis represents SHAP values (contribution to prediction), and point colors indicate feature values (red = high, blue = low). Each point represents one sample. For example, advanced age (red points) typically has positive SHAP values, indicating that aging increases CVD risk; similarly, low eGFR values (blue points) also have positive SHAP values, suggesting that decreased kidney function increases CVD risk.
Fig. 6.
SHAP force plot analysis of individual patient CVD risk prediction. (A) Force plot of a patient with CVD, themodel generated a predicted CVD probability of 61.9%. (B) Force plot of a patient without CVD, the model generated a predicted CVD probability of 36.2%. (A). Force plot for a positive CVD prediction. This patient has a predicted CVD risk of 61.9% (vs. baseline expectation of 36.6%). Red arrows indicate risk-increasing features, while blue arrows show risk-decreasing features. Major risk contributors include: chronic kidney disease (+ 0.0457), former smoking status (+ 0.0526), and lower total cholesterol (+ 0.0525). Major protective factors include AST level (−0.047) and mild alcohol consumption (−0.0416). Arrow length represents feature impact magnitude. (B). Force plot for a negative CVD prediction. This patient has a predicted CVD risk of 36.2% (close to baseline expectation of 36.6%). Major risk factors include hypertension (+ 0.0336) and anti-diabetic medication use (+ 0.0337), while major protective factors include platelet count (−0.0372) and mild alcohol consumption (−0.0358). This indicates a relatively balanced profile of risk and protective factors.
Results
Sorting the importance of CVD risk factors in T2DM patients with the Boruta algorithm
In the report generated by the Boruta algorithm, variables in the green area were identified as important factors that play a significant role in the model. Variables in the red area were deemed unimportant. A total of 32 factors were evaluated, among which 22 factors were confirmed to have a significant impact on CVD risk, while 10 factors were excluded, as shown in Fig. 2.
Baseline characteristics of participants
A total of 4,015 participants were included in this study. Analysis of the participants’ general information showed significant differences in multiple health indicators between the cardiovascular disease (CVD) group and the non-CVD group. The average age of the CVD group was 66.92 years (SE = 0.47), significantly higher than the 56.97 years (SE = 0.36) in the non-CVD group (P < 0.0001). The creatinine level in the CVD group was 100.34 µmol/L (SE = 2.72), also significantly higher than the 80.57 µmol/L (SE = 1.14) in the non-CVD group (P < 0.0001). Additionally, other biomarkers such as uric acid, BUN, ALT, albumin, TC, and LDL also showed significant differences between the two groups (P < 0.0001). eGFR showed a significant difference between the two groups, with the CVD group at 71.12 (SE = 1.09) and the non-CVD group at 88.01 (SE = 0.54) (P < 0.0001). Hypertension was more prevalent in CVD patients, 83.48% in the CVD group and 67.34% in the non-CVD group (P < 0.0001). Chronic kidney disease (CKD) was present in 56.66% of the CVD group compared with 35.05% of the non-CVD group (P < 0.0001). Demographic factors such as sex, race, and marital status were also associated with cardiovascular health. The proportion of men in the CVD group was 57.36%, which was higher than 50.50% in the non-CVD group (P = 0.002). Ethnicity was also significantly different between the groups (P < 0.0001). In terms of marital status, 53.45% were married in the CVD group and 57.36% in the non-CVD group (P < 0.0001). There were also significant differences between the two groups in education and PIR (poverty income ratio), with P values of 0.03 and 0.003, respectively. BMI was not significantly different between the two groups (P = 0.45). Smoking habits and alcohol intake were also significantly different between the two groups (P < 0.0001). Non-smokers were relatively high in the non-CVD group (51.62%) compared to 42.54% in the CVD group, as shown in Table 1.
Table 1.
Baseline characteristics of participants.
| Variables | Total(n = 4015) | Non-CVD(n = 3016) | CVD(n = 999) | P-value |
|---|---|---|---|---|
| Age, mean (SE) | 59.29(0.32) | 56.97(0.36) | 66.92(0.47) | < 0.0001 |
| Creatinine, mean (SE) | 85.17(1.10) | 80.57(1.14) | 100.34(2.72) | < 0.0001 |
| Uric acid, mean (SE) | 348.21(2.06) | 341.93(2.21) | 368.88(4.72) | < 0.0001 |
| BUN, mean (SE) | 5.72(0.05) | 5.38(0.05) | 6.84(0.13) | < 0.0001 |
| ALT, mean (SE) | 27.91(0.38) | 28.93(0.46) | 24.52(0.50) | < 0.0001 |
| AST, mean (SE) | 26.45(0.31) | 26.78(0.32) | 25.36(0.69) | 0.05 |
| Albumin, mean (SE) | 4.15(0.01) | 4.16(0.01) | 4.09(0.02) | < 0.0001 |
| TC, mean (SE) | 4.91(0.03) | 4.99(0.03) | 4.63(0.05) | < 0.0001 |
| LDL, mean (SE) | 2.82(0.02) | 2.90(0.03) | 2.56(0.04) | < 0.0001 |
| GGT, mean (SE) | 38.51(1.45) | 38.76(1.75) | 37.68(1.79) | 0.65 |
| FPG, mean (SE) | 8.41(0.07) | 8.44(0.08) | 8.31(0.13) | 0.42 |
| HbA1c, mean (SE) | 6.99(0.03) | 6.99(0.04) | 6.99(0.07) | 0.97 |
| eGFR, mean (SE) | 84.08(0.51) | 88.01(0.54) | 71.12(1.09) | < 0.0001 |
| Lymphocyte, mean (SE) | 2.04(0.02) | 2.07(0.02) | 1.95(0.03) | 0.002 |
| Platelet, mean (SE) | 243.59(1.59) | 248.57(1.79) | 227.17(2.72) | < 0.0001 |
| Neutrophils, mean (SE) | 4.50(0.04) | 4.45(0.04) | 4.68(0.06) | 0.003 |
| WC, mean (SE) | 109.80(0.38) | 109.37(0.47) | 111.25(0.79) | 0.05 |
| Hemoglobin, mean (SE) | 14.24(0.04) | 14.29(0.04) | 14.08(0.08) | 0.02 |
| Bilirubin, mean (SE) | 11.86(0.13) | 11.82(0.14) | 12.01(0.21) | 0.41 |
| TG, mean (SE) | 1.97(0.05) | 1.97(0.05) | 1.97(0.07) | 0.99 |
| Sex, n(%) | 0.002 | |||
| Male | 2096 (52.20) | 1523 (50.50) | 573 (57.36) | |
| Female | 1919 (47.80) | 1493 (49.50) | 426 (42.64) | |
| Race, n(%) | < 0.0001 | |||
| Mexican American | 803 (20.00) | 659 (21.85) | 144 (14.41) | |
| Non-Hispanic Black | 925 (23.04) | 704 (23.34) | 221 (22.12) | |
| Non-Hispanic White | 1586 (39.50) | 1095 (36.31) | 491 (49.15) | |
| Other | 701 (17.46) | 558 (18.50) | 143 (14.31) | |
| Marital, n(%) | < 0.0001 | |||
| Married | 2264 (56.39) | 1730 (57.36) | 534 (53.45) | |
| Never Married | 341 (8.49) | 291 (9.65) | 50 (5.01) | |
| Divorced | 489 (12.18) | 342 (11.34) | 147 (14.71) | |
| Unmarried but have/had partner | 921 (22.94) | 653 (21.65) | 268 (26.83) | |
| Education, n(%) | 0.03 | |||
| Less than high School | 1453 (36.19) | 1071 (35.51) | 382 (38.24) | |
| High school or equivalent | 948 (23.61) | 699 (23.18) | 249 (24.92) | |
| College or above | 1614 (40.20) | 1246 (41.31) | 368 (36.84) | |
| PIR, n(%) | 0.003 | |||
| < 1.3 | 1362 (33.92) | 995 (32.99) | 367 (36.74) | |
| 1.3–3.5 | 1656 (41.25) | 1239 (41.08) | 417 (41.74) | |
| > 3.5 | 997 (24.83) | 782 (25.93) | 215 (21.52) | |
| BMI, n(%) | 0.45 | |||
| < 25 | 594 (14.79) | 463 (15.35) | 131 (13.11) | |
| 25–30 | 1258 (31.33) | 955 (31.66) | 303 (30.33) | |
| > 30 | 2163 (53.87) | 1598 (52.98) | 565 (56.56) | |
| Smoke, n(%) | < 0.0001 | |||
| Never | 1982 (49.36) | 1557 (51.62) | 425 (42.54) | |
| Former | 1357 (33.80) | 950 (31.50) | 407 (40.74) | |
| Now | 676 (16.84) | 509 (16.88) | 167 (16.72) | |
| Alcohol, n(%) | < 0.0001 | |||
| Never | 739 (18.41) | 565 (18.73) | 174 (17.42) | |
| Former | 1106 (27.55) | 750 (24.87) | 356 (35.64) | |
| Mild | 1235 (30.76) | 927 (30.74) | 308 (30.83) | |
| Moderate | 393 (9.79) | 323 (10.71) | 70 (7.01) | |
| Heavy | 542 (13.50) | 451 (14.95) | 91 (9.11) | |
| Hypertension, n(%) | < 0.0001 | |||
| Yes | 2865 (71.36) | 2031 (67.34) | 834 (83.48) | |
| No | 1150 (28.64) | 985 (32.66) | 165 (16.52) | |
| Anti Diabetic, n(%) | < 0.0001 | |||
| Yes | 2242 (55.84) | 1601 (53.08) | 641 (64.16) | |
| No | 1773 (44.16) | 1415 (46.92) | 358 (35.84) | |
| COPD, n(%) | < 0.0001 | |||
| Yes | 261 (6.50) | 150 (4.97) | 111 (11.11) | |
| No | 3754 (93.50) | 2866 (95.03) | 888 (88.89) | |
| CKD, n(%) | < 0.0001 | |||
| Yes | 1623 (40.42) | 1057 (35.05) | 566 (56.66) | |
| No | 2392 (59.58) | 1959 (64.95) | 433 (43.34) |
Date are presented as mean (SE) or n(%);PIR: Poverty income ratio, BMI: Body mass index, HbA1c: Glycosylated hemoglobin, BUN: Blood urea nitrogen, CKD: Chronic kidney disease, BMI: Body mass index, TG: Triglyceride, TC: Total cholesterol, LDL: Low density lipoprotein, FPG: Fasting plasma glucose, eGFR: Estimated glomerular filtration rate, WC: Waist circumference, ALT: Alanine aminotransferase, AST: Aspartate aminotransferase, GGT: Gamma-glutamyl transferase, COPD: Chronic obstructive pulmonary disease.
Machine learning model comparison
The machine learning predictive model for CVD risk in patients with T2DM performed as shown in Fig. 3A and B for the training set and validation set. In the training set, the KNN algorithm performed the best, with an area under the ROC curve of 1, accuracy of 0.97, sensitivity of 0.97, specificity of 0.97, F1 score of 0.95, recall of 0.97, and area under the PR curve of 0.99. Following closely was XgBoost, with an area under the ROC curve of 0.75, accuracy of 0.68, sensitivity of 0.75, specificity of 0.66, F1 score of 0.54, recall of 0.75, and area under the PR curve of 0.5. In the validation set, the MLP model showed the best performance, with an area under the ROC curve of 0.74, accuracy of 0.64, sensitivity of 0.74, specificity of 0.61, F1 score of 0.51, recall of 0.74, and area under the PR curve of 0.44. XgBoost performed in the validation set with an area under the ROC curve of 0.72, accuracy of 0.62, sensitivity of 0.71, specificity of 0.60, F1 score of 0.48, recall of 0.71, and area under the PR curve of 0.43. In contrast, KNN performed poorly in the test set, with an area under the ROC curve of 0.64, accuracy of 0.65, sensitivity of 0.45, specificity of 0.71, F1 score of 0.39, recall of 0.45, and area under the PR curve of 0.34. It is important to note that the KNN model exhibited an AUC of 1.00 in the training set, which indicates substantial overfitting. This perfect performance on training data but significantly lower performance on testing data (AUC = 0.64) suggests that the model memorized the training examples rather than learning generalizable patterns. In contrast, although XGBoost showed more modest training performance (AUC = 0.75), its validation performance (AUC = 0.72) represented a smaller drop-off, indicating better generalization to unseen data. This balance between model complexity and generalizability was a key factor in our selection of XGBoost as the optimal model for clinical implementation.
Figure 4A and B show the ROC curve analysis for the training and validation sets.Decision curve analysis (DCA) is an important tool for evaluating the clinical value of prediction models by measuring net benefit across different threshold probabilities. As shown in Fig. 4C, the XGBoost model demonstrates optimal net benefit performance in the threshold range of 10–40%, which aligns with clinical decision thresholds for CVD risk management in T2DM patients. Within this range, the XGBoost model’s curve is notably higher than other models and the reference lines (‘Treat All’ and ‘Treat None’ lines), indicating that using the XGBoost model for risk prediction in this clinically relevant decision range provides the greatest clinical benefit. The net benefit calculation considers the trade-off between the benefit of correctly identifying true positive cases and avoiding false positive results, thus higher net benefit values indicate that the model can maximize identification of high-risk patients requiring preventive measures while minimizing unnecessary interventions. Figure 4D presents the calibration curve of the XGBoost model, where the diagonal line represents the ideal calibration reference line. The proximity of the calibration curve to the ideal diagonal line reflects the consistency between predicted probabilities and actual observed probabilities. As shown, the XGBoost model’s calibration curve closely follows the ideal calibration line, particularly in the moderate risk range (20%−60%), indicating that the model accurately estimates CVD risk probabilities in T2DM patients. Good calibration performance is critical for clinical application, as it ensures that physicians and patients can rely on the model’s predicted risk probabilities to guide treatment decisions and risk communication.
Explanation of machine learning models
We used SHAP analysis to evaluate the importance of various features and their impact on the prediction results in the XgBoost model. In the analysis, SHAP values greater than zero indicate an increased risk of CVD in T2DM patients, and the higher the SHAP value, the greater the risk. The results are shown in Fig. 5A and B. The analysis indicates that age is the most critical variable affecting the risk of CVD in T2DM patients, having the highest SHAP value. Additionally, hypertension, chronic kidney disease (CKD), smoking, and alcohol consumption were also identified as factors that increase the risk of CVD in T2DM patients.
At the individual level, the model (Fig. 6) aims to provide a personalized interpretation of the prediction results by inputting the actual values of the features. It not only shows the predicted probability of CVD risk in T2DM patients but also reveals the impact of individual characteristics on the model outcomes. Figure 6A presents a case of a positive prediction, with a predicted probability of CVD risk at 61.9%, while Fig. 6B displays a case of a negative prediction, with a predicted probability of 36.2%.
Web calculator implementation
We developed a deployable web platform that employs the SHAP algorithm to analyze feature importance and constructed a calculator for the top ten variables related to CVD risk in patients with T2DM (website: https://cvdshiny.shinyapps.io/shiny_cls2_1model_fastshap/). This platform provides a convenient online tool for assessing the CVD risk in T2DM patients. Users simply need to input their clinical feature data into the designated text boxes to easily obtain the desired prediction results (Fig. 7).
Fig. 7.
Screenshot of the Shiny application, which now incorporates an integrated usage guide and a visible disclaimer, supporting responsible interpretation of output.
Discussion
This study aimed to develop a machine learning-based cardiovascular disease (CVD) risk prediction model for patients with type 2 diabetes mellitus (T2DM) and validated it using the NHANES database. The results indicated that the XGBoost model demonstrated good performance in predicting CVD risk for T2DM patients, with an area under the ROC curve of 0.77 (95% CI: 0.75–0.78). This finding is consistent with previous research, suggesting that machine learning models have certain advantages in predicting CVD risk34,35. Some studies have employed various machine learning models to predict CVD risk and found that the predictive performance of the XGBoost model outperformed the traditional Framingham risk score36.
Through Boruta feature selection, we identified 22 factors significantly associated with cardiovascular disease (CVD) risk in patients with type 2 diabetes mellitus (T2DM). These factors include traditional CVD risk factors such as age, hypertension, chronic kidney disease (CKD), smoking, and alcohol consumption, as well as some biomarkers like creatinine, uric acid, and estimated glomerular filtration rate (eGFR). These biomarkers reflect common renal function abnormalities and metabolic disorders in T2DM patients, which may be closely related to the occurrence and progression of CVD. SHAP analysis further revealed the specific impact of these factors on prediction outcomes. For instance, age emerged as the most important predictive factor with the highest SHAP value, which is consistent with previous research indicating that age is a significant determinant of CVD risk37. Hypertension and CKD are also important risk factors, reflecting that T2DM patients often have multiple comorbidities that can synergistically increase CVD risk38,39. The results of this study suggest that actively controlling blood pressure and delaying the progression of CKD may help reduce CVD risk in T2DM patients. Notably, lifestyle factors such as smoking and alcohol consumption were also identified as important risk factors, indicating that changing harmful lifestyle habits may lower CVD risk in T2DM patients40–42.
In this study, the average age of the CVD group was significantly higher than that of the non-CVD group (66.92 years vs. 56.97 years, P < 0.0001). This indicates that age is an important risk factor for the occurrence of CVD in T2DM patients. With increasing age, vascular elasticity gradually declines, and the severity of atherosclerosis worsens, leading to an increased risk of CVD43. The prevalence of hypertension and CKD in the CVD group was also significantly higher than in the non-CVD group (hypertension: 82.35% vs. 65.36%, P < 0.0001; CKD: 52.01% vs. 31.36%, P < 0.0001). These results highlight the importance of controlling hypertension and delaying the progression of CKD to reduce the risk of CVD in T2DM patients44,45. Furthermore, the smoking rate in the CVD group was significantly higher than that in the non-CVD group (P < 0.0001), suggesting that smoking cessation is a crucial measure for preventing CVD in T2DM patients. Smoking can lead to endothelial injury, inflammatory responses, and thrombosis, thereby increasing the risk of CVD46.
In this study, we developed and validated several machine learning models to predict CVD risk in patients with T2DM. Among the six machine learning algorithms tested, XGBoost demonstrated the best overall performance with moderate discriminatory power (validation AUC = 0.72). While this performance level is not excellent, it represents a clinically meaningful ability to stratify risk in this complex patient population. It is worth noting that predicting cardiovascular outcomes in diabetic patients is inherently challenging due to the multifactorial nature of the disease and complex interactions between risk factors. Similar studies in this field have reported comparable AUC values ranging from 0.70 to 0.80 for machine learning models predicting cardiovascular outcomes in diabetic populations.
Our analysis revealed an important methodological consideration: the KNN model showed perfect discrimination (AUC = 1.00) on the training data but performed substantially worse on the validation data (AUC = 0.64), indicating severe overfitting. This highlights the critical importance of robust validation procedures when developing prediction models. XGBoost, while not achieving the highest training performance, demonstrated more consistent performance across training and testing datasets, suggesting better generalizability to new patients. This balance between model complexity and generalization ability is crucial for clinical applications where reliability across diverse patient populations is essential.
The moderate predictive performance of our best model (XGBoost) also underscores the need for continued refinement of cardiovascular risk prediction in T2DM patients. While our model offers improvements over traditional approaches, the AUC of 0.72 indicates that there remains significant opportunity for enhancing predictive accuracy through the incorporation of additional biomarkers, longitudinal data, or alternative modeling approaches in future research.XGBoost is a gradient boosting algorithm that iteratively trains multiple weak learners and combines them into a strong learner. Regularization techniques used during the training process can effectively prevent overfitting and improve the model’s generalization ability47. At the same time, XGBoost can efficiently handle missing data, identify nonlinear relationships, and perform feature selection48. Additionally, XGBoost supports parallel computing, which can accelerate the model’s training speed49.
Another highlight of this study is the development of an online prediction model (https://cvdshiny.shinyapps.io/shiny_cls2_1model_fastshap/) that is based on the key risk factors identified through SHAP analysis, making it convenient for healthcare professionals to use. By inputting patients’ clinical characteristic data, healthcare providers can quickly obtain predictions of CVD risk, thus providing a reference for clinical decision-making. This online prediction tool can help doctors identify high-risk patients and formulate individualized intervention measures, such as enhancing blood glucose control, managing blood pressure, regulating blood lipids, and improving lifestyle factors. Furthermore, this online prediction tool can also be used for patient education, helping patients understand their own CVD risk and actively participate in disease management.
This study has certain limitations. Firstly, The diagnosis of CVD in this study was based on self-reported outcomes, which introduces the risk of misclassification bias due to inaccurate or incomplete reporting. Such misclassification may occur if participants fail to report events, misremember medical history, or confuse CVD with other conditions. This bias can have several implications.①during model training, inaccurate labels increase label noise and may weaken the observed associations between predictors and the true presence of CVD, potentially limiting the model’s ability to fully capture the relationships between risk factors and outcomes.②the resulting predictive models may have decreased sensitivity or specificity, as outcome misclassification can obscure distinctions between cases and controls. Consequently, performance metrics such as AUC or net benefit may be underestimated.③if the pattern of misclassification in the study cohort differs from that in external populations, the model’s generalizability could be reduced. Therefore, future studies should prioritize validating these prediction models in external cohorts using objectively measured or adjudicated CVD outcomes.Secondly, the population of this study mainly consists of Americans, and the results may not be applicable to other countries or regions. Different populations may have variations in genetic background, lifestyle, dietary habits, etc., which could impact CVD risk. Therefore, future research needs to validate the findings of this study across different racial and cultural backgrounds.Thirdly, This study is based on cross-sectional data from NHANES, which limits our ability to draw direct causal inferences between predictors and CVD. Since both exposure and outcome information are collected simultaneously, it is impossible to establish temporal precedence; therefore, observed associations cannot confirm the directionality or causality of relationships between risk factors and CVD. Our findings should thus be interpreted as statistical associations within the population rather than evidence of cause-and-effect. Further longitudinal and prospective studies are needed to clarify these relationships.Fourthly, although the Boruta algorithm streamlines feature selection and supports model interpretability, it may omit variables that are clinically important but only weakly associated with cardiovascular disease in the current dataset. As Boruta prioritizes features based on statistical significance within the training data, nuanced or context-dependent risk factors with modest associations may be excluded. Despite our use of a relatively liberal significance threshold, some subtle yet clinically relevant predictors may still have been missed. This highlights the need to complement statistical feature selection with domain knowledge when interpreting model results.Therefore, absence of a variable from the final model does not imply a lack of clinical importance.In future work, combining statistical feature selection with clinical expertise or using hybrid selection methods may help capture a broader spectrum of relevant predictors.Fifthly, our study did not perform external validation of the XGBoost model using an independent dataset. As a result, the generalizability of our findings beyond the NHANES population remains uncertain. The model’s predictive performance and clinical utility in other populations or settings have not been established. Therefore, caution is warranted when applying these results outside the NHANES framework. Future research should prioritize external validation with independent and diverse cohorts to enhance the robustness and applicability of the developed model.In addition, this study included a limited number of clinical variables, which might have left out some essential risk factors. For instance, socioeconomic status, psychological factors, and family history may also influence CVD risk. Future research could incorporate more variables to enhance the predictive accuracy of the model.Finally, although deployment on Shinyapps.io enhances accessibility and usability, the platform’s performance and data security standards may not fully satisfy the stringent requirements of clinical practice. Before considering widespread adoption in routine healthcare settings, comprehensive evaluations of platform stability, privacy protection, and regulatory compliance are necessary. Future adaptations should explore migration to institutionally managed or HIPAA-compliant solutions to ensure patient data confidentiality and support sustainable clinical integration.
Future research can validate the findings of this study in larger, more diverse populations. Additionally, other machine learning algorithms or ensemble learning methods could be explored to further improve predictive accuracy. For example, deep learning algorithms could be used to learn more complex feature representations, or ensemble methods could be employed to combine predictions from multiple models. More importantly, prospective studies are required to verify the predictive capability and clinical utility of the model. Prospective research can track CVD events in patients, allowing for the assessment of the model’s predictive accuracy and clinical benefits. Simultaneously, future studies could evaluate the model’s application value in various clinical settings, such as primary healthcare institutions and community health service centers.
Conclusion
In summary, this study demonstrates that the predictive model based on Boruta feature selection and XGBoost machine learning shows promising results in assessing CVD risk in T2DM patients. This model provides a valuable tool for clinical practice, aiding early intervention and precision treatment to improve cardiovascular health in T2DM patients. By using this model, clinicians can better identify high-risk patients and develop personalized interventions, thereby reducing the risk of CVD and improving patient outcomes. Furthermore, this model can serve as a reference for public health decision-making, assisting in the formulation of more effective CVD prevention strategies.
Supplementary Information
Below is the link to the electronic supplementary material.
Author contributions
CMX and CYF wrote the main manuscript text. FCS and WLD prepared Table 1 and Supplementary Table 1.FCS and CMF prepared Figs. 1, 2, 3, 4, 5, 6 and 7. All authors reviewed the manuscript and approved the submitted version.
Data availability
The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Ethics statement
The study was conducted according to the Declaration of Helsinki. All information from the NHANES program is freely available to the public and therefore does not require approval from the Medical Ethics Committee Committee.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally to this work: Chunming Xu and Fachao Shi.
References
- 1.Wong, N. D. & Sattar, N. Cardiovascular risk in diabetes mellitus: epidemiology, assessment and prevention. Nat. Reviews Cardiol.20 (10), 685–695 (2023). [DOI] [PubMed] [Google Scholar]
- 2.Yun, J. S. & Ko, S. H. Current trends in epidemiology of cardiovascular disease and cardiovascular risk management in type 2 diabetes. Metab. Clin. Exp.123, 154838 (2021). [DOI] [PubMed] [Google Scholar]
- 3.Caussy, C., Aubin, A. & Loomba, R. The relationship between type 2 diabetes, NAFLD, and cardiovascular risk. Curr. Diab. Rep.21 (5), 15 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Onat, A., Dönmez, I., Karadeniz, Y., Cakır, H. & Kaya, A. Type-2 diabetes and coronary heart disease: common physiopathology, viewed from autoimmunity. Expert Rev. Cardiovasc. Ther.12 (6), 667–679 (2014). [DOI] [PubMed] [Google Scholar]
- 5.Zhou, Z. et al. Canagliflozin and stroke in type 2 diabetes mellitus. Stroke50 (2), 396–404 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Pandey, A., Khan, M. S., Patel, K. V., Bhatt, D. L. & Verma, S. Predicting and preventing heart failure in type 2 diabetes. Lancet Diabetes Endocrinol.11 (8), 607–624 (2023). [DOI] [PubMed] [Google Scholar]
- 7.Shi, W., Huang, X., Schooling, C. M. & Zhao, J. V. Red meat consumption, cardiovascular diseases, and diabetes: a systematic review and meta-analysis. Eur. Heart J.44 (28), 2626–2635 (2023). [DOI] [PubMed] [Google Scholar]
- 8.Athyros, V. G. et al. Diabetes and lipid metabolism. Horm. (Athens Greece). 17 (1), 61–67 (2018). [DOI] [PubMed] [Google Scholar]
- 9.Joseph, J. J. & Golden, S. H. Type 2 diabetes and cardiovascular disease: what next? Current opinion in endocrinology, diabetes. Obes.21 (2), 109–120 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zhang, P., Gao, J., Pu, C. & Zhang, Y. Apolipoprotein status in type 2 diabetes mellitus and its complications (Review). Mol. Med. Rep.16 (6), 9279–9286 (2017). [DOI] [PubMed] [Google Scholar]
- 11.Iadecola, C. & Parikh, N. S. Framingham general cardiovascular risk score and cognitive impairment: the power of foresight. J. Am. Coll. Cardiol.75 (20), 2535–2537 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Almeda-Valdes, P., Cuevas-Ramos, D., Mehta, R., Gomez-Perez, F. J. & Aguilar-Salinas, C. A. UKPDS risk engine, Decode and diabetes PHD models for the Estimation of cardiovascular risk in patients with diabetes. Curr. Diabetes. Rev.6 (1), 1–8 (2010). [DOI] [PubMed] [Google Scholar]
- 13.Fatima, F. et al. Comparison of Cvd risk assessment via Qrisk®2 vs Reynolds risk score in inflammatory joint diseases. J. Ayub Med. Coll. Abbottabad: JAMC. 34 (4), 843–848 (2022). [DOI] [PubMed] [Google Scholar]
- 14.Ridker, P. M., Buring, J. E., Rifai, N. & Cook, N. R. Development and validation of improved algorithms for the assessment of global cardiovascular risk in women: the Reynolds risk score. Jama297 (6), 611–619 (2007). [DOI] [PubMed] [Google Scholar]
- 15.Intensive blood-glucose. Control with sulphonylureas or insulin compared with conventional treatment and risk of complications in patients with type 2 diabetes (UKPDS 33). UK prospective diabetes study (UKPDS) group. Lancet (London England). 352 (9131), 837–853 (1998). [PubMed] [Google Scholar]
- 16.Hemann, B. A., Bimson, W. F. & Taylor, A. J. The Framingham risk score: an appraisal of its benefits and limitations. Am. Heart Hosp. J.5 (2), 91–96 (2007). [DOI] [PubMed] [Google Scholar]
- 17.Billingsley, H. E., Heiston, E. M., Bellissimo, M. P., Lavie, C. J. & Carbone, S. Nutritional aspects to cardiovascular diseases and type 2 diabetes mellitus. Curr. Cardiol. Rep.26 (3), 73–81 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Dong, J. et al. Machine learning model for early prediction of acute kidney injury (AKI) in pediatric critical care. Crit. Care. (London, England). 25 (1), 288 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Stark, G. F., Hart, G. R., Nartowt, B. J. & Deng, J. Predicting breast cancer risk using personal health data and machine learning models. PloS One. 14 (12), e0226765 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Alaa, A. M., Bolton, T., Di Angelantonio, E., Rudd, J. H. F. & van der Schaar, M. Cardiovascular disease risk prediction using automated machine learning: A prospective study of 423,604 UK biobank participants. PloS One. 14 (5), e0213653 (2019). [DOI] [PMC free article] [PubMed]
- 21.Thoresen, M. Logistisk regresjon-anvendt Og Anvendelig [Logistic regression-applied and applicable]. Tidsskr Nor Laegeforen.137(19) (2017). [DOI] [PubMed]
- 22.Mousavi, S. M. & Beroza, G. C. Deep-learning seismology. Science377 (6607), eabm4470 (2022). [DOI] [PubMed] [Google Scholar]
- 23.Geva, S. & Sitte, J. Adaptive nearest neighbor pattern classification. IEEE Trans. Neural Netw.2 (2), 318–322 (1991). [DOI] [PubMed] [Google Scholar]
- 24.Char, D. S. & Burgart, A. Machine-learning implementation in clinical anesthesia: opportunities and challenges. Anesth. Analg. 130 (6), 1709–1712 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Salmanpour, M. R., Shamsaei, M. & Rahmim, A. Feature selection and machine learning methods for optimal identification and prediction of subtypes in parkinson’s disease. Comput. Methods Programs Biomed.206, 106131 (2021). [DOI] [PubMed] [Google Scholar]
- 26.Pang, S. et al. The CMLA score: A novel tool for early prediction of renal replacement therapy in patients with cardiogenic shock. Curr. Probl. Cardiol.49 (12), 102870 (2024). [DOI] [PubMed] [Google Scholar]
- 27.Degenhardt, F., Seifert, S. & Szymczak, S. Evaluation of variable selection methods for random forests and omics data sets. Brief. Bioinform. 23 (2), bbab553 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Speiser, J. L., Miller, M. E., Tooze, J. & Ip, E. A comparison of random forest variable selection methods for classification prediction modeling. Expert Syst. Appl.161, 113524 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Saleem, J. et al. Application of the Boruta algorithm to assess the multidimensional determinants of malnutrition among children under five years living in Southern punjab, Pakistan. BMC Public. Health. 24 (1), 167 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Yu, B. et al. The non-high-density lipoprotein cholesterol to high-density lipoprotein cholesterol ratio (NHHR) as a predictor of all-cause and cardiovascular mortality in US adults with diabetes or prediabetes: NHANES 1999–2018. BMC Med.22 (1), 317 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Liao, J. et al. Association between estimated glucose disposal rate and cardiovascular diseases in patients with diabetes or prediabetes: a cross-sectional study. Cardiovasc. Diabetol.24 (1), 13 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Huang, Q. et al. Characterisation of cardiovascular disease (CVD) incidence and machine learning risk prediction in middle-aged and elderly populations: data from the China health and retirement longitudinal study (CHARLS). BMC Public. Health. 25 (1), 518 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Lundberg, S. M. & Lee, S-I. A unified approach to interpreting model predictions.In: proceedings of the 31st international conference on neural information processing systems. Long Beach, California, USA: Curran Associates Inc.; :4768–4777. (2017).
- 34.Lee, J. et al. Prediction of cardiovascular complication in patients with newly diagnosed type 2 diabetes using an xgboost/gru-ode-bayes-based machine-learning algorithm. Endocrinol. Metabolism (Seoul Korea). 39 (1), 176–185 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Das, S., Rahman, R. & Talukder, A. Determinants of developing cardiovascular disease risk with emphasis on type-2 diabetes and predictive modeling utilizing machine learning algorithms. Medicine103 (49), e40813 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Lv, H. et al. Machine Learning-driven models to predict prognostic outcomes in patients hospitalized with heart failure using electronic health records: retrospective study. J. Med. Internet. Res.23 (4), e24996 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Rodgers, J. L. et al. Cardiovascular risks associated with gender and aging. J Cardiovasc Dev Disease.6(2) (2019). [DOI] [PMC free article] [PubMed]
- 38.Agarwal, R. et al. Cardiovascular and kidney outcomes with finerenone in patients with type 2 diabetes and chronic kidney disease: the FIDELITY pooled analysis. Eur. Heart J.43 (6), 474–484 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Burnier, M. & Damianaki, A. Hypertension as cardiovascular risk factor in chronic kidney disease. Circul. Res.132 (8), 1050–1063 (2023). [DOI] [PubMed] [Google Scholar]
- 40.Mukamal, K. J. The effects of smoking and drinking on cardiovascular disease and risk factors. Alcohol Res. Health: J. Natl. Inst. Alcohol Abuse Alcoholism. 29 (3), 199–202 (2006). [PMC free article] [PubMed] [Google Scholar]
- 41.Ng, R., Sutradhar, R., Yao, Z., Wodchis, W. P. & Rosella, L. C. Smoking, drinking, diet and physical activity-modifiable lifestyle risk factors and their associations with age to first chronic disease. Int. J. Epidemiol.49 (1), 113–130 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Rosoff, D. B., Davey Smith, G., Mehta, N., Clarke, T. K. & Lohoff, F. W. Evaluating the relationship between alcohol consumption, tobacco use, and cardiovascular disease: A multivariable Mendelian randomization study. PLoS Med.17 (12), e1003410 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Villella, E. & Cho, J. S. Effect of aging on the vascular system plus monitoring and support. Surg. Clin. North. Am.95 (1), 37–51 (2015). [DOI] [PubMed] [Google Scholar]
- 44.Yen, F. S., Wei, J. C., Chiu, L. T., Hsu, C. C. & Hwu, C. M. Diabetes, hypertension, and cardiovascular disease development. J. Translational Med.20 (1), 9 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zoccali, C. et al. Cardiovascular complications in chronic kidney disease: a review from the European renal and cardiovascular medicine working group of the European renal association. Cardiovascular. Res.119 (11), 2017–2032 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Kondo, T., Nakano, Y., Adachi, S. & Murohara, T. Effects of tobacco smoking on cardiovascular disease. Circulation Journal: Official J. Japanese Circulation Soc.83 (10), 1980–1985 (2019). [DOI] [PubMed] [Google Scholar]
- 47.Borstelmann, S. M. Machine learning principles for radiology investigators. Acad. Radiol.27 (1), 13–25 (2020). [DOI] [PubMed] [Google Scholar]
- 48.Dong, J., Peng, L., Yang, X., Zhang, Z. & Zhang, P. XGBoost-based intelligence yield prediction and reaction factors analysis of amination reaction. J. Comput. Chem.43 (4), 289–302 (2022). [DOI] [PubMed] [Google Scholar]
- 49.Liu, H., Hu, D., Li, H. & Oguz, I. Medical image segmentation using deep learning. In Machine Learning for Brain Disorders (ed. Colliot O, Humana). 391-434 (2023). [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.






