Skip to main content
The International Journal of Angiology : Official Publication of the International College of Angiology, Inc logoLink to The International Journal of Angiology : Official Publication of the International College of Angiology, Inc
. 2025 Jul 14;35(1):50–63. doi: 10.1055/a-2644-4444

Machine Learning-Based Ensemble Predictive Model for Cardiovascular Disease Prevention

Neeraj Kumar 1,2, Rekha Agarwal 1,✉, Lokesh Kumar Sharma 3, Rashmi Vashisth 1
PMCID: PMC12909081  PMID: 41705083

Abstract

Cardiovascular diseases (CVDs) are a primary cause of death globally, with an increasing incidence in India. Machine learning (ML) has emerged as a viable approach for CVD prediction; however, dataset size and generalizability limit model robustness. This study aims to develop an enhanced ML prediction model for CVD detection using ensemble methods. Six datasets were considered, including 7,916 records with clinical parameters. The records were classified into Dataset 1 ( n  = 3,676) and Dataset 2 ( n  = 4,240) based on available features to establish a feature set. Dataset 1 underwent analysis utilizing two approaches: binary classification of target variable (0: absence of CVD, 1: presence of CVD) and multiclass classification of target variable (based on CVD severity). Likewise, Dataset 2 underwent further analysis using binary classification of target variable (risk of CVD in 10 years). Identical data preprocessing and exploratory data analysis steps were performed for both dataset groups. Subsequently, 18 ML algorithms were used to develop distinct models for both dataset groups, from which LazyPredict picked the top 10 performing models. The Voting Classifier was used to build an ensemble model to integrate the models and enhance predictive performance. In the case of Dataset 1, our framework was obtained an accuracy of 96.5% in binary classification and 85.5% in multiclass classification. Similarly, our framework achieved an accuracy of 81.18% for Dataset 2. Utilizing ensemble modeling and an extensive dataset, our framework surpasses traditional and existing ML models in predicting stability, mitigating bias and improving decision support in CVD detection.

Keywords: cardiovascular disease, machine learning, ensemble learning, early detection, risk stratification


Cardiovascular diseases (CVDs) remain to be the predominant cause of morbidity and death globally, imposing a substantial burden on health care systems and economies. Recent studies indicate that CVDs constitute almost one-third of worldwide mortality, with ischemic heart disease and stroke being the primary contributors. 1 2 The increasing incidence of CVDs is influenced by variables like aging demographics, metabolic disorders, and inactive lifestyles, requiring enhanced diagnostic and prognostic methodologies. The present situation in India is especially alarming owing to the elevated prevalence of premature CVD-related fatalities. The transition to noncommunicable diseases from infectious diseases has led to a rapid rise in cardiovascular risk factors, including dyslipidemia, hypertension, and diabetes. 3

Despite advancements in medical therapies and bioinformatics techniques, 4 limited access to specialist treatments and late-stage diagnoses and remain critical challenges, especially in rural and economically disadvantaged regions. 5 The increasing impact highlights the immediate need for early detection methods that may identify individuals at risk before serious effects arise. 1 Traditional diagnostic methods for CVD, including electrocardiography (ECG), echocardiography, and angiography, occasionally require specialized expertise and equipment, making them less accessible in resource-limited environments. 6 Furthermore, risk prediction instruments such as the Framingham Risk Score 7 and the atherosclerotic cardiovascular disease 8 risk estimator mostly rely on conventional clinical criteria, which may insufficiently capture the complexities of disease progression. 9 The variety of CVD presentations, together with lifestyle variations and individual genetic vulnerability complicates early detection. 10 Consequently, there is a growing interest in using bioinformatics and machine learning (ML) approaches 11 to enhance predictive accuracy and facilitate early intervention. 12

Recent advancements in ML have shown have shown promising results in CVD risk assessment, using algorithms such as decision trees, support vector machines (SVMs), and deep neural networks to analyze huge datasets. 5 9 These models integrate several clinical parameters, including demographic data, ECG signals, and biochemical markers, to identify patterns indicative of CVD. 10 Nonetheless, a significant difficulty in ML methodologies is the limited sample size of accessible datasets, which might raise apprehensions about model overfitting and bias and constrain the robustness and generalizability of prediction models. 5 Securing sufficient training data and rectifying class imbalances are essential for enhancing model performance and therapeutic relevance. 2 The present study seeks to create an improved prediction model for CVD risk stratification, considering the constraints of current diagnostic procedures and the capabilities of ML approaches.

Review of Literature

We examined several recent studies that have employed ML to predict and diagnose CVDs, underscoring its increasing significance for early identification and risk stratification of CVD. The studies examine numerous ML methods, such as deep learning, ensemble classifiers, and hybrid methodologies, to assess diverse clinical and demographic characteristics for precise CVD prediction.

For instance, Tiwari et al introduced an ensemble framework that integrates numerous ML methods, such as ExtraTrees Classifier, Random Forest, and XGBoost, to predict CVDs. 13 The dataset used included Hungarian, Cleveland, Long Beach VA, Switzerland, and Statlog datasets, emphasizing variables like Maximum Heart Rate Achieved, Serum Cholesterol, Chest Pain Type, and Fasting Blood Sugar. The suggested stacked ensemble classifier attained an accuracy of 92.34%, surpassing the performance of individual models. Similarly, El Bialy et al conducted research proposing a majority vote-based ensemble model that integrated classifiers like Naïve Bayes, Decision Trees, and SVMs for the diagnosis of heart disorders. 14 The model undergoes assessment on two benchmark datasets. The ensemble model attained classification accuracies over 90%, with the inclusion of Naïve Bayes resulting in an accuracy of 92%. Nevertheless, the studies lack comprehensive insights into computational complexity, which might impact scalability in practical applications.

Weng et al examined the use of photoplethysmography (PPG) data and deep learning models to forecast the probability of major adverse cardiovascular events (MACE). 15 The deep learning algorithm used age, sex, smoking status, and PPG data to predict the likelihood of MACE within a decade. The deep learning model attained a C-statistic of 71.1%, equivalent to conventional office-based risk assessments. The concept employs noninvasive PPG signals, obtainable via cellphones, therefore enabling extensive, cost-effective screening. Likewise, Miao et al create an ensemble ML model with an adaptive boosting approach for the detection of coronary heart disease. 16 The model is used on four datasets: Hungarian Institute of Cardiology, Switzerland University Hospital, Cleveland Clinic Foundation, and Long Beach Medical Center. The ensemble model attained accuracies of 89.12% for the Hungarian dataset, 96.72% for the Switzerland dataset, 80.14% for the Cleveland dataset, and 77.78% for the Long Beach dataset. Bashir et al introduced a clinical decision support system called MV5, which used a majority vote-based ensemble of classifiers, including Memory-Based Learner, Naïve Bayes, SVMs, and Decision Trees for predicting heart disease. 17 MV5 attained an accuracy of 88.52%, F-measure of 88.85%, and sensitivity of 86.96%, and specificity of 90.83%. Additionally, Atallah and Al-Mousa used datasets from the UCI ML Repository and implemented individual ML techniques such as K-Nearest Neighbor, Logistic Regression, Stochastic Gradient Descent (SGD), and Random Forest. 18 Subsequently, the individual models were integrated into an ensemble approach, whereby classification was determined by the models' majority vote (hard voting). The research revealed that the hard voting ensemble model demonstrated the best accuracy of 90% in predicting heart disease, surpassing individual classifiers. Nonetheless, the heterogeneity of performance across datasets indicates a need for more model refinement and validation across these studies.

These results emphasize the growing importance of ML in improving CVD prediction, while also exposing constraints related to dataset size, generalizability, and heterogeneity in model performance. Consequently, we aim to enhance model accuracy and mitigate sample size limitations by consolidating data from various additional sources. Furthermore, we want to assess the influence of critical clinical characteristics on predictive results, offering insights into the most pertinent indicators for early identification. Our methodology utilizes ensemble learning approaches to improve reliability and reduce bias, hence facilitating more effective CVD risk assessment and preventative interventions. 1 12

Materials and Methods

The workflow adopted for this research has been demonstrated in Fig. 1 .

Fig. 1.

Fig. 1

Overview of the CVD prediction workflow: data preprocessing, exploratory data analysis, model selection, ensemble learning, and model assessment. CVD, cardiovascular disease.

Data Sources and Description

The inclusion criteria for this study consisted of all publicly accessible datasets providing clinical feature information (as described below) relevant to CVD prediction. Datasets lacking comprehensive clinical characteristics, those not publicly accessible, and those concentrating on comorbidities related to CVDs were omitted from the study. These criteria ensured the uniformity and relevancy of the data included in model construction. Subsequently, we retrieved data from nine publicly available datasets: Cleveland HF (CLHF), 19 UCI HF (UCIHF), 20 Statlog HF, 21 Framingham HF, 7 Mendeley HF, 22 IEEE HF, 23 PhysioNet HF, 24 PTB ECG, 25 and MIT ECG, 26 as detailed in Table 1 .

Table 1. Overview of datasets utilized in the study including dataset description, case–control sample distribution, and associated features.

Serial number Name of dataset Description Number of records Cases Controls Number of features Names of features
1 Cleveland HF (CLHF) Contains 76 attributes; commonly used subset includes 14 features. UCI archive (tabular) 303 165 138 14 Age, sex, chest pain type, resting blood pressure, serum cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina, ST depression, slope of ST segment, number of major vessels, thalassemia, diagnosis of heart disease
2 UCI HF (UCIHF) Similar to Cleveland dataset; used for heart disease prediction studies. UCI archive (tabular) 303 165 138 14 Age, sex, chest pain type, resting blood pressure, serum cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina, ST depression, slope of ST segment, number of major vessels, thalassemia, diagnosis of heart disease
3 Statlog HF Derived from Cleveland dataset; used in the Statlog project (tabular) 270 120 150 13 Age, sex, chest pain type, resting blood pressure, serum cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina, ST depression, slope of ST segment, number of major vessels, thalassemia
4 Framingham HF Longitudinal study; data collected over multiple exams. BioLINCC (tabular) 4,434 1,924 2,510 15 Sex, age, education, current smoker, cigarettes per day, BP medication, prevalent stroke, prevalent hypertension, diabetes, total cholesterol, systolic BP, diastolic BP, BMI, heart rate, glucose
5 Mendeley HF Collected from an Indian hospital; useful for early-stage heart disease detection (tabular) 1,000 500 500 12 Age, sex, chest pain type, resting blood pressure, serum cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina, ST depression, slope of ST segment, diagnosis of heart disease
6 IEEE HF Comprehensive dataset; used for machine learning applications in heart disease prediction (tabular) 1,000 500 500 12 Age, sex, chest pain type pressure, serum cholesterol, fasting blood sugar, resting ECG, max heart rate, exercise-induced angina, ST depression, slope of ST segment, diagnosis of heart disease
7 PhysioNet HF Part of the PhysioNet/CinC Challenge 2016; focuses on heart sound classification (audio) 3,125 1,562 1,563 Varies Heart sound recordings
8 PTB ECG High-resolution 15-lead ECGs; include various heart diseases (signal) 549 368 181 15 ECG leads: I, II, III, V2, V3, V4, V5, V6, VX, VY, VZ
9 MIT ECG Contains 48 half-hour excerpts of two-channel ambulatory ECG recordings (signal) 48 25 23 2 ECG

Abbreviations: BMI, body mass index; BP, blood pressure; ECG, electrocardiography.

These datasets provide diverse patient records, incorporating various clinical and physiological features relevant to CVD prediction. Out of the nine datasets, three were excluded due to difference in data type (one focused on heart sound classification (auditory readings) and two focused on ECG recordings [Signal readings]). Subsequently, we divided the remaining six datasets into two groups based on the differences in the features available with the dataset and their final outcome. The first group (referred as Dataset 1 hereafter) consisted of the CLHF, 19 UCIHF, 20 Statlog HF, 21 Mendeley HF, 22 and IEEE HF 23 datasets having 14 common attributes while the remaining one dataset (Framingham HF 7 ) demonstrated 16 attributes (referred as Dataset 2 hereafter). The division was done since Dataset 1 aimed to predict the presence of CVD and its severity, and Dataset 2 aimed to predict whether CVD will occur or not in the next 10 years (hence, the difference in their focus of research and corresponding attributes considered). The distribution of data and resulting number of samples in both datasets have been summarized in Fig. 2 .

Fig. 2.

Fig. 2

Overview of the sample selection criteria and categorization considered for the study.

Dataset 1 comprised 13 critical clinical attributes that affect cardiovascular health, namely age, gender, chest pain type (CP), resting blood pressure (Trestbps), serum cholesterol (Chol), fasting blood sugar (FBS), resting electrocardiographic findings, maximum heart rate achieved (Thalach), exercise-induced angina (Exang), ST depression (Oldpeak), slope of the peak exercise ST segment (Slope), number of significant vessels (CA), and thallium stress test results (Thal). Categorical variables including CP, Resting ECG, Slope, CA, and Thal were transformed into numerical formats to enhance model training, as described below:

  • Age: recorded as a continuous variable in years.

  • Gender: encoded as a binary variable, with 1 representing male and 0 representing female.

  • Chest pain type (CP): categorized into four classes—0 for characteristic angina, 1 for atypical angina, 2 for nonanginal pain, and 3 for asymptomatic cases.

  • Resting blood pressure (Trestbps): measured in millimeters of mercury (mmHg).

  • Serum cholesterol (Chol): recorded in milligrams per deciliter (mg/dL).

  • Fasting blood sugar (FBS): encoded as a binary variable, with 1 indicating levels above 120 mg/dL and 0 indicating levels below the threshold.

  • Resting electrocardiographic findings: classified into three categories - 0 for normal results, 1 for ST-T wave abnormalities suggestive of ischemia, and 2 for left ventricular hypertrophy.

  • The maximum heart rate achieved (Thalach): recorded as a continuous variable.

  • Exercise-induced angina (Exang): encoded as 1 if present and 0 if absent.

  • ST depression induced by exercise relative to rest (Oldpeak): measured as a continuous variable.

  • The slope of the peak exercise ST segment (Slope): categorized into three values - 0 for an upsloping trend, 1 for a flat trend, and 2 for a downsloping trend.

  • Number of major vessels (CA) visualized through fluoroscopy: represented as an integer ranging from 0 to 3.

  • Thallium stress test results (Thal): classified into three categories—0 for normal blood flow, 1 for a fixed defect indicating permanent damage, and 2 for a reversible defect indicating temporary ischemia.

Similarly, Dataset 2 comprised 15 critical clinical attributes that affect cardiovascular health, namely age, gender, chol, education, currentSmoker, cigsPerDay, BPMeds, prevalentStroke, prevalentHyp, diabetes, sysBP, diaBP, BMI, heartrate, and glucose, as described below:

  • Age: recorded as a continuous variable in years.

  • Gender: encoded as a binary variable, with 1 representing male and 0 representing female.

  • chol: recorded in milligrams per deciliter (mg/dL).

  • Education: classified into four categories – 0 to 4, depending on the level of formal education attained by an individual.

  • currentSmoker: encoded as a binary variable, with 1 representing “currently smoking” and 0 representing “currently not smoking”.

  • cigsPerDay: measured as a continuous variable.

  • BPMeds: encoded as a binary variable, with 1 representing “currently taking medication for blood pressure” and 0 representing “currently not taking medication for blood pressure”.

  • prevalentStroke: encoded as a binary variable, with 1 indicating that “the person has had a stroke in the past” and 0 indicating that “the person has not had a stroke in the past”.

  • prevalentHyp: encoded as a binary variable, with 1 indicating that “the person has a history of hypertension”, and 0 indicating that “the person does not a history of hypertension”.

  • diabetes: encoded as a binary variable, with 1 indicating that “the person has diabetes”, and 0 indicating that “the person does not have diabetes”.

  • sysBP: systolic blood pressure (recorded as a continuous variable).

  • diaBP: diastolic blood pressure (recorded as a continuous variable).

  • BMI: body mass index (recorded as a continuous variable).

  • heartRate: heartbeats per minute (recorded as a continuous variable).

  • glucose: blood sugar level (mg/dL) (recorded as a continuous variable).

Exploratory Data Analysis and Data Preprocessing

EDA was used to evaluate feature distributions, identify patterns, and examine correlations between clinical data and CVD severity. Visualizations using the matplotlib 27 and seaborn 28 libraries illustrated relationships among variables. Missing values were addressed using Multivariate Imputation by Chained Equations (MICE). 29 This technique imputed missing data points based on existing correlations among given variables, yielding a more precise imputation compared to the conventional techniques like mean or median replacements. MICE uses patterns within the dataset to retain its integrity and prevent the loss of important patient information. This step was essential for maintaining uniformity across all attributes while making sure that the prediction models were trained on reliable data.

Dataset 1 was subjected to modeling using two distinct approaches. The first method encoded the target variable (Inference) as binary (0: Absence of disease, 1: Presence of disease), whereas the second method classified the target variable into multiple classes based on CVD severity (five levels of CVD severity: 0 for absence of disease, 1 for low severity, 2 for moderate severity, 3 for significant severity, and 4 for extreme severity). Similarly, Dataset 2 was utilized for modeling by encoding the target variable (Inference) as binary (0: No risk of CVD in next 10 years, 1: Potential risk of CVD in next 10 years). This enabled a comparative examination of model performance in Dataset 1 (binary and multiclass classification) and Dataset 2 (binary), with all preprocessing and modeling techniques being identical for both.

Model Selection and Evaluation

An extensive set of 18 individual ML models was utilized for both dataset groups, including LGBMClassifier, XGBClassifier, CatBoostClassifier, ExtraTreesClassifier, VotingClassifier, RandomForestClassifier, BaggingClassifier, DecisionTreeClassifier, GradientBoostingClassifier, QuadraticDiscriminantAnalysis, AdaBoostClassifier, LinearDiscriminantAnalysis, LogisticRegression, GaussianNB, MLPClassifier (DNN), MLPClassifier (ANN), KNeighborsClassifier, and SVC . Subsequently, LazyPredict 30 was used to rigorously evaluate several ML algorithms to rank and identify the best 10 models according to their predictive performance. Each model was assessed for accuracy, precision, recall, and other critical performance parameters to ascertain its efficacy in diagnosing the occurrence/severity of CVD. The selection approach guaranteed that only the top-performing classifiers were used for further analysis and ensemble modeling.

Addressing Class Imbalance

Due to the unequal distribution of target variable values in both dataset groups, class imbalance was mitigated by the use of the Synthetic Minority Over-sampling Technique (SMOTE). 31 This approach produced synthetic examples for underrepresented severity levels, ensuring that the model was trained on an adequately balanced dataset. To prevent data leakage, SMOTE was included in each iteration of Stratified K-Fold Cross-Validation, maintaining uniformity across training and validation sets. By handling class imbalance, the model eliminated bias towards majority classes so that predictions were representative across all severity levels.

Ensemble Learning Utilizing a Voting Classifier

A soft voting classifier was developed (using scikit-learn , 32 xgboost , 33 and lightgbm 34 libraries) to further increase prediction accuracy by combining the predictions of the best-performing models (based on accuracy) for both dataset groups. The ensemble method integrated many classifiers for their individual strengths, enabling enhanced robustness and generalizability. The soft voting technique allocated probabilistic weights to each model's predictions, increasing the balance between accuracy and recall.

Assessment of Models and Performance Indicators

The developed models were evaluated using several performance metrics, including accuracy, precision, recall, F1-score, and receiver-operating characteristic (ROC)-area under the curve (AUC). These metrics provide a comprehensive evaluation of the model's ability to distinguish between various severity levels of CVD. Stratified K-fold cross-validation improved evaluation robustness and avoided the risk of overfitting or underfitting by dividing the dataset into balanced subgroups. The evaluation validated the performance of the ensemble learning approach, demonstrating that the model accurately classified the severity of CVD across varied patient data.

Results

Data Description and Preprocessing

A total of 7,916 records were obtained based on inclusion and exclusion criteria and divided into Dataset 1 ( n  = 3,676) and Dataset 2 ( n  = 4,240), as outlined in Supplementary File 1 (available in online version).

Dataset 1 exhibited an age range of 20 to 80 years, with a mean age of 52.56 years. Trestbps ranged from 0 to 200, with a mean of 137.46; chol varied from 0 to 603, averaging 241.23; thalach spanned from 60 to 202, with an average of 142.35; and oldpeak fluctuated between −2.6 and 6.2, with a mean of 1.42 ( Fig. 3 ). Similarly, in Dataset 2, the age varied from 32 to 70 years, with a mean of 49.58 years. Chol ranged from 107 to 696, with a mean of 236.70; cigsPerDay varied from 0 to 70, averaging 9.005; sysBP ranged from 83.5 to 295, with a mean of 132.35; diaBP varied from 48 to 142.5, averaging 82.89; BMI spanned from 15.54 to 56.8, with an average of 25.008; heart rate varied from 44 to 143, with a mean of 75.8; and glucose levels ranged from 40 to 394, averaging 81.96 ( Fig. 4 ).

Fig. 3.

Fig. 3

Distribution of numerical features with corresponding correlation matrix illustrating feature interrelationships for Dataset 1.

Fig. 4.

Fig. 4

Distribution of numerical features with corresponding correlation matrix illustrating feature interrelationships for Dataset 2.

Subsequently, missing values were ascertained and imputed. The correlations among the features were calculated. Subsequently, missing values were imputed using MICE, and the preprocessed dataset was subjected to model selection using LazyPredict .

Model Selection Using LazyPredict

LazyPredict analysis revealed that XGBClassifier (86%) and LGBMClassifier (85%) had the maximum accuracy, indicating that gradient-boosting models were particularly appropriate for the dataset. The RandomForestClassifier (84%) and ExtraTreesClassifier (84%) exhibited robust performance, demonstrating the efficacy of ensemble learning techniques. Conventional models such as Logistic Regression (75%) and Linear Discriminant Analysis (75%) exhibited reasonable performance, whereas Naïve Bayes and AdaBoost (74%) had somewhat poorer accuracy. The lowest accuracies were recorded in the PassiveAggressiveClassifier (61%) and DummyClassifier (45%), indicating that the dataset needed advanced learning algorithms for optimum prediction. The accuracy and F1 scores for all models considered by LazyPredict are demonstrated in Fig. 5 . Consequently, the top 10 models chosen for their predictive performance were LGBMClassifier , 35 XGBClassifier , 36 RandomForestClassifier , 37 Support Vector Classification (SVC) , 38 ExtraTreesClassifier , 39 DecisionTreeClassifier , 40 QuadraticDiscriminantAnalysis , 41 ExtraTreeClassifier , 42 GaussianNB , 43 and AdaBoostClassifier . 44

Fig. 5.

Fig. 5

The accuracy and F1 scores for all models considered by LazyPredict .

Model Training and Evaluation

The outcomes were evaluated distinctly for Dataset 1 (multiclass classification [CVD severity levels] and binary classification [CVD vs No CVD]) and Dataset 2 (risk of CVD in 10 years). Feature scaling and ensemble methods were further examined to improve model performance. Prior to model training, an examination of class distribution indicated a substantial imbalance in target variable instances, potentially resulting in biased model predictions. To rectify this, SMOTE was used to artificially create synthetic samples in the minority class, ensuring balanced representation. This step significantly enhanced recall scores, especially in models susceptible to class imbalances, like Random Forest and Neural Networks.

Performance in Dataset 1

Multiclass Classification

XGBoost (86.1%), LightGBM (86.2%), and CatBoost (85.9%) were identified as the leading models for predicting CVD severity levels, exhibiting high accuracy and AUC-ROC values of about 0.84 to 0.85 ( Fig. 6A ). Conventional models like Logistic Regression (66.3%) and Linear Discriminant Analysis (66.5%) exhibited subpar performance owing to the dataset's intricate, nonlinear nature. The use of feature scaling significantly improved models such as SVC (from 42.7 to 79.2%) and KNN (from 57.4 to 72.4%), which exhibit considerable sensitivity to variations in feature values. Furthermore, the ensemble model (VotingClassifier), integrating models determined from LazyPredict , exhibited impressive performance, with an accuracy of 85.5% and an ROC value of 0.91 ( Fig. 6B ). The ensemble method efficiently used the advantages of individual models, decreasing variation and enhancing stability in multiclass classification.

Fig. 6.

Fig. 6

ROC-curves for ( A ) multiclass multimodel classification; ( B ) multiclass ensemble model classification; ( C ) binary multimodel classification; and ( D ) binary ensemble model classification.

Binary Classification

In the binary classification framework (0: No CVD, 1: CVD), models attained markedly superior accuracy across all classifiers as compared with multiclass classification models. XGBoost (97.0%), LightGBM (96.9%), and CatBoost (96.5%) exhibited exceptional predictive performance, with AUC-ROC values over 0.99 ( Fig. 6C ), hence assuring nearly ideal class separation. Conventional models, such as Logistic Regression (84.8%) and Gaussian Naïve Bayes (84.5%), had significantly superior performance in the binary classification context compared with multiclass classification. Furthermore, the ensemble model significantly improved classification robustness, attaining 96.5% accuracy and exhibiting its capacity to generalize across diverse patient characteristics. SMOTE significantly contributed to this enhancement, especially in recall scores, guaranteeing that forecasts for the minority class were not neglected, which is vital for clinical dependability. The ensemble model enhanced classification robustness ( Fig. 6D ), attaining 96.5% accuracy and exhibiting its capacity to generalize across various patient profiles. The integration of multiple models guaranteed stability, minimizing the risk of overfitting while preserving elevated sensitivity and specificity.

Performance in Dataset 2

The performance evaluation for Dataset 2 ( Fig. 7 ), intended for predicting CVD risk over a decade, showed strong classification accuracy across numerous models. The CatBoostClassifier attained the greatest accuracy at 81.18%, closely followed by the VotingClassifier at 81.08% and the LGBMClassifier at 81.06%, emphasizing the efficacy of gradient boosting methods for long-term risk prediction. The RandomForestClassifier and ExtraTreesClassifier demonstrated commendable performance, with accuracy ratings beyond 80%, suggesting that ensemble approaches are especially appropriate for this dataset. Notably, GaussianNB and Quadratic Discriminant Analysis demonstrated high recall values (61.5 and 61.2%, respectively), indicating its potential effectiveness in reducing false negatives, which is essential for CVD risk evaluation. Nonetheless, models like Logistic Regression and SVC exhibited comparatively worse performance, indicating the dataset's complexity and the need for more sophisticated algorithms. The findings suggest that ensemble learning methods, especially boosting-based classifiers, provide the most dependable predictions for long-term CVD risk assessment, serving as a valuable instrument for early detection and intervention. Performance metrics of all models (including individual models and ensemble voting classifier) for both dataset groups have been summarized in Table 2 .

Fig. 7.

Fig. 7

ROC curves (Dataset 2) for ( A ) multimodel classification; ( B ) ensemble model classification.

Table 2. Performance metrics of all models (including individual models and ensemble voting classifier) for multiclass and binary classification approaches.
Dataset 1
Model performance metrics—multiclass classification
Model Accuracy Precision F1 score Recall
LGBMClassifier 0.8615 0.5535 0.5525 0.5571
XGBClassifier 0.8610 0.5460 0.5463 0.5520
CatBoostClassifier 0.8591 0.5516 0.5470 0.5485
ExtraTreesClassifier 0.8558 0.5910 0.5425 0.5226
VotingClassifier 0.8550 0.5546 0.5514 0.5551
RandomForestClassifier 0.8531 0.5664 0.5622 0.5641
BaggingClassifier 0.8460 0.5580 0.5578 0.5640
DecisionTreeClassifier 0.8221 0.5370 0.5268 0.5228
GradientBoostingClassifier 0.8123 0.5222 0.5330 0.5554
QuadraticDiscriminantAnalysis 0.7247 0.4399 0.4503 0.5004
AdaBoostClassifier 0.7111 0.4409 0.4465 0.4813
LinearDiscriminantAnalysis 0.6654 0.4118 0.4142 0.5052
LogisticRegression 0.6627 0.4139 0.4171 0.5129
GaussianNB 0.6505 0.4109 0.4001 0.4984
MLPClassifier (DNN) 0.6093 0.4036 0.3840 0.4550
MLPClassifier (ANN) 0.5963 0.4140 0.3687 0.4559
KNeighborsClassifier 0.5740 0.3427 0.3235 0.3554
SVC 0.4274 0.3373 0.2718 0.3878
Model performance metrics—binary classification
Model Accuracy Precision F1 score Recall
XGBClassifier 0.9695 0.9696 0.9693 0.9692
LGBMClassifier 0.9690 0.9690 0.9688 0.9688
VotingClassifier 0.9655 0.9656 0.9652 0.9650
CatBoostClassifier 0.9652 0.9653 0.9650 0.9648
RandomForestClassifier 0.9641 0.9640 0.9639 0.9639
ExtraTreesClassifier 0.9608 0.9609 0.9606 0.9604
BaggingClassifier 0.9565 0.9562 0.9563 0.9567
DecisionTreeClassifier 0.9369 0.9366 0.9365 0.9365
GradientBoostingClassifier 0.9268 0.9264 0.9265 0.9268
AdaBoostClassifier 0.8887 0.8884 0.8883 0.8890
QuadraticDiscriminantAnalysis 0.8588 0.8582 0.8580 0.8582
LinearDiscriminantAnalysis 0.8509 0.8504 0.8503 0.8510
MLPClassifier 0.8504 0.8552 0.8483 0.8476
LogisticRegression 0.8479 0.8474 0.8473 0.8480
GaussianNB 0.8455 0.8451 0.8447 0.8452
MLPClassifier 0.8390 0.8445 0.8379 0.8401
KNeighborsClassifier 0.8085 0.8077 0.8074 0.8074
SVC 0.6956 0.7070 0.6949 0.7025
Dataset 2—model performance metrics
Model Accuracy Precision F1 score Recall
CatBoostClassifier 0.811792 0.580317 0.55806 0.5525
VotingClassifier 0.810849 0.588248 0.56755 0.5609
LGBMClassifier 0.810613 0.581169 0.56087 0.555
RandomForestClassifier 0.809198 0.595936 0.57939 0.572
ExtraTreesClassifier 0.806368 0.588589 0.57309 0.5659
XGBClassifier 0.80283 0.568283 0.55342 0.5491
BaggingClassifier 0.798585 0.578036 0.5671 0.5625
GradientBoostingClassifier 0.794575 0.582893 0.5776 0.5749
GaussianNB 0.772406 0.595006 0.60223 0.6159
QuadraticDiscriminantAnalysis 0.766981 0.589975 0.59649 0.6122
AdaBoostClassifier 0.765094 0.582348 0.58812 0.5996
DecisionTreeClassifier 0.723821 0.537253 0.53852 0.5497
MLPClassifier (DNN) 0.685613 0.561698 0.54906 0.5921
LogisticRegression 0.676179 0.57642 0.56702 0.6344
LinearDiscriminantAnalysis 0.672642 0.575532 0.56497 0.6336
MLPClassifier (ANN) 0.644811 0.568681 0.51543 0.5863
KNeighborsClassifier 0.643396 0.545629 0.52809 0.5807
SVC 0.638443 0.582589 0.55551 0.6555

Discussion

The presented artificial intelligence (AI)-enabled approach to predict CVD demonstrated strong performance in both binary and multiclass classification models. Employing ensemble approach, namely the VotingClassifier , significantly enhanced prediction accuracy and model generalizability relative to individual classifiers. This improvement is especially important in clinical settings because prompt and precise identification of CVD may significantly affect patient outcomes. Existing studies (as discussed above) have examined numerous ML methodologies for CVD prediction, including deep learning frameworks, majority voting ensemble techniques, and AI-based risk assessment models. Although these works provide encouraging outcomes, they often need substantial computing resources, intricate feature engineering, or encounter challenges related to class imbalance. Moreover, several previous research have insufficiently addressed the challenges linked to smaller datasets, resulting in diminished generalizability in real-world clinical applications. Our methodology efficiently overcomes these restrictions by including several classifiers, using boosting-based models such as XGBoost and LightGBM, and exploiting SMOTE to alleviate class imbalance. The contrasting findings of Dataset 1 and Dataset 2 reveal distinct performance patterns across the evaluated models, reflecting discrepancies in prediction objectives and data features. Dataset 1 concentrated on multiclass and binary classification, attaining a prediction performance of 85.5% for multiclass classification and 96.5% accuracy for binary classification. The AUC-ROC values of 0.99 for binary classification and 0.85 for multiclass classification highlight the effectiveness of our strategy in distinguishing between various severity levels of CVD. The ensemble models, especially the VotingClassifier, demonstrated improved generalizability by effectively utilizing the strengths of individual classifiers to increase stability and accuracy. This enhanced performance is crucial for clinical applications, encouraging precise risk assessment and personalized treatment strategies. Conversely, Dataset 2 focused on forecasting the ten-year risk of CVD, shifting the classification task from immediate diagnosis to long-term prognosis. Performance trends varied, with models such as GaussianNB (AUC-ROC: 0.690) and ExtraTreesClassifier (AUC-ROC: 0.683) depicting proficiency in identifying long-term risk patterns. In comparison to Dataset 1, Dataset 2 exhibited better recall and sensitivity in models like GradientBoostingClassifier and RandomForestClassifier, indicating superior capabilities to identify at-risk individuals over an extended period. However, the overall accuracy in Dataset 2 was observed to diminish, perhaps owing to the complexities of long-term risk prediction, whereby environmental factors and external lifestyle affect outcomes beyond the available data. Nevertheless, the adaptability of ensemble models was demonstrated in both datasets, affirming their dependability for diverse prediction purposes. Our findings highlight the need to ascertain appropriate modeling methods based on the specific objectives of CVD prediction, whether for short-term diagnosis or long-term risk assessment. Despite these advantages, some limitations persist. Our work relies on publicly available data, which may introduce inaccuracies owing to differences in data collection techniques and demographic representation. Secondly, even though SMOTE efficiently addresses class imbalance, the production of synthetic data may not adequately encapsulate the complexities of real-life clinical scenarios. Finally, although ensemble techniques increase overall performance, they also elevate computing complexity, restricting real-time implementation in clinical environments. Future work may concentrate on incorporating additional clinical factors and real-time data to augment model efficacy. Furthermore, expanding the research to include bigger, multi-institutional datasets with varied demographic representation would enhance the validation of the generalizability of our results. The integration of time-series data with longitudinal patient records may enhance prediction skills, enabling proactive solutions for high-risk people. Our technique signifies a substantial advancement in using ML for precise, scalable, and interpretable CVD risk assessment.

Conclusion

Our research effectively developed an advanced ML model for CVD prediction and risk stratification, with superior accuracy and AUC-ROC scores in both binary and multiclass classification. By integrating data from several sources and using ensemble techniques, the model exhibited strong predictive efficacy. The findings emphasize the significant potential of ML in enhancing precision medicine for CVD, with future efforts focused on increasing dataset diversification and validating the model for wider clinical use.

Acknowledgments

We would like to acknowledge Amity University, Uttar Pradesh for providing us with the opportunity to perform our research. We would also like to thank the Indian Council of Medical Research for providing the essential guidance.

Conflict of Interest None declared.

Authors' Contributions

N.K., R.A., and R.V. were involved in the conceptualization of the study. Methodology was developed by N.K., R.A., and L.K.S. Formal analysis and investigation were carried out by N.K., L.K.S., and R.V. N.K. prepared the original draft of the manuscript, while R.A. and R.V. contributed to reviewing and editing the draft. N.K. and R.A. provided the necessary resources, and R.A. supervised the overall preparation of the manuscript.

Data Availability Statement

The data used for this study are detailed in the publication, accompanied by references and supplementary material.

Supplementary Material

10-1055-a-2644-4444-s20250019.xlsx (544.3KB, xlsx)

Supplementary Material

Supplementary Material

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

10-1055-a-2644-4444-s20250019.xlsx (544.3KB, xlsx)

Supplementary Material

Supplementary Material

Data Availability Statement

The data used for this study are detailed in the publication, accompanied by references and supplementary material.


Articles from The International Journal of Angiology : Official Publication of the International College of Angiology, Inc are provided here courtesy of Thieme Medical Publishers

RESOURCES