Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Mar 6;16:16206. doi: 10.1038/s41598-026-41042-z

A preprocessing-enhanced stacking classifier for generalized cardiovascular disease detection across diverse datasets

Adeel Ashraf 1, Adven Masih 1,✉, Aysha Saddiqa 1, Jabar Mahmood 2, Aitzaz Ali 3,✉, Mohamed Shabbir Hamza Abdulnabi 4, Daniel Musafiri Balungu 5, Xu Ying 6
PMCID: PMC13201779  PMID: 41792193

Abstract

Cardiovascular diseases (CVDs) remain a major global health challenge, requiring early-detection models that are both accurate and generalizable across diverse data settings. This study introduces a preprocessing-enhanced stacking ensemble for binary CVD prediction, explicitly designed to improve robustness under heterogeneous feature distributions. The preprocessing pipeline incorporates feature transformation, derived attribute construction, encoding, and K-modes clustering, all applied after strict train–test separation to preserve evaluation validity. The proposed stacking architecture integrates three complementary tree-based base learners i.e., Random Forest, Decision Tree, and Extra Trees with Logistic Regression as a meta-learner to aggregate out-of-fold predictions. The framework was evaluated on three heterogeneous datasets. The ensemble achieved accuracies of 93.26%, 72%, and 99% on Datasets I, II, and III, respectively, with performance stability confirmed using 95% confidence intervals across five random seeds. Statistical significance analysis using McNemar’s test demonstrated that the proposed model significantly outperformed several strong baselines (p < 0.05), including Random Forest, Logistic Regression, and XGBoost on Dataset I; CNN on Dataset II; and Decision Tree and Logistic Regression on Dataset III. These results indicate that the proposed framework maintains consistent performance across varying data modalities, noise levels, and feature structures.

Keywords: CVDs, ML, Ensemble learning, Stacking, Clustering, preprocessing

Subject terms: Computational biology and bioinformatics, Diseases, Engineering, Health care, Mathematics and computing

Introduction

Cardiovascular diseases remain the leading cause of mortality worldwide, accounting for nearly one-third of global deaths and placing a substantial burden on healthcare systems, economies, and patient quality of life1–4. Early diagnosis plays a critical role in reducing morbidity and mortality; however, traditional diagnostic approaches such as electrocardiography and echocardiography often rely on subjective clinical interpretation and may fail to detect early-stage or asymptomatic conditions5. These limitations underscore the need for objective, data-driven tools capable of identifying subtle clinical patterns that precede overt disease.

Machine learning (ML) has emerged as a powerful complement to conventional diagnostics by enabling the automatic extraction of complex relationships from clinical data6. Despite significant progress, a major challenge “generalizability” persists. Many existing CVD prediction studies achieve high accuracy but rely on a single dataset, limited preprocessing pipelines, or narrowly defined patient cohorts7. As a result, their performance often degrades when applied to new populations or datasets that differ in demographics, feature distributions, or measurement standards. Recent biomedical ML research has repeatedly highlighted this limitation, emphasizing the need for robust, reproducible frameworks that can maintain performance across heterogeneous datasets8.

Motivated by these gaps, this study proposes a preprocessing-enhanced stacking classifier designed to improve generalizable CVD detection across diverse data sources. Unlike prior works that primarily focus on model architecture, our approach integrates a comprehensive preprocessing pipeline comprising transformation, categorization, attribute engineering, label encoding, and clustering to standardize feature representations before model training. We benchmark the proposed stacking model against traditional ML methods (Decision Tree, Logistic Regression, Random Forest), ensemble techniques (XGBoost), and deep learning architectures (Feedforward Neural Network, Convolutional Neural Network and TabNet) to assess its robustness.

By evaluating performance across three distinct datasets, each with different clinical variables and patient characteristics, the study provides a broader and more realistic assessment of model generalizability compared with prior research. Moreover, our approach directly addresses current limitations in biomedical ML by demonstrating how enhanced preprocessing and ensemble learning can strengthen predictive consistency across heterogeneous clinical environments. Through this contribution, the study advances ML-based CVD diagnostics toward more reliable, scalable, and clinically applicable solutions.

Related studies

Cardiovascular diseases remain a major global health concern, and the integration of artificial intelligence (AI), particularly ML, has emerged as a promising approach to improving prediction and diagnosis. Numerous studies have explored various ML and deep learning (DL) techniques for CVD prediction, often reporting high accuracy on specific datasets. However, despite these advances, many existing studies face significant limitations in terms of generalizability. Most works rely on a single dataset often the Cleveland dataset or use narrowly focused, homogenous cohorts, making it difficult to determine how well the proposed models perform across diverse populations, feature distributions, or clinical settings. Furthermore, many studies apply preprocessing steps post-splitting or without explicit leakage controls, potentially inflating performance metrics and limiting real-world applicability. These gaps highlight the need for methodologies that rigorously evaluate performance across varied datasets and ensure robustness against data leakage, thereby strengthening the rationale for the present study.

In the domain of ML, several studies have employed a wide range of algorithms to identify patterns in CVD data and predict disease outcomes. For instance, Garavand et al.9 applied artificial neural networks (ANN) to a dataset of 303 records with 13 attributes, achieving an accuracy of 88%. Ahmad et al.10 evaluated SVM, Random Forest (RF), and XGBoost, with SVM achieving 87.91% accuracy on the Cleveland dataset. Akkaya et al.11 assessed eight ML classifiers on the same dataset and reported that K-nearest neighbor (KNN) performed best at 85.6%. Tougui et al.12 similarly reported Random Forest as the top classifier with 87.64% accuracy, while another voting ensemble of Naïve Bayes and logistic regression achieved 87.41%13. Subanya and Rajalaxmi14 combined SVM with the Artificial Bee Colony (ABC) algorithm for feature optimization, yielding 86.76% accuracy. Mokeddem et al.15 employed genetic algorithms with Naïve Bayes and SVM, reporting accuracies of 85.50% and 83.82%, respectively. Other studies, including Khanna et al.16 and Kumar et al.17, compared classical models such as logistic regression and decision tree C4.5, reporting accuracies of 84.80% and 83.40%. More recently, Javid et al.18 further demonstrated the effectiveness of gradient boosting methods, particularly XGBoost, for CVD risk prediction, reinforcing the value of tree-based ensembles as strong baselines. While these works demonstrate the predictive potential of ML, the majority evaluate models only on the Cleveland dataset or similarly structured datasets. This narrow evaluation scope makes it challenging to determine whether the models would perform reliably on data with differing feature spaces, demographic compositions, or noise patterns. As such, their reported performance may not translate effectively into real-world clinical environments.

On the DL front, CNNs and related architectures have shown strong performance in CVD prediction across several datasets. Singhal et al.19 achieved 95% accuracy on the Cleveland dataset using a CNN with three convolutional layers. Dutta et al.20 obtained an accuracy of 81.78% on a National Health and Nutrition Examination Survey based dataset using a multi-layer CNN. A deep belief network (DBN) classifier applied to ECG image data achieved 95% accuracy21. Other CNN-based studies reported 76.9% accuracy in imbalanced settings13 and 90.088% with feature augmentation15. A hybrid CNN-LSTM model reached 73.52% accuracy on a CVD dataset22. The theoretical underpinnings of recurrent architectures for temporal modeling are supported by mathematical frameworks by23, who demonstrated the use of special series with recurrently computed coefficients for solving nonlinear evolution equations. Despite these promising results, most DL studies share similar limitations: reliance on single-site datasets, absence of cross-dataset evaluation, and limited attention to domain shifts that commonly occur in real-world data. Saranya, K et al.24 proposed “DenseNet-ABiLSTM,” a hybrid deep learning model for detecting cardiac arrhythmias—a key cardiovascular condition from PPG signals. Their approach combines DenseNet for feature extraction with an Attention-based Bidirectional LSTM to capture temporal patterns in cardiac data, enabling accurate multiclass arrhythmia classification for potential use in continuous cardiovascular monitoring.

Hybrid ML–DL models have also been developed to enhance prediction performance. Mehmood et al.25 proposed CardioHelp, a hybrid DL model that achieved 97% accuracy on a local dataset using convolutional layers. Tarawneh and Embarak26 used a combination of Naïve Bayes, SVM, decision tree, neural network, and KNN, reporting improved performance through combined modeling. Bhavekar and Goswami27 presented an RNN–LSTM hybrid model for cardiac disease classification, showing strong performance through optimized feature extraction. Subhadra and Vikas28 evaluated decision trees, logistic regression, gradient boosting, and MLP models, reporting 93.39% accuracy on the Cleveland dataset. Recently, Sadr et al. (2024) proposed a holistic ensemble integrating ML (KNN, XGBoost) with DL (CNN-LSTM) models evaluated on two public datasets and one local dataset, achieving accuracies of 80.25% and 95.85%29.

Although these hybrid models attempt to improve generalizability by incorporating multiple algorithms, they still largely depend on limited, homogenous datasets and do not systematically address cross-dataset structural differences, such as varying feature spaces or inconsistent clinical variable availability. More importantly, many studies do not explicitly test for or mitigate data leakage, which can lead to overly optimistic performance claims.

Collectively, existing research demonstrates rapid methodological progress but also underscores persistent gaps related to dataset diversity, leakage control, and real-world generalization. These limitations reinforce the need for the present study, which aims to evaluate a preprocessing-enhanced stacking classifier across multiple heterogeneous datasets while ensuring rigorous methodological transparency.

Methodology

This research aims to advance the early detection of CVDs by proposing a preprocessing-enhanced stacking classifier that is explicitly designed to provide a robust and generalizable predictive model across multiple datasets. As CVDs remain a leading cause of global mortality, the development of reliable and scalable diagnostic support systems is critical for enabling timely clinical intervention and improving patient outcomes. Traditional diagnostic approaches are often time-consuming and heavily dependent on clinicians’ expertise for interpreting heterogeneous medical data, which limits their efficiency and scalability in large-scale healthcare settings.

To address these challenges, this study adopts an ensemble machine learning framework supported by a carefully structured and leakage-free preprocessing pipeline. Following an initial data cleaning stage, the dataset is partitioned into training and testing subsets prior to any feature engineering, ensuring strict isolation of test data. All subsequent preprocessing operations such as data transformation, categorization, attribute combining, label encoding, normalization, and clustering are learned exclusively from the training data and then consistently applied to the test set using fixed parameters. This design enables a fair evaluation of model generalization. The complete systematic methodology of the study is illustrated in Fig. 1.

Fig. 1.

Fig. 1

Detailed methodology of the study.

Data collection

This study utilizes three datasets i.e., two public and one proprietary local, structured as two-dimensional collections with categorical and numerical variables, featuring missing values, redundant entries, and imbalanced classes across two classes. Each includes patient data, medical examination results, and subjective patient information, enabling robust model training and validation. Table 1 compares the datasets in detail.

Table 1.

Comparative overview of structural characteristics across Dataset I, II, and III, highlighting differences in scale, accessibility, and data quality.

Property Dataset III Dataset II Dataset I
Origin Proprietary (local) Publicly sourced Publicly sourced
Availability Permission-based Open access Open access
Data Structure Record-based Record-based Record-based
Instances 60,515 entries 70,000 entries 1,190 entries
Attributes 14 variables 12 variables 12 variables
Attribute Nature Categorical + Numerical Categorical + Numerical Categorical + Numerical
Incomplete Values Present None Present
Redundant Entries Contains duplicates Contains duplicates No duplicates
Class Categories 2 distinct 2 distinct 2 distinct
Class Balance Skewed Skewed Skewed

Dataset I utilized under this study was sources from Kaggle. It is a structured cardiovascular dataset comprising 1190 clinically assessed patient records. It includes 12 features categorized into objective demographics (age, gender), examination metrics (blood pressure, cholesterol, ECG), patient-reported symptoms (chest pain, exercise-induced angina), and a binary target variable (cardiovascular disease). Clinical measurements feature specialized annotations including ST-depression ischemia quantification and 4-class angina characterization, providing granular insights for predictive modeling of cardiac risk. The complete dataset I detail is provided in Table 2.

Table 2.

Attribute specification for CVD prediction modeling - objective, examination, and subjective variables of Data I.

Sr. Feature name Feature type Data type Data distribution
1 Age Objective Feature int (years) 53.72 ± 9.36
2 Gender Objective Feature categorical (0: female, 1: male) 0:1 (76.39% : 23.61%)
3 Chest Pain Subjective Feature categorical (1: typical angina, 2: atypical angina, 3: non-angina pain, 4: asymptomatic)

1:2:3:4

(5.55%: 18.15% : 23.78% : 52.52%)

4 Resting Blood Pressure Examination Feature int (mmHg) 132.15 ± 18.37
5 Cholesterol Examination Feature int (mg/dL) 210.36 ± 101.42
6 Fasting Blood Sugar Examination Feature binary (0: normal, 1: elevated) 0:1 (78.66% : 21.34%)
7 Resting ECG Examination Feature categorical (0: normal, 1: ST-T abnormality, 2: LV hypertrophy) 0:1:2 (57.48% : 15.21% : 27.31%)
8 Max Heart Rate Examination Feature int (bpm) 139.73 ± 25.52
9 Exercise-Induced Angina Subjective Feature binary (0: no, 1: yes) 0:1 (67.26% : 38.74%)
10 ST Depression (Exercise) Examination Feature float (mm) 0.92 ± 1.09
11 ST Segment Slope (Peak Exercise) Examination Feature categorical (0: normal, 1: upsloping, 2: flat, 3: downsloping) 0:1:2:3 (0.08% : 44.02% : 48.91% : 6.81%)
12 Cardiovascular Disease Target Feature categorical (0: absence, 1: presence) 0:1 (47.14% : 52.86%)

Dataset II, sourced from Kaggle and known as the Cardiovascular Disease Dataset, comprises 70,000 patient records with various attributes designed to predict and analyze CVD. The detailed description involving features types such as objective data, examination results, and subjective information is presented in Table 3. It includes 12 columns in total, with the first 11 columns featuring attributes such as age, gender, height, weight, systolic blood pressure (ap_hi), diastolic blood pressure (ap_lo), cholesterol level (gluc), smoking status (smoke), alcohol intake, and physical activity. The final column, cardio, serves as the target variable, indicating the presence (1) or absence (0) of cardiovascular disease, forming a binary classification problem.

Table 3.

Attribute specification for CVD prediction modeling - objective, examination, and subjective variables of Data II.

Sr. Feature name Feature type Data type Data distribution
1 Age Objective Feature int (days) 53.34 ± 6.76 (years)
2 Height Objective Feature int (cm) 164.36 ± 8.21
3 Weight Objective Feature float (kg) 74.21 ± 14.40
4 Gender Objective Feature categorical code (1: female, 2: male) 1:2 (65.04% : 34.96%)
5 Systolic Blood Pressure Examination Feature int (ap_hi) 128.82 ± 154.01
6 Diastolic Blood Pressure Examination Feature int (ap_lo) 96.63 ± 188.47
7 Cholesterol Examination Feature int (1: normal, 2: above normal, 3: well above normal) 1:2:3 (74.84%:13.64%:11.52%)
8 Glucose Examination Feature int (1: normal, 2: above normal, 3: well above normal) 1:2:3 (84.97% : 7.41% : 7.62%)
9 Smoking Subjective Feature binary (0: non-smoker, 1: smoker) 0:1 (91.19% : 8.81%)
10 Alcohol Intake Subjective Feature binary (0: non-alcoholic, 1: alcoholic) 0:1 (94.62% : 5.38%)
11 Physical Activity Subjective Feature binary (0: inactive, 1: active) 0:1 (19.63% : 80.37%)
12 Cardio Target Feature binary (0: No CVD, 1: CVD) 0:1 (50.03% : 49.57%)

Whereas Dataset III is a cardiovascular health dataset collected at Farooq Hospital DHA, Lahore, Pakistan. It contains 150,000 patients records with 14 features as presented in Table 4, categorized into objective demographics (sex, BMI, age groups), subjective self-reports (general health perception, sleep patterns, 6 chronic conditions, smoking/e-cigarette behaviors), and clinical examination (blood pressure status). The dataset features specialized encodings including a 5-tier ordinal health scale, 4-category smoking classification, and 13-stage age stratification, enabling granular analysis of lifestyle-disease interactions across life stages. Particular emphasis is placed on behavioral gradations through differentiated current smoking categories and e-cigarette usage frequencies.

Table 4.

Attribute specification for CVD prediction modeling - objective, examination, and subjective variables of Data III (local).

Sr. Feature name Feature type Data type Data distribution
1 Sex Objective Feature categorical (0: Male, 1: Female) 0 : 1 (48.96% : 51.04%)
2 GeneralHealth Subjective Feature ordinal (1: Excellent, 2: Very Good, 3: Good, 4: Fair, 5: Poor) 1:2:3:4:5 (16.63%:3.75%: 32.33%:13.25%: 4.04% )
3 Blood Pressure Examination Feature categorical (0: Normal, 1: High) 0:1 ( 23.46% : 76.54%)
4 SleepHours Subjective Feature int (hours per night) 7.01 ± 1.49
5 HadAngina Subjective Feature binary (0: No, 1: Yes) 0:1 (94.54% : 5.46%)
6 HadAsthma Subjective Feature binary (0: No, 1: Yes) 0:1 (84.32% : 15.68%)
7 HadCOPD Subjective Feature binary (0: No, 1: Yes) 0:1 (92.09% : 7.91%)
8 HadDepressiveDisorder Subjective Feature binary (0: No, 1: Yes) 0:1 (77.95% : 22.05%)
9 HadDiabetes Subjective Feature binary (0: No, 1: Yes) 0:1:2:3 (83.43%:13.38%:2.32%:0.87%)
10 SmokerStatus Subjective Feature categorical (0: Never smoked, 1: Former smoker, 2: Current smoker (some days), 3: Current smoker (every day)) 0:1:2:3 (58.33% : 28.60% : 3.72% : 9.35%)
11 ECigaretteUsage Subjective Feature categorical (0: Never used, 1: Not at all (now), 2: Use some days, 3: Use every day) 0:1:2:3 (74.30%:19.56%:3.26%:2.87%)
12 BMI Objective Feature float (kg/m²) 28.21 ± 6.45
13 AgeCategory Objective Feature ordinal (1:18–24, 2:25–29, 3:30–34, 4:35–39, 5:40–44, 6:45–49, 7:50–54, 8:55–59, 9:60–64, 10:65–69, 11:70–74, 12:75–79, 13:80+) 53.0 ± 18.62 (years)
14 Target Target Feature binary (0: No, 1: Yes) 0:1 (38.35% : 61.65%)

Data preprocessing

Data preprocessing is a vital step in developing machine learning and deep learning models, as it ensures that the datasets are clean, well-scaled, and representative of underlying patterns to enhance model performance and prediction accuracy. Under this study several key preprocessing techniques were applied to prevent overfitting and improve generalization across all datasets. These included handling missing values, removing outliers, and conducting comprehensive feature engineering, such as feature transformation, categorization, and attribute combination. Additionally, label encoding, and clustering, were also utilized to reduce noise and bolster model robustness, with clustering serving as the main component to refine data structures. Further details of each considered preprocessing step are listed below;

  • Data cleaning: All datasets were thoroughly examined for missing values. Datasets I and II contained no missing entries and therefore required no imputation or data removal. In contrast, Dataset III (local) contained a total of 4484 missing records, which were subsequently removed prior to analysis.

  • Removing outliers/duplicates: In the preprocessing phase, a comprehensive analysis was conducted across the three datasets to address data quality issues. For Dataset I, no outliers or duplicate entries were identified, thus requiring no corrective measures. In contrast, Dataset II exhibited the presence of various outliers (i.e., 1299), particularly in features such as systolic blood pressure (ap_hi), diastolic blood pressure (ap_lo), weight, and height. These outliers were mitigated by restricting data points to the 2.5th to 97.5th percentile range, similarly, 10,981 duplicate entries were also detected in dataset III during data cleaning. Following this preprocessing, the size of Dataset II was reduced to 68,701 records. Similarly, Dataset III after missing values and duplicate removal, resulted in a reduction from its initial size of 150,000 to 134,550 records, ensuring enhanced data integrity for subsequent modeling.

  • Feature transformation: The “age” column was transformed from days to years by dividing by 365 and rounding to the nearest integer particularly in data II where age was recorded in days. This conversion ensured that data is in a more meaningful and understandable format for modeling approaches, facilitating better prediction accuracy.

  • Feature categorization: The “age” feature across all three datasets was restructured into distinct age groups to enhance the modeling process, prompted by the collection of age data in ranges within Dataset III. This transformation involved categorizing the “age” column into broader age group categories, replacing individual age values with a new feature named “age_group.” This adjustment was applied uniformly across all datasets, with the “age_group” feature defined by age classes instead of numerical values, representing various age brackets.

  • To simplify the dataset and reduce dimensionality of dataset II in particular, new attributes were created by combining existing features. For example, for dataset II, Body Mass Index (BMI) was derived from the weight and height columns and categorized into six groups (0 to 5). Similarly, Mean Arterial Pressure (MAP) was calculated using the ap_hi (systolic blood pressure) and ap_lo (diastolic blood pressure) columns and also categorized into six groups. These transformations made the dataset more interpretable and less complex, enhancing its suitability for predictive modeling.

  • Label encoding: Label encoding serves to transform categorical data into numerical values, rendering it suitable for machine learning and deep learning models. This study implemented both Label Encoding and Ordinal Encoding across all three datasets to facilitate this conversion. By assigning unique numerical values starting from 0, the process ensures compatibility with the models while mitigating potential misinterpretations where higher values might suggest increased significance. Specifically, features such as “gender” in Dataset I & III, and “HadAngina” “HadAsthma”, “HadAngina” etc., in Dataset III were encoded, enhancing the models’ ability to process these categorical variables effectively across the diverse datasets.

  • Clustering: Clustering highlights the hidden groupings in the data that are not explicitly present in the original features, facilitating enhanced data analysis. Instead of the model trying to learn the structure from raw variables alone, the cluster label provides a shortcut to a pattern discovered by unsupervised learning, which can reduce model complexity. In this study, the cluster label derived via K-Modes was included as a categorical feature to encode latent structural groupings in the data, enriching the feature space for enhanced predictive performance. This approach aligns with established frameworks in feature engineering where clustering-based feature augmentation has been shown to benefit classification models30. The elbow curve method was applied to each dataset as presented in Fig. 2 to determine the optimal number of clusters. Subsequently, a new feature indicating cluster groups, represented as values 0 and 1, was added to enrich the feature set across all datasets, improving the robustness of the subsequent modeling process.

Fig. 2.

Fig. 2

Elbow method for optimal number of clusters on dataset I(a), II(b) and III(c).

Proposed model stacking classifier

Stacking is an ensemble learning technique that combines the predictions of multiple base models using a meta-learner to achieve superior predictive performance compared to individual models31. Unlike simple ensemble methods like bagging or boosting, stacking leverages a meta-model to learn how to optimally combine the outputs of diverse base models, capturing their complementary strengths32. The proposed stacking classifier integrates three base models i.e., Random Forest, Extra Trees, and Decision Tree at level-0 and Logistic Regression as the final estimator at level-1 as shown in Fig. 3.

Fig. 3.

Fig. 3

Architecture of the proposed stacking ensemble model.

Mathematical description of the stacking model

Let the training dataset be denoted as:

graphic file with name d33e1156.gif 1

where Inline graphic represents the feature vector and Inline graphic is the target label for binary classification.

Base models

The stacking model integrates predictions from multiple base classifiers to enhance predictive performance. The base models used in this stacking classifier are:

Random forest classifier

A Random Forest consists of Inline graphic decision trees, where each tree Inline graphic (for Inline graphic ) is trained on a bootstrap sample of the training data, and features are randomly sampled at each split to reduce correlation between trees. For a given input Inline graphic, the Random Forest outputs a probability vector for class Inline graphic.

graphic file with name d33e1200.gif 2

where Inline graphic is the probability estimate from the Inline graphic-th tree for class Inline graphic.

Extra trees classifier

Similar to Random Forest, the Extra Trees Classifier uses Inline graphic trees. However, it introduces additional randomness by selecting split points randomly for each feature, rather than optimizing splits. The output probability for class Inline graphic is.

graphic file with name d33e1233.gif 3
Decision tree classifier

A single Decision Tree with maximum depth Inline graphic partitions the feature space into regions based on feature thresholds. For an input Inline graphic, the tree assigns a class probability based on the proportion of training samples in the leaf node containing Inline graphic.

graphic file with name d33e1256.gif 4

Stacking classifier

The Stacking Classifier combines the predictions of the base models using a meta-learner, in this case, a Logistic Regression model. Let the base models be indexed by Inline graphic, where Inline graphic 3 (Random Forest, Extra Trees, Decision Tree). For an input Inline graphic, each base model Inline graphic produces a probability vector by following Eq. 5, where Inline graphic is the number of classes.

graphic file with name d33e1289.gif 5

Meta-features

The predictions from the base models are used as meta-features for the meta-learner. For a given input Inline graphic, the meta-feature vector is constructed by concatenating the probability outputs of all base models:

graphic file with name d33e1301.gif 6

Thus, if there are Inline graphic classes, the meta-feature vector Inline graphic has dimensionality Inline graphic.

Meta-learner (Logistic Regression)

The meta-learner is a Logistic Regression model with Inline graphic-regularization (penalty Inline graphic, regularization strength Inline graphic, solver liblinear, and maximum iterations 100). It takes the metafeatures Inline graphic as input and outputs the final probability for each class Inline graphic :

graphic file with name d33e1344.gif 7

where: Inline graphic is the weight vector for class Inline graphic, Inline graphic is the bias term for class Inline graphic, The weights and biases are learned by minimizing the regularized log-loss (cross-entropy loss):

Here, Inline graphic is an indicator variable (1 if the true class of the Inline graphic-th sample is Inline graphic otherwise), and Inline graphic is the Inline graphic-regularization term.

Training and prediction process

The proposed stacking framework follows a structured two-stage training process, incorporating internal k-fold cross-validation to ensure unbiased meta-feature generation and prevent information leakage:

  1. Base model training and out-of-fold meta-feature generation: Each base model is first trained on the original training dataset Inline graphic. To generate the meta-features required for the second stage, the stacking classifier applies k-fold cross-validation (typically Inline graphic) on the training data. In each fold, a base model is trained on Inline graphic folds and used to predict the held-out fold. This process produces out-of-fold (OOF) predictions for every instance in the training set. For each sample Inline graphic, the prediction Inline graphicfrom the base model Inline graphicis generated only from the model trained without including that sample, ensuring that no information from Inline graphicinfluences its own prediction. These OOF predictions constitute the meta-feature vector.

graphic file with name d33e1430.gif 9

which serves as the input for training the meta-learner.

  • 2.

    Meta-learner training: In the second stage, the meta-learner (Logistic Regression) is trained on the newly constructed dataset Inline graphic comprising the OOF meta-features and their corresponding true labels. After this training step, all base models are retrained using the entire training set, and the meta-learner operates on their final predictions during inference. This ensures that the full information in the training data is utilized while maintaining the integrity and non-leakage guarantees provided by the OOF cross-validation procedure.

This two-stage training process aligns with standard stacking methodologies and ensures that the model achieves robust generalization, minimizes overfitting, and adheres to best practices for ensemble learning in biomedical machine learning applications.

Whereas for predicting the class of a new input Inline graphic, each base model Inline graphic first computes its corresponding probability vector Inline graphic. These probability vectors from all base models are then concatenated to form the meta-feature vector Inline graphic. The meta-learner processes this meta-feature vector to compute the final class probabilities Inline graphic for each class Inline graphic, and the class with the highest probability is selected as the predicted class following Eq. (10) below:

graphic file with name d33e1480.gif 10

This mathematical formulation ensures that the stacking classifier leverages the strengths of diverse base models while using Logistic Regression to optimally combine their predictions, enhancing overall classification performance.

Performance metrics

This study assesses the performance of the implemented models using various evaluation metrics, including accuracy, precision, recall, F1-score, and confusion matrix. Precision, which measures the proportion of true positives (TP) relative to all predicted positives (TP + false positives (FP)), is crucial in scenarios where minimizing false positives is essential due to their high cost. Recall, calculated as the ratio of TP to all actual positives (TP + false negatives (FN)), is critical in applications like medical diagnoses, where failing to identify positive cases carries significant consequences. Accuracy reflects the overall correctness of predictions but can be misleading for imbalanced datasets, making the F1-score, which balances precision and recall, a more robust metric. The confusion matrix provides a detailed breakdown of predictions, where TP denotes correctly identified positives, FP represents negatives incorrectly classified as positives, FN indicates positives misclassified as negatives, and TN refers to correctly classified negatives.

graphic file with name d33e1490.gif 11
graphic file with name d33e1496.gif 12
graphic file with name d33e1501.gif 13
graphic file with name d33e1508.gif 14

Experimental setting

The experiments were conducted using a combination of software tools, machine learning libraries, and computing resources to ensure reliable model development and evaluation. Microsoft Excel was initially used for dataset cleaning and formatting, while Python served as the primary programming environment for data preprocessing, model training, and evaluation. Essential Python libraries including NumPy, pandas, matplotlib, Scikit-learn, XGBoost, PyTorch-TabNet, and Keras/TensorFlow enabled efficient implementation of both traditional machine learning algorithms and deep learning architectures. All experiments were executed in Google Colaboratory, a cloud-based Jupyter environment that provides GPU-accelerated computation when available. The hardware configuration for local testing included an Intel Core i5 processor (8 GB RAM) running on Ubuntu OS, which offered adequate computational support for model development and verification.

All three datasets were partitioned into training and testing sets using an 80:20 ratio. This split ensured consistent evaluation while preventing any overlap between training and testing instances. For the deep learning models (FFNN, CNN, TabNet), an additional validation set was created from the training portion to monitor learning progress and mitigate overfitting.

Hyperparameter optimization was systematically performed for each model to ensure a fair and comparable evaluation across all datasets. Recognizing that no single configuration is universally optimal, key hyperparameters were selected based on empirical performance, computational efficiency, and established best-practice guidelines. For traditional machine learning models, tuning emphasized parameters such as the number of estimators, maximum depth, splitting criteria, and class weighting to balance performance and generalization.

For deep learning architectures, the primary considerations included network depth, activation functions, learning rate, batch size, dropout, and the number of training epochs. Specifically, the ANN employed a three-layer Dense network (64 → 32 → 1) with ReLU activations and Sigmoid output, optimized with Adam and trained over 32 epochs with a batch size of 32. The CNN model comprised two stacked Conv1D layers (filters = 64 and 128, kernel_size = 2) with MaxPooling1D, followed by Dense(256) and Dropout(0.1), trained using Adam (lr = 0.001) for 32 epochs with a batch size of 16.

The TabNet model utilized attentive feature masking with entmax sparsity, controlled learning via Adam optimizer (lr = 1e−3), and incorporated 5 decision steps, n_d = 32, n_a = 32, with a maximum of 100 epochs and patience of 10 to ensure robust convergence. For the proposed preprocessing-enhanced stacking classifier, the base learners were configured with major hyperparameters: Random Forest (n_estimators = 300, max_depth = None, n_jobs = − 1), Extra Trees (n_estimators = 550, max_depth = None, n_jobs = − 1), and Decision Tree (max_depth = 10, min_samples_split = 2, min_samples_leaf = 4). The Logistic Regression meta-learner was configured with L2 regularization (C = 100), solver = liblinear, and max_iter = 100 to ensure stable and efficient training while integrating predictions from diverse base models.

A consolidated overview of the major hyperparameters for all traditional ML models, deep neural networks, and the proposed stacking classifier is provided in Table 5. This unified configuration ensures transparency, reproducibility, and methodological clarity across the experimental workflow.

Table 5.

Detailed Hyperparameter summary of all the considered modeling approaches.

Model Major hyperparameters
Decision Tree max_depth = 7, criterion = “gini” (default), random_state = 42
Random Forest n_estimators = 100, min_impurity_decrease = 0.01, bootstrap = True, class_weight = “balanced”, random_state = 42
XGBoost n_estimators = 100, max_depth = 3, learning_rate = 0.01, objective = “binary: logistic”, random_state = 42
Logistic Regression penalty = “l1” or “l2”, solver = “saga” or “liblinear”, max_iter = 1000, C = 100*
Artificial Neural Network (FFNN) Layers: Dense(64 → 32 → 1), Activation: ReLU (hidden), Sigmoid (output), optimizer = “adam”, loss = “binary_crossentropy”, epochs = 32, batch_size = 32
Convolutional Neural Network Conv1D: filters = 64 & 128, kernel_size = 2; MaxPooling1D(pool_size = 2); Dense(256), Dropout(0.1); optimizer = Adam(lr = 0.001), loss = “binary_crossentropy”, epochs = 32, batch_size = 16
TabNet n_d = 32, n_a = 32, n_steps = 5, gamma = 1.5, lambda_sparse = 1e-4, optimizer = Adam(lr = 1e-3), mask_type = “entmax”, max_epochs = 100, batch_size = 256, patience = 10, virtual_batch_size = 128
Proposed Stacking Model

Base Models: Random Forest: n_estimators = 300, max_depth = None, random_state = 42, n_jobs = − 1, Extra Trees: n_estimators = 550, max_depth = None, random_state = 42, n_jobs = − 1, Decision Tree: max_depth = 10, min_samples_split = 2, min_samples_leaf = 4, random_state = 42

Meta-Learner: Logistic Regression: C = 100, penalty = “l2”, solver = “liblinear”, max_iter = 100, random_state = 42

Performance evaluation

The detailed results obtained on Dataset I, summarized in Table 6, demonstrate notable performance differences among machine learning, deep learning, and ensemble-based models for binary CVD classification. Among the traditional machine learning approaches, Stacking Classifier achieved the highest accuracy (92.85%), followed closely by the CNN (87.81%) and FFNN (86.1%), while the Decision Tree model attained an accuracy of 85.3%. Whereas, TabNet achieved a balanced performance with an accuracy of 87.3%. The ensemble-based XGBoost model recorded an accuracy of 83.6%, but showed comparatively lower recall (0.86) for the negative class. In contrast, the proposed stacking classifier significantly outperformed all baseline models, achieving an accuracy of 92.85% and consistently high precision, recall, and F1-scores (≈ 0.93) for both classes.

Table 6.

Model performance comparison following the dataset I.

Model group Algorithm Precision Recall F1 score Accuracy (%)
Machine learning DT 0 0.89 0.79 0.84 85.3
1 0.83 0.91 0.87
LR 0 0.86 0.83 0.85 85.7
1 0.85 0.88 0.87
RF 0 0.82 0.87 0.84 85.2
1 0.88 0.83 0.85
Deep learning FFNN 0 0.86 0.85 0.85 86.1
1 0.87 0.87 0.87
CNN 0 0.88 0.86 0.87 87.81
1 0.88 0.90 0.89
TabNet 0 0.89 0.83 0.86 87.3
1 0.86 0.91 0.88
Ensemble XGBoost 0 0.81 0.86 0.83 83.6
1 0.87 0.82 0.84
Stacking 0 0.91 0.95 0.93 92.85
1 0.95 0.91 0.93

Table 7 summarizes the performance of machine learning, deep learning, and ensemble models on the 70,000-record Dataset II. Overall, all models exhibit closely clustered performance, indicating the challenging and heterogeneous nature of the dataset. Among machine learning approaches, Decision Tree achieved the highest accuracy (71.4%), while deep learning models such as FFNN (71.8%) and TabNet (71.5%) delivered comparable results; CNN performed notably worse (59.51%), reflecting its limited suitability for tabular CVD data. The proposed preprocessing-enhanced stacking model attained an accuracy of 71.5%, matching the best-performing baselines and consistently maintaining balanced precision, recall, and F1-scores across both classes, highlighting its robustness and reliable generalization.

Table 7.

Model performance comparison following the dataset II.

Model group Algorithm Precision Recall F1 score Accuracy (%)
Machine learning DT 0 0.70 0.75 0.73 71.4
1 0.73 0.68 0.70
LR 0 0.69 0.77 0.73 71.1
1 0.74 0.65 0.69
RF 0 0.68 0.75 0.72 70.4
1 0.72 0.64 0.68
Deep Learning FFNN 0 0.72 0.72 0.72 71.8
1 0.72 0.72 0.72
CNN 0 0.57 0.76 0.65 59.51
1 0.63 0.42 0.50
TabNet 0 0.71 0.75 0.73 71.5
1 0.73 0.68 0.70
Ensemble XGBoost 0 0.70 0.75 0.72 71.1
1 0.73 0.67 0.70
Stacking (proposed) 0 0.70 0.75 0.73 71.5
1 0.73 0.68 0.70

Interestingly, the performance of all models on Dataset III (local dataset), as summarized in Table 8, demonstrates exceptionally high classification effectiveness. Traditional machine learning models, particularly Decision Tree (99.4%) and Random Forest (99.5%), achieved near-perfect accuracy with precision, recall, and F1-scores approaching 1.00 for both classes. Similarly, deep learning models i.e., FFNN, CNN, and TabNet exhibited consistently strong performance, each attaining approximately 99.5% accuracy with balanced class-wise metrics, indicating highly reliable discrimination on this dataset. Among ensemble approaches, XGBoost and the proposed preprocessing-enhanced stacking model also achieved 99.5% accuracy, with stable precision, recall, and F1-scores (≈ 0.99–1.00) across both classes. These results confirm that the stacking framework effectively integrates the strengths of diverse base learners while maintaining robust and consistent predictions.

Table 8.

Performance comparison on Dataset III (local data).

Model group Algorithm Precision Recall F1 score Accuracy (%)
Machine learning DT 0 1.00 0.99 0.99 99.4
1 0.99 1.00 1.00
LR 0 0.86 0.92 0.89 90.8
1 0.95 0.90 0.92
RF 0 1.00 0.99 0.99 99.5
1 0.99 1.00 0.99
Deep Learning FFNN 0 1.00 0.99 0.99 99.5
1 0.99 1.00 1.00
CNN 0 1.00 0.99 0.99 99.5
1 0.99 1.00 1.00
TabNet 0 1.00 0.99 0.99 99.5
1 0.99 1.00 1.00
Ensemble XGBoost 0 1.00 0.99 0.99 99.5
1 0.99 1.00 1.00
Stacking (proposed) 0 1.00 0.99 0.99 99.5
1 0.99 1.00 1.00

In contrast, Logistic Regression showed comparatively lower performance, achieving an accuracy of 90.8%, with reduced precision and F1-scores, reflecting its limited ability to model complex, non-linear decision boundaries present in the data. Overall, the uniformly high performance across most models suggests that Dataset III is well-structured and highly separable, enabling multiple learning paradigms including the proposed stacking approach to achieve near-optimal results, while also underscoring the importance of model capacity in capturing underlying data characteristics.

To assess the robustness and statistical reliability of the proposed model beyond single-run evaluations, performance was analyzed across multiple random initializations. Accordingly, Table 9 reports a comparative evaluation of top-performing models on Datasets I, II, and III, expressed as mean values with 95% confidence intervals over five different random seed values (0, 7, 42, 123, and 2024). For Dataset I, the proposed stacking model achieves the highest overall performance, attaining an accuracy of 0.9328 ± 0.0059 and an F1-score of 0.9359 ± 0.0059, outperforming all individual machine learning, deep learning, and ensemble baselines while also exhibiting low variance across seeds. This indicates strong predictive capability and stable generalization on a moderately complex dataset.

Table 9.

Performance comparison of top baseline contenders and proposed approach with 95% confidence interval following 5 different seed values (0, 7, 42, 123 and 2024) under data I, II and III.

Dataset Model Accuracy Precision Recall F1-score
 I Decision Tree 0.9013 ± 0.0042 0.9133 ± 0.0032 0.8988 ± 0.0100 0.9060 ± 0.0045
Random Forest 0.8424 ± 0.0189 0.8735 ± 0.0056 0.8214 ± 0.0422 0.8462 ± 0.0223
XGBoost 0.8908 ± 0.0084 0.8917 ± 0.0067 0.8903 ± 0.0085 0.8903 ± 0.0085
Logistic Regression 0.8571 ± 0.0000 0.8538 ± 0.000 0.8810 ± 0.0000 0.8672 ± 0.0000
FFNN 0.8918 ± 0.0125 0.8971 ± 0.0107 0.8988 ± 0.0176 0.8979 ± 0.0122
CNN 0.8782 ± 0.0146 0.8620 ± 0.0407 0.9206 ± 0.0372 0.8891 ± 0.0092
Stacking 0.9328 ± 0.0059 0.9453 ± 0.0040 0.9266 ± 0.0100 0.9359 ± 0.0059
II Decision Tree 0.7143 ± 0.0000 0.7270 ± 0.0000 0.6768 ± 0.0000 0.7010 ± 0.0000
Random Forest 0.7005 ± 0.0043 0.7279 ± 0.0034 0.6306 ± 0.0204 0.6756 ± 0.0102
XGBoost 0.7114 ± 0.0000 0.7263 ± 0.0000 0.6688 ± 0.0000 0.6964 ± 0.0000
Logistic Regression 0.7116 ± 0.0000 0.7357 ± 0.0000 0.6510 ± 0.0000 0.6908 ± 0.0000
FFNN 0.7174 ± 0.0020 0.7352 ± 0.0184 0.6726 ± 0.0417 0.7015 ± 0.0143
CNN 0.5915 ± 0.0004 0.6393 ± 0.0025 0.4005 ± 0.0047 0.4925 ± 0.0029
Stacking 0.7150 ± 0.0002 0.7279 ± 0.0005 0.6773 ± 0.0007 0.7017 ± 0.0002
III Decision Tree 0.9947 ± 0.00002 0.9924 ± 0.0000 0.9991 ± 0.00004 0.9957 ± 0.00002
Random Forest 0.9953 ± 0.0000 0.9924 ± 0.000 1.0000 ± 0.000 0.9962 ± 0.0000
XGBoost 0.9953 ± 0.000 0.9924 ± 0.0000 1.0000 ± 0.0000 0.9962 ± 0.000
Logistic Regression 0.9089 ± 0.00004 0.9460 ± 0.000003 0.9032 ± 0.00006 0.9241 ± 0.00003
FFNN 0.9953 ± 0.0000 0.9924 ± 0.0000 1.000 ± 0.0000 0.9962 ± 0.000
CNN 0.9952 ± 0.00013 0.9925 ± 0.00004 0.9998 ± 0.000257 0.9961 ± 0.0001
Stacking 0.9952 ± 0.0000 0.9924 ± 0.0000 0.9998 ± 0.0000 0.9961 ± 0.0000

On Dataset II, all models demonstrate closely clustered performance with relatively narrow confidence intervals, reflecting the challenging and heterogeneous nature of the dataset. The proposed stacking approach achieves an accuracy of 0.7150 ± 0.0002 and an F1-score of 0.7017 ± 0.0002, performing comparably to the strongest baselines (Decision Tree, FFNN, and XGBoost) while maintaining consistent precision–recall balance and minimal variability across runs. For Dataset III, near-ceiling performance is observed across most models, with accuracies exceeding 99.4% and extremely small confidence intervals, suggesting that the dataset is highly separable. In this setting, the proposed stacking model attains an accuracy of 0.9952 ± 0.0000 and an F1-score of 0.9961 ± 0.0000, matching the performance of state-of-the-art baselines. While these results highlight the effectiveness and stability of the proposed method, the uniformly high scores across models indicate that performance is largely driven by dataset characteristics rather than architectural complexity.

Overall, the results demonstrate that the proposed stacking framework delivers consistent, statistically stable, and competitive performance across datasets of varying complexity, achieving clear improvements on Dataset I, performance parity on Dataset II, and robust behavior on Dataset III without evidence of instability across random initializations.

To assess the statistical significance of the observed performance differences, McNemar’s test (chi-square with continuity correction) was applied to compare the proposed stacking ensemble against each baseline model across three datasets. The results presented in Table 10 indicate that stacking achieved statistically significant superiority (p < 0.05) over several key competitors: notably, over Random Forest, Logistic Regression, and XGBoost in Dataset I; over CNN in Dataset II; and over Decision Tree and Logistic Regression in Dataset III. Although not every comparison reached significance for example, against Decision Tree in Dataset II or against XGBoost and CNN in Dataset III, stacking was never outperformed at a statistically significant level by any baseline. This consistent, non-inferior performance underscores the reliability of stacking as a robust and high-performing unified model across diverse datasets.

Table 10.

Comparative performance and statistical significance of the proposed stacking ensemble against baseline models across three datasets.

Dataset Model A (Stacking) Model B b (A correct, B wrong) c (A wrong, B correct) Chi-square p-value
I Stacking Decision Tree 9 5 0.6429 0.4227
Stacking Random Forest 22 7 6.7586 0.0093
Stacking Logistic Regression 22 7 6.7586 0.0093
Stacking XGBoost 25 6 10.4516 0.0012
Stacking Feedforward NN 18 8 3.1154 0.0776
Stacking CNN 12 6 1.3889 0.2386
II Stacking Decision Tree 451 441 0.0908 0.7632
Stacking Logistic Regression 531 484 2.0847 0.1488
Stacking XGBoost 597 547 2.0988 0.1474
Stacking Feedforward NN 402 425 0.5852 0.4443
Stacking CNN 3167 1482 609.9927 < 0.0001
III Stacking Decision Tree 14 3 5.8824 0.0153
Stacking Logistic Regression 2359 38 2245.4735 < 0.0001
Stacking XGBoost 0 4 2.2500 0.1336
Stacking Feedforward NN 0 4 2.2500 0.1336
Stacking CNN 4 4 0.1250 0.7237

Significance values are in bold.

Discussion

This study conducted a comprehensive evaluation of traditional ML, DL, and ensemble-based approaches for cardiovascular disease detection across three datasets with varying characteristics. The objective was not solely to maximize peak accuracy on individual datasets, but to identify a robust and generalizable model capable of maintaining stable performance across diverse data conditions. The results highlight how dataset size, noise level, feature composition, and structural separability significantly influence model behavior and comparative effectiveness.

Traditional ML models

Traditional ML models, including Decision Tree, Logistic Regression, and Random Forest, demonstrated dataset-dependent performance. On Dataset I, DT and RF achieved competitive results, with DT showing relatively strong accuracy and balanced F1-scores, indicating its ability to capture non-linear relationships in moderately structured data. However, LR consistently underperformed relative to tree-based models, reflecting its limited capacity to model complex decision boundaries. On Dataset II, which represents a large-scale and heterogeneous setting, the performance of traditional ML models declined, with accuracies clustering around 70–71% and reduced F1-scores, suggesting sensitivity to noise and feature interactions. In contrast, on Dataset III, which is comparatively clean and well-structured, DT and RF achieved near-ceiling performance (≈ 99.5%), indicating that traditional ML approaches can be highly effective when class separability is strong. These observations suggest that while traditional ML models remain effective in structured environments, their robustness across heterogeneous datasets is limited.

Deep learning techniques

Deep learning models, including FFNN, CNN, and TabNet, exhibited mixed but informative behavior across datasets. On Dataset I, FFNN and CNN achieved competitive performance, although improvements over strong ML baselines were moderate. On Dataset II, FFNN and TabNet marginally outperformed several traditional ML models, indicating improved representation learning in larger datasets; however, CNN showed noticeably lower performance, reinforcing prior findings that convolutional architectures may be less suitable for tabular clinical data without strong spatial correlations. On Dataset III, all DL models achieved near-perfect accuracy and F1-scores, similar to traditional and ensemble approaches. This convergence suggests that model architecture becomes less critical when the dataset is highly separable. Overall, DL methods demonstrate reasonable generalization potential, but their computational complexity and inconsistent gains across datasets limit their standalone suitability as a universal solution.

Ensemble methods

Ensemble-based approaches, particularly XGBoost and the proposed preprocessing-enhanced stacking model, exhibited the most consistent behavior across datasets. On Dataset I, stacking achieved the highest accuracy and F1-score among all evaluated models, with low variance across random seeds, demonstrating strong generalization. On Dataset II, although absolute performance was constrained (71.5% accuracy), stacking matched or slightly exceeded the strongest baselines while maintaining balanced precision–recall trade-offs, indicating robustness under challenging conditions. On Dataset III, stacking achieved near-ceiling performance comparable to XGBoost, RF, and DL models. Importantly, statistical validation using McNemar’s test showed that stacking achieved significant improvements over multiple baselines on Datasets I and III and was never significantly outperformed on any dataset. These findings suggest that stacking effectively integrates complementary decision patterns from heterogeneous learners, leading to improved reliability rather than dataset-specific overfitting.

Cross-approach insights and generalization

Across all experiments, the stacking classifier demonstrated the most stable and non-inferior performance across datasets with varying complexity. While some models achieved comparable or marginally better results on individual datasets, stacking consistently balanced accuracy, precision, recall, and F1-score, particularly in challenging scenarios such as Dataset II. CNN emerged as a strong contender in specific settings but lacked consistency across all datasets, whereas traditional ML models showed strong performance only under favorable data conditions. These results underscore that generalization, robustness, and statistical stability are more informative than peak accuracy when developing clinically relevant CVD detection systems. By maintaining competitive performance across heterogeneous datasets and demonstrating statistically validated reliability, the proposed stacking framework represents a practical and dependable approach for real-world CVD risk prediction.

Comparative analysis with state-of-the-art approaches

The proposed preprocessing-enhanced stacking classifier was benchmarked against prior studies on Datasets I and II (Tables 11 and 12). On Dataset I, the model achieved 92.86% accuracy, outperforming most traditional and early ensemble methods reported in earlier works30,33, while remaining competitive with more complex deep and hybrid architectures such as the 5-layer CNN and CNN–LSTM–based frameworks19,29. This demonstrates that combining enhanced preprocessing with stacking can yield strong performance without excessive model complexity. While on Dataset II, the proposed approach attained 71.50% accuracy, reflecting the increased scale and heterogeneity of the dataset. Prior optimization-driven approaches including GA-ANN, TSTO, and TPTM-HANN-GA reported accuracies between approximately 73–74%34,35,36, while hybrid deep ensemble methods achieved higher performance under dataset-specific configurations29. However, unlike these studies, the present work applies a unified pipeline across heterogeneous datasets, prioritizing robustness and reproducibility rather than dataset-specific tuning. As a result, the stacking framework maintains balanced class-wise performance and statistical stability despite the increased variability of Dataset II. Overall, these comparisons suggest that while certain optimization-focused methods may achieve higher peak accuracy under tailored settings34,35,36, the proposed stacking framework provides consistent and generalizable performance across datasets of varying structure, supporting its suitability for practical cardiovascular disease detection in heterogeneous clinical environments.

Table 11.

Accuracy comparison of proposed Stacking model with the prior studies on Data I.

Sources Best model Accuracy (%)
Acharya 30 KNN 82.00
Kumar et al. 17 Decision Tree C4.5 83.40
Khanna et al. 16 Logistic Regression 84.80
Mokeddem et al. 15 GA + Naive Bayes 85.50
Akkaya et al. 11 KNN 85.60
Subanya and Rajalaxmi 14 SVM 86.76
Amin et al. 13 Voting classifier with Naïve Bayes and Logistic Regression 87.41
Tougui et al. 12 Random Forest 87.64
Ahamad et al. 10 SVM 87.91
Nazari et al. 33 Ensemble model based on GA 88.43
Singhal et al. 19 5-layer CNN 95.00
Sadr et al. 29 CNN-LSTM + KNN + XGB 95.85
Proposed Preprocessing Enhanced Stacking 92.86

Table 12.

Accuracy comparison of proposed Stacking model with the prior studies on Data II.

Sources Best model Accuracy (%)
Arroyo et al. 34 GA-ANN 73.43
Lin et al. 35 TSTO 74.14
Li et al. 36 TPTM-HANN-GA 74.25
Sadr et al. 29 CNN-LSTM + XGB + KNN 80.25
Proposed Preprocessing Enhanced Stacking 71.50

Limitations

While the proposed preprocessing-enhanced stacking classifier demonstrates robust and consistent performance across all three datasets, several limitations warrant acknowledgment to contextualize the findings and guide future research.

First, the substantial heterogeneity in feature composition, measurement scales, and data collection protocols across the datasets ranging from clinical/examination records (Dataset I) to anthropometric/lifestyle data (Dataset II) and survey‑based responses (Dataset III) limits direct harmonization and comparability of results. This diversity, while valuable for evaluating model robustness in varied real‑world settings, prevents the derivation of a single optimal feature set and requires that generalizability be interpreted within the specific structure and constraints of each dataset.This challenge is consistent with prior optimization-driven CVD prediction studies that report dataset-specific performance variability when feature distributions differ across populations34–36.

Second, dataset‑specific characteristics meaningfully influenced model performance. Notably, Dataset III is highly structured and clean, leading to near‑perfect performance across several classifiers. Although this validates the approach under controlled conditions, it may overestimate generalizability to routine clinical environments where data are often noisy, incomplete, and heterogeneous. Moreover, Dataset II revealed performance trade‑offs, such as XGBoost’s high accuracy but comparatively lower F1‑score, and Random Forest’s stability at the cost of reduced recall for certain classes. These observations highlight sensitivity to dataset‑specific factorsincluding class imbalance, feature distribution shifts, and population heterogeneitywhich were not explicitly optimized in the current study.

Third, the computational complexity of ensemble and deep learning models, including the CNN and the proposed stacking classifier, may hinder deployment in resource‑constrained clinical environments. Although these models deliver improved predictive performance, further optimization for efficiency, real‑time inference, and lower power consumption would be necessary for practical integration into clinical workflows.

Looking forward, future work should prioritize improving generalizability through cross‑institutional validation, privacy‑preserving synthetic data generation, and federated learning frameworks. Additionally, enhancing model reproducibility, reliability, and interpretability through explainable AI techniques and transparent reporting will be essential for the successful translation of machine learning‑based cardiovascular diagnostic systems into clinical practice.

Conclusion

This study presents a systematic evaluation of machine learning, deep learning, and ensemble models for cardiovascular disease prediction across three heterogeneous datasets, employing a carefully designed preprocessing pipeline. The proposed stacking ensemble consistently achieved the highest and most robust performance, with accuracy of 93% in Dataset I, 72% in Dataset II, and near‑perfect results (99%) in Dataset III. McNemar’s test confirmed that the stacking model significantly outperformed multiple strong baselines including Random Forest, Logistic Regression, and XGBoost in Dataset I, and CNN in Dataset II demonstrating statistically reliable superiority. The strength of the stacking classifier stems from its ability to integrate diverse base learners while leveraging an optimized preprocessing stage that includes feature transformation, clustering‑derived representations, and importance‑based feature selection. This design enables the model to capture complementary patterns across clinical, lifestyle, and survey‑based data types, yielding generalizable performance even under dataset‑specific challenges such as class imbalance and feature heterogeneity.

Despite these results, certain limitations persist, including dataset disparities that prevent full harmonization, the computational overhead of ensemble methods, and performance trade‑offs on noisier real‑world data. Future work should prioritize cross‑institutional validation, efficiency optimization for clinical deployment, and the incorporation of explainable AI techniques to enhance interpretability and trust.

Acknowledgements

The authors would like to thank the administration of Farooq Hospital DHA, Lahore, Pakistan, for generously providing access to the locally collected dataset used in this study. Their support and cooperation were invaluable to the successful completion of this research.

Author contributions

A.A. and A.M. contributed equally and share first authorship. A.M. led conceptualization, software development, investigation, and supervision. A.A. and A.S. drafted the manuscript, with D.M.B. and J.M. contributing to review and editing. X.Y. provided clinical domain expertise, cross-dataset rigor, addressed major reviewer concerns, and contributed to manuscript revision and reproducibility support. M.S.H.A. and A.A*. contributed to project administration and overall coordination.

Funding

No funding was acquired under this study.

Data availability

Our experiments utilized three datasets: two publicly available from Kaggle can be obtained following the specific citation, and one locally collected dataset, which can be freely accessed for academic use upon request. While the detailed code files can be freely accessed on Github through https://github.com/AyshaSaddiqa/CVC-detection.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Adven Masih, Email: adven.masih@uskt.edu.pk.

Aitzaz Ali, Email: ali@apu.edu.my.

References

  • 1.World Health Organization. cardiovascular diseases (CVDs), WHO Fact Sheets, 2021. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds).
  • 2.Roth, G. A. et al. Global, regional, and national burden of cardiovascular diseases for 10 causes, 1990 to 2015. J. Am Coll. Cardiol.70 (1), 1–25 (2017). 10.1016/j.jacc.2017.04.052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Heidenreich, P. A. et al. Forecasting the future of cardiovascular disease in the United States: A policy statement from the American Heart Association. Circulation123(8), 933–944 (2011). 10.1161/CIR.0b013e31820a55f5 [DOI] [PubMed] [Google Scholar]
  • 4.Stewart, J. et al. Psychological and social impacts of cardiovascular disease: A review. Eur. J. Prev. Cardiol.23 (14), 1492–1500 (2016). 10.1177/2047487315616997 [Google Scholar]
  • 5.Al’Aref, S. J. et al. Machine learning of clinical variables and coronary artery calcium scoring for the prediction of obstructive coronary artery disease. Eur. Heart J.40 (4), 359–367 (2019). 10.1093/eurheartj/ehy568 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Weng, S F. et al. Can machine-learning improve cardiovascular risk prediction using routine clinical data. PLoS ONE. 12 (4), e0174944 (2017). 10.1371/journal.pone.0174944. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Damen, J. A. et al. May., Prediction models for cardiovascular disease risk in the general population: Systematic review. BMJ353, i2416 (2016). 10.1136/bmj.i2416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Khondoker, M et al. Ensemble learning for cardiovascular disease prediction. J. Med. Syst.44 (12), 1–10 (2020). 10.1007/s10916-020-01670-1. [Google Scholar]
  • 9.Garavand, A., Emami, H., Rabiei, R., Pishgahi, M. & Vahidi-Asl, M. Designing the coronary artery disease registry with data management processes approach: A comparative systematic review in selected registries. Int. Cardiovasc. Res. J.14 (1), e100498 (2020). 10.5812/icrj.100498 [Google Scholar]
  • 10.Ahamad, G. N. et al. Influence of optimal hyperparameters on the performance of machine learning algorithms for predicting heart disease. Processes11 (3), 734 (2023). 10.3390/pr11030734 [Google Scholar]
  • 11.Akkaya, B., Sener, E. & Gursu, C. A comparative study of heart disease prediction using machine learning techniques. In Proceedings of the International Congress on Human-Computer Interaction, Optimization, and Robotic Applications (HORA), Ankara, Turkey 1–6 (2022). 10.1109/HORA55278.2022.9799999
  • 12.Tougui, I, Jilbab, A. & Mhamdi, J. E. Heart disease classification using data mining tools and machine learning techniques. Health Technol. 10(5), 1137–1144 (2020). 10.1007/s12553-020-00438-9. [Google Scholar]
  • 13.Amin, M. S., Chiam, Y. K. & Varathan, K. D. Identification of significant features and data mining techniques in predicting heart disease. Telematics Inf.36, 82–93 (2019). 10.1016/j.tele.2018.11.007 [Google Scholar]
  • 14.Subanya, B. & Rajalaxmi, R. Feature selection using artificial bee colony for cardiovascular disease classification. In Proceedings of International Conference on Electronics and Communication Systems (ICECS), Coimbatore, India 1–6 (2014). 10.1109/ECS.2014.6892741
  • 15.Mokeddem, S., Atmani, B. & Mokaddem, M. Supervised feature selection for diagnosis of coronary artery disease based on genetic algorithm, arXiv preprint. (2013). Available: https://arxiv.org/abs/1305.6046
  • 16.Khanna, D. et al. Comparative study of classification techniques (SVM, logistic regression and neural networks) to predict the prevalence of heart disease. Int. J. Mach. Learn. Comput.5 (5), 414–418. (2015). 10.7763/IJMLC.2015.V5.541 [Google Scholar]
  • 17.Kumar, M. N., Koushik, K. & Deepak, K. Prediction of heart diseases using data mining and machine learning algorithms and tools. Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol.3(3), 887–898 (2018). [Google Scholar]
  • 18.Javid, I., Zafar, A. & Ghafoor, A. Cardiovascular disease risk prediction using gradient boosting techniques. New. Generation Comput.41 (4), 1–18. 10.1007/s00354-023-00234-1 (2023). [Google Scholar]
  • 19.Singhal, S., Kumar, H. & Passricha, V. Prediction of heart disease using CNN. Am. Int. J. Res. Sci. Technol. Eng. Math.23 (1), 257–261 (2018). [Google Scholar]
  • 20.Dutta, A et al. An efficient convolutional neural network for coronary heart disease prediction. Expert Syst. Appl.159, 113408. (2020). 10.1016/j.eswa.2020.113408. [Google Scholar]
  • 21.Acharya, U R. et al. A deep convolutional neural network model to classify heartbeats. Comput. Biol. Med.89, 389–396. (2017). 10.1016/j.compbiomed.2017.08.022. [DOI] [PubMed] [Google Scholar]
  • 22.Samavat, T. & Hojatzadeh, E. Programs for Prevention and Control of Cardiovascular Diseases (Ministry of Health, 2012).
  • 23.Filimonov, M. & Masih, A. Presentation of special series with computed recurrently coefficients of solutions of nonlinear evolution equations. J. Phys: Conf. Ser.722(1) 012040. (2016). 10.1088/1742-6596/722/1/012040 [Google Scholar]
  • 24.Saranya, K et al. DenseNet-ABiLSTM: Revolutionizing multiclass arrhythmia detection and classification using hybrid deep learning approach leveraging. Int. J. Comput. Intell. Syst., 18(1), 1–15 . (2025). 10.1007/s44196-025-00765-z. [Google Scholar]
  • 25.Mehmood, A et al. Prediction of heart disease using deep convolutional neural networks. Arab. J. Sci. Eng.46 (4), 3409–3422. (2021). 10.1007/s13369-020-05105-1. [Google Scholar]
  • 26.Tarawneh, M. & Embarak, O. Hybrid approach for heart disease prediction using data mining techniques. In Proceedings of International Conference on Emerging Internetworking, Data & Web Technologies, Cham, Switzerland 447–454 (2019). 10.1007/978-3-030-12839-5_41
  • 27.Bhavekar, G. S. & Goswami, A. D. A hybrid model for heart disease prediction using recurrent neural network and long short-term memory. Int. J. Inf. Technol.14 (4), 1781–1789. (2022). 10.1007/s41870-022-00896-7 [Google Scholar]
  • 28.Subhadra, K. & Vikas, B. Neural network based intelligent system for predicting heart disease. Int. J. Innov. Technol. Explor. Eng.8 (5), 484–487 (2019). [Google Scholar]
  • 29.Sadr, H, Salari, A. & Ashoobi, M. T. cardiovascular disease diagnosis: A holistic approach using the integration of machine learning and deep learning models. Eur. J. Med. Res.29, 455. (2024). 10.1186/s40001-024-02044-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Piernik, M. & Morzy, T. A study on using data clustering for feature extraction to improve the quality of classification. Knowl. Inf. Syst.63, 1771–1805. 10.1007/s10115-021-01572-6 (2021). [Google Scholar]
  • 31.Wolpert, D. H. Stacked generalization. Neural Netw.5 (2), 241–259. 10.1016/S0893-6080(05)80023-1 (1992). [Google Scholar]
  • 32.Pedregosa, F. et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
  • 33.Nazari, M., Emami, H., Rabiei, R., Hosseini, A. & Rahmatizadeh, S. Detection of cardiovascular diseases using data mining approaches: Application of an ensemble-based model. Cogn. Comput.16 (5), 2264–2278. 10.1007/s12559-024-10306-z (2024). [Google Scholar]
  • 34.Arroyo, J. C. T. & Delima, A. J. P. An optimized neural network using genetic algorithm for cardiovascular disease prediction. J. Adv. Inf. Technol.13 (1), 95–99. 10.12720/jait.13.1.95-99 (2022). [Google Scholar]
  • 35.Lin, C. M. & Lin, Y. S. Utilizing a two-stage Taguchi method and artificial neural network for the precise forecasting of cardiovascular disease risk. Bioengineering10 (11), 1286 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Lin, C. M. & Lin, Y. S. TPTM-HANN-GA A novel hyperparameter optimization framework integrating the Taguchi method, an artificial neural network, and a genetic algorithm for the precise prediction of cardiovascular disease risk. Mathematics12 (9), 1303 (2024). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Our experiments utilized three datasets: two publicly available from Kaggle can be obtained following the specific citation, and one locally collected dataset, which can be freely accessed for academic use upon request. While the detailed code files can be freely accessed on Github through https://github.com/AyshaSaddiqa/CVC-detection.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES