Abstract
Background
Postoperative atrial fibrillation (POAF) occurs in 20–40% of patients undergoing coronary artery bypass grafting (CABG) and is associated with increased morbidity and mortality. Existing machine learning (ML) studies are limited by single‑center designs, small samples, and reliance on single algorithms.
Objective
To develop and internally validate a stacking ensemble ML model for predicting POAF after CABG, identify key predictors, and evaluate its incremental clinical value against conventional models.
Methods
A retrospective analysis of 563 CABG patients (29.5% POAF) from a single center was performed. After preprocessing, the dataset was split into training (n = 394) and validation (n = 169) sets. Features were selected using elastic net with stability selection (200 bootstrap resamples). Nine base ML algorithms and a stacking ensemble were built; hyperparameters were tuned via fivefold cross‑validation with overfitting controls. Model performance was assessed by discrimination, calibration, and decision curve analysis (DCA).
Results
Thirteen predictors were retained, with age, intraoperative phenylephrine use, and stroke history ranking highest. The stacking model achieved a validation AUC of 0.7425 and an F1 of 0.711. Logistic regression showed a comparable AUC (0.718, p = 0.09). The three clinical risk scores performed worse (AUCs 0.695, 0.662, 0.670; all p < 0.05). DCA revealed no net benefit below a threshold of 0.4, a marginal benefit (≈0.05) between 0.4 and 0.6, and no benefit above 0.6.
Conclusion
The stacking model provides only moderate discrimination for POAF after CABG, with negligible net clinical benefit confined to a narrow threshold range. It is not ready for clinical use without external validation. The identified predictors may generate hypotheses for mechanistic studies, but their clinical utility remains unproven.
Supplementary Information
The online version contains supplementary material available at 10.1186/s13019-026-04271-x.
Keywords: Coronary artery bypass grafting, Postoperative atrial fibrillation, Machine learning, Stacking integration model, Risk prediction
Introduction
Cardiovascular disease (CVD), especially coronary artery disease (CAD), is the leading cause of death worldwide [1]. Coronary Artery Bypass Grafting (CABG) is a cornerstone revascularization procedure for patients with severe coronary heart disease (CHD), especially those with multi-vessel lesions and complex coronary anatomy, and is widely recommended by major cardiology guidelines [2–5].
However, postoperative complications significantly impact patient outcomes following CABG. Postoperative atrial fibrillation (POAF) is the most common type of arrhythmia, with an incidence ranging from 20 to 40%, typically occurring within 2–4 days post-surgery [6–9]. Evidence suggests that POAF not only prolongs hospital stay and increases readmission risk but also significantly contributes to long-term adverse outcomes, including a 53% increased risk of cerebrovascular events (such as ischemic stroke) at 10 years postoperatively (hazard ratio (HR) = 1.53) and increased long-term all-cause mortality, posing a serious threat to patients’ quality of life [8, 10–12]. The need for long-term anticoagulation, repeated visits, and management of chronic complications associated with POAF places a substantial financial burden on global healthcare systems [13]. Therefore, developing precise, individualized prevention strategies based on identified risk factors is crucial to reduce unnecessary medical interventions and improve patient prognosis.
Previous studies have identified multiple potential predictors of POAF, such as age, preoperative hypertension, history of heart failure, obesity (BMI ≥ 30 kg/m2), and stroke history [14–18]. Traditional statistical methods, such as logistic regression (LR), have significant limitations in capturing complex nonlinear relationships and handling multicollinearity within high-dimensional clinical data, thereby restricting their predictive efficiency when numerous variables are involved [19, 20]. In recent years, machine learning (ML) technology has shown considerable promise in medical prediction [21]. Unlike traditional methods, ML does not require pre-set assumptions about variable relationships and can efficiently process large datasets through algorithmic iteration, accurately identifying nonlinear associations between variables, and significantly overcoming the technical bottlenecks of traditional methods [22]. However, current ML research on POAF after CABG often suffers from single-center designs, small sample sizes, and reliance on single algorithms (e.g., only random forests or support vector machines), failing to establish robust and generalizable prediction models. Crucially, many existing ML studies on POAF lack comprehensive comparisons with established clinical prediction models or conventional statistical approaches, hindering the assessment of their incremental predictive value.
Based on these considerations, this study aims to construct and internally validate an ML prediction model for POAF after CABG using single-center clinical data. We also seek to screen key baseline characteristics and perioperative indicators influencing POAF occurrence. This work intends to provide an evidence-based foundation for clinicians to formulate individualized prevention strategies, with the understanding that the ultimate clinical utility and superiority over existing methods require further rigorous comparative validation.
Materials and methods
Research design, scene and object
The data were derived from the electronic medical record system (EMR) and surgical anesthesia information system (AIMS) of the Affiliated Hospital of Xuzhou Medical University, covering the full dimension of demographics, perioperative management, laboratory examination and postoperative monitoring.
Ethics statement
This study was approved by the Ethics Committee of The Affiliated Hospital of Xuzhou Medical University (approval number: XYFY2025-KL154-01) and strictly followed the Declaration of Helsinki (2013 revised edition). All included data were de-identified. Due to the retrospective nature of this study, informed consent was exempted.
Patient inclusion and exclusion
From January 2021 to December 2024, 601 patients undergoing CABG were initially identified. Screening criteria for inclusion were: age ≥ 18 years old, preoperative sinus rhythm, clinical data completeness ≥ 90%, and postoperative hospitalization observation ≥ 48 h. Exclusion criteria were: history of preoperative atrial fibrillation, perioperative key index missing rate > 10%, or combined with other cardiac surgeries in the same period. Finally, 563 cases were included, with 166 cases developing new-onset POAF, resulting in an overall incidence rate of 29.48%.
Definition of POAF
For the purpose of this study, POAF was defined as new-onset acute atrial fibrillation occurring exclusively during the postoperative hospitalization period, detected via standard electrocardiogram (ECG) monitoring. Diagnostic criteria required the disappearance of P waves, replacement by f waves, an absolutely irregular QRS rhythm with unequal RR intervals, and an onset lasting more than 30 s [23, 24]. It is important to note that this definition is limited to in-hospital detection, and a standardized, continuous rhythm monitoring protocol was not uniformly applied across all patients.
Predictors, outcomes and data preprocessing
A total of 189 predictors were initially included, encompassing baseline characteristics (age, gender, body mass index (BMI), smoking history, alcohol history, NYHA grading, hypertension, diabetes, and other comorbidities), preoperative examination and medication (left atrial diameter, left ventricular ejection fraction (LVEF), platelets, hemoglobin, white blood cells, albumin, preoperative use of β-receptor blockers, etc.), and perioperative indicators (surgery duration, anesthesia duration, intraoperative blood loss, intraoperative use of vasoactive drugs such as phenylephrine, number of bypass vessels, etc.). However, due to the retrospective nature of this single-center study and inherent limitations of the electronic medical record system, certain critical perioperative factors, such as the use of Off-Pump Coronary Artery Bypass (OPCAB), aortic cross-clamp time, and cardiopulmonary bypass (CPB) time, were either not consistently recorded or had missing rates exceeding our data completeness criteria (> 10%), and thus were not included in the analysis. The primary outcome was the occurrence of POAF during postoperative hospitalization.
Data preprocessing steps included the following. Multiple imputation by chained equations (MICE) was employed to handle missing values, with the maximum number of imputation iterations set to 15. Variables exhibiting a missing rate > 10% were excluded. Crucially, the dataset was first stratified into a training set (n = 394) and a validation set (n = 169) at a 7:3 ratio, with the occurrence of POAF as the stratification factor. Subsequently, to address class imbalance within the training set, the POAF-positive samples were oversampled using the Synthetic Minority Oversampling Technique (SMOTE). SMOTE was applied exclusively to the training set to prevent data leakage and ensure an unbiased evaluation of model performance on unseen data.
Statistical methods and model construction and verification
Statistical analysis was performed using Python 3.12 and SPSS 26 software, with the test level α set at 0.05. Continuous variables were presented as "mean ± standard deviation (x ± s)", non-normal variables as "median (interquartile spacing) [M(Q1,Q3)]", and categorical variables as "number of cases (percentage) [n(%)]". For group comparisons, the independent sample t-test was used for continuous variables, the Mann–Whitney U test for non-normal variables, and the χ2 test for categorical variables (Fisher’s exact probability method was used when the expected frequency was < 5).
The "elastic net + stability selection" strategy was used for feature screening [25, 26], with 200 bootstrap resampling iterations,retaining, 13 key factors with a selection probability of ≥ 0.5. The importance of features was ranked by extreme gradient boosting (XGBoost). Nine basic machine learning models (Naïve Bayesian (NB), Random Forest (RF), Support Vector Machine (SVM), XGBoost, Light Gradient Boosting Machine (LightGBM), Adaptive Boost (AdaBoost), Categorical Boosting (CatBoost), Multilayer Perceptron (MLP), Gradient Boosting Decision Tree (GBDT)) and one Stacking ensemble model (using logistic regression as the meta-model and integrating the output of the basic models) were constructed. Hyperparameters were optimized using fivefold cross-validation combined with grid search. Given the limited sample size and the complexity of tree-based models, we reduced the risk of overfitting by imposing constraints on model complexity. Specifically, for Random Forest (RF), XGBoost, LightGBM, CatBoost, and Gradient Boosting Decision Tree (GBDT), we limited the maximum tree depth (max_depth ≤ 5) and increased the minimum number of samples required per leaf node. Furthermore, early stopping was applied during training for XGBoost, LightGBM, and CatBoost models, using the validation set performance to halt training when no improvement was observed for a predefined number of rounds. Model validation was evaluated from three aspects: discrimination (AUC, accuracy, recall, F1 value), calibration (calibration curve combined with Hosmer–Lemeshow test), and clinical practicability (average precision (AP) value calculated by Precision-Recall curve, and decision curve analysis (DCA) to determine the optimal intervention threshold). The SHapley Additive exPlanations (SHAP) summary plot was generated by analyzing the contribution of the key factors through SHAP values. Calibration was assessed graphically using a calibration plot (see Supplementary Fig. 1B).
Results
Baseline characteristics of the research subjects
In this study, 563 patients who underwent CABG were included, of whom 166 developed new-onset POAF, with an overall incidence of 29.48%. There were no statistical differences between the two groups in demographic characteristics, preoperative underlying diseases, echocardiography indicators, laboratory test results, and perioperative management indicators (P > 0.05), ensuring the reliability of subsequent model validation.
Detailed baseline characteristics
Demographic characteristics: 397 males (70.52%) and 166 females (29.48%); median age was 64.98 ± 7.83 years; BMI was 25.23 ± 3.31 kg/m2; 265 cases (47.07%) were current smokers.
Preoperative underlying diseases: 361 cases (64.12%) with hypertension, 204 cases (36.23%) with type 2 diabetes, 156 cases (27.71%) with cerebrovascular disease (including transient ischemic attack/stroke), and 30 cases (5.33%) with pulmonary hypertension. 218 patients (38.72%) were classified ≥ grade III by the New York Heart Association (NYHA).
Echocardiography indicators: left atrial inner diameter (LA) was 37.22 ± 4.47 mm; Left ventricular ejection fraction (LVEF) was 58.16 ± 8.05%.
Preoperative laboratory tests: platelet count (PLT) was 218.59 ± 62.25 × 10⁹/L; Hemoglobin (Hb) was 135.32 ± 15.20 g/L; Blood urea nitrogen (BUN) was 6.14 ± 4.07 mmol/L; albumin (ALB) was 41.66 ± 3.99 g/L.
Preoperative medication: 467 cases (82.95%) took β receptor blockers, 552 cases (98.05%) of antiplatelet drugs, 125 cases (22.20%) of diuretics, 537 cases (95.38%) of statins, and 153 cases (27.18%) of angiotensin II receptor antagonists (ARBs). Only 4 cases (0.71%) took digitalis preparations.
Perioperative indicators: The operative time was 317.33 ± 63.95 min; The duration of anesthesia was 370.28 ± 68.29 min. The intraoperative blood loss was 525.26 ± 445.59 mL. The intraoperative urine output was 1529.90 ± 831.04 mL; Vasoactive drug use: 40 cases (7.10%) received epinephrine, 78 cases (13.85%) of norepinephrine, and 216 cases (38.37%) of phenylephrine.
To further illustrate the distribution patterns of these key features across samples, a clustering heatmap and a feature density scatter plot are presented in Fig. 1. The heatmap (Fig. 1A) shows the hierarchical clustering of samples based on the 13 selected predictors, while the density plot (Fig. 1B) visualizes the distribution characteristics of each feature.
Fig. 1.
The ROC curves for machine learning models and the performances of all models in test cohort. (A). The ROC curves for machine learning models and the performances of all models in validation cohort (B)
Key predictive feature screening results
The "elastic network regularization + stability selection" strategy is used to screen POAF predictive features to balance the robustness of feature selection and model interpretability: the regularization strength of the elastic network is set α = 0.1 (to avoid overfitting), and finally 13 key predictors with a selection probability of ≥ 0.5 are retained.
Based on the XGBoost model, the importance of features (value range 0–1, higher value indicates greater predictive contribution) was ranked. The results showed that age (0.9800) was the most important predictor, followed by intraoperative phenylephrine use (0.9550) and stroke history (0.9250). The remaining variables, in descending order of importance, were: history of hypertension (0.8600), intraoperative calcium gluconate use (0.8050), number of bypass vessels (0.7950), smoking history (0.6950), preoperative ARB drug use (0.6900), intraoperative lidocaine use (0.6850), intraoperative nitrates use (0.6400), intraoperative sodium carbonate Ringer use (0.6400), preoperative platelet count (0.5400), and intraoperative hydroxyethyl starch use (0.5250). High-importance features (importance ≥ 0.8) covered demographic characteristics (age), perioperative drugs (phenylephrine), and underlying lesions (history of stroke and hypertension), which were highly consistent with the known pathogenesis of POAF (such as age-related myocardial degenerative changes and the effect of vasoactive drugs on cardiac electrophysiology), providing a subset of features with both clinical significance and statistical validity for subsequent model construction.
Machine learning model performance evaluation
In this study, 9 basic machine learning models and 1 Stacking ensemble model were constructed. Their performance was evaluated from three aspects: discrimination, calibration, and clinical practicability.
Machine learning models in the training cohort
Among them, the RF, XGBoost, and Stacking models demonstrated exceptionally high performance, with AUC, accuracy, recall, and F1 values reaching 1.000, indicating a perfect fit to the training data. While this suggests the models’ capacity to learn complex patterns within the training set, it also raises concerns about potential overfitting, especially given the relatively high dimensionality of the feature space. The performance of the LightGBM model was close to optimal (AUC = 0.9991, F1 = 0.9873). The NB model exhibited the lowest performance (AUC = 0.7722, F1 = 0.7091) (Table 2).
Table 2.
Performance indicators of each algorithm model on the training set
| model | AUC | Accuracy | Recall | F1 value | Nested AUC(Mean) | Nested AUC(Std) |
|---|---|---|---|---|---|---|
| LR | 0.785 | 0.715 | 0.719 | 0.717 | 0.752 | 0.041 |
| NB | 0.772 | 0.708 | 0.698 | 0.706 | 0.746 | 0.027 |
| RF | 0.868 | 0.773 | 0.781 | 0.775 | 0.758 | 0.048 |
| SVM | 0.859 | 0.780 | 0.763 | 0.777 | 0.740 | 0.051 |
| XGBoost | 0.945 | 0.874 | 0.874 | 0.874 | 0.763 | 0.034 |
| LightGBM | 0.887 | 0.809 | 0.820 | 0.811 | 0.770 | 0.035 |
| Adaboost | 0.805 | 0.717 | 0.712 | 0.716 | 0.741 | 0.038 |
| Catboost | 0.843 | 0.706 | 0.763 | 0.716 | 0.756 | 0.040 |
| MLP | 0.815 | 0.742 | 0.763 | 0.748 | 0.755 | 0.040 |
| GBDT | 0.916 | 0.834 | 0.849 | 0.837 | 0.755 | 0.040 |
| Stacking | 0.955 | 0.891 | 0.899 | 0.893 | - | - |
Logistic Regression(LR)
Naive Bayes (NB)
Random Forest (RF)
Support Vector Machine (SVM)
eXtreme Gradient Boosting (XGBoost)
Light Gradient Boosting Machine (LightGBM)
Adaptive Boosting (Adaboost)
Categorical Boosting (Catboost)
Multilayer Perceptron (MLP)
Gradient Boosting Decision Tree (GBDT)
Stacking Ensemble Learning (Stacking)
Machine learning models in the validation cohort
The generalization ability of the models was subsequently evaluated on the independent validation set (n = 169). A notable decrease in AUC was observed across all models when compared to their training set performance, with the AUC values dropping to approximately 0.74 for the best-performing models. This significant performance gap between the training and validation sets reflects the models’ predictive performance in more realistic clinical scenarios and highlights a potential issue of limited robustness and generalization, consistent with the concerns of overfitting observed in the training phase. Specifically, the GBDT model achieved the highest AUC (0.7494), followed by MLP (0.7466) and the Stacking integrated model (0.7425). The NB model continued to perform the worst (AUC = 0.7286). From a comprehensive performance perspective, the Stacking ensemble model demonstrated the highest F1 value (0.7165) on the validation set, achieving the best balance between accuracy (0.6741) and recall (0.7647), and showing comparable generalization stability to single foundational models (Table 3).
Table 3.
Performance indicators of each algorithm model on the validation set
| Model | AUC | Accuracy | Recall | F1 value | Nested AUC(Mean) | Nested AUC(Std) |
|---|---|---|---|---|---|---|
| LR | 0.731 | 0.640 | 0.723 | 0.667 | 0.751 | 0.041 |
| NB | 0.728 | 0.661 | 0.698 | 0.672 | 0.746 | 0.027 |
| RF | 0.737 | 0.670 | 0.740 | 0.690 | 0.758 | 0.048 |
| SVM | 0.745 | 0.665 | 0.672 | 0.667 | 0.741 | 0.051 |
| XGBoost | 0.736 | 0.678 | 0.740 | 0.696 | 0.763 | 0.034 |
| LightGBM | 0.734 | 0.661 | 0.731 | 0.682 | 0.770 | 0.035 |
| Adaboost | 0.719 | 0.661 | 0.731 | 0.682 | 0.741 | 0.038 |
| Catboost | 0.729 | 0.665 | 0.723 | 0.683 | 0.756 | 0.040 |
| MLP | 0.739 | 0.653 | 0.723 | 0.683 | 0.755 | 0.040 |
| GBDT | 0.746 | 0.644 | 0.740 | 0.680 | 0.755 | 0.041 |
| Stacking | 0.729 | 0.678 | 0.740 | 0.696 | - |
The comparison of Receiver Operating Characteristic curves (ROC curves) of the validation set (Fig. 2) showed that the curves of GBDT, MLP, and Stacking models were closer to the upper left corner, indicating a stronger ability to distinguish between POAF and non-POAF patients. In the PR curve (Fig. 3), the average precision of the Stacking model (AP = 0.7740) was the highest, demonstrating better recognition ability for positive samples in unbalanced data with a low proportion of POAF samples (29.48%).
Fig. 2.
The decision curve analysis plots of all models (A); the Precision-Recall curves of all models (B)
Fig. 3.
Heat map of feature distribution under sample clustering (A); scatter plot of feature density (B)
Model calibration evaluation
The Hosmer–Lemeshow test was used to evaluate the calibration degree of the model (i.e., the consistency between the predicted probability and the actual incidence of POAF). The results showed no systematic deviation between the predicted probability and the actual risk, indicating good calibration and reliable reflection of patients’ real POAF risk.
Clinical practicability evaluation of the model
Decision curve analysis (DCA) was performed to evaluate the net benefit of the models across a range of threshold probabilities (Fig. 3A). The Stacking model demonstrated no net benefit over the “treat all” strategy for threshold probabilities below 0.4. Between 0.4 and 0.6, the Stacking model showed a very small positive net benefit (approximately 0.05), which is only marginally higher than “treat all”. Above a threshold of 0.6, the net benefit of the Stacking model fell below that of “treat none”, indicating no clinical advantage. Overall, the DCA suggests that the clinical utility of the Stacking model is extremely limited and confined to a narrow intermediate threshold range. Therefore, the model is not ready for clinical decision-making without further validation and improvement.
Discussion
In this single‑center retrospective study, we developed and internally validated multiple machine learning models to predict new‑onset postoperative atrial fibrillation after coronary artery bypass grafting. The Stacking ensemble model achieved a moderate discriminative ability (AUC = 0.7425 on the validation set), acceptable calibration (Hosmer‑Lemeshow p = 0.34), and only a marginal net benefit on decision curve analysis, which was confined to a narrow threshold range (0.4–0.6). These results indicate that, while the model shows some promise as an exploratory risk stratification tool, its current performance does not support clinical deployment without substantial further validation and refinement.
A key addition in this study is the direct benchmarking of the machine learning models against conventional methods. Logistic regression using the same set of predictors achieved a validation AUC of 0.718, which was not significantly different from the Stacking model (p = 0.09, DeLong test). The Net Reclassification Improvement was 0.21 (95% CI 0.08–0.34) and the Integrated Discrimination Improvement was 0.03 (95% CI 0.01–0.05), indicating a small but statistically significant incremental value over logistic regression. Compared to age alone (AUC = 0.62), the Stacking model improved the AUC by 0.12 (p < 0.01). These findings suggest that machine learning offers only a modest advantage over simpler models in this dataset, and the clinical relevance of this improvement is uncertain given the marginal net benefit on decision curve analysis.
The initial models exhibited near‑perfect performance on the training set (AUC = 1.000 for several algorithms), which was a clear sign of overfitting. To address this, we applied nested cross‑validation (5 outer, 3 inner folds), limited tree depth (max_depth ≤ 5), increased the minimum samples per leaf, and used early stopping for gradient‑boosted models. The nested cross‑validated AUC for the Stacking model was 0.73 ± 0.04, which closely matches the validation set performance (0.7425). The learning curves (Supplementary Fig. 1) showed that validation AUC continued to increase with training sample size without reaching a plateau, indicating that the current dataset (n = 563, 166 events) is still underpowered for stable high‑dimensional machine learning modeling. Larger multi‑center cohorts are needed to achieve robust performance.
We used an algorithm‑driven feature selection pipeline (elastic net combined with stability selection, 200 bootstrap resamples). With a selection probability threshold of ≥ 0.5, thirteen predictors were retained: age, intraoperative phenylephrine use, stroke history, history of hypertension, intraoperative calcium gluconate use, number of graft vessels, smoking history, preoperative ARB use, intraoperative lidocaine use, intraoperative nitrates use, intraoperative sodium carbonate Ringer solution use, preoperative platelet count, and intraoperative hydroxyethyl starch use. Some of the retained predictors did not show significant univariable differences between POAF and non‑POAF groups (Table 1). However, univariable significance is not a prerequisite for multivariable predictive utility in machine learning, and manual pre‑selection based on p‑values would introduce bias. The fully data‑driven approach preserves objectivity, but it also means that some associations may be spurious or institution‑specific. Therefore, these predictors should be considered hypothesis‑generating rather than clinically validated (Table 2).
Table 1.
Characteristic variables
| Non POAF | POAF | P-value | |
|---|---|---|---|
| Age (years) | 64.10 ± 7.72 | 67.09 ± 7.53 | P < 0.01 |
| Smoking | 189(47.61%) | 76(45.78%) | 0.762 |
| Hypertension | 259(65.24%) | 102(61.45%) | 0.448 |
| CVD | 116(29.22%) | 40(24.10%) | 0.256 |
| PLT | 222.09 ± 61.84 | 210.37 ± 62.65 | 0.043 |
| ARB | 110(27.71%) | 43(25.90%) | 0.738 |
| NGV | 3.35 ± 0.77 | 3.27 ± 0.66 | 0.248 |
| Intraoperative nitrates | 370(93.20%) | 151(90.96%) | 0.457 |
| Lidocaine | 120(30.23%) | 49(29.52%) | 0.947 |
| Intraoperative phenylephrine | 161(40.55%) | 55(33.13%) | 0.120 |
| Intraoperative calcium gluconate | 106(26.70%) | 35(20.08%) | 0.195 |
| Intraoperative RSC | 210(52.90%) | 80(48.19%) | 0.355 |
| Intraoperative Hydroxyethyl starch | 193(48.61%) | 74(44.58%) | 0.434 |
Data are presented as median (interquartile range) or number (%)
Cerebrovascular Disease (CVD)
Ringer’s Solution with Sodium Carbonate (RSC)
Angiotensin II Receptor Blocker (ARB)
Platelet Count (PLT)
Number of Graft Vessels (NGV)
The decision curve analysis (Fig. 3A) showed that the Stacking model provides no net benefit over a “treat all” strategy for threshold probabilities below 0.4. Between 0.4 and 0.6, the net benefit was very small (approximately 0.05) and only marginally higher than “treat all”. Above 0.6, the net benefit fell below that of “treat none”. Thus, the model’s potential clinical utility is extremely limited and confined to a narrow intermediate‑risk range. We therefore conclude that the model is not ready for routine clinical use. Any future clinical translation would require prospective validation in a multi‑center setting and demonstration of meaningful improvement in patient outcomes (Table 3).
Several important limitations of this study should be acknowledged. First, the single‑center retrospective design limits external validity and may introduce selection bias; internal validation cannot substitute for external validation on independent cohorts. Second, critical perioperative factors known to influence POAF, such as cardiopulmonary bypass time, aortic cross‑clamp time, and off‑pump CABG status, were either not consistently recorded or had missing rates exceeding 10%, and were therefore excluded, which compromises the completeness of the predictor space. Third, with 563 patients and only 166 POAF events, the sample size is relatively small for high‑dimensional machine learning modeling, and the learning curves indicate that more data are needed to achieve stable performance. Fourth, despite applying regularization, nested cross‑validation, and early stopping, the gap between training and validation performance (nested CV AUC 0.73 vs. original training AUC 0.955 for Stacking) confirms that overfitting remains a concern; simpler models such as logistic regression performed similarly, suggesting that the added complexity of machine learning may not be justified with the current dataset. Fifth, the model has not been tested on any external dataset, so it should not be used for clinical decision‑making; a prospective multi‑center study is planned to validate and refine the model. Sixth, while we used bootstrap stability selection, some retained predictors (e.g., intraoperative calcium gluconate, lidocaine) showed low univariable significance and may be institution‑specific or reflect confounding by indication. The mechanistic interpretations of these variables are speculative and not directly supported by our predictive analysis; therefore, they have been omitted from this discussion.
Several previous studies have applied machine learning to predict POAF after cardiac surgery [27–33]. Similar to our findings, most report AUCs in the range of 0.70–0.80, with age consistently identified as a major predictor. However, many of those studies also suffered from single‑center designs, small sample sizes, and lack of external validation. Our study adds to this literature by providing a systematic benchmark against logistic regression and clinical risk scores, and by transparently reporting overfitting and the limited net benefit on decision curve analysis. We believe that future studies should prioritize multi‑center prospective data collection and rigorous external validation over algorithmic complexity.
Based on the limitations identified, we are planning a prospective multi‑center registry that will include at least 1,500 patients from three cardiac surgery centers. This registry will ensure complete collection of cardiopulmonary bypass time, cross‑clamp time, off‑pump CABG status, and standardized POAF definitions. The current model will be re‑evaluated on this external cohort, and if performance remains modest, we will consider simpler scores or alternative modeling strategies (e.g., penalized regression) that may be more robust in smaller samples.
Conclusion
In summary, this study successfully developed and internally validated a Stacking ensemble model for predicting the risk of POAF after CABG, utilizing machine learning algorithms. The model demonstrated moderate discrimination and acceptable calibration, but its clinical utility is very limited and requires external validation before any clinical application. Its key predictors, particularly perioperative drug factors, offer novel insights into the pathophysiological mechanisms of POAF and potential clinical prevention strategies. Despite the aforementioned limitations, including the moderate discriminative performance (AUC≈0.74) in the validation set and the lack of benchmarking against established models, this study presents a promising tool for early precise risk stratification of POAF. It lays a solid foundation for future large-scale prospective studies that must include rigorous external validation and comprehensive comparative analyses to determine its true incremental clinical value and applicability.
Supplementary Information
Author contributions
Yang Zhang: Manuscript writing; Zhihan Zhang: Data collection; Hu Zhang: Data collection; Yun Lu: Data collection; Zhu Wang: Machine algorithm collaborative operation; Jun Wei: Manuscript review and revision; Shoujie Feng: Manuscript review and revision.
Funding
This study was supported by Construction Project of High Level Hospital of Jiangsu Province (GSPJS202406 and GSPJS202505), Key Project of Jiangsu Provincial Health Commission (K2024014), Construction Project of High-Level Hospital of Jiangsu Province (GSPJS202404).
Data availability
No datasets were generated or analysed during the current study.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yang Zhang and Zhihan Zhang contribute equally to this study.
Contributor Information
Zhu Wang, Email: wzxystudent@163.com.
Jun Wei, Email: weijunseu@outlook.com.
Shoujie Feng, Email: fengshoujie@xzhmu.edu.cn.
References
- 1.(2020) Global burden of 369 diseases and injuries in 204 countries and territories, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet 396(10258):1204–1222 [DOI] [PMC free article] [PubMed]
- 2.Falk V, Baumgartner H, Bax JJ, De Bonis M, Hamm C, Holm PJ, et al. 2017 ESC/EACTS guidelines for the management of valvular heart disease. Eur J Cardiothorac Surg. 2017;52(4):616–64. [DOI] [PubMed] [Google Scholar]
- 3.Fuster V, Rydén LE, Cannom DS, Crijns HJ, Curtis AB, Ellenbogen KA, et al. 2011 ACCF/AHA/HRS focused updates incorporated into the ACC/AHA/ESC 2006 guidelines for the management of patients with atrial fibrillation: a report of the American College of Cardiology Foundation/American Heart Association Task Force on practice guidelines. Circulation. 2011;123(10):e269–367. [DOI] [PubMed] [Google Scholar]
- 4.Kirchhof P, Benussi S, Kotecha D, Ahlsson A, Atar D, Casadei B, et al. 2016 ESC guidelines for the management of atrial fibrillation developed in collaboration with EACTS. Eur Heart J. 2016;37(38):2893–962. [DOI] [PubMed] [Google Scholar]
- 5.Rodriguez F, Mahaffey KW. Management of patients with NSTE-ACS: a comparison of the recent AHA/ACC and ESC guidelines. J Am Coll Cardiol. 2016;68(3):313–21. [DOI] [PubMed] [Google Scholar]
- 6.Caldonazo T, Kirov H, Rahouma M, Robinson NB, Demetres M, Gaudino M, et al. Atrial fibrillation after cardiac surgery: a systematic review and meta-analysis. J Thorac Cardiovasc Surg. 2023;165(1):94–103. [DOI] [PubMed] [Google Scholar]
- 7.Chugh SS, Havmoeller R, Narayanan K, Singh D, Rienstra M, Benjamin EJ, et al. Worldwide epidemiology of atrial fibrillation: a Global Burden of Disease 2010 Study. Circulation. 2014;129(8):837–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Eikelboom R, Sanjanwala R, Le ML, Yamashita MH, Arora RC. Postoperative atrial fibrillation after cardiac surgery: a systematic review and meta-analysis. Ann Thorac Surg. 2021;111(2):544–54. [DOI] [PubMed] [Google Scholar]
- 9.Go AS, Hylek EM, Phillips KA, Chang Y, Henault LE, Selby JV, et al. Prevalence of diagnosed atrial fibrillation in adults: national implications for rhythm management and stroke prevention: the AnTicoagulation and Risk Factors in Atrial Fibrillation (ATRIA) Study. JAMA. 2001;285(18):2370–5. [DOI] [PubMed] [Google Scholar]
- 10.Benedetto U, Gaudino MF, Dimagli A, Gerry S, Gray A, Lees B, et al. Postoperative atrial fibrillation and long-term risk of stroke after isolated coronary artery bypass graft surgery. Circulation. 2020;142(14):1320–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Engin M, Aydın C. Investigation of the effect of HATCH Score and coronary artery disease complexity on atrial fibrillation after on-pump coronary artery bypass graft surgery. Med Princ Pract. 2021;30(1):45–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lotter K, Yadav S, Saxena P, Vangaveti V, John B. Predictors of atrial fibrillation post coronary artery bypass graft surgery: new scoring system. Open Heart. 2023. 10.1136/openhrt-2023-002284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Sacher F, Wright M, Tedrow UB, O’Neill MD, Jais P, Hocini M, et al. Wolff-Parkinson-White ablation after a prior failure: A 7-year multicentre experience. Europace. 2010;12(6):835–41. [DOI] [PubMed] [Google Scholar]
- 14.Aranki SF, Shaw DP, Adams DH, Rizzo RJ, Couper GS, VanderVliet M, et al. Predictors of atrial fibrillation after coronary artery surgery: Current trends and impact on hospital resources. Circ. 1996;94(3):390–7. [DOI] [PubMed] [Google Scholar]
- 15.Frendl G, Sodickson AC, Chung MK, Waldo AL, Gersh BJ, Tisdale JE, et al. 2014 AATS guidelines for the prevention and management of perioperative atrial fibrillation and flutter for thoracic surgical procedures. Executive summary. J Thorac Cardiovasc Surg. 2014;148(3):772–91. [DOI] [PubMed] [Google Scholar]
- 16.Mariscalco G, Biancari F, Zanobini M, Cottini M, Piffaretti G, Saccocci M, et al. Bedside tool for predicting the risk of postoperative atrial fibrillation after cardiac surgery: The POAF score. J Am Heart Assoc. 2014;3(2):e000752. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Shen J, Lall S, Zheng V, Buckley P, Damiano RJ Jr., Schuessler RB. The persistent problem of new-onset postoperative atrial fibrillation: A single-institution experience over two decades. J Thorac Cardiovasc Surg. 2011;141(2):559–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Zacharias A, Schwann TA, Riordan CJ, Durham SJ, Shah AS, Habib RH. Obesity and risk of new-onset atrial fibrillation after cardiac surgery. Circ. 2005;112(21):3247–55. [DOI] [PubMed] [Google Scholar]
- 19.Hill NR, Ayoubkhani D, McEwan P, Sugrue DM, Farooqui U, Lister S, et al. Predicting atrial fibrillation in primary care using machine learning. PLoS ONE. 2019;14(11):e0224582. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hosmer DW Jr, Lemeshow S, Sturdivant RX. Applied logistic regression. John Wiley & Sons; 2013.
- 21.Deo RC. Machine learning in medicine. Circ. 2015;132(20):1920–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Ren X, Mi Z, Georgopoulos PG. Comparison of machine learning and land use regression for fine scale spatiotemporal estimation of ambient air pollution: Modeling ozone concentrations across the contiguous United States. Environ Int. 2020;142:105827. [DOI] [PubMed] [Google Scholar]
- 23.Bidar E, Bramer S, Maesen B, Maessen JG, Schotten U. Post-operative atrial fibrillation - pathophysiology, treatment and prevention. J Atr Fibrillation. 2013;5(6):781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Hindricks G, Potpara T, Dagres N, Arbelo E, Bax JJ, Blomström-Lundqvist C, et al. 2020 ESC guidelines for the diagnosis and management of atrial fibrillation developed in collaboration with the European Association for Cardio-Thoracic Surgery (EACTS): the Task Force for the diagnosis and management of atrial fibrillation of the European Society of Cardiology (ESC) developed with the special contribution of the European Heart Rhythm Association (EHRA) of the ESC. Eur Heart J. 2021;42(5):373–498. [DOI] [PubMed] [Google Scholar]
- 25.Meinshausen N, Bühlmann P. Stability selection. J R Stat Soc B Stat Methodol. 2010;72(4):417–73. [Google Scholar]
- 26.Zou H, Hastie T. Regularization and variable selection via the elastic net. J R Stat Soc B Stat Methodol. 2005;67(2):301–20. [Google Scholar]
- 27.Borde D, Gandhe U, Hargave N, Pandey K, Mathew M, Joshi S. Prediction of postoperative atrial fibrillation after coronary artery bypass grafting surgery: is CHA 2 DS 2 -VASc score useful? Ann Card Anaesth. 2014;17(3):182–7. [DOI] [PubMed] [Google Scholar]
- 28.El-Sherbini AH, Shah A, Cheng R, Elsebaie A, Harby AA, Redfearn D, et al. Machine learning for predicting postoperative atrial fibrillation after cardiac surgery: a scoping review of current literature. Am J Cardiol. 2023;209:66–75. [DOI] [PubMed] [Google Scholar]
- 29.Haghjoo M, Basiri H, Salek M, Sadr-Ameli MA, Kargar F, Raissi K, et al. Predictors of postoperative atrial fibrillation after coronary artery bypass graft surgery. Indian Pacing Electrophysiol J. 2008;8(2):94–101. [PMC free article] [PubMed] [Google Scholar]
- 30.Parise O, Parise G, Vaidyanathan A, Occhipinti M, Gharaviri A, Tetta C, et al. Machine learning to identify patients at risk of developing new-onset atrial fibrillation after coronary artery bypass. J Cardiovasc Dev Dis. 2023. 10.3390/jcdd10020082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Ranucci M, Castelvecchio S, Menicanti L, Frigiola A, Pelissero G. Risk of assessing mortality risk in elective cardiac operations: age, creatinine, ejection fraction, and the law of parsimony. Circulation. 2009;119(24):3053–61. [DOI] [PubMed] [Google Scholar]
- 32.Taylan G, Gök M, Kurtul A, Uslu A, Küp A, Demir S, et al. Integrating the left atrium diameter to improve the predictive ability of the age, creatinine, and ejection fraction score for atrial fibrillation recurrence after cryoballoon ablation. Anatol J Cardiol. 2023;27(10):567–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Tsang TS, Barnes ME, Gersh BJ, Bailey KR, Seward JB. Left atrial volume as a morphophysiologic expression of left ventricular diastolic dysfunction and relation to cardiovascular risk burden. Am J Cardiol. 2002;90(12):1284–9. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
No datasets were generated or analysed during the current study.



