Abstract
Background
Atrial fibrillation (AF) increases the risk of cerebrovascular accidents (CVA), underscoring the need for accurate risk prediction. Traditional tools like CHA₂DS₂-VASc use limited clinical variables, often yielding imprecise estimates. Machine learning (ML) can offer potential for improved prediction by integrating broader clinical data. This study assesses the methods, performance, and real-world applicability of ML models for CVA prediction in AF.
Methods
A systematic search of PubMed, Embase, and Web of Science was conducted in September 2024. A random-effects meta-analysis on eligible studies pooled the predictive performance (AUROC) of ML models and the CHA₂DS₂-VASc score. Comparative analyses were conducted between the two, alongside subgroup analyses by ML algorithm type and validation strategy. This study was prospectively registered in PROSPERO (CRD42024539628).
Results
Seventeen studies were included. Tree-based models were most frequently used, followed by logistic regression and neural networks. The pooled AUROC for ML models was 0.66 (95% CI: 0.64–0.68), comparable to CHA₂DS₂-VASc (0.64; 95% CI: 0.62–0.66; P = 0.327). Sensitivity analyses confirmed robustness of estimates, and no significant differences were found between algorithm types. Models evaluated only on training data had significantly higher AUROCs than those using validation sets (P = 0.002), suggesting performance inflation. Moreover, only two studies employed external validation, highlighting a critical gap in methodological rigor. Notably, heterogeneity arising from diversity in model designs, populations, outcomes, and CVA prevalence (as identified by meta-regression) limited the generalizability of our findings.
Conclusion
Current ML models perform comparably to traditional tools for CVA risk prediction in AF patients; however, clinical applicability warrants methodological standardization and external validation in diverse populations.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12872-026-05508-2.
Keywords: Atrial fibrillation, Stroke, Cerebrovascular accident, Risk prediction, Machine learning
Background
Cerebrovascular accidents (CVA), commonly known as stroke, are one of the leading causes of death and disability worldwide [1]. Atrial fibrillation (AF) is the culprit for 15–20% of strokes. In older adults, this proportion increases, with AF contributing to nearly one-third of all strokes [2]. Currently, traditional risk stratification tools, such as the CHA₂DS₂-VASc score, are widely adopted into clinical practice for stroke risk prediction in patients with AF [3] and are endorsed by major guidelines, including the European Society of Cardiology (ESC) and the American Heart Association (AHA) [4–6]. Despite its widespread application, the major weakness of the CHA₂DS₂-VASc score is that it is based on a limited number of clinical variables. This may lead to oversimplifying stroke risk, not fully reflecting the complexity of the unique profile of the individual patient. Moreover, its precision has been diminished in several subpopulations, including younger AF patients or those with multiple comorbidities [7], with an over- or underestimation of stroke risk because of that imbalance [8, 9]. Incorporating a broader range of clinical data, genetic information, and biomarkers via machine learning (ML) could significantly improve stroke risk prediction, allowing for more tailored treatment strategies in AF management. As clinical practice increasingly integrates precision medicine, ML-based models offer the potential to refine decision-making, particularly in identifying high-risk AF patients who may benefit from early anticoagulation therapy [10, 11].
ML can analyze complex datasets for the detection of hidden patterns; thus, personalized and precise stroke risk assessment can be provided, optimizing preventive strategies in AF patients. Application of ML has already shown promise in stroke prediction among populations. Several ML algorithms, such as decision trees, support vector machines, and deep learning models, have been applied in observational studies for the prediction of stroke based on factors involving hypertension, diabetes, and prior cardiovascular events [12–18]. Given these successes, ML might find its place in an AF population as well. For instance, a study utilizing convolutional neural networks (CNN) and long short-term memory (LSTM) models demonstrated superior performance in AF detection, potentially leading to earlier intervention and better stroke prevention outcomes [19]. Importantly, ML models can incorporate AF-specific data, leading to more accurate and tailored risk stratification compared to traditional tools like the CHA₂DS₂-VASc score.
Despite the growing interest in applying machine learning for stroke prediction, there remains a lack of consolidated evidence focused on its role in AF populations. Numerous studies have independently explored different algorithms and datasets, yet the results remain fragmented, often using heterogeneous methods and patient cohorts. This systematic review and meta-analysis synthesize the current literature to assess the predictive performance and clinical applicability of ML models in this context. A central aim is to compare their predictive performance against the CHA₂DS₂-VASc score, the current standard for stroke risk stratification in AF. By identifying the most promising approaches and common limitations, this review provides a clearer picture of the current landscape and informs directions for future research and model development.
Methods
Study selection
This systematic review adhered to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines [20] (Table S1) and was prospectively registered in PROSPERO (CRD42024539628). While the original protocol specified a qualitative synthesis, the decision to conduct a meta-analysis was made after identifying a sufficient number of studies reporting AUROC with comparable outcomes and predictors.
We conducted a comprehensive literature search in PubMed, Embase, and Web of Science databases on September 18, 2024. The search strategy combined terms related to atrial fibrillation, cerebrovascular accidents, and ML, with the full strategy provided in Table S2.
Inclusion criteria encompassed original full-text articles in peer-reviewed journals that focused on the general population of atrial fibrillation patients and developed or validated ML models for predicting ischemic CVA. Ischemic CVA was defined to include ischemic stroke, transient ischemic attack (TIA), or systemic embolism, as specified by the original studies. Primary intracerebral hemorrhage was not considered an eligible outcome unless it was part of a composite endpoint that included ischemic CVA. For studies that reported broader or non-specific “stroke” outcomes, we extracted and reported the authors’ exact outcome definitions. We excluded studies that primarily focused on hemorrhagic events, used only traditional statistical methods without ML components, focused solely on feature selection without model development, or were review articles, case reports, or conference abstracts.
Two reviewers independently screened titles and abstracts of all identified studies. Full texts of potentially eligible studies were then assessed by the same reviewers. Any disagreements were resolved through discussion with a third reviewer.
Data extraction
Data extraction was performed independently by two reviewers using a standardized form. Extracted information included study characteristics (author, year, country, sample size, follow-up duration), patient characteristics (age, sex, inclusion/exclusion criteria), ML algorithms trained, methods used for training and validation, and comparisons with traditional risk scores. Any disagreements were resolved through discussion.
Risk of bias
The included studies were assessed using the PROBAST (Prediction model Risk Of Bias ASsessment Tool) [21] for risk of bias (ROB). PROBAST is specifically designed to evaluate the ROB and applicability of studies that develop, validate, or update prediction models for diagnosis and prognosis. It consists of four domains: participants, predictors, outcome, and analysis. Two reviewers independently handled the ROB assessment, with disagreements being resolved through consensus or consultation with a senior reviewer.
Meta-analysis
We conducted a meta-analysis of studies reporting the area under the receiver operating characteristic curve (AUROC) for prediction models. To enhance comparability across studies, we restricted our meta-analysis to models developed using commonly available demographic and clinical variables (e.g., age, sex, comorbidities) and excluded models relying on imaging, blood biomarkers, or device-derived features. All analyses were performed using a random-effects model with the Metafor package in R (version 4.5.0; R Foundation for Statistical Computing, Vienna, Austria) [22]. Since AUROC values are constrained between 0.5 and 1.0, they were logit-transformed to linearize the scale for meta-analytic purposes using the formula:
![]() |
All analyses were performed on the logit-transformed data, and only for forest plot demonstrations, the values were back-transformed. When a study developed more than one prediction model (e.g., applying different ML algorithms or study populations), each model was considered as an independent entry in the analysis. AUROCs of the CHA₂DS₂-VASc scores were pooled separately and then compared with the pooled AUROCs of ML models. We quantified heterogeneity using I2 and τ2 statistics for all meta-analyses.
Quantitative subgroup analyses were conducted according to the general class of ML algorithms, including support vector machines (SVM), logistic regression (LR), neural networks (NN), and tree-based methods (e.g., random forests, gradient boosting machines, and decision trees). Subgroup differences were evaluated using a Q-test for heterogeneity. In addition, an exploratory subgroup analysis was performed according to whether studies reported performance only on the training dataset or also on a separate validation set (internal or external). For this validation-strategy comparison, we did not apply formal statistical tests; instead, subgroup-specific pooled AUROCs were reported descriptively to illustrate patterns in performance and potential sources of heterogeneity, rather than to support inferential comparisons.
Several sensitivity analyses were conducted to assess the robustness of our findings. First, we restricted the analysis to studies with a low risk of bias in applicability, thereby excluding those with heterogeneous outcome definitions (e.g., combining ischemic CVA with systemic embolism) to improve outcome comparability. Second, to address potential dependency bias and reduce the risk of underestimating uncertainty or inflating precision due to including multiple models per study, we repeated the meta-analysis using only the best-performing ML model from each study. Third, we examined the within-study difference in discriminatory performance between ML models and the CHA₂DS₂-VASc score. For studies reporting AUROCs for both, we calculated the difference (ΔAUROC = AUROC_ML − AUROC_CHA₂DS₂-VASc) using the best-performing ML model. These differences were then pooled using a random-effects meta-analysis, enabling a more direct comparison while minimizing the confounding effect of between-study heterogeneity.
Pooled estimates from the top-performing models were used to perform univariate meta-regression to examine study-level characteristics, including CVA prevalence and sample size, as potential sources of heterogeneity affecting model performance. Small-study effects were explored by visual inspection of funnel plots for (i) the within-study difference in AUROC between ML models and CHA₂DS₂-VASc (ΔAUROC), (ii) the AUROC of the best-performing ML model from each study, and (iii) the AUROC of the CHA₂DS₂-VASc score.
Results
Study selection
Our initial search strategy identified 2,677 articles. After removing duplicates and screening titles and abstracts, 40 articles were selected for full-text review. We excluded 23 articles based on the following criteria: not being a full-length article (n = 1), not focusing on a general AF population (n = 6), not estimating CVA occurrence risk (n = 7), and using non-ML models (n = 4). Finally, 5 articles were excluded for focusing on identifying predictive features rather than developing ML models for CVA prediction [23–27]. Ultimately, 17 articles met our eligibility criteria (Fig. 1).
Fig. 1.

PRISMA flowchart illustrating the study selection process. AF: atrial fibrillation; CVA: cerebrovascular accidents; ML: machine learning
Study characteristics
As demonstrated in Table 1, all studies were published between 2015 and 2024. The 17 included studies utilized data from diverse populations: United States (n = 3) [28–30], China (n = 3) [31–33], Japan (n = 3) [34–36], Korea (n = 1) [37], Australia (n = 1) [10), Italy (n = 1) [38], United Kingdom (n = 1) [38], India (n = 1) [40], and international registries (n = 4) like ORBIT-AF, GARFIELD-AF, APHRS-AF registries [11, 40–42]. The review encompassed 913,570 AF patients, with a mean proportion of 42.54% females. The mean frequency of events across studies was 7.82%. However, it is important to note that the definition of events varied across studies: Most studies modelled ischemic stroke as the primary outcome [10, 11, 31–34, 37–39], while others used broader ischemic cerebrovascular outcomes that included TIA and/or systemic embolism [36, 41, 42]. One study included a secondary dataset in which the outcome was defined as primary ischemic stroke or secondary hemorrhagic transformation of an ischemic stroke or systemic embolism [42]. Moreover, the inclusion criteria for AF patients were heterogeneous, with some studies being more stringent in excluding patients with valvular AF [10, 34–36, 38, 40], taking anticoagulants [29, 31, 32, 36], or those with coexisting cardiovascular conditions. In contrast, one study investigated exclusively patients taking vitamin K antagonists [41], and two others reported results for both groups of patients with or without anticoagulant therapy [34, 38].
Table 1.
Core features of the reviewed studies
| Country (Database) |
Sampling Criteria; Type of Stroke |
Sample Size (N); F/U (M) | Female (%); Mean Age (Y) | Input Features (N) | |
|---|---|---|---|---|---|
| Letham et al. 2015 [28] | USA (Medicare beneficiaries dataset) |
Inclusion: AF patients; N/S |
12,586; 12 |
66.4%; N/S |
N/S (4148) |
| Li et al. 2016a [31] | China (CAFR registry) |
Inclusion: AF patients Exclusion: on warfarin or RFA treatment; Ischemic |
1864; 24 |
N/S; N/S |
Demographics, clinicals laboratory tests, and life styles (107) |
| Li et al. 2016b [32] | China (CAFR registry) |
Inclusion: AF patients Exclusion: on warfarin or RFA treatment; Ischemic |
1864; 24 |
N/S; N/S |
Demographics, clinicals laboratory tests, and life styles (167) |
| Leonarduzzi et al. 2018 [34] | Japan |
Inclusion: AF patients (more than 1 year, no sinus rhythm, no planned sinus rhythm restoration) Exclusion: complete AV block, sustained ventricular tachycardia, ventricular ectopy > 5%, cardiac pacemakers, paroxysmal AF, rhythm-control drugs, ….; Ischemic |
173; 47 ± 35 |
N/S; N/S |
RR intervals dynamics in 24-h Holter records (characterized by the scattering transform into 4 final features) |
| Han et al. 2019 [29] | USA |
Inclusion: AF patients having 30 consecutive days of remote monitoring preceding the index stroke Exclusion: on oral anticoagulation.; Excluded TIA or stroke that is due to occlusion of the carotid artery or its branches |
3185; N/S |
Cases: 98.6%; 69.0 ± 11.0 Controls: 98.5%; 67.6 ± 10.3 |
Daily AF burden signature in 30 days. (the cumulative amount of time during a day that patient meets AF detection criteria, using remote monitoring), |
| Goto et al. 2020 [41] | International (ORBIT-AF and GARFIELD-AF registries) |
Inclusion: newly diagnosed AF patients treated with vitamin K antagonists, with 3 PT-INR over the first 30 days after prescription; Ischemic Stroke, TIA or Systemic Embolism |
4708; 12 |
44.3%; 72.1 ± 9.9 |
PT-INR measurement patterns |
| Loring et al. 2020 [42] | International (ORBIT-AF registry) |
Inclusion: AF patients; Stroke |
22,760; 24 |
42%; 73 (65–80) |
Demographics and clinicals |
| International (GARFIELD-AF registry) |
Inclusion: newly diagnosed AF patients having at least one additional risk factor for stroke; Primary ischemic or secondary hemorrhagic ischemic stroke, or systemic embolism |
52,032; 24 |
44%; 71 (63–78) |
||
| Watanabe et al. 2021 [35] | Japan (J-RHYTHM Registry) |
Inclusion: nonvalvular AF patients, Exclusion: patients with mitral valve stenosis and/or prosthetic heart valve history; Ischemic strokes, TIA, or systemic embolisms |
7406; 24 |
29.2%; 69.8 ± 10.0 |
Demographics, clinicals, and laboratory tests (42) |
| Jung et al. 2022 [37] | Korea (KNHIS registry) |
Inclusion: AF patients; Ischemic |
754,949; 60 |
Cases: 50.3%; 71.5 ± 9.5 Controls: 40.8%; 64.6 ± 13.3 |
Demographics, clinicals (65) |
| Lu et al. 2022 [10] | Australia |
Inclusion: nonvalvular AF patients, Exclusion: patients with mitral valve stenosis and/or prosthetic heart valve history; Ischemic |
9670; 12 |
46%; 76.9 |
Clinical, and laboratory tests rural residence, specialty admission (Cardiology), and length of hospital stay |
| Nishi et al. 2022 [36] | Japan (Fushimi AF registry) |
Inclusion: nonvalvular AF patients Exclusion: on anti-coagulation therapy; Cerebral infarction, TIA, Systemic thromboembolism |
1757; Median: 4.5 |
43.1%; 73 (66–81) |
Demographic, clinicals, laboratory tests, chest X-ray, and echocardiography (168) |
| Li et al. 2023 [33] | China |
Inclusion: AF patients aged 45–85 years Exclusion: neuropsychiatric or ophthalmic disorders; Ischemic |
250; 12 |
Cases: N/S; 60.1 ± 9.1 Controls: N/S; 65.2 ± 9.8 |
Fundus images at 548, 605, and 810 nm wavelengths |
| Bernardini et al. 2024 [38] | Italy (START-Registry) |
Inclusion: nonvalvular AF patients treated with oral anticoagulation Exclusion: life expectancy < 6 months, refusal of the patient to undergo follow-up visits; Ischemic |
11,078; 18 |
45.7%; 77 (18–99) |
Demographic, clinicals, laboratory tests, and therapeutic range expected (in cases of treatment with VKA) |
| Chen et al. 2024 [40] | India (KERALA-AF registry) Asia (APHRS-AF registry) |
Inclusion: non-valvular AF patients; Ischemic |
2101; 12 |
46.5%; 68 (60–76) |
Demographic, clinicals, laboratory tests, and imaging features |
| Gao et al. 2024 [30] | USA |
Inclusion: AF patients aged > = 18 years; N/S |
4716; 24 |
43.1%; N/S |
Demographic, clinicals, laboratory tests, and CHA2DS2-VASc features (445) |
| Lu et al. 2024 [11] | International (GLORIA-AF registry) |
Inclusion: AF patients; N/S |
25,656; 12 |
44.8%; 70.3 |
Demographics, clinicals, laboratory tests (110) |
| Papadopoulou et al. 2024 [39] | UK (UK-Biobank) |
Inclusion: AF patients; Ischemic |
6300; N/S |
N/S; N/S |
Demographics, clinicals, laboratory tests, phenotypes, and lifestyle (129) |
AF Atrial fibrillation, APHRS-AF registry Asia Pacific Heart Rhythm Society Atrial Fibrillation registry, CAFR registry Chinese Atrial Fibrillation Registry, GARFIELD-AF registry Global Anticoagulant Registry in the Field–Atrial Fibrillation, GLORIA-AF registry Global Registry on Long-Term Oral Antithrombotic Treatment in Atrial Fibrillation, ICH Intracranial hemorrhage, KNHIS registry Korean National Health Insurance Service registry, MI Myocardial infarction, N/S Not specified, ORBIT-AF registry Outcomes Registry for Better Informed Treatment of Atrial Fibrillation, PT-INR prothrombin time international normalized ratio, RFA Radiofrequency ablation, START registry Survey on Anticoagulated AF Patients Receiving Therapy, TIA transient ischemic attack
Models development and evaluation
Data preprocessing and feature selection
Data preprocessing techniques varied across studies. Data cleaning was explicitly mentioned in two studies [36, 40]. Missing data imputation was addressed in several studies using single imputation methods, either with a constant value [10] or statistical methods [32, 35, 36, 42]. However, two studies used a more advanced multiple imputation method, the Multivariate Imputation by Chained Equations (MICE) technique [11, 39]. Only one study by Li et al. used imaging data, which explicitly described image processing focusing on fundus images [33] (Table 2).
Table 2.
Training and validation process of the include studies
| Algorithms Used | Feature Selection Methods (Final Features Count) |
Data Preprocessing; Explainability Analysis |
Cross Validation; Validation Type (train to test ratio) | Non ML Models for Comparison | Performance Metrics | Newly Identified Potential Risk Factors | |
|---|---|---|---|---|---|---|---|
| Letham et al. 2015 [28] |
- Bayesian Rule Lists - CART - C5.0 - LASSO - SVM - RF - BRL-post |
N/A |
N/S; Used interpretable algorithms |
3 or 5 folds; None |
CHADS2, CHA2DS2-VASc | AUROC | N/S |
| Li et al. 2016a [31] | LR |
- None (107) - Filtering by information gain (11) - Wrapper for AUROC (23) - LASSO (C = 0.1) (21) |
N/S; N/A |
10 folds; None |
Framingham, CHA2DS2-VASc | AUROC, AUPR | N/S |
| Li et al. 2016b [32] |
- LR - Cox - Naïve Bayes - CART - RF |
Excluded features with excessive missing data and then: - Chi-squared filter (p < 0.001) - Information gain (IG > 0.001) - correlation-based feature subset selection - Wrapper for AUROC - LASSO (C = 0.1) |
Statistical missing data imputation; used interpretable algorithms |
10 folds; Internal (3:2) |
Framingham, CHA2DS2-VASc | AUROC, AUPR | LV posterior wall thickness, LVEF, Total cholesterol, MI, ICH, Medication for ventricular rate control, Years since last TIA, Years since diabetes diagnosis, Years since PSVT |
| Leonarduzzi et al. 2018 [34] | SVM | Wilcoxon rank sum tests were used to identify significant scattering coefficients |
N/S; N/A |
5 folds; None |
CHA2DS2-VASc | AUROC | irregularity measurements from RR dynamics (characterized by the scattering transform) |
| Han et al. 2019 [29] |
- CNN (3 layer) - RF - LASSO |
N/A |
N/S; N/A |
10 folds; Internal (7:3) |
CHA2DS2-VASc | AUROC, Sensitivity, Specificity | Daily AF burden signature |
| Goto et al. 2020 [41] | CNN + LSTM (multiple layer) | N/A |
N/S; N/A |
N/A; Internal (7:3) |
Time in therapeutic range | AUROC, Sensitivity, Specificity, Accuracy | PT-INR measurement patterns |
| Loring et al. 2020 [42] |
- RF - GB - Multi-layer—NN - Single-layer NN |
N/A |
Statistical missing data imputation; N/A |
15 folds; External on GARFIELD-AF data (4:1) |
Stepwise LR | AUROC | N/S |
|
15 folds; External on ORBIT-AF data (4:1) | |||||||
| Watanabe et al. 2021 [35] |
- RF - stepwise LR |
Sequential forward floating selection |
Statistical missing data imputation; N/A |
5 folds; Internal (4:1) |
CHADS2 | AUROC | N/S |
| Jung et al. 2022 [37] |
- Attention- Based Deep NN (3 layer) - LR - XGB - RF |
LR (48) |
N/S; N/A |
10 folds; Internal (8:2) |
- CHA2DS2-VASc - Stepwise LR |
AUROC, Precision, Recall, F1 score | N/S |
| Lu et al. 2022 [10] |
- Multilabel GB - Multilayer NN - SVM |
N/A |
Constant missing data imputation; N/A |
N/S folds; Internal (3:1) |
CHA2DS2-VASc | AUROC, Precision, Recall, F1 score, Brier score | N/S |
| Nishi et al. 2022 [36] | - Cat Boost | Excluded features with > 30% missing data, selection by permutation importance using different ML models (RF, LR, SVM, NN, and NB) (14) |
Data cleaning, statistical missing data imputation; N/A |
5 folds; Internal (4:3) |
CHADS2, CHA2DS2-VASc | AUROC, Sensitivity, Specificity Accuracy, AUPR | N/S |
| Li et al. 2023 [33] |
3 models of deep NN: - Convolutional, - Residual - Squeeze-and-Excitation |
N/A |
Image processing; N/A |
N/A; None |
N/S | AUROC, Sensitivity, Specificity Accuracy, PPV, NPV | Multi- spectrum fundus images |
| Bernardini et al. 2024 [38] |
- Stepwise LR - GBDT - Naive Multi-Task—NN Multi-Gate Mixture of Experts NN |
using SHAP |
N/S; SHAP on GBDT model outputs |
5 folds stratified; None |
CHA2DS2-VASc | AUROC | Hemoglobin, Platelet Count, BMI |
| Chen et al. 2024 [40] |
- LightGBM - RF - ML LR - SVM - Multilayer NN |
Excluded features with > 50% missing data, Pearson correlation analysis and variance inflation factor did not find any strong collinearity (12) |
Data cleaning, multiple imputation; SHAP |
5 folds; External (7:3) |
CHA2DS2-VASc | AUROC, Sensitivity, Specificity Accuracy, Precision, Recall, F1 score, G mean | CKD, Hypertension, Diuretic use, elevated AST, enlarged LA, MV involvement, AF Treatment Strategy (rhythm vs rate) |
| Gao et al. 2024 [30] | LightGBM | not explained (7) |
N/S; LIME |
fivefold; Temporal (7:3) |
CHADS2, CHA2DS2-VASc | AUROC, EOD, FPO, DI, DP | N/S |
| Lu et al. 2024 [11] | Multi-label GB | Permutation features importance |
MICE; N/A |
N/A; Internal (7:3) |
CHA2DS2-VASc | AUROC, Sensitivity, Specificity, NRI | |
| Papadopoulou et al. 2024 [39] |
- XGB - LightGBM - RF - Deep NN - SVM - LASSO |
Excluded features with > 25% missing data, Removal of correlated variables (Pearson correlation > 0.8), Deletion of features with variance below 90%, Recursive feature elimination with cross-validation |
MICE and fixed under-sampling technique; SHAP |
tenfold stratified; Internal (N/S) |
CHA2DS2-VASc | AUROC, Accuracy, Precision, Recall, F1 score | Creatinine, Glycated Hemoglobin, Monocytes |
AF Atrial fibrillation, BRL Bayesian Rule Lists, CART Classification and Regression Trees, CKD Chronic kidney disease, CNN Convolutional Neural Network, DI Disparate Impact, DP Fairness Performance, EOD Equal Opportunity Difference, FPO False Positive Rate, GB Gradient Boosting, GBDT Gradient Boosted Decision Trees, LA Left atrium, LASSO Least Absolute Shrinkage and Selection Operator, LIME Local Interpretable Model-Agnostic Explanations, LR Logistic Regression, LV Left ventricle, LVEF Left ventricular ejection fraction, LSTM Long Short-Term Memory, MICE Multivariate Imputation by Chained Equations, N/A Not applicable, N/S Not specified, NN Neural Network, NPV Negative predictive value, NRI Net reclassification index, PPV Positive predictive value, PSVT Paroxysmal supraventricular tachycardia, PT-INR Prothrombin time international normalized ratio, RF Random Forest, SHAP SHapley Additive exPlanations, SVM Support Vector Machine, XGB Extreme Gradient Boosting
Feature selection was a crucial step in model development across studies, addressing the risk of overfitting from large initial feature sets (up to 167 in one study [31]). Researchers employed various methods, ranging from simple techniques like eliminating features with high percentages of missing values [36, 39, 40], identifying statistically significant features [31, 32, 34, 40], or removing features with little variance [39], to more advanced approaches such as permutation feature importance [11, 40], sequential forward floating selection [35], or recursive feature elimination after preprocessing and with cross-validation [39]. The resulting final feature sets varied widely, from 4 [34] to 48 [37] features. Notably, two earlier studies by Li et al. [31, 32] compared multiple feature selection techniques, including information gain, wrapper algorithms for AUROC, Lasso regularization, and correlation-based methods. Their findings suggested that wrapper algorithms for AUROC consistently yielded superior model performance, providing valuable insight for future studies in this field.
Machine learning algorithms
The studies employed a diverse array of ML algorithms. Traditional ML methods were widely used, with tree-based methods, including RF, decision trees, Classification and Regression Trees (CART), and various Gradient Boosting algorithms (GB, XGBoost, LightGBM, CatBoost), appearing in thirteen studies [10, 11, 28–30, 32, 35–40, 42]. The other popular algorithms were LR and its variants (e.g., L1-regularized LR) employed in eleven studies [11, 28, 29, 31, 32, 35, 37–40]. Deep learning approaches were prominent in nine studies [10, 29, 33, 37–42], featuring various NN architectures. These can be classified into basic, specialized, and task-specific categories. Basic NNs include multi-layer [10, 29, 33, 37–42], single-layer [42], and deep neural networks [39]. Specialized NNs encompass Convolutional NNs [29, 33, 41], Long Short-Term Memory networks (LSTM) [41], attention-based Deep NNs [37], Residual NNs [33], and Squeeze-and-Excitation NNs [33]. Task-specific NNs include Naive Multi-Task NNs and Multi-Gate Mixture of Experts NNs [38]. Probabilistic models were less commonly explored, like Bayesian Rule Lists in a study [28]. Overall, LR, RF, and various forms of NNs emerged as the most frequently used algorithms across the reviewed studies.
Validation approaches
To evaluate the model's generalization ability, detect overfitting, and provide a robust method for tuning hyperparameters and comparing models, cross-validation was employed in several studies. The number of folds ranged from 5 to 15. The most common approach was five-fold [28, 30, 34–36, 38, 40] or tenfold cross-validation [29, 31, 37, 39], but Loring et al. [42] used 15 folds with two sample sizes of 22,760 and 52,032. Only two studies [39] considered stratified cross-validation. However, it is noteworthy that for imbalanced datasets like CVA incidents in the AF population, stratification of each fold is often suggested to ensure that each fold has a similar proportion of each class, making the evaluation more representative.
While some studies only reported the performance results of models during training, others reported the results of model performance on an unseen part of the same dataset (internal validation) [10, 11, 29, 30, 32, 35, 37, 39–42]. A more rigorous evaluation was done only in two studies by Loring et al. [42] and Chen et al. [40] that reported the results of model performance on a dataset different from the training one (external validation). Gao et al. [30] incorporated temporal validation, reporting performance metrics on data collected after the training period to assess the model's ability to generalize to future patients [30]. Finally, comparisons with traditional risk scores like CHA₂DS₂-VASc [10, 11, 28–32, 34–40], CHADS₂ [28, 30, 35], Framingham [31, 32], or even time in therapeutic range [41] were also conducted. Additionally, some studies sought to compare their ML models with other statistical methods like logistic regression [37, 42] or Cox hazard [32] models.
A few studies in this review made efforts to enhance the interpretability and explainability of their ML models for predicting CVA in atrial fibrillation patients. Letham et al. (2015) and Li et al. (2016b) focused on using inherently interpretable models such as Bayesian Rule Lists and decision trees [28, 32]. Bernardini et al. (2024) and Chen et al. (2024) employed SHAP (SHapley Additive exPlanations) to explain their model outputs, particularly for gradient boosting decision tree models [38, 40]. SHAP assigns importance values to each feature, providing insights into the model's decision-making process. Gao et al. [30] utilized LIME (Local Interpretable Model-agnostic Explanations) for their LightGBM model, offering another approach to understanding individual predictions [30]. The studies that used feature selection methods inherently provide some level of interpretability by identifying the most relevant predictors.
Model performance assessments
The diversity of metrics used across studies reflects the complexity of evaluating ML models for CVA risk prediction in AF patients. While AUROC was universally reported as a standard measure for discrimination, many researchers recognized the need for additional metrics to provide a more comprehensive assessment of model performance (Table 2). Accuracy, reported in five studies [33, 36, 39–41], offered an overall measure of correct predictions. However, given the rarity of CVA events in AF populations, accuracy alone can be misleading: Models might achieve high accuracy by predominantly predicting the majority class (CVA negative), potentially overlooking critical positive cases. To address this limitation, several studies reported sensitivity and specificity [29, 33, 37, 40, 41], providing insight into the models' ability to correctly identify positive and negative cases, respectively. Recognizing the clinical implications of prediction outcomes, Li et al. [33] reported Positive Predictive Value (PPV) and Negative Predictive Value (NPV), indicating the probabilities of true positive and true negative results. To better assess the performance on the minority class (CVA positive), four studies [11, 37, 39, 40] reported the pack of precision (equivalent to PPV), recall (equivalent to sensitivity), and the F1 score (a balanced measure of precision and recall). The Area Under the Precision-Recall curve (AUPR) was reported in three other studies [31, 32, 36], providing a more informative measure for imbalanced datasets compared to AUROC alone.
Reporting of calibration and clinical utility was limited across the included studies. Only four of 17 studies formally assessed calibration, typically using calibration plots [29, 35, 42] or Brier scores [10]. Although many models output predicted probabilities, only a few of them directly compared these with observed event rates [29, 35, 42]. Similarly, clinical utility was evaluated sparsely: four studies reported reclassification metrics such as the NRI [10, 11, 29, 42], several presented Kaplan–Meier curves or hazard ratios by model-defined risk groups or simple threshold-based scenarios [10, 11, 36, 40, 41], and only a single investigation incorporated formal decision-curve net benefit analysis [11]. Overall, three studies provided both some calibration assessment and at least one measure related to clinical utility, three provided partial information in one of these domains, and the remaining 11 studies did not adequately address absolute risk calibration or clinical usefulness (Table S4).
Predictive features
The reviewed studies incorporated a diverse range of predictive features, extending beyond those utilized in traditional risk scores. While most studies included standard demographic information and clinical variables such as comorbidities, some focused on specific types of clinical data, yielding notable insights.
Leonarduzzi et al. demonstrated the predictive power of irregularity measurements from RR dynamics in 24-h Holter records, characterized by the scattering transform. This approach showed superior CVA risk prediction compared to the CHA₂DS₂-VASc score [34]. Han et al. utilized daily AF burden signatures over 30 days as a predictive feature, again outperforming the CHA₂DS₂-VASc score in CVA risk prediction [29]. Goto et al. incorporated PT-INR measurement patterns over time, demonstrating superior prediction of CVA or systemic embolic events compared to the Time in Therapeutic Range metric [41]. In a novel approach, Li et al. employed multi-spectrum fundus images as predictive features, highlighting the potential of ocular imaging data in risk assessment [33].
Other studies expanded the feature set by supplementing standard demographic information and comorbidities with additional data types. Six studies [10, 30, 32, 38–40] included medication data in their models. Five of them [30, 32, 38–40] also incorporated laboratory and echocardiography data.
These comprehensive approaches identified several potential predictors of CVA risk in AF patients. Chen et al. highlighted the importance of chronic kidney disease, hypertension, diuretic use, and AF treatment strategy [40]. They also found elevated aspartate aminotransferase (AST) to be a potential marker of increased CVA risk [40]. Hemoglobin, glycated hemoglobin, monocyte, and platelet count were also distinguished by two studies [38, 39]. Li et al. (2016b) identified left ventricular posterior wall thickness, left ventricular ejection fraction (LVEF), left atrium enlargement, and mitral valve involvement as significant predictors [32].
Anticoagulation use was also considered as a predictive factor in some studies, although its handling varied considerably (Table S3). Four studies explicitly excluded patients on oral anticoagulants (OAC) at baseline [29, 31, 32, 36], while two limited their cohorts to OAC-treated patients [38, 41]. Six studies included both treated and untreated patients [10, 11, 34, 35, 40, 41], and five did not clearly describe OAC eligibility criteria. Baseline OAC status was commonly reported in mixed or OAC-only cohorts and was occasionally incorporated as a predictor or used for subgroup analyses (e.g., OAC vs. no OAC, or vitamin K antagonist [VKA] vs. direct oral anticoagulant [DOAC]) [10, 11, 34, 38, 42]. However, several studies using contemporary data failed to report baseline OAC status altogether. On the other hand, incident OAC use during follow-up was rarely quantified or integrated into model development. Most studies did not mention it; in the few that did, it was explicitly excluded from modeling [11, 29]. Only two VKA-based cohorts considered treatment intensity by incorporating time in therapeutic range (TTR) [35, 41], and no study accounted for OAC adherence or persistence. Overall, OAC was either absent from the predictor set, treated as a static baseline binary variable, or used solely to define or stratify subgroups. No model treated OAC as a time-varying exposure, and we judged that none of the studies adequately managed this complexity. A few studies reported OAC-related findings, including improved stroke prediction in untreated patients, identification of TTR or OAC use as important predictive features, and similar ML model performance across OAC strata. However, such analyses were uncommon across the included studies.
Risk of bias
The ROB across the included studies varied, with a mixture of low, high, and unclear risks across key domains. Notably, four studies included systemic thromboembolism in addition to CVA [35, 36, 41, 42], contributing to high ROB in applicability relative to this review’s focus (Table 3).
Table 3.
Risk of bias in the included studies
| Risk of bias | Applicability | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Participants | Predictors | Outcomes | Analysis | Overall | Participants | Predictors | Outcomes | Overall | |
| Letham et al. 2015 [28] | H/D | L | U | H/DV | H | L | L | U | U |
| Li et al. 2016a [31] | L | L | L | H/DV | H | L | L | L | L |
| Li et al. 2016b [32] | L | L | L | L | L | L | L | L | L |
| Leonarduzzi et al. 2018 [34] | L | L | L | H/V | H | L | L | L | L |
| Han et al. 2019 [29] | L | L | U | L | U | L | L | U | U |
| Goto et al. 2020 [41] | L | L | L | L | L | L | L | H | H |
| Loring et al. 2020 [42] | H/D | L | H/V | H/D | H | L | L | L | L |
| Watanabe et al. 2021 [35] | L | L | L | L | L | L | L | H | H |
| Jung et al. 2022 [37] | H/D | L | L | L | H | L | L | L | L |
| Lu et al. 2022 [10] | L | L | L | L | L | L | L | L | L |
| Nishi et al. 2022 [36] | L | L | L | L | L | L | L | H | H |
| Li et al. 2023 [33] | L | L | L | H/V | H | L | L | L | L |
| Bernardini et al. 2024 [38] | L | L | L | H/V | H | L | L | L | L |
| Chen et al. 2024 [40] | L | L | U | L | U | L | L | U | U |
| Gao et al. 2024 [30] | H/D | L | U | H/D | H | L | L | U | U |
| Papadopoulou et al. 2024 [39] | H/D | L | L | L | H | L | L | L | L |
| Lu et al. 2024 [11] | H/D | L | L | L | H | L | L | L | L |
L Low risk, U Unclear risk, H High risk, H/D High risk in development, H/V High risk in validation, H/DV High risk in development and validation
Meta-analysis of models’ performance
A random-effects meta-analysis was performed, incorporating 13 studies comprising 46 ML model records that utilized standard demographic and clinical predictors and reported AUROC values (Fig. 2). The pooled AUROC for ML models was 0.66 (95% CI: 0.64–0.68), indicating moderate discriminative ability for stroke risk prediction. For studies that included the CHA₂DS₂-VASc score as a comparator, the pooled AUROC was 0.64 (95% CI: 0.62–0.66). The difference in pooled AUROCs between ML models and the CHA₂DS₂-VASc score was not statistically significant (P = 0.327), suggesting that ML-based approaches are comparable to, but not yet superior to, established clinical risk stratification tools (Figures S1, S2).
Fig. 2.
Forest plot summarizing meta-analysis results, including the main analysis, subgroup analyses, and sensitivity analyses. The last column P-values represent within-subgroup differences for subgroup analyses and differences compared to the overall ML pooled estimate in sensitivity analyses. AUROC = area under the receiver operating characteristic curve; CI = confidence interval
Subgroup analysis based on ML algorithm type (tree-based models, SVMs, LRs, and NNs) did not reveal significant heterogeneity in performance (P = 0.111). SVM models showed the highest pooled AUROC (0.68; 95% CI: 0.60–0.75), followed by tree-based models (0.66; 95% CI: 0.64–0.69). However, the number of SVM-based models was limited (n = 5), and this small sample size may reduce the stability and generalizability of the pooled estimate (Figure S3-S6). In a descriptive subgroup analysis by validation strategy, studies that reported only training-set performance had a pooled AUROC of 0.70 (95% CI 0.66–0.74), whereas studies reporting performance on a validation set (internal or external) had a pooled AUROC of 0.64 (95% CI 0.62–0.66) (Figures S7, S8). This pattern is consistent with expected optimism when models are evaluated on their development data; however, the comparison is based on heterogeneous, non-random groups of studies and should be interpreted cautiously.
Sensitivity analyses revealed no significant difference in pooled AUROCs compared to the overall estimate when restricting the analysis to either the best-performing model from each study (P = 0.151, Figure S9) or studies with low ROB (P = 0.969, Figure S10). However, when the best models from each study were compared specifically to the pooled performance of the CHA₂DS₂-VASc score, the difference was statistically significant, favoring ML models (AUROC = 0.69; 95% CI: 0.66–0.72; P = 0.009). This finding was further supported by the third sensitivity analysis, which examined the within-study difference in discriminatory performance between the best-performing ML models and the CHA₂DS₂-VASc score. In this analysis, ML models demonstrated significantly higher discrimination (ΔAUROC = 0.05; 95% CI: 0.03–0.07; P < 0.0001), although the magnitude of this difference is unlikely to be clinically meaningful (Figure S11).
Heterogeneity was estimated to be high across all analyses. Meta-regression revealed no significant association between AUROC and sample size (P = 0.800), but identified a significant positive association with stroke prevalence in study populations (logit coefficient: 0.053; 95% CI: 0.035–0.071; P < 0.0001). However, stroke prevalence was not the sole contributor to heterogeneity, and adjusting for it did not substantially reduce overall heterogeneity. The remaining variability is likely attributable to differences in study design, patient populations, outcome definitions, ML algorithms, and model development and validation methods, all of which are discussed throughout the manuscript.
Visual inspection of the funnel plots did not reveal a marked asymmetry suggestive of strong small-study effects for (i) the within-study difference in AUROC between ML models and CHA₂DS₂-VASc (ΔAUROC), (ii) the AUROC of the best-performing ML model from each study, or (iii) the AUROC of the CHA₂DS₂-VASc score (Fig. 3, Figures S12–S13). However, given the limited number of studies and substantial between-study heterogeneity, these analyses should be regarded as exploratory and were not used as primary evidence of publication bias.
Fig. 3.
Funnel plot assessing potential small-study effects for the ΔAUROC (within-study difference in AUROC between ML models and CHA₂DS₂-VASc)
Discussion
To our knowledge, this is the first meta-analysis to evaluate the performance of ML models for predicting stroke in patients with AF. Based on 17 included studies, ML models demonstrated predictive performance comparable to traditional clinical tools, though with considerable variability. Key findings include: (i) in sensitivity analyses limited to the best-performing model per study and to within-study comparisons with the CHA₂DS₂-VASc score, ML models demonstrated statistically greater discrimination; however, the absolute gain in AUROC was modest and of uncertain clinical relevance; (ii) a wide range of ML algorithms were applied, with tree-based models being the most commonly used; (iii) although SVMs showed the highest pooled AUROC, this was based on a limited number of records, while tree-based models, more frequently implemented, demonstrated consistently strong performance; and (iv) several studies explored novel features such as RR dynamics in 24-h Holter records, AST levels, and left ventricular structural indices which warrant further validation before clinical adoption. This systematic review underscores the potential future role of ML in stroke prediction, pending further validation.
Stroke is considered the most devastating outcome of AF, with more than 100,000 strokes in the year in the US attributed to AF [43]. Hence, the prediction of this outcome has a higher importance, and this is the main reason traditional risk scores have been developed and designed. In this regard, the CHA₂DS₂-VASc score has been designed but failed to take several other risk factors, such as chronic kidney disease, AF type, and electrocardiographic features, into account [44–46]. Moreover, the lack of distinguishability between risk factors is another limitation of this score since all risk factors have the same weight and contribution to the overall score. This highlights an area where ML models could contribute to enhanced stroke prediction, if validated. ML-based predictive models have the advantage of controlling variables’ covariance and encompassing a wide range of variables [47]. Moreover, the conventional risk models for stroke in patients with AF, except for a limited model on AF-prespecified burden thresholds [48], do not incorporate measures of AF burden or duration and hence, are not able to report the changes in burden over time. In our study, several ML models were implemented for the prediction of stroke after AF; however, a large heterogeneity was observed among them. These were seen in populations, methods, verification methods, variables implemented, and evaluation metrics.
One of the main criticisms of ML models is the lack of transparency in the process, which makes feature interactions and intermediate steps poorly understood. Hence, these are considered black boxes. Among the ML models assessed in our included studies, tree-based models were the most frequently applied and demonstrated consistently strong predictive performance. While these models, such as RF, are also black boxes, they have the advantage of providing us with the most important variables that contribute the most [49, 50]. RF model randomly shuffles the data for one of the variables for one feature at a time over the whole dataset and calculates the drop in each metric. This might lead to the selection of variables not commonly used in traditional risk factors or logistic regression models. This is mainly because, unlike regression models, the RF model uses a non-parametric decision tree approach, which emphasizes the fact that an increase in these variables does not necessarily lead to a monotonic change in the risk. On the other hand, it should be noted that the fact that novel features or patterns could not be interpreted in these models might question their use for clinicians.
Several predictive features have been investigated and reported in the included studies. For instance, in the investigation by Watanabe et al. age was identified as the most important feature [35]. This was similar to the traditional risk factors, such as the CHA₂DS₂-VASc score, in which age is an important contributor and was observed in previous studies as well [51–53]. Interestingly, total cholesterol was the second top contributor to the risk of thromboembolism in patients with AF. While this is not usually added to traditional risk estimators for stroke in patients with AF, dyslipidemia is a well-established risk factor for stroke [54]. Additionally, its involvement in coronary heart disease and atherosclerosis both have associations with the development of AF [55]. Another important finding was the role of heart rate variability in the prediction of stroke in patients with AF, which was observed in the Leonarduzzi et al. study [34]. This was previously observed in the Watanabe et al. investigation as well, highlighting the need for focusing on the Holter electrocardiogram for the prediction of stroke in these high-risk populations [56].
Although anticoagulation use is central to stroke prevention in AF, its modeling in ML-based risk prediction was limited and inconsistent across studies. Most models treated OAC as a static baseline variable or subgrouping factor (e.g., VKA vs. DOAC), with few exceptions like Goto et al. (2020), which incorporated TTR to reflect treatment quality [41]. However, even these studies did not capture important real-world dynamics such as treatment initiation, adherence, switching between agents, or time-varying exposure. This contrasts with other cardiovascular ML applications, where time-updated predictors such as medication use or vital signs have been successfully integrated using methods like recurrent neural networks or landmarking [57–59]. The failure to account for these treatment dynamics may lead to misestimation of stroke risk, particularly in patients who start, stop, or inconsistently use anticoagulation. In contrast to traditional scores like CHA₂DS₂-VASc, which assume untreated risk, ML models have the potential to incorporate treatment trajectories and improve clinical relevance. Future models should consider linking to prescription or EHR data and applying time-varying or causal frameworks to more accurately reflect the evolving role of anticoagulation in stroke prevention.
Based on our systematic review of literature on the prediction of stroke in AF cases, a majority of the included studies used only internal validation. This is while proper validation is an undetachable part of a model design and is crucial for the generalizability of the models. On the other hand, only investigations of Loring et al. [42] and Chen et al. [40] used external validation as an important part of benefiting from the models in clinical practice, and 5 studies did not incorporate any separate set for validation. An additional observation was that studies reporting only training-set performance tended to show higher pooled AUROCs than those reporting performance on a separate validation set. This aligns with the well-recognized issue that performance estimates can be optimistically biased when models are tested on the same data used for training, potentially overstating their predictive utility in real-world settings. However, this contrast is based on non-comparable groups of studies and remains susceptible to confounding, selective reporting, and residual non-independence when multiple models are reported per study; we therefore regard these subgroup estimates as descriptive and hypothesis-generating rather than definitive.
Robust evaluation of calibration and clinical utility is essential in machine-learning-based prognostic modelling, as calibration determines whether predicted probabilities reflect true event rates, while clinical utility analyses establish whether a model confers meaningful benefit over existing decision strategies [60, 61]. Without these evaluations, neither the accuracy of absolute risk estimates nor the practical value of implementing a model in clinical workflows can be ascertained. These considerations are particularly critical for stroke prediction in atrial fibrillation, where treatment decisions depend directly on well-calibrated absolute risk estimates and demonstrable net clinical benefit. However, among the included studies, calibration was only selectively examined, and clinical utility was rarely assessed through formal decision-analytic methods, limiting the ability to determine whether the proposed models are suitable for guiding anticoagulation decisions or improving current risk-stratification practices.
Another methodological consideration is that several studies developed multiple ML models using the same underlying cohort. Treating these models as independent data points in a meta-analysis can underestimate uncertainty due to within-study correlation. To assess the potential impact of this, we conducted a sensitivity analysis at the study level, including only the best-performing ML model from each study. This analysis produced effect estimates and confidence intervals broadly consistent with those of the original model-level meta-analysis, suggesting that our conclusions are not substantially affected by within-study dependence. Notably, under this more conservative specification, ML models still demonstrated statistically higher discrimination than CHA₂DS₂-VASc; however, the absolute gain in AUROC was modest and reflects selective reporting of the best-performing model in each study. We therefore interpret this finding cautiously and do not consider it sufficient evidence of clear or clinically meaningful superiority of current ML approaches over CHA₂DS₂-VASc.
Taken together, our findings highlight several priorities for future work on ML-based stroke prediction in AF. First, new models should be developed and externally validated in large, diverse, and prospectively assembled cohorts with harmonized outcome definitions (e.g., clearly defined ischemic stroke/TIA vs broader composites) to reduce the current heterogeneity in case-mix and endpoints. Second, future studies should move beyond discrimination and routinely report calibration, absolute risk estimates, and decision-analytic measures such as decision-curve analysis, so that the clinical utility of ML models can be judged in realistic treatment-threshold scenarios. Third, dynamic aspects of AF and its management, including AF burden, temporal patterns of anticoagulation initiation and discontinuation, treatment intensity (e.g., TTR), and switching between OAC agents, should be incorporated using time-varying or causal modeling frameworks rather than static baseline indicators. Fourth, there is a need to systematically evaluate and validate novel data sources and features identified in this review (e.g., Holter-derived RR dynamics, AF burden signatures, longitudinal laboratory and imaging markers, fundus images), ideally with attention to parsimony and clinical interpretability. Finally, future research should prioritize transparent model development and reporting (including handling of missing data, feature selection strategies, and explainability methods such as SHAP or LIME) and progress towards prospective impact studies or pragmatic trials to determine whether integrating ML-based tools into clinical workflows meaningfully improves decision-making and patient outcomes compared with existing risk scores such as CHA₂DS₂-VASc.
However, there are also limitations to this study that need to be addressed. First, high heterogeneity among the studies in terms of population, design, models used, and metrics reported threatens the generalizability of these findings and emphasizes the need for large-scale multi-center trials in real-world settings. Moreover, as demonstrated by our meta-regression, stroke prevalence appears to influence reported AUROCs, further limiting the comparability of model performance across studies with differing baseline risks. Our assessment of small-study effects using funnel plots is likewise constrained by the limited number of studies, substantial heterogeneity, and the bounded, context-dependent nature of AUROC, and should therefore be interpreted as exploratory rather than definitive evidence regarding publication bias. Finally, as noted, the lack of external validation in most of the included studies not only limits the models’ application in clinical settings but also has a profound impact on our reported pooled result.
Conclusion
In this systematic review and meta-analysis, ML models for stroke prediction in AF showed overall discrimination similar to CHA₂DS₂-VASc, with only modest incremental gains in sensitivity analyses. Marked heterogeneity in populations, outcome definitions, predictors, and modeling strategies, together with limited external validation and poor reporting of calibration and clinical utility, currently limits confidence in these models and does not support routine clinical implementation. Future studies should focus on methodologically robust development and external validation in diverse cohorts, incorporate dynamic information such as AF burden and anticoagulation exposure, and formally evaluate calibration and net clinical benefit to determine whether ML-based risk prediction tools can offer meaningful advantages over existing risk scores.
Supplementary Information
Acknowledgements
None.
Abbreviations
- AF
Atrial Fibrillation
- AHA
American Heart Association
- AUROC
Area Under the Receiver Operating Characteristic Curve
- AUPR
Area Under the Precision-Recall Curve
- AST
Aspartate Aminotransferase
- CART
Classification and Regression Trees
- CHA₂DS₂-VASc
Congestive Heart Failure, Hypertension, Age ≥ 75 years, Diabetes Mellitus, Stroke, Vascular Disease, Age 65–74 years, Sex Category (female)
- CHADS₂
Congestive Heart Failure, Hypertension, Age ≥ 75 years, Diabetes Mellitus, Stroke/TIA
- CNN
Convolutional Neural Network
- CVA
Cerebrovascular Accident
- ESC
European Society of Cardiology
- DOAC
Direct oral anticoagulant
- GB
Gradient Boosting
- LIME
Local Interpretable Model-agnostic Explanations
- LR
Logistic Regression
- LSTM
Long Short-Term Memory networks
- ML
Machine Learning
- MICE
Multivariate Imputation by Chained Equations
- NN
Neural Networks
- NPV
Negative Predictive Value
- OAC
Oral anticoagulant
- PPV
Positive Predictive Value
- PT-INR
Prothrombin Time International Normalized Ratio
- PRISMA
Preferred Reporting Items for Systematic Reviews and Meta-Analyses
- PROBAST
Prediction model Risk Of Bias ASsessment Tool
- ROB
Risk of Bias
- RF
Random Forests
- SHAP
SHapley Additive exPlanations
- SVM
Support Vector Machines
- TIA
Transient Ischemic Attack
- TTR
Time in Therapeutic Range
- VKA
Vitamin K Antagonist
Authors’ contributions
A.Az., AH.B., M.M., and Z.H. drafted the original manuscript. Mo.T., Z.H., and Ma.T. reviewed and edited the manuscript. Ma.T., Mo.T., and A.Ar. conceptualized the study and developed the methodology. A.Az., AH.B., M.M., P.F. and Z.H.curated the data. A.Ar. and Z.H. provided resources. A.Ar. created visualizations. Mo.T. managed the project administration. Ma.T. supervised the study. P.F. performed data analysis. Finally, all authors reviewed the manuscript.
Funding
None.
Data availability
The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Morvarid Taebi and Alireza Arvin contributed equally to this project and share the authorship.
References
- 1.Global, regional, and national burden of stroke and its risk factors, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet Neurol. 2021;20(10):795-820. [DOI] [PMC free article] [PubMed]
- 2.Migdady I, Russman A, Buletko AB. Atrial fibrillation and ischemic stroke: a clinical review. Semin Neurol. 2021;41(4):348–64. [DOI] [PubMed] [Google Scholar]
- 3.Gažová A, Leddy JJ, Rexová M, Hlivák P, Hatala R, Kyselovič J. Predictive value of CHA2DS2-VASc scores regarding the risk of stroke and all-cause mortality in patients with atrial fibrillation (CONSORT compliant). Medicine (Baltimore). 2019;98(31):e16560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Casselman F, Coca A, De Caterina R, Deftereos S, Dobrev D, Ferro JM, et al. 2016 ESC Guidelines for the management of atrial fibrillation developed in collaboration with EACTS. Eur Heart J. 2016;37:2893–962. [DOI] [PubMed] [Google Scholar]
- 5.January CT, Wann LS, Calkins H, Chen LY, Cigarroa JE, Cleveland JC, et al. 2019 AHA/ACC/HRS Focused Update of the 2014 AHA/ACC/HRS Guideline for the Management of Patients With Atrial Fibrillation: A Report of the American College of Cardiology/American Heart Association Task Force on Clinical Practice Guidelines and the Heart Rhythm Society in Collaboration With the Society of Thoracic Surgeons. Circulation. 2019;140(2):e125–51. [DOI] [PubMed] [Google Scholar]
- 6.Cheung CC, Nattel S, Macle L, Andrade JG. Management of atrial fibrillation in 2021: an updated comparison of the current CCS/CHRS, ESC, and AHA/ACC/HRS guidelines. Can J Cardiol. 2021;37(10):1607–18. [DOI] [PubMed] [Google Scholar]
- 7.Jegatheswaran J, Hundemer GL, Massicotte-Azarniouch D, Sood MM. Anticoagulation in patients with advanced chronic kidney disease: walking the fine line between benefit and harm. Can J Cardiol. 2019;35(9):1241–55. [DOI] [PubMed] [Google Scholar]
- 8.Jia X, Levine GN, Birnbaum Y. The CHA(2)DS(2)-VASc score: not as simple as it seems. Int J Cardiol. 2018;257:92–6. [DOI] [PubMed] [Google Scholar]
- 9.Sun Y, Ling Y, Chen Z, Wang Z, Li T, Tong Q, et al. Finding low CHA2DS2-VASc scores unreliable? Why not give morphological and hemodynamic methods a try? Front Cardiovasc Med. 2022;9:1032736. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Lu J, Hutchens R, Hung J, Bennamoun M, McQuillan B, Briffa T, et al. Performance of multilabel machine learning models and risk stratification schemas for predicting stroke and bleeding risk in patients with non-valvular atrial fibrillation. Comput Biol Med. 2022;150. [DOI] [PubMed]
- 11.Lu J, Bisson A, Bennamoun M, Zheng Y, Sanfilippo F, Hung J, et al. Predicting multifaceted risks using machine learning in atrial fibrillation: insights from GLORIA-AF study. Eur Heart J Dig Health. 2024. [DOI] [PMC free article] [PubMed]
- 12.Emon MU, Keya MS, Meghla TI, Rahman MM, Al Mamun MS, Kaiser MS, editors. Performance analysis of machine learning approaches in stroke prediction. 2020 4th international conference on electronics, communication and aerospace technology (ICECA). IEEE; 2020.
- 13.JalajaJayalakshmi V, Geetha V, Ijaz MM, editors. Analysis and prediction of stroke using machine learning algorithms. 2021 International Conference on Advancements in Electrical, Electronics, Communication, Computing and Automation (ICAECA). IEEE; 2021.
- 14.Dritsas E, Trigka M. Stroke risk prediction with machine learning techniques. Sensors. 2022;22(13):4670. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Heo J, Yoon JG, Park H, Kim YD, Nam HS, Heo JH. Machine learning–based model for prediction of outcomes in acute stroke. Stroke. 2019;50(5):1263–5. [DOI] [PubMed] [Google Scholar]
- 16.Azam MS, Habibullah M, Rana HK. Performance analysis of various machine learning approaches in stroke prediction. Int J Comput Appl. 2020;175(21):11–5. [Google Scholar]
- 17.Wu Y, Fang Y. Stroke prediction with machine learning methods among older Chinese. Int J Environ Res Public Health. 2020;17(6):1828. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Dev S, Wang H, Nwosu CS, Jain N, Veeravalli B, John D. A predictive analytics approach for stroke prediction using machine learning and neural networks. Healthcare Analytics. 2022;2:100032. [Google Scholar]
- 19.Liaqat S, Dashtipour K, Zahid A, Assaleh K, Arshad K, Ramzan N. Detection of atrial fibrillation using a machine learning approach. Information. 2020;11(12):549. [Google Scholar]
- 20.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372. [DOI] [PMC free article] [PubMed]
- 21.Wolff RF, Moons KG, Riley RD, Whiting PF, Westwood M, Collins GS, et al. Probast: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51–8. [DOI] [PubMed] [Google Scholar]
- 22.Viechtbauer W, Viechtbauer MW. Package ‘metafor’. The comprehensive R Archive network Package ‘metafor. 2015.
- 23.Ding Q, Xing J, Bai F, Shao W, Hou K, Zhang S, et al. C1QC, VSIG4, and CFD as potential peripheral blood biomarkers in atrial fibrillation-related cardioembolic stroke. Oxidative Med Cell Longevity. 2023;2023. [DOI] [PMC free article] [PubMed]
- 24.Guo J, Zhou Y, Zhou B, Guo J, Zhou Y, Zhou B. Development and validation of a new nomogram model for predicting acute ischemic stroke in elderly patients with non-valvular atrial fibrillation: a single-center cross-sectional study. Clin Interv Aging. 2024;19:67–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Jiang C, Chen T, Du X, Li X, He L, Lai Y, et al. A simple and easily implemented risk model to predict 1-year ischemic stroke and systemic embolism in Chinese patients with atrial fibrillation. Chin Med J. 2021;134(19):2293–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kostev K, Wu T, Wang Y, Chaudhuri K, Reeve R, Tanislav C. Predicting the risk of ischemic stroke in patients treated with novel oral anticoagulants: a machine learning approach. Neuroepidemiology. 2021;55(5):387–92. [DOI] [PubMed] [Google Scholar]
- 27.Lip G, Tran G, Genaidy A, Marroquin P, Estes C, Landsheft J, et al. Improving dynamic stroke risk prediction in non-anticoagulated patients with and without atrial fibrillation: comparing common clinical risk scores and machine learning algorithms. Eur Heart J Quality Care Clin Outcomes. 2022;8(5):548–56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Letham B, Rudin C, McCormick T, Madigan D, Letham B, Rudin C, et al. Interpretable classifiers using rules and Bayesian analysis: building a better stroke prediction model. Ann Appl Stat. 2015;9(3):1350–71. [Google Scholar]
- 29.Han L, Askari M, Altman R, Schmitt S, Fan J, Bentley J, et al. Atrial fibrillation burden signature and near-term prediction of stroke a machine learning analysis. Circulation Cardiovasc Quality Outcomes. 2019;12(10). [DOI] [PMC free article] [PubMed]
- 30.Gao JF, Mar P, Tang ZZ, Chen GH. Fair prediction of 2-year stroke risk in patients with atrial fibrillation. J Am Med Inform Association. 2024. [DOI] [PMC free article] [PubMed]
- 31.Li X, Liu H, Du X, Hu G, Xie G, Zhang P. Using frequent item set mining and feature selection methods to identify interacted risk factors - the atrial fibrillation case study. Stud Health Technol Inform. 2016;228:562–6. [PubMed] [Google Scholar]
- 32.Li X, Liu H, Du X, Zhang P, Hu G, Xie G, et al. Integrated machine learning approaches for predicting ischemic stroke and thromboembolism in atrial fibrillation. AMIA Annu Symp Proc. 2016;2016:799–807. [PMC free article] [PubMed] [Google Scholar]
- 33.Li H, Gao M, Song H, Wu X, Li G, Cui Y, et al. Predicting ischemic stroke risk from atrial fibrillation based on multi-spectral fundus images using deep learning. Front Cardiovasc Med. 2023;10. [DOI] [PMC free article] [PubMed]
- 34.Leonarduzzi R, Abry P, Wendt H, Kiyono K, Yamamoto Y, Watanabe E, et al. Scattering transform of heart rate variability for the prediction of ischemic stroke in patients with atrial fibrillation. Methods Inf Med. 2018;57(3):141–5. 10.3414/ME17-02-0006. [DOI] [PubMed] [Google Scholar]
- 35.Watanabe E, Noyama S, Kiyono K, Inoue H, Atarashi H, Okumura K, et al. Comparison among random forest, logistic regression, and existing clinical risk scores for predicting outcomes in patients with atrial fibrillation: a report from the J-RHYTHM registry. Clin Cardiol. 2021;44(9):1305–15. 10.1002/clc.23688. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Nishi H, Oishi N, Ogawa H, Natsue K, Doi K, Kawakami O, et al. Predicting cerebral infarction in patients with atrial fibrillation using machine learning: The Fushimi AF registry. J Cerebral Blood Flow Metabolism. 2022;42(5):746–56. 10.1177/0271678X211063802. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Jung S, Song M, Lee E, Bae S, Kim Y, Lee D, et al. Predicting ischemic stroke in patients with atrial fibrillation using machine learning. Front Biosci Landmark. 2022;27(3). [DOI] [PubMed]
- 38.Bernardini A, Bindini L, Antonucci E, Berteotti M, Giusti B, Testa S, et al. Machine learning approach for prediction of outcomes in anticoagulated patients with atrial fibrillation. Int J Cardiol. 2024;407. [DOI] [PubMed]
- 39.Papadopoulou A, Harding D, Slabaugh G, Marouli E, Deloukas P. Prediction of atrial fibrillation and stroke using machine learning models in UK Biobank. Heliyon. 2024;10(7):e28034. 10.1016/j.heliyon.2024.e28034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Chen Y, Gue Y, Calvert P, Gupta D, McDowell G, Azariah JL, et al. Predicting stroke in Asian patients with atrial fibrillation using machine learning: a report from the KERALA-AF registry, with external validation in the APHRS-AF registry. Current problems in cardiology. 2024;49(4):102456. [DOI] [PubMed]
- 41.Goto S, Goto S, Pieper K, Bassand J, Camm A, Fitzmaurice D, et al. New artificial intelligence prediction model using serial prothrombin time international normalized ratio measurements in atrial fibrillation patients on vitamin K antagonists: GARFIELD-AF. Eur Heart J Cardiovasc Pharmacother. 2020;6(5):301–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Loring Z, Mehrotra S, Piccini J, Camm J, Carlson D, Fonarow G, et al. Machine learning does not improve upon traditional regression in predicting outcomes in atrial fibrillation: an analysis of the ORBIT-AF and GARFIELD-AF registries. Europace. 2020;22(11):1635–44. 10.1093/europace/euaa172. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wolf PA, Abbott RD, Kannel WB. Atrial fibrillation: a major contributor to stroke in the elderly. The Framingham Study. Arch Intern Med. 1987;147(9):1561–4. [PubMed] [Google Scholar]
- 44.Olesen JB, Lip GYH, Kamper AL, Hommel K, Køber L, Lane DA, et al. Stroke and bleeding in atrial fibrillation with chronic kidney disease. N Engl J Med. 2012;367(7):625–35. [DOI] [PubMed] [Google Scholar]
- 45.Steinberg BA, Hellkamp AS, Lokhnygina Y, Patel MR, Breithardt G, Hankey GJ, et al. Higher risk of death and stroke in patients with persistent vs. paroxysmal atrial fibrillation: results from the ROCKET-AF Trial. Eur Heart J. 2015;36(5):288–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.O’Neal WT, Howard VJ, Kleindorfer D, Kissela B, Judd SE, McClure LA, et al. Interrelationship between electrocardiographic left ventricular hypertrophy, QT prolongation, and ischaemic stroke: the reasons for geographic and racial differences in stroke study. EP Europace. 2016;18(5):767–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Lip GY, Genaidy A, Tran G, Marroquin P, Estes C, Sloop S. Improving stroke risk prediction in the general population: a comparative assessment of common clinical rules, a new multimorbid index, and machine-learning-based algorithms. Thromb Haemost. 2022;122(01):142–50. [DOI] [PubMed] [Google Scholar]
- 48.Boriani G, Glotzer TV, Santini M, West TM, De Melis M, Sepsi M, et al. Device-detected atrial fibrillation and risk for stroke: an analysis of >10,000 patients from the SOS AF project (Stroke preventiOn Strategies based on Atrial Fibrillation information from implanted devices). Eur Heart J. 2014;35(8):508–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Fisher A, Rudin C, Dominici F. All models are wrong, but many are useful: learning a variable's importance by studying an entire class of prediction models simultaneously. J Mach Learn Res. 2019;20. [PMC free article] [PubMed]
- 50.Khalaji A, Behnoush AH, Jameie M, Sharifi A, Sheikhy A, Fallahzadeh A, et al. Machine learning algorithms for predicting mortality after coronary artery bypass grafting. Front Cardiovasc Med. 2022;9. [DOI] [PMC free article] [PubMed]
- 51.Kim T-H, Yang P-S, Yu HT, Jang E, Uhm J-S, Kim J-Y, et al. Age threshold for ischemic stroke risk in atrial fibrillation. Stroke. 2018;49(8):1872–9. [DOI] [PubMed] [Google Scholar]
- 52.Chao T-F, Wang K-L, Liu C-J, Lin Y-J, Chang S-L, Lo L-W, et al. Age threshold for increased stroke risk among patients with atrial fibrillation: a nationwide cohort study from Taiwan. J Am Coll Cardiol. 2015;66(12):1339–47. [DOI] [PubMed] [Google Scholar]
- 53.Mitrousi K, Lip GYH, Apostolakis S. Age as a risk factor for stroke in atrial fibrillation patients: implications in thromboprophylaxis in the era of novel oral anticoagulants. J Atr Fibrillation. 2013;6(1):783. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Kopin L, Lowenstein CJ. Dyslipidemia. Ann Intern Med. 2017;167(11):ITC81–96. [DOI] [PubMed] [Google Scholar]
- 55.Li Z-Z, Du X, Guo X-y, Tang R-b, Jiang C, Liu N, et al. Association between blood lipid profiles and atrial fibrillation: a case-control study. Med Sci Monit. 2018;24:3903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Watanabe E, Kiyono K, Hayano J, Yamamoto Y, Inamasu J, Yamamoto M, et al. Multiscale entropy of the heart rate variability for the prediction of an ischemic stroke in patients with permanent atrial fibrillation. PLoS One. 2015;10(9):e0137144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Yao X, Abraham NS, Alexander GC, Crown W, Montori VM, Sangaralingham LR, et al. Effect of adherence to oral anticoagulants on risk of stroke and major bleeding among patients with atrial fibrillation. J Am Heart Assoc. 2016;5(2):e003074. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Nguyen HT, Vasconcellos HD, Keck K, Reis JP, Lewis CE, Sidney S, et al. Multivariate longitudinal data for survival analysis of cardiovascular event prediction in young adults: insights from a comparative explainable study. BMC Med Res Methodol. 2023;23(1):23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Yu J, Yang X, Deng Y, Krefman AE, Pool LR, Zhao L, et al. Incorporating longitudinal history of risk factors into atherosclerotic cardiovascular disease risk prediction using deep learning. Sci Rep. 2024;14(1):2554. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Steyerberg EW, Vergouwe Y. Towards better clinical prediction models: seven steps for development and an ABCD for validation. Eur Heart J. 2014;35(29):1925–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Shah ND, Steyerberg EW, Kent DM. Big data and predictive analytics: recalibrating expectations. JAMA. 2018;320(1):27–8. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.



