Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2024 Nov 7;14:27088. doi: 10.1038/s41598-024-78120-z

Machine learning for outcome prediction in patients with non-valvular atrial fibrillation from the GLORIA-AF registry

Martha Joddrell 1,2,✉, Wahbi El-Bouri 1,2, Stephanie L Harrison 1,2, Menno V Huisman 3, Gregory Y H Lip 1,2,5, Yalin Zheng 1,4; GLORIA-AFinvestigators
PMCID: PMC11544011  PMID: 39511367

Abstract

Clinical risk scores that predict outcomes in patients with atrial fibrillation (AF) have modest predictive value. Machine learning (ML) may achieve greater results when predicting adverse outcomes in patients with recently diagnosed AF. Several ML models were tested and compared with current clinical risk scores on a cohort of 26,183 patients (mean age 70.13 (standard deviation 10.13); 44.8% female) with non-valvular AF. Inputted into the ML models were 23 demographic variables alongside comorbidities and current treatments. For one-year stroke prediction, ML achieved an area under the curve (AUC) of 0.653 (95% confidence interval 0.576–0.730), compared to the CHADS2 and CHA2DS2-VASc scores performance of 0.587 (95% CI 0.559–0.615) and 0.535 (95% CI 0.521–0.550), respectively. Using ML for one-year major bleed prediction increased the AUC from 0.537 (95% CI 0.518–0.557) generated by the HAS-BLED score to 0.677 (95% CI 0.619–0.724). ML was able to predict one-year and three-year all-cause mortality with an AUC of 0.734 (95% CI 0.696–0.771) and 0.742 (95% CI 0.718–0.766). In this study a significant improvement in performance was observed when transitioning from clinical risk scores to machine learning-based approaches across all applications tested. Obtaining precise prediction tools is desirable for increased interventions to reduce event rates.

Trial Registryhttps://www.clinicaltrials.gov; Unique identifier: NCT01468701, NCT01671007, NCT01937377.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-78120-z.

Subject terms: Interventional cardiology, Mathematics and computing

Introduction

Risk stratification scores are used to determine the likelihood of an outcome occurring, guiding appropriate treatment and therapy interventions. Existing methods for stroke and major bleeding prediction in patients with atrial fibrillation (AF) are typically developed using traditional statistical approaches, hence, in many cases under-perform and are “oversimplified”1. One study demonstrated that four commonly-used cardiovascular risk stratification tools overestimated the intended outcome risk by 8-67% in women and 37-154% in men2.

Clinical scores are often designed to be easy-to-use and make fast predictions without the need for extensive training. However, these scores are frequently created with strict requirements including stringent feature composition that limit predictive ability3. Lack of generalisation of these tools results from their development on outdated cohorts. With life expectancy increasing and treatment options progressing, patients today have different factors contributing to disease development, with this change known as `data shift’4.

AF is the most frequent arrhythmia worldwide5, and is often asymptomatic but carries an increased mortality and morbidity risk6. To reduce the risk of stroke, AF patients are usually recommended oral anticoagulation, which needs to be balanced against a potential increase in bleeding likelihood7,8.

Approaches exploring improvement of current risk scores have begun exploiting machine learning (ML), which embeds greater complexity and interactions of information. Recent studies have shown improvement through ML when predicting AF9,10 with subsequent validation using an external data11, predicting the risk of developing AF post-stroke12, and predicting outcomes in individuals with AF, including numerous studies on stroke prediction13.

In contrast, some have reported that ML does not always improve prediction. A 2022 review concluded a limited indication that ML can move beyond classical scores or basic logistic regression when attempting to predict a specific AF-associated risk14. Auxiliary analyses showed that ML did not achieve superior performance over clinical scores when predicting stroke, major bleeding, or mortality on two AF-based registries15. Therefore, further studies are needed to confirm the appropriateness, applicability, and performance on a range of diverse datasets with associated and extensive model validation.

In the present study, validated clinical risk scores for stroke and major bleeding are compared with ML-based approaches that incorporate a more complex formulation between a larger set of features to assess absolute difference in approaches. Comparative performance is deemed in terms of area under the curve (AUC), sensitivity, and specificity, with optimal models selected by highest AUC value. Prediction of all-cause mortality using ML models are also reported.

Methods

Many studies provide a comprehensive overview of current clinical risk scores for AF-associated outcomes16,17. Recommendations for standard reporting18,19 are followed to ensure transparency, including a critical appraisal’s suggestions for cardiovascular artificial intelligence (AI) applications20.

Study population

The GLORIA-AF registry21 is an observational, prospective cohort of 37,235 patients with recently diagnosed non-valvular AF at risk of a stroke. Data were collected between 2011 and 2020 from a global registry program located in over 600 locations within 21 countries. For the purpose of this analysis, not all patients were included. Material S1 details the inclusion and exclusion criteria. Figure 1 demonstrates the final cohort of 26,183 used for this analysis. As the GLORIA-AF registry contains hundreds of variables, for this analysis 23 patient history, demographic, comorbidities, and medication variables were manually selected based on clinical expertise. All 23 variables (Table 1) were included in the model.

Fig. 1.

Fig. 1

Patient inclusion criteria resulting in 26,183 patients.

Table 1.

Variable characteristics that will be included in the models, along with the associated mean, standard deviation, and percentage of the variable that contains missing data.

Baseline characteristics Total cohort (% population) Mean (SD) Missing (% population)

Age (years)

 18–44

 45–54

 55–64

 65–74

 75–90

423 (1.6)

1549 (6.0)

4769 (18.2)

9485 (36.2)

9541 (36.4)

70.13 (10.13) 416 (1.6)

Gender

 Female

 Male

11,733 (44.8)

14,450 (55.2)

0 (0.0)
Heart rate 80.1 (21.3) 210 (0.8)
Systolic blood pressure 132.2 (18.7) 196 (0.8)
Diastolic blood pressure 78 (12) 199 (0.8)
Height (cm) 168.2 (10.3) 265 (1.0)
Weight (kg) 81.3 (20.6) 219 (0.8)
BMI 28.6 (6.3) 285 (1.0)

Treatment Group

 Apixaban

 ASA

 Dabigatran

 Edoxaban

 Rivaroxaban

 VKA

 Antiplatelets other than ASA

 None

4505 (17.2)

2163 (8.3)

8722 (33.3)

332 (1.3)

4015 (15.3)

4836 (18.5)

213 (0.8)

1386 (22.8)

11 (0.04)

Region

 Africa/Middle East

 Asia

 Europe

 Latin America

 North America

314 (1.2)

4905 (18.7)

12,993 (49.6)

2007 (7.7)

5964 (22.8)

0 (0.0)

Race

 Arab/Middle East

 Asian

 Black/Afro-Caribbean

 White

 Other

331 (1.3)

4557 (17.4)

453 (1.7)

18,200 (69.5)

895 (3.4)

1747 (6.7)

Alcohol use

 No alcohol

 Less than 1 drink/week

 1–7 drinks/week

 More than 7 drinks/week

11,411 (43.6)

6312 (24.1)

5094 (19.4)

1770 (6.8)

1596 (6.1)

Smoking status

 Never smoked

 Current smoker

 Ex-smoker

15,155 (57.9)

2443 (9.3)

7750 (29.6)

835 (3.2)
Types of AF 0 (0.0)
Paroxysmal 14,543 (55.5)
Persistent 9000 (34.4)
Permanent 2649 (10.1)
Hypertension 19,671 (75.7) 56 (0.2)
Diabetes mellitus 6067 (23.2)
Coronary artery disease 4932 (18.8) 662 (2.5)
Peripheral artery disease 756 (2.9) 196 (0.7)
Previous thromboembolism 3889 (14.9) 0 (0.0)
Previous myocardial infarction 2492 (9.5) 18 (0.07)
Previous deep vein thrombosis 303 (1.2) 315 (1.2)
Complex aortic plaque 225 (1.0) 5938 (22.7)
Respiratory disease 2923 (11.2) 270 (1.0)
Outcomes
 Major bleed 873 (3.3) 0 (0.0)
 Stroke 681 (2.6) 0 (0.0)
 All-cause death 2328 (8.9) 0 (0.0)

GLORIA-AF is a global registry, and the study was approved by local institutional review boards at each participating centre. There were multiple participating centres, in which the protocol was approved; these are listed in ClinicalTrials.gov (NCT01468701, NCT01671007, and NCT019373770). All participants provided informed consent. All studies were performed in accordance with the Declaration of Helsinki.

Outcomes

Three outcomes are predicted in the subsequent analysis; stroke, major bleeding, and all-cause mortality. Characterisation of these outcomes is provided in Material S2.

Data pre-processing

Data were initially split by a 70:30 ratio into training and testing sets before further pre-processing on the training cohort. Multiple imputation (m = 5, maxit = 50) was performed to avoid a substantial reduction in sample size. Random Over-Sampling Examples (ROSE) was implemented to counteract the outcome variable’s large class imbalance. More detail on both multiple imputation and ROSE method is provided in Material S3. Normalisation was applied to numerical variables and one-hot encoding to categorical variables with more than two levels.

Current clinical risk scores for stroke and major bleeding

The CHADS2score was developed to estimate the risk of stroke in AF patients22. A total of 6 points can be obtained (Congestive heart failure + 1, Hypertension + 1, Age Inline graphic 75 + 1, Diabetes mellitus + 1, previous Stroke/TIA history + 2) via this score. Patients with a score of Inline graphic2 are generally considered for anti-coagulation. The CHA2DS2-VASc score8, also calculates the risk of stroke; included are the same risk factors except for: Age +1 when 65–74 and +2 Inline graphic 75, Vascular disease +1 (including prior myocardial infraction, peripheral artery disease, or aortic plaque), and Sex: female +1. Here, a total of 9 points can be summed; again, oral anti-coagulation is recommended for anyone with a score of 2 or more.

The HAS-BLED Score23 is a 9 point-based score used to determine the risk of major bleeding in patients prescribed anticoagulation. Within this score, a summation of the risk factors includes: Hypertension + 1, Abnormal renal or liver function + 1–2, Stroke history + 1, Bleeding predisposition or tendency + 1, Labile INR + 1, Age Inline graphic 65 (Elderly) +1, and Drugs (for bleeding predisposition such as aspirin or NSAIDs) or alcohol + 1–2.

When calculating the performance of clinical risk scores, any patient classified as high risk by these scores within the specified study period would receive a prediction of stroke. Although the moderate classes still carry a risk of stroke, for the purpose of this analysis they were assigned ‘no stroke’ which allows insight into how the score performs under certain conditions - else, ~ 100% of patients would have a positive prediction if the combination of moderate- and high-risk was used.

Machine learning approaches

ML classification models included logistic regression (LR)24, random forest (RF)25, linear discriminant analysis (LDA)26, naive Bayes (NB)27, eXtreme Gradient Boosting (XGB)28, and neural network (NN)29.

All models were trained with 10-fold cross validation. For appropriate comparison with the clinical risk scores, both 3-year (total dataset) and 1-year (event within 12 months) cohorts were tested. The models were assessed using AUC, sensitivity, and specificity but performance was primarily determined based on the AUC. DeLong’s statistical significance test for comparing AUCs was used to differentiate improvements between methods at the 95th percentile. AUC confidence intervals at the 95% level were also reported. Analyses were conducted using python and R. More information on model development is provided in Material S4.

Results

Population characteristics

The mean age of the population was 70.13 (standard deviation (SD) 10.13) and 44.8% were female. Table 1 provides further detail on the overall population. Within the supplementary material, the population characteristics of those who had the outcome of stroke (Material S7) and major bleeding (Material S8) is provided. All 23 variables (not including the outcomes) in Table 1 were used to build the ML models.

Clinical risk scores

Of those with complete data (n = 26,183), 681 (2.6%) patients had an outcome of stroke over the 3-year study period. However, 22,520 patients (86%) had been classified as high risk by the CHA2DS2-VASc score (score Inline graphic 2). Contrasting, 14,994 (57.3%) patients were deemed high risk (score Inline graphic 2) by the CHADS2 score. Over the total study period, the number of patients who experienced a major bleed was 873 (3.3%). Consequently, the HAS-BLED score deemed 2,323 (8.9%) patients as high risk (score Inline graphic 3). Table 2 displays the number of patients obtaining each level of these current clinical risk scores. The mean (SD) of the CHADS2, CHA2DS2-VASc, and HAS-BLED scores were 1.9 (1.2), 3.2 (1.5), and 1.4 (0.9) respectively. Figure 2 demonstrates the complexity of predicting events using these simple risk categorisation methods.

Table 2.

Performance metrics of the CHADS2, CHA2DS2-VASc, and HAS-BLED clinical risk scores.

Sensitivity Specificity AUC 95% CI (AUC)

1-year

 CHADS

 CHA2DS2-VASc

HAS-BLED

0.761

0.948

0.155

0.413

0.123

0.92

0.587

0.535

0.537

(0.559, 0.615)

(0.521, 0.55)

(0.518, 0.557)

3-year

 CHADS2

 CHA2DS2-VASc

 HAS-BLED

0.738

0.933

0.125

0.415

0.123

0.921

0.577

0.528

0.522

(0.557, 0.6)

(0.517, 0.54)

(0.51, 0.535)

Fig. 2.

Fig. 2

Bar plot displaying the proportion of patients receiving moderate/high risk outcomes from the clinical risk scores, stratified by the proportion that had the outcome.

The 1- and 3-year AUCs calculated for stroke risk by the CHADS2 score were 0.587 (95% CI 0.559–0.615) and 0.577 (95% CI 0.557-0.600), respectively and for the CHA2DS2-VASc score, 0.535 (95% CI 0.521–0.550) and 0.528 (95% CI 0.517–0.540). When assessing 1-year stroke prediction as the two scores were designed for, the CHADS2 risk score was able to correctly classify 0.761 of actual strokes (recall), however it only achieved a precision value of 0.013 as it over-predicted the number of patients who would have a stroke, classifying almost half of the patients as stroke (50.4%) when in reality this value was much lower (0.026). Additionally, it obtained a 1-year/3-year sensitivity of 0.761/0.738 and specificity of 0.413/0.415. CHA2DS2-VASc performed similarly with a recall of 0.948 and precision value of 0.011. However, this score classified a significantly larger number of patients as predicted stroke (75.2%) when in reality the true rate of stroke in the observed study duration was 2.6%. The associated sensitivities and specificities for 1-year/3-year prediction were 0.948/0.933 and 0.123/0.123, respectively.

When estimating major bleeding risk, the HAS-BLED score obtained a 1-year AUC of 0.537 (95% CI 0.518-557), decreasing marginally to 0.522 (95% CI 0.510–0.535) for the 3-year cohort. The HAS-BLED score had a specificity of 0.921, compared with a sensitivity value of 0.125 for 3-year prediction, representing a very low rate of outcome detection. Similarly, at 1-year prediction the sensitivity and specificity values achieved were 0.155 and 0.920, respectively. The precision of this score for 1- and 3-year prediction was reported at 0.030 and 0.054, respectively, demonstrating over-prediction of the outcome compared with the true event rate.

Performance of the CHADS2 and CHA2DS2-VASc scores for 1-year stroke risk was similar to 3-year stroke risk (AUCs CHADS2 1-/3-year: 0.587/0.577, p-value 0.547 and AUCs CHA2DS2-VASc 1-/3-year: 0.535/0.528, p-value 0.452). No statistically significant difference was observed between the AUCs of 1- and 3-year prediction of major bleeding (p-value 0.215; AUC 1-year 0.537, 3-year 0.522). Figure 3 displays the ROC curves for the clinical risk scores. Subsequently, it was tested to determine if altering the numerical threshold categorising patients into ‘stroke’/‘no stroke’ had any effect on performance of the methods (Materials S5).

Fig. 3.

Fig. 3

1-year and 3-year clinical risk score (CHADS2, CHA2DS2-VASc, and HAS-BLED) receiver operating characteristic (ROC) curves for the prediction of stroke and major bleeding.

Machine learning models

Figure 4 displays the ROC curves for all models. When determining performance, it is important to consider the threshold boundary for classification. In practice, sensitivity or specificity will be optimized to reflect the consequences and clinical priorities of an application. Material S6 provides a more detailed explanation of the need for a trade-off.

Fig. 4.

Fig. 4

1-year and 3-year machine learning ROCs for prediction of stroke, major bleed, and mortality.

Table 3 reports the metrics when sensitivity and specificity have been balanced by thresholding for each model. Beyond that, the thresholds have been adjusted to maximise sensitivity whilst attempting to keep the specificity at an acceptable rate (> 50%), to increase the degree of events captured. Typically, the Youden Index is used to determine the appropriate threshold, however it attempts to balance both sensitivity and specificity, which for this application is not ideal. Table 4 gives the performance of the machine learning models maximised for either sensitivity or specificity.

Table 3.

Performance of 1-year and 3-year ML models for the prediction of stroke, major bleed, and all-cause mortality.

Balanced Optimised AUC 95% CI AUC
Sensitivity Specificity Sensitivity Specificity
1-year stroke
 Logistic regression 0.623 0.617 0.698 0.530 0.653 (0.576, 0.73)
 Random forest 0.623 0.615 0.660 0.602 0.634 (0.556, 0.712)
 Linear discriminant analysis 0.623 0.615 0.698 0.528 0.653 (0.577, 0.73)
 Naïve Bayes 0.604 0.561 0.642 0.542 0.625 (0.541, 0.709)
 XGBoost 0.604 0.565 0.642 0.527 0.633 (0.549, 0.717)
Neural network 0.566 0.533 0.566 0.533 0.562 (0.478, 0.646)
3-year stroke
 Logistic regression 0.602 0.594 0.720 0.504 0.653 (0.603, 0.704)
 Random forest 0.627 0.604 0.695 0.519 0.650 (0.597, 0.703)
 Linear discriminant analysis 0.602 0.594 0.720 0.504 0.654 (0.604, 0.704)
 Naïve Bayes 0.610 0.606 0.703 0.513 0.626 (0.578, 0.674)
 XGBoost 0.636 0.633 0.703 0.573 0.652 (0.599, 0.705)
 Neural network 0.576 0.554 0.619 0.501 0.596 (0.544, 0.649)
1-year major bleed
 Logistic regression 0.618 0.616 0.708 0.522 0.677 (0.62, 0.735)
 Linear discriminant analysis 0.629 0.614 0.708 0.524 0.677 (0.619, 0.734)
 Naïve Bayes 0.607 0.565 0.719 0.506 0.643 (0.583, 0.703)
 XGBoost 0.652 0.641 0.753 0.527 0.662 (0.605, 0.719)
 Neural network 0.640 0.631 0.742 0.506 0.670 (0.615, 0.724)
3-year major bleed
 Logistic regression 0.621 0.607 0.703 0.505 0.655 (0.616, 0.695)
Random forest 0.621 0.610 0.736 0.503 0.656 (0.616, 0.696)
 Linear discriminant analysis 0.621 0.610 0.703 0.501 0.655 (0.616, 0.695)
Naïve Bayes 0.593 0.590 0.698 0.500 0.629 (0.589, 0.669)
 XGBoost 0.599 0.587 0.670 0.501 0.633 (0.594, 0.672)
 Neural network 0.637 0.632 0.714 0.507 0.649 (0.61, 0.688)
1-year mortality
 Logistic regression 0.667 0.650 0.818 0.510 0.733 (0.695, 0.771)
 Linear discriminant analysis 0.660 0.657 0.818 0.511 0.734 (0.696, 0.771)
 Naïve Bayes 0.648 0.626 0.755 0.501 0.680 (0.638, 0.721)
 XGBoost 0.673 0.600 0.686 0.59 0.686 (0.647, 0.725)
 Neural network 0.660 0.656 0.786 0.500 0.716 (0.677, 0.756)
3-year mortality
 Logistic regression 0.684 0.683 0.827 0.505 0.742 (0.719, 0.766)
 Linear discriminant analysis 0.684 0.682 0.827 0.507 0.742 (0.719, 0.766)
 Naïve Bayes 0.616 0.612 0.733 0.500 0.656 (0.63, 0.682)
 XGBoost 0.689 0.647 0.754 0.598 0.719 (0.695, 0.744)
 Neural network 0.674 0.668 0.817 0.504 0.729 (0.705, 0.753)

Left: balanced sensitivity and specificity, right: optimised sensitivity. Bold highlights the highest AUC of each application. 

Table 4.

Performance of 1-year and 3-year machine learning models for the prediction of stroke, major bleed, and all-cause mortality.

~ 90% sensitivity ~ 80% specificity AUC 95% CI AUC
Sensitivity Specificity Sensitivity Specificity
1-year stroke
  Logistic regression 0.906 0.248 0.377 0.849 0.653 (0.576, 0.730)
 Random forest 0.887 0.219 0.415 0.79 0.634 (0.556, 0.712)
 Linear discriminant analysis 0.906 0.247 0.377 0.850 0.653 (0.577, 0.730)
 Naïve Bayes 0.906 0.184 0.330 0.792 0.625 (0.541, 0.709)
 XGBoost 0.905 0.233 0.263 0.811 0.633 (0.549, 0.717)
 Neural network 0.887 0.157 0.302 0.809 0.562 (0.478, 0.646)
3-year stroke
 Logistic regression 0.907 0.348 0.398 0.806 0.653 (0.603, 0.704)
Random forest 0.881 0.204 0.415 0.797 0.650 (0.597, 0.703)
 Linear discriminant analysis 0.907 0.249 0.415 0.799 0.654 (0.604, 0.704)
Naïve Bayes 0.907 0.213 0.331 0.800 0.626 (0.578, 0.674)
 XGBoost 0.890 0.216 0.398 0.797 0.652 (0.599, 0.705)
 Neural network 0.907 0.172 0.339 0.806 0.596 (0.544, 0.649)
1-year major bleed
 Logistic regression 0.910 0.290 0.438 0.805 0.677 (0.620, 0.735)
 Linear discriminant analysis 0.898 0.303 0.438 0.805 0.677 (0.619, 0.734)
 Naïve Bayes 0.910 0.187 0.404 0.812 0.643 (0.583, 0.703)
 XGBoost 0.880 0.266 0.202 0.895 0.662 (0.605, 0.719)
 Neural network 0.910 0.306 0.360 0.822 0.670 (0.615, 0.724)
3-year major bleed
 Logistic regression 0.901 0.268 0.423 0.802 0.655 (0.616. 0.695)
 Random forest 0.896 0.236 0.346 0.813 0.656 (0.616, 0.695)
 Linear discriminant analysis 0.901 0.271 0.423 0.800 0.655 (0.616, 0.695)
 Naïve Bayes 0.890 0.226 0.385 0.800 0.629 (0.589, 0.669)
 XGBoost 0.901 0.235 0.363 0.802 0.633 (0.594, 0.672)
Neural network 0.901 0.219 0.352 0.800 0.649 (0.610, 0.688)
1-year mortality
 Logistic regression 0.906 0.372 0.503 0.798 0.733 (0.695, 0.771)
 Linear discriminant analysis 0.906 0.373 0.503 0.799 0.734 (0.696, 0.771)
 Naïve Bayes 0.906 0.208 0.403 0.804 0.680 (0.638, 0.721)
 XGBoost 0.912 0.308 0.365 0.814 0.686 (0.647, 0.725)
 Neural network 0.906 0.315 0.484 0.804 0.716 (0.677, 0.756)
3-year mortality
 Logistic regression 0.902 0.370 0.522 0.800 0.742 (0.719, 0.766)
 Linear discriminant analysis 0.902 0.370 0.518 0.800 0.742 (0.719, 0.766)
 Naïve Bayes 0.902 0.280 0.370 0.802 0.656 (0.630, 0.682)
 XGBoost 0.906 0.334 0.436 0.817 0.719 (0.695, 0.744)
 Neural network 0.902 0.363 0.473 0.800 0.729 (0.705, 0.753)

Left ~ 90% sensitivity and ~ 80% specificity; right optimised sensitivity. Bold highlights the highest AUC of each application. 

When predicting stroke across the total cohort (3-year), the greatest AUC of 0.654 (95% CI 0.604–0.704) was generated by linear discriminant analysis. Improvement over the CHADS2 and CHA2DS2-VASc scores, which obtained respective 1-year AUCs of 0.570 (95% CI 0.559–0.615) and 0.535 (95% CI 0.521–0.550) was achieved using logistic regression, giving an AUC of 0.653 (95% CI 0.576–0.730). Most influential towards prediction of stroke within 1-year was age, race, and previous thromboembolism; two of which are incorporated within the existing clinical risk scores. For 3-year stroke, age, systolic blood pressure, previous thromboembolism, and complex aortic plaque were found to be significant predictors.

Achieving significantly higher performance than the HAS-BLED score (1-year AUC 0.537 95% CI 0.518–0.557; 3-year AUC 0.522 95% CI 0.510–0.535), the best ML model for 1-year and 3-year prediction of major bleeding was found through linear discriminant analysis and random forests, respectively, with corresponding AUCs of 0.677 (95% CI 0.619–0.724) and 0.656 (95% CI 0.616–0.696). Due to the HAS-BLED score being heavily skewed towards specificity for this dataset, the optimised sensitivities found through machine learning greatly improved from 0.155 to 0.708 for 1-year prediction and 0.125 to 0.736 when predicting 3-year likelihood. For prediction of both 1- and 3-year major bleeding, the two most influential predictors were age and region.

All the best performing ML models for each application achieved a statistically significant improvement (p-value < 0.05) in AUC over their comparative clinical score, with the exception of the 1-year application of the CHADS2 score (p-value = 0.114).

Mortality prediction of 1- and 3-year resulted in AUCs of 0.734 (95% CI 0.696–0.771) and 0.742 (95% CI 0.719–0.766), respectively, both through linear discriminant analysis. Predictors influential to mortality across both 1- and 3-year tasks were age, diastolic blood pressure, treatment of dabigatran or VKA, and the presence of respiratory disease.

Discussion

A statistically significant increase using ML was obtained through all 3-year models and over both the 1-year CHA2DS2-VASc score and HAS-BLED score. Due to the cohort being comprised of AF patients entailing a greater risk of stroke, scores such as CHA2DS2-VASc will reflect this, classifying almost all patients as high risk. Additionally, the developed ML techniques are similarly hindered in their ability to classify low-risk patients, despite achieving greater performance.

Although ML comes with increased methodology complexity, discovery of alternative influential variables not currently incorporated into existing methods has been identified. As ML can embed an intricate combination of more risk factors, it allows for greater reflection of everyone - which is advantageous in personalised medicine.

AI is generating superior performance for clinical risk scores; however, this paper shows that these models are not faultless in all populations. Conflicting debates still occur regarding ML performance, and it is vital that similar studies continue to highlight concerns such as robustness and appropriateness in mediocre-performing cohorts.

This study was conducted within the constraints of available data, however external validation is a crucial step in ensuring a model’s generalisability in future cohorts which is a future avenue for this work. Due to conflicting reports of ML performance in which discrepancy mostly stems from the varying appropriateness of datasets, be it: data size, quality, or relevance, additional applications should use divergent sources to solidify the debate. Since previous work has shown greater predictive ability of outcomes using ML, use of GLORIA-AF for external validation of is encouraged.

As GLORIA-AF is a global registry, it is advantageous for reducing the impact of bias, however the vast geographical span may limit learning trends that predict a certain outcome in a community. Given the number of outcomes for stroke and major bleeding is low in this population, it is unsurprising that performance is limited. An increase in performance when predicting all-cause mortality may stem from the larger proportion of reported outcomes - evidence for increased data collection to capture more positive cases. Additionally, a combination of GLORIA-AF with a larger database may yield greater calibration and superior results.

When working with clinical data, the presence of missing data is unavoidable. Although multiple imputation was suitable for this application, it may not match the assumptions of alternative implementations. Hence, for validation studies, the process of handling missing data will need to be re-evaluated.

A final limitation rests on the desideratum of these findings. One contribution of this study is the application to a differing population; those with newly-diagnosed AF who have predominately been anti-coagulated. Since anti-coagulation is known to reduce the risk of stroke and increase the risk of bleeding, the cohort may not accurately reflect true risk in populations not on anticoagulation or those with varied usage patterns. Facilitation on reducing their risk are uncertain and may require interventions such as dosage alteration. As the nature of the data containing primarily high-risk patients would have made a comparison of methods improbable, the choice to assign high-risk with an outcome ‘stroke’ and moderate-risk as ‘no stroke’ should be noted as a limitation when considering these findings.

AF patients at higher risk for embolic strokes are also at a higher risk for stroke due to atherosclerosis, but given the limitations of the GLORIA-AF data, it was not possible to fully ascertain whether these patients that had a stroke while on anticoagulants are actually those with a residual embolic risk despite anticoagulation, or those with a higher thrombotic risk.

Despite these limitations, there are numerous beneficial clinical implications of updating existing model’s including gaining a statistically significant increase in predictive performance, which can allow more timely interventions. Additionally, by embedding the most current and relevant data, tailored risk management can improve through individualised stratification. Models such as the ones in this study can be seamlessly integrated into clinical work flow and electronic health record (EHR) systems to enable real-time risk assessment.

Asides from building upon clinical risk scores, ML is now being embedded into mobile health (mHealth) applications to aid ‘real time’ patient care. Attributing to advances in cloud computing and accessibility to novel technologies, mHealth research and implementation is rapidly increasing. Through devices such as fitness watches, patients can be monitored continuously. Applications are also being implemented for diagnosis of AF30. Further development could transition into automated detection or prediction of changes in dynamic risk through time-series analysis. A recent review investigated mHealth apps relating to cardiovascular disease and reported on 38 studies for purposes including wearables for diagnosis31, emphasising this growing shift towards assistant automated healthcare.

Conclusion

Current clinical risk scores for assessing the risk of stroke and major bleeding remain modest in performance hence spurring the development of more precise approaches. Machine learning has the potential to improve prediction by incorporating a more complex combination of risk factors. Any gain in outcome prediction will result in greater aversion of events but also reduce over-prescription to those without need of preventative treatment.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 1 (52.7KB, docx)

Acknowledgements

The authors would like to thank all participants of the registry and all affiliated study personnel. Data has been made available through Vivli, Inc., however they did not approve, contribute or are responsible for the publication contents.

Author contributions

Conceptualization, W.E.B., S.L.H., M.V.H., G.Y.H.L. and Y.Z.; methodology, M.J., W.E.B. and Y.Z.; software, M.J., G.Y.H.L. and Y.Z.; validation, M.J.; formal analysis, M.J.; investigation, M.J., W.E.B., S.L.H., M.V.H., G.Y.H.L. and Y.Z.; resources, S.L.H., M.V.H., G.Y.H.L. and Y.Z.; data curation, M.V.H. and G.Y.H.L.; writing—original draft preparation, M.J.; writing—review and editing, W.E.B., S.L.H., M.V.H., G.Y.H.L. and Y.Z.; visualization, M.J.; supervision, W.E.B., S.L.H., G.Y.H.L. and Y.Z.; project administration, W.E.B., S.L.H., M.V.H., G.Y.H.L. and Y.Z.; funding acquisition, G.Y.H.L. All authors have read and agreed to the published version of the manuscript.

Funding

Boehringer Ingelheim GmbH sponsored the GLORIA-AF registry.

Data availability

The data that support the findings of this study are available from Boehringer Ingelheim but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available. Data are however available upon reasonable request and with permission of Boehringer Ingelheim (https://trials.boehringer-ingelheim.com/).

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Siontis, K. C. et al. How will machine learning inform the clinical care of atrial fibrillation?. Circul. Res.127(1), 155–169 (2020). [DOI] [PubMed] [Google Scholar]
  • 2.DeFilippis, A. P. et al. An analysis of calibration and discrimination among multiple cardiovascular risk scores in a modern multiethnic cohort. Ann. Intern. Med.162 (4), 266–275 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Ustun, B. & Rudin, C. Learning optimised risk scores. J. Mach. Learnign Res.20 (150), 1–75 (2019). [Google Scholar]
  • 4.Webb, G. I. et al. Analyzing concept drift and shift from sample data. Data Min. Knowl. Disc32, 1179–1199 (2018). [Google Scholar]
  • 5.Benjamin, E. J. et al. Heart disease and stroke statistics-2019 update: A report from the American Heart Association. Circulation139(10), e56–e528 (2019). [DOI] [PubMed] [Google Scholar]
  • 6.Wolf, P. A., Abbott, R. D. & Kannel, W. B. Atrial fibrillation as an independent risk factor for stroke: The Framingham study. Stroke22(8), 983–988 (1991). [DOI] [PubMed] [Google Scholar]
  • 7.Chao, T. F. et al. 2021 Focused update consensus guidelines of the asia pacific heart rhythm society on stroke prevention in atrial fibrillation: Executive summary. Thromb. Haemost. 122(1), 20–47 (2022). [DOI] [PMC free article] [PubMed]
  • 8.Lip, G. Y. et al. Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor-based approach: The euro heart survey on atrial fibrillation. Chest137(2), 263–272 (2010). [DOI] [PubMed] [Google Scholar]
  • 9.Tiwari, P. et al. Assessment of a machine learning model applied to harmonized electronic health record data for the prediction of incident atrial fibrillation. JAMA Netw. Open3(1). e1919396 (2020). [DOI] [PMC free article] [PubMed]
  • 10.Hill, N. R. et al. Predicting atrial fibrillation in primary care using machine learning. PLoS One14(11) (2019). [DOI] [PMC free article] [PubMed]
  • 11.Sekelj, S. et al. Detecting undiagnosed atrial fibrillation in UK primary care: Validation of a machine learning prediction algorithm in a retrospective cohort study. Eur. J. Prev. Cariol.28 (6), 598–605 (2021). [DOI] [PubMed] [Google Scholar]
  • 12.Zheng, X. et al. Using machine learning to predict atrial fibrillation diagnosed after ischemic stroke. Int. J. Cardiol.347, 21–27 (2022). [DOI] [PubMed] [Google Scholar]
  • 13.Watanabe, E. et al. Comparison among random forest, logistic regression, and existing clinical risk scores for predicting outcomes in patients with atrial fibrillation: A report from the J-RHYTHM registry. Clin. Cardiol.44 (9), 1305–1315 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Isaksen, J. et al. Artificial intelligence for the detection, prediction, and management of atrial fibrillation. Herzschrittmacherther. Elektrophysiol.33 (1), 34–41 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Loring, Z. et al. Machine learning does not improve upon traditional regression in predicting outcomes in atrial fibrillation: An analysis of the ORBIT-AF and GARFIELD-AF registries. Eurospace. 22 (11), 1635–1644 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.O’Neal, W. T. & Alonso, A. The appropriate use of risk scores in the prediction of atrial fibrillation. J. Thorac. Disease8 (10), E1391–E1394 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Brieger, D. & Freedman, B. Decoding stroke risk scores in atrial fibrillation: Still more work to dO. Eur. Heart J.42 (15), 1486–1488 (2021). [DOI] [PubMed] [Google Scholar]
  • 18.Stevens, L. M. et al. Recommendations for reporting machine learning analyses in clinical research. Cardiovasc. Qual. Outcomes13, e00-6556 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.de Hond, A. et al. Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: A scoping review. NPJ Digit. Med.5(1). (2022). [DOI] [PMC free article] [PubMed]
  • 20.Smeden, M. et al. Critical appraisal of artificial intelligence-based prediction models for cardiovascular disease. Eur. Heart J.43 (31), 2921–2930 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Huisman, M. V. et al. Design and rationale of Global Registry on long-term oral antithrombotic treatment in patients with Atrial Fibrillation: A global registry program on long-term oral antithrombotic treatment in patients with atrial fibrillation. Am. Heart J.167(3), 329–334 (2014). [DOI] [PubMed] [Google Scholar]
  • 22.Gage, B. F. et al. Validation of clinical classification schemes for predicting stroke: Results from the National Registry of Atrial Fibrillation. JAMA285(22), 2864–2870 (2001). [DOI] [PubMed] [Google Scholar]
  • 23.Pisters, R. et al. A novel user-friendly score (HAS-BLED) to assess 1-year risk of major bleeding in patients with atrial fibrillation: The Euro Heart Survey. Chest138 (5), 1093–1100 (2010). [DOI] [PubMed] [Google Scholar]
  • 24.Cox, D. R. The regression analysis of binary sequences. J. Roy. Stat. Soc.: Ser. B (Methodol.)20 (2), 215–232 (1958). [Google Scholar]
  • 25.Ho, T. K. Random decision forests. In Proceedings of 3rd International Conference on Document Analysis and Recognition vol. 1, pp. 278–282 (1995).
  • 26.Fisher, R. A. The use of multiple measurements in taxonomic problems. Annals Eugenics7 (2), 179–188 (1936). [Google Scholar]
  • 27.Webb, G. I., Keogh, E. & Miikkulainen, R. Naive Bayes. Encyclopedia Mach Learn 15, 713–714 (2010).
  • 28.Chen, T. & Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–694 (2016).
  • 29.McCulloch, W. S. & Pitts, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biophys.5 (4), 115–133 (1943). [PubMed] [Google Scholar]
  • 30.Bonini, N. et al. Mobile health technology in atrial fibrillation. Expert Rev. Med. Dev.19 (4), 327–340 (2022). [DOI] [PubMed] [Google Scholar]
  • 31.Cruz-Ramos, N. A. et al. mHealth apps for self-management of cardiovascular diseases: A scoping review. Healthcare10(2), 322 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (52.7KB, docx)

Data Availability Statement

The data that support the findings of this study are available from Boehringer Ingelheim but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available. Data are however available upon reasonable request and with permission of Boehringer Ingelheim (https://trials.boehringer-ingelheim.com/).


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES