Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2026 Jun 22;26:334. doi: 10.1186/s12911-026-03636-5

Machine learning based prediction models for first stroke in community primary care: a systematic review and meta-analysis

Zijiao Zhang 1, Yuru Zhang 1, Zheng Li 2, Chenghua Tian 3, Jun Liang 3,4,5, Jun Xie 6,#, Jianbo Lei 7,8,9,✉,#
PMCID: PMC13540839  PMID: 42332728

Abstract

Background

Stroke is the third leading cause of death and the fourth leading cause of disability globally, particularly in low- and middle-income countries, where the disease burden is increasing significantly. Risk prediction of first-stroke is crucial for stroke prevention. Although artificial intelligence (AI)-based prediction models have rapidly developed in recent years, their clinical utility in primary community healthcare settings remains unclear. This study aimed to systematically review and analyze the application of algorithms, especially machine learning (ML)-based ones, in first-stroke risk prediction models within community settings.

Methods

We searched PubMed, EMBASE, and the Cochrane Library from inception to 30 June 2024 for studies developing or validating multivariable stroke risk models for primary care. We analyzed AUC/C-statistic values with 95% confidence intervals (CI) and conducted a meta-analysis.

Results

A total of 43 studies encompassing 93 models were included. The Cox regression model was the most common traditional method (n = 32, 74.42% of traditional models). Among ML models, 25 algorithms were identified, with logistic regression being the most frequent (n = 8, 18.60%). The meta-analysis showed a pooled AUC of 0.79 (95% CI: 0.77–0.81) for the Cox regression model, while logistic regression, random forest, and eXtreme Gradient Boosting models yielded 0.76 (95% CI: 0.64–0.89), 0.77 (95% CI: 0.61–0.93), and 0.77 (95% CI: 0.59–0.94), respectively. Notably, 76.7% of the included studies had a high risk of bias, and 20.9% raised high concerns regarding their applicability to community primary healthcare settings. ML models did not outperform traditional regression models, with significant performance variability observed.

Conclusion

Despite the increasing use of ML models for first-stroke risk prediction, their clinical utility in community primary healthcare remains unverified. Current ML models show no significant advantage over traditional regression methods and face challenges in algorithm innovation, data standardization, and external validation. Furthermore, the prevalence of methodological flaws and applicability concerns among existing studies underscores the need for caution. Future research should focus on standardized validation and real-world clinical analysis to develop effective stroke prevention tools.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1186/s12911-026-03636-5.

Keywords: First stroke, Prediction model, Community, Prevention, Machine learning, Meta-analysis

Background

Stroke is the third most common cause of death in the global burden of disease and the fourth leading cause of disability adjusted life years loss [1]. Over 100 million people worldwide have experienced a stroke, and its incidence rate increases significantly with age [2]. Particularly in low- and middle-income countries, the burden of stroke is rapidly rising [3]. Additionally, stroke treatment presents numerous challenges. Not only is the recovery period lengthy [4], but approximately half of stroke survivors experience varying levels of disability. These sequelae significantly diminish the quality of life for patients and pose substantial challenges to both their families and social support systems [5]. Therefore, intensifying stroke prevention efforts, optimizing treatment and rehabilitation strategies, and reducing the financial and psychological burdens on patients and their families have emerged as crucial pressing issues in the current medical field.

Primary prevention remains the most effective strategy for stroke control. Research indicates that more than 85% of strokes are preventable using preventive measures [1, 6]. Currently, guidelines for primary stroke prevention recommend the use of risk prediction models to identify individuals at high risk for cardiovascular disease, including stroke [7]. These models are not only widely used for the prevention of cardiovascular diseases [8], but also for prognostic evaluation of diseases such as tumors [9, 10] and various health risk assessment fields, offering unique advantages [11]. In particular, these risk prediction models can effectively identify high-risk individuals who would greatly benefit from early intervention. Through interventions such as dietary adjustments and lifestyle modifications, it is possible to significantly reduce stroke incidence and alleviate the associated social and economic burden, thereby improving public health outcomes.

Numerous studies have explored stroke risk prediction models, emphasizing the development of more precise and reliable prediction methods. However, traditional regression-based models, while interpretable, often rely on pre-defined linear relationships and may not fully capture the complex, non-linear interactions among diverse risk factors (e.g., lifestyle, genetics, comorbidities). In recent years, machine learning (ML) technology has become prominent in stroke risk prediction precisely because it can automatically handle high-dimensional data and model these intricate, non-linear patterns without strong prior assumptions [12]. For example, ML algorithms such as support vector machines, neural networks, and decision trees have demonstrated superior performance in cardiovascular risk assessment by identifying subtle risk factor combinations that traditional models might miss, thereby markedly enhancing predictive accuracy [13].

Although many studies have been devoted to improving prediction accuracy, the real-world effectiveness of existing models, especially in community primary care settings, remains uncertain. Previous reviews have evaluated cardiovascular disease risk prediction accuracy in the general population [14, 15], while stroke, as only a part of the composite outcome, may be overlooked as a specific predictor required for independent disease. Different predictive factors may have opposite predictive effects in different disease components, leading to mutual cancellation of predictive contributions [16, 17]. For instance, the Framingham Heart Study highlighted significant differences in how serum cholesterol impacts the risk of coronary heart disease and stroke [18]. Moreover, reviews of independent stroke outcomes often focus on specific hospital patient groups, such as those with diabetes [19], atrial fibrillation [20] and dialysis [21]. However, research on the applicability of these models to the general population in community primary healthcare environments is relatively insufficient [22]. Although some studies have explored the effectiveness of ML modeling algorithms [23], they have not fully evaluated the feasibility and practicality of these models in community primary healthcare practices.

Therefore, drawing on the design principles of community-based heart failure prediction models [24], this study aims to comprehensively evaluate prediction models tailored for community primary care to assess individual first stroke risk using routine data collection. The ultimate goal was to determine the models that can be most effectively applied in these environments. Further, building on prior research, this study broadens its approach by incorporating the latest AI technology in literature, without being confined to specific algorithms. Specifically, this study, through comparative analysis of traditional and ML models, along with a meta-analysis, seeks to identify risk assessment and prevention strategies that offer greater precision and are better aligned with everyday healthcare needs for primary care.

Methods

This systematic review follows recommendations outlined in Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement [25] as well as Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: checklist for systematic reviews and meta-analyses (TRIPOD-SRMA) [26]. Our protocol was registered on PROSPERO (CRD42024581870).

Search strategy

We searched the PubMed、EMBASE and Cochrane Library databases from inception through 30 June 2024. We used a combination of keywords and subject headings related to stroke and prediction models, and the search was limited to the English-language (Supplementary Appendix P2-3 for details). We also performed forward and backward citation searching for more studies that are relevant.

Inclusion criteria

To be eligible for inclusion a study had to be an original study in human adults (≥ 18 years of age), develop and/or validate a prediction model(s) for first stroke based on multivariable analysis, and be written in English. Additionally, the predictive variables were easily accessible in the community primary care environment. (Supplementary Appendix P3-5 for details)

Study selection and data extraction

During screening, to ensure consistency of the results, two researchers independently screened the titles/abstracts and full texts of 50 randomly selected identical articles according to inclusion and exclusion criteria. A kappa value of ≥ 0.75 was achieved before proceeding with the remaining literature. Disagreements were resolved by a third researcher. The checklist for critical appraisal and data extraction for systematic reviews of prediction modeling studies (CHARMS) was used for data extraction [27]. For model evaluation metrics, C-statistic or the Area Under the Receiver Operating Characteristic Curve (AUC) and the corresponding 95% confidence interval (95% CI) were extracted.

Data synthesis and statistical analysis

Continuous variables are transferred as means, and categorical variables as percentages. We used a significance level of 0.05 to assess the statistical significance of all analyses. For evaluating model performance, we primarily focus on two core indicators: the C-statistic and AUC, which together reflect the model’s ability to distinguish between possible outcomes. Given that various studies use different performance metrics, such as the C-statistic and AUC, we employed specific methods to thoroughly summarize and analyze these two types of indicators.

Considering the significant differences in model construction, predictor selection, cohort composition, and result time frames among studies, we anticipate a high degree of heterogeneity among the included studies. Therefore, when summarizing the effect values of each study, we used a random-effects model and visually presented the results through a forest plot [28]. When the study only provided AUC values without reporting their 95% confidence intervals and standard error, we estimated the standard error using Hanley and McNeil’s method [29], combined with AUC values, sample size, and event ratio, and conducted a meta-analysis based on this. For the meta-analysis of C-statistic, we used Bayesian method for comprehensive analysis [30]. In addition, we also calculated a 95% prediction interval (PI) to comprehensively evaluate the heterogeneity between studies and the possible fluctuation range of model performance across different studies or datasets [31, 32]. A wider prediction interval often implies significant differences in predicted values between different studies or datasets, indicating a high degree of heterogeneity. Meanwhile, to quantify heterogeneity among studies, we use the Cochran’s Q test (where significant heterogeneity was defined as p ≤ 0·10 or I²>50%) [33].

During the data analysis process, we used Excel 2021 and R 4.4.0 software for data organization and analysis. We conducted meta-analyses in R using the metafor and metamisc package (R Foundation for Statistical Computing 4.4.0) [34–36].

Quality assessment

We used the prediction model risk of bias assessment tool (PROBAST) to assess the risk of bias (ROB) and the applicability of the included literature [37].

Results

Study selection

A total of 12,476 studies were identified in the initial search. Following the screening of title and abstract, and full text for eligibility, 34 studies met inclusion and exclusion criteria. We searched the reference list and found another 9 studies that met the criteria. Therefore, a total of 43 articles were ultimately included in this study (Fig. 1).

Fig. 1.

Fig. 1

Study selection

Characteristics of included studies

A total of 93 models were included in the study, including 43 regression models and 50 machine learning models. The study involved populations from 19 countries, with the top three countries being China (28.6%), the United States (10.7%), and Japan (8.9%). Three of the studies used datasets from multiple countries. The sample size of participants ranged from 1,203 to 5,715,311, with an age range of 18 to 97 years old, and the proportion of males ranged from 33.91% to 100%. (Supplementary Appendix Table S1 for details)

Characteristics of included prediction models

Among the prediction models, 39.5% included development and external validation, 41.9% focused on development, and 18.6% on external validation. Traditional methods like Cox regression were used in 74.42% of studies, while machine learning showed diversity with logistic regression (LR) (n = 8, 18.60%), random forest (RF) (n = 7, 16.28%), and decision trees (DT) (n = 4, 9.30%) as common. All models reported C-statistic or AUC; calibration methods included Hosmer-Lemeshow test (40.74%), calibration plot (14.81%), observation expectation ratio (14.81%), and others; 20 studies (46.52%) lacked calibration details (Figure S1). In the analysis of predictors, Key predictors included age, smoking status, diabetes, systolic blood pressure, and gender. (Supplementary Appendix Table S2-4 for details)

Risk of bias assessment

We used PROBAST to evaluate Risk of Bias, and the results showed that, overall, 76.7% of the studies had a high risk of bias (Fig. 2), mainly due to the high risk of bias in the field of analytical methods (76.7%), with the highest being due to improper handling of missing data (58.2%). Additionally, 20.9% of the studies were classified as having high applicability risk. (Supplementary Appendix Table S5 and Figure S2 for details)

Fig. 2.

Fig. 2

Risk of bias across all included studies

Meta-analysis

Firstly, we grouped machine learning models into three groups based on different algorithms and summarized three models that met the meta-analysis criteria: RF model, eXtreme Gradient Boosting (eXGBoost) model, and LR model. Specifically, the summary AUC value of the RF model is 0.77 (95% CI 0.61–0.93), the eXGBoost model is 0.77 (95% CI 0.59–0.94), and the LR model is 0.76 (95% CI 0.64–0.89). (Fig. 3) These AUC values collectively indicate that the three machine learning models mentioned above have demonstrated moderate predictive accuracy in identifying the risk of first-time stroke onset. However, it is worth noting that the aggregated results of all models exhibit significant heterogeneity, which further reveals the differences in predictive performance among different models.

Fig. 3.

Fig. 3

Summary AUC effect values of three machine learning models

Secondly, for traditional regression models, we also conducted a summary analysis, specifically based on the Cox proportional hazards regression algorithm, with a summary AUC value of 0.79 (95% CI 0.77–0.81). (Supplementary Material Figure S3) From this result, it can be seen that the limited evidence currently does not clearly demonstrate that machine learning models have significant performance advantages compared to traditional models.

In addition, during our data collection process, we found that some traditional models were validated based on at least two or more different population datasets, and we specifically conducted a meta-analysis on the C-statistical values of these models. There are four specific models: FSRS (Framingham Stroke Risk Score) -2017 model [38], Q-Stroke model [39], China PAR stroke risk model [40], and Stroke Riskometer™ model [41]. Figure 4 shows the aggregated effect values and 95% confidence intervals of these four models. Only the FSRS-2017 model and the China PAR stroke risk model have acceptable discriminative performance, with C-statistical values of 0.68 (95% CI 0.66–0.71) and 0.76 (95% CI 0.65–0.85), respectively. When the prediction time window is further limited, only the FSRS-2017 model has an acceptable discriminative power in predicting 10-year stroke risk with a total effect value of 0.70 (95% CI 0.68, 0.71) (Supplementary Material Figure S4). Furthermore, when we further restricted the inclusion of models in the meta-analysis to meet the “low” risk of bias, there were no models available for summary analysis.

Fig. 4.

Fig. 4

Summary c-statistical values of 4 traditional model

Discussions

Main research results

This study aims to evaluate the effectiveness and reliability of a stroke risk prediction model specifically designed for community primary healthcare environments. Through literature review, we observed a notable increase in the number of studies on first-time stroke risk prediction models in recent years among the 43 included studies, with the majority originating from China. This trend within our dataset aligns with the reported high stroke incidence in China. In order to systematically evaluate the performance of existing models, we screened 43 studies and extracted 93 risk prediction models for comprehensive analysis.

Firstly, in terms of model algorithms, the most common algorithm is still the traditional regression algorithm. There are a total of 25 machine learning algorithms involved, but there is a relative lack of exploration and innovation in new algorithms or models, and homogenization is also prominent. Although machine learning algorithms cover a wide range of techniques from logistic regression, decision trees to advanced ensemble learning methods such as random forests, gradient boosting decision trees, etc., the exploration and innovation of new algorithms or models are relatively insufficient. Many so-called ‘new methods’ often only involve fine-tuning or combining classic algorithms, lacking true innovation. For example, N-SRS (Nonlinear Stroke Risk Score) [42] combines the Optimal Classification Tree (OCT), it is still essentially a variant based on decision trees; CoxNET [43] combines Cox proportional hazards model and elastic network regularization, it does not break away from the traditional survival analysis framework. Furthermore, the phenomenon of homogenization is clearly manifested in researchers’ tendency to use widely validated and accepted algorithms, such as random forests, support vector machines, etc., while ignoring other algorithms that may be equally effective or even better. It may also reduce the comparability of research results and make it difficult to distinguish the true differences and innovative points of different studies. In addition, only 17 studies (39.53%) conducted external validation after model development, and none of the studies further evaluated the model by applying it to a real medical environment. The lack of validation in real-world medical environments has become a key factor limiting the reliability, accuracy, and clinical applicability of models.

Secondly, in the meta-analysis section, we specifically focused on the performance of machine learning models compared to traditional regression models. The results indicate that the performance advantage of machine learning models is not significant compared to traditional regression models. Specifically, the regression model demonstrated robust predictive ability, with a total AUC value of 0.79 (95% CI 0.77–0.81). The overall AUC value of the machine learning models is between 0.76 and 0.77, indicating that their discriminative ability is still moderate. However, the performance results of machine learning models vary greatly, some models perform well, while others fall short of expectations. The confidence interval range of the aggregated effect values (95% CI 0.59–0.94) of machine learning models is much larger than that of traditional models (95% CI 0.77–0.81). This may be highly correlated with the quality and processing methods of the training data available for each study. In addition, in the included studies, we found that the machine learning model trained by Chun et al. (2021) [44] based on 500,000 Chinese adult data had certain advantages over traditional models (the AUROC of the Gradient Boosted Trees model in males and females were 0.833 and 0.836, respectively, while the Framingham stroke risk model using traditional algorithms in males and females were 0.781 and 0.772, respectively). However, the comprehensive research conclusion of Hong et al. (2023) [43], which included 62,482 participants, showed that the novel machine learning technique did not significantly improve the model discrimination accuracy. Therefore, it is currently unclear whether machine learning models are superior to traditional regression models.

Furthermore, when we evaluate the bias risk of the model, we find that the processing and analysis of complex datasets are the main issues, especially the improper handling of missing datasets, which is a common problem in predictive modeling [45]. This improper handling leads to a high risk of bias in most models, which may weaken the reliability of research results.

Finally, we also found that an increasing number of scholars are working on developing more refined and personalized predictive models, providing new perspectives for model optimization by incorporating factors such as gender, race, and stroke sub-types. Chun et al. (2022) [46] emphasized the differences in the proportion of risk factors in different sub-types of stroke, which not only deepened our understanding of stroke risk factors, but also provided valuable insights for developing more precise primary prevention strategies.

Theoretical and practical value

Theoretically, the stroke risk prediction model contributes to the advancement of disease prevention theory by integrating machine learning algorithms with routine clinical data, thereby enabling a more nuanced understanding of risk stratification in primary care settings [47]. This integration supports the paradigm of precision medicine by offering individualized risk assessments that can inform targeted prevention strategies, even in resource-constrained environments.

Practically, the model enhances community healthcare delivery by identifying high-risk individuals who may otherwise remain undetected, facilitating early and cost-effective interventions. It empowers primary care providers to implement proactive, data-driven health management, which can reduce the burden on specialized healthcare services and improve population health outcomes. Moreover, the model’s reliance on routinely collected data ensures its scalability and sustainability in real-world primary care settings.

Strengths and limitations

We employed a detailed search strategy to ensure the extensive collection of relevant research and included only models validated in the general population to ensure reliability particularly for primary prevention in community healthcare environments. However, there are still some limitations to this study. Firstly, due to the inability to obtain model calibration metrics with sufficient data, our meta-analysis is limited to the discriminative metrics of the models. Secondly, this study is restricted to English literature, potentially overlooking significant non-English research [48]. Finally, we found that the studies included in the meta-analysis had high heterogeneity (I ²>99%). The reason for this is that traditional randomized trial meta-analyses use strict study designs with low heterogeneity, while predictive model meta-analyses have high heterogeneity due to the diversity of eligible study designs, wide population coverage, and increased complexity of required statistical methods [30]. Moreover, the heterogeneity of previous meta-analyses of similar risk prediction models is generally high (I²>50%, up to 99.9%) [49, 50].

Implications for future research

This review found that there are still significant limitations in the applicability of current initial stroke risk prediction models in community primary healthcare systems, and their clinical value still needs to be fully validated through real-world studies. Existing models face multiple challenges in terms of algorithm innovation, data standardization, and external validation validity. Firstly, most studies remain in the theoretical modeling stage and lack prospective clinical cohort validation; Secondly, there is still a lack of effective solutions for the integration of heterogeneous data across regions and populations; Furthermore, the dynamic predictive ability and clinical interpretability of existing models urgently need to be improved. Importantly, while this review quantitatively synthesized model discrimination (AUC/C-statistic), it is critical to recognize that discrimination alone is insufficient for clinical deployment. A model’s ability to distinguish between individuals who will develop the outcome and those who will not does not indicate the accuracy of its predicted probabilities. Calibration—the agreement between predicted and observed outcomes—is equally vital for clinical decision-making, as poorly calibrated models may lead to inappropriate risk stratification and subsequent clinical actions in primary care settings. Unfortunately, as noted in our limitations, the underreporting of calibration metrics in the primary studies precluded a quantitative synthesis of this aspect. Based on the TRIPOD + AI declaration framework [51], evidence-based analysis suggests that future research should focus on reconstructing a validation system that is suitable for clinical practice scenarios: while exploring new machine learning algorithms, it is necessary to establish a multi center clinical validation platform and integrate predictive models with grassroots diagnosis and treatment pathways for integrated testing. For the important issue of clinical validation of these models, relevant literature systematically and comprehensively elaborates on the principles, methods, and steps of evaluating or validating predictive model studies in clinical practice [52–54]. It is recommended that future research follow the guidelines of these literature and provide practical predictive tools for primary stroke prevention through standardized model calibration and clinical decision impact analysis.

Conclusion

Although a large number of machine learning models have been developed to predict the risk of initial stroke, the clinical application value of these models in community primary healthcare environments still needs further validation. Despite the promising predictive capabilities demonstrated in many studies, the effective deployment of these models in community primary care is hindered by significant challenges related to algorithm innovation, data standardization, and external validation validity. These limitations currently prevent a definitive conclusion about their real-world predictive value and practical usefulness. Future research should achieve the development of truly effective and practical stroke prevention and prediction tools through standardized model validation and real clinical application analysis.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (572.5KB, docx)

Acknowledgements

The authors thank all researchers for their supports.

Abbreviations

AI

Artificial intelligence

ML

Machine learning

PRISMA

Preferred Reporting Items for Systematic Reviews and Meta analysis

TRIPOD

Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis

AUC

Area Under the Receiver Operating Characteristic Curve

CI

Confidence intervals

RF

Random Forest

XGBoost

eXtreme Gradient Boosting

LR

Logistic Regression

FSRS

Framingham Stroke Risk Score

GBDT

Gradient boosting decision trees

N-SRS

Nonlinear Stroke Risk Score

Author contributions

ZZJ and XJ assisted in study conception and design. ZYR, LZ, TCH, and LJ assist in search and quality assessment. ZZJ and ZYR conducted data analysis. ZZJ wrote the initial draft, while XJ and LJB determined study design, revised the manuscript and final quality improvement.

Funding

This work was financially supported by the Natural Science Foundation of China [Grant number 72574006]; the Zhejiang Provincial Natural Science Foundation of China [Grant number LY22H180001]; the Key Research and Development Program of Zhejiang Province [Grant number 2024C03215]. The funding bodies played no role in the design of the study and collection, analysis, and interpretation of data and in writing the manuscript.

Data availability

All data generated or analysed during this study are included in this published article [and its supplementary information files].

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Jun Xie and Jianbo Lei contributed equally to this work.

References

  • 1.GBD 2021 Stroke Risk Factor Collaborators. Global, regional, and national burden of stroke and its risk factors, 1990–2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Neurol. 2024;23:973–1003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.World Stroke Organization. https://www.world-stroke.org/world-stroke-day-campaign/about-stroke/impact-of-stroke. Accessed 18 Jan 2025.
  • 3.Owolabi MO, Thrift AG, Mahal A, Ishida M, Martins S, Johnson WD, et al. Primary stroke prevention worldwide: translating evidence into action. Lancet Public Health. 2022;7:e74–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Crichton SL, Bray BD, McKevitt C, Rudd AG, Wolfe CDA. Patient outcomes up to 15 years after stroke: survival, disability, quality of life, cognition and mental health. J Neurol Neurosurg Psychiatry. 2016;87:1091–8. [DOI] [PubMed] [Google Scholar]
  • 5.Feigin VL, Owolabi MO, World Stroke Organization–Lancet Neurology Commission Stroke Collaboration Group. Pragmatic solutions to reduce the global burden of stroke: a World Stroke Organization-Lancet Neurology Commission. Lancet Neurol. 2023;22:1160–206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Sarikaya H, Ferro J, Arnold M. Stroke prevention–medical and lifestyle measures. Eur Neurol. 2015;73:150–7. [DOI] [PubMed] [Google Scholar]
  • 7.Goldstein LB, Bushnell CD, Adams RJ, Appel LJ, Braun LT, Chaturvedi S, et al. Guidelines for the Primary Prevention of Stroke. Stroke. 2011;42:517–84. [DOI] [PubMed] [Google Scholar]
  • 8.D’Agostino RB, Vasan RS, Pencina MJ, Wolf PA, Cobain M, Massaro JM, et al. General cardiovascular risk profile for use in primary care: the Framingham Heart Study. Circulation. 2008;117:743–53. [DOI] [PubMed] [Google Scholar]
  • 9.Huang C, Huang J, He Y, Zhao Q, Ming W-K, Duan X, et al. Competing-risks model for predicting the prognosis of patients with angiosarcoma based on the SEER database of 3905 cases. Holist Integ Oncol. 2024;3:13. [Google Scholar]
  • 10.Chen S, Shi C, Li B, Li L. Development of fibrotic gene signature and construction of a prognostic model in melanoma. Holist Integ Oncol. 2023;2:11. [Google Scholar]
  • 11.Adams ST, Leveson SH. Clinical prediction rules. BMJ. 2012;344:d8312. [DOI] [PubMed] [Google Scholar]
  • 12.Deo RC. Machine Learning in Medicine. Circulation. 2015;132:1920–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Amann J. Machine LearningMachine learning (ML) in StrokeStroke Medicine: Opportunities and Challenges for Risk PredictionRisk prediction and PreventionPrevention. In: Jotterand F, Ienca M, editors. Artificial Intelligence in Brain and Mental Health: Philosophical, Ethical & Policy Issues. Cham: Springer International Publishing; 2021. pp. 57–71. [Google Scholar]
  • 14.Damen JAAG, Hooft L, Schuit E, Debray TPA, Collins GS, Tzoulaki I, et al. Prediction models for cardiovascular disease risk in the general population: systematic review. BMJ. 2016;353:i2416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Liu W, Laranjo L, Klimis H, Chiang J, Yue J, Marschner S, et al. Machine-learning versus traditional approaches for atherosclerotic cardiovascular risk prognostication in primary prevention cohorts: a systematic review and meta-analysis. Eur Heart J Qual Care Clin Outcomes. 2023;9:310–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Gondrie MJA, Janssen KJM, Moons KGM, van der Graaf Y. A simple adaptation method improved the interpretability of prediction models for composite end points. J Clin Epidemiol. 2012;65:946–53. [DOI] [PubMed] [Google Scholar]
  • 17.Ferreira-González I, Busse JW, Heels-Ansdell D, Montori VM, Akl EA, Bryant DM, et al. Problems with use of composite end points in cardiovascular trials: systematic review of randomised controlled trials. BMJ. 2007;334:786. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Stokes J, Kannel WB, Wolf PA, Cupples LA, D’Agostino RB. The relative importance of selected risk factors for various manifestations of cardiovascular disease among men and women from 35 to 64 years old: 30 years of follow-up in the Framingham Study. Circulation. 1987;75(6 Pt 2):V65–73. [PubMed] [Google Scholar]
  • 19.Chowdhury MZI, Yeasmin F, Rabi DM, Ronksley PE, Turin TC. Predicting the risk of stroke among patients with type 2 diabetes: a systematic review and meta-analysis of C-statistics. BMJ Open. 2019;9:e025579. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Xiong Q, Chen S, Senoo K, Proietti M, Hong K, Lip GYH. The CHADS2 and CHA2DS2-VASc scores for predicting ischemic stroke among East Asian patients with atrial fibrillation: A systemic review and meta-analysis. Int J Cardiol. 2015;195:237–42. [DOI] [PubMed] [Google Scholar]
  • 21.de Jong Y, Ramspek CL, van der Endt VHW, Rookmaaker MB, Blankestijn PJ, Vernooij RWM, et al. A systematic review and external validation of stroke prediction models demonstrates poor performance in dialysis patients. J Clin Epidemiol. 2020;123:69–79. [DOI] [PubMed] [Google Scholar]
  • 22.Xu W, Huang J, Yu Q, Yu H, Pu Y, Shi Q. A systematic review of the status and methodological considerations for estimating risk of first ever stroke in the general population. Neurol Sci. 2021;42:2235–47. [DOI] [PubMed] [Google Scholar]
  • 23.Yuhan Deng S, Liu Z, Wang Y, Wang B, Liu. The Effect of Machine Learning Model for Predicting Stroke Risk in the General Population Based on Structured Data: A Systematic Review and Meta-Analysis. Chin J Stroke. 2022;17:1189–97. [Google Scholar]
  • 24.Nadarajah R, Younsi T, Romer E, Raveendra K, Nakao YM, Nakao K, et al. Prediction models for heart failure in the community: A systematic review and meta-analysis. Eur J Heart Fail. 2023;25:1724–38. [DOI] [PubMed] [Google Scholar]
  • 25.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Snell KIE, Levis B, Damen JAA, Dhiman P, Debray TPA, Hooft L, et al. Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: checklist for systematic reviews and meta-analyses (TRIPOD-SRMA). BMJ. 2023;381:e073538. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Moons KGM, de Groot JAH, Bouwmeester W, Vergouwe Y, Mallett S, Altman DG, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11:e1001744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Tufanaru C, Munn Z, Stephenson M, Aromataris E. Fixed or random effects meta-analysis? Common methodological issues in systematic reviews of effectiveness. Int J Evid Based Healthc. 2015;13:196–207. [DOI] [PubMed] [Google Scholar]
  • 29.Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143:29–36. [DOI] [PubMed] [Google Scholar]
  • 30.Debray TPA, Damen JAAG, Snell KIE, Ensor J, Hooft L, Reitsma JB, et al. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460. [DOI] [PubMed] [Google Scholar]
  • 31.Snell KI, Ensor J, Debray TP, Moons KG, Riley RD. Meta-analysis of prediction model performance across multiple studies: Which scale helps ensure between-study normality for the C-statistic and calibration measures? Stat Methods Med Res. 2018;27:3505–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Debray TP, Damen JA, Riley RD, Snell K, Reitsma JB, Hooft L, et al. A framework for meta-analysis of prediction model studies with binary and time-to-event outcomes. Stat Methods Med Res. 2019;28:2768–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327:557–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Debray T, De Jong V, metamisc. Meta-analysis of diagnosis and prognosis research studies. 2012;:0.4.0.
  • 35.Null RCTR, Team R, Null RCT, Core Writing T, Null R, Team R, et al. R: A language and environment for statistical computing. Computing. 2011;1:12–21. [Google Scholar]
  • 36.Conducting Meta-Analyses. in R with the metafor Package | Journal of Statistical Software. https://www.jstatsoft.org/article/view/v036i03. Accessed 18 Jan 2025.
  • 37.Moons KGM, Wolff RF, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: A Tool to Assess Risk of Bias and Applicability of Prediction Model Studies: Explanation and Elaboration. Ann Intern Med. 2019;170:W1–33. [DOI] [PubMed] [Google Scholar]
  • 38.Dufouil C, Beiser A, McLure LA, Wolf PA, Tzourio C, Howard VJ, et al. Revised Framingham Stroke Risk Profile to Reflect Temporal Trends. Circulation. 2017;135:1145–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hippisley-Cox J, Coupland C, Brindle P. Derivation and validation of QStroke score for predicting risk of ischaemic stroke in primary care and comparison with other risk scores: a prospective open cohort study. BMJ. 2013;346:f2573. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Xing X, Yang X, Liu F, Li J, Chen J, Liu X, et al. Predicting 10-Year and Lifetime Stroke Risk in Chinese Population. Stroke. 2019;50:2371–8. [DOI] [PubMed] [Google Scholar]
  • 41.Parmar P, Krishnamurthi R, Ikram MA, Hofman A, Mirza SS, Varakin Y, et al. The Stroke Riskometer(TM) App: validation of a data collection tool and stroke risk predictor. Int J Stroke. 2015;10:231–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Orfanoudaki A, Chesley E, Cadisch C, Stein B, Nouh A, Alberts MJ, et al. Machine learning provides evidence that stroke risk is not linear: The non-linear Framingham stroke risk score. PLoS ONE. 2020;15:e0232414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Hong C, Pencina MJ, Wojdyla DM, Hall JL, Judd SE, Cary M, et al. Predictive Accuracy of Stroke Risk Prediction Models Across Black and White Race, Sex, and Age Groups. JAMA. 2023;329:306–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Chun M, Clarke R, Cairns BJ, Clifton D, Bennett D, Chen Y, et al. Stroke risk prediction using machine learning: a prospective cohort study of 0.5 million Chinese adults. J Am Med Inf Assoc. 2021;28:1719–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Nijman S, Leeuwenberg AM, Beekers I, Verkouter I, Jacobs J, Bots ML, et al. Missing data is poorly handled and reported in prediction model studies using machine learning: a literature review. J Clin Epidemiol. 2022;142:218–29. [DOI] [PubMed] [Google Scholar]
  • 46.Chun M, Clarke R, Zhu T, Clifton D, Bennett DA, Chen Y, et al. Development, validation and comparison of multivariable risk scores for prediction of total stroke and stroke types in Chinese adults: a prospective study of 0.5 million adults. Stroke Vasc Neurol. 2022;7:328–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Jameson JL, Longo DL. Precision medicine–personalized, problematic, and promising. N Engl J Med. 2015;372:2229–34. [DOI] [PubMed] [Google Scholar]
  • 48.Morrison A, Polisena J, Husereau D, Moulton K, Clark M, Fiander M, et al. The effect of English-language restriction on systematic review-based meta-analyses: a systematic review of empirical studies. Int J Technol Assess Health Care. 2012;28:138–44. [DOI] [PubMed] [Google Scholar]
  • 49.Feng Y, Wang AY, Jun M, Pu L, Weisbord SD, Bellomo R, et al. Characterization of Risk Prediction Models for Acute Kidney Injury: A Systematic Review and Meta-analysis. JAMA Netw Open. 2023;6:e2313359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Smith LA, Oakden-Rayner L, Bird A, Zeng M, To M-S, Mukherjee S, et al. Machine learning and deep learning predictive models for long-term prognosis in patients with chronic obstructive pulmonary disease: a systematic review and meta-analysis. Lancet Digit Health. 2023;5:e872–81. [DOI] [PubMed] [Google Scholar]
  • 51.TRIPOD + AI statement. updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:q902. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Collins GS, Dhiman P, Ma J, Schlussel MM, Archer L, Van Calster B, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ. 2024;384:e074819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Riley RD, Archer L, Snell KIE, Ensor J, Dhiman P, Martin GP, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ. 2024;384:e074820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Riley RD, Snell KIE, Archer L, Ensor J, Debray TPA, van Calster B, et al. Evaluation of clinical prediction models (part 3): calculating the sample size required for an external validation study. BMJ. 2024;384:e074821. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (572.5KB, docx)

Data Availability Statement

All data generated or analysed during this study are included in this published article [and its supplementary information files].


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES