Abstract
Background:
Ischemic stroke accounts for 3.71 million deaths annually worldwide. Yet current risk prediction models demonstrate modest discrimination with reported C-statistics typically ranging from 0.63 to 0.75 and cannot reliably identify which individuals within high-risk populations will experience first ischemic stroke (FIS). We developed GRaph-based Analysis for Stroke Prediction (GRASP), a novel multimodal approach to predict FIS likelihood among individuals at risk who share similar vascular risk profiles.
Methods:
We analyzed n=1,226 UK Biobank participants having at least one established vascular risk factor associated with FIS, including 317 (25.9%) who developed FIS during follow-up (mean time to FIS onset: 9.5 years post-recruitment). Participants were stratified into at-risk with future ischemic stroke (AR-FIS) and at-risk with no ischemic stroke (AR-NIS). GRASP employed Graph Attention Network architecture incorporating 136 clinically relevant variables across demographic, clinical, lifestyle, and neuroimaging modalities. Performance was compared against baseline vascular risk factor models and conventional machine learning classifiers.
Results:
GRASP achieved superior performance in distinguishing AR-FIS from AR-NIS populations with AUC-ROC of 0.82 (95% CI: 0.78–0.87) and PR-AUC of 0.73 (95% CI: 0.67–0.80). This represented a 15.3% improvement in AUC-ROC and 27.9% improvement in PR-AUC over the best conventional multimodal classifier (XGBoost: AUC-ROC 0.69, PR-AUC 0.60) and substantial improvements over baseline vascular risk factor models (AUC-ROC 0.55, PR-AUC 0.28). Age-stratified analysis revealed optimal performance in younger participants (≤55 years: AUC-ROC 0.82) compared to older participants (>55 years: AUC-ROC 0.75).
Conclusions:
GRASP demonstrated substantial improvements in FIS risk prediction among at-risk populations through integration of multimodal data within a graph-based framework. GRASP also significantly outperformed traditional risk-factors based models in distinguishing which individuals with similar vascular risk profile will develop FIS, offering enhanced clinical risk stratification capabilities.
Keywords: First Ischemic Stroke, prognostic model, Multimodal, graph-neural networks
Introduction
Ischemic stroke is known to be a predominant subtype of stroke and represents a significant cause of disability and mortality globally {1,2}. While advances in stroke prediction and management of risk factors (including hypertension, diabetes, or atrial fibrillation) have led to meaningful reductions in stroke incidences, the lack of biomarkers that can reliably identify which individuals within high-risk populations will go on to experience a first ischemic stroke (FIS) remains a challenge. Timely identification of patients with high risk of FIS can help reduce disease severity, ensuring early intervention of modifiable factors preventing FIS and improved patient outcomes. A wide variety of risk prediction models, such as the Framingham Stroke Risk Profile (FSRP) {3}, QRISK {4}, CHADS2 score {5}, and CHA2DS2-VASc score {6}, have been broadly adopted as clinical tools for stroke risk estimation. These models predominantly rely on vascular risk factors (VRFs), including demographic and modifiable predictors such as age, gender, hypertension, diabetes, hyperlipidemia, smoking, and obesity {7,8}. While widely accepted, these tools are typically built upon additive or univariate frameworks, which may not comprehensively capture the nonlinear and multifactorial interactions that underline stroke pathophysiology {9,10}. Consequently, these models demonstrate modest discrimination, with reported C-statistics typically ranging from 0.63 to 0.75 across clinical populations {7,8,11}. For instance, FSRP and QRISK3 have shown sensitivity between 63–73% and specificity between 52–54% in prior studies {8}. Additionally, while these models are widely used for stroke risk estimation in the general population to identify individuals at elevated risk due to vascular risk factors, their ability to further discriminate which individuals within clinically high-risk groups will go on to experience a FIS remains modest.
These risk-prediction models also do not incorporate multimodal markers such as neuroimaging biomarkers, which can offer additional prognostic value. For example, white matter hyperintensities (WMH), silent infarcts, and microbleeds, have all been associated with elevated stroke risk even after adjusting for conventional risk scores {12–17}. As such, individuals with similar clinical profiles may exhibit vastly different cerebrovascular burdens which current models are not equipped to detect. These challenges underscore the need for multimodal approaches that could integrate richer, granular, and complementary sources of information to improve FIS risk predictions.
To overcome these limitations, machine learning (ML) and deep learning (DL) approaches have gained traction, offering the ability to model complex, nonlinear relationships among heterogeneous inputs{18}. For instance, graph-based models, particularly graph neural networks (GNNs), have emerged as powerful DL tools for integrating heterogeneous biomedical data and modeling complex yet complementary relationships between features {19,20}. GNNs represent patients as nodes, with associated clinical, imaging, or genetic features encoded as node attributes, and encode similarity-based relationships as edges allowing the model to learn both individual and population level interactions. GNNs have demonstrated superior performance over conventional ML methods in various disease prediction tasks, including those in cardiovascular and neurological domains {19–22}.
In this work, we present GRaph-based Analysis for Stroke Prediction (GRASP) – a new multimodal model to predict likelihood of individuals who are at risk of FIS from those with similar VRFs but no subsequent stroke. Since stroke is a multifactorial disease culminating from various factors including pathophysiology, lifestyle, and environmental factors, we hypothesized that by modeling complex nonlinear interactions across imaging, clinical, and environmental factors, GRASP will comprehensively capture risk-factors at different granularities and thus will significantly outperform traditional risk models in predicting FIS among the at-risk population. GRASP represents one of the first efforts to integrate phenotypic, genetic, imaging, and lifestyle data available from the UK Biobank cohort within a Graph Attention Network (GAT) architecture for individualized FIS risk prediction {23}.
Methods
Study Design and Population
The study utilized a retrospective cohort leveraging data from the UK biobank (UKB). The UKB is a large-scale biomedical database and research resource containing in-depth genetic, lifestyle, and health information from approximately 500,000 participants aged 40 – 69 at recruitment (de-tailed information is available at https://biobank.ndph.ox.ac.uk/showcase/). The UK Biobank (UKB) provides extensive clinical, phenotypic, and imaging data {24}. Detailed descriptions are provided in the Supplementary Materials (S1).
Figure 1 shows the inclusion and exclusion criteria for the data analyzed in this study. Inclusion criteria: The availability of (1) asymptomatic brain MRI, (2) the structural MRI-derived diffusivity and volumetric cortical and sub-cortical measures and (3) existing VRFs (age, gender, hypertension, diabetes, smoking) or an FIS event in the later timepoints (within a 15-year timeframe). Exclusion criteria: The subjects having prior history of stroke or transient ischemic attack (TIA), history of cancer, history of trauma or bleed, or history of other degenerative neurological conditions at the time of recruitment or UKB MRI brain protocol acquisition were excluded from the study. This ensured that the study population was representative of individuals who are considered at risk of FIS without confounding conditions.
Figure 1:

Overview of cohort construction, feature composition, and demographic-outcome distribution. a) Flowchart illustrating the inclusion criteria and final sample used for model development and analysis. b) Distribution of input features grouped by domain and subdomain, highlighting the heterogeneity across clinical, lifestyle, imaging, and demographic variables. c) Stroke outcome distribution stratified by sex, with corresponding average time-to-stroke (in years) among stroke cases shown above each group. This panel supports both class balance interpretation and demographic insight into time-to-event variation.
The final dataset comprised 1,226 participants, including 640 males (52.2%) and 586 females (47.8%) with a mean age of 55.5 ± 7.3 years, all presenting with at least one established vascular risk factor associated with FIS. Within this cohort, 317 participants (25.9%) developed FIS during subsequent follow-up assessments (mean time to FIS onset: 9.5 years post-recruitment). Participants were stratified into two primary groups: at-risk with future ischemic stroke (AR-FIS) and at-risk with no ischemic stroke (AR-NIS). Demographic characteristics and vascular risk factors are summarized in Figure 1.
Data Preparation and Feature Selection
Following cohort selection, 136 clinically relevant variables were chosen from the UKB’s feature catalog (>3,000 variables) based on established pathophysiological mechanisms of FIS and evidence from prior literature {3–7, 13–17}. This literature-guided approach balanced clinical relevance with computational feasibility, encompassing vascular risk factors, biomarkers, and neuroimaging parameters linked to cerebrovascular disease and FIS prediction. The selected variables were aligned into our feature matrix denoted as where is the number of participants and is the number of features. The corresponding binary outcome variable was denoted as , where indicates AR-FIS, and denotes AR-NIS.
The feature set encompassed two primary data modalities. The first consisted of structured clinical, demographic, and lifestyle information referred to as FEHR. This included demographic attributes (e.g. age, sex, ethnicity), lifestyle factors (e.g. smoking history, alcohol intake, Townsend Deprivation Index, physical activity measured as walking (≥ 4 days = 1; < 4 days = 0), sleep issues), and clinical indicators (e.g. hypertension, hyperlipidemia, diabetes status, body mass index (BMI), history of Myocardial Infarction, as well as Parental stroke history, please see Supplementary Material for comprehensive list). The second modality compromised IDPs denoted as FImg, obtained from multiple MRI modalities. These included regional measures of brain structure and microstructure, such as mean fractional anisotropy (FA) and mean diffusivity (MD) from diffusion tensor imaging (DTI) across 92 regions of interest, mean magnetic susceptibility values from 16 brain regions via susceptibility-weighted imaging (SWI), and total volumes of deep and periventricular white matter hyperintensities derived from T1 and T2-FLAIR sequences.
To ensure consistency in our cohort, only features with complete data across all participants were retained. This preprocessing step ensured that the final dataset was free from missing values, enabling more robust downstream modeling without the need for imputation or exclusion due to data sparsity. Furthermore, clinical features were encoded where necessary, such as binary encoding for categorical variables (e.g., history of hypertension: yes = 1, no = 0). A complete list of the 136 features incorporated in our classification models is provided in Supplementary Material excel.
Dimensionality reduction and feature selection were subsequently performed using Least Absolute Shrinkage and Selection Operator (LASSO) regression. 5-fold cross-validation LassoCV was used to determine the optimal regularization strength parameter . This yielded a sparse subset of features with non-zero coefficients, forming the reduced matrix , where . The reduced set consisted of 12 multimodal features, including 5 imaging-derived features from set and 7 clinical and parental history features from set. The complete retained feature list, along with corresponding LASSO coefficients indicating relative feature importance, is provided in Supplementary Table S2A. This approach reduced overfitting risk—particularly given the high dimensionality of imaging features—while improving model interpretability and computational efficiency. The refined matrix and labels were used as input for all downstream models.
Experimental Design
GRaph-based Analysis for Stroke Prediction (GRASP)
To represent complex relationships among clinical, lifestyle, and neuroimaging features in predicting FIS, GRASP uses a GAT architecture, enabling the model to learn both individual-level feature representations and topological relationships between at-risk individuals based on shared multimodal phenotypes. By leveraging the PyTorch Geometric library, the graph is built {25}, where each node represents a patient and each edge connects them. The node features are given by and the binary classification label vector indicates whether a participant later developed FIS (AR-FIS) or remained stroke-free (AR-NIS).
Graph Construction
To model inter-subject similarity, we constructed a population graph using a mutual k-nearest neighbors (k-NN) graph construction strategy.
Algorithm 1: Graph Construction via Mutual kNN.
| Input: Feature matrix X ∈ ℝN×F′number of neighbors k, training node set T, class labels y |
| Output: |
| Graph G = (V, E) |
| Standardize all features of X and compute Euclidean distances. |
| for each node i ∈ V do |
| Identify the k nearest neighbors Nk (i) |
| end for |
| for all node pairs i ≠ j do |
| if j ∈ Nk (i) and i ∈ Nk (j) then |
| Add undirected edge (i, j) to E. |
| end if |
| end for |
| for each i ∈ T belonging to the minority class do |
| Add up to k−4 additional edges to nearest neighbors from the same class. |
| end for |
| return G |
For each node , we retrieved its nearest neighbors identified in the feature space. We then examine each neighbor of to determine whether is also listed as a neighbor . If this bidirectional relationship holds, the pair () is added to the mutual neighbors list, and an undirected edge between nodes and was retained only if the neighbor was mutual (i.e., appeared in ’s k-NN list and vice versa). This mutual criterion reduces spurious connections and improves graph robustness in high-dimensional feature spaces. Evaluation of k values between 6 and 12 demonstrated that increasing k monotonically increased graph density (average node degree increasing from ~22 at to ~46 at ), while global edge homophily showed a modest decline (from 0.77 to 0.76), indicating increased neighborhood mixing at higher k values. Based on this tradeoff, was selected as a balance between sufficient connectivity (average degree ) and preservation of class-specific neighborhood structure.
To address class imbalance, particularly for the minority class (AR-FIS) an edge-boosting strategy was implemented only on training nodes, wherein up to additional connections were added between training samples of the AR-FIS group based on nearest-neighbor similarity {26}. This helped propagate minority class signals more effectively through the graph. The structural effects of edge boosting on graph connectivity and class homophily were quantitatively assessed. The edge boosting increased average node degree and class homophily within the AR-FIS group while preserving overall graph topology. In the training graph, the average node degree for AR-FIS increased from 23.50 to 35.01, and class-specific homophily increased from 0.54 to 0.78 following edge boosting. Global edge homophily showed a modest increase from 0.77 to 0.82, indicating improved minority-class connectivity without excessive distortion of the graph structure.
Model Architecture and Learning Framework
The GRASP model consisted of transformer-style attention layers that aggregated information from neighboring nodes using multi-head attention mechanisms. A two-layer GAT architecture was selected to balance model complexity with the dimensionality of the feature space (), preventing overfitting while enabling sufficient representation learning. Each attention layer was followed by a ReLU activation function, and a dropout layer () was incorporated to mitigate overfitting. The final output layer produced class probabilities for binary FIS risk stratification using a log-SoftMax function. Focal loss was employed specifically to address the class imbalance between AR-FIS (25.9%) and AR-NIS (74.1%) populations by dynamically weighting harder-to-classify examples, and the training was optimized using Adam Optimizer with weight decay regularization.
Data splitting was performed at the node level where nodes were divided into training and test sets, with approximately 60% allocated for training and 40% reserved for testing. During training, the labels of the test nodes were masked to prevent data leakage, ensuring that the model learned solely from the labeled training nodes. The model was trained using transductive learning, with predictions made on the full graph including unseen test nodes. Model training utilized early stopping with F1-score monitoring on the validation set, achieving convergence within 200 – 300 epochs. All models were trained on an NVIDIA RTX GPU with training time under one hour. Model performance was evaluated on the held-out test set using metrics such as accuracy, AUC-ROC, F1-score {27}, and PR-AUC {28}.
Building on GRASP’s multimodal graph-based framework, we conducted experiments to compare its performance against traditional vascular risk factor models and standard machine learning classifiers. These experiments also explored risk prediction across clinical subgroups, allowing us to assess the added value of integrating diverse data and graph modeling in ischemic stroke risk stratification. Table 1 summarizes the experiments and their corresponding feature sets.
Table 1:
Summary of Models, Feature Inputs, and Descriptions
| Model Name | Feature Set | Model Type |
|---|---|---|
| GRASP | F′ | Graph Attention Network (GAT) |
| Baseline VRF | FVRF | Gradient Boosted Decision Trees |
| Multimodal ML classifiers (XGB, RF, SVM) | F′ | Gradient Boosted Decision Trees, Ensemble Tree-Based Classifier, Kernel-Based Classifier |
| GRASPAG1 (AG1: Age group ≤ 55 years) | F1′ | Graph Attention Network (GAT) |
| GRASPAG2 (AG2: Age group > 55 years) | F2′ | Graph Attention Network (GAT) |
Comparative Strategies
Comparison of GRASP with the Baseline VRF model
As a baseline, we developed a classification model built using only established VRFs, aligned with the FSRP {3}. These features are a subset of the structured EHR-derived data, denoted . Specifically, includes demographic variables (e.g., age, sex, ethnicity), lifestyle factors (e.g., smoking status, alcohol consumption, physical activity), and clinical measures (e.g., hypertension, diabetes, hyperlipidemia, BMI, systolic and diastolic blood pressure, and cholesterol levels). Three ML classifiers – XGBoost (XGB) {29}, Random Forest (RF) {30}, and Support Vector Machine (SVM) {31} were trained on this feature subset, with the 60/40 train-test split. Hyperparameters were optimized using randomized search with 5-fold cross-validation. Performance was evaluated using Area under the receiver operating characteristic curve (AUC-ROC), and precision-recall curve (PR-AUC) metrics.
Comparison of GRASP with Multimodal ML classifiers
For this comparative model, we assessed the predictive value of integrating multimodal data. Using the full multimodal feature matrix the same three ML classifiers: XGB, RF, and SVM were evaluated to determine the best and most robust ML model. Given the imbalance between our curated AR-FIS and AR-NIS groups, Synthetic Minority Over-sampling Technique (SMOTE) was applied to the training data for all conventional ML models to address class imbalance. Each model was optimized using randomized search hyperparameter tuning with 5-fold cross-validation to select optimal hyperparameters. Final model performance was evaluated on the held-out test set. Performance evaluation was done by employing metrics such as AUC-ROC and PR-AUC to identify the best multimodal ML model.
Age stratified analysis within at-risk populations for prediction of FIS (GRASPAG1, GRASPAG2)
Motivated by known age-related differences in FIS risk and etiology, we performed age stratified subgroup analysis to investigate potential disparities. The cohort was stratified based on the mean age of the selected population (55.55 years), resulting in two subgroups: Age ≤ 55 years ( 572) with its corresponding matrix , where ; and Age years with its matrix , where . Detailed list of selected features is available in supplementary section S2B. These groups were independently evaluated using the top-performing classification model. Lasso-based feature selection was performed separately for each subgroup to identify the most informative predictors, which were then used to train and test models within each subgroup. For age-stratified analyses, evaluation of k values demonstrated that increasing k monotonically increased graph density in both subgroups, with average node degree rising from ~22 to ~38 (≤55: 22.18→38.28; >55: 22.23→38.26). In parallel, global edge homophily showed a modest decline with increasing k (≤55: 0.788→0.759; >55: 0.754→0.739), indicating increased neighborhood mixing at higher k values. Based on this tradeoff, k = 8 was selected to balance sufficient connectivity with preservation of outcome-specific neighborhood structure in each subgroup. Performance metrics, AUC-ROC, PR-AUC, accuracy and recall were compared to assess model efficacy across subgroups and to explore potential age-based disparities in FIS risk classification.
Statistical Analysis
Model performance was assessed using multiple evaluation metrics including area under the receiver operating characteristic curve (AUC-ROC), precision-recall area under the curve (PR-AUC), accuracy, precision, recall, and F1-score. Statistical robustness of performance estimates was evaluated through bootstrap resampling with 1,000 iterations on the test set to generate 95% confidence intervals for key metrics.
Statistical significance of performance differences between models was assessed using paired statistical tests on the held-out test set. AUC-ROC comparisons were performed using the DeLong test, which accounts for the correlation between models evaluated on the same subjects and provides asymptotically exact p-values {32}. PR-AUC comparisons were evaluated using paired bootstrap resampling with 1,000 iterations, where test subjects were resampled with replacement and PR-AUC was recomputed for each model on the same resampled subjects. Two-sided p-values for bootstrap comparisons were computed based on the proportion of iterations where the performance difference crossed zero.
Risk profile analysis was conducted through examination of feature co-occurrence patterns within model defined risk strata. Patients were stratified into high-risk and low-risk subgroups, and weighted feature co-occurrence networks were constructed to summarize relationships between features within each strata. Chord plot visualizations (Figures 5, S1 and S2) were used to illustrate patterns of jointly expressed features. These visualizations may provide clinically interpretable insights into multimodal risk factors’ clustering patterns rather than individual feature rankings. Full methodological details and subgroup analysis are provided in Supplementary Section S3.
Figure 5:

Chord plots illustrating co-occurrence patterns of clinical and imaging features in high-risk (left) and low-risk (right) populations. Each segment represents a feature, and chords between segments represents the joint activation of features within the cohort. Frequent intra- and cross-domain co-occurrence patterns are observed among vascular comorbidities and neuroimaging markers in the high-risk group, while low-risk individuals exhibit sparser, less clustered patterns.
Model interpretability and feature space visualization were assessed using t-distributed Stochastic Neighbor Embedding (t-SNE) projections of learned representations from each classification approach {33}. t-SNE analysis enabled comparison of class separation quality between the baseline VRF model, best multimodal ML classifier, and GRASP model by visualizing the clustering patterns of AR-FIS and AR-NIS populations in reduced dimensional space. All statistical analyses were performed using Python 3.8 with scipy.stats for hypothesis testing and scikit-learn for bootstrap sampling. Statistical significance was set at = 0.05 for all tests.
Results
GRASP Model Performance
The GRASP model demonstrated robust performance in distinguishing between AR-FIS and AR-NIS populations, achieving an AUC-ROC of 0.82 (95% CI: 0.78–0.87) and PR-AUC of 0.73 (95% CI: 0.67–0.80). For AR-FIS classification, GRASP achieved a recall of 0.68, precision of 0.60, and F1-score of 0.64. For AR-NIS classification, the model demonstrated superior performance with precision of 0.88, recall of 0.84 and F1-score of 0.86.
Comparative Strategies
Comparison of GRASP with the Baseline VRF model
On evaluating the performance of GRASP alongside benchmark vascular risk factors () model, GRASP (AUC-ROC 0.82; PR-AUC 0.73) demonstrated substantially superior performance in AR-FIS versus AR-NIS discrimination. In contrast, the XGB classifier trained on feature set yielded an AUC-ROC of 0.55, PR-AUC of 0.28, with similar performance as SVM (AUC-ROC: 0.55) and outperforming RF classifiers (AUC-ROC: 0.53) in distinguishing AR-FIS versus AR-NIS.
Comparison of GRASP with Multimodal ML classifiers
Among conventional machine learning approaches applied to the multimodal feature set , the multimodal ML-XGB classifier demonstrated superior performance among conventional classifiers based on balanced accuracy validation. The multimodal ML-XGB classifier achieved an AUC-ROC of 0.69 (95% CI: 0.647–0.780) and PR-AUC of 0.60 (95% CI: 0.466–0.662) on the test set, outperforming multimodal ML-RF (AUC-ROC: 0.72) and multimodal ML-SVM (AUC-ROC: 0.69) classifiers in AR-FIS versus AR-NIS classification. The GRASP model significantly outperformed the best-performing conventional classifier (multimodal ML-XGB), demonstrating a 15.3% improvement in AUC-ROC and 27.9% improvement in PR-AUC.
Age-Stratified Analysis Within At-Risk Populations for Prediction of FIS (GRASPAG1, GRASPAG2):
Performance analysis across age groups revealed differential model effectiveness. For the GRASPAG1, lasso feature selection resulted in , which included a combination of and features. This model achieved optimal performance with an AUC-ROC of 0.82, PR-AUC of 0.78, and overall accuracy of 75%. Within this younger subgroup, recall was 0.67 for AR-FIS and 0.78 for AR-NIS cases.
For the , lasso feature selection resulted in , which included a combination of and features. This model performance declined moderately, yielding an AUC-ROC of 0.75, PR-AUC of 0.65, and accuracy of 71%. Recall for AR-FIS was 0.66, while AR-NIS recall was 0.72.
Model Interpretability and Feature Space Analysis
t-SNE visualization of learned feature representations revealed distinct clustering patterns across modeling approaches (Figure 5). The baseline VRF and multimodal ML-XGB classifier exhibited partial overlap between AR-FIS and AR-NIS groups, indicating suboptimal class separation in the learned feature space. In contrast, the GRASP model demonstrated well-defined, distinct clusters between AR-FIS and AR-NIS populations, indicating superior discriminatory learning capabilities.
Risk Profile Co-occurrence Analysis
Chord plot analysis (GRASP) revealed distinct co-occurrence patterns between high-risk and low-risk populations [Figure 5]. In high-risk, risk profiles were clustered across both clinical and imaging modalities, with history of hypertension (H-HT) showing the most frequent co-occurrences. Neuroimaging markers such as mean MD and FA also formed densely interconnected clusters. In contrast, low-risk populations exhibited primarily intra-modal connections with weaker cross-modal links. Full feature-level and subgroup-specific co-occurrence patterns are provided in the Supplementary Materials (S3, S3A, S3B S3C and figures S1 and S2).
Statistical Robustness Assessment
Bootstrapping analysis (n=1000) confirmed the statistical robustness of performance metrics. Bootstrap confidence intervals for key metrics confirmed robust performance: GRASP AUC-ROC of 0.82 (95% CI: 0.78–0.87) and PR-AUC of 0.73 (95% CI: 0.67–0.80) versus multimodal ML-XGB classifier AUC-ROC of 0.69 (95% CI: 0.647–0.780) and PR-AUC of 0.60 (95% CI: 0.466–0.662). Paired statistical comparisons on the held-out test set demonstrated that GRASP significantly outperformed both multimodal XGBoost and VRF-based models. Compared with XGBoost, GRASP achieved a higher AUC-ROC (ΔAUC = 0.074, p = 0.0012) and PR-AUC (ΔPR-AUC = 0.112, p < 0.001). Improvements were even more pronounced relative to VRF-XGB (ΔAUC = 0.161, p < 0.001; ΔPR-AUC = 0.226, p < 0.001). Detailed results reported in Supplementary Table S4.
Discussion
Ischemic stroke remains a major cause of disability and mortality worldwide. Although advances in stroke prediction have contributed to declining incidence, reliable biomarkers to identify which high-risk individuals will experience a first ischemic stroke (FIS) are still lacking. Here, we present GRASP, a multimodal graph-based model designed to differentiate individuals at risk of FIS from those with comparable vascular risk factors but no subsequent stroke. We hypothesized that by modeling complex nonlinear interactions across imaging, clinical, and environmental factors, GRASP will capture the multilayered determinants of stroke risk and outperform traditional models in predicting FIS. Our results demonstrated that the GRASP model, a graph attention network approach integrating multimodal clinical and neuroimaging data, achieved superior performance in distinguishing individuals at risk of developing first ischemic stroke (AR-FIS) from those who remain stroke-free (AR-NIS) within at-risk populations. GRASP achieved an AUC-ROC of 0.82 (95% CI: 0.78–0.87) and PR-AUC of 0.73 (95% CI: 0.67–0.80), representing a ~18.8% improvement in AUC-ROC and ~21.7% improvement in PR-AUC compared to the best-performing conventional machine learning classifier (XGB: AUC-ROC 0.69).
While GRASP demonstrated clear performance gains over conventional machine-learning approaches, it is important to note that different strategies were used to address class imbalance across models. Traditional machine-learning baselines relied on sample-space augmentation through SMOTE, whereas GRASP mitigated imbalance through graph-based mechanisms, including focal loss and targeted edge enhancement among minority-class nodes during training {26}. These approaches may not be directly comparable. Notably, GRASP achieved consistent improvements across both global discrimination metrics (AUC-ROC) and minority-sensitive metrics (PR-AUC, recall), indicating that its performance gains extend beyond class rebalancing and reflect the advantages of graph-based representation learning in capturing structured inter-subject relationships. These results highlight the ability of graph-based neural networks in capturing complex feature relationships for FIS risk stratification. The integration of multimodal features proved critical for effective risk prediction, with GRASP demonstrating substantial improvements over unimodal models based solely on vascular risk factors (~49.1% improvement in AUC-ROC and ~160.7% improvement in PR-AUC compared to baseline VRF approaches). Risk co-occurrence analysis demonstrated that AR-FIS populations exhibit clustered risk profiles across clinical and neuroimaging domains rather than isolated individual risk factors, highlighting the value of graph-based modeling approaches that can capture complex inter-feature relationships.
Our findings are in line with previous studies demonstrating improvement in disease-risk prediction using ML-based models over traditional risk scores {27}, especially when leveraging high-dimensional datasets. For instance, Kakadiaris et al. {18} developed a cardiovascular ML-based risk calculator using SVM which achieved an overall AUC-ROC of 0.92 vs ACC/AHA (American College of Cardiology/American Heart Association) risk calculator AUC-ROC of 0.71 for Cardio-Vascular Diseases which was a combination of events including myocardial infarction, fatal CHD, stroke, and stroke death. A study by Jung et al. {34} leveraged a large-scale Korean National Health Insurance (KNHIS) dataset (N ≈ 500, 000) to develop a deep learning model for predicting ischemic stroke in atrial fibrillation (AF) patients. Their attention-based neural network, trained on 65 features, achieved an AUROC of 0.727—significantly outperforming the traditional CHA2DS2-VASc score (AUROC = 0.651) and other machine learning models. Similar to our findings, key predictors that were identified included hypertension and diabetes, in addition to heart failure and peripheral vascular disease. Similarly, Chun et al., {35} leveraging data from over 500,000 Chinese adults in the China Kadoorie Biobank, compared Cox regression with machine learning methods, including gradient boosted trees (GBT), for predicting stroke risk across short- and long-term intervals. GBT consistently outperformed other models, achieving AUROC scores of 0.833 for men and 0.836 for women. An ensemble model that combined GBT and Cox regression further improved accuracy and specificity. Notably, the most important features for stroke prediction included age, systolic blood pressure (SBP), diastolic blood pressure (DBP), hypertension treatment status, and physical activity levels, with minor variations between men and women. Our GRASP model extends these findings by demonstrating that graph-based neural networks, which explicitly model complex interdependencies within multimodal data, can achieve even higher performance levels (AUC-ROC 0.83) specifically for first ischemic stroke risk stratification.
Risk co-occurrence analysis through chord plot visualization (Figure 5) revealed clustering patterns that align with established stroke risk factors while providing new insights into their patterns of co-occurrence across features. The analysis demonstrated that AR-FIS populations exhibit interconnected risk profiles rather than isolated individual risk factors, supporting a multifactorial approach to stroke risk assessment.
The clinical risk factors observed in our co-occurrence analysis align with variables established in both atrial fibrillation and general population studies, including those incorporated into clinical scores like CHA2DS2-VASc and QRISK3 {4,11}. Clinical risk factors showed frequent co-occurrence patterns, with vascular comorbidities forming distinct clusters. History of hypertension, diabetes, and deep vein thrombosis demonstrated frequent co-occurrence, consistent with established cardiovascular risk factor clustering {7}. Additionally, familial predisposition (parental stroke history) showed notable connections with other clinical risk factors, reinforcing findings from previous studies highlighting the strong predictive value of family history in stroke risk stratification {7,13}.
Our inclusion of MRI features derived from routine clinical MR protocols—diffusion-weighted imaging (DWI) {36}, susceptibility-weighted imaging (SWI) {37}, and white matter hyperintensities (WMH) volumes and locations {38}—captured brain changes linked to stroke risk as demonstrated in prior neuroimaging studies. Neuroimaging markers revealed distinct co-occurrence patterns among diffusion tensor imaging parameters, including mean diffusivity and fractional anisotropy measures across multiple brain regions. The MRI features identified by our model represent neurovascular mechanisms that precede FIS. WMH burden reflects small vessel injury that leads to demyelination and gliosis whereby greater white matter burden indicates reduced cerebrovascular reserve and greater ischemic vulnerability. DTI metrics provides complementary evidence of compromised microstructural integrity, with increased MD and decreased FA indicating axonal and myelin injury that further increases the vulnerability to IS. SWI-derived susceptibility indicates iron deposition and microbleeds that further compromises vascular integrity and increases ischemic risk. The cluster of factors, i.e. WMH with accompanying higher MD, lower FA and increased susceptibility, leads to impaired vascular integrity thus lowering the threshold for the occurrence of a FIS. GRASP effectively captures this inter-modal clustering that could potentially represent a stroke risk signature that can be captured through routine clinical imaging. This clustering aligns with previous research demonstrating that white matter microstructural changes associated with stroke risk represent widespread alterations in neural connectivity rather than isolated regional abnormalities {39–41}. While GRASP demonstrates strong predictive performance using multimodal features from routine clinical data and neuroimaging, systematic implementation requires consideration of infrastructure requirements. The DTI-derived metrics (FA, MD) require more comprehensive diffusion protocols than standard clinical DWI sequences, and their systematic extraction requires computational infrastructure that may vary across healthcare settings. Several approaches could enhance accessibility across diverse clinical settings and clinical translation. These include simplified models using only the most discriminative features, tiered frameworks where advanced imaging is reserved for cases requiring refined risk assessment, or hybrid approaches integrating GRASP with established clinical risk scores to leverage complementary strengths.
Cross-modal co-occurrence patterns between clinical and neuroimaging features are consistent with prior studies, such as Sun et al., 15} who demonstrated improved prediction performance by integrating imaging biomarkers with clinical risk scores. Our results extend these findings by showing that systematic vascular pathology manifests as detectable brain microstructural changes prior to clinical stroke onset.
Stratifying by age revealed notable differences in both model performance and key predictive features. Age stratification revealed notable differences in model performance and feature composition. GRASP achieved higher AUC-ROC and precision in the ≤55 years subgroup (GRASPAG1), where LASSO retained 19 features (), indicating greater separability between AR-FIS and AR-NIS individuals. In contrast, the >55 years subgroup (GRASPAG2) showed more moderate performance, with a total of 9 retained features (). GRASPAG1 retained a diverse set of vascular risk factors (e.g., deep vein thrombosis, myocardial infarction, diabetes), lifestyle variables, and neuroimaging biomarkers capturing white matter microstructure (FA/MD metrics) and subcortical susceptibility. In contrast, GRASPAG2 retained common age-associated conditions (e.g., hypertension, atrial fibrillation) and demographic factors, while diffusion- and susceptibility-based imaging markers were no longer selected, suggesting reduced outcome-specific contrast in this population.
Graph-based clustering metrics further supported this pattern. AR-FIS class homophily was lower in older individuals compared to younger patients (0.47 vs. 0.58), indicating less cohesive outcome-specific clustering. Together, differences in retained feature composition and graph structure suggest that age-related clinical heterogeneity may have reduces discriminative contrast between outcome groups, rather than reflecting methodological limitations of the modeling framework. This interpretation is consistent with prior large-scale studies reporting age-dependent variation in stroke risk prediction performance, with reduced discrimination in older populations attributed to multimorbidity and overlapping risk profiles [42].
The risk co-occurrence analysis revealed distinct age-related patterns in feature clustering among high-risk individuals. In younger adults, co-occurrence patterns were centered on hypertension and myocardial infarction, accompanied by diffusion-based markers of white matter integrity in the posterior thalamic radiation. In contrast, older adults exhibited co-occurrence patterns characterized by familial stroke history, atrial fibrillation, and microstructural white matter alterations in the uncinate fasciculus. These differences are consistent with established age-related shifts in stroke risk profiles; wherein traditional vascular risk factors remain prevalent across age groups but cluster differently with advancing age {43}. Collectively, these observations highlight age-dependent heterogeneity in risk factor interactions and underscore the potential value of age-adapted modeling strategies for improved risk stratification {44}.
Our feature analysis further highlighted the complex interplay of traditional vascular risk factors (VRFs), multimorbidity, lifestyle factors, and advanced neuroimaging markers in FIS classification. The GRASP model consistently selected a broader range of both clinical and imaging-derived features as contributors to FIS risk. Classic VRFs such as treated hypertension remained prominent, alongside familial stroke history {7,13}. In addition, diffusion-based MRI markers, including fractional anisotropy (FA) and mean diffusivity (MD) within key white matter tracts, captured subtle microstructural alterations that contributed complementary information for FIS risk classification, reinforcing the growing relevance of neuroimaging markers in stroke risk assessment {39–41}. In another study conducted on the general population from the China Kadoorie Biobank (CKB, N ≈ 500, 000), it was shown that the Framingham Stroke Risk Profile (FSRP) achieved an AUC-ROC of 0.78 for total FIS prediction but exhibited poor calibration and underestimated absolute risks {35}. Our findings using GRASP suggest that integrating structured relationships between clinical, multimorbidity, and neuroimaging features may uncover latent risk factors, offering a promising avenue for improving FIS risk classification.
Our study has limitations. First, the analysis was conducted using a single large-scale cohort from the UK Biobank, and model performance was evaluated using a fixed train–test split with bootstrap resampling to assess stability. While this approach demonstrates internal consistency and robustness to sampling variations within our cohort, it does not substitute for external validation. The generalizability of the proposed model beyond the UK Biobank cohort remains to be established. Second, the model relies on structured clinical and imaging data obtained at baseline without accounting for temporal variations in risk factors, which may influence long-term FIS risk assessment. While baseline-only models are common in stroke risk prediction, incorporating longitudinal information may further improve performance, particularly in older or clinically complex populations. Third, the practical implementation of GRASP in routine clinical settings faces several challenges. The neuroimaging features include DTI-derived metrics (FA, MD) that require more comprehensive diffusion protocols than standard clinical DWI sequences. While advanced diffusion imaging is increasingly available in clinical and research settings, acquisition protocols and the computational infrastructure for systematic feature extraction may vary across healthcare institutions, potentially limiting immediate applicability in some clinical environments.
Future work will prioritize (a) validation of the model in independent, multi-center cohorts from diverse geographic regions and healthcare institutions, (b) exploration of temporal validation using prospectively collected data, and (c) integration with established clinical risk scores and extension of the framework to incorporate longitudinal follow-up data and additional modalities. Specifically, future extensions of GRASP will explore: (1) the integration of temporal risk trajectories including longitudinal blood pressure measurements, antithrombotic medication exposure histories, and incident comorbidities; (2) emerging biomarkers such as retinal imaging markers of microvascular disease and advanced neuroimaging metrics (e.g., cerebral perfusion imaging); (3) development of simplified or tiered models to enhance accessibility across diverse clinical settings; (4) hybrid approaches combining GRASP with traditional clinical risk scores; and (5) prospective implementation studies to assess clinical decision support utility, cost-effectiveness, and impact on patient outcomes in real-world healthcare environments.
Supplementary Material
Figure 2:

Methodology Workflow Overview. (A) Overview of the GRASP framework, a graph-based architecture that leverages multimodal features for stroke risk prediction. (B) Data curation and lasso feature selection process to build the final dataset. (C(a)) Comparative evaluation of model performance across baseline (VRF-only), machine learning, and GRASP models, including subgroup analysis stratified by age. (C(b)) Statistical interpretation pipeline comprising dimensionality reduction (t-SNE), bootstrapped performance estimation, and feature risk co-occurrence analysis for high- and low-risk individuals.
Figure 3:

a-b) Receiver Operating Characteristic (ROC) and Precision-Recall curves illustrating the performance of different models distinguishing between AR-FIS and AR-NIS cases. c-e) Performance comparison of baseline VRF, multimodal ML-XGB classifier, GRASP, and . The bar plots depict AUC-ROC, PR-AUC, precision and recall. F) Bootstrapped confidence intervals for multimodal ML-XGB classifier and GRASP models.
Figure 4:

t-SNE plots depicting the distribution of AR-FIS and AR-NIS cases across different models. The visualizations highlight the clustering behavior of data points based on features learned by the XGB-based multimodal model and the GRASP model, with distinct separations observed between the two populations, demonstrating the models’ effectiveness in distinguishing AR-FIS from AR-NIS cases.
Highlights:
GRASP: multimodal graph attention network predicting first ischemic stroke risk.
GRASP achieved AUC-ROC 0.82, outperforming XGBoost (0.69) and VRF model (0.55).
Multimodal data improved stroke risk stratification in at-risk populations.
Stronger discrimination was observed in younger participants (≤55 years, AUC-ROC 0.82).
Acknowledgments
This research has been conducted using the UK Biobank Resource under Application Number 101084.
Sources of Funding
We would like to acknowledge the following funding sources: R01NS123378, R01NS105646, P50HD105353, 1U01CA248226-01, 1R01CA264017-01A1, 3U01CA248226-03S1, Department of Defense/PRCRP Career Development Award W81XWH-18-1-0404, The Dana Foundation David Mahoney Neuroimaging Program, The V Foundation Translational Research Award, The Johnson & Johnson WiSTEM2D Award, and Pilot Funding, Department of Radiology, University of Wisconsin-Madison.
Footnotes
Competing Interests
Dr. Pallavi Tiwari reports being a Founder and holding equity in LivAi Inc., and serving as a Scientific Consultant for Johnson & Johnson Inc.
Juhi Desai & Drs. Vivek, Veena, and Nagesh report no disclosures.
Declaration of interests
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:
Pallavi Tiwari reports financial support was provided by National Institutes of Health. Pallavi Tiwari reports financial support was provided by National Institutes of Health National Cancer Institute. Pallavi Tiwari reports financial support was provided by Department of Defense - PRCRP Career Development Award. Pallavi Tiwari reports financial support was provided by The Dana Foundation David Mahoney Neuroimaging Program. Pallavi Tiwari reports financial support was provided by The V Foundation Translational Research Award. Pallavi Tiwari reports financial support was provided by The Johnson & Johnson WiSTEM2D Award. Vivek Prabhakaran reports was provided by National Institutes of Health. Vivek Prabhakaran reports financial support was provided by National Institute of Child Health and Human Development. Veena Nair reports financial support was provided by National Institutes of Health. Veena Nair reports financial support was provided by National Institute of Child Health and Human Development. Nagesh Adluru reports financial support was provided by National Institutes of Health. Nagesh Adluru reports financial support was provided by National Institute of Child Health and Human Development. Marwa Ismail reports financial support was provided by University of Wisconsin-Madison. Pallavi Tiwari reports a relationship with LivAi Inc that includes: equity or stocks. Pallavi Tiwari reports a relationship with Johnson & Johnson that includes: consulting or advisory. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Ethics Statement
This research was conducted using data from the UK Biobank Resource under Application Number 101084. UK Biobank has obtained Research Tissue Bank approval from the North West Multi-centre Research Ethics Committee (REC reference: 11/NW/0382). This approval allows researchers to use the UK Biobank data without requiring separate project-specific ethical approval, provided the research is conducted under the terms of the approved access.
All UK Biobank participants provided informed consent for their data to be used for health-related research. The current study complied with the UK Biobank Ethics and Governance Framework, the Access Procedures, and all applicable institutional and national guidelines and regulations. All data were de-identified prior to access by the research team, and no attempt was made to reidentify individuals.
No direct contact with participants occurred, and no additional ethics approval was required for this secondary analysis of anonymized data.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Data Availability
The data that support the findings of this study are available through application to the UK Biobank (www.ukbiobank.ac.uk), subject to their institutional data access procedures. Derived data supporting the results are available from the corresponding author upon reasonable request, in compliance with UK Biobank data-sharing policies.
Code Availability
The code used for data preprocessing, model implementation, and analysis in this study is available from the corresponding author upon reasonable request. A version of the model framework will be made available on a public repository following peer-reviewed publication.
References:
- 1.Association, A. H. (2024, Jan). 2024 Heart Disease and Stroke Statistics Update Fact Sheet. American Heart Association. Retrieved from https://www.heart.org/-/media/PHD-Files-2/Science-News/2/2024-Heart-and-Stroke-Stat-Update/2024-Statistics-At-A-Glance-final_2024.pdf?sc_lang= [Google Scholar]
- 2.Wu Simia et al. Global burden of stroke: dynamic estimates to inform action. The Lancet Neurology, 23:952–953, 2022. [DOI] [PubMed] [Google Scholar]
- 3.Bos Daniel, Ikram M. Arfan, Leening Maarten J. G., and Ikram M. Kamran. The revised framingham stroke risk profile in a primary prevention population. Circulation, 135(22):2207–2209, 2017. [DOI] [PubMed] [Google Scholar]
- 4.Hippisley-Cox Julia, Coupland Carol, and Brindle Peter. Development and validation of Qrisk3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ, 357, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Gage Brian F., Waterman Amy D., Shannon William, Boechler Michael, Rich Michael W., and Radford Martha J.. Validation of clinical classification schemes for predicting stroke results from the national registry of atrial fibrillation. JAMA, 285(22):2864–2870, 06 2001. [DOI] [PubMed] [Google Scholar]
- 6.Chen Lin Y., Norby Faye L., Chamberlain Alanna M., MacLehose Richard F., Bengtson Lindsay G.S., Lutsey Pamela L., and Alonso Alvaro. Cha2d2VASc score and stroke prediction in atrial fibrillation in whites, blacks, and hispanics. Stroke, 50(1):28–33, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Boehme Amelia K., Esenwa Charles, and Elkind Mitchell S.V.. Stroke risk factors, genetics, and prevention — circulation research. Circulation Research, 120(3), 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Goldstein Larry B., Adams Robert, Alberts Mark J., Appel Lawrence J., Brass Lawrence M., Bushnell Cheryl D., Culebras Antonio, DeGraba Thomas J., Gorelick Philip B., Guyton John R., Hart Robert G., Howard George, Kelly-Hayes Margaret, Nixon (Ian) J.V., and Sacco Ralph L.. Primary prevention of ischemic stroke. Stroke, 37(6):1583–1633, 2006. [DOI] [PubMed] [Google Scholar]
- 9.Orfanoudaki Agni, Chesley Emma, Cadisch Christian, Stein Barry, Nouh Amre, Alberts Mark J., and Bertsimas Dimitris. Machine learning provides evidence that stroke risk is not linear: The non-linear Framingham stroke risk score. PLOS ONE, 15(5):1–20, 05 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wei Xu, Huang Jiuyi, Qingsong Yu, Hongfan Yu, Yang Pu, and Qiuling Shi. A systematic review of the status and methodological considerations for estimating risk of first ever stroke in the general population. Neurological Sciences, 2021. [DOI] [PubMed] [Google Scholar]
- 11.Hong Chuan, Pencina Michael J., Wojdyla Daniel M., Hall Jennifer L., Judd Suzanne E., Cary Michael, Engelhard Matthew M., Berchuck Samuel, Xian Ying, Sr D’Agostino Ralph, Howard George, Kissela Brett, and Henao Ricardo. Predictive accuracy of stroke risk prediction models across black and white race, sex, and age groups. JAMA, 329(4):306–317, 01 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Vermeer Sarah E., den Heijer Tom, Koudstaal Peter J., Oudkerk Matthijs, Hofman Albert, and Breteler Monique M.B.. Incidence and risk factors of silent brain infarcts in the population-based Rotterdam scan study. Stroke, 34(2):392–396, 2003. [DOI] [PubMed] [Google Scholar]
- 13.Debette Stéphanie, Schilling Sabrina, Duperron Marie-Gabrielle, Larsson Susanna C., and Markus Hugh S.. Clinical significance of magnetic resonance imaging markers of vascular brain injury: A systematic review and meta-analysis. JAMA Neurology, 76(1):81–94, 01 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Arsava EM. The role of MRI as a prognostic tool in ischemic stroke. J Neurochem, 123, 2012. https://pubmed.ncbi.nlm.nih.gov/23050639/. [DOI] [PubMed] [Google Scholar]
- 15.Sun J, Sui Y, Chen Y, Lian J, and Wang W. Predicting acute ischemic stroke using the revised framingham stroke risk profile and multimodal magnetic resonance imaging. Front Neurol, 14, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Iwasa Kenichi, Onoda Keiichi, Takamura Masahiro, Takayoshi Hiroyuki, Mitaki Shingo, Yamaguchi Shuhei, and Nagai Atsushi. Development of a stroke risk score with mri asymptomatic brain lesions attributes to evaluate prognostic vascular events. Journal of the Neurological Sciences, 448:120642, 2023. [DOI] [PubMed] [Google Scholar]
- 17.Gallacher Katie I., Jani Bhautesh D., Hanlon Peter, Nicholl Barbara I., and Mair Frances S.. Multimorbidity in stroke. Stroke, 50(7):1919–1926, 2019. [DOI] [PubMed] [Google Scholar]
- 18.Kakadiaris Ioannis A., Vrigkas Michalis, Yen Albert A., Kuznetsova Tatiana, Budoff Matthew, and Naghavi Morteza. Machine learning outperforms acc/aha cvd risk calculator in mesa. Journal of the American Heart Association, 7(22):e009476, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Sarah Parisot, Ktena Sofia Ira Ferrante Enzo, et al. Spectral graph convolutions for population-based disease prediction. In Medical Image Computing and Computer Assisted Intervention MICCAI 2017, pages 177–185. Springer International Publishing, 2017. [Google Scholar]
- 20.Zhou J, Cui G, Zhang Z, Yang C, Liu Z, Wang L, … & Sun, M. (2020). Graph neural networks: A review of methods and applications. AI Open, 1, 57–81. 10.1016/j.aiopen.2021.01.001 [DOI] [Google Scholar]
- 21.Gupta Akshat, Gupta Anubha, Shetty Manu, Goyal Dixit, Girish MP, and Gupta Mohit. CardioRiskNet: Attention-based CVAE-enabled GCN for risk prediction in STEMI. In ICASSP 2025 – 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2025. [Google Scholar]
- 22.Tong Tong, Gray Katherine, Gao Qinquan, Chen Liang, and Rueckert Daniel. Multi-modal classification of alzheimer’s disease using nonlinear graph fusion. Pattern Recognition, 63:171–181, 2017. [Google Scholar]
- 23.Sudlow C, Gallacher J, Allen N, Beral V, Burton P, Danesh J, Downey P, Elliott P, Green J, Landray M, Liu B, Matthews P, Ong G, Pell J, Silman A, Young A, Sprosen T, Peakman T, Collins R. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015. Mar 31;12(3):e1001779. doi: 10.1371/journal.pmed.1001779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Littlejohns J, Holliday J, and Gibson LM et al. The UK Biobank imaging enhancement of 100,000 participants: rationale, data collection, management and future directions. Nature Communications, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Fey Matthias and Lenssen Jan E.. Fast graph representation learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019. [Google Scholar]
- 26.Zhao T, Zhang X, & Wang S (2021). GraphSMOTE: Imbalanced Node Classification on Graphs with Graph Neural Networks. In WSDM 2021 - Proceedings of the 14th ACM International Conference on Web Search and Data Mining (pp. 833–841). (WSDM 2021 - Proceedings of the 14th ACM International Conference on Web Search and Data Mining). Association for Computing Machinery, Inc. 10.1145/3437963.3441720 [DOI] [Google Scholar]
- 27.Chinchor Nancy and Sundheim Beth M. Muc-5 evaluation metrics. In Fifth Message Understanding Conference (MUC-5): Proceedings of a Conference Held in Baltimore, Maryland, August 25–27, 1993, 1993. [Google Scholar]
- 28.Buckland Michael and Gey Fredric. The relationship between recall and precision. Journal of the American society for information science, 45(1):12–19, 1994. [Google Scholar]
- 29.Chen Tianqi and Guestrin Carlos. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, page 785–794. ACM, August 2016. [Google Scholar]
- 30.Ho Tin Kam. Random decision forests. In Proceedings of 3rd international conference on document analysis and recognition, volume 1, pages 278–282. IEEE, 1995 [Google Scholar]
- 31.Vapnik Vladimir. The nature of statistical learning theory. Springer science & business media, 1999. [Google Scholar]
- 32.DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics. 1988. Sep;44(3):837–45. [PubMed] [Google Scholar]
- 33.van der Maaten Laurens and Hinton Geoffrey.. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008. [Google Scholar]
- 34.Jung Seonwoo, Song Min-Keun, Lee Eunjoo, Bae Sejin, Kim Yeon-Yong, Lee Doheon, Myoung Jin Lee, and Sunyong Yoo. Predicting ischemic stroke in patients with atrial fibrillation using machine learning. FBL, 27(3), 2022. [DOI] [PubMed] [Google Scholar]
- 35.Matthew Chun, Robert Clarke, Cairns Benjamin J Clifton David, Derrick Bennett, Yiping Chen, Yu Guo, Pei Pei, Jun Lv, Canqing Yu, Ling Yang, Liming Li, Zhengming Chen, Tingting Zhu, and the China Kadoorie Biobank Collaborative Group. Stroke risk prediction using machine learning: a prospective cohort study of 0.5 million chinese adults. Journal of the American Medical Informatics Association, 28(8):1719–1727, 05 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Yang Siyu, Zhou Yihao, Wang Feng, He Xuesong, Cui Xuan, Cai Shaojie, Zhu Xingyan, and Wang Dongyan. Diffusion tensor imaging in cerebral small vessel disease applications: opportunities and challenges. Frontiers in Neuroscience, 18, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Werring DJ, Coward LJ, Losseff NA, Jäger HR, and Brown MM. Cerebral microbleeds are common in ischemic stroke but rare in tia. Neurology, 65(12):1914–1918, 2005. [DOI] [PubMed] [Google Scholar]
- 38.Ghaznawi Rashid, Geerlings Mirjam I., Jaarsma-Coes Myriam, Hendrikse Jeroen, de Bresser Jeroen NathoeMD HM Bots ML Emmelot MH de Borst GJ Kappelle LJ Leiner T de Jong PA Lely AT van der Kaaij NP Ruigrok Y Verhaar MC Westerink J on behalf of the UCC-Smart Study Group Visseren FLJ Asselbergs FW. Association of white matter hyperintensity markers on mri and long-term risk of mortality and ischemic stroke. Neurology, 96(17):e2172–e2183, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Maillard Pauline, Mitchell Gary F., Himali Jayandra J., Beiser Alexa, Fletcher Evan, Tsao Connie W., Pase Matthew P., Satizabal Claudia L., Vasan Ramachandran S., Seshadri Sudha, and DeCarli Charles. Aortic stiffness increased white matter free water and altered microstructural integrity. Stroke, 48(6):1567–1573, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Guey Stéphanie, Lesnik Oberstein Saskia A.J., Tournier-Lasserve Elisabeth, and Chabriat Hugues. Hereditary cerebral small vessel diseases and stroke: A guide for diagnosis and management. Stroke, 52(9):3025–3032, 2021. [DOI] [PubMed] [Google Scholar]
- 41.Wilson S, Harkness Duncan, and Kirsty et al. Cerebral microbleeds and stroke risk after ischaemic stroke or transient ischaemic attack: a pooled analysis of individual patient data from cohort studies. The Lancet Neurology, 18(9):653–665, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Hong C, Pencina MJ, Wojdyla DM, et al. Predictive Accuracy of Stroke Risk Prediction Models Across Black and White Race, Sex, and Age Groups. JAMA. 2023;329(4):306–317. doi: 10.1001/jama.2022.24683 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Yahya Tamer, Jilani Mohammad Hashim, Khan Safi U., Mszar Reed, Hassan Syed Zawahir, Blaha Michael J., Blankstein Ron, Virani Salim S., Johansen Michelle C., Vahidy Farhaan, Cainzos-Achirica Miguel, Nasir Khurram. Stroke in young adults: Current trends, opportunities for prevention and pathways forward, American Journal of Preventive Cardiology, Volume 3, 2020, ISSN 2666–6677 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.George MG, Tong X, Bowman BA. Prevalence of Cardiovascular Risk Factors and Strokes in Younger Adults. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data that support the findings of this study are available through application to the UK Biobank (www.ukbiobank.ac.uk), subject to their institutional data access procedures. Derived data supporting the results are available from the corresponding author upon reasonable request, in compliance with UK Biobank data-sharing policies.
The code used for data preprocessing, model implementation, and analysis in this study is available from the corresponding author upon reasonable request. A version of the model framework will be made available on a public repository following peer-reviewed publication.
