Skip to main content
Frontiers in Bioinformatics logoLink to Frontiers in Bioinformatics
. 2026 Sep 10;6:1875388. doi: 10.3389/fbinf.2026.1875388

Evaluation of functional annotation-informed and ancestry-specific polygenic risk scores for ischemic stroke

Nicole D Armstrong 1,*, Vinodh Srinivasasainagendra 2, Amit Patki 2, Lavanya Pilla 2, Chad Aldridge 3, Alana C Jones 1, Ulrich Broeckel 4, Ethan M Lange 5, Leslie A Lange 5, Pankaj Arora 6, Nita A Limdi 7, Hirotaka Iwaki 8,9, Lana Sargent 8,10, Joohyun Kim 11, Maggie C Y Ng 11, Josep M Mercader 12,13,14,15,16, Keith L Keene 17,18, Bradford B Worrall 3,19, Hemant K Tiwari 2, Marguerite R Irvin 1
PMCID: PMC13601841  PMID: 42787420

Abstract

Introduction

Ischemic stroke (IS) is a leading cause of morbidity and mortality, and predicting events remains challenging. Polygenic risk scores (PRSs) aggregate genetic variants, but performance varies across ancestries and remains modest, particularly in non-European populations.

Methods

We compared Bayesian PRS methods incorporating functional annotations and expanded linkage disequilibrium reference panels for IS prediction. PRSs were generated using PRS-CS with HapMap3 and TagIt panels, and SBayesRC integrating functional priors. Scores were optimized in 1,403 European ancestry (EA) and 6,342 African ancestry (AA) participants from the Reasons for Geographic and Racial Differences in Stroke (REGARDS) study and independently validated in All of Us participants (78,155 EA and 29,594 AA participants), a REGARDS holdout cohort (11,459 EA participants), and the Genetics of Hypertension Associated Treatments (GenHAT) study (6,908 AA participants). Predictive performance was evaluated using logistic regression adjusted for age, sex, the top 10 genetic principal components, and antihypertensive treatment assignment where applicable. Performance was assessed using odds ratios per standard deviation (OR/SD), liability-scale R2, and area under the curve (AUC).

Results

Among EA participants, SBayesRC demonstrated the largest effect size and liability R2 in the REGARDS holdout validation cohort (OR/SD = 1.24, 95% CI 1.12–1.36, R2 = 2.42%). PRS-CS using the expanded TagIt panel (OR/SD = 1.11, 95% CI 1.00–1.22) showed numerically stronger performance than the HapMap3 panel (OR/SD = 1.09, 95% CI 0.98–1.20). Performance was consistent across the All of Us EA cohort, with OR/SD ranging from 1.10 to 1.16 across methods. Among AA participants, associations and predictive metrics were attenuated, with no significant associations observed in All of Us. A modest association was observed for PRS-CS-TagIt in GenHAT (for an all-stroke outcome). PRSs modestly improved discrimination beyond covariate models among EA participants (AUCs up to 66.6%), but showed little to no improvement among AA participants (AUCs 62.7%–66.7%). High PRS percentile groups (e.g., top 2% versus 98%) showed modest stroke enrichment with high specificity and limited sensitivity.

Conclusion

Functional annotation-informed PRSs and expanded reference panels modestly improved prediction, particularly among EA participants. Limited performance among AA participants underscores the need for larger, more diverse genomic datasets and integrative models incorporating genetic, clinical, lifestyle, and social determinants of health.

Keywords: functional annotation, genetic epidemiology, ischemic stroke, polygenic risk score, precision medicine

Introduction

Stroke is a leading cause of death and disability globally, resulting in more than 7 million deaths per year (Palaniappan et al., 2026). Beyond mortality, stroke imposes a substantial public health burden (Taylor et al., 1996) through prolonged hospitalization, long-term rehabilitation, loss of independence, and decreased quality of life for survivors and their families (Godwin et al., 2013; Alhusayni and Alzahrani, 2025; Strilciuc et al., 2021; Donkor, 2018). Despite advances in prevention and acute treatment, predicting who will experience a stroke remains challenging. Improved risk prediction tools that can identify high-risk individuals earlier could enable targeted prevention strategies, optimize clinical management, and ultimately reduce the individual and societal burden of stroke across diverse populations.

While well-established behavioral and clinical risk factors, such as body mass index, hypertension, type 2 diabetes, and cigarette smoking, play a major role in stroke risk (Chen et al., 2014), there is compelling evidence for a heritable component. Twin and family studies indicate that ischemic stroke has heritability estimates ranging from 16% to 40%, depending on the subtype and population studied (Bevan et al., 2012; Traylor et al., 2017) Recent genome-wide association studies (GWAS), including large consortia efforts such as MEGASTROKE (Malik et al., 2018), GIGASTROKE (Mishra et al., 2022), and COMPASS (Keene et al., 2020), have identified dozens of loci associated with ischemic stroke (IS) and its subsequent subtypes. However, the majority of these variants individually confer modest effect sizes, and their translational potential for clinical risk prediction remains limited.

Polygenic risk scores (PRSs) offer a promising framework to translate GWAS findings into individualized risk prediction by aggregating the effects of thousands to millions of common variants across the genome (Lewis and Vassos, 2017). Several statistical methods have been developed to improve PRS performance, with Bayesian approaches, such as PRS-CS, demonstrating improved accuracy by incorporating genome-wide variants and applying continuous shrinkage priors to variant effect sizes (Ge et al., 2019). However, PRS performance varies significantly across ancestral populations, in part due to differences in linkage disequilibrium (LD) patterns, allele frequencies, and underrepresentation of non-European ancestry populations in GWAS summary statistics (Martin et al., 2017). In stroke specifically, PRS research remains comparatively sparse, with approximately 30 scores overall in the Polygenic Score (PGS) catalog (Lambert et al., 2021; Lambert et al., 2024), in contrast to over 70 scores developed for coronary artery disease. Furthermore, recent stroke PRS studies, such as those by Abraham et al. (Abraham et al., 2019) and Mishra et al. (Mishra et al., 2022) have shown that current scores, as well as composite meta-scores, offer only modest incremental value beyond clinical risk factors, particularly in non-European populations.

To address this gap, we evaluated whether incorporating functional annotations and expanded ancestry-specific LD reference panels improves IS PRS performance. We hypothesized that leveraging functional priors would enhance the prioritization of putatively causal variants and improve PRS performance, particularly in non-European populations. We generated and compared two Bayesian PRS approaches: (i) PRS-CS (Ge et al., 2019) using the standard HapMap3 (HM3) LD reference panels and with an expanded, multi-ancestry LD reference panel (TagIt) developed by the D-PRIME consortium (Huerta-Chagoya et al., 2026), and (ii) SBayesRC (Zheng et al., 2024), which integrates functional annotations. These methods were assessed using individual-level data in 1,403 self-reported European ancestry (EA) and 6,342 self-reported African ancestry (AA) participants from the Reasons for Geographic and Racial Differences in Stroke (REGARDS) study. Validation was performed in independent cohorts, including 78,155 (898 ischemic stroke cases) EA participants from the All of Us Program, 11,459 (444 IS cases) holdout EA REGARDS participants, 29,594 (311 IS cases) AA All of Us participants, and 6,908 (366 any stroke) AA participants from the Genetics of Hypertension Associated Treatments (GenHAT) study. For contextual benchmarking, we also evaluated two previously published stroke PRSs (Abraham et al. (Abraham et al., 2019) and Mishra et al. (Mishra et al., 2022)).

Methods

Study populations

The Reasons for Geographic and Racial Differences in Stroke (REGARDS) study

REGARDS is a national, population-based cohort established to investigate incident stroke and related risk factors. Between 2003 and 2007, 30,239 self-reported AA and EA adults aged ≥45 years were enrolled from all 48 contiguous U.S. states and the District of Columbia (Howard et al., 2005). At baseline, participants completed a computer-assisted telephone interview followed by an in-home visit that included blood pressure measurement, phlebotomy, and anthropometrics. REGARDS is ongoing, with participants contacted every 6 months to ascertain incident stroke and secondary cardiovascular outcomes.

Suspected events identified through surveillance were adjudicated through medical record review by trained physicians using the World Health Organization definition: focal neurological deficits lasting >24 h or imaging evidence consistent with acute stroke (Howard et al., 2005; Howard et al., 2011). All ischemic stroke events occurring on or before 30 September 2019, were eligible for inclusion. Participants with a history of stroke at baseline were excluded.

The All of Us Research Program

All of Us is a prospective cohort designed to advance precision medicine by recruiting a diverse population across the United States. Enrollment of adults aged ≥18 years began in May 2018 across more than 340 sites. Participants contribute electronic health record (EHR) data, complete baseline demographic and lifestyle surveys, and undergo physical measurements and biospecimen collection (All of Us Research Program Investigators et al., 2019; Ramirez et al., 2022). EHR data include diagnoses, procedures, medications, and laboratory results.

For the present analysis, we used the Controlled Tier version 7 data release (April 2023). The genomic dataset available in version 7 included 245,388 individuals with short-read whole-genome sequencing data (All of Us Research Program Genomics and I, 2024). To harmonize the age distribution with REGARDS, analyses were restricted to participants aged ≥45 years. The reported results comply with the All of Us Data and Statistics Dissemination Policy, which disallows disclosure of group counts <20.

The Genetics of Hypertension Associated Treatments (GenHAT) study

GenHAT was an ancillary pharmacogenomics study of the Antihypertensive and Lipid-Lowering Treatment to Prevent Heart Attack Trial (ALLHAT). ALLHAT was a double-blind, randomized clinical trial conducted from 1994 to 2002 that enrolled >42,000 adults aged ≥55 years with hypertension and at least one additional cardiovascular risk factor. Participants were randomized to chlorthalidone, lisinopril, amlodipine, or doxazosin (Davis et al., 1996; The ALLHAT Officers and Coordinators for the ALLHAT Collaborative Research Group, 2000).

The original GenHAT cohort included 39,114 participants (12,520 self-reported AA participants) (Arnett et al., 2002). A subsequent GWAS ancillary study generated genome-wide genotype data on a subset of AA participants randomized to chlorthalidone or lisinopril (Armstrong et al., 2022; Armstrong et al., 2021). Because ischemic and hemorrhagic stroke were not reported separately in ALLHAT/GenHAT, GenHAT analyses used an all-stroke phenotype.

Stroke endpoint

REGARDS

IS events in REGARDS were identified through biannual surveillance and adjudicated via medical record review. For participants aligned to Medicare claims, additional ischemic stroke events were identified from inpatient claims using ICD-9 codes 433. x1 or 434. x1 or ICD-10 codes I63. xx in the primary diagnosis position. Hospitalizations occurring within 1 day of each other were merged into a single episode of care. Participants with self-reported stroke at baseline were excluded.

All of Us

To harmonize phenotype definitions with REGARDS, analyses in All of Us were restricted to participants aged ≥45 years. IS cases were defined as individuals with an inpatient encounter associated with ICD-9 codes 433. x1 or 434. x1, or ICD-10 codes I63. xx. Stroke codes identified only in non-inpatient encounters, SNOMED stroke codes, and hemorrhagic stroke codes (ICD-9,430. x, 431. x; ICD-10 I60. x, I61. x) were excluded from both case and control sets.

GenHAT

Incident stroke in GenHAT was defined as a persistent neurological deficit resulting from arterial occlusion or rupture, lasting >24 h (unless fatal), or with CT/MRI evidence consistent with acute stroke, excluding stroke due to trauma, infection, tumor, or other non-ischemic causes. Events occurring perioperatively were included (The ALLHAT Officers and Coordinators for the ALLHAT Collaborative Research Group, 2000). In the primary ALLHAT study, the type of stroke (specifically differentiating between ischemic and hemorrhagic) was not systematically classified or reported as part of the trial’s pre-specific secondary outcomes; therefore, stroke was analyzed as an all-stroke phenotype in this cohort.

Genomic data

REGARDS optimization cohort

A subset of 10,551 AA and EA REGARDS participants underwent genotyping using the Illumina Infinium AMR/AFR (MEGA) BeadChip array. Following previously described quality control procedures (Armstrong et al., 2021), internal duplicates and HapMap controls were removed; samples were excluded when genetic sex did not match self-reported sex. Variants were excluded if located on sex chromosomes, were non–biallelic, had ambiguous strand alignment, violated Hardy–Weinberg equilibrium (p < 1.00E-12), had minor allele frequency (MAF) <5%, or had missingness >10%. Genotypes were imputed using the Trans-omics for Precision Medicine (TOPMed) Release 2 (Freeze 8) reference panel (Das et al., 2016).

Ancestry principal components (PCs) were generated using EIGENSTRAT (Price et al., 2006). Primary ancestry groups were defined based on self-reported ancestry, with genetic PCs included as covariates to account for within population substructure. After quality control (QC), 6,342 AA (422 ischemic stroke cases; 5,920 controls) and 1,403 EA (343 ischemic stroke cases; 1,060 controls) were included in the PRS optimization cohort.

Independent validation cohorts

REGARDS (Holdout)

A separate, independent batch of REGARDS genetic data included 12,118 participants of European ancestry genotyped on the Illumina Infinium Global Diversity Array-9 (GDA) with custom content targeting neurodegenerative disease-focused variants (Bandres-Ciga et al., 2024). Variants were excluded if they were located on sex chromosomes, had ambiguous strands, were multi-allelic, were indels, violated Hardy Weinberg Equilibrium (HWE, p < 1.00e-05), and/or had a missing rate >5%. Individuals were excluded based on sex discrepancies (genotyped versus self-reported), internal duplicates, or those with <98% call rates (Armstrong et al., 2026). Genotypes were imputed using the TOPMed release 3 reference panel. PCs were generated using EIGENSTRAT. After QC, 11,255 EA participants (439 ischemic stroke cases and 10,816 controls) remained for validation analysis.

All of Us

Short-read whole-genome sequencing (WGS) data from the All of Us Research Program were obtained from the Controlled Tier version 7 release (All of Us Research Program Genomics Investigators, 2024) and analyzed using the program’s standardized genomic pipeline (e.g., variant calling, QC, and imputation). Participants with low sequencing quality, discordant sex, excess heterozygosity, or contamination were removed. Only participants with WGS and complete and relevant phenotype data (n = 107,749; ∼27.5% AA) were included in analyses.

GenHAT

Genotyping was performed using the Illumina Infinium Multi-Ethnic AMR/AFR BeadChip (>1.4 M markers), and imputation used the TOPMed Release 2 (Freeze 8) reference panel. Detailed genotyping and imputation quality control procedures have been previously published (Armstrong et al., 2022; Armstrong et al., 2021). Imputation to TOPMed Release 2 (Freeze 8) was performed using Minimac4, retaining variants with imputation quality r2 > 0.5. After QC and restriction to participants with available phenotype and genetic data, 6,908 AA GenHAT participants were included in validation analyses.

Construction and optimization of IS PRSs

Bayesian PRS methods were constructed using GWAS summary statistics and external LD reference panels. Ancestry-specific GWAS summary statistics were obtained from the GIGASTROKE consortium (Mishra et al., 2022), including results from 1,296,908 EA (GCST90104540) and 20,924 AA (GCST90104550) individuals.

To assess potential sample overlap between the discovery GWAS and validation cohorts, available cohort descriptions from GIGASTROKE consortium studies were reviewed. To our knowledge, REGARDS, All of Us, and GenHAT participants included in the present analyses were not included in the ancestry-specific GIGASTROKE discovery summary statistics that we used for PRS construction. Therefore, all datasets were considered independent of the discovery GWAS.

Prior to PRS calculation, summary statistics and target genotype datasets were harmonized based on chromosome position (hg38) and reference/alternate alleles using the McCarthy Group Tools (Rayner, , 2011). Ambiguous strand variants, allele mismatches, and variants unavailable in the summary statistics, as well as the target genotyped datasets, were excluded during pre-processing. Only variants passing genotype QC and available after harmonization, were retained for PRS calculation.

PRS-CS

PRS-CS is a Bayesian polygenic prediction method that estimates posterior SNP effect sizes using GWAS summary statistics and an external LD reference panel (Ge et al., 2019). For PRS-CS analyses, we used the HM3 LD reference panels for EA and AA populations. SNP weights were generated under four values of the global shrinkage parameter (phi, ϕ): 1.0, 1.0E-02, 1.0E-04, and 1.0E-06. PRSs generated using each ϕ were evaluated separately within each ancestry group.

To evaluate whether expanded LD reference panels improve PRS, PRS-CS was additionally applied using the TagIt LD reference panel developed by the D-PRISM Consortium (Huerta-Chagoya et al., 2026). The TagIt panel expands the HM3 SNP set to increase variant density and improve LD estimation across diverse populations (Huerta-Chagoya et al., 2026). As with PRS-CS–HM3, four PRS-CS-TagIt PRSs were generated using the same grid of ϕ values.

SBayesRC

SBayesRC is a Bayesian summary-statistics-based method that incorporates GWAS data with functional genomic annotations, including coding, conserved, regulatory, and LD-related features (Zheng et al., 2024). SBayesRC extends SBayesR (Lloyd-Jones et al., 2019) by modeling all imputed common SNPs simultaneously and incorporating annotations through an annotation-dependent multicomponent mixture prior that allows annotations to influence both the probability that a SNP is causal and the distribution of its effect size (Zheng et al., 2024). SBayesRC requires GWAS summary statistics and an LD reference panel as inputs. Heritability and annotation parameters are estimated internally within the Bayesian framework. Outputs include posterior SNP effect estimates for PRS construction, posterior inclusion probabilities, and annotation-specific genetic architecture parameters.

In total, 18 PRSs were generated in the REGARDS optimization cohort (nine per ancestry group): four PRS-CS–HM3 scores, four PRS-CS–TagIt scores, and one SBayesRC score.

Statistical analysis

Baseline characteristics of study participants were summarized separately by self-reported ancestry group and cohort. Categorical variables were presented as counts and percentages, and continuous variables were summarized as mean (standard deviation, SD) as appropriate. Statistical analyses were conducted separately within each ancestry group. All PRSs were standardized to a mean of zero and an SD of 1 within ancestry group before analysis. Associations between PRS and incident stroke were evaluated using logistic regression models, with ORs and 95% confidence intervals (CIs) estimated per SD increase in the PRS (OR/SD). Base models (covariate-only) were adjusted for age, sex, and first 10 ancestry PCs. For GenHAT, the antihypertensive treatment randomization group was also included as a covariate due to the randomized trial design of ALLHAT. A PRS only model evaluated the association of IS (or all-stroke in GenHAT) with each individual PRS. A full model was developed for each PRS, evaluating the association between the outcome and PRS, while adjusting for the base model covariates (age, sex, PCs, and randomization group where applicable). All statistical analyses were performed using R version 4.4.2 (R Foundation for Statistical Computing). Two-sided p-values <0.05 were considered statistically significant.

Optimization of PRS performance

All PRSs were initially evaluated within the REGARDS optimization cohort to select the best-performing PRS-CS parameterization within each ancestry group. For PRS-CS-HM3 and PRS-CS-TagIt, performance was assessed across four global shrinkage parameters (ϕ = 1.0, 1.0E-02, 1.0E-04, 1.0E-06). Selection of the optimal PRS-CS score was based on multiple performance metrics including the proportion of variation in the trait explained by the PRS on the liability scale (liability-scale R2), OR/SD, and area under the receiver operating characteristic curve (AUC), with R2 as the primary selection metric. AUCs were calculated for the covariates-only model (age, sex, and PCs), the PRS-only model, and the full model. The optimization cohort was used exclusively for PRS parameter selection and was not considered an independent assessment of predictive performance.

Evaluation of PRS performance

The selected PRS-CS-HM3 and PRS-CS-TagIt scores for each ancestry group, along with the SBayesRC score, were subsequently evaluated in independent validation cohorts, including the REGARDS holdout cohort, All of Us, and GenHAT. Primary performance metrics included OR/SD, liability-scale R2, and AUC. AUC was evaluated for the covariate-only model, PRS-only model, and full model. Incremental discrimination was assessed using the change in AUC (ΔAUC), calculated as the difference between the full model and the covariate-only model.

The association between PRS and stroke risk was additionally evaluated by comparing individuals in the top 2%, 5%, and 10% of the PRS distribution with the remainder of the population. Sensitivity, specificity, adjusted positive predictive value (PPV), and adjusted negative predictive value (NPV) were estimated using population stroke prevalences of 2.7% and 4.3% for EA and AA, respectively (Imoisili et al., 2024) were used to calculate adjusted PPV, adjusted NPV, and the liability-scale R2.

Assessment of model calibration and incremental predictive performance

Additional evaluation of predictive performance was conducted in the independent REGARDS holdout and GenHAT validation cohorts. Calibration was assessed using the Brier score for the covariate-only and full models, the change in Brier score (ΔBrier), calibration intercept, and calibration slope, using the R package CalibrationCurves (Van Calster et al., 2016). Calibration curves were generated by comparing observed and predicted stroke risk across deciles of predicted risk and by plotting observed stroke risk against mean predicted risk across deciles of predicted risk for each PRS and ancestry group.

Continuous net reclassification improvement (NRI) quantified the improvement in risk classification associated with addition of the PRS by comparing the proportion of individuals with and without stroke. Integrated discrimination improvement (IDI) was calculated as the difference in discrimination slopes between the full models (covariate + PRS) and covariate-only models, reflecting the improvement in separation of predicted risks between stroke cases and non-cases after addition of the PRS. Confidence intervals for continuous NRI and IDI were estimated using bootstrap resampling (200 iterations) within the REGARDS holdout and GenHAT validation cohorts. To formally compare discrimination across PRS methods, receiver-operating characteristic (ROC) curves were generated using the pROC package in R (Robin et al., 2011), for each PRS-enhanced model within the independent REGARDS holdout and GenHAT validation cohorts. Pairwise comparisons of full model AUCs between correlated ROC curves were performed using DeLong’s test. Comparisons were conducted between PRS approaches within each ancestry group between (i) PRS-CS-HM3 and PRS-CS-TagIt, (ii) PRS-CS-HM3 and SBayesRC, and (iii) PRS-CS-TagIt and SBayesRC to evaluate whether any observed differences in discrimination between PRS methods were statistically significant. P-values were adjusted using the Holm method within each ancestry-specific validation cohort to control the family-wise error rate.

Application of external PRS for comparative assessment

We evaluated two publicly available ischemic stroke PRSs retrieved from the Polygenic Score (PGS) Catalog (Lambert et al., 2021; Lambert et al., 2024). The Abraham et al. score (PGS000039) is a metaPRS composed of 19 component scores, and the Mishra et al. score (PGS002724) is an integrative metaPRS comprising 22 component polygenic scores. These published scores were applied to the REGARDS holdout, GenHAT, and All of Us validation cohorts to provide a point of comparison for predictive performance. Because these external scores differ substantially in construction methodology and component variants, direct statistical comparisons with the newly generated PRSs were not performed.

Results

Study populations and PRS construction

A total of 133,657 participants (42,844 AA) were included across the REGARDS (optimization and holdout sub-cohorts), GenHAT, and All of Us studies after applying quality control filters and restricting analyses to individuals with complete covariate and phenotype data (Figure 1). Participant demographics are summarized in Table 1. The REGARDS optimization cohort included 6,342 AA participants (mean age 63.6 ± 9.2 years; 60.9% female) and 1,403 EA participants (mean age 68.3 ± 10.4 years; 42.3% female) from REGARDS. Independent validation cohorts included 29,594 AA and 78,155 EA participants from All of Us, 6,908 AA participants from GenHAT, and 11,255 EA participants from the REGARDS holdout cohort (mean age range: 63.3–68.8 years; female representation: 51.6%–58.2%).

FIGURE 1.

Flowchart depicting ancestry-specific genetic risk score optimization using summary statistics. Inputs include PRS-CS-HM3, PRS-CS-TagIt, and SBayesRC, each utilizing ancestry-specific GWAS summary statistics and reference panels. PRS-CS methods evaluate performance across different settings, selecting scores that maximize liability R squared, odds ratio, or AUC. SBayesRC estimates scores directly. Optimized polygenic risk scores are then independently validated in three study cohorts: REGARDS Holdout (EA), GenHAT (AA), and All of Us (EA and AA), each with case and control numbers specified.

Overview of study design, PRS development and optimization, and validation. Ancestry-specific (EA and AA) summary statistics from the GIGASTROKE consortium were used to construct PRS using three methods: PRS-CS with the HM3 reference panel (PRS-CS-HM3), PRS-CS with the TagIt reference panel (PRS-CS-TagIt), and SBayesRC. For PRS-CS, global shrinkage parameters (phi, ɸ) of 1.0E-06, 1.0E-04, 1.0E-02, and 1 were evaluated, and the parameter that maximized the liability-scale R2, had the strongest parameter estimate (odds ratio), and the largest PRS AUC, was selected as the optimized PRS (ɸ = 1.0E-02). For SBayesRC, the optimal PRS is output directly. Optimized PRS were evaluated in three independent cohorts (no participant overlaps with the discovery GWAS used for PRS construction or with the REGARDS optimization cohort). Abbreviations: AA-African ancestry; EA-European ancestry; IS-ischemic stroke; EHR-electronic health records; AUC-area under the curve.

TABLE 1.

Sample characteristics of optimization and testing cohorts.

Study Ancestry Age, mean ± SD Female sex, % N cases N controls N total
Optimization
REGARDS AA 63.6 ± 9.2 60.9 422 5,920 6,342
REGARDS EA 68.3 ± 10.4 42.3 343 1,060 1,403
Validation/Testing
Study Ancestry Age, mean ± SD Female sex, % N cases N controls N total
All of us AA 63.3 ± 8.1 53.9 311 29,283 29,594
All of us EA 68.8 ± 10.0 58.2 898 77,257 78,155
GenHAT AA 66.2 ± 7.8 55.4 366 6,542 6,908
REGARDS holdout EA 64.0 ± 9.1 51.6 439 10,816 11,255

Abbreviations: AA- african ancestry; EA- european ancestry.

PRSs were generated using three Bayesian approaches: PRS-CS-HM3, PRS-CS-TagIt, and SBayesRC, stratified by ancestry. Variant overlap across validation cohorts was high for all PRS approaches, supporting reproducibility of PRS calculation across datasets (Supplementary Table S1). Among AA participants, PRS-CS-HM3 included 983,900 variants in the optimization cohort (>99% overlap in validation cohorts), PRS-CS-TagIt included 1,368,789 variants (>99.99% overlap in validation cohorts), and SBayesRC included 6,064,174 variants with 89% overlap in All of Us and 98% overlap in GenHAT. Among EA participants, PRSs generated in the optimization cohort included 1,093,847 variants for PRS-CS-HM3 (>98% overlap in the validation cohorts), 1,368,236 variants for PRS-CS-TagIt (>97% overlap in validation cohorts), and 7,356,518 variants for SBayesRC (89% and 97% overlap in All of Us and REGARDS holdout cohort, respectively).

PRS optimization and performance in independent validation cohorts

PRS parameter optimization was performed exclusively within the REGARDS optimization cohorts, and a ϕ of 1.00E-02 was selected for PRS-CS-HM3 (AA and EA) and PRS-CS-TagIt (AA and EA) (Supplementary Table S2; Supplementary Table S3). Estimates from the optimization cohorts (AA and EA) were not considered independent estimates of predictive performance.

Primary validation results for the selected PRSs are presented in Table 2, Supplementary Table S4, Supplementary Table S5, and Figure 2. Across validation cohorts, associations and predictive metrics were consistently stronger among EA participants compared with AA participants. Among EA participants, SBayesRC demonstrated the largest point estimates for effect size and variance explained among the evaluated PRS approaches. In All of Us participants, the SBayesRC score was associated with incident IS with an OR/SD of 1.16 (95% CI = 1.12–1.19) and R2-liability of 0.59%. Similar patterns were observed in the REGARDS holdout validation cohort, where SBayesRC demonstrated an OR/SD of 1.24 (95% CI: 1.12–1.36) and liability R2 of 2.42%. PRS-CS-derived scores showed consistent but modest associations across EA validation cohorts (Table 2; Figure 2; Supplementary Table S4).

TABLE 2.

Validation performance per 1 standard deviation of ischemic stroke PRS.

Ancestry Study Method Or (95% CI) P AUC PRS (95% CI)
AA All of us v7 PRS-CS-HM3 1 1.04 (0.99, 1.09) 1.30E-01 51.71% (48.08%, 55.34%)
PRS-CS-TagIt 1 1.03 (0.98, 1.08) 2.94E-01 49.97% (46.23%, 53.71%)
SBAYESRC 0.99 (0.94, 1.05) 7.91E-01 50.26% (46.70%, 53.81%)
GenHAT 2 PRS-CS-HM3 1 1.11 (0.98, 1.27) 1.06E-01 51.11% (48.10%, 54.12%)
PRS-CS-TagIt 1 1.14 (1.03, 1.26) 9.48E-03 53.68% (50.69%, 56.68%)
SBAYESRC 1.01 (0.89, 1.14) 8.86E-01 50.15% (46.97%, 53.32%)
EA All of us v7 PRS-CS-HM3 1 1.10 (1.07, 1.14) 3.25E-11 53.88% (51.77%, 56.00%)
PRS-CS-TagIt 1 1.10 (1.07, 1.13) 8.43E-10 53.85% (51.76%, 55.94%)
SBAYESRC 1.16 (1.12, 1.19) 1.62E-22 54.47% (52.36%, 56.57%)
REGARDS (holdout) PRS-CS-HM3 1 1.09 (0.98, 1.20) 9.97E-02 52.05% (49.29%, 54.82%)
PRS-CS-TagIt 1 1.11 (1.00, 1.22) 4.76E-02 52.18% (49.40%, 54.97%)
SBAYESRC 1.24 (1.12, 1.36) 1.81E-05 55.16% (52.49%, 57.82%)
1

Optimal phi = 1.00E-02.

2

GenHAT, is evaluating ischemic stroke PRS, on all-stroke phenotype.

Abbreviations: AA- african ancestry; EA- european ancestry; OR-odds, ratio; CI- confidence interval; AUC- area under the curve; HM3- HapMap3.

FIGURE 2.

Bar chart comparing polygenic risk score (PRS) performance by R squared liability for African American (AA) and European American (EA) groups using PRS-CS-TagIt, PRS-CS-HM3, and SBAYESRC methods, with AA group showing highest R squared in GenHAT (blue) and EA group in REGARDS (green). Legend indicates studies: All of Us (red), REGARDS (green), GenHAT (blue).

R2 liability comparison per 1 SD of PRS among testing/validation cohorts. Liability-scale R2 represents the proportion of variance in ischemic stroke risk explained by each PRS per 1 standard deviation (SD) increase in PRS. Results are shown separately for African ancestry (AA, Left Panel) and European ancestry (EA, Right Panel) participants across independent testing/validation cohorts, including All of Us (AA and EA), REGARDS holdout (EA), and GenHAT (AA). PRS performance was evaluated for three Bayesian PRS approaches: PRS-CS using the TagIt linkage disequilibrium (LD) reference panel (PRS-CS-TagIt), PRS-CS using the HapMap3 LD reference panel (PRS-CS-HM3), and SBayesRC incorporating functional annotations. Higher liability-scale R2 values indicate greater variance explained by the PRS.

In contrast, associations and predictive metrics were attenuated among AA participants. In All of Us AA participants, associations for the generated PRSs showed modest associations and were close to the null, with ORs close to 1.0, and R2-liability values below 0.1%. In GenHAT, PRS-CS-TagIt provided the strongest association among the novel scores (OR/SD = 1.14, 95% CI: 1.03–1.26, R2-liability = 0.87%), although the magnitude of association remained lower than that observed in EA cohorts (Table 2; Figure 2; Supplementary Table S5).

PRS percentile-based risk stratification

Associations comparing individuals in the highest PRS percentile groups (e.g., 2%–3% strata) with the remainder of the population revealed modest enrichment of stroke events (Supplementary Tables S4,S5). Exploratory thresholds of 3% for EA and 2% for AA were selected based on the distribution of performance metrics observed during PRS optimization in REGARDS.

Among EA participants from All of Us, the top 3% (versus bottom 97%) of the PRS distribution showed increased stroke risk across all PRSs. PRS-CS-TagIt, PRS-CS-HM3, and SBayesRC demonstrated ORs of 1.34 (95% CI: 1.15–1.57, p = 1.52E-04), 1.44 (95% CI: 1.24–1.67, p = 2.65 E−06), and 1.44 (95% CI: 1.24–1.67, p = 2.35E-06), respectively. Similar enrichment was observed in the REGARDS EA holdout cohort, with the strongest associations observed for the PRS-CS-TagIt (OR = 1.71, 95% CI: 1.08–2.71, p = 2.11E-02) and SBayesRC (OR = 1.82, 95% CI: 1.16–2.84, p = 8.80E-03) scores (Supplementary Table S4).

Among AA participants, threshold-based associations (comparing the top 2% of the PRS distribution to the bottom 98%) were less consistent across cohorts. In All of Us AA participants, some nominal associations were observed for PRS-CS-HM3 (OR = 1.49, 95% CI: 1.10–2.02, p = 9.17 × 10−3), whereas PRS-CS-TagIt (OR = 1.21, 95% CI 0.88–1.66, p = 2.34E-01) and SBayesRC (OR = 1.20, 95% CI: 0.87–1.66, p = 2.73E-01) were not statistically significant. No PRS demonstrated robust predictive performance in GenHAT (Supplementary Table S5).

Across cohorts, sensitivity values remained limited, ranging from 1% to 25%, while specificity was consistently high (≥90%). Adjusted PPVs were modest, while adjusted NPVs were stable across threshold cut points. Among EA participants, the covariate model achieved AUCs of 66.0% in All of Us and 62.4% in the REGARDS holdout cohort. Addition of PRSs modestly increased AUCs to 66.3%–66.6% for All of Us and 62.6%–63.8% in the REGARDS holdout cohort, depending on the PRS method. Among AA participants, the covariate model achieved AUCs of 66.6% in All of Us and 62.7% in GenHAT, with little or no improvement following inclusion of the PRSs (62.7%–66.7%). Complete per SD or percentile-based estimates for each method, ancestry group, and cohort are provided in Supplementary Tables S4,S5.

Incremental predictive performance, calibration, and comparison of PRS approaches

Addition of PRS to covariate-adjusted models resulted in modest improvements in discrimination beyond covariates alone. Across validation cohorts, ΔAUC values were small, ranging from 0.002% to 1.33% among EA participants and were smaller than 0.38% among AA participants. The largest improvements in discrimination were observed for SBayesRC in EA participants, although absolute improvements remained modest (Supplementary Tables S4,S5).

Calibration analyses conducted in the REGARDS EA holdout and GenHAT AA validation cohorts demonstrated minimal changes in model performance after addition of PRS. Brier scores showed small absolute improvements for PRS-enhanced models compared to covariate-only models, with the largest reduction observed for SBayesRC in the REGARDS holdout cohort. Calibration slopes were approximately one and calibration intercepts were close to zero across PRS models, suggesting reasonable agreement between predicted and observed risk (Supplementary Table S6). Calibration plots demonstrating observed versus predicted stroke risk across risk deciles are provided in Supplementary Figures S1–S6.

Adding the PRSs to our covariate-only models did not meaningfully improve our ability to classify patient stroke risk. In the REGARDS EA holdout cohort, the SBayesRC method showed the largest improvement, but the actual change in predicted probabilities was negligible (IDI = 0.002, 95% CI 0.001–0.003). All other PRS methods showed similarly minimal improvements over the basic covariates. In the GenHAT cohort, the metrics showed almost no change, and the results were not statistically significant (Supplementary Table S6).

Formal comparisons of PRS discrimination were performed using DeLong’s test comparing ROC curves among PRS-CS-HM3, PRS-CS-TagIt, and SBayesRC. In the REGARDS EA holdout cohort, SBayesRC showed higher discrimination compared to PRS-CS-HM3, although this relationship was not significant when adjusting for multiple comparisons (pHolm= 0.136). Comparisons between PRS-CS-HM3 and PRS-CS-TagIt and PRS-CS-TagIt and SBayesRC were not statistically significant. In GenHAT AA participants, no significant differences in discrimination were observed between PRS approaches (all p > 0.30; Supplementary Table S7). These findings suggest that although some approaches demonstrated numerically improved performance, differences between PRS methods were generally modest. Thus, while expanded LD reference panels and functional annotation-informed approaches showed numerical improvements in some settings, evidence for statistically significant superiority of one PRS method over another was limited.

Comparison with previously published stroke PRSs

Two previously published ischemic stroke PRSs from the Polygenic Score Catalog (Abraham et al. and Mishra et al.) were evaluated in secondary analyses. Among EA participants, both external PRSs demonstrated positive associations with IS risk, with similar performance between the two scores in All of Us and modest but statistically significant associations in the REGARDS holdout cohort. Addition of the external PRSs resulted in small improvements in discrimination beyond covariate-only models. Among AA participants, associations were weaker and less consistent across cohorts. The Mishra PRS demonstrated modest associations in All of Us AA participants (OR/SD = 1.17, 95% CI: 1.01–1.36) and GenHAT participants (OR/SD = 1.38, 95% CI: 1.19–1.60), although improvements in discrimination remained small, consistent with findings from the newly generated PRSs (Supplementary Table S8).

Discussion

PRSs have become an increasingly appealing tool in cerebrovascular research, driven by the availability of large-scale GWAS summary statistics and the development of advanced Bayesian modeling approaches. Despite this rapid progress, clinical translation for IS remains limited in comparison to other outcomes (Kurniansyah et al., 2022; Patel et al., 2023), in part because of modest predictive performance and persistent disparities across ancestry groups (Neumann et al., 2021; Moreno-Grau et al., 2024; Abraham et al., 2021). Although methods that incorporate functional annotations or leverage more representative LD reference panels have the potential to improve prediction accuracy and enhance portability across populations (Zeng and Visscher, 2025), their utility for IS PRSs has not been comprehensively evaluated. GWAS have identified hundreds of loci associated with vascular and cerebrovascular diseases, and PRS provide a scalable way to aggregate the polygenic signal (Jagodic et al., 2025). In large consortia studies, IS PRS have shown modest associations that persist after adjustment for major clinical factors, supporting their potential role in early risk stratification (Mishra et al., 2022). However, few studies have systematically evaluated and compared multiple PRS methods for IS prediction in AA adults, despite the disproportionate burden of stroke (Owolabi et al., 2017; Madsen et al., 2024; Howard et al., 2007) in this population. Improving the performance and equitable application of PRS in AA adults is therefore critical to avoid widening existing disparities in genetic risk assessment and future precision prevention strategies. In the present study, we examined whether adding functional annotations and expanding the LD reference panel could improve PRS performance for incident stroke compared with conventional PRS-CS-HM3 approaches among EA and AA adults.

Consistent with prior work suggesting that causal common variants for complex traits are often shared across ancestry groups (Hou et al., 2023), functional annotations have been proposed to help distinguish true causal SNPs from correlated non-causal markers (Li et al., 2024; Weissbrod et al., 2020). However, incorporating annotations increases model complexity and may dilute effect estimates when the number of included variants is large (Zhuang et al., 2024). In our analysis, the annotation-informed PRS (SBayesRC) incorporated approximately three times more SNPs than the standard PRS-CS models. Although functional priors led to a slight improvement in prediction for both EA and AA participants, the magnitude of improvement was modest and insufficient to meaningfully reduce cross-ancestry disparities or outperform previously published scores. One explanation may be that most stroke-associated variants lie in regulatory regions, rather than coding sequences (Malik et al., 2018; Jagodic et al., 2025), limiting the utility of annotation frameworks that prioritize coding or defined functional categories. These results highlight that, while annotation-informed approaches have potential in stroke genomics, careful consideration and model complexity are critical to achieve meaningful improvements in PRS performance across diverse populations.

Prediction performance remained substantially lower among AA participants across all evaluated approaches. This aligns with numerous studies documenting reduced PRS transferability to African ancestry populations (Ju et al., 2022; Martin et al., 2019). The primary drivers are well documented and include underrepresentation in GWAS, differences in LD architecture, and a higher proportion of population-specific variants that are not adequately tagged by current panels or methods (Kachuri et al., 2024). To address this, we evaluated the D-PRISM Consortium’s expanded TagIt LD panel, which includes >1 million additional variants and better represents global genetic diversity. Although PRS-CS-TagIt demonstrated numerically improved performance compared with PRS-CS-HM3 in some analyses, improvements were generally modest especially in AAs and did not translate into consistent statistically significant differences in discrimination based on formal ROC comparisons. These findings suggest that expanded LD reference panels may provide incremental improvements in PRS performance but cannot fully overcome limitations related to ancestry differences in GWAS summary statistic sample size, LD structure, and allele frequency distributions. Notably, our study was focused on the role of functional annotations and updated LD reference information in within ancestry settings. Thus, we did not leverage available multi-ancestry methods such as PRS-CSx (Ruan et al., 2022), Genomic Prediction through Transfer Learning (GPTL) (Wu et al., 2026), BridgePRS (Hoggart, et al., 2024) (Hoggart et al., 2024), or PROSPER (polygenic risk scores based on ensemble of penalized regression models) (Zhang et al., 2024), which should be considered in future IS PRS efforts.

PRS research for IS lags behind other cardiovascular conditions, such as hypertension (Kurniansyah et al., 2022) or coronary artery disease (Patel et al., 2023), where multi-ancestry PRS have been more extensively evaluated and generally show larger effect sizes (often with effect estimates >1.5 per SD (Agbaedeng et al., 2021; King et al., 2022; Mars et al., 2023)). Although our models performed similarly to previously published scores, absolute prediction remained low. This difference likely reflects the complex and heterogeneous biology of stroke, which includes multiple etiologic mechanisms and clinically distinct subtypes. Therefore, even substantial improvements in genetic prediction methodology may yield only incremental gains when applied to IS as a broad phenotype.

Several published stroke PRSs have been constructed as metaPRSs or “mega-scores,” including those developed by Abraham et al. and Mishra et al. Rather than being derived from a single GWAS, these scores integrate multiple trait-specific PRS representing genetically correlated cardiovascular and cardiometabolic phenotypes, such as blood pressure, lipid traits, and atrial fibrillation. This multi-trait architecture allows metaPRS to capture a broader range of biological pathways contributing to ischemic stroke, consistent with the heterogeneous and multifactorial etiology of the disease. Because their construction differs from single-trait or functionally informed PRS, direct comparisons across score types are inherently challenging, and observed differences in predictive performance should be interpreted cautiously. Integrating multiple trait-specific signals may also enhance robustness across cohorts, as the aggregated score is less dependent on any single discovery dataset or genetic architecture. Consistent with this concept, our evaluation of previously published metaPRS approaches demonstrated comparable or modestly improved performance in some settings; however, these scores also did not overcome the broader limitations in cross-ancestry prediction.

A strength of our study is the use of rigorously adjudicated stroke outcomes and harmonized genetic data in a well-characterized U.S. cohort with substantial representation of AA adults for the generation of the PRS. Additionally, the direct comparison of functional priors and LD panel choice within a consistent analytic workflow provides unique insight into the relative contributions of each methodological component. We further evaluated clinical predictive performance beyond traditional association metrics by examining discrimination improvement, calibration, and reclassification measures. Although some approaches improved predictive metrics numerically, the incremental improvement beyond covariates alone was modest, with covariate-only model AUCs ranging from 62.4% to 66.6% and covariate-plus-PRS model AUCs ranging from 62.6% to 66.7% across validation cohorts and PRS methods, highlighting the challenges of translating stroke PRSs into clinically meaningful prediction tools.

However, this study is not without limitations. First, while we relied on single-ancestry summary statistics from the largest-to-date stroke GWAS for both EA and AA populations, our sample size, particularly for the AA analyses, was limited. Larger GWAS efforts, particularly those including greater representation of historically underrepresented populations, will be necessary to improve PRS development and portability. Second, PRS parameter selection for PRS-CS-HM3 and PRS-CS-TagIt was performed within the REGARDS optimization cohort. Although this approach allowed selection of ancestry-specific model parameters prior to independent evaluation, estimates obtained during optimization may be subject to optimism due to model selection. Therefore, optimization cohort results were used only for parameter selection and were not interpreted as independent estimates of predictive performance. Third, we evaluated ancestry-specific rather than multi-ancestry approaches. Although ancestry-specific models allowed direct assessment of population-specific performance, they do not address whether jointly modeling diverse populations may improve PRS portability or predictive performance. Fourth, we focused strictly on IS as a single disease and not subtype-specific IS specifically due to GWAS summary statistic constraints, acknowledging that IS subtypes have distinct genetic architectures and that subtype distributions may differ across ancestries. Fifth, we focused on EA and AA populations due to the composition of REGARDS, which may limit generalizability to other ancestry groups. Sixth, our inability to distinguish ischemic from hemorrhagic stroke somewhat limits the interpretation of the results in GenHAT, however IS known to make up ∼85–90% of stroke cases in the US (Martin et al., 2025); therefore, we felt the inclusion of this data in our validation set was warranted. Finally, while we aimed to harmonize the ischemic stroke phenotype, we note that there is likely heterogeneity in our cases and controls based on the differing study designs (e.g., REGARDS is a population-based longitudinal study with adjudicated outcomes; ALLHAT/GenHAT is a randomized clinical trial with an all-stroke outcome; All of Us through electronic medical records).

In summary, we demonstrate that Bayesian approaches incorporating expanded LD reference panels and functional annotations provided modest improvements in IS performance among EA individuals, but yielded limited gains beyond covariates alone. However, despite these incremental gains, model performance for AA participants remained substantially lower, underscoring persistent ancestry-related disparities. IS risk is shaped by a complex interplay of genetic, clinical, environmental, and social factors, and genetic risk only represents one dimension of this larger framework. As a result, developing PRS that focus on clinical subtypes of IS and then integrating those scores with clinical risk factors, lifestyle patterns, and social determinants will likely be needed to achieve meaningful improvements in prediction. Overall, continued expansion of diverse genomic resources and methodological innovation remains essential to advancing the clinical utility of PRS for stroke prevention and risk stratification.

Acknowledgments

The authors thank the other investigators, the staff, and the participants of the REGARDS study for their valuable contributions. A full list of participating REGARDS investigators and institutions can be found at: https://www.uab.edu/soph/regardsstudy/. Further, the authors thank the participants of the All of Us Research Program, without whom this work would not be possible. This research was conducted using data accessed through the All of Us Researcher Workbench (https://www.researchallofus.org/data-tools/workbench), supported by the National Institutes of Health.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. The REGARDS genetic study (R01HL136666, MRI, LAL) and polygenic risk score study (R35HL155466, MRI) were supported by the NHLBI. The parent REGARDS study was supported by cooperative agreement U01 NS041588, co-funded by the National Institute of Neurological Disorders and Stroke (NINDS) and the National Institute on Aging (NIA, ZO1 AG000949). The content is solely the responsibility of the authors and does not necessarily represent the official views of the NINDS or the NIA. This research was supported, in part, by the Intramural Research Program of the National Institutes of Health (NIH). The contributions of the NIH author(s) are considered Works of the United States Government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services. The GenHAT genetics study was supported by NHLBI (R01HL123782). Other funding sources include internal pilot funding provided by the University of Alabama at Birmingham School of Public Health Office of Research (NDA).

Footnotes

Edited by: Sare Verstockt, Metabolism and Ageing, Belgium

Reviewed by: Gerald Mboowa, Makerere University, Uganda

Hao Wu, Michigan State University, United States

Data availability statement

The REGARDS (Study Accession: phs002719.v1.p1) and GenHAT (Study Accession: phs002716.v1.p1) phenotypic and genetic data used in this study are available through the National Center for Biotechnology Information (NCBI) database of Genotypes and Phenotypes (dbGaP) to approved researchers following applicable data access procedures. This study also used Controlled Tier data from the All of Us Research Program through an institutional Data Use and Registration Agreement (DURA) which allows UAB researchers who have completed required training, and approvals to access individual level data through the All of Us Researcher Workbench. Individual-level participant data cannot be made publicly available outside the platform. The REGARDS holdout sample used for validation analyses is supported by the National Institute on Aging (NIA) and the National Institute of Neurological Disorders and Stroke (NINDS) through the Center for Alzheimer’s and Related Dementias (CARD) and is made available through the AD Knowledge Portal (https://adknowledgeportal.synapse.org) in accordance with applicable data access policies.

Summary statistics from the GIGASTROKE genome-wide association study used for polygenic risk score development are publicly available through the NHGRI-EBI GWAS Catalog (ID: 36180795). Previously published polygenic scores evaluated in this study are publicly available through the Polygenic Score Catalog, including the Abraham et al. score (PGS Catalog ID: PGS000039) and the Mishra et al. score (PGS Catalog ID: PGS002724).The polygenic risk scores developed in this study are found under the publication ID PGP000839 (score IDs PGS019947-PGS019952). Scripts and code for PRS construction are available through GitHub (https://github.com/uabgenepi/stroke_prs/wiki).

Ethics statement

The studies involving humans were approved by University of Alabama at Birmingham (IRB-300006721). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

NA: Writing – original draft, Writing – review and editing, Funding acquisition, Visualization, Formal Analysis, Software, Conceptualization, Methodology, Data curation, Investigation. VS: Formal Analysis, Data curation, Methodology, Writing – review and editing. AP: Writing – review and editing, Software, Data curation, Formal Analysis. LP: Data curation, Writing – review and editing. CA: Writing – review and editing. AJ: Methodology, Writing – review and editing. UB: Writing – review and editing. EL: Writing – review and editing, Methodology. LL: Resources, Funding acquisition, Writing – review and editing. PA: Writing – review and editing. NL: Writing – review and editing, Resources. HI: Writing – review and editing, Resources, Data curation. LS: Funding acquisition, Writing – review and editing, Resources, Data curation. JK: Writing – review and editing, Resources. MN: Resources, Writing – review and editing. JM: Resources, Writing – review and editing. KK: Writing – review and editing. BW: Writing – review and editing. HT: Conceptualization, Writing – review and editing, Supervision, Methodology. MI: Resources, Writing – original draft, Writing – review and editing, Conceptualization, Funding acquisition, Supervision.

Conflict of interest

Author HI was employed by DataTecnica LLC.

The remaining author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fbinf.2026.1875388/full#supplementary-material

Table1.xlsx (60.6KB, xlsx)

References

  1. Abraham G., Malik R., Yonova-Doing E., Salim A., Wang T., Danesh J., et al. (2019). Genomic risk score offers predictive performance comparable to clinical risk factors for ischaemic stroke. Nat. Commun. 10 (1), 5819. 10.1038/s41467-019-13848-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Abraham G., Rutten-Jacobs L., Inouye M. (2021). Risk prediction using polygenic risk scores for prevention of stroke and other cardiovascular diseases. Stroke 52 (9), 2983–2991. 10.1161/strokeaha.120.032619 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Agbaedeng T. A., Noubiap J. J., Mofo Mato E. P., Chew D. P., Figtree G. A., Said M. A., et al. (2021). Polygenic risk score and coronary artery disease: a meta-analysis of 979,286 participant data. Atherosclerosis 333, 48–55. 10.1016/j.atherosclerosis.2021.08.020 [DOI] [PubMed] [Google Scholar]
  4. Alhusayni A. I., Alzahrani A. H. (2025). Length of stay in hospital and rehabilitation centers after stroke in Arab countries and Saudi Arabia: a systematic review and meta-analysis. Ann. Saudi Med. 45 (4), 256–269. 10.5144/0256-4947.2025.256 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. All of Us Research Program Investigators Denny J. C., Rutter J. L., Goldstein D. B., Philippakis A., Smoller J. W., et al. (2019). The “All of us” research program. N. Engl. J. Med. 381 (7), 668–676. 10.1056/NEJMsr1809937 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. All of Us Research Program Genomics Investigators (2024). Genomic data in the all of us research program. Nature 627 (8003), 340–346. 10.1038/s41586-023-06957-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Armstrong N. D., Srinivasasainagendra V., Patki A., Tanner R. M., Hidalgo B. A., Tiwari H. K., et al. (2021). Genetic contributors of incident stroke in 10,700 African Americans with hypertension: a meta-analysis from the genetics of hypertension associated treatments and reasons for geographic and racial differences in stroke studies. Front. Genet. 12, 781451. 10.3389/fgene.2021.781451 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Armstrong N. D., Srinivasasainagendra V., Chekka L. M. S., Nguyen N. H. K., Nahid N. A., Jones A. C., et al. (2022). Genetic contributors of efficacy and adverse metabolic effects of chlorthalidone in African Americans from the genetics of hypertension associated treatments (GenHAT) study. Genes (Basel) 13 (7), 1260. 10.3390/genes13071260 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Armstrong N. D., Srinivasasainagendra V., Pilla L., Gangaraju R., Hunt P. W., Nance R. M., et al. (2026). Genomic risk prediction of type 2 diabetes in people living with and without HIV. Sci. Rep. 16 (1), 3078. 10.1038/s41598-025-31471-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Arnett D. K., Boerwinkle E., Davis B. R., Eckfeldt J., Ford C. E., Black H. (2002). Pharmacogenetic approaches to hypertension therapy: design and rationale for the genetics of hypertension associated treatment (GenHAT) study. Pharmacogenomics J. 2 (5), 309–317. 10.1038/sj.tpj.6500113 [DOI] [PubMed] [Google Scholar]
  11. Bandres-Ciga S., Faghri F., Majounie E., Koretsky M. J., Kim J., Levine K. S., et al. (2024). NeuroBooster array: a genome-wide genotyping platform to study neurological disorders across diverse populations. Mov. Disord. 39 (11), 2039–2048. 10.1002/mds.29902 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Bevan S., Traylor M., Adib-Samii P., Malik R., Paul N. L. M., Jackson C., et al. (2012). Genetic heritability of ischemic stroke and the contribution of previously reported candidate gene and genomewide associations. Stroke 43 (12), 3161–3167. 10.1161/STROKEAHA.112.665760 [DOI] [PubMed] [Google Scholar]
  13. Chen X., Zhou L., Zhang Y., Yi D., Liu L., Rao W., et al. (2014). Risk factors of stroke in Western and Asian countries: a systematic review and meta-analysis of prospective cohort studies. BMC Public Health 14, 776. 10.1186/1471-2458-14-776 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Das S., Forer L., Schönherr S., Sidore C., Locke A. E., Kwong A., et al. (2016). Next-generation genotype imputation service and methods. Nat. Genet. 48 (10), 1284–1287. 10.1038/ng.3656 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Davis B. R., Cutler J. A., Gordon D. J., Furberg C. D., Wright J. T., Cushman W. C., et al. (1996). Rationale and design for the antihypertensive and lipid lowering treatment to prevent heart attack trial (ALLHAT). ALLHAT research group. Am. J. Hypertens. 9 (4 Pt 1), 342–360. 10.1016/0895-7061(96)00037-4 [DOI] [PubMed] [Google Scholar]
  16. Donkor E. S. (2018). Stroke in the 21(st) century: a snapshot of the burden, epidemiology, and quality of life. Stroke Res. Treat. 2018, 3238165. 10.1155/2018/3238165 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Ge T., Chen C. Y., Ni Y., Feng Y. C. A., Smoller J. W. (2019). Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat. Commun. 10 (1), 1776. 10.1038/s41467-019-09718-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Godwin K. M., Ostwald S. K., Cron S. G., Wasserman J. (2013). Long-term health-related quality of life of stroke survivors and their spousal caregivers. J. Neurosci. Nurs. 45 (3), 147–154. 10.1097/JNN.0b013e31828a410b [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Hoggart C. J., Choi S. W., García-González J., Souaiaia T., Preuss M., O’Reilly P. F. (2024). BridgePRS leverages shared genetic effects across ancestries to increase polygenic risk score portability. Nat. Genet. 56 (1), 180–186. 10.1038/s41588-023-01583-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Hou K., Ding Y., Xu Z., Wu Y., Bhattacharya A., Mester R., et al. (2023). Causal effects on complex traits are similar for common variants across segments of different continental ancestries within admixed individuals. Nat. Genet. 55 (4), 549–558. 10.1038/s41588-023-01338-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Howard V. J., Cushman M., Pulley L., Gomez C. R., Go R. C., Prineas R. J., et al. (2005). The reasons for geographic and racial differences in stroke study: objectives and design. Neuroepidemiology 25 (3), 135–143. 10.1159/000086678 [DOI] [PubMed] [Google Scholar]
  22. Howard G., Labarthe D. R., Hu J., Yoon S., Howard V. J. (2007). Regional differences in African Americans' high risk for stroke: the remarkable burden of stroke for southern African Americans. Ann. Epidemiol. 17 (9), 689–696. 10.1016/j.annepidem.2007.03.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Howard V. J., Kleindorfer D. O., Judd S. E., McClure L. A., Safford M. M., Rhodes J. D., et al. (2011). Disparities in stroke incidence contributing to disparities in stroke mortality. Ann. Neurol. 69 (4), 619–627. 10.1002/ana.22385 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Huerta-Chagoya A., Kim J., Mandla R., Lu Y., Suzuki K., Petty L. E., et al. (2026). Multi-ancestry polygenic risk scores for the prediction of type 2 diabetes and complications in diverse ancestries. Lancet Diabetes Endocrinol. 14 (7), 558–569. 10.1016/S2213-8587(25)00405-X [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Imoisili O. E., Chung A., Tong X., Hayes D. K., Loustalot F. (2024). Prevalence of stroke - behavioral risk factor surveillance system, United States, 2011-2022. MMWR Morb. Mortal. Wkly. Rep. 73 (20), 449–455. 10.15585/mmwr.mm7320a1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Jagodic A., Zivalj D., Krsek A., Baticic L. (2025). Genetic architecture of ischemic stroke: insights from genome-wide association studies and beyond. J. Cardiovasc Dev. Dis. 12 (8), 281. 10.3390/jcdd12080281 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Ju D., Hui D., Hammond D. A., Wonkam A., Tishkoff S. A. (2022). Importance of including non-european populations in large human genetic studies to enhance precision medicine. Annu. Rev. Biomed. Data Sci. 5, 321–339. 10.1146/annurev-biodatasci-122220-112550 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kachuri L., Chatterjee N., Hirbo J., Schaid D. J., Martin I., Kullo I. J., et al. (2024). Principles and methods for transferring polygenic risk scores across global populations. Nat. Rev. Genet. 25 (1), 8–25. 10.1038/s41576-023-00637-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Keene K. L., Hyacinth H. I., Bis J. C., Kittner S. J., Mitchell B. D., Cheng Y. C., et al. (2020). Genome-wide association study meta-analysis of stroke in 22 000 individuals of African descent identifies novel associations with stroke. Stroke 51 (8), 2454–2463. 10.1161/strokeaha.120.029123 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. King A., Wu L., Deng H. W., Shen H., Wu C. (2022). Polygenic risk score improves the accuracy of a clinical risk score for coronary artery disease. BMC Med. 20 (1), 385. 10.1186/s12916-022-02583-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Kurniansyah N., Goodman M. O., Kelly T. N., Elfassy T., Wiggins K. L., Bis J. C., et al. (2022). A multi-ethnic polygenic risk score is associated with hypertension prevalence and progression throughout adulthood. Nat. Commun. 13 (1), 3549. 10.1038/s41467-022-31080-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Lambert S. A., Gil L., Jupp S., Ritchie S. C., Xu Y., Buniello A., et al. (2021). The polygenic score catalog as an open database for reproducibility and systematic evaluation. Nat. Genet. 53 (4), 420–425. 10.1038/s41588-021-00783-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Lambert S. A., Wingfield B., Gibson J. T., Gil L., Ramachandran S., Yvon F., et al. (2024). Enhancing the polygenic score catalog with tools for score calculation and ancestry normalization. Nat. Genet. 56 (10), 1989–1994. 10.1038/s41588-024-01937-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Lewis C. M., Vassos E. (2017). Prospects for using risk scores in polygenic medicine. Genome Med. 9 (1), 96. 10.1186/s13073-017-0489-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Li Y., Xiao J., Ming J., Zeng Y., Cai M. (2024). Funmap: integrating high-dimensional functional annotations to improve fine-mapping. Bioinformatics 41 (1), btaf017. 10.1093/bioinformatics/btaf017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Lloyd-Jones L. R., Zeng J., Sidorenko J., Yengo L., Moser G., Kemper K. E., et al. (2019). Improved polygenic prediction by Bayesian multiple regression on summary statistics. Nat. Commun. 10 (1), 5086. 10.1038/s41467-019-12653-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Madsen T. E., Ding L., Khoury J. C., Haverbusch M., Woo D., Ferioli S., et al. (2024). Trends over time in stroke incidence by race in the greater Cincinnati northern Kentucky stroke study. Neurology 102 (3), e208077. 10.1212/wnl.0000000000208077 [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. The ALLHAT Officers and Coordinators for the ALLHAT Collaborative Research Group (2000). Major cardiovascular events in hypertensive patients randomized to doxazosin vs chlorthalidone: the antihypertensive and lipid-lowering treatment to prevent heart attack trial (ALLHAT). JAMA 283 (15), 1967–1975. 10.1001/jama.283.15.1967 [DOI] [PubMed] [Google Scholar]
  39. Malik R., Chauhan G., Traylor M., Sargurupremraj M., Okada Y., Mishra A., et al. (2018). Multiancestry genome-wide association study of 520,000 subjects identifies 32 loci associated with stroke and stroke subtypes. Nat. Genet. 50 (4), 524–537. 10.1038/s41588-018-0058-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Marston N. A., Pirruccello J. P., Melloni G. E. M., Koyama S., Kamanu F. K., Weng L. C., et al. (2023). Predictive utility of a coronary artery disease polygenic risk score in primary prevention. JAMA Cardiol. 8 (2), 130–137. 10.1001/jamacardio.2022.4466 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Martin A. R., Gignoux C. R., Walters R. K., Wojcik G. L., Neale B. M., Gravel S., et al. (2017). Human demographic history impacts genetic risk prediction across diverse populations. Am. J. Hum. Genet. 100 (4), 635–649. 10.1016/j.ajhg.2017.03.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Martin A. R., Kanai M., Kamatani Y., Okada Y., Neale B. M., Daly M. J. (2019). Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet. 51 (4), 584–591. 10.1038/s41588-019-0379-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Martin S. S., Aday A. W., Allen N. B., Almarzooq Z. I., Anderson C. A. M., Arora P., et al. (2025). 2025 heart diseases and stroke statistics: a report of US and global data from the American heart association. Circulation 151 (8), e41–e660. 10.1161/CIR.0000000000001303 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Mishra A., Malik R., Hachiya T., Jürgenson T., Namba S., Posner D. C., et al. (2022). Stroke genetics informs drug discovery and risk prediction across ancestries. Nature 611 (7934), 115–123. 10.1038/s41586-022-05165-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Moreno-Grau S., Vernekar M., Lopez-Pineda A., Mas-Montserrat D., Barrabés M., Quinto-Cortés C. D., et al. (2024). Polygenic risk score portability for common diseases across genetically diverse populations. Hum. Genomics 18 (1), 93. 10.1186/s40246-024-00664-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Neumann J. T., Riaz M., Bakshi A., Polekhina G., Thao L. T. P., Nelson M. R., et al. (2021). Predictive performance of a polygenic risk score for incident ischemic stroke in a healthy older population. Stroke 52 (9), 2882–2891. 10.1161/STROKEAHA.120.033670 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Owolabi M., Sarfo F., Howard V. J., Irvin M. R., Gebregziabher M., Akinyemi R., et al. (2017). Stroke in Indigenous Africans, African Americans, and European Americans: interplay of racial and geographic factors. Stroke 48 (5), 1169–1175. 10.1161/strokeaha.116.015937 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Palaniappan L. P., Allen N. B., Almarzooq Z. I., Anderson C. A., Arora P., Avery C. L., et al. (2026). 2026 heart disease and stroke statistics: a report of US and global data from the American heart association. Circulation 153 (9), e275–e906. 10.1161/cir.0000000000001412 [DOI] [PubMed] [Google Scholar]
  49. Patel A. P., Wang M., Ruan Y., Koyama S., Clarke S. L., Yang X., et al. (2023). A multi-ancestry polygenic risk score improves risk prediction for coronary artery disease. Nat. Med. 29 (7), 1793–1803. 10.1038/s41591-023-02429-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Price A. L., Patterson N. J., Plenge R. M., Weinblatt M. E., Shadick N. A., Reich D. (2006). Principal components analysis corrects for stratification in genome-wide association studies. Nat. Genet. 38 (8), 904–909. 10.1038/ng1847 [DOI] [PubMed] [Google Scholar]
  51. Ramirez A. H., Sulieman L., Schlueter D. J., Halvorson A., Qian J., Ratsimbazafy F., et al. (2022). The all of us research program: data quality, utility, and diversity. Patterns (N Y) 3 (8), 100570. 10.1016/j.patter.2022.100570 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Rayner N. W. M. M. I. (2011). Development and Use of a Pipeline to Generate Strand and Position Information for Common Genotyping Chips. [Google Scholar]
  53. Robin X., Turck N., Hainard A., Tiberti N., Lisacek F., Sanchez J. C., et al. (2011). pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinforma. 12, 77. 10.1186/1471-2105-12-77 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Ruan Y., Lin Y. F., Feng Y. C. A., Chen C. Y., Lam M., Guo Z., et al. (2022). Improving polygenic prediction in ancestrally diverse populations. Nat. Genet. 54 (5), 573–580. 10.1038/s41588-022-01054-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Strilciuc S., Alecsandra Grad D., Radu C., Chira D., Stan A., Ungureanu M., et al. (2021). The economic burden of stroke: a systematic review of cost of illness studies. J. Med. Life 14 (5), 606–619. 10.25122/jml-2021-0361 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Taylor T. N., Davis P. H., Torner J. C., Holmes J., Meyer J. W., Jacobson M. F. (1996). Lifetime cost of stroke in the United States. Stroke 27 (9), 1459–1466. 10.1161/01.str.27.9.1459 [DOI] [PubMed] [Google Scholar]
  57. Traylor M., Rutten-Jacobs L., Curtis C., Patel H., Breen G., Newhouse S., et al. (2017). Genetics of stroke in a UK African ancestry case-control study: south London ethnicity and stroke study. Neurol. Genet. 3 (2), e142. 10.1212/NXG.0000000000000142 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Van Calster B., Nieboer D., Vergouwe Y., De Cock B., Pencina M. J., Steyerberg E. W. (2016). A calibration hierarchy for risk models was defined: from utopia to empirical data. J. Clin. Epidemiol. 74, 167–176. 10.1016/j.jclinepi.2015.12.005 [DOI] [PubMed] [Google Scholar]
  59. Weissbrod O., Hormozdiari F., Benner C., Cui R., Ulirsch J., Gazal S., et al. (2020). Functionally informed fine-mapping and polygenic localization of complex trait heritability. Nat. Genet. 52 (12), 1355–1363. 10.1038/s41588-020-00735-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Wu H., Pérez-Rodríguez P., Boehnke M., Cui Y., Liang X., Vazquez A. I., et al. (2026). Improving polygenic score prediction for underrepresented groups through transfer learning. Nat. Commun. 17 (1), 1973. 10.1038/s41467-026-68696-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Zeng J., Visscher P. M. (2025). Harnessing functional annotation to improve the accuracy and transferability of polygenic scores. Nat. Rev. Genet. 26 (12), 805–806. 10.1038/s41576-025-00893-4 [DOI] [PubMed] [Google Scholar]
  62. Zhang J., Zhan J., Jin J., Ma C., Zhao R., O’Connell J., et al. (2024). An ensemble penalized regression method for multi-ancestry polygenic risk prediction. Nat. Commun. 15 (1), 3238. 10.1038/s41467-024-47357-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Zheng Z., Liu S., Sidorenko J., Wang Y., Lin T., Yengo L., et al. (2024). Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries. Nat. Genet. 56 (5), 767–777. 10.1038/s41588-024-01704-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Zhuang Y., Kim N. Y., Fritsche L. G., Mukherjee B., Lee S. (2024). Incorporating functional annotation with bilevel continuous shrinkage for polygenic risk prediction. BMC Bioinforma. 25 (1), 65. 10.1186/s12859-024-05664-2 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table1.xlsx (60.6KB, xlsx)

Data Availability Statement

The REGARDS (Study Accession: phs002719.v1.p1) and GenHAT (Study Accession: phs002716.v1.p1) phenotypic and genetic data used in this study are available through the National Center for Biotechnology Information (NCBI) database of Genotypes and Phenotypes (dbGaP) to approved researchers following applicable data access procedures. This study also used Controlled Tier data from the All of Us Research Program through an institutional Data Use and Registration Agreement (DURA) which allows UAB researchers who have completed required training, and approvals to access individual level data through the All of Us Researcher Workbench. Individual-level participant data cannot be made publicly available outside the platform. The REGARDS holdout sample used for validation analyses is supported by the National Institute on Aging (NIA) and the National Institute of Neurological Disorders and Stroke (NINDS) through the Center for Alzheimer’s and Related Dementias (CARD) and is made available through the AD Knowledge Portal (https://adknowledgeportal.synapse.org) in accordance with applicable data access policies.

Summary statistics from the GIGASTROKE genome-wide association study used for polygenic risk score development are publicly available through the NHGRI-EBI GWAS Catalog (ID: 36180795). Previously published polygenic scores evaluated in this study are publicly available through the Polygenic Score Catalog, including the Abraham et al. score (PGS Catalog ID: PGS000039) and the Mishra et al. score (PGS Catalog ID: PGS002724).The polygenic risk scores developed in this study are found under the publication ID PGP000839 (score IDs PGS019947-PGS019952). Scripts and code for PRS construction are available through GitHub (https://github.com/uabgenepi/stroke_prs/wiki).


Articles from Frontiers in Bioinformatics are provided here courtesy of Frontiers Media SA

RESOURCES