Skip to main content
Alzheimer's & Dementia logoLink to Alzheimer's & Dementia
. 2026 Jul 28;22(8):e71529. doi: 10.1002/alz.71529

Cross‐ancestry polygenic risk scores enhance Alzheimer's disease risk prediction in multiethnic cohorts

Meri Okorie 1,2, Caroline Jonson 1,3,4,5, Alexis P Oddi 1,3, Patricia A Castruita 1,3, Brian Fulton‐Howard 6,7, Kristine Yaffe 1,2,8,9,10, Jennifer S Yokoyama 1,3,11, Chinedu Udeh‐Momoh 12,13, Shea J Andrews 1,2,✉; for the Alzheimer's Disease Sequencing Project and the Healthy Aging Brain Study—Health Disparities
PMCID: PMC13415749  PMID: 42522059

Abstract

INTRODUCTION

Genome‐wide association studies (GWAS) have identified 80+ genetic loci associated with Alzheimer's disease (AD), enabling the development of polygenic risk scores (PRS). However, the predictive accuracy of PRS in diverse populations remains low. Here, we evaluated the predictive accuracy of single‐, multi‐, and cross‐ancestry AD‐PRS models across multi‐ancestral populations.

METHODS

We used AD GWAS summary statistics from European, African, Admixed American, and East Asian populations to construct AD‐PRS for each target population. Model performance was assessed by estimating odds ratios, R2, and area under the curve.

RESULTS

The cross‐ancestry Bayesian PRS model demonstrated the highest predictive performance in non‐European populations. It was significantly associated with poorer cognitive function, lower Aβ42 cerebrospinal fluid levels, and the most severe category of Aβ and tau neuropathological burden.

DISCUSSION

Inclusive genetic datasets and cross‐ancestry PRS models are needed to enhance the transportability of AD‐PRS across multi‐ancestral populations.

Keywords: AD‐PRS, Alzheimer's disease, polygenic risk scores, population genetics

Highlights

  • Single‐ancestry polygenic risk score (PRS) is only predictive in participants of European populations.

  • Cross‐ancestry PRS improves risk predictions in non‐European participants.

  • Cross‐ancestry PRS is associated with abnormal Aβ and tau pathology and cognitive decline

  • Cross‐ancestry PRS is associated with the Alzheimer's disease (AD) latent variable in a multi‐ancestral cohort.

1. BACKGROUND

The abundance of genome‐wide association studies (GWAS) and the growing number of disease‐associated genetic variants have enabled the development of polygenic risk scores (PRS). PRS calculate an individual's genetic liability for a given trait or a disease by aggregating all the variants weighted by their effect size estimates derived from GWAS summary statistics. 1 In some diseases, at‐risk individuals identified through PRS exhibit up to 20‐fold higher than the carrier frequency of rare monogenic mutations conferring comparable risk. 2 These findings highlight the potential of PRS to identify at‐risk individuals who may not be captured through family history or monogenic variants, creating opportunities to modify risk factors and personalize care before clinical symptoms emerge. 3

Alzheimer's disease (AD) GWAS have greatly advanced our understanding of AD genetic architecture, with heritability estimated to be 40‐60% 4 , 5 , 6 , 7 and facilitate the development of AD‐PRS. AD‐PRS show good predictive accuracy, with the best‐performing model having an area under the curve (AUC) of 0.74 while adjusting for apolipoprotein E (APOE) ε2 and ε4 status. 8 The predictive accuracy further increases to 81% when applied to neuropathologically confirmed cases and controls or when using extreme PRS cutoffs. 9 , 10 The high predictive accuracy of AD‐PRS demonstrates its potential utility for risk assessment. However, clinical translation has been slow, partly due to inconsistent performance across studies that raises concerns about generalizability. 11 , 12 , 13 , 14 , 15

A major driver of the lack of generalizability of AD‐PRS stems from the lack of diversity in GWAS training datasets. Most AD‐PRS models have been derived from European ancestry GWAS, given the scarcity of large studies in non‐European or admixed populations. 16 , 17 As a result, AD‐PRSs derived from European datasets show markedly reduced predictive accuracy when applied to non‐European ancestry populations, with performance declining as genetic distance from the training population increases. 16 , 17 , 18 For instance, AD‐PRS derived from African American GWAS datasets outperformed those based on European or multi‐ancestry AD‐PRS in the African cohort, with some variability across different regions. 13 Similarly, including variants identified specifically in African populations has been shown to improve the accuracy of PRS across ancestrally diverse and admixed populations. 15 In some cases, AD‐PRS derived from European GWAS data remains associated with AD risk in non‐European populations by careful tuning of pruning and thresholding (P + T) parameters; 18 however, such optimizations may not be practical for routine clinical use. These trends highlight the limited portability of European‐derived PRS and the potential to exacerbate existing health disparities if such models are applied indiscriminately. 16

The need to address these disparities is particularly urgent given the disproportionate burden of ADRD among older Black and Latinx individuals compared to non‐Latinx White individuals. 19 Yet, relatively few studies have systematically evaluated the performance of different AD‐PRS models across ancestrally diverse populations, largely due to the limited availability of training and validation datasets. By increasing GWAS sample sizes of historically underrepresented populations and matching the statistical power of European ancestry GWAS, the accuracy of PRS can improve. 15 , 20 , 21 , 22 , 23 , 24 Although current non‐European AD GWAS are notably smaller than those of European ancestry, multi‐ancestry PRS derived from smaller ancestry‐matched GWAS have been shown to outperform PRS derived from larger European GWAS. 25 Alternatively, cross‐ancestry PRS methods aim to jointly model GWAS summary statistics from multiple ancestry groups to improve PRS prediction within a specified target ancestry. 22 , 23

The present study evaluates the predictive accuracy of AD‐PRS computed using single‐, multi‐, and cross‐ancestry PRS models in multi‐ancestral populations. We expect that AD‐PRS will exhibit reduced predictive accuracy in populations not represented in the training data. Beyond disease diagnosis, we also assessed the clinical validity of PRS by examining whether the best‐performing model was associated with AD endophenotypes, including biomarkers and cognitive domains within the A/T/N framework. 26 , 27 This evaluation provides insight into whether PRS can capture early disease‐related processes, complementing risk stratification for AD.

2. METHODS

2.1. Datasets and quality control

2.1.1. GWAS datasets

We utilized summary statistics from the largest publicly available ancestry‐specific GWAS datasets on AD and related dementias (ADRD) to date, including participants of European (Bellenguez et al. 5 : 39,106 cases, 46,828 proxy cases, 401,577 controls; FinnGen Release 6: 7329 cases, 131,102 controls), African (Kunkle et al. 28 : 2748 cases, 5222 controls), East Asian (Shigemizu et al.: 3962 cases, 4074 controls 29 ), and Caribbean Hispanic (Columbia University Study: 1088 cases, 1152 controls) descent. We also utilized multi‐ancestry meta‐analysis (MAMA) GWAS summary statistics, combining data from the five ancestry‐specific studies listed above (Lake et al.: 54,233 AD cases, 46,828 proxy cases, 543,127 controls 30 ) to construct a multi‐ancestry PRS.

2.1.2. Alzheimer's Disease Sequencing Project

Data from the Alzheimer's Disease Sequencing Project (ADSP) Release 4 whole genome sequencing dataset were accessed from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS) (https://dss.niagads.org/). The participants in the ADSP are recruited from large, existing family‐based and case/control studies that provide comprehensive phenotypic and genetic data. The ADSP has made a concerted effort to include participants from multiracial and multiethnic groups, as many earlier genetic studies have predominantly involved individuals of European descent. AD diagnosis was harmonized across the cohorts by the ADSP Phenotype Harmonization Consortium (ADSP‐PHC). 31 The ADSP‐PHC also harmonized AD endophenotype measurements, including cognitive function (language, memory, executive functions; total N = 11,046), cerebrospinal fluid (CSF) biomarkers (Tau, pTau181, Aβ42; total N = 815), and neuropathology (Consortium to Establish a Registry for Alzheimer's Disease [CERAD], N = 2727; Thal Phase, N = 1186; BRAAK staging, N = 2725). 32 , 33 For exclusion, A total of 515 participants diagnosed with mild cognitive impairment (MCI) were excluded from the analysis. An additional 117 participants from a family‐based study that overlapped with case/control studies were removed. Additionally, 488 samples were removed due to missing genotype data, and 5763 participants were removed due to missing diagnosis information.

RESEARCH IN CONTEXT
  1. Systematic review: Using diverse genome‐wide association studies (GWAS) datasets to construct Alzheimer's disease (AD) polygenic risk score (PRS) is a promising yet underexplored approach to improve risk prediction accuracy across different populations. Integrating diverse base GWAS datasets and evaluating various PRS models can enhance model performance, but such approaches have not yet been widely applied to multi‐ancestral cohorts to measure AD risks or abnormalities in biomarkers.

  2. Interpretation: Incorporation of ancestrally diverse base GWAS datasets enhanced the association between PRS and AD risk across multiple populations. Leveraging both these diverse discovery datasets and a Bayesian framework markedly improved model performance, extending its potential clinical applicability beyond AD case‐control classification to the prediction of biomarker abnormalities.

  3. Future directions: Future research should prioritize the validation of cross‐ancestry PRS models in larger and more heterogeneous populations, alongside systematic benchmarking against an expanding repertoire of PRS methodologies. Clinical implementation of AD‐PRS will require rigorous validation in large, diverse as well as community‐based populations to ensure reproducibility and generalizability, thereby enhancing its translational relevance.

A standardized quality control QC pipeline was applied to the whole genome sequencing (WGS) datasets using default settings in GenoTools. 34 Variants failing to meet predefined quality thresholds were excluded, including those with a minimum depth of coverage below 10, genotype quality below 20, and minor allele frequency (MAF) below 0.01. Variants with missing rsIDs were assigned a chromosome position ID. To further assess quality, sample‐level and variant‐level QC were performed using GenoTools to undergo a series of QC checks. Sample‐level QC includes call rate (≥ 0.95), sex concordance (male: F ≥ 0.99, female: F ≥ 0.03), and relatedness between samples (cutoff = 0.0884). Variant‐level QC was performed to examine non‐random missingness by haplotypes, Hardy–Weinberg equilibrium (≤ 1 x 10−4), and genotype missingness (< 0.05). 2743 participants were removed due to sex discrepancy, 359 participants were removed due to heterozygosity pruning, and 3164 participants were removed due to the relatedness check. The African GWAS summary statistics were derived from participants in the Alzheimer's Disease Genetic Consortium (ADGC), which includes sample overlaps with those from the ADSP. To retain sample independence between the base and the target datasets and avoid PRS inflation caused by sample overlaps, 35 relationship inference was performed using software, KING. 36 to identify and exclude samples based on relatedness. Pairwise samples with kinship coefficients ≥ 0.354, 36 2776 participants, were removed from the ADSP target datasets. After the main GenoTools filters and excluding participants with missing data, 19,398 participants (7111 cases and 12,287 controls) remained for the downstream analysis. Additionally, WHICAP participants (N = 442) were excluded from the ADSP to remove potential overlapping participants with the Caribbean Hispanic GWAS. The Stage I Bellenguez et al. GWAS includes European Alzheimer & Dementia Biobank, European Association of Development Research and Training Institutes, GERAD, Bonn, RS1&2, Genome Research at Fundació Ace, DemGene, the CCHS study, NxC, UK Biobank, which do not overlap with the ADSP cohorts.

Genetic ancestry was determined using a principal component analysis (PCA) with the 1000 Genome Project (1KG) + Human Genome Diversity Project (HGDP) as the reference panel, followed by classification via a random forest algorithm implemented in pgsc_calc. 37 Out of 19,398 participants after initial exclusion (eMethods), 8043 (41.5%) participants were classified as 1KG + HGDP‐EUR‐like (EUR; European), 1945 (10.0%) as 1KG + HGDP‐AFR‐like (AFR; African), 6901 (35.6%) as 1KG + HGDP‐AMR‐like (AMR; Admixed American), 2365 (12.2%) as 1KG + HGDP‐CSA‐like (CSA; Central/South Asian), 72 (0.37%) as 1KG + HGDP‐EAS‐like (EAS; East Asian), and 72 (0.37%) participants were assigned as 1KG + HGDP‐MID‐like (MID; Middle Eastern) ancestry. Participants of EAS and MID were removed due to small sample size, while participants of CSA were excluded due to imbalanced case‐control ratios (18 cases: 2347 controls). Global genetic admixture for each sample was estimated using ADMIXTURE version 1.3. 38 Briefly, we identified a representative reference panel composed of samples distributed by the 1KG. The reference panel and target ADSP dataset were subset to include only overlapping single‐nucleotide polymorphisms (SNPs). We performed an unsupervised admixture analysis of the reference dataset, setting K = 5. The target samples were then projected onto the population structure learned from the reference panel.

2.2. Health and Brain Aging Study—Health Disparities

2.2.1. Participants

The Health and Aging Brain Study—Health Disparities (HABS‐HD) cohort comprises participants from Black, Mexican American, and non‐Hispanic White populations recruited at the University of North Texas Health Science Center, Fort Worth, Texas, USA. 39 , 40 The study collected information on demographics, AD biomarkers, neuroimaging, clinical history, and genomics, providing a comprehensive resource for AD and related dementia research. Eligibility to participate in the study was restricted to generally healthy individuals without serious mental or medical conditions, who were willing to undergo sample collections and neuroimaging procedures. We utilized participant data from the HABS‐HD release 5 baseline visit 1. We further applied the exclusion criteria to remove participants younger than 55 years old, those without genetic data, and those with plasma biomarker values identified as outliers. Outliers were detected following the published methodology using the interquartile range approach. 41 All HABS‐HD participants (and/or their legal guardians) signed written informed consent to participate in the study.

2.2.2. AD endophenotypes

HABS‐HD participants have undergone extensive phenotyping for AD endophenotypes, including plasma A/T/N biomarkers, brain morphometry, amyloid and tau positron emission tomography (PET), and neuropsychological testing, as previously described. 39 , 40 For downstream analyses of each endophenotype, we applied the following approaches:

Plasma biomarkers: Fasting blood samples were collected, and the levels of plasma biomarkers (i.e., Aβ42, Aβ40, pTau217, total Tau, NfL) were quantified using commercially available kits, Quanterix, for all the participants of the HABS‐HD. Aβ40, Aβ42, pTau217, Tau, and NfL values were natural log‐transformed to reduce skewness, and then standardized to z‐scores (mean = 0, SD = 1) across the study samples.

Brain morphometry: Cortical thickness was defined as the surface area–weighted average of cortical thickness in the right and left entorhinal cortex, fusiform, and inferior and middle temporal cortex. Hippocampal volumes were determined by taking the mean volume of the right and left hippocampal volumes determined on a T1‐weighted volume scan.

Neuropsychological testing: Composite scores for each cognitive domain (memory, language, and executive function) were constructed by averaging z‐score‐normalized test scores from the neuropsychological battery. The memory domain was assessed using immediate and delayed recall from the Wechsler Memory Scale (WMS‐III) Logical Memory and the Spanish English Verbal Learning Test (SEVLT). The language domain was assessed using Letter Fluency and Animal Naming tests. The executive function domain was assessed using the WMS‐III Digit Span and the Trail Making Test, Parts A and B. In addition, the Mini‐Mental State Examination (MMSE) and Clinical Dementia Rating (CDR) scale were administered to all participants as part of the neuropsychological assessment.

Amyloid PET: For amyloid PET scan variables, the global standardized uptake value ratio (SUVR) was derived by normalizing tracer uptake to the whole cerebellum. Amyloid PET positivity was defined using an SUVR threshold of 1.08.

2.2.3. Genome‐wide genotyping and QC

Participants were genotyped on the Illumina Global Screening Array (GSA), and genotype data underwent stringent QC checks using an in‐house Snakemake Pipeline. 42 Variants were excluded if the call rate was < 0.95 and not in Hardy–Weinberg equilibrium (p < 1 × 10−6), and samples were excluded if the call rate was < 0.95; discordant sex was reported based on X chromosome heterozygosity, excessive/insufficient heterozygosity, and cryptic relatedness. Related individuals were determined within and across cohorts by identity‐by‐descent (IBD) using KING. 36 with individuals excluded based on a proportion of IBD < 0.1875, corresponding to less than halfway between second‐ and third‐degree relatives. Single‐nucleotide polymorphisms that were not directly genotyped were imputed on the TOPMed Imputation Server. 43 Ancestry groups were imputed separately using all ethnicities of the TOPMed reference panel (Version R3) to allow ancestry‐specific imputation quality (R2) estimation, with Eagle used for phasing and Minimac3 used for imputation. Rather than imputing the joint dataset, ancestry‐specific imputation was performed to preserve accuracy and reduce poorly imputed variants in individual ancestry groups. 44 Following imputation, poorly imputed (R2 < 0.3) or rare (MAF < 0.01) variants were removed, and the ancestry groups were merged for joint analysis. Following this merger, variants with low call rate due to differential imputation (< 95%) were removed, and then samples with low call rates (< 95%) were removed. To note, different genetic QC software were used for the ADSP and the HABS‐HD to accommodate differences in genotyping platforms and data types; accordingly, platform‐specific and widely accepted QC thresholds were applied for each cohort. These differences reflect standard best practices for WGS versus array‐based data and are not expected to materially affect downstream analyses. The final dataset contained 2559 samples with ∼22 M high‐quality SNPs per individual.

Of the total of 2559 participants that were included in the downstream statistical analysis, 1040 were of European ancestry, 891 were of Admixed American ancestry, and 574 were of African ancestry based on genetic ancestry inference (Refer to ancestry assignment in 2.3.2).

2.3. PRS construction:

2.3.1. PRS scorefile generation

PRSice was used to implement P + T to construct four AD‐PRS, stratified by EUR ancestry and MAMA ancestry, at two p‐value thresholds. P + T generates scorefiles for SNPs that exceed a given p‐value threshold (PT ) and its corresponding effect size estimated from an independent GWAS. Scorefiles for the single‐ (EUR) and multi‐ancestry (MAMA) PRS models were constructed using SNPs with a p‐value < 0.1 (P T0.1) and genome‐wide significance (GWS) (P TGWS). SNPs from APOE regions ± 500 kb (GRCh38, chr19:bp44405791‐45409393) were excluded from all the base datasets. A p‐value threshold of 0.1 in the construction of PRS has been previously shown to have the best predictive accuracy and will be tested in this study. 8 Linkage disequilibrium (LD) clumping using 1KG + HGDP LD reference panels, with an r2 threshold of 0.1 and clumping distance of 250 kb, was used to select independent SNPs. For the single‐ancestry PRS model, the 1KG + HGDP reference panel was filtered for European‐only samples, and the whole 1KG + HGDP reference panel was used for the multi‐ancestry PRS model.

PRS‐CSx was used to construct a scorefile for the cross‐ancestry AD‐PRS model, a method that improves cross‐population polygenic prediction by jointly modeling GWAS summary statistics from multiple populations. 23 The global shrinkage parameter (phi) was set to “auto”, which allows for automatic estimation of the effect size from the datasets. Additionally, the “meta” parameter was set to “true” to combine the SNP effect sizes across populations using an inverse‐variance‐weighted meta‐analysis. For PRS‐CSx, SNPs were restricted to the HapMap3 panel as the software was designed for HapMap3 variants, whereas PRSice obtains independent SNPs using a P + T approach. Both software used the 1KG reference panel to maintain the LD estimation consistent between the two software. The summary statistics for ancestry‐specific AD GWAS were used to construct the scorefile. SNPs from APOE regions ± 500 kb (GRCh37, chr19:bp44909053‐45912650) were excluded from the final scorefile prior to PRS calculation. Liftover was automatically applied during PRS calculation to convert scorefile build to GRCh38. The same methodology was applied to the HABS‐HD cohort with the same parameters and base GWAS datasets.

2.3.2. PRS calculation and ancestry normalization

To mitigate ancestry‐related confounding due to differences in PRS distribution and to enable more accurate PRS comparisons across ancestrally diverse populations, 45 we used pgsc_calc 37 to construct PRS and normalize for ancestry differences in the ADSP and the HABS‐HD cohort using the PCA method. 46 PCA was performed on the 1KG reference panel to derive principal components (PCs) representing global genetic variation. Target individuals were then projected onto these PCs to determine their position in genetic ancestry space. A RandomForest classifier, trained on the reference panel's PCA loadings (default: 10 PCs), was used to predict the reference population to which each target sample was most similar. 47

Following PC projection and ancestry assignment, PRS were normalized by ancestry. First, a regression model predicted PRS from PC loadings, and the residuals were used to center the PRS distribution at zero across ancestry groups (znorm1). A second regression on the squared residuals adjusted for variance, standardizing the PRS spread to approximately one (znorm2). All five PRS models underwent ancestry normalization, and znorm2 scores were used for the downstream analysis. Importantly, phenotypic information (e.g., case–control status) was not used in the normalization process. This approach removes PRS variation attributable to population stratification driven by allele frequency and LD shifts while retaining the GWAS‐derived SNP weights and rankings of individual risks within ancestry groups. This normalization step scales PRS to enable inter‐ancestry comparisons, while still capturing disease‐relevant genetic risks.

This study was conducted and reported in accordance with the PRS reporting guidelines; a completed checklist is provided in supplementary files.

2.4. Statistical analysis

2.4.1. AD risk

Logistic regression models were used to evaluate the associations of EUR, MAMA, and PRS‐CSx PRS models with AD, adjusting for age at baseline, sex, APOE genotypes, and 10 PCs. The APOE status was encoded as a categorical variable, categorized as APOE ε2+ carriers (ε2/ε2, ε2/ε3) and ε4+ carriers (ε2/ε4, ε3/ε4, or ε4/ε4), with the ε3/ε3 allele serving as the reference group. Model performance was evaluated via area under the receiver operating characteristics curves (AUCROC) to assess model discrimination. R2redux 48 was used to calculate and statistically compare the differences in R2 for the AD outcome between models. Calibration of AD‐PRS was assessed to evaluate the estimated probability of AD against the observed probability using PRS as a predictor.

2.4.2. Endophenotypes

In the ADSP, multiple linear regression models were fit to assess the relationship between PRS‐CSx and (i) cerebrospinal fluid biomarkers (i.e., pTau181, total tau, and Aβ42), and (ii) cognitive domain scores (i.e., language, memory, and executive function). Models were adjusted for age, sex, APOE genotype, PC1‐10, and study cohort. Models for cognitive outcomes were additionally adjusted for years of education. Ordinal logistic regression models, performed using the polr function in the MASS R package, were fitted for Thal phase, CERAD score, and Braak staging as ordinal outcomes, adjusting for age, sex, APOE genotype, PC1‐10, and study cohort. All p‐values were adjusted for multiple testing at 5% false discovery rate (FDR). Genetic ancestry stratification was performed for cognitive outcomes. CSF biomarkers and other cognitive measures had insufficient sample sizes for ancestry‐stratified analyses.

Similarly, in the HABS‐HD, multiple linear regression models were fit to assess the associations between PRS‐CSx and (i) plasma biomarkers (i.e., Aβ40, Aβ42, pTau217, Tau, and NfL), (ii) neuroimaging variables (i.e., cortical thickness and hippocampal volume), (iii) composite scores for each domain (i.e., language, memory, and executive function), and (iv) global cognition tests (i.e., CDR and MMSE). Models were adjusted for age, sex, APOE genotype, and PC1‐4. Additionally, body mass index (BMI) and estimated glomerular filtration rate (eGFR) were accounted for in the plasma biomarkers. Intracranial volume was accounted for in the neuroimaging variables. Years of education were accounted for in the neuropsychological test battery variables. Both raw and FDR‐adjusted p‐values (5%) were reported.

Continuous and binary outcomes were modeled using linear and logistic regressions, respectively. PRS effect estimates are interpreted as the change in the outcome per one SD increase in PRS. For ordinal outcomes, ordinal logistic regression was performed, and predicted probabilities represent the estimated likelihood of being in each severity category based on PRS values.

3. RESULTS

3.1. Patient characteristics

Of 16,461 participants in the ADSP, 9645 were classified as cognitively normal, and 6816 were diagnosed with AD (Table 1). The cohort had a mean age of 68 ± 10 years and consisted of 60% females. EUR ancestry was the most common (49%), followed by AMR (40%) and AFR participants (11%), respectively. The majority of people with AD were APOE ε4 allele carriers (54%). A similar profile is presented in the HABS‐HD, with 61% female, with a mean age of 69 ± 8 years. The cohort primarily consists of EUR (41%), AMR (35%), and AFR participants (22%) (Table 1, Table S1, Table S2).

TABLE 1.

Characteristics of the ADSP and the HABS‐HD cohort, including age, sex, genetic ancestry, APOE genotype, and endophenotypes

Cohort ADSP HABS‐HD
Characteristic Cognitively normal AD Cognitively normal MCI + dementia
Sample size a 9645 6816 1855 704
Age, years 68 (11) 73 (9) 67 (7.3) 68 (8.5)
Sex
Female 6088 (63%) 3986 (58%) 1194 (64%) 371 (53%)
Male 3557 (37%) 2830 (42%) 661 (36%) 333 (47%)
Ancestry
European 3637 (30%) 4406 (62%) 845 (47%) 195 (28%)
African 1377 (11%) 470 (6.8%) 330 (18%) 244 (35%)
Admixed American 4631 (38%) 1941 (28%) 637 (35%) 254 (37%)
APOE genotype
e3/e3 6195 (64%) 2843 (42%) 1019 (65%) 317 (56%)
e2+ 932 (9.7%) 282 (4.1%) 165 (11%) 65 (11%)
e4+ 2518 (26%) 3691 (54%) 390 (25%) 189 (33%)
Unknown – – 281 133
CSF Biomarkers
Aβ42 0.32 (0.97) −0.64 (0.87) – –
pTau −0.25 (0.86) 0.62 (0.95) – –
Tau −0.25 (0.90) 0.57 (0.91) – –
Plasma biomarkers
Aβ42/Aβ40 – – 0.02 (0.99) −0.02 (1.02)
pTau – – −0.10 (0.90) 0.26 (1.19)
Tau – – −0.02 (0.97) 0.07 (1.08)
Neuroimaging
Cortical thickness – – 2.75 (0.13) 2.69 (0.17)
Hippocampal volume – – 0.12 (0.92) −0.32 (1.13)
PET b
Aβ SUVR – – 1.04 (0.15) 1.37 (0.24)
Aβ positivity – – 195 (19%) 18 (90%)
Cognitive domains
Memory 0.67 (0.53) −0.66 (0.82) 0.33 (0.67) −0.85 (0.78)
Language 0.43 (0.70) −0.32 (0.80) 0.24 (0.75) −0.62 (0.84)
Executive function 0.45 (0.82) −0.41 (0.88) −0.05 (0.48) 0.12 (0.72)
Global cognition
MMSE – – 28.20 (2.11) 25.37 (4.33)
CDR – – 0.00 (0.00) 2.07 (2.17)
Neuropathology c
Thal phase
None 46 (29%) 21 (3%) – –
Phase 1 33 (21%) 10 (1%) – –
Phase 2 15 (9.4%) 16 (2%) – –
Phase 3 32 (20%) 54 (8%) – –
Phase 4 24 (15%) 121 (17%) – –
Phase 5 9 (6%) 475 (68%) – –
CERAD score
None 377 (58%) 47 (3%) – –
Sparse/possible 167 (26%) 65 (4%) – –
Moderate/probable 71 (11%) 320 (21%) – –
Definite/frequent 30 (5%) 1086 (72%) – –
BRAAK stage
None 31 (5%) 8 (0.5%) – –
Stage I 117 (18%) 17 (1%) – –
Stage II 179 (28%) 37 (2%) – –
Stage III 182 (28%) 80 (5%) – –
Stage IV 110 (17%) 188(12%) – –
Stage V 22 (3%) 458 (30%) – –
Stage VI 3 (0.5%) 732 (48%) – –

Abbreviations: AD, Alzheimer's disease; ADSP, Alzheimer's Disease Sequencing Project; APOE, apolipoprotein E; CDR, Clinical Dementia Rating; CERAD, Consortium to Establish a Registry for Alzheimer's Disease; HABS‐HD, Health and Aging Brain Study—Health Disparities; MMSE, Mini‐Mental State Examination; PET, positron emission tomography; SUVR, standardized uptake value ratio.

a

Mean (SD); n (%).

b

N = 1029; Aβ PET positivity from available PET data; % by diagnosis.

c

Available sample sizes; CERAD: N = 2713, Thal Phase: N = 1175, BRAAK staging: N = 2711.

The PCA projection of the ADSP revealed a high admixture profile among AMR participants. Most of these participants clustered into two distinct patterns, separated from the EUR group, while some overlapped with the AFR group (Figure 1(B), Figure S1). Consistent with this, the admixture analysis showed that AMR participants exhibited genetic heterogeneity, with contributions from EUR, Amerindian, and AFR ancestries (Figure 1(C)). To account for ancestry‐related differences in genetic architecture, we examined the distribution of all the PRS models across ancestral groups. Following ancestry normalization, the distributions were well‐aligned and approximately normally distributed across groups, confirming that ancestral confounding was mitigated (Figure S2).

FIGURE 1.

FIGURE 1

Genetic ancestry and global admixture analysis in the ADSP cohort. (A) PCA of genetic ancestry using the 1000 Genomes (1KG) reference panel. ADSP participants (black dots) are overlaid onto 1KG reference populations (square dots). (B) PCA of ADSP participants only. (C) Global ancestry proportions estimated using admixture analysis, illustrating varying degrees of admixture across ancestry groups. Same ancestry labels were used for all analyses. ADSP, Alzheimer's Disease Sequencing Project; PCA, principal component analysis.

3.2. EUR single‐ancestry PRS

The European‐only, single‐ancestry PRS model showed robust associations with AD risk in EUR participants at both P TGWS (odds ratio [OR] [95% confidence interval {CI}] = 1.4[1.37–1.52], p = 5.53 × 10−42) and P T0.1 (OR[95%CI] = 1.38[1.31–1.47], p = 1.69 × 10−27). R2 values also declined as the genetic distance for AFR and AMR participants increased from the EUR participants (R2 = 0.013, 0.018, 0.026, respectively) (Table S3). Association of single‐ancestry PRS remained statistically significant in the AMR participants at both PTGWS (OR[95%CI] = 1.17[1.10–1.24], p = 5.59 × 10−7) and P T0.1 (OR[95%CI] = 1.14[1.06–1.21], p = 1.69 × 10−4) with similar effect sizes. It showed no association with AD risk in the AFR participants, both at P TGWS (OR[95%CI] = 1.08[0.96–1.21], p = 0.22) and P T0.1 (OR[95%CI] = 0.93[0.81–1.04], p = 0.22) thresholds (Figure 2; Table S4).

FIGURE 2.

FIGURE 2

(A) Odds ratio of ancestry‐normalized single‐ (EUR), multi‐ (MAMA), and cross‐ancestry (PRS‐CSx) PRS models (right x‐axis) stratified by genetic ancestry (top y‐axis). p‑Values were adjusted for multiple testing using the Benjamini–Hochberg FDR procedure, controlling the FDR at 5%. (B) Comparison of model performance across ancestry groups. (Left) R2 for EUR PRS (GWS), MAMA PRS (GWS), and PRS‐CSx PRS models in the participants of EUR, AMR, and AFR ancestry groups. Error bars represent 95% confidence intervals. (Right) Differences in R2 between EUR versus MAMA PRS, EUR versus PRS‐CSx PRS, and MAMA versus PRS‐CSx PRS models. Error bars represent 95% confidence intervals of the difference. AFR, African; AMR, Admixed American; EUR, European; FDR, false discovery rate; PRS, polygenic risk scores.

3.3. MAMA multi‐ancestry PRS

When leveraging the MAMA GWAS summary statistics, multi‐ancestry PRS performance improved substantially in the AFR and AMR participants while remaining comparable in the EUR participants. In the EUR participants, the multi‐ancestry PRS showed similar associations as EUR PRS with AD risk for P TGWS (OR[95%CI] = 1.44[1.37–1.52], p = 5.53 × 10−42) and P T0.1 (OR[95%CI] = 1.38[1.31–1.47], p = 1.69 × 10−27). PT0.1 model showed a stronger association with AD risk than P TGWS model in the AMR participants (P TGWS (OR[95%CI] = 1.17[1.10–1.24], p = 1.27 × 10−6) and P T0.1 (OR[95%CI] = 1.34[1.26–1.41], p = 1.71 × 10−23). The effect was even larger in the AFR participants with P T0.1 model, demonstrating 80% increase in the odds of developing AD, while P TGWS model remained nonsignificant (P TGWS (OR[95%CI] = 1.13[1.00–1.27], p = 0.053) and P T0.1 (OR[95%CI] = 1.80[1.62–2.01], p = 7.30 × 10−27) (Figure 2; Table S4).

3.4. PRS‐CSx cross‐ancestry PRS

The PRS‐CSx cross‐ancestry model showed similar trends as PT0,1 MAMA PRS models in regards to the association with AD. It was associated with the highest AD risk in the AFR participants (OR[95%CI] = 1.71[1.52–1.93], p = 1.45 × 10−18), while remaining significant in the AMR and EUR participants (OR[95%CI] = 1.29[1.21–1.37], p = 1.76 × 10−15 and OR[95%CI] = 1.31[1.24–1.38], p = 7.86 × 10−22, respectively) (Figure 2). The model also achieved the highest R2 values in the AFR and AMR participants (R2 = 0.047, 0.010, respectively) compared to other models (Figure 2; Table S3). Differences in R2 were also significant between PRS‐CSx and EUR or MAMA PRS models in the AFR participants (ΔR2 p = 4.6 × 10−7 and 1.3 × 10−6, respectively), showcasing PRS‐CSx strong predictive performance. A less prominent yet similar trend was seen in the AMR participants (ΔR2 p = 0.016, 0.023, EUR–PRS‐CSx, MAMA–PRS‐CSx) (Table S3).

Calibration analyses further illustrate that while all models calibrate best in the EUR participants, and maintain high accuracy in the AMR participants, predictive accuracy decays progressively with increasing genetic distance from Europeans, with the greatest calibration gains in AFR participants arising from ancestry‐normalized MAMA and PRS‐CSx PRS models (Figures S3, S4).

3.5. AD endophenotypes: ADSP

To measure PRS performance beyond their established utility in predicting AD risk, we further evaluated the association of the best‐predictive PRS model, PRS‐CSx, with cognitive domains (memory, language, and executive function) and CSF biomarkers (Tau, pTau181, and Aβ42). PRS‐CSx demonstrated strong associations with poorer cognitive performance across multiple domains (β[95% CI]: Memory = 0.94[0.93–0.95], p = 6.69 × 10−18, Language = 0.95[0.94–0.97], p = 1.10 × 10−12, Executive function = 0.95[0.94–0.97], p = 1.18 × 10−9) in the total cohort (Table S5).

PRS‐CSx demonstrated significant associations with poorer cognitive functions in participants of European ancestry (Memory (β[95% CI] = ‐0.08[−0.099, −0.062], p = 3.71 × 10−17); Language (β[95% CI] = −0.059[−0.077, −0.042], p = 2.87 × 10−10); and Executive function (β[95% CI] = ‐0.056[−0.075, −0.037], p = 1.22 × 10−8). In AFR and AMR participants, PRS‐CSx was associated with memory language, respectively. Incremental R2 values were marginal, ranging from 1.8 × 10−4 to 8.7 × 10−3 in AFR and 3.3 × 10−4 to 1.6 × 10−3 in AMR participants across the domains (Table S5).

Similarly, PRS‐CSx was associated with lower Aβ42 CSF levels (β[95% CI] = 0.91[0.85–0.98], p = 0.015) (Figure 3), but was not associated with Tau or pTau181. Lower CSF levels of Aβ42 indicate a worse endophenotype for AD. Furthermore, higher PRS‐CSx values were associated with an increased probability of being in the most severe category of CERAD, Thal, and BRAAK classification (e.g., “Frequent/Definite” for CERAD, “Phase 5” for Thal, and “Stage VI” for Braak), suggesting a strong relationship between PRS‐CSx and advanced AD pathology (Figure 3; Table S5).

FIGURE 3.

FIGURE 3

Associations of the best‐predictive PRS model, cross‐ancestry PRS‐CSx, with (A) Cognitive function (N = 10,903; AFR: 409, AMR: 2,888, EUR: 7,606) and (B) CSF biomarkers (N = 808) in the ADSP. Darker colors indicate statistical significance after correction at a 5% FDR. Pooled analyses were performed for CSF biomarkers due to limited sample sizes precluding ancestry‐stratified analyses. (C) Predicted probabilities of neuropathological staging outcomes by PRS‐CSx. CERAD: N = 2713, Thal Phase: N = 1175, BRAAK staging: N = 2711. The sample sizes reported in Table 1 vary, as certain participants contributing to the endophenotype analyses lack an AD diagnosis. AFR, African; AMR, Admixed American; CERAD, Consortium to Establish a Registry for Alzheimer's Disease; CSF, cerebrospinal fluid; EUR, European; FDR, false discovery rate; PRS, polygenic risk scores.

3.6. AD endophenotypes: HABS‐HD

We further validated the association of the cross‐ancestry PRS in HABS‐HD using AD endophenotypes. Higher PRS‐CSx was significantly associated with lower levels of plasma levels of Aβ42, higher levels of pTau217, and higher Aβ PET global SUVR in the total cohort. The association with pTau217 and Aβ global SUVR was significant in the participants of European ancestry, retaining significance after multiple testing correction. There were no other ancestry‐specific effects in the AFR and AMR participants (Table S6).

4. DISCUSSION

We evaluated the predictive accuracy of single‐, multi‐, and cross‐ancestry AD‐PRS models in participants of African, Admixed American, and European ancestry within the ADSP and the HABS‐HD cohort to assess PRS transportability. Single‐ancestry PRS showed the strongest performance in participants of European ancestry but declined notably in participants of Admixed American and African ancestry, especially among individuals of African ancestry, consistent with previous reports highlighting ancestry‐specific limitations in PRS‐based risk stratification. 8 , 49 , 50 , 51 This highlights reduced accuracy and portability of a single‐ancestry PRS model across ancestries, likely driven by differences in LD structure, allele frequency, and genetic architecture. 52 , 53 , 54 , 55 , 56 In contrast, multi‐ and cross‐ancestry PRS models demonstrated improved predictive accuracy in African and Admixed American ancestry groups. When models incorporated a large number of SNPs from multi‐ancestry GWAS (e.g., MAMA PT0.1, PRS‐CSx), both associations with AD risk and explained variance (R2) improved, indicating that including more ancestry‐matched SNPs captures better risk profiles in African and Admixed populations. The absence of this trend for single‐ancestry EUR PRS likely reflects the limited portability of European‐derived variants to non‐European populations with more variable genetic architectures. This supports previous findings that PRS constructed from smaller, ancestry‐matched GWAS can outperform those derived from larger European datasets. 25

Although the non‐European GWAS datasets used in this study were substantially smaller than the European GWAS dataset, incorporating base datasets that match the ancestry of the target populations improved overall PRS performance. In contrast, one study reported similar performance across single‐, multi‐, and cross‐ancestry AD‐PRS models using the MVP data. 57 This discrepancy may be due to differences in SNP selections. For example, their single‐ancestry PRS included all GWS SNPs, whereas our approach applied LD pruning to select independent SNPs. As a result, incorporating additional SNPs from multi‐ancestry GWAS in their framework may have contributed only marginal effects. Differences in cohort characteristics may also play a role, as AD cases and controls in MVP were defined using International Classification of Diseases (ICD) codes rather than clinical or biomarker‐confirmed diagnoses.

Beyond AD diagnosis, we evaluated the clinical validity of PRS in predicting abnormalities in AD endophenotypes for Aβ, Tau, and cognitive function. The best‐performing model, PRS‐CSx, was associated with CSF Aβ42 levels and cognitive domains of memory, executive function, and language. Notably, PRS‐CSx was not associated with CSF pTau181 or total tau in the ADSP cohort, which may stem from sampling variability, data harmonization differences, and/or small sample sizes. Higher levels of plasma pTau217 were significant in the HABSHD cohort, in addition to associations with lowered levels of plasma Aβ42. Significant associations with key plasma biomarkers of AD pathology suggest that polygenic risk prediction may capture upstream pathogenic processes, including Aβ and tau pathways. Higher PRS‐CSx was also associated with the most severe neuropathology burdens across CERAD, Thal phase, and Braak stage in the ADSP cohort. In contrast, associations with lower or intermediate stages were weak or absent, suggesting that PRS‐CSx predicts advanced pathology but may be less sensitive to earlier or moderate changes. Although we observed significant associations between AD‐PRS and endophenotypes, the effect sizes of PRS still remain modest, explaining only a small fraction of variance in cognitive, neuroimaging, and biomarker outcomes, especially when looking at incremental R2. This is consistent with previous studies that, while PRS captures cumulative genetic effects, other factors such as age, sex, APOE genotype often explain more variance in AD phenotypes than PRS alone. 8 , 58 However, individuals in the extreme tail of the PRS distribution exhibit substantially elevated disease risk, reportedly as high as a 30‐fold increase among those in the top decile. 2 , 57 By defining thresholds that identify this high‐risk group, PRS can still be clinically meaningful for identifying individuals who may benefit from further testing and more regular screenings.

Ancestry normalization emerged as a critical step in improving prediction accuracy for underrepresented groups by accounting for population‐specific genetic structures that differ in allele frequency and LD. Although several PRS methods have been developed to include ancestry stratification or ancestry‐specific SNP effects to address bias in cross‐population prediction, 22 , 23 , 30 , 59 , 60 , 61 an explicit normalization of the PRS itself based on global ancestry is still not widely implemented for AD‐PRS. Shifts in PRS distributions that arise from differences in MAF and LD complicate the assessment of risks across populations and studies. By employing a reproducible framework such as the pgsc_calc Nextflow pipeline, genetic ancestry estimation and ancestry normalization can be performed effectively and seamlessly. Integrating ancestry normalization into PRS construction enables more equitable genetic prediction tools that better reflect the genetic architecture of ancestrally diverse populations. Given the high degree of genetic heterogeneity within a single superpopulation 62 (e.g., African and admixed populations), ancestry normalization is a critical step to account for differential risk profiles that arise from ancestral differences rather than true genetic risk.

Several limitations should be acknowledged in our study. First, a lack of representation of people of South and East Asian ancestry limited our ability to assess the cross‐ancestry transportability of the PRS models. The dataset is also limited to participants in the United States, and how multi‐ and cross‐ancestry perform in non‐European populations from other continents is unknown. Furthermore, the currently available GWAS summary statistics from minority populations still don't reflect the full genetic diversity and substructures that exist within one superpopulation. This calls for the necessities to continue expanding and improving the diversity and sample sizes of future GWAS to more fully represent global and intra‐ancestry genetic variation. Second, we did not incorporate genotype–environment (GxE) interactions, a framework that explains the interplay of environmental, socioeconomic, and lifestyle factors that can modify genetic risk. These factors are known to contribute additively or multiplicatively to AD susceptibility across populations. 63 , 64 , 65 Third, future iterations of such studies should prioritize larger, more representative cohorts and integrate extensive phenotypic data from biobanks (e.g., All of Us, UK biobanks) to enhance both the calibration and clinical adaptability of risk predictions. Finally, the application of other advanced data‐driven models, such as LDpred2, 66 PolyPred+, 69 and BridgPRS, 67 and the incorporation of SNP heritability estimates could further refine the predictive power and cross‐population transferability of AD‐PRS. 68 With the rapid proliferation of PRS methods, 69 future studies may focus on benchmarking across all the methods in large, diverse cohorts as well as community‐based populations.

Despite these limitations, our study demonstrates several key strengths. We systematically evaluated single‐, multi‐, and cross‐ancestry PRS models across participants of African, Admixed American, and European ancestries, providing a comprehensive assessment of PRS transferability in diverse cohorts. By incorporating ancestry‐specific normalization and leveraging multi‐ancestry GWAS summary statistics, we improved predictive accuracy in underrepresented populations and highlighted the importance of including ancestry‐matched SNPs in risk models. Furthermore, by extending the evaluation beyond AD diagnosis to biomarkers and cognitive endophenotypes within the A/T/N framework, we provide evidence for the clinical validity of PRS in capturing biologically relevant AD pathology. The broader implications of our findings align with the objectives of the eMERGE Consortium's Genome Informed Risk Assessment (GIRA) initiative, which emphasizes the integration of PRS into clinical care while ensuring analytical validity and health equity across diverse populations. 70 By leveraging diverse datasets and applying ancestry‐specific normalization to PRS calculations, our methodology enhances risk stratification accuracy, supporting the growing consensus that such approaches are crucial for equitable and effective genomic medicine. 16 , 20 , 60

In summary, this study highlights methodological innovations of leveraging multi‐ancestry base GWAS summary statistics, applying ancestry normalization, and extending evaluation beyond diagnosis to AD endophenotypes that can enhance both the accuracy and clinical relevance of AD‐PRS. Importantly, our results call for equitable genomic medicine through increasing genetic representation and moving beyond Eurocentric risk models to reflect the full spectrum of human genetic diversity. By demonstrating practical strategies to improve prediction in underrepresented populations, this work provides a foundation for more inclusive risk assessment tools that can be translated into meaningful clinical applications.

CONFLICT OF INTEREST STATEMENT

C.J.’s participation in this project was part of a competitive contract awarded to DataTecnica LLC by the National Institutes of Health to support open science research. J.S.Y. serves on the scientific advisory board for the Epstein Family Alzheimer's Research Collaboration and the Charleston Conference on Alzheimer's Disease and is the editor‐in‐chief of npj Dementia. Other co‐authors report no conflicts of interest. Author disclosures are available in the supporting information. Author disclosures are available in the Supporting Information.

CONSENT STATEMENT

All participants in the ADSP provided written informed consent, and institutional review board (IRB) approvals were obtained by each contributing cohort. All aspects of the HABS‐HD study protocol are managed by the North Texas Regional IRB. All HABS‐HD partners (and or legal guardian) undergo informed consent and provide written informed authorization to engage in the research study.

Supporting information

Supporting Information: alz71529‐sup‐0001‐figuresS1‐S4.docx

ALZ-22-e71529-s001.docx (4.6MB, docx)

Supporting Information: alz71529‐sup‐0002‐tablesS1‐S6.xlsx

ALZ-22-e71529-s004.xlsx (29.5KB, xlsx)

Supporting Information: alz71529‐sup‐0003‐SuppMat.docx

ALZ-22-e71529-s002.docx (17.6KB, docx)

Supporting Information: alz71529‐sup‐0004‐SuppMat.pdf

ALZ-22-e71529-s003.pdf (583.4KB, pdf)

ACKNOWLEDGMENTS

The authors thank all the participants, staff, researcher teams, and partners of the Alzheimer's Disease Sequencing Project (ADSP). The ADSP is comprised of two Alzheimer's Disease (AD) genetics consortia and three National Human Genome Research Institute (NHGRI) funded Large Scale Sequencing and Analysis Centers (LSAC). The two AD genetics consortia are the Alzheimer's Disease Genetics Consortium (ADGC) funded by NIA (U01 AG032984), and the Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) funded by NIA (R01 AG033193), the National Heart, Lung, and Blood Institute (NHLBI), other National Institute of Health (NIH) institutes and other foreign governmental and non‐governmental organizations. The Discovery Phase analysis of sequence data is supported through UF1AG047133 (to Drs. Schellenberg, Farrer, Pericak‐Vance, Mayeux, and Haines); U01AG049505 to Dr. Seshadri; U01AG049506 to Dr. Boerwinkle; U01AG049507 to Dr. Wijsman; and U01AG049508 to Dr. Goate and the Discovery Extension Phase analysis is supported through U01AG052411 to Dr. Goate, U01AG052410 to Dr. Pericak‐Vance and U01 AG052409 to Drs. Seshadri and Fornage.

Sequencing for the Follow Up Study (FUS) is supported through U01AG057659 (to Drs. PericakVance, Mayeux, and Vardarajan) and U01AG062943 (to Drs. Pericak‐Vance and Mayeux). Data generation and harmonization in the Follow‐up Phase is supported by U54AG052427 (to Drs. Schellenberg and Wang). The FUS Phase analysis of sequence data is supported through U01AG058589 (to Drs. Destefano, Boerwinkle, De Jager, Fornage, Seshadri, and Wijsman), U01AG058654 (to Drs. Haines, Bush, Farrer, Martin, and Pericak‐Vance), U01AG058635 (to Dr. Goate), RF1AG058066 (to Drs. Haines, Pericak‐Vance, and Scott), RF1AG057519 (to Drs. Farrer and Jun), R01AG048927 (to Dr. Farrer), and RF1AG054074 (to Drs. Pericak‐Vance and Beecham).

The ADGC cohorts include: Adult Changes in Thought (ACT) (U01 AG006781, U19 AG066567), the Alzheimer's Disease Research Centers (ADRC) (P30 AG062429, P30 AG066468, P30 AG062421, P30 AG066509, P30 AG066514, P30 AG066530, P30 AG066507, P30 AG066444, P30 AG066518, P30 AG066512, P30 AG066462, P30 AG072979, P30 AG072972, P30 AG072976, P30 AG072975, P30 AG072978, P30 AG072977, P30 AG066519, P30 AG062677, P30 AG079280, P30 AG062422, P30 AG066511, P30 AG072946, P30 AG062715, P30 AG072973, P30 AG066506, P30 AG066508, P30 AG066515, P30 AG072947, P30 AG072931, P30 AG066546, P20 AG068024, P20 AG068053, P20 AG068077, P20 AG068082, P30 AG072958, P30 AG072959), the Chicago Health and Aging Project (CHAP) (R01 AG11101, RC4 AG039085, K23 AG030944), Indiana Memory and Aging Study (IMAS) (R01 AG019771), Indianapolis Ibadan (R01 AG009956, P30 AG010133), the Memory and Aging Project (MAP) (R01 AG17917), Mayo Clinic (MAYO) (R01 AG032990, U01 AG046139, R01 NS080820, RF1 AG051504, P50 AG016574), Mayo Parkinson's Disease controls (NS039764, NS071674, 5RC2HG005605), University of Miami (R01 AG027944, R01 AG028786, R01 AG019085, IIRG09133827, A2011048), the Multi‐Institutional Research in Alzheimer's Genetic Epidemiology Study (MIRAGE) (R01 AG09029, R01 AG025259), the National Centralized Repository for Alzheimer's Disease and Related Dementias (NCRAD) (U24 AG021886), the National Institute on Aging Late Onset Alzheimer's Disease Family Study (NIA‐ LOAD) (U24 AG056270), the Religious Orders Study (ROS) (P30 AG10161, R01 AG15819), the Texas Alzheimer's Research and Care Consortium (TARCC) (funded by the Darrell K Royal Texas Alzheimer's Initiative), Vanderbilt University/Case Western Reserve University (VAN/CWRU) (R01 AG019757, R01 AG021547, R01 AG027944, R01 AG028786, P01 NS026630, and Alzheimer's Association), the Washington Heights‐Inwood Columbia Aging Project (WHICAP) (RF1 AG054023), the University of Washington Families (VA Research Merit Grant, NIA: P50AG005136, R01AG041797, NINDS: R01NS069719), the Columbia University Hispanic Estudio Familiar de Influencia Genetica de Alzheimer (EFIGA) (RF1 AG015473), the University of Toronto (UT) (funded by Wellcome Trust, Medical Research Council, Canadian Institutes of Health Research), and Genetic Differences (GD) (R01 AG007584). The CHARGE cohorts are supported in part by National Heart, Lung, and Blood Institute (NHLBI) infrastructure grant HL105756 (Psaty), RC2HL102419 (Boerwinkle) and the neurology working group is supported by the National Institute on Aging (NIA) R01 grant AG033193.

The CHARGE cohorts participating in the ADSP include the following: Austrian Stroke Prevention Study (ASPS), ASPS‐Family study, and the Prospective Dementia Registry‐Austria (ASPS/PRODEM‐Aus), the Atherosclerosis Risk in Communities (ARIC) Study, the Cardiovascular Health Study (CHS), the Erasmus Rucphen Family Study (ERF), the Framingham Heart Study (FHS), and the Rotterdam Study (RS). ASPS is funded by the Austrian Science Fond (FWF) grant number P20545‐P05 and P13180 and the Medical University of Graz. The ASPS‐Fam is funded by the Austrian Science Fund (FWF) project I904), the EU Joint Programme – Neurodegenerative Disease Research (JPND) in frame of the BRIDGET project (Austria, Ministry of Science) and the Medical University of Graz and the Steiermärkische Krankenanstalten Gesellschaft. PRODEM‐Austria is supported by the Austrian Research Promotion agency (FFG) (Project No. 827462) and by the Austrian National Bank (Anniversary Fund, project 15435. ARIC research is carried out as a collaborative study supported by NHLBI contracts (HHSN268201100005C, HHSN268201100006C, HHSN268201100007C, HHSN268201100008C, HHSN268201100009C, HHSN268201100010C, HHSN268201100011C, and HHSN268201100012C). Neurocognitive data in ARIC is collected by U01 2U01HL096812, 2U01HL096814, 2U01HL096899, 2U01HL096902, 2U01HL096917 from the NIH (NHLBI, NINDS, NIA and NIDCD), and with previous brain MRI examinations funded by R01‐HL70825 from the NHLBI. CHS research was supported by contracts HHSN268201200036C, HHSN268200800007C, N01HC55222, N01HC85079, N01HC85080, N01HC85081, N01HC85082, N01HC85083, N01HC85086, and grants U01HL080295 and U01HL130114 from the NHLBI with additional contribution from the National Institute of Neurological Disorders and Stroke (NINDS). Additional support was provided by R01AG023629, R01AG15928, and R01AG20098 from the NIA. FHS research is supported by NHLBI contracts N01‐HC‐25195 and HHSN268201500001I. This study was also supported by additional grants from the NIA (R01s AG054076, AG049607 and AG033040 and NINDS (R01 NS017950). The ERF study as a part of EUROSPAN (European Special Populations Research Network) was supported by European Commission FP6 STRP grant number 018947 (LSHG‐CT‐2006‐01947) and also received funding from the European Community's Seventh Framework Programme (FP7/2007‐2013)/grant agreement HEALTH‐F4‐ 2007‐201413 by the European Commission under the programme “Quality of Life and Management of the Living Resources” of 5th Framework Programme (no. QLG2‐CT‐2002‐ 01254). High‐throughput analysis of the ERF data was supported by a joint grant from the Netherlands Organization for Scientific Research and the Russian Foundation for Basic Research (NWO‐RFBR 047.017.043). The Rotterdam Study is funded by Erasmus Medical Center and Erasmus University, Rotterdam, the Netherlands Organization for Health Research and Development (ZonMw), the Research Institute for Diseases in the Elderly (RIDE), the Ministry of Education, Culture and Science, the Ministry for Health, Welfare and Sports, the European Commission (DG XII), and the municipality of Rotterdam. Genetic data sets are also supported by the Netherlands Organization of Scientific Research NWO Investments (175.010.2005.011, 911‐03‐012), the Genetic Laboratory of the Department of Internal Medicine, Erasmus MC, the Research Institute for Diseases in the Elderly (014‐93‐015; RIDE2), and the Netherlands Genomics Initiative (NGI)/Netherlands Organization for Scientific Research (NWO) Netherlands Consortium for Healthy Aging (NCHA), project 050‐060‐810. All studies are grateful to their participants, faculty and staff. The content of these manuscripts is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health or the U.S. Department of Health and Human Services.

The FUS cohorts include: the Alzheimer's Disease Research Centers (ADRC) (P30 AG062429, P30 AG066468, P30 AG062421, P30 AG066509, P30 AG066514, P30 AG066530, P30 AG066507, P30 AG066444, P30 AG066518, P30 AG066512, P30 AG066462, P30 AG072979, P30 AG072972, P30 AG072976, P30 AG072975, P30 AG072978, P30 AG072977, P30 AG066519, P30 AG062677, P30 AG079280, P30 AG062422, P30 AG066511, P30 AG072946, P30 AG062715, P30 AG072973, P30 AG066506, P30 AG066508, P30 AG066515, P30 AG072947, P30 AG072931, P30 AG066546, P20 AG068024, P20 AG068053, P20 AG068077, P20 AG068082, P30 AG072958, P30 AG072959), Alzheimer's Disease Neuroimaging Initiative (ADNI) (U19AG024904), Amish Protective Variant Study (RF1AG058066), Cache County Study (R01AG11380, R01AG031272, R01AG21136, RF1AG054052), Case Western Reserve University Brain Bank (CWRUBB) (P50AG008012), Case Western Reserve University Rapid Decline (CWRURD) (RF1AG058267, NU38CK000480), CubanAmerican Alzheimer's Disease Initiative (CuAADI) (3U01AG052410), Estudio Familiar de Influencia Genetica en Alzheimer (EFIGA) (5R37AG015473, RF1AG015473, R56AG051876), Genetic and Environmental Risk Factors for Alzheimer Disease Among African Americans Study (GenerAAtions) (2R01AG09029, R01AG025259, 2R01AG048927), Gwangju Alzheimer and Related Dementias Study (GARD) (U01AG062602), Hillblom Aging Network (2014‐A‐004‐NET, R01AG032289, R01AG048234), Hussman Institute for Human Genomics Brain Bank (HIHGBB) (R01AG027944, Alzheimer's Association “Identification of Rare Variants in Alzheimer Disease”), Ibadan Study of Aging (IBADAN) (5R01AG009956), Longevity Genes Project (LGP) and LonGenity (R01AG042188, R01AG044829, R01AG046949, R01AG057909, R01AG061155, P30AG038072), Mexican Health and Aging Study (MHAS) (R01AG018016), Multi‐Institutional Research in Alzheimer's Genetic Epidemiology (MIRAGE) (2R01AG09029, R01AG025259, 2R01AG048927), Northern Manhattan Study (NOMAS) (R01NS29993), Peru Alzheimer's Disease Initiative (PeADI) (RF1AG054074), Puerto Rican 1066 (PR1066) (Wellcome Trust (GR066133/GR080002), European Research Council (340755)), Puerto Rican Alzheimer Disease Initiative (PRADI) (RF1AG054074), Reasons for Geographic and Racial Differences in Stroke (REGARDS) (U01NS041588), Research in African American Alzheimer Disease Initiative (REAAADI) (U01AG052410), the Religious Orders Study (ROS) (P30 AG10161, P30 AG72975, R01 AG15819, R01 AG42210), the RUSH Memory and Aging Project (MAP) (R01 AG017917, R01 AG42210Stanford Extreme Phenotypes in AD (R01AG060747), University of Miami Brain Endowment Bank (MBB), University of Miami/Case Western/North Carolina A&T African American (UM/CASE/NCAT) (U01AG052410, R01AG028786), and Wisconsin Registry for Alzheimer's Prevention (WRAP) (R01AG027161 and R01AG054047).

The four LSACs are: the Human Genome Sequencing Center at the Baylor College of Medicine (U54 HG003273), the Broad Institute Genome Center (U54HG003067), The American Genome Center at the Uniformed Services University of the Health Sciences (U01AG057659), and the Washington University Genome Institute (U54HG003079). Genotyping and sequencing for the ADSP FUS is also conducted at John P. Hussman Institute for Human Genomics (HIHG) Center for Genome Technology (CGT).

Biological samples and associated phenotypic data used in primary data analyses were stored at Study Investigators institutions, and at the National Centralized Repository for Alzheimer's Disease and Related Dementias (NCRAD, U24AG021886) at Indiana University funded by NIA. Associated Phenotypic Data used in primary and secondary data analyses were provided by Study Investigators, the NIA funded Alzheimer's Disease Centers (ADCs), and the National Alzheimer's Coordinating Center (NACC, U24AG072122) and the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS, U24AG041689) at the University of Pennsylvania, funded by NIA. Harmonized phenotypes were provided by the Alzheimer's Disease Sequencing Project Phenotype Harmonization Consortium (ADSP‐PHC), funded by NIA (U24 AG074855, U01 AG068057 and R01 AG059716) and Ultrascale Machine Learning to Empower Discovery in Alzheimer's Disease Biobanks (AI4AD, U01 AG068057). This research was supported in part by the Intramural Research Program of the National Institutes of health, National Library of Medicine. Contributors to the Genetic Analysis Data included Study Investigators on projects that were individually funded by NIA, and other NIH institutes, and by private U.S. organizations, or foreign governmental or nongovernmental organizations.

The ADSP Phenotype Harmonization Consortium (ADSP‐PHC) is funded by NIA (U24 AG074855, U01 AG068057 and R01 AG059716). The harmonized cohorts within the ADSP‐PHC include: the Anti‐Amyloid Treatment in Asymptomatic Alzheimer's study (A4 Study), a secondary prevention trial in preclinical Alzheimer's disease, aiming to slow cognitive decline associated with brain amyloid accumulation in clinically normal older individuals. The A4 Study is funded by a public‐private‐philanthropic partnership, including funding from the National Institutes of Health‐National Institute on Aging, Eli Lilly and Company, Alzheimer's Association, Accelerating Medicines Partnership, GHR Foundation, an anonymous foundation and additional private donors, with in‐kind support from Avid and Cogstate. The companion observational Longitudinal Evaluation of Amyloid Risk and Neurodegeneration (LEARN) Study is funded by the Alzheimer's Association and GHR Foundation. The A4 and LEARN Studies are led by Dr. Reisa Sperling at Brigham and Women's Hospital, Harvard Medical School and Dr. Paul Aisen at the Alzheimer's Therapeutic Research Institute (ATRI), University of Southern California. The A4 and LEARN Studies are coordinated by ATRI at the University of Southern California, and the data are made available through the Laboratory for Neuro Imaging at the University of Southern California. The participants screening for the A4 Study provided permission to share their de‐identified data in order to advance the quest to find a successful treatment for Alzheimer's disease. We would like to acknowledge the dedication of all the participants, the site personnel, and all of the partnership team members who continue to make the A4 and LEARN Studies possible. The complete A4 Study Team list is available on: a4study.org/a4‐study‐team.; the Adult Changes in Thought study (ACT), U01 AG006781, U19 AG066567; Alzheimer's Disease Neuroimaging Initiative (ADNI): Data collection and sharing for this project was funded by the ADNI (National Institutes of Health Grant U01 AG024904) and DOD ADNI (Department of Defense award number W81XWH‐12‐2‐0012). ADNI is funded by the National Institute on Aging, the National Institute of Biomedical Imaging and Bioengineering, and through generous contributions from the following: AbbVie, Alzheimer's Association; Alzheimer's Drug Discovery Foundation; Araclon Biotech; BioClinica, Inc.; Biogen; Bristol‐Myers Squibb Company; CereSpir, Inc.; Cogstate; Eisai Inc.; Elan Pharmaceuticals, Inc.; Eli Lilly and Company; EuroImmun; F. Hoffmann‐La Roche Ltd and its affiliated company Genentech, Inc.; Fujirebio; GE Healthcare; IXICO Ltd.;Janssen Alzheimer Immunotherapy Research & Development, LLC.; Johnson & Johnson Pharmaceutical Research & Development LLC.; Lumosity; Lundbeck; Merck & Co., Inc.;Meso Scale Diagnostics, LLC.; NeuroRx Research; Neurotrack Technologies; Novartis Pharmaceuticals Corporation; Pfizer Inc.; Piramal Imaging; Servier; Takeda Pharmaceutical Company; and Transition Therapeutics. The Canadian Institutes of Health Research is providing funds to support ADNI clinical sites in Canada. Private sector contributions are facilitated by the Foundation for the National Institutes of Health (www.fnih.org). The grantee organization is the Northern California Institute for Research and Education, and the study is coordinated by the Alzheimer's Therapeutic Research Institute at the University of Southern California. ADNI data are disseminated by the Laboratory for Neuro Imaging at the University of Southern California; Estudio Familiar de Influencia Genetica en Alzheimer (EFIGA): 5R37AG015473, RF1AG015473, R56AG051876; Memory & Aging Project at Knight Alzheimer's Disease Research Center (MAP at Knight ADRC): The Memory and Aging Project at the Knight‐ADRC (Knight‐ADRC). This work was supported by the National Institutes of Health (NIH) grants R01AG064614, R01AG044546, RF1AG053303, RF1AG058501, U01AG058922 and R01AG064877 to Carlos Cruchaga. The recruitment and clinical characterization of research participants at Washington University was supported by NIH grants P30AG066444, P01AG03991, and P01AG026276. Data collection and sharing for this project was supported by NIH grants RF1AG054080, P30AG066462, R01AG064614 and U01AG052410. We thank the contributors who collected samples used in this study, as well as patients and their families, whose help and participation made this work possible. This work was supported by access to equipment made possible by the Hope Center for Neurological Disorders, the Neurogenomics and Informatics Center (NGI: https://neurogenomics.wustl.edu/) and the Departments of Neurology and Psychiatry at Washington University School of Medicine; National Alzheimer's Coordinating Center (NACC): The NACC database is funded by NIA/NIH Grant U24 AG072122. NACC data are contributed by the NIA‐funded ADRCs: P30 AG062429 (PI James Brewer, MD, PhD), P30 AG066468 (PI Oscar Lopez, MD), P30 AG062421 (PI Bradley Hyman, MD, PhD), P30 AG066509 (PI Thomas Grabowski, MD), P30 AG066514 (PI Mary Sano, PhD), P30 AG066530 (PI Helena Chui, MD), P30 AG066507 (PI Marilyn Albert, PhD), P30 AG066444 (PI John Morris, MD), P30 AG066518 (PI Jeffrey Kaye, MD), P30 AG066512 (PI Thomas Wisniewski, MD), P30 AG066462 (PI Scott Small, MD), P30 AG072979 (PI David Wolk, MD), P30 AG072972 (PI Charles DeCarli, MD), P30 AG072976 (PI Andrew Saykin, PsyD), P30 AG072975 (PI David Bennett, MD), P30 AG072978 (PI Neil Kowall, MD), P30 AG072977 (PI Robert Vassar, PhD), P30 AG066519 (PI Frank LaFerla, PhD), P30 AG062677 (PI Ronald Petersen, MD, PhD), P30 AG079280 (PI Eric Reiman, MD), P30 AG062422 (PI Gil Rabinovici, MD), P30 AG066511 (PI Allan Levey, MD, PhD), P30 AG072946 (PI Linda Van Eldik, PhD), P30 AG062715 (PI Sanjay Asthana, MD, FRCP), P30 AG072973 (PI Russell Swerdlow, MD), P30 AG066506 (PI Todd Golde, MD, PhD), P30 AG066508 (PI Stephen Strittmatter, MD, PhD), P30 AG066515 (PI Victor Henderson, MD, MS), P30 AG072947 (PI Suzanne Craft, PhD), P30 AG072931 (PI Henry Paulson, MD, PhD), P30 AG066546 (PI Sudha Seshadri, MD), P20 AG068024 (PI Erik Roberson, MD, PhD), P20 AG068053 (PI Justin Miller, PhD), P20 AG068077 (PI Gary Rosenberg, MD), P20 AG068082 (PI Angela Jefferson, PhD), P30 AG072958 (PI Heather Whitson, MD), P30 AG072959 (PI James Leverenz, MD); National Institute on Aging Alzheimer's Disease Family Based Study (NIA‐AD FBS): U24 AG056270; Religious Orders Study (ROS): P30AG10161,R01AG15819, R01AG42210; Memory and Aging Project (MAP—Rush): R01AG017917, R01AG42210; Minority Aging Research Study (MARS): R01AG22018, R01AG42210; Washington Heights/Inwood Columbia Aging Project (WHICAP): RF1 AG054023;and Wisconsin Registry for Alzheimer's Prevention (WRAP): R01AG027161 and R01AG054047. Additional acknowledgments include the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS, U24AG041689) at the University of Pennsylvania, funded by NIA.

Data for this study were prepared, archived, and distributed by the National Institute on Aging Alzheimer's Disease Data Storage Site (NIAGADS) at the University of Pennsylvania (U24‐ AG041689), funded by the National Institute on Aging.

The NACC database is funded by NIA/NIH Grant U24 AG072122. NACC data are contributed by the NIA‐funded ADRCs: P30 AG062429 (PI James Brewer, MD, PhD), P30 AG066468 (PI Oscar Lopez, MD), P30 AG062421 (PI Bradley Hyman, MD, PhD), P30 AG066509 (PI Thomas Grabowski, MD), P30 AG066514 (PI Mary Sano, PhD), P30 AG066530 (PI Helena Chui, MD), P30 AG066507 (PI Marilyn Albert, PhD), P30 AG066444 (PI John Morris, MD), P30 AG066518 (PI Jeffrey Kaye, MD), P30 AG066512 (PI Thomas Wisniewski, MD), P30 AG066462 (PI Scott Small, MD), P30 AG072979 (PI David Wolk, MD), P30 AG072972 (PI Charles DeCarli, MD), P30 AG072976 (PI Andrew Saykin, PsyD), P30 AG072975 (PI David Bennett, MD), P30 AG072978 (PI Neil Kowall, MD), P30 AG072977 (PI Robert Vassar, PhD), P30 AG066519 (PI Frank LaFerla, PhD), P30 AG062677 (PI Ronald Petersen, MD, PhD), P30 AG079280 (PI Eric Reiman, MD), P30 AG062422 (PI Gil Rabinovici, MD), P30 AG066511 (PI Allan Levey, MD, PhD), P30 AG072946 (PI Linda Van Eldik, PhD), P30 AG062715 (PI Sanjay Asthana, MD, FRCP), P30 AG072973 (PI Russell Swerdlow, MD), P30 AG066506 (PI Todd Golde, MD, PhD), P30 AG066508 (PI Stephen Strittmatter, MD, PhD), P30 AG066515 (PI Victor Henderson, MD, MS), P30 AG072947 (PI Suzanne Craft, PhD), P30 AG072931 (PI Henry Paulson, MD, PhD), P30 AG066546 (PI Sudha Seshadri, MD), P20 AG068024 (PI Erik Roberson, MD, PhD), P20 AG068053 (PI Justin Miller, PhD), P20 AG068077 (PI Gary Rosenberg, MD), P20 AG068082 (PI Angela Jefferson, PhD), P30 AG072958 (PI Heather Whitson, MD), P30 AG072959 (PI James Leverenz, MD).

Data collection and sharing for this project was funded by the Alzheimer's Disease Neuroimaging Initiative (ADNI) (National Institutes of Health Grant U01 AG024904) and DOD ADNI (Department of Defense award number W81XWH‐12‐2‐0012). ADNI is funded by the National Institute on Aging, the National Institute of Biomedical Imaging and Bioengineering, and through generous contributions from the following: AbbVie, Alzheimer's Association; Alzheimer's Drug Discovery Foundation; Araclon Biotech; BioClinica, Inc.; Biogen; Bristol‐Myers Squibb Company; CereSpir, Inc.; Cogstate; Eisai Inc.; Elan Pharmaceuticals, Inc.; Eli Lilly and Company; EuroImmun; F. Hoffmann‐La Roche Ltd and its affiliated company Genentech, Inc.; Fujirebio; GE Healthcare; IXICO Ltd.;Janssen Alzheimer Immunotherapy Research & Development, LLC.; Johnson & Johnson Pharmaceutical Research & Development LLC.; Lumosity; Lundbeck; Merck & Co., Inc.;Meso Scale Diagnostics, LLC.; NeuroRx Research; Neurotrack Technologies; Novartis Pharmaceuticals Corporation; Pfizer Inc.; Piramal Imaging; Servier; Takeda Pharmaceutical Company; and Transition Therapeutics. The Canadian Institutes of Health Research is providing funds to support ADNI clinical sites in Canada. Private sector contributions are facilitated by the Foundation for the National Institutes of Health (www.fnih.org). The grantee organization is the Northern California Institute for Research and Education, and the study is coordinated by the Alzheimer's Therapeutic Research Institute at the University of Southern California. ADNI data are disseminated by the Laboratory for Neuro Imaging at the University of Southern California.

The authors would like to thank all the participants, staff, researcher teams, and partners of the Health & Aging Brain Study—Health Disparities (HABS‐HD). The HABS‐HD is supported by the National Institute on Aging of the National Institutes of Health under Award Numbers R01AG054073, R01AG058533, R01AG070862, P41EB015922, and U19AG078109. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. We gratefully acknowledge the contributions of our study partners and their families, whose help and participation made this work possible. C.J. is supported in part by the NIH Intramural Center for Alzheimer's and Related Dementias (CARD), project NIH‐NIA ZIAAG000534. J.S.Y. receives funding from NIH‐NIA R01AG062588, R01AG057234, P30AG062422, P01AG019724, and U19AG079774, NIH‐NINDS U54NS123985; the Rainwater Charitable Foundation; the Alzheimer's Association; the Global Brain Health Institute; Genentech; the French Foundation; and the Mary Oakley Foundation. This work was conducted using the National Alzheimer's Coordinating Center Uniform Dataset under application 10238; the Alzheimer's Disease Neuroimaging Initiative under application SJA; and the Alzheimer's Disease Sequencing Project under application 10050. SJA is supported by the National Alzheimer's Coordinating Center New Investigator Award (U24AG072122).

DATA AVAILABILITY STATEMENT

European summary statistics from Bellenguez et al. 2022 5 were accessed through the National Human Genome Research Institute‐European Bioinformatics Institute using accession number GCST90027158. Finnish European summary statistics from FinnGen Release 6 were accessed through https://www.finngen.fi/en/access_results. African summary Statistics from Kunkle et al. 2021 28 were accessed through NIAGADS (https://www.niagads.org/) using accession number NG00100. East Asian summary statistics from Shigemizu et al. 2021 29 were accessed through the National Bioscience Database Center (NBDC) at the Japan Science and Technology Agency (JST) at https://humandbs.biosciencedbc.jp/en/ through accession number hum0237.v1.gwas.v1. Caribbean Hispanics summary statistics from the Columbia University Study of Caribbean Hispanics and Late‐Onset Alzheimer's disease were accessed via application to dbGaP accession number phs000496.v1.p1. All codes developed and used are available at https://github.com/AndrewsLabUCSF/AFRPRS. Information regarding all the software and reference datasets used in the analysis can also be found in the repository.

REFERENCES

  • 1. Choi SW, Mak TS‐H, O'Reilly PF. Tutorial: a guide to performing polygenic risk score analyses. Nat Protoc. 2020;15:2759‐2772. doi: 10.1038/s41596-020-0353-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Khera AV, Chaffin M, Aragam KG. Genome‐wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations. Nat Genet. 2018;50:1219‐1224. doi: 10.1038/s41588-018-0183-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Lewis CM, Vassos E. Polygenic risk scores: from research tools to clinical instruments. Genome Med. 2020;12:1‐11. doi: 10.1186/s13073-020-00742-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Karlsson IK, Escott‐Price V, Gatz M, et al. Measuring heritable contributions to Alzheimer's disease: polygenic risk score analysis with twins. Brain Commun. 2022;4(1):fcab308. doi: 10.1093/braincomms/fcab308 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Bellenguez C, Küçükali F, Jansen IE, et al. New insights into the genetic etiology of Alzheimer's disease and related dementias. Nat Genet. 2022;54:412‐436. doi: 10.1038/s41588-022-01024-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Wightman DP, Jansen IE, Savage JE, et al. A genome‐wide association study with 1,126,563 individuals identifies new risk loci for Alzheimer's disease. Nat Genet. 2021;53:1276‐1282. doi: 10.1038/s41588-021-00921-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Baker E, Leonenko G, Schmidt KM, et al. What does heritability of Alzheimer's disease represent?. PLoS One. 2023;18:e0281440. doi: 10.1371/journal.pone.0281440 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Leonenko G, Baker E, Stevenson‐Hoare J, et al. Identifying individuals with high risk of Alzheimer's disease using polygenic risk scores. Nat Commun. 2021;12:4506. doi: 10.1038/s41467-021-24082-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Escott‐Price V, Jones L. Genomic profiling and diagnostic biomarkers in Alzheimer's disease. Lancet Neurol. 2017;16:582‐583. doi: 10.1016/S1474-4422(17)30202-8 [DOI] [PubMed] [Google Scholar]
  • 10. Escott‐Price V, Sims R, Bannister C, et al. Common polygenic variation enhances risk prediction for Alzheimer's disease. Brain. 2015;138:3673‐3684. doi: 10.1093/brain/awv268 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Osterman MD, Song YE, Lynn A, et al. Founder population‐specific weights yield improvements in performance of polygenic risk scores for Alzheimer disease in the Midwestern Amish. Hum Genet Genomics Adv. 2023;4:100241. doi: 10.1016/j.xhgg.2023.100241 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Sariya S, Felsky D, Reyes‐Dumeyer D, et al. Polygenic risk score for Alzheimer's disease in Caribbean Hispanics. Ann Neurol. 2021;90:366‐376. doi: 10.1002/ana.26131 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Kamiza AB, Toure SM, Vujkovic M, et al. Transferability of genetic risk scores in African populations. Nat Med. 2022;28:1163‐1166. doi: 10.1038/s41591-022-01835-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Zhou X, Chen Y, Ip FCF, et al. Genetic and polygenic risk score analysis for Alzheimer's disease in the Chinese population. Alzheimers Dement Diagn Assess Dis Monit. 2020;12:e12074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Cavazos TB, Witte JS. Inclusion of variants discovered from diverse populations improves polygenic risk score transferability. Hum Genet Genomics Adv. 2020;2:100017. doi: 10.1016/j.xhgg.2020.100017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Wang Y, Tsuo K, Kanai M, Neale BM, Martin AR. Challenges and opportunities for developing more generalizable polygenic risk scores. Annu Rev Biomed Data Sci. 2022;5:293‐320. doi: 10.1146/annurev-biodatasci-111721-074830 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Ding Yi, Hou K, Xu Z, et al. Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature. 2023;618:774‐781. doi: 10.1038/s41586-023-06079-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Jung S‐H, Kim H‐R, Chun MY, et al. Transferability of Alzheimer disease polygenic risk score across populations and its association with Alzheimer disease‐related phenotypes. JAMA Netw Open. 2022;5:e2247162. doi: 10.1001/jamanetworkopen.2022.47162 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. 2025 Alzheimer's disease facts and figures, Alzheimers Dement. 2025;21(4):e70235. doi: 10.1002/alz.70235 [DOI] [Google Scholar]
  • 20. Wang Y, Guo J, Ni G, Yang J, Visscher PM, Yengo L. Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations. Nat Commun. 2020;11:3865. doi: 10.1038/s41467-020-17719-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Bocher O, Gilly A, Park Y‐C, Zeggini E, Morris AP. Bridging the diversity gap: analytical and study design considerations for improving the accuracy of trans‐ancestry genetic prediction. Hum Genet Genomics Adv. 2023;4:100214. doi: 10.1016/j.xhgg.2023.100214 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Hoggart C, Choi SW, García‐González J, Souaiaia T, Preuss M, O'Reilly P. BridgePRS: a powerful trans‐ancestry polygenic risk score method. Preprint at bioRxiv. doi: 10.1101/2023.02.17.528938 (2023). [DOI] [Google Scholar]
  • 23. Ruan Y, Lin Y‐F, Feng Y‐CA, et al. Improving polygenic prediction in ancestrally diverse populations. Nat Genet. 2022;54:573‐580. doi: 10.1038/s41588-022-01054-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Márquez‐Luna C, Loh Po‐Ru. Multiethnic polygenic risk scores improve risk prediction in diverse populations. Genet Epidemiol. 2017;41:811‐823. doi: 10.1002/gepi.22083 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Wang Y, Kanai M, Tan T, et al. Polygenic prediction across populations is influenced by ancestry, genetic architecture, and methodology. Cell Genomics. 2023;3:100408. doi: 10.1016/j.xgen.2023.100408 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Jack CR, Bennett DA, Blennow K, et al. A/T/N: an unbiased descriptive classification scheme for Alzheimer disease biomarkers. Neurology. 2016;87:539‐547. doi: 10.1212/WNL.0000000000002923 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Jack CR, Andrews JS, Beach TG, et al. Revised criteria for diagnosis and staging of Alzheimer's disease: Alzheimer's Association Workgroup. Alzheimers Dement. 2024;20:5143‐5169. doi: 10.1002/alz.13859 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Kunkle BW, Schmidt M, Klein H‐U, et al. Novel Alzheimer disease risk loci and pathways in African American individuals using the African genome resources panel. JAMA Neurol. 2021;78:1‐13. doi: 10.1001/jamaneurol.2020.3536 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Shigemizu D, Mitsumori R, Akiyama S, et al. Ethnic and trans‐ethnic genome‐wide association studies identify new loci influencing Japanese Alzheimer's disease risk. Transl Psychiatry. 2021;11:151. doi: 10.1038/s41398-021-01272-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Lake J, Warly Solsberg C, Kim JJ, et al. Multi‐ancestry meta‐analysis and fine‐mapping in Alzheimer's disease. Mol Psychiatry. 2023;28:3121‐3132. doi: 10.1038/s41380-023-02089-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. NIAGADS/ADSPIntegratedPhenotypes: this script is intended to provide approved users of the Alzheimer's disease sequencing project (ADSP) dataset (ng00067) with a starting point for combining and using sample and phenotype information to generate a genetically unique list of whole genome samples to use in an Alzheimer's disease case/control analysis. https://github.com/NIAGADS/ADSPIntegratedPhenotypes/tree/main
  • 32. Leung YY, Lee W‐P, Kuzma AB, et al. Alzheimer's disease sequencing project release 4 whole genome sequencing dataset. Alzheimers Dement. 2025;21:e70237. doi: 10.1002/alz.70237 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Mukherjee S, Choi S‐E, Lee ML, et al. Cognitive domain harmonization and cocalibration in studies of older adults. Neuropsychology. 2023;37:409‐423. doi: 10.1037/neu0000835 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Vitale D, Koretsky M, Kuznetsov N, et al. GenoTools: an open‐source Python Package for efficient genotype data quality control and analysis. Preprint at bioRxiv. 2024. doi: 10.1101/2024.03.26.586362 . bioRxiv [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Ellis CA, Oliver KL, Harris RV, et al. Inflation of polygenic risk scores caused by sample overlap and relatedness: examples of a major risk of bias. Am J Hum Genet. 2024;111(9):1806‐1809. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Relationship inference in KING. https://www.kingrelatedness.com/manual.shtml
  • 37. Lambert SA, Wingfield B, Gibson JT, et al. Enhancing the polygenic score catalog with tools for score calculation and ancestry normalization. Nat Genet. 2024;56:1989‐1994. doi: 10.1038/s41588-024-01937-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Alexander DH, Novembre J, Lange K. Fast model‐based estimation of ancestry in unrelated individuals. Genome Res. 2009;19:1655‐1664. doi: 10.1101/gr.094052.109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. O'Bryant SE, Johnson LA, Barber RC, et al. The Health & Aging Brain among Latino Elders (HABLE) study methods and participant characteristics. Alzheimers Dement Diagn Assess Dis Monit. 2021;13:e12202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Petersen ME, Zhang F, Hall J, et al. Characterization of plasma AT(N) biomarkers among a racial and ethnically diverse community‐based cohort: an HABS‐HD study. Alzheimers Dement Transl Res Clin Interv. 2025;11:e70045. doi: 10.1002/trc2.70045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Timsina J, Ali M, Do A, et al. Harmonization of CSF and imaging biomarkers in Alzheimer's disease: need and practical applications for genetics studies and preclinical classification. Neurobiol Dis. 2024;190:106373. doi: 10.1016/j.nbd.2023.106373 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Andrews SJ, Fulton‐Howard B, O'Reilly P, Marcora E. Causal associations between modifiable risk factors and the Alzheimer's phenome. Ann Neurol. 2021;89:54‐65. doi: 10.1002/ana.25918 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. TOPMed Imputation Server. https://imputation.biodatacatalyst.nhlbi.nih.gov/#!
  • 44. Shi M, Tanikawa C, Munter HM, et al. Genotype imputation accuracy and the quality metrics of the minor ancestry in multi‐ancestry reference panels. Brief Bioinform. 2024;25:bbad509. doi: 10.1093/bib/bbad509 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Vargas LB, Meyer MC, Konigsberg IR, et al. Ancestry calibration of polygenic risk scores improves risk stratification and effect estimation in African American adults. Preprint at medRxiv. 2025. doi: 10.1101/2025.06.18.25329573 [DOI] [Google Scholar]
  • 46. Khera AV, Chaffin M, Zekavat SM, et al. Whole‐genome sequencing to characterize monogenic and polygenic contributions in patients hospitalized with early‐onset myocardial infarction. Circulation. 2019;139:1593‐1602. doi: 10.1161/CIRCULATIONAHA.118.035658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Zhang D, Dey R, Lee S. Fast and robust ancestry prediction using principal component analysis. Bioinformatics. 2020;36:3439‐3446. doi: 10.1093/bioinformatics/btaa152 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Momin MdM, Lee S, Wray NR, Lee SH. Significance tests for R2 of out‐of‐sample prediction using polygenic scores. Am J Hum Genet. 2023;110:349‐358. doi: 10.1016/j.ajhg.2023.01.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Martin AR, Kanai M, Kamatani Y, Okada Y, Neale BM. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat Genet. 2019;51:584‐591. doi: 10.1038/s41588-019-0379-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Petrovski S, Goldstein DB. Unequal representation of genetic variation across ancestry groups creates healthcare inequality in the application of precision medicine. Genome Biol. 2016;17:157. doi: 10.1186/s13059-016-1016-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Martin AR, Gignoux CR, Walters RK, et al. Human demographic history impacts genetic risk prediction across diverse populations. Am J Hum Genet. 2017;100:635‐649. doi: 10.1016/j.ajhg.2017.03.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. 1000 Genomes Project Consortium , Auton A, Brooks LD, et al. A global reference for human genetic variation. Nature. 2015;526:68‐74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Graham SE, Clarke SL, Wu K‐HH, et al. The power of genetic diversity in genome‐wide association studies of lipids. Nature. 2021;600:675‐679. doi: 10.1038/s41586-021-04064-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Wall JD, Pritchard JK. Haplotype blocks and linkage disequilibrium in the human genome. Nat Rev Genet. 2003;4:587‐597. doi: 10.1038/nrg1123 [DOI] [PubMed] [Google Scholar]
  • 55. Jakobsson M, Scholz SW, Scheet P, et al. Genotype, haplotype and copy‐number variation in worldwide human populations. Nature. 2008;451:998‐1003. doi: 10.1038/nature06742 [DOI] [PubMed] [Google Scholar]
  • 56. Lee JS, Bae JS, Park B‐L, et al. Association analysis of TEC polymorphisms with aspirin‐exacerbated respiratory disease in a Korean population. Genomics Inform. 2014;12:58‐63. doi: 10.5808/GI.2014.12.2.58 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Nicolas A, Sherva R, Grenier‐Boley B, et al. Transferability of European‐derived Alzheimer's disease polygenic risk scores across multiancestry populations. Nat Genet. 2025:1‐13. doi: 10.1038/s41588-025-02227-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Andrews SJ, Renton AE, Fulton‐Howard B, Podlesny‐Drabiniok A, Marcora E. The complex genetic architecture of Alzheimer's disease: novel insights and future directions. eBioMedicine. 2023;90:104511. doi: 10.1016/j.ebiom.2023.104511 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Chen C‐Y, Han J, Hunter DJ, Kraft P, Price AL. Explicit modeling of ancestry improves polygenic risk scores and BLUP prediction. Genet Epidemiol. 2015;39:427‐438. doi: 10.1002/gepi.21906 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Ge T, Irvin MR, Patki A, et al. Development and validation of a trans‐ancestry polygenic risk score for type 2 diabetes in diverse populations. Genome Med. 2022;14:70. doi: 10.1186/s13073-022-01074-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Sun Q, Rowland BT, Chen J, et al. Improving polygenic risk prediction in admixed populations by explicitly modeling ancestral‐specific effects via GAUDI. Preprint at bioRxiv. 2022. doi: 10.1101/2022.10.06.511219 [DOI] [PMC free article] [PubMed]
  • 62. Majara L, Kalungi A, Koen N, et al. Low and differential polygenic score generalizability among African populations due largely to genetic diversity. Hum Genet Genomics Adv. 2023;4:100184. doi: 10.1016/j.xhgg.2023.100184 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Boye C, Nirmalan S, Ranjbaran A, Luca F. Genotype × environment interactions in gene regulation and complex traits. Nat Genet. 2024:1‐12. doi: 10.1038/s41588-024-01776-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Jirtle RL, Skinner MK. Environmental epigenomics and disease susceptibility. Nat Rev Genet. 2007;8:253‐262. doi: 10.1038/nrg2045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Migliore L, Coppedè F. Genetics, environmental factors and the emerging role of epigenetics in neurodegenerative diseases. Mutat Res Mol Mech Mutagen. 2009;667:82‐97. [DOI] [PubMed] [Google Scholar]
  • 66. Privé F, Arbel J, Vilhjálmsson BJ. LDpred2: better, faster, stronger. Bioinformatics. 2021;36:5424‐5431. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Hoggart CJ, Choi SW, García‐González J, Souaiaia T, Preuss M. BridgePRS leverages shared genetic effects across ancestries to increase polygenic risk score portability. Nat Genet. 2024;56:180‐186. doi: 10.1038/s41588-023-01583-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68. Yang J, Lee SH, Goddard ME, Visscher PM. GCTA: a tool for genome‐wide complex trait analysis. Am J Hum Genet. 2011;88:76‐82. doi: 10.1016/j.ajhg.2010.11.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Jayasinghe D, Eshetie S, Beckmann K, Benyamin B, Lee SH. Advancements and limitations in polygenic risk score methods for genomic prediction: a scoping review. Hum Genet. 2024;143:1401‐1431. doi: 10.1007/s00439-024-02716-8 [DOI] [PubMed] [Google Scholar]
  • 70. Linder JE, Allworth A, Bland HT, et al. Returning integrated genomic risk and clinical recommendations: the eMERGE study. Genet Med Off J Am Coll Med Genet. 2023;25:100006. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting Information: alz71529‐sup‐0001‐figuresS1‐S4.docx

ALZ-22-e71529-s001.docx (4.6MB, docx)

Supporting Information: alz71529‐sup‐0002‐tablesS1‐S6.xlsx

ALZ-22-e71529-s004.xlsx (29.5KB, xlsx)

Supporting Information: alz71529‐sup‐0003‐SuppMat.docx

ALZ-22-e71529-s002.docx (17.6KB, docx)

Supporting Information: alz71529‐sup‐0004‐SuppMat.pdf

ALZ-22-e71529-s003.pdf (583.4KB, pdf)

Data Availability Statement

European summary statistics from Bellenguez et al. 2022 5 were accessed through the National Human Genome Research Institute‐European Bioinformatics Institute using accession number GCST90027158. Finnish European summary statistics from FinnGen Release 6 were accessed through https://www.finngen.fi/en/access_results. African summary Statistics from Kunkle et al. 2021 28 were accessed through NIAGADS (https://www.niagads.org/) using accession number NG00100. East Asian summary statistics from Shigemizu et al. 2021 29 were accessed through the National Bioscience Database Center (NBDC) at the Japan Science and Technology Agency (JST) at https://humandbs.biosciencedbc.jp/en/ through accession number hum0237.v1.gwas.v1. Caribbean Hispanics summary statistics from the Columbia University Study of Caribbean Hispanics and Late‐Onset Alzheimer's disease were accessed via application to dbGaP accession number phs000496.v1.p1. All codes developed and used are available at https://github.com/AndrewsLabUCSF/AFRPRS. Information regarding all the software and reference datasets used in the analysis can also be found in the repository.


Articles from Alzheimer's & Dementia are provided here courtesy of Wiley

RESOURCES