Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Jun 20;17:7806. doi: 10.1038/s41467-026-74461-7

PGS Browser: a public platform for personalized polygenic score analysis and interpretation

Nikita Kolosov 1,2,3, Mary P Reeve 3,4, Pietro Della Briotta Parolo 3, Mitja I Kurki 3,4,5; FinnGen, Vincent Llorens 3, Timo Petteri Sipila 3, Adam Herman 1, Ivan Molotkov 1,2,3, Mervi Aavikko 3, Samuli Ripatti 3, Aarno Palotie 3,4, Mark J Daly 3,4,5, Mykyta Artomov 1,2,3,
PMCID: PMC13439439  PMID: 42323305

Abstract

Polygenic scores (PGSs) quantify individual genetic susceptibility to complex diseases and can identify high-risk individuals well before clinical onset. Their clinical translation, however, requires population-based reference resources, standardized benchmarking, and accessible tools for translating individual scores into disease likelihood. In this article, we systematically evaluate 3168 PGS models, primarily from the PGS Catalog, in 473,681 FinnGen participants, placing all models on a common performance scale to enable cross-model and cross-trait comparison. For each PGS, we create ancestry-adjusted reference distributions, providing a biobank-scale resource for interpreting individual scores. We perform phenome-wide association studies for each PGS, identifying 439,070 significant phenotypic associations, demonstratin g that integrating multiple scores improves predictive performance for most complex diseases, and providing public access to 11 top-performing interactive time-to-event models. All resources are accessible through the PGS Browser (pgs.nchigm.org), which offers a population-aware framework for score interpretation and lays groundwork for the clinical application of PGSs.

Subject terms: Genetic association study, Predictive medicine, Prognostic markers


Benchmarking of 3,168 publicly available polygenic risk score models in 473,681 individuals provides access to performance metrics, PheWAS atlas and disease predictive models via the PGS browser (https://pgs.nchigm.org).

Introduction

Polygenic scores (PGSs) aggregate the effects of multiple genetic variants, derived from independent genome-wide association studies (GWAS), to estimate inherited susceptibility to complex traits. Each PGS model is defined by the specific variants it includes and corresponding effect sizes. Applied in a research context, PGSs have demonstrated substantial utility for risk stratification, disease prediction, and exploration of shared genetic etiology across phenotypes through phenome-wide association studies (PheWASs)16.

The Polygenic Score Catalog7 is a centralized repository of 3688 PGS models with associated metadata and contributor-supplied performance metrics, spanning 1450 traits. PGS Catalog has been pivotal in advancing the field by promoting standardized reporting practices8. However, most published models are accompanied by heterogeneous validation metrics derived from completely different cohorts, complicating model selection even within a single disorder. In the absence of standardized benchmarking that puts models on a common performance scale, choosing an optimal model for a given disease remains challenging. To date, large-scale external validation has been conducted only in a cohort of non-European ancestry9, but because most PGSs were originally developed in European cohorts, such assessments provide only partial insight.

Unlike individual pathogenic variants, which can be directly linked to disease risk, the absolute value of a PGS has no inherent meaning and must be interpreted relative to a population-based reference distribution. Such interpretation requires large-scale genetic datasets and careful adjustment for confounding factors, particularly population structure10,11. Pipelines have been developed to automate reproducible PGS calculation and reference-based interpretation12, but by default they rely on small public datasets such as the 1000 Genomes13 and Human Genome Diversity Project14, whose limited size and lack of clinical information constrain their value as population references. This gap underscores the need for advanced, biobank-linked, reference frameworks to enable more informative interpretation of individual PGSs.

The demand for such frameworks is further amplified by predictive gains delivered by both individual scores and predictive models that integrate multiple PGSs3,1517. However, these models remain scattered across studies and are rarely released in directly usable formats, creating substantial technical barriers that have prevented researchers and clinicians from integrating such technologies into practice18. Making such models accessible through public platforms could help bridge this gap. However, this is often not feasible in practice, as it would require sharing individual-level genetic and clinical data with the platform, which are typically restricted in biobanks and clinical settings, particularly in jurisdictions with strict data-sharing policies. Privacy-preserving frameworks that rely on aggregated data or enable preprocessing of sensitive information locally provide a practical alternative, allowing dissemination of predictive models without exposing individual-level data.

In this study, we systematically evaluated 3168 PGS models from the PGS Catalog across 4739 clinical traits in 473,681 FinnGen participants19. For each score, we computed a comprehensive set of performance metrics, enabling cross-model and cross-trait comparisons, identifying those with the greatest translational potential. Across 10,531 PheWAS, we discovered 439,070 significant PGS-phenotype associations and showed that combining multiple scores improves prediction for most complex diseases. We then generated ancestry-adjusted reference distributions for all 3168 scores from a biobank-scale cohort, making previously non-sharable resources publicly accessible. This framework allows researchers beyond FinnGen, including those without access to large biobanks, to contextualize individual scores against biobank-scale distributions. Building on these insights, we developed 11 top-performing time-to-event models for major disorders by integrating PGSs with demographic factors. We validated these models in non-White British cohorts of the UK Biobank (UKBB)20 and released them in interactive form, thereby lowering technical barriers and enabling interpretation of individual scores in the context of patient-specific characteristics. All resources, together with other experimental results, are freely available through the PGS Browser (https://pgs.nchigm.org)—a public platform for personalized polygenic score analysis and interpretation.

Results

PGS Catalog and FinnGen data harmonization

We aimed to evaluate the performance of polygenic score models from the PGS Catalog in distinguishing individuals with and without a lifetime history of each trait, using FinnGen as an independent validation cohort (Fig. 1a). FinnGen includes genetic and phenotypic data for 473,681 participants (Release R11), representing ~10% of Finland’s population, including 19,947 of non-Finnish ancestry19. Its large size, detailed longitudinal records, and minimal overlap with PGS development samples (in contrast to the substantial overlap in UK Biobank; Supplementary Fig. 1a, b) make FinnGen a well-suited resource for external validation of PGS Catalog models.

Fig. 1. Harmonization of PGS Catalog and FinnGen data.

Fig. 1

a Overview of the study workflow, including source of data, harmonization of PGS Catalog and FinnGen data, key analysis, and results of the study. b Number of retained PGS models as a function of the threshold applied to the proportion of matched variants. Vertical dashed line depicts the selected 75% threshold; Horizontal dashed line depicts the remaining 3025 PGS models after applying this threshold. c Distribution of variant counts across retained PGS models on the log10 scale; the dashed column predominantly represents models restricted to HapMap3 variants. d Distribution of retained PGS models by publication year. e Numbers of PGS models and matched FinnGen endpoints retained for different types of analyses. PGSC - PGS Catalog, FG - FinnGen.

We first assessed the compatibility of genetic variants between PGS models and the 21,342,234 variants available in FinnGen (Methods, PGS Catalog and FinnGen harmonization). Of the 3688 models in the PGS Catalog, 3025 (82%) achieved at least 75% variant overlap with FinnGen and were retained for downstream analyses (Fig. 1b). The median number of variants per model was 7372 (range: 1–10,318,272), and only around 9% of models were based on HapMap3 variants commonly used for PGS construction21 (Fig. 1c). Most analyzed models (73%) were published after 2022 (Fig. 1d).

We next harmonized phenotype definitions by manually matching PGS trait descriptions to FinnGen disease endpoints, which are derived from nationwide health registries and International Classification of Diseases (ICD) codes (Methods, FinnGen disease endpoints). Direct endpoint matches were identified for 1308 models, whereas 1717 models lacked a corresponding FinnGen endpoint, most often because they targeted quantitative traits that were not available in FinnGen at the time of analysis (e.g., brain volume and cystatin C level; Supplementary Fig. 1c).

To avoid inflated performance estimates, we performed multiple rounds of manual review to identify and annotate models with sample overlap between FinnGen and the corresponding PGS development sets (Supplementary Fig. 2). This process identified 460 models (15%) with overlap. These models were flagged in the results database and excluded from benchmarking and predictive modeling analyses (Fig. 1e). Of the remaining 2565 models without sample overlap, 779 had matched binary FinnGen endpoints. Because multiple PGS models often correspond to the same phenotype, these 779 models collectively covered 221 binary FinnGen endpoints (Supplementary Data 1). The remaining 1786 non-overlapping models lacked corresponding binary endpoints.

We calculated scores for all 3025 retained PGS Catalog models. Of these, 460 were excluded from benchmarking because of sample overlap with FinnGen, 779 non-overlapping models were linked to binary FinnGen endpoints and used for evaluation, and the remaining 1786 were retained for other experiments (Fig. 1e).

In addition to the PGS Catalog models, some analyses also included 143 PGS models developed by the FinnGen core team (see Methods, PGS Catalog, and FinnGen harmonization). These were processed identically to the PGS Catalog models, increasing the total number of analyzed models in relevant experiments to 3168.

Performance evaluation of PGS Catalog models in the FinnGen cohort

We analyzed 779 PGS models from the PGS Catalog matched to 221 FinnGen binary endpoints, omitting those with identified sample overlap. Each score was assigned to one of eight broad trait categories defined in the catalog (Fig. 2a)7. FinnGen endpoints were preprocessed to reflect disease prevalence in the Finnish population (Methods, FinnGen disease endpoints). For each model, we quantified discriminative ability using the unadjusted area under the receiver-operating characteristic curve (ROC AUC) and tested association with the corresponding endpoint using logistic regression, adjusting for sex, age, the first six principal components (PCs), and genotyping array (Fig. 2a).

Fig. 2. Evaluation of PGS Catalog models using the FinnGen cohort.

Fig. 2

a ROC AUC distributions for eight PGS categories derived from PGS Catalog annotations. Boxplots summarize the distribution of ROC AUC values across distinct PGS models within each category. Numbers above boxplots indicate the corresponding number of PGS models. The dashed box highlights the ROC AUC distribution for cancer PGSs. In panels a, b, boxplots show the median (notch), interquartile range (box bounds), and 1.5 × interquartile range (whiskers); points beyond the whiskers are shown as outliers. b ROC AUC distributions for the top ten cancer PGSs. Multiple distinct PGS models were available for each cancer type, with varying ROC AUC values. Boxplots summarize these distributions, and individual dots indicate the model achieving the highest ROC AUC for the selected cancer type. Numbers above boxplots indicate the number of PGS models available for the corresponding cancer type. c ROC AUC distribution for the 157 best-performing scores. Individual dots represent the percentage of scores exceeding the corresponding ROC AUC thresholds (dashed lines). d Top 15 PGSs across all categories ranked by unadjusted ROC AUC. Points represent mean ROC AUC values, and error bars indicate 95% confidence intervals estimated from 300 bootstrap resamples of the FinnGen dataset. Exact case and control counts for each endpoint are provided in Supplementary Data 2. Arrows here and below indicate PGS models developed by Privé F. et al. (PGP000263 [https://www.pgscatalog.org/publication/PGP000263/]). e Top 15 PGSs across all categories ranked by the proportion of liability-scale variability explained by the PGS, calculated as the difference between the full model (including the PGS) and the null model (excluding the PGS). Points represent liability-scale R2 values, and error bars indicate 95% confidence intervals estimated from 50 bootstrap resamples of the endpoint-specific FinnGen datasets. Exact case and control counts for each endpoint are provided in Supplementary Data 4. M.N. - malignant neoplasm.

To identify the most predictive PGSs for each endpoint, we selected the significantly associated model (p < 1.06 × 10−5) with the highest ROC AUC. Endpoints without significant associations were excluded. For example, among 13 models for testicular cancer, PGS00079622 was significantly associated (OR = 1.77, CI:1.63–1.92) and achieved the best ROC AUC of 0.69 (CI:0.67–0.71; Fig. 2b). ORs in this case and hereafter are reported per standard deviation unit increase in PGS. Other top-performing cancer scores included polycythemia vera (PGS00181023; AUC = 0.69; CI:0.68–0.71), prostate cancer (PGS00201623; AUC = 0.64; CI:0.64–0.65) and melanoma (PGS00076624; AUC = 0.63; CI:0.62–0.64). Across all traits, we identified 157 best-performing PGSs (Supplementary Fig. 3 and Supplementary Data 2). Only six (3.8%) achieved a ROC AUC of 0.70 or higher (Fig. 2c). The highest-performing scores included coeliac disease (PGS00185623; AUC = 0.84, CI:0.84–0.85), ankylosing spondylitis (PGS00187623; AUC = 0.82, CI:0.81–0.83) and disorders of iron metabolism (PGS00182323; AUC = 0.80, CI:0.78–0.83; Fig. 2d). Most of these top scores originated from Privé F. et al. 23, were based on UK Biobank summary statistics, and generated with penalized regression.

Despite similar overall performance in some cases, we observed substantial individual-level variability among models for the same disease. For example, two hypertension scores (PGS00270125 and PGS00204723) yielded comparable AUCs of 0.60 and 0.59 (Supplementary Fig. 4a, b) and relative risks (RR; 90th vs 50th%) of 1.36 and 1.32 but placed largely distinct individuals in the top 2.5% of the distribution, with only 30% overlap (Supplementary Fig. 4c–e). This discordance persisted even among top-performing scores. For instance, two coeliac disease models (PGS00210723 and PGS00185623) with AUCs of 0.83 and 0.84 and RR of 10.43 and 11.7 shared only 18% overlap in the bottom 2.5% quantile, which can be partly explained by the multimodal shape of the distributions (Supplementary Fig. 5). Thus, neither similar ROC AUC nor RR for the two PGS models for the same disease, guarantees concordant presence of the same individuals in tailed percentiles, which is consistent with previous reports26. The primary drivers of discordance were methodological differences in GWAS summary statistics and PGS construction methods selection, even when models were built from the same GWAS (Supplementary Fig. 6). These findings highlight the importance of considering individual-level variability when selecting or applying a PGS for risk stratification26,27.

To supplement the PGS model selection process by informing discordance and complementarity among scores, we provide a Spearman rank-correlation matrix for all 3025 scores (Data availability). Highly correlated PGSs tend to rank individuals similarly and are therefore largely redundant, in which case the model with the strongest predictive performance may be selected. In contrast, pairs of PGSs that both show strong performance for the same phenotype but low correlation are likely to capture partially distinct components of genetic liability, in this case models could potentially be used together to improve lifetime disease risk assessment.

For each best-performing model, we also estimated variance explained on the liability scale28. Notably, disorders of iron metabolism (e.g., haemochromatosis) reached R2 of 0.72 (CI: 0.58–0.87), which could be explained by high heritability of such conditions and, at the same time, partly, due to low prevalence of this endpoint in FinnGen (<0.1%). Other top-performing scores included Ankylosing spondylitis (0.38, CI:0.36–0.40), Coeliac disease (0.33, CI:0.32–0.34), Congenital deficiency of clotting factors (0.28, CI:0.20–0.35) and Type 1 diabetes (0.24, CI: 0.23–0.25) (Fig. 2e, Supplementary Fig. 7, and Supplementary Data 4).

In addition, following PGS reporting standards8, we calculated several complementary performance measures, including odds ratios (OR), true positive rate (TPR; recall), true negative rate (specificity), positive predictive value (PPV; precision), standardized PPV, false discovery rate (FDR), and relative and absolute risks. Because most of these metrics are threshold-dependent, the choice of cutoff typically depends on the specific scientific or clinical context. To facilitate standardized comparison between top-performing models, we used the 90th percentile of the PGS distribution as a reference threshold for classification measures, which is commonly applied in the PGS literature2932. We emphasize, however, that this cutoff is provided for comparative purposes only and is not intended to define clinical decision thresholds for the scores discussed in this study. To allow more flexible interpretation, we provide an interactive interface where some of these metrics can be explored across the full range of percentile thresholds (see Data Availability). Accordingly, we report these measures for the top-performing scores at the 90th percentile threshold. For example, coeliac disease demonstrated one of the highest TPR of 0.58 (90%) and relative risk of 3.23 (90 vs 50%), OR of 2.73 (Supplementary Fig. 8 and Supplementary Data 5). For a detailed description and interpretation of these metrics, see Supplementary Information, PGS evaluation measures.

Atlas of PGS-based phenome-wide association studies

We performed a phenome-wide association study (PheWAS) for 3168 polygenic scores across 4739 FinnGen endpoints using four complementary designs (Fig. 3a; Methods, Phenome-wide association study): intact (n = 3168), survival (n = 3168) which used Cox proportional-hazards models instead of logistic regression; exclusion (n = 1172), which excluded individuals with the target phenotype to mitigate phenotypic hitchhiking5; noMHC (n = 3023), which removed all variants from the major histocompatibility complex (MHC) locus. The exclusion and noMHC designs provided important validation and interpretive support for the primary intact analysis33.

Fig. 3. Atlas of PGS-based phenome-wide association studies.

Fig. 3

a Overview of the four distinct PheWAS designs used, together with the corresponding numbers of studies and experiment-wide significant associations. In the exclusion design, cases of the target phenotype were removed to mitigate confounding by phenotypic overlap. In the noMHC design, variants within the MHC region were removed from the PGS to evaluate associations driven exclusively by non-MHC variants. b The curve shows the proportion of PGSs with at least a given number of significant phenome-wide associations. The horizontal dashed line marks 50%, corresponding to the median proportion of PheWAS studies, and the vertical dashed line marks the median number of significant phenome-wide associations per PGS. c PheWAS results for the top 15 PGSs and Depression. Numbers and color gradients indicate the number of significant phenome-wide associations within each FinnGen endpoint category. d PheWAS results for the best-performing depression PGS (PGS000907). Each point represents the odds ratio estimated for one FinnGen endpoint, and horizontal error bars indicate 95% confidence intervals. Odds ratios and corresponding confidence intervals were obtained from adjusted logistic regression models as described in Methods. The asterisk marks the target Depression endpoint (F5_DEPRESSIO). The second point is absent because cases with F5_DEPRESSIO were removed in the exclusion design. e Combined PheWAS results for the depression endpoint, showing associated non-target PGSs. Each point represents the odds ratio estimated for one PGS, and horizontal error bars indicate 95% confidence intervals. Odds ratios were obtained from the same adjusted logistic regression model as in panel (d).

In the intact design, 3001 of 3168 PGSs had at least one phenome-wide significant association (Bonferroni threshold: p < 0.05/4739). The median number of associations per score was 81 (mean = 206, max = 1632; Fig. 3b). Endpoints were grouped using ICD-10-based tags curated by the FinnGen team (e.g., I9, circulatory system; Supplementary Data 6). Most scores showed their strongest associations within the target-trait category but frequently revealed additional associations in other domains (Fig. 3c). Collectively, these analyses produced the PGS-PheWAS atlas, cataloging all PGS-endpoint associations in FinnGen.

Importantly, such an atlas can be explored in two contrasting applications. The first, PGS-centric application, is intended to characterize a given PGS in terms of its target and secondary phenotypic associations. For example, when examining PGS of interest, this approach can identify other non-target endpoints that are associated with that score and partially share genetic etiology. The second, endpoint-centric application, is designed to identify PGSs associated with a given FinnGen endpoint. This application supports genetic risk-factor prioritization for the selected endpoint and can serve, for example, as a feature pre-selection step for multi-PGS disease prediction models. Thus, the same phenotype can be explored from PGS and FinnGen endpoint perspectives.

For example, the top-performing depression score (PGS00090734) was strongly associated with the FinnGen depression endpoint (F5_DEPRESSIO; OR = 1.35, CI:1.34–1.36) and showed 1256 additional significant associations (Fig. 3c, d). These included 116 digestive-system, 110 musculoskeletal and connective-tissue, and 82 mental-health endpoints. The most associated traits included bipolar disorder (OR = 1.40, CI:1.37–1.43), schizophrenia or delusion (OR = 1.31, CI:1.29–1.34) and panic disorder (OR = 1.39, CI:1.36–1.43). Significant, and biologically relevant, associations were also observed with hypothyroidism (OR = 1.24, CI:1.22–1.25), irritable bowel syndrome (OR = 1.19, CI:1.17–1.21), and cardiovascular disease (OR = 1.10, CI:1.09–1.11), consistent with earlier studies3539.

We next compared exclusion and noMHC with the intact design to assess whether secondary associations were driven by the primary trait or the MHC region33. For depression, MHC removal had no effect, as the PGS did not include MHC variants. Excluding depression cases reduced effect sizes but did not eliminate associations, confirming their robustness (Fig. 3d).

Alternatively, FinnGen depression endpoint (F5_DEPRESSIO), showed significant association with PGSs of schizophrenia (OR = 1.22, CI:1.20–1.23), esophagitis (OR = 1.10, CI:1.09–1.11), and hypertension (OR = 1.08, CI:1.07–1.09). Conversely, negative associations were observed for scores of income (OR = 0.89, CI:0.89–0.91), lifespan (OR = 0.94, CI:0.93–0.95) and height (OR = 0.95, CI:0.95–0.97; Fig. 3e).

We provide access to the full atlas of 10,531 PheWASs, covering all study designs, through an interactive web application (see Data availability).

Integration of multiple PGSs improves disease prediction

To determine which traits benefit from integrating multiple PGSs, it was first essential to exclude any undetected sample overlap between score-development cohorts and FinnGen40. We therefore restricted analyses to six publications, covering nearly half of the PGS Catalog, and manually confirmed that none of them included FinnGen samples (Fig. 4a). From these studies, we extracted 1514 models spanning 110 matched FinnGen disease endpoints.

Fig. 4. Overview of key steps in the multi-PGS modeling experiments.

Fig. 4

a Publication included in the downstream analysis. Six publications were retained, together accounting for approximately 50% of the retained PGS Catalog models and considered free of FinnGen sample overlap. b Schematic overview of the training and testing workflow for the three predictive model classes used in subsequent analyses. c Illustration of how relative feature importance was defined and visualized in the panels of Fig. 5.

FinnGen was divided into training and testing sets (Fig. 4b). In the training set, we fit three models: a null model (sex, age, and six genetic PCs), a best-PGS model (null + single top-performing PGS), and a multi-PGS model (null + an optimal combination of PGSs). Elastic-net logistic regression was used for binary outcomes and elastic-net Cox proportional hazards for time-to-event prediction (Methods, Risk prediction models). PGSs with non-zero elastic-net coefficients were considered informative (Fig. 4c).

For disease-status prediction, 80 of 110 endpoints (73%) showed significant ROC AUC improvement when the best single PGS was added to the null model (median ΔAUC = 0.015, interquartile range (IQR): 0.009–0.031; Fig. 5a, Supplementary Fig. 9, and Supplementary Data 7). The largest gains were observed for coeliac disease (ΔAUC = 0.28, CI: 0.27–0.30) and type 1 diabetes (ΔAUC = 0.19, CI: 0.18–0.19). Multi-PGS models further improved performance for 101 endpoints (92%; median ΔAUC = 0.025, IQR: 0.015–0.041), reaching ΔAUC of 0.29 (CI: 0.27–0.30) for coeliac disease and ΔAUC of 0.19 (CI: 0.18–0.19) for type 1 diabetes. Compared with best-PGS models, multi-PGS models showed significant gains for 87 endpoints (79%; median ΔAUC = 0.015, IQR: 0.009–0.022), with the largest for obesity and hyperalimentation (ΔAUC = 0.077, CI: 0.076–0.079), rheumatoid arthritis (ΔAUC = 0.05, CI: 0.04–0.06), and Crohn’s disease (ΔAUC = 0.04, CI:0.03–0.05; Supplementary Fig. 10). Many of these improvements were driven by both target and non-target scores. For instance, in inflammatory bowel disease (IBD), key contributors included scores for IBD and ulcerative colitis (target), as well as rheumatoid arthritis, polycythemia vera, and eosinophil count (non-target; Fig. 5b).

Fig. 5. Integration of multiple PGSs improves disease prediction.

Fig. 5

a Top 25 elastic-net logistic regression models for binary disease-status classification. The middle panel shows incremental ROC AUC, defined as the difference between the best single-PGS or multi-PGS model and the null model for each endpoint. Each point represents adjusted ROC AUC for one binary FinnGen endpoint, and error bars indicate 95% confidence intervals estimated from 150 bootstrap resamples of the endpoint-specific FinnGen test sets. Exact case and control counts for the training and test sets for each endpoint are provided in Supplementary Data 7. The side panels show the relative contributions of four feature types to the final multi-PGS model. The dashed box highlights the feature-importance profile for inflammatory bowel disease (IBD). b Relative feature importances for the optimal PGS combination selected for IBD. Bars show the contributions of individual PGSs retained in the final elastic-net model. c Top 25 CoxNet models for time-to-event prediction. The middle panel shows incremental adjusted time-dependent ROC AUC, defined as the difference between the best single-PGS or multi-PGS model and the null model across the 1–10-year prediction horizon. Each point represents one FinnGen endpoint, and error bars indicate 95% confidence intervals estimated from 150 bootstrap resamples of the endpoint-specific FinnGen test sets. Exact case and control counts for the training and test sets for each endpoint are provided in Supplementary Data 8. The dashed box highlights the feature-importance profile for substance abuse. d Relative feature importances for the optimal PGS combination selected for substance abuse. Bars show the contributions of individual PGSs retained in the final CoxNet model. Asterisks in panels (b, d) indicate scores with negative effects on the corresponding disease endpoint.

We applied the same approach to time-to-event prediction using elastic-net Cox models (Supplementary Data 8). Compared with a null model, the best PGS delivered a significant increase in time-dependent (td) ROC AUC for 29 of 110 endpoints (26%; median ΔAUC = 0.029, IQR:0.022–0.067). Multi-PGS models extended this to 42 endpoints (38%; median ΔAUC = 0.038, IQR:0.026–0.067), with the largest gains seen for coeliac disease (ΔAUC = 0.27, CI:0.25–0.29), obesity and related hyperalimentation disorders (ΔAUC = 0.15, CI:0.14–0.16), and type 1 diabetes (ΔAUC = 0.14, CI:0.09–0.15; Fig. 5c and Supplementary Fig. 11). Direct comparison of multi-PGS with best-PGS models revealed improvements only for 15 endpoints (13%; median ΔAUC = 0.034, IQR:0.026–0.045), including obesity (ΔAUC = 0.07, CI:0.06–0.09), sleep disorders (ΔAUC = 0.06, CI:0.04–0.08), and diabetic retinopathy (ΔAUC = 0.05, CI:0.03–0.07; Supplementary Fig. 12). As with binary outcomes, these gains were often driven by non-target scores. For example, prediction of substance-abuse onset benefited from integrating PGSs for smoking status, age at first sexual intercourse and neuroticism score (Fig. 5d).

For all 110 diseases, we provide lists of contributing PGSs and their elastic-net weights for both disease status and time-to-event models (Supplementary Data 9, 10).

Reference PGS distributions

To enable external use while preserving data privacy regulations, we parameterized each PGS distribution into shareable percentiles. These percentiles allow distributions to be reconstructed and applied as reference resources for risk stratification. To prevent population-structure bias in percentile interpretation, all 3168 scores were ancestry-adjusted following established methodology10. Each score was residualized on genetic principal components (PCs) to remove population-structure effects, making distributions comparable across ancestral groups (Fig. 6a, b). For each PGS, we first computed predicted scores based on genetic PCs alone. Several scores, such as those for height (r = 0.62; PGS0029896), intracranial aneurysm (r = 0.59; PGS00340741), and atopic dermatitis (r = 0.54; PGS00275542), showed a substantial correlation with population structure (Fig. 6c and Supplementary Data 11). Subtracting this predicted component from the raw scores eliminated ancestry-driven variation, leaving adjusted scores uncorrelated with PCs (Supplementary Fig. 13). For example, the raw breast cancer PGS (PGS00000143) overestimated percentile assignments in African-ancestry individuals (Fig. 6d), whereas the adjusted score produced comparable percentiles across ancestries (Fig. 6e).

Fig. 6. Ancestry-adjusted polygenic scores.

Fig. 6

a Projection of 473,681 FinnGen samples onto the first two principal components derived from the 1000 Genomes Project. Colored points represent 1000 G reference samples, and shaded areas indicate the distribution of FinnGen samples by inferred ancestry. b Overview of the PGS-ancestry-adjustment procedure. c Correlation between raw PGS values and predicted PGS values, representing the portion of genetic risk explained solely by population structure; the dashed brown line represents zero correlation. d Breast-cancer PGS distributions for Finnish and African-ancestry groups, highlighting how a given raw PGS value corresponds to different percentiles across populations due to population structure. e Adjusted breast cancer PGS distributions for the same populations. After adjustment, the distributions are better aligned, making percentile-based interpretation more comparable across ancestries.

We next recalculated ROC AUCs for ancestry-adjusted scores and compared them to those obtained for unadjusted PGSs (Supplementary Data 3). Only a small minority of scores showed significant differences, and the observed changes were modest in magnitude (Supplementary Fig. 14).

All of the ancestry-adjusted PGS reference distributions for 3168 scores are available via an interactive web application (see Data availability)

Public predictive models

To make FinnGen-derived predictive models applicable across major continental populations, we ancestry-adjusted PGSs and refit time-to-event models using the adjusted scores. This eliminated the need for explicit principal-component covariates, allowing the models to be applied externally with only minimal inputs: sex, current age, and the individual’s PGS percentile, which can be computed locally (Supplementary Information, Public predictive models). To reduce identifiability risk, we used percentiles (an aggregated measure) rather than raw individual-level values. Twenty-two best-PGS models met performance criteria—each achieved a time-dependent ROC AUC above 0.65 in internal FinnGen testing, showed excellent calibration (D-calibration44 p > 0.99), and improved ΔAUC by at least 0.01 over a null model (Supplementary Fig. 15). These models were then validated in UK Biobank Non-White British participants (n = 78,336) (Supplementary Fig. 16; Supplementary Information, UK Biobank cohort).

We first applied ancestry adjustment to the 22 PGSs evaluated in the UK Biobank validation cohort using parameters estimated in FinnGen and then assessed whether adjustment for population structure was successful (R2 < 0.05; Supplementary Fig. 17). Of the 22 models considered, only the skin cancer score remained significantly correlated with principal components after adjustment (R2 = 0.192, 95% CI: 0.187–0.197) and was therefore excluded from the public release. We then evaluated the performance of the corresponding PGSs and CoxNet models across ancestry groups (Supplementary Figs. 18 and 19). Eleven CoxNet models achieved a tdROC AUC greater than 0.6 in at least three of the six ancestry groups, while also using PGSs with ROC AUC values above 0.6. The best-performing models included palmar fascial fibromatosis (tdROC AUC = 0.84), malignant neoplasm of prostate (tdROC AUC = 0.81), and atrial fibrillation/flutter (tdROC AUC = 0.8; Supplementary Fig. 20).

Polygenic score browser (PGS browser)

All PGS models matched to the FinnGen cohort (n = 3168), along with results from phenome-wide association studies (n = 10,540), PGS Catalog model evaluations, interactive ancestry-adjusted reference distributions (n = 3168), and time-to-event predictive models (n = 11) were made accessible through a dedicated web application—the PGS Browser. (Fig. 7; Data availability). Additionally, we provide pgsb-cli—a Docker-based command-line tool that runs locally and does not share any individual data with the external server (Supplementary Information, PGS Browser command-line tool; Code availability). Given a PGS model downloaded from the PGS Browser and a patient’s genotype data, it automatically matches variants, computes the ancestry-adjusted PGS, and reports standardized values and percentiles. Only these aggregated percentiles need to be sent to the web interface for predictive models, keeping all raw genetic data on the user’s machine (Fig. 7).

Fig. 7. Polygenic score browser (PGS browser).

Fig. 7

An overview of the data types and functionalities available in the PGS Browser: the top level illustrates the available data and predictive models, the middle level highlights elements of the graphical interface, and the bottom level shows the schematic results displayed to the user; pgsb-cli—PGS Browser command-line interface.

Discussion

Our study provides comprehensive biobank-scale benchmarking of the PGS Catalog and introduces open-access reference resources, including ancestry-adjusted PGS distributions, publicly available time-to-event models, and a large atlas of PGS-based phenome-wide studies, to accelerate prospective clinical translation and research use of polygenic scores.

Systematic evaluation of most publicly available PGS models within a single biobank-scale cohort enabled reliable cross-model comparison on a common performance scale. This analysis revealed the highest-performing scores for target traits, identified optimal score combinations for numerous diseases, and highlighted key technical and practical challenges.

One key challenge arises from the scale of the analysis. Large-scale PGS evaluation requires extensive automation, which depends on standardized formats and robust metadata. The PGS Catalog largely meets these needs—its submission standards and detailed annotations make most scores accessible and reusable. However, our analysis revealed a recurring issue, that is, hidden sample overlap between source GWAS datasets and FinnGen, leading to inflated performance estimates. In many cases where the target-trait ROC AUC exceeded 0.65, we confirmed unreported inclusion of FinnGen or its legacy cohorts in the source GWAS. These findings show that, even under current Catalog guidelines, cohort provenance reporting, particularly for scores derived from meta-analyses, remains incomplete.

Beyond this reporting bias, benchmarking demonstrated that only a small fraction of models achieved ROC AUC >0.7, even in a European cohort. Most of these top performers were immune-related scores, often derived from UK Biobank summary statistics. This is probably explainable by the broad phenotype coverage of a biobank and predominantly European ancestry of evaluation datasets. Such ancestry composition also represents an important limitation of the current study. The modest number of non-European samples in FinnGen restricts our ability to reliably assess predictive performance and guide model selection across other continental populations. Additionally, while ancestry adjustment using a fixed PC space derived from diverse reference panels (e.g., the 1000 G) improves cross-ancestry comparability of PGS percentiles and helps reduce ancestry-driven artifacts in downstream predictive models, such calibration does not address the fundamental limitation of reduced predictive accuracy when PGS models derived primarily from European cohorts are applied to other populations. Extending benchmarking and model optimization to larger and more ancestrally diverse cohorts will therefore be an important direction for future work.

From a translational perspective, we observed substantial variability in individual risk estimates produced by different PGSs for the same trait. Such heterogeneity complicates individual-level prediction for selected disorders, undermining confidence in risk stratification27. This variation partly arises from methodological factors, including the source GWAS, scoring algorithm, development cohort, and parameter choices, as well as the stochastic nature of certain inference methods (e.g., PRS-CS45 and LDPRed246). Even when models achieve comparable overall performance, these differences can yield discordant predictions for the same individual. Transitioning PGSs into clinical use will therefore require standardized benchmarking, transparent methodological reporting, and explicit consideration of individual-level variability. To support informed model selection, we provide a rank-correlation matrix for all scores, offering a reference for the expected degree of discordance within each trait.

Combining target-trait PGSs with scores for genetically related, non-target traits improved predictive performance for the majority of diseases examined in binary classification settings and over one-third in time-to-event predictions. Our phenome-wide association analyses identified numerous cross-trait associations, suggesting that shared genetic architecture potentially can be leveraged to improve disease prediction. Across 110 diseases, multi-PGS models consistently outperformed corresponding single-score models, with notable AUC gains for obesity, rheumatoid arthritis, Crohn’s disease, and several other conditions. As large-scale PGS pipelines12 and integrative modeling frameworks16,47 continue to advance, such cross-trait, multi-score approaches are likely to become increasingly powerful and operationally feasible, offering a way to explain some of the currently unexplained phenotypic variability.

Finally, our work addresses another barrier to PGS translation: the lack of large-scale, public, population-based reference frameworks for score interpretation. Although many PGS models are publicly available, the reference distributions required to interpret individual scores typically rely on small WGS datasets or large biobank datasets that are subject to strict data-access restrictions. Using data from nearly half a million FinnGen participants, we generated ancestry-adjusted reference distributions for all 3168 scores and 11 public time-to-event prediction models. Importantly, these reference distributions can be shared as aggregated summary resources without exposing individual-level data. By making these reference distributions and predictive models accessible through the PGS Browser platform, we provide a framework that enables researchers without direct access to large biobank cohorts to interpret individual PGS values. All of these results are available through the PGS Browser—a public platform that aims to democratize polygenic risk assessment, introduces a framework to communicate biobank-linked predictive models, and establishes essential groundwork for the responsible integration of PGS into clinical care.

Methods

Study cohort

For the majority of experiments, we used the FinnGen cohort (Release 11). The dataset includes 473,682 Finnish individuals, containing both genetic and clinical data19. Experimental findings, predictive models, and PGS distributions are shared under the authorization of the FinnGen Scientific Committee (project number F_2020_073). All participants provided written informed consent. For external validation of predictive models, we used the UK Biobank. We restricted analyses to 78,336 participants classified by the UK Biobank as non-White British (Data-Field 22006). All participants provided written informed consent.

Genotyping and imputation

FinnGen participants were genotyped using Illumina and Affymetrix arrays (Illumina and Thermo Fisher Scientific). Genotype data underwent standard sample- and variant-level quality control, followed by phasing, imputation, and post-imputation filtering according to the FinnGen analysis pipeline. Protocol details are available at 10.17504/protocols.io.nmndc5e. UK Biobank participants were genotyped using the UK BiLEVE or UK Biobank Axiom arrays. We used the centrally processed imputed genotype release provided by UK Biobank after cohort-wide quality control, phasing, and imputation20.

PGS catalog and FinnGen harmonization

Variants from each of the 3688 PGS models were downloaded, matched to FinnGen variant data, and aggregated into effect-size matrices using PGS Catalog utilities (see Supplementary Information, PGS Models Processing). We evaluated all models in their original form, intentionally avoiding substitution of unmatched variants with proxies in strong linkage disequilibrium. This approach minimized the introduction of additional variability and enabled direct assessment of the original model performance. Some analyses additionally included 143 PGS models developed by the FinnGen core team (https://github.com/FINNGEN/CS-PRS-pipeline), which were preprocessed in the same way as the PGS Catalog models.

PGS models developed from cohorts that potentially overlapped with FinnGen or its legacy cohorts (as reported in the source GWAS) were flagged accordingly (see Supplementary Data 12). Individual PGS scores were calculated using PLINK 2.0’s --score function48. Reported traits for each model were matched to FinnGen endpoint descriptions via regular expressions, and all matches were manually confirmed.

Ancestry prediction

FinnGen includes 453,734 participants of Finnish ancestry and 19,947 participants of non-Finnish ancestry, often treated as PCA outliers and excluded from prior analyses. To characterize them, we projected all individuals into the 1000 Genomes PCA space and used a gradient-boosting classifier trained on the first three PCs (tenfold cross-validation) to assign continental ancestry. Samples with more than 65% of predicted probability were labeled accordingly; all others were marked Others (see Supplementary Information, Ancestry prediction in FinnGen cohort).

PGS-ancestry adjustment

Because raw PGS values often differ across populations due to allele frequency differences, linkage disequilibrium structure, and other technical factors, raw PGS differences may partially reflect ancestry variation rather than relative genetic risk within a population. To make PGS percentiles comparable between populations, we implemented an ancestry-adjustment procedure following Hao et al10.

PGSadjusted=PGSrawPGSpredicted 1

Note that, because the reference PC basis is defined using the 1000 G dataset, the same projection and adjustment procedure applied to FinnGen samples can also be applied to external datasets in the same way. To facilitate reproducibility, we provide an open-source command-line tool, pgsb-cli (see Code availability), which automates projection into the reference PC space and the subsequent adjustment steps. Importantly, this procedure does not address differences in predictive accuracy of PGS models across ancestry groups, rather it reduces ancestry-driven artifacts in downstream predictive models and makes percentiles comparable between different ancestries.

FinnGen disease endpoints

From FinnGen Release 11, we utilized 4739 endpoints, each with more than six cases. These endpoints represent binary features indicating the presence or absence of specific conditions, with diagnoses made on the basis of the International Classification of Diseases (ICD-8, ICD-9, and ICD-10) criteria. Detailed inclusion/exclusion criteria and case/control definitions are available via the Risteys portal (https://risteys.finngen.fi/). All endpoints were grouped into 38 categories, with corresponding tags (e.g., I9, F5), as defined by the FinnGen core team (Supplementary Data 6). For PGS evaluation and predictive modeling, samples excluded from the control set were reintroduced to better reflect the control population and disease prevalence in the Finnish population. For PheWAS experiments, we kept the endpoints intact.

Phenome-wide association study

For each PheWAS, we fitted a logistic-regression model with LDLT decomposition using the fastglm R package to speed up computations49. We adjusted the relationship between the standardized PGS and each endpoint for sex, age at the end of follow-up, PCs 1–6, and six genotyping-array binary variables:

logitP=β0+βPGSPGS+βAgeAge+βSexSex+i=16βPCiPCi+j=16βArrjArrj 2

where the primary focus was βPGS, with other factors included to address potential confounding. Additionally, we computed unadjusted ROC AUC for each PGS-endpoint pair using the bigsnpr R package50. For gender-specific traits, we excluded the variable sex. To handle multiple comparisons for assessment of associations for a single PGS, we set a Bonferroni threshold at p value <1.06 × 10−5 (0.05/4,739). When determining the experiment-wide number of significant associations, we have used Bonferroni adjustment for the total number of comparisons − 0.05/(4739 endpoints × 3168 PGS models). In addition, we performed three alternative PheWAS designs51: (1) exclusion—in which individuals with the target phenotype were removed to test whether the secondary associations remain after removing the effect of the primary phenotype5; (2) noMHC—in which variants in the MHC region (chr6: 28.5–33.5 Mb, GRCh38) were excluded and scores were recalculated before conducting a new PheWAS on the altered scores. Comparing these results with the original design (unaltered scores and binary endpoints) allows us to identify associations driven solely by the MHC locus or phenotypic hitchhiking; and (3) survival—in which logistic regression was replaced with Cox proportional-hazards models (survival R package) to use time-to-event as the response variable. Additionally, for survival model we substituted age at the end of follow-up by baseline age.

Associations involving PGSs with sample overlap with Finnish cohorts were excluded from the main results but are provided, with full annotation, in the public data release (see Data availability).

Risk prediction models

For binary classification, we used logistic regression with an elastic-net penalty, implemented in the scikit-learn Python package (v1.5.1). Regularization strength (C) and the elastic-net mixing parameter (L1-L2 ratio) were optimized through grid-search cross-validation (cv = 5), using ROC AUC as the scoring metric. The total number of samples used for training (70%) and testing (30%) for each endpoint are reported in a corresponding table (Supplementary Data 7).

For time-to-event prediction, we utilized Cox’s proportional-hazards model with an elastic-net penalty, implemented in the scikit-survival Python package (v0.23.2). Follow-up years were used as the time scale, calculated by subtracting baseline age from either event age or the end of follow-up. For each trait, individuals with baseline ages greater than event ages were excluded, which resulted in a different number of samples used for training (70%) and testing (30%) in each individual model. The number of samples used in the training and test sets for each endpoint are reported in Supplementary Data 8. Regularization strength (alpha) and the elastic-net mixing parameter (l1 ratio) were optimized using grid-search cross-validation (cv = 5), with mean time-dependent ROC AUC (1–10 years) as the scoring metric.

Note that both approaches utilized an elastic-net regularization to identify an optimal combination of PGSs in multi-PGS models. In each case, hyperparameters (the L1-L2 mixing parameter and regularization strength) were tuned via cross-validation. This framework is well-suited for settings with many correlated PGSs. The ridge (L2) component stabilizes coefficient estimates for correlated variables, while the lasso (L1) component encourages sparsity. As a result, groups of correlated PGSs may be retained together with reduced coefficients when this improves predictive performance, while redundant predictors may be excluded when a sparser solution performs better. In this context, we define an optimal set as the combination of PGSs that maximizes predictive performance under the specified regularization and cross-validation scheme.

Importantly, PGS-derived risk estimates derived from these models represent probabilistic risk stratification rather than deterministic predictions for a given individual. Uncertainty arises from multiple sources, including finite GWAS discovery sample sizes, the probabilistic nature of PGS methods and multi-PGS model construction. As a result, individual-level predictions should be interpreted as risk estimates within a population context rather than precise estimates of an individual’s absolute disease probability.

Feature importances were derived from the weights of the elastic-net models. As all features were standardized before model fitting, their importances were directly comparable (Supplementary Information, Disease risk prediction models).

To compare performance differences between sets of models (e.g., null vs. multi), we used a paired Wilcoxon signed-rank test from the R stats package (v4.4.2) and paired bootstrap test for within endpoint comparisons.

Graphical interface

The web application was primarily created using the Shiny R package (v1.10.0). Unless specified otherwise, all analyses and software development were conducted using R 4.3.2.

Ethics

Individuals in FinnGen provided informed consent for biobank research, based on the Finnish Biobank Act. Alternatively, separate research cohorts, collected prior to the Finnish Biobank Act came into effect (in September 2013) and the start of FinnGen (August 2017), were collected based on study-specific consents and later transferred to the Finnish biobanks after approval by Fimea (Finnish Medicines Agency) and the National Supervisory Authority for Welfare and Health. Recruitment protocols followed the biobank protocols approved by Fimea. The Coordinating Ethics Committee of the Hospital District of Helsinki and Uusimaa (HUS) statement number for the FinnGen study is Nr HUS/990/2017. Further details on permit and biobank decision numbers are available in the FinnGen Ethics Statement in the Supplementary Information.

All UK Biobank participants also provided informed consent, and participation was voluntary. UK Biobank has ethical approval from the North West Multi-center Research Ethics Committee as a Research Tissue Bank (REC reference 16/NW/0274), under which most approved studies can be conducted without separate project-specific ethical approval. The present analyses were performed under UK Biobank Application 88907.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

41467_2026_74461_MOESM2_ESM.pdf (185.4KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1-13 (1.1MB, xlsx)
Reporting Summary (3.1MB, pdf)

Source data

Source Data (677.5KB, xlsx)

Acknowledgements

We thank the participants and investigators of the FinnGen study and the UK Biobank. Following biobanks are acknowledged for delivering biobank samples to FinnGen: Auria Biobank (www.auria.fi/biopankki), THL Biobank (www.thl.fi/biobank), Helsinki Biobank (www.helsinginbiopankki.fi), Biobank Borealis of Northern Finland (https://www.ppshp.fi/Tutkimus-ja-opetus/Biopankki/Pages/Biobank-Borealis-briefly-in-English.aspx), Finnish Clinical Biobank Tampere (www.tays.fi/en-US/Research_and_development/Finnish_Clinical_Biobank_Tampere), Biobank of Eastern Finland (www.ita-suomenbiopankki.fi/en), Central Finland Biobank (www.ksshp.fi/fi-FI/Potilaalle/Biopankki), Finnish Red Cross Blood Service Biobank (www.veripalvelu.fi/verenluovutus/biopankkitoiminta), Terveystalo Biobank (www.terveystalo.com/fi/Yritystietoa/Terveystalo-Biopankki/Biopankki/), and Arctic Biobank (https://www.oulu.fi/en/university/faculties-and-units/faculty-medicine/northern-finland-birth-cohorts-and-arctic-biobank). All Finnish Biobanks are members of BBMRI.fi infrastructure (www.bbmri.fi). Finnish Biobank Cooperative - FINBB is the coordinator of BBMRI-ERIC operations in Finland. The Finnish biobank data can be accessed through the Fingenious® services (https://site.fingenious.fi/en/) managed by FINBB. N.K. and M.A. would like to thank the High-Performance Computing core at Abigail Wexner Research Institute for their support of the computational infrastructure for the project.

Author contributions

Study design and conceptualization—N.K., M.P.R., M.J.D., M.A. Data analysis—N.K., M.P.R., P.D.B.P., M.I.K., V.L., T.P.S., I.M. and M.A. Statistical model design and manuscript writing—N.K. and M.A. Web application design—N.K. Web application hosting support—A.H. Funding acquisition—M.A., M.J.D., A.P. and S.R. Project supervision—M.A. Manuscript editing and approval—all authors

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.

Funding

N.K., I.M., and M.A. disclose support for this work from the Aging Biology Foundation. The FinnGen project is funded by two grants from Business Finland (HUS 4685/31/2016 and UH 4386/31/2016) and by the following industry partners: AbbVie Inc., AstraZeneca UK Ltd, Biogen MA Inc., Bristol Myers Squibb (including Celgene Corporation and Celgene International II Sàrl), Genentech Inc., Merck Sharp & Dohme LLC, Pfizer Inc., GlaxoSmithKline Intellectual Property Development Ltd., Sanofi US Services Inc., Maze Therapeutics Inc., Janssen Biotech Inc., Novartis Pharma AG, and Boehringer Ingelheim International GmbH. UK Biobank is funded by the Medical Research Council, Wellcome, the Department of Health, the Scottish Government, the Welsh Assembly Government, the British Heart Foundation, Cancer Research UK, Diabetes UK, the National Institute for Health and Care Research (NIHR), and the Northwest Regional Development Agency.

Data availability

The individual-level FinnGen data used in this study were obtained under FinnGen Proposal Number F_2020_073 (https://www.finngen.fi/en). The FinnGen data are available under restricted access because they contain sensitive genetic and health information protected by participant consent, ethical approval, and Finnish and European data protection regulations, including the General Data Protection Regulation (GDPR). Thus, the raw FinnGen data are protected and are not publicly available due to data privacy laws. Details on the data-access restrictions and policies are listed within the FinnGen flagship publication - Kurki et al., Nature, 2023. This research has been conducted using the UK Biobank Resource under Application Number 88907 (https://www.ukbiobank.ac.uk/). The individual-level UK Biobank data used in this study are available under restricted access because they contain sensitive participant-level genetic and phenotypic information and are governed by UK Biobank access policies. Access can be obtained by bona fide researchers through an application to the UK Biobank. The raw UK Biobank data are protected and are not publicly available due to participant privacy and data governance restrictions. All PGS models analyzed in this study are publicly available through the PGS Catalog [https://www.pgscatalog.org/] (PGS Catalog) using their corresponding score identifiers (e.g., PGS000907 and are also accessible through the PGS Browser [https://pgs.nchigm.org] (PGS Browser). The processed data generated in this study, including matched PGS models, corresponding metadata, PheWAS results, predictive models, and tutorials, are publicly available through the PGS Browser [https://pgs.nchigm.org] (PGS Browser). The provisional patent (USPTO serial no. 63/716,306) limits the ability to use the PGS browser and predictive models for profit. For commercial/for profit usage of the PGS browser, please reach out to the corresponding author for licensing. Source data are provided with this paper.

Code availability

The pgsb-cli software is available from GitHub (https://github.com/ArtomovLab/PGS_Browser) and archived at Zenodo (10.5281/zenodo.20060604).

Competing interests

A provisional patent has been filed with respect to the findings described in this manuscript (USPTO serial no. 63/716,306). Mark J. Daly is a founder of Maze Therapeutics. The remaining authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

A list of authors and their affiliations appears at the end of the paper.

A full list of members and their affiliations appears in the Supplementary Information.

Contributor Information

Mykyta Artomov, Email: mykyta.artomov@nationwidechildrens.org.

FinnGen:

Nikita Kolosov, Mary P. Reeve, Mitja I. Kurki, Pietro Della Briotta Parolo, Vincent Llorens, Timo Petteri Sipila, Ivan Molotkov, Mervi Aavikko, Samuli Ripatti, Aarno Palotie, Mark J. Daly, and Mykyta Artomov

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-74461-7.

References

  • 1.Leppert, B. et al. A cross-disorder PRS-pheWAS of 5 major psychiatric disorders in UK Biobank. PLoS Genet.16, e1008185 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Mars, N. et al. Polygenic and clinical risk scores and their impact on age at onset and prediction of cardiometabolic diseases and common cancers. Nat. Med.26, 549–557 (2020). [DOI] [PubMed] [Google Scholar]
  • 3.Patel, A. P. et al. A multi-ancestry polygenic risk score improves risk prediction for coronary artery disease. Nat. Med.29, 1793–1803 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Richardson, T. G., Harrison, S., Hemani, G. & Davey Smith, G. An atlas of polygenic risk score associations to highlight putative causal relationships across the human phenome. Elife8, e43657 (2019). [DOI] [PMC free article] [PubMed]
  • 5.Fritsche, L. G. et al. Cancer PRSweb: an online repository with polygenic risk scores for major cancer traits and their evaluation in two independent biobanks. Am. J. Hum. Genet.107, 815–836 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ma, Y., Patil, S., Zhou, X., Mukherjee, B. & Fritsche, L. G. ExPRSweb: an online repository with polygenic risk scores for common health-related exposures. Am. J. Hum. Genet.109, 1742–1760 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Lambert, S. A. et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nat. Genet.53, 420–425 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wand, H. et al. Improving reporting standards for polygenic scores in risk prediction studies. Nature591, 211–219 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Sun, T.-H. et al. Utility of polygenic scores across diverse diseases in a hospital cohort for predictive modeling. Nat. Commun.15, 1–12 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Hao, L. et al. Development of a clinical polygenic risk score assay and reporting workflow. Nat. Med.28, 1006–1013 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kachuri, L. et al. Principles and methods for transferring polygenic risk scores across global populations. Nat. Rev. Genet.25, 8–25 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lambert, S. A. et al. Enhancing the polygenic score catalog with tools for score calculation and ancestry normalization. Nat. Genet.56, 1989–1994 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Auton, A. et al. A global reference for human genetic variation. Nature526, 68–74 (2015). [DOI] [PMC free article] [PubMed]
  • 14.Bergström, A. et al. Insights into human genetic variation and population history from 929 diverse genomes. Science367, eaay5012 (2020). [DOI] [PMC free article] [PubMed]
  • 15.Khera, A. V. et al. Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations. Nat. Genet.50, 1219–1224 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Misra, A. et al. Instability of high polygenic risk classification and mitigation by integrative scoring. Nat. Commun.16, 1584 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Jung, H. et al. Integration of risk factor polygenic risk score with disease polygenic risk score for disease prediction. Commun. Biol.7, 1–13 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Roberts, M. C., Holt, K. E., Del Fiol, G., Baccarelli, A. A. & Allen, C. G. Precision public health in the era of genomics and big data. Nat. Med.30, 1865–1873 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature613, 508–518 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature562, 203–209 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Privé, F., Arbel, J., Aschard, H. & Vilhjálmsson, B. J. Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores. HGG Adv.3, 100136 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Kachuri, L. et al. Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction. Nat. Commun.11, 6084 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Privé, F. et al. Portability of 245 polygenic scores when derived from the UK Biobank and applied to 9 ancestry groups from the same cohort. Am. J. Hum. Genet.109, 373 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Bakshi, A. et al. Genomic risk score for melanoma in a prospective study of older individuals. J. Natl. Cancer Inst.113, 1379–1385 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Weissbrod, O. et al. Leveraging fine-mapping and multipopulation training data to improve cross-population polygenic risk scores. Nat. Genet.54, 450–458 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Abramowitz, S. A. et al. Evaluating performance and agreement of coronary heart disease polygenic risk scores. JAMA10.1001/jama.2024.23784 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ding, Y. et al. Large uncertainty in individual polygenic risk score estimation impacts PRS-based risk stratification. Nat. Genet.54, 30–39 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Lee, S. H., Goddard, M. E., Wray, N. R. & Visscher, P. M. A better coefficient of determination for genetic profile analysis. Genet. Epidemiol.36, 214–224 (2012). [DOI] [PubMed] [Google Scholar]
  • 29.Middha, P. et al. Polygenic risk score for ulcerative colitis predicts immune checkpoint inhibitor-mediated colitis. Nat. Commun.15, 2568 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.He, Y.-Q. et al. A polygenic risk score for nasopharyngeal carcinoma shows potential for risk stratification and personalized screening. Nat. Commun.13, 1966 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Tamlander, M. et al. Genome-wide polygenic risk scores for colorectal cancer have implications for risk-based screening. Br. J. Cancer130, 651–659 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Moreno-Grau, S. et al. Polygenic risk score portability for common diseases across genetically diverse populations. Hum. Genomics18, 93 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Reeve, M. et al. Autoimmune hypothyroidism GWAS reveals independent autoimmune and thyroid-specific contributions and an inverse relation with cancer risk. Res. Sq.10.21203/rs.3.rs-4626646/v1 (2024).
  • 34.Campos, A. I. et al. Understanding genetic risk factors for common side effects of antidepressant medications. Commun. Med.1, 45 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Ma, Y., Wang, M. & Zhang, Z. The association between depression and thyroid function. Front. Endocrinol.15, 1454744 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Hage, M. P. & Azar, S. T. The link between thyroid function and depression. J. Thyroid Res.2012, 590648 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Dhar, A. K. & Barton, D. A. Depression and the link with cardiovascular disease. Front. Psychiatry7, 33 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Kwapong, Y. A. et al. Association of depression and poor mental health with cardiovascular disease and suboptimal cardiovascular health among young adults in the United States. J. Am. Heart Assoc.12, e028332 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Piovani, D., Armuzzi, A. & Bonovas, S. Association of depression with incident inflammatory bowel diseases: a systematic review and meta-analysis. Inflamm. Bowel Dis.30, 573–584 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Ellis, C. A. et al. Inflation of polygenic risk scores caused by sample overlap and relatedness: examples of a major risk of bias. Am. J. Hum. Genet.111, 1805–1809 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Bakker, M. K. et al. Genetic risk score for intracranial aneurysms: prediction of subarachnoid hemorrhage and role in clinical heterogeneity. Stroke54, 810–818 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Mars, N. et al. Systematic comparison of family history and polygenic risk across 24 common diseases. Am. J. Hum. Genet.109, 2152–2162 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Michailidou, K. et al. Large-scale genotyping identifies 41 new loci associated with breast cancer risk. Nat. Genet.45, 353–61 (2013). 361e1–2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Haider, H., Hoehn, B., Davis, S. & Greiner, R. Effective ways to build and evaluate individual survival distributions. J. Mach. Learn. Res. 21, 1–63 (2020).
  • 45.Ge, T., Chen, C.-Y., Ni, Y., Feng, Y.-C. A. & Smoller, J. W. Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat. Commun.10, 1776 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Privé, F., Arbel, J. & Vilhjálmsson, B. J. LDpred2: better, faster, stronger. Bioinformatics36, 5424–5431 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Truong, B. et al. Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases. Cell Genom.4, 100523 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Purcell, S. et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet.81, 559–575 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Marschner, I. Glm2: Fitting generalized linear models with convergence problems. R J3, 12 (2011). [Google Scholar]
  • 50.Privé, F., Aschard, H., Ziyatdinov, A. & Blum, M. G. B. Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr. Bioinformatics34, 2781–2787 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Reeve, M. P. et al. Oral and non-oral lichen planus show genetic heterogeneity and differential risk for autoimmune disease and oral cancer. Am. J. Hum. Genet.111, 1047–1060 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

41467_2026_74461_MOESM2_ESM.pdf (185.4KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1-13 (1.1MB, xlsx)
Reporting Summary (3.1MB, pdf)
Source Data (677.5KB, xlsx)

Data Availability Statement

The individual-level FinnGen data used in this study were obtained under FinnGen Proposal Number F_2020_073 (https://www.finngen.fi/en). The FinnGen data are available under restricted access because they contain sensitive genetic and health information protected by participant consent, ethical approval, and Finnish and European data protection regulations, including the General Data Protection Regulation (GDPR). Thus, the raw FinnGen data are protected and are not publicly available due to data privacy laws. Details on the data-access restrictions and policies are listed within the FinnGen flagship publication - Kurki et al., Nature, 2023. This research has been conducted using the UK Biobank Resource under Application Number 88907 (https://www.ukbiobank.ac.uk/). The individual-level UK Biobank data used in this study are available under restricted access because they contain sensitive participant-level genetic and phenotypic information and are governed by UK Biobank access policies. Access can be obtained by bona fide researchers through an application to the UK Biobank. The raw UK Biobank data are protected and are not publicly available due to participant privacy and data governance restrictions. All PGS models analyzed in this study are publicly available through the PGS Catalog [https://www.pgscatalog.org/] (PGS Catalog) using their corresponding score identifiers (e.g., PGS000907 and are also accessible through the PGS Browser [https://pgs.nchigm.org] (PGS Browser). The processed data generated in this study, including matched PGS models, corresponding metadata, PheWAS results, predictive models, and tutorials, are publicly available through the PGS Browser [https://pgs.nchigm.org] (PGS Browser). The provisional patent (USPTO serial no. 63/716,306) limits the ability to use the PGS browser and predictive models for profit. For commercial/for profit usage of the PGS browser, please reach out to the corresponding author for licensing. Source data are provided with this paper.

The pgsb-cli software is available from GitHub (https://github.com/ArtomovLab/PGS_Browser) and archived at Zenodo (10.5281/zenodo.20060604).


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES