Skip to main content
Medizinische Genetik logoLink to Medizinische Genetik
. 2026 Jul 8;38(3):157–168. doi: 10.1515/medgen-2026-3009

An introduction to polygenic scores – methodological basics and recent advances

Hannah Klinkhammer 1, Andreas Mayr 2, Carlo Maj 3,✉
PMCID: PMC13340550  PMID: 42416883

Abstract

Polygenic scores (PGS) allow the estimation of genetic predisposition to complex diseases and traits. Based on results from genome-wide association studies (GWAS) and data from large deeply phenotyped population biobanks, numerous PGS have been developed in recent years. These scores summarize the combined effects of many genetic variants and can support risk stratification for multifactorial diseases based on an individual’s genetic susceptibility. In this review, we introduce the basic principles and methods of PGS and explain how they are generated and applied. We outline their potential translational role, particularly in the context of precision medicine and personalized risk stratification for preventive measures, in combination with established clinical risk factors. At the same time, we discuss important limitations, including limited generalizability across populations and the issue of missing heritability. Finally, we highlight current methodological developments and future perspectives for the integration of PGS into clinical practice.

Introduction

In the past few decades, genome-wide association studies (GWAS) have established the standard analytical framework for identifying common genetic variants associated with complex traits and diseases [1, 2]. Large-scale GWAS have identified thousands of associated loci across the human genome, demonstrating that the genetic architecture of most complex phenotypes is highly polygenic, with numerous variants of small individual effects contributing to the overall genetic liability [3].

This progress has been enabled by major advances in genotyping and sequencing technologies [4, 5]. The decreasing cost of high-density single nucleotide polymorphism (SNP) arrays has facilitated the collection and analysis of very large population cohorts, allowing systematic assessment of common variant effects [6, 7]. More recently, improvements in sequencing and the development of large-scale biobanks and international consortia have enabled genome-wide data generation at unprecedented scale and cost-efficiency [8–12]. As a result, genetic association studies have expanded from analyses involving a few hundred individuals to those including millions, achieving the statistical power necessary to detect associations with increasingly modest effect sizes [8].

The initial objective of GWAS was the identification of genetic loci associated with disease risk to provide mechanistic insight and inform potential therapeutic targets [1]. This approach has been successful in mapping reproducible associations and revealing biological pathways relevant to human traits and diseases [13]. However, an important limitation soon became evident: the identification of a risk locus does not necessarily translate into predictive utility at the individual level. For most complex traits, the effect size of a single common variant is small, and its contribution to disease risk is very small when considered in isolation [1]. A few notable exceptions exist, including the APOE ε4 allele in Alzheimer’s disease, variants in ™TO associated with obesity, and common missense and loss-of-function variants in PCSK9 affecting lipid levels, which show detectable phenotypic differences among individuals carrying different genotypes [14 – 17]. Another well-known example of a comparatively oligogenic architecture is represented by autoimmune diseases, in which variants within the HLA region can explain a substantial proportion of SNP-based heritability [18]. Nevertheless, such examples are rare, and for most phenotypes genetic susceptibility arises from the combined influence of many small-effect variants [2, 3].

This recognition has led to the development of polygenic scores (PGS), also known as polygenic risk scores (PRS) when designed to estimate susceptibility to disease [19, 20]. A PGS aggregates the effects of multiple genetic variants into a single quantitative measure of genetic liability, typically computed as a weighted sum of allele dosages using GWAS-derived effect estimates [21]. The availability of large-scale biobank datasets and high-dimensional genotype–phenotype data has further enabled the development of improved prediction models, including those based on direct genotype-to-phenotype regression, which better approximate the true distribution of genetic effects [22 – 24].

The objective of this review is to provide a concise methodological and conceptual overview of PGS; from their derivation and statistical assumptions to their current limitations and emerging translational applications. We describe the main methodological approaches for PGS estimation and interpretation, outline the major sources of publicly available PGS repositories, refer to currently implemented models in clinical practice, and discuss key challenges and future directions for their robust and generalizable implementation in translational research and clinical applications.

Methods for PGS development

In order to compute PGS, a structured approach for combining the small effects of many genetic variants into a single measure is needed. As the goal is prediction, a PGS could be generally derived via a multivariable regression model and correspond to a structured linear predictor

graphic file with name j_medgen-2026-3009_fig_003.jpg

where the PGS of person i is derived as the weighted sum of the number of risk alleles of variants j = 1, … p (xij) with effect estimates βj. However, this is statistically and computationally challenging because most traits are influenced by a very large number of variants and because variants in close proximity in the genome are correlated through linkage disequilibrium (LD). Therefore, most commonly used PGS methods rely on combining GWAS summary statistics (e.g., βj that have been estimated separately) and adjust for LD structure to derive an individual-level genetic score in an independent target cohort with genotype data. As such, development of a PGS with these methods involves many statistical estimates (GWAS, LD matrix, PGS effect estimates) and is sensitive to high quality data. Therefore, it is crucial to rely on well powered GWAS that incorporate rigorous quality controls and take further adjustments, in particular for genetic ancestry, into account. Beyond summary-statistics-based approaches, PGS can be derived using individual-level training in large biobank datasets, leveraging joint modeling of variants while accounting for LD within the training cohort.

During the last decade, a wide range of PGS methods have emerged spanning from simple clumping and thresholding approaches (PRSice, [25]) to highly complex deep convolutional networks [26]. Methods can be generally categorized into two groups: GWAS-based and individual-level data based approaches (Figure 1) while the majority falls into the first category. Here, PGS are constructed as a weighted sum of effect alleles with weights taken from the estimated effect sizes βj of the variants in GWAS. One of the first PGS algorithms was clumping and thresholding (implemented in PRSice). To overcome estimation problems due to LD, SNPs are selected for the PGS based on a p-value threshold (thresholding) and a threshold for pairwise correlation within defined windows (clumping). Both thresholds are typically optimized on a validation set or via cross-validation. This simplistic approach neglects a more complex LD structure which led to the development of more advanced algorithms that are based on external population-matched LD reference panels. While lassosum [27] uses penalized regression, most algorithms in this category rely on Bayesian statistics, such as LDpred [28] and PRScs [29]. Those methods differ in the assumed priors and generally include all SNPs in the PGS but modify the effect estimates from GWAS. For example, LDpred assumes a proportion p of causal variants and adapts a Gaussian mixture prior, meaning that the effect sizes of the causal variants follow a normal distribution and the effect sizes of the remaining variants are theoretically 0. Using a Bayesian approach, this results in posterior means for the effect size of each variant that are typically different from 0.

Instead of relying on GWAS and subsequently attempting to adjust univariate effect estimates βj from summary statistics with advanced methods to account for LD, in large cohorts the βj for PGS can nowadays also be derived directly from individual-level data via multivariable regression models [29]. From a methodological perspective, these methods can be seen as more adequate for the polygenic situation at hand, as they hence account for LD already naturally when estimating the βj in the first place.

In recent years, statistical learning methods relying on penalized regression via the lasso [22] or statistical boosting [23] have emerged but also Bayesian approaches based on individual-level data have been implemented [31]. Via simultaneous consideration of all SNPs, those methods can yield better PGS estimates and more accurate predictions but rely on individual-level data from large cohorts [23, 32]. In comparison, via GWAS one can easily leverage statistical power by combining results from millions of individuals from multiple different cohorts. The landscape of PGS methods have further been enlarged to multi-PGS approaches (e.g. [33]), multi-ancestry PGS (e.g. [34]) and the integration of a priori knowledge such as functional annotations (e.g. [35]).

The choice of the best method depends on various factors and most importantly on the available data. In the beginnings of PGS, genotype-phenotype associations were particularly available as summary statistics from GWAS which are easily shareable without exposing individual information. More recently, various biobanks were established that employ managed user access to ensure compliance with data sharing agreements.

Figure 1:

Figure 1:

PGS workflow

Overview of the end-to-end workflow for polygenic score (PGS) construction and use. PGS models can be developed using training/validation data either from GWAS summary statistics or from individual-level genotype–phenotype data, with key steps including effect size estimation, variant selection/shrinkage (often accounting for LD), and optional adjustment for covariates. Alternatively, previously published scores (e.g., from the PGS Catalog) can be selected and reused. The resulting PGS is then computed in an independent target cohort and integrated with clinical risk factors, environmental and lifestyle variables, and other confounders for prediction, risk stratification, and downstream analyses.

Among the most prominent, UK Biobank comprises data of approximately half a million individuals [6], however covering mostly European ancestry. Over time, further biobanks were established including FinnGen [36], Biobank Japan [37], deCODE Genetics [38] and the Estonian Biobank [39]. Lastly, data of the All of Us program became available [11] which aims at enrolling at least one million individuals from the US and focuses on building a diverse biobank with respect to historically under-represented groups.

For quantitative traits that are comparatively easy to measure at scale (e.g., blood biomarkers or anthropometric traits), the gain in methodological accuracy enabled by directly modeling variants jointly in large biobanks often translates into improved predictive performance compared with summary-statistics-based approaches. However, for phenotypes that are harder to define reliably in population biobanks – such as neuropsychiatric traits, clinically adjudicated outcomes, or early-onset and rare conditions that are underrepresented in biobanks – PGS construction based on GWAS from deeply phenotyped clinical cohorts assembled by disease-specific consortia is often preferred, as these datasets provide more accurate and harmonized phenotype definitions. Moreover, individual-level data are frequently difficult to share across cohorts due to consent and data-protection constraints, such that GWAS meta-analyses and their summary statistics often represent the only broadly available input for downstream PGS construction.

Beyond the choice between GWAS summary statistics and individual-level training, the optimal PGS model also depends on the underlying genetic architecture. Current frameworks range from (near-)omnigenic/infinitesimal models, which posit that effects are widely distributed across the genome and therefore apply broad shrinkage while retaining a large fraction of SNPs, to sparse models that emphasize variable selection and concentrate weight on a comparatively small set of variants. Sparse approaches tend to perform better when the genetic signal is more concentrated (low-to-moderate polygenicity), whereas highly polygenic or near-omnigenic traits are often captured more accurately by models that assume many non-zero effects. Indeed, some traits are characterized by extremely polygenic architectures, such as height [40], whereas others are more strongly influenced by a small number of loci with larger effects and therefore exhibit a comparatively sparse architecture, for example lipoprotein(a) levels [41], which are largely driven by variation at the LPA locus.

Availability and downstream application of PGS models

For most PGS methods, the final PGS is a weighted sum of effect alleles, i.e. the model is characterized by the included SNPs, the defined effect allele and the corresponding effect size βj. Therefore, PGS models are easily shareable and are not subject to privacy concerns. As a result, large databases of published PGS models have formed such as the PGS catalog (https://www.pgscatalog.org/, [42]). Resources like this enable the evaluation of polygenic factors in smaller cohorts without the need to construct a new PGS. Given a target data set for analysis with a specific trait of interest, the choice of a suitable PGS from the PGS catalog depends on different features. One major decision point is the training population of a PGS. In the optimal case, the ancestry of the training population (GWAS and PGS training) match the ancestry of the target data set. Most PGS, however, are derived from individuals of European ancestry and will therefore show strongly decreased performance on individuals of different ancestry. In the PGS catalogue, the ancestry of the training population is indicated for each PGS. Furthermore, it is important to find a PGS with a matching phenotype definition. While this can be straight-forward for well-defined traits like height and BMI or disease conditions based on standardized ICD codes, it can be more challenging for phenotypes with a less structured definitions (e.g., behavioural traits) or that are characterized by a specific age of onset or dichotomized based on a (biomarker) threshold (e.g., obesity based on BMI). Finally, one should select a PGS with maximal SNP coverage in the target dataset; this is usually not a problem with high-density imputed genotypes or whole-genome sequencing, but can be limiting for samples genotyped on older arrays or with targeted genotyping/sequencing panels. After selecting a suitable PGS, it can be computed on the target set, e.g. via plink [43, 44]. To ensure a high quality, thorough quality control of the genotype data of the target set should be carried out including appropriate filters for minor allele frequency, genotype missing rate and Hardy-Weinberg-equilibrium. Challenges in the calculation mainly arise due to missing genotypes which can partly be replaced by proxy-SNPs or differences in the specification of the reference allele which can be re-coded.

Typically, the PGS of an individual cannot be interpreted directly but is compared to the distribution in a defined reference population. Additionally, different performance metrics should be taken into account. For quantitative phenotypes, the coefficient of determination R2 is a common choice. It is often interpreted as the proportion of explained variance, however, its definition on test data is ambiguous [45]. Most often, it’s defined as the squared correlation of the PGS and the observed phenotype in the test set. As such, it can be interpreted as the proportion of explained variance on the test set when fitting a linear regression with the PGS as the only covariate. Noteworthy, it captures solely association of the PGS and the phenotype, not calibration which could be assessed via metrics such as the mean squared error of prediction or the mean absolute error of prediction. When interpreting R2 is a common choice. It is often interpreted as the proportion of explained variance, however, its definition on test data is ambiguous [45]. Most often, it’s defined as the squared correlation of the PGS and the observed phenotype in the test set. As such, it can be interpreted as the proportion of explained variance on the test set when fitting a linear regression with the PGS as the only covariate. Noteworthy, it captures solely association of the PGS and the phenotype, not calibration which could be assessed via metrics such as the mean squared error of prediction or the mean absolute error of prediction. When interpreting R2 values, it is crucial to take SNP-based heritability estimates as an upper bound into account. Most complex traits are not only influenced by polygenic signal but also by rare or structural variants as well as clinical and environmental factors. Methods like LD score regression provide estimates of the polygenic heritability of complex traits which is typically modest [46, 47]. Putting the R2 in relation to this value increases interpretability. For qualitative outcomes, the R2 can be replaced by Nagelkerke’s R2 values, it is crucial to take snp-based heritability estimates as an upper bound into account. Most complex traits are not only influenced by polygenic signal but also by rare or structural variants as well as clinical and environmental factors. Methods like LD score regression provide estimates of the polygenic heritability of complex traits which is typically modest [46, 47]. Putting the R2 in relation to this value increases interpretability. For qualitative outcomes, the R2 can be replaced by Nagelkerke’s R2 which measures the explained variability on a liability scale. Also the area under the receiver operator characteristic curve (AUC) is a common discriminatory measure for binary outcomes. For time-to-event data, this principle is represented via the C-index. For both, AUC and C-index, a value of 0.5 indicates no discriminatory power while a value of 1 refers to perfect discrimination of cases and controls. To take additionally calibration into account, the Brier score (binary) and integrated Brier score (time-to-event) are popular alternatives with lower scores indicating a better calibration and accuracy.

Due to the architecture of complex traits, in the subsequent analysis, PGS are typically integrated into (statistical) models alongside confounding factors (such as sex, age and genetic ancestry) as well as additional risk factors including environmental, lifestyle and clinical variables. Typically, the performance of PGS is then further assessed by incremental value, i.e. improvement of R2, AUC or C-index by the full model compared to a baseline model excluding the PGS. Furthermore, standardized effect estimates (β, odds ratio or hazard ratio per standard deviation) can be derived and quantify the polygenic component. Lastly, individuals are often grouped into PGS strata (e.g., based on PGS deciles) that enable more categorical comparisons (e.g., high PGS vs. medium PGS).

Figure 2:

Figure 2:

Risk stratification by PGS

PGS-based stratification using quantiles of the standardized PGS distribution (left). Individuals are grouped into PGS strata (e.g., extreme low, intermediate, and extreme high percentiles), and the corresponding age-dependent cumulative disease risk is shown (right), highlighting increasing lifetime risk with higher PGS strata.

Among the first fields of application, polygenic contributions have been quantified in psychiatric diseases including schizophrenia and bipolar disorder [48]. Khera et al. showed that in some common and complex diseases, individuals with a high PGS show equivalent risk compared to individuals with a monogenic condition [49]. These studies gave insights into a new understanding of genetic predisposition and explained parts of observed heritability that could not be explained before [50 – 52]. Across studies, individuals in different strata of the PGS distribution show age-dependent differences in disease prevalence, supporting the use of PGS quantiles for risk stratification (Figure 2).

PGS studies are highly heterogeneous in both design and methodology, highlighting the need for standardized reporting to improve comparability [53]. A PGS report should clearly describe the study design, research objective, and outcome definition, as well as the study population, including recruitment strategy, demographic and clinical characteristics, and ancestral background. Any incorporated datasets, such as GWAS summary statistics, should be explicitly specified. Methodological reporting should include data pre-processing steps, including genetic quality control and handling of missing data. PGS development and evaluation should be conducted in independent samples. Accordingly, studies constructing a new PGS should detail cohort splitting procedures, the applied PGS method, and computational pipelines. For published scores, identifiers such as PMID or PGS Catalogue Score ID should be provided. Finally, reports should describe all statistical models, predictor variables, software, PGS distributions, and performance metrics relevant to the research aim. Limitations should be acknowledged, and code and, where possible, data should be made publicly available to ensure transparency and reproducibility.

Translation of PGS into clinical practice

PGS can be translated into clinical care mainly as risk stratification tools rather than standalone prognostic markers. They estimate an individual’s relative genetic liability within a reference population and are clinically useful when combined with baseline risks and established predictors to guide threshold-based decisions on prevention, screening, or follow-up. The most advanced applications concern primarily the prevention of common multifactorial diseases [49]. For coronary artery disease and related cardiometabolic traits, integration of disease-specific PGS into established risk algorithms improves discrimination and, more importantly, reclassification of individuals at the borders of guideline-defined risk categories [54].

In oncology, PGS are increasingly embedded into multivariable models to enable risk-adapted screening. Breast cancer risk calculators such as BOADICEA/CanRisk now incorporate a PGS alongside classical risk factors and family history [55, 56]. Trial data from risk-stratified screening, such as WISDOM, further demonstrate that adding a PGS leads to clinically relevant changes in screening recommendations [57]. For prostate cancer, studies such as BARCODE1 have shown that targeting MRI and biopsy to men in the highest decile of a prostate cancer PGS yields a high detection rate of clinically significant tumours, including cancers that would not have fulfilled biopsy criteria under PSA- and MRI-based protocols alone [58].

A second clinical use case is refining risk in carriers of pathogenic variants in monogenic disease genes with incomplete penetrance. For breast and colorectal cancer, PGS can stratify carriers of moderate-penetrance variants such as CHEK2 or PMS2, with higher PGS strata approaching risks observed in carriers of established high-penetrance variants, such as BRCA1/2 or MLH1/MSH2, and lower strata closer to population-level risks [59, 60].PGS are also being explored as stratification tools in pharmacogenomics and therapeutic studies [61]. Post-hoc analyses of cardiovascular outcome trials have shown that patients with high polygenic risk for coronary artery disease derive greater absolute and relative benefit from intensive LDL-lowering therapy than those with low genetic risk, despite similar lipid reductions [62]. In parallel, genotype-based restrictions on single variants, such as APOE ε4 status for lecanemab in Alzheimer’s disease, illustrate how genetic markers can be tied directly to labelling and prescribing decisions [63].

Different implementation projects based on standardized electronic health record–linked care pathways, such as the eMERGE study, illustrate the integration of PGS into clinical workflows for common complex diseases, including coronary heart disease, cancer predisposition, and metabolic disorders [64]. In parallel, population programmes like Our Future Health plan to compute PGS for several diseases at scale and to investigate how these scores can inform risk-stratified prevention strategies [65].

Beyond classical clinical care settings, PGS have also been proposed for embryo selection in the context of in vitro fertilization, where embryos are ranked according to PGS for common diseases or even non-disease traits. At present, this use is highly controversial: for most traits, the proportion of variance explained by current PGS is modest and robust prospective validation is lacking. In addition, such applications raise substantial concerns about the potential for eugenic misuse and inequities in reproductive medicine, and are therefore viewed critically in current international ethical debates [66 – 69]. Nevertheless, several commercial providers, particularly in the US, have begun to implement PGS-based embryo selection, underlining the need for clear professional guidance and regulatory frameworks as polygenic scoring computation becomes technically straightforward and widely available [70].

Looking further ahead, conceptual discussions have also started to address the possibility of polygenic genome editing in embryos or germ cells. Simulation studies suggest that editing a limited number of variants could, in theory, produce relevant shifts in polygenic risk for several common diseases and related risk factors, with the potential to substantially reduce disease prevalence and to yield both individual-level benefits and reductions in health-care costs [71].

Challenges and limitations

Despite new and larger data cohorts as well as continuous methodological progress (see previous section), there still remain substantial challenges:

Overall, the predictive performance of PGS remains limited [72]. For most complex traits, current models explain only a modest proportion of phenotypic variance, and their accuracy varies substantially across traits, cohorts, and ancestral backgrounds [73, 74]. Several factors contribute to this variability. Effect-size estimates used in PGS construction are typically derived from cohorts in populations of European ancestry, which limits transferability to other populations due to differences in LD structure, allele frequencies, and gene–environment interactions [73]. These issues have both methodological and ethical implications, as the application of PGS across ancestries may exacerbate disparities in predictive performance and clinical utility [74]. Thus, one of the main challenges at the current stage is the limited transferability of PGS across populations [75]. Potential solutions to this challenge come from two different angles. On the one hand, there is a need for more diverse cohorts and more cohorts with participants with other than European ancestry to develop population-specific scores [76]. On the other hand, also the methods for the development of PGS need to be adapted in order to yield also scores that are better to transfer or update when moving from one cohort to another and that can also work well with individuals with admixed ancestry [77, 78].

Importantly, PGS performance is influenced not only by ancestry differences, but also by additional sources of mismatch between the training dataset and the target population in which the score is applied. This applies, for example, to the translation of estimates from case-control studies to the population, where disease prevalence needs to be considered when interpreting genetic liability, heritability, and the expected strength of PGS-based risk stratification [79]. It is also relevant for population-scale applications, because many current and future scores are derived from large biobanks that often do not represent the general population. For example, the UK Biobank is a volunteer-based cohort recruited among middle-aged and older adults and is affected by healthy-volunteer bias, with participants being, on average, healthier, more educated, and socioeconomically less deprived than the underlying population [80]. This may also limit the capture of age-specific genetic effects, for instance when genetic factors play a larger role in early-onset phenotypes. More generally, participation-related biases may distort genetic associations and downstream estimates used for PGS development and evaluation.

Another key limitation is the incomplete capture of heritability by common variants [81–83]. Even the largest GWAS explain only part of the overall genetic contribution to complex traits, leaving a portion of “missing heritability” that may be attributable to rare variants, structural variation, gene-gene and gene-environment interactions, and regulatory mechanisms not adequately represented in current models [82, 84]. Recently, methodological developments have been made on the integration of rare variants [85], dominance effects [86] or gene-gene interactions [87]. Furthermore, environmental and lifestyle factors strongly modulate disease risk, and their interplay with genetic predisposition remains an active area of research [88–90]. Integrating environmental, epigenetic, and multi-omic information with genetic predictors is expected to improve both the predictive accuracy and biological interpretability of polygenic models [91, 92].

Finally, a key harmonization challenge arises from the inherently relative nature of PGS when they are implemented in clinical health-care systems that require standardized operating procedures (SOPs). By construction, a PGS is defined with respect to the allele frequency distribution, sequencing or genotyping platform, and imputation pipeline of the training population; consequently, the same numerical score does not necessarily correspond to the same absolute disease risk across cohorts or health-care systems. Robust clinical deployment therefore requires harmonized calibration frameworks that allow comparison of baseline hazards and enable systematic re-calibration when models are transferred to new datasets. In this context, large-scale, population-based national health-care datasets as well as clear reporting standards [53] may play a key role in supporting harmonization and cross-system comparability of PGS. In parallel, interoperable implementation pipelines and common data models are needed to support the integration of PGS with harmonized clinical, biomarker, and lifestyle data. Such infrastructures are essential to ensure that PGS-based decision support can be deployed consistently across centers, rather than resulting in locally defined and non-comparable scales of “genetic risk”.

Conclusion and outlook

Recent years have seen substantial progress both regarding methodological advances for the development of PGS and substantial investments in population-based genetic cohorts. These developments should now help in the future to translate PGS from an academic concept towards routine clinical care. While there are first and promising developments in this direction [55, 60, 93], there are still some challenges ahead.

Moving forward will require not only more sophisticated methods that can take epistasis, non-linear (dominant) effects, interactions with environmental factors into account but also models that are better fine-mapped on the causal variants to enhance the transferability of PGS from one population to another. Beyond predictive discrimination, future work should increasingly focus on calibration, individual uncertainty quantification, and robust estimation across heterogeneous populations. An important aspect here could be the inclusion of even more advanced statistical modeling frameworks like distributional regression [94] or quantile regression [95, 96], that relate more aspects (beyond the mean [97]) of the distribution of the phenotype to the genetic variants. Lately, these methods have been also adapted to estimate PGS for the variance of a phenotype, which could even help to identify candidates for gene-environment interactions [98, 99]. Prospective cohorts also offer opportunities to develop longitudinal and time-to-event polygenic models that account for dynamic risk trajectories of the life course.

Additionally, there is a need for more diverse cohorts for PGS development with a focus on participants with non-European or admixed ancestry. To ensure that PGS are accepted both among patients and healthcare professionals, it is essential to make sure that the benefits are accessible also to populations currently underrepresented in genetic cohorts. In addition to predictive performance also under shifts of the distribution, interpretability of the models and transparent reporting will be critical for clinical acceptance and regulatory evaluation.

An important aspect of future research on translating PGS into clinical practice might be on how to include them in integrated clinical decision support systems, where the PGS for various phenotypes could be incorporated as baseline genetic liability that may complement or interact with variables from the clinical history or environment of the patient. These systems then need to focus on a clear clinical decision point, where the prediction really can help to guide potential preventive or therapeutic interventions. To assess the performance of these systems, there is a strong need for pragmatic clinical trials that help to externally validate the combination of different risk factors and PGS. This includes randomized controlled trials which can evaluate the benefit of PGS but are rarely conducted so far [57, 100].

To realize the full potential of PGS in clinical practice, interdisciplinary collaboration between clinical and methodological experts across disciplines is essential, alongside increasing genetic diversity in large-scale GWAS and biobanks to better represent population diversity and enable the development of more generalizable risk models. In parallel, technical best practices and reporting standards are needed to ensure portability, standardization, and reproducibility. Such standards will be essential for a reliable integration of PGS into clinical decision-making.

Biographies

graphic file with name j_medgen-2026-3009_cv_001.jpg

Ph. Dr. Hannah Klinkhammer

graphic file with name j_medgen-2026-3009_cv_002.jpg

Prof. Dr. Andreas Mayr

graphic file with name j_medgen-2026-3009_cv_003.jpg

PD Dr. Carlo Maj

Affiliations

1Institute for Medical Biometry and Statistics, Marburg University, Marburg, Germany

2Center for Human Genetics, Marburg University, Marburg, Germany

Footnotes

Research ethics: Not applicable.

Informed consent: Not applicable.

Author contributions: All authors have accepted responsibility for the entire content of this manuscript and approved its submission.

Use of Large Language Models, AI and Machine Learning Tools: ChatGPT was used to improve language; all changes were reviewed and approved by the authors.

Conflict of interest: The authors state no conflict of interest.

Research funding: The work on this article was suppported by the German Research Foundation (DFG) under project number 534238115. Hannah Klinkhammer receives funding from the German Research Foundation (DFG) under project numer 574440500.

Data availability: Not applicable.

Contributor Information

Ph. Dr. Hannah Klinkhammer, Email: hannah.klinkhammer@uni-marburg.de.

Prof. Dr. Andreas Mayr, Email: andreas.mayr@uni-marburg.de.

PD Dr. Carlo Maj, Email: carlo.maj@uni-marburg.de.

References

  • [1].Tam V, Patel N, Turcotte M, Bossé Y, Paré G, Meyre D. Benefits and limitations of genome-wide association studies. Nat Rev Genet. 2019 Aug;20(8) pp. 467–84. doi:10.1038/s41576-019-0127-1 PubMed PMID: 31068683. [DOI] [PubMed]
  • [2].Visscher PM, Wray NR, Zhang Q, Sklar P, McCarthy MI, Brown MA. 10 Years of GWAS Discovery: Biology, Function, and Translation. Am J Hum Genet. 2017 Jul 6;101(1) pp. 5–22. et al. doi:10.1016/j.ajhg.2017.06.005 PubMed PMID: 28686856; PubMed Central PMCID: PMC5501872. [DOI] [PMC free article] [PubMed]
  • [3].Boyle EA, Li YI, Pritchard JK. An Expanded View of Complex Traits: From Polygenic to Omnigenic. Cell. 2017 Jun 15;169(7) pp. 1177–86. doi:10.1016/j.cell.2017.05.038 PubMed PMID: 28622505; PubMed Central PMCID: PMC5536862. [DOI] [PMC free article] [PubMed]
  • [4].Jenkins S, Gibson N. High-throughput SNP genotyping. Comp Funct Genomics. 2002;3(1) pp. 57–66. doi:10.1002/cfg.130 PubMed PMID: 18628885; PubMed Central PMCID: PMC2447245. [DOI] [PMC free article] [PubMed]
  • [5].Goodwin S, McPherson JD, McCombie WR. Coming of age: ten years of next-generation sequencing technologies. Nat Rev Genet. 2016 May 17;17(6) pp. 333–51. doi:10.1038/nrg.2016.49 PubMed PMID: 27184599; PubMed Central PMCID: PMC10373632. [DOI] [PMC free article] [PubMed]
  • [6].Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018 Oct;562(7726) pp. 203–9. et al. doi:10.1038/s41586-018-0579-z PubMed PMID: 30305743; PubMed Central PMCID: PMC6786975. [DOI] [PMC free article] [PubMed]
  • [7].Buniello A, MacArthur JAL, Cerezo M, Harris LW, Hayhurst J, Malangone C. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 2019 Jan 8;47(D1) pp. D1005–12. et al. doi:10.1093/nar/gky1120 PubMed PMID: 30445434; PubMed Central PMCID: PMC6323933. [DOI] [PMC free article] [PubMed]
  • [8].Abdellaoui A, Yengo L, Verweij KJH, Visscher PM. 15 years of GWAS discovery: Realizing the promise. Am J Hum Genet. 2023 Feb 2;110(2) pp. 179–94. doi:10.1016/j.ajhg.2022.12.011 PubMed PMID: 36634672; PubMed Central PMCID: PMC9943775. [DOI] [PMC free article] [PubMed]
  • [9].Kanai M, Akiyama M, Takahashi A, Matoba N, Momozawa Y, Ikeda M. Genetic analysis of quantitative traits in the Japanese population links cell types to complex human diseases. Nat Genet. 2018 Mar;50(3) pp. 390–400. et al. doi:10.1038/s41588-018-0047-6 PubMed PMID: 29403010. [DOI] [PubMed]
  • [10].Wei CY, Yang JH, Yeh EC, Tsai MF, Kao HJ, Lo CZ. Genetic profiles of 103,106 individuals in the Taiwan Biobank provide insights into the health and history of Han Chinese. NPJ Genomic Med. 2021 Feb 11;6(1) p. 10. et al. doi:10.1038/s41525-021-00178-9 PubMed PMID: 33574314; PubMed Central PMCID: PMC7878858. [DOI] [PMC free article] [PubMed]
  • [11].Genomic data in the All of Us Research Program. Nature. 2024 Mar;627(8003) pp. 340–6. doi:10.1038/s41586-023-06957-x PubMed PMID: 38374255; PubMed Central PMCID: PMC10937371. [DOI] [PMC free article] [PubMed]
  • [12].Thareja G, Al-Sarraj Y, Belkadi A, Almotawa M. , Suhre K. Whole genome sequencing in the Middle Eastern Qatari population identifies genetic associations with 45 clinically relevant traits. Nat Commun. 2021 Feb 23;12(1) p. 1250. et al. doi:10.1038/s41467-021-21381-3 PubMed PMID: 33623009; PubMed Central PMCID: PMC7902658. [DOI] [PMC free article] [PubMed]
  • [13].Hindorff LA, Sethupathy P, Junkins HA, Ramos EM, Mehta JP, Collins FS. Potential etiologic and functional implications of genome-wide association loci for human diseases and traits. Proc Natl Acad Sci U S A. 2009 Jun 9;106(23) pp. 9362–7. et al. doi:10.1073/pnas.0903103106 PubMed PMID: 19474294; PubMed Central PMCID: PMC2687147. [DOI] [PMC free article] [PubMed]
  • [14].Williams DM, Heikkinen S, Hiltunen M. , Davies NM, Anderson EL. The proportion of Alzheimer’s disease attributable to apolipoprotein E. NPJ Dement. 2026;2(1) p. 1. doi:10.1038/s44400-025-00045-9 PubMed PMID: 41522467; PubMed Central PMCID: PMC12789039. [DOI] [PMC free article] [PubMed]
  • [15].Dina C, Meyre D, Gallina S, Durand E, Körner A, Jacobson P. Variation in FTO contributes to childhood obesity and severe adult obesity. Nat Genet. 2007 Jun;39(6) pp. 724–6. et al. doi:10.1038/ng2048 PubMed PMID: 17496892. [DOI] [PubMed]
  • [16].Scartezini M, Hubbart C, Whittall RA, Cooper JA, Neil AHW, Humphries SE. The PCSK9 gene R46L variant is associated withlower plasma lipid levels and cardiovascular risk in healthy U.K. men. Clin Sci. 2007 Dec;113(11) pp. 435–41. doi:10.1042/CS20070150 PubMed PMID: 17550346. [DOI] [PubMed]
  • [17].Cohen JC, Boerwinkle E, Mosley TH, Hobbs HH. Sequence Variations in PCSK9, Low LDL, and Protection against Coronary Heart Disease. N Engl J Med. 2006 Mar 23;354(12) pp. 1264–72. doi:10.1056/NEJMoa054013. [DOI] [PubMed]
  • [18].Ramos PS, Shedlock AM, Langefeld CD. Genetics of autoimmune diseases: insights from population genetics. J Hum Genet. 2015 Nov;60(11) pp. 657–64. doi:10.1038/jhg.2015.94 PubMed PMID: 26223182; PubMed Central PMCID: PMC4660050. [DOI] [PMC free article] [PubMed]
  • [19].Crouch DJM, Bodmer WF. Polygenic inheritance, GWAS, polygenic risk scores, and the search for functional variants. Proc Natl Acad Sci U S A. 2020 Aug 11;117(32) pp. 18924–33. doi:10.1073/pnas.2005634117 PubMed PMID: 32753378; PubMed Central PMCID: PMC7431089. [DOI] [PMC free article] [PubMed]
  • [20].Sun J, Wang Y, Folkersen L, Borné Y, Amlien I, Buil A. Translating polygenic risk scores for clinical use by estimating the confidence bounds of risk prediction. Nat Commun. 2021 Sep 6;12(1) p. 5276. et al. doi:10.1038/s41467-021-25014-7 PubMed PMID: 34489429; PubMed Central PMCID: PMC8421428. [DOI] [PMC free article] [PubMed]
  • [21].Choi SW, Mak TSH, O’Reilly PF. Tutorial: a guide to performing polygenic risk score analyses. Nat Protoc. 2020 Sep;15(9) pp. 2759–72. doi:10.1038/s41596-020-0353-1 PubMed PMID: 32709988; PubMed Central PMCID: PMC7612115. [DOI] [PMC free article] [PubMed]
  • [22].Qian J, Tanigawa Y, Du W, Aguirre M, Chang C, Tibshirani R. A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression with application to the UK Biobank. PLOS Genet. 2020 Oct 23;16(10) p. e1009141. et al. doi:10.1371/journal.pgen.1009141. [DOI] [PMC free article] [PubMed]
  • [23].Klinkhammer H, Staerk C, Maj C, Krawitz PM, Mayr A. A statistical boosting framework for polygenic risk scores based on large-scale genotype data. Front Genet. 2023 Jan 10;13. doi:10.3389/fgene.2022.1076440. [DOI] [PMC free article] [PubMed]
  • [24].Zhang Q, Privé F, Vilhjálmsson B, Speed D. Improved genetic prediction of complex traits from individual-level data or summary statistics. Nat Commun. 2021 Jul 7;12(1) p. 4192. doi:10.1038/s41467-021-24485-y PubMed PMID: 34234142; PubMed Central PMCID: PMC8263809. [DOI] [PMC free article] [PubMed]
  • [25].Euesden J, Lewis CM, O’Reilly PF. PRSice: Polygenic Risk Score software. Bioinformatics. 2015 May 1;31(9) pp. 1466–8. doi:10.1093/bioinformatics/btu848 PubMed PMID: 25550326; PubMed Central PMCID: PMC4410663. [DOI] [PMC free article] [PubMed]
  • [26].Kelemen M, Xu Y, Jiang T, Zhao JH, Anderson CA, Wallace C. Performance of deep-learning-based approaches to improve polygenic scores. Nat Commun. 2025 Jun 2;16(1) p. 5122. et al. doi:10.1038/s41467-025-60056-1 PubMed PMID: 40456720; PubMed Central PMCID: PMC12130321. [DOI] [PMC free article] [PubMed]
  • [27].Mak TSH, Porsch RM, Choi SW, Zhou X, Sham PC. Polygenic scores via penalized regression on summary statistics. Genet Epidemiol. 2017 Sep;41(6) pp. 469–80. doi:10.1002/gepi.22050 PubMed PMID: 28480976. [DOI] [PubMed]
  • [28].Vilhjálmsson BJ, Yang J, Finucane HK, Gusev A, Lindström S, Ripke S. Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores. Am J Hum Genet. 2015 Oct 1;97(4) pp. 576–92. et al. doi:10.1016/j.ajhg.2015.09.001 PubMed PMID: 26430803; PubMed Central PMCID: PMC4596916. [DOI] [PMC free article] [PubMed]
  • [29].Ge T, Chen CY, Ni Y, Feng YCA, Smoller JW. Polygenic predictionvia Bayesian regression and continuous shrinkage priors. Nat Commun. 2019 Apr 16;10(1) p. 1776. doi:10.1038/s41467-019-09718-5 PubMed PMID: 30992449; PubMed Central PMCID: PMC6467998. [DOI] [PMC free article] [PubMed]
  • [30].Maj C, Staerk C, Borisov O, Klinkhammer H, Wai Yeung M, Krawitz P. Statistical learning for sparser fine-mapped polygenic models: The prediction of LDL-cholesterol. Genet Epidemiol. 2022 Dec;46(8) pp. 589–603. et al. doi:10.1002/gepi.22495 PubMed PMID: 35938382. [DOI] [PubMed]
  • [31].Moser G, Lee SH, Hayes BJ, Goddard ME, Wray NR, Visscher PM. Simultaneous discovery, estimation and prediction analysis of complex traits using a bayesian mixture model. PLoS Genet. 2015 Apr;11(4) p. e1004969. doi:10.1371/journal.pgen.1004969 PubMed PMID: 25849665; PubMed Central PMCID: PMC4388571. [DOI] [PMC free article] [PubMed]
  • [32].Ohta R, Tanigawa Y, Suzuki Y, Kellis M, Morishita S. A polygenic score method boosted by non-additive models. Nat Commun. 2024 May 29;15(1) p. 4433. doi:10.1038/s41467-024-48654-x PubMed PMID: 38811555; PubMed Central PMCID: PMC11522481. [DOI] [PMC free article] [PubMed]
  • [33].Albiñana C, Zhu Z, Schork AJ, Ingason A, Aschard H, Brikell I. Multi-PGS enhances polygenic prediction by combining 937 polygenic scores. Nat Commun. 2023 Aug 5;14(1) p. 4702. et al. doi:10.1038/s41467-023-40330-w PubMed PMID: 37543680; PubMed Central PMCID: PMC10404269. [DOI] [PMC free article] [PubMed]
  • [34].Patel AP, Wang M, Ruan Y, Koyama S, Clarke SL, Yang X. A multi-ancestry polygenic risk score improves risk prediction for coronary artery disease. Nat Med. 2023 Jul;29(7) pp. 1793–803. et al. doi:10.1038/s41591-023-02429-x PubMed PMID: 37414900; PubMed Central PMCID: PMC10353935. [DOI] [PMC free article] [PubMed]
  • [35].Zheng Z, Liu S, Sidorenko J, Wang Y, Lin T, Yengo L. Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries. Nat Genet. 2024 May;56(5) pp. 767–77. et al. doi:10.1038/s41588-024-01704-y PubMed PMID: 38689000; PubMed Central PMCID: PMC11096109. [DOI] [PMC free article] [PubMed]
  • [36].Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023 Jan;613(7944) pp. 508–18. et al. doi:10.1038/s41586-022-05473-8. [DOI] [PMC free article] [PubMed]
  • [37].Nagai A, Hirata M, Kamatani Y, Muto K, Matsuda K, Kiyohara Y. Overview of the BioBank Japan Project: Study design and profile. J Epidemiol. 2017 Mar;27(3S) pp. S2–8. et al. doi:10.1016/j.je.2016.12.005 PubMed PMID: 28189464; PubMed Central PMCID: PMC5350590. [DOI] [PMC free article] [PubMed]
  • [38].Gudbjartsson DF, Helgason H, Gudjonsson SA, Zink F, Oddson A, Gylfason A. Large-scale whole-genome sequencing of the Icelandic population. Nat Genet. 2015 May;47(5) pp. 435–44. et al. doi:10.1038/ng.3247. [DOI] [PubMed]
  • [39].Leitsalu L, Haller T, Esko T, Tammesoo ML, Alavere H, Snieder H. Cohort Profile: Estonian Biobank of the Estonian Genome Center, University of Tartu. Int J Epidemiol. 2015 Aug;44(4) pp. 1137–47. et al. doi:10.1093/ije/dyt268 PubMed PMID: 24518929. [DOI] [PubMed]
  • [40].Bicknell LS, Hirschhorn JN, Savarirayan R. The genetic basis of human height. Nat Rev Genet. 2025 Sep;26(9) pp. 604–19. doi:10.1038/s41576-025-00834-1. [DOI] [PubMed]
  • [41].Mack S, Coassin S, Rueedi R, Yousri NA, Seppälä I, Gieger C. A genome-wide association meta-analysis on lipoprotein (a) concentrations adjusted for apolipoprotein (a) isoforms[S] J Lipid Res. 2017 Sep 1;58(9) pp. 1834–44. et al. doi:10.1194/jlr.M076232. [DOI] [PMC free article] [PubMed]
  • [42].Lambert SA, Gil L, Jupp S, Ritchie SC, Xu Y, Buniello A. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nat Genet. 2021 Apr;53(4) pp. 420–5. et al. doi:10.1038/s41588-021-00783-5 PubMed PMID: 33692568; PubMed Central PMCID: PMC11165303. [DOI] [PMC free article] [PubMed]
  • [43].Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D. PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. Am J Hum Genet. 2007 Sep 1;81(3) pp. 559–75. et al. doi:10.1086/519795 PubMed PMID: 17701901. [DOI] [PMC free article] [PubMed]
  • [44].Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience. 2015;4. p. 7. doi:10.1186/s13742-015-0047-8 PubMed PMID: 25722852; PubMed Central PMCID: PMC4342193. [DOI] [PMC free article] [PubMed]
  • [45].Staerk C, Klinkhammer H, Wistuba T, Maj C, Mayr A. Generalizability of polygenic prediction models: how is the R2 defined on test data? BMC Med Genomics. 2024 May 16;17(1) p. 132. doi:10.1186/s12920-024-01905-8. [DOI] [PMC free article] [PubMed]
  • [46].Tang M, Wang T, Zhang X. A review of SNP heritability estimation methods. Brief Bioinform. 2022 May 1;23(3) p. bbac067. doi:10.1093/bib/bbac067. [DOI] [PubMed]
  • [47].Bulik-Sullivan BK, Loh PR, Finucane HK, Ripke S, Yang J, Patterson N. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat Genet. 2015 Mar;47(3) pp. 291–5. et al. doi:10.1038/ng.3211. [DOI] [PMC free article] [PubMed]
  • [48].Purcell SM, Wray NR, Stone JL, Visscher PM, O’Donovan MC. Common polygenic variation contributes to risk of schizophrenia and bipolar disorder. Nature. 2009 Aug 6;460(7256) pp. 748–52. et al. doi:10.1038/nature08185 PubMed PMID: 19571811; PubMed Central PMCID: PMC3912837. [DOI] [PMC free article] [PubMed]
  • [49].Khera AV, Chaffin M, Aragam KG, Haas ME, Roselli C, Choi SH. Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations. Nat Genet. 2018 Sep;50(9) pp. 1219–24. et al. doi:10.1038/s41588-018-0183-z PubMed PMID: 30104762; PubMed Central PMCID: PMC6128408. [DOI] [PMC free article] [PubMed]
  • [50].Yengo L, Vedantam S, Marouli E, Sidorenko J, Bartell E, Sakaue S. A saturated map of common genetic variants associated with human height. Nature. 2022 Oct;610(7933) pp. 704–12. et al. doi:10.1038/s41586-022-05275-y PubMed PMID: 36224396; PubMed Central PMCID: PMC9605867. [DOI] [PMC free article] [PubMed]
  • [51].Plomin R, von Stumm S. The new genetics of intelligence. Nat Rev Genet. 2018 Mar;19(3) pp. 148–59. doi:10.1038/nrg.2017.104 PubMed PMID: 29335645; PubMed Central PMCID: PMC5985927. [DOI] [PMC free article] [PubMed]
  • [52].Brandt M. Finding missing heritability in complex traits. Nat Genet. 2025 Dec;57(12) p. 2942. doi:10.1038/s41588-025-02455-0 PubMed PMID: 41366078. [DOI] [PubMed]
  • [53].Wand H, Lambert SA, Tamburro C, Iacocca MA, O’Sullivan JW, Sillari C. Improving reporting standards for polygenic scores in risk prediction studies. Nature. 2021 Mar;591(7849) pp. 211–9. et al. doi:10.1038/s41586-021-03243-6 PubMed PMID: 33692554; PubMed Central PMCID: PMC8609771. [DOI] [PMC free article] [PubMed]
  • [54].Inouye M, Abraham G, Nelson CP, Wood AM, Sweeting MJ, Dudbridge F. Genomic Risk Prediction of Coronary Artery Disease in 480,000 Adults. JACC. 2018 Oct 16;72(16) pp. 1883–93. et al. doi:10.1016/j.jacc.2018.07.079. [DOI] [PMC free article] [PubMed]
  • [55].Archer S, Babb de Villiers C, Scheibl F, Carver T, Hartley S, Lee A. Evaluating clinician acceptability of the prototype CanRisk tool for predicting risk of breast and ovarian cancer: A multi-methods study. PloS One. 2020;15(3) p. e0229999. et al. doi:10.1371/journal.pone.0229999 PubMed PMID: 32142536; PubMed Central PMCID: PMC7059924. [DOI] [PMC free article] [PubMed]
  • [56].Lee A, Mavaddat N, Wilcox AN, Cunningham AP, Carver T, Hartley S. BOADICEA: a comprehensive breast cancer risk prediction model incorporating genetic and nongenetic risk factors. Genet Med. 2019 Aug 1;21(8) pp. 1708–18. et al. doi:10.1038/s41436-018-0406-9. [DOI] [PMC free article] [PubMed]
  • [57].Esserman LJ, Fiscalini AS, Naeim A, Van’t Veer LJ, Kaster A, Scheuner MT. Risk-Based vs Annual Breast Cancer Screening: The WISDOM Randomized Clinical Trial. JAMA. 2025 Dec 12. p. e2524784. et al. doi:10.1001/jama.2025.24784 PubMed PMID: 41385349; PubMed Central PMCID: PMC12701531. [DOI] [PMC free article] [PubMed]
  • [58].McHugh JK, Bancroft EK, Saunders E, Brook MN, McGrowder E, Wakerell S. Assessment of a Polygenic Risk Score in Screening for Prostate Cancer. N Engl J Med. 2025 Apr 10;392(14) pp. 1406–17. et al. doi:10.1056/NEJMoa2407934 PubMed PMID: 40214032; PubMed Central PMCID: PMC7617604. [DOI] [PMC free article] [PubMed]
  • [59].Hassanin E, May P, Aldisi R, Spier I, Forstner AJ, Nöthen MM. Breast and prostate cancer risk: The interplay of polygenic risk, rare pathogenic germline variants, and family history. Genet Med Off J Am Coll Med Genet. 2022 Mar;24(3) pp. 576–85. et al. doi:10.1016/j.gim.2021.11.009 PubMed PMID: 34906469. [DOI] [PubMed]
  • [60].Hassanin E, Spier I, Bobbili DR, Aldisi R, Klinkhammer H, David F. Clinically relevant combined effect of polygenic background, rare pathogenic germline variants, and family history on colorectal cancer incidence. BMC Med Genomics. 2023 Mar 5;16(1) p. 42. et al. doi:10.1186/s12920-023-01469-z. [DOI] [PMC free article] [PubMed]
  • [61].Singh S, Stocco G, Theken KN, Dickson A, Feng Q, Karnes JH. Pharmacogenomics polygenic risk score: Ready or not for prime time? Clin Transl Sci. 2024 Aug;17(8) p. e13893. et al. doi:10.1111/cts.13893 PubMed PMID: 39078255; PubMed Central PMCID: PMC11287822. [DOI] [PMC free article] [PubMed]
  • [62].Marston NA, Kamanu FK, Nordio F, Gurmu Y, Roselli C, Sever PS. Predicting Benefit From Evolocumab Therapy in Patients With Atherosclerotic Disease Using a Genetic Risk Score: Results From the FOURIER Trial. Circulation. 2020 Feb 25;141(8) pp. 616–23. et al. doi:10.1161/CIRCULATIONAHA.119.043805 PubMed PMID: 31707849; PubMed Central PMCID: PMC8058781. [DOI] [PMC free article] [PubMed]
  • [63].van Dyck CH, Swanson CJ, Aisen P, Bateman RJ, Chen C, Gee M. Lecanemab in Early Alzheimer’s Disease. N Engl J Med. 2023 Jan 5;388(1) pp. 9–21. et al. doi:10.1056/NEJMoa2212948 PubMed PMID: 36449413. [DOI] [PubMed]
  • [64].Linder JE, Allworth A, Bland HT, Caraballo PJ, Chisholm RL, Clayton EW. Returning integrated genomic risk and clinical recommendations: The eMERGE study. Genet Med Off J Am Coll Med Genet. 2023 Apr;25(4) p. 100006. et al. doi:10.1016/j.gim.2023.100006 PubMed PMID: 36621880; PubMed Central PMCID: PMC10085845. [DOI] [PMC free article] [PubMed]
  • [65].Cook MB, Sanderson SC, Deanfield JE, Reddington F, Roddam A, Hunter DJ. Our Future Health: a unique global resource for discovery and translational research. Nat Med. 2025 Mar;31(3) pp. 728–30. et al. doi:10.1038/s41591-024-03438-0 PubMed PMID: 39838119. [DOI] [PubMed]
  • [66].Furrer RA, Barlevy D, Gandhi A, Carmi S, Lencz T, Pereira S. Survey of U.S. reproductive medicine clinicians’ attitudes on polygenic embryo screening. NPJ Genomic Med. 2025 Dec 1;10(1) p. 79. et al. doi:10.1038/s41525-025-00530-3 PubMed PMID: 41326419; PubMed Central PMCID: PMC12669692. [DOI] [PMC free article] [PubMed]
  • [67].Forzano F, Antonova O, Clarke A, de Wert G, Hentze S, Jamshidi Y. The use of polygenic risk scores in pre-implantation genetic testing: an unproven, unethical practice. Eur J Hum Genet EJHG. 2022 May;30(5) pp. 493–5. et al. doi:10.1038/s41431-021-01000-x PubMed PMID: 34916614; PubMed Central PMCID: PMC9090769. [DOI] [PMC free article] [PubMed]
  • [68].Grebe TA, Khushf G, Greally JM, Turley P, Foyouzi N, Rabin-Havt S. Clinical utility of polygenic risk scores for embryo selection: A points to consider statement of the American College of Medical Genetics and Genomics (ACMG) Genet Med Off J Am Coll Med Genet. 2024 Apr;26(4) p. 101052. et al. doi:10.1016/j.gim.2023.101052 PubMed PMID: 38393332. [DOI] [PubMed]
  • [69].Turley P, Meyer MN, Wang N, Cesarini D, Hammonds E, Martin AR. Problems with Using Polygenic Scores to Select Embryos. N Engl J Med. 2021 Jun 30;385(1) pp. 78–86. et al. doi:10.1056/NEJMsr2105065. [DOI] [PMC free article] [PubMed]
  • [70].Capalbo A, de Wert G, Mertes H, Klausner L, Coonen E, Spinella F. Screening embryos for polygenic disease risk: a review of epidemiological, clinical, and ethical considerations. Hum Reprod Update. 2024 Oct 1;30(5) pp. 529–57. et al. doi:10.1093/humupd/dmae012 PubMed PMID: 38805697; PubMed Central PMCID: PMC11369226. [DOI] [PMC free article] [PubMed]
  • [71].Visscher PM, Gyngell C, Yengo L, Savulescu J. Heritable polygenic editing: the next frontier in genomic medicine? Nature. 2025 Jan;637(8046) pp. 637–45. doi:10.1038/s41586-024-08300-4 PubMed PMID: 39779842; PubMed Central PMCID: PMC11735401. [DOI] [PMC free article] [PubMed]
  • [72].Martínez-Minguet D, Noel R, G Simón A, Pastor Ó. Challenges in clinical translation of polygenic risk score analyses: A systematic review. Genet Med Off J Am Coll Med Genet. 2026 Feb;28(2) p. 101662. doi:10.1016/j.gim.2025.101662 PubMed PMID: 41384390. [DOI] [PubMed]
  • [73].Martin AR, Gignoux CR, Walters RK, Wojcik GL, Neale BM, Gravel S. Human Demographic History Impacts Genetic Risk Prediction across Diverse Populations. Am J Hum Genet. 2017 Apr 6;100(4) pp. 635–49. et al. doi:10.1016/j.ajhg.2017.03.004 PubMed PMID: 28366442; PubMed Central PMCID: PMC5384097. [DOI] [PMC free article] [PubMed]
  • [74].Martin AR, Kanai M, Kamatani Y, Okada Y, Neale BM, Daly MJ. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat Genet. 2019 Apr;51(4) pp. 584–91. doi:10.1038/s41588-019-0379-x PubMed PMID: 30926966; PubMed Central PMCID: PMC6563838. [DOI] [PMC free article] [PubMed]
  • [75].Kachuri L, Chatterjee N, Hirbo J, Schaid DJ, Martin I, Kullo IJ. Principles and methods for transferring polygenic risk scores across global populations. Nat Rev Genet. 2024 Jan;25(1) pp. 8–25. et al. doi:10.1038/s41576-023-00637-2. [DOI] [PMC free article] [PubMed]
  • [76].Sirugo G, Williams SM, Tishkoff SA. The Missing Diversity in Human Genetic Studies. Cell. 2019 May 2;177(4) p. 1080. doi:10.1016/j.cell.2019.04.032 PubMed PMID: 31051100. [DOI] [PMC free article] [PubMed]
  • [77].Wang Y, Guo J, Ni G, Yang J, Visscher PM, Yengo L. Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations. Nat Commun. 2020 Jul 31;11(1) p. 3865. doi:10.1038/s41467-020-17719-y PubMed PMID: 32737319; PubMed Central PMCID: PMC7395791. [DOI] [PMC free article] [PubMed]
  • [78].Tanigawa Y, Kellis M. Power of inclusion: Enhancing polygenic prediction with admixed individuals. Am J Hum Genet. 2023 Nov 2;110(11) pp. 1888–902. doi:10.1016/j.ajhg.2023.09.013 PubMed PMID: 37890495; PubMed Central PMCID: PMC10645553. [DOI] [PMC free article] [PubMed]
  • [79].Pain O, Gillett AC, Austin JC, Folkersen L, Lewis CM. A tool for translating polygenic scores onto the absolute scale using summary statistics. Eur J Hum Genet EJHG. 2022 Mar;30(3) pp. 339–48. doi:10.1038/s41431-021-01028-z PubMed PMID: 34983942; PubMed Central PMCID: PMC8904577. [DOI] [PMC free article] [PubMed]
  • [80].Schoeler T, Speed D, Porcu E, Pirastu N, Pingault JB, Kutalik Z. Participation bias in the UK Biobank distorts genetic associations and downstream analyses. Nat Hum Behav. 2023 Jul;7(7) pp. 1216–27. doi:10.1038/s41562-023-01579-9 PubMed PMID: 37106081; PubMed Central PMCID: PMC10365993. [DOI] [PMC free article] [PubMed]
  • [81].Plomin R. Commentary: missing heritability, polygenic scores, and gene-environment correlation. J Child Psychol Psychiatry. 2013 Oct;54(10) pp. 1147–9. doi:10.1111/jcpp.12128 PubMed PMID: 24007418; PubMed Central PMCID: PMC4033839. [DOI] [PMC free article] [PubMed]
  • [82].Uricchio LH. Evolutionary perspectives on polygenic selection, missing heritability, and GWAS. Hum Genet. 2020 Jan;139(1) pp. 5–21. doi:10.1007/s00439-019-02040-6 PubMed PMID: 31201529; PubMed Central PMCID: PMC8059781. [DOI] [PMC free article] [PubMed]
  • [83].Yang J, Benyamin B, McEvoy BP, Gordon S, Henders AK, Nyholt DR. Common SNPs explain a large proportion of the heritability for human height. Nat Genet. 2010 Jul;42(7) pp. 565–9. et al. doi:10.1038/ng.608 PubMed PMID: 20562875; PubMed Central PMCID: PMC3232052. [DOI] [PMC free article] [PubMed]
  • [84].Wainschtein P, Jain D, Zheng Z. , Assessing the contribution of rare variants to complex trait heritability from whole-genome sequence data. Nat Genet. 2022 Mar;54(3) pp. 263–73. et al. doi:10.1038/s41588-021-00997-7 PubMed PMID: 35256806; PubMed Central PMCID: PMC9119698. [DOI] [PMC free article] [PubMed]
  • [85].Williams J, Chen T, Hua X, Wong W, Yu K, Kraft P. Integrating Common and Rare Variants Improves Polygenic Risk Prediction Across Diverse Populations [Internet] Epidemiology; 2024 [cited 2026 Feb 25] et al. Available from: http://medrxiv.org/lookup/doi/10.1101/2024.11.05.24316779. doi:10.1101/2024.11.05.24316779. [DOI] [PMC free article] [PubMed]
  • [86].Heyne HO, Karjalainen J, Karczewski KJ, Lemmelä SM, Zhou W. , Mono- and biallelic variant effects on disease at biobank scale. Nature. 2023 Jan;613(7944) pp. 519–25. et al. doi:10.1038/s41586-022-05420-7 PubMed PMID: 36653560; PubMed Central PMCID: PMC9849130. [DOI] [PMC free article] [PubMed]
  • [87].Fu B, Pazokitoroudi A, Shi Z, Kar A, Xue A, Anand A. A biobank-scale test of marginal epistasis reveals genome-wide signals of polygenic interaction effects. Nat Genet. 2025 Dec;57(12) pp. 3175–84. et al. doi:10.1038/s41588-025-02411-y PubMed PMID: 41366086; PubMed Central PMCID: PMC12695669. [DOI] [PMC free article] [PubMed]
  • [88].Miao J, Lin Y, Wu Y, Zheng B, Schmitz LL, Fletcher JM. A quantile integral linear model to quantify genetic effects on phenotypic variability. Proc Natl Acad Sci U S A. 2022 Sep 27;119(39) p. e2212959119. et al. doi:10.1073/pnas.2212959119 PubMed PMID: 36122202; PubMed Central PMCID: PMC9522331. [DOI] [PMC free article] [PubMed]
  • [89].Wray NR, Wijmenga C, Sullivan PF, Yang J, Visscher PM. Common Disease Is More Complex Than Implied by the Core Gene Omnigenic Model. Cell. 2018 Jun 14;173(7) pp. 1573–80. doi:10.1016/j.cell.2018.05.051 PubMed PMID: 29906445. [DOI] [PubMed]
  • [90].Herrera-Luis E, Benke K, Volk H, Ladd-Acosta C, Wojcik GL. Gene-environment interactions in human health. Nat Rev Genet. 2024 Nov;25(11) pp. 768–84. doi:10.1038/s41576-024-00731-z PubMed PMID: 38806721; PubMed Central PMCID: PMC12288441. [DOI] [PMC free article] [PubMed]
  • [91].Garg M, Karpinski M, Matelska D, Middleton L, Burren OS, Hu F. Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank. Nat Genet. 2024 Sep;56(9) pp. 1821–31. et al. doi:10.1038/s41588-024-01898-1 PubMed PMID: 39261665; PubMed Central PMCID: PMC11390475. [DOI] [PMC free article] [PubMed]
  • [92].Monti R, Eick L, Hudjashov G, Läll K, Kanoni S, Wolford BN. Evaluation of polygenic scoring methods in five biobanks shows larger variation between biobanks than methods and finds benefits of ensemble learning. Am J Hum Genet. 2024 Jul 11;111(7) pp. 1431–47. et al. doi:10.1016/j.ajhg.2024.06.003 PubMed PMID: 38908374; PubMed Central PMCID: PMC11267524. [DOI] [PMC free article] [PubMed]
  • [93].Torkamani A, Wineinger NE, Topol EJ. The personal and clinical utility of polygenic risk scores. Nat Rev Genet. 2018 Sep;19(9) pp. 581–90. doi:10.1038/s41576-018-0018-x PubMed PMID: 29789686. [DOI] [PubMed]
  • [94].Rigby RA, Stasinopoulos DM. Generalized Additive Models for Location, Scale and Shape. J R Stat Soc Ser C Appl Stat. 2005 Jun 1;54(3) pp. 507–54. doi:10.1111/j.1467-9876.2005.00510.x.
  • [95].Koenker R, Hallock KF. Quantile Regression. J Econ Perspect. 2001 Nov 1;15(4) pp. 143–56. doi:10.1257/jep.15.4.143.
  • [96].Klinkhammer H, Staerk C, Maj C, Krawitz PM, Mayr A. Genetic Prediction Modeling in Large Cohort Studies via Boosting Targeted Loss Functions. Stat Med. 2024 Dec 10;43(28) pp. 5412–30. doi:10.1002/sim.10249. [DOI] [PMC free article] [PubMed]
  • [97].Kneib T. Beyond mean regression. Stat Model. 2013 Aug;13(4) pp. 275–303. doi:10.1177/1471082X13494159.
  • [98].Miao J, Lin Y, Wu Y, Zheng B, Schmitz LL, Fletcher JM. A quantile integral linear model to quantify genetic effects on phenotypic variability. Proc Natl Acad Sci. 2022 Sep 27;119(39) p. e2212959119. et al. doi:10.1073/pnas.2212959119. [DOI] [PMC free article] [PubMed]
  • [99].Wu Q, Klinkhammer H, Kunwar K, Staerk C, Maj C, Mayr A. Detecting gene–environment interactions to guide personalized intervention: Boosting distributional regression for polygenic scores. Proc Natl Acad Sci. 2026 Apr 7;123(14) p. e2529164123. doi:10.1073/pnas.2529164123. [DOI] [PMC free article] [PubMed]
  • [100].Rodosthenous RS, Viiri LE, Carson A, Cajuso T, Corbetta A, Jones S. Effect of a Dietary Intervention on Weight Loss in Adults with High vs Low Genetic Predisposition to Higher BMI: A Randomized Diet Intervention Trial [Internet] medRxiv; 2025 [cited 2026 May 15] et al. p. 2025.10.06.25337395. Available from: https://www.medrxiv.org/content/10.1101/2025.10.06.25337395v1. doi:10.1101/2025.10.06.25337395.

Articles from Medizinische Genetik are provided here courtesy of De Gruyter

RESOURCES