Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Apr 24;17:5772. doi: 10.1038/s41467-026-72185-2

Integrating common and rare variants improves polygenic risk prediction across diverse populations

Jacob Williams 1,✉, Tony Chen 2, Xing Hua 1,3, Wendy Wong 1, Kai Yu 1, Peter Kraft 1, Xihao Li 4,5,✉, Haoyu Zhang 1,✉
PMCID: PMC13323733  PMID: 42031785

Abstract

PRSs predict complex traits by aggregating genetic effects across the genome, yet most models focus on common variants, overlooking rare variants that may contribute to hidden heritability. Here, we develop RICE, a PRS framework integrating both common and rare variants to improve genetic risk prediction across diverse ancestries. RICE constructs separate PRSs: for common variants, it integrates methods using ensemble learning; for rare variants, it uses gene-level testing with functional annotations and penalized regression. We evaluate RICE using simulated datasets and sequencing data from UK Biobank and All of Us, involving up to 740 million genetic variants from 361,939 individuals across diverse ancestries and 11 complex traits. In real data analysis, RICE improves predictive accuracy compared to leading common variant methods for traits with distinct rare variant architectures, particularly lipids and height. For lipid traits, incorporating rare variants increased R2 by up to ~11.2% in Europeans and ~60.7% in African ancestry compared to common variant PRS alone. Notably, for lipid traits, RICE captures substantial predictive signal beyond established high-penetrance genes, validating its ability to leverage the broader polygenic architecture of rare variation.

Subject terms: Statistical methods, Computational models, Genome-wide association studies, Quantitative trait


Most PRS models focus on common variants, overlooking rare variant contributions. Here, the authors present RICE, a PRS framework that integrates common and rare genetic variants to strengthen polygenic risk prediction across ancestries.

Introduction

Polygenic risk scores (PRSs) aggregate the effect sizes of numerous genetic variants across the genome to predict an individual’s risk of developing complex traits and diseases1–6. PRSs have shown promise in clinical applications, enabling personalized risk prediction and informing prevention strategies for conditions such as cardiovascular diseases, type 2 diabetes, and breast cancer7–10. By identifying individuals at higher genetic risk, PRSs can facilitate early interventions and improve patient outcomes.

Existing PRS methods focus primarily on common genetic variants, which are well-represented in large genome-wide association studies (GWAS). Due to their higher allele frequencies, common variants provide robust effect size estimates from large-scale GWAS, supporting the development of reliable risk prediction models. Advances in genotyping technologies, imputation methods, and the availability of extensive GWAS data have contributed to the success of common variant-based PRS models. To further refine risk prediction, methods accounting for linkage disequilibrium (LD) between variants, such as clumping and thresholding11–13, penalization-based models10,14–16, and Bayesian approaches17–21, have been developed. Recent approaches that jointly model data across multiple ancestries have improved performance in diverse cohorts22–26.

However, rare variants (minor allele frequency [MAF] <1%) are often excluded from PRSs due to low statistical power in single-variant association testing, historical limitations in large-scale sequencing data, limited integration methods, and challenges in accurately estimating their effect sizes. While some rare variants have larger, well-documented effect sizes (e.g., mutations in BRCA1/227, APOB28, PCSK9 28,29 or BSN30), many others tend to have modest effects that collectively contribute to disease risk31–35. Excluding these variants may underestimate risk, as the cumulative effect of multiple modest-effect rare variants can be significant36,37. Incorporating rare variants into PRSs can enhance risk prediction and provide a more comprehensive understanding of genetic risk, especially for complex traits influenced by numerous genetic factors.

Advancements in whole-genome sequencing (WGS) and whole-exome sequencing (WES) biobanks, such as UK Biobank (UKB) and All of Us Research Program (AoU), have greatly expanded opportunities for evaluating rare variants in both coding and noncoding regions32,38–41. To address analytical challenges posed by rare variants, methods such as the burden test, sequence kernel association test (SKAT), and variant-Set Test for Association using Annotation infoRmation (STAAR)42–47 have been developed to analyze groups of rare variants collectively. Approaches like STAARpipeline leverage functional annotations to define variant sets for both coding and noncoding rare variants46,48–51, improving statistical power and interpretability of identified genes. Despite these advancements in association testing for rare variants, methods that simultaneously integrate rare variants into PRSs and support multi-ancestry data are notably limited. In parallel, several studies have explored the use of rare variant burden scores for risk prediction. For example, Lali et al. developed RV-EXCALIBER to integrate rare variant burden into cardiovascular disease prediction52, while Chan et al. introduced the Genome-wide Rare Variant Score for autism spectrum disorders53. These approaches highlight the potential of rare variant burden scores in disease risk stratification, but they have not been broadly extended to integrate rare and common variants within a unified, multi-ancestry framework, nor have they been systematically evaluated across multiple large-scale biobanks and a wide range of complex traits.

To address these limitations, we present RICE (polygenic Risk predictions Integrating Common and rarE variants), a framework designed for biobank sequencing data that efficiently integrates both common and rare variants. RICE combines ensemble learning techniques with advanced association testing methods to construct comprehensive PRSs that capture genetic risk from both common and rare variants. RICE significantly enhances PRS, providing a more accurate and inclusive approach to genetic risk prediction.

We evaluated RICE using extensive simulated datasets and large-scale WES, Imputed, and WGS data from UKB and AoU, with up to 740 million genetic variants from 361,939 unrelated individuals across African (AFR), Admixed American or Latino (AMR), European (EUR), Middle Eastern (MID) and South Asian (SAS) ancestries. Our analyses spanned 11 complex traits, including height, body mass index (BMI), breast cancer, and type 2 diabetes (T2D). Simulation studies demonstrated that RICE effectively models rare variant signals, achieving substantial improvements in PRS accuracy across all ancestries compared to leading common variant methods. In real data analyses, RICE improved predictive accuracy compared to top existing common variant PRS methods for traits with distinct rare variant architectures, particularly lipid traits and height, demonstrating the utility of a comprehensive common and rare variant PRS framework.

Results

Method overview

RICE predicts risk of complex traits and diseases by incorporating common (RICE-CV) and rare variants (RICE-RV). The framework has three steps (Fig. 1): (1) RICE-CV combines multiple PRSs for common variants using ensemble learning to generate a single PRS; (2) RICE-RV identifies significant rare variant sets conditioned on RICE-CV, creates PRSs based on burden scores, and combines them via ensemble learning; (3) RICE-CV and RICE-RV PRSs are jointly evaluated in a regression model with covariates. RICE requires three independent datasets: (1) training for generating GWAS summary statistics, rare variant p values, and PRS models; (2) tuning for optimizing parameters; (3) validation for assessing prediction performance.

Fig. 1. Overview of the RICE framework for polygenic risk prediction.

Fig. 1

RICE integrates common and rare variants in three steps: a RICE-CV (common variant PRS) is generated by combining PRSs from multiple existing methods, such as clumping and thresholding (CT), LDpred2, and Lassosum2 for single-ancestry data, or CT-SLEB, JointPRS, and PROSPER for multi-ancestry data. These PRSs are then combined using ensemble learning (LASSO and ridge regression) to produce a single robust common variant PRS. b RICE-RV (rare variant PRS) identifies significant rare variant sets using STAARpipeline ( p<1×10−3) conditioned on the RICE-CV PRS. These significant sets are collapsed into burden scores and modeled using LASSO and ridge regression. The resulting PRSs are combined through ensemble learning (LASSO and ridge regression) to create a single rare variant PRS. c Final evaluation: The pre-trained RICE-CV and RICE-RV models (with combination weights fixed from the tuning step) are applied to the validation set to evaluate overall predictive performance (R2 or AUC). Additionally, a regression model is used solely to estimate individual standardized effect sizes (Beta per SD) of the components, adjusting for covariates such as age, sex, and principal components. The data are divided into three independent sets: a training set used to derive genome-wide association study (GWAS) summary statistics and perform rare variant association testing; a tuning set used to optimize model parameters and train the ensemble models; and a validation set used to evaluate the final predictive performance of the integrated PRS.

In the first step (Fig. 1a), RICE-CV uses the training dataset to identify common variants (MAF > 0.01 in any ancestry group) and compute ancestry-specific GWAS summary statistics. Using these summary statistics, existing PRS methods are used to generate multiple PRSs. To leverage the advantages of different methodological approaches, we select a representative method from each category: clumping-based approach, penalization-based approach, and Bayesian-based approach. For single-ancestry data (e.g., primarily EUR populations), PRSs are obtained from clumping and thresholding (CT)11–13, LDpred254, and Lassosum215. For multi-ancestry data, PRSs are obtained from CT-SLEB22, JointPRS55, and PROSPER25. After computing PRSs from different approaches under various tuning parameters, RICE-CV performs ensemble learning 54 to combine all PRSs into a single robust common variant PRS, using LASSO and ridge regression as base learners.

In the second step (Fig. 1b), RICE-RV uses the training set to fit a baseline model adjusting for covariates, including top 10 principal components (PCs), sex, age, age squared, and RICE-CV to prioritize rare variant signals independent of common variant effects. Using residuals from this baseline model, RICE-RV performs variant set analyses with STAARpipeline to identify significant rare variant sets48, which defines rare variant sets using functional categories within each protein-coding gene. In WES datasets, RICE-RV conducts gene-centric analysis with rare variants in coding region only. In WGS datasets, RICE-RV analyzes rare variants in both coding and noncoding regions. Significant rare variant sets are selected as those with a STAAR-Burden p value less than 1×10−3. These sets are then collapsed into burden scores, which approximately represent the combined allele frequencies (Methods). These burden scores are jointly modeled using LASSO and ridge regression under different tuning parameters. PRSs are then generated on the tuning dataset using these burden scores and weights estimated from different penalized regression models. Finally, the PRSs are combined through ensemble learning with LASSO and ridge regression as base learners, to create a single rare variant PRS.

To evaluate RICE, in the final step (Fig. 1c), RICE-CV and RICE-RV are jointly evaluated on the validation set using a linear or logistic regression model, adjusting for covariates including top 10 PCs, sex, age, and age squared. The estimated regression coefficients of RICE-CV and RICE-RV are reported separately as measures of predictive performance.

Evaluation of prediction performance

PRSs are traditionally evaluated using metrics like the coefficient of determination (R2) for continuous traits or area under the receiver operating characteristic curve (AUC) for binary traits2. These measures assess how well the PRS explains variation in a trait on average across the entire population. However, because rare variants are present in only a relatively small subset of individuals, their impact may not significantly influence these average-based metrics. Despite this, rare variants can have substantial effects on individuals who carry them, which may be overlooked when using traditional evaluation methods.

To better capture the contribution of rare variants, we report an additional, complementary metric: the estimated regression coefficient of the standardized PRS on the standardized trait (Methods). For continuous traits, this coefficient represents the expected change in standardized phenotype per standard deviation (SD) increase in PRS. For binary traits, it corresponds to the log odds ratio (log OR) per SD increase in PRS, a widely used and interpretable measure in clinical genetics and genetic epidemiology. Throughout this manuscript, we refer to this as the standardized effect size of PRS, denoted as “Beta per SD of PRS” for continuous traits and “log OR per SD of PRS” for binary traits. As detailed in the Supplementary Note, this coefficient can also be expressed as the square root of the heritability explained by the PRS. We use this metric to quantify the effect sizes of both common and rare variant PRSs, particularly when rare variants may have limited influence on average-based metrics, offering an interpretable per-SD effect that complements standard PRS measures like R2 or AUC. RICE’s predictive performance was evaluated using three complementary metrics: (1) standardized effect size, for assessing statistical significance and relative improvements; (2) R2 (continuous traits) or AUC (binary traits), for absolute predictive accuracy; and (3) trait differences across PRS quantiles, to illustrate potential clinical relevance in the UKB and AoU datasets.

Figure 2 illustrates the estimated coefficients of the standardized PRS for both RICE-CV and RICE-RV using high-density lipoprotein cholesterol (HDL) levels from the UKB WGS dataset. For RICE-CV, a one-unit increase in the PRS is associated with a 0.389 standard deviation (SD) increase in HDL levels (β1=0.389), while a one-unit increase in the RICE-RV PRS corresponds to a 0.107 SD increase (β2=0.107). Notably, the distribution of RICE-RV PRS values is more dispersed compared to RICE-CV, with approximately 8.1% of individuals exhibiting PRS values greater than five units above the population mean. This demonstrates that while rare variants affect a smaller segment of the population, their impact on those individuals can be substantial.

Fig. 2. Comparison of standardized PRSs from RICE-CV and RICE-RV for high-density lipoprotein cholesterol (HDL).

Fig. 2

Standardized PRSs for HDL were computed using UK Biobank whole genome sequencing data for 13,839 individuals of European ancestry from the validation dataset. The left panel shows the standardized observed HDL levels plotted against the standardized PRS from RICE-CV (common variants), while the right panel shows the same comparison using RICE-RV (rare variants). In both figures, the blue dashed line represents the upper 10% quantile cutoff. The estimated regression coefficients (β1 for RICE-CV and β2 for RICE-RV) are displayed in orange.

Simulation Study Results

We evaluated RICE’s ability to detect and predict rare variant effects using extensive simulation studies based on UKB WES data with 140,080 unrelated individuals (Methods). Genetic ancestries were inferred using 1000 Genomes Phase 3 Project 3 data (1000G, Methods)56. Since over 90% of individuals were of EUR ancestry, we designed simulations with an EUR-only training dataset (N = 98,343), while the tuning (N = 20,869) and validation datasets (N = 20,868) included individuals from AFR, AMR, EUR, and SAS ancestries (sample size details in Supplementary Data 1).

To evaluate RICE’s performance under realistic genetic architectures, we calibrated the relative contribution of rare variants by adopting a 1:12 ratio of rare variant burden heritability to common variant heritability, consistent with recent median estimates based on 22 complex traits 36. For our single-chromosome simulation framework, we set the common variant heritability (hCV2) at 5% to ensure sufficient signal strength on chromosome 22 for evaluating model mechanics, such as weight estimation and prevention of overfitting, within a manageable computational burden. Accordingly, the rare variant heritability (hRV2) was set at 0.42%.

RICE-CV was constructed by applying single-ancestry common variant PRS methods: CT, LDpred2, and Lassosum2 to the EUR training data. We compared the performance of RICE-CV and RICE-RV to these methods across different ancestries. The rare variant analysis was conducted using STAARpipeline with protein-coding genes within coding regions. RICE was evaluated across various simulation settings, including different causal variant proportions (1%, 5%, and 20%), causality structures of causal rare variant sets (a proportion or all rare variants within a set are causal), training sample size (N = 49,173 and 98,343), and effect size distributions (strong negative selection and no negative selection, Methods). RICE-CV consistently demonstrated robust performance (Fig. 3 and Supplementary Fig. 1). By leveraging the strengths of different common variant PRS methods, RICE-CV achieved higher prediction accuracy than any single approach. On average, RICE-CV yielded the highest standardized effect size across all ancestries and traits, showing consistent gains over the best existing single-method common variant PRS.

Fig. 3. Simulation results comparing the predictive performance of PRSs for four ancestral groups from the UK Biobank (UKB).

Fig. 3

The training data (N = 98,343) included only individuals of European ancestry (EUR), while the tuning (N = 20,869) and validation datasets (N = 20,868) contained individuals of African (AFR), Admixed American or Latino (AMR), European (EUR), and South Asian (SAS) ancestries (Supplementary Data 1). Simulations assumed a common variant heritability of 0.05 and a rare variant set heritability of 4.17×10−3, under the assumption of no negative selection effect size distribution (Methods). Causal proportions for both common variants and rare variant sets varied across three levels: 0.01 (top), 0.05 (middle), and 0.2 (bottom). Causal rare-variant sets were selected, and all variants within selected sets were assumed to contribute to the burden. Data were generated using unrelated individuals from UK Biobank whole-exome sequencing data (WES), with simulation based on chromosome 22. For each simulation scenario, 100 simulated traits were generated and results shown are the mean across the 100 validation-set evaluations. PRS performance is reported as the “Beta of PRS per standard deviation (SD)”, derived from the regression model Y~PRS×β, with β representing the effect of standardized PRS on the standardized outcome (Methods). For RICE, the model used was Y~PRSCV×βCV+PRSRV×βRV. Statistical significance of the RICE-RV component was assessed using percentile bootstrap confidence intervals (10,000 resamples), testing the two-sided alternative HA: βRV≠0; *** indicates the lower bound of the 99% bootstrap confidence interval (CI) > 0 and ** indicates the lower bound of the 95% bootstrap CI > 0. Exact bootstrap p values and CI bounds are provided in the Source Data file. Source data are provided as a Source Data file.

RICE-RV also performed strongly, effectively identifying causal rare variant signals and predicting their effects across both EUR and non-EUR populations (Fig. 3 and Supplementary Fig. 1). When added to RICE-CV, RICE-RV consistently contributed to independent predictive information. The standardized effect sizes for RICE-RV were significantly greater than zero for EUR across most scenarios (p<0.05), whereas in non-EUR populations, significance was generally observed only with larger sample sizes and higher proportions of causal variants. Notably, RICE-RV showed particularly strong performance in AFR individuals, as the randomly selected causal variants and rare variant sets across the genome led to higher average burden scores (combined allele frequencies) in this ancestry (Supplementary Fig. 2), driving more notable rare variant prediction performance. Overall, RICE-CV achieved the highest prediction accuracy among common variant PRSs across all populations in most scenarios, while RICE-RV correctly identified causal rare variants and accurately predicted their impact.

UKB imputed + WES and WGS results

We applied RICE to compute PRSs for 11 complex traits, with six continuous traits: BMI, HDL, height, low-density lipoprotein (LDL), log-transformed triglycerides (log(TG)), and total cholesterol (TC), and five binary traits: asthma, breast cancer, coronary artery disease (CAD), prostate cancer, and T2D. Training data included 98,103 EUR individuals, while tuning (N = 20,869) and validation (N = 20,868) datasets included individuals from AFR, AMR, EUR, and SAS ancestries (sample size details in Supplementary Data 2, 3; variant counts in Supplementary Data 9). We conducted association analyses as follows: for common variants, we adjusted for top 10 PCs, sex, age and age squared. For rare variant association testing, we further included RICE-CV as an additional covariate to ensure independence between RICE-RV and RICE-CV (Methods). Genomic control metrics indicated well-calibrated analyses, with λ1000 ranging between 1 to 1.018. LD score regression intercepts57 further confirmed minimal population stratification (intercept < 1.2) across all traits (Supplementary Data 4). Manhattan and QQ plots for common and rare variant association tests are available for both analyses: Imputed + WES (Supplementary Figs. 3, 4), and WGS (Supplementary Figs. 5, 6).

We used two analytical configurations: (1) Imputed + WES, using imputed genotypes for common variants and WES data for rare variants; and (2) WGS, using WGS for both common and rare variant analyses. Across both configurations, RICE-CV consistently outperformed existing common variant PRS methods for continuous traits, with average R2 improvements of 9.4% over the best alternative method per trait-ancestry combination (Supplementary Data 6, 7). Incorporating rare variants via RICE-RV yielded further gains, particularly for lipid traits (HDL, LDL, log(TG), and TC), where improvements were driven by genes with established roles in lipid metabolism58–60(e.g., APOC3 aggregated weight for HDL = 0.516, 3.8 standard deviations above the average aggregated effect size; full model weights are available on Harvard Dataverse as described in the Data Availability statement). For height, RICE-RV provided modest but statistically significant improvements in EUR and AMR, while BMI showed limited gains likely due to its polygenic architecture with smaller rare variant effects, and binary traits exhibited no improvements, probably owing to lower case numbers. Below, we detail results by each configuration.

In the UKB Imputed + WES analysis, we utilized imputed genotypes for common variants to ensure comprehensive genome-wide coverage, while retaining WES-based rare variants. Among common variant methods, RICE-CV generally achieved the largest standardized effect sizes across continuous traits (standardized effect sizes in Fig. 4 and Supplementary Fig. 7a; R2/AUC in Supplementary Figs. 7b, c; quantile differences in Supplementary Fig. 8). Among common variant methods, RICE-CV generally achieved the largest standardized effect sizes across continuous traits (Fig. 4). For rare variants, RICE-RV showed significant standardized effect sizes (p<0.05) for all lipid traits across ancestries and for height in EUR and AMR (Fig. 4). Relative to the best alternative method, RICE (CV + RV) improved R2 for lipid traits in EUR by 4.9–8.0% (Supplementary Fig. 7c). Notable non-EUR gains included 29.8% for TC in AFR, 25.9% for log(TG) in AMR, and 25.5% for HDL in AMR (Supplementary Fig. 7c). For height, the rare variant component was statistically significant but small (e.g., β=0.039 in EUR and 0.038 in AMR), and corresponding R2 gains were marginal, leaving RICE’s performance not significantly different from the best alternative for this trait, consistent with a highly polygenic rare variant architecture that yields small effects at the population level. Quantile stratification across ancestries showed clear gradients for lipid traits in EUR, with significant differences between extreme quantiles (Supplementary Fig. 8b, d–f).

Fig. 4. Predictive performance of ancestry-adjusted PRSs for continuous traits across four ancestral groups from UK Biobank (UKB) imputed + whole-exome sequencing (WES) data.

Fig. 4

The traits analyzed include body mass index (BMI), height, high-density lipoprotein cholesterol (HDL), low-density lipoprotein cholesterol (LDL), the natural logarithm of triglyceride cholesterol (log(TG)), and total cholesterol (TC). Results are shown for individuals of African (AFR), Admixed American or Latino (AMR), European (EUR), and South Asian (SAS) ancestries. The training data consisted solely of individuals of European ancestry, while tuning and validation sets included all four ancestries. Full sample size details for each ancestry are provided in Supplementary Data 2. Statistical significance of the RICE-RV component was assessed using percentile bootstrap confidence intervals (10,000 resamples), testing the two-sided alternative HA:βRV≠0; *** indicates the lower bound of the 99% bootstrap confidence interval (CI) > 0 and ** indicates the lower bound of the 95% bootstrap CI > 0. Exact bootstrap p values and CI bounds are provided in the Source Data file. Source data are provided as a Source Data file.

To test whether rare variants flag individuals missed by common variant PRS, we cross-classified participants by RICE-CV (top 10% vs. bottom 90%) and RICE-RV (top 5% vs. bottom 95%), yielding four strata (Methods). For HDL, stratifying by RICE-RV quantiles shows significant differences across RICE-CV quantiles (Fig. 5a). Further, 6.3% of individuals in the top phenotype decile were captured only by high RICE-RV despite low RICE-CV (Fig. 5b). Group-level comparisons confirmed that the low-CV/high-RV stratum had significantly higher HDL than the low-CV/low-RV group (p = 3.41 × 10⁻14; Fig. 5c). Similar patterns were seen for other lipid traits but not for BMI, consistent with the standardized effect-size results (Supplementary Fig. 9). These findings show that the rare variant component of RICE can identify individuals with extreme lipid profiles who would be missed by common variant PRS alone.

Fig. 5. Joint stratification of common and rare variant PRS identifies discrepant high-risk individuals.

Fig. 5

Analyses were performed in the UK Biobank (UKB) imputed genotype and whole-exome sequencing (WES) validation dataset (n = 18,150). High-density lipoprotein cholesterol (HDL) was residualized for age, age2, sex, and the first 10 genetic principal components, then standardized; PRSs were ancestry-adjusted as described in Methods. a Mean standardized HDL levels ± 1× standard error (SE) are plotted stratified by PRS quantiles for RICE-CV (common variants) on the x-axis and RICE-RV (rare variants) by color (blue: below 5%, gray: 30–70%, pink: above 95%). b Participants were cross-classified into four strata based on RICE-CV (high = top 10% vs low = bottom 90%) and RICE-RV (high = top 5% vs low = bottom 95%): Low CV/Low RV (n = 1248), High CV/Low RV (n = 435), Low CV/High RV (n = 115), and High CV/High RV (n = 35). Pie chart shows the proportion of individuals in each stratum among the top decile of observed HDL values (n = 1833). c Mean standardized HDL for each stratum with 95% confidence intervals (mean ± 1.96×SE). Pairwise p values were calculated using two-sided Welch’s t-tests comparing each stratum to the Low CV/Low RV group; exact p values and sample sizes are provided in the Source Data file. Source data are provided as a Source Data file.

To quantify the benefit of genome-wide rare variant signal beyond established high-penetrance genes, we compared RICE-RV to a gene-restricted baseline that used only burden scores from LDLR, APOB, and PCSK9 (with identical ensemble learning) for HDL, LDL, log(TG), and TC. Across ancestries, RICE-RV consistently outperformed this baseline, yielding 196% higher standardized effect size for Europeans (β per SD) (Supplementary Fig. 10), indicating contributions from additional modest-effect rare variant sets. We assessed sensitivity to the STAAR-Burden inclusion threshold by varying the p value cut-off from 1 × 10⁻5 to 1 × 10⁻2; performance was stable with no monotonic improvement at more liberal thresholds (Supplementary Fig. 11). Balancing computational cost and signal capture, we retained 1 × 10⁻3 for the main analyses.

We also assessed the computational efficiency of the RICE framework (Supplementary Data 10). Ensemble learning for RICE-CV and RICE-RV required on average 1.73 compute hours per trait (~1.1% of total compute time). Rare variant association testing accounted for the largest compute share (106.7 h, 381 parallel jobs, 68.6%), followed by LDpred2/Lassosum2 modeling (28.4 h, 18.9%).

Using the UKB WGS (with comprehensive variant coverage) reproduced the overall patterns seen in Imputed + WES (standardized effect sizes in Supplementary Fig. 12a, c; R2/AUC in Supplementary Fig. 12b, d; quantile stratification in Supplementary Fig. 13). Among common variant methods, RICE-CV generally achieved the largest standardized effect sizes across continuous traits. For rare variants, RICE-RV showed significant effects (p<0.05) for height and all lipid traits across ancestries, with BMI only being significant for SAS (Supplementary Fig. 12c). Relative to the best alternative method, RICE (CV + RV) increased R2 in EUR by 7.5% (height), 8.6% (HDL), 11.2% (LDL), 9.6% (log(TG)), and 10.6% (TC) (Supplementary Fig. 12d). Notable non-EUR gains included 60.7% for log(TG) in AFR and 29.8% for HDL in AMR. In contrast, binary traits again showed limited improvement. In EUR, stratifying individuals by RICE-RV quantiles revealed clear lipid gradients: for HDL, the lowest quantile averaged ~0.22 SD lower and the highest ~0.32 SD higher than the middle group (Supplementary Fig. 13), indicating that rare variants help identify individuals with markedly extreme lipid profiles that common variant PRS alone may miss. Due to smaller validation samples, trends were less stable in non-EUR groups.

We compared (1) regression-based adjustment for population structure (Methods; Supplementary Note) versus (2) ancestry-specific standardization using group-specific means/SDs. Performance was comparable (Supplementary Fig. 14). To avoid explicit ancestry categorization while adequately controlling structure, we adopted the regression-based approach in the main analyses.

When comparing UKB imputed data (~1.48 million common variants) to UKB WGS (~5.47 million common variants), we observed that further expansion to WGS did not yield additional improvements over imputed data for RICE-RV (Fig. 6, Supplementary Fig. 19). This indicates high accuracy of imputation for common variants in UKB. For RICE-RV, expanding rare variants from ~17 million in WES to 734 million in WGS did not result in noticeable predictive gains (Fig. 6, Supplementary Fig. 15). The lack of improvement may be due to rare variant signals being predominantly located in coding regions, with variants located in noncoding regions contributing less independent signal. A further breakdown of RICE-RV results by coding and noncoding regions revealed that variants in noncoding regions provided limited predictive signal compared to coding region variants (Supplementary Fig. 16).

Fig. 6. Comparison of ancestry-adjusted PRSs from RICE-CV and RICE-RV using UK Biobank (UKB) in two configurations: imputed + WES (imputed genotypes for common variants and whole exome sequencing data for rare variants) and WGS (whole-genome sequencing data for both common and rare variants).

Fig. 6

Traits include six continuous traits: body mass index (BMI), height, high-density lipoprotein cholesterol (HDL), low-density lipoprotein cholesterol (LDL), the natural logarithm of triglyceride (log(TG)), and total cholesterol (TC), and five binary traits: asthma, breast cancer, coronary artery disease (CAD), prostate cancer, and type 2 diabetes (T2D). The training data included only individuals of European ancestry, while the tuning and validation sets contained individuals from all four ancestries. This figure shows results for European ancestry individuals, and results for all other ancestries are provided in Supplementary Fig. 15. Full sample size details for each ancestry are provided in Supplementary Data 2 and 3. Source data are provided as a Source Data file.

All of Us results

We applied RICE and compared its performance against other existing multi-ancestry PRS methods across six continuous traits in the AoU WGS dataset (Methods). Covariate adjustments followed similar procedures as UKB analyses. The training dataset (N = 155,611) consisted of individuals from AFR, AMR, and EUR ancestries, while the tuning (N = 33,342) and validation (N = 33,363) datasets included AFR, AMR, EAS, EUR, MID, and SAS populations (sample size details in Supplementary Data 5, variant count in Supplementary Data 9). Common variants were defined as those with MAF > 0.01 in any of the three ancestries (AFR, AMR, or EUR) in the training dataset. Rare variant analyses were constrained to the exome region for computational efficiency (Methods). The genomic inflation factor was well controlled, with λ1000 ranging from 1 to 1.003 (Supplementary Data 4), with LD score regression intercepts57 confirming minimal population stratification across traits and ancestries. Manhattan and QQ plots for common and rare variant association tests are provided in Supplementary Figs. 17, 18.

Performance patterns largely mirrored UKB (standardized effect sizes in Fig. 7; R2 in Supplementary Fig. 19; quantile differences in Supplementary Fig. 20). Among common variant methods, RICE-CV showed robust performance across most trait–ancestry pairs, typically matching or modestly exceeding the best alternatives (Fig. 7). Two exceptions were HDL in MID and BMI in SAS, where JointPRS performed better, likely reflecting smaller validation sample sizes in these populations. The rare variant component RICE-RV contributed significantly (p<0.05) for several traits, with the clearest signals in lipids and height, though patterns varied by ancestry. Significance was observed in: EUR (HDL, height, LDL, log(TG)); AFR (HDL, height, log(TG), TC); AMR (HDL, height, LDL, log(TG)); EAS (HDL, height, log(TG)); MID (HDL, TC); and SAS (log(TG) only). No ancestry showed a significant RV effect for BMI. Consistent with this, the RV contribution to standardized effects was substantial for lipids, for example, for HDL, the RICE-RV coefficient was ~26–31% of the RICE-CV coefficient in EUR and ~24–40% across non-EUR ancestries; for log(TG), the corresponding ratios were ~22–36.5% (EUR) and ~32–70% (non-EUR). For explained variance, the full RICE model (CV + RV) improved R2 by an average of 1.7% over the best existing method across trait-ancestry pairs (Supplementary Fig. 19). Notable gains included 14.2% for height in EUR, 43.2% for HDL in EAS, and 18.5% for TC in AFR. In contrast to UKB, we did not observe significant R2 gains for lipid traits in EUR within AoU, likely reflecting smaller training sample sizes (average N = 65,549 per lipid trait in AoU vs. N = 91,356 in UKB).

Fig. 7. Predictive performance of ancestry-adjusted PRSs for continuous traits across six ancestral groups from the All of Us (AoU) whole-exome sequencing (WES) data.

Fig. 7

Traits include body mass index (BMI), height, high-density lipoprotein cholesterol (HDL), low-density lipoprotein cholesterol (LDL), the natural logarithm of triglyceride (log(TG)), and total cholesterol (TC). Results are shown for individuals of African (AFR), Admixed American or Latino (AMR), East Asian (EAS), European (EUR), Middle Eastern (MID), and South Asian (SAS) ancestries. Training data consisted of individuals of EUR, AFR, and AMR ancestry, while tuning and validation sets included individuals from all six ancestries. Full sample size details for each ancestry and dataset are provided in Supplementary Data 5. Statistical significance of the RICE-RV component was assessed using percentile bootstrap confidence intervals (10,000 resamples), testing the two -sided alternative HA:βRV≠0; *** indicates the lower bound of the 99% bootstrap confidence interval (CI) > 0 and ** indicates the lower bound of the 95% bootstrap CI > 0. Exact bootstrap p values and CI bounds are provided in the Source Data file. Source data are provided as a Source Data file.

Larger AFR and AMR sample sizes in AoU enabled clearer phenotypic gradients when stratifying by RICE-RV quantiles (below 5%, 30–70%, above 95%) for height and lipid traits (Supplementary Fig. 20b–f). For example, AFR individuals in the top 95% quantile of height were on average 0.07 units higher than those in the middle group and 0.41 units higher than those in the bottom 5% quantile. However, separations in EAS, SAS, and MID individuals were less pronounced due to smaller sample sizes (Supplementary Fig. 20b–f).

Evaluation of AoU PRS on UKB

To assess the cross-dataset portability of RICE PRSs, we applied models trained on AoU data (using AFR, AMR, and EUR individuals for training; Methods) to the UKB validation dataset and compared performance against validation within AoU (standardized effect sizes in Fig. 8; R2 in Supplementary Fig. 21). Analyses focused on six continuous traits across AFR, AMR, EUR, and SAS ancestries (full sample sizes in Supplementary Data 2 and 5). RICE PRSs demonstrated strong portability, with effect sizes from AoU-trained models remaining robust when evaluated in UKB (Fig. 8). Average standardized effect sizes for RICE-CV were 0.241 in AoU and 0.247 in UKB across traits and ancestries, while those for RICE-RV were 0.059 in AoU and 0.053 in UKB. RICE-RV retained significant associations (p<0.05) in UKB for many of the same traits as in AoU, including HDL, height, LDL, and log(TG) in EUR; HDL, height, log(TG), and TC in AFR; HDL, height, LDL, and log(TG) in AMR; and log(TG) in SAS (Fig. 8). No significance was observed for BMI in either dataset. Notably, R2 values were often higher in UKB than in AoU, with an average relative increase of 45.0% for the full RICE model (Supplementary Fig. 21). This enhanced performance in UKB may reflect superior phenotype quality or measurement consistency in that cohort, highlighting RICE’s generalizability across diverse datasets.

Fig. 8. Predictive performance of RICE trained on All of Us (AoU) data and evaluated on both AoU and UK Biobank (UKB) validation datasets.

Fig. 8

Analyzed traits include body mass index (BMI), height, high-density lipoprotein cholesterol (HDL), low-density lipoprotein cholesterol (LDL), log-transformed triglycerides (log(TG)), and total cholesterol (TC). Results are shown for African (AFR), Admixed American/Latino (AMR), European (EUR), and South Asian (SAS) ancestries. Training used EUR, AFR, and AMR individuals from AoU (sample sizes in Supplementary Data 5), with validation on AFR, AMR, EUR, and SAS individuals from AoU and all UKB individuals (Supplementary Data 2). Statistical significance of the RICE-RV component was assessed using percentile bootstrap confidence intervals (10,000 resamples), testing the two -sided alternative HA:βRV≠0; *** indicates the lower bound of the 99% bootstrap confidence interval (CI) > 0 and ** indicates the lower bound of the 95% bootstrap CI > 0. Exact bootstrap p values and CI bounds are provided in the Source Data file. Source data are provided as a Source Data file.

Discussion

Understanding the role of rare genetic variants is crucial for unraveling the genetic architecture of complex traits. We introduced RICE, a PRS framework that integrates both common and rare genetic variants to enhance genetic risk prediction across diverse ancestries. Using large-scale sequencing data from UKB and AoU studies, RICE significantly improves predictive accuracy compared to leading common variant PRS methods, particularly for traits with distinct rare variant architectures. For lipid traits, RICE achieved substantial gains (e.g., up to ~8% increase in R2), enabling more precise risk stratification across multiple populations.

Our findings showed that RICE-CV, the common variant component, consistently delivered robust predictive accuracy, often slightly exceeding or matching existing common variant PRS methods across ancestries and traits (Figs. 4, 7, Supplementary Figs. 7, 12, and 19). By combining multiple PRS approaches, RICE-CV captured a broader spectrum of polygenic signals. This ensemble framework highlights that while no single method is best for all scenarios22,25,55, integrating diverse approaches can optimize predictions across populations.

Importantly, RICE-RV enhanced predictive power by incorporating rare variants, particularly improving risk prediction for lipid-related traits, while showing limited gains for BMI (Figs. 4, 7, Supplementary Figs. 7, 12, and 19). These results highlight that the predictive value of rare variants is not universal but strongly dependent on both the underlying genetic architecture and available statistical power. For lipid traits, our findings align with previous heritability studies indicating that rare variants contribute more significantly to certain traits36. For instance, Weiner et al.36 estimated that the ratio between exome burden heritability and common variant heritability was approximately 0.231 for LDL, but only 0.048 for BMI. In our models, this substantial improvement reflected RICE’s ability to capture distinct rare variant architectures, aggregating signals from both established high-impact genes (e.g., APOB, LDLR) and broader polygenic rare variation (Supplementary Fig. 10). In contrast, the negligible improvements for BMI suggest that its genetic architecture is likely dominated by infinitesimal common variant effects, or that rare variant effects are too diffuse to be effectively captured by gene-centric burden testing. Furthermore, for binary traits or diseases, the limited gains should be interpreted in the context of statistical power. Unlike quantitative traits, where the effective sample size includes the entire cohort, analyses in population-based cohorts like the UK Biobank are restricted to the subset of cases. The resulting lower case counts substantially reduce the power to accurately estimate rare variant weights, limiting the utility of RICE-RV in this specific study design.

Traditional evaluation metrics like R2 may underestimate the impact of rare variants because they focus on average effects across the population. The regression coefficient of the standardized PRS on the standardized trait provides a direct measure of the PRS effect size (equivalent to correlation and subject to attenuation bias due to PRS measurement error61), offering an interpretable assessment of predictive performance (e.g., per-SD effect). Our analyses demonstrated that, despite affecting fewer individuals, rare variants can have significant effects on trait values, reinforcing the importance of including them in PRS models.

We emphasize that the reporting of standardized effect sizes is intended to provide a measure of the expected phenotypic shift for individual carriers, an interpretation that is often more intuitive in clinical settings than variance explained. However, it is important to acknowledge the mathematical coupling between these metrics; for a standardized PRS, the Beta per SD is equivalent to the square root of the heritability explained by the score. Consequently, for rare variants, these effect sizes should be interpreted with caution regarding population-level impact. A substantial per-SD effect can be observed even when the overall R2 is modest, as the low frequency nature of rare variants inherently limits the total variance they can explain across the entire population.

One contribution of our study is the comprehensive assessment of rare variants’ contributions across large and diverse datasets. Rare variants are often overlooked in risk prediction models due to complexity and challenges in modeling them1,2. Existing efforts have focused on combining high-penetrance genes with common variant PRSs62–64, as seen in breast cancer risk prediction models that include genes such as BRCA1, BRCA2, PALB2, CHEK2, and ATM 62. While these approaches effectively integrate high-effect genes, they may miss modest-risk rare variant sets. By incorporating functional annotations using the STAARpipeline48, RICE-RV can identify rare variant sets beyond those described in existing literature, leveraging larger sample sizes and diverse ancestries to capture effects of rare variants across functionally distinct regions (e.g., predicted loss-of-function (pLoF) or deleterious missense variants). This enables us to capture the varying effect sizes of rare variants across functional regions, as shown in previous studies where loss-of-function variants often exhibit large effects65–68.

Our results further indicated that rare variant signals contributing to PRS were largely captured in coding regions, while the contribution of rare variants in noncoding regions was comparatively weaker (Fig. 6). This observation aligns with findings from recent association testing studies69,70. For example, Gaynor et al.70 reported that the gene-based association signals from WGS data differed by only 1% from those derived from WES plus imputation, suggesting that the most actionable signals for prediction are heavily concentrated in coding regions. Similarly, Karczewski et al.41,69 demonstrated that rare coding variants, particularly pLoF variants, are significantly enriched for disease associations compared to synonymous or non-coding variants. However, this concentration of predictive signal contrasts with recent heritability estimates. Wainschtein et al.71 recently estimated that non-coding regions account for the majority (~79%) of total rare variant heritability due to their vast size. This apparent discrepancy likely reflects the structural limitations of burden-based aggregation for capturing diffuse signals. RICE relies on aggregating variants to amplify signal. In coding regions, the high per-variant heritability enrichment allows burden tests to effectively distinguish risk carriers. In contrast, the heritability in non-coding regions is likely more diffuse and intermixed with a vast number of neutral variants; aggregating these sparser signals over broad genomic windows likely dilutes the true effects with noise, rendering the resulting burden scores non-predictive. Thus, while non-coding regions harbor substantial latent heritability, coding regions currently offer the most accessible signal architecture for effective polygenic risk modeling.

An important feature of RICE to provide a single PRS for common variants applicable across ancestries. Existing methods like CT-SLEB22 and PROSPER25 are often optimized for specific ancestries, requiring ancestry-specific tuning. While these methods may achieve high performance in a target ancestry, they complicate clinical translation by necessitating prior knowledge of a patient’s genetic ancestry. In contrast, RICE-CV uses a mixed-ancestry tuning dataset that combines results from multiple methods, yielding a single PRS applicable across ancestries. Our results showed that this approach achieved higher or comparable performance to the best alternative methods in nearly all traits and ancestries (Fig. 7), simplifying its potential clinical utility.

Our cross-dataset evaluation further demonstrated RICE’s robustness, with AoU-trained PRSs performing comparably or better when applied to UKB, despite differences in sequencing technologies (WGS in AoU vs. imputed + WES in UKB) and reduced variant overlap. Effect sizes remained stable, and R2 often increased in UKB (Fig. 8; Supplementary Fig. 21), potentially due to phenotype quality or measurement consistency in UKB, while rare variant signals showed consistent results with only minor attenuation, highlighting RICE’s consistent predictive performance across real-world datasets. This portability supports RICE’s utility in federated or multi-biobank settings, where training and application cohorts differ.

Our analyses demonstrated that incorporating rare variants in the coding region enhances the prediction of lipid traits, with genes like APOC3, APOB, and LDLR emerging as key contributors in our models, consistent with existing findings in lipid metabolism72–74. For example, rare loss-of-function variants in APOC3 have been associated with lower triglyceride levels and reduced coronary artery disease risk74. This alignment with known biological mechanisms reinforces the validity of our results, highlighting the relevance of these genes in clinical risk assessment and therapeutic strategies, while also offering new opportunities for biological insights. A comparison to high-penetrance genes alone (LDLR, APOB, PCSK9) showed that RICE-RV’s genome-wide approach adds substantial value for lipids (Supplementary Fig. 10).

Despite the promising results, our study has several limitations. First, regarding modeling simplifications: RICE aggregates rare variant into burden scores. While efficient, this approach assumes equal weighting of all variants, potentially overlooking differences in pathogenicity46–48,75–78. Furthermore, while the STAAR framework can detect non-linear signals (e.g., via SKAT47), RICE prioritizes STAAR-Burden (STAAR-B) to align with its additive prediction model; this ensures consistency but may fail to capture complex epistatic interactions, suggesting a need for future non-linear integration methods. Second, regarding data constraints: computational limits restricted rare variant inclusion to sets with p<1×10−3, though sensitivity analyses showed stable performance across thresholds (Supplementary Fig. 11). Additionally, limited case counts for binary outcomes currently restrict utility, necessitating larger disease-specific consortia. Future applications in these cohorts could employ metrics like the net reclassification index to better assess improvements in high-risk individual identification and clinical decision-making. Finally, regarding generalizability: while RICE leverages individual-level data for optimal conditioning, it can adapt to external summary statistics (e.g., large-scale GWAS for CV and STAARpipeline results for RV). However, challenges may arise if these summary statistics are from disparate samples, potentially compromising independence between components. Future extensions could incorporate orthogonalization techniques to maintain component independence when using disparate summary statistics, or adopt summary statistics-based fine-tuning approaches79–82 to eliminate the requirement for an independent tuning data.

In conclusion, we developed RICE, a PRS framework that integrates both common and rare variants within a single model to improve genetic risk prediction across diverse ancestries. Our study demonstrates that incorporating rare coding variants enhances the predictive power of PRSs for lipid-related traits, particularly for individuals carrying these variants who may experience significant impacts on trait values. By providing open-source software for RICE, we offer a practical tool to advance polygenic risk prediction and contribute to precision medicine.

Methods

We developed RICE, a framework that integrates both common and rare genetic variants to enhance polygenic risk prediction across diverse populations. RICE has three primary components: (1) RICE-CV, which constructs a robust PRS for common variants using ensemble regression; (2) RICE-RV, which identifies and incorporates rare variant signals into the PRS; and (3) a final evaluation step that combines the PRSs from RICE-CV and RICE-RV within a regression model to produce an integrated risk prediction (Fig. 1).

Ethics

This study used de-identified data from the UK Biobank under approved application 52008 and from the All of Us Research Program under controlled-tier access via the Researcher Workbench. UK Biobank has research ethics approval from the North West Centre for Research Ethics Committees (REC reference 11/NW/0382). All of Us participants are consented under the All of Us research protocol, approved by the All of Us Institutional Review Board. Additional information on All of Us IRB oversight is available at https://allofus.nih.gov/about/who-we-are/institutional-review-board-irb-of-all-of-us. Analyses were conducted using de-identified data within the respective secure analysis environments.

Data processing and standardization

All analyses employ a three-way data split into independent training, tuning, and validation datasets. The training dataset is used to compute GWAS summary statistics, train common variant PRS models, perform rare variant association testing, compute burden scores, and train rare variant PRS models using penalized regression (LASSO and ridge regression). The tuning dataset is used to train the ensemble learning model, combining PRSs generated by different methods and tuning parameters, using LASSO and ridge regression as base learners. The final validation dataset is a fully independent, held-out test set used only to evaluate the performance of the final PRSs.

To ensure consistency and comparability across ancestries during the tuning and validation stages, we implemented specific standardization procedures for phenotypes and PRSs. For the initial association testing of common variants in GWAS and rare variants in STAARpipeline within the training dataset, we used the original phenotype values. These analyses adjusted for covariates including the top 10 PCs, sex, age, and age squared. Association testing of common variants was performed with REGENIE83. Rare variant association testing using STAARpipeline48 was performed with linear regression for continuous traits and logistic regression for binary traits, aligning with standard association testing practices.

During the ensemble learning for RICE-CV and RICE-RV in the tuning dataset and fitting LASSO and ridge regression models for burden scores in the training dataset (detailed in the later sections), continuous phenotypes were adjusted by regressing out the same set of covariates to obtain residuals. These residuals were used as inputted outcomes in the mentioned procedures. Binary traits remained on their original scale throughout all stages.

In the validation stage, we standardized the residualized continuous phenotypes within each ancestry group to have a mean of 0 and a variance of 1. PRSs generated from existing methods, RICE-CV, and RICE-RV are also standardized within each ancestry to have a mean of 0 and a variance of 1, following methods outlined in prior works4,10 and described in the Supplementary Note. This standardization ensured that both phenotypes and PRSs were on a common scale across ancestries, facilitating fair comparisons during performance evaluation. All PRS methods evaluated in the manuscript followed identical evaluation procedure.

RICE-CV: common variant PRS modeling

The RICE-CV component has two main steps: (1) PRS training and (2) ensemble learning.

PRS training

In RICE-CV, common variants are defined as variants with MAF >0.01 in any of the genetically inferred ancestry groups. Let u^kl denote the estimated effect size of the k-th genetic variant in the l-th ancestry group, with skl as its standard error. We construct the common variant PRS using several established PRS methods.

For single-ancestry analyses (e.g., in a primarily European population), we employ methods such as CT11–13, LDpred254, and Lassosum215, which represent clumping-based, Bayesian, and penalization-based approaches, respectively. For multi-ancestry analyses, we apply CT-SLEB22, JointPRS55, and PROSPER25, covering the same methodological categories. Each method generates multiple PRSs for each individual in the tuning dataset by varying its tuning parameters. Detailed implementations of each method are provided in the “Existing PRS Methods” section.

Ensemble learning

The ensemble learning step combines the candidate PRSs to optimize predictive performance. Specifically, we use LASSO and ridge regression as base learners within a generalized linear regression model54,83.

For continuous traits, the ensemble model is trained to minimize the cross-validated mean squared error (MSE). The process is as follows: for LASSO regression, we train a model on the tuning dataset using the PRS features as predictors. The model minimizes the objective function:

minw(1)1n∑i=1nYi−∑j=1Jwj(1)PRSij2+λ1∑j=1J∣wj(1)∣ 1

where Yi is the residualized phenotypes for individual i regressing out for covariates, PRSij is the j-th candidate PRS for individual i, and w(1)=(w1(1),…,wJ(1))T is the vector of weights estimated for each PRS. Here, n denotes the tuning data sample size, J represents the total number of candidate PRSs, and λ1 is the LASSO regularization parameter.

For ridge regression, we minimize the objective function:

minw21n∑i=1nYi−∑j=1Jwj2PRSij2+λ2∑j=1J(wj2)2, 2

where w(2)=(w1(2),…,wJ(2))T are the weights estimated for each PRS and λ2 is the ridge regularization parameter.

The optimal regularization parameters λ1 and λ2 are selected by minimizing the cross-validated MSE (default fold = 10) over a default grid of regularization parameters provided in glmnet84. With optimized weights for each base learner, we generate predictions for each individual in the tuning dataset as: Y^i(1)=∑j=1Jwj(1)PRSij for LASSO and Y^i(2)=∑j=1Jwj(2)PRSij for ridge. The ensemble learning algorithm then estimates the optimal weights α=α1,α2T to combine these predictions from the base learners using linear regression. This is done by solving:

minα1n∑i=1nYi−∑m=12αmY^i(m)2 3

The final ensemble prediction for individual i is a weighted combination of base learner predictions: Y^i=α1Y^i(1)+α2Y^i(2)=∑j=1J(α1wj1+α2wj2)PRSij.

Similarly for binary traits, the ensemble model combines predictions from LASSO and ridge regression using a generalized linear regression model. For LASSO logistic regression, we minimize the objective function:

minw(1)−1n∑i=1nYilogpi+1−Yilog1−pi2+λ1∑j=1J∣wj(1)∣ 4

where pi=σ(∑j=1Jwj(1)PRSij) is the predicted probability for individual i, with σx=exp(x)/{1+expx} being the logistic function and Yi is the binary phenotype (0 or 1) for individual i.

For ridge logistic regression, we minimize the objective function:

minw(2)−1n∑i=1nYilogpi+1−Yilog1−pi2+λ2∑j=1J(wj(2))2 5

where pi=σ(∑j=1Jwj(2)PRSij). Optimal regularization parameters λ1 and λ2 are selected by minimizing the cross-validated AUC (default fold = 10). Using the optimal weights for each base learner, the predicted log odds are generated for LASSO and ridge (Y^(1) and Y^(2) respectively). The ensemble learning algorithm estimates the optimal weights of the predictions, α=(α1,α2)T, by minimizing the objective function:

minα−1n∑i=1nYilogpi+1−Yilog1−pi2 6

where pi=σ(∑j=12αjY^i(j)). The final prediction of the ensemble learning algorithm is then: Y^CV=α1Y^(1)+α2Y^(2)=α1w(1)PRSCV+α2w(2)PRSCV, where PRSCV=[PRS1,PRS2,…,PRSJ].

RICE-RV: rare variant PRS modeling

The RICE-RV component focuses on identifying and integrating signals from rare variants into the PRS to enhance genetic risk prediction. It has four main steps: (1) fitting the null model and performing rare variant association testing, (2) modeling significant rare variant sets using penalized regression, and (4) combining PRSs through ensemble learning. Steps 1–2 are based on STAARpipeline focusing on association testing. Steps 3–4 focuses on building the PRS model using the select rare variants sets.

Null model and association testing

We first fit a null model using STAARpipeline to adjust for covariates and the predicted common variant PRS from RICE-CV. This step ensures that rare variant signals are independent of common variant effects. For each individual i, let Yi denote the phenotype of interest, which can be continuous (e.g., height) or binary (e.g., disease status). We model the conditional mean of Yi using:

gμi=α0+QiTα 7

where gμi is the link function (identify for continuous traits, logit for binary traits). α0 is the intercept term, Qi=(Qi1,…,Qiq)T represent q covariates, such as age, sex, PCs and the common variant PRS from RICE-CV (Y^CV,i).

After fitting the null model, we perform rare variant association testing using STAARpipeline, which is specifically designed for large-scale sequencing studies. STAARpipeline adjusts for population structure, corrects for case-control imbalances, and incorporates functional annotations to increase the power of rare variant tests.

Rare variants are grouped into biologically meaningful sets based on functional categories, increasing the power to detect associations by aggregating variants likely to affect gene function. To capture orthogonal signals, RICE-RV uses STAARpipeline to perform association tests on rare variant sets, conditional on the RICE-CV PRS. STAARpipeline defines gene-centric rare variant sets by functional categories. For variants in coding regions, it analyzes five categories per gene: (1) putative loss of function (stop gain, stop loss and splice); (2) missense, (3) disruptive missense, (4) putative loss of function and disruptive missense; and (5) synonymous. For variants in noncoding regions, it tests eight categories: (1) promoter overlaid with CAGE sites, (2) promoter overlaid with DHS sites, (3) enhancer overlaid with CAGE sites, (4) enhancer overlaid with DHS sites, (5) UTR, (6) upstream region, (7) downstream region, and (8) ncRNA. In our study, we use variants in coding regions for WES data and variants in coding and noncoding regions for WGS data.

Within each set, STAARpipeline performs annotation-weighted tests, aggregating information across variants and weighting each variant based on its functional annotation. Multiple test statistics, including burden, SKAT, and ACAT-V, are computed, and an omnibus p value is generated using ACAT to combine correlated p values. Within RICE-RV, we specifically restrict our identification to the p values derived from the annotation-weighted burden tests, referred to as STAAR-B(1,1) p values in STAARpipeline to ensure that the identified rare variant signals align with the additive burden-based architecture used in the downstream RICE prediction model. We identify significant rare variant sets as those with STAAR-B (1,1) p values less than p<1×10−3.

Burden score construction and penalization regression

For each significant rare variant set, we compute a burden score for each individual to summarize the collective effect of the rare variants within that set. Let K denote the total number of significant rare variant sets identified. For the k-th rare variant set (k=1, 2,…,K), let mk be the number of rare variants in the k-th set. For individual i, let Gik=(Gi1k,Gi2k,…,Gimkk)T denote the vector of genotypes for the rare variants within the k-th set. Here, Gijk represents the genotype of individual i for the j-th rare variant in set k, coded as the number of minor alleles (e.g., 0, 1, or 2). The burden score for individual i for the k-th rare variant set is defined as:

bik=∑j=1mkGijk=GikT1mk 9

where 1mk is a vector of ones of length mk. This burden score bik represents the total number of minor alleles carried by individual i across all rare variants in the k-th rare variant set.

We then model the relationship between phenotype Yi and burden scores using penalized regression methods, specifically LASSO and ridge regression. For continuous traits, we estimate the coefficients γ=γ1,γ2,…,γKT by minimizing the following objective function:

minγ1ntr∑i=1ntrYi−∑k=1Kγkbik2+λPγ 10

where Yi is the residualized phenotypes regressing out for covariates, ntr is the total of number individuals in the training dataset, and λ is the regularization parameter controlling the penalty strength. Pγ is the penalty function (∑k=1K∣γk∣ for LASSO or ∑k=1Kγk2 for ridge regression).

For binary traits, we use a logistic regression framework with penalization, estimating γ by minimizing:

minγ−1ntr∑i=1ntrYilogpi+1−Yilog1−pi2+λPγ 11

where pi=σ(∑k=1Kγkbik) is the predicted probability for individual i, with σx=exp(x)/{1+expx} being the logistic function. For fitting the penalized regression models on burden scores, we used the glmnet84 package in R (version 4.1.8). The PRS for specific individual i given the estimated γ^ under a specific penalization λ, can be calculated as PRSi=∑k=1Kγ^kbik.

Ensemble learning to combine PRSs for RICE-RV

After computing the burden score PRSs, we follow a similar ensemble learning step on the tuning dataset as in RICE-CV to compute the final PRS for RICE-RV. Using a generalized linear regression model, we combine predictions from LASSO and ridge regression to optimizing predictive performance. For rare variants, we select regularization parameters λ1 (for Lasso) and λ2 (for ridge) via a grid search on the full tuning dataset. We use default grid values provided by glmnet84 and choose the optimal regularization parameter based on R2 or AUC. The final PRS for RICE-RV is then denoted as: Y^RV=PRSRV wRV, where Y^RV is the predicted response from the rare variants. PRSRV=[PRS1,PRS2,…,PRSJ] represents the matrix of PRSs generated by based on the burden score, where the j-th PRS, PRSj=(PRS1j,..,PRSnj)T, corresponds to a PRS derived from a specific penalized regression model for n individuals in the tuning dataset. wRV is the vector of optimal weights derived from the ensemble learning model.

Evaluation of RICE

We evaluated RICE using standardized effect sizes and R2/AUC. First, to establish the combined RICE score, we model the joint prediction on the tuning dataset. For continuous traits, we fit: Y=θ^1Y^CV+θ^2Y^RV, and for binary traits: logit(P(Y=1))=α0+θ^1Y^CV+θ^2Y^RV. Here Y^CV and Y^RV are the standardized RICE-CV and RICE-RV PRSs, respectively. Crucially, the weights θ^1 and θ^2 are estimated exclusively in the tuning dataset and fixed. These fixed weights are then applied to the validation dataset to compute the final score. R2 (for continuous traits) and AUC (for binary traits) are calculated in the validation dataset, adjusting for covariates.

Second, to report standardized effect sizes of RICE-CV and RICE-RV, we fit regression models directly in the validation dataset. This step is intended solely to quantify component effect sizes in an independent sample. For continuous traits, we fit linear regression: Y=β^1Y^CV+β^2Y^RV. For binary traits, we fit logistic regression logit(P(Y=1))=α0+β^1Y^CV+β^2Y^RV+QiTα, where Qi represent other covariates, such as age, sex, PCs. These estimated coefficients β^1 and β^2 are the reported standardized effect sizes.

We assessed the significance of the standardized effect size, R2, and AUC using bootstrap confidence intervals. For the standardized effect-size of RICE-RV, we performed 10,000 bootstrap resamples, and tested whether the coefficient differed significantly from zero at the 95% and 99% confidence levels. For R2 and AUC, we conducted pairwise comparisons with RICE and the best alternative method for each trait-ancestry combination. Using 10,000 bootstrap resamples, we tested whether the pairwise difference in R2 or AUC were significantly different from zero at the 95% and 99% confidence levels.

Existing PRS methods

CT11–13 is a method that first removes variants in high LD with an index variant that has the lowest p value within a specified base-pair window. From the remaining set of index SNPs, multiple p value thresholds are applied to generate PRSs. The optimal PRS and corresponding p value threshold are then chosen based on performance in the tuning dataset. We implemented CT using plink1.985 with an LD threshold r2 of 0.5 (--clump-r2), a window size of 500 kb, and nine p value thresholds ranging from 5×10−8 to 1.

Lassosum215 implements LASSO regression using GWAS summary statistics and LD information. Lassosum2 penalizes effect sizes according to LD by minimizing the objective function: βT1−sR+sIβ−βTr+λ1∣∣β∣∣1, where β is the vector of effect sizes, R is the LD matrix, I is the identity matrix, r is the vector of marginal correlations from summary statistics, and s is a tuning parameter to ensure convexity. The parameter λ1 controls L1 penalty, with∣β∣1 denote the L1 norm of β. We implemented Lassosum2 with 10 different values of s (ranging from 0.5 to 1000) and 30 values of λ1. As in CT, the optimal combination of s and λ1 was chosen based on predictive performance in the tuning dataset.

LDpred254 utilizes a Bayesian framework with a spike-and-slab prior for variant effects βj. The prior density is given by p(βj∣π)=(1−π)δ(βj=0)+π×N(0,h2Mπ), where δ is the Dirac delta function, π is the proportion of casual variants, h2 is the total heritability, M is the total number of variants, and π is the proportion of nonzero variants. Posterior effect sizes are estimated using Markov Chain Monte Carlo (MCMC) sampling. We conducted a grid search over 15 values of h2 (from 0.1 to 1.5) times the heritability estimated by LD score regression, 17 values of π ranging from 1×10−4 to 1, and set the sparse setting to “False”. Both LDpred2 and Lassosum2 were implemented in R using the bigsnpr package (version 1.12.6). The optimal values of h2 and π were selected using the tuning dataset.

CT-SLEB22 is a recent multi-ancestry extension of CT method that consists of three steps: two-dimensional CT, calibration of regression coefficients using Empirical Bayes, and ensemble learning. Two-dimensional CT and Empirical Bayes are implemented between pairs of ancestries, designating one as the target population and the other as the reference population. For our analysis of the AoU data, which included three ancestries in the training dataset (AFR, AMR, and EUR), CT-SLEB was implemented twice, treating AFR and AMR as target populations with EUR as the reference population. We further implemented the ensemble learning step on the tuning data to include both sets of PRSs, producing a single prediction for all ancestries. The hyperparameters required for CT-SLEB are identical to those used in CT. We implemented CT-SLEB with eight different p value thresholds ranging from 5×10−8 to 1, two window sizes (50 kb and 100 kb), and six LD threshold (r2 = 0.01, 0.05, 0.1, 0.2, 0.5, 0.8).

JointPRS55 is a multi-ancestry PRS method that extends the PRS-CSx24 method. JointPRS assumes variant effect size βj across multiple populations using a correlated Gaussian prior: βj~N(0,ψjMΣM), where MΣM accounts for correlations across ancestries and differences in the sample size, and ψj is a variant-specific hyper-parameter with a hyper-prior ψj~Gamma(1,δj) and δj~Gamma(12,ϕ). We implemented JointPRS-auto version, which assumes ϕ~Cauchy(0, 1). JointPRS-auto requires only GWAS summary statistics from the training data and does not need tuning data for parameter estimation. LD information was provided using the 1000 G reference panel included with the PRS-CSx24 software.

PROSPER25 is an ensemble penalized regression multi-ancestry PRS method. It involves three steps: (1) performing single-ancestry Lassosum2 to estimate optimal parameters, second, (2) constructing a joint analysis across ancestries using penalized regression, and (3) combining all PRSs from the second step using ensemble learning with SuperLearner. We implemented single ancestry Lassosum2 with the default five values for s and five values for λ1, as described in the Lassosum2 section. The joint multi-ancestry analysis was performed similarly, using default values for the penalty parameters for λ and c following PROSPER GitHub guidance. Ensemble regression was performed on the tuning data using SuperLearner.

Simulation study

We conducted large-scale simulations using simulated traits generated from chromosome 22 of the UKB WES dataset. The training dataset included only individuals with EUR ancestry, with sample sizes of 49,173 or 98,343. The remaining individuals, comprising those with EUR, AFR, AMR, and SAS ancestry, were evenly divided into tuning and validation sets.

Causality was assigned to both randomly selected common variants (MAF > 0.01 in the UKB WES dataset) and to rare variant sets. The proportion of causality varied among three levels: 1%, 5%, and 20%. We explored two different strategies of assigning causality to rare variants within a causal rare variant set. First, all rare variants within a rare variant set are causal and used in the construction of the causal burden score. Second, a proportion (πi~Uniform0.1, 0.9)) of rare variants within causal rare variant set i are used in construction of the causal burden score. The heritability was kept constant across simulations, with the common variant heritability (hCV2) to 5% to ensure sufficient statistical power to evaluate model mechanics within a single-chromosomal simulation framework. To reflect realistic genetic architectures, we calibrated the rare variant heritability based on recent empirical estimates of the ratio between burden and common variant heritability (median around 1:1236). Accordingly, we set rare variant heritability hRV2=0.42%. Additionally, we considered two scenarios of negative selection: strong negative selection and no negative selection.

Let uCV,q denote the standardized effect size of the qth causal common variant, and uRV,k denote the standardized effect size for the k-th causal rare variant burden. Under strong negative selection scenario, the standardized effect sizes were drawn from normal distributions: uCV,q≈N0,h2CVCCV,uRV,k≈N(0,h2RVCRV), where CCV and CRV are the number of causal common variants and rare variants, respectively.

Phenotypes were simulated using the following linear model:

Yi=∑q=1CCVGiqσ^2quCV,q+∑k=1CRVbikσ^2kuRV,k+ϵi 12

where Yi is the phenotype for individual i, Giq is the genotype of individual i at the q-th causal common variant, bik is the burden score of individual i for the k-th causal rare variant set, and ϵi is the residual error term. The variance σ^2q and σ^2k are defined as follows: σ^2q=2fq(1−fq), where fq is the effect allele frequency for the q-th causal common variant, and σ^2k is calculated similarly for the k-th rare variant sets. Under the no negative selection scenario, we set σ^2q=σ^2k=1. Each simulation scenario had 100 simulated traits, and the final presented results is the average of the 100 validation dataset results.

UKB imputed + WES and WGS analysis

We analyzed the UKB imputed (Field #22828), WES (Field #23156), and WGS (Fields #24304 and #24305) datasets, focusing on participants from four ancestries: AFR, AMR, EUR, and SAS. EAS ancestry was excluded due to limited sample size in the UKB data, which led to unstable performance in the evaluation. Our analyses included six continuous traits: BMI, height, HDL, LDL, log(TG), and TC, and five binary traits: asthma, breast cancer, CAD, prostate cancer, and T2D.

Genetically inferred ancestry groups were defined using the 1000 Genomes Project Phase 3 (1000G) dataset56, which includes 2,504 individuals across five ancestry groups: 661 AFR, 347 AMR, 504 EAS, 503 EUR, and 489 SAS. We first performed principal component analysis on all 2,504 samples using plink v2.085, using a set of selected variants provided from gnomAD to capture population structure. Using the first five PCs, we trained a random forest classifier to accurately assign each individual in the 1000 G dataset to their respective ancestry group. Next, we projected the UKB WES, Imputed, and WGS datasets onto this PC space and applied the pre-trained random forest classifier to assign genetically-inferred ancestry groups to each UKB sample based on maximum probability.

Participants were randomly divided into independent training, tuning, and validation sets using stratified sampling within each ancestry, allocating 70%, 15%, and 15% of the total sample size, respectively. Since EUR is the majority ancestry in UKB, we restricted the training dataset to EUR participants, while both EUR and non-EUR participants were included in the tuning and validation sets. Detailed sample sizes for each group are provided in Supplementary Data 2 and 3. Standard quality control procedures were applied to both common and rare variants following previous studies38,76,86–88 (Supplementary Note). All analyses, including GWAS, rare variant association testing, and ensemble learning within RICE-CV and RICE-RV, were adjusted for control covariates: the first 10 PCs, sex, age, and age squared. PCs used were the pre-computed ones provided by UKB (Field #22009).

We evaluated PRS prediction performance on the validation set across five methods: CT, LDpred2, Lassosum2, RICE-CV, and RICE-RV. Optimal hyperparameters for CT, LDpred2, and Lassosum2 were chosen based on performance in the tuning dataset. Both phenotypes and PRSs were standardized to have mean 0 and variance 1 within each ancestry, as described in the Phenotype and PRS Standardization section. Performance of the standardized PRSs was assessed by fitting a linear regression model between the standardized phenotype and standardized PRS. For binary traits, PRS standardization followed the same procedure as for continuous traits, and prediction performance was evaluated using logistic regression with control covariates.

For the CT, LDpred2, and Lassosum2 methods for common variant PRS construction, LD reference panels were estimated from 3,000 individuals randomly sampled from the training dataset for both Imputed and WGS analyses. For LDpred2 and Lassosum2 in the imputed data analysis, LD was estimated using variants included in the HapMap3 (HM3) and Multi-Ethnic Genotyping Array (MEGA). In contrast, for the WGS analysis, computational constraints restricted LDpred2 and Lassosum2 to the subset of variants included in the HM3 array. CT was implemented using all variants available in the reference panel without restriction. For rare variant PRS construction, the STAARpipeline differed based on sequencing data: for the WES data (used in the Imputed + WES configuration), we analyzed rare variants located in the coding region of protein-coding genes. For the WGS data, we included rare variants located in both coding and noncoding regions of protein-coding genes.

AoU analysis

We analyzed the AoU dataset using WGS data version 7.1. For common variants, we used the Allele Count/Allele Frequency (ACAF) Threshold callset provided in the AoU Researcher Workbench, which include variants with MAF > 0.01 or allele count > 100 in any of the computed ancestry groups. Given that the training dataset included only AFR, AMR, and EUR ancestries, we further restricted common variants to those with MAF > 0.01 in any of these three ancestries. For rare variants, we focused on exonic regions using the “Exome” callset from the AoU Researcher Workbench. We employed the genetically inferred ancestry groups provided by AoU, based on variants from gnomAD69, the Human Genome Diversity Project89, and 1000G56. The ancestries included in our analysis were AFR, AMR, EAS, MID, EUR, and SAS. We examined six continuous traits: BMI, height, HDL, LDL, log(TG), and TC. Genetically inferred ancestry groups (EUR, EAS, SAS, AMR, MID, AFR) were directly provided by the All of Us Research Program40. For AoU, we computed PCs based on pre-computed PCA loadings calculated from the samples of the 1000G reference panel.

Data were randomly split using stratified sampling within ancestries into training, tuning, and validation, comprising 70%, 15%, and 15% of total sample size, respectively. Since AFR, AMR, and EUR are the majority ancestries in the AoU database, we included only these three ancestries in the training dataset. For the tuning and validation datasets, we included all available ancestries in AoU. The detailed sample size breakdown is provided in Supplementary Data 5. Both common and rare variants underwent standard quality control procedures (Supplementary Note). All analyses, including GWAS, rare variant association testing, and ensemble learning within RICE-CV and RICE-RV, were adjusted for control covariates: the first 10 PCs, sex, age, and age squared.

We compared PRS prediction performance on the validation set across five methods: CT-SLEB, JointPRS, PROSPER, RICE-CV, and RICE-RV. Following a similar procedure as in the UKB analyses, we standardized the phenotypes and PRSs to have a mean of 0 and variance of 1 within each ancestry. The detailed standardization procedure is provided in the Phenotype and PRS Standardization section. The performance of the standardized PRSs was assessed by fitting a linear regression model between the standardized phenotype and the standardized PRS.

LD reference data for JointPRS and PROSPER was publicly provided and built using the 1000G reference data. Specifically, the provided LD for PROSPER was constrained to the HM3 + MEGA array, while the provided LD for JointPRS was constrained to the HM3 array. Consequently, our analyses using these methods were limited to these specific variant subsets. In contrast, for CT-SLEB, the LD reference was derived directly from a random subset of 3000 individuals from each ancestry (AFR, AMR and EUR) in the training dataset. This approach allowed us to implement CT-SLEB on the full set of available common variants. RICE-RV was implemented using rare variants located in coding regions of protein-coding genes utilizing STAARpipeline.

Cross-dataset portability analysis

To evaluate PRS portability, we applied AoU-trained models (using AFR, AMR, and EUR ancestries for training, as described above) to the UKB validation dataset. Common variants were scored using UKB imputed genotypes, and rare variants using UKB WES data, with the same quality control and covariate adjustments (top 10 PCs, sex, age, age squared). RICE models were trained using all available variants in the AoU dataset. When applying these models to the UKB validation dataset, we utilized the subset of variants present in both datasets (details on variant counts in Supplementary Data 9). Performance was assessed on the six shared continuous traits (BMI, height, HDL, LDL, log(TG), TC) across AFR, AMR, EUR, and SAS ancestries, using the standardized effect size, R2, and bootstrap significance (10,000 resamples) as in other evaluations.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

41467_2026_72185_MOESM1_ESM.pdf (14.7MB, pdf)

Supplementary Figures and Supplementary Notes

41467_2026_72185_MOESM2_ESM.pdf (21.7KB, pdf)

Description of Additional Supplementary Files

Supplementary Data (98.3KB, xlsx)
Reporting Summary (107.5KB, pdf)

Source data

Source Data (879.2KB, xlsx)

Acknowledgements

The analysis utilized the high-performance computation Biowulf cluster at National Institutes of Health (NIH), USA. This research has been conducted using the UK Biobank Resource under Application Number 52008. We gratefully acknowledge All of Us participants for their contributions, without whom this research would not have been possible. We also thank the National Institutes of Health’s All of Us Research Program for making available the participant level data examined in this study. We want to thank Jin Jin, Jingning Zhang, and Xiaoyu Wang for sharing their code for PRSs approaches. This research was supported by NIH Intramural Research Program (J.W., X.H., W.W., K.Y., P.K. and H.Z.), NIH grants 1R01HL173044 and 1R01AG085581 (X.L.), the research start-up funds from the Department of Biostatistics and the Department of Genetics at the University of North Carolina at Chapel Hill (X.L.), and NIH Training Grant T32GM135117 and NSF Graduate Research Fellowship DGE-2140743 (T.C.). This research was supported in part by the Intramural Research Program of the NIH. The contributions of the NIH authors were made as part of their official duties as NIH federal employees, are in compliance with agency policy requirements, and are considered Works of the United States Government. However, the findings and conclusions presented in this paper are those of the authors and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.

Author contributions

J.W., X.L., and H.Z. conceived the project. J.W. carried out all data analyses with supervision from X.L. and H.Z.; T.C. provided code and performed phenotype querying for the All of Us analysis; X.H. processed and cleaned rare variants in the All of Us analysis; J.W., X.L., and H.Z. drafted the manuscript, and T.C., X.H., W.W., P.K., and K.Y. provided comments. All co-authors reviewed and approved the final version of the manuscript.

Peer review

Peer review information

Nature Communications thanks Jian Zeng, and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Data availability

UK Biobank phenotype data, WES data (Field #23156), imputed data (Field #22828), and WGS data (Fields #24304 and #24305) can be accessed through the UK Biobank research analysis platform [https://ukbiobank.dnanexus.com/landing]. All data used in this research are publicly available to registered researchers through the UKB data-access protocol and who are listed as collaborators on UKB-approved access applications. All of Us phenotype data, WES data, and WGS data can be accessed through the All of Us research workbench [https://workbench.researchallofus.org/login] (version 7.1). All data used in this research are publicly available to registered researchers with controlled tier access through the All of Us data-access protocol. Data generated in this study are largely available in the Supplementary Data or Source Data files. Large results generated in this study, including common-variant GWAS summary statistics, rare-variant association test summary statistics, and the RICE-CV and RICE-RV model weight files, are available through a Harvard Dataverse dataset90. Source data are provided with this paper.

Code availability

Simulation and data analyses code are archived on Zenodo (v1.0.091). The corresponding GitHub repository is https://github.com/jwilliams10/RareVariantPRS. Software and tutorial to implement RICE are archived on Zenodo (v1.0.092) and available on GitHub [https://github.com/jwilliams10/RICE].

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors jointly supervised this work: Xihao Li, Haoyu Zhang.

Contributor Information

Jacob Williams, Email: jacob.williams@nih.gov.

Xihao Li, Email: xihaoli@unc.edu.

Haoyu Zhang, Email: haoyu.zhang2@nih.gov.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-72185-2.

References

  • 1.Chatterjee, N., Shi, J. & García-Closas, M. Developing and evaluating polygenic risk prediction models for stratified disease prevention. Nat. Rev. Genet.17, 392–406 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Kachuri, L. et al. Principles and methods for transferring polygenic risk scores across global populations. Nat. Rev. Genet.25, 8–25 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Patel, A. P. et al. A multi-ancestry polygenic risk score improves risk prediction for coronary artery disease. Nat. Med.29, 1793–1803 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ge, T. et al. Development and validation of a trans-ancestry polygenic risk score for type 2 diabetes in diverse populations. Genome Med.14, 1–16 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zhang, H. et al. Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses. Nat. Genet.52, 572–581 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Mavaddat, N. et al. Polygenic risk scores for prediction of breast cancer and breast cancer subtypes. Am. J. Hum. Genet.104, 21–34 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Shieh, Y. et al. Breast cancer screening in the precision medicine era: risk-based screening in a population-based trial. J. Natl. Cancer Inst.109, djw290 (2017). [DOI] [PubMed]
  • 8.Hao, L. et al. Development of a clinical polygenic risk score assay and reporting workflow. Nat. Med.28, 1006–1013 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Linder, J. E. et al. Returning integrated genomic risk and clinical recommendations: the eMERGE study. Genet. Med. 25, 100006 (2023). [DOI] [PMC free article] [PubMed]
  • 10.Chen, T. et al. Genomic insights for personalised care in lung cancer and smoking cessation: motivating at-risk individuals toward evidence-based health practices. EBioMedicine110, 105441 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Wray, N. R., Goddard, M. E. & Visscher, P. M. Prediction of individual genetic risk to disease from genome-wide association studies. Genome Res.17, 1520–1528 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Purcell, S. M. et al. Common polygenic variation contributes to risk of schizophrenia and bipolar disorder. Nature460, 748–752 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Privé, F., Vilhjálmsson, B. J., Aschard, H. & Blum, M. G. B. Making the most of clumping and thresholding for polygenic scores. Am. J. Hum. Genet.105, 1213–1221 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Newcombe, P. J., Nelson, C. P., Samani, N. J. & Dudbridge, F. A flexible and parallelizable approach to genome-wide polygenic risk scores. Genet. Epidemiol.43, 730–741 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Privé, F., Arbel, J., Aschard, H. & Vilhjálmsson, B. J. Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores. Hum. Genet. Genom. Adv.3, 100136 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mak, T. S. H., Porsch, R. M., Choi, S. W., Zhou, X. & Sham, P. C. Polygenic scores via penalized regression on summary statistics. Genet. Epidemiol.41, 469–480 (2017). [DOI] [PubMed] [Google Scholar]
  • 17.Vilhjálmsson, B. J. et al. Modeling linkage disequilibrium increases accuracy of polygenic risk scores. Am. J. Hum. Genet.97, 576–592 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Lloyd-Jones, L. R. et al. Improved polygenic prediction by Bayesian multiple regression on summary statistics. Nat. Commun.10, 1–11 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ge, T., Chen, C. Y., Ni, Y., Feng, Y. C. A. & Smoller, J. W. Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat. Commun.10, 1–10 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Song, S., Jiang, W., Hou, L. & Zhao, H. Leveraging effect size distributions to improve polygenic risk scores derived from summary statistics of genome-wide association studies. PLoS Comput. Biol. 16, e1007565 (2020). [DOI] [PMC free article] [PubMed]
  • 21.Zhou, G. & Zhao, H. A fast and robust Bayesian nonparametric method for prediction of complex traits using summary statistics. PLoS Genet.17, e1009697 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Zhang, H. et al. A new method for multiancestry polygenic prediction improves performance across diverse populations. Nat. Genet.55, 1757–1768 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Hoggart, C. J. et al. BridgePRS leverages shared genetic effects across ancestries to increase polygenic risk score portability. Nat. Genet.56, 180–186 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Ruan, Y. et al. Improving polygenic prediction in ancestrally diverse populations. Nat. Genet.54, 573–580 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Zhang, J. et al. An ensemble penalized regression method for multi-ancestry polygenic risk prediction. Nat. Commun.15, 1–14 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Jin, J. et al. MUSSEL: Enhanced Bayesian polygenic risk prediction leveraging information across multiple ancestry groups. Cell Genom.4, 100539 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Breast Cancer Association Consortium Breast cancer risk genes—association analysis in more than 113,000 women. N. Engl. J. Med.384, 428–439 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Peloso, G. M. et al. Rare protein-truncating variants in APOB, lower low-density lipoprotein cholesterol, and protection against coronary heart disease. Circ. Genom. Precis. Med.12, e002376 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Dron, J. S. et al. Association of rare protein-truncating DNA variants in APOB or PCSK9 with low-density lipoprotein cholesterol level and risk of coronary heart disease. JAMA Cardiol.8, 258–267 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhao, Y. et al. Protein-truncating variants in BSN are associated with severe adult-onset obesity, type 2 diabetes and fatty liver disease. Nat. Genet.56, 579–584 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Selvaraj, M. S. et al. Whole genome sequence analysis of blood lipid levels in >66,000 individuals. Nat. Commun.13, 1–18 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Wang, Q. et al. Rare variant contribution to human disease in 281,104 UK Biobank exomes. Nature597, 527–532 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hawkes, G. et al. Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height. Nat. Commun.15, 1–11 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Flannick, J. et al. Exome sequencing of 20,791 cases of type 2 diabetes and 24,440 controls. Nature570, 71–76 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Jurgens, S. J. et al. Rare coding variant analysis for human diseases across biobanks and ancestries. Nat. Genet.56, 1811–1820 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Weiner, D. J. et al. Polygenic architecture of rare coding variation across 394,783 exomes. Nature614, 492 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Wainschtein, P. et al. Assessing the contribution of rare variants to complex trait heritability from whole-genome sequence data. Nat. Genet.54, 263–273 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Halldorsson, B. V. et al. The sequences of 150,119 genomes in the UK Biobank. Nature607, 732–740 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature599, 628–634 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Bick, A. G. et al. Genomic data in the All of Us Research Program. Nature627, 340–346 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Zhao, Y. et al. Population-scale gene-based analysis of whole-genome sequencing provides insights into metabolic health. Nat. Genet.57, 2436–2444 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Morris, A. P. & Zeggini, E. An evaluation of statistical approaches to rare variant analysis in genetic association studies. Genet. Epidemiol.34, 188–193 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Madsen, B. E. & Browning, S. R. A groupwise association test for rare mutations using a weighted sum statistic. PLoS Genet. 5, e1000384 (2009). [DOI] [PMC free article] [PubMed]
  • 44.Li, B. & Leal, S. M. Methods for detecting associations with rare variants for common diseases: application to analysis of sequence data. Am. J. Hum. Genet.83, 311–321 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Morgenthaler, S. & Thilly, W. G. A strategy to discover genes that carry multi-allelic or mono-allelic risk for common diseases: a cohort allelic sums test (CAST). Mutat. Res.615, 28–56 (2007). [DOI] [PubMed] [Google Scholar]
  • 46.Li, X. et al. Dynamic incorporation of multiple in silico functional annotations empowers rare variant association analysis of large whole-genome sequencing studies at scale. Nat. Genet.52, 969–983 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Wu, M. C. et al. Rare-variant association testing for sequencing data with the sequence kernel association test. Am. J. Hum. Genet.89, 82 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Li, Z. et al. A framework for detecting noncoding rare variant associations of large-scale whole-genome sequencing studies. Nat. Methods19, 1599 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Frankish, A. et al. GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Res.47, D766–D773 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Harrow, J. et al. GENCODE: the reference human genome annotation for The ENCODE Project. Genome Res.22, 1760–1774 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Zhou, H. et al. FAVOR: functional annotation of variants online resource and annotator for variation across the human genome. Nucleic Acids Res.51, D1300–D1311 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Lali, R. et al. Calibrated rare variant genetic risk scores for complex disease prediction using large exome sequence repositories. Nat. Commun.12, 1–15 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Chan, A. J. S. et al. Genome-wide rare variant score associates with morphological subtypes of autism spectrum disorder. Nat. Commun.13, 1–16 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Privé, F., Arbel, J. & Vilhjálmsson, B. J. LDpred2: better, faster, stronger. Bioinformatics36, 5424–5431 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Xu, L. et al. JointPRS: A data-adaptive framework for multi-population genetic risk prediction incorporating genetic correlation. Nat. Commun.16, 1–20 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Auton, A. et al. A global reference for human genetic variation. Nature526, 68–74 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Bulik-Sullivan, B. et al. LD score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat. Genet.47, 291–295 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Jiang, L. et al. The distribution and characteristics of LDL receptor mutations in China: a systematic review. Sci. Rep. 5, 1–11 (2015). [DOI] [PMC free article] [PubMed]
  • 59.Shen, H. et al. Familial defective apolipoprotein B-100 and increased low-density lipoprotein cholesterol and coronary artery calcification in the old order amish. Arch. Intern. Med.170, 1850–1855 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Pollin, T. I. et al. A null mutation in human APOC3 confers a favorable plasma lipid profile and apparent cardioprotection. Science322, 1702–1705 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.van Kippersluis, H. et al. Overcoming attenuation bias in regressions using polygenic indices. Nat. Commun.14, 1–16 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Lee, A. et al. BOADICEA: a comprehensive breast cancer risk prediction model incorporating genetic and nongenetic risk factors. Genet. Med.21, 1708–1718 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Darst, B. F. et al. Combined effect of a polygenic risk score and rare genetic variants on prostate cancer risk. Eur. Urol.80, 134–138 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Dornbos, P. et al. A combined polygenic score of 21,293 rare and 22 common variants improves diabetes diagnosis based on hemoglobin A1C levels. Nat. Genet.54, 1609–1614 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Lek, M. et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature536, 285–291 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.MacArthur, D. G. et al. A systematic survey of loss-of-function variants in human protein-coding genes. Science (1979)335, 823–828 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Gerasimavicius, L., Livesey, B. J. & Marsh, J. A. Loss-of-function, gain-of-function and dominant-negative mutations have profoundly different effects on protein structure. Nat. Commun.13, 1–15 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Fiziev, P. P. et al. Rare penetrant mutations confer severe risk of common diseases. Science380, eabo1131 (2023). [DOI] [PubMed]
  • 69.Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature581, 434–443 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Gaynor, S. M. et al. Yield of genetic association signals from genomes, exomes and imputation in the UK Biobank. Nat. Genet.56, 2345–2351 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Wainschtein, P. et al. Estimation and mapping of the missing heritability of human phenotypes. Nature649, 1219–1227 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Khlebus, E. et al. Multiple rare and common variants in APOB gene locus associated with oxidatively modified low-density lipoprotein levels. PLoS ONE14, e0217620 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Goldstein, J. L. & Brown, M. S. A century of cholesterol and coronaries: from plaques to genes to statins. Cell161, 161 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Jørgensen, A. B., Frikke-Schmidt, R., Nordestgaard, B. G. & Tybjærg-Hansen, A. Loss-of-function mutations in APOC3 and risk of ischemic vascular disease. N. Engl. J. Med.371, 32–41 (2014). [DOI] [PubMed] [Google Scholar]
  • 75.He, Z., Xu, B., Lee, S. & Ionita-Laza, I. Unified sequence-based association tests allowing for multiple functional annotations and meta-analysis of noncoding variation in metabochip data. Am. J. Hum. Genet.101, 340–352 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Li, X. et al. Powerful, scalable and resource-efficient meta-analysis of rare variant associations in large whole genome sequencing studies. Nat. Genet.55, 154–164 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.McCaw, Z. R. et al. An allelic-series rare-variant association test for candidate-gene discovery. Am. J. Hum. Genet.110, 1330–1342 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Clarke, B. et al. Integration of variant annotations using deep set networks boosts rare variant association testing. Nat. Genet.56, 2271–2280 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Zhao, Z. et al. Optimizing and benchmarking polygenic risk scores with GWAS summary statistics. Genome Biol.25, 1–28 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zhao, Z. et al. One score to rule them all: regularized ensemble polygenic risk prediction with GWAS summary statistics. bioRxiv10.1101/2024.11.27.625748 (2024).
  • 81.Jiang, W., Chen, L., Girgenti, M. J. & Zhao, H. Tuning parameters for polygenic risk score methods using GWAS summary statistics from training data. Nat. Commun.15, 1–15 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Zhao, Z. et al. PUMAS: fine-tuning polygenic risk scores with GWAS summary statistics. Genome Biol.22, 1–19 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Mbatchou, J. et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet.53, 1097–1103 (2021). [DOI] [PubMed] [Google Scholar]
  • 84.Friedman, J., Hastie, T. & Tibshirani, R. Regularization paths for generalized linear models via coordinate descent. J. Stat. Softw.33, 1 (2010). [PMC free article] [PubMed] [Google Scholar]
  • 85.Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience4, s13742-015 (2015). [DOI] [PMC free article] [PubMed]
  • 86.Szustakowski, J. D. et al. Advancing human genetics research and drug discovery through exome sequencing of the UK Biobank. Nat. Genet.53, 942–948 (2021). [DOI] [PubMed] [Google Scholar]
  • 87.Van Hout, C. V. et al. Exome sequencing and characterization of 49,960 individuals in the UK Biobank. Nature586, 749–756 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Li, X. et al. Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data. Cell Genom.5, 101009 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Cavalli-Sforza, L. L. The Human Genome Diversity Project: past, present and future. Nat. Rev. Genet.6, 333–340 (2005). [DOI] [PubMed] [Google Scholar]
  • 90.Williams, J. Data for Integrating Common and Rare Variants Improves Polygenic Risk Prediction Across Diverse Populations. Harvard Dataverse 10.7910/DVN/VCV7RZ (2026). [DOI] [PMC free article] [PubMed]
  • 91.Williams, J. jwilliams10/RareVariantPRS: v1.0.0. Zenodo 10.5281/zenodo.18788730 (2026).
  • 92.Williams, J. jwilliams10/RICE: v1.0.0. Zenodo 10.5281/zenodo.19186468 (2026).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

41467_2026_72185_MOESM1_ESM.pdf (14.7MB, pdf)

Supplementary Figures and Supplementary Notes

41467_2026_72185_MOESM2_ESM.pdf (21.7KB, pdf)

Description of Additional Supplementary Files

Supplementary Data (98.3KB, xlsx)
Reporting Summary (107.5KB, pdf)
Source Data (879.2KB, xlsx)

Data Availability Statement

UK Biobank phenotype data, WES data (Field #23156), imputed data (Field #22828), and WGS data (Fields #24304 and #24305) can be accessed through the UK Biobank research analysis platform [https://ukbiobank.dnanexus.com/landing]. All data used in this research are publicly available to registered researchers through the UKB data-access protocol and who are listed as collaborators on UKB-approved access applications. All of Us phenotype data, WES data, and WGS data can be accessed through the All of Us research workbench [https://workbench.researchallofus.org/login] (version 7.1). All data used in this research are publicly available to registered researchers with controlled tier access through the All of Us data-access protocol. Data generated in this study are largely available in the Supplementary Data or Source Data files. Large results generated in this study, including common-variant GWAS summary statistics, rare-variant association test summary statistics, and the RICE-CV and RICE-RV model weight files, are available through a Harvard Dataverse dataset90. Source data are provided with this paper.

Simulation and data analyses code are archived on Zenodo (v1.0.091). The corresponding GitHub repository is https://github.com/jwilliams10/RareVariantPRS. Software and tutorial to implement RICE are archived on Zenodo (v1.0.092) and available on GitHub [https://github.com/jwilliams10/RICE].


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES