Skip to main content
Nature Portfolio logoLink to Nature Portfolio
. 2026 May 8;58(6):1226–1236. doi: 10.1038/s41588-026-02597-9

Empirically determined baseline masking strategies and other considerations for gene-level burden tests

Trang Nguyen 1,2, Ryan Koesterer 1,3, Poeya Haydarlou 2, Peter Dornbos 1,3,4,5, Satoshi Yoshiji 1,6,7,8,9, Alex Llamas 1,3, Dongkeun Jang 1, Patrick Smadbeck 1, Annie Moriondo 1, Quy Hoang 1, Oliver Ruebenacker 1, Connie R Bezzina 2, Patrick Ellinor 10,11,12,13, Sean J Jurgens 2,10,11, Noël P Burtt 1, Jason Flannick 1,3,4,
PMCID: PMC13263148  PMID: 42104092

Abstract

Rare-variant association studies typically perform gene-level tests in which coding variants are filtered (or ‘masked’) and aggregated based on functional annotation and allele frequency. Through a systematic literature review, we cataloged 664 masks used across 234 studies and found that masking strategies (that is, sets of masks) rarely repeat across studies and are rarely justified. To quantify their impact on association results, we applied all previously employed strategies to 54 traits within 189,947 UK Biobank exomes. Here we find that the number of significant associations greatly depends on the masking strategy (ranging from 58 to 2,523 associations), which is a key reason for the modest overlap (<30%) of associations between separate published analyses of this dataset. We empirically determine masking strategies with high discovery power for low-frequency and rare variant gene-level associations across numerous datasets and traits, and we use these to explore the impact of other factors on burden test results. These findings offer a baseline strategy in burden tests to increase study power and replicability, addressing one source of inconsistency in previous studies.

Subject terms: Computational biology and bioinformatics, Genetic association study, Population genetics


Systematic comparisons of different masking approaches for rare variant association tests across 54 traits in UK Biobank highlight strategies for gene-level burden analyses that increase study power and replicability.

Main

Gene-level association tests, in which low-frequency and rare coding variants are grouped within a gene and tested for aggregate association with a trait, are central to whole-exome sequencing (WES) and whole-genome sequencing (WGS) association studies13. Various gene-level tests have been developed2 to estimate phenotypic associations for grouped variants. Some tests, like the sequence kernel association test (SKAT)4 or SKAT-O5, maintain power when variants have variable effect sizes5. Most studies, however, perform burden tests6,7 due to their simplicity and interpretability for drug discovery8,9 and functional follow-up3. Because burden tests assume that all included variants have the same impact on gene function, selecting included variants is especially important10.

Most burden tests include putative loss-of-function (pLoF) variants11,12, which have a high probability of disrupting gene function and can be bioinformatically predicted13. Theory suggests, however, that also including damaging missense alleles can triple effective sample size10. Early WES studies defined four ‘masks’ for including a missense variant in a burden test according to which of five bioinformatic algorithms predict it as damaging14,15 and its minor allele frequency (MAF), a variable shown to considerably impact burden test power10. Over time, many more bioinformatic prediction algorithms have been developed16,17, substantially increasing the number of potential masks18. Moreover, because most studies test multiple masks, which we refer to as a ‘masking strategy’, the number of potential study designs is far greater than the number of potential masks. Although in principle these options present opportunities for methodological innovation or study-specific masking criteria, in practice many WES studies cite previous publications1923 to justify their masking strategy.

In this study, we investigated how masking strategies used in previous studies varied in their definitions and discovery power, constructed masking strategies with empirically consistent and high discovery power across a variety of traits and datasets, and (c) contextualized these strategies alongside other variables that could affect gene-level association results. We addressed these questions by using WES data and conducting gene-level burden tests for 65 traits in the UK Biobank (~190 K and ~440 K samples)24, 56 traits in the All of Us Research Program (~415 K samples)25, and 12 traits in the AMP T2D GENES consortium (~40 K samples)26. Our results demonstrate substantial variability across existing masking strategies that impedes interpretation and comparison of studies. We suggest, as alternatives, two baseline masking strategies that can serve as a benchmark against which new masks or additional parameters in burden testing can be evaluated and that provide a sensible default for future studies.

Results

Previous masking strategies used in gene-level association studies

We surveyed mask usage from 2012 to 2024 by searching PubMed with keywords for coding variant association analyses and filtering results for citation rates and content (Fig. 1a, Supplementary Table 1a and Methods), identifying 234 studies describing gene-level association tests. These studies collectively used 664 masks (~2.84 per study; Supplementary Table 2a), 601 of which were clearly defined, employing 45 bioinformatic annotations and 25 allele frequency/count filters (Fig. 1a and Supplementary Table 2b). We mapped all mask definitions to a common software pipeline using the Variant Effect Predictor (VEP)27, dbNSFP28 and UK Biobank frequency filters, slightly altering definitions as needed (Methods). We removed 141 masks that we either could not implement with this pipeline or that contained non-coding variants or only synonymous variants (Methods). These filters resulted in 460 masks (Fig. 1a and Supplementary Table 2c,d), applied as part of 154 masking strategies within 169 publications (Methods).

Fig. 1. Literature review of masks and masking strategies.

Fig. 1

a, Summary of the literature review. b, Number of publications that employed the four MAF groups over the years. c, Number of publications that employed six groups of bioinformatic annotations over the years. d, Usage frequency of masks and masking strategies after mask harmonization in VEP. e, Overlap of Bonferroni-significant associations published by three high-profile studies of the UK Biobank for 46 available traits. f, Overlap of Bonferroni-significant associations produced by the masking strategies of three high-profile studies of the UK Biobank for 54 analyzed traits using 189,947 UK Biobank samples under the same association testing procedure. MAF, minor allele frequency; misIndels, missense + indels; pLoFmis, predicted loss-of-function + missense; pLoFdamMis, predicted loss-of-function + damaging missense; damMis, damaging missense; pLoF, predicted loss-of-function.

The 460 masks grouped broadly into 24 categories, defined by six bioinformatic annotation types and four MAF groups (Supplementary Tables 2a and 3). Masks with low-frequency filters (MAF < 1%) appeared the most (92 studies, 183 masks), with rare (MAF < 0.1%) and ultra-rare (MAF < 0.01%) mask usage increasing since 2019 (Fig. 1b), consistent with increasing sample sizes making it possible to detect significant associations with rarer variants. In terms of bioinformatic annotations, pLoF variants (83 studies, 116 masks) and pLoF or damaging missense variants (pLoFdamMis; 71 studies, 150 masks) have remained the most popular types (Fig. 1c). Collectively, 22 (13% of) studies excluded pLoF variants to only focus on medium-impact variants2931, 14 (8.3% of) studies only included pLoF variants and the majority (78.7%) of studies included both pLoF and non-pLoF variants (Supplementary Table 2a).

Strikingly, 78.2% of masks and 92.2% of masking strategies were used in only one publication (Fig. 1d). This, if anything, understates the lack of consistency in mask usage, as without mask harmonization the non-repeated fractions were even larger (91.7% of masks and 98.5% of masking strategies). Furthermore, most studies did not justify their masking choice and instead referenced previous publications (Methods), suggesting that this lack of consistency is not motivated by principled study-specific selection criteria. As an example of the impact of this inconsistency on reported associations, three recent large-scale studies22,23,32 of the UK Biobank WES dataset each used unique masking strategies (Supplementary Table 2a and Methods) and reported largely non-overlapping associations: only 28.2% of reported associations were shared by all three studies, and only 50% were shared by at least two (Fig. 1e). The overlap of associations increased only slightly when we compared the three masking strategies under an identical association testing procedure for identical UK Biobank samples (Methods): 35.6% of the associations were identified by all three strategies, and 51.4% were identified by at least two (Fig. 1f). These results suggested that inconsistency across masking strategies is a potentially substantial source of inconsistency in reported burden test results.

Variability in associations detected by previously employed masking strategies

To examine more fully the extent to which differences in mask definition lead to differences in association results, we conducted burden tests for each of the 298 distinct masks (from the 460 masks reviewed above) across 55 quantitative traits and 17,605,686 variants within 189,947 (~190 K) UK Biobank whole-exome sequences (Supplementary Table 4 and Methods). We excluded analyses with high genomic inflation and low-count genes and masks, after which 271 masks (146 unique masking strategies within 163 studies) and 54 traits remained for analysis (Supplementary Tables 5a,b and 6 and Methods). As a negative control, burden tests of synonymous variants with MAF < 0.1% yielded only 7 exome-wide significant (P < 2.5 × 10−6) associations across 54 traits. By contrast, the other masks collectively produced 6,563 significant associations, ranging per mask from 3 (missense variants predicted damaging by PolyPhen2 with MAF < 0.001%; ref. 19) to 2,706 (pLoF variants/missense variants predicted damaging by SIFT33 or PolyPhen2 (refs. 34,35) Supplementary Fig. 1a and Supplementary Table 7a). Wide variation across masks was observed for nearly every trait (Supplementary Fig. 1b), within each of the 24 mask categories, as well as between the categories (F-statistic = 22.27, P = 1.23 × 10−47 for all associations; Fig. 2a).

Fig. 2. Number of significant associations for 54 traits in the ~190 K UK Biobank analysis.

Fig. 2

a, Average number of total, low-frequency and rare variant significant associations in 24 mask categories. Error bars represent standard deviations. b, Number of total, low-frequency and rare variant Bonferroni-significant associations detected by the masking strategy used in each of the 163 publications. Dashed lines represent the number of significant associations produced by the ‘brute-force’ 271-mask strategy. misIndels, missense + indels; pLoFmis, predicted loss-of-function + missense; pLoFdamMis, predicted loss-of-function + damaging missense; damMis, damaging missense; pLoF, predicted loss-of-function; lowFreq, low-frequency.

Notably, many of these associations resulted from masks that included common variants (Supplementary Table 7a), which due to linkage disequilibrium (LD) can tag variants outside of the mask (Supplementary Fig. 1c, Supplementary Table 7b and Methods). Because most WES studies instead aim to identify low-frequency and rare variant associations, we conducted two further analyses to filter significant associations to those detectable via only low-frequency variants (MAF < 1%; low-frequency associations) or rare variants (MAF < 0.1%; rare associations; Methods). The number of significant low-frequency and rare variant associations still varied widely across masks (ranging from 3 to 440 for low frequency and 3 to 289 for rare; Supplementary Fig. 1a and Supplementary Table 7a), within and between mask categories (Fig. 2a), and for nearly every trait (Supplementary Fig. 2a,b). We note that, as previously demonstrated36, frequency filters reduced but did not eliminate LD between these associations and common variants (Supplementary Fig. 2c,d and Supplementary Table 7b); conditional analysis is still required to demonstrate independence between rare variant gene-level and common-variant associations.

Across the 146 masking strategies (in 163 publications), the number of Bonferroni-corrected significant associations varied nearly as widely as it did across the 271 masks (Fig. 2b and Supplementary Tables 2a and 8), ranging from 58 (a two-mask strategy37 aggregating pLoF and missense singletons) to 2,523 (a two-mask strategy35 aggregating pLoF and damaging missense variants at two MAF thresholds). Variation was lower but still substantial for significant low-frequency (58 to 563) and rare (58 to 352) variant associations (Fig. 2b and Supplementary Tables 2a and 8). There was no clear trend in the number of significant associations versus the year in which the masking strategy was defined (Fig. 2b and Supplementary Table 8).

Potential baseline masking strategies

The large number of previously used masking strategies, lack of study-specific rationale for them and high dependence of reported associations on masking strategy motivated us to explore ‘baseline’ masking strategies that could be used as sensible starting points in future burden test analyses of low-frequency and rare variants (Fig. 3a–j). One potential baseline option is to mimic recent and high-profile studies, such as the large-scale UK Biobank WES analyses22,23,32 (Fig. 3a). These ‘high-profile’ masking strategies detected more associations than most previously employed strategies, ranking in the 50th to 83rd percentiles for low-frequency variant associations (0.94-fold to 1.34-fold more than the ‘average’ strategy, defined as a proxy strategy that detected the average number of significant associations across all previously employed strategies) and 66th to 88th percentiles for rare variant associations (1.17-fold to 1.43-fold more than the average strategy; Supplementary Table 8).

Fig. 3. Overview of the potential baseline masking strategies.

Fig. 3

a, Masking strategies from three high-profile UK Biobank studies. b, Brute-force masking strategies consisting of 271 previously employed masks. c, Masking strategies derived from the clustering of the 271 masks based on the variants they included. d, Masking strategies derived from the clustering of the 271 masks based on the minor allele frequencies of the variants they included. e, Masking strategies produced by the greedy algorithm when it was applied to the 271 masks. f, Masking strategies produced by the greedy algorithm when it was applied to the expanded set of 424 masks (271 previously employed masks and 153 new masks). g, Final proposed baseline masking strategies. h, Example application of the proposed baseline masking strategies. i, Robust performance of the baseline masking strategies in various contexts. j, Other considerations for gene-level burden tests.

On the other extreme, a naïve (but comprehensive) masking strategy would include every mask ever defined (Fig. 3b). This ‘brute-force’ strategy, even after Bonferroni correction for 271 masks, produced 1.9-fold more low-frequency and 1.8-fold more rare variant significant associations than the average strategy (Fig. 4a, Supplementary Table 8 and Methods). In fact, it detected at least 1.07-fold more low-frequency variant significant associations and 1.03-fold more rare variant significant associations than any previous strategy (Supplementary Table 8 and Methods).

Fig. 4. Number of Bonferroni-significant associations produced by potential baseline masking strategies in the ~190 K UK Biobank analysis.

Fig. 4

a, Across all 54 traits. Orange dashed line represents low-frequency variant associations, and blue dashed line represents rare variant associations from the average strategy. b, Low-frequency variant associations for each trait. c, Rare variant associations for each trait. See Supplementary Table 4 for phenotype descriptions.

A logical improvement on the brute-force 271-mask strategy would be to remove redundant masks from it. We performed mask principal-component analysis (PCA)38 (based on the variants they included) and clustering, producing 10-mask clusters (Supplementary Fig. 3a and Methods). Two ‘clustered’ 10-mask strategies with one mask per cluster (Fig. 3c, Supplementary Table 9a–c and Methods) produced more low-frequency and rare variant Bonferroni-significant associations than did the brute-force strategy (Fig. 4a and Supplementary Table 10). Although these clusters were heterogeneous (silhouette score = 0.51; Supplementary Fig. 3b), two ‘subclustered’ 37-mask strategies determined by a second round of PCA (based on variant MAFs) and clustering (Fig. 3d and Methods) produced slightly fewer Bonferroni-significant associations than did the 10-mask strategies, despite the greater homogeneity of the subclusters (mean of silhouette scores = 0.79, standard deviation = 0.12; Supplementary Table 9d), due to the cost of increased Bonferroni correction outweighing the benefit of new discoveries from more masks.

As a more direct alternative to maximize significant associations produced by a masking strategy, we used a greedy algorithm to add masks to the strategy based on additional significant associations identified in our UK Biobank 54 trait analysis (Fig. 3e and Methods). We validated the robustness of this approach via a cross-validation analysis (Supplementary Fig. 4, Supplementary Table 11 and Methods). This approach yielded two ‘greedy’ 26-mask strategies (Supplementary Fig. 5a,b and Supplementary Tables 10 and 12) that produced 97 more low-frequency variant significant associations (15.3% increase) and 70 more rare variant significant associations (19.1% increase) than did the clustered 10-mask strategies (Fig. 4a and Supplementary Table 10).

Finally, we investigated potential gains from including masks not employed in the literature. We derived two types of new masks: 117 masks from applying stricter MAF thresholds to the 271 previously employed masks, and 36 masks combining pLoF variants with missense variants predicted to be damaging by different fractions of bioinformatic algorithms (Supplementary Table 6 and Methods). This expanded mask library did in fact detect 969 additional associations (ignoring Bonferroni correction) not found by any of the 271 previously used masks. Repeating the greedy algorithm including the expanded library, we obtained two ‘expanded greedy’ 22-mask and 35-mask strategies (Fig. 3f and Supplementary Table 13) that detected the most low-frequency and rare variant Bonferroni-significant associations of any strategy we explored (Fig. 4a and Supplementary Table 10) with consistent performance across traits (Supplementary Fig. 6a,b).

Proposed baseline masking strategies

It is possible that many researchers may be hesitant to test 22 or 35 masks in practice. As an alternative, our analyses indicate much smaller 8-mask greedy strategies captured approximately 95% as many associations as did the expanded greedy strategies (Fig. 4a and Supplementary Table 10). The smaller 8-mask strategies detected twice as many associations as the average strategy and showed consistent proportional increase across traits (Fig. 4b,c).

We therefore propose these two 8-mask strategies as baseline strategies for future gene-level burden analyses (Table 1). Although they represent a 2.9-fold increase in the number of masks relative to a typical previous study, compared to testing one mask (using the REGENIE software39), we found almost no difference in analyst setup time, almost no increase in memory usage (Fig. 5a,b) and a modest (1.2-fold and 2.2-fold for rare and low-frequency variant masks, respectively) increase in CPU time (Fig. 5c,d).

Table 1.

Proposed baseline masking strategies

Baseline strategy Mask Mask summary Number of significant associations Number of significant associations added to the masking strategy Number of significant associations cumulated in the masking strategy
Low frequency new_damaging_og25.0_01 Low-frequency (MAF < 1%) pLoF variants or missense variants predicted to be damaging by 25% of the algorithms 349 349 349
32853339_m1 Low-frequency (MAF < 1%) high-impact or (likely) pathogenic variants in ClinVar 245 99 448
29177435_m1 Low-frequency (MAF < 1%) exonic variants 258 84 532
29378355_m1.0_01 Low-frequency (MAF < 1%) pLoF or missense variants with MAF > 0.5% 148 58 590
36327219_m3 Low-frequency (MAF < 1%) high-impact, indels or damaging missense predicted by 5 algorithms 324 49 639
new_damaging_og50.0_01 Low-frequency (MAF < 1%) pLoF variants or missense variants predicted to be damaging by 50% of the algorithms 309 31 670
24507775_m6.0_01 Low-frequency (MAF < 1%) pLoF variants or damaging missense predicted by PolyPhen2 304 29 699
32141622_m7.0_01 Low-frequency (MAF < 1%) pLoF or splice sites 140 19 718
Rare 31383942_m4 Rare (MAF < 0.1%) pLoF or damaging missense predicted by REVEL 232 232 232
37348876_m8 Rare (MAF < 0.1%) pLoF or damaging variants predicted by CADD 225 64 296
31383942_m10 Rare (MAF < 0.1%) pLoF or (likely) pathogenic variants in ClinVar 207 53 349
31118516_m5.0_001 Rare (MAF < 0.1%) pLoF or damaging missense predicted by 5 algorithms 210 24 373
36411364_m4.0_001 Rare (MAF < 0.1%) pLoF or damaging missense predicted by REVEL 220 15 388
34183866_m1 Ultra-rare (MAF < 0.01%) pLoF, indels or damaging missense predicted by PolyPhen2, MAF < 0.01% 162 13 401
30269813_m4 Rare (MAF < 0.1%) pLoF, indels or damaging missense predicted by PolyPhen2 223 11 412
34216101_m3.0_001 Rare (MAF < 0.1%) damaging missense predicted by PolyPhen2 47 10 422

This table shows the summary of the proposed baseline masking strategies and the number of significant associations for 54 continuous traits in the ~190 K UK Biobank analysis. MAF, minor allele frequency; pLoF, predicted loss-of-function. See Supplementary Table 6 for detailed mask definitions.

Fig. 5. Computational cost of conducting burden tests for alanine aminotransferase when using the baseline masks together versus separately in REGENIE.

Fig. 5

a, Maximum memory usage for low-frequency variant masks. b, Maximum memory usage for rare variant masks. c, Total compute time for low-frequency variant masks. d, Total compute time for rare variant masks. MB, megabytes. See Table 1 for mask descriptions.

Example application of the baseline masking strategies

As a practical application, we used our proposed baseline strategies to conduct exome-wide analyses of the 54 traits in the full 439,829 (~440 K) sample UK Biobank24 WES dataset (Fig. 3h and Methods). Across 46 traits shared with three previous large-scale UK Biobank WES studies22,23,32, we identified 114 significant associations not previously identified (Supplementary Table 14a), most (95.6%) with P values exceeding exome-wide significance by at least an order of magnitude. Gene set enrichment analyses across 4,203 gene sets from the Mouse Genomic Informatics database40 demonstrated the biological coherence and plausibility of these associations in aggregate (Supplementary Table 14b,c and Methods).

The 114 novel associations included genes within three categories: known ‘effector genes’41 for the associated traits, new effector gene candidates and genes with no prior genetic links to the associated trait. For example, we observed a significant burden association between eosinophil count association and rare variants in PRG2 (P = 1.1 × 10−9), encoding a major component of the crystalline core of the eosinophil granule (https://www.ncbi.nlm.nih.gov/gene/5553), consistent with GWAS associations for eosinophil count42 nearby PRG2 (rs548854, P = 8.2 × 10−20; rs639509, P = 5.1 × 10−25). Similarly, we observed an association between hemoglobin A1C (HbA1C) and rare variants in PTPN11 (beta = −0.57, P = 1.5 × 10−8), the coding gene for SHP2 protein (https://www.ncbi.nlm.nih.gov/gene/5781), consistent with HbA1C GWAS associations43 (rs11066309, beta = −0.03, P = 3.9 × 10−26; rs1029850317, beta = −0.02, P = 2.7 × 10−18). A recent review discussed the potential protective or promoting roles of SHP2 on insulin resistance44, whereas another study identified the overexpression of PTPN11 in patients with type 2 diabetes complicated with colorectal cancer45.

One example of a novel association outside of a GWAS locus is between platelet count and rare variants in PTPN6 (P = 1.9 × 10−9), a biologically plausible association given the role of PTPN6 in signaling pathways of hematopoietic cells (https://www.ncbi.nlm.nih.gov/gene/5777), platelet signal transduction46 and inflammatory processes47, as well as animal model evidence (in mice, PTPN6 inhibited platelet apoptosis and necroptosis48, and in zebrafish, ptpn6 knockdown impaired the innate immune system49). Another example is an association between forced vital capacity and rare variants in ASXL1 (P = 2.4 × 10−13), studied extensively for its role in myeloid diseases50 and Bohring-Opitz syndrome51. Supporting this association, a study found that Asxl1 ablation in mice caused defective lung maturation52, and abnormal alveoli formation was associated with ASXL1 variants in an autopsy of a newborn who died of Bohring-Opitz syndrome53.

Robust performance of the baseline masking strategies in various contexts

To evaluate the performance of the baseline strategies in datasets other than the UK Biobank, we first analyzed 45 continuous traits using 410,400 samples from the All of Us (AoU) research program25 (Fig. 3i, Supplementary Table 15a and Methods). Compared to the high-profile masking strategies employed in the three large-scale UK Biobank WES studies22,23,32, the baseline strategies produced 74 to 126 more low-frequency variant associations (1.5-fold to 2.16-fold increases; Extended Data Fig. 1a) and 40 to 46 more rare variant associations (1.31-fold to 1.37-fold increases; Extended Data Fig. 1b). An additional analysis of 12 continuous traits in 40,054 samples from the AMP T2D GENES consortium26 (Supplementary Table 15b and Methods) showed mostly similar results, although one high-profile strategy identified 16 more rare variant associations (17 gained, 1 lost) than did the rare variant baseline strategy (Extended Data Fig. 1c,d); the 17 gained associations were detected by a rare variant mask containing all functional coding variants32, were exclusively between immunoglobulin genes and fasting glucose and/or fasting insulin (Supplementary Table 15c) and had no supporting evidence from GWAS54 or experimental studies.

Extended Data Fig. 1. Number of Bonferroni-significant associations produced by the baseline masking strategies versus three ‘high-profile’ strategies in various contexts.

Extended Data Fig. 1

a,b, Low-frequency variant associations (a) and rare variant associations (b) across 45 continuous traits in the AoU Research Program. c,d, Low-frequency variant associations (c) and rare variant associations (d) across 12 continuous traits in the AMP T2D GENES cohort. e,f, Low-frequency variant associations (e) and rare variant associations (f) across 11 binary traits in the UK Biobank. g,h, Low-frequency variant associations (g) and rare variant associations (h) across 11 binary traits in the AoU Research Program. Note: PMID: 36778668 did not employ any rare variant masks.

The baseline strategies also detected more associations for binary traits. Across 11 diseases of varying case frequencies (0.08%–32%) in the UK Biobank (Supplementary Table 16a and Methods), the baseline strategies produced 1.21-fold more low-frequency variant associations (Extended Data Fig. 1e) and 1.32-fold to 1.5-fold more rare variant associations (Extended Data Fig. 1f) than did the high-profile strategies. Results were similar for analyses of the same 11 diseases in the AoU25 dataset (Supplementary Table 16b, Extended Data Fig. 1g,h and Methods).

We next evaluated the discovery power of the baseline strategies across different genetic architectures, typically represented by degrees of polygenicity55. Since polygenicity defined by the number of independent common SNPs55 did not reliably reflect the number of significant genes detected by testing the burden of low-frequency and rare variants (Extended Data Fig. 2a,b and Methods), we grouped the 54 continuous traits and 11 binary traits into four ‘polygenicity’ groups based on the number of significantly associated genes within the UK Biobank. Across these groups, the baseline strategies consistently produced more significant low-frequency and rare variant associations than did the high-profile strategies (Extended Data Fig. 2c,d).

Extended Data Fig. 2. The discovery power of the baseline and high-profile masking strategies for traits across the genetic architecture spectrum in the ~440 K UK Biobank analysis.

Extended Data Fig. 2

a, The estimated common-variant polygenicity for 12 traits published in O’Connor et al.55 (orange line), and number of significant genes detected by any low-frequency variant mask in the four masking strategies for the same 12 traits (blue bars). b, The estimated common-variant polygenicity for 12 traits published in O’Connor et al.55 (orange line), and number of significant genes detected by any rare variant mask in the four masking strategies for the same 12 traits (blue bars). c, Proportion of significant low-frequency variant associations detected by the baseline masking strategies compared to the high-profile masking strategies in each group of traits. d, Proportion of significant rare variant associations detected by the baseline masking strategies compared to the high-profile masking strategies in each group of traits. Error bars represent the 95% confidence interval (CI). The sample sizes represent the number of traits in each group. Note: PMID: 36778668 did not employ any rare variant mask.

We did not have adequate power to evaluate the performance of the baseline strategies across different ancestries. However, we found that, in general, across ancestries mask composition was highly concordant (Supplementary Fig. 7a–c and Methods) and the number of variants per gene was highly correlated (Supplementary Fig. 7d and Methods). These findings complement the high discovery power of the baseline masks in AMP T2D GENES (Extended Data Fig. 1c,d) and AoU (Extended Data Fig. 1a,b,g,h), which are majority non-European, to suggest that properties of the baseline strategies is likely to be transferrable to non-European ancestries22,56,57.

Other considerations for gene-level burden tests

Although masking strategies are a major consideration for burden tests, other variables are also important. We used the baseline masking strategies as a foundation to examine the impact of five such variables (Fig. 3j): (a) burden test software, (b) level of variant aggregation, (c) transcripts included in analysis, (d) emphasis on mask simplicity and interpretability and (e) trade-offs between using low-frequency versus rare masks.

First, we compared burden test results for 5 continuous traits and 5 binary traits within the UK Biobank using both SAIGE-GENE + (ref. 58) and REGENIE39 (Methods). Burden test z-scores had a near-perfect correlation (Pearson’s R = 0.99, P < 1 × 10−324), consistent with previous observations58, leading to a substantial overlap of significant associations per mask (Supplementary Table 17a). REGENIE identified 50 more associations for continuous traits, whereas SAIGE-GENE+ identified 22 more associations for binary traits. Of the additional SAIGE-GENE+ associations, 6 failed to converge in REGENIE and 14 were associations with cardiomyopathy (case frequency = 0.7%). These results are consistent with SAIGE-GENE+’s intended design to optimize rare variant tests for unbalanced case-control phenotypes58.

Second, we reconducted burden tests aggregating only variants in conserved protein domains (Methods). Associations substantially overlapped those obtained from gene-level aggregation (mean Jaccard similarity score = 0.86, standard deviation = 0.08; Supplementary Table 17b), although with slightly fewer (91% as many) achieving significance, likely due to lower variant counts leading to lower effective sample size.

Third, we conducted burden tests with four different approaches for transcript-level variant aggregation: testing only the canonical transcript20, testing the most experimentally supported transcript18, annotating variants according to the ‘most severe’ transcript consequence22 and testing each transcript individually (Methods). Significant associations overlapped substantially (Extended Data Fig. 3a and Supplementary Table 17c), with the ‘most severe’ approach producing the most associations (4,654; Extended Data Fig. 3b). Testing all transcripts individually (and correcting for multiple tests; Methods) identified 98% as many associations as the ‘most severe’ approach, although this approach was computationally expensive (76,288 transcripts versus 19,474 genes). On the other hand, testing the canonical transcripts only produced 80% as many associations as the ‘most severe’ approach, missing out on many genes whose biological impacts were most pronounced in a non-canonical isoform. Finally, testing the most experimentally supported transcripts (the approach used throughout our study) identified 96% as many associations as the ‘most severe’ approach while producing results at the transcript level.

Extended Data Fig. 3. Transcript analysis results for 54 continuous traits using ~440 K UK Biobank samples.

Extended Data Fig. 3

a, Average Jaccard similarity scores in terms of the significant associations across 8 baseline rare variant masks for each pair of transcript choices. b, Overlap of significant associations produced across 8 baseline rare variant masks when we annotated and collapsed variants according to four transcript choices: TSL (most experimentally supported transcripts), canonical transcripts, each of all available transcripts, and using variants’ most severe consequence annotations across all transcripts.

Fourth, we compared the baseline strategies to others that prioritized mask interpretability and simplicity. First, reapplying the greedy algorithm to 53 masks we manually annotated as ‘interpretable’ (Methods) produced strategies in which each mask was simply defined, but it was unclear whether the resultant strategies were more interpretable than the baseline strategies (Supplementary Table 18). Two even more ‘simple’ low-frequency and rare variant masking strategies (each containing only four ‘nested’ masks; Supplementary Table 18) detected substantially fewer (86.7% as many low-frequency and 86.1% as many rare variant) associations than did the baseline strategies (Extended Data Fig. 4a,b). Second, as loss-of-function variants are the most interpretable variant class, we re-applied the greedy algorithm while forcing it to include pLoF masks (Methods). The analysis yielded two new 8-mask ‘pLoF-inclusive’ strategies that consisted of largely the same masks as the baseline strategies (Supplementary Table 18 and Extended Data Fig. 4a,b). These strategies may be worth employing if detecting all pLoF associations is of high importance.

Extended Data Fig. 4. Number of Bonferroni-significant associations produced by the ‘simple’ masking strategies, ‘pLoF-inclusive’ masking strategies, and baseline masking strategies across 54 continuous traits using ~440 K UK Biobank samples.

Extended Data Fig. 4

a, Low-frequency variant associations. b, Rare variant associations.

Finally, we found that the low-frequency variant masks produced substantially (1.83-fold to 1.86-fold) more Bonferroni-significant associations than did the rare variant masks in all three groups of strategies (Extended Data Fig. 5a–c), albeit with some of these likely due to LD with common variants (Supplementary Fig. 2c,d and Supplementary Table 7b). Employing both low-frequency and rare variant masks (and Bonferroni correcting for all masks) recovered almost all associations detected by the low-frequency or rare variant masks alone (Extended Data Fig. 5a–c), balancing between discovery power and independence from common variants.

Extended Data Fig. 5. Number of Bonferroni-significant associations across 54 traits using ~440 K UK Biobank samples produced by the low-frequency variant masks only, rare variant masks only, and both low-frequency and rare variant masks.

Extended Data Fig. 5

a, Baseline masking strategies. b, Simple masking strategies. c, pLoF-inclusive masking strategies.

Discussion

Although rare variant association studies have produced many important discoveries22,23,32, the lack of standards for variant filtering and grouping limits the consistency and reproducibility of gene-level associations. This inconsistency is demonstrated by the large number of masks and masking strategies employed by previous studies that have never been used again (Fig. 1d). Such a situation could be desirable if studies focused on study-specific mask design or burden test interpretation, but our literature review suggests that instead mask usage is typically an afterthought (Methods), despite (as our analyses show) a high dependence of mask choice on reported gene-level associations and discovery power. When mask choice is not the focus of a study, our proposed baseline masking strategies represent a reasonable default choice to achieve robust discovery power across different datasets and traits (Fig. 4a, Extended Data Figs. 1 and 2 and Supplementary Fig. 7). Our follow-up analyses further clarify the impact that other burden testing choices have on gene-level association results (Extended Data Figs. 35).

Our study has several limitations. First, our PubMed search strategy did not capture all masks ever employed; this biases our reported variability of mask usage downwards and may understate the problem of mask inconsistency. Second, we assumed that the new associations detected by the baseline masking strategies were mostly true positives rather than statistical noise; this claim is supported by their overlap with GWAS signals and our gene set enrichment analyses, but independent replication would increase confidence in their quality. Third, although we found that it is rare for multiple masks to produce significant but opposite effects for the same gene-trait association (6% of associations in an average strategy, and none in the rare variant baseline strategy), we did not explore in-depth the association patterns across masks within each strategy; careful examination into why one or more masks detected a signal would benefit the interpretation of its biological significance. Fourth, our results rely on bioinformatic predictions of variant functional impact and would needed to be updated if, in the future, substantially better methods become available for predicting or measuring variant impact. Finally, our proposed baseline strategies are empirically determined, which limits the insights they provide into disease genetic architecture and study interpretability.

Our study nonetheless suggests several clear recommendations for gene-level association studies. First, researchers should be cognizant of the impact that masks have on their results and transparently report and justify their masking strategy. Second, both REGENIE and SAIGE-GENE+ produce highly concordant results, with REGENIE detecting slightly more associations for continuous traits and SAIGE-GENE+ slightly more associations for binary traits. Third, we recommend annotating and aggregating variants according to experimentally supported transcripts to balance specificity, efficiency and discovery power. Fourth, most researchers should likely avoid using masks that include common variants, as these are largely redundant with GWAS and swamp true rare variant signals. Finally, we recommend our baseline masking strategies as a benchmark for newly proposed masking strategies and as a sensible default for researchers who are not interested in mask tuning and/or who prioritize other research areas (for example, association patterns across masks, GWAS effector gene candidate evaluations or biological interpretations of novel findings). These are not intended to be rigid prescriptions, but their adoption when appropriate will increase study replicability and transparency.

Methods

Our research complies with all relevant ethical regulations. Details are provided below in the respective section of each cohort (UK Biobank, AMP T2D GENES, All of Us).

Literature review of variant masks and masking strategies

We searched PubMed (2024-02-28) for scientific articles published between 2012 and 2024 with two sets of keywords: (1) rare variants AND (burden OR skat) and (2) ‘whole-exome sequencing’ AND (gene-level OR gene-burden OR ‘coding variants’), retrieving 2094 articles (Supplementary Table 1a). We included papers published between 2012 and 2022 that were cited at least 2.5 times a year (609 publications) and papers published between 2023 and 2024 that were cited at least once (119 publications). We reviewed each paper and manually filtered them to 215 relevant publications based on title and content (Supplementary Note). We then added 19 other studies that we identified through additional ad hoc searches to a total of 234 publications (Supplementary Table 1a).

Data and samples used

For the ~190 K sample analysis, we obtained genotypes for 200,631 consented samples from the UK Biobank24 (application ID: 41189) October 2020 version of the ‘Population level exome OQFE variants, PLINK format – interim 200k release (data field: 23155)’. For the ~440 K sample analysis, we used the UK Biobank Research Analysis Platform (RAP) to obtain genotypes for 469,818 consented samples from the UK Biobank ‘Population level exome OQFE variants, * format – final release’ (* = BGEN/PLINK/pVCF, data fields: 23159/23158/231257). We pre-filtered the variants of the PLINK (data field: 23158) version by removing the low-quality variants identified in the UK Biobank provided helper file ‘ukb23158_500k_OQFE.90pct10dp_qc_variants.txt’.

We selected 28 continuous phenotypes with different genetic architectures previously analyzed for rare variant burden heritability19. We then added 27 more phenotypes in various phenotype groups (eg, anthropometric, cardiometabolic, hematological) to increase the phenotype pool (Supplementary Table 4). Because UK Biobank participants had multiple visits, we took the mean of all measurements and inverse-normalized the mean values to ensure normality before association testing.

Data quality control

We followed standard variant and sample quality control (QC) procedures59 (Supplementary Note). Briefly, we first removed variants with call rate < 0.9, heterozygosity = 1, or Hardy–Weinberg equilibrium P < 1 × 10−6. A restricted set of variants was used as markers for sample QC (SNPs; call rate ≥ 0.98, MAF ≥ 0.01, and Hardy–Weinberg equilibrium P ≥ 1 × 10−6) and pruned using PLINK60,61 (https://www.cog-genomics.org/plink) to independent SNPs (-indep-pairwise 1000 kb 1 0.2).

For sample ancestry inference, we merged the genotypes with reference genotypes on a set of 5,835 known ancestry informative SNPs, derived their principal components (PCs) using FlashPCA262, and applied the K-nearest neighbor (KNN) method63 using the ‘knn’ function from the ‘class’ package in R64 (Supplementary Fig. 8, Supplementary Table 19a and Supplementary Note).

For sample QC, we first assessed duplicates and cryptic relatedness using the KING v2.2.8 relationship inference software (https://www.kingrelatedness.com) in the ~190 K analysis and PLINK261 v2.3a (https://www.cog-genomics.org/plink/2.0) in the ~440 K analysis, excluding one sample from each of 27 pairs identified as duplicates (kinship > 0.4). We also checked for genotypic/clinical data agreement for sex using the ‘impute_sex’ method in Hail v0.2.95 (https://github.com/hail-is/hail) with default settings, and found no disagreements between genotypic and clinical sex. We then conducted sample outlier detection according to 10 sample-by-variant metrics (Supplementary Table 19b), calculated using Hail and converted to principal-component adjusted residuals of the metrics (PCARM) using the restricted sample QC marker variants. We clustered the samples into Gaussian distributed subsets with respect to each PCARM using the software Klustakwik (https://klusta.readthedocs.io/en/latest/), removing samples that did not fit into any Gaussian distributed set of samples. We applied an analogous clustering and removal procedure to principal components of the PCARMs to identify samples that were outliers across multiple metrics (Supplementary Fig. 9).

In summary, in the ~190 K analysis, we removed 181 samples as outliers, and in the ~440 K analysis, we removed 188 samples as outliers (Supplementary Table 19c). Only samples of inferred European ancestry with non-missing phenotypes were used in each association analysis (Supplementary Table 4).

Variant annotation

We annotated variants with the VEP27 v110.1, LOFTEE13 plug-in v1.0.3 (~190 K analysis) and v1.0.4 (~440 K analysis) and dbNSFP28 plug-in v4.4a. For all burden test analyses in all datasets and traits, we prioritized the annotations of variants in the well-supported transcripts by applying the following parameter in VEP following a previous publication18: ‘-flag_pick -pick_order tsl,biotype,appris,rank,ccds,canonical,length’ unless otherwise noted. Then we only kept the annotation of each variant where ‘PICK = = 1’.

Construction of previously employed masks

From the 664 masks in shortlisted studies, we removed 63 masks that did not have well-documented definitions (Supplementary Table 2a). We then harmonized masks by translating the definitions into the closest equivalent in VEP. Specifically, (a) all MAF values were calculated from UK Biobank samples; (b) if PolyPhen234 was employed without specifying which version, we included both PolyPhen2-HDIV and PolyPhen2-HVAR; (c) if splice sites were included without further explanation, we included both splice donor variants and splice acceptor variants and excluded splice donor 5th base variants; (d) nonsynonymous with little explanation was considered missense; and (e) MAF less than or equal to a threshold was converted to MAF less than the threshold. We removed masks that could not be reconstructed using VEP annotations with reasonable minor adjustments (Supplementary Table 2a). Some examples include variants annotated as pathogenic to a specific trait based on experts’ opinions or certain databases, variants previously identified as causal, and variants located in constrained regions. After additional filtering based on variant classes and QC for association testing (Supplementary Note), 271 masks (making up 146 unique masking strategies within 163 studies) remained for downstream analyses.

Construction of new masks

We derived three types of new masks. First, we applied MAF < 1% and MAF < 0.1% thresholds to the 271 previously employed masks. Second, we systematically combined three MAF thresholds (all MAF, MAF < 1% and MAF < 0.1%) with high-confidence pLoF variants predicted by LOFTEE13 and missense variants predicted to be damaging by at least 25%, 50%, 75% and 90% of the 39 bioinformatic algorithms with rank scores in dbNSFP28,65 (original scores; Supplementary Tables 2d, 6 and 20a,b). Third, we followed the second approach but defined damaging scores based on ‘composite’ annotations derived from either PCA38 or independent-components analysis66 of the bioinformatic algorithms (Supplementary Tables 2d, 6 and 20a,c–e and Supplementary Note).

Association testing

We conducted single-variant and gene-level burden association tests using REGENIE39 (v3.1.2 and v3.2.5) adjusting for sex, genotyping array and the first five PCs as covariates. We used the single-variant association results for conditional analysis of burden tests (Supplementary Note).

In the ~190 K analysis, we conducted 16,445 gene-level burden tests for 298 previously employed masks and a synonymous mask as a negative control across 55 phenotypes using REGENIE39 (v3.1.2) implemented in dig-loam-stream pipeline (https://github.com/broadinstitute/dig-loam-stream). We removed 6 highly inflated masks with an average lambda across 55 phenotypes > 1.2 (Supplementary Table 5a), as well as all analyses for height as its association results were much more inflated than for the other 54 phenotypes (Supplementary Table 5b). As a follow-up analysis, we conducted 864 gene-level burden tests for the same 54 phenotypes and 16 proposed baseline masks (see below) using ~440 K UK Biobank samples and obtained published results from three high-profile UK Biobank studies22,23,32 (Supplementary Fig. 10a–c and Supplementary Note).

All association tests were two-sided. A gene-level association in each mask was considered (exome-wide) significant with P < 2.5 × 10−6. The Bonferroni significance threshold for a masking strategy was set at P < 2.5 × 10−6/number of masks.

Exploration of potential baseline masking strategies and our proposal

Potential baseline masking strategies

Reference masking strategies. For each of the total, low-frequency and rare variant association analyses, we defined the ‘best’ masking strategy previously employed as the masking strategy used in a past study that detected the largest number of significant associations. Similarly, we defined the ‘average’ masking strategy previously employed as a proxy strategy that detected the average number of significant associations across all strategies used in past studies.

‘High-profile’ masking strategies. The high-profile masking strategies were used in three recent high-profile WES studies of the UK Biobank22,23,32.

‘Brute-force’ masking strategy. The brute-force strategy refers to using all 271 masks to count the number of significant associations.

‘Clustered’ and ‘subclustered’ masking strategies. We applied PCA in scikit-learn67, using the binary inclusion of UK Biobank variants as features for each mask (variant membership clustering). We then applied k-means clustering in scikit-learn67 to cluster the masks based on the first 9 PCs that explained 90% of the total variance (Supplementary Table 9a). The cost function of k-means clustering reduced minimally after 10 clusters, suggesting that 10 was the optimal number of clusters (Supplementary Fig. 3e). Because there was substantial heterogeneity within each cluster (silhouette score = 0.51; Supplementary Fig. 3f), we further applied PCA and k-means to each cluster and identified 37 subclusters (variant MAF clustering; Supplementary Table 9b). In each of the clusters and subclusters, we selected the mask with the largest number of significant associations as the ‘representative’ mask (Supplementary Table 9c,d).

‘Greedy’ and ‘expanded greedy’ masking strategies. We applied a divide-and-conquer approach to group the masking strategies into subgroups consisting of the same number of masks, that is, having the same size m (m [1.271]), and hence, the same Bonferroni correction penalty. Next, to determine which masking strategy detected the largest number of significant associations (highest-yield strategy) within each subgroup of size m, we applied a depth-first search as follows: (1) for each mask, obtain significant associations (P < 2.5 × 10−6/m), (2) order the masks based on the number of significant associations, (3) pick the mask with the largest number of significant associations, (4) remove the significant associations already detected by the selected mask from the remaining masks, (5) reorder the masks based on the number of significant associations and (6) repeat steps 3 to 5 until m masks were selected. Finally, we compared the number of significant associations detected by the highest-yield strategies of all 271 subgroups and selected the overall highest-yield strategy. We defined this as the greedy masking strategy.

To examine the robustness of the greedy algorithm, we conducted a cross-validation test for each of the total, low-frequency and rare variant association analyses. For each phenotype p, we applied the greedy algorithm to obtain the trait-specific greedy masking strategy that detected the maximum number of significant associations. Then, we applied this method to the other 53 traits altogether and compared the significant associations for p detected by the 53-trait greedy masking strategy to the maximum number of significant associations for p (Supplementary Fig. 4, Supplementary Table 11 and Supplementary Note). We applied the greedy algorithm to the original library of 271 masks and the expanded library of 424 masks to obtain the greedy and expanded greedy masking strategies (Supplementary Note).

Compute time and cost of running the baseline masking strategies

To run 8 masks in each of the baseline strategies at once for ALT levels, we added a fourth column to the REGENIE annotation files. The third column in the annotation files contained the mask identifiers but the fourth column contained a list of all masks. As a result, REGENIE output mask-specific burden test results for all masks together. We extracted compute time from the ‘Elapsed time’ provided in REGENIE’s standard log files. We extracted the maximum memory usage from the ‘ru_maxrss’ provided by the high-compute clusters by running the ‘qacct’ command after the job completion.

Gene set enrichment analyses

In the ~440 K analysis, for each gene/trait pair, we applied the ‘minimum P-value test’ to aggregate the mask-level P values produced by each of the low-frequency and rare variant baseline strategies into one aggregate P value18. In short, the minimum P value test calculates the correlations between masks using the allele frequencies of the variants included in the masks, infers the ‘effective’ number of masks, and corrects the strongest P value based on the effective number of masks following the family-wise error rate formula. We then conducted gene set enrichment analysis for 4,203 gene sets from the Mouse Genome Informatics database40 by applying a one-sided Wilcoxon rank-sum test to evaluate whether genes in a gene set ranked significantly higher than in the matched background gene set provided in a previous publication18. We finally applied false discovery rate correction and considered a gene set to be significantly enriched for a trait if the false discovery rate-adjusted P < 0.05.

Analysis of 11 binary traits in the UK Biobank

We selected 11 diseases with a wide range of case frequencies in the ~440 K samples of the UK Biobank for additional analyses (Supplementary Table 16a). We converted the PheCode of each disease to a list of ICD-10 codes following previously established maps (https://github.com/PheWAS/PheWAS). We assigned each sample as a case if at least one of the ICD-10 codes was present for a disease. Data QC (variant QC, ancestry inference and sample QC), variant annotation and variant grouping were similar to what was described above. We conducted burden tests on these 11 diseases using the baseline masking strategies in REGENIE. For step 1, the following parameters were set: ‘–bsize 200 –bt’. For step 2, the following parameters were set: ‘-minMAC 1 -aaf-bins 0.5 -bsize 200 -build-mask sum -bt -write-mask-snplist’.

Analysis of 12 continuous traits in the AMP T2D GENES cohort

We obtained the data (40,054 samples, 29,283 (73.1%) of which were non-European) from the AMP T2D GENES consortium26 (https://www.niddk.nih.gov/about-niddk/research-areas/diabetes/accelerating-medicines-partnership-type-2-diabetes). All participants provided informed consent, and all samples were approved for use at their respective institution15,18,6870. We annotated the variants and ran burden tests using the masks in the baseline masking strategies on 12 continuous traits (Supplementary Table 15b), following the procedures for the analyses of the continuous traits in the UK Biobank detailed above.

Analysis of 56 traits in the All of Us Research Program

We used the V8 release of AoU, with QC as described in the documentation (https://support.researchallofus.org/hc/en-us/articles/29390274413716-All-of-Us-Genomic-Quality-Report), resulting in PLINK format datasets of 414,830 samples. We applied additional QC procedures and retained 410,400 samples, of which 315,536 samples had electronic health record data available (Supplementary Note). Access to the AoU resource was granted by the AoU Institutional Review Board, and use of AoU data was approved under a data use agreement between Amsterdam University Medical Center and the AoU Research Program.

We extracted the phenotypic values for 11 binary traits (Supplementary Table 16b) and 45 quantitative traits (Supplementary Table 15a). For the binary traits, only participants with electronic health record data were included. We mapped the PheCode of each trait to a set of ICD-10 codes following previously established maps (https://github.com/PheWAS/PheWAS). For quantitative traits (Supplementary Table 15a), we excluded missing values, retained only the most recent measurements, harmonized units, excluded outliers >6 standard deviations from the mean, and excluded physiologically implausible values. Blood pressure, LDL cholesterol and apolipoprotein B values were adjusted for medication use (Supplementary Note). All quantitative traits were then subjected to rank-based inverse normal transformation.

For burden tests, for step 1 of REGENIE (null model fitting), we used genome-wide common variants from the ACAF files (provided by AoU), followed by variant QC (Supplementary Note) and a pruning step. We used default parameters for quantitative traits, whereas for binary traits, the ‘-loocv’ approach, the ‘-bt’ flag and the ‘-write-null-firth’ flag were applied. For step 2 of REGENIE (association testing), we used approximate Firth’s regression models with back-correction of standard errors; burden masks were constructed using the ‘-sum’ option.

Analysis of traits of varying genetic architecture

We used stratified LD fourth moments regression55 (S-LD4M) to estimate common-variant polygenicity, defined as the effective number of independently associated SNPs (Me). We extracted the log10MeCommon for 12 traits (displayed in Table 1 of O’Connor et al.55) that were available in our analyses of the UK Biobank samples. The size of the union (number of unique genes) was used as a proxy for low-frequency/rare variant polygenicity.

Analysis of coverage of variants of masks across ancestries

To examine variant coverage of masks across different ancestries, we reconducted the PCA analysis on the binary variant membership of the 271 masks using ancestry-specific samples for 6 ancestries. For each of the first 8 PCs, we then calculated cosine similarity between 15 pairs of ancestries, resulting in 120 (8 × 15) comparisons. To assess the relative burden of genes in the same mask across ancestries, we calculated the number of variants per gene within each mask and each ancestry. We then applied Spearman’s correlation to each mask for each pair of ancestries, resulting in 4,065 (271 × 15) comparisons.

Analysis of burden test software

We compared the burden test results (using the ~440 K UK Biobank samples) produced by REGENIE versus SAIGE-GENE+ (Supplementary Note) using the baseline rare masks for five continuous traits (ALP, eGFR-cystatin C, HbA1C, LDL cholesterol, platelet volume) and 5 binary traits (hypertension, type 2 diabetes, atrial fibrillation or flutter, cardiomyopathy, hereditary anemia).

Analysis of burden tests on conserved domains

To run the burden tests on only the conserved domains (instead of the full gene length), we first re-ran VEP to re-annotate the variants using an additional flag ‘-domains’. This added a ‘DOMAINS’ column in the VEP output, representing any protein domains, provided by PDBe71, Pfam72, Prosite73 and InterPro74, that overlap a variant. We then excluded variants that had empty (represented by ‘-’ in VEP) values in the ‘DOMAINS’ column.

Analysis of transcript-level burden tests

We re-ran VEP to obtain transcript-level variant annotations following three commonly used approaches:

  1. Selecting the annotation in the transcripts whose structures are well-supported by experimental evidence (the main approach used throughout our study, by using ‘-flag_pick -pick_order tsl,biotype,appris,rank,ccds,canonical,length’ in VEP, and selecting rows where ‘PICK’ == ‘1’).

  2. Selecting the annotation in the canonical transcripts (by using ‘-canonical’ option and selecting only rows where ‘CANONICAL’ == ‘1’).

  3. Selecting the most severe consequence across all transcripts (by using ‘-flag_pick -pick_order rank,tsl,biotype,appris,ccds,canonical,length’ in VEP, and selecting rows where ‘PICK’ == ‘1’).

Finally, we also kept all variant annotations, collapsed variants into each of all available transcripts for each gene, and ran burden tests on each of the transcripts. We then applied a ‘minimum P-value test’18 to correct for the multiple transcripts in each gene to obtain an overall gene-level P value for each gene/trait/mask association.

Analysis of alternative masking strategies to the baseline

We defined ‘interpretable’ masks as those consisting of only pLoF variants (defined as high-confidence by LOFTEE), missense variants predicted to be damaging by a straightforward selection of bioinformatic algorithms (eg, PolyPhen2, REVEL, CADD), or all functional coding variants (Supplementary Table 6). We designed two ‘simple’ masking strategies to include four ‘nested’ masks that included either low-frequency (MAF < 1%) or rare (MAF < 0.1%) variants: pLoF mask (defined as high-confidence by LOFTEE), pLoF or missense variants predicted to be damaging by 50% of the algorithms, pLoF or missense variants predicted to be damaging by 25% of the algorithms, and all exonic variants (Supplementary Table 18). We constructed the ‘pLoF-inclusive’ masking strategies by including pLoF masks by default and reapplying the greedy algorithm to the other 423 masks in the expanded library (Supplementary Note).

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Online content

Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41588-026-02597-9.

Supplementary information

Supplementary Information (4.9MB, pdf)

Supplementary Notes and Figs. 1–11.

Reporting Summary (1.7MB, pdf)
Peer Review File (2MB, pdf)
Supplementary Tables (686KB, xlsx)

Supplementary Tables 1–22.

Acknowledgements

This research has been conducted using the UK Biobank Resource under application number 41189. We thank the UK Biobank resources and its participants for their essential contribution to the conduct of this research. We gratefully acknowledge the contributions of the All of Us participants, whose involvement made this research possible. We also thank the National Institutes of Health’s All of Us Research Program for granting access to the participant data utilized in this study. T.N., J.F., R.K., O.R., D.J., P.S., A.M., Q.H. and N.P.B. were supported by National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) grant UM1 DK105554-06 and National Human Genome Research Institute (NHGRI) grant U24 HG011453-01. J.F. and A.L. were supported by NIDDK grant R01 DK125490. C.R.B. was supported by funding from the Pathfinder Cardiogenomics program of the European Innovation Council of the European Union (DCM-NEXT and Nav1.5-CARED projects). S.J.J. was supported by funding from the Dutch Heart Foundation (03-007-2022-0035) and the Amsterdam University Medical Center. C.R.B. and S.J.J. were also supported through the KIC program funded by the Dutch Research Council and Dutch Heart Foundation (PRECISE-CVD project).

Extended data

Author contributions

T.N. and J.F. developed the study design, performed the analyses and wrote the paper. R.K. and T.N. performed data QC and association tests in the UK Biobank and AMP T2D GENES datasets. P.H. performed QC and statistical analyses in the All of Us dataset, under supervision of S.J.J. S.J.J., P.D. and S.Y. provided important feedback on the research questions and result interpretations. A.L., D.J., P.S., A.M., Q.H., O.R., C.R.B., P.E. and N.P.B. participated in discussions and provided expertise on the development of the script and browser.

Peer review

Peer review information

Nature Genetics thanks Heiko Runz and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.

Data availability

All genotyping data for the ~190 K analysis was accessed from the UK Biobank (Application ID: 41189). All genotyping data for the ~440 K analysis is accessible on the UK Biobank Research Analysis Platform (https://www.ukbiobank.ac.uk/use-our-data/research-analysis-platform/). Variant group files for the UK Biobank samples and gene-level burden results using the UK Biobank and All of Us samples are publicly accessible on a web browser (https://hugeamp.org/research.html?pageid=proposed_baseline_masks_home). The web browser was developed using Bring Your Own Results services (BYOR; https://byor.science) provided by the Knowledge Portal Network at the Broad Institute of MIT and Harvard.

Code availability

We processed individual-level data using KING v2.2.8 and v2.3.2 (https://www.kingrelatedness.com), PLINK v1.9 and v2.3a (https://www.cog-genomics.org/plink) and Hail v0.2.95 (https://github.com/hail-is/hail). We annotated the variants using VEP v110.1 (https://github.com/Ensembl/ensembl-vep). We performed the burden tests using REGENIE v3.1.2 and v3.2.5 (https://rgcgithub.github.io/regenie/) and SAIGE-GENE+ v1.0.9 (https://saigegit.github.io/SAIGE-doc/). We also developed a publicly available script (https://github.com/broadinstitute/genemasker; https://zenodo.org/records/19161631) to annotate a user-provided variant list and create group files for the baseline strategies, with options to customize different approaches for variant annotation and aggregation.

Competing interests

As of April 2026, P.D. is an employee and stockholder of Regeneron Pharmaceuticals. The other authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Extended data

is available for this paper at 10.1038/s41588-026-02597-9.

Supplementary information

The online version contains supplementary material available at 10.1038/s41588-026-02597-9.

References

  • 1.Kiezun, A. et al. Exome sequencing and the genetic basis of complex traits. Nat. Genet.44, 623–630 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Lee, S., Abecasis, G. R., Boehnke, M. & Lin, X. Rare-variant association analysis: study designs and statistical tests. Am. J. Hum. Genet.95, 5–23 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Majithia, A. R. et al. Rare variants in PPARG with decreased activity in adipocyte differentiation are associated with increased risk of type 2 diabetes. Proc. Natl Acad. Sci. USA111, 13127–13132 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wu, M. C. et al. Rare-variant association testing for sequencing data with the sequence kernel association test. Am. J. Hum. Genet.89, 82–93 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Lee, S. et al. Optimal unified approach for rare-variant association testing with application to small-sample case-control whole-exome sequencing studies. Am. J. Hum. Genet.91, 224–237 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Li, B. & Leal, S. M. Methods for detecting associations with rare variants for common diseases: application to analysis of sequence data. Am. J. Hum. Genet.83, 311–321 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Price, A. L. et al. Pooled association tests for rare variants in exon-resequencing studies. Am. J. Hum. Genet.86, 832–838 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Akbari, P. et al. Sequencing of 640,000 exomes identifies GPR75 variants associated with protection from obesity. Science373, eabf8683 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zhou, S. et al. Converging evidence from exome sequencing and common variants implicates target genes for osteoporosis. Nat. Genet.55, 1277–1287 (2023). [DOI] [PubMed] [Google Scholar]
  • 10.Zuk, O. et al. Searching for missing heritability: designing rare variant association studies. Proc. Natl Acad. Sci. USA111, E455–E464 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Romeo, S. et al. Rare loss-of-function mutations in ANGPTL family members contribute to plasma triglyceride levels in humans. J. Clin. Invest.119, 70–79 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Cohen, J. et al. Low LDL cholesterol in individuals of African descent resulting from frequent nonsense mutations in PCSK9. Nat. Genet.37, 161–165 (2005). [DOI] [PubMed] [Google Scholar]
  • 13.Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature581, 434–443 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Purcell, S. M. et al. A polygenic burden of rare disruptive mutations in schizophrenia. Nature506, 185–190 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Fuchsberger, C. et al. The genetic architecture of type 2 diabetes. Nature536, 41–47 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mahmood, K. et al. Variant effect prediction tools assessed using independent, functional assay-based datasets: implications for discovery and diagnostics. Hum. Genomics11, 10 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Liu, Y., Yeung, W. S. B., Chiu, P. C. N. & Cao, D. Computational approaches for predicting variant impact: An overview from resources, principles to applications. Front. Genet.13, 981005 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Flannick, J. et al. Exome sequencing of 20,791 cases of type 2 diabetes and 24,440 controls. Nature570, 71–76 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Weiner, D. J. et al. Polygenic architecture of rare coding variation across 394,783 exomes. Nature614, 492–499 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Butler-Laporte, G. et al. Exome-wide association study to identify rare variants influencing COVID-19 outcomes: results from the Host Genetics Initiative. PLoS Genet.18, e1010367 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Praveen, K. et al. Population-scale analysis of common and rare genetic variation associated with hearing loss in adults. Commun. Biol.5, 540 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature599, 628–634 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Karczewski, K. J. et al. Systematic single-variant and gene-based association testing of thousands of phenotypes in 394,841 UK Biobank exomes. Cell Genomics2, 100168 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature562, 203–209 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.All of Us Research Program Genomics Investigators. Genomic data in the All of Us Research Program. Nature627, 340–346 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Costanzo, M. C. et al. Accelerating medicines partnership in type 2 diabetes and common metabolic diseases: collaborating to maximize the value of genetic and genomic data. Diabetes74, 1089–1098 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.McLaren, W. et al. The Ensembl Variant Effect Predictor. Genome Biol.17, 122 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Liu, X., Li, C., Mou, C., Dong, Y. & Tu, Y. dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med.12, 103 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Jordan, C. T. et al. Rare and common variants in CARD14, encoding an epidermal regulator of NF-kappaB, in psoriasis. Am. J. Hum. Genet.90, 796–808 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Gaare, J. J. et al. Rare genetic variation in mitochondrial pathways influences the risk for Parkinson’s disease. Mov. Disord.33, 1591–1600 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Armstrong, M. J. et al. Role of TET1-mediated epigenetic modulation in Alzheimer’s disease. Neurobiol. Dis.185, 106257 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Wang, Q. et al. Rare variant contribution to human disease in 281,104 UK Biobank exomes. Nature597, 527–532 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sim, N.-L. et al. SIFT web server: predicting effects of amino acid substitutions on proteins. Nucleic Acids Res.40, W452–W457 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Adzhubei, I., Jordan, D. M. & Sunyaev, S. R. Predicting functional effect of human missense mutations using PolyPhen-2. Curr. Protoc. Hum. Genet.7, Unit7.20 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kirino, Y. et al. Targeted resequencing implicates the familial Mediterranean fever gene MEFV and the toll-like receptor 4 gene TLR4 in Behçet disease. Proc. Natl Acad. Sci. USA110, 8134–8139 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Jurgens, S. J. et al. Adjusting for common variant polygenic scores improves yield in rare variant association analyses. Nat. Genet.55, 544–548 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Laver, T. W. et al. Evaluation of evidence for pathogenicity demonstrates that BLK, KLF11, and PAX4 should not be included in diagnostic testing for MODY. Diabetes71, 1128–1136 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Jolliffe, I. T. & Cadima, J. Principal component analysis: a review and recent developments. Philos. Transact. A Math. Phys. Eng Sci.374, 20150202 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Mbatchou, J. et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet.53, 1097–1103 (2021). [DOI] [PubMed] [Google Scholar]
  • 40.Baldarelli, R. M. et al. Mouse Genome Informatics: an integrated knowledgebase system for the laboratory mouse. Genetics227, iyae031 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Costanzo, M. C. et al. Realizing the promise of genome-wide association studies for effector gene prediction. Nat. Genet.57, 1578–1587 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Chen, M.-H. et al. Trans-ethnic and ancestry-specific blood-cell genetics in 746,667 individuals from 5 global populations. Cell182, 1198–1213.e14 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Sinnott-Armstrong, N. et al. Genetics of 35 blood and urine biomarkers in the UK Biobank. Nat. Genet.53, 185–194 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Saint-Laurent, C. et al. The tyrosine phosphatase SHP2: a new target for insulin resistance? Biomedicines10, 2139 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Sun, M. et al. PTPN11 is a potential biomarker for type 2 diabetes mellitus complicated with colorectal cancer. Sci. Rep.14, 25155 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Senis, Y. A. Protein-tyrosine phosphatases: a new frontier in platelet signal transduction. J. Thromb. Haemost.11, 1800–1813 (2013). [DOI] [PubMed] [Google Scholar]
  • 47.Kiratikanon, S., Chattipakorn, S. C., Chattipakorn, N. & Kumfu, S. The regulatory effects of PTPN6 on inflammatory process: reports from mice to men. Arch. Biochem. Biophys.721, 109189 (2022). [DOI] [PubMed] [Google Scholar]
  • 48.Jiang, J. et al. Platelet ITGA2B inhibits caspase-8 and Rip3/Mlkl-dependent platelet death through PTPN6 during sepsis. iScience26, 107414 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Kanwal, Z. et al. Deficiency in hematopoietic phosphatase ptpn6/Shp1 hyperactivates the innate immune system and impairs control of bacterial infections in zebrafish embryos. J. Immunol.190, 1631–1645 (2013). [DOI] [PubMed] [Google Scholar]
  • 50.Gao, X. et al. Role of ASXL1 in hematopoiesis and myeloid diseases. Exp. Hematol.115, 14–19 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Lin, I. et al. Multiomics of Bohring-Opitz syndrome truncating ASXL1 mutations identify canonical and noncanonical Wnt signaling dysregulation. JCI Insight8, e167744 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Moon, S. et al. Asxl1 exerts an antiproliferative effect on mouse lung maturation via epigenetic repression of the E2f1-Nmyc axis. Cell Death Dis.9, 1118 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Arioka, M. et al. ASXL1-related Bohring-Optiz syndrome complicated by persistent neonatal pulmonary hypertension and abnormal alveoli formation. Eur. J. Med. Genet.72, 104978 (2024). [DOI] [PubMed] [Google Scholar]
  • 54.Nguyen, T. et al. A resource of ‘bottom-line’ variant associations for 1,281 complex traits by integrating data across published genome-wide association studies. Preprint at Research Square10.21203/rs.3.rs-8585052/v1 (2026).
  • 55.O’Connor, L. J. et al. Extreme polygenicity of complex traits is explained by negative selection. Am. J. Hum. Genet.105, 456–476 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Hindy, G. et al. Rare coding variants in 35 genes associate with circulating lipid levels—a multi-ancestry analysis of 170,000 exomes. Am. J. Hum. Genet.109, 81–96 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Jurgens, S. J. et al. Rare coding variant analysis for human diseases across biobanks and ancestries. Nat. Genet.56, 1811–1820 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Zhou, W. et al. SAIGE-GENE+ improves the efficiency and accuracy of set-based rare variant association tests. Nat. Genet.54, 1466–1469 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Weale, M. E. in Genetic Variation 1st edn (eds Barnes, M. R. & Breen, G.) Vol. 628, 341–372 (Humana, 2010).
  • 60.Purcell, S. et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet.81, 559–575 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience4, 7 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Abraham, G., Qiu, Y. & Inouye, M. FlashPCA2: principal component analysis of Biobank-scale genotype datasets. Bioinformatics33, 2776–2778 (2017). [DOI] [PubMed] [Google Scholar]
  • 63.Kramer, O. in Dimensionality Reduction with Unsupervised Nearest Neighbors (ed. Kramer, O.) 13–23 (Springer, 2013).
  • 64.Venables, W. N., Ripley, B. D. Modern Applied Statistics with S 4th edn (Springer, 2002).
  • 65.Jurgens, S. J. et al. Analysis of rare genetic variation underlying cardiometabolic diseases and traits among 200,000 individuals in the UK Biobank. Nat. Genet.54, 240–250 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Hyvärinen, A. Independent component analysis: recent advances. Philos. Transact. A Math. Phys. Eng Sci.371, 20110534 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
  • 68.Fu, W. et al. Analysis of 6,515 exomes reveals the recent origin of most human protein-coding variants. Nature493, 216–220 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Lohmueller, K. E. et al. Whole-exome sequencing of 2,000 Danish individuals and the role of rare coding variants in type 2 diabetes. Am. J. Hum. Genet.93, 1072–1086 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Williams, A. L. et al. Sequence variants in SLC16A11 are a common risk factor for type 2 diabetes in Mexico. Nature506, 97–101 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Armstrong, D. R. PDBe: improved findability of macromolecular structure data in the PDB. Nucleic Acids Res.48, D335–D343 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Finn, R. D. et al. Pfam: the protein families database. Nucleic Acids Res.42, D222–D230 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Hulo, N. et al. The PROSITE database. Nucleic Acids Res.34, D227–D230 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Apweiler, R. et al. The InterPro database, an integrated documentation resource for protein families, domains and functional sites. Nucleic Acids Res.29, 37–40 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (4.9MB, pdf)

Supplementary Notes and Figs. 1–11.

Reporting Summary (1.7MB, pdf)
Peer Review File (2MB, pdf)
Supplementary Tables (686KB, xlsx)

Supplementary Tables 1–22.

Data Availability Statement

All genotyping data for the ~190 K analysis was accessed from the UK Biobank (Application ID: 41189). All genotyping data for the ~440 K analysis is accessible on the UK Biobank Research Analysis Platform (https://www.ukbiobank.ac.uk/use-our-data/research-analysis-platform/). Variant group files for the UK Biobank samples and gene-level burden results using the UK Biobank and All of Us samples are publicly accessible on a web browser (https://hugeamp.org/research.html?pageid=proposed_baseline_masks_home). The web browser was developed using Bring Your Own Results services (BYOR; https://byor.science) provided by the Knowledge Portal Network at the Broad Institute of MIT and Harvard.

We processed individual-level data using KING v2.2.8 and v2.3.2 (https://www.kingrelatedness.com), PLINK v1.9 and v2.3a (https://www.cog-genomics.org/plink) and Hail v0.2.95 (https://github.com/hail-is/hail). We annotated the variants using VEP v110.1 (https://github.com/Ensembl/ensembl-vep). We performed the burden tests using REGENIE v3.1.2 and v3.2.5 (https://rgcgithub.github.io/regenie/) and SAIGE-GENE+ v1.0.9 (https://saigegit.github.io/SAIGE-doc/). We also developed a publicly available script (https://github.com/broadinstitute/genemasker; https://zenodo.org/records/19161631) to annotate a user-provided variant list and create group files for the baseline strategies, with options to customize different approaches for variant annotation and aggregation.


Articles from Nature Genetics are provided here courtesy of Nature Publishing Group

RESOURCES