Abstract
Variants in the familial hypercholesterolemia gene LDLR—the most important genetic driver of cardiovascular disease—can raise circulating low-density lipoprotein (LDL) cholesterol concentrations and increase the risk of premature atherosclerosis. Definitive classifications are lacking for nearly half of clinically-encountered LDLR missense variants, limiting interventions that reduce disease burden. Here, we tested the impact of ~17,000 (nearly all possible) LDLR missense coding variants on both LDLR cell-surface abundance and LDL uptake, yielding sequence–function maps that recapitulate known biochemistry, offer functional insights, and provide evidence for interpreting clinical variants. Functional scores correlated with hyperlipidemia phenotypes in prospective human cohorts and augmented polygenic scores to improve risk inference, highlighting the potential of this resource to accelerate familial hypercholesterolemia diagnosis and improve patient outcomes.
Heterozygous familial hypercholesterolemia (HeFH)—defined by an increase in circulating low-density lipoprotein (LDL) cholesterol—is among the most common and severe genetic causes of cardiovascular disease (1, 2). The early adoption of LDL-lowering therapeutics can prevent atherosclerosis and prolong life by reducing the risk of myocardial infarction and all-cause mortality.
The LDL receptor (LDLR) protein is the primary mechanism by which LDL is cleared from the circulation: LDLR binds to Apolipoprotein B100 (the main structural component of LDL) via its ligand-binding domain, the complex is endocytosed, and LDL is degraded (3–5).
Genetic variants in LDLR can disrupt the clearance of LDL, accounting for ~ 80% of molecularly diagnosed HeFH cases (6). Identification of a pathogenic LDLR variant facilitates accurate diagnosis and enables prognoses and therapeutic management beyond standard lipid profiling alone (6–8), with potential to extend these benefits to at-risk relatives identified through cascade screening (9, 10). HeFH is one of three diseases deemed Tier 1 Genomics Applications by the Center for Disease Control and Prevention (CDC) because of the clear benefits arising from early pathogenic variant detection (11). Most LDLR coding variants, however, lack a definitive clinical classification (12); moreover, even where variants have been classified, quantitative estimates of variant impact are often unavailable. This currently limits opportunities for early diagnosis and patient risk stratification for HeFH.
Experimental functional studies can provide strong evidence for clinical variant classification (13). Although a subset of clinically-observed variants have already been assayed (14–19), nearly all possible single-nucleotide variants (in LDLR and all other genes) already exist in the human population (20, 21). We therefore built on technological advances that have enabled the measurement of variant impacts at scale (22) to comprehensively measure the functional impact of nearly all possible single LDLR amino acid substitutions. We assayed both cellular LDL uptake and LDLR cell-surface abundance as quantitative cellular read-outs relevant to HeFH pathophysiology. The resulting sequence–function maps not only reflect our current understanding of LDLR function, but also reveal unexpected biochemical insights, and offer the potential to inform clinical variant interpretation and improve patient risk estimation for HeFH.
Results
A sequence–function map of LDLR
To develop a multiplexed in vitro assay of LDLR variant functions, we first focused on the role of LDLR in the binding and internalization of LDL particles—the mechanism underlying disease-related dysfunction for many LDLR variants (23). We adapted an assay that uses dye-labelled LDL particles that fluoresce most strongly at endosomal pH (24) to quantify cellular LDL internalization (Fig. 1, A and B). Because LDLR expression and the ability to take up LDL are observed across many cell types (25), we adopted HeLa cells as a tractable model that could enable a large-scale study (see Methods). We next generated a set of barcoded cDNA libraries collectively containing nearly all possible single amino acid substitutions in LDLR. To generate heterogeneous pools in which each cell expresses a single full-length LDLR cDNA variant, we introduced libraries at a genomically integrated Bxb1 “landing pad” (26) in HeLa cells lacking endogenous LDLR. Cells bearing a single integrated LDLR variant cDNA construct were selected for LDL uptake via fluorescence-activated cell sorting. Functional scores were calculated by comparing barcode abundance in selected and unselected populations and scaled such that scores of 0 and 1 correspond to the median behavior of nonsense and synonymous variants. We thereby measured LDL uptake function for 16677 (97%; post-quality filtering) of the 17200 possible coding substitutions in LDLR (Fig. 1C; data S1).
Fig. 1. A sequence–function map of the familial hypercholesterolemia gene LDLR.

(A) Overview of the experimental workflow and (B) phenotypic selections used in this study. (C) Functional scores measuring LDL uptake for all possible amino acid substitutions in LDLR. Color scale is shown at the bottom right: damaging (blue); normal (white); above normal (red); unmeasured (gray); reference amino acid (yellow). The vertical axis represents each possible amino acid (one-letter code) grouped by hydrophobic (hyd.) and polar residues (pol.; +/−). Error bars indicate standard error. (D) Score distributions for synonymous (green), missense (gray), and nonsense (blue) variants. (E) Functional scores as a measure of log (minor allele frequency) for variants present in the UK Biobank (dark gray), All of Us (light gray), and both cohorts (white).MAF, minor allele frequency; Hyd, hydrophobic; Pol, polar; UKB, UK Biobank; AoU, All of Us.
Multiple analyses supported the overall quality of this dataset. Score distributions for nonsense and synonymous variants were distinct (P < 0.001; Fig. 1D). Missense variants were bimodally distributed, with modes aligning to those of synonymous and nonsense variants. Functional scores were correlated with measures of conservation (27) (Pearson’s r = 0.5, P < 0.001; fig. S1A); protein instability (28) (r = −0.45, P < 0.001; fig. S1B); and with computational predictions of variant dysfunction from AlphaMissense (29) (Pearson’s r = −0.63, P < 0.001; fig. S1C), as well as functional scores from an independent in vitro study of LDLR variant function (15) (r = 0.51, P < 0.001; fig. S1D). Substitutions that lowered LDL uptake were depleted at high minor allele frequencies in the UK Biobank (30) and All of Us cohorts (31) (Fig. 1E). For example, variants with low (< 0.5) functional scores tended to be less common (μMAF = 1.7e-05) than those with high (≥ 0.5) scores (μMAF = 9.24e-05; P < 0.001, Mann-Whitney U).
Relating sequence–function maps to biochemistry
The ligand-binding domain—composed of seven LDLR type A (LA) modules—interacts with Apolipoprotein B100 (ApoB100; the main structural component of LDL) to facilitate lipoprotein clearance. Each of these LA modules harbors three specific disulfide bridges (six cysteines) and a single Ca2+ ion that is maintained in an octahedral cage (5). These residues, which are conserved both evolutionarily and between LA modules, are critical for both structure and function in LDL internalization (32, 33) [with the possible exception of LA1 and LA2 (34–36)]. Within the LA modules that were most evidently important for LDL uptake in our assay (LA3–5 and LA7; fig. S1, E and F), cysteine residues and acidic Ca2+-coordinating residues (except for those contributing a backbone-carboxyl group) were intolerant to substitution (for example, see LA4 and LA5 residues in fig. S2, A to D). Aromatic residues contributing a backbone-carboxyl group (such as W165 in LA4 and W214 in LA5) were also intolerant to substitution: this is likely due to their direct interaction with ApoB100 in LDL (36). Our functional scores thus recapitulate the key main biochemical features known for LA modules 3–5 and 7.
Surprisingly, substitutions in the remaining LA modules (LA1, LA2, and LA6) generally did not affect LDL uptake in our study (fig. S1E): only 8% of missense variants in these modules showed reduced function compared to 38% in LA modules 3–5 and 7 (P < 0.001, Mann-Whitney U). Although LA modules are often interpreted as repeats sharing similar functions, it is increasingly apparent that their functions are divergent (37). Reported structures of the LDLR-ApoB100 complex (36) suggest that while LA modules 3, 4, 5, and 7 appeared closely bound to the β-belt of ApoB100, LA1 and LA2 were either unbound or unresolved, and LA6 showed a smaller interaction interface (Fig. 2, A to D). Moreover, FH-associated variants in ApoB100 were specifically located at binding sites for LA3–5 and LA7. Our results are thus consistent with atomic-level details of the receptor-ligand interaction in suggesting that LA modules 1, 2, and 6 may be less important than other LA modules for ApoB100 binding.
Fig. 2. Functional scores capture atomic-level details of the LDLR-ApoB100 interaction.

(A) Ribbon diagram of LDLR (LA2 to EGF-A) bound to ApoB100 (residues 1833–3892; yellow) (PDB 9BDE). LDLR is colored by functional score (residue mean) scaled as in Fig. 1C. Insets show (B) LA5, (C) LA6, and (D) LA7-EGF-A (partial). ApoB100 LDLR binding sites A and B are shown in cyan; Ca2+ ions are shown in gray; FH variants in ApoB100 are annotated in red.
Our sequence–function map further captured known biochemistry beyond the LA modules. For instance, the β-propellor domain [thought to mediate conformational changes in the receptor (5)] appeared broadly intolerant to substitution, and the transmembrane domain of LDLR was, as expected, intolerant to polar residues. Disruption of the NPxY motif in the cytoplasmic tail [required for internalization of the receptor-ligand complex (38)] also impaired LDL uptake.
To distinguish substitutions that directly disrupt LDL uptake from those that act by lowering cell-surface levels of LDLR (for example, by affecting trafficking or causing misfolding and degradation), we also measured LDLR cell-surface abundance impacts at scale. The resulting map captured 16105 (94% of all possible) coding substitutions (fig. S3 & fig. S4, A to F; data S2). While the LA7 module, which adopts a fixed conformation (5, 36), appeared concordant between maps, nearly all (95%) of the amino acid substitutions within LA modules 1–6 were tolerated in the abundance assay (fig. S4, G and H) compared to 78% in the LDL uptake assay (Fig. 3; P < 0.001, Mann-Whitney U). After excluding these LA modules, cell-surface abundance scores were well correlated with functional scores (r = 0.84, P < 0.001; fig. S4F), suggesting that disease-causing variants in LA modules 1–6 disrupt receptor function through mechanisms other than altering LDLR stability and trafficking to the cell surface.
Fig. 3. Mapping variant function across readouts highlights functional patterns in the ligand-binding domain.

Functional scores for all possible amino acid substitutions in the LDLR ligand-binding domain (LA modules 1–7) measuring (top) LDLR cell-surface abundance; (middle) LDL uptake measured without VLDL (− VLDL); and (bottom) LDL uptake with a 4:1 stoichiometric excess of VLDL (+ VLDL). Functional scores are scaled as in Fig. 1C. Conserved disulfide-forming residues (:) and acidic Ca2+-coordinating (*) residues are annotated. Error bars indicate standard error. VLDL, very low-density lipoprotein.
Assessing VLDL-dependent variant impacts on LDL uptake
Although our observation that LA modules 1, 2, and 6 are tolerant to substitutions was supported by the recent ApoB-LDLR structure and by patterns of pathogenic variation in ApoB, we struggled to reconcile this observation with the fact that all seven modules are well conserved (32, 33) and known to harbor pathogenic missense variants (12). Given that LDLR also interacts with very low-density lipoprotein (VLDL), we hypothesized that LA modules 1, 2, and 6 serve in VLDL (rather than LDL) uptake. We therefore measured LDL uptake in the presence of a stoichiometric excess of exogenous VLDL, capturing the impact of 6106 (98%) substitutions in the ligand-binding domain (fig. S5, A to D; data S3). While LA1 substitutions appeared tolerated both with and without excess VLDL, several substitutions in LA2 and LA6 showed LDL uptake that was decreased, but only in the presence of excess VLDL (Fig. 3): 23% of missense variants in LA modules 2 and 6 showed reduced LDL uptake in the presence of VLDL compared to only 7% when measuring LDL uptake in the absence of other lipoprotein subtypes (P < 0.001, Mann-Whitney U). Moreover, in the presence of VLDL, substitutions that reduced function in LA2 and LA6 matched those at homologous positions in LA3–5 and LA7 (annotated in Fig. 3). The dependence of LDL uptake on VLDL was confirmed for a pathogenic variant (C95S) in LA2, and was not observed for a negative control pathogenic variant (C52Y) in LA1 (fig. S5, E to G). These variants were also assayed in the presence of other lipoprotein subtypes (for example, chylomicrons and intermediate-density lipoproteins), but the impact on LDL uptake was only observed in the presence of VLDL (fig. S5, H to M). Taken together, our functional maps suggest a role for modules LA2 and LA6 that could be both lipoprotein-specific and qualitatively different from that of LA modules 3–5 and 7. Although we propose one possible model (see Discussion), a mechanistic understanding of these findings—along with any in vivo implications they may have—will require further study.
Potential clinical applications of functional scores for identifying pathogenic variants
Functional scores for clinically-reported [ClinVar (39)] LDLR variants were bimodally distributed (fig. S6, A and B). Pathogenic variants scored significantly lower than benign variants both for annotated variants derived from ClinVar (P < 0.001, Mann-Whitney U; fig. S6A) and for a subset of variants annotated by the ClinGen FH Variant Curation Expert Panel (VCEP) (40) without prior use of functional evidence (PS3 or BS3) (P = 0.005, Mann-Whitney U; data S4). That more than 20% of variants lacking a definitive classification showed reduced function (fig. S6, B and C; leftmost peak) suggested the potential value of our scores in identifying additional pathogenic variants.
We calibrated functional scores to derive, for each score, both a likelihood of pathogenicity (a quantitative measure of evidence strength that can be applied within a Bayesian framework) and the corresponding cardinal evidence strength levels applied under current variant classification guidelines from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) (table S1) (41–44). Under current ClinGen guidelines for classifying HeFH variants (40), scores from our function map can be considered as level 1 evidence, in that they capture the complete spectrum of LDLR activity. Although the abundance map does not capture all functions, we argue that evidence towards pathogenicity (but not benignity) from this dataset should also be considered as level 1, given that the disruption of surface abundance represents disruption of overall LDLR function. Our sequence–function map provided evidence for 6425 LDLR missense variants, with 2022 variants receiving strong or very-strong evidence towards pathogenicity (fig. S6, D and E; data S5). We similarly calibrated LDLR surface abundance scores, excluding LA1–6 for which variation was generally tolerated in this map (fig. S6, F and G; data S5), yielding evidence towards pathogenicity for 1915 variants, of which 95% (1819) received at least strong evidence. Calibration of the LDL uptake map also indicated evidence towards benignity (all at the supporting level) for 4172 variants. Because measurements of LDLR surface abundance do not capture the full spectrum of LDLR functions, we did not derive evidence towards benignity from the abundance map. Of the 1053 variants that received evidence from both datasets, 85% (898) received evidence that was concordant with respect to pathogenicity or benignity, and no discordant variant received more than supporting evidence towards either pathogenicity or benignity. Together, LDLR function and abundance scores provided evidence for 7287 missense variants in LDLR—including over 400 clinical missense variants for which a definitive classification was not available (fig. S6, H and I; data S5).
Potential clinical applications of functional scores for quantifying HeFH risk
We next sought to evaluate the potential use of functional scores in quantifying individual risk of HeFH. We applied diagnostic criteria from the Dutch Lipid Clinic Network (DLCN) (45, 46) to participants in the prospective UK Biobank cohort, assigning HeFH status where a DLCN score indicated either “probable” HeFH or “definite” HeFH. (To avoid circularity, we did not apply any DLCN points based on the presence of a genetic variant.) Participants carrying a variant with a ‘damaging’ functional score (bottom 15th percentile of all scores) showed a high risk of HeFH (OR 21.9, 95% CI 14.4 to 33.2), comparable to the risk conferred from carrying an annotated pathogenic (pathogenic or likely pathogenic) variant (OR 22.5, 95% CI 16.7 to 30.4) (Fig. 4A, table S2). Carriers of variants with a ‘normal’ score (top 70th percentile) exhibited an HeFH risk (OR 1.3, 95% CI 1.1 to 1.5) that was lower than that observed among all LDLR variant carriers collectively (OR 1.6, 95% CI 1.4 to 1.8) and equivalent to the risk among participants carrying a benign (benign or likely benign) variant (OR 1.3, 95% CI 1.1 to 1.5).
Fig. 4. Functional scores enable HeFH risk quantification and diagnosis.

(A) Risk (odds ratio) of “probable” familial hypercholesterolemia (HeFH) for UK Biobank participants who are carrying LDLR variants of a given type: nonsynonymous variants (top); ClinVar variants lacking a definitive classification (VUS) (middle); or pathogenic (P/LP) variants (bottom). In each, risk is relative to noncarriers and is stratified by functional score. Dotted line denotes the odds ratio of HeFH among all LDLR variant carriers. (B) Receiver operating characteristic (ROC) curve describing the ability of functional scores to distinguish between two groups of LDLR variant carriers in the UK Biobank: those defined as having “definite” HeFH under Dutch Lipid Clinic Network criteria; and those not suspected of having HeFH, without considering genotype for either group (AUROC = 0.87.).
Functional scores further stratified risk among those carrying previously-known pathogenic variants. Relative to all pathogenic variant carriers, the risk of HeFH among carriers of a pathogenic variant with a damaging score was approximately 4-fold higher (OR 44.3, 95% CI 27.4 to 71.5) than the risk among carriers of normal-scoring pathogenic variants (OR 10.5, 95% CI 6.0 to 18.3) (Fig. 4A). We observed a similar increase in risk when restricting our analysis to VCEP-classified pathogenic variants (OR 24.1, 95% CI 11.8 to 49.3 vs OR 6.7, 95% 3.2, 14.2). Together, these results suggest that our quantitative functional evidence can augment existing variant pathogenicity annotations towards better inferring HeFH risk.
Potential clinical applications of functional scores for diagnosing HeFH
Under the DLCN criteria for the diagnosis of HeFH, the presence of a “functional mutation” in LDLR is by itself sufficient to support a probable HeFH diagnosis (45, 46). Towards identifying a clinically applicable score threshold for a “functional mutation”, we defined two subsets of UK Biobank participants who carry LDLR missense variants: i) participants who would receive a definitive diagnosis of HeFH under DLCN criteria (> 8 points) even without attributing points for genotype; and ii) carriers who would not be suspected of having HeFH (< 3 points), again without considering genotype. We next calculated a receiver operating characteristic curve describing the ability of functional scores to distinguish between these two participant subsets (AUROC = 0.87; Fig. 4B). Establishing a score threshold based on a false positive rate (FPR) of 5% (functional score ≤ 0.04) could offer a potential means for clinicians to identify “functional mutations” in the context of current diagnostic frameworks. We note that at this score threshold, participants exhibited a risk of probable HeFH (OR 21.2, 95% CI 13.8 to 32.8) that is equivalent to carrying a clinically annotated pathogenic variant (OR 22.5, 95% CI 16.7 to 30.4).
We next identified participants in the UK Biobank carrying a rare LDLR coding variant that either lacked a definitive clinical annotation in ClinVar or had no annotation whatsoever and applied DLCN criteria as above (without accounting for genotype). Of these 5190 carriers, 19 were assigned probable HeFH. Using the above threshold to meet the standard of “functional mutation” under the DLCN criteria, 116 additional participants would be moved to probable HeFH, and 132 would reach definite HeFH, thereby increasing the number of participants with at least a probable diagnosis of HeFH by over 10-fold. We note that our application of the DLCN criteria was likely conservative given that the relevant family history and clinical phenotypes (for example, xanthomas) will have often not been ascertained or, if so, reported in electronic health records (47, 48).
Potential for refining risk inference for hypercholesterolemia and coronary artery disease
Having related our functional scores to variant pathogenicity and HeFH diagnosis, we next sought to explore their potential utility for predicting HeFH-related phenotypes based on personal genome information in population-scale cohorts. We found functional scores to be significantly correlated with dyslipidemia phenotypes including elevated LDL-C (LDL-C ≥ 190 mg/dL) (Spearman’s ρ = −0.49, P = 0.01; Fig. 5A; see also fig. S7, A to C) and use of cholesterol-lowering medication (ρ = −0.47, P = 0.02; fig. S7D) among participants in the UK Biobank. Moreover, LDL-C levels were increased by 39 mg/dL among participants carrying low-scoring variants, and by 108 mg/dL among those carrying the most extreme-scoring variants, relative to the population mean of 146 mg/dL (Fig. 5B). We next examined markers of vascular health and cardiovascular disease (49, 50) and found that participants carrying a damaging variant were at increased risk, relative to those carrying normal-scoring variants, of having aortic valve stenosis (OR 2.1, 95% CI 1.2 to 3.7) or high (top quintile) intima-media thickness of the carotid artery (OR 2.2, 95% CI 1.1 to 4.3).
Fig. 5. Refining risk estimation for HeFH-related phenotypes.

(A) The correlation of functional scores with the frequency of elevated LDL-C (LDL-C ≥ 190 mg/dL) (Spearman’s ρ = −0.49; P = 0.01) among UK Biobank participants. The probability of each phenotype was assessed within a given functional score window (0.05 +/− 0.025) with > 5 participants per window. The dotted line denotes the mean of this measure among all LDLR variant carriers. (B) Elevation in LDL-C (mg/dL) among participants in the UK Biobank carrying an LDLR variant grouped by functional score; elevation is shown relative to the UK Biobank population average (146 mg/dL). Error bars denote standard error. The adjusted odds ratio of (C) hyperlipidemia (LDL-C ≥ 190 mg/dL) for participants stratified by functional score [normal (gray) and damaging (black)] and polygenic risk score quartile (for LDL-C and CAD, respectively). Odds ratios are adjusted for age, sex, BMI, and ethnicity. Non-carriers with intermediate PRS are used as reference. The dotted line denotes an OR of 1. (D) Time-to-hyperlipidemia diagnosis for UK Biobank participants—stratified by functional score, PRS, and both—is shown for participants with a: 25th percentile PRS (dotted gray); 75th percentile PRS (dotted pink); damaging functional score (solid pink); or 75th percentile PRS and damaging functional score (solid red). Time-to-diagnosis among all participants is shown as reference (dotted black). (E and F) Same as (C and D) assessing the odds ratio coronary artery disease (CAD) and CAD-free survival, respectively. Error bars in (C to F) indicate the 95th percentile CI. LDL, low-density lipoprotein; OR, odds ratio; HeFH, familial hypercholesterolemia; CAD, coronary artery disease; PRS, polygenic risk score; Norm., normal; Dam., damaging.
We next assessed the effects of polygenic background and functional scores on the risk of hyperlipidemia (LDL-C ≥ 190 mg/dL). UK Biobank participants with a high (top quartile) polygenic risk score (PRS) for LDL-C (51) showed a ~2-fold (Adj. OR 2.3, 95% CI 2.3 to 2.4) increased risk of hyperlipidemia, with lower PRS (bottom quartile) indicating a protective effect (Adj. OR 0.4, 95% CI 0.3 to 0.4) (fig. S7E; table S3). By comparison, rare damaging functional variants conferred a nearly 3-fold increase in risk (Adj. OR 6.1, 95% CI 5.0 to 7.4) (fig. S7F; table S3). To evaluate the combined effects of common and rare genotypes on risk, we again grouped participants by PRS and compared carriers harboring rare (MAF ≤ 0.01) damaging LDLR variants to those with normal-scoring variants. Previous evaluations have shown that the risk captured by PRS did not meaningfully depend on variation within monogenic HeFH genes (LDLR, APOB, or PCSK9) or on the LDL cholesterol uptake pathway (52, 53). Among participants carrying a damaging LDLR variant, the risk of hyperlipidemia was nearly 3-fold higher for those in the highest quartile of PRS (Adj. OR 8.4, 95% CI 5.9 to 11.9) compared to those in the lowest quartile (Adj. OR 2.9, 95% CI 1.7 to 4.9) (Fig. 5C; table S4). Thus, while variants with damaging functional scores conferred a more severe risk of hyperlipidemia, PRS further informed participant risk regardless of LDLR genotype.
We next examined the effect of genotype on the age of onset for hyperlipidemia. After first finding that functional scores correlated with the age at hyperlipidemia diagnosis (based on ICD10 code E78.0; ρ = 0.43, P = 0.03; fig. S7C), we again stratified participants by functional score and PRS and estimated event-free survival. PRS again showed modest effect on time-to-diagnosis compared to damaging LDLR variants, but those with both a damaging LDLR variant and a PRS score in the highest quartile showed the earliest onset of hyperlipidemia (Fig. 5D).
Given the established relationship between HeFH (and hyperlipidemia more broadly) and coronary artery disease (CAD), one might expect functional scores to also stratify CAD risk. Indeed, although the effect was less strong than that observed for either hyperlipidemia or HeFH, both CAD PRS (51) (fig. S7G) and functional scores (fig. S7H) predicted CAD risk (table S3). Likewise, the presence of a damaging LDLR variant conferred added risk even among participants with a high-risk PRS (Fig. 5, E; table S4).
Together, our results support previous suggestions (52–55) that: phenotypes related to LDL-C concentrations can be markedly affected by the aggregated impact of many common variants or by single LDLR variants of large effect; aggregated common and single rare variants can convey quantitatively similar levels of risk (Fig. 5F); and common and rare variants can predict higher risks together than can either alone.
To assess whether these findings were reproducible where demographics and ancestry are more diverse, we extended our analysis to the All of Us cohort. Although the number of participants with both sequenced genomes and LDL-C measurements (~145k) was 3-fold lower than that of the UK Biobank, reducing statistical power, we reached similar conclusions with respect to hyperlipidemia and HeFH risk (fig. S8, A to H; table S2 to S4). Functional scores correlated with the frequency of hyperlipidemia (ρ = −0.89; P = 0.02; fig. S8A; see also fig. S8B); the use of cholesterol-lowering medication (ρ = −0.94; P = 0.005; fig. S8C); and the age at hyperlipidemia diagnosis (ICD-10 E78.0) (ρ = 0.99; P < 0.001; fig. S8D). Carriers of variants with a damaging score again showed a high risk of HeFH (OR 12.2; 95% CI 6.3 to 23.9), comparable to carrying an annotated pathogenic variant (OR 16.6; 95%CI 12.2 to 22.7) (fig. S8E; table S2), and HeFH risk was increased more than 2-fold (OR 38.2; 95% CI 19.1 to 76.6) among those carrying a pathogenic variant that received a damaging functional score. We again found that both (population-appropriate) PRS (56) and LDLR functional scores could each quantitatively predict hyperlipidemia risk (fig. S8, F and G; table S3) and that the combination of this common and rare variant evidence yielded stronger risk predictions than did either evidence type alone (fig. S8H; table S4). Participants carrying a damaging LDLR variant had a more than 2-fold higher risk of having hyperlipidemia if they had a PRS in the highest quartile (Adj. OR 8.6, 95% CI 3.8 to 19.6) compared to those in the lowest quartile (Adj. OR 3.8, 95% CI 1.2 to 12.7). Thus, our findings demonstrate the potential of our LDLR functional scores to improve risk inference for HeFH, even in populations with diverse ethnicity and ancestry.
Discussion
A confirmed genetic diagnosis of HeFH can have a profound impact on patient care and outcomes. Unfortunately, HeFH is underdiagnosed and undertreated (57–59) in large part due to our current inability to confidently identify the pathogenicity of most missense variants in LDLR. This highlights the pressing need for functional evidence that can facilitate diagnosis and enable lifesaving early therapeutic intervention.
Here, we systematically assessed the functional impacts of LDLR variation on both cellular LDL uptake and LDLR surface abundance, providing comprehensive sequence–function maps of LDLR missense variant function. Beyond observing striking differences in the apparent importance of different LA modules for LDL internalization, a comparison of the function and abundance maps enables a systematic identification of variants impacting function independently from protein trafficking and abundance, promising to inform future studies. For example, we noted a subtle pattern of hypermorphic variants in the surface abundance map, but not the function map, at the site of the NPxY motif which is required for both clathrin-mediated uptake of the LDLR-LDL complex (38) and for ubiquitination of LDLR by the IDOL protein (60, 61). This observation could be explained by residues that escape IDOL-dependent degradation. Although these apparently-hypermorphic variants were observed in too few UK Biobank and All of Us participants to assess their effects on physiologic LDL-C levels, future cohorts could enable our resource to inform sequence-structure-function relationships related to hypermorphic variants and downstream clinical phenotypes.
Within the ACMG/AMP framework for classifying variant pathogenicity, we calibrated the likelihood of pathogenicity corresponding to each functional score. By providing evidence for over 7,000 variants (including ~400 currently annotated as VUS), we expect that this resource will contribute to the adjudication of variants with ambiguous classifications, including the as-yet-unseen rare variants that will continue to emerge. To further enable the translation of experimental measurements to the clinic, we used the prospective UK Biobank cohort to propose a threshold for classifying a “functional mutation” under the DLCN criteria for HeFH diagnosis. Beyond diagnosing HeFH for patients with a newly-discovered LDLR variant, each definitive variant classification enables cascade screening, with the potential to identify additional family members (~ 50% of 1st-degree and 25% of 2nd-degree relatives) for whom the risk of cardiovascular disease can then also be lowered (9).
We further showed that functional scores could be integrated with PRSs to further improve risk inference for hyperlipidemia. This is consistent with previous reports that well-validated PRS evaluation can explain variability in phenotypes and outcomes among carriers of the same LDLR variant and inform more personalized care (52–55). Although clinical guidelines for the integrated use of PRS in HeFH are lacking, a comprehensive assessment of genetic risk (from both common and rare variation) early in life offers the potential to stratify disease risk independently from clinical presentation and improve allocation of resources for increased follow-up, advanced screening, and second-line therapeutics for primary prevention (for example, PCSK9 inhibitors)—with the goal of limiting lifetime exposure to LDL cholesterol and reducing risk of CAD (62–64).
Extending this approach to the All of Us cohort, using a population-appropriate estimate of common variant risk (56), demonstrated the applicability of integrated risk prediction in more diverse cohorts for which new sources of evidence are critical. Indeed, participants in All of Us were ~ 1.3 times more likely than participants in the UK Biobank to harbor a variant that either lacked a definitive classification in ClinVar or had no clinical annotation whatsoever. The ancestry-agnostic evidence we provide here has the potential to ameliorate the currently disproportionately high number of VUSs reported in populations of non-European ancestry (65).
We acknowledge certain limitations of our study. First, because our assays relied on the expression of variant cDNAs, they excluded noncoding variants and missed some potential impacts of coding variants (for example, on splicing). Although most current VUS annotations are coding variants, large-scale assays could be further expanded, for instance, to capture the impacts of both coding and non-coding variants on splicing (66, 67). Indeed, large-scale studies of variant effects in the LDLR promoter have already been reported (68). Moreover, our cDNA maps complement smaller-scale variant assays based on mutagenesis at the endogenous LDLR locus (14), in that differences between these studies could point to non-coding impacts of coding variants.
Second, we note that some of the apparent impacts of substitutions in our surface abundance map may have arisen because they form the epitope for the antibody used in our study. This said, we could observe no regional patterns of impactful substitutions in the surface abundance map that were not also apparent in the function map, suggesting that the key epitope-forming residues are also important for LDL internalization.
Third, the cellular model (HeLa) and culture conditions used in our assay do not perfectly correspond to the (variable) physiologic conditions in which the LDL receptor functions. This may explain why many LA2 and LA6 variant impacts were observed in the context of excess VLDL. One model for these findings is that these modules are not required for VLDL binding, but are required for VLDL uptake, such that VLDL becomes a better competitive inhibitor in the presence of an LA2 or LA6 deficiency. Although we did not observe this phenomenon with the addition of other lipoprotein subtypes, further examination of in vitro environments (including a physiologic ratio of LDL to VLDL in the presence of other serum factors) would add nuance to our knowledge of LDLR functions. Although we did not provide calibrated variant classification evidence for variants in LA1, LA2, and LA6 (described above), future variant classification efforts may consider the mounting functional and biochemical evidence suggesting that LA modules appear heterogeneous in their function and potentially their importance for disease risk.
Moreover, we have assessed only two functions (albeit those most directly implicated in human disease) of LDLR, which may also operate through other mechanisms to influence atherosclerosis.
Estimates of individual-level risk are often limited by incomplete penetrance or variable expressivity of the causal variant (69, 70) due to polygenic, epigenetic, or gene-environment interactions. Indeed, LDL-C concentrations often vary between individuals harboring the same variants. Estimating variant-level heterogeneity is complicated by the rarity of most variants: indeed, more than half of LDLR coding variants in the UK Biobank and All of Us cohorts collectively were observed in two or fewer participants. Future scalable in vitro or even individual subject-level ex vivo assays may enable modeling of incomplete penetrance or variable expressivity (71, 72).
The resource of LDLR sequence–function maps we present here should improve personalized genomic medicine for HeFH—one of the three CDC Tier 1 conditions for which it is recommended that incidental findings be reported back to physicians and patients. Large-scale functional studies are already available for genes associated with the other two Tier 1 conditions: BRCA1 (73–78) and BRCA2 (79–81) for Hereditary Breast and Ovarian Cancer; and MSH2 (82) for Lynch Syndrome. More broadly, although large-scale variant effect data are currently available for less than 1% of the ~ 4000 human proteins currently associated with genetic diseases (83), the availability of such studies for cardiometabolic conditions is growing (84–90), with a roadmap now emerging towards a comprehensive atlas of variant effects (91–93).
Materials and methods
Ethics approval and consent to participate:
Consent to participate was obtained via the UK Biobank and All of Us projects, and datasets were analyzed in accordance with the associated data use agreements. The transfer of human data was approved and overseen by The UK Biobank Ethics Advisory Committee (Project ID: 51135). This study was performed in alignment with the ethical principles outlined in the All of Us Policy on the Ethical Conduct of Research.
In the All of Us Research Program, informed consent for all participants is conducted in person or through an eConsent platform that includes primary consent, Health Insurance Portability and Accountability Act authorization for research EHRs and consent for return of genomic results. The protocol was reviewed by the Institutional Review Board (IRB) of the All of Us Research Program. The All of Us Institutional Review Board follows the regulations and guidance of the National Institutes of Health Office for Human Research Protections for all studies.
Oligonucleotides
Oligonucleotide and gRNA sequences are listed in data S6.
Cell culture and transfection
HeLa cells were cultured in DMEM supplemented with 10% FBS and 1% penicillin-streptomycin (all from Wisent). Cells were passaged at 1:10 every 3 days. All transfections were performed using the Neon Electroporation system (ThermoFisher Scientific).
Plasmids
The Tet promoter of the dAAVS1-TetBxb1-BFP plasmid (26) was replaced with the EF1α promoter from PL-sin-EF1α-eGFP (Addgene #21320) using Gibson Assembly (94) to create the dAAVS1-EF1α-BFP plasmid. A Gateway-compatible attB integration construct (pDEST-Bxb1-Rec v2) was made by replacing the Kozak sequence and PTEN open reading frame of attB-PTEN-IRES-mCherry with the 1873bp Gateway attR1-CcdB-attR2 cassette from pDEST-AD-CYH2. The pDONR-LDLR-BC plasmid was created by PCR amplification of the LDLR ORF from pDONR-LDLR (human ORFeome V8.1 (95)) with primers LDLR-BC-F and LDLR-BC-R. The amplicon was combined with pDONR223 in a standard Gateway BP reaction. The pDONR-LDLR-BC plasmid was used as a template for the creation of the coding variant libraries (see below).
Cell line generation
The HeLa LP line was generated by transfecting HeLa T-Rex cells with equal proportions of dAAVS1-EF1α-BFP, AAVS1-TALEN-L and AAVS1-TALEN-R (Addgene #59025 and #59026) using JetPRIME reagent (Polyplus Sartorius) as previously described (26). Transfections were cultured for ~ two weeks and were enriched for BFP-positive cells using fluorescence-activated cell sorting (FACS), cloned (by FACS), and validated as previously described (26). The HeLa LP line was used to generate the HeLa LP-LDLRKO line (see below).
The HeLa LP-LDLRKO line was generated using CRISPR-Cas9 genome editing (96) (PX459 V2.0; Addgene plasmid # 62988) targeting exon 3 of LDLR (guide sequence GGAAATGCATCTCCTACAAG from the Toronto KnockOut v3 library (97); LDLR-6-F and LDLR-6-R). HeLa LP cells were transfected and treated one day post-transfection with 200 μg/mL puromycin for two days prior to cloning by FACS. Expanded clones were assessed by immunoblotting, by targeted sequencing, and based on their diminished ability to uptake fluor-conjugated LDL and reduced cell-surface expression of LDLR (fig. S9, A and B).
Library generation and integration
Variant libraries were generated using an updated version of the POPCode method such that flanking 5’ and 3’ sequence tags are incorporated to enable selective PCR amplification of the mutagenized strand (98). Here, a single oligonucleotide containing a central NNK degeneracy was designed for each of the 860 residues in LDLR as described (99) (Eurofins Genomics), resulting in random codon substitutions at each position. Mutagenesis was focused in turn on each of five regions of the LDLR cDNA (ranging in lengths from 147 to 189 residues), using corresponding oligonucleotide pools. Oligonucleotides were pooled for each region. Each pool was phosphorylated as follows: 1× PNK buffer; 1 mM ATP; 300 pmol of oligonucleotide pool; 10 U Polynucleotide Kinase (NEB); 50 μl total volume; 30:00 at 37°C. Mutagenized DNA libraries were generated as follows: 1) 90 ng pDONR-LDLR-BC; 1.5 μM M13ext_5’Tag2_F oligo; 1.5 μM phosphorylated 3’Tag_Hyb oligo (for region 5 only); 9.9 μM oligo pool; 20 μl total volume; 95°C for 3:00; cool to 25°C at a rate of 0.1°C/s. 2) 5 μl of reaction (1) was combined with 5 μl 2X Phusion Mastermix (NEB); 2:00:00 at 50°C. 3) To reaction (2) add 1.5 μl 10X T4 DNA Ligation buffer; 200 U T4 DNA Ligase (NEB); 3 μl H2O; 45°C for 20:00. 4) To enrich for newly-synthesized variant DNA: 1 μl reaction (3); 1× Phusion Mastermix (NEB); 0.4 μM M13Rext (for single tag POPCode), or 0.4 μM 5’tag2_F1 oligo and 0.4 μM 3’TagR_amp (for double tag POPCode); 50 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 65°C for 0:30; c) 72°C for 2:30; 20 cycles of a-c; 72°C for 5:00. Products were electrophoresed on a 1% agarose gel and purified (Mackery Nagel #740609.50). Gateway B1 and B2 recombination sites and (25-mer) degenerate barcodes were added as follows: 1 μl of purified product from (4); 1× Phusion Mastermix (NEB), 0.4 μM each of oligos POP_BC_F1 and POP_BC_R1; 50 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 56°C for 0:30; c) 72°C for 2:30; 5 cycles of a-c; d) 98°C for 0:15; e) 72°C for 2:30; 15 cycles of d-e; 72°C for 7:00. Products were again electrophoresed on a 1% agarose gel and purified.
Variant pools were Gateway-cloned into pDONR223, and plasmid DNA was extracted from ~200k pooled bacterial transformants (Qiagen #12362) for long-read sequencing analysis (see below). Post-QC pDONR variant libraries were Gateway-cloned into pDEST-Bxb1-Rec v2. Plasmid DNA was extracted from ~40k transformants and used for transfections.
PacBio library construction, sequencing, and analysis
Sequencing libraries for pDONR LDLR variant pools were generated either using a PCR-amplified linear template (regions 1 and 2) or using a PCR-free approach (regions 3, 4, and 5). Linear templates for regions 1 and 2 were generated by PCR as follows: 1× Phusion Mastermix (NEB); 10 ng LDLR POPCode plasmid DNA; 0.4 μM M13F and 0.4 μM M13R; 50 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 57°C for 0:30; c) 72°C for 3:00; 16 cycles of a-c; 72°C for 7:00. Products were run on a 1% agarose gel and purified (Mackery Nagel #740609.50). Linear products were immediately purified using PacBio AMPure beads following the manufacturer’s recommended protocol and libraries were generated using the PacBio SMRTbell Prep Kit V2.0. Linear template for regions 3, 4, and 5 were generated by restriction enzyme digest of plasmid DNA as follows: 15 μg plasmid DNA: 1× Cutsmart buffer: 30 U PsiI and 30 U SapI: 100 μl total volume; 37°C for 3:00:00. The ~2.7kb band was gel purified as above and further purified using the PacBio AMPure beads. Libraries were generated using the SMRTbell Prep kit V3.0.
Libraries were sequenced on the PacBio Sequel IIe at The Centre for Applied Genomics, Toronto Canada. PacBio sequencing data was post-processed using DeepConsensus (100) and filtered to exclude sequencing reads with RQ < 0.998. Consensus barcode-genotype associations were then calculated using Pacybara (101). Briefly, Pacybara groups and aligns reads based on barcode and genotype similarity while accounting for sequencing error, while allowing for the detection and automatic exclusion of non-unique barcodes and chimeric reads.
Integration and selection of variant libraries
On day 1, cells were transfected (1) by electroporation with pCAG-NLS-HA-Bxb1 (Addgene #51271) and incubated for 2 days in standard conditions. On day 3, 8 M transfected cells (from 1) were trypsinized and again transfected (2) by electroporation with pDEST-Bxb1-Rec v2 variant libraries and incubated for 5–7 days. Integration efficiencies were evaluated by flow cytometry. Cells containing an integrated variant (based on loss of BFP) were selected by FACS and were expanded for 2–4 days.
LDL uptake and surface abundance assays
Approximately 2 days prior to selection, variant-expressing cell libraries (described above) were seeded to 30% confluency. One day prior to selection, cells were serum-deprived by replacing standard media (described above) with DMEM supplemented with 1% penicillin-streptomycin and 5% lipoprotein-depleted FBS (Kalen Biomedical).
To measure LDL uptake, media was supplemented with 5 μg/mL pHrodo Green-LDL (ThermoFisher Scientific; L34355) for 1 hour before isolating the top 20th percentile of cells by FACS based on pHrodo Green fluorescence intensity (see fig. S10 and S11 for example).
To measure LDL uptake in the presence of very low-density lipoprotein (VLDL), media was again supplemented with pHrodo Green-LDL, but also with a stoichiometric excess (1:4) of exogenous (unlabelled) VLDL (Athens Research & Technology; 12–16-221204). To assess the effects of all lipoprotein subtypes on pHrodo Green-LDL uptake by LDLR, we also tested the impact of media supplemented with one of: intermediate-density lipoprotein (IDL); unlabeled LDL; and chylomicrons (Athens Research & Technology; 12–16-090412, 12–16-120412, and 12–16-030825, respectively).
To measure LDLR cell-surface abundance, 20 M cells were incubated with 16 μL of anti-human LDLR antibody (MAB2148) for 30 minutes at room temperature. Cells were washed in 4°C PBS containing 10% FBS and incubated with 16 μL of secondary antibody (Goat anti-Mouse IgG Alexa Fluor 488; Abcam #ab150113) for 30 minutes at room temperature. Cells were again washed in 4°C PBS, and the top 25th percentile of cells was isolated by FACS based on Alexa Fluor 488 fluorescence intensity.
In each assay, sorted cells and unsorted controls were seeded and expanded to confluency prior to collection for gDNA extraction (Qiagen # 13343). Genomic DNA was used as template to amplify integrated cassettes for Illumina library preparation (described below).
Post-selection library preparation, sequencing, and analysis
The Bxb1-integrated LDLR cassette was amplified by PCR using oligonucleotides specific to genomically-integrated constructs. For each condition 16 reactions of the following reactions were generated as follows: 1 μg genomic DNA; 1× Q5 Mastermix (NEB); 0.4 μM Bxb-Amp-EF1α-F8new and 0.4 μM Bxb-Amp-R1; 50 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 56°C for 0:30; c) 72°C for 2:30; 25 cycles of a-c; 72°C for 7:00. Products from the above 16 reactions were pooled and run on a 1% agarose gel and purified (Mackery Nagel #740609.50).
From each sample, 5’ and 3’ barcodes were individually amplified using oligos flanked by Illumina adapter sequences, then further amplified with primers adding Illumina indices and grafting sequences. The reactions were as follows: 10 ng cassette DNA, 1× Phusion Mastermix, 0.4 μM oligonucleotides; 50 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 58°C for 0:30; c) 72°C for 0:30; 25 cycles of a-c; 72°C for 5:00. Products were diluted 1:10 and 1 μl was used as a template in the following reaction: 1× Phusion Mastermix, 0.4 μM oligonucleotides; 20 μl total volume. PCR conditions were as follows: 98°C for 0:30; a) 98°C for 0:15; b) 65°C for 0:30; c) 72°C for 2:00; 7 cycles of a-c; 72°C for 5:00. Products were gel-purified (Mackery Nagel #740609.50), quantified, and sequenced on an Illumina NextSeq 500 as described (102).
Barcode sequences were extracted from sequencing reads and clustered and tallied using Bartender (103). Barcode-genotype associations were obtained from the Pacybara output as above. Pre- and post-selection frequencies were calculated by dividing counts by sequencing depth and determining the mean and standard deviation of frequencies across four technical replicates for each barcode. Log ratios of post- to pre-selection frequency were calculated for each barcode in each library. For each unique variant (codon change), we averaged the log ratios for all barcodes corresponding to single-mutant clones carrying that variant, and estimated measurement error in the combined value using the delta method to propagate errors from error estimates of the input values. Where no single mutant clones were available, we tested whether > 1 multi-mutant clones containing the variant of interest existed, and if so, averaged across their log ratios and similarly estimated error. Next, for each unique amino-acid change, we averaged codon-level scores and similarly estimated error. Finally, scores were re-scaled such that the median nonsense variant score was 0 and the median synonymous variant score was 1. The above procedure was performed separately for upstream and downstream barcodes and the results merged.
Biochemical analysis
Conservation measures for LDLR were obtained from ConSurf (27) (PDB 1N7D). Stability measures (ddG) were derived using MutateX (104), which estimates variant impact on protein stability using FoldX (28).
Functional scores for LDLR were mapped to structures of LDLR using the TileSeqMave v1.0 package (https://github.com/rothlab/tileseqMave). 3D structures were obtained for the extracellular domain of LDLR (1N7D) (5) and the ligand-binding domain of LDLR bound to ApoB100 (9BDE) (36). PDBs colored by functional score (scaled as in Fig. 1C) are provided in data S7 (1N7D) and data S8 (9BDE).
Selection of clinical variants and evidence calibration
We obtained clinically annotated LDLR variants from ClinVar, restricting to those with at least a 1-star review status, considering variants annotated as either pathogenic or likely pathogenic variants to be pathogenic and variants annotated as either benign or likely benign to be benign.
We calibrated the evidentiary value of our functional scores for clinical variant interpretation as previously described (43, 44). We note that this calibration was restricted to missense variants only. Briefly, we separately estimated the probability distributions of both pathogenic and benign missense variants using kernel density estimation and used these to obtain, for any given functional score, the log-likelihood ratio of pathogenicity (LLR) for a missense variant with this score. We do not report evidence calibrations for variants in LA modules 1, 2, and 6 for the function map or for modules 1–6 for the surface-abundance map. Each LLR represents a quantitative measure of evidence strength that can be applied within a Bayesian framework or converted to the corresponding cardinal level of ACMG/AMP evidence strength using a modified version of previous approaches (41, 42). In assigning evidence strength, we enforced monotonicity for the lowest and highest scores: We did not assert evidence towards pathogenicity where a lower score had received evidence towards benignity, and vice versa. We further note that the conversion to ACMG/AMP evidence strength assumes a prior probability that is based on a clinically observed missense variant in a person suspected of a relevant genetic disease, and that this procedure should not otherwise be applied to LDLR variants. Finally, we note that this calibration procedure has not been endorsed by a clinical authority, and that guidance for clinical use of these data is under active consideration by the ClinGen Familial Hypercholesterolemia Variant Curation Expert Panel (on which DRT, JKW, RAH and FPR serve as members).
Processing of population-level cohorts
Analysis in the UK Biobank cohort was conducted with whole-exome sequencing data from the final exome release (469,779 participants, UK Biobank field ID 23157). Data was extracted and processed as previously described (105). Briefly, coding variants were extracted from OQFE VCF files (field ID 23157), and filtering was adapted from the UK Biobank (106). Variants were mapped to the canonical LDLR transcript (ENST00000558518). Data transfer was approved and overseen by the UK Biobank Ethics Advisory Committee (project ID 51135). We excluded participants who had withdrawn from the UK Biobank study from our analysis. Participant phenotypes and ICD codes were extracted from the UK Biobank cohort (discussed below). Where participants had multiple measurements from repeat assessments, measurements from the first assessment were retained. Where participants had multiple LDLR variants, the lowest-scoring variant was retained.
Analysis of the All of Us cohort was conducted with short-read whole-genome sequencing data from the All of Us Controlled Tier Dataset v8 (414,349 participants). Post-sequencing variant and sample QC were performed by the All of Us Data and Research Center. Data were extracted and processed as previously described (105). Briefly, coding variants were extracted from Hail MatrixTables (version 0.2.130, Hail Team) and filtered for Phred quality score > 20, individual missingness < 10%, minimum read coverage depth of 7, and presence in at least one participant passed the allele balance threshold of 0.20. Where participants had > 1 LDLR variant, only the lowest-scoring variant was retained.
Polygenic risk scores for LDL-C and CAD
For UK Biobank participants, polygenic risk scores for LDL [Standard PRS for low-density lipoprotein cholesterol (LDL_SF); UK Biobank field ID 26250] and CAD [Standard PRS for coronary artery disease (CAD); UK Biobank field ID 26227] were obtained from the UK Biobank (51). To further ensure that risk assessment using the PRS developed for UKB participants is independent of risk assessment using LDLR functional scores, we only considered rare (MAF ≤ 0.01) variation in LDLR that is unlikely to have contributed to the PRS.
For All of Us participants, polygenic risk scores for hyperlipidemia were calculated for each participant using available weights (56). Polygenic risk scores were calibrated to participant ancestry as previously described (56) (fig. S12).
Phenotype selection
Cholesterol-lowering medication use in the UK Biobank was ascertained as previously described (107) and used to correct participant LDL-C measurements where appropriate (UK Biobank field ID 30780). Clinical diagnosis of hyperlipidemia was determined based on the ICD-10 (UK Biobank field ID 41270) code E78.0; similarly, age at diagnosis was obtained from the date of record entry. In All of Us, cholesterol-lowering medication use was ascertained as in the UK Biobank; participant LDL-C was divided by 0.7 to correct for medication use (7). Aortic stenosis was ascertained from ICD-10 codes I35.0 and I35.2 as previously described (108). Carotid intima-media thickness of the carotid artery was taken as the average across four readings [120° and 150° for the right carotid artery (UK Biobank field IDs 22671 and 22674, respectively); and 210° and 240° for the left carotid artery (UK Biobank field IDs 22677 and 22680) ] as previously described (109).
Coronary artery disease was ascertained based on hospital records (ICD-10) (UK Biobank field ID 41270), death registry (UK Biobank field ID 40001), and operation and procedure codes (OPCS4) (UK Biobank field ID 41272) similarly to previous work (53). Peripheral artery disease (PVD) status was assigned based on hospital records. The age at which participants had a first occurrence of CAD or PVD was determined based on the date of the first recorded entry. ICD-10 and OPCS4 codes for CAD and PVD are listed in table S5.
To identify participants with suspected heterozygous familial hypercholesterolemia (HeFH) in the UK Biobank, we applied diagnostic criteria from the Dutch Lipid Clinic Network (DLCN) based on available biomarkers and electronic records (45–48). We adapted these criteria to consider evidence from family history, clinical history, clinical presentations, and LDL-C levels. Family history and clinical presentations were ascertained from ICD codes. Clinical history considered the occurrence of CAD and PVD (determined as above). In the UK Biobank, the pre-defined health outcome for myocardial infarction (UK Biobank field ID 42000) was also considered. Participant LDL-C levels were corrected for medication use as above. The criteria and points ascribed to each finding are outlined in table S6. Participants were broadly assigned HeFH if they received ≥ 6 points (corresponding to a diagnosis of at least “probable” HeFH under the DLCN criteria). Towards defining a score threshold for a “functional mutation” under the DLCN, we defined two groups: those receiving > 8 points (corresponding to a “definite” diagnosis) without attributing points for a “functional mutation”; and variant carriers receiving < 3 points (corresponding to “unlikely” HeFH), again without attributing points for a “functional mutation”. Our adapted framework differs from the DLCN in that, to avoid circularity, we do not assign points for variants in any of APOB, LDLR, or PCSK9 and thus our estimates of HeFH status are conservative.
Phenotypes in the All of Us cohort were selected to match those from the UK Biobank set. Where participants had multiple measurements from repeat assessments, only measurements from the initial assessment were retained. For each quantitative trait, units were harmonized, and non-physiologic values were removed.
Statistical analysis
To assess the ability of functional scores to distinguish between participants with and without HeFH under DLCN criteria (described above), we calculated a receiver operating characteristic curve (Python Scikit roc_curve, auc) and derived a score threshold corresponding to a 5% false positive rate.
To assess the correlation of functional scores with hyperlipidemia phenotypes, the probability (or mean where applicable) was assessed within functional score windows (0.05 +/− 0.025) that included at least 5 participants. In All of Us, correlation was similarly assessed, but because All of Us prohibits the publication of data representing fewer than 20 participants, we evaluated wider windows (0.3 +/− 0.15) that included at least 20 participants.
To estimate the effect of participant genotype on the risk of HeFH, participants were stratified by damaging (≤ 15%ile) or normal (≥ 70th %ile) functional score and by clinical variant classification (VUS or P/LP). Odds ratios were calculated relative to noncarriers.
To assess participant risk of hyperlipidemia and CAD, the odds ratio of each was calculated by logistic regression (Python Statsmodels Logit) correcting for age, sex, body mass index (BMI) and self-reported ethnicity. To assess the effects of: i) functional score, participants were stratified by normal or damaging LDLR variant (as above), with noncarriers as reference; ii) polygenic risk score, participants were stratified by top and bottom quartile of PRS, with intermediate quartiles as reference; and iii) both functional score and polygenic risk, participants in the top and bottom quartiles of PRS were further stratified whether they carried a normal or damaging LDLR variant, with noncarriers in the intermediate quartile as reference.
The time-to-event for hyperlipidemia was estimated by Kaplan-Meier survival analysis (Python Lifelines KaplanMeierFitter). We considered the age of diagnosis for hyperlipidemia (ICD-10 E78.0) and CAD (as described above) from participant electronic health records. Observations were censored based on participant age at the last known record entry.
Supplementary Material
Acknowledgements:
We are grateful to Michèle S Palkovits for her thoughtful contributions to this work throughout its many stages. We thank Annie Bang and Michael Parsons from the LTRI Flow Cytometry facility for their help with experimental design and implementation. We also thank the many participants of the UK Biobank and All of Us cohorts, as well as those who envisioned, developed, and continue to support these exceptional resources. Access to the UK Biobank was made available under a material transfer agreement with Sinai Health Systems. Access to All of Us data occurred solely within the secure Researcher Workbench under Vanderbilt University Medical Center’s Data Use and Registration Agreement; no biospecimens or data were transferred.
Funding:
R01HL164675 (FPR, DMR, EAA, CAM, VNP, AMG, BMK).
The One Brave Idea Initiative (jointly funded by the American Heart Association, Verily Life Sciences LLC, and Astra-Zeneca, Inc. with additional support from Quest Diagnostics) (CAM, FPR).
Canada Foundation for Innovation (FPR).
Canadian Institutes of Health Research Foundation Grant FDN159926 (FPR).
The Canada Excellence Research Chairs Program (FPR).
University of Toronto Precision Medicine Initiative (PRiME) Fellowship PRMF2022 (DRT).
National Human Genome Research Institute of the National Institutes of Health Center of Excellence in Genomic Science Initiative RM1HG010461 (DMF, FPR).
NHGRI Impact of Genomic Variation on Function Initiative UM1HG011989 (FPR).
EU Horizon RIA: 101155885–2, FH-EARLY (MB, SGP).
La Caixa Foundation (LCF/PR/HP23/52330032) (MB, SGP).
NIH from R01DK106236, P30 DK116074 (to the Stanford Diabetes Research Center), R01DK116750, R01DK137889, R01DK120565 (JWK), and NHGRI K08HG014001 (MCL).
The All of Us Research Program is supported by the National Institutes of Health, Office of the Director: Regional Medical Centers: 1 OT2 OD026549; 1 OT2 OD026554; 1 OT2 OD026557; 1 OT2 OD026556; 1 OT2 OD026550; 1 OT2 OD 026552; 1 OT2 OD026553; 1 OT2 OD026548; 1 OT2 OD026551; 1 OT2 OD026555; IAA #: AOD 16037; Federally Qualified Health Centers: HHSN 263201600085U; Data and Research Center: 5 U2C OD023196; Biobank: 1 U24 OD023121; The Participant Center: U24 OD023176; Participant Technology Systems Center: 1 U24 OD023163; Communications and Engagement: 3 OT2 OD023205; 3 OT2 OD023206; and Community Partners: 1 OT2 OD025277; 3 OT2 OD025315; 1 OT2 OD025337; 1 OT2 OD025276.
Footnotes
Competing interests:
SGP: MONCYTE Health (founder, chief scientific officer, stock).
VNP: Lexeo Therapeutics (SAB); BioMarin (consultant, sponsored research); Constantiam Biosciences (clinical advisor); Borealis (consultant).
EAA: Personalis (founder); Deepcell (founder); Svexa (founder); Candela (founder); Parameter Health (founder); Saturnus Bio (founder); SequenceBio (advisor); Foresite Labs (advisor); Pacific Biosciences (advisor); Versant Ventures (advisor); AstraZeneca (non-executive director); Svexa (non-executive director); Pacific Biosciences (stock); Personalis (stock); AstraZeneca (stock); Illumina (collaborative support in kind); Pacific Biosciences (collaborative support in kind); Oxford Nanopore (collaborative support in kind); Cache DNA (collaborative support in kind); Cellsonics (collaborative support in kind).
JWK: Chief Research Advisor for the Family Heart Foundation (no financial interest); Mammoth Biosciences (consultant); Wave Life Sciences (consultant).
ETC: Helix (employee).
CAM: Atman Health (founder and shareholder), Tanaist (founder and shareholder), TMA Precision Health (scientific advisor and shareholder), Everyone Medicines (scientific advisor), LifeMD (non-executive director and shareholder).
FPR: Constantiam Biosciences (scientific advisor and shareholder); Ranomics, Inc. (shareholder); SeqWell, Inc. (shareholder).
All other authors declare that they have no competing interests.
Data and materials availability:
All data are available in the main text or the supplementary materials. Correspondence and requests for materials should be addressed to F.P.R.
References and Notes:
- 1.Brunham LR, Ruel I, Aljenedil S, Rivière J-B, Baass A, Tu JV, Mancini GBJ, Raggi P, Gupta M, Couture P, Pearson GJ, Bergeron J, Francis GA, McCrindle BW, Morrison K, St-Pierre J, Henderson M, Hegele RA, Genest J, Goguen J, Gaudet D, Paré G, Romney J, Ransom T, Bernard S, Katz P, Joy TR, Bewick D, Brophy J, Canadian Cardiovascular Society position statement on familial hypercholesterolemia: Update 2018. Can. J. Cardiol. 34, 1553–1563 (2018). [DOI] [PubMed] [Google Scholar]
- 2.Gidding SS, Champagne MA, de Ferranti SD, Defesche J, Ito MK, Knowles JW, McCrindle B, Raal F, Rader D, Santos RD, Lopes-Virella M, Watts GF, Wierzbicki AS, American Heart Association Atherosclerosis, Hypertension, and Obesity in Young Committee of Council on Cardiovascular Disease in Young, Council on Cardiovascular and Stroke Nursing, Council on Functional Genomics and Translational Biology, and Council on Lifestyle and Cardiometabolic Health, The Agenda for Familial Hypercholesterolemia: A Scientific Statement From the American Heart Association. Circulation 132, 2167–2192 (2015). [DOI] [PubMed] [Google Scholar]
- 3.Brown MS, Goldstein JL, Receptor-mediated control of cholesterol metabolism. Science 191, 150–154 (1976). [DOI] [PubMed] [Google Scholar]
- 4.Brown MS, Goldstein JL, A receptor-mediated pathway for cholesterol homeostasis. Science 232, 34–47 (1986). [DOI] [PubMed] [Google Scholar]
- 5.Rudenko G, Henry L, Henderson K, Ichtchenko K, Brown MS, Goldstein JL, Deisenhofer J, Structure of the LDL receptor extracellular domain at endosomal pH. Science 298, 2353–2358 (2002). [DOI] [PubMed] [Google Scholar]
- 6.Sturm AC, Knowles JW, Gidding SS, Ahmad ZS, Ahmed CD, Ballantyne CM, Baum SJ, Bourbon M, Carrié A, Cuchel M, de Ferranti SD, Defesche JC, Freiberger T, Hershberger RE, Hovingh GK, Karayan L, Kastelein JJP, Kindt I, Lane SR, Leigh SE, Linton MF, Mata P, Neal WA, Nordestgaard BG, Santos RD, Harada-Shiba M, Sijbrands EJ, Stitziel NO, Yamashita S, Wilemon KA, Ledbetter DH, Rader DJ, Convened by the Familial Hypercholesterolemia Foundation, Clinical genetic testing for familial hypercholesterolemia: JACC scientific expert panel. J. Am. Coll. Cardiol. 72, 662–680 (2018). [DOI] [PubMed] [Google Scholar]
- 7.Khera AV, Won H-H, Peloso GM, Lawson KS, Bartz TM, Deng X, van Leeuwen EM, Natarajan P, Emdin CA, Bick AG, Morrison AC, Brody JA, Gupta N, Nomura A, Kessler T, Duga S, Bis JC, van Duijn CM, Cupples LA, Psaty B, Rader DJ, Danesh J, Schunkert H, McPherson R, Farrall M, Watkins H, Lander E, Wilson JG, Correa A, Boerwinkle E, Merlini PA, Ardissino D, Saleheen D, Gabriel S, Kathiresan S, Diagnostic Yield and Clinical Utility of Sequencing Familial Hypercholesterolemia Genes in Patients With Severe Hypercholesterolemia. J. Am. Coll. Cardiol. 67, 2578–2589 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Clarke SL, Tcheandjieu C, Hilliard AT, Lee KM, Lynch J, Chang K-M, Miller D, Knowles JW, O’Donnell C, Tsao PS, Rader DJ, Wilson PW, Sun YV, Gaziano JM, Assimes TL, VA Million Veteran Program, Coronary artery disease risk of familial hypercholesterolemia genetic variants independent of clinically observed longitudinal cholesterol exposure. Circ. Genom. Precis. Med. 15, e003501 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Umans-Eckenhausen MAW, Defesche JC, Sijbrands EJG, Scheerder RL, Kastelein JJP, Review of first 5 years of screening for familial hypercholesterolaemia in the Netherlands. Lancet 357, 165–168 (2001). [DOI] [PubMed] [Google Scholar]
- 10.Knowles JW, Rader DJ, Khoury MJ, Cascade screening for familial hypercholesterolemia and the use of genetic testing. JAMA 318, 381 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Centers for Disease Control and Prevention (CDC), Tier 1 Genomics Applications and their Importance to Public Health (2014). https://archive.cdc.gov/www_cdc_gov/genomics/implementation/toolkit/tier1.htm.
- 12.Iacocca MA, Chora JR, Carrié A, Freiberger T, Leigh SE, Defesche JC, Kurtz CL, DiStefano MT, Santos RD, Humphries SE, Mata P, Jannes CE, Hooper AJ, Wilemon KA, Benlian P, O’Connor R, Garcia J, Wand H, Tichy L, Sijbrands EJ, Hegele RA, Bourbon M, Knowles JW, ClinGen FH Variant Curation Expert Panel, ClinVar database of global familial hypercholesterolemia-associated DNA variants. Hum. Mutat. 39, 1631–1640 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Richards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, Grody WW, Hegde M, Lyon E, Spector E, Voelkerding K, Rehm HL, Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine 17, 405–423 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ryu J, Barkal S, Yu T, Jankowiak M, Zhou Y, Francoeur M, Phan QV, Li Z, Tognon M, Brown L, Love MI, Bhat V, Lettre G, Ascher DB, Cassa CA, Sherwood RI, Pinello L, Joint genotypic and phenotypic outcome modeling improves base editing variant effect quantification. Nat. Genet. 56, 925–937 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Islam MM, Tamlander M, Hlushchenko I, Ripatti S, Pfisterer SG, Large-scale functional characterization of low-density lipoprotein receptor gene variants improves risk assessment in cardiovascular disease. JACC Basic Transl. Sci, doi: 10.1016/j.jacbts.2024.10.006 (2024). [DOI] [Google Scholar]
- 16.Thormaehlen AS, Schuberth C, Won H-H, Blattmann P, Joggerst-Thomalla B, Theiss S, Asselta R, Duga S, Merlini PA, Ardissino D, Lander ES, Gabriel S, Rader DJ, Peloso GM, Pepperkok R, Kathiresan S, Runz H, Systematic cell-based phenotyping of missense alleles empowers rare variant association studies: a case for LDLR and myocardial infarction. PLoS Genet. 11, e1004855 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Graça R, Alves AC, Zimon M, Pepperkok R, Bourbon M, Functional profiling of LDLR variants: Important evidence for variant classification: Functional profiling of LDLR variants. J. Clin. Lipidol. 16, 516–524 (2022). [DOI] [PubMed] [Google Scholar]
- 18.Graça R, Zimon M, Alves AC, Pepperkok R, Bourbon M, High-throughput microscopy characterization of rare LDLR variants. JACC Basic Transl. Sci. 8, 1010–1021 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Alves AC, Azevedo S, Benito-Vicente A, Graça R, Galicia-Garcia U, Barros P, Jordan P, Martin C, Bourbon M, LDLR variants functional characterization: Contribution to variant classification. Atherosclerosis 329, 14–21 (2021). [DOI] [PubMed] [Google Scholar]
- 20.Weile J, Roth FP, Multiplexed assays of variant effects contribute to a growing genotype–phenotype atlas. Hum. Genet. 137, 665–678 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Shirts BH, Pritchard CC, Walsh T, Family-specific variants and the limits of human genetics. Trends Mol. Med. 22, 925–934 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Tabet D, Parikh V, Mali P, Roth FP, Claussnitzer M, Scalable Functional Assays for the Interpretation of Human Genetic Variation. Annu. Rev. Genet. 56, 441–465 (2022). [DOI] [PubMed] [Google Scholar]
- 23.Defesche JC, Gidding SS, Harada-Shiba M, Hegele RA, Santos RD, Wierzbicki AS, Familial hypercholesterolaemia. Nat Rev Dis Primers 3, 17093 (2017). [DOI] [PubMed] [Google Scholar]
- 24.Ritter P, Yousefi K, Ramirez J, Dykxhoorn DM, Mendez AJ, Shehadeh LA, LDL cholesterol uptake assay using live cell imaging analysis with cell health monitoring. J. Vis. Exp, doi: 10.3791/58564-v (2018). [DOI] [Google Scholar]
- 25.Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, Olsson I, Edlund K, Lundberg E, Navani S, Szigyarto CA-K, Odeberg J, Djureinovic D, Takanen JO, Hober S, Alm T, Edqvist P-H, Berling H, Tegel H, Mulder J, Rockberg J, Nilsson P, Schwenk JM, Hamsten M, von Feilitzen K, Forsberg M, Persson L, Johansson F, Zwahlen M, von Heijne G, Nielsen J, Pontén F, Proteomics. Tissue-based map of the human proteome. Science 347, 1260419 (2015). [DOI] [PubMed] [Google Scholar]
- 26.Matreyek KA, Stephany JJ, Fowler DM, A platform for functional assessment of large variant libraries in mammalian cells. Nucleic Acids Res. 45, e102 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Yariv B, Yariv E, Kessel A, Masrati G, Chorin AB, Martz E, Mayrose I, Pupko T, Ben-Tal N, Using evolutionary data to make sense of macromolecules with a “face-lifted” ConSurf. Protein Sci. 32, e4582 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Schymkowitz J, Borg J, Stricher F, Nys R, Rousseau F, Serrano L, The FoldX web server: an online force field. Nucleic Acids Res. 33, W382–8 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Cheng J, Novati G, Pan J, Bycroft C, Žemgulytė A, Applebaum T, Pritzel A, Wong LH, Zielinski M, Sargeant T, Schneider RG, Senior AW, Jumper J, Hassabis D, Kohli P, Avsec Ž, Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492 (2023). [DOI] [PubMed] [Google Scholar]
- 30.Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, Motyer A, Vukcevic D, Delaneau O, O’Connell J, Cortes A, Welsh S, Young A, Effingham M, McVean G, Leslie S, Allen N, Donnelly P, Marchini J, The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.All of Us Research Program Investigators, Denny JC, Rutter JL, Goldstein DB, Philippakis A, Smoller JW, Jenkins G, Dishman E, The “All of Us” Research Program. N. Engl. J. Med. 381, 668–676 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Blacklow SC, Versatility in ligand recognition by LDL receptor family proteins: advances and frontiers. Curr. Opin. Struct. Biol. 17, 419–426 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Jeon H, Blacklow SC, Structure and physiologic function of the low-density lipoprotein receptor. Annu. Rev. Biochem. 74, 535–562 (2005). [DOI] [PubMed] [Google Scholar]
- 34.Esser V, Limbird LE, Brown MS, Goldstein JL, Russell DW, Mutational analysis of the ligand binding domain of the low density lipoprotein receptor. J. Biol. Chem. 263, 13282–13290 (1988). [PubMed] [Google Scholar]
- 35.Russell DW, Brown MS, Goldstein JL, Different combinations of cysteine-rich repeats mediate binding of low density lipoprotein receptor to two different proteins. J. Biol. Chem. 264, 21682–21688 (1989). [PubMed] [Google Scholar]
- 36.Reimund M, Dearborn AD, Graziano G, Lei H, Ciancone AM, Kumar A, Holewinski R, Neufeld EB, O’Reilly FJ, Remaley AT, Marcotrigiano J, Structure of apolipoprotein B100 bound to the low-density lipoprotein receptor. Nature 638, 829–835 (2025). [DOI] [PubMed] [Google Scholar]
- 37.Yamamoto T, Ryan RO, Domain swapping reveals that low density lipoprotein (LDL) type A repeat order affects ligand binding to the LDL receptor. J. Biol. Chem. 284, 13396–13400 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Chen WJ, Goldstein JL, Brown MS, NPXY, a sequence often found in cytoplasmic tails, is required for coated pit-mediated internalization of the low density lipoprotein receptor. J. Biol. Chem. 265, 3116–3123 (1990). [PubMed] [Google Scholar]
- 39.Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K, Katz K, Liu C, Maddipatla Z, Malheiro A, McDaniel K, Ovetsky M, Riley G, Zhou G, Holmes JB, Kattman BL, Maglott DR, ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062–D1067 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Chora JR, Iacocca MA, Tichý L, Wand H, Kurtz CL, Zimmermann H, Leon A, Williams M, Humphries SE, Hooper AJ, Trinder M, Brunham LR, Costa Pereira A, Jannes CE, Chen M, Chonis J, Wang J, Kim S, Johnston T, Soucek P, Kramarek M, Leigh SE, Carrié A, Sijbrands EJ, Hegele RA, Freiberger T, Knowles JW, Bourbon M, ClinGen Familial Hypercholesterolemia Expert Panel, The Clinical Genome Resource (ClinGen) Familial Hypercholesterolemia Variant Curation Expert Panel consensus guidelines for LDLR variant classification. Genet. Med. 24, 293–306 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Brnich SE, Abou Tayoun AN, Couch FJ, Cutting GR, Greenblatt MS, Heinen CD, Kanavy DM, Luo X, McNulty SM, Starita LM, Tavtigian SV, Wright MW, Harrison SM, Biesecker LG, Berg JS, Clinical Genome Resource Sequence Variant Interpretation Working Group, Recommendations for application of the functional evidence PS3/BS3 criterion using the ACMG/AMP sequence variant interpretation framework. Genome Med. 12, 3 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Tavtigian SV, Greenblatt MS, Harrison SM, Nussbaum RL, Prabhu SA, Boucher KM, Biesecker LG, ClinGen Sequence Variant Interpretation Working Group (ClinGen SVI), Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genet. Med. 20, 1054–1060 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Floyd BJ, Weile J, Kannankeril PJ, Glazer AM, Reuter CM, MacRae CA, Ashley EA, Roden DM, Roth FP, Parikh VN, Proactive Variant Effect Mapping Aids Diagnosis in Pediatric Cardiac Arrest. Circ Genom Precis Med 16, e003792 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.van Loggerenberg W, Sowlati-Hashjin S, Weile J, Hamilton R, Chawla A, Sheykhkarimli D, Gebbia M, Kishore N, Frésard L, Mustajoki S, Pischik E, Di Pierro E, Barbaro M, Floderus Y, Schmitt C, Gouya L, Colavin A, Nussbaum R, Friesema ECH, Kauppinen R, To-Figueras J, Aarsand AK, Desnick RJ, Garton M, Roth FP, Systematically testing human HMBS missense variants to reveal mechanism and pathogenic variation. Am. J. Hum. Genet. 110, 1769–1786 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Benn M, Watts GF, Tybjaerg-Hansen A, Nordestgaard BG, Familial hypercholesterolemia in the Danish general population: prevalence, coronary artery disease, and cholesterol-lowering medication. J. Clin. Endocrinol. Metab. 97, 3956–3964 (2012). [DOI] [PubMed] [Google Scholar]
- 46.World Health Organization, Familial hypercholesterolemia: report of a second WHO Consultation. (1999).
- 47.Mues KE, Bogdanov AN, Monda KL, Yedigarova L, Liede A, Kallenbach L, How well can familial hypercholesterolemia be identified in an electronic health record database? Clin. Epidemiol. 10, 1667–1677 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Safarova MS, Liu H, Kullo IJ, Rapid identification of familial hypercholesterolemia from electronic health records: The SEARCH study. J. Clin. Lipidol. 10, 1230–1239 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Rämö JT, Jurgens SJ, Kany S, Choi SH, Wang X, Smirnov AN, Friedman SF, Maddah M, Khurshid S, Ellinor PT, Pirruccello JP, Rare genetic variants in LDLR, APOB, and PCSK9 are associated with aortic stenosis. Circulation 150, 1767–1780 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.O’Leary DH, Polak JF, Kronmal RA, Manolio TA, Burke GL, Wolfson SK Jr, Carotid-artery intima and media thickness as a risk factor for myocardial infarction and stroke in older adults. Cardiovascular Health Study Collaborative Research Group. N. Engl. J. Med. 340, 14–22 (1999). [DOI] [PubMed] [Google Scholar]
- 51.Thompson DJ, Wells D, Selzam S, Peneva I, Moore R, Sharp K, Tarran WA, Beard EJ, Riveros-Mckay F, Giner-Delgado C, Palmer D, Seth P, Harrison J, Futema M, McVean G, Plagnol V, Donnelly P, Weale ME, Genomics England Research Consortium, UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits, bioRxiv (2022). 10.1101/2022.06.16.22276246. [DOI] [Google Scholar]
- 52.Khera AV, Chaffin M, Aragam KG, Haas ME, Roselli C, Choi SH, Natarajan P, Lander ES, Lubitz SA, Ellinor PT, Kathiresan S, Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations. Nat. Genet. 50, 1219–1224 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Fahed AC, Wang M, Homburger JR, Patel AP, Bick AG, Neben CL, Lai C, Brockman D, Philippakis A, Ellinor PT, Cassa CA, Lebo M, Ng K, Lander ES, Zhou AY, Kathiresan S, Khera AV, Polygenic background modifies penetrance of monogenic variants for tier 1 genomic conditions. Nat. Commun. 11, 3635 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Paquette M, Trinder M, Ruel I, Guay S-P, Hegele RA, Genest J, Brunham LR, Baass A, Polygenic risk score for coronary artery disease predicts atherosclerotic cardiovascular disease in familial hypercholesterolemia. J. Clin. Lipidol. 19, 595–604 (2025). [DOI] [PubMed] [Google Scholar]
- 55.Trinder M, Cermakova L, Ruel I, Baass A, Paquette M, Wang J, Kennedy BA, Hegele RA, Genest J, Brunham LR, Influence of polygenic background on the clinical presentation of familial hypercholesterolemia. Arterioscler. Thromb. Vasc. Biol. 44, 1683–1693 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Lennon NJ, Kottyan LC, Kachulis C, Abul-Husn NS, Arias J, Belbin G, Below JE, Berndt SI, Chung WK, Cimino JJ, Clayton EW, Connolly JJ, Crosslin DR, Dikilitas O, Velez Edwards DR, Feng Q, Fisher M, Freimuth RR, Ge T, GIANT Consortium, All of Us Research Program Glessner JT, Gordon AS, Patterson C, Hakonarson H, Harden M, Harr M, Hirschhorn JN, Hoggart C, Hsu L, Irvin MR, Jarvik GP, Karlson EW, Khan A, Khera A, Kiryluk K, Kullo I, Larkin K, Limdi N, Linder JE, Loos RJF, Luo Y, Malolepsza E, Manolio TA, Martin LJ, McCarthy L, McNally EM, Meigs JB, Mersha TB, Mosley JD, Musick A, Namjou B, Pai N, Pesce LL, Peters U, Peterson JF, Prows CA, Puckelwartz MJ, Rehm HL, Roden DM, Rosenthal EA, Rowley R, Sawicki KT, Schaid DJ, Smit RAJ, Smith JL, Smoller JW, Thomas M, Tiwari H, Toledo DM, Vaitinadin NS, Veenstra D, Walunas TL, Wang Z, Wei W-Q, Weng C, Wiesner GL, Yin X, Kenny EE, Selection, optimization and validation of ten chronic disease polygenic risk scores for clinical implementation in diverse US populations. Nat. Med. 30, 480–487 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.deGoma EM, Ahmad ZS, O’Brien EC, Kindt I, Shrader P, Newman CB, Pokharel Y, Baum SJ, Hemphill LC, Hudgins LC, Ahmed CD, Gidding SS, Duffy D, Neal W, Wilemon K, Roe MT, Rader DJ, Ballantyne CM, Linton MF, Duell PB, Shapiro MD, Moriarty PM, Knowles JW, Treatment gaps in adults with heterozygous familial hypercholesterolemia in the United States: Data from the CASCADE-FH Registry. Circ. Cardiovasc. Genet. 9, 240–249 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Nordestgaard BG, Chapman MJ, Humphries SE, Ginsberg HN, Masana L, Descamps OS, Wiklund O, Hegele RA, Raal FJ, Defesche JC, Wiegman A, Santos RD, Watts GF, Parhofer KG, Hovingh GK, Kovanen PT, Boileau C, Averna M, Borén J, Bruckert E, Catapano AL, Kuivenhoven JA, Pajukanta P, Ray K, Stalenhoef AFH, Stroes E, Taskinen M-R, Tybjærg-Hansen A, European Atherosclerosis Society Consensus Panel, Familial hypercholesterolaemia is underdiagnosed and undertreated in the general population: guidance for clinicians to prevent coronary heart disease: consensus statement of the European Atherosclerosis Society. Eur. Heart J. 34, 3478–90a (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Kramer AI, Christian S, Bartels K, Vaizman N, Hegele RA, Brunham LR, Genetic testing for familial hypercholesterolemia: The current state of its implementation in Canada. CJC Open 6, 1395–1402 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Aldworth H, Hooper NM, Post-translational regulation of the low-density lipoprotein receptor provides new targets for cholesterol regulation. Biochem. Soc. Trans. 52, 431–440 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Zelcer N, Hong C, Boyadjian R, Tontonoz P, LXR regulates cholesterol uptake through Idol-dependent ubiquitination of the LDL receptor. Science 325, 100–104 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Sarraju A, Knowles JW, Genetic testing and risk scores: Impact on familial Hypercholesterolemia. Front. Cardiovasc. Med 6, 5 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Goldberg AC, Hopkins PN, Toth PP, Ballantyne CM, Rader DJ, Robinson JG, Daniels SR, Gidding SS, de Ferranti SD, Ito MK, McGowan MP, Moriarty PM, Cromwell WC, Ross JL, Ziajka PE, National Lipid Association Expert Panel on Familial Hypercholesterolemia, Familial hypercholesterolemia: screening, diagnosis and management of pediatric and adult patients: clinical guidance from the National Lipid Association Expert Panel on Familial Hypercholesterolemia. J. Clin. Lipidol 5, S1–8 (2011). [DOI] [PubMed] [Google Scholar]
- 64.Jacob EO, Hegele RA, How reliable are polygenic risk scores for risk prediction in patients with heart disease? Expert Rev. Mol. Diagn. 23, 105–107 (2023). [DOI] [PubMed] [Google Scholar]
- 65.Dawood M, Fayer S, Pendyala S, Post M, Kalra D, Patterson K, Venner E, Muffley LA, Fowler DM, Rubin AF, Posey JE, Plon SE, Lupski JR, Gibbs RA, Starita LM, Robles-Espinoza CD, Coyote-Maestas W, Gallego Romero I, Using multiplexed functional data to reduce variant classification inequities in underrepresented populations. Genome Med. 16, 143 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Baeza-Centurion P, Miñana B, Valcárcel J, Lehner B, Mutations primarily alter the inclusion of alternatively spliced exons. Elife 9 (2020). [Google Scholar]
- 67.Julien P, Miñana B, Baeza-Centurion P, Valcárcel J, Lehner B, The complete local genotype–phenotype landscape for the alternative splicing of a human exon. Nat. Commun. 7, 1–8 (2016). [Google Scholar]
- 68.Kircher M, Xiong C, Martin B, Schubach M, Inoue F, Bell RJA, Costello JF, Shendure J, Ahituv N, Saturation mutagenesis of twenty disease-associated regulatory elements at single base-pair resolution. Nat. Commun. 10, 3583 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.McClellan J, King M-C, Genetic heterogeneity in human disease. Cell 141, 210–217 (2010). [DOI] [PubMed] [Google Scholar]
- 70.Wright CF, West B, Tuke M, Jones SE, Patel K, Laver TW, Beaumont RN, Tyrrell J, Wood AR, Frayling TM, Hattersley AT, Weedon MN, Assessing the pathogenicity, penetrance, and expressivity of putative disease-causing variants in a population setting. Am. J. Hum. Genet. 104, 275–286 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.O’Neill MJ, Sala L, Denjoy I, Wada Y, Kozek K, Crotti L, Dagradi F, Kotta M-C, Spazzolini C, Leenhardt A, Salem J-E, Kashiwa A, Ohno S, Tao R, Roden DM, Horie M, Extramiana F, Schwartz PJ, Kroncke BM, Continuous Bayesian variant interpretation accounts for incomplete penetrance among Mendelian cardiac channelopathies. Genet. Med. 25, 100355 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Kroncke BM, Smith DK, Zuo Y, Glazer AM, Roden DM, Blume JD, A Bayesian method to estimate variant-induced disease penetrance. PLoS Genet. 16, e1008862 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Findlay GM, Daza RM, Martin B, Zhang MD, Leith AP, Gasperini M, Janizek JD, Huang X, Starita LM, Shendure J, Accurate classification of BRCA1 variants with saturation genome editing. Nature 562, 217–222 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Starita LM, Islam MM, Banerjee T, Adamovich AI, Gullingsrud J, Fields S, Shendure J, Parvin JD, A Multiplex Homology-Directed DNA Repair Assay Reveals the Impact of More Than 1,000 BRCA1 Missense Substitution Variants on Protein Function. Am. J. Hum. Genet. 103, 498–508 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Starita LM, Young DL, Islam M, Kitzman JO, Gullingsrud J, Hause RJ, Fowler DM, Parvin JD, Shendure J, Fields S, Massively Parallel Functional Analysis of BRCA1 RING Domain Variants. Genetics 200, 413–422 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Kweon J, Jang A-H, Shin HR, See J-E, Lee W, Lee JW, Chang S, Kim K, Kim Y, A CRISPR-based base-editing screen for the functional assessment of BRCA1 variants. Oncogene 39, 30–35 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Diabate M, Islam MM, Nagy G, Banerjee T, Dhar S, Smith N, Adamovich AI, Starita LM, Parvin JD, DNA repair function scores for 2172 variants in the BRCA1 amino-terminus. PLoS Genet. 19, e1010739 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Adamovich AI, Diabate M, Banerjee T, Nagy G, Smith N, Duncan K, Mendoza Mendoza E, Prida G, Freitas MA, Starita LM, Parvin JD, The functional impact of BRCA1 BRCT domain variants using multiplexed DNA double-strand break repair assays. Am. J. Hum. Genet. 109, 618–630 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Sahu S, Galloux M, Southon E, Caylor D, Sullivan T, Arnaudi M, Zanti M, Geh J, Chari R, Michailidou K, Papaleo E, Sharan SK, Saturation genome editing-based clinical classification of BRCA2 variants. Nature 638, 538–545 (2025). [DOI] [PubMed] [Google Scholar]
- 80.Sahu S, Sullivan TL, Mitrophanov AY, Galloux M, Nousome D, Southon E, Caylor D, Mishra AP, Evans CN, Clapp ME, Burkett S, Malys T, Chari R, Biswas K, Sharan SK, Saturation genome editing of 11 codons and exon 13 of BRCA2 coupled with chemotherapeutic drug response accurately determines pathogenicity of variants. PLoS Genet. 19, e1010940 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Ikegami M, Kohsaka S, Ueno T, Momozawa Y, Inoue S, Tamura K, Shimomura A, Hosoya N, Kobayashi H, Tanaka S, Mano H, High-throughput functional evaluation of BRCA2 variants of unknown significance. Nat. Commun. 11, 2573 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Jia X, Burugula BB, Chen V, Lemons RM, Jayakody S, Maksutova M, Kitzman JO, Massively parallel functional testing of MSH2 missense variants conferring Lynch syndrome risk. Am. J. Hum. Genet. 108, 163–175 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Amberger JS, Bocchini CA, Scott AF, Hamosh A, Omim.org: leveraging knowledge across phenotype-gene relationships. Nucleic Acids Res. 47, D1038–D1043 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Majithia AR, Tsuda B, Agostini M, Gnanapradeepan K, Rice R, Peloso G, Patel KA, Zhang X, Broekema MF, Patterson N, Duby M, Sharpe T, Kalkhoven E, Rosen ED, Barroso I, Ellard S, UK Monogenic Diabetes Consortium, S. Kathiresan, Myocardial Infarction Genetics Consortium, S. O’Rahilly, UK Congenital Lipodystrophy Consortium, Chatterjee K, Florez JC, Mikkelsen T, Savage DB, Altshuler D, Prospective functional classification of all possible missense variants in PPARG. Nat. Genet. 48, 1570–1575 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Longwell CK, Hanna S, Hartrampf N, Sperberg RAP, Huang P-S, Pentelute BL, Cochran JR, Identification of N-terminally diversified GLP-1R agonists using saturation Mutagenesis and chemical design. ACS Chem. Biol 16, 58–66 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Gersing S, Cagiada M, Gebbia M, Gjesing AP, Coté AG, Seesankar G, Li R, Tabet D, Weile J, Stein A, Gloyn AL, Hansen T, Roth FP, Lindorff-Larsen K, Hartmann-Petersen R, A comprehensive map of human glucokinase variant activity. Genome Biol. 24, 97 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Glazer AM, Kroncke BM, Matreyek KA, Yang T, Wada Y, Shields T, Salem J-E, Fowler DM, Roden DM, Deep Mutational Scan of an SCN5A Voltage Sensor. [Preprint] (2020). 10.1161/circgen.119.002786. [DOI] [Google Scholar]
- 88.Kozek KA, Glazer AM, Ng C-A, Blackwell D, Egly CL, Vanags LR, Blair M, Mitchell D, Matreyek KA, Fowler DM, Knollmann BC, Vandenberg JI, Roden DM, Kroncke BM, High-throughput discovery of trafficking-deficient variants in the cardiac potassium channel KV11.1. [Preprint] (2020). 10.1016/j.hrthm.2020.05.041. [DOI] [Google Scholar]
- 89.O’Neill MJ, Ng C-A, Aizawa T, Sala L, Bains S, Winbo A, Ullah R, Shen Q, Tan C-Y, Kozek K, Vanags LR, Mitchell DW, Shen A, Wada Y, Kashiwa A, Crotti L, Dagradi F, Musu G, Spazzolini C, Neves R, Bos JM, Giudicessi JR, Bledsoe X, Gamazon ER, Lancaster MC, Glazer AM, Knollmann BC, Roden DM, Weile J, Roth F, Salem J-E, Earle N, Stiles R, Agee T, Johnson CN, Horie M, Skinner JR, Ackerman MJ, Schwartz PJ, Ohno S, Vandenberg JI, Kroncke BM, Multiplexed assays of variant effect and automated patch clamping improve KCNH2-LQTS variant classification and cardiac event risk stratification. Circulation 150, 1869–1881 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Muhammad A, Calandranis ME, Li B, Yang T, Blackwell DJ, Harvey ML, Smith JE, Daniel ZA, Chew AE, Capra JA, Matreyek KA, Fowler DM, Roden DM, Glazer AM, High-throughput functional mapping of variants in an arrhythmia gene, KCNE1, reveals novel biology. Genome Med. 16, 73 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Fowler DM, Adams DJ, Gloyn AL, Hahn WC, Marks DS, Muffley LA, Neal JT, Roth FP, Rubin AF, Starita LM, Hurles ME, An Atlas of Variant Effects to understand the genome at nucleotide resolution. Genome Biol. 24, 147 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Fowler D, Adams D, Ahituv N, Bock C, Bolognesi B, Chanock S, Cheng J, Cho J, Davis M, del Angel G, Doench J, Ellingford J, Estrada K, Kyle Fahr K, Fayer S, Gallego Romero I, Glazer A, Huangfu D, Landrum M, Lehner B, Lossie A, Martin M-J, Massart M, Nelson S, O’Donnell-Luria A, Parikh V, Pesaran T, Posey J, Powell B, Roden D, Rubin A, Semple RK, Simola D, Singh H, Starita L, Tejura M, Vockley J, Wagner A, Won H, Birney E, Cool J, Hurles M, Hutter C, Muffley L, Rehm HL, Roth F, Atlas of Variant Effects 2030 Roadmap: resolving human variants of uncertain significance. Zenodo; [Preprint] (2025). 10.5281/ZENODO.15420414. [DOI] [Google Scholar]
- 93.Glazer AM, Tabet DR, Parikh VN, Kroncke BM, Cote AG, Yamamoto Y, Wang Q, Muhammad A, Lancaster MC, O’Neill MJ, Weile J, Yang T, Macrae CA, Ashley EA, Roth FP, Roden DM, Creating an atlas of variant effects to resolve variants of uncertain significance and guide cardiovascular medicine. Nat. Rev. Cardiol, doi: 10.1038/s41569-025-01201-7 (2025). [DOI] [Google Scholar]
- 94.Gibson DG, Young L, Chuang R-Y, Venter JC, Hutchison CA 3rd, Smith HO, Enzymatic assembly of DNA molecules up to several hundred kilobases. Nat. Methods 6, 343–345 (2009). [DOI] [PubMed] [Google Scholar]
- 95.Yang X, Boehm JS, Yang X, Salehi-Ashtiani K, Hao T, Shen Y, Lubonja R, Thomas SR, Alkan O, Bhimdi T, Green TM, Johannessen CM, Silver SJ, Nguyen C, Murray RR, Hieronymus H, Balcha D, Fan C, Lin C, Ghamsari L, Vidal M, Hahn WC, Hill DE, Root DE, A public genome-scale lentiviral expression library of human ORFs. Nat. Methods 8, 659–661 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Ran FA, Hsu PD, Wright J, Agarwala V, Scott DA, Zhang F, Genome engineering using the CRISPR-Cas9 system. Nat. Protoc. 8, 2281–2308 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Hart T, Tong AHY, Chan K, Van Leeuwen J, Seetharaman A, Aregger M, Chandrashekhar M, Hustedt N, Seth S, Noonan A, Habsid A, Sizova O, Nedyalkova L, Climie R, Tworzyanski L, Lawson K, Sartori MA, Alibeh S, Tieu D, Masud S, Mero P, Weiss A, Brown KR, Usaj M, Billmann M, Rahman M, Constanzo M, Myers CL, Andrews BJ, Boone C, Durocher D, Moffat J, Evaluation and design of genome-wide CRISPR/SpCas9 knockout screens. G3 (Bethesda) 7, 2719–2727 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Axakova A, Ding M, Cote AG, Subramaniam R, Senguttuvan V, Zhang H, Weile J, Douville SV, Gebbia M, Al-Chalabi A, Wahl A, Reuter J, Hurt J, Mitchell AA, Fradette S, Andersen PM, van Loggerenberg W, Roth FP, Landscapes of missense variant impact for human superoxide dismutase 1. Am. J. Hum. Genet. 112, 2295–2315 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Weile J, Sun S, Cote AG, Knapp J, Verby M, Mellor JC, Wu Y, Pons C, Wong C, van Lieshout N, Yang F, Tasan M, Tan G, Yang S, Fowler DM, Nussbaum R, Bloom JD, Vidal M, Hill DE, Aloy P, Roth FP, A framework for exhaustively mapping functional missense variants. Mol. Syst. Biol. 13, 957 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Baid G, Cook DE, Shafin K, Yun T, Llinares-López F, Berthet Q, Belyaeva A, Töpfer A, Wenger AM, Rowell WJ, Yang H, Kolesnikov A, Ammar W, Vert J-P, Vaswani A, McLean CY, Nattestad M, Chang P-C, Carroll A, DeepConsensus improves the accuracy of sequences with a gap-aware sequence transformer. Nat. Biotechnol. 41, 232–238 (2023). [DOI] [PubMed] [Google Scholar]
- 101.Weile J, Ferra G, Boyle G, Pendyala S, Amorosi C, Yeh C-L, Cote AG, Kishore N, Tabet D, van Loggerenberg W, Rayhan A, Fowler DM, Dunham MJ, Roth FP, Pacybara: accurate long-read sequencing for barcoded mutagenized allelic libraries. Bioinformatics 40 (2024). [Google Scholar]
- 102.Sun S, Weile J, Verby M, Wu Y, Wang Y, Cote AG, Fotiadou I, Kitaygorodsky J, Vidal M, Rine J, Ješina P, Kožich V, Roth FP, A proactive genotype-to-patient-phenotype map for cystathionine beta-synthase. Genome Med. 12, 13 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Zhao L, Liu Z, Levy SF, Wu S, Bartender: a fast and accurate clustering algorithm to count barcode reads. Bioinformatics 34, 739–747 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Tiberti M, Terkelsen T, Degn K, Beltrame L, Cremers TC, da Piedade I, Di Marco M, Maiani E, Papaleo E, MutateX: an automated pipeline for in silico saturation mutagenesis of protein structures and structural ensembles. Brief. Bioinform. 23 (2022). [Google Scholar]
- 105.Tabet DR, Kuang D, Lancaster MC, Li R, Liu K, Weile J, Coté AG, Wu Y, Hegele RA, Roden DM, Roth FP, Benchmarking computational variant effect predictors by their ability to infer human traits. Genome Biol. 25, 172 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Szustakowski JD, Balasubramanian S, Kvikstad E, Khalid S, Bronson PG, Sasson A, Wong E, Liu D, Wade Davis J, Haefliger C, Katrina Loomis A, Mikkilineni R, Noh HJ, Wadhawan S, Bai X, Hawes A, Krasheninina O, Ulloa R, Lopez AE, Smith EN, Waring JF, Whelan CD, Tsai EA, Overton JD, Salerno WJ, Jacob H, Szalma S, Runz H, Hinkle G, Nioi P, Petrovski S, Miller MR, Baras A, Mitnaul LJ, Reid JG, UKB-ESC Research Team, Advancing human genetics research and drug discovery through exome sequencing of the UK Biobank. Nat. Genet. 53, 942–948 (2021). [DOI] [PubMed] [Google Scholar]
- 107.Sinnott-Armstrong N, Tanigawa Y, Amar D, Mars N, Benner C, Aguirre M, Venkataraman GR, Wainberg M, Ollila HM, Kiiskinen T, Havulinna AS, Pirruccello JP, Qian J, Shcherbina A, FinnGen, Rodriguez F, Assimes TL, Agarwala V, Tibshirani R, Hastie T, Ripatti S, Pritchard JK, Daly MJ, Rivas MA, Genetics of 35 blood and urine biomarkers in the UK Biobank. Nat. Genet. 53, 185–194 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Helgadottir A, Thorleifsson G, Gretarsdottir S, Stefansson OA, Tragante V, Thorolfsdottir RB, Jonsdottir I, Bjornsson T, Steinthorsdottir V, Verweij N, Nielsen JB, Zhou W, Folkersen L, Martinsson A, Heydarpour M, Prakash S, Oskarsson G, Gudbjartsson T, Geirsson A, Olafsson I, Sigurdsson EL, Almgren P, Melander O, Franco-Cereceda A, Hamsten A, Fritsche L, Lin M, Yang B, Hornsby W, Guo D, Brummett CM, Abecasis G, Mathis M, Milewicz D, Body SC, Eriksson P, Willer CJ, Hveem K, Newton-Cheh C, Smith JG, Danielsen R, Thorgeirsson G, Thorsteinsdottir U, Gudbjartsson DF, Holm H, Stefansson K, Genome-wide analysis yields new loci associating with aortic valve stenosis. Nat. Commun. 9, 987 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Mitra S, Biswas RK, Hooijenga P, Cassidy S, Nova A, De Ciutiis I, Wang T, Kroeger CM, Stamatakis E, Masedunskas A, De Caterina R, Cagigas ML, Fontana L, Carotid intima-media thickness, cardiovascular disease, and risk factors in 29,000 UK Biobank adults. Am. J. Prev. Cardiol. 22, 101011 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data are available in the main text or the supplementary materials. Correspondence and requests for materials should be addressed to F.P.R.
