Abstract
Background:
Hypertension comprises a heterogeneous range of phenotypes. We asked whether underlying genetic structure could explain a part of this heterogeneity.
Methods:
Our study sample comprised N=198,148 FinnGen participants (56% women, mean age 58 years) and N=21,168 well-phenotyped FINRISK participants (53% women, mean age 50 years). First, we identified genetic hypertension components with an unsupervised Bayesian non-negative matrix factorization algorithm using public genome-wide association data for 144 genetic hypertension variants and 16 clinical traits. For these components, we computed their (1) cross-sectional associations with clinical traits in FINRISK using linear regression and (2) longitudinal associations with incident adverse outcomes in FinnGen using Cox regression.
Results:
We observed 4 genetic hypertension components corresponding to recognizable clinical phenotypes: Obesity (high body mass index), Dyslipidemia (low high-density lipoprotein cholesterol and high triglycerides), Hypolipidemia (low low-density lipoprotein cholesterol and low total cholesterol), and Short Stature. In FINRISK, all hypertension components had robust associations with their respective clinical characteristics. In FinnGen, the Obesity component was associated with increased diabetes risk (hazard ratio per 1 SD increase 1.08, Bonferroni corrected confidence interval 1.05–1.10) and the Hypolipidemia component with increased autoimmune disease risk (hazard ratio per 1 SD increase 1.05, Bonferroni corrected confidence interval 1.03–1.07). In addition, all hypertension components were related to both hypertension and cardiovascular disease.
Conclusions:
Our unsupervised analysis demonstrates that the genetic basis of hypertension can be understood as a mixture of 4 broad, clinically interpretable components capturing disease heterogeneity. These components could be used to stratify individuals into specific genetic subtypes and, therefore, to benefit personalized healthcare and pharmaceutical research.
Keywords: hypertension, genetics, association, risk factors
Introduction
To treat hypertension and prevent its complications effectively, the clinician should be able to provide individualized treatment to hypertensive patients. However, because the heterogeneity of hypertension is poorly understood, treatment strategies of this complex polygenic disease remain largely impersonalized.1
To elucidate the heterogeneity of hypertension, we and others have previously identified phenotypic hypertension components based on clinical characteristics.2—4 However, the inconsistent results between these studies suggest that phenotypic hypertension components are sensitive to differences in study samples and statistical methods. Moreover, clinical characteristics change with time and, therefore, need to be measured at regular intervals. In contrast to clinical characteristics, genetic variants are fixed and known to predict the onset of polygenic diseases such as type 2 diabetes.5 Based on genetic variants, both Ma et al.6 (N=1,187) and Luo et al.7 (N=660) subtyped hypertensive individuals and uncovered genetic hypertension components with differing echocardiographic measurements. While these two studies have pioneered a classification of hypertension genetics, they are limited in sample size and by phenotype data that were mainly limited to echocardiographic variables.
Genome wide association studies (GWASs) are commonly used for estimating the associations between millions of genetic variants and a phenotype such as hypertension. For example, the UK Biobank8 alone hosts publicly available GWAS results from over 13,000,000 single nucleotide polymorphisms (SNPs) and 7,000 phenotypes derived from up to 500,000 individuals. Therefore, an opportunity exists to leverage existing high quality SNP-phenotype association data for a detailed genetic analysis. Following this idea, Udler et al.9 applied an innovative subtyping algorithm to GWAS results for type 2 diabetes and its risk factors, revealing 5 genetic type 2 diabetes components with distinct clinical characteristics and outcomes. However, no prior study has used existing GWAS results to assess the heterogeneity in hypertension genetics.
An improved understanding of heterogeneity in hypertension genetics is needed to personalize hypertension treatment. We aimed to identify genetic hypertension subtypes by applying a subtyping algorithm, Bayesian non-negative matrix factorization (bNMF), while leveraging publicly available GWAS results for hypertension and its risk factors. We then assessed the clinical significance of these components by estimating their associations with clinical characteristics and adverse outcomes in an independent cohort of almost 200,000 individuals.
Methods
Because of the sensitive nature of the data collected for this study, requests to access the data set from qualified researchers trained in human subject confidentiality protocols may be submitted through the Finnish Biobanks’ FinnBB portal (https://finbb.fi/) for FinnGen and at https://www.thl.fi/biobank/researchers for FINRISK. The Coordinating Ethics Committee of the Hospital District of Helsinki and Uusimaa approved both FinnGen and FINRISK study protocols. All participants gave informed written consent. Figure 1 shows a flow chart for the current study, and the methods, including explicit mathematical formulas and R code for bNMF, are provided in the Data Supplement.
Figure 1. Flow chart for the current study.
After applying Bayesian non-negative matrix factorization (bNMF) on genome-wide association study (GWAS) summary statistics, we observed 4 genetic hypertension components. Downstream analyses indicated that these components have robust clinical interpretations consistent with previous knowledge about hypertension (Figures and Tables). *bNMF weights for variants and GWAS traits in each genetic hypertension component, †literature-based associations between genetic hypertension components and GWAS traits, ‡cross-sectional associations between genetic hypertension components and clinical traits using linear regression, §longitudinal associations between genetic hypertension components and disease end points using Cox models.
Results
Genetic Hypertension Component Characteristics from GWAS Summary Statistics
After applying bNMF to 144 hypertension SNPs and 16 hypertension-related traits for 1,000 iterations, the number of latent genetic hypertension components was 2 in 28 iterations, 3 in 462 iterations, and 4 in 510 iterations, meaning that a 4-component solution (Figure 2) was the most supported by the data. The 4-component solution appeared identical to the 3-component solution (Figure I in the Data Supplement) apart from a dyslipidemia-like component (associations to high-density lipoprotein [HDL] cholesterol and triglycerides) that was not present in the 3-component solution. While the 3-component solution given by bNMF was also a good fit to the data, we considered this additional dyslipidemia-like component to be clinically important because hypertension and dyslipidemia often appear together in metabolic syndrome.10 Therefore, we chose the 4-component solution for subsequent analyses as it was supported by both the data and clinical experience.
Figure 2. Associations between genetic hypertension components and hypertension-related traits for the 4-component solution.
After applying Bayesian non-negative matrix factorization (bNMF) to 144 hypertension SNPs and 16 clinical traits, we chose the 4-component solution for subsequent analyses. The bNMF method provides weights that link the original traits to the hypertension components, allowing us to interpret their clinical meaning. A high weight between a component and a plus/minus-signed trait indicates a positive/negative component-trait association, respectively. The rows and columns of the weight matrix have been sorted using the R function hclust with Euclidean distance and the Ward agglomeration method (ward.D). BMI indicates body mass index; CRP, C-reactive protein; eGFR, estimated glomerular filtration rate; HbA1c, glycated hemoglobin; HDLc, high-density lipoprotein cholesterol; LDLc, low-density lipoprotein cholesterol; and UACR, urine albumin-to-creatinine ratio.
The highest weighted traits in each component were height (component 1), BMI and waist circumference (component 2), low-density lipoprotein (LDL) cholesterol and total cholesterol (component 3), and HDL cholesterol and triglycerides (component 4) (Figure 2), while the highest weighted loci were LCORL (component 1), FTO (component 2), SH2B3 (component 3), and KLF14 (component 4). The genetic variants and their corresponding weights in each of the four components are described in Table 1 and Tables I–V in the Data Supplement.
Table 1.
Genetic Variants Included in Genetic Hypertension Components.
| Obesity N Loci = 28 | Dyslipidemia N Loci = 14 | Hypolipidemia N Loci = 27 | Short Stature N Loci = 25 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ID (Locus) | Chr | bNMF Weight | ID (Locus) | Chr | bNMF Weight | ID (Locus) | Chr | bNMF Weight | ID (Locus) | Chr | bNMF Weight |
| rs55872725 (FTO) | 16 | 5.14 | rs6974400 (KLF14) | 7 | 2.61 | rs3184504 (SH2B3) | 12 | 2.23 | 4:17940040_CT_C (LCORL) | 4 | 2.35 |
| rs9972727 (PRSS36) | 16 | 1.56 | rs13179413 (C5orf67) | 5 | 2.23 | rs198851 (HIST1H4C) | 6 | 1.35 | 6:31610189_TAAAG_T (BAG6) | 6 | 2.2 |
| rs144020965 (REXO1) | 19 | 1.35 | rs645040 (MSL2) | 3 | 1.51 | 6:127163736_ACT_A (RSPO3) | 6 | 1.35 | rs12656497 (NPR3) | 5 | 1.83 |
| rs2977324 (HNF4G) | 8 | 1.27 | rs2287922 (RASIP1) | 19 | 1.27 | rs77924615 (PDILT) | 16 | 1.32 | rs62071306 (TBX2) | 17 | 1.32 |
| rs58241742 (SPDYE5) | 7 | 1.17 | rs7117386 (PTPRJ) | 11 | 1.10 | rs34592089 (BANK1) | 4 | 1.29 | rs10493818 (GTF2B) | 1 | 1.27 |
| rs645040 (MSL2) | 3 | 1.16 | rs3184504 (SH2B3) | 12 | 1.07 | rs11072508 (CSK) | 15 | 1.11 | rs3918226 (NOS3) | 7 | 1.13 |
| rs7118770 (HSD17B12) | 11 | 1.14 | rs557675 (OVOL1) | 11 | 0.97 | rs880315 (CASZ1) | 1 | 1.07 | rs17637472 (PHB) | 17 | 1.07 |
| rs1845034 (RP11–397G17.1) | 6 | 1.11 | rs80226362 (INPP5A) | 10 | 0.75 | rs740746 (ADRB1) | 10 | 0.99 | rs2287922 (RASIP1) | 19 | 1.05 |
| rs17419291 (TMEM161B) | 5 | 1.11 | rs9972727 (PRSS36) | 16 | 0.73 | rs2643826 (SLC4A7) | 3 | 0.98 | rs11605518 (ARNTL) | 11 | 0.9 |
| rs4841436 (LOC102723313) | 8 | 1.08 | rs6564889 (CMIP) | 16 | 0.66 | rs557675 (OVOL1) | 11 | 0.97 | rs13405815 (FOSL2) | 2 | 0.89 |
| rs557675 (OVOL1) | 11 | 1.05 | rs6923212 (FHL5) | 6 | 0.65 | rs4883481 (LIMA1) | 12 | 0.97 | rs12324159 (EXD1) | 15 | 0.84 |
| rs2627323 (ABHD17C) | 15 | 1.01 | rs13405815 (FOSL2) | 2 | 0.58 | rs6464165 (PRKAG2) | 7 | 0.83 | rs7117386 (PTPRJ) | 11 | 0.81 |
| rs11770148 (MAD1L1) | 7 | 1.00 | rs17637472 (PHB) | 17 | 0.57 | rs8027450 (FURIN) | 15 | 0.81 | rs6773655 (MECOM) | 3 | 0.8 |
| rs765141229 (PKP4) | 2 | 0.99 | rs11072508 (CSK) | 15 | 0.56 | rs72831343 (C10orf107) | 10 | 0.80 | rs11373776 (BAHCC1) | 17 | 0.76 |
| rs8070737 (ZZEF1) | 17 | 0.91 | rs1275988 (KCNK3) | 2 | 0.79 | rs28394055 (CYP11B2) | 8 | 0.74 | |||
| rs1507153 (IRAK1BP1) | 6 | 0.90 | rs1658490 (BICC1) | 10 | 0.76 | rs11556924 (ZC3HC1) | 7 | 0.73 | |||
| rs301799 (RERE) | 1 | 0.87 | rs3790604 (WNT2B) | 1 | 0.73 | rs7763350 (ZNF318) | 6 | 0.69 | |||
| rs587752534 (GTF2IRD2) | 7 | 0.85 | rs3918226 (NOS3) | 7 | 0.72 | rs301799 (RERE) | 1 | 0.68 | |||
| rs11605518 (ARNTL) | 11 | 0.82 | rs4980379 (LSP1) | 11 | 0.70 | rs4841436 (LOC102723313) | 8 | 0.64 | |||
| rs745360648 (FOXP2) | 7 | 0.81 | rs9508495 (SLC7A1) | 13 | 0.67 | rs448385 (RUNX3) | 1 | 0.64 | |||
| rs6564889 (CMIP) | 16 | 0.80 | rs57139556 (PLEKHG1) | 6 | 0.66 | rs139179002 (RIMKLB) | 12 | 0.63 | |||
| rs7498127 (ZNF423) | 16 | 0.71 | rs11428568 (DPEP1) | 16 | 0.66 | rs645040 (MSL2) | 3 | 0.61 | |||
| rs10105774 (NCALD) | 8 | 0.68 | rs145158522 (SORCS3) | 10 | 0.63 | rs1658490 (BICC1) | 10 | 0.57 | |||
| rs80226362 (INPP5A) | 10 | 0.64 | rs7763350 (ZNF318) | 6 | 0.63 | rs55770580 (LOC100996447) | 3 | 0.57 | |||
| rs61862567 (ADK) | 10 | 0.62 | rs62481856 (LINC02577) | 7 | 0.61 | rs3184504 (SH2B3) | 12 | 0.56 | |||
| rs35783704 (ZFPM2) | 8 | 0.61 | rs7246865 (MYO9B) | 19 | 0.60 | ||||||
| rs6031431 (JPH2) | 20 | 0.61 | rs2443708 (ATG7) | 3 | 0.59 | ||||||
| rs13179413 (C5orf67) | 5 | 0.58 | |||||||||
Based on genome-wide association study summary statistics, we selected 144 variants for bNMF, which yielded weights for 4 distinct components: Obesity, Dyslipidemia, Hypolipidemia, and Short Stature. To reduce noise in subsequent analyses, we imposed a post hoc bNMF weight cutoff of 0.55, leaving us with 73 unique variants shown in this table. For details of all 144 variants, see Table V in the Data Supplement. bNMF indicates Bayesian non-negative matrix factorization; and Chr, chromosome.
Literature-based associations between hypertension component genetic risk scores (GRSs; weight cutoff 0.55, Figure II in the Data Supplement) and GWAS traits suggested the following clinical interpretations for the hypertension components: Short Stature (component 1), Obesity (component 2), Hypolipidemia (component 3), and Dyslipidemia (component 4) (Figure 3, Table 2). We will use these names for the hypertension components in the rest of the manuscript.
Figure 3. Defining characteristics of genetic hypertension components.
Bars represent standardized effect sizes (Z-score) for associations between hypertension component genetic risk scores and selected traits. For each hypertension component, we highlighted the trait with the strongest association to the component in black. We derived the effect sizes from genome-wide association study summary statistics using a genetic risk score weighted with inverse variances (see Supplemental Methods for details). BMI indicates body mass index; LDLc, low-density lipoprotein cholesterol; and TG, triglycerides.
Table 2.
Literature-Based Associations of Hypertension Component Genetic Risk Scores and Selected Traits.
| Obesity N Loci = 28 | Dyslipidemia N Loci = 14 | Hypolipidemia N Loci = 27 | Short Stature N Loci = 25 | |||||
|---|---|---|---|---|---|---|---|---|
| Trait | Beta | P Value | Beta | P Value | Beta | P Value | Beta | P Value |
| BMI | 0.0163 | 1.8×10 −240 | 0.0043 | 1.8×10 −10 | −0.0019 | 3.2×10−4 | 0.0053 | 1.0×10 −26 |
| CRP | 0.0059 | 1.2×10 −31 | 0.0051 | 1.5×10 −13 | −0.0003 | 0.54 | 0.0049 | 9.1×10 −22 |
| eGFR | 0.0000 | 0.87 | −0.0005 | 1.0×10−7 | −0.0007 | 7.7×10 −19 | −0.0001 | 9.2×10−2 |
| Glucose | 0.0031 | 2.8×10−9 | 0.0007 | 0.30 | 0.0001 | 0.92 | 0.0012 | 2.6×10−2 |
| HbA1c | 0.0018 | 5.4×10−4 | −0.0004 | 0.55 | −0.0029 | 1.9×10−6 | 0.0004 | 0.45 |
| HDLc | −0.0075 | 5.7×10 −55 | −0.0148 | 1.7×10 −112 | −0.0002 | 0.76 | −0.0043 | 1.1×10 −18 |
| Height | −0.0019 | 8.2×10−8 | −0.0051 | 3.4×10 −27 | 0.0008 | 3.6×10−2 | −0.0139 | <10 −300 |
| LDLc | −0.0015 | 2.6×10−3 | −0.0001 | 0.89 | −0.0107 | 6.6×10 −89 | −0.0019 | 2.2×10−4 |
| Neuroticism | −0.0003 | 0.54 | 0.000 | 0.95 | −0.0008 | 0.16 | 0.0008 | 0.13 |
| Smoking | 0.0012 | 4.0×10−6 | 0.001 | 1.0×10−2 | −0.0003 | 0.32 | 0.0003 | 0.23 |
| TC | −0.0027 | 2.7×10−8 | −0.0024 | 3.2×10−4 | −0.0103 | 3.6×10 −85 | −0.0024 | 9.2×10−7 |
| TG | 0.0059 | 4.6×10 −34 | 0.0166 | 3.1×10 −136 | −0.0036 | 8.7×10 −12 | 0.0047 | 2.7×10 −21 |
| UACR | 0.0010 | 2.9×10−2 | 0.0024 | 2.6×10−5 | 0.0019 | 8.2×10−5 | 0.0033 | 9.6×10 −11 |
| Urate | 0.0094 | 7.1×10 −26 | 0.0129 | 4.1×10 −29 | 0.0125 | 1.1×10 −40 | 0.0095 | 4.5×10 −21 |
| WC (f) | 0.0161 | 4.5×10 −128 | 0.0064 | 4.8×10 −12 | −0.0023 | 1.5×10−3 | 0.0014 | 3.6×10−2 |
| WC (m) | 0.0142 | 5.3×10 −85 | 0.0021 | 3.1×10−2 | −0.0007 | 0.40 | 0.0014 | 6.2×10−2 |
We derived literature-based associations between hypertension components and selected traits from genome-wide association study summary statistics using inverse-variance weighted genetic risk scores (see Supplemental Methods for details). We used the following sources and traits: Chronic Kidney Disease Genetics (eGFR, UACR, and urate), Meta-Analyses of Glucose and Insulin-related traits Consortium (HbA1c) and the UK Biobank (CRP, glucose, HDLc, LDLc, TC, TG, BMI, height, WC (f/m), neuroticism, and smoking). P values <7.8×10−10 and corresponding effect sizes are bolded, representing a Bonferroni correction of 16 traits × 4 components with respect to the genome-wide significance threshold of 5×10−8. BMI indicates body mass index; CRP, C-reactive protein; eGFR, estimated glomerular filtration rate; HbA1c, glycated hemoglobin; HDLc, high-density lipoprotein cholesterol; LDLc, low-density lipoprotein cholesterol; TC, total cholesterol; TG, triglycerides; UACR, urine albumin-to-creatinine ratio; and WC (f/m), female/male waist circumference.
Clinical Associations of Genetic Hypertension Components in FinnGen and FINRISK
In FINRISK (N=21,168; 53% women, mean age at baseline 50 years), all hypertension component GRSs were associated with their respective clinical traits (Table 3). A 1 SD increase in the Obesity component GRS was related to 0.29 kg/m2 (Bonferroni corrected confidence interval [CI], 0.19–0.40; P=5.0×10−21) higher BMI. A 1 SD increase in the Dyslipidemia component GRS was related to 0.013 mmol/L (CI, 0.005–0.022; P=1.2×10−7) lower HDL cholesterol and 1.7 % (CI, 0.7–2.8; P=6.2×10−8) higher triglycerides. A 1 SD increase in the Hypolipidemia component GRS was related to 0.026 mmol/L (CI, 0.006–0.047; P=1.8×10−5) lower LDL cholesterol and 0.028 mmol/L (CI, 0.005–0.051; P=4.5×10−5) lower total cholesterol. A 1 SD increase in the Short Stature component GRS was related to a 0.28 cm (CI, 0.14–0.42; P=1.4×10−9) lower height. The top decile categories of hypertension component GRSs displayed heterogeneous clinical trait means, with Obesity and Short Stature components displaying a statistically significant difference compared to the reference group (Table VI in the Data Supplement). The systolic blood pressure (SBP) GRS was also associated with clinical traits: a 1 SD increase was related to 0.14 kg/m2 (0.04–0.25; P=2.6×10−6) higher BMI, 3.1% (0.4–5.8; P=9.2×10−5) higher C-reactive protein, 0.25 cm (0.11–0.39; P=1.3×10−9) lower height, 0.011 mmol/L (0.002–0.019; P=1.8×10−5) lower HDL cholesterol, and 2.2% (1.1–3.3; P=1.2×10−11) higher triglycerides (Table 4).
Table 3.
Cross-Sectional Associations Between Hypertension Genetic Risk Scores and Clinical Traits.
| Obesity N Loci = 25 | Dyslipidemia N Loci = 14 | Hypolipidemia N Loci = 27 | Short Stature N Loci = 21 | SBP GRS N Loci = 1,098,015 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Trait | Beta (CI) | P Value | Beta (CI) | P Value | Beta (CI) | P Value | Beta (CI) | P Value | Beta (CI) | P Value |
| BMI (kg/m2) | 0.2919 (0.1878, 0.3960) | 5.0×10 −21 | −0.0289 (−0.1332, 0.0755) | 0.35 | −0.0524 (−0.1569, 0.0521) | 9.2×10−2 | −0.0162 (−0.1207, 0.0884) | 0.60 | 0.1468 (0.0420, 0.2516) | 2.6×10 −6 |
| log CRP (mg/L) | 0.0239 (−0.0019, 0.0496) | 1.8×10−3 | −0.0006 (−0.0264, 0.0251) | 0.93 | 0.0031 (−0.0227, 0.0289) | 0.68 | −0.0081 (−0.0177, 0.0339) | 0.29 | 0.0301 (0.0043, 0.0560) | 9.2×10 −5 |
| eGFR (mL/min/1.73 m2) | 0.1321 (−0.1756, 0.4398) | 0.15 | −0.1571 (−0.4649, 0.1507) | 8.6×10−2 | −0.0592 (−0.3647, 0.2491) | 0.52 | 0.3282 (0.0200, 0.6364) | 3.5×10 −4 | 0.0608 (−0.2484, 0.3701) | 0.51 |
| Height (m) | −0.0008 (−0.0022, 0.0006) | 4.5×10−2 | −0.0010 (−0.0024, 0.0004) | 2.0×10−2 | −0.0006 (−0.0020, 0.0008) | 0.12 | −0.0028 (−0.0042, −0.0014) | 9.2×10 −12 | −0.0025 (−0.0039, −0.0011) | 1.3×10 −9 |
| HDLc (mmol/L) | −0.0084 (−0.0169, 0.0000) | 8.3×10−4 | −0.0133 (−0.0218, −0.0049) | 1.2×10 −7 | 0.0000 (−0.0085, 0.0084) | 0.99 | −0.0025 (−0.0109, 0.0060) | 0.32 | −0.0108 (−0.0193, −0.0023) | 1.8×10 −5 |
| LDLc (mmol/L) | 0.0078 (−0.0127, 0.0283) | 0.20 | −0.0028 (−0.0234, 0.0177) | 0.64 | −0.0263 (−0.0469, −0.0057) | 1.8×10 −5 | −0.0017 (−0.0222, 0.0189) | 0.79 | 0.0005 (−0.0202, 0.0211) | 0.94 |
| TC (mmol/L) | 0.0065 (−0.0165, 0.0295) | 0.34 | −0.0050 (−0.0279, 0.0180) | 0.47 | −0.0280 (−0.0510, −0.0050) | 4.5×10 −5 | −0.0017 (−0.0248, 0.0213) | 0.80 | 0.0046 (−0.0185, 0.0277) | 0.50 |
| log TG (mmol/L) | 0.0099 (0.0007, 0.0211) | 1.7×10−3 | 0.0170 (0.0065, 0.0275) | 6.2×10 −8 | −0.0045 (−0.0151, 0.0061) | 0.15 | 0.0041 (−0.0065, 0.0146) | 0.20 | 0.0214 (0.0108, 0.0320) | 1.2×10 −11 |
The study sample consisted of 21,168 individuals from FINRISK 1997, 2002, 2007, and 2012 cohorts. We calculated betas from linear regression models adjusted for age, sex, cohort year, genotyping batch, and the first 4 genetic principal components. P values <7.8×10−4 and corresponding betas are bolded, representing a Bonferroni correction of 16 traits × 4 components. BMI indicates body mass index; CI, confidence interval with Bonferroni corrected confidence level 7.8×10−4; CRP, C-reactive protein; eGFR, estimated glomerular filtration rate; GRS, genetic risk score; HDLc, high-density lipoprotein cholesterol; LDLc, low-density lipoprotein cholesterol; TC, total cholesterol; and TG, triglycerides.
Table 4.
Longitudinal Associations Between Hypertension Genetic Risk Scores and Incident Disease End Points.
| Obesity N Loci = 25 | Dyslipidemia N Loci = 14 | Hypolipidemia N Loci = 27 | Short Stature N Loci = 21 | SBP GRS N Loci = 1,098,015 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| End Point | Cases / Controls | Hazard Ratio (CI) | P Value | Hazard Ratio (CI) | P Value | Hazard Ratio (CI) | P Value | Hazard Ratio (CI) | P Value | Hazard Ratio (CI) | P Value |
| Hypertension | 44,472 / 153,676 | 1.07 (1.05–1.08) | 1.0×10 −40 | 1.05 (1.04–1.07) | 4.9×10 −27 | 1.15 (1.13–1.16) | 8.2×10 −178 | 1.09 (1.07–1.11) | 9.9×10 −74 | 1.43 (1.41–1.46) | <10 −300 |
| Cardiovascular disease | 23,562 / 174,586 | 1.03 (1.01–1.05) | 2.2×10 −6 | 1.04 (1.01–1.06) | 1.4×10 −7 | 1.07 (1.04–1.09) | 5.3×10 −22 | 1.05 (1.03–1.07) | 1.8×10 −14 | 1.14 (1.11–1.16) | 2.1×10 −80 |
| Coronary heart disease | 16,419 / 181,729 | 1.03 (1.01–1.06) | 3.3×10 −5 | 1.04 (1.01–1.07) | 2.0×10 −7 | 1.07 (1.04–1.10) | 5.5×10 −17 | 1.05 (1.03–1.08) | 4.6×10 −11 | 1.16 (1.13–1.19) | 9.7×10 −74 |
| Stroke | 9,819 / 188,329 | 1.03 (1.00–1.07) | 1.2×10−3 | 1.03 (0.99–1.06) | 8.5×10−3 | 1.06 (1.03–1.10) | 4.6×10 −9 | 1.05 (1.01–1.08) | 8.5×10 −6 | 1.10 (1.06–1.14) | 1.1×10 −20 |
| Heart failure | 11,454 / 186,694 | 1.05 (1.02–1.09) | 8.4×10 −8 | 1.03 (0.99–1.06) | 4.9×10−3 | 1.05 (1.02–1.09) | 2.2×10 −8 | 1.04 (1.00–1.07) | 2.1×10 −4 | 1.15 (1.11–1.18) | 2.1×10 −45 |
| Type 2 diabetes | 26,404 / 171,744 | 1.08 (1.05–1.10) | 6.1×10 −33 | 1.02 (1.00–1.05) | 1.5×10 −4 | 0.99 (0.97–1.01) | 0.21 | 1.02 (1.00–1.04) | 1.8×10 −4 | 1.12 (1.10–1.15) | 1.7×10 −76 |
| Chronic kidney disease | 2,362 / 195,786 | 1.08 (1.01–1.16) | 1.9×10 −4 | 1.04 (0.97–1.12) | 4.2×10−2 | 1.06 (0.99–1.14) | 2.8×10−3 | 1.04 (0.97–1.11) | 7.6×10−2 | 1.12 (1.04–1.21) | 5.0×10 −8 |
| Autoimmune diseases | 34,312 / 163,836 | 1.01 (0.99–1.02) | 0.35 | 1.03 (1.02–1.05) | 8.3×10 −10 | 1.05 (1.03–1.07) | 5.1×10 −21 | 1.02 (1.00–1.04) | 1.4×10 −4 | 1.01 (0.99–1.03) | 8.9×10−2 |
The study sample consisted of 198,148 individuals from FinnGen. We calculated hazard ratios per 1 standard deviation increases from Cox proportional hazards models with age as the time scale. We adjusted the models for sex, DNA sample collection year, genotyping batch, and the first 10 genetic principal components. P values <7.8×10−4 and corresponding hazard ratios are bolded, representing a Bonferroni correction of 16 traits × 4 components. CI indicates confidence interval with Bonferroni corrected confidence level 7.8×10−4; SBP GRS, genetic risk score for systolic blood pressure.
In FinnGen (N=198,148; 56% women, mean age at the end of follow-up 58 years), all hypertension component GRSs were associated with hypertension and cardiovascular disease (Table 4). Hazard ratios (HR) per 1 SD increases for hypertension ranged from 1.05 to 1.15, with the Hypolipidemia component demonstrating the strongest association (HR 1.15, CI 1.13–1.16, P=8.2×10−178). For type 2 diabetes, the HRs per 1 SD increases in GRSs were 1.08 (CI, 1.05–1.10; P=6.1×10−33) for the Obesity component, 1.02 (CI, 1.00–1.05; P=1.5×10−4) for the Dyslipidemia component, and 1.02 (CI, 1.00–1.04; P=1.8×10−4) for the Short Stature component. For chronic kidney disease, the HR was 1.08 (CI, 1.01–1.16; P=1.9×10−4) for the Obesity component. For autoimmune diseases, the HRs were 1.03 (CI, 1.02–1.05; P=8.3×10−10) for the Dyslipidemia component, 1.05 (CI, 1.03–1.07; P=5.1×10−21) for the Hypolipidemia component, and 1.02 (CI, 1.00–1.04; P=1.4×10−4) for the Short Stature component. The SBP GRS was strongly associated with all end points, except for autoimmune diseases (Table 4). All GRSs were associated with history of beta blocker, calcium channel blocker, renin-angiotensin-aldosterone system inhibitor, and thiazide diuretic use (Table VII in the Data Supplement).
Discussion
GRSs have been widely used to predict polygenic disease onset.11 For hypertension, we have previously shown that stratifying healthy individuals based on an overall GRS ranging from “low-risk” to “high-risk” could improve our ability to predict hypertension onset in the clinic.12 However, even a highly accurate overall GRS cannot capture the heterogeneity in clinical characteristics and trajectories observed in individuals with hypertension. To elucidate this heterogeneity, we sought to partition genetic hypertension risk into components that each capture distinct phenotypes in addition to hypertension. Our unsupervised data-driven analysis demonstrates the existence of 4 genetic hypertension components: Obesity, Dyslipidemia, Hypolipidemia, and Short Stature. After forming a GRS for each component separately, we assessed their longitudinal associations in 198,148 individuals and cross-sectional associations in 21,168 individuals, confirming that all 4 components captured different aspects of genetic hypertension risk. Therefore, our study underlines the potential of GRSs to not only predict disease onset, but to also give insight into disease heterogeneity and facilitate subtyping.
The strongest weighted SNPs of our genetic hypertension components reside in loci that have been previously linked to the clinical traits describing each component. The Obesity component is dominated by FTO - a central genetic driver of obesity and obesity-mediated hypertension.13,14 Similarly, the strongest weighted SNP in the Hypolipidemia component, located in SH2B3 (LNK), has strong associations to decreased LDL and total cholesterol.15 In the Dyslipidemia component, KLF14 has been shown to associate with a dyslipidemic phenotype,16 while C5orf67 has recently been implicated in the metabolic syndrome.17 In the Short Stature component, LCORL has been linked with prenatal growth, postnatal stature, and peak height velocity in infants.18,19 All of these loci are well known and have previously been scrutinized in isolation. However, the genetic classification we propose here ties them together, for the first time, into a single framework of hypertension genetics.
The literature-based clinical characteristics of our genetic hypertension components are consistent with previous epidemiological observations of hypertensive individuals. The Obesity and Dyslipidemia components followed their respective clinical definitions: the Obesity component was associated with increased BMI and waist circumference, while the Dyslipidemia component was associated with decreased HDL cholesterol and increased triglycerides. These two components are highly comprehensible because hypertension often appears together with obesity and dyslipidemia in metabolic syndrome.20 The Hypolipidemia component was associated with decreased LDL and total cholesterol, meaning that a subset of hypertension risk variants is associated with a lower level of these lipids. This surprising link between hypertension and hypolipidemia is supported by two previous observations: first, that SH2B3/LNK links hypertension with chronic inflammation20 and second, that chronic inflammation is linked with hypolipidemia.21 Indeed, variants rs3184504 (SH2B3/LNK) and rs34592089 (BANK1) in the hypolipidemia component have been linked to rheumatoid arthritis and Crohn’s disease, respectively.22,23 The Short Stature component was driven by a negative association to height. While an association between short stature and elevated blood pressure may seem unintuitive, Bourgeois et al. demonstrated a robust association between elevated SBP and short stature in a multi-ethnic cohort of 13,000 individuals.24 Not only are these 4 clinical phenotypes supported by the epidemiological literature, but our independent unsupervised analysis indicates that they may arise naturally from hypertension genetics.
In our independent study sample, the genetic hypertension components displayed clinical associations consistent with their literature-based characteristics. Cross-sectional component-trait associations in FINRISK (N=21,168) replicated all salient component characteristics (Table 3). Longitudinal associations between disease end points and hypertension components in FinnGen (N=198,148) further solidified their clinical significance (Table 4): all hypertension components were associated with both hypertension and cardiovascular disease, and the Obesity component had the strongest association to diabetes. The SBP GRS, despite being based on 50,000 times more SNPs than the hypertension component GRSs, displayed weaker associations to clinical traits and completely lacked associations to LDL and total cholesterol. Interestingly, the Hypolipidemia component had a robust association to autoimmune diseases, which was also missed by the SBP GRS and – together with the SNP-level observations discussed above - supports the role of chronic inflammation as a mediator between hypertension and the Hypolipidemia component.25
To our knowledge, only two prior studies have assessed genetic subtyping of hypertension. The first study by Ma et al.6 used a traditional subtyping algorithm, non-negative matrix factorization (NMF), to find genetic hypertension components in 1,187 hypertensive participants of the Hypertension Genetic Epidemiology Network26 (HyperGEN). HyperGEN is a cross-sectional family study aiming to identify and characterize the genetic basis of familial hypertension. Ma et al. observed two genetic components with different associations to echocardiographic measurements: component 1 was characterized by lower septal and lateral left ventricular early diastolic relaxation velocity, lower global longitudinal strain, and lower atrial strain rate than component 2. Because echocardiographic measurements are unavailable in FINRISK, we were unable to compare the component characteristics of Ma et al. to those of the current study. However, the discrepancy in the number of components could be explained by differences in methodology: Ma et al. used raw allele counts to derive genetic components, while we included multi-trait information about hypertension-related traits. The second study by Luo et al.7 proposed a hybrid NMF method for combining phenotypic and genotypic information to find hypertension components. Luo et al. used hybrid NMF on 660 hypertensive participants of the HyperGEN study and demonstrated its superiority over other similar methods (not including bNMF). While hybrid NMF is an exciting step towards more powerful subtyping methods, high-powered applications of hybrid NMF would require large study cohorts with both extensive phenotyping and genotyping, while our approach uses results from existing, publicly available GWASs.
Although our study has several strengths, such as the use of public GWAS results, an independent study sample of almost 200,000 individuals with long-term follow-up, and a subtyping algorithm suitable for polygenic diseases, its results should be viewed within the context of its limitations. First, we used GWASs based primarily on individuals with European ancestry, meaning that additional studies are needed to determine the generalizability of our genetic hypertension components to other ancestries. Second, we recognize that our genetic hypertension components are influenced by the specific choice of traits that we included into our analysis. Including more traits from better powered GWASs in the future could reveal genetic hypertension components that are outside the reach of the current study. Third, we included hypertension-associated SNPs into our analysis based on GWAS results from the UK biobank GWAS (N=340,000) instead of the meta-analysis by Evangelou et al.27 with both the UK biobank and the International Consortium of Blood Pressure (ICBP; N=760,000). We did this to avoid overfitting, because some FINRISK cohorts are included in ICBP. Finally, we accounted for linkage disequilibrium between hypertension-associated SNPs by only including independent hypertension SNPs into the analysis. Adjusting effect sizes for linkage disequilibrium instead could allow for the inclusion of more genetic variants, thus increasing the power to detect genetic components. However, we are not aware of a subtyping algorithm that can currently achieve this.
Our results suggest that the genetic basis for hypertension is a mixture of 4 broad components corresponding to distinct clinical characteristics and incident adverse outcomes. Only 73 genetic variants are required to calculate all 4 components for any given individual. While these low-resolution components already offer a proof of concept data-driven genetic risk stratification of individuals for hypertension reseach, their higher resolution refinements could benefit personalized healthcare in the future via personalized risk prediction and genetic risk counseling. To achieve this, future developments in genetic subtyping methodology should incorporate the increased power offered by modern polygenic risk scores comprising millions of SNPs.
Supplementary Material
Acknowledgments:
We acknowledge the work of Dr Jaegil Kim in developing and implementing bNMF. We thank the participants and investigators of the FinnGen and UK Biobank studies for their invaluable contributions to this work. Contributors of FinnGen are acknowledged individually in Supplemental Methods.
The following biobanks are acknowledged for delivering biobank samples to FinnGen: Auria Biobank (www.auria.fi/biopankki), THL Biobank (www.thl.fi/biobank), Helsinki Biobank (www.helsinginbiopankki.fi), Biobank Borealis of Northern Finland (https://www.ppshp.fi/Tutkimus-ja-opetus/Biopankki/Pages/Biobank-Borealis-briefly-in-English.aspx), Finnish Clinical Biobank Tampere (www.tays.fi/en-US/Research_and_development/Finnish_Clinical_Biobank_Tampere), Biobank of Eastern Finland (www.ita-suomenbiopankki.fi/en), Central Finland Biobank (www.ksshp.fi/fi-FI/Potilaalle/Biopankki), Finnish Red Cross Blood Service Biobank (www.veripalvelu.fi/verenluovutus/biopankkitoiminta) and Terveystalo Biobank (www.terveystalo.com/fi/Yritystietoa/Terveystalo-Biopankki/Biopankki/). All Finnish Biobanks are members of BBMRI.fi infrastructure (www.bbmri.fi). Finnish Biobank Cooperative -FINBB (https://finbb.fi/) is the coordinator of BBMRI-ERIC operations in Finland. The Finnish biobank data can be accessed through the Fingenious® services (https://site.fingenious.fi/en/) managed by FINBB.
Sources of Funding:
This work has been funded by the Academy of Finland (grant n:o 321351 to TN, 295741 to LL), the Finnish Medical Foundation, the Finnish Foundation for Cardiovascular Research, the Emil Aaltonen Foundation, the Hospital District of Southwest Finland, and the University of Turku.
The FinnGen project is funded by two grants from Business Finland (HUS 4685/31/2016 and UH 4386/31/2016) and the following industry partners: AbbVie Inc., AstraZeneca UK Ltd, Biogen MA Inc., Bristol Myers Squibb (and Celgene Corporation & Celgene International II Sàrl), Genentech Inc., Merck Sharp & Dohme Corp, Pfizer Inc., GlaxoSmithKline Intellectual Property Development Ltd., Sanofi US Services Inc., Maze Therapeutics Inc., Janssen Biotech Inc, and Novartis AG.
Disclosures:
VS has received a honorarium from Sanofi for consulting. He also has ongoing research collaboration with Bayer Ltd. (All unrelated to the present study).
Nonstandard Abbreviations and Acronyms
- bNMF
Bayesian non-negative matrix factorization
- GRS
genetic risk score
- GWAS
genome-wide association study
- HDL
high-density lipoprotein
- HyperGEN
Hypertension Genetic Epidemiology Network
- LDL
low-density lipoprotein
- NMF
non-negative matrix factorization
- SBP
systolic blood pressure
- SNP
single nucleotide polymorphism
References:
- 1.Cabrera CP, Ng FL, Nicholls HL, Gupta A, Barnes MR, Munroe PB, Caulfield MJ. Over 1000 genetic loci influencing blood pressure with multiple systems and tissues implicated. Hum Mol Genet. 2019;28:R151–R161. doi: 10.1093/hmg/ddz197 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Guo Q, Lu X, Gao Y, Zhang J, Yan B, Su D, Song A, Zhao X, Wang G. Cluster analysis: a new approach for identification of underlying risk factors for coronary artery disease in essential hypertensive patients. Sci Rep. 2017;7:43965. doi: 10.1038/srep43965 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Katz DH, Deo RC, Aguilar FG, Selvaraj S, Martinez EE, Beussink-Nelson L, Kim KA, Peng J, Irvin MR, Tiwari H, et al. Phenomapping for the identification of hypertensive patients with the myocardial substrate for heart failure with preserved ejection fraction. J Cardiovasc Transl Res. 2017;10:275–284. doi: 10.1007/s12265-017-9739-z [DOI] [PubMed] [Google Scholar]
- 4.Vaura FC, Salomaa VV, Kantola IM, Kaaja R, Lahti L, Niiranen TJ. Unsupervised hierarchical clustering identifies a metabolically challenged subgroup of hypertensive individuals. J Clin Hypertens (Greenwich). 2020;22:1546–1553. doi: 10.1111/jch.13984 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Mars N, Koskela JT, Ripatti P, Kiiskinen TJ, Havulinna AS, Lindbohm JV, Ahola-Olli A, Kurki M, Karjalainen J, Palta P, et al. Polygenic and clinical risk scores and their impact on age at onset and prediction of cardiometabolic diseases and common cancers. Nat Med. 2020;26:549–557. doi: 10.1038/s41591-020-0800-0 [DOI] [PubMed] [Google Scholar]
- 6.Ma Y, Jiang H, Shah SJ, Arnett D, Irvin MR, Luo Y. Genetic-based hypertension subtype identification using informative SNPs. Genes. 2020;11:1265. doi: 10.3390/genes11111265 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Luo Y, Mao C, Yang Y, Wang F, Ahmad FS, Arnett D, Irvin MR, Shah SJ. Integrating hypertension phenotype and genotype with hybrid non-negative matrix factorization. Bioinformatics. 2019;35:1395–1403. doi: 10.1093/bioinformatics/bty804 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Sudlow C, Gallacher J, Allen N, Beral V, Burton P, Danesh J, Downey P, Elliott P, Green J, Landray M, et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLOS Med. 2015;12:e1001779. doi: 10.1371/journal.pmed.1001779 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Udler MS, Kim J, von Grotthuss M, Bonàs-Guarch S, Cole JB, Chiou J, Anderson CD, Boehnke M, Laakso M, Atzmon G, et al. Type 2 diabetes genetic loci informed by multi-trait associations point to disease mechanisms and subtypes: a soft clustering analysis. PLoS Med. 2018;15:e1002654. doi: 10.1371/journal.pmed.1002654 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Eckel RH, Alberti KGMM, Grundy SM, Zimmet PZ.The metabolic syndrome. Lancet. 2010;375:181–183. doi: 10.1016/S0140-6736(09)61794-3 [DOI] [PubMed] [Google Scholar]
- 11.Lewis CM, Vassos E. Polygenic risk scores: from research tools to clinical instruments. Genome Med. 2020;12:44. doi: 10.1186/s13073-020-00742-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Vaura F, Kauko A, Suvila K, Havulinna AS, Mars N, Salomaa V, FinnGen, Cheng S, Niiranen T.Polygenic risk scores predict hypertension onset and cardiovascular risk. Hypertension. 2021;77:1119–1127. doi: 10.1161/HYPERTENSIONAHA.120.16471 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Fawcett KA, Barroso I. The genetics of obesity: FTO leads the way. Trends Genet. 2010;26:266–274. doi: 10.1016/j.tig.2010.02.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.He D, Fu M, Miao S, Hotta K, Chandak GR, Xi B. FTO gene variant and risk of hypertension: a meta-analysis of 57,464 hypertensive cases and 41,256 controls. Metabolism. 2014;63:633–639. doi: 10.1016/j.metabol.2014.02.008 [DOI] [PubMed] [Google Scholar]
- 15.Hoffmann TJ, Theusch E, Haldar T, Ranatunga DK, Jorgenson E, Medina MW, Kvale MN, Kwok P, Schaefer C, Krauss RM, et al. A large electronic-health-record-based genome-wide study of serum lipids. Nat Genet. 2018;50:401–413. doi: 10.1038/s41588-018-0064-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Small KS, Todorčević M, Civelek M, Moustafa JSE, Wang X, Simon MM, Fernandez-Tajes J, Mahajan A, Horikoshi M, Hugill A, et al. Regulatory variants at KLF14 influence type 2 diabetes risk via a female-specific effect on adipocyte size and body composition. Nat Genet. 2018;50:572. doi: 10.1038/s41588-018-0088-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lind L.Genetic determinants of clustering of cardiometabolic risk factors in U.K. biobank. Metab Syndr Relat Disord. 2020;18:121–127. doi: 10.1089/met.2019.0096 [DOI] [PubMed] [Google Scholar]
- 18.Horikoshi M, Yaghootkar H, Mook-Kanamori DO, Sovio U, Taal HR, Hennig BJ, Bradfield JP, Pourcain BS, Evans DM, Charoen P, et al. New loci associated with birth weight identify genetic links between intrauterine growth and adult height and metabolism. Nat Genet. 2013;45:76–82. doi: 10.1038/ng.2477 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Sovio U, Bennett AJ, Millwood IY, Molitor J, O’Reilly PF, Timpson NJ, Kaakinen M, Laitinen J, Haukka J, Pillas D, et al. Genetic determinants of height growth assessed longitudinally from infancy to adulthood in the northern finland birth cohort 1966. PLoS Genet. 2009;5:e1000409. doi: 10.1371/journal.pgen.1000409 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Dale BL, Madhur MS. Linking inflammation and hypertension via LNK/SH2B3. Curr Opin Nephrol Hypertens. 2016;25:87–93. doi: 10.1097/MNH.0000000000000196 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Elmehdawi RR. Hypolipidemia: a word of caution. Libyan J Med. 2008;3:84–90. doi: 10.3402/ljm.v3i2.4764 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Stahl EA, Raychaudhuri S, Remmers EF, Xie G, Eyre S, Thomson BP, Li Y, Kurreeman FAS, Zhernakova A, Hinks A, et al. Genome-wide association study meta-analysis identifies seven new rheumatoid arthritis risk loci. Nat Genet. 2010;42:508–514. doi: 10.1038/ng.582 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Liu JZ, van Sommeren S, Huang H, Ng SC, Alberts R, Takahashi A, Ripke S, Lee JC, Jostins L, Shah T, et al. Association analyses identify 38 susceptibility loci for inflammatory bowel disease and highlight shared genetic risk across populations. Nat Genet. 2015;47:979–986. doi: 10.1038/ng.3359 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Bourgeois B, Watts K, Thomas DM, Carmichael O, Hu FB, Heo M, Hall JE, Heymsfield SB. Associations between height and blood pressure in the United States population. Medicine (Baltimore). 2017;96:e9233. doi: 10.1097/MD.0000000000009233 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Duan L, Rao X, Sigdel KR. Regulation of inflammation in autoimmune disease. J Immunol Res. 2019;2019:7403796. doi: 10.1155/2019/7403796 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Williams RR, Rao DC, Ellison RC, Arnett DK, Heiss G, Oberman A, Eckfeldt JH, Leppert MF, Province MA, Mockrin SC, et al. NHLBI family blood pressure program: methodology and recruitment in the HyperGEN network. Hypertension genetic epidemiology network. Ann Epidemiol. 2000;10:389–400. doi: 10.1016/s1047-2797(00)00063-6 [DOI] [PubMed] [Google Scholar]
- 27.Evangelou E, Warren HR, Mosen-Ansorena D, Mifsud B, Pazoki R, Gao H, Ntritsos G, Dimou N, Cabrera CP, Karaman I, et al. Genetic analysis of over 1 million people identifies 535 new loci associated with blood pressure traits. Nat Genet. 2018;50:1412–1425. doi: 10.1038/s41588-018-0205-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.UK Biobank. GWAS round 2 results. Neale lab Web site. Accessed January 24, 2022. http://www.nealelab.is/uk-biobank
- 29.Pulit SL, de With SAJ, de Bakker PIW. Resetting the bar: statistical significance in whole-genome sequencing-based association studies of global populations. Genet Epidemiol. 2017;41:145–151. doi: 10.1002/gepi.22032 [DOI] [PubMed] [Google Scholar]
- 30.Myers TA, Chanock SJ, Machiela MJ. LDlinkR: an R package for rapidly calculating linkage disequilibrium statistics in diverse populations. Front Genet. 2020;11:157. doi: 10.3389/fgene.2020.00157 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Wuttke M, Li Y, Li M, Sieber KB, Feitosa MF, Gorski M, Tin A, Wang L, Chu AY, Hoppmann A, et al. A catalog of genetic loci associated with kidney function from analyses of a million individuals. Nat Genet. 2019;51:957–972. doi: 10.1038/s41588-019-0407-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Tin A, Marten J, Halperin Kuhns VL, Li Y, Wuttke M, Kirsten H, Sieber KB, Qiu C, Gorski M, Yu Z, et al. Target genes, variants, tissues and transcriptional pathways influencing human serum urate levels. Nat Genet. 2019;51:1459–1474. doi: 10.1038/s41588-019-0504-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Teumer A, Li Y, Ghasemi S, Prins BP, Wuttke M, Hermle T, Giri A, Sieber KB, Qiu C, Kirsten H, et al. Genome-wide association meta-analyses and fine-mapping elucidate pathways influencing albuminuria. Nat Commun. 2019;10:4130. doi: 10.1038/s41467-019-11576-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Dupuis J, Langenberg C, Prokopenko I, Saxena R, Soranzo N, Jackson AU, Wheeler E, Glazer NL, Bouatia-Naji N, Gloyn AL, et al. New genetic loci implicated in fasting glucose homeostasis and their impact on type 2 diabetes risk. Nat Genet. 2010;42:105–116. doi: 10.1038/ng.520 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Wheeler E, Leong A, Liu C-T, Hivert M-F, Strawbridge RJ, Podmore C, Li M, Yao J, Sim X, Hong J, et al. Impact of common genetic determinants of hemoglobin A1c on type 2 diabetes risk and diagnosis in ancestrally diverse populations: a transethnic genome-wide meta-analysis. PLoS Med. 2017;14:e1002383. doi: 10.1371/journal.pmed.1002383 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Folkersen L, Fauman E, Sabater-Lleal M, Strawbridge RJ, Franberg M, Sennblad B, Baldassarre D, Veglia F, Humphries SE, Rauramaa R, et al. Mapping of 79 loci for 83 plasma protein biomarkers in cardiovascular disease. PLoS Genet. 2017;13:e1006706. doi: 10.1371/journal.pgen.1006706 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Lee DD, Sebastian Seung H. Learning the parts of objects by non-negative matrix factorization. Nature. 1999;401:788–791. doi: 10.1038/44565 [DOI] [PubMed] [Google Scholar]
- 38.Tan VYF, Févotte C. Automatic relevance determination in nonnegative matrix factorization with the β-divergence. IEEE Trans Pattern Anal Mach Intell. 2013;35:1592–1605. doi: 10.1109/TPAMI.2012.240 [DOI] [PubMed] [Google Scholar]
- 39.Yaghootkar H, Scott RA, White CC, Zhang W, Speliotes E, Munroe PB, Ehret GB, Bis JC, Fox CS, Walker M, et al. Genetic evidence for a normal-weight “metabolically obese” phenotype linking insulin resistance, hypertension, coronary artery disease, and type 2 diabetes. Diabetes. 2014;63:4369–4377. doi: 10.2337/db14-0318 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Borodulin K, Tolonen H, Jousilahti P, Jula A, Juolevi A, Koskinen S, Kuulasmaa K, Laatikainen T, Mannisto S, Peltonen M, et al. Cohort profile: the national FINRISK study. Int J Epidemiol. 2017;47:696–696i. doi: 10.1093/ije/dyx239 [DOI] [PubMed] [Google Scholar]
- 41.Parn K, Isokallio MA, Fontarnau JN, Palotie A, Ripatti S, Palta P. Genotype imputation workflow v3.0 v1. Protocols.io Web site. Accessed January 24, 2022. doi: 10.17504/protocols.io.nmndc5e [DOI]
- 42.Sequencing Initiative Suomi (SISu). SISu Web site. Accessed January 24, 2022. http://sisuproject.fi
- 43.FinnGen. FinnGen Documentation of R5 release. Gitbook.io Web site. Accessed January 24, 2022. https://finngen.gitbook.io/documentation/
- 44.Lambert SA, Gil L, Jupp S, Ritchie SC, Xu Y, Buniello A, McMahon A, Abraham G, Chapman M, Parkinson H, et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nat Genet. 2021;53:420–425. doi: 10.1038/s41588-021-00783-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Therneau TM. Survival: a package for survival analysis in R. R package version 3.2–13. https://CRAN.R-project.org/package=survival. Accessed January 24, 2022.
- 46.Kassambara A, Kosinski M, Biecek P. Survminer: drawing survival curves using ‘ggplot2’. R package version 0.4.9. https://CRAN.R-project.org/package=survminer. Accessed January 24, 2022.
- 47.Schoenfeld D.Partial residuals for the proportional hazards regression model. Biometrika. 1982;69:239–241. doi: 10.1093/biomet/69.1.239 [DOI] [Google Scholar]
- 48.FinnGen. FinnGen R5 Phenotype Data. Accessed January 24, 2022. https://r5.risteys.finngen.fi/
- 49.Kim H, Westerman K, Kim J. Pipeline for GWAS clustering using Bayesian non-negative matrix factorization. GitHub. Accessed January 24, 2022. https://github.com/gwas-partitioning/bnmf-clustering [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



