Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2026 Mar 19;113(4):794–808. doi: 10.1016/j.ajhg.2026.02.015

Genetics of skeletal proportions across two different populations

Eric Bartell 1,2,3, Kuang Lin 4, Kristin Tsuo 1,2,5, Wei Gan 4,6, Sailaja Vedantam 2,3, Joanne B Cole 2,3,7,8,9,10,11, John M Baronas 3, Loic Yengo 12, Eirini Marouli 13, Tiffany Amariuta 14,15, Zhengming Chen 4, Liming Li 16,17,18; GIANT consortium; China Kadoorie Biobank Collaborative Group, Nora E Renthal 3, Christina M Jacobsen 3, Rany M Salem 19, Robin G Walters 4, Joel N Hirschhorn 1,2,3,
PMCID: PMC13087470  NIHMSID: NIHMS2158455  PMID: 41861830

Summary

Human height can be divided into sitting height and leg length. These measures reflect the growth of different parts of the skeleton whose relative proportions are captured by the ratio of sitting to total height (the sitting height ratio [SHR]). Height is a highly heritable trait, and its genetic basis has been well studied. However, the genetic determinants of skeletal proportion are much less well characterized. Expanding substantially on past work, we performed a genome-wide association study (GWAS) of the SHR in ∼450,000 individuals with European ancestry and ∼100,000 individuals with East Asian ancestry from the UK and China Kadoorie Biobanks, respectively. We identified 565 loci independently associated with the SHR, including all genomic regions implicated in prior GWASs in these ancestries. While SHR loci significantly overlap height-associated loci (p < 0.001), the fine-mapped SHR signals are often distinct from height signals. We also used fine-mapped signals to identify 36 credible sets with heterogeneous effects across ancestries. Lastly, we used the SHR, sitting height, and leg length to identify genetic variation acting on specific body regions rather than on overall human height.

Keywords: sitting height ratio, height, growth, body proportion, GWAS, genome-wide association study, fine-mapping, gene set enrichment analysis, gene prioritization, trans-ancestry


Bartell et al. identify variation at hundreds of genetic loci that are associated with the sitting height ratio, a measurement of body proportion. Many are distinct from genetic associations with height. Additionally, a small number of these genetic signals have differing effects when measured in European- vs. East Asian-ancestry samples.

Introduction

Human height has long been studied as a model polygenic phenotype due to its high heritability,1,2,3,4,5 polygenicity,6 and ease of accurate measurement. Recent work from the GIANT consortium, combining data from 5.4 million individuals across 5 major population groups, maps genomic regions that account for nearly all of the heritability attributable to common variation (h2variant, estimated to be approximately 51%), identifying over 12,000 signals associated with height.5

Standing height reflects the sum of the sizes of different skeletal components, whose relative proportions vary across individuals and populations. Despite the genetics of height being well studied, much less is known about the genetics of body proportions.7,8 Here, we hypothesize that investigating the genetics of body proportion will provide additional insights into our understanding of human skeletal growth and skeletal growth disorders. Furthermore, as height is frequently used as a model polygenic trait, parsing height into its components can serve as a model of genetic investigations into a set of polygenic traits with overlapping underlying biology.

We have previously studied the sitting height ratio (SHR), the ratio of sitting height (SH) to standing height,7 as well as protein levels for growth-related genes, to begin to understand the underlying biology at selected height loci.9 Our prior genome-wide association study (GWAS) of the SHR,7 in 3,545 African American individuals and 21,590 individuals of European (EUR) ancestry, identified 6 genome-wide significant associations, including the locus encoding the growth-related protein IGFBP-3. The substantial heritability of the SHR (estimated as 26% previously7) suggests that a larger sample size, combined with techniques such as statistical fine-mapping for multiple height-related traits, could improve the identification of causal variants and improve our understanding of associations with height and height-related phenotypes. Furthermore, fine-mapping has not been performed to date on height-related traits, which could help further define the specific consequences of causal variants, especially in the scenario where there are multiple genetic associations with distinct phenotypic consequences in the same locus.9,10

The vast majority of genetic research has focused on individuals of EUR ancestry, although in recent years, GWASs involving non-EUR ancestries have become more frequent, with increasing sample sizes.5,11,12,13,14 Polygenic trait and disease biology is likely shared across ancestries, as demonstrated by estimated genome-wide genetic correlations across ancestries that are often consistent with unity—including for height.15 Thus, expanding the range of studies in non-EUR ancestries is important not only to improve the predictive power and thus the clinical utility of GWASs in populations with diverse ancestries but also to meta-analyze genetic associations from diverse populations to better define the location and population specificity of causal variants.16 In previous work,16 height-associated loci have been fine-mapped in both EUR and non-EUR ancestries. However, the presence of multiple signals at each locus, combined with differing linkage disequilibrium (LD) patterns across ancestries, made it challenging to confidently identify genetic associations with differential effects across ancestries.

Here, we perform a GWAS and meta-analysis of several height-related traits—the SHR, leg length (LL), and SH—in large biobanks from the United Kingdom (UK Biobank [UKB]) and China (China Kadoorie Biobank [CKB]). We evaluate genetic correlations between phenotypes and ancestries and identify biological processes underlying these traits using gene set enrichment analysis (GSEA). We perform fine-mapping in the UKB to identify 95% credible sets (CSs), enabling us to better characterize height signals based on their relationships with signals for other skeletal phenotypes and also to identify height CSs present in the UKB with evidence of significantly different effects in the CKB. Together, these results provide additional insights into the genetics and biology of skeletal growth and exemplify a set of possible approaches for evaluating GWASs from groups of related phenotypes and multiple ancestries.

Material and methods

GWAS in UKB

The GWAS was performed on phenotype height, SH, LL (height − SH), and SHR (SH/height) using 206,812 males and 245,109 females of EUR ancestry from the UKB, with over 77,753,733 autosomal variants passing quality control (QC) filters. Phenotypes were adjusted for age, age2, and sex (for joint analyses of males and females). The SHR GWAS was performed in multiple ways, either using the same covariates as other phenotypes or additionally correcting for BMI or height and BMI. Genotyping array number and 10 genetic principal components were additionally included as covariates in the GWAS. The GWAS was performed using BOLT-LMM v.2.3.4. Additionally, GWASs of independent halves of the UKB (n = 188,575 and 188,597), the UKB size matched with the CKB (n = 72340, 72321) subset from the halves, and thirds of the UKB (n = 125,740, 125,735, and 125,748) were performed and were used in supplemental polygenic risk score (PRS) analyses. The individual-level genotype and phenotype data are available from the UKB by application at https://www.ukbiobank.ac.uk/. The work on UKB was reviewed and deemed exempt by the Boston Children's Hospital Institutional Review Board (IRB).

SHR EUR meta-analysis

The UKB GWAS results were meta-analyzed with previously published SHR genome-wide association data. The previous GWAS of the SHR used data from six studies, for a combined sample size of 21,590 individuals of EUR ancestry and 3,545 African Americans. Imputation for all six cohorts was performed using a 1000 Genomes reference panel, and only variants with an information score of >0.6 were included in the meta-analysis. From the UKB study, variants with an information score of >0.3 were used. METAL v.2011-03-2517 was used.

GWAS in CKB

The CKB (https://www.ckbiobank.org) is a population-based prospective cohort. Blood samples of adult subjects were collected from 10 geographically defined regions of China in 2004–2008. Hard genotyping was conducted on custom Affymetrix Axiom arrays, with 511,885 variants passing QC. Genotypes were phased using SHAPEIT3 v.4.12 and imputed into the 1000 Genomes Phase 3 reference with IMPUTE4 v.4.r265 as described previously.18

The four traits are standing height, SH, LL (standing height – SH), and SHR (SH/standing height). Across the 20 regions by sex strata, the SHR was linearly regressed against age, age2, year of birth (as a categorical variable), height, and BMI. The three other traits were regressed against age and age2. The residues were rank-based inverse normal transformed and used for the GWAS.

The CKB genotype data were split into 5 different sets. Set A is the full set of 95,674 subjects. Set B is the relative-free subset of set A. The individuals were selected via the command “plink2 –king-cutoff 0.05.” The set has 72,471 subjects. Set C is the relative-free subset of set A − B. Again, the samples were selected via the command “plink2 –king-cutoff 0.05.” Set C has 15,620 subjects. Set D is a randomly selected subset of set B. Its size is 36,235. Set E is the set B – D, with 36,236 subjects.

We used BOLT-LMM v.2.3.2 for the GWAS of the transformed residuals. The covariates were the genotype array version and the sample disease ascertainments. The fixed-effects meta-analysis across 20 regions, stratified by gender, was performed using METAL.

Ethical approval was obtained from the Oxford Tropical Research Ethics Committee, University of Oxford (UK, 025-04); the Ethical Review Committees of the Chinese Centre for Disease Control and Prevention (Beijing, China 005/2004), Chinese Academy of Medical Sciences; and the Institutional Review Board (IRB) at Peking University.

Trans-ethnic fixed-effects meta-analysis

Fixed-effects meta-analysis was performed between GWASs performed in the UKB and CKB using METAL. Variants included in only 1 of the 2 populations were included in the reported summary statistics. 473,511 EUR subjects and 72,473 CKB subjects were combined for a total sample size of 545,984.

GSEA for multiple traits in each population

Multi-marker analysis of genomic annotation (MAGMA) analysis (v.1.09) was performed using GWAS summary statistics calculated in the UKB and separately in the CKB based on reconstituted gene sets (data-driven expression prioritized integration for complex traits [DEPICT]19). A Bonferroni-corrected significance threshold (corrected for the number of gene sets) was used to identify significant gene sets at p < 0.05/14,462. 1000 Genomes was used as the reference panel for gene significance, using East Asian (EAS) and EUR ancestries for the CKB and UKB, respectively.

GWAS subsets

GWASs of multiple sizes were performed as shown in Table S7.

Lead variant proximity

The proximity of lead variants identified for each phenotype was compared. A null distribution was generated by matching each lead variant with up to 1,000 null variants using the SNPSnap software,20 matching on default parameters of minor-allele frequency (MAF; ±5%), gene density (±50%), distance to nearest gene (±50%), and LD buddies (±50% using an r2 > 0.5).

Height-relevant gene sets

Height-relevant gene sets from Online Mendelian Inheritance in Man (OMIM) and Blueprint Genetics were used to validate fine-mapping results. Height-relevant genes from OMIM were manually curated.2,5,21 The short-stature genetic testing panel from Blueprint Genetics was used,22 and additionally, genes were classified by their clinical indications as acting in a body-region-proportionate (changes in height associated with gene dysfunction tend not to specifically target the upper or lower body) or body-region-disproportionate (changes in height associated with gene dysfunction tend to act primarily in one body region) way. Syndrome implications were classified using omim.org.

Round zone specificity

Mouse gene expression specificity to different parts of the growth plate was calculated as previously described.23 Human orthologs for mouse genes were identified using the BioMart package, Ensembl build 105. CS scores for expression specificity were identified for CSs overlapping a gene; averages were used when overlapping multiple genes. Human genes mapping to multiple mouse used averages.

Winner’s curse correction

Variants passing genome-wide significance were identified in summary statistics from the GWAS performed in half of the UKB or half of the CKB; the effect size of these variants was shown to be inflated, mostly near the significance threshold when compared to the unrelated second halves (Figure S21). As such, effect sizes were corrected for the winner’s curse.24,25 Optimal correction was achieved when using linear regression p values from BOLT-LMM. After showing the efficacy of winner’s curse correction, the method was applied to the full UKB and CKB GWASs.

Fine-mapping

Fine-mapping was performed using sum of single effects (SuSiE) v.0.8.0 on 1 megabase pair (Mb) windows surrounding each lead variant for unrelated GWASs of phenotype height, SHR, SH, LL, and SHR unadjusted for either height or BMI (SHR-unadjusted) measured in the UKB. The parameters maximum number of non-zero effects L = 20 and maximum iterations max_iter = 1,000 were set, and SuSiE was used to estimate scaled prior and residual variances. LDstore v.2.0 was used to calculate LD in the 1 Mb windows from unrelated EUR individuals in the UKB.

For overlapping 1 Mb windows, additional analysis was performed. Overlapping windows were merged, and fine-mapping was performed on the entire combined region jointly. In order to save on computational time and space, many variants were removed; variants were only included in the model if either their p value for the GWAS of interest was less than 0.01 or if the variant was identified as being a part of a CS in the 1 Mb fine-mapping analysis. As a result, 23.1% of variants were included in fine-mapping analyses of merged regions greater than 1 Mb. A small number of fine-mapping analyses using all variants of a merged region were performed and compared with fine-mapping performed with the limited variant set, with high concordance of results (see Note S2 and Figure S11).

Effect-size comparison

Effect sizes at lead variants in the CKB and UKB were compared. Winner’s-curse-corrected trait-increasing effect sizes for UKB and CKB lead variants were compared with effect sizes for the variants calculated in the second study, and their relationship was quantified using least-squares regression and Deming regression. Cochran’s Q and I2 were calculated for lead variants identified in either the UKB or CKB by comparing winner’s-curse-corrected effect sizes with the variant’s effect size in the second study. I2 vs. Cochran’s Q p values are shown in Figures S20B and S20C.

Fine-mapped CSs were evaluated for ancestral heterogeneity (Figure 3B). If all variants in a CS were heterogeneous at Bonferroni-corrected significance (Cochran’s Q, p < 0.05/CSs) when comparing betas estimated in the UKB and CKB, then the CS was classified as heterogeneous.

Figure 3.

Figure 3

Fine-mapping to identify ancestral heterogeneity

(A) Cross-ancestry genetic correlation as evaluated using Popcorn between GWASs derived from the UKB and CKB across phenotype height, SH, LL, and the SHR adjusted for height and BMI. Standard errors for estimates are shown. Estimates of the genetic impact correlation (⍴g) for each phenotype are not significantly different from 1.

(B) Schematic of heterogeneous credible set (CS) identification.

(C) Variants in heterogeneous CSs identified for height. The UKB GWAS is fine-mapped using the sum of single effects (SuSiE) model, and effect estimates for variants in credible sets were compared with estimates from the CKB using Cochran’s Q. CSs that were entirely composed of heterogeneous variants were nominated as heterogeneous credible sets, and their standard errors shown.

(D) For height, limiting to most simple CSs (maximum CSs per locus = 1 [top] and maximum variants per CS = 1 [bottom]) and then relaxing each constraint modifies the threshold for heterogeneity significance of CSs as multiple testing increases. Cross-ancestry heterogeneous CSs were only observed in loci with multiple credible sets.

(E) Locus where fine-mapped variant effects differ between the CKB and UKB. Effects are measured using X2/N, a sample-size-invariant heuristic. Variants in credible sets are marked with colored circles, and variants in the 3 heterogeneous credible sets are marked with triangles.

Trans-ancestry genetic correlation

Popcorn15 v.0.9.9 was run on each phenotype to identify trans-ancestry genetic correlations. The “fit” command using pre-computed 1,000 genome LD references for each of the EUR and EAS ancestral groups was run.

PRS in UKB and CKB

PRS estimation was performed using the pruning-and-thresholding (P+T) framework implemented in PRSice26 v.2.2.12. PRSs were derived for phenotype height, SHR, SH, LL, and SHR-unadjusted and performed independently for the CKB and UKB. The GWAS of half of the individuals from the study was used for effect sizes; individuals from an independent, unrelated quarter were used to find the best-fit p threshold, and prediction accuracy was evaluated in the final unrelated quarter. Variants with an INFO < 0.3, a MAF < 0.005, or ambiguous alleles were filtered.

Proxy-informed trans-ethnic PRSs

Trans-ethnic PRSs were generated to improve prediction using the P+T framework as implemented in PRSice.26 The following algorithm (shown in Figure S23B) was used to generate the list of variants used for prediction: the GWAS was performed in study 1, half 1; naive PRSs used direct GWAS summary statistics. Variants were filtered based on the MAF and INFO in both studies (half 1 from each study); this list of variants was used for initial, filtered PRSs. The variant list was then clumped by significance, and LD proxies with an r2 > threshold 1 to a lead variant were identified using study 1 ancestry. LD proxies with an r2 > threshold 2 to this new list were next identified using study 2 ancestry. This final variant list was clumped and PRSs generated using p values from study 2, half 1. LD information from the UK10K study was used for UKB proxy identification steps, and LD information from 1KG EAS was used for CKB proxy identification steps. RSID identification was performed using hg19 dbSNP 151,27 and all ambiguous variants (A/T or G/C variants) were removed from the analysis prior to clumping. For analyses where threshold 2 = 1, the second proxy identification step was skipped; for analyses where threshold 2 = “window,” boundaries were defined using variants from the first proxy identification step, and all variants within these boundaries were included. Independently, for analyses where threshold 1 = “1 Mb” or another distance, lead variants were identified, and nearby variants within the described window (i.e., 1 Mb) were included for study 2 pruning.

Results

GWAS of skeletal growth phenotypes in larger numbers of individuals of largely EUR ancestry recapitulates prior results and reveals associations

We have previously performed a GWAS meta-analysis of the SHR (Figure 1A)7 of 21,590 subjects of EUR ancestry. To more fully describe the genetic basis of the SHR, we extended this study to include 451,921 individuals of EUR ancestry from the UKB (our analyses of individuals from the CKB of largely EAS ancestry are described in subsequent sections). We performed genome-wide association in the individuals from the UKB, adjusting the SHR for BMI and standing height (as in our prior work) along with other standard covariates (see material and methods). We identified 618 non-overlapping loci by pruning signals within 500 kilobase pair (kb) windows and merging overlapping windows (Figure 1B; GWAS summary statistics are available in the GWAS Catalog [GWAS Catalog: GCST90728584, GCST90728585, GCST90728586, GCST90728587, GCST90728588]). The observed genomic inflation factor was consistent with previous studies of polygenic traits at this sample size, and the LD score regression (LDSC) ratio reflects that our study has acceptable levels of inflation (p < 5 × 10−8, λGC = 1.92, ratio = 0.125 for the SHR; Figure S1). Briefly, the LDSC ratio is calculated as the regression intercept − 1 divided by the average chi-squared statistic − 1; this value estimates the fraction of GWAS inflation that is not due to polygenic heritability.28

Figure 1.

Figure 1

Basic GWAS and simple comparisons

(A) Schematic illustrating the phenotypes studied in this manuscript, including height, sitting height (SH), leg length (LL), and sitting height ratio (SHR), which is SH/height; for most analyses, the SHR is adjusted for height and BMI.

(B) Manhattan plot of the meta-analysis of GWAS results for SHR from UK Biobank (UKB), China Kadoorie Biobank (CKB), and our previous study.7 The −log10 of p values for the association of variants with the SHR is plotted, with variants ordered by chromosome and position. Genome-wide significance (p < 5 × 10−8) is indicated by the dashed line. Variants reaching genome-wide significance that are also within 35 kb of a height lead variant (as defined in Yengo et al.5) are indicated by red dots, and genome-wide significant variants falling outside those windows are indicated by blue dots.

(C) Phenotypic (above the diagonal) and genotypic (below the diagonal) correlations for height-related traits are shown; positive correlations are red, and negative correlations are blue. Values are estimated in both the UKB (top value) and CKB (bottom).

(D) Each row is a gene set identified by enrichment analysis (GSEA) of the SHR using the multi-marker analysis of genomic annotation (MAGMA) algorithm (top 2.5% of prioritized gene sets), and each column is a gene prioritized by the Polygenic Priority Score (PoPs) algorithm (top 2.5% of prioritized gene sets). The squares in the heatmap indicate the likelihood of membership of the gene in the corresponding gene set (gene-pathway Z score supplied by the data-driven expression prioritized integration for complex traits [DEPICT] reconstituted gene sets19). Row and column annotations in the left and top margins indicate whether the gene set or gene, respectively, was also prioritized by similar analysis of each of the other height-related phenotypes (more darkly colored indicates more strongly prioritized).

(E and F) Zoomed regions show clusters of related prioritized genes and gene sets: (E) those more prioritized by SH and LL than by height (indicated by lighter blue squares in the margins) and, similarly, (F) those prioritized across phenotypes (indicated by dark blue, green, and orange squares in the margins).

We meta-analyzed these results with results from Chan et al.7; this EUR-specific meta-analysis of 473,511 individuals from 7 studies identified 553 non-overlapping SHR-associated loci that reached genome-wide significance. Of the five loci that reached genome-wide significance in the previous GWAS of EUR-ancestry individuals, four were genome-wide significant in the UKB and replicated in the meta-analysis (rs6931421, rs1722141, rs882367, and rs228836), and one was only nominally significant in the UKB (rs140449984, p = 4.1 × 10−3) and therefore showed only suggestive statistical evidence in the meta-analysis (p = 2.5 × 10−5). The association at variant rs140449984 was primarily driven by one of the six cohorts (ARIC) in the previous GWAS.

To evaluate the overlap of other skeletal growth phenotypes related to the SHR, we performed a GWAS in the UKB of SH, LL, and height. Additionally, we performed a GWAS of SHR-unadjusted, primarily to enable a sensitivity analysis that tested for evidence of collider bias induced by the adjustment for height, discussed below. Collectively, we identified 705 SH loci, 724 LL loci, 765 height loci, and 643 SHR-unadjusted loci. We observed a few instances of heterogeneity in a sex-stratified GWAS (Figure S3) but focused our analyses on the GWAS of both sexes combined. To help frame our expectations of the relationships between GWAS results from these skeletal growth phenotypes, we evaluated the genetic correlations between height, LL, SH, SHR-unadjusted, and SHR. As expected from the strong phenotypic correlation, we observed strong genetic correlations28 between height and either LL or SH (genetic correlation [rg] = 0.91 [SE = 0.004] for LL and 0.84 [SE = 0.0071] for SH). The estimated genetic correlation between height and SHR-unadjusted is lower than for LL and SH but still strongly significant (rg = −0.39, p = 2.3 × 10−46); the negative correlation is consistent with the stronger correlation of height with LL than with SH. After adjustment of the SHR for height and BMI, the estimated genetic correlation with height is not significantly different from zero (rg = −0.032, p = 0.35), suggesting that this adjusted phenotype provides a genetically independent assessment of an aspect of skeletal growth and that genetic associations from the (adjusted for height and BMI) SHR GWAS could reveal distinct biological mechanisms of skeletal development. Other relationships between pairs of skeletal growth phenotypes are shown in Figure 1C.

Because of the strong genetic and phenotypic correlations between height, SH, and LL, we expected that many alleles would have consistent directions of effect on these three height-related phenotypes. Indeed, we observed a strong directional concordance between associations of height, SH, and LL in our results and in prior GWAS analyses of height5 (Figures S2A–S2C). However, concordance between height and the SHR was lower, consistent with the low genetic correlation between these two traits. Of note, the low correlation between effect-size estimates for the SHR and height (r2 = 0.01; Figure S2D) provides initial evidence that collider bias does not strongly influence the SHR GWAS results.

Genetic and phenotypic comparison of skeletal growth phenotypes

We next sought to estimate the heritability of the SHR in our data and to evaluate the shared genetic component between our phenotypes via genome-wide genetic correlation analysis. To this end, we used LDSC29 to estimate the heritability of the SHR. The variance explained by autosomal variants in the 473,511 individuals of EUR ancestry is 0.29 (SE = 0.013; Figure S4A), which is not significantly different from the estimate from the previous smaller study of the SHR (heritability [h2] = 0.26, SE = 0.026). Heritability estimates for both ancestries can be found in Figure S4A.

To further investigate the genetic relationship between growth-related phenotypes and height, we evaluated the proximity of lead variants for different traits and for height. As expected, given the phenotypic overlap and the strong genetic correlation, lead variants for SH and LL were highly likely to be near height lead variants, much more so than null variants (matched using SNPSnap20; see material and methods; p < 0.001 in 1,000 permutations; Figures S5A and S5B). However, despite there being essentially no genetic correlation between the SHR and height, lead variants associated with the SHR were also significantly more likely than null variants to be near height lead variants (p < 0.001 in 1,000 permutations; Figure 2A). Interestingly, despite an expectation that associations would be more strongly associated with height than the SHR (as height is more heritable than the SHR), variant associations with the SHR were sometimes more strongly associated than with height at a particular height-associated locus. As an extreme example, the variant chr4:103188709C>T (rs13107325, a lead variant for multiple traits, marked with the red box in Figure S6) is more significantly associated with the SHR than height; this is due to the fact that the alternate allele is associated with a decrease in SH (p = 1.7 × 10−36) and, surprisingly, a twice-larger increase in LL (p = 9.5 × 10−132). Instances where we observe stronger SHR associations likely point to loci that have disproportionate and unexpected (as per phenotypic and genotypic correlations) effects on body proportion. In addition, even when only considering SHR associations near height signals (within 35 kb), 14% of associations have low LD with the nearby height signals (r2 < 0.2 with height lead variant; Figure S7A), implying that these shared genetic loci associated with the SHR and height harbor associations with skeletal growth phenotypes independent from those having effects on height.

Figure 2.

Figure 2

Fine-mapping across phenotypes

(A) Number of lead variants (red bar) within 35 kb of height signals of association (COJO lead variants5), compared with distribution of 1,000 control sets of random variants matched using SNPsnap (histogram).

(B) Examples of height CSs that are “far” (left), “near, shared” (middle), and “near, distinct” (right) from SHR CSs. Colors mark variant inclusion in a credible set (CS), and X marks a variant in a CS for both phenotypes. Regional genes are annotated; “” indicates that SLC39A8 is an OMIM gene.

(C) Relationship between height and SHR CSs is plotted with distance (base pairs [bp]) on the x axis and LD (measured with Pearson r2) on the y axis. Each point corresponds to one height CS. For CSs containing more than one variant, the minimum distance and maximum LD between variants in the height CS and nearby SHR CSs are plotted. Colors mark point density. CSs are classified by their proximity as “near, shared” (282 in the top left box, LD between CSs of >0.8 and distance of <30 kb), “near, distinct” (189 in the bottom left box, LD of <0.2 and distance of <30 kb), or “far” (1,566 in the bottom right box, LD of <0.2 and distance of >100 kb); the remainder are “unclassified” (525).

(D) For height signals proximal to SHR signals, associations with SH and LL are plotted, with the LL Z score on the y axis and the SH Z score on the x axis. “Near, shared” CSs are plotted in red, “near, distinct” are in black, and “far” are in blue. 260/282 “near, shared” signals maintain significance for association with the SHR after Bonferroni correction.

(E) Overlap between the regions spanned by two different categories of height CSs (“near, shared” and “far”) and categories of genes underlying skeletal growth disorders is shown. Black bars indicate 95% confidence intervals. OMIM: manually annotated list of genes implicated in short stature derived from the OMIM database; short stature: growth disorders and skeletal dysplasia panel used for genetic testing from Blueprint Genetics.

Because the SHR in our analyses is adjusted for height, we considered the possibility that the proximity of height and SHR signals could be due to collider bias induced by this adjustment. However, variants associated with SHR-unadjusted were located near height associations more often than those associated with the SHR (adjusted for height and BMI). Under the expectation of collider bias due to adjustment of the SHR for height, one would expect associations with SHR (adjusted for height and BMI) to be more proximal to height associations than SHR-unadjusted associations, but the trend is actually in the opposite direction (one-sided Mann-Whitney U p = 0.95). This result further suggests that the proximity of SHR signals and height signals is not due to collider bias (Figure S7B). In addition, we constructed PRSs for the SHR. Notably, the PRS of the SHR (adjusted for height and BMI) is a far better predictor of the SHR (adjusted for height and BMI) than a PRS constructed for height (r2 = 0.137 vs. r2 = 0.001), consistent with SHR (adjusted for height and BMI) and height being two distinct phenotypes rather than being related largely through collider bias (Figure S8). As we sought to investigate genetic mechanisms of distinct growth-related phenotypes, the remaining analysis of the SHR will focus primarily on the SHR (adjusted for height and BMI) rather than SHR-unadjusted, as the latter is significantly genetically correlated with height and therefore more redundant with height.

Understanding biological pathways underpinning shared and distinct mechanisms of growth-related phenotypes

To explore the genes and biological pathways underlying skeletal growth-related phenotypes, GSEAs were performed using MAGMA30 applied to the meta-analyses of the UKB and CKB (see material and methods). In line with observational and genetic correlation results, gene-set significance for LL, SH, and height were highly overlapping; height and LL gene-set significance were qualitatively the most similar (Figure S9; Table S2). Of note, the SHR and height also shared enriched pathways, despite being genetically uncorrelated, consistent with the proximity of SHR and height signals. Examples of enriched pathways shared by all four phenotypes include “ossification” and “decreased length of the long bones.” Interestingly, GSEA of the SHR also prioritized spine morphology-related pathways (e.g., “abnormal cervical vertebrae morphology” and “abnormal sternebra morphology”), which were not as clearly prioritized for height. We also applied Polygenic Priority Score (PoPs),31 which prioritizes genes based on similarity to genome-wide enrichment (Table S1). PoPs prioritized multiple clearly growth-related genes (e.g., GHR, COL1A1, and ACAN) (Figures 1D and F) for all four phenotypes. We also sought to highlight biological factors specifically affecting skeletal proportion but not height; mechanisms affecting only the SHR would need to have balancing effects on SH and LL in order to have no effect on height. We observed that a cluster of HOX genes was more strongly prioritized for the SHR and SH than for height or LL (Figures 1D, 1E, and S10) and were prioritized for pathways relating to vertebral morphology and branching morphogenesis. When applied to genes selectively prioritized by the SHR but not height, Gene Ontology (GO) identified “negative regulation of animal organ morphogenesis” (GO: 0110111, p = 3 × 10−4) and “negative regulation of Wnt signaling pathway involved in dorsal/ventral axis specification” (GO: 2000054, p = 3 × 10−4), among other relevant pathways (prioritized list in Table S3). Overall, GSEA of skeletal growth phenotypes primarily concluded that broad-strokes biology is shared between phenotypes, and further method development and analysis are needed to clearly articulate differences.

Genetic fine-mapping of associations in UKB reveals over 1,300 causal loci for skeletal proportion phenotypes

Our observations suggest that, within associated regions shared between the SHR and height, distinct variants in low LD affect each phenotype. To evaluate this further, we performed statistical fine-mapping to more precisely examine the independent effects of associated variants on these related phenotypes in shared loci. Due to limitations of existing fine-mapping methods that have been shown to result in substantial miscalibration when applied to meta-analyses,32 fine-mapping was performed by applying SuSiE33 (see material and methods) to UKB GWAS summary statistics, identifying 95% CSs. After merging overlapping loci (where lead variants were within 1 Mb of each other), the 618 SHR-associated, 765 height-associated, 705 SH-associated, and 724 LL-associated loci had average sizes of 2.81, 2.19, 2.5, and 2.5 Mb, respectively. To computationally enable fine-mapping even at >1 Mb loci containing many variants, analyses were limited to a subset of variants per locus (those with p < 0.01 or identified as part of a CS in an analysis of an adjoining 1 Mb window; see material and methods and Figure S11). For the SHR, there were on average 62,021 variants per locus, of which an average of 15,064 (24%) per locus were considered in the fine-mapping model. Across all loci, an average of 56 variants were members of 2.2 CSs per locus, for a total of 1,327 CSs (Figure S12). Smaller CS sizes indicate greater confidence in the causal variant; from the 1,327 CSs identified, 645 (49%) contained 5 or fewer variants. Loci with many CSs are more complex, suggesting many independent effects acting on the SHR, while loci with fewer CSs are simpler, with likely fewer independent effects; the majority of loci (460) contained 5 or fewer CSs. Corresponding statistics for fine-mapping of each of the other traits can be found in Table S4, and full variant-level summary statistics can be found in Table S5. Of the fine-mapped SHR loci, 17 contain only one CS comprising a single variant; 3 of these variants are annotated as coding, in SCARF2, MTMR11, and FBXL14. Interestingly, rs9680797 is a nonsynonymous variant (P661L) in SCARF2, which underlies Van den Ende-Gupta syndrome (VDEGS34; characterized by multiple skeletal phenotypes). Notably, this approach to find simple loci containing one or a few associated coding variants identifies strong links between variant/gene and phenotype but misses other important biology where the fine-mapping signals are more complex or will highlight noncoding variants.

Comparative fine-mapping analysis reveals heterogeneity and shared effects across growth phenotypes

The CSs for two genetically uncorrelated skeletal phenotypes—height and the SHR—enabled us to more accurately classify the fine-mapped height signals based on their proximity to (or overlap with) CSs for the SHR. As shown in Figure 2C, some height CSs were identified both in close proximity to (<30 kb) and sometimes overlapping SHR CSs, with either high (r2 > 0.8) or low (r2 < 0.2) levels of LD. By contrast, other height CSs were relatively distant (>100 kb) from SHR CSs in genomic distance, with low LD. Using these metrics, we classified groups of CSs into “near, shared,” “near, distinct,” “far,” and “unclassified” (see material and methods). As a specific example of a “near, shared” CS on chromosome 4 (Figure 2B, middle), a CS overlapping the SLC39A8 gene is shared between height and the SHR (and also SH and LL). As expected from the shared effect on height and the SHR, this specific CS exhibits opposing effects on SH and LL. Specifically, the allele of the single-variant CS associated with increasing height is also associated with an increase in LL but a (smaller) decrease in SH, leading to a decreased SHR. Indeed, most height CSs in the category of “near, shared” with the SHR that also affect SH and LL have disproportionate or (occasionally) opposing effects on SH and LL, despite the observation that genome-wide SH and LL are highly genetically correlated (Figures 2D and S14; see material and methods). In contrast with the SLC39A8 locus, associations near TET2 exemplify a locus with a height CS near but distinct from an SHR CS (Figure 2B, right). Association testing conditioning on the SHR CS yielded an attenuated but significant association between the height CS SNP and height (rs143847362 p = 3.7 × 10−70 before and 1.1 × 10−18 after conditioning on the SHR CS). Association signals can be both in LD but distinct: in this locus, there are distinct conditionally independent associations with these two growth-related phenotypes. SH and LL CSs that are similarly classified by their proximity are visualized in Figure S15.

We then sought to characterize the proximity of different categories of height-associated CSs to height-related genes. To do this, we used two manually curated gene lists (see material and methods): height-related genes from the OMIM database as described in Lango Allen et al.,2 Yengo et al.,5 and Lui et al.21 and genes from commercially available diagnostic tests for genetic causes of short stature and skeletal disorders.22 Height CSs “near, shared” to SHR CSs more often overlapped height-related genes than height CSs “far” from SHR CSs (Figure 2E); this observation persists even when subsetting “far” CSs to match “near, shared” GWAS association p values with height (Figure S16).

The “genetic testing” genes were further classified by whether they underlie syndromes of short stature with proportional effects on the upper and lower body (e.g., growth hormone insensitivity35,36) or disproportional effects (e.g., spondyloepimetaphyseal dysplasia37) (see material and methods and Table S6). Because the SHR is a measure of skeletal proportion, we considered whether the overlap with short stature was driven by genes underlying syndromes with disproportionate short stature. We compared the number of genes that overlapped “near, shared” and “far” CSs in these 2 groups of short-stature genes. We hypothesized that CSs unique to height with no association with the SHR would predominantly be discovered near genes that affect growth in a proportional manner, and indeed, fewer of the “far” CSs, relative to the “near, shared” CSs, overlap genes related to disproportionate short stature than overlap short-stature proportionate genes; however, this observation did not reach statistical significance (chi-squared test, p = 0.09; Figure S17). Notably, “near, distinct” CSs were not significantly different from “near, shared” or “far” CSs in their relationships with OMIM genes or proportionate/disproportionate short-stature genes, so further work will be needed to understand how “near, distinct” CSs might represent distinct biology.

Our previous work had shown that genes implicated by height GWASs showed expression specificity for the round layer of the growth plate.23 We tested if this observation replicated among our categories of fine-mapped CSs. Genes overlapping either height CSs “near, shared” to SHR CSs or height CSs “far” from SHR CSs had higher specificity of expression in the round zone than other expressed genes (unpaired t test p = 2.1 × 10−13 and 3.2 × 10−8, respectively). Furthermore, these “near, shared” CSs had significantly higher round zone expression specificity than the “far” CSs (p = 2.4 × 10−4; Figure S18). These findings were specific to the round zone; due to the definition of specificity, “near, shared” CSs had proportionately lower specificity for expression in the hypertrophic and flat zones than “far” CSs (p = 0.008 for round zone vs. p = 0.03, 0.25 for hypertrophic and flat zones, respectively; Figure S18). The same phenomenon was observed previously by de Leeuw et al.,30 who showed, using conditional analysis, that specificity for the round layer was driving correlation with GWAS associations. Thus, by multiple measures—proximity to height-related genes and genes expressed more specifically in the round zone—these “near, shared” height loci provide more relevant areas of focus for further interrogation of the biology of human growth.

Comparison of GWAS of skeletal phenotypes across ancestries indicates largely shared overall genetic architecture

We sought to understand the extent to which the genetic architecture of growth-related phenotypes is shared or differs between populations. We observed that the population means in the UKB are statistically significantly different between ancestries for each of the skeletal growth phenotypes, consistent with prior studies.7,38,39 For example, individuals with EAS ancestry had a raw mean SHR of 0.539, in contrast to a mean of 0.530 among individuals with EUR ancestry (unpaired t test p = 4 × 10−97 after adjusting for covariates; Figure S19A; comparisons for other phenotypes are shown in Figures S19B–S19E and sex-stratified differences are shown in Figures S19F–S19M). While group differences, even for heritable traits, certainly do not need to have a genetic basis, this observation, plus our prior focus on populations with largely EUR ancestry, prompted us to explore the genetic basis of skeletal growth phenotypes in a population of largely non-EUR ancestry.

To expand our investigation of skeletal growth phenotypes into a large sample of largely non-EUR ancestry, we performed a GWAS of skeletal growth phenotypes in the CKB, using 72,471 unrelated individuals of EAS ancestry (see material and methods). A GWAS was first performed in each of 10 subpopulations (reflecting geographic regions in mainland China), and then inverse-variance-weighted meta-analysis was performed (GWAS summary statistics are available in the GWAS Catalog [GWAS Catalog: GCST90727382, GCST90727383, GCST90727384, GCST90727385]). We identified 57 SHR-associated, 74 SH-associated, 96 LL-associated, and 128 height-associated loci using p value pruning with a 500 kb window (p < 5 × 10−8, λGC = 1.177 for the SHR). As an early indication of similar genetic bases for these phenotypes across ancestries, the majority (74/128) of CKB lead SNPs were within 35 kb of a UKB lead SNP, and the great majority (125/128) of CKB lead SNPs were within 500 kb of a UKB lead SNP.

To further investigate the extent of shared genetic effects between these two cohorts, a fixed-effects meta-analysis of the CKB and UKB using METAL17 was performed (max n = 524,392), identifying 594 SHR-associated, 672 SH-associated, 686 LL-associated, and 732 height-associated loci when including only variants present in both studies. We directly compared effect sizes of lead variants from the UKB and CKB GWASs after correcting for winner’s curse, a technical cause of differences in effect sizes between cohorts (see Note S1 and Figures S20 and S21). Despite correction for the winner’s curse, effect sizes at lead variants ascertained in the UKB were larger than when effects at these variants were estimated in the CKB. This finding could again be due to the fact that lead variants likely tag causal signals in the discovery population better than in the replication population or to true differences in effect sizes of causal variants.

Thus, an approach focusing on causal signals implicated by fine-mapping might better distinguish between these possibilities. Because we had previously published a smaller GWAS involving populations of diverse ancestries for the SHR,7 we also performed a meta-analysis of the SHR, including these prior data, UKB, and CKB (total n = 545,984), identifying 565 non-overlapping loci reaching genome-wide significance when limiting to variants present in all studies (Figure 1B; p < 5 × 10−8, λGC = 1.31; Figure S1B). To permit direct comparisons between large samples from substantially different ancestries, we focus on GWAS results from the CKB and UKB.

To address one possible source of differences in effect sizes, we used LDSC to calculate the heritability of each phenotype in the CKB. Notably, we observed lower heritability in the CKB than in the UKB for each phenotype. We tested whether these differences could stem from smaller sample sizes in the CKB by estimating heritability for GWASs of each phenotype in a UKB cohort downsampled to match the CKB sample size (Figure S4A); still, heritability estimates were significantly higher in the UKB than in the CKB (p < 2.9 × 10−20 for height). To further understand this difference, we first noted that the lower heritability estimate for the CKB is based on summary statistics of a meta-analysis of multiple subpopulations within the CKB. We therefore evaluated heritability in CKB subpopulations; despite the lower heritability estimate for height in the CKB meta-analysis, the estimated heritabilities of height in each CKB subpopulation were not significantly different from the UKB, with higher point estimates (Figure S4B).

To test the hypothesis that the lower heritability in the CKB meta-analysis compared with individual sub-studies could be due to systematically variable effect sizes across subpopulations, perhaps due to gene-environment interactions, we investigated the impact of the urban/rural origin of CKB sub-studies on effect estimates at associated variants. We compared effect sizes in each urban sub-study from the CKB with a meta-analysis of either the remaining urban sub-studies or the rural sub-studies (and similarly compared each rural sub-study with meta-analyses of either rural or urban sub-studies). At lead SNPs identified in the UKB, the correlation of effect sizes with the matching meta-analyses (urban-urban or rural-rural) was higher than the correlation with mismatched meta-analyses (t test p ≤ 1 × 10−4 for all 4 comparisons; Figure S4C). We conclude that the lower heritability estimate from the CKB meta-analysis (but not the CKB sub-studies) is likely due to the inclusion of heterogeneous CKB sub-studies with systematically varying effect sizes in the meta-analysis rather than a data quality issue, although other biobank-specific factors could also contribute. Importantly, cross-ancestry genetic correlations for each phenotype, estimated using Popcorn15 (see material and methods), did not show significant differences between the UKB and CKB, or between the CKB meta-analysis and the individual CKB subpopulations, across any of the growth-related phenotypes (⍴g was not statistically different from 1) (Figures 3A and S4D). Thus, samples with strong genetic correlation (in this case, the CKB meta-analysis compared with either the UKB or the CKB individual subpopulations) can, as expected, still have differences in estimated heritability. Because genome-wide genetic correlation is calculated across millions of variants, these estimates of strong genetic correlation do not preclude differences at individual loci, such as heterogeneity in the effect size or even direction of effect of a causal allele identified in the fine-mapping analysis presented above.

Comparing effect sizes of fine-mapped loci across ancestries identifies ancestrally distinct causal associations

Because our initial analyses evaluating differences in effect sizes were limited solely to lead signals discovered in the UKB (Figures 2A and S5–S7), these do not conclusively demonstrate true differences in effects of underlying causal variants. To search for regions of association, discovered in the UKB, where differences in observed effect sizes between the UKB and CKB were more likely due to true differences in causal variant effects, rather than the effects of differing LD between lead and causal variants, we further analyzed the UKB-discovered fine-mapped CSs. UKB CSs were proposed to be “ancestrally heterogeneous” (A-Het) if all variants in the CS showed significant heterogeneity of effect size between the UKB and CKB (described in the schematic in Figure 3B). CSs labeled as A-Het were concluded to be more likely to have differential effects at causal variants in the UKB and CKB (Figure 3C). Specifically, we reasoned that if the causal variant was in the CS and the causal variant had similar effect sizes in the UKB and CKB, then at least one variant in the CS should not have heterogeneity of effect between the UKB and CKB. Of 8,554 CSs identified across 4 phenotypes in the UKB, 53 CSs in 44 loci were labeled as A-Het (p for heterogeneity < 0.05/# total CSs per phenotype), and 9 loci contained more than one A-Het CS. Interestingly, 9 A-Het CSs were shared between phenotypes, and one was shared between 3 phenotypes: height, SH, and LL. Compared to CSs that included variants without evidence of heterogeneity in effect size, A-Het CSs had fewer variants per CS and were located in loci with more CSs; variant frequencies in A-Het CSs were more common than for the full set of variants in all CSs. These observations are likely in part due to the conservative definition of A-het CSs, which was designed to avoid spurious identification of loci with apparent differences in effect size that could in fact be explained by LD differences across populations. Examples of CSs with evidence of heterogeneity across ancestries are provided in Figures 3E and S22.

When restricting our analyses to the 14 SHR-associated loci containing a single CS comprising a single variant, only 1 significant A-Het locus was identified, despite having partially alleviated the multiple-testing burden. This single-variant CS (chr3:188365278A>G, rs4686991, [GRCh 37]) is associated with the SHR in the UKB and CKB in opposite directions (CKB marginal beta = 0.0401 [SE = 0.0204, p = 0.0496] vs. UKB marginal beta = −0.024792 [SE = 0.003156, p = 4.8 × 10−16]). This variant lies within an intron for the LPP gene encoding LPP, a LIM-domain-containing protein that may have roles in cell-cell adhesion or cell motility40; the CS also falls ∼41 kb upstream of MIR28, a microRNA (miRNA) involved in cell proliferation.41 Relaxing filtering for any of the phenotypes to consider loci with more than 1 CS identified additional A-Het CSs, but no additional A-Het CSs were identified when allowing for more than 1 variant in a CS (Figure 3D). In total, we identified a small number of CSs showing differences in variant effects between ancestries, only one of which is a single-variant CS and a few of which are in less complex loci with few CSs. Although our definition of A-Het CSs is conservative to avoid spurious examples of apparent heterogeneity in effect size, these results are consistent with our genome-wide genetic correlations showing shared genetics between ancestries and emphasize the importance of considering LD when evaluating ancestral differences. An approach considering only lead variants, as illustrated earlier in this work, is a poor proxy for—and likely overestimates—differences in causal effects between ancestries.

Discussion

The inherited basis of human height has been studied for over a century42; this includes some of the largest GWASs to date, with a recent near-complete mapping of common variant contributions to the heritability of height.5 However, much work remains to translate our new descriptions of the genetic basis of human height to specific insights about the biology driving human growth. Approaches to achieve this goal have included computational analyses of associated loci3,9 or functional experiments in relevant models.43,44,45,46,47 However, both of these types of studies are hindered by the uncertainty around the causal variant(s) at associated loci, and neither considers potential variability in phenotypic consequences of causal variants. A complementary approach to progressing from GWAS to biology is to integrate results from multiple related phenotypes. Past work involving GWASs of multiple related phenotypes has aimed at improving power for discovery,48 identifying loci with overlapping effects on multiple phenotypes,49,50 or classifying loci based on their effect on related phenotypes.51 These approaches, however, are not well equipped to handle complex loci containing multiple signals—as is often the case for polygenic phenotypes such as height—that could potentially have distinct effects on different phenotypes. In this work, we generated GWAS data for height-related phenotypes and then focused on using these data to disentangle putative causal signals within loci, including identifying loci with nearby but independent effects on related phenotypes.

To enable this multi-phenotype approach to interrogate the biology of human growth, our work increased the number of loci found to be associated with the SHR from 6 (in our earlier study) to 565. From these data, we prioritized biologically relevant genes and gene sets and fine-mapped signals specific to, or shared between, height-related phenotypes. By comparing results from two major ancestries, we could describe individual variant-level differences in effect sizes and identify loci where effect sizes are consistently different across CSs of potentially causal variants; these loci comprise a small minority of associated loci.

Leveraging associations with each of four height-related phenotypes, we could evaluate the shared and distinct components of each. Height-associated loci are observed both near and far from associations with each of the other three phenotypes, with LL-associated loci showing the greatest proximity. Additionally, identified height-associated signals5 are often located near loci with more significant associations with the SHR, SH, or LL than height, despite being ascertained on height, which has greater power for detection due to higher heritability in the UKB. This emphasizes the potential value of these height-related traits in understanding underlying skeletal growth biology. Height is highly genetically and phenotypically correlated with SH and LL but not with the SHR (when corrected for height), which is in turn moderately correlated with SH and LL (as expected). Additionally, SH and LL are only moderately correlated (expected given the variability in the SHR). GSEA identifies gene sets primarily related to growth and structural development (e.g., limb development and collagen binding) and to general biology (e.g., transcription factor binding) for each of the four traits. In addition, well-known genes related to growth were prioritized, and gene and gene-set prioritization for each trait can be found in Tables S1 and S2. SHR GSEA prioritized mostly genes and pathways shared with the other phenotypes, but, in particular, HOX genes (which are known to be implicated in the development of body structures52) were prioritized more strongly for the SHR than for height.

To evaluate the relationship between height-related phenotypes at a more granular level, we performed fine-mapping for loci associated with each of the 4 phenotypes and identified many variants that are members of a CS for more than one trait. In addition to expected shared signals between related phenotypes, we also identify CSs specific to a single phenotype; interestingly, many phenotype-specific CSs were observed to be in close proximity but distinct from CSs of other phenotypes. Notably, CSs near and shared between height and the SHR, compared to height-only CSs, were enriched for multiple metrics of relevance to the biology of growth. Additionally, we observe CSs for distinct phenotypes near each other; the presence of distinct but proximal genetic drivers is consistent with prior work.3,9,43,44

Our previous study7 evaluating the genetics of SHR demonstrated its high heritability and its potential use as a tool for understanding height-associated variants. Here, we augmented these results with GWASs using the UKB and CKB, increasing the total sample size from 25,135 to 545,984. In addition to increasing power, including multiple ancestries allowed us to ask questions about differences in putatively causal loci and their effect sizes across ancestral groups. Using GWASs of individuals from the two ancestries with the largest sample sizes, we looked to characterize similarity in effect sizes across ancestries genome wide. To explore this, we performed fine-mapping (in the better-powered EUR GWAS), nominating CSs where we are reasonably confident (alpha > 0.95) that the CS of the variants contains a truly causal variant. We reasoned that if we observe heterogeneity in effect size between ancestries at all potentially causal variants in the CS, we can be more confident that differing LD patterns are a much less likely explanation for the difference in effect size. We identified a small proportion of CSs (0.6% for height) with heterogeneous effect sizes across ancestries for all variants in the CS. Although it is possible that the same causal variants are present in each ancestry but have different effect sizes due to differing genetic background (potentially due to variant-variant interactions, which remain challenging to detect53) or that there is an ancestry-specific causal variant, it is also possible that these observed differences in effect are due to interactions with population-specific environmental factors or differences in sample ascertainment. Of note, the CSs showing complete heterogeneity of effect size were mostly seen in loci with many different signals of association, raising the possibility that some of this heterogeneity could be explained by insufficiently precise fine-mapping at these complex loci, leading the actual causal variant to not actually be in the CS.

This work is subject to some limitations. Notably, we observe distinct estimates of heritability between CKB sub-studies (individually consistent with estimates in the UKB) and a meta-analysis of CKB sub-studies (a lower estimate than in the UKB), contrary to the default expectation of shared driving genetics. Although surprising, this observation could reflect distinct environmental exposures in rural and urban environments, combined with gene × environment interactions. If this were true, some sets of variants would contribute more or less strongly to heritability in urban and rural sub-studies, and this would lower the estimate of heritability in the CKB meta-analysis. These differences in contributions to heritability could be due to a spectrum of possibilities, from a small number of variants having highly distinct effect sizes to polygenic differences where many variants have marginally distinct effects on phenotype. Future work could take a similar approach to that described here to identify and evaluate variants with genetic effects that differ between urban and rural CKB populations.

Our work based on fine-mapping is subject to additional limitations, some of which we discuss here. Fine-mapping can be sensitive to which variants are included in the model,16,54 and our computational approach to performing fine-mapping of regions larger than 1 Mb limited the number of variants included in the model to a subset of variants more likely to be causal. However, our comparisons across both phenotypes and populations remain similar when considering only loci of size 1 Mb, implying that our specific implementation for larger loci did not substantially bias the findings. It is also possible that some causal variants are not in our dataset (due to low frequency, type of variant, and/or filtering based on poor imputation/missingness), resulting in a CS being incorrectly designated as heterogeneous even though the missing causal variant has consistent effect sizes across ancestries. Additionally, this study only describes fine-mapping analyses performed on EUR-ancestry GWASs. Future work must be done to extend this approach to populations with additional ancestries, especially as sample sizes continue to grow, in order to ensure that findings from this and other genetic studies are equitable for all people. In particular, because of differences in available sample sizes, our approach is only powered to identify the heterogeneity of effect sizes at loci discovered in the UKB, as opposed to loci identified specifically in the CKB. For instance, we chose not to analyze the approximately 2,364 individuals with EAS ancestry in the UKB to preserve the homogeneity of individual GWAS analyses,55 and further method development would be needed to best integrate these samples to maximize equity of discovery. Finally, our approach, which requires heterogeneity at all variants in a CS, is likely biased to identify heterogeneity of effect size in small rather than larger CSs; additionally, this definition is conservative, and it remains possible that further analysis will uncover instances of effect heterogeneity between populations. The conservative nature of this approach and the presence of remaining caveats reflect the difficulty of estimating differences in variant effects across ancestries. Notably, in loci where there are multiple signals, nearby signals can be hard to distinguish and, without fine-mapping, could be errantly classified as heterogeneous in the context of differing LD between populations; therefore, we were conservative in our approach, relying on fine-mapped CSs and requiring all variants to show evidence of heterogeneity. Future work leveraging larger sample sizes, particularly in the replication dataset (where large standard errors around point estimates limit the power of heterogeneity testing), could be more sensitive in identifying instances of ancestral heterogeneity. Additionally, larger power in discovery, improvements in fine-mapping (such as well-validated methods for functionally aware fine mapping), or discovery in an ancestry with smaller LD blocks (e.g., African ancestry) could each reduce CS size, improving test power by reducing the number of variants tested for heterogeneity.

In summary, we have explored the genetics of height-related skeletal growth phenotypes—the SHR, SH, and LL, with substantially increased power to identify genetic associations with these phenotypes. We identify shared biology across phenotypes through GSEA, prioritizing relevant genes and pathways. We additionally identify shared and distinct genetic effects between phenotypes through the use of statistical fine-mapping and use this to categorize height-associated variants into different phenotypic categories that have different biological properties. We also propose a preliminary fine-mapping-based approach to identify loci where effects at causal variants are likely to differ between ancestries. As sample sizes continue to increase and fine-mapping methods continue to improve, these and similar approaches will enable a more thorough dissection of causal genes and variants, their phenotypic consequences, and any differences in effect sizes across ancestries. Taken together, this work illustrates the advantages of integrating genetic association data across related polygenic phenotypes and multiple ancestries to better understand underlying genetic and biological mechanisms.

Data and code availability

GWAS summary statistics per phenotype for the UKB and UKB-CKB-dbGaP meta-analysis summary statistics are available in the GWAS Catalog: the accession numbers for the UKB and meta-analysis summary statistics reported in this paper are GWAS Catalog: GCST90728584, GCST90728585, GCST90728586, GCST90728587, GCST90728588 and are available at https://www.joelhirschhornlab.org/additional-results. CKB summary stats will be viewable on the CKB Pheweb (pheweb.ckbiobank.org), the GWAS Catalog, and upon pending approval at https://www.joelhirschhornlab.org/additional-results. The accession numbers for the CKB summary statistics reported in this paper are GWAS Catalog: GCST90727382, GCST90727383, GCST90727384, GCST90727385. Analysis scripts will be made available on GitHub and are available upon request.

Acknowledgments

This research was conducted using the UK Biobank Resource under application 11898. We thank participants of the UK Biobank for making this work possible. This work was funded in part by NIH awards T32 HG002295 and R01DK075787. The CKB baseline survey and the first re-survey were supported by the Kadoorie Charitable Foundation in Hong Kong. Long-term follow-up was supported by the Wellcome Trust (212946/Z/18/Z, 202922/Z/16/Z, 104085/Z/14/Z, and 088158/Z/09/Z), the National Natural Science Foundation of China (82192901, 82192904, and 82192900), and the National Key Research and Development Program of China (2016YFC0900500). DNA extraction and genotyping were funded by GlaxoSmithKline and the UK Medical Research Council (MC-PC-13049 and MC-PC-14135). The project is supported by core funding from the UK Medical Research Council (MC_UU_00017/1, MC_UU_12026/2, and MC_U137686851), Cancer Research UK (C16077/A29186 and C500/A16896), and the British Heart Foundation (CH/1996001/9454) to the Clinical Trial Service Unit and Epidemiological Studies Unit and to the MRC Population Health Research Unit at Oxford University. We thank the participants in CKB and the members of the survey teams in each of the 10 regional centers and the project development and management teams based at Beijing, Oxford, and the 10 regional centers. China’s National Health Insurance provides electronic linkage to all hospital treatments. The content is the sole responsibility of the authors and does not necessarily reflect the official views of the NIH.

Declaration of interests

E.B. is currently an employee and shareholder at Empirico. J.N.H. held equity for Camp4 Therapeutics. W.G. is an employee and shareholder at Novo Nordisk. S.V. is an employee and shareholder at Regeneron.

Published: March 19, 2026

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2026.02.015.

Web resources

Supplemental information

Document S1. Figures S1–S23, Tables S4 and S7, and Notes S1–S3
mmc1.pdf (5.6MB, pdf)
Table S1. Gene prioritization by PoPs
mmc2.xlsx (4.3MB, xlsx)
Table S2. Gene set prioritization by MAGMA
mmc3.xlsx (2.2MB, xlsx)
Table S3. GO prioritization for SHR-specific genes
mmc4.xlsx (50.1KB, xlsx)
Table S5. Fine-mapping using SuSiE
mmc5.xlsx (15.2MB, xlsx)
Table S6. Short-stature genetic testing genes
mmc6.xlsx (26.9KB, xlsx)
Document S2. Article plus supplemental information
mmc7.pdf (12.3MB, pdf)

References

  • 1.Yang J., Bakshi A., Zhu Z., Hemani G., Vinkhuyzen A.A.E., Lee S.H., Robinson M.R., Perry J.R.B., Nolte I.M., van Vliet-Ostaptchouk J.V., et al. Genetic variance estimation with imputed variants finds negligible missing heritability for human height and body mass index. Nat. Genet. 2015;47:1114–1120. doi: 10.1038/ng.3390. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Lango Allen H., Estrada K., Lettre G., Berndt S.I., Weedon M.N., Rivadeneira F., Willer C.J., Jackson A.U., Vedantam S., Raychaudhuri S., et al. Hundreds of variants clustered in genomic loci and biological pathways affect human height. Nature. 2010;467:832–838. doi: 10.1038/nature09410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Wood A.R., Esko T., Yang J., Vedantam S., Pers T.H., Gustafsson S., Chu A.Y., Estrada K., Luan J. ’an, Kutalik Z., et al. Defining the role of common variation in the genomic and biological architecture of adult human height. Nat. Genet. 2014;46:1173–1186. doi: 10.1038/ng.3097. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Yengo L., Sidorenko J., Kemper K.E., Zheng Z., Wood A.R., Weedon M.N., Frayling T.M., Hirschhorn J., Yang J., Visscher P.M., GIANT Consortium Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry. Hum. Mol. Genet. 2018;27:3641–3649. doi: 10.1093/hmg/ddy271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Yengo L., Vedantam S., Marouli E., Sidorenko J., Bartell E., Sakaue S., Graff M., Eliasen A.U., Jiang Y., Raghavan S., et al. A saturated map of common genetic variants associated with human height. Nature. 2022;610:704–712. doi: 10.1038/s41586-022-05275-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.O’Connor L.J., Schoech A.P., Hormozdiari F., Gazal S., Patterson N., Price A.L. Extreme Polygenicity of Complex Traits Is Explained by Negative Selection. Am. J. Hum. Genet. 2019;105:456–476. doi: 10.1016/j.ajhg.2019.07.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Chan Y., Salem R.M., Hsu Y.-H.H., McMahon G., Pers T.H., Vedantam S., Esko T., Guo M.H., Lim E.T., et al. GIANT Consortium Genome-wide Analysis of Body Proportion Classifies Height-Associated Variants by Mechanism of Action and Implicates Genes Important for Skeletal Development. Am. J. Hum. Genet. 2015;96:695–708. doi: 10.1016/j.ajhg.2015.02.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Kun E., Javan E.M., Smith O., Gulamali F., de la Fuente J., Flynn B.I., Vajrala K., Trutner Z., Jayakumar P., Tucker-Drob E.M., et al. The genetic architecture of the human skeletal form. bioRxiv. 2023 doi: 10.1101/2023.01.03.521284. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Bartell E., Fujimoto M., Khoury J.C., Khoury P.R., Vedantam S., Astley C.M., Hirschhorn J.N., Dauber A. Protein QTL analysis of IGF-I and its binding proteins provides insights into growth biology. Hum. Mol. Genet. 2020;29:2625–2636. doi: 10.1093/hmg/ddaa103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Hernández N., Soenksen J., Newcombe P., Sandhu M., Barroso I., Wallace C., Asimit J.L. The flashfm approach for fine-mapping multiple quantitative traits. Nat. Commun. 2021;12:6147. doi: 10.1038/s41467-021-26364-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Nagai A., Hirata M., Kamatani Y., Muto K., Matsuda K., Kiyohara Y., Ninomiya T., Tamakoshi A., Yamagata Z., Mushiroda T., et al. Overview of the BioBank Japan Project: Study design and profile. J. Epidemiol. 2017;27:S2–S8. doi: 10.1016/j.je.2016.12.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Chen Z., Chen J., Collins R., Guo Y., Peto R., Wu F., Li L., China Kadoorie Biobank CKB collaborative group China Kadoorie Biobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. Int. J. Epidemiol. 2011;40:1652–1666. doi: 10.1093/ije/dyr120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Conti D.V., Darst B.F., Moss L.C., Saunders E.J., Sheng X., Chou A., Schumacher F.R., Olama A.A.A., Benlloch S., Dadaev T., et al. Trans-ancestry genome-wide association meta-analysis of prostate cancer identifies new susceptibility loci and informs genetic risk prediction. Nat. Genet. 2021;53:65–75. doi: 10.1038/s41588-020-00748-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Mahajan A., Spracklen C.N., Zhang W., Ng M.C.Y., Petty L.E., Kitajima H., Yu G.Z., Rüeger S., Speidel L., Kim Y.J., et al. Multi-ancestry genetic study of type 2 diabetes highlights the power of diverse populations for discovery and translation. Nat. Genet. 2022;54:560–572. doi: 10.1038/s41588-022-01058-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Brown B.C., Asian Genetic Epidemiology Network Type 2 Diabetes Consortium. Ye C.J., Price A.L., Zaitlen N. Transethnic Genetic-Correlation Estimates from Summary Statistics. Am. J. Hum. Genet. 2016;99:76–88. doi: 10.1016/j.ajhg.2016.05.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Kanai M., Ulirsch J.C., Karjalainen J., Kurki M., Karczewski K.J., Fauman E., Wang Q.S., Jacobs H., Aguet F., Ardlie K.G., et al. Insights from complex trait fine-mapping across diverse populations. medRxiv. 2021 doi: 10.1101/2021.09.03.21262975. [DOI] [Google Scholar]
  • 17.Willer C.J., Li Y., Abecasis G.R. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26:2190–2191. doi: 10.1093/bioinformatics/btq340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Walters R.G., Millwood I.Y., Lin K., Valle D.S., McDonnell P., Hacker A., Avery D., Cai N., Kretzschmar W.W., Ansari M.A., et al. Genotyping and population structure of the China Kadoorie Biobank. Cell Genom. 2023;3 doi: 10.1016/j.xgen.2023.100361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pers T.H., Karjalainen J.M., Chan Y., Westra H.-J., Wood A.R., Yang J., Lui J.C., Vedantam S., Gustafsson S., Esko T., et al. Biological interpretation of genome-wide association studies using predicted gene functions. Nat. Commun. 2015;6:5890. doi: 10.1038/ncomms6890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Pers T.H., Timshel P., Hirschhorn J.N. SNPsnap: a Web-based tool for identification and annotation of matched SNPs. Bioinformatics. 2015;31:418–420. doi: 10.1093/bioinformatics/btu655. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Lui J.C., Nilsson O., Chan Y., Palmer C.D., Andrade A.C., Hirschhorn J.N., Baron J. Synthesizing genome-wide association studies and expression microarray reveals novel genes that act in the human growth plate to modulate height. Hum. Mol. Genet. 2012;21:5193–5201. doi: 10.1093/hmg/dds347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Blueprint Genetics. 2021. Genetic testing for Growth related and skeletal disorders.https://www.blueprintgenetics.com/tests/no-test-type/comprehensive-growth-disorders-skeletal-dysplasias-and-disorders-panel/ [Google Scholar]
  • 23.Renthal N.E., Nakka P., Baronas J.M., Kronenberg H.M., Hirschhorn J.N. Genes with specificity for expression in the round cell layer of the growth plate are enriched in genomewide association study (GWAS) of human height. J. Bone Miner. Res. 2021;36:2300–2308. doi: 10.1002/jbmr.4408. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Palmer C., Pe’er I. Statistical correction of the Winner’s Curse explains replication variability in quantitative trait genome-wide association studies. PLoS Genet. 2017;13 doi: 10.1371/journal.pgen.1006916. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Zhong H., Prentice R.L. Bias-reduced estimators and confidence intervals for odds ratios in genome-wide association studies. Biostatistics. 2008;9:621–634. doi: 10.1093/biostatistics/kxn001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Choi S.W., O’Reilly P.F. PRSice-2: Polygenic Risk Score software for biobank-scale data. GigaScience. 2019;8 doi: 10.1093/gigascience/giz082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Sherry S.T., Ward M., Sirotkin K. dbSNP—Database for Single Nucleotide Polymorphisms and Other Classes of Minor Genetic Variation. Genome Res. 1999;9:677–679. [PubMed] [Google Scholar]
  • 28.Bulik-Sullivan B., Finucane H.K., Anttila V., Gusev A., Day F.R., Loh P.-R., Duncan L., Perry J.R.B., Robinson E.B., et al. Patterson N., Genetic Consortium for Anorexia Nervosa of the Wellcome Trust Case Control Consortium 3 An atlas of genetic correlations across human diseases and traits. Nat. Genet. 2015;47:1236–1241. doi: 10.1038/ng.3406. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Bulik-Sullivan B.K., Loh P.-R., Finucane H.K., Ripke S., Yang J., Schizophrenia Working Group of the Psychiatric Genomics Consortium. Patterson N., Daly M.J., Price A.L., Neale B.M. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat. Genet. 2015;47:291–295. doi: 10.1038/ng.3211. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.de Leeuw C.A., Mooij J.M., Heskes T., Posthuma D. MAGMA: generalized gene-set analysis of GWAS data. PLoS Comput. Biol. 2015;11 doi: 10.1371/journal.pcbi.1004219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Weeks E.M., Ulirsch J.C., Cheng N.Y., Trippe B.L., Fine R.S., Miao J., Patwardhan T.A., Kanai M., Nasser J., Fulco C.P., et al. Leveraging polygenic enrichments of gene features to predict genes underlying complex traits and diseases. Nat Genet. 2023;55:1267–1276. doi: 10.1038/s41588-023-01443-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kanai M., Elzur R., Zhou W., Global Biobank Meta-analysis Initiative. Daly M.J., Finucane H.K. Meta-analysis fine-mapping is often miscalibrated at single-variant resolution. Cell Genom. 2022;2 doi: 10.1016/j.xgen.2022.100210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Wang G., Sarkar A., Carbonetto P., Stephens M. A Simple New Approach to Variable Selection in Regression, with Application to Genetic Fine Mapping. J. R. Stat. Soc., Ser. B Stat. Methodol. 2020;82:1273–1300. doi: 10.1111/rssb.12388. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Anastasio N., Ben-Omran T., Teebi A., Ha K.C.H., Lalonde E., Ali R., Almureikhi M., Der Kaloustian V.M., Liu J., Rosenblatt D.S., et al. Mutations in SCARF2 are responsible for Van Den Ende-Gupta syndrome. Am. J. Hum. Genet. 2010;87:553–559. doi: 10.1016/j.ajhg.2010.09.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Laron Z. Lessons from 50 years of study of Laron syndrome. Endocr. Pract. 2015;21:1395–1402. doi: 10.4158/EP15939.RA. [DOI] [PubMed] [Google Scholar]
  • 36.dos Santos C.C., Gavish H., Buchwald M. Fanconi anemia revisited: old ideas and new advances. Stem Cell. 1994;12:142–153. doi: 10.1002/stem.5530120202. [DOI] [PubMed] [Google Scholar]
  • 37.Cormier-Daire V. Spondylo-epi-metaphyseal dysplasia. Best Pract. Res. Clin. Rheumatol. 2008;22:33–44. doi: 10.1016/j.berh.2007.12.009. [DOI] [PubMed] [Google Scholar]
  • 38.Eveleth P.B., Tanner J.M. Cambridge University Press; 1991. Worldwide Variation in Human Growth. [Google Scholar]
  • 39.Bogin B., Varela-Silva M.I. Leg length, body proportion, and health: a review with a note on beauty. Int. J. Environ. Res. Publ. Health. 2010;7:1047–1075. doi: 10.3390/ijerph7031047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Kuriyama S., Yoshida M., Yano S., Aiba N., Kohno T., Minamiya Y., Goto A., Tanaka M. LPP inhibits collective cell migration during lung cancer dissemination. Oncogene. 2016;35:952–964. doi: 10.1038/onc.2015.155. [DOI] [PubMed] [Google Scholar]
  • 41.Schneider C., Setty M., Holmes A.B., Maute R.L., Leslie C.S., Mussolin L., Rosolen A., Dalla-Favera R., Basso K. MicroRNA 28 controls cell proliferation and is down-regulated in B-cell lymphomas. Proc. Natl. Acad. Sci. USA. 2014;111:8185–8190. doi: 10.1073/pnas.1322466111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Galton F. Regression Towards Mediocrity in Hereditary Stature. J. Anthropol. Inst. G. B. Ireland. 1886;15:246–263. [Google Scholar]
  • 43.Muthuirulan P., Capellini T.D. Complex Phenotypes: Mechanisms Underlying Variation in Human Stature. Curr. Osteoporos. Rep. 2019;17:301–323. doi: 10.1007/s11914-019-00527-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Baronas J.M., Bartell E., Eliasen A., Doench J.G., Yengo L., Vedantam S., Marouli E., Kronenberg H.M., Hirschhorn J.N., Renthal N.E. Genome-wide CRISPR screening of chondrocyte maturation newly implicates genes in skeletal growth and height-associated GWAS loci. Cell Genom. 2023;3 doi: 10.1016/j.xgen.2023.100299. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Welter D., MacArthur J., Morales J., Burdett T., Hall P., Junkins H., Klemm A., Flicek P., Manolio T., Hindorff L., Parkinson H. The NHGRI GWAS Catalog, a curated resource of SNP-trait associations. Nucleic Acids Res. 2014;42:D1001–D1006. doi: 10.1093/nar/gkt1229. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Lappalainen T., MacArthur D.G. From variant to function in human disease genetics. Science. 2021;373:1464–1468. doi: 10.1126/science.abi8207. [DOI] [PubMed] [Google Scholar]
  • 47.Amariuta T., Luo Y., Gazal S., Davenport E.E., van de Geijn B., Ishigaki K., Westra H.-J., Teslovich N., Okada Y., Yamamoto K., et al. IMPACT: Genomic Annotation of Cell-State-Specific Regulatory Elements Inferred from the Epigenome of Bound Transcription Factors. Am. J. Hum. Genet. 2019;104:879–895. doi: 10.1016/j.ajhg.2019.03.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Turley P., Walters R.K., Maghzian O., Okbay A., Lee J.J., Fontana M.A., Nguyen-Viet T.A., Wedow R., Zacher M., Furlotte N.A., et al. Multi-trait analysis of genome-wide association summary statistics using MTAG. Nat. Genet. 2018;50:229–237. doi: 10.1038/s41588-017-0009-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Shi H., Mancuso N., Spendlove S., Pasaniuc B. Local Genetic Correlation Gives Insights into the Shared Genetic Architecture of Complex Traits. Am. J. Hum. Genet. 2017;101:737–751. doi: 10.1016/j.ajhg.2017.09.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Giambartolomei C., Vukcevic D., Schadt E.E., Franke L., Hingorani A.D., Wallace C., Plagnol V. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 2014;10 doi: 10.1371/journal.pgen.1004383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Udler M.S., Kim J., von Grotthuss M., Bonàs-Guarch S., Cole J.B., Chiou J., et al. Christopher D Anderson on behalf of METASTROKE and the ISGC. Boehnke M., Laakso M., Atzmon G. Type 2 diabetes genetic loci informed by multi-trait associations point to disease mechanisms and subtypes: A soft clustering analysis. PLoS Med. 2018;15 doi: 10.1371/journal.pmed.1002654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Hajirnis N., Mishra R.K. Homeotic Genes: Clustering, Modularity, and Diversity. Front. Cell Dev. Biol. 2021;9 doi: 10.3389/fcell.2021.718308. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Westerman K.E., Majarian T.D., Giulianini F., Jang D.-K., Miao J., Florez J.C., Chen H., Chasman D.I., Udler M.S., Manning A.K., Cole J.B. Variance-quantitative trait loci enable systematic discovery of gene-environment interactions for cardiometabolic serum biomarkers. Nat. Commun. 2022;13:3993. doi: 10.1038/s41467-022-31625-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Ulirsch J.C. Harvard University; 2022. Identification and Interpretation of Causal Genetic Variants Underlying Human Phenotypes. [Google Scholar]
  • 55.Constantinescu A.-E., Mitchell R.E., Zheng J., Bull C.J., Timpson N.J., Amulic B., Vincent E.E., Hughes D.A. A framework for research into continental ancestry groups of the UK Biobank. Hum. Genomics. 2022;16:3. doi: 10.1186/s40246-022-00380-5. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S23, Tables S4 and S7, and Notes S1–S3
mmc1.pdf (5.6MB, pdf)
Table S1. Gene prioritization by PoPs
mmc2.xlsx (4.3MB, xlsx)
Table S2. Gene set prioritization by MAGMA
mmc3.xlsx (2.2MB, xlsx)
Table S3. GO prioritization for SHR-specific genes
mmc4.xlsx (50.1KB, xlsx)
Table S5. Fine-mapping using SuSiE
mmc5.xlsx (15.2MB, xlsx)
Table S6. Short-stature genetic testing genes
mmc6.xlsx (26.9KB, xlsx)
Document S2. Article plus supplemental information
mmc7.pdf (12.3MB, pdf)

Data Availability Statement

GWAS summary statistics per phenotype for the UKB and UKB-CKB-dbGaP meta-analysis summary statistics are available in the GWAS Catalog: the accession numbers for the UKB and meta-analysis summary statistics reported in this paper are GWAS Catalog: GCST90728584, GCST90728585, GCST90728586, GCST90728587, GCST90728588 and are available at https://www.joelhirschhornlab.org/additional-results. CKB summary stats will be viewable on the CKB Pheweb (pheweb.ckbiobank.org), the GWAS Catalog, and upon pending approval at https://www.joelhirschhornlab.org/additional-results. The accession numbers for the CKB summary statistics reported in this paper are GWAS Catalog: GCST90727382, GCST90727383, GCST90727384, GCST90727385. Analysis scripts will be made available on GitHub and are available upon request.


Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES