Summary
Hepatocellular carcinoma (HCC) is the third leading cause of cancer death globally, often arising on a background of cirrhosis. Here, we aimed to establish genetic drivers of all-cause HCC across ancestries in a large meta-analysis. We included 15 cohorts comprising 17,697 HCC affected individuals and 2,715,683 control subjects in this meta-analysis. We found 15 genome-wide significant (p < 5 × 10−8) germline loci, including in/near GCKR, MTTP, ADH5, 8q24.21 (nearest gene MYC), MAP3K9, and GABPB2, and a further two loci found on transcriptome- and regulome-wide association analyses. MAP3K9, TERT, and GABPB2 variants act independently of cirrhosis on both colocalization analysis and sensitivity analyses. There was significant ancestral heterogeneity in six loci including variants in the HLA locus that had divergent effects on HCC risk between East Asian and European ancestries. Fine mapping identified 11 potentially causal coding variants, including p.Leu446Pro (c.1337T>C) in GCKR and p.Asp423Glu (c.1269C>T) in MEN1. MEN1, 8q24.21 (nearest gene MYC), and TERT are all involved in the β-catenin pathway transactivation complex. Transcriptome-wide analysis identified enrichment of germline-encoded DHRS1 in HCC. Regulome-wide analysis replicated the germline signal for EPHA2 and found a chromatin-accessible region containing genes ZNF367 and HABP4. Finally, we demonstrated that population-level genetic architecture for HCC overlaps with steatotic and viral liver disease, and individuals with genetic risk for lower body mass index have higher risk of HCC. Genetic risk for HCC is determined by germline susceptibility to β-catenin pathway activation and cirrhosis. HCC is driven by both heterogeneous and homogeneous genetic factors across ancestries.
Keywords: liver cancer, cirrhosis, steatosis, genomics
Graphical abstract

Germline variation contributes to hepatocellular carcinoma risk through susceptibility to β-catenin pathway activation and cirrhosis. This multi-ancestry meta-analysis of 17,697 affected individuals identifies 15 GWAS-significant loci and two further loci through transcriptome- and regulome-wide analyses, including 8q24 (nearest gene MYC), MAP3K9, and DHRS1.
Introduction
Hepatocellular carcinoma (HCC [MIM: 114550]) is the third most common cause of cancer death globally.1 It typically arises on a background of cirrhosis, which can be caused by any etiology of liver disease.2 HCC may also develop in non-cirrhotic individuals, particularly in chronic hepatitis B virus (HBV) infection and in metabolic-dysfunction-associated steatotic liver disease (MASLD). Identification of genetic risk factors is critical to mitigating the rapidly growing burden of HCC.3
Several previous genome-wide association studies (GWASs) have discovered a total of 35 independent risk loci in individuals of European and East Asian genetic ancestry.4,5,6,7 Recent GWASs have demonstrated that combining etiologies of liver disease can increase power to detect variants that perturb common pathways in the mechanisms of liver fibro-inflammation.8 For example, an all-cause cirrhosis GWAS identified p.Thr165Pro (c.493A>G) in MTARC1 (MIM: 614126) as a protective variant, which was not apparent in single-etiology analyses.9
Two GWAS meta-analyses of HCC have expanded our understanding of germline variation that drives HCC.6,7 Genes that contribute to HCC can be broadly divided into those that act via cirrhosis and those that are independent of liver fibrosis. Variants in TM6SF2 (MIM: 606563) and PNPLA3 (MIM: 609567) are the two most replicated genes that increase risk of cirrhosis liver fibrosis, primarily by increasing hepatic lipid accumulation,10 whereas variants in TERT (MIM: 187270) appear to directly affect HCC development without any mediation via chronic-liver-disease-related pathways.7,11
Here, we performed a multi-ancestry meta-analysis (MAMA) of HCC GWASs, which included transcriptome-wide and regulome-wide association analyses. In doing so, we identified 15 loci as putative risk genes for HCC and established a baseline for future studies with inclusion of diverse ancestries.
Material and methods
Study design and cohort descriptions
We used a single-stage joint meta-analysis design to maximize statistical power: (1) MR-MEGA for heterogeneous effects across ancestries; and (2) sample-size weighted analysis via METAL for homogeneous effects across all ancestries as well as within European (EUR) and East Asian (EAS) ancestries. We used datasets representing the four different ancestry groups EUR, EAS, Admixed American (AMR), and African (AFR), as well as two cohorts of mixed ancestry (which were >70% European, Table S1). GWAS results of Trépo et al.4 were acquired through GWAS Catalog (NHGRI: GCST90092003). Summary statistics for Hassan et al.5 were previously reported and provided via direct collaboration with the authors. UK Biobank (UKBB), DeCode, and Intermountain summary statistics were obtained from https://www.decode.com/summarydata/ and are reported previously.12 The FinnGen GWAS summary statistics were acquired from FinnGen Release 9 (https://www.finngen.fi/en). Taiwan Precision Medicine Initiative (TPMI) summary statistics were acquired from https://pheweb.ibms.sinica.edu.tw/pheno/155.1.13,14 China Kadoorie Biobank (CKBB) summary statistics were acquired from https://pheweb.ckbiobank.org./download/c22.15 BioBank Japan (BBJ) summary statistics were acquired from GWAS Catalog (NHGRI: GCST90018638).16 Korean BioBank (KoGES) summary statistics were acquired from https://koges.leelabsg.org/.17 Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial summary statistics were acquired from https://exploregwas.cancer.gov/plco-atlas.18 Millions Veterans Program (MVP) summary statistics were acquired from https://ftp.ncbi.nlm.nih.gov/dbgap/studies/phs002453/analyses/GIA/ as three separate cohorts for EUR, AMR, and AFR, and are reported previously.19 All of Us summary statistics were generated through analyses for this study.20,21
GRCh38 were used for analysis and reported throughout. Where GRCh37 was required by specific tools (e.g., FUMA), LiftOver from MungeSumstats22 was used for conversion.
Ethics
This study is a meta-analysis of existing summary-level GWAS data. All primary studies included in this meta-analysis were conducted in accordance with the ethical standards of the responsible relevant institutions, and informed consent was obtained from all participants as described in the original publications. No new data collection or participant recruitment was performed for this study, and therefore no additional ethical approval was required.
All of Us
Inclusion criteria and genotyping of the All of Us cohort is described in detail elsewhere.21 In brief, All of Us is a population-level dataset of approximately 350,000 individuals from the United States. Analysis of the All of Us Research Program dataset was conducted under Research Project ID 65834, in accordance with the All of Us Research Program policies and covered by the All of Us IRB approval. Individuals were included in the current analysis if there was available data on age, sex, and presence or absence of HCC. ICD diagnostic codes (C22.0-ICD03, 155.0-ICD9, and C22.8-ICD10) were used to identify HCC cases. The GWAS was run using Hail23 adjusted for age, sex, and the first nine principal components (PCs) of genetic ancestry.
MAMA
We performed MAMA of GWAS results using MR-MEGA24 and METAL.25 MR-MEGA accounts for differences in population structure through meta-regression, whereby axes of genetic variation generate covariates for inclusion in the meta-analysis. Four PCs were included in the analysis. MR-MEGA has reduced power to detect associations for variants with homogeneous effects across populations. It is therefore recommended to run MR-MEGA in conjunction with another method, so we used sample-size weighted METAL (as has been reported in two recent meta-analyses6,7) to detect homogeneous allelic effects.
Prior to meta-analysis, all datasets were harmonized to genome build hg38 using MungeSumstats22 and R 4.5.026 and included in NCBI GRCh38 (2013-12-17) reference panel. Sex-chromosome data were not available for all cohorts; therefore, only autosomal variants were kept in the results.
To assess potential sample overlap between cohorts of the same ancestry, summary statistics were converted to GRCh37 using MungeSumstats and munged against HapMap3 SNPs using munge_sumstats.py. Pairwise genetic covariance intercepts (gcov_int) were then calculated for all EUR cohort pairs (BCM, deCODE, FinnGen, Intermountain, MVP-EUR, Trépo et al.,4 and UKBB) and all EAS cohort pairs (BBJ, CKBB, KoGES, and TPMI) using ldsc.py with EUR and EAS LD score reference panels, respectively. A gcov_int near zero or within the 95% confidence interval (gcov_int ± 1.96 × gcov_int_se) was considered to indicate no meaningful sample overlap (Table S2).
In total, 11,122,532 variants were included in MR-MEGA analyses, as this method has a cohort-number requirement that varies based on the number of axes of variation. Number of variants included in METAL analyses were: 33,182,383 for all cohorts, 22,029,853 for EUR, and 7,735,436 for EAS. Genomic inflations were measured for all cohorts and the meta-analyses. Genomic control was applied to meta-analysis using METAL. All inflation was nominal and below 1.04 for all cohorts and results (highest for MR-MEGA with λ = 1.0375 due to increased degrees of freedom in the test (Figure S1). GWAS significance was set at p < 5 × 10−8 and suggestive threshold set at p < 1 × 10−6. For variants to be considered GWAS-significant in all ancestry analysis using METAL, they must also reach at least suggestive significance with MR-MEGA. Ancestral (pANC-HET) and residual heterogeneity (pRES-HET) p values from MR-MEGA were assessed at all loci reaching suggestive significance (p < 1 × 10−6, n = 76). A Bonferroni-corrected threshold of p < 6.6 × 10−4 (0.05/76) was applied to declare significant heterogeneity in each analysis.
We identified genomic risk loci within our meta-analysis results using Functional Mapping and Annotation (FUMA) v.1.5.2.27,28 FUMA identifies independent significant SNPs in the GWAS results. FUMA identifies independent significant SNPs by clumping variants at r2 < 0.1, then defines locus boundaries by merging linkage disequilibrium (LD) blocks of all SNPs in r2 ≥ 0.6 with any independent significant SNP within 250 kb.
All significant SNPs were compared to the known HCC risk variants reported in previous analyses, including the preprint by Vujkovic et al.7 Forest plots and Manhattan plots were generated using R 4.5.0. QQ plots were generated by FUMA.
Colocalization
To determine whether variants were shared between HCC and cirrhosis, we performed colocalization analysis using summary statistics from the largest available GWAS of cirrhosis.29 Summary statistics were standardized using MungeSumstats and converted into variant call format using gwas2vcf for analysis using gwasglue.30 Evidence of colocalization due to a common causal variant was defined by PP.H4 ≥0.8.
Fine mapping
Fine mapping was performed using MR-MEGA, which derives a natural log of Bayes factor (lnBF) per SNP in favor of association, as previously described.31 All significant and suggestive SNPs were included for fine mapping. Posterior probabilities (PPs) of driving the association signal at each locus were calculated from the Bayes factor, as previously described.31 Credible sets of fewer than five SNPs with sum PP greater than 0.95 were accepted as putative causal variants. Fine-mapped variants were run through Variant Effect Predictor (VEP) to determine functional consequences.32
MAGMA: Gene ontology and tissue enrichment
Results were analyzed using MAGMA33 to identify Gene Ontology term enrichment and gene expression data from tissues in GTEx v.8, performed via the GENE2FUNC function of FUMA with default parameters.28 Mapped genes were analyzed using gene set analysis for ontology terms from MSigDB v.7 and gene-property analysis for tissue specificity. Results were adjusted for multiple tests using Bonferroni false discovery rate (FDR) correction with alpha of 0.05. The significant ontology terms were filtered for terms that share the same signals. Tissue-level enrichment analysis was carried out using the preprocessed GTEx gene-expression dataset provided by FUMA investigators.
Transcriptome-wide and regulome-wide association studies
Transcriptome-wide association studies (TWASs) and regulome-wide association studies (RWASs) were performed using FUSION,34 which uses tissue-specific (or cancer-specific) expression quantitative trait loci (eQTLs) to determine the effects of germline variation on transcriptome profile (TWAS) or chromatin accessibility (RWAS).35 TWAS was performed using (1) GTEx v8 liver precomputed weights and (2) HCC precomputed weights (“LIHC”), as provided by the FUSION investigators, with 1000 Genomes data reference panels. RWAS was performed using known chromatin accessibility profiles for HCC (“LIHC”). All results underwent Benjamini-Hochberg FDR correction with alpha of 0.05.
LDSC regression analysis across traits
We performed linkage disequilibrium score regression (LDSC) analysis to compare the similarity in genetic architecture across traits. Summary statistics for the largest publicly available GWAS were obtained for cirrhosis (NHGRI: GCST90319878),29 liver fat (NHGRI: GCST90016673),36 alanine transaminase (ALT; NHGRI: GCST90018943),16 hepatitis B viral infection (HBV; NHGRI: GCST90479751),19 hepatitis C virus infection (HCV; NHGRI: GCST90018805),16 body mass index (BMI; NHGRI: GCST009004),37 type 2 diabetes mellitus (T2DM [MIM: 125853]; NHGRI: GCST90444202),38 upper gastrointestinal (GI) cancer (gastric and esophageal; NHGRI: GCST90011807),39 and colorectal cancer (NHGRI: GCST90129505).40 Summary statistics were formatted and converted to GRCh37 using MungeSumstats, then aligned to the 1000 Genomes reference panel. We then ran pairwise regression (--rg) comparisons between all traits.
Results
GWAS meta-analyses identify 15 significant loci
We used four datasets from EAS, seven from EUR, one from AFR, one from AMR, and two from mixed-ancestry cohorts (Figure 1 and Table S1). Overall, this included 17,697 HCC affected individuals and 2,715,683 control subjects in the MAMA. There was no evidence of overlap between cohorts within ancestries based on LDSCs, which were close to zero or within the 95% confidence interval (Table S2).
Figure 1.
Study overview
(A) Fifteen cohorts from four genetic ancestries were included in the meta-analysis, totaling 17,329 hepatocellular carcinoma (HCC) cases.
(B) Post-GWAS analysis included colocalization with cirrhosis, multi-marker analysis of genomic annotation (MAGMA) for Gene Ontology and tissue enrichment, fine mapping, and transcriptome-/regulome-wide association studies.
Following data harmonization and mapping to genome build hg38, MAMA was performed using meta-regression of multi-ethnic genetic association (MR-MEGA) and METAL (including single-ancestry analyses for EUR and EAS cohorts). As previously described, MR-MEGA has greater power to detect heterogeneous effects across different cohorts as well as distinguishing variation in effect estimates from ancestry-level genetic variation. Whereas METAL had better power for identification of homogeneous allelic effects. MR-MEGA also facilitates calculation of posterior probabilities for fine mapping across ancestries using allele frequencies in the different cohorts.
In combination, we found 15 loci that met GWAS significance of p < 5 × 10−8 (Table 1; Figures 2, S1, and S2; Tables S3 and S4). Out of 15 total GWAS-significant loci, nine were observed across both MR-MEGA and METAL models.
Table 1.
GWAS-significant loci associated with HCC from MR-MEGA and METAL meta-analysis models
| rsID | Nearest gene | Genomic position (GRCh38) | Effects | Beta (SE) [MR-MEGA] | p(MR-MEGA) | p(METAL-all) | p(METAL-EUR) | p(METAL-EAS) | p(ANC-HET) | p(RES-HET) | EAF |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Significant | |||||||||||
| rs924204 | EPHA2 | g.16187431A>G (GenBank: NC_000001.11) | +++++-++++-+?+- | 0.10 (0.02) | 9.20e−13 | 7.12e−13 | 1.36e−13 | 9.47e−3 | 5.43e−4 | 1.23e−1 | 0.65 |
| rs77018356 | GABPB2 | g.151162485G>C (GenBank: NC_000001.11) | ?++++?-+-+-+??? | 2.90 (3.88) | 4.91e−6 | 1.02e−6 | 3.84e−8 | 2.06e−3 | 8.97e−2 | 8.76e−3 | 0.06 |
| rs780094 | GCKR | g.27518370T>C (GenBank: NC_000002.12) | -----++-----?-+ | −0.08 (0.01) | 2.22e−7 | 8.69e−9 | 8.46e−7 | 9.07e−3 | 2.22e−1 | 4.79e−1 | 0.57 |
| rs4089 | HSD17B13 | g.87300193C>G (GenBank: NC_000004.12) | ----+-------?-+ | −0.15 (0.02) | 1.78e−20 | 2.82e−17 | 5.74e−18 | 4.10e−2 | 2.01e−9 | 1.24e−1 | 0.3 |
| rs6822348 | ADH5 | g.99132743A>T (GenBank: NC_000004.12) | -----?------??? | −0.12 (0.10) | 3.12e−7 | 4.44e−8 | 3.43e−8 | 3.69e−5 | 9.90e−1 | 6.72e−1 | 0.76 |
| rs112850336 | MTTP | g.99492512C>T (GenBank: NC_000004.12) | -----+-+---+?-- | −0.17 (0.04) | 2.57e−6 | 2.01e−8 | 6.83e−5 | 4.77e−5 | 5.27e−2 | 5.65e−1 | 0.22 |
| rs10069690 | TERT | g.1279675C>T (GenBank: NC_000005.10) | ----++--------+ | −0.15 (0.02) | 2.50e−20 | 1.84e−21 | 6.40e−16 | 5.77e−5 | 5.07e−2 | 8.04e−3 | 0.25 |
| rs1800562 | HFE | g.26092913G>A (GenBank: NC_000006.12) | ?+++-?+?++-+-?? | 0.12 (2.89) | 1.74e−13 | 5.96e−13 | 5.58e−14 | 9.83e−4 | 2.26e−2 | 9.51e−4 | 0.05 |
| rs1811359 | HLA | g.33084605G>C (GenBank: NC_000006.12) | +---++----++?++ | 0.04 (0.01) | 4.38e−26 | 8.29e−17 | 1.18e−8 | 5.39e−29 | 1.70e−14 | 8.76e−1 | 0.46 |
| rs1561929 | MYC | g.128555127C>T (GenBank: NC_000008.11) | ?++++?++++++??? | 0.71 (1.70) | 1.00e−6 | 2.96e−8 | 8.77e−8 | 4.28e−4 | 6.35e−1 | 5.48e−1 | 0.88 |
| rs77052733 | MAP3K9 | g.70819179C>T (GenBank: NC_000014.9) | -+++--++++-+?++ | 0.14 (0.04) | 1.54e−6 | 2.87e−8 | 1.22e−6 | 6.82e−5 | 2.97e−3 | 1.55e−1 | 0.07 |
| rs8107974 | TM6SF2 | g.19277691A>T (GenBank: NC_000019.10) | +++++-++++++?++ | 0.31 (0.03) | 1.95e−48 | 1.09e−36 | 2.00e−47 | 1.63e−2 | 5.70e−20 | 1.74e−2 | 0.09 |
| rs12983038 | IFNL3 | g.39250484G>A (GenBank: NC_000019.10) | ++++--+?++++?+- | 0.15 (0.03) | 1.91e−17 | 3.41e−15 | 6.51e−8 | 5.02e−7 | 3.83e−3 | 4.84e−2 | 0.13 |
| rs429358 | APOE | g.44908684T>C (GenBank: NC_000019.10) | -----+--------- | −0.17 (0.01) | 4.93e−16 | 3.87e−14 | 1.13e−16 | 3.16e−2 | 9.75e−5 | 8.42e−1 | 0.13 |
| rs738408 | PNPLA3 | g.43928850C>T (GenBank: NC_000022.11) | +++++-+++++++++ | 0.33 (0.02) | 1.06e−121 | 2.35e−96 | 4.41e−96 | 6.87e−13 | 1.42e−33 | 3.18e−2 | 0.31 |
| Suggestive | |||||||||||
| rs2642438 | MTARC1 | g.220796686A>G (GenBank: NC_000001.11) | -+++++++++++?+- | 0.07 (0.02) | 7.01e−5 | 2.57e−4 | 5.98e−7 | 2.05e−2 | 1.07e−2 | 2.24e−2 | 0.77 |
| rs7601754 | STAT4 | g.191075725G>A (GenBank: NC_000002.12) | --+++-++++--?-- | −0.01 (0.01) | 7.85e−6 | 5.43e−3 | 3.11e−4 | 1.43e−7 | 7.00e−5 | 8.11e−1 | 0.8 |
| rs59128789 | CD81 | g.2370530T>C (GenBank: NC_000011.10) | ++--++?-+--+?+- | 0.01 (0.08) | 2.97e−6 | 6.83e−4 | 3.20e−4 | 2.08e−3 | 1.34e−5 | 3.34e−1 | 0.04 |
| rs112635299 | SERPINA1 | g.94371805G>T (GenBank: NC_000014.9) | ?+++-?+?+++-??? | 8.45 (3.89) | 1.63e−7 | 1.11e−6 | 1.93e−7 | 3.98e−4 | 3.64e−2 | 5.65e−1 | 0.02 |
“Effects” indicates the direction of effect on meta-analysis for each study, where “+” represents a positive β, “-” a negative β, and “?” where this variant was not included in the dataset. Studies included under “Effects” in order (from left to right): Biobank Japan, Hassan et al.,5 DeCode, FinnGen, Intermountain, KoGES, PLCO, Trépo et al.,4 UKBB, MVP (EUR), MVP (AFR), MVP (AMR), All of Us, TPMI, and CKBB. CHR, chromosome; BP, basepair; NEA, non-effect allele; EA, effect allele. Beta (SE), allelic effect in log odds ratio from MR-MEGA analysis; p(MR-MEGA): two-sided p value of association from MR-MEGA (chi-squared test with df = 4); p(METAL-all), two-sided p value of association from sample-size weighted analysis using METAL with all ancestries; p(METAL-EUR/EAS), two-sided p value of association from sample-size weighted analysis using METAL with European (EUR) or East Asian (EAS) ancestries; p(ANC-HET), p value for the two-sided ancestral heterogeneity test (chi-squared test with df = 3); p(RES-HET), p value for the two-sided residual heterogeneity test (chi-squared test with df = 3); EAF, effect allele frequency. Data in bold are significant p values (p < 5 × 10−8 for the two-sided association tests, p < 6.6 × 10−4 for the heterogeneity tests).
Figure 2.
Miami plot of the meta-analysis results from 17,697 HCC affected individuals and 2,715,683 control subjects
Results from MR-MEGA (top) and METAL (including all ancestries, bottom), with genome-wide significant (p < 5 × 10−8) loci in red and top suggestive loci in orange.
Three GWAS-significant loci (GCKR [MIM: 600842], ADH-1B [MIM: 103720]/-4 [MIM: 103740]/-5 [MIM: 103710], and MTTP [MIM: 157147]) are well-established risk genes for liver disease. GCKR (lead SNP rs780094 [g.27518370T>C] [GenBank: NC_000002.12]) is associated with MASLD,10 and plasma protein levels of GCKR are associated with HCC on Mendelian randomization analyses,7 though not with cirrhosis (Table S5). The ADH-1B/-4/-5 locus (lead SNP rs6822348 [g.99132743A>T] [GenBank: NC_000004.12]) is a genome-wide significant locus for cirrhosis,29 confirmed through colocalization analysis (probability of shared causal variant [PP.H4.abf] = 0.90, Table S5), which may be mediated through the effect on alcohol intake (Figure 3A). The causal gene in this region remains unclear as, while typically annotated as ADH1B or ADH4, rs6822348 has a liver-specific eQTL for ADH5 (−0.17 normalized effect size, p = 8.1 × 10−8, from GTEx). It should be noted that Vujkovic et al.7 identified a different, coding lead variant in ADH1B, rs1229984 (g.99318162T>A [GenBank: NC_000004.12]) (c.143A>T [GenBank: NM_000668.6] [p.His48Arg]) encoding p.His48Arg to be associated with cirrhosis, which is not in LD with rs6822348 (R2 < 0.1), suggesting there may be more than one effect in this locus.
Figure 3.
Region plots of genome-wide significant associations
LocusZoom plots of genome-wide significant (p < 5 × 10−8) loci demonstrating individual study effects for ADH5 (A), MTTP (B), 8q24.21 (nearest gene MYC) (C), and MAP3K9 (D).
MTTP (lead SNP rs112850336 [g.99492512C>T] [GenBank: NC_000004.12]) is central to hepatic triglyceride metabolism, which is genome-wide significant for MASLD.10 MTTP is less than 1 Mb upstream of ADH5, but these genes were demonstrated to have independent effects on liver fat in the MASLD GWAS.10 Here, we observed separate peaks, and the lead variants (rs6822348 near ADH5 and rs112850336 near MTTP) were not in LD (R2 = 0.1, Figures 3A and 3B).
SNP rs1561929 (g.128555127C>T [GenBank: NC_000008.11)]) demonstrated consistent effects across all cohorts where present. The nearest protein-coding gene is MYC (MIM: 190080) (Figure 3C), located approximately 1 Mb away within the 8q24 locus. MYC is amplified or has gain-of-function mutations in 227 out of 373 (61%) HCC individuals based on The Cancer Genome Atlas Program (TCGA) data. MYC is a common essential gene (including liver cancer cell lines) in DepMap41 and has a probability of having a loss-of-function intolerant (pLI) score of 1.0.42
Variants in MAP3K9 (MIM: 600136) lead SNP rs77052733 (g.70819179C>T [GenBank: NC_000014.9]) were also positively associated with HCC (Figure 3D). 110 of 373 (29.5%) HCC individuals from TCGA have heterozygous deletion of MAP3K9.
Lastly, we identified a GWAS-significant locus for EUR ancestry at near GABPB2 (MIM: 621284) lead SNP rs77018356 (g.151162485G>C [GenBank: NC_000001.11]) (Figures S3A and S3B). There was no evidence of colocalization with variants causing cirrhosis for 8q24.21 (nearest gene MYC), MAP3K9, EPHA2 (MIM: 176946), or GABPB2 loci.
Our analyses also replicated loci in SERPINA1 (MIM: 107400), CD81 (MIM: 186845), STAT4 (MIM: 600558), and MTARC1 at least suggestive significance (p < 1 × 10−6), in addition to 57 additional loci (Table S4; Figures S3C and S3D). Prioritization of these loci based on pLI found 13 genes with pLI 0.99–1, including HNF4G (MIM: 605966), MEN1 (MIM: 613733) (Figures S3E and S3F), and BRINP3, which may be worthy of further study.
Cirrhosis-independent effects
Most signals found are in loci previously linked to development of cirrhosis; therefore, we ran a sensitivity analysis in two cohorts with cirrhotic controls (Trépo et al.4 and Hassan et al.5). We found consistent effects in APOE (MIM: 107741), EPHA2, GABPB2, HSD17B13 (MIM: 612127), MAP3K9, PNPLA3, TERT, and TM6SF2, suggesting that these loci increase risk of HCC even compared to cirrhotic controls (Table S6). There were no variants included in both studies suitable as a proxy for rs1800562 (g.26092913G>A [GenBank: NC_000006.12]) in HFE (MIM: 613609). The effects of ADH5, GCKR, HLA region, IFNL3 (MIM: 607402), and 8q24.21 (nearest gene MYC) were attenuated, suggesting that some of their effect may be via underlying cirrhosis, although there was reduced power in this two-study sensitivity analysis.
MAMA identifies heterogeneous effects across ancestries
Nine loci showed homogeneous effects across different ancestries (pANC-HET > 6.6 × 10−4), including variants in or near TERT, GCKR, ADH5, MAP3K9, MTTP, and 8q24.21 (Figures 4A–4C and Table 1). For six loci (PNPLA3, TM6SF2, EPHA2, HSD17B13, APOE, and HLA), we found significant ancestral heterogeneity (pANC-HET < 6.6 × 10−4) without residual heterogeneity (pRES-HET > 6.6 × 10−4), suggesting that the results are due to population structural differences instead of other confounders (Figure 4D).
Figure 4.
Forest plots of genome-wide significant associations
Forest plots of genome-wide significant (p < 5 × 10−8) loci (top) with forest plots demonstrating individual study effects (below) for ADH5 (A), MTTP (B), rs1561929 (8q24.21 nearest gene MYC) (C), and PNPLA3 (D). rs738408 near PNPLA3 demonstrated significant ancestral heterogeneity (p < 6.6 × 10−4). BBJ, BioBank Japan; CKBB, China Kadoorie BioBank; KoGES, Korean BioBank; MVP, Million Veterans Program; PLCO, Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial; TPMI, Taiwan Precision Medicine Initiative; UKBB, UK Biobank. Annotated with genetic ancestry: AFR, African; AMR, Admixed American; EAS, East Asian; EUR, European.
Of note, the HLA locus (lead SNP rs1811359; g.33084605G>C [GenBank: NC_000006.12], pANC-HET = 1.7 × 10−14) has strongly positive-effect estimates in EAS cohorts while having negative-effect estimates in EUR ancestry cohorts (Figure S3G). The AFR and AMR cohorts demonstrated neutral effects at the HLA locus.
Five loci had significant effects in EUR ancestry cohorts and MAMA but no effect in EAS cohorts. The TM6SF2 locus (lead SNP rs8107974, pANC-HET = 5.7 × 10−20) has strongly positive in AMR and EUR ancestry cohorts but no effect in AFR and EAS cohorts (pMETAL-EAS = 0.016, Figure S3H). Similarly, the EPHA2, APOE, HFE, and HSD17B13 loci had no significant effects in EAS cohorts (p > 1 × 10−5) but were directionally consistent with genome-wide significant effects in EUR cohorts. These observations likely reflect the influence of the variants on cirrhosis, as the primary mediator of HCC, in their respective ancestries.
Fine mapping identifies 11 missense variants
MR-MEGA was used to generate posterior probabilities for fine mapping, which found 28 significant or suggestive loci that had fewer than five variants in the 95% credible set. Within this, we found putative causal coding variants in 11 loci (Tables 2 and S7). This includes the well-established missense mutations in PNPLA3, TM6SF2, HFE, APOE, SERPINA1, and MTARC1. In addition, identified GCKR p.Leu446Pro (c.1337T>C) as the likely causal variant; this missense mutation is strongly associated with insulin resistance and hepatic steatosis.
Table 2.
Coding variants associated with population risk of HCC identified from fine mapping
| chr:start:end | Variant | p value | Gene | Protein change | CADD |
|---|---|---|---|---|---|
| 22:43927717:44014293 | g.43928847C>G (GenBank: NC_000022.11) (c.444C>G [GenBank: NM_025225.3] [p.Ile148Met]) | 1.06e−121 (MR-MEGA) | PNPLA3 | p.Ile148Met | 13.4 |
| 19:18952750:20375946 | g.19268740C>T (GenBank: NC_000019.10) (c.499G>A [GenBank: NM_001001524.3] [p.Glu167Lys]) | 1.95e−48 (MR-MEGA) | TM6SF2 | p.Glu167Lys | 24.6 |
| 19:44883210:44925202 | g.44908684T>C (GenBank: NC_000019.10) (c.388T>C [GenBank: NM_000041.4] [p.Cys130Arg]) | 4.93e−16 (MR-MEGA) | APOE | p.Cys130Arg | 16.6 |
| 6:25526091:27552168 | g.26092913G>A (GenBank: NC_000006.12) (c.845G>A [GenBank: NM_000410.4] [p.Cys282Tyr]) | 1.74e−13 (MR-MEGA) | HFE | p.Cys282Tyr | 27 |
| 2:27375230:27529596 | g.27508073T>C (GenBank: NC_000002.12) (c.1337T>C [GenBank: NM_001486.4] [p.Leu446Pro]) | 8.69e−9 (METAL-all) | GCKR | p.Leu446Pro | 10.99 |
| 10:127419693:127419693 | g.127419693G>A (GenBank: NC_000010.11) (c.4657G>A [GenBank: NM_001380.5] [p.Glu1553Lys]) | 3.35e−7 (MR-MEGA) | DOCK1 | p.Glu1553Lys | 23.6 |
| 19:8359282:8386830 | g.8369239G>C (GenBank: NC_000019.10) (c.568G>C [GenBank: NM_139314.3] [p.Glu190Gln]) | 4.93e−7 (MR-MEGA) | ANGPTL4 | p.Glu190Gln | 19.01 |
| 1:220796686:220800419 | g.220796686A>G (GenBank: NC_000001.11) (c.493A>G [GenBank: NM_022746.4] [p.Thr165Pro]) | 5.98e−7 (METAL-EUR) | MTARC1 | p.Thr165Pro | 18.3 |
| 6:31354452:31359387 | g.31355519A>G (GenBank: NC_000006.12) (c.693T>C [GenBank: NM_005514.8] [p.Gly231=]) | 4.38e−26 (MR-MEGA) | HLA-B | p.Gly231= | – |
| 14:94206394:94378610 | g.94378610C>T (GenBank: NC_000014.9) (c.1096G>A [GenBank: NM_000295.5] [p.Glu366Gln]) | 1.63e−7 (MR-MEGA) | SERPINA1 | p.Glu366Gln | 24.2 |
| 11:64800220:64813774 | g.64805130G>A (GenBank: NC_000011.10) (c.1269C>T [GenBank: NM_000244.4] [p.Asp423Glu]) | 2.25e−6 (MR-MEGA) | MEN1 | p.Asp423Glu | 19.25 |
Posterior probabilities were calculated from Bayes scores from MR-MEGA meta-analysis results for all significant and suggestive loci. Loci were filtered for those with fewer than five variants per 95% credible set. Ensembl Variant Effect Predictor was used to identify common coding variants from those credible sets. CADD, combined annotation-dependent depletion score; CHR, chromosome.
Across suggestive loci, we found MEN1 p.Asp423Glu (c.1269C>T) and ANGPTL4 (MIM: 605910) p.Glu190Gln (c.568G>C) missense variants. MEN1 is a tumor-suppressor gene, where heterozygous loss-of-function mutations cause multiple endocrine neoplasia type 1. ANGPTL4 is a circulating protein that regulates angiogenesis and inactivation of lipoprotein lipase and therefore influences serum lipid levels.43,44
DHRS1 is associated with HCC on transcriptome-wide analysis
To move from statistical association to candidate gene and mechanism, we applied two complementary functional analyses to the meta-analysis results. The GWAS identifies genomic loci associated with HCC risk at a population level but does not itself indicate which genes or regulatory elements mediate that risk. TWAS analysis leverages tissue-specific eQTL data from liver and HCC tissue to test whether germline variation at GWAS loci influences gene expression. RWAS analysis extends this by asking whether germline variants alter chromatin accessibility in HCC tissue, identifying regulatory mechanisms that may mediate risk independently of coding variation. Together, GWAS provides locus discovery, TWAS provides transcriptomic functional annotation, and RWAS provides regulatory annotation: three complementary layers of evidence linking germline variation to HCC risk.
The TWAS using FUSION (Table S9) replicated the above observations for EPHA2 and TM6SF2. In addition, we found a TWAS signal in DHRS1 (MIM: 610410), lead GWAS-SNP rs2295308 (g.24292599A>C [GenBank: NC_000014.9]) (rs2295308; intergenic/upstream of DHRS1, p-adjTWAS = 6.3 × 10−5) based on MR-MEGA meta-analysis results and HCC-specific eQTL (peQTL = 5.1 × 10−5). DHRS1 is thought to have oxidoreductase activity,45 although its function in the liver or malignancy has not yet been characterized.
The RWAS discovered signals for the above GWAS-significant genes HSD17B13, HFE, TM6SF2, and PNPLA3 (Table S10). In addition, we found a highly significant chromatin accessible region at chr9:96436312–96436813 (lead GWAS-SNP rs16911041 (g.96516265C>A [GenBank: NC_000009.12]) (rs16911041; intergenic), p-adjRWAS = 2.7 × 10−24). The nearest protein-coding genes are ZNF367 (MIM: 610160), a transcription factor with tumor-suppressor properties,46 and HABP4 (MIM: 617369), which is implicated in familial colorectal cancer.47
Gene set analysis shows enrichment in liver
We ran MAGMA for Gene Ontology and tissue-level and single-cell expression data.33 Enriched Gene Ontology sets in MSigDB v.7.0 were dominated by those related to antigen presentation given mapping to multiple HLA genes (Figure S4A and Table S8). In addition, we observed enrichment of the β-catenin-TCF transactivating complex based on TERT, MEN1, and histone genes (p-adj = 0.03) and GWAS Catalog signatures related to chronic liver disease. Differentially expressed genes were preferentially expressed in liver tissue, based on GTEx data (Figure S4B).
HCC shares genetic architecture with chronic liver disease but not other gastrointestinal cancers
Finally, we used LD score regression to quantify the similarity in genetic architecture between HCC, liver disorders, and other gastrointestinal cancers (Figure 5). There was a strong positive genomic correlation between HCC and cirrhosis (rg = 0.87, p = 5.0 × 10−10), HBV (rg = 0.84, p = 0.003), HCV (rg = 0.65, p = 0.004), and MASLD (rg = 0.50, p = 6.0 × 10−4), reflecting the underlying drivers of HCC in the population. There was no significant correlation between HCC and upper GI cancer (esophageal and gastric) or colorectal cancer. There were shared genetic determinants between HCC and T2DM (rg = 0.25, p = 6.8 × 10−5), while there was inverse association between HCC and BMI (rg = −0.17, p = 6.0 × 10−4), suggesting that genetically lower BMI is linked to increased genetic risk of HCC.
Figure 5.
Similarity in genetic architecture between HCC and other gastrointestinal traits
LD score regression results between HCC (from this study) and previously published GWASs for chronic liver disease,16,19,29,36 metabolic disease (type 2 diabetes mellitus [T2DM],38 body mass index [BMI],37 and upper gastrointestinal [GI] cancer [esophageal and gastric]),39 and colorectal carcinoma.40
Discussion
This large GWAS meta-analysis of all-cause HCC has discovered evidence for germline involvement of several genes involved in HCC, most notably 8q24 (nearest gene MYC) and MAP3K9. We also found genome-wide significant germline associations for HCC in GCKR, ADH5, and MTTP. In addition, we have dissected diverse effects across ancestries and have leveraged these results to fine-map 11 missense variants. Finally, we demonstrate the importance of germline risk for cirrhosis, T2DM, and low BMI as drivers of population risk for HCC.
The 8q24.21 locus has been recognized as a gene desert harboring multiple independent cancer risk loci for prostate,48,49 colorectal,50 and breast cancer, none of which align to protein-coding genes.51,52 Functional studies have demonstrated that these loci act as tissue-specific long-range enhancers that physically interact with MYC through chromatin looping over several hundred kilobases. The independent signals in different cancer types are not in LD with each other, illustrating that this region harbors multiple tissue-specific regulatory elements. This is of particular relevance in HCC, as MYC is a transcription factor that lies downstream of multiple cell-surface signaling pathways, including the WNT-β-catenin pathway,53 which is robustly implicated in HCC development.54,55,56 Variation in 8q24 MYC confers increased susceptibility to WNT signaling in other malignancies.57 We propose that rs1561929 represents an HCC-specific signal at 8q24 operating through an analogous mechanism. It is unclear whether its effect is dependent on cirrhosis or not; there was no evidence of colocalization, but it was not significant on sensitivity analysis using only cirrhotic controls.
The mitogen-activated protein kinase kinase kinases (MAP3K) are a group of proteins that act downstream of cellular stress signals and growth factor receptors. They function as part of a signaling cascade that ultimately leads to pro-growth effects. MAP3K9 primarily transduces cellular stress through to JUN and GATA4, which influences mitochondrial cytochrome c release and, therefore, apoptosis.58,59 Mutations in MAP3K9 have been identified in lung cancer and melanoma,60,61 as well as in HCC, via the TCGA dataset. It is not known precisely which stress signals regular MAP3K9 in the liver, although interactions with a variety of microRNAs have been proposed.62,63
Prevalence of HCC varies greatly across the globe, from 1 in 100,000 in Morocco to 16 in 100,000 in the Republic of Korea.64 While this is driven by differences in the prevalence of chronic viral hepatitis, alcohol intake, and obesity, our analysis has uncovered ancestral heterogeneity in germline contribution to HCC risk. Specific HLA loci showed divergent effects in European and East Asian ancestries. Also, in individuals of East Asian ancestry, we observed no significant effect from TM6SF2, which was the most significant locus in Europeans.65 Some of these differences may reflect the underlying etiology of HCC in these regions (i.e., viral hepatitis versus steatotic liver disease).
Genetically driven accumulation of hepatic lipid droplets has been associated with severity of almost every studied liver disease. In this analysis, we found evidence for six genes (MTTP, PNPLA3, TM6SF2, HSD17B13, GCKR, and APOE) that are all genome-wide risk genes for MASLD.10 Gene-environment studies have previously shown that the effect of such variants is amplified by the presence of obesity, T2DM, and excess alcohol consumption.66 The stronger effects observed in European cohorts may reflect the higher prevalence of steatotic liver disease in these populations.67 It should also be noted that non-cirrhotic HCC is well described in steatotic liver disease.68 These data underscore the importance of public health interventions to tackle HCC.
The principal strengths of this study are its size and TWAS/RWAS analyses. The integration of GWAS, TWAS, and RWAS provides three complementary layers of functional evidence. EPHA2 was identified in both the GWAS and RWAS, providing independent genomic and regulatory support. DHRS1 emerged specifically through TWAS, demonstrating that germline-encoded transcriptomic variation can identify candidate genes not captured by positional mapping alone. The ZNF367/HABP4 locus at chromosome 9 was identified exclusively through RWAS, representing a chromatin-accessible region that would not have been detected by a GWAS or TWAS. Further mechanistic work is required to understand the function of these loci in HCC development.
This study would be improved by inclusion of individual-participant-level data, which would allow for mediation analyses as performed by Vujkovic et al.7 Also, most cohorts used electronic health records for identification of cases, which carries the risk of misclassification. GWAS in the All of Us cohort was performed using Wald logistic regression, which can produce inflated test statistics under case/control imbalance. Although the genomic inflation for this cohort was nominal (λ = 1.008) and All of Us contributed no genome-wide significant loci to the meta-analysis, future analyses should consider Firth correction or saddlepoint approximation methods. In addition, we have been unable to robustly identify the causal gene at 8q24.21. We propose that MYC is the most likely gene for the rs1561929 variant, but functional work is required to establish this.
Despite the inclusion of multiple ancestries, individuals of European ancestry are still over-represented and account for 49% of HCC cases in this analysis. Individuals of African descent are particularly under-represented, with only 6% of all affected individuals and no data from the continent of Africa, which has some of the highest rates of HCC.69 We are also lacking data from the Middle East and South America, where HCC are major causes of cancer death. Addition of data from population biobank studies from these areas will strengthen future MAMA.
Most cohorts included in this analysis had healthy population controls; therefore, this meta-analysis reflects the population-level drivers of HCC, and many loci are already implicated in the development of underlying liver disease (e.g., TM6SF2). Our sensitivity analysis with cirrhosis controls and colocalization studies demonstrated that most significant loci, including MAP3K9, appear to act independently of cirrhosis. Larger meta-analyses with cirrhotic control individuals or etiology-specific studies may identify further treatment targets.
Conclusion
HCC is driven by heterogeneous and homogeneous genetic factors across ancestries. Germline susceptibility to β-catenin pathway activation influences HCC risk via MEN1, TERT, and, potentially, the 8q24 locus (mediated by MYC). MAP3K9 and DHRS1 are HCC risk genes worthy of further investigation. Genes involved in hepatic lipid droplet formation, including MTTP, contribute to the risk of HCC. There is a need for further ancestral diversity in liver GWASs.
Data code and availability
Raw data for Hassan et al.5 are available via dbGaP phs001744.v1.p1. All other summary statistics are publicly available via the aforementioned links. Full summary statistics from this analysis can be found on GWAS Catalog (https://www.ebi.ac.uk/gwas/home) using the accession numbers NHGRI: GCST90860789, GCST90860790, GCST90860791, and GCST90860792. All code used in analysis is available at https://github.com/jmann01/HCC-gwas.
Acknowledgments
J.P.M. is supported by grants from Birmingham Health Partners, NIHR ACL (CL-2022-09-005), The Royal Society (RG\R1\241312), BSPGHAN-Guts (BSPGHAN2023_01), ESPGHAN, Royal College of Physicians Dame Sheila Sherlock Bursary, and Little Princess Trust (CCLGA 2024 10 Mann). S.S. is supported by a Cancer Research UK Advanced Clinician Scientist Fellowship (C53575/A29959). We thank all the participants, investigators, and funders who contributed to the studies included in this meta-analysis. The research was carried out at the National Institute for Health and Care Research (NIHR) Birmingham Biomedical Research Centre (BRC).
Declaration of interests
V.L.C. received grants from AstraZeneca, Takeda, Ipsen, and KOWA, paid to the University of Michigan. S.S. has received honorarium from Astra Zeneca.
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.xhgg.2026.100639.
Web resources
China Kadoorie Biobank (CKBB) PheWeb, https://pheweb.ckbiobank.org/download/c22
deCODE/UK Biobank/Intermountain summary statistics, https://www.decode.com/summarydata/
DepMap: The Cancer Dependency Map, https://depmap.org/portal/
FinnGen summary statistics, https://www.finngen.fi/en
GWAS Catalog, https://www.ebi.ac.uk/gwas/
gwasglue, https://github.com/MRCIEU/gwasglue
Hail genomic analysis framework, https://github.com/hail-is/hail
HCC GWAS analysis code (this study), https://github.com/jmann01/HCC-gwas
Korean BioBank (KoGES) summary statistics, https://koges.leelabsg.org/
Million Veterans Program (MVP) summary statistics, https://ftp.ncbi.nlm.nih.gov/dbgap/studies/phs002453/analyses/GIA/
Online Mendelian Inheritance in Man (OMIM), https://www.omim.org
PLCO Atlas, https://exploregwas.cancer.gov/plco-atlas
Taiwan Precision Medicine Initiative (TPMI) PheWeb, https://pheweb.ibms.sinica.edu.tw/pheno/155.1
Supplemental information
References
- 1.Bray F., Laversanne M., Sung H., Ferlay J., Siegel R.L., Soerjomataram I., Jemal A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin. 2024;74:229–263. doi: 10.3322/caac.21834. [DOI] [PubMed] [Google Scholar]
- 2.Singal A.G., Kanwal F., Llovet J.M. Global trends in hepatocellular carcinoma epidemiology: implications for screening, prevention and therapy. Nat. Rev. Clin. Oncol. 2023;20:864–884. doi: 10.1038/s41571-023-00825-3. [DOI] [PubMed] [Google Scholar]
- 3.Llovet J.M., Pinyol R., Kelley R.K., El-Khoueiry A., Reeves H.L., Wang X.W., Gores G.J., Villanueva A. Molecular pathogenesis and systemic therapies for hepatocellular carcinoma. Nat. Cancer. 2022;3:386–401. doi: 10.1038/s43018-022-00357-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Trépo E., Caruso S., Yang J., Imbeaud S., Couchy G., Bayard Q., Letouzé E., Ganne-Carrié N., Moreno C., Oussalah A., et al. Common genetic variation in alcohol-related hepatocellular carcinoma: a case-control genome-wide association study. Lancet Oncol. 2022;23:161–171. doi: 10.1016/S1470-2045(21)00603-3. [DOI] [PubMed] [Google Scholar]
- 5.Hassan M.M., Li D., Han Y., Byun J., Hatia R.I., Long E., Choi J., Kelley R.K., Cleary S.P., Lok A.S., et al. Genome-wide association study identifies high-impact susceptibility loci for HCC in North America. Hepatology. 2024;80:87–101. doi: 10.1097/HEP.0000000000000800. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ghouse J., Gellert-Kristensen H., O’Rourke C.J., Seidelin A.-S., Thorleifsson G., Sveinbjörnsson G., Tragante V., Konkwo C., Brancale J., Vilarinho S., et al. Genome-wide meta-analysis identifies nine loci associated with higher risk of hepatocellular carcinoma development. JHEP Rep. 2025;7 doi: 10.1016/j.jhepr.2025.101485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Vujkovic M., Kaplan D.E., Ghouse J., Loza B.-L., Brancale J., Lewis A., Zhang D.Y., Levin M.G., Veatch O.J., Johnson J.P., et al. Germline variants influence chronic liver disease progression through distinct pathways. medRxiv. 2025 doi: 10.1101/2025.09.16.25335186. Preprint at. [DOI] [Google Scholar]
- 8.Verweij N., Haas M.E., Nielsen J.B., Sosina O.A., Kim M., Akbari P., De T., Hindy G., Bovijn J., Persaud T., et al. Germline Mutations in CIDEB and Protection against Liver Disease. N. Engl. J. Med. 2022;387:332–344. doi: 10.1056/NEJMoa2117872. [DOI] [PubMed] [Google Scholar]
- 9.Emdin C.A., Haas M.E., Khera A.V., Aragam K., Chaffin M., Klarin D., Hindy G., Jiang L., Wei W.Q., Feng Q., et al. A missense variant in Mitochondrial Amidoxime Reducing Component 1 gene and protection against liver disease. PLoS Genet. 2020;16 doi: 10.1371/journal.pgen.1008629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Chen Y., Du X., Kuppa A., Feitosa M.F., Bielak L.F., O’Connell J.R., Musani S.K., Guo X., Kahali B., Chen V.L., et al. Genome-wide association meta-analysis identifies 17 loci associated with nonalcoholic fatty liver disease. Nat. Genet. 2023;55:1640–1650. doi: 10.1038/s41588-023-01497-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Buch S., Innes H., Lutz P.L., Nischalke H.D., Marquardt J.U., Fischer J., Weiss K.H., Rosendahl J., Marot A., Krawczyk M., et al. Genetic variation in TERT modifies the risk of hepatocellular carcinoma in alcohol-related cirrhosis: results from a genome-wide case-control study. Gut. 2023;72:381–391. doi: 10.1136/gutjnl-2022-327196. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Sveinbjornsson G., Ulfarsson M.O., Thorolfsdottir R.B., Jonsson B.A., Einarsson E., Gunnlaugsson G., Rognvaldsson S., Arnar D.O., Baldvinsson M., Bjarnason R.G., et al. Multiomics study of nonalcoholic fatty liver disease. Nat. Genet. 2022;54:1652–1663. doi: 10.1038/s41588-022-01199-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Yang H.-C., Kwok P.-Y., Li L.-H., Liu Y.-M., Jong Y.-J., Lee K.-Y., Wang D.W., Tsai M.F., Yang J.H., Chen C.H., et al. The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies. Nature. 2025;648:117–127. doi: 10.1038/s41586-025-09680-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Chen H.-H., Chen C.-H., Hou M.-C., Fu Y.-C., Li L.-H., Chou C.-Y., Yeh E.C., Tsai M.F., Chen C.H., Yang H.C., et al. Population-specific polygenic risk scores for people of Han Chinese ancestry. Nature. 2025;648:128–137. doi: 10.1038/s41586-025-09350-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Walters R.G., Millwood I.Y., Lin K., Schmidt Valle D., McDonnell P., Hacker A., Avery D., Edris A., Fry H., Cai N., et al. Genotyping and population characteristics of the China Kadoorie Biobank. Cell Genom. 2023;3 doi: 10.1016/j.xgen.2023.100361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Sakaue S., Kanai M., Tanigawa Y., Karjalainen J., Kurki M., Koshiba S., Narita A., Konuma T., Yamamoto K., Akiyama M., et al. A cross-population atlas of genetic associations for 220 human phenotypes. Nat. Genet. 2021;53:1415–1424. doi: 10.1038/s41588-021-00931-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Nam K., Kim J., Lee S. Genome-wide study on 72,298 individuals in Korean biobank data for 76 traits. Cell Genom. 2022;2 doi: 10.1016/j.xgen.2022.100189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Machiela M.J., Huang W.-Y., Wong W., Berndt S.I., Sampson J., De Almeida J., Abubakar M., Hislop J., Chen K.L., Dagnall C., et al. GWAS Explorer: an open-source tool to explore, visualize, and access GWAS summary statistics in the PLCO Atlas. Sci. Data. 2023;10:25. doi: 10.1038/s41597-022-01921-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Verma A., Huffman J.E., Rodriguez A., Conery M., Liu M., Ho Y.-L., Kim Y., Heise D.A., Guare L., Panickan V.A., et al. Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program. Science. 2024;385 doi: 10.1126/science.adj1182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.All of Us Research Program Genomics Investigators Genomic data in the All of Us Research Program. Nature. 2024;627:340–346. doi: 10.1038/s41586-023-06957-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.All of Us Research Program Investigators. Denny J.C., Rutter J.L., Goldstein D.B., Philippakis A., Smoller J.W., Jenkins G., Dishman E. The “All of Us” Research Program. N. Engl. J. Med. 2019;381:668–676. doi: 10.1056/NEJMsr1809937. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Murphy A.E., Schilder B.M., Skene N.G. MungeSumstats: a Bioconductor package for the standardization and quality control of many GWAS summary statistics. Bioinformatics. 2021;37:4593–4596. doi: 10.1093/bioinformatics/btab665. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.GitHub - hail-is/hail Cloud-native genomic dataframes and batch computing [Internet] https://github.com/hail-is/hail GitHub.
- 24.Mägi R., Horikoshi M., Sofer T., Mahajan A., Kitajima H., Franceschini N., McCarthy M.I., COGENT-Kidney Consortium, T2D-GENES Consortium. Morris A.P. Trans-ethnic meta-regression of genome-wide association studies accounting for ancestry increases power for discovery and improves fine-mapping resolution. Hum. Mol. Genet. 2017;26:3639–3650. doi: 10.1093/hmg/ddx280. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Willer C.J., Li Y., Abecasis G.R. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26:2190–2191. doi: 10.1093/bioinformatics/btq340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.R Core Team . R Foundation for Statistical Computing; Vienna, Austria: 2019. A Language and Environment for Statistical Computing. [Google Scholar]
- 27.Watanabe K., Taskesen E., van Bochoven A., Posthuma D. Functional mapping and annotation of genetic associations with FUMA. Nat. Commun. 2017;8:1826. doi: 10.1038/s41467-017-01261-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Watanabe K., Umićević Mirkov M., de Leeuw C.A., van den Heuvel M.P., Posthuma D. Genetic mapping of cell type specificity for complex traits. Nat. Commun. 2019;10:3222. doi: 10.1038/s41467-019-11181-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Ghouse J., Sveinbjörnsson G., Vujkovic M., Seidelin A.-S., Gellert-Kristensen H., Ahlberg G., Tragante V., Rand S.A., Brancale J., Vilarinho S., et al. Integrative common and rare variant analyses provide insights into the genetic architecture of liver cirrhosis. Nat. Genet. 2024;56:827–837. doi: 10.1038/s41588-024-01720-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.GitHub - MRCIEU/gwasglue Linking GWAS data to analytical tools in R [Internet] https://github.com/MRCIEU/gwasglue GitHub.
- 31.Kim J.J., Vitale D., Otani D.V., Lian M.M., Heilbron K., Iwaki H., Lake J., Solsberg C.W., Leonard H., Makarious M.B., et al. Multi-ancestry genome-wide association meta-analysis of Parkinson’s disease. Nat. Genet. 2024;56:27–36. doi: 10.1038/s41588-023-01584-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.McLaren W., Gil L., Hunt S.E., Riat H.S., Ritchie G.R.S., Thormann A., Flicek P., Cunningham F. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122. doi: 10.1186/s13059-016-0974-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.de Leeuw C.A., Mooij J.M., Heskes T., Posthuma D. MAGMA: generalized gene-set analysis of GWAS data. PLoS Comput. Biol. 2015;11 doi: 10.1371/journal.pcbi.1004219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Gusev A., Ko A., Shi H., Bhatia G., Chung W., Penninx B.W.J.H., Jansen R., de Geus E.J.C., Boomsma D.I., Wright F.A., et al. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 2016;48:245–252. doi: 10.1038/ng.3506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Grishin D., Gusev A. Allelic imbalance of chromatin accessibility in cancer identifies candidate causal risk variants and their mechanisms. Nat. Genet. 2022;54:837–849. doi: 10.1038/s41588-022-01075-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Liu Y., Basty N., Whitcher B., Bell J.D., Sorokin E.P., van Bruggen N., Thomas E.L., Cule M. Genetic architecture of 11 organ traits derived from abdominal MRI using deep learning. eLife. 2021;10 doi: 10.7554/eLife.65554. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Pulit S.L., Stoneman C., Morris A.P., Wood A.R., Glastonbury C.A., Tyrrell J., Yengo L., Ferreira T., Marouli E., Ji Y., et al. Meta-Analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry. Hum. Mol. Genet. 2019;28:166–174. doi: 10.1093/hmg/ddy327. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Huerta-Chagoya A., Schroeder P., Mandla R., Li J., Morris L., Vora M., Alkanaq A., Nagy D., Szczerbinski L., Madsen J.G.S., et al. Rare variant analyses in 51,256 type 2 diabetes cases and 370,487 controls reveal the pathogenicity spectrum of monogenic diabetes genes. Nat. Genet. 2024;56:2370–2379. doi: 10.1038/s41588-024-01947-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Rashkin S.R., Graff R.E., Kachuri L., Thai K.K., Alexeeff S.E., Blatchins M.A., Cavazos T.B., Corley D.A., Emami N.C., Hoffman J.D., et al. Pan-cancer study detects genetic risk variants and shared genetic basis in two large cohorts. Nat. Commun. 2020;11 doi: 10.1038/s41467-020-18246-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Fernandez-Rozadilla C., Timofeeva M., Chen Z., Law P., Thomas M., Schmit S., Díez-Obrero V., Hsu L., Fernandez-Tajes J., Palles C., et al. Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and east Asian ancestries. Nat. Genet. 2023;55:89–99. doi: 10.1038/s41588-022-01222-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.DepMap Portal DepMap: The Cancer Dependency Map Project at Broad Institute. https://depmap.org/portal/
- 42.Chen S., Francioli L.C., Goodrich J.K., Collins R.L., Kanai M., Wang Q., Alföldi J., Watts N.A., Vittal C., Gauthier L.D., et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625:92–100. doi: 10.1038/s41586-023-06045-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Gealekman O., Burkart A., Chouinard M., Nicoloro S.M., Straubhaar J., Corvera S. Enhanced angiogenesis in obesity and in response to PPARgamma activators through adipocyte VEGF and ANGPTL4 production. Am. J. Physiol. Endocrinol. Metab. 2008;295:E1056–E1064. doi: 10.1152/ajpendo.90345.2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Mysling S., Kristensen K.K., Larsson M., Kovrov O., Bensadouen A., Jørgensen T.J., Olivecrona G., Young S.G., Ploug M. The angiopoietin-like protein ANGPTL4 catalyzes unfolding of the hydrolase domain in lipoprotein lipase and the endothelial membrane protein GPIHBP1 counteracts this unfolding. eLife. 2016;5 doi: 10.7554/eLife.20958. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zemanová L., Navrátilová H., Andrýs R., Šperková K., Andrejs J., Kozáková K., Meier M., Möller G., Novotná E., Šafr M., et al. Initial characterization of human DHRS1 (SDR19C1), a member of the short-chain dehydrogenase/reductase superfamily. J. Steroid Biochem. Mol. Biol. 2019;185:80–89. doi: 10.1016/j.jsbmb.2018.07.013. [DOI] [PubMed] [Google Scholar]
- 46.Jain M., Zhang L., Boufraqech M., Liu-Chittenden Y., Bussey K., Demeure M.J., Wu X., Su L., Pacak K., Stratakis C.A., Kebebew E. ZNF367 inhibits cancer progression and is targeted by miR-195. PLoS One. 2014;9 doi: 10.1371/journal.pone.0101423. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Gray-McGuire C., Guda K., Adrianto I., Lin C.P., Natale L., Potter J.D., Newcomb P., Poole E.M., Ulrich C.M., Lindor N., et al. Confirmation of linkage to and localization of familial colon cancer risk haplotype on chromosome 9q22. Cancer Res. 2010;70:5409–5418. doi: 10.1158/0008-5472.CAN-10-0188. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Yeager M., Orr N., Hayes R.B., Jacobs K.B., Kraft P., Wacholder S., Minichiello M.J., Fearnhead P., Yu K., Chatterjee N., et al. Genome-wide association study of prostate cancer identifies a second risk locus at 8q24. Nat. Genet. 2007;39:645–649. doi: 10.1038/ng2022. [DOI] [PubMed] [Google Scholar]
- 49.Haiman C.A., Patterson N., Freedman M.L., Myers S.R., Pike M.C., Waliszewska A., Neubauer J., Tandon A., Schirmer C., McDonald G.J., et al. Multiple regions within 8q24 independently affect risk for prostate cancer. Nat. Genet. 2007;39:638–644. doi: 10.1038/ng2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Pomerantz M.M., Ahmadiyeh N., Jia L., Herman P., Verzi M.P., Doddapaneni H., Beckwith C.A., Chan J.A., Hills A., Davis M., et al. The 8q24 cancer risk variant rs6983267 shows long-range interaction with MYC in colorectal cancer. Nat. Genet. 2009;41:882–884. doi: 10.1038/ng.403. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Ghoussaini M., Song H., Koessler T., Al Olama A.A., Kote-Jarai Z., Driver K.E., Pooley K.A., Ramus S.J., Kjaer S.K., Hogdall E., et al. Multiple loci with different cancer specificities within the 8q24 gene desert. J. Natl. Cancer Inst. 2008;100:962–966. doi: 10.1093/jnci/djn190. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Jia L., Landan G., Pomerantz M., Jaschek R., Herman P., Reich D., Yan C., Khalid O., Kantoff P., Oh W., et al. Functional enhancers at the gene-poor 8q24 cancer-linked locus. PLoS Genet. 2009;5 doi: 10.1371/journal.pgen.1000597. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Dang C.V. MYC on the path to cancer. Cell. 2012;149:22–35. doi: 10.1016/j.cell.2012.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Calvisi D.F., Conner E.A., Ladu S., Lemmer E.R., Factor V.M., Thorgeirsson S.S. Activation of the canonical Wnt/beta-catenin pathway confers growth advantages in c-Myc/E2F1 transgenic mouse model of liver cancer. J. Hepatol. 2005;42:842–849. doi: 10.1016/j.jhep.2005.01.029. [DOI] [PubMed] [Google Scholar]
- 55.Kim M., Lee H.C., Tsedensodnom O., Hartley R., Lim Y.-S., Yu E., Merle P., Wands J.R. Functional interaction between Wnt3 and Frizzled-7 leads to activation of the Wnt/beta-catenin signaling pathway in hepatocellular carcinoma cells. J. Hepatol. 2008;48:780–791. doi: 10.1016/j.jhep.2007.12.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Bisso A., Filipuzzi M., Gamarra Figueroa G.P., Brumana G., Biagioni F., Doni M., Ceccotti G., Tanaskovic N., Morelli M.J., Pendino V., et al. Cooperation Between MYC and β-Catenin in Liver Tumorigenesis Requires Yap/Taz. Hepatology. 2020;72:1430–1443. doi: 10.1002/hep.31120. [DOI] [PubMed] [Google Scholar]
- 57.Tuupanen S., Turunen M., Lehtonen R., Hallikas O., Vanharanta S., Kivioja T., Björklund M., Wei G., Yan J., Niittymäki I., et al. The common colorectal cancer predisposition SNP rs6983267 at chromosome 8q24 confers potential to enhanced Wnt signaling. Nat. Genet. 2009;41:885–890. doi: 10.1038/ng.406. [DOI] [PubMed] [Google Scholar]
- 58.Durkin J.T., Holskin B.P., Kopec K.K., Reed M.S., Spais C.M., Steffy B.M., Gessner G., Angeles T.S., Pohl J., Ator M.A., Meyer S.L. Phosphoregulation of mixed-lineage kinase 1 activity by multiple phosphorylation in the activation loop. Biochemistry. 2004;43:16348–16355. doi: 10.1021/bi049866y. [DOI] [PubMed] [Google Scholar]
- 59.Xu Z., Maroney A.C., Dobrzanski P., Kukekov N.V., Greene L.A. The MLK family mediates c-Jun N-terminal kinase activation in neuronal apoptosis. Mol. Cell Biol. 2001;21:4713–4724. doi: 10.1128/MCB.21.14.4713-4724.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Fawdar S., Trotter E.W., Li Y., Stephenson N.L., Hanke F., Marusiak A.A., Edwards Z.C., Ientile S., Waszkowycz B., Miller C.J., Brognard J. Targeted genetic dependency screen facilitates identification of actionable mutations in FGFR4, MAP3K9, and PAK5 in lung cancer. Proc. Natl. Acad. Sci. USA. 2013;110:12426–12431. doi: 10.1073/pnas.1305207110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Stark M.S., Woods S.L., Gartside M.G., Bonazzi V.F., Dutton-Regester K., Aoude L.G., Chow D., Sereduk C., Niemi N.M., Tang N., et al. Frequent somatic mutations in MAP3K5 and MAP3K9 in metastatic melanoma identified by exome sequencing. Nat. Genet. 2011;44:165–169. doi: 10.1038/ng.1041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Jun L., Xuhong L., Hui L. Circ_SIPA1L1 promotes osteosarcoma progression via miR-379-5p/MAP3K9 axis. Cancer Biother. Radiopharm. 2023;38:604–618. doi: 10.1089/cbr.2020.3891. [DOI] [PubMed] [Google Scholar]
- 63.Xia J., Cao T., Ma C., Shi Y., Sun Y., Wang Z.P., Ma J. MiR-7 suppresses tumor progression by directly targeting MAP3K9 in pancreatic cancer. Mol. Ther. Nucleic Acids. 2018;13:121–132. doi: 10.1016/j.omtn.2018.08.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Wu T., Du M., Zhang T., Chen X., Liu Z. Comparisons of global incidence and risk factor profiles of hepatocellular carcinoma and intrahepatic cholangiocarcinoma. Ann. Hepatol. 2026;31 doi: 10.1016/j.aohep.2025.102139. [DOI] [PubMed] [Google Scholar]
- 65.Kozlitina J., Smagris E., Stender S., Nordestgaard B.G., Zhou H.H., Tybjærg-hansen A., Vogt T.F., Hobbs H.H., Cohen J.C. Exone-wide association study identifies TM6SF2 variant that confers susceptibility to nonalcoholic fatty liver disease. Nat. Genet. 2014;46:352–356. doi: 10.1038/ng.2901. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Stender S., Kozlitina J., Nordestgaard B.G., Tybjærg-Hansen A., Hobbs H.H., Cohen J.C. Adiposity amplifies the genetic risk of fatty liver disease conferred by multiple loci. Nat. Genet. 2017;49:842–847. doi: 10.1038/ng.3855. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Feng G., Targher G., Byrne C.D., Yilmaz Y., Wai-Sun Wong V., Adithya Lesmana C.R., Adams L.A., Boursier J., Papatheodoridis G., El-Kassas M., et al. Global burden of metabolic dysfunction-associated steatotic liver disease, 2010 to 2021. JHEP Rep. 2025;7 doi: 10.1016/j.jhepr.2024.101271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Tan D.J.H., Tamaki N., Kim B.K., Wijarnpreecha K., Aboona M.B., Faulkner C., Kench C., Salimi S., Sabih A.H., Lim W.H., et al. Prevalence of low FIB-4 in MASLD-related hepatocellular carcinoma: A multicentre study. Aliment. Pharmacol. Ther. 2025;61:278–285. doi: 10.1111/apt.18346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Tan E.Y., Danpanichkul P., Yong J.N., Yu Z., Tan D.J.H., Lim W.H., Koh B., Lim R.Y.Z., Tham E.K.J., Mitra K., et al. Liver cancer in 2021: Global burden of disease study. J. Hepatol. 2025;82:851–860. doi: 10.1016/j.jhep.2024.10.031. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.





