Skip to main content
PLOS Genetics logoLink to PLOS Genetics
. 2026 Sep 18;22(9):e1012308. doi: 10.1371/journal.pgen.1012308

Identifying shared polygenic risk across cancers

Jiaqi Hu 1, Maiyier Muheyati 2, Leqi Xu 2, Andrew DeWan 1, Hongyu Zhao 2,3,*
Editor: Heather J Cordell4
PMCID: PMC13600620  PMID: 42758802

Abstract

Background

Shared genetic susceptibility across cancers has been reported but is generally modest at the genome-wide level. Whether such shared polygenic risk exhibits structured convergence at regional or functional levels remains unclear. We investigated shared genetic risk across cancers by integrating local genetic correlation analyses with cross-cancer polygenic risk score (PRS) associations.

Methods

We estimated pairwise local genetic correlations across 16 specific cancers and one pan-cancer phenotype using SUPERGNOVA. Genome regions harboring multiple cancers with mutually correlated local genetic effects were annotated using genetic correlations with non-cancer phenotypes and associations from the GWAS Catalog. In parallel, cross-cancer PRS associations were evaluated, and significant cancer pairs were identified. Genome-wide PRSs for selected pairs were further decomposed into pleiotropy-informed and pathway-specific components to assess functional enrichment of shared polygenic risk.

Results

Genome-wide genetic correlation analyses identified 20 significantly correlated cancer pairs, whereas local analyses revealed 82 regions with shared genetic signals across 66 cancer pairs. Five regions exhibited mutually correlated cancer clusters, with enrichment in functional domains such as inflammatory functions. Cross-cancer PRS analyses identified five cancer pairs with shared polygenic risk. Decomposition of PRSs indicated that these cross-cancer associations were enriched in specific pleiotropy groups and immune-related pathways rather than reflecting diffuse genome-wide overlap.

Conclusion

Our findings demonstrate that although shared genetic susceptibility across cancers is limited at the genome-wide level, it becomes evident when examined at regional and polygenic scales. Integrating local genetic correlation and PRS decomposition analyses reveals structured patterns of shared genetic risk, providing a framework for investigating cross-cancer polygenic susceptibility.

Author summary

Many cancers are known to share inherited genetic risk, but it remains unclear where this genetic overlap occurs and what biological mechanisms underlie it. Identifying these shared genetic influences may reveal common processes involved in cancer development and could ultimately improve cancer risk prediction and prevention. In this study, we analyzed genetic data from 24 specific cancer types and a pan-cancer phenotype to investigate shared inherited cancer risk. In addition to assessing genetic overlap across the genome, we examined specific genomic regions and groups of genes involved in related biological functions. We found that most cancers showed only limited shared genetic risk at the genome-wide level. However, many pairs of cancers shared genetic influences within particular genomic regions and biological pathways. These findings suggest that different cancers may arise through overlapping genetic mechanisms, even when their overall genetic similarity is low. Overall, this study provides a framework for identifying shared genetic susceptibility across cancer types and offers new insight into biological processes that contribute to cancer development, with potential implications for future cancer prevention and risk prediction.

Introduction

Genetic factors contribute substantially to cancer susceptibility. Evidence from a large Nordic twin cohort demonstrates that approximately one-third of the variance in overall cancer risk is attributable to heritable factors (heritability [h2] = 33%, 95% CI: 30%-37%) [1]. Genome-wide association studies (GWASs) have identified numerous common variants associated with risk across diverse cancer types [2–4]. Notably, some variants exhibit pleiotropy, i.e., they are independently associated with two or more cancers. For example, rs78378222 in TP53 has been implicated in the risk of breast cancer, skin cancer, and several other malignancies [2,5]. These observations suggest shared genetic architecture across cancers and highlight the importance of systematically characterizing cross-cancer genetic overlap.

Efforts to quantify shared genetic susceptibility across cancers—typically measured through genetic correlations— have consistently shown that genome-wide overlap is generally modest. Sampson et al. (2015) evaluated pairwise genetic correlations among 13 cancers and identified significant correlations for only a few pairs, such as bladder and lung cancers, indicating limited shared heritable risk at the genome-wide level [6]. Subsequent studies leveraging larger GWAS datasets have reported similarly low genetic correlations [7,8]. Nevertheless, some cancer pairs exhibit local signals despite weak or nonsignificant genome-wide correlations. For instance, shared risk at the chromosome 9p21 region has been observed across several cancers [7], suggesting that genetic overlap may be concentrated within specific genomic regions rather than distributed broadly across the genome. Furthermore, polygenic risk scores (PRSs), which aggregate effects of risk variants, have revealed cross-cancer associations that are not captured by the global genetic correlation estimate [9], pointing to additional cross-cancer relationships.

Although prior studies have quantified the generally modest genetic overlap across cancers and highlighted that shared signals may be confined to specific genomic regions, the underlying biological mechanisms of this overlap remain insufficiently understood. In particular, it is unclear whether the variants or loci shared across cancers converge on common oncogenic pathways or reflect distinct, context-dependent mechanisms. Moreover, most existing analyses have examined cancer pairs in isolation, limiting the ability to characterize broader patterns of shared genetic architecture. As a result, a comprehensive, network-level understanding of how multiple cancers are genetically interconnected has yet to be fully established.

In this study, we aimed to characterize shared genetic susceptibility across cancers by integrating regional correlation analyses with polygenic risk evaluation. We decomposed the genome-wide sharing into local signals through local genetic correlations and PRS decompositions. Our results confirmed that the shared genetic architecture across cancers on the genome-wide scale is modest, but the overlapped genetic risk is enriched in certain functions, such as immune-metabolic pathways.

Methods

Overview of methods

Using GWAS summary statistics for 24 specific cancers and one pan-cancer phenotype, we first characterized shared genetic susceptibility across cancers by estimating genetic correlations at both genome-wide and regional scales. To provide functional context for regional genetic sharing, we further annotated correlated regions by evaluating their genetic correlations with non-cancer phenotypes and by linking variants within these regions to previously reported associations in the GWAS Catalog.

We next examined pairwise polygenic overlap using PRSs in UK Biobank (UKB) participants. Genome-wide PRSs were used to identify cancer pairs exhibiting significant cross-cancer associations. To further investigate the potential sources of shared polygenic risk, PRSs were decomposed into components with different functions, enabling attribution of cross-cancer associations to pleiotropy groups and biological pathways.

GWAS summary statistics

We downloaded GWAS summary statistics for 24 specific cancers and one pan-cancer phenotype, restricting analyses to studies of European ancestry and excluding datasets that included UKB participants [2–4,10,11]. Sample sizes and study details for all cancer phenotypes are provided in S1 Table. Briefly, the breast cancer GWAS included 133,384 cases and 113,789 controls [2], the prostate cancer GWAS included 79,194 cases and 61,112 controls [3], the lung cancer GWAS comprised 44,069 cases and 68,712 controls [4], and the ovarian cancer GWAS consisted of 25,509 cases and 40,941 controls [10]. The remaining 21 cancer GWASs were derived from the FinnGen R12 release [11], which used a shared set of cancer-free controls (except for sex-specific cancers including cervix uteri and testis cancers) and included between 152 cases for nasopharyngeal cancer and 121,495 cases for the pan-cancer phenotype. The pan-cancer was defined as diagnosis of any malignant neoplasm. FinnGen is a large-scale national research initiative integrating genotype data from over 500,000 Finnish biobank participants with comprehensive longitudinal health registry data to study disease susceptibility.

All summary statistics underwent quality control procedures. We excluded non-autosomal variants, insertion-deletion polymorphisms, and variants not present in the UKB genotype. For multi-allelic variants, a single allele was retained based on the smallest association p-value. The cleaned summary statistics were then harmonized and processed using LD score regression (LDSC) [12] to estimate SNP-based heritability, restricting analyses to variants with minor allele frequency (MAF) ≥ 0.001. Cancers with heritability not significantly larger than 0 were excluded from downstream analyses.

UK biobank data

Individual-level data from the UKB were used for PRS analyses. UKB is a large prospective cohort study that has enrolled over 500,000 participants aged 40–69 years across the United Kingdom [13]. Cancer case status was ascertained using a combination of self-reported diagnoses and linked hospital records, incorporating International Classification of Diseases, Ninth and Tenth Revisions (ICD-9 and ICD-10), as well as Office for National Statistics Classification of Interventions and Procedures (OPCS) codes. The specific codes used to define cancer outcomes are provided in S2 Table. Individuals with any cancer diagnosis were classified as cases, and controls were defined as participants without a recorded diagnosis of any cancers.

Analyses were restricted to unrelated participants of genetically inferred European ancestry [14]. We used phase III UKB genotype data, in which participants were genotyped using either the UK BiLEVE Axiom Array or the UK Biobank Axiom Array, covering approximately 820,000 variants. Genotypes were centrally imputed using the 1000 Genomes Project and Haplotype Reference Consortium (HRC) reference panels, resulting in approximately 93 million variants per individual. Standard quality control procedures were applied, retaining autosomal variants with an imputation quality score > 0.3, a Hardy-Weinberg equilibrium p-value > 1 × 10-5, and an MAF > 0.05.

Genetic correlations

Genome-wide genetic correlations across cancers were estimated using cleaned GWAS summary statistics via GNOVA [15], which accounts for sample overlaps between GWASs and performs pairwise correlation analyses. Bonferroni correction was applied to identify statistically significant cancer pairs.

Local genetic correlations were further evaluated using SUPERGNOVA [16], a method that provides stable estimates of regional genetic correlations [17]. The genome was partitioned into more than 2,000 approximately independent linkage disequilibrium (LD) blocks, and block-specific genetic correlations were computed for each cancer pair. Statistical significance was assessed using Bonferroni correction within each cancer pair. To identify shared correlated regions, we hypothesized that regions exhibiting local genetic correlations may contribute to shared risk across multiple cancers in a coordinated manner. Accordingly, we identified genomic regions having correlated cancer clusters involving at least three distinct cancer types. These regions were retained for downstream analyses.

Identified regions were annotated to characterize broader biological relevance in two complementary ways. First, we evaluated correlations with non-cancer phenotypes. GWAS summary statistics for 3,950 non-cancer traits, restricted to individuals of genetically inferred European ancestry, were obtained from the Neale Lab UKB analyses [18]. For each selected region, phenotypes with enriched local heritability (p-value < 0.05) were identified using LAVA [19]. Local genetic correlations between cancers in the region and the selected non-cancer phenotypes were then estimated using SUPERGNOVA, with significance assessed using Bonferroni correction. Second, variants within the regions of interest were extracted and annotated using the SNP2GENE function in FUMA [20] with the default positional mapping settings. Mapped genes were further examined using the GENE2FUNC module to identify previously reported associations in the GWAS Catalog. Through this approach, regional genetic correlations were linked to known variant-phenotype associations.

PRS development and decomposition

To evaluate cross-cancer polygenic associations, we constructed genome-wide PRSs for each specific cancer. PRS weights were inferred from cleaned GWAS summary statistics using Summary statistics-based Dirichlet Process Regression (SDPR) [21], a Bayesian shrinkage method that does not require parameter tuning and has demonstrated stable predictive performance across multiple cancers [22]. SDPR employs a Dirichlet process prior on SNP effect-size variances, resulting in continuous shrinkage of effect estimates toward zero without explicitly assuming a fixed proportion of causal variants. Individual-level PRSs were calculated in the UKB participants as the weighted sum of risk alleles using PLINK 1.9 [23]. All PRSs were standardized to have a mean of zero and a standard deviation of one. Predictive performance was evaluated using the area under the receiver operating characteristic curve (AUC), and the statistical significance was assessed accordingly. Associations between PRSs and cancer outcomes were summarized using odds ratios (ORs).

To identify cancer pairs sharing polygenic risk, we conducted pairwise association analyses. For each cancer pair (i, j), cancer i was modeled as the outcome, and its association with the PRS for cancer j was evaluated using logistic regression, adjusting for the PRS for cancer i, age at recruitment, sex (except for sex-specific cancers), and the top 10 genetic principal components (PCs). The inclusion of PRS for cancer i was intended to evaluate whether the PRS for cancer j provided information on the risk of cancer i beyond that explained by the PRS for cancer i. This procedure was repeated across all cancer pairs, and statistically significant associations were identified. Because cancer comorbidity may be influenced by factors beyond shared common polygenic susceptibility, including rare variants, somatic alterations, and environmental exposures, we conducted sensitivity analyses excluding individuals diagnosed with both cancers to assess the robustness of the observed associations. Specifically, analyses were repeated among participants diagnosed with cancer i but not cancer j. Cancer pairs that remained significant after exclusion of comorbid cases were considered to share polygenic risk independent of clinical co-occurrence.

To investigate potential mechanisms underlying shared polygenic risk, we decomposed PRSs into functionally informed component PRSs using two complementary strategies. These approaches aimed to attribute across-cancer PRS associations to broader pleiotropic or biological processes rather than individual variants.

First, we applied our recently developed pleiotropy-decomposed PRS (PD-PRS) framework, which decomposed the genome-wide PRS into pleiotropy-informed components based on local genetic correlations [24]. Briefly, the genome was partitioned into approximately 2,000 independent LD blocks. Block-specific genetic correlations between cancers and 47 non-cancer phenotypes—grouped into 12 pleiotropy groups [25]—were estimated using SUPERGNOVA [16]. These groups were defined using domain knowledge of phenotype relationships, with their biological relevance further supported by significant genetic correlations among grouped phenotypes. Variants located within blocks showing significant correlations (p-value < 0.05) were assigned to the corresponding pleiotropy groups. For each pleiotropy group, a PD-PRS was constructed using variants assigned to that group and SDPR-derived genome-wide PRS weights, implemented with PLINK v1.9 [23]. An additional PD-PRS was defined using variants not assigned to any pleiotropy groups. In total, 13 PD-PRSs were generated for each cancer.

For each selected cancer pair, logistic regression models were fitted with cancer i (excluding individuals with cancer j) as the outcome and each PD-PRS for cancer j as the exposure, adjusting for the PRS of cancer i, age at recruitment, sex (except for sex-specific cancers), and the top 10 genetic PCs. To quantify the contribution of each pleiotropy group to the across-cancer association, we evaluated attenuation in explained variance using Nagelkerke’s R2. Specifically, a full model including the PRS for cancer j (Equation (1)) was compared with reduced models substituting either the PD-PRS (Equation (2)) or the remaining PD-PRS constructed using variants not included in the PD-PRS (Equation (3)).

Canceri (Excluding patients with cancerj)~PRScancerj+PRScanceri+covariates (1)
Canceri (Excluding patients with cancerj)~PDPRScancerj+PRScanceri+covariates (2)
Canceri (Excluding patients with cancerj)~RemainingPDPRScancerj+PRScanceri+covariates (3)

Nagelkerke's R2 was calculated for each model. The ΔR2 for excluding PD-PRS was calculated as the difference between R2 of equation (1) and of equation (3), and the ΔR2 for excluding remaining PD-PRS was the difference between R2 of equation (1) and of equation (2). The attenuation in R2 was computed as the difference between ΔR2 the values for PD-PRS and for the remaining PD-PRS. To test the null hypothesis that the across-cancer association cannot be attributed to the PD-PRS, a permutation test was conducted with 1,000 permutations while preserving the associations between the outcome, the PRS for cancer i, and covariates. Specifically, we first fitted the null model including the PRS for cancer i and all covariates and estimated the corresponding outcome probabilities. Permuted outcomes were then generated by resampling from these fitted probabilities while preserving the overall number of cases. The association between the PD-PRS and the permuted outcomes was subsequently evaluated to obtain the empirical null distribution.

Second, we decomposed PRSs into pathway-specific PRSs (PS-PRSs), which assigned SNPs to genes and then labelled with pre-defined pathways, to evaluate shared biological mechanisms between cancers. A total of 8,004 autosomal genes from 364 KEGG pathways [26] were mapped to genetic variants using positional mapping with a ± 10-kilobase window. For each pathway, a PS-PRS was calculated using pathway-specific variant sets and SDPR-derived weights, implemented in PLINK 1.9 [23]. Across-cancer associations for PS-PRSs were evaluated using logistic regression models analogous to those described above.

Statistical analyses

All statistical analyses were conducted using R version 4.2, except where noted.

Results

An overview of the study design and analytic workflow is shown in Fig 1.

Fig 1. Overview of the study design.

Fig 1

This figure summarizes the analytic framework used to characterize shared polygenic risk across cancers. Panel A shows the genetic correlation analyses based on GWAS summary statistics. Genome-wide genetic correlations were estimated pairwise using GNOVA, and local genetic correlations were calculated using SUPERGNOVA to identify genomic regions harboring correlated cancer clusters. Regions with shared signals were further annotated using genetic correlations with non-cancer traits and variant mappings from the GWAS Catalog. Panel B illustrates the PRS analyses. Genome-wide cancer PRSs were constructed using SDPR and evaluated in cross-cancer association analyses to identify cancer pairs with shared polygenic risk. For selected pairs, PRSs were further decomposed into pleiotropy-informed and pathway-specific components to assess functional enrichment of shared genetic risk.

GWAS summary statistics

We downloaded and cleaned GWAS summary statistics for 24 specific cancers and one pan-cancer phenotype. SNP-based heritability was estimated using LDSC [12]. Seventeen cancers showed significantly enriched heritability except for testis cancer, stomach cancer, oral cavity cancer, cervix uteri cancer, acute myeloid leukemia, nasopharyngeal cancer, chronic myeloid leukemia, and oropharyngeal cancer. Among the cancers with significant heritability estimates, prostate and breast cancers showed the highest estimates, h2 = 0.13 ± 0.02; and h2 = 0.13 ± 0.01, respectively, whereas acute lymphocytic leukemia (ALL) had the lowest (h2 = 0.002 ± 0.001) (S1 Table).

Genetic correlations

Using GNOVA, we estimated pairwise genome-wide genetic correlations across 15 specific cancers and one pan-cancer phenotype. One more cancer, ALL, was excluded from this analysis due to negative heritability estimate in GNOVA. After Bonferroni correction for multiple testing, 20 cancer pairs showed significant correlations (p-value ≤ 4.17 × 10-4; Fig 2). Of these, 10 involved the pan-cancer phenotype, suggesting that some genetic risk factors contributing to individual cancers also influence aggregated cancer susceptibility. An additional 6 correlations involved non-melanoma skin cancer (NMSC). The remaining four correlated pairs were melanoma skin cancer (MSC)-non-Hodgkin lymphoma (NHL), breast-colorectal cancer, breast-lung cancer, and breast-ovarian cancers.

Fig 2. Genome-wide genetic correlations across cancers.

Fig 2

Heatmap showing pairwise genome-wide genetic correlations among 15 specific cancers and one pan-cancer phenotype estimated using GNOVA. Each cell represents the genetic correlation coefficient between a pair of cancers, with red indicating positive correlations and blue indicating negative correlations. Asterisks denote correlations that remain significant after Bonferroni correction (p-value ≤ 4.17 × 10-4. Consistent with prior studies, significant genome-wide correlations were observed for a limited number of cancer pairs, with the pan-cancer phenotype and NMSC showing correlations with multiple specific cancers. Overall, genome-wide genetic sharing across cancers was modest.

Furthermore, we evaluated local genetic correlations across 16 specific cancers. Following pair-specific Bonferroni correction, 82 unique genomic regions with significant local genetic correlations were identified across 66 cancer pairs (S3 Table). Among these pairs, nine also showed significant genome-wide genetic correlations. In contrast, the MSC-NHL pair, which was significantly genetically correlated, did not show any significant local genetic correlations. This discrepancy may reflect limited power to detect local genetic correlations for these cancers, a polygenic shared genetic architecture distributed across many genomic regions, or a combination of both factors. Of note, 57 cancer pairs showed significant local genetic correlations in the absence of genome-wide significance, suggesting that global correlation estimates may fail to capture regional genetic sharing.

Across the 82 identified regions, seven regions showed mutually correlated cancer clusters involving at least three cancer types (Table 1). One region on chromosome 2 (chr2:201572564–202829668; 2q33) was excluded because Chronic Lymphocytic Leukemia and Small Lymphocitis Leukemia (CLL) was a clinical subtype of Non-Hodgkin Lymphoma (NHL), and one region on chromosome 6 (chr6:32424108–32682443; 6p21) was excluded due to overlap with the major histocompatibility complex (MHC) since its extensive and complex LD structure can lead to unstable estimates and complicate the interpretation of local genetic correlation analyses. The remaining five regions were carried forward for functional annotation. Specifically, we examined local genetic correlations between cancers and non-cancer phenotypes and mapped variants within each region to known associations in the GWAS Catalog. Four of the five regions (excluding chr1:182294372–183796074; 1q25) showed at least one non-cancer phenotype that was significantly correlated with all cancers in the corresponding cluster (S4-S7 Tables). Variants within these regions were further mapped to gene sets using FUMA (S8 Table).

Table 1. Selected regions with mutually correlated cancer clusters.

Region (chr:start:end) Cytoband Mutually correlated pair(s)
1:182294372-183796074 1q25 Colorectal-NMSC-Prostate
2:201572564-202829668 2q33 CLL-NHL-NMSC
6:32424108-32682443 6p21 CLL-NHL-NMSC
Bladder-NHL-NMSC
6:167178790-168548525 6q27 Lung-NHL-NMSC
8:128166556-128542444 8q24 Breast-Colorectal-Prostate
9:18660695-19129349 9p22 Lung-Prostate-ThyGland
11:63154309-66835194 11q12 Colorectal-NHL-NMSC
Breast-Colorectal-NMSC

Chr: chromosome; NMSC: non-melanoma skin cancer; CLL: chronic lymphocytic leukemia and small lymphocitis leukemia; NHL: non-Hodgkin lymphoma; CML: chronic myeloid leukemia; ThyGland: thyroid gland cancer.

For the region chr6:167178790–168548525 (6q27), a mutually correlated cluster involving lung cancer, NHL, and NMSC was identified. All three cancers were significantly correlated with total protein levels in this region (S4 Table). Although total protein itself is not an established cancer risk factor, it reflects the combined abundance of circulating proteins and may capture underlying immune and inflammatory processes. This interpretation is supported by previous prospective proteomic studies demonstrating associations between numerous plasma proteins and the future risk of lung cancer and NHL [27]. In addition, variants within this region were enriched for 28 GWAS catalog traits, including thyroid function-related phenotypes such as thyroid-stimulating hormone levels and hyperthyroidism (S8 Table). This region has previously been associated with basal cell carcinoma [28], and may represent a locus with broader relevance to cancer risk.

In the chr8:128166556–128542444 (8q24) region, three cancers—breast cancer, colorectal cancer, and prostate cancer—showed shared correlations with six non-cancer phenotypes, including sibling history of prostate cancer (S5 Table). Variants were enriched for four cancers including urinary bladder carcinoma, breast carcinoma, CLL, and prostate carcinoma (S8 Table). The clustering of breast, colorectal, and prostate cancers at 8q24 is consistent with previous reports [8], further supporting the role of this region as a shared cancer susceptibility locus.

For chr9:18660695–19129349 (9p22), three cancers (lung cancer, prostate cancer, and thyroid gland cancer) were significantly correlated with 18 non-cancer phenotypes, including measures of lung function and blood cell counts (S6 Table). Variants in this region were mapped to 15 GWAS traits, including small-cell lung carcinoma (S8 Table). This region is an established risk region for ovarian cancer [29], and our results further extended it beyond a single cancer.

In the chr11:63154309–66835194 (11q12) region, four cancers (colorectal cancer, breast cancer, NHL, and NMSC) were correlated with 30 non-cancer phenotypes, notably obesity-related traits and blood cell indices (S7 Table), and variants were similarly enriched for obesity-related GWAS associations (S8 Table). This region has previously been associated with breast, prostate, and ovarian cancer risk [8,30]. Our findings further supported its role as a shared cancer susceptibility locus.

In contrast, for chr1:182294372–183796074 (1q25), no non-cancer phenotypes were significantly correlated across all cancers in the cluster. Variants in this region were mapped to seven GWAS traits, including chronotype-related phenotypes (S8 Table). Similar associations have been established with colorectal and prostate cancers previously [31,32], whereas our findings suggest a broader role for this locus in shared cancer susceptibility.

PRS analyses

PRS weights were inferred for 16 specific cancers using SDPR, and individual-level PRSs were calculated for UKB participants of European ancestry. Discriminative performance was evaluated using AUC, and relative risk was summarized using ORs. Results are shown in S9 Table. The number of cases ranged from 124 for ALL to 25,666 for NMSC. Fifteen PRSs except ALL PRS showed AUCs and ORs significantly different from the null hypothesis (p-value < 0.05) and were retained for subsequent analysis.

Pairwise correlations among the 15 PRSs are shown in S1A Fig. Most PRSs were significantly correlated (p-value < 4.76 × 10-4), although the magnitudes of correlation were modest. Statistical significance was likely influenced by the large number of shared controls, whereas the small effect sizes reflected limited genome-wide genetic overlap across cancers. We further examined overlap among individuals in the top 10% of each cancer (S1B Fig). Consistent with the modest correlation coefficients, the proportion of shared high-risk individuals was approximately 10% for most cancer pairs. An exception was observed for NMSC and MSC, for which nearly 20% overlap was detected. Overall, these results indicate that genome-wide polygenic overlap across cancers is present but generally limited in magnitude.

We next assessed associations between 15 cancer-specific PRSs and specific cancer outcomes in the UKB using a two-step approach. In the first step, we tested cross-cancer associations by modeling each cancer outcome as a function of the PRS for another cancer, while adjusting for the target cancer's PRS. Among the 345 cancer-PRS pairs examined, 18 showed significant associations after Bonferroni correction (p-value < 1.45 × 10-4; Table 2). In the second step, we assessed whether these associations could be explained by clinical comorbidity by excluding individuals diagnosed with both cancers. After Bonferroni correction (p-value < 2.78 × 10-3), seven cancer-PRS pairs remained statistically significant (Table 2). These included associations between bladder cancer risk and the lung cancer PRS, as well as associations between lung cancer risk and the bladder cancer PRS. To further evaluate whether the identified shared polygenic risk can be fully attributed to smoking, we conducted stratified analyses for the bladder–lung cancer pair by smoking status. The association between bladder cancer and the lung cancer PRS remained significant among both ever and never smokers. In contrast, associations between lung cancer and the bladder cancer PRS were no longer significant in either stratum, but the effect estimates were similar to those observed in the pooled analysis (S10 Table), indicating potential loss of power in stratified analyses. Our results suggest that smoking-related mechanisms alone are unlikely to account for the observed genetic overlap between bladder and lung cancers. Based on these results, subsequent analyses focused on six cancer-PRS pairs: bladder cancer-lung cancer PRS, lung cancer-bladder cancer PRS, breast cancer-NMSC PRS, MSC-NMSC PRS, NMSC-CLL PRS, and NMSC-lung cancer PRS. Because CLL is a clinical subtype of non-Hodgkin lymphoma, the NMSC-CLL pair was excluded from further analyses.

Table 2. Significant cross-cancer associations for genome-wide PRSs.

Step 1 Step 2
Cancer PRS OR (95% CI) P-value OR (95% CI) P-value
Bladder Lung 1.09 (1.06-1.12) 3.03E-08 1.08 (1.05-1.11) 6.93E-07
Bladder Prostate 1.07 (1.04-1.1) 1.78E-05 0.98 (0.95-1.01) 1.77E-01
Breast NMSC 1.06 (1.04-1.07) 5.99E-14 1.03 (1.01-1.04) 2.36E-04
CLL NHL 1.16 (1.09-1.22) 3.39E-07 1.13 (1.07-1.2) 4.31E-05
Colorectal Breast 1.04 (1.02-1.06) 1.44E-04 1.01 (0.99-1.03) 2.08E-01
Colorectal Prostate 1.05 (1.03-1.07) 1.14E-05 1.01 (0.99-1.03) 5.33E-01
Kidney Breast 1.08 (1.04-1.12) 1.12E-04 1.06 (1.02-1.1) 4.02E-03
Kidney NMSC 1.09 (1.05-1.13) 1.33E-05 1.05 (1.01-1.09) 1.43E-02
Lung Bladder 1.05 (1.02-1.08) 1.39E-04 1.04 (1.02-1.07) 1.12E-03
Lung NMSC 1.06 (1.03-1.09) 7.37E-06 1.04 (1.01-1.06) 8.69E-03
MSC NMSC 1.17 (1.14-1.21) 1.16E-24 1.09 (1.06-1.13) 6.03E-07
NMSC Breast 1.03 (1.02-1.05) 3.49E-07 1.01 (0.99-1.02) 4.51E-01
NMSC CLL 1.04 (1.03-1.06) 3.49E-10 1.04 (1.02-1.05) 5.85E-08
NMSC Lung 1.05 (1.03-1.06) 6.69E-12 1.04 (1.03-1.06) 4.02E-10
NMSC Prostate 1.03 (1.01-1.04) 7.04E-05 0.98 (0.97-1) 1.34E-02
Ovary Breast 1.12 (1.08-1.17) 5.64E-08 1.07 (1.02-1.12) 3.44E-03
Prostate Colorectal 1.04 (1.02-1.06) 3.30E-05 1.03 (1.01-1.05) 3.00E-03
Prostate NMSC 1.05 (1.03-1.07) 7.49E-08 1.01 (0.99-1.03) 4.18E-01

PRS: polygenic risk score; OR: odds ratio; CI: confidence interval.

Bold: significant after Bonferroni correction in step 2 (p-value < 2.78 × 10 -3 ).

PRSs involved in the six selected cancer pairs—lung cancer PRS, bladder cancer PRS, NMSC PRS, and CLL PRS—were involved for the following decompositions. Specifically, we applied two decomposition frameworks to the four selected genome-wide PRSs to further characterize sources of shared polygenic risk. Specifically, PRSs were decomposed into pleiotropy-informed PD-PRSs and pathway-specific PS-PRSs, and cross-cancer associations were evaluated using these component scores.

PD-PRSs.

The genome-wide PRSs for bladder cancer, lung cancer, NMSC, and CLL were decomposed into 14 PD-PRSs, with the number of SNPs in each component summarized in S2 Fig. For all three cancers, the largest proportion of variants was assigned to the other PD-PRS, which comprised variants located in genomic regions showing no detectable genetic correlation with the 47 non-cancer phenotypes. Associations between PD-PRSs and their corresponding cancers were evaluated using logistic regression models with covariate adjustment (S3 Fig). All PD-PRSs for lung cancer and NMSC were significantly associated with their respective cancers after Bonferroni correction (p-value < 8.93 × 10-4; S3C-S3D Fig). For bladder cancer, associations with alcohol consumption- and hypertension-related PD-PRSs failed to reach significance (S3A Fig). For CLL, two PD-PRSs—alcohol consumption and chronic kidney disease—were not significantly associated and were excluded from the following analyses (S3B Fig). Despite containing the largest number of variants, the other PD-PRS did not consistently yield the strongest associations with cancer risk. Instead, higher ORs were often observed for PD-PRSs corresponding to specific pleiotropy groups, suggesting that cancer-associated genetic risk may be enriched invariants shared with non-cancer phenotypes. For bladder cancer, the neuropsychiatric disease PD-PRS showed the strongest association (OR = 1.29, 95% CI: 1.26-1.33), implicating enrichment of related function in the development of bladder cancer. For CLL, the autoimmune disease PD-PRS showed the strongest association (OR = 1.45, 95% CI: 1.38-1.54), supporting an immune-related component of genetic susceptibility. For lung cancer, the diabetes PD-PRS exhibited the largest effect size (OR = 1.27, 95% CI: 1.23-1.30), consistent with evidence linking metabolic dysfunction, insulin resistance, and diabetes-related pathways to cancer development and progression [33]. For NMSC, the neuropsychiatric disease PD-PRS showed the strongest association (OR = 1.46, 95% CI: 1.44–1.48), suggesting that genetic components enriched for neuropsychiatric disease may also contribute to NMSC susceptibility. Although the underlying mechanisms remain unclear, previous studies have reported shared genetic architecture between neurological disorders such as Parkinson’s disease and skin cancers, potentially reflecting overlap in pigmentation biology, neural crest development, DNA damage response, and immune regulation [34].

Further, we analyzed cross-cancer associations (excluding comorbidity) using PD-PRSs for the six cancer pairs identified in prior analyses. Among 82 tested associations, 30 were statistically significant (p-value < 0.05; Fig 3 and S11 Table). For the bladder-lung cancer pair, the diabetes PD-PRS for lung cancer showed the strongest association with bladder cancer risk (OR = 1.06, 95% CI: 1.03-1.09), consistent with its prominent role in lung cancer susceptibility. While lung cancer-bladder cancer PRS pair, the diabetes PD-PRS for bladder cancer was not associated with lung cancer risk (OR = 1.02, 95% CI: 0.99-1.04), and the lipids PD-PRS showed the strongest association (OR = 1.04, 95% CI: 1.01-1.07). For the breast-NMSC pair, four PD-PRSs showed significant associations with comparable effect sizes. Notably, the neuropsychiatric PD-PRS, which showed the strongest association with NMSC, was not associated with breast cancer but instead showed its strongest association with MSC, suggesting cancer-specific patterns of pleiotropy. For the NMSC-CLL pair, two CLL PD-PRSs—others and smoking PD-PRSs—were significantly associated with NMSC, although the association with the smoking PD-PRS was inverse (OR = 0.99, 95% CI: 0.97-1). In addition, seven lung cancer PD-PRSs, including diabetes PD-PRS, were significantly associated with NMSC risk. Together, these findings suggested that cross-cancer PRS associations are driven by specific pleiotropic components rather than by uniform genome-wide effects.

Fig 3. Cross-cancer associations of PD-PRS.

Fig 3

ORs and 95% confidence intervals for associations between PD-PRSs and cancer outcomes after excluding individuals with comorbid cancers are shown for six cancer pairs of interest. Each panel corresponds to a target cancer outcome, and PD-PRSs were derived from the genome-wide PRS of the paired cancer. Asterisks indicate nominally significant associations (p-value < 0.05). Across pairs, only a subset of PD-PRS components showed significant associations, indicating that cross-cancer polygenic overlap is concentrated in specific pleiotropy-defined trait domains rather than uniformly distributed across the genome.

To formally assess whether cross-cancer associations could be attributed to specific PD-PRSs, we evaluated attenuation in explained variance using differences in Nagelkerke’s R2. Among the 30 significant associations, 26 showed significantly different ΔR2 values between the PD-PRS and the remaining PRS (permutation p-value < 0.05; Fig 4). In 23 of these associations, the PD-PRS explained a larger proportion of the variance than the remaining PRS, suggesting enrichment of shared genetic risk within specific pleiotropy components. For bladder cancer risk, lung cancer PD-PRSs related to atherosclerotic disease, chronic kidney disease, diabetes, and neuropsychiatric disease showed significantly greater ΔR2, while for lung cancer, all five associated bladder cancer PD-PRSs related to atherosclerotic disease, autoimmune disease, chronic kidney disease, lipids, and neuropsychiatric disease showed significantly greater ΔR2. All four PD-PRSs for NMSC associated with breast cancer, as well as all seven PD-PRSs for NMSC associated with MSC, demonstrated larger ΔR2 values. In contrast, for the NMSC-CLL pair, only the other PD-PRS showed greater variance explained, suggesting that additional or uncharacterized mechanisms may contribute to this association. Overall, decomposing genome-wide cancer PRSs into pleiotropy-informed components enabled quantification of relative contributions of specific phenotype groups to cross-cancer associations.

Fig 4. Attenuation of explained variance (ΔR2) for PD-PRSs.

Fig 4

For each selected cancer pair, the proportion of variance explained ΔR2 by the PD-PRS (red) was compared with that explained by the corresponding remaining PRS (blue). Each panel represents one selected cancer-PRS pair. Asterisks indicate PD-PRS components for which the difference in ΔR2 between the PD-PRS and remaining PRS was statistically significant based on the permutation testing (1,000 permutations). Although the overall variance explained by cross-cancer PRSs was modest, selected PD-PRS components accounted for a larger share of the explained variance, indicating enrichment of shared polygenic risk within specific functions.

PS-PRSs.

To examine biologically informed polygenic sharing, variants included in the PRSs for bladder cancer, CLL, lung cancer, and NMSC were mapped to 8,003 autosomal genes and subsequently assigned to 364 KEGG pathways. The number of variants contributing to each PS-PRS is summarized in S12 Table. The pathways containing the largest numbers of variants were metabolic pathways, pathways in cancer, and pathways related to neurodegenerative diseases. PS-PRSs were calculated for UKB participants of European ancestry as weighted sums of pathway-specific variants, using SDPR-derived genome-wide weights.

Associations between PS-PRSs and their corresponding cancers were evaluated using logistic regression. After Bonferroni correction, 258 PS-PRSs were significantly associated with their target cancers (p-value < 3.42 × 10-5), including 21 for bladder cancer, 49 for CLL, 31 for lung cancer, and 174 for NMSC (S13 Table). We next evaluated cross-cancer PS-PRS associations for the six cancer pairs identified in prior analyses, excluding individuals with comorbid diagnoses. Using a nominal significance threshold (p-value < 0.05), 70 significant cross-cancer associations were identified (Fig 5 and S14 Table).

Fig 5. Significant cross-cancer associations of PS-PRS.

Fig 5

ORs and 95% confidence intervals for 70 significant associations (p-value < 0.05) between PS-PRSs and cancer outcomes after excluding individuals with comorbid cancers are shown for five pairs of interest. Each panel corresponds to a target cancer outcome, and PS-PRSs were derived from the genome-wide PRS of the paired cancer. Across all PS-PRSs, only subset of PS-PRSs showed significant cross-cancer associations, suggesting the genetic overlap between cancers might converge to specific functional pathways.

To assess overlap among pathways, we examined gene sharing using UpSet plots and Jaccard similarity indices (S4 Fig). For the bladder-lung cancer pair, six lung cancer PS-PRSs showed significant associations with bladder cancer risk. Gene overlap across these pathways was limited, except for the gastric cancer pathway (S4A–S4B Fig). These pathways are broadly clustered into cancer-related and infection-related categories, indicating that polygenic overlap between bladder and lung cancer is distributed across multiple functional groups rather than driven by a single shared pathway. Similar pathways were found for associations between bladder cancer PS-PRS and lung cancer risk (S4E–S4F Fig), further supporting the roles of these pathways in genetic overlap for these two cancers. For the NMSC-MSC pair, 39 NMSC PS-PRSs were significantly associated with MSC. Although gene overlap across pathways was generally limited, clusters of pathways related to neurodegenerative diseases and metabolic processes were observed, indicating convergence at the functional category level rather than at the gene level (S4G–S4H Fig). Only two CLL PS-PRSs were significantly associated with NMSC. These pathways did not share genes (S4I–S4J Fig), suggesting that shared polygenic risk between CLL and NMSC may arise from distinct biological processes rather than common pathway-level effects. For the NMSC–lung cancer pair, nine lung cancer PS-PRSs were significantly associated with NMSC (S4K–S4L Fig). While gene overlap across pathways remained limited, several functional clusters were evident, including infection-related immune pathways, cancer-related pathways, and other immune-associated processes.

Overall, PS-PRS analyses indicate that cross-cancer polygenic overlap is distributed across multiple biological pathways with limited gene-level overlap. These results suggest that shared genetic susceptibility across cancers reflects convergence within broader functional categories rather than dependence on specific genes or pathways.

Discussion

In this study, we characterized shared genetic architecture across cancers by integrating genetic correlation analyses with PRS-based approaches. Local genetic correlation analyses revealed clusters of multiple cancers sharing regional genetic susceptibility that was not apparent from genome-wide correlations alone. Complementary cross-cancer PRS analyses identified a limited number of cancer pairs with shared polygenic risk. Decomposition of genome-wide PRSs into pleiotropy- and pathway-informed components further indicated that cross-cancer associations are driven by structured enrichment of shared genetic effects across specific non-cancer trait groups and biological pathways, rather than by uniform genome-wide overlap. To conclude, we confirmed limited average genetic sharing at the genome-wide scale across cancers and found that these signals were enriched for specific functions, such as immune-related pathways.

Our observation of a limited number of correlated cancer pairs and cross-cancer PRS associations reinforces prior evidence that polygenic risk sharing across cancers is generally modest [6–9]. Importantly, the cancer pairs identified by our analyses recapitulate well-established relationships among cancers, supporting the validity of our analytic framework. For example, genome-wide genetic correlations between breast cancer and colorectal cancer, as well as between breast cancer and lung cancer, have been reported previously [7]. Similarly, the association between bladder cancer risk and the lung cancer PRS identified in our study is consistent with findings from Sampson et al. (2015), which reported both a significant genetic correlation between these cancers and cross-cancer PRS associations driven by a subset of risk variants [6]. We also observed a unidirectional association in which the PRS for CLL was associated with NMSC, whereas the NMSC PRS was not associated with CLL, a pattern consistent with prior work [35]. Together, these findings confirmed the limited genetic sharing globally across cancers and advocated regional and functional analyses.

By integrating local genetic correlation analyses with functionally informed PRS decomposition, we demonstrate that shared polygenic risk across cancers, while limited in magnitude, converges on specific functions. Such a regional and polygenic organization is not apparent from genome-wide correlation analyses alone. For example, beyond global genetic correlations, we identified a multi-cancer regional cluster on chromosome 8 (chr8:128166556–128542444; 8q24) shared by breast, colorectal, and prostate cancers. This region corresponds to 8q24, a well-established cancer susceptibility locus that can be missed by genome-wide correlation metrics despite its relevance across cancer types [7]. By identifying 8q24 through local correlation clustering and annotating it with non-cancer traits, our analysis places this locus within a broader framework of shared regional genetic susceptibility across cancers. Consistent with prior biological knowledge, this region contains multiple long non-coding RNAs and has been linked to cancer risk through regulation of the proto-oncogene MYC [36]. In addition, cancers within this cluster showed shared genetic correlations with hypertension-related traits, aligning with previous evidence linking MYC to blood pressure regulation [37].

Decomposition of genome-wide PRSs into pleiotropy-informed and pathway-specific components further extends current understanding of cross-cancer PRS associations. Specifically, our PD-PRS analyses indicate that the shared polygenic signal between bladder and lung cancers is enriched for genetic components associated with atherosclerotic diseases and neuropsychiatric phenotypes (Fig 3). Complementary PS-PRS analyses highlight enrichment in immune- and virus-related pathways (Fig 5). Although these findings do not implicate a single causal mechanism, they suggest that the shared genetic susceptibility between bladder and lung cancers may arise from interconnected biological processes involving multiple functions. These findings provide a structured framework for interpreting cross-cancer PRS associations and suggest that shared polygenic risk is concentrated within specific functional domains rather than distributed uniformly across the genome.

We note several limitations of our study. First, our analyses focused on polygenic risk derived from common variants and did not incorporate rare variants or somatic mutations, which play important roles in cancer susceptibility and progression. We also did not evaluate variant-level pleiotropy, which may provide complementary insight into shared cancer risk. Second, some established pleiotropic cancer loci may not have been detected because of limitations of the predefined regional partitioning, as illustrated by the TERT gene being split across two adjacent regions (chr5:10,056-1,267,356 and chr5:1,270,983-1,762,678), or because of limited power and heterogeneity in the input GWAS summary statistics, as observed for the ABO gene. Third, this study was designed to be hypothesis-generating, and independent validation will be necessary to confirm the identified patterns of shared genetic architecture. Fourth, both genetic correlations and cross-cancer PRS associations were modest in magnitude, indicating that shared polygenic risk explains only a limited proportion of cancer variance. Fifth, cancer phenotypes included in this study may be heterogeneous with respect to tumor subtype and clinical characteristics, which could attenuate cancer-specific patterns of shared genetic architecture. In addition, statistical power varied across cancer types because of substantial differences in sample size, potentially limiting the detection of shared genetic signals for less prevalent cancers. Furthermore, the cancer GWASs and UKB analyses included prevalent cancer cases diagnosed before genotyping. Because germline genetic risk is fixed at birth, this design is unlikely to substantially affect the estimated inherited susceptibility. Nevertheless, future studies restricted to incident cancer cases would be valuable to further evaluate the robustness of our findings. Additionally, differences in genetic ancestry across the included GWASs, particularly between FinnGen and other European-ancestry studies, may have influenced the estimated genetic correlations. Future studies using more ancestrally homogeneous datasets will be valuable for validating and refining these estimates. Sixth, the use of cancer-free controls in several source GWASs may have modestly influenced estimates of cross-cancer genetic correlation because controls were depleted of risk alleles associated with other cancer types. Future studies using alternative control definitions may help further evaluate the impact of this design choice on pleiotropy estimates. Additionally, the exclusion of individuals with multiple cancer diagnoses in sensitivity analyses may have reduced the ability to detect certain pleiotropic effects, as these individuals could be enriched for shared genetic susceptibility across cancers. Of note, the relatively modest sample sizes of several source GWASs may have limited statistical power to detect shared genetic signals, particularly for less common cancers. Finally, all analyses were restricted to individuals of European ancestry, limiting generalizability to other populations.

In conclusion, our results show that shared genetic architecture across cancers is generally modest at the genome-wide level but becomes apparent when examined at regional scales. Local genetic correlation analyses identified multi-cancer genomic regions that are not captured by global correlations, while cross-cancer PRS analyses highlighted a limited set of cancer pairs with polygenic overlap. Annotations of correlated regions and PRS further indicated that this shared risk is structured and enriched in specific genetic domains rather than uniformly distributed across the genome. Together, these findings refine current understanding of cross-cancer genetic sharing and provide an integrative framework for studying how shared polygenic susceptibility manifests across cancer types.

Supporting Information

S1 Table. GWAS summary statistics.

(XLSX)

pgen.1012308.s001.xlsx (10.6KB, xlsx)
S2 Table. Cancer codes in the UK Biobank.

(XLSX)

pgen.1012308.s002.xlsx (9.6KB, xlsx)
S3 Table. Significant local genetic correlations.

(XLSX)

pgen.1012308.s003.xlsx (17.4KB, xlsx)
S4 Table. Correlations with non-cancer phenotypes in chr6:167178790-168548525.

(XLSX)

pgen.1012308.s004.xlsx (7.7KB, xlsx)
S5 Table. Correlations with non-cancer phenotypes in chr8:128166556–128542444.

(XLSX)

pgen.1012308.s005.xlsx (8.2KB, xlsx)
S6 Table. Correlations with non-cancer phenotypes in chr9:18660695–19129349.

(XLSX)

pgen.1012308.s006.xlsx (8.9KB, xlsx)
S7 Table. Correlations with non-cancer phenotypes in chr11:63154309–66835194.

(XLSX)

pgen.1012308.s007.xlsx (9.6KB, xlsx)
S8 Table. GWAS Catalog associations for selected regions.

(XLSX)

pgen.1012308.s008.xlsx (16.5KB, xlsx)
S9 Table. Associations between PRS and target cancers.

(XLSX)

S10 Table. Associations between bladder-lung cancer pair stratified by smoking status.

(XLSX)

pgen.1012308.s010.xlsx (7.8KB, xlsx)
S11 Table. Cross-cancer associations of PD-PRSs.

(XLSX)

pgen.1012308.s011.xlsx (11.7KB, xlsx)
S12 Table. Number of SNPs included in PS-PRSs.

(XLSX)

pgen.1012308.s012.xlsx (26.7KB, xlsx)
S13 Table. Associations between PS-PRSs and target cancers.

(XLSX)

pgen.1012308.s013.xlsx (20.2KB, xlsx)
S14 Table. Cross-cancer associations of PS-PRSs.

(XLSX)

pgen.1012308.s014.xlsx (30.2KB, xlsx)
S1 Fig. Correlations and shared high-risk proportions across PRSs.

(A) Heatmap of pairwise correlations among 15 cancer PRSs. Each cell represents the Pearson correlation coefficient, with red indicating positive correlations and blue indicating negative correlations. Statistically significant correlations (p-value < 0.05) are marked with asterisks in the upper triangle, while correlation coefficients are shown in the lower triangle. Although many PRSs were significantly correlated, the magnitude of correlations was generally modest. (B) Heatmap of the proportion of shared high-risk participants (top 10% PRS) across 15 cancer PRSs. Each cell represents the proportion of individuals classified as high risk for both PRSs, with numeric values displayed in the lower triangle.

(TIF)

pgen.1012308.s015.tif (1.8MB, tif)
S2 Fig. Number of SNPs included in PD-PRSs.

This figure displays the number of SNPs included in each of the 14 PD-PRSs for bladder cancer (red), CLL (green), lung cancer (blue), and NMSC (purple). Each bar corresponds to one PD-PRS, and bar height indicates the number of SNPs assigned to that component. Across all three cancers, the other PD-PRS contained the largest number of SNPs, followed by the diabetes-related PD-PRS.

(TIF)

pgen.1012308.s016.tif (616.4KB, tif)
S3 Fig. Associations between PD-PRSs and target cancers.

This figure shows the ORs and 95% CIs for associations between PD-PRSs and their corresponding target cancers for (A) bladder cancer, (B) CLL, (C) lung cancer, and (D) NMSC. Associations that remained significant after Bonferroni correction (p-value < 8.93 × 10-4) are indicated with asterisks. All PD-PRSs were significantly associated with the corresponding target cancer outcome, with exception of the alcohol consumption- and hypertension and blood pressure-related PD-PRSs for bladder cancer and alcohol consumption– and chronic kidney disease–related PD-PRSs for CLL.

(TIF)

pgen.1012308.s017.tif (1.4MB, tif)
S4 Fig. Gene overlap across PS-PRSs.

This figure summarizes gene overlap across PS-PRSs using UpSet plots (A, C, E, G, I, K) and pairwise Jaccard similarity heatmaps (B, D, F, H, J, L) for six cancer pairs: (A-B) lung cancer PS-PRSs associated with bladder cancer; (C-D) NMSC PS-PRSs associated with breast cancer; (E-F) bladder cancer PS-PRSs associated with lung cancer; (G-H)NMSC PS-PRSs associated with MSC; (I-J) CLL PS-PRSs associated with NMSC; and (K-L) lung cancer PS-PRSs associated with NMSC. Across all pairs, overlap of genes across PS-PRSs was limited, and Jaccard similarity coefficients were generally low, indicating that cross-cancer polygenic overlap is distributed across distinct pathway-level gene sets rather than driven by a shared set of individual genes.

(TIF)

pgen.1012308.s018.tif (9.5MB, tif)

Acknowledgments

We conducted the research using the UK Biobank resource under an approved data request (ref: 29900). We thank many GWAS consortia for making their GWAS summary data publicly accessible. The funder had no role in the design of the study; the collection, analysis, or interpretation of the data; or the writing of the manuscript and decision to submit it for publication. We want to acknowledge the participants and investigators of the FinnGen study.

Data Availability

The individual genotype and phenotype data underlying this article were provided by the UK Biobank by permission (ref: 29900), and the instructions to apply for the data can be found at https://www.ukbiobank.ac.uk/enable-your-research/apply-for-access. The GWAS summary statistics were downloaded from publicly available databases, and the information on related articles was available in Methods. The summary-level data (e.g. PRS weights) are available on Zenodo (https://zenodo.org/records/22694018).

Funding Statement

This was supported in part by National Institute of Health (NIH; https://www.nih.gov/) grant R01 HG012735 and P50 CA196530 to HZ. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Mucci LA, Hjelmborg JB, Harris JR, Czene K, Havelick DJ, Scheike T, et al. Familial Risk and Heritability of Cancer Among Twins in Nordic Countries. JAMA. 2016;315(1):68–76. doi: 10.1001/jama.2015.17703 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Zhang H, Ahearn TU, Lecarpentier J, Barnes D, Beesley J, Qi G, et al. Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses. Nat Genet. 2020;52(6):572–81. doi: 10.1038/s41588-020-0609-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Schumacher FR, Al Olama AA, Berndt SI, Benlloch S, Ahmed M, Saunders EJ, et al. Association analyses of more than 140,000 men identify 63 new prostate cancer susceptibility loci. Nat Genet. 2018;50: 928–36. [DOI] [PMC free article] [PubMed]
  • 4.McKay JD, Hung RJ, Han Y, Zong X, Carreras-Torres R, Christiani DC, et al. Large-scale association analysis identifies new lung cancer susceptibility loci and heterogeneity in genetic susceptibility across histological subtypes. Nat Genet. 2017;49(7):1126–32. doi: 10.1038/ng.3892 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wang Y, Wu X-S, He J, Ma T, Lei W, Shen Z-Y. A novel TP53 variant (rs78378222 A > C) in the polyadenylation signal is associated with increased cancer susceptibility: evidence from a meta-analysis. Oncotarget. 2016;7(22):32854–65. doi: 10.18632/oncotarget.9056 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Sampson JN, Wheeler WA, Yeager M, Panagiotou O, Wang Z, Berndt SI, et al. Analysis of Heritability and Shared Heritability Based on Genome-Wide Association Studies for Thirteen Cancer Types. J Natl Cancer Inst. 2015;107(12):djv279. doi: 10.1093/jnci/djv279 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Jiang X, Finucane HK, Schumacher FR, Schmit SL, Tyrer JP, Han Y, et al. Shared heritability and functional enrichment across six solid cancers. Nat Commun. 2019;10(1):431. doi: 10.1038/s41467-018-08054-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lindström S, Wang L, Feng H, Majumdar A, Huo S, Macdonald J, et al. Genome-wide analyses characterize shared heritability among cancers and identify novel cancer susceptibility regions. J Natl Cancer Inst. 2023;115(6):712–32. doi: 10.1093/jnci/djad043 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Graff RE, Cavazos TB, Thai KK, Kachuri L, Rashkin SR, Hoffman JD, et al. Cross-cancer evaluation of polygenic risk scores for 16 cancer types in two large cohorts. Nat Commun. 2021;12(1):970. doi: 10.1038/s41467-021-21288-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Phelan CM, Kuchenbaecker KB, Tyrer JP, Kar SP, Lawrenson K, Winham SJ, et al. Identification of 12 new susceptibility loci for different histotypes of epithelial ovarian cancer. Nat Genet. 2017;49(5):680–91. doi: 10.1038/ng.3826 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613(7944):508–18. doi: 10.1038/s41586-022-05473-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bulik-Sullivan BK, Loh P-R, Finucane HK, Ripke S, Yang J, Schizophrenia Working Group of the Psychiatric Genomics Consortium, et al. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat Genet. 2015;47(3):291–5. doi: 10.1038/ng.3211 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sudlow C, Gallacher J, Allen N, Beral V, Burton P, Danesh J, et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12(3):e1001779. doi: 10.1371/journal.pmed.1001779 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Xu L, Zhou G, Jiang W, Zhang H, Dong Y, Guan L, et al. JointPRS: A data-adaptive framework for multi-population genetic risk prediction incorporating genetic correlation. Nat Commun. 2025;16(1):3841. doi: 10.1038/s41467-025-59243-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lu Q, Li B, Ou D, Erlendsdottir M, Powles RL, Jiang T, et al. A Powerful Approach to Estimating Annotation-Stratified Genetic Covariance via GWAS Summary Statistics. Am J Hum Genet. 2017;101(6):939–64. doi: 10.1016/j.ajhg.2017.11.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhang Y, Lu Q, Ye Y, Huang K, Liu W, Wu Y, et al. SUPERGNOVA: local genetic correlation analysis reveals heterogeneous etiologic sharing of complex traits. Genome Biol. 2021;22(1):262. doi: 10.1186/s13059-021-02478-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zhang C, Zhang Y, Zhang Y, Zhao H. Benchmarking of local genetic correlation estimation methods using summary statistics from genome-wide association studies. Brief Bioinform. 2023;24(6):bbad407. doi: 10.1093/bib/bbad407 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Howrigan DP, Abbott L, rkwalters, Palmer D, Francioli L, Hammerbacher J. Nealelab/UK_Biobank_GWAS: v2. Zenodo; 2023. doi: 10.5281/ZENODO.8011558 [DOI] [Google Scholar]
  • 19.Werme J, van der Sluis S, Posthuma D, de Leeuw CA. An integrated framework for local genetic correlation analysis. Nat Genet. 2022;54(3):274–82. doi: 10.1038/s41588-022-01017-y [DOI] [PubMed] [Google Scholar]
  • 20.Watanabe K, Taskesen E, van Bochoven A, Posthuma D. Functional mapping and annotation of genetic associations with FUMA. Nat Commun. 2017;8(1):1826. doi: 10.1038/s41467-017-01261-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Zhou G, Zhao H. A fast and robust Bayesian nonparametric method for prediction of complex traits using summary statistics. PLoS Genet. 2021;17(7):e1009697. doi: 10.1371/journal.pgen.1009697 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hu J, Ye Y, Zhou G, Zhao H. Using clinical and genetic risk factors for risk prediction of 8 cancers in the UK Biobank. JNCI Cancer Spectr. 2024;8(2):pkae008. doi: 10.1093/jncics/pkae008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559–75. doi: 10.1086/519795 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Hu J, Ye Y, Zhang C, Ruan Y, Natarajan P, Zhao H. Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk. PLoS Comput Biol. 2025;21(6):e1013191. doi: 10.1371/journal.pcbi.1013191 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hu J, Zhou G, Zhao H, DeWan AT. Leveraging pleiotropy to improve genetic risk prediction across diseases. medRxiv. 2025. 10.1101/2025.06.16.25329688 [DOI] [PMC free article] [PubMed]
  • 26.Kanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 2000;28(1):27–30. doi: 10.1093/nar/28.1.27 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Papier K, Atkins JR, Tong TYN, Gaitskell K, Desai T, Ogamba CF, et al. Identifying proteomic risk factors for cancer using prospective and exome analyses of 1463 circulating proteins and risk of 19 cancers in the UK Biobank. Nat Commun. 2024;15(1):4010. doi: 10.1038/s41467-024-48017-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Chahal HS, Wu W, Ransohoff KJ, Yang L, Hedlin H, Desai M, et al. Genome-wide association study identifies 14 novel risk alleles associated with basal cell carcinoma. Nat Commun. 2016;7:12510. doi: 10.1038/ncomms12510 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Song H, Ramus SJ, Tyrer J, Bolton KL, Gentry-Maharaj A, Wozniak E, et al. A genome-wide association study identifies a new ovarian cancer susceptibility locus on 9p22.2. Nat Genet. 2009;41(9):996–1000. doi: 10.1038/ng.424 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Kar SP, Beesley J, Amin Al Olama A, Michailidou K, Tyrer J, Kote-Jarai Zs, et al. Genome-Wide Meta-Analyses of Breast, Ovarian, and Prostate Cancer Association Studies Identify Multiple New Susceptibility Loci Shared by at Least Two Cancer Types. Cancer Discov. 2016;6(9):1052–67. doi: 10.1158/2159-8290.CD-15-1227 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Nam RK, Zhang WW, Loblaw DA, Klotz LH, Trachtenberg J, Jewett MAS, et al. A genome-wide association screen identifies regions on chromosomes 1q25 and 7p21 as risk loci for sporadic prostate cancer. Prostate Cancer Prostatic Dis. 2008;11(3):241–6. doi: 10.1038/sj.pcan.4501010 [DOI] [PubMed] [Google Scholar]
  • 32.Dimopoulou O, Fuller H, Richmond RC, Bouras E, Hayes B, Dimou N, et al. Mendelian randomization study of sleep traits and risk of colorectal cancer. Sci Rep. 2025;15(1):13478. doi: 10.1038/s41598-024-83693-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Pearson-Stuttard J, Papadimitriou N, Markozannes G, Cividini S, Kakourou A, Gill D, et al. Type 2 Diabetes and Cancer: An Umbrella Review of Observational and Mendelian Randomization Studies. Cancer Epidemiol Biomarkers Prev. 2021;30(6):1218–28. doi: 10.1158/1055-9965.EPI-20-1245 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Dube U, Ibanez L, Budde JP, Benitez BA, Davis AA, Harari O, et al. Overlapping genetic architecture between Parkinson disease and melanoma. Acta Neuropathol. 2020;139(2):347–64. doi: 10.1007/s00401-019-02110-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Besson C, Moore A, Wu W, Vajdic CM, de Sanjose S, Camp NJ, et al. Common genetic polymorphisms contribute to the association between chronic lymphocytic leukaemia and non-melanoma skin cancer. Int J Epidemiol. 2021;50(4):1325–34. doi: 10.1093/ije/dyab042 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Wilson C, Kanhere A. 8q24.21 Locus: A Paradigm to Link Non-Coding RNAs, Genome Polymorphisms and Cancer. Int J Mol Sci. 2021;22(3):1094. doi: 10.3390/ijms22031094 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Youn EK, Cho HM, Jung JK, Yoon G-E, Eto M, Kim JI. Pathologic HDAC1/c-Myc signaling axis is responsible for angiotensinogen transcription and hypertension induced by high-fat diet. Biomed Pharmacother. 2023;164:114926. doi: 10.1016/j.biopha.2023.114926 [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Kent W Hunter, Heather J Cordell

12 May 2026

PGENETICS-D-26-00272

Identifying shared polygenic risk across cancers

PLOS Genetics

Dear Dr. Zhao,

Thank you for submitting your manuscript to PLOS Genetics. After careful consideration, we feel that it has merit but does not fully meet PLOS Genetics's publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Jul 11 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosgenetics@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pgenetics/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

* A letter that responds to each point raised by the editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.

* A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

* An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, competing interests statement, or data availability statement, please make these updates within the submission form at the time of resubmission. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only the individual author can complete the verification step; PLOS staff cannot verify ORCID iDs on behalf of authors.

We look forward to receiving your revised manuscript.

Kind regards,

Kent W. Hunter

Section Editor

PLOS Genetics

Kent Hunter

Section Editor

PLOS Genetics

Aimée Dudley

Editor-in-Chief

PLOS Genetics

Anne Goriely

Editor-in-Chief

PLOS Genetics

Journal Requirements:

1) Please ensure that the CRediT author contributions listed for every co-author are completed accurately and in full.

At this stage, the following Authors/Authors require contributions: jiaqi hu. Please ensure that the full contributions of each author are acknowledged in the "Add/Edit/Remove Authors" section of our submission form.

The list of CRediT author contributions may be found here: https://journals.plos.org/plosgenetics/s/authorship#loc-author-contributions

2) Please provide an Author Summary. This should appear in your manuscript between the Abstract (if applicable) and the Introduction, and should be 150-200 words long. The aim should be to make your findings accessible to a wide audience that includes both scientists and non-scientists. Sample summaries can be found on our website under Submission Guidelines:

https://journals.plos.org/plosgenetics/s/submission-guidelines#loc-parts-of-a-submission

3) Please upload all main figures as separate Figure files in .tif or .eps format. For more information about how to convert and format your figure files please see our guidelines:

https://journals.plos.org/plosgenetics/s/figures

4) We have noticed that you have uploaded Supporting Information files, but you have not included a list of legends. Please add a full list of legends for your Supporting Information files after the references list.

5) Thank you for stating that "The summary-level data (e.g. PRS weights) will be available on the PGS catalog (https://www.pgscatalog.org/) once published." We strongly recommend all authors deposit their data before acceptance, as the process can be lengthy and hold up publication timelines. Please note that, though access restrictions are acceptable now, your entire minimal dataset will need to be made freely accessible if your manuscript is accepted for publication. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board.

6) Your current Financial Disclosure states, "The author(s) received no specific funding for this work.".

However, your funding information on the submission form indicates receiving funds.Please ensure that the funders and grant numbers match between the Financial Disclosure field and the Funding Information tab in your submission form. Note that the funders must be provided in the same order in both places as well.

Please amend your detailed Financial Disclosure statement. This is published with the article. It must therefore be completed in full sentences and contain the exact wording you wish to be published.

1) Please clarify all sources of financial support for your study. List the grants, grant numbers, and organizations that funded your study, including funding received from your institution. Please note that suppliers of material support, including research materials, should be recognized in the Acknowledgements section rather than in the Financial Disclosure

2) State the initials, alongside each funding source, of each author to receive each grant. For example: "This work was supported by the National Institutes of Health (####### to AM; ###### to CJ) and the National Science Foundation (###### to AM)."

3) State what role the funders took in the study. If the funders had no role in your study, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

4) If any authors received a salary from any of your funders, please state which authors and which funders.

7) Please revise your current Competing Interest statement to the standard "The authors have declared that no competing interests exist."

Note: If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

Reviewers' comments:

Reviewer's Responses to Questions

Reviewer #1: This paper uses publicly available summary genome-wide association data and individual-level data from the UK Biobank to assess the genetic correlations between cancers and decompose these correlations into putative functional components. The top-line findings are in line with previous studies [PMIDs 36929942 30683880 26464424 27432226]: although the overall genetic correlations between cancers are weak, there are regions with stronger correlations among multiple cancers (e.g. TERT, HLA, 8q24, ABO, others discussed below).The authors use these local correlations between cancers and correlations between cancers and other traits to assess the mechanisms driving these correlations. These analyses take an approach that has previously been used to understand the polygenic contribution to individual traits (including at least one cancer) [41315867 38374256 38443691 https://www.medrxiv.org/content/10.1101/2025.05.15.25327708v1] and apply it to multiple traits.

The paper is hard to follow in some places and could use more details in the methods and results sections (specifics below). Some of the authors’ interpretations of their results are “too little” (e.g. tautological restatements of the empirical results) and some are “too much,” latching on to one possible explanation for the results, without considering alternatives (again, specifics below). There are also some analytic choices that I would not have made, and I didn’t find the authors’ justification compelling: notably, the focus on analyses removing individuals with multiple cancers, which shifts the focus away from shared genetic contributions to distinct genetic and environmental contributions.

But the biggest limitation of this paper is the limited power in most of the input GWAS and the UK Biobank.The GWAS summary statistics for 12 of the 24 cancers they considered come from a single biobank, FinnGen, with small case numbers (median of 1,894 cases, mean of 3,780 cases). As a consequence, one plausible explanation for many of these results is “small sample size”—especially any result that highlights the absence of a correlation.

There are larger GWAS summary statistics available for many of these cancers, see e.g https://github.com/xueyaowunci/cancer-context-fine-mapping. Results from analyses on larger GWAS will be more robust.

Specific comments follow

1. Methods:

1.1 The details in the methods section are a little thin, which makes it difficult for the reader to assess or interpret the results without reading cited methodological or data resources. The authors don’t need to recapitulate all of the information in the original papers, but some important high-level details would be very helpful.

For example:

–Provide URLs and date accessed and/or persistent identifier for cancer GWAS summary statistics;

–Provide an overview of how the FinnGen PanCancer GWAS was run (any record of any cancer ever versus no record of any cancer ever?)

–Clarify the definition of controls in the UK Biobank. “...participants without a recorded diagnosis of cancer”--does that mean diagnosis of *any* cancer, or of the target cancer?

–In FinnGen and UKB, are cases incident after blood draw or a mix of prevalent and incident cases?

–Does SDPR generate “infinitessimal” PRSes (i.e. every variant has a non-zero weight although some may be extremely shrunk) or does it learn a sparsity parameter (i.e. some percentage of variants get zero weight)?

1.2. On FinnGen, other than sample size:

–Use of FinnGen’s “PanCancer” results may complicate interpretation: if the outcome in the PanCacer GWAS is diagnosis with any cancer, then of course cancers A, B, and C will be genetically correlated with the PanCancer trait, even if the genetic contributions to A, B, and C are orthogonal.

–I wonder whether the correlations with cancers whose summary stats do not come from FinnGen are deflated due to mis-match in genetic variation between sampled populations. Minor allele frequencies and linkage disequilibrium patterns will differ between large aggregate European ancestry samples and Finnish samples. This is always an issue when comparing summary statistics from different studies, and the further the genetic distance of the compared studies, the bigger the potential problem. FinnGen-EUR is not as large a mismatch as EAS-EUR, e.g., but might be enough to matter (esp combined with small sample size). Bringing in other large European ancestry GWAS summary statistics might help here, but the potential confounding of genetic correlations by sample ancestry differences should be mentioned as a limitation.

1.3. On the use of “supercontrols:”

Using “supercontrols” (individuals without a personal history of *any* cancer) in the FinnGen individual cancer analyses (and perhaps also in the UK Biobank analyses) will, in principle, induce associations between the target cancer and variants related to cancers other than the target cancer. (Folks never diagnosed with cancers B, C, and D likely carry fewer risk alleles for cancers B, C, and D.) This could induce genetic correlations among cancers, especially more common cancers (ahem, NMSC). The magnitude of this potential bias is unclear, but the possibility should be noted.

1.4. When assessing the association between the PRS for cancer A and cancer B, the authors adjust for the PRS for cancer B. The authors give no justification for this adjustment, and I’m not sure it is appropriate in this context. For variants that are truly associated with both cancers, would this adjustment (inappropriately) weaken the correlation between the PRS for A and cancer B? For variants that are only associated with cancer A but happen to be linkage disequilibrium (LD) with a variant truly associated with B, would this procedure actually remove the confounding due to LD? Those are honest Qs. Please provide some motivation for this adjustment and justification that it does what it should be doing.

1.5. Similarly, I didn’t understand the justification for removing individuals with multiple cancers. The authors note that “associations could be driven by comorbidity between cancers,” but, as I read that, “comorbidity” is a synonym for pleiotropy and hence a feature, not a bug. It‘s part of the phenomenon motivating this whole exercise. Maybe the authors are trying to rule out the situation where G -> cancer A -> cancer B, but (a) that’s still of interest, (b) I suspect that is not a widespread phenomenon, (c) if that’s the rationale, the order of diagnoses matters and that is not considered here, and (d) many folks with multiple cancers got them because of some underlying dysfunction that is independently raising risk for both cancers–excluding those people is excluding the people who are very informative for understanding pleiotropy. Not only are the genetic correlations estimated after removing the overlap noisier, they are systematically attenuated towards the null–which I think is an artifact of removing individuals with multiple cancers.

1.6. How were pleiotropy groups defined? How is genetic correlation with a pleiotropy group estimated? I’m interested in reading the preprint (ref. 24), but until I do I have to guess at what was done here.

1.7. Comparing Nagelkerke’s R2 between Equations 1, 2 and 3:

–How was missing data (if any) handled? If different numbers of individuals were included in fitting models 1-3, comparing Nagelkerke’s R2s will be invalid.

–Was the permutation test done so as to preserve the correlation between cancer i (but not j) and age, sex, and top 10 PCs? If so, how? If not, this procedure is invalid, because the observed data includes any real-world associations between cancer i and the covariates but the permuted data does not.

1.8. “Unless otherwise specified, multiple testing was controlled using Bonferroni correction.” Just cut this, as the Bonferroni correction used varies from analysis to analysis (as the number of tests varies from analysis to analysis). You’ve done a good job saying what the thresholds used were for various analyses and why, so this blanket statement is unnecessary..

2. Results:

2.1. Seems to be a formatting error at the bottom of page 12. Some text has gone mising?

2.2. Page 13. “Two cancer pairs with significant genome-wide correlations did not show any significant local genetic contributions, which suggests the shared genetic signal for these pairs is likely distributed across the genome rather than driven by specific regions.” Does it? It also seems likely that this could be due to limited power. These two pairs have the smallest number of cases among the pairs with significant genome-wide correlations. Absence of evidence is not evidence of absence.

2.3. Page 14 and Table 1. Including a “landmark” (gene or cytoband) in Table 1 will help many readers who haven’t memorized chr:pos for all regions of interest, esp known multi-cancer loci (e.g. 1q25, 2q33/CASP8, HLA, 8q24/MYC, 6q27, 9p22, 11q12). The cancers for 2q33 are repeated; typo?

Why exclude the HLA-overlapping region from downstream analysis?

Worth commenting here or in discussion on what’s known already about these loci: e.g. Lindstrom JNCI found individual variants at 2q33 pleiotropic for breast, prostate, and melanoma; at HLA with lung, prostate, head and neck, colorectal, endometrial, and lung cancers; at 8q24 with breast, prostate, and colorectal; and at 11q12 with breast and prostate cancer; Kar Cancer Discovery 2016 found evidence for pleiotropy with breast, prostate and ovarian at 11q12; Kong (https://www.biorxiv.org/content/10.64898/2026.01.06.697866v1 )reviews evidence for associations at 2q33 with breast cancer, melanoma, NMSC,and lung cancer; etc.

More striking is the absence of some well-documented multi-cancer loci, such as TERT and ABO. Why were these not identified? Is this a result of the statistical method, the input data, or both?

2.4. Page 14. For genetic correlations with GWAS results from the Neale lab, please cite the Phenotype Code. For GWAS catalog look ups, please cite the EFO ID.

2.5. Page 14. “Variants were enriched for four cancer-related GWAS traits.” These are not cancer-related traits, they are cancers. You can say “Variants were enriched for four cancers” and list them.

2.6. Page 15. “three cancers were correlated,” “four cancers were correlated”: list them.

2.7. Page 16. “To further evaluate potential confounding by smoking…” (a) Is this confounding or mediation? If the genetic contributions to bladder cancer and lung cancer are distinct and are not correlated with smoking, then we should not see a genetic correlation between the cancers. If the genetic components are correlated with smoking, then there would be a genetic correlation–via the mechanism of smoking (which is still of interest). (b) Failure to find “significant” associations between lung cancer and the bladder cancer PRS may just be a power thing; the effect estimates are similar across smoking strata and in the pooled analysis. I don’t know that removing the lung-bladder PRS pair from downstream analysis is warranted; nor do I think you can conclude based on p-values (one of which just missed “significance”) that these results “indicate context specific patterns of polygenic association.”

2.8. Page 17. It took me several reads and flipping back and forth between pages to understand why you only focused on the lung, NMSC and CLL PRS. For thick people like me, you might make this explicit. “We selected the PRS involved in these cancer-PRS pairs (lung, NMSC, and CLL PRS) for downstream analysis. We applied two decomposition frameworks to these PRS to further characterize…”

2.9. Page 18. “Consistent with metabolic contributions to lung cancer.” Reference? Are metabolic changes (e.g. metabolic syndrome) causally related to lung cancer (which is how I read “contributes to”) or are they associated with lung cancer (markers of latent disease development)?

2.10. Page 19. “To formally assess…” I have concerns about the validity of these analyses (see 1.7 above). And even if they were valid, I am not sure what they add beyond the PD-PRS results presented in the previous paragraph.

3. Discussion, page 23: “the shared polygenic signal between bladder and lung cancers is enriched… for inflammatory traits, such as cardiovascular and neuropsychiatric phenotypes, which have been implicated in cancer susceptibility.” (a) Neuropsychiatric traits are inflammatory? Maybe, in part, as these are complex traits that involve multiple systems. But using that standard there are few complex diseases that are not inflammatory. More importantly, because the definition of neuropsychiatric traits is not given here (it’s in a reference preprint), it’s hard to evaluate this statement. (b) Language is unclear. “...which have been implicated…”: what does “which” refer to here? Inflammatory traits? Or cardiovascular and neuropsychiatric traits? (c) Are these the best references here? They are not specific to bladder or lung cancer, and focus on shared CVD & cancer risk factors (not CVD and cancer per se) or on neuroinflammation specifically (as opposed to neuropsychiatric traits broadly)? There’s also a bit of conflation of cancer (treatment) sequelae and cancer risk factors in those refs.

Reviewer #2: This manuscript investigates the shared genetic risk across cancers. Through genetic correlation and polygenic risk score analyses, the paper demonstrated that although genome-wide sharing is modest, there are specific regions with stronger sharing of cross-cancer polygenic risk. Overall, the paper addresses the interesting question and uses rigorous methods. The manuscript is generally well written. However, some of the findings are hard to interpret. I have the following comments.

1. Some text at the bottom of page 12, GWAS summary statistics, seems to be missing due to document formatting issues. Please add back the text.

2. Page 14, paragraphs 2 and 3: The authors mentioned local genetic correlation between cancers total protein levels. However, the interpretation is unclear. Is there previous literature that discusses the link between cancer and total protein levels? What are plausible hypotheses?

3. Page 15, paragraph 3: Variants chr1:182294372–183796074 were mapped to chronotype-related phenotypes. The authors should comment on the connection between chronotype and cancer.

4. Page 18, paragraph 1: “For NMSC, the neuropsychiatric disease PD-PRS showed the

strongest association (OR = 1.46, 95% CI: 1.44–1.48), suggesting enrichment of shared genetic effects between NMSC and neuropsychiatric phenotypes at the polygenic level.” This finding seems difficulty to interpret – what are the hypotheses explaining shared genetic architecture between NMSC and neuropsychiatric disorders?

5. The acronyms of some cancers, such as CLL, NHL, are not defined in the main text. Although they are defined in the supplements, it would improve readability to define them in the main text where they first appear.

6. The font size in Figures 3 and 5 is very small. Consider increasing it to improve legibility.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics  data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: No:  The publicly available GWAS results are cited, but URLs and accession numbers are not given.

Reviewer #2: None

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

Figure resubmission:

Reproducibility:

To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Decision Letter 1

Kent W Hunter, Heather J Cordell

26 Aug 2026

Dear Dr Zhao,

We are pleased to inform you that your manuscript entitled "Identifying shared polygenic risk across cancers" has been editorially accepted for publication in PLOS Genetics. Congratulations!

Before your submission can be formally accepted and sent to production you will need to complete our formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Please note: the accept date on your published article will reflect the date of this provisional acceptance, but your manuscript will not be scheduled for publication until the required changes have been made.

Once your paper is formally accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you’ve already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosgenetics@plos.org.

In the meantime, please log into Editorial Manager at https://www.editorialmanager.com/pgenetics/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production and billing process. Note that PLOS requires an ORCID iD for all corresponding authors. Therefore, please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field.  This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager.

If you have a press-related query, or would like to know about making your underlying data available (as you will be aware, this is required for publication), please see the end of this email. If your institution or institutions have a press office, please notify them about your upcoming article at this point, to enable them to help maximise its impact. Inform journal staff as soon as possible if you are preparing a press release for your article and need a publication date.

Thank you again for supporting open-access publishing; we are looking forward to publishing your work in PLOS Genetics!

Yours sincerely,

Heather J Cordell

Academic Editor

PLOS Genetics

Kent Hunter

Section Editor

PLOS Genetics

Aimée Dudley

Editor-in-Chief

PLOS Genetics

Anne Goriely

Editor-in-Chief

PLOS Genetics

www.plosgenetics.org

BlueSky: @plos.bsky.social

----------------------------------------------------

Comments from the reviewers (if applicable):

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #2: The authors have addressed my comments.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics  data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #2: None

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #2: No

----------------------------------------------------

Data Deposition

If you have submitted a Research Article or Front Matter that has associated data that are not suitable for deposition in a subject-specific public repository (such as GenBank or ArrayExpress), one way to make that data available is to deposit it in the Dryad Digital Repository. As you may recall, we ask all authors to agree to make data available; this is one way to achieve that. A full list of recommended repositories can be found on our website.

The following link will take you to the Dryad record for your article, so you won't have to re‐enter its bibliographic information, and can upload your files directly:

http://datadryad.org/submit?journalID=pgenetics&manu=PGENETICS-D-26-00272R1

More information about depositing data in Dryad is available at http://www.datadryad.org/depositing. If you experience any difficulties in submitting your data, please contact help@datadryad.org for support.

Additionally, please be aware that our data availability policy requires that all numerical data underlying display items are included with the submission, and you will need to provide this before we can formally accept your manuscript, if not already present.

----------------------------------------------------

Press Queries

If you or your institution will be preparing press materials for this manuscript, or if you need to know your paper's publication date for media purposes, please inform the journal staff as soon as possible so that your submission can be scheduled accordingly. Your manuscript will remain under a strict press embargo until the publication date and time. This means an early version of your manuscript will not be published ahead of your final version. PLOS Genetics may also choose to issue a press release for your article. If there's anything the journal should know or you'd like more information, please get in touch via plosgenetics@plos.org.

Acceptance letter

Kent W Hunter, Heather J Cordell

PGENETICS-D-26-00272R1

Identifying shared polygenic risk across cancers

Dear Dr Zhao,

We are pleased to inform you that your manuscript entitled "Identifying shared polygenic risk across cancers" has been formally accepted for publication in PLOS Genetics! Your manuscript is now with our production department and you will be notified of the publication date in due course.

The corresponding author will soon be receiving a typeset proof for review, to ensure errors have not been introduced during production. Please review the PDF proof of your manuscript carefully, as this is the last chance to correct any errors. Please note that major changes, or those which affect the scientific understanding of the work, will likely cause delays to the publication date of your manuscript.

Soon after your final files are uploaded, unless you have opted out or your manuscript is a front-matter piece, the early version of your manuscript will be published online. The date of the early version will be your article's publication date. The final article will be published to the same URL, and all versions of the paper will be accessible to readers.

For Research Articles, you will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

Thank you again for supporting PLOS Genetics and open-access publishing. We are looking forward to publishing your work!

With kind regards,

Janani Seenivasan

PLOS Genetics

On behalf of:

The PLOS Genetics Team

Carlyle House, Carlyle Road, Cambridge CB4 3DN | United Kingdom

plosgenetics@plos.org | +44 (0) 1223-442823

plosgenetics.org | Twitter: @PLOSGenetics

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Table. GWAS summary statistics.

    (XLSX)

    pgen.1012308.s001.xlsx (10.6KB, xlsx)
    S2 Table. Cancer codes in the UK Biobank.

    (XLSX)

    pgen.1012308.s002.xlsx (9.6KB, xlsx)
    S3 Table. Significant local genetic correlations.

    (XLSX)

    pgen.1012308.s003.xlsx (17.4KB, xlsx)
    S4 Table. Correlations with non-cancer phenotypes in chr6:167178790-168548525.

    (XLSX)

    pgen.1012308.s004.xlsx (7.7KB, xlsx)
    S5 Table. Correlations with non-cancer phenotypes in chr8:128166556–128542444.

    (XLSX)

    pgen.1012308.s005.xlsx (8.2KB, xlsx)
    S6 Table. Correlations with non-cancer phenotypes in chr9:18660695–19129349.

    (XLSX)

    pgen.1012308.s006.xlsx (8.9KB, xlsx)
    S7 Table. Correlations with non-cancer phenotypes in chr11:63154309–66835194.

    (XLSX)

    pgen.1012308.s007.xlsx (9.6KB, xlsx)
    S8 Table. GWAS Catalog associations for selected regions.

    (XLSX)

    pgen.1012308.s008.xlsx (16.5KB, xlsx)
    S9 Table. Associations between PRS and target cancers.

    (XLSX)

    S10 Table. Associations between bladder-lung cancer pair stratified by smoking status.

    (XLSX)

    pgen.1012308.s010.xlsx (7.8KB, xlsx)
    S11 Table. Cross-cancer associations of PD-PRSs.

    (XLSX)

    pgen.1012308.s011.xlsx (11.7KB, xlsx)
    S12 Table. Number of SNPs included in PS-PRSs.

    (XLSX)

    pgen.1012308.s012.xlsx (26.7KB, xlsx)
    S13 Table. Associations between PS-PRSs and target cancers.

    (XLSX)

    pgen.1012308.s013.xlsx (20.2KB, xlsx)
    S14 Table. Cross-cancer associations of PS-PRSs.

    (XLSX)

    pgen.1012308.s014.xlsx (30.2KB, xlsx)
    S1 Fig. Correlations and shared high-risk proportions across PRSs.

    (A) Heatmap of pairwise correlations among 15 cancer PRSs. Each cell represents the Pearson correlation coefficient, with red indicating positive correlations and blue indicating negative correlations. Statistically significant correlations (p-value < 0.05) are marked with asterisks in the upper triangle, while correlation coefficients are shown in the lower triangle. Although many PRSs were significantly correlated, the magnitude of correlations was generally modest. (B) Heatmap of the proportion of shared high-risk participants (top 10% PRS) across 15 cancer PRSs. Each cell represents the proportion of individuals classified as high risk for both PRSs, with numeric values displayed in the lower triangle.

    (TIF)

    pgen.1012308.s015.tif (1.8MB, tif)
    S2 Fig. Number of SNPs included in PD-PRSs.

    This figure displays the number of SNPs included in each of the 14 PD-PRSs for bladder cancer (red), CLL (green), lung cancer (blue), and NMSC (purple). Each bar corresponds to one PD-PRS, and bar height indicates the number of SNPs assigned to that component. Across all three cancers, the other PD-PRS contained the largest number of SNPs, followed by the diabetes-related PD-PRS.

    (TIF)

    pgen.1012308.s016.tif (616.4KB, tif)
    S3 Fig. Associations between PD-PRSs and target cancers.

    This figure shows the ORs and 95% CIs for associations between PD-PRSs and their corresponding target cancers for (A) bladder cancer, (B) CLL, (C) lung cancer, and (D) NMSC. Associations that remained significant after Bonferroni correction (p-value < 8.93 × 10-4) are indicated with asterisks. All PD-PRSs were significantly associated with the corresponding target cancer outcome, with exception of the alcohol consumption- and hypertension and blood pressure-related PD-PRSs for bladder cancer and alcohol consumption– and chronic kidney disease–related PD-PRSs for CLL.

    (TIF)

    pgen.1012308.s017.tif (1.4MB, tif)
    S4 Fig. Gene overlap across PS-PRSs.

    This figure summarizes gene overlap across PS-PRSs using UpSet plots (A, C, E, G, I, K) and pairwise Jaccard similarity heatmaps (B, D, F, H, J, L) for six cancer pairs: (A-B) lung cancer PS-PRSs associated with bladder cancer; (C-D) NMSC PS-PRSs associated with breast cancer; (E-F) bladder cancer PS-PRSs associated with lung cancer; (G-H)NMSC PS-PRSs associated with MSC; (I-J) CLL PS-PRSs associated with NMSC; and (K-L) lung cancer PS-PRSs associated with NMSC. Across all pairs, overlap of genes across PS-PRSs was limited, and Jaccard similarity coefficients were generally low, indicating that cross-cancer polygenic overlap is distributed across distinct pathway-level gene sets rather than driven by a shared set of individual genes.

    (TIF)

    pgen.1012308.s018.tif (9.5MB, tif)
    Attachment

    Submitted filename: Response to review.docx

    pgen.1012308.s019.docx (5.1MB, docx)

    Data Availability Statement

    The individual genotype and phenotype data underlying this article were provided by the UK Biobank by permission (ref: 29900), and the instructions to apply for the data can be found at https://www.ukbiobank.ac.uk/enable-your-research/apply-for-access. The GWAS summary statistics were downloaded from publicly available databases, and the information on related articles was available in Methods. The summary-level data (e.g. PRS weights) are available on Zenodo (https://zenodo.org/records/22694018).


    Articles from PLOS Genetics are provided here courtesy of PLOS

    RESOURCES