Abstract
Aims/hypothesis
Type 2 diabetes has been associated with increased risk of colorectal cancer (CRC), but the specific diabetogenic pathways contributing to this risk remain unclear.
Methods
We analysed individual-level data from 129,420 participants of European ancestry in the Genetics and Epidemiology of Colorectal Cancer Consortium (GECCO) and the Colon Cancer Family Registry (CCFR), comprising 58,531 patients with CRC and 70,889 control participants. We applied eight validated partitioned polygenic scores (PPSs) representing distinct diabetogenic processes: two related to relative insulin secretion insufficiency and six to insulin resistance with varying degrees of preserved insulin secretion. Adjusted ORs and 95% CIs for CRC and early-onset CRC were estimated using conditional logistic regression.
Results
PPSs reflecting insulin resistance-linked hyperinsulinaemia, particularly those related to lipodystrophy, body fat and obesity, were associated with higher odds of CRC (p<0.001 for each). Compared with individuals in the lowest decile, those in the highest decile had ORs of 1.09 (95% CI 1.05, 1.14), 1.13 (1.09, 1.18) and 1.15 (1.10, 1.20) for lipodystrophy, body fat and obesity, respectively. In contrast, PPSs for insulin secretion insufficiency or insulin resistance without hyperinsulinaemia were not associated with CRC. The obesity-related hyperinsulinaemic insulin resistance PPS was also associated with higher odds of early-onset CRC (OR for top vs bottom decile=1.26; 95% CI 1.13, 1.41), with the strongest association among those with obesity (OR 1.75; 95% CI 1.46, 2.11; p value for interaction with BMI <0.001).
Conclusions/interpretation
Diabetogenic processes characterised by insulin resistance-linked hyperinsulinaemia were associated with increased odds of CRC, including early-onset disease. These findings offer new insights into diabetes–CRC pathogenesis, and may inform targeted prevention strategies.
Graphical Abstract
Supplementary Information
The online version contains peer-reviewed but unedited supplementary material available at 10.1007/s00125-025-06631-z.
Keywords: Colorectal cancer, Epidemiology, Genetics, Polygenic scores, Precision health, Type 2 diabetes
Introduction
The global rise in type 2 diabetes presents a major public health challenge, not only due to its cardiometabolic complications but also its association with increased cancer risk, particularly for colorectal cancer (CRC) [1, 2]. Epidemiological studies have consistently demonstrated a positive association between type 2 diabetes and CRC, with meta-analyses reporting a 1.2- to 1.5-fold increased risk among individuals with type 2 diabetes [3]. Although shared risk factors such as ageing, lifestyle behaviours, obesity and social determinants of health probably contribute to this link [4, 5], growing evidence suggests that the dysglycaemic environment that is characteristic of type 2 diabetes may play a direct mechanistic role in CRC risk [6, 7].
Type 2 diabetes is a clinically and aetiologically heterogeneous disease, in which diverse pathophysiological processes affecting distinct tissues converge into a common phenotype of dysglycaemia [8]. This heterogeneity is reflected in varied clinical presentations, differential responses to therapy, and the likelihood of developing complications [8, 9]. It is therefore plausible that specific diabetogenic processes may confer differential CRC risk. Supporting this hypothesis, a recent Mendelian randomisation study demonstrated a causal association between hyperinsulinaemia, but not other glycaemic traits, and CRC risk [10]. Pre-clinical studies have shown that hyperinsulinaemia activates insulin and IGF signalling pathways that promote proliferation, inhibit apoptosis and impair genomic stability, thereby fostering colorectal carcinogenesis [11–13]. However, elevated fasting insulin arises from various underlying mechanisms involving insulin production, secretion and sensitivity processes. Recent work partitioning hyperinsulinaemia into mechanistic clusters has shown that not all elevations in insulin confer equal risk for diabetes or related cardiometabolic complications [14], and these processes may have a differential impact on CRC. Observational studies further reported an elevated CRC risk in the early stages of type 2 diabetes, potentially implicating a metabolic phenotype that is characterised by insulin resistance with compensatory hyperinsulinaemia as a driver of tumorigenesis [5, 15]. Nonetheless, findings across studies remain inconsistent, with some reporting increased CRC risk among individuals with longer diabetes duration [16], underscoring the need for more granular investigation into the specific diabetogenic processes that contribute to CRC susceptibility.
Advances in human genetics and computational biology now offer new opportunities to investigate these relationships at scale. By integrating findings from genome-wide association studies (GWAS) with functional annotations at single-cell resolution, type 2 diabetes-associated genetic variants can be clustered into biologically plausible pathways. This approach enables the construction of partitioned polygenic scores (PPSs) that capture distinct diabetogenic mechanisms [17]. The largest GWAS for type 2 diabetes to date identified eight such scores: two related to relative insulin secretion insufficiency and six representing various forms of insulin resistance with different degrees of hyperinsulinaemia [18]. These PPSs have demonstrated differential associations with cardiometabolic outcomes. For example, polygenic scores for hyperinsulinaemic insulin resistance have been shown to be associated with increased coronary artery disease risk, whereas scores for insulin resistance without hyperinsulinaemia were associated with reduced renal function [19, 20].
In this study, we leveraged individual-level data from 129,420 participants of European ancestry in the Genetics and Epidemiology of Colorectal Cancer Consortium (GECCO) and the Colon Cancer Family Registry (CCFR) to test the hypothesis that distinct diabetogenic processes differentially influence CRC risk. We also investigated the potential impact of these processes on early-onset CRC and their interaction with established risk factors.
Methods
Study population
We included individual-level data pooled from 129,420 participants from 52 studies from GECCO and the CCFR that were initiated between 1976 and 2003, and enrolled patients with CRC diagnosed between 1976 and 2011 with matched control participants. A list of the included studies and a short description of each is provided in electronic supplementary material (ESM) Table 1. In brief, for prospective cohorts, nested case–control sets were constructed by fixing the cohort at a timepoint upon which risk set sampling was employed to select cases and controls. For other case–control studies, population-based controls were used. For all studies, controls were matched for age, sex and race/ethnicity, and on additional factors in some studies. All participants provided written or verbal informed consent, and studies were reviewed and approved by their respective institutional review boards or ethics committees.
Type 2 diabetes PPSs
Detailed descriptions of genotyping, imputation and quality control for included studies have been published previously [21]. To generate PPSs, we used data from a previous GWAS that grouped established type 2 diabetes-associated loci according to shared molecular and physiological similarities [18]; this represents the most comprehensive and up-to-date classification of type 2 diabetes-associated variants into biologically informed pathways. We constructed eight PPSs representing distinct pathophysiological processes: two related to relative insulin secretion insufficiency and six related to insulin resistance with various degrees of preserved insulin secretion. The insulin resistance scores were further categorised into two physiological subtypes based on physiological similarities: (1) those indicative of insulin resistance-linked hyperinsulinaemia, reflecting primarily adipose tissue alterations such as lipodystrophy, body fat, obesity or metabolic syndrome, and (2) those characterised by insulin resistance without hyperinsulinaemia, capturing processes of impaired liver/lipid metabolism or residual dysglycaemia [18]. The clustering approach applied assigned each genetic variant to a single pathway-specific score; therefore, they do not overlap across scores. However, multiple independent SNPs located within the same genomic region (or near the same gene) may appear in different pathways if they have distinct trait associations. Each resulting PPS comprises a different number of genetic variants, ranging from 389 for residual glycaemia to just three for impaired liver/lipid metabolism. A full list of variants and their nearest genes is provided in the ESM Data. The scores were generated by multiplying a variant’s genotype dosage by its respective weight using PLINK, version 2.0 [22] (https://zzz.bwh.harvard.edu/plink/). Polygenic scores were standardised to a mean of zero and SD of 1 to allow comparisons across scores computed in this study with different numbers of genetic variants.
CRC ascertainment
CRC cases were defined as individuals with invasive colorectal adenocarcinoma, and confirmed through review of medical records, pathology reports or death certificates according to individual study cohort procedures. CRC diagnoses were classified according to the International Classification of Diseases for Oncology, 3rd edition (ICD-O-3), using the following topography codes: C18.0–C18.9 (colon), C19.9 (rectosigmoid junction) and C20.9 (rectum). The specific procedures and sources used for CRC ascertainment in each cohort included in this study are detailed in ESM Table 1. A secondary outcome was early-onset CRC, defined as CRC diagnosis before the age of 50 years. All ascertainment methods adhered to standardised protocols to ensure consistency and diagnostic accuracy across study populations.
Statistical analysis
We developed a prespecified protocol, including definitions of exposures, outcomes and covariates, and a statistical analysis plan, prior to data analysis (ESM Methods). We summarised continuous measurements by using means with SD or medians with IQR, and present categorical observations as frequency and percentages.
We used conditional logistic regression to estimate ORs and 95% CI for CRC for each of the eight generated PPSs. PPSs were modelled as a continuous variable (per 1 SD increase), and we also generated categories of the scores based on deciles of the distribution. Models were adjusted for age (continuous: years), sex (categorical: men, women) and the ten first ancestry-derived principal components (categorical).
In sensitivity analyses, we adjusted our models for potential confounders, including BMI (continuous: kg/m2; model 2), family history of CRC (categorical: yes, no; model 3) and environmental exposures including smoking (categorical: ever smoked: yes, no) and alcohol intake (categorical: non-drinker, 1–28 g/day or >28 g/day; model 4). We present both complete-case analyses and analyses using multiple imputation by chained equations with five imputations. All covariates in the primary analysis and the outcome were included in the multiple imputation procedure, and estimates generated from each imputed dataset were combined using Rubin’s rules [23]. To avoid potential inflated type I error rates owing to overlap between the genetic discovery dataset, from which genetic variants and weights to generate the PPSs were identified, and the testing dataset for association with CRC, we performed a sensitivity analysis by excluding 41,459 individuals (n=11,947 cases and n=29,512 control participants) from five studies that were included in the main GWAS discovery study. We performed several subgroup analyses to test the robustness of our primary findings. We also performed a subgroup analysis according to CRC subsites (proximal colon, distal colon and rectum). In addition, we performed stratified analyses by age group (60 years or younger, older than 60 years), sex (men, women), BMI (<25 kg/m2, 25–30 kg/m2, ≥30 kg/m2) and history of type 2 diabetes (yes, no).
In secondary analyses, we used early-onset CRC, defined as a CRC diagnosis before the age of 50 years. For these analyses, we adjusted for the same covariates used in previous models, except age. We performed subgroup analyses for early-onset CRC, and if these analyses yielded divergent results, we tested for interactions between the strata variable and genetic risk on the odds of early-onset CRC.
A two-sided α level of 0.05/8 emerging from the eight tested PPSs was used to determine statistical significance (p<0.006). All statistical analyses were performed using R software, version 4.0.3 (R Foundation, Vienna, Austria).
Results
Among the 129,420 participants with available genetic data from 52 study sets in GECCO and CCFR included in this analysis, 58,531 were patients with CRC and 70,889 were control participants. Women comprised 48% of the study sample (n=62,195). The median age was 63 years (IQR 56–69 years), and the median BMI was 26.6 kg/m2 (IQR 24.1–29.7) (Table 1). Due to the case–control matching design, CRC cases and controls were similar with respect to age, sex and BMI. However, CRC patients were more likely to report a family history of CRC (14.3% vs 9.9%) and a diagnosis of type 2 diabetes (9.4% vs 7.0%) (Table 1).
Table 1.
Characteristics of the included participants
| All participants (n=129,420) | CRC patients (n=58,531) | Control participants (n=70,889) | |
|---|---|---|---|
| Age, years | 63 (56–69) | 64 (57–70) | 62 (55–68) |
| Missing | 1504 (1.2) | 356 (0.6) | 1148 (1.6) |
| Sex, female | 62,195 (48.1) | 27,052 (46.2) | 35,143 (49.6) |
| Missing | - | - | - |
| BMIa, kg/m2 | 26.6 (24.1–29.7) | 26.6 (24.2–29.9) | 26.5 (24.0–29.5) |
| Normal weight | 36,164 (27.9) | 14,589 (24.9) | 21,575 (30.4) |
| Overweight | 45,262 (35.0) | 18,879 (32.3) | 26,383 (37.2) |
| Obesity | 24,850 (19.2) | 11,034 (18.9) | 13,816 (19.5) |
| Missing | 23,144 (17.9) | 14,029 (24.0) | 9115 (12.9) |
| Type 2 diabetes | |||
| Yes | 10,446 (8.1) | 5472 (9.4) | 4974 (7.0) |
| Missing | 25,352 (19.6) | 14,236 (24.3) | 11,116 (15.7) |
| Family history of CRC | |||
| Yes | 15,426 (11.9) | 8387 (14.3) | 7039 (9.9) |
| Missing | 35,926 (27.8) | 14,377 (24.6) | 21,549 (30.4) |
| Smoking | |||
| Ever smoked | 56,410 (43.6) | 25,149 (43.0) | 31,261 (44.1) |
| Missing | 20,939 (16.2) | 12,552 (21.4) | 8387 (11.8) |
| Alcohol intake | |||
| Non-drinker | 35,845 (27.7) | 16,412 (28.0) | 19,433 (27.4) |
| 1–28 g/day | 49,953 (38.6) | 19,125 (32.7) | 30,828 (43.5) |
| >28 g/day | 13,031 (10.1) | 5404 (9.2) | 7627 (10.8) |
| Missing | 30,591 (23.6) | 17,590 (30.1) | 13,001 (18.3) |
| CRC diagnosis site | |||
| Proximal | 16,366 (28.0) | ||
| Distal | 14,845 (25.4) | ||
| Rectal | 16,350 (27.9) | ||
| Missing | 10,970 (18.7) | ||
Values are medians (IQR) for continuous variables, and n (%) for categorical variables
aBMI categories: normal weight, 18.5–24.9 kg/m2; overweight, 25.0–29.9 kg/m2; obesity, ≥30 kg/m2
The eight PPSs showed weak correlation with one another in our study sample (Pearson correlation coefficients <0.05 for all; ESM Table 2), indicating that they capture distinct diabetogenic processes. Despite their weak correlation, we showed that all PPSs were significantly associated with increased odds of type 2 diabetes in this study sample (p<0.001 for all), with adjusted ORs per 1 SD increase in PPS ranging from 1.33 (95% CI 1.31, 1.36) for the beta cell dysfunction score with elevated proinsulin levels to 1.07 (95% CI 1.04, 1.09) for the score related to liver/lipid metabolism (ESM Table 3).
We next examined the associations between PPSs and CRC. PPSs for pathophysiological processes related to insulin resistance-linked hyperinsulinaemia, specifically those reflecting mechanisms for lipodystrophy, body fat and obesity, were significantly associated with increased odds of CRC (p<0.001 for all) (Fig. 1). In contrast, PPSs for relative insulin secretion insufficiency alone and insulin resistance in the absence of hyperinsulinaemia showed no significant associations (Fig. 1). Participants in the highest decile of the insulin resistance-linked hyperinsulinaemia scores had higher odds of CRC compared with those in the lowest decile, with adjusted ORs of 1.13 (95% CI 1.09, 1.18) for the body fat distribution score, 1.15 (95% CI 1.10, 1.20) for the obesity score, 1.09 (95% CI 1.05, 1.14) for the lipodystrophy score, and a non-significant 1.04 (95% CI 1.00, 1.08) for the metabolic syndrome score (Fig. 2).
Fig. 1.
Association between PPSs that denote distinct diabetogenic processes and CRC. BMI-unadjusted and BMI-adjusted ORs and 95% CIs for the association between a 1 SD increase in PPSs and CRC. The reported physiological effects are based on the prior characterisation of these scores in the discovery study by Suzuki et al [18]. In brief, the top four scores with positive insulin secretion and negative insulin sensitivity correspond to the insulin resistance-linked hyperinsulinaemia phenotype. The middle two scores with negative insulin secretion and insulin sensitivity correspond to a phenotype of insulin resistance without hyperinsulinaemia. The bottom two scores with negative insulin secretion and positive insulin sensitivity correspond to a phenotype of relative insulin secretion insufficiency without insulin resistance. Details on the genetic variants and weights used to construct the scores are available in the ESM Data. The BMI-unadjusted model included age, sex and the first ten ancestry-derived principal components. The BMI-adjusted model included these same covariates plus BMI. Adjusted false discovery rate: p<0.006. PI, proinsulin
Fig. 2.
Association between PPSs for insulin resistance-linked hyperinsulinaemia and CRC odds. Adjusted ORs and 95% CIs for the association between PPSs for insulin resistance-linked hyperinsulinaemia and CRC. The PPSs shown in (a–d) are from body fat, obesity, lipodystrophy and metabolic syndrome distributions, respectively. The dots are the crude OR for pathway-specific PPS centiles (error bars indicate ±1 SE). The reference is the 50th PPS centile. The fitted line represents ORs adjusted for age, sex and the first ten ancestry-derived principal components (model 1). Participants in the highest decile of these scores, compared with those in the lowest decile, had adjusted ORs of 1.13 (95% CI 1.09, 1.18; p=1.06 × 10−9) for the body fat distribution score (a), 1.15 (95% CI 1.10, 1.20; p=1.92 × 10−11) for the obesity score (b), 1.09 (95% CI 1.05, 1.14; p=1.36 × 10−5) for the lipodystrophy score (c), and 1.04 (95% CI 1.00, 1.08; p=0.08 [non-significant]) for the metabolic syndrome score (d). Adjusted false discovery rate: p<0.006
We performed several sensitivity analyses to assess the robustness of these findings. The observed associations between insulin resistance-linked hyperinsulinaemia PPSs and CRC remained unchanged after further adjustment for family history of CRC or for smoking status and alcohol intake (ESM Table 4). The results were also consistent when analyses were conducted using multiple imputation by chained equations instead of complete-observations analyses (ESM Table 5). Further, we performed additional analyses by excluding 41,459 individuals (n=11,947 CRC patients and n=29,512 control participants) from five studies that were included in the main type 2 diabetes GWAS discovery study to avoid potential inflated type I error rates owing to overlap between the discovery and the testing datasets. These analyses yielded consistent associations between insulin resistance-linked hyperinsulinaemia PPSs and CRC, with no evidence of attenuation due to type I error (ESM Table 6). These associations were also consistent across strata defined by age, sex, BMI category and cancer subsite (proximal colon, distal colon and rectum) (ESM Fig. 1). In addition, we performed further analyses restricted to participants with or without previous clinical diagnosis of type 2 diabetes. These analyses showed evidence of significant associations between insulin resistance-linked hyperinsulinaemia PPSs and CRC only among people without a history of type 2 diabetes (ESM Table 7).
In secondary analyses restricted to patients with early-onset CRC (n=6228, almost 11% of the total number of CRC cases), we showed that the obesity-related insulin resistance-linked hyperinsulinaemia PPS was significantly associated with increased odds of early-onset CRC, with an adjusted OR of 1.06 (95% CI 1.04, 1.10) per 1 SD increase (ESM Table 8). The distribution of insulin resistance-linked hyperinsulinaemia PPSs was similar between early-onset and late-onset CRC (ESM Fig. 2). Comparing individuals in the highest vs lowest decile of the obesity-related PPS, the adjusted OR was 1.26 (95% CI 1.13, 1.41). The observed association remained unchanged after further adjustment for additional covariates (ESM Table 8). Analyses stratified for sex and cancer subsite revealed consistent associations (ESM Table 9), but there was evidence of effect modification by BMI (p<0.001) (Fig. 3). Among individuals with normal weight, there was no significant differences in the odds of early-onset CRC between those at the top vs bottom decile of the obesity-related PPS (adjusted OR 1.02; 95% CI 0.88, 1.18). In contrast, the association was stronger when comparing the extremes of genetic susceptibility among individuals with overweight (adjusted OR 1.27; 95% CI 1.14, 1.42), and strongest among those with obesity (adjusted OR 1.75; 95% CI 1.46, 2.11) (Fig. 3). Among the 1097 early-onset CRC cases occurring in this group of individuals with obesity, 69 cases (6.3%) were observed among those in the lowest decile of the obesity-related PPS, whereas 175 (16.0%) were among those in the highest decile.
Fig. 3.
Association between the obesity-related PPSs for individuals in various BMI categories and early-onset CRC. Adjusted ORs and 95% CIs for the association between the obesity-related PPS and early-onset CRC among individuals from three BMI categories: (a) normal weight (18.5–24.9 kg/m2), (b) overweight (25.0–29.9 kg/m2) and (c) obesity (≥30 kg/m2). The obesity-related PPS values were divided into 30 equal-sized bins; the dots are crude OR per bin (error bars indicate ±1 SE). The reference is the mean PPS of each BMI group. Early CRC onset is defined as a CRC diagnosis at less than 50 years old. The lines are model-predicted ORs (with 95% CI) adjusted for sex and the first ten ancestry-derived principal components. A multiplicative interaction term was included in the model to identify interactions between BMI and PPS (p value for the interaction term=1.60 × 10−5). Compared with the lowest decile of the PPS, individuals in the highest decile had adjusted ORs of 1.75 (95% CI 1.46, 2.11), 1.27 (95% CI 1.14, 1.42) and 1.02 (95% CI 0.88, 1.18) for the obesity (c), overweight (b) and normal weight (a) categories, respectively
Discussion
In this large-scale study of over 129,000 individuals from the GECCO and CCFR consortia, we demonstrate that diabetogenic processes characterised by insulin resistance-linked hyperinsulinaemia are associated with increased odds of CRC. In contrast, we found no evidence of significant associations for pathways related to relative insulin secretion insufficiency or insulin resistance without hyperinsulinaemia. In addition, we showed that the PPS for obesity-related hyperinsulinaemic insulin resistance was associated with early-onset CRC, particularly among individuals with elevated BMI. Taken together, these findings advance the mechanistic understanding of the relationship between type 2 diabetes and CRC, and highlight the potential for targeted risk stratification and prevention.
Our study builds on extensive epidemiological evidence demonstrating an association between type 2 diabetes and increased CRC risk [3, 5, 6, 15, 16, 24] by dissecting the contribution of specific diabetogenic processes. Previous observational studies have yielded conflicting results, with some reporting stronger associations with shorter type 2 diabetes duration, implicating early metabolic disturbances such as hyperinsulinaemia in CRC [5, 6], while others have found stronger links with longer diabetes duration, possibly reflecting cumulative metabolic damage [16]. These divergent findings probably reflect the clinical and molecular heterogeneity of type 2 diabetes, with varying degrees of residual insulin secretion and insulin sensitivity across individuals and over time [25, 26]. A recent Mendelian randomisation study supported a causal effect of hyperinsulinaemia on increased CRC risk [10]. These results are aligned with those of previous experimental studies that demonstrated that hyperinsulinaemia can enhance colorectal epithelial proliferation and tumorigenesis through adipose tissue-released factors and impaired insulin and IGF signalling [11–13]. Chronic low-grade inflammation, which frequently accompanies insulin resistance-linked hyperinsulinaemia, may represent an additional mechanism linking diabetogenic pathways to CRC risk [27], and should be considered in future mechanistic and epidemiological investigations. Our findings expand on these observations by providing suggestive evidence that insulin resistance-linked hyperinsulinaemia processes, particularly those mediated through adipose tissue mechanisms, are associated with increased odds of CRC, including early-onset disease.
A major strength of this study is the use of PPSs to investigate the diabetogenic mechanisms underlying the associations between type 2 diabetes and CRC. These scores were constructed using genetic variants that have established associations with specific metabolic phenotypes that have been validated through tissue-specific epigenomic and gene expression data, offering biologically meaningful representations of type 2 diabetes pathophysiology. This integrative approach yields functionally grounded representations of the underlying metabolic pathways implicated in type 2 diabetes pathophysiology. Our findings showing that insulin resistance-linked-hyperinsulinaemia mediated through obesity, lipodystrophy and body fat mechanisms is associated with increased odds of CRC represent an advance over previous studies showing that hyperinsulinaemia is causally associated with CRC. Of particular interest, we observed that certain well-characterised genes that are implicated in hyperinsulinaemia, such as INSR, IRS1 and IGF1R, are exclusively represented within clusters defined by insulin resistance-linked hyperinsulinaemia, which supports the biological plausibility of our findings, and provides a mechanistic context for understanding the observed associations.
A novel and clinically important aspect of this study is the link to early-onset CRC, a form of the disease with increasing incidence and poorly understood aetiology [6, 28]. Emerging evidence indicates that early-onset CRC may present unique mutational, epigenetic and metabolic features. A previous study from our consortium that included 6176 patients with early-onset CRC and 65,829 control participants identified new early-onset CRC susceptibility genes related to insulin signalling and immune/infection-related pathways [29]. Others have reported that early-onset CRC may exhibit distinct molecular features, including microsatellite instability, a CpG island methylator phenotype, and mutations in TP53 and PTEN, while BRAF mutations are less common than in older-onset CRC [30]. These molecular features probably interact with environmental and metabolic factors, manifesting in susceptible individuals. In our study, only the obesity-related hyperinsulinaemic insulin resistance PPS was associated with early-onset CRC. This association was independent of behavioural factors such as smoking and alcohol intake but was significantly modified by BMI: individuals with obesity and high genetic susceptibility had a 75% higher odds of early-onset CRC compared with those at lowest genetic risk. No evidence of associations was observed among normal-weight individuals. These findings highlight the importance of gene–BMI interactions in early-onset CRC, and suggest that genetically determined metabolic dysfunction may play a key role in early tumorigenesis in this increasingly common CRC subtype.
Although the effect sizes observed in our study were modest, with ORs ranging from 1.1 to 1.3 when comparing individuals in the top vs bottom deciles of the PPSs, these estimates are broadly consistent with the 1.2–1.5-fold increased CRC risk reported in large observational studies of type 2 diabetes [3]. Reporting effect size based on extremes of genetic risk is supported by evidence from a recent population-based prostate cancer screening programme. In that study, targeted screening of participants in the highest 10% of genetic risk identified clinically actionable prostate cancer in 55.1% (n=103) of cases, and cancer would not have been detected in 74 of these participants (71.8%) according to current prostate cancer diagnosis pathways [31]. Our study highlights the potential utility of PPSs to help identify individuals at elevated CRC risk, particularly those who do not meet current age-based screening guidelines. Younger adults with obesity and those with metabolically adverse genetic profiles may represent an under-recognised high-risk population that could benefit from targeted screening, as reflected in the recent international management guidelines initiative for early-onset CRC [32]. Although our work was not designed to inform clinical risk prediction or screening strategies, these insights contribute to a deeper understanding of the shared pathophysiology between diabetogenic processes and CRC, and highlight opportunities for eventual mechanistically informed screening and prevention strategies.
Several limitations merit consideration. First, replication in independent studies is needed to confirm the generalisability of our findings. Second, the study population was primarily of European ancestry, which may limit applicability to other populations. However, the latest multi-ancestry GWAS of CRC has shown limited population-specific heterogeneity [21]. Furthermore, the pathway-specific PPSs used in this study have shown consistent associations with cardiovascular outcomes across ancestries, suggesting that the underlying biological mechanisms may be broadly conserved [18]. Third, while the case–cohort design employed here is well suited for aetiological inference, it may limit direct applicability to the broader population. Future prospective observational studies are necessary to better assess the predictive capability of these PPSs in CRC risk. Fourth, hyperinsulinaemia itself can act as a causal driver of adiposity, suggesting a bi-directional relationship between insulin secretion and body fat. This complexity has implications for interpreting our findings, as genetic determinants that promote hyperinsulinaemia may increase CRC risk both directly, through mitogenic signalling pathways, and indirectly, through their effect on adiposity. Our stratified analyses by BMI provide some insight, showing effect maximisation in individuals with high BMI, consistent with the presence of both direct and indirect effects. However, future studies incorporating longitudinal metabolic and imaging data are essential to disentangle these pathways and clarify the relative contributions of hyperinsulinaemia and adiposity to CRC risk. Finally, we lacked data on potentially important confounders or effect modifiers, such as socioeconomic status, dietary habits, healthcare access and CRC screening behaviours (e.g. colonoscopy use), which could influence the observed associations [5]. Future studies should aim to incorporate these variables to further refine risk estimates.
In conclusion, our findings suggest that diabetogenic processes characterised by insulin resistance-linked hyperinsulinaemia are associated with increased odds of CRC, including early-onset disease. In contrast, no such associations were observed for relative insulin secretion insufficiency or insulin resistance without hyperinsulinaemia, underscoring the specificity of the implicated diabetogenic processes in CRC pathogenesis. Furthermore, we observed a synergistic interaction between genetic susceptibility and elevated BMI, underscoring the importance of gene–environment interactions in early-onset CRC. While replication in ancestrally diverse populations and further evaluation of clinical utility are needed, these results have important public health and clinical implications. They support a use of a more mechanistically informed approach to CRC prevention and early detection, particularly in individuals with obesity or adverse genetic risk profiles.
Supplementary Information
Below is the link to the electronic supplementary material.
Abbreviations
- CCFR
Colon Cancer Family Registry
- CRC
Colorectal cancer
- GECCO
Genetics and Epidemiology of Colorectal Cancer Consortium
- GWAS
Genome-wide association study
- PPS
Partitioned polygenic scores
Acknowledgements
The CCFR graciously acknowledges the generous contributions of the study participants, the dedication of study staff and the financial support from the US National Cancer Institute, without which this important registry would not exist. The authors would like to thank the study participants and staff of the Seattle Colon Cancer Family Registry and the Hormones and Colon Cancer study (CORE Studies). Acknowledgements pertaining to the specific studies included are provided in the ESM (Funding support and acknowledgements from participating studies).
Data availability
The data underlying the generation of the global polygenic score for type 2 diabetes are available from Suzuki et al [18]. Information including the procedures to obtain and access the data and codes used in this study in the GECCO and CCFR consortia is described at https://research.fredhutch.org/peters/en/genetics-and-epidemiology-of-colorectal-cancer-consortium.html. The scripts used to analyse GECCO and CCFR data presented in this manuscript are available upon request by contacting the corresponding author.
Funding
GECCO is supported by the National Cancer Institute (NCI), National Institutes of Health (NIH), US Department of Health and Human Services (U01 CA137088, R01 CA201407, R01 CA273198). Genotyping/sequencing services were provided by the Center for Inherited Disease Research (CIDR) under contract numbers HHSN268201700006I and HHSN268201200008I. This research was funded in part through NIH/NCI Cancer Center support grant P30 CA015704. The Scientific Computing Infrastructure at Fred Hutchinson Cancer Center is funded by ORIP grant S10OD028685. This work was delivered as part of the Team PROSPECT (ATC), supported by the Cancer Grand Challenges partnership funded by Cancer Research UK (CGCATF-2023/100036), the NCI (OT2CA297680), the Bowelbabe Fund for Cancer Research UK and the Institut National du Cancer. CCFR is supported in part by funding from the NCI/NIH (award U01 CA167551). Support for case ascertainment was provided in part by the Surveillance, Epidemiology and End Results (SEER) Program, US state cancer registries in Arizona, Colorado, Minnesota, North Carolina and New Hampshire, and the Victoria Cancer Registry (Australia) and Ontario Cancer Registry (Canada). The CCFR Set-1 (Illumina 1M/1M-Duo) and Set-2 (Illumina Omni1-Quad) scans were supported by NIH awards U01 CA122839 and R01 CA143237. The CCFR Set-3 (Affymetrix Axiom CORECT Set array) scan was supported by NIH awards U19 CA148107 and R01 CA81488. The CCFR Set-4 (Illumina OncoArray 600K SNP array) scan was supported by NIH award U19 CA148107 and by the CIDR, which is funded by a grant from the NIH to Johns Hopkins University (contract number HHSN268201200008I). Additional funding for the OFCCR was provided through award GL201-043 from the Ontario Research Fund, award 112746 from the Canadian Institutes of Health Research, a Cancer Risk Evaluation (CaRE) Program grant from the Canadian Cancer Society, and through generous support from the Ontario Ministry of Research and Innovation. Additional support for the CCFR was in part through NCI/NIH awards U01/U24 CA074794 and R01 CA076366. The content of this manuscript does not necessarily reflect the views or policies of the NCI, NIH or any of the collaborating centres in the CCFR, nor does mention of trade names, commercial products or organisations imply endorsement by the US Government, any cancer registry or the CCFR. JM was supported by Novo Nordisk Foundation grant NNF23SA0084103, an EFSD/Novo Nordisk Foundation Future Leaders Award (number 0094134) and the European Union (HORIZON-EIC-2023-PATHFINDERCHALLENGES-01-101161509). However, the views and opinions expressed are those of the author(s) only and do not necessarily reflect those of the European Union or European Innovation Council and SME Executive Agency (EISMEA). Neither the European Union nor the granting authority can be held responsible for these views. ATC is supported by NCI grant R01 R35CA253178 and an American Cancer Society Professorship. MS-G is supported by grants NIDDK UM1 DK078616 and 1K99DK139461-01A1. EG is funded as an American Cancer Society Clinical Research Professor (CRP-23-1014041). Further information on funding is provided in the ESM (Funding support and acknowledgements from participating studies).
Authors’ relationships and activities
JM is an Associate Editor for Diabetologia but played no role in the evaluation of this manuscript. UP was a consultant with AbbVie and her family hold individual stocks for the following companies: Amazon, Boeing Company, BioNTech, BYD Company Limited, Crowdstrike Holdings Inc., CureVac, Google/Alphabet, Microsoft Corp, MicroStrategy Inc., NVIDIA Corp and Stellantis. The remaining authors declare that there are no other relationships or activities that might bias, or be perceived to bias, this work.
Contribution statement
The study was conceptualised by MS, JM and EG. XZ validated and verified the data. The methodology was developed by JM, MS-G, UP, AIP and MS. JM wrote the original draft and visualised the findings; the draft was reviewed and edited by MS-G, XZ and MS. All authors contributed to the interpretation of data, reviewed the manuscript, and agreed to the final version of the manuscript. JM is the guarantor of this work, and, as such, had full access to all the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Mingyang Song and Jordi Merino contributed equally to this study.
References
- 1.Ong KL, Stafford LK, McLaughlin SA et al (2023) Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the Global Burden of Disease Study 2021. Lancet 402(10397):203–234. 10.1016/S0140-6736(23)01301-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Peeters PJHL, Bazelier MT, Leufkens HGM, De Vries F, De Bruin ML (2015) The risk of colorectal cancer in patients with type 2 diabetes: associations with treatment stage and obesity. Diabetes Care 38(3):405–502. 10.2337/dc14-1175 [DOI] [PubMed] [Google Scholar]
- 3.Guraya SY (2015) Association of type 2 diabetes mellitus and the risk of colorectal cancer: a meta-analysis and systematic review. World J Gastroenterol 21(19):6026–6031. 10.3748/wjg.v21.i19.6026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Dekker E, Tanis PJ, Vleugels JLA, Kasi PM, Wallace MB (2019) Colorectal cancer. Lancet 394(10207):1467–1480. 10.1016/S0140-6736(19)32319-0 [DOI] [PubMed] [Google Scholar]
- 5.Lawler T, Walts ZL, Steinwandel M et al (2023) Type 2 diabetes and colorectal cancer risk. JAMA Netw Open 6(11):e2343333. 10.1001/jamanetworkopen.2023.43333 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Samuel SM, Varghese E, Büsselberg D (2025) Complexity of insulin resistance in early-onset colorectal cancer. Cancer Cell 43(5):797–802. 10.1016/J.CCELL.2025.03.033 [DOI] [PubMed] [Google Scholar]
- 7.Gallagher EJ, LeRoith D (2020) Hyperinsulinaemia in cancer. Nat Rev Cancer 20(11):629–644. 10.1038/s41568-020-0295-5 [DOI] [PubMed] [Google Scholar]
- 8.Tuomi T, Santoro N, Caprio S, Cai M, Weng J, Groop L (2014) The many faces of diabetes: a disease with increasing heterogeneity. Lancet 383(9922):1084–1094. 10.1016/S0140-6736(13)62219-9 [DOI] [PubMed] [Google Scholar]
- 9.McCarthy MI (2017) Painting a new picture of personalised medicine for diabetes. Diabetologia 60(5):793–799. 10.1007/s00125-017-4210-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Murphy N, Song M, Papadimitriou N et al (2022) Associations between glycemic traits and colorectal cancer: a Mendelian randomization analysis. J Natl Cancer Inst 114(5):740–752. 10.1093/jnci/djac011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wang M, Yang Y, Liao Z (2020) Diabetes and cancer: epidemiological and biological links. World J Diabetes 11(6):227–238. 10.4239/wjd.v11.i6.227 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Li M, Chi X, Wang Y, Setrerrahmane S, Xie W, Xu H (2022) Trends in insulin resistance: insights into mechanisms and therapeutic strategy. Signal Transduct Target Ther 7(1):1–25. 10.1038/s41392-022-01073-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zhang AMY, Magrill J, de Winter TJJ et al (2019) Endogenous hyperinsulinemia contributes to pancreatic cancer development. Cell Metab 30(3):403–404. 10.1016/j.cmet.2019.07.003 [DOI] [PubMed] [Google Scholar]
- 14.Sevilla-González M, Smith K, Wang N et al (2025) Heterogeneous effects of genetic variants and traits associated with fasting insulin on cardiometabolic outcomes. Nat Commun 16(1):2569. 10.1038/S41467-025-57452-Y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Pang Y, Kartsonaki C, Guo Y et al (2018) Diabetes, plasma glucose and incidence of colorectal cancer in Chinese adults: a prospective study of 0.5 million people. J Epidemiol Community Health 72(10):919–925. 10.1136/jech-2018-210651 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Zaccardi F, Ling S, Brown K, Davies M, Khunti K (2023) Duration of type 2 diabetes and incidence of cancer: an observational study in England. Diabetes Care 46(11):1923–1930. 10.2337/DC23-1013 [DOI] [PubMed] [Google Scholar]
- 17.Udler MS, McCarthy MI, Florez JC, Mahajan A (2019) Genetic risk scores for diabetes diagnosis and precision medicine. Endocr Rev 40(6):1500–1520. 10.1210/er.2019-00088 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Suzuki K, Hatzikotoulas K, Southam L et al (2024) Genetic drivers of heterogeneity in type 2 diabetes pathophysiology. Nature 627(8003):347–357. 10.1038/s41586-024-07019-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Udler MS, Kim J, von Grotthuss M et al (2018) Type 2 diabetes genetic loci informed by multi-trait associations point to disease mechanisms and subtypes: a soft clustering analysis. PLoS Med 15(9):e1002654. 10.1371/journal.pmed.1002654 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Dicorpo D, Leclair J, Cole JB et al (2022) Type 2 diabetes partitioned polygenic scores associate with disease outcomes in 454,193 individuals across 13 cohorts. Diabetes Care 45(3):674–683. 10.2337/dc21-1395 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Fernandez-Rozadilla C, Timofeeva M, Chen Z et al (2022) Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and east Asian ancestries. Nat Genet 55(1):89–99. 10.1038/s41588-022-01222-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ (2015) Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4:7. 10.1186/s13742-015-0047-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Rubin DB (1987) Multiple imputation for nonresponse in surveys. Wiley, New York [Google Scholar]
- 24.Ma Y, Yang W, Song M et al (2018) Type 2 diabetes and risk of colorectal cancer in two large U.S. prospective cohorts. Br J Cancer 119(11):1436–1442. 10.1038/S41416-018-0314-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Nair ATN, Wesolowska-Andersen A, Brorsson C et al (2022) Heterogeneity in phenotype, disease progression and drug response in type 2 diabetes. Nat Med 28(5):982–988. 10.1038/s41591-022-01790-7 [DOI] [PubMed] [Google Scholar]
- 26.Schön M, Prystupa K, Mori T et al (2024) Analysis of type 2 diabetes heterogeneity with a tree-like representation: insights from the prospective German Diabetes Study and the LURIC cohort. Lancet Diabetes Endocrinol 12(2):119–131. 10.1016/S2213-8587(23)00329-7 [DOI] [PubMed] [Google Scholar]
- 27.Tsilidis KK, Branchini C, Guallar E, Helzlsouer KJ, Erlinger TP, Platz EA (2008) C-reactive protein and colorectal cancer risk: a systematic review of prospective studies. Int J Cancer 123(5):1133–1140. 10.1002/IJC.23606 [DOI] [PubMed] [Google Scholar]
- 28.Akimoto N, Ugai T, Zhong R et al (2021) Rising incidence of early-onset colorectal cancer – a call to action. Nat Rev Clin Oncol 18(4):230–243. 10.1038/S41571-020-00445-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Laskar RS, Qu C, Huyghe JR et al (2024) Genome-wide association studies and Mendelian randomization analyses provide insights into the causes of early-onset colorectal cancer. Ann Oncol 35(6):523. 10.1016/J.ANNONC.2024.02.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Ugai T, Haruki K, Harrison TA et al (2022) Molecular characteristics of early-onset colorectal cancer according to detailed anatomical locations: comparison to later-onset cases. Am J Gastroenterol 118(4):712. 10.14309/AJG.0000000000002171 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.McHugh JK, Bancroft EK, Saunders E et al (2025) Assessment of a polygenic risk score in screening for prostate cancer. N Engl J Med 392(14):1406–1417. 10.1056/NEJMOA2407934 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Cavestro GM, Mannucci A, Balaguer F et al (2023) Delphi Initiative for Early-Onset Colorectal Cancer (DIRECt) international management guidelines. Clin Gastroenterol Hepatol 21(3):581-603.e33. 10.1016/j.cgh.2022.12.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying the generation of the global polygenic score for type 2 diabetes are available from Suzuki et al [18]. Information including the procedures to obtain and access the data and codes used in this study in the GECCO and CCFR consortia is described at https://research.fredhutch.org/peters/en/genetics-and-epidemiology-of-colorectal-cancer-consortium.html. The scripts used to analyse GECCO and CCFR data presented in this manuscript are available upon request by contacting the corresponding author.





