Skip to main content
AACR Open Access logoLink to AACR Open Access
. 2026 May 5;35(9):1555–1564. doi: 10.1158/1055-9965.EPI-26-0020

The Landscape of Genomic and Socioeconomic Variables in Patients with Colorectal Cancer Based on Genetic Ancestry

Preethi Srinivasan 1,*, Sara L Bristow 1, Fernando L Mendez 1, Shifra Krinshpun 1, Adham Jurdi 1, Minetta C Liu 1, Matthew Rabinowitz 1, Jeffrey Wall 2, Carlos D Bustamante 2, Alexander G Ioannidis 3,4, Francisco M De La Vega 2,3, Breeana L Mitchell 1, Alexey Aleshin 1, Johannes G Reiter 1,*
PMCID: PMC13530983  PMID: 42084482

Abstract

Background:

Despite differences in tumor alterations across genetic ancestries, investigations of the colorectal cancer molecular landscape have used self-reported ethnicity instead of genetic ancestry.

Methods:

We used tumor and matched normal whole-exome sequencing data from 16,388 patients with stage I to IV colorectal cancer to investigate colorectal cancer’s germline and somatic molecular landscape and the potential influence of socioeconomic factors (Distressed Communities Index, DCI) across diverse genetic ancestries. Genetic ancestry determined via supervised local ancestry inference included African (AFR, N = 1,697), Native American (AMR, N = 1,291), East Asian (EAS, N = 2,247), European (EUR, N = 9,726), Levantine Middle Eastern (LME, N = 1,192), and South Asian (SAS, N = 184).

Results:

Microsatellite instability (MSI) was the most common form of hypermutation (80.8%), higher in the EUR genetic ancestry than in the AFR, AMR, and EAS genetic ancestry. Among germline findings, positive results were most common in high-penetrance genes associated with Lynch syndrome. Enrichment patterns included MLH1 (SAS) and PMS2 (AFR). There were significant differences in the frequency of driver mutations in APC, BRAF, KRAS, TP53, and PIK3CA between the EUR and other ancestry groups in both MSI and microsatellite stable tumors. Mutational signatures suggested enrichment of reactive oxygen species and POLE in AFR, colibactin in EAS, and aflatoxin and NTHL1 in SAS. DCI scores differed by ancestry (higher distress in AFR/AMR than in EUR), but driver mutation frequencies did not vary across DCI quintiles.

Conclusions:

Genetic ancestry shapes hereditary risk, tumor biology, and environmental exposures.

Impact:

These findings suggest that incorporating ancestry into screening, trials, and precision oncology may improve equity, though outcome-linked prospective studies and implementation research are warranted.

Introduction

Colorectal cancer pathogenesis is a slow process spanning years to decades that overrides the cellular mechanisms responsible for regulating cell growth and proliferation (1, 2). Both environmental and hereditary factors can increase the risk of colorectal cancer (3), with more than half of environmental risk factors considered modifiable (e.g., diet, physical activity, alcohol consumption). However, family history is considered the strongest risk factor, with an estimated 2 to 4x relative risk (RR) to the general population (4). Although colorectal cancer is estimated to be the fourth most common cancer in newly diagnosed cases and the second leading cause of mortality in 2025 in the United States (5), colorectal cancer incidence and mortality rates vary by self-reported race/ethnicity: Black people have the highest rates (incidence: 41.7 per 100,000; mortality: 17.6 per 100,000), White people have intermediate rates (incidence: 35.7 per 100,000; mortality: 13.1 per 100,000), and Hispanic/Latinos have the lowest rates (incidence: 32.5; mortality: 10.7; ref. 4). These racial and ethnic differences in overall survival are likely driven by many factors. First, race and ethnicity are social constructs, capturing lived experiences, discrimination, and environmental exposures. In fact, a recent study has shown that outcomes in patients with colorectal cancer were driven by socioeconomic status, tumor biology, and tumor location (6), suggesting that differences in lived experience and in the genetic and molecular landscape of colorectal cancer are both factors that should be considered. Further work on increasing access to screening, diagnosis, and treatment and understanding the underlying molecular drivers of tumor biology is needed to improve equity in patient outcomes.

The molecular landscape of colorectal cancer is marked with significant racial and ethnic variation in both germline and somatic variants. For example, the proportion of individuals with a pathogenic/likely pathogenic germline variant (PGV) is as low as 3.3% in Japanese patients (7) and ranges from 10% to 15% in a more heterogeneous population in the US (8, 9). Similarly, differences in somatic mutations have been observed across racial and ethnic groups (10, 11). The connection between germline risk, tumor characteristics (including somatic mutations), and precision treatment options has begun to emerge, allowing healthcare providers to tailor therapy decisions based on a patient’s unique molecular profile. However, because the molecular profile can be different based on ancestry, some therapies may be inaccessible for some populations. For example, patients with colorectal cancer having high microsatellite instability (MSI) may be eligible for immunotherapy. MSI is more common in White patients than in Black/African American patients (12). Further, PGVs in genes associated with microsatellite stability are more common in White than in Black/African American patients with colorectal cancer (13). Consequently, Black/African American patients with colorectal cancer are less likely to be eligible for immunotherapy as a potential treatment option. Another precision medicine therapy that has shown benefit for patients with colorectal cancer who carry the KRASG12C somatic variant is targeted therapy with sotorasib (14). However, the prevalence of this mutation is significantly lower in Asian populations than in Black and White populations (15). Consequently, equitable access and relevance of precision medicine may differ globally, highlighting the need for population-specific genomic data and diversity in clinical trials to ensure treatment strategies are appropriately tailored and considered.

Historically, investigations into biological differences across racial and ethnic groups have relied on self-reported ethnicity rather than genetic ancestry. Although the concordance between self-reported race/ethnicity and genetic ancestry is relatively high (16), specific populations, such as Hispanic/Latinos, are known to be an admixture of Native American (AMR), African (AFR), and European (EUR) ancestries, the contributions of each varying by geographical location (1720). In fact, recent work has shown that most individuals with admixed genetic ancestry, even those with a predominantly European background, self-identify with and report a minority racial or ethnic group (21). Recent studies have therefore used genetic ancestry algorithms to explore the somatic landscape of colorectal cancer across various genetic ancestries, reporting differences in observed actionable targets, molecular subtypes, and patient outcomes (12, 2225).

Because US-based clinical trials rarely report all major race and ethnicity subgroups or include subgroup analyses based on race and ethnicity (26), a comprehensive analysis of germline and somatic findings has not been extensively explored in non-White patients. Furthermore, socioeconomic status is a critical confounding variable that may bias observed genomic differences, as disparities in morbidity and mortality are linked to social determinants of health (27, 28). Socioeconomic variables, such as race and ethnicity, age, education level, insurance status, and income level, have been shown to affect colorectal cancer screening rates in the United States (4). Therefore, this study used tumor and matched normal whole-exome sequencing (WES) data from more than 16,000 patients to provide a comprehensive view of the germline and somatic molecular landscape of colorectal cancer, as well as the potential influence of socioeconomic factors across diverse genetic ancestries.

Materials and Methods

Study design and data abstraction

Patients eligible for inclusion in this study had to meet the following criteria: (i) diagnosed with stage I-IV colorectal cancer, (ii) had matched normal/tumor WES data available as a part of assay design for circulating tumor DNA (ctDNA) detection (Signatera, Natera, Inc.), and (iii) had testing completed between July 2019 and February 2023.

In total, 16,388 patients met these criteria and were analyzed for genetic ancestry using the available matched normal samples. De-identified patient information, including age, sex, disease stage, microsatellite status, tumor mutational burden (TMB), and the Distressed Communities Index (DCI), was abstracted and included in these analyses. Self-reported race and ethnicity was not available for these patients. All methods were carried out per Natera’s IRB-approved protocol (#23-067-GT), with informed consent waived per 45 CFR 46.116(d) and 45 CFR 46.117(c; ref. 2). This study was conducted in accordance with the Declaration of Helsinki.

Assigning genetic ancestry

Germline WES data were analyzed to determine genetic ancestry via supervised ancestry inference (Neural ADMIXTURE) using a reference panel of AFR, AMR, East Asian (EAS), EUR, Levantine Middle Eastern (LME), Oceanic (OCE), and South Asian (SAS; Supplementary Fig. S1; refs. 29, 30). Patients were then categorized into genetic ancestries based on thresholds. Patients with AMR >10% were categorized as Latino (31). Patients who had LME as the dominant genetic ancestry or who had dominant EUR and >10% LME were categorized as Ashkenazi Jewish (32). Patients with dominant EAS and OCE >5% or dominant OCE were categorized as Pacific Islander. Patients with >50% of AFR, EAS, EUR, or SAS were categorized as AFR, EAS, non-Ashkenazi Jewish EUR, or SAS. Patients who could not be categorized using these criteria were classified as the two top dominant genetic ancestries. Only patients categorized as AFR, AMR, EAS, EUR, LME, and SAS were included in the study.

Identification of PGVs in colorectal cancer–associated genes

Germline analysis was limited to a total of 22 genes with an established risk of increasing colorectal cancer risk (Supplementary Table S1; refs. 33, 34). Gene- and variant-level penetrance was determined based on the RR of cancer in carriers compared with the general population (as reported in these studies), with cancer risk (relative to the general population) estimated to be >10% as high penetrance, 2% to 10% as moderate penetrance, and <2% as low penetrance. Genes with an established, but undefined, increased cancer risk were categorized as uncertain.

Germline variants were restricted to single-nucleotide variants (SNVs) and short insertions/deletions (indels), which were identified using the GATK HaplotypeCaller (bioRxiv 10.1101/201178). The variants underwent cohort-level genotyping and annotation using a variant effect predictor (VEP, version 105, RRID:SCR_007931; ref. 35), which included population frequency annotations from gnomAD (RID:SCR_014964) and known pathogenicity reports from ClinVar (RRID:SCR_006169; ref. 36). Variants with a gnomAD allele frequency of >2% or located outside the coding region of the gene were excluded, with the threshold selected because the study population spans populations that are less well characterized in public reference datasets. Variants outside of known common variants that remained recurrent in the study cohort (>3% minor allele frequency) were excluded as they were suspected to be artifacts. Additionally, variants presumed to be due to clonal hematopoiesis were filtered and excluded based on their allele frequency in the tumor and matched normal WES data (3739).

Germline variants in colorectal cancer–associated genes were classified as benign, likely benign, variant of uncertain significance (VUS), likely pathogenic, or pathogenic based on current guidelines from the American College of Medical Genetics and the Association for Molecular Pathology using the NCBI’s ClinVar (36, 40). Variants described as pathogenic or likely pathogenic were categorized as pathogenic in this study. Given that ClinVar can have multiple, sometimes contradictory, submissions for the same variant, variants defined as pathogenic could not have any “benign” annotations. Truncating mutations in known tumor suppressor genes (classification based on OncoKB, RRID:SCR_014782) were also considered pathogenic. Finally, for VUS calls, we assessed several in silico predictors and used the BayesDel score above 0.9 to re-classify them as pathogenic (41).

Identification of somatically derived alterations

WES data from formalin-fixed paraffin-embedded tumor tissue were analyzed for SNVs, which were called using the standard tissue variant calling pipeline used for the bespoke ctDNA assay (Signatera; ref. 42). Differences in somatic alterations in known colorectal cancer driver genes reported by OncoKB (RRID:SCR_014782) were considered (43).

Purity and ploidy estimations were carried out using FACETS, an allele-specific copy number calling method (44). Microsatellite status was determined using MSIsensor (45). Consistent with previous studies (46), patients with a cutoff of 3.5 or higher were categorized as high microsatellite instability (MSI-high) status. All other patients were classified as microsatellite stable (MSS). TMB was inferred based on the somatic variants within the targeted WES regions (germline variants were excluded). Tumor-derived SNVs were analyzed for mutational signature decomposition using SigProfiler V3.4 (47). Signatures with the same underlying etiology were merged for downstream analyses. Since aging is usually the most common mutational process, our analyses used the most active mutational process other than aging, with at least 20% contribution from somatic mutations for comparisons.

Geographic and DCI analyses

Using the home address ZIP code, patients residing in the United States were assigned to a census region (West, South, Northeast, or Midwest) and a county type (suburban, small urban, moderate rural, nonmetro rural, large urban, or exurban).

Individuals were also categorized based on their communities’ overall level of economic prosperity using the census-based socioeconomic index called the DCI, which was obtained from the Economic Innovation Group (https://eig.org/download-dci-data/). The DCI score is a composite measure that uses seven indicators of economic well-being: (i) no high school diploma, (ii) housing vacancy rate, (iii) adults not working, (iv) poverty rate, (v) median income ratio, (vi) percentage change in employment, and (vii) percentage change in business establishments. Each community is ranked on each measure, then normalized into a score ranging from 0 (most prosperous) to 100 (most distressed). For this study, the 2022 edition of the DCI was used, which was derived from the US Census Bureau’s ACS 5-year Estimates covering the years 2016 to 2020 and the Business Patterns datasets for 2016 and 2020. Patients were ranked into five DCI quintiles based on their DCI score: 1, prosperous (score 0–19); 2, comfortable (score 20–39); 3, middle (score 40–59); 4, at risk (score 60–79); and 5, distressed (score 80–100).

Statistical analysis

The EUR population was used as the reference group for all one-way comparisons and statistical analyses. Standard statistical tests were used, including the Mann–Whitney U test, Fisher’s exact test, and χ2 test, depending on the data type (e.g., categorical vs. numerical). Multiple hypothesis testing was corrected using false discovery rate (FDR) correction with the Benjamini–Hochberg method. Patient demographic and clinical information, including age at diagnosis, sex, disease stage, microsatellite status, hypermutation status, and hypermutation mechanism, was summarized across genetic ancestry groups. Differences in PGV rates by penetrance and gene between EUR and other genetic ancestries were calculated using Fisher’s exact test (P ≤ 0.05, adjusted for multiple hypothesis testing). For analysis of mutational signatures, P values were adjusted for multiple hypothesis testing and an FDR threshold of 0.2 was used.

Results

Cohort description

In total, 16,388 patients were eligible for inclusion in this analysis. Among these patients, 79.2% were older than 50 years (n = 12,984), 52.9% were male (n = 8,663), and 38.3% had stage I/II disease (n = 6,273). Following genetic ancestry analysis, 16,337 (99.7%) were assigned to a single genetic ancestry group. The remaining 51 patients were assigned to more than one ancestry group (n = 37) or to an ancestry group not included in this analysis (n = 14).

Similar to the overall cohort, the majority of patients across ancestry groups were older than 50 years at the time of testing (AFR: 77.8%, LME: 81.9%, EAS: 83.3%, EUR: 79.4%, AMR: 73.3%, and SAS: 63.6%; Table 1). The proportion of males in the overall cohort was similar for the LME (53.1%), EAS (53.9%), EUR (53.3%), and AMR (53.1%) ancestral groups; the proportion was slightly lower in the AFR group (47.2%) and higher in the SAS group (67.4%).

Table 1.

Patient characteristics.

AFR AMR EAS EUR LME SAS
N 1,697 1,291 2,247 9,726 1,192 184
Age at testing, years, median (Q1, Q3) 61.4
(50.9, 69.5)
58.7
(49.3, 67.2)
64.9
(53.7, 73.3)
62.6
(51.6, 71.8)
64.9
(53.4, 73.9)
53.4
(45.8, 65)
Age at testing, years, n (%)
 <35 54 (3.2) 48 (3.7) 52 (2.3) 268 (2.8) 27 (2.3) 0
 35–50 321 (18.9) 294 (22.8) 324 (14.4) 1,728 (17.8) 186 (15.6) 60 (32.6)
 >50 1,321 (77.8) 946 (73.3) 1,871 (83.3) 7,719 (79.4) 976 (81.9) 117 (63.6)
 Unknown 1 (<0.1) 3 (0.2) 0 11 (0.1) 3 (0.3) 7 (3.8)
Sex, n (%)
 Female 891 (52.5) 605 (46.9) 1,031 (45.9) 4,523 (46.5) 554 (46.5) 60 (32.6)
 Male 801 (47.2) 685 (53.1) 1,212 (53.9) 5,181 (53.3) 633 (53.1) 124 (67.4)
 Unknown 5 (0.3) 1 (0.1) 4 (0.2) 22 (0.2) 5 (0.4) 0
Disease stage
 I 143 (8.4) 109 (8.4) 259 (12.2) 818 (8.4) 108 (9.1) 16 (8.8)
 II 496 (29.2) 364 (28.2) 628 (29.5) 2,892 (29.7) 359 (30.1) 59 (32.6)
 III 673 (39.7) 504 (39) 942 (44.3) 3,933 (40.4) 488 (40.9) 73 (40.3)
 IV 320 (18.9) 251 (19.4) 163 (7.7) 1,742 (17.9) 208 (17.5) 33 (18.2)
 Unknown 65 (3.8) 63 (4.9) 136 (6.4) 341 (3.5) 29 (2.4) 0
TMB, median (Q1, Q3) 4.5
(2, 6.7)
4.1
(1.8, 6.2)
5.4
(4.3, 7)
4.6
(2.1, 7.1)
4.4
(2.1, 6.7)
4.3
(1.9, 6.8)

Similar percentages of patients had stage I/II disease across all ancestry groups (AFR: 37.7%, AMR: 36.6%, EAS: 41.7%, EUR: 38.1%, LME: 39.2%, and SAS: 41.4%). Across all cases, we find that early-stage disease was more common in MSI tumors than in MSS tumors (P < 0.001), and when limiting to MSS tumors without PGVs in Lynch syndrome genes and MSI tumors with PGVs in Lynch syndrome genes, we find a similar difference (P < 0.001; Supplementary Fig. S2). Although differences in TMB levels were observed (EUR vs. EAS, P < 0.001; Fig. 1A), they appeared to be attributable to tumor purity (EUR vs. EAS, P < 0.001; Supplementary Fig. S2).

Figure 1.

Figure 1.

Tumor characteristics across genetic ancestries. A, TMB. For each genetic ancestry, TMB is reported in the box and whisker plots, which report the median (line), the first quartile (bottom of box), the third quartile (top of box), and 1.5 times the IQR (whiskers). Dots represent outliers. B, Hypermutation rates. Stacked bar charts report the proportion of patients with hypermutated tumors for each genetic ancestry to represent the mechanism of action. C, MSI rates by age at the time of testing. For each genetic ancestry, the MSI rate was calculated across three age groups (<35 years, 35–50 years, and >50 years) and reported as a bar chart. #, P < 0.05; *, P < 0.001.

The hypermutation rate among all patients was 16.1% (n = 2,636). Hypermutated tumors were significantly more frequent in the EUR group (17.9%) than in the AFR (15.1%, P = 0.006), LME (15.4%, P = 0.033), EAS (11.1%, P < 0.001), and AMR (13.2%, P < 0.001) groups (Fig. 1B). MSI tumors accounted for 80.8% of all hypermutated tumors (n = 2,129/2,636). POLE and other mechanisms were responsible for the remaining 7.7% and 11.5% of hypermutated tumors, respectively. POLE rates were significantly lower in the EUR group than in the AFR group (EUR: 1.1%, AFR: 2.1%, P = 0.0074). Among patients with hypermutation but without MSI or POLE, common causes of hypermutation included thiopurine chemotherapy treatment (Supplementary Fig. S3).

MSI rates were significantly higher in the EUR ancestry group (14.6%) than in the AFR (11.4%, P < 0.001), EAS (8.2%, P < 0.001), and AMR (11.4%, P = 0.001) ancestry groups, though there was no difference compared with the LME (12.6%) and SAS (16.8%) groups. Further, MSI rates varied by age at the time of testing (Fig. 1C). MSI rates decreased with increasing age for patients in the AFR, EAS, AMR, and SAS ancestry groups. For the LME and EUR groups, MSI rates were higher in patients older than 50 years than in those between 35 to 50 years. Among patients younger than 35 years, no significant differences in MSI rates were observed. MSI rates in the 35- to 50-year-old group were significantly higher in the AMR (12.9%, P = 0.04) and SAS (28.3%, P = 0.00003) groups than in the EUR (15.7%) group. Conversely, MSI rates in the >50-year-old group were significantly lower in the AFR (10.7%, P = 2 × 10−6), EAS (7.8%, P = 2.2 × 10−20), and AMR (10.4%, P = 1 × 10−5) groups than in the EUR (16%) group.

Germline findings

Overall, a PGV in a colorectal cancer–associated gene was identified in 111 (6.5%) patients of AFR ancestry, 101 (8.5%) patients of LME ancestry, 100 (4.5%) patients of EAS ancestry, 737 (7.6%) of EUR ancestry, 105 (8.1%) of AMR ancestry, and 30 (16.3%) patients of SAS ancestry. PGV rates were highest in high-penetrance genes (Fig. 2A), with significantly higher PGV rates in the AFR group (4.6%, P < 0.001) and in the SAS group (12%, P < 0.001) than in the EUR group (3.1%). Notably, genes associated with Lynch syndrome genes were most common among high-penetrance genes (AFR: 85.9%, LME: 78.4%, EAS: 78.4%, EUR: 80%, AMR: 72.7%, and SAS: 86.4%). The high rate of PGVs in the SAS ancestry group was primarily driven by MLH1 (Fig. 2B), but there was no enrichment of any individual P/LP variants in this cohort. Among non–Lynch syndrome high-penetrance genes, PGVs in APC were commonly observed across all ancestry groups (Fig. 2C).

Figure 2.

Figure 2.

Germline findings across genetic ancestries. A, PGV rates by gene penetrance. Across all genetic ancestries, the proportion of patients with a P/LP in a high-, moderate-, or low-penetrance gene was calculated. Comparisons were made with the EUR group as the reference. *, P < 0.001. B, PGV rates in high-penetrance Lynch syndrome genes. Bar plot reporting the percentage of patients, by genetic ancestry, with a PGV in a high-penetrance Lynch syndrome gene. C, PGV rates in high-penetrance genes other than Lynch syndrome. Bar plot reporting the percentage of patients, by genetic ancestry, with a PGV in a non–Lynch syndrome high-penetrance gene. Note that the y-axis scale is different for each graph. D, PGV rates in moderate- and low-penetrance genes. Bar plot reporting the percentage of patients, by genetic ancestry, with a PGV in a moderate- or low-penetrance gene. E, Colorectal cancer risk associations by gene across ancestries. Volcano plot reporting the log-based fold change of the PGV rates in individual genes (penetrance indicated in parentheses; H: high, L: low, M: moderate, U: unknown) observed in AFR, ASJ, EAS, LAT, and SAS compared with EUR (x-axis) vs. the −log P values (y-axis). PGV rates that were significantly different (FDR P < 0.05) are colored according to the genetic ancestry. Dots in the volcano plot with a log fold change less than 0 indicate that the PGV rate was higher in the EUR cohort than in the other genetic ancestry. In contrast, log fold changes greater than 0 indicate that the PGV rate was higher in the non-EUR genetic ancestry. Nonsignificant changes are represented as gray dots.

Among moderate-penetrance genes, the EUR cohort (1.12%) had a higher PGV rate than the AFR (0.35%, P < 0.001) and EAS (0.17%, P < 0.001) cohorts. The only moderate-penetrance gene with PGVs across all groups was in CHEK2 (Fig. 2D), and the higher PGV rate in the EUR cohort was driven by the common NM_007194.4:c.1100delC (p.T410Mfs*15) variant. With the exclusion of this variant, the difference in PGV rate was no longer observed (Supplementary Fig. S4). PGV rates in low-penetrance genes were significantly higher in the EUR cohort (2.3%) relative to the AFR (0.82%, P < 0.001) and EAS (0.27%, P < 0.001) cohorts. These differences were driven by the p.G393D and p.Y176C variants in MUTYH, commonly observed in the EUR population. Exclusion of this variant eliminated the differences with the AFR and EAS cohorts but revealed a significantly lower PGV rate in the EUR cohort than in the LME cohort (P < 0.001; Supplementary Fig. S4), which was driven by the founder variant APC p.I1307K.

The impact of age on PGV rates by gene penetrance was evaluated. Among all patients, we observed a strong enrichment of high-penetrance variants in the <50 years age group compared with the 50- to 65-year and >65-year-old age groups (χ2P < 0.001; Supplementary Fig. S5). However, we did not see a significant difference in the rates of moderate, low, and uncertain variants. Across ancestries, the significant differences seen in the EUR versus AFR and SAS in high-penetrance genes were driven by patients younger than 50 years old (Supplementary Fig. S5). Among moderate- and low-penetrance genes, there were no differences between AFR and EUR in any age group, suggesting that the difference observed in the overall AFR versus EUR group is due to contributions across all ages. However, the differences between EUR and EAS in moderate- and low-penetrance genes were observed in patients older than 65 and 50 years, respectively.

Enrichment of genes with PGVs was evaluated across genetic ancestries (Fig. 2E). Among high-penetrance genes, MLH1 was enriched in the SAS group compared with the EUR group (fold change: 10.9, P < 0.001) and the AMR group compared with the EUR group (fold change: 2.1, P = 0.035). Additionally, PMS2 was enriched in the AFR versus the EUR group (fold change: 2.7, P = 0.005). Among moderate-penetrance genes, CHEK2 was enriched in the EUR group compared with both the AFR group (fold change: 0.33, P = 0.032) and the EAS group (fold change: 0.12, P < 0.001). These differences were not observed when the common CHEK2 variant, NM_007194.4:c.1100delC (p.T410Mfs*15), was excluded from analysis (Supplementary Fig. S4). Patients in the EUR ancestry group had monoallelic PGVs in MUTYH (low penetrance) compared with the AFR group (fold change: 0.34, P < 0.001) and the EAS group (fold change: 0.09, P < 0.001). These differences were not observed when the common MUTYH variants, p.G393D and p.Y176C, were excluded from analysis (Supplementary Fig. S4). The low-penetrance I1307K APC allele was enriched in the LME group compared with the EUR group (fold change: 15.5, P < 0.001).

Somatic findings

When examining the impact of known drivers on colorectal cancer progression, significant differences were observed in the proportion of patients with variants in driver genes depending on whether the tumor was MSI or non-hypermutated and MSS (Fig. 3). Among patients with MSI tumors, mutations in APC were more common in those with AFR ancestry than in those with EUR ancestry (AFR: 60.1%, EUR: 46.7%, P < 0.05). The prevalence of BRAF mutations was significantly higher in the EUR group (51.5%) than in all other ancestry groups (AFR: 31.2%, P < 0.001; LME: 40.6%, P < 0.05; EAS: 41.1%, P < 0.05; AMR: 26.5%, P < 0.001; and SAS: 2.9%, P < 0.001). Conversely, the prevalence of KRAS mutations was significantly lower in EUR patients (22.5%) than in AFR (39.4%, P < 0.001), LME (32%, P < 0.05), AMR (35.8%, P < 0.05), and SAS patients (44.1%, P < 0.05).

Figure 3.

Figure 3.

Somatic landscape. Bar plots reporting percentages of patients with MSI tumors (A) or with MSS, non-hypermutated tumors (B) that have a mutation in specific driver genes (identified by tumor WES), stratified by genetic ancestry. Comparisons were made with the EUR group as the reference. #, P < 0.05; *, P < 0.001.

Among patients with MSS tumors, APC mutations were significantly less prevalent in EUR patients (76%) compared with AFR (82.8%, P < 0.001) and EAS patients (83.9%, P < 0.001). BRAF mutations were significantly more common in EUR patients (6.6%) than in AFR (4%, P < 0.05), EAS (4%, P < 0.001), and AMR patients (4.1%, P < 0.05). Mutations in KRAS were more common in the AFR ancestry than in the EUR ancestry (54.5% vs. 41.5%, P < 0.001). PIK3CA mutations were significantly more common in the EUR ancestry than in the EAS ancestry (16.5% vs. 14%, P < 0.05). Compared with the EUR ancestry (69.1%), TP53 mutations were significantly lower in the AFR ancestry (65.6%, P < 0.05) and significantly higher in the EAS ancestry (79.5%, P < 0.001). In addition to these genes, a significant difference in enrichment of mutations in MAP3K7 was observed for the EAS group when compared with the EUR group (EAS: 4.1%, EUR: 2%, P < 0.05).

Mutational signature analysis reveals enrichment of environmental exposures

Analysis of mutational signatures from tumor WES data (adjusted FDR P < 0.2) revealed that defective DNA mismatch repair signatures were enriched in the EUR cohort compared with the EAS (fold change: 0.826, P = 0.0043) and AFR cohorts (fold change: 0.858, P = 0.13; Fig. 4), which was driven by hypermutated tumors (EAS: fold change: 0.494, P < 0.001; AFR: fold change 0.517, P < 0.001). Compared with the EUR cohort, the reactive oxygen species (fold change: 2.044, P < 0.001) and POLE signatures (fold change: 1.666, P = 0.13) were enriched in the AFR cohort. The enrichment of the reactive oxygen species signature was driven by non-hypermutated tumors (fold change: 2.06, P < 0.001), whereas the enrichment of the POLE signature was driven by hypermutated tumors (fold change: 2.48, P < 0.001). Similarly, mutational signatures consistent with colibactin exposure were enriched in the EAS cohort compared with EUR (fold change: 2.125, P = 0.13). Further, hypermutated tumors in the EAS cohort were enriched with POLE relative to EUR hypermutated tumors (fold change: 2.58, P < 0.001). In the SAS cohort, mutational signatures consistent with aflatoxin exposure (fold change: 3.014, P = 0.18) and NTHL1 (fold change: 17.806, P = 0.16) were observed more frequently than in the EUR cohort.

Figure 4.

Figure 4.

Mutational signature analysis. Volcano plot reporting the log-based fold change of the frequency of mutational signature groups observed in AFR, LME, EAS, AMR, and SAS compared with EUR (x-axis) vs. the −log P values (y-axis). Mutational signatures that were significantly different (FDR P < 0.2) are colored according to the genetic ancestry. Dots in the volcano plot with a log fold change less than 0 were enriched in the EUR cohort compared with the other genetic ancestry. In contrast, log fold changes greater than 0 indicated that the mutational signature was enriched in the non-EUR genetic ancestry. Nonsignificant changes are represented as gray dots.

Additionally, the association of mutational signatures with PGVs in Lynch syndrome genes, MUTYH, and POLE/POLD1 relative to sporadic cases of colorectal cancer (i.e., no PGVs) was explored. Among patients with PGVs in Lynch syndrome, a strong enrichment of defective DNA mismatch repair signatures (MLH1: fold change: 275, P < 0.001; MSH2: fold change: 123, P < 0.001; MSH6: fold change: 140, P < 0.001, PMS2: fold change: 89, P < 0.001) and the thiopurine chemotherapy treatment signature (MLH1: fold change: 10.2, P < 0.001; MSH2: fold change: 56.5, P < 0.001; MSH6: fold change: 55, P < 0.001, PMS2: fold change: 47, P < 0.001) was observed. The reactive oxygen species (fold change: 2,570, P < 0.001) and MUTYH signatures (fold change: 1,201, P < 0.001) were enriched in patients with PGVs in MUTYH. In patients with PGVs in POLE, a slight enrichment of defective DNA mismatch repair (fold change: 10.7, P < 0.001) and thiopurine chemotherapy treatment signatures (fold change: 105, P < 0.001) was observed.

Differences by geography, county type, and DCI scores

To understand whether there were differences in geographic distribution by genetic ancestry, patients residing in the United States were grouped according to four census regions and county types. By region and county type, the distribution of patients was significantly different based on genetic ancestry (P < 0.001 for both census region and county type). Patients in the AFR, EUR, AMR, and SAS ancestry groups were most likely located in the South and Midwest. Those in the LME ancestry group were most commonly in the South, West, and Northeast. Patients with EAS ancestry were most commonly located in the South and West (Supplementary Fig. S6A). Across county types, large urban and suburban areas were the most common among AFR, LME, EAS, and SAS patients. Patients of EUR ancestry were most commonly located in large urban and moderate rural county types, whereas those of AMR ancestry were most commonly located in large urban and small urban counties (Supplementary Fig. S6B).

The relationship between genetic ancestry and socioeconomic measures (i.e., DCI, which estimates a community’s overall level of economic prosperity) was also evaluated. DCI scores were significantly different between EUR and all other genetic ancestry groups. For example, the AFR and AMR groups each had higher DCI scores than the EUR group, indicating that patients in these groups were more likely to live in areas with higher levels of economic distress (Supplementary Fig. S6C). Compared with patients older than 50 years, DCI scores were significantly lower in patients between 35 to 50 years (P < 0.001) and in patients younger than 30 years (P = 0.021; Supplementary Fig. S6D). However, DCI scores did not vary based on mutation rates in colorectal cancer driver genes nor across disease stages (Supplementary Fig. S6E and S6F).

Discussion

In this extensive study encompassing more than 16,000 patients, we comprehensively explored the germline and somatic molecular landscapes across diverse genetic ancestries. Our findings significantly contribute to the growing body of evidence that colorectal cancer tumor molecular heterogeneity is influenced by genetic ancestry (12, 23, 25, 48). Additionally, this study is the first to examine the differences in germline findings across genetic ancestry groups, as opposed to previous comparisons based on self-reported race and ethnicity focused on early-onset colorectal cancer (49, 50).

The intersection of germline and somatic mutations is critical as we consider the evolution of precision medicine so clinicians can select treatments based on a patient’s unique tumor and germline molecular profile (51). For patients with hypermutated tumors, immune checkpoint inhibitors have been shown to improve outcomes in a subset of patients (52, 53). Although increased screening and prophylactic options exist for patients who carry PGVs but do not yet have colorectal cancer, studies show that patients with colorectal cancer having PGVs may have a shorter time to a second event (54) and could benefit from germline-directed therapeutic interventions (8, 55). Furthermore, germline profiles influence the biological process of carcinogenesis, having an impact on clinical trial inclusion/exclusion criteria and predicting interventional efficacy (56). Thus, both tumor-derived alterations and germline findings can guide treatment decisions and improve patient outcomes.

When comparing results to other studies that evaluated differences across genetic ancestries, we find generally consistent results across similar analyses (12, 23, 25, 48). For example, the enrichment of APC and KRAS mutations and depletion of BRAF mutations in AFR was also observed in other studies (12, 22, 25). However, PIK3CA mutations were found to be enriched in AFR patients in other studies (12, 22), but not in this analysis, though the proportion of AFR cases (MSI or MSS) with PIK3CA mutations was higher in AFR patients than in EUR patients. This is the first report to find a significant difference in TP53 mutations in AFR patients compared with EUR patients. Among patients with AMR ancestry, we found that KRAS mutations were more common than with EUR ancestry. A study investigating tumor-derived alterations in genetic ancestries associated with the Latino population observed enrichment of different driver genes depending on the combination of AMR, EUR, and AFR descent (47). Similar to this cohort, patients with high proportions of AMR ancestry had lower rates of MSI, and the most commonly mutated genes were APC, TP53, and KRAS (though not evaluated by the MSS/MSI status). Among EAS cases, the lower proportion of BRAF and PIK3CA mutations and higher proportion of TP53 mutations have been previously reported (22). Although significant depletions of BRAF mutations have been reported previously (22, 25), no significant observations were seen in this analysis, though the proportion of BRAF mutations were lower in EAS patients than in EUR patients. Interestingly, APC mutations were enriched in EAS cases, which has not been previously observed. Amuzu and colleagues (22) found that APC mutations were depleted, whereas we observed the frequency of APC mutations was higher in SAS than EUR in MSS tumors. The differences in this observation may be because the analysis in this study separated MSI and MSS tumors. In fact, we do observe a lower, but not significant, frequency of APC mutations in MSS tumors among SAS patients. Unique to this study, we found that MSI tumors in both LME and SAS cohorts had increased proportions of KRAS mutations and decreased proportions of BRAF mutations.

This study also evaluated the prevalence of PGVs in this cohort, revealing interesting trends of PGVs in colorectal cancer risk genes with both age and mutational signatures. Consistent with other reports that hereditary colorectal cancer risk is associated with earlier disease onset, we observe similar findings. However, only high-penetrance genes in the AFR and SAS cohorts relative to EUR were found to be enriched. We further found several correlations between mutational signatures and patients with PGVs in several genes. Although many were to be expected (i.e., defective mismatch repair signatures and PGVs in Lynch syndrome genes and POLE), an enrichment of the thiopurine chemotherapy treatment was observed in patients with PGVs. This enrichment has been seen in other cohorts of patients with colorectal cancer (57), though not in the context of patients with PGVs in these genes. It has been shown that thiopurine-induced mutagenesis is accelerated by MMR deficiency (58). Further, as thiopurine treatment is common in individuals with inflammatory bowel disease (59), and both the treatment and this condition are associated with an increased risk for colorectal cancer (59, 60), it is possible that this mutational signature is correlated due to comorbidities and past treatments. However, the lack of personal history of inflammatory bowel disease and thiopurine treatment precludes any analysis on the correlation of these factors in our cohort.

To our knowledge, this is the first study to report on a comprehensive molecular landscape of colorectal cancer in patients of SAS ancestry. Despite a small sample size and limited power for statistical analysis, we observed high hypermutation rates and germline findings in Lynch syndrome–associated genes within this group, which have significant clinical implications for treatment decisions, as a greater proportion of SAS patients may be eligible for immunotherapy. Notably, analysis of mutational signatures revealed that exposure to colibactin and aflatoxin was associated with the EAS and SAS groups, respectively. Consistent with the findings of this report, a recent study demonstrated colibactin as an early colorectal cancer marker enriched in Japanese patients (61) and aflatoxin exposure to be associated with poor outcomes in South Asian regions (62, 63). Notably, aflatoxin exposure has also recently been shown to be correlated in Latino patients with gastric cancer (64). Future prospective studies that evaluate mutational signatures in the context of longitudinal exposures, including early-life exposures and migratory histories, may help elucidate the correlation between aflatoxin exposure and colorectal cancer and other gastrointestinal tumors.

Notably, we observed differences between the genetic ancestry groups and potential social determinants of health, including geography, neighborhood, and DCI score. These findings align with studies examining economic prosperity across patient-reported race and ethnicity (65, 66). Across all genetic ancestries, most patients lived in the South, with additional patients in the Midwest and Northeast. Similarly, most patients lived in large urban neighborhoods, with further enrichment in urban and moderate rural regions. A national survey indicated that the colorectal cancer burden was highest in the South (4), suggesting a potential regional factor that was not limited to a single genetic ancestry and that may have contributed to the increased burden. Regarding DCI, patients of AFR or AMR genetic ancestry were more likely to reside in areas with higher economic distress than those of EUR genetic ancestry. However, evaluating mutational frequencies of tumor-derived alterations did not reveal differences across the DCI quintiles. We observed trends in tumor-derived alterations in driver mutations and somatic signatures across genetic ancestries; however, larger studies are needed to have sufficient statistical power to identify differences adjusted for the strong association between genetic ancestry and DCI.

The findings from this study are based on a retrospective cohort drawn from commercially ordered testing. This cohort may differ from the broader colorectal cancer population, potentially introducing selection bias. For example, the EAS group had a higher percentage of patients with stage I disease than the other groups; additional studies in other populations would be needed to validate this enrichment. Furthermore, the cross-sectional design precluded linkage to treatment or longitudinal outcomes, limiting conclusions about the clinical consequences of the observed molecular differences. Future studies incorporating patient outcomes are needed to provide further insight into the clinical benefit of using the molecular profile of colorectal cancer. The analysis in this study relied exclusively on genetic ancestry, which inherently does not take the social determinants that may have shaped the observed molecular phenotypes, which may lead to biases in the results. We had attempted to use the current ZIP code–level DCI scores to evaluate socioeconomic status to evaluate these potential biases. However, this measurement is a community-level surrogate, in which a single ZIP code could encompass multiple DCI scores, cannot account for individual-level social determinants or non-US patients, and may therefore underestimate the nuance in economic disparity analyses. Further, a single DCI score at testing may not have captured the longitudinal socioeconomic factors that affect health, including migratory histories, early-life environments, and exposure windows. DCI scores are also unable to differentiate which patients were US-born versus foreign-born. Prospective studies with patient-level information regarding socioeconomic factors over time would be needed to gain a deeper understanding.

In summary, this study, encompassing more than 16,000 patients with colorectal cancer, provided a deep and comprehensive analysis of germline and somatic molecular alterations across diverse genetic ancestries. The findings affirm the molecular heterogeneity of colorectal cancer and underscore the influence of genetic ancestry on tumor biology and hereditary risk, both of which can inform treatment decisions by using precision medicine approaches. This study contributes to the evidence by characterizing the AMR genetic ancestry group and reporting on hereditary risks associated with colorectal cancer, together with somatic mutations, across genetic ancestries. Taken together, these findings demonstrate that understanding the underlying molecular landscape of colorectal cancer may help further personalize treatment decisions.

Supplementary Material

Supplemental Table S1

Supplemental Table S1. Genes included in germline analysis, by penetrance

Supplemental Figure S1

Supplemental Figure S1. Genetic ancestry by patient

Supplemental Figure S2

Supplemental Figure S2. Stage distribution based on microsatellite status across all patients and tumor purity by genetic ancestry

Supplemental Figure S3

Supplemental Figure S3. Other mechanisms of hypermutation

Supplemental Figure S4

Supplemental Figure S4. Germline findings across genetic ancestries with common EUR germline variants excluded from analysis

Supplemental Figure S5

Supplemental Figure S5. Germline findings by age across genetic ancestries

Supplemental Figure S6

Supplemental Figure S6. Geographic and socioeconomic status

Footnotes

Note: Supplementary data for this article are available at Cancer Epidemiology, Biomarkers & Prevention Online (http://cebp.aacrjournals.org/).

Contributor Information

Preethi Srinivasan, Email: prsrinivasan@natera.com.

Johannes G. Reiter, Email: jreiter@natera.com.

Data Availability

To protect the privacy and confidentiality of the patients in this study, clinical and genomic data are not made publicly available. The data generated in this study are available upon request from the corresponding authors. Any request will be reviewed within a time frame of 2 to 3 weeks by corresponding authors to verify whether the request is subject to any intellectual property or confidentiality obligations. All data shared will be de-identified.

Authors’ Disclosures

P. Srinivasan reports employment and stock ownership at Natera. S. Krinshpun reports grants and other support from Natera during the conduct of the study. A. Jurdi reports other support from Natera during the conduct of the study. M.C. Liu reports other support from Natera outside the submitted work. M. Rabinowitz reports other support from Natera during the conduct of the study, as well as other support from MyOme and Medici Therapeutics outside the submitted work. J. Wall reports other support from Natera during the conduct of the study. C.D. Bustamante reports serving as a consultant, shareholder, and board member of Galatea Bio. A.G. Ioannidis reports personal fees from Galatea Bio outside the submitted work. F.M. De La Vega reports other support from Galatea Bio outside the submitted work. J.G. Reiter reports personal fees from Natera outside the submitted work. No disclosures were reported by the other authors.

Authors’ Contributions

P. Srinivasan: Conceptualization, resources, data curation, software, formal analysis, validation, investigation, visualization, methodology, writing–review and editing. S.L. Bristow: Investigation, visualization, methodology, writing–original draft, project administration, writing-review and editing. F.L. Mendez: Data curation, software, validation, investigation, methodology, writing–review and editing. S. Krinshpun: Resources, data curation, supervision, methodology, project administration, writing–review and editing. A. Jurdi: Supervision, validation, investigation, project administration, writing–review and editing. M.C. Liu: Resources, supervision, validation, investigation, project administration, writing–review and editing. M. Rabinowitz: Resources, supervision, investigation, writing–review and editing. J. Wall: Resources, data curation, software, supervision, methodology, writing–review and editing. C.D. Bustamante: Conceptualization, resources, data curation, software, methodology, writing–review and editing. A.G. Ioannidis: Resources, data curation, software, supervision, validation, methodology, writing–review and editing. F.M. De La Vega: Conceptualization, resources, data curation, software, supervision, validation, methodology, writing–review and editing. B.L. Mitchell: Conceptualization, supervision, validation, investigation, methodology, project administration, writing–review and editing. A. Aleshin: Conceptualization, resources, supervision, methodology, project administration, writing–review and editing. J.G. Reiter: Conceptualization, resources, supervision, investigation, methodology, project administration, writing–review and editing.

References

  • 1. Li J, Ma X, Chakravarti D, Shalapour S, DePinho RA. Genetic and biological hallmarks of colorectal cancer. Genes Dev 2021;35:787–820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Pancione M, Remo A, Colantuoni V. Genetic and epigenetic events generate multiple pathways in colorectal cancer progression. Patholog Res Int 2012;2012:509348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Johnson CM, Wei C, Ensor JE, Smolenski DJ, Amos CI, Levin B, et al. Meta-analyses of colorectal cancer risk factors. Cancer Causes Control 2013;24:1207–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. American Cancer Society . Colorectal cancer facts & figures 2023–2025. Atlanta (GA): American Cancer Society. [cited 2026 Apr 28]. Available from:https://www.cancer.org/content/dam/cancer-org/research/cancer-facts-and-statistics/colorectal-cancer-facts-and-figures/colorectal-cancer-facts-and-figures-2023.pdf. [Google Scholar]
  • 5. Siegel RL, Kratzer TB, Giaquinto AN, Sung H, Jemal A. Cancer statistics, 2025. CA Cancer J Clin 2025;75:10–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Yousef M, Yousef A, Chowdhury S, Fanaeian MM, Knafl M, Peterson J, et al. Molecular, socioeconomic, and clinical factors affecting racial and ethnic disparities in colorectal cancer survival. JAMA Oncol 2024;10:1519–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Fujita M, Liu X, Iwasaki Y, Terao C, Mizukami K, Kawakami E, et al. Population-based screening for hereditary colorectal cancer variants in Japan. Clin Gastroenterol Hepatol 2022;20:2132–41.e9. [DOI] [PubMed] [Google Scholar]
  • 8. Uson PLS Jr, Riegert-Johnson D, Boardman L, Kisiel J, Mountjoy L, Patel N, et al. Germline cancer susceptibility gene testing in unselected patients with colorectal adenocarcinoma: a multicenter prospective study. Clin Gastroenterol Hepatol 2022;20:e508–28. [DOI] [PubMed] [Google Scholar]
  • 9. You YN, Borras E, Chang K, Price BA, Mork M, Chang GJ, et al. Detection of pathogenic germline variants among patients with advanced colorectal cancer undergoing tumor genomic profiling for precision medicine. Dis Colon Rectum 2019;62:429–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Holowatyj AN, Wen W, Gibbs T, Seagle HM, Keller SR, Edwards DRV, et al. Racial/ethnic and sex differences in somatic cancer gene mutations among patients with early-onset colorectal cancer. Cancer Discov 2023;13:570–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Kamran SC, Xie J, Cheung ATM, Mavura MY, Song H, Palapattu EL, et al. Tumor mutations across racial groups in a real-world data registry. JCO Precis Oncol 2021;5:1654–8. [DOI] [PubMed] [Google Scholar]
  • 12. Myer PA, Lee JK, Madison RW, Pradhan K, Newberg JY, Isasi CR, et al. The genomics of colorectal cancer in populations with African and European ancestry. Cancer Discov 2022;12:1282–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Wyatt Castillo RB, Nielsen SM, Chen E, Heald B, Ellsworth RE, Esplin ED, et al. Disparate rates of germline variants in cancer predisposition genes in African American/Black compared with non-Hispanic White individuals between 2015 and 2022. JCO Precis Oncol 2024;8:e2300715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Hong DS, Fakih MG, Strickler JH, Desai J, Durm GA, Shapiro GI, et al. KRASG12C inhibition with sotorasib in advanced solid tumors. N Engl J Med 2020;383:1207–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Nassar AH, Adib E, Kwiatkowski DJ. Distribution of KRASG12C somatic mutations across race, sex, and cancer type. N Engl J Med 2021;384:185–7. [DOI] [PubMed] [Google Scholar]
  • 16. Mersha TB, Abebe T. Self-reported race/ethnicity in the age of genomic research: its potential impact on understanding health disparities. Hum Genomics 2015;9:1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Bryc K, Durand EY, Macpherson JM, Reich D, Mountain JL. The genetic ancestry of African Americans, Latinos, and European Americans across the United States. Am J Hum Genet 2015;96:37–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. De Oliveira TC, Secolin R, Lopes-Cendes I. A review of ancestrality and admixture in Latin America and the caribbean focusing on native American and African descendant populations. Front Genet 2023;14:1091269. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Price AL, Patterson N, Yu F, Cox DR, Waliszewska A, McDonald GJ, et al. A genomewide admixture map for Latino populations. Am J Hum Genet 2007;80:1024–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Zakharia F, Basu A, Absher D, Assimes TL, Go AS, Hlatky MA, et al. Characterizing the admixed African ancestry of African Americans. Genome Biol 2009;10:R141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Banda Y, Risch N. The complex relationship of genetic ancestry with self-reported race/ethnicity. Genet Epidemiol 2025;49:e70019. [DOI] [PubMed] [Google Scholar]
  • 22. Amuzu S, Xie AX, Bai X, Pekala KR, Pickersgill NA, Ma D, et al. Meta-analysis reveals differences in somatic alterations by genetic ancestry across common cancers. Nat Genet 2025;57:2655–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Arora K, Tran TN, Kemel Y, Mehine M, Liu YL, Nandakumar S, et al. Genetic ancestry correlates with somatic differences in a real-world clinical cancer sequencing cohort. Cancer Discov 2022;12:2552–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Jiagge E, Jin DX, Newberg JY, Perea-Chamblee T, Pekala KR, Fong C, et al. Tumor sequencing of African ancestry reveals differences in clinically relevant alterations across common cancers. Cancer Cell 2023;41:1963–71.e3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Rhead B, Hein DM, Pouliot Y, Guinney J, De La Vega FM, Sanford NN. Association of genetic ancestry with molecular tumor profiles in colorectal cancer. Genome Med 2024;16:99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Loree JM, Anand S, Dasari A, Unger JM, Gothwal A, Ellis LM, et al. Disparity of race reporting and representation in clinical trials leading to cancer drug approvals from 2008 to 2018. JAMA Oncol 2019;5:e191870. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Kawachi I, Daniels N, Robinson DE. Health disparities by race and class: why both matter. Health Aff (Millwood) 2005;24:343–52. [DOI] [PubMed] [Google Scholar]
  • 28. Williams DR, Mohammed SA, Leavell J, Collins C. Race, socioeconomic status, and health: complexities, ongoing challenges, and research opportunities. Ann N Y Acad Sci 2010;1186:69–101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Mantes AD, Montserrat DM, Bustamante CD, Giró-I-Nieto X, Ioannidis AG. Neural ADMIXTURE for rapid genomic clustering. Nat Comput Sci 2023;3:621–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Parikh VN, Ioannidis AG, Jimenez-Morales D, Gorzynski JE, De Jong HN, Liu X, et al. Deconvoluting complex correlates of COVID-19 severity with a multi-omic pandemic tracking strategy. Nat Commun 2022;13:5107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Grinde KE, Brown LA, Reiner AP, Thornton TA, Browning SR. Genome-wide significance thresholds for admixture mapping studies. Am J Hum Genet 2019;104:454–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Atzmon G, Hao L, Pe’er I, Velez C, Pearlman A, Palamara PF, et al. Abraham’s children in the genome era: major Jewish diaspora populations comprise distinct genetic clusters with shared Middle Eastern Ancestry. Am J Hum Genet 2010;86:850–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Ceyhan-Birsoy O, Jayakumaran G, Kemel Y, Misyura M, Aypar U, Jairam S, et al. Diagnostic yield and clinical relevance of expanded genetic testing for cancer patients. Genome Med 2022;14:92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Natera Inc . Empower for cancer care or preventative settings. 2025[cited 2025 Dec 22]. Available from:https://www.natera.com/oncology/empower-hereditary-cancer-test/clinicians/#pg-menu-tabs.
  • 35. McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GR, Thormann A, et al. The Ensembl Variant Effect Predictor. Genome Biol 2016;17:122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Landrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, et al. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res 2014;42:D980–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Bolton KL, Ptashkin RN, Gao T, Braunstein L, Devlin SM, Kelly D, et al. Cancer therapy shapes the fitness landscape of clonal hematopoiesis. Nat Genet 2020;52:1219–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Coombs CC, Zehir A, Devlin SM, Kishtagari A, Syed A, Jonsson P, et al. Therapy-related clonal hematopoiesis in patients with non-hematologic cancers is common and associated with adverse clinical outcomes. Cell Stem Cell 2017;21:374–82.e4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Steensma DP. Clinical consequences of clonal hematopoiesis of indeterminate potential. Blood Adv 2018;2:3404–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Richards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med 2015;17:405–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Tian Y, Pesaran T, Chamberlin A, Fenwick RB, Li S, Gau CL, et al. REVEL and BayesDel outperform other in silico meta-predictors for clinical variant classification. Sci Rep 2019;9:12752. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Reinert T, Henriksen TV, Christensen E, Sharma S, Salari R, Sethi H, et al. Analysis of plasma cell-free DNA by ultradeep sequencing in patients with stages I to III colorectal cancer. JAMA Oncol 2019;5:1124–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Chakravarty D, Gao J, Phillips SM, Kundra R, Zhang H, Wang J, et al. OncoKB: a precision oncology knowledge base. JCO Precis Oncol 2017;1:1–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Shen R, Seshan VE. FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing. Nucleic Acids Res 2016;44:e131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Niu B, Ye K, Zhang Q, Lu C, Xie M, McLellan MD, et al. MSIsensor: microsatellite instability detection using paired tumor-normal sequence data. Bioinformatics 2014;30:1015–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Middha S, Zhang L, Nafa K, Jayakumaran G, Wong D, Kim HR, et al. Reliable pan-cancer microsatellite instability assessment by using targeted next-generation sequencing data. JCO Precis Oncol 2017;1:1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Alexandrov LB, Kim J, Haradhvala NJ, Huang MN, Tian Ng AW, Wu Y, et al. The repertoire of mutational signatures in human cancer. Nature 2020;578:94–101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Matejcic M, Teer JK, Hoehn HJ, Diaz DB, Shankar K, Gong J, et al. Colorectal tumors in diverse patient populations feature a spectrum of somatic mutational profiles. Cancer Res 2025;85:1928–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Dharwadkar P, Greenan G, Stoffel EM, Burstein E, Pirzadeh-Miller S, Lahiri S, et al. Racial and ethnic disparities in germline genetic testing of patients with young-onset colorectal cancer. Clin Gastroenterol Hepatol 2022;20:353–61.e3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Seagle HM, Keller SR, Tavtigian SV, Horton C, Holowatyj AN. Clinical multigene panel testing identifies racial and ethnic differences in germline pathogenic variants among patients with early-onset colorectal cancer. J Clin Oncol 2023;41:4279–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Muhar AM, Velaro AJ, Prananda AT, Nugraha SE, Halim P, Syahputra RA. Precision medicine in colorectal cancer: genomics profiling and targeted treatment. Front Pharmacol 2025;16:1532971. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Ambrosini M, Rousseau B, Manca P, Artz O, Marabelle A, André T, et al. Immune checkpoint inhibitors for POLE or POLD1 proofreading-deficient metastatic colorectal cancer. Ann Oncol 2024;35:643–55. [DOI] [PubMed] [Google Scholar]
  • 53. André T, Shiu KK, Kim TW, Jensen BV, Jensen LH, Punt C, et al. Pembrolizumab in microsatellite-instability-high advanced colorectal cancer. N Engl J Med 2020;383:2207–18. [DOI] [PubMed] [Google Scholar]
  • 54. Roila N, Pithukpakorn M, Thuwajit C, Trakarnsanga A, Chantharasamee J. Pathogenic germline variants among Thai patients with colorectal cancer: a study in Genomics Thailand Project. J Clin Oncol 2024;42(16_suppl):e22534. [Google Scholar]
  • 55. Stadler ZK, Maio A, Chakravarty D, Kemel Y, Sheehan M, Salo-Mullen E, et al. Therapeutic implications of germline testing in patients with advanced cancers. J Clin Oncol 2021;39:2698–709. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Chatrath A, Ratan A, Dutta A. Germline variants that affect tumor progression. Trends Genet 2021;37:433–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Banerjee S, Zhang X, Kuang S, Wang J, Li L, Fan G, et al. Comparative analysis of clonal evolution among patients with right- and left-sided colon and rectal cancer. iScience 2021;24:102718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Brady SW, Gout AM, Zhang J. Therapeutic and prognostic insights from the analysis of cancer mutational signatures. Trends Genet 2022;38:194–208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Sato Y, Tsujinaka S, Miura T, Kitamura Y, Suzuki H, Shibata C. Inflammatory bowel disease and colorectal cancer: epidemiology, etiology, surveillance, and management. Cancers (Basel) 2023;15:4154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Wewer MD, Letnar G, Andersen KK, Malham M, Wewer V, Seidelin JB, et al. Thiopurines and the risk of cancer in patients with inflammatory bowel disease and reference individuals without inflammatory bowel disease: a Danish nationwide cohort study (1996–2018). Clin Gastroenterol Hepatol 2025;23:1030–8. [DOI] [PubMed] [Google Scholar]
  • 61. Díaz-Gay M, Dos Santos W, Moody S, Kazachkova M, Abbasi A, Steele CD, et al. Geographic and age variations in mutational processes in colorectal cancer. Nature 2025;643:230–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Ismail A, Naeem I, Gong YY, Routledge MN, Akhtar S, Riaz M, et al. Early life exposure to dietary aflatoxins, health impact and control perspectives: a review. Trends Food Sci Techn 2021;112:212–24. [Google Scholar]
  • 63. Umar A, Bhatti HS, Honey SF. A call for aflatoxin control in Asia. CABI Agric Biosci 2023;4:27. [Google Scholar]
  • 64. Toal TW, Estrada-Florez AP, Polanco-Echeverry GM, Sahasrabudhe RM, Lott PC, Suarez-Olaya JJ, et al. Multiregional sequencing analysis reveals extensive genetic heterogeneity in gastric tumors from latinos. Cancer Res Commun 2022;2:1487–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Gandhi K, Lim E, Davis J, Chen JJ. Racial-ethnic disparities in self-reported health status among US adults adjusted for sociodemographics and multimorbidities, National Health and Nutrition Examination Survey 2011–2014. Ethn Health 2020;25:65–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Sangaramoorthy M, Shariff-Marco S, Conroy SM, Yang J, Inamdar PP, Wu AH, et al. Joint associations of race, ethnicity, and socioeconomic status with mortality in the Multiethnic Cohort Study. JAMA Netw Open 2022;5:e226370. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Table S1

Supplemental Table S1. Genes included in germline analysis, by penetrance

Supplemental Figure S1

Supplemental Figure S1. Genetic ancestry by patient

Supplemental Figure S2

Supplemental Figure S2. Stage distribution based on microsatellite status across all patients and tumor purity by genetic ancestry

Supplemental Figure S3

Supplemental Figure S3. Other mechanisms of hypermutation

Supplemental Figure S4

Supplemental Figure S4. Germline findings across genetic ancestries with common EUR germline variants excluded from analysis

Supplemental Figure S5

Supplemental Figure S5. Germline findings by age across genetic ancestries

Supplemental Figure S6

Supplemental Figure S6. Geographic and socioeconomic status

Data Availability Statement

To protect the privacy and confidentiality of the patients in this study, clinical and genomic data are not made publicly available. The data generated in this study are available upon request from the corresponding authors. Any request will be reviewed within a time frame of 2 to 3 weeks by corresponding authors to verify whether the request is subject to any intellectual property or confidentiality obligations. All data shared will be de-identified.


Articles from Cancer Epidemiology, Biomarkers & Prevention are provided here courtesy of American Association for Cancer Research

RESOURCES