Key Points
Question
Is genetic ancestry associated with colorectal cancer (CRC) burden and age-related risk patterns, and can ancestry and clinical features be integrated into a multiethnic CRC risk-prediction model?
Findings
In this cohort study of 316 624 All of Us participants, 2914 of whom developed CRC, European ancestry was significantly associated with greater overall CRC burden, whereas American admixed-Latino and East Asian ancestry were significantly associated with earlier diagnosis and higher age-specific risk. A multiethnic extreme gradient boosting prediction model demonstrated strong discrimination and calibration.
Meaning
These findings suggest that genetic ancestry is associated with meaningful differences in CRC burden and age-specific risk, and a multiethnic extreme gradient boosting model may complement CRC risk stratification and screening.
This cohort study examines associations of genetic ancestry with colorectal cancer burden, age at diagnosis, and age-specific risk and examines whether ancestry and clinical features can be integrated into a multiethnic risk-prediction model.
Abstract
Importance
Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited.
Objective
To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model.
Design, Setting, and Participants
This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome.
Exposures
Genetically inferred ancestry categories and principal components.
Main Outcomes and Measures
Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost).
Results
Among 316 624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172 327 [54.4%] of European ancestry; 191 705 female [61.2%]; 121 585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated.
Conclusions and Relevance
In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.
Introduction
Each year approximately 2 million people globally receive a diagnosis of colorectal cancer (CRC), a leading cause of cancer mortality and a major burden on individuals and health systems worldwide.1,2 CRC incidence and outcomes reflect genetics, biology, lifestyle, and social determinants of health, with community behaviors and environmental factors further contributing to disparities in risk and prognosis; understanding how these factors interact may help explain the disproportionate CRC burden observed in certain population subgroups.3
Race and ethnicity, which dominate the literature on CRC disparities, represent social constructs that capture the cultural and lived experience but imperfectly reflect genetic similarity and often exclude admixed individuals from disparities research.4,5,6 Genetic ancestry complements these measures by more precisely capturing shared genetic background and is increasingly relevant as the frequency and effect of cancer-associated genetic variants, and of tumor biology itself, differ across ancestral groups.7 Recent work has linked genetic ancestry to breast and prostate cancer risk. Other studies8,9,10,11,12 have reported that individuals of European ancestry had higher skin cancer incidence and later diagnosis than individuals of African, American admixed-Latino, East Asian, and Middle Eastern descent. In CRC, however, studies explicitly assessing genetic ancestry remain few and are typically restricted to selected subpopulations, small sample sizes, and a narrow range of ancestry comparisons.1,4,13
To address this gap, the objective of the current study was to characterize associations between genetic ancestry and CRC across complementary frameworks in a large, multiethnic cohort. Crude and covariate-adjusted associations, age at diagnosis, and attained-age time-to-event analyses were used, accounting for differential follow-up and the competing risk of death. We also developed a multiethnic CRC prediction model, evaluated its performance across ancestry groups, and identified the principal contributors to risk stratification.
Methods
Data Source, Cohort, and Variable Capture
We analyzed All of Us Research Program Controlled Tier data, freeze version 8, restricted to participants with short-read whole-genome sequencing (srWGS) and linked electronic health records (EHRs), available from July 1986 to October 2023 (median [IQR] follow-up, 133.1 [57.1-186.5] months).14 CRC was ascertained from EHR condition records (International Classification of Diseases, Ninth Revision codes 153-154; International Statistical Classification of Diseases and Related Health Problems, Tenth Revision codes C18-C20) with the earliest qualifying diagnosis as the index date; follow-up, index-age derivation, and exclusions are detailed in the eMethods in Supplement 1.
Clinical variables (CRC, prior non-CRC cancer, and type 2 diabetes) were derived from EHR diagnosis records, and cumulative aspirin exposure was derived from EHR medication records. Self-reported race and ethnicity and social-determinant variables (education, income, health insurance, and housing concern) were reported by participants in All of Us surveys. County-level Social Vulnerability Index (SVI) was assigned from linked residential geography. Genetic ancestry and principal components (PCs) were computed by All of Us from srWGS. Participants with hereditary polyposis or Lynch syndrome were excluded to focus on sporadic CRC (eFigure 1 in Supplement 1).4 This cohort study followed Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guidelines and was deemed exempt from review and the need for informed consent by the Ohio State University institutional review board.15
Outcomes
The primary outcome was any CRC. Secondary analyses classified cases as colon vs rectal cancer (colon cancer was further subcategorized as right-sided, left-sided, or other or overlapping) and as early-onset (age <50 years) or late-onset (age ≥50 years).16,17
Genetic Ancestry and Race and Ethnicity
All of Us assigned categorical genetic ancestry by projecting srWGS genotypes into a 16-component PC space built from Human Genome Diversity Project and 1000 Genomes Project reference panels and classifying by genetic similarity, yielding African, American admixed-Latino, East Asian, European, Middle Eastern, South Asian, and other (not matching a reference group).18 Because descriptive, time-to-event, and stratified analyses required 20 or more cases per cell, sparse ancestries were combined as East Asian, South Asian, and Middle Eastern. Self-reported race and ethnicity—recorded by participants in All of Us surveys—was harmonized into African, Asian, European, Hispanic, Middle Eastern, Pacific, multiracial, and other or unknown (American Indian or Alaska Native, none of these, and missing or declined responses [prefer not to answer, skip, or do not know]). Because several categories contained fewer than 20 CRC cases within ancestry strata, Asian, Middle Eastern, Pacific, and other or unknown were combined for stratified models.12 Genetic ancestry and self-reported race and ethnicity were treated as distinct constructs and modeled separately; continental ancestry probabilities were omitted from adjusted models because of collinearity with PCs, which were modeled continuously (per 1 SD).
Covariates and Missing Data
Baseline characteristics—defined as participant status at study cohort entry—included index age, sex at birth, self-reported race and ethnicity, prior cancer history, type 2 diabetes, educational attainment, annual household income, health insurance, housing concern, county SVI, body mass index (BMI; calculated as weight in kilograms divided by height in meters squared), smoking pack-years, family history of CRC and of polyps, and cumulative aspirin exposure.2,4 Missing covariates were handled by multiple imputation by chained equations (10 imputed datasets; estimates pooled by Rubin rules).19,20 To prevent information leakage in prediction models, imputation was estimated from training observations only and was applied to impute both training and tests sets. Variable coding, missingness, imputation methods, and Monte-Carlo precision are provided in the eMethods in Supplement 1.
Statistical Analysis
Baseline characteristics were summarized by genetic ancestry and compared using χ2 or Fisher exact and Kruskal-Wallis tests, as appropriate. Crude ancestry-CRC associations were estimated with 2-sided Fisher exact tests as odds ratios (ORs) with 95% CIs, comparing each ancestry with all others combined. Statistical significance was defined as P < .05. Adjusted associations used multivariable logistic regression pooled across imputations; this approach used the full cohort but did not account for differential follow-up.19,20,21 Time-to-event analyses used attained age as the timescale with delayed entry, so each participant contributed risk time only from their earliest All of Us visit (entry) to the age at CRC diagnosis, competing death, or censoring (exit); 38 995 participants without observable at-risk time were excluded from only these analyses (eMethods in Supplement 1).22,23 Treating death as a competing event, we reported 3 complementary estimates, each addressing a distinct question: cumulative incidence functions, compared across ancestry with the Gray test, cause-specific hazard models, and Fine-Gray subdistribution hazard models. All hazard models used European ancestry as the reference and were fit unadjusted and covariate-adjusted, pooled across imputations; full specifications and estimated interpretation are provided in the eMethods in Supplement 1.
CRC Prediction Model
Models classified any CRC vs no CRC over available follow-up in the overall, European, and non-European genetic ancestry cohorts, as a risk-enrichment rather than fixed-horizon tool. Penalized least absolute shrinkage and selection operator (LASSO) logistic regression and extreme gradient boosting (XGBoost) were trained on an 80% split with 5-fold cross-validation and evaluated on the held-out 20%; the primary XGBoost used raw, missingness-aware data (native missing-data handling), and imputed-data XGBoost served as a sensitivity analysis.24,25,26,27,28 We emphasized threshold-independent discrimination (receiver operating characteristic [ROC] and precision-recall [PR] area under the curve [AUC]) and calibration (Brier score, intercept, and slope); a single Youden-index operating threshold from out-of-fold training predictions was applied unchanged to the test set. The 95% CIs for all performance metrics were obtained by stratified bootstrap (2000 replicates; for multiply-imputed LASSO models, 500 replicates per imputation pooled across the 10 datasets).12 Hyperparameter grids and metric definitions are described in eMethods in Supplement 1.
All of Us Reporting and Data Analysis Tools
Per All of Us policy, cell counts less than 20 were suppressed; tests were 2-sided (α = .05) with Benjamini-Hochberg correction where applicable. Analyses were conducted February to June 2026 in the All of Us Researcher Workbench (RStudio version 2024.4.0.735 and R version 4.4.0; R Project for Statistical Computing).
Results
Descriptive Characteristics of Genetic Ancestries
Among 316 624 participants, 2914 (0.9%) developed CRC. A total of 172 327 individuals (54.4%) were of European genetic ancestry; 144 297 individuals (45.6%) were of non-European genetic ancestry, including 61 922 African (19.6%); 52 856 American admixed-Latino (16.7%); 9932 East Asian, South Asian, and Middle Eastern (3.1%); and 19 587 other (6.2%). The median (IQR) age was 56.3 (40.2-68.2) years; 191 705 participants (61.2%) were female and 121 585 (38.8%) were male (Table 1; eTables 1 and 2 in Supplement 1). Self-reported race and ethnicity was predominantly African (55 039 individuals [17.4%]), European (168 082 individuals [53.1%]), and Hispanic (58 799 individuals [18.6%]). Compared with participants with non-European genetic ancestry, participants with European genetic ancestry were older (mean [SD] age, 49.7 [16.3] vs 58.4 [17.1] years), and more frequently had a cancer history (12 510 individuals [8.7%] vs 33 341 individuals [19.3%]), health insurance (124 505 individuals [90.0%] vs 165 829 individuals [97.5%]), and a college education (40 379 individuals [29.0%] vs 98 841 individuals [58.1%]), but less often had type 2 diabetes (33 322 individuals [23.1%] vs 29 390 individuals [17.1%]). Participants of non-European genetic ancestry more often reported income in the less than $10 000 to $25 000 range (49 656 individuals [48.8%] vs 26 112 individuals [17.4%]), housing concerns (30 491 individuals [21.6%] vs 19 406 individuals [11.4%]), and residence in high-SVI counties (59 586 individuals [44.3%] vs 41 048 individuals [24.5%]) compared with participants with European genetic ancestry. In contrast, participants with European genetic ancestry were more likely than those of non-European genetic ancestry to report heavy smoking (>40 pack-years; 6506 individuals [4.6%] vs 1968 individuals [2.0%]) and a family history of CRC (18 018 individuals [17.0%] vs 5005 individuals [11.2%]) or polyps (17 429 individuals [18.2%] vs 3978 individuals [9.1%]).
Table 1. Baseline Characteristics of the Analytic All of Us Cohort, by Genetic Ancestry.
| Characteristic | Participants, No. (%) | P valuea | ||
|---|---|---|---|---|
| Overall (N = 316 624) | European (n = 172 327) | Non-European (n = 144 297) | ||
| Index age, median (IQR), y | 56.3 (40.2-68.2) | 61.3 (45.0-71.9) | 50.8 (35.9-62.3) | <.001 |
| Sex at birth | ||||
| Female | 191 705 (61.2) | 102 909 (60.2) | 88 796 (62.3) | <.001 |
| Male | 121 585 (38.8) | 67 931 (39.8) | 53 654 (37.7) | |
| Self-reported race and ethnicity | ||||
| American admixed-Latino | 11 421 (3.6) | 4325 (2.5) | 7096 (4.9) | <.001 |
| African | 55 039 (17.4) | 29 (<0.1) | 55 010 (38.1) | |
| Asian, Middle Eastern, Pacific, other, or unknownb | 23 283 (7.4) | 5008 (2.9) | 18 275 (12.7) | |
| European | 168 082 (53.1) | 161 430 (93.7) | 6652 (4.6) | |
| Hispanic | 58 799 (18.6) | 1535 (0.9) | 57 264 (39.7) | |
| Cancer history | 45 851 (14.5) | 33 341 (19.3) | 12 510 (8.7) | <.001 |
| Type 2 diabetes | 62 712 (19.8) | 29 390 (17.1) | 33 322 (23.1) | <.001 |
| Body mass index categoryc | ||||
| <18.5 | 3546 (1.1) | 1833 (1.1) | 1713 (1.2) | <.001 |
| 18.5 to <25.0 | 82 058 (26.2) | 48 482 (28.5) | 33 576 (23.5) | |
| 25.0 to <30.0 | 97 538 (31.2) | 55 415 (32.6) | 42 123 (29.5) | |
| ≥30.0 | 129 518 (41.4) | 64 185 (37.8) | 65 333 (45.8) | |
| Family history of colorectal cancer | 23 023 (15.3) | 18 018 (17.0) | 5005 (11.2) | <.001 |
| Family history of polyps | 21 407 (15.4) | 17 429 (18.2) | 3978 (9.1) | <.001 |
P values were calculated by the Kruskal-Wallis rank sum test or Pearson χ2 test.
Other or unknown refers to American Indian or Alaska Native, none of these, and missing or declined responses (prefer not to answer, skip, or do not know).
Body mass index is calculated as weight in kilograms divided by height in meters squared.
Genetic Ancestry in All of Us
Among individuals assigned a continental ancestry, the mean (SD) proportion of genome matching for assigned ancestry exceeded 97% among African (98.4% [3.7%]), American admixed-Latino (97.3% [5.4%]), East Asian (99.6% [2.0%]), European (97.6% [5.1%]), and South Asian (99.6% [2.6%%]) individuals and was 89.6% (7.2%) for Middle Eastern individuals. Concordance between genetic ancestry and self-reported race and ethnicity exceeded 90% for East Asian–Asian, European-European, and American admixed-Latino individuals, followed by South Asian–Asian (2505 individuals [89.4%]), African-African (51 004 individuals [87.2%]), and Middle Eastern–Middle Eastern (530 individuals [82.6%]) (eTable 3 in Supplement 1) groups. Among individuals in the other group, those self-reporting as European (5701 individuals [29.1%]), Hispanic (6477 individuals [33.1%]), or African (948 individuals [44.5%]) race or ethnicity had the highest genomic proportions of European (mean [SD], 61.2% [18.2%]), American admixed-Latino (mean [SD], 48.0% [19.5%]), and African (mean [SD], 44.5% [20.1%]) ancestry, respectively (eTables 4-6 in Supplement 1).
PC1 separated African ancestry, PC2 separated East Asian ancestry, PC3 separated American admixed-Latino ancestry, and PC4 separated South Asian ancestry from the other ancestries (Figure 1). PC1, PC2, PC3, and PC4 captured 97.6% of the variance among the 16 retained ancestry PCs (eTable 7 in Supplement 1).
Figure 1. Scatterplot of Genetic Ancestry Groups Across Genomic Principal Component (PC) Spaces.

Each point is a participant, colored by assigned genetic ancestry; axes are standardized PCs 1 to 4.
Genetic Ancestry–Based Comparisons for CRC Age at Diagnosis
The age at CRC diagnosis was oldest among individuals with European ancestry (median [IQR], 63.4 [53.9-71.2] years), with earlier diagnoses among American admixed-Latino (8.4 years earlier), East Asian (7.9 years earlier), African (5.0 years earlier), and other (4.7 years earlier) groups (eTable 8 and eFigure 2 in Supplement 1). The European cohort also had the oldest age at colon cancer diagnosis (median [IQR], 64.8 [55.9-72.3] years) (American admixed-Latino, 9.0 years earlier; other, 6.4 years earlier; African, 6.0 years earlier). Individuals in the American admixed-Latino group developed CRC 3.5 years earlier than the African group and 3.7 years earlier than the other group, with consistent patterns across anatomic subtypes.
Genetic Ancestry and CRC: Enrichment and Adjusted Associations
In crude (Fisher exact) comparisons, European ancestry was associated with 50% higher odds of CRC (OR, 1.50; 95% CI, 1.39-1.62), including colon (OR, 1.42; 95% CI, 1.29-1.56) and rectal (OR, 1.66; 95% CI, 1.45-1.91) cancer. In contrast, African (OR, 0.77; 95% CI, 0.70-0.85) and American admixed-Latino (OR, 0.63; 95% CI, 0.56-0.71) ancestries were associated with lower odds than all other ancestries combined (eTable 9 in Supplement 1).
On multivariable logistic regression (eTable 10 in Supplement 1), older age (adjusted OR [aOR], 1.00; 95% CI, 1.00-1.01), prior cancer history (aOR, 1.61; 95% CI, 1.48-1.77), type 2 diabetes (aOR, 1.95; 95% CI, 1.79-2.13), health insurance (aOR, 2.15; 95% CI, 1.65-2.80), smoking pack-years (aOR, 1.01; 95% CI, 1.00-1.01), and aspirin exposure (aOR, 1.06; 95% CI, 1.05-1.08) were independently associated with higher CRC odds. In contrast, female sex (aOR, 0.86; 95% CI, 0.79-0.93) and housing concern (aOR, 0.76; 95% CI, 0.67-0.86) were associated with lower odds (Table 2). Among PCs, PC4 was positively associated with CRC (aOR, 1.29; 95% CI, 1.15-1.44). In the extended model, family history of polyps (aOR, 1.52; 95% CI, 1.34-1.73) and of CRC (aOR, 2.82; 95% CI, 2.42-3.29) were associated with CRC, without altering other estimates.
Table 2. Multivariable Logistic Regression of Any Colorectal Cancer Across All, European, and Non-European Cohorts.
| Variable | All | European | Non-European | |||
|---|---|---|---|---|---|---|
| aOR (95% CI)a | P valueb | aOR (95% CI)a | P valueb | aOR (95% CI)a | P valueb | |
| Age at index (years) | 1.00 (1.00-1.01) | <.001 | 1.00 (1.00-1.01) | .05 | 1.01 (1.00-1.01) | .002 |
| Sex: female (vs male) | 0.86 (0.79-0.93) | <.001 | 0.85 (0.77-0.93) | <.001 | 0.85 (0.74-0.97) | .01 |
| Self-reported race and ethnicity | ||||||
| African | 0.82 (0.64-1.05) | .11 | NA | NA | NA | NA |
| Asian | 0.88 (0.44-1.77) | .73 | 8.48 (1.05-68.71) | .05 | 0.56 (0.35-0.91) | .02 |
| European | 1.06 (0.90-1.25) | .48 | NA | NA | 1.17 (0.88-1.55) | .28 |
| Hispanic | 0.90 (0.74-1.10) | .31 | NA | NA | NA | NA |
| Cancer history | 1.61 (1.48-1.77) | <.001 | 1.39 (1.25-1.55) | <.001 | 2.35 (2.02-2.74) | <.001 |
| Type 2 diabetes | 1.95 (1.79-2.13) | <.001 | 1.68 (1.51-1.87) | <.001 | 2.08 (1.81-2.39) | <.001 |
| Annual household income | 0.98 (0.95-1.00) | .03 | 0.97 (0.95-1.00) | .03 | NA | NA |
| Education level | 1.01 (0.98-1.06) | .46 | NA | NA | NA | NA |
| Health insurance | 2.15 (1.65-2.80) | <.001 | 1.59 (1.08-2.33) | .02 | 2.47 (1.72-3.54) | <.001 |
| Housing concern | 0.76 (0.67-0.86) | <.001 | 0.69 (0.58-0.83) | <.001 | 0.83 (0.70-0.98) | .03 |
| County Social Vulnerability Index | NA | NA | NA | NA | 1.33 (1.04-1.70) | .02 |
| Smoking pack-years | 1.01 (1.00-1.01) | <.001 | 1.01 (1.00-1.01) | .004 | 1.01 (1.00-1.01) | <.001 |
| Body mass index | 0.98 (0.98-0.99) | <.001 | NA | NA | 0.98 (0.97-0.99) | <.001 |
| Aspirin exposure (per 1 y increase) | 1.06 (1.05-1.08) | <.001 | 1.06 (1.05-1.08) | <.001 | 1.06 (1.04-1.09) | <.001 |
| Principal component per 1 SD | ||||||
| 1 | 0.89 (0.76-1.04) | .14 | NA | NA | NA | NA |
| 2 | 1.02 (0.88-1.19) | .80 | NA | NA | 0.90 (0.81-1.00) | .04 |
| 3 | 0.97 (0.84-1.11) | .62 | NA | NA | 1.02 (0.94-1.12) | .59 |
| 4 | 1.29 (1.15-1.44) | <.001 | NA | NA | 1.22 (1.10-1.35) | <.001 |
| 5 | 1.01 (0.95-1.08) | .67 | NA | NA | NA | NA |
| 6 | 0.87 (0.80-0.96) | .004 | 1.18 (0.79-1.75) | .43 | 0.93 (0.87-0.99) | .02 |
| 7 | 0.99 (0.91-1.08) | .83 | 1.17 (0.77-1.77) | .47 | NA | NA |
| 8 | 0.98 (0.92-1.05) | .63 | 1.03 (0.92-1.15) | .65 | NA | NA |
| 9 | NA | NA | 1.04 (0.98-1.10) | .24 | NA | NA |
| 11 | NA | NA | 1.05 (1.00-1.10) | .05 | NA | NA |
| 12 | 0.94 (0.90-0.98) | .001 | NA | NA | 0.91 (0.86-0.97) | .002 |
| 13 | 0.96 (0.92-1.00) | .04 | NA | NA | NA | NA |
| 14 | NA | NA | NA | NA | 1.06 (1.00-1.12) | .07 |
| Extended model | ||||||
| Family history of colon polyps | 1.52 (1.34-1.73) | <.001 | 1.52 (1.32-1.77) | <.001 | 1.45 (1.14-1.83) | .003 |
| Family history of colorectal cancer | 2.82 (2.42-3.29) | <.001 | 2.71 (2.31-3.19) | <.001 | 3.04 (2.45-3.76) | <.001 |
Abbreviations: aOR, adjusted odds ratio; NA, not applicable.
aORs with 95% CIs were calculated from models including genomic principal components, pooled across 10 imputed datasets. European-specific cells are empty where a variable was nonestimable.
P values are 2-sided.
The largest European vs non-European differences were for prior cancer history (European, aOR, 1.39; 95% CI, 1.25-1.55; non-European, aOR, 2.35; 95% CI, 2.02-2.74) and health insurance (European, aOR, 1.59; 95% CI, 1.08-2.33; non-European, aOR, 2.47; 95% CI, 1.72-3.54) (Table 2). Annual household income was inversely associated with CRC among European participants only (aOR, 0.97; 95% CI, 0.94-1.00), and county-level SVI was positively associated with CRC among non-European participants only (aOR, 1.33; 95% CI, 1.04-1.70).
Genetic Ancestry and CRC: Time-to-Event With Competing Risks
Cumulative incidence functions, treating death as a competing event, demonstrated the lowest CRC incidence among those of European ancestry through age 75 years (2.1%; 95% CI, 2.0%-2.2%) compared with African (2.4%; 95% CI, 2.1%-2.7%), American admixed-Latino (2.5%; 95% CI, 2.2%-2.9%), other (2.6%; 95% CI, 2.1%-3.0%), and East Asian (2.9%; 95% CI, 1.8%-3.7%) ancestry groups. By age 90 years, European ancestry (5.0%; 95% CI, 4.5%-5.5%) was surpassed only by American admixed-Latino ancestry (5.2%; 95% CI, 3.1%-7.2%) (Gray test P < .001) (eFigure 3 in Supplement 1). Patterns were similar for colon cancer inclusive of different colonic locations (Gray test P < .05) (eTable 11 in Supplement 1) but not rectal cancer.
On cause-specific hazard models, American admixed-Latino (hazard ratio [HR], 1.30; 95% CI, 1.14-1.47) and East Asian (HR, 1.43; 95% CI, 1.06-1.94) groups had higher risk of CRC than the European group, and these findings persisted after adjustment (American admixed-Latino, adjusted HR [aHR], 1.31; 95% CI, 1.13-1.52; East Asian, aHR, 1.61; 95% CI, 1.19-2.18) (eTable 12 in Supplement 1). On the subdistribution scale, East Asian (HR, 1.41; 95% CI, 1.07-1.87), American admixed-Latino (HR, 1.33; 95% CI, 1.18-1.49), other (HR, 1.24; 95% CI, 1.06-1.44), and African (HR, 1.13; 95% CI, 1.02-1.26) ancestry groups demonstrated higher cumulative CRC incidence than the European ancestry group, suggesting that competing mortality was associated with cumulative CRC burden across ancestry groups; adjusted estimates remained independently associated with CRC incidence (East Asian, aHR, 1.60; 95% CI, 1.18-2.18; American admixed-Latino, aHR, 1.33; 95% CI, 1.14-1.55; other, aHR, 1.23; 95% CI, 1.05-1.45) (eTable 13 in Supplement 1).
LASSO Prediction Models
Penalized LASSO models demonstrated modest discrimination. The held-out test ROC AUC was 0.733 (95% CI, 0.711-0.755) for the all-ancestry model, 0.701 (95% CI, 0.673-732) for the European model, and 0.759 (95% CI, 0.724-0.792) for the non-European model, with low PR AUCs of 0.028 (95% CI, 0.022-0.037) for the all-ancestry model, 0.030 (95% CI, 0.023-0.041) for the European model, and 0.023 (95% CI, 0.017-0.034) for the non-European model, reflecting the rarity of the outcome (eTable 14 in Supplement 1). Family history of CRC and polyps, prior cancer history, type 2 diabetes, health insurance, aspirin exposure, smoking pack-years, and BMI were selected consistently across imputations; PC4 was retained across the all and European models, whereas PC7, PC9, and PC11 were retained across non-European models (eTable 15 in Supplement 1).
XGBoost Prediction Models
The primary raw, missingness-aware XGBoost models discriminated well, with test ROC AUCs of 0.898 (95% CI, 0.882 to 0.912) for the all-ancestry model, 0.900 (95% CI, 0.882 to 0.918) for the European model, and 0.869 (95% CI, 0.839 to 0.896) for the non-European model. The PR AUCs were 0.389 (95% CI, 0.340 to 0.443) for the all-ancestry model, 0.338 (95% CI, 0.296 to 0.379) for the European model, and 0.227 (95% CI, 0.172 to 0.288) for the non-European model (Table 3; eFigure 4 in Supplement 1). The all-ancestry cohort model was selected as the primary model for its broad applicability and was well calibrated (Brier score, 0.0072; 95% CI, 0.0070 to 0.0075; calibration intercept, 0.011; 95% CI, −0.026 to 0.047; calibration slope, 1.086; 95% CI, 1.034 to 1.138).
Table 3. Performance of the Primary Raw (Nonimputed) Extreme Gradient Boosting Models for Incidence of Any Colorectal Cancera.
| Group and dataset | ROC AUC (95% CI) | PR AUC (95% CI) | PPV (precision) (95% CI) | NPV (95% CI) | Sensitivity (recall) (95% CI) | Specificity (95% CI) | F1 score (95% CI) | Accuracy (95% CI) | Threshold |
|---|---|---|---|---|---|---|---|---|---|
| All | |||||||||
| CV | 0.886 (0.878-0.894) | 0.325 (0.306-0.345) | 0.064 (0.062-0.066) | 0.997 (0.997-0.997) | 0.720 (0.702-0.738) | 0.902 (0.901-0.903) | 0.118 (0.115-0.121) | 0.901 (0.899-0.902) | 0.015 |
| Test | 0.898 (0.882-0.912) | 0.338 (0.296-0.379) | 0.065 (0.061-0.068) | 0.997 (0.997-0.998) | 0.729 (0.691-0.765) | 0.903 (0.900-0.905) | 0.119 (0.113-0.125) | 0.901 (0.899-0.903) | 0.015 |
| European | |||||||||
| CV | 0.892 (0.882-0.902) | 0.373 (0.349-0.399) | 0.065 (0.064-0.067) | 0.997 (0.997-0.997) | 0.779 (0.757-0.800) | 0.876 (0.875-0.878) | 0.121 (0.117-0.124) | 0.875 (0.874-0.877) | 0.013 |
| Test | 0.900 (0.882-0.918) | 0.389 (0.340-0.443) | 0.063 (0.059-0.066) | 0.997 (0.997-0.998) | 0.775 (0.732-0.816) | 0.877 (0.874-0.880) | 0.116 (0.110-0.123) | 0.876 (0.872-0.879) | 0.013 |
| Non-European | |||||||||
| CV | 0.855 (0.840-0.869) | 0.225 (0.195-0.255) | 0.032 (0.030-0.033) | 0.998 (0.997-0.998) | 0.705 (0.674-0.736) | 0.847 (0.845-0.849) | 0.061 (0.058-0.064) | 0.846 (0.844-0.848) | 0.007 |
| Test | 0.869 (0.839-0.896) | 0.227 (0.172-0.288) | 0.036 (0.034-0.039) | 0.998 (0.997-0.998) | 0.752 (0.693-0.807) | 0.847 (0.843-0.851) | 0.069 (0.064-0.075) | 0.846 (0.842-0.850) | 0.007 |
Abbreviations: AUC, area under the curve; CV, cross-validated; NPV, negative predictive value; PPV, positive predictive value; PR, precision recall; ROC, receiver operating characteristic.
CV and held-out test discrimination, classification, and calibration metrics for all, European, and non-European cohorts are shown, each with 95% CIs from stratified bootstrap (2000 replicates).
Applied across subgroups, the model retained strong discrimination in European (ROC AUC, 0.908; 95% CI, 0.891-0.923), non-European (ROC AUC, 0.881; 95% CI, 0.853-0.907), other (ROC AUC, 0.941; 95% CI, 0.905-0.969), American admixed-Latino (ROC AUC, 0.878; 95% CI, 0.827-0.924), and African (ROC AUC, 0.859; 95% CI, 0.818-0.898) cohorts, as well as across clinical subsets (ROC AUC; early-onset, 0.867; 95% CI, 0.827-0.904; late-onset, 0.905; 95% CI, 0.890-0.920; colon, 0.891; 95% CI, 0.872-0.909; rectal, 0.911; 95% CI, 0.889-0.930) (eTable 16 in Supplement 1). Shapley additive explanations ranking identified smoking pack-years, family history of CRC, age, prior cancer history, family history of polyps, and type 2 diabetes as the dominant contributors, with aspirin exposure, county SVI, educational attainment, BMI, and PC2 and/or PC4 also being associated with higher risk (Figure 2). In sensitivity analyses using imputed data, discrimination was substantially lower (test ROC AUC, all ancestry, 0.772; 95% CI, 0.754-0.789; European, 0.723; 95% CI, 0.697-0.747; non-European, 0.759; 95% CI, 0.728-0.789) (eTable 17 in Supplement 1).
Figure 2. Shapley Additive Explanations (SHAP) Summary Plots of the Most Influential Variables in the All-Ancestry Cohort Extreme Gradient Boosting (XGBoost) Models.

Panels show the raw, missingness-aware model (A) and the imputed sensitivity model (B) for any colorectal cancer (CRC). Each point is a participant; position reflects the SHAP contribution to predicted risk and color the feature value. BMI indicates body mass index; PC, principal component; SVI, Social Vulnerability Index.
Discussion
The pathogenesis of CRC is often a slow process, spanning years to decades, that depends on the interplay of genetic, biological, behavioral, and environmental factors; understanding these interactions may help explain its disproportionate burden across population subgroups.1,5,7,21,29,30 In this large, multiethnic All of Us cohort study, we demonstrated meaningful ancestry-related differences in CRC burden and age-related patterns while highlighting the substantial associations of demographic, clinical, and behavioral factors. Individuals of European ancestry accounted for the largest proportion of CRC cases and demonstrated enrichment in crude CRC odds, whereas individuals of African, American admixed-Latino, and East Asian descent received diagnoses at younger ages, and attained-age analyses suggested higher age-specific CRC hazard among individuals of American admixed-Latino and East Asian ancestry. Despite these patterns, established clinical, behavioral, and family-history variables dominated risk prediction. Penalized logistic regression yielded modest performance, whereas XGBoost with native missing-data handling performed substantially better, balancing strong discrimination, good calibration, and reasonable generalizability across ancestry strata.
These findings add to a limited and heterogeneous literature on genetic ancestry and CRC in which prior studies report ancestry-related differences in age at diagnosis, variant profiles, and risk but are restricted to selected populations or few comparisons.1,4,5,13,31 For example, compared with European ancestry, individuals of African ancestry have been reported to receive diagnoses at younger ages and to carry higher frequency of KRAS, APC, and PIK3CA variants.5,13 In the current cohort, individuals of European ancestry had higher crude odds of receiving a CRC diagnosis, yet African, American admixed-Latino, and East Asian individuals received diagnoses, earlier and American admixed-Latino and East Asian individuals had higher age-specific CRC hazards than European individuals. Therefore, cumulative burden and age-specific risk did not follow the same pattern: European populations may accumulate more cases later in life owing to a larger older at-risk population, whereas some non-European groups, particularly American admixed-Latino, may experience earlier onset or higher relative risk at younger attained ages, underscoring the importance of evaluating both cumulative burden and attained-age risk.22,23 These divergent patterns also demonstrate why genetic ancestry, rather than self-reported race and ethnicity alone, merits assessment relative to CRC risk.
The value of assessing genetic ancestry alongside self-reported race and ethnicity lies in its capacity to capture differences in tumor biology and hereditary risk that race does not define. Srinivasan et al7 reported that genetic ancestry influences hereditary risk, tumor biology, and environmental exposures with driver-variant frequencies differing across ancestries. Specifically, KRAS is more common in American admixed-Latino than European populations, and higher frequencies of TP53 and lower frequencies of BRAF are found among East Asian populations. The specific combination of ancestry proportions within a population can shape the enrichment of CRC driver genes.7,32 Genetic ancestry can also characterize admixed individuals, who are often excluded from disparity research. In the current cohort, individuals in the other genetic ancestry group most commonly self-identified as Hispanic (33.1%), followed by European (29.1%); among both African and European ancestries, the second-highest genomic proportion was consistently American admixed-Latino, reflecting geographic admixture. Of note, the previously described race and ethnicity-related CRC incidence—highest among Black, then White, then Hispanic or Latino individuals—was not mirrored by our ancestry-based findings, underscoring that ancestry and social experience are interconnected but nonequivalent constructs that may be important relative to equity in outcomes and trial enrollment.5,3,13,21,33,34,35 These signals are novel and set the stage for further work assessing genetic ancestry and somatic variant simultaneously in relation to CRC outcomes.7,32 Interestingly, despite ancestry signals being important, prediction was driven primarily by nonancestral factors.
Prediction analyses confirmed that established clinical and family-history variables contributed most strongly; LASSO selection frequency and Shapley Additive Explanations rankings consistently prioritized family history of CRC and of polyps, smoking burden, age, prior cancer history, type 2 diabetes, BMI, and aspirin exposure—supporting biological plausibility—whereas genomic PCs and self-reported race and ethnicity contributed more modestly.12,36 A global genotype-principal component analysis framework incorporated genetic information without multiple ancestry indicator variables and allowed inclusion of admixed individuals. Although penalized logistic regression provided only modest discrimination, XGBoost captured the CRC signal far more effectively because it better accommodated nonlinear relationships, interactions, multicollinearity, and incomplete real-world data with several top continuous variables, demonstrating nonlinear associations that may partly explain its superiority (eFigure 5 in Supplement 1).12,37,38 Certain associations nonetheless warrant caution, such as the positive aspirin-CRC association, which likely reflects confounding by indication rather than harm, and socioeconomic signals, such as health insurance, which more plausibly reflect access and detection rather than causal effects.2,3,12,39,40,41,42,43,44 Although PR AUC remained modest in absolute terms, the model substantially improved over baseline prevalence in this highly imbalanced dataset, supporting its use not as a stand-alone diagnostic but as a supplementary risk-enrichment or triage tool to identify individuals who may benefit from earlier or more intensive screening.4,45,46,47,48
Limitations
These findings should be interpreted in light of several limitations. Although uniquely diverse, All of Us was not population based, and restriction to participants with whole-genome sequencing and linked EHR data may have introduced selection bias. Because some covariates were survey derived and occasionally measured near the time of diagnosis, temporality could not always be ensured, introducing possible reverse causation.4 Although using 10 imputed datasets improved precision, the fraction of missing information remained high for the most incompletely captured variables, so associations and selection involving these variables should be interpreted cautiously. The prediction models classified any CRC vs no CRC over variable-length rather than fixed follow-up; differential follow-up and incomplete EHR ascertainment may have also misclassified some future cases as noncases and biased performance toward the null. The binary framing did not fully capture the time-dependent nature of CRC risk—limitations that were partly addressed through the complementary time-to-event and competing risks analyses. Furthermore, locus-specific ancestry was not assessed, while molecular tumor characteristics, and tumor stage, which could provide more granular insight, were unavailable.
Conclusions
In this cohort study of 316 624 All of Us participants, genetic ancestry was associated with meaningful differences in CRC burden, age at diagnosis, and age-specific risk, although estimates were driven primarily by established clinical, behavioral, and family-history factors. The pooled multiethnic XGBoost prediction model achieved strong performance and may serve as a useful adjunct for risk enrichment within existing CRC screening strategies.
eMethods.
eTable 1. Baseline characteristics of the analytic cohort across all genetic ancestry groups
eTable 2. Social determinants of health and smoking characteristics, by EUR vs non-EUR genetic ancestry
eTable 3. Concordance between genetic ancestry and self-reported race/ethnicity
eTable 4. Continental genomic ancestry composition within the OTH genetic ancestry group, by self-reported race/ethnicity
eTable 5. Genomic ancestry composition among participants with lower-confidence ancestry assignment, by genetic ancestry and self-reported race/ethnicity
eTable 6. Mean continental ancestry composition across genetic ancestry and self-reported race/ethnicity groups
eTable 7. Variance explained by genomic principal components
eTable 8. Age at CRC diagnosis across genetic ancestry groups, with pairwise comparisons
eTable 9. Genetic ancestry enrichment for colorectal, colon, and rectal cancer (Fisher exact tests vs all other ancestries)
eTable 10. Univariable logistic regression of any CRC across All, EUR, and non-EUR cohorts
eTable 11. Cumulative incidence of CRC by attained age across genetic ancestry groups
eTable 12. Attained-age cause-specific hazard models of CRC by genetic ancestry (EUR reference)
eTable 13. Fine-Gray subdistribution hazard models of CRC by genetic ancestry (EUR reference)
eTable 14. Mean performance of LASSO models for predicting any CRC across 10 imputed datasets
eTable 15. Variable selection frequency and coefficient direction across imputed LASSO models
eTable 16. Performance of the primary multiethnic XGBoost model across ancestry and clinical subgroups
eTable 17. Performance of imputed XGBoost models for predicting any CRC across All, EUR, and non-EUR cohorts
eFigure 1. Flowchart of cohort construction and exclusions for the analytic study population
eFigure 2. Age at diagnosis by genetic ancestry across CRC subtypes
eFigure 3. Cumulative incidence of CRC by attained age across genetic ancestry groups
eFigure 4. SHAP summary plots comparing feature importance in the raw EUR and non-EUR XGBoost models
eFigure 5. Restricted cubic spline associations of continuous predictors with CRC risk
eReferences
Data Sharing Statement
References
- 1.Fernandez-Rozadilla C, Timofeeva M, Chen Z, et al. Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and East Asian ancestries. Nat Genet. 2023;55(1):89-99. doi: 10.1038/s41588-022-01222-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Nafisi S, Støer NC, Veierød MB, et al. Low-dose aspirin and prevention of colorectal cancer: evidence from a nationwide registry-based cohort in Norway. Am J Gastroenterol. 2024;119(7):1402-1411. doi: 10.14309/ajg.0000000000002695 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Eng C, Holowatyj AN. Colorectal cancer genomics by genetic ancestry. Cancer Discov. 2022;12(5):1187-1188. doi: 10.1158/2159-8290.CD-22-0217 [DOI] [PubMed] [Google Scholar]
- 4.Pérez-Mayoral J, Soto-Salgado M, Shah E, et al. Association of genetic ancestry with colorectal tumor location in Puerto Rican Latinos. Hum Genomics. 2019;13(1):12. doi: 10.1186/s40246-019-0196-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Rhead B, Hein DM, Pouliot Y, Guinney J, De La Vega FM, Sanford NN. Association of genetic ancestry with molecular tumor profiles in colorectal cancer. Genome Med. 2024;16(1):99. doi: 10.1186/s13073-024-01373-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Gravlee CC. How race becomes biology: embodiment of social inequality. Am J Phys Anthropol. 2009;139(1):47-57. doi: 10.1002/ajpa.20983 [DOI] [PubMed] [Google Scholar]
- 7.Srinivasan P, Bristow SL, Mendez FL, et al. The landscape of genomic and socioeconomic variables in patients with colorectal cancer based on genetic ancestry. Cancer Epidemiol Biomarkers Prev. Published online May 5, 2026. doi: 10.1158/1055-9965.EPI-26-0020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Al-Alem U, Rauscher G, Shah E, et al. Association of genetic ancestry with breast cancer in ethnically diverse women from Chicago. PLoS One. 2014;9(11):e112916. doi: 10.1371/journal.pone.0112916 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ricks-Santi LJ, Apprey V, Mason T, et al. Identification of genetic risk associated with prostate cancer using ancestry informative markers. Prostate Cancer Prostatic Dis. 2012;15(4):359-364. doi: 10.1038/pcan.2012.19 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Fejerman L, Stern MC, Ziv E, et al. Genetic ancestry modifies the association between genetic risk variants and breast cancer risk among Hispanic and non-Hispanic white women. Carcinogenesis. 2013;34(8):1787-1793. doi: 10.1093/carcin/bgt110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Irizarry-Ramírez M, Kittles RA, Wang X, et al. Genetic ancestry and prostate cancer susceptibility SNPs in Puerto Rican and African American men. Prostate. 2017;77(10):1118-1127. doi: 10.1002/pros.23368 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.D’Antonio M, Gonzalez Rivera WG, Greenes RA, Gymrek M, Frazer KA. A highly accurate risk factor-based XGBoost multiethnic model for identifying patients with skin cancer. Nat Commun. 2025;16(1):9542. doi: 10.1038/s41467-025-64556-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Myer PA, Lee JK, Madison RW, et al. The genomics of colorectal cancer in populations with African and European ancestry. Cancer Discov. 2022;12(5):1282-1293. doi: 10.1158/2159-8290.CD-21-0813 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Denny JC, Rutter JL, Goldstein DB, et al. ; All of Us Research Program Investigators . The “All of Us” research program. N Engl J Med. 2019;381(7):668-676. doi: 10.1056/NEJMsr1809937 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Vandenbroucke JP, von Elm E, Altman DG, et al. ; STROBE Initiative . Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): explanation and elaboration. PLoS Med. 2007;4(10):e297. doi: 10.1371/journal.pmed.0040297 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Baran B, Mert Ozupek N, Yerli Tetik N, Acar E, Bekcioglu O, Baskin Y. Difference between left-sided and right-sided colorectal cancer: a focused review of literature. Gastroenterology Res. 2018;11(4):264-273. doi: 10.14740/gr1062w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Low EE, Demb J, Liu L, et al. Risk factors for early-onset colorectal cancer. Gastroenterology. 2020;159(2):492-501.e7. doi: 10.1053/j.gastro.2020.01.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.All of Us Research Program . Genomic research data quality report: All of Us curated data repository (CDR) release C2024Q3R3. February 3, 2025. Accessed July 1, 2026. https://support.researchallofus.org/hc/en-us/articles/29390274413716-All-of-Us-Genomic-Quality-Report-Archived-C2024Q3R9-v8
- 19.van Buuren S, Groothuis-Oudshoorn K. mice: Multivariate imputation by chained equations in R. J Stat Software. 2011;45(3):1-67. doi: 10.18637/jss.v045.i03 [DOI] [Google Scholar]
- 20.Yuan YC. Multiple imputation for missing data: concepts and new development. January 2005. Accessed July 1, 2026. https://www.researchgate.net/publication/228574397_Multiple_Imputation_for_Missing_Data_Concepts_and_New_Development
- 21.Hein DM, Deng W, Bleile M, et al. Racial and ethnic differences in genomic profiling of early onset colorectal cancer. J Natl Cancer Inst. 2022;114(5):775-778. doi: 10.1093/jnci/djac014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Cain KC, Harlow SD, Little RJ, et al. Bias due to left truncation and left censoring in longitudinal studies of developmental and disease processes. Am J Epidemiol. 2011;173(9):1078-1084. doi: 10.1093/aje/kwq481 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Griffin BA, Anderson GL, Shih RA, Whitsel EA. Use of alternative time scales in Cox proportional hazard models: implications for time-varying environmental exposures. Stat Med. 2012;31(27):3320-3327. doi: 10.1002/sim.5347 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Liu L, Hou J; CKM-AD/ADRD Risk Prediction Study Group . Lasso and XGBoost-enabled prediction models for sensory dysfunction, biological age, and APOE genotype in cognitive decline risk assessment. Alzheimer Dement. 2025;21(S2):e107762. doi: 10.1002/alz70856_107762 [DOI] [Google Scholar]
- 25.Hu P, Liu Y, Li Y, et al. A comparison of LASSO regression and tree-based models for delayed cerebral ischemia in elderly patients with subarachnoid hemorrhage. Front Neurol. 2022;13:791547. doi: 10.3389/fneur.2022.791547 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Schindele A, Krebold A, Heiss U, et al. Interpretable machine learning for thyroid cancer recurrence prediction: leveraging XGBoost and SHAP analysis. Eur J Radiol. 2025;186:112049. doi: 10.1016/j.ejrad.2025.112049 [DOI] [PubMed] [Google Scholar]
- 27.Dong B, Zhang H, Duan Y, Yao S, Chen Y, Zhang C. Development of a machine learning-based model to predict prognosis of alpha-fetoprotein-positive hepatocellular carcinoma. J Transl Med. 2024;22(1):455. doi: 10.1186/s12967-024-05203-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: KDD ’16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery; 2016. doi: 10.1145/2939672.2939785 [DOI] [Google Scholar]
- 29.Keku TO, Dulal S, Deveaux A, Jovov B, Han X. The gastrointestinal microbiota and colorectal cancer. Am J Physiol Gastrointest Liver Physiol. 2015;308(5):G351-G363. doi: 10.1152/ajpgi.00360.2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Tuomisto AE, Mäkinen MJ, Väyrynen JP. Systemic inflammation in colorectal cancer: underlying factors, effects, and prognostic significance. World J Gastroenterol. 2019;25(31):4383-4404. doi: 10.3748/wjg.v25.i31.4383 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Hernandez-Suarez G, Sanabria MC, Serrano M, et al. Genetic ancestry is associated with colorectal adenomas and adenocarcinomas in Latino populations. Eur J Hum Genet. 2014;22(10):1208-1216. doi: 10.1038/ejhg.2013.310 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Banda Y, Risch N. The complex relationship of genetic ancestry with self-reported race/ethnicity. Genet Epidemiol. 2025;49(8):e70019. doi: 10.1002/gepi.70019 [DOI] [PubMed] [Google Scholar]
- 33.Holowatyj AN, Wen W, Gibbs T, et al. Racial/ethnic and sex differences in somatic cancer gene mutations among patients with early-onset colorectal cancer. Cancer Discov. 2023;13(3):570-579. doi: 10.1158/2159-8290.CD-22-0764 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Fei-Zhang DJ, Bentrem DJ, Wayne JD, Hou L, Fei P, Pawlik TM. Associations of social vulnerability and race-ethnicity with gastrointestinal cancers in the United States. Cancer Med. 2025;14(5):e70591. doi: 10.1002/cam4.70591 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Florez JC, Price AL, Campbell D, et al. Strong association of socioeconomic status with genetic ancestry in Latinos: implications for admixture studies of type 2 diabetes. Diabetologia. 2009;52(8):1528-1536. doi: 10.1007/s00125-009-1412-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Amariuta T, Ishigaki K, Sugishita H, et al. Improving the trans-ancestry portability of polygenic risk scores by prioritizing variants in predicted cell-type-specific regulatory elements. Nat Genet. 2020;52(12):1346-1354. doi: 10.1038/s41588-020-00740-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Montomoli J, Romeo L, Moccia S, et al. ; RISC-19-ICU Investigators . Machine learning using the extreme gradient boosting (XGBoost) algorithm predicts 5-day delta of SOFA score at ICU admission in COVID-19 patients. J Intensive Med. 2021;1(2):110-116. doi: 10.1016/j.jointm.2021.09.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Li J, Luo Y, Dong M, et al. Tree-based risk factor identification and stroke level prediction in stroke cohort study. Biomed Res Int. 2023;2023:7352191. doi: 10.1155/2023/7352191 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Orchard SG, Polekhina G, Zalcberg J, et al. Cancer incidence and mortality with aspirin in older adults: follow-up of the ASPREE Trial. JAMA Oncol. 2026;12(3):285-294. doi: 10.1001/jamaoncol.2025.6196 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Dwan K, Altman DG, Arnaiz JA, et al. Systematic review of the empirical evidence of study publication bias and outcome reporting bias. PLoS One. 2008;3(8):e3081. doi: 10.1371/journal.pone.0003081 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Pearson-Stuttard J, Papadimitriou N, Markozannes G, et al. Type 2 diabetes and cancer: an umbrella review of observational and mendelian randomization studies. Cancer Epidemiol Biomarkers Prev. 2021;30(6):1218-1228. doi: 10.1158/1055-9965.EPI-20-1245 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Holowatyj AN, Ruterbusch JJ, Rozek LS, Cote ML, Stoffel EM. Racial/ethnic disparities in survival among patients with young-onset colorectal cancer. J Clin Oncol. 2016;34(18):2148-2156. doi: 10.1200/JCO.2015.65.0994 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Khalil M, Woldesenbet S, Shaw S, et al. Impact of neighborhood socioeconomic trajectories on gastrointestinal cancer care: a SEER-Medicare analysis. Ann Surg Oncol. 2025;32(10):7587-7595. doi: 10.1245/s10434-025-17764-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Fan Q, Nogueira L, Yabroff KR, Hussaini SMQ, Pollack CE. Housing and cancer care and outcomes: a systematic review. J Natl Cancer Inst. 2022;114(12):1601-1618. doi: 10.1093/jnci/djac173 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi: 10.1371/journal.pone.0118432 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Shan R, Li X, Chen J, et al. Interpretable machine learning to predict the malignancy risk of follicular thyroid neoplasms in extremely unbalanced data: retrospective cohort study and literature review. JMIR Cancer. 2025;11:e66269. doi: 10.2196/66269 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Richardson E, Trevizani R, Greenbaum JA, Carter H, Nielsen M, Peters B. The receiver operating characteristic curve accurately assesses imbalanced datasets. Patterns (N Y). 2024;5(6):100994. doi: 10.1016/j.patter.2024.100994 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Augustus GJ, Ellis NA. Colorectal cancer disparity in African Americans: risk factors and carcinogenic mechanisms. Am J Pathol. 2018;188(2):291-303. doi: 10.1016/j.ajpath.2017.07.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
eMethods.
eTable 1. Baseline characteristics of the analytic cohort across all genetic ancestry groups
eTable 2. Social determinants of health and smoking characteristics, by EUR vs non-EUR genetic ancestry
eTable 3. Concordance between genetic ancestry and self-reported race/ethnicity
eTable 4. Continental genomic ancestry composition within the OTH genetic ancestry group, by self-reported race/ethnicity
eTable 5. Genomic ancestry composition among participants with lower-confidence ancestry assignment, by genetic ancestry and self-reported race/ethnicity
eTable 6. Mean continental ancestry composition across genetic ancestry and self-reported race/ethnicity groups
eTable 7. Variance explained by genomic principal components
eTable 8. Age at CRC diagnosis across genetic ancestry groups, with pairwise comparisons
eTable 9. Genetic ancestry enrichment for colorectal, colon, and rectal cancer (Fisher exact tests vs all other ancestries)
eTable 10. Univariable logistic regression of any CRC across All, EUR, and non-EUR cohorts
eTable 11. Cumulative incidence of CRC by attained age across genetic ancestry groups
eTable 12. Attained-age cause-specific hazard models of CRC by genetic ancestry (EUR reference)
eTable 13. Fine-Gray subdistribution hazard models of CRC by genetic ancestry (EUR reference)
eTable 14. Mean performance of LASSO models for predicting any CRC across 10 imputed datasets
eTable 15. Variable selection frequency and coefficient direction across imputed LASSO models
eTable 16. Performance of the primary multiethnic XGBoost model across ancestry and clinical subgroups
eTable 17. Performance of imputed XGBoost models for predicting any CRC across All, EUR, and non-EUR cohorts
eFigure 1. Flowchart of cohort construction and exclusions for the analytic study population
eFigure 2. Age at diagnosis by genetic ancestry across CRC subtypes
eFigure 3. Cumulative incidence of CRC by attained age across genetic ancestry groups
eFigure 4. SHAP summary plots comparing feature importance in the raw EUR and non-EUR XGBoost models
eFigure 5. Restricted cubic spline associations of continuous predictors with CRC risk
eReferences
Data Sharing Statement
