Abstract
Overall survival (OS) of colorectal cancer (CRC) patients remains suboptimal, especially in advanced disease. This study aimed to construct and validate explainable machine learning (ML) models using routine blood indices for accurate CRC prognosis across multicenter cohorts. The training cohort included 850 CRC patients (demographic and routine blood data) from Union; validation cohorts were 403 patients (Hefei) and 217 (Shihezi). Seven time-to-event models and SHapley Additive exPlanation (SHAP) (for interpretation) were used. Among the evaluated models, the random survival forest (RSF) algorithm demonstrated superior predictive performance. RSF algorithm demonstrated high discriminatory performance in the Union test cohort with AUCs of 0.768, 0.775, and 0.731 for 1-year, 2-year and 3-year OS, which was sustained in external validation cohorts: Hefei (0.820, 0.805, 0.775) and Shihezi (0.651, 0.706, 0.747). SHAP analysis identified CEA, CA125, age, MPV, CA19-9, INR and monocyte that contributed to the accurate prediction of RSF model. This study provides an innovative strategy for the convenient and accurate prediction of survival outcome of CRC individuals based on routine blood laboratory indices. RSF model helps oncologists to early identify CRC patients with high risk of death and provides a basis for personalized treatment.
Subject terms: Cancer, Computational biology and bioinformatics, Oncology, Risk factors
Introduction
Colorectal cancer (CRC) is a predominant tumor in the digestive tract, with a high incidence in clinical practice, causing a heavy burden on the medical system globally1. Despite significant progress in the screening and treatment of CRC, the overall survival outcome of individuals with advanced-stage is still less favorable. Unfortunately, metastatic CRC remains a main cause of cancer-related death worldwide2. There are approximately 1.85 million cases diagnosed with CRC, and 850,000 deaths attributed to CRC annually3. The identification of reliable prognostic markers facilitates risk stratification in CRC, which not only optimizes treatment planning but also ultimately improves patient prognosis.
The prognostication of CRC has traditionally depended on the TNM staging system, but some CRC individuals with the same TNM stage displayed different prognosis4, indicating that more reliable prognostic biomarkers need to be designed. Liquid biopsy, an innovative noninvasive technique, has attracted considerable attention from clinicians and researchers over the past few decades. Recently, its widespread application in CRC management has established it as an integral component of precision medicine5. DNA methylation, gut microbes, CTCs and exosomes were reported to have a close correlation with the prognosis of CRC individuals6–9. Beyond the high cost, these novel biomarkers have not been integrated into routine clinical practice. Consequently, there is an urgent imperative to conduct comprehensive research on liquid biopsies for surveillance of patient survival.
In clinical practice, a number of blood-based indicators are routinely employed, several of which have demonstrated clear prognostic value in CRC. Notably, serum carcinoembryonic antigen (CEA) is a recommended prognostic biomarker, and elevated CEA levels indicate disease progression10. A recent study from Yang et al.11 reported that the C-reactive protein-albumin-lymphocyte index was independently correlated with the prognosis among CRC patients. Our previous study demonstrated that systemic inflammation response index, derived from blood routine, is a promising prognostic biomarker in CRC individuals who experienced surgery12. Given the prognostic value of existing clinical serum biomarkers, we therefore aimed to construct prognostic models by integrating all available blood indices using machine learning methods. This approach holds the potential to open new avenues for predicting survival in individuals with CRC.
In the present study, we attempted to fully leverage the prognostic significance of massive routine blood indexes to create an efficient prognostic model. Importantly, we constructed the predictive models based on seven time-to-event models to ensure the high accuracy of the predictive model. The SHapley additive exPlanation (SHAP) model provides global and local explanations by calculating the marginal contribution of each feature to the model’s output13, thus we utilized this method to exhibit the feature importance and explain the prediction results of the model. Finally, we validated the interpretable machine learning methods based on routine blood indexes to predict the risk of death in CRC individuals from another two independent clinical cohorts. Hence, this multi-center study will assist clinicians in identifying CRC patients with high risk, and provide a significant clue for the personalized treatment plans, ultimately improving the prognosis of CRC patients.
Results
Baseline clinical features of CRC individuals
Four thousand, four hundred and ten patients with CRC from three independent medical centers were initially screened, and a total of 1470 CRC individuals (Union cohort: 850; Hefei cohort: 403; Shihezi cohort: 217) were finally confirmed into this analysis according to the inclusion and exclusion criteria. The detailed enrollment flowchart of CRC patients was listed in Fig. 1. The demographic data and blood test results were vividly shown in Table S1. The distributions of most blood test data were largely comparable among the three cohorts. In the Union cohort, there were 186 patients (21.9%) died during the survival period. The mean age was 57.9 ± 12.1 years, and 59.5% were male patients. In the Hefei cohort, there were 54 patients (13.4%) died during the survival period. The mean age for them was 57.8 ± 11.6 years, and 57.8% were male patients. Finally, there were 65 patients (30.0%) died during the survival period. The mean age was 60.2 ± 12.6 years, and 63.1% were male patients in the Shihezi cohort.
Fig. 1.

The selection process of CRC patients from three medical centers.
Model selection and performance comparison
The Union cohort was used as the training set (70%), and the internal validation set (30%), and Hefei and Shihezi cohorts were used as the external validation sets. Four demographic features and 41 blood indexes were initially included in this analysis, and D-dimer was excluded as the high missing rate. Seven time-to-event models [Cox models with penalization (CoxPH), random survival forests (RSF), support vector machines (SVM), gradient boosting survival analysis (GBM), ridge regression (ridge), elastic net (E-net) and survival tree (tree)] were applied to predict the survival outcome of CRC individuals. The comprehensive workflow of this AI study is visually shown in Fig. 2.
Fig. 2.

Schematic representation of the AI study workflow.
In the training phase (Union cohort), the RSF model achieved the highest discrimination with a C-index of 0.721 (95% CI: 0.654–0.782) among the seven models; thus, the RSF model was selected as the optimal model for subsequent feature optimization (Fig. 3a). Importantly, the RSF model also exhibited robust generalizability in independent external validation. In the Hefei cohort, the RSF model maintained excellent performance with a C-index of 0.743 (95% CI: 0.676–0.811). Similarly, in the geographically distinct Shihezi cohort, the model achieved a C-index of 0.650 (95% CI: 0.576–0.716) (Table S3).
Fig. 3. Predictive performance of seven machine learning algorithms in the Union cohort, and prognostic utility of RSF model for colorectal cancer patients.

a The values of C-index with 95%CIs obtained by seven machine learning algorithms in the Union cohort. Time-dependent ROC curves of the RSF model for predicting 1-year, 2-year, and 3-year overall survival in the b Union cohort, c Hefei cohort, and d Shihezi cohort.
Comprehensive assessment of the RSF model
To further validate the clinical reliability of the RSF model, we conducted a multi-dimensional assessment based on discrimination, calibration, clinical utility, and risk stratification capabilities. Time-dependent ROC analysis revealed excellent predictive performance of RSF (Fig. 3b–d). In the Union training cohort (Fig. 3b), the model achieved time-dependent area under the curve (tAUC) values of 0.768, 0.775, and 0.731 for 1-year, 2-year, and 3-year OS, respectively. The Hefei cohort (Fig. 3c) yielded robust tAUCs of 0.820 (1-year), 0.805 (2-year), and 0.775 (3-year). The Shihezi cohort (Fig. 3d) also demonstrated satisfactory discrimination with tAUCs of 0.651, 0.706, and 0.747 for 1-year, 2-year, and 3-year OS, respectively.
We additionally conducted decision curve analysis (DCA) to comprehensively assess the clinical utility of the RSF model. The results demonstrated that, over a wide range of clinically reasonable threshold probabilities, the RSF model offered a greater net benefit than the two extreme strategies of “treat-all” or “treat-none” (Fig. 4). This consistent advantage was observed at 1, 2, and 3 years, and across all three medical cohorts, indicating that utilizing the RSF model to guide clinical decisions could enhance patient management.
Fig. 4. Decision curve analysis (DCA) of the RSF model for colorectal cancer individuals.

DCA quantifying the clinical usefulness of the RSF model in the (a 1-year, b 2-year, c 3-year) Union, (d 1-year, e 2-year, f 3-year) Hefei, and (g 1-year, h 2-year, i 3-year) Shihezi cohorts.
In terms of model calibration, the calibration plots revealed strong consistency between predicted survival probabilities and observed outcomes across all three cohorts (Fig. 5a–c). The curves for 1-year, 2-year, and 3-year survival outcomes of the CRC individuals were nearly aligned with the ideal 45-degree diagonal. The detailed parameters, such as intercept, slope, integrated calibration index (ICI) and E50, of the calibration curves in three medical centers are exhibited in Fig. S1. This supports the reliability of the RSF model in providing individualized prognostic predictions.
Fig. 5. Calibration curves and Kaplan-Meier curves of the RSF models in different cohorts.

Calibration curves of the RSF model for 1-, 2-, and 3-year overall survival in the a Union cohort, b Hefei cohort, and c Shihezi cohort. Kaplan–Meier survival analysis for risk stratification based on the RSF model. Significant survival differences were observed in the d Union (P < 0.001), e Hefei (P < 0.001), and f Shihezi (P = 0.0054) cohorts. The distributions of the high and low risk patients with CRC in the three medical cohorts are displayed in the g Union, h Hefei, and i Shihezi cohorts.
Finally, CRC patients were stratified into high-risk and low-risk groups based on the best cut-off value derived from the log-rank chi-square test in the Union training cohort (Fig. S2). High-risk patients with CRC exhibited significantly worse overall survival compared to low-risk patients in the Union cohort (P < 0.001, Fig. 5d), Hefei cohort (P < 0.001, Fig. 5e), and Shihezi cohort (P =0.0054, Fig. 5f). The distributions of the high and low risk patients with CRC in the three medical cohorts are displayed in Fig. 5g–i.
Model interpretations with SHAP
Since a machine-learning model is often regarded as a “black box,” we utilized the SHAP method to reasonably explain the RSF model. SHAP analysis could interpret the RSF model at the global and local levels. The global explanation described the overall functionality of the RSF model. The contributions of the clinical features to the predictive model were assessed using the average SHAP values (Fig. 6a), and the top 7 clinical indexes were CEA, CA125, age, MPV, CA19-9, INR, and monocyte (Fig. 6b).
Fig. 6. Model interpretation using SHAP analysis.

a Beeswarm plot showing the distribution of SHAP values for each selected feature. b Bar chart ranking features by mean absolute SHAP value, identifying CEA, CA125, and age as the top three predictors.
Subsequently, we performed a SHAP-based ablation study using a forward feature selection strategy. The predictive performance of the RSF model was evaluated within the Union training cohort using the mean 5-fold cross-validated C-index, while variables were added sequentially according to their SHAP importance ranking. As illustrated in Fig. S3, the mean C-index peaked at 0.72 when the top 7 features were included. Accordingly, the final optimal RSF model was constructed using these seven features: CEA, CA125, age, MPV, CA19-9, INR, and monocyte.
Comparison with the TNM staging system
As illustrated in Fig. S4a, the RSF model consistently achieved higher C-indices compared to the TNM staging system across all three cohorts, indicating superior overall discrimination to the TNM staging system (Fig. S4). Furthermore, we utilized net reclassification improvement (NRI) along with integrated discrimination improvement (IDI) to quantify the improvement in risk prediction accuracy of the RSF model at the 1-year benchmark. As shown in Fig. S5a, the RSF model yielded positive NRI values of 0.400 in the Union cohort, 0.379 in the Hefei cohort, and 0.125 in the Shihezi cohort. Similarly, the IDI analysis (Fig. S5b) confirmed the robustness of the RSF model, with positive improvements of 0.029 (Union), 0.034 (Hefei), and 0.019 (Shihezi). Collectively, these metrics demonstrate that integrating routine blood indices via RSF model provides a more precise prognostic assessment than TNM staging system alone.
Subgroup analyses
As RSF model obtained the high predictive performance in Union cohort, and we performed subgroup analyses to assess its predictive ability of machine learning models among specific population based on sex, age bands, TNM stage, primary tumor location and adjuvant therapy. The detailed C-indexes and 95% CIs of RSF model in subgroups are listed in (Fig. S6). In the Union cohort, the C-index remained consistently high across most subgroups. Notably, even in high-risk sub-populations, such as patients with TNM stage III–IV or those receiving adjuvant chemotherapy, the RSF model maintained satisfactory discriminative ability in both the Hefei and Shihezi cohorts. The P values for formal interaction tests among subgroup analysis are listed in Fig. S7, and no significant difference of C-indexes was detected among the subgroups in three medical cohorts.
Discussion
This multi-center study leverages seven advanced time-to-event models to accurately predict the risk of death for CRC individuals. This study demonstrates the excellent discrimination ability and calibration capacity of RSF in both training and validation cohorts. The three medical centers (Union cohort, Hefei cohort and Shihezi cohort) are distributed in different areas of China, but the predictive performances of RSF were still encouraging, highlighting the robustness and universality of the blood data-based predictive model. Importantly, the SHAP method provides a transparent explanation of the random forest, revealing CEA, CA125, age, MPV, CA19-9, INR, and monocyte as the dominant predictors of prognosis among CRC individuals. In brief, these results provide an innovative strategy for the convenient and accurate prediction of the survival outcome of CRC, emphasizing the importance of routine blood laboratory indices in combination with machine-learning techniques.
As an accurate prognosis prediction is critical for personalized treatment and care in the cancer population, a number of studies have investigated survival analysis. Wang et al.14 created a survival nomogram based on pathology, clinical factors, radiomics features and immune microenvironment to predict the survival outcome of CRC patients with lung metastasis, and this survival nomogram displayed outstanding predictive accuracy for survival outcome (AUC:0.860). Skrede et al.15 developed a useful prognostic marker for CRC patients after surgical resection via pathology images, and this pathological prognostic marker could also be used to guide the selection of cancer treatment. Xiao et al.16 developed a histology-based model via deep learning to predict 5-year recurrence risk among CRC individuals, but this model achieved AUCs of 0.833 in the validation cohort and 0.715 in the external cohort. Despite the misdiagnosis of the pathological images, the pathological images were not available in primary hospitals. Hence, developing a predictive model with high accuracy based on blood indexes is critical to improve the diagnosis and treatment capabilities of primary hospitals.
Wu et al.17 used 58 routine blood biochemical indices to construct a diagnostic model for gastric cancer (GC), and obtained the diagnostic accuracy of 0.9138. Our previous work18 also underscored the blood indices that were routinely measured in clinical practice, and used these blood indices to create a predictive model for HER2 mutation. Li et al.19 reported that a combination of CEA, CA19-9, and CA125 has significantly improved the predictive performance of the prognosis in CRC individuals, but the predictive performance is only 0.774. Our recent work20 developed the diagnostic models for patients with young-onset CRC based on clinical variables, and also obtained nice predictive accuracy. In the present work, we included 4 demographic features and 41 clinical indices derived from blood tests for model construction. After the selection via 7 time-to-event models, the independent blood indices could also well predict the overall survival of CRC patients. The highest predictive performance was RSF in the training set and in the validation cohorts. The present work highlights the importance of blood indices, such as CEA, CA125, age, MPV, CA19-9, INR, monocyte, while they were easily ignored by some oncologists in the clinical settings.
Our study found that CEA, CA125, age, MPV, CA19-9, INR, and monocyte were the strong predictive factors for the overall survival of CRC individuals. Due to the heavy medical burden, prognostic biomarkers should be cost-effective. The CEA, CA125, age, MPV, CA19-9, INR, and monocyte indices were routinely measured in most primary hospitals, and were different from the ctDNA, exosome and microbes, which were unified and only available in several top hospitals. Importantly, the common blood indices support a rapid and real-time prediction, which is really easy to use in clinical practice. Dealing with the massive blood data from thousands of CRC patients is time-consuming and labor-intensive, and artificial intelligence could solve this problem.
Artificial intelligence (AI) technology, represented by machine learning and deep learning, has achieved remarkable progress in multi-omics and model construction, enhancing the precise diagnosis and accurate prognostic prediction of cancer patients21. Our previous study22 used AI models to accurately diagnose and predict the survival outcomes of patients with gastric cancer, which will offer assistance to choose appropriate treatment to improve the survival status of gastric cancer patients. In this present analysis, we used seven time-to-event models to construct the prognostic models based on the blood indices, and validated the prognostic models with two independent cohorts. Fortunately, the prognostic models all obtained good results, and the RSF model almost reached good predictive performance and accuracy in three medical centers. In brief, machine learning methods make the massive blood data more efficient and accurate in predicting survival outcomes among CRC individuals.
Although this is a multicenter clinical research, several limitations still existed in the present study. First, the clinical data were all retrospectively collected via an electronic medical system, thus some bias might be inevitable. Then, we only collected the clinical data on admission, and the dynamic changes of these data were not available, which might also possess impact on the predictive performance. Finally, over a 10-year period, inter-site variability was observed across the three medical cohorts due to differences in laboratory instruments, assays, and reference ranges. Hence, the prospective clinical trials combined with advanced AI technologies are needed to further validate the predictive models based on routine blood data.
Based on routine blood indexes and multicenter data, we successfully developed and validated the explainable machine learning models to predict the survival outcomes among CRC individuals. The RSF model displayed superior predictive ability in both training and external validation cohorts. Our study provides an innovative strategy for the convenient and accurate prediction of the survival outcome of CRC based on routine blood laboratory indices, but more prospective clinical trials are needed to validate the predictive models based on routine blood data.
Methods
Study cohort
This retrospective multi-center cohort study in China was carried out in CRC individuals for the development and validation of the prediction model. The training cohort consisted of CRC patients admitted to Wuhan Union Hospital from 2013 to 2017. The validation cohorts consisted of CRC patients admitted to Second People’s Hospital of Hefei (Hefei cohort) from 2017 to 2023, and CRC patients admitted to Shihezi University Hospital (Shihezi cohort) from 2017 to 2023. The inclusion criteria were as follows: (1) the diagnosis of CRC was based on pathology; (2) all the CRC patients received surgical resection; (3) CRC patients with a hospital stay length of more than 48 h. The exclusion criteria were as follows: (1) patients with severe renal disease, cardiovascular disease or metabolic disease; (2) patients with a history of other malignant tumors; (3) patients with targeted immune therapy before surgical resection; (4) CRC individuals with key information missing. Finally, a total of 1470 CRC individuals (Union cohort: 850; Hefei cohort: 403; Shihezi cohort:217) were formally included in this multi-center study. The Union cohort was used as the training set (70%), and internal validation set (30%), and Hefei and Shihezi cohorts were used as the external validation sets. This study protocol was in line with the principles of the Declaration of Helsinki. This multi-center study was approved by the Ethics Committee of Wuhan Union Hospital (No. 2018-S377), Second People’s Hospital of Hefei (No. 2023-S127) and Shihezi University Hospital (No. KJ2023-351-01). All the CRC individuals were willing to take part in this clinical research and gave their informed consent to this clinical research. The research outcome of this multi-center study is the overall survival, and the follow-up endpoint in Hefei and Shihezi cohorts was December 31, 2024, and the follow-up endpoint in the Union cohort was December 31, 2021.
Data process and feature selection
In order to develop an easy-to-use prediction model based on a milliliter of blood, we only collected demographic data and blood test indexes of each patient from the electronic medical record system. The demographic data consist of body mass index (BMI), age, gender and smoking. The blood test indexes consist of blood routine, coagulation function, liver and kidney function and tumor markers. Indicators obtained from blood routine include white blood cell count, neutrophils, lymphocytes, monocytes, eosinophils, red blood cell count, platelet (PLT) count, hemoglobin (HGB) level. Indexes obtained from biochemical tests were alanine aminotransferase (ALT), aspartate aminotransferase (AST), alkaline phosphatase (ALP), lactate dehydrogenase (LDH), gamma-glutamyl transpeptidase (GGT), triglyceride (TG), total cholesterol (TC), high-density lipoprotein-cholesterol (HDL-C), low-density lipoprotein-cholesterol (LDL-C), total bile acid (TBA), and blood urea nitrogen (BUN). Features from blood coagulation were prothrombin time (PT), activated partial thromboplastin time (APTT), international normalized ratio (INR) and D-dimer. Finally, serum tumor markers consisted of carcinoembryonic antigen (CEA), carbohydrate antigen (CA)125, CA19-9 and CA724. TNM stage was also obtained for the assessment of tumor staging. For CRC patients with repeated hospital visits, the blood data from their first visit within 48 h were selected for subsequent analysis. The detailed features included in this analysis were listed in Table S3.
Clinical variables with more than 65% missingness in the Union training cohort were excluded before model development, and CRC patients with more than 20% missingness across the retained candidate features were excluded before imputation. Missing values in the remaining variables were imputed using an iterative imputation model fitted only in the Union training cohort, and the fitted model was subsequently applied unchanged to the Union internal validation cohort and the two external validation cohorts. After imputation, continuous variables were normalized using training-derived scaling parameters only, whereas binary variables, including gender and smoking, were retained as binary indicators. No preprocessing parameter was refit in the external cohorts, thereby ensuring a strictly locked evaluation and preventing data leakage.
Development and validation of the machine-learning models
We applied seven time-to-event models (CoxPH, RSF, survival tree, E-Net, GBM, Ridge, and SVM) to predict the prognosis of the CRC individuals. The detailed descriptions of these machine-learning algorithms were reported in a previous AI research23. To ensure model robustness and prevent overfitting, hyperparameters of the RSF model were optimized using grid search, with 5-fold cross-validation within the training cohort, which are listed in Table S4.
Model assessment and explanation
We calculated and compared the C-index of all seven models across the training cohort (Union) and two external validation cohorts (Hefei and Shihezi). The predictive model demonstrating the highest C-index across the three medical cohorts was identified as the optimal model for further analysis.
| 1 |
Where Ti and Tj represent the survival times of CRC patient i and patient j, ηi and ηj stand for their predicted risk scores.
TdROC curves were generated to assess the performance of the model at 1-, 2-, and 3-year intervals, with the tAUC calculated to quantify predictive accuracy. Calibration curves were plotted to rate the consistency between model-predicted probabilities and actual occurrence probabilities. Ideally, the curve should align with the 45◦ diagonal line. Second, survival calibration across consistent time points was examined using calibration plots and quantitative calibration measures (calibration intercept, slope, ICI and E50), with bootstrap-based 95% confidence intervals provided. DCA was performed to quantify the net benefit of machine learning models across different threshold probabilities, with comparisons made against the “treat-all” and “treat-none” strategies.
Interpreting machine learning models is vital to solving the “black box” problem. The SHAP method is a feature-based interpretability technique that could rank the importance of input features and interpret the results of the machine learning models. The SHAP method could well enhance the interpretability and transparency of the machine learning models. The SHAP method explains the predictive models based on local and global perspectives24. The local interpretation could display a specific prediction for a CRC patient by inputting the specific features. The global interpretation could give accurate attribution values for each blood index within a machine learning model to show the relationships between input features and the survival outcome of CRC patients.
A SHAP-guided forward feature selection procedure was performed entirely within the Union training cohort to identify the most parsimonious feature subset. First, an RSF model was fitted using the full candidate feature set in the Union training cohort only. Feature importance was then ranked according to the mean absolute SHAP values derived from this training-only model. Next, features were added sequentially according to this ranking, and the optimal subset size was defined as the number of top-ranked features that yielded the highest mean 5-fold cross-validated C-index within the Union training cohort. After the optimal subset had been fixed, the final RSF model was refit on the full Union training cohort and then evaluated on the Union internal validation cohort and the two external validation cohorts without further feature selection.
Statistical analysis
All statistical analyses in this study were conducted using Python (version 3.12.11) and SPSS software (version 21.0). The detailed information of package versions and random seeds used in this analysis was listed in Table S5. The NRI and IDI were calculated to quantify the improvement in prognostic discrimination of the RSF model compared with the TNM staging system at a fixed time point t.
| 2 |
| 3 |
where and t denote the predicted event probabilities at time t from the RSF model and the TNM staging system, respectively. D(t) = 1 indicates patients who experienced death by time t, whereas D(t) = 0 indicates patients who did not experience death by time t. The overbar denotes the mean predicted event probability within the corresponding patient group. Subgroup analyses were performed based on common clinical factors to assess the predictive ability of machine learning models among a specific population. Kaplan–Meier survival curves were plotted to visualize the survival differences between the two groups. Only when P < 0.05 at both sides was then deemed as statistically significant.
Supplementary information
Acknowledgements
This work was supported by the National Nature Science Foundation of China (Nos. 82403144 and 82505327), and the China Postdoctoral Science Foundation-Hubei Joint Support Program (No. 2025T068HB). Hubei Provincial Health Commission Young Talent Project (No. WJ2025Q008), Hubei Provincial Natural Science Foundation of China (No. JCZRQN202500265), Taishan Scholars Program of Shandong Province (No. ts202507367). We specifically thanked Professor Zhichao Zuo, from Xiangtan University, for his professional AI guidance.
Author contributions
Y.C. and S.T. conceptualized the study and defined the methodology. H.C., Q.L., and Y.C. provided theoretical guidance. J.L., Y.C., H.C., Y.H., and D.S. acquired funding for the projects that provided the data for this study. Y.C., B.Y., J.L., P.Z., Z.S., and L.Q. administered the projects that provided the data for this study. J.L. and Q.L. conducted the statistical analysis and developed the Python codes. Q.L. provided methodological feedback. Y.C. and S.T. wrote the original draft of the manuscript. All authors reviewed and edited the manuscript.
Data availability
Individual participant data in this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.
Code availability
Model construction and validation were performed in Python (version 3.12.11), and codes were made publicly available (https://github.com/cyh532/CRC.git).
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Shan Tian, Jinxiao Li, Qian Liu, Yugang Hu.
Contributor Information
Dong Sun, dsun@email.sdfmu.edu.cn.
Huan Cao, Email: caohuan2016ty@163.com.
Yinghao Cao, Email: yinghaocao@hust.edu.cn.
Supplementary information
The online version contains Supplementary material available at https://doi.org/10.1038/s41746-026-02781-5.
References
- 1.Shevach, J. W. et al. Established cancer predisposition genes in single and multiple cancer diagnoses. JAMA Oncol.11, 1222–1230 (2025). [DOI] [PMC free article] [PubMed]
- 2.Sullo, F. et al. Personalized therapy in metastatic colorectal cancer: biomarker-driven use of biologics. Expert Opin. Biol. Ther.25, 947–965 (2025). [DOI] [PubMed] [Google Scholar]
- 3.Biller, L. H. & Schrag, D. Diagnosis and treatment of metastatic colorectal cancer: a review. JAMA325, 669–685 (2021). [DOI] [PubMed] [Google Scholar]
- 4.Chen, K., Collins, G., Wang, H. & Toh, J. W. T. Pathological features and prognostication in colorectal cancer. Curr. Oncol.28, 5356–5383 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Raza, A. et al. Dynamic liquid biopsy components as predictive and prognostic biomarkers in colorectal cancer. J. Exp. Clin. Cancer Res.41, 99 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Müller, D. & Győrffy, B. DNA methylation-based diagnostic, prognostic, and predictive biomarkers in colorectal cancer. Biochim. Biophys. Acta Rev. Cancer1877, 188722 (2022). [DOI] [PubMed] [Google Scholar]
- 7.John Kenneth, M. et al. Diet-mediated gut microbial community modulation and signature metabolites as potential biomarkers for early diagnosis, prognosis, prevention and stage-specific treatment of colorectal cancer. J. Adv. Res.52, 45–57 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Hui, J. et al. Regulatory role of exosomes in colorectal cancer progression and potential as biomarkers. Cancer Biol. Med.20, 575–598 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Jiang, M. et al. Detection and clinical significance of circulating tumor cells in colorectal cancer. Biomark. Res.9, 85 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Campos-da-Paz, M., Dórea, J. G., Galdino, A. S., Lacava, Z. G. M. & de Fatima Menezes Almeida Santos, M. Carcinoembryonic antigen (CEA) and hepatic metastasis in colorectal cancer: update on biomarker for clinical and biotechnological approaches. Recent Pat. Biotechnol.12, 269–279 (2018). [DOI] [PubMed] [Google Scholar]
- 11.Yang, M. et al. Association between C-reactive protein-albumin-lymphocyte (CALLY) index and overall survival in patients with colorectal cancer: from the investigation on nutrition status and clinical outcome of common cancers study. Front. Immunol.14, 1131496 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Cao, Y. et al. Levels of systemic inflammation response index are correlated with tumor-associated bacteria in colorectal cancer. Cell Death Dis.14, 69 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Hou, F. et al. Development and validation of an interpretable machine learning model for predicting the risk of distant metastasis in papillary thyroid cancer: a multicenter study. EClinicalMedicine77, 102913 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Wang, R. et al. Development of a novel combined nomogram model integrating deep learning-pathomics, radiomics and immunoscore to predict postoperative outcome of colorectal cancer lung metastasis patients. J. Hematol. Oncol.15, 11 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Skrede, O.-J. et al. Deep learning for prediction of colorectal cancer outcome: a discovery and validation study. Lancet395, 350–360 (2020). [DOI] [PubMed] [Google Scholar]
- 16.Xiao, H. et al. Predicting 5-year recurrence risk in colorectal cancer: development and validation of a histology-based deep learning approach. Br. J. Cancer130, 951–960 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Wu, J. et al. GCdiscrimination: identification of gastric cancer based on a milliliter of blood. Brief. Bioinform.22, 536–544 (2021). [DOI] [PubMed] [Google Scholar]
- 18.Tian, S. et al. Prediction of HER2 status via random forest in 3257 Chinese patients with gastric cancer. Clin. Exp. Med.23, 5015–5024 (2023). [DOI] [PubMed] [Google Scholar]
- 19.Li, C. et al. Prediction models of colorectal cancer prognosis incorporating perioperative longitudinal serum tumor markers: a retrospective longitudinal cohort study. BMC Med.21, 63 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Zhen, J. et al. Development and validation of machine learning models for young-onset colorectal cancer risk stratification. NPJ Precis. Oncol.8, 239 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Elguoshy, A., Zedan, H. & Saito, S. Machine learning-driven insights in cancer metabolomics: from subtyping to biomarker discovery and prognostic modeling. Metabolites15, 514 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Huang, B. et al. Accurate diagnosis and prognosis prediction of gastric cancer using deep learning on digital pathological images: a retrospective multicentre study. EBioMedicine73, 103631 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Yuan, Y. et al. Personalized prediction for recurrence of cystitis glandularis: insights from SHAP and machine learning models. Transl. Androl. Urol.14, 808–819 (2025). [DOI] [PMC free article] [PubMed]
- 24.Ponce-Bobadilla, A. V., Schmitt, V., Maier, C. S., Mensing, S. & Stodtmann, S. Practical guide to SHAP analysis: explaining supervised machine learning model predictions in drug development. Clin. Transl. Sci.17, e70056 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Individual participant data in this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.
Model construction and validation were performed in Python (version 3.12.11), and codes were made publicly available (https://github.com/cyh532/CRC.git).
