Graphical abstract
Keywords: Gastrointestinal cancers, Cell-free DNA methylation, Diagnosis prediction, Cancer signal origin
Highlights
-
•
High Accuracy: SPOGIT was demonstrated robust performance in multicenter validation set (88.1 % sensitivity, 91.2 % specificity, n = 1,079).
-
•
Early Detection: The sensitivities of AA, high-risk preGC and low-risk preGC were 56.5 %, 62.4 % and 49.5 % respectively.
-
•
Better Outcomes: SPOGIT may reduce late-stage diagnoses by 92 % and boost 5-year survival rates by 27.02–30.47 %, enabling earlier intervention.
-
•
Minimally Invasive & Accessible: Requires only 10 mL of blood and <30 ng cfDNA input, ideal for low-resource healthcare settings.
Abstract
Background
Gastrointestinal (GI) tract cancers are the second leading cause of cancer-related mortality, often due to late detection. There is a critical need for non-invasive, highly sensitive biomarkers for early-stage cancer and precancerous lesion detection to enable timely intervention. This study aimed to develop a blood-based method specifically optimized for early GI cancer detection and population-level screening.
Methods
Using large-scale public tissue methylation data and the Twist probe cfDNA profiles, we developed SPOGIT (Screening for the Presence of Gastrointestinal Tumors), a multi-algorithm model (Logistic Regression/Transformer/MLP/Random Forest/SGD/SVC) for early GI cancer detection. The model was rigorously validated through an internal (n = 83) and multicenter external validation (386 cancers/113 controls/580 precancers), with an interception model assessing its clinical potential.
Results
SPOGIT demonstrated high accuracy in detecting GI cancers, with a sensitivity of 88.1 % and a specificity of 91.2 %. Notably, it effectively identified early-stage (0-II) cancers with 83.1 % sensitivity. The model also showed significant potential for intercepting premalignant progression, detecting advanced adenomas (AA) and gastric precancerous lesions with sensitivities of 56.5 % and up to 62.4 %, respectively. In the external independent validation cohort, a complementary model CSO (Cancer Signal Origin) demonstrated an accuracy of 83 % for colorectal cancer and 71 % for gastric cancer. Most importantly, simulation analyses projected that SPOGIT implementation could significantly reduce late-stage diagnoses and increase 5-year survival rate by 27.02 % through early interception.
Conclusions
This study introduced a novel, dual-model blood architecture (SPOGIT/CSO) that enables highly accurate, early detection of GI cancers and their precursors. By facilitating timely clinical intervention, SPOGIT/CSO represented a paradigm-shifting strategy with the potential to significantly improve patient survival outcomes and transform GI cancer management.
Introduction
Gastrointestinal cancers (GI cancers) represent a spectrum of malignancies including esophageal, gastric, colorectal, hepatocellular, and pancreatic carcinomas. Among these, gastric cancer (GC) and colorectal cancer (CRC) stand out as particularly critical targets for early detection, with GC ranking as the 4th most common cancer globally and CRC as the 3rd, together accounting for approximately 20 % of global cancer mortality [1,2]. Colorectal and gastric cancers share significant molecular similarities, including common genetic alterations and epigenetic signatures [3]. These adenocarcinomas of the digestive tract present unique diagnostic challenge compared to other GI malignancies like hepatocellular and pancreatic carcinomas that are often detectable through ultrasound or CT at earlier stages. In contrast, early gastric and colorectal lesions frequently evade radiographic detection, and the invasive nature, procedural risks, and resource demands of gastroscopy/colonoscopy severely limit their population-wide applicability, particularly in low-resource settings applicability [4,5]. Current non-invasive alternatives—including serum biomarkers (CEA, CA19-9) and stool-based tests (gFOBT, FIT)—exhibit suboptimal sensitivity (CEA 54.7 %, CA19-9 59 %, gFOBT 31 %-79 %, FIT 71.3 %) [[6], [7], [8]]for early lesions, highlighting the urgent need for minimally-invasive, cost-effective screening solutions with high acceptability for large-scale implementation in diverse healthcare settings.
Advances in liquid biopsy technologies have reinvigorated the search for non-invasive biomarkers, with circulating tumor DNA (ctDNA) methylation emerging as a transformative tool. Unlike genetic mutations, methylation alterations occur early in carcinogenesis and exhibit tissue-specific patterns, making them ideal candidates for early detection [[9], [10], [11]]. Preliminary successes, such as the Methylation CpG Tandem Amplification and Sequencing (MCTA-Seq) achieving 67 % sensitivity for gastric cancer [12] and ColonSecure demonstrating 86.4 % sensitivity for colorectal malignancies [13], highlight the potential of methylation-based approaches. Individuals often face risks of multiple cancers, making single-cancer screening costly and prone to cumulative false positives, leading to unnecessary diagnostics. Multi-cancer detection biomarkers with high sensitivity offer greater clinical utility, and recent studies have highlighted the feasibility of blood-based multi-cancer detection [[14], [15], [16]]. However, existing cfDNA-based methods for multi-cancer detection struggle to balance sensitivity and specificity across heterogeneous tumor types, as evidenced by the THUNDER study’s 70.6 % sensitivity in detecting the six cancers and 81.5 % sensitivity of the GutSeer in screening five cancers [15,17,18]. Gastrointestinal malignancies, while exhibiting considerable heterogeneity akin to other tumor types, share a common origin in the adenomatous epithelium of the digestive tract. This shared embryological derivation results in overlapping pathogenic mechanisms, histopathological features, clinical presentations, diagnostic and therapeutic strategies, and prognostic determinants. Emerging evidence indicates that DNA methylation-driven oncogenesis is a conserved molecular hallmark across gastrointestinal cancers. This epigenetic uniformity provides a compelling rationale for developing highly sensitive and specific methylation-based biomarkers for the early detection of gastrointestinal neoplasms.
To achieve the objective, we integrated multiple machine learning approaches to develop SPOGIT (Screening for the Presence of Gastrointestinal Tumors)—a blood-based methylation detection model specifically optimized for gastrointestinal malignancies. Leveraging data from 4.5 million methylation sites across normal and gastrointestinal tissues, combined with plasma analysis from training and validation cohorts, we identified a diagnostic panel of 311 differentially methylated regions (DMRs) and a 230 DMRs tissue-origin tracing panel. To establish an optimal diagnostic model for gastrointestinal cancers, we systematically evaluated 12 machine learning algorithms, including logistic regression, random forest, and support vector machines. Through ensemble learning techniques, we integrated the strengths of individual models to develop SPOGIT. This integrated approach enabled the simultaneous achievement of high specificity (91.2 %) to minimize false positives, exceptional sensitivity for early-stage cancers (Stage 0-II detection rate: 83.1 %), and clinically significant detection capability for precancerous lesions (advanced adenoma detection: 56.5 %; high risk precursor gastric cancer: 62.4 %) in the external validation set (386 cancers/113 controls/580 precancers). The integrated CSO (Cancer Signal Origin) model further achieved 83 % accuracy in colorectal cancer localization, guiding subsequent clinical decision-making. The interception model, developed using real-world data on the incidence and 5-year survival rates of GI cancers, highlights the potential of SPOGIT in facilitating stage shift from III/IV to I/II and enhancing 5-year survival rates. This underscores SPOGIT's promise as a straightforward and cost-efficient approach for large-scale gastrointestinal cancer screening.
Methods
Study design and samples
The SPOGIT study is a multicenter, retrospective case-control study designed to detect gastrointestinal tumors early through cfDNA methylation analysis. It comprises three phases: methylation panel design, model training and validation, and independent validation. The primary endpoints are the sensitivity and specificity of the SPOGIT model and the overall accuracy of CSO prediction in the independent validation cohort, while secondary endpoints include sensitivity and specificity stratified by cancer stage and type. Participants for the training and validation cohorts were recruited from the Seventh Medical Center of PLA General Hospital, Dongying People's Hospital, and the First Hospital of Longyan, Fujian Medical University, with the external validation cohort enrolled from these hospitals during a different time period (Table 1, Table S1). The inclusion criteria consisted of individuals aged over 18 who were capable of providing written informed consent and a blood sample for detection. Healthy control subjects were required to have negative results from colonoscopy or pathology and no history of other cancers. Patients with colorectal cancer and precancerous lesions were required to have a confirmed diagnosis via colonoscopy and/or pathology, with no concurrent malignancies affecting other systems or organs. Blood was collected from participants before any form of treatment (surgery, radiotherapy, or chemotherapy). Patients assigned to the AA (advanced adenoma) group had to meet one of the following criteria: (1) Tubular adenoma or SSL (sessile serrated lesion) larger than 10 mm in diameter; (2) Tubulovillous adenoma; (3) High-grade intraepithelial neoplasia. And we defined Cp (common polyp) as a tubular adenoma or SSL smaller than 10 mm without high-risk histological features [19]. Following the national and international expert consensus on gastric precancerous conditions and lesions, we stratified patients into two risk categories [20]. Patients with mucosal atrophy and/or intestinal metaplasia in both antral and corpus mucosa or gastric intraepithelial neoplasia, were assigned to the high risk pre-GC group and patients with gastritis or intestinal metaplasia confined to the gastric antrum and incisura angularis are classified as low risk pre-GC.
Table 1.
Participant demographics and baseline characteristics.
| Cohort | Training cohort | Validation cohort | Independent validation | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Group | Cancer (n = 135) | Non-cancer (n = 66) | Cancer (n = 58) | Non-cancer (n = 25) | Cancer (n = 386) | Non-cancer (n = 113) | AA (n = 230) | CP (n = 79) | high-preGC(n = 85) | low-preGC(n = 186) |
| Age, years | ||||||||||
| Mean (SD) | 64 (9) | 55 (9) | 65 (9) | 57 (10) | 64 (10) | 58 (11) | 60 (9) | 54 (9) | 62 (9) | 56 (10) |
| Median (Q1, Q3) | 66 (58, 70) | 54 (47, 62) | 67 (59, 71) | 55 (50, 66) | 66 (55, 73) | 57 (50, 66) | 60 (53, 67) | 54 (49, 62) | 64 (55, 69) | 55 (49, 64) |
| Min, max | 39, 80 | 40, 80 | 40, 79 | 40, 77 | 32, 80 | 39, 80 | 39, 80 | 40, 80 | 45, 80 | 40, 80 |
| Sex, n (%) | ||||||||||
| Female | 47 (35 %) | 37 (56 %) | 18 (31 %) | 12(48 %) | 140 (36 %) | 67 (59 %) | 92 (40 %) | 34 (43 %) | 32 (38 %) | 85 (46 %) |
| Male | 88 (65 %) | 29 (44 %) | 40 (69 %) | 13 (52 %) | 246 (64 %) | 46 (41 %) | 138 (60 %) | 45 (57 %) | 53 (62 %) | 101 (54 %) |
| Clinical cancer stage, n (%) | ||||||||||
| 0 | 16(12 %) | 10(17 %) | 29(8 %) | |||||||
| I | 24(18 %) | 5(9 %) | 54(14 %) | |||||||
| II | 30(22 %) | 9(16 %) | 71(18 %) | |||||||
| III | 29(21 %) | 17(29 %) | 123(32 %) | |||||||
| IV | 20(15 %) | 8(14 %) | 56(14 %) | |||||||
| Missing | 16(12 %) | 9(16 %) | 53(14 %) | |||||||
| T | ||||||||||
| T0-2 | 22(16 %) | 6(10 %) | 75(19 %) | |||||||
| T3-4 | 57(42 %) | 27(47 %) | 200(50 %) | |||||||
| T miss | 56(41 %) | 25(43 %) | 111(30 %) | |||||||
| N, n (%) | ||||||||||
| N0 | 31(23 %) | 7(12 %) | 113(30 %) | |||||||
| N1-N3 | 36(27 %) | 21(36 %) | 118(30 %) | |||||||
| Missing | 68(50 %) | 30(52 %) | 155(40 %) | |||||||
| Metastasis, n (%) | ||||||||||
| M0 | 81(60 %) | 32(55 %) | 212(55 %) | |||||||
| M1 | 17(13 %) | 8(14 %) | 45(11 %) | |||||||
| Missing | 37(27 %) | 18(31 %) | 129(34 %) | |||||||
Exclusion criteria were as follows: (1) Participants who were pregnant or lactating. (2) Patients with a history of multiple primary cancers. (3) Cancer patients who had previously received surgery, radiotherapy, chemotherapy, targeted therapy, or other anti-tumor interventions. (4) A history of familial FAP (adenomatous polyposis) or IBD (inflammatory bowel disease) discovered during previous colonoscopy. (5) Recipients of organ, bone marrow, or stem cell transplants. (6) Individuals who had participated in an interventional clinical trial and had taken an experimental drug within the past 30 days. (7) Individuals deemed unsuitable for this trial by the investigator. An individual was considered unsuitable if they withdrew consent, developed health issues prior to blood collection, or were assessed by the investigator as having complex medical conditions that could pose potential safety risks. (8) Blood samples that failed quality control, including samples with non-compliant collection or processing, repeated freeze–thaw cycles, hemolysis, or microbial contamination. The study received ethical approval from all participating centers, and written informed consent was obtained from all participants.
Design of diagnostic and cancer origin panels
DNA methylation data from The Cancer Genome Atlas (TCGA, Infinium Human Methylation 450 K array) and methylome data from Gene Expression Omnibus (GEO) datasets (GSE139404, GSE48684, GSE103186), including colorectal cancer (CRC) tissues (n = 392), adjacent/normal colorectal tissues (n = 104), gastric cancer (GC) tissues (n = 383), and normal gastric tissues (n = 63), were used to identify cancer-specific and tissue-specific CpG sites. For the diagnostic panel, 201 plasma samples (67 CRC, 68 GC, 66 normal) from the training cohort were analyzed. Adjacent CpG sites (<50 bp apart) with a Pearson correlation coefficient >0.5 were merged, and differentially methylated regions (DMRs) between cancer and healthy samples were identified using Wilcoxon tests (p < 0.05, |Δβ| > 0.01). Least Absolute Shrinkage and Selection Operator (LASSO) regression was then applied to construct a panel for gastrointestinal cancer diagnosis. For the tissue origin panel, tissue-specific CpG sites were identified using TCGA data for gastric and colorectal cancer tissues, followed by analysis of plasma samples (67 CRC, 68 GC) to pinpoint DMRs specific to each cancer type using Wilcoxon tests (p < 0.05, |Δβ| > 0.01). LASSO regression was subsequently used to refine these regions, producing a panel for tracing the tissue origin of gastrointestinal cancers.
Sample collection, processing, and sequencing
Peripheral blood samples were collected from participants using EDTA-K2 vacuum tubes. Plasma was separated via a standardized two-step centrifugation protocol: an initial centrifugation at 1600 × g for 10 min at 4 °C to isolate plasma, followed by a second centrifugation of the supernatant at 16,000 × g for 10 min at 4 °C to remove residual cells and debris. The processed plasma was aliquoted and stored at −80 °C until use; all samples were transported on dry ice to prevent thawing.
cfDNA was extracted from plasma using the MagMAX™ Cell-Free DNA Isolation Kit (A29319, Thermo Fisher, Waltham, MA, USA), with strict precautions to avoid repeated freeze–thaw cycles to prevent cfDNA degradation. The concentration of isolated cfDNA was measured using fluorometric quantification with the Qubit Fluorometer 4.0 (Thermo Fisher) and the Qubit dsDNA HS Assay Kit (Thermo Fisher). Fragment size profiles of cfDNA were assessed using the 2100 Bioanalyzer (Agilent Inc., Santa Clara, CA, USA) with the Agilent High-Sensitivity DNA Kit. CfDNA samples yielding <5 ng or exhibiting quality issues were excluded.
Libraries were quantified using the Qubit dsDNA HS Assay Kit, and 4 or 8 individual libraries were pooled per capture using the Twist Human Methylome Panel or Twist Methyl Custom Panel, respectively. Hybridization and capture were conducted with the Fast Hybridization Target Enrichment Kit from Twist Bioscience. The pooled libraries were sequenced on an Illumina NovaSeq 6000 or HiSeq X Ten platform using paired-end 150-bp reads, with a target average coverage depth of >1000x per sample.
Methylation data processing
The raw image data files acquired through high-throughput sequencing were subjected to base calling using the bcl2fastq software, resulting in the conversion to raw sequencing data stored in the fastq file format. Rigorous quality control measures were applied to the on-machine raw data, facilitating the exclusion of reads of inferior quality. Bismark (Krueger, 2011, invoking Bowtie2 at the core) was used for the alignment analysis of methylation data with the reference genome. After the acquisition of alignment results, Bismark (Krueger, 2011) was once again utilized for the identification of methylation sites. Differential methylation analysis for paired samples (tumor vs. adjacent normal) was conducted employing the paired t-test method, defined as those with a methylation difference ≥10 % and a q-value (adjusted p-value) < 0.05. For unpaired samples (case vs. control), the Wilcoxon rank-sum test was employed with the same significance thresholds.
Construction of gastrointestinal cancer diagnostic model (SPOGIT)
To develop the SPOGIT model, we first constructed diagnostic models using 12 machine learning methods [21]. These algorithms were then integrated based on their respective weights, from which six were ultimately selected and applied to methylation data from a gastrointestinal cancer diagnostic panel in the training cohort. Finally, these selected algorithms were used to build an ensemble model for gastrointestinal cancer screening. A model minimizing binary deviation criteria was built through 10-fold cross-validation. Model performance was assessed using receiver operating characteristic (ROC) curves generated by the R package “pROC (1.18.0),” with the area under the curve (AUC) and sensitivity/specificity as evaluation metrics. The optimal cutoff value was determined using the maximum Youden index: test results exceeding the cutoff indicated cancer signal detection, while results below the cutoff indicated no cancer signal. Model stability was evaluated in the validation cohort, and diagnostic performance was assessed in the external validation cohort.
Construction of cancer signal origin (CSO) prediction model
Machine learning was applied to methylation data from the gastrointestinal cancer tissue origin panel in the training cohort. A logistic regression model was constructed to predict tumor signal origin in the positive group, ensuring consistency with the primary origin of each sample. Probabilities for gastric and colorectal cancer origins were normalized to sum to 100 %. The model's accuracy (percentage of correctly predicted samples in each category) was used to evaluate CSO model performance. Stability was assessed in the validation cohort, while external validation measured its diagnostic utility.
The interception model evaluates the benefits of SPOGIT in real-world setting
The clinical information required for the interception model includes: (1) the incidence and stage distribution of gastrointestinal tumors [22,23]; (2) the 5-year survival rate based on follow-up data from previous studies [24,25]; and (3) the dwell time in each stage of these cancers, estimated using a robust modeling algorithm that simulates three tumor growth scenarios: slow growth (3–7 years in stage I, depending on the cancer type), fast growth (2–4 years), and aggressive growth (1–2 years) [26]. Illustrated in Fig. S5, the model assumes cancers progress sequentially through stages I to IV, with SPOGIT's detectability at each stage calculated as the sum of newly detected cases and previously missed cases from earlier stages.
Statistical analysis
All statistical analyses were conducted using SPSS 26.0 or R 4.1.3. For continuous variables, the total number of participants (n), mean, standard deviation (SD) or standard error (SE), median, first quartile (Q1), third quartile (Q3), minimum, and maximum were calculated. For categorical variables, the number and percentage of participants in each category were reported. A two-tailed p-value < 0.05 was considered statistically significant. ROC analysis, performed using the “ROC” function in the R package “pROC,” generated AUC and 95 % confidence intervals (CI). Confidence intervals for proportions were calculated using the Clopper-Pearson method. Heatmaps were created with the R package “ComplexHeatmap,” and t-SNE analysis was performed using the R package “Rtsne”.
Results
Participant disposition
This study analyzed blood samples from 321 CRC, 258 GC, 204 normal, 230 AA (advanced adenoma), 79 Cp (Common polyp), and 85 high risk-preGC (precursor gastric cancer), 186 low risk − preGC participants across three stages (Fig. 1 and Table 1). Samples from 96 CRC, 97 GC, and 91 normal participants, collected during the same period, were used for marker selection, model development, and validation (Targeted methylation sequencing was performed using Twist probes covering >4 million CpG sites). Independent validation was conducted using samples from 225 CRC, 161 GC, 113 normal, 230 AA, 79 Cp, and 85 high risk − preGC, 186 low risk − preGC participants collected during a different period (Targeted methylation sequencing for specifically identifying gastrointestinal cancer-specific differentially methylated positions (DMPs). Baseline demographic characteristics were consistent across the training, validation, and independent validation cohorts (Table 1). The baseline clinicopathological characteristics of patients in the training, validation, and independent validation cohorts were listed in Table S1.
Fig. 1.
Study design and patient enrollment. This study analyzed DNA methylation data from colorectal cancer (CRC) and gastric cancer (GC) tissue samples from the TCGA (The Cancer Genome Atlas) and GEO (Gene Expression Omnibus) datasets to identify cancer-specific and tissue-specific CpG sites. Based on plasma samples from the training set, a diagnostic panel containing 311 differentially methylated regions (DMRs) and a tissue origin panel with 230 DMRs were constructed. Using these panels, SPOGIT (Screening for the Presence of Gastrointestinal Tumors)/CSO (Cancer Signal Origin) models were developed with multi-machine learning integration.
Panel Construction
Large-scale DNA methylation data from TCGA gastrointestinal tumors and GEO datasets (GSE139404, GSE48684, GSE103186) identified 39,661 differentially methylated sites (DMSs) for CRC and 43,872 for GC (Fig. S1A-B). Methylation patterns between cancer and normal tissues were distinctly separated (Fig. S1E-F), and t-SNE analysis confirmed this distinction (Fig. S1C-D). For the diagnostic panel, we performed targeted methylation sequencing on plasma samples from the training cohort (67 CRC, 68 GC, and 66 normal controls) using Twist probes covering >4 million CpG sites, followed by Wilcoxon rank-sum differential analysis (significance threshold: p < 0.05 with |Δβ| > 0.1). Adjacent DMSs (<50 bp apart) with a Pearson correlation coefficient >0.5 were merged, resulting in 17,614 differentially methylated regions (DMRs) (Fig. S2A). LASSO regression reduced these to a final diagnostic panel of 311 DMRs (Fig. S2B). For the tissue origin panel, differential methylation analysis of TCGA gastrointestinal tissue data using chAMP (p < 0.05, |Δβ| > 0.2) identified 2,100 hypermethylated sites specific to CRC and 5086 to GC (Fig. 4A). t-SNE analysis demonstrated clear clustering of CRC and GC samples based on these sites (Fig. 4B). Further Wilcoxon differential analysis ((p < 0.05, |Δβ| > 0.01) on 67 CRC and 68 GC samples in the training cohort yielded 357 CRC-specific and 1024 GC-specific hypermethylated DMRs (Fig. S3A). LASSO regression refined these to a final tissue origin panel comprising 230 DMRs (Fig. S3B).
Fig. 4.
Performance of the CSO (Cancer Signal Origin) Model. (A) Hierarchical clustering heatmap showcasing differences in tissue-specific methylation regions among GC, CRC, and Non-GC/Non-CRC samples; (B) t-SNE visualization based on tissue-specific methylation regions for tumor origin prediction; (C) Confusion matrix displaying the accuracy of tumor origin predictions by the CSO model across training, validation, and independent validation cohorts; (D-E) Stacked bar plots illustrating the conditional CSO probabilities (y-axis) derived from CSE decomposition for each sample (x-axis). Samples are grouped by cancer type, with the x-axis ordered by the highest conditional probabilities corresponding to their own cancer type; (F-G) Accuracy of the CSO model stratified by tumor stage. Abbreviations: CSO, cancer signal origin; GC, gastric cancer; CRC, colorectal cancer; t-SNE, t-stochastic neighbor embedding.
Cancer signal detection for GI cancers
Based on methylation data from a 311-DMR gastrointestinal cancer diagnostic panel, utilizing methylation data from a 311-DMR gastrointestinal cancer diagnostic panel, we employed 12 distinct machine learning algorithms to construct and validate diagnostic models for gastrointestinal malignancies. Different models exhibit varying sensitivity and specificity in tumor prediction (Fig. S4A), with even more pronounced differences in their efficacy for detecting precancerous lesions (Fig. 2A,B). Among these models, Logistic Regression demonstrated the highest predictive performance for GI cancers (AUC = 0.95), followed by the MLP model (AUC = 0.93) and the Transformer model (AUC = 0.93) (Table S2). The discriminative performance metrics of all 12 models are detailed in Supplementary Table S2. Fig. 2C–E present the ROC curves of the top five performing models for advanced-stage gastric cancer (GC III-IV), early-stage gastric cancer (GC 0-II) and high risk-preGC. The models’ performance for low risk preGC is presented in Fig. S4B. Similarly, Fig. 2F–H display the ROC curves of the top five models for: colorectal cancer and advanced colorectal adenomas. Notably, certain models exhibited superior performance in cancer detection (e.g., Logistic Regression) but showed limited efficacy in identifying precancerous lesions. Conversely, other models (e.g., Transformer) demonstrated enhanced diagnostic accuracy for precancerous conditions. To optimize both sensitivity and specificity for cancer and precancerous lesion detection, we performed weighted ensemble integration of the top-performing models. The performance metrics of 11 integrated models are summarized in Supplementary Table S3, with the top six weighted models showing the most robust combined performance, forming our final SPOGIT model.
Fig. 2.
Performance of Machine learning models to predict GI cancers and precancerous lesions. (A-B) Model performance for cancer diagnosis in high risk-preGC(A) and AA(B). The diagnostic performance of the twelve models, including AdaBoost, GaussianProcess, GradientBoosting, KNeighbors, LogisticRegression, MLP, RandomForest, SGD, SVC, XGB, transformer, pretrain + transformer was tested in the external validation set. The accuracy, along with sensitivity and specificity, was shown. (C-E) ROC curves of the top five best-performing ML models in distinguishing GC III-IV(A), GC 0-II(B), high risk-pre-GC(C) from normal individuals. (F-H) ROC curve analysis evaluating the the top five models' ability to distinguish CRC III-IV, CRC 0-II, AA patients from normal individuals. Abbreviations: ROC, receiver operating characteristic; GC, gastric cancer; CRC, colorectal cancer. MLP, Multilayer perceptron; SGD, Stochastic Gradient Descent; SVC, Stochastic Gradient Descent.
In the training and validation set, methylation scores for the GC and CRC groups were significantly higher than those of the normal group (P < 0.001) (Fig. 3A). The model demonstrated a sensitivity of 91 % for diagnosing GI cancers (Fig. S5 A-D). ROC curve analysis in the validation set revealed AUC values of 0.963 (0.914–1.000) for GC, 0.966 (0.927–1.000) for CRC, and 0.964 (0.930–0.999) for GI cancers (Fig. 3B). Further evaluation in an external validation cohort yielded AUC values of 0.951 (0.925–0.976) for GC, 0.939 (0.914–0.963) for CRC, and 0.944 (0.922–0.965) for GI cancers (Fig. 3C). In this cohort, most GI cancer samples exhibited methylation scores above the threshold, while the majority of normal samples fell below it (Fig. 3D). And the SPOGIT model achieved sensitivities of 88 %for GI cancers, with a specificity of 91 % (Fig. 3F). Notably, SPOGIT tracked histopathological progressions. The AUC values for AA and CRC stages 0-II were 0.785(0.736–0.834) and 0.927 (0.893–0.960), respectively (Fig. 3E), while for low risk-preGC, high risk-preGC and GC stages 0-II, the AUC values were 0.747(0.692–0.802), 0.845 (0.793–0.898) and 0.933 (0.891–0.976), respectively (Fig. 3E). Methylation scores increased with lesion severity in colorectal lesions (CRC 0-II > AA > Cp > Normal), and GC lesions (GC 0-II > high risk-preGC > low risk-preGC > Normal), with corresponding improvements in sensitivity (Fig. 3G,H).
Fig. 3.
Development and Validation of the SPOGIT Model. (A) Methylation scores for Normal, GC, CRC, and All GI cancers in the training and validation set; (B) ROC curve analysis evaluating the SPOGIT model's ability to distinguish CRC, GC, GI cancerpatients from normal individuals in the validation set; (C) ROC curve analysis of the SPOGIT model in the external validation cohort; (D) Methylation scores for Normal, GC, CRC, and All GI cancers in the external validation cohort, with dashed lines indicating model thresholds; (E) Performance of the SPOGIT model in differentiating PreGC, AA, GC 0-II, CRC 0-II and normal individuals; (F) Sensitivity and specificity of the SPOGIT model in the external validation cohort; (G)Model scores for normal individuals and various pathological stages of gastrointestinal tumors; (H) Sensitivity of the SPOGIT model in detecting early-stage gastrointestinal tumors (GC 0-II and CRC 0-II) and precancerous gastrointestinal lesions (AA, and high risk-preGC). (I) Comparative ROC analysis of SPOGIT model versus conventional serum tumor markers in the external validation cohort; (J) Comparative DCA of SPOGIT Model versus Conventional Serum Tumor Markers. Abbreviations: SPOGIT, screening for the presence of gastrointestinal tumor; ROC, receiver operating characteristic; GC, gastric cancer; CRC, colorectal cancer; AA, advanced adenoma; PreGC, gastric precancerous lesions; DCA, Decision curve analysis.
SPOGIT standed out among its blood-based peers, demonstrating significant advantages in detecting both gastrointestinal cancer (88.1 % vs. 55.7 %–86.1 %) and, critically, precancerous lesions (58 % vs. 12.5 %–47.1 %) (Table 2). Comparative analyses revealed the SPOGIT model exhibited significantly superior diagnostic performance (AUC: 0.944 vs 0.569–0.818) to conventional serum biomarkers including CEA, CA125, and CA19-9 (Fig. 3I). Decision curve analysis (DCA) demonstrated its markedly higher net benefit across clinically relevant threshold probabilities (20–60 %) (Fig. 3J), suggesting substantial advantages over existing biomarkers for clinical decision-making in gastrointestinal cancer management.
Table 2.
Comparing the performance of SPOGIT with other cfDNA and biomarker-based tests.
| Test’s name [Ref] | Tested in | Target(s) tested | Sensitivity for precancerous lesions | Sensitivity for cancer | Specitivity |
|---|---|---|---|---|---|
| SPOGIT | Blood | 311 differentially methylated regions | 56.3 % for AA | 88.1 % | 91.2 % |
| 62.4 % for high risk pre-GC | |||||
| Colorectal cancer screening [30] | Blood | Multiomics(genomic, epigenomic, and proteomic biomarkers) | 12.5 % (11.3 %–13.8) | 79.2 % (68.4 %–86.9 %) | 91.5 % (91.2 % − 91.9 %) |
| Colorectal Cancer Screening [46] | Feces | Fecal hemoglobin and the methylated DNA markers (LASS4, LRRC4, PPP2R5C, ZDHHC1) | 43.4 % (41.3 %–45.6 %) | 93.9 % (87.1 %–97.7 %) | 90.6 % (90.1 % − 91.0 %) |
| Fecal hemoglobin | 23.3 % (21.5 %–25.2 %) | 67.3 % (57.1 %–76.5 %) | 94.8 % (94.4 % − 95.1 %) | ||
| ColoSense [35] | Feces | Fecal immunochemical test (FIT), concentration of 8 mRNA targets, and participant-reported smoking status | 45.9 % (42 %–50 %) | 94.4 % (81 %–99 %) | 87.9 % (87 % − 89 %) |
| ColonAiQ [31] | Blood | 6 differentially methylated regions (SEPT9, BCAT1, IKZF1, BCAN, VAV3, SEOT9 region2) | 42.1 % (33.1 %–51.5 %) | 86.1 % | 91.9 % |
| Colorectal cancer screening [32] | Blood | cfDNA genomic alterations, aberrant methylation status, and fragmentomic patterns | 13.2 % (11.3 %–15.3 %) | 83.1 % (72.2 %–90.3 %) | 89.6 % (88.8 % − 90.3 %) |
| ColoClear [34] | Feces | KRAS mutation, aberrant methylation of the BMP3 and NDRG4 promoter region, ACTB and B2M (DNA quantity), hemoglobin immunoassay | 63.5 % (58.3 %–68.3 %) | 91.9 % (86.8 %–95.3 %) | 90.3 % (88.7 %–91.7 %) |
| Gastric Cancer Screening [33] | serum | The levels of serum EFNA1 and MMP13 | 55.7 % | 90.2 % | |
| GutSeer [17] (gastrointestinal cancers screening) | Blood | 1,656 methylation markers and fragmentomics features | 47.1 % for colorectal | 81.5 % (77.1 %−85.9 %) | 94.4 % (92.4 %−96.5 %) |
| 38.9 % for esophageal | |||||
| 21.4 % for gastric |
CSO prediction
Accurate origin prediction is essential for multi-cancer detection [27]. The CSO (Cancer Signal Origin) model, an integral component of the dual-model architecture (SPOGIT/CSO), complements cancer detection with precise tissue localization using methylation profiles from a 230-DMR gastrointestinal origin-tracing panel. Demonstrating prediction accuracy for colorectal cancer (CRC) of 100 %, 100 %, and 83 % in the training, validation, and external independent validation cohorts, respectively. For gastric cancer (GC), the accuracy rates were 100 %, 83 %, and 71 %, respectively. The reduced accuracy in Cancer Signal Origin (CSO) for GC may be attributed to the presence of highly heterogeneous samples, including signet-ring cell carcinoma other poorly differentiated carcinoma (Table S1), within the external validation cohort. (Fig. 4C). The CSO model correctly predicted the origin of the majority of GI samples, with a probability exceeding 50 % (Fig. 4D-E). Furthermore, the model showed higher accuracy for GC in stages II-IV compared to stages 0-I, while CRC staging did not affect the accuracy of origin prediction (Fig. 4F-G, Table S4). This hierarchical prediction capability synergizes with SPOGIT's detection sensitivity, enabling clinicians to prioritize diagnostic pathways—particularly valuable given SPOGIT's 88 % pan-GI cancer sensitivity in external validation. The integrated framework thus bridges cancer presence detection with actionable anatomical localization, addressing a critical gap in multi-cancer screening paradigms. Fig. 5 was showed the general clinical application workflow for the SPOGIT/CSO model. For cost-effectiveness, we recommend screening for all individuals aged 40–75. Individuals who test negative with the SPOGIT analysis should continue with routine screening. For those who tested positive, the CSO model was used to predict the origin of tumor, which guided the subsequent referral for either a gastroscopy or a colonoscopy.
Fig. 5.
The Proposed clinical workflow for the SPOGIT/CSO test. A blood sample from an eligible individual is analyzed by the SPOGIT model. A negative result leads to a recommendation for routine screening. A positive result triggers the CSO model to predict the tumor origin from the same data, directly guiding a targeted referral for either gastroscopy or colonoscopy.
The interception model evaluates the potential clinical benefits of SPOGIT/CSO
The interception model was used to estimate stage shift and 5-year survival rates for the SPOGIT/CSO model. Under usual care, 57.1 % of gastrointestinal (GI) cancers would be diagnosed at late stages (III/IV), while SPOGIT/CSO screening could reduce this proportion to 4.4 % (Fig. 6A, Table 3). For patients initially diagnosed at late stages under usual care, the diagnosis model could identify 92 %-100 % of these cases at early stages (I/II) across scenarios ranging from slow to aggressive tumor growth (Fig. 6B, Table 3). Additionally, SPOGIT/CSO achieved a stage shift of 82 %–99 % from stages II-IV to stage I (Table 3). As a result, the 5-year survival rate for GI cancer patients improved by 27.02 %–30.47 % (Fig. 6C, Table 3). Fig. 6D further illustrates the impact of stage shift on mortality. The intercepted cancer incidence included a significant portion of cases that would have been diagnosed at late stages, most of which were shifted to early stages, leading to improved mortality rates (Fig. 6D). Furthermore, participation rates significantly influenced the clinical benefits of the model. Fig. 6E–F show that as screening participation rates increased, the diagnosis of late-stage cancers (III/IV) and cancer mortality decreased more significantly (Fig. 6E–F). The potential clinical benefits of SPOGIT/CSO in gastric and colorectal cancers were similar to the overall results, significantly increasing the stage shift from stages III-IV to I-II and improving the 5-year survival rates for gastric and colorectal cancer patients (Fig. S7-S8, Table S5-S6). These projections position SPOGIT/CSO not merely as a diagnostic tool but as a population-level intervention capable of fundamentally altering GI cancer trajectories through preclinical interception.
Fig. 6.
Potential Clinical Benefits of the SPOGIT Model. (A–C) Using the incidence and survival data of gastrointestinal cancer in China along with the SPOGIT diagnostic model to construct an interception model, estimating the stage distribution (A), stage migration (B) and 5-year survival rate (C) for three dwell time scenarios; (D) Corresponding stages and clinical outcomes of gastrointestinal cancer before and after interception; (E–F) Changes in the reduction of late-stage diagnoses (E) and mortality (F) per 100,000 screenable individuals for gastrointestinal tumors as screening participation rates increase. SPOGIT, screening for the presence of gastrointestinal tumor.
Table 3.
Estimated stage shift and 5-year survival rate of gastrointestinal tumor.
| GI cancers |
SPOGIT |
|||
|---|---|---|---|---|
| Cancer growth scenario | Slow | Fast | AggFast | |
| Incidence (per 100 K) pre | ||||
| Stage I | 10.95 | |||
| Stage II | 15.66 | |||
| Stage III | 23.27 | |||
| Stage IV | 12.14 | |||
| Incidence (per 100 K) post | ||||
| Stage I | 61.52 | 59.01 | 52.65 | |
| Stage II | 0.47 | 2.76 | 6.62 | |
| Stage III | 0.00 | 0.19 | 2.32 | |
| Stage IV | 0.00 | 0.03 | 0.40 | |
| Stage shift from III/IV to I/II | 1.00 | 0.99 | 0.92 | |
| Stage shift from II − IV to I | 0.99 | 0.94 | 0.82 | |
| 5-year survival rate (%) | ||||
| Pre | 60.15 | 60.15 | 60.15 | |
| Post | 90.63 | 89.99 | 87.18 | |
| Increase (%) | 30.47 | 29.83 | 27.02 | |
Discussion
The scarcity and heterogeneity of cfDNA from early-stage cancers limit the sensitivity and consistency of current MCED (multi-cancer early detection) tests, creating a major barrier to effective screening [28,29]. Breaking through the barriers of early detection is a critical objective in this domain. Artificial intelligence (AI) provides a promising solution, with the potential to enhance detection performance by improving biomarker selection, optimizing library conversion efficiency, and constructing complex, multi-dimensional models. In our research, we performed targeted methylation sequencing of training-set plasma samples using Twist probes covering >4 million CpG sites, combined with large-scale public tissue methylation datasets, followed by comprehensive data mining with multiple machine learning algorithms to ensure optimal inclusion of differential methylation loci. This approach ultimately identified 311 gastrointestinal cancer-specific DNA methylation biomarkers, which were integrated with six optimized machine learning algorithms to develop the SPOGIT model. In a large-scale external validation involving 1079 clinical samples (46 % precancerous lesions), the model demonstrated robust performance, achieving 88.1 % sensitivity for gastrointestinal malignancies, 56.5 % for AA, and 62.4 % for high risk pre-GC, with a specificity of 91.2 %.
SPOGIT stood out among blood-based tests, demonstrating significant advantages in detecting both gastrointestinal cancer (88.1 % vs. 55.7 %-86.1 %) and, critically, precancerous lesions (58 % vs. 12.5 %–47.1 %) [17,[30], [31], [32], [33]] (Table 3). Its comparison with top-tier fecal tests revealed a complex landscape of trade-offs. While ColoClear reported slightly higher sensitivity for cancer (91.9 %) and precancerous lesions (63.5 %), this result was largely attributable to its study cohort being enriched with pre-screened, high-risk individuals [34]. The data of SPOGIT, derived from a general population, was thus more representative for widespread screening. Furthermore, SPOGIT offered higher specificity than the multi-target fecal test ColoSense (91.2 % vs. 87.9 %), implying a lower false-positive rate at the cost of slightly lower cancer sensitivity (88.1 % vs. 94.4 %) [35]. However, these nuanced performance metrics might be outweighed by the decisive real-world advantage of blood tests: patient adherence. The convenience of integrating sample collection into routine clinical visits, which bypassed the user aversion associated with fecal collection, had been shown to increase screening participation by 11 percentage points in a randomized clinical trial of individuals overdue for screening [36,37]. The long-term benefited driven by such higher screening adherence could potentially compensate for a minor sensitivity gap, ultimately leading to superior cancer prevention and control outcomes in practical application [38,39].
The SPOGIT model demonstrated significantly improved discriminative capacity for gastrointestinal malignancies relative to conventional serum biomarkers, achieving an AUC of 0.944 (95 % CI: 0.922–0.965) compared to CEA (0.648; 95 % CI: 0.582–0.714), CA19-9 (0.569; 95 % CI: 0.492–0.646), and CA125 (0.782; 95 % CI: 0.720–0.844). Decision curve analysis revealed superior clinical utility, with SPOGIT providing greater net benefit across the clinically relevant 20–60 % probability threshold range. These findings position SPOGIT as both a diagnostically superior screening tool and a clinically actionable decision-making aid for endoscopic referral in gastrointestinal cancer detection. This diagnostic capability was complemented by the CSO model's tissue-of-origin prediction, which achieved 83 % accuracy for CRC and 71 % for GC localization, with performance enhancement in advanced gastric cancers (74.2 % vs. 67.6 % in early-stage) reflecting stage-dependent methylation pattern consolidation.
Tissue-of-origin analysis is critical for multi-cancer early detection, as the accuracy of tracing directly impacts the alignment of subsequent diagnostic procedures [[40], [41], [42]]. The CSO component amplified this impact by enabling targeted diagnostic pathways, thereby reducing unnecessary procedures and accelerating treatment initiation. In this study, we identified 230 DMRs specific to gastrointestinal tissues and plasma from GI cancer patients and developed the CSO model to predict the cancer signal origin. In an external independent validation cohort, the origin-tracing accuracy of the CSO model was significantly lower in GC than in CRC (71 % vs. 83 %). We speculated that this performance difference stemmed primarily from disease stage and tumor heterogeneity. On the one hand, tumor origin prediction accuracy typically increases with more advanced disease stages. Consistent with this, our model showed higher accuracy in stage II–IV GC compared to stage 0–I GC, whereas no such staging effect was observed in CRC. Given the higher proportion of early-stage cases in the GC cohort than in the CRC cohort (29.4 % vs. 22.2 %; Table S1), this factor likely increased the difficulty of origin tracing for gastric cancer. On the other hand, the inherent high heterogeneity of gastric cancer itself posed another key challenge. Our cohort included pathologically complex subtypes such as signet ring cell carcinoma and mixed-type gastric cancer, whose variable molecular features present a severe test to the tracing capabilities of model. To address the complexity of gastric cancer, future work should focus on deeper molecular profiling. This will require using high-resolution technologies (such as single-cell and spatial omics) to perform in-depth analysis of the roots of its heterogeneity, coupled with the use of liquid biopsies for non-invasive, dynamic tracking to monitor its evolution during treatment. Ultimately, the integration and prediction of this vast amount of data using artificial intelligence will guide innovative diagnostic and therapeutic strategies, thereby enabling the precise, personalized management of this highly heterogeneous tumor.
The integrated clinical value of this dual architecture manifested most profoundly in mortality reduction. Simulating the application of our SPOGIT model in real-world settings indicated a significant shift toward earlier-stage detection in both gastric and colorectal cancers. This shift was projected to enhance 5-year survival rates by approximately 40 % for gastric cancer and 22 % for colorectal cancer following curative surgery (Table S5-S6). This clinical benefit further translated into considerable public health value, centered on strong potential of the method for cost-effective, large-scale population screening. Its cost-effectiveness was manifested on two levels. First, as a convenient and non-invasive initial screening tool, it could effectively triage the population, precisely allocating expensive endoscopic resources to high-risk individuals who test positive. Second, by significantly shifting diagnoses from the costly late stages to curable early stages, it could fundamentally alleviate the immense economic burden of cancer, offering particular benefits for resource-limited regions. Underpinning this potential was the model's design for large-scale application, it required only 10 ml of blood, a reduced cfDNA input (<30 ng), and fewer DMRs. These streamlined technical requirements were not just a matter of feasibility but also for lowering costs and simplifying logistics for widespread implementation [[43], [44], [45]]. the SPOGIT diagnostic model demonstrated substantial advantages for the early detection of GI cancers and hold promise as a novel screening approach.
Despite these promising results, several limitations of our study had to be acknowledged. First, while the SPOGIT model demonstrated strong performance in detecting GI cancers, the accuracy of its accompanying CSO model for GC remained inadequate and required further improvement. Future research will aim to integrate high-resolution technologies, such as dynamic multi-omics monitoring, to further enhance the overall robustness of cfDNA-based diagnostic and origin-tracing models. Second, the model's positive predictive value (PPV), false positive rate (FPR), and long-term impact on clinical interventions in an asymptomatic, average-risk population have yet to be definitively established through large-scale, prospective screening cohort studies. Third, as a potential screening tool, its optimal clinical implementation strategy remains to be explored, including its ideal testing frequency and a comprehensive cost-effectiveness analysis. Future modeling studies will need to comprehensively consider multiple factors, such as test performance, natural disease progression, patient adherence, and healthcare costs, to provide an evidence-based recommendation for the optimal screening interval.
Translating a liquid biopsy technology like the SPOGIT/CSO paradigm into clinical screening applications presents significant ethical and regulatory challenges. Ethically, the core task is managing the impact of result uncertainty on asymptomatic individuals. This requires ensuring participants fully comprehend the risks of false positives and negatives, while a clear, efficient management pathway for positive results is crucial to minimize patient anxiety. Issues of overdiagnosis and biodata privacy must also be addressed. Concurrently, establishing standardized operating procedures (SOPs) and stringent quality control (QC) is a prerequisite for the reliable, reproducible, and safe implementation of the test for widespread use.
In conclusion, we developed two robust and sensitive models based on targeted cfDNA methylation detection, the SPOGIT model for early detection of GI cancers and the CSO model for predicting tumor signal origin. The SPOGIT/CSO architecture effectively bridged the critical gap between initial detection and subsequent clinical action by providing high GI-specific sensitivity and guiding definitive diagnostic referrals. This paradigm held the potential to significantly reduce late-stage cancer incidence and mortality through early interception, thereby setting a new standard for precision screening. Realizing this potential, however, requires a rigorous translational journey. This includes standardizing the assay for clinical-grade robustness and designing pivotal trials in collaboration with regulatory authorities. Ultimately, the definitive clinical utility of this paradigm must be established in a large-scale, prospective screening study of an average-risk population, with a demonstrable reduction in late-stage cancer incidence and mortality as primary endpoints. Through this validation pathway, the SPOGIT/CSO approach is poised to transform the landscape of gastrointestinal cancer screening and deliver substantial public health benefits.
Data availability statement
All of the data supporting this work will be made available from the corresponding author upon reasonable request.
CRediT authorship contribution statement
Lingqin Zhu: Conceptualization, Formal analysis, Investigation, Methodology, Writing – original draft. Shuye Lin: Conceptualization, Formal analysis, Investigation, Methodology. Junfeng Xu: Conceptualization, Formal analysis, Investigation, Methodology. Jianwei Yu: Investigation. Fangli Men: Investigation. Dongliang Yu: Investigation. Xianzong Ma: Investigation. Ju Tian: Investigation. Hui Xie: Investigation. Linghui Duan: Investigation. Xin Wang: Investigation. Shuyang Sun: Methodology. Chenguang Li: Methodology. Shu Li: Methodology. Qian Kang: Methodology. Mengyu Jia: Methodology. Xueqin Lin: Methodology. Qiaoqiao Lin: Methodology. Lijuan Lin: Methodology. Xiang Yi: Formal analysis. Ruiru Wang: Formal analysis. Wei Guo: Formal analysis. Xueqing Gong: Formal analysis. Jianqiu Sheng: Conceptualization, Formal analysis, Methodology, Supervision. Ni Guo: Conceptualization, Formal analysis, Methodology. Shiqian Lan: Conceptualization, Formal analysis, Methodology, Supervision. Peng Jin: Conceptualization, Formal analysis, Methodology. Yuqi He: Conceptualization, Formal analysis, Funding acquisition, Methodology, Supervision, Writing – review & editing.
Ethics approval and consent to participate
The study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of the Seventh Medical Center of PLA General Hospital (Approval No. 2016-70, 2020-78), Dongying People's Hospital (Approval No. DYYW-2019-002-01), and the First Hospital of Longyan, Fujian Medical University (Approval No. 2021-k0001). Informed consents were obtained from all participants involved in the study.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgements
This research was funded by grants from the National Natural Science Foundation of China (Grant no. 82273245), Beijing Hospitals Authority Clinical Medicine Development of special funding support (Grant no. ZLRK202531), the Capital's Funds for Health Improvement and Research (Grant no. 2022-1-5082), the Beijing Tongzhou District Science and Technology Program (Grant no. KJ2024CX042) and Beijing Hospitals Authority Youth Programme (QML20231602). Sponsored by Dongying City Natural Science Foundation (Grant no. 2023ZR026). Sponsored by Longyan City Science and Technology Plan Project (Grant number: 2023LYF17024). The authors acknowledge the use of Biorender that is used to create Figure 1.
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.jare.2025.10.038.
Contributor Information
Ni Guo, Email: guoni1974@sina.com.
Shiqian Lan, Email: lan4000@163.com.
Peng Jin, Email: jinpeng@301hospital.com.cn.
Yuqi He, Email: endohe@163.com.
Appendix A. Supplementary material
The following are the Supplementary data to this article:
References
- 1.Bray F., Laversanne M., Sung H., et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–263. doi: 10.3322/caac.21834. [DOI] [PubMed] [Google Scholar]
- 2.Wang S., Zheng R., Li J., et al. Global, regional, and national lifetime risks of developing and dying from gastrointestinal cancers in 185 countries: a population-based systematic analysis of GLOBOCAN. Lancet Gastroenterol Hepatol. 2024;9(3):229–237. doi: 10.1016/S2468-1253(23)00366-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Abdel Hamid M., Pammer L.M., Oberparleiter S., et al. Multidimensional differences of right- and left-sided colorectal cancer and their impact on targeted therapies. npj Precis Oncol. 2025;9(1):116. doi: 10.1038/s41698-025-00892-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hayman C.V., Vyas D. Screening colonoscopy: the present and the future. World J Gastroenterol. 2021;27(3):233–239. doi: 10.3748/wjg.v27.i3.233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Bretthauer M., Løberg M., Wieszczy P., et al. Effect of colonoscopy screening on risks of colorectal cancer and related death. N Engl J Med. 2022;387(17):1547–1556. doi: 10.1056/NEJMoa2208375. [DOI] [PubMed] [Google Scholar]
- 6.Gómez-Molina R., Suárez M., Martínez R., Chilet M., Bauça J.M., Mateo J. Utility of stool-based tests for colorectal cancer detection. A comprehensive review. Healthcare (Basel, Switzerland) 2024;12(16) doi: 10.3390/healthcare12161645. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Liang J.Q., Li T., Nakatsu G., et al. A novel faecal Lachnoclostridium marker for the non-invasive diagnosis of colorectal adenoma and cancer. Gut. 2020;69(7):1248–1257. doi: 10.1136/gutjnl-2019-318532. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Sun Q., Long L. Diagnostic performances of methylated septin9 gene, CEA, CA19-9 and platelet-to-lymphocyte ratio in colorectal cancer. BMC Cancer. 2024;24(1):906. doi: 10.1186/s12885-024-12670-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Malla M., Loree J.M., Kasi P.M., Parikh A.R. Using circulating tumor DNA in colorectal cancer: current and evolving practices. J Clin Oncol. 2022;40(24):2846–2857. doi: 10.1200/JCO.21.02615. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zhou H., Zhu L., Song J., et al. Liquid biopsy at the frontier of detection, prognosis and progression monitoring in colorectal cancer. Mol Cancer. 2022;21(1):86. doi: 10.1186/s12943-022-01556-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Schwarzenbach H., Hoon D.S., Pantel K. Cell-free nucleic acids as biomarkers in cancer patients. Nat Rev Cancer. 2011;11(6):426–437. doi: 10.1038/nrc3066. [DOI] [PubMed] [Google Scholar]
- 12.Ren J., Lu P., Zhou X., et al. Genome-scale methylation analysis of circulating cell-free DNA in gastric cancer patients. Clin Chem. 2022;68(2):354–364. doi: 10.1093/clinchem/hvab204. [DOI] [PubMed] [Google Scholar]
- 13.Zhao F., Bai P., Xu J., et al. Efficacy of cell-free DNA methylation-based blood test for colorectal cancer screening in high-risk population: a prospective cohort study. Mol Cancer. 2023;22(1):157. doi: 10.1186/s12943-023-01866-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Liu M.C., Oxnard G.R., Klein E.A., Swanton C., Seiden M.V. Sensitive and specific multi-cancer detection and localization using methylation signatures in cell-free DNA. Ann Oncol. 2020;31(6):745–759. doi: 10.1016/j.annonc.2020.02.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Gao Q., Lin Y.P., Li B.S., et al. Unintrusive multi-cancer detection by circulating cell-free DNA methylation sequencing (THUNDER): development and independent validation studies. Ann Oncol. 2023;34(5):486–495. doi: 10.1016/j.annonc.2023.02.010. [DOI] [PubMed] [Google Scholar]
- 16.Klein E.A., Richards D., Cohn A., et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann Oncol. 2021;32(9):1167–1177. doi: 10.1016/j.annonc.2021.05.806. [DOI] [PubMed] [Google Scholar]
- 17.Huang A., Guo D.Z., Su Z.X., et al. GUIDE: a prospective cohort study for blood-based early detection of gastrointestinal cancers using targeted DNA methylation and fragmentomics sequencing. Mol Cancer. 2025;24(1):163. doi: 10.1186/s12943-025-02367-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Nicholson B.D., Oke J., Virdee P.S., et al. Multi-cancer early detection test in symptomatic patients referred for cancer investigation in England and Wales (SYMPLIFY): a large-scale, observational cohort study. Lancet Oncol. 2023;24(7):733–743. doi: 10.1016/S1470-2045(23)00277-2. [DOI] [PubMed] [Google Scholar]
- 19.Sullivan B.A., Lieberman D.A. Colon polyp surveillance: separating the wheat from the chaff. Gastroenterology. 2024;166(5):743–757. doi: 10.1053/j.gastro.2023.11.305. [DOI] [PubMed] [Google Scholar]
- 20.Pimentel-Nunes P., Libânio D., Marcos-Pinto R., et al. Management of epithelial precancerous conditions and lesions in the stomach (MAPS II): European Society of Gastrointestinal Endoscopy (ESGE), European Helicobacter and Microbiota Study Group (EHMSG), European Society of Pathology (ESP), and Sociedade Portuguesa de Endoscopia Digestiva (SPED) guideline update 2019. Endoscopy. 2019;51(4):365–388. doi: 10.1055/a-0859-1883. [DOI] [PubMed] [Google Scholar]
- 21.Wang H., Fu T., Du Y., et al. Scientific discovery in the age of artificial intelligence. Nature. 2023;620(7972):47–60. doi: 10.1038/s41586-023-06221-2. [DOI] [PubMed] [Google Scholar]
- 22.Han B., Zheng R., Zeng H., et al. Cancer incidence and mortality in China, 2022. J Natl Cancer Center. 2024;4(1):47–53. doi: 10.1016/j.jncc.2024.01.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zeng H., Ran X., An L., et al. Disparities in stage at diagnosis for five common cancers in China: a multicentre, hospital-based, observational study. Lancet Public Health. 2021;6(12):e877–e887. doi: 10.1016/S2468-2667(21)00157-2. [DOI] [PubMed] [Google Scholar]
- 24.Li H., Zhang H., Zhang H., Wang Y., Wang X., Hou H. Survival of gastric cancer in China from 2000 to 2022: a nationwide systematic review of hospital-based studies. J Glob Health. 2022;12:11014. doi: 10.7189/jogh.12.11014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Zeng H., Chen W., Zheng R., et al. Changing cancer survival in China during 2003-15: a pooled analysis of 17 population-based cancer registries. Lancet Glob Health. 2018;6(5):e555–e567. doi: 10.1016/S2214-109X(18)30127-X. [DOI] [PubMed] [Google Scholar]
- 26.Hubbell E., Clarke C.A., Aravanis A.M., Berg C.D. Modeled reductions in late-stage cancer with a multi-cancer early detection test. Cancer Epidemiol Biomarkers Prev. 2021;30(3):460–468. doi: 10.1158/1055-9965.EPI-20-1134. [DOI] [PubMed] [Google Scholar]
- 27.van der Pol Y., Mouliere F. Toward the early detection of cancer by decoding the epigenetic and environmental fingerprints of cell-free DNA. Cancer Cell. 2019;36(4):350–368. doi: 10.1016/j.ccell.2019.09.003. [DOI] [PubMed] [Google Scholar]
- 28.Hanahan D. Hallmarks of cancer: new dimensions. Cancer Discov. 2022;12(1):31–46. doi: 10.1158/2159-8290.CD-21-1059. [DOI] [PubMed] [Google Scholar]
- 29.Liu J., Dai L., Wang Q., et al. Multimodal analysis of cfDNA methylomes for early detecting esophageal squamous cell carcinoma and precancerous lesions. Nat Commun. 2024;15(1):3700. doi: 10.1038/s41467-024-47886-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Shaukat A., Burke C.A., Chan A.T., et al. Clinical validation of a circulating tumor DNA-based blood test to screen for colorectal cancer. J Am Med Assoc. 2025;334(1):56–63. doi: 10.1001/jama.2025.7515. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Cai G., Cai M., Feng Z., et al. A multilocus blood-based assay targeting circulating tumor DNA methylation enables early detection and early relapse prediction of colorectal cancer. Gastroenterology. 2021;161(6):2053–2056.e2052. doi: 10.1053/j.gastro.2021.08.054. [DOI] [PubMed] [Google Scholar]
- 32.Chung D.C., Gray D.M., 2nd, Singh H., et al. A cell-free DNA blood-based test for colorectal cancer screening. N Engl J Med. 2024;390(11):973–983. doi: 10.1056/NEJMoa2304714. [DOI] [PubMed] [Google Scholar]
- 33.Chu L.Y., Wu F.C., Guo H.P., et al. Combined detection of serum EFNA1 and MMP13 as diagnostic biomarker for gastric cancer. Sci Rep. 2024;14(1):15957. doi: 10.1038/s41598-024-65839-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Hu Y.T., Chen X.F., Zhai C.B., et al. Clinical evaluation of a multitarget fecal immunochemical test-sDNA test for colorectal cancer screening in a high-risk population: a prospective, multicenter clinical study. MedComm. 2023;4(4):e345. doi: 10.1002/mco2.345. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Barnell E.K., Wurtzler E.M., La Rocca J., et al. Multitarget stool RNA test for colorectal cancer screening. J Am Med Assoc. 2023;330(18):1760–1768. doi: 10.1001/jama.2023.22231. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Shaukat A., Levin T.R. Current and future colorectal cancer screening strategies. Nat Rev Gastroenterol Hepatol. 2022;19(8):521–531. doi: 10.1038/s41575-022-00612-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Liles E.G., Coronado G.D., Perrin N., et al. Uptake of a colorectal cancer screening blood test is higher than of a fecal test offered in clinic: a randomized trial. Cancer Treat Res Commun. 2017;10:27–31. [Google Scholar]
- 38.Meester RG, Piscitello A, Baldo L, Liang PSJG. 539 Higher assumed adherence to blood-based vs stool-based colorectal cancer screening compensates for potential lower advanced adenoma sensitivity. 2024;166(5):S-122.
- 39.Mannucci A., Goel A. Stool and blood DNA tests for colorectal cancer screening. N Engl J Med. 2024;390(23):2224. doi: 10.1056/NEJMc2404924. [DOI] [PubMed] [Google Scholar]
- 40.Mattox A.K., Douville C., Wang Y., et al. The origin of highly elevated cell-free DNA in healthy individuals and patients with pancreatic, colorectal, lung, or ovarian cancer. Cancer Discov. 2023;13(10):2166–2179. doi: 10.1158/2159-8290.CD-21-1252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Conway A.M., Pearce S.P., Clipson A., et al. A cfDNA methylation-based tissue-of-origin classifier for cancers of unknown primary. Nat Commun. 2024;15(1):3292. doi: 10.1038/s41467-024-47195-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Nguyen T.H., Doan N.N.T., Tran T.H., et al. Tissue of origin detection for cancer tumor using low-depth cfDNA samples through combination of tumor-specific methylation atlas and genome-wide methylation density in graph convolutional neural networks. J Transl Med. 2024;22(1):618. doi: 10.1186/s12967-024-05416-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Mo S., Ye L., Wang D., et al. Early detection of molecular residual disease and risk stratification for stage I to III colorectal cancer via circulating tumor DNA methylation. JAMA Oncol. 2023;9(6):770–778. doi: 10.1001/jamaoncol.2023.0425. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Mannucci A., Goel A. Stool and blood biomarkers for colorectal cancer management: an update on screening and disease monitoring. Mol Cancer. 2024;23(1):259. doi: 10.1186/s12943-024-02174-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Kandimalla R., Xu J., Link A., et al. EpiPanGI Dx: a cell-free DNA methylation fingerprint for the early detection of gastrointestinal cancers. Clin Cancer Res. 2021;27(22):6135–6144. doi: 10.1158/1078-0432.CCR-21-1982. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Imperiale T.F., Porter K., Zella J., et al. Next-generation multitarget stool DNA test for colorectal cancer screening. N Engl J Med. 2024;390(11):984–993. doi: 10.1056/NEJMoa2310336. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All of the data supporting this work will be made available from the corresponding author upon reasonable request.







