Skip to main content
Communications Medicine logoLink to Communications Medicine
. 2026 Feb 17;6:166. doi: 10.1038/s43856-026-01437-5

Machine learning–based cfDNA fragmentation profiling using automated capillary electrophoresis for early detection of hepatocellular carcinoma

Sasimol Udomruk 1, Songphon Sutthitthasakul 1,2, Nuttida Bunsermvicha 3, Kanokwan Pinyopornpanish 4, Dumnoensun Pruksakorn 1,5, Phasit Charoenkwan 6, Petlada Yongpitakwattana 1, Kanlaya Khounkaew 1, Thanapak Jaimalai 1, Treephum Duangsan 7, Santhasiri Orrapin 1, Sutpirat Moonmuang 1, Pitiporn Noisagul 1, Arnat Pasena 1, Pathacha Suksakit 1, Ratikorn Gamngoen 1, Pimpisa Teeyakasem 1, Chaiyut Charoentum 4, Sarawut Kongkarnka 8, Worakitti Lapisatepun 3,✉, Parunya Chaiyawat 1,✉
PMCID: PMC13022191  PMID: 41703285

Abstract

Background

Early detection of hepatocellular carcinoma (HCC) remains a significant clinical challenge due to the limited sensitivity of current surveillance tools, alpha-fetoprotein (AFP) and ultrasound. Recently, cell-free DNA (cfDNA) fragmentation analysis has shown promise in cancer detection; however, current sequencing-based approaches remain costly and unsuitable for large-scale screening.

Methods

Here, we introduce a predictive model for early HCC detection called “CEliver” (CfDNA-based automated capillary Electrophoresis method for Liver cancer screening), a model leveraging high-dimensional fragmentation profiling from the intensity distribution of cfDNA fragment lengths using automated capillary electrophoresis. We developed CF-2D features, a computational framework that extracts over 300 quantitative features from electropherogram data, including cfDNA concentration, dominant fragment sizes, two-dimensional shape descriptors, and short-to-long fragment ratios. We integrated these features with AFP levels to build the CEliver model, developed in 111 individuals and validated in an independent cohort of 69 subjects.

Results

Here we show the CF-2D profiles differ significantly between HCC patients and high-risk individuals. The CEliver model achieves 98% sensitivity across all HCC cases, and 96% sensitivity with 99% specificity for early-stage HCC (stage 0/A), substantially outperforming AFP (60% overall sensitivity, 35% for early-stage). In external validation, CEliver shows 88% sensitivity and 100% specificity.

Conclusions

CEliver provides a practical and accurate strategy for early HCC detection. By enabling high-dimensional cfDNA fragmentomics analysis on a widely accessible electrophoresis platform, it bridges the gap between research-grade cfDNA technologies and real-world clinical implementation. This method represents a simple and scalable approach that could potentially be applied in HCC surveillance.

graphic file with name 43856_2026_1437_Figa_HTML.jpg

Subject terms: Diagnostic markers, Cancer screening


Udomruk et al. develop CEliver, a machine learning model that analyzes high-dimensional cfDNA fragmentation profiles using automated capillary electrophoresis for early Hepatocellular Carcinoma (HCC) detection. The model identifies early-stage HCC with higher accuracy than standard AFP-based screening.

Plain language summary

Hepatocellular carcinoma (HCC) is a liver cancer that is often diagnosed too late for effective treatment to be used. Here, we developed a computational model to detect HCC early called “CEliver”. The method analyses DNA fragments found in the blood. We developed the model using data from over 100 patients. CEliver accurately distinguishes early-stage HCC from people at risk of disease and performs better than existing diagnosis methods. This simple and scalable approach could be applied for large-scale population screening, helping at-risk people receive earlier diagnosis and treatment, potentially improving survival outcomes.

Introduction

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with rising incidence1,2. Prognosis depends on tumor stage at diagnosis, with early-stage detection achieving a 5-year survival rate of 70–80%3. Over 80% of HCC patients have underlying liver diseases, particularly hepatitis B virus (HBV)4. In 2019, an estimated 296 million people had chronic HBV infection, which resulted in a global public health burden. In Thailand, 50% of HCC cases are attributed to chronic HBV infection, reflecting a critical health concern5. This high-risk group is the primary target for surveillance6.

Guidelines recommend HCC screening every 6 months using ultrasound (US) ± serum alpha-fetoprotein (AFP). Evidence from randomized controlled trials in high-risk Chinese populations show screening improves early detection by 60% and reduces mortality7. However, current methods have limitations. Serum AFP has poor sensitivity ( < 40%) for early-stage HCC8. Ultrasound has high specificity ( > 90%) but low sensitivity (21–47%) for tumors <2 cm, especially in cirrhotic livers9. Combining AFP and US modestly improves sensitivity to 67%10. These limitations highlight the need for more sensitive early detection methods, particularly in HCC screening clinics.

Cell-free DNA (cfDNA) analysis has emerged as a promising approach for cancer detection11–13. Various genetic and non-genetic cfDNA characteristics allow precise cancer identification14. Among these, cfDNA fragment size analysis has been extensively developed for early detection. In cancer patients, cfDNA fragmentation patterns differ due to changes in nucleosome regulation15,16 and DNase enzymatic activity17. Typically, cfDNA peaks at 167 bp18, while cancer-derived cfDNA is shorter, with an increased proportion of small fragments19. Studies confirm that HCC patients have shorter cfDNA fragments than healthy individuals. A genome-wide cfDNA fragmentomics model achieved 88% sensitivity and 98% specificity for early-stage HCC detection20–23. Additionally, enriching cfDNA fragments shorter than 150 bp improves tumor fraction detection22, further supporting the significance of shortened cfDNA fragment size in HCC.

Several hypotheses explain cfDNA shortening in HCC. One suggests epigenetic alterations influence fragment size, as hypomethylation leads to loosely packed DNA, making it more susceptible to nucleases, producing shorter fragments24. Another factor is DNase enzyme dysfunction. Deficiency in DNASE1L3 (Deoxyribonuclease 1-Like 3) is associated with an increase in short cfDNA fragments in both mice and humans25. Notably, DNASE1L3 expression and activity are reduced in HCC patients26. These processes might contribute to the presence of short cfDNA fragments in the plasma of HCC patients.

Various cfDNA analysis platforms based on fragmentomics enhance early cancer detection. Lapin M. used automated capillary electrophoresis to identify cfDNA fragment size patterns in pancreatic cancer patients27. Cristiano et al. developed a high-resolution sequencing platform for genome-wide cfDNA fragmentation analysis, revealing diverse fragment size profiles in cancer patients and strong early detection performance in the early detection of several cancer types28. Furthermore, studies confirm the ability of genome-wide fragmentation models to distinguish early-stage HCC from high-risk individuals21,29, Recently, the Fragle model, a deep learning model generated from the intensity of cfDNA fragment size distribution using low pass whole genome sequencing, demonstrated high performance in multi-cancer detection including liver cancer30. This evidence highlights cfDNA fragment size as a potential key feature for HCC detection.

To date, cfDNA fragmentation analyses have relied on WGS, which is cost-prohibitive for large-scale screening31. Our method enables high-dimensional fragmentation profiling based on intensity distribution of cfDNA-fragment lengths using a commonly available platform, automated capillary electrophoresis. We developed the CEliver model (CfDNA-based automated capillary Electrophoresis method for Liver cancer screening), integrating cfDNA fragment features with clinical factors. Constructed using machine learning, the model demonstrates strong diagnostic performance in distinguishing early HCC from high-risk individuals. Overall, the CEliver model shows great potential for early HCC detection and prognosis prediction in high-risk groups.

Methods

Cohort design

This study included two cohorts. First, 71 high-risk individuals and 40 HCC patients were used for model development. Second, an independent cohort comprising 27 HCC patients, 30 high-risk donors, and 12 healthy donors was enrolled for external validation (Fig. 1). All samples were collected prospectively during 2022–2024 at Maharaj Nakorn Chiang Mai hospital under the Suandok Repository Unit (SRU) System, the official human biobank at Chiang Mai University (Ethics approval code ORT- 2563–07122). All research was conducted in accordance with the Declaration of Helsinki and approved by the Research Ethics Committee of Faculty of Medicine, Chiang Mai University (Ethics approval code ORT- 2563–07122). Written informed consent was obtained from all subjects.

Fig. 1. The model development and external validation.

Fig. 1

Model development includes 71 high-risk individuals and 40 HCC patients. cfDNA is extracted and fragment size profiles are analyzed using automated capillary electrophoresis. A total of 332 CF-2D features is extracted. Feature selection is performed using SelectKBest. The CEliver model is established using LightGBM, incorporating CF-2D features and AFP to distinguish HCC patients from high-risk individuals. The dataset is split into training and test sets (80:20). The model outputs a prediction score ranging from 0 to 1, representing the probability of HCC. Finally, external validation is performed using an independent cohort comprising 27 HCC cases, 30 high-risk individuals, and 12 healthy donors.

The diagnosed HCC patients were defined by the imaging test according to accepted standard AASLD (American Association for the Study of Liver Diseases) criteria or pathologic confirmation. The tumor stage was decided on BCLC staging guidelines version 202232. All high-risk patients were recruited from two HCC screening clinics, including those with chronic hepatitis B or C or liver cirrhosis. The inclusion criteria also required: 1) no history of cancer; 2) no benign diseases; and 3) being cancer-free for at least one year after the sample collection date. AFP levels were reported by routine clinical laboratories belonging to hospital. The acceptable clinical data were collected following the policy of the Data Governance Faculty of Medicine Chiang Mai University data repository (Ethics approval code FAC-MED-2565-09352).

cfDNA extraction

The 10 ml of blood samples were collected in K2-EDTA tube. Within 2 h after blood draw, plasma was separated by centrifugation at 1600 × g, 10 min. The aliquots of plasma were stored at 80 °C until processed for cfDNA extraction. Freeze–thaw cycles should be avoided. Before cfDNA extraction, plasma was centrifuged by high-speed centrifugation at 6000 × g for 10 min to remove leukocytes and cell debris. The 2 mL of plasma was then used for cfDNA extraction using the QIAamp Circulating Nucleic Acid Kit (Qiagen) with optimized manufacturer’s protocols. The QIAamp DNA Mini kit and the QIAamp DNA Blood Mini kit were used to extract matched tumor gDNA and blood DNA (PBMC), respectively.

Automated capillary electrophoresis

The quantitative cfDNA was determined using the QIAxcel Advanced System with a QIAxcel DNA High Resolution Kit, according to the manufacturer’s instructions. The cfDNA fragment size and concentration were analyzed by QIAxcel screengel software based on DNA marker 100 bp to 2.5 kb and the alignment marker 15 bp/3 kb. Importantly, the accuracy and precision of the running system were qualified using a positive fragment size control, 150 bp of double-strand DNA. Individual cfDNA samples were measured by an independent three-time analysis.

CF-2D features extraction

For model development, we generated 332 cfDNA fragmentation features called “CF-2D” from the electropherogram data. The electropherogram is a graph generated by automated capillary electrophoresis that shows the amount and size of cfDNA fragments, with peak positions indicating cfDNA fragment size and peak heights representing the fluorescence intensity of cfDNA (Supplementary Fig. 1). From these raw electropherogram data, we extracted quantitative fragmentation features for model inputs, including 1 feature of cfDNA Concentration, 1 feature of main cfDNA Fragment size, 20 two-dimensional (2D) intensity (size distribution) features, and 310 F-ratio features representing the ratio of short-to-long fragments. The main cfDNA fragment size was defined as the average dominant peak, while cfDNA concentration was calculated in ng/mL of plasma.

The 2D intensity features were established by dividing the cfDNA fragment size range (51–250 bp) into 20 bins of 10 base pairs each: 51–60, 61–70, 71–80,…, up to 241–250 bp. Each interval captured the signal intensity, thereby reflecting the correlation between fragment size and cfDNA abundance. To enhance the model’s accuracy and learning capability, 310 additional features, referred to as F-ratio features, were generated based on short-to-long cfDNA fragment ratio. This ratio is not fixed but varies depending on the binning scheme or interval settings applied during fragment size analysis.

The F-ratio features were derived from 20 bins of the 2D features set using three feature engineering methods, including 1) Sliding window fraction (19 features), 2) Probability fraction with regular dividing (252 features), 3) Probability fraction with irregular dividing (39 features).

First, we generated sliding window fraction features from the fragment intensity across 20 predefined bins as follows: fraction₁ = sum(2–20)⁻¹, fraction₂ = sum(1–2)/sum(3–20), fraction₃ = sum(1–3)/sum(4–20), fraction₄ = sum(1–4)/sum(5–20), …, and finally, fraction₁₉ = sum(1–19)/fraction₂₀, resulting in a total of 19 features. The sliding window fraction features were calculated as:

Fk=∑i=1kBi∑j=k+120Bj,k=1,2,…,19

This approach yielded a total of 19 features, where each Fk represents the ratio of cumulative intensity of shorter fragments (bins 1–k) to that of longer fragments (bins k + 1–20). Here, Bi represents the index of bins contributing to the cumulative sum of shorter fragments (i=1,…,20) and Bj represents the index of bins corresponding to longer fragments used in the denominator. This method allows the model to capture progressive patterns in fragment distribution, providing features that reflect the relative abundance of short versus long cfDNA fragments.

For regular dividing, 2D features fractions were grouped into 4 sets based on divisible numbers: 1, 2, 4, and 5. Within each set, the occurrence probability of the short-to-long ratio (S/L ratio), defined as the proportion of shorter fragments divided by longer fragments, was calculated, yielding 252 distinct features.

The irregular division sets were defined using 2D features groupings based on indivisible numbers, including 3, 6, 7, 8, and 9. A sliding window technique was applied to complete the short-to-long (S/L) fragment ratios across irregular sets. For example, in the set divided by 3, the S/L ratios were calculated using a sliding window across every three consecutive windows, staring form the first set and moving forward. The same approach was applied for sets divided by 6, 7, 8, and 9 windows, respectively. In total 39 S/L ratio features were generated using this irregular division method (Fig. 2).

Fig. 2. An explanation of CF-2D features generation.

Fig. 2

First, two-dimensional (2D) features are generated by dividing the cfDNA fragment size range (51–250 bp) into 20 bins of 10 bp each. F-ratio features are then derived from these 20 bins using three feature engineering approaches: (1) sliding window fractions (19 features), (2) probability fractions with regular binning (252 features), and (3) probability fractions with irregular binning (39 features).

5. Model development

To establish the CEliver model, a machine learning classifier was developed using LightGBM (Light Gradient Boosting Machine; LGBM), incorporating the engineered CF-2D features and AFP levels to distinguish between HCC patients and high-risk individuals. The prediction cohort consisted of 71 high-risk individuals and 40 HCC cases. Using a stratified sampling approach, the dataset was separated into training and testing sets at an 80:20 ratio and normalized by the Min-Max scaler. Feature selection was performed using SelectKBest. Specifically, the top 150 features were selected based on their scores with p-values < 0.05 from a t-test, indicating statistically significant associations with HCC. These selected features were then used in the training process with the LightGBM classifier.

All hyperparameters were optimized through a grid search algorithm with 5-fold cross-validation based on the area under the ROC curve (AUC), following model evaluation by test set. Model performance on the test set was then evaluated using accuracy as the primary performance metric. Additional metrics, including the F1 score and Matthew’s correlation coefficient (MCC), were computed to assess classification balance and potential bias.

The model output was a prediction score ranging from 0 to 1, representing the probability of HCC presence (with scores closer to 0 indicating non-HCC and scores closer to 1 indicating HCC). Finally, an external validation cohort comprising 27 HCC cases, 30 high-risk individuals, and 12 healthy donors was used to evaluate the performance of the fixed CEliver model.

Whole-genome sequencing

All whole genome sequencing (WGS) workflow was performed by Macrogen Inc. company. The 10 ng of plasma cfDNA was used for WGS-library preparation by TruSeq Nano DNA (350) kit, according to the manufacturer. The library conditions of cfDNA were prepared without fragmentation and size selection. The amplification with light PCR cycles (6–10 cycles) was performed after adapter ligation. The sequenced reads were generated on Novaseq 6000 platforms (Illumina) with 150-bp paired-end reads at 40x coverage. At the same time, WGS data of gDNA from matched peripheral blood mononuclear cells (PBMCs), a matched normal control, was also constructed at 30x coverage.

To identify the variants, the fastq sequence data was analyzed using Genome Analysis Toolkit (GATK), a global variant calling pipeline. First, the quality of fastq files was checked by FASTQC (RRID:SCR_014583). After adapter removal, paired-end sequence reads were aligned to the human reference genome version GRCh38 using BWA-mem tool (RRID:SCR_010910). To avoid measurement bias and reduce error, Mark Duplicates (Picard Tools) was used to mark the set of duplicate reads, which were excluded from downstream analysis. Given a matched normal, the somatic single nucleotide variants (SNVs) and Indel mutation were identified through Mutect2 variant caller.

Statistics and reproducibility

All cfDNA profiles analyzed in this study were derived from 67 HCC patients, 101 high-risk individuals, and 12 healthy controls. Each cfDNA sample was performed in triplicate run analysis using automated capillary electrophoresis to ensure reproducibility and data accuracy, and the average of the three runs was used for subsequent statistical analyses. All statistical analyses were performed using GraphPad Prism version 9.4.1 (GPS-1145384-ELPE-5E7B2; RRID: SCR_002798). The cfDNA profiles including cfDNA level and cfDNA fragment size was represented by median and Interquartile range; IQR. The normality of each dataset was assessed by Shapiro-Wik test. The Mann-Whitney U test was used to compare the cfDNA profiles, clinical data, and CEliver score between high-risk and HCC patients. For model development, the ANOVA F-value was used to identify significant feature differences between HCC and high-risk patients. The best cut-off of CEliver scores for distinguishing HCC from high-risk was optimized by ROC curve analysis, achieving the highest sensitivity and specificity. For all statistical tests, a p-value < 0.05 was considered statistically significant.

Ethics statement

All research was conducted in accordance with the Declaration of Helsinki and approved by the Research Ethics Committee of Faculty of Medicine, Chiang Mai University (Ethics approval code ORT- 2563–07122 and FAC-MED-2565-09352). Written consent was given in writing by all subjects.

Results

Clinical cohort and biological characteristics

To develop the HCC screening prediction model, 111 participants were recruited into the prediction cohort, including 71 high-risk individuals and 40 patients with HCC at various stages. The majority of high-risk participants had underlying chronic HBV hepatitis with or without cirrhosis, accounting for 83% of the cohort. Among the 40 HCC patients, the distribution of Barcelona Clinic Liver Cancer (BLCL) stage was as follows: 65% were in the early stage (0/A), 15% in the intermediate stage (B), and 30% in the advanced stage (C/D). In the validation cohort, a significant portion of high-risk patients also had HBV-related liver diseases (80%). Additionally, over 60% of the HCC cases in the cohort were in the early stages. Further clinical characteristics of the patients are provided in Table 1.

Table 1.

Main characteristics of patients in this study

Prediction cohort Validation cohort
Patient characteristics High-risk HCC P value Healthy High-risk HCC

P value

(high-risk vs HCC)

Subjects number 71 40 12 30 27
Gender (%)
Male 44 (62%) 32 (80%) 2 (17%) 24 (89%) 18 (67%)
Female 27 (38%) 8 (20%) 10 (83%) 6 (11%) 9 (33%)
Age at diagnosis (years)
Median 59 62 51 58 65
Range 39-80 33-76 23-71 31-81 17-81
Underling of liver disease (%)
Hepatitis B 59 (83%) 26 (65%) N/A 24 (80%) 11 (41%)
Hepatitis C 4 (5.7%) 6 (15%) N/A 1 (3%) 7 (26%)
Alcohol-associated 0 (0%) 6 (15%) N/A 1 (3%) 3 (11%)
Other 8 (11.3%) 2 (5%) N/A 4 (14%)s 6 (22%)
BCLC stage (%)
0/A N/A 23 (57.5%) N/A N/A 17 (63%)
B N/A 6 (15%) N/A N/A 3 (11%)
C N/A 9 (22.5%) N/A N/A 3 (11%)
D N/A 2 (5%) N/A N/A 4 (15%)
Tumor size
<2 cm N/A 5 (12.5%) N/A N/A 6 (24%)
≥ 2 cm N/A 35 (87.5%) N/A N/A 19 (76%)
Number of tumors
1 N/A 26 (65%) N/A N/A 16 (62%)
2-3 N/A 7 (17.5%) N/A N/A 7 (27%)
multifocal N/A 7 (17.5%) N/A N/A 3 (11%)
cfDNA level (ng/mL plasma)
Median 6.3 (3.9) 18.9 (47.3) <0.001 6.1 (3.8) 6.0 (5.0) 21 (17.5) <0.001
cfDNA size distribution (bp)
Median 168 (6.0) 161 (6.6) <0.001 168 (3.0) 176 (9.5) 164 (11.0) <0.001
AFP (ng/mL)
Median 2.7 (1.6) 165 (6196.0) <0.001 2.0 (0.3) 1.9 (0.8) 21 (363.1) <0.001
AST (U/L)
Median 24 (13.5) 123 (123.3) <0.001 24 (8.3) 25 (12.0) 63 (64.5) <0.001
ALT (U/L)
Median 20 (12.5) 84 (106.0) <0.001 18 (20.5) 24 (7.8) 38 (94.5) 0.0027

The data for cfDNA, AFP, AST, and ALT were represented as the median with IQR. A p-value < 0.05 was considered statistically significant.

cfDNA characterization and clinical data

To investigate cfDNA characteristics of high-risk individuals and HCC patients, we analyzed cfDNA level and fragment size distribution using the automated capillary electrophoresis method. We found that plasma cfDNA concentrations in HCC patients were significantly higher than high-risk individuals (Fig. 3A). Moreover, cfDNA levels tend to increase with HCC stage. The median cfDNA level in high-risk patients was 6.3 ng/mL plasma (range: 1.3–49.8; IQR 3.9). The cfDNA level was significantly increased to 10.5 ng/mL plasma (range: 2.5–217.9; IQR 26.4) in early stage (0/A) and 38.40 ng/mL plasma (range 3.3–480.8: IQR 58.7) in intermediate to advanced stage (B/C/D) (Fig. 3B). For cfDNA fragment size distribution analysis, we observed shorter cfDNA lengths in HCC patients compared to high-risk group and these lengths tended to be shorter in more advanced HCC stages (Fig. 3A, B). The median cfDNA fragment size in the high-risk group was 168 bp (range: 159–209 bp; IQR: 6.0). In early- and advanced-stage HCC, the median fragment sizes were 163 bp (range: 151–174 bp; IQR: 7.5) and 162 bp (range: 133–174 bp; IQR: 6.6), respectively. Moreover, we also investigated the cfDNA profiles in healthy, high-risk, and HCC. The cfDNA profiles of high-risk individuals more closely resembled that of healthy controls than that of HCC patients, which is consistent with the non-malignant status of this group (Supplementary Fig. 2A, B).

Fig. 3. The cfDNA characterization in HCC patients.

Fig. 3

The comparison of cfDNA level, cfDNA fragment size, and AFP level between (A) HCC patients and High-risk, B BCLC stage. Plasma cfDNA levels are significantly higher in HCC patients than in high-risk individuals and increase with HCC stage. cfDNA fragment size distributions are shorter in HCC patients and decrease further with advancing stage. C ROC analysis indicates limited early HCC detection performance for cfDNA fragment size, cfDNA level, and AFP when used individually. D The distribution of cfDNA fragmentation pattern according to two-dimension analysis in High-risk and HCC patients showed that HCC-derived cfDNA is more fragmented than cfDNA from high-risk individuals. Statistical significance is indicated as follows: p < 0.05 (*), p < 0.01 (**), p < 0.001 (***), and p < 0.0001 (****). The (ns) are provided for comparisons that were not statistically significant.

Serum AFP levels were significantly higher in HCC patients compared to high-risk individuals, with elevated AFP levels strongly associated with more advanced stages of HCC. Interestingly, cfDNA levels and size distribution showed consistency with serum AFP levels. However, when focusing on early-stage HCC patients, cfDNA fragment size exhibited a significant difference between high-risk individuals and early HCC cases, while cfDNA levels and AFP did not show statistical significance (Fig. 3B). Notably, over 69% of early-stage HCC patients had AFP levels below the typical screening cut-off of 20 ng/mL (Fig. 3B). Based on ROC analysis, the individals of cfDNA level, cfDNA fragment size, and AFP demonstrated low performance for detecting early-stage HCC from high-risk groups (Fig. 3C, Table 2). These findings suggest that cfDNA level or fragment size alone can distinguish high-risk individuals from HCC patients; however, integrating these markers with clinical data may improve the detection of early-stage HCC.

Table 2.

The HCC detection performance of CEliver, cfDNA level, cfDNA fragment size and AFP

Groups No. of patients Model/Test Performance
Sensitivity Specificity PPV NPV AUC (95% CI)
Model Development
All HCC High-risk = 71 AFP 60% 98% 92% 81% 0.82 (0.72–0.92)
HCC = 40 cfDNA level - - - - 0.80 (0.71–0.89)
cfDNA fragment - - - - 0.78 (0.70–0.88)
CEliver model 98% 99% 98% 99% 0.98 (0.93–1.00)
Early HCC High-risk = 71 AFP 35% 98% 80% 82% 0.73 (0.58–0.88)
(0/A) Stage 0/A = 23 cfDNA level - - - - 0.73 (0.60–0.86)
cfDNA fragment - - - - 0.74 (0.63–0.86)
CEliver model 96% 99% 95% 99% 0.96 (0.89–1.00)
External Validation
All HCC High-risk = 30 AFP 52% 100% 100% 70% 0.76 (0.63–0.89)
HCC = 27 CEliver model 85% 100% 100% 91% 0.93 (0.85–1.00)
Early HCC High-risk = 30 AFP 47% 100% 100% 77% 0.73 (0.57–0.90)
(0/A) Stage 0/A = 17 CEliver model 88% 100% 100% 100% 0.94 (0.85–1.00)

cfDNA 2-dimension feature

To enhance detection sensitivity, we developed the 2-dimension features (2D features) analysis to captures both fragment size and cfDNA intensity simultaneously. We examine the distribution and density of cfDNA in each 10 bp per window starting with 50 to 250 bp, going up to 20 windows. The high density of cfDNA fragments in HCC patients significantly increased into shorter fragments as compared to high-risk cfDNA, especially for cfDNA fragment sizes smaller than the 167 bp peak, which represents the typical length of DNA wrapped around by a mono-nucleosome complex. Heatmap analysis of cfDNA fragmentation patterns strongly suggests that HCC-derived cfDNA is more fragmented than cfDNA from high-risk individuals (Fig. 3D). This finding supports the potential use of cfDNA fragmentation profiles in distinguishing HCC from high-risk group.

In this study, we aimed to improve the sensitivity and specificity of a prediction model based on automated capillary electrophoresis data. A total of 332 CF-2D features, extracted from electropherogram data, along with AFP levels, were incorporated into the dataset.

Machine learning model development

To develop a machine learning model to detect HCC from high-risk patients, all data was learned under Light Gradient Boosting Machine (LightGBM) by repeated 5-fold validation. The model performance metrics, including accuracy, sensitivity, specificity, F1 score, and MCC, are summarized in Supplementary Table 2. The 150 features, including 1 of AFP level and 149 of CF-2D features, were presented to be significant in distinguishing between the high-risk group and HCC in a machine learning model. These features were then utilized to be the final input data for CEliver model establishment. The pattern of the top 15 significant features differed between high-risk and HCC groups, reflecting the model’s design to identify features with consistent differences between these populations and emphasizing the biological relevance of cfDNA fragmentation signatures associated with HCC (Fig. 4A). Overall, this dataset covers a comprehensive set of cfDNA features and clinical parameters that can increase the accuracy of the HCC detection algorithm.

Fig. 4. CEliver model establishment.

Fig. 4

A A unique pattern of 15 significant features distinguishes cfDNA profiles between HCC patients and high-risk individuals. B Comparison of CEliver scores between HCC patients and high-risk individuals shows that CEliver scores were higher in HCC patients than in high-risk individuals. C ROC curve analysis comparing the performance on distinguishing HCC and high-risk individuals using cfDNA fragment size (purple curve), cfDNA level (blue curve), AFP (green curve), and CEliver model from the training and test set during model development (orange curve). D ROC curve analysis comparing the performance of CEliver model (training and test set) with that of the traditional AFP method for early HCC detection. the CEliver model outperforms AFP for early HCC detection and demonstrates superior performance compared with cfDNA fragment size and cfDNA level alone. E Correlation of CEliver score in other clinical parameters shows that CEliver scores are associated with AFP levels but not with gender, age, or tumor size. Statistical significance is indicated as follows: p < 0.05 (*), p < 0.01 (**), p < 0.001 (***), and p < 0.0001 (****). The (ns) are provided for comparisons that were not statistically significant.

CEliver model for HCC detection

The CEliver score, which ranges from 0.00 to 1.00, was constructed based on the distance to HCC, with values closer to 1.00 indicating a stronger association with HCC. According to our predictive model, the CEliver score was found to be lower in high-risk individuals (median 0.03; ranging 0.01–0.58; IQR 0.04), instead patients with HCC had a significantly higher score (median 0.94; ranging 0.63–0.99: IQR 0.11) (Fig. 4B). Furthermore, the CEliver score indicated a trend toward a gradual increase in advanced HCC compared to early HCC, with median score of 0.98 and 0.88, respectively. The correlation between CEliver score and other clinical indicators demonstrated that the CEliver score was associated with AFP level, but there was no correlation with gender, age, and tumor size (Fig. 4E).

The performance of the CEliver model, derived from the training and test sets, demonstrated an excellent AUC of 0.98 (95% CI: 0.93–1.00) for distinguishing HCC patients from high-risk individuals, according to ROC analysis. In contrast, the performance of using cfDNA level or cfDNA fragment length alone, without the integrative model, was notably lower, with AUCs of 0.80 (95% CI: 0.71–0.89) and 0.79 (95% CI: 0.70–0.88), respectively. These results were comparable to the conventional serum AFP measurement, which yielded an overall AUC of 0.82 (95% CI: 0.72–0.92) (Fig. 4C). Notably, the CEliver model maintained strong performance in predicting early-stage HCC (stage 0/A), which remains the most challenging to detect, achieving an AUC of 0.96 (95% CI: 0.89–1.00) (Fig. 4D). While, AFP measurement performance declined substantially, with an AUC of 0.73 (95% CI: 0.58–0.88).

Diagnostic performance of the CEliver model

The optimal cutoff value of the CEliver score (0.4) was determined using Youden’s Index derived from the ROC curve analysis, corresponding to the point maximizing the sum of sensitivity and specificity. In model construction, CEliver demonstrated effective diagnostic performance, achieving an accuracy of 98%, sensitivity of 98%, specificity of 99%, positive predictive value (PPV) of 98%, and negative predictive value (NPV) of 99%. Only one HCC patient was undetectable, resulting in high accuracy with a sensitivity of 98% (39/40) in all stages, 96% in early stage (22/23), and 100% in advanced stage (17/17) at a specificity of 99% (Fig. 5A and B). To investigate the practical implications of the CEliver model for HCC detection, we compare the efficacy of the CEliver model with the AFP level at 20 ng/mL, which is a recommended routine for HCC screening. In our cohort, the sensitivity of AFP level in detection of all HCC stages was found to be 60% (24/40) and dropped to only 35% (8/23) in patients with early stage at a specificity of 98% (Fig. 5B).

Fig. 5. The predictive performance of CEliver model.

Fig. 5

A The diagnostic performance of the CEliver model in the prediction phase demonstrates high accuracy (98%), sensitivity (98%), and specificity (99%). B Comparison of the sensitivity of AFP and the CEliver model shows that CEliver achieves higher sensitivity for early-stage HCC detection than AFP. Error bars represent the mean with ± 95% confidence interval (CI). C Stratified by TNM stage, the CEliver model effectively detects tumors in all patients with tumor sizes <2 cm (5/5), whereas AFP shows a detection rate of only 40% (2/5). Error bars represent the mean with ± 95% CI. D Clinical timeline of a high-risk patient who tested positive using the CEliver model.

The performance of CEliver model in case of small tumor size was also evaluated to assess the model’s performance in detecting small ( ≤ 2 cm) HCC lesions, which remain a major clinical challenge, as these cases are often AFP-negative or missed by ultrasound screening. The CEliver model is effective in detecting tumors in all patients with tumor sizes less than 2 cm (5/5), whereas the AFP level only provided a 40% detection rate (2/5) (Fig. 5C). Interestingly, one of the high-risk patients developed a 1 cm hepatic nodule 11 months after testing positive. Seventeen months later, the nodule showed an increase in size, and at 24 months, the patient was diagnosed with HCC through MRI confirmation (Fig. 5D). These findings showed that our model effectively identified early HCC in high-risk individuals, including those with single nodules smaller than 2 cm. Furthermore, the CEliver model successfully detected HCC in high-risk patients prior to clinical diagnosis.

External validation of the CEliver model in a separate cohort

To evaluate the predictive performance of the model, the fixed CEliver model was validated using an independent cohort comprising 27 HCC patients, 30 high-risk individuals (19 with viral-related chronic liver disease and 11 with liver cirrhosis), and 12 healthy volunteers. The HCC median score was 0.89, which was higher than that of the high-risk groups, which had median scores of 0.19. All healthy donors had undetectable scores, with a median score of 0.13 (Fig. 6A). Using a cutoff score of 0.4, the CEliver model identified 23 out of 27 HCC patients as positive, while correctly classifying high-risk patients and healthy individuals as non-cancerous. As a result, the accuracy was 94%, with a sensitivity of 85%, specificity of 100%, PPV of 100%, and NPV of 91% at AUC 0.93 (Fig. 6B, Table 2).

Fig. 6. External validation of CEliver model.

Fig. 6

A CEliver scores in healthy donors, high-risk individuals, and HCC patients, with higher median scores in HCC. B External validation shows 94% accuracy, 85% sensitivity, and 100% specificity. C Sankey plot illustrates successful detection of 15/17 early-stage HCC cases. D Compared with AFP, CEliver improves early-stage HCC detection, identifying 7/9 AFP-negative cases (78%). E Across BCLC stages, CEliver achieves higher sensitivity (88%) than AFP (47%). Error bars represent the mean with ± 95% CI. F Stratified by TNM stage, CEliver detects tumors <2 cm with 83% sensitivity (5/6), whereas AFP detects none. Error bars represent the mean with ± 95% CI. Statistical significance: *p < 0.05, **p < 0.01, ***p < 0.001, **p < 0.0001; ns, not significant.

Additionally, we evaluated the performance of the CEliver model in the validation cohort compared to AFP methods. Focusing on early-stage HCC, our predictive algorithm identified 7 out of 9 early HCC patients (78%) with negative AFP levels (Fig. 6D). Notably, CEliver successfully detected 15 out of 17 early-stage HCC cases, highlighting a high sensitivity for early-stage HCC (Fig. 6C), achieving a sensitivity of 88%, compared to AFP alone, which showed a sensitivity of only 47% (Fig. 6E). We also assessed the model’s ability to detect small tumor lesions. CEliver successfully identified patients with tumors smaller than 2 cm at sensitivity of 83% (5 out of 6), whereas AFP failed to detect any of these cases (Fig. 6F, Supplementary Fig. 3C). These findings suggest that the CEliver model could serve as an effective tool to integrate to the HCC surveillance program, particularly for identifying early-stage HCC cases that may go undetected by standard methods.

Conceptual verification of cfDNA fragment size on genome-wide scale

To investigate whether the shortened cfDNA fragments primarily originate from cancer cells, cfDNA from three HCC cases were analyzed using the whole-genome sequencing (WGS). Mutant cfDNA was distinguished from wild-type cfDNA using Jvarkit (Biostar322664; RRID:SCR_002580). Fragment pattern analysis revealed that mutant cfDNA fragments were predominantly shifted toward shorter lengths compared to wild-type cfDNA fragments. Specifically, a higher proportion of mutant cfDNA fragments were observed within the 140–160 bp range, whereas non-mutant cfDNA fragments had an average size of 166 bp (Fig. 7A). These findings provide evidence that tumor-derived cfDNA is associated with shorter fragment sizes.

Fig. 7. Conceptual verification of cfDNA fragment size on genome-wide scale.

Fig. 7

A Distribution of cfDNA fragment reads carrying somatic mutations (red) and wild-type alleles (green) from HCC01, HCC02, and HCC03 patients. Mutant cfDNA fragments are enriched in shorter fragment size ranges. B Genome-wide cfDNA fragmentation patterns assessed using the short (90–150 bp) to long (151–250 bp) fragment ratio in healthy individuals, HBV-related hepatitis patients, and HCC patients. HCC samples from both public datasets and our cohort show different fragmentation profiles compared with non-cancer individuals.

To demonstrate that the cfDNA fragmentation patterns observed in the CEliver model are biologically consistent with the well-established genome-wide short/long (S/L) ratio concept, we analyzed genome-wide cfDNA fragmentation patterns in 6 HCC cases and 3 high-risk patients. The ratio of short (90–150 bp) to long (151–250 bp) cfDNA fragments across the genome was determined using WGS data from HCC samples in our cohort, along with data from the FinaleDB database. Our analysis revealed that HCC patients exhibited a higher short-to-long (S/L) ratio compared to those with HBV-related hepatitis and healthy donors, as indicated by the cfDNA WGS data from FinaleDB (Fig. 7B). Alterations in fragmentation profiles were observed in HCC patients from both the database and our HCC samples, which were markedly different from those in non-cancer individuals. This finding confirms that the cfDNA short/long ratio as a critical element in constructing a predictive model for HCC.

Discussion

An earlier diagnosis of HCC not only increases the likelihood of timely medical intervention but also broadens the range of therapeutic options, including surgical resection, liver transplantation, and local ablation, all of which contribute to improved survival rates in HCC patients33–36. In this study, we identified a distinct distribution of shorter cfDNA fragment sizes in HCC patients compared to non-HCC patients (those with chronic liver diseases with or without cirrhosis). While using cfDNA fragment size alone demonstrates high discriminatory power in distinguishing HCC patients from healthy individuals, it has lower accuracy in differentiating HCC from high-risk individuals (Supplementary Fig. 2A, B). This limitation likely arises from the similarity in cfDNA profiles between high-risk individuals and early-stage HCC patients. Certain clinical variables, particularly inflammation, may act as significant confounders in high-risk viral hepatitis patients, potentially influencing cfDNA fragmentation patterns37. Consequently, achieving high sensitivity in detecting early-stage HCC among high-risk individuals remains a substantial challenge.

To enhance the sensitivity and accuracy of early-stage HCC detection, we introduced a CF-2D features analysis, designed to provide a high-dimensional profiling of cfDNA fragmentation analysis from the automated electrophoresis platform. By integrating CF-2D features with clinical data, we developed the CEliver predictive model. The top 15 key features effectively distinguished HCC patients from high-risk patients. One remarkable case involved a high-risk patient who later developed HCC within a year, exhibiting CF-2D patterns that closely aligned with those observed in HCC patients. Furthermore, according to our fixed model, the CEliver test demonstrated strong diagnostic performance, achieving a sensitivity of 88% (15 out of 17 early HCC cases detected), highlighting its potential for early HCC diagnosis.

Another major challenge of early-stage HCC detection is the notably low sensitivity of widely available ultrasound (US), with detection rates ranging from 21–47%10. Nodules that are undetectable by surveillance US often occur in tumors smaller than 2.0 cm38, especially in cirrhotic livers and liver nodules located in difficult area such as segment 7,8 or 4a39. A meta-analysis of 756 articles including 4250 patients demonstrated a high rate of inadequate US, highlighting the limitations of current HCC surveillance program40. Additionally, detecting these small tumors is challenging with cfDNA detection due to the low proportion of circulating tumor DNA41. Interestingly, our analysis using tumor size as a continuous variable revealed that cfDNA concentration increased and fragment size decreased with larger tumor burden, reinforcing the notion that tumor-derived cfDNA reflects both disease extent and cellular turnover, consistent with the biological hypothesis (Supplementary Fig. 3A and 3B). Moreover, the CEliver model performed excellently in identifying HCC patients with tumor sizes smaller than 2 cm (10/11 cases), effectively overcoming the limitations of ultrasound. In comparison, AFP, the primary biomarker for HCC detection, identified only 2 positive cases (2/11 cases) (Supplementary Fig. 3C). Notably, the CEliver model successfully detected HCC in eight out of nine patients with tumor volume smaller than 2 cm who were also negative for AFP (Supplementary Fig. 3D). Moreover, the CEliver model effectively detected HCC in patients with low tumor fractions that were missed by copy number variation, a genetic-based analysis (Supplementary Table 4). The CEliver model has the potential to address critical clinical challenges in HCC detection.

Importantly, our approach shares key conceptual similarities with the Fragle model, which is a deep learning framework derived from the density distribution of cfDNA fragment lengths using low-pass whole-genome sequencing. Their model achieved high performance in multi-cancer detection. This study provides strong evidence supporting the relevance of high-dimensional cfDNA intensity and fragment size distribution features in characterizing cancer patients30. Moreover, previous studies have demonstrated the successful application of genome-wide fragmentation analysis for HCC detection, achieving a sensitivity of 88% and a specificity of 98% in a high-risk group. Most studies identified cfDNA fragment ranges of 90-150 bp and 151-220 bp as key for effectively distinguishing HCC29. By performing whole-genome sequencing analysis of cfDNA, we also observed distinct patterns in the short (90-150 bp) to long (151-220 bp) ratio of cfDNA fragments in HCC patients compared to high-risk and healthy individuals (Fig. 7). Interestingly, two of the top 15 features in CF-2D, PR1_7_15 and PR1_9_15, defined by the short/long fragment ratios within the ranges of 111–120 bp to 181–190 bp and 131–140 bp to 181–190 bp, align with the common short-to-long ratio for assessing cfDNA distribution across the genome. These results confirm that the cfDNA fragment size feature used in the CELiver model reflects unique cfDNA fragmentation patterns in HCC.

Compared to other non-NGS approaches, qPCR-based analysis showed a lower cfDNA integrity index, defined as the ratio of long fragments ( > 205 bp) to short fragments (110 bp), in HCC patients compared to those with chronic liver disease, regardless of viral or non-viral etiology. The combination of cfDNA integrity and AFP levels demonstrated a sensitivity and specificity of 68% and 67%, respectively, with an AUC of 0.6742. Elzehery et al.43 reported that the integrity index of ALU 247/115 achieved a sensitivity of 70% and a specificity of 88% for distinguishing HCC patients from those with liver cirrhosis, with an AUC of 0.856. Detection using a hotspot mutation panel of four gene loci, TP53 (c.747 G > T), CTNNB1 (c.121 A > G, c.133 T > C), and TERT (c.1-124 C > T) via ddPCR in 48 HCC patients exhibited a sensitivity of 56% (27/48) and specificity of 86%44. In contrast, the CEliver model demonstrated markedly superior performance compared to other non-NGS approaches, achieving 88% sensitivity, 100% specificity, and an AUC of 0.95 for early HCC detection (Supplementary Table 5).

Although both qPCR/ddPCR and CEliver assays can generally be completed within a one-day turnaround time, qPCR and ddPCR workflows are technically more complex, requiring multiple primer/probe sets, enzymatic reagents, and stringent quality control steps, which increase both labor demand and overall assay cost. In contrast, the CEliver assay employs a single-step electrophoretic cfDNA profiling workflow with minimal reagent consumption and straightforward computational analysis, thereby reducing the per-sample cost and simplifying operation. These findings highlight that the CEliver model represents a significant breakthrough in non-NGS-based diagnostics for liver cancer and has strong potential for scalable clinical implementation.

While this study represents a potential improvement in HCC screening, several considerations should be addressed. One limitation is the small sample size of individuals with HCC, particularly in the validation cohort. In the high-risk population, most individuals were recruited from viral hepatitis clinics. The high proportion of HBV- and HCV-related liver diseases in our learning set may affect the model’s accuracy when applied to other etiologies, such as alcoholic liver disease. Therefore, this model may be more suitable for implementation in HCC screening clinics associated with viral hepatitis, particularly in region with high HBV prevalence, such as Asia. Epidemiological data from China have shown that over 84% of HCC cases and 77% of cirrhosis are attributable to HBV infection45. Similarly, a recent meta-analysis including 39,050 HBV patients in Southeast Asia, with Thailand as the primary contributor, reported a 45% prevalence of HBV-associated HCC46. Consistent with our cohort, more than 65% of HCC cases had underlying HBV-related liver disease.

Although HCC most commonly develops in patients with cirrhosis. It can also arise in non-cirrhotic patients with chronic HBV infection. Currently, screening is recommended for men aged ≥40 years and women aged ≥50 years, as the annual incidence of HCC in these subgroups exceeds 0.2%47. This incidence threshold has been widely accepted as the cost-effectiveness benchmark for initiating surveillance, indicating that regular screening becomes economically justified when the expected benefit outweighs the healthcare expenditure48.

Future studies focusing on the expansion of larger external validation cohorts, particularly for early-stage HCC patients, along with multi-site research, are essential to confirm the precision and reliability of the model algorithm. Moreover, we plan to include high-risk individuals with various risk factors to reduce selection bias and better evaluate the model’s predictive performance across diverse etiologies of liver disease. This effort involves the recruitment of additional patients with metabolic liver diseases, including non-alcoholic fatty liver disease (NAFLD) and non-alcoholic steatohepatitis (NASH), which represent an increasing trend in liver disease prevalence in the future. These analyses will provide more robust evidence of the model’s applicability in real-world clinical settings. In addition to its high accuracy, our model is affordable and easy to analyze, aligning with the screening program requirements outlined by the World Health Organization. Therefore, evaluating the cost per detection or cost per quality-adjusted life year, longitudinal follow-up, as well as survival analysis will be one of our next key objectives in promoting the practical and widespread adoption of the model in HCC screening clinics

Conclusions

The CEliver model, which utilizes an electrophoresis approach, demonstrated excellent accuracy and sensitivity in the early diagnosis of HCC in high-risk groups. With its high performance, non-invasive nature, ease of use, and affordability, the CEliver model offers a promising option for widespread HCC screening in high-risk populations.

Supplementary information

43856_2026_1437_MOESM2_ESM.pdf (28.1KB, pdf)

Description of Additional Supplementary files

Supplementary Data 1 (38.4KB, xlsx)

Acknowledgements

This research project was supported by the Fundamental Fund 2024, Chiang Mai University, the Faculty of Medicine, Chiang Mai University, and the National Research Council of Thailand (NRCT) (grant no. N42A670184), Thailand. We would like to thank all the patients who participated in this study. The graphical abstract, as well as Figs. 1 and 2 were created using BioRender.com.

Author contributions

S. Udomruk: Conceptualization, data curation, formal analysis, investigation, methodology, project administration, validation, visualization, funding acquisition, writing–original draft, writing–review and editing. S. Sutthitthasakul: Formal analysis, software, methodology, validation. N. Bunsermvicha: Resources, data curation, formal analysis. K. Pinyopornpanish, C. Charoentum, S. Kongkarnka, and D. Pruksakorn: Resources, data curation. P. Charoenkwan, T. Jaimalai, T. Duangsan and P. Noisagul: Software, methodology. P. Yongpitakwattana and K. Khounkaew: Investigation, methodology. S. Orrapin and S. Moonmuang: Investigation, validation. A. Pasena, P. Teeyakasem, P. Suksakit and R. Gamngoen: Resources, data curation. W. Lapisatepun: Conceptualization, supervision, resources, validation, funding acquisition, writing–review and editing. P. Chaiyawat: Conceptualization, supervision, methodology, funding acquisition, project administration, resources, validation, visualization, writing–original draft, writing–review and editing. All authors have read and agreed to the published version of this manuscript.

Peer review

Peer review information

Communications Medicine thanks Jilei Liu and the other anonymous reviewer(s) for their contribution to the peer review of this work.

Data availability

Source data underlying the analyses in the main figures are available in Supplementary Data. Data underlying Fig. 3A–C are provided as raw data, while data underlying Fig. 3D are provided as a 2D heatmap plot. Source data for Figs. 4B, 4C, 4D, 5A, and 5B are provided as CEliver scores.

Sequencing data and other sensitive information are stored on secure institutional servers at the Faculty of Medicine, Chiang Mai University, Thailand. Access to these data is restricted due to ethical and institutional policies. Requests for access should be directed to Dr. Chaiyawat (parunya.chaiyawat@cmu.ac.th) and will be considered subject to institutional approval procedures and policies.

Code availability

To facilitate reproducibility and further research, all scripts for feature extraction, model training, and prediction, as well as a small demonstration dataset, are publicly available on GitHub: 10.5281/zenodo.1829735749.

Competing interests

Sasimol Udomruk, Songphon Sutthitthasakul, Worakitti Lapisatepun, and Parunya Chaiyawat report a petty patent application (Thai-2403001380) currently under review, licensed to CELiver, for the cfDNA-based automated capillary electrophoresis method for liver cancer screening. All other authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Worakitti Lapisatepun, Email: worakitti.l@cmu.ac.th.

Parunya Chaiyawat, Email: parunya.chaiyawat@cmu.ac.th.

Supplementary information

The online version contains supplementary material available at 10.1038/s43856-026-01437-5.

References

  • 1.Rumgay, H. et al. Global burden of primary liver cancer in 2020 and predictions to 2040. J. Hepatol.77, 1598–1606 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Samant, H., Amiri, H. S. & Zibari, G. B. Addressing the worldwide hepatocellular carcinoma: epidemiology, prevention and management. J. Gastrointest. Oncol.12, S361–s373 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Calderon-Martinez, E. et al. Prognostic scores and survival rates by etiology of hepatocellular carcinoma: a review. J. Clin. Med. Res.15, 200–207 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Konyn, P., Ahmed, A. & Kim, D. Current epidemiology in hepatocellular carcinoma. Expert Rev. Gastroenterol. Hepatol.15, 1295–1307 (2021). [DOI] [PubMed] [Google Scholar]
  • 5.Chonprasertsuk, S. & Vilaichone, R. K. Epidemiology and treatment of hepatocellular carcinoma in Thailand. Jpn J. Clin. Oncol.47, 294–297 (2017). [DOI] [PubMed] [Google Scholar]
  • 6.Kanwal, F. & Singal, A. G. Surveillance for hepatocellular carcinoma: current best practice and future direction. Gastroenterology157, 54–64 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Zhang, B. H., Yang, B. H. & Tang, Z. Y. Randomized controlled trial of screening for hepatocellular carcinoma. J. Cancer Res. Clin. Oncol.130, 417–422 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Parikh, N. D. et al. Biomarkers for the Early Detection of Hepatocellular Carcinoma. Cancer Epidemiol. Biomark. Prev.29, 2495–2503 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Park, H. J. & Kim, S. Y. Imaging modalities for hepatocellular carcinoma surveillance: expanding horizons beyond ultrasound. J. Liver Cancer20, 99–105 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Tzartzeva, K. et al. Surveillance imaging and alpha fetoprotein for early detection of hepatocellular carcinoma in patients with cirrhosis: a meta-analysis. Gastroenterology154, 1706–1718.e1701 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Francini, E., Nuzzo, P. V. & Fanelli, G. N. Cell-free DNA: unveiling the future of cancer diagnostics and monitoring. Cancers16, 662 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Gao, Q. et al. Circulating cell-free DNA for cancer early detection. Innovation3, 100259 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Telekes, A. & Horváth, A. The role of cell-free DNA in cancer treatment decision making. Cancers14, 10.3390/cancers14246115 (2022). [DOI] [PMC free article] [PubMed]
  • 14.Moser, T., Kühberger, S., Lazzeri, I., Vlachos, G. & Heitzer, E. Bridging biological cfDNA features and machine learning approaches. Trends Genet.39, 285–307 (2023). [DOI] [PubMed] [Google Scholar]
  • 15.An, Y. et al. DNA methylation analysis explores the molecular basis of plasma cell-free DNA fragmentation. Nat. Commun.14, 287 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhou, Q. et al. Epigenetic analysis of cell-free DNA by fragmentomic profiling. Proc. Natl. Acad. Sci. USA119, e2209852119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zhou, Z. et al. Fragmentation landscape of cell-free DNA revealed by deconvolutional analysis of end motifs. Proc. Natl. Acad. Sci. USA120, e2220982120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Snyder, M. W., Kircher, M., Hill, A. J., Daza, R. M. & Shendure, J. Cell-free DNA comprises an in vivo nucleosome footprint that informs its tissues-of-origin. Cell164, 57–68 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Udomruk, S., Orrapin, S., Pruksakorn, D. & Chaiyawat, P. Size distribution of cell-free DNA in oncology. Crit. Rev. Oncol. Hematol.166, 103455 (2021). [DOI] [PubMed] [Google Scholar]
  • 20.Chen, L. et al. Cell-free DNA testing for early hepatocellular carcinoma surveillance. EBioMedicine100, 104962 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Jin, C. et al. Characterization of fragment sizes, copy number aberrations and 4-mer end motifs in cell-free DNA of hepatocellular carcinoma for enhanced liquid biopsy-based cancer detection. Mol. Oncol.15, 2377–2389 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Nguyen, V. C. et al. Fragment length profiles of cancer mutations enhance detection of circulating tumor DNA in patients with early-stage hepatocellular carcinoma. BMC Cancer23, 233 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Zhang, X. et al. Ultrasensitive and affordable assay for early detection of primary liver cancer using plasma cell-free DNA fragmentomics. Hepatology76, 317–329 (2022). [DOI] [PubMed] [Google Scholar]
  • 24.Wang, J. et al. Altered cfDNA fragmentation profile in hypomethylated regions as diagnostic markers in breast cancer. Epigenetics Chromatin16, 33 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Serpas, L. et al. Dnase1l3 deletion causes aberrations in length and end-motif frequencies in plasma DNA. Proc. Natl. Acad. Sci. USA116, 641–649 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Li, B. et al. DNASE1L3 inhibits proliferation, invasion and metastasis of hepatocellular carcinoma by interacting with β-catenin to promote its ubiquitin degradation pathway. Cell Prolif.55, e13273 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Lapin, M. et al. Fragment size and level of cell-free DNA provide prognostic information in patients with advanced pancreatic cancer. J. Transl. Med.16, 300 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Cristiano, S. et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature570, 385–389 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Foda, Z. H. et al. Detecting liver cancer using cell-free DNA fragmentomes. Cancer Discov.13, 616–631 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhu, G. et al. A deep-learning model for quantifying circulating tumour DNA from the density distribution of DNA-fragment lengths. Nat. Biomed. Eng.9, 307–319 (2025). [DOI] [PubMed] [Google Scholar]
  • 31.WHO Screening criteria.
  • 32.Reig, M. et al. BCLC strategy for prognosis prediction and treatment recommendation: The 2022 update. J. Hepatol.76, 681–693 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Drefs, M. et al. Changes of long-term survival of resection and liver transplantation in hepatocellular carcinoma throughout the years: A meta-analysis. Eur. J. Surg. Oncol.50, 107952 (2024). [DOI] [PubMed] [Google Scholar]
  • 34.Toubert, C. et al. Prolonged survival after recurrence in HCC resected patients using repeated curative therapies: Never give up! Cancers15, 10.3390/cancers15010232 (2022). [DOI] [PMC free article] [PubMed]
  • 35.Schoenberg, M. B. et al. Resection or transplant in early hepatocellular carcinoma. Dtsch Arztebl Int114, 519–526 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Zhang, X. et al. Surgical treatment improves overall survival of hepatocellular carcinoma with extrahepatic metastases after conversion therapy: a multicenter retrospective study. Sci. Rep.14, 9745 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Nakano, K. et al. Fragmentation of cell-free DNA is induced by upper-tract urothelial carcinoma-associated systemic inflammation. Cancer Sci.112, 168–177 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Kim, Y. Y. et al. Failure of hepatocellular carcinoma surveillance: inadequate echogenic window and macronodular parenchyma as potential culprits. Ultrasonography38, 311–320 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Tsang, S. H. et al. High-intensity focused ultrasound ablation of liver tumors in difficult locations. Int J. Hyperth.38, 56–64 (2021). [DOI] [PubMed] [Google Scholar]
  • 40.Hong, S. B. et al. Inadequate ultrasound examination in hepatocellular carcinoma surveillance: a systematic review and meta-analysis. J. Clin. Med.10, 10.3390/jcm10163535 (2021). [DOI] [PMC free article] [PubMed]
  • 41.Bredno, J., Lipson, J., Venn, O., Aravanis, A. M. & Jamshidi, A. Clinical correlates of circulating cell-free DNA tumor fraction. PLoS One16, e0256436 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kumar, S. et al. Evaluation of the cell-free DNA integrity index as a liquid biopsy marker to differentiate hepatocellular carcinoma from chronic liver disease. Front Mol. Biosci.9, 1024193 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Elzehery, R. et al. Circulating cell-free DNA and DNA integrity as molecular diagnostic tools in hepatocellular carcinoma. Am. J. Clin. Pathol.158, 254–262 (2022). [DOI] [PubMed] [Google Scholar]
  • 44.Huang, A. et al. Detecting circulating tumor DNA in hepatocellular carcinoma patients using droplet digital PCR is feasible and reflects intratumoral heterogeneity. J. Cancer7, 1907–1914 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Yang, D. H. et al. Hepatocellular carcinoma progression in hepatitis B virus-related cirrhosis patients receiving nucleoside (acid) analogs therapy: a retrospective cross-sectional study. World J. Gastroenterol.27, 2025–2038 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Rabaan, A. A. et al. Prevalence of hepatocellular carcinoma in hepatitis B population within Southeast Asia: a systematic review and meta-analysis of 39,050 participants. Pathogens12, 10.3390/pathogens12101220 (2023). [DOI] [PMC free article] [PubMed]
  • 47.Mittal, S. et al. Role of age and race in the risk of hepatocellular carcinoma in veterans with hepatitis B virus infection. Clin. Gastroenterol. Hepatol.16, 252–259 (2018). [DOI] [PubMed] [Google Scholar]
  • 48.Sachar, Y., Brahmania, M., Dhanasekaran, R. & Congly, S. E. Screening for hepatocellular carcinoma in patients with hepatitis B. Viruses13, 10.3390/v13071318 (2021). [DOI] [PMC free article] [PubMed]
  • 49.Sutthitthasakul, S. Celiver: v1.0.0, <10.5281/zenodo.18297357> (2026).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

43856_2026_1437_MOESM2_ESM.pdf (28.1KB, pdf)

Description of Additional Supplementary files

Supplementary Data 1 (38.4KB, xlsx)

Data Availability Statement

Source data underlying the analyses in the main figures are available in Supplementary Data. Data underlying Fig. 3A–C are provided as raw data, while data underlying Fig. 3D are provided as a 2D heatmap plot. Source data for Figs. 4B, 4C, 4D, 5A, and 5B are provided as CEliver scores.

Sequencing data and other sensitive information are stored on secure institutional servers at the Faculty of Medicine, Chiang Mai University, Thailand. Access to these data is restricted due to ethical and institutional policies. Requests for access should be directed to Dr. Chaiyawat (parunya.chaiyawat@cmu.ac.th) and will be considered subject to institutional approval procedures and policies.

To facilitate reproducibility and further research, all scripts for feature extraction, model training, and prediction, as well as a small demonstration dataset, are publicly available on GitHub: 10.5281/zenodo.1829735749.


Articles from Communications Medicine are provided here courtesy of Nature Publishing Group

RESOURCES