Skip to main content
Cell Reports Methods logoLink to Cell Reports Methods
. 2024 Oct 14;4(10):100877. doi: 10.1016/j.crmeth.2024.100877

Cell-free DNA end characteristics enable accurate and sensitive cancer diagnosis

Jia Ju 1,5, Xin Zhao 2,5, Yunyun An 1,5, Mengqi Yang 1,5, Ziteng Zhang 2,5, Xiaoyi Liu 1, Dingxue Hu 1, Wanqiu Wang 1, Yuqi Pan 1, Zhaohua Xia 3, Fei Fan 4, Xuetong Shen 1, Kun Sun 1,6,
PMCID: PMC11573786  PMID: 39406232

Summary

The fragmentation patterns of cell-free DNA (cfDNA) in plasma can potentially be utilized as diagnostic biomarkers in liquid biopsy. However, our knowledge of this biological process and the information encoded in fragmentation patterns remains preliminary. Here, we investigated the cfDNA fragmentomic characteristics against nucleosome positioning patterns in hematopoietic cells. cfDNA molecules with ends located within nucleosomes were relatively shorter with altered end motif patterns, demonstrating the feasibility of enriching tumor-derived cfDNA in patients with cancer through the selection of molecules possessing such ends. We then developed three cfDNA fragmentomic metrics after end selection, which showed significant alterations in patients with cancer and enabled cancer diagnosis. By incorporating machine learning, we further built high-performance diagnostic models, which achieved an overall area under the curve of 0.95 and 85.1% sensitivity at 95% specificity. Hence, our investigations explored the end characteristics of cfDNA fragmentomics and their merits in building accurate and sensitive cancer diagnostic models.

Keywords: liquid biopsy, cfDNA fragmentomics, N-index, machine learning, plasma

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • Selecting cfDNA with ends discordant to nucleosomes enriches tumor-derived signals

  • We develop 3 diagnostic biomarkers based on cfDNA end selection

  • We validate using multiple datasets of various cancer types

  • We develop the EXCEL diagnosis model with 85.1% sensitivity for cross-tumor dataset

Motivation

Fragmentation patterns of plasma cell-free DNA (cfDNA) are emerging directions in cancer diagnostic biomarkers. In a previous study, we found that, in tumors, the nucleases tend to cut within nucleosomes during apoptotic DNA fragmentation to generate short-sized cfDNA molecules. We hypothesized that the tumor-derived cfDNA in patients with cancer could be enriched by picking up the molecules with ends located within the nucleosomes. We further analyzed alterations in the cfDNA fragmentomic patterns to assess their potential as a cancer biomarker.


Ju et al. explore the end characteristics in cfDNA for cancer liquid biopsy. In patients with cancer, the selection of cfDNA with ends discordant to hematopoietic nucleosomes could enrich the tumor-derived ones and uncover biomarkers; further incorporation with machine learning enables the development of highly accurate and sensitive diagnostic models.

Introduction

Plasma circulating cell-free DNA (cfDNA) is an intensively studied analyte for non-invasive prenatal testing, cancer diagnosis, and treatment surveillance (i.e., “cancer liquid biopsy”).1,2,3,4 In healthy subjects, the hematopoietic system is the major contributor of the cfDNA pool5,6,7; however, in various physiological/pathological conditions (e.g., cancer), the affected tissues (e.g., tumors) contribute a remarkable proportion of cfDNA to plasma.6,7,8 Researchers have discovered that cfDNA molecules are generated through a non-random procedure that is highly related to their tissues of origin.2,4,9,10,11,12 For instance, in patients with cancer, the tumor-derived cfDNA molecules are different from the background ones (i.e., from the hematopoietic system) in a variety of ways, such as end motif sequence usages (e.g., CCCA, which is linked to DNASE1L3),13,14,15,16 ending preferences,17,18,19,20 and coverage patterns in regulatory regions.11,12,21,22 In a recent study, Cristiano et al. showed that one could obtain high-performance diagnostic models through integrating cfDNA fragmentomics features with machine learning approaches.23 As a result, analysis of cfDNA fragmentomics has become an emerging direction in cancer liquid biopsy, and novel diagnostic biomarkers are under active development to explore their promising translational values.2,4,24,25

In previous studies, we and others demonstrated that the fragmentation of cfDNA is closely related to the nucleosome structure.9,11,12,17,18,26 In particular, we investigated the molecular basis of cfDNA fragmentation and revealed an inherent nexus among nucleosome structure, cfDNA ends, and cfDNA size profile; we also demonstrated that the major peak in the cfDNA size profile corresponds to DNA wrapped around an intact nucleosome with a linker, while short cfDNA fragments (e.g., tumor- and fetus-derived ones) are preferably cut within the nucleosomes.17,18 In addition, we found that the overall cfDNA ending profiles of patients with cancer are significantly different from control subjects and further developed an end-based metric, the E-index, for cancer diagnosis.17 However, the differences between non-hematopoietic cfDNA ends and the background ones are still largely uncharacterized. In this study, we investigated the characteristics of cfDNA relative to nucleosomes (either derived from hematopoietic cells or directly deduced from cfDNA coverage patterns) to further explore the biology of cfDNA fragmentomics as well as its translational values.

Results

Principle of end characteristics in cfDNA

Figure 1 illustrates the principle of cfDNA ending patterns relative to the nucleosome structure. In control subjects, as the majority of cfDNA molecules are released from the hematopoietic system,5,6,7 their ends should be relatively concordant with the nucleosome positioning of the blood cells, i.e., a high proportion of the cfDNA ends should be located in the linker regions. In patients with cancer, the concordances between cfDNA ends and hematopoietic nucleosomes would be decreased due to the presence of tumor-derived cfDNA molecules, which are short in size (i.e., more likely to be cut within the nucleosomes)18,27,28,29 and originate from non-hematopoietic tissues with different nucleosome positioning patterns. Hence, one could pick up the cfDNA molecules whose ends are discordant with blood cells’ nucleosome positioning (i.e., located within hematopoietic nucleosomes) to enrich and quantify the non-hematopoietic ones (e.g., tumor derived). To quantitatively measure the concordance between cfDNA and hematopoietic nucleosomes, we defined a metric named the “N-index” as the proportion of cfDNA molecules whose ends are located within hematopoietic nucleosomes (Figures 1 and S1; see STAR Methods). Hence, higher N-index values would indicate an elevated discordance between cfDNA ends and hematopoietic nucleosome positioning patterns, and vice versa. As a result, in patients with cancer, higher N-index values could be expected due to the presence of tumor-derived cfDNA, which might promise cancer diagnosis. In addition, we also investigated the changes in two widely studied fragmentomic features, cfDNA size and end motif pattern, with end selection (Figure S1; see STAR Methods) to further explore the biological insights as well as potential translational values of end selection. Note that in the current study, we utilized the nucleosome protecting regions in the GM12878 cell line (lymphoid lineage) determined by an MNase-seq (micrococcal nuclease digestion with deep sequencing) experiment as the nucleosome positions9,17; in addition, the nucleosome track deduced from the cfDNA coverage pattern in a previous study9 was also analyzed as a supplementary track to validate the results.

Figure 1.

Figure 1

Principle of end selection in cfDNA

Plasma cell-free DNA (cfDNA) is a mixture of hematopoietic (blue lines) and non-hematopoietic molecules (e.g., tumor derived; orange lines). cfDNA prefers to be cut in the linkers between nucleosomes; as the nucleosome positioning patterns are different among cell types, cutting ends in non-hematopoietic cfDNA (i.e., linkers in non-hematopoietic cells) could be located in regions annotated as nucleosomes in hematopoietic cells (gray shadow). Hence, it would be possible to enrich the tumor-derived cfDNA by picking up those molecules with ends located within nucleosomes (i.e., end selection). We further investigated the effect on cfDNA fragmentomics as well as explored the translational values (i.e., cancer diagnosis) of end selection.

cfDNA fragmentomic characteristics with end selection

We first investigated the fragmentomic characteristics of cfDNA in a panel of healthy subjects retrieved from our previous study.17 As shown in Figures 2A–2C and Table S1, the overall size distribution of cfDNA after end selection was similar to that without end selection, and the characteristic features, e.g., a peak at 166 bp and 10 bp periodicity below 143 bp,17 were preserved; however, the proportions of fragments no bigger than 150 bp were increased, and the proportions of reads with a CCCA end motif were also significantly increased (both p < 10−5, Wilcoxon signed-rank tests).

Figure 2.

Figure 2

Alterations in cfDNA fragmentomics with end selection

(A) Size distribution of cfDNA with end selection in a healthy subject.

(B and C) Proportions of (B) short fragments or (C) fragments with CCCA end motif pre- and post-end selection in 24 healthy subjects. Each dot represents one sample, and its pre- (black) and post-end selection (red) cfDNA fragmentomic features are linked by lines.

(D and E) Changes in the percentages of (D) short fragments (ΔS150) and (E) fragments with CCCA end motif (ΔMCCCA) with end selection between controls and patients with hepatocellular carcinoma (HCC) in various stages.

(F) Proportion of cfDNA fragments with ends located within nucleosomes (N-index) between controls and patients with HCC.

In (B) and (C), p values were calculated using Wilcoxon signed-rank tests. In (D)–(F), p values were calculated using Mann-Whitney U tests.

To explore the effect of end selection in patients with cancer, 56 pre-treatment patients with hepatocellular carcinoma (HCC) (31, 10, and 15 were in stage A [early], B [intermediate], and C [late], respectively, according to the Barcelona Clinic Liver Cancer staging system30) and 24 age- and gender-matched non-cancerous subjects (both p > 0.05, Mann-Whitney U test and chi-squared test for age and gender, respectively; Table S2) were recruited and analyzed in this study. In this cohort, the proportion of fragments no bigger than 150 bp and the proportions of reads with a CCCA end motif were significantly increased in all samples with end selection (Figures S2A–S2D), which was consistent with the observations in healthy subjects; however, we found that the levels of increase were different between HCC samples and controls. Hence, two metrics were defined to investigate the differences introduced by end selection: ΔS150 measures the change in the percentage of fragments no bigger than 150 bp with end selection and ΔMCCCA measures the change in the percentage of CCCA end motif usage (Figure S1; see STAR Methods). As shown in Figures 2D, 2E, S2E, and S2F and Table S1, ΔS150 values were significantly higher in HCC samples than controls, while ΔMCCCA values showed an opposite trend. In addition, the N-index values were significantly higher in HCC samples than controls (Figure 2F) and positively correlated with estimated tumor DNA fractions (Figure S2G); moreover, in HCC samples, both the genomic regions with detectable copy-number aberrations and estimated tumor DNA fractions were significantly elevated with end selection (Figures S2H and S2I). Considering that tumor-derived cfDNA molecules are shorter27 with lower CCCA end motif usage,13 the results demonstrated that the tumor-derived cfDNA molecules were indeed enriched with end selection. Notably, end selection based on the 5′ end only (see STAR Methods) or using the supplementary nucleosome track showed consistent results (Figure S3; Table S1). In addition, no significant differences in N-index, ΔS150, or ΔMCCCA values were observed among non-cancerous samples generated in different batches (Figure S4).

Diagnosis of HCC with end selection

Considering that, with end selection, the N-index, ΔS150, and ΔMCCCA values were all significantly different between patients with HCC and controls (Figures 2D–2F), we investigated the diagnostic potential of these metrics using receiver operating characteristics (ROC) analysis. As shown in Figures 3A–3C and Table S1, all N-index, ΔS150, and ΔMCCCA values were capable of differentiating patients with HCC from controls. For instance, areas under the curves (AUCs) for the N-index in differentiating patients with HCC from controls ranged from 0.72 to 0.94 for different stages (all p < 0.01, Z-tests). Notably, the ROCs were comparable between ΔS150 and directly measuring the size patterns of cfDNA (p > 0.1 for all stages, DeLong tests; Figure S5A), while they were significantly improved in ΔMCCCA compared to directly measuring the CCCA end motif usages (p < 0.05 for all stages, DeLong tests; Figure S5B). In addition, we found that the combination of these metrics could further improve the diagnostic performance, especially sensitivity in the detection of early-stage cancers. For instance, for detecting patients with stage A HCC at 100% specificity, the sensitivities of ΔS150 and ΔMCCCA alone were 29.0% and 51.6%, respectively (Figures 3B and 3C), while 64.5% sensitivity was achieved when these two metrics were utilized in a straightforward combination, as illustrated in Figure 3D.

Figure 3.

Figure 3

End selection for diagnosis of HCC

(A–C) Receiver operating characteristic (ROC) curves of (A) N-index, (B) ΔS150, and (C) ΔMCCCA metrics in differentiating patients with HCC from controls.

(D) An illustration approach to combine ΔS150 and ΔMCCCA for improved diagnosis of HCC.

(E and F) In Liang et al. dataset, the (E) N-index values between patients with HCC and controls and (F) ROC curves of N-index and ΔS150 metrics for diagnosis.

(G and H) In Zhang et al. dataset, the (G) N-index values between patients with HCC and controls and (H) ROC curves of N-index, ΔS150, and ΔMCCCA for diagnosis.

p values for areas under the curve (AUCs) were calculated using Z-tests. p values comparing N-index values (E and G) were calculated using Mann-Whitney U tests.

To validate the results, two external HCC datasets were collected and analyzed: the Liang et al. dataset31 contained cfDNA samples from 10 controls and 10 patients with HCC and the Zhang et al. dataset32 contained cfDNA whole-genome bisulfite sequencing (WGBS) data for 37 control subjects and 8 patients with HCC. As shown in Figures 3E–3H and S5C–S5F and Table S1, concordant results to our HCC cohort were obtained: in both datasets, patients with HCC showed significantly higher N-index and ΔS150 values than controls (all p < 0.01, Mann-Whitney U tests), and the AUCs for these metrics in differentiating patients with HCC from controls ranged from 0.87 to 0.93 (both p < 10−5, Z-test). ΔMCCCA values were significantly decreased in patients with HCC compared with controls in the Zhang et al. dataset (Figure S5F), demonstrating its capacity to differentiate between samples in these two categories (Figure 3H); a consistent trend could also be observed in the Liang et al. dataset, though the difference did not reach statistical significance. Together, the results demonstrated that our metrics could serve as biomarkers for diagnosis of HCC.

End selection for diagnosis of other cancer types

To explore whether the N-index could work for other cancers besides HCC, a large dataset containing 343 colorectal cancer (CRC) samples in various stages and 274 non-cancerous controls (including benign lesions) were collected from Walker et al.33 As shown in Figures 4A, 4B, S6A, and S6B and Table S1, the N-index and ΔS150 metrics were significantly increased in patients with CRC compared to controls (all p < 0.01, Mann-Whitney U tests) and showed the capacity to differentiate patients with CRC of different stages from the non-cancerous controls (all p < 0.01, Z-tests). ΔMCCCA values were significantly decreased in patients with CRC of stages II–IV than controls and demonstrated a capacity for diagnosis (Figures S6C and S6D).

Figure 4.

Figure 4

The N-index metric for pan-cancer diagnosis

(A and B) N-index values between controls and patients with colorectal cancer of different stages in Walker et al. dataset (A) and corresponding receiver operator characteristic (ROC) curves for cancer diagnosis (B).

(C and D) N-index values between controls and patients with various types of cancer in Song et al. dataset (C) and ROC curves of N-index, ΔS150, and ΔMCCCA for cancer diagnosis (D). Due to the limited sample size for each cancer type, all cancer samples were grouped together for diagnostic analyses.

(E and F) N-index values between controls and patients with various types of cancer in Cristiano et al. dataset (E) and ROC curves of N-index for diagnoses of each cancer type (F).

In (A), (C), and (E), p values were calculated using Mann-Whitey U tests. In (B), (D), and (F), p values were calculated using Z-tests.

Inspired by the results on HCC and CRC, we further investigated the potential of these three metrics for pan-cancer diagnosis. To this end, two datasets were analyzed: the Song et al. dataset34 contained 6 controls and 39 patients with cancer from 7 types and the Cristiano et al. dataset23 consisted of 215 controls and 208 patients from 7 cancer types of various stages. As expected, in the Song et al. dataset, all N-index, ΔS150, and ΔMCCCA values were significantly different in patients with cancer compared to controls with diagnostic potential (Figures 4C, 4D, S6E, and S6F; Table S1). In the Cristiano et al. dataset, significantly elevated N-index values were observed in all cancer types except gastric cancer (all p < 10−5, Mann-Whitney U tests; Figures 4E and S6G–S6J), which enabled diagnosis in these cancer types (all p < 10−5, Z-tests; Figure 4F); similarly, ΔS150 and ΔMCCCA values were also significantly altered in these cancer types and showed a capacity for cancer diagnosis (Figures S6K–S6N). The results demonstrated that all our N-index, ΔS150, and ΔMCCCA metrics held high potential as pan-cancer diagnostic biomarkers.

Enhancing diagnostic performance with machine learning

Previous studies have shown that high-performance diagnostic models could be built with machine learning, e.g., the DELFI (DNA evaluation of fragments for early interception) approach23; we wondered whether we could further improve the diagnostic performance by integrating machine learning into our metrics. To this end, we divided the genome into 5 Mb windows; for each sample, we calculated the N-index values per window and then employed a gradient boosting decision tree (GBDT) algorithm to utilize the N-index values for differentiating patients with cancer from controls. We named this approach EXCEL (examination of cfDNA with end selection), which generated models with scores to classify whether the samples were cancerous or not. We estimated the performance characteristics of the EXCEL models by 10-fold cross-validation repeated 10 times, as in Cristiano et al.23 Briefly, the EXCEL model on N-index values per 5 Mb windows showed an overall AUC value of 0.95 (95% confidence interval: 0.93–0.97; Figure 5A; Table S3), which was comparable to DELFI (0.94; p = 0.19, DeLong test) and significantly higher than genome-wide N-index metrics (p < 10−5, DeLong test). The AUCs were at least 0.92 for all cancers and stages (Figures 5B, 5C, and S7A). Moreover, EXCEL detected 154 out of 181 patients with cancer (sensitivity = 85.1%, 95% confidence interval: 79.0%–89.9%; Table 1) at 95% specificity, which was higher than the DELFI algorithm (79.8%) on the same dataset; similarly, an 84.8% sensitivity for patients with resectable (stage I–III) tumors is achieved (Tables 1 and S3), which was also higher than DELFI (79.2%). In addition, we also built EXCEL models using ΔS150 or ΔMCCCA metrics per 5 Mb window and achieved overall AUC values of 0.88 and 0.81, respectively, in differentiating cancer samples from controls, which were significantly higher than those using the genome-wide metrics without machine learning (Figures S7B–S7I).

Figure 5.

Figure 5

Cancer diagnosis using EXCEL

(A) Receiver operator characteristic (ROC) curve of EXCEL model using N-index values per 5 Mb window in a cohort with 215 healthy individuals and 181 patients with cancer. ROC curves for DELFI and genome-wide N-index metrics on this cohort are also shown.

(B and C) ROC curves of the EXCEL model for (B) each cancer type versus non-cancerous controls and (C) cancer samples in each stage.

Dashed lines in each image indicated 95% specificity. p values for areas under the curve (AUCs) were calculated using Z-tests.

Table 1.

Performance of EXCEL model with N-index metric for cancer detection

Individuals analyzed No. individuals analyzed 95% specificity
No. individuals detected Sensitivity (%) 95% CI (%)
Healthy 215 10
Cancer 181 154 85.1 79.0–89.9

Type

 Bile duct 26 21 80.8 60.6–93.4
 Breast 54 47 87.0 75.1–94.6
 Colorectal 27 22 81.5 61.9–93.7
 Lung 12 12 100.0 73.5–100.0
 Ovarian 28 25 89.3 71.8–97.7
 Pancreatic 34 27 79.4 62.1–91.3

Stage

 I 37 29 78.4 61.8–90.2
 II 103 86 83.5 74.9–90.1
 III 25 25 100.0 86.3–100.0
 IV 16 14 87.5 61.7–98.4

Discussion

In this proof-of-concept study, we investigated the biology of plasma cfDNA ends regarding the nucleosome context and explored the translational values of end selection in cancer diagnosis. End selection could enrich the short fragments in cfDNA with altered CCCA end motif usage (Figure 2); more importantly, changes in cfDNA size and motif patterns with end selection were different between control subjects and patients with cancer. Hence, in cancer samples, we observed higher increases in short fragments (i.e., ΔS150 values) and lower increases in CCCA end motif usage (i.e., ΔMCCCA values) with end selection, phenomena that were highly consistent with the fact that tumor-derived cfDNA molecules are shorter in size along with having lower CCCA end motif usage. In addition, the proportion of cfDNA fragments with discordant ends to nucleosomes (i.e., N-index value), either derived from GM12878 cells or deduced using cfDNA coverage patterns, was higher in cancer samples, and the tumor DNA fractions were also elevated in cfDNA fragments with end selection (Figure S2), suggesting that end selection could enrich tumor-derived cfDNA in patients with cancer. Interestingly, an increase of N-index and ΔS150 values and a decrease of the ΔMCCCA value were observed in most cancers, except for gastric cancer in the Cristiano et al. dataset. In fact, compared to controls, these gastric cancer samples showed a significantly decreased proportion of short fragments (i.e., their cfDNA molecules were longer than controls; Figure S6J), which was inconsistent with the shortness nature of tumor-derived cfDNA molecules; hence, in these gastric cancer samples, there would be other sources contributing to long cfDNA molecules, which might be the reason for the abnormalities in the N-index, ΔS150, and ΔMCCCA metrics. Together, these results reveal that end selection introduces significant changes in various cfDNA fragmentomic characteristics and justifies the importance of end patterns in cfDNA biology with inherent links to cfDNA size.

In addition, we also demonstrated the translational values of N-index, ΔS150, and ΔMCCCA metrics as promising biomarkers for pan-cancer diagnosis (Figures 3 and 4). In fact, solid tumors investigated in this study all possess a non-hematopoietic origin; in addition, dysregulation in chromatin remodelers is common in tumors, which leads to altered nucleosome positioning.35 These biological principles might serve as the underlying mechanism supporting our metrics for pan-cancer diagnosis. On the other hand, although these three metrics are also based on end analysis of cfDNA like our previous E-index metric,17 they model the cfDNA end pattern from different perspectives. Of note, these three metrics take the biological background of cfDNA (i.e., nucleosome positioning) into the model instead of a panel of controls, suggesting a platform-independent advantage over E-index and DELFI metrics. In addition, Mouliere et al. demonstrated the translational value of cfDNA size selection for liquid biopsy.29 In vitro size selection could be powerful in enriching tumor-derived cfDNA for precision oncology, though it requires specific experimental procedures; for in silico size selection, as short fragments (i.e., ≤150 bp) only account for ∼15%–20% of total cfDNA, one may only have ∼5 million reads for downstream analyses for samples with ∼30 million reads (e.g., most samples in the Cristiano et al. dataset), which would adversely affect the reliability and stability of cfDNA fragmentomic profiling and limit its applicability to low-pass sequencing datasets. As a contrast, we showed that end selection could also enrich tumor-derived cfDNA molecules, and in most datasets, the N-index values are around 0.9 (Figures 2, 3, and 4), meaning that our end selection would recover ∼90% of reads for downstream analyses. Furthermore, using a cohort of ∼400 samples, we showed that integrating these metrics with machine learning approaches achieved significantly improved diagnostic performances in pan-cancer (Figures 5 and S7; Table 1). Hence, these results demonstrated that the three metrics developed in this work could all serve as valuable elements in the construction of large multi-biomarker, high-performance diagnostic models.36

Limitations of the study

Due to the lack of high-quality hematopoietic nucleosome positioning profiles currently, the nucleosome track from the GM12878 cell line was utilized as a surrogate in this study. Although this track had been widely used in cfDNA fragmentomic analyses, it is far from perfect, with certain defects. For instance, GM12878 cells are rather different from primary B cells,37 and DNA cut by MNase could be different from that cut by endogenous nucleases,38 and this track could not capture the heterogeneity of hematopoietic cells contributing to the plasma cfDNA pool (which includes lymphoid, myeloid, and erythroid lineages).6,39 We employed a supplementary track deduced from cfDNA coverage patterns to validate the results (Figure S3), while studies incorporating more types of nucleosome tracks are critical to comprehensively investigate the relationship between hematopoietic nucleosome profiles and cfDNA fragmentomics as well as further optimize the diagnostic models. In addition, systematic biases across datasets, usually caused by discrepancies in experimental protocols and sequencing platforms, could significantly affect cfDNA fragmentomic features, e.g., leading to baseline differences in controls among datasets (Figures 2, 3, and 4). In fact, such pre-analytical biases are known to adversely affect the generalizability of diagnostic models.40,41,42,43 Therefore, coherent and unified sample processing protocols, sequencing platforms, and advanced computational approaches would be essential to overcome dataset-specific biases toward the development of high-performance cross-dataset diagnostic models. On the other hand, generalizable diagnostic models might be achieved through substantially increasing the sample size such that the technical noises and biases could be well recognized and properly handled during model building.

Resource availability

Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Kun Sun (sunkun@szbl.ac.cn).

Materials availability

This study did not generate new unique reagents.

Data and code availability

  • Raw sequencing data generated in this study have been deposited in the Genome Sequence Archive in the National Genomics Data Center (GSA-Human) under accession number HRA005521. The data are under controlled access due to patient consent restrictions, and applications for data access should be sent to Kun Sun (sunkun@szbl.ac.cn). This work also analyzes existing publicly available data. The accession numbers for the datasets are listed in the key resources table.

  • Computational programs and scripts to reproduce the results are publicly available at https://github.com/hellosunking/cfDNA-end-selection and https://zenodo.org/doi/10.5281/zenodo.13579355.

  • Any additional information required to reanalyze the data reported in this work paper is available from the corresponding author upon request.

Acknowledgments

This work was supported by the National Key R&D Program of China (2022YFA0912700), the Guangdong Basic and Applied Basic Research Foundation (2023B1515120073), the National Natural Science Foundation of China (82300656), the Science, Technology and Innovation Commission of Shenzhen Municipality (JCYJ20220530163215034), and the Major Program of Shenzhen Bay Laboratory (S241101004). We would like to thank Ms. Qi Wang for technical assistance, the SZBL HPC Core for computational support, and the SZBL Sequencing Core for DNA sequencing support.

Author contributions

Conception and design, K.S.; study supervision, K.S.; development of methodology, J.J., Y.A., M.Y., and K.S.; patient recruitment and clinical data analysis, Z.Z. and X.Z.; acquisition of data, all authors; analysis and interpretation of data, all authors; writing of the manuscript, M.Y., J.J., and K.S.

Declaration of interests

Y.A., J.J., and K.S. have filed patent applications based on the method developed in this work.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data

CfDNA dataset 1 An et al.17 NGDC GSA-Human: HRA002250
CfDNA dataset 2 This paper NGDC GSA-Human: HRA005521
CfDNA dataset 3 Liang et al.31 CNGBdb CNSA: CNP0000680
CfDNA dataset 4 Zhang et al.32 NGDC GSA: CRA001537
CfDNA dataset 5 Walker et al.33 NCBI SRA: PRJNA755688
CfDNA dataset 6 Song et al.34 NCBI GEO: GSE81314
CfDNA dataset 7 Cristiano et al.23 FinaleDB: STUDY Cristiano et al., 201923
Nucleosome track for GM12878 cell line Zhao et al.44 NucMap database: hsNuc0390101
Supplementary nucleosome track Snyder et al.9 NCBI GEO: GSE71378

Software and algorithms

Ktrim Sun et al.45 https://github.com/hellosunking/Ktrim/
Bowtie2 Langmead et al.46 https://github.com/BenLangmead/bowtie2
Msuite2 Li et al.47,48 https://github.com/hellosunking/Msuite2
ichorCNA Adalsteinsson et al.49 https://github.com/broadinstitute/ichorCNA
End selection and fragmentomic analysis This paper https://github.com/hellosunking/cfDNA-end-selection
https://zenodo.org/doi/10.5281/zenodo.13579355

Experimental model and study participant details

Ethics approval and sample processing

This study had been approved by the Ethics Committee of Shenzhen Bay Laboratory and Ethics Committee of The Third People’s Hospital of Shenzhen. Participants were recruited from The Third People’s Hospital of Shenzhen with written informed consents (N = 80; age range: 24–78, median:51; female proportion: 26.3%; Table S2). Notably, most of the non-cancerous controls were Hepatitis B Virus carriers and/or suffered from benign liver lesions. For each subject, 10 mL peripheral blood was collected using EDTA-containing tubes and processed within 4 h50; briefly, blood samples were centrifuged at 1,600 g, 4°C for 15 min, and then the plasma portion was harvested and re-recentrifuged at 16,000 g, 4°C for 15 min to remove blood cells17,50; plasma samples were stored at −80°C until further usage. For each sample, 600 μL plasma was used to extract cfDNA with MagPure Circulating DNA Mini Kit (Magen, #IVD5432), and DNA library was constructed using VAHTS Universal DNA Library Prep Kit for MGI (Vazyme, #NDM607-02) following the manufacturer’s instructions. The DNA libraries were used to prepare a single-stranded circular DNA using VAHTS Circularization Kit for MGI (Vazyme, #NM201-01) and subjected to DNA nanoball (DNB) generation, and then sequenced on an MGISEQ-2000 (MGI) sequencer with MGISEQ-2000RS Sequencing Reagent (MGI) in paired-end 100 bp mode.

Method details

sequencing data processing

Raw paired-end sequencing data was processed as reported previously.17 Briefly, cfDNA whole genome sequencing data was preprocessed using Ktrim software (version 1.5.0)45 and aligned to the NCBI GRCh38 human reference genome (patch 13) using Bowtie2 software (version 2.3.5.1)46; PCR duplicates (i.e., sequencing reads with identical start and end positions) were then removed using in-house programs,43 and only the uniquely mapped reads with mapping quality scores larger than or equal to 30 and located on the autosomes were used in the downstream fragmentomic analyses. Tumor DNA fractions in HCC samples were deduced using ichorCNA algorithm49 with the non-cancerous samples as reference (“panel of normal”). WGBS data was analyzed using Msuite2 software (version 2.2.0)47,48 against GRCh38 human reference genome. In addition, to eliminate the dosage discrepancy in sex chromosomes between male and female subjects, only autosomes were used in downstream analyses. Note that Song et al. study provided both hydroxymethylation-capture and whole-genome sequencing data,34 while only the whole-genome sequencing data was used in this study.

Nucleosome tracks and end selection

Two nucleosome tracks were investigated in this study. The primary nucleosome track was obtained from GM12878 cell line (lymphoblastoid lineage).9,17,51 This nucleosome track was downloaded from NucMap database,44 which compiled high-depth MNase-seq data from ENCODE project52 using DANPOS algorithm with default parameters53 and recorded ∼13.9 million nucleosome-protected regions. The supplementary nucleosome track was obtained from Snyder et al. that was directly deduced using cfDNA coverage pattern.9 The “CH01” track (which contained ∼12.9 million nucleosome-protected regions) was used and the genomic coordinates were translated from NCBI GRCh37 human reference genome to GRCh38 using UCSC genome browser’s “liftOver” utility.54 For both nucleosome tracks, we defined nucleosome-protected regions as ±73 bp around the annotated center loci in the nucleosome track as in previous studies.9,17,51

To perform end selection, for each sample, we extracted the ends of each cfDNA fragment, and then compared them to the nucleosome track independently and picked up those fragments with at least one end located within the nucleosome-protected regions. We defined N-index metric as the proportion of cfDNA fragments passed this end selection procedure as shown in the following formula (Figure S1A):

N-index=NumberoffragmentswithendslocatedinnuclesomeprotectedregionsTotalnumberoffragments

Note that we also explored an alternative implementation that only picked up the fragments whose 5′-ends were located within the nucleosome-protected regions (where the 3′-ends were ignored during the analysis), and the results on the HCC cohort were shown in Figure S3.

In addition, we investigated the size distribution and end motif patterns of the fragments after end selection and compared these fragmentomic features to the original data without end selection. Briefly, the sizes of cfDNA molecules were determined as the distances between the genomic coordinates of the 2 outmost ends, and the sequences of 4 nucleotides from the outmost ends in the reference genome were extracted to calculate 4-mer motif frequencies.13 CfDNA molecules no bigger than 150 bp (i.e., ≤150 bp) in size were considered as “short fragments”; for motif analysis, we focused on CCCA as it was the most widely studied motif with proven diagnostic values.13,15,55 Hence, two metrics were defined to quantify the changes in these two features after end selection: ΔS150 measured the change in percent of fragments no bigger than 150 bp (Figure S1B), and ΔMCCCA measured the change in percent of CCCA end motif usage (Figure S1C) as shown in the following formular.

  • ΔS150 = Percent of fragments no bigger than 150 bp post end selection - Percent of fragments no bigger than 150 bp prior end selection

  • ΔMCCCA = Percent of fragments with CCCA end motif post end selection - Percent of fragments with CCCA end motif prior end selection

For cancer diagnoses, changes in fragmentomic profiles after end selection and N-index values were compared across samples of different types of cancer versus controls, and the performances for diagnoses were evaluated using ROC analyses.

Machine learning-based diagnostic models

The human genome was divided into 5 Mb non-overlapping windows; windows with extremely low reads (e.g., telomeres, centromeres) were discarded and 542 windows remained for downstream analyses (covering approximately 2.7 Gb of the genome). For each sample, we performed end selection and calculated the N-index, ΔS150, and ΔMCCCA values for each window. The gradient boosted decision tree (GBDT) algorithm, implemented by the package “gbm” (version 2.1.9) and “caret” (version 6.0–94) in R software (version 4.2.0), was employed to build three diagnostic models on N-index, ΔS150, or ΔMCCCA values separately (parameters: n.trees = 200, interaction.depth = 3, shrinkage = 0.3, n.minobsinnode = 15). Features that displayed correlation above 0.9 to each other or had near zero variance within training dataset were removed during cross-validation as in Cristiano et al.23 Performance characteristics were evaluated by 10-fold cross-validation repeated 10 times as in Cristiano et al.23 Confidence intervals for sensitivity with specificity fixed at 95% were calculated using exact binomial tests implemented by the “binom.test” function in R software. Note that due to aberrant fragmentomic features in gastric cancer samples (Figure S6), these samples were omitted from the analyses; the rest cancer samples (N = 181) were used in building EXCEL models, and DELFI scores in these samples were extracted to perform ROC analysis for comparison with EXCEL.

Quantification and statistical analysis

All procedures involving statistical analysis were described in the STAR Methods section, and number of samples used in each test was described in the results sections.

Published: October 14, 2024

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.crmeth.2024.100877.

Supplemental information

Document S1. Figures S1‒S7
mmc1.pdf (3.2MB, pdf)
Table S1. Details of N-index, ΔS150, and ΔMCCCA analyzed in this study, related to Figures 2, 3, and 4
mmc2.xlsx (214.3KB, xlsx)
Table S2. Clinical information for samples recruited in this study, related to Figure 2 and STAR Methods
mmc3.xlsx (12.2KB, xlsx)
Table S3. EXCEL scores on the Cristiano et al. dataset, related to Figure 5 and Table 1
mmc4.xlsx (54.6KB, xlsx)
Document S2. Article plus supplemental information
mmc5.pdf (7.4MB, pdf)

References

  • 1.Wan J.C.M., Massie C., Garcia-Corbacho J., Mouliere F., Brenton J.D., Caldas C., Pacey S., Baird R., Rosenfeld N. Liquid biopsies come of age: towards implementation of circulating tumour DNA. Nat. Rev. Cancer. 2017;17:223–238. doi: 10.1038/nrc.2017.7. [DOI] [PubMed] [Google Scholar]
  • 2.Lo Y.M.D., Han D.S.C., Jiang P., Chiu R.W.K. Epigenetics, fragmentomics, and topology of cell-free DNA in liquid biopsies. Science. 2021;372:eaaw3616. doi: 10.1126/science.aaw3616. [DOI] [PubMed] [Google Scholar]
  • 3.Hawkes N. Cancer survival data emphasise importance of early diagnosis. BMJ. 2019;364:l408. doi: 10.1136/bmj.l408. [DOI] [PubMed] [Google Scholar]
  • 4.van der Pol Y., Mouliere F. Toward the Early Detection of Cancer by Decoding the Epigenetic and Environmental Fingerprints of Cell-Free DNA. Cancer Cell. 2019;36:350–368. doi: 10.1016/j.ccell.2019.09.003. [DOI] [PubMed] [Google Scholar]
  • 5.Lui Y.Y.N., Chik K.W., Chiu R.W.K., Ho C.Y., Lam C.W.K., Lo Y.M.D. Predominant hematopoietic origin of cell-free DNA in plasma and serum after sex-mismatched bone marrow transplantation. Clin. Chem. 2002;48:421–427. [PubMed] [Google Scholar]
  • 6.Sun K., Jiang P., Chan K.C.A., Wong J., Cheng Y.K.Y., Liang R.H.S., Chan W.K., Ma E.S.K., Chan S.L., Cheng S.H., et al. Plasma DNA tissue mapping by genome-wide methylation sequencing for noninvasive prenatal, cancer, and transplantation assessments. Proc. Natl. Acad. Sci. USA. 2015;112:E5503–E5512. doi: 10.1073/pnas.1508736112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Moss J., Magenheim J., Neiman D., Zemmour H., Loyfer N., Korach A., Samet Y., Maoz M., Druid H., Arner P., et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease. Nat. Commun. 2018;9:5068. doi: 10.1038/s41467-018-07466-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Loyfer N., Magenheim J., Peretz A., Cann G., Bredno J., Klochendler A., Fox-Fisher I., Shabi-Porat S., Hecht M., Pelet T., et al. A DNA methylation atlas of normal human cell types. Nature. 2023;613:355–364. doi: 10.1038/s41586-022-05580-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Snyder M.W., Kircher M., Hill A.J., Daza R.M., Shendure J. Cell-free DNA comprises an in vivo nucleosome footprint that informs its tissues-of-origin. Cell. 2016;164:57–68. doi: 10.1016/j.cell.2015.11.050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Chandrananda D., Thorne N.P., Bahlo M. High-resolution characterization of sequence signatures due to non-random cleavage of cell-free DNA. BMC Med. Genom. 2015;8:29. doi: 10.1186/s12920-015-0107-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sun K., Jiang P., Cheng S.H., Cheng T.H.T., Wong J., Wong V.W.S., Ng S.S.M., Ma B.B.Y., Leung T.Y., Chan S.L., et al. Orientation-aware plasma cell-free DNA fragmentation analysis in open chromatin regions informs tissue of origin. Genome Res. 2019;29:418–427. doi: 10.1101/gr.242719.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ulz P., Thallinger G.G., Auer M., Graf R., Kashofer K., Jahn S.W., Abete L., Pristauz G., Petru E., Geigl J.B., et al. Inferring expressed genes by whole-genome sequencing of plasma DNA. Nat. Genet. 2016;48:1273–1278. doi: 10.1038/ng.3648. [DOI] [PubMed] [Google Scholar]
  • 13.Jiang P., Sun K., Peng W., Cheng S.H., Ni M., Yeung P.C., Heung M.M.S., Xie T., Shang H., Zhou Z., et al. Plasma DNA End-Motif Profiling as a Fragmentomic Marker in Cancer, Pregnancy, and Transplantation. Cancer Discov. 2020;10:664–673. doi: 10.1158/2159-8290.CD-19-0622. [DOI] [PubMed] [Google Scholar]
  • 14.Jin X., Wang Y., Xu J., Li Y., Cheng F., Luo Y., Zhou H., Lin S., Xiao F., Zhang L., et al. Plasma cell-free DNA promise monitoring and tissue injury assessment of COVID-19. Mol. Genet. Genom. 2023;298:823–836. doi: 10.1007/s00438-023-02014-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Guo W., Chen X., Liu R., Liang N., Ma Q., Bao H., Xu X., Wu X., Yang S., Shao Y., et al. Sensitive detection of stage I lung adenocarcinoma using plasma cell-free DNA breakpoint motif profiling. EBioMedicine. 2022;81:104131. doi: 10.1016/j.ebiom.2022.104131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Jin C., Liu X., Zheng W., Su L., Liu Y., Guo X., Gu X., Li H., Xu B., Wang G., et al. Characterization of fragment sizes, copy number aberrations and 4-mer end motifs in cell-free DNA of hepatocellular carcinoma for enhanced liquid biopsy-based cancer detection. Mol. Oncol. 2021;15:2377–2389. doi: 10.1002/1878-0261.13041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.An Y., Zhao X., Zhang Z., Xia Z., Yang M., Ma L., Zhao Y., Xu G., Du S., Wu X., et al. DNA methylation analysis explores the molecular basis of plasma cell-free DNA fragmentation. Nat. Commun. 2023;14:287. doi: 10.1038/s41467-023-35959-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Sun K., Jiang P., Wong A.I.C., Cheng Y.K.Y., Cheng S.H., Zhang H., Chan K.C.A., Leung T.Y., Chiu R.W.K., Lo Y.M.D. Size-tagged preferred ends in maternal plasma DNA shed light on the production mechanism and show utility in noninvasive prenatal testing. Proc. Natl. Acad. Sci. USA. 2018;115:E5106–E5114. doi: 10.1073/pnas.1804134115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Jiang P., Sun K., Tong Y.K., Cheng S.H., Cheng T.H.T., Heung M.M.S., Wong J., Wong V.W.S., Chan H.L.Y., Chan K.C.A., et al. Preferred end coordinates and somatic variants as signatures of circulating tumor DNA associated with hepatocellular carcinoma. Proc. Natl. Acad. Sci. USA. 2018;115:E10925–E10933. doi: 10.1073/pnas.1814616115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chan K.C.A., Jiang P., Sun K., Cheng Y.K.Y., Tong Y.K., Cheng S.H., Wong A.I.C., Hudecova I., Leung T.Y., Chiu R.W.K., Lo Y.M.D. Second generation noninvasive fetal genome analysis reveals de novo mutations, single-base parental inheritance, and preferred DNA ends. Proc. Natl. Acad. Sci. USA. 2016;113:E8159–E8168. doi: 10.1073/pnas.1615800113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ulz P., Perakis S., Zhou Q., Moser T., Belic J., Lazzeri I., Wölfler A., Zebisch A., Gerger A., Pristauz G., et al. Inference of transcription factor binding from cell-free DNA enables tumor subtype prediction and early detection. Nat. Commun. 2019;10:4666. doi: 10.1038/s41467-019-12714-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Esfahani M.S., Hamilton E.G., Mehrmohamadi M., Nabet B.Y., Alig S.K., King D.A., Steen C.B., Macaulay C.W., Schultz A., Nesselbush M.C., et al. Inferring gene expression from cell-free DNA fragmentation profiles. Nat. Biotechnol. 2022;40:585–597. doi: 10.1038/s41587-022-01222-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Cristiano S., Leal A., Phallen J., Fiksel J., Adleff V., Bruhm D.C., Jensen S.Ø., Medina J.E., Hruban C., White J.R., et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature. 2019;570:385–389. doi: 10.1038/s41586-019-1272-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.An Y., Fan F., Jiang X., Sun K. Recent Advances in Liquid Biopsy of Brain Cancers. Front. Genet. 2021;12:720270. doi: 10.3389/fgene.2021.720270. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Gai W., Sun K. Epigenetic Biomarkers in Cell-Free DNA and Applications in Liquid Biopsy. Genes. 2019;10:32. doi: 10.3390/genes10010032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Han D.S.C., Lo Y.M.D. The Nexus of cfDNA and Nuclease Biology. Trends Genet. 2021;37:758–770. doi: 10.1016/j.tig.2021.04.005. [DOI] [PubMed] [Google Scholar]
  • 27.Underhill H.R., Kitzman J.O., Hellwig S., Welker N.C., Daza R., Baker D.N., Gligorich K.M., Rostomily R.C., Bronner M.P., Shendure J. Fragment length of circulating tumor DNA. PLoS Genet. 2016;12:e1006162. doi: 10.1371/journal.pgen.1006162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Lo Y.M.D., Chan K.C.A., Sun H., Chen E.Z., Jiang P., Lun F.M.F., Zheng Y.W., Leung T.Y., Lau T.K., Cantor C.R., Chiu R.W.K. Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus. Sci. Transl. Med. 2010;2:61ra91. doi: 10.1126/scitranslmed.3001720. [DOI] [PubMed] [Google Scholar]
  • 29.Mouliere F., Chandrananda D., Piskorz A.M., Moore E.K., Morris J., Ahlborn L.B., Mair R., Goranova T., Marass F., Heider K., et al. Enhanced detection of circulating tumor DNA by fragment size analysis. Sci. Transl. Med. 2018;10:eaat4921. doi: 10.1126/scitranslmed.aat4921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Reig M., Forner A., Rimola J., Ferrer-Fàbrega J., Burrel M., Garcia-Criado Á., Kelley R.K., Galle P.R., Mazzaferro V., Salem R., et al. BCLC strategy for prognosis prediction and treatment recommendation: The 2022 update. J. Hepatol. 2022;76:681–693. doi: 10.1016/j.jhep.2021.11.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Liang H., Li F., Qiao S., Zhou X., Xie G., Zhao X., Zhang Y., Wu K. Whole-genome sequencing of cell-free DNA yields genome-wide read distribution patterns to track tissue of origin in cancer patients. Clin. Transl. Med. 2020;10:e177. doi: 10.1002/ctm2.177. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zhang H., Dong P., Guo S., Tao C., Chen W., Zhao W., Wang J., Cheung R., Villanueva A., Fan J., et al. Hypomethylation in HBV integration regions aids non-invasive surveillance to hepatocellular carcinoma by low-pass genome-wide bisulfite sequencing. BMC Med. 2020;18:200. doi: 10.1186/s12916-020-01667-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Walker N.J., Rashid M., Yu S., Bignell H., Lumby C.K., Livi C.M., Howell K., Morley D.J., Morganella S., Barrell D., et al. Hydroxymethylation profile of cell-free DNA is a biomarker for early colorectal cancer. Sci. Rep. 2022;12:16566. doi: 10.1038/s41598-022-20975-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Song C.X., Yin S., Ma L., Wheeler A., Chen Y., Zhang Y., Liu B., Xiong J., Zhang W., Hu J., et al. 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages. Cell Res. 2017;27:1231–1242. doi: 10.1038/cr.2017.106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Mittal P., Roberts C.W.M. The SWI/SNF complex in cancer - biology, biomarkers and therapy. Nat. Rev. Clin. Oncol. 2020;17:435–448. doi: 10.1038/s41571-020-0357-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Jamshidi A., Liu M.C., Klein E.A., Venn O., Hubbell E., Beausang J.F., Gross S., Melton C., Fields A.P., Liu Q., et al. Evaluation of cell-free DNA approaches for multi-cancer early detection. Cancer Cell. 2022;40:1537–1549.e12. doi: 10.1016/j.ccell.2022.10.022. [DOI] [PubMed] [Google Scholar]
  • 37.Piroeva K.V., McDonald C., Xanthopoulos C., Fox C., Clarkson C.T., Mallm J.P., Vainshtein Y., Ruje L., Klett L.C., Stilgenbauer S., et al. Nucleosome repositioning in chronic lymphocytic leukemia. Genome Res. 2023;33:1649–1661. doi: 10.1101/gr.277298.122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Zhu D., Wang H., Wu W., Geng S., Zhong G., Li Y., Guo H., Long G., Ren Q., Luan Y., et al. Circulating cell-free DNA fragmentation is a stepwise and conserved process linked to apoptosis. BMC Biol. 2023;21:253. doi: 10.1186/s12915-023-01752-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Lam W.K.J., Gai W., Sun K., Wong R.S.M., Chan R.W.Y., Jiang P., Chan N.P.H., Hui W.W.I., Chan A.W.H., Szeto C.C., et al. DNA of erythroid origin is present in human plasma and informs the types of anemia. Clin. Chem. 2017;63:1614–1623. doi: 10.1373/clinchem.2017.272401. [DOI] [PubMed] [Google Scholar]
  • 40.Mathios D., Johansen J.S., Cristiano S., Medina J.E., Phallen J., Larsen K.R., Bruhm D.C., Niknafs N., Ferreira L., Adleff V., et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat. Commun. 2021;12:5060. doi: 10.1038/s41467-021-24994-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Bohers E., Viailly P.J., Jardin F. cfDNA Sequencing: Technological Approaches and Bioinformatic Issues. Pharmaceuticals. 2021;14:596. doi: 10.3390/ph14060596. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Parpart-Li S., Bartlett B., Popoli M., Adleff V., Tucker L., Steinberg R., Georgiadis A., Phallen J., Brahmer J., Azad N., et al. The Effect of Preservative and Temperature on the Analysis of Circulating Tumor DNA. Clin. Cancer Res. 2017;23:2471–2477. doi: 10.1158/1078-0432.CCR-16-1691. [DOI] [PubMed] [Google Scholar]
  • 43.Liu X., Yang M., Hu D., An Y., Wang W., Lin H., Pan Y., Ju J., Sun K. Systematic biases in reference-based plasma cell-free DNA fragmentomic profiling. Cell Rep. Methods. 2024;4:100793. doi: 10.1016/j.crmeth.2024.100793. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Zhao Y., Wang J., Liang F., Liu Y., Wang Q., Zhang H., Jiang M., Zhang Z., Zhao W., Bao Y., et al. NucMap: a database of genome-wide nucleosome positioning map across species. Nucleic Acids Res. 2019;47:D163–D169. doi: 10.1093/nar/gky980. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Sun K. Ktrim: an extra-fast and accurate adapter- and quality-trimmer for sequencing data. Bioinformatics. 2020;36:3561–3562. doi: 10.1093/bioinformatics/btaa171. [DOI] [PubMed] [Google Scholar]
  • 46.Langmead B., Salzberg S.L. Fast gapped-read alignment with Bowtie 2. Nat. Methods. 2012;9:357–359. doi: 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Sun K., Li L., Ma L., Zhao Y., Deng L., Wang H., Sun H. Msuite: A High-Performance and Versatile DNA Methylation Data-Analysis Toolkit. Patterns. 2020;1:100127. doi: 10.1016/j.patter.2020.100127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Li L., An Y., Ma L., Yang M., Yuan P., Liu X., Jin X., Zhao Y., Zhang S., Hong X., Sun K. Msuite2: All-in-one DNA methylation data analysis toolkit with enhanced usability and performance. Comput. Struct. Biotechnol. J. 2022;20:1271–1276. doi: 10.1016/j.csbj.2022.03.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Adalsteinsson V.A., Ha G., Freeman S.S., Choudhury A.D., Stover D.G., Parsons H.A., Gydush G., Reed S.C., Rotem D., Rhoades J., et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat. Commun. 2017;8:1324. doi: 10.1038/s41467-017-00965-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Meddeb R., Pisareva E., Thierry A.R. Guidelines for the Preanalytical Conditions for Analyzing Circulating Cell-Free DNA. Clin. Chem. 2019;65:623–633. doi: 10.1373/clinchem.2018.298323. [DOI] [PubMed] [Google Scholar]
  • 51.Gaffney D.J., McVicker G., Pai A.A., Fondufe-Mittendorf Y.N., Lewellen N., Michelini K., Widom J., Gilad Y., Pritchard J.K. Controls of nucleosome positioning in the human genome. PLoS Genet. 2012;8:e1003036. doi: 10.1371/journal.pgen.1003036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Kundaje A., Kyriazopoulou-Panagiotopoulou S., Libbrecht M., Smith C.L., Raha D., Winters E.E., Johnson S.M., Snyder M., Batzoglou S., Sidow A. Ubiquitous heterogeneity and asymmetry of the chromatin environment at regulatory elements. Genome Res. 2012;22:1735–1747. doi: 10.1101/gr.136366.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Chen K., Xi Y., Pan X., Li Z., Kaestner K., Tyler J., Dent S., He X., Li W. DANPOS: dynamic analysis of nucleosome position and occupancy by sequencing. Genome Res. 2013;23:341–351. doi: 10.1101/gr.142067.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Kent W.J., Sugnet C.W., Furey T.S., Roskin K.M., Pringle T.H., Zahler A.M., Haussler D. The human genome browser at UCSC. Genome Res. 2002;12:996–1006. doi: 10.1101/gr.229102. Article published online before print in May 2002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Chen L., Abou-Alfa G.K., Zheng B., Liu J.F., Bai J., Du L.T., Qian Y.S., Fan R., Liu X.L., Wu L., et al. Genome-scale profiling of circulating cell-free DNA signatures for early detection of hepatocellular carcinoma in cirrhotic patients. Cell Res. 2021;31:589–592. doi: 10.1038/s41422-020-00457-7. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1‒S7
mmc1.pdf (3.2MB, pdf)
Table S1. Details of N-index, ΔS150, and ΔMCCCA analyzed in this study, related to Figures 2, 3, and 4
mmc2.xlsx (214.3KB, xlsx)
Table S2. Clinical information for samples recruited in this study, related to Figure 2 and STAR Methods
mmc3.xlsx (12.2KB, xlsx)
Table S3. EXCEL scores on the Cristiano et al. dataset, related to Figure 5 and Table 1
mmc4.xlsx (54.6KB, xlsx)
Document S2. Article plus supplemental information
mmc5.pdf (7.4MB, pdf)

Data Availability Statement

  • Raw sequencing data generated in this study have been deposited in the Genome Sequence Archive in the National Genomics Data Center (GSA-Human) under accession number HRA005521. The data are under controlled access due to patient consent restrictions, and applications for data access should be sent to Kun Sun (sunkun@szbl.ac.cn). This work also analyzes existing publicly available data. The accession numbers for the datasets are listed in the key resources table.

  • Computational programs and scripts to reproduce the results are publicly available at https://github.com/hellosunking/cfDNA-end-selection and https://zenodo.org/doi/10.5281/zenodo.13579355.

  • Any additional information required to reanalyze the data reported in this work paper is available from the corresponding author upon request.


Articles from Cell Reports Methods are provided here courtesy of Elsevier

RESOURCES