Abstract
Background
Differentiating benign prostatic hyperplasia (BPH) from prostate cancer (PCa) remains a diagnostic challenge due to overlapping clinical features and the limited specificity of prostate-specific antigen (PSA) testing. Recent studies highlight the promise of DNA methylation patterns in cell-free DNA (cfDNA) as non-invasive biomarkers for prostate disease stratification. This study aimed to develop a targeted methylation-based diagnostic model capable of distinguishing PCa from BPH using cfDNA. We hypothesized that specific genomic regions exhibit consistent hyper- or hypomethylation signatures that could serve as robust diagnostic markers.
Materials and methods
We applied a custom panel targeting 748 predefined differentially methylated regions (DMRs) to cfDNA from 365 participants (PCa: 230; BPH: 135). Methylation was quantified by targeted enzymatic methyl-sequencing. Following feature selection, a 43 DMR-based model was built using a stacking ensemble framework (Random Forest + XGBoost with stacking). Data were split by stratified sampling into training/test sets and a validation cohort. Uncertainty was quantified via 1,000-iteration non-parametric bootstraps. Functional relevance was assessed by gene ontology (GO) enrichment stratified by genomic context.
Results
The diagnostic model exhibited robust discriminative performance across both test and validation cohorts. In the test set, the model achieved an area under the receiver operating characteristic curve (ROC-AUC) of 0.99 and overall accuracy of 0.95, with balanced sensitivity (0.96) and specificity (0.93). Consistent results were observed in the validation set, with a ROC-AUC of 0.98, overall accuracy of 0.93, sensitivity of 0.94, and specificity of 0.93. Importantly, in the validation cohort, the model retained a high sensitivity of 0.94 (95 % CI: 0.84-0.98), underscoring its potential clinical applicability. GO analysis revealed promoter-proximal hypermethylation enriched for developmental programs, transcription factor activity, and chromatin remodeling, whereas gene body hypermethylation mapped to RNA metabolic processes and intracellular transport, supporting biological plausibility.
Conclusions
Integrating cfDNA methylation with machine learning delivers a minimally invasive, clinically actionable assay that robustly differentiates PCa from BPH. These results motivate prospective, multi-center validation and position cfDNA methylation profiling as a scalable strategy for prostate disease stratification.
Keywords: benign prostatic hyperplasia, biomarker, cfDNA methylation, machine learning, non-invasive diagnostics, prostate cancer
1. Introduction
Prostate cancer (PCa) remains the second most diagnosed malignancy and a leading cause of cancer-related deaths among men globally, presenting a substantial burden on healthcare systems worldwide.1 Although early detection significantly improves clinical outcomes, the differentiation between PCa and benign prostatic hyperplasia (BPH), a common non-malignant enlargement of the prostate, continues to pose a diagnostic dilemma.2 This challenge primarily stems from overlapping clinical features and the limited specificity of current diagnostic modalities such as serum prostate-specific antigen (PSA) testing and transrectal ultrasound-guided biopsy.3 PSA-based screening, while sensitive, is notorious for its high false-positive rate, often resulting in unnecessary invasive procedures and overtreatment.4,5 These limitations underscore the critical need for novel, precise, and minimally invasive diagnostic tools capable of distinguishing malignant from benign prostate conditions.
Epigenetic modifications, particularly DNA methylation, have emerged as promising biomarkers for cancer detection due to their early occurrence in tumorigenesis and relative stability in bodily fluids.6 Aberrant DNA methylation, including region-specific hypermethylation and hypomethylation, has been extensively reported in PCa, implicating key regulatory genes and pathways involved in carcinogenesis.7, 8, 9 Despite this, a systematic, comparative analysis of methylation patterns between PCa and BPH remains scarce, particularly when leveraging high-dimensional computational approaches capable of discerning subtle yet biologically significant epigenetic alterations.
In recent years, the integration of machine learning techniques with epigenomic data has opened new frontiers in precision diagnostics. Machine learning algorithms offer the capacity to detect complex, non-linear patterns within large-scale DNA methylation datasets, enabling the identification of distinct epigenetic signatures associated with disease phenotypes.10 Applying these advanced methodologies to the context of prostate disease holds promise for developing a robust, non-invasive diagnostic framework that enhances stratification and clinical decision-making. Herein, we hypothesize that specific CpG regions exhibit unique methylation signatures that are characteristic of either PCa or BPH. Using machine learning–based analysis of cell-free DNA (cfDNA) methylation profiles, this study aims to identify and validate differential methylation patterns as reliable diagnostic biomarkers. The findings are expected to support the development of a clinically applicable assay with enhanced specificity, reducing unnecessary prostate biopsies and improving patient management.
2. Materials and methods
2.1. Panel design/patient recruitment and sample collection
As described in our previous study,11 diagnosed 53 patients sample generated functional methylome (85Mbp methylation panel) sequencing data (BPH-11, PCa-42) were recruited in the present study. The methylation data fromTCGA datasets (https://portal.gdc. cancer.gov/) were analyzed by limma (R package) along with the in-house data to select differentially methylated CpG sites (Benjamini–Hochberg-corrected FDR <0.05). The methylation data from the GEO dataset with 656 normal white blood cell samples under the accession code GSE4027912 were used to remove hypermethylated CpG sites in the hematopoietic lineage (>0.1). CpG sites on the X and Y chromosomes were excluded, while those linked to common cancers were retained, resulting in 161,984 CpG sites covering ∼2.7 Mb of the genome across six cancer types (ovarian, lung, colorectal, pancreatic, liver, and esophageal). The panel was originally designed for early detection and tissue-of-origin classification of multiple tumor types.
A total of 365 patients were enrolled, including 230 with PCa and 135 with BPH. Plasma-derived cfDNA was collected from each participant, and all patients provided written informed consent for specimen use and genetic analysis.
2.2. DNA extraction
Whole blood was collected in anticoagulant-treated tubes, and plasma was separated by centrifugation at 1,600×g for 15 min, followed by a second spin to remove residual cells. The clarified plasma was transferred to clean tubes and stored at −80 °C. cfDNA was extracted from plasma using the Qiagen Circulating Nucleic Acid Kit (Qiagen, Hilden, Germany), quantified with a PicoGreen assay on a Qubit 4.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA, USA), and assessed for integrity using a 4200 TapeStation (Agilent Technologies, Santa Clara, CA, USA).
2.3. Sequencing data generation
We used cfDNA samples with ≥60 % intactness and 20 ng input to construct sequencing libraries using the NEBNext Enzymatic Methyl-seq Kit (New England Biolabs, Ipswich, MA, USA).13 Libraries were enriched with the LiquidSCAN_mDNA-P panel targeting 0.2 Mbp of prostate cancer–related regions and hybridized using the Twist NGS Methylation Detection System (Twist Biosciences, San Francisco, CA, USA) per the manufacturer's protocol. Libraries were diluted to 1.6 nM, pooled in equal volumes, denatured, and subjected to cluster generation. Paired-end sequencing (150 bp) was performed on a NovaSeq 6000 platform (Illumina, San Diego, CA, USA).
2.4. Sequencing data processing/quality control
Sequencing reads were trimmed using Trim Galore (v0.6.6, paired-end mode) with an additional 10 bp removed from the 5′ end to eliminate methylation bias and low-quality bases. Trimmed reads were aligned to the hg19 reference genome using Bismark (v0.22.3), followed by deduplication to reduce PCR bias. Methylation extraction was performed with Bismark's methylation extractor, generating CpG reports, from which only CpG sites within targeted panel regions were used for downstream analysis. Of the 365 sequencing datasets, samples with mean target coverage <300 × or usable bases on target <8 % were excluded, leaving 300 samples (PCa: 184; BPH: 116) for analysis. Among PCa cases, 50 were reserved for validation, and the remaining 250 were split into training and test sets with balanced group representation: training (PCa: 100, BPH: 57), test (PCa: 34, BPH: 19), and validation (PCa: 50, BPH: 40).
After quality control and cohort allocation, cfDNA library properties were assessed to confirm suitability for methylation analysis. Fragment length profiles displayed the expected cfDNA pattern, featuring a dominant mono-nucleosomal peak (∼150–200 bp) and a long-tail distribution representing sub- and di-/oligo-nucleosomal fragments (Fig. 1A). Across all retained samples (n = 300), the mean fragment length was 284.2 bp (median 227 bp; Fig. 1B). Over 90 % of fragments exceeded the 3 bp threshold, confirming high-quality library preparation without adapter dimer contamination. The observed nucleosome-associated fragmentation patterns further validated dataset integrity for downstream biomarker discovery.
Figure 1.

Distribution of CpG density and region length within the targeted methylation panel.
(A) Histogram of region lengths (bp) across the targeted panel. Most regions span fewer than 400 bp, with an average length of 284.2 bp and a median length of 227 bp (red and blue dashed lines, respectively). These distributions indicate that the custom panel was designed to balance sufficient CpG density with compact region sizes, optimizing sensitivity for cfDNA methylation profiling.
(B) Histogram showing the distribution of CpG sites per targeted region. The majority of regions harbor fewer than 30 CpGs, with a mean of 23.9 CpGs and a median of 17 CpGs per region (red and blue dashed lines, respectively).
CpG reports were processed using an in-house Python script to quantify cytosine conversion. For each locus, total counts were calculated as the sum of converted and unconverted cytosines, and sites with total counts <30 were excluded. Filtered CpG reports from all samples were then concatenated to identify loci present across the dataset.
2.5. Identification of region and data analysis
2.5.1. DNA methylation quantification
Raw methylation signals, represented by methylated (meth) and unmethylated (unmeth) read counts per CpG locus, were quantified as β-values (β), calculated as the proportion of methylated reads:
To enhance statistical robustness and stabilize variance across β-value distributions, a logit transformation was applied to compute M-values,14 incorporating a pseudo count (α = 0.05) to avoid by zero:
CpG-level β- and M-values were aggregated by clinical subgroups (Gleason score, PSA, age, stage, and diagnosis) through summing methylated and unmethylated counts before recalculation. Region-level β- and M-values were then computed by aggregating CpG sites within each genomic region. This hierarchical approach (CpG → region) minimized noise from low-coverage loci while preserving region-specific methylation patterns.
2.5.2. Identification of differentially methylated regions (DMRs)
DMRs between BPH and PCa were identified by comparing region-level M-values. Two-sided Mann–Whitney U tests were applied to account for the non-normal distribution of methylation values, yielding test statistics (U) and p-values for regions with valid data in both groups. Effect sizes were summarized as the difference in median M-values between PCa and BPH:
| ΔM = median (MPCa) − median (MBPH) |
Regions with p ≤ 0.05 were considered statistically significant. To ensure biological relevance, effect size thresholds were applied: regions with |ΔM| ≥ 0.2 were defined as candidate DMRs.
2.6. Reference materials and limit of detection (LoD) validation
Analytical sensitivity was assessed via an LoD study using the Seraseq® Methylated ctDNA Mutation Mix (SeraCare, Milford, MA, USA). Reference materials with defined methylation and mutation profiles were serially diluted to 0 %, 0.05 %, 0.1 %, 10 %, and 100 % variant allele fractions (VAFs). Each dilution was prepared in triplicate and quantified using the Qubit dsDNA HS Assay (Thermo Fisher Scientific, Waltham, MA, USA).
2.7. Machine learning-based classification
To evaluate the predictive value of methylation features, supervised machine learning models were built using region-level M-values and encoded clinical covariates (PSA, Age). Missing values were imputed with a sentinel (−4.5). Data were partitioned into validation (PCa: 50, BPH: 40) and training/test sets (PCa: 134, BPH: 76), with the latter split 75:25 into training (PCa: 100, BPH: 57) and test (PCa: 34, BPH: 19) cohorts. All features were standardized to zero mean and unit variance.
2.7.1. Correction for class imbalance
The Synthetic Minority Oversampling Technique (SMOTE) was applied to the training dataset to generate synthetic minority-class samples and address class imbalance.15
2.7.2. Model development and optimization
Two supervised classifiers were trained: Random Forest (RF) and Extreme Gradient Boosting (XGBoost).
-
•
RF hyperparameters (number of trees, maximum depth, minimum samples per split/leaf, and feature selection strategy) were tuned via grid search.16
-
•
XGBoost hyperparameters (earning rate, estimators, depth, subsampling, column sampling, gamma, L1/L2 regularization, and minimum child weight) were optimized using randomized search.17
Both optimization procedures employed k-fold cross-validation, with the area under the receiver operating characteristic curve (ROC-AUC) used as the primary performance metric.
2.7.3. Ensemble learning
Optimized RF and XGBoost models were integrated using a stacking ensemble approach, employing logistic regression as the meta-learner with ten-fold cross-validation and passthrough of original features.18
2.8. Software and packages
All analyses were performed in Python (version 3.10) using the following packages:
3. Results
3.1. cfDNA methylation signatures distinguishing PCa from BPH
To identify diagnostic methylation markers for PCa, a targeted panel of 754 DMRs was developed from 1,251 tumor-specific DMRs identified in tissue samples. A subset of 643 genes was selected to create a cfDNA-optimized plasma panel. Validation via LoD testing using serial dilutions (0 %, 0.05 %, 0.1 %, 10 %, 100 %) demonstrated a detection sensitivity of 93.98 % at 0.1 % methylation, confirming its ultra-sensitive cfDNA performance (Table S1 and Fig. S1).
A total of 365 plasma-derived cfDNA samples (PCa: 230; BPH: 135) were analyzed using enzymatic methyl-sequencing with the Twist Targeted Methylation Workflow. PCa cases were stratified by Gleason score (6–9) into groups of 57, 120, 29, and 24 samples, respectively (Table 1). After quality control filtering (coverage >300×, usable bases >8 %), 300 samples were retained for analysis (PCa: 184; BPH: 116).
Table 1.
Baseline Characteristics of patients included in the study
Clinical characteristics of benign prostatic hyperplasia (BPH) and prostate cancer (PCa) patients stratified by Gleason score. The table summarizes the number of samples, serum prostate-specific antigen (PSA) levels (ng/mL, mean ± SD), and age (years, mean ± SD).
| Diagnosis | Gleason Score | No. of Samples | PSA level (ng/ml) Mean ± SD | Age Mean ± SD | ||
|---|---|---|---|---|---|---|
| BPH | BPH | 135 | 7.74 | 7.42 | 71.46 | 7.39 |
| PCa | 6 | 57 | 6.66 | 6.89 | 68.35 | 5.02 |
| PCa | 7 | 120 | 11.23 | 8.09 | 70.18 | 6.32 |
| PCa | 8 | 29 | 7.84 | 7.04 | 71.14 | 6.53 |
| PCa | 9 | 24 | 9.35 | 7.50 | 72.50 | 7.23 |
Abbreviations: BPH, benign prostatic hyperplasia; PCa, prostate cancer; PSA, prostate-specific antigen; SD, standard deviation.
Sequencing produced comparable read depths between groups (mean: 3.88 × 107 for PCa; 2.94 × 107 for BPH). Alignment to the hg19 reference genome yielded mean alignment rates of 16.1 % (PCa) and 18.0 % (BPH), with average deduplicated coverage exceeding 600 × per sample, ensuring robust methylation quantification. Enzymatic conversion efficiency, estimated from non-CpG (CHH and CHG) contexts, was consistently high (median: 99.2 %; IQR: 98.8–99.5 %; 95th percentile: 99.7 %), and all libraries exceeded 91 % efficiency, confirming the reliability of the conversion process (Table S2).
Across 300 samples, methylation signals were detected at 10,275 CpG sites corresponding to 748 predefined DMRs (Fig. 2A). These loci showed high coverage uniformity and reproducibility across samples, confirming the panel's suitability for liquid biopsy applications. Comparative analysis between PCa and BPH identified 466 significant CpGs spanning 43 DMRs, which constituted the core diagnostic feature set (Fig. 2B).
Figure 2.

Differential methylation profiles distinguishing prostate cancer (PCa) from benign prostatic hyperplasia (BPH). (A) Heatmap of methylation profiles across 10,275 CpGs within 748 predefined regions. Each row represents a CpG site and each column corresponds to an individual sample (PCa, n = 184; BPH, n = 116). Samples are annotated with diagnostic group (PCavs. BPH), clinical stage, and Gleason score. Distinct clustering patterns reveal global hypermethylation in PCa compared with BPH. (B) Heatmap of methylation profiles across 466 CpGs within 43 regions selected based on significant differential methylation (ΔM-value ≥0.2). The refined panel demonstrates a sharper segregation between PCa and BPH cohorts, with tumor samples exhibiting higher methylation levels (blue) relative to benign controls (green). Clinical annotations further confirm the association between methylation signatures and pathological features.
3.2. Gene ontology analysis
To investigate the functional significance of cfDNA-derived methylation changes, Gene Ontology (GO) enrichment analysis was conducted on genes linked to hypermethylated DMRs. DMRs were categorized by genomic context—upstream regulatory regions (≤5 kb from transcription start sites), gene bodies (first exon to 3′ UTR), and full gene spans—to enable context-specific functional interpretation.
GO enrichment within the biological process category revealed a clear dichotomy: upstream hypermethylation was enriched for transcriptional repression, embryonic development, and cell fate commitment—hallmarks of promoter-associated silencing.22 In contrast, gene body hypermethylation was linked to RNA processing, DNA repair, and metabolic homeostasis, reflecting roles in transcriptional fine-tuning and post-initiation regulation (Fig. 3A).23
Figure 3.

Gene Ontology (GO) enrichment analysis of differentially methylated genes distinguishing prostate cancer (PCa) from benign prostatic hyperplasia (BPH).
(A)Enriched biological process categories. Significant enrichment was observed in processes related to chromatin organization, regulation of gene expression, and morphogenesis, supporting the functional impact of aberrant methylation on cancer-associated pathways.
(B)Enriched cellular component categories. Genes were predominantly associated with the nuclear chromatin, nucleosome, and transcription regulator complexes, highlighting the central role of nuclear architecture and chromatin remodeling in prostate cancer epigenetic reprogramming.
(C)Enriched molecular function categories. The differentially methylated genes were significantly associated with transcription factor activity, DNA-binding, and RNA polymerase II regulatory functions, indicating potential epigenetic regulation of transcriptional machinery.
Cellular component analysis supported this distinction, with upstream hypermethylated genes localized to nuclear compartments (chromatin, transcription factor complexes, nucleoplasm), while gene body–methylated genes were enriched in cytoplasmic structures (vesicle membranes, Golgi apparatus, cytoskeleton), implicating functions in intracellular transport and structural organization (Fig. 3B).
Under the molecular function category, upstream methylation was associated with DNA binding, chromatin remodeling, and transcription factor activity—consistent with epigenetic repression—whereas gene body hypermethylation was enriched for catalytic activity, RNA binding, and enzyme regulation, indicating modulation of gene expression rather than complete silencing (Fig. 3C).24
Together, these findings underscore the context-dependent functions of DNA methylation: upstream hypermethylation serves as a transcriptional gatekeeper, while gene body methylation contributes to transcript stability, alternative splicing, and metabolic adaptation—processes likely pivotal in prostate tumorigenesis and progression.25
3.3. Identification of methylation marker between BPH and PCa
Differentially methylated regions (DMRs) were identified using region-level M-values derived from cfDNA. Statistical comparisons between PCa and BPH groups were performed using two-sided Mann-Whitney U tests. Regions with p-values ≤0.05 and absolute median M-value differences (|ΔM|) ≥ 0.2 were considered statistically and biologically significant. To ensure regional robustness and reduce noise from isolated CpG effects, only regions containing at least four significant CpG sites were retained. This procedure yielded 43 high-confidence DMRs, which were fixed prior to model training. Importantly, machine learning algorithms were not used for marker discovery; instead, they were applied exclusively in the classification stage to integrate the selected methylation features and estimate predictive probabilities.
Before marker selection, genomic distribution analysis showed that 63.9 % of DMRs were in gene body regions and 36.1 % in upstream regulatory elements (Fig. 4A). Among these, 43 DMRs were hypermethylated in PCa relative to BPH, indicating tumor-specific methylation gains. Overall, 30.2 % of hypermethylated DMRs mapped to upstream regions, including promoters (27.9 %) and 5′ UTRs (2.3 %), while 69.8 % were within gene bodies, notably in first introns (18.6 %), 3′ UTRs (7.0 %), and coding exons (7.0 %) (Fig. 4B). This distribution aligns with region-specific methylation functions proposed in current epigenetic models.
Figure 4.

Genomic feature distribution of CpG sites included in the targeted methylation panel.(A) Overall distribution of all CpG sites (n = 10,275) across genomic features. The majority of CpGs were located within gene bodies (63.9 %), followed by upstream regulatory regions (36.1 %), with smaller fractions mapping to introns, untranslated regions (UTRs), and exons.(B) Distribution of CpGs retained after feature selection (n = 466 sites across 43 regions). Compared with the total panel, the selected CpGs were enriched in gene bodies (69.8 %) and first introns (18.6 %), whereas the relative proportion of upstream regulatory sites was reduced (30.2 %). This redistribution suggests that discriminatory methylation signals between prostate cancer and BPH preferentially reside within intragenic regions rather than promoters.
Serum PSA levels were dichotomized into normal or high based on age-specific thresholds (40–49: ≤2.5 ng/mL; 50–59: ≤3.5; 60–69: ≤4.5; 70–79: ≤6.5; ≥80: ≤7.2). Using this age-adjusted classification, distinct groupwise differences in PSA status were observed, suggesting a potential association between PSA levels and methylation-based indicators (Fig. 5A). A similar pattern was noted under stage-adjusted analysis (Fig. 5B).
Figure 5.

Distribution of age-and stage-adjusted PSA levels in prostate cancer patients.
Serum PSA values were dichotomized into normal versus high using a priori age-specific thresholds (40–49: ≤2.5 ng/mL; 50–59: ≤3.5; 60–69: ≤4.5; 70–79: ≤6.5; ≥80: ≤7.2) (A) Proportion of individuals with high versus normal prostate-specific antigen (PSA) levels stratified by age category. The prevalence of elevated PSA was markedly higher in younger patients (50s) and older patients (90s), whereas intermediate age groups (60s–80s) showed a more balanced distribution, with a notable shift toward normal PSA levels in the 70s.
(B) Distribution of high versus normal PSA levels across clinical stage categories. Clinical stages were defined according to the AJCC TNM system (8th edition): stage 0–I (organ-confined disease without extraprostatic extension), stage II (localized tumors with higher Gleason grade or PSA levels), stage III (locally advanced disease with extracapsular extension or seminal vesicle invasion), and stage IV (regionally invasive or metastatic disease). Elevated PSA was most frequently observed in stage 2 and stage 3 patients, whereas early (stage 0–1) and late (stage 4) stages demonstrated a more balanced distribution. Error bars represent the standard error of the mean (SEM). These findings highlight distinct age-and stage-specific patterns of PSA elevation, supporting the potential clinical utility of stratified PSA assessment in prostate cancer risk evaluation.
Comparative analysis of cfDNA methylation profiles between PCa (n = 184) and BPH (n = 116) identified a subset of differentially methylated genes (DMGs) with significant hypermethylation in PCa. Using thresholds of |ΔM| ≥ 0.2 and p < 0.05, several loci showed strong effect sizes. Notably, GATA2, COL9A3, WNT6, STK10, NT5DC2, RASSF10, and MAP7 displayed the largest ΔM values, indicating marked methylation differences between diagnostic groups (Fig. 6A). Selected DMGs with |ΔM| ≥ 0.2 were further visualized as volcano and bar plots (Fig. 6B and C), and CpG-wise methylation distributions were illustrated as box plots by group (Fig. S2).
Figure 6.

Differential methylation analysis between prostate cancer (PCa) and benign prostatic hyperplasia (BPH).(A) Boxplots of M-values for representative differentially methylated genes (DMGs) across PCa(n = 184) and BPH (n = 116) cohorts. Several loci, including GATA2, COL9A3, WNT6, STK10, NT5DC2, RASSF10, and MAP7, exhibited marked hypermethylation in PCa relative to BPH.(B) Volcano plot illustrating differential methylation at the regional level. Regions with ΔM-value ≥0.2 and p < 0.05 (red dots) were considered significantly differentially methylated, highlighting enrichment of hypermethylated loci in PCa.(C) Ranked bar plot of top candidate genes based on absolute ΔM-values. Genes surpassing the ΔM-value cutoff of 0.2 (dashed red line) were prioritized as robust biomarkers, with GATA2 showing the strongest differential signal.
3.4. Classification via machine learning
To assess the diagnostic potential of cfDNA methylation biomarkers, a supervised machine learning framework was developed using 43 DMRs and serum PSA as input features. After stringent quality control (coverage, base quality, on-target rate), 300 cfDNA samples were retained (PCa: 184; BPH: 116). The dataset was divided into a training/test cohort (n = 210; PCa: 134, BPH: 76) and an independent validation cohort (n = 90; PCa: 50, BPH: 40) using stratified random sampling to preserve class balance.
Feature scaling was applied via z-score normalization. Two ensemble-capable algorithms—RF and XGBoost—were trained with hyperparameter tuning via randomized search and 10-fold cross-validation. Sensitivity (recall) was prioritized as the objective to minimize false negatives, a key clinical requirement for early cancer detection.
Model performance was evaluated on both test and validation sets. The final stacking ensemble, combining probabilistic outputs from RF and XGBoost via a logistic regression meta-learner, achieved ROC-AUCs of 0.99 and 0.98 in the test and validation sets, respectively. The model showed excellent accuracy (test: 0.95; validation: 0.93) with balanced sensitivity (test: 0.96; validation: 0.94) and specificity (test: 0.93; validation: 0.93). These results highlight strong robustness and generalizability of the cfDNA-based classifier for distinguishing PCa from BPH.
To contextualize these results relative to conventional clinical assessment, we additionally evaluated a classifier based on serum PSA alone. The PSA classifier demonstrated limited discriminative performance, yielding a sensitivity of 0.59, specificity of 0.67, and overall accuracy of 0.62 (Fig. S3). Notably, this model resulted in a substantial number of false-negative prostate cancer cases, underscoring the limitations of PSA-based stratification when used in isolation.
In contrast, the cfDNA methylation-based model achieved markedly superior performance, with sensitivity of 0.96 and specificity of 0.93 in the test cohort, corresponding to a pronounced reduction in false-negative classifications. These findings indicate that cfDNA methylation features are the primary drivers of improved classification performance and provide substantial diagnostic information beyond serum PSA alone.
In the independent validation cohort, the optimized cfDNA-based model achieved a sensitivity of 0.94 (95 % CI: 0.84-0.98) and specificity of 0.93 (95 % CI: 0.80-0.97) (Fig. 7), supporting its suitability for screening-oriented applications. Bootstrap resampling (1,000 iterations) was used to derive empirical 95 % confidence intervals for AUC, sensitivity, and specificity, providing a non-parametric assessment of model stability and reproducibility.
Figure 7.

Receiver operating characteristic (ROC) analysis of the stacking ensemble model. ROC curves illustrate the diagnostic performance of the stacking model integrating Random Forest and XGBoost classifiers. In the independent test cohort, the model achieved near-perfect discrimination (AUC = 0.99), with a sensitivity of 0.96 and a specificity of 0.93 at the optimal threshold. In the validation cohort, the model maintained robust performance (AUC = 0.98), with sensitivity of 0.94 and specificity of 0.93. The diagonal dashed line indicates random classification. These results highlight the high generalizability and clinical potential of the ensemble methylation-based classifier for distinguishing prostate cancer from benign prostatic hyperplasia.
Given the known diagnostic ambiguity associated with intermediate PSA levels, we further performed a subgroup analysis within the PSA gray zone. Notably, the cfDNA methylation-based classifier achieved complete separation between PCa and BPH cases in this subgroup (Fig. S4), with no observed false-positive or false-negative classifications. Although this analysis was conducted on a limited sample size, the result highlights the strong discriminative potential of cfDNA methylation profiling in clinically challenging scenarios where PSA-based stratification is insufficient.
Collectively, these results support the translational feasibility of a minimally invasive, cfDNA-based diagnostic assay capable of discriminating PCa from BPH with high accuracy, sensitivity, and clinical applicability.26
4. Discussion
This study presents the development and validation of a cfDNA methylation–based diagnostic framework capable of differentiating PCa from BPH. Using a targeted panel of 43 high-confidence DMRs, our ensemble machine learning model achieved excellent discriminative performance, with an AUC of 0.98 and a sensitivity of 0.94 at a clinically relevant specificity of 93 %. Importantly, comparative analysis demonstrated that the cfDNA methylation-based model substantially outperformed a PSA-only classifier, particularly by reducing false-negative prostate cancer classifications. Together, these results demonstrate the clinical feasibility of cfDNA methylation profiling as a minimally invasive, blood-based strategy for prostate disease stratification.
A major strength of this work lies in translating tissue-derived epigenomic findings into a cfDNA-compatible diagnostic assay optimized for plasma-based detection. Notably, the majority of cfDNA-derived hypermethylated regions did not overlap with previously reported tumor tissue markers, revealing a distinct epigenetic landscape in circulating tumor DNA. This discordance likely reflects the unique biological context of cfDNA release and clearance, which is influenced by apoptosis, immune interactions, and DNA fragmentation biases. Collectively, these findings indicate that cfDNA methylation signatures provide complementary, rather than redundant, molecular information to tissue-based profiles and help explain their added diagnostic value beyond conventional clinical markers such as PSA.
Functional annotation of the identified hypermethylated regions revealed enrichment in pathways related to transcriptional regulation, chromatin remodeling, and morphogenesis, underscoring their biological relevance.27 Notably, promoter-proximal hypermethylation was enriched in developmental and transcription factor–related genes, consistent with canonical tumor suppressor silencing, whereas intragenic hypermethylation was associated with RNA metabolic and catalytic processes. This pattern suggests that methylation exerts region-specific regulatory effects, consistent with emerging models of spatially defined epigenetic control in cancer.28
When compared with our previously published tissue-derived methylation dataset, the 43 cfDNA-derived hypermethylated genes were largely non-overlapping, including the absence of previously reported genes such as RAD51, CD9, FZD7, STARD13, and YTHDF1.29, 30, 31 This finding further supports the independent biological origin and complementary diagnostic potential of cfDNA methylation profiles. Furthermore, while several known methylation biomarkers—including SOX7, DHRS4L2, TACC2, FLNC, AOX1, HAPLN3, SLC18A2, and CTBP2—were detected, they did not rank among the top discriminative features. This observation suggests that our cfDNA-based approach captures novel, clinically informative methylation signals not emphasized in previous tissue-focused studies.32
Importantly, the integrated cfDNA model addresses one of the most significant challenges in PCa diagnostics, the limited specificity of PSA testing, which often leads to unnecessary biopsies and overdiagnosis. By incorporating cfDNA methylation as an orthogonal biomarker, our model substantially enhanced diagnostic accuracy while maintaining high specificity. This advantage was particularly evident in patients within the PSA gray zone, a clinically challenging subgroup in which PSA-based stratification is least informative. In this subgroup, cfDNA methylation profiling achieved complete separation between PCa and BPH cases, highlighting its potential utility in resolving diagnostic ambiguity. Nevertheless, we acknowledge that this subgroup analysis was conducted on a limited sample size and should be interpreted cautiously until validated in larger prospective cohorts.
Several limitations should be acknowledged. First, cross-sectional design restricts longitudinal inference; prospective studies are required to evaluate diagnostic performance for early detection, disease monitoring, and treatment response. Second, although the targeted panel was optimized for cfDNA, it may not fully capture intratumoral heterogeneity or stromal contributions. Third, while this study included samples from multiple institutions, further validation in multi-ethnic populations will be essential to confirm the generalizability and clinical translatability of the model.
Future research should aim to integrate cfDNA methylation with orthogonal data modalities such as fragmentomics, transcriptomics, and proteomics to enhance predictive accuracy and enable mechanistic interpretation. Longitudinal cfDNA monitoring may further improve biomarker specificity for clinically actionable endpoints such as treatment selection or relapse detection. Ultimately, large-scale, regulatory-grade prospective validation will be critical for clinical implementation of cfDNA-based diagnostic assays.
In conclusion, this study introduces a robust and clinically translatable cfDNA methylation model that accurately discriminates PCa from BPH. Beyond diagnostic performance, our findings highlight the distinct biological features of cfDNA methylation and establish a scalable framework for next-generation, non-invasive prostate cancer diagnostics.
Ethical approval
This study was approved by the institutional review board of Seoul St. Mary's Hospital (IRB No. KC20SNSI0988) Asan Medical Center (IRB No. S2021-0500-0002), and Ewha Womans University Mokdong Hospital (IRB No. EUMC 2023-10-038).
Conflicts of interests
The authors declare no competing interests.
Acknowledgment
This work was supported by a Korea Medical Device Development Fund grant funded by the Korean government (the Ministry of Science and ICT, the Ministry of Trade, Industry, and Energy, the Ministry of Health and Welfare, and the Ministry of Food and Drug Safety, RS-2020-KD000018).
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.prnil.2026.02.003.
Appendix A. Supplementary data
The following are the Supplementary data to this article:
References
- 1.Chu F., Chen L., Guan Q., Chen Z., Ji Q., Ma Y., et al. Global burden of prostate cancer: Age-period-cohort analysis from 1990 to 2021 and projections until 2040. World J Surg Oncol. 2025:23. doi: 10.1186/s12957-025-03733-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.McNally C.J., Ruddock M.W., Moore T., McKenna D.J. Biomarkers that differentiate benign prostatic hyperplasia from prostate cancer: a literature review. Cancer Manag Res. 2020;12:5225–5241. doi: 10.2147/CMAR.S250829. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Logozzi M., Angelini D.F., Giuliani A., Mizzoni D., Di Raimo R., Maggi M., et al. Increased plasmatic levels of PSA-expressing exosomes distinguish prostate cancer patients from benign prostatic hyperplasia: a prospective study. Cancers (Basel) 2019;11:1449. doi: 10.3390/cancers11101449. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Lumbreras B., Parker L.A., Caballero-Romeu J.P., Gómez-Pérez L., Puig-García M., López-Garrigós, et al. Variables associated with false-positive PSA results: a cohort study with real-world data. Cancers (Basel) 2022;15:261. doi: 10.3390/cancers15010261. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Bernal-Soriano M.C., Parker L.A., López-Garrigos M., Hernández-Aguado I., Caballero-Romeu J.P., Gómez-Pérez L., et al. Factors associated with false negative and false positive results of prostate-specific antigen (PSA) and the impact on patient health. Medicine (Baltim) 2019;98 doi: 10.1097/MD.0000000000017451. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Constâncio V., Nunes S.P., Henrique R., Jerónimo C. DNA methylation-based testing in liquid biopsies as detection and prognostic biomarkers for the four major cancer types. Cells. 2020;9:624. doi: 10.3390/cells9030624. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Muletier R., Bourgne C., Guy L., Dougé A. DNA methylation in prostate cancer: clinical implications and potential applications. Cancer Med. 2025;14 doi: 10.1002/cam4.70528. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Chen S., Petricca J., Ye W., Guan J., Zeng Y., Cheng N., et al. The cell-free DNA methylome captures distinctions between localized and metastatic prostate tumors. Nat Commun. 2022;13 doi: 10.1038/s41467-022-34012-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Shin H.J., Hua J.T., Li H. Recent advances in understanding DNA methylation of prostate cancer. Front Oncol. 2023;13:1182727. doi: 10.3389/fonc.2023.1182727. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Panagopoulou M., Karaglani M., Manolopoulos V.G., Iliopoulos I., Tsamardinos I., Chatzaki E. Deciphering the methylation landscape in breast cancer: diagnostic and prognostic biosignatures through automated machine learning. Cancers (Basel) 2021;13:1677. doi: 10.3390/cancers13071677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kim S.S., Lee S.C., Lim B., Shin S.H., Kim M.Y., Kim S.Y., et al. DNA methylation biomarkers distinguishing early-stage prostate cancer from benign prostatic hyperplasia. Prostate Int. 2023;11:113–121. doi: 10.1016/j.prnil.2023.01.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Loyfer N., Magenheim J., Peretz A., Cann G., Bredno J., Klochendler A., et al. A DNA methylation atlas of normal human cell types. Nature. 2023;613:355–364. doi: 10.1038/s41586-022-05580-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Vaisvila R., Ponnaluri V.K.C., Sun Z., Langhorst B.W., Saleh L., Guan S., et al. Enzymatic methyl sequencing detects DNA methylation at single-base resolution from picograms of DNA. Genome Res. 2021;31:1280–1289. doi: 10.1101/gr.266551.120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Du P., Zhang X., Huang C.C., Jafari N., Kibbe W.A., Hou L., et al. Comparison of beta-value and M-value methods for quantifying methylation levels by microarray analysis. BMC Bioinf. 2010;11:587. doi: 10.1186/1471-2105-11-587. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Fernández A., García S., Herrera F., Chawla N.V. SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary. J Artif Intell Res. 2018;61:863–905. doi: 10.1613/JAIR.1.11192. [DOI] [Google Scholar]
- 16.Probst P., Wright M.N., Boulesteix A.L. Hyperparameters and tuning strategies for random forest. WIREs Data Min Knowl Discov. 2018;9 doi: 10.1002/widm.1301. [DOI] [Google Scholar]
- 17.Chen T., Guestrin C. XGBoost: a scalable tree boosting system. Proc ACM SIGKDD Int Conf Knowl Discov Data Min. 2016;22:785–794. doi: 10.1145/2939672.2939785. [DOI] [Google Scholar]
- 18.Bai J., Wang S., Xu Q., Zhu J., Li Z., Lai K., et al. Intelligent regional subsurface prediction based on limited borehole data and interpretability stacking technique of ensemble learning. Bull Eng Geol Environ. 2024;83 doi: 10.1007/s10064-024-03758-y. [DOI] [Google Scholar]
- 19.Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–2830. [Google Scholar]
- 20.Chen T., Guestrin C. Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 2016. XGBoost: a scalable tree boosting system; pp. 785–794. [Google Scholar]
- 21.Lemaître G., Nogueira F., Aridas C.K. Imbalanced-learn: a Python toolbox to tackle the curse of imbalanced datasets in machine learning. J Mach Learn Res. 2017;18:1–5. [Google Scholar]
- 22.Ehrlich M., Lacey M. DNA hypermethylation in disease: mechanisms and clinical relevance. Epigenetics. 2019;14:1141–1163. doi: 10.1080/15592294.2019.1638701. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Wang Q., Xiong F., Wu G., Liu W., Chen J., Wang B., et al. Gene body methylation in cancer: molecular mechanisms and clinical applications. Clin Epigenet. 2022;14 doi: 10.1186/s13148-022-01382-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Héberlé É., Bardet A.F. Sensitivity of transcription factors to DNA methylation. Essays Biochem. 2019;63:727–741. doi: 10.1042/EBC20190033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Long M.D., Smiraglia D.J., Campbell M.J. The genomic impact of DNA CpG methylation on gene expression; relationships in prostate cancer. Biomolecules. 2017;7:15. doi: 10.3390/biom7010015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Bryzgunova O., Bondar A., Ruzankin P., Laktionov P., Tarasenko A., Kurilshikov A., et al. Locus-specific methylation of GSTP1, RNF219, and KIAA1539 genes with single molecule resolution in cell-free DNA from healthy donors and prostate tumor patients: application in diagnostics. Cancers (Basel) 2021;13:6234. doi: 10.3390/cancers13246234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Ando M., Saito Y., Xu G., Bui NQ., Medetgul-Ernar K., Pu M., et al. Chromatin dysregulation and DNA methylation at transcription start sites associated with transcriptional repression in cancers. Nat Commun. 2019;10:2188. doi: 10.1038/s41467-019-09937-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Pidsley R., Lam D., Qu W., Peters T.J., Luu P.L., Korbie, et al. Comprehensive methylome sequencing reveals prognostic epigenetic biomarkers for prostate cancer mortality. Clin Transl Med. 2022;12:e1030. doi: 10.1002/ctm2.1030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Hu X., Xuan H., Du H., Jiang H., Huang J. Down-regulation of CD9 by methylation decreased bortezomib sensitivity in multiple myeloma. PLoS One. 2014;9 doi: 10.1371/journal.pone.0095765. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Yang A., Yan S., Yin Y., Chen C., Tang X., Ran M., et al. FZD7, regulated by non-CpG methylation, plays an important role in immature porcine Sertoli cell proliferation. Int J Mol Sci. 2023;24:6179. doi: 10.3390/ijms24076179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zeng C., Li H., Liang W., Chen J., Zhang Y., Zhang H., et al. Loss of STARD13 contributes to aggressive phenotype transformation and poor prognosis in papillary thyroid carcinoma. Endocrine. 2024;83:127–141. doi: 10.1007/s12020-023-03468-7. [DOI] [PubMed] [Google Scholar]
- 32.Oloomi M., Moazzezy N., Bouzari S. Comparing blood versus tissue-based biomarker expression in breast cancer patients. Heliyon. 2020;6 doi: 10.1016/j.heliyon.2020.e03728. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
