Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 Jun 22:2026.06.17.732357. [Version 1] doi: 10.64898/2026.06.17.732357

nanoASM: Long-Read Allele-Specific DNA Methylation Profiling Enables Functional Annotation of Regulatory Noncoding Variants in Human Prostate Tissues

Yijun Tian 1, Jodie Wong 1, Shannon K McDonnell 2, Hua Zhong 3,4, Lang Wu 3,4, Nicholas B Larson 2, Brandon J Manley 5, Liang Wang 1,*
PMCID: PMC13320949  PMID: 42395488

Summary

Long-read nanopore sequencing enables simultaneous detection of germline variation and native DNA base modifications on individual DNA molecules, providing a unique opportunity to investigate allele-specific epigenetic regulation. Here, we performed whole-genome nanopore sequencing on normal and tumor prostate tissues to characterize differential methylation, methylation entropy, and allele-specific methylation (ASM) associated with noncoding genetic variants. Genome-wide analysis identified extensive cancer-associated differentially methylated regions (DMRs), with hypermethylated DMRs significantly enriched near transcription start sites and transcriptional regulatory regions. Integration with transcriptomic datasets revealed strong inverse relationships between promoter methylation and gene expression, while 5-hydroxymethylcytosine (5hmC) levels positively correlated with transcriptional activity across gene bodies. Using fragment-level methylation patterns enabled by long-read sequencing, we further quantified methylation entropy incorporating both 5mCG and 5hmCG states. Cancer-hypermethylated DMRs exhibited markedly reduced entropy, consistent with clonal fixation of methylation states during tumor progression. Entropy profiling across chromatin annotations demonstrated maximal epigenetic heterogeneity at partially modified enhancer-associated regions. To investigate cis-regulatory genetic effects, we developed a simple ASM framework (nanoASM) that can partition sequencing reads by allelic state and identifies allele-specific DMRs directly from long-read data. Compared with conventional population-level mQTL analysis, ASM demonstrated substantially improved statistical efficiency by leveraging within-individual contrasts and reducing sample-level heterogeneity. Although germline single nucleotide polymorphisms (SNPs) were largely shared between normal and tumor tissues, ASM patterns differed substantially, with tumor-associated ASM regions displaying significantly larger genomic span and stronger allelic methylation differences. Comparative analysis with TCGA prostate mQTL and GTEx prostate eQTL datasets demonstrated substantial concordance between ASM directionality and downstream transcriptional effects, particularly for variants located within DMRs and near transcription start sites. At the IRX4 prostate cancer risk locus, ASM identified an androgen-responsive regulatory domain overlapping AR ChIP-seq and H3K27ac peaks, nominating rs6885084 as a candidate functional variant. At the PSCA locus, ASM anchored by rs4736369 was associated with allele-specific methylation, chromatin activation, transcript abundance, and isoform usage. Together, these findings establish nanopore-based ASM analysis as a powerful approach for resolving functional noncoding variants and their regulatory domains they control in prostate cancer.

Introduction

The completion of the Human Genome Project (HGP) marked a major milestone in biomedical research, providing a nearly complete reference sequence of the human genome.13 Although the vast majority of the genome is identical among individuals, phenotypic diversity is largely driven by single nucleotide polymorphisms (SNPs).4,5 Deciphering the functional impact of these genetic variants, however, remains a formidable challenge.68 Genome-wide association studies (GWAS) have identified thousands of SNPs associated with complex traits and diseases, yet most reside in noncoding regions where their biological interpretation remains elusive.9,10 To link genetic associations with molecular function, expression quantitative trait locus (eQTL) mapping has been widely applied to connect germline variants with gene expression,11,12 revealing that SNPs can modulate transcription by altering enhancer activity,13 promoter accessibility,14 or chromatin organization.15 Nevertheless, transcript-level associations alone provide only a partial view of regulatory complexity, as epigenetic modifications — particularly DNA methylation — play a pivotal role in shaping transcriptional outcomes.

Methylation quantitative trait locus (mQTL) analysis extends these efforts by identifying variants that influence DNA methylation patterns, revealing how sequence differences can alter local chromatin states and gene regulation.1618 Despite its strong functional implications, current mQTL approaches are constrained by several methodological limitations. Although mQTL studies have identified numerous statistically significant associations, the majority rely on microarray-based methylation data in which the signal corresponds to isolated single-CpG sites rather than broader functional regions.1921 Combined with the effects of linkage disequilibrium (LD), this produces a large number of associations distributed broadly across the genome, many of which offer limited power to prioritize functionally relevant variants. 18

One strategy to more effectively pinpoint functional noncoding variants is to evaluate their regulatory potential by examining the local methylation pattern flanking each allele.22 Such allele-specific information can indicate whether a variant perturbs proximal epigenetic states and thereby influences regulatory activity — but this requires the simultaneous quantification of base modifications and variant alleles within the same DNA molecules. Bisulfite-based methods combined with microarray or next-generation sequencing (NGS) technologies have enabled genome-wide methylation profiling, yet they carry inherent limitations. Bisulfite conversion destroys allelic information at cytosine-containing SNPs, complicating variant detection. Short-read sequencing further restricts the analysis of co-methylation patterns among distal CpG sites, and PCR amplification can distort native methylation ratios. Critically, bisulfite-based approaches cannot distinguish 5-methylcytosine (5mC) from its oxidized derivative 5-hydroxymethylcytosine (5hmC)23, which serves as an intermediate in active demethylation24,25 and plays independent regulatory roles in transcription and development.26

To overcome these limitations, we employed nanopore long-read sequencing, which directly detects base modifications and sequence variation on the same DNA molecules without bisulfite conversion. We aggregated read-level methylation by the allele carried on each read, obtaining one methylation estimate per allele per individual (reference vs. alternative), and compared allele-specific methylation (ASM) differences across individuals. Conceptually, this framework operates at the allele level rather than on bulk sample averages: heterozygous individuals contribute two allele-specific observations under a shared genetic background, while homozygous individuals contribute genotype-dosage information. This allele-resolved design reduces the influence of sample-level heterogeneity relative to conventional population-level mQTL analysis, thereby increasing effective power to detect cis effects — particularly for common variants.

Methods

Cohort information and case selection

As demonstrated in our previous study, a germline risk variant can induce allele-specific epigenetic differences that are substantially altered during tumorigenesis.14 This observation led us to hypothesize that inherited genetic variation may contribute to cancer-associated epigenetic remodeling and that such effects can be systematically identified through allele-specific methylation analysis. Therefore, in addition to normal prostate tissues, we included prostate cancer samples to investigate how genetic regulation of DNA methylation is maintained, altered, or amplified in the cancer state. Prostate genomic DNA and tissue samples used for nanopore sequencing were collected from two institutions: Mayo Clinic (dbGaP accession phs000985.v2.p1, “Functional Significance of Prostate Cancer Risk-SNPs”) and Moffitt Cancer Center (“Total Cancer Care Biobanking”). The informed-consent process was completed with all participants, and the study was approved by the Institutional Review Boards of both institutions.

For the Mayo Clinic cohort, archived genomic DNA (N = 30) was used directly for nanopore sequencing library preparation. Prostate tissue was retrieved from archived fresh-frozen collections obtained from patients who underwent radical prostatectomy or cystoprostatectomy. For each normal prostate tissue sample, H&E-stained slides were reviewed and required to meet the following criteria: (1) absence of cancer cells; (2) absence of benign prostatic hyperplasia; (3) high representation of normal prostatic epithelial glands (≥40% of all cells); and (4) low fraction of lymphocytic infiltrate (≤20% of all cells).

For the Moffitt Cancer Center cohort, primary prostate tumor (N = 15) and paired adjacent normal tissue (N = 15) were obtained from patients with prostate cancer receiving surgical treatment under protocols approved by the Moffitt Cancer Center Institutional Review Board (Total Cancer Care protocol [TCC], MCC# 50468; Advarra IRB Pro00000971).

To ensure a uniform genetic background, all patients included in this study are self-reported as Caucasian male.

Tissue processing and genomic DNA extraction

For the Mayo Clinic cohort, tissue preparation and genomic DNA extraction procedures have been described previously 27. For the Moffitt Cancer Center cohort, fresh-frozen tissue samples were pulverized, and genomic DNA was extracted using the QIAamp DNA Blood Kit (Qiagen). Prior to proteinase K digestion, RNase A was added to the tissue lysate to ensure thorough RNA removal. DNA quality was assessed by the 260/280 absorbance ratio and total DNA yield.

Nanopore library preparation and multiplexed sequencing

Prior to library preparation, genomic DNA samples were sheared to 8–10 kb fragments using a g-TUBE (Covaris). Sheared DNA was end-repaired and used for library preparation with the Native Barcoding Kit (SQK-LSK114). The multiplexed library was loaded onto an R10.4.1 PromethION flow cell (FLO-PRO114M) for long-read sequencing data acquisition. After sequencing, raw signal was base-called using Dorado with the model dna_r10.4.1_e8.2_400bps_sup@v5.0.0 to infer both genomic sequence and CpG base modification probabilities. Base calling produced a BAM file per sample containing aligned reads with per-CpG modification probability tags, which were used for all downstream analyses.

Nanopore sequencing quality control

Quality control was applied at two levels. First, read-level quality was assessed by the median per-read quality score, and samples with a median score below Q20 were excluded. Second, tissue composition was evaluated by CpG methylation deconvolution.28 To ensure that each sample was dominated by prostate epithelial cells, we calculated a purity ratio defined as the fraction of prostate epithelium divided by the fraction of the second most abundant cell type. Samples were retained only if they met both criteria: prostate epithelium fraction >40% and purity ratio ≥1.5. Applying these filters, three samples from the Moffitt cohort (Cancer.P10, Normal.P06, Normal.P08) and eight samples from the Mayo Clinic cohort (s86, s460, s307, s160, s315, s424, s1139, s112) were excluded. A summary of sequencing metrics for all retained samples is provided in Supplementary Table 1.

Copy number variation analysis

Copy number variation (CNV) analysis was performed using CNVkit29 to assess genome-wide DNA integrity. Because this analysis relies solely on alignment depth, all samples passing per-read quality thresholds were included regardless of methylation deconvolution status. A pooled reference was constructed by combining BAM files from all normal samples across both cohorts to derive per-bin copy number estimates. This reference was then applied to all prostate tumor samples. CNVkit was run in whole-genome sequencing mode (--method wgs), with very low-coverage bins excluded prior to segmentation (--drop-low-coverage). Genome-wide CNV patterns were visualized as a sample-wise heatmap, and scatter plots were used to illustrate focal copy number gains in selected samples.

Differential methylation region (DMR) identification between tumor and normal tissue

Following Dorado base calling, CpG-level methylation profiles were summarized into BEDGRAPH format using modkit (https://github.com/nanoporetech/modkit). CpG sites were retained if covered by ≥10 reads, and the resulting profiles were used as input for de novo DMR detection with metilene.30 Significant DMRs were defined by a Bonferroni-adjusted q-value ≤0.05 and an absolute tumor-normal methylation difference ≥10%. For downstream Genomic Regions Enrichment of Annotations Tool (GREAT) analysis, significant DMRs were stratified into hypermethylated and hypomethylated subsets. All tested genomic regions, irrespective of statistical significance, were used as the background reference for GREAT analysis.

Methylation entropy analysis

To evaluate epigenetic heterogeneity in prostate tissue, we performed methylation entropy analysis. Methylation entropy quantifies the diversity of methylation patterns by enumerating all possible methylation states across a fixed number of adjacent CpG sites. Entropy profiles were calculated for every set of three consecutive CpG sites using the modkit entropy function. An entropy value of zero indicates a uniform methylation pattern across reads, whereas higher values reflect increased heterogeneity and more complex methylation configurations. Results were stratified by the genomic context of the central CpG site (within or outside CpG islands) and by the genomic distance spanned by each CpG triplet. Given the large proportion of zero-entropy observations, downstream analyses were restricted to CpG triplets with non-zero entropy values.

Variant calling and read-level allele-specific methylation analysis

Variant calling was performed per sample using PEPPER-Margin-DeepVariant, a pipeline31 designed to improve SNP calling accuracy in homopolymer regions. Called variants were annotated against dbSNP build 153, and annotated SNPs were filtered by variant quality score (QUAL ≥10) and variant allele fraction (VAF ≥ 0.1). Filtered VCF files were merged across samples and further refined by minimum allele count (--min-ac 3), heterozygous sample allele depth (≥3 heterozygous samples with FORMAT/DP ≥10, FORMAT/AD[REF] ≥5), and post-merge variant quality (QUAL ≥30). Only biallelic SNPs were retained for allele-specific methylation (ASM) analysis.

To prepare methylation signals for allele-level comparisons, individual BAM files were processed with the modkit call-mods subcommand, which binarizes per-CpG modification probabilities to 0% or 100% using sample-specific thresholds. This step ensures consistent methylation calls across samples. Processed BAM files were then merged with sample identifiers encoded in the read group (RG) tag.

Each input variant’s coordinate and allele information were used to label reads in the merged BAM with a variant allele tag (VA), indicating whether each read carries the reference or alternative allele. Labeled reads were processed with modkit pileup with sample (RG) and allele (VA) partitioning enabled, generating per-allele BEDGRAPH files. DMRs between reference and alternative allele groups were identified de novo using metilene with default settings. Prior to segmentation, samples with ≥90% missing CpG methylation values across a given region were excluded to prevent instability from sparse coverage. The full script developed has been deposited in the nanoASM GitHub repository (https://github.com/Yijun-Tian/nanoASM).

Results

Quality Control and Initial Characterization of Nanopore Long-Read Sequencing Data

Following nanopore sequencing, we performed quality control at multiple levels. Per-read quality score distributions across all samples from both the Mayo Clinic and Moffitt cohorts showed peaks well above Q20, consistent with the expected performance of the R10.4.1 PromethION platform (Figure S1A). Notably, within the Moffitt cohort, prostate cancer tissue samples exhibited significantly lower per-read quality scores compared with their paired normal counterparts (KS-test p < 2.2×10−16; Figure S1B). We hypothesize that this reflects a greater diversity of complex or atypical base modifications in cancer cells, which may reduce base calling confidence. Consistent with this interpretation, our previous work22 demonstrated that the concordance between nanopore and bisulfite sequencing is systematically lower in cancer cell lines than in normal cell lines, suggesting a similar underlying phenomenon.

To characterize cellular heterogeneity in the sequencing data, we performed cell-type deconvolution using nanopore-derived CpG methylation profiles referenced against the ATLAS reference pane.28 The inferred cell-type composition confirmed that most samples were dominated by prostate epithelial cells (Figure S1C). However, three samples from the Moffitt cohort (Cancer.P10, Normal.P06, Normal.P08) and eight samples from the Mayo Clinic cohort (s86, s460, s307, s160, s315, s424, s1139, s112) - highlighted in red in Figure S1C — showed insufficient prostate epithelial fractions and were excluded from downstream analyses.

Copy number variation (CNV) analysis was performed on both normal and tumor prostate tissues using alignment depth information. Despite most tumors being derived from early-stage surgical resections, cancer tissues exhibited substantially greater copy number variation than matched normal tissues (Figure S2A). For instance, a log2 ratio scatter plot of chromosome 8 in patient 06 revealed the characteristic pattern of 8p loss and 8q gain/amplification, a hallmark of prostate cancer (Figure S2B).

Genome-wide Identification of Differentially Methylated Regions in Prostate Cancer

Using the high-quality nanopore long-read sequencing data, we performed de novo identification of differentially methylated regions (DMRs) between normal and cancer tissues. After annotating these DMRs using GREAT analysis, we observed that DMRs exhibiting higher methylation in cancer intersected significantly more transcription start site (TSS)–proximal regions compared with DMRs that were more highly methylated in normal tissues (Figure 1A).

Figure 1. Characterization of differentially methylated regions between prostate cancer and normal tissue.

Figure 1.

(A) Distribution of region–gene associations by distance to the nearest transcription start site (TSS) for hypermethylated DMRs (Cancer > Normal, red), hypomethylated DMRs (Cancer < Normal, blue), and background genomic regions (gray). (B) Distribution of the number of genes associated per DMR for hypermethylated (red), hypomethylated (blue), and background (gray) regions. (C) Gene Ontology (GO) Molecular Function enrichment analysis of hypermethylated (top, red) and hypomethylated (bottom, blue) DMRs, performed using GREAT. Bar lengths represent −log10(hypergeometric p-value); values are labeled at the right of each bar. (D) Scatter plots of Spearman correlation between CpG methylation and gene expression as a function of distance to TSS, derived from TCGA PRAD normal (top) and cancer (bottom) samples. Red dots represent hypermethylated DMRs (Cancer > Normal) and blue dots represent hypomethylated DMRs (Cancer < Normal). (E) Heatmap of DMR methylation levels across Moffitt cohort samples (normal, green; tumor, orange). Selected DMR coordinates are labeled on the left. DMRs clearly stratify tumor from normal samples, with hypermethylated regions (red) predominantly enriched in cancer and hypomethylated regions (blue) enriched in normal tissue. (F) Heatmap of expression levels of DMR annotated genes in TCGA PRAD samples (normal, green; tumor, orange). Selected known prostate cancer–associated genes (ZNF154, AOX1, DLEC1, RASSF1, WNT3A, GPX3, TWIST1) are labeled on the left. Color scale represents methylation z-score (−2 to +2).

When further annotating DMRs to known human genes, we found that 228 cancer-hypermethylated DMRs and only three normal-hypermethylated DMRs did not overlap with any gene regions (Figure 1B). Functional annotations using Gene Ontology (GO) Molecular Function revealed that cancer-hypermethylated DMRs were significantly enriched in genomic regions associated with transcriptional regulation, including sequence-specific DNA binding and RNA polymerase II regulatory regions. In contrast, DMRs hypermethylated in normal tissues were enriched in regions associated with ion channel receptor activity, aminotransferase activity, and related molecular functions (Figure 1C). Based on the DMR-to-gene mapping, we then leveraged RNA expression profiles from the TCGA prostate cancer cohort (PRAD) to examine the relationship between gene expression and DMR methylation differences. In both the PRAD normal and cancer subgroups, the strongest correlations between gene expression and DNA methylation were consistently enriched near transcription start sites, particularly for DMRs that were hypermethylated in cancer tissues. This observation is consistent with the transcription regulatory functions identified in our functional annotation analysis (Figure 1D). To highlight representative DMR–gene associations, we visualized the aggregated methylation levels of the most significant DMRs (p ≤ 0.00001) in the Moffitt cohort (Figure 1E), alongside the corresponding gene expression levels in matched tumor-normal pairs from the TCGA prostate cancer cohort (Figure 1F). Several genes previously reported to exhibit recurrent methylation alterations in prostate cancer were highlighted in the heatmap, including RASSF1,32,33 TWIST1,34,35 AOX1,36,37 GPX3,3840 DLEC1,41 ZNF154,42 and WNT5A.43,44 As expected, the genes showed over-expression in TCGA-PRAD tumor tissue mostly showed hypomethylation in Moffitt tumor tissues.

RNA expressions associated with epigenetic profiling detected in nanopore DNA sequencing.

Current models suggest that DNA modifications act as stable and functional epigenetic markers that are strongly correlated with gene expression. However, most approaches used to investigate DNA modifications rely on capture-based methods, and many studies have been conducted primarily in mouse models or in vitro systems.4547 By enabling single-base resolution detection of both 5mCG and 5hmCG in native genomic DNA, nanopore sequencing in our cohort provides an opportunity to re-evaluate these established relationships in human prostate tissues.

We first visualized the average 5mCG (Figure 2A) and 5hmCG (Figure 2B) modification levels across genes grouped by their expression levels based on GTEx prostate tissue data (n = 231). Notably, the average 5mCG level in transcription start site (TSS) regions decreased with increasing gene expression (Figure 2C), whereas the 5hmCG level in gene body regions increased with gene expression (Figure 2D). These patterns suggest that the two epigenetic modifications may play distinct regulatory roles across different genomic regions.

Figure 2. Genome-wide 5mCG and 5hmCG profiles across gene bodies stratified by expression level.

Figure 2.

(A) Meta-gene profiles of 5mCG percentage across gene bodies (TSS to TES, ±1500 bp flanking) for normal (blue, top) and cancer (red, bottom) prostate tissue samples, stratified by GTEx RNA-seq expression expressivity bins (<1%, 1–5%, 5–10%, 10–30%, 30–50%, 50–70%, ≥70%). The number of genes in each bin is indicated above each panel. (B) Meta-gene profiles of 5hmCG percentage across gene bodies for normal (blue, top) and cancer (red, bottom) prostate tissue samples, using the same expression percentile stratification as in (A). (C) Spearman correlation between 5mCG levels and gene expression percentile at the TSS (±200 bp, solid lines) and across the gene body (dashed lines) for cancer (red) and normal (blue) tissue. (D) Spearman correlation between 5hmCG levels and gene expression percentile at the TSS (±200 bp, solid lines) and across the gene body (dashed lines) for cancer (red) and normal (blue) tissue. (E) Heatmap of 5hmCG-based GSVA enrichment scores for significant HALLMARK gene sets (top, p ≤ 0.05) and C1 positional gene sets (bottom, p ≤ 0.05) across Moffitt cancer and normal samples. Red indicates positive enrichment and green indicates negative enrichment. (F) Heatmap of 5hmCG levels across individual mitochondrial genes in cancer (left) and normal (right) Moffitt samples. Mitochondrial genes in the leading edge are labeled in red while those are not labeled in black.

Given that 5hmCG has been reported to be strongly associated with transcriptional activity, we further performed gene set enrichment analysis (GSEA) using gene-level aggregated 5hmCG measurements across the HALLMARK (h) and positional (c1) gene set collections. Interestingly, multiple hallmark pathways and autosomal positional gene sets exhibited higher 5hmCG levels in normal tissues compared with cancer tissues (Figure 2E). In contrast, mitochondrial gene sets displayed significantly higher 5hmCG levels in cancer tissues relative to normal tissues (Figure 2F).

Entropy-based characterization of CpG methylation landscapes

In addition to analyzing the 5hmCG and 5mCG information on a per CpG-site basis, we adopted methylation entropy in evaluating the methylation diversity from the nanopore sequencing data. Methylation entropy recognizes epigenetic diversities defined by combinatory modifications including 5hmCG, 5mCG and CG without modifications to describe the epigenetic diversity in each set of fixed CpG site triplets (Figure 3A). At the genome-wide level, regional entropy was comparable between normal and cancer samples, indicating a highly similar background distribution of CpG methylation configurations (Figure 3B). Consistent with this, no significant differences were observed after stratifying by CpG island status (Figure S3A) or gene-centric annotations (Figure S3B). However, when focusing on normal–cancer DMRs, a directional difference was observed: regions that become hypermethylated in cancer (DMR normal < cancer) showed lower entropy in cancer, whereas hypomethylated regions (DMR cancer < normal) did not show a consistent change (Figure 3B).

Figure 3. Methylation entropy characterizes epigenetic heterogeneity across DMRs and chromatin states.

Figure 3.

(A) Schematic illustrating the calculation of methylation entropy (ME) for a set of three adjacent CpG sites. Each row represents a distinct methylation pattern observed across sequencing reads, with filled circles (black: 5mCG, gray: 5hmCG, white: unmodified CG) indicating the modification state at each position. The probability of each pattern is enumerated, and ME is computed as −0.2 × Σ p(i) × log2p(i). (B) Paired dot plots comparing regional median methylation entropy between matched normal (green) and cancer (red) samples from the Moffitt cohort, stratified by genomic context: background CpG triplets (n = 100,957; left), hypomethylated DMRs in cancer (DMR normal > cancer, n = 2,879; middle), and hypermethylated DMRs in cancer (DMR cancer > normal, n = 1,654; right). (C) Hexbin density scatter plots of methylation entropy versus 5mCG percentage (top row) and 5hmCG percentage (bottom row) for the normal tissue sample from patient P01, stratified by genomic context: background (left), DMR normal > cancer (middle), and DMR cancer > normal (right). Color scale indicates bin count density. Spearman correlation coefficients (ρ) are shown for each panel; values highlighted in red indicate significant differences from the background correlation. (D) Equivalent hexbin density scatter plots as in (C) for the matched cancer tissue sample from patient P01. (E) Hexbin density scatter plots of methylation entropy versus 5mCG percentage stratified by histone modification chromatin state, including background, H3K27me3 (repressive), H3K36me3 (gene body), H3K27ac (active enhancer), H3K4me1 (poised enhancer), and H3K4me3 (active promoter). Spearman ρ values are shown for each context. (F) Hexbin density scatter plots of methylation entropy versus 5hmCG percentage across the same chromatin state categories as in (E).

To further characterize this pattern, we examined the relationship between entropy and modification levels in a representative sample (Figure 3CD). In regions corresponding to DMR normal < cancer, the correlation between 5mCG and entropy was positive (ρ = 0.42) in normal tissue but became negative (ρ = −0.12) in cancer. This shift indicates a change in how CpG methylation levels relate to local configuration variability. In the same regions, entropy showed a consistent positive association with 5hmC, suggesting that variation in entropy tracks more closely with 5hmCG than with 5mCG. In contrast, background region sets showed more stable or weaker relationships. Consistent with their genomic distribution (Figure 1C), hypermethylated DMRs are enriched near transcription start sites and regulatory elements, whereas hypomethylated DMRs are more broadly distributed. Together, these results describe a context-dependent change in the relationship between methylation level and entropy in cancer.

Finally, we assessed whether CpG methylation and entropy profiles alone could recapitulate underlying chromatin context (Figure 3EF). Using nanopore-derived signals, regions stratified by histone modification annotations from normal prostate tissue displayed distinct methylation–entropy patterns consistent with their regulatory states. More specifically, repressive (H3K27me3) and gene body (H3K36me3) chromatin states displayed increased density associated with strong 5mCG signals (>0.75), weak 5hmCG signals (<0.1), and low methylation entropy (0-0.5). Promoter regions marked by H3K4me3 showed low levels of both 5mCG (<0.25) and 5hmCG (<0.1) and correspondingly low entropy (0-0.1). Interestingly, enhancer-associated histone modifications (H3K27ac and H3K4me1) exhibited intermediate 5mCG (0.25-0.75) and high 5hmCG (0.1-0.25) levels together with high entropy (0.5-1.0). Across all contexts, entropy consistently peaked at intermediate modification levels on 5mCG, with reduced entropy observed toward fully methylated or unmethylated states (Figure 3CF). This intermediate-state dependence was reproducible across samples and genomic annotations, indicating that CpG configuration diversity is maximized in partially modified regions.

Consistent patterns were also observed in a cancer type–specific chromatin accessibility framework (Figure S3C), where regions defined by accessibility across tumor types exhibited the same enrichment of high entropy at intermediate methylation levels. Together, these results highlight the intermediate methylation state as a key determinant of local CpG configuration variability and support the use of methylation–entropy profiles as informative proxies for underlying chromatin state.

Theoretical Evaluation of Statistical Efficiency Between ASM and Population-Based mQTL testing

To improve sensitivity for detecting cis-regulatory methylation effects, we implemented an allele-specific methylation (ASM) framework based on within-individual allelic contrasts. In heterozygous individuals, both alleles are measured within the same cellular and experimental context, allowing shared sample-level variation to be removed through paired comparison. As a result, the residual variation is primarily driven by technical noise, leading to a substantial reduction in variance relative to population-level analyses.

This property contrasts with traditional mQTL approaches, which rely on between-individual comparisons of bulk methylation levels and are therefore more susceptible to sample-level heterogeneity and measurement variability. In addition, ASM analysis leverages sequencing reads that capture both allele identity and CpG methylation states, often aggregating information across multiple CpGs within a local genomic region. This aggregation further stabilizes methylation estimates and reduces effective measurement noise. The advantage is particularly pronounced in long-read sequencing data, where haplotype information and contiguous CpG methylation patterns are directly observed within individual reads.

The efficiency gain of ASM increases with the proportion of heterozygous individuals in the cohort, as more samples contribute informative allelic contrasts. Under a Gaussian approximation, and defining bulk methylation as the average of allele-specific measurements, the relative statistical efficiency of ASM compared to population-level testing can be approximated as

R(p)4p(1p)r+121ρ

where p is the allele frequency, r is the ratio of sample-level to technical variance, and ρ represents the correlation of allele-specific technical errors within individuals. This formulation yields more conservative but more realistic conditions under which ASM provides improved statistical power.

In particular, for common variants (p≈0.5) and settings with substantial sample-level variability (large r), ASM can achieve markedly higher efficiency than population-based mQTL analysis (Figure S4A). Consequently, comparable statistical power can be obtained with substantially smaller sample sizes (Figure S4BC). Full derivations and simulation results are provided in Supplementary Note 1.

Allele-specific Methylation Identification from Nanopore Long-read DNA sequencing

Given the evidence of the improved statistical power in the ASM framework over traditional mQTL approaches, we sought to implement a pipeline to identify the read-level allele-specific methylation region from Nanopore sequencing BAM files with base modification tags. For each heterozygote variant, we labeled and grouped the nanopore long reads according to the allele they carried. After labeling, we performed the de-novo DMR identification within the reference and alternative allele group (Figure 4A). Given any variant, we were able to use this simple method to identify differential methylation domains among the read spanning regions across the whole genome. From 4,076,763 and 5,788,657 input variants, we identified 406,011 (Figure 4B) and 391,918 (Figure 4C) allele-specific methylation events within the tumor and normal subgroups respectively. From these findings, we further defined shared ASM events as those associated with the same SNP and overlapping differential methylation regions between tumor and normal. To evaluate the correlation in effect size between the ASM and mQTL method, we intersected the ASM findings with the Pancan-meQTL results from TCGA prostate cancer cohort by the SNP id and the DMR-to-probe inclusion. We found majority of overlapped associations showed the same trend of allelic effect (Figure 4D4E). Interestingly, although we observed large overlaps in the input SNPs between normal and tumor subgroups (Figure 4F), SNPs identified with ASM (Figure 4G) and the relevant ASM (Figure 4H) overlapping is relatively small. We further compared the genomic size and the absolute methylation difference of the DMR between normal and cancer subgroups and found that cancer ASM showed significantly larger DMR size (Figure 4I) and stronger allelic methylation difference (Figure 4J) than normal ASM. We had also explored whether allele-related CpG site change could influence DMR methylation direction. We observed that when a SNP is in proximity to a DMR, allele forming a CpG site is associated with hypermethylation, especially in those ASM events identified in normal tissue (Figure 4K).

Figure 4. Genome-wide allele-specific methylation analysis using nanopore long-read sequencing.

Figure 4.

(A) Schematic of the allele-specific methylation (ASM) analysis framework. Multiplexed nanopore long-read sequencing simultaneously captures base modifications and sequence variants on individual DNA molecules. After basecalling, variant calling, call-mods processing, and demultiplexing, reads are labeled by the allele they carry (reference or alternative). Per-allele methylation signals are aggregated across samples, and de novo DMRs between reference and alternative allele groups are identified by binary segmentation. The ASM delta (Di = Yref − Yalt) quantifies the methylation difference between alleles. (B-C) Genome-wide Manhattan-style scatter plot of absolute ASM delta (%) for all significant ASM-DMRs identified in prostate cancer (B) and normal (C) tissue, plotted by chromosomal position. Green dots represent ASM-DMRs unique to cancer; red dots represent ASM-DMRs shared between cancer and normal tissue; Blue dots represent ASM-DMRs unique to normal tissue. (D) Scatter plot comparing ASM delta (%) in cancer tissue against TCGA PRAD mQTL beta coefficients. (E) Scatter plot comparing ASM delta (%) in normal tissue against TCGA PRAD mQTL beta coefficients. (F) Proportional Venn diagram showing the overlap of input SNPs (n = 4,007,251) between cancer (red, n = 4,076,763) and normal (blue, n = 5,788,657) tissue analyses. (G) Proportional Venn diagram showing the overlap of SNPs (n = 46,276) with significant ASM between cancer (red, n = 330,683) and normal (blue, n = 328,395) tissue. (H) Proportional Venn diagram showing the overlap of ASM-DMR genomic intervals (n = 39,176) between cancer (red, n = 366,835) and normal (blue, n = 352,742) tissue. (I) Density distributions of ASM-DMR region size (bp, log2 scale) for private (left) and shared (right) ASM-DMRs in cancer (red) and normal (blue) tissue. Dashed vertical lines indicate median values. (J) Density distributions of absolute ASM delta (%) for private (left) and shared (right) ASM-DMRs in cancer (red) and normal (blue) tissue. (K) Stacked Sankey plots showing the allelic distribution in CpG site (Alt allele, teal; Ref allele, pink) and the direction of allelic-methylation (hyper-methylation allele: Ref or Alt) for ASM-DMRs stratified by tissue context (cancer only, normal only, shared in cancer and shared in normal) and by SNP-to-DMR distance (reside: SNP within DMR; 0–200 bp: proximal; >200 bp: distal). Numbers within bars indicate the percentage of ASM-DMRs in each category. SNPs residing within or proximal to DMRs show a strong tendency for the alternative allele to be the hypermethylated allele, while this asymmetry diminishes with increasing SNP-to-DMR distance.

Consistency of ASM across mQTL and eQTL signals

Having observed the high consistency between ASM and mQTL signals (Figure 4DE), we next sought to better understand where and why these signals differ, and to extend this comparison to eQTL effects. To this end, we first defined discrepancy between mQTL and ASM based on the agreement between DMR-level methylation differences (DMR delta) and mQTL effect sizes (Figure 5A). For each SNP–DMR (ASM) and SNP-CpG (mQTL) pair, discrepancy was defined as the inconsistency in direction between allele-specific regional methylation delta and probe-level mQTL effect size. Interestingly, discrepancy between TCGA PRAD mQTLs and cancer-derived ASM showed a clear distance-dependent pattern (Figure 5B). Discrepancy decreased sharply from very short SNP–CpG distances (<50 bp) to intermediate distances (~75–150 bp), and then gradually increased with increasing distance, reaching higher levels in distal regions (>2000 bp). In contrast, discrepancy decreased monotonically with increasing CpG density, with higher discrepancy observed in CpG-sparse regions and lower discrepancy in CpG-dense regions (density > 0.07-0.08).

Figure 5. Concordance and discrepancy between nanopore ASM and bulk mQTL/eQTL effect sizes.

Figure 5.

(A) Schematic defining the metrics used to compare ASM-DMR delta values with bulk mQTL and eQTL effect sizes. The ASM delta (Δ methylation = allele A − allele G) is compared with the mQTL methylation beta coefficient or eQTL normalized enrichment score (NES). (B) TCGA mQTL and cancer ASM discrepancy rate as a function of SNP-to-CpG distance (left, binned in bp) and local CpG density (right). (C) Bar plot comparing the eQTL concordance rate between SNPs residing within ASM-DMRs (dark purple, N = variant in DMR) and SNPs outside ASM-DMRs (light purple, N = variant outside DMR). (D) Concordance rate as a function of absolute SNP-to-TSS distance (x-axis, log scale) across nine gene-relative annotation contexts (1–5 kb upstream, 5′UTR, promoter, first exon, 3′UTR, exon, CDS, intron–exon boundary, exon–intron boundary). Red lines represent SNPs residing inside ASM-DMRs and blue lines represent SNPs outside ASM-DMRs. Shaded bands indicate 95% confidence intervals. The number of SNP–gene pairs (N) is shown for each context in matching colors. (E) Correlation between ASM delta and eQTL NES (y-axis) as a function of relative SNP position to the DMR center (x-axis, DMR-centered with outside regions compressed), across the same nine gene-relative annotation contexts as in (D). Red lines show the running correlation with 95% confidence intervals (shaded). Gray shading marks the interior of the DMR (positions −10 to +10).

Having defined discrepancy between ASM and mQTL signals (Figure 5AB), we next examined whether allele-specific methylation captures downstream regulatory effects on gene expression by comparing ASM direction and magnitude with GTEx eQTL signals. Concordance was defined based on the agreement between the direction of allele-specific methylation (DMR delta) and the direction of gene expression changes (GTEx prostate tissue eQTL normalized effect size, NES), such that concordant pairs reflect consistent regulatory effects, i.e., higher methylation associated with lower gene expression (Figure 5A). Using this framework, variants located within DMRs showed modest but significantly higher concordance with eQTL directionality compared to variants outside DMR (p = 0.028; Figure 5C), suggesting that SNPs residing within ASM-defined regions are more likely to reflect functional regulatory effects on gene expression. We next examined how concordance varies with SNP-to-TSS distance across genomic annotations (Figure 5D). Concordance showed a robust decay with increasing distance from transcription start sites, with the strongest agreement observed for variants proximal to TSS and progressively diminishing at distal regions. This pattern was consistent across most genomic features, indicating that proximity to canonical regulatory elements remains a strong determinant of expression-associated allele-specific epigenetic effects. To test whether the relative position of a SNP to DMR influences the allele-specific methylation to GTEx effect, we evaluated the correlation between allele-specific methylation (ASM delta) and GTEx (NES) as a function of SNP position relative to DMRs (Figure 5E). Interestingly, the correlation pattern was genomic-context specific: promoter- and 5UTR-associated DMRs showed consistently negative ASM–eQTL correlations across most relative positions, whereas 3UTR-associated DMRs displayed weaker positive correlation and in distant SNP-to-DMR positions.

Hotspot ASM at the IRX4 locus nominates a candidate androgen-responsive regulatory element

The IRX4 locus has been previously reported as a prostate cancer GWAS risk locus and is also associated with gene expression variation (eQTL), indicating its potential regulatory role. After examining the ASM pattern in this region, we observed that an ASM cluster with larger methylation differences were concentrated within a relatively narrow genomic interval (chr5:1888662-1890105), while weaker signals were more dispersed (Figure 6A). To further characterize the regulatory context, we integrated AR ChIP–seq data (PRJNA752082) and found that this interval overlapped AR binding peaks that were higher in prostate adenocarcinoma than in normal prostate (Figure 6B). In Castration-Resistant Prostate Cancer (CRPC) tissue (PRJNA777879), AR binding at this site increased upon DHT treatment and was accompanied by elevated H3K27ac signal (Figure 6C), consistent with androgen-responsive regulatory activity. Notably, the local SNP structure lacks a clear LD block (Figure 6D), and ASM signals are observed across multiple nearby variants. Among these, rs6885084 shows a cancer-specific methylation difference, consistent with the cancer-associated changes in AR ChIP–seq profiles. This difference is further supported by read-level methylation patterns, which show a larger allelic methylation difference in cancer samples than in normal samples (Figure 6E), highlighting rs6885084 as a candidate variant for further functional investigation.

Figure 6. Case study of a prostate cancer GWAS locus on chromosome 5 with allele-specific methylation and androgen receptor binding.

Figure 6.

(A) Locus overview showing IRX4 gene models (top), mQTL, eQTL, and GWAS prostate cancer risk SNPs (colored ticks), and ASM-DMRs for each input SNP (rows). Each row corresponds to one SNP (rsID labeled on left); colored rectangles indicate ASM-DMR intervals, with fill color representing methylation difference (%) between alleles. Cancer SNPs are shown in red, normal SNPs in blue, and shared SNPs in green. Normal DMRs are outlined in blue and cancer DMRs in red. (B) AR ChIP-seq signal tracks from prostate adenocarcinoma (top, red) and normal prostate (bottom, blue) tissue samples at the same locus. The orange-shaded region (chr5:1,888,662–1,890,105) marks a DMR cluster overlapping a prominent AR binding peak that is substantially stronger in cancer than in normal tissue. (C) AR ChIP-seq and H3K27ac ChIP-seq tracks from two castration-resistant prostate cancer (CRPC) (CRPC35 and CRPC167) under vehicle control and DHT stimulation conditions. The orange-shaded DMR cluster region shows strong DHT-induced AR binding and H3K27ac enrichment, indicating androgen-responsive enhancer activation. (D) Linkage disequilibrium (LD) heatmap of all input SNPs across the locus, with r2 values represented by color intensity (white = low LD, red = high LD). (E) Methylation heatmap for the lead ASM-DMR (chr5:1,888,792–1,890,105; FDR = 3.92×10−2) stratified by allele at rs6885084 (G vs. C, green arrow). Each row represents the per-site methylation profile of one allele in one patient sample, with red indicating higher methylation and blue indicating lower methylation.

The ASM-associated locus linked PSCA gene expression to rs4736369

PSCA gene is being evaluated as a therapeutic target in metastatic castration-resistant prostate cancer (mCRPC).48 We therefore examined whether allele-specific methylation (ASM) at this locus was associated with PSCA regulation. Across the tested variants, we identified a prominent ASM signal centered on rs4736369 in both normal and cancer samples (Figure 7A). The corresponding ~500-bp DMR showed methylation differences of 50.7% and 53.4% in normal and cancer cohorts, respectively (Figure 7B). Although no significant eQTL signal was detected in the GTEx normal prostate cohort, multiple significant eQTL and isoform-QTL associations were observed in the MAYO prostate cohort (prostate cancer–adjacent normal tissue) (Figure 7C). Gene set enrichment analysis further indicated that PSCA expression was associated with multiple prostate-related pathways (Figure 7D). In sashimi plot, RNA-seq coverage profiles showed differences in junction usage between genotype groups, with higher coverage and junction reads observed in C/C compared to T/T samples (Figure 7E), consistent with alternative splicing at this locus. Leveraging long-read phasing information from nanopore sequencing, we performed haplotype-resolved analysis in heterozygous individuals. Allele directed haplotype-specific expression analysis, integrating nanopore DNA and RNA-seq data, showed that variants phased with the rs4736369 C allele are associated with higher RNA expression (Figure 7F). Consistently, ChIP–seq data (PRJNA540151) from heterozygous prostate tissue demonstrated higher H3K4me3 signal on the rs4736369 C allele than on the T allele (Figure 7G).

Figure 7. Allele-specific methylation at a prostate cancer GWAS locus on chromosome 8 links rs4736369 to transcriptional regulation.

Figure 7.

(A) Locus overview showing PSCA gene models (top), eQTL SNPs (blue ticks), and ASM-DMR intervals for each input SNP (rows). Each row corresponds to one SNP (rsID labeled on left); colored rectangles indicate ASM-DMR intervals with fill color representing methylation difference (%) between alleles. Cancer SNPs are shown in red, normal SNPs in blue, and shared SNPs in green. Normal DMRs are outlined in blue and cancer DMRs in red. The lead SNP rs4736369 (green arrow) is highlighted. (B) Per-CpG site methylation heatmaps for the ASM-DMRs at rs4736369 in normal and cancer tissue, stratified by allele group (T vs. C). Each row represents the per-site methylation profile of one allele in one patient sample; red indicates higher methylation and blue indicates lower methylation. (C) Boxplots of normalized gene expression for five transcripts at this locus (ENSG00000167653, ENST00000301258, ENST00000513264, ENST00000505305, ENST00000510969) stratified by rs4736369 genotype (TT, TC, CC; sample sizes in parentheses). (D) Hallmark pathway enrichment analysis for genes significantly associated with the PSCA eQTL (FDR < 0.1) at this locus. Bar length represents −log10(FDR) (red) and normalized enrichment score (NES, blue). (E) Sashimi plots of RNA-seq read coverage and splice junction usage at the locus, generated from BAM files merged by rs4736369 genotype: homozygous T/T (top, red) and homozygous C/C (bottom, orange). Arc heights and labeled scores indicate splice junction read counts.

(F) Haplotype-directed allele-specific RNA-seq coverage at rs4736369, where haplotype phasing was determined from nanopore long-read DNA sequencing. Each dot represents one Mayo Clinic sample, colored by sample ID, plotted as the allele fraction of the C allele at the rs4736369 position in DNA (top) and RNA-seq (bottom); dot size in the RNA-seq panel reflects SNP read depth. DNA allele fractions cluster near 0, 0.5, or 1.0, confirming germline genotype assignments. RNA-seq allele fractions show systematic deviation from 0.5 in heterozygous samples, indicating allele-specific expression consistent with the eQTL direction and the allele-specific methylation observed in (B). (G) H3K4me3 ChIP-seq tracks at the locus for matched normal and tumor tissue from two patients (1820 and 2208). The red-shaded region marks the ASM-DMR interval. Allele counts at rs4736369 (C/T) are indicated for each sample. H3K4me3 signal is present right near the DMR in both normal and tumor tissue, indicating that this region overlaps an active promoter, and the signal intensity varies between samples with different allele compositions, supporting allele-specific chromatin activity at this locus.

Discussion

Nanopore sequencing uses deep neural networks with alignment-free decoding to directly infer nucleotide sequences and base modifications from raw ionic current signals, enabling robust measurement of regional DNA methylation. Leveraging these single-molecule, multimodal data, we first characterized genome-wide tumor–normal differential methylation regions and their genomic and functional features. Our whole-genome nanopore sequencing design enabled comprehensive profiling of cancer-associated methylation, revealing strong enrichment of DNA-binding and transcriptional regulatory functions. These findings support a model in which early-stage prostate cancer–associated hypermethylation preferentially targets CpG islands, whereas hypomethylation occurs more broadly across the genome.4952

To further assess the relationship between DNA modification and transcriptional activity, we examined modification patterns across genes stratified by prostate tissue-specific expression levels derived from GTEx RNA-seq data. Both 5mCG and 5hmCG exhibited clear expression-dependent distributions, characterized by pronounced depletion around transcription start sites (TSS) and enrichment across gene bodies. These patterns suggest that DNA modifications may serve as surrogate markers of transcriptional activity. To this end, we explored the functional relevance of these observations by performing GSEA using gene-level aggregated 5hmCG signals. Consistent with previous reports 5355, hallmark and chromosomal gene sets exhibited modestly higher 5hmCG levels in normal samples compared to tumor tissues. In contrast, enrichment analysis of positional gene sets (C1) revealed that mitochondrial genes showed significantly elevated 5hmCG levels in cancer relative to normal tissue, consistent with increased oxidative phosphorylation activity56 and mitochondrial metabolism57 observed in tumors. This distinct pattern may reflect reduced cellular heterogeneity and lower background variability in the mitochondrial genome compared to autosomal regions, thereby enabling more robust detection of modification differences even in small sample sizes. In addition to mtDNA copy number,58,59 mutation60,61 and 5mCG methylation,62,63 these observations raise the possibility that mitochondrial 5hmCG modification patterns could serve as potential biomarkers for tumor detection, particularly in liquid biopsy or circulating tumor DNA (ctDNA) settings.

Beyond per-site methylation measurements, long-read sequencing enables the analysis of fragment-level combinations of CpG sites, allowing quantification of regional methylation entropy. This fragment-level view captures coordinated methylation patterns and provides a measure of epigenetic heterogeneity that cannot be inferred from average methylation levels alone. Notably, conventional methylation entropy metrics are typically derived from unmodified cytosine and 5mCG states,64,65 whereas the entropy framework used in this study incorporates both 5mCG and 5hmCG signals, thereby capturing additional layers of epigenetic information. Contrary to the intuition that cancer epigenomes are more disordered, we observed a significant reduction in methylation entropy specifically within DMRs that gained methylation in cancer tissue, while entropy in hypomethylated or unchanged DMRs remained unaltered. The entropy reduction at cancer hypermethylated DMRs reflects the clonal fixation of a uniform 5mCG state across the tumor cell population — a process underpinned by two converging mechanisms. First, 5hmCG is broadly lost in cancer as a consequence of TET downregulation66 or dysfunction.67,68 Second, once 5hmCG is depleted, the remaining hemi-5mCG at replication forks is faithfully restored by DNMT1, which has a strong and well-characterized preference for hemi-methylated over hemi-hydroxymethylated substrates.45 The 5mCG state is therefore heritably propagated through cell division while 5hmCG is passively diluted, and the expanding tumor clone converges on a low-entropy, uniformly methylated pattern.46,47 Across different histone modifications, the highest entropy was colocated at enhancer markers, which are characterized by partial demethylation and enrichment of 5hmCG,69,70 generating a mixed modification landscape across reads. This confirms that entropy is not a linear marker of chromatin activity but specifically reports on modification heterogeneity, which is maximal at regulatory elements poised between methylated and demethylated states.

The structured methylation patterns revealed by expression-dependent analyses and entropy profiling suggest that epigenetic states are not purely stochastic but are shaped by underlying mechanisms, including local germline variation. To enable this analysis, we developed a simple allele-specific methylation (nanoASM) pipeline to systematically identify and quantify allele-specific methylation events from nanopore long-read sequencing data. A key assumption underlying this method is that functionally relevant germline variants exert local (cis) effects on the epigenetic landscape, preferentially influencing nearby CpG sites rather than distal regions. Although many mQTL studies17,7174 have demonstrated predominantly cis-acting genetic effects of SNPs on nearby CpG sites, the distance thresholds used to define cis associations vary widely across studies, typically ranging from 250 kb to 1 Mb. This variability partly reflects inherent limitations of genotype-to-single-CpG association designs. We have previously discussed potential biases associated with this approach.22 In this study, we further observed that discrepancies between genetic and methylation signals were greater when SNPs were in proximity to the target CpG sites. This may be explained by probe-binding artifacts in array-based platforms, where SNPs within or near probe sequences can introduce allele-specific hybridization biases, thereby affecting methylation measurements. Similar effects have been reported previously.20,21

In our nanoASM framework, we propose an approach that aggregates allelic methylation signals across local 8–20 kb regions flanking germline variants, enabling robust detection of ASM and improving the identification of regulatory SNPs. Interestingly, although the germline SNPs were largely shared between normal and cancer tissues, the corresponding ASM signals were mostly not reproduced across conditions, indicating substantial differences in the epigenetic landscape between normal and cancer states. We found that part of this discrepancy could be attributed to loss of heterozygosity (LOH) in cancer, while a larger fraction likely reflects genuine differences in epigenetic regulation between normal and tumor tissues. Notably, cancer-derived ASM regions exhibited broader genomic span and larger methylation differences compared to normal tissues. This observation is consistent with previous next-generation sequencing (NGS) studies 7578 and suggests that epigenetic dysregulation in cancer operates at a larger regional scale. Leveraging nanopore long-read sequencing, which preserves haplotype-resolved allele information, we further confirmed that SNPs disrupting CpG sites are associated with local hypomethylation of DMRs, whereas SNPs creating CpG sites are associated with regional hypermethylation.79,80

Another advantage of variant-level ASM analysis is that, compared to haplotype-based approaches, it reduces diffuse effects from imprinting regions and better isolates the local impact of individual alleles on nearby methylation patterns. 81,82. In addition, haplotype-based phasing can result in heterogeneous block structures across individuals, complicating cross-sample comparisons and may introduce inconsistency in statistical analyses.83 In contrast, variant-level analysis provides a consistent and well-defined unit of comparison across samples. Interestingly, in conventional mQTL analyses, we observed that SNPs located further away from CpG sites exhibit greater discrepancy between mQTL effect sizes and allele-specific methylation differences, compared to proximal SNP–CpG pairs. We hypothesize that this pattern is driven by two factors. First, long-range SNP–CpG associations are more likely to be influenced by LD, where the tested SNP may not be the causal variant, leading to allele switching and inconsistent effect directions.84 Second, mQTL estimates are often derived from single CpG measurements from array-based data, which are more susceptible to technical noise and site-specific variability, in contrast to allele-specific methylation that aggregates signals across reads and local CpG contexts.85

A central challenge in post-GWAS functional genomics is the identification of both the causal regulatory variant and the functional domain it controls, tasks that conventional association-based approaches struggle to resolve simultaneously due to LD confounding and the absence of direct functional readout 86. The two loci presented here illustrate two distinct but complementary utilities of variant-level ASM analysis in translational cancer epigenomics, highlighting it a unified framework to address both challenges. At the IRX4 locus, the ASM-DMR anchored by rs6885084 delineates an androgen-responsive enhancer domain, evidenced by its precise colocalization with cancer-specific AR ChIP-seq occupancy and H3K27ac signal that is further potentiated by DHT stimulation in CRPC cell lines. This is consistent with the well-established enrichment of sequencedependent ASM at active regulatory elements, including enhancers and transcription factor binding sites. 82,87 Importantly, the amplification of the allelic methylation difference in cancer tissue relative to normal tissue — a pattern documented more broadly across cancer types77 suggests that tumor-specific epigenetic remodeling does not create ASM de novo but rather potentiates a pre-existing constitutive cis-regulatory signal. At the PSCA locus, the analytical logic shifts from domain delineation to causal variant resolution. By examining allelic methylation differences, ASM analysis identifies rs4736369 as the likely causal cis-regulatory variant, a conclusion supported by the convergent evidence of allele-specific H3K4me3 occupancy, and allele-specific expression validated concordantly across nanopore long-read and Illumina short-read platforms. This exemplifies the use of ASM as a post-GWAS fine-mapping strategy, where the SNP anchoring the local allelic methylation difference is a cis-acting functional candidate — providing orthogonal, functionally grounded evidence that statistical fine-mapping methods based on LD alone cannot achieve. Taken together, the two loci establish that variant-level ASM analysis serves a dual purpose: it identifies the spatial boundaries of functional regulatory domains and resolves the identity of causal regulatory variants within them.

Supplementary Material

Supplement 1
media-1.xlsx (16.8KB, xlsx)
Supplement 2
media-2.docx (21.5KB, docx)
3

Acknowledgements

This study has been supported by the National Institutes of Health [R01CA250018 and R01CA212097, to L.Wang. and R01CA263494 to L.Wu]. The funders had no role in study design, data collection and analysis, publication decision, or manuscript preparation.

Footnotes

Conflicts of Interest Statement

L.Wu. provided consulting service to Pupil Bio Inc. and reviewed manuscripts for Gastroenterology Report, not related to this study, and received honorarium. No potential conflicts of interest were disclosed by other authors.

References

  • 1.Venter J.C., Adams M.D., Myers E.W., Li P.W., Mural R.J., Sutton G.G., Smith H.O., Yandell M., Evans C.A., Holt R.A., et al. (2001). The sequence of the human genome. Science 291, 1304–1351. 10.1126/science.1058040. [DOI] [PubMed] [Google Scholar]
  • 2.Lander E.S., Linton L.M., Birren B., Nusbaum C., Zody M.C., Baldwin J., Devon K., Dewar K., Doyle M., FitzHugh W., et al. (2001). Initial sequencing and analysis of the human genome. Nature 409, 860–921. 10.1038/35057062. [DOI] [PubMed] [Google Scholar]
  • 3.International Human Genome Sequencing, C. (2004). Finishing the euchromatic sequence of the human genome. Nature 431, 931–945. 10.1038/nature03001. [DOI] [PubMed] [Google Scholar]
  • 4.Sachidanandam R., Weissman D., Schmidt S.C., Kakol J.M., Stein L.D., Marth G., Sherry S., Mullikin J.C., Mortimore B.J., Willey D.L., et al. (2001). A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature 409, 928–933. 10.1038/35057149. [DOI] [PubMed] [Google Scholar]
  • 5.Genomes Project C., Auton A., Brooks L.D., Durbin R.M., Garrison E.P., Kang H.M., Korbel J.O., Marchini J.L., McCarthy S., McVean G.A., and Abecasis G.R. (2015). A global reference for human genetic variation. Nature 526, 68–74. 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Zhang F., and Lupski J.R. (2015). Non-coding genetic variants in human disease. Hum Mol Genet 24, R102–110. 10.1093/hmg/ddv259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ward L.D., and Kellis M. (2012). Interpreting noncoding genetic variation in complex traits and human disease. Nat Biotechnol 30, 1095–1106. 10.1038/nbt.2422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tian Y., Liu Q., Yu S., Chu Q., Chen Y., Wu K., and Wang L. (2020). NRF2-Driven KEAP1 Transcription in Human Lung Cancer. Mol Cancer Res 18, 1465–1476. 10.1158/1541-7786.MCR-20-0108. [DOI] [PubMed] [Google Scholar]
  • 9.Morris J.A., Caragine C., Daniloski Z., Domingo J., Barry T., Lu L., Davis K., Ziosi M., Glinos D.A., Hao S., et al. (2023). Discovery of target genes and pathways at GWAS loci by pooled single-cell CRISPR screens. Science 380, eadh7699. 10.1126/science.adh7699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Claussnitzer M., Cho J.H., Collins R., Cox N.J., Dermitzakis E.T., Hurles M.E., Kathiresan S., Kenny E.E., Lindgren C.M., MacArthur D.G., et al. (2020). A brief history of human disease genetics. Nature 577, 179–189. 10.1038/s41586-019-1879-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Consortium G.T., Laboratory D.A., Coordinating Center -Analysis Working, G., Statistical Methods groups-Analysis Working, G., Enhancing, G.g., Fund, N.I.H.C., Nih/Nci, Nih/Nhgri, Nih/Nimh, Nih/Nida, et al. (2017). Genetic effects on gene expression across human tissues. Nature 550, 204–213. 10.1038/nature24277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Consortium G.T. (2020). The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318–1330. 10.1126/science.aaz1776. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ahmed M., Soares F., Xia J.H., Yang Y., Li J., Guo H., Su P., Tian Y., Lee H.J., Wang M., et al. (2021). CRISPRi screens reveal a DNA methylation-mediated 3D genome dependent causal mechanism in prostate cancer. Nat Commun 12, 1781. 10.1038/s41467-021-21867-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tian Y., Soupir A., Liu Q., Wu L., Huang C.C., Park J.Y., and Wang L. (2022). Novel role of prostate cancer risk variant rs7247241 on PPP1R14A isoform transition through allelic TF binding and CpG methylation. Hum Mol Genet 31, 1610–1621. 10.1093/hmg/ddab347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tian Y., Dong D., Wang Z., Wu L., Park J.Y., consortium P., Wei G.H., and Wang L. (2023). Combined CRISPRi and proteomics screening reveal a cohesin-CTCF-bound allele contributing to increased expression of RUVBL1 and prostate cancer progression. Am J Hum Genet 110, 1289–1303. 10.1016/j.ajhg.2023.07.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Sheng X., Guan Y., Ma Z., Wu J., Liu H., Qiu C., Vitale S., Miao Z., Seasock M.J., Palmer M., et al. (2021). Mapping the genetic architecture of human traits to cell types in the kidney identifies mechanisms of disease and potential treatments. Nat Genet 53, 1322–1333. 10.1038/s41588-021-00909-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Oliva M., Demanelis K., Lu Y., Chernoff M., Jasmine F., Ahsan H., Kibriya M.G., Chen L.S., and Pierce B.L. (2023). DNA methylation QTL mapping across diverse human tissues provides molecular links between genetic variation and complex traits. Nat Genet 55, 112–122. 10.1038/s41588-022-01248-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Min J.L., Hemani G., Hannon E., Dekkers K.F., Castillo-Fernandez J., Luijk R., Carnero-Montoro E., Lawson D.J., Burrows K., Suderman M., et al. (2021). Genomic and phenotypic insights from an atlas of genetic effects on DNA methylation. Nat Genet 53, 1311–1321. 10.1038/s41588-021-00923-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Stefansson O.A., Sigurpalsdottir B.D., Rognvaldsson S., Halldorsson G.H., Juliusson K., Sveinbjornsson G., Gunnarsson B., Beyter D., Jonsson H., Gudjonsson S.A., et al. (2024). The correlation between CpG methylation and gene expression is driven by sequence variants. Nat Genet 56, 1624–1631. 10.1038/s41588-024-01851-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Meeks G.L., Henn B.M., and Gopalan S. (2023). Genetic differentiation at probe SNPs leads to spurious results in meQTL discovery. Commun Biol 6, 1295. 10.1038/s42003-023-05658-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Zhi D., Aslibekyan S., Irvin M.R., Claas S.A., Borecki I.B., Ordovas J.M., Absher D.M., and Arnett D.K. (2013). SNPs located at CpG sites modulate genome-epigenome interaction. Epigenetics 8, 802–806. 10.4161/epi.25501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Tian Y., McDonnell S.K., Wu L., Larson N.B., and Wang L. (2026). Fine mapping regulatory variants by characterizing native CpG methylation with nanopore long-read sequencing. HGG Adv 7, 100532. 10.1016/j.xhgg.2025.100532. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Skvortsova K., Zotenko E., Luu P.L., Gould C.M., Nair S.S., Clark S.J., and Stirzaker C. (2017). Comprehensive evaluation of genome-wide 5-hydroxymethylcytosine profiling approaches in human DNA. Epigenetics Chromatin 10, 16. 10.1186/s13072-017-0123-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wu X., and Zhang Y. (2017). TET-mediated active DNA demethylation: mechanism, function and beyond. Nat Rev Genet 18, 517–534. 10.1038/nrg.2017.33. [DOI] [PubMed] [Google Scholar]
  • 25.Smith Z.D., Hetzel S., and Meissner A. (2025). DNA methylation in mammalian development and disease. Nat Rev Genet 26, 7–30. 10.1038/s41576-024-00760-8. [DOI] [PubMed] [Google Scholar]
  • 26.Wu F., Li X., Looso M., Liu H., Ding D., Gunther S., Kuenne C., Liu S., Weissmann N., Boettger T., et al. (2023). Spurious transcription causing innate immune responses is prevented by 5-hydroxymethylcytosine. Nat Genet 55, 100–111. 10.1038/s41588-022-01252-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Thibodeau S.N., French A.J., McDonnell S.K., Cheville J., Middha S., Tillmans L., Riska S., Baheti S., Larson M.C., Fogarty Z., et al. (2015). Identification of candidate genes for prostate cancer-risk SNPs utilizing a normal prostate tissue eQTL data set. Nat Commun 6, 8653. 10.1038/ncomms9653. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Loyfer N., Magenheim J., Peretz A., Cann G., Bredno J., Klochendler A., Fox-Fisher I., Shabi-Porat S., Hecht M., Pelet T., et al. (2023). A DNA methylation atlas of normal human cell types. Nature 613, 355–364. 10.1038/s41586-022-05580-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Talevich E., Shain A.H., Botton T., and Bastian B.C. (2016). CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12, e1004873. 10.1371/journal.pcbi.1004873. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Juhling F., Kretzmer H., Bernhart S.H., Otto C., Stadler P.F., and Hoffmann S. (2016). metilene: fast and sensitive calling of differentially methylated regions from bisulfite sequencing data. Genome Res 26, 256–262. 10.1101/gr.196394.115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Shafin K., Pesout T., Chang P.C., Nattestad M., Kolesnikov A., Goel S., Baid G., Kolmogorov M., Eizenga J.M., Miga K.H., et al. (2021). Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat Methods 18, 1322–1332. 10.1038/s41592-021-01299-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Valentini V., Santi R., Silvestri V., Saieva C., Roviello G., Amorosi A., Comperat E., Ottini L., and Nesi G. (2025). CD44 Methylation Levels in Androgen-Deprived Prostate Cancer: A Putative Epigenetic Modulator of Tumor Progression. Int J Mol Sci 26. 10.3390/ijms26062516. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Majumdar K., Silva R., Perry A.S., Watson R.W., Rau A., Jaffrezic F., Murphy T.B., and Gormley I.C. (2024). A novel family of beta mixture models for the differential analysis of DNA methylation data: An application to prostate cancer. PLoS One 19, e0314014. 10.1371/journal.pone.0314014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang C., Xu X., Wang T., Lu Y., Lu Z., Wang T., and Pan Z. (2024). Clinical performance and utility of a noninvasive urine-based methylation biomarker: TWIST1/Vimentin to detect urothelial carcinoma of the bladder. Sci Rep 14, 7941. 10.1038/s41598-024-58586-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Detilleux D., Spill Y.G., Balaramane D., Weber M., and Bardet A.F. (2022). Pan-cancer predictions of transcription factors mediating aberrant DNA methylation. Epigenetics Chromatin 15, 10. 10.1186/s13072-022-00443-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Wu J., Wei Y., Li T., Lin L., Yang Z., and Ye L. (2023). DNA Methylation-Mediated Lowly Expressed AOX1 Promotes Cell Migration and Invasion of Prostate Cancer. Urol Int 107, 517–525. 10.1159/000522634. [DOI] [PubMed] [Google Scholar]
  • 37.Aldakheel F.M., Alnajran H., Alduraywish S.A., Mateen A., Alqahtani M.S., and Syed R. (2025). Analysing DNA methylation and transcriptomic signatures to predict prostate cancer recurrence risk. Discov Oncol 16, 110. 10.1007/s12672-025-01833-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wang H., Luo K., Tan L.Z., Ren B.G., Gu L.Q., Michalopoulos G., Luo J.H., and Yu Y.P. (2012). p53-induced gene 3 mediates cell death induced by glutathione peroxidase 3. J Biol Chem 287, 16890–16902. 10.1074/jbc.M111.322636. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Maldonado L., Brait M., Loyo M., Sullenberger L., Wang K., Peskoe S.B., Rosenbaum E., Howard R., Toubaji A., Albadine R., et al. (2014). GSTP1 promoter methylation is associated with recurrence in early stage prostate cancer. J Urol 192, 1542–1548. 10.1016/j.juro.2014.04.082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Kaminska K., Bialkowska A., Kowalewski J., Huang S., and Lewandowska M.A. (2019). Differential gene methylation patterns in cancerous and non-cancerous cells. Oncol Rep 42, 43–54. 10.3892/or.2019.7159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Zhang L., Zhang Q., Li L., Wang Z., Ying J., Fan Y., He Q., Lv T., Han W., Li J., et al. (2015). DLEC1, a 3p tumor suppressor, represses NF-kappaB signaling and is methylated in prostate cancer. J Mol Med (Berl) 93, 691–701. 10.1007/s00109-015-1255-5. [DOI] [PubMed] [Google Scholar]
  • 42.Zhang W., Shu P., Wang S., Song J., Liu K., Wang C., and Ran L. (2018). ZNF154 is a promising diagnosis biomarker and predicts biochemical recurrence in prostate cancer. Gene 675, 136–143. 10.1016/j.gene.2018.06.104. [DOI] [PubMed] [Google Scholar]
  • 43.Wilkinson E.J., Raspin K., Malley R.C., Donovan S., Nott L.M., Holloway A.F., and Dickinson J.L. (2024). WNT5A is a putative epi-driver of prostate cancer metastasis to the bone. Cancer Med 13, e70122. 10.1002/cam4.70122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wang Q., Williamson M., Bott S., Brookman-Amissah N., Freeman A., Nariculam J., Hubank M.J., Ahmed A., and Masters J.R. (2007). Hypomethylation of WNT5A, CRIP1 and S100P in prostate cancer. Oncogene 26, 6560–6565. 10.1038/sj.onc.1210472. [DOI] [PubMed] [Google Scholar]
  • 45.Adam S., Klingel V., Radde N.E., Bashtrykov P., and Jeltsch A. (2023). On the accuracy of the epigenetic copy machine: comprehensive specificity analysis of the DNMT1 DNA methyltransferase. Nucleic Acids Res 51, 6622–6633. 10.1093/nar/gkad465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Meir Z., Mukamel Z., Chomsky E., Lifshitz A., and Tanay A. (2020). Single-cell analysis of clonal maintenance of transcriptional and epigenetic states in cancer cells. Nat Genet 52, 709–718. 10.1038/s41588-020-0645-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Brocks D., Assenov Y., Minner S., Bogatyrova O., Simon R., Koop C., Oakes C., Zucknick M., Lipka D.B., Weischenfeldt J., et al. (2014). Intratumor DNA methylation heterogeneity reflects clonal evolution in aggressive prostate cancer. Cell Rep 8, 798–806. 10.1016/j.celrep.2014.06.053. [DOI] [PubMed] [Google Scholar]
  • 48.Dorff T.B., Blanchard M.S., Adkins L.N., Luebbert L., Leggett N., Shishido S.N., Macias A., Del Real M.M., Dhapola G., Egelston C., et al. (2024). PSCA-CAR T cell therapy in metastatic castration-resistant prostate cancer: a phase 1 trial. Nat Med 30, 1636–1644. 10.1038/s41591-024-02979-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Yegnasubramanian S., Haffner M.C., Zhang Y., Gurel B., Cornish T.C., Wu Z., Irizarry R.A., Morgan J., Hicks J., DeWeese T.L., et al. (2008). DNA hypomethylation arises later in prostate cancer progression than CpG island hypermethylation and contributes to metastatic tumor heterogeneity. Cancer Res 68, 8954–8967. 10.1158/0008-5472.CAN-07-6088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Li M., Paik H.I., Balch C., Kim Y., Li L., Huang T.H., Nephew K.P., and Kim S. (2008). Enriched transcription factor binding sites in hypermethylated gene promoters in drug resistant cancer cells. Bioinformatics 24, 1745–1748. 10.1093/bioinformatics/btn256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Cho N.Y., Kim J.H., Moon K.C., and Kang G.H. (2009). Genomic hypomethylation and CpG island hypermethylation in prostatic intraepithelial neoplasm. Virchows Arch 454, 17–23. 10.1007/s00428-008-0706-6. [DOI] [PubMed] [Google Scholar]
  • 52.Wong J., Tian Y., Patel M.S., Avasthi K., Hanson C., Larsen M., Ampaw E., Gutowski R., Fadlullah M.Z.H., Finkelstein J., et al. (2025). Plasma cell-free DNA methylation-based prognosis in metastatic castrate-resistant prostate cancer. NPJ Precis Oncol 10, 29. 10.1038/s41698-025-01232-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Kharat S.S., and Sharan S.K. (2025). 5-Hydroxymethylcytosine: a key epigenetic mark in cancer and chemotherapy response. Epigenetics Chromatin 18, 73. 10.1186/s13072-025-00636-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Fernandez A.F., Bayon G.F., Sierra M.I., Urdinguio R.G., Torano E.G., Garcia M.G., Carella A., Lopez V., Santamarina P., Perez R.F., et al. (2018). Loss of 5hmC identifies a new type of aberrant DNA hypermethylation in glioma. Hum Mol Genet 27, 3046–3059. 10.1093/hmg/ddy214. [DOI] [PubMed] [Google Scholar]
  • 55.Carella A., Tejedor J.R., Garcia M.G., Urdinguio R.G., Bayon G.F., Sierra M., Lopez V., Garcia-Torano E., Santamarina-Ojeda P., Perez R.F., et al. (2020). Epigenetic downregulation of TET3 reduces genome-wide 5hmC levels and promotes glioblastoma tumorigenesis. Int J Cancer 146, 373–387. 10.1002/ijc.32520. [DOI] [PubMed] [Google Scholar]
  • 56.Uslu C., Kapan E., and Lyakhovich A. (2024). Cancer resistance and metastasis are maintained through oxidative phosphorylation. Cancer Lett 587, 216705. 10.1016/j.canlet.2024.216705. [DOI] [PubMed] [Google Scholar]
  • 57.Du H., Xu T., Yu S., Wu S., and Zhang J. (2025). Mitochondrial metabolism and cancer therapeutic innovation. Signal Transduct Target Ther 10, 245. 10.1038/s41392-025-02311-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.van der Pol Y., Moldovan N., Ramaker J., Bootsma S., Lenos K.J., Vermeulen L., Sandhu S., Bahce I., Pegtel D.M., Wong S.Q., et al. (2023). The landscape of cell-free mitochondrial DNA in liquid biopsy for cancer detection. Genome Biol 24, 229. 10.1186/s13059-023-03074-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Mair R., Mouliere F., Smith C.G., Chandrananda D., Gale D., Marass F., Tsui D.W.Y., Massie C.E., Wright A.J., Watts C., et al. (2019). Measurement of Plasma Cell-Free Mitochondrial Tumor DNA Improves Detection of Glioblastoma in Patient-Derived Orthotopic Xenograft Models. Cancer Res 79, 220–230. 10.1158/0008-5472.CAN-18-0074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Vikramdeo K.S., Anand S., Sudan S.K., Pramanik P., Singh S., Godwin A.K., Singh A.P., and Dasgupta S. (2023). Profiling mitochondrial DNA mutations in tumors and circulating extracellular vesicles of triple-negative breast cancer patients for potential biomarker development. FASEB Bioadv 5, 412–426. 10.1096/fba.2023-00070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Fliss M.S., Usadel H., Caballero O.L., Wu L., Buta M.R., Eleff S.M., Jen J., and Sidransky D. (2000). Facile detection of mitochondrial DNA mutations in tumors and bodily fluids. Science 287, 2017–2019. 10.1126/science.287.5460.2017. [DOI] [PubMed] [Google Scholar]
  • 62.Ma Y., Du J., Chen M., Gao N., Wang S., Mi Z., Wei X., and Zhao J. (2023). Mitochondrial DNA methylation is a predictor of immunotherapy response and prognosis in breast cancer: scRNA-seq and bulk-seq data insights. Front Immunol 14, 1219652. 10.3389/fimmu.2023.1219652. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Wong J., Muralidhar R., Wang L., and Huang C.C. (2025). Epigenetic modifications of cfDNA in liquid biopsy for the cancer care continuum. Biomed J 48, 100718. 10.1016/j.bj.2024.100718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Lee D., Koo B., Yang J., and Kim S. (2023). Metheor: Ultrafast DNA methylation heterogeneity calculation from bisulfite read alignments. PLoS Comput Biol 19, e1010946. 10.1371/journal.pcbi.1010946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Fu Y., Timp W., and Sedlazeck F.J. (2025). Computational analysis of DNA methylation from long-read sequencing. Nat Rev Genet 26, 620–634. 10.1038/s41576-025-00822-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Yang H., Liu Y., Bai F., Zhang J.Y., Ma S.H., Liu J., Xu Z.D., Zhu H.G., Ling Z.Q., Ye D., et al. (2013). Tumor development is associated with decrease of TET gene expression and 5-methylcytosine hydroxylation. Oncogene 32, 663–669. 10.1038/onc.2012.67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Thomson J.P., Ottaviano R., Unterberger E.B., Lempiainen H., Muller A., Terranova R., Illingworth R.S., Webb S., Kerr A.R., Lyall M.J., et al. (2016). Loss of Tet1-Associated 5-Hydroxymethylcytosine Is Concomitant with Aberrant Promoter Hypermethylation in Liver Cancer. Cancer Res 76, 3097–3108. 10.1158/0008-5472.CAN-15-1910. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Kudo Y., Tateishi K., Yamamoto K., Yamamoto S., Asaoka Y., Ijichi H., Nagae G., Yoshida H., Aburatani H., and Koike K. (2012). Loss of 5-hydroxymethylcytosine is accompanied with malignant cellular transformation. Cancer Sci 103, 670–676. 10.1111/j.1349-7006.2012.02213.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Serandour A.A., Avner S., Oger F., Bizot M., Percevault F., Lucchetti-Miganeh C., Palierne G., Gheeraert C., Barloy-Hubler F., Peron C.L., et al. (2012). Dynamic hydroxymethylation of deoxyribonucleic acid marks differentiation-associated enhancers. Nucleic Acids Res 40, 8255–8265. 10.1093/nar/gks595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Hon G.C., Song C.X., Du T., Jin F., Selvaraj S., Lee A.Y., Yen C.A., Ye Z., Mao S.Q., Wang B.A., et al. (2014). 5mC oxidation by Tet2 modulates enhancer activity and timing of transcriptome reprogramming during differentiation. Mol Cell 56, 286–297. 10.1016/j.molcel.2014.08.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Villicana S., Castillo-Fernandez J., Hannon E., Christiansen C., Tsai P.C., Maddock J., Kuh D., Suderman M., Power C., Relton C., et al. (2023). Genetic impacts on DNA methylation help elucidate regulatory genomic processes. Genome Biol 24, 176. 10.1186/s13059-023-03011-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Smith A.K., Katrinli S., Maihofer A.X., Aiello A.E., Baker D.G., Boks M.P., Brick L.A., Chen C.Y., Dalvie S., Fani N., et al. (2025). Cell-type-specific and inflammatory DNA methylation patterns associated with PTSD. Brain Behav Immun 128, 540–548. 10.1016/j.bbi.2025.04.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Kassam I., Tan S., Gan F.F., Saw W.Y., Tan L.W., Moong D.K.N., Soong R., Teo Y.Y., and Loh M. (2021). Genome-wide identification of cis DNA methylation quantitative trait loci in three Southeast Asian Populations. Hum Mol Genet 30, 603–618. 10.1093/hmg/ddab038. [DOI] [PubMed] [Google Scholar]
  • 74.Do C., Shearer A., Suzuki M., Terry M.B., Gelernter J., Greally J.M., and Tycko B. (2017). Genetic-epigenetic interactions in cis: a major focus in the post-GWAS era. Genome Biol 18, 120. 10.1186/s13059-017-1250-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Zhao J., Zhao Z., Li H., Yu H., Lin H., Chen H., Li X., Liu D., Wang Y., and Wang G. (2025). CanASM: a comprehensive database for genome-wide allele-specific DNA methylation identification and annotation in cancer. BMC Genomics 26, 648. 10.1186/s12864-025-11849-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Hansen K.D., Timp W., Bravo H.C., Sabunciyan S., Langmead B., McDonald O.G., Wen B., Wu H., Liu Y., Diep D., et al. (2011). Increased methylation variation in epigenetic domains across cancer types. Nat Genet 43, 768–775. 10.1038/ng.865. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Do C., Dumont E.L.P., Salas M., Castano A., Mujahed H., Maldonado L., Singh A., DaSilva-Arnold S.C., Bhagat G., Lehman S., et al. (2020). Allele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs. Genome Biol 21, 153. 10.1186/s13059-020-02059-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Batra R.N., Lifshitz A., Vidakovic A.T., Chin S.F., Sati-Batra A., Sammut S.J., Provenzano E., Ali H.R., Dariush A., Bruna A., et al. (2021). DNA methylation landscapes of 1538 breast cancers reveal a replication-linked clock, epigenomic instability and cis-regulation. Nat Commun 12, 5406. 10.1038/s41467-021-25661-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Shoemaker R., Deng J., Wang W., and Zhang K. (2010). Allele-specific methylation is prevalent and is contributed by CpG-SNPs in the human genome. Genome Res 20, 883–889. 10.1101/gr.104695.109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Bell C.G., Gao F., Yuan W., Roos L., Acton R.J., Xia Y., Bell J., Ward K., Mangino M., Hysi P.G., et al. (2018). Obligatory and facilitative allelic variation in the DNA methylome within common disease-associated loci. Nat Commun 9, 8. 10.1038/s41467-017-01586-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Kerkel K., Spadola A., Yuan E., Kosek J., Jiang L., Hod E., Li K., Murty V.V., Schupf N., Vilain E., et al. (2008). Genomic surveys by methylation-sensitive SNP analysis identify sequence-dependent allele-specific DNA methylation. Nat Genet 40, 904–908. 10.1038/ng.174. [DOI] [PubMed] [Google Scholar]
  • 82.Tycko B. (2010). Allele-specific DNA methylation: beyond imprinting. Hum Mol Genet 19, R210–220. 10.1093/hmg/ddq376. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Abante J., Fang Y., Feinberg A.P., and Goutsias J. (2020). Detection of haplotype-dependent allele-specific DNA methylation in WGBS data. Nat Commun 11, 5238. 10.1038/s41467-020-19077-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Do C., Lang C.F., Lin J., Darbary H., Krupska I., Gaba A., Petukhova L., Vonsattel J.P., Gallagher M.P., Goland R.S., et al. (2016). Mechanisms and Disease Associations of Haplotype-Dependent Allele-Specific DNA Methylation. Am J Hum Genet 98, 934–955. 10.1016/j.ajhg.2016.03.027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Fan Y., Vilgalys T.P., Sun S., Peng Q., Tung J., and Zhou X. (2019). IMAGE: high-powered detection of genetic effects on DNA methylation using integrated methylation QTL mapping and allele-specific analysis. Genome Biol 20, 220. 10.1186/s13059-019-1813-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Schaid D.J., Chen W., and Larson N.B. (2018). From genome-wide associations to candidate causal variants by statistical fine-mapping. Nat Rev Genet 19, 491–504. 10.1038/s41576-018-0016-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Kreibich E., Kleinendorst R., Barzaghi G., Kaspar S., and Krebs A.R. (2023). Single-molecule footprinting identifies context-dependent regulation of enhancers by DNA methylation. Mol Cell 83, 787–802 e789. 10.1016/j.molcel.2023.01.017. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
media-1.xlsx (16.8KB, xlsx)
Supplement 2
media-2.docx (21.5KB, docx)
3

Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES