Abstract
DNA methylation is an important biological process in epigenetics, and many methods have been developed to profile DNA methylation. Recent studies have employed Oxford Nanopore long-read sequencing for DNA methylation detection, presenting an alternative to the widely utilized Infinium arrays and short-read methods. In this study, we evaluate the performance of Nanopore sequencing in DNA methylation detection by comparing it to Illumina’s MethylationEPIC microarray (EPIC) and Enzymatic Methyl-Sequencing (EM-Seq). The initial comparison was between Nanopore and the EPIC array using blood samples (n = 4). Among the ~850,000 CpG sites covered by both methods, we observed high concordance (Pearson correlation coefficient, r ≥ 0.94 across all four samples). After downsampling Nanopore data from an average coverage of 26.4 reads per site to 10 reads per site, the correlation in CpG methylation remained high (r ≥ 0.93). When comparing Nanopore and EM-Seq using brain tissue samples (n = 4), lower correlation of CpG methylation (r: 0.79 – 0.88) was detected between Nanopore and EM-Seq, which can be attributed to biased and reduced coverage of hypomethylated CpG sites by EM-Seq. We also investigated additional features of Nanopore sequencing, such as native DNA sequencing that can differentiate 5mC and 5hmC, as well as haplotype phasing. Overall, the Nanopore platform exhibited a high degree of concordance with the EPIC array and provided more uniform genomic coverage than EM-Seq. This study provides insights for researchers in selecting appropriate DNA methylation detection methods, considering factors such as cost, DNA input, and the complexity of downstream analysis.
Keywords: Nanopore, EM-Seq, DNA methylation, long-read sequencing, 5mC/5hmC, haplotype phasing
Graphical Abstract

We compared DNA methylation profiles obtained from Nanopore sequencing with those from the MethylationEPIC array (blood, n = 4) and EM-Seq (caudate, n = 4), and observed high concordance along with unique features captured exclusively by Nanopore sequencing.
1. Introduction
DNA methylation, one of the most extensively studied epigenetic mechanisms, primarily occurs at CpG sites1,2. It plays a crucial role in regulating gene expression, genomic imprinting, aging, and cellular differentiation3–6. Various methods have been developed to detect DNA methylation, most of which rely on bisulfite treatment that converts unmethylated cytosines to uracil while leaving methylated cytosines unchanged. Illumina Methylation assay and bisulfite sequencing are the two most widely used for DNA methylation profiling using bisulfite treatment. Recently, some other approaches such as EM-seq and third-generation sequencing platforms7 are also widely used to detect cytosine modifications8,9.
Illumina’s MethylationEPIC array assesses DNA methylation at approximately 850,000 (EPIC v1) or 930,000 (EPIC v2) CpG sites. This high-throughput approach allows researchers to explore methylation changes across regulatory-centric genomic regions10. Although the EPIC array offers high throughput11, reproducibility12,13, and cost-effectiveness14, its limitations include predetermined CpG site selection10, and suboptimal performance in repetitive regions14.
Enzymatic Methyl-seq15 (EM-Seq) is a short-read sequencing approach capable of detecting cytosine modifications at single base-resolution using enzymatic reactions to selectively protect methylated cytosines while deaminating unmethylated ones, enhancing methylome profiling accuracy by reducing DNA damage and GC-bias often associated with bisulfite treatment14. However, as with other short-read technologies16, EM-Seq may face challenges in GC-rich regions14, which are crucial to understanding DNA methylation patterns.
Long-read sequencing technologies17, such as PacBio18 and Nanopore, are capable of generating reads exceeding 20 kilobases19,20, facilitating comprehensive methylation profiling across the genome. Nanopore sequencing, in particular, natively reads the DNA sequence without requiring21 amplification or chemical conversion, providing unique advantages for epigenomic studies. Although there are some studies that compared the Oxford Nanopore22 Technologies platform to other DNA methylation profile methods8,23, most of these studies focused on comparing to the EPIC microarray in blood at high coverage. This study benchmarks the Nanopore platform for DNA methylation detection, comparing its performance with the EPIC microarray and EM-Seq, in both blood and brain tissue. We evaluated the performance and cost-effectiveness of Nanopore sequencing at reduced read depths. We further investigated the unique features of Nanopore sequencing including 5-Hydroxymethylcytosine and haplotype-specific methylation. Using metrics such as CpG site coverage, CpG methylation correlation, and methodological attributes, we aimed to provide researchers with a robust framework for selecting the most suitable technology for their study design with consideration of cost, DNA input, and complexity.
2. Results
2.1. Nanopore Demonstrates High Methylation Concordance with EPIC Array
DNA methylation levels were assessed in blood samples from four participants using both the EPIC array and Nanopore sequencing platform (Supplementary Table 1). Nanopore sequencing yielded a mean read depth of 26.4x across samples, with individual coverages of 23.3x, 27.4x, 29.3x, and 25.4x. Concordance between CpG methylation levels obtained from Nanopore sequencing and the EPIC array was evaluated by examining the correlation of CpG methylation proportions (β value) across both platforms. Figure 1A illustrates high concordance (Pearson correlation coefficient r = 0.96) between the EPIC array and Nanopore sequencing for Sample C at the initial coverage of 29.3x. In the bimodal methylation distribution observed with the EPIC array, a slight inward shift of each peak is attributed to background fluorescence and the array’s regularization offset24.
Figure 1.

Heatmaps of concordance of methylation between Nanopore and MethylationEPIC for sample C. Sites analyzed had at least 10x Nanopore depth. The top and left density plots show the distribution of CpG methylation proportion. (A) Original data: 29.3x coverage (r = 0.96) (B) downsampled to 15x coverage (r = 0.95), (C) downsampled to 12.5x coverage (r = 0.94).
Table 1 shows the correlation for all samples at varying coverage levels. Within-platform correlations across different biological samples were highly robust for both Nanopore (r: 0.92–0.93) and the EPIC array (r: 0.96–0.99), indicating strong internal consistency for both technologies when processing blood samples (Supplementary Table 2). The overall categorical agreement across the four samples ranged from 0.698 to 0.812 (Supplementary Table 3). This agreement was further supported by Bland-Altman plots comparing methylation levels between the platforms (Supplementary Figure 1).
Table 1.
CpG sites detected by Nanopore with comparison to EPIC array at the full and downsampled coverage levels.
| Coverage Condition | Sample (Coverage) | CpG sites Detected (10x) | Genome Coverage (%) | CpG sites Overlapped with EPIC (%) | Correlation Coefficient |
|---|---|---|---|---|---|
| Original | A (23.3x) | 27,574,362 | 97.44 | 844,938 (98.24) | 0.97 |
| B (27.4x) | 28,196,708 | 99.64 | 855,425 (99.46) | 0.97 | |
| C (29.3x) | 27,916,885 | 98.65 | 851,477 (99.00) | 0.96 | |
| D (25.4x) | 28,134,525 | 99.42 | 857,148 (99.66) | 0.94 | |
| 15x Downsampling | A | 23,724,607 | 83.83 | 746,515 (86.80) | 0.96 |
| B | 24,553,466 | 86.76 | 772,847 (89.86) | 0.96 | |
| C | 24,327,396 | 85.96 | 761,980 (88.60) | 0.95 | |
| D | 25,091,744 | 88.66 | 788,615 (91.69) | 0.94 | |
| 12.5x Downsampling | A | 19,797,250 | 69.95 | 630,620 (73.32) | 0.96 |
| B | 21,177,884 | 74.83 | 675,144 (78.50) | 0.96 | |
| C | 20,736,998 | 73.28 | 655,383 (76.20) | 0.95 | |
| D | 21,822,914 | 77.11 | 697,948 (81.15) | 0.94 | |
| 10x Downsampling | A | 13,362,154 | 47.22 | 429,393 (49.93) | 0.96 |
| B | 15,107,112 | 53.38 | 487,728 (56.71) | 0.96 | |
| C | 14,464,996 | 51.11 | 460,406 (53.53) | 0.95 | |
| D | 15,816,119 | 55.89 | 515,879 (59.98) | 0.94 |
The cost of Nanopore sequencing is currently higher than that of both array-based platforms and short-read sequencing. To assess its potential cost-effectiveness, we evaluated the performance of Nanopore sequencing at lower read depths by downsampling reads to simulate reduced coverage. CpG methylation concordance between platforms remained robust under downsampling of Nanopore sequencing data, with Pearson correlation coefficient (r) consistently greater than 0.93 (mean = 0.96, 95% CI = (0.94, 0.98); Figure 1B–C, Table 1). However, fewer CpG sites in the EPIC array were covered in Nanopore sequencing with at least 10x coverage after downsampling. At 15x coverage, an average of 89.3% of CpG sites present on the EPIC array were retained in the Nanopore data, whereas at 12.5x average coverage, retention averaged 77.3%. The retention rate decreased to 50% when the simulated coverage was 10x (Table 1).
2.2. Nanopore Yields Superior Coverage Uniformity Compared to EM-Seq
Short-read sequencing has been criticized for biased coverage, particularly in GC-rich regions. We hypothesized that Nanopore sequencing may provide more uniform coverage across the genome, especially in GC-rich regions. Comparisons of DNA methylation profiles from Nanopore and EM-Seq were performed in brain tissue from four additional participants (Supplementary Table 1). Brain tissue was selected due to its high levels of 5hmC25, which can be used to investigate the unique features of Nanopore. The average sequencing coverage was 17.2x for EM-Seq (sample coverages: 17.2x, 16.3x, 17.0x, and 18.5x) and 18.2x for Nanopore sequencing (sample coverages: 17.9x, 18.1x, 18.3x, and 18.4x). Although the mean coverage is similar between Nanopore and EM-Seq, Nanopore sequencing demonstrated a higher coverage at CpG sites across samples (Figure 2A) and Nanopore consistently detected more CpG sites and at greater read depths (Supplementary Table 4). Nanopore sequencing yielded a more uniform coverage distribution, with a higher proportion of CpG sites achieving coverage levels closer to the mean.
Figure 2.

(A) Number of CpG sites covered with minimum read depth for each of the four samples for EM-Seq (Orange) and Nanopore (Blue). (B) Coverage distribution at CpGs for each platform across genomic regions for sample 1.
EM-Seq showed lower coverage relative to Nanopore in GC enriched regions linked to methylation, including exons, coding sequences (CDS), untranslated regions (UTRs), promoters, CpG islands, shores, and shelves (Figure 2B). This disparity was quantified via odds ratio analysis (Supplementary Figure 2). To further characterize sequencing biases, we plotted GC percentage against mean normalized coverage (Supplementary Figure 3), as well as GC percentage against mean methylation (Supplementary Figure 4).
Figure 3 presents comparative heatmaps depicting CpG methylation overlap between Nanopore and EM-Seq, demonstrating the concordance between the two platforms. At a minimal coverage of 10x, the mean Pearson’s correlation coefficient between Nanopore and EM-Seq was 0.82 (CI: 0.80–0.83).. Within-platform analysis revealed a distinct divergence: Nanopore maintained high consistency across different brain samples (r: 0.90–0.92), whereas EM-seq exhibited noticeably lower within-platform correlations (r: 0.74–0.76) (Supplementary Table 2). In the binned categorical analysis, the overall agreement across the four samples ranged from 0.842 to 0.875 (Supplementary Table 3).
Figure 3.

Heatmap of concordance of methylation between EM-Seq and Nanopore for sample 1. The top and left density plot shows the distribution of CpG methylation proportion in EM-Seq and Nanopore. (A) All CpG sites detected by both techniques (1x, r = 0.87). (B) Sites analyzed with at least 10x (r = 0.81).
To further evaluate the concordance between the two platforms at a locus-specific level, we examined the well-characterized brain-related gene BDNF (Supplementary Figure 5). This visualization confirms the overall agreement in methylation signals between the methods, while also highlighting the advantage of ONT sequencing in successfully deconvoluting 5mC and 5hmC signals that remain combined in standard EM-Seq data.
Uneven coverage across the genome in both datasets also introduces banding artifacts in the heatmap. Although the bimodal methylation distribution could be captured with EM-Seq at 1x coverage (Figure 3A), the unmethylated peak of the distribution was missing at 10x coverage, which reveals that EM-Seq struggles to capture CpG sites with low methylation levels at sufficient depth (Figure 3B).
Coverage and methylation profiles were further compared around transcription start sites (TSS) for EM-Seq and Nanopore sequencing (Figure 4). Nanopore sequencing achieves more uniform coverage than EM-Seq (Figure 4A). Both methods detect comparable CpG site numbers at 1x depth (ratio close to 1), but the number of CpGs detected by EM-Seq is less than Nanopore at 10x minimum coverage threshold especially around TSS regions. The ratio of CpG sites detected (EM-Seq / Nanopore) dropped from about 0.6 to 0.30 (Figure 4B). Figures 4C and 4D illustrate a coverage bias in EM-Seq. The coverage difference between CpGs with high methylation (β ≥ 0.9) and low methylation (β ≤ 0.1) was larger in EM-Seq (~5x, 50%) compared with Nanopore (~2x, 13%). When we compared coverage thresholds within each platform’s data, the β value is similar for Nanopore (Figure 4E) but different for EM-Seq (Figure 4F) as many CpGs with lower β value were filtered out because of low coverage. The unmethylated peak disappears under a 10x coverage threshold (Figure 3B).
Figure 4.

(A) Mean coverage of CpG sites ± 2 kb around transcription start sites (TSS) for Nanopore (blue) and EM-Seq (orange). (B) Ratio of CpG sites detected (EM-Seq / Nanopore) around TSS for CpGs with at least 10x (blue) and 1x (red) for both techniques. Mean coverage of CpG sites with high methylation (β ≥ 0.9) and low methylation (β ≤ 0.1) in the window ± 2kb around TSS for Nanopore (C) and EM-Seq (D). Mean β value at CpG sites around TSS with at least 10x (dashed line) and 1x (solid line) for Nanopore (E) and EM-Seq (F).
2.3. Expanded Methylation Profiling: Distinguishing 5mC/5hmC and Resolving Haplotypes
Compared to EM-Seq and the EPIC array, Nanopore sequencing has a few new features. Nanopore sequencing distinguishes between 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) without requiring supplementary assays (Figure 5A, 5B). In our study, EM-Seq and the EPIC array profiled total methylation (5mC + 5hmC). While 5hmC-specific profiling can be achieved with EM-Seq when combined with an additional oxidative bisulfite (oxBS) assay, this was not performed here and is not part of the default EM-Seq workflow. The overall distribution of total methylation appears similar between blood and brain tissue; notably, brain tissue exhibits a higher proportion of 5hmC compared to the blood samples (Supplementary Table 5). For both 5mC and 5hmC, modification levels in the same tissue were consistent across samples within the same genomic regions.
Figure 5.

β value distribution for (A) blood and (B) brain tissue, broken down by methylation type, 5mC(green), 5hmC (red), and total methylation (blue). The 5hmC level was higher in brain tissue, while the total methylation level was similar. Scatter plots show the β values of individual CpG sites for two distinct haplotypes (Haplotype 1 in gold, Haplotype 2 in blue). Each point represents the methylation level at a specific CpG locus. (C) Displays a fully haplotype specific methylation region, suggesting strong allelic methylation differences between haplotypes. (D) Example of partial haplotype-specific methylation, indicating variability between haplotypes.
Additionally, the long-read capabilities of Nanopore sequencing facilitate effective haplotype phasing using read-based methods, providing insights into haplotype-specific methylation patterns (Figure 5C, 5D). Across the eight samples, the mean number of haplotype-specific regions detected ranged from 6,964 to 14,672 (M = 11,149, SD = 2,816, Supplementary Table 6).
3. Discussion
3.1. Downsampling Reveals High Methylation Concordance Between Nanopore and EPIC Arrays at Lower Read Depths
The EPIC array offers targeted coverage of approximately 850,000 CpG sites (expanded to about 930,000 CpG sites in EPIC v2), focusing on regions of established biological significance. Due to its affordability, demonstrated reliability26, and analytical accessibility, the EPIC array remains a widely preferred platform for large-scale methylation studies27 and clinical applications28. However, the EPIC array's restricted genome coverage limits its suitability for comprehensive whole-genome investigations.
In our study, we observed a high concordance across nearly all CpGs covered by the EPIC array in the original dataset. The curvature in Figure 1 reflects differences in beta value calculations. EPIC derives continuous values from fluorescence intensities. The beta value rarely yields exact 0s or 1s (the two peaks in the density plot are a little off 0 and 1). In contrast, Nanopore relies on discrete read counts, yielding exact extremes of 0s and 1s. To assess Nanopore’s performance across different sequencing depths, we downsampled the read depth and examined its correlation with EPIC at varying mean coverage levels. We observed high concordance of CpG methylation between EPIC array and Nanopore sequencing even at very low mean coverage. At a mean coverage of 12.5x, more than 20 million CpGs were covered with at least 10 reads, and we retained 75% of the original CpGs detected. The downsampling analysis suggested that we can capture most EPIC assay sites with high consistency and with cost-effective low coverage with Nanopore sequencing. Although lower-depth Dorado basecalling risks individual 5mC/5hmC misclassifications, our downsampling analysis shows that aggregated population-level signals at 10x coverage remain highly concordant with EPIC arrays.
These results suggest that Nanopore sequencing can reliably capture methylation patterns within regulatory regions at reduced coverage, underscoring its potential for cost-effective methylation analysis in large-scale studies.
3.2. High Methylation Concordance but Divergent Coverage Profiles Between Nanopore and EM-Seq
In this study, four brain tissue samples were sequenced using both short-read (EM-Seq) and long-read (Nanopore) approaches, yielding an average coverage of 17.2x for EM-Seq and 18.2x for Nanopore. We observed a moderately high correlation in CpG methylation patterns between platforms (r ≥ 0.8), suggesting potential platform-specific biases. Although lower than the correlation observed between the EPIC array and Nanopore, the similarity between EM-Seq and Nanopore results is consistent with previous findings by Vaisvila et al.15, who reported similar correlation between EM-Seq replicates.
Although WGBS remains the reference standard for methylome profiling, we selected EM-Seq as our short-read comparator because it provides single-base resolution without bisulfite conversion: mitigating conversion-induced DNA damage, reducing GC bias, and supporting lower DNA input relative to WGBS15. If Nanopore sequencing were compared with WGBS, we would expect results similar to those observed with EM-Seq, but with improved coverage in GC-rich regions.
EM-Seq may exhibit two types of coverage biases. We observed lower read depth in GC-enriched regions (Figure 2B) and at CpG sites with low DNA methylation levels (Figure 4D). The first bias is associated with PCR bias and short-read sequencing29. However, the cause of the second bias remains unclear. This bias is also present in Nanopore sequencing, albeit to a far smaller degree (Figure 4C). The lower read depth at CpG sites with low methylation levels significantly impacts methylation measurements, particularly when applying a minimum read depth cutoff for EM-Seq. Many CpGs with low methylation levels may be filtered out due to insufficient coverage (Figures 3B, 4F). To address this issue, it is important to achieve a high read depth using EM-Seq. While Nanopore offers superior coverage uniformity and haplotype resolution, EM-Seq remains advantageous for its lower cost and higher throughput, making it well-suited for large-scale or resource-limited methylation studies.
3.3. Nanopore Sequencing Reveals Tissue-Specific 5hmC and Enables Haplotype-Phased Methylation
Nanopore sequencing of native DNA provides substantial advantages over conventional microarrays and short-read sequencing technologies, particularly for epigenomic research. One of the primary advantages of Nanopore technology is its capacity to detect multiple DNA methylation modifications, including 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) which cannot be discriminated in EPIC array or EM-Seq, directly from native DNA samples. While total methylation level was similar, 5hmC level is much higher in brain tissue than blood (Figure 5A–B), which plays an important role in neurodevelopmental and neurodegenerative disorders30. Beyond these, Nanopore sequencing can identify additional modifications31, including N6-methyladenine (6mA) and N4-methylcytosine (4mC), enabling more comprehensive epigenetic profiling of complex samples. DNA methylation plays a pivotal role in understanding intricate biological processes such as cancer progression32,33 and aging34, with recent studies highlighting the importance of 5hmC in carcinogenesis35 and R-loop formation36, which are associated with gene regulation and genome stability.
The long-read capability of Nanopore sequencing further enhances its applicability by enabling accurate haplotype phasing in parallel with methylation analysis37. Using read-based phasing tools38, maternal and paternal alleles can be distinguished by analyzing reads that contain multiple heterozygous variants, with longer reads capturing larger haplotype blocks. With phasing information, haplotype-specific methylation could be identified (Figure 5C–D), which helps us better understand biological features such as genomic imprinting and X-chromosome inactivation. Haplotype-specific methylation enables the identification of allele-specific methylation, which can help elucidate the influence of genetic variation on DNA modifications.
This study has several limitations. The sample size was small (n = 4 per comparison), limiting generalizability and statistical power. No technical replicates were included due to the high cost of sequencing; reproducibility for comparable protocols have been evaluated in prior studies16,39–41. Thus, the observed concordance and correlation may not be fully generalizable to larger, more heterogeneous populations. Although Nanopore sequencing distinguishes 5mC and 5hmC, this study did not explore additional modifications (e.g., 6mA, 4mC).
4. Conclusions
This study provides a comprehensive comparison of Oxford Nanopore Technologies’ long-read sequencing platform with Illumina’s MethylationEPIC array and Enzymatic Methyl-Sequencing (EM-Seq) for DNA methylation detection. Our findings indicate that Nanopore sequencing exhibits a high degree of concordance with the EPIC array (r ≥ 0.94) across nearly all shared CpG sites, even under simulated conditions of reduced sequencing coverage. These results suggest that Nanopore sequencing can reliably capture methylation patterns within key regulatory regions, potentially facilitating cost-effective analyses at lower read depths. When compared with EM-Seq, a moderate concordance (r: 0.79 – 0.88) was observed, with EM-Seq exhibiting limitations in accurately covering undermethylated regions and CpG sites located within GC-rich genomic regions. Given EM-Seq’s lower cost, increasing EM-Seq depth could enhance its performance, narrowing the methylation detection gap and broadening its epigenomic utility. Differences in input requirements, per-sample cost, computational burden, and achievable resolution across platforms are detailed in Supplementary Table 7. Nanopore sequencing demonstrates highconcordance with other technologies and introduces fewer biases associated with chemical or enzymatic conversion; however, it currently has higher costs and requires more high-quality DNA. Compared with the other two technologies, Nanopore sequencing also requires high-performance GPUs, increased storage capacity, and greater bioinformatics computational overhead. Ultimately, the ability of Nanopore sequencing to achieve high concordance with established methods while providing single-molecule resolution and direct methylation detection underscores its potential to reshape the landscape of DNA methylation analysis.
5. Materials & Methods
5.1. DNA Samples
For the EPIC and Nanopore comparison, DNA samples were isolated from human blood samples. Four donors (samples A, B, C, and D) were used for the EPIC microarray, and Nanopore native DNA long-read sequencing.
For the EM-Seq and Nanopore comparison, DNA samples were isolated from human post-mortem caudate tissue. Four donors (samples 1, 2, 3, and 4) were used for EM-Seq and Illumina short-read sequencing, and Nanopore native DNA long-read sequencing. Details of sample information can be found in Supplementary Table 1.
5.2. Enzymatic-Methyl Sequencing
After following the standard EM-Seq protocol, paired-end short-reads were sequenced using Illumina. Bismark42, with default parameters, handled mapping, reference alignment, and methylation calling.
5.3. Nanopore DNA Sequencing
DNA samples were quantified with both Qubit and TapeStation. Native long-read DNA libraries were prepared with the Native Barcoding Kit 24 V14 (SQK-NBD114.24) (Oxford Nanopore Technologies) following the manufacturer’s instructions. 1000 ng of total DNA per sample was used as input. Two samples were sequenced on one PromethION R10.4.1 flow cell (FLO-PRO114M). Sequencing was performed on a PromethION 24 for up to 96 hours. Flow cells were washed and reloaded twice during the run.
5.4. Nanopore Data Processing
Super accurate (SUP) model was performed with Dorado v0.7.2 Oxford Nanopore Technologies (https://github.com/nanoporetech/dorado) on the generated pod5 files with model dna_r10.4.1_e8.2_400bps_sup@v5.0.0 and modification models dna_r10.4.1_e8.2_400bps_sup@v5.0.0_5mC_5hmC@v1 and dna_r10.4.1_e8.2_400bps_sup@v5.0.0_6mA@v1 using default parameters. Computational processing was performed on the Big Red 200 supercomputer at Indiana University using an NVIDIA A100 GPU. The analysis of a single flow cell required approximately 40 hours of wall-clock time on a single compute node utilizing 32 cores and 64 GB of allocated memory (observed peak usage: ~46 GB). The basecalled unaligned BAM files generated from Dorado were next processed with the Nanopore wf-human-variation pipeline v2.0.0 (https://github.com/epi2me-labs/wf-human-variation). Whole-genome methylation pileup files were generated with Modkit (https://github.com/nanoporetech/modkit) using the default parameters.
5.5. Downstream analysis
The downstream analysis was conducted using R43 version 4.3.1. For the assessment of concordance between methylation proportions across methods, we calculated the Pearson correlation coefficient for each method's methylation proportion on the same sample, enabling a direct comparison of quantitative methylation measurements. For all cross-platform comparisons (Nanopore vs. EPIC and Nanopore vs. EM-Seq), we used total methylation (5mC + 5hmC) from Nanopore sequencing to ensure compatibility, as EM-Seq and the EPIC array jointly capture both 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) without distinction. Separate quantification of 5mC and 5hmC was performed only in platform-specific analyses (Figure 5A–B, Supplementary Table 5). Heatmaps and correlation tables were generated to visualize these associations across samples. To ensure uniform coverage across samples for unbiased comparisons, downsampling was applied. This was performed using the rbinom function in R, which randomly removed individual reads until reaching the target mean coverage, thus simulating a controlled reduction in read depth while preserving the original distribution of methylation proportions. Only the CpG sites with a minimal coverage of 10x were included in the downsampling analysis. To evaluate categorical agreement of DNA methylation detection across platforms, continuous beta values were binned into three states: Low (0–20%), Medium (20–80%), and High (80–100%). Position-matched loci were filtered for 10x coverage. Concordance was assessed using 3×3 confusion matrices for Nanopore vs. EPIC and Nanopore vs. EM-Seq (Supplementary Table 3). Overall agreement and categorical agreement were calculated to evaluate concordance within methylation bins.
Genomic attribute calculations were conducted using the GenomicRanges44 package (v1.54.1), with all genomic annotations aligned to the GRCh38 reference genome. Annotations for genes and transcripts were derived from GENCODE (UCSC) Release v47, focusing on protein-coding genes and their associated transcripts to streamline the analysis toward functionally relevant regions. Exon and intron boundaries were defined according to GENCODE v47 (GrCh38.p14), and adjustments to genomic coordinates ensured compatibility with GRCh38, preserving alignment accuracy across all samples. Haplotype specific methylation regions (HSMRs) between haplotypes of each sample were identified using the DSS45 package (v2.50.1). Only CpG sites with a read depth of at least 5 on each phased haplotype were utilized for the HSMR analysis. The HSMRs were computed based on a smoothed methylation proportion to mitigate noise, with p-values adjusted for multiple testing to control the false discovery rate. All visualizations were created with ggplot246 (v3.5.0) for scatter plots, bar plots, and other graphical summaries. The half-violin plots were created with gghalves47 (v0.1.4). ComplexHeatmap48 (v2.18) was used for complex heatmap visualizations that provide hierarchical clustering and detailed annotations of methylation patterns across regions and methods.
Supplementary Material
Highlights.
DNA methylation profiles measured by Nanopore sequencing show very high concordance with the MethylationEPIC array and strong agreement with EM-Seq, even at reduced sequencing coverage.
Nanopore sequencing provides more uniform genome-wide coverage than EM-Seq at comparable sequencing depth.
Unique Nanopore features, including 5mC/5hmC discrimination and haplotype phasing, offer extended biological insights beyond those achievable with short-read or array-based methods.
Funding
Support for this work was provided by the National Institutes of Health R25HG012325,5U10AA008401, and P30 AG010133.
Declaration of Competing Interest
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Dr. Saykin receives support from multiple NIH grants (P30 AG010133, P30 AG072976, R01 AG019771, R01 AG057739, U19 AG024904, R01 LM013463, R01 AG068193, R01 AG092591, T32 AG071444, U01 AG068057, U01 AG072177, U19 AG074879, as well as U24 AG074855). He has also received in-kind support from Avid Radiopharmaceuticals, a subsidiary of Eli Lilly (PET tracer precursor) and Gates Ventures, LLC and Sanofi (Proteomics panel assays on IADRC and KBASE participants as part of the Global Neurodegeneration Proteomics Consortium), gift funds (GV) supporting technical contributions to the GRIP platform, funding to IU by the Alzheimer’s Drug Discovery Foundation’s Diagnostics Accelerator (ADDF), and he has participated in Scientific Advisory Boards (Bayer Oncology, Bristol Myers Squibb, Eisai, Novo Nordisk, and Siemens Medical Solutions USA, Inc) and an Observational Study Monitoring Board (MESA, NIH NHLBI), as well as External Advisory Committees for multiple NIA grants. He also serves as Editor-in-Chief of Brain Imaging and Behavior, a Springer-Nature Journal.
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
CRediT authorship contribution statement
Steven Brooks: Writing – original draft, software, formal analysis. Melissa Robins: Data curation, validation. Hongyu Gao: Data curation. Xuhong Yu: Data curation. Kwangsik Nho: Data curation. Andrew Saykin: Data curation. Yunlong Liu: Writing – review & editing, funding acquisition. Gang Peng: Writing – review & editing, supervision, conceptualization.
All authors reviewed the manuscript for important scientific content and approved the final version.
Data Availability
The methylation data can be accessed at this repository: DOI: 10.5281/zenodo.17378987.
To protect sensitive human genotype information, raw POD5 files were not shared publicly, but we have provided all processed methylation data necessary to fully reproduce the analyses and results presented in this study
The software written to analyze and visualize the results can be accessed at this repository: https://github.com/PengBioinformaticsLab/EvaluationNanoporeDNAMethylation.
References
- 1.Mattei AL, Bailly N & Meissner A DNA methylation: a historical perspective. Trends Genet. 38, 676–707 (2022). [DOI] [PubMed] [Google Scholar]
- 2.Moore LD, Le T & Fan G DNA Methylation and Its Basic Function. Neuropsychopharmacology 38, 23–38 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Dhar GA, Saha S, Mitra P & Nag Chaudhuri R DNA methylation and regulation of gene expression: Guardian of our health. The Nucleus 64, 259–270 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Elhamamsy AR Role of DNA methylation in imprinting disorders: an updated review. J. Assist. Reprod. Genet 34, 549–562 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Jones MJ, Goodman SJ & Kobor MS DNA methylation and healthy human aging. Aging Cell 14, 924–932 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Suelves M, Carrió E, Núñez-Álvarez Y & Peinado MA DNA methylation dynamics in cellular commitment and differentiation. Brief. Funct. Genomics 15, 443–453 (2016). [DOI] [PubMed] [Google Scholar]
- 7.Searle B, Müller M, Carell T & Kellett A Third-Generation Sequencing of Epigenetic DNA. Angew. Chem 135, e202215704 (2023). [DOI] [PubMed] [Google Scholar]
- 8.Flynn R et al. Evaluation of nanopore sequencing for epigenetic epidemiology: a comparison with DNA methylation microarrays. Hum. Mol. Genet 31, 3181–3190 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Larson NB, Oberg AL, Adjei AA & Wang L A Clinician’s Guide to Bioinformatics for Next-Generation Sequencing. J. Thorac. Oncol. Off. Publ. Int. Assoc. Study Lung Cancer 18, 143–157 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Pidsley R et al. Critical evaluation of the Illumina MethylationEPIC BeadChip microarray for whole-genome DNA methylation profiling. Genome Biol. 17, 208 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Noguera-Castells A, García-Prieto CA, Álvarez-Errico D & Esteller M Validation of the new EPIC DNA methylation microarray (900K EPIC v2) for high-throughput profiling of the human DNA methylome. Epigenetics 18, 2185742 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Christiansen SN et al. Reproducibility of the Infinium methylationEPIC BeadChip assay using low DNA amounts. Epigenetics 17, 1636–1645 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zhang W et al. Critical evaluation of the reliability of DNA methylation probes on the Illumina MethylationEPIC v1.0 BeadChip microarrays. Epigenetics 19, 2333660 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Guanzon D, Ross JP, Ma C, Berry O & Liew YJ Comparing methylation levels assayed in GC-rich regions with current and emerging methods. 2023.09.06.556603 Preprint at 10.1101/2023.09.06.556603 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Vaisvila R et al. Enzymatic methyl sequencing detects DNA methylation at single-base resolution from picograms of DNA. Genome Res. 31, 1280–1289 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Foox J et al. The SEQC2 epigenomics quality control (EpiQC) study. Genome Biol. 22, 332 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.van Dijk EL et al. Genomics in the long-read sequencing era. Trends Genet. 39, 649–671 (2023). [DOI] [PubMed] [Google Scholar]
- 18.Eid J et al. Real-Time DNA Sequencing from Single Polymerase Molecules. Science 323, 133–138 (2009). [DOI] [PubMed] [Google Scholar]
- 19.Cook R et al. The long and short of it: benchmarking viromics using Illumina, Nanopore and PacBio sequencing technologies. Microb. Genomics 10, 001198 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hon T et al. Highly accurate long-read HiFi sequencing data for five complex genomes. Sci. Data 7, 399 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wang Y, Zhao Y, Bollas A, Wang Y & Au KF Nanopore sequencing technology, bioinformatics and applications. Nat. Biotechnol 39, 1348–1365 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Jain M, Olsen HE, Paten B & Akeson M The Oxford Nanopore MinION: delivery of nanopore sequencing to the genomics community. Genome Biol. 17, 239 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.de Abreu AR et al. Comparison of current methods for genome-wide DNA methylation profiling. Epigenetics Chromatin 18, 57 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Du P et al. Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysis. BMC Bioinformatics 11, 587 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Li W & Liu M Distribution of 5-hydroxymethylcytosine in different human tissues. J. Nucleic Acids 2011, 870726 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Harrison A & Parle-McDermott A DNA Methylation: A Timeline of Methods and Applications. Front. Genet 2, 74 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Schmidt M, Maié T, Dahl E, Costa IG & Wagner W Deconvolution of cellular subsets in human tissue based on targeted DNA methylation analysis at individual CpG sites. BMC Biol. 18, 178 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Galbraith K & Snuderl M DNA methylation as a diagnostic tool. Acta Neuropathol. Commun 10, 71 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Dohm JC, Lottaz C, Borodina T & Himmelbauer H Substantial biases in ultra-short read data sets from high-throughput DNA sequencing. Nucleic Acids Res. 36, e105 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Cheng Y, Bernstein A, Chen D & Jin P 5-Hydroxymethylcytosine: A new player in brain disorders? Exp. Neurol 268, 3–9 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Kong Y, Mead EA & Fang G Navigating the pitfalls of mapping DNA and RNA modifications. Nat. Rev. Genet 24, 363–381 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Ebrahimi V et al. Epigenetic modifications in gastric cancer: Focus on DNA methylation. Gene 742, 144577 (2020). [DOI] [PubMed] [Google Scholar]
- 33.Mangelinck A & Mann C Chapter One - DNA methylation and histone variants in aging and cancer. in International Review of Cell and Molecular Biology (eds. Weyemi U & Galluzzi L) vol. 364 1–110 (Academic Press, 2021). [DOI] [PubMed] [Google Scholar]
- 34.Jiang S & Guo Y Epigenetic Clock: DNA Methylation in Aging. Stem Cells Int. 2020, 1047896 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Zacapala-Gómez AE et al. TET enzymes and 5hmC epigenetic mark: new key players in carcinogenesis and progression in gynecological cancers. Eur. Rev. Med. Pharmacol. Sci 28, 1123–1134 (2024). [DOI] [PubMed] [Google Scholar]
- 36.Xu X et al. 5hmC modification regulates R-loop accumulation in response to stress. Front. Psychiatry 14, 1198502 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Kolmogorov M et al. Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation. Nat. Methods 20, 1483–1492 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Martin M et al. WhatsHap: fast and accurate read-based phasing. Preprint at 10.1101/085050 (2016). [DOI] [Google Scholar]
- 39.Kaur D et al. Comprehensive evaluation of the Infinium human MethylationEPIC v2 BeadChip. Epigenetics Commun. 3, 6 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Silva C, Machado M, Ferrão J, Sebastião Rodrigues A & Vieira L Whole human genome 5’-mC methylation analysis using long read nanopore sequencing. Epigenetics 17, 1961–1975 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Zhuang BC et al. Comparison of Infinium MethylationEPIC v2.0 to v1.0 for human population epigenetics: considerations for addressing EPIC version differences in DNA methylation-based tools. bioRxiv 2024.07.02.600461 (2024) doi: 10.1101/2024.07.02.600461. [DOI] [Google Scholar]
- 42.Krueger F & Andrews SR Bismark: a flexible aligner and methylation caller for Bisulfite-Seq applications. Bioinforma. Oxf. Engl 27, 1571–1572 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.R: The R Project for Statistical Computing. https://www.r-project.org/.
- 44.Lawrence M et al. Software for Computing and Annotating Genomic Ranges. PLOS Comput. Biol 9, e1003118 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Park Y & Wu H Differential methylation analysis for BS-seq data under general experimental design. Bioinformatics 32, 1446–1453 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Wickham H Ggplot2. (Springer International Publishing, Cham, 2016). doi: 10.1007/978-3-319-24277-4. [DOI] [Google Scholar]
- 47.Tiedemann F gghalves: Compose Half-Half Plots Using Your Favourite Geoms. (2022). [Google Scholar]
- 48.Gu Z, Eils R & Schlesner M Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847–2849 (2016). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The methylation data can be accessed at this repository: DOI: 10.5281/zenodo.17378987.
To protect sensitive human genotype information, raw POD5 files were not shared publicly, but we have provided all processed methylation data necessary to fully reproduce the analyses and results presented in this study
The software written to analyze and visualize the results can be accessed at this repository: https://github.com/PengBioinformaticsLab/EvaluationNanoporeDNAMethylation.
