Skip to main content
BMC Medical Genomics logoLink to BMC Medical Genomics
. 2025 Jul 29;18:122. doi: 10.1186/s12920-025-02193-6

Cost-effective promoter methylation analysis via target long-read bisulfite sequencing: a case study in severe preterm birth

Silvana Pereyra 1,, Angela Sardina 1, Rita Neumann 2, Celia May 2, Rossana Sapiro 3, Bernardo Bertoni 1, Mónica Cappetta 1
PMCID: PMC12309061  PMID: 40731354

Abstract

Background

DNA methylation plays a critical role in the dynamics of gene expression regulation and the development of various disorders. Whole-genome bisulfite sequencing can provide single-base resolution of CpG methylation levels and is the “gold standard” for DNA methylation quantification, but its high cost limits its widespread application. In contrast, targeted sequencing provides an optimal, cost-effective solution when focusing on specific candidate regions while providing sufficient sequencing depth. Here, we present a targeted bisulfite sequencing approach in which nanopore sequencing is used to study the methylation status of regions of interest.

Methods

We applied this workflow to study the promoters of candidate genes associated with severe preterm delivery in a Latin American population. We amplified fragments greater than 1 kilobase in length from 12 genes via long PCR. Each sample was barcoded and pooled for sequencing in MinION flow cells.

Results

This approach achieves high sequencing depths, ensuring robust DNA methylation (DNAm) estimates. We detected significant hypomethylation of MIR155HG and hypermethylation of the ANKRD24 gene promoter in severe preterm birth samples, which is concordant with previously reported gene expression changes.

Conclusions

This approach represents a scalable and cost-effective method suitable for targeted promoter methylation profiling across several samples. Our study provides a proof-of-concept for larger studies, demonstrating the broad applicability and scalability of our assay to any locus of interest. These features render the method especially apt for clinical diagnostics and precision medicine.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12920-025-02193-6.

Keywords: Target sequencing, Preterm birth, DNA methylation, Gene promoter, long-read sequencing

Background

Recent studies highlight the critical role of epigenetic modifications, including DNA methylation, in regulating gene expression and contributing to the development of various diseases. DNA methylation (DNAm), the addition of a methyl group to the 5-prime position of a cytosine base in cytosine‒phosphate‒guanosine (CpG) dinucleotides, regulates gene expression by altering DNA accessibility to the transcriptional machinery in gene promoters and regulatory sequences. CpG islands are regions with a high density of CpGs, and methylation of CpG island promoters plays an essential role in gene regulation and transcriptional repression [1]. DNA methylation alters chromatin structure and affects other epigenetic marks, such as histone modifications. DNA methylation is also implicated in repressing transposable elements, playing a fundamental role in the maintenance of genomic stability [2]. DNA methylation is tightly regulated and affected by environmental factors; therefore, aberrant DNA methylation can lead to various human diseases [3]. In this context, the detection of differential methylation is crucial for understanding the causes of differential gene expression and, potentially, human disorders.

Several methods are available for quantifying DNA methylation. Mostly, DNA methylation analysis relies on microarrays that focus on mapping CpG-rich regions, gene promoters, and cis-regulatory elements. However, even though this method accounts for more than 97% of RefSeq data, it captures only approximately 3% of the DNA methylation sites in the human genome, with a median of 18 sites per gene [4]. This results in a substantial loss of valuable information, and it is therefore very likely that many of the precise locations where aberrant DNAm drives the onset of a particular disease are either not included in present analyses or are only somewhat indexed by proximal sites. The current ‘gold standard’ technique is whole-genome bisulfite sequencing, which is an alternative, more thorough method that includes all 28 million CpGs after bisulfite DNA treatment. Nonetheless, it is expensive, and it is often difficult to obtain the necessary depth for specific sites, making accurate quantification of DNA methylation difficult [5].

To overcome this limitation, targeted sequencing is optimal when a candidate region is defined. This approach ensures the necessary sequencing depth in specific regions of interest, providing a more reliable and comprehensive assessment of DNA methylation. This method offers the potential for cost reduction and includes various techniques, such as bait capture, nanopore Cas9-targeted sequencing (nCATS), adaptive sampling, and bisulfite-treated DNA amplification [68].

Reduced representation bisulfite sequencing (RRBS) [9] allows assessment of a fraction of CpG sites (1.5–2 million) at low cost but often results in uneven coverage and targets nonvariable regions [10]. Capture techniques (e.g., CpGiant, TruSeqEpic) use second-generation short-read technology to characterize bisulfite-treated targets [11]. Hybridization capture-based enrichment using biotinylated probes can be used to capture oligos that recover targeted regions, but these probes can be expensive [12]. Alternatively, nanopore sequencing can directly detect DNA methylation without the need for additional sample processing, as is required in sodium bisulfite procedures, and can examine DNA methylation patterns directly from genomic DNA [6]. However, as DNA methylation marks are not preserved during PCR amplification, only native genomic DNA can be analyzed, requiring an additional step to convert it into a targeted assay (i.e., nCATS). Nevertheless, nCATS can be an expensive and labor-intensive method, especially in undeveloped countries, as it requires high amounts of DNA (3–5 µg) [13] and does not achieve high enrichment [14].

Bisulfite sequencing remains the gold standard for methylation analysis. Bisulfite treatment results in highly fragmented DNA, which typically prevents amplification of fragments larger than 300–500 bp [15]. Larger size ranges, up to 1,500 bp, have been achieved but only by employing commercial kits [16]. Even so, long bisulfite-treated DNA amplification has the potential to be a low-cost approach, enabling the analysis of numerous regions and scalability for cost-effective examination of a large number of samples.

This approach is ideal for population studies on complex diseases, such as preterm birth (PTB). As with other complex diseases, PTB is regulated by changes in DNA methylation [17]. It is defined as the delivery of an infant before 37 weeks of gestation and is the result of the interaction of genetic and environmental components, representing a multifactorial syndrome [18]. PTB can be classified as severe (< 34 weeks of gestation) or moderate (34–37 weeks). Preterm birth can occur spontaneously or be induced by a physician through the induction of labor or cesarean section, which may be medically indicated. Despite advances in understanding the risk factors and mechanisms underlying preterm birth, the underlying basis of PTB is still not fully understood, primarily because of the complex interplay of genetic, environmental, and host factors. Therefore, it is important to develop multiple approaches that integrate genetic, transcriptomic, epigenetic, and epidemiological data to identify biomarkers that will enable advances in translational medicine.

As with other diseases, most genetic research on preterm birth has been conducted on people of European descent [19, 20]. The lack of ethnic diversity in studies of preterm birth means that published results are not necessarily generalizable to other populations, as distinct demographic histories can result in unique genetic variants associated with a given phenotype [21]. Thus, exploring genetic and epigenetic diversity in underrepresented populations is crucial for obtaining a comprehensive understanding of health and disease in diverse human populations. To this end, our group has made efforts in integrated genomics and epigenetics studies on complex diseases such as cancer and preterm birth, aiming to identify biomarkers in Latin American populations [2226].

In this study, we developed an analysis pipeline to investigate the methylation status of multiple gene promoters. We tested this workflow on candidate genes for spontaneous severe preterm birth in a Uruguayan population. We selected genes previously found to be differentially expressed in chorioamniotic tissue during preterm labor [26]. Our results demonstrate the use of a cost-effective approach to analyze DNA methylation in targeted regions.

Methods

Study population

The controls and cases were term and spontaneous severe preterm deliveries, respectively, from unrelated offspring of women receiving obstetrical care at the Centro Hospitalario Pereira Rossell, Montevideo, Uruguay, who were recruited in a previous study described in Pereyra et al. [26]. Briefly, severe preterm chorioamnion tissues were collected immediately after labor from pregnancies complicated by birth before 33 weeks of gestational age (GA). Term chorioamnion tissues were obtained from uncomplicated pregnancies delivered after 37 weeks of GA.

The ethics committee of the Facultad de Medicina of the Universidad de la República (Uruguay) and the ethics committee of the Centro Hospitalario Pereira Rossell (Uruguay) both approved the study protocol (Facultad de Medicina Board resolution N° 071140−000907−11). Written informed consent was obtained from mothers prior to the collection of biological material.

Maternal history of previous preterm births was collected through direct questioning of the mothers after delivery. Clinical and obstetric data were obtained from the Perinatal Information System, which consists of basic perinatal clinical records developed by the Latin American Center for Perinatology (CLAP) under WHO/PAHO [27]. A subset of epidemiological variables collected from the participants was examined to assess potential differences among groups. Gestational age, newborn weight, maternal age and previous gestations were analyzed as continuous variables and were assessed via Student’s t-test, whereas newborn sex, premature rupture of membranes (PROM), intrauterine growth restriction, preeclampsia, anemia, hypertension, and smoking were considered qualitative variables and were assessed via chi-square tests. All the statistical analyses were carried out via the R programming language (http://www.rproject.org/). Statistical significance was set at p < 0.05.

Bisulfite-treated DNA amplification

DNA was extracted from chorioamniotic membranes via standardized salting-out methods. Five hundred nanograms of DNA per sample were bisulfite-treated via the Zymo EZ−96 DNA methylation kit (Zymo Research, Orange, CA, USA). A bisulfite sequencing PCR approach was applied to the promoters of 12 candidate genes (Supplementary Table S1). Candidate genes were selected on the basis of the highest absolute logarithmic fold change values reported by Pereyra et al. [26].

For each gene, a promoter region was selected for amplification and analysis. When possible, a CpG island was included in the selected region. To optimize the amplification of bisulfite-treated DNA, long and nested PCR techniques were employed to amplify fragments approximately 1 kb in length. Oligonucleotides for gene promoters were designed in Methyl Primer Express™ Software v1.0 (Applied Biosystems, USA), following the methods of Wojdacz et al. [28]. Universal tail sequences provided by Oxford Nanopore Technologies (forward primer: 5’-TTTCTGTTGGTGCTGATATTGC−3’, reverse primer: 5’-ACTTGCCTGTCGCTCTATCTTC−3’) were added to the second round of PCR primers at the 5’ end. The designed oligonucleotides were checked on the BiSearch Web Server [15]. In total, the targeted regions cover ~ 10 kb.

Amplification was performed via an Applied Biosystems Veriti™ Thermal Cycler (Thermo Fisher Scientific, Waltham, MA, USA) with the following PCR conditions for the first PCR round: 1 cycle at 96 °C for 5 s, a gene-specific annealing temperature (Supplementary Table S1) for 1 min and 64 °C for 4 min, followed by 35 cycles at 95 °C for 20 s, a gene-specific annealing temperature for 30 s, and 64 °C for 2 min. The second round of PCR was incubated at 96 °C for 1 min, followed by 32 cycles of 96 °C for 20 s, gene-specific annealing for 30 s, and 64 °C for 1:30 min. First, PCR rounds were performed in 10 µl reaction volumes, containing 1X in-house PCR buffer [29], 12 mM Tris, 0.3 µM each forward and reverse gene primer, 1.2 units of HighTaq DNA polymerase (Bioron, Germany), 0.06 units of Velocity DNA polymerase (Bioline, UK) and 50 ng of template DNA. The second round of PCR was performed in a 20 µl volume, which contained 1X in-house PCR buffer, 12.5 mM Tris, 0.4 µM each forward and reverse gene primer, 2.4 units of HighTaq DNA polymerase (Bioron, Germany), 0.12 units of Velocity DNA polymerase (Bioline, UK) and 0.5 µL of the first round of PCR product as a template.

Nanopore sequencing and methylation calling

We sequenced 6 DNA chorioamniotic tissue samples from term newborns and 5 from severe preterm newborns in MinION flow cell to analyze 12 gene promoter regions. For nanopore sequencing, equal quantities of each amplicon per sample were pooled together. Libraries were prepared with a total of 300 ng of each pooled barcoded amplicon via a PCR Barcoding Kit (SQK-PBK004) according to the manufacturer’s instructions. All sample libraries were multiplexed on a MinION device via a MinION flow cell (FLO-MIN106, R9.4.1, Oxford Nanopore) via MinKnow v21.06.13 software. The samples were sequenced for a total of 20 h.

The raw FAST5 files produced by MinION were basecalled under high-accuracy mode via the ONT basecaller Guppy v5.0.16 (Oxford Nanopore Technologies). The raw sequences were preprocessed to remove barcodes in MinKnow. A minimum quality score cutoff of 7 was used for filtering low-quality reads, employing Nanofilt [30]. The trimmed files were then aligned to the human reference genome (GRCh38) using Bismark v0.16.3 [31], which was set to interpret only methylation in a ‘CG’ context, with the following parameters optimized for long sequenced reads: --local --score_min G,1,5.4. Reads were visualized with the Integrative Genomics Viewer (IGV) browser [32]. The aligned reads were further processed with Bismark Methylation Extractor (default parameters) to yield Bismark coverage files.

CpG site methylation analyses

All further analyses were carried out via the R programming language (http://www.r-project.org/). We filtered the CpG sites retaining those with more than 30 reads, as reported in the Bismark coverage files. To remove spurious reads from library preparation and potential mapping artifacts, we used the GenomicRanges R package [33] to filter reads on the basis of their genomic position for all selected genes. The Wilcoxon rank sum test was performed to identify differentially methylated CpG sites between the term and severe preterm groups, either individually or in groups. Specifically, for each gene, we identified a subset of CpGs common across samples, filtered by sequencing coverage. This approach was chosen to focus on gene-level trends, given the biological question of interest.

The p-values were adjusted for multiple testing using false discovery rate (FDR) estimation, with a threshold p value of < 0.05. For data exploration, we performed Pearson’s correlation analysis and principal component analysis on sites that covered all samples at 30× or greater via the Methylkit package [34] implemented in R. Cluster analysis of the differentially methylated CpG sites was performed via unsupervised hierarchical clustering with complete linkage and Euclidean distance as a measure of similarity between samples.

A single-linkage clustering algorithm was implemented in the R/Bioconductor package BiSeq [35] to detect differentially methylated regions (DMRs). The data were reduced to CpG sites covered in 50% of the samples, comprising at least 7 CpG sites close enough to each other (maximum 100 bp). To reduce bias due to outlier coverage, the coverage of CpG sites was limited to the 90% quantile. Differently methylated clusters were detected with a false discovery rate of 0.05. Within the differentially methylated clusters, differentially methylated CpG sites were detected with an FDR of 0.05. Differentially methylated CpG sites were grouped in a DMR when they were less than 100 bp away from each other.

Combined reduction dimension analysis

Multidimensional scaling (MDS) analysis was employed using all CpG sites that passed quality and coverage filtering, as a dimension reduction technique to explore the dataset for patterns in the methylation values for the 11 samples assayed with the R package edgeR [36]. Additionally, when possible, the methylation and expression levels within the same samples were compared. Since seven of the studied samples had previously undergone RNA sequencing analysis (Supplementary Table S2) (SRP139931; [26]), it was possible to combine methylation and RNA sequencing data into an MDS analysis for the 12 genes analyzed in the present study.

Results

A complete analysis pipeline was designed for the analysis of target regions. Bisulfite PCR primers were carefully designed for 12 targeted regions (Table 1, Supplementary Table S1), and we successfully applied this strategy to assess single-base methylation status with high sensitivity and increased coverage.

Table 1.

Chromosomal and relative positions of the amplified regions.

Gene Chromosomal position of amplified region Amplicon size (bp) Relative position to TSS CpG context Number of analyzed CpGs Mean Coverage
ACCS chr11:44,065,877–44,066,839 962 (−392) - (+570) Island 72 81.24
ANKRD24 chr19:4,198,696–4,199,647 951 (+607) - (+1558) shore 22 284.01
BIRC3 chr11:102,317,048–102,317,759 711 (−435) - (+276) Island 39 151.23
CXCL2 chr4:74,098,538–74,099,133 595 (+62) - (+657) Island 93 701.37
EGR3 chr8:22,694,093–22,695,046 953 (−2682) - (−1729) Island 65 61.68
GK chrX:30,652,839–30,653,338 499 (−583) - (−84) Island 54 1000.89
MAMDC2 chr9:70,043,447–70,044,211 764 (−133) - (+631) Island 79 1182.61
MIR155HG chr21:25,562,022–25,562,619 597 (−122) - (+475) Island 60 302.17
NRN1 chr6:6,003,310–6,004,084 774 (−59) - (+715) Island 138 612.79
PIK3AP1 chr10:96,720,099–96,720,913 814 (+415) - (−399) Island 65 68.73
STEAP1 chr7:90,153,483–90,154,614 1131 (−985) - (+146) Island 84 431.86
WNT1 chr12:48,976,387–48,977,557 1170 (−1934) - (−764) shore 17 87.14

As a proof of concept for our targeted long-read bisulfite sequencing approach, a case‒control sample of preterm birth was analyzed. A total of 6 chorioamniotic tissue samples from term newborns and 5 from severe preterm newborns were analyzed. A subset of the epidemiologic variables analyzed is shown in Supplementary Table S3. There were no significant differences among the groups in terms of sex, maternal age, number of previous gestations, PROM or anemia (p > 0.05). Preterm infants had significantly lower birth weights and earlier gestational ages at delivery. None of the mothers actively smoked cigarettes during pregnancy in this study, and the proportion of mothers who smoked passively was not significantly different between the PTB patients and the controls. No mothers had previously had preterm births, nor did they present with hypertension, preeclampsia, intrauterine growth restriction or hemorrhage during pregnancy.

The samples were barcoded, all amplicons were pooled together, and four Oxford Nanopore sequencing Minion runs were performed (Supplementary Table S2). After alignment and filtering reads with quality scores greater than 7, a total of 60,914 reads were mapped to targeted genes. After filtering per coverage (> 30 reads per site), a total of 788 unique CpG sites were available for analyses of the targeted genes. These CpG sites are distributed among promoter genes, as shown in Table 1. The average sequencing depth per CpG was 619x, whereas the mean coverage per sequencing run varied between 325x and 834x. Taken together, the mean coverage per site and gene ranged from 61x to 1182x (Table 1). Owing to different PCR efficiencies and the quality of some DNA samples from frozen tissues, sequencing data were not obtained for all the genes in all the samples.

To explore patterns in methylation status, we employed descriptive tools. Hierarchical clustering of sample methylation profiles, via either Pearson’s correlation distance or principal component analysis, did not group samples on the basis of their categories (Supplementary Figure S1). Additionally, an MDS analysis based on methylation data did not detect distinct patterns among term controls and severe preterm births, even when combined with transcriptome data (Supplementary Figure S2).

Our approach allowed us to determine the single-base methylation status. Methylation percentage levels (%5mC/5mC + C) per CpG per gene showed high variability across gene promoters (Supplementary Figure S3). Despite the low sample number, a DMR analysis was performed, and we detected one differentially methylated region between severe preterm births and controls in the CXCL2 gene, chr4:74,099,080–74,099,111, comprising 3 CpGs (FDR < 0.05). In addition, individual CpGs with potential differential methylation were identified across gene promoters; however, due to the small sample size, these CpGs did not reach statistical significance after multiple testing correction (FDR).

Genes with significant differences in methylation between groups were identified. Methylation values are presented as percentages of all CpGs grouped by each gene and condition. Specifically, significant differences in the methylation of 7 gene promoters were detected between cases and controls: ANKRD24, MIR155HG, CXCL2, NRN1, STEAP1, MAMDC2 and GK (Wilcoxon rank sum test, adjusted FDR p value < 0.05; Fig. 1, Supplementary Table S4). The changes in methylation observed here (Fig. 1A) were compared with the RNA expression levels and found to be concordant for the MIR155HG and ANKRD24 genes but not for the remaining analyzed genes (Fig. 1B). In the MIR155HG promoter, we detected hypomethylation in severe preterm births (Fig. 2), whereas mRNAs were previously reported to be overexpressed in this group (Fig. 1B). In contrast, the opposite pattern was observed for ANKRD24 [26].

Fig. 1.

Fig. 1

(A) DNA methylation level (%5mC/5mC + C) of grouped CpG sites per gene, which showed significant differences in chorioamniotic membranes between severe preterm and term patients. The boxes represent the interquartile ranges, and the lines across the boxes indicate the median values. Statistically significant differences between severe preterm and term patients were determined via the Wilcoxon rank sum test and adjusted by the false discovery rate (FDR) (p < 0.05). (B) Transcriptome counts and log2FC values of severe preterm births and term controls per gene, as reported previously [26]

Fig. 2.

Fig. 2

Methylation percentage per CpG site in term controls and severe preterm births for MIR155HG. Each row represents a different CpG site, while columns represent individual samples. Unsupervised hierarchical clustering with complete linkage and Euclidean distance was used to group samples on the basis of their methylation profiles. CpG sites are annotated relative to the transcription start sites (TSSs). Sequencing data were not obtained for all the samples

To assess whether combining methylation and gene expression data can discriminate between cases and controls, a combined reduction dimension analysis was conducted using the seven genes that exhibited significant methylation differences. MDS analysis of the methylation data of these genes did not clearly cluster samples on the basis of their condition (Supplementary Figure S4). However, severe preterm samples presented less variability and were more tightly grouped than term control samples were. Upon integrating both methylation and transcriptome data for samples where both datasets were available, this pattern persisted. In contrast, the transcriptome data analyzed for this gene set revealed a discordant pattern, with control samples clustering together and preterm samples showing greater variability.

Discussion

In this study, we report an accurate, robust, cost- and time-effective method to characterize DNA methylation in targeted regions via nanopore sequencing. PCR amplicons derived from bisulfite-treated DNA samples from severe preterm and term newborns were pooled and sequenced on MinION flow cells. Custom-designed primers were used to target selected regions of 12 candidate gene promoters.

The targeted bisulfite sequencing approach presented here holds great promise for its use in medical diagnostics, as it offers several advantages over alternative approaches. The main one is that it is far less expensive than whole-genome bisulfite sequencing because only the targeted region methylation data are generated, and downstream data analysis is much simpler.

This approach also provides advantages over other targeted sequencing approaches. Notably, our approach achieves exceptionally high coverage per gene, which allows for more accurate quantification and detection of methylation changes. We reached up to 1100x in total and up to 834x per run. Read depth has a crucial impact on the accuracy of DNAm estimates, particularly at sites with intermediate DNAm levels, which may be erroneously classified as methylated or unmethylated at low read depths [37]. Masser et al. [38] reported that a 1000x read depth is necessary for reliable methylation estimates. Our approach ensures sufficient depth at each site in all samples, ensuring reliable DNAm estimation. Our approach therefore performs better than nCATS assays, which often yield much lower coverages with considerable variance, ranging from means of 350x to as low as < 10x [3942].

Generally, because bisulfite-treated DNA has a high amount of uracil homopolymers, it is often difficult to amplify, yielding a typically low bisulfite PCR size range (~ 300–500 bp) [16]. However, the use of long PCR conditions and the subsequent use of nanopore technology allowed us to successfully amplify bisulfite-treated DNA fragments over 1 kb in length. On the other hand, sodium bisulfite degrades DNA through the pyrimidination of unmethylated cytosines. Consequently, whole-genome bisulfite sequencing exhibits low sequencing coverage in GC-rich regions [43]. However, our approach is capable of sequencing entire CpG islands with high coverage, making it an excellent method for analyzing regulatory regions of gene expression.

In addition, long amplicon sequencing by ONT gives a more thorough spatial and continuous view of DNA methylation patterns across entire promoter regions—including full-length CpG islands. Thus, it greatly enhances the resolution of methylation profiling while allowing more accurate allele-specific methylation analysis. Short read sequencing or EPIC array requires genotypic data and more complicated bioinformatics approaches to infer allele-specific patterns [44].

The EPIC methylation array is one of the most widely used technologies for DNA methylation studies; however, it interrogates significantly fewer CpG sites per promoter region compared to our method. Our method enables the tailored design of multiple overlapping amplicons per region, ensuring comprehensive CpG coverage across all the CpG islands and regulatory elements in and around the promoters of interest. The maximum number of CpG sites investigated by the EPIC array on the promoter regions examined in this study was 10, at the ACCS gene promoter—whereas other promoter regions had fewer probes or were not covered at all. This highlights a critical limitation of array-based approaches when focusing on specific regulatory regions of interest. In comparison, our approach provides more flexibility and higher resolution, which makes it better suited for detailed epigenetic studies of selected candidate loci.

The approach presented here allows the simultaneous analysis of the DNAm of multiple candidate regions in large cohorts. This method is highly customizable and scalable. The number of amplicons to be sequenced can be comfortably increased without compromising performance. Given the high coverage obtained, increasing the number of barcoded samples running on each flow cell is feasible, and the driving costs are even lower. Also, nanopore flow cells allow sequential runs of multiple sequencing libraries in the same flow cell by washing and reloading, further reducing per-sample sequencing costs. Importantly, in our experiment, each flow cell was underutilized, allowing users to potentially scale up the assay to optimize resources more efficiently. MinION is the sequencer with the lowest capital investment on the market and potentially the lowest cost per sample.

This makes our approach ideal for research laboratories analyzing DNAm with limited resources, a common scenario in geographic regions with underrepresented populations in genomic studies. Facilitating access to research methods for studying genetic and epigenetic variation in underrepresented populations is crucial; without data from diverse populations, our understanding of genomic influences on health would be incomplete, potentially exacerbating health disparities and inequities.

One key advantage of nanopore sequencing is the simultaneous interrogation of the native DNA molecule’s genome and epigenome. Nevertheless, for human whole-genome applications, PromethION flow cells are needed, as they have just enough throughput to cover the human genome at low depth (https://nanoporetech.com/resource-centre/workflow-human-variant-calling). This prevents multiplexing several samples on a single flow cell, and thus increases costs per sample.

Alternatively, adaptive sampling is a PCR-free method for analyzing DNAm directly from native DNA in targeted regions [45, 46]. Adaptive sampling does not require treating samples with sodium bisulfite or designing or running PCRs. However, it is currently not suitable for barcoding, increasing costs. For example, Payne et al. [47] achieved a maximum mean coverage of only 15x, while using 3 barcodes in a MinION flow cell.

In our study, our method permitted the analysis of DNA methylation directly in the target tissue, with a level of read depth and resolution not usually possible when other techniques are used. This high coverage enabled a fine-scale characterization of methylation patterns in key regions implicated in preterm birth. Epigenetic regulation has been shown to be a relevant mechanism for PTB in regulatory tissues [17], but no studies have been performed on chorioamniotic membranes to date. Differences in DNA methylation between term and preterm births have been reported in fetal amnios [48, 49], placenta [50], maternal blood at birth [51], and the umbilical cord and blood [52, 53]. To the best of our knowledge, this work is the first targeted long-read sequencing effort using Nanopore technology to study a Latin American population and the first to analyze DNAm data from chorioamniotic membranes for PTB.

Our approach allowed us to report methylation levels for samples of chorioamniotic membranes and contrast the methylation differences found with existing transcriptome data in the same samples. We identified two genes, MIR155HG and ANKRD24, whose direction of DNA methylation changes in the promoter is concordant with their reported gene expression changes. Notably, the methylation levels of MIR155HG CpGs were sufficient to discriminate between preterm and control samples, with the exception of one outlier sample (Fig. 2). MIR155HG, which encodes a precursor RNA of microRNA−155 (miRNA−155), plays a role in regulating inflammatory responses and is implicated in pathologies such as preeclampsia and glioma [52, 5456]. Aberrant miR−155−5p overexpression in maternal serum samples during gestation is one of the best predictive biomarkers of preeclampsia, even before the onset of symptoms [56]. These findings highlight the importance of epigenetic regulation in inflammatory pathologies such as PTB and cancer.

The lack of concordance for the remaining analyzed genes may be due to the presence of different gene regulatory mechanisms for DNAm, such as noncoding RNA regulation, histone modification or chromatin remodeling. In addition, our design only analyzed a limited promoter region per gene (approximately 1 kb, specifically targeting CpG islands), which may not necessarily represent the DNAm levels of the entire gene promoter, let al.one the enhancers associated with the gene. In the future, a larger region should be analyzed to detect changes, as DNAm may not be homogeneous along the promoter region [57].

Furthermore, one of the main limitations of our study is the small number of samples analyzed, inherent to the proof-of-concept design of this work. This limitation is clear in MDS results (Supplementary Figure S4) since methylation data from the seven differentially methylated genes do not clearly separate samples by condition. In order to draw robust biological conclusions from methylation patterns for our dataset, a larger number of high-quality samples should be included in future studies. Also, amplification results differed in terms of PCR efficiencies across regions, presumably because of varying GC content. This highlights the importance of DNA quality: for CpG islands spanning areas with very different GC content, excellent-quality genomic DNA is required.

The workflow outlined here to perform targeted long-read methylation sequencing is optimal for its use in medical diagnostics owing to its high sensitivity and coverage, which is also a cost-effective approach. It would be particularly useful as a tool to identify epigenetic biomarkers for complex diseases. High-resolution methylation data for specific genomic regions could help develop early diagnostic assays or design targeted personalized assays. For example, methylation patterns of important regulatory genes, such as MIR155HG, could indicate inflammatory pathologies, including PTB and preeclampsia, and thus enhance clinical outcomes through timely interventions.

Conclusions

This study serves as a proof-of-concept for the design of larger studies. We demonstrated this in a pilot study with a small sample size, which limited the ability to detect differential methylation between groups in this case but nevertheless shows the utility and versatility of this approach.

We demonstrate that combining bisulfite DNA treatment with pooled long-read sequencing is a scalable and cost-effective way to evaluate DNAm in several targeted regions and several samples in parallel. This cost-effectiveness, combined with the flexibility of amplicon design and the minimal laboratory infrastructure required, makes this approach a good method for translation research and clinical applications.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 1 (14.3KB, png)
Supplementary Material 2 (16.3KB, eps)
Supplementary Material 3 (790.1KB, eps)
Supplementary Material 4 (16.5KB, eps)
Supplementary Material 5 (380.3KB, docx)

Acknowledgements

The authors are indebted to the participating mothers, whose generosity and cooperation have made this study possible.

Abbreviations

bp

Base pair

CLAP

Latin American Center for Perinatology

CpG

Cytosine-phosphate-guanosine

DMR

Differentially methylated region

DNAm

DNA methylation

FDR

False discovery rate

GA

Gestational age

IGV

Integrative Genomics Viewer

kb

Kilobase

MDS

Multidimensional scaling

nCATS

Nanopore Cas9-targeted sequencing

PROM

Premature rupture of membranes

PTB

Preterm birth

RRBS

Reduced representation bisulfite sequencing

TSS

Transcription start site

Author contributions

SP and MC conceived and designed the project. SP, MC, AS, and BB performed the epigenetic experiments and analyzed and interpreted the experimental data. CM and RN contributed to the development of long PCR experimental protocols. RS and BB contributed to the statistical analysis and provided feedback on the manuscript, and SP and MC wrote the manuscript. All the authors read and approved the final manuscript.

Funding

This study was partly supported by Programa de Desarrollo de las Ciencias Básicas (PEDECIBA) and Universidad de la República (UDELAR) to SP, MC and BB. MC, BB and RS are supported by Sistema Nacional de Investigadores, Agencia Nacional de Investigación e Innovación (ANII, Uruguay).

Data availability

The dataset generated and analyzed during the current study is publicly available in the NCBI BioProject repository under the accession number PRJNA1081111 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1081111). Main R scripts are available at https://github.com/spereyra/cost-effective-promoter-methylation.git.

Declarations

Ethics approval and consent to participate

The study was conducted according to the guidelines of the Declaration of Helsinki and approved by both the Ethics Committee of the Facultad de Medicina of the Universidad de la República (Uruguay) and the Ethics Committee of the Centro Hospitalario Pereira Rossell (Uruguay) (Facultad de Medicina Board resolution N° 071140-000907‐11). Written informed consent to participate has been obtained from all the participants and the parents/legal guardians of minors.

Consent for publication

Not applicable.

Competing interests

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as potential conflicts of interest.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Goldberg AD, Allis CD, Bernstein E. Epigenetics: A landscape takes shape. Cell. 2007;128:635–8. [DOI] [PubMed] [Google Scholar]
  • 2.Greenberg MVC, Bourc’his D. The diverse roles of DNA methylation in mammalian development and disease. Nat Rev Mol Cell Biol. 2019;20:590–607. [DOI] [PubMed] [Google Scholar]
  • 3.Ortiz-Barahona V, Joshi RS, Esteller M. Use of DNA methylation profiling in translational oncology. Semin Cancer Biol. 2022;83:523–35. [DOI] [PubMed] [Google Scholar]
  • 4.Noguera-Castells A, García-Prieto CA, Álvarez-Errico D, Esteller M. Validation of the new EPIC DNA methylation microarray (900K EPIC v2) for high-throughput profiling of the human DNA methylome. Epigenetics. 2023;18:2185742. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Flynn R, Washer S, Jeffries AR, Andrayas A, Shireby G, Kumari M, et al. Evaluation of nanopore sequencing for epigenetic epidemiology: a comparison with DNA methylation microarrays. Hum Mol Genet. 2022;31:3181–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Gilpatrick T, Lee I, Graham JE, Raimondeau E, Bowen R, Heron A, et al. Targeted nanopore sequencing with Cas9-guided adapter ligation. Nat Biotechnol. 2020;38:433–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Loose M, Malla S, Stout M. Real-time selective sequencing using nanopore technology. Nat Methods. 2016;13:751–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Simpson JT, Workman RE, Zuzarte PC, David M, Dursi LJ, Timp W. Detecting DNA cytosine methylation using nanopore sequencing. Nat Methods. 2017;14:407–10. [DOI] [PubMed] [Google Scholar]
  • 9.Gu H, Smith ZD, Bock C, Boyle P, Gnirke A, Meissner A. Preparation of reduced representation bisulfite sequencing libraries for genome-scale DNA methylation profiling. Nat Protoc. 2011;6:468–81. [DOI] [PubMed] [Google Scholar]
  • 10.Morselli M, Farrell C, Rubbi L, Fehling HL, Henkhaus R, Pellegrini M. Targeted bisulfite sequencing for biomarker discovery. Methods. 2021;187:13–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Tanić M, Moghul I, Rodney S, Dhami P, Vaikkinen H, Ambrose J, et al. Comparison and imputation-aided integration of five commercial platforms for targeted DNA methylome analysis. Nat Biotechnol. 2022;40:1478–87. [DOI] [PubMed] [Google Scholar]
  • 12.Kacmarczyk TJ, Fall MP, Zhang X, Xin Y, Li Y, Alonso A, et al. Same difference: comprehensive evaluation of four DNA methylation measurement platforms. Epigenetics Chromatin. 2018;11:21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Gombert S, Jahn K, Pathak H, Burkert A, Schmidt G, Wiehlmann L, et al. Comparison of methylation estimates obtained via minion nanopore sequencing and Sanger bisulfite sequencing in the TRPA1 promoter region. BMC Med Genomics. 2023;16:257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Sun Z, Vaisvila R, Hussong L-M, Yan B, Baum C, Saleh L, et al. Nondestructive enzymatic deamination enables single-molecule long-read amplicon sequencing for the determination of 5-methylcytosine and 5-hydroxymethylcytosine at single-base resolution. Genome Res. 2021;31:291–300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tusnády GE, Simon I, Váradi A, Arányi T. BiSearch: primer-design and search tool for PCR on bisulfite-treated genomes. Nucleic Acids Res. 2005;33:e9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Yang Y, Sebra R, Pullman BS, Qiao W, Peter I, Desnick RJ, et al. Quantitative and multiplexed DNA methylation analysis using long-read single-molecule real-time bisulfite sequencing (SMRT-BS). BMC Genomics. 2015;16:350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Park B, Khanam R, Vinayachandran V, Baqui AH, London SJ, Biswal S. Epigenetic biomarkers and preterm birth. Environ Epigenetics. 2020;6:dvaa005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Romero R, Dey SK, Fisher SJ. Preterm labor: one syndrome, many causes. Science. 2014;345:760–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Fatumo S, Chikowore T, Choudhury A, Ayub M, Martin AR, Kuchenbaecker K. A roadmap to increase diversity in genomic studies. Nat Med. 2022;28:243–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sirugo G, Williams SM, Tishkoff SA. The missing diversity in human genetic studies. Cell. 2019;177:26–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Martin AR, Gignoux CR, Walters RK, Wojcik GL, Neale BM, Gravel S, et al. Human demographic history impacts genetic risk prediction across diverse populations. Am J Hum Genet. 2020;107:788–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Brignoni L, Cappetta M, Colistro V, Sans M, Artagaveytia N, Bonilla C, et al. Genomic diversity in sporadic breast Cancer in a Latin American population. Genes. 2020;11:1272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Cappetta M, Berdasco M, Hochmann J, Bonilla C, Sans M, Hidalgo PC, et al. Effect of genetic ancestry on leukocyte global DNA methylation in cancer patients. BMC Cancer. 2015;15:434. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cappetta M, Fernandez L, Brignoni L, Artagaveytia N, Bonilla C, López M, et al. Discovery of novel DNA methylation biomarkers for non-invasive sporadic breast cancer detection in the Latino population. Mol Oncol. 2021;15:473–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Pereyra S, Bertoni B, Sapiro R. Interactions between environmental factors and maternal-fetal genetic variations: strategies to elucidate risks of preterm birth. Eur J Obstet Gynecol Reprod Biol. 2016;202:20–5. [DOI] [PubMed] [Google Scholar]
  • 26.Pereyra S, Sosa C, Bertoni B, Sapiro R. Transcriptomic analysis of fetal membranes reveals pathways involved in preterm birth. BMC Med Genomics. 2019;12:53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.De Mucio B, Abalos E, Cuesta C, Carroli G, Serruya S, Giordano D, et al. Maternal near miss and predictive ability of potentially life-threatening conditions at selected maternity hospitals in Latin America. Reprod Health. 2016;13:134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Wojdacz TK, Dobrovic A, Hansen LL. Methylation-sensitive high-resolution melting. Nat Protoc. 2008;3:1903–8. [DOI] [PubMed] [Google Scholar]
  • 29.Kauppi L, May CA, Jeffreys AJ. Analysis of meiotic recombination products from human sperm. In: Keeney S, editor. Meiosis: volume 1, molecular and genetic methods. Totowa, NJ: Humana; 2009. pp. 323–55. [DOI] [PubMed] [Google Scholar]
  • 30.De Coster W, D’Hert S, Schultz DT, Cruts M, Van Broeckhoven C. NanoPack: visualizing and processing long-read sequencing data. Bioinformatics. 2018;34:2666–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Krueger F, Andrews SR. Bismark: a flexible aligner and methylation caller for Bisulfite-Seq applications. Bioinformatics. 2011;27:1571–2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Thorvaldsdóttir H, Robinson JT, Mesirov JP. Integrative genomics viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform. 2013;14:178–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Lawrence M, Huber W, Pagès H, Aboyoun P, Carlson M, Gentleman R, et al. Software for computing and annotating genomic ranges. PLoS Comput Biol. 2013;9:e1003118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Akalin A, Kormaksson M, Li S, Garrett-Bakelman FE, Figueroa ME, Melnick A, et al. MethylKit: a comprehensive R package for the analysis of genome-wide DNA methylation profiles. Genome Biol. 2012;13:R87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hebestreit K, Dugas M, Klein H-U. Detection of significantly differentially methylated regions in targeted bisulfite sequencing data. Bioinformatics. 2013;29:1647–53. [DOI] [PubMed] [Google Scholar]
  • 36.Robinson MD, McCarthy DJ, Smyth GK. EdgeR: a bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26:139–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Seiler Vellame D, Castanho I, Dahir A, Mill J, Hannon E. Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation. BMC Genomics. 2021;22:446. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Masser DR, Berg AS, Freeman WM. Focused, high accuracy 5-methylcytosine quantitation with base resolution by benchtop next-generation sequencing. Epigenetics Chromatin. 2013;6:33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Skowronek D, Pilz RA, Bonde L, Schamuhn OJ, Feldmann JL, Hoffjan S, et al. Cas9-Mediated nanopore sequencing enables precise characterization of structural variants in CCM genes. Int J Mol Sci. 2022;23:15639. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Wieting J, Jahn K, Bleich S, Frieling H, Deest M. A targeted long-read sequencing approach questions the association of OXTR methylation with high-functioning autism. Clin Epigenetics. 2023;15:195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Alfano M, De Antoni L, Centofanti F, Visconti VV, Maestri S, Degli Esposti C, et al. Characterization of full-length CNBP expanded alleles in myotonic dystrophy type 2 patients by Cas9-mediated enrichment and nanopore sequencing. eLife. 2022;11:e80229. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Iyer SV, Kramer M, Goodwin S, McCombie WR. ACME: an Affinity-based Cas9 mediated enrichment method for targeted nanopore sequencing. i oRxiv. 2022;:2022.02.03.478550. 10.1101/2022.02.03.478550
  • 43.Guanzon D, Ross JP, Ma C, Berry O, Liew YJ. Comparing methylation levels assayed in GC-rich regions with current and emerging methods. BMC Genomics. 2024;25:741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Montano C, Timp W. Evolution of genome-wide methylation profiling technologies. Genome Res. 2025;35:572–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Kovaka S, Fan Y, Ni B, Timp W, Schatz MC. Targeted nanopore sequencing by real-time mapping of Raw electrical signal with UNCALLED. Nat Biotechnol. 2021;39:431–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Payne A, Holmes N, Clarke T, Munro R, Debebe BJ, Loose M. Readfish enables targeted nanopore sequencing of gigabase-sized genomes. Nat Biotechnol. 2021;39:442–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Payne A, Munro R, Holmes N, Moore C, Carlile M, Loose M. Barcode aware adaptive sampling for GridION and promethion Oxford nanopore sequencers. bioRxiv 2022;:2021.12.01.470722. 10.1101/2021.12.01.470722
  • 48.Parets SE, Conneely KN, Kilaru V, Fortunato SJ, Syed TA, Saade G, et al. Fetal DNA methylation associates with early spontaneous preterm birth and gestational age. PLoS ONE. 2013;8:e67489. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Kim J, Pitlick MM, Christine PJ, Schaefer AR, Saleme C, Comas B, et al. Genome-wide analysis of DNA methylation in human Amnion. ScientificWorldJournal. 2013;2013:678156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Schoorlemmer J, Macías-Redondo S, Strunk M, Ramos-Ruíz R, Calvo P, Benito R, et al. Altered DNA methylation in human placenta after (suspected) preterm labor. Epigenomics. 2020;12:1769–82. [DOI] [PubMed] [Google Scholar]
  • 51.You Y-A, Kwon EJ, Hwang H-S, Choi S-J, Choi SK, Kim YJ. Elevated methylation of the vault RNA2–1 promoter in maternal blood is associated with preterm birth. BMC Genomics. 2021;22:528. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Wu Y, Lin X, Lim IY, Chen L, Teh AL, MacIsaac JL, et al. Analysis of two birth tissues provides new insights into the epigenetic landscape of neonates born preterm. Clin Epigenetics. 2019;11:26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Wang X-M, Tian F-Y, Fan L-J, Xie C-B, Niu Z-Z, Chen W-Q. Comparison of DNA methylation profiles associated with spontaneous preterm birth in placenta and cord blood. BMC Med Genomics. 2019;12:1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Mahesh G, Biswas R. MicroRNA–155: A master regulator of inflammation. J Interferon Cytokine Res. 2019;39:321–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Peng L, Chen Z, Chen Y, Wang X, Tang N. MIR155HG is a prognostic biomarker and associated with immune infiltration and immune checkpoint molecules expression in multiple cancers. Cancer Med. 2019;8:7161–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Srinivasan S, Treacy R, Herrero T, Olsen R, Leonardo TR, Zhang X, et al. Discovery and verification of extracellular MiRNA biomarkers for Non-invasive prediction of Pre-eclampsia in asymptomatic women. Cell Rep Med. 2020;1:100013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Irizarry RA, Ladd-Acosta C, Wen B, Wu Z, Montano C, Onyango P, et al. The human colon cancer methylome shows similar hypo- and hypermethylation at conserved tissue-specific CpG Island Shores. Nat Genet. 2009;41:178–86. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (14.3KB, png)
Supplementary Material 2 (16.3KB, eps)
Supplementary Material 3 (790.1KB, eps)
Supplementary Material 4 (16.5KB, eps)
Supplementary Material 5 (380.3KB, docx)

Data Availability Statement

The dataset generated and analyzed during the current study is publicly available in the NCBI BioProject repository under the accession number PRJNA1081111 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1081111). Main R scripts are available at https://github.com/spereyra/cost-effective-promoter-methylation.git.


Articles from BMC Medical Genomics are provided here courtesy of BMC

RESOURCES