Abstract
DNA methylation profiling by Illumina methylation array‐based methods has revolutionized the molecular classification and diagnosis of brain tumors. A significant barrier to adopting these methods in a clinical environment is the requirement for specialized scanners, which results in high additional costs and a larger laboratory footprint. DNA sequencing‐based alternatives are attractive because most clinical molecular pathology laboratories already use sequencers for other molecular assays. This study aimed to compare the utility of the newly developed sequencing‐based enzymatic methyl sequencing (EM‐seq) method paired with the Twist Human Methylome panel for brain tumor classification with standard Infinium Methylation BeadChip‐based methods. We used DNA from fresh‐frozen or formalin‐fixed, paraffin‐embedded (FFPE) brain cancer samples from 19 patients and 1 control sample to construct DNA libraries covering 3.98 million CpG sites. We developed and validated a bioinformatics pipeline to analyze target‐enriched EM‐seq (TEEM‐seq) data in comparison with standard array‐based methods for tumor classification and copy number profiling. We found high concordance between TEEM‐seq and traditional methods, with high correlation coefficients (>0.98) between FFPE replicates. We successfully classified tumor samples into the expected molecular classes with robust prediction scores (>0.82). We observed that FFPE samples required a sequencing depth of at least 35x to achieve consistently high and reliable prediction scores. The TEEM‐seq method has the potential to complement existing tumor classification methods and lower the barriers for the adoption of methylation profiling in routine clinical use.
Keywords: brain tumor, classification, EM‐seq, methylation sequencing, targeted enriched
Target‐enriched enzymatic methylation sequencing (TEEM‐seq), a DNA methylation sequencing method using the Twist Human Methylome panel, shows comparable performance to array‐based tumor methylation profiling. The method is compatible with both frozen and formalin‐fixed paraffin‐embedded (FFPE) samples and is a viable approach to tumor methylation and copy number profiling.

1. INTRODUCTION
DNA methylation profiling is an important method used to distinguish tumors of different molecular subtypes. Many human tumors are characterized by distinct DNA methylation signatures, which are thought to represent a combination of their cell of origin and their genomic driver abnormalities [1, 2]. These DNA methylation patterns have been used to resolve controversial histopathologic tumor entities, identify novel tumor types, and reveal molecular substructures within defined tumor types [3]. Supervised classification models have been introduced into the clinical environment to improve diagnostic accuracy and facilitate molecular subtyping [4, 5]. Comprehensive classification models have been established for human brain tumors and sarcomas [4, 5, 6]. In addition to improving diagnostic precision, supervised models have a significant advantage over conventional tumor classification approaches. Additional tumor examples or types can be accumulated over time and added to the training dataset to improve the scope and performance of models. Existing models are expected to expand through the addition of novel molecularly defined tumor types and be extended to include other major tumor groups [4, 5].
The most common platform for DNA methylation profiling uses Illumina Human Infinium BeadChip arrays (e.g., 450 K arrays or EPIC v1.0 BeadChip arrays). This platform shows comparable performance when used with DNA derived from frozen tissue or from formalin‐fixed, paraffin‐embedded (FFPE) tissue [7], is robust for use with fragmented nucleic acids, and shows high reproducibility with DNA input as low as 125 ng [8]. The high‐fidelity performance of platforms with FFPE‐derived DNA is critical to clinical adoption, as FFPE tissues are the most commonly used format for clinical samples. FFPE tissues can also be readily quality‐controlled for tumor content by the evaluation of matched hematoxylin and eosin‐stained sections [9, 10].
A significant barrier to the adoption of DNA methylation profiling for clinical classification purposes is the need for specialized instrumentation (e.g., iSCAN array reader, robotic liquid handlers, or Nextseq550 system), which are expensive and have a large laboratory footprint. These limitations particularly apply to under‐resourced centers, which may use alternative sequencers to perform various molecular tests and may lack the capacity or test volume to justify purchasing a dedicated array reader. Introducing alternative, flexible sequencing‐based methods to perform tumor DNA methylation profiling with existing laboratory instrumentation may facilitate more widespread adoption.
Genome‐wide DNA methylation can be captured by using alternative methods such as DNA methylation sequencing or reduced representation bisulfite sequencing (RRBS) [11, 12]. DNA methylation patterns from whole‐genome bisulfite sequencing (WGBS) show high fidelity with array‐based patterns [13, 14]. Unfortunately, using WGBS on FFPE material is challenging and the method is expensive on a per‐sample basis, making its routine adoption difficult. Recent methods utilizing nanopore sequencing for methylation profiling have shown promise [15, 16, 17] but lack the wide adoption of traditional sequencing platforms in the clinical environment and are challenging to implement for FFPE testing. RRBS may not be as amenable as other methods are for input into machine learning classification models because of the variability in the genomic regions captured [18]. WGBS and RRBS both utilize bisulfite conversion for methylation detection. Enzymatic methylation conversion offers several advantages over bisulfite conversion. First, enzymatic conversion is less harsh on DNA than bisulfite conversion, reducing fragmentation and improving genomic coverage. Second, this method enables distinction between 5mC and 5hmC through using different enzymes. Enzymatic methylation conversion is a new method that requires further optimization and validation. Capture‐based methylation sequencing platforms can hypothetically increase the sensitivity and specificity of detecting regions of interest.
In this study, we used enzymatic methylation sequencing combined with the Twist Human Methylome panel to perform target‐enriched enzymatic methylation sequencing (TEEM‐seq) on FFPE samples. We established bioinformatics pipelines for brain tumor classification and copy number profiling using TEEM‐seq data. We further validated the utility of these methods using diagnostic samples from patients with brain tumors, demonstrating comparable performance to data from Illumina Human Epic arrays.
2. MATERIALS AND METHODS
2.1. Sample selection
Samples from 19 patients with previously diagnosed brain tumors were selected for analysis. Of these, 16 tumors were entities presented in the reference data set used to train the classifier (positive tumor samples). In contrast, the other tumors were not presented within the reference data set and served as negative control samples. All tumors had previously been classified using Infinium MethylationEPIC v1.0 array. This study was approved by the St. Jude Institutional Review Board. Commercially available human DNA (G1471, Promega) was used as a normal control.
2.2. TEEM‐seq assay
The genomic DNA or FFPE‐derived DNA samples (DNA integrity number: 1.5–2.9) from 20 known diagnostic samples from patients with brain tumors were fragmented to an average insert size of 240–290 bp by a Covaris E220 focused‐ultrasonicator (Table S1). Five samples were processed with technical replicates (total n = 26). The library was constructed from the fragmented DNA using the NEBNext enzymatic methyl‐seq library kit (#103496, Twist Bioscience) and quantified with the Qubit 1xds DNA HS assay kit (#Q33230, Thermo Fisher Scientific). The library sizes were evaluated with a high‐sensitivity D1000 screen tape assay (5067‐5584, Agilent Technologies). An equal amount of DNA from eight libraries was pooled for target enrichment with the Twist Human Methylome panel (#105517, Twist Bioscience) by following the Twist targeted methylation sequencing protocol. The final enriched enzyme‐converted libraries were sent for 150 bp paired‐end sequencing on the NovaSeq6000 platform by the Hartwell Center at St. Jude Children's Research Hospital.
2.3. TEEM‐seq analysis pipeline
The TEEM‐seq analysis pipeline was developed and followed based on open‐source tools and packages (Figure S1). Briefly, the reference genome (hg38_noAltHla_UCSC.fa) was indexed and the dictionary file was generated using bwa v0.7.17, samtools v1.15.1, and Picard v2.21.2. Quality control was performed on pretrimmed reads by using fastqc v0.11.8 and multiqc v1.14. Illumina standard adapters were trimmed using Trim Galore v0.6.6 and cutadapt v1.2.1. Trimmed reads were then aligned to the reference genome using bwa‐meth v0.2.2, samtools v1.15.1, and sambamba v0.7.1. Duplicate reads were identified and tagged in the bam files by using Picard with the following parameters: CREATE_INDEX = false MAX_RECORDS_IN_RAM = 1000 SORTING_COLLECTION_SIZE_RATIO = 0.15 ASSUME_SORT_ORDER = coordinate OPTICAL_DUPLICATE_PIXEL_DISTANCE = 100. Performance metrics such as Fold‐80 base penalty, hybrid selection library size, percent duplicates, and percent off‐bait were generated by Picard HSMetrics. Additional performance metrics such as GC bias metrics, insert size metrics, and alignment summary metrics were also collected. CpG calls were generated by MethylDackel v0.5.1 with the following parameters: ‐minDepth 10 ‐maxVariantFrac 0.15. Downstream analysis was performed using R methylKit v1.22.0 and python Keras v2.12.0.
2.4. Unsupervised analysis
Unsupervised dimensionality reduction analysis of methylation data was performed using t‐distributed stochastic neighbor embedding (t‐SNE) with Rtsne (v.0.15). The t‐SNE analysis was used to visualize the separation of the 26 test samples against the 2801 reference tumor samples (GSE90496) [4]. In brief, principal components were calculated using the 32,000 most variably methylated probes, as measured by the standard deviation of the probe‐level beta values across samples. The same probes were used for principal component analysis. The number of statistically significant principal components was computed by agDimension function in the PCDimension package (v.1.1.11). Principal component scores for all statistically significant components (k = 14) were used for t‐SNE analysis [19]. The following non‐default parameters were used: theta = 0.1, pca = F, perplexity = 18, max_iter = 5000.
Hierarchical analysis was performed on replicates from different batches to evaluate the reproducibility of the EM‐seq data by using the dendextend R package (v1.17.1). The Euclidean distance was calculated among the samples with Ward.D2 as the agglomeration method. Principal component analysis was performed using the R factoextra package (v1.0.7) to examine the batch effects.
2.5. Correlation analysis
A correlation matrix for methylation profiles from control samples was developed to assess the concordance between TEEM‐seq and the Illumina EPIC array platform. The correlation matrix with correlation coefficients and the scatter plots with 95% confidence ellipses were produced using the R packages psych (v2.4.3) and methylkit (v1.28.0). The correlation hypothesis tests for all comparisons were adjusted for multiple comparisons using the Holm method.
2.6. Training and testing the brain tumor classification neural network
A deep neural network (DNN) model was built on 90% of 2801 brain tumor samples from GSE90496 [4] to validate the clinical utility of the EM‐seq methylation profiles. We extracted 325,139 intersecting CpGs from 4 platforms: Illumina Human 450 K, Illumina MethylationEPICv1.0, Illumina MethylationEPICv2.0 arrays, and the next‐generation EM‐seq platform. A custom model was trained with those intersecting probes and was tested on a 90:10 split for 250 epochs using the Stochastic Gradient Descent optimization method with a learning rate of 0.0001, a decay of 1e‐6, and a momentum of 0.9. This process was repeated five times by sampling the whole data sets to create 5 random training and hold‐out test sets. Classification results from the most optimized model were reported. The classification calls obtained from a previously deployed clinical neural network classification model [20] were compared to the outputs of the custom model on both methylation array and methylation sequencing data.
2.7. Down‐sampling analysis
To investigate the level of sequencing depth at which the classification results for FFPE samples remain reliable, we randomly selected the sequenced reads of these samples to 10%, 15%, and 25% of the original data. Two FFPE samples (S1 and S2) along with one FF sample (S3, control) were selected for the down‐sampling experiment. Sampling was done using the seqtk command. The down‐sampled files were then analyzed using the same pipeline.
2.8. Copy number variation analysis
DNA copy number variation (CNV) from high‐throughput DNA sequencing data was inferred and visualized using a Python library CNVkit v0.9.1. DNA CNV from the microarray methylation data was inferred by using the conumee Bioconductor package in R with the default settings [21]. The combined intensities of all available CpG probes were normalized against those of pediatric non‐neoplastic adrenal gland controls (n = 12) and using a linear‐regression approach. Copy number gains and losses were designated using a mean segment value threshold of −0.18 or 0.18 for copy number loss or gain, respectively.
3. RESULTS
3.1. TEEM‐seq data generated using FFPE tissues are highly reproducible across replicates and batches
We designed a pilot study with technical replicates to examine the variability of the TEEM‐seq data generated using the Twist Human Methylome panel on FFPE tissues. We performed experiments in multiple batches and used the same controls across experiments. In the first two batches, we used four known brain tumors and one genomic DNA control. Each sample had at least one technical replicate for a total of 15 samples. Hierarchical clustering analysis showed that all technical replicates were clustered together (Figure 1A). Principal component analysis showed that batch effects on these samples were minimal (Figure 1B). The controls from the three batches had a correlation coefficient of >0.97 (Figure 1C). Together, these results showed that there were minimal technical variations or batch effects within TEEM‐seq data generated from the Twist Human Methylome panel with FFPE tissues.
FIGURE 1.

Analysis of technical replicates and batch effects among the 15 TEEM‐seq samples. (A) Hierarchical clustering using Euclidean distance and Ward.D2 agglomeration. (B) Principal component analysis using the first 2 principal components. (C) A matrix of correlation coefficients and scatter plots of all the beta values obtained from the five control replicates. On the scatter plots, blue indicates uncorrelated regions, yellow indicates highly correlated regions, and green indicates variably correlated regions. The red lines represent linear regression fits. LOWESS, locally weighted scatterplot smoothing; TEEM‐seq, target‐enriched enzymatic methyl sequencing.
3.2. Methylation profiles measured by TEEM‐seq and Illumina microarray platforms are robust
We compared the methylation profiles from TEEM‐Seq and Illumina EPIC arrays to examine whether the methylation status remained consistent across platforms. In the third batch of experiments, we added 15 known tumors for a total of 20 individual tumors across the 3 batches. Figure 2 shows the correlation plot of the TEEM‐seq samples and their corresponding EPIC v1.0 profile. Each TEEM‐seq control profile had a correlation coefficient of at least 0.93 with the EPIC control profile (Figure 2). We also performed t‐SNE analysis to compare the test sample data set with the 2801 reference tumor data set generated from the Illumina HumanMethylation 450 K arrays (Figure 3A) [4]. The TEEM‐seq samples clustered with the 450 K platform samples according to their known methylation classes (Figure 3B).
FIGURE 2.

Pearson correlation of control sample methylation profiles obtained by TEEM‐seq or EPIC array. The 785,445 CpG sites shared between TEEM‐seq and EPIC v1.0 array were used for paired comparisons among five TEEM‐seq control samples and one EPIC v1.0 control sample. Correlation coefficients were calculated for each pair. Histograms of the beta values for each sample are depicted along the diagonal. The scatter plots for each pair of control samples are plotted with 95% confidence overlayed as a red ellipse. On the scatter plots, the ellipses representing the 95% confidence and the best‐fitting LOWESS regression line (red line) are shown without the individual data points. ***p < 0.0001. LOWESS, locally weighted scatterplot smoothing; TEEM‐seq, target‐enriched enzymatic methyl sequencing.
FIGURE 3.

Unsupervised analysis of the 26 samples with the brain tumor reference data set. (A) t‐distributed stochastic neighbor embedding (t‐SNE) analysis showing the 91 known methylation classes for the 2801 tumor reference data set (GSE90496). (B) t‐SNE analysis of the 26 samples compared to the reference data set (dimmed in gray). The samples are color‐coded according to the methylation class that they clustered with. Methylation class abbreviations are described in Table S2.
3.3. Development of a robust multi‐platform DNN model for brain tumor classification
To validate the clinical utility of the TEEM‐seq methylation profiles, we built a DNA methylation‐based DNN model based on the 2801 samples from the reference GSE90496 data set [4]. We trained the DNN model by using 325,139 probes that overlapped among Illumina Human 450K, MethylationEPIC v1.0, MethylationEPICv2.0 arrays, and TEEM‐seq platforms (Figure 4). The model achieved an average of 97% accuracy for the hold‐out test set after five samplings, with an average of 97% precision and recall (Table 1 and Figures S2 and S3).
FIGURE 4.

Architecture and validation scheme of the deep neural network brain tumor classifier. A deep neural network (DNN) model was built on 90% of 2801 brain‐tumor samples from GSE90496. The model was tested and trained on a 90:10 split for 250 epochs using the Stochastic Gradient Decent optimization method. Only 325,139 CpGs that overlapped among methylation profiles obtained from Illumina EPICv2.0, EPIC, 450 K, or TEEM‐seq were used to train the DNN. The hold‐out test sets included 280 reference tumor samples (10% of the 2801 samples) and 20 non‐replicate TEEM‐seq samples. TEEM‐seq, target‐enriched enzymatic methyl sequencing.
TABLE 1.
Results from five independent training and testing sets for validation of the deep neural network model.
| Sampling round | Precision | Recall | F1‐score | Accuracy (%) |
|---|---|---|---|---|
| #1 | 0.9672 | 0.9656 | 0.9618 | 96.56 |
| #2 | 0.9596 | 0.9562 | 0.9499 | 95.62 |
| #3 | 0.9639 | 0.9687 | 0.9648 | 96.87 |
| #4 | 0.9777 | 0.9781 | 0.9760 | 97.81 |
| #5 | 0.9843 | 0.9843 | 0.9820 | 98.43 |
We then validated the model on 20 independent tumors and achieved 100% accuracy as all predicted results matched their diagnostic labels (Figure 5A). We found that the classification scores obtained from the TEEM‐seq and Illumina EPIC array platforms were highly correlated (correlation coefficient, 0.99%; Figure 5B). The three negative control samples had prediction scores <0.85, whereas most other known tumor samples had prediction scores higher than 0.9 (Figure 5B). These results suggested that the methylation sequencing data obtained by applying the Twist Human Methylome panel on FFPE tissues provide sufficient and robust measurement features for brain tumor classification.
FIGURE 5.

Deep neural network performance on the 20 non‐replicate TEEM‐seq samples. (A) Confusion matrix showing the diagnosis and predicted results. The diagonal numbers represent the frequency of the methylation calls. A darker color indicates a higher number of samples in the specific methylation class. (B) Scatter plot showing the predicted scores for each sample using the TEEM‐seq (round) or the EPIC (triangle) platform. The colors represent the methylation classes according to the reference set. TEEM‐seq, target‐enriched enzymatic methyl sequencing.
3.4. Assessment of classification results with down‐sampling experiment
To determine the minimum sequencing depth needed for reliable classification of FFPE samples, we performed a down‐sampling experiment. Sequencing reads were randomly reduced to 10%, 15%, and 25% of the total for two FFPE samples (S1 and S2) and one fresh‐frozen (FF) sample (S3), which served as a control. The classifier maintained high prediction accuracy for the FF sample, with scores remaining at 0.99 even at 10% down‐sampling (equivalent to 30x sequencing depth) (Figure 6).
FIGURE 6.

Downsampling results from TEEM‐seq. Methylation class predicted score versus the sequencing coverage for S1 (red), S2 (green), and S3 (blue) after 10 (circle), 15 (triangle), 25 (square) percent downsampling of the original sequencing reads. FF: Fresh frozen; FFPE: Formalin‐fixed paraffin‐embedded tissue.
In contrast, the FFPE samples, which typically have lower DNA quality, showed reduced prediction scores at lower sequencing depths. At 25% down‐sampling, equivalent to sequencing depths of 32x (S1) and 38x (S2), prediction scores dropped from 0.91 and 0.96 to 0.67 and 0.88, respectively (Figure 6). This highlights the greater sensitivity of FFPE samples to lower sequencing coverage. These results suggested a minimum sequencing depth of 38x is essential for FFPE samples to achieve reliable and consistent classification results. This depth compensates for the challenges of degraded DNA in FFPE samples and ensures sufficient data quality for accurate predictions.
3.5. Copy number data concordance
Because genome‐wide copy number data can also be generated from array‐based DNA methylation profiles, we compared the consistency of CNV calls generated by the Illumina Human MethylationEPIC v1.0 array platform and the TEEM‐seq method. While both platforms showed high concordance in identifying patterns of copy number gains and losses (Figure 7A,B) and detected clinically significant copy number abnormalities, TEEM‐seq demonstrated distinct advantages. The enzymatic approach used in TEEM‐seq preserved DNA integrity, resulting in superior data quality and uniform genome‐wide coverage. This enabled more robust detection of copy number changes at the whole chromosome/arm level and focal regions. For example, while gains in chromosome 2 and losses in chromosome 9 (sample 2) were detected by both platforms (Figure 7C,D), the TEEM‐seq method offered enhanced resolution for focal copy number changes, such as MYCN amplifications, with greater accuracy and sensitivity (Figure 7E,F). These findings highlight the potential of using the TEEM‐seq method on FFPE samples to provide more comprehensive genomic insights than array‐based methods.
FIGURE 7.

Copy number variant profiles from TEEM‐seq and EPIC array. Copy number variant plots were produced using segmentation data generated from either (A) Illumina Human EPIC v1.0 arrays or (B) TEEM‐seq. (C, D) Zoom in on the broad and local changes on S2 using data from EPIC arrays and TEEM‐seq, respectively. Red, blue, and white bars represent copy gain (log2 copy ratio >0.18), copy loss (log2 copy ratio < −0.18), and copy neutral regions (log2 copy ratio between −0.18 and 0.18), respectively. (E, F) Local amplification of MYCN in chromosome 2. TEEM‐seq, target‐enriched enzymatic methyl sequencing.
4. DISCUSSION
This study evaluated the feasibility of using TEEM‐seq with the Twist Human Methylome panel on FFPE tissues to generate methylation data for diagnostic tumor classification. TEEM‐seq offers several theoretical advantages over traditional array‐based methods, including the preservation of more intact DNA, reducing GC bias, and providing more comprehensive genome‐wide coverage. These features make TEEM‐seq particularly well‐suited for working with degraded samples like FFPE tissues.
The primary goal was to expand access to DNA methylation profiling by implementing a method that is compatible with any laboratory equipped for next‐generation sequencing, eliminating the need for specialized array scanners. Our results demonstrated that TEEM‐seq produced reliable and reproducible methylation profiles that were highly concordant with Illumina EPIC array data. Using shared CpG sites, we developed a robust brain tumor classification model, which performed comparably when applied to data from either platform. These findings underscore TEEM‐seq's potential as a superior, versatile, and clinically viable alternative for methylation‐based tumor classification.
FFPE tissues are essential for clinical diagnostics because of their availability and ease of quality control through matched histology. Tumor methylation profiling in clinical settings is typically done on FFPE material using Illumina BeadChip arrays. Our study shows that TEEM‐seq with the Twist Human Methylome panel is a viable alternative with several advantages. It enables broader access for laboratories with sequencing capabilities and ensures compatibility with existing array‐based referent data. This eliminates the need to produce referent data for TEEM‐seq‐specific models and will facilitate the integration of rare tumor types to expand the referent data set. Consequently, TEEM‐seq has the potential to accelerate advancements in tumor classification and diagnostics by enabling the integration of methylation data across diverse platforms and research settings.
This work extends previous studies utilizing EM‐seq for DNA methylation profiling [22, 23, 24]. As the introduction of EM‐seq [22], several different implementations have been described. For instance, low‐passage whole genome EM‐seq has been deployed for cell‐free DNA testing [25]. Rubenstein et al. showed that TEEM‐seq could be performed on specific regions of targeted genes using DNA extracted from the blood of passerine birds [23]. The Twist Human Methylation panel has advantages for clinical DNA methylation profiling in that it covers most of the regions interrogated by the EPIC array and is compatible with FFPE testing; however, previous implementations have focused on unsupervised analysis and non‐tumor profiling [26]. Our current data supports these previous findings, showing a close correlation between TEEM‐seq‐derived beta values and array‐derived values. We also extend these findings to demonstrate the potential to apply this technology directly for robust clinical tumor classification [26].
A high level of reproducibility is critical for assays used in clinical and research applications. We observed that TEEM‐seq methylation data were consistent across technical replicates and across batches. We also found high concordance between the methylation profiles obtained by TEEM‐seq and those acquired through the established Illumina microarray platforms. The high correlation between these two platforms underscores the robustness of TEEM‐seq, reinforcing its credibility as an alternative to traditional methylation profiling methods. We used TEEM‐seq methylation profiles to accurately classify various brain tumor types. This finding suggests that performing TEEM‐seq on captured‐based methylation regions can inform personalized medicine, where accurate tumor classification is crucial for tailored treatment strategies.
A key advantage of our DNA methylation analysis approach is its cost‐effectiveness. Comprehensive studies that use traditional methods such as WGBS are often prohibitively expensive for many research and clinical laboratories [27]. In contrast, the TEEM‐seq and Twist Human Methylome panel method provides a more affordable alternative without compromising the depth or accuracy of the analysis. We estimate a slight increase in per sample cost for TEEM‐seq compared to methylation array testing (under $70 USD). Importantly, our approach avoids the requirement of additional expensive hardware (e.g., iScan system or specific sequencer Nextseq550) which is required to read traditional DNA methylation arrays.
This method could also be useful in situations where DNA quality or limits of detection may be challenging. For instance, recent reports have utilized methylation profiling on cell‐free DNA which is subject to some of the same quality problems as FFPE DNA. Although this was not the focus of the current study, the TEEM‐seq methods described here could be adopted for methylation profiling in cell‐free DNA. Compared to methods such as low‐pass whole genome EM‐seq, there could be theoretic gains in coverage at individual sites and a possible improved limit of detection in cfDNA, but that should be evaluated empirically.
Our method has several limitations compared to traditional array‐based approaches. First, because of the library preparation, there is a slight increase in hands‐on bench time (approximately a half day) compared to array testing. Also, uneven sequencing read distribution and site dropout, especially at lower sequencing coverage, can hinder tumor classification accuracy, with imputation performance decreasing as dropout increases. This issue is more pronounced with FFPE‐derived DNA, which showed a nearly threefold reduction in coverage compared to FF material and could lead to slightly higher sequencing costs for FFPE samples. To address this, separating FFPE and FF samples during multiplexing may help ensure sufficient coverage for FFPE. While direct nanopore sequencing has been effective for rapid methylation‐based classification, it performs best with FF material, though recent advances suggest the potential for FFPE [17, 28]. Integrating these advancements with TEEM‐seq could enhance CpG site coverage and reduce sequencing costs for FFPE samples.
5. CONCLUSION
In conclusion, our study demonstrates the feasibility of obtaining DNA methylation data from FFPE samples by performing TEEM‐seq with the Twist Human Methylome panel. Enzymatically converted DNA was more intact than DNA treated with sodium bisulfite, enabling less input DNA, reducing GC bias, providing longer sequencing reads, and more thorough genome‐wide coverage. This method sets a new standard in cost‐effective epigenetic analysis for brain tumor classification. Our approach has the potential for broader adoption in research and clinical diagnostics of brain tumors to enable personalized and effective healthcare solutions.
AUTHOR CONTRIBUTIONS
QTT developed the bioinformatic pipeline, analyzed data, interpreted results, and drafted the manuscript. SJ and RGT performed the experiments and collected the data. MZA developed the cross‐platform classifier. LW provided the DNA. LW and CGM edited the manuscript. BAO conceptualized the project, interpreted the results, and drafted the manuscript. All authors have read and edited the manuscript.
FUNDING INFORMATION
This work was partially supported by the National Cancer Institute support Grant (P30 CA021765) and the American Lebanese Syrian Associated Charities (ALSAC).
CONFLICT OF INTEREST STATEMENT
The authors declare no conflict of interest.
Supporting information
Figure S1. Bioinformatic analysis pipeline. The EM‐seq analysis pipeline was followed and developed based on open‐source tools and packages. Briefly, the hg38 reference genome was indexed and the dictionary file was generated using bwa, samtools, and Picard. Quality control was performed on pretrimmed reads by using fastqc and multiqc. Illumina standard adapters were trimmed using Trim Galore and cutadapt. Trimmed reads were then aligned to the reference genome using bwa‐meth, samtools, and sambamba. Duplicate reads were identified and tagged in the bam files by using Picard. Performance metrics such as Fold‐80 base penalty, HS library size, percent duplicates, and percent off‐bait were generated by Picard HSMetrics. Additional performance metrics such as GC bias metrics, insert size metrics, and alignment summary metrics were also collected. CpG calls were generated by MethylDackel.
Figure S2. Correlation matrix for all 15 replicate samples. A matrix of correlation coefficients and scatter plots of all the beta values obtained from all 15 replicates. On the scatter plots, blue indicates uncorrelated regions, yellow indicates highly correlated regions, and green indicates variably correlated regions. The red lines represent linear regression fits. LOWESS, locally weighted scatterplot smoothing.
Figure S3. Performance of the most optimized deep neural network model on the hold‐out test set. (A) The received operating characteristic curve of the model. (B) Confusion matrix of the predicted labels versus the true label matching frequency.
Table S1. Clinical and molecular features of cohort.
Table S2. Abbreviations for DNA methylation classes.
ACKNOWLEDGMENTS
We would like to thank the Hartwell Center for performing sequencing.
Tran QT, Jia S, Alom MZ, Wang L, Mullighan CG, Tatevossian RG, et al. Validation of target‐enriched enzymatic methylation sequencing for brain tumor classification from formalin‐fixed paraffin embedded‐derived DNA . Brain Pathology. 2025;35(5):e70000. 10.1111/bpa.70000
DATA AVAILABILITY STATEMENT
The data utilized in this study is available from the corresponding authors upon reasonable request for non‐commercial use.
REFERENCES
- 1. Moran S, Martinez‐Cardus A, Sayols S, Musulen E, Balana C, Estival‐Gonzalez A, et al. Epigenetic profiling to classify cancer of unknown primary: a multicentre, retrospective analysis. Lancet Oncol. 2016;17:1386–1395. 10.1016/S1470-2045(16)30297-2 [DOI] [PubMed] [Google Scholar]
- 2. Yang J, Wang Q, Zhang ZY, Long L, Ezhilarasan R, Karp JM, et al. DNA methylation‐based epigenetic signatures predict somatic genomic alterations in gliomas. Nat Commun. 2022;13:4410. 10.1038/s41467-022-31827-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Liu C, Tang H, Hu N, Li T. Methylomics and cancer: the current state of methylation profiling and marker development for clinical care. Cancer Cell Int. 2023;23:242. 10.1186/s12935-023-03074-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Capper D, Jones DTW, Sill M, Hovestadt V, Schrimpf D, Sturm D, et al. DNA methylation‐based classification of central nervous system tumours. Nature. 2018;555:469–474. 10.1038/nature26000 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Koelsche C, Schrimpf D, Stichel D, Sill M, Sahm F, Reuss DE, et al. Sarcoma classification by DNA methylation profiling. Nat Commun. 2021;12:498. 10.1038/s41467-020-20603-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Ferreyra Vega S, Olsson Bontell T, Corell A, Smits A, Jakola AS, Caren H. DNA methylation profiling for molecular classification of adult diffuse lower‐grade gliomas. Clin Epigenetics. 2021;13:102. 10.1186/s13148-021-01085-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Hovestadt V, Remke M, Kool M, Pietsch T, Northcott PA, Fischer R, et al. Robust molecular subgrouping and copy‐number profiling of medulloblastoma from small amounts of archival tumour material using high‐density DNA methylation arrays. Acta Neuropathol. 2013;125:913–916. 10.1007/s00401-013-1126-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Christiansen SN, Andersen JD, Kampmann ML, Liu J, Andersen MM, Tfelt‐Hansen J, et al. Reproducibility of the Infinium methylationEPIC BeadChip assay using low DNA amounts. Epigenetics. 2022;17:1636–1645. 10.1080/15592294.2022.2051861 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Heyn H, Esteller M. DNA methylation profiling in the clinic: applications and challenges. Nat Rev Genet. 2012;13:679–692. 10.1038/nrg3270 [DOI] [PubMed] [Google Scholar]
- 10. Siegel EM, Berglund AE, Riggs BM, Eschrich SA, Putney RM, Ajidahun AO, et al. Expanding epigenomics to archived FFPE tissues: an evaluation of DNA repair methodologies. Cancer Epidemiol Biomarkers Prev. 2014;23:2622–2631. 10.1158/1055-9965.EPI-14-0464 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Bonora G, Rubbi L, Morselli M, Ma F, Chronis C, Plath K, et al. DNA methylation estimation using methylation‐sensitive restriction enzyme bisulfite sequencing (MREBS). PLoS One. 2019;14:e0214368. 10.1371/journal.pone.0214368 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Liu Y, Han Y, Zhou L, Pan X, Sun X, Liu Y, et al. A comprehensive evaluation of computational tools to identify differential methylation regions using RRBS data. Genomics. 2020;112:4567–4576. 10.1016/j.ygeno.2020.07.032 [DOI] [PubMed] [Google Scholar]
- 13. Northcott PA, Buchhalter I, Morrissy AS, Hovestadt V, Weischenfeldt J, Ehrenberger T, et al. The whole‐genome landscape of medulloblastoma subtypes. Nature. 2017;547:311–317. 10.1038/nature22973 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Teh AL, Pan H, Lin X, Lim YI, Patro CP, Cheong CY, et al. Comparison of methyl‐capture sequencing vs. Infinium 450K methylation array for methylome analysis in clinical samples. Epigenetics. 2016;11:36–48. 10.1080/15592294.2015.1132136 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Drexler R, Brembach F, Sauvigny J, Ricklefs FL, Eckhardt A, Bode H, et al. Unclassifiable CNS tumors in DNA methylation‐based classification: clinical challenges and prognostic impact. Acta Neuropathol Commun. 2024;12:9. 10.1186/s40478-024-01728-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Kuschel LP, Hench J, Frank S, Hench IB, Girard E, Blanluet M, et al. Robust methylation‐based classification of brain tumours using nanopore sequencing. Neuropathol Appl Neurobiol. 2023;49:e12856. 10.1111/nan.12856 [DOI] [PubMed] [Google Scholar]
- 17. Vermeulen C, Pages‐Gallego M, Kester L, Kranendonk MEG, Wesseling P, Verburg N, et al. Ultra‐fast deep‐learned CNS tumour classification during surgery. Nature. 2023;622:842–849. 10.1038/s41586-023-06615-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Arreaza G, Qiu P, Pang L, Albright A, Hong LZ, Marton MJ, et al. Pre‐analytical considerations for successful next‐generation sequencing (NGS): challenges and opportunities for formalin‐fixed and paraffin‐embedded tumor tissue (FFPE) samples. Int J Mol Sci. 2016;17(9):1579. 10.3390/ijms17091579 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. van der Maaten L, Hinton G. Visualizing Data using t‐SNE. J Mach Learn Res. 2008;9:2579–2605. [Google Scholar]
- 20. Tran QT, Breuer A, Lin T, Tatevossian R, Allen SJ, Clay M, et al. Comparison of DNA methylation based classification models for precision diagnostics of central nervous system tumors. NPJ Precis Oncol. 2024;8:218. 10.1038/s41698-024-00718-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Hovestadt VZM. Conumee: enhanced copy‐number variation analysis using Illumina methylation arrays v.1.4.2 R package v.0.99.4. 2015.
- 22. Kim SS, Lee SC, Lim B, Shin SH, Kim MY, Kim SY, et al. DNA methylation biomarkers distinguishing early‐stage prostate cancer from benign prostatic hyperplasia. Prostate Int. 2023;11:113–121. 10.1016/j.prnil.2023.01.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Rubenstein DR, Solomon J. Target‐enriched enzymatic methyl sequencing: flexible, scalable and inexpensive hybridization capture for quantifying DNA methylation. PLoS One. 2023;18:e0282672. 10.1371/journal.pone.0282672 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Vaisvila R, Ponnaluri VKC, Sun Z, Langhorst BW, Saleh L, Guan S, et al. Enzymatic methyl sequencing detects DNA methylation at single‐base resolution from picograms of DNA. Genome Res. 2021;31:1280–1289. 10.1101/gr.266551.120 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. An Y, Zhao X, Zhang Z, Xia Z, Yang M, Ma L, et al. DNA methylation analysis explores the molecular basis of plasma cell‐free DNA fragmentation. Nat Commun. 2023;14:287. 10.1038/s41467-023-35959-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Hubert JN, Iannuccelli N, Cabau C, Jacomet E, Billon Y, Serre RF, et al. Detection of DNA methylation signatures through the lens of genomic imprinting. Sci Rep. 2024;14:1694. 10.1038/s41598-024-52114-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Hovestadt V, Jones DT, Picelli S, Wang W, Kool M, Northcott PA, et al. Decoding the regulatory landscape of medulloblastoma using DNA methylation sequencing. Nature. 2014;510:537–541. 10.1038/nature13268 [DOI] [PubMed] [Google Scholar]
- 28. Afflerbach AK, Albers A, Appelt A, Schweizer L, Paulus W, Bockmayr M, et al. Nanopore sequencing from formalin‐fixed paraffin‐embedded specimens for copy‐number profiling and methylation‐based CNS tumor classification. Acta Neuropathol. 2024;147:74. 10.1007/s00401-024-02731-z [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figure S1. Bioinformatic analysis pipeline. The EM‐seq analysis pipeline was followed and developed based on open‐source tools and packages. Briefly, the hg38 reference genome was indexed and the dictionary file was generated using bwa, samtools, and Picard. Quality control was performed on pretrimmed reads by using fastqc and multiqc. Illumina standard adapters were trimmed using Trim Galore and cutadapt. Trimmed reads were then aligned to the reference genome using bwa‐meth, samtools, and sambamba. Duplicate reads were identified and tagged in the bam files by using Picard. Performance metrics such as Fold‐80 base penalty, HS library size, percent duplicates, and percent off‐bait were generated by Picard HSMetrics. Additional performance metrics such as GC bias metrics, insert size metrics, and alignment summary metrics were also collected. CpG calls were generated by MethylDackel.
Figure S2. Correlation matrix for all 15 replicate samples. A matrix of correlation coefficients and scatter plots of all the beta values obtained from all 15 replicates. On the scatter plots, blue indicates uncorrelated regions, yellow indicates highly correlated regions, and green indicates variably correlated regions. The red lines represent linear regression fits. LOWESS, locally weighted scatterplot smoothing.
Figure S3. Performance of the most optimized deep neural network model on the hold‐out test set. (A) The received operating characteristic curve of the model. (B) Confusion matrix of the predicted labels versus the true label matching frequency.
Table S1. Clinical and molecular features of cohort.
Table S2. Abbreviations for DNA methylation classes.
Data Availability Statement
The data utilized in this study is available from the corresponding authors upon reasonable request for non‐commercial use.
