Skip to main content
RNA Biology logoLink to RNA Biology
. 2026 Feb 9;23(1):1–15. doi: 10.1080/15476286.2026.2627968

Advanced deep learning strategies in nanopore RNA sequencing

Crystal Ling a,#, Benjamin Lebeau a,#, Kwoh Chee Keong b,✉, Melissa Fullwood a,✉
PMCID: PMC12915787  PMID: 41663212

ABSTRACT

The epitranscriptome comprises chemical modifications found on RNA molecules that play essential roles in co- and post-transcriptional gene regulation. Dysregulation of these modifications has been implicated in various diseases, fuelling interest in evaluating them as emerging biomarkers and therapeutic targets. Nanopore direct RNA sequencing provides a powerful platform for profiling diverse RNA modifications at single-molecule resolution, but the complexity of the signals requires advanced computational approaches for interpretation. Artificial intelligence, particularly deep learning (DL), has become central to this effort. While classical DL architectures such as convolutional and recurrent neural networks have been widely applied, more recent approaches employ specialized learning frameworks and ensemble strategies to address challenges of data scarcity, noise, and biological variability while providing higher resolution output. In this review, we summarize these developments and highlight future multidisciplinary opportunities at the intersection of artificial intelligence and biology for characterizing the epitranscriptome obtained with direct RNA nanopore sequencing.

KEYWORDS: Deep learning, nanopore sequencing, direct RNA sequencing, epitranscriptome, RNA modifications

1. Introduction

The epitranscriptome is a co- and post-transcriptional gene regulatory layer, which comprises chemical modifications that govern key aspects of RNA metabolism, including stability [1,2], splicing [3], export [4,5], localization [6], and translation [7]. More than 170 distinct chemical modifications have been identified across diverse RNA biotypes [8–10], with N6-methyladenosine (m6A), 5-methylcytosine (m5C), pseudouridine (Ψ) and adenine-to-inosine (A-to-I) editing among the most extensively studied. When dysregulated, these modifications impair cellular functions and contribute to aberrant transcriptional networks [2,7,11] and signalling pathways [12,13] that underlie disease development [14,15]. Distinct epitranscriptomic signatures have been identified across a range of diseases, including cancer [9,16–18], metabolic diseases [19], cardiovascular diseases [20,21], and psychiatric disorders [12], highlighting their emerging value as diagnostic biomarkers [22] and therapeutic targets [9,23,24].

Nanopore sequencing is a third-generation sequencing technology that has revolutionized the epitranscriptomic field by enabling the simultaneous profiling of multiple RNA modifications at single-molecule resolution [25,26]. In this approach, polyadenylated native RNA molecules are translocated through nanopores by motor proteins, thereby generating characteristic ionic current signals that correspond to the underlying nucleotide sequence [26]. These raw signals are then interpreted by artificial intelligence models in a process known as basecalling to reconstruct the full RNA sequence [26]. As chemically modified nucleotide sequences generate distinct electrical signatures compared to unmodified sequences, RNA modifications can be inferred from signal profiles, leading to the development of modification-aware basecallers that infer both sequence and modification contexts [25]. This stands in contrast to conventional immunoprecipitation [27,28] and chemical labelling [29] methods, which are constrained by limited quantitative resolution [30,31], cross-reactivity challenges [30,32], and their ability to profile only a single modification at a time indirectly on enriched cDNA fragments [25]. Central to this breakthrough in epitranscriptomic profiling has been the rapid development of computational tools for modification detection, which have progressed from early statistical approaches [33–36], through classical machine learning (ML) models [37–39], and more recently to sophisticated deep learning frameworks [40–43]. In addition to basecalling and detecting modified bases, computational algorithms have also been developed for read error correction [44,45], denoising and signal segmentation [46], signal-to-sequence alignment for adaptive sampling [47–49], identification of nascent RNA transcripts [50] and inference of RNA structural features [51,52].

Deep learning, a subfield of machine learning that employs multi-layered neural networks to automatically learn hierarchical representations from raw data [53], has emerged as a powerful approach for capturing the complex, non-linear signal patterns inherent to nanopore RNA sequencing. Unlike statistical modelling and classical machine learning models which depend on restrictive distributional assumptions [33,34] and manual feature engineering [37,38], deep learning learns directly from raw signals [54], thereby capturing subtle feature representations [54] and enhancing generalizability across diverse contexts [55] without being limited to localized predictions. This makes deep learning particularly well suited for nanopore data, where signals are inherently noisy due to stochastic ionic current fluctuations [56], variable translocation rates [47,57], and pore-to-pore variability [57]. These inherent variabilities often mask subtle differences between modified and unmodified reads [43,58], limiting the sensitivity of manual annotation or simple statistical approaches. Deep learning models can learn these subtle differences, distinguishing not only between modified and unmodified reads but also among different modification types [43], while producing probabilistic predictions at single-molecule resolution [43]. Hence, deep learning approaches are especially promising for nanopore signal analysis, as they are adept at managing intrinsic variability and capturing subtle modification-dependent signal features across diverse k-mer sequence contexts.

More recently, advanced strategies such as embedding RNA modification tasks into specialized deep learning frameworks and ensemble approaches have been developed to address key challenges in nanopore sequencing. Specialized learning frameworks extend beyond conventional supervision paradigms by modifying the training process or label structure to address specific data challenges or application constraints. For example, the multi-instance learning framework m6Anet [41] estimates single-molecule modification probabilities from site-level labels; transfer learning in TandemMod [42] overcomes the scarcity of high-quality training data for multi-modification detection; and the one-class classification strategy in nanoDoc2 [40] enables de novo discovery of RNA modified sites without prior labels. In parallel, ensemble learning has emerged as another promising direction, encompassing strategies such as bagging, boosting, and stacking, where multiple base models are combined to improve overall model robustness [59,60]. These approaches are particularly valuable in contexts with limited ground truth and imbalanced datasets [59], such as detecting epitranscriptomic modifications on viral RNA, which is typically less abundant than host transcripts [61]. For instance, a mixed-weight neural bagging ensemble was used to identify m6A modifications in Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) RNA [61]. This review aims to summarize these emerging trends, with the goal of catalysing future cross-disciplinary advances in computational methodologies for epitranscriptomic biology. We focus on approaches for modification detection post-basecalling, rather than on the architectures of modification-aware basecalling models, which have been comprehensively reviewed elsewhere [62,63].

2. Deep-learning pipeline and supervision paradigms in nanopore RNA sequencing

Most deep learning methods for nanopore signal analysis begin with a preprocessing step known as ‘resquiggling’ (or event alignment), which aligns raw ionic current signals to a reference sequence at the k-mer level through signal segmentation and probabilistic alignment [41,43,47] (Figure 1(a)). Several algorithms have been developed for this purpose [46,47,64,65], each adopting distinct architectures with different trade-offs in accuracy and efficiency. Nanopolish employs a hidden Markov model with dynamic programming to compute the most probable alignment [64], Tombo statistically compares observed signals to precomputed pore models [65], and Uncalled4 utilizes a banded dynamic time warping algorithm guided by basecaller metadata to achieve faster and more memory-efficient raw signal alignment [47]. In most deep learning pipelines, such signal-to-reference alignment provides the foundation for various signal analyses ranging from modification detection [41,66] to training novel modification-aware basecalling models [67].

Figure 1.

Figure 1.

Deep learning pipeline in nanopore RNA sequencing. (a) The workflow typically begins with event alignment (or ‘resquiggling’), which maps raw current signals to reference sequences. Various resquiggling algorithms are utilized, including Nanopolish [64], Tombo [65], and Uncalled4 [47]. Then, in conventional machine learning methods, manually engineered features including signal mean and standard deviation are used to distinguish modified from unmodified sites [37,38]. In deep learning, models directly learn feature representations from the raw signals without manual feature extraction [53]. (b) Supervision paradigms in nanopore sequencing typically include supervised [43] and semi-supervised [40,41] methods, which rely on training datasets derived from in vitro modified transcripts [40,42,43], in vivo data from knockout/knockdown promoting or depleting RNA modifications [68,69], from high-throughput sequencing datasets [41,70].

Deep learning models for nanopore analysis are mostly trained under supervised [43] or semi-supervised [40,41] paradigms (Figure 1b). Fully unsupervised approaches remain less common, and while they have shown utility for de novo discovery of modified sites [71,72], they generally lack quantitative precision, nucleotide-level resolution and modification-type specificity [40]. In contrast, supervised methods often produce higher resolution modification-type specificity by leveraging high quality labelled datasets [43]. In cases where fully labelled training datasets are scarce, semi-supervised training leverages both labelled and unlabelled data to improve generalization and model robustness [41].

In supervised learning, each training instance requires accurate ground-truth labels, typically obtained through in vitro transcription (IVT) [37,43], in vivo knockout/knockdown (KO/KD) [68,69,73] models or orthogonal high-throughput sequencing (HTS) datasets [41,70]. While IVT provides controlled training data, it can be limited in sequence and modification diversity compared to endogenous RNA [68]. In contrast, KO/KD strategies can introduce off-target effects that inadvertently perturb additional modification types, thereby confounding interpretations [74]. In some cases, in vivo transcription has also been leveraged, enabling training on naturally occurring m6A-modified transcripts under authentic biological conditions and capturing the full diversity of endogenous sequences and modification contexts [68]. Model performance depends on the choice of method, relative transcript abundance and the quality of training data, while expanding dataset diversity in future developments will be crucial for improving representation learning, enhancing basecalling accuracy and model generalizability [55].

For more advanced biological investigations where ground-truth labels are often unavailable, weakly supervised strategies have been utilized to accommodate partial annotations and noisy labels. For example, m6Anet utilized site-level modification labels obtained from immunoprecipitation experiments for modelling read-level signals to make probabilistic predictions that generalize well across diverse cell lines [41]. NanoDoc2 represents another hybrid weakly supervised approach in which a CNN with WaveNet architecture and an encoder is first trained in a supervised manner to learn embeddings of raw nanopore signals for all possible 6-mer sequences [40]. These embeddings are then applied to WT and IVT RNA, where unsupervised clustering and statistical tests in the embedding space were used to detect subtle differences between modified and unmodified signal distributions for de novo discovery of modified sites (Figure 1b). Weakly supervised strategies offer a practical way to harness partial or noisy annotations and overcome the persistent challenge of scarce ground-truth labels and are expected to become increasingly central to the development of future advanced nanopore RNA models.

3. Classical deep learning methods

Expanding on the basic densely connected neural network (DNN), which comprises multiple layers of interconnected neurons, classical deep learning architectures include convolutional neural networks (CNNs), which extract local features across different regions of the input, recurrent neural networks (RNNs) and their variants – such as long short-term memory (LSTM) and gated recurrent unit (GRU) – which are well-suited for modelling sequential dependencies, and autoencoders, which generate compact, low-dimensional representations of data (Figure 2). Transformer architectures, leveraging self-attention mechanisms to capture both local and long-range dependencies have also been utilized (Figure 2). Among these, CNN and RNN architectures have been most widely utilized in model development for post-basecalling analysis, with CNNs effectively capturing local current fluctuations within k-mer contexts and RNNs modelling the sequential dependencies of ionic current traces (Table 1). Transformer-based architectures are increasingly used in basecalling model development, leveraging self-attention mechanisms to achieve improved performance [75].

Figure 2.

Figure 2.

Schematic representation of classical deep learning architectures. Convolutional Neural Networks capture spatial hierarchies in signals through successive convolution and pooling operations, making them well-suited for extracting features from nanopore signals. Recurrent Neural Networks model sequential dependencies by processing inputs in order, which enables interpretation of temporal signal patterns. Transformers leverage self-attention to capture long-range dependencies in sequences without recurrent connections. Autoencoders learn compressed latent representations of input data through an encoder-decoder framework, facilitating tasks such as signal denoising and dimensionality reduction. Among these, CNN and RNN remain the most common architectures for post-basecalling analyses.

Table 1.

Comparison of deep learning approaches for nanopore RNA sequencing.

Method Model Training Data Model Output Cross-species performance Benchmark/
Validation
CHEUI [43] CNN for predicting m6A and m5C modified sites in model 1, binary classifier in model 2 In vitro transcribed sequences Simultaneous prediction of
m6A and m5C modified sites at read-level resolution; differential methylation analysis
Mouse (embryonic brain tissue) and human (WT and NSUN2-KO HeLa cells) EpiNano, NanoRMS, Nanocompore, xPore, Tombo
NanoSPA [69] Feedforward neural network Nanopore sequencing of in vivo wild-type HeLa cells with m6A motifs identified using m6A-SAC-seq Simultaneous prediction of m6A and Ψ sites. N/D GLORI, m6Anet
DENA [68] BiLSTM In vivo wild-type and m6A-deficient A. thaliana Identification and quantification of m6A modified sites at single-nucleotide and isoform resolution A. thaliana mutants and human Nanom6A, SCARLET, LEAD-m6A-seq
m6Anet [41] Multiple instance learning based neural network miCLIP, m6ACE-seq and nanopore sequencing of human HCT116 and HEK293T cell lines. Prediction of m6A modified sites with read-level m6A modification probabilities. Human (HEK293T, HCT116), Arabidopsis VIR-1 mutant, Curlcake sequences Tombo, EpiNano, MINES, nanom6A
nanoDoc2 [40] CNN with WaveNet architecture In vitro transcribed sequences Prediction of modified sites N/D N/D
TandemMod [42] CNN, BiLSTM, Attention Unit In vitro transcribed sequences, Curlcake and ELIGOS public datasets Prediction of modified sites, types (m6A, m1A, m5C, m7G, hm5C, inosine, Ψ), and associated probability scores Wild-type and METTL3 knockout HEK293T cell line, A549, HCT116, MCF7, HepG2, K562, HeLa human cell lines, Yeast, Rice. Nanom6A, m6Anet, m6A-seq, miCLIP, m6ACE-seq.
Mixed-Weight Neural Bagging [61] Hybrid ensemble framework combining LightGBM and LSTM models through a neural bagging strategy. In vitro transcribed sequences Prediction of m6A modified sites on viral SARS-CoV-2 RNA. N/D SVM, RF, DT, ET, LightGBM, RNN, GRU, LSTM, Bagging-DT, Bagging-LightGBM; validated on in vivo SARS-CoV-2 (Vero cells).

Deep Learning methods are evaluated by model design, training data sources, output type, cross-species performance, and whether benchmarking of model or validation of predicted results have been performed. Abbreviations: SVM (Support Vector Machine), RF (Random Forest), DT (Decision Trees), ET (Extra Trees), RNN (Recurrent Neural Network), GRU (Gated Recurrent Unit), CNN (Convolutional Neural Network), BiLSTM (Bidirectional Long Short-Term Memory), LSTM (Long Short-Term Memory), N/D (Not Demonstrated).

Conventional methods struggle to resolve multiple RNA modifications simultaneously as modification-specific signal shifts are often subtle and highly context-dependent across different k-mer sequences [43]. By distinguishing m6A and m5C methylation signals, deep learning models have shown promise for advancing single-molecule co-modification profiling [43]. CHEUI employs a two-stage CNN for profiling co-occurring m6A and m5C modifications at single-molecule resolution [43]. Model 1 learns local signal features from nanopore traces by representing each candidate site within a 9-mer context (NNNN[A|C]NNNN) [43]. Separate Model 1 networks were trained for m6A and m5C using IVT-derived datasets, with positive datasets consisting of sequences in which all adenines or cytosines were replaced by m6A and m5C respectively, and negative datasets consisting of unmodified transcripts [43]. This training approach does not recapitulate naturally occurring methylation patterns and may result in overfitting, but it ensures label accuracy. Observed signals from overlapping 5-mers are segmented and normalized into fixed-length vectors, compared against expected signals from the Nanopolish k-mer model, and the resulting observed values and absolute differences are used as inputs to the CNN [43]. Model 1 outputs a per-read modification probability at a candidate site, enabling direct comparison of modification levels across conditions. The distribution of these read-level probabilities can then be aggregated and passed to Model 2, which predicts both the site-level probability of modification and its stoichiometry. Using hybrid CNN architectures, CHEUI distinguishes m6A from m5C signals, enabling co-modification profiling at single-molecule resolution. NanoSPA is a machine learning-based method that utilizes neural networks for m6A and Ψ co-profiling [69]. Unlike CHEUI, it employs m6A motif-specific feedforward neural networks (or multilayer perceptrons) trained on hand-engineered features derived from basecalling errors to operate within the same feature space as NanoPsu, a previously developed model for predicting Ψ [76]. Leveraging shared features extends the framework from Ψ-only in NanoPsu to simultaneous m6A and Ψ analysis in NanoSPA, while preserving the interpretability of the input features. Deep learning models are instrumental for profiling multiple modifications simultaneously by distinguishing their signal signatures, thereby advancing nanopore RNA analysis towards a more faithful representation of endogenous RNA than studying each modification in isolation.

Unlike earlier approaches trained on in vitro synthetic RNA, DENA is trained on wild-type and m6A-deficient in vivo transcribed RNA in Arabidopsis model, thereby enabling it to capture signal deviations across diverse endogenous motifs at natural modification stoichiometries [68]. Architecturally, it employs a Bidirectional Long Short-Term Memory (BiLSTM) neural network to model contextual and sequential dependencies in signals both upstream and downstream of candidate sites. Unlike models trained on in vitro transcribed RNAs, where electrical signals are often artificially saturated by heavily modified sequences aimed at amplifying electrical differences [58,68], DENA leverages endogenous RNA transcripts to capture natural, subtle signal variations without inflated signal deviations [68]. This design allows DENA to resolve interactions between isoform identity and to detect subtle shifts in raw signal patterns, enabling isoform-aware non-modified A versus m6A profiling at single-nucleotide resolution [68]. In contrast, MINES is a machine learning-based model for isoform-level m6A detection that employs a random forest classifier trained on hand-engineered signal-derived features and immunoprecipitation-derived labels [39]. While random forests are robust to noise and capable of modelling non-linear relationships, their reliance on predefined features limits their ability to capture subtle sequential dependencies as effectively [77]. Nevertheless, random forests, as implemented in MINES, represent an advance over the statistical learning approaches of EpiNano, which relied on support vector machine classifiers trained on basecalling error profiles of in vitro transcribed sequences rather than primary signal deviations. Hence, EpiNano demonstrated limited generalizability and is unable to distinguish m6A sites within k-mers containing more than one adenosine [73,77,78]. Furthermore, its reliance on basecalling errors also suggests that the method may become obsolete as basecalling accuracy continues to improve [77].

Deep learning models were also developed to identify non-canonical nucleotide analogs, such as 5-ethynyl uridine (5-EU), which are metabolically incorporated into newly synthesized RNA for tracking nascent transcription [50,79]. This approach enables the study of RNA kinetics and dynamics, including transcript half-lives and turnover. RNAkinet is a deep learning framework that distinguishes 5-EU-labelled nascent transcripts from mature pre-existing molecules, thereby quantifying isoform-specific half-lives and metabolic dynamics at the single-molecule level [50]. This provides a powerful means to investigate post-transcriptional regulation and RNA dynamics at single-molecule resolution. RNAkinet employs a hybrid CNN-RNN architecture trained on human cell line, which demonstrated robust cross-species generalizability in mouse datasets [50]. In comparison, nano-ID, which uses a dense neural network coupled with kinetic modelling and trained on engineered signal features to classify nascent and mature transcripts [79], showed limited generalizability when tested on independent datasets [50].

4. Specialized learning frameworks

Beyond standard deep learning strategies, various specialized deep learning paradigms have emerged to overcome challenges relating to data scarcity [41], semi-labelled data [41], and the need for robust generalization across different modification types or species [42]. These specialized approaches enable effective learning under limited labelling [41], support weak supervision [41], facilitate anomaly and outlier detection [40], and promote knowledge reuse [42]. In the context of RNA nanopore sequencing, strategies including multiple instance learning, where models are trained with group-level labels rather than instance-level labels [41], transfer learning, where representations from pretrained models are adapted to new tasks [42], and deep one-class classification, where models are trained exclusively on baseline examples for identifying anomalies [40] have been implemented (Figure 3).

Figure 3.

Figure 3.

Specialized deep learning strategies extend beyond classical architectures to address unique challenges in nanopore signal analysis. Multiple Instance Learning is a weakly supervised strategy where labels are assigned to groups of signals rather than individual events. It is useful for tasks where ground truth at single-molecule resolution is unavailable. One Class Classification is a strategy focused on distinguishing normal from abnormal signals using only positive training data, making it well-suited for detecting novel modified sites in nanopore reads. Transfer Learning adapts models trained on large datasets to related tasks with limited data, enabling efficient reuse of pretrained architectures for new nanopore applications such as basecalling of new RNA modifications or in different organisms.

m6Anet is a neural network-based method that applies a Multiple Instance Learning (MIL) framework to detect m6A modification sites at both the site and read level, addressing the challenge of missing per-molecule ground-truth annotations of m6A sites [41]. In MIL, labels are assigned to groups of instances (known as ‘bags’) rather than to individual instances [41]. A bag is labelled positive if it contains at least one positive instance, and negative, if all instances are negative [41]. This allows the model to learn from weak, group-level supervision even when fine-grained labels are unavailable. Hence, by leveraging MIL and grouping reads that align to the same genomic site, m6Anet can infer modifications even without explicit per-molecule annotations. This allows the model to not only identify modified sites, but also to estimate modification stoichiometry by quantifying the proportion of modified reads at each site [41]. In contrast, xPore models signal distributions of modified and unmodified reads using a statistical mixture model to estimate site-level modification status and stoichiometry but cannot assign read-level modification probabilities [33]. While NanoRMS extracts per-read signal features and applies machine learning classifiers to assign read-level modification status and estimate stoichiometry, its performance is highly dependent on the quality and discriminative power of the chosen features [37]. By leveraging MIL, m6Anet overcomes these limitations, outperforming the limitations of existing computational methods, matching the accuracy of experimental approaches, and generalizing across different cell types and organisms [41].

NanoDoc2 adopts a WaveNet-based convolutional neural network, originally developed for modelling human speech waveforms, to represent raw ionic current traces from nanopore sequencing [40,80]. The WaveNet architecture, which captures both local signal fluctuations and long-range dependencies in sequential waveform-like inputs, is well-suited for nanopore signals that share these characteristics. The model is trained exclusively on unmodified signal data, which serves as a baseline for learning the expected signal distribution for each 6-mer sequence context [40]. By framing modification detection as an anomaly detection task, nanoDoc2 employs a deep one-class classification strategy [81] and a scoring function that quantifies deviations from baseline signals for flagging candidate modified sites. By leveraging abundant unmodified data, this approach enables the detection of putative modifications directly as deviations from expected patterns, making it broadly generalizable across diverse epitranscriptomic marks and sequence contexts. However, orthogonal validation remains necessary to determine the precise chemical identity of modifications, and integration with complementary models would be required for stoichiometry estimation.

TandemMod is a deep learning framework that applies transfer learning to enable the detection of multiple RNA modification types including m6A, m5C, m1A, hm5C, m7G, Ψ and inosine, even with limited available training data [42]. In transfer learning, a model pre-trained on larger datasets may be reused by adapting its learned representations or features to improve its performance on a related task with limited task-specific data [82]. By leveraging shared features across tasks, this strategy significantly reduced the amount of training data and computational resources required while maintaining high accuracy [42]. In practice, TandemMod is first trained on in vitro transcribed m5C datasets and then fine-tuned with much smaller datasets for other modifications. The model uses a hybrid architecture of 1D CNN, a BiLSTM, and attention layers to capture both temporal and contextual dependencies in nanopore signals. TandemMod has been validated on both in vitro and in vivo datasets in rice under diverse conditions and further demonstrated to generalize to human datasets. TandemMod addresses the challenges of limited or incomplete labelled data for many RNA modifications by harnessing transfer learning, thereby enabling efficient adaptation of models across multiple modification types [42]. Unlike traditional approaches that require separate training for each modification, TandemMod provides a feasible and scalable solution for comprehensive epitranscriptome profiling.

5. Ensemble deep learning strategies

Ensemble learning methods integrate multiple models to reduce overfitting, thereby improving both accuracy and generalization [59]. In neural ensembles, multiple sub-models are trained, and different weights can be assigned to their predictions to balance contributions and prevent bias from a dominant architecture [60]. In bioinformatics, ensemble neural networks have been proven effective in overcoming the challenges of small sample sizes, high dimensionality, imbalanced class distributions, and noisy or heterogeneous data [59]. These properties make ensemble approaches particularly suitable for nanopore sequencing signals [47,57]. In cases where data are limited and imbalanced, such as viral RNA reads which typically represent only a small fraction of total reads compared to host transcripts [61], ensemble methods may offer robustness and generalizability.

In supervised ensemble learning, the three most common strategies are bagging, boosting, and stacking (Figure 4). In bagging (or bootstrap aggregating), multiple models are trained independently in parallel on different subsets of the data, and their predictions are aggregated to make the final prediction. In boosting, models are trained sequentially, with each iterative model focusing on data misclassified or poorly predicted by the previous model. In stacking (or stacked generalization), diverse base models are trained on the same dataset, and their predictions are then used as input features for a meta-learner, which produces the final prediction (Figure 4).

Figure 4.

Figure 4.

Ensemble strategies combine the strengths of multiple models to improve robustness and predictive performance. Bagging trains multiple models in parallel on different subsets of the data and aggregates their predictions, reducing variance and improving stability. Boosting sequentially trains models where each model focuses on correcting the errors of the previous, enhancing accuracy for challenging signal patterns. Stacking integrates predictions from diverse base models through a meta-learner, enabling complementary architectures to be combined for enhanced performance.

There is great interest in profiling viral RNA modifications, given their potential roles in key processes such as viral-host interactions [83], evolution [83], replication [84], immune responses [85], and splicing efficiency [86]. Direct RNA sequencing has enabled full-length, site-specific mapping of m6A in HIV-1 [87]. However, for cytoplasmic RNA viruses such as chikungunya [84] and dengue [85], the mechanism underlying m6A acquisition remains unclear, given that the methyltransferase machinery is localized in the nucleus [88]. Conflicting results have been reported [89,90], and while some studies report extensive viral m6A modifications, direct RNA sequencing indicates that these may be absent altogether [88]. Beyond m6A, other modifications, including Ψ and A-to-I editing, have also garnered attention [83,91]. Accurately detecting these modifications continues to pose challenges, necessitating orthogonal validation strategies and methodological advances for precise mapping [88,92].

Shorter and less abundant RNA reads, such as some viral RNAs, yield noisy and low-coverage nanopore signals that are highly prone to background noise and variability [83,93]. These factors make accurate modified base assignment particularly challenging. Ensemble approaches help mitigate these issues by enhancing learning and improving the performance of modification detection. Mixed-weight neural bagging addresses this challenge by integrating two complementary classifiers: a feature-based LightGBM model that extracts engineered descriptors from nanopore signals, and an LSTM model that captures sequential dependencies on raw current traces [61]. Both networks were trained via bagging and integrated using a mixed-weight strategy, in which a neural network assigns optimized weights to balance their contributions, thereby improving robustness and accuracy in identifying m6A modifications from low-abundance viral RNA [61]. This ensemble deep learning framework merging feature-driven models and raw signal models through neural-weighted bagging provides a robust solution for detecting m6A modification in SARS-CoV-2 RNA [61], although broader experimental validation across multiple viruses and host systems remains necessary.

Although not leveraging nanopore signals, sequence-based ensemble methods offer complementary relevance for distinguishing the modified sites of m6A from the structurally similar m1A isomer [43] at the motif level. Distinguishing between these isomers from nanopore signals is challenging due to their subtle differences, which is further compounded by limitations in training data where antibody-based approaches often suffer from false positives and poor resolution due to the limited specificity of immunoprecipitation [94,95] and knockout/knockdown approaches [96]. Most sequence-based methods encode RNA using multi-encoding integrations, such as one-hot, Word2Vec, and RNA embeddings to capture sequence contexts without relying on raw current signal data [97–100]. For example, DeepPromise [100] combines CNNs across enhanced nucleic acid composition, one-hot, and RNA embeddings for simultaneous m6A and m1A site prediction with high predictive performance for m1A but modest performance for m6A across human, mouse, and yeast datasets. Similarly, EMDLP employs dilated CNN-BiLSTM architecture with soft voting on one-hot, RNA word, and RGloVe encodings for high-accuracy classification of m6A and m1A modified sites in human datasets, outperforming single-branch models [99]. Gene2vec [98] and EDLm6APred [97] likewise employ CNN voting and weighted BiLSTM ensembles respectively, on diverse embeddings for human and mouse m6A modified sites. Although sequence-based methods do not leverage signal deviations for identifying modifications at nucleotide and isoform levels, future hybrids fusing sequence motifs with nanopore-derived features may eventually enhance isomer resolution and generalizability to enable higher-resolution epitranscriptomic perspectives.

6. Future directions

Future advances in RNA modification detection using nanopore sequencing will depend on addressing several outstanding challenges. More than 170 distinct chemical modifications exist, with many modifications sharing high structural similarity, which makes distinguishing signals challenging [101]. While progress has been made with methods such as TandemMod, which extend detection to modification types with more limited datasets, further advances are needed to achieve higher resolution multiplexed analyses. More sophisticated ensemble approaches may enhance generalizability and accuracy, while specialized learning strategies designed to overcome data scarcity may help address current limitations [100]. At the algorithmic level, while current basecalling models, can detect modifications, distinct modifications are processed separately, except for the latest version of Oxford Nanopore Technologies (ONT) proprietary Dorado basecaller, which increases overall computational time. Future basecallers capable of simultaneously identifying multiple modifications would accelerate analysis and enable more comprehensive profiling of the RNA modification code.

Many current studies lack standardized benchmarking and cross-platform validation, making it difficult to identify a clear ‘state of the art’ and to compare model performance across methods and platforms. Addressing this gap will require standardization of evaluation metrics for modified basecalling (accuracy/AUROC and calibration) and for stoichiometric estimation (correlation and absolute error), together with systematic head-to-head benchmarking as key priority. Confidence in modification callers should be further strengthened by validating results with orthogonal biological assays, which would reduce false positives and mitigate basecaller-specific biases. Performing such validation across basecallers, sequencing depths and defined stoichiometric mixtures will clarify how unsupervised approaches compare to supervised models in accuracy and robustness, and whether downstream single-modification callers offer superior accuracy for rare and complex modifications relative to faster, integrated simultaneous multi-modification calling by the latest version of Dorado, recommended by ONT and widely utilized by Nanopore users. Benchmarking efforts should also explicitly evaluate isomer discrimination and include rigorous assessment across computational tools using rationally designed oligonucleotide models or well-established reference datasets, while accounting for intrinsic biases of each platform [77].

Developing interpretable models presents another promising direction, linking predictions to meaningful biological features, such as signal characteristics or sequence motif, that drive modification calls. Approaches such as attention weight visualization, saliency mapping, feature attribution, or perturbation analyses could reveal aspects of signal or sequence most informative for each prediction, enhancing both computational calls and hypothesis generation, and uncovering mechanistic insights into determinants of RNA modification calls.

Beyond existing strategies, other specialized learning paradigms merit exploration. Multi-task learning may be particularly relevant, as modifications often share overlapping motifs and contextual features, allowing models to learn common representations, while improving prediction across individual tasks. Meta-learning frameworks, which adapt quickly to new tasks with few labelled examples, could be applied to modifications with scarce training data. Similarly, curriculum learning, where models are trained on progressively more complex signals, may yield better generalization across diverse contexts.

Future work in ensemble learning approaches could adopt more advanced strategies for RNA modification detection. Bayesian deep ensembles may be particularly useful for low-coverage viral RNA or rare modification detection, where robust uncertainty estimates are more informative than binary predictions. Mixture of expert models, in which networks specialize in distinct sequence contexts, and a gating network selects the most appropriate expert, align well with the context-dependent nature of nanopore signals. Hierarchical ensembles that integrate predictions across multiple levels may further improve robustness. Importantly, these methods should be coupled with orthogonal validation, tested for cross-species generalizability, and developed with interpretability in mind to maximize biological relevance [58,88].

Finally, multiplexed profiling of RNA modifications at single-cell resolution remains an exciting frontier. As current protocols are designed for the capture of polyadenylated RNA, typically mRNAs, most single-cell RNA-seq workflows on Oxford Nanopore platforms remain indirect. RNA molecules are reverse transcribed to cDNA to be barcoded and sequenced on DNA flow cells [102,103], thereby failing to capture native RNA modifications. While WarpDemuX and SeqTagger recently enabled multiplexed direct RNA sequencing by classifying reads using unique adapter-barcode signals with lightweight machine learning models, important constraints remain [104,105]. These methods were restricted to polyadenylated RNA and performed poorly on short reads [104,105], although the latest update of WarpDemux-tRNA was reported to have addressed these issues [106]. Barcode-based adaptive sampling is limited to the older RNA002 sequencing model and is not yet available for the newer RNA004 sequencing chemistry [104]. In addition, less than twelve adapter barcodes for multiplexing on RNA004 sequencing chemistry are presently supported – far fewer than combinatorial barcoding schemes on short-read single-cell platforms, which scale exponentially across multiple split-pool ligation rounds [107]. Consequently, multiplexing native RNA molecules requires further development before it can routinely resolve cell-to-cell epitranscriptomic heterogeneity and provide new biological insights into regulatory and disease mechanisms.

7. Conclusions

Nanopore RNA sequencing has ushered in a new era of epitranscriptomic studies, with AI playing a central role. Foundational deep learning architectures have not only enabled feature extraction from nanopore signals for single-modification profiling but also for detecting complex co-modification patterns that were previously inaccessible using short-read platforms. Building on these advances, more sophisticated deep learning approaches have emerged to address the unique challenges of nanopore sequencing. By embedding models within specialized learning frameworks and ensemble strategies, these methods have proven promising in overcoming limitations in training data availability and extending applicability across diverse experimental contexts, thereby unlocking novel biological insights. Looking ahead, several key directions are expected to shape the field, including multiplexed RNA modification profiling, the development of interpretable models, the establishment of consensus practices and benchmarking standards, and continued improvements in model accuracy. Expanding on the use of ensemble strategies and specialized learning frameworks will further enhance robustness of nanopore RNA analysis. Together, these advances will not only refine the computational toolkit for nanopore analysis but also accelerate the discovery of epitranscriptomic principles governing RNA biology in health and disease. With the transformative synergy of AI and nanopore sequencing, the future of the field is full of exciting possibilities and is poised to unlock new dimensions of RNA regulation and functions in unprecedented ways.

Acknowledgments

Following initial conceptualization and drafting by the author, ChatGPT-4o was used to identify language and formatting errors.

Conceptualization – C.C.Y.L and B.L.; Visualization – C.C.Y.L; Writing, original draft – C.C.Y.L; Writing, reviewing & editing – C.C.Y.L and B. L; Funding acquisition – M.J.F and K.C.K; Supervision – M.J.F and K.C.K.

Funding Statement

This research is supported by the National Research Foundation, Singapore, under its AI Singapore Programme [AISG Award No: AISG3-GV-2023–014] and by the Ministry of Education, Singapore, under its [Academic Research Fund Tier 1 (RG38/23)], both awarded to M.J.F. (PI) and K.C.K. (Co-I).

Disclosure statement

No potential conflict of interest was reported by the author(s).

Data availability statement

Data sharing is not applicable to this article as no new data were created or analysed in this study.

References

  • [1].Rücklé C, Körtel N, Basilicata MF, et al. Rna stability controlled by m6a methylation contributes to x-to-autosome dosage compensation in mammals. Nat Struct Mol Biol. 2023;30(8):1207–1215. doi: 10.1038/s41594-023-00997-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [2].Lee Y, Choe J, Park OH, et al. Molecular mechanisms driving mRNA degradation by m6a modification. Trends Genet. 2020;36(3):177–188. doi: 10.1016/j.tig.2019.12.007 [DOI] [PubMed] [Google Scholar]
  • [3].Mendel M, Delaney K, Pandey RR, et al. Splice site m6A methylation prevents binding of U2AF35 to inhibit RNA splicing. Cell. 2021;184(12):3125–3142. e25. doi: 10.1016/j.cell.2021.03.062 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Garbo S, D’Andrea D, Colantoni A, et al. m6A modification inhibits miRNAs’ intracellular function, favoring their extracellular export for intercellular communication. Cell Rep. 2024;43(6):114369. doi: 10.1016/j.celrep.2024.114369 [DOI] [PubMed] [Google Scholar]
  • [5].Dominissini D, Rechavi G.. 5-methylcytosine mediates nuclear export of mRNA. Cell Res. 2017;27(6):717–719. doi: 10.1038/cr.2017.73 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [6].Li J, Xie G, Tian Y, et al. Rna m6a methylation regulates dissemination of cancer cells by modulating expression and membrane localization of β-catenin. Mol Ther. 2022;30(4):1578–1596. doi: 10.1016/j.ymthe.2022.01.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Wang H, Hu X, Huang M, et al. Mettl3-mediated mRNA m6A methylation promotes dendritic cell activation. Nat Commun. 2019;10(1):1898. doi: 10.1038/s41467-019-09903-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Crespo-García E, Bueno-Costa A, Esteller M. Single-cell analysis of the epitranscriptome: rNA modifications under the microscope. RNA Biol. 2024;21(1):1–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Esteller M, Pandolfi PP. The epitranscriptome of noncoding RNAs in cancer. Cancer Discov. 2017;7(4):359–368. doi: 10.1158/2159-8290.CD-16-1292 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Wiener D, Schwartz S. The epitranscriptome beyond m6a. Nat Rev Genet. 2021;22(2):119–131. doi: 10.1038/s41576-020-00295-8 [DOI] [PubMed] [Google Scholar]
  • [11].Patil DP, Chen C-K, Pickering BF, et al. m6A RNA methylation promotes XIST-mediated transcriptional repression. Nature. 2016;537(7620):369–373. doi: 10.1038/nature19342 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Engel M, Eggert C, Kaplick PM, et al. The role of m6a/m-RNA methylation in stress response regulation. Neuron. 2018;99(2):389–403.e9. doi: 10.1016/j.neuron.2018.07.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].An X, Wang R, Lv Z, et al. Wtap-mediated m6a modification of FRZB triggers the inflammatory response via the Wnt signaling pathway in osteoarthritis. Exp Mol Med. 2024;56(1):156–167. doi: 10.1038/s12276-023-01135-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Liu J, Eckert MA, Harada BT, et al. m6A mRNA methylation regulates AKT activity to promote the proliferation and tumorigenicity of endometrial cancer. Nat Cell Biol. 2018;20(9):1074–1083. doi: 10.1038/s41556-018-0174-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Li R, Zhao H, Huang X, et al. Super-enhancer RNA m6A promotes local chromatin accessibility and oncogene transcription in pancreatic ductal adenocarcinoma. Nat Genet. 2023;55(12):2224–2234. doi: 10.1038/s41588-023-01568-8 [DOI] [PubMed] [Google Scholar]
  • [16].Jiang Q, Crews LA, Holm F, et al. Rna editing-dependent epitranscriptome diversity in cancer stem cells. Nat Rev Cancer. 2017;17(6):381–392. doi: 10.1038/nrc.2017.23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Lee AC, Lee Y, Choi A, et al. Spatial epitranscriptomics reveals A-to-I editome specific to cancer stem cell microniches. Nat Commun. 2022;13(1):2540. doi: 10.1038/s41467-022-30299-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].de Mendonça Fernandes GM, Wang W, Ahmadian SS, et al. Epitranscriptomic analysis reveals clinical and molecular signatures in glioblastoma. Acta Neuropathol Commun. 2025;13(1):74. doi: 10.1186/s40478-025-01966-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Matsumura Y, Wei F-Y, Sakai J. Epitranscriptomics in metabolic disease. Nat Metab. 2023;5(3):370–384. doi: 10.1038/s42255-023-00764-4 [DOI] [PubMed] [Google Scholar]
  • [20].Wu S, Zhang S, Wu X, et al. m6A RNA methylation in cardiovascular diseases. Mol Ther. 2020;28(10):2111–2119. doi: 10.1016/j.ymthe.2020.08.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].Gatsiou A, Stellos K. Rna modifications in cardiovascular health and disease. Nat Rev Cardiol. 2023;20(5):325–346. doi: 10.1038/s41569-022-00804-8 [DOI] [PubMed] [Google Scholar]
  • [22].Relier S, Amalric A, Attina A, et al. Multivariate analysis of RNA chemistry marks uncovers epitranscriptomics-based biomarker signature for adult diffuse glioma diagnostics. Anal Chem. 2022;94(35):11967–11972. doi: 10.1021/acs.analchem.2c01526 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].Yang C, Han H, Lin S. Rna epitranscriptomics: a promising new avenue for cancer therapy. Mol Ther. 2022;30(1):2–3. doi: 10.1016/j.ymthe.2021.12.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Cerneckis J, Ming G-L, Song H, et al. The rise of epitranscriptomics: recent developments and future directions. Trends in pharmacological sciences. Trends Pharmacol Sci. 2024;45(1):24–38. doi: 10.1016/j.tips.2023.11.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Diensthuber G, Novoa EM. Charting the epitranscriptomic landscape across RNA biotypes using native RNA nanopore sequencing. Molecular Cell. 2025;85(2):276–289. doi: 10.1016/j.molcel.2024.12.014 [DOI] [PubMed] [Google Scholar]
  • [26].Deamer D, Akeson M, Branton D. Three decades of nanopore sequencing. Nat Biotechnol. 2016;34(5):518–524. doi: 10.1038/nbt.3423 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Dominissini D, Moshitch-Moshkovitz S, Schwartz S, et al. Topology of the human and mouse m6A RNA methylomes revealed by m6A-seq. Nature. 2012;485(7397):201–206. doi: 10.1038/nature11112 [DOI] [PubMed] [Google Scholar]
  • [28].Meyer KD, Saletore Y, Zumbo P, et al. Comprehensive analysis of mRNA methylation reveals enrichment in 3’ UTRs and near stop codons. Cell. 2012;149(7):1635–1646. doi: 10.1016/j.cell.2012.05.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Xie Y, Han S, Li Q, et al. Transcriptome-wide profiling of N 6-methyladenosine via a selective chemical labeling method. Chem Sci. 2022;13(41):12149–12157. doi: 10.1039/D2SC03181G [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [30].Garcia-Campos MA, Edelheit S, Toth U, et al. Deciphering the “m(6)A code” via antibody-independent quantitative profiling. Cell. 2019;178(3):731–747.e16. doi: 10.1016/j.cell.2019.06.013 [DOI] [PubMed] [Google Scholar]
  • [31].Chen H-X, Zhang Z, Ma D-Z, et al. Mapping single-nucleotide m6A by m6A-REF-seq. Methods. 2022;203:392–398. doi: 10.1016/j.ymeth.2021.06.013 [DOI] [PubMed] [Google Scholar]
  • [32].Meyer KD. DART-seq: an antibody-free method for global m6A detection. Nat Methods. 2019;16(12):1275–1280. doi: 10.1038/s41592-019-0570-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Pratanwanich PN, Yao F, Chen Y, et al. Identification of differential RNA modifications from nanopore direct RNA sequencing with xPore. Nat Biotechnol. 2021;39(11):1394–1402. doi: 10.1038/s41587-021-00949-w [DOI] [PubMed] [Google Scholar]
  • [34].Leger A, Amaral PP, Pandolfini L, et al. Rna modifications detection by comparative nanopore direct RNA sequencing. Nat Commun. 2021;12(1):7198. doi: 10.1038/s41467-021-27393-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Jenjaroenpun P, Wongsurawat T, Wadley TD, et al. Decoding the epitranscriptional landscape from native RNA sequences. Nucleic Acids Res. 2020;49(2):e7–e7. doi: 10.1093/nar/gkaa620 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Abebe JS, Price AM, Hayer KE, et al. Drummer—rapid detection of RNA modifications through comparative nanopore sequencing. Bioinformatics. 2022;38(11):3113–3115. doi: 10.1093/bioinformatics/btac274 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Begik O, Lucas MC, Pryszcz LP, et al. Quantitative profiling of pseudouridylation dynamics in native RNAs with nanopore sequencing. Nat Biotechnol. 2021;39(10):1278–1291. doi: 10.1038/s41587-021-00915-6 [DOI] [PubMed] [Google Scholar]
  • [38].Gao Y, Liu X, Wu B, et al. Quantitative profiling of n6-methyladenosine at single-base resolution in stem-differentiating xylem of Populus trichocarpa using nanopore direct RNA sequencing. Genome Biol. 2021;22(1):22. doi: 10.1186/s13059-020-02241-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39].Lorenz DA, Sathe S, Einstein JM, et al. Direct RNA sequencing enables m(6)a detection in endogenous transcript isoforms at base-specific resolution. RNA. 2020;26(1):19–28. doi: 10.1261/rna.072785.119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Ueda H, Dasgupta B, Yu BY. Rna modification detection using nanopore direct rna sequencing and nanodoc2. Methods Mol Biol. 2023;2632:299–319. [DOI] [PubMed] [Google Scholar]
  • [41].Hendra C, Pratanwanich PN, Wan YK, et al. Detection of m6a from direct RNA sequencing using a multiple instance learning framework. Nat Methods. 2022;19(12):1590–1598. doi: 10.1038/s41592-022-01666-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Wu Y, Shao W, Yan M, et al. Transfer learning enables identification of multiple types of RNA modifications using nanopore direct RNA sequencing. Nat Commun. 2024;15(1):4049. doi: 10.1038/s41467-024-48437-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [43].Acera Mateos P, J Sethi A, Ravindran A, et al. Prediction of m6A and m5C at single-molecule resolution reveals a transcriptome-wide co-occurrence of RNA modifications. Nat Commun. 2024;15(1):3899. doi: 10.1038/s41467-024-47953-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].Wang L, Qu L, Yang L et al. An error correction method of nanopore sequencing data using deep learning. In: 2020 13th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI); 2020; Chengdu, China; p. 23–28. doi: 10.1109/CISP-BMEI51763.2020.9263622 [DOI] [Google Scholar]
  • [45].Wang L, Qu L, Yang L, et al. Nanoreviser: an error-correction tool for nanopore sequencing based on a deep learning algorithm. Front Genet. 2020;11 - 2020. doi: 10.3389/fgene.2020.00900 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [46].Cheng G, Vehtari A, Cheng L. Raw signal segmentation for estimating RNA modification from nanopore direct RNA sequencing data. eLife Sciences Publications, Ltd; 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Kovaka S, Hook PW, Jenike KM, et al. Uncalled4 improves nanopore DNA and RNA modification detection via fast and accurate signal alignment. Nat Methods. 2025;22(4):681–691. doi: 10.1038/s41592-025-02631-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [48].Shivakumar VS, Ahmed OY, Kovaka S, et al. Sigmoni: classification of nanopore signal with a compressed pangenome index. Bioinformatics. 2024;40(Sup plement_1):i287–i296. doi: 10.1093/bioinformatics/btae213 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [49].Firtina C, Mansouri Ghiasi N, Lindegger J, et al. RawHash: enabling fast and accurate real-time analysis of raw nanopore signals for large genomes. Bioinformatics. 2023;39(Sup plement_1):i297–i307. doi: 10.1093/bioinformatics/btad272 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Martinek V, Martin J, Belair C, et al. Deep learning and direct sequencing of labeled RNA captures transcriptome dynamics. NAR Genomics Bioinf. 2024;6(3). doi: 10.1093/nargab/lqae116 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [51].Aw JGA, Lim SW, Wang JX, et al. Determination of isoform-specific RNA structure with nanopore long reads. Nat Biotechnol. 2021;39(3):336–346. doi: 10.1038/s41587-020-0712-z [DOI] [PubMed] [Google Scholar]
  • [52].Stephenson W, Razaghi R, Busan S, et al. Direct detection of RNA modifications and structure using single-molecule nanopore sequencing. Cell Genomics. 2022;2(2):100097. doi: 10.1016/j.xgen.2022.100097 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53].LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015;521(7553):436–444. doi: 10.1038/nature14539 [DOI] [PubMed] [Google Scholar]
  • [54].Dematties D, Wen C, Pérez MD, et al. Deep learning of nanopore sensing signals using a bi-path network. ACS Nano. 2021;15(9):14419–14429. doi: 10.1021/acsnano.1c03842 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [55].Wang Z, Liu Z, Fang Y, et al. Training data diversity enhances the basecalling of novel RNA modification-induced nanopore sequencing readouts. Nat Commun. 2025;16(1):679. doi: 10.1038/s41467-025-55974-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [56].Xi G, Su J, Ma J, et al. A robust signal processing program for nanopore signals using dynamic correction threshold with compatible baseline fluctuations. Analyst. 2025;150(7):1386–1397. doi: 10.1039/D4AN01384K [DOI] [PubMed] [Google Scholar]
  • [57].Rang FJ, Kloosterman WP, de Ridder J. From squiggle to basepair: computational approaches for improving nanopore sequencing read accuracy. Genome Biol. 2018;19(1):90. doi: 10.1186/s13059-018-1462-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [58].Zhong Z-D, Xie Y-Y, Chen H-X, et al. Systematic comparison of tools used for m6A mapping from nanopore direct RNA sequencing. Nat Commun. 2023;14(1). doi: 10.1038/s41467-023-37596-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [59].Cao Y, Geddes TA, Yang JYH, et al. Ensemble deep learning in bioinformatics. Nat Mach Intell. 2020;2(9):500–508. doi: 10.1038/s42256-020-0217-y [DOI] [Google Scholar]
  • [60].Ganaie MA, Hu M, Malik AK, et al. Ensemble deep learning: a review. Eng Appl Artif Intel. 2022;115:105151. doi: 10.1016/j.engappai.2022.105151 [DOI] [Google Scholar]
  • [61].Liu R, Ou L, Sheng B, et al. Mixed-weight neural bagging for detecting m(6)a modifications in SARS-CoV-2 RNA sequencing. IEEE Trans Biomed Eng. 2022;69(8):2557–2568. doi: 10.1109/TBME.2022.3150420 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Wick RR, Judd LM, Holt KE. Performance of neural network basecalling tools for Oxford Nanopore sequencing. Genome Biol. 2019;20(1):129. doi: 10.1186/s13059-019-1727-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [63].Pagès-Gallego M, de Ridder J. Comprehensive benchmark and architectural analysis of deep learning models for nanopore sequencing basecalling. Genome Biol. 2023;24(1):71. doi: 10.1186/s13059-023-02903-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [64].Loman NJ, Quick J, Simpson JT. A complete bacterial genome assembled de novo using only nanopore sequencing data. Nat Methods. 2015;12(8):733–735. doi: 10.1038/nmeth.3444 [DOI] [PubMed] [Google Scholar]
  • [65].Stoiber M. De novo identification of DNA modifications enabled by genome-guided nanopore signal processing. BioRxiv. 2017:094672. doi: 10.1101/094672 [DOI] [Google Scholar]
  • [66].Wang L, Li T, Zhou Y. Accurate prediction of multiple RNA modifications from nanopore direct RNA sequencing data with RNANO. BioRxiv. 2025:2025.03. 01.640267. doi: 10.1101/2025.03.01.640267 [DOI] [Google Scholar]
  • [67].Bai G, Dhillon N, Felton C, et al. Smadd-seq: probing chromatin accessibility with small molecule DNA intercalation and nanopore sequencing. Nucleic Acids Res. 2025;53(14). doi: 10.1093/nar/gkaf671 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [68].Qin H, Ou L, Gao J, et al. DENA: training an authentic neural network model using nanopore sequencing data of Arabidopsis transcripts for detection and quantification of N6-methyladenosine on RNA. Genome Biol. 2022;23(1):25. doi: 10.1186/s13059-021-02598-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [69].Huang S, Wylder AC, Pan T. Simultaneous nanopore profiling of mRNA m6A and pseudouridine reveals translation coordination. Nat Biotechnol. 2024;42(12):1831–1835. doi: 10.1038/s41587-024-02135-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [70].Zhang Y, Hamada M. DeepM6ASeq: prediction and characterization of m6A-containing sequences using deep learning. BMC Bioinf. 2018;19(19):524. doi: 10.1186/s12859-018-2516-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [71].Ding H, Bailey AD, Jain M, et al. Gaussian mixture model-based unsupervised nucleotide modification number detection using nanopore-sequencing readouts. Bioinformatics. 2020;36(19):4928–4934. doi: 10.1093/bioinformatics/btaa601 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [72].Vujaklija I, Biđin S, Volarić M, et al. Detecting a wide range of epitranscriptomic modifications using a nanopore-sequencing-based computational approach with 1d score-clustering. Nucleic Acids Res. 2024;53(1). doi: 10.1093/nar/gkae1168 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [73].Liu H, Begik O, Lucas MC, et al. Accurate detection of m6A RNA modifications in native RNA sequences. Nat Commun. 2019;10(1):4079. doi: 10.1038/s41467-019-11713-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [74].Tahiliani M, Koh KP, Shen Y, et al. Conversion of 5-methylcytosine to 5-hydroxymethylcytosine in mammalian DNA by MLL partner TET1. Science. 2009;324(5929):930–935. doi: 10.1126/science.1170116 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [75].Li Q, Sun C, Wang D, et al. Gcrtcall: a transformer based basecaller for nanopore RNA sequencing enhanced by gated convolution and relative position embedding via joint loss training. Front Genet. 2024;15:1443532. doi: 10.3389/fgene.2024.1443532 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [76].Huang S, Zhang W, Katanski CD, et al. Interferon inducible pseudouridine modification in human mRNA by quantitative nanopore profiling. Genome Biol. 2021;22(1):330. doi: 10.1186/s13059-021-02557-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [77].Maestri S, Furlan M, Mulroney L, et al. Benchmarking of computational methods for m6A profiling with nanopore direct RNA sequencing. Brief Bioinform. 2024;25(2). doi: 10.1093/bib/bbae001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [78].Parker MT, Knop K, Sherwood AV, et al. Nanopore direct RNA sequencing maps the complexity of Arabidopsis mRNA processing and m(6)a modification. Elife. 2020;9:9. doi: 10.7554/eLife.49658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [79].Maier KC, Gressel S, Cramer P, et al. Native molecule sequencing by nano-ID reveals synthesis and stability of RNA isoforms. Genome Res. 2020;30(9):1332–1344. doi: 10.1101/gr.257857.119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [80].Rethage D, Pons J, Serra X. A Wavenet for speech denoising. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2018. IEEE. p. 5069–5073. [Google Scholar]
  • [81].Ruff L. Deep one-class classification. In: Jennifer D, Andreas K, editors. Proceedings of the 35th International Conference on Machine Learning. PMLR: Proceedings of Machine Learning Research; 2018. p. 4393–4402. [Google Scholar]
  • [82].Theodoris CV, Xiao L, Chopra A, et al. Transfer learning enables predictions in network biology. Nature. 2023;618(7965):616–624. doi: 10.1038/s41586-023-06139-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [83].Wang W, Jin Y, Xie Z, et al. When animal viruses meet N(6)-methyladenosine (m(6)a) modifications: for better or worse? Vet Res. 2024;55(1):171. doi: 10.1186/s13567-024-01424-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [84].DeJesus VA. The role of m6A RNA methylation in Chikungunya virus replication. University of Leeds; 2024. [Google Scholar]
  • [85].Zhang Y, Guo J, Gao Y, et al. Dynamic transcriptome analyses reveal m6a regulated immune non-coding RNAs during dengue disease progression. Heliyon. 2023;9(1):e12690. doi: 10.1016/j.heliyon.2022.e12690 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [86].Gokhale NS, McIntyre ABR, Mattocks MD, et al. Altered m6a modification of specific cellular transcripts affects Flaviviridae infection. Mol Cell. 2020;77(3):542–555.e8. doi: 10.1016/j.molcel.2019.11.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [87].Baek A, Lee G-E, Golconda S, et al. Single-molecule epitranscriptomic analysis of full-length HIV-1 RNAs reveals functional roles of site-specific m6As. Nat Microbiol. 2024;9(5):1340–1355. doi: 10.1038/s41564-024-01638-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [88].Baquero-Pérez B, Yonchev ID, Delgado-Tejedor A, et al. N6-methyladenosine modification is not a general trait of viral RNA genomes. Nat Commun. 2024;15(1):1964. doi: 10.1038/s41467-024-46278-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [89].Ye F, Chen ER, Nilsen TW. Kaposi’s sarcoma-associated herpesvirus utilizes and manipulates RNA N(6)-adenosine methylation to promote lytic replication. J Virol. 2017;91(16). doi: 10.1128/JVI.00466-17 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [90].Tan B, Liu H, Zhang S, et al. Viral and cellular N(6)-methyladenosine and N(6),2’-O-dimethyladenosine epitranscriptomes in the KSHV life cycle. Nat Microbiol. 2018;3(1):108–120. doi: 10.1038/s41564-017-0056-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [91].Tsai K, Cullen BR. Epigenetic and epitranscriptomic regulation of viral replication. Nat Rev Microbiol. 2020;18(10):559–570. doi: 10.1038/s41579-020-0382-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [92].Tan L, Guo Z, Wang X, et al. Utilization of nanopore direct RNA sequencing to analyze viral RNA modifications. mSystems. 2024;9(2):01163–23. doi: 10.1128/msystems.01163-23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [93].Katopodi X-L, Begik O, Novoa EM. Toward the use of nanopore RNA sequencing technologies in the clinic: challenges and opportunities. Nucleic Acids Res. 2025;53(5). doi: 10.1093/nar/gkaf128 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [94].Grozhik AV. Antibody cross-reactivity accounts for widespread appearance of m1a in 5’utrs. Nat Commun. 2019;10(1):5126. doi: 10.1038/s41467-019-13146-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [95].Helm M, Lyko F, Motorin Y. Limited antibody specificity compromises epitranscriptomic analyses. Nat Commun. 2019;10(1):5669. doi: 10.1038/s41467-019-13684-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [96].Wei J, Liu F, Lu Z, et al. Differential m(6)a, m(6)a(m), and m(1)a demethylation mediated by FTO in the cell nucleus and cytoplasm. Mol Cell. 2018;71(6):973–985.e5. doi: 10.1016/j.molcel.2018.08.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [97].Zhang L, Li G, Li X, et al. Edlm6APred: ensemble deep learning approach for mRNA m6A site prediction. BMC Bioinf. 2021;22(1):288. doi: 10.1186/s12859-021-04206-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [98].Zou Q, Xing P, Wei L, et al. Gene2vec: gene subsequence embedding for prediction of mammalian N6-methyladenosine sites from mRNA. RNA. 2019;25(2):205–218. doi: 10.1261/rna.069112.118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [99].Wang H, Liu H, Huang T, et al. Emdlp: ensemble multiscale deep learning model for RNA methylation site prediction. BMC Bioinf. 2022;23(1):221. doi: 10.1186/s12859-022-04756-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [100].Chen Z, Zhao P, Li F, et al. Comprehensive review and assessment of computational methods for predicting RNA post-transcriptional modification sites from RNA sequences. Brief Bioinform. 2020;21(5):1676–1696. doi: 10.1093/bib/bbz112 [DOI] [PubMed] [Google Scholar]
  • [101].Zhang Y, Lu L, Li X. Detection technologies for RNA modifications. Exp & Mol Med. 2022;54(10):1601–1616. doi: 10.1038/s12276-022-00821-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [102].Philpott M, Watson J, Thakurta A, et al. Nanopore sequencing of single-cell transcriptomes with scCOLOR-seq. Nat Biotechnol. 2021;39(12):1517–1520. doi: 10.1038/s41587-021-00965-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [103].You Y, Prawer YDJ, De Paoli-Iseppi R, et al. Identification of cell barcodes from long-read single-cell RNA-seq with BLAZE. Genome Biol. 2023;24(1):66. doi: 10.1186/s13059-023-02907-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [104].van der Toorn W, Bohn P, Liu-Wei W, et al. Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing. Nat Commun. 2025;16(1):3742. doi: 10.1038/s41467-025-59102-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [105].Pryszcz LP, Diensthuber G, Llovera L, et al. Rapid and accurate demultiplexing of direct RNA nanopore sequencing data with SeqTagger. Genome Res. 2025;35(4):956–966. doi: 10.1101/gr.279290.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [106].van der Toorn W, Naarmann-de Vries I, Liu-Wei W, et al. WarpDemuX-tRNA: barcode multiplexing for nanopore tRNA sequencing. Nucleic Acids Res. 2025;53(17). doi: 10.1093/nar/gkaf873 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [107].Kuijpers L. Split pool ligation-based single-cell transcriptome sequencing (SPLiT-seq) data processing pipeline comparison. BMC Genomics. 2024;25(1):361. doi: 10.1186/s12864-024-10285-3 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing is not applicable to this article as no new data were created or analysed in this study.


Articles from RNA Biology are provided here courtesy of Taylor & Francis

RESOURCES