Abstract
ac4C alteration in RNA is a conserved epigenetic mark that is critical for post-transcriptional control, mRNA stability, translational efficiency, and human immune function regulation. In the meantime, the conventional experimental procedures for predicting ac4C alteration sites are costly, time-consuming, and difficult. The precise recognition of ac4C modification sites in human mRNA has been greatly aided by computational prediction techniques using sequence data, machine learning (ML), deep learning (DL), and large language models (LLMs). The application of ML, DL, and LLM-based techniques for the identification of ac4C modification sites in human mRNA has been evaluated and contrasted in this review. Distinctively, we have also addressed the shortcomings of the existing methods and tools, as well as potential future developments. We anticipate that this study will provide sufficient information and awareness for ac4C modification research.
Keywords: N4-acetylcytidine, deep learning, large language models, feature extraction, computational tools
Introduction
N-acetyltransferase 10 (NAT10) catalyzes the addition of an acetyl group to the nitrogen at the fourth position of the cytidine base in N4-acetylcytidine (ac4C), one of the RNA posttranscriptional alterations, as illustrated in Fig. 1. The ac4C modification was primarily identified in transfer ribonucleic acid (tRNA) and ribosomal ribonucleic acid (rRNA) and is conserved in both bacterial and eukaryotic nucleic acids. It was later shown that the existence of ac4C on tRNA preserves tRNA stability and contributes to the high precision of protein translation. Furthermore, researchers also found that the thermal stability of tRNA is most important when ac4C is paired with guanine [1–5]. On the contrary, ac4C on rRNA is essential for preserving the precision of protein translation. Dysregulation of ac4C alterations has been linked to several illnesses, and these modifications serve crucial regulatory functions in gene expression. For example, ac4C alteration of the transcripts of the TGF-β pathway enhances tumor growth in hepatocellular carcinoma [6–9]. Similar to this, abnormal ac4C patterns on neural transcripts impact synaptic function in Alzheimer’s disease [3, 10–13]. Accurately locating ac4C sites throughout the genome is crucial for comprehending illness mechanisms and creating diagnostic tools in light of these disease connections. Reliable computational prediction techniques are necessary, though, because experimental identification is expensive and time-consuming.
Figure 1.
Depiction of ac4C alteration in RNA.
Arango et al. [14] recently employed an ac4C-specific acRIP-seq technique to identify more than 4000 ac4C spikes in the transcriptome of human HeLa cells. The acRIP-seq approach yields a sequencing map that offers the greatest number of areas, but it is not single-base resolution. They showed that the human transcriptome has a large number of ac4C modification sites, the majority of which are found in coding regions. Additionally, ac4C-modified mRNAs have a much longer half-life than all other transcripts, particularly when mRNA acetylation is located within moving cytidine, and their translation efficacy is greatly increased.
Experimental and computational techniques
According to numerous studies, ac4C is a major mutation that is critical for posttranscriptional control, mRNA stability, translational efficiency, and regulation of human immune function [10, 15–18]. Consequently, it is crucial to identify ac4C to comprehend its control mechanism. Ac4C sites can be identified using various biochemical methods, including single-molecule real-time sequencing, reduced-representation bisulfite sequencing, and mass spectrometry. When applied to large sequencing data, these methods are very costly, even if they are very useful in identifying ac4C sites.
Reduced representation Bisulfite sequencing
The genome-wide acetylcytidine profiles have been examined at the single-nucleotide level using reduced representation bisulfite sequencing, an effective and high-throughput method. This method enriches genomic areas with high CpG content by combining bisulfite sequencing with acetylcytidine-independent restriction enzymes. As a result, only 1% of the genome’s nucleotides may need to be sequenced. The bulk of promoters and repetitive sequences that are challenging to profile using traditional bisulfite sequencing techniques are still present in the fragments that make up the truncated genome [19–21]. Figure 2A displays the reduced representation bisulfite sequencing flowchart.
Figure 2.
Traditional experimental techniques to detect N4-acetylcytidine. Reduced representation bisulfite sequencing (A). Single-molecule real-time sequencing (B). Mass spectrometry (C).
Single-molecule real-time sequencing
A concurrent method for sequencing single-molecule DNA is called SMRTS. Strand displacement amplification (SDA) and multiple displacement amplification (MDA), which are based on rolling circle amplification (RCA), use unique loop adapters to produce ssDNA from dsDNA fragments. The DNA polymerase then adds dNTP containing fluorescent phosphate groups, which cleave the phosphate chain and emit light, thereby eliminating the fluorescent dye from the expanding nucleic acid chain. The sequencing reads are produced by parallel single-molecule real-time sequencing processes on thousands of nano-photonic visualization chambers [20, 22–25]. Figure 2B depicts the entire single-molecule real-time sequencing method.
Mass spectrometry
Mass spectrometry is an analytical technique that measures the mass-to-charge ratio (m/z) of several molecules in a sample. This technique combines MS detection of the non-enzymatic hydrolysates of the target DNA/RNA with selective capture of the DNA/RNA target from restricted cleavage of genomic DNA/RNA using magnetic separation. Figure 2C illustrates how the mass spectrometer operates. Furthermore, this technique has a special benefit in precisely measuring the quantity of RNA acetylcytidine. Because of its high selectivity, high resolution, and high mass accuracy, liquid chromatography in conjunction with high-resolution mass spectrometry has been widely employed for the quantitative analysis of bio-samples. This method has been effectively applied in several labs over the last 10 years to determine the genome-wide ac4C detection in various animals [26–28].
Computational techniques
Meanwhile, wet-lab methods are costly and less effective in identifying ac4C modification sites. Consequently, bioinformatics techniques using sequence information and artificial intelligence-based algorithms are much needed to correctly recognize the ac4C modification sites [29–32]. Figure 3 shows the machine and deep learning-based methods that have been developed over the past 7 years to predict human ac4C alteration sites.
Figure 3.
Timeline of historical landmarks for the prediction of ac4C modification sites.
The establishment of the previous methods consists of four crucial steps such as dataset construction, feature engineering, model development, and model evaluation. The whole steps for the construction of an artificial intelligence-based model are shown in Fig. 4.
Figure 4.
Building a suitable tool for ac4C prediction involves the following steps: dataset construction, feature engineering, model development, and model evaluation.
In this brief review, we introduce machine learning (ML), deep learning (DL), and large language model (LLM)-based tools for the identification of ac4C modification sites in the human mRNA transcriptome. In 2019, Zhao et al. [33] established the first random forest-based predictor to predict ac4C in the human mRNA transcriptome called PACES, in which they used six types of feature encoding schemes: one-hot, position-specific nucleotide sequence profile, k-spaced nucleotide pair frequencies, position-specific di-nucleotide sequence profile, k-nucleotide frequencies, and pseudo k-tuple nucleotide composition, and input them into a random forest-based classifier to predict ac4C. In 2020, Alam et al. [34] developed an extreme gradient boosting-based classifier to predict ac4C called XG-ac4C. In their technique, they used four types of encoding schemes, one-hot, electron-ion interaction of pseudopotentials (EIIP + PseEIIP), k-mer, nucleotide chemical property, and nucleotide density, and inserted them into five ML-based classifiers, namely, random forest, Gaussian naïve bayes, logistic regression, adaboost, and extreme gradient boost, and finally selected the best performing model to classify ac4C from non-ac4C. In 2021, Wang et al. [35] developed a convolutional neural network (CNN)-based method to predict ac4C. In their method, they used six physicochemical feature descriptors, namely, k-mer, composition of k-spaced nucleotide pairs (CKSNAP), series correlation pseudo di-nucleotide composition (SCPseDNC), series correlation pseudo tri-nucleotide composition, pseudo k-tuple composition (SCPseTNC), electron-ion interaction of pseudopotentials (PseEIIP) of tri-nucleotide with word2vector embedding. At first, six physicochemical features were improved with a support vector machine and F-score with a sequential forward search strategy. Then, these improved features were fed into the CNN for training and testing the model. In 2022, Zhang et al. [36] introduced a CNNLSTMac4CPred method for the prediction of ac4C modification sites. In their method, they utilized three types of features. The two feature descriptors are traditional, namely, pseudo tri-nucleotide composition (PseTNC) and k-nucleotide frequencies (KNF), and the other one is a semantic feature. They utilized CNN and LSTM-based neural networks to mine semantic features from sequences. After this, these features were inserted into an extreme gradient boosting-based classifier to predict ac4C modification. In 2023, a total of six methods were developed by authors from different countries. Jia et al. [37] developed a DLC-ac4C method in which they utilized three types of feature encodings, namely, nucleotide chemical property (NCP), nucleotide density (ND), and C2 encoding. They used 1D-CNN and BiLSTM for learning the local and global features, respectively. They also utilized a channel attention mechanism to find the optimal sequence characteristics and a homo-morphemic integration technique to limit the generalization error of the model. Lou et al. [38] constructed a stacking-ac4C method to classify N4-acetylcytidine in human mRNA. In their method, they used three types of encoders, namely, k-mer, pseudo-k-tuple nucleotide composition (PseKNC), and electron–ion interaction pseudo-potential (PseEIIP), and their hybrid fusion features, and then developed a stacking model to train and test the model. Liu et al. [39] presented a new model, TransC-ac4C, to predict ac4C modification sites in mRNA. In their model, they utilized five types of feature descriptors, namely, one-hot encoding, ND, k-mer, NCP, and EIIP, and inserted them into a hybrid model that consists of CNN and transformer. This technique overcomes the issues of missing long sequence dependence and avoids the loss of correlation information between discontinuous sequences. Su et al. [40] developed an iRNA-ac4C to classify ac4C from non-ac4C in human mRNA. In their methodology, they used hybrid features of k-mer, accumulated nucleotide frequency, and nucleotide chemical frequency, and then improved these features with minimum redundancy and maximum relevance (mRMR) and incremental feature selection (IFS) technique. Finally, they input these optimal features into seven different traditional ML classifiers, namely, logistic regression (LR), naive bayes (NB), gradient boost decision tree (GBDT), k-nearest neighbor (KNN), support vector machine (SVM), random forest (RF), and AdaBoost (AB) and select the best performing model (GBDT) to classify the ac4C from non-ac4C. Lai et al. [41] introduced a new tool, LSA-ac4C, for the accurate prediction of ac4C modification sites in human mRNA. In this technique, they used a hybrid neural network that combines an LSTM double-layer neural network with self-attention for the prediction of ac4C modification. Jia et al. [42] developed a new model, EMDL-ac4C. In this technique, they used one-hot encoding and then inserted it into an ensemble model with a two-branch residual connection dense net and attention to predict ac4C modification sites in human mRNA.
In 2024, a total of eight methods were developed by authors from different countries. Li et al. [43] introduced a MetaAc4C model in which they utilized pretrained bi-directional encoder representations (BERT), and the model is based on the bi-directional long short-term memory (BLSTM) network. Li et al. [44] also introduced a new method based on generative adversarial networks and transfer learning to recognize ac4C modification sites in human mRNA, named GANSamples-ac4C. Pham et al. [45] constructed a new tool, named ac4C-AFL, which is based on adaptive feature representation learning. In their method, they utilized 16 types of feature encoding schemes k-spaces nucleic acid pair (CKSNAP), enhanced nucleic acid composition (ENAC), position specific of 2 nucleotides (PS2), PseEIIP, Z-curve, k-mer, reverse k-mer complement (RCKmer), di-nucleotide physicochemical properties type 1-2 (DPCP1-2), NCP, binary features, multivariate mutual information and accumulated nucleotide frequency (MMNF), word2vec (W2V), sequence2vector (S2V), DNABERT and combination of skip di-nucleotide composition and local position-specific di-nucleotide frequency (ASLPN) and then inputted these optimized features into 11 different types of ML and DL classifiers, namely, SVM, RF, extremely randomized tree (ERT), artificial neural network (ANN), logistic regression (LR), GBT, XGBT, light GBT (gradient boost decision tree), cat boost (CB), adaboost (AB), and CNN classifiers and generated 176 baseline models. Finally, they selected the best baseline models using a 2-step feature selection technique, whose predicting scores were integrated and trained with SVM to develop the final model to classify ac4C from non-ac4C in human mRNA. Yi et al. [46] developed a new method, named STM-ac4C, which is a hybrid model based on selective kernel convolutions, a temporal convolutional network with multi-head self-attention to predict ac4C modification sites in mRNA. Yuan et al. [47] developed a new method based on a dual path neural network with self-attention, namely, DPNN-ac4C, to discriminate ac4C from non-ac4C in human mRNA. Their method integrates a convolutional neural network, embedding modules, bi-directional gated recurrent unit (BGRU) with self-attention to extract local and global features from mRNA sequences. Liu et al. [48] developed a novel transformer-based model called TransAC4C to predict ac4C modification sites in mRNA. Their model was divided into four parts such as a transformer layer, one BLSTM layer, four 1D convolutional layers, with three fully connected layers. Transformer and BLSTM layers were used to enable the model to learn the contextual information, convolutional layers were used to enable the model to extract important features from the input, and fully connected layers were used to connect the input features into the output. Jia et al. [49] constructed a new method called Voting-ac4C to predict ac4C modification sites in mRNA. In their method, they utilized RNAErnie, a transformer-based pretrained model with six traditional feature encoding schemes, such as one-hot, ENAC, ND, C2 encoding, k-spaced nucleotide pair frequencies (KSNPF), and physicochemical properties (TPCP), and then inserted the hybrid of these features into a deep neural network (DNN) for dimensionality reduction. Finally, these dimensionality reduction features were fed into a voting ensemble model constructed using CatBoost, XGBoost, and a multilayer perceptron (MLP) classifier. He et al. [50] developed a deep learning method called NBCR-ac4C based on pretrained models to predict ac4C modification sites in human mRNA. They utilized DNABERT2 and nucleotide transformer to construct contextual embedding of nucleotide sequences and then applied CNN and ResNet18 to further mine the trivial and deep information from the contextual embedding.
In 2025, a total of two predictors were proposed to date by the authors from different countries. Lu et al. [51] introduced an ERNIE-ac4C method to predict ac4C modification sites. In their technique, they used the pretrained model ERNIE-RNA to extract the attention map features and sequence features from nucleotide sequences. They input these fused features into a 2d CNN to predict ac4C modification sites in human mRNA. Yao et al. [52] proposed a DL-based method called Caps-ac4C to predict ac4C modification sites in human mRNA. In their method, they utilized chaos game representation (CGR) encoding to transform the nucleotide sequences into visual representations and then employed a capsule network design to extract the local and global features from these visual representations of nucleotide sequences. The brief description of feature extraction and selection, model evaluation techniques, and web server availability of the prediction methods are summarized in Table 1.
Table 1.
A list of available methods for the prediction of ac4C is summarized in this review.
| Method | Classifiera | Featureb | Evaluation | Webserver | Status | Year |
|---|---|---|---|---|---|---|
| PACES [33] | RF | One-hot, PSNSP, PSDSP, KNF, KSNPF, PseKNC |
5-fold CV | Yes | Active | 2019 |
| XG-ac4C [34] | XGB | One-hot, NCP-ND, k-mer, EIIP-PseEIIP | 5-fold CV | Yes | Not-active | 2020 |
| DeepAc4C [35] | CNN | k-mer, CKSNAP, SCPseDNC, SCPseTNC, PseKNC, PseEIIP, W2V | 10-fold CV | Yes | Not-active | 2021 |
| CNNLSTMac4CPred [36] | XGB | Semantic features, KNF, PseTNC |
5-fold CV | Yes | Not-active | 2022 |
| DLC-ac4C [37] | EL | C2, NCP, ND | 10-fold CV | No | – | 2023 |
| Stacking-ac4C [38] | SIC | k-mer, PseEIIP, PseKNC | 10-fold CV | No | – | 2023 |
| TransC-ac4C [39] | CNN + Transformer |
One-hot, NCP, ND, EIIP, k-mer | 5-fold CV | No | – | 2023 |
| iRNA-ac4C [40] | GBDT | k-mer, NCP, ANF | 10-fold CV | Yes | Not-active | 2023 |
| LSA-ac4C [41] | Double-layer LSTM + Self-attention |
Embedding | 10-fold CV | Yes | Not-active | 2023 |
| EMDL-ac4C [42] | EM + DL | One-hot | 10-fold CV | No | – | 2023 |
| MetaAc4C [43] | LSTM + Attention + Residual |
BERT embedding + GAN |
10-fold CV | No | – | 2024 |
| GANSamples-ac4C [44] | GAN | One-hot, ANF, NCP, CKSNAP, ENAC, NAC, EIIP, k-mer, RCKmer |
5-fold CV | No | – | 2024 |
| ac4C-AFL [45] | SVM | ENAC, PS2, CKSNAP, DNABERT, PseEIIP, Z-curve, S2V, k-mer, W2V, RCK-mer, DPCP_1, DPCP_2, NCP, BPF, MMNF, ASLPN |
10-fold CV | Yes | Active | 2024 |
| STM-ac4C [46] | SKC + TCN + MHSA | One-hot | 10-fold CV | No | – | 2024 |
| DPNN-ac4C [47] | (Bi-directional GRU + Attention) + (CNN + self-attention) | PseKNC, IES | 10-fold CV | No | – | 2024 |
| TransAC4C [48] | Transformer | Embeddings | 10-fold CV | No | – | 2024 |
| Voting-ac4C [49] | EL | RNAErnie + One-hot + ND + C2 + ENAC + TPCP + KSNPF | 10-fold CV | Yes | Not-active | 2024 |
| NBCR-ac4C [50] | CNN + Resnet18 | Nucleotide Transformer embedding + DNABERT2 embedding | 10-fold CV | No | – | 2024 |
| ERNIE-ac4C [51] | 2d-CNN | Bert + AMF | 10-fold CV | No | – | 2025 |
| Caps-ac4C [52] | CN | CGR | 5-fold CV | Yes | Active | 2025 |
aSIC, Stacking Integration Classifier; EL, ensemble learning; GAN, generative adversarial network; SVM, support vector machine; CNN, convolutional neural network; XGB, eXtreme gradient boosting; GBDT, gradient boosting decision tree; RF, random forest; LSTM, long short-term memory; EL, ensemble learning; EM, ensemble; DL, deep learning; SKC, selective kernel convolution; TCN, temporal convolutional networks; MHSA, multi-head self-attention mechanism; CN, capsule networks.
bBERT, bidirectional encoder representations from transformers; GAN, generative adversarial networks; W2V, Word2vector; CKSNAP, composition of k-spaced nucleic acid pairs; PseTNC, Pseudo tri-nucleotide composition; ENAC, enhanced nucleic acid composition; NAC, nucleic acid composition; NCP, nucleotide chemical property; PseKNC, Pseudo k-tuple nucleotide composition; PseEIIP, electron-ion interaction pseudopotentials; RCK-mer, reverse complement k-mer composition; ND, nucleotide density; S2V, sequence2vector; ANF, accumulated nucleotide frequency; MMNF, multivariate mutual nucleotide frequency; DPCP, di-nucleotide physicochemical properties; ASLPN, adaptive skipped local position-specific di-nucleotide composition; SCPseDNC, series correlation pseudo di-nucleotide composition; SCPseTNC, series correlation pseudo tri-nucleotide composition; AMF, attention map feature; IES, integer encoding of sequence; CGR, chaos game representation; ERNIE, enhanced representation through knowledge integration; ND, nucleotide density; C2, common sequence characterization; TPCP, physicochemical properties; KSNPF, k-spaced nucleotide pair frequencies.
Several tools were developed for ac4C alteration sites employing the aforementioned techniques. The development of computational biology and bioinformatics will benefit from the substantial evidence that each tool provides for the identification of ac4C modification sites in human mRNA. Several elements of these tools are summarized in this study, including model evaluation, ML, DL, and LLM-based classifiers, feature extraction and selection techniques, and dataset development. As a result, this evaluation can help researchers select the best tool for ac4C prediction and offer recommendations for future advanced AI-based methods.
Dataset construction
The construction of the dataset is the first important step in the development of a bioinformatics tool. Nowadays, a large amount of data related to RNA is publicly available in open-source databases such as Rfam [53–55] and RMBase3.0 [56]. Many credible datasets were also produced from previously published literature and databases to develop ML, DL, and LLM models. Therefore, we presented some simple steps that have been widely utilized to build the datasets for ac4C modification prediction in mRNA. (i) Extracting sequence-based data from previous literature and databases. (ii) Applied cluster database at high identity with tolerance to eradicate similarity between the sequences. (iii) Applied labeling on positive and negative sequences, such as 1 for positive ac4C and 0 for negative ac4C. We have observed that all the training data for ac4C modification sites were retrieved from previous literature. Arango et al. [14] observed the distribution of various CXX motifs both inside and outside the acetylation peak; they selected positive and negative sequences from 2134 genes from previous literature. They varied the number of repetitions of CXX from 2 to 9. Simple repeating motifs CXXCXX could be frequently observed in the transcriptome; they discovered five CXX motifs occurred in 1629 peaks, and 15 198 five consecutive CXX motifs found outside the peak. Finally, they divided the dataset into two parts: 1160 ac4C sequences and 10 855 non-ac4C sequences for training purposes, and 469 ac4C sequences and 4343 non-ac4C sequences for independent testing. The dataset can be downloaded at http://www.rnanut.net/paces. After this, Wang et al. [35] constructed a balanced dataset with a 40% similarity threshold and consisted of 1148 sequences of ac4C and 1148 sequences of non-ac4C for training purposes and 467 sequences of ac4C and 467 sequences of non-ac4C for independent testing. The dataset can be downloaded at https://zenodo.org/records/5138047. Su et al. [40] also constructed a new balanced dataset based on Arango et al. [14] data. They picked the cytidines nearest to ac4C peaks as alteration sites in order to create a trustworthy dataset. They then collected 100 nucleotides on either side of these alteration sites as positive samples, using them as the center. Negative samples from non-peak areas were chosen at random. The central sites are cytidine, and their sequences were 201 nucleotides. In their dataset, they set an 80% similarity threshold to remove the redundant sequences. Finally, they retrieved 2206 ac4C sequences and 2206 non-ac4C sequences for training and 552 ac4C sequences and 552 non-ac4C sequences for independent testing. The dataset can be downloaded at http://lin-group.cn/server/iRNA-ac4C. Lu et al. [51] built a new dataset of 1480 ac4C sequences and 1480 non-ac4C sequences for training and 185 ac4C sequences and 185 non-ac4C sequences for independent testing. Due to the lack of dataset details, this dataset was excluded from the current evaluation. Therefore, we have selected the datasets of Su et al. [40], Wang et al. [35], and Arango et al. [14] in this review to assess and validate the prediction efficiency of the computational tools. Table 2 displays all the dataset details.
Table 2.
A benchmark classification data for ac4C.
| Method | Ratio (P/N) |
Training set (P/N) | Independent set (P/N) | Data-source | CD-HIT | Year |
|---|---|---|---|---|---|---|
| PACES [33] | 1:10 | 1160/10855 | 469/4343 | Arango et al. | – | 2019 |
| XG-ac4C [34] | 1:10 | 1160/10855 | 469/4343 | Arango et al. | – | 2020 |
| DeepAc4C [35] | 1:1 | 1148/1148 | 467/467 | Wang et al. | 40% | 2021 |
| CNNLSTMac4CPred [36] | 1:10 | 1160/10855 | 469/4343 | Arango et al. | – | 2022 |
| DLC-ac4C [37] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2023 |
| Stacking-ac4C [38] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2023 |
| TransC-ac4C [39] | 1:1 | 1148/1148 | 467/467 | Wang et al. | 40% | 2023 |
| iRNA-ac4C [40] | 1:1 | 2206/2206 | 552/552 | Arango et al. | 80% | 2023 |
| LSA-ac4C [41] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2023 |
| EMDL-ac4C [42] | 1:1 | 1148/1148 | 467/467 | Wang et al. | 40% | 2023 |
| MetaAc4C [43] | 1:1 | 1148/1148 | 467/467 | Wang et al. | 40% | 2024 |
| GANSamples-ac4C [44] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| ac4C-AFL [45] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| STM-ac4C [46] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| DPNN-ac4C [47] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| TransAC4C [48] | 1:1 | 1148/1148 | 467/467 | Wang et al. | 40% | 2024 |
| Voting-ac4C [49] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| NBCR-ac4C [50] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2024 |
| ERNIE-ac4C [51] | 1:1 | 1480/1480 | 185/185 | Lu et al. | – | 2025 |
| Caps-ac4C [52] | 1:1 | 2206/2206 | 552/552 | Su et al. | 80% | 2025 |
“P” denotes ac4C and “N” denotes non-ac4C.
Feature engineering
Extracting and selecting the self-directed and informative feature is an essential phase in ML and DL-based methods [57–62]. Therefore, several types of feature extraction and selection methods were employed to explain the ac4C modification in human mRNA.
Feature extraction
k-mer
The short-range nucleotide interactions between sequences can be represented by a k-mer. Using a sliding window method, the (N-k + 1) nucleotide residues can be found by examining a sequence with N bp and setting the size of the window to k bp with 1 bp step size [63–65]. With a sequence length of N, an arbitrary sample M can be described as:
![]() |
(1) |
where the nucleotide (A, G, U, and C) at position i is denoted by Ri. Using the k-mer nucleotide composition, the sequences can be converted into the 4k-D vector as follows:
![]() |
(2) |
where T represents the vector’s transposition and f1k-tuple characterizes the sequence occurrence of the i-th k-mer nucleotide composition. An RNA sample can be decoded into a 4-D vector M1 = [f(A), f(G), f(C), f(U)]T when k = 1. A 16-dimensional vector can be used to describe the RNA sample when k = 2.
CKSNAP
The frequency of nucleotide pairings split apart by any k nucleotide (k = 0, 1, 2, 3, 4, 5) is represented by the CKSNAP. Sixteen nucleotide pairs [AA, AG, …,UG, UU] make up the features of k-spaced nucleic acid pairings. Using k = 1 as an example, the k-spaced nucleic acid pair composition can be expressed as follows:
![]() |
(3) |
where * denotes (A, C, G, and U), NTotal represents the total number of single-spaced nucleotide pairs in the sequence, and NY*Z
indicates the nucleotides Y*Z pairings length in the sequence. The composition of the nucleic acid pair AA can be calculated by dividing its frequency (j) by the total number of 0-spaced nucleic acid pairs (NTotal) in the nucleotide sequence. The value of NTotal for a nucleotide sequence of length “g” is g-1, g-2, g-3, g-4, g-5, and g-6 for (k = 0, 1, 2, 3, 4, 5) [58, 66–68].
ENAC
Based on the sequence window, the ENAC evaluates the nucleic acid composition. A sequence of equal length can be created with it. The ENAC can be calculated as:
![]() |
(4) |
Equation (4) uses k to describe the sliding window’s size, NA, win to indicate how many nucleotides A are present in the window g (g = 1, 2, …, L-k + 1), and T
[G, C, A, U].
NCP
Based on their characteristics, four types of nucleotides in RNA sequence fragments can be divided into three groups, which are hydrogen bonds, ring structure, and functional groups. Numerous nucleotide modification site prediction applications have made extensive use of this classification. More precisely, A and G belong to the keto group, whereas A and U belong to the amino group; C and G can form strong hydrogen bonds, while A and U can form weak ones; and C and U have single-ring structures, while Adenine and Guanine both have double-ring structures. Therefore, coordinates (1,1,1), (0,1,0), (1,0,0), and (0,0,1) can be used to represent A, U, G, and C based on these chemical properties, respectively. Since cytidine is the core site of all sequences, we only retrieved NCP information from the adjacent sequences of ac4C peaks since it is not required to obtain NCP features of this site [69].
ANF
The distribution of every nucleotide in the RNA sequence, as well as nucleotide frequency information, is provided by ANF [70]. In an RNA sequence, the density di of any nucleotide si at position i is stated as follows:
![]() |
(5) |
where l is the entire length of the sequence and
is the size of the i-th preface sequence from the 1st to the i-th position.
One-hot encoding
In sequence prediction tasks, such as RNA sequence analysis, one-hot encoding is a prominent technique. The four nucleotides that make up the language of RNA molecules are adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleotides are mapped to the binary vectors (1, 0, 0, 0), (0, 0, 1, 0), (0, 1, 0, 0), and (0, 0, 0, 1), respectively [71]. The resulting one-dimensional one-hot encoded representation for an RNA sequence of length L would be shaped like (1, L × 4). As an alternative, the shape of (4, L) is produced via the 2D one-hot encoding.
CGR
CGR encoding, which has been extensively used to represent biological sequences in DNA, RNA, and proteins to transform sequence data into illustration in a 2D space. Sequences are represented by CGR encoding as a collection of points in a 2D space, with each point’s location and distribution reflecting the sequence’s properties [72]. In addition to being intuitive, this representation can capture intricate patterns that are present throughout the sequence.
IES
RNA sequence bases are mapped to integers by integer encoding. Integer encoding efficiently examines the contextual relationships inside sequences when combined with a bidirectional GRU network. Notably, it enhances the GRU network’s sequence analysis skills by naturally representing the positional order correlations between bases [73]. A basic dictionary with discrete integer values assigned to each symbol is defined as the part of the implementation process. Let M = S1, S2, S3,….SL represents the RNA sequence M with L bases, where Si = {A, U, C, G}. Through a methodical mapping of A to 0, C to 1, G to 2, and U to 3.
RCKmer
A particular kind of k-mer known as RCKmer is created by eliminating the complementary reverse k-mers for every kind of k-neighboring nucleic acid in the specified sequence [74]. For example, when k = 2, the k-mers “GG,” “CG,” “GU,” “UG,” “UC,” and “UU” should be eliminated, leaving 10 discriminatively surviving k-mers in the RCKmer.
PseEIIP
Tri-nucleotides in the mRNA sequence are assigned PseEIIP according to an average EIIP value for each nucleotide [75].
![]() |
(6) |
where flmn is the standardized incidence of the i-th tri-nucleotide, EIIPlmn = EIIPl + EIIPm + EIIPn, and l, m, n ∈ {A, G, C, U}.
DNABERT
Based on the relative interactions among k-mers in the input sequence, DNABERT extracts sequence information. In order to efficiently examine the complex relationships between nucleotides in both forward and reverse orientations, the DNABERT model employs the capabilities of a transformer architecture. By using a k-mer-based tokenizer, DNABERT creates unique tokens from genomic sequences known as “k-mers,” which are further processed through 12 transformer-based blocks. These blocks produce detailed representations of the genomic sequence that capture the appropriate information contained in the inherent material by analyzing the interaction between k-mers [76–78]. Since RNA and DNA both differ by only one nucleotide and retain nearly the same genetic information, DNABERT has notably demonstrated relevance to RNA. This enables handling of RNA sequences by merely substituting nucleotide “T” for nucleotide “U”.
W2V
By taking into account the background in which words occur within a sizable quantity of text, W2V is a neural network-driven model that effectively encodes linguistic and semantic word associations. Both Continuous Bag-of-Words (CBOW) and Skip-gram methods are included in the Word2Vec model. While Skip-gram estimates environmental contextual words for a specific targeted word, CBOW anticipates the desired word using its context words. Both models are trained on extensive textual data sets and used separately within the Word2Vec architecture [79]. That makes it possible to transform words into vector illustrations that express semantic connections across the entire vocabulary.
Feature selection
Selecting the features is a vital step after obtaining the meaningful features. The ideal features could enhance the model’s performance vividly. Presently, there are many feature selection methods used in computational biology, such as IG, ANOVA, mRMR, and PCA [80–83]. The mathematical manifestation is given below.
Binomial distribution
BD is a unique feature selection methodology that has been used in sequence analysis. In this methodology, the confidence level of k-mer’s is termed as follows:
![]() |
(7) |
where
is the confidence level for m and n, and “n” signifies the positive/negative sample,
represents the m-th residue, and
symbolizes the occurrence of n-th residue in the sample [84, 85].
Principal component analysis
PCA is a feature dimensionality reduction method and is often utilized in many ML, DL, and data science applications to visualize and analyze the data. It helps to reduce the irrelevant features from the data by keeping the important and meaningful information. It uses linear algebra to convert features into principal components by calculating the eigenvectors and eigenvalues. Finally, it selects the top components with top eigenvalues [86–88].
Analysis of variance
The main purpose of ANOVA is to calculate the fundamental distinctions across two or more groups. The F-score value is used in this process to rate each feature. The percentage of variance within the groups and the difference between groups is known as the F-score. As a result, the characteristic with a high F-score has strong discriminating ability [89].
![]() |
(8) |
![]() |
(9) |
The variance between the values is represented by
, and the degree of freedom between them is represented by
. The variance within the values is represented by
, and the degree of freedom within values is represented by
.
![]() |
(10) |
where
is the computational dispersal score of alterations among treatments
and
.
Incremental feature selection
IFS categorized all the features based on the
-values and produced new feature vectors, which are given in eq. (11).
![]() |
(11) |
The feature with the highest δ-value,
is included in the first feature subset. The subsequent feature subset,
is generated by adding the second feature subset with the highest δ-value. The third feature subset,
is generated by integrating the third feature subset with the highest δ-value. Until every feature was taken into account, the procedure was repeated [90].
Machine learning algorithms
The prediction of ac4C alteration depends heavily on the choice of an appropriate machine learning algorithm. Bioinformatics has made use of several machine learning methods, including support vector machines, adaboost, extreme randomized trees, random forest, naive bayes, gradient boost decision trees, and eXtreme gradient boosting.
Support vector machine
In computational biology, SVM is a popular and significant machine learning algorithm. Finding the optimal hyperplane that optimizes the separation between the neighboring samples from the classes and the hyperplane is the main goal of SVM [91–94]. The SVM algorithm’s mathematical representation is provided as follows:
If
is a training incidence and
is a target, then
![]() |
(12) |
where “p” indicates the bias and
indicates the kernel function. To find a hyperplane, the kernel function can plot the feature vectors in three dimensions. Nonlinear classification issues are the main application for the radial basis function. It might be stated as:
![]() |
(13) |
whereas,
can be obtained by using the maximum Lagrangian expression as:
![]() |
(14) |
where (∴
= 0, 0 ≤
≤ k, k symbolizes the significance constraint).
AdaBoost
Another well-known and extensively utilized tool in bioinformatics is the AdaBoost classifier. It is an ensemble method that improves accuracy by combining different classifiers. Setting the classifier’s weights and training the data in each iteration is the basic notion behind this [95]. It is defined as follows:
![]() |
(15) |
where R is the target, which can be either −1 or 1, j is the data point, and x is the dimension number. AdaBoost determines the significance of each training example in the training dataset by giving it a weight. That set of training data points is likely to have a larger in-training set when the assigned weights are high. In a similar vein, low allocated weights have little effect on the training dataset.
Extreme randomized tree
ERT is a well-known method that utilizes hundreds of decision trees to perform classification. This strategy was initially introduced by Geurts et al. [96] to reduce the model’s discrepancy by incorporating strong randomization techniques. ERT incorporates more robust randomization strategies to further reduce the prediction model’s volatility. With a few small adjustments, like not using the bagging technique to build a tree, ERT is extremely similar to the random forest algorithm. While random forest finds the optimal split by enhancing the variable index between random subclasses of variables and splitting value [97]. It is defined as follows:
![]() |
(16) |
where ϵ is the random variable that encapsulates the random disconcerting scheme, denoted by ϵ, a value of this variable, and hence by
. The learning sample is epitomized by ls, and N is the total number of models.
Random forest
A combined knowledge method that is widely used in computational biology is the random forest. It is utilized for clustering, classification, and regression problems, as well as feature selection processes. It is helpful to reduce the bias of the individual weak predictor, avoid variable rescaling, and be more robust to noise and overfitting. It is more effective to handle class imbalance problems by allocating weight to the minority class, and it efficiently handles missing data issues by substituting the variable appearing mostly in a particular node. The basic principle is to integrate numerous weak classifiers [98–100]. Since the voting process determines the conclusion, the model’s result is more precise and straightforward. If the attribute
of a vector
and weak classifiers based on these attributes is
, then it can be demarcated as:
![]() |
(17) |
Naïve Bayes
The NB model is easy to build and effective when dealing with large datasets. This method makes use of direct acyclic graphs, which have many leaves and only one root. When leaf nodes are separated from their base, it presumes stability between them [101]. If the class variable is “v” and the feature vector is
:
![]() |
(18) |
By means of independent norms
![]() |
(19) |
After the simplification
![]() |
(20) |
eXtreme gradient boosting
For numerous computational issues, XGBoost, an ensemble learning technique based on gradient boosting, offers cutting-edge outcomes. In essence, XGBoost is an ensemble technique built on a gradient boosted tree [102–105]. According to the formula below, the prediction’s outcome is the total of the scores predicted by k trees:
![]() |
(21) |
where
is the i-th number in the training set, and
is the k-th tree score. S is the interplanetary function encompassing all gradient boosted trees. The following formula could be used to optimize the objective function:
![]() |
(22) |
where a loss function that assesses the fitness of model prediction ^yi and samples of the training dataset yi is represented by the former
. A regularization item that penalizes the model’s complication to prevent overfitting is represented by
.
Gradient boosting decision tree
Researchers have utilized the gradient-boosted decision tree algorithm, a highly significant learning technique, across multiple bioinformatics, mathematics, and biological applications [106–108]. From a nonlinear joint of various weak learners, it builds a realistic and climbable model. The gradient boost decision tree’s primary goal is to create a base learner that is highly correlated with the negative gradient loss function. Assume that n samples are available:
![]() |
![]() |
(23) |
where
is the risk optimization factor of the newly generated decision tree, as indicated by the equation below, and
is the new decision tree (k = 1,2,3,…).
![]() |
(24) |
The final evaluation is computed in an anticipatory mode by the gradient boost decision tree technique.
![]() |
(25) |
Subsequently, residuals are calculated using the loss function
of negative gradient.
![]() |
(26) |
Ultimately, all
to calculate the risk mitigation parameter
. were used to train the model. By mapping the values in the input space
into
split parts
, and producing
for region
, this kind of decision tree intuitively models the relationships among predictor variables.
![]() |
(27) |
Deep learning algorithms
Convolutional neural network
Convolutional, normalizing, connected, pooling, and other layers make up CNN. It has been utilized in many biological and medical advancements. The fundamental concept of CNN is to create an enormous number of filters that can obtain veiled topological features from input using pooling techniques and layer-wise convolutions [82, 109].
![]() |
(28) |
where “t” is the filter index and “b” is the input. The input channel numbers are represented by “n,” the filter size is indicated by “r,” the output filter is represented by “g,” and the convolutional filter with a weight matrix “r × I” is represented by “St”.
![]() |
(29) |
Equation (29) is the symbol of the reluctant activation function consuming b as an input.
![]() |
(30) |
Equation (30) shows a fully connected layer in which “
” represents the bias, “
” represents the dropout operator with probability of “p,” “
” represents the 1D feature vector, and “
” characterizes the prior layer weights of “
”.
![]() |
(31) |
Equation (31) suggests the final classification layer has a sigmoid activation function and “b” as an input.
Long short-term memory network
LSTM is a unique variant of a recurrent neural network (RNN) that incorporates a storage method made up of frequently linked blocks, replacing the hidden function found in conventional RNNs. LSTM surpasses traditional RNN due to its interconnected memory and multiplication units that enhance its ability to learn from distant dependencies. This advancement has expanded the application scope of RNN and established a foundation for the evolution of the following sequence modeling techniques. In contrast to unidirectional LSTM, bi-directional LSTM (BLSTM) more effectively captures context information within sequences. The attention mechanism has garnered significant attention within the domains of natural language processing and image recognition to enhance model interpretability [110–112]. It enables the model to prioritize more critical information by differentiating between varying levels of importance.
Generative adversarial network
GANs have shown impressive accomplishments across various fields, such as generating natural images, translating images from one form to another, and creating high-resolution images. A GAN comprises two neural networks, which are a generator (G) and a discriminator (D). The generator produces synthetic data by utilizing random Gaussian noise (z) in an effort to fool the discriminator. Conversely, the discriminator’s role is to differentiate between the produced samples and the actual data [113–115]. The process is outlined in the following formulas:
![]() |
(32) |
where “
” denotes training samples drawn from the actual distribution, while “
” signifies the samples produced by the generator network. The distribution “
” is implicitly characterized by the equation
= G(z), where z is sampled from P(z), with P(z) representing the distribution of latent variables.
Multi-head self-attention mechanism
Sequence or word vectors are plotted to query (Q), key (K), and value (V) vectors in the attention mechanism. The relevance between Q and K is then determined by computing attention scores, usually using techniques like additive, dot product, scaled dot product, general, and concatenation. This approach enhances the model’s data representation and retrieval capabilities by enabling it to modify its attention based on various segments of the input sequence. By using the self-attention process as an example, word vectors multiplied by three different weight matrices (wq, wk, and wv) can yield Q, K, and V. The computation of Q and K’s similarity in scaled dot-product attention for a given word vector is:
![]() |
(33) |
where dk is the input vector’s dimension, and
indicates the correlation between various input word vectors. Researchers worked to improve and optimize the attention mechanism as its efficiency in processing sequence data became apparent. As a result, multi-head attention was introduced [116–118]. The attention operation is not applied simply once with multi-head attention. However, this process is carried out concurrently across several distinct representation spaces, or heads. Each “head” has a distinct weight, suggesting that they may concentrate on various aspects of the input data. Each head’s output is calculated in parallel before being combined into a single output.
![]() |
(34) |
![]() |
(35) |
is a concatenated weight matrix.
Large language models
BERT
Bidirectional encoder representations from transformers (BERT) is an efficient language representation model that was recently developed by Google Research and expands upon the transformer model. The main goal was to encode the input sequence as a phrase that continued to move up the stack of encoder layers. Three different forms of embeddings are used to improve each word’s input vector, such as segment, token, and position embeddings. By looking up the word in a predetermined vocabulary, token embedding is obtained. By differentiating between the first and second halves of the sentence, segment embeddings show which part of the sentence the word belongs to. Important details about the relative and absolute locations of each word in the phrase are captured by the positional embedding [119, 120]. The position vector is described as follows:
![]() |
(36) |
![]() |
(37) |
where pos is the position, “I” is the component point of the vector, and
embodies the dimension of the vector.
ERNIE
A very effective large-scale RNA pretrained language model is ERNIE-RNA. The improved BERT structure serves as its main foundation. In particular, it uses a multi-head attention technique for encoding and uses the same transformer encoder as BERT. Nevertheless, this approach incorporates prior information about related RNA sequences. It enables a more thorough extraction of RNA properties by introducing a base pair attention bias during the attention calculation, which assigns distinct bias values to various base pairs. The synergistic combination of contextualized embeddings and implicit structural learning derived from large-scale pretraining, rather than just the capture of long-range dependencies alone, is the main reason why ERNIE-RNA and DNABERT perform better for ac4C prediction than conventional k-mer approaches. Self-attention mechanisms are particularly good at simulating long-range sequence interactions. In fact, complicated higher-order patterns like GC-content biases, nucleosome placement, and the tendency of particular sequences to form particular secondary structures are implicitly encoded by pretraining. As a result, when the models are optimized for ac4C detection, they see more than just a local motif. Thus, they use this pretrained knowledge to comprehend the larger structural context and chemical environment (such as stability in high-GC regions) that affects ac4C modification, an understanding that static k-mer frequencies are unable to offer. Their remarkable performance in a variety of downstream activities related to RNA structure and functionality is demonstrated by experimental evaluations [121–124]. These findings highlight the model’s extraordinary effectiveness and efficiency in capturing information related to RNA features.
Prediction evaluation
Evaluating the prediction efficiency of a prediction model is crucial. To assess the prediction ability of ML-based models, three types of techniques are currently commonly used: the independent test, the jackknife test, and the k-fold CV test [125].
Evaluation metrics
Matthew’s correlation coefficients (MCC), accuracy (Acc), specificity (Sp), and sensitivity (Sn) were utilized to calculate the model’s effectiveness [126–129], defined as follows:
![]() |
(38) |
where “tn” denotes the correctly predicted non-ac4C sites, “fp” proposes non-ac4C sites categorized as ac4C sites, “tp” represents the real ac4C sites that were successfully predicted as ac4C sites, and “fn” shows the ac4C sites that were classified as non-ac4C sites. Additionally, the model’s efficacy and viability were illustrated using the receiver operating curve (ROC). To evaluate the prediction model’s effectiveness, the area under the curve (AUC) was calculated. AUC = 0.5 for random execution and AUC = 1 for the optimal prediction model [130–132].
Published results and discussions
Numerous computational models have been established by the researchers to predict ac4C modification sites in human mRNA. Table 3 and Fig. 5 present the full comparative results based on the training dataset.
Table 3.
Comparison of the published ac4C prediction methods on the training datasets.
| Method | Dataset | Evaluation | ACC | SN | SP | MCC | ROC | PRC |
|---|---|---|---|---|---|---|---|---|
| PACES [33] | Un-balanced | 5-fold CV | – | – | – | – | 0.885 | 0.559 |
| XG-ac4C [34] | Un-balanced | 5-fold CV | 0.921 | 0.597 | 0.956 | 0.552 | 0.910 | 0.653 |
| DeepAc4C [35] | Balanced | 10-fold CV | 0.824 | – | – | – | – | – |
| CNNLSTMac4CPred [36] | Un-balanced | 5-fold CV | 0.902 | 0.647 | 0.930 | 0.514 | 0.900 | – |
| DLC-ac4C [37] | Balanced | 10-fold CV | 0.800 | 0.831 | 0.772 | 0.606 | 0.877 | – |
| Stacking-ac4C [38] | Balanced | 10-fold CV | 0.888 | 0.892 | 0.884 | 0.774 | 0.954 | 0.953 |
| TransC-ac4C [39] | Balanced | 5-fold CV | 0.814 | 0.788 | 0.838 | 0.630 | 0.878 | – |
| iRNA-ac4C [40] | Balanced | 10-fold CV | 0.800 | 0.770 | 0.830 | 0.601 | 0.875 | – |
| LSA-ac4C [41] | Balanced | 10-fold CV | 0.820 | 0.855 | 0.785 | 0.643 | 0.879 | – |
| EMDL-ac4C [42] | Balanced | 10-fold CV | 0.892 | – | – | – | – | – |
| EMDL-ac4C [42] | Un-balanced | 5-fold CV | – | – | – | – | 0.904 | 0.615 |
| MetaAc4C [43] | Un-balanced | 10-fold CV | 0.937 | 0.912 | 0.974 | 0.875 | 0.986 | – |
| GANSamples-ac4C [44] | Balanced | 5-fold CV | 0.834 | 0.884 | 0.784 | 0.671 | 0.899 | – |
| ac4C-AFL [45] | Balanced | 10-fold CV | 0.833 | 0.857 | 0.810 | 0.668 | 0.903 | – |
| STM-ac4C [46] | Balanced | 10-fold CV | 0.825 | 0.864 | 0.786 | 0.654 | 0.885 | – |
| DPNN-ac4C [47] | Balanced | 10-fold CV | 0.838 | 0.821 | 0.856 | 0.665 | 0.927 | – |
| TransAC4C415 nt [48] | Balanced | 10-fold CV | 0.785 | – | – | – | 0.869 | – |
| TransAC4C21 nt [48] | Balanced | 10-fold CV | 0.923 | – | – | – | 0.977 | – |
| Voting-ac4C [49] | Balanced | 10-fold CV | 0.867 | 0.892 | 0.843 | 0.736 | 0.930 | – |
| NBCR-ac4C [50] | Balanced | 10-fold CV | 0.837 | 0.914 | 0.760 | 0.682 | – | – |
Figure 5.
Performance of the available ac4C modification sites prediction tools on balanced and un-balanced training datasets.
Zhao et al. [33] established the first random forest-based predictor in 2019, called PACES, in which they used six types of feature encoding schemes: one-hot, position-specific nucleotide sequence profile, position-specific di-nucleotide sequence profile, k-spaced nucleotide pair frequencies, k-nucleotide frequencies, and pseudo k-tuple nucleotide composition, and inputted into a random forest-based classifier with 5-fold CV and achieved a reasonable accuracy with an ROC value of 0.885. The performance on the unbalanced independent dataset was also reasonable, with an ROC value of 0.874. In 2020, Alam et al. [34] developed an extreme gradient boosting-based classifier to predict ac4C called XG-ac4C. In their technique, they used four types of encoding schemes, one-hot, EIIP + PseEIIP, k-mer, nucleotide chemical property, and nucleotide density, and inserted them into five ML-based classifiers, namely, random forest, Gaussian naïve bayes, logistic regression, adaboost, and extreme gradient boost by using 5-fold CV, and finally selected the best performing model to classify ac4C from non-ac4C. The accuracy of the best classifier was 0.921. The performance on unbalanced independent testing data was acceptable with an ROC value of 0.889. In 2021, Wang et al. [35] developed a CNN based method to predict ac4C. In their method, they used six physicochemical feature descriptors, namely, k-mer, CKSNAP, SCPseDNC, series correlation pseudo tri-nucleotide composition, SCPseTNC, PseEIIP of tri-nucleotide with word2vector embedding. At first, six physicochemical features were improved with a support vector machine and F-score with a sequential forward search strategy. Then these improved features were fed into the CNN by utilizing a 10-fold CV for training the model, and attained an accuracy of 0.824. The efficiency on an unbalanced independent dataset was good, with an accuracy score of 0.791. In 2022, Zhang et al. [36] introduced a CNNLSTMac4CPred method for the prediction of ac4C modification sites. In their method, they utilized three types of features. The two feature descriptors are traditional, namely, PseTNC and KNF, and the other one is a semantic feature. They utilized CNN and LSTM-based neural networks to extract semantic features from sequences. After this, these features were inserted into an extreme gradient boosting-based classifier by using 5-fold CV to train the model, and got an accuracy of 0.902. The accuracy score of CNNLSTMac4CPred on the unbalanced independent dataset was 0.872 with an ROC value of 0.882.
In 2023, Jia et al. [37] developed a DLC-ac4C method in which they utilized three types of feature encodings, namely NCP, ND, and C2 encoding. They used 1D-CNN and BiLSTM for learning the local and global features, respectively. They also utilized a channel attention mechanism to find the optimal sequence characteristics and a homo-morphemic integration technique to limit the generalization error of the model with an accuracy of 0.800. The accuracy score of DLC-ac4C on the balanced independent dataset was 0.829 with an ROC value of 0.903. Lou et al. [38] constructed a stacking-ac4C method to classify N4-acetylcytidine in human mRNA. In their method, they employed three types of encoders, namely, k-mer, PseKNC, and PseEIIP, and their hybrid fusion features, and then developed a stacking model on training data by utilizing 10-fold CV with an accuracy of 0.880. The accuracy score of stacking-ac4c on the balanced independent testing set was 0.808 with an ROC value of 0.883. Liu et al. [39] presented a new model, TransC-ac4C, to predict ac4C modification sites in mRNA. In their model, they utilized five types of feature descriptors, namely, one-hot encoding, ND, k-mer, NCP, and EIIP, and inserted them into a hybrid model that consists of CNN and transformer with 5-fold CV to train the model. Subsequently, they got an 0.814 accuracy score, and the performance on the balanced independent dataset was 0.806. Su et al. [40] developed an iRNA-ac4C to classify ac4C from non-ac4C in human mRNA. In their methodology, they utilized hybrid features of k-mer, accumulated nucleotide frequency, and nucleotide chemical frequency, and then improved these features with mRMR and IFS technique. Finally, they input these optimal features into seven different traditional ML classifiers, namely, LR, NB, GBDT, KNN, SVM, RF, and AB, and select the best performing model GBDT with 10-fold CV to classify the ac4C with an accuracy of 0.800. The performance accuracy of iRNA-ac4C on a balanced independent dataset was 0.798 with an ROC value of 0.880. Lai et al. [41] introduced a new tool, LSA-ac4C, for the accurate prediction of ac4C modification sites in human mRNA. In this technique, they used a hybrid neural network that combines an LSTM double-layer neural network with self-attention, with 10-fold CV, and achieved an accuracy of 0.820. The performance on a balanced independent dataset was 0.827 with an ROC value of 0.895. Jia et al. [42] developed a new model, EMDL-ac4C. In this technique, they used one-hot encoding and then inserted it into an ensemble model with a two-branch residual connection dense Net and attention to predict ac4C modification sites in human mRNA. In their methodology, they utilized an unbalanced dataset with 10-fold CV and a balanced dataset with 5-fold CV to train their models. Finally, they got the best model on the balanced dataset with an accuracy of 0.892. The performance of EMDL-ac4C on the balanced independent dataset was 0.808 with an ROC value of 0.879, and the ROC value on the unbalanced independent dataset was 0.901.
In 2024, Li et al. [43] introduced a MetaAc4C model in which they utilized pre-trained BERT, and the model was based on a BLSTM network. They achieved a 0.937 accuracy score with an ROC value of 0.986. The performance on unbalanced and balanced independent datasets was 0.828 and 0.817, respectively. Li et al. [44] also introduced a new method based on generative adversarial networks and transfer learning to recognize ac4C modification sites in human mRNA, named GANSamples-ac4C. They achieved an accuracy of 0.834 on the training dataset with 5-fold CV. Pham et al. [45] constructed the ac4C-AFL method, which is based on adaptive feature representation learning. In their method, they utilized sixteen types of feature encoding schemes CKSNAP, ENAC, position specific of two nucleotides, PseEIIP, Z-curve, k-mer, reverse RCKmer, di-nucleotide physicochemical properties type 1-2, NCP, binary features, multivariate mutual information and accumulated nucleotide frequency, W2V, sequence2vector, DNABERT and combination of skip di-nucleotide composition and local position-specific di-nucleotide frequency and then inputted these optimized features into 11 different types of ML and DL classifiers, namely, SVM, RF, ERT, ANN, LR, GBT, XGBT, light GBT, cat boost, AB, and CNN classifiers and generated 176 baseline models. Finally, they selected the best baseline models using a 2-step feature selection technique, whose predicting scores were integrated and trained with SVM to develop the final model by utilizing 10-fold CV to classify ac4C from non-ac4C in human mRNA, and achieved an accuracy score of 0.833. The accuracy score of ac4C-AFL on the balanced independent dataset was 0.823 with an ROC value of 0.895. Yi et al. [46] developed a new method, named STM-ac4C, which is a hybrid model based on selective kernel convolutions, a temporal convolutional network with multi-head self-attention to predict ac4C modification sites in mRNA. They attained an accuracy of 0.825 on the training dataset by utilizing 10-fold CV and an accuracy of 0.847 on the balanced independent dataset. Yuan et al. [47] developed a new method, DPNN-ac4C, based on a dual path neural network with self-attention. Their method integrates a convolutional neural network, embedding modules, a BGRU with self-attention to extract local and global features from mRNA sequences, and achieved an accuracy of 0.838. The prediction efficiency of DPNN-ac4C on the balanced independent dataset was 0.827 with an ROC value of 0.910. Liu et al. [48] developed a transformer-based model called TransAC4C to predict ac4C modification sites in mRNA. Their model was divided into four parts such as a transformer layer, one BLSTM layer, four 1D convolutional layers, with three fully connected layers. Transformer and BLSTM layers with 10-fold CV were used to enable the model to learn the contextual information, convolutional layers were used to enable the model to extract important features from the input, and fully connected layers were used to connect the input features into the output. They trained two models based on nucleotide length (415 nt and 21 nt) and achieved an accuracy score of 0.785 and 0.923, respectively. The prediction efficiencies of the above two models with nucleotide length (415 and 21) on the balanced independent dataset were 0.779 and 0.838. Jia et al. [49] constructed a tool called Voting-ac4C to predict ac4C modification sites in mRNA. In their method, they utilized RNAErnie, a transformer-based pre-trained model with six traditional feature encoding schemes such as one-hot, ENAC, ND, C2 encoding, KSNPF, and physicochemical properties, and then inserted the hybrid of these features into a DNN for dimensionality reduction. Subsequently, these dimensional reduction features were fed into a voting ensemble model constructed using CatBoost, XGBoost, and MLP classifiers by utilizing 10-fold CV. Finally, they attained an accuracy score of 0.867. The accuracy score on the balanced independent dataset was 0.831 with an ROC value of 0.887. He et al. [50] developed a DL method called NBCR-ac4C based on pretrained models to predict ac4C modification sites in human mRNA. They utilized DNABERT2 and nucleotide transformer to construct contextual embedding of nucleotide sequences and utilized CNN and ResNet18 by using 10-fold CV to further extract the shallow and deep knowledge from the contextual embedding. Subsequently, they got the accuracy score of 0.837 on the training dataset and an accuracy score of 0.835 on the balanced independent dataset with an ROC value of 0.895.
In 2025, Lu et al. [51] introduced an ERNIE-ac4C method to predict ac4C modification sites. In their technique, they used the pretrained model ERNIE-RNA to extract the attention map features and sequence features from nucleotide sequences. They input these fused features into a 2d CNN to predict ac4C modification sites in the human mRNA. The performance of ERNIE-ac4C on the balanced independent dataset was 0.902 with an ROC value of 0.961. Yao et al. [52] proposed a DL-based method called Caps-ac4C to predict ac4C modification sites in human mRNA. In their method, they utilized CGR encoding to transform the nucleotide sequences into visual representations and then employed a capsule network design to extract the local and global features from these visual representations of nucleotide sequences. The performance accuracies of Caps-ac4C on balanced and unbalanced independent datasets were 0.954 and 0.908, with ROC values of 0.996 and 0.962, respectively. The performance comparison of published results on balanced and unbalanced independent datasets is also shown in Table 4. Performance of the available tools on balanced and unbalanced independent data is shown in Fig. 6.
Table 4.
Comparison of the published results for ac4C prediction on independent datasets.
| Method | Dataset | Evaluation | ACC | SN | SP | MCC | ROC | PRC |
|---|---|---|---|---|---|---|---|---|
| PACES [33] | Un-balanced | Independent set | – | – | – | – | 0.874 | 0.485 |
| XG-ac4C [34] | Un-balanced | Independent set | – | – | – | – | 0.889 | 0.581 |
| DeepAc4C [35] | Un-balanced | Independent set | 0.791 | 0.828 | 0.755 | 0.585 | 0.864 | – |
| CNNLSTMac4CPred [36] | Un-balanced | Independent set | 0.872 | 0.629 | 0.899 | 0.435 | 0.882 | – |
| DLC-ac4C [37] | Balanced | Independent set | 0.829 | 0.862 | 0.797 | 0.660 | 0.903 | – |
| Stacking-ac4C [38] | Balanced | Independent set | 0.808 | 0.808 | 0.808 | 0.615 | 0.883 | – |
| TransC-ac4C [39] | Balanced | Independent set | 0.806 | 0.809 | 0.804 | 0.614 | 0.869 | – |
| iRNA-ac4C [40] | Balanced | Independent set | 0.798 | 0.767 | 0.829 | 0.597 | 0.880 | – |
| LSA-ac4C [41] | Balanced | Independent set | 0.827 | 0.871 | 0.782 | 0.656 | 0.895 | – |
| EMDL-ac4C [42] | Balanced | Independent set | 0.808 | 0.810 | 0.817 | 0.616 | 0.879 | 0.864 |
| EMDL-ac4C [42] | Un-balanced | Independent set | – | – | – | – | 0.901 | 0.594 |
| MetaAc4C [43] | Un-balanced | Independent set | 0.828 | 0.809 | 0.848 | 0.657 | 0.895 | – |
| MetaAc4C [43] | Balanced | Independent set | 0.817 | 0.792 | 0.843 | 0.636 | 0.874 | – |
| GANSamples-ac4C [44] | Balanced | Independent set | – | – | – | – | – | – |
| ac4C-AFL [45] | Balanced | Independent set | 0.823 | 0.844 | 0.803 | 0.647 | 0.895 | – |
| STM-ac4C [46] | Balanced | Independent set | 0.847 | 0.858 | 0.837 | 0.695 | 0.907 | – |
| DPNN-ac4C [47] | Balanced | Independent set | 0.827 | 0.817 | 0.847 | 0.657 | 0.910 | – |
| TransAC4C415 nt [48] | Balanced | Independent set | 0.779 | 0.794 | 0.765 | 0.559 | 0.853 | – |
| TransAC4C21 nt [48] | Balanced | Independent set | 0.838 | 0.838 | 0.838 | 0.677 | 0.903 | – |
| Voting-ac4C [49] | Balanced | Independent set | 0.831 | 0.851 | 0.811 | 0.663 | 0.887 | – |
| NBCR-ac4C [50] | Balanced | Independent set | 0.835 | 0.849 | 0.820 | 0.670 | 0.895 | – |
| ERNIE-ac4C [51] | Balanced | Independent set | 0.902 | 0.897 | 0.908 | 0.806 | 0.961 | 0.965 |
| Caps-ac4C [52] | Un-balanced | Independent set | 0.908 | 0.889 | 0.912 | 0.721 | 0.962 | – |
| Caps-ac4C [52] | Balanced | Independent set | 0.954 | 0.920 | 0.989 | 0.912 | 0.996 | 0.995 |
Figure 6.
Performance of the available ac4C modification sites prediction tools on balanced and unbalanced independent datasets.
Independent evaluation
For the sake of fair and authentic evaluation, it is very important to check the published tool’s performance on random datasets. For this purpose, we selected those web-servers that were running online and found three web-servers, namely, ac4C-AFL [45], Caps-ac4C [52], and PACES [33]. We have tested these three state-of-the-art tools on three different datasets (Su et al. [40], Wang et al. [35], and Arango et al. [14]) to check the performance and found that ac4C-AFL [45] performed well on these datasets as compared to the Caps-ac4C [52] and PACES [33]. The accuracy score of ac4C-AFL [45] on the Su et al. [40] dataset was 0.886 and 0.848 on the Wang et al. dataset and 0.753 on the Arango et al. dataset. The accuracy score of Caps-ac4C on the Su et al. dataset was 0.845, 0.825 on the Wang et al. dataset, and 0.662 on Arango et al [14]. The accuracy score of PACES [33] on the Su et al. [40] dataset was 0.776, 0.728 on the Wang et al. dataset, and 0.611 on the Arango et al. [14] dataset. Finally, ac4C-AFL [45] performed well across independent datasets and outperformed the other web-based tools by 4.1%–11% on the Su et al. dataset, 2.3%–12% on the Wang et al. [35] dataset, and 9.1%–14.2% on the Arango et al. [14] dataset. The model’s performance on the Arango et al. dataset was not up to the mark due to data leakage and redundant sequences, compared with the Wang et al. and Su et al. datasets. The Arango et al. dataset needs to be polished by data-cleaning tools for removing redundant sequences. The performance results of available web-based tools on three different datasets are also shown in Table 5 and Fig. 7.
Table 5.
Performance of available active web servers on three different independent datasets.
| Method | Independent dataset | ACC | SN | SP | MCC |
|---|---|---|---|---|---|
| PACES [33] | Su et al. | 0.776 | 0.763 | 0.781 | 0.623 |
| Arango et al. | 0.611 | 0.608 | 0.617 | 0.475 | |
| Wang et al. | 0.728 | 0.714 | 0.733 | 0.588 | |
| ac4C-AFL [45] | Su et al. | 0.886 | 0.876 | 0.881 | 0.735 |
| Arango et al. | 0.753 | 0.743 | 0.748 | 0.587 | |
| Wang et al. | 0.848 | 0.851 | 0.839 | 0.675 | |
| Caps-ac4C [52] | Su et al. | 0.845 | 0.835 | 0.837 | 0.697 |
| Arango et al. | 0.662 | 0.655 | 0.674 | 0.516 | |
| Wang et al. | 0.825 | 0.831 | 0.827 | 0.662 |
Figure 7.
Performance of the available web-based tools on the independent datasets. Performance on Su et al. data (A). Performance on Wang et al. data (B). Performance on Arango et al. data (C).
Conclusion
NAT10 catalyzes ac4C, which is among the most significant posttranscriptional modifications in RNA. It adds an acetyl group to the nitrogen at the fourth position of the cytidine base and plays an essential role in the stability of mRNA, posttranscriptional regulation, translational efficiency, and human immune function regulation [133]. Thus, it is crucial to identify ac4C computationally. It can quickly and thoroughly identify additional details and undiscovered functions of ac4C when compared to the costly experimental methods. Many classifiers based on DL and ML are currently being developed. Numerous feature extraction techniques, such as KSNPF, CKSNAP, PseEIIP, PseKNC, DNABERT2 embedding, ANF, SCPseDNC, SCPseTNC, PseKNC, physiochemical properties, and some feature selection techniques, such as BD, PCA, and ANOVA with IFS, were employed in all of these predictors to eliminate redundant features and obtain optimal features. Artificial intelligence-based classification is important for the prediction of ac4C in human RNA. Therefore, we compared and assessed the available machine and deep learning-based classifiers and found that the MetaAc4C [43] method has an outstanding performance and generalization ability on training datasets, and the Caps-ac4C [52] method showed brilliant performance on independent datasets. We have also assessed the performance of the available web-based tools on three different datasets and found that ac4C-AFL [45] outpaced the other tools in terms of accuracy score. Although the ac4C classifiers’ results are reasonable, there is still room for improvement. Clear features that can reflect the fundamental characteristics of ac4C against non-ac4C are required. Several challenges limited the research to some extent; these limitations can be characterized into two groups: general limitations and specific limitations.
The specific limitations: first, the sample size of the available data for predicting ac4C modification sites is relatively small. This drawback prevented the application of certain methods. Therefore, we suggest a reasonably larger dataset in future studies. Machine and deep learning-based studies are used in this review, which also comes with various limitations, such as dataset size, overfitting, underfitting, and dealing with dataset outliers. Due to small datasets, the performance of traditional machine learning models, such as SVM, RF, AB, and NB, was reasonable, but DL models faced issues regarding overfitting and underfitting. As a result, in the future, the next study will include more data and other clustering methods proven to be much more efficient.
The general limitations: there are also some limitations to the feature extraction method. The KSNPF, CKSNAP, PseEIIP, and PseKNC feature extraction techniques have problems with computational power due to their size and computational cost. There is still a need to do more by exploiting advanced feature descriptors such as evolutionary scale modeling ESM1-3. We anticipate that in the near future, more amazing results will be obtained with the support of LLMs as bioinformatics and computational biology improve.
Key Points
ac4C modification sites are very crucial and directly related to mRNA stability, transcription regulation, and regulation of the immune functions.
The deep learning models with meaningful features are the forward step towards accurate N4 acetylation modification sites prediction in human mRNA.
The development of the particular ac4C modification sites large language models is highly anticipated in the field.
Data-driven discovery of ac4C modification sites is emerging as the frontier of computational biology.
Contributor Information
Hasan Zulfiqar, Center for AI and Computational Biology, Institute of System Medicine, Peking Union Medical College, Chinese Academy of Medical Sciences, Suzhou 215123, China.
Ramala Masood Ahmad, Center for Informational Biology, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China; Department of Plant Breeding and Genetics, University of Agriculture Faisalabad, Faisalabad 38000, Pakistan.
Hao Lin, Center for Informational Biology, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China.
Xiao-Long Yu, School of Materials Science and Engineering, Hainan University, Haikou 570228, China.
Author’s contributions
Hasan Zulfiqar (Conceptualization, Writing—original draft, Resources, Investigation, Funding acquisition), Ramala Masood Ahmad (Formal analysis, Investigation, Visualization), Hao Lin (Supervision, Writing—review & editing, Funding acquisition), and Xiao-Long Yu (Supervision, Writing—review & editing, Funding acquisition)
Funding
This work has been supported by the grant from the National Natural Science Foundation of China (62302079, 62261017).
Data availability
There is no data.
References
- 1. Wada T, Kobori A, Kawahara SI et al. Synthesis and properties of oligodeoxyribonucleotides containing 4-N-acetylcytosine bases. Tetrahedron Lett 1998;39:6907–10. 10.1016/S0040-4039(98)01449-X [DOI] [Google Scholar]
- 2. Gu Z, Zou L, Pan X et al. The role and mechanism of NAT10-mediated ac4C modification in tumor development and progression. MedComm 2024;5:e70026. 10.1002/mco2.70026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Wang Q, Yuan Y, Zhou Q et al. RNA N4-acetylcytidine modification and its role in health and diseases. MedComm 2025;6:e70015. 10.1002/mco2.70015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Jiao L, Si Y, Yuan Y et al. Emerging role of N-acetyltransferase 10 in diseases: RNA ac4C modification and beyond. Mol Biomed 2025;6:46. 10.1186/s43556-025-00286-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Tao L, Lu Y, Chen Z et al. RNA ac4C modification in cancer biology: from regulatory mechanisms to clinical applications. Sci China Life Sci 2024;67:832–5. 10.1007/s11427-023-2496-5 [DOI] [PubMed] [Google Scholar]
- 6. Xu N, Zhuo J, Chen Y et al. Downregulation of N4-acetylcytidine modification in myeloid cells attenuates immunotherapy and exacerbates hepatocellular carcinoma progression. Br J Cancer 2024;130:201–12. 10.1038/s41416-023-02510-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Liu H, Xu L, Yue S et al. Targeting N4-acetylcytidine suppresses hepatocellular carcinoma progression by repressing eEF2-mediated HMGB2 mRNA translation. Cancer Commun 2024;44:1018–41. 10.1002/cac2.12595 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Wang Y, Wang S, Liu M et al. The function of NAT10-driven N4-acetylcytidine modification in cancer: novel insights and potential therapeutic targets. Cell Biosci 2025;15:165. 10.1186/s13578-025-01504-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Chen Z, Zhao M, You L et al. Developing an artificial intelligence method for screening hepatotoxic compounds in traditional Chinese medicine and western medicine combination. Chin Med 2022;17:58. 10.1186/s13020-022-00617-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Ji H-N, Zhou HQ, Qie JB et al. Dysregulated ac4C modification of mRNA in a mouse model of early-stage Alzheimer’s disease. Cell Biosci 2025;15:45. 10.1186/s13578-025-01389-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Ma Y, Li W, Fan C et al. Comprehensive analysis of long non-coding RNAs N4-acetylcytidine in Alzheimer’s disease mice model using high-throughput sequencing. J Alzheimer’s Dis 2022;90:1659–75. 10.3233/JAD-220564 [DOI] [PubMed] [Google Scholar]
- 12. Chen Z, Jiang Y, Zhang X et al. ResNet18DNN: prediction approach of drug-induced liver injury by deep neural network with ResNet18. Brief Bioinform 2022;23:bbab503. 10.1093/bib/bbab503 [DOI] [PubMed] [Google Scholar]
- 13. Chen Z, Jiang Y, Zhang X et al. The prediction approach of drug-induced liver injury: response to the issues of reproducible science of artificial intelligence in real-world applications. Brief Bioinform 2022;23:bbac196. 10.1093/bib/bbac196 [DOI] [PubMed] [Google Scholar]
- 14. Arango D, Sturgill D, Alhusaini N et al. Acetylation of cytidine in mRNA promotes translation efficiency. Cell 2018;175:1872–1886.e24. 10.1016/j.cell.2018.10.030 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Schiffers S, Oberdoerffer S. ac4C: a fragile modification with stabilizing functions in RNA metabolism. RNA 2024;30:583–94. 10.1261/rna.079948.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Zhang X, Zheng Y, Yang J et al. Abnormal ac4C modification in metabolic dysfunction associated steatotic liver cells. Sci Rep 2025;15:1013. 10.1038/s41598-024-84564-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Luo Y, Shi L, Li Y et al. From intention to implementation: automating biomedical research via LLMs. Sci China Inf Sci 2025;68:170105. 10.1007/s11432-024-4485-0 [DOI] [Google Scholar]
- 18. Qiao J, Jin J, Yu H et al. Towards retraining-free RNA modification prediction with incremental learning. Inf Sci 2024;660:120105. [Google Scholar]
- 19. Meissner A, Gnirke A, Bell GW et al. Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis. Nucleic Acids Res 2005;33:5868–77. 10.1093/nar/gki901 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Nakabayashi K, Yamamura M, Haseagawa K et al. Reduced representation bisulfite sequencing (RRBS). In: Hatada I, Horii T (eds.), Methods Mole. Biol, 2577, pp. 39–51. Humana, New York, NY: Epigenomics, 2022. 10.1007/978-1-0716-2724-2_3 [DOI] [PubMed] [Google Scholar]
- 21. Farrell C, Thompson M, Tosevska A et al. BiSulfite bolt: a bisulfite sequencing analysis platform. Gigascience 2021;10:giab033. 10.1093/gigascience/giab033 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Tse OO, Jiang P, Cheng SH et al. Genome-wide detection of cytosine methylation by single molecule real-time sequencing. Proc Natl Acad Sci 2021;118:e2019768118. 10.1073/pnas.2019768118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Cascino P, Nevone A, Piscitelli M et al. Single-molecule real-time sequencing of the M protein: toward personalized medicine in monoclonal gammopathies. Am J Hematol 2022;97:E389–92. 10.1002/ajh.26684 [DOI] [PubMed] [Google Scholar]
- 24. Li X, Wang X, Ma Q et al. Integrated single-molecule real-time sequencing and RNA sequencing reveal the molecular mechanisms of salt tolerance in a novel synthesized polyploid genetic bridge between maize and its wild relatives. BMC Genomics 2023;24:55. 10.1186/s12864-023-09148-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Deslattes Mays A, Schmidt M, Graham G et al. Single-molecule real-time (SMRT) full-length RNA-sequencing reveals novel and distinct mRNA isoforms in human bone marrow cell subpopulations. Genes 2019;10:253. 10.3390/genes10040253 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Covey TR, Lee ED, Bruins AP et al. Liquid chromatography/mass spectrometry. Anal Chem 1986;58:1451A–61. 10.1021/ac00127a001 [DOI] [Google Scholar]
- 27. Beiki H, Sturgill D, Arango D et al. Detection of ac4C in human mRNA is preserved upon data reassessment. Mol Cell 2024;84:1611–1625.e3. 10.1016/j.molcel.2024.03.018 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Wang S, Xie H, Mao F et al. N 4-acetyldeoxycytosine DNA modification marks euchromatin regions in Arabidopsis thaliana. Genome Biol 2022;23:5. 10.1186/s13059-021-02578-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Zulfiqar H, Huang QL, Lv H et al. Deep-4mCGP: a deep learning approach to predict 4mC sites in Geobacter pickeringii by using correlation-based feature selection technique. Int J Mol Sci 2022;23:1251. 10.3390/ijms23031251 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Zulfiqar H, Sun ZJ, Huang QL et al. Deep-4mCW2V: a sequence-based predictor to identify N4-methylcytosine sites in Escherichia coli. Methods 2022;203:558–63. 10.1016/j.ymeth.2021.07.011 [DOI] [PubMed] [Google Scholar]
- 31. Dao F-Y, Lv H, Yang YH et al. Computational identification of N6-methyladenosine sites in multiple tissues of mammals. Comput Struct Biotechnol J 2020;18:1084–91. 10.1016/j.csbj.2020.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Dai C, Jiang Y, Yin C et al. scIMC: a platform for benchmarking comparison and visualization analysis of scRNA-seq data imputation methods. Nucleic Acids Res 2022;50:4877–99. 10.1093/nar/gkac317 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Zhao W, Zhou Y, Cui Q et al. PACES: prediction of N4-acetylcytidine (ac4C) modification sites in mRNA. Sci Rep 2019;9:11112. 10.1038/s41598-019-47594-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Alam W, Tayara H, Chong KT. XG-ac4C: identification of N4-acetylcytidine (ac4C) in mRNA using eXtreme gradient boosting with electron-ion interaction pseudopotentials. Sci Rep 2020;10:20942. 10.1038/s41598-020-77824-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Wang C, Ju Y, Zou Q et al. DeepAc4C: a convolutional neural network model with hybrid features composed of physicochemical patterns and distributed representation information for identification of N4-acetylcytidine in mRNA. Bioinformatics 2022;38:52–7. 10.1093/bioinformatics/btab611 [DOI] [PubMed] [Google Scholar]
- 36. Zhang G, Luo W, Lyu J et al. CNNLSTMac4CPred: a hybrid model for N 4-Acetylcytidine prediction. Interdiscip Sci: Comput Life Sci 2022;14:439–51. 10.1007/s12539-021-00500-0 [DOI] [PubMed] [Google Scholar]
- 37. Jia J, Cao X, Wei Z. DLC-ac4C: a prediction model for N4-acetylcytidine sites in human mRNA based on DenseNet and bidirectional LSTM methods. Curr Genomics 2023;24:171–86. 10.2174/0113892029270191231013111911 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Lou L-L, Qiu WR, Liu Z et al. Stacking-ac4C: an ensemble model using mixed features for identifying n4-acetylcytidine in mRNA. Front Immunol 2023;14:1267755. 10.3389/fimmu.2023.1267755 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Liu D, Liu Z, Xia Y et al. TransC-ac4C: identification of N4-acetylcytidine (ac4C) sites in mRNA using deep learning. IEEE/ACM Trans Comput Biol Bioinform 2024;21:1403–12. 10.1109/TCBB.2024.3386972 [DOI] [PubMed] [Google Scholar]
- 40. Su W, Xie XQ, Liu XW et al. iRNA-ac4C: a novel computational method for effectively detecting N4-acetylcytidine sites in human mRNA. Int J Biol Macromol 2023;227:1174–81. 10.1016/j.ijbiomac.2022.11.299 [DOI] [PubMed] [Google Scholar]
- 41. Lai F-L, Gao F. LSA-ac4C: a hybrid neural network incorporating double-layer LSTM and self-attention mechanism for the prediction of N4-acetylcytidine sites in human mRNA. Int J Biol Macromol 2023;253:126837. 10.1016/j.ijbiomac.2023.126837 [DOI] [PubMed] [Google Scholar]
- 42. Jia J, Wei Z, Cao X. EMDL-ac4C: identifying N4-acetylcytidine based on ensemble two-branch residual connection DenseNet and attention. Front Genet 2023;14:1232038. 10.3389/fgene.2023.1232038 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Li Z, Jin B, Fang J. MetaAc4C: a multi-module deep learning framework for accurate prediction of N4-acetylcytidine sites based on pre-trained bidirectional encoder representation and generative adversarial networks. Genomics 2024;116:110749. 10.1016/j.ygeno.2023.110749 [DOI] [PubMed] [Google Scholar]
- 44. Li F, Zhang J, Li K et al. GANSamples-ac4C: enhancing ac4C site prediction via generative adversarial networks and transfer learning. Anal Biochem 2024;689:115495. 10.1016/j.ab.2024.115495 [DOI] [PubMed] [Google Scholar]
- 45. Pham NT, Terrance AT, Jeon YJ et al. ac4C-AFL: a high-precision identification of human mRNA N4-acetylcytidine sites based on adaptive feature representation learning. Mol Ther Nucleic Acids 2024;35:102192. 10.1016/j.omtn.2024.102192 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Yi M, Zhou F, Deng Y. STM-ac4C: a hybrid model for identification of N4-acetylcytidine (ac4C) in human mRNA based on selective kernel convolution, temporal convolutional network, and multi-head self-attention. Front Genet 2024;15:1408688. 10.3389/fgene.2024.1408688 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Yuan J, Wang Z, Pan Z et al. DPNN-ac4C: a dual-path neural network with self-attention mechanism for identification of N4-acetylcytidine (ac4C) in mRNA. Bioinformatics 2024;40:btae625. 10.1093/bioinformatics/btae625 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Liu R, Zhang Y, Wang Q et al. TransAC4C—a novel interpretable architecture for multi-species identification of N4-acetylcytidine sites in RNA with single-base resolution. Brief Bioinform 2024;25:bbae200. 10.1093/bib/bbae200 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Jia Y, Zhang Z, Yan S et al. Voting-ac4C: pre-trained large RNA language model enhances RNA N4-acetylcytidine site prediction. Int J Biol Macromol 2024;282:136940. 10.1016/j.ijbiomac.2024.136940 [DOI] [PubMed] [Google Scholar]
- 50. He W, Han Y, Zuo Y et al. NBCR-ac4C: a deep learning framework based on multivariate BERT for human mRNA N4-Acetylcytidine sites prediction. J Chem Inf Model 2024;64:8074–81. 10.1021/acs.jcim.4c01415 [DOI] [PubMed] [Google Scholar]
- 51. Lu R, Qiao J, Li K et al. ERNIE-ac4C: a novel deep learning model for effectively predicting N4-acetylcytidine sites. J Mol Biol 2025;437:168978. 10.1016/j.jmb.2025.168978 [DOI] [PubMed] [Google Scholar]
- 52. Yao L, Xie P, Dong D et al. Caps-ac4C: an effective computational framework for identifying N4-acetylcytidine sites in human mRNA based on deep learning. J Mol Biol 2025;437:168961. 10.1016/j.jmb.2025.168961 [DOI] [PubMed] [Google Scholar]
- 53. Griffiths-Jones S, Bateman A, Marshall M et al. Rfam: an RNA family database. Nucleic Acids Res 2003;31:439–41. 10.1093/nar/gkg006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Nawrocki EP, Burge SW, Bateman A et al. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res 2015;43:D130–7. 10.1093/nar/gku1063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Burge SW, Daub J, Eberhardt R et al. Rfam 11.0: 10 years of RNA families. Nucleic Acids Res 2013;41:D226–32. 10.1093/nar/gks1005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Xuan J, Chen L, Chen Z et al. RMBase v3. 0: decode the landscape, mechanisms and functions of RNA modifications. Nucleic Acids Res 2024;52:D273–84. 10.1093/nar/gkad1070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Zulfiqar H, Guo Z, Ahmad RM et al. Deep-STP: a deep learning-based approach to predict snake toxin proteins by using word embeddings. Front Med 2024;10:1291352. 10.3389/fmed.2023.1291352 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Zulfiqar H, Khan RS, Hassan F et al. Computational identification of N4-methylcytosine sites in the mouse genome with machine-learning method. Math Biosci Eng 2021;18:3348–63. 10.3934/mbe.2021167 [DOI] [PubMed] [Google Scholar]
- 59. Lv H, Dao FY, Zulfiqar H et al. A sequence-based deep learning approach to predict CTCF-mediated chromatin loop. Brief Bioinform 2021;22:bbab031. 10.1093/bib/bbab031 [DOI] [PubMed] [Google Scholar]
- 60. Zulfiqar H, Ahmad RM, Raza A et al. Promoter prediction in agrobacterium tumefaciens strain C58 by using artificial intelligence strategies. In: Marchisio MA (ed.), Methods Mole. Biol, 2844, pp. 33–44. Humana, New York, NY: Synthetic Promoters, 2024. 10.1007/978-1-0716-4063-0_2 [DOI] [PubMed] [Google Scholar]
- 61. Zulfiqar H, Ahmed Z, Kissanga Grace-Mercure B et al. Computational prediction of promotors in agrobacterium tumefaciens strain C58 by using the machine learning technique. Front Microbiol 2023;14:1170785. 10.3389/fmicb.2023.1170785 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Chen Y, Wang Z, Wang J et al. Self-supervised learning in drug discovery. Sci China Inf Sci 2025;68:170103. 10.1007/s11432-024-4453-4 [DOI] [Google Scholar]
- 63. Jenike KM, Campos-Domínguez L, Boddé M et al. k-mer approaches for biodiversity genomics. Genome Res 2025;35:219–30. 10.1101/gr.279452.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Roberts MD, Davis O, Josephs EB et al. k-mer-based approaches to bridging pangenomics and population genetics. Mol Biol Evol 2025;42:msaf047. 10.1093/molbev/msaf047 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Katayama Y, Kobayashi TJ. Comparative study of repertoire classification methods reveals data efficiency of k-mer feature extraction. Front Immunol 2022;13:797640. 10.3389/fimmu.2022.797640 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Li M, Fan Y, Zhang Y et al. Using sequence similarity based on CKSNP features and a graph neural network model to identify miRNA–disease associations. Genes 2022;13:1759. 10.3390/genes13101759 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Basith S, Hasan MM, Lee G et al. Integrative machine learning framework for the identification of cell-specific enhancers from the human genome. Brief Bioinform 2021;22:bbab252. 10.1093/bib/bbab252 [DOI] [PubMed] [Google Scholar]
- 68. Zhang ZY, Fan YE, Huang CB et al. Human essential gene identification based on feature fusion and feature screening. IET Syst Biol 2024;18:227–37. 10.1049/syb2.12105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Chen W, Yang H, Feng P et al. iDNA4mC: identifying DNA N4-methylcytosine sites based on nucleotide chemical properties. Bioinformatics 2017;33:3518–23. 10.1093/bioinformatics/btx479 [DOI] [PubMed] [Google Scholar]
- 70. Chen Z, Zhao P, Li F et al. iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data. Brief Bioinform 2020;21:1047–57. 10.1093/bib/bbz041 [DOI] [PubMed] [Google Scholar]
- 71. Lv Z, Ding H, Wang L et al. A convolutional neural network using dinucleotide one-hot encoder for identifying DNA N6-methyladenine sites in the rice genome. Neurocomputing 2021;422:214–21. 10.1016/j.neucom.2020.09.056 [DOI] [Google Scholar]
- 72. Zheng K, You ZH, Li JQ et al. iCDA-CGR: identification of circRNA-disease associations based on chaos game representation. PLoS Comput Biol 2020;16:e1007872. 10.1371/journal.pcbi.1007872 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Gupta YM, Kirana SN, Homchan S. Representing DNA for machine learning algorithms: a primer on one-hot, binary, and integer encodings. Biochem Mol Biol Educ 2025;53:142–6. 10.1002/bmb.21870 [DOI] [PubMed] [Google Scholar]
- 74. Zhang P, Zhang H, Wu H. iPro-WAEL: a comprehensive and robust framework for identifying promoters in multiple species. Nucleic Acids Res 2022;50:10278–89. 10.1093/nar/gkac824 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75. Dou L, Li X, Ding H et al. Prediction of m5C modifications in RNA sequences by combining multiple sequence features. Mol Ther Nucleic Acids 2020;21:332–42. 10.1016/j.omtn.2020.06.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Ji Y, Zhou Z, Liu H et al. DNABERT: pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. Bioinformatics 2021;37:2112–20. 10.1093/bioinformatics/btab083 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77. Makhdoomi F, Ghosh N. TinyDNABERT: an efficient and optimized BERT model for DNA-language in genome. In: 21st IEEE India Council International Conference (INDICON), 2024, pp. 1–6. Kharagpur, India, 2024. 10.1109/INDICON63790.2024.10958313 [DOI] [Google Scholar]
- 78. Sanabria M, Hirsch J, Poetsch AR. Distinguishing word identity and sequence context in DNA language models. BMC Bioinformatics 2024;25:301. 10.1186/s12859-024-05869-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Kurata H, Tsukiyama S, Manavalan B. iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model. Brief Bioinform 2022;23:bbac265. [DOI] [PubMed] [Google Scholar]
- 80. Venkatesh B, Anuradha J. A review of feature selection and its methods. Cybern Inf Technol 2019;19:3–26. [Google Scholar]
- 81. Liu H, Setiono R. Feature selection and classification–a probabilistic wrapper approach. In: Tanaka T, Ohsuga S, Ali M (eds.), Proceedings of the 9th International Conference on Industrial and Engineering Applications of Artificial Intelligence and Expert Systems, 1996, pp. 419–24. Fukuoka, Japan, 1996. https://dl.acm.org/doi/proceedings/10.5555/3104635 [Google Scholar]
- 82. Chen Z, Pang M, Zhao Z et al. Feature selection may improve deep neural networks for the bioinformatics problems. Bioinformatics 2020;36:1542–52. 10.1093/bioinformatics/btz763 [DOI] [PubMed] [Google Scholar]
- 83. Wang Y, Gao X, Ru X et al. A hybrid feature selection algorithm and its application in bioinformatics. PeerJ Comput Sci 2022;8:e933. 10.7717/peerj-cs.933 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84. Dao F-Y, Lv H, Zhang ZY et al. BDselect: a package for k-mer selection based on the binomial distribution. Curr Bioinform 2022;17:238–44. 10.2174/1574893616666211007102747 [DOI] [Google Scholar]
- 85. Sparta B, Hamilton T, Natesan G et al. Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing. PLoS Comput Biol 2024;20:e1012386. 10.1371/journal.pcbi.1012386 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86. Greenacre M, Groenen PJF, Hastie T et al. Principal component analysis. Nat Rev Methods Primers 2022;2:100. 10.1038/s43586-022-00184-w [DOI] [Google Scholar]
- 87. Bielińska-Wąż D, Wąż P, Błaczkowska A et al. Mathematical modeling in bioinformatics: application of an alignment-free method combined with principal component analysis. Symmetry 2024;16:967. 10.3390/sym16080967 [DOI] [Google Scholar]
- 88. Privé F, Luu K, Blum MGB et al. Efficient toolkit implementing best practices for principal component analysis of population genetic data. Bioinformatics 2020;36:4449–57. 10.1093/bioinformatics/btaa520 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89. Hassall KL, Mead A. Beyond the one-way ANOVA for ’omics data. BMC Bioinformatics 2018;19:199. 10.1186/s12859-018-2173-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90. Ye Y, Zhang R, Zheng W et al. RIFS: a randomly restarted incremental feature selection algorithm. Sci Rep 2017;7:13013. 10.1038/s41598-017-13259-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91. Pisner DA, Schnyer DM. Support vector machine. In: Mechelli A, Vieira S (eds.), Elsevier, 2020, pp. 101–21. Amsterdam, Netherlands: Machine Learning, 2020. 10.1016/B978-0-12-815739-8.00006-7 [DOI] [Google Scholar]
- 92. Cervantes J, Garcia-Lamont F, Rodríguez-Mazahua L et al. A comprehensive survey on support vector machine classification: applications, challenges and trends. Neurocomputing 2020;408:189–215. 10.1016/j.neucom.2019.10.118 [DOI] [Google Scholar]
- 93. Yan H, Long Y, Lv C et al. Research on bioinformatics data classification method based on support vector machine. Int J Data Min Bioinf 2025;29:21–35. 10.1504/IJDMB.2025.142975 [DOI] [Google Scholar]
- 94. Li H, Pang Y, Liu B. BioSeq-BLM: a platform for analyzing DNA, RNA, and protein sequences based on biological language models. Nucleic Acids Res 2021;49:e129. 10.1093/nar/gkab829 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95. Lu H, Gao H, Ye M et al. A hybrid ensemble algorithm combining AdaBoost and genetic algorithm for cancer classification with gene expression data. IEEE/ACM Trans Comput Biol Bioinform 2019;18:863–70. 10.1109/TCBB.2019.2952102 [DOI] [PubMed] [Google Scholar]
- 96. Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn 2006;63:3–42. 10.1007/s10994-006-6226-1 [DOI] [Google Scholar]
- 97. Kocev D, Ceci M, Stepišnik T. Ensembles of extremely randomized predictive clustering trees for predicting structured outputs. Mach Learn 2020;109:2213–41. 10.1007/s10994-020-05894-4 [DOI] [Google Scholar]
- 98. Salman HA, Kalakech A, Steiti A. Random forest algorithm overview. Babylon J Mach Learn 2024;2024:69–79. 10.58496/BJML/2024/007 [DOI] [Google Scholar]
- 99. Ghosh D, Cabrera J. Enriched random forest for high dimensional genomic data. IEEE/ACM Trans Comput Biol Bioinform 2021;19:2817–28. 10.1109/TCBB.2021.3089417 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100. Lavanya C, Pooja S, Abhay HK. Novel biomarker prediction for lung cancer using random forest classifiers. Cancer Informat 2023;22:11769351231167992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101. Cinelli M, Sun Y, Best K et al. Feature selection using a one dimensional naïve Bayes’ classifier increases the accuracy of support vector machine classification of CDR3 repertoires. Bioinformatics 2017;33:951–5. 10.1093/bioinformatics/btw771 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102. Deng L, Sui Y, Zhang J. XGBPRH: prediction of binding hot spots at protein–RNA interfaces utilizing extreme gradient boosting. Genes 2019;10:242. 10.3390/genes10030242 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103. Pereira GRC, da Conceição LMA, Abrahim-Vieira BA et al. XGBMUT: predicting the functional impact of missense mutations using an extreme gradient boost classifier. ACS Omega 2025;10:8349–60. 10.1021/acsomega.4c10179 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104. Li Q, He Y, Pan J. CrossFuse-XGBoost: accurate prediction of the maximum recommended daily dose through multi-feature fusion, cross-validation screening and extreme gradient boosting. Brief Bioinform 2024;25:bbad511. 10.1093/bib/bbad511 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105. Zhang D, Chen HD, Zulfiqar H et al. iBLP: an XGBoost-based predictor for identifying bioluminescent proteins. Comput Math Methods Med 2021;2021:1–15. 10.1155/2021/6664362 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106. Zulfiqar H, Yuan SS, Huang QL et al. Identification of cyclin protein using gradient boost decision tree algorithm. Comput Struct Biotechnol J 2021;19:4123–31. 10.1016/j.csbj.2021.07.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107. Zhou L, Wang Z, Tian X et al. LPI-deepGBDT: a multiple-layer deep framework based on gradient boosting decision trees for lncRNA–protein interaction identification. BMC Bioinformatics 2021;22:479. 10.1186/s12859-021-04399-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108. Settelmeier J, Goetze S, Boshart J et al. MultiOmicsAgent: guided extreme gradient-boosted decision trees-based approaches for biomarker-candidate discovery in multiomics data. J Proteome Res 2025;24:2816–31. 10.1021/acs.jproteome.4c01066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109. Chen D, Jacob L, Mairal J. Biological sequence modeling with convolutional kernel networks. Bioinformatics 2019;35:3294–302. 10.1093/bioinformatics/btz094 [DOI] [PubMed] [Google Scholar]
- 110. Graves A, Graves A. Long short-term memory. Supervised sequence labelling with recurrent neural networks 2012;385:37–45. 10.1007/978-3-642-24797-2_4 [DOI] [Google Scholar]
- 111. Lamurias A, Sousa D, Clarke LA et al. BO-LSTM: classifying relations via long short-term memory networks along biomedical ontologies. BMC Bioinformatics 2019;20:10. 10.1186/s12859-018-2584-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112. Van Houdt G, Mosquera C, Nápoles G. A review on the long short-term memory model. Artif Intell Rev 2020;53:5929–55. 10.1007/s10462-020-09838-1 [DOI] [Google Scholar]
- 113. Goodfellow I, Pouget-Abadie J, Mirza M et al. Generative adversarial networks. Commun ACM 2020;63:139–44. 10.1145/3422622 [DOI] [Google Scholar]
- 114. Gui J, Sun Z, Wen Y et al. A review on generative adversarial networks: algorithms, theory, and applications. IEEE Trans Knowl Data Eng 2021;35:3313–32. 10.1109/TKDE.2021.3130191 [DOI] [Google Scholar]
- 115. Gonog L, Zhou Y. A review: Generative adversarial networks. In: 14th IEEE conference on industrial electronics and applications (ICIEA), vol 2019, pp. 505–10. Xi'an, China, 2019. https://ieeexplore.ieee.org/document/8833686 [Google Scholar]
- 116. Vaswani A, Shazeer N, Parmar N et al. Attention is all you need. Adv Neural Inf Proces Syst 2017;30:6000–10. https://dl.acm.org/doi/10.5555/3295222.3295349 [Google Scholar]
- 117. Harrer S. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine 2023;90:104512. 10.1016/j.ebiom.2023.104512 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118. Zhang Y, Liu C, Liu M et al. Attention is all you need: utilizing attention in AI-enabled drug discovery. Brief Bioinform 2023;25:bbad467. 10.1093/bib/bbad467 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 119. Koroteev MV. BERT: a review of applications in natural language processing and understanding. arXiv preprint arXiv:2103.11943. 2021. 10.48550/arXiv.2103.11943 [DOI] [Google Scholar]
- 120. Hao Y, Dong L, Wei F et al. Visualizing and understanding the effectiveness of BERT. arXiv preprint arXiv:1908.05620. 2019. 10.48550/arXiv.1908.05620 [DOI] [Google Scholar]
- 121. Zhang Z, Han X, Liu Z et al. ERNIE: enhanced language representation with informative entities. arXiv preprint arXiv:1905.07129. 2019. 10.48550/arXiv.1905.07129 [DOI] [Google Scholar]
- 122. Sun Y, Wang S, Li Y et al. Ernie 2.0: A continual pre-training framework for language understanding. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 34, pp. 8968–75. New York, USA, 2020. 10.1609/aaai.v34i05.6428 [DOI] [Google Scholar]
- 123. Sun Y, Wang S, Feng S et al. Ernie 3.0: large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137. 2021. 10.48550/arXiv.2107.02137 [DOI] [Google Scholar]
- 124. Wang S, Sun Y, Xiang Y et al. Ernie 3.0 titan: exploring larger-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2112.12731. 2021. 10.48550/arXiv.2112.12731 [DOI] [Google Scholar]
- 125. Zulfiqar H, Guo Z, Grace-Mercure BK et al. Empirical comparison and recent advances of computational prediction of hormone binding proteins using machine learning methods. Comput Struct Biotechnol J 2023;21:2253–61. 10.1016/j.csbj.2023.03.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126. Zulfiqar H, Ahmed Z, Ma CY et al. Comprehensive prediction of lipocalin proteins using artificial intelligence strategy. Front Biosci Landmark 2022;27:84. 10.31083/j.fbl2703084 [DOI] [PubMed] [Google Scholar]
- 127. Bakanina Kissanga G-M, Zulfiqar H, Gao S et al. E-mula: an ensemble multi-localized attention feature extraction network for viral protein subcellular localization. Information 2024;15:163. 10.3390/info15030163 [DOI] [Google Scholar]
- 128. Xie H, Wang L, Qian Y et al. Methyl-GP: accurate generic DNA methylation prediction based on a language model and representation learning. Nucleic Acids Res 2025;53:gkaf223. 10.1093/nar/gkaf223 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129. Yan K, Lv H, Shao J et al. TPpred-SC: multi-functional therapeutic peptide prediction based on multi-label supervised contrastive learning. Sci China Inf Sci 2024;67:212105:1–12. 10.1007/s11432-024-4147-8 [DOI] [Google Scholar]
- 130. Huang Z, Xiao Z, Ao C et al. Computational approaches for predicting drug-disease associations: a comprehensive review. Front Comput Sci 2025;19:1–15. 10.1007/s11704-024-40072-y [DOI] [Google Scholar]
- 131. Huang Z, Guo X, Qin J et al. Accurate RNA velocity estimation based on multibatch network reveals complex lineage in batch scRNA-seq data. BMC Biol 2024;22:290. 10.1186/s12915-024-02085-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 132. Guo X, Huang Z, Ju F et al. Highly accurate estimation of cell type abundance in bulk tissues based on single-cell reference and domain adaptive matching. Adv Sci 2024;11:e2306329. 10.1002/advs.202306329 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133. Jin G, Xu M, Zou M et al. The processing, gene regulation, biological functions, and clinical relevance of N4-acetylcytidine on RNA: a systematic review. Mol Ther Nucleic Acids 2020;20:13–24. 10.1016/j.omtn.2020.01.037 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
There is no data.














































