Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 May 25;27(3):bbag263. doi: 10.1093/bib/bbag263

ac4C modification sites prediction in human mRNA: a complete review

Hasan Zulfiqar 1,✉, Ramala Masood Ahmad 2,3, Hao Lin 4,✉, Xiao-Long Yu 5,✉
PMCID: PMC13200550  PMID: 42184113

Abstract

ac4C alteration in RNA is a conserved epigenetic mark that is critical for post-transcriptional control, mRNA stability, translational efficiency, and human immune function regulation. In the meantime, the conventional experimental procedures for predicting ac4C alteration sites are costly, time-consuming, and difficult. The precise recognition of ac4C modification sites in human mRNA has been greatly aided by computational prediction techniques using sequence data, machine learning (ML), deep learning (DL), and large language models (LLMs). The application of ML, DL, and LLM-based techniques for the identification of ac4C modification sites in human mRNA has been evaluated and contrasted in this review. Distinctively, we have also addressed the shortcomings of the existing methods and tools, as well as potential future developments. We anticipate that this study will provide sufficient information and awareness for ac4C modification research.

Keywords: N4-acetylcytidine, deep learning, large language models, feature extraction, computational tools

Introduction

N-acetyltransferase 10 (NAT10) catalyzes the addition of an acetyl group to the nitrogen at the fourth position of the cytidine base in N4-acetylcytidine (ac4C), one of the RNA posttranscriptional alterations, as illustrated in Fig. 1. The ac4C modification was primarily identified in transfer ribonucleic acid (tRNA) and ribosomal ribonucleic acid (rRNA) and is conserved in both bacterial and eukaryotic nucleic acids. It was later shown that the existence of ac4C on tRNA preserves tRNA stability and contributes to the high precision of protein translation. Furthermore, researchers also found that the thermal stability of tRNA is most important when ac4C is paired with guanine [1–5]. On the contrary, ac4C on rRNA is essential for preserving the precision of protein translation. Dysregulation of ac4C alterations has been linked to several illnesses, and these modifications serve crucial regulatory functions in gene expression. For example, ac4C alteration of the transcripts of the TGF-β pathway enhances tumor growth in hepatocellular carcinoma [6–9]. Similar to this, abnormal ac4C patterns on neural transcripts impact synaptic function in Alzheimer’s disease [3, 10–13]. Accurately locating ac4C sites throughout the genome is crucial for comprehending illness mechanisms and creating diagnostic tools in light of these disease connections. Reliable computational prediction techniques are necessary, though, because experimental identification is expensive and time-consuming.

Figure 1.

Schematic illustration of ac4C modification in RNA molecules.

Depiction of ac4C alteration in RNA.

Arango et al. [14] recently employed an ac4C-specific acRIP-seq technique to identify more than 4000 ac4C spikes in the transcriptome of human HeLa cells. The acRIP-seq approach yields a sequencing map that offers the greatest number of areas, but it is not single-base resolution. They showed that the human transcriptome has a large number of ac4C modification sites, the majority of which are found in coding regions. Additionally, ac4C-modified mRNAs have a much longer half-life than all other transcripts, particularly when mRNA acetylation is located within moving cytidine, and their translation efficacy is greatly increased.

Experimental and computational techniques

According to numerous studies, ac4C is a major mutation that is critical for posttranscriptional control, mRNA stability, translational efficiency, and regulation of human immune function [10, 15–18]. Consequently, it is crucial to identify ac4C to comprehend its control mechanism. Ac4C sites can be identified using various biochemical methods, including single-molecule real-time sequencing, reduced-representation bisulfite sequencing, and mass spectrometry. When applied to large sequencing data, these methods are very costly, even if they are very useful in identifying ac4C sites.

Reduced representation Bisulfite sequencing

The genome-wide acetylcytidine profiles have been examined at the single-nucleotide level using reduced representation bisulfite sequencing, an effective and high-throughput method. This method enriches genomic areas with high CpG content by combining bisulfite sequencing with acetylcytidine-independent restriction enzymes. As a result, only 1% of the genome’s nucleotides may need to be sequenced. The bulk of promoters and repetitive sequences that are challenging to profile using traditional bisulfite sequencing techniques are still present in the fragments that make up the truncated genome [19–21]. Figure 2A displays the reduced representation bisulfite sequencing flowchart.

Figure 2.

Diagram showing three experimental techniques for detecting N4-acetylcytidine, including bisulfite sequencing, single-molecule real-time sequencing, and mass spectrometry.

Traditional experimental techniques to detect N4-acetylcytidine. Reduced representation bisulfite sequencing (A). Single-molecule real-time sequencing (B). Mass spectrometry (C).

Single-molecule real-time sequencing

A concurrent method for sequencing single-molecule DNA is called SMRTS. Strand displacement amplification (SDA) and multiple displacement amplification (MDA), which are based on rolling circle amplification (RCA), use unique loop adapters to produce ssDNA from dsDNA fragments. The DNA polymerase then adds dNTP containing fluorescent phosphate groups, which cleave the phosphate chain and emit light, thereby eliminating the fluorescent dye from the expanding nucleic acid chain. The sequencing reads are produced by parallel single-molecule real-time sequencing processes on thousands of nano-photonic visualization chambers [20, 22–25]. Figure 2B depicts the entire single-molecule real-time sequencing method.

Mass spectrometry

Mass spectrometry is an analytical technique that measures the mass-to-charge ratio (m/z) of several molecules in a sample. This technique combines MS detection of the non-enzymatic hydrolysates of the target DNA/RNA with selective capture of the DNA/RNA target from restricted cleavage of genomic DNA/RNA using magnetic separation. Figure 2C illustrates how the mass spectrometer operates. Furthermore, this technique has a special benefit in precisely measuring the quantity of RNA acetylcytidine. Because of its high selectivity, high resolution, and high mass accuracy, liquid chromatography in conjunction with high-resolution mass spectrometry has been widely employed for the quantitative analysis of bio-samples. This method has been effectively applied in several labs over the last 10 years to determine the genome-wide ac4C detection in various animals [26–28].

Computational techniques

Meanwhile, wet-lab methods are costly and less effective in identifying ac4C modification sites. Consequently, bioinformatics techniques using sequence information and artificial intelligence-based algorithms are much needed to correctly recognize the ac4C modification sites [29–32]. Figure 3 shows the machine and deep learning-based methods that have been developed over the past 7 years to predict human ac4C alteration sites.

Figure 3.

Timeline illustrating key developments in methods for predicting ac4C modification sites.

Timeline of historical landmarks for the prediction of ac4C modification sites.

The establishment of the previous methods consists of four crucial steps such as dataset construction, feature engineering, model development, and model evaluation. The whole steps for the construction of an artificial intelligence-based model are shown in Fig. 4.

Figure 4.

Workflow diagram outlining the main steps in developing an ac4C prediction tool, including dataset construction, feature engineering, model development, and evaluation.

Building a suitable tool for ac4C prediction involves the following steps: dataset construction, feature engineering, model development, and model evaluation.

In this brief review, we introduce machine learning (ML), deep learning (DL), and large language model (LLM)-based tools for the identification of ac4C modification sites in the human mRNA transcriptome. In 2019, Zhao et al. [33] established the first random forest-based predictor to predict ac4C in the human mRNA transcriptome called PACES, in which they used six types of feature encoding schemes: one-hot, position-specific nucleotide sequence profile, k-spaced nucleotide pair frequencies, position-specific di-nucleotide sequence profile, k-nucleotide frequencies, and pseudo k-tuple nucleotide composition, and input them into a random forest-based classifier to predict ac4C. In 2020, Alam et al. [34] developed an extreme gradient boosting-based classifier to predict ac4C called XG-ac4C. In their technique, they used four types of encoding schemes, one-hot, electron-ion interaction of pseudopotentials (EIIP + PseEIIP), k-mer, nucleotide chemical property, and nucleotide density, and inserted them into five ML-based classifiers, namely, random forest, Gaussian naïve bayes, logistic regression, adaboost, and extreme gradient boost, and finally selected the best performing model to classify ac4C from non-ac4C. In 2021, Wang et al. [35] developed a convolutional neural network (CNN)-based method to predict ac4C. In their method, they used six physicochemical feature descriptors, namely, k-mer, composition of k-spaced nucleotide pairs (CKSNAP), series correlation pseudo di-nucleotide composition (SCPseDNC), series correlation pseudo tri-nucleotide composition, pseudo k-tuple composition (SCPseTNC), electron-ion interaction of pseudopotentials (PseEIIP) of tri-nucleotide with word2vector embedding. At first, six physicochemical features were improved with a support vector machine and F-score with a sequential forward search strategy. Then, these improved features were fed into the CNN for training and testing the model. In 2022, Zhang et al. [36] introduced a CNNLSTMac4CPred method for the prediction of ac4C modification sites. In their method, they utilized three types of features. The two feature descriptors are traditional, namely, pseudo tri-nucleotide composition (PseTNC) and k-nucleotide frequencies (KNF), and the other one is a semantic feature. They utilized CNN and LSTM-based neural networks to mine semantic features from sequences. After this, these features were inserted into an extreme gradient boosting-based classifier to predict ac4C modification. In 2023, a total of six methods were developed by authors from different countries. Jia et al. [37] developed a DLC-ac4C method in which they utilized three types of feature encodings, namely, nucleotide chemical property (NCP), nucleotide density (ND), and C2 encoding. They used 1D-CNN and BiLSTM for learning the local and global features, respectively. They also utilized a channel attention mechanism to find the optimal sequence characteristics and a homo-morphemic integration technique to limit the generalization error of the model. Lou et al. [38] constructed a stacking-ac4C method to classify N4-acetylcytidine in human mRNA. In their method, they used three types of encoders, namely, k-mer, pseudo-k-tuple nucleotide composition (PseKNC), and electron–ion interaction pseudo-potential (PseEIIP), and their hybrid fusion features, and then developed a stacking model to train and test the model. Liu et al. [39] presented a new model, TransC-ac4C, to predict ac4C modification sites in mRNA. In their model, they utilized five types of feature descriptors, namely, one-hot encoding, ND, k-mer, NCP, and EIIP, and inserted them into a hybrid model that consists of CNN and transformer. This technique overcomes the issues of missing long sequence dependence and avoids the loss of correlation information between discontinuous sequences. Su et al. [40] developed an iRNA-ac4C to classify ac4C from non-ac4C in human mRNA. In their methodology, they used hybrid features of k-mer, accumulated nucleotide frequency, and nucleotide chemical frequency, and then improved these features with minimum redundancy and maximum relevance (mRMR) and incremental feature selection (IFS) technique. Finally, they input these optimal features into seven different traditional ML classifiers, namely, logistic regression (LR), naive bayes (NB), gradient boost decision tree (GBDT), k-nearest neighbor (KNN), support vector machine (SVM), random forest (RF), and AdaBoost (AB) and select the best performing model (GBDT) to classify the ac4C from non-ac4C. Lai et al. [41] introduced a new tool, LSA-ac4C, for the accurate prediction of ac4C modification sites in human mRNA. In this technique, they used a hybrid neural network that combines an LSTM double-layer neural network with self-attention for the prediction of ac4C modification. Jia et al. [42] developed a new model, EMDL-ac4C. In this technique, they used one-hot encoding and then inserted it into an ensemble model with a two-branch residual connection dense net and attention to predict ac4C modification sites in human mRNA.

In 2024, a total of eight methods were developed by authors from different countries. Li et al. [43] introduced a MetaAc4C model in which they utilized pretrained bi-directional encoder representations (BERT), and the model is based on the bi-directional long short-term memory (BLSTM) network. Li et al. [44] also introduced a new method based on generative adversarial networks and transfer learning to recognize ac4C modification sites in human mRNA, named GANSamples-ac4C. Pham et al. [45] constructed a new tool, named ac4C-AFL, which is based on adaptive feature representation learning. In their method, they utilized 16 types of feature encoding schemes k-spaces nucleic acid pair (CKSNAP), enhanced nucleic acid composition (ENAC), position specific of 2 nucleotides (PS2), PseEIIP, Z-curve, k-mer, reverse k-mer complement (RCKmer), di-nucleotide physicochemical properties type 1-2 (DPCP1-2), NCP, binary features, multivariate mutual information and accumulated nucleotide frequency (MMNF), word2vec (W2V), sequence2vector (S2V), DNABERT and combination of skip di-nucleotide composition and local position-specific di-nucleotide frequency (ASLPN) and then inputted these optimized features into 11 different types of ML and DL classifiers, namely, SVM, RF, extremely randomized tree (ERT), artificial neural network (ANN), logistic regression (LR), GBT, XGBT, light GBT (gradient boost decision tree), cat boost (CB), adaboost (AB), and CNN classifiers and generated 176 baseline models. Finally, they selected the best baseline models using a 2-step feature selection technique, whose predicting scores were integrated and trained with SVM to develop the final model to classify ac4C from non-ac4C in human mRNA. Yi et al. [46] developed a new method, named STM-ac4C, which is a hybrid model based on selective kernel convolutions, a temporal convolutional network with multi-head self-attention to predict ac4C modification sites in mRNA. Yuan et al. [47] developed a new method based on a dual path neural network with self-attention, namely, DPNN-ac4C, to discriminate ac4C from non-ac4C in human mRNA. Their method integrates a convolutional neural network, embedding modules, bi-directional gated recurrent unit (BGRU) with self-attention to extract local and global features from mRNA sequences. Liu et al. [48] developed a novel transformer-based model called TransAC4C to predict ac4C modification sites in mRNA. Their model was divided into four parts such as a transformer layer, one BLSTM layer, four 1D convolutional layers, with three fully connected layers. Transformer and BLSTM layers were used to enable the model to learn the contextual information, convolutional layers were used to enable the model to extract important features from the input, and fully connected layers were used to connect the input features into the output. Jia et al. [49] constructed a new method called Voting-ac4C to predict ac4C modification sites in mRNA. In their method, they utilized RNAErnie, a transformer-based pretrained model with six traditional feature encoding schemes, such as one-hot, ENAC, ND, C2 encoding, k-spaced nucleotide pair frequencies (KSNPF), and physicochemical properties (TPCP), and then inserted the hybrid of these features into a deep neural network (DNN) for dimensionality reduction. Finally, these dimensionality reduction features were fed into a voting ensemble model constructed using CatBoost, XGBoost, and a multilayer perceptron (MLP) classifier. He et al. [50] developed a deep learning method called NBCR-ac4C based on pretrained models to predict ac4C modification sites in human mRNA. They utilized DNABERT2 and nucleotide transformer to construct contextual embedding of nucleotide sequences and then applied CNN and ResNet18 to further mine the trivial and deep information from the contextual embedding.

In 2025, a total of two predictors were proposed to date by the authors from different countries. Lu et al. [51] introduced an ERNIE-ac4C method to predict ac4C modification sites. In their technique, they used the pretrained model ERNIE-RNA to extract the attention map features and sequence features from nucleotide sequences. They input these fused features into a 2d CNN to predict ac4C modification sites in human mRNA. Yao et al. [52] proposed a DL-based method called Caps-ac4C to predict ac4C modification sites in human mRNA. In their method, they utilized chaos game representation (CGR) encoding to transform the nucleotide sequences into visual representations and then employed a capsule network design to extract the local and global features from these visual representations of nucleotide sequences. The brief description of feature extraction and selection, model evaluation techniques, and web server availability of the prediction methods are summarized in Table 1.

Table 1.

A list of available methods for the prediction of ac4C is summarized in this review.

Method Classifiera Featureb Evaluation Webserver Status Year
PACES [33] RF One-hot,
PSNSP, PSDSP, KNF, KSNPF,
PseKNC
5-fold CV Yes Active 2019
XG-ac4C [34] XGB One-hot, NCP-ND, k-mer, EIIP-PseEIIP 5-fold CV Yes Not-active 2020
DeepAc4C [35] CNN k-mer, CKSNAP, SCPseDNC, SCPseTNC, PseKNC, PseEIIP, W2V 10-fold CV Yes Not-active 2021
CNNLSTMac4CPred [36] XGB Semantic
features, KNF, PseTNC
5-fold CV Yes Not-active 2022
DLC-ac4C [37] EL C2, NCP, ND 10-fold CV No – 2023
Stacking-ac4C [38] SIC k-mer, PseEIIP, PseKNC 10-fold CV No – 2023
TransC-ac4C [39] CNN
+
Transformer
One-hot, NCP, ND, EIIP, k-mer 5-fold CV No – 2023
iRNA-ac4C [40] GBDT k-mer, NCP, ANF 10-fold CV Yes Not-active 2023
LSA-ac4C [41] Double-layer
LSTM
+
Self-attention
Embedding 10-fold CV Yes Not-active 2023
EMDL-ac4C [42] EM + DL One-hot 10-fold CV No – 2023
MetaAc4C [43] LSTM
+
Attention
+
Residual
BERT embedding
+
GAN
10-fold CV No – 2024
GANSamples-ac4C [44] GAN One-hot, ANF, NCP, CKSNAP,
ENAC, NAC, EIIP, k-mer,
RCKmer
5-fold CV No – 2024
ac4C-AFL [45] SVM ENAC, PS2, CKSNAP,
DNABERT, PseEIIP, Z-curve, S2V,
k-mer, W2V,
RCK-mer, DPCP_1, DPCP_2, NCP, BPF, MMNF, ASLPN
10-fold CV Yes Active 2024
STM-ac4C [46] SKC + TCN + MHSA One-hot 10-fold CV No – 2024
DPNN-ac4C [47] (Bi-directional GRU + Attention) + (CNN + self-attention) PseKNC, IES 10-fold CV No – 2024
TransAC4C [48] Transformer Embeddings 10-fold CV No – 2024
Voting-ac4C [49] EL RNAErnie + One-hot + ND + C2 + ENAC + TPCP + KSNPF 10-fold CV Yes Not-active 2024
NBCR-ac4C [50] CNN + Resnet18 Nucleotide Transformer embedding + DNABERT2 embedding 10-fold CV No – 2024
ERNIE-ac4C [51] 2d-CNN Bert + AMF 10-fold CV No – 2025
Caps-ac4C [52] CN CGR 5-fold CV Yes Active 2025

aSIC, Stacking Integration Classifier; EL, ensemble learning; GAN, generative adversarial network; SVM, support vector machine; CNN, convolutional neural network; XGB, eXtreme gradient boosting; GBDT, gradient boosting decision tree; RF, random forest; LSTM, long short-term memory; EL, ensemble learning; EM, ensemble; DL, deep learning; SKC, selective kernel convolution; TCN, temporal convolutional networks; MHSA, multi-head self-attention mechanism; CN, capsule networks.

bBERT, bidirectional encoder representations from transformers; GAN, generative adversarial networks; W2V, Word2vector; CKSNAP, composition of k-spaced nucleic acid pairs; PseTNC, Pseudo tri-nucleotide composition; ENAC, enhanced nucleic acid composition; NAC, nucleic acid composition; NCP, nucleotide chemical property; PseKNC, Pseudo k-tuple nucleotide composition; PseEIIP, electron-ion interaction pseudopotentials; RCK-mer, reverse complement k-mer composition; ND, nucleotide density; S2V, sequence2vector; ANF, accumulated nucleotide frequency; MMNF, multivariate mutual nucleotide frequency; DPCP, di-nucleotide physicochemical properties; ASLPN, adaptive skipped local position-specific di-nucleotide composition; SCPseDNC, series correlation pseudo di-nucleotide composition; SCPseTNC, series correlation pseudo tri-nucleotide composition; AMF, attention map feature; IES, integer encoding of sequence; CGR, chaos game representation; ERNIE, enhanced representation through knowledge integration; ND, nucleotide density; C2, common sequence characterization; TPCP, physicochemical properties; KSNPF, k-spaced nucleotide pair frequencies.

Several tools were developed for ac4C alteration sites employing the aforementioned techniques. The development of computational biology and bioinformatics will benefit from the substantial evidence that each tool provides for the identification of ac4C modification sites in human mRNA. Several elements of these tools are summarized in this study, including model evaluation, ML, DL, and LLM-based classifiers, feature extraction and selection techniques, and dataset development. As a result, this evaluation can help researchers select the best tool for ac4C prediction and offer recommendations for future advanced AI-based methods.

Dataset construction

The construction of the dataset is the first important step in the development of a bioinformatics tool. Nowadays, a large amount of data related to RNA is publicly available in open-source databases such as Rfam [53–55] and RMBase3.0 [56]. Many credible datasets were also produced from previously published literature and databases to develop ML, DL, and LLM models. Therefore, we presented some simple steps that have been widely utilized to build the datasets for ac4C modification prediction in mRNA. (i) Extracting sequence-based data from previous literature and databases. (ii) Applied cluster database at high identity with tolerance to eradicate similarity between the sequences. (iii) Applied labeling on positive and negative sequences, such as 1 for positive ac4C and 0 for negative ac4C. We have observed that all the training data for ac4C modification sites were retrieved from previous literature. Arango et al. [14] observed the distribution of various CXX motifs both inside and outside the acetylation peak; they selected positive and negative sequences from 2134 genes from previous literature. They varied the number of repetitions of CXX from 2 to 9. Simple repeating motifs CXXCXX could be frequently observed in the transcriptome; they discovered five CXX motifs occurred in 1629 peaks, and 15 198 five consecutive CXX motifs found outside the peak. Finally, they divided the dataset into two parts: 1160 ac4C sequences and 10 855 non-ac4C sequences for training purposes, and 469 ac4C sequences and 4343 non-ac4C sequences for independent testing. The dataset can be downloaded at http://www.rnanut.net/paces. After this, Wang et al. [35] constructed a balanced dataset with a 40% similarity threshold and consisted of 1148 sequences of ac4C and 1148 sequences of non-ac4C for training purposes and 467 sequences of ac4C and 467 sequences of non-ac4C for independent testing. The dataset can be downloaded at https://zenodo.org/records/5138047. Su et al. [40] also constructed a new balanced dataset based on Arango et al. [14] data. They picked the cytidines nearest to ac4C peaks as alteration sites in order to create a trustworthy dataset. They then collected 100 nucleotides on either side of these alteration sites as positive samples, using them as the center. Negative samples from non-peak areas were chosen at random. The central sites are cytidine, and their sequences were 201 nucleotides. In their dataset, they set an 80% similarity threshold to remove the redundant sequences. Finally, they retrieved 2206 ac4C sequences and 2206 non-ac4C sequences for training and 552 ac4C sequences and 552 non-ac4C sequences for independent testing. The dataset can be downloaded at http://lin-group.cn/server/iRNA-ac4C. Lu et al. [51] built a new dataset of 1480 ac4C sequences and 1480 non-ac4C sequences for training and 185 ac4C sequences and 185 non-ac4C sequences for independent testing. Due to the lack of dataset details, this dataset was excluded from the current evaluation. Therefore, we have selected the datasets of Su et al. [40], Wang et al. [35], and Arango et al. [14] in this review to assess and validate the prediction efficiency of the computational tools. Table 2 displays all the dataset details.

Table 2.

A benchmark classification data for ac4C.

Method Ratio
(P/N)
Training set (P/N) Independent set (P/N) Data-source CD-HIT Year
PACES [33] 1:10 1160/10855 469/4343 Arango et al. – 2019
XG-ac4C [34] 1:10 1160/10855 469/4343 Arango et al. – 2020
DeepAc4C [35] 1:1 1148/1148 467/467 Wang et al. 40% 2021
CNNLSTMac4CPred [36] 1:10 1160/10855 469/4343 Arango et al. – 2022
DLC-ac4C [37] 1:1 2206/2206 552/552 Su et al. 80% 2023
Stacking-ac4C [38] 1:1 2206/2206 552/552 Su et al. 80% 2023
TransC-ac4C [39] 1:1 1148/1148 467/467 Wang et al. 40% 2023
iRNA-ac4C [40] 1:1 2206/2206 552/552 Arango et al. 80% 2023
LSA-ac4C [41] 1:1 2206/2206 552/552 Su et al. 80% 2023
EMDL-ac4C [42] 1:1 1148/1148 467/467 Wang et al. 40% 2023
MetaAc4C [43] 1:1 1148/1148 467/467 Wang et al. 40% 2024
GANSamples-ac4C [44] 1:1 2206/2206 552/552 Su et al. 80% 2024
ac4C-AFL [45] 1:1 2206/2206 552/552 Su et al. 80% 2024
STM-ac4C [46] 1:1 2206/2206 552/552 Su et al. 80% 2024
DPNN-ac4C [47] 1:1 2206/2206 552/552 Su et al. 80% 2024
TransAC4C [48] 1:1 1148/1148 467/467 Wang et al. 40% 2024
Voting-ac4C [49] 1:1 2206/2206 552/552 Su et al. 80% 2024
NBCR-ac4C [50] 1:1 2206/2206 552/552 Su et al. 80% 2024
ERNIE-ac4C [51] 1:1 1480/1480 185/185 Lu et al. – 2025
Caps-ac4C [52] 1:1 2206/2206 552/552 Su et al. 80% 2025

“P” denotes ac4C and “N” denotes non-ac4C.

Feature engineering

Extracting and selecting the self-directed and informative feature is an essential phase in ML and DL-based methods [57–62]. Therefore, several types of feature extraction and selection methods were employed to explain the ac4C modification in human mRNA.

Feature extraction

k-mer

The short-range nucleotide interactions between sequences can be represented by a k-mer. Using a sliding window method, the (N-k + 1) nucleotide residues can be found by examining a sequence with N bp and setting the size of the window to k bp with 1 bp step size [63–65]. With a sequence length of N, an arbitrary sample M can be described as:

graphic file with name DmEquation1.gif (1)

where the nucleotide (A, G, U, and C) at position i is denoted by Ri. Using the k-mer nucleotide composition, the sequences can be converted into the 4k-D vector as follows:

graphic file with name DmEquation2.gif (2)

where T represents the vector’s transposition and f1k-tuple characterizes the sequence occurrence of the i-th k-mer nucleotide composition. An RNA sample can be decoded into a 4-D vector M1 = [f(A), f(G), f(C), f(U)]T when k = 1. A 16-dimensional vector can be used to describe the RNA sample when k = 2.

CKSNAP

The frequency of nucleotide pairings split apart by any k nucleotide (k = 0, 1, 2, 3, 4, 5) is represented by the CKSNAP. Sixteen nucleotide pairs [AA, AG, …,UG, UU] make up the features of k-spaced nucleic acid pairings. Using k = 1 as an example, the k-spaced nucleic acid pair composition can be expressed as follows:

graphic file with name DmEquation3.gif (3)

where * denotes (A, C, G, and U), NTotal represents the total number of single-spaced nucleotide pairs in the sequence, and NY*ZInline graphic indicates the nucleotides Y*Z pairings length in the sequence. The composition of the nucleic acid pair AA can be calculated by dividing its frequency (j) by the total number of 0-spaced nucleic acid pairs (NTotal) in the nucleotide sequence. The value of NTotal for a nucleotide sequence of length “g” is g-1, g-2, g-3, g-4, g-5, and g-6 for (k = 0, 1, 2, 3, 4, 5) [58, 66–68].

ENAC

Based on the sequence window, the ENAC evaluates the nucleic acid composition. A sequence of equal length can be created with it. The ENAC can be calculated as:

graphic file with name DmEquation4.gif (4)

Equation (4) uses k to describe the sliding window’s size, NA, win to indicate how many nucleotides A are present in the window g (g = 1, 2, …, L-k + 1), and T  Inline graphic [G, C, A, U].

NCP

Based on their characteristics, four types of nucleotides in RNA sequence fragments can be divided into three groups, which are hydrogen bonds, ring structure, and functional groups. Numerous nucleotide modification site prediction applications have made extensive use of this classification. More precisely, A and G belong to the keto group, whereas A and U belong to the amino group; C and G can form strong hydrogen bonds, while A and U can form weak ones; and C and U have single-ring structures, while Adenine and Guanine both have double-ring structures. Therefore, coordinates (1,1,1), (0,1,0), (1,0,0), and (0,0,1) can be used to represent A, U, G, and C based on these chemical properties, respectively. Since cytidine is the core site of all sequences, we only retrieved NCP information from the adjacent sequences of ac4C peaks since it is not required to obtain NCP features of this site [69].

ANF

The distribution of every nucleotide in the RNA sequence, as well as nucleotide frequency information, is provided by ANF [70]. In an RNA sequence, the density di of any nucleotide si at position i is stated as follows:

graphic file with name DmEquation5.gif (5)

where l is the entire length of the sequence and Inline graphic is the size of the i-th preface sequence from the 1st to the i-th position.

One-hot encoding

In sequence prediction tasks, such as RNA sequence analysis, one-hot encoding is a prominent technique. The four nucleotides that make up the language of RNA molecules are adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleotides are mapped to the binary vectors (1, 0, 0, 0), (0, 0, 1, 0), (0, 1, 0, 0), and (0, 0, 0, 1), respectively [71]. The resulting one-dimensional one-hot encoded representation for an RNA sequence of length L would be shaped like (1, L × 4). As an alternative, the shape of (4, L) is produced via the 2D one-hot encoding.

CGR

CGR encoding, which has been extensively used to represent biological sequences in DNA, RNA, and proteins to transform sequence data into illustration in a 2D space. Sequences are represented by CGR encoding as a collection of points in a 2D space, with each point’s location and distribution reflecting the sequence’s properties [72]. In addition to being intuitive, this representation can capture intricate patterns that are present throughout the sequence.

IES

RNA sequence bases are mapped to integers by integer encoding. Integer encoding efficiently examines the contextual relationships inside sequences when combined with a bidirectional GRU network. Notably, it enhances the GRU network’s sequence analysis skills by naturally representing the positional order correlations between bases [73]. A basic dictionary with discrete integer values assigned to each symbol is defined as the part of the implementation process. Let M = S1, S2, S3,….SL represents the RNA sequence M with L bases, where Si = {A, U, C, G}. Through a methodical mapping of A to 0, C to 1, G to 2, and U to 3.

RCKmer

A particular kind of k-mer known as RCKmer is created by eliminating the complementary reverse k-mers for every kind of k-neighboring nucleic acid in the specified sequence [74]. For example, when k = 2, the k-mers “GG,” “CG,” “GU,” “UG,” “UC,” and “UU” should be eliminated, leaving 10 discriminatively surviving k-mers in the RCKmer.

PseEIIP

Tri-nucleotides in the mRNA sequence are assigned PseEIIP according to an average EIIP value for each nucleotide [75].

graphic file with name DmEquation6.gif (6)

where flmn is the standardized incidence of the i-th tri-nucleotide, EIIPlmn = EIIPl + EIIPm + EIIPn, and l, m, n ∈ {A, G, C, U}.

DNABERT

Based on the relative interactions among k-mers in the input sequence, DNABERT extracts sequence information. In order to efficiently examine the complex relationships between nucleotides in both forward and reverse orientations, the DNABERT model employs the capabilities of a transformer architecture. By using a k-mer-based tokenizer, DNABERT creates unique tokens from genomic sequences known as “k-mers,” which are further processed through 12 transformer-based blocks. These blocks produce detailed representations of the genomic sequence that capture the appropriate information contained in the inherent material by analyzing the interaction between k-mers [76–78]. Since RNA and DNA both differ by only one nucleotide and retain nearly the same genetic information, DNABERT has notably demonstrated relevance to RNA. This enables handling of RNA sequences by merely substituting nucleotide “T” for nucleotide “U”.

W2V

By taking into account the background in which words occur within a sizable quantity of text, W2V is a neural network-driven model that effectively encodes linguistic and semantic word associations. Both Continuous Bag-of-Words (CBOW) and Skip-gram methods are included in the Word2Vec model. While Skip-gram estimates environmental contextual words for a specific targeted word, CBOW anticipates the desired word using its context words. Both models are trained on extensive textual data sets and used separately within the Word2Vec architecture [79]. That makes it possible to transform words into vector illustrations that express semantic connections across the entire vocabulary.

Feature selection

Selecting the features is a vital step after obtaining the meaningful features. The ideal features could enhance the model’s performance vividly. Presently, there are many feature selection methods used in computational biology, such as IG, ANOVA, mRMR, and PCA [80–83]. The mathematical manifestation is given below.

Binomial distribution

BD is a unique feature selection methodology that has been used in sequence analysis. In this methodology, the confidence level of k-mer’s is termed as follows:

graphic file with name DmEquation7.gif (7)

where Inline graphic is the confidence level for m and n, and “n” signifies the positive/negative sample, Inline graphic represents the m-th residue, and Inline graphic symbolizes the occurrence of n-th residue in the sample [84, 85].

Principal component analysis

PCA is a feature dimensionality reduction method and is often utilized in many ML, DL, and data science applications to visualize and analyze the data. It helps to reduce the irrelevant features from the data by keeping the important and meaningful information. It uses linear algebra to convert features into principal components by calculating the eigenvectors and eigenvalues. Finally, it selects the top components with top eigenvalues [86–88].

Analysis of variance

The main purpose of ANOVA is to calculate the fundamental distinctions across two or more groups. The F-score value is used in this process to rate each feature. The percentage of variance within the groups and the difference between groups is known as the F-score. As a result, the characteristic with a high F-score has strong discriminating ability [89].

graphic file with name DmEquation8.gif (8)
graphic file with name DmEquation9.gif (9)

The variance between the values is represented by Inline graphic, and the degree of freedom between them is represented by Inline graphic. The variance within the values is represented by Inline graphic, and the degree of freedom within values is represented by Inline graphic.

graphic file with name DmEquation10.gif (10)

where Inline graphic is the computational dispersal score of alterations among treatments Inline graphicand Inline graphic.

Incremental feature selection

IFS categorized all the features based on theInline graphic-values and produced new feature vectors, which are given in eq. (11).

graphic file with name DmEquation11.gif (11)

The feature with the highest δ-value, Inline graphic is included in the first feature subset. The subsequent feature subset, Inline graphic is generated by adding the second feature subset with the highest δ-value. The third feature subset, Inline graphic is generated by integrating the third feature subset with the highest δ-value. Until every feature was taken into account, the procedure was repeated [90].

Machine learning algorithms

The prediction of ac4C alteration depends heavily on the choice of an appropriate machine learning algorithm. Bioinformatics has made use of several machine learning methods, including support vector machines, adaboost, extreme randomized trees, random forest, naive bayes, gradient boost decision trees, and eXtreme gradient boosting.

Support vector machine

In computational biology, SVM is a popular and significant machine learning algorithm. Finding the optimal hyperplane that optimizes the separation between the neighboring samples from the classes and the hyperplane is the main goal of SVM [91–94]. The SVM algorithm’s mathematical representation is provided as follows:

If Inline graphic is a training incidence and Inline graphic is a target, then

graphic file with name DmEquation12.gif (12)

where “p” indicates the bias and Inline graphic indicates the kernel function. To find a hyperplane, the kernel function can plot the feature vectors in three dimensions. Nonlinear classification issues are the main application for the radial basis function. It might be stated as:

graphic file with name DmEquation13.gif (13)

whereas, Inline graphic can be obtained by using the maximum Lagrangian expression as:

graphic file with name DmEquation14.gif (14)

where (∴ Inline graphic = 0, 0 ≤ Inline graphic ≤ k, k symbolizes the significance constraint).

AdaBoost

Another well-known and extensively utilized tool in bioinformatics is the AdaBoost classifier. It is an ensemble method that improves accuracy by combining different classifiers. Setting the classifier’s weights and training the data in each iteration is the basic notion behind this [95]. It is defined as follows:

graphic file with name DmEquation15.gif (15)

where R is the target, which can be either −1 or 1, j is the data point, and x is the dimension number. AdaBoost determines the significance of each training example in the training dataset by giving it a weight. That set of training data points is likely to have a larger in-training set when the assigned weights are high. In a similar vein, low allocated weights have little effect on the training dataset.

Extreme randomized tree

ERT is a well-known method that utilizes hundreds of decision trees to perform classification. This strategy was initially introduced by Geurts et al. [96] to reduce the model’s discrepancy by incorporating strong randomization techniques. ERT incorporates more robust randomization strategies to further reduce the prediction model’s volatility. With a few small adjustments, like not using the bagging technique to build a tree, ERT is extremely similar to the random forest algorithm. While random forest finds the optimal split by enhancing the variable index between random subclasses of variables and splitting value [97]. It is defined as follows:

graphic file with name DmEquation16.gif (16)

where ϵ is the random variable that encapsulates the random disconcerting scheme, denoted by ϵ, a value of this variable, and hence by Inline graphic. The learning sample is epitomized by ls, and N is the total number of models.

Random forest

A combined knowledge method that is widely used in computational biology is the random forest. It is utilized for clustering, classification, and regression problems, as well as feature selection processes. It is helpful to reduce the bias of the individual weak predictor, avoid variable rescaling, and be more robust to noise and overfitting. It is more effective to handle class imbalance problems by allocating weight to the minority class, and it efficiently handles missing data issues by substituting the variable appearing mostly in a particular node. The basic principle is to integrate numerous weak classifiers [98–100]. Since the voting process determines the conclusion, the model’s result is more precise and straightforward. If the attribute Inline graphic of a vector Inline graphic and weak classifiers based on these attributes is Inline graphic, then it can be demarcated as:

graphic file with name DmEquation17.gif (17)

Naïve Bayes

The NB model is easy to build and effective when dealing with large datasets. This method makes use of direct acyclic graphs, which have many leaves and only one root. When leaf nodes are separated from their base, it presumes stability between them [101]. If the class variable is “v” and the feature vector is Inline graphic:

graphic file with name DmEquation18.gif (18)

By means of independent norms

graphic file with name DmEquation19.gif (19)

After the simplification

graphic file with name DmEquation20.gif (20)

eXtreme gradient boosting

For numerous computational issues, XGBoost, an ensemble learning technique based on gradient boosting, offers cutting-edge outcomes. In essence, XGBoost is an ensemble technique built on a gradient boosted tree [102–105]. According to the formula below, the prediction’s outcome is the total of the scores predicted by k trees:

graphic file with name DmEquation21.gif (21)

where Inline graphic is the i-th number in the training set, and Inline graphic is the k-th tree score. S is the interplanetary function encompassing all gradient boosted trees. The following formula could be used to optimize the objective function:

graphic file with name DmEquation22.gif (22)

where a loss function that assesses the fitness of model prediction ^yi and samples of the training dataset yi is represented by the former Inline graphic. A regularization item that penalizes the model’s complication to prevent overfitting is represented by Inline graphic.

Gradient boosting decision tree

Researchers have utilized the gradient-boosted decision tree algorithm, a highly significant learning technique, across multiple bioinformatics, mathematics, and biological applications [106–108]. From a nonlinear joint of various weak learners, it builds a realistic and climbable model. The gradient boost decision tree’s primary goal is to create a base learner that is highly correlated with the negative gradient loss function. Assume that n samples are available:

graphic file with name DmEquation23.gif
graphic file with name DmEquation24.gif (23)

where Inline graphic is the risk optimization factor of the newly generated decision tree, as indicated by the equation below, and Inline graphic is the new decision tree (k = 1,2,3,…).

graphic file with name DmEquation25.gif (24)

The final evaluation is computed in an anticipatory mode by the gradient boost decision tree technique.

graphic file with name DmEquation26.gif (25)

Subsequently, residuals are calculated using the loss function Inline graphic of negative gradient.

graphic file with name DmEquation27.gif (26)

Ultimately, all Inline graphicto calculate the risk mitigation parameter Inline graphic. were used to train the model. By mapping the values in the input space Inline graphic into Inline graphic split parts Inline graphic, and producing Inline graphic for region Inline graphic, this kind of decision tree intuitively models the relationships among predictor variables.

graphic file with name DmEquation28.gif (27)

Deep learning algorithms

Convolutional neural network

Convolutional, normalizing, connected, pooling, and other layers make up CNN. It has been utilized in many biological and medical advancements. The fundamental concept of CNN is to create an enormous number of filters that can obtain veiled topological features from input using pooling techniques and layer-wise convolutions [82, 109].

graphic file with name DmEquation29.gif (28)

where “t” is the filter index and “b” is the input. The input channel numbers are represented by “n,” the filter size is indicated by “r,” the output filter is represented by “g,” and the convolutional filter with a weight matrix “r × I” is represented by “St”.

graphic file with name DmEquation30.gif (29)

Equation (29) is the symbol of the reluctant activation function consuming b as an input.

graphic file with name DmEquation31.gif (30)

Equation (30) shows a fully connected layer in which “Inline graphic” represents the bias, “Inline graphic” represents the dropout operator with probability of “p,” “Inline graphic” represents the 1D feature vector, and “Inline graphic” characterizes the prior layer weights of “Inline graphic”.

graphic file with name DmEquation32.gif (31)

Equation (31) suggests the final classification layer has a sigmoid activation function and “b” as an input.

Long short-term memory network

LSTM is a unique variant of a recurrent neural network (RNN) that incorporates a storage method made up of frequently linked blocks, replacing the hidden function found in conventional RNNs. LSTM surpasses traditional RNN due to its interconnected memory and multiplication units that enhance its ability to learn from distant dependencies. This advancement has expanded the application scope of RNN and established a foundation for the evolution of the following sequence modeling techniques. In contrast to unidirectional LSTM, bi-directional LSTM (BLSTM) more effectively captures context information within sequences. The attention mechanism has garnered significant attention within the domains of natural language processing and image recognition to enhance model interpretability [110–112]. It enables the model to prioritize more critical information by differentiating between varying levels of importance.

Generative adversarial network

GANs have shown impressive accomplishments across various fields, such as generating natural images, translating images from one form to another, and creating high-resolution images. A GAN comprises two neural networks, which are a generator (G) and a discriminator (D). The generator produces synthetic data by utilizing random Gaussian noise (z) in an effort to fool the discriminator. Conversely, the discriminator’s role is to differentiate between the produced samples and the actual data [113–115]. The process is outlined in the following formulas:

graphic file with name DmEquation33.gif (32)

where “Inline graphic” denotes training samples drawn from the actual distribution, while “Inline graphic” signifies the samples produced by the generator network. The distribution “Inline graphic” is implicitly characterized by the equation Inline graphic= G(z), where z is sampled from P(z), with P(z) representing the distribution of latent variables.

Multi-head self-attention mechanism

Sequence or word vectors are plotted to query (Q), key (K), and value (V) vectors in the attention mechanism. The relevance between Q and K is then determined by computing attention scores, usually using techniques like additive, dot product, scaled dot product, general, and concatenation. This approach enhances the model’s data representation and retrieval capabilities by enabling it to modify its attention based on various segments of the input sequence. By using the self-attention process as an example, word vectors multiplied by three different weight matrices (wq, wk, and wv) can yield Q, K, and V. The computation of Q and K’s similarity in scaled dot-product attention for a given word vector is:

graphic file with name DmEquation34.gif (33)

where dk is the input vector’s dimension, and Inline graphic indicates the correlation between various input word vectors. Researchers worked to improve and optimize the attention mechanism as its efficiency in processing sequence data became apparent. As a result, multi-head attention was introduced [116–118]. The attention operation is not applied simply once with multi-head attention. However, this process is carried out concurrently across several distinct representation spaces, or heads. Each “head” has a distinct weight, suggesting that they may concentrate on various aspects of the input data. Each head’s output is calculated in parallel before being combined into a single output.

graphic file with name DmEquation35.gif (34)
graphic file with name DmEquation36.gif (35)

Inline graphic is a concatenated weight matrix.

Large language models

BERT

Bidirectional encoder representations from transformers (BERT) is an efficient language representation model that was recently developed by Google Research and expands upon the transformer model. The main goal was to encode the input sequence as a phrase that continued to move up the stack of encoder layers. Three different forms of embeddings are used to improve each word’s input vector, such as segment, token, and position embeddings. By looking up the word in a predetermined vocabulary, token embedding is obtained. By differentiating between the first and second halves of the sentence, segment embeddings show which part of the sentence the word belongs to. Important details about the relative and absolute locations of each word in the phrase are captured by the positional embedding [119, 120]. The position vector is described as follows:

graphic file with name DmEquation37.gif (36)
graphic file with name DmEquation38.gif (37)

where pos is the position, “I” is the component point of the vector, and Inline graphic embodies the dimension of the vector.

ERNIE

A very effective large-scale RNA pretrained language model is ERNIE-RNA. The improved BERT structure serves as its main foundation. In particular, it uses a multi-head attention technique for encoding and uses the same transformer encoder as BERT. Nevertheless, this approach incorporates prior information about related RNA sequences. It enables a more thorough extraction of RNA properties by introducing a base pair attention bias during the attention calculation, which assigns distinct bias values to various base pairs. The synergistic combination of contextualized embeddings and implicit structural learning derived from large-scale pretraining, rather than just the capture of long-range dependencies alone, is the main reason why ERNIE-RNA and DNABERT perform better for ac4C prediction than conventional k-mer approaches. Self-attention mechanisms are particularly good at simulating long-range sequence interactions. In fact, complicated higher-order patterns like GC-content biases, nucleosome placement, and the tendency of particular sequences to form particular secondary structures are implicitly encoded by pretraining. As a result, when the models are optimized for ac4C detection, they see more than just a local motif. Thus, they use this pretrained knowledge to comprehend the larger structural context and chemical environment (such as stability in high-GC regions) that affects ac4C modification, an understanding that static k-mer frequencies are unable to offer. Their remarkable performance in a variety of downstream activities related to RNA structure and functionality is demonstrated by experimental evaluations [121–124]. These findings highlight the model’s extraordinary effectiveness and efficiency in capturing information related to RNA features.

Prediction evaluation

Evaluating the prediction efficiency of a prediction model is crucial. To assess the prediction ability of ML-based models, three types of techniques are currently commonly used: the independent test, the jackknife test, and the k-fold CV test [125].

Evaluation metrics

Matthew’s correlation coefficients (MCC), accuracy (Acc), specificity (Sp), and sensitivity (Sn) were utilized to calculate the model’s effectiveness [126–129], defined as follows:

graphic file with name DmEquation39.gif (38)

where “tn” denotes the correctly predicted non-ac4C sites, “fp” proposes non-ac4C sites categorized as ac4C sites, “tp” represents the real ac4C sites that were successfully predicted as ac4C sites, and “fn” shows the ac4C sites that were classified as non-ac4C sites. Additionally, the model’s efficacy and viability were illustrated using the receiver operating curve (ROC). To evaluate the prediction model’s effectiveness, the area under the curve (AUC) was calculated. AUC = 0.5 for random execution and AUC = 1 for the optimal prediction model [130–132].

Published results and discussions

Numerous computational models have been established by the researchers to predict ac4C modification sites in human mRNA. Table 3 and Fig. 5 present the full comparative results based on the training dataset.

Table 3.

Comparison of the published ac4C prediction methods on the training datasets.

Method Dataset Evaluation ACC SN SP MCC ROC PRC
PACES [33] Un-balanced 5-fold CV – – – – 0.885 0.559
XG-ac4C [34] Un-balanced 5-fold CV 0.921 0.597 0.956 0.552 0.910 0.653
DeepAc4C [35] Balanced 10-fold CV 0.824 – – – – –
CNNLSTMac4CPred [36] Un-balanced 5-fold CV 0.902 0.647 0.930 0.514 0.900 –
DLC-ac4C [37] Balanced 10-fold CV 0.800 0.831 0.772 0.606 0.877 –
Stacking-ac4C [38] Balanced 10-fold CV 0.888 0.892 0.884 0.774 0.954 0.953
TransC-ac4C [39] Balanced 5-fold CV 0.814 0.788 0.838 0.630 0.878 –
iRNA-ac4C [40] Balanced 10-fold CV 0.800 0.770 0.830 0.601 0.875 –
LSA-ac4C [41] Balanced 10-fold CV 0.820 0.855 0.785 0.643 0.879 –
EMDL-ac4C [42] Balanced 10-fold CV 0.892 – – – – –
EMDL-ac4C [42] Un-balanced 5-fold CV – – – – 0.904 0.615
MetaAc4C [43] Un-balanced 10-fold CV 0.937 0.912 0.974 0.875 0.986 –
GANSamples-ac4C [44] Balanced 5-fold CV 0.834 0.884 0.784 0.671 0.899 –
ac4C-AFL [45] Balanced 10-fold CV 0.833 0.857 0.810 0.668 0.903 –
STM-ac4C [46] Balanced 10-fold CV 0.825 0.864 0.786 0.654 0.885 –
DPNN-ac4C [47] Balanced 10-fold CV 0.838 0.821 0.856 0.665 0.927 –
TransAC4C415 nt [48] Balanced 10-fold CV 0.785 – – – 0.869 –
TransAC4C21 nt [48] Balanced 10-fold CV 0.923 – – – 0.977 –
Voting-ac4C [49] Balanced 10-fold CV 0.867 0.892 0.843 0.736 0.930 –
NBCR-ac4C [50] Balanced 10-fold CV 0.837 0.914 0.760 0.682 – –

Figure 5.

Bar chart comparing the performance of ac4C prediction tools on balanced and imbalanced training datasets.

Performance of the available ac4C modification sites prediction tools on balanced and un-balanced training datasets.

Zhao et al. [33] established the first random forest-based predictor in 2019, called PACES, in which they used six types of feature encoding schemes: one-hot, position-specific nucleotide sequence profile, position-specific di-nucleotide sequence profile, k-spaced nucleotide pair frequencies, k-nucleotide frequencies, and pseudo k-tuple nucleotide composition, and inputted into a random forest-based classifier with 5-fold CV and achieved a reasonable accuracy with an ROC value of 0.885. The performance on the unbalanced independent dataset was also reasonable, with an ROC value of 0.874. In 2020, Alam et al. [34] developed an extreme gradient boosting-based classifier to predict ac4C called XG-ac4C. In their technique, they used four types of encoding schemes, one-hot, EIIP + PseEIIP, k-mer, nucleotide chemical property, and nucleotide density, and inserted them into five ML-based classifiers, namely, random forest, Gaussian naïve bayes, logistic regression, adaboost, and extreme gradient boost by using 5-fold CV, and finally selected the best performing model to classify ac4C from non-ac4C. The accuracy of the best classifier was 0.921. The performance on unbalanced independent testing data was acceptable with an ROC value of 0.889. In 2021, Wang et al. [35] developed a CNN based method to predict ac4C. In their method, they used six physicochemical feature descriptors, namely, k-mer, CKSNAP, SCPseDNC, series correlation pseudo tri-nucleotide composition, SCPseTNC, PseEIIP of tri-nucleotide with word2vector embedding. At first, six physicochemical features were improved with a support vector machine and F-score with a sequential forward search strategy. Then these improved features were fed into the CNN by utilizing a 10-fold CV for training the model, and attained an accuracy of 0.824. The efficiency on an unbalanced independent dataset was good, with an accuracy score of 0.791. In 2022, Zhang et al. [36] introduced a CNNLSTMac4CPred method for the prediction of ac4C modification sites. In their method, they utilized three types of features. The two feature descriptors are traditional, namely, PseTNC and KNF, and the other one is a semantic feature. They utilized CNN and LSTM-based neural networks to extract semantic features from sequences. After this, these features were inserted into an extreme gradient boosting-based classifier by using 5-fold CV to train the model, and got an accuracy of 0.902. The accuracy score of CNNLSTMac4CPred on the unbalanced independent dataset was 0.872 with an ROC value of 0.882.

In 2023, Jia et al. [37] developed a DLC-ac4C method in which they utilized three types of feature encodings, namely NCP, ND, and C2 encoding. They used 1D-CNN and BiLSTM for learning the local and global features, respectively. They also utilized a channel attention mechanism to find the optimal sequence characteristics and a homo-morphemic integration technique to limit the generalization error of the model with an accuracy of 0.800. The accuracy score of DLC-ac4C on the balanced independent dataset was 0.829 with an ROC value of 0.903. Lou et al. [38] constructed a stacking-ac4C method to classify N4-acetylcytidine in human mRNA. In their method, they employed three types of encoders, namely, k-mer, PseKNC, and PseEIIP, and their hybrid fusion features, and then developed a stacking model on training data by utilizing 10-fold CV with an accuracy of 0.880. The accuracy score of stacking-ac4c on the balanced independent testing set was 0.808 with an ROC value of 0.883. Liu et al. [39] presented a new model, TransC-ac4C, to predict ac4C modification sites in mRNA. In their model, they utilized five types of feature descriptors, namely, one-hot encoding, ND, k-mer, NCP, and EIIP, and inserted them into a hybrid model that consists of CNN and transformer with 5-fold CV to train the model. Subsequently, they got an 0.814 accuracy score, and the performance on the balanced independent dataset was 0.806. Su et al. [40] developed an iRNA-ac4C to classify ac4C from non-ac4C in human mRNA. In their methodology, they utilized hybrid features of k-mer, accumulated nucleotide frequency, and nucleotide chemical frequency, and then improved these features with mRMR and IFS technique. Finally, they input these optimal features into seven different traditional ML classifiers, namely, LR, NB, GBDT, KNN, SVM, RF, and AB, and select the best performing model GBDT with 10-fold CV to classify the ac4C with an accuracy of 0.800. The performance accuracy of iRNA-ac4C on a balanced independent dataset was 0.798 with an ROC value of 0.880. Lai et al. [41] introduced a new tool, LSA-ac4C, for the accurate prediction of ac4C modification sites in human mRNA. In this technique, they used a hybrid neural network that combines an LSTM double-layer neural network with self-attention, with 10-fold CV, and achieved an accuracy of 0.820. The performance on a balanced independent dataset was 0.827 with an ROC value of 0.895. Jia et al. [42] developed a new model, EMDL-ac4C. In this technique, they used one-hot encoding and then inserted it into an ensemble model with a two-branch residual connection dense Net and attention to predict ac4C modification sites in human mRNA. In their methodology, they utilized an unbalanced dataset with 10-fold CV and a balanced dataset with 5-fold CV to train their models. Finally, they got the best model on the balanced dataset with an accuracy of 0.892. The performance of EMDL-ac4C on the balanced independent dataset was 0.808 with an ROC value of 0.879, and the ROC value on the unbalanced independent dataset was 0.901.

In 2024, Li et al. [43] introduced a MetaAc4C model in which they utilized pre-trained BERT, and the model was based on a BLSTM network. They achieved a 0.937 accuracy score with an ROC value of 0.986. The performance on unbalanced and balanced independent datasets was 0.828 and 0.817, respectively. Li et al. [44] also introduced a new method based on generative adversarial networks and transfer learning to recognize ac4C modification sites in human mRNA, named GANSamples-ac4C. They achieved an accuracy of 0.834 on the training dataset with 5-fold CV. Pham et al. [45] constructed the ac4C-AFL method, which is based on adaptive feature representation learning. In their method, they utilized sixteen types of feature encoding schemes CKSNAP, ENAC, position specific of two nucleotides, PseEIIP, Z-curve, k-mer, reverse RCKmer, di-nucleotide physicochemical properties type 1-2, NCP, binary features, multivariate mutual information and accumulated nucleotide frequency, W2V, sequence2vector, DNABERT and combination of skip di-nucleotide composition and local position-specific di-nucleotide frequency and then inputted these optimized features into 11 different types of ML and DL classifiers, namely, SVM, RF, ERT, ANN, LR, GBT, XGBT, light GBT, cat boost, AB, and CNN classifiers and generated 176 baseline models. Finally, they selected the best baseline models using a 2-step feature selection technique, whose predicting scores were integrated and trained with SVM to develop the final model by utilizing 10-fold CV to classify ac4C from non-ac4C in human mRNA, and achieved an accuracy score of 0.833. The accuracy score of ac4C-AFL on the balanced independent dataset was 0.823 with an ROC value of 0.895. Yi et al. [46] developed a new method, named STM-ac4C, which is a hybrid model based on selective kernel convolutions, a temporal convolutional network with multi-head self-attention to predict ac4C modification sites in mRNA. They attained an accuracy of 0.825 on the training dataset by utilizing 10-fold CV and an accuracy of 0.847 on the balanced independent dataset. Yuan et al. [47] developed a new method, DPNN-ac4C, based on a dual path neural network with self-attention. Their method integrates a convolutional neural network, embedding modules, a BGRU with self-attention to extract local and global features from mRNA sequences, and achieved an accuracy of 0.838. The prediction efficiency of DPNN-ac4C on the balanced independent dataset was 0.827 with an ROC value of 0.910. Liu et al. [48] developed a transformer-based model called TransAC4C to predict ac4C modification sites in mRNA. Their model was divided into four parts such as a transformer layer, one BLSTM layer, four 1D convolutional layers, with three fully connected layers. Transformer and BLSTM layers with 10-fold CV were used to enable the model to learn the contextual information, convolutional layers were used to enable the model to extract important features from the input, and fully connected layers were used to connect the input features into the output. They trained two models based on nucleotide length (415 nt and 21 nt) and achieved an accuracy score of 0.785 and 0.923, respectively. The prediction efficiencies of the above two models with nucleotide length (415 and 21) on the balanced independent dataset were 0.779 and 0.838. Jia et al. [49] constructed a tool called Voting-ac4C to predict ac4C modification sites in mRNA. In their method, they utilized RNAErnie, a transformer-based pre-trained model with six traditional feature encoding schemes such as one-hot, ENAC, ND, C2 encoding, KSNPF, and physicochemical properties, and then inserted the hybrid of these features into a DNN for dimensionality reduction. Subsequently, these dimensional reduction features were fed into a voting ensemble model constructed using CatBoost, XGBoost, and MLP classifiers by utilizing 10-fold CV. Finally, they attained an accuracy score of 0.867. The accuracy score on the balanced independent dataset was 0.831 with an ROC value of 0.887. He et al. [50] developed a DL method called NBCR-ac4C based on pretrained models to predict ac4C modification sites in human mRNA. They utilized DNABERT2 and nucleotide transformer to construct contextual embedding of nucleotide sequences and utilized CNN and ResNet18 by using 10-fold CV to further extract the shallow and deep knowledge from the contextual embedding. Subsequently, they got the accuracy score of 0.837 on the training dataset and an accuracy score of 0.835 on the balanced independent dataset with an ROC value of 0.895.

In 2025, Lu et al. [51] introduced an ERNIE-ac4C method to predict ac4C modification sites. In their technique, they used the pretrained model ERNIE-RNA to extract the attention map features and sequence features from nucleotide sequences. They input these fused features into a 2d CNN to predict ac4C modification sites in the human mRNA. The performance of ERNIE-ac4C on the balanced independent dataset was 0.902 with an ROC value of 0.961. Yao et al. [52] proposed a DL-based method called Caps-ac4C to predict ac4C modification sites in human mRNA. In their method, they utilized CGR encoding to transform the nucleotide sequences into visual representations and then employed a capsule network design to extract the local and global features from these visual representations of nucleotide sequences. The performance accuracies of Caps-ac4C on balanced and unbalanced independent datasets were 0.954 and 0.908, with ROC values of 0.996 and 0.962, respectively. The performance comparison of published results on balanced and unbalanced independent datasets is also shown in Table 4. Performance of the available tools on balanced and unbalanced independent data is shown in Fig. 6.

Table 4.

Comparison of the published results for ac4C prediction on independent datasets.

Method Dataset Evaluation ACC SN SP MCC ROC PRC
PACES [33] Un-balanced Independent set – – – – 0.874 0.485
XG-ac4C [34] Un-balanced Independent set – – – – 0.889 0.581
DeepAc4C [35] Un-balanced Independent set 0.791 0.828 0.755 0.585 0.864 –
CNNLSTMac4CPred [36] Un-balanced Independent set 0.872 0.629 0.899 0.435 0.882 –
DLC-ac4C [37] Balanced Independent set 0.829 0.862 0.797 0.660 0.903 –
Stacking-ac4C [38] Balanced Independent set 0.808 0.808 0.808 0.615 0.883 –
TransC-ac4C [39] Balanced Independent set 0.806 0.809 0.804 0.614 0.869 –
iRNA-ac4C [40] Balanced Independent set 0.798 0.767 0.829 0.597 0.880 –
LSA-ac4C [41] Balanced Independent set 0.827 0.871 0.782 0.656 0.895 –
EMDL-ac4C [42] Balanced Independent set 0.808 0.810 0.817 0.616 0.879 0.864
EMDL-ac4C [42] Un-balanced Independent set – – – – 0.901 0.594
MetaAc4C [43] Un-balanced Independent set 0.828 0.809 0.848 0.657 0.895 –
MetaAc4C [43] Balanced Independent set 0.817 0.792 0.843 0.636 0.874 –
GANSamples-ac4C [44] Balanced Independent set – – – – – –
ac4C-AFL [45] Balanced Independent set 0.823 0.844 0.803 0.647 0.895 –
STM-ac4C [46] Balanced Independent set 0.847 0.858 0.837 0.695 0.907 –
DPNN-ac4C [47] Balanced Independent set 0.827 0.817 0.847 0.657 0.910 –
TransAC4C415 nt [48] Balanced Independent set 0.779 0.794 0.765 0.559 0.853 –
TransAC4C21 nt [48] Balanced Independent set 0.838 0.838 0.838 0.677 0.903 –
Voting-ac4C [49] Balanced Independent set 0.831 0.851 0.811 0.663 0.887 –
NBCR-ac4C [50] Balanced Independent set 0.835 0.849 0.820 0.670 0.895 –
ERNIE-ac4C [51] Balanced Independent set 0.902 0.897 0.908 0.806 0.961 0.965
Caps-ac4C [52] Un-balanced Independent set 0.908 0.889 0.912 0.721 0.962 –
Caps-ac4C [52] Balanced Independent set 0.954 0.920 0.989 0.912 0.996 0.995

Figure 6.

Bar chart comparing the performance of ac4C prediction tools on balanced and imbalanced independent datasets.

Performance of the available ac4C modification sites prediction tools on balanced and unbalanced independent datasets.

Independent evaluation

For the sake of fair and authentic evaluation, it is very important to check the published tool’s performance on random datasets. For this purpose, we selected those web-servers that were running online and found three web-servers, namely, ac4C-AFL [45], Caps-ac4C [52], and PACES [33]. We have tested these three state-of-the-art tools on three different datasets (Su et al. [40], Wang et al. [35], and Arango et al. [14]) to check the performance and found that ac4C-AFL [45] performed well on these datasets as compared to the Caps-ac4C [52] and PACES [33]. The accuracy score of ac4C-AFL [45] on the Su et al. [40] dataset was 0.886 and 0.848 on the Wang et al. dataset and 0.753 on the Arango et al. dataset. The accuracy score of Caps-ac4C on the Su et al. dataset was 0.845, 0.825 on the Wang et al. dataset, and 0.662 on Arango et al [14]. The accuracy score of PACES [33] on the Su et al. [40] dataset was 0.776, 0.728 on the Wang et al. dataset, and 0.611 on the Arango et al. [14] dataset. Finally, ac4C-AFL [45] performed well across independent datasets and outperformed the other web-based tools by 4.1%–11% on the Su et al. dataset, 2.3%–12% on the Wang et al. [35] dataset, and 9.1%–14.2% on the Arango et al. [14] dataset. The model’s performance on the Arango et al. dataset was not up to the mark due to data leakage and redundant sequences, compared with the Wang et al. and Su et al. datasets. The Arango et al. dataset needs to be polished by data-cleaning tools for removing redundant sequences. The performance results of available web-based tools on three different datasets are also shown in Table 5 and Fig. 7.

Table 5.

Performance of available active web servers on three different independent datasets.

Method Independent dataset ACC SN SP MCC
PACES [33] Su et al. 0.776 0.763 0.781 0.623
Arango et al. 0.611 0.608 0.617 0.475
Wang et al. 0.728 0.714 0.733 0.588
ac4C-AFL [45] Su et al. 0.886 0.876 0.881 0.735
Arango et al. 0.753 0.743 0.748 0.587
Wang et al. 0.848 0.851 0.839 0.675
Caps-ac4C [52] Su et al. 0.845 0.835 0.837 0.697
Arango et al. 0.662 0.655 0.674 0.516
Wang et al. 0.825 0.831 0.827 0.662

Figure 7.

Spider charts showing the performance of web-based ac4C prediction tools on three independent datasets from different studies.

Performance of the available web-based tools on the independent datasets. Performance on Su et al. data (A). Performance on Wang et al. data (B). Performance on Arango et al. data (C).

Conclusion

NAT10 catalyzes ac4C, which is among the most significant posttranscriptional modifications in RNA. It adds an acetyl group to the nitrogen at the fourth position of the cytidine base and plays an essential role in the stability of mRNA, posttranscriptional regulation, translational efficiency, and human immune function regulation [133]. Thus, it is crucial to identify ac4C computationally. It can quickly and thoroughly identify additional details and undiscovered functions of ac4C when compared to the costly experimental methods. Many classifiers based on DL and ML are currently being developed. Numerous feature extraction techniques, such as KSNPF, CKSNAP, PseEIIP, PseKNC, DNABERT2 embedding, ANF, SCPseDNC, SCPseTNC, PseKNC, physiochemical properties, and some feature selection techniques, such as BD, PCA, and ANOVA with IFS, were employed in all of these predictors to eliminate redundant features and obtain optimal features. Artificial intelligence-based classification is important for the prediction of ac4C in human RNA. Therefore, we compared and assessed the available machine and deep learning-based classifiers and found that the MetaAc4C [43] method has an outstanding performance and generalization ability on training datasets, and the Caps-ac4C [52] method showed brilliant performance on independent datasets. We have also assessed the performance of the available web-based tools on three different datasets and found that ac4C-AFL [45] outpaced the other tools in terms of accuracy score. Although the ac4C classifiers’ results are reasonable, there is still room for improvement. Clear features that can reflect the fundamental characteristics of ac4C against non-ac4C are required. Several challenges limited the research to some extent; these limitations can be characterized into two groups: general limitations and specific limitations.

The specific limitations: first, the sample size of the available data for predicting ac4C modification sites is relatively small. This drawback prevented the application of certain methods. Therefore, we suggest a reasonably larger dataset in future studies. Machine and deep learning-based studies are used in this review, which also comes with various limitations, such as dataset size, overfitting, underfitting, and dealing with dataset outliers. Due to small datasets, the performance of traditional machine learning models, such as SVM, RF, AB, and NB, was reasonable, but DL models faced issues regarding overfitting and underfitting. As a result, in the future, the next study will include more data and other clustering methods proven to be much more efficient.

The general limitations: there are also some limitations to the feature extraction method. The KSNPF, CKSNAP, PseEIIP, and PseKNC feature extraction techniques have problems with computational power due to their size and computational cost. There is still a need to do more by exploiting advanced feature descriptors such as evolutionary scale modeling ESM1-3. We anticipate that in the near future, more amazing results will be obtained with the support of LLMs as bioinformatics and computational biology improve.

Key Points

  • ac4C modification sites are very crucial and directly related to mRNA stability, transcription regulation, and regulation of the immune functions.

  • The deep learning models with meaningful features are the forward step towards accurate N4 acetylation modification sites prediction in human mRNA.

  • The development of the particular ac4C modification sites large language models is highly anticipated in the field.

  • Data-driven discovery of ac4C modification sites is emerging as the frontier of computational biology.

Contributor Information

Hasan Zulfiqar, Center for AI and Computational Biology, Institute of System Medicine, Peking Union Medical College, Chinese Academy of Medical Sciences, Suzhou 215123, China.

Ramala Masood Ahmad, Center for Informational Biology, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China; Department of Plant Breeding and Genetics, University of Agriculture Faisalabad, Faisalabad 38000, Pakistan.

Hao Lin, Center for Informational Biology, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China.

Xiao-Long Yu, School of Materials Science and Engineering, Hainan University, Haikou 570228, China.

Author’s contributions

Hasan Zulfiqar (Conceptualization, Writing—original draft, Resources, Investigation, Funding acquisition), Ramala Masood Ahmad (Formal analysis, Investigation, Visualization), Hao Lin (Supervision, Writing—review & editing, Funding acquisition), and Xiao-Long Yu (Supervision, Writing—review & editing, Funding acquisition)

Funding

This work has been supported by the grant from the National Natural Science Foundation of China (62302079, 62261017).

Data availability

There is no data.

References

  • 1. Wada  T, Kobori  A, Kawahara  SI  et al. Synthesis and properties of oligodeoxyribonucleotides containing 4-N-acetylcytosine bases. Tetrahedron Lett  1998;39:6907–10. 10.1016/S0040-4039(98)01449-X [DOI] [Google Scholar]
  • 2. Gu  Z, Zou  L, Pan  X  et al. The role and mechanism of NAT10-mediated ac4C modification in tumor development and progression. MedComm  2024;5:e70026. 10.1002/mco2.70026 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Wang  Q, Yuan  Y, Zhou  Q  et al. RNA N4-acetylcytidine modification and its role in health and diseases. MedComm  2025;6:e70015. 10.1002/mco2.70015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Jiao  L, Si  Y, Yuan  Y  et al. Emerging role of N-acetyltransferase 10 in diseases: RNA ac4C modification and beyond. Mol Biomed  2025;6:46. 10.1186/s43556-025-00286-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Tao  L, Lu  Y, Chen  Z  et al. RNA ac4C modification in cancer biology: from regulatory mechanisms to clinical applications. Sci China Life Sci  2024;67:832–5. 10.1007/s11427-023-2496-5 [DOI] [PubMed] [Google Scholar]
  • 6. Xu  N, Zhuo  J, Chen  Y  et al. Downregulation of N4-acetylcytidine modification in myeloid cells attenuates immunotherapy and exacerbates hepatocellular carcinoma progression. Br J Cancer  2024;130:201–12. 10.1038/s41416-023-02510-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Liu  H, Xu  L, Yue  S  et al. Targeting N4-acetylcytidine suppresses hepatocellular carcinoma progression by repressing eEF2-mediated HMGB2 mRNA translation. Cancer Commun  2024;44:1018–41. 10.1002/cac2.12595 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Wang  Y, Wang  S, Liu  M  et al. The function of NAT10-driven N4-acetylcytidine modification in cancer: novel insights and potential therapeutic targets. Cell Biosci  2025;15:165. 10.1186/s13578-025-01504-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Chen  Z, Zhao  M, You  L  et al. Developing an artificial intelligence method for screening hepatotoxic compounds in traditional Chinese medicine and western medicine combination. Chin Med  2022;17:58. 10.1186/s13020-022-00617-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Ji  H-N, Zhou  HQ, Qie  JB  et al. Dysregulated ac4C modification of mRNA in a mouse model of early-stage Alzheimer’s disease. Cell Biosci  2025;15:45. 10.1186/s13578-025-01389-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Ma  Y, Li  W, Fan  C  et al. Comprehensive analysis of long non-coding RNAs N4-acetylcytidine in Alzheimer’s disease mice model using high-throughput sequencing. J Alzheimer’s Dis  2022;90:1659–75. 10.3233/JAD-220564 [DOI] [PubMed] [Google Scholar]
  • 12. Chen  Z, Jiang  Y, Zhang  X  et al. ResNet18DNN: prediction approach of drug-induced liver injury by deep neural network with ResNet18. Brief Bioinform  2022;23:bbab503. 10.1093/bib/bbab503 [DOI] [PubMed] [Google Scholar]
  • 13. Chen  Z, Jiang  Y, Zhang  X  et al. The prediction approach of drug-induced liver injury: response to the issues of reproducible science of artificial intelligence in real-world applications. Brief Bioinform  2022;23:bbac196. 10.1093/bib/bbac196 [DOI] [PubMed] [Google Scholar]
  • 14. Arango  D, Sturgill  D, Alhusaini  N  et al. Acetylation of cytidine in mRNA promotes translation efficiency. Cell  2018;175:1872–1886.e24. 10.1016/j.cell.2018.10.030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Schiffers  S, Oberdoerffer  S. ac4C: a fragile modification with stabilizing functions in RNA metabolism. RNA  2024;30:583–94. 10.1261/rna.079948.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Zhang  X, Zheng  Y, Yang  J  et al. Abnormal ac4C modification in metabolic dysfunction associated steatotic liver cells. Sci Rep  2025;15:1013. 10.1038/s41598-024-84564-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Luo  Y, Shi  L, Li  Y  et al. From intention to implementation: automating biomedical research via LLMs. Sci China Inf Sci  2025;68:170105. 10.1007/s11432-024-4485-0 [DOI] [Google Scholar]
  • 18. Qiao  J, Jin  J, Yu  H  et al. Towards retraining-free RNA modification prediction with incremental learning. Inf Sci  2024;660:120105. [Google Scholar]
  • 19. Meissner  A, Gnirke  A, Bell  GW  et al. Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis. Nucleic Acids Res  2005;33:5868–77. 10.1093/nar/gki901 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Nakabayashi  K, Yamamura  M, Haseagawa  K  et al. Reduced representation bisulfite sequencing (RRBS). In: Hatada I, Horii T (eds.), Methods Mole. Biol, 2577, pp. 39–51. Humana, New York, NY: Epigenomics, 2022. 10.1007/978-1-0716-2724-2_3 [DOI] [PubMed] [Google Scholar]
  • 21. Farrell  C, Thompson  M, Tosevska  A  et al. BiSulfite bolt: a bisulfite sequencing analysis platform. Gigascience  2021;10:giab033. 10.1093/gigascience/giab033 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Tse  OO, Jiang  P, Cheng  SH  et al. Genome-wide detection of cytosine methylation by single molecule real-time sequencing. Proc Natl Acad Sci  2021;118:e2019768118. 10.1073/pnas.2019768118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Cascino  P, Nevone  A, Piscitelli  M  et al. Single-molecule real-time sequencing of the M protein: toward personalized medicine in monoclonal gammopathies. Am J Hematol  2022;97:E389–92. 10.1002/ajh.26684 [DOI] [PubMed] [Google Scholar]
  • 24. Li  X, Wang  X, Ma  Q  et al. Integrated single-molecule real-time sequencing and RNA sequencing reveal the molecular mechanisms of salt tolerance in a novel synthesized polyploid genetic bridge between maize and its wild relatives. BMC Genomics  2023;24:55. 10.1186/s12864-023-09148-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Deslattes Mays  A, Schmidt  M, Graham  G  et al. Single-molecule real-time (SMRT) full-length RNA-sequencing reveals novel and distinct mRNA isoforms in human bone marrow cell subpopulations. Genes  2019;10:253. 10.3390/genes10040253 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Covey  TR, Lee  ED, Bruins  AP  et al. Liquid chromatography/mass spectrometry. Anal Chem  1986;58:1451A–61. 10.1021/ac00127a001 [DOI] [Google Scholar]
  • 27. Beiki  H, Sturgill  D, Arango  D  et al. Detection of ac4C in human mRNA is preserved upon data reassessment. Mol Cell  2024;84:1611–1625.e3. 10.1016/j.molcel.2024.03.018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Wang  S, Xie  H, Mao  F  et al. N 4-acetyldeoxycytosine DNA modification marks euchromatin regions in Arabidopsis thaliana. Genome Biol  2022;23:5. 10.1186/s13059-021-02578-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Zulfiqar  H, Huang  QL, Lv  H  et al. Deep-4mCGP: a deep learning approach to predict 4mC sites in Geobacter pickeringii by using correlation-based feature selection technique. Int J Mol Sci  2022;23:1251. 10.3390/ijms23031251 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Zulfiqar  H, Sun  ZJ, Huang  QL  et al. Deep-4mCW2V: a sequence-based predictor to identify N4-methylcytosine sites in Escherichia coli. Methods  2022;203:558–63. 10.1016/j.ymeth.2021.07.011 [DOI] [PubMed] [Google Scholar]
  • 31. Dao  F-Y, Lv  H, Yang  YH  et al. Computational identification of N6-methyladenosine sites in multiple tissues of mammals. Comput Struct Biotechnol J  2020;18:1084–91. 10.1016/j.csbj.2020.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Dai  C, Jiang  Y, Yin  C  et al. scIMC: a platform for benchmarking comparison and visualization analysis of scRNA-seq data imputation methods. Nucleic Acids Res  2022;50:4877–99. 10.1093/nar/gkac317 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Zhao  W, Zhou  Y, Cui  Q  et al. PACES: prediction of N4-acetylcytidine (ac4C) modification sites in mRNA. Sci Rep  2019;9:11112. 10.1038/s41598-019-47594-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Alam  W, Tayara  H, Chong  KT. XG-ac4C: identification of N4-acetylcytidine (ac4C) in mRNA using eXtreme gradient boosting with electron-ion interaction pseudopotentials. Sci Rep  2020;10:20942. 10.1038/s41598-020-77824-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Wang  C, Ju  Y, Zou  Q  et al. DeepAc4C: a convolutional neural network model with hybrid features composed of physicochemical patterns and distributed representation information for identification of N4-acetylcytidine in mRNA. Bioinformatics  2022;38:52–7. 10.1093/bioinformatics/btab611 [DOI] [PubMed] [Google Scholar]
  • 36. Zhang  G, Luo  W, Lyu  J  et al. CNNLSTMac4CPred: a hybrid model for N 4-Acetylcytidine prediction. Interdiscip Sci: Comput Life Sci  2022;14:439–51. 10.1007/s12539-021-00500-0 [DOI] [PubMed] [Google Scholar]
  • 37. Jia  J, Cao  X, Wei  Z. DLC-ac4C: a prediction model for N4-acetylcytidine sites in human mRNA based on DenseNet and bidirectional LSTM methods. Curr Genomics  2023;24:171–86. 10.2174/0113892029270191231013111911 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Lou  L-L, Qiu  WR, Liu  Z  et al. Stacking-ac4C: an ensemble model using mixed features for identifying n4-acetylcytidine in mRNA. Front Immunol  2023;14:1267755. 10.3389/fimmu.2023.1267755 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Liu  D, Liu  Z, Xia  Y  et al. TransC-ac4C: identification of N4-acetylcytidine (ac4C) sites in mRNA using deep learning. IEEE/ACM Trans Comput Biol Bioinform  2024;21:1403–12. 10.1109/TCBB.2024.3386972 [DOI] [PubMed] [Google Scholar]
  • 40. Su  W, Xie  XQ, Liu  XW  et al. iRNA-ac4C: a novel computational method for effectively detecting N4-acetylcytidine sites in human mRNA. Int J Biol Macromol  2023;227:1174–81. 10.1016/j.ijbiomac.2022.11.299 [DOI] [PubMed] [Google Scholar]
  • 41. Lai  F-L, Gao  F. LSA-ac4C: a hybrid neural network incorporating double-layer LSTM and self-attention mechanism for the prediction of N4-acetylcytidine sites in human mRNA. Int J Biol Macromol  2023;253:126837. 10.1016/j.ijbiomac.2023.126837 [DOI] [PubMed] [Google Scholar]
  • 42. Jia  J, Wei  Z, Cao  X. EMDL-ac4C: identifying N4-acetylcytidine based on ensemble two-branch residual connection DenseNet and attention. Front Genet  2023;14:1232038. 10.3389/fgene.2023.1232038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Li  Z, Jin  B, Fang  J. MetaAc4C: a multi-module deep learning framework for accurate prediction of N4-acetylcytidine sites based on pre-trained bidirectional encoder representation and generative adversarial networks. Genomics  2024;116:110749. 10.1016/j.ygeno.2023.110749 [DOI] [PubMed] [Google Scholar]
  • 44. Li  F, Zhang  J, Li  K  et al. GANSamples-ac4C: enhancing ac4C site prediction via generative adversarial networks and transfer learning. Anal Biochem  2024;689:115495. 10.1016/j.ab.2024.115495 [DOI] [PubMed] [Google Scholar]
  • 45. Pham  NT, Terrance  AT, Jeon  YJ  et al. ac4C-AFL: a high-precision identification of human mRNA N4-acetylcytidine sites based on adaptive feature representation learning. Mol Ther Nucleic Acids  2024;35:102192. 10.1016/j.omtn.2024.102192 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Yi  M, Zhou  F, Deng  Y. STM-ac4C: a hybrid model for identification of N4-acetylcytidine (ac4C) in human mRNA based on selective kernel convolution, temporal convolutional network, and multi-head self-attention. Front Genet  2024;15:1408688. 10.3389/fgene.2024.1408688 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Yuan  J, Wang  Z, Pan  Z  et al. DPNN-ac4C: a dual-path neural network with self-attention mechanism for identification of N4-acetylcytidine (ac4C) in mRNA. Bioinformatics  2024;40:btae625. 10.1093/bioinformatics/btae625 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Liu  R, Zhang  Y, Wang  Q  et al. TransAC4C—a novel interpretable architecture for multi-species identification of N4-acetylcytidine sites in RNA with single-base resolution. Brief Bioinform  2024;25:bbae200. 10.1093/bib/bbae200 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Jia  Y, Zhang  Z, Yan  S  et al. Voting-ac4C: pre-trained large RNA language model enhances RNA N4-acetylcytidine site prediction. Int J Biol Macromol  2024;282:136940. 10.1016/j.ijbiomac.2024.136940 [DOI] [PubMed] [Google Scholar]
  • 50. He  W, Han  Y, Zuo  Y  et al. NBCR-ac4C: a deep learning framework based on multivariate BERT for human mRNA N4-Acetylcytidine sites prediction. J Chem Inf Model  2024;64:8074–81. 10.1021/acs.jcim.4c01415 [DOI] [PubMed] [Google Scholar]
  • 51. Lu  R, Qiao  J, Li  K  et al. ERNIE-ac4C: a novel deep learning model for effectively predicting N4-acetylcytidine sites. J Mol Biol  2025;437:168978. 10.1016/j.jmb.2025.168978 [DOI] [PubMed] [Google Scholar]
  • 52. Yao  L, Xie  P, Dong  D  et al. Caps-ac4C: an effective computational framework for identifying N4-acetylcytidine sites in human mRNA based on deep learning. J Mol Biol  2025;437:168961. 10.1016/j.jmb.2025.168961 [DOI] [PubMed] [Google Scholar]
  • 53. Griffiths-Jones  S, Bateman  A, Marshall  M  et al. Rfam: an RNA family database. Nucleic Acids Res  2003;31:439–41. 10.1093/nar/gkg006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Nawrocki  EP, Burge  SW, Bateman  A  et al. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res  2015;43:D130–7. 10.1093/nar/gku1063 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Burge  SW, Daub  J, Eberhardt  R  et al. Rfam 11.0: 10 years of RNA families. Nucleic Acids Res  2013;41:D226–32. 10.1093/nar/gks1005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Xuan  J, Chen  L, Chen  Z  et al. RMBase v3. 0: decode the landscape, mechanisms and functions of RNA modifications. Nucleic Acids Res  2024;52:D273–84. 10.1093/nar/gkad1070 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Zulfiqar  H, Guo  Z, Ahmad  RM  et al. Deep-STP: a deep learning-based approach to predict snake toxin proteins by using word embeddings. Front Med  2024;10:1291352. 10.3389/fmed.2023.1291352 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Zulfiqar  H, Khan  RS, Hassan  F  et al. Computational identification of N4-methylcytosine sites in the mouse genome with machine-learning method. Math Biosci Eng  2021;18:3348–63. 10.3934/mbe.2021167 [DOI] [PubMed] [Google Scholar]
  • 59. Lv  H, Dao  FY, Zulfiqar  H  et al. A sequence-based deep learning approach to predict CTCF-mediated chromatin loop. Brief Bioinform  2021;22:bbab031. 10.1093/bib/bbab031 [DOI] [PubMed] [Google Scholar]
  • 60. Zulfiqar  H, Ahmad  RM, Raza  A  et al. Promoter prediction in agrobacterium tumefaciens strain C58 by using artificial intelligence strategies. In: Marchisio MA (ed.), Methods Mole. Biol, 2844, pp. 33–44. Humana, New York, NY: Synthetic Promoters, 2024. 10.1007/978-1-0716-4063-0_2 [DOI] [PubMed] [Google Scholar]
  • 61. Zulfiqar  H, Ahmed  Z, Kissanga Grace-Mercure  B  et al. Computational prediction of promotors in agrobacterium tumefaciens strain C58 by using the machine learning technique. Front Microbiol  2023;14:1170785. 10.3389/fmicb.2023.1170785 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Chen  Y, Wang  Z, Wang  J  et al. Self-supervised learning in drug discovery. Sci China Inf Sci  2025;68:170103. 10.1007/s11432-024-4453-4 [DOI] [Google Scholar]
  • 63. Jenike  KM, Campos-Domínguez  L, Boddé  M  et al. k-mer approaches for biodiversity genomics. Genome Res  2025;35:219–30. 10.1101/gr.279452.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Roberts  MD, Davis  O, Josephs  EB  et al. k-mer-based approaches to bridging pangenomics and population genetics. Mol Biol Evol  2025;42:msaf047. 10.1093/molbev/msaf047 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Katayama  Y, Kobayashi  TJ. Comparative study of repertoire classification methods reveals data efficiency of k-mer feature extraction. Front Immunol  2022;13:797640. 10.3389/fimmu.2022.797640 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Li  M, Fan  Y, Zhang  Y  et al. Using sequence similarity based on CKSNP features and a graph neural network model to identify miRNA–disease associations. Genes  2022;13:1759. 10.3390/genes13101759 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Basith  S, Hasan  MM, Lee  G  et al. Integrative machine learning framework for the identification of cell-specific enhancers from the human genome. Brief Bioinform  2021;22:bbab252. 10.1093/bib/bbab252 [DOI] [PubMed] [Google Scholar]
  • 68. Zhang  ZY, Fan  YE, Huang  CB  et al. Human essential gene identification based on feature fusion and feature screening. IET Syst Biol  2024;18:227–37. 10.1049/syb2.12105 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Chen  W, Yang  H, Feng  P  et al. iDNA4mC: identifying DNA N4-methylcytosine sites based on nucleotide chemical properties. Bioinformatics  2017;33:3518–23. 10.1093/bioinformatics/btx479 [DOI] [PubMed] [Google Scholar]
  • 70. Chen  Z, Zhao  P, Li  F  et al. iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data. Brief Bioinform  2020;21:1047–57. 10.1093/bib/bbz041 [DOI] [PubMed] [Google Scholar]
  • 71. Lv  Z, Ding  H, Wang  L  et al. A convolutional neural network using dinucleotide one-hot encoder for identifying DNA N6-methyladenine sites in the rice genome. Neurocomputing  2021;422:214–21. 10.1016/j.neucom.2020.09.056 [DOI] [Google Scholar]
  • 72. Zheng  K, You  ZH, Li  JQ  et al. iCDA-CGR: identification of circRNA-disease associations based on chaos game representation. PLoS Comput Biol  2020;16:e1007872. 10.1371/journal.pcbi.1007872 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Gupta  YM, Kirana  SN, Homchan  S. Representing DNA for machine learning algorithms: a primer on one-hot, binary, and integer encodings. Biochem Mol Biol Educ  2025;53:142–6. 10.1002/bmb.21870 [DOI] [PubMed] [Google Scholar]
  • 74. Zhang  P, Zhang  H, Wu  H. iPro-WAEL: a comprehensive and robust framework for identifying promoters in multiple species. Nucleic Acids Res  2022;50:10278–89. 10.1093/nar/gkac824 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Dou  L, Li  X, Ding  H  et al. Prediction of m5C modifications in RNA sequences by combining multiple sequence features. Mol Ther Nucleic Acids  2020;21:332–42. 10.1016/j.omtn.2020.06.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76. Ji  Y, Zhou  Z, Liu  H  et al. DNABERT: pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. Bioinformatics  2021;37:2112–20. 10.1093/bioinformatics/btab083 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. Makhdoomi  F, Ghosh  N. TinyDNABERT: an efficient and optimized BERT model for DNA-language in genome. In: 21st IEEE India Council International Conference (INDICON), 2024, pp. 1–6. Kharagpur, India, 2024. 10.1109/INDICON63790.2024.10958313 [DOI] [Google Scholar]
  • 78. Sanabria  M, Hirsch  J, Poetsch  AR. Distinguishing word identity and sequence context in DNA language models. BMC Bioinformatics  2024;25:301. 10.1186/s12859-024-05869-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79. Kurata  H, Tsukiyama  S, Manavalan  B. iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model. Brief Bioinform  2022;23:bbac265. [DOI] [PubMed] [Google Scholar]
  • 80. Venkatesh  B, Anuradha  J. A review of feature selection and its methods. Cybern Inf Technol  2019;19:3–26. [Google Scholar]
  • 81. Liu  H, Setiono  R. Feature selection and classification–a probabilistic wrapper approach. In: Tanaka T, Ohsuga S, Ali M (eds.), Proceedings of the 9th International Conference on Industrial and Engineering Applications of Artificial Intelligence and Expert Systems, 1996, pp. 419–24. Fukuoka, Japan, 1996. https://dl.acm.org/doi/proceedings/10.5555/3104635 [Google Scholar]
  • 82. Chen  Z, Pang  M, Zhao  Z  et al. Feature selection may improve deep neural networks for the bioinformatics problems. Bioinformatics  2020;36:1542–52. 10.1093/bioinformatics/btz763 [DOI] [PubMed] [Google Scholar]
  • 83. Wang  Y, Gao  X, Ru  X  et al. A hybrid feature selection algorithm and its application in bioinformatics. PeerJ Comput Sci  2022;8:e933. 10.7717/peerj-cs.933 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84. Dao  F-Y, Lv  H, Zhang  ZY  et al. BDselect: a package for k-mer selection based on the binomial distribution. Curr Bioinform  2022;17:238–44. 10.2174/1574893616666211007102747 [DOI] [Google Scholar]
  • 85. Sparta  B, Hamilton  T, Natesan  G  et al. Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing. PLoS Comput Biol  2024;20:e1012386. 10.1371/journal.pcbi.1012386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Greenacre  M, Groenen  PJF, Hastie  T  et al. Principal component analysis. Nat Rev Methods Primers  2022;2:100. 10.1038/s43586-022-00184-w [DOI] [Google Scholar]
  • 87. Bielińska-Wąż  D, Wąż  P, Błaczkowska  A  et al. Mathematical modeling in bioinformatics: application of an alignment-free method combined with principal component analysis. Symmetry  2024;16:967. 10.3390/sym16080967 [DOI] [Google Scholar]
  • 88. Privé  F, Luu  K, Blum  MGB  et al. Efficient toolkit implementing best practices for principal component analysis of population genetic data. Bioinformatics  2020;36:4449–57. 10.1093/bioinformatics/btaa520 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89. Hassall  KL, Mead  A. Beyond the one-way ANOVA for ’omics data. BMC Bioinformatics  2018;19:199. 10.1186/s12859-018-2173-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90. Ye  Y, Zhang  R, Zheng  W  et al. RIFS: a randomly restarted incremental feature selection algorithm. Sci Rep  2017;7:13013. 10.1038/s41598-017-13259-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91. Pisner  DA, Schnyer  DM. Support vector machine. In: Mechelli A, Vieira S (eds.), Elsevier, 2020, pp. 101–21. Amsterdam, Netherlands: Machine Learning, 2020. 10.1016/B978-0-12-815739-8.00006-7 [DOI] [Google Scholar]
  • 92. Cervantes  J, Garcia-Lamont  F, Rodríguez-Mazahua  L  et al. A comprehensive survey on support vector machine classification: applications, challenges and trends. Neurocomputing  2020;408:189–215. 10.1016/j.neucom.2019.10.118 [DOI] [Google Scholar]
  • 93. Yan  H, Long  Y, Lv  C  et al. Research on bioinformatics data classification method based on support vector machine. Int J Data Min Bioinf  2025;29:21–35. 10.1504/IJDMB.2025.142975 [DOI] [Google Scholar]
  • 94. Li  H, Pang  Y, Liu  B. BioSeq-BLM: a platform for analyzing DNA, RNA, and protein sequences based on biological language models. Nucleic Acids Res  2021;49:e129. 10.1093/nar/gkab829 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95. Lu  H, Gao  H, Ye  M  et al. A hybrid ensemble algorithm combining AdaBoost and genetic algorithm for cancer classification with gene expression data. IEEE/ACM Trans Comput Biol Bioinform  2019;18:863–70. 10.1109/TCBB.2019.2952102 [DOI] [PubMed] [Google Scholar]
  • 96. Geurts  P, Ernst  D, Wehenkel  L. Extremely randomized trees. Mach Learn  2006;63:3–42. 10.1007/s10994-006-6226-1 [DOI] [Google Scholar]
  • 97. Kocev  D, Ceci  M, Stepišnik  T. Ensembles of extremely randomized predictive clustering trees for predicting structured outputs. Mach Learn  2020;109:2213–41. 10.1007/s10994-020-05894-4 [DOI] [Google Scholar]
  • 98. Salman  HA, Kalakech  A, Steiti  A. Random forest algorithm overview. Babylon J Mach Learn  2024;2024:69–79. 10.58496/BJML/2024/007 [DOI] [Google Scholar]
  • 99. Ghosh  D, Cabrera  J. Enriched random forest for high dimensional genomic data. IEEE/ACM Trans Comput Biol Bioinform  2021;19:2817–28. 10.1109/TCBB.2021.3089417 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100. Lavanya  C, Pooja  S, Abhay  HK. Novel biomarker prediction for lung cancer using random forest classifiers. Cancer Informat  2023;22:11769351231167992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101. Cinelli  M, Sun  Y, Best  K  et al. Feature selection using a one dimensional naïve Bayes’ classifier increases the accuracy of support vector machine classification of CDR3 repertoires. Bioinformatics  2017;33:951–5. 10.1093/bioinformatics/btw771 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102. Deng  L, Sui  Y, Zhang  J. XGBPRH: prediction of binding hot spots at protein–RNA interfaces utilizing extreme gradient boosting. Genes  2019;10:242. 10.3390/genes10030242 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103. Pereira  GRC, da Conceição  LMA, Abrahim-Vieira  BA  et al. XGBMUT: predicting the functional impact of missense mutations using an extreme gradient boost classifier. ACS Omega  2025;10:8349–60. 10.1021/acsomega.4c10179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104. Li  Q, He  Y, Pan  J. CrossFuse-XGBoost: accurate prediction of the maximum recommended daily dose through multi-feature fusion, cross-validation screening and extreme gradient boosting. Brief Bioinform  2024;25:bbad511. 10.1093/bib/bbad511 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105. Zhang  D, Chen  HD, Zulfiqar  H  et al. iBLP: an XGBoost-based predictor for identifying bioluminescent proteins. Comput Math Methods Med  2021;2021:1–15. 10.1155/2021/6664362 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106. Zulfiqar  H, Yuan  SS, Huang  QL  et al. Identification of cyclin protein using gradient boost decision tree algorithm. Comput Struct Biotechnol J  2021;19:4123–31. 10.1016/j.csbj.2021.07.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107. Zhou  L, Wang  Z, Tian  X  et al. LPI-deepGBDT: a multiple-layer deep framework based on gradient boosting decision trees for lncRNA–protein interaction identification. BMC Bioinformatics  2021;22:479. 10.1186/s12859-021-04399-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108. Settelmeier  J, Goetze  S, Boshart  J  et al. MultiOmicsAgent: guided extreme gradient-boosted decision trees-based approaches for biomarker-candidate discovery in multiomics data. J Proteome Res  2025;24:2816–31. 10.1021/acs.jproteome.4c01066 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109. Chen  D, Jacob  L, Mairal  J. Biological sequence modeling with convolutional kernel networks. Bioinformatics  2019;35:3294–302. 10.1093/bioinformatics/btz094 [DOI] [PubMed] [Google Scholar]
  • 110. Graves  A, Graves  A. Long short-term memory. Supervised sequence labelling with recurrent neural networks  2012;385:37–45. 10.1007/978-3-642-24797-2_4 [DOI] [Google Scholar]
  • 111. Lamurias  A, Sousa  D, Clarke  LA  et al. BO-LSTM: classifying relations via long short-term memory networks along biomedical ontologies. BMC Bioinformatics  2019;20:10. 10.1186/s12859-018-2584-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112. Van Houdt  G, Mosquera  C, Nápoles  G. A review on the long short-term memory model. Artif Intell Rev  2020;53:5929–55. 10.1007/s10462-020-09838-1 [DOI] [Google Scholar]
  • 113. Goodfellow  I, Pouget-Abadie  J, Mirza  M  et al. Generative adversarial networks. Commun ACM  2020;63:139–44. 10.1145/3422622 [DOI] [Google Scholar]
  • 114. Gui  J, Sun  Z, Wen  Y  et al. A review on generative adversarial networks: algorithms, theory, and applications. IEEE Trans Knowl Data Eng  2021;35:3313–32. 10.1109/TKDE.2021.3130191 [DOI] [Google Scholar]
  • 115. Gonog  L, Zhou  Y. A review: Generative adversarial networks. In: 14th IEEE conference on industrial electronics and applications (ICIEA), vol 2019, pp. 505–10. Xi'an, China, 2019. https://ieeexplore.ieee.org/document/8833686 [Google Scholar]
  • 116. Vaswani  A, Shazeer  N, Parmar  N  et al. Attention is all you need. Adv Neural Inf Proces Syst  2017;30:6000–10. https://dl.acm.org/doi/10.5555/3295222.3295349 [Google Scholar]
  • 117. Harrer  S. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine  2023;90:104512. 10.1016/j.ebiom.2023.104512 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118. Zhang  Y, Liu  C, Liu  M  et al. Attention is all you need: utilizing attention in AI-enabled drug discovery. Brief Bioinform  2023;25:bbad467. 10.1093/bib/bbad467 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119. Koroteev  MV. BERT: a review of applications in natural language processing and understanding. arXiv preprint arXiv:2103.11943. 2021. 10.48550/arXiv.2103.11943 [DOI] [Google Scholar]
  • 120. Hao  Y, Dong  L, Wei  F  et al. Visualizing and understanding the effectiveness of BERT. arXiv preprint arXiv:1908.05620. 2019. 10.48550/arXiv.1908.05620 [DOI] [Google Scholar]
  • 121. Zhang  Z, Han  X, Liu  Z  et al. ERNIE: enhanced language representation with informative entities. arXiv preprint arXiv:1905.07129. 2019. 10.48550/arXiv.1905.07129 [DOI] [Google Scholar]
  • 122. Sun  Y, Wang  S, Li  Y  et al. Ernie 2.0: A continual pre-training framework for language understanding. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 34, pp. 8968–75. New York, USA, 2020. 10.1609/aaai.v34i05.6428 [DOI] [Google Scholar]
  • 123. Sun  Y, Wang  S, Feng  S  et al. Ernie 3.0: large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137. 2021. 10.48550/arXiv.2107.02137 [DOI] [Google Scholar]
  • 124. Wang  S, Sun  Y, Xiang  Y  et al. Ernie 3.0 titan: exploring larger-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2112.12731. 2021. 10.48550/arXiv.2112.12731 [DOI] [Google Scholar]
  • 125. Zulfiqar  H, Guo  Z, Grace-Mercure  BK  et al. Empirical comparison and recent advances of computational prediction of hormone binding proteins using machine learning methods. Comput Struct Biotechnol J  2023;21:2253–61. 10.1016/j.csbj.2023.03.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126. Zulfiqar  H, Ahmed  Z, Ma  CY  et al. Comprehensive prediction of lipocalin proteins using artificial intelligence strategy. Front Biosci Landmark  2022;27:84. 10.31083/j.fbl2703084 [DOI] [PubMed] [Google Scholar]
  • 127. Bakanina Kissanga  G-M, Zulfiqar  H, Gao  S  et al. E-mula: an ensemble multi-localized attention feature extraction network for viral protein subcellular localization. Information  2024;15:163. 10.3390/info15030163 [DOI] [Google Scholar]
  • 128. Xie  H, Wang  L, Qian  Y  et al. Methyl-GP: accurate generic DNA methylation prediction based on a language model and representation learning. Nucleic Acids Res  2025;53:gkaf223. 10.1093/nar/gkaf223 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129. Yan  K, Lv  H, Shao  J  et al. TPpred-SC: multi-functional therapeutic peptide prediction based on multi-label supervised contrastive learning. Sci China Inf Sci  2024;67:212105:1–12. 10.1007/s11432-024-4147-8 [DOI] [Google Scholar]
  • 130. Huang  Z, Xiao  Z, Ao  C  et al. Computational approaches for predicting drug-disease associations: a comprehensive review. Front Comput Sci  2025;19:1–15. 10.1007/s11704-024-40072-y [DOI] [Google Scholar]
  • 131. Huang  Z, Guo  X, Qin  J  et al. Accurate RNA velocity estimation based on multibatch network reveals complex lineage in batch scRNA-seq data. BMC Biol  2024;22:290. 10.1186/s12915-024-02085-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132. Guo  X, Huang  Z, Ju  F  et al. Highly accurate estimation of cell type abundance in bulk tissues based on single-cell reference and domain adaptive matching. Adv Sci  2024;11:e2306329. 10.1002/advs.202306329 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133. Jin  G, Xu  M, Zou  M  et al. The processing, gene regulation, biological functions, and clinical relevance of N4-acetylcytidine on RNA: a systematic review. Mol Ther Nucleic Acids  2020;20:13–24. 10.1016/j.omtn.2020.01.037 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

There is no data.


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES