Skip to main content
Molecular Therapy. Nucleic Acids logoLink to Molecular Therapy. Nucleic Acids
. 2019 Aug 5;18:16–23. doi: 10.1016/j.omtn.2019.07.019

Selecting Essential MicroRNAs Using a Novel Voting Method

Xiaoqing Ru 1,2,5, Peigang Cao 3,5, Lihong Li 2, Quan Zou 1,4,∗
PMCID: PMC6727015  PMID: 31479921

Abstract

Among the large number of known microRNAs (miRNAs), some miRNAs play negligible roles in cell regulation. Therefore, selecting essential miRNAs is an important initial step for a deeper understanding of miRNAs and their functions. In this study, we generated 60 classification models by combining 12 representative feature extraction methods and 5 commonly used classification algorithms. The optimal model for essential miRNA classification that we obtained is based on the Mismatch feature extraction method combined with the random forest algorithm. The F-Measure, area under the curve, and accuracy values of this model were 93.2%, 96.7%, and 93.0%, respectively. We also found that the distribution of the positive and negative examples of the first few features greatly influenced the classification results. The feature extraction methods performed best when the differences between the positive and negative examples were obvious, and this led to better classification of essential miRNAs. Because each classifier’s predictions for the same sample may be different, we employed a novel voting method to improve the accuracy of the classification of essential miRNAs. The performance results showed that the best classification results were obtained when five classification models were used in the voting. The five classification models were constructed based on the Mismatch, pseudo-distance structure status pair composition, Subsequence, Kmer, and Triplet feature extraction methods. The voting result was 95.3%. Our results suggest that the voting method can be an important tool for selecting essential miRNAs.

Keywords: miRNA, biological function, classification, feature extraction, voting

Introduction

MicroRNAs (miRNAs) are short noncoding RNAs that are found widely in eukaryotes.1 Their breadth and diversity indicate that they have a very wide variety of biological functions. They are involved in many important biological processes in cells, including regulating the expression of genes that encode proteins involved in biological development,2, 3, 4 cell proliferation,5 differentiation,6 and apoptosis.7 miRNAs are associated with cancer8, 9, 10 and other diseases.11, 12, 13, 14, 15 Drugs that target genes have been developed based on miRNA gene silencing and have been applied to some previously incurable diseases that threaten human health.16, 17, 18, 19, 20 miRNAs also play important roles in cell adaptation to abnormal environments, such as freezing, dehydration, and hypoxia.21, 22, 23 Because of the many biological functions of miRNAs, a lot of attention has been given to miRNA-related problems in bioinformatics.24, 25, 26, 27, 28, 29, 30

Accurate identification of miRNA sequences is one such problem that has achieved good results. For example, in 2013, Wei et al.31 constructed a classifier to identify miRNAs using a high-quality negative set and reported a classification accuracy rate of 93%. In 2015, Peace et al.1 proposed a framework for improving miRNA prediction in non-human genomes using sequence conservation and phylogenetic distance information. Their framework uses accuracy, sensitivity, and specificity parameters to obtain species-specific predictions. In 2016, Jiang et al.5 used a backpropagation neural network algorithm to identify miRNAs in Arabidopsis. In their model, the precision and recall rates were 95% and 96%, respectively; however, these results do not make much sense for the in-depth study of miRNAs. The reasons for this failure were likely because of the recent dramatic increase in known miRNAs (e.g., miRBase [Release 22.1: October 2018] contains 38,589 miRNA sequences from 271 species32) and the proposal that some miRNAs or miRNA families have negligible effects in cell development.33 Therefore, to efficiently study the biological mechanisms of miRNAs, it is necessary to detect essential miRNAs from among the many other miRNAs.

Two important factors that influence miRNA prediction results are the feature extraction method and classification algorithm selected. A good feature extraction method will fully express the sequence information. The existing methods for RNA feature extraction can be divided into four categories: those based on ribonucleic acid composition, autocorrelation, pseudo ribonucleic acid composition, or predicted structure composition.34, 35 Methods based on RNA sequence composition include basic kmer (Kmer), Mismatch, and Subsequence.31, 36, 37, 38, 39 Kmer31 represents RNA sequences as the frequency of occurrence of k adjacent bases and is the simplest of the three methods. Methods based on autocorrelation include dinucleotide-based auto-covariance (DAC),40 dinucleotide-based cross-covariance (DCC),41 dinucleotide-based auto-cross-covariance (DACC; a combination of DAC and DCC),42 Moran autocorrelation (MAC),43 Geary autocorrelation (GAC),44 and normalized Moreau-Broto autocorrelation (NMBAC).45 Methods based on the pseudo-RNA composition46, 47 include general parallel correlation pseudo-dinucleotide composition (PC-PseDNC-General) and its variant general series correlation pseudo-dinucleotide composition (SC-PseDNC-General). The equations used in PC-PseDNC-General and SC-PseDNC-General differ in that they calculate the correlation factors that reflect the sequence or the order correlations, respectively, among all of the consecutive dinucleotides along an RNA sequence.42 Methods based on the predicted structure composition include local structure-sequence triplet elements (Triplet),48, 49, 50 pseudo-structure status composition (PseSSC),26 and pseudo-distance structure status pair composition (PseDPC).13

The most commonly used classification algorithms are random forest and support vector machine. Random forest51, 52, 53, 54, 55, 56, 57 can be considered an integrated algorithm that reduces the one-sidedness and inaccuracy of a single decision tree by combining multiple different decision trees. Support vector machine13, 48, 58, 59, 60, 61, 62, 63, 64, 65, 66 maximizes the classification of positive and negative examples by constructing a hyperplane. Other machine learning algorithms also have been used for classification and recognition, such as neural networks,5, 67, 68 Naive Bayes,69, 70 evolutionary algorithms,71 and ensemble learning.72, 73, 74, 75, 76

The aims of this study were: (1) to construct a classification model by combining 12 different feature extraction algorithms and 5 classification algorithms to find the most suitable model for essential miRNA classification; (2) to explore the distribution of positive and negative examples under different feature extraction methods, and to determine the influence of distribution differences between positive and negative examples on classification results; and (3) to further improve the classification accuracy of essential miRNAs through a novel voting method. The performance of the optimal classification model shows the validity of our conclusions and methods.

Results and Discussion

Determine the Parameters for Kmer, Mismatch, and Subsequence

For these three feature extraction algorithms, the parameter k, which has 4k-dimensional features, has to be set. For k = 1, the extracted features do not represent the complete sequence information. For k > 5, the extracted features will have more than 1,024 dimensions. When the dimensions are very high, the computational time can be very long, and over-fitting phenomenon and dimensionality disaster may occur. To avoid these problems, we set k = 2, 3, and 4. The results for each of the methods on the pre-miRNA dataset are shown in Table 1. Each performance value was taken from the best classification model under each method.

Table 1.

Performances of the Three Feature Extraction Methods with Different k Values


Kmer
Mismatch
Subsequence
K Values Sn Sp ACC Sn Sp ACC Sn Sp ACC
2 67.0 80.6 73.9 96.4 89.7 93.0 91.7 88.6 90.1
3 76.4 85.2 80.9 94.1 93.1 93.6 96.4 87.5 90.7
4 83.5 90.9 87.2 89.4 92.0 90.7 90.5 88.6 89.5

ACC, accuracy; Sn, sensitivity; Sp, specificity.

In the Kmer-based model, the performance was best for k = 4. In the models that used the Mismatch and Subsequence feature extraction methods, all three k values produced similar results that were better than the classification results obtained with the Kmer-based model (Table 1). Smaller k values will require a shorter computational time, so k = 2 was selected as the best value for the Mismatch- and Subsequence-based models.

Selection of the Best Classification Model for Predicting Essential miRNAs

A total of 12 feature extraction methods were used in this study. The models based on Kmer, Mismatch, and Subsequence involve parameter setting, as described in Determine the Parameters for Kmer, Mismatch, and Subsequence. We combined the 12 feature extraction methods with 5 commonly used classification algorithms to obtain 60 classification models. The accuracy of these models on the pre-miRNA dataset is shown in Figure 1.

Figure 1.

Figure 1

Accuracy of All the Classification Models

The accuracy of each of the classification models varied depending on the feature extraction method that was used (Figure 1). Three classification models had accuracies >90%, namely, Mismatch + random forest, PseDPC + support vector machine, and Subsequence + Logistic. Detailed performance information is shown in Table 2.

Table 2.

Performances of the Three Best Classification Models

Methods Dimensions F-Measure (%) AUC (%) ACC (%)
Mismatch 16 93.2 96.7 93.0
PseDPC 515 92.6 93.0 93.0
Subsequence 16 90.2 95.1 90.1

ACC, accuracy; AUC, area under the curve.

The Mismatch + random forest model, which has low dimensionality and good performance, was considered the optimal model for predicting essential miRNAs in the dataset (Table 2).

Representation of Important Features of the Classification Models

We explored the distribution of positive and negative examples under the different feature extraction methods. As shown in Figure 1, the accuracy was <70% for all of the classification models when combined with the NMBAC and SC-PseDNC-General feature extraction methods, indicating these methods were not suitable for predicting essential miRNAs. From among the remaining 50 classification models, those with the best classification performance under each feature extraction method were selected. Among the 10 selected models, those based on the Mismatch, Triplet, and PC-PseDNC-General methods all showed better performances when combined with the random forest algorithm. On the basis of the ANOVA of these three feature extraction methods, we took out the first four-dimensional features that had the greatest influence on the classification results. The obtained distribution of positive and negative examples is shown in Figure 2. Clearly, the difference between the positive and negative examples is larger with Mismatch than with the other two methods. This indicates that a feature extraction method that produces a large difference in the distribution of positive and negative samples contributes to a better final classification result.

Figure 2.

Figure 2

Distribution of Positive and Negative Examples of Important Features Obtained Using Different Extraction Methods

Optimal Voting Results

To achieve higher prediction accuracy, we used a novel voting method to predict all samples based on the results from the 10 selected classification models described in Representation of Important Features of the Classification Models. The predictions of these 10 models for all samples are shown in Figure 3.

Figure 3.

Figure 3

Distribution of Mispredicted Samples for Each Classification Model

We obtained four types of voting results from the 10 classification models, and the best results for each type are shown in Table 3. Each voting process eliminates two classification models, and the eliminated models have strong correlations. For example, from the distribution of the erroneously predicted samples shown in Figure 3, the classification models based on the GAC and MAC feature extraction methods had strong correlations, and the correct or mispredicted samples were almost the same, so these methods were eliminated in the second type of voting. Excluding the most relevant classification models was beneficial to the final voting result, as shown in Table 4.

Table 3.

Best Results for Four Types of Voting

Category Model Included No. of Samples that Were Mispredicted Voting Results (%)
1 Mismatch + PseDPC + Subsequence,Kmer + Triplet + PseSSC, DACC,GAC + MAC 14 91.9
2 Mismatch + PseDPC + Subsequence,Kmer + Triplet + PseSSC + DACC 9 94.7
3 Mismatch + PseDPC + Subsequence,Kmer + Triplet 8 95.3
4 Mismatch + PseDPC + Subsequence 9 94.7

Table 4.

Verifying the Impact of Model Relevance on Voting

Reserved Classification Model No. of Samples that Were Mispredicted Voting Results (%)
Mismatch + PseDPC + Subsequence 9 94.7
Mismatch + PseDPC + Kmer 14 91.9
Mismatch + PseDPC + Triplet 9 94.7

As shown in Table 3, the results obtained by voting on the classification model based on the Mismatch, PseDPC, Subsequence, Kmer, and Triplet feature extraction methods were the best (accuracy rate of 95.3%). The accuracy of voting was higher than the accuracy of the case alone.

The predictions of the classification model based on the Triplet method were worse than those of the model based on Kmer (Figure 3), but after participating in the voting with the model based on Mismatch and PseDPC, the number of samples that were mispredicted with Triplet was less than the number with Kmer. This is because the model based on Kmer had a stronger correlation with the other two classification models. We chose to vote with the classification model based on Mismatch and PseDPC because these two feature extraction methods performed best in all of the classification models, and the correlation between them was very low.

The results shown in Tables 3 and 4 fully demonstrate that the novel voting method proposed in this study can achieve excellent results for the selection of essential miRNAs.

Conclusions

The aim of this study was to select essential miRNAs from a large number of miRNA sequences, thus making the study of the biological mechanisms of miRNAs more efficient. We used known mouse miRNAs as the dataset. We used different feature extraction methods to represent these data, then combined the extracted features with different classification algorithms to construct classification models. The final classification result was determined by a novel voting method, which gave a final voting result of 95.3%. This result showed that this method was effective in identifying the essential miRNAs in the dataset. In future work, we will focus on detecting new essential miRNAs, analyzing their function, and exploring the relationship between new miRNAs and diseases.27, 77

Materials and Methods

The general pipeline used in this study is shown in Figure 4.

Figure 4.

Figure 4

Flowchart of the Processes Used in This Study

Acquisition of Datasets

Acquisition of essential pre-miRNA sequences: miRNA genes produce primary miRNA (pri-miRNA) sequences that are 300–1,000 nt long. The pri-miRNAs are processed to precursor miRNA (pre-miRNA) sequences that are 60–70 nt long. Mature miRNAs, which are 20–24 nt long, are formed from pre-miRNAs by the action of enzymes.26, 48, 78 The hairpin structure of pre-miRNAs is an important feature that is widely used to identify miRNAs.48 In this study, we used pre-miRNA sequences from miRBase (http://www.mirbase.org/). We collected a total of 91 pre-miRNA sequences that are essential in mice from Bartel’s 2018 review of metazoan miRNAs.79 The 91 pre-miRNA sequences were from several families.

Acquisition of pre-miRNA sequences of unknown importance: We downloaded all the mice pre-miRNA sequences (a total of 1,234) from miRBase. All the sequences that belonged to the same families as the 91 essential miRNAs were excluded. The remaining pre-miRNAs (1,090) were considered to be non-essential pre-miRNAs.

To shorten the computational time and to ensure better performance results, we removed redundant sequences80 and obtained a final dataset that contained 85 essential and 88 non-essential pre-miRNA sequences.

Feature Extraction Methods

Many methods have been used for extracting features of RNA sequences. In this study, we used a number of feature extraction methods on the selected pre-miRNA sequences. We expressed a pre-miRNA sequence as R = r1, r2, r3, r4, r5, …, rL, where ri ∈{A, U, G, C},33 and L is the sequence length. These features can be generated easily using the repRNA web server81 and the BioSeq-Analysis platform.82

Kmer and Mismatch

Kmer is a simple feature extraction method in which k indicates the number of bases in a subsequence. For a given k value, there will be 4k seed sequences.31 For example, for k = 2, there are 16 subsequences, AA, AC, AU, AG, UA, UU, …, CC. Mismatch calculates the number of occurrences of subsequences containing k adjacent bases and uses a parameter m (maximum allowed error match = 0 ≤ m < k).83 For k = 2, m is 0, 1. For example, for the subsequence AC, the A– and –C subsequences are all regarded as AC.

Subsequence

Subsequence also counts the number of occurrences of a subsequence, but, in particular, it takes into account the length of the sequence that is eligible and the influence factor ∂.84 This method allows for spacing between the bases in a given subsequence. For example, there are four ways in which a search for AAC in a sequence is eligible: AAC, AAXXXC, AXXXAC, and AXXAC (where X can be U or G). For this example, the number of occurrences of AAC can be expressed as: 1 + ∂6 + ∂6 + ∂5. For AAC, where there is no gap between the bases, the length is considered to be 1.

Triplet

Triplet extracts features based on secondary structure information in the sequence. Two states are considered for each nucleotide: matched (represented by “(” or “)” for A paired with U and G paired with C) and unmatched (represented by “.”). The possible structural forms for any three adjacent nucleotides are: “(((,” “((.,” “(.(,” “.((,” “(..,” “.(., ” “..(,” and “….” For the four bases (A, U, G, C) in a pre-miRNA sequence, there will be 8 × 4 different triplets. Triplet counts the number of occurrences of these 32 triplets in a sequence.48

PseDPC

Pre-miRNA sequences can be represented by 10 secondary structure states: A, C, G, U, A–U, U–A, G–C, C–G, G–U, and U–G. Within a given range of distance thresholds, the PseDPC algorithm counts the frequency of occurrence of the two-two combination between the structural states. When the distance d is equal to 0, only the frequencies of each of the 10 states are counted. When d is ≥ 1, there will be 100 combinations of the 10 states, so the frequency of occurrence of the 100 combinations is counted. Therefore, 10 + 100d dimensional features can be extracted. Then the features of the free energy between the two structural states of a combination are extracted, giving φ features, where φ is the highest counted rank of the structural correlation along an RNA chain.13

Voting Process

The results produced by each classifier were integrated using a voting process as follows. First, each classification algorithm will correctly or incorrectly predict each of the samples (a sequence is considered a sample). Correctly predicted samples are not marked, the samples that were mispredicted are marked with “+,” and the number of classifiers is represented by n. Second, when n is a singular number ≥3, the number of times each sample is erroneously predicted under n classifiers is counted and recorded as m, where m ≥ n + 1/2; the sample is considered to be a sample that was mispredicted. Third, when n is a double number ≥4, the number of times each sample is mispredicted under the n classifiers is counted, and the sample corresponding to m = n/2 is selected. Suppose that there are z samples that meet the requirements, and the z samples are predicted differently by the different classifiers, the classifier with the highest number of mispredicted samples is eliminated. Then, n is singular and step 2 can be repeated. Fourth, when n satisfies the condition of being a singular number, steps 2 and 3 can count the number of samples considered to be mispredicted in various voting processes, from which the final voting result can be derived.

Author Contributions

X.R. implemented the experiments and drafted the manuscript. P.C., L.L., and Q.Z. initiated the idea, conceived the whole process, and finalized the paper. All authors have read and approved the final manuscript.

Acknowledgments

The work was supported by the National Key R&D Program of China (grant 2018YFC0910405) and the Natural Science Foundation of China (grant 61771331).

References

  • 1.Peace R.J., Biggar K.K., Storey K.B., Green J.R. A framework for improving microRNA prediction in non-human genomes. Nucleic Acids Res. 2015;43:e138. doi: 10.1093/nar/gkv698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.La Torre A., Georgi S., Reh T.A. Conserved microRNA pathway regulates developmental timing of retinal neurogenesis. Proc. Natl. Acad. Sci. USA. 2013;110:E2362–E2370. doi: 10.1073/pnas.1301837110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Cheng L., Sun J., Xu W., Dong L., Hu Y., Zhou M. OAHG: an integrated resource for annotating human genes with multi-level ontologies. Sci. Rep. 2016;6:34820. doi: 10.1038/srep34820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Hu Y., Zhao T., Zhang N., Zang T., Zhang J., Cheng L. Identifying diseases-related metabolites using random walk. BMC Bioinformatics. 2018;19(Suppl 5):116. doi: 10.1186/s12859-018-2098-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jiang L., Zhang J., Xuan P., Zou Q. BP Neural Network Could Help Improve Pre-miRNA Identification in Various Species. BioMed Res. Int. 2016;2016:9565689. doi: 10.1155/2016/9565689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Le M.T., Xie H., Zhou B., Chia P.H., Rizk P., Um M., Udolph G., Yang H., Lim B., Lodish H.F. MicroRNA-125b promotes neuronal differentiation in human cells by repressing multiple targets. Mol. Cell. Biol. 2009;29:5290–5305. doi: 10.1128/MCB.01694-08. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Körner C., Keklikoglou I., Bender C., Wörner A., Münstermann E., Wiemann S. MicroRNA-31 sensitizes human breast cells to apoptosis by direct targeting of protein kinase C epsilon (PKCepsilon) J. Biol. Chem. 2013;288:8750–8761. doi: 10.1074/jbc.M112.414128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jiang L., Xiao Y., Ding Y., Tang J., Guo F. Discovering Cancer Subtypes via an Accurate Fusion Strategy on Multiple Profile Data. Front. Genet. 2019;10:20. doi: 10.3389/fgene.2019.00020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Yu L., Zhao J., Gao L. Predicting Potential Drugs for Breast Cancer based on miRNA and Tissue Specificity. Int. J. Biol. Sci. 2018;14:971–982. doi: 10.7150/ijbs.23350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Pavithra D., Sabitha K., Rajkumar T. Identification of small molecule inhibitors for differentially expressed miRNAs in gastric cancer. Comput. Biol. Chem. 2018;77:442–454. doi: 10.1016/j.compbiolchem.2018.07.013. [DOI] [PubMed] [Google Scholar]
  • 11.Jiang Q., Wang Y., Hao Y., Juan L., Teng M., Zhang X., Li M., Wang G., Liu Y. miR2Disease: a manually curated database for microRNA deregulation in human disease. Nucleic Acids Res. 2009;37:D98–D104. doi: 10.1093/nar/gkn714. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Cheng L., Hu Y. Human Disease System Biology. Curr. Gene Ther. 2018;18:255–256. doi: 10.2174/1566523218666181010101114. [DOI] [PubMed] [Google Scholar]
  • 13.Liu B., Fang L., Liu F., Wang X., Chou K.C. iMiRNA-PseDPC: microRNA precursor identification with a pseudo distance-pair composition approach. J. Biomol. Struct. Dyn. 2016;34:223–235. doi: 10.1080/07391102.2015.1014422. [DOI] [PubMed] [Google Scholar]
  • 14.Liu G., Xu Y., Jiang Y., Zhang L., Feng R., Jiang Q. PICALM rs3851179 variant confers susceptibility to Alzheimer’s disease in Chinese population. Mol. Neurobiol. 2017;54:3131–3136. doi: 10.1007/s12035-016-9886-2. [DOI] [PubMed] [Google Scholar]
  • 15.Hu Y., Zhao T., Zang T., Zhang Y., Cheng L. Identification of Alzheimer’s Disease-Related Genes Based on Data Integration Method. Front. Genet. 2019;9:703. doi: 10.3389/fgene.2018.00703. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Kelly P.S., Gallagher C., Clynes M., Barron N. Conserved microRNA function as a basis for Chinese hamster ovary cell engineering. Biotechnol. Lett. 2015;37:787–798. doi: 10.1007/s10529-014-1751-7. [DOI] [PubMed] [Google Scholar]
  • 17.Jiang Q., Jin S., Jiang Y., Liao M., Feng R., Zhang L., Liu G., Hao J. Alzheimer’s Disease Variants with the Genome-Wide Significance are Significantly Enriched in Immune Pathways and Active in Immune Cells. Mol. Neurobiol. 2017;54:594–600. doi: 10.1007/s12035-015-9670-8. [DOI] [PubMed] [Google Scholar]
  • 18.Liu G., Zhao Y., Jin S., Hu Y., Wang T., Tian R., Han Z., Xu D., Jiang Q. Circulating vitamin E levels and Alzheimer’s disease: a Mendelian randomization study. Neurobiol. Aging. 2018;72:189.e1–189.e9. doi: 10.1016/j.neurobiolaging.2018.08.008. [DOI] [PubMed] [Google Scholar]
  • 19.Liu G., Zhang Y., Wang L., Xu J., Chen X., Bao Y., Hu Y., Jin S., Tian R., Bai W. Alzheimer’s Disease rs11767557 Variant Regulates EPHA1 Gene Expression Specifically in Human Whole Blood. J. Alzheimers Dis. 2018;61:1077–1088. doi: 10.3233/JAD-170468. [DOI] [PubMed] [Google Scholar]
  • 20.Liu G., Wang T., Tian R., Hu Y., Han Z., Wang P., Zhou W., Ren P., Zong J., Jin S., Jiang Q. Alzheimer’s Disease Risk Variant rs2373115 Regulates GAB2 and NARS2 Expression in Human Brain Tissues. J. Mol. Neurosci. 2018;66:37–43. doi: 10.1007/s12031-018-1144-9. [DOI] [PubMed] [Google Scholar]
  • 21.Biggar K.K., Kornfeld S.F., Maistrovski Y., Storey K.B. MicroRNA regulation in extreme environments: differential expression of microRNAs in the intertidal snail Littorina littorea during extended periods of freezing and anoxia. Genomics Proteomics Bioinformatics. 2012;10:302–309. doi: 10.1016/j.gpb.2012.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Biggar K.K., Storey K.B. Evidence for cell cycle suppression and microRNA regulation of cyclin D1 during anoxia exposure in turtles. Cell Cycle. 2012;11:1705–1713. doi: 10.4161/cc.19790. [DOI] [PubMed] [Google Scholar]
  • 23.Wu C.W., Biggar K.K., Storey K.B. Dehydration mediated microRNA response in the African clawed frog Xenopus laevis. Gene. 2013;529:269–275. doi: 10.1016/j.gene.2013.07.064. [DOI] [PubMed] [Google Scholar]
  • 24.Jiang L., Xiao Y., Ding Y., Tang J., Guo F. FKL-Spa-LapRLS: an accurate method for identifying human microRNA-disease association. BMC Genomics. 2018;19(Suppl 10):911. doi: 10.1186/s12864-018-5273-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Cheng L., Hu Y., Sun J., Zhou M., Jiang Q. DincRNA: a comprehensive web-based bioinformatics toolkit for exploring disease associations and ncRNA function. Bioinformatics. 2018;34:1953–1956. doi: 10.1093/bioinformatics/bty002. [DOI] [PubMed] [Google Scholar]
  • 26.Liu B., Fang L., Liu F., Wang X., Chen J., Chou K.C. Identification of real microRNA precursors with a pseudo structure status composition approach. PLoS ONE. 2015;10:e0121501. doi: 10.1371/journal.pone.0121501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Jiang Q., Wang G., Jin S., Li Y., Wang Y. Predicting human microRNA-disease associations based on support vector machine. Int. J. Data Min. Bioinform. 2013;8:282–293. doi: 10.1504/ijdmb.2013.056078. [DOI] [PubMed] [Google Scholar]
  • 28.Wang G., Luo X., Wang J., Wan J., Xia S., Zhu H., Qian J., Wang Y. MeDReaders: a database for transcription factors that bind to methylated DNA. Nucleic Acids Res. 2018;46(D1):D146–D151. doi: 10.1093/nar/gkx1096. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wang G., Wang F., Huang Q., Li Y., Liu Y., Wang Y. Understanding Transcription Factor Regulation by Integrating Gene Expression and DNase I Hypersensitive Sites. BioMed Res. Int. 2015;2015:757530. doi: 10.1155/2015/757530. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Gong W., Huang Y., Xie J., Wang G., Yu D., Sun X. Genome-wide identification and characterization of conserved and novel microRNAs in grass carp (Ctenopharyngodon idella) by deep sequencing. Comput. Biol. Chem. 2017;68:92–100. doi: 10.1016/j.compbiolchem.2017.02.010. [DOI] [PubMed] [Google Scholar]
  • 31.Wei L., Liao M., Gao Y., Ji R., He Z., Zou Q. Improved and Promising Identification of Human MicroRNAs by Incorporating a High-Quality Negative Set. IEEE/ACM Trans. Comput. Biol. Bioinform. 2014;11:192–201. doi: 10.1109/TCBB.2013.146. [DOI] [PubMed] [Google Scholar]
  • 32.Kozomara A., Birgaoanu M., Griffiths-Jones S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019;47(D1):D155–D162. doi: 10.1093/nar/gky1141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Alvarez-Saavedra E., Horvitz H.R. Many families of C. elegans microRNAs are not essential for development or viability. Curr. Biol. 2010;20:367–373. doi: 10.1016/j.cub.2009.12.051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Jiang Q., Ma R., Wang J., Wu X., Jin S., Peng J., Tan R., Zhang T., Li Y., Wang Y. LncRNA2Function: a comprehensive resource for functional investigation of human lncRNAs based on RNA-seq data. BMC Genomics. 2015;16(Suppl 3):S2. doi: 10.1186/1471-2164-16-S3-S2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Cheng L., Wang P., Tian R., Wang S., Guo Q., Luo M., Zhou W., Liu G., Jiang H., Jiang Q. LncRNA2Target v2.0: a comprehensive database for target genes of lncRNAs in human and mouse. Nucleic Acids Res. 2019;47(D1):D140–D144. doi: 10.1093/nar/gky1051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Leslie C.S., Eskin E., Cohen A., Weston J., Noble W.S. Mismatch string kernels for discriminative protein classification. Bioinformatics. 2004;20:467–476. doi: 10.1093/bioinformatics/btg431. [DOI] [PubMed] [Google Scholar]
  • 37.El-Manzalawy Y., Dobbs D., Honavar V. Predicting flexible length linear B-cell epitopes. Comput. Syst. Bioinformatics Conf. 2008;7:121–132. [PMC free article] [PubMed] [Google Scholar]
  • 38.Liu B., Liu F., Wang X., Chen J., Fang L., Chou K.-C. Pse-in-One: a web server for generating various modes of pseudo components of DNA, RNA, and protein sequences. Nucleic Acids Res. 2015;43:W65–W71. doi: 10.1093/nar/gkv458. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Chen W., Yang H., Feng P., Ding H., Lin H. iDNA4mC: identifying DNA N4-methylcytosine sites based on nucleotide chemical properties. Bioinformatics. 2017;33:3518–3523. doi: 10.1093/bioinformatics/btx479. [DOI] [PubMed] [Google Scholar]
  • 40.Guo Y., Yu L., Wen Z., Li M. Using support vector machine combined with auto covariance to predict protein-protein interactions from protein sequences. Nucleic Acids Res. 2008;36:3025–3030. doi: 10.1093/nar/gkn159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Dong Q., Zhou S., Guan J. A new taxonomy-based protein fold recognition approach based on autocross-covariance transformation. Bioinformatics. 2009;25:2655–2662. doi: 10.1093/bioinformatics/btp500. [DOI] [PubMed] [Google Scholar]
  • 42.Chen W., Zhang X., Brooker J., Lin H., Zhang L., Chou K.C. PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions. Bioinformatics. 2015;31:119–120. doi: 10.1093/bioinformatics/btu602. [DOI] [PubMed] [Google Scholar]
  • 43.Horne D.S. Prediction of protein helix content from an autocorrelation analysis of sequence hydrophobicities. Biopolymers. 1988;27:451–477. doi: 10.1002/bip.360270308. [DOI] [PubMed] [Google Scholar]
  • 44.Sokal R.R., Thomson B.A. Population structure inferred by local spatial autocorrelation: an example from an Amerindian tribal population. Am. J. Phys. Anthropol. 2006;129:121–131. doi: 10.1002/ajpa.20250. [DOI] [PubMed] [Google Scholar]
  • 45.Feng Z.P., Zhang C.T. Prediction of membrane protein types based on the hydrophobic index of amino acids. J. Protein Chem. 2000;19:269–275. doi: 10.1023/a:1007091128394. [DOI] [PubMed] [Google Scholar]
  • 46.Feng C.Q., Zhang Z.Y., Zhu X.J., Lin Y., Chen W., Tang H., Lin H. iTerm-PseKNC: a sequence-based tool for predicting bacterial transcriptional terminators. Bioinformatics3024762535. 2019:1469–1477. doi: 10.1093/bioinformatics/bty827. [DOI] [PubMed] [Google Scholar]
  • 47.Dao F.Y., Lv H., Wang F., Feng C.Q., Ding H., Chen W., Lin H. Identify origin of replication in Saccharomyces cerevisiae using two-step feature selection technique. Bioinformatics. 2019;35:2075–2083. doi: 10.1093/bioinformatics/bty943. [DOI] [PubMed] [Google Scholar]
  • 48.Xue C., Li F., He T., Liu G.P., Li Y., Zhang X. Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine. BMC Bioinformatics. 2005;6:310. doi: 10.1186/1471-2105-6-310. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zhu X.J., Feng C.Q., Lai H.Y., Chen W., Lin H. Predicting protein structural classes for low-similarity sequences by evaluating different features. Knowl. Based Syst. 2019;163:787–793. [Google Scholar]
  • 50.Tan J.X., Li S.H., Zhang Z.M., Chen C.X., Chen W., Tang H., Lin H. Identification of hormone binding proteins based on machine learning methods. Math. Biosci. Eng. 2019;16:2466–2480. doi: 10.3934/mbe.2019123. [DOI] [PubMed] [Google Scholar]
  • 51.Yao Y., Li X., Liao B., Huang L., He P., Wang F., Yang J., Sun H., Zhao Y., Yang J. Predicting influenza antigenicity from Hemagglutintin sequence data based on a joint random forest method. Sci. Rep. 2017;7:1545. doi: 10.1038/s41598-017-01699-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Cutler A., Cutler D., Stevens J. Random Forests. Machine Learning. 2011;45:157–176. [Google Scholar]
  • 53.Ding Y., Tang J., Guo F. Predicting protein-protein interactions via multivariate mutual information of protein sequences. BMC Bioinformatics. 2016;17:398. doi: 10.1186/s12859-016-1253-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Liu B., Yang F., Huang D.S., Chou K.-C. iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNC. Bioinformatics. 2018;34:33–40. doi: 10.1093/bioinformatics/btx579. [DOI] [PubMed] [Google Scholar]
  • 55.Yu L., Su R., Wang B., Zhang L., Zou Y., Zhang J., Gao L. Prediction of Novel Drugs for Hepatocellular Carcinoma Based on Multi-Source Random Walk. IEEE/ACM Trans. Comput. Biol. Bioinformatics. 2017;14:966–977. doi: 10.1109/TCBB.2016.2550453. [DOI] [PubMed] [Google Scholar]
  • 56.Cheng L., Jiang Y., Ju H., Sun J., Peng J., Zhou M., Hu Y. InfAcrOnt: calculating cross-ontology term similarities using information flow by a random walk. BMC Genomics. 2018;19(Suppl 1):919. doi: 10.1186/s12864-017-4338-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Cheng L., Shi H., Wang Z., Hu Y., Yang H., Zhou C., Sun J., Zhou M. IntNetLncSim: an integrative network analysis method to infer human lncRNA functional similarity. Oncotarget. 2016;7:47864–47874. doi: 10.18632/oncotarget.10012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Hearst M.A., Dumais S.T., Osuna E., Platt J., Scholkopf B. Support vector machines. IEEE Intell. Syst. 1998;13:18–28. [Google Scholar]
  • 59.Ding Y., Tang J., Guo F. Identification of drug-target interactions via multiple information integration. Inf. Sci. 2017;418–419:546–560. [Google Scholar]
  • 60.Yang H., Lv H., Ding H., Chen W., Lin H. iRNA-2OM: A Sequence-Based Predictor for Identifying 2′-O-Methylation Sites in Homo sapiens. J. Comput. Biol. 2018;25:1266–1277. doi: 10.1089/cmb.2018.0004. [DOI] [PubMed] [Google Scholar]
  • 61.Yang W., Zhu X.J., Huang J., Ding H., Lin H. A brief survey of machine learning methods in protein sub-Golgi localization. Curr. Bioinform. 2019;14:234–240. [Google Scholar]
  • 62.Liu Y., Wang X., Liu B. A comprehensive review and comparison of existing computational methods for intrinsically disordered protein and region prediction. Brief. Bioinform. 2019;20:330–346. doi: 10.1093/bib/bbx126. [DOI] [PubMed] [Google Scholar]
  • 63.Chen W., Lv H., Nie F., Lin H. i6mA-Pred: Identifying DNA N6-methyladenine sites in the rice genome. Bioinformatics. 2019;35:2796–2800. doi: 10.1093/bioinformatics/btz015. [DOI] [PubMed] [Google Scholar]
  • 64.Sun Y., Xiong Y., Xu Q., Wei D. A hadoop-based method to predict potential effective drug combination. BioMed Res. Int. 2014;2014:196858. doi: 10.1155/2014/196858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.He J., Fang T., Zhang Z., Huang B., Zhu X., Xiong Y. PseUI: Pseudouridine sites identification based on RNA sequence information. BMC Bioinformatics. 2018;19:306. doi: 10.1186/s12859-018-2321-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Zhao Y., Wang F., Juan L. MicroRNA Promoter Identification in Arabidopsis Using Multiple Histone Markers. BioMed Res. Int. 2015;2015:861402. doi: 10.1155/2015/861402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Song T., Rodríguez-Patón A., Zheng P., Zeng X. Spiking Neural P Systems with Colored Spikes. IEEE Trans. Cogn. Dev. Syst. 2018;10:1106–1115. doi: 10.1109/TNB.2018.2873221. [DOI] [PubMed] [Google Scholar]
  • 68.Cabarle F.G.C., Adorna H.N., Jiang M., Zeng X. Spiking Neural P Systems With Scheduled Synapses. IEEE Trans. Nanobioscience. 2017;16:792–801. doi: 10.1109/TNB.2017.2762580. [DOI] [PubMed] [Google Scholar]
  • 69.Feng P.M., Ding H., Chen W., Lin H. Naïve Bayes classifier with feature selection to identify phage virion proteins. Comput. Math. Methods Med. 2013;2013:530696. doi: 10.1155/2013/530696. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Feng P.M., Lin H., Chen W. Identification of antioxidants from sequence information using naïve Bayes. Comput. Math. Methods Med. 2013;2013:567529. doi: 10.1155/2013/567529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Xu H., Zeng W., Zhang D., Zeng X. MOEA/HD: A Multiobjective Evolutionary Algorithm Based on Hierarchical Decomposition. IEEE Trans. Cybern. 2019;49:517–526. doi: 10.1109/TCYB.2017.2779450. [DOI] [PubMed] [Google Scholar]
  • 72.Wei L., Xing P., Zeng J., Chen J., Su R., Guo F. Improved prediction of protein-protein interactions using novel negative samples, features, and an ensemble classifier. Artif. Intell. Med. 2017;83:67–74. doi: 10.1016/j.artmed.2017.03.001. [DOI] [PubMed] [Google Scholar]
  • 73.Wei L., Wan S., Guo J., Wong K.K. A novel hierarchical selective ensemble classifier with bioinformatics application. Artif. Intell. Med. 2017;83:82–90. doi: 10.1016/j.artmed.2017.02.005. [DOI] [PubMed] [Google Scholar]
  • 74.Wei L., Chen H., Su R. M6APred-EL: A Sequence-Based Predictor for Identifying N6-methyladenosine Sites Using Ensemble Learning. Mol. Ther. Nucleic Acids. 2018;12:635–644. doi: 10.1016/j.omtn.2018.07.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.You R., Zhang Z., Xiong Y., Sun F., Mamitsuka H., Zhu S. GOLabeler: improving sequence-based large-scale protein function prediction by learning to rank. Bioinformatics. 2018;34:2465–2473. doi: 10.1093/bioinformatics/bty130. [DOI] [PubMed] [Google Scholar]
  • 76.Xiong Y., Wang Q., Yang J., Zhu X., Wei D.Q. PredT4SE-Stack: Prediction of Bacterial Type IV Secreted Effectors From Protein Sequences Using a Stacked Ensemble Method. Front. Microbiol. 2018;9:2571. doi: 10.3389/fmicb.2018.02571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Jiang Q., Hao Y., Wang G., Juan L., Zhang T., Teng M., Liu Y., Wang Y. Prioritization of disease microRNAs through a human phenome-microRNAome network. BMC Syst. Biol. 2010;4(Suppl 1):S2. doi: 10.1186/1752-0509-4-S1-S2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Jiang P., Wu H., Wang W., Ma W., Sun X., Lu Z. MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features. Nucleic Acids Res. 2007;35:W339–W344. doi: 10.1093/nar/gkm368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Bartel D.P. Metazoan MicroRNAs. Cell. 2018;173:20–51. doi: 10.1016/j.cell.2018.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zou Q., Lin G., Jiang X., Liu X., Zeng X. Sequence clustering in bioinformatics: an empirical study. Brief. Bioinform. 2018;10:1106–1115. doi: 10.1093/bib/bby090. [DOI] [PubMed] [Google Scholar]
  • 81.Liu B., Liu F., Fang L., Wang X., Chou K.-C. repRNA: a web server for generating various feature vectors of RNA sequences. Mol. Genet. Genomics. 2016;291:473–481. doi: 10.1007/s00438-015-1078-7. [DOI] [PubMed] [Google Scholar]
  • 82.Liu B. BioSeq-Analysis: a platform for DNA, RNA and protein sequence analysis based on machine learning approaches. Brief. Bioinform. 2017 doi: 10.1093/bib/bbx165. 2017, bbx165. [DOI] [PubMed] [Google Scholar]
  • 83.Luo L., Li D., Zhang W., Tu S., Zhu X., Tian G. Accurate Prediction of Transposon-Derived piRNAs by Integrating Various Sequential and Physicochemical Features. PLoS ONE. 2016;11:e0153268. doi: 10.1371/journal.pone.0153268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Lodhi H., Saunders C., Shawe-Taylor J., Cristianini N., Watkins C. Text Classification using String Kernels. J. Mach. Learn. Res. 2002;2:419–444. [Google Scholar]

Articles from Molecular Therapy. Nucleic Acids are provided here courtesy of The American Society of Gene & Cell Therapy

RESOURCES