Skip to main content
Springer logoLink to Springer
. 2025 Aug 1;115(4):102. doi: 10.1007/s11103-025-01632-3

TISCalling: leveraging machine learning to identify translational initiation sites in plants and viruses

Ming-Ren Yen 1, Ya-Ru Li 4, Chia-Yi Cheng 2,3, Ting-Ying Wu 1,, Ming-Jung Liu 4,5,
PMCID: PMC12316744  PMID: 40748408

Abstract

The recognition of translational initiation sites (TISs) offers complementary insights into identifying genes encoding novel proteins or small peptides. Conventional computational methods primarily identify Ribo-seq-supported TISs and lack the capacity of systematic and global identification of TIS, especially for non-AUG sites in plants. Additionally, these methods are often unsuitable for evaluating the importance of mRNA sequence features for TIS determination. In this study, we present TISCalling, a robust framework that combines machine learning (ML) models and statistical analysis to identify and rank novel TISs across eukaryotes. TISCalling generalized and ranks important features common to multiple plant and mammalian species while identifying kingdom-specific features such as mRNA secondary structures and “G”-nucleotide contents. Furthermore, TISCalling achieved high predictive power for identifying novel viral TISs. Importantly, TISCalling provides prediction scores for putative TIS along plant transcripts, enabling prioritization of those of interest for further validation. We offer TISCalling as a command-line-based package [https://github.com/yenmr/TISCalling], capable of generating prediction models and identifying key sequence features. Additionally, we provide web tools [https://predict.southerngenomics.org/TISCalling/] for visualizing pre-computed potential TISs, making it accessible to users without programming experience. The TISCalling framework offers a sequence-aware and interpretable approach for decoding genome sequences and exploring functional proteins in plants and viruses.

Supplementary Information

The online version contains supplementary material available at 10.1007/s11103-025-01632-3.

Keywords: Translational control, Translation initiation site, Open-reading frames, Gene annotation, Machine learning prediction

Key message

TISCalling offers a computational framework that integrates machine learning and sequence analysis to enable de novo prediction of translational initiation sites (TISs), independent of ribosome profiling (Ribo-seq) datasets.

Supplementary Information

The online version contains supplementary material available at 10.1007/s11103-025-01632-3.

Introduction

The translation of mRNA into protein is at the heart of gene expression. Ribosomes find the correct translation initiation sites (TISs), which determine the protein-coding potential of mRNA and control timely and accurate protein production in response to developmental and environmental cues. Current annotation based on in silico prediction was biased to genes that canonically initiate from AUG sites and encode large proteins with known functional domains (Yandell & Ence 2012; Kearse & Wilusz 2017; Hsu & Benfey 2018). Emerging evidence highlights the prevalence of non-canonical translational events, including those from upstream open reading frames (uORFs), translated regions on non-coding RNAs, and features of translation initiation sites from non-AUG codons in plants and plant viruses (von Arnim et al. 2014; Hsu & Benfey 2018; Fang & Liu 2023; Wu et al. 2024b). Thus, a reliable workflow for annotating TISs is essential for decoding genomes and identifying protein-coding genes.

Ribosome sequencing (Ribo-seq), a technique used to globally profile translating ribosome positions, offers in vivo evidence for identifying TISs and open reading frames (ORFs) across genomes (Ingolia et al. 2009, 2012). The application of translation inhibitors enhances the resolution of Ribo-seq by selectively enriching specific types of ribosomes: cycloheximide (CHX) stabilizes ribosomes during initiation and elongation whereas Lactimidomycin (LTM) predominantly stalls ribosomes around initiation sites (Schneider-Poetsch et al. 2010). This specificity makes LTM particularly effective for identifying in vivo TISs and inferring ORFs with higher resolution (Schneider-Poetsch et al. 2010; Lee et al. 2012; Li & Liu 2020). Bioinformatics tools such as RiboTaper and Count, in-frame Percentage and Site (CiPS) utilize ribosome phasing patterns and the magnitudes of ribosome occupancy in CHX-treated Ribo-seq signals to identify AUG TISs and their corresponding ORFs (Calviello et al. 2016; Wu et al. 2019, 2024a). Additionally, TIS hunter (Ribo-TISH) employs LTM- and/or CHX-treated Ribo-seq datasets to detect both AUG and non-AUG initiation sites and their associated ORFs (Zhang et al. 2017; Wu et al. 2024). While these tools provide valuable insights, they rely entirely on Ribo-seq datasets with two focusing specifically on AUG codons for TIS screening. This dependency may limit their ability to comprehensively capture the translational events across genomes. Additionally, Ribo-seq and proteomics/peptidomics resources remain relatively scarce compared to the extensive resources for RNA-seq and protein-DNA interaction sequencing datasets. Therefore, a bioinformatics tool independent of Ribo-seq datasets would serve as a general and comprehensive means of identifying TISs and ORFs across genomes.

Ribo-seq-independent methods for identifying novel AUG and non-AUG TISs rely on either sequence conservation or prediction models. For the former, the comparative-genomic approach has identified dozens of AUG- and non-AUG TISs and their corresponding ORFs across eudicot plant species, primarily located in the 5’ untranslated regions (5’UTRs) of genes (van der Horst et al. 2019). However, the reliance on conservation scores and statistical power limits its ability to identify short and non-conserved TIS-initiated ORFs (Couso 2015). Alternatively, PreTIS, a linear-regression-based prediction model, utilizes mRNA sequence as the sole input to profile potential AUG and non-AUG TISs in 5’UTRs of human and mouse genes (referred to as the de novo prediction of the TISs) (Supplementary Table S1) (Reuter et al. 2016). While promising, PreTIS was developed using previously identified AUG and non-AUG TIS in the 5’UTRs of human and mouse genes, leaving its applicability to plants and its effectiveness at profiling TISs within CDSs uncertain. Our prior study focusing on plant AUG- and non-AUG TISs in 5’UTRs and CDSs has successfully generated an ML pipeline to identify TIS-prediction models and key sequences associating with translation initiation mechanisms in Arabidopsis and tomato (Supplementary Table S1) (Wu, TY et al., 2024). Nevertheless, the generated models have not been used for de novo TIS prediction in plant and plant virus species. Additionally, no web tool was provided for visualizing the de novo predicted TISs, which would have made them accessible to users without programming experience. Furthermore, a command-line package for building customized TIS prediction models for specific datasets and species of interest was also unavailable (Supplementary Table S1). Together, the available Ribo-seq-dependent and independent approaches are constrained by their individual limitations and unable to deliver a comprehensive and systematic profile of AUG- and nonAUG-initiated translational events across different genic regions—5’ and 3’UTRs and within CDS regions—or across entire plant transcriptome. Developing more integrative and versatile tools remains a critical challenge for fully decoding translational landscapes.

Here, we introduce TISCalling, a sequence-aware, ML-based framework for TIS predictor that enables the identification of key mRNA sequences regulating TIS recognition and the systemic profiling of potential TISs along transcripts (Fig. 1). Using publicly available datasets of the in vivo TISs identified through Ribo-seq (Lee et al. 2012; Willems et al. 2017; Li & Liu 2020), we developed predictive models to identify both AUG and non-AUG TISs in plants (e.g. Arabidopsis and tomato) and mammals (e.g. human and mouse). With the TIS-predictive model generated from each species, we retrieved the feature weights of the input features, reflecting their contribution and importance to model performance, and for revealing TIS recognition mechanisms across species. In addition, the model with the best TIS-prediction performance was applied to perform de novo TIS prediction analyses (i.e. computing TIS prediction scores using the mRNA sequences as the only input) to profile potential TIS sites with high prediction scores in UTRs of plant stress-related genes, non-coding RNAs, and even viral genomes. To ensure accessibility and flexibility, we provided the TISCalling framework as a command-line-based tool on Github [https://github.com/yenmr/TISCalling], enabling users to identify mRNA features and generate prediction models and outcomes specific to their datasets and species of interest. Additionally, by integrating TIS models into a user-friendly web tool [https://predict.southerngenomics.org/TISCalling], the TISCalling facilitates the visualization of TISs along genes. By providing a sequence-aware approach independent of Ribo-seq datasets, TISCalling supports the discovery of AUG and nonAUG TISs and small ORFs, contributing to the decoding of plant and viral genomes.

Fig. 1.

Fig. 1

The TISCalling framework, along with its features and applications, is designed for identifying potential TISs in plants and viruses. A The TISCalling framework included three parts: the collection of Ribo-seq supported translation initiation sites (TISs) and genome information (blue box), the generation of sequence features and prediction models and the addition of the built-in models (i.e., the integration of the pre-trained model from the model generation step into the TIS prediction step) (green box), and the identification of the key features important for model performance and the use of built-in models for reporting TIS prediction scores (yellow box). B Illustration of TISCalling model construction. The 80% of the collected true-positive (TP) and true-negative (TN) TIS datasets (gray cylinders) are used as training datasets for sequence feature collection and model training. The 20% TP and TN datasets are used to evaluate the performance of the trained models. (More details in Supplementary Fig. S2A and Methods.) C TISCalling includes a function of identifying sequence features associated with TIS predictions. The feature weights assigned during model training can be retrieved from a selected TIS prediction model (i.e., one selected from previously trained TISCalling-generated models B) and are ranked to infer the importance of sequence features for TIS predictions. D TISCalling includes a function for de novo TIS prediction along sequences. The pre-trained model (i.e., a training model from the model generation step B) computes the prediction scores for the target TIS site using only the mRNA sequences flanked by the site of interest as input. E Model recommendations for TIS identification. The flowchart illustrates the recommended predictive models with/without target TIS types for identifying plant and viral TISs. (More details of A-D in Supplementary Fig. S2A, B and Methods.)

Materials and methods

Dataset collection

We collected datasets of novel (i.e. unannotated) translation initiation sites (TISs) with significant translation initiation activity, referred to as true positive (TP) datasets, from tomato and Arabidopsis LTM-treated ribosome profiling data as reported by Li and Liu (2020). The TP TIS datasets for human HEK293 cells and mouse MEF cells were retrieved from the previous study of Lee et al. (2012). Additionally, we gathered the publicly available TP TIS datasets from various plant and virus studies (See Supplementary Dataset S1 for full datasets). Briefly, for Arabidopsis, novel TIS data included those reported by (Wu, TY et al., 2024) and those associated with non-coding open-reading frames (ORFs), downstream ORFs, upstream ORFs (uORFs), and within coding regions (CDSs) as described in (Wu, HL et al., 2024a). In tomato, novel TISs from uORFs and small ORFs in de novo assembled transcripts were retrieved from Wu et al. (2019). For human and plant viruses, novel TIS datasets were sourced from cytomegalovirus (HCMV) (Stern-Ginossar et al. 2012), SARS-CoV-2 (Finkel et al. 2021), and Tomato yellow leaf curl Thailand virus (TYLCTHV; genus Begomovirus; Chiu et al. (2022).

To construct a dataset of true negative (TN) TISs, we collected both ATG and near-cognate codon sites for each positive TIS in our dataset. These sites were located upstream of the most downstream TP TIS within the same transcript and were not marked as TP TISs, as described by (Wu, TY et al., 2024). This methodology generated robust TP and TN datasets, enabling an accurate assessment of our model’s ability to distinguish between true translation initiation sites and potential false positives.

Feature collection

We focused on mRNA sequence information around the TP and TN TISs and extracted 1,240 features for each TIS, which were divided into three categories: 5 known features, such as the Kozak sequence, TIS codon usage, and adjacent flanking sequences (Kozak 1984; Kozak 1989; Noderer et al. 2014; Reuter et al. 2016; Zhang, S et al., 2017; Diaz de Arce et al. 2018; Li & Liu 2020); 18 ORF-related features, including mononucleotide contents, secondary RNA structures within or upstream of ORFs, and ORF sizes; and 1,217 contextual features detailing the nucleotide/amino acid frequency of k-mers within a 201-nt region centered on a TIS. Feature collection mainly followed the classifications and methodologies outlined in Wu, TY et al. (2024), incorporating slight modifications to existing frameworks of known, contextual, ORF and TIS codon usage features. Each category is described below in detail:

Position-weight matrix (PWM)

PWM-related features were derived to represent the relationship between the flanking sequence context and TIS translational efficiency. Features such as "PWM" were based on a PWM matrix generated from the flanking sequences (positions -15 to + 10) of a given TP group with in vivo translation initiation activity.

Noderer translational efficiency

The feature was derived from (Noderer et al. 2014), which analyzed the flanking sequence context from positions -6 to + 5 around the AUG translational start.

Kozak sequence context

The Kozak sequence context was categorized into two levels: strong Kozak (A or G at -3 and G at + 4), and no Kozak context.

TIS-codon usage

The TIS codon usage was determined as previously described with minor modifications (Zhang et al. 2017). Briefly, for a given codon, the proportion of target codon sites among all identified TP TISs was normalized to the proportion of target codon sites among all codon sites in the transcript regions of all annotated genes. The corresponding log₂ ratio was then computed and referred to as the feature value of TIS codon usage bias. Codons with negative values were excluded.

ORF features

These features included the A/T/C/G mononucleotide content in the upstream regions and within the ORF regions of TP/TN TIS-initiated ORFs, as well as their ORF lengths. Additionally, features included the distance from TP/TN TISs to the upstream region of the nearest start and stop codons.

Minimum free energy (MFE) of mRNA secondary structure

MFEs were calculated using the RNAfold program (Lorenz et al. 2011) for 80-bp regions centered on TISs, with a sliding window of 20 nt and a step size of 10 nt. Furthermore, the magnitudes of MFE differences were normalized across nearby regions to highlight the structural dynamics around the TISs.

Contextual features

We quantified the frequency of all possible k-mers within a 201-nt window centered on the TIS site, with k = 1 for position-specific k-mers and k = 3 for codon and respective amino acid k-mers (see illustration of feature definition and information extraction in Supplementary Fig. S1). Specifically, the sequence regions examined include the full window (-99 to + 102 nt relative to the TIS), the upstream region (-99 to -1 nt upstream of the TIS), the downstream region (+ 4 to + 102 nt downstream of the TIS), the upstream in-frame region (-99 to -1 nt upstream and in-frame to TIS), the downstream in-frame region (+ 4 to + 102 nt downstream and in-frame to TIS). There are four groups of nucleotide/amino acid features: 780 features of “nucleotide position” group (A/T/C/G identify in a specified position in the full-window region), 12 features of “nucleotide frequency” group (A/T/C/G frequency in the full-window, upstream and downstream regions), 320 features of “codon frequency” group (64 codons among 5 regions mentioned above) and 105 features of “amino acid frequency” group (20 amino acids and 1 stop-codon among 5 regions mentioned above). This contributed to a total of 1,217 contextual features.

Feature selection for model training and evaluation

In addition to TIS dataset collection and feature generation (as described above), the framework of building the TIS-prediction models involved the steps of feature selection, model training, and model evaluation (Supplementary Fig. S2). This pipeline was adapted from previous studies with minor modifications (Reuter et al. 2016; Wu, TY et al., 2024). The TP and TN TIS datasets were first split into training (80%) and testing (20%) subdatasets separately, repeating 10 times to generate 10 sets of TP and TN TIS sub-datasets. The training subdatasets were applied for feature selection and model generation; the testing subdatasets were set aside for model evaluation. For each TP and TN subdatasets, we balanced the number of TPs and TNs through random sampling without replacement. Note that the TISs with less than 100 bp of flanking sequences in either direction were excluded from the TP and TN TIS datasets to avoid the noise in the feature selection step.

Feature selection

For the training subdatasets, we performed feature selection, considering research indicating that an excess of features can degrade model performance (Bzdok & Meyer-Lindenberg 2018). We applied the Wilcoxon rank-sum test with Bonferroni correction to identify significant differences between TP and TN TIS groups. We then calculated Pearson correlation coefficients (r) among contextual features and retained the 50 most significant (smallest adjusted p-values) and uncorrelated (r < 0.7) contextual features, along with non-contextual features with adjusted p-values below 0.05, for model training.

Feature values were then scaled using the “MinMaxScaler” from the scikit-learn package, rescaling them to a range of 0 to 1. This scaling ensured that all features contributed equally to the model, preventing those with larger numerical ranges from dominating others. The scaling factor of each feature produced here was applied during the processes of model evaluation and model-computed TIS prediction scores (Fig. 1; Supplementary Fig. S2A,B).

To assess feature importance for TP and TN TISs, we calculated Cohen’s d, which quantifies the effect size of the difference between two group means relative to their standard deviations. In this context, a higher Cohen’s d value indicates stronger discriminative power between TPs and TNs.

Model training

An ML pipeline was described previously (Uygun et al. 2019; Wu et al. 2021). Briefly, we used scikit-learn version 1.5.0 in Python version 3.10.14 to train and evaluate the models. We constructed and trained ML models using two different algorithms to assess their efficacy in predicting TISs in the training segments: Logistic Regression (LR), and Support Vector Machine (SVM). To optimize these models, we employed a grid search for each algorithm to identify the optimal settings, coupled with fivefold cross-validation. The hyperparameters tested were as follows: LR: C = [0.01, 0.1, 1, 10, 100], intercept_scaling = [0.1, 0.5, 1, 2, 5, 10], penalty = ["l2"]. SVMlinear: kernel = ["linear"], C = [0.01, 0.1, 1, 10, 100].

Model evaluation

In the testing and evaluation phase, features from the balanced testing subdatasets were scaled. The performance of trained models was assessed via the model metrics whereas the model with the highest F1 score was considered as the best model. To assess model robustness and feature significance, we applied this ML workflow across ten randomly balanced TP and TN subdatasets in the testing segments.

De novo TIS prediction

The input for de novo TIS prediction (Fig. 1D; Supplementary Fig. S2B) consisted of mRNA sequences in FASTA format, providing sequence information for potential TIS sites at AUG and non-AUG codons (i.e., near-cognate codons) along with their flanking regions. For a given triplet of interest, 1,240 features were generated via customized scripts and scaled to provide the specific feature information required by the selected prediction model (pre-trained model) to generate TIS prediction scores. The chosen pre-trained model is the training model with the best performance (i.e., the model with its performance from the model evaluation step and the information of the model algorithm and parameters and the obtained feature weights from the model generation step; Supplementary Fig. S2A). The output included the predicted probability along with detailed site information, such as gene ID, position, and the sequence surrounding each potential site (See Supplementary Fig. S2B for a workflow).

Note that there are the TISs with the annotated regions upstream or downstream of itself less than 100 nts (Supplementary Table S2), which cause the missing values in feature extraction steps and block the model-computing TIS prediction scores. Thus, the upstream or downstream regions less than 100 bp were padded with ‘N’ bases for score computation. All the features were generated via customized scripts provided in the TISCalling package (https://github.com/yenmr/TISCalling).

The feature generation/scaling, model training/evaluation and computing TIS prediction scores were conducted via customized scripts and the scikit-learn version 1.5.0 in Python version 3.10.14 provided in the TISCalling package (https://github.com/yenmr/TISCalling).

Construction of JBrowse and visualization

Visualization

JBrowse 2 was employed to display TIS prediction scores along genes across multiple species, including the genomes of Arabidopsis thaliana, and Solanum lycopersicum (tomato), (Diesh et al. 2023). It was also used to visualize genome-encoded protein sequences and gene annotations. Publicly available Ribo-seq datasets with LTM treatment from Arabidopsis and tomato, as described previously (Willems et al. 2017; Li & Liu 2020; Wu, TY et al., 2024), and the Ribo-seq datasets with CHX treatment from heat-treated pollens (Poidevin et al. 2021) were provided.

Data preparation

The TIS prediction scores for all triplets along the sequences of each gene were computed (referred to the “De novo TIS prediction” part), and the coordinates of each predicted TIS triplet were mapped to their respective genomes and saved in BigWig format to enable efficient visualization. Aligned Ribo-seq and RNA-seq reads (i.e., bam/bedgraph files) from the mapping process were converted to BedGraph files. To standardize read depth across samples, coverage profiles were normalized to reads per million (RPM). BigWig files were subsequently generated from BedGraph files using the bedGraphToBigWig tool from the UCSC binary utilities (Kent et al. 2010).

JBrowse configuration

A JSON configuration file was generated to facilitate the display of all data tracks in JBrowse. This configuration included sequence files, annotation files, RNA-seq and Ribo-seq abundance tracks, and tracks for TIS prediction scores. To streamline visualization, tracks from the positive and negative strands of genes were merged. The AUG TIS-prediction model reports AUG triplets with scores > 0.5, whereas the non-AUG TIS-prediction model reports triplets at either AUG or near-cognate codons with scores > 0.8.

An user manual with the step-by-step instructions of using TISCalling web tool to visualize the pre-computed potential TISs along genes is provided [https://predict.southerngenomics.org/TISCalling/TISCalling_Manual.pdf].

Detection of TIS activity in vivo

The generation of expression constructs with the target-TIS driven reporter genes, the introduction of protein expression constructs into tobacco using an Agrobacterium-mediated transient expression system, and the detection of protein expression were performed as described previously (Li & Liu 2020), with minor modifications in the design of expression constructs. Specifically, for a given target TIS site of the non-coding RNAs, the regions ranging from the 5’end of the transcript to the 3’end of the target TIS-encoded CDS regions were amplified from Arabidopsis cDNAs by PCR and fused in-frame with a reporter gene encoding GFP, which was used previously (Li & Liu 2020; Chiu et al. 2022). The primers used are listed in Supplementary Table S3.

Results

TISCalling framework: machine learning models for identifying key mRNA features and predicting potential TISs across species

To develop a machine-learning (ML) model for predicting TISs across eukaryotic species, we utilized publicly available datasets of in vivo AUG and non-AUG TISs from Arabidopsis (At), tomato (Sl), human (Hs) and mouse (Mm). These datasets, derived from LTM-seq (Lee et al. 2012; Willems et al. 2017; Li & Liu 2020), employed ribosome profiling with LTM treatment, which enriched ribosome stalled at translation initiation sites during the first round of translation, (Schneider-Poetsch et al. 2010) (Supplementary Dataset S1). The TISs identified across 5’UTRs and CDS regions were classified into four groups: 5’UTR-AUG, 5’UTR-nonAUG, CDS-AUG, and CDS-nonAUG (Wu, TY et al., 2024). Sequence information from the 201-bp regions flanking each TISs was extracted to build predictive models using logistic regression (LR) and support vector machine (SVM) algorithms (Figs. 1,2A; Supplementary Fig. S2; see Methods). To train a model and assess the model performance, 80% of TP and TN datasets were applied for feature collections and model training processes, whereas 20% of datasets were set aside for evaluating the model performance and selection of the best model (Fig. 1B, Supplementary Fig. S2A; see Methods). The prediction models for 5’UTR-AUG, 5’UTR-nonAUG, and CDS-AUG TISs in Arabidopsis and tomato achieved high performances, with AUROC > 0.90. However, the CDS-nonAUG models showed reduced performance, with AUROC > 0.57 (Fig. 2A; Supplementary Dataset S2 for additional performance metrics). Consequently, we focused on the high-performing models of 5’UTR-AUG, 5’UTR-nonAUG, CDS-AUG TISs in the subsequent analyses. We observed comparable results for human and mouse models (Fig. 2A). When testing cross-species applicability, models trained on Arabidopsis performed best on predicting tomato TISs, followed by human and mouse TISs (Fig. 2B, At). Similarly, tomato-derived models performed well on other plant species (Fig. 2B, Sl). Human and mouse models showed high cross-predictive accuracy with AUROC > 0.73 when applied to each other and > 0.68 for Arabidopsis and tomato TIS (overall mean, Fig. 2B, Hs and Mm). These findings indicate that TIS models are broadly applicable across species, with within-kingdom predictions outperforming cross-kingdom predictions (Supplementary Dataset S2 for full results).

Fig. 2.

Fig. 2

The TISCalling framework-generated prediction models for both AUG and non-AUG TISs across plants and mammals. A The performance shown in AUROC scores of the logistic regression (LR) and support-vector-machine (SVM) algorisms in predicting TISs identified from Arabidopsis (At), tomato (Sl), human (Hs) and mouse (Mm). The identified TISs were categorized into four groups based on their locations (i.e., within 5’UTR and main CDS) and the codons of the TIS sites (i.e., AUG and non-AUG codons). B Cross-species prediction performance, using models trained on one species to predict TISs in another. The model performance was shown for the four TIS categories as described in panel A. C Cross-TIS prediction performance, using a model trained on one TIS group to predict another TIS group within the same species, as indicated in panel A. D Input features ranked by importance scores derived from the models predicting four types of TISs from different species. Higher ranks indicate greater importance, with features colored according to the difference in mean feature values between true positives (TPs) and true negatives (TNs), as represented by Cohen’s d effect size. Higher effect sizes mean that the mean feature values of TPs are greater than those of TNs

One innovative feature of the TISCalling workflow is its ability to rank and locate key sequence features, likely associating with TIS initiation mechanisms, for the species of interest (Fig. 1A,C). To demonstrate its efficacy, we analyzed the model with the best TIS prediction performance to retrieve the feature weights of the input features, reflecting their contribution and importance to model performance. Several top-ranked features were found in all 5’UTR-AUG, 5’UTR-nonAUG, CDS-AUG TIS models and shared across four species of Arabidopsis, tomato, human and mouse (Figs. 1A,C). These include the features of PWM (i.e., the TIS-proximal sequences spanning positions -15 to + 10), CU-related sequences, and the avoidance of upstream AUG sites and short AUA sequences near a TIS (features marked in red in Fig. 2D, see Methods for feature collections and Supplementary Dataset S3 for detailed feature information). Interestingly, the feature of “C” at the -12 position was top-ranked for mammalian, but not for plant, TIS prediction models and showed a significant depletion pattern in TIS TPs (large circle and blue Cohen’s d value; features marked in blue in Fig. 2D). These observations indicated that the “C” at the -12 position is an important and specific-feature for mammalian TIS prediction, consistent with previous findings (Reuter et al. 2016). In addition, the “G”-nucleotide content in the TIS-upstream regions was top-ranked in both plant and mammal TIS models (features marked in blue in Fig. 2D) but showed opposite enrichment patterns (blue and red Cohen’s d values in Fig. 2D), suggesting distinct sequence preferences between eukaryotic species. These observations highlighted the presence of common sequence features shared across mammals and plants and also the mammal- and plant-specific ones, suggesting the shared and lineage-specific TIS recognition mechanisms. In line with the previous observation of the moderate prediction performance of cross-species TIS predictions between mammals and plants (Fig. 2A), these results indicate both conserved and divergent sequence features governing translation initiation across kingdoms.

To investigate the generality of TIS recognition mechanisms across four TIS types, we assessed cross-prediction performance using models derived from one TIS type to predict the remaining three types. Overall, cross-TIS type prediction achieved AUROC values > 0.7 (mean, Fig. 2C) but performed less accurately compared to within-TIS type prediction (Fig. 2A). For instance, the 5’UTR-nonAUG model from Arabidopsis achieved ~ 0.94 for predicting 5’UTR-nonAUG TISs in tomato but only ~ 0.85 when predicting Arabidopsis 5’UTR-AUG TISs. These results suggest a combination of common and specific sequence features associated with different TIS types. For example, the avoidance of upstream AUG sites was observed in both 5’UTR-AUG and 5’UTR-nonAUG TIS groups, Kozak-related features were more prominent in 5’UTR nonAUG TISs (Fig. 2D).

Together, the TISCalling pipeline successfully generated predictive models for AUG and non-AUG TISs in 5’UTRs and CDS regions within and across species. In addition to robust prediction, it revealed conserved and species-specific sequence features underlying TIS recognition, offering novel insights into translational regulation in plants and animals.

TISCalling framework robustly profiles potential TISs in plant transcripts

Our TIS prediction model was developed using TISs experimentally supported from LTM-treated Ribo-seq datasets of Arabidopsis suspension cells and tomato leaves and statistically identified via TIS/ORF-identification methods (Supplementary Dataset S1) (Lee et al. 2012; Willems et al. 2017; Machkovech et al. 2019; Li & Liu 2020). To evaluate model robustness across datasets generated using different treatments, we incorporated additional TIS datasets derived from both LTM- and CHX-treated Ribo-seq experiments and diverse identification pipelines (Supplementary Dataset S1) (Wu et al. 2019; Li & Liu 2020; Hiragori et al. 2023). LTM preferentially stalls ribosomes during the first round of translation, providing high-resolution Ribo-seq signals around initiation sites, while cycloheximide (CHX) globally blocks elongating ribosomes across transcripts. Our Arabidopsis 5’UTR-AUG models achieve AUROC > 0.69 when predicting TISs identified using LTM-seq from Arabidopsis seedlings and CHX-seq data from Arabidopsis and tomato seedlings and through RiboTaper and CiPS pipelines (Fig. 3E, Known TIS types). Similar results were observed for 5’UTR-nonAUG and CDS-AUG models (Fig. 3E, Known TIS types). Strong correlations in prediction performance between Arabidopsis and tomato models (rho = 0.95, Fig. 3E, Known TIS types) further supported the cross-species applicability of TISCalling models in plants (Fig. 2B). Moreover, the models successfully identified AUG-TISs within 3’UTRs as well as in non-coding and de novo-assembled transcripts, demonstrating their versatility across diverse transcript types (Fig. 3E, Others; see Supplementary Dataset S2 for full results).

Fig. 3.

Fig. 3

TIS prediction models using flanking mRNA sequences robustly identify plant TISs. A The prediction scores of a given triplet along the Solyc02g023990 transcript were generated via the plant CDS-AUG-TIS model and were based on the sequences of the triplet and its flanking 200-bp regions. The AUG triple sites with a prediction score > 0.8 along the wild-type (WT) mRNA sequences were shown, including the annotated AUG TIS (aTIS) and a downstream in-frame AUG TIS (dTIS) encoding a mitochondria-localized protein. mt: the AUG dTIS was mutated to CUU triple in the mutant. The LTM plots showed the read density (reads per million mapped reads; RPM) derived from tomato LTM-treated Ribo-seq datasets. In the gene models (top), light and dark gray boxes indicate UTRs and annotated CDSs, while orange boxes indicate dTIS-initiated ORF, which encodes a protein isoform predicted to be localized to mitochondria. B Localization of Solyc02g023990-GFP protein with translation driven by the wild-type CDS (left panel) and dTIS-mutated CDS (middle panels). The aTIS, dTISs, and mutated TISs are indicated by a black arrow, orange arrow, and blue cross, respectively. Mitochondria marker: CD3-992 (Nelson et al. 2007). The results were derived from a previous study (Li & Liu 2020). C As in A, but for the Solyc03g096920 transcript with an upstream AUG TIS (uTIS). mCU: the mutations at CU-rich sequences of the upstream 100-nt region of the uTIS. The results were derived from a previous study (Wu, TY et al., 2024). D Expression of the Solyc03g096920-GFP proteins initiated from the uTIS indicated in C, with translation driven by the upstream 100-nt WT or mCU sequences. Vector: tobacco leaves infiltrated with agrobacteria containing the expression vector (i.e., the GFP plasmid without a target gene sequence). E Prediction performance of the four TIS-type models (from tomato (Sl) and Arabidopsis (At), as indicated in Fig. 2A) in predicting TISs derived from previous studies and identified via different TIS/ORF identification pipelines including RiboTaper and CiPS and a TIS identification method (Machkovech et al. 2019; Wu et al. 2019; Wu, HL et al., 2024a; Wu, TY et al., 2024). These previously identified TISs were categorized into two groups: known TIS types including 5’UTR-AUG, 5’UTR-nonAUG, CDS-AUG and CDS-nonAUG TISs (left panel) and others including the AUG TISs in 3’UTRs, non-coding RNAs and de novo-assembled transcripts (right panel). Spearman’s rank correlation coefficient (rho) and the corresponding p-values were shown

The other innovative feature of the TISCalling is its ability to perform de novo TIS prediction (i.e., using the pre-trained models with mRNA sequences as the sole input, independent of Ribo-seq datasets, to profile all potential TISs along sequences; more details in Fig. 1D, Supplementary Fig. S2B and Methods). Focusing on Solyc02g023990, which contains a novel CDS-AUG TIS driving the translation of a mitochondria-localized protein (Fig. 3B) (Li & Liu 2020), the model identified an AUG triplet within the CDS with a prediction score > 0.93 (orange peak in “WT” plot in Fig. 3A). This prediction aligns with experimental evidence showing higher LTM read signals at the same site (“LTM” plot in Fig. 3A). When the CDS-AUG TIS was mutated to CUU, the prediction score dropped, corresponding to the abolished protein translation and diminished mitochondrial localization signals (“mt” in Fig. 3A and 3B). For Solyc03g096920, the model identified two potential TISs in the 5’UTR with prediction scores > 0.96, one of which matched the experimentally identified 5’UTR-AUG TIS (orange peak in “WT” in Fig. 3C) (Wu, TY et al., 2024). Mutating the CU-rich sequence upstream of this 5’UTR-AUG TIS led to reduced prediction scores (“mCU” in Fig. 3C), consistent with the importance of CU-rich sequence features for TIS prediction (Fig. 2D) and reduced protein production (Fig. 3D) (Wu, TY et al., 2024).

Together, these results demonstrated the robustness of the prediction models in identifying TISs across datasets, species, and even at the single-gene level. By analyzing sequence features, the TISCalling framework provides a powerful sequence-aware tool for identifying potential AUG and non-AUG TISs.

TISCalling framework effectively profiles potential TISs in viral transcripts

Viruses rely on host translational machinery for protein synthesis, intriguing us to investigate whether TIS prediction models developed for hosts could be applied to viral genomes. Indeed, the TIS prediction models derived from Arabidopsis successfully predicted TISs reported in begomoviruses (Chiu et al. 2022) (Fig. 4A, At and Sl), despite slightly lower prediction performance compared to that in the host (Fig. 2). Similarly, human and mouse TIS models, especially the 5’UTR-nonAUG and CDS-AUG models, predicted TISs reported in human cytomegalovirus and SARS-CoV-2 genomes (Stern-Ginossar et al. 2012; Finkel et al. 2021)(right two heatmaps in Fig. 4A; Supplementary Dataset S2 for full results). These results highlight the versatility of the TISCalling framework in profiling TISs in the viral genomes using models trained on host TISs.

Fig. 4.

Fig. 4

Viral TISs can be predicted using host TIS predictive models. A The prediction performance (AUROC scores) of using the four TIS type-generated prediction models from tomato (Sl), Arabidopsis (At), human (Hs), and mouse (Mm) (as indicated in Fig. 2A) to predict the TISs along viral genomes. The viral TISs were derived from different public datasets of plant begomoviruses, human cytomegalovirus and SARS-CoV-2 viruses (Stern-Ginossar et al. 2012; Finkel et al. 2021; Chiu et al. 2022) (See Supplementary Dataset S1 for the TIS dataset information). B Expression of AV2-FLAG proteins during virus infections with a plant begomovirus infectious clone, containing a FLAG protein insertion at the C-terminus of AV2. Results derived from a previous study (Chiu et al. 2022) were shown for begomovirus infectious clone with WT sequences or specific mutations at nucleotide 135 (M135; M135AUG- > GCG) and/or 189 (M189AUG- > GCG). In the gene models (top), the black line indicates viral genomic DNAs, boxes represent annotated CDS of viral AV1 and AV2 genes, and the inserted FLAG region. The annotated TIS at position 87 and novel TISs at positions 135 and 189 of the AV2 gene were highlighted. *: WT and M189 samples were diluted 20-fold to equalize the protein abundance for presentation clarity. C-F As shown in Fig. 3A, but for the read density at LTM (C) and mRNA (D) levels along viral genomes. The Arabidopsis CDS-AUG TIS model-generated prediction scores are based on WT (E) and M135/M189 mutant (F) sequences. G As described in B, but showing FLAG-tagged BV2 proteins (orange arrows) expressed during virus infection. In the gene model (top), the BV1 TIS at position 174 and two novel AUG TISs of BV2 at positions 208 and 229 were highlighted. Results were shown for infectious clones with the WT sequence or the specific mutations at positions 208 (M208; M208AUG- > GCG) and/or 229 (M229; M229AUG- > GCG). H-J As shown in C-F, but presenting the results along the BV2-flanking regions of the viral genomes. The results of immunoblotting and LTM plots were derived from a previous study (Chiu et al. 2022)

At the single-gene level, previous studies identified novel TIS activities at positions 135 and 189 of AV2 genes and at positions 208 and 229 of BV2 genes on begomovirus genomes through LTM-seq analyses (black arrows in Fig. 4C,H) (Chiu et al. 2022). Consistent with these findings, the TISCalling framework assigned relatively high prediction scores to most of these TIS sites (0.58, 0.41, 0.15, 0.34 for each site, respectively) (orange peaks in Fig. 4E,J). Mutating these TIS sites (AUG → GCG) in AV2 and BV2 genes resulted in lower prediction scores (red peaks in Fig. 4E,F,J,K, Supplementary Dataset S4 for full results), and correlated with reduced protein expression due to the loss of initiation activities at these sites (Fig. 4B,G). These findings demonstrate that the TISCalling framework, developed for host plants and mammals, can effectively profile potential TISs in viral transcripts, providing insights into translation regulation in viral genomes.

TISCalling framework identifies potential TISs in stress-responsive genes and ncRNAs

The characterization of upstream open reading frames (uORFs), key regulatory elements for the translation of stress-related genes (von Arnim et al. 2014; Urquidi Camacho et al. 2020; Zhang et al. 2020; Wu, TY et al., 2024), remains limited due to annotation bias and the lack of translational evidence (Couso 2015; Hsu & Benfey 2018). Using the function of de novo TIS prediction in the TISCalling framework (Fig. 1D), we identified a predicted AUG TIS in the 5’UTR of Arabidopsis Heat Shock Transcription Factor B1 (AtHSFB1; Supplementary Fig. S3F), corresponding to the TIS of the known uORF1 (Pajerowska-Mukhtar et al. 2012; Zhu et al. 2012). This TIS/uORF1 did not show significant CHX nor LTM signals in Arabidopsis suspension cells or heat stress (HS)-treated pollen cells (Supplementary Fig. S3D,E) (Willems et al. 2017; Poidevin et al. 2021), likely due to tissue-specific gene expressions. Further analysis of the TIS-prediction scores revealed three additional potential TISs: one at an AUG codon and two at AUA and CUG codons (orange peaks in Supplementary Fig. S3F). These sites displayed strong LTM and CHX signals but were previously unannotated (Supplementary Fig. S3D,E). Similarly, for other heat-responsive genes, Heat Shock Transcription Factor B2B (AtHSFB2B) and Heat Shock protein 70 (AtHSP70), the TISCalling models predicted multiple AUG and non-AUG TISs (Supplementary Fig. S3). Some of these sites exhibited weak LTM/CHX signals in Arabidopsis suspension cells and pollen cells (Willems et al. 2017; Poidevin et al. 2021), likely due to sequencing depth (Supplementary Fig. S3). The tomato orthologous of these HSF genes also exhibited potential TISs in their 5’UTRs (Supplementary Fig. S4), underscoring the role of these uORFs-associated TISs in regulating these heat stress responsive genes. Notably, the non-AUG TIS model identified several potential TIS sites supported by LTM/CHX signals (orange peaks in Supplementary Fig. S3). These findings highlight the utility of TISCalling in uncovering previously uncharacterized TISs and uORFs, expanding the horizon of translational regulation in stress responses.

Non-coding RNAs (ncRNAs), once classified as no coding potential, are now known to encode small peptides and proteins in plants (Sruthi et al. 2022; Wang et al. 2023). Using the TISCalling framework, we identified potential TISs with high prediction scores, suggesting their ability to initiate ORFs in ncRNAs (red boxes in Fig. 5A,B). Immunoblotting analysis showed protein signals of ORFs initiated from the TISCalling-profiled TISs on different non-coding RNAs (Fig. 5B). In addition, some of them were supported by LTM and CHX signals (Fig. 5) and previous findings (Hsu et al. 2016). The integration of TIS prediction scores with ribosome profiling datasets highlights the ability of TISCalling to systematically profile both annotated and uncharacterized AUG and non-AUG TIS/ORFs in ncRNAs. Together, these results demonstrate the versatility of TISCalling in accurately predicting TISs across plant heat-stress-responsive genes and ncRNAs. By incorporating LTM/CHX-based ribosome profiling datasets, the framework provides a powerful tool for uncovering hidden translational activities in protein-coding and non-coding RNAs.

Fig. 5.

Fig. 5

A Case study: the potential TISs in plant non-coding RNA genes using TISCalling A The screenshot illustrates using the web-based TISCalling tool to visualize the predicted AUG-TISs, the LTM signals (red box and the zoom in), and the corresponding ORFs along four non-coding RNA (ncRNA) (orange arrows). B The TISCalling-web screenshots illustrate three predicted TISs on ncRNA genes, as well as the immunoblot analyses using the GFP antibody to detect the protein expression from these predicted TIS-encoded ORFs, which are in-frame fused with the GFP gene (See details in Methods). Vector: tobacco leaves infiltrated with agrobacteria containing the expression vector (i.e. the plasmid without target gene sequences)

A web- and package-based TISCalling tool visualizing and predicting potential TISs and ORFs

To facilitate the explorations of TISs at both genome-wide and single-gene scales, we developed a web-based TISCalling tool [https://predict.southerngenomics.org/TISCalling/]. This user-friendly JBrowse-based platform integrates precomputed TIS prediction scores with publicly available LTM/CHX-based Ribo-seq datasets for Arabidopsis and tomato, covering both protein-coding and non-coding transcripts. To visualize gene models with corresponding TIS prediction scores and LTM/CHX signals, users only need to 1) input a target gene ID, 2) select the species, 3) choose TIS prediction models and Ribo-seq datasets. The tool supports the uploading of Ribo-seq datasets for comparison with those from additional experimental conditions. This interactive interface allows users to profile and visualize potential TISs and their corresponding ORFs for target genes of interest (Fig. 5; Supplementary Fig. S4).

To expand the application of TISCalling to diverse species, we also present the framework as a Python-based package, providing detailed documentation including the step-by-step instructions, all customized scripts and an easy-to-run command-line interface [https://github.com/yenmr/TISCalling]. This package enables the generation of custom prediction models using user-defined datasets and genomes, reporting feature weights of key sequence features, and applying the build training model to predict TISs across species.

Together, TISCalling, with its web tool and Python package, provides an efficient and robust approach for profiling predicted AUG and non-AUG TISs. The sequence-aware prediction capabilities, independent from Ribo-seq datasets, and flexibility and expansibility of integrating additional omics data make it a versatile tool for exploring coding regions at both gene and genome scales.

Discussion

The annotation of protein-coding genes often overlooks the translation events involving short coding regions or initiation sites from AUG and non-AUG codons, leaving important aspects of genomic sequences underexplored. To address this, we propose TISCalling, a flexible framework designed to systematically investigate potential TISs along transcripts (Fig. 1). This framework features sequence-aware machine learning (ML)-based TIS-prediction models capable of robustly predicting AUG and non-AUG TISs independently of Ribo-seq datasets while identifying key sequence features associated with TIS recognition mechanisms in plants and mammals (Fig. 1A,C,2). By analyzing mRNA sequence information, the TIS-prediction models generate prediction scores for AUG and non-AUG triplets, enabling the profiling of potential TISs in plant and viral transcripts and the identification of ORFs in protein-coding and non-coding RNAs (Figs. 1A,D,35). Beyond our previous study of employing ML approaches to reveal TIS initiation mechanisms (Wu, TY et al., 2024), the ML-based TISCalling framework is available as a user-friendly web tool and a command-line Python package, offering users to build custom TIS prediction models, analyze mRNA sequences across genomes, profile and visualize potential TISs at both single-gene and genome-wide scales, further supporting its the broad applicability (Fig. 1,5; Supplementary Table S1). While intended as an alternative and complementary approach to established TIS-identification methods in plants (Calviello et al. 2016; Reuter et al. 2016; Zhang et al. 2017; Machkovech et al. 2019; van der Horst et al. 2019; Wu, HL et al., 2024a), TISCalling aims to provide a sequence-aware platform for profiling potential TISs comprehensively. We believe it can offer novel insights into the mechanisms by which ribosomes recognize AUG and non-AUG TISs across species and potentially shed light on previously uncharacterized sequence features in plants and viruses.

Our TIS-prediction model, built on sequence information, computes a prediction score for potential TISs (Figs. 1,35). Independent of experimental datasets, this approach complements the Ribo-seq-supported TIS identification methods (i.e., those with significant signals and/or phasing patterns from Ribo-seq datasets) and comparative genomics analyses (Calviello et al. 2016; Zhang, P et al., 2017; van der Horst et al. 2019; Wu, HL et al., 2024a) and serves as a stand-alone tool for the research community when experimental datasets are unavailable. By analyzing the likelihood of each triplet to function as a TIS, our model enables the generation of the datasets for all possible AUG and non-AUG ORFs across genomes. Despite overall robustness of the model, a known limitation is the potential for false-positive predictions, particularly in non-AUG models (Fig. 2). Nevertheless, the TIS-prediction model in TISCalling framework is intended to generate comprehensive datasets of potential TISs/ORFs, which can then be refined by integrating Ribo-seq data. This integration improves sensitivity and resolution, enabling a more accurate assessment of translational activity and the biological functions of genomic and genic regions. To prioritize TIS candidates for downstream validation, the LTM/CHX signals from Ribo-seq analyses should be considered, enabling targeted functional analysis of promising TISs.

The performance of TIS-prediction models varies across different regions of transcripts, reflecting the complexity of translational regulation. The 5’UTR-AUG, 5’UTR-nonAUG and CDS-nonAUG models generally outperformed and non-AUG ones (Fig. 2A). This discrepancy may be attributed to the noise of the collected TP/TN TIS datasets of CDS-nonAUG. First, multiple translational regulations occur within CDSs, such as translation initiation/elongation and ribosome stalling (Merret et al. 2015; Benitez-Cantos et al. 2020; Collart & Weiss 2020; Iwakawa et al. 2021). The elongating ribosome stalls or slowly moves within CDSs can induce ribosome traffic, contributing to the signals detected by Ribo-seq, and increasing the noise in TP TIS datasets. Second, compared to AUG codons, non-AUG codons have lower TIS activity (Kearse & Wilusz 2017; Li & Liu 2020). Together with the features of the high rRNA contamination and lower sequencing read-depth of Ribo-seq (Wang & Mao 2023), the nonAUG TISs with in vivo TIS activity might be missed in TP datasets and grouped into TNs, leading to the high noise in the TN datasets. In addition, our model considers only 100-bp upstream of TISs, which is likely sufficient for scanning ribosomes recognizing TISs in 5’UTRs but may be less suitable for TISs within CDS, where ribosomes scan longer regions from the 5’end to the CDSs of the transcripts. To address these limitations, deeper sequencing read-depth in Ribo-seq and exploring the ribosome pausing sites via disome-seq (Wang & Mao 2023) would improve the sensitivity of identifying in vivo non-AUG TISs within CDS regions and facilitate the generation of CDS-nonAUG TIS-prediction models.

Similarly, predicting TISs in viral genomes presents unique challenges. For viral TISs, the highest prediction performance reached an AUROC of ~ 0.6, suggesting room for improvement. The current viral TIS predictions rely on models built from host data. However, viruses are known to utilize diverse translational strategies, such as internal ribosome entry site (IRES) and secondary structures, to initiate translation from unconventional sites, maximizing protein diversity of their compact genomes (Lozano & Martinez-Salas 2015; Miras et al. 2017; Jaafar & Kieft 2019). The collection of more in vivo viral TIS sites from diverse plant viruses, representing and covering the information for viral TISs, and the incorporation of these sequences and structural features, representing the viral translational strategies, will refine the development of plant viral TIS-prediction models in the future.

The TISCalling framework represents a sequence-aware tool, independent of Ribo-seq datasets, designed to profile both AUG and non-AUG initiation sites at the genome and single-gene levels. By integrating prediction models with visualization capabilities, TISCalling offers a versatile approach for exploring translational initiation mechanisms across diverse genomic regions. The framework is accessible through a command-line Python package on Github and a user-friendly JBrowse tool to visualize and profile potential TISs across plants and viruses with simplicity, flexibility, and expandability. This advancement opens new opportunities for identifying potential AUG- and non-AUG-initiated ORFs across genomes. Moreover, it enhances our ability to uncover the biological significance of previously unexplored genomic regions and offers insights into the mechanisms driving translation initiation in plant and viral species. With its flexibility to integrate additional omics datasets and expand to other species, TISCalling serves as a valuable resource for advancing genome annotation and understanding translational regulation in diverse contexts.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

We thank the core facilities of AS-BCST and IPMB Bioinformatics Core for high-performance computing services. We acknowledge Biorender (https://www.biorender.com) for creating Fig. 1A-D. This research was financially supported by grants of AS-CDA-111-L06 to Ming-Jung Liu, AS-CDA-113-L03 to Ting-Ying Wu, and NSTC 112-2313-B-002-018-MY3 to Chia-Yi Cheng.

Author contribution

M.-R. Y. performed the computational analyses and developed the command-line-based and web-based tools. Y.-R. Li performed the experimental analyses. C.-Y. C. developed the command-line-based and web-based tools and revised the paper. T.-Y. W. and M.-J. L. conceived/designed the research, interpreted the results, and wrote the paper.

Funding

Open access funding provided by Academia Sinica. Academia Sinica, AS-CDA-113-L03 to Ting-Ying Wu; AS-CDA-111-L06 to Ming-Jung Liu; NSTC 112-2313-B-002-018-MY3 to Chia-Yi Cheng

Data availability

The LTM- and CHX-based Ribo-seq in Arabidopsis (suspension cells), Arabidopsis pollens and tomato were retrieved from NCBI Gene Expression Omnibus (GEO; https://www.ncbi.nlm.nih.gov/geo/) under accession number GSE145795, GSE88790, GSE143311.

Declarations

Competing interests

The authors declare that they have no conflict of interest.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Ting-Ying Wu, Email: tingying@gate.sinica.edu.tw.

Ming-Jung Liu, Email: mjliu@as.edu.tw.

References

  1. Benitez-Cantos MS, Yordanova MM, O’Connor PBF, Zhdanov AV, Kovalchuk SI, Papkovsky DB, Andreev DE, Baranov PV (2020) Translation initiation downstream from annotated start codons in human mRNAs coevolves with the Kozak context. Genome Res 30(7):974–984 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Bzdok D, Meyer-Lindenberg A (2018) Machine Learning for Precision Psychiatry: Opportunities and Challenges. Biol Psychiatry Cogn Neurosci Neuroimaging 3(3):223–230 [DOI] [PubMed] [Google Scholar]
  3. Calviello L, Mukherjee N, Wyler E, Zauber H, Hirsekorn A, Selbach M, Landthaler M, Obermayer B, Ohler U (2016) Detecting actively translated open reading frames in ribosome profiling data. Nat Methods 13(2):165–170 [DOI] [PubMed] [Google Scholar]
  4. Chiu CW, Li YR, Lin CY, Yeh HH, Liu MJ (2022) Translation initiation landscape profiling reveals hidden open-reading frames required for the pathogenesis of tomato yellow leaf curl Thailand virus. Plant Cell 34(5):1804–1821 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Collart MA, Weiss B (2020) Ribosome pausing, a dangerous necessity for co-translational events. Nucleic Acids Res 48(3):1043–1055 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Couso JP (2015) Finding smORFs: getting closer. Genome Biol 16(1):189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Diaz de Arce AJ, Noderer WL, Wang CL (2018) Complete motif analysis of sequence requirements for translation initiation at non-AUG start codons. Nucleic Acids Res 46(2):985–994 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Diesh C, Stevens GJ, Xie P, De Jesus MT, Hershberg EA, Leung A, Guo E, Dider S, Zhang J, Bridge C et al (2023) JBrowse 2: a modular genome browser with views of synteny and structural variation. Genome Biol 24(1):74 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Fang JC, Liu MJ (2023) Translation initiation at AUG and non-AUG triplets in plants. Plant Sci 335:111822 [DOI] [PubMed] [Google Scholar]
  10. Finkel Y, Mizrahi O, Nachshon A, Weingarten-Gabbay S, Morgenstern D, Yahalom-Ronen Y, Tamir H, Achdout H, Stein D, Israeli O et al (2021) The coding capacity of SARS-CoV-2. Nature 589(7840):125–130 [DOI] [PubMed] [Google Scholar]
  11. Hiragori Y, Takahashi H, Karino T, Kaido A, Hayashi N, Sasaki S, Nakao K, Motomura T, Yamashita Y, Naito S et al (2023) Genome-wide identification of Arabidopsis non-AUG-initiated upstream ORFs with evolutionarily conserved regulatory sequences that control protein expression levels. Plant Mol Biol 111(1–2):37–55 [DOI] [PubMed] [Google Scholar]
  12. Hsu PY, Benfey PN (2018) Small but Mighty: Functional peptides encoded by small ORFs in plants. Proteomics 18(10):e1700038 [DOI] [PubMed] [Google Scholar]
  13. Hsu PY, Calviello L, Wu HL, Li FW, Rothfels CJ, Ohler U, Benfey PN (2016) Super-resolution ribosome profiling reveals unannotated translation events in Arabidopsis. Proc Natl Acad Sci U S A 113(45):E7126–E7135 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Ingolia NT, Ghaemmaghami S, Newman JR, Weissman JS (2009) Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science 324(5924):218–223 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Ingolia NT, Brar GA, Rouskin S, McGeachy AM, Weissman JS (2012) The ribosome profiling strategy for monitoring translation in vivo by deep sequencing of ribosome-protected mRNA fragments. Nat Protoc 7(8):1534–1550 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Iwakawa HO, Lam AYW, Mine A, Fujita T, Kiyokawa K, Yoshikawa M, Takeda A, Iwasaki S, Tomari Y (2021) Ribosome stalling caused by the Argonaute-microRNA-SGS3 complex regulates the production of secondary siRNAs in plants. Cell Rep 35(13):109300 [DOI] [PubMed] [Google Scholar]
  17. Jaafar ZA, Kieft JS (2019) Viral RNA structure-based strategies to manipulate translation. Nat Rev Microbiol 17(2):110–123 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Kearse MG, Wilusz JE (2017) Non-AUG translation: a new start for protein synthesis in eukaryotes. Genes Dev 31(17):1717–1731 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Kent WJ, Zweig AS, Barber G, Hinrichs AS, Karolchik D (2010) BigWig and BigBed: enabling browsing of large distributed datasets. Bioinformatics 26(17):2204–2207 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Kozak M (1984) Point mutations close to the AUG initiator codon affect the efficiency of translation of rat preproinsulin in vivo. Nature 308(5956):241–246 [DOI] [PubMed] [Google Scholar]
  21. Kozak M (1989) Context effects and inefficient initiation at non-AUG codons in eucaryotic cell-free translation systems. Mol Cell Biol 9(11):5073–5080 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Lee S, Liu B, Lee S, Huang SX, Shen B, Qian SB (2012) Global mapping of translation initiation sites in mammalian cells at single-nucleotide resolution. Proc Natl Acad Sci U S A 109(37):E2424-2432 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Li YR, Liu MJ (2020) Prevalence of alternative AUG and non-AUG translation initiators and their regulatory effects across plants. Genome Res 30(10):1418–1433 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Lorenz R, Bernhart SH, Honer Zu Siederdissen C, Tafer H, Flamm C, Stadler PF, Hofacker IL. 2011. ViennaRNA Package 2.0. Algorithms Mol Biol 6: 26. [DOI] [PMC free article] [PubMed]
  25. Lozano G, Martinez-Salas E (2015) Structural insights into viral IRES-dependent translation mechanisms. Curr Opin Virol 12:113–120 [DOI] [PubMed] [Google Scholar]
  26. Machkovech HM, Bloom JD, Subramaniam AR (2019) Comprehensive profiling of translation initiation in influenza virus infected cells. PLoS Pathog 15(1):e1007518 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Merret R, Nagarajan VK, Carpentier MC, Park S, Favory JJ, Descombin J, Picart C, Charng YY, Green PJ, Deragon JM et al (2015) Heat-induced ribosome pausing triggers mRNA co-translational decay in Arabidopsis thaliana. Nucleic Acids Res 43(8):4121–4132 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Miras M, Miller WA, Truniger V, Aranda MA (2017) Non-canonical translation in plant RNA viruses. Front Plant Sci 8:494 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Nelson BK, Cai X, Nebenfuhr A (2007) A multicolored set of in vivo organelle markers for co-localization studies in Arabidopsis and other plants. Plant J 51(6):1126–1136 [DOI] [PubMed] [Google Scholar]
  30. Noderer WL, Flockhart RJ, Bhaduri A, Diaz de Arce AJ, Zhang J, Khavari PA, Wang CL (2014) Quantitative analysis of mammalian translation initiation sites by FACS-seq. Mol Syst Biol 10(8):748 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Pajerowska-Mukhtar KM, Wang W, Tada Y, Oka N, Tucker CL, Fonseca JP, Dong XN (2012) The HSF-like transcription factor TBF1 is a major molecular switch for plant growth-to-defense transition. Curr Biol 22(2):103–112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Poidevin L, Forment J, Unal D, Ferrando A (2021) Transcriptome and translatome changes in germinated pollen under heat stress uncover roles of transporter genes involved in pollen tube growth. Plant Cell Environ 44(7):2167–2184 [DOI] [PubMed] [Google Scholar]
  33. Reuter K, Biehl A, Koch L, Helms V (2016) PreTIS: a tool to predict non-canonical 5’ UTR translational initiation sites in human and mouse. PLoS Comput Biol 12(10):e1005170 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Schneider-Poetsch T, Ju J, Eyler DE, Dang Y, Bhat S, Merrick WC, Green R, Shen B, Liu JO (2010) Inhibition of eukaryotic translation elongation by cycloheximide and lactimidomycin. Nat Chem Biol 6(3):209–217 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Sruthi KB, Menon A, P A, Vasudevan Soniya E. 2022. Pervasive translation of small open reading frames in plant long non-coding RNAs. Front Plant Sci 13: 975938. [DOI] [PMC free article] [PubMed]
  36. Stern-Ginossar N, Weisburd B, Michalski A, Le VT, Hein MY, Huang SX, Ma M, Shen B, Qian SB, Hengel H et al (2012) Decoding human cytomegalovirus. Science 338(6110):1088–1093 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Urquidi Camacho RA, Lokdarshi A, von Arnim AG (2020) Translational gene regulation in plants: a green new deal. Wiley Interdiscip Rev RNA 11(6):e1597 [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Uygun S, Azodi CB, Shiu SH (2019) Cis-regulatory code for predicting plant cell-type transcriptional response to high salinity. Plant Physiol 181(4):1739–1751 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. van der Horst S, Snel B, Hanson J, Smeekens S (2019) Novel pipeline identifies new upstream ORFs and non-AUG initiating main ORFs with conserved amino acid sequences in the 5’ leader of mRNAs in Arabidopsis thaliana. RNA 25(3):292–304 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. von Arnim AG, Jia Q, Vaughn JN (2014) Regulation of plant translation by upstream open reading frames. Plant Sci 214:1–12 [DOI] [PubMed] [Google Scholar]
  41. Wang Q, Mao Y (2023) Principles, challenges, and advances in ribosome profiling: from bulk to low-input and single-cell analysis. Adv Biotechnolog 1(4):6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Wang Z, Cui Q, Su C, Zhao S, Wang R, Wang Z, Meng J, Luan Y (2023) Unveiling the secrets of non-coding RNA-encoded peptides in plants: a comprehensive review of mining methods and research progress. Int J Biol Macromol 242(Pt 3):124952 [DOI] [PubMed] [Google Scholar]
  43. Willems P, Ndah E, Jonckheere V, Stael S, Sticker A, Martens L, Van Breusegem F, Gevaert K, Van Damme P (2017) N-terminal proteomics assisted profiling of the unexplored translation initiation landscape in arabidopsis thaliana. Mol Cell Proteomics 16(6):1064–1080 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Wu HL, Song G, Walley JW, Hsu PY (2019) The tomato translational landscape revealed by transcriptome assembly and ribosome profiling. Plant Physiol 181(1):367–380 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Wu TY, Goh H, Azodi CB, Krishnamoorthi S, Liu MJ, Urano D (2021) Evolutionarily conserved hierarchical gene regulatory networks for plant salt stress response. Nat Plants 7(6):787–799 [DOI] [PubMed] [Google Scholar]
  46. Wu HL, Ai Q, Teixeira RT, Nguyen PHT, Song G, Montes C, Elmore JM, Walley JW, Hsu PY (2024a) Improved super-resolution ribosome profiling reveals prevalent translation of upstream ORFs and small ORFs in Arabidopsis. Plant Cell 36(3):510–539 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Wu HL, Jen J, Hsu PY (2024b) What, where, and how: regulation of translation and the translational landscape in plants. Plant Cell 36(5):1540–1564 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Wu TY, Li YR, Chang KJ, Fang JC, Urano D, Liu MJ (2024c) Modeling alternative translation initiation sites in plants reveals evolutionarily conserved cis-regulatory codes in eukaryotes. Genome Res 34(2):272–285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Yandell M, Ence D (2012) A beginner’s guide to eukaryotic genome annotation. Nat Rev Genet 13(5):329–342 [DOI] [PubMed] [Google Scholar]
  50. Zhang P, He D, Xu Y, Hou J, Pan BF, Wang Y, Liu T, Davis CM, Ehli EA, Tan L et al (2017a) Genome-wide identification and differential analysis of translational initiation. Nat Commun 8(1):1749 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Zhang S, Hu H, Jiang T, Zhang L, Zeng J (2017b) TITER: predicting translation initiation sites by deep learning. Bioinformatics 33(14):i234–i242 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Zhang T, Wu A, Yue Y, Zhao Y (2020) uORFs: important Cis-regulatory elements in plants. Int J Mol Sci 21(17):6238 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Zhu X, Thalor SK, Takahashi Y, Berberich T, Kusano T (2012) An inhibitory effect of the sequence-conserved upstream open-reading frame on the translation of the main open-reading frame of HsfB1 transcripts in Arabidopsis. Plant Cell Environ 35(11):2014–2030 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The LTM- and CHX-based Ribo-seq in Arabidopsis (suspension cells), Arabidopsis pollens and tomato were retrieved from NCBI Gene Expression Omnibus (GEO; https://www.ncbi.nlm.nih.gov/geo/) under accession number GSE145795, GSE88790, GSE143311.


Articles from Plant Molecular Biology are provided here courtesy of Springer

RESOURCES