Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 May 19.
Published in final edited form as: Dev Cell. 2025 Jan 15;60(10):1498–1515.e8. doi: 10.1016/j.devcel.2024.12.038

The regulatory landscape of 5′ UTRs in translational control during zebrafish embryogenesis

Madalena M Reimão-Pinto 1,6, Sebastian M Castillo-Hair 2,3, Georg Seelig 2,4, Alexander F Schier 1,5,6
PMCID: PMC12092203  NIHMSID: NIHMS2049895  PMID: 39818206

Summary

The 5′ UTRs of mRNAs are critical for translation regulation during development, but their in vivo regulatory features are poorly characterized. Here, we report the regulatory landscape of 5′ UTRs during early zebrafish embryogenesis using a massively parallel reporter assay of 18,154 sequences coupled to polysome profiling. We found that the 5′ UTR suffices to confer temporal dynamics to translation initiation, and identified 86 motifs enriched in 5′ UTRs with distinct ribosome recruitment capabilities. A quantitative deep learning model, DaniO5P, identified a combined role for 5′ UTR length, translation initiation site context, upstream AUGs and sequence motifs on ribosome recruitment. DaniO5P predicts the activities of maternal and zygotic 5′ UTR isoforms and indicates that modulating 5′ UTR length and motif grammar contributes to translation initiation dynamics. This study provides a first quantitative model of 5′ UTR-based translation regulation in development and lays the foundation for identifying the underlying molecular effectors.

Keywords: Massively parallel reporter assay, 5′ UTR, translation initiation control, zebrafish embryogenesis, deep learning model

In brief

Reimão-Pinto et al. introduce thousands of different mRNA reporters into zebrafish embryos to investigate the contribution of 5′ UTR features to translation initiation during embryogenesis. The work defines the regulatory potential of 5′ UTRs and develops a deep learning model that predicts translational dynamics based on the 5′ UTR sequence.

Graphical Abstract

graphic file with name nihms-2049895-f0008.jpg

Introduction

Translational control orchestrates the timing and extent of protein synthesis during embryogenesis13. For canonical eukaryotic translation to start, the ribosomal pre-initiation complex (PIC) is recruited to the 5′ cap and then scans the 5′ UTR until it reaches a suitable start codon. Only then is a translationally competent ribosome assembled4. Thus, the 5′ UTR serves as a point of control for selective mRNA translation and translation initiation can be rate-limiting for protein production57. Despite their essential role in translation initiation, 5′ UTRs have not been comprehensively characterized.

5′ UTRs harbor elements such as RNA structures and short upstream open reading frames (uORFs) that impact translational output79, or sequence motifs that can recruit regulatory trans-acting factors like RNA-binding proteins (RBPs)10,11. Additional complexity arises from 5′ UTR transcript variants of the same gene. For example, in yeast, 5′ UTR variants resulting from alternative transcription start site (TSS) utilization drive protein levels during meiotic differentiation12. Zebrafish embryogenesis is also characterized by TSS switching13,14, but to what extent 5′ UTR isoform switching impacts the dynamics of translation during the MZT is unknown.

More generally, we lack a systematic understanding of the regulatory roles of 5′ UTRs at transcriptomic scale during vertebrate embryogenesis. This in part because translation is influenced by features outside the 5′ UTR, such as codon composition, 3′ UTR regulatory elements or poly(A) tail length1,2,15. Previous studies found that transcripts with 5′ and 3′ UTR variants can display distinct translatability, but could not determine the regulatory contributions of the 5′ UTR sequence alone12,1623. Massively parallel reporter assays (MPRAs)2432 allow to interrogate the regulatory capability of thousands of pre-defined sequences independently of other transcript features, and have for example identified 3′ UTR sequences that regulate mRNA degradation and polyadenylation during the zebrafish MZT3335.

To systematically determine the contribution of zebrafish 5′ UTRs to translation initiation during early embryogenesis, we developed an in vivo 5′ UTR MPRA coupled to polysome profiling. We assayed mRNA reporters harboring 5′ UTRs representing over 10,000 expressed genes, including 5′ UTR isoforms arising from alternative TSS usage. Our work shows that the 5′ UTR sequence alone is sufficient to regulate ribosome recruitment dynamics by over 100-fold during embryogenesis. The MPRA captured the broad inhibitory impact of uORFs, the optimal translation initiation site context sequence (AAACAUG), and motifs enriched in 5′ UTRs that display distinct translational behaviors during embryogenesis. Dozens of motifs are uncharacterized, while others match consensus motifs of RBPs in other species. Our analyses identify a previously unappreciated effect of 5′ UTR length on ribosome recruitment during embryogenesis and show that switching 5′ UTR isoforms display distinct regulatory capabilities. To quantify the contribution of 5′ UTR features, we developed a deep learning model, DaniO5P. By integrating the contribution of sequence features and 5′ UTR length, DaniO5P predicts the differential regulatory activities of 5′ UTR isoforms. This study provides a comprehensive resource to study the role of 5′ UTR motifs in translation initiation during development and introduces DaniO5P as a powerful tool to predict the regulatory effects of 5′ UTRs.

Results

Construction of a 5′ UTR reporter library

The study of 5′ UTR activity relies on accurate annotation of TSSs and start codons. However, genome-wide transcriptome analyses do not precisely annotate TSSs, and databases provide reviewed 5′ UTR sequences for only a limited number of zebrafish transcripts8,3638. To recover zebrafish 5′ UTRs, we used a publicly available cap analysis of gene expression (CAGE) during zebrafish development14. CAGE determines TSSs of transcribed and 5′ capped mRNAs at single nucleotide resolution39. We re-mapped CAGE data of 12 zebrafish developmental stages14 and matched CAGE-recovered TSS genomic coordinates with coding sequence (CDS) start site coordinates of annotated transcripts (Figures 1A and S1A; STAR Methods). We recovered 13,309 5′ UTR sequences of 10,354 genes expressed during the first 33 hours of development (Table S1), with lengths ranging from 15 to 2,813 nucleotides (nts) and a median of 141 nts (Figure 1B), comparable to the length of previously confidently annotated zebrafish 5′ UTRs8. To identify 5′ UTR switching isoforms (Figure S1B), we employed CAGEr 40, which determines significant shifts in TSS usage across developmental stages13,14,40 (STAR Methods). Approximately 10% of the expressed genes (n = 1,075) switched 5′ UTR isoforms (Table S1), with the majority of the switching events occurring during the major wave of zygotic genome activation (ZGA), as previously described13.

Figure 1. A reporter library to characterize 5′UTR-mediated translational regulation during zebrafish embryogenesis.

Figure 1.

(A) Strategy for 5′ UTR recovery. CAGE data14 was used to recover transcription start sites (TSSs) and integrated with annotated coding sequence (CDS) start sites. (B) Violin plot showing 5′ UTR length distribution. Boxplot center line: median; box limits: upper/lower quartiles. (C) mRNA library design. See Figure S1 and Table S1.

To investigate how 5′ UTRs affect translation initiation independently of other mRNA regions, we designed a reporter library where 5′ UTRs are placed in an otherwise constant mRNA context. Endogenous 5′ UTRs between 15 and 238 nts in length (n = 9,863, ~74% of the library) were included fully. The remaining longer 5′ UTRs (n = 3,446) were split into shorter sequences to accommodate oligo synthesis limitations (Figure S1C), henceforth called split-segment 5′ UTRs. This approach resulted in a library consisting of 18,154 unique sequences representing 13,309 zebrafish 5′ UTRs (Figure S1F; Table S1), including 5′ UTR switching variants (n = 2,721) (Figure S1B). The library was synthesized and in vitro transcribed to give rise to an mRNA pool where 5′ UTRs are downstream of a common adaptor, and upstream of a superfolder GFP (sfGFP) reporter, a 3′ UTR devoid of known regulatory elements33, and a 36-nt-long poly(A) tail characteristic of adenylated maternal mRNAs41 (Figures S1D and E; STAR Methods). Our finally library contained > 99% of the designed 5′ UTR sequences (Figures S1FH).

In vivo zebrafish 5′ UTR MPRA uncovers translation initiation regulation

To systematically determine the impact of 5′ UTRs on translation initiation during early zebrafish embryogenesis, we devised an in vivo 5′ UTR MPRA coupled to polysome profiling and next-generation sequencing (Figure 2A). We focused on early embryonic stages when the embryo transitions from a mass of cells undergoing successive coordinated mitotic divisions (pre-gastrulation) to asynchronous cell divisions accompanied by lengthening of the cell-cycle and cell fate specification (onset of gastrulation ~ 5.5 hours post-fertilization; hpf)42,43. We injected the 5′ UTR mRNA library into 1-cell stage embryos and collected them at the 64-cell (2 hpf), sphere (4 hpf), shield (6 hpf) and bud (10 hpf) stages, and repeated this experiment on different days for a total of three independent replicates. The reporter library was readily translated, evidenced by increasing sfGFP fluorescence from the 64-cell stage onward (Figure S2A). We then performed polysome profiling (Figure 2B) and observed an increase in polysomes as embryos develop, indicative of a gradual increase in global translation, as previously reported44. We collected fractions corresponding to monosomes (80S), low molecular weight (LMW) and high molecular weight (HMW) polysomes, extracted RNA from each fraction and from total embryo RNA (total) at each timepoint (2, 4, 6 and 10 hpf) (Figures 2B and S2B), and prepared next-generation sequencing libraries (STAR Methods). We calculated relative abundances (transcript per million, TPM) for input (total) and polysome fractionated (80S, LMW and HMW) samples throughout the time-course (Figure 2C; Table S1), and found high inter-replicate correlation (Pearson’s r = 0.75-0.99; Figure S2D).

Figure 2. 5′ UTR MPRA uncovers translation initiation regulation.

Figure 2.

(A) Schematic of the 5′ UTR MPRA. (B) Polysome profiling traces from one replicate. (C) Ribosome recruitment scores (RRSs) calculated for each fraction and timepoint (TPM, transcript per million). (D) Heatmap representing mean log2 transformed RRS (log2 RRS) values of 5′ UTRs (rows, n = 17,879), organized into nine hierarchical clusters. (E) Density plots of log2 RRS values for three clusters displaying distinct ribosome recruitment dynamics. (F) Log2 RRS values for 80S, LMW and HMW fractions of three example 5′ UTRs (mean ± SD, n = 3). (G) Fluorescence microscopy images of embryos injected with egfl6-, nup43-, or hnrnpl-5′UTR-sfGFP reporters and control dextran red dye. TL – transmitted light. Scale bar = 250 μm. (H) Swarm plot of sfGFP fluorescence normalized to dextran intensity (y axis in log2 scale); bars: mean value; each dot represents one embryo (n = 50) from 2 injection rounds, **** p-value < 0.0001 Welch’s ANOVA with Dunnett’s T3 multiple comparisons test of log2 transformed data. See Figure S2, Table S2 and STAR Methods.

To determine differential ribosome occupancy of reporter mRNAs, we calculated the ribosome recruitment score (RRS)24, defined as the ratio between reporter abundance in the ribosome-bound fraction and the total RNA pool (Figure 2C; STAR Methods). This calculation yielded RRS values for each fraction at each timepoint (RRS80S, RRSLMW and RRSHMW at 2, 4, 6 and 10 hpf) (Table S2), which reflect the proportion of reporter transcripts engaged in active translation.

We observed that the 5′ UTR sequence alone is sufficient to modulate translation initiation during early zebrafish embryogenesis (Figure 2D and S2E). Hierarchical clustering of RRS values shows the translational effects mediated by the 5′ UTR, which can impair (log2 RRS < 0) or promote (log2 RRS > 0) ribosome recruitment. Importantly, these differences were not a mere consequence of differential reporter availability for translation in the mRNA pool (Figure S2F), indicating that active translational regulatory mechanisms are at play. We identified groups of 5′ UTRs that confer different translational behaviors (Figure 2E; Table S2). For instance, sequences in Cluster 3 (n = 2,232) led to relatively constant reporter representation across fractions until 6 hpf, whereas 5′ UTRs in Cluster 6 (n = 659) promoted ribosome recruitment while selectively dampening initiation at the shield (6 hpf) stage. 5′ UTRs in Cluster 8 (n = 51) led to preferential monosome occupancy of the reporter transcript. These data uncover the regulatory potential that lies in 5′ UTRs and their dynamic regulation throughout embryogenesis.

To test whether RRS is in fact a measure of protein expression, we co-injected either a poorly (5′ UTR egfl6), intermediately (5′ UTR nup43) or well (5′ UTR hnrnpl) ribosome-recruited reporter with a fluorescent dextran dye as injection control and measured relative fluorescence intensity during embryogenesis (Figures 2FH). As expected, embryos injected with 5′ UTR-egfl6-sfGFP resulted in lower sfGFP accumulation throughout the time-course, whereas 5′ UTR-hnrnpl-sfGFP yielded 12-fold higher sfGFP by the end of gastrulation (sfGFP half-life ~ 20 hrs45). The 5′ UTR-nup43-sfGFP reporter displayed an intermediate level of protein output, in agreement with initiation being rate-limiting for protein synthesis in this biological context. Taken together, the MPRA systematically determined the impact of 5′ UTRs on translation initiation during early zebrafish embryogenesis.

5′ UTR regulatory elements shape translation initiation in early development

We first evaluated the relationship between known 5′ UTR cis-acting elements68 and RRS. We focused our analysis on the HMW fraction at the bud stage (10 hpf), when global translational activity is high44 (Figure 2B). Previous studies identified uORFs as prevalent features in vertebrate 5′ UTRs that regulate translation4649. Using ORFik50, we determined that around 61% of the assayed 5′ UTRs contain predicted AUG-initiated uORFs (STAR Methods; Table S2). We observed that 5′ UTRs with a single uORF displayed significantly lower RRSHMW (p-value < 2.2x10−16, Wilcoxon rank sum test) values than those without (Figure 3A). The presence of additional uORFs further impaired ribosome recruitment in a uORF number-dependent manner, albeit to a lesser extent. This result agrees with a broad inhibitory effect of uORFs on downstream CDS translation in zebrafish embryos47,48, presumably by detachment of the scanning PIC49. Nonetheless, uORF-containing 5′ UTRs display a broad range of RRSHMW values, consistent with a context-dependent effect of uORFs on scanning47,49 and with a combinatorial impact of other 5′ UTR regulatory features on translation initiation.

Figure 3. 5′ UTR-mediated regulation of translation initiation and stability during early development.

Figure 3.

(A) Impact of upstream open reading frame (uORF) number on ribosome recruitment. Violin plots of log2 RRSHMW values for 5′ UTRs binned by uORF count. Boxplot center line: median; box limits: upper/lower quartiles. (B) Effect of 5′ UTR length on ribosome recruitment. Each dot represents one 5′ UTR (n = 17,906), adjusted R2 values for best-fit to a nonlinear generalized additive model and corresponding p-value are presented. (C) Relationship between length-adjusted structure formation propensity and ribosome recruitment. Adjusted R2 value for best-fit to a linear model and corresponding p-value are presented. (D) Position weight matrixes of translation initiation site (TIS) consensus contexts. Mean nucleotide frequency at positions −4 to −1 of the canonical sfGFP translation start site for 5′ UTRs with lowest (bottom 10%, n = 1,791; left panel) and highest RRSHMW values (top 10%, n = 1,791; right panel) at 10 hpf. Grey AUG represents sfGFP translation start site. (E) Heatmaps of normalized reporter abundance changes across the time-course (left), respective log2 RRSHMW values at 10 hpf (center) and 5′ UTR length (right). Reporters hierarchically clustered by normalized abundances. (F) Mean normalized abundance changes for 5′ UTR reporters in each of the stability clusters presented in (E). See STAR Methods, Figure S3 and Table S3.

5′ UTR length has been implicated in translation efficiency in mammalian cells, with contradictory observations made between in vivo and in vitro systems5154. Our data shows a negative relationship between RRSHMW and 5′ UTR length (Figure 3B, Adj. R2 = 0.82, p-value < 2x10−16), indicating an overarching impact of 5′ UTR length on translation initiation during embryogenesis. This length dependency persisted in 5′ UTRs without predicted uORFs (Figure S3A, Adj. R2 = 0.84, p-value < 2x10−16).

Stable structures in the 5′ UTR can impact translation in vitro and in vivo24,25,5561. To determine the effect of RNA structure on ribosome occupancy, we predicted 5′ UTR secondary structures using two complementary approaches: by calculating sequence-based ensemble free energies for all 5′ UTRs (NUPACK)62 and by using the deep-learning-based MXfold2 algorithm63 (Table S3). RNA folding propensity scales almost linearly with length (Figure S3B), which does not necessarily imply increased structural complexity64. In fact, the likelihood of RNA base-paring within a 5′ UTR sequence does not depend on its length (Figure S3C). Thus, we determined length-adjusted free energies (STAR Methods; Table S3), and found no relationship with RRSHMW (Figure 3C, Adj. R2 = 0, p-value = 0.1391). Similarly, we found only a very weak positive effect of average base-pairing probability on initiation capability (Figure S3D, Adj. R2 = 0.034, p-value < 2x10−16). Accordingly, 5′ UTRs with the same length displayed broad folding scores but no significant relationship between their MXfold2 scores and respective RRSHMW values (Figure S3E). Nonetheless, we find examples of 5′ UTRs with the same length with different ribosome recruitment capabilities, perhaps due to their distinct structural features; but also 5′ UTRs with similar RRSHMW despite markedly different RNA folding scores. We also observed a very weak correlation between 5′ UTR GC content and RRSHMW (Figure S3F, Adj. R2 = 0.008; p-value < 2x10−16), suggesting that GC-richness per se does not greatly affect the initiation process in vivo.

We next focused on the sequence context preceding the translation initiation site (TIS) of the main ORF. Landmark studies showed that this sequence is not random, and that CRCCAUGG (where R denotes a purine) is the optimal consensus sequence for translation initiation (the Kozak sequence) in mammals65,66. The zebrafish genome is characterized by the TIS consensus context NRNCAUGG6769 In particular, AAACAUG is associated with high translation rates49,68. To analyze at large scale the effect of the TIS on RRS, we selected reporters from the top and bottom RRSHMW deciles (Table S3), and calculated nucleotide frequencies for the four nucleotides preceding the sfGFP ORF. Strikingly, this analysis identified 5′ UTRs with the AAACAUG context among the best ribosome recruited-reporters (Figure 3D, top 10%). In contrast, poorest ribosome-recruited 5′ UTRs displayed no clear TIS sequence (nucleotide frequency ranging from approximately 20% to 28%), but a higher frequency of uridine (U) at positions −4 to −1 (Figure 3D, bottom 10%) relative to that of all MPRA 5′ UTRs (Figure S3G). Interestingly, 75% (1,339 out of 1,791) of the sequences in the bottom 10% corresponded to split-segment 5′ UTRs, whereas only full-length sequences were present in the top decile (n = 1,791). When comparing TIS nucleotide frequencies for full-length (n = 9,756; mean length of 109 nt) or split-segment 5′ UTRs (n = 8,150; mean length of 172 nt) (see also Figure S1C), we observed that while native 5′ UTRs displayed a clear optimal TIS context (AAACAUG), split-segment sequences showed comparatively higher representation of Us at positions −4 to −1 (Figure S3H). These analyses indicate that the presence of Us in this region is detrimental for translation initiation in zebrafish, which we could experimentally determine by including 5′ UTRs with non-endogenous initiation contexts. These results are similar for the 80S and LMW fractions (Table S3).

Finally, we asked if distinct initiation rates imposed by 5′ UTR-mediated regulation impact reporter mRNA stability. We calculated normalized reporter stabilities based on changes in mRNA levels throughout the time-course (Table S3; see Methods) and identified groups of reporters with distinct decay dynamics (Figure 3E and F). Most (stability clusters 1, 3 and 4) but not all sequences displayed a large decrease in abundance at 6 hpf, consistent with the timing of the onset of translational-dependent mRNA destabilization70,71. Reporters with slower decay kinetics (stability clusters 5 and 6) displayed generally high RRSs, in agreement with translation initiation being sufficient to affect mRNA stability32,7276, independently of the CDS and 3′ UTR sequences. Collectively, the MPRA comprehensively quantified the impact of 5′ UTR structure, length, uORFs and TIS sequence context on translation initiation and mRNA decay dynamics.

The MPRA identifies 5′ UTR cis-regulatory motifs

To identify cis-regulatory motifs that may drive dynamic ribosome recruitment across developmental stages, we took a two-step approach. First, we performed motif enrichment analysis on 5′ UTR sequences consistently found in the top or bottom deciles of each fraction (80S, LMW and HMW) (Figures 4A and S4A; Table S3). We used the MEME tool suite77, which enables de novo motif discovery and scanning of known RBP motifs78 (STAR Methods). This analysis yielded 61 unique motifs (5-12 nt-long): 30 are de novo predicted (Table S4) and the remaining match the consensus binding sequence of known RBPs in other species7882.

Figure 4. In vivo MPRA identifies 5′ UTR cis-regulatory motifs.

Figure 4.

(A) Motif enrichment analysis strategy. (B) Fuzzy c-means clustering of log2 RRS values for 80S, LMW and HMW fractions. Colors represent normalized differences in cluster means (log2 RRS values) for each timepoint (normalized to 2 hpf); clusters include 5′ UTRs with membership score ≥ 0.7. Clusters differ from the ones in Figure 2. (C) Representative motifs enriched in 5′ UTRs consistently found in the top 10% (“enhancing”), (D) bottom 10% (“repressive”), or (E) with high cluster membership scores (≥ 0.7) (“dynamic”). See STAR Methods, Figure S4 and Table S4.

Next, we performed unsupervised soft clustering analysis83,84 (STAR Methods) of RRS values across developmental time (2, 4, 6 and 10 hpf) for each fraction (80S, LMW and HMW), which identified clusters of 5′ UTRs that display coordinated changes in ribosome recruitment (Figure 4B; Table S3). We manually grouped clusters with qualitatively similar dynamics (Figures S4B and S4C) and ran motif enrichment analysis using the MEME tool suite77. This analysis returned an additional 25 motifs associated with temporal RRS dynamics, of which 14 are de novo predicted motifs (Table S4).

We categorized motifs as “enhancing”, “repressive” or “dynamic” based on whether they were enriched in the top or bottom RRS deciles, or in any of the dynamic-based clusters (Figures 4CE and S4DF; Table S4). Motifs identified in opposing deciles were largely non-overlapping, and similar motifs were commonly enriched across different fractions. For example, CA-rich motifs (e.g. ACACACA, MAUCCAR and AMAWACA) were consistently enriched in the top decile of all fractions (Figures 4C and S4D; Table S4), whereas G- and GC-rich motifs (e.g. GAKGAGGRRGAG and GMGCGCKCGSYC) and pyrimidine-rich motifs (e.g. UCUCUCUCUYUC and CCCUCUCYCYCY) were enriched in 5′ UTRs associated with poor ribosome recruitment (Figures 4D and S4E; Table S4). For some cases, the motif identified was enriched in only one of the fractions, as was the case for “enhancing” motifs HGGAGAA (top 10% 80S fraction), ACUUCCGG (top 10% LMW fraction) and AGUUGUUCC (top 10% HMW) (Figures 4C and S4D), or “repressive” motifs AUUUUUU (bottom 10% 80S fraction), GGGAGGG (bottom 10% LMW) and CAGAAGAGCAGC (bottom 10% HMW) (Figures 4D and S4E).

Among the “dynamic” motifs, UGU-rich motifs enriched in 5′ UTRs that display a decrease in RRS at the bud (10 hpf) stage stood out, namely UGUGUGUGUGUG, UGUKURUKU and UUUGUUU, as each was independently identified in distinct clusters that display analogous RRS dynamics (Figure 4E; Table S4). Shorter motifs (5 to 7 nt-long) may also contribute to the temporal regulation of transcript ribosome load. For example, the motif CCCGCC (80S cluster 1) is associated with monosome recruitment that gradually increased during gastrulation whereas the motif CCUYCCC (LMW cluster 3) is enriched in 5′ UTRs with similar temporal behavior but that led to the recruitment of 2 to 4 ribosomes per transcript. In contrast, motifs with GA dinucleotides (e.g. GAGAGARAGAGA, AGAGAAA and GAAGAAG) are associated with dynamic recruitment to polysomal fractions (Figure S4F; Table S4). Note that some motifs with high similarity can be categorized as both enhancing/repressive and dynamic (e.g., the “enhancing” ACACACA motif and the “dynamic” CACACACACACA motif) (Table S4). In summary, we identified 86 motifs (5-12 nt-long) enriched in 5′ UTRs associated with distinct translation dynamics, 44 of which are uncharacterized motifs with putative roles in translation initiation control.

DaniO5P predicts the effect of 5′ UTR sequences on translation dynamics

To quantify how 5′ UTR features regulate translation and mRNA abundance, we trained and interpreted neural network predictors using the MPRA data. Following our previous work in human cells26, we summarized reporter translation via the mean ribosome load (MRL), by taking the mean transcript abundances in each polysome fraction weighted by their approximate mean ribosome number (Figure 5A) (STAR Methods). Similarly, to summarize the mRNA degradation state, we used the change in total mRNA abundance during each time interval (Δlog2Xtiti−1). As with RRS, log2 MRLti and Δlog2Xtiti−1 were highly dependent on 5′ UTR length. In fact, a simple polynomial “length model” explained 38-86% of their variance (Figures 5B, 5C and S5A; Table S4). To capture the influence of sequence features beyond length, we trained an ensemble of 10 convolutional neural networks (CNNs) to simultaneously predict log2 MRLti and Δlog2Xtiti−1 residuals (observed - length model prediction) from 5′ UTR sequences (Figure 5B) (STAR Methods). Remarkably, in sequences held out from training, CNN predictions explained up to 53% of the remaining variation, and the combined length & CNN model accounted for 60-93% of the total log2 MRLti and Δlog2Xtiti−1 variation (Figure 5C). We termed the combined model Danio Optimus 5-Prime (DaniO5P). DaniO5P captured observed MRL dynamics, explaining, for example, 73% of the change in log2 MRL between 2 and 10 hpf, whereas the length model alone could only explain 51% of the variation (Figures 5C, 5D and S5B). Inspired by recent work on splicing modeling85, we explored introducing NUPACK-calculated secondary structure information as additional inputs to the CNN (STAR Methods), but these did not improve model performance (Table S4). Overall, these results indicate that 5′ UTR length and sequence element grammar cooperate to fine tune translation initiation.

Figure 5. DaniO5P identifies sequence determinants of translation dynamics.

Figure 5.

(A) MRL calculation. (B) DaniO5P model, where predictions arise from the combination of a polynomial “length model” and an ensemble of 10 CNNs. (C) Prediction performance separated by contributions of the length and sequence model. (D) Examples of 5′ UTRs with dynamics not explained by length alone, showing CNN predictions account for these differences. (E) Nucleotide contributions (SHAP values) of the unc119b 5′ UTR to log2 MRL at 2 hpf, 10 hpf, their difference, and Δlog2 X between 0 and 2 hpf. The y axis scale was adjusted differently for the TIS context region and upstream regions. (F) Sum of nucleotide contributions within the TIS context region (−4 to −1) and outside (upstream of −4) for every sequence in the MPRA. (G) Motifs extracted from DaniO5P and their average contribution to log2MRLti and Δ log2Xtiti−1 at all timepoints. See Figure S5 and STAR Methods.

To recover sequence features learned by DaniO5P, we calculated the contribution of each nucleotide in every MPRA sequence to log2 MRLti and Δlog2Xtiti−1 predictions (STAR Methods). Contribution scores recapitulated expected sequence features such as uAUGs, which contribute negatively to MRL at all timepoints (Figures 5G and S5C), and the impact of the TIS context (positions −4 to −1), consistent with our frequency-based analysis (Figure 3D). Interestingly, the effect of some features changed throughout embryogenesis. For example, long stretches of Gs in the 5′ UTR of unc119b (Figure 5E) and CU repeats in prkaa1 (Figure S5E) negatively contribute to MRL with a markedly stronger effect at 2 hpf compared to 10 hpf. In contrast, the presence of an adenosine (A) at position −3 of the main sfGFP ORF had a stronger enhancing contribution at 10 hpf than at 2 hpf (Figures S5D and 5E). Furthermore, G-repeats had an enhancing effect over stability until 2 hpf, whereas U repeats had the opposite effect (Figures 5E and S5D). Across all sequences, the TIS context (positions −4 to −1) showed a stronger contribution to log2 MRL and Δ log2 X at later stages of embryogenesis (10 hpf compared to 2 hpf), whereas nucleotides upstream contribute more strongly at earlier stages (Figure 5F).

Finally, we characterized sequence motifs learned by DaniO5P. We extracted motifs from its convolutional filters, clustered them to eliminate redundancy, and filtered based on their reproducibility across independently trained models, resulting in 13 motifs (STAR Methods). We next calculated their average contribution to model predictions, and found groups of motifs that contribute to changes in MRL throughout embryogenesis (pyrimidine-rich motifs, G-rich motifs), or to overall enhancement (U-rich motifs, CA-rich motifs) or dampening (uAUGs) (Figure 5G; Table S4). Motifs also had a strong effect over stability between 0 and 2 hpf, with G-rich and U-rich displaying the strongest enhancing and destabilizing effect, respectively. The most MRL inhibitory uAUGs had a strong TIS context, that is, with A/G at position −3. CU-rich motifs become less inhibitory at later stages (10 hpf) (Figure 5G; Table S4), agreeing with findings reported by a complementary 5′ UTR MPRA method28. To assess whether motifs or nucleotide frequency were responsible for the observed behavior, we shuffled sequences upstream of their −5 position to preserve nucleotide composition and TIS context while destroying motifs. Predictions on 100 shufflings were on average closer to what would be expected from length alone (Figures S5CF), consistent with regulation being driven by motifs. Additionally, some motifs exhibit positional dependence. For example, pyrimidine- and G-rich motifs more strongly contributed to MRL when located closer to the 5′ end, in contrast to U-rich or uAUGs (Figures S5G and S5H). Secondary structure unpaired probabilities within motif sites had little influence over their contribution (Figure S5I), further suggesting that secondary structure has a limited effect over 5′ UTR-mediated regulation.

Pair-wise comparison of motifs identified by DaniO5P or by sequence enrichment analysis pinpoints motifs with significant sequence similarity (p-value < 0.05) (Figure S6A; Table S4). Motifs with high similarity score (p < 1x10−4) generally displayed analogous effects on predicted MRL and measured RRS (Figures 4CE, S4DF and 5G). For example, model and RRS analyses indicate that pyrimidine-rich motifs and G-rich motifs are inhibitory. Similarly, both analyses pinpoint CA-rich motifs as enhancing that can be associated with an increase in ribosome recruitment as embryogenesis progresses. Thus, DaniO5P is remarkably predictive of 5′ UTR activity and provides a powerful approach to define the interaction of distinct 5′ UTR features on translation initiation and stability regulation.

Switching 5′ UTR isoforms display different translation initiation capabilities during the MZT

TSS switching during the zebrafish MZT can generate maternal and zygotic transcript isoforms with different 5′ UTRs that encode the same gene product (Figure 6A). We hypothesized that switching 5′ UTRs could provide an additional layer to gene expression regulation by modulating translation initiation. Analysis of differential TSS utilization identified cases of 5′ UTR isoform switching events (Figure 6A; Tables S1 and S5), as has been described previously13,14. Transcript 5′ UTR isoforms display a similar length distribution (median = 126 nts; Figure S6B), indicating that switching does not favor shortening nor lengthening of the 5′ UTR. TSS switching is bi-directional13 and tends to alter 5′ UTR length by a narrow size window (Figure S6C; Table S5), with a mean length shift of 37 nts between maternal and zygotic isoform pairs.

Figure 6. Switching 5′ UTR isoforms display different translation initiation capabilities.

Figure 6.

(A) 5′ UTR isoform switches during the MZT. (B) Difference in log2 RRSHMW values at 10 hpf between zygotic and maternal isoforms pairs. Maternal 5′ UTR isoforms leading to higher ribosome recruitment (Δlog2(RRSHMW) ≤ −2) are marked as “top-ranked maternal” (n = 33); zygotic 5′ UTR isoforms leading to higher ribosome recruitment (Δlog2(RRSHMW) ≥ 2) are marked as “top-ranked zygotic” (n = 54). Only 5′ UTRs shorter than 239 nts were considered. (C) Violin plots of length distributions for top-ranked maternal (n = 33), top-ranked zygotic (n = 54) and other (n = 2,580) switching isoforms. (D) Violin plot of the length distribution and bar blot of uORF number for top-ranked maternal isoforms and corresponding zygotic pairs, and (E) top-ranked zygotic isoforms and corresponding maternal pairs. (F) Isoform switches shortening 5′ UTR length and removing uORFs enhance ribosome recruitment. In all violin plots: Boxplot center line: median; box limits: upper/lower quartiles. **** p-value < 0.001, two-sample rank sum Wilcoxon test. See Figure S6, Table S4 and S5.

To determine whether 5′ UTR isoforms confer distinct translation initiation capabilities, we analyzed 5′ UTRs that were assayed as full uninterrupted sequences (n = 1,405 out of 2,721, that range between 15-238 nts). We considered isoform pairs of maternally deposited and zygotically re-expressed transcripts (maternal and zygotic pairs) and of zygotic transcript isoforms expressed throughout embryogenesis (ZGA and post-ZGA pairs) (Figure S6B). First, we calculated the difference in RRS (Δlog2(RRSHMW)) for the HMW fraction at the end of gastrulation between zygotic and maternal switching isoforms and found that 5′ UTR isoform switching leads to up to 2 orders of magnitude differences in ribosome recruitment (n = 507 5′ UTR pairs; Figure 6B and Table S5). Switching resulted in higher translatability of the reporter bearing the maternal 5′ UTR (Δlog2(RRSHMW) < 0) or the zygotic 5′ UTR (Δlog2(RRSHMW) > 0), indicating that isoform switching does not impose a unidirectional effect on translation initiation capability.

5′ UTR length and uORF number impact ribosome recruitment (Figures 3 and 5). We therefore selected 5′ UTR pairs that displayed largest differences in ribosome recruitment (absolute Δlog2(RRSHMW) ≥ 2), including 33 5′ UTRs for which the maternal isoform was better initiated (Δlog2(RRSHMW) ≤ −2, “top-ranked maternal”) and 54 5′ UTRs for which the zygotic isoform was better initiated (Δlog2(RRSHMW) ≥ 2, “top-ranked zygotic”) (Figure 6B; Table S5), and determined their 5′ UTR length and number of predicted uORFs. As anticipated, top-ranked isoforms were shorter than all other 5′ UTR variants regardless of it being of maternal or zygotic origin (p-value < 0.001, Wilcoxon rank sum test; Figure 6C). Moreover, we observed that the 5′ UTR length of the top-ranked maternal isoform was significantly shorter than its zygotically-expressed counterpart; the same was true for top-ranked zygotic isoforms and their maternal isoform pairs (Figures 6D and E). Similarly, top-ranked 5′ UTR isoforms had fewer uORFs than their zygotic or maternal variant pairs (Figures 6D and 6E), implying that TSS switching can result in the exclusion of inhibitory uORFs. We made similar observations for isoforms arising from switching events throughout ZGA (ZGA and postZGA 5′ UTR switching pairs, n = 330) (Figures S6BG; Table S5). These data demonstrate that modulation of 5′ UTR length and uORF number by alternative TSS switching is a simple yet powerful mechanism to regulate translation during embryogenesis (Figure 6F).

To explore whether isoform switching alters motif grammar, we used MAST86 to search for the occurrence of motifs identified by the MPRA, finding motif matches (p < 10−4) in 228 of the 2,721 switching 5′ UTRs (Table S4). We identified several cases where switching alters motif occurrence between 5′ UTR isoforms. For example, maternal and zygotic scarb2c 5′ UTR isoforms displayed distinct RRS patterns, with the zygotic isoform resulting in preferential ribosome recruitment during embryogenesis (Figure 7A). Isoform switching of scarb2c leads to the inclusion of a 12 nt stretch matching a UGU-rich motif (p-value = 7.8x10−7) and two uORFs in the zygotic 5′ UTR (Figure 7C). The zygotic isoform of cfl1l resulted in higher ribosome recruitment, particularly at the bud stage (10 hpf) (Figure 7D), and switching shortens the maternal 5′ UTR to eliminate two motifs (GA motif, p-value = 5.7x10−5 and GC-rich motif, p-value = 6.7x10−5) (Figure 7F). DaniO5P more accurately predicted the MRL of maternal and zygotic 5′ UTR isoforms than the length model alone (Figures 7B and E) and pointed to motif-like segments as contributors to the distinct translational behaviors (Figure S7A). These and additional examples (Figures S7B and S7C) showcase how length alone cannot fully explain the regulatory capacity of 5′ UTRs: the maternal scarb2c 5′ UTR and the zygotic jpt2 5′ UTR are of comparable length (129 and 124 nts), and so are the zygotic 5′ UTR isoforms of jpt2 and ube2q2 (92 and 91 nts), yet their translational behaviors are distinct. Our data suggest that TSS switching impacts transcript isoform translation initiation by modulating not only 5′ UTR length and uORF number, but also motif grammar.

Figure 7. Switching 5′ UTRs result in distinct protein output dynamics.

Figure 7.

(A, D) Line plots of log2 RRS values for 80S, LMW and HMW fractions for maternal and zygotic switching 5′ UTR pairs (scarb2c and cfl1l) with distinct translational dynamics. (B, E) Corresponding measured and predicted log2 MRL for length and full DaniO5P models. (C, F) Schematics of motif matches with p-values from MEME suite77. (G, J) Swarm plots of sfGFP fluorescence normalized to dextran intensity (y axis in log2 scale); bars: mean value; each dot represents one embryo (n = 50) from 2 injection rounds. (H, K) Increase in normalized sfGFP fluorescence relative to 2 hpf. (I, J) Fluorescence microscopy images of embryos injected with maternal or zygotic scarb2c- or cfl1l-5′UTR-sfGFP reporters and control dextran red dye. TL – transmitted light. Scale bar = 250 μm. See STAR Methods.

Finally, we asked whether differences in 5′ UTR isoform ribosome recruitment dynamics result in distinct protein synthesis output. We co-injected 1-cell stage embryos with either the maternal or the zygotic 5′ UTR isoform of scarb2c and cfl1l transcripts and a dextran dye as a control, and measured relative sfGFP protein accumulation during the developmental time-course (Figures 7GL) (STAR Methods). We observed that translational output of the sfGFP reporter reflects the distinct ribosome recruitment capabilities of the 5′ UTR isoforms, resulting in protein accumulation to different extents over time. These experiments support that modulating of translation initiation by alternative 5′ UTR isoforms may confer an additional layer of temporal control to protein expression as embryos develop.

Discussion

Our comprehensive in vivo characterization of the role of 5′ UTRs in translation initiation during zebrafish embryogenesis provides five main conclusions. First, 5′ UTR sequences are sufficient to regulate the temporal dynamics of translation (Figure 2). Second, the combined effects of uORFs, 5′ UTR length and TIS context contribute to translation initiation capability (Figure 3). Third, known conserved and uncharacterized motifs are enriched in 5′ UTRs displaying differential ribosome recruitment during embryogenesis (Figures 4 and 5). Fourth, shifts in 5′ UTR length and motif grammar modulate translation initiation during the zebrafish MZT (Figures 6 and 7). Finally, the deep learning model DaniO5P predicts ribosome recruitment and stability solely based on the 5′ UTR sequence (Figures 5 and 6) and provides a powerful tool to dissect the regulatory rules of 5′ UTR elements.

5′ UTR regulatory features acting during embryogenesis

The 5′ UTR MPRA recapitulated the widespread negative impact of uORFs on ribosome recruitment in vivo8,2426,29,4648,87,88 (Figure 3A), and showed that canonical uAUGs with a strong TIS sequence context (purine at position −3 and/or G at +4) are most detrimental for ribosome recruitment (Figure 5G). The data reaffirms the importance of the consensus context AAACAUG for efficient TIS recognition49,66,68 in zebrafish and indicates that the presence of a U nucleotide at positions −4 to −2 is particularly detrimental (Figure 3D). Our results suggest that their impact on translation is conserved across vertebrates, since mutagenesis of the TIS to 4Us (positions −5 to −1) abolished protein synthesis from a reporter plasmid in mammalian cells66, and Us at positions −3 and −2 are universally absent in the consensus TIS of annotated CDSs69,89. A complementary MPRA method also pinpointed A/C nucleotides as enhancing and U at position −3 as particularly detrimental for initiation in zebrafish embryos28. Interestingly, DaniO5P indicates that positions −4 to −1 are less deterministic of ribosome recruitment at earlier stages of embryogenesis (2 hpf) (Figure 5F). It is conceivable that the limited availability of free ribosomes at early stages of zebrafish embryogenesis44 amplifies the regulatory effects of other 5′ UTR sequence features on translation initiation, and perhaps reflects prominent RBP-mediated regulation at early stages of development90.

Notably, we find that 5′ UTR length has a major impact on translation initiation in vivo (Figures 3B and 5C). Once the PIC engages with the mRNA 5′ end, scanning starts and continues in a processive manner throughout the 5′ UTR49,58. In a scenario where the scanning PIC remains tethered to the 5′ cap, which would block the entry of a new PIC, 5′ UTR length would be limiting for translation efficiency51,54,91,92. It will be informative to define which 5′ UTRs promote cap-dependent (where length would be expected to limit initiation) and cap-independent ribosome recruitment. Through several analyses, we found no evidence of 5′ UTR secondary structure regulating translation initiation at a global level. However, comparing secondary structure metrics for sequences of different lengths is not straightforward: free energy is, on average, proportionally higher for a longer sequence64,93, a relationship we confirmed in our data (Figure S3B). Thus, we compared length-adjusted free energy (Figure 3C) and average unpaired probability (Figure S3D) to RRS, and failed to find any strong dependency. Furthermore, the predictive accuracy of DaniO5P did not improve when adding per-base unpaired probability and MFE structure information (Table S4), and motif contributions were found to not depend on the unpaired probability of their matching sites (Figure S5I). These data agree with work showing that the scanning ribosome can melt relatively stable secondary structures56,60, and a study in zebrafish embryos indicating that globally mRNA structures are not major determinants of translation due to the ribosome’s ability to remodel RNA structures in vivo94. Nevertheless, specific structured elements may be influential but poorly represented in our dataset as to arise from our global analyses (Figure S3E). The regulatory effect of specific structures has been clearly demonstrated in other contexts8,24,58,76,92,95,96 and we anticipate that future MPRAs probing for in vivo 5′ UTR RNA structure will determine the interplay between structural elements with uORFs, start codons, motif occurrence, flanking sequence context and their relative positions in the 5′ UTR on mRNA translation. Finally, the MPRA data shows that initiation rate impacts mRNA stability during early development. Our observation that reporters with higher RRS are more stable is in agreement with a protective role for translation initiation on mRNA decay32,72 and contrasts with work in cell lines supporting that a higher ribosome flux triggers mRNA destabilization7476. Future work will need to address the translation-dependent and - independent mechanism(s) at play during embryogenesis and assess the contribution of 5′ UTR motifs to translation initiation control and mRNA decay.

A resource to discover 5′ UTR-binding RBPs

Our study defines a set of motifs enriched in 5′ UTRs with distinct ribosome recruitment dynamics (Figures 4 and 5). RBPs and their sequence specificities are highly conserved11,78,79,9799, but their roles in embryogenesis are poorly understood. Many of the motifs identified by our MPRA match the consensus sequence of known RBPs in other species (Table S4), and homologs of those RBPs are expressed during zebrafish embryogenesis (Figure S4G). Recent work in cell lines has identified common motifs in human 5′ UTRs100,101, suggesting some degree of conservation across vertebrates.

Among the motifs enriched in 5′ UTRs, we find for example the consensus binding sequence for IGF2BP2 and PCBP2 RBPs, which have been described to regulate translation via 5′ UTR-binding102106. Notably, the majority of enriched RBP motifs match those of proteins with roles in pre-mRNA splicing, such as SR proteins and heterogeneous nuclear ribonucleoproteins. Splicing factors can play splicing-independent roles in translation regulation107114, and previous RNA interactome capture experiments showed that RBPs that dynamically bind to cytoplasmic mRNAs during the fly and zebrafish MZT are enriched for proteins involved in mRNA splicing90,115. We postulate that they may be playing broad uncharacterized roles in translational control via 5′ UTR-binding during early embryogenesis. Our work provides a comprehensive resource to analyze the roles of 5′ UTR motifs in translational control and to determine the identity and function of the associated RBPs.

Regulatory potential of 5′ UTR isoforms

Our study demonstrates that the regulatory potential of 5′ UTRs can be tuned by isoform switching during embryogenesis (Figures 6 and 7). These findings extend previous studies in mammalian and yeast cells that reported that longer 5′ UTR isoforms are associated with lower translation1720,116 and that the inclusion of inhibitory uORFs in extended 5′ UTRs reduced translation of the main ORF12,23,117120. MPRA measurements and DaniO5P predictions indicate that 5′ UTR sequence shortening provides an elegant and efficient mechanism to modulate protein synthesis capability. Finally, zebrafish switching 5′ UTR variants can display different motif grammar, and DaniO5P indicates that changes to motif composition impacts their translational effect. Interestingly, cancer cells take advantage of TSS switches to regulate translational output121,122. Moreover, two studies have reported 5′ UTR isoform-specific translation by binding of an RBP to one of the 5′ UTR isoforms but not the other123,124. Thus, differential RBP binding to 5′ UTR variants might contribute to isoform-specific translational regulation to coordinate protein synthesis during zebrafish embryogenesis.

DaniO5P as a resource for studying 5′ UTR-mediated regulation

DaniO5P integrates the effects of sequence length, TIS context, and motifs to make accurate predictions on translation and mRNA abundance kinetics (Figures 5AD). Previous deep learning predictors of 5′ UTR regulation have modeled translation in one condition26,27,125, whereas here we used a multi-task approach to simultaneously predict translation and stability over developmental time. Accordingly, we found regulatory motifs that influence both phenomena in complex, dynamic ways (Figure 5G), as well as complex dependencies on other sequence features such as motif position (Figures S5G and H). Nevertheless, a fraction of the data variability remains to be explained (Figure 5C), and our current model is restricted to 5′ UTRs of up to 238 bp. Future modeling work may capture additional regulatory motifs and more nuanced interactions with other sequence elements. Using more complex architectures such as transformers or foundational pretrained models126 may partially overcome this gap, but ultimately more data specific to zebrafish embryogenesis may be required.

We are releasing DaniO5P along with easy-to-follow example tutorials showing how to make predictions on arbitrary 5′ UTRs. Furthermore, we are including precomputed per-nucleotide, per output contribution scores for all 5′ UTRs in the MPRA, and example code on how to visualize them. We anticipate these resources to be useful in the following scenarios: 1) When studying endogenous 5′ UTRs, our precomputed contribution scores can be used to generate hypotheses about cis-regulatory elements (Figures 5E and S5CF). 2) Similarly, the effects of mutations can be assessed in silico to select examples for experimental testing. 3) DaniO5P can be used in an “active learning” setting: mutated and/or fully synthetic sequences can be designed to be maximally informative for further model training. Deep learning models can take advantage of large datasets well beyond the limits of oligo synthesis127, but some training examples are more useful than others. Using models to guide the design of training datasets has been used to study cis-regulatory elements in developing mouse retinas128, and is therefore an intriguing potential application of DaniO5P. 4) DaniO5P can be used with generative algorithms129,130 to design synthetic 5′ UTRs with desired gene expression patterns over developmental time, to further study 5′ UTR regulation but also as an initial testbed for potential therapeutics27,131.

Limitations of study

The cloning and sequencing strategy requires the presence of a common adaptor at the most 5′ end of the reporter, preventing the recovery of motifs with 5′-end positional dependent activity, such as the 5′ terminal oligopyrimidine (TOP) motif132. Some TSS switching events establish a 5′ TOP motif (Figure S7D), warranting further investigation. Since 5′ UTR motif functionality can be masked by the 3′ UTR (and vice-versa)133, it is possible that the regulatory effects of 5′ UTRs recovered by the MPRA may differ when in their endogenous transcript context. This study does not account for epitranscriptomic marks134 like N6-methyladenosine (m6A), which was shown to promote cap-independent translation initiation135. Finally, polysome fractionation does not allow to distinguish between actively translating and inactive ribosomes nor out-of-frame translation.

Resource Availability

Lead Contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Madalena M. Reimão-Pinto (madalena.pinto@unibas.ch).

Materials Availability

The sequence of the plasmid used for MPRA library cloning is available in the Mendeley Repository associated with this study.

Data and Code Availability

Hight-throughput sequencing datasets have been deposited in the Sequence Read Archive repository (SRA) (BioProject identifier PRJNA1056276, link https://www.ncbi.nlm.nih.gov/sra/PRJNA1056276). Code for the DaniO5P model is accessible at https://github.com/castillohair/DaniO5P.

Any additional information required to reanalyze the data reported in this work paper is available from the Lead Contact upon request.

STAR Methods

Experimental model and study participant details

Zebrafish husbandry and embryo rearing

Zebrafish (Danio rerio, TLAB strain) were grown in standard conditions (28°C at a 14/10 hour light/dark cycle). The TLAB strain was generated by crossing zebrafish natural variant TL (Tupfel Longfin) with AB strain. Zebrafish embryos were collected from TLAB strain crosses, incubated in E3 medium (5 mM NaCl, 0.17 mM KCl, 0.33 mM CaCl2, 0.33 mM MgSO4, pH 7.2) in standard conditions (28°C at a 14/10 hour light/dark cycle) and staged as described139. All experiments were conducted according to federal guidelines for animal research and approved by the Kantonales Veterinäramt of Kanton Basel-Stadt (under the Animal Holding License Form 1035H).

Method details

Alignment and mapping of CAGE data

Raw sequenced CAGE tags (27 bps) from a published zebrafish 12 developmental stages time-course14 were downloaded and mapped to a reference zebrafish genome (GRCz11, excluding alternate loci scaffolds) using Bowtie137 with default parameters as described in Haberle et at., 201413. All analyses were conducted in the R statistical computing environment140 using Bioconductor141 software packages.

CTSS calling

CAGE tag-defined transcription start site (CTSS) calling and positional frequency determination was performed using the CAGEr software package40. To enable comparison between multiple samples, raw tag count was normalized to a reference power-law distribution142 based on a total of 106 tags and α = 1.2, resulting in normalized tags per million. Low fidelity TSSs supported by less than 2 normalized tag counts were filtered out. Individual TSSs were clustered into tag clusters (TCs) using the data-driven parametric clustering method paraclu143 (clusters merged to non-overlapping), and clusters supported by a normalized signal > 5 tpm in at least one developmental stage were included. TCs for which the ratio between the maximal and minimal density was lower than 2 (i.e. stability < 2) and longer than 200 bps were excluded. TCs across all developmental stages within 100 bp of each other were aggregated into a single promoter region.

Recovery of 5′ UTR sequences

Gene annotations corresponding to the zebrafish genome build GRCz11 were retrieved from the Ensembl database (Ensembl 100 release)144. CAGE-recovered TCs for which the dominant TSS is more than 500 bp away from the closest annotated TSS were excluded. To retrieve precise and stage specific TSS data, we considered the dominant (most frequently used) TSS positions for each of the 12 developmental stages as the 1st nucleotide position of the 5′ UTR. The nucleotide position immediately upstream of the ATG triplet of annotated coding sequence (CDS) start sites was considered the last nucleotide position of the 5′ UTR sequence. Coding transcripts lacking an annotated CDS start site or for which the dominant TSS falls downstream of the annotated CDS start site were excluded. Similarly, transcripts for which the dominant TSS falls within an intron were not considered for analysis, since we could not confidently determine the downstream used splice acceptor site. We assume these correspond to expressed splice variants. To identify differential TSS usage across samples, we scored “shifting promoters” (shifting score threshold = 0.6, K-S test, FDR ≤ 0.01) using the CAGEr package40. 5′ UTR sequences were extracted using BEDTools136.

5′ UTR library design and synthesis

The distance between the exit channel of the ribosome and the P site is ~12 nts, thus a 5′ UTR length of <15 nts is expected to produce a 48S PIC complex in which the m7G cap will be situated in the mRNA exit channel and hamper engagement of the exit channel with nucleotides 5′ of the AUG, impairing their translation6. As such, recovered 5′ UTRs shorter than 15 nts were filtered out. For transcripts undergoing switches in TSS usage, we included all switching unique 5′ UTR variants (n = 2,721) (Figure S1B). For transcripts with more than one dominant TSS within a promoter region but that did not score as switching, we considered the longest 5′ UTR isoform (n = 11,620) (Figure S1A). After filtering steps, the UTR library represented 10,354 genes and 11,445 transcripts. The interested reader can find all recovered 5′ UTRs in the Mendeley Data repository associated to this work. For 5′ UTRs ≤ 238 nt long (n = 9,863), the full “endogenous” sequence was considered. For 5′ UTRs > 238 nt long (3,446), sequences were split into equally long stretches flanked by 10 nt of flanking overlapping regions (Figure S1C). After splitting, the 5′ UTRs consisted of a total of 18,154 unique sequences ranging from 15 to 238 nt in length. For library amplification and cloning purposes, common upstream (PBS1, GAAGAGTAGCCTGCAGATAGAC, 22 nts) and downstream (PBS2, ATGGTGTCTAAAGGAGAGGAGC, 22 nts) sequences were introduced to flank each 5′ UTR and a third common sequence (PBS0, TGGTTGATTACGGTCGCA, 18 nt) was introduced at the most 3′ UTR of the sequence (Figure S1D). In cases where the original 5′ UTR was shorter that 238 nt in length, a random sequence (“filler”, in pink color) was added between PBS2 and PBS0 for a total length of 300 nts. The 300-nt-long pooled oligonucleotide library (n = 18,154) was synthesized by Twist Bioscience.

5′ UTR library cloning and mRNA synthesis

The library cloning strategy was adapted from the STARR-seq cloning protocol145. A pMB1vector (Twist Bioscience) custom construct was designed to include an SP6 promoter followed by a landing site for directional cloning of the 5′ UTR library, the sfGFP open reading frame and a 3′ UTR zebrafish sequence (BG1, 110 nt long) devoid of known regulatory sequence motifs33 (pMB1-BG1 vector) (Figure S1D and E). The synthetic oligonucleotide pool (Twist Bioscience) was amplified via the two primer handles (PBS0 and PBS1, 14 cycles) to yield a 300 bp long PCR product, followed by a low-cycle nested PCR (PBS1 and PBS2, 7 cycles) to generate the 5′ UTR library ranging from 15 (59 bp PCR product) to 238 nt-long (282 bp PCR product) sequences using the KAPA HiFi HotStart ReadyMix (Roche). The PCR product (5′ UTR insert pool) was purified using Angecourt AMPure XP beads (1.8 bead volume to PCR reaction volume) with the addition of isopropanol to 45%. The pMB1-BG1 vector was digested with BbsI-HF (NEB) to remove a filler sequence flanked by inverted BbsI restriction sequences and to allow for directional cloning. The digested vector was purified with the QIAquick PCR purification kit (QIAGEN) and purified again with the MinElute PCR purification (QIAGEN). The 5′ UTR insert pool was cloned into the BbsI-digested pMB1-BG1 vector between the SP6 sequence and the sfGFP ORF, to restore the 5′ UTR of the sfGFP. Homologous recombination directional cloning was performed using In-Fusion HD (Takara Bio). The cloned DNA library was electroporated into electrocompetent MegaX DH10B bacteria (Invitrogen) at a 100x library coverage. The colony lawn resulting from transformation was harvested and plasmids extracted using the Plasmid Plus Giga kit (QIAGEN). The plasmid library was PCR amplified using a SP6 forward primer (Fwd_IVT_SP6: 5′ CACGCATCTGGAATAAGGAAGTGC 3′) and a 3′ UTR-specific reverse primer containing a 36 nt-long T overhang (Rev_IVT_BG1orig: 5′ TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTCCTGTGAGTCCCATGGGTTTAAG 3′), followed by DpnI treatment (1h 37°C) and PCR product purification using Angecourt AMPure XP beads (0.8 bead volume to PCR reaction volume) (Figures S1D and S1E). The PCR product encompassed 99.9% (18,142/18,154) of the designed 5′ UTR sequences with lengths ranging from 15 to 238 nts (Figures S1F and S1G). This PCR template was used for in vitro transcription (mMessage Machine Sp6 kit, Thermo Fisher) by incubation on a thermocycler machine for 12h at 30°C. After incubation, DNase treatment was performed according to kit instructions (mMessage Machine SP6 Transcription kit, Thermo Fisher) and the IVT mRNA product was purified using the RNA Clean & Concentrator kit (Zymo Research). The resulting IVT product is a 5′ capped reporter mRNA pool consisting of variant 5′ UTRs driving translation of an invariant sfGFP ORF, the 3′ UTR and a 36 nt-long polyA tail (Figure S1H). The mRNA library faithfully represents 5′ UTR sequences and consists of almost the entirety of sequences synthesized. Transcript abundances in the in vitro transcribed mRNA library range from ~1 to ~400 TPM (25th percentile = 23 TPM; 50th percentile = 44 TPM; 75th percentile = 76 TPM) with shorter 5′ UTRs less well represented (Table S1), mirroring the less efficient cloning of shorter oligonucleotide sequences.

Image acquisition and quantification

Embryos were injected with 1 nL of an injection mix directly into the cell at the 1-cell stage with a microinjection needle (Sutter Instruments) and allowed to develop in standard conditions. For 5′ UTR mRNA library pool experiments, 80 pg/embryo were injected. For single-reporter injections, 40 pg 5′ UTR-sfGFP test reporter were co-injected with 2 ng fluorescent red dextran dye (D1868, Invitrogen) as injection control per embryo. Embryos were collected at the desired developmental stage and placed on a custom-made agarose mold with squared indents for placing and aligning the embryos. For fluorescence intensity quantification, zebrafish embryo images were acquired using an upright ZEISS Axiozoom coupled to an Axiocam 503 color/mono digital camera (14-bit depth) in black & white color mode with fixed laser power (red laser power 85% and 300 ms exposure; green laser power 65% and 300 ms exposure), fixed zoom and fixed exposure time for red mRF12 (590/612) and green AF488 (493/517) channels. Two rounds of single-reporter injections (embryos from two different clutches) were performed for fluorescence intensity quantifications. For switching UTR pairs, two needles were prepared in parallel and embryos from the same clutch were injected with either of the 5′ UTR isoform reporters to exclude possible differences in developmental time. Images were quantified using Fiji (Image J) using a macro for automated thresholding and channel fluorescence intensity measurement. For each image, the script automatically segments the image based on red channel intensity (control dextran dye) using FIJI’s auto thresholding tool (RenyiEntropy) and outputs csv files with mean fluorescence intensities of red and green channels for that region of interest (and a small invariant region for background correction). The mean fluorescence values outputted were then used for calculating normalized mean intensity ratios (sfGFP/dextran). A total of 25 embryos per injection round were quantified, for a total of 50 embryos per timepoint, per reporter. The FIJI macro, outputted csv files and excel files with normalized fluorescence intensity values calculated (related to Figures 2 H and 7G, H, J and K) are deposited in the Mendeley Data repository associated to this work.

Zebrafish embryo injections and polysome profiling

Embryos were injected at the 1-cell stage with 80 pg of the 5′ UTR mRNA library pool and allowed to develop in standard conditions. Embryos were staged and collected at the 64-cell (2 hpf), sphere (4 hpf), shied (6 hpf) and bud (10 hpf) stage. For each replicate sample, staged embryos were pooled and flash-frozen in liquid nitrogen and stored at −80°C. Three independent time-courses were performed (one time-course per replicate, on three different days), where in each day samples for the 2, 4, 6 and 10 hpf timepoints were collected. For each polysome profiling experiment, 100 embryos injected with 5′ UTR mRNA library were used. Immediately before polysome gradient preparation, embryos were lysed in 450 μL polysome gradient buffer (20 mM Tris-Cl pH 7.5, 30 mM MgCl2, 100 mM NaCl, 0.25 % Igepal-630 [v/v], 0.1% sodium deoxycholate (w/v), 100 μg/mL cycloheximide, 0.5 mM dithiothreitol [DTT], 1 mg/mL heparin and 40U/mL recombinant RNasin ribonuclease inhibitor [Promega]) with ~30 strokes of a motor-driven “B” pestle while on ice. After lysis, 7.2U of TURBO DNase (Invitrogen) were added and the embryo lysate was incubated for 15 min at 4°C on a rotating mixer. Samples were centrifuged at 13,000 rpm for 10 min at 4°C, and 400 μL of clarified lysate were loaded onto a 5-50% (w/v) sucrose gradient prepared in TMS buffer (20 mM Tris-Cl pH 7.5, 5 mM MgCl2, 140 mM NaCl, 100 μg/mL cycloheximide, 1 mM DTT). The remaining 50 μL of embryo lysate were kept as input sample. Gradients were centrifuged in a SW40 Ti rotor (Beckman) at 4°C for 150 min and profiles were analyzed using a Gradient Station (Biocomp Instruments) with continuous recording of optical density (OD) at 260 nm.

Polysome fraction processing

Fractions of 1200 μL were individually collected and SDS was immediately added to a final concentration of 1%. Fractions corresponding to the 80S peak (fractions 4 + 5), to low-molecular weight (fractions 6 + 7) or high-molecular weight (fractions 8 + 9) polysomes were pooled together in 15mL RNase-free tubes (Ambion) (Figure S2C). A volume of 2.4 mL TRI Reagent (Sigma) was added per tube, vortexed and stored at −80°C. To the withheld input sample (50 μL), 1 mL TRI Reagent (Sigma) was added, vortexed and stored at −80°C. For total RNA extraction, samples were thawed at room temperature. For fractionated samples, half of the sample (~2.4 mL) was transferred to a fresh RNase free 5mL Eppendorf tube and an extra 1.25 mL of TRI Reagent (Sigma) was added to each, for a final ratio of ~2:1 TRI Reagent:sample. Samples were mixed by vortexing and incubated at room temperature for 5 min. To each 5 mL tube, 500 μL of chloroform (without isoamyl alcohol or other additives) were added (or 200 μL to the input sample) and samples were vortexed. Samples were centrifuged at 12,000 g (an Eppendorf 5427 R centrifuge with a rotor for 5 mL tubes) for 25 min at 4°C. The aqueous phase was transferred to a fresh 5 mL Eppendorf tube and an equal volume of pure ethanol (> 99%) was added followed by vortexing. The Direct-zol RNA Miniprep kit (Zymo Research) was used for RNA purification. The mixture was transferred into a Zymo-Spin IICR Column in a collection tube and centrifuged for 30 sec at 15,000 rpm. The procedure was repeated until all sample was passed through the column (~ 850 μL at a time). The Direct-zol RNA Miniprep kit protocol was followed, excluding the DNase treatment step. The final sample was eluted in 25 μL DNase/RNase-free water, 2 μL of sample were collected for gel quality control and the remaining was stored at −80°C. One sample of total RNA from non-injected embryos was included as a negative control for the procedure described below.

Library preparation and high-throughput sequencing

A total of 50 sequencing libraries were prepared: replica A (input, fractions 4 + 5, 6+7 and 8+9) at 2, 4, 6 and 10 hpf, replica B (input, fractions 4 + 5, 6+7 and 8+9) at 2, 4, 6 and 10 hpf, replica C (input, fractions 4 + 5, 6+7 and 8+9) at 2, 4, 6 and 10 hpf, as well as libraries of the pooled PCR template used for library in vitro transcription and the uninjected 5′ UTR mRNA library to determine library complexity. For samples corresponding to replicas A-C and the uninjected 5′ UTR mRNA library, an RT reaction containing 11 μL of RNA sample and 1 μL transcript specific primer (2 μM RT_TSP primer: 5′ GTTCAGACGTGTGCTCTTCCGATCTGAACTTGTGGCCGTTCACGTCTCCAT 3′) targeting the beginning of the sfGFP open reading frame was performed (55°C for 1h, then 70°C for 15 min and hold at 4°C) using Superscript III Reverse Transcriptase (Thermo). This was followed by an RNase treatment to remove template mRNA by adding 0.2 μL RNase A (10 mg/ml; Thermo Scientific EN0531) per 20 μL RT reaction, and incubation for 1h at 37°C. The cDNA was purified with AMPure XP beads (1.8 vol beads to 1 vol cDNA volume) and eluted in 22 μL DNase/RNase-free water. Then, 20 μL pure cDNA sample were used for UMI introduction by linear PCR using KAPA HiFi HotStart ReadyMix (Roche) at the second strand synthesis step (P5_UMI_fwd: 5′ CACGACGCTCTTCCGATCTNNNNNNNNNNGAAGAGTAGCCTGCAGATAGA*C 3′, where * denotes a phosphorothioate bond), analogous to the procedure described in Neumayr et al., 2019. The AMPure XP beads purification step was repeated (0.8 vol beads to 1 vol PCR product) to remove excess primer and eluted in 22 μL DNase/RNase-free water. Library amplification was performed with primers containing TruSeq indexes comptible with Illumina high-throughput sequencing (Illumina i5: 5′ AATGATACGGCGACCACCGAGATCTACAC-i5-ACACTCTTTCCCTACACGACGCTCTTCCGATCT 3′ Illumina i7: 5′ CAAGCAGAAGACGGCATACGAGAT-i7-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT 3′). First, qPCR reactions using a fraction of the UMI-tagged PCR product as input (1.25 μL in 25 μL qPCR reaction) were performed to determine the ideal number of cycles used for amplification of each library. The qPCR reactions were ran using The KAPA HiFi HotStart ReadyMix (Roche) with EvaGreen Dye (Biotium). All samples were amplified in parallel (40 cycles), and the cycle number at which the middle of the exponential amplification phase was reached was noted (CNtest) for each sample. The final library amplification was performed using 10 μL UMI-tagged PCR product and the KAPA HiFi HotStart ReadyMix (Roche), using the same cycling conditions and a total number of cycles corresponding to CNtest - 3 tailored to each sample. The PCR product (ranging between 263-486 bp in length) was purified with AMPure XP beads (0.8 vol beads to 1 vol PCR product) and 5 μL were analysed by gel electrophoresis. For generating a library of the pooled PCR template (template for in vitro transcription of the 5′ UTR mRNA library), a PCR reaction targeting the variant 5′ UTR sequence of the 5′ UTR library plasmid DNA pool (Fwd_IVT_SP6: 5′ CACGCATCTGGAATAAGGAAGTGC 3′ and NGS_TSP_rev: 5′GTTCAGACGTGTGCTCTTCCGATCTGAACTTGTGGCCGTTCACGT3′) was performed. The PCR product was treated with DpnI and purified with AMPure XP beads (0.8 vol beads to 1 vol PCR product), followed by UMI introduction by linear PCR and library preparation as described above. Library quality control and quantification was performed using a 5300 Fragment Analyzer System (Agilent) and the HS NGS Fragment kit (1-6000 bp) (Agilent, DNF-474-0500). The 50 libraries were sequenced with 100 nt paired-end reads (PE100) with uneven loading for the forward read in a NovaSeq machine. Library quality control and high-throughput sequencing was performed by the Genomics Facility Basel (D-BSSE/ETHZ).

Sequence processing, TPM and RRS calculations

Sequencing data was filtered to retain only sequences with a UMI immediately followed by the common adaptor sequence (GAAGAGTAGCCTGCAGATAGAC, colored in yellow in Figure S1D) (see also Table S1, sheet sequencing_stats). UMIs were extracted using UMI-tools138. Adaptors (invariant common adaptor and the sfGFP sequences flanking the variant 5′ sequence) were partially trimmed using the non-internal adaptor trimming method of Cutadapt146, leaving each 5′ UTR sequence flanked by 10 invariant bps. Leaving parts of the invariant upstream and downstream flanking sequences allows to map 5′ UTR isoforms that differ in size but share much of their sequence. Up to 2 mismatches were allowed for the remaining common adaptor sequence and 3 for the remaining sfGFP sequence. Bowtie2137 was used to align retained reads to a reference set containing all 5′ UTRs in the synthetic oligonucleotide library (18,154 sequences) flanked by the common adaptor sequence (upstream) and sfGFP open reading frame sequence (downstream) using the “end-to-end” and “very-sensitive” parameters. Multi-mapped reads were removed by applying a MAPQ>=2 and flag 256 filter using SAM-tools147. Reads were deduplicated using UMItools138 and the number of reads mapped to each 5′ UTR entry were recorded. For determining TPM values, the number of mapped reads for each 5′ UTR entry were divided by the total number of mapped reads in the library and multiplied by 106 (+ one pseudo count). For determining RRS values, TPM values for each 5′ UTR entry in the library fraction sample were divided by the TPM values of the respective 5′ UTR entry in the input sample (see Figure 2C and S2C), in a manner analogous to that described in Niederer et al. 24. Because RRS is normalized to total reporter abundance at each timepoint, possible transcript-specific differences in mRNA stability or library preparation bias for different reporter lengths are factored out. A TPM ≥ 2 cutoff was applied for RRS calculations. RRS values were log2 transformed, and the arithmetic mean of log2RRS for each 5′ UTR entry of all three replicate samples was calculated. Entries for which at least one of the replicate values was NA were excluded.

Calculation of estimated abundances

Our estimated abundances quantify the relative change of each reporter with respect to its abundance in the uninjected pool of mRNAs (IVT library, which we termed 0 hpf). To make these values easier to interpret, we normalize changes in abundance at each time interval such that the most stable transcripts have an apparent change of zero. This is similar to previous normalization approaches in cell cultures under non-changing conditions 32,148, but here we allow the stable controls to be different at each time interval since the zebrafish embryo is a highly dynamic system. Specifically, we first defined a minimal TPM expression threshold for the control uninjected 5′ UTR mRNA library (TPM >3). Then, for each interval between timepoints ti−1 and ti (e.g. 0 and 2 hpf) we calculated the change in Δ log2TPMtiti−1 for all reporters, identified the 100 most stable transcripts (i.e. with the highest Δ log2 TPM) and subtracted Δ log2 TPM values of all other transcripts from the mean of these stable controls. The resulting “normalized” change in abundance Δ log2Xtiti−1 is lower or equal to zero, which reflects the intuitive notion that no transcripts should increase their abundance over time. Finally, we estimate abundances at each timepoint (log2Xti) by summing all the corrected changes in abundances from time zero to ti (Figure 3).

RNA secondary structure

RNA folding scores for each assayed 5′ UTR sequence were predicted by inputting 5′ UTR FASTA sequences into MXfold2 (https://github.com/mxfold/mxfold2, version mxfold2-0.1.1) employing the default parameters trained from TrainSetA and TrainSetB63. A local installation of NUPACK62 version 4.0.1.1 was used to calculate ensemble free energies, MFE secondary structures, and per-base unpaired probabilities for the analyses in Figures 3, S3, and S5. The RNA model was set up with a temperature of 28 °C. We included in the analysis the constant 5′ adapter present in our reporter, as well as the first 100 bases of the sfGFP CDS. Length-adjusted free energies relate the structure’s free energy to that of a 100 nt-long segment 64,93. These were calculated by dividing the ensemble free energies by the input sequence length (adapter length + variable 5′ UTR length + 100), and multiplying by 100.

uORF predictions

The number of predicted uORFs for each assayed 5′ UTR sequence was determined with the R Package ORFik (https://bioconductor.org/packages/release/bioc/html/ORFik.html)50. Upstream ORFs consisting of a uAUG and a downstream in-frame stop codon were searched in the 5′ to 3′ direction in all three possible reading frames using the “minimumLength=0” filter (START+STOP = 6 bp minimum ORF length).

Data subsetting for 5′ UTR sequence analysis

Data subsetting was performed by quantile analysis and by fuzzy-c-means clustering (Figure 4A). Deciles of log2 RRS values for each fraction (80S, LMW and HMW) at each developmental timepoint (2, 4, 6, and 10 hpf) were determined. 5′ UTR entries that were consistently found either in the top 10% or bottom 10% deciles across all timepoints for either the 80S, LMW or HMW fractions were filtered and used for motif-based sequence analysis (see Table S3). Fuzzy-c-means clustering was performed using the Mfuzz package84, using the weighted k-nearest neighbor method to replace missing values and setting the fuzzifier clustering parameter m to 2.502999. Values were filtered a posteriori by setting a membership score threshold of α ≥ 0.7. 5′ UTR entries displaying analogous changes in RRS across developmental time were grouped together for motif-based sequence analysis (see Figure S4C).

Motif-based sequence analysis

FASTA sequences of 5′ UTR entries in the upper quantiles, lower quantiles and fuzzy c-means clusters were generated. The XSTREME149 motif discovery and enrichment analysis (version 5.5.4) data submission form from the MEME suite online tool77 was used for extracting motifs (accessible via https://meme-suite.org/). Sequences obtained by quantile and clustering analysis were used for motif search, and sequences in the MPRA library (n = 18,154) were used as control sequences. XSTREME E-value threshold was set ≤ 0.05, motif width was set from 5 to 12 nts, the model of control sequences was used as background model and sequences were aligned to their right ends. STREME and MEME E-values were set to default. Output motifs and respective E-values are in Table S4 (sheet Motif_enrichment_analysis). For searching for identified motifs (Table S4) in 5′ UTR switching variants, motifs were uploaded onto MAST86 and searched using sequence E-value ≤ 10 threshold, removing motifs that are too similar to others and sorting motifs by best combined matches. In parallel, FIMO150 was used for individual motif scanning using a match p-value < 1x10−4 threshold. P-values of motifs matched presented in Figures 6, 7 and S7 correspond to individual motif p-values. Motif matches are in Table S4 (sheets MAST_motif_matches and FIMO_motif_matches). For MEME-derived and CNN-derived motif comparison (Figure S6A; Table S4 sheet Tomtom_comparison), Tomtom151 was ran using the Pearson correlation coefficient function and a E-value < 1000 threshold.

MRL calculation

The mean ribosome load (MRL) is meant to represent the ribosome loading of a transcript by a single number. MRL at each timepoint was calculated according to the equation in Figure 5A, where n80S, nLMW, and nHMW are estimated mean number of ribosomes in the 80S, LMW, and HMW fractions, respectively. In contrast to the original MRL calculation26, our definition uses the total RNA TPM in the denominator instead of the unweighted TPM sum across fractions.

DaniO5P length model

We fit second order polynomials of the form a2x2 + a1x + a0, where x is the UTR length, to each one of the following: log2MRLti for 2, 4, 6, and 10 hpf, and Δ log2Xtiti−1 for the time intervals 0-2 hpf, 2-4 hpf, 4-6 hpf, and 10-6 hpf. Polynomial parameters are shown in Table S4.

DaniO5P CNN model

We trained 10 convolutional neural network (CNN) models to predict, given an arbitrary input 5′ UTR sequence, the residuals from the length model (measurement – length model prediction) for all log2MRLti and Δlog2Xtiti−1 outputs. Input sequences were one-hot encoded with a maximum length of 238 nt, with shorter sequences padded to the left with zeros. The architecture of each CNN was a small VGG-16152: input sequences are passed through a series of convolutional blocks, each with two convolutional and one max pooling layer, followed by a fully connected layer and a linear layer with one output for each prediction. Hyperparameters were optimized individually based on performance on the validation set of chromosomal split 0 (see below). Final hyperparameter values were: number of convolutional blocks: 3, convolutional filter size: 7, number of filters in the first convolutional layer: 128, dropout in the convolutional layers: 0.1, number of units in the fully connected layer: 150, dropout in the fully connected layer: 0, all activations were ReLu. Model training was performed in python 3.9 using tensorflow 2.4 with the Adam optimizer, a mean square error penalty, and a decayed learning rate scheduler. Only sequences with a minimum TPM of 3 in the uninjected IVT library were used for CNN training (17,879 total).

To generate an ensemble of diverse models and make sure the generalization ability of each individual model is accurately assessed, a cross-fold training strategy based on chromosome splits was used: For each model, sequences from two or three chromosomes were held out from training (test set, 1.8k sequences on average), two or three more were used for early stopping (validation set, 1.8k sequences on average) and the remaining ones were used for training (14.3k sequences on average). This split was designed such that each chromosome was part of the test set of exactly one model, while trying to maintain the number of sequences in the training, validation, and test sets similar across models. Designing such splits is a version of the “multiway number partition problem”, for which we used the prtpy python package (https://github.com/coin-or/prtpy). Table S4 lists the chromosomes and number of sequences in the training, validation, and test set of each model. To avoid overfitting during training, performance on the validation set was evaluated after each epoch, and training was stopped when this performance failed to improve for 10 epochs. To evaluate model performance (Figures 5C and S5B), each individual model was used to generate predictions on its own test set. For all other calculations in this manuscript, “model output” refers to the average across all 10 ensemble models.

Additional models were attempted where the inputs were augmented with NUPACK-precomputed secondary structure metrics. In one attempt, we incorporated per-base unpaired probabilities as an additional input channel (input dimensions: 238x5). In the second model, we incorporated unpaired probabilities and a one-hot encoding of the dot-bracket minimum free energy structure (input dimensions: 238x8). Hyperparameters were separately optimized for each model. Resulting performance compared to the base CNN model can be found in Table S4.

Calculation of nucleotide contribution scores

We obtained contribution scores with respect to all 8 model outputs of all nucleotides within all 17,951 MPRA sequences used for model training and evaluation. To this end, we used a custom version of DeepSHAP (https://github.com/castillohair/shap/tree/castillohair/genomics_modcommit#f77513e2e05eb63f4d3b17ec9f9d3569c930ad02), which can work with tensorflow 2 and generate hypothetical contribution scores needed for motif discovery via TFModisco153 though this was not the motif discovery method we ultimately used (see below). As a background, we calculated the dinucleotide frequencies across all MPRA sequences and precomputed 25 random sequences for each length between 15 and 238. Nucleotide contributions for differences in outputs (e.g. log2(MRL_10hpf) – log2(MRL_2hpf)) were obtained by subtracting the corresponding output contributions.

CNN motif discovery

Motif discovery was performed in several steps. 1) Convolutional filter motif extraction: for each of the 128 convolutional filters in the second convolutional layer (first convolutional block, before the max pooling layer), we accumulated the 100 sequence fragments (seqlets) that resulted in maximal filter activation across all MPRA sequences, and used them to generate a filter position weight matrix (PWM). Seqlet length was set to 13, the receptive field after two convolutions with filter size 7. We required seqlets to be entirely contained within a sequence, therefore this analysis may not effectively find motifs influential at the very beginning or end of the 5′ UTR. This process resulted in 1280 filter motifs (128 motifs per model x 10 models). 2) Motif clustering: to account for motif redundancy within and across models, PWMs were clustered using a standalone version of RSAT matrix clustering154, modified by us to remove the ability to use PWM reverse complements (https://github.com/castillohair/matrix-clustering_stand-alone). To improve clustering results, we first separated motifs that contained AUGs from the rest (AUG likelihood threshold = 0.83333) and clustered these sets separately with different normalized correlation thresholds (AUG motifs: 0.45, non-AUG motifs: 0.6) but otherwise identical parameters (correlation threshold = 0, minimum number of aligned positions = 4, linkage method = average). This resulted in 6 AUG and 575 non-AUG motif clusters. 3) Motif filtering and post-processing: Reproducible and robust motifs should emerge independently in models trained on different data splits. Thus, motif clusters were required to contain motif filters originating from multiple individual models (9 or more for AUG motifs, 5 or more for non-AUG motifs), and those that did not were discarded. This resulted in 3 and 10 AUG and non-AUG motifs, respectively. Finally, motif cluster PWMs were trimmed at one base from the first position with information content greater than 0.3 on both sides to produce the final list of motifs.

Finally, we calculated contributions of each motif occurrence in all MPRA sequences as follows: first, we scanned motif cluster PWMs against all 5′ UTRs using FIMO with p-value < 1x10−4. Then, for each motif match, we summed contribution scores over the matching positions. Thus, we obtained for each motif, a set of motif matches across the MPRA library and their predicted contributions to each DaniO5P output. The average contributions of each motif can be found in Figure 5G, and comparisons of contributions with motif position and unpaired probabilities are in Figure S5GI.

Quantification and statistical analysis

Statistical analyses and plotting

Statistical analyses relative to plots presented in Figures 3, 6CE, S2D, S3, S6E and S6G were performed in R version 4.3.1140. Statistical analyses relative to plots presented in Figure 2H were performed in GraphPad Prism. Hierarchical clustering (Figures 2D and S2E) was performed with the R package pheatmap (https://cran.r-project.org/web/packages/pheatmap/pheatmap.pdf), Fuzzy c-means clustering was performed with the R package Mfuzz84 and motif logos were generated with the R package SeqLogo155. Statistical analyses relative to CAGE data were performed using the R package CAGEr40. Motif p-values, E-values and PWMs were outputted from MEME suite tool77 analyses. For CNN discovered motifs, PWMs were generated in Python with the package logomaker (https://logomaker.readthedocs.io/en/latest/, version 0.8). Analyses and plots relative to Figures 5, S5 and S6HI were generated in Python 3.9, with the exception of nucleotide contributions (DeepSHAP) which used Python 3.6. All other plots were generated using R or GraphPad Prism. Final plots were beautified on Adobe Illustrator.

Supplementary Material

1
2
3
4
5
6

Key Resources Table

REAGENT or RESOURCE SOURCE IDENTIFIER
Bacterial and virus strains
MegaX DH10B T1R Electrocomp Cells Invitrogen C640003
Chemicals, peptides, and recombinant proteins
TRI-Reagent® Sigma Aldrich T9424
AMPure XP beads Agencourt A63881
SuperScript III Reverse Transcriptase Thermo Fisher Scientific 18080044
KAPA HiFi HotStart ReadyMix Roche KK2602
BbsI-HF NEB R3539S
Dextran, Tetramethylrhodamine, 10000 MW Invitrogen D1868
Critical commercial assays
In-Fusion HD Cloning Kit Takara Bio 639650
Plasmid Plus Giga kit QIAGEN 12991
MinElute PCR purification kit QIAGEN 28004
mMessage Machine SP6 Transcription kit Thermo Fisher Scientific AM1340
RNA Clean & Concentrator kit Zymo Research R1016
Direct-zol RNA Miniprep kit Zymo Research R2052
Deposited data
5′ UTR mRNA library sequencing data This paper SRA data: PRJNA1056276
Embryo imaging data and fluorescence intensity calculations This paper Mendeley Data Repository: 10.17632/rf98wbhzgf.1
DaniO5P model weights This paper Mendeley Data Repository: 10.17632/rf98wbhzgf.1
NUPACK structure predictions This paper Mendeley Data Repository: 10.17632/rf98wbhzgf.1
Experimental models: Organisms/strains
Zebrafish (Danio rerio): TL/AB N/A N/A
Oligonucleotides
5′ UTR oligo library (18,154 sequences, see Table S1) This paper N/A
PBS1_adapter: 5′ GAAGAGTAGCCTGCAGATAGAC 3′ This paper N/A
PBS0_rev: 5′ TGCGACCGTAATCAACCA 3′ This paper N/A
PBS2_sfGFP: 5′ GCTCCTCTCCTTTAGACACCAT 3′ This paper N/A
Fwd_IVT_SP6: 5′ CACGCATCTGGAATAAGGAAGTGC 3′
Rev_IVT_BG1orig: 5′ TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTCCTGTGAGTCCCATGGGTTTAAG 3′ This paper N/A
NGS_TSP_rev: 5′ GTTCAGACGTGTGCTCTTCCGATCTGAACTTGTGGCCGTTCACGT 3′ This paper N/A
RT_TSP: 5′ GTTCAGACGTGTGCTCTTCCGATCTGAACTTGTGGCCGTTCACGTCTCCAT 3′ This paper N/A
P5_UMI_fwd (* = phosphorothioate bond, N’s = UMI): 5′ CACGACGCTCTTCCGATCTNNNNNNNNNNGAAGAGTAGCCTGCAGATAGA*C 3′ This paper N/A
Illumina i5: 5′ AATGATACGGCGACCACCGAGATCTACAC-i5-ACACTCTTTCCCTACACGACGCTCTTCCGATCT 3′ This paper N/A
Illumina i7: 5′ CAAGCAGAAGACGGCATACGAGAT-i7-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT 3′ This paper N/A
Recombinant DNA
Plasmid: pTwist Amp High Copy (pMB1vector) for MPRA library cloning Twist Bioscience Mendeley Data Repository: 10.17632/rf98wbhzgf.1
Software and algorithms
CAGEr Haberle et al. 40 https://www.bioconductor.org/packages/release/bioc/html/CAGEr.html
BEDTools suite Quinlan and Hall 136 https://bedtools.readthedocs.io/en/latest/content/bedtools-suite.html
Bowtie2 Langmead and Salzberg 137 https://bowtie-bio.sourceforge.net/bowtie2/index.shtml
UMI-tools Smith et al. 138 https://umi-tools.readthedocs.io/en/latest/
ORFik Tjeldnes et al. 50 https://www.bioconductor.org/packages/release/bioc/html/ORFik.html
MXfold2 Sato et al. 63 https://github.com/mxfold/mxfold2
NUPACK Fornace et al. 62 https://www.nupack.org/
Mfuzz Kumar and M 84 https://www.bioconductor.org/packages/release/bioc/html/Mfuzz.html
MEME suite Bailey et al. 77 https://meme-suite.org/meme/
DaniO5P This paper https://github.com/castillohair/DaniO5P
Other
Fiji macro for automated imaging data analysis This paper Mendeley Data Repository: 10.17632/rf98wbhzgf.1
Mapped high-throughput sequencing reads This paper Mendeley Data Repository: 10.17632/rf98wbhzgf.1

Highlights.

  • In vivo MPRA systematically interrogates the regulatory roles of endogenous 5′ UTRs

  • 5′ UTRs are sufficient to regulate ribosome recruitment during zebrafish embryogenesis

  • The MPRA identifies 5′ UTR cis-regulatory sequence motifs

  • The DaniO5P model quantifies the contribution of 5′ UTR features to translation

Acknowledgments

We thank Dr. Sunil Shetty for polysome profiling training and assistance collecting polysome fractions; Dr. Andrea Pauli and her lab for advice regarding polysome profiling; Dr. Michal Rabani for feedback regarding initial MPRA data analysis; Dr. Matthias Muhar, Dr. Júlia Batki and Dr. Annika Nichols for advice throughout the project and for critical reading of the manuscript; the Schier lab and the Basel RNA Club for input and discussions; the Biozentrum Core Facilities for technical support and fish husbandry. High-throughput sequencing was performed at the Genomics Facility Basel. This project was funded by the Marie Skłodowska-Curie Individual Fellowship (No 898218) under the H2020 research and innovation programme and the EMBO long-term fellowship (ALTF 691-2019) awarded to M.M.R-P., the Allen Discovery Center for Cell Lineage Tracing and the H2020 ERC Advanced grant (No 834788) to A.F.S., and the NIH Award R33CA255893 and NSF Award 2021552 to G. S.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Declaration of interests

G.S. is on the SAB of Modulus Therapeutics and is a co-founder of Parse Biosciences. A.F.S is in the advisory board of Developmental Cell. The authors declare no other competing interests.

References

  • 1.Teixeira FK, and Lehmann R (2019). Translational Control during Developmental Transitions. Cold Spring Harb Perspect Biol 11. 10.1101/cshperspect.a032987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Vastenhouw NL, Cao WX, and Lipshitz HD (2019). The maternal-to-zygotic transition revisited. Development 146. 10.1242/dev.161471. [DOI] [PubMed] [Google Scholar]
  • 3.Yang G, Xin Q, and Dean J (2024). Degradation and translation of maternal mRNA for embryogenesis. Trends Genet 40, 238–249. 10.1016/j.tig.2023.12.008. [DOI] [PubMed] [Google Scholar]
  • 4.Merrick WC, and Pavitt GD (2018). Protein Synthesis Initiation in Eukaryotic Cells. Cold Spring Harb Perspect Biol 10. 10.1101/cshperspect.a033092. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Araujo PR, Yoon K, Ko D, Smith AD, Qiao M, Suresh U, Burns SC, and Penalva LO (2012). Before It Gets Started: Regulating Translation at the 5′ UTR. Comp Funct Genomics 2012, 475731. 10.1155/2012/475731. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Hinnebusch AG (2011). Molecular mechanism of scanning and start codon selection in eukaryotes. Microbiol Mol Biol Rev 75, 434–467, first page of table of contents. 10.1128/MMBR.00008-11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Hinnebusch AG, Ivanov IP, and Sonenberg N (2016). Translational control by 5′-untranslated regions of eukaryotic mRNAs. Science 352, 1413–1416. 10.1126/science.aad9868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Leppek K, Das R, and Barna M (2018). Functional 5′ UTR mRNA structures in eukaryotic translation regulation and how to find them. Nat Rev Mol Cell Biol 19, 158–174. 10.1038/nrm.2017.103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Renz PF, Valdivia-Francia F, and Sendoel A (2020). Some like it translated: small ORFs in the 5′UTR. Exp Cell Res 396, 112229. 10.1016/j.yexcr.2020.112229. [DOI] [PubMed] [Google Scholar]
  • 10.Hentze MW, Castello A, Schwarzl T, and Preiss T (2018). A brave new world of RNA-binding proteins. Nat Rev Mol Cell Biol 19, 327–341. 10.1038/nrm.2017.130. [DOI] [PubMed] [Google Scholar]
  • 11.Corley M, Burns MC, and Yeo GW (2020). How RNA-Binding Proteins Interact with RNA: Molecules and Mechanisms. Mol Cell 78, 9–29. 10.1016/j.molcel.2020.03.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Cheng Z, Otto GM, Powers EN, Keskin A, Mertins P, Carr SA, Jovanovic M, and Brar GA (2018). Pervasive, Coordinated Protein-Level Changes Driven by Transcript Isoform Switching during Meiosis. Cell 172, 910–923 e916. 10.1016/j.cell.2018.01.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Haberle V, Li N, Hadzhiev Y, Plessy C, Previti C, Nepal C, Gehrig J, Dong X, Akalin A, Suzuki AM, et al. (2014). Two independent transcription initiation codes overlap on vertebrate core promoters. Nature 507, 381–385. 10.1038/nature12974. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Nepal C, Hadzhiev Y, Previti C, Haberle V, Li N, Takahashi H, Suzuki AM, Sheng Y, Abdelhamid RF, Anand S, et al. (2013). Dynamic regulation of the transcription initiation landscape at single nucleotide resolution during vertebrate embryogenesis. Genome Res 23, 1938–1950. 10.1101/gr.153692.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Despic V, and Neugebauer KM (2018). RNA tales - how embryos read and discard messages from mom. J Cell Sci 131. 10.1242/jcs.201996. [DOI] [PubMed] [Google Scholar]
  • 16.Sterne-Weiler T, Martinez-Nunez RT, Howard JM, Cvitovik I, Katzman S, Tariq MA, Pourmand N, and Sanford JR (2013). Frac-seq reveals isoform-specific recruitment to polyribosomes. Genome Res 23, 1615–1623. 10.1101/gr.148585.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Floor SN, and Doudna JA (2016). Tunable protein synthesis by transcript isoforms in human cells. Elife 5. 10.7554/eLife.10921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Blair JD, Hockemeyer D, Doudna JA, Bateup HS, and Floor SN (2017). Widespread Translational Remodeling during Human Neuronal Differentiation. Cell Rep 21, 2005–2016. 10.1016/j.celrep.2017.10.095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wong QW, Vaz C, Lee QY, Zhao TY, Luo R, Archer SK, Preiss T, Tanavde V, and Vardy LA (2016). Embryonic Stem Cells Exhibit mRNA Isoform Specific Translational Regulation. PLoS One 11, e0143235. 10.1371/journal.pone.0143235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wang X, Hou J, Quedenau C, and Chen W (2016). Pervasive isoform-specific translational regulation via alternative transcription start sites in mammals. Mol Syst Biol 12, 875. 10.15252/msb.20166941. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Dieudonne FX, O’Connor PB, Gubler-Jaquier P, Yasrebi H, Conne B, Nikolaev S, Antonarakis S, Baranov PV, and Curran J (2015). The effect of heterogeneous Transcription Start Sites (TSS) on the translatome: implications for the mammalian cellular phenotype. BMC Genomics 16, 986. 10.1186/s12864-015-2179-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Arribere JA, and Gilbert WV (2013). Roles for transcript leaders in translation and mRNA decay revealed by transcript leader sequencing. Genome Res 23, 977–987. 10.1101/gr.150342.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Tresenrider A, Morse K, Jorgensen V, Chia M, Liao H, van Werven FJ, and Unal E (2021). Integrated genomic analysis reveals key features of long undecoded transcript isoform-based gene repression. Mol Cell 87,2231–2245 e2211. 10.1016/j.molcel.2021.03.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Niederer RO, Rojas-Duran MF, Zinshteyn B, and Gilbert WV (2022). Direct analysis of ribosome targeting illuminates thousand-fold regulation of translation initiation. Cell Syst 13, 256–264 e253. 10.1016/j.cels.2021.12.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Cuperus JT, Groves B, Kuchina A, Rosenberg AB, Jojic N, Fields S, and Seelig G (2017). Deep learning of the regulatory grammar of yeast 5′ untranslated regions from 500,000 random sequences. Genome Res 27, 2015–2024. 10.1101/gr.224964.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Sample PJ, Wang B, Reid DW, Presnyak V, McFadyen IJ, Morris DR, and Seelig G (2019). Human 5′ UTR design and variant effect prediction from a massively parallel translation assay. Nat Biotechnol 37, 803–809. 10.1038/s41587-019-0164-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Castillo-Hair S, Fedak S, Wang B, Linder J, Havens K, Certo M, and Seelig G (2024). Optimizing 5′UTRs for mRNA-delivered gene editing using deep learning. Nat Commun 15, 5284. 10.1038/S41467-024-49508-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Strayer EC, Krishna S, Lee H, Vejnar C, Beaudoin J-D, and Giraldez AJ (2023). NaP-TRAP, a novel massively parallel reporter assay to quantify translation control. bioRxiv, 2023.2011.2009.566434. 10.1101/2023.11.09.566434. [DOI] [Google Scholar]
  • 29.May GE, Akirtava C, Agar-Johnson M, Micic J, Woolford J, and McManus J (2023). Unraveling the influences of sequence and position on yeast uORF activity using massively parallel reporter systems and machine learning. Elife 12. 10.7554/eLife.69611. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Dvir S, Velten L, Sharon E, Zeevi D, Carey LB, Weinberger A, and Segal E (2013). Deciphering the rules by which 5′-UTR sequences affect protein expression in yeast. Proc Natl Acad Sci U S A 110, E2792–2801. 10.1073/pnas.1222534110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Lim Y, Arora S, Schuster SL, Corey L, Fitzgibbon M, Wladyka CL, Wu X, Coleman IM, Delrow JJ, Corey E, et al. (2021). Multiplexed functional genomic analysis of 5′ untranslated region mutations across the spectrum of prostate cancer. Nat Commun 12, 4217. 10.1038/S41467-021-24445-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Jia L, Mao Y, Ji Q, Dersh D, Yewdell JW, and Qian SB (2020). Decoding mRNA translatability and stability from the 5′ UTR. Nat Struct Mol Biol 27, 814–821. 10.1038/s41594-020-0465-x. [DOI] [PubMed] [Google Scholar]
  • 33.Rabani M, Pieper L, Chew GL, and Schier AF (2017). A Massively Parallel Reporter Assay of 3′ UTR Sequences Identifies In Vivo Rules for mRNA Degradation. Mol Cell 68, 1083–1094 e1085. 10.1016/j.molcel.2017.11.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Vejnar CE, Abdel Messih M, Takacs CM, Yartseva V, Oikonomou P, Christiano R, Stoeckius M, Lau S, Lee MT, Beaudoin JD, et al. (2019). Genome wide analysis of 3′ UTR sequence elements and proteins regulating mRNA stability during maternal-to-zygotic transition in zebrafish. Genome Res 29, 1100–1114. 10.1101/gr.245159.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Xiang K, Ly J, and Bartel DP (2024). Control of poly(A)-tail length and translation in vertebrate oocytes and early embryos. Dev Cell 59, 1058–1074 e1011. 10.1016/j.devcel.2024.02.007. [DOI] [PubMed] [Google Scholar]
  • 36.Akirtava C, May GE, and McManus CJ (2022). False-positive IRESes from Hoxa9 and other genes resulting from errors in mammalian 5′ UTR annotations. Proc Natl Acad Sci U S A 119, e2122170119. 10.1073/pnas.2122170119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Baranasic D, Hortenhuber M, Balwierz PJ, Zehnder T, Mukarram AK, Nepal C, Varnai C, Hadzhiev Y, Jimenez-Gonzalez A, Li N, et al. (2022). Multiomic atlas with functional stratification and developmental dynamics of zebrafish cis-regulatory elements. Nat Genet 54, 1037–1050. 10.1038/S41588-022-01089-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Lawson ND, Li R, Shin M, Grosse A, Yukselen O, Stone OA, Kucukural A, and Zhu L (2020). An improved zebrafish transcriptome annotation for sensitive and comprehensive detection of cell type-specific genes. Elife 9. 10.7554/eLife.55792. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Shiraki T, Kondo S, Katayama S, Waki K, Kasukawa T, Kawaji H, Kodzius R, Watahiki A, Nakamura M, Arakawa T, et al. (2003). Cap analysis gene expression for high-throughput analysis of transcriptional starting point and identification of promoter usage. Proc Natl Acad Sci U S A 100, 15776–15781. 10.1073/pnas.2136655100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Haberle V, Forrest AR, Hayashizaki Y, Carninci P, and Lenhard B (2015). CAGEr: precise TSS data retrieval and high-resolution promoterome mining for integrative analyses. Nucleic Acids Res 43, e51. 10.1093/nar/gkv054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Subtelny AO, Eichhorn SW, Chen GR, Sive H, and Bartel DP (2014). Poly(A)-tail profiling reveals an embryonic switch in translational control. Nature 508, 66–71. 10.1038/nature13007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Marlow FL (2020). Setting up for gastrulation in zebrafish. Curr Top Dev Biol 136, 33–83. 10.1016/bs.ctdb.2019.08.002. [DOI] [PubMed] [Google Scholar]
  • 43.Pinheiro D, and Heisenberg CP (2020). Zebrafish gastrulation: Putting fate in motion. Curr Top Dev Biol 136, 343–375. 10.1016/bs.ctdb.2019.10.009. [DOI] [PubMed] [Google Scholar]
  • 44.Leesch F, Lorenzo-Orts L, Pribitzer C, Grishkovskaya I, Roehsner J, Chugunova A, Matzinger M, Roitinger E, Belacic K, Kandolf S, et al. (2023). A molecular network of conserved factors keeps ribosomes dormant in the egg. Nature 613, 712–720. 10.1038/s41586-022-05623-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.He L, Binari R, Huang J, Falo-Sanjuan J, and Perrimon N (2019). In vivo study of gene expression with an enhanced dual-color fluorescent transcriptional timer. Elife 8. 10.7554/eLife.46181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Calvo SE, Pagliarini DJ, and Mootha VK (2009). Upstream open reading frames cause widespread reduction of protein expression and are polymorphic among humans. Proc Natl Acad Sci U S A 106, 7507–7512. 10.1073/pnas.0810916106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Chew GL, Pauli A, and Schier AF (2016). Conservation of uORF repressiveness and sequence features in mouse, human and zebrafish. Nat Commun 7, 11663. 10.1038/ncomms11663. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Johnstone TG, Bazzini AA, and Giraldez AJ (2016). Upstream ORFs are prevalent translational repressors in vertebrates. EMBO J 35, 706–723. 10.15252/embj.201592759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Giess A, Torres Cleuren YN, Tjeldnes H, Krause M, Bizuayehu TT, Hiensch S, Okon A, Wagner CR, and Valen E (2020). Profiling of Small Ribosomal Subunits Reveals Modes and Regulation of Translation Initiation. Cell Rep 31, 107534. 10.1016/j.celrep.2020.107534. [DOI] [PubMed] [Google Scholar]
  • 50.Tjeldnes H, Labun K, Torres Cleuren Y, Chyzynska K, Swirski M, and Valen E (2021). ORFik: a comprehensive R toolkit for the analysis of translation. BMC Bioinformatics 22, 336. 10.1186/s12859-021-04254-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Bohlen J, Fenzl K, Kramer G, Bukau B, and Teleman AA (2020). Selective 40S Footprinting Reveals Cap-Tethered Ribosome Scanning in Human Cells. Mol Cell 79, 561–574 e565. 10.1016/j.molcel.2020.06.005. [DOI] [PubMed] [Google Scholar]
  • 52.Kozak M. (1991). Effects of long 5′ leader sequences on initiation by eukaryotic ribosomes in vitro. Gene Expr 1, 117–125. [PMC free article] [PubMed] [Google Scholar]
  • 53.Kozak M. (1988). Leader length and secondary structure modulate mRNA function under conditions of stress. Mol Cell Biol 8, 2737–2744. 10.1128/mcb.8.7.2737-2744.1988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Chappell SA, Edelman GM, and Mauro VP (2006). Ribosomal tethering and clustering as mechanisms for translation initiation. Proc Natl Acad Sci U S A 103, 18077–18082. 10.1073/pnas.0608212103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Bairn SB, and Sherman F (1988). mRNA structures influencing translation in the yeast Saccharomyces cerevisiae. Mol Cell Biol 8, 1591–1601. 10.1128/mcb.8.4.1591-1601.1988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Kozak M. (1986). Influences of mRNA secondary structure on initiation by eukaryotic ribosomes. Proc Natl Acad Sci U S A 83, 2850–2854. 10.1073/pnas.83.9.2850. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Pelletier J, and Sonenberg N (1985). Insertion mutagenesis to increase secondary structure within the 5′ noncoding region of a eukaryotic mRNA reduces translational efficiency. Cell 40, 515–526. 10.1016/0092-8674(85)90200-4. [DOI] [PubMed] [Google Scholar]
  • 58.Wang J, Shin BS, Alvarado C, Kim JR, Bohlen J, Dever TE, and Puglisi JD (2022). Rapid 40S scanning and its regulation by mRNA structure during eukaryotic translation initiation. Cell 185, 4474–4487 e4417. 10.1016/j.cell.2022.10.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Babendure JR, Babendure JL, Ding JH, and Tsien RY (2006). Control of mammalian translation by mRNA structure near caps. RNA 12, 851–861. 10.1261/rna.2309906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Kozak M. (1989). Circumstances and mechanisms of inhibition of translation by secondary structure in eucaryotic mRNAs. Mol Cell Biol 9, 5134–5142. 10.1128/mcb.9.11.5134-5142.1989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Ringner M, and Krogh M (2005). Folding free energies of 5′-UTRs impact post-transcriptional regulation on a genomic scale in yeast. PLoS Comput Biol 1, e72. 10.1371/journal.pcbi.0010072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Fornace ME, Huang J, Newman CT, Porubsky NJ, Pierce MB, and Pierce NA (2022). NUPACK: analysis and design of nucleic acid structures, devices, and systems. ChemRxiv, 10.26434/chemrxiv-2022-xv98l, 2022. [DOI] [Google Scholar]
  • 63.Sato K, Akiyama M, and Sakakibara Y (2021). RNA secondary structure prediction using deep learning with thermodynamic integration. Nat Commun 12, 941. 10.1038/S41467-021-21194-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Trotta E (2014). On the normalization of the minimum free energy of RNAs by sequence length. PLoS One 9, e113380. 10.1371/journal.pone.0113380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Kozak M. (1987). An analysis of 5′-noncoding sequences from 699 vertebrate messenger RNAs. Nucleic Acids Res 15, 8125–8148. 10.1093/nar/15.20.8125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Kozak M. (1986). Point mutations define a sequence flanking the AUG initiator codon that modulates translation by eukaryotic ribosomes. Cell 44, 283–292. 10.1016/0092-8674(86)90762-2. [DOI] [PubMed] [Google Scholar]
  • 67.Nakagawa S, Niimura Y, Gojobori T, Tanaka H, and Miura K (2008). Diversity of preferred nucleotide sequences around the translation initiation codon in eukaryote genomes. Nucleic Acids Res 36, 861–871. 10.1093/nar/gkm1102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Grzegorski SJ, Chiari EF, Robbins A, Kish PE, and Kahana A (2014). Natural variability of Kozak sequences correlates with function in a zebrafish model. PLoS One 9, e108475. 10.1371/journal.pone.0108475. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Hernandez G, Osnaya VG, and Perez-Martinez X (2019). Conservation and Variability of the AUG Initiation Codon Context in Eukaryotes. Trends Biochem Sci 44, 1009–1021. 10.1016/j.tibs.2019.07.001. [DOI] [PubMed] [Google Scholar]
  • 70.Bazzini AA, Del Viso F, Moreno-Mateos MA, Johnstone TG, Vejnar CE, Qin Y, Yao J, Khokha MK, and Giraldez AJ (2016). Codon identity regulates mRNA stability and translation efficiency during the maternal-to-zygotic transition. EMBO J 35, 2087–2103. 10.15252/embj.201694699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Mishima Y, and Tomari Y (2016). Codon Usage and 3′ UTR Length Determine Maternal mRNA Stability in Zebrafish. Mol Cell 61, 874–885. 10.1016/j.molcel.2016.02.027. [DOI] [PubMed] [Google Scholar]
  • 72.Chan LY, Mugler CF, Heinrich S, Vallotton P, and Weis K (2018). Non-invasive measurement of mRNA decay reveals translation initiation as the major determinant of mRNA stability. Elife 7. 10.7554/eLife.32536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Schwartz DC, and Parker R (1999). Mutations in translation initiation factors lead to increased rates of deadenylation and decapping of mRNAs in Saccharomyces cerevisiae. Mol Cell Biol 19, 5247–5256. 10.1128/MCB.19.8.5247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Bicknell AA, Reid DW, Licata MC, Jones AK, Cheng YM, Li M, Hsiao CJ, Pepin CS, Metkar M, Levdansky Y, et al. (2024). Attenuating ribosome load improves protein output from mRNA by limiting translation-dependent mRNA decay. Cell Rep 43, 114098. 10.1016/j.celrep.2024.114098. [DOI] [PubMed] [Google Scholar]
  • 75.Mercier BC, Labaronne E, Cluet D, Guiguettaz L, Fontrodona N, Bicknell A, Corbin A, Wencker M, Aube F, Modolo L, et al. (2024). Translation-dependent and -independent mRNA decay occur through mutually exclusive pathways defined by ribosome density during T cell activation. Genome Res 34, 394–409. 10.1101/gr.277863.123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Dave P, Roth G, Griesbach E, Mateju D, Hochstoeger T, and Chao JA (2023). Single-molecule imaging reveals translation-dependent destabilization of mRNAs. Mol Cell 83, 589–606 e586. 10.1016/j.molcel.2023.01.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Bailey TL, Johnson J, Grant CE, and Noble WS (2015). The MEME Suite. Nucleic Acids Res 43, W39–49. 10.1093/nar/gkv416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Ray D, Kazan H, Cook KB, Weirauch MT, Najafabadi HS, Li X, Gueroussov S, Albu M, Zheng H, Yang A, et al. (2013). A compendium of RNA-binding motifs for decoding gene regulation. Nature 499, 172–177. 10.1038/nature12311. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Gerstberger S, Hafner M, and Tuschl T (2014). A census of human RNA-binding proteins. Nat Rev Genet 15, 829–845. 10.1038/nrg3813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Dominguez D, Freese P, Alexis MS, Su A, Hochman M, Palden T, Bazile C, Lambert NJ, Van Nostrand EL, Pratt GA, et al. (2018). Sequence, Structure, and Context Preferences of Human RNA Binding Proteins. Mol Cell 70, 854–867 e859. 10.1016/j.molcel.2018.05.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Lambert N, Robertson A, Jangi M, McGeary S, Sharp PA, and Burge CB (2014). RNA Bind-n-Seq: quantitative assessment of the sequence and structural binding specificity of RNA binding proteins. Mol Cell 54, 887–900. 10.1016/j.molcel.2014.04.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Van Nostrand EL, Freese P, Pratt GA, Wang X, Wei X, Xiao R, Blue SM, Chen JY, Cody NAL, Dominguez D, et al. (2020). A large-scale binding and functional map of human RNA-binding proteins. Nature 583, 711–719. 10.1038/s41586-020-2077-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Futschik ME, and Carlisle B (2005). Noise-robust soft clustering of gene expression time-course data. J Bioinform Comput Biol 3, 965–988. 10.1142/s0219720005001375. [DOI] [PubMed] [Google Scholar]
  • 84.Kumar L, and M EF, (2007). Mfuzz: a software package for soft clustering of microarray data. Bioinformation 2, 5–7. 10.6026/97320630002005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Liao SE, Sudarshan M, and Regev O (2023). Deciphering RNA splicing logic with interpretable machine learning. Proc Natl Acad Sci U S A 120, e2221165120. 10.1073/pnas.2221165120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Bailey TL, and Gribskov M (1998). Combining evidence using p-values: application to sequence homology searches. Bioinformatics 14, 48–54. 10.1093/bioinformatics/14.1.48. [DOI] [PubMed] [Google Scholar]
  • 87.Lin Y, May GE, Kready H, Nazzaro L, Mao M, Spealman P, Creeger Y, and McManus CJ (2019). Impacts of uORF codon identity and position on translation regulation. Nucleic Acids Res 47, 9358–9367. 10.1093/nar/gkz681. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Zhang H, Wang Y, and Lu J (2019). Function and Evolution of Upstream ORFs in Eukaryotes. Trends Biochem Sci 44, 782–794. 10.1016/j.tibs.2019.03.002. [DOI] [PubMed] [Google Scholar]
  • 89.Shabalina SA, Ogurtsov AY, Rogozin IB, Koonin EV, and Lipman DJ (2004). Comparative analysis of orthologous eukaryotic mRNAs: potential hidden functional signals. Nucleic Acids Res 32, 1774–1782. 10.1093/nar/gkh313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Despic V, Dejung M, Gu M, Krishnan J, Zhang J, Herzel L, Straube K, Gerstein MB, Butter F, and Neugebauer KM (2017). Dynamic RNA-protein interactions underlie the zebrafish maternal-to-zygotic transition. Genome Res 27, 1184–1194. 10.1101/gr.215954.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Shirokikh NE, and Preiss T (2018). Translation initiation by cap-dependent ribosome recruitment: Recent insights and open questions. Wiley Interdiscip Rev RNA 9, e1473. 10.1002/wrna.1473. [DOI] [PubMed] [Google Scholar]
  • 92.Livingston NM, Kwon J, Valera O, Saba JA, Sinha NK, Reddy P, Nelson B, Wolfe C, Ha T, Green R, et al. (2023). Bursting translation on single mRNAs in live cells. Mol Cell 83, 2276–2289 e2211. 10.1016/j.molcel.2023.05.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Zhang BH, Pan XP, Cox SB, Cobb GP, and Anderson TA (2006). Evidence that miRNAs are different from other RNAs. Cell Mol Life Sci 63, 246–254. 10.1007/S00018-005-5467-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Beaudoin JD, Novoa EM, Vejnar CE, Yartseva V, Takacs CM, Kellis M, and Giraldez AJ (2018). Analyses of mRNA structure dynamics identify embryonic gene regulatory programs. Nat Struct Mol Biol 25, 677–686. 10.1038/s41594-018-0091-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Xiang Y, Huang W, Tan L, Chen T, He Y, Irving PS, Weeks KM, Zhang QC, and Dong X (2023). Pervasive downstream RNA hairpins dynamically dictate start-codon selection. Nature 621, 423–430. 10.1038/s41586-023-06500-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Gonzalez-Sanchez AM, Castellanos-Silva EA, Diaz-Figueroa G, and Cate JHD (2024). JUN mRNA translation regulation is mediated by multiple 5′ UTR and start codon features. PLoS One 19, e0299779. 10.1371/journal.pone.0299779. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Gebauer F, Schwarzl T, Valcarcel J, and Hentze MW (2021). RNA-binding proteins in human genetic disease. Nat Rev Genet 22, 185–198. 10.1038/s41576-020-00302-y. [DOI] [PubMed] [Google Scholar]
  • 98.Matia-Gonzalez AM, Laing EE, and Gerber AP (2015). Conserved mRNA-binding proteomes in eukaryotic organisms. Nat Struct Mol Biol 22, 1027–1033. 10.1038/nsmb.3128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Beckmann BM, Horos R, Fischer B, Castello A, Eichelbaum K, Alleaume AM, Schwarzl T, Curk T, Foehr S, Huber W, et al. (2015). The RNA-binding proteomes from yeast to man harbour conserved enigmRBPs. Nat Commun 6, 10127. 10.1038/ncomms10127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Lavallee-Adam M, Cloutier P, Coulombe B, and Blanchette M (2017). Functional 5′ UTR motif discovery with LESMoN: Local Enrichment of Sequence Motifs in biological Networks. Nucleic Acids Res 45, 10415–10427. 10.1093/nar/gkx751. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Zheng W. Fong JHC, Wan YK, Chu AHY, Huang Y. Wong ASL, and Ho JWK. (2023). Discovery of regulatory motifs in 5′ untranslated regions using interpretable multi-task learning models. Cell Syst 14, 1103–1112 e1106. 10.1016/j.cels.2023.10.011. [DOI] [PubMed] [Google Scholar]
  • 102.Nielsen J, Christiansen J, Lykke-Andersen J, Johnsen AH, Wewer UM, and Nielsen FC (1999). A family of insulin-like growth factor II mRNA-binding proteins represses translation in late development. Mol Cell Biol 19, 1262–1270. 10.1128/MCB.19.2.1262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Dai N, Rapley J, Angel M, Yanik MF, Blower MD, and Avruch J (2011). mTOR phosphorylates IMP2 to promote IGF2 mRNA translation by internal ribosomal entry. Genes Dev 25, 1159–1172. 10.1101/gad.2042311. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Dai N, Zhao L, Wrighting D, Kramer D, Majithia A, Wang Y, Cracan V, Borges-Rivera D, Mootha VK, Nahrendorf M, et al. (2015). IGF2BP2/IMP2-Deficient mice resist obesity through enhanced translation of Ucp1 mRNA and Other mRNAs encoding mitochondrial proteins. Cell Metab 21, 609–621. 10.1016/j.cmet.2015.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Ji X, Jha A, Humenik J, Ghanem LR, Kromer A, Duncan-Lewis C, Traxler E, Weiss MJ, Barash Y, and Liebhaber SA (2021). RNA-Binding Proteins PCBP1 and PCBP2 Are Critical Determinants of Murine Erythropoiesis. Mol Cell Biol 41, e0066820. 10.1128/MCB.00668-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Smirnova VV, Shestakova ED, Bikmetov DV, Chugunova AA, Osterman IA, Serebryakova MV, Sergeeva OV, Zatsepin TS, Shatsky IN, and Terenin IM (2019). eIF4G2 balances its own mRNA translation via a PCBP2-based feedback loop. RNA 25, 757–767. 10.1261/rna.065623.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Kelley RL, Wang J, Bell L, and Kuroda MI (1997). Sex lethal controls dosage compensation in Drosophila by a non-splicing mechanism. Nature 387, 195–199. 10.1038/387195a0. [DOI] [PubMed] [Google Scholar]
  • 108.Bashaw GJ, and Baker BS (1997). The regulation of the Drosophila msl-2 gene reveals a function for Sex-lethal in translational control. Cell 89, 789–798. 10.1016/s0092-8674(00)80262-7. [DOI] [PubMed] [Google Scholar]
  • 109.Gebauer F, Merendino L, Hentze MW, and Valcarcel J (1998). The Drosophila splicing regulator sex-lethal directly inhibits translation of male-specific-lethal 2 mRNA. RNA 4, 142–150. [PMC free article] [PubMed] [Google Scholar]
  • 110.Sanford JR, Gray NK, Beckmann K, and Caceres JF (2004). A novel role for shuttling SR proteins in mRNA translation. Genes Dev 18, 755–768. 10.1101/gad.286404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Maslon MM, Heras SR, Bellora N, Eyras E, and Caceres JF (2014). The translational landscape of the splicing factor SRSF1 and its role in mitosis. Elife 3, e02028. 10.7554/eLife.02028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Palangat M, Anastasakis DG, Fei DL, Lindblad KE, Bradley R, Hourigan CS, Hafner M, and Larson DR (2019). The splicing factor U2AF1 contributes to cancer progression through a noncanonical role in translation regulation. Genes Dev 33, 482–497. 10.1101/gad.319590.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Ueno T, Taga Y, Yoshimoto R, Mayeda A, Hattori S, and Ogawa-Goto K (2019). Component of splicing factor SF3b plays a key role in translational control of polyribosomes on the endoplasmic reticulum. Proc Natl Acad Sci U S A 116, 9340–9349. 10.1073/pnas.1901742116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Long JC, and Caceres JF (2009). The SR protein family of splicing factors: master regulators of gene expression. Biochem J 417, 15–27. 10.1042/BJ20081501. [DOI] [PubMed] [Google Scholar]
  • 115.Sysoev VO, Fischer B, Frese CK, Gupta I, Krijgsveld J, Hentze MW, Castello A, and Ephrussi A (2016). Global changes of the RNA-bound proteome during the maternal-to-zygotic transition in Drosophila. Nat Commun 7, 12128. 10.1038/ncomms12128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Rojas-Duran MF, and Gilbert WV (2012). Alternative transcription start site selection leads to large differences in translation activity in yeast. RNA 18, 2299–2305. 10.1261/rna.035865.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Chen J, Tresenrider A, Chia M, McSwiggen DT, Spedale G, Jorgensen V, Liao H, van Werven FJ, and Unal E (2017). Kinetochore inactivation by expression of a repressive mRNA. Elife 6. 10.7554/eLife.27417. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Capell A, Fellerer K, and Haass C (2014). Progranulin transcripts with short and long 5′ untranslated regions (UTRs) are differentially expressed via posttranscriptional and translational repression. J Biol Chem 289, 25879–25889. 10.1074/jbc.M114.560128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Hollerer I, Barker JC, Jorgensen V, Tresenrider A, Dugast-Darzacq C, Chan LY, Darzacq X, Tjian R, Unal E, and Brar GA (2019). Evidence for an Integrated Gene Repression Mechanism Based on mRNA Isoform Toggling in Human Cells. G3 (Bethesda) 9, 1045–1053. 10.1534/g3.118.200802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Brar GA, Yassour M, Friedman N, Regev A, Ingolia NT, and Weissman JS (2012). High-resolution view of the yeast meiotic program revealed by ribosome profiling. Science 335, 552–557. 10.1126/science.1215110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Weber R, Ghoshdastider U, Spies D, Dure C, Valdivia-Francia F, Forny M, Ormiston M, Renz PF, Taborsky D, Yigit M, et al. (2023). Monitoring the 5′UTR landscape reveals isoform switches to drive translational efficiencies in cancer. Oncogene 42, 638–650. 10.1038/s41388-022-02578-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Sugimoto Y, and Ratcliffe PJ (2022). Isoform-resolved mRNA profiling of ribosome load defines interplay of HIF and mTOR dysregulation in kidney cancer. Nat Struct Mol Biol 29, 871–880. 10.1038/s41594-022-00819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Aeschimann F, Kumari P, Bartake H, Gaidatzis D, Xu L, Ciosk R, and Grosshans H (2017). LIN41 Post-transcriptionally Silences mRNAs by Two Distinct and Position-Dependent Mechanisms. Mol Cell 65, 476–489 e474. 10.1016/j.molcel.2016.12.010. [DOI] [PubMed] [Google Scholar]
  • 124.Popovitchenko T, Park Y, Page NF, Luo X, Krsnik Z, Liu Y, Salamon I, Stephenson JD, Kraushar ML, Volk NL, et al. (2020). Translational derepression of Elavl4 isoforms at their alternative 5′ UTRs determines neuronal development. Nat Commun 11, 1674. 10.1038/s41467-020-15412-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Karollus A, Avsec Z, and Gagneur J (2021). Predicting mean ribosome load for 5′UTR of any length using deep learning. PLoS Comput Biol 17, e1008982. 10.1371/journal.pcbi.1008982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Chen J, Hu Z, Sun S, Tan Q, Wang Y, Yu Q, Zong L, Hong L, Xiao J, Shen T, et al. (2022). Interpretable RNA Foundation Model from Unannotated Data for Highly Accurate RNA Structure and Function Predictions. arXiv:2204.00300v5. 10.48550/arXiv.2204.00300. [DOI] [Google Scholar]
  • 127.Sun C, Shrivastava A, Singh S, and Gupta A (2017). Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. arXiv:1707.02968v2. 10.48550/arXiv.1707.02968. [DOI] [Google Scholar]
  • 128.Friedman RZ, Ramu A, Lichtarge S, Wu Y, Tripp L, Lyon D, Myers CA, Granas DM, Gause M, Corbo JC, et al. (2024). Active learning of enhancer and silencer regulatory grammar in a developing neural tissue. bioRxiv, 2023.2008.2021.554146. 10.1101/2023.08.21.554146. [DOI] [Google Scholar]
  • 129.Linder J, and Seelig G (2021). Fast activation maximization for molecular sequence design. BMC Bioinformatics 22, 510. 10.1186/s12859-021-04437-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Linder J, Bogard N, Rosenberg AB, and Seelig G (2020). A Generative Neural Network for Maximizing Fitness and Diversity of Synthetic DNA and Protein Sequences. Cell Syst 11, 49–62 e16. 10.1016/j.cels.2020.05.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Castillo-Hair SM, and Seelig G (2022). Machine Learning for Designing Next-Generation mRNA Therapeutics. Acc Chem Res 55, 24–34. 10.1021/acs.accounts.1c00621. [DOI] [PubMed] [Google Scholar]
  • 132.Meyuhas O, and Kahan T (2015). The race to decipher the top secrets of TOP mRNAs. Biochim Biophys Acta 1849, 801–811. 10.1016/j.bbagrm.2014.08.015. [DOI] [PubMed] [Google Scholar]
  • 133.Theil K, Herzog M, and Rajewsky N (2018). Post-transcriptional Regulation by 3′ UTRs Can Be Masked by Regulatory Elements in 5′ UTRs. Cell Rep 22, 3217–3226. 10.1016/j.celrep.2018.02.094. [DOI] [PubMed] [Google Scholar]
  • 134.Seo KW, and Kleiner RE (2021). Mechanisms of epitranscriptomic gene regulation. Biopolymers 112, e23403. 10.1002/bip.23403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Meyer KD, Patil DP, Zhou J, Zinoviev A, Skabkin MA, Elemento O, Pestova TV, Qian SB, and Jaffrey SR (2015). 5′ UTR m(6)A Promotes Cap-Independent Translation. Cell 163, 999–1010. 10.1016/j.cell.2015.10.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Quinlan AR, and Hall IM (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Langmead B, and Salzberg SL (2012). Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357–359. 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Smith T, Heger A, and Sudbery I (2017). UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res 27, 491–499. 10.1101/gr.209601.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Kimmel CB, Ballard WW, Kimmel SR, Ullmann B, and Schilling TF (1995). Stages of embryonic development of the zebrafish. Dev Dyn 203, 253–310. 10.1002/aja.1002030302. [DOI] [PubMed] [Google Scholar]
  • 140.Team RC (2021). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna. https://www.R-project.org. [Google Scholar]
  • 141.Gentleman RC, Carey VJ, Bates DM, Bolstad B, Dettling M, Dudoit S, Ellis B, Gautier L, Ge Y, Gentry J, et al. (2004). Bioconductor: open software development for computational biology and bioinformatics. Genome Biol 5, R80. 10.1186/gb-2004-5-10-r80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Balwierz PJ, Carninci P, Daub CO, Kawai J, Hayashizaki Y, Van Belle W, Beisel C, and van Nimwegen E (2009). Methods for analyzing deep sequencing expression data: constructing the human and mouse promoterome with deepCAGE data. Genome Biol 10, R79. 10.1186/gb-2009-10-7-r79. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Frith MC, Valen E, Krogh A, Hayashizaki Y, Carninci P, and Sandelin A (2008). A code for transcription initiation in mammalian genomes. Genome Res 18, 1–12. 10.1101/gr.6831208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Kinsella RJ, Kahari A, Haider S, Zamora J, Proctor G, Spudich G, Almeida-King J, Staines D, Derwent P, Kerhornou A, et al. (2011). Ensembl BioMarts: a hub for data retrieval across taxonomic space. Database (Oxford) 2011, bar030. 10.1093/database/bar030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Neumayr C, Pagani M, Stark A, and Arnold CD (2019). STARR-seq and UMI-STARR-seq: Assessing Enhancer Activities for Genome-Wide-, High-, and Low-Complexity Candidate Libraries. Curr Protoc Mol Biol 128, e105. 10.1002/cpmb.105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Martin M (2011). Cutadapt removes adapter sequences from high-throughput sequencing reads. 2011 17, 3. 10.14806/ej.17.1.200. [DOI] [Google Scholar]
  • 147.Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, and Genome Project Data Processing, S. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079. 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Viegas JO, Fishman L, Meshorer E, and Rabani M (2023). Calculating RNA degradation rates using large-scale normalization in mouse embryonic stem cells. STAR Protoc 4, 102534. 10.1016/j.xpro.2023.102534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Grant CE, and Bailey TL (2021). XSTREME: Comprehensive motif analysis of biological sequence datasets. bioRxiv, 2021.2009.2002.458722. 10.1101/2021.09.02.458722. [DOI] [Google Scholar]
  • 150.Grant CE, Bailey TL, and Noble WS (2011). FIMO: scanning for occurrences of a given motif. Bioinformatics 27, 1017–1018. 10.1093/bioinformatics/btr064. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Gupta S, Stamatoyannopoulos JA, Bailey TL, and Noble WS (2007). Quantifying similarity between motifs. Genome Biol 8, R24. 10.1186/gb-2007-8-2-r24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Simonyan K, and Zisserman A (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. 10.48550/arXiv.1409.1556. [DOI] [Google Scholar]
  • 153.Shrikumar A, Tian K, Avsec Ž, Shcherbina A, Banerjee A, Sharmin M, Nair S, and Kundaje A (2018). Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5. 10.48550/arXiv.1811.00416. [DOI] [Google Scholar]
  • 154.Castro-Mondragon JA, Jaeger S, Thieffry D, Thomas-Chollier M, and van Helden J (2017). RSAT matrix-clustering: dynamic exploration and redundancy reduction of transcription factor binding motif collections. Nucleic Acids Res 45, e119. 10.1093/nar/gkx314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Bembom O, and Ivanek R (2023). seqLogo: Sequence logos for DNA sequence alignments. https://bioconductor.org/packages/seqLogo [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
3
4
5
6

Data Availability Statement

Hight-throughput sequencing datasets have been deposited in the Sequence Read Archive repository (SRA) (BioProject identifier PRJNA1056276, link https://www.ncbi.nlm.nih.gov/sra/PRJNA1056276). Code for the DaniO5P model is accessible at https://github.com/castillohair/DaniO5P.

Any additional information required to reanalyze the data reported in this work paper is available from the Lead Contact upon request.

RESOURCES