Abstract
Streptomycetes, Gram-positive bacteria with huge and GC-rich genomes provide an ample example of codon usage bias taken to the extreme. Particularly, in all sequenced to date streptomycete genomes leucyl codon TTA is the rarest one. It is present (usually once or twice) in 70–200 out of 7000–8000 coding sequences that make up a typical streptomycete genome. tRNALeuUAA of streptomycetes, encoded by the bldA gene, has been shown to be present in mature form only after the onset of morphological differentiation and activation of secondary metabolism. Consequently, during the early stages of cell growth, the translation of genes carrying the TTA codon can be interrupted due to the absence of tRNALeuUAA. Several reports show that mutations of TTA to synonymous codons in certain genes indeed relieve their expression from bldA dependence. However, the deletion of bldA does not always arrest the expression of TTA-containing genes. The nucleotides T/C downstream of TTA were suggested, in 2002, to favor TTA mistranslation. We tested this hypothesis using sizable datasets derived from individual Streptomyces genome and a subset of TTA+ genes for secondary metabolism known for their active expression. Our results revealed nucleotide biases downstream of NNA codons family, such as the preference for C and the avoidance of A. Yet, none of the observed biases was sufficient to claim a special case for TTA codon. Hence, the issue of codon context and TTA codon mistranslation in Streptomyces deserves further elaboration.
Electronic supplementary material
The online version of this article (10.1007/s12088-020-00902-6) contains supplementary material, which is available to authorized users.
Keywords: bldA-regulation, Streptomyces, UUA codon, Morphological differentiation
Introduction
In most organisms there is synonymous codon usage bias. The mechanisms that favor unequal usage of codons are a subject of much discussion. Mutational bias due to mechanistic forces and translational selection are both thought to contribute to this phenomenon [1–5]. In this regard, Streptomyces—a genus within the class Actinobacteria—is an interesting case of extremely rare usage of leucyl codon TTA [6–8]. TTA codon in this species has evolved into a component of regulatory switch that appears to control (among others) several salient phenotypic traits, such as sporulation and production of antibiotics [6]. Streptomycetes are therefore an attractive model to explore various factors that influence codon usage.
Studies of morphological mutants of Streptomyces coelicolor initiated in late 1960s [9] eventually revealed gene bldA for tRNALeuUAA as an important player in colony development and antibiotic production [10–12]. It has to be noted that TTA (decoded by bldA tRNA) is the rarest codon in GC-rich actinobacterial genomes [8] which is absent in essential genes. In the model strain S. coelicolor A(3)2, whose genome harbors 142 TTA-containing (TTA+) genes, bldA deficiency arrests aerial mycelium formation (the so called “bald” phenotype) and causes major-to-infinite delay of the onset of secondary metabolite production presumably due to being unable to properly translate the UUA-containing mRNAs [13]. These effects are mediated by a few TTA+ genes encoding pleiotropic regulators and regulators dedicated to antibiotic production. Functions of the majority of TTA+ genes in S. coelicolor and other species are not understood.
While bldA regulation is a well-documented phenomenon [14–16], there is evidence that certain TTA+ genes can escape translational arrest in bldA mutants and be expressed at significant levels [17, 18]. The existence of a subset of genes that are TTA+ but are still translated in bldA deficient mutants is paradoxical in the light of the currently accepted mechanism of bldA-mediated regulation. In 2002, Trepanier and coauthors [17] offered a theoretical explanation for this discordance. Briefly, they attribute the difference in expression of TTA+ genes in bldA mutants to the nature of the first nucleotide downstream of the TTA codon (hereafter referred to as the N1 position; see [19]). If either A or G follow TTA, then the ribosome is able to mistranslate the codon in-frame. In contrast, if the N1 is either a C or T, the frameshifting occurs, most likely resulting in abortive translation. Thus, in the absence of tRNALeuUAA, only mRNAs containing the quadruplets (N1N2N3|N1) TTAC/T will not lead to a functional protein. Currently the mechanisms of codon-specific reading frame maintenance are thoroughly studied for proline codons, and they hinge on proper post-transcriptional modification of tRNA isoacceptors and nature of nucleotide downstream of CCY codons [20]. However, there are no reports that such mechanisms operate on leucine codons. At the time of Trepanier’s publication, there was a very limited set of TTA+ genes to substantiate this context-based hypothesis [21]. Here we investigate this issue using extensive datasets and focusing on several functional classes of genes. We show that the nucleotide N1 is biased in a set of NNA codons in general, and no special or a more pronounced codon context bias is present for TTA. Our results imply that further bioinformatics and experimental investigations of the TTA codon context and mistranslation both within and beyond the genus Streptomyces are worthwhile efforts.
Methods
Sources of Data
The annotated sequences of selected Streptomyces genomes were obtained at the NCBI website (ftp://ftp.ncbi.nih.gov/genomes/Bacteria/). Streptomyces antibiotic biosynthetic genes were taken from MIBiG database (https://mibig.secondarymetabolites.org/) and screened for TTA codons. Summary of genomic sequences used in this study can be reviewed in Electronic Supplementary Materials (ESM, Table S1).
Analysis of TTA and Its Vicinity
Computational biology approaches, used to analyze dicodon usage in Anaconda software [22, 23], were adopted in this study. Perl scripts utilizing BioPerl modules [24] were written to search and analyze TTA codons and proximal regions. Residual dicodon frequencies were calculated as detailed in [22]. Briefly, the entire sets of annotated coding sequences of Streptomyces genomes were downloaded from GenBank and used to generate with Anaconda a di-codon counts table. The observed counts were subjected to statistical analysis to reveal those codon pairs that occurred either more (“preferred”) or less (“rejected”) frequently than would be expected for random co-occurrence of two codons. This analysis relies on z-score-type Pearson test to derive residual dicodon values. The value was considered significant (e.g., it can be construed as evidence of preference or rejection for a given di-codon) if it absolute value is greater than 3.
Input data (di-codon counts) used to arrive at Table 1 were obtained from analysis of Streptomyces ORFomes with Anaconda software; they are given in supplementary Excel Table S2. For each class of the dicodon (N1N2N3|N1N2N3) we calculated the distribution of nucleotides for each combination of codon positions and compared it with the distribution expected from background nucleotide frequencies, assuming Bernoulli process of nucleotide occurence.
Table 1.
Statistical analysis of nucleotide combination frequencies within dicodons of streptomycete ORFomes
| Factorsa | Sum of squares | Mean of squares | F value |
|---|---|---|---|
| N1|N1 | 26.4 | 1.76 | 8.64 |
| N1|N2 | 24.5 | 1.63 | 7.99 |
| N1|N3 | 40.6 | 2.71 | 13.53 |
| N2|N1 | 42.9 | 2.86 | 14.34 |
| N2|N2 | 13.8 | 0.92 | 4.45 |
| N2|N3 | 38.4 | 2.56 | 12.77 |
| N3|N1 | 211 | 14.07 | 90.01 |
| N3|N2 | 69.9 | 4.66 | 24.21 |
| N3|N3 | 38.8 | 2.59 | 12.89 |
Raw data are available in supplementary Excel Table S3 (global dicodon counts)
aNi|Nj—(nucleotide at position i in the first codon)|(nucleotide at position j in the second codon; d.f. =15 for all entries; Pr(> F) = 2 × 10− 16 except for N2|N2 (< 1.99 × 10− 8 for the latter)
Results
Global Pattern of N3|N1 Nucleotide Context at the Codon Boundary in Streptomyces
A general picture of codon contexts within Streptomyces genomes is needed prior to the analysis of TTA context. We resorted to a computational approach of Santos and coworkers (2005), who developed specialized bioinformatic application, Anaconda, to analyze primary genome sequences [22]. In this approach a null model assumes that there is no codon context bias, as expected for the random codon pair distribution, taking into account background nucleotide and codon frequencies. Such a codon pair would be shown as a black pixel on heat map (no bias). All deviations from the null model captured by the application would be independent of GC content and codon frequencies. Complete genomes of 79 experimentally studied streptomycetes (ESM Table S1) were used to calculate residual dicodon frequencies and then to infer nucleotide contexts. The data obtained from the analysis of 79 Streptomyces ORFomes were averaged and represented in the form of heat map (Fig. 1). The map depicts dependence of N1 nucleotide of second codon on the nature of N3 nucleotide of the first codon. Pronounced deviations from the null model (no context bias) were seen when N1 is either A or C.
Fig. 1.

N3–N1 nucleotide context map deduced from dicodon analysis of 79 streptomycete ORFomes. Logs of dicodon frequencies were the input data to build the map. Rows and columns were sorted to group codons ending with a particular nucleotide (N3; rows) and next codons starting with a particular nucleotide (N1; columns). Green cells correspond to preferred contexts (occur more frequently as it would be expected from background nucleotide frequencies) and red cells to rejected ones. Values that are not statistically significant are colored black (see also the main text). The color scale represents the full range of values of residuals for streptomycete codon context (color figure online)
We analyzed the distribution of all combinations of nucleotides Ni in dicodons to reveal the most significant contributors to N3|N1 context bias. The data, summarized in Table 1, showed that each nucleotide combination is significantly biased within the dicodon. However, N3|N1 frequency is the biggest contributor to the observed context bias.
No Distinct Differences in N1 Frequencies Downstream of TTA Codons, as Compared to Control Datasets
Next we focused on the analysis of TTA context within the analyzed Streptomyces genomes. First, we asked what codons are most similar to TTA in their context bias. Clustering analysis revealed that TTA, GTA and CTA are the closest in this regard (Fig. 2). The frequencies of quadruplets occurrence based on the aforementioned codons were found (Table 2), and four-sample test for equality of proportions has been carried out on this dataset [25]. The rationale behind this analysis was as follows. To reveal a selective force that skews N1 usage downstream of TTA, one has to analyze TTAN quadruplet proportions to the closest (in terms of the context) quadruplets which are not thought as regulatory ones. Our null hypothesis is that the four N1 counts have the same proportions regardless of the preceding codon. As a result, it was revealed that, on a genome-wide scale, proportions of TTAN and CTAN quadruplets are not significantly different (p < 0.05). Note that GTAN quadruplets have different usage proportions compared to TTAN.
Fig. 2.

Clustering analysis of N1 codon contexts. The dendrogram was generated using the neighbor-joining approach for quadruplets, as compared to intragenic position of all TTA codons in six streptomycete genomes
Table 2.
Counts of occurrence of TTA-, GTA- and CTA-based quadruplets in Streptomyces ORFomes
| A | T | G | C | |
|---|---|---|---|---|
| TTA | 18 | 25 | 45 | 143 |
| CTA | 59 | 86 | 193 | 638 |
| GTA | 204 | 420 | 676 | 5591 |
Sum of given TTA-based quadruplet found with 79 genomes divided by 79, and averagedSum of given TTA-based quadruplet found with 79 genomes divided by 79, and averaged
We further attempted to analyze quadruplet biases via a different approach, based on the analysis of TTA population from a single best studied genome of S. coelicolor. There are 152 TTA codons in the latter. We generated a control set of genes lacking TTA (TTA- genes) that match Streptomyces coelicolor TTA+ genes in terms of genomic location and overall gene length. Distributions of the quadruplets were calculated for these datasets and their significance was assessed using nonparametric Tukey’s test. These data are represented in Figs. S1 and S2, Electronic Supplementary Materials (ESM). Overall, results of this analysis agreed with the first approach, i.e., some biases were observed, particularly for C-ending quadruplets, but those were not TTA codon-based.
No Significant Variance in Codon and Quadruplet Intragenic Position
Since the global quadruplet frequency values weren’t informative, we modified a previously described method [26] to assess the evolutionary pressure on coding sequences based on intragenic location, to examine the pressure on TTAN tetraplets. In the absence of positional evolutionary pressure, , and under the conditions of active bldA-regulation, the quadruplets TTAN would be randomly scattered across coding sequences, so that the cell would waste more resources producing long, potentially toxic, nonfunctional peptides early in its life cycle. However, it has already been demonstrated that the intragenic location of TTA is biased towards the start codon, indicating that positional constraints act at least on the codon alone [8, 26]. On a genomic scale, appears to be a significant factor that alleviates the metabolic burden of the cell under the conditions of bldA-regulation. Hence, from the initial observation that TTA codons are location-biased within genes, we wanted to go step further and assess whether the pool of quadruplets TTA(C/T) is more positionally biased than TTA(G/A) in genomes of streptomycetes in which bldA-regulation was shown to take place. If so, it would imply that selects not only for TTA location, but also for its context.
On a genome-wide scale, the random distribution of codons within genes will be graphically represented as a diagonal line on the plot (see “expected” line on Fig. 3a). As the location of TTAs becomes more biased towards the start codon, it increases the area under the curve representing the cumulative codon (or quadruplet) frequency (Fig. 3). We reasoned here that the difference in areas under the expected and observed curves is a result of action of positional evolutionary pressure , which can be approximated using the trapezoidal rule:
where N is the number of equally-spaced subsequences into which the studied coding sequences were divided. In this way we can measure and compare for different codons or quadruplets in the genomes. Under conditions of no positional pressure Δω = 0, then f(x)obs would be equal to f(x)exp. Any features (quadruplets in our case) biased towards either the start or stop codons will have either positive or negative Δω values, respectively. We computed Δω values for 79 Streptomyces genomes; no clear-cut pattern can be deduced for the distribution of bases in N1 when analyzing TTA+ pool genome-wide (Fig. 3b).
Fig. 3.
No significant differences in intragenic positions of TTA-based quadruplets in streptomycete genomes. Examples of genome-wide skews of TTA and associated quadruplets are shown for S. albus J1074 and S. coelicolor M145 (a). Different symbols indicate different quadruplets (see graphical inset in the figure). Null distribution of codons (no bias; Exp) is indicated by clear circles. Box plot chart summarizing the distribution of Δω values for 79 Streptomyces genomes (b). Error bars indicate confidence interval (95%)
Focusing on the Genes for Secondary Metabolism
The failure to reveal TTA codon context patterns can be attributed to the nature of the dataset. Namely, little is known about the function(s) of the majority of TTA+ genes in S. coelicolor [27]; most of them shows no marginal expression level (see Fig. S3, ESM, and [28]). As long as we cannot ascribe a certain molecular or visual phenotype to majority of TTA+ genes, it is possible that many of them are under no positional evolutionary pressure. We therefore decided to analyze a subset of TTA+ genes that are actively expressed. Genes involved in antibiotic production fit our requirements. These genes form a quasi-random group (i.e. they control biosynthesis of numerous structurally and biosynthetically unrelated small molecules), many of them contain TTA, and it is possible to cherry pick those involved in the active production of natural products (thus warranting that corresponding genes are efficiently transcribed and translated). Along these lines of reasoning, we screened sequenced Streptomyces antibiotic biosynthesis gene clusters deposited in MIBiG database [29] for the presence of TTA codons (and TTAN quadruplets) in their genes. We revealed 810 TTA+ genes, of which 170 were annotated as transcriptional regulators. The lists of genes are given in ESM Excel Table S3. We noted that some selected genes contained more than one TTA, both as TTAT/C and TTAA/G. However, the presence of a single TTAT/C along with TTAA/G quadruplets could be enough to ensure that a gene is under bldA-regulation. Out 1066 TTA codons present in 810 genes set, 828 (77%) were in the form of TTAT/C (617 out of 828 were TTAC). Similar pattern was observed for 170 regulatory genes (75% TTAT/C). These numbers are in line with the overall occurrence of TTAN quadruplets in Streptomyces genomes (see Table 2).
Discussion
In this work we took bioinformatics approach to test the possibility that UUA translation in the absence of cognate tRNALeuUAA can be influenced by the nature of a nucleotide directly downstream of TTA. When analyzing the aggregated protein-coding sequences of 79 streptomycete genomes, we revealed biased usage of C downstream of NNA codons: that is, C is used more often than the null model would predict. Therefore, TTAC quadruplets indeed appear to be more frequent than expected, but this phenomenon is likely not caused by selective pressure acting specifically on that codon. For the pool of TTA+ genes from S. coelicolor, we observed no fourth base bias following TTA codons and there was no apparent intragenic positional bias in any of the TTA-based quadruplets. TTAA is an exception from the said above, although this result is probably due to its rarity in the GC-rich S. coelicolor genome. Nevertheless, all TTA+ genes of S. coelicolor are not an optimal sample, since most of these genes are functionally silent [27, 28], and their codon usage/intragenic bias might be distorted for many reasons. We therefore focused on a subset of TTA+ genes containing only regulatory and structural genes for antibiotic production by Streptomyces. We picked genes that control production of known metabolites, so that extraneous biases in silent genes could be ruled out. Here we also detected shifts to C/T downstream of TTA, but to the extent not larger than that observed for other NNA codons. Overall, our results are consistent with the initial assumption of Trepanier and coworkers [17]. However, an important amendment that our work makes is that these biases are a part of a larger phenomenon.
Analysis of sizable and diverse collections of streptomycete coding sequences reveals little-understood dicodon biases affecting wide set of codons, including TTA. The biases might be caused by a myriad of selective constraints imposed on flow of genetic information– from replication to translation [22, 30]. We note that some of these biases were observed already in other non-Streptomyces bacteria, such as rare U3|A1 pairs (see Fig. 1 and [31, 32]), while largely favored A3|C1 context appears to be a peculiar feature of Streptomyces. It is interesting to define codon context rules in this genus and understand their underlying forces. Further explanations of TTA codon mistranslation are to be sought. For example, more complex forms of codon context can be at play, so that UUA translation can be influenced by the triplets immediately down- and upstream of UUA. An example of such scenario has been recently demonstrated for the Salmonella flgM gene [33].
Electronic Supplementary Material
Below is the link to the electronic supplementary material
Acknowledgements
B.O. thanks for Grant support of the National Research Fund of Ukraine (F80) and Ministry of Education and Science of Ukraine (BG-80F). M.A. thanks the Swiss National Science Foundation for research funding (Grant 31003A_182330/1).
Compliance with Ethical Standards
Conflict of interest
On behalf of all authors, the corresponding author declares that there is no conflict of interest.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Plotkin JB, Kudla G. Synonymous but not the same: the causes and consequences of codon bias. Nat Rev Genet. 2011;12:32–42. doi: 10.1038/nrg2899. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Sauna ZE, Kimchi-Sarfaty C. Understanding the contribution of synonymous mutations to human disease. Nat Rev Genet. 2011;12:683–691. doi: 10.1038/nrg3051. [DOI] [PubMed] [Google Scholar]
- 3.Gingold H, Pilpel Y. Determinants of translation efficiency and accuracy. Mol Syst Biol. 2011;7:481. doi: 10.1038/msb.2011.14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Shabalina SA, Spiridonov NA, Kashina A. Sounds of silence: synonymous nucleotides as a key to biological regulation and complexity. Nucl Acids Res. 2013;41:2073–2094. doi: 10.1093/nar/gks1205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Brule CE, Grayhack EJ. Synonymous codons: choose wisely for expression. Trends Genet. 2017;33:283–297. doi: 10.1016/j.tig.2017.02.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Chater KF. Streptomyces inside-out: a new perspective on the bacteria that provide us with antibiotics. Philos Trans R Soc Lond B Biol Sci. 2006;361:761–768. doi: 10.1098/rstb.2005.1758. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chater KF, Chandra G. The use of the rare UUA codon to define "expression space" for genes involved in secondary metabolism, development and environmental adaptation in streptomyces. J Microbiol. 2008;46:1–11. doi: 10.1007/s12275-007-0233-1. [DOI] [PubMed] [Google Scholar]
- 8.Zaburannyy N, Ostash B, Fedorenko V. TTA Lynx: a web-based service for analysis of actinomycete genes containing rare TTA codon. Bioinformatics. 2009;25:2432–2443. doi: 10.1093/bioinformatics/btp402. [DOI] [PubMed] [Google Scholar]
- 9.Hopwood D. Genetic analysis and genome structure in Streptomyces coelicolor. Bacteriol Rev. 1967;31:373–403. doi: 10.1128/BR.31.4.373-403.1967. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Merrick M. A morphological and genetic mapping study of bald colony mutants of Streptomyces coelicolor. J Gen Microbiol. 1976;96:299–315. doi: 10.1099/00221287-96-2-299. [DOI] [PubMed] [Google Scholar]
- 11.Piret J, Chater K. Phage-mediated cloning of bldA, a region involved in Streptomyces coelicolor morphological development, and its analysis by genetic complementation. J Bacteriol. 1985;163:965–972. doi: 10.1128/JB.163.3.965-972.1985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lawlor E, Baylis H, Chater K. Pleiotropic morphological and antibiotic deficiencies result from mutations in a gene encoding a tRNA-like product in Streptomyces coelicolor A3(2) Genes Dev. 1987;1:1305–1310. doi: 10.1101/gad.1.10.1305. [DOI] [PubMed] [Google Scholar]
- 13.White J, Bibb M. bldA dependence of undecylprodigiosin production in Streptomyces coelicolor A3(2) involves a pathway-specific regulatory cascade. J Bacteriol. 1997;179:627–633. doi: 10.1128/jb.179.3.627-633.1997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Rebets Y, Ostash B, Fukuhara M, Nakamura T, Fedorenko V. Expression of the regulatory protein LndI for landomycin E production in Streptomyces globisporus 1912 is controlled by the availability of tRNA for rare UUA codon. FEMS Microbiol Lett. 2006;256:30–37. doi: 10.1111/j.1574-6968.2005.00087.x. [DOI] [PubMed] [Google Scholar]
- 15.den Hengst CD, Tran NT, Bibb MJ, Chandra G, Leskiw BK, Buttner MJ. Genes essential for morphological development and antibiotic production in Streptomyces coelicolor are targets of BldD during vegetative growth. Mol Microbiol. 2010;78:361–379. doi: 10.1111/j.1365-2958.2010.07338.x. [DOI] [PubMed] [Google Scholar]
- 16.Higo A, Horinouchi S, Ohnishi Y. Strict regulation of morphological differentiation and secondary metabolism by a positive feedback loop between two global regulators AdpA and BldA in Streptomyces griseus. Mol Microbiol. 2011;81:1607–1622. doi: 10.1111/mmi.12160. [DOI] [PubMed] [Google Scholar]
- 17.Trepanier N, Jensen S, Alexander D, Leskiw B. The positive activator of cephamycin C and clavulanic acid production in Streptomyces clavuligerus is mistranslated in a bldA mutant. Microbiology. 2002;148:643–656. doi: 10.1099/00221287-148-3-643. [DOI] [PubMed] [Google Scholar]
- 18.Makitrinskyy R, Ostash B, Tsypik O, Rebets Y, Doud E, Meredith T, Luzhetskyy A, Bechthold A, Walker S, Fedorenko V. Pleiotropic regulatory genes bldA, adpA and absB are implicated in production of phosphoglycolipid antibiotic moenomycin. Open Biol. 2013;3:1–13. doi: 10.1098/rsob.130121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Fedorov A, Saxonov S, Gilbert W. Regularities of context-dependent codon bias in eukaryotic genes. Nucl Acids Res. 2002;30:1192–1197. doi: 10.1093/nar/30.5.1192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hou Y-M, Masuda I, Gamper H. Codon-specific translation by m1G37 methylation of tRNA. Front Genet. 2019;9:713. doi: 10.3389/fgene.2018.00713. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.López-García MT, Santamarta I, Liras P. Morphological differentiation and clavulanic acid formation are affected in a Streptomyces clavuligerus adpA-deleted mutant. Microbiology. 2010;156:2354–2365. doi: 10.1099/mic.0.035956-0. [DOI] [PubMed] [Google Scholar]
- 22.Moura G, Pinheiro M, Silva R, Miranda I, Afreixo V, Dias G, Freitas A, Oliveira JL, Santos MA. Comparative context analysis of codon pairs on an ORFeome scale. Genome Biol. 2005;6:R28. doi: 10.1186/gb-2005-6-3-r28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Pinheiro M, Afreixo V, Moura G, Freitas A, Santos MA, Oliveira JL. Statistical, computational and visualization methodologies to unveil gene primary structure features. Methods Inf Med. 2006;45:163–168. doi: 10.1055/s-0038-1634061. [DOI] [PubMed] [Google Scholar]
- 24.Stajich J, et al. The Bioperl toolkit: Perl modules for the life sciences. Genome Res. 2002;12:1611–1618. doi: 10.1101/gr.361602. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Newcombe R. Interval estimation for the difference between independent proportions: comparison of eleven methods. Stat Med. 1998;17:873–890. doi: 10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I. [DOI] [PubMed] [Google Scholar]
- 26.Fuglsang A. Intragenic position of UUA codons in streptomycetes. Microbiology. 2005;151:3150–3152. doi: 10.1099/mic.0.28352-0. [DOI] [PubMed] [Google Scholar]
- 27.Li W, Wu J, Tao W, Zhao C, Wang Y, He X, Chandra G, Zhou X, Deng Z, Chater KF, Tao M. A genetic and bioinformatic analysis of Streptomyces coelicolor genes containing TTA codons, possible targets for regulation by a developmentally significant tRNA. FEMS Microbiol Lett. 2007;266:20–28. doi: 10.1111/j.1574-6968.2006.00494.x. [DOI] [PubMed] [Google Scholar]
- 28.Jeong Y, Kim JN, Kim MW, et al. The dynamic transcriptional and translational landscape of the model antibiotic producer Streptomyces coelicolor A3(2) Nat Commun. 2016;7:11605. doi: 10.1038/ncomms11605. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Medema MH, Kottmann R, Yilmaz P, Cummings M, Biggins JB, Blin K, et al. Minimum information about a biosynthetic gene cluster. Nat Chem Biol. 2015;11:625–631. doi: 10.1038/nchembio.1890. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhou Z, Dang Y, Zhou M, Li L, Yu CH, Fu J, Chen S, Liu Y. Codon usage is an important determinant of gene expression levels largely through its effects on transcription. Proc Natl Acad Sci USA. 2016;113:E6117–E6125. doi: 10.1073/pnas.1606724113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Moura G, Pinheiro M, Arrais J, Gomes AC, Carreto L, Freitas A, Oliveira JL, Santos MA. Large scale comparative codon-pair context analysis unveils general rules that fine-tune evolution of mRNA primary structure. PLoS ONE. 2007;2:e847. doi: 10.1371/journal.pone.0000847. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Tats A, Tenson T, Remm M. Preferred and avoided codon pairs in three domains of life. BMC Genom. 2008;9:463. doi: 10.1186/1471-2164-9-463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Chevance FFV, Hughes KT. Case for the genetic code as a triplet of triplets. Proc Natl Acad Sci USA. 2017;114:4745–4750. doi: 10.1073/pnas.1614896114. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.

