Skip to main content
mBio logoLink to mBio
. 2024 Mar 21;15(4):e00333-24. doi: 10.1128/mbio.00333-24

Proteins à la carte: riboproteogenomic exploration of bacterial N-terminal proteoform expression

Igor Fijalkowski 1,#, Valdes Snauwaert 1, Petra Van Damme 1,✉,#
Editors: Joerg Vogel2, Joseph Thomas Wade3
PMCID: PMC11005335  PMID: 38511928

ABSTRACT

In recent years, it has become evident that the true complexity of bacterial proteomes remains underestimated. Gene annotation tools are known to propagate biases and overlook certain classes of truly expressed proteins, particularly proteoforms—protein isoforms arising from a single gene. Recent (re-)annotation efforts heavily rely on ribosome profiling by providing a direct readout of translation to fully describe bacterial proteomes. In this study, we employ a robust riboproteogenomic pipeline to conduct a systematic census of expressed N-terminal proteoform pairs, representing two isoforms encoded by a single gene raised by annotated and alternative translation initiation, in Salmonella. Intriguingly, conditional-dependent changes in relative utilization of annotated and alternative translation initiation sites (TIS) were observed in several cases. This suggests that TIS selection is subject to regulatory control, adding yet another layer of complexity to our understanding of bacterial proteomes.

IMPORTANCE

With the emerging theme of genes within genes comprising the existence of alternative open reading frames (ORFs) generated by translation initiation at in-frame start codons, mechanisms that control the relative utilization of annotated and alternative TIS need to be unraveled and our molecular understanding of resulting proteoforms broadened. Utilizing complementary ribosome profiling strategies to map ORF boundaries, we uncovered dual-encoding ORFs generated by in-frame TIS usage in Salmonella. Besides demonstrating that alternative TIS usage may generate proteoforms with different characteristics, such as differential localization and specialized function, quantitative aspects of conditional retapamulin-assisted ribosome profiling (Ribo-RET) translation initiation maps offer unprecedented insights into the relative utilization of annotated and alternative TIS, enabling the exploration of gene regulatory mechanisms that control TIS usage and, consequently, the translation of N-terminal proteoform pairs.

KEYWORDS: alternative translation initiation, N-terminal proteoforms, N-terminal proteomics, riboproteogenomics, (retapamulin-assisted) ribosome profiling (Ribo-RET)

INTRODUCTION

In recent years, the rapid accumulation of novel genome sequencing efforts in bacteria has rendered manual genome annotation unfeasible. Despite continuous improvements, automatic genome annotation pipelines have repeatedly been shown to vastly underestimate the true complexity of bacterial genomes (1, 2). Annotation census efforts have revealed that among over 100 annotated prokaryotic genomes, up to 60% of genes suffered from incorrectly assigned translation initiation sites (TISs) with the longest possible open reading frame (ORF) often being preferentially selected in silico (3). This fact highlights the particular difficulty in accurately and comprehensively assigning TISs. Annotation pipelines have also demonstrated limited capacity in discovering novel protein families, as clearly evident for notoriously under annotated small ORFs (sORFs) (46), posing a risk of propagating errors persisting in current annotations (1, 7). Moreover, the currently employed tools exhibit relatively poor (~70%) agreement between predictions, and not all newly submitted genomes utilize the same state-of-the-art annotation tools (8). To add to this, the accuracy of prediction algorithms can be highly dependent on the properties of individual genomes, such as their guanine-cytosine (GC) content, impacting not only general gene detection but also TIS assignment. These inaccuracies further lead to erroneous predictions that do not accurately represent actual protein-coding genes or regions (9). Importantly, in the case of bacterial genomes, multiple TIS annotations within a single gene or transcript are generally not considered. Altogether, we and others have demonstrated that all these factors contribute to the persistent underestimation of the true complexity of bacterial genomes and their encoding proteomes, even in case of well-studied organisms (10).

The shortcomings of currently applied genome annotation strategies have gradually been addressed by continuous, experimental data-driven reannotation efforts. Owed to the recent developments in genomics, ribosome profiling (Ribo-seq) has provided an unprecedented view on bacterial translational landscapes by sequencing of mRNA fragments encapsulated within actively translating ribosomes (1113). This technology has not only demonstrated pervasive translation outside of annotated genes but has also allowed for the correction of existing annotation errors (7, 1116). The most recent additions to the bacterial genomics toolkit, retapamulin-assisted ribosome profiling (Ribo-RET) and antimicrobial peptide oncocin (Onc112) ribosome profiling, have enabled an in-depth examination of bacterial translation initiation landscapes (17, 18). With the recent advent of translation initiation mapping, high-throughput investigation of alternative TIS selection became feasible in bacterial models, similar to lactimidomycin-based TIS profiling permitting analogous studies in eukaryotic models (11). Such a systematic effort allows us to fully appreciate the repertoire of molecular protein forms expressed, including proteoforms originating from alternative in-frame TIS selection (and not proteolysis)—here referred to as N-terminal (Nt-)proteoforms—which posed a particular challenge to formerly employed detection methods (15, 16). In the absence of reliable genomics data on bacterial translation initiation, the repertoire of technologies suitable for proteoform-targeted studies has been limited. On the protein level, alternative Nt-proteoforms often share the great majority of their identifiable peptides, especially since an alternative TIS is often contained within the 3’ part of the coding sequence (CDS), rendering its unambiguous proteomic identification challenging. Therefore, unique Nt-peptides are often the sole means of confident Nt-proteoform detection. To this end, Nt-proteomics methods, jointly referred to as N-terminomics, have been the most successful in elucidating the landscape of Nt-proteome variation in bacteria (16). Complicating the assignment of N-termini to bacterial TIS further is the typical absence of Nt-protein modifications at mature bacterial protein N-termini, which, in the case of eukaryotic N-termini, serve as proxies for translation initiation (e.g., N-terminal acetylation). Previously, a subset of 11 genes displaying evidence of annotated as well as alternative initiation has been found in Escherichia coli utilizing a potent N-terminal proteomics technique (COFRADIC) in conjunction with actinonin treatment, effectively leaving true N-termini marked with formyl groups (15, 16). Studies that aim at correcting genome and concomitantly TIS annotations are thus undisputedly of central importance for improving our understanding of bacterial systems biology (1921). Utilizing complementary translation (initiation) and proteomics (e.g., N-terminomics) data, an approach referred to as riboproteogenomics, has already shown great promise in this regard (15, 16, 22). Yet, further improvements in proteomics validation of genomic findings still remain an important challenge (10, 23).

Albeit well-established in eukaryotic models (14, 24, 25), only scarce examples of alternative Nt-proteoform production in bacteria have been described to date (26). However, a recent landmark study, reporting on E. coli Ribo-RET data sets, discovered 42 genes displaying evidence of annotated as well as alternative translation initiation, thereby reporting on the translation of gene-specific Nt-proteoform pairs at the genome-wide level for the first time (17). Moreover, limited studied cases reporting on the properties of individual members of Nt-proteoform pairs have uncovered both structural and functional implications for such alternative proteoforms (27, 28). It is conceivable that alternative translation initiation in bacteria can bear similar functions to previously described mammalian and plant alternative initiation events, allowing for a rapid response to environmental cues, regulating subcellular localization and protein stability, and serving distinct functions within multiprotein complexes (10, 25, 29, 30). In Salmonella, it has been shown that the simultaneous expression of two SsaQ proteoforms is required for the effective formation of the type III secretion system (T3SS) utilized for effector delivery into host cells (31). Nt-proteoforms were further shown to engage in unique protein–protein interactions (PPIs) and in the formation of alternative protein complexes, as reviewed by Fijalkowska et al. (10). In cyanobacteria, it has been shown that CcmM Nt-proteoforms serve distinct functions in the maturation of the quaternary Rubisco complex. Specifically, the short proteoform of CcmM (short CcmM or CcmMS) serves as structural component of Rubisco supercomplex, while the annotated, longer CcmM proteoform (CcmML) serves as an anchor of the complex in the carboxysome lumen (32). Additionally, the Nt-truncated E. coli proteoform CheAS engages in both phosphorylating and dephosphorylating complexes of the chemotaxis pathway, while CheAL acts exclusively as a kinase (33). Moreover, two alternative TISs in the E. coli infB gene were shown to give rise to translation initiation factor 2 α and β isoforms, respectively (34). As a final representative example, the Nt-proteoforms GALLSS and GALLSL of Agrobacterium rhizogenes were shown to encode two essential components of the T-DNA complex required for plant transformation (35). Despite these intriguing findings, only a few systematic efforts have been made to comprehensively catalogue the repertoire of expressed bacterial Nt-proteoforms and to study the regulatory mechanisms governing their potential conditional (e.g., growth-phase-specific) expression. Given the functional implications of alternative Nt-proteoform expression, their investigation is crucial to elucidate the intricacies of bacterial systems biology and, by extension, to fully grasp the pathogenicity mechanisms displayed by infectious bacteria.

In this work, we conducted a systemic riboproteogenomic quest to explore the genome of the intensively studied model bacterial pathogen Salmonella enterica serovar Typhimurium (S. Typhimurium) for paired Nt-proteoform expression. Similar to most bacterial species, and despite major recent advances made by us and others, our understanding of the true genome complexity of S. Typhimurium is still incomplete (4, 16, 22, 3638). In S. Typhimurium, Nt-proteoform mapping revealed about 50 genes potentially encoding Nt-proteoform pairs, laying the foundation for their functional characterization and possible functional diversification. Specifically, Ribo-seq and Ribo-RET data generated across a panel of complementary growth conditions were repurposed to more comprehensively capture the range of expressed proteoforms—the full coding potential of the bacterial genome. By studying distinct bacterial growth phases and various environmental stresses, the value of using Ribo-RET data for studying regulation of Nt-proteoform expression was investigated, providing a first hint toward conditional proteoform expression alongside the potential diversifying biological functions of proteoforms. With the functional relevance of these understudied genomic elements, the study of so far uncharted translation (initiation) events opens a novel avenue into improving our understanding of bacterial systems.

RESULTS

Ribo-RET data for the detection of alternative initiation events leading to bacterial proteoform expression

We previously employed complementary Ribo-seq and proteomics data to understand the biases inherent in proteomics for the detection of sORF-encoded polypeptides (SEPs) from their encoding sORFs (39), a category that is often under-annotated in current genome annotations (5). More specifically, besides total shotgun proteomics data, we obtained matching Ribo-seq data by use of the translation inhibitors retapamulin (RET) and chloramphenicol (CAM), allowing the determination of ribosome occupancy during the initiation (Ribo-RET) and elongation phases of translation, respectively. These complementary data sets were generated from diverse growth conditions, reflecting various bacterial growth phases and environmental stresses (see Materials and Methods and Fijalkowski et al. [39]). In this study, we repurposed the Ribo-seq data specifically to discover expressed Nt-proteoform pairs. For such pairs, in addition to the commonly observed translation initiation at the annotated translation initiation site (referred to as database-annotated TIS or dbTIS) and corresponding to the annotated protein-coding region, Ribo-RET evidence additionally indicated the occurrence of alternative translation initiation (aTIS). Collectively, these TIS events give rise to the expression of Nt-proteoform pairs. As tools utilizing signal distribution have proven effective in delineating novel, true ORFs (15, 16), and our recently published gene detection pipeline selected the longest ORF in the case of in-frame reading frames, implying a bias against calling of in-frame ORFs (39), we conducted ORF calling on these data sets without imposing any length-based filters, yet adhering to stringent thresholds for ORF identification. This revised ORF calling strategy enabled us to identify a set of 26 curated genes, each encoding a pair of putative N-terminal proteoforms. Particularly for extended proteoforms like SseLL (Fig. 1), continuous elongating translation (Ribo-seq) signal upstream of the annotated TIS can be observed. This makes it challenging to detect independent translation initiation at the annotated TIS when expression levels of the long and annotated proteoforms do not significantly differ. Likewise, an increased intensity of the Ribo-seq signal around initiation (39) does not always serve as a definitive metric for discovering truncated proteoforms, as variations in translation speed (e.g., translation stalling) can also lead to local signal accumulation. These shortcomings can, however, be mitigated using the relatively high specificity of the translation initiation signal (Ribo-RET [39]), previously shown to facilitate the Ribo-seq-assisted discovery of alternative proteoforms, and thus concomitantly the concurrent identification of ORFs encoding multiple Nt-proteoforms.

Fig 1.

Fig 1

Ribo-seq unveils translation of Nt-proteoform pairs in S. Typhimurium. Ribo-RET (RET) and Ribo-seq signals from representative growth conditions are presented. (A) ssaQ under anaerobic shock, (B) sseL in MEP, and (C) mrcB in MEP. Called TIS are flagged, and the corresponding translation initiation region of alternative TIS (in orange) is indicated.

After manual curation, involving a meticulous inspection of raw and positional reads around the identified TIS using a genome browser (Integrative Genomics Viewer [IGV]) (40) (see Materials and Methods), more than 50% of the called proteoform pairs (50 in total) or 26 proteoform pairs exhibited robust and compelling Ribo-RET signal for both the database-annotated and the alternative translation initiation sites called. The highly confident Nt-proteoform pairs central to this study are depicted in Fig 2A ; Fig S1. This set includes proteoform pairs where the non-annotated member represents either an N-terminal extension (#6) or an N-terminal truncation (#20). Among these novel N-terminal proteoforms, 20 initiated at an alternative AUG start codon, while six initiated at a near-cognate start codon, such as GUG (#1), CUG (#4), or AUU (#1) (Fig. 2B). Thus, approximately one quarter of the newly discovered proteoforms initiates at near-cognate start codons, a higher frequency than observed for annotated TIS, where 88%, 9.2%, and 2.7% initiate at AUG, GUG, and UUG, respectively.

Fig 2.

Fig 2

High-confidence N-terminal S. Typhimurium proteoform pairs. (A) Relationship between the identity and size (expressed as a percentage of the annotated proteoform length) of the newly identified member and annotated proteoform of an N-terminal proteoform pair. N-terminal extensions and N-terminal truncations are represented in dark and light orange, respectively. (B) Usage of initiation codons in the newly discovered N-terminal proteoforms. Note that translation initiation at non-near-cognate in-frame, downstream TIS was not considered. (C) Shine-Dalgarno (SD)-like context strength for alternative (aTIS) versus database-annotated translation initiation sites (dbTIS) with an asterisk indicating a P value < 0.01 (t-test). Corresponding free energy predicted values are listed in Table S1.

The study by Meydan et al., which explored the alternative translatome of E. coli, encompassed 42 genes exhibiting conserved evidence of translation initiation at an in-frame downstream aTIS when comparing two strains (17). In the context of our confident Nt-proteoform pairs, we found evidence for the homologous expression of six proteoforms in our S. Typhimurium data sets previously reported in E. coli (mrcB, pmbA, speA, yebG, clpB, and infB) (17). Although potential indications for homologous members of additional proteoform pairs reported by Meydan et al. (17) were found in our data set (e.g., arcB and slyB), the calling of matching aTIS may have been missed by our gene detection tools for various reasons, including low expression of the alternative proteoform (e.g., arcB) or the omission of translation initiation at in-frame, downstream non-near-cognate start codons.

Conservation and experimental proteomic data in support of newly discovered proteoform pairs

When comparing the Shine-Dalgarno (SD)-like sequence initiation context, alternative initiation events detected in the context of Nt-proteoform pairs exhibit a significantly stronger initiation context compared to their annotated counterparts (P value < 0.01, as detailed in Table S1 and illustrated in Fig. 2C). Additionally, for all six non-annotated extended Nt-proteoform members of identified pairs, the ConSurf calculated conservation scores (41) of the coding sequence remain high throughout the extended sequence of the proteoform beyond the annotated ORF (Fig. 3). The ConSurf analysis involved a standard approach of homologous gene search for each candidate proteoform, followed by conservation scoring using the Rate4Site algorithm and subsequent phylogenetic tree reconstruction, as previously described (41). Moreover, (homology-inferred) UniProt database (unreviewed) entries for all six non-annotated extensions were identified in SL1344 and/or related LT2 and 14028s strains (e.g., SseLL [e.g., A0A719CWE6], SpeAL [A0A719A111], FruKL [Q7CQ78], PagCL [P23988], MotAL [A0A718V5M5], and RpoDL [A0A0F6B6Z4]). These findings support the true translation initiation and coding potential of the Nt-extended proteoforms identified by Ribo-RET.

Fig 3.

Fig 3

Nucleotide sequence conservation of newly discovered N-terminal extended proteoforms constituting proteoform pairs. Nucleotide sequence conservation remains high throughout the coding sequence encoding the N-terminal extensions (portion between the red and blue lines, with the red and blue lines indicating the first nucleotide [nt.] of the alternative [aTIS] and database-annotated TIS [dbTIS], respectively) of the six Nt-extended proteoforms of proteoform-pair-encoding genes identified, namely speA, sseL, fruK, pagC, motA, and rpoD, as visualized using ConSurf by plotting calculated nucleotide conservation scores (41). Only the first 120 nt. of the annotated coding sequence are shown for simplicity; hence, the missing blue line corresponding to the first nt. matching the dbTIS of rpoD.

As ribosomal occupation does not always relate to genuine protein production (42), experimental mass-spectrometry (MS)-based detection is a common go-to method to validate newly discovered translation events inferred from Ribo-seq. Initially, the theoretical MS detectability of Nt-proteoform indicative tryptic peptides was assessed using the advanced proteolytic peptide predictor algorithm AP3 (43) (Fig. S2). For univocally identifying proteoform members constituting a proteoform pair, and because the vast majority of their identifiable peptides are shared (i.e., the peptide coverage is on average over 90% identical when considering the two members of the Nt-proteoform pairs identified here), Nt-peptides may often serve as the sole means of confident Nt-proteoform detection by acting as proxies of translation initiation. Our analyses highlight that finding unique proteomic support is challenged by the scarcity of MS-detectible proteoform-specific tryptic peptides. Specifically, in the specific case of the YebGL, the theoretical N-terminal peptide detectability is very low, making direct and univocal MS detection of YebGL unlikely (Fig. 4A).

Fig 4.

Fig 4

Riboproteogenomic identification and blotting-based validation of the expression of YebG and FruK S. Typhimurium Nt-proteoform pairs. (A) AP3-derived detectability scores (43) of the YebGL and FruKL proteoforms (with the small proteoform region past the orange line). Peptide detectability is represented in a color code according to the AP3 scale: high detectability (scores greater than 0.9) is indicated in green, medium detectability (scores between 0.4 and 0.9) in light green, and low detectability (scores less than 0.4) in pale red. (B) Ribo-RET (RET) and Ribo-seq signal for yebG and fruK in salt shock (NaCl) is shown. Called TIS are flagged, and the corresponding translation initiation region of the alternative TIS (orange) is indicated. (C) HiBiT blotting of corresponding S. Typhimurium SL1344 wild-type, yebG::HiBiT KmR and yebGG-51A::HiBiT KmR (Chr: 1,933,051 G > A), fruKATG-1CTT::HiBiT KmR (Chr: 2,303,857 – Chr: 2,303,859 ATG > CTT), fruKGTG-82CTT::HiBiT KmR (Chr: 2,303,857 – Chr: 2,303,859 GTG > CTT), and fruKATG-1CTT,GTG-82CTT::HiBiT CmR and fruK::HiBiT KmR protein extracts (ESP, OD600 2.0 cultures) using chemiluminescence detection validates the (selective) expression of the YebG and FruK proteoforms. Arrows indicate sizes of corresponding large (black arrow) and small (white arrow) Nt-proteoform members (YebGL [12.0 kDa], YebGS [10.2 kDa], FruKL [35.0 kDa], and FruKS [32.0 kDa]). (*) A band detected by streptavidin blotting (S680) and corresponding to an endogenously biotinylated S. Typhimurium protein served as loading control.

Nonetheless, to explore riboproteogenomic matching support for our confident list of Nt-proteoform pairs, complementary proteomics data were examined (15, 16, 22). When interrogating Ribo-seq matching shotgun and N-terminal proteomics data, four translation initiation-indicative Nt-peptides (Nt-peptides of RpoDS, CspEL, and YebGL) were found. Besides, proteogenomic peptides exclusively mapped to the unique extensions of PagCL, FruKL, SpeAL, PmbAL, MrcBL, OmpXL, HemCL, ClpBL, InfBL, YdgAL, and SL1344_1071L (Table S1) were identified. Collectively, proteomic support for 14 proteoform members of the 26 confident Nt-proteoform pairs has been discovered. Of note, complementary data for the expression of the homologous PagC and OmpX proteoform pairs come from the observation that annotated OmpXL and newly discovered PagCL proteoform sequences display 37% sequence identity (Clustal O alignment, v1.2.4), with conservation of the TIS matching the extended proteoforms.

To obtain complementary evidence of Nt-proteoform pair expression, especially in cases where finding MS-supportive evidence is unlikely (e.g., YebG, Fig. 4A), we pursued the validation of endogenous expression for the YebG, ClpB, and FruK proteoform pairs using HiBiT blotting (44). Ribo-RET evidence suggested translation initiation from the annotated dbTIS (AUG) and initiation at a downstream in-frame AUG for all three cases (Fig. 4B and 5A). The small luminescent peptide tag HiBiT was chromosomally introduced as a C-terminal tag through recombineering, allowing for sensitive detection of endogenous protein expression. In the case of HiBiT-tagged YebG and ClpB in control cells, two distinct molecular weight (MW) bands were observed—with the lower band being fourfold less intense in the case of yebG and ninefold less in the case of clpB—at the corresponding molecular weights of the YebG and ClpB proteoforms (Fig. 4C and 5B). For FruK—likely due (in part) to the relatively small molecular weight difference (3 kDa) and/or the difference in expression of the two FruK proteoforms—no two discrete bands were apparent (Fig. 4C). Endogenous site-specific mutagenesis by means of oligo-mediated allelic replacement (OMAR) confirmed that the lower MW band of HiBiT-tagged YebG corresponds to YebGS (45). Specifically, the Ribo-RET-called ATG aTIS (encoding M17) was mutated to ATA (encoding I17) resulting in the exclusive expression of YebGL (Fig. 4C). Interestingly, mutating the Ribo-RET-called GTG dbTIS (encoding M27) matching FruKS to the non-near-cognate start codon CTT did not abolish FruK expression overall but resulted in a fivefold expression reduction, indicative of the exclusive expression of the newly discovered FruKL and the expression of FruKL at lower levels—especially in control cells—as compared to FruKS, in line with the corresponding Ribo-RET data (Fig. 4B). This was also apparent from the slightly higher MW band over the (composite) band detected in the fruk::HiBiT control strain. Mutation of both TIS completely abolished FruK expression. Although we have obtained conclusive support for the expression of N-terminal (Nt)-proteoform pairs, some validation efforts highlight the need for further refinement in the resolution of TIS calling. Specifically, in the case of ClpB (ClpBS), we identified an aTIS corresponding to Met143. However, mutating this clpB aTIS still resulted in the expression of a ClpB proteoform pair. In line with the previously experimentally confirmed aTIS of ClpB corresponding to Val149 (46), it was only when mutating this corresponding GTG TIS to CTT that exclusive expression of ClpBL was observed, confirming the true nature of this TIS (Fig. 5B).

Fig 5.

Fig 5

Riboproteogenomic identification and blotting-based validation of the expression ClpB Nt-proteoform pairs in S. Typhimurium under regulatory control. (A) Ribo-RET (RET) and Ribo-seq signal for clpB in salt shock (NaCl) is shown. Called TIS are flagged, and the corresponding translation initiation region of the alternative TIS (corresponding to Met143) is indicated (orange). The previously confirmed aTIS of ClpBS, corresponding to Val149, is flagged in blue. (B) HiBiT blotting of corresponding S. Typhimurium SL1344 wild-type, clpB::HiBiT KmR, clpBATG-427CTT::HiBiT KmR (Chr: 2,804,674 – Chr: 2,804,676 GTG > CTT), and clpBGTG-445CTT::HiBiT KmR (Chr: 2,804,656 – Chr: 2,804,658 GTG > CTT) protein extracts (ESP, OD600 2.0 cultures) using chemiluminescence detection validates the (selective) expression of the ClpB proteoform pair. Arrows indicate sizes of corresponding large (black arrow) and small (white arrow) Nt-proteoform members (ClpBL [96.7 kDa], ClpBS [81.0 kDa]). (*) A band detected by streptavidin blotting (S680) and corresponding to an endogenously biotinylated S. Typhimurium protein served as loading control. (C) Ribo-RET tracks of clpB shown in the same scales reveal differential expression of the two proteoforms when comparing MEP versus Salmonella pathogenicity island 2-inducing (InSPI2) growth conditions.

Ribo-RET peak intensity can serve as a proxy for Nt-proteoform expression

Although the general read distribution in ribosome profiling data sets is non-uniform and exhibits 5’ and 3’ polarization, especially in case of bacterial Ribo-seq data sets (11, 13, 47), these data sets have been successfully used for accurate protein synthesis quantification (13). However, to ensure unbiased analysis, the outer 5’ and 3’ codon occupancies of CDSs must be trimmed, preventing the inherent increased read density at these positions from influencing the analysis (14). Accumulation at the 5’ of the coding sequence originates from the fact that while elongation is effectively blocked, initiation ensues, causing ribosome accumulation at the start. Due to the substantial overlap in their sequences, quantifying Nt-proteoform expression of expressed Nt-proteoform pairs is complicated and has been unexplored thus far. However, as evident from the metagene plots displaying typical read distribution in Ribo-seq and Ribo-RET samples, and with the narrow initiation peaks of Ribo-RET data denoting effective translation initiation (39), Ribo-RET data do not suffer from the constraints observed in the case of elongating Ribo-seq data when distinguishing overlapping ORFs. This property makes Ribo-RET data potentially suitable for estimating translation in the case of overlapping or dual-encoding ORFs, as previously demonstrated for analogous lactimidomycin data in eukaryotes (47).

Interestingly, by correlating Ribo-seq and Ribo-RET data sets with matching protein abundance estimates obtained through proteomics, we demonstrate that Ribo-seq and Ribo-RET data correlate with proteomics-determined protein abundances to a similar extent. For instance, as illustrated for the MEP condition, the Pearson coefficient is 0.58 for Ribo-seq and 0.617 for Ribo-RET, respectively (Fig. 6A and B). This indicates that the Ribo-RET signal serves as a quantitative measure of translation and can therefore be exploited for the quantification of protein expression. This property is of particular interest in the context of quantifying N-terminal proteoform expression, as well as the quantification of translation products originating from overlapping CDSs and sORFs.

Fig 6.

Fig 6

Quantifying protein expression using Ribo-seq and Ribo-RET data. (A) Correlation between Ribo-seq determined translation measure and protein steady-state abundance (iBAQ) measured by proteomics in the matching MEP growth condition. (B) Correlation between Ribo-RET determined translation initiation measures and protein steady-state abundance (iBAQ) measured by proteomics in the matching growth conditions (MEP). Analysis has been performed for 3,027 annotated S. Typhimurium protein identifications (39).

Differential expression of proteoforms

The ability to measure proteoform expression effectively within and between biological conditions lays at the foundation for building a functional understanding of their biological relevance. Leveraging the quantitative aspect of Ribo-RET signal intensity, we quantified the expression of Nt-proteoforms across the investigated growth conditions. This approach allowed us to explore differential Nt-proteoform expression, providing an initial insight into their (conditional) dependency of expression (Fig. 7).

Fig 7.

Fig 7

Expression of the members of N-terminal proteoform pairs based on Ribo-RET across a series of growth conditions. (A) Expression of annotated proteoforms across investigated growth conditions as compared to expression in MEP. (B) Expression of alternative proteoforms across investigated growth conditions as compared to MEP. (C) Expression ratio between the alternative and the annotated proteoform members of proteoform pairs across investigated growth conditions. ANA, anaerobic growth; LEP, late exponential phase; InSPI2, SPI2-inducing conditions (growth in low-pH PCN medium); low Mg2+, InSPI2 growth in low magnesium containing PCN.

Interestingly, besides the condition-specific expression patterns observed for some genes encoding multiple proteoforms (Table S1), we additionally noted that individual members of certain proteoform pairs exhibit varying expression ratios across different growth conditions. While the origin—and thus expression—of alternative N-terminal proteoforms can potentially be explained by the expression of alternative transcripts, the analysis of complementary RNA sequencing data (16) combined with the review of previously published efforts to map transcriptional start sites (TSSs) in the S. Typhimurium SL1344 strain (48) indicate that for 25 out of the 26 N-terminal proteoforms under study, both the annotated as well as the alternative translation initiation sites are situated within annotated or experimentally detected transcripts (Table S1). Therefore, differential expression of the Nt-proteoform pairs under study is likely primarily explained at the level of translation and suggests that the selection of dbTIS versus aTIS, along with the corresponding expressed proteoforms, is subject to differential translational control. Intriguingly, in the case of the ClpB, YrbG, and, YdgA proteoform pairs, the expression of a specific proteoform member was notably strongly favored under infection-relevant (InSPI2) conditions (Table S1 ; Fig. S5C). This observation aligns with previous experimental findings during heat shock in E. coli, where ClpB proteoforms similarly showed an increased ClpBS:ClpBL expression ratio under thermal stress conditions (49).

Potential influence on protein localization and physiochemical properties of alternative N-terminal sequences

Using available metadata of annotated proteoforms, including EffectiveDB (50) and PSORTb (51) predicted protein localizations, we identified seven genes (sseL, ssaQ, ydgA, yfhG, mgtC, ompX, and pagC) that could potentially exhibit distinct subcellular localizations among their constituting Nt-proteoform members. In the case of the type III effector (T3E) deubiquitinase SseL, the newly identified 23-amino acid extended proteoform SseLL is predicted to contain a T3E secretion signal in its N-terminal extension (EffectiveT3 model 2.0.1, score 1.0). This signal promotes secretion/translocation of theT3E SseLL proteoform via the T3SS. Consequently, annotated SseLS is likely to remain localized in the bacterial cytosol due to the absence of this signal peptide. Supporting this observation, Niemann et al. reported the translocation of SseL fused to a Cya reporter only when including a part of its so-called 5′ UTR (52). Similarly, the Salmonella pathogenicity island 2 (SPI2)-encoded T3SS protein SsaQL localizes to the basal body of the flagellum, while SsaQS is predicted to reside in the cytoplasm, potentially playing a role in T3SS assembly similar to its homolog SpaO (53). Furthermore, the truncated version of the lipoprotein encoded by yfhG may lose its outer membrane localization, and MgtCS might not be (efficiently) integrated within the membrane compared to their longer counterparts. Lastly, the truncated DNA topoisomerase YdgAS is predicted to lose its membrane attachment.

Furthermore, we computed a series of physiochemical properties for both the alternative and the annotated forms of the 26 proteoform pairs identified in this study (Table S2; Fig. 8). Despite the low average impact on protein length (Fig. 1A), with only five proteoforms (TraS, SsaQ, YdgA, YrbG, and CBW17167) showing a reduction of over 50% in length, differences in N-terminal sequences can lead to altered proteoform properties, as (localized) sequence characteristics contained within these N-terminal sequences can vary considerably (Table S2; Fig. 8).

Fig 8.

Fig 8

Physiochemical properties for alternative and annotated forms of Nt-proteoform pairs may differ. Molecular weight (MW, Da), aliphatic index, instability index, GRAVY hydrophobicity score, and isoelectric point of alternative (matching aTIS, dark yellow) and annotated (matching dbTIS, red) proteoforms are presented. Lines in the line plots (upper panels) are in black only when differences between the members of a pair are significant (t-test, P value < 0.01) (Table S2).

As an illustrative example, the 29-amino acid extension of PagCL (Fig. 9A) exhibits a higher scaled hydropathy when compared to its annotated counterpart (Fig. 9B). SignalP 6.0 (54) predicts a Sec signal peptide with cleavage predicted between Ala23 and Asp24 in PagCL, resulting in the generation of a proteolytic PagC proteoform with N-terminal proteomics support (i.e., the neo-N-terminus D24TNAFSVGYAQSKVQDFKNIR was identified by means of N-terminal proteomics) (15, 39). Similarly, for homologous OmpX, a Sec signal peptide was predicted spanning amino acid sequences 1–23. Consequently, the annotated PagCS and the newly identified OmpXS proteoforms both lack this signal peptide and cleavage site, potentially leading to differential subcellular localization of these truncated proteoforms.

Fig 9.

Fig 9

Identification and characterization of PagC proteoforms. (A) Matching Ribo-RET and Ribo-seq evidence indicative of the translation of PagC Nt-proteoform pairs in S. Typhimurium. Ribo-seq signal of late exponential growth phase (LEP) is shown for pagC. (B) Differing amino acid sequence properties between extended (PagCL) and annotated (PagCS, sequence beyond the vertical red line) PagC proteoforms are illustrated here in the hydropathy plot.

DISCUSSION

N-terminal proteoforms have been identified across a diverse range of species, spanning humans, rodents, plants, viruses, and bacteria (10, 14, 25, 55, 56). Albeit relatively uncharted in case of bacteria, the exploration of bacterial proteoforms that underwent detailed functional characterization has proven to play crucial roles in essential aspects of bacterial physiology, as recently reviewed (10, 26). From the essential coexpression of two SsaQ proteoforms in Salmonella, also identified in our study, facilitating effective infection (31), to the differential regulation of phosphorylation by CheA proteoforms in E. coli (33), N-terminal proteoforms arising from (alternative) translation initiation events have been shown to fulfill distinct functional roles.

Bacterial genome-wide translatome studies and, particularly, translation initiation profiling (Ribo-RET) offer a means to comprehensively detect bacterial proteoforms of translational origin (17). However, existing tools for genome (re-)annotation using translatomics data face challenges in Nt-proteoform discovery, often favoring longer predictions (15, 16). In light of recent advances in genome (re-)annotation efforts exposing the underappreciated complexity of bacterial genomes (4, 6, 21, 55, 5760), studying the selective expression of N-terminal proteoforms represents a crucial step toward a better understanding of bacterial proteoform biology. In this study, we confidently identified 26 N-terminal proteoform pairs in the model bacterial pathogen S. Typhimurium, conducted in silico characterization, and validated proteoform expression using proteomics and HiBiT blotting-based detection. Furthermore, we demonstrate for the first time that translation initiation metrics can be used to effectively quantify proteoform expression across biological conditions, allowing for their functional characterization while providing the first (differential) expression atlas of N-terminal proteoforms in bacteria to date. Despite the substantial overlap between homologous proteoforms detected in the current and previous landmark E. coli study (17), observed discrepancies, especially in TIS calling resolution (e.g., the previously reported and experimentally confirmed aTIS of clpB corresponding to Val149 [46] while we called an aTIS corresponding to Met143 [Table S1 ; Fig. 5]), suggest the need for further improvements in gene detection algorithms to fully utilize the potential of Ribo-RET data sets especially since rigorous determination of the initiation triplet is critical for future TIS mutagenesis studies. Although increased resolution in TIS calling may be anticipated with recent advances in computational technology and deep learning (61), conservation analysis or N-terminomics efforts could also contribute to TIS refinements.

On the other hand, technical constrains persist in bacterial ribosome profiling techniques due to the use of more sequence-biased nucleases (MNase), resulting in inconsistent footprint lengths and more error-prone P-site assignment, lowering data resolution compared to its eukaryotic counterpart RNase I. In practical terms, ribosomal footprint coverage of multiple, neighboring (near-cognate) initiation codons can complicate confident delineation of the selected TIS and, consequently, correct TIS assignment. Moreover, the effective proteomics validation of the discriminative N-terminal part of proteoforms is challenging due to the scarcity of MS-detectable peptides (Fig. 4A; Fig. S2) (43), presenting a major hurdle that should be addressed by future technological advances. Currently, targeted proteomics strategies employing actinonin treatment combined with N-terminomics hold the highest potential for effective proteome-wide scale validation efforts of bacterial Nt-proteoform expression at single amino acid resolution (30, 62).

At an individual level, TIS mutagenesis proves effective for steering selective proteoform expression, opening up possibilities for differential functional proteoform studies, including interactomics (30) and phenotypic analysis (e.g., impact on virulence [31]). We demonstrate, for the first time, that steering and monitoring the endogenous expression of individual proteoforms can be accomplished effectively using multiplexed recombineering (i.e., ssDNA oligo and dsDNA-based recombineering). Scar-free TIS mutagenesis through OMAR (45) and concomitant introduction of a protein tag (e.g., HiBiT [44]) integrated in a selectable dsDNA cassette was successfully performed using multiplexed recombineering approaches (44, 63) (Fig. 4 and 5).

Nonetheless, validating the expression of Nt-proteoform pairs poses considerable challenges. These challenges arise due to frequently subtle differences in molecular weights, which can be attributed to closely spaced TIS (e.g., YebG) or post-translational proteolytic processing of specific proteoform members (e.g., PagC), or simply by significant disparities in proteoform abundance (e.g., FruK). Moreover, the impact of mutation introduced at the TIS on potential (efficiency of) alternative translation initiation cannot be overlooked, as the usage of alternative start sites has been reported in such scenarios (64) and also shown here in the case of FruK (Fig. 4C).

Despite the challenging nature of distinguishing Nt-proteoform identities using proteomic means, recent developments in proteome-wide localization studies allow, at least for the subset of detectable variants, the study of their differential localization (65). A simplified differential ultracentrifugation protocol could possibly shed more light on the (differential) localization of proteoforms at a proteome-wide scale. Moreover, similar to human proteoform efforts, investigating the potential impact on proteoform stability could provide valuable insights (25). Given the overall robust expression levels of alternative proteoforms detected in this study, their regulation, including unraveling the mechanisms driving alternative proteoform selection, remains an interesting avenue for future research, especially considering the absence of leaky scanning in prokaryotes. This exploration will aid to elucidate the functional implications of alternative transcription, alternative translation initiation, and potentially the action of alternative ribosomes (66) on the (selective) expression of alternative proteoforms, as well as the coordination of their functioning.

In addition to the themes outlined above, our findings unveil a notable enrichment for genes encoding proteoform pairs among (inner membrane) proteins that are predisposed to form higher-order homo- or hetero-oligomeric structures (e.g., FruK, ClpB, MotA, MrcB). Particularly, the MrcB proteoform pair is marked by a 42-amino acid extension in the longer variant preceding the transmembrane anchor and comprising a disordered N-terminal region enriched with a basic stretch of lysine and arginine, followed by an acidic sequence dominated by glutamate and aspartate residues. Notwithstanding the proven identical enzymatic functions of MrcB proteoforms, this unique polar region points to alternative modes of membrane interaction as shown for its homologous E. coli proteoform (26, 67, 68). Similarly, outer membrane β-barrel proteins like OmpX and PagC exhibit an interesting proteoform dichotomy; their shorter proteoform counterparts lack the signal peptide present in the longer forms, leading to a mature proteoform length that may be indistinguishable from the short proteoform due to post-translational signal processing. Furthermore, a subset of proteoform-encoding genes, such as SsaQ (and related SpaO), OrgA, and SseL, is implicated in T3E secretion and tends to exhibit conservation that is largely restricted to the Salmonella genus, highlighting their specialized roles in pathogenicity. Intriguingly, our proteoform-mapping effort also spotlights the kil/phd toxin/antitoxin system, wherein the two members encoded within the same cistronic unit showcase expression of proteoform pairs, pointing to a nuanced regulatory mechanism at play within this toxin–antitoxin interaction.

The expanding catalog of bacterial proteoforms of translational origin necessitates future validation and functional characterization efforts. Much like genome (re-)annotation efforts in general, effective detection of bacterial N-terminal proteoforms lays the foundation for such future studies. Using the workflow presented in this study, N-terminal proteoforms can be detected in other bacterial species, as previously shown in E. coli (17), thereby providing more comprehensive pictures of (conditional) translation initiation landscapes while aiding our understanding of bacterial proteoform biology. Given that various proteins involved in bacterial pathogenicity display alternative Nt-proteoform expression, gaining insights on their biological function can moreover enhance our understanding of bacterial infections.

MATERIALS AND METHODS

Bacterial culture conditions used for riboproteogenomics data analysis

The S. enterica serovar Typhimurium wild-type (WT) strain SL1344 (69) (genotype: hisG46, phenotype: His(-); biotype 26i) was acquired from the Salmonella Genetic Stock Center (SGSC, Calgary, Canada; cat# 438 [69]). Bacterial growth conditions, as previously detailed by Fijalkowski et al. (39), were performed in liquid Lennox broth (LB) growth medium (10 g/L Bacto tryptone, 5 g/L Bacto yeast extract, 5 g/L NaCl) or various formulations of phosphate carbon nitrogen (PCN) medium (70) (InSPI2; Salmonella pathogenicity island 2-inducing condition; pH 5.8, 0.4 mM Pi), as described previously (39). These conditions encompassed early exponential growth phase (EEP; OD600 0.1), mid-exponential growth phase (MEP; OD600 0.3), late exponential growth phase (LEP; OD600 1.0), early stationary phase (ESP, OD600 2.0), and late stationary phase (LSP, OD600 2.0 + 6 h of extra growth). Additionally, bacterial cultures at MEP were subjected to osmotic shock (0.3 M NaCl shock for 10 min), anaerobic shock, SPI2-inducing PCN (InSPI2; pH5.8, 0.4 mM Pi), low-magnesium SPI2-inducing PCN (low Mg2+; pH5.8, 0.4 mM Pi) containing low levels (10 µM) of magnesium sulfate or MEP-grown InSPI2 cultures subjected to nitric oxide shock (addition of Spermine NONOate to a final concentration of 250 µM for 20 min). All PCN media were supplemented with 5 mM final concentration (f.c.) histidine. These diverse growth conditions were chosen to capture a comprehensive snapshot of the riboproteogenomic landscape of S. Typhimurium. Ribosome profiling and Ribo-RET analyses were performed in biological duplicates as described by Fijalkowski et al. (39).

Plasmids and oligos

The pORTMAGE-2 helper plasmid (71) obtained from Addgene (plasmid #72677) encodes Gam, Beta, and Exo, in addition to a dominant negative mutL (E32K) allele. The expression of these elements is collectively controlled by the temperature-sensitive cl857 repressor. All oligonucleotides utilized in this study are detailed in Table S3 and were procured from Integrated DNA Technologies (IDT, Leuven, Belgium). Standard desalted purification was applied to 25 nmol oligos, while high-performance liquid chromatography (HPLC) purification was implemented for 100 nmol modified oligos (IDRT, Coralville, IA, USA). The oligos were resuspended in water, achieving a final concentration of 100 µM. Single-stranded DNA (ssDNA) oligos used for OMAR were designed using the Mage Oligo Design Tool (MODEST) (72). Modified ssDNA oligos, featuring two phosphorothioate linkages between the ultimate 5’ and 3’ nucleotides, and occasionally incorporating a 2-fluoro-uridine modified base within the range of five bases around the mismatched nucleotide responsible for the introduction of the point mutation, were synthetized for TIS mutagenesis.

Ribo-RET data analysis

Sequencing files underwent demultiplexing through the bcl2fastq software (Illumina). Sample-specific fastq.gz files, originating from individual lanes, were concatenated using Unix’s “cat” command. NextFlex library-introduced Unique molecular identifiers (UMI) were extracted from individual reads using a custom Python script. PCR bias normalization was performed using a previously established normalization procedure (47). Normalized data underwent a two-step trimming process using cutadapt. The first step involved the removal of standard Illumina adapters (cutadapt -q 20 -m 25 -e 0.2), followed by the secondary removal of UMIs (cutadapt -u 4 -u -4). Data were mapped to indexed ribosomal RNA (rRNA) sequences using Bowtie (bowtie -t -n 2 -p 6 –best). rRNA sequences retrieved from Ensembl and Genebank were supplemented with additional sequences of tRNAs, RNA subunits of nucleoproteins, and non-coding RNAs (ncRNAs) to the rRNA index. Reads aligning to these sequences were excluded from further analysis. The remaining reads were mapped to the S. Typhimurium SL1344 genome (Ensembl, GCA_000210855.2) using Bowtie (-t -n 2 -p 6 -m 1 --best --strata –sam). The resulting SAM files were converted to BAM format and were sorted using Samtools (73). RiboWalz (74), recognized for superior performance compared to plastid (75) used in previous studies (16), was employed for ribosomal P-site assignment for all reads. Positional data were counted using a custom Python script, normalized to reads per million (RPM) values, and were further processed for the creation of positional read occupancy tables in Python and R tidyverse environment. The Reads Per Kilobase of transcript per Million reads mapped (RPKM) values were then calculated for all Ensembl annotated genes. Separate BedGRaph files were generated using a custom Python script for effective data visualization, as detailed by Fijalkowski et al. (39). Sequencing statistics and metagene analysis has been reported by Fijalkowski et al. (39).

Ribo-seq inferred ORF detection and quantification of proteoform translation

Detailed procedures for the generation and sequencing (statistics) of ribosome profiling libraries of S. Typhimurium SL1344 libraries have been described previously (39). ORF delineation from Ribo-RET was essentially performed as described by Fijalkowski et al. (39). Notably, data filtration was omitted, allowing the unbiased selection of multiple in-frame ORFs and thus multiple alternative translation initiation events per ORF. Sliding window signal detection was executed using a custom Python script. A global average coverage per codon served as the background cutoff, subtracted from each genomic position before retapamulin signal detection. This ORF calling strategy yielded a set of 50 putative N-terminal proteoform pairs. The candidates were manually curated, involving detailed inspection of raw and positional reads surrounding the identified TIS for evidence of robust (>5 RPM) initiation signal clearly distinguished from surrounding background signal, precise alignment of raw read center with the putative TIS (within 9nt window from Ribo-Walz [74] inferred ribosomal P-site) and overall sequencing quality (sequencing quality score >30 and mapping quality score MAPQ > 30, including only uniquely mapped reads). Following this curation, over 50% (26) high-confidence proteoform pairs were retained. The curated high-confidence ORF list underwent further processing using custom R and Python scripts, as previously detailed (39). A differential expression analysis for ORFs detected across all investigated Ribo-seq conditions was conducted using the DESeq R package (76). Corrected retapamulin intensity values used in the expression analysis were derived as the integral of the intensity of the retapamulin peak detected normalized against the detected average expression intensity within the body of the gene (background).

Conservation analysis

We conducted a comprehensive evolutionary conservation analysis of the alternative proteoforms using the ConSurf pipeline (41). Initially, the nucleotide sequences of the putative proteoforms were BLAST-searched against the UNIREF-90 database to identify homologous sequences. Redundant homologous sequences were excluded, and the remaining sequences were aligned using MAFFT, facilitating the reconstruction of a phylogenetic tree. The phylogenetic tree, alongside calculated evolutionary distances and the multiple sequence alignment, was then utilized in the Rate4Site algorithm allowing for Bayesian calculation of evolutionary rates. After normalization and binning, the conservation scores were determined on a scale from 1 (indicating high variability) to 9 (denoting strong evolutionary conservation) (Fig. 3). For truncated proteoforms, where the alternative start codon is embedded within a generally highly conserved coding sequence, we assessed the conservation of the start codon and its associated Shine-Dalgarno context across 100 genomes. These genomes are representative of all major branches in the Enterobacteriaceae phylogeny, including the 10 most complete genomes from Escherichia, Shigella, Salmonella, Citrobacter, Enterobacter, Klebsiella, Erwinia, Serratia, Yersinia, Morganella, and Proteus, as available in the NCBI database. The genome sequences were aligned using ClustalW (v.2.1). The conservation of the start codon and its associated Shine-Dalgarno sequence was categorized into three groups: sequences (partially) conserved within the Salmonella genus, sequences conserved within the Salmonella genus, and those conserved across Enterobacteria (Table S1).

TIS mutagenesis and HiBiT tagging

Endogenous 3’ HiBiT-tagged strains were constructed using λ red-mediated recombineering essentially following the protocol outlined by Datsenko and Wanner (77). Briefly, pORTMAGE-2-transformed S. Typhimurium cells, with λ red expression induced by incubating the culture for 15 min at 42°C in a warm water bath at 250 rpm, were made electrocompetent. Electroporation of these cells, grown to an OD600 of 0.6, was conducted with a linear PCR-editing substrate designed for HiBiT tagging and introducing the kanamycin- (KmR) or chloramphenicol- (CmR) resistance cassette from pKD4 (accession number #7632, CGSS, Yale, USA) or pKD3, respectively (#45604, Addgene). PCR primers were designed to amplify the antibiotic-resistance cassette with 5’ and 3’ 50 bp homology arms complementary to the stop-codon flanking regions of yebG, clpB, and fruK using primers oPVDL2113 and oPVDL2114, oPVDL2191 and oPVDL2192, and oPVDL2309 and oPVDL2310, respectively (Table S3), and with the reverse primer incorporating an additional in-frame HiBiT tag (44) (11 aa tag MVSGWRLFKKIS). For TIS mutagenesis, a modified ssDNA oligo (40 pmol) was added to the electroporation reaction for multiplexed recombineering (yebG: oPVDL2063; clpB: oPVDL2496; fruK: oPVDL2419; Table S3). The ssDNA oligo contained two phosphorothioate linkages (indicated by an asterisk) (and a 2-fluoro-uridine modified base [2FU]) to create yebG aTIS (G-51A; Chr: 1,933,051 G > A), fruK dbTIS (GTG-1CTT; Chr: 2,303,857 – Chr: 2,303,859 GTG > CTT), and clpB aTIS (GTG-445CTT; Chr: 2,804,656 – Chr: 2,804,658 GTG > CTT) mutant strains exclusively expressing YebGL, FruKL, and ClpBL, respectively. Immediately following electroporation, bacterial cells were recovered in pre-warmed (28°C) Super Optimal broth with Catabolite repression (SOC, consisting of 2% tryptone, 0.5% yeast extract, 10 mM NaCl, 2.5 mM KCl, 5 mM MgCl2, 10 mM MgSO4, and 20 mM glucose) media at 28°C and were incubated at 180 rpm for 3 h. Mutant colonies were selected after plating the cell suspension and ~20 h of incubation (upside down) on LB agar plates supplemented with 50 µg/mL kanamycin or 25 µg/mL chloramphenicol at 28°C, and another round of re-streaking. The success of TIS mutagenesis and HiBiT tagging was confirmed through colony PCR using primer sequences indicated in Table S3 and through subsequent Sanger sequencing of the resulting PCR products (Eurofins Genomics).

Total protein extraction and blotting analysis

WT and (TIS-mutated) HiBiT-tagged S. Typhimurium SL1344 strains (yebG::HiBiT KmR, clpB::HiBiT KmR, and fruK::HiBiT KmR) were cultured overnight in liquid cultures (OD600 4–5). Subsequently, cultures were diluted 1:200 to an OD600 of 0.02 in 12-mL LB-Kan medium in a T25 cell culture flask with a ventilated cap. The flasks were incubated at 37°C with agitation (180 rpm) until reaching ESP (OD600 2.0). Cultures (5 mL) (corresponding to ~1.25 10^9 bacterial cells) were harvested by centrifugation for 10 min at 6,000 g (4°C), and the supernatant was discarded. The bacterial pellets were washed in ice-cold D-PBS (1X Gibco, Cat. #14190144), transferred to a 1.5-mL Eppendorf tube, and centrifuged for 5 min at 6,000 g. Finally, the supernatant was discarded, and the bacterial pellets were stored at −80°C until further processing. The pellets were resuspended in 400 µL urea lysis buffer (9M urea, 50 mM ammonium bicarbonate NH4HCO3 [pH 7.9]) to achieve a resulting protein concentration of approximately 2 mg/mL. Next, samples were subjected to mechanical disruption by three cycles of freezing in liquid nitrogen and thawing in a water bath at room temperature. Additionally, two cycles of probe sonication (50% duty cycle; 30 s of sonication applying 1-s pulses) were executed using a Branson Sonifier 250 (Ultrasonic convertor) at an amplitude of 50. Cleared supernatant was obtained by centrifugation for 10 min at 16,000 g (4°C). Protein concentration was determined using the Bradford method (Bio-Rad DC Protein assay kit, Bio-Rad, Cat. #5000006) following the manufacturer’s instructions. 2X Tricine Sample Buffer (Bio-Rad, #1610739) containing 2% of β-mercaptoethanol was added to the samples, and equivalent amounts of total protein were analyzed by 1D SDS-PAGE on a 16.5% precast polyacrylamide Tris-Tricine-Criterion gel using 1X Tris/Tricine/SD buffer (100 mM Tris, 100 mM Tricine, 0.1% SDS, pH 8.3, 10X stock solution) (Bio-Rad, #1610744) at 150 V for 1 h 45 min. Proteins were subsequently transferred to a 0.45-µm nitrocellulose membrane for 30 min at 100 V using transfer buffer [380.71 mM Tris-HCl (Tris(hydroxymethyl)aminomethane hydrochloride) and 485.2 mM boric acid]. Next, the membrane was rinsed in Tris-buffered saline Tween-20 (TBS-T) (1X Tris-buffered saline [38.07 mM Tris-HCl and 148.87 mM NaCl] and 0.1% Tween-20, pH 7.5) for 15 min on an orbital shaking platform. LgBiT protein (Promega, Cat. #N401C; 1:200) was incubated with the membrane in TBS-T for 1 h at room temperature. Nano-Glo Luciferase Assay Substrate (Promega, Cat. #N113A; 1:2,000) was added and incubated for 5 min. Chemiluminescence was detected using the Odyssey LI-COR Fc infrared imaging system (LI-COR Biosciences, Odyssey Fc Imager model n° 2800). For subsequent loading control detection, the membrane was blocked for 30 min at room temperature using a 1:1 Intercept blocking buffer (LI-COR, Cat. #27-60001)/1X TBS-T, followed by 30-min incubation with streptavidin-Alexa Fluor 680 Conjugate (Invitrogen, Cat. #S32358; 1/5,000). Fluorescent detection of endogenously biotinylated protein served as the loading control and was performed using the Odyssey LI-COR Fc. Signal quantification was carried out using the Odyssey LI-COR Fc infrared imaging system analysis application (Empiria Studio Software), with normalization of the signal automatically conducted using a user-defined background area.

ACKNOWLEDGMENTS

This work was made possible through the support of the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (PROPHECY grant agreement no. 803972 to P.V.D.), the Fonds Wetenschappelijk Onderzoek (FWO-Vlaanderen, project number G051120N to P.V.D.), and by the Ghent University Concerted Research Actions (grant BOF23/GOA/001 to P.V.D.). I.F. received support from a personal postdoctoral fellowship granted by FWO-Vlaanderen, with grant number 12I5517N.

The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

I.F. and P.V.D. conceived the study. I.F. and P.V.D. designed the experiments. I.F., V.S., and P.V.D. performed the experiments. I.F., V.S., and P.V.D. analyzed the data. I.F. and P.V.D. wrote the manuscript, and all authors contributed to finalizing the manuscript text and gave approval to the final version of the manuscript.

Contributor Information

Petra Van Damme, Email: Petra.VanDamme@UGent.be.

Joerg Vogel, University of Würzburg, Würzburg, Germany.

Joseph Thomas Wade, Wadsworth Center, Albany, New York, USA.

DATA AVAILABILITY

The proteomics data used in this study have been deposited in the PRIDE repository under the accession number PXD029391. The Ribo-seq and Ribo-RET data sets and their associated metadata (16) are available through the Open Science Framework at https://osf.io/h3vxz/?view_only=8e69e3e04f2d43119595237436b42389, with a DOI of 10.17605/OSF.IO/H3VXZ, and at https://osf.io/3u2hd/?view_only=8fd6750f2eb44d9594aa89a69501ef9c, with a DOI of 10.17605/OSF.IO/3U2HD.

SUPPLEMENTAL MATERIAL

The following material is available online at https://doi.org/10.1128/mbio.00333-24.

Figure S1. mbio.00333-24-s0001.pdf.

Ribo-seq reveals translation of Nt-proteoform pairs in S. Typhimurium.

mbio.00333-24-s0001.pdf (470.5KB, pdf)
DOI: 10.1128/mbio.00333-24.SuF1
Figure S2. mbio.00333-24-s0002.tif.

Peptide detectability scores for the longest proteoform of identified N-terminal proteoform pairs.

mbio.00333-24-s0002.tif (687.6KB, tif)
DOI: 10.1128/mbio.00333-24.SuF2
Legends. mbio.00333-24-s0003.docx.

Supplemental figure and table legends.

mbio.00333-24-s0003.docx (36.6KB, docx)
DOI: 10.1128/mbio.00333-24.SuF3
Table S2. mbio.00333-24-s0004.xlsx.

Physiochemical properties analyses of the members of identified N-terminal proteoform pairs identified in S. Typhimurium.

mbio.00333-24-s0004.xlsx (20.9KB, xlsx)
DOI: 10.1128/mbio.00333-24.SuF4
Table S3. mbio.00333-24-s0005.xlsx.

Primer sequences used.

mbio.00333-24-s0005.xlsx (14.6KB, xlsx)
DOI: 10.1128/mbio.00333-24.SuF5

ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.

REFERENCES

  • 1. Wood DE, Lin H, Levy-Moonshine A, Swaminathan R, Chang YC, Anton BP, Osmani L, Steffen M, Kasif S, Salzberg SL. 2012. Thousands of missed genes found in bacterial genomes and their analysis with COMBREX. Biol Direct 7:37. doi: 10.1186/1745-6150-7-37 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Warren AS, Archuleta J, Feng WC, Setubal JC. 2010. Missing genes in the annotation of prokaryotic genomes. BMC Bioinformatics 11:131. doi: 10.1186/1471-2105-11-131 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Nielsen P, Krogh A. 2005. Large-scale prokaryotic gene prediction and comparison to genome annotation. Bioinformatics 21:4322–4329. doi: 10.1093/bioinformatics/bti701 [DOI] [PubMed] [Google Scholar]
  • 4. Baek J, Lee J, Yoon K, Lee H. 2017. Identification of unannotated small genes in Salmonella. G3 (Bethesda) 7:983–989. doi: 10.1534/g3.116.036939 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Gray T, Storz G, Papenfort K. 2022. Small proteins; big questions. J Bacteriol 204:e0034121. doi: 10.1128/JB.00341-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Fijalkowski I, Peeters MKR, Van Damme P. 2021. Small protein enrichment improves proteomics detection of sORF encoded polypeptides. Front Genet 12:713400. doi: 10.3389/fgene.2021.713400 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Fields AP, Rodriguez EH, Jovanovic M, Stern-Ginossar N, Haas BJ, Mertins P, Raychowdhury R, Hacohen N, Carr SA, Ingolia NT, Regev A, Weissman JS. 2015. A regression-based analysis of Ribosome-profiling data reveals a conserved complexity to mammalian translation. Mol Cell 60:816–827. doi: 10.1016/j.molcel.2015.11.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Tripp HJ, Sutton G, White O, Wortman J, Pati A, Mikhailova N, Ovchinnikova G, Payne SH, Kyrpides NC, Ivanova N. 2015. Toward a standard in structural genome annotation for prokaryotes. Stand Genomic Sci 10:45. doi: 10.1186/s40793-015-0034-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Marcellin E, Licona-Cassani C, Mercer TR, Palfreyman RW, Nielsen LK. 2013. Re-annotation of the Saccharopolyspora erythraea genome using a systems biology approach. BMC Genomics 14:699. doi: 10.1186/1471-2164-14-699 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Fijalkowska D, Fijalkowski I, Willems P, Van Damme P. 2020. Bacterial riboproteogenomics: the era of N-terminal proteoform existence revealed. FEMS Microbiol Rev 44:418–431. doi: 10.1093/femsre/fuaa013 [DOI] [PubMed] [Google Scholar]
  • 11. Ingolia NT, Brar GA, Stern-Ginossar N, Harris MS, Talhouarne GJS, Jackson SE, Wills MR, Weissman JS. 2014. Ribosome profiling reveals pervasive translation outside of annotated protein-coding genes. Cell Rep 8:1365–1379. doi: 10.1016/j.celrep.2014.07.045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Ingolia NT, Ghaemmaghami S, Newman JRS, Weissman JS. 2009. Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science 324:218–223. doi: 10.1126/science.1168978 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Ingolia N.T, Lareau LF, Weissman JS. 2011. Ribosome profiling of mouse embryonic stem cells reveals the complexity and dynamics of mammalian proteomes. Cell 147:789–802. doi: 10.1016/j.cell.2011.10.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Crappe J, Ndah E, Koch A, Steyaert S, Gawron D, De Keulenaer S, De Meester E, De Meyer T, Van Criekinge W, Van Damme P, Menschaert G. 2015. PROTEOFORMER: deep proteome coverage through ribosome profiling and MS integration. Nucleic Acids Res 43:e29. doi: 10.1093/nar/gku1283 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Giess A, Jonckheere V, Ndah E, Chyżyńska K, Van Damme P, Valen E. 2017. Ribosome signatures aid bacterial translation initiation site identification. BMC Biol 15:76. doi: 10.1186/s12915-017-0416-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Ndah E, Jonckheere V, Giess A, Valen E, Menschaert G, Van Damme P. 2017. REPARATION: ribosome profiling assisted (re-)annotation of bacterial genomes. Nucleic Acids Res 45:e168. doi: 10.1093/nar/gkx758 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Meydan S, Marks J, Klepacki D, Sharma V, Baranov PV, Firth AE, Margus T, Kefi A, Vázquez-Laslop N, Mankin AS. 2019. Retapamulin-assisted ribosome profiling reveals the alternative bacterial proteome. Mol Cell 74:481–493. doi: 10.1016/j.molcel.2019.02.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Weaver J, Mohammad F, Buskirk AR, Storz G. 2019. Identifying small proteins by ribosome profiling with stalled initiation complexes. mBio 10:e02819-18. doi: 10.1128/mBio.02819-18 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Poptsova MS, Gogarten JP. 2010. Using comparative genome analysis to identify problems in annotated microbial genomes. Microbiology (Reading) 156:1909–1917. doi: 10.1099/mic.0.033811-0 [DOI] [PubMed] [Google Scholar]
  • 20. Haft DH, DiCuccio M, Badretdin A, Brover V, Chetvernin V, O’Neill K, Li W, Chitsaz F, Derbyshire MK, Gonzales NR, Gwadz M, Lu F, Marchler GH, Song JS, Thanki N, Yamashita RA, Zheng C, Thibaud-Nissen F, Geer LY, Marchler-Bauer A, Pruitt KD. 2018. RefSeq: an update on prokaryotic genome annotation and curation. Nucleic Acids Res 46:D851–D860. doi: 10.1093/nar/gkx1068 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Bland C, Hartmann EM, Christie-Oleza JA, Fernandez B, Armengaud J. 2014. N-Terminal-oriented proteogenomics of the marine bacterium roseobacter denitrificans Och114 using N-Succinimidyloxycarbonylmethyl)tris(2,4,6-trimethoxyphenyl)phosphonium bromide (TMPP) labeling and diagonal chromatography. Mol Cell Proteomics 13:1369–1381. doi: 10.1074/mcp.O113.032854 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Willems P, Fijalkowski I, Van Damme P. 2020. Lost and found: re-searching and re-scoring proteomics data AIDS genome annotation and improves proteome coverage. mSystems 5. doi: 10.1128/mSystems.00833-20 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Gawron D, Gevaert K, Van Damme P. 2014. The proteome under translational control. Proteomics 14:2647–2662. doi: 10.1002/pmic.201400165 [DOI] [PubMed] [Google Scholar]
  • 24. Thomas D, Plant LD, Wilkens CM, McCrossan ZA, Goldstein SAN. 2008. Alternative translation initiation in rat brain yields K2P2.1 potassium channels permeable to sodium. Neuron 58:859–870. doi: 10.1016/j.neuron.2008.04.016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Gawron D, Ndah E, Gevaert K, Van Damme P. 2016. Positional proteomics reveals differences in N-terminal proteoform stability. Mol Syst Biol 12:858. doi: 10.15252/msb.20156662 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Meydan S, Vázquez-Laslop N, Mankin AS. 2018. Genes within genes in bacterial genomes. Microbiol Spectr 6. doi: 10.1128/microbiolspec.RWR-0020-2018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Seol JH, Yoo SJ, Kim KI, Kang MS, Ha DB, Chung CH. 1994. The 65-kDa protein derived from the internal translational initiation site of the clpA gene inhibits the ATP-dependent protease Ti in Escherichia coli. J Biol Chem 269:29468–29473. doi: 10.1016/S0021-9258(18)43903-8 [DOI] [PubMed] [Google Scholar]
  • 28. Ozin AJ, Costa T, Henriques AO, Moran CP. 2001. Alternative translation initiation produces a short form of a spore coat protein in Bacillus subtilis. J Bacteriol 183:2032–2040. doi: 10.1128/JB.183.6.2032-2040.2001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Willems P, Ndah E, Jonckheere V, Van Breusegem F, Van Damme P. 2021. To new beginnings: riboproteogenomics discovery of N-terminal proteoforms in Arabidopsis thaliana. Front Plant Sci 12:778804. doi: 10.3389/fpls.2021.778804 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Jonckheere V, Van Damme P. 2021. N-terminal acetyltransferase Naa40p whereabouts put into N-terminal proteoform perspective. Int J Mol Sci 22:3690. doi: 10.3390/ijms22073690 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Yu XJ, Liu M, Matthews S, Holden DW. 2011. Tandem translation generates a chaperone for the Salmonella type III secretion system protein SsaQ. J Biol Chem 286:36098–36107. doi: 10.1074/jbc.M111.278663 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Long BM, Badger MR, Whitney SM, Price GD. 2007. Analysis of carboxysomes from Synechococcus PCC7942 reveals multiple Rubisco complexes with carboxysomal proteins CcmM and CcaA. J Biol Chem 282:29323–29335. doi: 10.1074/jbc.M703896200 [DOI] [PubMed] [Google Scholar]
  • 33. Wang H, Matsumura P. 1997. Phosphorylating and dephosphorylating protein complexes in bacterial chemotaxis. J Bacteriol 179:287–289. doi: 10.1128/jb.179.1.287-289.1997 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Plumbridge JA, Deville F, Sacerdot C, Petersen HU, Cenatiempo Y, Cozzone A, Grunberg-Manago M, Hershey JW. 1985. Two translational initiation sites in the infB gene are used to express initiation factor IF2 alpha and IF2 beta in Escherichia coli. EMBO J 4:223–229. doi: 10.1002/j.1460-2075.1985.tb02339.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Hodges LD, Lee LY, McNett H, Gelvin SB, Ream W. 2009. The Agrobacterium rhizogenes GALLS gene encodes two secreted proteins required for genetic transformation of plants. J Bacteriol 191:355–364. doi: 10.1128/JB.01018-08 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Kröger C, Colgan A, Srikumar S, Händler K, Sivasankaran SK, Hammarlöf DL, Canals R, Grissom JE, Conway T, Hokamp K, Hinton JCD. 2013. An infection-relevant transcriptomic compendium for Salmonella enterica serovar Typhimurium. Cell Host Microbe 14:683–695. doi: 10.1016/j.chom.2013.11.010 [DOI] [PubMed] [Google Scholar]
  • 37. Kröger C, Dillon SC, Cameron ADS, Papenfort K, Sivasankaran SK, Hokamp K, Chao Y, Sittka A, Hébrard M, Händler K, Colgan A, Leekitcharoenphon P, Langridge GC, Lohan AJ, Loftus B, Lucchini S, Ussery DW, Dorman CJ, Thomson NR, Vogel J, Hinton JCD. 2012. The transcriptional landscape and small RNAs of Salmonella enterica serovar Typhimurium. Proc Natl Acad Sci USA 109:E1277–86. doi: 10.1073/pnas.1201061109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Srikumar S, Kröger C, Hébrard M, Colgan A, Owen SV, Sivasankaran SK, Cameron ADS, Hokamp K, Hinton JCD. 2015. RNA-Seq brings new insights to the intra-macrophage transcriptome of Salmonella Typhimurium. PLoS Pathog 11:e1005262. doi: 10.1371/journal.ppat.1005262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Fijalkowski I, Willems P, Jonckheere V, Simoens L, Van Damme P. 2022. Hidden in plain sight: challenges in proteomics detection of small ORF-encoded polypeptides. Microlife 3:uqac005. doi: 10.1093/femsml/uqac005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Thorvaldsdóttir H, Robinson JT, Mesirov JP. 2013. Integrative genomics viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform 14:178–192. doi: 10.1093/bib/bbs017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Ashkenazy H, Abadi S, Martz E, Chay O, Mayrose I, Pupko T, Ben-Tal N. 2016. ConSurf 2016: an improved methodology to estimate and visualize evolutionary conservation in macromolecules. Nucleic Acids Res 44:W344–50. doi: 10.1093/nar/gkw408 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Fremin BJ, Bhatt AS. 2020. Structured RNA contaminants in bacterial Ribo-Seq. mSphere 5:e00855-20. doi: 10.1128/mSphere.00855-20 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Gao Z, Chang C, Yang J, Zhu Y, Fu Y. 2019. AP3: an advanced proteotypic peptide predictor for targeted proteomics by incorporating peptide digestibility. Anal Chem 91:8705–8711. doi: 10.1021/acs.analchem.9b02520 [DOI] [PubMed] [Google Scholar]
  • 44. Schwinn MK, Machleidt T, Zimmerman K, Eggers CT, Dixon AS, Hurst R, Hall MP, Encell LP, Binkowski BF, Wood KV. 2018. CRISPR-mediated tagging of endogenous proteins with a luminescent peptide. ACS Chem Biol 13:467–474. doi: 10.1021/acschembio.7b00549 [DOI] [PubMed] [Google Scholar]
  • 45. Wang HH, Xu G, Vonner AJ, Church G. 2011. Modified bases enable high-efficiency oligonucleotide-mediated allelic replacement via mismatch repair evasion. Nucleic Acids Res 39:7336–7347. doi: 10.1093/nar/gkr183 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Nagy M, Guenther I, Akoyev V, Barnett ME, Zavodszky MI, Kedzierska-Mieszkowska S, Zolkiewski M. 2010. Synergistic cooperation between two ClpB isoforms in aggregate reactivation. J Mol Biol 396:697–707. doi: 10.1016/j.jmb.2009.11.059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. McGlincy NJ, Ingolia NT. 2017. Transcriptome-wide measurement of translation by ribosome profiling. Methods 126:112–129. doi: 10.1016/j.ymeth.2017.05.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Ramachandran VK, Shearer N, Thompson A. 2014. The primary transcriptome of Salmonella enterica serovar Typhimurium and its dependence on ppGpp during late stationary phase. PLoS One 9:e92690. doi: 10.1371/journal.pone.0092690 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Chow IT, Baneyx F. 2005. Coordinated synthesis of the two ClpB isoforms improves the ability of Escherichia coli to survive thermal stress. FEBS Lett 579:4235–4241. doi: 10.1016/j.febslet.2005.06.054 [DOI] [PubMed] [Google Scholar]
  • 50. Eichinger V, Nussbaumer T, Platzer A, Jehl MA, Arnold R, Rattei T. 2016. EffectiveDB--updates and novel features for a better annotation of bacterial secreted proteins and Type III, IV, VI secretion systems. Nucleic Acids Res 44:D669–74. doi: 10.1093/nar/gkv1269 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Yu NY, Wagner JR, Laird MR, Melli G, Rey S, Lo R, Dao P, Sahinalp SC, Ester M, Foster LJ, Brinkman FSL. 2010. PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics 26:1608–1615. doi: 10.1093/bioinformatics/btq249 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Niemann GS, Brown RN, Mushamiri IT, Nguyen NT, Taiwo R, Stufkens A, Smith RD, Adkins JN, McDermott JE, Heffron F. 2013. RNA type III secretion signals that require Hfq. J Bacteriol 195:2119–2125. doi: 10.1128/JB.00024-13 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Lara-Tejero M, Qin Z, Hu B, Butan C, Liu J, Galán JE. 2019. Role of SpaO in the assembly of the sorting platform of a Salmonella type III secretion system. PLoS Pathog 15:e1007565. doi: 10.1371/journal.ppat.1007565 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Teufel F, Almagro Armenteros JJ, Johansen AR, Gíslason MH, Pihl SI, Tsirigos KD, Winther O, Brunak S, von Heijne G, Nielsen H. 2022. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nat Biotechnol 40:1023–1025. doi: 10.1038/s41587-021-01156-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Berry IJ, Steele JR, Padula MP, Djordjevic SP. 2016. The application of terminomics for the identification of protein start sites and proteoforms in bacteria. Proteomics 16:257–272. doi: 10.1002/pmic.201500319 [DOI] [PubMed] [Google Scholar]
  • 56. Davis RG, Park H-M, Kim K, Greer JB, Fellers RT, LeDuc RD, Romanova EV, Rubakhin SS, Zombeck JA, Wu C, Yau PM, Gao P, van Nispen AJ, Patrie SM, Thomas PM, Sweedler JV, Rhodes JS, Kelleher NL. 2018. Top-down proteomics enables comparative analysis of brain proteoforms between mouse strains. Anal Chem 90:3802–3810. doi: 10.1021/acs.analchem.7b04108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Clauwaert J, Menschaert G, Waegeman W. 2019. DeepRibo: a neural network for precise gene annotation of prokaryotes by combining ribosome profiling signal and binding site patterns. Nucleic Acids Res 47:e36. doi: 10.1093/nar/gkz061 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Fuchs S, Kucklick M, Lehmann E, Beckmann A, Wilkens M, Kolte B, Mustafayeva A, Ludwig T, Diwo M, Wissing J, Jänsch L, Ahrens CH, Ignatova Z, Engelmann S. 2021. Towards the characterization of the hidden world of small proteins in Staphylococcus aureus, a proteogenomics approach. PLoS Genet 17:e1009585. doi: 10.1371/journal.pgen.1009585 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Richardson EJ, Watson M. 2013. The automatic annotation of bacterial genomes. Brief Bioinform 14:1–12. doi: 10.1093/bib/bbs007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Storz G, Wolf YI, Ramamurthi KS. 2014. Small proteins can no longer be ignored. Annu Rev Biochem 83:753–777. doi: 10.1146/annurev-biochem-070611-102400 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Clauwaert J, Menschaert G, Waegeman W. 2019. DeepRibo: a neural network for precise gene annotation of prokaryotes by combining ribosome profiling signal and binding site patterns. Nucleic Acids Res 47:e36. doi: 10.1093/nar/gkz061 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Impens F, Rolhion N, Radoshevich L, Bécavin C, Duval M, Mellin J, García Del Portillo F, Pucciarelli MG, Williams AH, Cossart P. 2017. N-terminomics identifies Prli42 as a membrane miniprotein conserved in Firmicutes and critical for stressosome activation in Listeria monocytogenes. Nat Microbiol 2:17005. doi: 10.1038/nmicrobiol.2017.5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Costantino N, Court DL. 2003. Enhanced levels of λ red-mediated recombinants in mismatch repair mutants. Proc Natl Acad Sci USA 100:15748–15753. doi: 10.1073/pnas.2434959100 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Morel-Deville F, Vachon G, Sacerdot C, Cozzone AJ, Grunberg-Manago M, Cenatiempo Y. 1990. Characterization of the translational start site for IF2 beta, a short form of Escherichia coli initiation factor IF2. Eur J Biochem 188:605–614. doi: 10.1111/j.1432-1033.1990.tb15441.x [DOI] [PubMed] [Google Scholar]
  • 65. Geladaki A, Kočevar Britovšek N, Breckels LM, Smith TS, Vennard OL, Mulvey CM, Crook OM, Gatto L, Lilley KS. 2019. Combining LOPIT with differential ultracentrifugation for high-resolution spatial proteomics. Nat Commun 10:331. doi: 10.1038/s41467-018-08191-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Chen YX, Xu ZY, Ge X, Sanyal S, Lu ZJ, Javid B. 2020. Selective translation by alternative bacterial ribosomes. Proc Natl Acad Sci U S A 117:19487–19496. doi: 10.1073/pnas.2009607117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Nicholas RA, Lamson DR, Schultz DE. 1993. Penicillin-binding protein 1B from Escherichia coli contains a membrane association site in addition to its transmembrane anchor. J Biol Chem 268:5632–5641. [PubMed] [Google Scholar]
  • 68. Kato J, Suzuki H, Hirota Y. 1984. Overlapping of the coding regions for alpha and gamma components of penicillin-binding protein 1 b in Escherichia coli. Mol Gen Genet 196:449–457. doi: 10.1007/BF00436192 [DOI] [PubMed] [Google Scholar]
  • 69. Hoiseth SK, Stocker BA. 1981. Aromatic-dependent Salmonella typhimurium are non-virulent and effective as live vaccines. Nature 291:238–239. doi: 10.1038/291238a0 [DOI] [PubMed] [Google Scholar]
  • 70. Löber S, Jäckel D, Kaiser N, Hensel M. 2006. Regulation of Salmonella pathogenicity island 2 genes by independent environmental signals. Int J Med Microbiol 296:435–447. doi: 10.1016/j.ijmm.2006.05.001 [DOI] [PubMed] [Google Scholar]
  • 71. Nyerges Á, Csörgő B, Nagy I, Bálint B, Bihari P, Lázár V, Apjok G, Umenhoffer K, Bogos B, Pósfai G, Pál C. 2016. A highly precise and portable genome engineering method allows comparison of mutational effects across bacterial species. Proc Natl Acad Sci U S A 113:2502–2507. doi: 10.1073/pnas.1520040113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Bonde MT, Klausen MS, Anderson MV, Wallin AIN, Wang HH, Sommer MOA. 2014. MODEST: a web-based design tool for oligonucleotide-mediated genome engineering and recombineering. Nucleic Acids Res 42:W408–W415. doi: 10.1093/nar/gku428 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. 2021. Twelve years of SAMtools and BCFtools. Gigascience 10. doi: 10.1093/gigascience/giab008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74. Lauria F, Tebaldi T, Bernabò P, Groen EJN, Gillingwater TH, Viero G. 2018. riboWaltz: optimization of ribosome P-site positioning in ribosome profiling data. PLoS Comput Biol 14:e1006169. doi: 10.1371/journal.pcbi.1006169 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Dunn JG, Weissman JS. 2016. Plastid: nucleotide-resolution analysis of next-generation sequencing and genomics data. BMC Genomics 17:958. doi: 10.1186/s12864-016-3278-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76. Love MI, Huber W, Anders S. 2014. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15:550. doi: 10.1186/s13059-014-0550-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. Datsenko KA, Wanner BL. 2000. One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc Natl Acad Sci USA 97:6640–6645. doi: 10.1073/pnas.120163297 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1. mbio.00333-24-s0001.pdf.

Ribo-seq reveals translation of Nt-proteoform pairs in S. Typhimurium.

mbio.00333-24-s0001.pdf (470.5KB, pdf)
DOI: 10.1128/mbio.00333-24.SuF1
Figure S2. mbio.00333-24-s0002.tif.

Peptide detectability scores for the longest proteoform of identified N-terminal proteoform pairs.

mbio.00333-24-s0002.tif (687.6KB, tif)
DOI: 10.1128/mbio.00333-24.SuF2
Legends. mbio.00333-24-s0003.docx.

Supplemental figure and table legends.

mbio.00333-24-s0003.docx (36.6KB, docx)
DOI: 10.1128/mbio.00333-24.SuF3
Table S2. mbio.00333-24-s0004.xlsx.

Physiochemical properties analyses of the members of identified N-terminal proteoform pairs identified in S. Typhimurium.

mbio.00333-24-s0004.xlsx (20.9KB, xlsx)
DOI: 10.1128/mbio.00333-24.SuF4
Table S3. mbio.00333-24-s0005.xlsx.

Primer sequences used.

mbio.00333-24-s0005.xlsx (14.6KB, xlsx)
DOI: 10.1128/mbio.00333-24.SuF5

Data Availability Statement

The proteomics data used in this study have been deposited in the PRIDE repository under the accession number PXD029391. The Ribo-seq and Ribo-RET data sets and their associated metadata (16) are available through the Open Science Framework at https://osf.io/h3vxz/?view_only=8e69e3e04f2d43119595237436b42389, with a DOI of 10.17605/OSF.IO/H3VXZ, and at https://osf.io/3u2hd/?view_only=8fd6750f2eb44d9594aa89a69501ef9c, with a DOI of 10.17605/OSF.IO/3U2HD.


Articles from mBio are provided here courtesy of American Society for Microbiology (ASM)

RESOURCES