Abstract
Quantitative information on protein abundance is crucial to understand biological processes and is therefore frequently gathered in proteomic studies. However, the quality of a quantitative proteomic dataset is greatly affected by the number of missing values, which need to be minimized to produce robust and meaningful data. In this context, small proteins (≤100 amino acids) pose specific analytical challenges, which hinder their efficient identification and quantitative characterization in complex proteomes. In this study, methods for sample preparation and MS-data processing are systematically evaluated for their contribution to identification and quantification of small proteins of Clostridioides difficile 630 Δerm. Results show that small protein enrichment can enhance the number of identified and quantified proteins also for low abundant small proteins. Through application of spectral libraries for identification of MS spectra the number of robustly quantified proteins is increased and a lower limit of their detection is reached. Additionally, the dataset presented here is currently the most comprehensive protein repository for C. difficile covering 84.7% of the predicted proteome and 61.4% of all predicted small proteins of this important pathogen.
Keywords: mass spectrometry, peptidomics, low molecular weight proteome, Clostridioides difficile, SEPs, sProteins, database search, spectral library
This study systematically evaluates methods for improving the identification and quantification of small proteins using bottom-up mass spectrometry, thereby providing the most comprehensive dataset of proteins from the important pathogen Clostridioides difficile.
Introduction
Investigation of proteins in living systems continues to yield valuable insights into various biological processes across all domains of life. About 15 years ago, the primary objective of most proteomic experiments was the identification of expressed proteins using mass spectrometry (MS) to provide a qualitative understanding of protein composition and protein–protein interactions. However, to capture most physiological changes resulting from perturbations in a biological system, it is crucial to acquire quantitative information on protein abundance. Meanwhile, quantitative measurements form the core of nearly every proteomic study and have become routine in laboratories worldwide. However, there are specific challenges associated with the accurate quantitative description of proteomes. In commonly used bottom-up approaches, proteins are digested by a specific protease and the resulting peptides are analysed by liquid chromatography–tandem mass spectrometry (LC–MS/MS) (Zhang et al. 2013). In most cases, successful protein identification requires only detection of a sufficient number of associated peptides in at least one of the samples. In contrast, valid quantitative values under at least two different conditions need to be available to compare protein abundances. In order to statistically validate the quantitative information, these quantitative values need to be present in multiple biological replicates. Hence, the quality of a quantitative proteomic dataset is greatly affected by the number of missing values contained. Missing values appear due to different reasons. A protein might simply not be expressed in a given condition or its abundance might be below the limit of detection for the given experimental setup. Moreover, the stochastic nature of data-dependent MS measurements and applied filters for protein quantification (e.g. the necessity of identifying at least two unique peptides per protein) may contribute to the extent of missing values in a dataset. Although missing values can be handled differentially during data processing (e.g. by removing corresponding protein hits, by using only valid values, or by imputation of missing values), this data processing can have a significant impact on the results of the analyses. Therefore, it is desirable to generate a robust quantitative dataset with the least possible number of missing values.
Proteins with a very small size (here defined as “small proteins” with a length ≤100 amino acids) pose specific analytical challenges. They can be easily lost during sample preparation and, signals from larger proteins often mask them during ionization, separation, and detection in MS. These issues hinder their efficient identification and quantitative characterization. This led to the underrepresentation of small proteins in current proteomic studies although it has been shown that these proteins are often involved in important physiological functions, such as metabolism (Krauspe et al. 2021, Alvarenga-Lucius et al. 2023), signaling (Weel-Sneve et al. 2013, Voichek et al. 2020), virulence (Arnison et al. 2013, Venturini et al. 2020, Edelmann and Berghoff 2022), multiresistance (Melior et al. 2020), or regulation of enzyme activity (Yin et al. 2019, Gutt et al. 2021, Song et al. 2022). To meet the analytical challenges of small proteins, which are most often related to their short sequence and low abundance in a total protein background, different strategies have been proposed to successfully enrich small proteins from various samples. These approaches include membrane or gel filtration (Hu et al. 2007, Klein et al. 2007, Petruschke et al. 2021), two-dimensional LC-based prefractionation (Cassidy et al. 2021), differential precipitation (Cassidy et al. 2019), as well as solid phase extraction (Bartel et al. 2020). As small proteins vary in their physicochemical properties, it is improbable that a single method can effectively enrich all small proteins in an organism. In fact, a study comparing various sample preparation workflows and enrichment methods for small proteins found that the most comprehensive set of identified small proteins is achieved by combining different approaches (Petruschke et al. 2021). Beside sample enrichment, the utilization of modified digestion protocols can aid in the discovery of small proteins (Bartel et al. 2020, Müller et al. 2010, Swaney et al. 2010, Meyer et al. 2014, Osbak et al. 2016, Giansanti et al. 2016).
To enhance the number of identified small proteins further, the simultaneous utilization of multiple search engines has proven advantageous (Bartel et al. 2020, Müller et al. 2010, Wen et al. 2015) which is likely attributed to variations in spectrum preprocessing and scoring functions implemented in different search algorithms, leading to slightly different sets of reported peptides (Searle et al. 2008). MS-based identification of peptides is achieved by matching fragment ion spectra to peptide sequences and builds the base for protein identification, quantification, and thus biological interpretation (Mallick and Kuster 2010). Currently, the standard approach is database searching (Sinitcyn et al. 2018), where masses observed in actually acquired fragment spectra are compared to masses of theoretical spectra using computational methods. In this process, intensity information is disregarded. One approach to leverage the intensity information is to assemble spectral libraries from previous peptide identification data. These experimental spectral libraries are hypothesis-free regarding the content of the spectra but consider qualitative and quantitative characteristics of fragment spectra (Schubert et al. 2015, Griss 2016, Shao and Lam 2017), thus providing numerous advantages in terms of peptide validation and robust identification in extensive datasets. As experimental spectral libraries have the capability to accommodate nonstandard peaks, such as neutral losses, their application usually results in a higher number of identified proteins with less missing values in quantitative datasets (Junker et al. 2018a, Hentschker et al. 2020). However, generation of experimental spectral libraries requires labor and time intensive MS experiments as any peptide in the sample that lacks a corresponding library spectrum would be lost in the analysis. Alternatively, deep learning methods can meanwhile predict fragment spectra based on the amino acid sequence of any peptide (Gessulat et al. 2019, Gabriels et al. 2019, Cox 2023). Indeed, it was shown already that the additional information contained in predicted spectral libraries increase the number and correctness of peptide identification in standard proteome experiments (Silva et al. 2019, Gessulat et al. 2019, Tiwary et al. 2019). However, most prediction tools are currently not trained on nontryptic data or longer peptides, which might be obtained from semispecific cleavage or due to the very small size of a protein, making their use for small protein discovery and quantification nonoptimal. Consequently, we are not aware of any published study investigating the benefit of predicted spectral libraries for (small) protein quantification.
Over the last two decades, the anaerobic bacterium Clostridioides difficile has emerged as the primary cause of antibiotic-associated diarrhea (Smits et al. 2016). The treatment of C. difficile infections (CDI) is confronted by high rates of recurrent infections (Johnson 2009) and escalating numbers of multidrug-resistant clinical isolates (Peng et al. 2017). The clinical difficulties associated with CDI have spurred considerable efforts to understand how C. difficile modulates its virulence. Currently, there is a comprehensive body of literature (Table S1) with evidence for translation of up to 2385 proteins (and up to 119 small proteins) in a single study (Trautwein-Schult et al. 2018). However, there is still no evidence yet for the active translation of 1077 proteins (28.5% of the predicted proteome) (Dannheim et al. 2017) including 203 small proteins (58% of the 350 yet annotated small proteins).
Besides their identification, the quantitative comparison of small protein abundances poses an even greater challenge as quantitative values are often based on only a single peptide, preventing any statistical analysis within the protein. Consequently, systemic investigations on robust quantification of small proteins by nontargeted MS are still lacking. In this study, we systematically evaluated sample processing protocols and the performance of sequence databases (SD), experimental spectral libraries and predicted spectral libraries for the identification and quantification of proteins with special emphasis on small proteins (≤100 amino acids) in C. difficile. Finally, with the extensive dataset generated in this study, we enhanced the number of high-confident MS evidence for C. difficile proteins and report the extended proteome of this important pathogen.
Materials and methods
Bacterial strain and cultivation
Clostridioides difficile 630 Δerm (Hussain et al. 2005) was grown either in brain heart infusion (BHI; Oxoid, Basingstoke, UK) or chemically defined medium (CDMM) (Neumann-Schaal et al. 2015) in an anaerobic chamber (Don Whitley Scientific Ltd., Bingley, UK) at 37°C to exponential or early stationary phase. For this purpose, spores were germinated in BHI containing 0.1% (w/v) taurocholate for at least 24 h and precultures were inoculated by adding 0.1% (v/v) germinated spore solution to cold BHI or CDMM, respectively. After ~16 h, precultures were diluted to an OD600nm of 0.05 in the corresponding medium. When the cultures reached the desired growth phase, samples were obtained by centrifugation at 10 000 × g and 4°C for 5 min, cell pellets were washed twice with ice-cold Tris–HCl buffer (50 mM, pH 8.0), and stored at −20°C. Growth experiments were carried out in three independent replicates.
Sample preparation
For protein extraction, cell pellets were resuspended in Tris–HCl buffer and mechanically disrupted with 0.1 mm diameter glass beads for five homogenization cycles (6.5 m/s2) of 30 s each in a FastPrep-24 5G instrument (MP Biomedicals, Irvine, USA). Samples were cooled on ice for 5 min between the cycles. Cell debris and glass beads were removed by a mild centrifugation step at 8000 × g and 4°C for 5 min, which ensured minimal loss of membrane-associated small proteins. Protein concentration was determined by Bradford assay using bovine serum albumin as an external calibrant. Aliquots of 500 µg protein amount were used for small protein enrichment by solid phase extraction columns as previously described (Bartel et al. 2020). The enriched small protein fraction and total protein samples were digested with trypsin using S-Trap spin columns (Protifi, Huntington, USA) as described in Kroniger et al. (2022) except that Tris(2-carboxyethyl)phosphine instead of dithiothreitol was used as the reducing agent. Generated peptides were purified with 100 µl Pierce C18 Tips and their concentration was determined by Pierce Quantitative Fluorometric Peptide Assay (both Thermo Fisher Scientific, Waltham, USA) according to the manufacturer’s protocols. Prior to MS measurement, retention time calibration peptides (iRT, Biognosys, Schlieren, Switzerland) were added to the samples and constant peptide amounts were injected for total protein or small protein-enriched samples within each experiment.
In order to enhance the number and quality of spectra obtained in the experimental spectral libraries, 175 synthetic tryptic peptides of 55 C. difficile proteins (Table S2) were purchased from JPT Peptide Technologies GmbH (Berlin, Germany) in 96-well plate format. 250 µg of each peptide was pooled and the resulting stock solution was diluted in 10% (v/v) acetonitrile and 0.1% (v/v) acetic acid to reach a final concentration of 1 ng/µl for each peptide. Finally, 10 µl of a subsequent 1:10 dilution in 0.1% (v/v) acetic acid were injected for each LC–MS/MS run resulting in 1 ng of each peptide on column.
To determine the limits of small protein quantification, protein extraction was performed from one biological replicate grown in BHI to the early stationary phase as described above. After protein concentration determination, samples were spiked with a mixture of six purified small proteins from Haloferax volcanii H119 (see Supplemental material) in technical triplicates. Spike-in concentrations were selected to cover the range of low- and medium-abundant proteins and were: 0.5, 1.0, 5.0, 10, 50, 100, 250, and 500 pg/µg C. difficile proteome. After adding the Haloferax proteins, sample aliquots were stored immediately at −80°C and processed replicate-wise on different days as described above for the other samples. For spectral library generation, a mixture of 1 µg of each of the six purified Haloferax proteins was digested as described above.
MS
LC–MS/MS analyses were performed with an EASY-nLC II liquid chromatography system coupled to an LTQ Orbitrap Velos (Thermo Fisher Scientific). Peptides were loaded on a self-packed analytical column (OD 360 µm, ID 100 µm, and length 20 cm) filled with 3 µm diameter C18 particles (Dr. Maisch HPLC GmbH). Peptides were eluted using a binary nonlinear gradient of 5%–99% acetonitrile in 0.1% acetic acid over 157 min at a flow rate of 300 nl/min, and subjected to electrospray ionization-based MS. A full scan in the Orbitrap with a resolution of 60 000 was followed by CID of the twenty most abundant precursor ions. MS/MS experiments were acquired in the linear ion trap.
LC–MS/MS analyses of synthetic peptides were performed as described above, except that the peptides were eluted by a shorter 90-min binary nonlinear gradient of 5%–99% acetonitrile in 0.1% acetic acid.
SD and spectral libraries
Spectra from MS raw files were identified either by searching them against a SD or against spectral libraries. The SD contained 3781 entries for C. difficile 630 Δerm (Dannheim et al. 2017). Different spectral libraries were applied, namely an experimental spectral library (ExSpLib), a spectral library generated by machine learning (MaLeSpLib) and a merged spectral library (merged SpLib) where spectra predicted by machine learning were added for those peptides, which were not already contained in the ExSpLib.
The MS raw files used for the generation of the ExSpLib were obtained from previously published proteomics studies or were generated in this work. An overview on used input MS data can be found in Table S2. Data for generation of the ExSpLib were search with MSGF-Plus as this algorithm resulted in highest numbers of PSMs for CID spectra (Bartel et al. 2020). The same SD for C. difficile 630 Δerm (Dannheim et al. 2017) as described above but with added common laboratory contaminations and reverse entries was applied. Trypsin was specified as protease thereby assuming strictly tryptic peptides for total protein samples and a semitryptic digest for small protein enriched samples. Oxidation on methionine and carbamidomethylation on cysteine residues were selected as variable modification to consider that not all samples used for ExSpLib generation were reduced and alkylated during their preparation. Spectral library creation was performed according to Schubert et al. (2015) with slight modifications. Briefly, spectra and their particular identifications were linked together and hits were combined to interact.pep.XML files. Afterwards the PeptideProphet algorithm (Keller et al. 2002) was applied to adjust peptide identification probabilities. The following parameters were applied: minimum peptide sequence length of seven amino acids, use accurate mass binning, using mass errors in ppm, inclusion of decoy hits to pin down the negative distribution, exclusion of charge states 1+ and higher than 4+. Identifying spectra were filtered for a false discovery rate (FDR) of 0.01 on the level of peptide–spectrum matches. The resulting spectra were then imported into a spectral library. Spectra corresponding to the same identified peptide ion were merged to generate a representative consensus spectrum for a particular peptide species. The ExSpLib was subsequently extended with the same number of spectra as decoy hits generated by a random mass shift of the precursor and shuffling the peptide sequence (Yen et al. 2011).
Machine learning spectral libraries (MaLeSpLib) were generated based on the abovementioned SD for C. difficile 630 Δerm (Dannheim et al. 2017) supplemented with common laboratory contaminants. All protein entries in the database file were in silico digested with trypsin using the generate-peptides function of the crux toolkit (McIlwain et al. 2014). For digestion, full-digest was assumed and peptides were filtered to contain up to two missed cleavage sites, a length between 6 and 55 amino acids, a minimal mass of 400 Da and a maximal mass of 10 000 Da. For protein N-termini both variants, with and without the initiating methionine, were considered. Using in-house Python scripts, the resulting peptide list was filtered to contain unique peptides only and converted to the input format of the Prosit spectrum prediction tool (Gessulat et al. 2019). According to Prosit’s prediction rules, all cysteine residues were assumed to be carbamidomethylated and peptide variants with up to three oxidized methionine residues were added if applicable. Doubly, triply, and quadruply charged precursor variants were included if the precursor mass-to-charge ratio was within the range of 200–2000 Th. Raw spectral libraries in the .msp format were created using an in-house instance of Koina and spectra were predicted with the Prosit-CID2020 model. The raw spectral libraries were imported into SpectraST (Lam et al. 2007), spectra were formally converted into consensus spectra and a shuffled decoy variant was added for each precursor. For decoy-to-protein mapping, a decoy.fasta file was generated containing reversed protein entries concatenated with the decoy peptides separated by a “KRKR” spacer sequence.
The final ExSpLib and MaLeSplib were used for all datasets in this study except for those acquired to determine the limit of small protein quantification in data-independent acquisition (DIA) mode, as these were generated after higher-energy collision-induced dissociation. The ExSpLib contained 102 234 spectra representing 66 192 peptide sequences, which include information for 2345 C. difficile proteins. The MaLeSpLib contained 2 294 652 spectra from 521 062 peptides of 3690 C. difficile proteins.
To construct the merged spectral library (merged SpLib), spectra predicted by machine learning were added for those peptide ions which were not already present in the ExSpLib by combining both spectral libraries without their decoy entries. After generating consensus spectra, shuffled decoy entries were generated as described above, mapped to the forward protein sequences and added to the merged SpLib. The final merged SpLib contained 2 309 744 spectra representing 532 129 peptide sequences of 3700 C. difficile proteins.
Data processing
In order to identify spectra with the help of a classical SD, the widely applied software MaxQuant (v 2.0.3.0) was used. MS raw data were searched against a C. difficile 630 Δerm database (Dannheim et al. 2017) with 3781 entries. Common laboratory contaminants and reverse entries were added by MaxQuant. MaxQuant was used with the following parameters: primary digest reagent, trypsin, acetylation N, K (+42.0106) and oxidation M (+15.9949) as variable modifications. Results were filtered for a 1% FDR on spectrum, peptide, and protein levels. Match between runs with default parameters was enabled. In this study, the minimal number of unique peptides required for protein identification was set to one to meet the special requirements for small protein detection. MaxLFQ-values (Cox et al. 2014) were determined for protein quantification.
For identification of spectra with different spectral libraries, SpectraST (Lam et al. 2007) was used. By application of PeptideProphet (Keller et al. 2002) and ProteinProphet (Nesvizhskii et al. 2003) search results were combined on peptide and protein level and the corresponding FDRs were calculated. Output data were filtered for 1% FDR on spectrum, peptide, and protein level. IonQuant (Yu et al. 2021) was used to determine MaxLFQ-values for protein quantification. At least one ion was required for quantification and match between runs was enabled. Protein identification originating from a single unique peptide were only accepted, if the corresponding MS/MS-spectrum contained at least five consecutive b- or y-ions.
Sample processing for DIA data, acquired to determine the limits of small protein quantification, is described in the Supplemental material.
Results
Experimental design
This study was designed to systematically evaluate different MS data processing strategies for their performance in protein identification and protein quantification. Moreover, a special focus was put on small proteins (length ≤100 amino acids) as this challenging class of proteins has been comparatively understudied in eukaryotic and prokaryotic species.
In order to provide a comprehensive dataset which allows both, robust estimation of method performance and mapping of the so far hidden “small proteome” in C. difficile, this anaerobe Gram-positive bacterium was cultured in complex BHI and defined chemical medium and harvested during exponential growth and in stationary phase (Fig. 1). Each sample was generated with three independent biological replicates and further subjected to total protein extraction and enrichment of small proteins prior to MS analyses. For enrichment of small proteins, solid phase extraction was selected as a well-established exemplary method. An overview on alternative methods to enrich small proteins can be found elsewhere (Cassidy et al. 2021).
Figure 1.
To generate the dataset used in this study C. difficile was grown in two media (first column, purple), harvested in two different growth phases (second column, red) with three biological replicates each (third column, blue). The resulting samples were prepared with two different methods (last column, orange) to obtain either the total protein extract or enrich for small proteins.
Obtained mass spectra were identified either by searching them against a SD or against spectral libraries. While application of SDs for processing of mass spectrometric data is still the mostly used method for protein identification and quantification, searches against spectral libraries have proven to result in higher protein identification rates and more robust protein quantification (Lam et al. 2007, Junker et al. 2018b, Fernández-Costa et al. 2020, Hentschker et al. 2020) and are meanwhile standard when analysing mass spectra acquired in DIA mode. In this study, different types of spectral libraries were applied, namely an experimental spectral library (ExSpLib) containing actually acquired mass spectra, a spectral library containing spectra predicted by machine learning (MaLeSpLib), and a merged spectral library (merged SpLib) where spectra predicted by machine learning were added for those peptides, which were not already contained in the ExSpLib. More details on the SD, the generation of the spectral libraries and their sizes can be found in the section “Material and methods” and in Table S2. Following the quality criteria recently reported in the field of peptidomics and small protein identification (Slavoff et al. 2013, D’Lima et al. 2017), protein identification based on one single peptide were only accepted if the corresponding MS/MS-spectrum contained at least five consecutive b- or y-ions. The results obtained from each data processing strategy were examined in terms of number of protein identification (proteins identified with at least one unique peptide in any of the samples), and number of quantified proteins (identified and with an available quantitative value in at least two out of three biological replicates of any condition defined by medium and growth phase). All results were analysed based on total proteins but also with special emphasis on small proteins (Tables S3 and S4). Protein groups, representing <2% of the hits in the dataset, were removed for the analyses as grouping is slightly different in the search algorithms applied during database or spectral library search.
Importance of sample enrichment for quantification of small proteins
While it is well known that enrichment of small proteins can result in higher identification rates and may also lead to more robust protein identification between replicates (Cassidy et al. 2021, Fabre et al. 2021), this study examined the effect of small protein enrichment on their quantification by processing all samples either as total protein extract or enriched for small proteins using solid phase extraction (Bartel et al. 2020) prior to MS analyses (Fig. 1). As expected, the number of proteins identified by combination of all three identification strategies was reduced after small protein enrichment which is depleting larger proteins from the sample (Fig. 2A). Although enrichment also results in an overall lower number of identified small proteins, 26 small proteins could only be identified after enrichment (Fig. 2B).
Figure 2.
(A) Number of identified proteins after preparation of the total protein sample (orange) or after small protein enrichment by solid phase extraction (blue). (B) Number of identified small proteins (≤100 amino acids) after preparation of the total protein sample (orange) or after small protein enrichment. (C) Number of quantified proteins in total protein extracts (orange) and solid-phase enriched samples (blue). (D) Number of quantified small proteins in total protein extracts (orange) and solid-phase enriched samples (blue). A condition is defined by the used medium (BHI or CDMM) and growth phase (exponential growth or stationary phase) (Fig. 1).
Surprisingly, the quantification rate of small proteins after small protein enrichment did not improve compared to that found without enrichment. Still, enrichment for small proteins allows quantification of proteins, which cannot be detected in the total protein extract. Together with previously published results (Cassidy et al. 2019, Bartel et al. 2020, Petruschke et al. 2020) the data suggest that complementary sample preparation methods might be beneficial for both small protein identification and quantification.
Comparison of protein identification strategies
Utilizing spectral libraries for protein identification has demonstrated increased identification rates and more reliable protein quantification as compared to searches against SDs (Lam et al. 2007, Junker et al. 2018b, Fernández-Costa et al. 2020, Hentschker et al. 2020). Hence, it was tempting to explore the performance of spectral libraries also for quantification of the difficult to analyse group of small proteins. Moreover, a comparison was also made for experimental and predicted spectral libraries as both of them have different advantages and limitations. The dataset of this study (Fig. 1) was therefore processed with different strategies, namely a SD search as well as searches against an experimental spectral library (ExSpLib) and a predicted spectral library (MaLeSpLib) (Table S3). It is important to note that the approach of applying predicted spectral libraries differs from analysis strategies used in software such as DIA-NN and MSBooster, where spectra are predicted on-the-fly and are compared with extracted fragment masses or, in the case of MSBooster, used to rescore initial peptide identifications. This contrasts with our approach, which utilizes precalculated spectral libraries as a complete reference for spectral matching.
In concordance with the current knowledge, the number of identified proteins from total protein extracts and small protein -enriched samples depended on the search strategy with higher identification rates for the spectral library approaches and the MaLeSpLib leading to the highest number of identified proteins (Fig. 3A). This was also reflected in the number of identified small proteins (Fig. 3B).
Figure 3.
(A) Total number of identified proteins after data processing with a SD (orange), an experimental spectral library (ExSpLib, green), or a predicted spectral library (MaLeSpLib, blue). (B) Number of identified small proteins (≤100 amino acids) after data processing with SD (orange), ExSpLib (green), or MaLeSpLib (blue). (C) Number of quantified proteins after application of different search strategies. (D) Number of quantified small proteins after application of different search strategies. A condition is defined by the used medium (BHI or CDMM) and growth phase (exponential growth or stationary phase).
Of note, the reliable detection of a protein in any of the replicates of a given sample was sufficient to render a protein “identified.” In contrast, the definitions applied in this study consider a protein to be “quantified” if it has been detected in at least two replicates of a given condition. Although the MaLeSpLib performed best in (small) protein identification, protein quantification rates were differentially influenced by the protein identification strategy. Whereas searches against a SD or ExSpLib resulted in high protein quantification rates of ≥80% for all proteins identified, a significantly smaller fraction of proteins could be quantified with MaLeSpLib (34.2%) (Fig. 3C). The effect of the applied protein identification strategy was even more pronounced when small proteins were in the focus of the analyses. Whereas quantification rates for small proteins were 66.3% and 75.3% for searches against SD and ExSpLib, respectively, only 22.2% of the small proteins were quantified with searches against the MaLeSpLib (Fig. 3D). In an attempt to understand the driving factors of these observations in more detail, the number of protein identifications, assigned unique peptides, and the fraction of missing values were binned for protein sizes (Fig. 4, Table S4).
Figure 4.
The number of protein identifications, assigned unique peptides per protein, and the fraction of missing values for proteins identified after data processing with a SD (left column, orange), an experimental spectral library (ExSpLib, middle column, green), or a predicted spectral library (MaLeSpLib, right column, blue) was binned for the protein’s length. Detailled data are provided in Table S4.
Although the application of the MaLeSpLib allowed for a high number of identified proteins over all size bins as spectra for more peptide sequences are contained in this library, the number of assigned unique peptides for these proteins is significantly lower than for those proteins identified with SD or ExSpLib (Fig. 4). Indeed, the spectra quality in the MaLeSpLib is significantly different from those contained in the ExSpLib (Fig. S1). Ultimately, this leads to occasional peptide identification associated with a high number of missing values in the MaLeSpLib dataset, which results in a comparable low fraction of quantified proteins with this search approach (Fig. 3C and D). Additionally, in order to be able to compare the quantification performances of the different search approaches, this study employed the IonQuant algorithm (Yu et al. 2021) to determine LFQ values from spectral library searches. However, successful peak integration on MS1 levels requires multiple detection of the precursor ion mass in the analytical window of the MS-measurements rendering determination of LFQ values especially challenging for low abundant peptides. We hypothesize that these effects would have been less pronounced if quantification methods based on spectral counting would have been applied.
In order to potentially combine the advantages of ExSpLib and MaLeSpLib, both types of spectral libraries were merged, whereby experimentally acquired spectra were supplemented with predicted spectra of peptide ions not yet included in the ExSpLib. Interestingly, application of the merged SpLib yielded only slightly lower quantification rates than ExSpLib (Fig. 3C and D). In detail, quantification rates did not drop to the same extend as observed after processing with the MaLeSpLib, which demonstrates that this reduction is not mainly rooted in the vastly increased search space.
Limits of small protein quantification
To determine the limits of small protein quantification with the different workflows, a protein sample obtained from exponential growing C. difficile in BHI was spiked in technical triplicates with six purified small proteins from H. volcanii H119 ranging from 38 to 78 amino acids in length (HVO_0758, HVO_2212, HVO_2753, HVO_2922, HVO_2983, and HVO_A0101) (Supplemental material and Table S2). Spike-in concentrations were selected to cover the range of low- and medium-abundant proteins in a background proteome, hence ranging from 0.5 to 500 pg of spike-in protein per µg of C. difficile proteins. All samples were subjected to total protein preparation or enrichment of small proteins by solid-phase extraction prior to MS analyses. Identification of mass spectra was achieved by SD search or by application of an experimental spectral library (ExSpLib), or a predicted spectral library (MaLeSpLib).
Whereas HVO_0758, HVO_2212, and HVO_2922 (56, 78, and 60 amino acids long, respectively) were quantified frequently (Table 1), HVO_2983 (38 AA) could not be identified in any of the samples. HVO_A0101 (61 AA) and HVO_2753 (62 AA) were only identified in one sample and were therefore not considered to be robustly quantified in this study.
Table 1.
Quantification limits of small proteins in a bacterial cell lysate. Data were obtained either for total protein extracts (total) or after enrichment of small proteins (SEP). Data were derived from triplicate MS/MS experiments in DDA mode. Identification of mass spectra was achieved by SD search or by application of an experimental spectral library (ExSpLib) or a predicted spectral library (MaLeSpLib). The lowest concentration, in which a protein could be quantified (providing a quantitative value in at least 2 of 3 technical replicates) is given in pg per µg background proteome and colored in blue where darker color shades indicate a lower limit of detection. The coefficient of determination (R2) for the correlation of protein concentration to peak area is represented by orange bars with longer bars indicating higher R2.
A higher number of small proteins could be quantified and their limit of detection was lower when samples were specifically enriched for small proteins (Table 1). Together with the fact, that small protein enrichment gave rise to proteins not yet detected in the nonenriched sample (see the section “Importance of sample enrichment for quantification of small proteins”), this emphasizes the benefits of sample enrichment for robust quantification of small proteins. Moreover, the observed quantitative values showed good linearity for the three frequently quantified proteins, demonstrating that the additional processing steps for small protein enrichment do not introduce biases in the comparison of their abundances (Table 1).
When comparing different search strategies in the context of quantification limits, it becomes evident that the application of spectral libraries can enhance the number of quantifiable small proteins by providing confident identification of their peptides in a given dataset, especially when no enrichment of small proteins was carried out (Table 1). This is in line with the data obtained from the comparison of different search strategies (Fig. 3C and D).
Of note, the benefit of enrichment before small protein quantification was also significant, when samples were analysed in DIA mode (Table S5). Moreover, also the number of quantifiable small proteins is enhanced when samples are acquired in DIA mode compared to DDA. More detailed results on the limits of small protein quantification in DIA mode can be found in the Supplemental material.
The extended (small) proteome of C. difficile
In this study, a comprehensive dataset was generated (Fig. 1), which was processed by searching acquired mass spectra either against a SD or against different spectral libraries, namely the ExSpLib, MaLeSpLib, and a merged SpLib obtained from the latter. Appending all search results of this study on the level of identified proteins resulted in 3202 protein hits of which 1840 proteins have been identified independently with at least two of the search approaches. As 3781 proteins have been predicted by the most recent genome analyses (Dannheim et al. 2017) this represents an 84.7% coverage of the predicted proteome of C. difficile 630 Δerm. Due to the lower sensitivity and the restricted search space for SD and ExSpLib as compared to the MaLeSpLib, respectively, the identification rate in the MaLeSpLib was much larger (Fig. 3A), resulting in a proteome coverage of 48.7% if only these proteins are considered, which have been identified with at least two search approaches. Therewith, the dataset presented here is, to our knowledge, the most comprehensive protein repository for C. difficile 630 Δerm to date (Table S1).
The dataset processed with the ExSpLib contained the highest number of robustly quantified proteins (quantified in at least two out of three biological replicates of any condition defined by medium and growth phase). Hence, the ExSpLib derived data for 1459 proteins were used as base for a quantitative analyses of protein abundance of which 1308 proteins (Table S6) could be robustly quantified in at least two of the conditions enabling analyses of differential protein abundance. Indeed, 1213 of these proteins have also been reported earlier in Otto et al. (2016), who compared protein abundances of C. difficile 630 Δerm grown in the same media as used in this study. Whereas quantitative data of both studies were consistent, we herewith add information on differential protein abundance for 95 more proteins including six small proteins with up to 100 amino acid length. Additionally, this dataset adds information on changes in protein abundance in two different growth phases (exponential and stationary phase). Another recently published dataset also compares proteins extracted from growing and nongrowing cells in BHI and minimal medium (Trautwein-Schult et al. 2018). Filtering this dataset for statistical significance (two-way ANOVA, p<0.01) and abundance changes of at least 1.74-fold (corresponding to log2FC>|0.8|) revealed only 56 differently abundant proteins after metabolic labeling. In contrast, the current study reports 363 differentially abundant proteins thereby adding robust quantitative information with less missing values to the current knowledge.
Most of these 363 proteins (212) show differential protein abundance caused by the different media (Fig. 5A), whereas 167 proteins were altered dependent on the growth phase and for 81 proteins the medium influenced the differential abundance in the growth phase (termed “interaction” in Fig. 5A). In each of these groups, the highest number of proteins function in energy production and conversion, translation, and amino acid metabolism and transport. The higher expression of enzymes involved in vitamin and purine biosynthesis and the reduced abundance of proteins involved in butanoate fermentation in minimal medium reported by Otto et al. (2016) as well as the observed comparable abundance of proteins with functions in DNA metabolism, protein synthesis, and the cell envelope in the different media is also detectable in the dataset of this study.
Figure 5.
(A) Functional categorization of differentially abundant proteins (p<0.01, log2FC>|0.8|) when comparing growth of C. difficile 630 Δerm in complex BHI or defined chemical medium (medium), in exponential versus stationary phase (growth) or by both factors (interaction). Functional categorization is based on COG categories (Galperin et al. 2025). (B) Heat map representing small proteins (≤100 amino acids), which have been robustly quantified (quantitative value in at least two out of three biological replicates [BR1-3]) in at least two of the examined conditions (defined as the combination of medium [BHI: complex medium, CDMM: minimal medium] and growth phase [E: exponential, S: stationary). Heat map tiles are colored according to the protein abundance represented by the MaxLFQ value. The color in front of the protein names represents the functional categorization with the same colors used as in panel A. The first five lines of the heat map represent small proteins with differential abundance (p<0.01, log2FC>|0.8|).
In this study, we present first-time evidence for the actual translation of 639 proteins, out of which 82 represent small proteins with up to 100 amino acid length. Additionally, the so far hidden “small proteome” of C. difficile 630 Δerm could be extended to 215 proteins out of which 100 could be identified with at least two search approaches. Considering the 350 predicted small proteins for this organism (Dannheim et al. 2017), the coverage of the small proteome is 61.4% (28.6% if identification by at least two search approaches were required). Data processing with the ExSpLib allowed robust quantification of 46 small proteins, of which 33 were quantified in at least two conditions and can thus be used for differential abundance analyses (Fig. 5B). Given the poor characterization of many small proteins, it is not surprising that for a considerable fraction of the quantified small proteins (8) no function can be predicted. Out of the remaining 25 small proteins 32% (8) are small ribosomal proteins fulfilling functions during translation and another 16% (4) are represented by cold shock proteins categorized in the group “transcription.” Only five small proteins show a differential abundance (p<0.01, log2FC>|0.8|) in this study. In line with the current knowledge, the small ribosomal proteins RpsS was depleted in stationary phase compared to exponential growth. In contrast, the L14E/L6E/L27E-like ribosomal protein CDIF630erm_00162 was higher abundant during exponential growth in BHI compared to stationary phase in the same medium suggesting an adaption of the ribosomal protein complex in these conditions. The translocase SecE shows a similar pattern with higher abundance during exponential growth than in stationary phase in both of the media pointing at a reorganization of the Sec-translocon during transition from growing to nongrowing conditions. The subunit VorC1 of the 3-methyl-2-oxobutanoate dehydrogenase functioning in energy production and conversion as well as CDIF630erm_01777, the IIB component of a lactose/cellobiose-family PTS system, showed medium-dependent differential abundance with enhanced protein amounts when cells have been grown in minimal medium. This is most probably linked to different nutrient availability in the two media.
Discussion
In this study, different sample preparation methods and data processing strategies have been evaluated for (small) protein identification and quantification. Moreover, MS evidence for the translation of 639 proteins in C. difficile 630 Δerm is provided for the first time and 61.4% of the “small proteome” of this important human pathogen could be mapped.
Although it is well known that small protein identification can benefit from dedicated sample preparation methods (reviewed in Cassidy et al. 2021), the effects on robust quantification of this challenging class of proteins was not yet studied systematically. Although in this study, the quantification rate of small proteins did not improve following their enrichment, it enabled the quantification of proteins that are undetectable in the total protein extract. Moreover, their enrichment did improve the quantification limit for the spiked small proteins. Still, unlike published earlier (Bartel et al. 2020), not only fewer proteins were identified after enrichment (Fig. 2A and B), but also fewer peptides/ions/spectra per proteins were detected independent of protein size. This effect might be due to different digestion methods used in Bartel et al. (2020, in solution digest) and in this study (S-Trap digest). Although it has been reported multiple times, that S-Trap digests result in efficient protein digestion, high numbers of protein identifications, as well as sensitive and reproducible protein quantification (Ludwig et al. 2018, Antelo-Varela et al. 2019), the application of this protocol on small protein enriched samples seems to cause significant protein loss. In any case, according to the current knowledge (Ma et al. 2016, Cardon et al. 2020, Petruschke et al. 2020, Wang et al. 2021), the heterogeneity of small proteins in terms of their physicochemical properties would most probably require the combination of multiple complementary methodologies to improve both the number of identified small proteins in a sample as well as their robust quantification in biological datasets.
Besides the application of multiple sample preparation methods, subcellular fractionation of bacterial cell samples might support an enhanced (small) proteome coverage. However, a higher degree of sample fractionation is usually accompanied by the requirement for higher amounts of sample material. According to PSORTb (Yu et al. 2010) 54.6% of the 3781 annotated C. difficile proteins are predicted to be located in the cytoplasm whereas only 26.1% and 1.1% of the proteins are assigned to the membrane (including cell wall proteins) and extracellular fraction, respectively. This already implies that a higher sample volume is necessary to prepare the same amount of cytosolic, membrane, and secreted (small) proteins. Indeed, in Bacillus subtilis, which has a comparable distribution of subcellular protein localization than C. difficile, 300 times more cell culture needs to be harvested in stationary phase to prepare the same amount of extracellular proteins than compared to cytosolic proteins. If extracellular proteins are harvested at timepoints with less active secretion and/or cell lyses, even a 1000 times higher culture volume is necessary (F. Grilli, personal communication). These considerations multiply with the fact, that enrichment of small proteins by solid phase extraction usually starts with 50 times more starting material than needed for a nonenriched sample. However, even without subcellular fractionation this study was able to report 3202 identified proteins. Assuming 3256 protein-coding genes that are expressed during growth in BHI in late-exponential phase (Lamm-Schmidt et al. 2021), this study almost covers the complete potential proteome.
If, like in this study, multiple data processing approaches are used to enhance the number of protein identifications and thus to achieve a more robust protein quantification, this raises the question whether all hits provided by different algorithms are valid. Indeed, it has already been reported that diverse spectrum preprocessing and scoring functions in different search algorithms lead to marginally different sets of reported peptides (Searle et al. 2008). In the current dataset, 346 135 spectra could be assigned to C. difficile proteins out of which 201 422 (58.2%) were identified by at least two of the applied search approaches. Only 18 022 spectra (8.9% of the spectra identified by multiple search strategies) were assigned to different proteins by the different search approaches.
To compare protein abundances of (small) proteins in an unbiased manner, robust datasets with the least number of missing values are anticipated. The systematic evaluation of proteomic workflows in this study has shown that enrichment for small proteins can reduce the limit of detection for this challenging protein class. Moreover, the application of spectral libraries for data processing did not only result in a lower limit of detection of small proteins (Table S5) but also enhanced the number of robustly quantified proteins (Fig. 3D). In order to generate quantitative datasets with an even higher number of robustly quantified proteins, DIA approaches (Ludwig et al. 2018) might be a valuable option. DIA workflows have shown to exhibit higher reproducibility and proteomic depth compared to data-dependent acquisition (DDA) methods (Bruderer et al. 2017, Collins et al. 2017, Vowinckel et al. 2018). Indeed, here we found that the number of quantifiable small proteins is enhanced when samples are acquired in DIA mode compared to DDA (Table S6).
In this study, 215 small proteins (length ≤100 amino acids) of C. difficile, representing 61.4% of the predicted “small proteome”, could be detected by MS. Also other recent studies, which focus on the identification of small proteins, report a high number of small proteins but were only able to cover 34.6%, 27.6%, and 7.6% of the predicted small proteins in their model organisms B. subtilis (Bartel et al. 2020), H. volcanii (Hadjeras et al. 2023a), and Sinorhizobium meliloti (Hadjeras et al. 2023b), respectively, by mass-spectrometry alone. Therewith the current study represents one of the most comprehensive “small proteomes”, which is most likely attributed to the combination of sample preparation methods and broad application of data processing approaches.
This study is currently limited to the identification of already annotated open reading frames. In order to facilitate the discovery of not-yet annotated small proteins, proteogenomic approaches need to be applied. Such proteogenomic tools integrate genomics and proteomics to identify previously unannotated proteins with the aim to enhance or refine genome annotations (Nesvizhskii 2014). Using a six-frame translation-based protein database, Slavoff et al. (2013) successfully detected 86 novel small proteins in a human cell line. The application of integrated proteogenomics search databases (iPtgxDB) allowed for identification of 22, 11, and 3 novel small proteins in Bartonella henselae (Omasits et al. 2017), S. meliloti (Hadjeras et al. 2023b), and B. subtilis (Bartel et al. 2020), respectively. However, databases used in proteogenomic approaches are usually huge which complicates prediction of spectra from these databases by machine learning. The required computational power results in long processing times to generate corresponding spectral libraries. Although the application of resulting spectral libraries is not more time demanding than other search approaches, the time needed to generate the spectral libraries might only pay of if the spectral library is used for many and/or very comprehensive experiments.
Supplementary Material
Acknowledgements
We acknowledge Connor Wichern for support in sample preparation. We would like to thank Stefan Kemnitz for establishing our in-house instance of the Koina server as well as the developers of Koina (Mathias Wilhelm and Ludwig Lautenbacher) for their support during this.
Contributor Information
Jürgen Bartel, Department of Microbial Proteomics, Institute of Microbiology, University of Greifswald, 17489 Greifswald, Germany.
Vaikhari Kale, Department of Microbial Proteomics, Institute of Microbiology, University of Greifswald, 17489 Greifswald, Germany.
Dennis Joshua Pyper, Center of Biomolecular Magnetic Resonance (BMRZ), Institute for Organic Chemistry and Chemical Biology, Goethe University Frankfurt, 60438 Frankfurt am Main, Germany.
Harald Schwalbe, Center of Biomolecular Magnetic Resonance (BMRZ), Institute for Organic Chemistry and Chemical Biology, Goethe University Frankfurt, 60438 Frankfurt am Main, Germany.
Sandra Maaß, Department of Microbial Proteomics, Institute of Microbiology, University of Greifswald, 17489 Greifswald, Germany.
Author contributions
S.M. conceptualized and designed the study; S.M., J.B., V.K., P.G., and D.J.P. contributed to methodology; S.M. investigated and visualized the study; S.M. wrote the original draft; S.M., J.B., D.J.P., and H.S. revised the manuscript; H.S. and S.M. supervised the study at their site; H.S. and S.M. were involved in funding acquisition. All authors read and approved the final version of the manuscript.
Conflicts of interest
The authors declare no conflict of interest.
Funding
This work was supported by the German Research Council (DFG) priority program (SPP) 2002 ‘Small proteins in Prokaryotes, an unexplored world’ (BE3869/5-1, BE3869/5–2, and SCHW701/21-1). Work at the Center for Biomolecular Magnetic Resonance is supported by the state of Hessen.
Data availability
The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium (http://proteomecentral.proteomexchange.org) via the PRIDE partner repository (Perez-Riverol et al. 2022) with the dataset identifier PXD065346.
References
- Alvarenga-Lucius L, Linhartová M, Schubert H et al. The high-light-induced protein SliP4 binds to NDH1 and photosystems facilitating cyclic electron transport and state transition in Synechocystis sp. PCC 6803. New Phytol. 2023;239:1083–97. 10.1111/nph.18987. [DOI] [PubMed] [Google Scholar]
- Antelo-Varela M, Bartel J, Quesada-Ganuza A et al. Ariadne’s thread in the analytical labyrinth of membrane proteins: integration of targeted and shotgun proteomics for global absolute quantification of membrane proteins. Anal Chem. 2019;91:11972–80. 10.1021/acs.analchem.9b02869. [DOI] [PubMed] [Google Scholar]
- Arnison PG, Bibb MJ, Bierbaum G et al. Ribosomally synthesized and post-translationally modified peptide natural products: overview and recommendations for a universal nomenclature. Nat Prod Rep. 2013;30:108–60. 10.1039/C2NP20085F. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bartel J, Varadarajan AR, Sura T et al. Optimized proteomics workflow for the detection of small proteins. J Proteome Res. 2020;19:4004–18. 10.1021/acs.jproteome.0c00286. [DOI] [PubMed] [Google Scholar]
- Bruderer R, Bernhardt OM, Gandhi T et al. Optimization of experimental parameters in data-independent mass spectrometry significantly increases depth and reproducibility of results. Mol Cell Proteomics. 2017;16:2296–309. 10.1074/mcp.RA117.000314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cardon T, Hervé F, Delcourt V et al. Optimized sample preparation workflow for improved identification of ghost proteins. Anal Chem. 2020;92:1122–9. 10.1021/acs.analchem.9b04188. [DOI] [PubMed] [Google Scholar]
- Cassidy L, Kaulich PT, Maaß S et al. Bottom-up and top-down proteomic approaches for the identification, characterization, and quantification of the low molecular weight proteome with focus on short open reading frame-encoded peptides. Proteomics. 2021;21:e2100008. 10.1002/pmic.202100008. [DOI] [PubMed] [Google Scholar]
- Cassidy L, Kaulich PT, Tholey A. Depletion of high-molecular-mass proteins for the identification of small proteins and short open reading frame encoded peptides in cellular proteomes. J Proteome Res. 2019;18:1725–34. 10.1021/acs.jproteome.8b00948. [DOI] [PubMed] [Google Scholar]
- Collins BC, Hunter CL, Liu Y et al. Multi-laboratory assessment of reproducibility, qualitative and quantitative performance of SWATH-mass spectrometry. Nat Commun. 2017;8:291. 10.1038/s41467-017-00249-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cox J. Prediction of peptide mass spectral libraries with machine learning. Nat Biotechnol. 2023;41:33–43. 10.1038/s41587-022-01424-w. [DOI] [PubMed] [Google Scholar]
- Cox J, Hein MY, Luber CA et al. Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ. Mol Cell Proteomics. 2014;13:2513–26. 10.1074/mcp.M113.031591. [DOI] [PMC free article] [PubMed] [Google Scholar]
- C Silva AS, Bouwmeester R, Martens L et al. Accurate peptide fragmentation predictions allow data driven approaches to replace and improve upon proteomics search engine scoring functions. Bioinformatics. 2019;35:5243–8. 10.1093/bioinformatics/btz383. [DOI] [PubMed] [Google Scholar]
- Dannheim H, Riedel T, Neumann-Schaal M et al. Manual curation and reannotation of the genomes of Clostridium difficile 630Δerm and Clostridium difficile 630. J Med Microbiol. 2017;66:286–93. 10.1099/jmm.0.000427. [DOI] [PubMed] [Google Scholar]
- D’Lima NG, Khitun A, Rosenbloom AD et al. Comparative proteomics enables identification of nonannotated cold shock proteins in E. coli. J Proteome Res. 2017;16:3722–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Edelmann D, Berghoff BA. A shift in perspective: a role for the Type I Toxin TisB as persistence-stabilizing factor. Front Microbiol. 2022;13:871699. 10.3389/fmicb.2022.871699. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fabre B, Combier J-P, Plaza S. Recent advances in mass spectrometry-based peptidomics workflows to identify short-open-reading-frame-encoded peptides and explore their functions. Curr Opin Chem Biol. 2021;60:122–30. 10.1016/j.cbpa.2020.12.002. [DOI] [PubMed] [Google Scholar]
- Fernández-Costa C, Martínez-Bartolomé S, McClatchy DB et al. Impact of the identification strategy on the reproducibility of the DDA and DIA results. J Proteome Res. 2020;19:3153–61. 10.1021/acs.jproteome.0c00153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gabriels R, Martens L, Degroeve S. Updated MS2PIP web server delivers fast and accurate MS2 peak intensity prediction for multiple fragmentation methods, instruments and labeling techniques. Nucleic Acids Res. 2019;47:W295–9. 10.1093/nar/gkz299. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Galperin MY, Vera Alvarez R, Karamycheva S et al. COG database update 2024. Nucleic Acids Res. 2025;53:D356–63. 10.1093/nar/gkae983. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gessulat S, Schmidt T, Zolg DP et al. Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning. Nat Methods. 2019;16:509–18. 10.1038/s41592-019-0426-7. [DOI] [PubMed] [Google Scholar]
- Giansanti P, Tsiatsiani L, Low TY et al. Six alternative proteases for mass spectrometry-based proteomics beyond trypsin. Nat Protoc. 2016;11:993–1006. 10.1038/nprot.2016.057. [DOI] [PubMed] [Google Scholar]
- Griss J. Spectral library searching in proteomics. Proteomics. 2016;16:729–40. 10.1002/pmic.201500296. [DOI] [PubMed] [Google Scholar]
- Gutt M, Jordan B, Weidenbach K et al. High complexity of glutamine synthetase regulation in Methanosarcina mazei: small protein 26 interacts and enhances glutamine synthetase activity. FEBS J. 2021;288:5350–73. 10.1111/febs.15799. [DOI] [PubMed] [Google Scholar]
- Hadjeras L, Bartel J, Maier L-K et al. Revealing the small proteome of Haloferax volcanii by combining ribosome profiling and small-protein optimized mass spectrometry. microLife. 2023;4:uqad001. 10.1093/femsml/uqad001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hadjeras L, Heiniger B, Maaß S et al. Unraveling the small proteome of the plant symbiont Sinorhizobium meliloti by ribosome profiling and proteogenomics. microLife. 2023;4:uqad012. 10.1093/femsml/uqad012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hentschker C, Maaß S, Junker S et al. Comprehensive spectral library from the pathogenic bacterium Streptococcus pneumoniae with focus on phosphoproteins. J Proteome Res. 2020;19:1435–46. 10.1021/acs.jproteome.9b00615. [DOI] [PubMed] [Google Scholar]
- Hu L, Li X, Jiang X et al. Comprehensive peptidome analysis of mouse livers by size exclusion chromatography prefractionation and nanoLC-MS/MS identification. J Proteome Res. 2007;6:801–8. 10.1021/pr060469e. [DOI] [PubMed] [Google Scholar]
- Hussain HA, Roberts AP, Mullany P. Generation of an erythromycin-sensitive derivative of Clostridium difficile strain 630 (630Δerm) and demonstration that the conjugative transposon Tn916ΔE enters the genome of this strain at multiple sites. J Med Microbiol. 2005;54:137–41. 10.1099/jmm.0.45790-0. [DOI] [PubMed] [Google Scholar]
- Johnson S. Recurrent Clostridium difficile infection: a review of risk factors, treatments, and outcomes. J Infect. 2009;58:403–10. 10.1016/j.jinf.2009.03.010. [DOI] [PubMed] [Google Scholar]
- Junker S, Maaß S, Otto A et al. Spectral library based analysis of arginine phosphorylations in Staphylococcus aureus. Mol Cell Proteomics. 2018;17:335–48. 10.1074/mcp.RA117.000378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Junker S, Maaß S, Otto A et al. Spectral library based approach to protein arginine phosphorylation in S. aureus. Mol Cell Proteomics. 2018b;17:335–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Keller A, Nesvizhskii AI, Kolker E et al. Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search. Anal Chem. 2002;74:5383–92. 10.1021/ac025747h. [DOI] [PubMed] [Google Scholar]
- Klein C, Aivaliotis M, Olsen JV et al. The low molecular weight proteome of Halobacterium salinarum. J Proteome Res. 2007;6:1510–8. 10.1021/pr060634q. [DOI] [PubMed] [Google Scholar]
- Krauspe V, Timm S, Hagemann M et al. Phycobilisome breakdown effector NblD is required to maintain the cellular amino acid composition during nitrogen starvation. J Bacteriol. 2022;204:JB0015821. 10.1128/jb.00158-21. [DOI] [Google Scholar]
- Kroniger T, Mehanny M, Schlüter R et al. Effect of iron limitation, elevated temperature, and florfenicol on the proteome and vesiculation of the fish pathogen Aeromonas salmonicida. Microorganisms. 2022;10:1735. 10.3390/microorganisms10091735. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lam H, Deutsch EW, Eddes JS et al. Development and validation of a spectral library searching method for peptide identification from MS/MS. Proteomics. 2007;7:655–67. 10.1002/pmic.200600625. [DOI] [PubMed] [Google Scholar]
- Lamm-Schmidt V, Fuchs M, Sulzer J et al. Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production. microLife. 2021;2:uqab004. 10.1093/femsml/uqab004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ludwig C, Gillet L, Rosenberger G et al. Data-independent acquisition-based SWATH-MS for quantitative proteomics: a tutorial. Mol Syst Biol. 2018;14:e8126. 10.15252/msb.20178126. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ludwig KR, Schroll MM, Hummon AB. Comparison of In-Solution, FASP, and S-Trap based digestion methods for bottom-up proteomic studies. J Proteome Res. 2018;17:2480–90. 10.1021/acs.jproteome.8b00235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma J, Diedrich JK, Jungreis I et al. Improved identification and analysis of small open reading frame encoded polypeptides. Anal Chem. 2016;88:3967–75. 10.1021/acs.analchem.6b00191. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mallick P, Kuster B. Proteomics: a pragmatic perspective. Nat Biotechnol. 2010;28:695–709. 10.1038/nbt.1658. [DOI] [PubMed] [Google Scholar]
- McIlwain S, Tamura K, Kertesz-Farkas A et al. Crux: rapid open source protein tandem mass spectrometry analysis. J Proteome Res. 2014;13:4488–91. 10.1021/pr500741y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Melior H, Maaß S, Li S et al. The leader peptide peTrpL forms antibiotic-containing ribonucleoprotein complexes for posttranscriptional regulation of multiresistance genes. mBio. 2020;11:e01027–20. 10.1128/mBio.01027-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meyer JG, Kim S, Maltby DA et al. Expanding proteome coverage with orthogonal-specificity α-lytic proteases. Mol Cell Proteomics. 2014;13:823–35. 10.1074/mcp.M113.034710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Müller SA, Kohajda T, Findeiss S et al. Optimization of parameters for coverage of low molecular weight proteins. Anal Bioanal Chem. 2010;398:2867–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nesvizhskii AI. Proteogenomics: concepts, applications and computational strategies. Nat Methods. 2014;11:1114–25. 10.1038/nmeth.3144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nesvizhskii AI, Keller A, Kolker E et al. A statistical model for identifying proteins by tandem mass spectrometry. Anal Chem. 2003;75:4646–58. 10.1021/ac0341261. [DOI] [PubMed] [Google Scholar]
- Neumann-Schaal M, Hofmann JD, Will SE et al. Time-resolved amino acid uptake of Clostridium difficile 630Δerm and concomitant fermentation product and toxin formation. BMC Microbiol. 2015;15:281. 10.1186/s12866-015-0614-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Omasits U, Varadarajan AR, Schmid M et al. An integrative strategy to identify the entire protein coding potential of prokaryotic genomes by proteogenomics. Genome Res. 2017;27:2083–95. 10.1101/gr.218255.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Osbak KK, Houston S, Lithgow KV et al. Characterizing the syphilis-causing Treponema allidum ssp. pallidum proteome using complementary mass spectrometry. PLOS Negl Trop Dis. 2016;10:e0004988. 10.1371/journal.pntd.0004988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Otto A, Maaß S, Lassek C et al. The protein inventory of Clostridium difficile grown in complex and minimal medium. Proteomics Clin Apps. 2016;10:1068–72. 10.1002/prca.201600069. [DOI] [Google Scholar]
- Peng Z, Jin D, Kim HB et al. Update on antimicrobial resistance in Clostridium difficile: resistance mechanisms and antimicrobial susceptibility testing. J Clin Microbiol. 2017;55:1998–2008. 10.1128/JCM.02250-16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Perez-Riverol Y, Bai J, Bandla C et al. The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res. 2022;50:D543–52. 10.1093/nar/gkab1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Petruschke H, Anders J, Stadler PF et al. Enrichment and identification of small proteins in a simplified human gut microbiome. J Proteomics. 2020;213:103604. 10.1016/j.jprot.2019.103604. [DOI] [PubMed] [Google Scholar]
- Petruschke H, Schori C, Canzler S et al. Discovery of novel community-relevant small proteins in a simplified human intestinal microbiome. Microbiome. 2021;9:55. 10.1186/s40168-020-00981-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schubert OT, Gillet LC, Collins BC et al. Building high-quality assay libraries for targeted analysis of SWATH MS data. Nat Protoc. 2015;10:426–41. 10.1038/nprot.2015.015. [DOI] [PubMed] [Google Scholar]
- Searle BC, Turner M, Nesvizhskii AI. Improving sensitivity by probabilistically combining results from multiple MS/MS search methodologies. J Proteome Res. 2008;7:245–53. 10.1021/pr070540w. [DOI] [PubMed] [Google Scholar]
- Shao W, Lam H. Tandem mass spectral libraries of peptides and their roles in proteomics research. Mass Spectrom Rev. 2017;36:634–48. 10.1002/mas.21512. [DOI] [PubMed] [Google Scholar]
- Sinitcyn P, Rudolph JD, Cox J. Computational methods for understanding mass spectrometry–based shotgun proteomics data. Annu Rev Biomed Data Sci. 2018;1:207–34. 10.1146/annurev-biodatasci-080917-013516. [DOI] [Google Scholar]
- Slavoff SA, Mitchell AJ, Schwaid AG et al. Peptidomic discovery of short open reading frame-encoded peptides in human cells. Nat Chem Biol. 2013;9:59–64. 10.1038/nchembio.1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smits WK, Lyras D, Lacy DB et al. Clostridium difficile infection. Nat Rev Dis Primers. 2016;2:16020. 10.1038/nrdp.2016.20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Song K, Baumgartner D, Hagemann M et al. AtpΘ is an inhibitor of FoF1 ATP synthase to arrest ATP hydrolysis during low-energy conditions in cyanobacteria. Curr Biol. 2022;32:136–148.e5. 10.1016/j.cub.2021.10.051. [DOI] [PubMed] [Google Scholar]
- Swaney DL, Wenger CD, Coon JJ. Value of using multiple proteases for large-scale mass spectrometry-based proteomics. J Proteome Res. 2010;9:1323–9. 10.1021/pr900863u. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tiwary S, Levy R, Gutenbrunner P et al. High-quality MS/MS spectrum prediction for data-dependent and data-independent acquisition data analysis. Nat Methods. 2019;16:519–25. 10.1038/s41592-019-0427-6. [DOI] [PubMed] [Google Scholar]
- Trautwein-Schult A, Maaß S, Plate K et al. A metabolic labeling strategy for relative protein quantification in Clostridioides difficile. Front Microbiol. 2018;9:2371. 10.3389/fmicb.2018.02371. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Venturini E, Svensson SL, Maaß S et al. A global data-driven census of Salmonella small proteins and their potential functions in bacterial virulence. microLife. 2020;1. 10.1093/femsml/uqaa002. [DOI] [Google Scholar]
- Voichek M, Maaß S, Kroniger T et al. Peptide-based quorum sensing systems in Paenibacillus polymyxa. Life Sci Allian. 2020;3:e202000847. 10.26508/lsa.202000847. [DOI] [Google Scholar]
- Vowinckel J, Zelezniak A, Bruderer R et al. Cost-effective generation of precise label-free quantitative proteomes in high-throughput by microLC and data-independent acquisition. Sci Rep. 2018;8:4346. 10.1038/s41598-018-22610-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang B, Hao J, Pan N et al. Identification and analysis of small proteins and short open reading frame encoded peptides in Hep3B cell. J Proteomics. 2021;230:103965. 10.1016/j.jprot.2020.103965. [DOI] [PubMed] [Google Scholar]
- Weel-Sneve R, Kristiansen KI, Odsbu I et al. Single transmembrane peptide DinQ modulates membrane-dependent activities. PLoS Genet. 2013;9:e1003260. 10.1371/journal.pgen.1003260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wen B, Du C, Li G et al. IPeak: an open source tool to combine results from multiple MS/MS search engines. Proteomics. 2015;15:2916–20. 10.1002/pmic.201400208. [DOI] [PubMed] [Google Scholar]
- Yen C-Y, Houel S, Ahn NG et al. Spectrum-to-spectrum searching using a proteome-wide spectral library. Mol Cell Proteomics. 2011;10F:M111.007666. [Google Scholar]
- Yin X, Wu Orr M, Wang H et al. The small protein MgtS and small RNA MgrR modulate the PitA phosphate symporter to boost intracellular magnesium levels. Mol Microbiol. 2019;111:131–44. 10.1111/mmi.14143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu F, Haynes SE, Nesvizhskii AI. IonQuant enables accurate and sensitive label-free quantification with FDR-controlled match-between-runs. Mol Cell Proteomics. 2021;20:100077. 10.1016/j.mcpro.2021.100077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu NY, Wagner JR, Laird MR et al. PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics. 2010;26:1608–15. 10.1093/bioinformatics/btq249. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Y, Fonslow BR, Shan B et al. Protein analysis by shotgun/bottom-up proteomics. Chem Rev. 2013;113:2343–94. 10.1021/cr3003533. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium (http://proteomecentral.proteomexchange.org) via the PRIDE partner repository (Perez-Riverol et al. 2022) with the dataset identifier PXD065346.






