Abstract
Background
The birth of new genes from non-coding sequences has been postulated to be preceded by a proto-gene phase, in which a sequence is translated into protein but does not exhibit hallmarks of a clear function. Despite the abundance of such proto-genes in bacterial genomes, the frequency of their emergence and whether they actually act as precursors of new genes in natural populations are still open questions.
Results
To address these issues, we applied a combination of transcriptomic, proteomic, and comparative genomic approaches to identify and analyze hundreds of novel bacterial protein-coding genes that have previously escaped annotation. These novel proteins, including many that are widely conserved across genera, display sequence properties indistinguishable from the non-coding regions of the genome, suggesting that the vast majority are evolving neutrally. Despite their abundance and high degree of taxonomic restriction, we were only able to rigorously establish the de novo emergence of one proto-gene within the history of Escherichia coli, highlighting the difficulty of detecting this mode of gene birth in bacterial genomes. Contrary to expectations, we discover that proto-genes emerge at a uniform rate across distant bacterial taxa despite significant differences in their genomic characteristics, suggesting the presence of taxon-specific mechanisms that regulate their origination and persistence.
Conclusions
Overall, our findings indicate that proto-genes regularly emerge in bacterial populations but that their sequence properties furnish little evidence that they serve as precursors to new genes.
Supplementary Information
The online version contains supplementary material available at 10.1186/s13059-025-03825-x.
Keywords: Proto-genes, De novo gene evolution, Bacteria, Mass spectrometry
Background
De novo gene evolution—the emergence of genes from non-coding sequences—has received considerable attention in recent years due to the growing number of genes reported to have originated by this process [1–3]. This mode of new gene origination is considered unlikely in bacteria on account of their paucity of non-coding DNA, despite the prevalence of novel, lineage-specific genes (“ORFans”) in most bacterial genomes [4, 5]. It has been hypothesized that de novo genes evolve from taxonomically restricted sequences that are transcribed and translated but are functionally ambiguous [6]. These prototype sequences, termed “proto-genes,” have been shown to emerge rapidly in experimental populations [7] and might be pervasive across bacterial genomes [8, 9]; however, their occurrence and contribution to the origination of new genes in natural populations have never been investigated.
The analysis of proto-genes, and novel genes lacking functional annotation in general, faces several technical challenges—the first being the manner in which they are experimentally assayed. Ribosome profiling-based surveys are generally too permissive and will potentially recognize artifacts due to stochastic expression or translation readthrough [10]. Alternatively, mass spectrometry (MS), although capable of providing a global assessment of the proteome, often fails to detect short, weakly expressed or highly hydrophobic proteins [11–13], which are characteristics common to new genes [14]. Additionally, MS databases designed to detect non-annotated proteins typically include all open reading frames found in the genome [15], and such large databases can lead to artifacts from false-positive identifications. Yet another difficulty arises when attempting to identify events of proto-gene emergence based on comparative genomics in that the initial determination of whether genes are confined to a particular taxonomic group is typically based on negative evidence—i.e., the failure to detect a homolog outside of focal taxa [16]. Owing to their short lengths, homology search procedures are particularly unreliable when it comes to recognizing proto-genes [17–20]. Computational approaches have been developed to detect newly emerged genes, but these rely on the preservation of syntenic disabling mutations across more than one outgroup lineages [21, 22]—likely an unattainable expectation in bacteria, whose genomes are purged of non-coding sequences due to a pervasive deletional bias [23, 24].
In this study, we circumvent these complications by developing new analytic procedures that integrate transcriptomic, Ribo-seq, and MS datasets for the discovery and evolutionary analysis of proto-genes in bacterial genomes. Despite significant differences in their genomic characteristics, we discover a uniform rate of proto-gene emergence across the bacterial species examined. Furthermore, we find that in terms of their sequence properties, the extant pool of novel proteins shows scant evidence of being new gene precursors.
Results
Novel non-annotated proteins in Escherichia coli detected by mass spectrometry
To amass a complete set of protein-coding sequences in the genome, we leveraged two comprehensive mass spectrometry (MS) datasets generated for E. coli strains REL606 and K-12 MG1655 at different growth phases and conditions [25–27] as well as new MS data that we generated for three additional strains of E. coli (ECOR 11, ECOR 27, ECOR 37) grown in rich media. Upon searching mass spectra against all ORFs encoded by the genome, we discovered that unannotated proteins and “decoys” (i.e., reversed protein sequences) were detected at comparable levels at the standard threshold of false discovery rate (whole proteome-level q value < 0.01). This situation was not improved even after applying a transcription-guided reduced search database to increase sensitivity (see Methods) or rescoring the peptide-spectra search results using deep-learning-based methods [28]. We therefore elected to manually analyze the fragmentation spectra of every unannotated peptide detected at a q value < 0.0001 (Fig. 1A), a threshold at which no decoy proteins were detected. This exhaustive analysis yielded a total of 39 novel proteins with at least one high-confidence peptide-spectra match (PSM) across the datasets (Fig. 1B, Additional file 1: Table S1).
Fig. 1.
Detection of unannotated proteins from MS data. A Annotated MS/MS scans of three representative peptides, one from each dataset. Gray peaks designated unassigned spectra, and M designates the precursor peptide. B Number of PSMs detected for each of the 39 unannotated proteins according to different detection thresholds. The number of individual mass spectrometry experiment within each dataset is mentioned within parentheses. C Length of total protein sequence covered by the detected peptide. Order of proteins is maintained from B. The protein marked with an asterisk (*) contains two detected contiguous peptides
Unannotated proteins averaged 59 amino acids in length, all but one of which manifested a single detectable peptide (Fig. 1C), half displayed one PSM at more stringent thresholds (Fig. 1B), and only two of the 39 unannotated proteins were detectable across more than one dataset (Additional file 1: Table S1). This mirrors the situation with annotated proteins that are identified on the basis of a single peptide, most of which (61.2%) were detected on the basis of just one high-confidence PSM.
Novel bacterial genes resemble non-coding regions of the genome
To account for limitations inherent to MS-based surveys for detecting unannotated or novel proteins [14, 29], we incorporated unannotated proteins (> 9 aa) identified in previous proteomic or translatomic investigations of bacterial genomes into our analyses [8, 9, 29–35] (Additional file 1: Table S3). Focusing on species having at least 150 novel proteins reported across various studies (Fig. 2), a total of 492, 108, and 588 proteins were compiled from Escherichia coli, Salmonella enterica, and Mycobacterium tuberculosis, respectively, after removing those that matched or resembled annotated genes. The vast majority of these proteins were detected by ribosome profiling, with a minority based on mass spectrometry (this study) or western blotting [34], with most unannotated proteins (456 of 492 in E. coli) being detected in only one dataset (Additional file 1: Table S2).
Fig. 2.
Workflow for detection, curation, and comparative genomic analysis of novel genes. E. coli is depicted as a representative example
To first answer questions concerning the precise genomic locations of novel genes, only the datasets of Stringer et al. [9] and Smith et al. [8] were used, as they constitute an unbiased sampling of unannotated proteins across the genome. Based on these results, the largest category of novel genes in E. coli consists of those embedded within annotated genes on the same strand, and strictly intergenic and antisense proteins constitute a minor portion to the total (Additional file 1: Table S2, Additional file 2: Fig. S1). This trend is not apparent in M. tuberculosis, in which intergenic, antisense, and same-sense embedded ORFs all occur in similar proportions.
Novel genes in E. coli show expression patterns intermediate to annotated and non-coding ORFs: when compared to non-coding ORFs, novel genes are expressed at higher levels and across a wider range of conditions (Fig. 3). Despite this difference, compositional properties of their protein sequences and codon distributions either do not differ significantly from the non-coding ORFs in a genome or exhibit a more extreme departure from annotated genes than non-coding ORFs (Fig. 4). This pattern is especially informative concerning the frequency of polar amino acid residues and codon adaptation indices, properties for which annotated and non-coding fractions of the genome display the most extreme biases.
Fig. 3.
Expression of non-annotated genes in comparison to annotated and non-coding ORFs. A Number of conditions in which different categories of genes are expressed according to two different thresholds. For unannotated (novel) genes, only those without overlap with an annotated gene in the same strand are shown. Datasets without at least 30 such genes have not been depicted. Transcription levels of untranslated ORFs are included as control. B Number of ORFs passing a minimum threshold of expression in at least one sample
Fig. 4.
Novel genes differ in sequence properties from annotated genes. Violin plots of sequences properties for A E. coli, B S. enterica, and C M. tuberculosis. Novel gene properties were compared with annotated and untranslated ORFs, with the color of stars depicting the category compared to which a significant difference was found. Statistical significance is indicated by the number of stars, with 1, 2, and 3 stars representing p values less than 0.05, 0.01, and 0.001, respectively
Taxon-specific trends in proto-gene emergence
Because novel bacterial genes are functionally ambiguous, in that they do not resemble annotated genes nor have been implicated in cellular function, the taxonomically restricted fraction of such elements can be considered the pool of bacterial “proto-genes” [6]. To characterize the properties and rate of emergence of these elements in different taxa, we identified all novel genes that are only found in their respective genera and species (hereafter called “ORFans”). A blastp search against all translated outgroup ORFs detected matches for only a small fraction of novel genes—33% for Escherichia and 26% for M. tuberculosis—prompting us to adopt a manual search strategy (Fig. 2) to identify the final list of ORFans. Only species-specific ORFans were investigated for M. tuberculosis, as its level of species-level divergence (i.e., average nucleotide identity to nearest outgroup taxon) is comparable to the genus-level divergence in Escherichia and Salmonella (80–83%).
Of the novel proteins identified from E. coli and M. tuberculosis, 48.3% and 34.5% were found to be restricted to the Escherichia genus and the M. tuberculosis species-complex, respectively (Additional file 1: Table S2). In contrast, just 16.7% of the novel genes in S. enterica were genus-specific ORFans, which might be attributable to the fact that 92.6% of these sequences (vs. 63.4% for E. coli) were partially or completely embedded in their annotated genes, which limits their divergence (Additional file 1: Table S2). Despite having more than twice the fraction of ORFans than S. enterica, M. tuberculosis has a similar number of proteins (86.6%) that overlapped annotated genes. Across all taxa considered, novel genes show a higher rate of divergence when compared to their annotated counterparts: only 2% (E. coli), 2.8% (S. enterica), and 3.9% (M. tuberculosis) of annotated genes were genus-specific, and even non-ORFan novel genes had a significantly more restricted taxonomic distribution than their annotated counterparts (Additional file 1: Table S2).
Our comparisons to outgroup genomes then identified the fraction of ORFans that trace to homologous non-coding sequences in outgroup genomes, an approach that diminishes the possibility of taxonomic restriction of genes being attributable to artifactual homology-detection failure. Not only did E. coli and M. tuberculosis have comparable numbers of ORFans (238 vs. 203), but a similar number of their ORFans showed synteny to homologous sequences in at least one outgroup genome (71 vs. 91, corresponding to 14.4% and 15.4% of all novel genes), indicating that ORFans arise in both taxa at similar rates. Similar fractions of outgroup-traceable genes are recovered from both taxa even when the requirement for synteny is removed (15.6% vs. 16.6%), when analyses are confined to ORFs that overlap annotated genes (13.4% vs. 13.9%), and when we include syntenic regions that share no homology with the ORFan (23.4% vs. 21.8%), all indicators that these finds are robust.
Because proto-genes are hypothesized to be the precursors to novel, lineage-specific genes, we investigated whether their sequence properties depart from non-coding regions of the genome and approach those of annotated genes. (Due to the paucity of detectable proto-genes in Salmonella, this analysis was confined to Escherichia and M. tuberculosis.) In neither taxon was there a significant departure in the codon or amino acid properties of proto-genes from those of non-coding ORFs in the genome, except in the direction opposite to that of annotated genes (Figs. 5 and 6). Additionally, newly emerged (i.e., ORFan-encoded) novel proteins and older novel proteins do not differ significantly in their expression patterns (Additional file 2: Fig. S2).
Fig. 5.

Comparison of sequence properties between annotated genes and novel genes across levels of conservation. A E. coli. B M. tuberculosis. Novel gene properties compared with both categories of annotated genes and untranslated ORFs, with the color of stars depicting the category compared to which a significant difference was found. Statistical significance in violin plots (top panels) is indicated by stars, with 1, 2, and 3 stars representing p values less than 0.05, 0.01, and 0.001, respectively
Fig. 6.
Sequence properties of annotated and novel genes contrasted with the untranslated regions of the genome. A E. coli. B M. tuberculosis
Emergence and spread of proto-genes within pangenomes
We next investigated the distribution of proto-genes in the bacterial pangenome by leveraging a large set of E. coli genomes that represent the taxonomic diversity of the species. Whereas annotated ORFan genes have restricted distributions within the pangenome (Additional file 2: Fig. S3), this pattern does not apply to proto-genes (Fig. 7). Of the 48 species-specific proto-genes, over half were present in all (or all but one) phylogroups. At the other extreme, five proto-genes are present in only a single strain and lack coding or non-coding homologs elsewhere in the pangenome, three of which had non-coding homologs outside the genus, suggesting acquisition via recent horizontal gene transfer.
Fig. 7.
Pangenome distribution of species-specific proto-genes in E. coli. Each column to the right of the phylogeny represents the distribution of a gene across lineages. Each tip in the phylogeny represents a lineage containing 10 related E. coli genomes [68]. A gene was considered present in a lineage if any of its genomes carried a homolog. (The four proto-genes with no homologs in any genome are not shown.) The below inset shows the sequence alignment of a de novo emerged proto-gene, in which the removal of a stop codon in E. coli K-12 MG1655 led to the formation of an open reading frame not present in outgroup genomes
We then sought evidence for de novo emergence of any proto-gene within the E. coli pangenome. Although a substantial fraction of proto-genes is traceable to non-coding sequences (Additional file 1: Table S2), only one met the criterion of having a shared disabling mutation in more than one outgroup lineage—a pattern indicative of de novo origin. The protein encoded by this ORF was 17AA in length, and its non-coding homologs were characterized either by a lack of start codon or presence of a stop codon in the coding region (Fig. 7).
Discussion
Based on evidence generated by mass spectrometry (MS), Ribo-seq, and western blots, bacterial genomes contain hundreds of novel, unannotated protein-coding genes whose existence would not be anticipated based on genome annotations and whose functions are unknown [8, 9, 15]. Across the bacterial taxa investigated, ~ 10–30% of the novel protein-coding genes are proto-genes that are restricted to the species in which they are found, suggesting they emerged recently in the history of the species. This mirrors the situation reported for the E. coli lineages propagated during experimental evolution [36, 37], in which proto-genes arise frequently over short timescales and persist for tens of thousands of generations [7].
Selection is considered the hallmark of protein functionality, and the effects of selection can be inferred either from the ratio of synonymous and nonsynonymous substitutions in aligned DNA sequences or, in the absence of homologs, by observing how sequence characteristics depart from those of the non-coding, non-functional fraction of the genome. By these approaches, previous studies detected signatures of selection in numerous unannotated bacterial protein-coding genes [8, 38], whereas others view the bulk of such unannotated proteins as evolving neutrally [9]. Despite their lack of clear functionality, as precursors of new genes, lineage-specific unannotated genes have been postulated to exhibit sequence properties distinct from the non-coding fraction of the genome [6]. We find that the codon and amino acid compositions specified by novel genes in both E. coli and M. tuberculosis are indistinguishable from the non-coding regions of the genome, and for some characteristics, they are even more dissimilar from annotated genes than are non-coding sequences due to their enrichment in intergenic regions.
The finding that novel genes do not clearly differ from the non-coding fraction of the genome is reinforced by their inconsistent translation. Most non-annotated proteins detected in the present study were detected only once across datasets, and we were unable to detect MS-based evidence for all but two of the novel proteins previously detected by Ribo-seq. This is in agreement with the fact that studies find MS-based evidence for few (or no) Ribo-seq-validated novel proteins [8, 39–41]. In that the transcription levels of novel genes are comparable to that of annotated genes, the inconsistent detection of proteins encoded by novel genes could result from low levels of translation [14], which, in turn, supplies little substrate for selection. The rarity of these proteins helps explain why their coding sequences do not manifest characteristics of annotated genes and instead mimic the untranslated regions of the genome. Moreover, we find that of the 37 E. coli proteins detected in more than one Ribo-seq dataset, 89% showed more robust transcription (TPM > 1 in at least 10 different growth conditions), compared to 63% of all novel genes. Due to the small number detected, it is not possible to draw general conclusions about the sequence properties and patterns of conservation of these more consistently translated proteins, except to note that about half were genus-specific.
To date, few surveys of unannotated bacterial proteins have considered MS evidence, and in most cases, very small numbers of novel proteins have been discovered [12, 39–41]. Similarly, our expansive search of MS datasets spanning 242 samples led us to validate only 39 unannotated proteins, with the majority supported by a single high-quality peptide-spectrum match (PSM) across all samples. These low numbers are due, in part, to technical and statistical considerations. For example, studies in yeast and bacteria suggest that a large fraction of short or unannotated proteins produce poorly detectable peptides by MS, and such proteins often cannot be reliably detected using false discovery thresholds [14, 29]—a concern that led to our manual validation of all spectra [8]. Combined with difficulties in detecting proteins with low rates of translation [14], these considerations suggest that the list of MS-validated proteins may substantially underestimate the total number of non-annotated proteins expressed in the cell.
Despite the reported abundance of ORFan genes in this species, we detected only one de novo emerged proto-gene in E. coli. A previous study estimated that only 0.2% of the 600,000 + ORFan gene families could be considered candidates for de novo gene emergence [42]. Together, these results highlight the difficulty of tracking new gene birth in bacteria, in which non-coding sequences, the putative raw material of new genes, are rapidly removed [23, 24] and phylogenetic inferences are complicated by high rates of gene exchange. The mode(s) of origin for the rest of the novel proteins remain unclear; however, their distributions across the pangenome provide a clue. Unlike horizontally transferred genes, which tend to be distributed in a few strains [43], most proto-genes are conserved across multiple E. coli phylogroups, suggesting that the modification of coding and/or non-coding sequences are the main drivers of proto-gene emergence in E. coli.
Based on evidence from two divergent species that vary greatly in their genomic characteristics, the rate of new gene emergence remains relatively constant across bacterial taxa. This result was unanticipated given the higher fraction of ORFans among annotated genes in M. tuberculosis, the ubiquity of non-canonical forms of translation, including leaderless translation, in this species [8], and the higher GC content of its genome leading to the more frequent formation of ribosome-binding site motifs. Cumulatively, these features, along with their higher median length of novel proteins (31AA, vs. 19AA in E. coli), are expected to yield a greater abundance of proto-genes in M. tuberculosis. These findings suggest that the birth and retention proto-genes is not a neutral process driven by mutation alone, but is subject some form of regulation. This leads to the possibility that larger fractions of the unannotated protein expression in E. coli constitute “noise” and fail to impact cell fitness [9], making them prone to evolve rapidly by neutral processes. In contrast, there appear to be mechanisms that control the expression of novel genes in M. tuberculosis.
Conclusions
Bacterial genomes have been hypothesized to contain substantial numbers of proto-genes that serve as substrates for new genes. Through the comprehensive analysis of MS and Ribo-seq datasets, we found that the timescale of the first step of protein-coding gene birth—i.e., non-coding sequences acquiring translation—is uniform across taxa, despite broad differences in their genomic characteristics and propensity for spurious translation. Whereas a substantial number of proto-genes were present in the E. coli genome, only one was found to have emerged by de novo processes, with the vast majority likely representing divergent members of coding or non-coding sequences.
Methods
To detect non-annotated proteins encoded by the E. coli genome, previously published mass spectrometry (MS) data from E. coli strains K-12 MG1655 and REL606 was obtained from [25–27]. In addition, new MS data was generated from whole-cell lysates of three strains in the ECOR collection (ECOR 11, ECOR 27, ECOR 37) [44]. These strains were chosen because they have the highest number of strain-specific annotated ORFans in their respective phylogroups, suggesting a greater likelihood of carrying newly emerged proto-genes. Whereas two biological replicates were included for each strain, two additional replicates were investigated to specifically detect small proteins found in ECOR 37, the strain with the highest number of genes in the ECOR collection.
MS data generation
To prepare whole-cell lysates, 5 mL of overnight cultures of E. coli grown in Luria–Bertani broth were centrifuged at 4000 g for 15 min at 4 °C. Cell pellets were resuspended in 400 µL of lysis buffer (2% sodium dodecyl sulfate, 0.1 M Tris–HCl, 0.1 M dithiothreitol) and incubated at 100 °C for 10 min. To further facilitate cell lysis, 200 mg of 0.1 mm zirconia-silicate beads (Cole-Parmer) was added to each tube and subjected to bead-beating for 10, 30-s intervals. Tubes were then centrifuged at 14,000g for 20 min to precipitate the beads, followed by the addition of 400 µL of 8 M urea to the resulting supernatant. To prepare lysates enriched for small proteins, whole-cell lysates from the ECOR 37 strain were extracted by the procedure above, which were then loaded on to Nanosep 30 K Omega ultracentrifugation filters (Pall) and centrifuged for 10 min, retaining the resulting flow-through.
Lysates were submitted to the University of Texas at Austin CBRS Biological Mass Spectrometry Facility (RRID:SCR_021728) for proteolytic digest and liquid chromatography-tandem mass spectrometry (LC–MS/MS). Proteins were digested overnight with Promega trypsin using the SP4 protocol [45], followed by desalting using Millipore U-C18 ZipTip pipette tips in accordance with the manufacturer’s protocol. Mass spectra (MS) were acquired using the Thermo Ultimate 3000 RSLCnano UPLC coupled to the Orbitrap Fusion, with MS/MS fragmentation performed using high-energy collision disassociation fragmentation [46]. Raw data has been submitted to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD062120 and 10.6019/PXD062120 [47].
MS database construction
To analyze mass spectra generated from the five strains of E. coli, we constructed databases of protein candidates by first extracting all open reading frames (ORFs) > 30 bp in all six frames of the respective genomes using the getorf tool [48]. Only ORFs beginning with ATG, GTG, or TTG start codons were considered, as all other codons account for only 0.05% of initiation events in E. coli [49]. These sequences were then dereplicated with the seqkit tool [50], yielding a total of 103,252 and 103,694 ORFs from the REL606 and K-12 MG1655 genomes, respectively. Any predicted ORF having the same stop codon as an annotated gene was replaced with the corresponding gene. For cases in which we detected a longer ORF than the one predicted by getorf, we included only the longer ORF. These sequences were used as databases for peptide-spectra searches for both previously published and newly generated MS data.
To maximize the sensitivity of the peptide-spectral search, we conducted a separate analysis of the REL606 genome by leveraging transcription data generated from 181 MS samples [25, 26]. For this analysis, amended databases for each sample were constructed by retaining only those ORFs showing evidence of transcription in each MS sample [25, 26]. RNA-seq raw reads were processed with trimmomatic [51], which removed adapter and low-quality bases from both ends, and reads shorter than 29 nucleotides in length were removed. The processed reads were mapped to the REL606 genome with bowtie2 using the “local” alignment option in “very sensitive” mode [52]. Reads mapping to each ORF were counted with htseq-count with the “nonunique-all” option [53]. Only ORFs with transcript per million (TPM) values > 0.5 were included in the database for the corresponding MS experiment. This processing resulted in a total of 3474 annotated genes and 69,791 ORFs that lacked annotation (termed “non-annotated”) present in at least one database, with the median count of 15,620 protein candidates in each database from which we uncovered proto-genes.
MS data analysis
The mass spectra generated in each experiment were searched against the corresponding databases using MS-GF + [54] using the following parameters: fully tryptic peptides were considered, with two missed cleavages allowed; a precursor mass tolerance set at 10 ppm; and carbamidomethylation of cysteine and oxidation of methionine were used as fixed and variable modifications, respectively. FDR values were calculated by enabling a target-decoy search, and proteins with at least one peptide with a proteome-wide q value < 0.0001 were retained as candidates, since no decoy proteins were detected at this threshold. In line with previous reports [14], we found that no statistical metric returned by MS-GF + or ms2rescore [28] can distinguish non-canonical proteins from decoy proteins. This result compelled us to validate all candidates by manually examining the fragmentation spectra corresponding to each peptide-spectral match using the Interactive Peptide Spectral Annotator [55]. Of the 207 candidates examined across all datasets, 39 proteins had at least one peptide which displayed good quality fragmentation spectra characterized by high-intensity peaks corresponding to multiple peptide fragments.
Novel protein datasets
In addition to the proteins identified by MS, we mined available datasets of studies that attempted to identify novel proteins across bacterial genomes. Considering only those studies published since 2015, we identified 33 ribosome profiling and/or MS-based analyses representing 21 bacterial species. For comparative purposes, we limited analyses to those species for which > 150 non-annotated proteins were identified across studies. Only three species—Escherichia coli, Salmonella enterica, and Mycobacterium tuberculosis—met this threshold. All but three of the studies conducted on these species were based on ribosome profiling, and two were based on MS (and one on western blots). The two MS-based studies were subsequently excluded as they identified a very small number of proteins (four in E. coli [56] and 18 in S. enterica [57]), and two additional E. coli studies were excluded due to the abnormally high number of proteins detected (n = 485) [58], or to the phylogenetic distance of the strain examined from the focal strains used in all other studies [59]. List of studies and relevant exclusion criteria are listed in Additional file 1: Table S3.
Curating novel proteins
Genomic coordinates of novel proteins were extracted from supplementary files of the respective studies, and only the sequences of proteins longer than 9 amino acids in length were retained. From this list, previously annotated proteins were excluded by a two-step process: Initially, each of the genomes from which these sequences were extracted were re-annotated with Prodigal, GeneMarks-2, and Balrog [60–62], and with SmORFinder, which is specifically designed to detect smaller proteins [63]. Several annotation programs were used to ensure depth in annotation coverage, and all genomes were subjected to the same annotation procedures to discount biases arising from mixing annotation methods [64]. Novel proteins that shared a stop codon with a protein annotated by these tools were removed. Next, all novel protein sequences were searched against annotated proteins of all ingroup and outgroup genomes (see below), and proteins with positive hits were excluded. These steps led to 492, 108, and 588 numbers of proteins for E. coli, S. enterica, and M. tuberculosis, respectively.
Identifying ORFans
To detect ORFans (i.e., novel proteins specific either to the genus or species in which they are found), we first constructed outgroup protein and genome databases. For each taxon, genus-excluded genome databases were assembled by compiling all complete bacterial genome sequences downloaded from GenBank (accessed 10/28/2024). Genomes were selected based on their associated taxonomic information and validated by calculating their average nucleotide identity to the respective focal strains. Species-specific outgroup databases were constructed by including genomes from the same genus, but of different species, from the AllTheBacteria database [65]. For protein databases, all annotated proteins and translated ORFs > 30 base pairs in length were compiled. Because genomes extracted from the AllTheBacteria database do not have associated gene annotations, genomes in the species-specific outgroup databases were subsequently annotated with Prodigal [59].
We then conducted blastp searches using novel proteins as queries against these outgroup protein databases. Proteins with matches in a database (query coverage > 60%, e-value < 0.001) were eliminated, yielding our initial list of taxon-specific ORFans. Due to their short lengths, genuine hits to novel proteins could go undetected by conventional blast search criteria. To account for this possibility, ORFans were subjected to a further search by first extracting their upstream and downstream flanking regions (the proximate 500-bp sequences and flanking annotated genes), which were queried against outgroup genomes. For genomes that contained regions homologous to both upstream and downstream sequences within a maximum distance of 10,000 bp, the sequence between the two homologous regions were extracted as potentially harboring homologs to the candidate ORFan. Nucleotide sequences of the candidates were then searched against these regions using a word size of 7 to account for their shorter lengths. All hits spanning at least 50% of the candidate were extracted, re-aligned with MAFFT [66], and visually evaluated to discount all cases in which a protein-coding hit (query coverage > 60%) was present in the outgroup. The same search strategy was implemented without the synteny-conservation requirement to account for cases in which homologs are found in outgroups, but in different genomic contexts.
Identifying de novo emerged proto-genes
To identify those proto-genes that arose de novo, we leveraged a phylogenetically well-resolved set of 450 E. coli genomes spanning the diversity of human-isolated bacterial lineages [67, 68]. All nucleotide sequences, either within or outside the species pangenome, with hits against genus-specific ORFans were re-aligned with MAFFT and manually examined. De novo status was inferred if non-coding homologs of the gene bore a shared disabling mutation in at least two outgroup taxa [22]. Lineage status for each E. coli genome was extracted from [68].
Categorizing and characterizing ORFs
Locations and overlaps of non-annotated ORFs relative to annotated genes were identified with the “intersect” utility of the bedtools package [69]. Annotated and non-annotated genes with outgroup hits exceeding the number of hits corresponding to the median percentile value of all non-ORFan genes were classified as “conserved.” Gene length, GC content, and amino acid frequencies were calculated with custom bash scripts, following the strategies outlined in [42, 70]. Amino acid metabolic costs were extracted from [71], and codon adaptation indices were calculated using the cai tool from EMBOSS package based on their species-specific codon usage data [48]. For statistical comparison between properties, pairwise Wilcoxon rank-sum tests were conducted with Bonferroni correction.
Supplementary Information
Additional file 1. Supplementary Tables 1–3.
Additional file 2. Supplementary Figures. 1–3.
Acknowledgements
We thank Kim Hammond for figure preparation and Dr. Zachary Ardern for their helpful comments. Mass spectrometry services were provided by the UT Austin Center for Biomedical Research Support Biological Mass Spectrometry Facility (RRID:SCR_021728).
Peer review information
Tim Sands was the primary editor of this article and managed its editorial process and peer review in collaboration with the rest of the editorial team. The peer-review history is available in the online version of this article.
Authors’ contributions
M.U. designed the study, performed the analyses and wrote and reviewed the manuscript. H. O. designed the study and wrote and reviewed the manuscript.
Funding
This work was supported by the National Institutes of Health (R35GM118038 to H.O.). The funders had no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Data availability
The mass spectrometry proteomics data generated in this study have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD062120 and 10.6019/PXD062120. Mass spectrometry and transcriptomics data obtained from published papers are cited in the Methods section. Scripts used for analyses conducted in this study are available at https://github.com/Hassan-1991/Protogene_properties.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Van Oss SB, Carvunis A-R. De novo gene birth. PLoS Genet. 2019;15:e1008160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Knowles DG, McLysaght A. Recent de novo origin of human protein-coding genes. Genome Res. 2009;19:1752–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zhang L, Ren Y, Yang T, Li G, Chen J, Gschwend AR, et al. Rapid evolution of protein diversity by de novo origination in Oryza. Nat Ecol Evol. 2019;3:679–90. [DOI] [PubMed] [Google Scholar]
- 4.Daubin V, Ochman H. Bacterial genomes as new gene homes: the genealogy of ORFans in E. coli. Genome Res. 2004;14:1036–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Karlowski WM, Varshney D, Zielezinski A. Taxonomically restricted genes in Bacillus may form clusters of homologs and can be traced to a large reservoir of noncoding sequences. Genome Biol Evol. 2023;15:evad023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Carvunis A-R, Rolland T, Wapinski I, Calderwood MA, Yildirim MA, Simonis N, et al. Proto-genes and de novo gene birth. Nature. 2012;487:370–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Uz-Zaman MH, D’Alton S, Barrick JE, Ochman H. Promoter recruitment drives the emergence of proto-genes in a long-term evolution experiment with Escherichia coli. PLoS Biol. 2024;22:e3002418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Smith C, Canestrari JG, Wang AJ, Champion MM, Derbyshire KM, Gray TA, et al. Pervasive translation in Mycobacterium tuberculosis. Elife. 2022. 10.7554/eLife.73980. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Stringer A, Smith C, Mangano K, Wade JT. Identification of novel translated small ORFs in Escherichia coli using complementary ribosome profiling approaches. J Bacteriol. 2021;204:JB0035221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Li C, Zhang J. Stop-codon read-through arises largely from molecular errors and is generally nonadaptive. PLoS Genet. 2019;15:e1008141. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ahrens CH, Wade JT, Champion MM. A practical guide to small protein discovery and characterization using mass spectrometry. J Bacteriol. 2022;204:e0035321. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Meier-Credo J, Heiniger B, Schori C, Rupprecht F, Michel H, Ahrens CH, et al. Detection of known and novel small proteins in Pseudomonas stutzeri using a combination of bottom-up and digest-free proteomics and proteogenomics. Anal Chem. 2023;95:11892–900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Kucharova V, Wiker HG. Proteogenomics in microbiology: taking the right turn at the junction of genomics and proteomics. Proteomics. 2014;14:2360–675. [DOI] [PubMed] [Google Scholar]
- 14.Wacholder A, Carvunis A-R. Biological factors and statistical limitations prevent detection of most noncanonical proteins by mass spectrometry. PLoS Biol. 2023;21:e3002409. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Omasits U, Varadarajan AR, Schmid M, Goetze S, Melidis D, Bourqui M, et al. An integrative strategy to identify the entire protein coding potential of prokaryotic genomes by proteogenomics. Genome Res. 2017;27:2083–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Entwistle S, Li X, Yin Y. Orphan genes shared by pathogenic genomes are more associated with bacterial pathogenicity. mSystems. 2019. 10.1128/mSystems.00290-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Weisman CM, Murray AW, Eddy SR. Many, but not all, lineage-specific genes can be explained by homology detection failure. PLoS Biol. 2020;18:e3000862. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McLysaght A, Hurst LD. Open questions in the study of de novo genes: what, how and why. Nat Rev Genet. 2016;17:567–78. [DOI] [PubMed] [Google Scholar]
- 19.Albà MM, Castresana J. On homology searches by protein blast and the characterization of the age of genes. BMC Evol Biol. 2007;7:53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Moyers BA, Zhang J. Phylostratigraphic bias creates spurious patterns of genome evolution. Mol Biol Evol. 2015;32:258–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Vakirlis N, Carvunis A-R, McLysaght A. Synteny-based analyses indicate that sequence divergence is not the main source of orphan genes. Elife. 2020;9. Available from: https://pubmed.ncbi.nlm.nih.gov/32066524/. [DOI] [PMC free article] [PubMed]
- 22.Vakirlis N, McLysaght A. Computational prediction of de novo emerged protein-coding genes. Methods Mol Biol. 2019;1851:63–81. [DOI] [PubMed] [Google Scholar]
- 23.Mira A, Ochman H, Moran NA. Deletional bias and the evolution of bacterial genomes. Trends Genet. 2001;17:589–96. [DOI] [PubMed] [Google Scholar]
- 24.Kuo C-H, Moran NA, Ochman H. The consequences of genetic drift for bacterial genome complexity. Genome Res. 2009;19:1450–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Houser JR, Barnhart C, Boutz DR, Carroll SM, Dasgupta A, Michener JK, et al. Controlled measurement and comparative analysis of cellular components in E. coli reveals broad regulatory changes in response to glucose starvation. PLoS Comput Biol. 2015;11:e1004400. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Caglar MU, Houser JR, Barnhart CS, Boutz DR, Carroll SM, Dasgupta A, et al. The E. coli molecular phenotype under different growth conditions. Sci Rep. 2017;7:1–15. [DOI] [PMC free article] [PubMed]
- 27.Mori M, Zhang Z, Banaei-Esfahani A, Lalanne J-B, Okano H, Collins BC, et al. From coarse to fine: the absolute Escherichia coli proteome under diverse growth conditions. Mol Syst Biol. 2021;17:e9536. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Declercq A, Bouwmeester R, Hirschler A, Carapito C, Degroeve S, Martens L, et al. MS2Rescore: data-driven rescoring dramatically boosts immunopeptide identification rates. Mol Cell Proteomics. 2022;21:100266. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Fijalkowski I, Willems P, Jonckheere V, Simoens L, Van Damme P. Hidden in plain sight: challenges in proteomics detection of small ORF-encoded polypeptides. MicroLife. 2022;3:uqac005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Venturini E, Svensson SL, Maaß S, Gelhausen R, Eggenhofer F, Li L, et al. A global data-driven census of Salmonella small proteins and their potential functions in bacterial virulence. MicroLife. 2020;1:uqaa002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Giess A, Jonckheere V, Ndah E, Chyżyńska K, Van Damme P, Valen E. Ribosome signatures aid bacterial translation initiation site identification. BMC Biol. 2017;15:76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Ndah E, Jonckheere V, Giess A, Valen E, Menschaert G, Van Damme P. Reparation: ribosome profiling assisted (re-)annotation of bacterial genomes. Nucleic Acids Res. 2017;45:e168. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Nakahigashi K, Takai Y, Kimura M, Abe N, Nakayashiki T, Shiwa Y, et al. Comprehensive identification of translation start sites by tetracycline-inhibited ribosome profiling. DNA Res. 2016;23:193–201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.VanOrsdel CE, Kelly JP, Burke BN, Lein CD, Oufiero CE, Sanchez JF, et al. Identifying new small proteins in Escherichia coli. Proteomics. 2018;18:e1700064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Weaver J, Mohammad F, Buskirk AR, Storz G. Identifying small proteins by ribosome profiling with stalled initiation complexes. MBio. 2019. 10.1128/mbio.02819-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Lenski RE, Rose MR, Simpson SC, Tadler SC. Long-term experimental evolution in Escherichia coli. I. adaptation and divergence during 2,000 generations. Am Nat. 1991;138:1315–41. [Google Scholar]
- 37.Tenaillon O, Barrick JE, Ribeck N, Deatherage DE, Blanchard JL, Dasgupta A, et al. Tempo and mode of genome evolution in a 50,000-generation experiment. Nature. 2016;536:165–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Fesenko I, Sahakyan H, Dhyani R, Shabalina SA, Storz G, Koonin EV. The hidden bacterial microproteome. Mol Cell. 2025;S1097–2765(25):00059–0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Yu S, Yang M, Xiong J, Zhang Q, Gao X, Miao W, et al. Proteogenomic analysis provides novel insight into genome annotation and nitrogen metabolism in Nostoc sp. PCC 7120. Microbiol Spectr. 2021;9:e00490–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Petruschke H, Schori C, Canzler S, Riesbeck S, Poehlein A, Daniel R, et al. Discovery of novel community-relevant small proteins in a simplified human intestinal microbiome. Microbiome. 2021;9:55. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Fuchs S, Kucklick M, Lehmann E, Beckmann A, Wilkens M, Kolte B, et al. Towards the characterization of the hidden world of small proteins in Staphylococcus aureus, a proteogenomics approach. PLoS Genet. 2021;17:e1009585. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Vakirlis N, Kupczok A. Large-scale investigation of species-specific orphan genes in the human gut microbiome elucidates their evolutionary origins. Genome Res. 2024;34:888–903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Treangen TJ, Rocha EPC. Horizontal transfer, not duplication, drives the expansion of protein families in prokaryotes. PLoS Genet. 2011;7:e1001284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Ochman H, Selander RK. Standard reference strains of Escherichia coli from natural populations. J Bacteriol. 1984;157:690–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Johnston HE, Yadav K, Kirkpatrick JM, Biggs GS, Oxley D, Kramer HB, et al. Solvent precipitation SP3 (SP4) enhances recovery for proteomics sample preparation without magnetic beads. Anal Chem. 2022;94:10320–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Gadush MV, Sautto GA, Chandrasekaran H, Bensussan A, Ross TM, Ippolito GC, et al. Template-assisted de novo sequencing of SARS-CoV-2 and influenza monoclonal antibodies by mass spectrometry. J Proteome Res. 2022;21:1616–27. [DOI] [PubMed] [Google Scholar]
- 47.Uz-Zaman MH, Ochman H. Propensity for proto-gene emergence in bacteria. 2025. ProteomeXchange Consortium via PRIDE. 10.6019/PXD062120.
- 48.Rice P, Longden I, Bleasby A. EMBOSS: the European Molecular Biology Open Software Suite. Trends Genet. 2000;16:276–7. [DOI] [PubMed] [Google Scholar]
- 49.Hecht A, Glasgow J, Jaschke PR, Bawazer LA, Munson MS, Cochran JR, et al. Measurements of translation initiation from all 64 codons in E. coli. Nucleic Acids Res. 2017;45:3615–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Shen W, Le S, Li Y, Hu F. SeqKit: a cross-platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS ONE. 2016;11:e0163962. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30:2114–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Langmead B, Salzberg SL. Fast gapped-read alignment with bowtie 2. Nat Methods. 2012;9:357–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Anders S, Pyl PT, Huber W. HTSeq–a Python framework to work with high-throughput sequencing data. Bioinformatics. 2015;31:166–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Kim S, Pevzner PA. MS-GF+ makes progress towards a universal database search tool for proteomics. Nat Commun. 2014;5:5277. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Brademan DR, Riley NM, Kwiecien NW, Coon JJ. Interactive peptide spectral annotator: a versatile web-based tool for proteomic applications. Mol Cell Proteomics. 2019;18:S193-201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.D’Lima NG, Khitun A, Rosenbloom AD, Yuan P, Gassaway BM, Barber KW, et al. Comparative proteomics enables identification of nonannotated cold shock proteins in E. coli. J Proteome Res. 2017;16:3722–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Willems P, Fijalkowski I, Van Damme P. Lost and found: re-searching and re-scoring proteomics data aids genome annotation and improves proteome coverage. mSystems. 2020. 10.1128/mSystems.00833-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Hücker SM, Ardern Z, Goldberg T, Schafferhans A, Bernhofer M, Vestergaard G, et al. Discovery of numerous novel small genes in the intergenic regions of the Escherichia coli O157:H7 Sakai genome. PLoS ONE. 2017;12:e0184119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Neuhaus K, Landstorfer R, Fellner L, Simon S, Schafferhans A, Goldberg T, et al. Translatomics combined with transcriptomics and proteomics reveals novel functional, recently evolved orphan genes in Escherichia coli O157:H7 (EHEC). BMC Genomics. 2016;17:133. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics. 2010;11:119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Sommer MJ, Salzberg SL. Balrog: a universal protein model for prokaryotic gene prediction. PLoS Comput Biol. 2021;17:e1008727. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Lomsadze A, Gemayel K, Tang S, Borodovsky M. Modeling leaderless transcription and atypical genes results in more accurate gene prediction in prokaryotes. Genome Res. 2018;28:1079–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Durrant MG, Bhatt AS. Automated prediction and annotation of small open reading frames in microbial genomes. Cell Host Microbe. 2021;29:121-131.e4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Weisman CM, Murray AW, Eddy SR. Mixing genome annotation methods in a comparative analysis inflates the apparent number of lineage-specific genes. Curr Biol. 2022;32:2632-2639.e2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Hunt M, Lima L, Anderson D, Bouras G, Hall M, Hawkey J, Schwengers O, Shen W, Lees JA, Iqbal Z. AllTheBacteria: all bacterial genomes assembled, available, and searchable. bioRxiv. 2024. 10.1101/2024.03.08.584059.
- 66.Katoh K, Misawa K, Kuma K-I, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 2002;30:3059–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Horesh G, Blackwell GA, Tonkin-Hill G, Corander J, Heinz E, Thomson NR. A comprehensive and high-quality collection of Escherichia coli genomes and their genes. Microb Genom. 2021;7(2):000499. [DOI] [PMC free article] [PubMed]
- 68.Zhao B, Lees JA, Wu H, Yang C, Falush D. Genealogical inference and more flexible sequence clustering using iterative-PopPUNK. Genome Res. 2023;33:988–98. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Quinlan AR, Hall IM. BEDtools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Yomtovian I, Teerakulkittipong N, Lee B, Moult J, Unger R. Composition bias and the origin of ORFan genes. Bioinformatics. 2010;26:996–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Akashi H, Gojobori T. Metabolic efficiency and amino acid composition in the proteomes of Escherichia coli and Bacillus subtilis. Proc Natl Acad Sci U S A. 2002;99:3695–700. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Citations
- Uz-Zaman MH, Ochman H. Propensity for proto-gene emergence in bacteria. 2025. ProteomeXchange Consortium via PRIDE. 10.6019/PXD062120.
Supplementary Materials
Additional file 1. Supplementary Tables 1–3.
Additional file 2. Supplementary Figures. 1–3.
Data Availability Statement
The mass spectrometry proteomics data generated in this study have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD062120 and 10.6019/PXD062120. Mass spectrometry and transcriptomics data obtained from published papers are cited in the Methods section. Scripts used for analyses conducted in this study are available at https://github.com/Hassan-1991/Protogene_properties.






