Skip to main content
Molecular & Cellular Proteomics : MCP logoLink to Molecular & Cellular Proteomics : MCP
. 2023 Aug 11;22(9):100631. doi: 10.1016/j.mcpro.2023.100631

What Can Ribo-Seq, Immunopeptidomics, and Proteomics Tell Us About the Noncanonical Proteome?

John R Prensner 1,2,, Jennifer G Abelin 3, Leron W Kok 4, Karl R Clauser 3, Jonathan M Mudge 5, Jorge Ruiz-Orera 6, Michal Bassani-Sternberg 7,8,9, Robert L Moritz 10, Eric W Deutsch 10, Sebastiaan van Heesch 4
PMCID: PMC10506109  PMID: 37572790

Abstract

Ribosome profiling (Ribo-Seq) has proven transformative for our understanding of the human genome and proteome by illuminating thousands of noncanonical sites of ribosome translation outside the currently annotated coding sequences (CDSs). A conservative estimate suggests that at least 7000 noncanonical ORFs are translated, which, at first glance, has the potential to expand the number of human protein CDSs by 30%, from ∼19,500 annotated CDSs to over 26,000 annotated CDSs. Yet, additional scrutiny of these ORFs has raised numerous questions about what fraction of them truly produce a protein product and what fraction of those can be understood as proteins according to conventional understanding of the term. Adding further complication is the fact that published estimates of noncanonical ORFs vary widely by around 30-fold, from several thousand to several hundred thousand. The summation of this research has left the genomics and proteomics communities both excited by the prospect of new coding regions in the human genome but searching for guidance on how to proceed. Here, we discuss the current state of noncanonical ORF research, databases, and interpretation, focusing on how to assess whether a given ORF can be said to be “protein coding.”

Keywords: Ribo-Seq, mass spectrometry, immunopeptidomics, microprotein, noncanonical ORF

Graphical Abstract

graphic file with name ga1.jpg

Highlights

  • Ribo-seq paired with proteomics-based methods optimally detects noncanonical ORFs.

  • Data quality and analytical pipelines impact the output of a Ribo-seq experiment.

  • Noncanonical ORF catalogs variably report both high- and low-stringency nominations.

  • A framework for standardized noncanonical ORF evidence will advance the field.

In Brief

The human genome encodes thousands of noncanonical ORFs along with protein-coding genes. As a nascent field, many questions about them remain: How many exist? Do they encode proteins? What evidence is needed for their verification? Central to these debates has been the advent of ribosome profiling (Ribo-Seq) to discern genome-wide ribosome occupancy and immunopeptidomics to detect peptides presented by major histocompatibility complex molecules. This article synthesizes the current state of noncanonical ORF research and proposes standards for their future investigation and reporting.


Defining the extent of RNA translation in the human genome—and the resulting proteins—has long been a major focus for biomedical research. Approximately 19,500 protein-coding genes, which produce ∼80,000 annotated protein coding isoforms, constitute the canonical proteome (1, 2, 3, 4, 5, 6). Yet, whether this catalog is comprehensive has recently undergone substantial debate spurred by sequencing-based advances in the analysis of ribosome translation, termed ribosome profiling (Ribo-Seq). Based on classical techniques used to isolate ribosome–RNA complexes, Ribo-Seq is an RNA sequencing–based approach that profiles ribosome-protected RNA fragments, precisely defining ORFs actively engaged by translating ribosomes (7, 8). As a tool to detect the translation of RNA, the precision of this methodology is unprecedented: from individual ribosome footprints, the exact codon being translated in a purified ribosome–RNA complex can be determined. Through the sequencing of hundreds of millions of ribosome footprints, a single Ribo-Seq experiment can therefore produce a detailed and accurate representation of a given sample’s translated RNAs, typically identifying ∼11,000 to 12,000 translated genes per sample (9, 10, 11), which is more similar to the ∼12,000 to 13,000 expressed protein-coding mRNAs detected in a given cell type (12) compared with the ∼9000 to 11,000 proteins per sample typically detected in mass spectrometry (MS) methods (13, 14).

In addition to confirming known protein coding sequences (CDSs), the high predictive power of Ribo-Seq has unveiled thousands of other genomic sites of ribosome translation. These are most commonly found within known mRNAs (i.e., different reading frames than canonical CDS regions) but also within transcripts annotated as long noncoding RNAs (lncRNAs), pseudogenes, or retroviral elements in the genome (7, 9, 11, 15, 16, 17, 18, 19, 20, 21, 22, 23). Ribo-Seq can also provide clues on previously missed N-terminal in-frame extensions to known CDSs, initiated at sites alternative to the classically annotated initiation codon (24, 25, 26, 27). The nomenclature and estimated abundance of noncanonical ORFs are listed in Figure 1A. For clarity, these ORFs are termed “noncanonical” to distinguish them from CDSs included in reference gene annotation—that is, Ensembl-GENCODE—even though their translation, to our knowledge, occurs through mechanisms of ribosome activity similar to that of CDSs. Throughout this text, the term “noncanonical ORF” is therefore defined as any ORF that is not an annotated CDS, an in-frame extension or truncation (either N-terminal or C-terminal), or an in-frame intron retention of an annotated CDS. For our purposes, we will be focusing on upstream ORFs (uORFs), upstream overlapping ORFs (uoORFs), internal ORFs that overlap the CDS but are translated in a different frame (intORFs), downstream overlapping ORFs (doORFs), downstream ORFs (dORFs), and lncRNA-ORFs (as in Fig. 1A). We will not discuss in depth ORFs that may be translated from pseudogenes (19), genomic retroviruses (28), or other repetitive sequences (29) (see Limitations section).

Fig. 1.

Fig. 1

An overview of noncanonical ORF types and detection methods.A, a schematic illustrating the standardized names of noncanonical ORF types, their relationship to known mRNAs, and current estimations of their abundance. B, generalized workflows for ribosome profiling (Ribo-Seq), tryptic proteome mass spectrometry, and human leukocyte antigen (HLA) immunopeptidomics. The schematic indicates general properties of sample preparation for these data types. CDS, coding sequence; dORF, downstream ORF; doORF, downstream overlapping ORF; intORF, internal ORF; lncRNA-ORF, ORF residing within an annotated lncRNA; uORF, upstream ORF; uoORF, upstream overlapping ORF.

Given these observations, the genomics community has been faced with the fundamental question: does the genome actually encode far more than the ∼19,500 protein-coding genes currently accepted as canonical? In response, there have been increasing efforts to corroborate the observations from Ribo-Seq using MS, with the overall conclusion that only a low percentage of noncanonical ORFs are detectable by conventional tryptic proteome methods employing liquid chromatography with tandem MS (LC–MS/MS) techniques (9, 15, 30, 31, 32, 33, 34). Yet, far more noncanonical ORFs appear to be detectable with immunopeptidomic approaches that profile peptides presented by the class I human leukocyte antigen (HLA-I) system (Fig. 1B) (34, 35, 36, 37, 38, 39). Moreover, independent of their protein-coding capacity, noncanonical ORFs may serve important roles in the regulation of mRNA translation (40, 41, 42). With these observations at hand, one of the central tasks for the proteomics and genomics communities alike is to develop a consensus understanding on what constitutes sufficient evidence of detection for a noncanonical ORF from each technology and how to standardize these assessments given the limitations of each methodology.

Types of Evidence for Noncanonical ORFs

Translated noncanonical ORFs can be detected by either Ribo-Seq or LC–MS/MS approaches, with examples of transition to canonical annotated protein-coding genes emerging from both. For example, translation of the signaling proteins, APELA (43), POLGARF (44, 45), TINCR (46), and the cardiac proteins, MYMX (47) and MRLN (48), was first identified using Ribo-Seq, whereas LC–MS/MS data provided the initial evidence for the translation products of uORFs in ASNSD1, MKKS, MIEF1, and SLC35A4 (30, 49).

Together, the combination of Ribo-Seq and LC–MS/MS is a powerful way to identify translated CDSs and ORFs (21, 50, 51, 52). Ribo-Seq does not directly detect proteins but rather provides evidence of ongoing nucleotide translation. By contrast, LC–MS/MS evidence for noncanonical ORFs takes the form of direct detection of peptides. In the case of conventional LC–MS/MS of cellular lysates, these peptides are typically tryptic, meaning they were generated by protein cleavage at the C-terminal side of a lysine or an arginine, or semitryptic, meaning they were generated by protein cleavage at the C-terminal side of a lysine or an arginine at one end of the peptide but not the other. However, many ORFs have now been observed in MS-based HLA-I immunopeptidomics data (18, 34, 36, 38, 53). Here, no tryptic digestion is employed. Instead, peptides containing the HLA-I peptide–binding motifs of the HLA-I allele expressed by a specific cell line or tissue are observed. A variety of lower-throughput approaches have also been used to assess translation of noncanonical ORFs, including generation of custom antibodies, expression of epitope-tagged ORF complementary DNAs, selective reaction monitoring, and radiolabeled in vitro translation (9, 17, 54, 55, 56).

While high-quality Ribo-Seq and LC–MS/MS tryptic proteome data on the same sample should be able to identify highly consistent sets of endogenous CDSs, Ribo-Seq is not able to pinpoint the responsible translation event for exogenous proteins, which originate from sources other than the sample’s own genetic material. Similarly, Ribo-Seq cannot detect or predict protein stability, folding, or post-translational modification (PTM). If there is a substantial discrepancy with MS detecting many additional proteins, then the quality of the Ribo-Seq library should be inspected (see later). It should also be noted that Ribo-Seq, like all sequencing-based methods, may not be able to resolve translation events in repetitive genomic regions, such as retrotransposons, pseudogenes, or genes with very high homology.

By contrast, Ribo-Seq will almost always detect many noncanonical ORFs that are not found by proteomics. This is due to several factors: both the nature of the data itself as well as technological differences in the methods that may impact the ability to detect lowly expressed molecules with high confidence. For example, all MS-based proteomics methods lack a PCR amplification step that is present in most nucleotide sequencing–based methods, which enables higher sensitivity at lower sample inputs. Regarding the nature of the data, Ribo-Seq has the ability to identify translating ribosome signatures in an unbiased way, which may confidently find ORFs less than eight amino acids long that are fundamentally challenging to identify by MS (15, 57). In fact, Ribo-Seq can confidently identify an ORF that is simply a start codon followed by a stop codon (i.e., Met∗) because the Ribo-Seq reads remain sufficiently long for unique genomic mapping (58).

Second, since some noncanonical ORFs are located in GC-rich promoters (such as uORFs), these may encode amino acid sequences that are enriched in arginine (CGU/CGC/CGA/CGG codons) and thus would be excessively cleaved by trypsin to small peptides that cannot be uniquely mapped to a single ORF. Whether use of alternative proteases (59) could improve noncanonical ORF detection in whole lysate proteomics is unclear.

Considerations and Quality Control Steps for the Data-Driven Discovery of Noncanonical Human ORFs

Differences in the nature of Ribo-Seq and LC–MS/MS-based tryptic proteome and immunopeptidome data collection also represent a source of substantial variability in the detection of noncanonical ORFs. Notably, while targeted proteome and immunopeptidome LC–MS/MS approaches may offer improved sensitivity, these require candidate noncanonical proteins of interest to be known prior to analysis. While each method uses high-throughput data generation to profile cellular translation comprehensively, the data have intrinsically different strengths and weaknesses that may result in discordance between them (Table 1).

Table 1.

Features and characteristics of methods to detect noncanonical ORF translation

Data type Molecule detected Digestion step? Target size of analyte Number of CDSs detected Number of ORFs detected Strengths Weaknesses
Ribo-Seq RNA bound within ribosomes RNase and DNase 28–30 nt 10,000–13,000 2000–200,000 Genomewide Does not detect proteins directly
No bias because of trypsin Cannot detect PTMs
Detects small and large CDSs Cannot inform post-translational protein regulation
Nucleotide-level precision Analysis pipelines may be discordant
Defines exact reading frame of ORF
LC–MS/MS Tryptic peptides Trypsin 8–25 amino acids 9000–11,000 10s to 100s Direct protein detection High false-positive rate without Ribo-Seq
Informs protein abundance Biased against small proteins
May detect PTMs Trypsin may bias protein representation
Proteome-wide Does not provide nucleotide-level precision
HLA immunopeptidomics HLA-presented peptide antigens None 8–12 amino acids 8000–10,000 1000–5000 Direct protein detection Does not inform protein stability
Enrichment for low-abundance, strong binders Does not indicate intracellular abundance
Proteome-wide HLA allele expression limits peptide representation
Can detect unstable translations Does not provide nucleotide-level precision
Does not require tryptic sites

Ribo-Seq

The quality of a Ribo-Seq dataset is most commonly evaluated using three considerations: codon periodicity, library complexity, and number of canonical CDSs identified.

Codon periodicity reflects the percentage of Ribo-Seq reads that correctly identify the known reading frame of CDSs (Fig. 2, AC). In a high-quality Ribo-Seq dataset, ≥70% of reads that are between 28 and 30 nucleotides in length map to the correct reading frame of known CDSs. The precise read length that displays the most preferable (the “cleanest”) signal can vary and depends on the sample type and the method of nuclease digestion used to eliminate cellular RNAs not bound within the translating ribosome. Because of limitations of the experimental technique as well as biological variation in ribosome occupancy, a codon periodicity above 90% is typically not attainable (60). A Ribo-Seq dataset with a codon periodicity <60% should ideally not be used for ORF discovery because of challenges with accurate identification of the reading frame (19, 60, 61). A periodicity between 60 and 70% is a gray zone where the data may be used in some cases with increased caution and stringency.

Fig. 2.

Fig. 2

Quality metrics of Ribo-Seq and stringency of ORF calling.A, an illustration showing codon periodicity as a central metric of Ribo-Seq library generation. Three illustrations indicate high-quality, borderline, and poor-quality Ribo-Seq libraries. B, an illustration representing high-stringency and low-stringency ORF calling. In the top case, a small number of reads map the the 3′UTR of an annotated mRNA, and only two-thirds of those 3′UTR reads support the same reading frame of a potential dORF nomination. In the middle and bottom cases, a potential intORF has varying read support evidence. The middle case shows clear evidence of an intORF by a large increase in reads mapping to the +2 reading frame midway through the CDS. In the bottom case, there is a smaller change in the reads mapping to the +2 reading frame. C, use of ribosome-stalling drug treatments to clarify translational start sites. Cultured cells are treated with homoharringtonine or lactimidomycin to stall ribosomes at the main translational start site of a given ORF, leading to a clearer resolution of the specific start codon. CDS, coding sequence; dORF, downstream ORF; intORF, internal ORF.

Library complexity refers to the number of unique RNA molecules sequenced and what fraction of these are ribosome footprints that map to CDSs. The challenge with a low complexity library is that the majority of the reads will be PCR duplicates. When the number of initially isolated footprints is limited (e.g., because of low quality of the input material or suboptimal sample processing), ultimately many duplicate copies of this limited number of footprints will be sequenced. This means that deeper sequencing of this library will yield no or only minimally more biologically distinct footprints. Typically, the majority of reads in such low-complexity libraries will come from nonfootprint sources, particularly intergenic and intronic contaminants (e.g., microsatellite repeat elements, ribosomal RNAs, or small RNAs that overlap gene regions), which are unintentionally isolated during the Ribo-Seq procedure because these RNA species are of a similar size to the ribosomal footprint and may have certain RNA structures (62, 63). In general, a Ribo-Seq library with sufficient complexity will have the majority of reads mapping to annotated and novel CDSs. In some cases, such as with degraded samples, there may be substantial intergenic noise or a higher fraction of RNA species that are normally restricted to the cell nucleus but yet still sufficient codon periodicity and library complexity in terms of unique RNA molecules that map to CDSs. Here, the challenge is to achieve sufficient sequencing depth to ensure adequate sampling of unique RNA molecules. While 150 million reads typically suffices for the analysis of a high-quality Ribo-Seq library, a “noisy”—yet usable—library may require very deep coverage (>400 million reads), which is mostly a consideration for the financial cost of the sequencing (60, 64, 65). For human Ribo-Seq libraries, typically 15 to 30% of the sequenced reads can be classified as ribosome footprints, and the rest is often discarded. For a library sequenced to a depth of 150 million reads, that would total to approximately 22.5 to 45 million ribosome footprints—a number comparable to a routinely sequenced RNA-Seq library. Of these, >80% should map to annotated CDSs (60), leaving ∼5 million ribosome footprints for ORF discovery.

The number of known CDSs identified is particularly important when one aims to provide a comprehensive view of all translated ORFs in a sample of interest. This metric relates both to the amount of noise in the library, the periodicity of the footprints, as well as the depth of the sequencing. A sufficiently sequenced Ribo-Seq library for a human sample with high periodicity should detect at least >9000 annotated CDSs and often >10,000 annotated CDSs (9, 10, 11, 18). Human sample Ribo-Seq libraries that do not reach this threshold—despite sufficiently deep sequencing and periodicity—should be used with caution, as the false-negative rate for detecting ORFs will be high (many ORFs will be missed). While Ribo-Seq-based ORF detection tools theoretically have a low false-negative rate, the confidence (false discovery rate [FDR]) with which an ORF or CDS is detected, the number of independent samples in which it can be found, and the translation rate of the ORF should always inform research decision-making. For instance, direct comparison of noncanonical ORF FDRs and translation rates, compared with those of canonical CDSs, can inform both the relative abundance of the ORF’s translation product and the degree of certainty with which the algorithm could nominate it.

Because de novo and ab initio RNA assemblies are technically challenging with the short nucleotide sequences (28–30 nt) obtained during a Ribo-Seq experiment, analysis of Ribo-Seq data requires alignment of the reads to a reference transcriptome, most commonly Ensembl or RefSeq though custom transcriptomes are also used in some cases. Statistical assessment of a noncanonical ORF nomination is inconsistent across computational methods, with some approaches calculating a p value for significance (e.g., RiboTaper (61), ORFquant (10), Ribo-TISH (66), PRICE (67), and RiboCode (68)) and other approaches computing confidence scores (e.g., RibORF (19), Ribotricer (69), ORF-RATER (70)). In addition, these methods are often based on fundamentally different modeling approaches, including hidden Markov (RiboHMM (20)), multitaper (RiboTaper (61)), transformer (DeepRibo, TIS Transformer (71, 72)), support vector machine (RibORF (19)), expectation-maximization (PRICE (19, 67)) models, among others. As such, different methods may be more appropriate for certain research questions, datasets, or desired ORF types.

As a consequence, two different algorithms can have differing ORF outputs for the same gene. This can be due to the level of stringency or the strengths and weaknesses of a particular ORF caller for a certain type of ORF or certain quality of data. For example, some ORF callers cannot detect ORFs with near cognate start codons, whereas others are better suited for the detection of overlapping reading frames where periodic footprint signals are mixed and hard to dissect. Other tools handle alternative splicing better. Depending on the research question, input data quality, species of interest, or annotation goals, combinations of ORF callers followed by curation of called ORFs may be necessary (see later in “How many noncanonical ORFs are there?”).

HLA-I and HLA-II Immunopeptidomics

In the past decade, interest in HLA-I and HLA-II presented peptides has become widespread across many areas of biomedical research, as a subset of HLA-presented peptides demonstrate antigenic properties and represent a class of potential therapeutic targets (73, 74, 75, 76). The application of HLA immunopeptidomics differs from tryptic proteome protocols, as these methods leverage native lysis buffer and antibody or affinity-tag enrichment steps to isolate HLA–peptide complexes from cell lysates (Fig. 1B) (77, 78). The peptides are naturally produced following degradation of endogenously expressed source proteins by cellular proteases and peptidases and the proteasome. As such, no tryptic digestion is used in immunopeptidome analyses, which may enable some noncanonical proteins to be detected by immunopeptidomics even if they cannot generate tryptic peptides. Therefore, regarding detection of noncanonical proteins, HLA immunopeptidome analysis has three advantages over tryptic proteome analysis: (1) each HLA allele has a distinct peptide-binding motif that presents specific subsets of peptides, which can then be detected with MS in the absence of digestion with a protease; (2) the HLA presentation pathway may have privileged access to proteins that are rapidly degraded as the half-life of HLA–peptide complexes (hours) are in general longer than the half-life of rapidly degraded proteins (minutes) (78, 79); and (3) HLA immunopeptidomics broadly samples endogenous proteins from all abundance levels including those from lower-abundance noncanonical ORFs (80, 81, 82). These advantages align with recent studies that have shown higher observation rates of noncanonical proteins in the HLA-I immunopeptidome compared with the tryptic proteome (39, 83).

Similar to tryptic proteome datasets, immunopeptidome datasets require strict quality control steps to ensure the data and analysis are of high quality. Peptide length, the presence of peptide-binding motifs, and predicted binding to HLA molecules coded by specific alleles are common quality control steps in immunopeptidomics workflows. Because HLA-I and HLA-II molecules have unique peptide-binding grooves that accommodate peptides of different lengths, peptide size is an important quality control metric of immunopeptidomics data. Specifically, HLA-I peptides are ∼8 to 12 amino acids long (mostly 9mers), whereas HLA-II peptides are generally 12 to 25mers (77). HLA-II peptides are also typically found in nested sets, while this is not a global feature of HLA-I peptides, and can also be used to quality control HLA-II immunopeptidome datasets. Furthermore, each individual person expresses different HLA alleles with distinct HLA-binding motifs, which influence which peptides are presented. Therefore, it is common to confirm that HLA allele–specific binding motifs of the expressed HLA molecules are present in the immunopeptidome data, and that peptides derived from canonical and noncanonical ORFs in a given dataset are predicted to bind to the expressed HLA molecules to a similar extent. A number of computational approaches (e.g., MHCflurry, NetMHCpan, MixMHCpred, ForestMHC, HLAthena) can be used to both predict HLA peptides and the strength of their binding to various HLA molecules (76, 84, 85, 86, 87, 88, 89). It is important to note that HLA-I binding prediction is currently more accurate compared with HLA-II binding prediction, as HLA-II motifs are more complex and large subsets of diverse HLA-II heterodimers are in the process of being characterized and the associated prediction algorithms are being further improved (90, 91, 92, 93).

Interestingly, peptides derived from noncanonical ORFs are much more abundant in HLA-I datasets compared with HLA-II datasets (18, 34, 36, 38, 39, 53, 94). HLA-I molecules usually present peptides derived from proteasome-mediated degradation of newly synthesized and other cellular proteins, and HLA-I presentation is tightly linked with protein synthesis and degradation rates. In contrast, HLA-II molecules, which are often expressed on professional antigen-presenting cells, present peptides derived from degradation of extracellular proteins that were taken up by the antigen-presenting cells or from endogenous proteins that are destined to be degraded in specialized vacuolar compartments of the endosome–lysosome system. Both HLA-I and HLA-II systems require trafficking to ensure peptide loading in the right compartment. For HLA-I, the peptides themselves are transported into the endoplasmic reticulum by a transporter associated with antigen processing, whereas in case of HLA-II, the source proteins must first reach the acidic compartments for degradation, for example, via receptor-mediated internalization or recycling of transmembrane proteins. Hence, the sources of HLA-II–presented peptides are often stable and abundant proteins.

Because of HLA-I binding constraints, and the short length of some noncanonical proteins, a noncanonical ORF is often represented by a single peptide in HLA-I immunopeptidome data, and therefore, additional quality control measures should be taken to support these identifications. To this end, a noncanonical protein subset-specific FDR threshold should be applied to each individual ORF type, rather than a global FDR (83, 95) because noncanonical ORF peptides represent a small fraction (typically <5%) of the overall immunopeptidome and individual ORF types vary considerably in their frequency. Thus, a global FDR can be excessively permissive for a small subpopulation and lead to higher false-positive identifications.

Beyond leveraging known HLA-specific peptide lengths, binding motifs, and subset-specific FDR, there are further quality metrics that can be applied to immunopeptidomics datasets when the focus is the identification of rare noncanonical proteins (96). The gold standard for supporting the identification of noncanonical peptides presented by HLA molecules is by comparing the retention time and MS/MS spectrum of an identified peptide with a synthetic peptide of the same amino acid sequence. However, it is often the case that hundreds of noncanonical peptides are identified in a single HLA-I immunopeptidome experiment, making the synthetic peptide confirmation for all potential noncanonical-derived HLA-I peptides not feasible. To overcome this challenge, it is now possible to compare the observed MS/MS spectra with predicted MS/MS spectra with tools such as Prosit (97). The comparison of the predicted and observed MS/MS spectra provides additional support for noncanonical peptide identification (98, 99). In addition, there are also multiple algorithms that can predict peptide retention times. The predicted retention time, using tools such as DeepLC or DeepRescore, can be compared with measured retention time for all peptides in a sample (canonical and noncanonical), as the correlation between predicted and observed retention time supports the LC–MS/MS identifications of noncanonical-derived peptides in immunopeptidomes (100, 101). Overall, deep learning–based prediction of peptide MS/MS spectra and retention time are powerful tools that help reduce the number of false-positive noncanonical peptide identifications in immunopeptidome datasets.

Tryptic Proteome LC–MS/MS

Rigorous standards for the analysis of LC–MS/MS tryptic proteome data have been established by the Human Proteome Organization/Human Proteome Project (HUPO/HPP) international consortium, as reviewed elsewhere (102, 103, 104), and these standards remain the expectation for researchers claiming identification of noncanonical ORF peptides (30). For claims of detection of proteins not previously detected, these guidelines require two nonnested and uniquely mapping peptides each of at least nine residues in length with a total extent of at least 18 amino acids and with high-quality peptide-spectrum matches (PSMs) upon manual inspection (30, 102, 104). Peptides may be from different samples but ideally should be reported in the same article to ensure consistency of data analysis, which is consistent with prior HUPO/HPP recommendations (102, 104). These PSMs should be provided in the form of universal spectrum identifiers so that the spectra can be easily examined by others (105).

Yet, consistent application of high-quality tryptic proteome data collection and analysis guidelines remains nonuniform across the research community. Proteogenomic studies looking for noncanonical ORFs without Ribo-Seq data—that is, by predicting and including all ORFs in RNA transcripts—have been plagued by high false-positive rates (30, 49, 106, 107, 108, 109), and initial efforts to inspect early claims of noncanonical ORF peptides concluded that “many of the spectral matches appear suspect” (30).

Moreover, while use of decoys is standard in tryptic proteome experiments to define global FDRs, decoys may be less useful for distinguishing true peptides for noncanonical ORFs. Indeed, Wacholder et al. (110) have concluded that decoy bias among noncanonical ORF products leads to inaccurate FDR estimates for short ORFs when decoys are created by reversing the complete protein sequence but not when excluding the initial Met from the reversal. Finally, efforts to identify noncanonical ORFs in tryptic proteome data must account for peptides instead being derived from canonical variants including single amino acid variants and splice-site peptides for alternative isoforms of known CDSs. The use of personalized proteogenomic database searches is not straightforward or used by all in the proteomics community.

Considering these factors, the general experience of the research community is that few noncanonical ORFs are found by conventional tryptic proteome LC–MS/MS analyses, and some of those are ultimately false-positive peptides (111, 112). In some cases, such ORFs are “undiscoverable” by tryptic proteome approaches, either because of the short length of noncanonical ORFs or intrinsic sequence features that do not produce LC–MS/MS observable tryptic peptides. For example, translation of repetitive amino acid sequences (e.g., glycine–leucine) has recently been described (29). Nevertheless, even approaches aimed at enriching for small proteins from cell lysates result in only modest increases in noncanonical ORF detection, rather than exponential increases (33). On the other hand, other enrichment techniques focused on PTMs (i.e., the acetylome, phosphoproteome, and ubiquitylome) have also reported noncanonical proteins and may provide both an alternative method to enrich for noncanonical proteins and also hint toward potential functional relevance of this subset of noncanonical proteins given the cellular roles of those PTMs (83).

Furthermore, data-independent acquisition-MS (DIA-MS) provides a potential opportunity to detect noncanonical ORF-derived peptides that have been reliably detected previously with high-quality spectra obtained with narrow isolation windows from a data-dependent acquisition approach. In DIA-MS, previously identified peptides are more reproducibly sampled by sequentially isolating and fragmenting peptides across the m/z range, which decreases stochastic sampling bias toward higher abundant species and may increase the chances of finding rare noncanonical ORFs (113). This approach has been used in conjunction with Ribo-Seq to claim detection of microproteins from noncanonical ORFs (50). Caution should remain with DIA approaches as fragmentation spectra are predominantly a mixture of multiple coisolated peptide ions in broader mass windows, rather than discrete isolated narrow mass ion windows. This results in blended spectra, often containing multiple low-abundance peptide ions, which can confuse DIA algorithms and that make manual verification extremely challenging.

Beyond technical limitations of MS, there are also biological factors that may make noncanonical ORFs less frequently observed in tryptic proteome LC–MS/MS datasets. To this end, there is increasing evidence that points toward intrinsic instability of proteins translated from noncanonical ORFs, resulting in their immediate degradation. Kesner et al. (114) used functional genomics approaches to demonstrate that the ribosome-associated BAG6 membrane protein may directly triage hydrophobic noncanonical ORF translations to the proteasome for degradation. Thus, it is possible that many noncanonical ORFs do not generate a stable protein product and might only be observable by immunopeptidomics or in tryptic proteome experiments with inhibition of the protein degradation mechanisms of a cell.

How Many Noncanonical Human ORFs are There?

The number of noncanonical ORFs encoded in the human genome remains highly speculative. To date, a limited number of human tissues and cell lines have been analyzed by Ribo-Seq, and proteogenomics studies that have aimed to incorporate ORFs derived from these datasets have been difficult to interpret because of numerous false positives. As such, while it is well-established that the human genome contains thousands of translated noncanonical ORFs, whether the precise number is closer to 10,000 or 100,000 remains a matter of debate. A further complication is that different research communities may not use a consistent definition of what types of ORFs we define as “noncanonical.” Yet, while analyses of more cell lines and tissues will certainly uncover additional noncanonical ORFs, there can be variable noncanonical ORF identifications even within analyses of the same cell line. Such variability reflects the equal—perhaps foremost—contribution of different analytical methods for noncanonical ORFs in the estimation of their prevalence.

The Number of Noncanonical ORFs

Most Ribo-Seq studies focusing on noncanonical ORFs report detection of several thousand ORFs, typically between 2000 and 8000 (9, 11, 15, 16, 18, 19, 20, 21, 51, 61, 115). Interestingly, this range seems relatively stable when comparing studies that employ only a few cell lines and broader analyses looking across many different human tissue types. To consolidate these findings, we have recently participated in an international consortium to aggregate 7264 high-confidence noncanonical ORFs and provided formalized annotations for them within the GENCODE gene annotation database (16). This GENCODE set demonstrates substantial overlap in the identification of certain types of ORFs, such as uORFs, across diverse datasets such as pancreatic progenitors, heart and stem cells, suggesting that perhaps the diversity of several ORF types may not be dramatically larger with the inclusion of more tissue types. In support of this, Ribo-Seq profiling of five human tissue types and six primary human cell types similarly reported 7767 ORFs in total (15). When subsetting this dataset for consistency with the inclusion criteria for the GENCODE catalog (i.e., removing ORFs below 16 amino acids in size, as well as ORFs without an AUG start codon), 2475 of 7767 ORFs remained, of which 1702 (±70%) were represented in the GENCODE catalog as well (supplemental Tables S1–S4).

While these studies have measured and determined noncanonical ORF translation directly from Ribo-Seq data, there are many other databases that have aggregated larger numbers of ORFs from a variety of sources, including both Ribo-Seq and in silico predictions. Among these, smProt (n = 327,995 human ORFs (116)), sORFs.org (n = 4,377,422 ORFs across humans, mouse, and fruit flies (117)), RPFdb (118, 119), and smORFunction (n = 617,462 human ORFs (120)) have compiled reported or putative noncanonical ORFs. Notably, OpenProt (121, 122) has two aspects to their database workflow: one that collates all predicted ORFs (n = 488,956) and a second that proposes 33,836 translated ORFs identified by a reanalysis of over a hundred Ribo-Seq datasets with the PRICE pipeline (67). When considering studies that have generated Ribo-Seq datasets to measure noncanonical ORF translation, there are also several efforts that have proffered exceptionally large numbers of directly detected ORFs—specifically, the nuORFdb (34) by Ouspenskaia et al. and the Human Brain Translatome Database (123) by Duffy et al., which propose numbers of >230,000 and >75,000 ORFs, respectively.

Why is There Such Discordance in the Number of Noncanonical ORFs Across Databases?

The interpretation of such dramatically different accounts of noncanonical ORF abundance remains a challenge. Indeed, given that there are currently only ∼60,000 Ensembl genes (including 19,827 protein-coding genes, 18,886 lncRNAs, 4864 small ncRNAs, 15,241 pseudogenes, and 2221 other RNAs in Ensembl, version 109.38), colossal datasets with >200,000 ORFs may be interpreted to suggest that every gene has upward of four distinct ORFs. In practice, these large datasets may include isoform variants (e.g., N-terminal extensions, C-terminal extensions, and intron retentions) that are not part of the reference proteome, and thus the number of noncanonical ORFs may be larger in some databases because of differences in how these isoforms are categorized.

While sample and data quality likely contribute to the variability in the numbers of noncanonical ORFs in some catalogs, differences in Ribo-Seq data analysis also account for much variation in prospective noncanonical ORFs. For example, biologically, there is some amount of stochastic or pervasive translation across all RNAs, which may relate to leaky ribosomal scanning (124, 125, 126) or transient interactions between ribosomes and RNAs as the ribosomes locate CDSs or RNAs accomplish proper folding (127, 128). Yet, the manner in which computational pipelines process Ribo-Seq data results in ORF calls that may be more or less stringent (Fig. 2B), resulting in different proportions of false-positive (stochastic) and false-negative (e.g., sample-specific) ORF calls (60, 129, 130). For example, RibORF (19), which uses a support vector machine and recommends a fixed cutoff score of 0.7, has been shown to produce the highest numbers of ORF calls of any tested algorithm in a recent benchmarking study (131). To confirm these differences directly, we have reanalyzed published high-quality Ribo-Seq data for six biological replicates of pancreatic progenitor cells differentiated from human embryonic stem cells (11) using four common ORF detection pipelines (ORFquant (10), PRICE (67), Ribo-TISH (66), and Ribotricer (69)), observing substantial variability in the number of ORFs called (∼10-fold difference from ∼50,000 to ∼500,000), the types of ORFs called, the length of the called ORFs, and the reproducibility with which ORFs could be detected across all six replicates (Fig. 3 and Experimental procedures section).

Fig. 3.

Fig. 3

ORF callers have different specialties and variable performance.A, stacked bar plot displaying all detected ORF categories per ORF caller. For each, the percentage of unique ORFs shared between at least one, three, or six replicates is shown. Please note that these are relative contributions to the total number of ORFs. The absolute numbers of ORF identifications can be inferred from C. B, density plots displaying the distribution of ORF lengths in nucleotides (excluding the stop codon) for unique ORFs shared between at least one, three, or six replicates. C, line graphs showing the numbers of unique ORFs detected by each tool shared between at least one, three, or six replicates. The x-axis denotes the percentage of overlap used to consider two ORFs being similar or not, with 100% overlap meaning that the detected ORF was fully identical between [x] number of replicates. Please note that the total numbers of ORFs detected per algorithm (y-axis) can differ by an order of magnitude. These numbers are given for each line, with numbers reflecting the total ORFs with 100% similarity between replicates (i.e., the end of each curve). D, genomic view of a short upstream ORF (uORF) in the STPBN1 gene indicating that ORF callers have variable affinity for certain types of ORFs. The top two tracks show the ribosomal P-site positions derived from the sequenced ribosome footprints, as processed independently from the sequencing data by the deterministic ORF caller ORFquant (top; red shading) and the probabilistic ORF caller PRICE (bottom; blue shading). The differently colored P-site bars indicate different reading frames (0, +1, and +2) on the same transcript, with bars in the same color indicating a shared in-frame codon movement by the ribosome. For this visualization, newly found ORF variations of the annotated CDS that could be assigned to predicted noncoding RNA isoforms (e.g., transcript biotype: “processed_transcript”), but matched CDS of SPTBN1 is not displayed. E, genomic view of a near-cognate start codon ORF in TUG1. Image and track details as in (E) above. CDS, coding sequence.

There may be specific reasons for the different performance characteristics of each algorithm. For example, the lower stringency of RibORF may be due to the fact that this pipeline considers uniformity of read coverage across the ORF, whereas Ribo-Seq is known to have a 5′ bias to read coverage. Therefore, RibORF may excessively promote intORFs and doORFs since the 5′ ends of these ORFs overlap annotated CDSs, which typically have higher read coverage independent of a periodic footprint signal that matches the correct reading frame. This is evident in nuORFdb (34) and the Human Brain Translatome Database (123): when analyzing the fraction of ORFs with an AUG-start resulting in an ORF ≥16 amino acids, doORFs and intORFs are 173-fold and 18-fold (respectively) higher in abundance compared with other major datasets (Fig. 4, supplemental Tables S5–S8). By contrast, uORFs are only three times more abundant (Fig. 4).

Fig. 4.

Fig. 4

An analysis of major noncanonical ORF databases.A, here, each dot reflects a dataset, and the Y-axis uses a log-10 scale to show the number of ORFs included that are ≥16 amino acids long and contain an AUG start codon. The GENCODE catalog reflects the summation of the studies by Ji et al. (19), Calviello et al. (61), Raj et al. (20), van Heesch et al. (9), Martinez et al. (21), Chen et al. (18) and Gaertner et al. (11) datasets as described (16). B, the number of ORFs per dataset compared with the number of samples profiled by Ribo-Seq. C, the number of ORFs per dataset compared with the number of unique cell types profiled by Ribo-Seq. D, the ratio of the number of ORFs per cell type compared with the number of ORFs per number of samples for each dataset. E, a bubble plot integrating the number of samples, number of different cell or tissue types, and the number of noncanonical ORFs found in each dataset.

It is also true that different computational pipelines may have different capacity to identify certain classes of noncanonical ORFs. For example, the deterministic multitaper-based statistical inference of significant periodic signal within predicted ORFs as performed by RiboTaper (61) and ORFquant (10) provides high-confidence detection of ORFs with an AUG start codon, but have not, to date, been optimized for non-AUG ORFs. In contrast, the probabilistic algorithm employed by PRICE (67) has enhanced ability to identify very short ORFs and non-AUG ORFs absent from other ORF callers (Fig. 3, B and E). Yet, when there are neighboring putative initiation codons (e.g., CUG and AUG), PRICE will generate larger numbers of putative ORFs that might require manual curation or further filtering. In addition, since annotated CDSs have generally more abundant Ribo-Seq read coverage, low-abundance out-of-frame reads may be more readily interpreted as an intORF with a non-AUG start codon by PRICE, whereas other ORF callers are less likely to consider these reads as sufficient evidence for a translated ORF. Thus, when applied to biological replicates of the same sample, PRICE produces the least consistent ORF calls compared with other pipelines, independent of initiation codon variability (Fig. 3, AC) (131). nuORFdb (34) and OpenProt (122) both employ PRICE in their analysis pipelines. It is important to note, however, that the specific research question being pursued should inform the types of ORF callers used: indeed, deterministic algorithms such as RiboTaper or ORFquant may miss intORFs or overlapping ORFs identified by PRICE because of the difficulty in resolving mixed periodicity signals of overlapping reading frames (Fig. 3A).

In summary, depending on the type of ORF one aims to find and the desired inclusiveness of ORFs one aims to output, one ORF caller might be better suited than another. Certain ORF callers outperform others in detecting specific ORF categories such as intORFs (Fig. 3A), very small ORFs (Fig. 3, B and D), or near cognate start codons (Fig. 3E), whereas others handle exon–exon junctions and longer ORFs better and/or provide better replicate behavior. These differences then lend to substantially different results when producing noncanonical ORF catalogs (Fig. 4).

Detection of Translational Start Sites

Determining the translational start site of an ORF remains a nuanced problem. While conventionally proteins have been annotated with AUG start sites, exceptions to this rule have long been known (132, 133), and noncanonical ORFs are more likely to employ non-AUG start sites (125, 134). In a typical Ribo-Seq experiment, identification of translational start sites from Ribo-Seq data is inferred based on two factors: sequencing coverage and the intrinsic restrictions of the computational pipeline (e.g., some algorithms only consider AUG start codons, as discussed previously). Yet, independent of the computational pipeline, there may be gaps in the sequencing coverage that lead to misidentification of the main translational initiation site (Fig. 2C). For experiments with cultured cells, use of small molecules that block ribosome elongation, such as homoharringtonine (135) or lactimidomycin (136), enables ribosome accumulation on translational initiation sites, which enables more precise determination of the start codon. Because of the difficulty in identifying noncanonical ORF start sites and the variability in computational approaches to start codon recognition (e.g., Fig. 3E), use of homoharringtonine or lactimidomycin with cultured cells is highly recommended. In frozen tissue samples, these compounds are no longer effective.

How to Select an ORF Sequence Database for MS Data Analysis?

Given the wide differences between the different databases for Ribo-Seq ORFs, one central question is how to use these databases, or which to use for any specific analysis? Because the size of the ORF output in a given database can vary enormously, users should base their decision on what scientific question they intend to pursue and evaluate carefully the suitability of the input Ribo-Seq data quality as well as the stringency with which ORF calling was performed. In general, high stringency databases provide high-confidence Ribo-Seq ORF detections, and thus peptides found mapping to these ORFs are more likely to reflect a true positive result. While these databases reduce false positives, it is at the expense of comprehensiveness, as the existing high stringency databases will yield more false negatives in the MS analysis. Low stringency databases provide a much larger set of Ribo-Seq ORFs but will yield more false positives—because of the lack of support from another orthogonal technique. If the ORFs are accompanied by Ribo-Seq quality metrics, it may be tractable to estimate the proportion of false positives and refilter the ORFs to suit one’s own purposes. These databases will provide a larger candidate search space for peptide alignment and may enable detection of true positive ORFs not present in the high stringency databases. Yet as described earlier, because of the concern for false-positive nominations, ORFs detected by MS searches should be closely inspected to verify integrity of both ORF call and peptide identification, as there will likely be cases of false-positive ORFs being supported by false-positive peptides. Ultimately, certain scientific questions may lend themselves to certain databases: for example, analyses of alternative N-terminal CDS extensions often emphasize non-AUG start sites (24), which may benefit from a Ribo-Seq analysis that employs the PRICE algorithm. Research efforts aimed to identify a maximal space of potential translation events may also favor a lower stringency database, with the caveat that any individual result should receive additional scrutiny. Alternatively, if the goal is to characterize a high-confidence unannotated microprotein, a high stringency database may be more desirable. Likewise, for reference annotation purposes and functional studies, we prefer more stringent workflows that yield reproducible ORF calls across samples (no false positives).

Are Noncanonical ORFs Proteins?

The term “protein” is conventionally used to refer to an amino acid sequence that produces a molecular structure that plays an intrinsic cellular role in maintaining normal cell biology. While some proteins may be unstable and rapidly degraded under certain conditions (e.g., beta-catenin), most proteins participate in cell biology when present in a stable form. Also, almost all annotated proteins show evidence of evolutionary conservation, structural folding, and domain architecture, and frequently also protein–protein interactions and/or interactions with nucleic acids.

According to this understanding of the term “protein,” it could be inferred that the vast majority of noncanonical ORFs do not encode proteins on the basis that they lack these characteristics. To our knowledge, microproteins from noncanonical ORFs also do not have paralogs within the proteome that might enable inferred protein functions. However, we see two additional considerations. First, it may be incorrect to assume that a protein that exists in the cell—even one that is detectable by MS—is therefore a functional molecule. It could be that the proteome contains a certain amount of nonfunctional translational “noise.” While it is difficult to prove the extent to which such translation occurs in normal cells, evidence from cancer cells shows abundant dysregulation of translation, exemplified by “aberrant” noncanonical proteins that lack evidence for function under normal physiological conditions (34, 35) as well as out-of-frame peptide byproducts of oncogene activity (137).

Second, the classical definition of protein “function” invokes the protein’s role in cellular processes that have been derived over time through evolution, which has been summarized as the maxim that “conservation = function.” This maxim has been central—but not universally required—for gene annotation projects, and the only canonical proteins currently within GENCODE that can be inferred to have evolved de novo in human or higher primates were initially detected in cancer cells (e.g., MYEOV (138) and HMHB1 (139)). Even so, evidence for the existence and function of de novo proteins under normal physiological conditions is accumulating (57, 140, 141, 142). Nonetheless, it remains true that most noncanonical ORFs display much higher rates of intrinsic disorder, fewer structural features, and lack amino acid constraint across evolution (17, 18, 140, 141, 143, 144, 145, 146, 147, 148, 149). While these features may be observed in diverse annotated proteins (e.g., intrinsically disordered regions of a given protein), their presence is predominant in noncanonical ORFs.

The absence of protein function as a criteria should not determine whether noncanonical ORFs are categorized as translational “noise.” Indeed, the function of many human proteins remains obscure, motivating multi-institutional efforts such as the Understudied Proteins Initiative (150) and the HPP Grand Challenge to define “a function or functions for every human protein” (151). In the case of noncanonical ORFs, because many may only exist as unstable peptides that are presented on the immunopeptidome, the question of whether potential recognition by T cells constitutes a molecular “function” becomes a central and partly philosophical debate for the research community. There is no current precedent to regard major histocompatibility complex presentation as a central “function” of a protein—as opposed to an ancillary observation for a protein that has additional roles in cell biology—and therefore, in the absence of additional experimental data on this question, we are disinclined to consider major histocompatibility complex presentation as proof that a noncanonical ORF has an intrinsic cellular role at this time.

The Interpretation of Peptide-Level Evidence of Ribo-Seq ORFs

How, then, should one interpret the peptide-level evidence for some noncanonical ORFs? High-quality tryptic proteome LC–MS/MS PSMs that survive rigorous manual inspection are strong evidence of true translation of a noncanonical ORF. With adequate evidence, therefore, tryptic proteome PSMs supporting noncanonical ORFs do indicate the possible existence of a translated protein, and these cases may reasonably be considered to be part of the cell proteome, similar to any other proteins.

When considering the larger number of noncanonical ORFs with peptide-level evidence in HLA immunopeptidomics but not tryptic proteome LC–MS/MS (18, 34, 36, 38, 152), firm conclusions are more difficult to draw. These noncanonical ORFs cannot be said to generate a true protein based on immunopeptidomics alone, considering that the HLA system is expected to present peptides resulting from translation products that are unstable and rapidly degraded, alongside those derived from canonical proteins. Yet, detection of an HLA-presented peptide does verify RNA translation in these cases, which distinguishes them from the majority of Ribo-Seq-detected noncanonical ORFs that are detected in neither tryptic proteome LC–MS/MS nor immunopeptidomics experiments. Therefore, these noncanonical ORFs can at least be said to be confirmed as both translated and presented by the HLA, as opposed to an artifact of the Ribo-Seq protocol.

A related question is how to interpret PSMs matching noncanonical ORFs that are not detected by Ribo-Seq, when the same sample is interrogated using both technologies. Because the sensitivity of Ribo-Seq is generally higher than MS-based methods, and because Ribo-Seq provides nucleotide-level precision for genome mapping, there are three possibilities here: first, these peptides may be false-positive identifications, second, the Ribo-Seq data exhibit a false-negative identification, or third, they may be derived from another source not included in the search space (e.g., aberrant splicing). None of these hypotheses has been rigorously evaluated at this time. One challenge is that many proteomics and immunopeptidomics experiments do not currently generate matched Ribo-Seq data for their samples, and thus it cannot be directly known if Ribo-Seq supports translation of that ORF. When considering unmatched analyses, it is also noted that, at present, proteomics and immunopeptidomics datasets cover a broader range of tissue and cell types than Ribo-Seq datasets.

A Proposed Framework to Classify the Translation of Noncanonical ORFs

Given the expanding volume of research on noncanonical ORFs, a shared vocabulary for the interpretation of their detection is a critical need in the genomics, translatomics, proteomics, and immunopeptidomics communities. Notably, there has been no formalized initiative to annotate noncanonical ORFs as protein-coding genes by major genome databases, although recent collaborative work has raised this point as a topic of interest (16). Historically, protein-coding genes have been annotated one by one in a manual process of careful data inspection, which may or may not have included protein-level evidence. At this time, noncanonical ORFs detected by tryptic proteome data would potentially be eligible for manual annotation as protein-coding genes. Yet, given the paucity of noncanonical ORFs in tryptic proteome data and their much greater abundance in HLA immunopeptidomic datasets, there is uncertainty about whether most noncanonical ORFs produce proteins in the classical sense, and whether immunopeptidomic evidence is equivalent to tryptic proteome data for the purposes of protein annotation.

We advocate both a cautious but open-minded approach to noncanonical ORF classification, summarized in Table 2. Notably, although most annotated proteins show evidence of amino acid constraint across species and most noncanonical ORFs do not, it is also unquestionably true that at least some proteins are lineage- or species-specific. Thus, we propose that de novo translations should be considered for annotation as protein coding. While recognizing that evolutionary analysis is a core part of gene annotation workflows in projects like GENCODE, we have not included conservation or constraint metrics as part of this proposed framework. The framework itself is oriented toward harmonizing subsequent dataset generation and analysis. In practice, it might be applied to classifying published datasets, and it is intended as a helpful tool for candidate prioritization rather than a guarantee that certain ORFs will be annotated by a genome database. We stress that researchers looking to move forward with potential annotation of a protein encoded by a noncanonical ORF should be able to provide the raw LC–MS/MS spectra for review.

Table 2.

A proposed framework to standardize levels of evidence of noncanonical ORFs

Tier Required supporting evidence Standardized outcome
Tier 1A Tryptic proteome LC–MS/MS (≥2 peptides according to HUPO/HPP criteria) “Protein candidate.” Consider discussing research findings with genome annotation databases for possible annotation.
Ribo-Seqa
Tier 1B HLA immunopeptidomics MS (≥2 observations; multiple high-confidence peptides from multiple distinct sources) “Presented”
Ribo-Seqa
Tier 2A Tryptic proteome LC–MS/MS (≥2 peptides not satisfying HUPO/HPP spacing criteria) “Detected”
Tryptic proteome LC–MS/MS (1 peptide)
Ribo-Seqa
Tier 2B HLA immunopeptidomics MS (1 observation) “Detected”
Ribo-Seqa
Tier 3 Any HLA immunopeptidomics or tryptic proteome LC–MS/MS evidence without Ribo-Seqa evidence “Putative,” consider alternative sources
Tier 4 Ribo-Seqa evidence without any proteomic evidence “Ribo-Seq ORF”
Tier 5 In silico prediction of an ORF on an expressed transcript without any Ribo-Seqa or proteomic evidence “Predicted”
a

From credible Ribo-Seq data with quality metrics meeting the guidelines suggested in this article. Ribo-Seq need not be performed on aliquots of the same samples analyzed by proteomics.

Our framework centers proposes these definitions for specific terminology:

  • Protein candidate”: a tier 1A noncanonical ORF can be regarded as translated into a protein candidate if it satisfies current HUPO/HPP guidelines for the detection of ≥2 uniquely mapping tryptic proteome peptides, as well as having evidence of translation by Ribo-Seq. Such candidates would be prioritized for further manual review by annotation groups.

  • Presented”: A presented noncanonical ORF (tier 1B) is one with multiple lines of evidence for its translation and presentation on HLA molecules. These ORFs are detected with multiple high-confidence peptides from multiple distinct samples for HLA immunopeptidomics data as well as having evidence of translation by Ribo-Seq.

  • Detected”: A detected noncanonical ORF is one with evidence of translation by Ribo-Seq as well as evidence of protein production by either (tier 2A) tryptic proteome LC–MS/MS (1 peptide or >1 peptide not satisfying HUPO/HPP guidelines for their spacing) or evidence of protein production by HLA immunopeptidomics with a single PSM (tier 2B).

  • Putative”: A putative noncanonical ORF (tier 3) is one with evidence of translation with tryptic proteome LC–MS/MS or HLA immunopeptidomics data but no evidence of translation in Ribo-Seq data. This discrepancy may alert to the possibility of false-positive MS identifications or false-negative absence in Ribo-Seq and therefore requires more investigation.

  • Ribo-Seq ORF”: A noncanonical ORF that is only detected in Ribo-Seq data but not elsewhere is considered a “Ribo-Seq ORF” (tier 4). These are likely to be the majority of cases. The number of these ORF nominations may be variable based on the stringency of the Ribo-Seq analysis and/or the quality of the input data.

  • Predicted”: A predicted noncanonical ORF (tier 5) is one that is computationally predicted in silico on an expressed RNA transcript but without current evidence in Ribo-Seq or MS datasets.

Experimental Procedures

Benchmarking and Comparing ORF Caller Performance on Replicate Ribo-Seq Datasets

Ribo-Seq Data Processing and Mapping

Ribosome profiling data of late pancreatic progenitor cells obtained from six independent differentiations of H1 human embryonic stem cells (11) were collected from the Gene Expression Omnibus database (GSE144682). For all analyses, the Ensembl primary DNA assembly (GRCh38) and the Ensembl human reference transcriptome (Ensembl v102) were used as reference. Quality control and trimming of the Ribo-Seq reads was done using Trim Galore 0.6.6 with the options “--length 25” and “--trim-n” (153). Next, contaminant RNA and DNA were removed using Bowtie2 2.4.2 by aligning reads to a contaminant file using the default options of Bowtie2 (154). The contaminant-depleted reads were aligned using STAR with the options “--twopassMode Basic,” “--outFilterMismatchNmax 2,” “--outFilterMultimapNmax 20,” “--limitOutSJcollapsed 10,000,000,” “--alignSJoverhangMin 1000,” and “--outSAMattributes All” (155). For PRICE, the option “--alignEndsType EndToEnd” was set as well. Also, the individual bamfiles were filtered using SAMtools 1.12 to exclude reads with a mapping quality lower than 5 (156).

ORF Calling With ORFquant

The function RiboseQC_analysis from RiboseQC 1.1 was run in R 4.1.2 with the options “read_subset” and “fast_mode” set to false (157). The output was used by the function run_ORFquant from ORFquant 1.02 in R with the default options (10). ORF calling with PRICE: Before using PRICE, a reference genome was created with the IndexGenome function of the Gedi framework 1.0.2. After the creation of the reference genome, PRICE 1.0.3b was run (67). A filtered list of ORFs detected by PRICE and a list of P-sites (called activity values by PRICE) were extracted from the outputted “orfs.cit” files using the Gedi Nashorn and ViewCIT functions, respectively. Because the start codon prediction is a separate step in the PRICE program, ORF coordinates from both before and after start codon prediction were available. We used the coordinates after start codon prediction. PRICE can also be run in a multisample mode by providing a text file with the bam file locations as input. This mode favors ORFs that occur in all samples during the ORF calling process and would likely enhance the reproducibility of ORF calls between replicates. To keep all ORF callers comparable, we did not use this mode. ORF calling with Ribo-TISH: From Ribo-TISH 0.2.7, the predict function was used to infer ORFs with the option “--longest” set (66). The output file contained only the genomic start and end coordinates and the transcript id of each ORF. The reference GTF was used to determine the exons within each ORF. ORF calling with Ribotricer: The Ribotricer 1.3.3 function prepare_orfs was first used with the options “--longest” and “--min_orf_length 9” (69). The option “--start_codons” was set to include all near cognate start codons with one base difference compared with ATG. Afterward, the function detect_orfs was used with the option “--phase_score_cutoff 0.440.”

Comparing ORF Callers

ORF calls were compared between algorithms for the types of ORF categories that were found, in how many replicates they were independently discovered, how ORF differed in length, and how reproducible and similar their detection was based on, for example, the percentage of ORF sequence overlap between replicate ORF calls. Before the analyses, data were converted to GRangesList objects in R with stop codons included in the coordinates. ORF categories were determined by comparing the start and end coordinates, and the transcript id of each ORF with the CDSs in the “gtf.rannot” object created by the ORFquant function “prepare_annotation_files.” ORFs were compared by their overlap, with different thresholds set for the required percentage of overlap. Two ORFs were considered to be similar if the exons of one ORF were fully contained within the exons of a second ORF, both codons had the same stop codon, and the first ORF covered at least the required percentage of overlap of the length of the second ORF. These overlap relations were recursive, such that a parent ORF could be the child of another ORF, and all three would be counted as one unique ORF.

Comparison of Published Ribo-Seq Datasets

We used publicly available datasets from GENCODE (16), Chothani et al. (15), Ouspenskaia et al. (34), and Duffy et al. (123) for comparisons of published reports of noncanonical ORFs that might encode microproteins. The GENCODE dataset itself is a metaanalysis of data from Ji et al. (19), Calviello et al. (61), Raj et al. (20), van Heesch et al. (9), Martinez et al. (21), Chen et al. (18), and Gaertner et al. (11); datasets employed are listed in supplemental Table S1. Source data for these datasets are listed in supplemental Table S2. To facilitate comparisons between studies, we extracted only noncanonical ORFs with a length of ≥16 amino acids and had an AUG start codon. For ORFs using a non-AUG start site, the first internal AUG start codon was identified and the amino acid sequence starting with that internal AUG was included for analysis if the resulting ORF was ≥16 amino acids long. ORFs were then analyzed for their replication across primary datasets. Since the GENCODE list represents a meta-analysis of other individual datasets, the presence of an ORF in the GENCODE list was not used as part of the analysis for ORF replication across primary datasets. Next, ORF calls were associated with one of the following six categories: lncRNA-ORF, uORF, uoORF, internal ORF, doORF, or dORF, according to the schema by Mudge et al. (16). Duffy et al. used the nomenclature “external” for doORF, and these ORFs were reclassified as doORF for this analysis; they used “internal” for intORFs, which were reclassified as intORFs for this analysis. For lncRNAs, Duffy et al. used the term “noncoding,” which included the biotypes “noncoding,” “lncRNA,” “antisense_RNA,” “misc_RNA,” “TEC,” and “processed_transcript,” which were included as part of the lncRNA-ORF designation for this study. For Ouspenskaia et al., we analyzed ORFs according to the authors’ designation of ORF “plotType,” reflecting their final classification. Ouspenskaia et al. used the term “3′ dORF” for dORF, “3′ overlap dORF” for doORF, “5′ overlap uORF” of uoORF, “5′ uORF” for uORF, “lncRNA” for lncRNA-ORF, and “out-of-frame” for intORF. Chothani et al. reported final ORF types of “dORF,” “doORF,” “ncORF,” “overlap_uORF,” “intORF,” and “uORF.” For Chothani et al., Duffy et al., and Ouspenskaia et al., ORFs that had a final classification of pseudogene were excluded from this analysis; however, these datasets variably reclassified some ORFs on pseudogene transcript biotypes as noncoding or lncRNA, and we did not refilter these ORFs beyond the original reclassifications provided by the authors. ORFs that switch a classification corresponding to a small RNA, tRNA, or rRNA species, such as “rRNA,” “snoRNA,” “tRNA,” “snRNA,” or “miRNA,” were excluded from this analysis. The number of cell types and/or tissue types for analyses of each ORF dataset was extracted from the source publication.

Limitations

With this work, we have endeavored to clarify how Ribo-Seq can be used for noncanonical ORF research. Yet, our focus has several important limitations. First, the vast majority of—but not all—translated peptides can be traced back to an RNA sequence. There may be peptides that derive from amino acid splicing within the proteasome during protein degradation (158), which would not be detectable in Ribo-Seq data. Second, there are also well-established protein CDSs that are difficult to resolve with Ribo-Seq and do not have optimized computational methods for their quantification. For example, translated pseudogenes, retroviruses, retrotransposons, and paralogous protein-coding genes may have high sequence homology that precludes unique mapping of the short ∼30 bp reads from a Ribo-Seq experiment, although multimapping reads will provide evidence of translation. These cases are not discussed here. This issue of short Ribo-Seq sequencing reads also highlights the potential role for emerging long-read sequencing technologies to enhance detection of noncanonical ORFs on alternative transcript forms (159), which we do not discuss. Finally, each individual’s genome (and particularly each cancer’s genome) has a unique range of germline or somatic single nucleotide variants that will impact the proteome: in this article, we have not addressed the importance of generating personalized reference genomes and proteomes for the analysis of microproteins and noncanonical ORFs.

Conclusions

The widespread description of noncanonical ORFs has sparked a paradigm shift in the perception of both the human genome and the proteome. Yet, as a field still in its infancy, this area of investigation is plagued by a lack of standardization, which may lead to imprecise analyses, ultimately leading to self-injurious confusion. While the proportion of noncanonical ORFs that encode a functional protein remains to be seen, a large fraction of them can be verified as translated by both MS-based and Ribo-Seq-based approaches. A central effort for the research community is now to build reputable databases and analysis pipelines to ensure rigor in this quickly expanding—and highly exciting—field while also enabling functional studies to proceed with confidence. Here, we have considered the technologies used to detect noncanonical ORFs and attempted to provide a framework for categorizing differing levels of evidence for them. Our work aims to coalesce the research community around a common terminology and shared set of database resources for noncanonical ORFs. Ultimately, we believe that the study of noncanonical ORFs, if pursued with proper precision, will prove invaluable to the global community of biomedical researchers.

Inclusion and Diversity

We support inclusive, diverse, and equitable research.

Code Availability

All codes used for these analyses as well as data visualization are available at https://bitbucket.org/vanHeeschLab/orfcaller_comparison.

Supplemental data

This article contains supplemental data.

Conflict of interest

The authors declare no competing interests.

Acknowledgments

Funding and additional information

J. R. P. acknowledges funding from the National Institutes of Health (NIH)/National Cancer Institute (grant no.: K08-CA263552-01A1), the Alex’s Lemonade Stand Foundation Young Investigator Award (grant no.: 21-23983), the St Baldrick’s Foundation Scholar Award (grant no.: 931638), the Musella Foundation for Brain Tumor Research, the DIPG/DMG Research Funding Alliance, and a Collaborative Pediatric Cancer Research Awards Program/Kids Join the Fight award (grant no.: 22FN23). E. W. D. and R. L. M acknowledge funding from the NIH grant R01 GM087221. E. W. D. acknowledges funding from the National Science Foundation grant DBI-1933311. R. L. M. acknowledges funding from the National Institutes of Health grant U19AG023122. S. v. H. acknowledges funding from Fonds Cancers, Stichting Reggeborgh, Stichting Bergh in het Zadel, and Stichting Villa Joep. J. G. A. and K. R. C. were supported in part by grants P01CA206978 from the NIH, U24CA270823, U01CA271402, and U24CA271075 from National Cancer Institute Clinical Proteomic Tumor Analysis Consortium program and from the Dr Miriam and Sheldon G. Adelson Medical Research Foundation. J. M. M. is supported by the Wellcome Trust (grant number 108749/Z/15/Z), the National Human Genome Research Institute of the US NIH under award number 2U41HG007234, and the European Molecular Biology Laboratory. The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH. Ensembl is a registered trademark of EMBL.

Author contributions

E. W. D., J. R. P., S. v. H., J. R.-O., J. M. M., M. B.-S., J. G. A., K. R. C., and R. L. M. conceptualization; E. W. D., J. R. P., S. v. H., J. R.-O., J. M. M., M. B.-S., J. G. A., K. R. C., L. W. K., and R. L. M. methodology; J. R. P. and L. W. K. formal analysis; J. R. P., J. R.-O., and L. W. K. data curation; E. W. D., J. R. P., S. v. H., J. R.-O., J. M. M., M. B.-S., J. G. A., K. R. C., and R. L. M. writing–original draft; E. W. D., J. R. P., S. v. H., J. R.-O., J. M. M., M. B.-S., J. G. A., K. R. C., L. W. K., and R. L. M writing–review & editing; J. R. P. and L. W. K. visualization; E. W. D., J. R. P., S. v. H., J. R.-O., J. M. M., M. B.-S., J. G. A., K. R. C., and R. L. M. supervision; E. W. D., J. R. P., and S. v. H. funding acquisition.

Supplementary Data

Supplemental Tables S1–S3
mmc1.xlsx (36.6MB, xlsx)
Supplemental Data
mmc2.docx (16.9KB, docx)

References

  • 1.Aebersold R., Agar J.N., Amster I.J., Baker M.S., Bertozzi C.R., Boja E.S., et al. How many human proteoforms are there? Nat. Chem. Biol. 2018;14:206–214. doi: 10.1038/nchembio.2576. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Tress M.L., Abascal F., Valencia A. Alternative splicing may not be the key to proteome complexity. Trends Biochem. Sci. 2017;42:98–110. doi: 10.1016/j.tibs.2016.08.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Blencowe B.J. The relationship between alternative splicing and proteomic complexity. Trends Biochem. Sci. 2017;42:407–408. doi: 10.1016/j.tibs.2017.04.001. [DOI] [PubMed] [Google Scholar]
  • 4.Sinitcyn P., Richards A.L., Weatheritt R.J., Brademan D.R., Marx H., Shishkova E., et al. Global detection of human variants and isoforms by deep proteome sequencing. Nat. Biotechnol. 2023 doi: 10.1038/s41587-023-01714-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Frankish A., Carbonell-Sala S., Diekhans M., Jungreis I., Loveland J.E., Mudge J.M., et al. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res. 2023;51:D942–D949. doi: 10.1093/nar/gkac1071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.UniProt Consortium UniProt: the universal protein knowledgebase in 2023. Nucleic Acids Res. 2023;51:D523–D531. doi: 10.1093/nar/gkac1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ingolia N.T., Ghaemmaghami S., Newman J.R.S., Weissman J.S. Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science. 2009;324:218–223. doi: 10.1126/science.1168978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.McGlincy N.J., Ingolia N.T. Transcriptome-wide measurement of translation by ribosome profiling. Methods. 2017;126:112–129. doi: 10.1016/j.ymeth.2017.05.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.van Heesch S., Witte F., Schneider-Lunitz V., Schulz J.F., Adami E., Faber A.B., et al. The translational landscape of the human heart. Cell. 2019;178:242–260.e29. doi: 10.1016/j.cell.2019.05.010. [DOI] [PubMed] [Google Scholar]
  • 10.Calviello L., Hirsekorn A., Ohler U. Quantification of translation uncovers the functions of the alternative transcriptome. Nat. Struct. Mol. Biol. 2020;27:717–725. doi: 10.1038/s41594-020-0450-4. [DOI] [PubMed] [Google Scholar]
  • 11.Gaertner B., van Heesch S., Schneider-Lunitz V., Schulz J.F., Witte F., Blachut S., et al. A human ESC-based screen identifies a role for the translated lncRNA in pancreatic endocrine differentiation. Elife. 2020;9 doi: 10.7554/eLife.58659. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Fagerberg L., Hallström B.M., Oksvold P., Kampf C., Djureinovic D., Odeberg J., et al. Analysis of the human tissue-specific expression by genome-wide integration of transcriptomics and antibody-based proteomics. Mol. Cell. Proteomics. 2014;13:397–406. doi: 10.1074/mcp.M113.035600. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Krug K., Jaehnig E.J., Satpathy S., Blumenberg L., Karpova A., Anurag M., et al. Proteogenomic landscape of breast cancer tumorigenesis and targeted therapy. Cell. 2020;183:1436–1456.e31. doi: 10.1016/j.cell.2020.10.036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Cao L., Huang C., Cui Zhou D., Hu Y., Lih T.M., Savage S.R., et al. Proteogenomic characterization of pancreatic ductal adenocarcinoma. Cell. 2021;184:5031–5052.e26. doi: 10.1016/j.cell.2021.08.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Chothani S.P., Adami E., Widjaja A.A., Langley S.R., Viswanathan S., Pua C.J., et al. A high-resolution map of human RNA translation. Mol. Cell. 2022;82:2885–2899.e8. doi: 10.1016/j.molcel.2022.06.023. [DOI] [PubMed] [Google Scholar]
  • 16.Mudge J.M., Ruiz-Orera J., Prensner J.R., Brunet M.A., Calvet F., Jungreis I., et al. Standardized annotation of translated open reading frames. Nat. Biotechnol. 2022;40:994–999. doi: 10.1038/s41587-022-01369-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Prensner J.R., Enache O.M., Luria V., Krug K., Clauser K.R., Dempster J.M., et al. Noncanonical open reading frames encode functional proteins essential for cancer cell survival. Nat. Biotechnol. 2021;39:697–704. doi: 10.1038/s41587-020-00806-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Chen J., Brunner A.-D., Cogan J.Z., Nuñez J.K., Fields A.P., Adamson B., et al. Pervasive functional translation of noncanonical human open reading frames. Science. 2020;367:1140–1146. doi: 10.1126/science.aay0262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ji Z., Song R., Regev A., Struhl K. Many lncRNAs, 5’UTRs, and pseudogenes are translated and some are likely to express functional proteins. Elife. 2015;4 doi: 10.7554/eLife.08890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Raj A., Wang S.H., Shim H., Harpak A., Li Y.I., Engelmann B., et al. Thousands of novel translated open reading frames in humans inferred by ribosome footprint profiling. Elife. 2016;5 doi: 10.7554/eLife.13328. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Martinez T.F., Chu Q., Donaldson C., Tan D., Shokhirev M.N., Saghatelian A. Accurate annotation of human protein-coding small open reading frames. Nat. Chem. Biol. 2020;16:458–468. doi: 10.1038/s41589-019-0425-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Aspden J.L., Eyre-Walker Y.C., Phillips R.J., Amin U., Mumtaz M.A.S., Brocard M., et al. Extensive translation of small open reading frames revealed by Poly-Ribo-seq. Elife. 2014;3 doi: 10.7554/eLife.03528. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Douka K., Birds I., Wang D., Kosteletos A., Clayton S., Byford A., et al. Cytoplasmic long noncoding RNAs are differentially regulated and translated during human neuronal differentiation. RNA. 2021;27:1082–1101. doi: 10.1261/rna.078782.121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Fedorova A.D., Kiniry S.J., Andreev D.E., Mudge J.M., Baranov P.V. Thousands of human non-AUG extended proteoforms lack evidence of evolutionary selection among mammals. Nat. Commun. 2022;13:7910. doi: 10.1038/s41467-022-35595-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Van Damme P., Gawron D., Van Criekinge W., Menschaert G. N-terminal proteomics and ribosome profiling provide a comprehensive view of the alternative translation initiation landscape in mice and men. Mol. Cell. Proteomics. 2014;13:1245–1261. doi: 10.1074/mcp.M113.036442. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Koch A., Gawron D., Steyaert S., Ndah E., Crappé J., De Keulenaer S., et al. A proteogenomics approach integrating proteomics and ribosome profiling increases the efficiency of protein identification and enables the discovery of alternative translation start sites. Proteomics. 2014;14:2688–2698. doi: 10.1002/pmic.201400180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Menschaert G., Van Criekinge W., Notelaers T., Koch A., Crappé J., Gevaert K., et al. Deep proteome coverage based on ribosome profiling aids mass spectrometry-based protein and peptide discovery and provides evidence of alternative translation products and near-cognate translation initiation events. Mol. Cell. Proteomics. 2013;12:1780–1790. doi: 10.1074/mcp.M113.027540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Griffin G.K., Wu J., Iracheta-Vellve A., Patti J.C., Hsu J., Davis T., et al. Epigenetic silencing by SETDB1 suppresses tumour intrinsic immunogenicity. Nature. 2021;595:309–314. doi: 10.1038/s41586-021-03520-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Al-Turki T.M., Griffith J.D. Mammalian telomeric RNA (TERRA) can be translated to produce valine–arginine and glycine–leucine dipeptide repeat proteins. Proc. Natl. Acad. Sci. U. S. A. 2023;120 doi: 10.1073/pnas.2221529120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Omenn G.S., Lane L., Lundberg E.K., Overall C.M., Deutsch E.W. Progress on the HUPO draft human proteome: 2017 metrics of the human proteome project. J. Proteome Res. 2017;16:4281–4287. doi: 10.1021/acs.jproteome.7b00375. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Schwaid A.G., Shannon D.A., Ma J., Slavoff S.A., Levin J.Z., Weerapana E., et al. Chemoproteomic discovery of cysteine-containing human short open reading frames. J. Am. Chem. Soc. 2013;135:16750–16753. doi: 10.1021/ja406606j. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Cao X., Khitun A., Na Z., Dumitrescu D.G., Kubica M., Olatunji E., et al. Comparative proteomic profiling of unannotated microproteins and alternative proteins in human cell lines. J. Proteome Res. 2020;19:3418–3426. doi: 10.1021/acs.jproteome.0c00254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Ma J., Diedrich J.K., Jungreis I., Donaldson C., Vaughan J., Kellis M., et al. Improved identification and analysis of small open reading frame encoded Polypeptides. Anal. Chem. 2016;88:3967–3975. doi: 10.1021/acs.analchem.6b00191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Ouspenskaia T., Law T., Clauser K.R., Klaeger S., Sarkizova S., Aguet F., et al. Unannotated proteins expand the MHC-I-restricted immunopeptidome in cancer. Nat. Biotechnol. 2022;40:209–217. doi: 10.1038/s41587-021-01021-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Chong C., Coukos G., Bassani-Sternberg M. Identification of tumor antigens with immunopeptidomics. Nat. Biotechnol. 2022;40:175–188. doi: 10.1038/s41587-021-01038-8. [DOI] [PubMed] [Google Scholar]
  • 36.Laumont C.M., Vincent K., Hesnard L., Audemard É., Bonneil É., Laverdure J.-P., et al. Noncoding regions are the main source of targetable tumor-specific antigens. Sci. Transl. Med. 2018;10 doi: 10.1126/scitranslmed.aau5516. [DOI] [PubMed] [Google Scholar]
  • 37.Laumont C.M., Perreault C. Exploiting non-canonical translation to identify new targets for T cell-based cancer immunotherapy. Cell. Mol. Life Sci. 2018;75:607–621. doi: 10.1007/s00018-017-2628-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Laumont C.M., Daouda T., Laverdure J.-P., Bonneil É., Caron-Lizotte O., Hardy M.-P., et al. Global proteogenomic analysis of human MHC class I-associated peptides derived from non-canonical reading frames. Nat. Commun. 2016;7 doi: 10.1038/ncomms10238. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Ruiz Cuevas M.V., Hardy M.-P., Hollý J., Bonneil É., Durette C., Courcelles M., et al. Most non-canonical proteins uniquely populate the proteome or immunopeptidome. Cell Rep. 2021;34 doi: 10.1016/j.celrep.2021.108815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Calvo S.E., Pagliarini D.J., Mootha V.K. Upstream open reading frames cause widespread reduction of protein expression and are polymorphic among humans. Proc. Natl. Acad. Sci. U. S. A. 2009;106:7507–7512. doi: 10.1073/pnas.0810916106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Johnstone T.G., Bazzini A.A., Giraldez A.J. Upstream ORFs are prevalent translational repressors in vertebrates. EMBO J. 2016;35:706–723. doi: 10.15252/embj.201592759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wu Q., Wright M., Gogol M.M., Bradford W.D., Zhang N., Bazzini A.A. Translation of small downstream ORFs enhances translation of canonical main open reading frames. EMBO J. 2020;39 doi: 10.15252/embj.2020104763. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Pauli A., Norris M.L., Valen E., Chew G.-L., Gagnon J.A., Zimmerman S., et al. Toddler: an embryonic signal that promotes cell movement via apelin receptors. Science. 2014;343 doi: 10.1126/science.1248636. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Khan Y.A., Jungreis I., Wright J.C., Mudge J.M., Choudhary J.S., Firth A.E., et al. Evidence for a novel overlapping coding sequence in POLG initiated at a CUG start codon. BMC Genet. 2020;21:25. doi: 10.1186/s12863-020-0828-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Loughran G., Zhdanov A.V., Mikhaylova M.S., Rozov F.N., Datskevich P.N., Kovalchuk S.I., et al. Unusually efficient CUG initiation of an overlapping reading frame in mRNA yields novel protein POLGARF. Proc. Natl. Acad. Sci. U. S. A. 2020;117:24936–24946. doi: 10.1073/pnas.2001433117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Boix O., Martinez M., Vidal S., Giménez-Alejandre M., Palenzuela L., Lorenzo-Sanz L., et al. pTINCR microprotein promotes epithelial differentiation and suppresses tumor growth through CDC42 SUMOylation and activation. Nat. Commun. 2022;13:6840. doi: 10.1038/s41467-022-34529-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Bi P., Ramirez-Martinez A., Li H., Cannavino J., McAnally J.R., Shelton J.M., et al. Control of muscle formation by the fusogenic micropeptide myomixer. Science. 2017;356:323–327. doi: 10.1126/science.aam9361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Anderson D.M., Anderson K.M., Chang C.-L., Makarewich C.A., Nelson B.R., McAnally J.R., et al. A micropeptide encoded by a putative long noncoding RNA regulates muscle performance. Cell. 2015;160:595–606. doi: 10.1016/j.cell.2015.01.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Slavoff S.A., Mitchell A.J., Schwaid A.G., Cabili M.N., Ma J., Levin J.Z., et al. Peptidomic discovery of short open reading frame-encoded peptides in human cells. Nat. Chem. Biol. 2013;9:59–64. doi: 10.1038/nchembio.1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Martinez T.F., Lyons-Abbott S., Bookout A.L., De Souza E.V., Donaldson C., Vaughan J.M., et al. Profiling mouse brown and white adipocytes to identify metabolically relevant small ORFs and functional microproteins. Cell Metab. 2023;35:166–183.e11. doi: 10.1016/j.cmet.2022.12.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Mackowiak S.D., Zauber H., Bielow C., Thiel D., Kutz K., Calviello L., et al. Extensive identification and analysis of conserved small ORFs in animals. Genome Biol. 2015;16:179. doi: 10.1186/s13059-015-0742-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Bazzini A.A., Johnstone T.G., Christiano R., Mackowiak S.D., Obermayer B., Fleming E.S., et al. Identification of small ORFs in vertebrates using ribosome footprinting and evolutionary conservation. EMBO J. 2014;33:981–993. doi: 10.1002/embj.201488411. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Chong C., Müller M., Pak H., Harnett D., Huber F., Grun D., et al. Integrated proteogenomic deep sequencing and analytics accurately identify non-canonical peptides in tumor immunopeptidomes. Nat. Commun. 2020;11:1293. doi: 10.1038/s41467-020-14968-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Huang N., Li F., Zhang M., Zhou H., Chen Z., Ma X., et al. An upstream open reading frame in phosphatase and tensin homolog encodes a circuit breaker of lactate metabolism. Cell Metab. 2021;33:454. doi: 10.1016/j.cmet.2021.01.008. [DOI] [PubMed] [Google Scholar]
  • 55.Na Z., Dai X., Zheng S.-J., Bryant C.J., Loh K.H., Su H., et al. Mapping subcellular localizations of unannotated microproteins and alternative proteins with MicroID. Mol. Cell. 2022;82:2900–2911.e7. doi: 10.1016/j.molcel.2022.06.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Jayaram D.R., Frost S., Argov C., Liju V.B., Anto N.P., Muraleedharan A., et al. Unraveling the hidden role of a uORF-encoded peptide as a kinase inhibitor of PKCs. Proc. Natl. Acad. Sci. U. S. A. 2021;118 doi: 10.1073/pnas.2018899118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Sandmann C.-L., Schulz J.F., Ruiz-Orera J., Kirchner M., Ziehm M., Adami E., et al. Evolutionary origins and interactomes of human, young microproteins and small peptides translated from short open reading frames. Mol. Cell. 2023;83:994–1011.e18. doi: 10.1016/j.molcel.2023.01.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Tanaka M., Sotta N., Yamazumi Y., Yamashita Y., Miwa K., Murota K., et al. The Minimum open reading frame, AUG-stop, Induces Boron-dependent ribosome stalling and mRNA degradation. Plant Cell. 2016;28:2830–2849. doi: 10.1105/tpc.16.00481. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Dau T., Bartolomucci G., Rappsilber J. Proteomics using protease alternatives to trypsin benefits from sequential digestion with trypsin. Anal. Chem. 2020;92:9523–9527. doi: 10.1021/acs.analchem.0c00478. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Calviello L., Ohler U. Beyond read-counts: Ribo-Seq data analysis to understand the functions of the transcriptome. Trends Genet. 2017;33:728–744. doi: 10.1016/j.tig.2017.08.003. [DOI] [PubMed] [Google Scholar]
  • 61.Calviello L., Mukherjee N., Wyler E., Zauber H., Hirsekorn A., Selbach M., et al. Detecting actively translated open reading frames in ribosome profiling data. Nat. Methods. 2016;13:165–170. doi: 10.1038/nmeth.3688. [DOI] [PubMed] [Google Scholar]
  • 62.Fremin B.J., Bhatt A.S. Structured RNA contaminants in bacterial Ribo-Seq. mSphere. 2020;5:e00855–e00920. doi: 10.1128/mSphere.00855-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Chung B.Y., Hardcastle T.J., Jones J.D., Irigoyen N., Firth A.E., Baulcombe D.C., et al. The use of duplex-specific nuclease in ribosome profiling and a user-friendly software package for Ribo-Seq data analysis. RNA. 2015;21:1731–1745. doi: 10.1261/rna.052548.115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Hsu P.Y., Calviello L., Wu H.-Y.L., Li F.-W., Rothfels C.J., Ohler U., et al. Super-resolution ribosome profiling reveals unannotated translation events in. Proc. Natl. Acad. Sci. U. S. A. 2016;113:E7126–E7135. doi: 10.1073/pnas.1614788113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Diament A., Tuller T. Estimation of ribosome profiling performance and reproducibility at various levels of resolution. Biol. Direct. 2016;11:24. doi: 10.1186/s13062-016-0127-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Zhang P., He D., Xu Y., Hou J., Pan B.-F., Wang Y., et al. Genome-wide identification and differential analysis of translational initiation. Nat. Commun. 2017;8:1749. doi: 10.1038/s41467-017-01981-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Erhard F., Halenius A., Zimmermann C., L’Hernault A., Kowalewski D.J., Weekes M.P., et al. Improved Ribo-Seq enables identification of cryptic translation events. Nat. Methods. 2018;15:363–366. doi: 10.1038/nmeth.4631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Xiao Z., Huang R., Xing X., Chen Y., Deng H., Yang X. De novo annotation and characterization of the translatome with ribosome profiling data. Nucleic Acids Res. 2018;46:e61. doi: 10.1093/nar/gky179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Choudhary S., Li W., D Smith A. Accurate detection of short and long active ORFs using Ribo-seq data. Bioinformatics. 2020;36:2053–2059. doi: 10.1093/bioinformatics/btz878. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Fields A.P., Rodriguez E.H., Jovanovic M., Stern-Ginossar N., Haas B.J., Mertins P., et al. A Regression-based analysis of ribosome-profiling data reveals a conserved complexity to mammalian translation. Mol. Cell. 2015;60:816–827. doi: 10.1016/j.molcel.2015.11.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Clauwaert J., Menschaert G., Waegeman W. DeepRibo: a neural network for precise gene annotation of prokaryotes by combining ribosome profiling signal and binding site patterns. Nucleic Acids Res. 2019;47:e36. doi: 10.1093/nar/gkz061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Clauwaert J., McVey Z., Gupta R., Menschaert G. TIS Transformer: remapping the human proteome using deep learning. NAR Genom. Bioinform. 2023;5:lqad021. doi: 10.1093/nargab/lqad021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Freudenmann L.K., Marcu A., Stevanović S. Mapping the tumour human leukocyte antigen (HLA) ligandome by mass spectrometry. Immunology. 2018;154:331–345. doi: 10.1111/imm.12936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Bassani-Sternberg M., Bräunlein E., Klar R., Engleitner T., Sinitcyn P., Audehm S., et al. Direct identification of clinically relevant neoepitopes presented on native human melanoma tissue by mass spectrometry. Nat. Commun. 2016;7 doi: 10.1038/ncomms13404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Shapiro I.E., Bassani-Sternberg M. The impact of immunopeptidomics: from basic research to clinical implementation. Semin. Immunol. 2023;66 doi: 10.1016/j.smim.2023.101727. [DOI] [PubMed] [Google Scholar]
  • 76.Abelin J.G., Keskin D.B., Sarkizova S., Hartigan C.R., Zhang W., Sidney J., et al. Mass spectrometry profiling of HLA-associated peptidomes in Mono-allelic cells enables more accurate epitope prediction. Immunity. 2017;46:315–326. doi: 10.1016/j.immuni.2017.02.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Purcell A.W., Ramarathinam S.H., Ternette N. Mass spectrometry–based identification of MHC-bound peptides for immunopeptidomics. Nat. Protoc. 2019;14:1687–1707. doi: 10.1038/s41596-019-0133-y. [DOI] [PubMed] [Google Scholar]
  • 78.Bassani-Sternberg M., Pletscher-Frankild S., Jensen L.J., Mann M. Mass spectrometry of human leukocyte antigen class I peptidomes reveals strong effects of protein abundance and turnover on antigen presentation. Mol. Cell. Proteomics. 2015;14:658–673. doi: 10.1074/mcp.M114.042812. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Yewdell J.W., Reits E., Neefjes J. Making sense of mass destruction: quantitating MHC class I antigen presentation. Nat. Rev. Immunol. 2003;3:952–961. doi: 10.1038/nri1250. [DOI] [PubMed] [Google Scholar]
  • 80.Yewdell J.W. Immunology. Hide and seek in the peptidome. Science. 2003;301:1334–1335. doi: 10.1126/science.1089553. [DOI] [PubMed] [Google Scholar]
  • 81.Blaha D.T., Anderson S.D., Yoakum D.M., Hager M.V., Zha Y., Gajewski T.F., et al. High-throughput stability screening of neoantigen/HLA complexes improves immunogenicity predictions. Cancer Immunol. Res. 2019;7:50–61. doi: 10.1158/2326-6066.CIR-18-0395. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Prevosto C., Usmani M.F., McDonald S., Gumienny A.M., Key T., Goodman R.S., et al. Allele-independent turnover of human leukocyte antigen (HLA) class Ia molecules. PLoS One. 2016;11 doi: 10.1371/journal.pone.0161011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Abelin J.G., Bergstrom E.J., Taylor H.B., Rivera K.D., Klaeger S., Xu C., et al. MONTE enables serial immunopeptidome, ubiquitylome, proteome, phosphoproteome, acetylome analyses of sample-limited tissues. bioRxiv. 2022 doi: 10.1101/2021.06.22.449417. [preprint] [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Boehm K.M., Bhinder B., Raja V.J., Dephoure N., Elemento O. Predicting peptide presentation by major histocompatibility complex class I: an improved machine learning approach to the immunopeptidome. BMC Bioinformatics. 2019;20:7. doi: 10.1186/s12859-018-2561-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Abelin J.G., Harjanto D., Malloy M., Suri P., Colson T., Goulding S.P., et al. Defining HLA-II ligand processing and binding rules with mass spectrometry enhances cancer epitope prediction. Immunity. 2019;51:766–779.e17. doi: 10.1016/j.immuni.2019.08.012. [DOI] [PubMed] [Google Scholar]
  • 86.Alvarez B., Reynisson B., Barra C., Buus S., Ternette N., Connelley T., et al. NNAlign_MA; MHC peptidome deconvolution for accurate MHC binding motif characterization and improved T-cell epitope predictions. Mol. Cell. Proteomics. 2019;18:2459–2477. doi: 10.1074/mcp.TIR119.001658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.O’Donnell T.J., Rubinsteyn A., Laserson U. MHCflurry 2.0: improved pan-allele prediction of MHC class I-presented peptides by incorporating antigen processing. Cell Syst. 2020;11:418–419. doi: 10.1016/j.cels.2020.09.001. [DOI] [PubMed] [Google Scholar]
  • 88.Jurtz V., Paul S., Andreatta M., Marcatili P., Peters B., Nielsen M. NetMHCpan-4.0: improved peptide-MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data. J. Immunol. 2017;199:3360–3368. doi: 10.4049/jimmunol.1700893. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Sarkizova S., Klaeger S., Le P.M., Li L.W., Oliveira G., Keshishian H., et al. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nat. Biotechnol. 2020;38:199–209. doi: 10.1038/s41587-019-0322-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Taylor H.B., Klaeger S., Clauser K.R., Sarkizova S., Weingarten-Gabbay S., Graham D.B., et al. MS-based HLA-II peptidomics combined with multiomics will aid the development of future immunotherapies. Mol. Cell. Proteomics. 2021;20 doi: 10.1016/j.mcpro.2021.100116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Chen B., Khodadoust M.S., Olsson N., Wagar L.E., Fast E., Liu C.L., et al. Predicting HLA class II antigen presentation through integrated deep learning. Nat. Biotechnol. 2019;37:1332–1343. doi: 10.1038/s41587-019-0280-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Racle J., Michaux J., Rockinger G.A., Arnaud M., Bobisse S., Chong C., et al. Robust prediction of HLA class II epitopes by deep motif deconvolution of immunopeptidomes. Nat. Biotechnol. 2019;37:1283–1286. doi: 10.1038/s41587-019-0289-6. [DOI] [PubMed] [Google Scholar]
  • 93.Shao X.M., Bhattacharya R., Huang J., Sivakumar I.K.A., Tokheim C., Zheng L., et al. High-throughput prediction of MHC class I and II Neoantigens with MHCnuggets. Cancer Immunol. Res. 2020;8:396–408. doi: 10.1158/2326-6066.CIR-19-0464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Lozano-Rabella M., Garcia-Garijo A., Palomero J., Yuste-Estevanez A., Erhard F., Martín-Liberal J., et al. Immunogenicity of non-canonical HLA-I tumor ligands identified through proteogenomics. bioRxiv. 2022 doi: 10.1101/2022.11.07.514886. [preprint] [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Erhard F., Dölken L., Schilling B., Schlosser A. Identification of the cryptic HLA-I immunopeptidome. Cancer Immunol. Res. 2020;8:1018–1026. doi: 10.1158/2326-6066.CIR-19-0886. [DOI] [PubMed] [Google Scholar]
  • 96.Lichti C.F., Vigneron N., Clauser K.R., Van den Eynde B.J., Bassani-Sternberg M. Navigating critical challenges associated with immunopeptidomics-based detection of proteasomal spliced peptide candidates. Cancer Immunol. Res. 2022;10:275–284. doi: 10.1158/2326-6066.CIR-21-0727. [DOI] [PubMed] [Google Scholar]
  • 97.Gessulat S., Schmidt T., Zolg D.P., Samaras P., Schnatbaum K., Zerweck J., et al. Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning. Nat. Methods. 2019;16:509–518. doi: 10.1038/s41592-019-0426-7. [DOI] [PubMed] [Google Scholar]
  • 98.Wilhelm M., Zolg D.P., Graber M., Gessulat S., Schmidt T., Schnatbaum K., et al. Deep learning boosts sensitivity of mass spectrometry-based immunopeptidomics. Nat. Commun. 2021;12:3346. doi: 10.1038/s41467-021-23713-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Declercq A., Bouwmeester R., Chiva C., Sabidó E., Hirschler A., Carapito C., et al. Updated MS2PIP web server supports cutting-edge proteomics applications. Nucleic Acids Res. 2023;51:W338–W342. doi: 10.1093/nar/gkad335. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Bouwmeester R., Gabriels R., Hulstaert N., Martens L., Degroeve S. DeepLC can predict retention times for peptides that carry as-yet unseen modifications. Nat. Methods. 2021;18:1363–1369. doi: 10.1038/s41592-021-01301-5. [DOI] [PubMed] [Google Scholar]
  • 101.Li K., Jain A., Malovannaya A., Wen B., Zhang B. DeepRescore: leveraging deep learning to improve peptide identification in immunopeptidomics. Proteomics. 2020;20 doi: 10.1002/pmic.201900334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Deutsch E.W., Lane L., Overall C.M., Bandeira N., Baker M.S., Pineau C., et al. Human proteome project mass spectrometry data interpretation guidelines 3.0. J. Proteome Res. 2019;18:4108–4116. doi: 10.1021/acs.jproteome.9b00542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Adhikari S., Nice E.C., Deutsch E.W., Lane L., Omenn G.S., Pennington S.R., et al. A high-stringency blueprint of the human proteome. Nat. Commun. 2020;11:5301. doi: 10.1038/s41467-020-19045-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Deutsch E.W., Overall C.M., Van Eyk J.E., Baker M.S., Paik Y.-K., Weintraub S.T., et al. Human proteome project mass spectrometry data interpretation guidelines 2.1. J. Proteome Res. 2016;15:3961–3970. doi: 10.1021/acs.jproteome.6b00392. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Deutsch E.W., Perez-Riverol Y., Carver J., Kawano S., Mendoza L., Van Den Bossche T., et al. Universal spectrum identifier for mass spectra. Nat. Methods. 2021;18:768–770. doi: 10.1038/s41592-021-01184-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Kim M.-S., Pinto S.M., Getnet D., Nirujogi R.S., Manda S.S., Chaerkady R., et al. A draft map of the human proteome. Nature. 2014;509:575–581. doi: 10.1038/nature13302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Oyama M., Kozuka-Hata H., Suzuki Y., Semba K., Yamamoto T., Sugano S. Diversity of translation start sites may define increased complexity of the human short ORFeome. Mol. Cell. Proteomics. 2007;6:1000–1006. doi: 10.1074/mcp.M600297-MCP200. [DOI] [PubMed] [Google Scholar]
  • 108.Volders P.J., Verheggen K., Menschaert G., Vandepoele K., Martens L., Vandesompele J., et al. An update on LNCipedia: a database for annotated human lncRNA sequences. Nucleic Acids Res. 2015;43:4363–4364. doi: 10.1093/nar/gkv295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Iyer M.K., Niknafs Y.S., Malik R., Singhal U., Sahu A., Hosono Y., et al. The landscape of long noncoding RNAs in the human transcriptome. Nat. Genet. 2015;47:199–208. doi: 10.1038/ng.3192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Wacholder A., Carvunis A.-R. Rare detection of noncanonical proteins in yeast mass spectrometry studies. bioRxiv. 2023 doi: 10.1101/2023.03.09.531963. [preprint] [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Verheggen K., Volders P.-J., Mestdagh P., Menschaert G., Van Damme P., Gevaert K., et al. Noncoding after all: biases in proteomics data do not Explain observed absence of lncRNA translation products. J. Proteome Res. 2017;16:2508–2515. doi: 10.1021/acs.jproteome.7b00085. [DOI] [PubMed] [Google Scholar]
  • 112.Bogaert A., Fijalkowska D., Staes A., Van de Steene T., Demol H., Gevaert K. Limited evidence for protein products of noncoding transcripts in the HEK293T cellular Cytosol. Mol. Cell. Proteomics. 2022;21 doi: 10.1016/j.mcpro.2022.100264. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Cassidy L., Kaulich P.T., Tholey A. Proteoforms expand the world of microproteins and short open reading frame-encoded peptides. iScience. 2023;26 doi: 10.1016/j.isci.2023.106069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Kesner J.S., Chen Z., Shi P., Aparicio A.O., Murphy M.R., Guo Y., et al. Noncoding translation mitigation. Nature. 2023;617:395–402. doi: 10.1038/s41586-023-05946-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Fritsch C., Herrmann A., Nothnagel M., Szafranski K., Huse K., Schumann F., et al. Genome-wide search for novel human uORFs and N-terminal protein extensions using ribosomal footprinting. Genome Res. 2012;22:2208–2218. doi: 10.1101/gr.139568.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Li Y., Zhou H., Chen X., Zheng Y., Kang Q., Hao D., et al. SmProt: a reliable repository with comprehensive annotation of small proteins identified from ribosome profiling. Genomics Proteomics Bioinformatics. 2021;19:602–610. doi: 10.1016/j.gpb.2021.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Olexiouk V., Crappé J., Verbruggen S., Verhegen K., Martens L., Menschaert G. sORFs.org: a repository of small ORFs identified by ribosome profiling. Nucleic Acids Res. 2016;44:D324–D329. doi: 10.1093/nar/gkv1175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Wang H., Yang L., Wang Y., Chen L., Li H., Xie Z. RPFdb v2. 0: an updated database for genome-wide information of translated mRNA generated from ribosome profiling. Nucleic Acids Res. 2019;47:D230–D234. doi: 10.1093/nar/gky978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Xie S.-Q., Nie P., Wang Y., Wang H., Li H., Yang Z., et al. RPFdb: a database for genome wide information of translated mRNA generated from ribosome profiling. Nucleic Acids Res. 2016;44:D254–D258. doi: 10.1093/nar/gkv972. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Ji X., Cui C., Cui Q. smORFunction: a tool for predicting functions of small open reading frames and microproteins. BMC Bioinformatics. 2020;21:455. doi: 10.1186/s12859-020-03805-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Brunet M.A., Brunelle M., Lucier J.-F., Delcourt V., Levesque M., Grenier F., et al. OpenProt: a more comprehensive guide to explore eukaryotic coding potential and proteomes. Nucleic Acids Res. 2019;47:D403–D410. doi: 10.1093/nar/gky936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Brunet M.A., Lucier J.-F., Levesque M., Leblanc S., Jacques J.-F., Al-Saedi H.R.H., et al. OpenProt 2021: deeper functional annotation of the coding potential of eukaryotic genomes. Nucleic Acids Res. 2021;49:D380–D388. doi: 10.1093/nar/gkaa1036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Duffy E.E., Finander B., Choi G., Carter A.C., Pritisanac I., Alam A., et al. Developmental dynamics of RNA translation in the human brain. Nat. Neurosci. 2022;25:1353–1365. doi: 10.1038/s41593-022-01164-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Smirnova V.V., Shestakova E.D., Nogina D.S., Mishchenko P.A., Prikazchikova T.A., Zatsepin T.S., et al. Ribosomal leaky scanning through a translated uORF requires eIF4G2. Nucleic Acids Res. 2022;50:1111–1127. doi: 10.1093/nar/gkab1286. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Andreev D.E., Loughran G., Fedorova A.D., Mikhaylova M.S., Shatsky I.N., Baranov P.V. Non-AUG translation initiation in mammals. Genome Biol. 2022;23:111. doi: 10.1186/s13059-022-02674-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Stacey S.N., Jordan D., Williamson A.J., Brown M., Coote J.H., Arrand J.R. Leaky scanning is the predominant mechanism for translation of human papillomavirus type 16 E7 oncoprotein from E6/E7 bicistronic mRNA. J. Virol. 2000;74:7284–7297. doi: 10.1128/jvi.74.16.7284-7297.2000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Duss O., Stepanyuk G.A., Puglisi J.D., Williamson J.R. Transient protein-RNA interactions guide nascent ribosomal RNA folding. Cell. 2019;179:1357–1369.e16. doi: 10.1016/j.cell.2019.10.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Karamyshev A.L., Karamysheva Z.N. Lost in translation: ribosome-associated mRNA and protein quality controls. Front. Genet. 2018;9:431. doi: 10.3389/fgene.2018.00431. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Gelhausen R., Müller T., Svensson S.L., Alkhnbashi O.S., Sharma C.M., Eggenhofer F., et al. RiboReport - benchmarking tools for ribosome profiling-based identification of open reading frames in bacteria. Brief. Bioinform. 2022;23 doi: 10.1093/bib/bbab549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Kiniry S.J., Michel A.M., Baranov P.V. Computational methods for ribosome profiling data analysis. Wiley Interdiscip. Rev. RNA. 2020;11 doi: 10.1002/wrna.1577. [DOI] [PubMed] [Google Scholar]
  • 131.Lei T., Chang Y., Yao C., Zhang H. A systematic evaluation revealed that detecting translated non-canonical ORFs from ribosome profiling data remains challenging. bioRxiv. 2022 doi: 10.1101/2022.12.11.520003. [preprint] [DOI] [PubMed] [Google Scholar]
  • 132.Blackwood E.M., Lugo T.G., Kretzner L., King M.W., Street A.J., Witte O.N., et al. Functional analysis of the AUG- and CUG-initiated forms of the c-Myc protein. Mol. Biol. Cell. 1994;5:597–609. doi: 10.1091/mbc.5.5.597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Prats H., Kaghad M., Prats A.C., Klagsbrun M., Lélias J.M., Liauzun P., et al. High molecular mass forms of basic fibroblast growth factor are initiated by alternative CUG codons. Proc. Natl. Acad. Sci. U. S. A. 1989;86:1836–1840. doi: 10.1073/pnas.86.6.1836. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Cao X., Slavoff S.A. Non-AUG start codons: expanding and regulating the small and alternative ORFeome. Exp. Cell Res. 2020;391 doi: 10.1016/j.yexcr.2020.111973. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Ingolia N.T., Lareau L.F., Weissman J.S. Ribosome profiling of mouse embryonic stem cells reveals the complexity and dynamics of mammalian proteomes. Cell. 2011;147:789–802. doi: 10.1016/j.cell.2011.10.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Lee S., Liu B., Lee S., Huang S.-X., Shen B., Qian S.-B. Global mapping of translation initiation sites in mammalian cells at single-nucleotide resolution. Proc. Natl. Acad. Sci. U. S. A. 2012;109:E2424–E2432. doi: 10.1073/pnas.1207846109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Champagne J., Pataskar A., Blommaert N., Nagel R., Wernaart D., Ramalho S., et al. Oncogene-dependent sloppiness in mRNA translation. Mol. Cell. 2021;81:4709–4721.e9. doi: 10.1016/j.molcel.2021.09.002. [DOI] [PubMed] [Google Scholar]
  • 138.Janssen J.W., Vaandrager J.W., Heuser T., Jauch A., Kluin P.M., Geelen E., et al. Concurrent activation of a novel putative transforming gene, myeov, and cyclin D1 in a subset of multiple myeloma cell lines with t(11;14)(q13;q32) Blood. 2000;95:2691–2698. [PubMed] [Google Scholar]
  • 139.Dolstra H., Fredrix H., Preijers F., Goulmy E., Figdor C.G., de Witte T.M., et al. Recognition of a B cell leukemia-associated minor histocompatibility antigen by CTL. J. Immunol. 1997;158:560–565. [PubMed] [Google Scholar]
  • 140.Ruiz-Orera J., Villanueva-Cañas J.L., Albà M.M. Evolution of new proteins from translated sORFs in long non-coding RNAs. Exp. Cell Res. 2020;391 doi: 10.1016/j.yexcr.2020.111940. [DOI] [PubMed] [Google Scholar]
  • 141.Vakirlis N., Vance Z., Duggan K.M., McLysaght A. De novo birth of functional microproteins in the human lineage. Cell Rep. 2022;41 doi: 10.1016/j.celrep.2022.111808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Broeils L.A., Ruiz-Orera J., Snel B., Hubner N., van Heesch S. Evolution and implications of de novo genes in humans. Nat. Ecol. Evol. 2023;7:804–815. doi: 10.1038/s41559-023-02014-y. [DOI] [PubMed] [Google Scholar]
  • 143.Erady C., Boxall A., Puntambekar S., Suhas Jagannathan N., Chauhan R., Chong D., et al. Pan-cancer analysis of transcripts encoding novel open-reading frames (nORFs) and their potential biological functions. NPJ Genom. Med. 2021;6:4. doi: 10.1038/s41525-020-00167-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Na Z., Luo Y., Cui D.S., Khitun A., Smelyansky S., Loria J.P., et al. Phosphorylation of a human microprotein promotes dissociation of biomolecular condensates. J. Am. Chem. Soc. 2021;143:12675–12687. doi: 10.1021/jacs.1c05386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.D’Lima N.G., Ma J., Winkler L., Chu Q., Loh K.H., Corpuz E.O., et al. A human microprotein that interacts with the mRNA decapping complex. Nat. Chem. Biol. 2017;13:174–180. doi: 10.1038/nchembio.2249. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Ruiz-Orera J., Messeguer X., Subirana J.A., Alba M.M. Long non-coding RNAs as a source of new peptides. Elife. 2014;3 doi: 10.7554/eLife.03523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Schlesinger D., Elsässer S.J. Revisiting sORFs: overcoming challenges to identify and characterize functional microproteins. FEBS J. 2022;289:53–74. doi: 10.1111/febs.15769. [DOI] [PubMed] [Google Scholar]
  • 148.Vakirlis N., Acar O., Hsu B., Castilho Coelho N., Van Oss S.B., Wacholder A., et al. De novo emergence of adaptive membrane proteins from thymine-rich genomic sequences. Nat. Commun. 2020;11:781. doi: 10.1038/s41467-020-14500-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Heames B., Buchel F., Aubel M., Tretyachenko V., Loginov D., Novák P., et al. Experimental characterization of de novo proteins and their unevolved random-sequence counterparts. Nat. Ecol. Evol. 2023;7:570–580. doi: 10.1038/s41559-023-02010-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 150.Kustatscher G., Collins T., Gingras A.-C., Guo T., Hermjakob H., Ideker T., et al. An open invitation to the understudied proteins initiative. Nat. Biotechnol. 2022;40:815–817. doi: 10.1038/s41587-022-01316-z. [DOI] [PubMed] [Google Scholar]
  • 151.Omenn G.S., Lane L., Overall C.M., Pineau C., Packer N.H., Cristea I.M., et al. The 2022 report on the human proteome from the HUPO human proteome project. J. Proteome Res. 2022;22:1024–1042. doi: 10.1021/acs.jproteome.2c00498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Kesner J.S., Chen Z., Aparicio A.A., Wu X. A unified model for the surveillance of translation in diverse noncoding sequences. bioRxiv. 2022 doi: 10.1101/2022.07.20.500724. [preprint] [DOI] [Google Scholar]
  • 153.Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.J. 2011;17:10–12. [Google Scholar]
  • 154.Langmead B., Salzberg S.L. Fast gapped-read alignment with Bowtie 2. Nat. Methods. 2012;9:357–359. doi: 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Dobin A., Davis C.A., Schlesinger F., Drenkow J., Zaleski C., Jha S., et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. doi: 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.Li H., Handsaker B., Wysoker A., Fennell T., Ruan J., Homer N., et al. The sequence alignment/map format and SAMtools. Bioinformatics. 2009;25:2078–2079. doi: 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157.Calviello L., Sydow D., Harnett D., Ohler U. Ribo-seQC: comprehensive analysis of cytoplasmic and organellar ribosome profiling data. bioRxiv. 2019 doi: 10.1101/601468. [preprint] [DOI] [Google Scholar]
  • 158.Paes W., Leonov G., Partridge T., Chikata T., Murakoshi H., Frangou A., et al. Contribution of proteasome-catalyzed peptide -splicing to viral targeting by CD8 T cells in HIV-1 infection. Proc. Natl. Acad. Sci. U. S. A. 2019;116:24748–24759. doi: 10.1073/pnas.1911622116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 159.Sun Y.H., Wang A., Song C., Shankar G., Srivastava R.K., Au K.F., et al. Single-molecule long-read sequencing reveals a conserved intact long RNA profile in sperm. Nat. Commun. 2021;12:1361. doi: 10.1038/s41467-021-21524-6. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Tables S1–S3
mmc1.xlsx (36.6MB, xlsx)
Supplemental Data
mmc2.docx (16.9KB, docx)

Articles from Molecular & Cellular Proteomics : MCP are provided here courtesy of American Society for Biochemistry and Molecular Biology

RESOURCES