Abstract
Candida albicans is a diploid pathogen known for its ability to live as a commensal fungus in healthy individuals but causing both superficial infections and disseminated candidiasis in immunocompromised patients where it is associated with high morbidity and mortality. Its success in colonizing the human host is attributed to a wide range of virulence traits that modulate interactions between the host and the pathogen, such as optimal growth rate at 37 °C, the ability to switch between yeast and hyphal forms, and a remarkable genomic and phenotypic plasticity. A fascinating aspect of its biology is a prominent heterogeneous proteome that arises from frequent genomic rearrangements, high allelic variation, and high levels of amino acid misincorporations in proteins. This leads to increased morphological and physiological phenotypic diversity of high adaptive potential, but the scope of such protein mistranslation is poorly understood due to technical difficulties in detecting and quantifying amino acid misincorporation events in complex protein samples. We have developed and optimized mass spectrometry and bioinformatics pipelines capable of identifying rare amino acid misincorporation events at the proteome level. We have also analyzed the proteomic profile of an engineered C. albicans strain that exhibits high level of leucine misincorporation at protein CUG sites and employed an in vivo quantitative gain-of-function fluorescence reporter system to validate our LC-MS/MS data. C. albicans misincorporates amino acids above the background level at protein sites of diverse codons, particularly at CUG, confirming our previous data on the quantification of leucine incorporation at single CUG sites of recombinant reporter proteins, but increasing misincorporation of Leucine at these sites does not alter the translational fidelity of the other codons. These findings indicate that the C. albicans statistical proteome exceeds prior estimates, suggesting that its highly plastic phenome may also be modulated by environmental factors due to translational ambiguity.
Keywords: translation fidelity, proteogenomics, CUG ambiguity, Candida albicans, mass spectrometry, bioinformatics
Graphical Abstract

Highlights
-
•
A proteogenomic pipeline identifies protein biosynthesis errors in Candida albicans.
-
•
Misincorporation of Leu at CUG sites detected at the proteome scale.
-
•
C. albicans misincorporates amino acids at diverse codons.
-
•
High level mistranslation is an intrinsic characteristic of C. albicans biology.
In Brief
A new proteogenomic pipeline based on mass spectrometry and bioinformatics enables the identification and quantification of amino acid misincorporations in Candida albicans. Results revealed diverse codon misincorporation patterns and a higher frequency of Leu misincorporation at CUG sites. Proteogenomic data was validated using fluorescence reporters expressed in an engineered strain that confirmed higher leucine misincorporation than WT strains. The data expands the understanding of C. albicans' proteomic complexity and its potential modulation by environmental factors.
All living organisms face challenging environments and rely on adaptation mechanisms to survive and thrive. To colonize multiple niches in the human body for instance, pathogens must face host immune defenses, high temperature, niches with different pH and diverse nutrient availabilities, and even therapeutic drugs (1). Immediate adaptation to these extracellular challenges normally relies on the activation of signaling pathways that trigger transcriptional responses modulating the fungal proteome. Other adaptive mechanisms (long term) involve selection of advantageous genetic mutations that can be passed to the next generations. Such mutations change all polypeptides encoded by mutant coding genes, that is, change the proteome in a quantitative and heritable manner and normally have visible phenotypic outcomes. Protein biosynthesis errors produced during mRNA translation (mistranslation) have emerged recently as alternative adaptive mechanisms that operate through proteome diversification and expanded functionality. Unlike genetic mutations, translational mutations occur randomly in a small percentage of the polypeptides of each protein, affect a very large number of proteins or even the entire proteome, are normally transient and may not be passed to the following generations. Importantly, mistranslation produces statistical proteins, that is, proteins that are a mixture of WT and mutated polypeptides where the latter are diverse and normally represent a small percentage of the total number of polypeptides of each protein.
These amino acid misincorporations have been looked as inevitable errors of the protein synthesis machinery that are effectively eliminated by the proteome quality control systems (2). However, the alternative hypothesis of adaptive translation, where polypeptide variants produced by codon mistranslations increase the functional repertoire of a proteome or facilitate the rewiring or amplification of signaling networks, enhancing the organisms capacity to respond and adapt quickly to environmental changes, is gaining increasing support (3, 4).
Recent works show that mistranslation is widespread in nature and that both single cell and multicellular organisms can take advantage of it, by regulating its levels, under specific physiological and environmental conditions (5, 6). For example, inactivation of the threonyl-tRNA synthase editing domain by oxidative stress or point mutations can lead to threonine-to-serine misincorporation in Escherichia coli and Mycoplasma spp (7, 8). E. coli exposed to aminoglycoside antibiotics, which interfere with the ribosome’s proofreading mechanism, or grown under amino acid starvation conditions mistranslates at high level (9). The archaeon Aeropyrum pernix, which grows optimally at 90 °C, undergoes inducible leucine-to-methionine substitution when incubated at 75 °C, leading to enhanced enzymatic activity of hyperthermophilic proteins, likely through an increase in the conformational flexibility required for protein function (10). Mycobacterium species exhibit increased mistranslation in response to mutations in the GatCAB enzyme complex, increasing their resistance to rifampicin (11). Remarkably, GatCAB natural mutations were found in clinical isolates from Mycobacterium tuberculosis, providing a mechanism that enhances the microorganism’s survival during infection (12). Mistranslation could further increase the sampling of genetic mutations by increasing tolerance to stressful agents. The artificial induction or suppression of generalized mistranslation levels affects E. coli’s early survival and resistance to DNA damage caused by ciprofloxacin (13).
The phenomenon of mistranslation has been extensively studied in the leucine CUG codon, which has been reassigned to either serine or alanine or ambiguously assigned to serine and leucine in several fungal species of the so called CTG clade (14). Serine-to-leucine (Ser→Leu) ambiguous decoding has been observed in Ascoidea asiatica (15), Candida maltosa (16), Candida albicans (17), and more recently, in the halotolerant yeast Debaryomyces hansenii (18). This unique translational event creates a Leu/Ser statistical proteome, that is, a proteome where each protein is a statistical average of polypeptides containing Leu or Ser at CUG sites (19). The number of different polypeptides of each protein is determined by the expression n2, where n is the number of CUG codons in the mRNA and the power 2 refers to both Leu and Ser. In C. albicans, the ambiguous decoding of CUG codons allows for the potential generation of >1.0 × 1011 different polypeptides from its 6226 coding genes (17).
The impact of CUG dual translation in C. albicans has been analyzed in the context of the commensal-pathogen transition, host–pathogen interactions, and immune evasion. Accumulating evidence suggests that the identity of the CUG codon can rewire protein–protein interactions (20, 21) and modulate protein activity in important virulence traits (22, 23). These CUG-encoded residues are often located in conserved regions of proteins involved in biofilm formation, mating, morphogenesis, adhesion, and signal transduction. For instance, the substitution of polar serine by nonpolar leucine in C. albicans leads to a decrease in the stability and activity of the Cek1 signaling kinase (24), an increased substrate adherence by cells expressing the adhesin Als3-Leu (25) and to temperature sensitivity of the translation initiation factor (eIF)4E, which mediates mRNA binding to the ribosome (26). Despite a high fitness cost in rich medium, C. albicans tolerates high level of leucine incorporation at CUG sites (27) and hypermistranslating strains are more resistant to oxidative stress and antifungals, display high phenotypic diversity (28), have higher adherence to substrates, and are less susceptible to phagocytosis (25). These observations are relevant given the interaction of C. albicans with its human host.
In C. albicans, CUG translation is facilitated by a hybrid tRNA(CAG)Ser that contains identity elements for both seryl-tRNA synthetase (SerRS) and leucyl-tRNA synthetase (LeuRS). This results in the aminoacylation of the Ser CUG–decoding tRNA with Ser (Ser-tRNA(CAG)Ser) or Leu (Leu-tRNA(CAG)Ser) and consequent incorporation of Ser or Leu at CUG codons at the ribosome A-site. The incorporation of Leu at CUG sites is possible because the mischarged Leu-tRNA(CAG)Ser is not edited by the LeuRS editing site and is not discriminated by the translation elongation factor1 (eEF1A) (29). Under normal physiological conditions, the tRNA(CAG)Ser is mainly aminoacylated with Ser by the SerRS (appx 97%) (17).
Different reporter systems have been employed to quantify CUG mistranslation. One approach relies on the GFP marker whose stability and activity depend on Leu incorporation at codon site 201 (27). Incorporation of Ser at position 201 leads to degradation and full inactivation of GFP. However, introduction of CUG_201 by site-directed mutagenesis leads to partial recovery of fluorescence which can be correlated with the levels of mistranslation (Leu misincorporation) by microscopy or flow cytometry. Another method involves mass spectrometry (MS) of a specific protein containing a single CUG site. The use of synthetic peptides for the calibration of the MS system enables the quantification of Ser and Leu peptide variants (17). However, these approaches are limited to single CUG sites, and accurate data on Leu and Ser incorporation on a proteome wide scale have been difficult to obtain. This is due to a lack of MS sensitivity to detect the rare peptides containing amino acid misincorporations that normally appear in complex protein samples, normally below 0.1% of the WT peptides, and to statistical and computational difficulties in dealing with the high noise-to-signal ratio of the MS datasets. However, recent works provide some evidence that these difficulties could be overcome and raise the hope that it may be possible to detect rare peptide variants in complex mixtures of peptides produced from total protein extracts and to identify isoforms that are absent from reference protein databases (30). In this study, we developed a pipeline based on PEAKS software algorithms (www.bioinfor.com, Bioinformatics Solutions Inc) (31) to detect and determine the frequency of low level amino acid misincorporations in complex peptide mixtures prepared from total protein extracts of C. albicans.
The study demonstrates that our MS-based method is sufficiently robust to identify rare mutant peptides and determine translation error frequencies in complex peptide mixtures produced from total protein extracts of C. albicans. We have validated our MS data using bioengineered strains of C. albicans whose leucine misincorporation at serine CUG sites was 20%, as determined by our GFP gain-of-function reporter system and protein single-site MS analysis (17, 27). The data show that CUG is not the only ambiguous codon in C. albicans and that the engineering of high levels of leucine misincorporation at CUG sites does not alter the translational accuracy of other codons. Our findings also indicate that further developments in protein sample preparation, protein fractionation, MS sensitivity, and computational data analysis are required to overcome the high noise-to-signal ratio associated with rare peptides and to accurately quantify translational amino acid misincorporations at the global proteome level. Finally, this work provides a promising opportunity to detect and determine the frequency of protein biosynthesis errors in different WT clinical strains of C. albicans and of other pathogens, as well as in other cell types, leading to a better understanding of the role of translational errors in microbial infections and diseases in general.
Experimental Procedures
Experimental Design and Statistical Rationale
The dataset reported in this study comprises 32 .raw files from two C. albicans strains with different levels of leucine misincorporation (T0 and T1 samples). Total protein extracts from each strain, grown in standard conditions, were separated by SDS-PAGE into eight fractions which were separately digested and injected in a QExactive Orbitrap (Thermo Fisher Scientific) mass spectrometer. Exclusion lists were generated for each fraction, corresponding to the most abundant peptides, and samples were re-injected for a second MS/MS run. A single biological experiment was initially used to produce a test dataset to develop and optimize the pipeline, and multiple publicly available datasets and various bioinformatics approaches were used to validate the workflow. The strains used to generate the dataset were previously characterized and their levels of leucine misincorporation at CUG sites were quantified using a gain-of-function fluorescent reporter system (27). The data was analyzed with PEAKS software (Bioinformatics Solutions Inc) to achieve peptide and protein identifications, differential protein expression, and determine codons’ error frequencies. Both peptides and proteins were filtered at 1% false discovery rate (FDR) which was estimated using a decoy-fusion approach enabled during PEAKS DB search (32). Only peptides with peptide-spectrum-matches (PSMs) above the −10lgP score threshold (related to the peptide FDR) are listed, and these filtered peptides are used as supporting peptides to infer protein identifications. This corresponds to an FDR at the PSM level of 0.4% when the default database is used and 0.5% for mistranslation analysis with the diploid database, for both T0 and T1 strains (Table 3 from Summary view available in the PEAKS HTLM report included in the Supplemental data). Label-free quantification was performed by PEAKS Q module and a significance score threshold of 20 (significance testing p-value of 0.01) was selected. PEAKS Q significance testing method was used to compare proteins from two groups as each group contained only one sample. PEAKS Q is similar to the Significance B method employed by Cox and Mann (33) and considers support peptide similarity. Bioinformatic analysis was performed in R.
Whole-Cell Lysis and Protein Extraction
C. albicans cells from strains T0 (WT) and T1 (27) were grown at 30 °C in YPD medium (2% glucose, 2% peptone, 1% yeast extract) for 16 h or 22 h, respectively, diluted at OD 0.1 in fresh medium and incubated at 30 °C until OD 1.0 to 1.5. Twenty milliliters of yeast suspension were harvested by centrifugation, washed 3× with PBS (137 mM NaCl, 2.7 mM KCl, 10 mM KH2PO4, 1.8 mM NaH2PO4, pH 7.4), and resuspended in lysis buffer (50 mM PBS pH 7.4, 1 mM EDTA, 5% glycerol, 1 mM PMSF and protease inhibitors from Roche). Cell lysis and protein extraction was achieved through mechanical lysis using glass beads and Precellys24 Homogenizer (3 cycles of 30 s at 5500 rpm, alternating with 2 min on ice). Protein quantification was carried out using Pierce BCA Protein Assay Kit.
Sample Preparation and Mass Spectrometry
Fifty micrograms of extracted proteins were resolved on 12% SDS-PAGE. The gel was fixed (40% methanol, 10% acetic acid) for 30 min, stained with Coomassie blue for 1h at room temperature, and destained with 25% methanol to remove the excess of dye. Protein bands were manually excised from the gel and transferred to Eppendorf tubes. Each sample/lane was divided into eight fractions for separate enzymatic digestion and MS injection. Gel pieces were sequentially washed with 25 mM ammonium bicarbonate (AMBIC), then with 50% acetonitrile (ACN) in 25 mM AMBIC (as many times as needed to remove the dye), and once with ACN (30 min each wash). Cysteine residues were reduced with 10 mM DTT in 25 mM AMBIC (45 min at 56 °C) and alkylated with 55 mM iodoacetamide in 25 mM AMBIC (30 min at RT in the dark). A new round of washes (25 mM AMBIC, 50% ACN in 25 mM AMBIC for 15 min, and ACN for 10 min) was performed, and gel pieces were dried and later rehydrated in the digestion buffer. Proteins were digested as recommended by Shevchenko et al. (34) with few modifications. Trypsin was added at an enzyme-to-substrate ratio of 1:50 (w/w) in 50 mM AMBIC. After 45 min on ice, the excess of buffer was removed, 50 mM AMBIC were added to cover the gel pieces, and the samples were incubated overnight at 37 °C. Extraction of tryptic peptides was achieved by washing once with 5% formic acid (FA) and twice with 5% FA in 50% ACN (20 min each incubation). Samples were dried and resuspended in 1% FA. Each fraction of tryptic peptides was injected and separately analyzed on an Ultimate 3000 HPLC system coupled to a QExactive Orbitrap (Thermo Fisher Scientific). The trap (5 mm × 300 μm) and the EASY-spray analytical (150 mm × 75 μm) columns used were C18 Pepmap100 (Dionex, LC Packings) with a particle size of 3 μm. Peptides were trapped at 30 μl/min in 96% solvent A (0.1% FA). Elution was achieved with the solvent B (0.1% FA/80% ACN (v/v)) at 300 nl/min. The 92 min gradient used was as follows: 0 to 3 min, 96% solvent A; 3 to 70 min, 4 to 25% solvent B; 70 to 90 min, 25 to 40% solvent B; 90 to 92 min, 90% solvent B; 90 to 100 min, 90% solvent B; 101 to 120 min, 96% solvent A. The mass spectrometer was operated at 1.7 kV in the data-dependent acquisition (DDA) mode. A MS2 method was used with a FT survey scan from 400 to 1600 m/z (resolution 70,000; automatic gain control target 1E6). The 10 most intense peaks were subjected to higher-energy collision dissociation fragmentation (resolution 17,500; automatic gain control target 5E4, normalized collision energy 28%, max. injection time 100 ms, dynamic exclusion 35 s).
To analyze the impact of using exclusion lists between additional runs of MS/MS acquisition, raw data from the first MS/MS injection was analyzed with Thermo Scientific Proteome Discoverer Software (https://www.thermofisher.com/pt/en/home/industrial/mass-spectrometry/liquid-chromatography-mass-spectrometry-lc-ms/lc-ms-software/multi-omics-data-analysis/proteome-discoverer-software.html) and lists were generated (one for each fraction) comprising the most abundant peptides. The same sample was reinjected in the spectrometer using these exclusion lists (limited to 5000 peptides per run) to avoid analysis of redundant peptides.
Mass Spectrometry Data Analysis
Mass spectrometry raw data was processed by PEAKS Studio XPro software. For analysis, raw data from different fractions of the same sample were combined. Carbamidomethylation was set as a fixed modification, and oxidation of methionine (M), deamidation (NQ), and protein N-acetylation were set as variable modifications. Searches were performed using mass tolerances of 5 ppm for precursor ions and 0.02 Da for fragment ions. Trypsin digestion was set as specific with up to two missed cleavages allowed. Samples’ raw data was initially searched against a haploid reference proteome database (6207 entries), obtained from the C. albicans diploid genome assembly 22 (35, 36) and available at Candida Genome Database (CGD), where a single allele represents each pair in the diploid genome: C_albicans_SC5314_version_A22-s07-m01-r149_default_protein.fasta, plus a list of common proteomics contaminants obtained from the common Repository of Adventitious Proteins (37). The significance of protein/peptide identifications was controlled to 1% FDR at both peptide and protein level. The list of peptides and identified proteins (protein-peptides and proteins files – Supplementary data) was used to evaluate proteome and protein coverage and analyze the number and frequency of each codon type assigned to the identified proteins and peptides in the samples. To assess C. albicans global mistranslation, the same raw data was searched against the diploid reference proteome database available also at CGD, which contains sequence information from both haplotypes A and B: C_albicans_SC5314_version_A22-s07-m01-r149_orf_trans_all.fasta, plus the common Repository of Adventitious Proteins list. To reduce database size and redundancy, we removed the haplotype B from protein duplicates with identical amino acid sequences and 20 entries from proteins encoded by mitochondrial genes (Supplemental Fig. S3). Unspecified posttranslation modifications (PTMs) were uncovered by PEAKS PTM algorithm in which all PTMs were selected except those marked as “Isotopic label” or “Chemical derivatives.” The remaining unmatched spectra were analyzed by SPIDER algorithm to identify amino acid substitutions. The significance of protein/peptide identifications was controlled to 1% FDR at both peptide and protein levels. PTMs’ localization accuracy was supported by AScore, meaning that the probability that the modification occurred at the reported position compared to other possible positions was at least 0.01. To confirm amino acid substitutions, a pair of b or y ions was found with at least 5% relative intensity showing fragmentation before and after the modified amino acid. Thus, peptides with at least one PTM above the Ascore (depicted in the PTM column at protein-peptides file – Supplementary data) and peptides with at least one substitution above 5% ion intensity (selected from the peptides list containing only mutations found by SPIDER – Supplementary data) were validated and maintained in the dataset. Peptides with insertions and deletions were also validated by the ion intensity filter. Finally, peptides with robust amino acid substitutions were further filtered according to the existence of their unmodified counterpart, using R scripts. Redundant peptide sequences due to allelic duplicates or paralogs (with same m/z and retention time) were also removed from the final dataset.
Error Frequency Determination
C. albicans genome sequence was obtained from CGD: C_albicans_SC5314_version_A22-s07-m01-r149_orf_coding.fasta which matched the diploid proteome database used for reference. Codons were assigned to amino acids from all validated peptides identified by the different algorithms from PEAKS software, using R scripts. The frequency of all codons was assessed, and the mistranslation frequency was calculated for each specific amino acid substitution/codon pair and for each codon independently on the destination amino acid, using the frequency of mutated codons per the frequency of all identified codons present in the final dataset. For example, for CUG codon: [frequency of CUG codons translated to other amino acids besides serine/frequency of all detected CUG codons in the sample, independently on the incorporated amino acid] x 100. For error frequency assessment, we disregarded substitutions indistinguishable from PTMs or artefacts in the Unimod database {http://www.unimod.org}: Ala→Ile/Leu (tri-Methylation); Asp→Glu (Methylation); Phe→Tyr (Oxidation); Asn→Asp/Gln (Deamidation/Methylation); Pro→Glu (di-Oxidation); Gln→Glu (Deamidation); Ser→Asp/Glu/Thr (Formylation/Acetylation/Methylation); Thr→Glu (Formylation).
Proteome Coverage and Codon Frequency Analysis
Proteome and protein sequence coverage and peptide/protein abundance information were obtained from proteins file (Supplementary data) exported after PEAKS DB search against the “haploid” C. albicans database. The area under the curve of the peptide feature found at the same m/z and retention time as the MS/MS scan (specified in Area column) was used as an indicator of peptide and protein abundance. The list of identified peptides (protein-peptides file – Supplementary data) was filtered according to the parameters described before and assigned to their codons using the matching genome sequence available at CGD: C_albicans_SC5314_version_A22-s07-m01-r149_default_coding. Global codon frequency and relative synonymous codon usage (RSCU) were assessed for all validated peptides, after PEAKS DB search against the “haploid” C. albicans database, using R scripts. For this, peptide variants (such as identical peptide sequences but with different PTMs or in different locations) and peptide duplicates due to multiple protein matching were excluded from the dataset. The RSCU from all identified proteins was obtained by Anaconda software using the nonstandard translation table for alternative yeast codon usage (38). Protein CUG content was obtained from CGD, and gene ontology analysis was performed using the GO Term Finder at the CGD website. Venn diagrams were obtained with the FunRich 3.1.3 program.
Label-free Quantification
An ID-directed label-free quantification provided by PEAKS Q module in PEAKSXpro software was used to compare the abundance of proteins from T0 and T1 strains. This approach is based on the relative abundance (MS1 feature area) of all identified peptide features detected in the two samples. Data were normalized to the total ion current and filtered according to quality (≥10), significance (≥20), and fold change (≥2). A volcano plot was obtained by PEAKS software and information regarding protein area (top-3 peptides) and the ratio between T0 and T1 was extracted from proteins.csv file.
Flow Cytometry
T0 and T1 strains expressing different variants of GFP (27) were grown at 30 °C in YPD medium for 16h or 22h, respectively, diluted at OD 0.15 in fresh medium and incubated at 30 °C until OD 1.0 to 1.5. Cells were recovered in PBS at 1:10 dilution, washed once, and filtered (0.45 μm) before being analyzed in a BD Accuri C6 Flow Cytometer. Fluorescence profiles (FL1) of 20,000 events for each sample were collected only on those cells appearing in R2 to exclude agglutinates and debris. The fluorescence intensity of each GFP variant was analyzed: GFP-Leu201 (positive control); GFP-Ser201 (negative control); GFP-Ser/Leu201 (CUG reporter). The assay was replicated three times, and Leucine incorporation was calculated for each one according to the formula: [(GFP-Ser/Leu201) – (GFP-Ser201)]/[(GFP-Leu201) – (GFP-Ser201)].
Results
Tools to Detect Amino Acid Misincorporations in C. albicans
Protein biosynthesis errors occur with estimated rates of 10−5 to 10−4 errors per incorporated amino acid in yeast (39), which is much higher than the mutation rates of DNA replication (10−9 to 10−10) (40, 41) and transcription (10−6) (42). Translational errors at frequencies of 1 misincorporation in 10,000 to 100,000 codons translated by the ribosome produce extremely low abundance peptides whose detection and identification are highly challenging. To overcome such limitations, misincorporations have been detected at single preselected amino acid sites by monitoring misincorporation of radioactive amino acids or by the recovery of function of mutant fluorescent or chemiluminescent reporter systems, where gain-of-function depends on misincorporation of the WT amino acid at the preselected sites (43).These methods are sensitive and accurate but are restricted to single codon positions of specific reporter proteins and do not provide a global view of amino acid misincorporations, which is critical to fully understand the biology of protein biosynthesis errors. Recent works carried out by our group and others have attempted to resolve this issue by analyzing amino acid misincorporations in highly complex mixtures of total protein extracts from bacteria, yeast, and mammalian cells, by MS (4, 9, 44). These studies produced promising data but also highlighted the need to develop new sample preparation and peptide fractionation methods as well as more robust computational methods to search the MS data space. In this study, we present a pipeline for the identification of amino acid misincorporations and for the determination of codon-associated error frequency at a global proteomic scale in the pathogenic fungus C. albicans by MS.
C. albicans total protein extracts were prepared using current methods and proteins were fractionated by gel electrophoresis into eight fractions. Proteins were digested with trypsin and the resulting peptide mixtures of each fraction were subjected to downstream LC-MS/MS analysis. Data analysis was carried out using the PEAKS DB algorithm, which integrates database search with de novo sequencing, to identify peptides and proteins. The haploid complement of features of the diploid Genome Assembly 22 available at the CGD was used as a reference database for this analysis.
To evaluate the quality and representativeness of the samples, we assigned each amino acid present in the final list of identified peptides to its corresponding codon and determined its overall frequency in the sample. The frequency pattern of abundant/rare codons and the RSCU of the WT sample (T0 strain) were similar to those of the reference database (Fig. 1 and Supplemental Fig. S1), indicating that the detected peptides accurately represented the global codon usage. However, some deviations were observed in rare codons. For example, 14 out of the 16 codons with a frequency in the reference genome below 6.5 codons per thousand (half of the codons’ frequency median) were underrepresented in the T0 strain sample, and the most underrepresented codons were rare codons. We then focused our analysis on the CUG codon, which is associated with ambiguous Ser/Leu translation in C. albicans. The CUG codon has a genomic frequency of 4.2 per thousand codons (0.33 RSCU), but the data of our samples showed a lower frequency of 1.8 per thousand codons (0.17 RSCU), that is, the MS approach used in this work misses peptides containing rarely used codons. Indeed, proteins encoded by genes containing rare codons are normally expressed at low level, as indicated by the codon adaptation index (17). Reassuringly, the peptides that were predominantly detected corresponded to abundant proteins that contained few CUG codons (average of 0.14 CUGs per proteins with area >1.0 × 108) while peptides of proteins with low associated total areas, where the average of CUG frequency was 3.15, were rarely detected (Fig. 2, A and B). In the specific case of the CUG peptides detection, sensitivity may have also been affected by a highly biased accumulation of these codons in genes encoding membrane and cell wall proteins, whose hydrophobicity and low solubility decreases their abundance in total protein extracts (22, 45). Accordingly, we observed an underrepresentation of membrane proteins in our MS sample, which contained fewer CUG proteins than the reference genomic database and a positive bias towards proteins containing a lower number of CUG codons (Fig. 2, C and D). These technical issues likely decreased the ability to detect amino acid substitutions at rare codon sites in our samples. However, we were still able to identify 54% of the C. albicans proteome, with most proteins (>60%) identified by more than four peptides and an average amino acid sequence coverage of 31% (Supplemental Fig. S2). We used this raw dataset to analyze the profile of amino acid substitutions in C. albicans by uncovering peptide variants with mass shifts relative to the genome-encoded sequences.
Fig. 1.
Codon frequency analysis. Dots in the graphic compare the relative synonymous codon usage (RSCU) for the reference proteome used for PEAKS DB search and for the T0 strain sample after codon assignment to all identified peptides (without duplicates). The columns highlight the RSCU deviations of the T0 strain sample regarding the RSCU obtained for the reference DB, showing their RSCU ratio. The dashed line indicates the expected ratio if no differences were observed. Arrows point to rare codons, defined as those whose frequency is lower than half of the median of all codon frequencies obtained for the DB used for reference and calculated by Anaconda software (6.5 codons per thousand). The box shows the total number of codons used for this analysis.
Fig. 2.
CUG codons are underrepresented in the peptides of WT MS sample.A, link between protein abundance and CUG content. B, link between protein abundance and number of identified peptides. C, Venn diagram and gene ontology of unidentified proteins. Proteins in both fractions have been analyzed regarding their CUG content. Venn tool available in the program FunRich 3.1.3. GO enrichment analysis of proteins absent from the MS sample, using GOTermFinder application at CGD (background: 6206 proteins from MS reference database). p-value cut-off: 0.05. ∗∗p < 0.01; ∗∗∗p < 0.001. D, analysis of identified proteins regarding their content on CUG codons.
An additional difficulty was related to the diploid and heterozygous nature of the C. albicans genome, which complicated the differentiation of translational errors from alternative amino acid incorporations associated with genomic allelic variations. To circumvent this issue, we used the diploid reference proteome available at CGD (Supplemental Fig. S3), which contains the translated sequences of both haplotypes (allele A and allele B), as a reference database for PSMs. To reduce redundancy and maintain sensitivity, we removed from the C. albicans diploid database all sequences from allele B when allele A of the same protein had an identical amino acid sequence. For spectra that did not match the pre-selected database during the de novo–assisted database search, the PEAKS software contains algorithms that detect PTMs (PEAKS PTM) and mutations (SPIDER). We used the pipeline described in Figure 3, assuming that mutated peptides (mistranslated) are less abundant than their unmutated counterparts and that PTMs with isobaric delta masses can better explain peptide mass shifts than amino acid misincorporations.
Fig. 3.
Workflow for mistranslation analysis. Protein extracts from actively growing cultures were resolved in SDS-PAGE and divided into eight fractions. Bands were manually excised from the gel, digested, and injected separately in the MS system. Raw MS/MS data was analyzed by PEAKS Xpro software. The list of identified peptides was filtered to remove duplicates and low-quality hits, and codon assignment and mistranslation frequency was calculated using R scripts. Numbers indicate the number of peptides identified or validated in each step of the pipeline from T0 (WT) strain.
We analyzed and filtered the list of all identified peptides (protein-peptides file – Supplementary data) from a C. albicans WT strain (T0 strain) using R scripts (see Experimental Procedures). Peptides with low-quality PTMs that did not pass the AScore filter were removed from the list, as well as peptide duplicates with the same amino acid sequence, matched to the same protein, but to a different allele. For correct codon-amino acid assignments, we took into consideration that peptides of identical sequences could be aligned to different proteins due to the presence of paralogs and conserved regions among protein families. To minimize overestimation of codon frequencies, a single protein per peptide was chosen for codon assignment and we also developed our scripts to record any potential bias in codon assignment. Additionally, for peptides that aligned to multiple regions of a protein due to the presence of repetitive regions, the program assumed the position closer to the N-terminal of the protein. Finally, to validate peptides with amino acid misincorporations, we used an ion intensity filter to localize the mutated amino acid (≥5%) with high accuracy and accepted only those peptides whose unmutated counterpart (WT peptide, correctly translated) was also detected in the sample.
Protein Biosynthesis Errors Occur at Different Codons
To evaluate global mistranslation, we utilized the SPIDER algorithm of the PEAKS software, which reconstructs new peptide variant sequences by combining homology search and de novo sequence tags (31). After applying database search and validation filters, the algorithm identified several amino acid substitutions in different codons of 540 peptides. Before assuming these variants as true mutations, we have considered the possibility that PTMs and mutations could produce identical mass shifts and that PTMs are more frequent than amino acid misincorporations (46). For instance, a mass shift of 14.02 Da can result from either methylation or Gly to Ala, Asp to Glu, Val to Ile/Leu, Ser to Thr, or Asn to Gln substitutions. To avoid inaccurate assignment of amino acid misincorporations, which would overestimate mistranslation events, we excluded any substitution whose associated mass shift could be explained by a PTM. For this, the PEAKS PTM identification algorithm was applied to the data upstream of the SPIDER algorithm and as expected, the number of misincorporations decreased (Fig. 4). By following this workflow and after removing redundant peptides due to multiple protein alignments, 94 unique mutated peptides were detected out of 48,474 peptides identified in the WT strain (T0). These mutations occurred in 27 different codons (44 substitution types), and the most frequent substitutions detected were alanine-to-glutamine (AlaGCU/GCC→Gln) and glycine-to-asparagine (GlyGGU→Asn). We also observed that serine codons were prone to aspartic acid misincorporation, although serine formylation can occur as an artefact (47), resulting in the same mass shift as Ser→Asp substitution.
Fig. 4.
Impact of using PEAKS PTM algorithm before SPIDER in the identification of amino acid substitutions. The matrix reveals the number and identity of substitutions found for each codon. Different posttranslational modifications with isobaric delta masses may explain the observed differences. Highlighted in blue boxes are amino acid substitutions in the CUG codon and Ser→Leu substitutions in serine codons.
In line with previous reports (48, 49), we observed that the codons prone to mistranslation were not necessarily the most frequently used in the peptides of our sample. Indeed, highly frequent codons like GAU, AAU, and CAA showed high translational accuracy, indicating that biosynthesis errors are not correlated with codon usage frequency (Supplemental Fig. S4). As expected, we were only able to detect two mistranslation events in rare codons, specifically HisCAC→Gln and SerCUG→Leu. The Ser/Leu ambiguous translation of CUG codons in C. albicans has been previously described and quantified using fluorescent reporters and MS/MS of recombinant reporter proteins (17, 27). Our new data confirmed this mistranslation event, and we observed no other substitutions at the CUG codon, even when PEAKS PTM was not used. Furthermore, the data indicates that this substitution only occurs at the CUG codon and not at any other of the six serine codons (Fig. 4).
The CUG Codon is Mistranslated at High Frequency in C. albicans
To gain a comprehensive understanding of the distribution of translational errors in the C. albicans proteome, each amino acid of the identified and validated peptides was assigned to the corresponding codon. Translational accuracy was assessed by calculating error frequency, defined as the number of times a specific codon is mistranslated relative to the total number of times it appears in the MS dataset. The global error frequency for the WT sample (T0 strain) was 1.08 × 10-4 (0.01%), considering the median of all codon-associated mistranslation frequencies (Fig. 5A). While GCC, GCU, and GGU codons were found to be associated with the most frequent amino acid misincorporations, their calculated error frequency of 0.08%, 0.06%, and 0.03%, respectively, was lower than the error frequency for the rare CUG codon (0.17%). We detected more than one amino acid misincorporated at both GCC and GCU codon sites, but Ser→Leu were the only substitutions detected at CUG sites (Fig. 5B). Thus, the CUG codon had the highest mistranslation frequency of all amino acid-codon pairs, (Fig. 5A, left side) and this result was maintained even when the error frequency associated with each codon independently of the misincorporated amino acid was computed (Fig. 5A, right side).
Fig. 5.
Comparative analysis of protein biosynthesis errors in strains with different levels of Leu misincorporation at CUG codons. Mistranslation frequency was assessed by LC-MS/MS analysis of whole cell lysates (A and B) and by flow cytometry to detect a fluorescent reporter (C). Error bars from flow cytometry data represent the SD of the mean from three independent experiments. A, distribution of mistranslation frequency calculated for each codon mutation either specific for one single amino-acid (left) or for all detected substitutions (right). The blue line indicates the median of all calculated mistranslation frequencies. CUG codon mistranslation frequency is highlighted in red. B, comparison between codons mistranslation frequencies of T0 and T1 strains. Graphics at the bottom detail which amino acids are being misincorporated and at what frequency in codons whose mistranslation frequency is closer to the one found for CUG in the WT strain (>0.08%).
Previous studies have quantified the relative levels of Leu and Ser incorporated at single CUG sites using mass spectrometry analysis of purified recombinant proteins (17) and fluorescence microscopy of strains transformed with fluorescent reporters (27). In the present study we provide a global overview of the frequency of CUG sites where Leu is incorporated relative to the total number of CUGs present in a large set of C. albicans proteins. Therefore, the error frequency determined herein is complementary to the codon specific error quantification previously obtained for C. albicans (Fig. 5C). Moreover, to validate our pipeline, we analyzed a Leu-CUG hypermistranslating strain (T1) engineered in our laboratory which misincorporates 20.61 ± 1.81% of Leu at the Ser-CUG sites, as quantified by fluorescence microscopy (27). Consistent with the higher level of CUG-Leu incorporation, we detected 98 Ser→Leu substitutions in peptides derived from 84 proteins of the T1 strain, contrasting with only two peptides in two proteins of the T0 control strain. Importantly, almost all CUG-related Leu misincorporations in the T1 strain occurred in proteins of the T0 strain where we could detect the WT CUG peptides (Ser-CUG incorporation only). This provides a clear indication that proteome scale detection of rare peptides containing amino acid misincorporations requires the development of more sensitive sample preparation, MS/MS, and computational methods.
To rule out the remote possibility that misincorporations at CUG sites resulted from genetic variation related to genetic manipulation of the T0 and T1 strains, we reanalyzed the genome sequencing data of these strains (27), focusing on the genes that encoded the peptides containing amino acid misincorporations in both strains. No allelic difference was detected at those specific sites (Supplemental Table S1), confirming that the mutations were true translational amino acid misincorporations. Moreover, a stringent analysis of putative error arising from codon misassignment related to alignment of peptides to more than one protein was also carried out. Four mutated peptides matched different proteins in the T0 strain, but only one mutated peptide (K.NQQ(sub A)AMNPANTVFDQ(sub A)K.R) with equal area, mass, and retention time matched with the same probability proteins that are genetically distinct, C1_13480W_A and C1_04300C_A. The second Q(sub A) could be attributed to either GCC or GCU codons, leading to potential alterations in codon frequencies. In the T1 strain, eight mutated peptides with different protein matches were identified, four of them at genetically different chromosomal loci: two involving GCC/GCU mismatches (with no change in the error frequency), and the other two peptides whose mismatches were assigned to CAC codons instead of CAU codons, leading to a change in the codon error frequency from 0.07% and 0.02%, respectively, to 0.04% for both. None of the identified CUG-containing peptides in either T0 or T1 were matched to more than one protein. The estimated CUG mistranslation frequency in the T1 strain (10.94%) was lower than that reported previously but similar to the in vivo values obtained by flow cytometry (12.74 ± 1.20%). By using this pipeline, we were thus able to discriminate two C. albicans strains with different CUG-related error frequencies.
The Artificial Increase of Leucine Incorporation is CUG-Specific
C. albicans CUG-hypermistranslating strains exhibit a high degree of phenotypic and genomic diversity, which impacts fungal adhesion to host substrates, immune recognition, and tolerance to antifungals (50). However, it is unclear whether these adaptive phenotypes result from Ser/Leu substitutions at CUG sites or from general deregulation of translational fidelity caused by the stress induced by Ser→Leu CUG mistranslation and/or by the heterologous expression of the Saccharomyces cerevisae tRNA(CAG)Leu and the GFP reporter. Our new data show a 3-fold increase in the number of mutated peptides in the T1 strain relative to the WT T0 strain (in a total sample of 48,474 peptides in T0 versus 50,186 peptides in T1), with an overall mistranslation frequency of 1.95 × 10-4 (0.02%) in T1 versus 1.08 × 10-4 (0.01%) in T0 (Fig. 5A). The analysis of the identity and frequency of substitutions confirms a bias towards mistranslation of CUG codons and the specificity of Leu incorporation at CUG sites. Indeed, Leu was not incorporated at any other of the six Ser codons (Fig. 5B and Supplemental Fig. S5), and the translational accuracy of other codons was not affected. Therefore, these data confirm for the first time that the phenotypes previously observed in the T1 strain (27, 51) are a direct result of increased Leu/Ser ambiguity at CUG sites.
Interestingly, the total number of CUG codons assigned to the T1 peptides was lower relative to the number detected in the control T0 strain (897 in T1 vs 1197 in T0). This, combined with the higher number of mistranslated codons (98 in T1 vs 2 in T0), explained the increase in misincorporation frequency from 0.17% in T0 to 10.93% in T1. To further clarify the result above, we estimated the RSCU in T0 and T1 strains using the detected peptides as described in the methods section. Unlike other codons that showed similar RSCU values in both strains, CUG codons were underrepresented in T1 (Fig. 6A), suggesting that Leu misincorporation at certain CUG sites may lead to protein misfolding and degradation, as previously analyzed in our laboratory (22). However, when analyzing the RSCU using the full sequence of all identified proteins, no differences were observed. The percentage of identified proteins with CUG-encoded residues was similar between T0 (60.5% with an average of 2.81 CUG codons per protein) and T1 (59.7%, average of 2.84). The major difference was the detection of peptides containing CUG-encoded residues, which were identified in 34.7% of the T0 proteins and only in 24.0% of the T1 proteins. Apart from this, there were no major differences found regarding proteome coverage, overall protein sequence coverage, and the number of identified peptides per protein (Supplemental Fig. S2). In line with this, the higher level of Leu incorporation at CUG sites in the T1 strain remodeled the proteome: 263 and 354 proteins were exclusive of T0 and T1, respectively. A GO enrichment analysis showed that T0-specific proteins were associated with ubiquitination and autophagy, and more than 80% of these proteins contained at least one CUG codon (Fig. 6B). Interestingly, the function of 40% of the T1-specific proteins is “unknown.” We used the PEAKS Q module from PEAKS Xpro software to perform a label-free quantification on the proteins common to both strains and found that 680 were downregulated and 247 were upregulated in the hypermistranslating strain relative to the T0 control (Fig. 6C). Although there were only slight differences in the total abundance of common proteins (Fig. 6D), we found that 62% of the downregulated proteins contained CUG codons, suggesting that the diminished abundance of these proteins in the T1 strain may impede the detection of some of their peptides.
Fig. 6.
Comparative proteogenomic analysis between T0 and T1 strains.A, codon frequency analysis. RSCU deviations of the T1 sample regarding the RSCU obtained for the WT (T0) sample, calculated from all detected peptides or from all identified proteins using Anaconda software. The dashed line indicates the expected ratio if no differences were observed. B, Venn diagram and gene ontology of T0-specific proteins. Common and exclusive proteins identified in T0 and T1 strains were analyzed regarding their CUG content. Venn tool available in the program FunRich 3.1.3. GO enrichment analysis of proteins absent from T1 sample, using GOTermFinder application at CGD (background: 3680 proteins identified in both strains). p-value cut-off: 0.05. ∗∗p < 0.01; ∗∗∗p < 0.001. C, volcano plot from label-free quantification of T1 relative to T0 strain obtained by PEAKS Q module. Markers for the proteins that are above the set significance threshold are displayed in red (for upregulated proteins regarding the reference strain T0) and green (for downregulated proteins). Up- and down-regulated proteins were separately analyzed regarding their CUG content. D, comparison between the abundance of differentially expressed proteins and their content on CUG codons.
Peptide Exclusion Lists as a Strategy to Increase the Detection of Mistranslated Peptides
To optimize our methodology and increase the detection of mistranslated peptides, we tested different approaches to acquire and analyze the MS/MS data. We lessened the stringency of our data analysis filters, namely by using a mass tolerance of 10 ppm instead of 5 ppm for precursor ions during searches, allowing a semi-specific cleavage instead of trypsin-specific cleavage and using a 1% ion intensity filter instead of 5% to validate amino acid substitutions. These variables increased slightly the number of detected peptides containing mutations but did not substantially alter the mistranslation frequencies, and the CUG codon remained the most ambiguous codon in C. albicans (Supplemental Fig. S6A).
We also investigated whether a more direct search for Leu misincorporation, by selecting Ser→Leu substitutions as a variable PTM when performing the DB search, would increase the number of peptides with Leu/Ser mutations. This strategy resulted in the identification of an additional CUG site mutation in the WT sample and four mutations associated to other, more frequent, serine codons. Out of the 76 new Ser→Leu mutations detected in the T1 strain, 70 were found to occur at CUG sites, suggesting a higher propensity for Leu misincorporation at CUG codons in this strain. The error frequency linked to the CUG codon was noticeably higher than that of other serine codons in both strains when employing this approach (Supplemental Fig. S6B).
To increase the sensitivity of our MS/MS methodology, we investigated the impact of using exclusion lists between additional runs of MS/MS acquisition, which would enable the detection of other less abundant peptides. When the samples were reinjected for a second analysis, excluding from MS/MS analysis the most abundant spectra/peptides already analyzed in the first MS/MS run, we identified 343 new proteins, detected almost 13,500 new peptides, and found 29 more substitution types (Fig. 7, A and B). The global mistranslation frequency did not change considerably, decreasing from 1.08 × 10−4 (0.011%) to 9.73 × 10−5 (0.010%). However, there was a substantial increase in alanine to glutamine (AlaGCU/GCC→Gln) substitutions, which was the most frequent substitution detected with only one MS analysis.
Fig. 7.
Impact of exclusion lists in the detection of peptide variants.A, Venn diagram showing the number of proteins and peptides identified and validated when doing the analysis with the data obtained only in the first MS injection (T0) and when using the data from both MS injections, before and after applying the exclusion lists (T0_EL). B, distribution of mistranslation frequency calculated for each codon mutation either specific for one single amino-acid (left) or for all detected substitutions (right). “n” denotes the number of different substitution types and the blue line indicates the median of all calculated mistranslation frequencies. L(S)CUG mistranslation frequency is highlighted in red and the top-3 codons with highest error frequency are labeled. C, Venn diagram showing the number of peptides identified and validated in the first MS injection (T0) and in the second MS injection after applying the exclusion lists (T0_2run). D, detailed information on peptides exclusively found in T0_2run data, correlating the protein abundance with the number of identified proteins, peptides, and peptides with amino acid substitutions.
The second MS/MS injection alone enabled the detection of 13,883 new peptides (Fig. 7C). We observed that the new peptides were mostly from highly abundant proteins that were previously identified in the first MS/MS analysis and that only 4% of these new peptides led to the identification of new proteins (Fig. 7D), suggesting that exclusion lists can improve peptide and protein identification sensitivity in the MS analysis of complex samples but do not substantially increase the number of peptides belonging to less abundant proteins. In fact, while the second MS/MS injection increased the detection of peptides with amino acid substitutions, they were all matched to proteins already identified without exclusion lists and mainly to those most abundant in the sample (area >1 × 108). Moreover, these new substitutions were found in codons previously identified as the most frequently associated with mistranslation such as GCU, GGU, and GCC. In summary, while this is a promising strategy to increase the identification of peptide variants, it may introduce bias towards mistranslation of highly abundant proteins (and frequent codons) that could compromise our approach for global error frequency calculation, while considerably increasing the cost of the assay.
Analysis of CUG-Associated Misincorporations in Publicly Available MS Data Repositories
We utilized our bioinformatics workflow to scrutinize the protein MS raw data of C. albicans that have been deposited in the Proteomics Identifications database (PRIDE public data repository). Over 80 samples (including replicates) from nine different projects were reanalyzed, and Leu-CUG misincorporations were detected in 18 samples from three projects (Supplemental Table S2). This limited Leu-CUG detection could be attributed to the different methodologies employed in sample preparation, fractionation, and in acquiring the MS datasets, which are optimized for each study's objectives, and highlights the technical difficulties in analyzing amino acid misincorporations by MS/MS. Indeed, the data from the PRIDE repository that we re-analyzed in this work did not use our fractionation methodology but employed a peptide-level fractionation approach. This implies that peptides from highly abundant proteins may become distributed across all fractions, thereby hampering the detection of peptides derived from low-expressed proteins expected to be enriched in misincorporations. Most MS-based studies focus on identifying and quantifying differentially expressed proteins in distinct strains, morphologies, or under diverse environmental conditions where the noise-to-signal ratio is normally low, ignoring the presence of peptide variants from protein isoforms that are rare in complex samples.
We identified a single peptide containing a CUG Ser→Leu substitution in six samples out of 52 samples (PXD020195), which represented several morphologies and culture conditions. These samples were obtained to generate a C. albicans spectral library for data independent acquisition (DIA) (52). We also detected this amino acid misincorporation in three replicates from a dataset (PXD031774) obtained from a mixture of five isogenic strains retrieved from a fluconazole-treated AIDS patient who suffered from recurrent oropharyngeal candidiasis (53). Finally, we re-analyzed a dataset (PXD027278) containing MS raw data from various C. albicans strains and growth forms. This dataset was originally produced to assess CUG mistranslation in various species and growth forms (54). In this work, Mühlhausen et al. refute the idea that C. albicans misincorporates Leu at CUG sites and argue against its prevalence as a mistranslation event. However, by using our bioinformatics pipeline, we have identified Leu-CUG incorporation events in several samples. Amino acid substitutions were also observed in other codons, but the frequency of CUG mistranslation stood as one of the highest among all samples. Additionally, we re-analyzed our own dataset (WT T0 strain) using our pipeline with the “unbiased” reference database described in the manuscript (albeit without removing sequence redundancy). This involved extracting all proteins containing CUG sites from our diploid database and modifying them to incorporate each one of the 19 amino acids at CUG positions (excluding isoleucine mutations). By employing this approach, we found that 95.8% of CUG sites were indeed translated as serine, confirming the reassignment of serine for the CUG codon in C. albicans, and the remaining 4.2% of CUG sites were assigned to other codons, predominantly threonine, aspartic acid, glutamine, and leucine/isoleucine (Supplemental Fig. S7). However, it is important to exercise caution when interpreting these results. Firstly, certain amino acid substitutions may introduce a mass shift that could potentially be attributed to PTMs, such as methylated serine, offering an alternative explanation for Ser→Thr substitution. Secondly, the database used in this approach contains numerous entries with redundant regions, which increases the search space and may diminish statistical power. Lastly, employing an altered and incomplete proteome for the database search, where proteins lacking CUG sites were removed, can potentially lead to peptide misidentification and introduce bias towards the expected mutations.
The findings presented in our study are thus in agreement with previous research conducted using various MS approaches, fluorescence, and in vitro tRNA charging methods (16, 17, 27), which collectively refute the study by Mühlhausen et al. (54). Despite the low sensitivity and high variability among datasets and even between replicates from the same dataset, we found that whenever CUG mistranslation was identified, the associated error frequency was always higher than the global error of the sample (Fig. 8). While we did observe some Ser→Leu substitutions associated with other serine codons (particularly UCA) in some samples, the error frequency was not meaningful. These results validated the specificity of Ser-Leu translation at CUG sites, as no other amino acid was detected in this codon.
Fig. 8.
Global mistranslation analysis using datasets deposited on PRIDE repository. Distribution of mistranslation frequencies calculated for each codon mutation either specific for one single amino-acid (left) or for all detected substitutions (right). The blue line indicates the median of all calculated mistranslation frequencies. CUG codon mistranslation frequency is highlighted in red.
In the present study, we further demonstrate that protein biosynthesis errors in multiple codons are common in C. albicans and that certain codons are more error-prone than others. Similar results were obtained in E. coli and Saccharomyces cerevisiae by Mordret et al. who developed a pipeline for the identification of amino acid substitutions based on MaxQuant algorithms and the use of spectral libraries (9). We reanalyzed the S. cerevisiae sample of these studies (six fractions of the peptide mixture separated by strong cation exchange (55)) and obtained a very similar result for the top 10 amino acid misincorporations (Supplemental Fig. S8). Although the type of amino acid misincorporations found in our study for C. albicans are substantially different, questioning the idea of a universal error pattern for mistranslation, there are many similarities with our own S. cerevisiae dataset (Supplemental Fig. S9).
These carefully controlled studies demonstrate the utility of our pipeline in the study of protein biosynthesis errors although they also highlight the impact that different sample preparation procedures and thus different datasets have on the detection of amino acid substitutions and on the identity of substitution types, as well as on other proteomic outcomes (56, 57).
Discussion
Over the past decade, the fields of proteomics and proteogenomics have seen significant improvements due to the development of more sensitive and high acquisition rate MS instruments, optimized software tools with user-friendly interfaces, and the availability of data in public repositories (58). While protein identification has been widely achieved through database search approaches requiring prior knowledge and availability of genome and/or proteome reference sequences (59), new strategies have been developed to detect novel peptides or peptide variants arising from protein isoforms absent from reference databases. To increase peptide identification, a common approach is the use of spectral libraries composed of previously observed and identified MS/MS spectra (60). Alternatively, de novo sequencing can be used to identify peptide-derived tandem mass spectra without prior knowledge of the sample or a predefined sequence database (61). While computationally intensive, de novo sequencing is unbiased and essential when no database is available. Search engines can also benefit from hybrid identification methods, combining de novo sequencing and database search to improve sensitivity and accuracy. This approach is exploited by PEAKS software which performs de novo sequencing of spectra before any database searching (32). A complete and comprehensive reference database is critical, as incomplete databases may lead to unidentified or misidentified peptides resulting in false positive identifications (62). For unmatched de novo sequences, new algorithms can identify peptide variants to increase peptide identity (31, 63).
Our goal was to detect and identify peptide mutations resulting from amino acid misincorporations in the human pathogenic fungus C. albicans, with the ambition of conducting a comprehensive analysis of protein biosynthesis errors in total protein extracts. To achieve this, we aimed to increase the sensitivity of our MS method while maintaining full codon representativeness. We opted to fractionate the samples to reduce their complexity and dynamic range using one-dimensional gel electrophoresis and obtain peptides via in-gel digestion of eight band segments for separate injection and independent sequencing analysis. This approach significantly enhances the depth of MS identification compared to other fractionation methods (64), enabling the detection of a wider range of peptide variants in complex mixtures and facilitating the removal of contaminants that could interfere with MS analysis (65). In addition, SDS-PAGE allows the visualization of proteins’ relative abundance and size and the study of virtually any protein samples, thanks to the highly effective solubilization properties of SDS. Alternative fractionation strategies at the peptide level such as high pH reversed-phase peptides fractionation or 2D LC-MS/MS offer potential enhancements in peptide detection, each with its own set of advantages and drawbacks depending on the experimental objectives and sample characteristics. However, we wondered whether protein-level fractionation would outperform peptide-level separation in the detection of low abundance peptides, because highly abundant proteins would be more effectively isolated from lower abundance proteins during protein-level fractionation. In our particular case, while mistranslation events can potentially occur in any protein during the process of translation, positions codified by rare codons present at higher frequency in low-abundance proteins are more susceptible to amino acid misincorporations (abundant proteins do not contain or contain very few rare codons). In fact, the rare CUG codon that is mistranslated in C. albicans is absent from most abundant proteins. Additionally, SDS-PAGE gel slice fractionation remains one of the most widely employed methods due to its robustness and performance on both low- and high-resolution instruments, making it suitable for numerous research groups (66).
To identify amino acid substitutions, we utilized PEAKS algorithms. We first performed a PEAKS PTM search before conducting a mutation search using SPIDER to avoid incorrectly matching of unexpected proteins or protein isoforms. We also included a list of common contaminants in our reference database and carefully selected protein isoforms resulting from allelic variations to prevent their misidentification as amino acid substitutions. However, the use of a reference database for proteogenomics has limitations, and some putative errors on codon assignment cannot be avoided. Although our strains are derived from the clinical isolate SC5314 whose sequence was used as a reference database, we cannot rule out the possibility that genetic variations may have occurred in our bioengineered strains that are not reflected at the proteome level. In addition, multiple peptide-protein matching to different proteins, different alleles, or even within the same protein, with genetic different sequences, can also be an important source of error.
In this study, we show that by analyzing the codon frequency of the peptides identified by MS/MS, we can obtain a relative quantification of codons-specific mistranslation at the proteome scale. However, the very low level of amino acid misincorporations detected remains a significant technical challenge to produce global maps of protein synthesis errors. In our analysis of C. albicans MS raw data available in the PRIDE archive, no Ser→Leu substitutions were found in most cases, even though up to three CUG-Leu incorporations were detected in some samples. This suggests that sample preparation, digestion, and fractionation for protein synthesis error detection requires important optimization. For example, simple lysis buffer devoid of detergents (making extraction of membrane proteins inefficient) or a single digestion enzyme (producing extremely short or long peptides that are not detected or produce poor quality data) can create a systematic bias and generate false negatives in the acquired datasets (46). As Leu and Ser are chemically distinct amino acids, Ser→Leu substitutions may lead to rapid protein turnover both in vivo and during sample preparation or to significant alterations in peptide behavior inside the MS system, further complicating the detection of mutant peptides. This is consistent with our observation that 62% of the downregulated proteins in T1 contain at least one CUG-encoded residue, potentially explaining why proteins in T1 are still identified, but their CUG-containing regions (peptides) are not.
The proteomic analysis conducted on a hypermistranslating strain (T1) supports the hypothesis that C. albicans tolerates very high levels of Leu incorporation at CUG sites, which is essential to explain the functional roles of Leu-incorporation on phenotypic diversity and pathogenesis (50). Our data further demonstrates the specificity of Leu-misincorporation to CUG codons and highlights the potential of MS in characterizing complex samples based on their level of protein biosynthesis errors. To optimize our pipeline, we conducted preliminary assays in which the total protein extract of T0, T1, and the background strain SC5314 was divided into four fractions (Supplemental Fig. S10). Despite the anticipated lower number of mutated peptides discovered compared to the use of eight fractions and acknowledging that both experiments are not exact replicates, the results remained consistent. Notably, a peptide containing a CUG Ser→Leu substitution (belonging to the Cdc60 protein) was identified in both T0 and SC5314, while this substitution occurred 35 times in the T1 strain. As in the eight-fraction experiment, the frequency of CUG mistranslation was the highest in all strains and there was a clear difference between the WT strains and the hypermistranslating strain. Additionally, the codons most frequently associated with amino acid substitutions were once again GCC and GCU (Ala to Gln substitution), GGU (Gly to Asn substitution), and CCA (Pro to Ala substitution).
Interestingly, one of the Leu-CUG peptides consistently detected belonged to the Cdc60 protein, which encodes the cytosolic leucyl tRNA synthetase (LeuRS). Notably, Cdc60 possesses a C-terminal domain where a CUG codon is localized, enabling it to recognize both the hybrid tRNA(CAG)Ser and its cognate tRNA(Leu) (67). Both the LeuRS-Ser and LeuRS-Leu isoforms catalyze activation and aminoacylation; however, the Leu isoform demonstrates higher activity (68). Also, SerRS isoforms, another critical component involved in CUG decoding and containing its own CUG codon, exhibit similar characteristics (22). To the best of our knowledge, this study provides the first evidence that both isoforms can coexist within live cells of C. albicans, thereby underscoring the importance of conducting a more comprehensive analysis to explore the full spectrum of CUG ambiguities.
In summary, our proteogenomics pipeline has successfully identified amino acid substitutions resulting from translational errors in the diploid fungus C. albicans, revealing varying levels of CUG mistranslation among different strains. This tool provides a valuable resource for assessing protein biosynthesis errors in various C. albicans strains, including clinical isolates, without the need of heterologous or synthetic reporters. The developed R scripts facilitate peptide filtration, codon assignment, and error frequency determination while also preserving intermediate information on duplicates and putative errors in codon assignment for further in-depth analysis if required. For a comprehensive analysis of mistranslation events in C. albicans, we recommend whole-cell proteome fractionation by gel electrophoresis (>8 fractions) for separate trypsin digestion and MS/MS analysis, the use of comprehensive cell lysis buffer to ensure sample representativeness and long LC runs. Incorporating de novo sequencing to enhance peptide detection, utilizing the diploid C. albicans database to account for allelic variations, and considering PTMs mass shifts are all relevant due to the high noise-to-signal ratio of the analysis. The proposed pipeline and the multiple analyses conducted in this study with both our own MS/MS dataset and publicly available proteomics data advance significantly our capacity to detect protein synthesis errors at the proteome level using mass spectrometry; however, they also highlight the large challenges of obtaining complete maps of protein synthesis errors and, more importantly, how far we still are from having a methodology to comprehensively quantify amino acid misincorporations at this scale. It is also crucial to address the significant variability on peptide identification among datasets and even within replicates from the same dataset. In the context of DDA, the process of selecting peptide precursors for fragmentation is semi-stochastic, intensity-based, and constrained to a predefined number of precursors. This variability results in a considerable diversity in the subset of identified peptides between samples or replicates, particularly affecting low-abundance peptides (69). To mitigate these challenges, alternative data acquisition workflows such as DIA have emerged. In DIA, all precursor ions within an m/z window are fragmented regardless of their intensity and the m/z window is systematically scanned across a mass range. It requires prior knowledge about the fragment ion spectra of targeted peptides, given by a spectral library previously generated through DDA, significantly improving the reproducibility of proteome quantification across runs and reducing the prevalence of missing values (70). However, it is important to note that DIA also comes with its own set of limitations on sensitivity and dynamic range. Despite the technical difficulties, we anticipate that expected increases in MS sensitivity, development of new sample preparation methods, and integration of artificial intelligence techniques in our pipelines will further enhance our capacity to explore in depth the translational misincorporation of amino acids into proteins (71). This is essential to better understand the biology of mistranslation.
Data Availability
The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE (72) partner repository with the dataset identifier PXD047025 and 10.6019/PXD047025. R scripts of the pipeline used for data processing are available at the GitHub repository: https://github.com/andreia-reis/proteogenomic_pipeline_calbicans and at Zenodo: https://doi.org/10.5281/zenodo.10651622.
Supplemental data
This article contains supplemental figures and tables.
Conflicts of interest
The authors declare that they have no conflicts of interest with the contents of this article.
Acknowledgments
Funding and additional information
This work was supported by FCT - Fundação para a Ciência e Tecnologia, I. P., by project reference UIDB/04501/2020 (https://doi.org/10.54499/UIDB/04501/2020) and UIDP/04501/2020 (https://doi.org/10.54499/UIDP/04501/2020). Additional support was granted by FEDER (Fundo Europeu de Desenvolvimento Regional) funds through the COMPETE 2020, Operational Programme for Competitiveness and Internationalization (POCI), and by Portuguese national funds via FCT under the projects PBE-POCI-01-0145-FEDER-031238, Varcal (2022.01376.PTDC), and FunResist (https://doi.org/10.54499/PTDC/BIA-MIC/1141/2021). This work was also supported by the Portuguese Roadmap of Research Infrastructures, under GenomePT (POCI-01–0145-FEDER-022184) and RNEM - Portuguese Mass Spectrometry Network (LISBOA-01-0145-FEDER-402-022125). I. C. is supported by national funds (OE), through FCT, I. P. (https://doi.org/10.54499/2021.00329.CEECIND/CP1659/CT0009). M. A. S. S. is supported by the European Union's Horizon 2020 research and innovation program under grant agreement no 857524.
Author contributions
I. C. writing–original draft; I. C., C. O., A. R., A. R. G., and S. A. methodology; I. C., P. D., A. R. B., R. V., G. M., and M. A. S. S. writing–review and editing; I. C., P. D., A. R. B., R. V., and M. A. S. S. conceptualization; I. C. data curation; P. D. A. R. B., G. M., and M. A. S. S. funding acquisition; P. D., G. M., and M. A. S. S. resources.
Contributor Information
Inês Correia, Email: inescorreia@ua.pt.
Manuel A.S. Santos, Email: mansilvasantos@uc.pt.
Supplementary Data
References
- 1.Alves R., Barata-Antunes C., Casal M., Brown A.J.P., Van Dijck P., Paiva S. Adapting to survive: how Candida overcomes host-imposed constraints during human colonization. PLoS Pathog. 2020;16 doi: 10.1371/journal.ppat.1008478. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Drummond D.A., Wilke C.O. The evolutionary consequences of erroneous protein synthesis. Nat. Rev. Genet. 2009;10:715–724. doi: 10.1038/nrg2662. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ribas de Pouplana L., Santos M.A., Zhu J.H., Farabaugh P.J., Javid B. Protein mistranslation: friend or foe? Trends Biochem. Sci. 2014;39:355–362. doi: 10.1016/j.tibs.2014.06.002. [DOI] [PubMed] [Google Scholar]
- 4.Mohler K., Ibba M. Translational fidelity and mistranslation in the cellular response to stress. Nat. Microbiol. 2017;2 doi: 10.1038/nmicrobiol.2017.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Ling J., O'Donoghue P., Söll D. Genetic code flexibility in microorganisms: novel mechanisms and impact on physiology. Nat. Rev. Microbiol. 2015;13:707–721. doi: 10.1038/nrmicro3568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Schwartz M.H., Pan T. Function and origin of mistranslation in distinct cellular contexts. Crit. Rev. Biochem. Mol. Biol. 2017;52:205–219. doi: 10.1080/10409238.2016.1274284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ling J., Söll D. Severe oxidative stress induces protein mistranslation through impairment of an aminoacyl-tRNA synthetase editing site. Proc. Natl. Acad. Sci. U. S. A. 2010;107:4028–4033. doi: 10.1073/pnas.1000315107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Li L., Boniecki M.T., Jaffe J.D., Imai B.S., Yau P.M., Luthey-Schulten Z.A., et al. Naturally occurring aminoacyl-tRNA synthetases editing-domain mutations that cause mistranslation in Mycoplasma parasites. Proc. Natl. Acad. Sci. U. S. A. 2011;108:9378–9383. doi: 10.1073/pnas.1016460108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Mordret E., Dahan O., Asraf O., Rak R., Yehonadav A., Barnabas G.D., et al. Systematic detection of amino acid substitutions in proteomes reveals mechanistic basis of ribosome errors and selection for translation fidelity. Mol. Cell. 2019;75:427–441.e425. doi: 10.1016/j.molcel.2019.06.041. [DOI] [PubMed] [Google Scholar]
- 10.Schwartz M.H., Pan T. Temperature dependent mistranslation in a hyperthermophile adapts proteins to lower temperatures. Nucleic Acids Res. 2016;44:294–303. doi: 10.1093/nar/gkv1379. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Javid B., Sorrentino F., Toosky M., Zheng W., Pinkham J.T., Jain N., et al. Mycobacterial mistranslation is necessary and sufficient for rifampicin phenotypic resistance. Proc. Natl. Acad. Sci. U. S. A. 2014;111:1132–1137. doi: 10.1073/pnas.1317580111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Su H.W., Zhu J.H., Li H., Cai R.J., Ealand C., Wang X., et al. The essential mycobacterial amidotransferase GatCAB is a modulator of specific translational fidelity. Nat. Microbiol. 2016;1 doi: 10.1038/nmicrobiol.2016.147. [DOI] [PubMed] [Google Scholar]
- 13.Samhita L., Raval P.K., Agashe D. Global mistranslation increases cell survival under stress in Escherichia coli. PLoS Genet. 2020;16 doi: 10.1371/journal.pgen.1008654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Krassowski T., Coughlan A.Y., Shen X.X., Zhou X., Kominek J., Opulente D.A., et al. Evolutionary instability of CUG-Leu in the genetic code of budding yeasts. Nat. Commun. 2018;9:1887. doi: 10.1038/s41467-018-04374-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Mühlhausen S., Schmitt H.D., Pan K.T., Plessmann U., Urlaub H., Hurst L.D., et al. Endogenous stochastic decoding of the CUG codon by competing ser- and Leu-tRNAs in Ascoidea asiatica. Curr. Biol. 2018;28:2046–2057.e2045. doi: 10.1016/j.cub.2018.04.085. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Suzuki T., Ueda T., Watanabe K. The 'polysemous' codon--a codon with multiple amino acid assignment caused by dual specificity of tRNA identity. EMBO J. 1997;16:1122–1134. doi: 10.1093/emboj/16.5.1122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Gomes A.C., Miranda I., Silva R.M., Moura G.R., Thomas B., Akoulitchev A., et al. A genetic code alteration generates a proteome of high diversity in the human pathogen Candida albicans. Genome Biol. 2007;8:R206. doi: 10.1186/gb-2007-8-10-r206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ochoa-Gutiérrez D., Reyes-Torres A.M., de la Fuente-Colmenares I., Escobar-Sánchez V., González J., Ortiz-Hernández R., et al. Alternative CUG codon usage in the halotolerant yeast. J. Fungi (Basel) 2022;8:970. doi: 10.3390/jof8090970. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Woese C.R. On the evolution of the genetic code. Proc. Natl. Acad. Sci. U. S. A. 1965;54:1546–1552. doi: 10.1073/pnas.54.6.1546. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Stynen B., Van Dijck P., Tournu H. A CUG codon adapted two-hybrid system for the pathogenic fungus Candida albicans. Nucleic Acids Res. 2010;38:e184. doi: 10.1093/nar/gkq725. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Côte P., Sulea T., Dignard D., Wu C., Whiteway M. Evolutionary reshaping of fungal mating pathway scaffold proteins. mBio. 2011;2:e00230-10. doi: 10.1128/mBio.00230-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Rocha R., Pereira P.J., Santos M.A., Macedo-Ribeiro S. Unveiling the structural basis for translational ambiguity tolerance in a human fungal pathogen. Proc. Natl. Acad. Sci. U. S. A. 2011;108:14091–14096. doi: 10.1073/pnas.1102835108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Sárkány Z., Silva A., Pereira P.J., Macedo-Ribeiro S. Ser or Leu: structural snapshots of mistranslation in Candida albicans. Front. Mol. Biosci. 2014;1:27. doi: 10.3389/fmolb.2014.00027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Fraga J.S., Sárkány Z., Silva A., Correia I., Pereira P.J.B., Macedo-Ribeiro S. Genetic code ambiguity modulates the activity of a C. albicans MAP kinase linked to cell wall remodeling. Biochim. Biophys. Acta Proteins Proteom. 2019;1867:654–661. doi: 10.1016/j.bbapap.2019.02.004. [DOI] [PubMed] [Google Scholar]
- 25.Miranda I., Silva-Dias A., Rocha R., Teixeira-Santos R., Coelho C., Gonçalves T., et al. Candida albicans CUG mistranslation is a mechanism to create cell surface variation. mBio. 2013;4 doi: 10.1128/mBio.00285-13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Feketová Z., Masek T., Vopálenský V., Pospísek M. Ambiguous decoding of the CUG codon alters the functionality of the Candida albicans translation initiation factor 4E. FEMS Yeast Res. 2010;10:558–569. doi: 10.1111/j.1567-1364.2010.00629.x. [DOI] [PubMed] [Google Scholar]
- 27.Bezerra A.R., Simões J., Lee W., Rung J., Weil T., Gut I.G., et al. Reversion of a fungal genetic code alteration links proteome instability with genomic and phenotypic diversification. Proc. Natl. Acad. Sci. U. S. A. 2013;110:11079–11084. doi: 10.1073/pnas.1302094110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Miranda I., Rocha R., Santos M.C., Mateus D.D., Moura G.R., Carreto L., et al. A genetic code alteration is a phenotype diversity generator in the human pathogen Candida albicans. PLoS One. 2007;2:e996. doi: 10.1371/journal.pone.0000996. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Santos M.A., Perreau V.M., Tuite M.F. Transfer RNA structural change is a key element in the reassignment of the CUG codon in Candida albicans. EMBO J. 1996;15:5060–5068. [PMC free article] [PubMed] [Google Scholar]
- 30.Dupree E.J., Jayathirtha M., Yorkey H., Mihasan M., Petre B.A., Darie C.C. A critical review of bottom-up proteomics: the good, the bad, and the future of this field. Proteomes. 2020;8:14. doi: 10.3390/proteomes8030014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Han Y., Ma B., Zhang K. SPIDER: software for protein identification from sequence tags with de novo sequencing error. J. Bioinform Comput. Biol. 2005;3:697–716. doi: 10.1142/s0219720005001247. [DOI] [PubMed] [Google Scholar]
- 32.Zhang J., Xin L., Shan B., Chen W., Xie M., Yuen D., et al. PEAKS DB: de novo sequencing assisted database search for sensitive and accurate peptide identification. Mol. Cell Proteomics. 2012;11 doi: 10.1074/mcp.M111.010587. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Cox J., Mann M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat. Biotechnol. 2008;26:1367–1372. doi: 10.1038/nbt.1511. [DOI] [PubMed] [Google Scholar]
- 34.Shevchenko A., Wilm M., Vorm O., Jensen O.N., Podtelejnikov A.V., Neubauer G., et al. A strategy for identifying gel-separated proteins in sequence databases by MS alone. Biochem. Soc. Trans. 1996;24:893–896. doi: 10.1042/bst0240893. [DOI] [PubMed] [Google Scholar]
- 35.Muzzey D., Schwartz K., Weissman J.S., Sherlock G. Assembly of a phased diploid Candida albicans genome facilitates allele-specific measurements and provides a simple model for repeat and indel structure. Genome Biol. 2013;14:R97. doi: 10.1186/gb-2013-14-9-r97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Skrzypek M.S., Binkley J., Binkley G., Miyasato S.R., Simison M., Sherlock G. The Candida Genome Database (CGD): incorporation of Assembly 22, systematic identifiers and visualization of high throughput sequencing data. Nucleic Acids Res. 2017;45:D592–D596. doi: 10.1093/nar/gkw924. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Mellacheruvu D., Wright Z., Couzens A.L., Lambert J.P., St-Denis N.A., Li T., et al. The CRAPome: a contaminant repository for affinity purification-mass spectrometry data. Nat. Methods. 2013;10:730–736. doi: 10.1038/nmeth.2557. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Pinheiro M., Afreixo V., Moura G., Freitas A., Santos M.A., Oliveira J.L. Statistical, computational and visualization methodologies to unveil gene primary structure features. Methods Inf. Med. 2006;45:163–168. [PubMed] [Google Scholar]
- 39.Kramer E.B., Vallabhaneni H., Mayer L.M., Farabaugh P.J. A comprehensive analysis of translational missense errors in the yeast Saccharomyces cerevisiae. RNA. 2010;16:1797–1808. doi: 10.1261/rna.2201210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Lynch M., Sung W., Morris K., Coffey N., Landry C.R., Dopman E.B., et al. A genome-wide view of the spectrum of spontaneous mutations in yeast. Proc. Natl. Acad. Sci. U. S. A. 2008;105:9272–9277. doi: 10.1073/pnas.0803466105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Zhu Y.O., Siegal M.L., Hall D.W., Petrov D.A. Precise estimates of mutation rate and spectrum in yeast. Proc. Natl. Acad. Sci. U. S. A. 2014;111:E2310–E2318. doi: 10.1073/pnas.1323011111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Chung C., Verheijen B.M., Navapanich Z., McGann E.G., Shemtov S., Lai G.J., et al. Evolutionary conservation of the fidelity of transcription. Nat. Commun. 2023;14:1547. doi: 10.1038/s41467-023-36525-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Tavares J., Assis-Santos F., Santos M.A. Proteomics analysis for amino acid misincorporation detection: Mini review. J. Proteomics Bioinform. 2018 doi: 10.4172/jpb.1000464. [DOI] [Google Scholar]
- 44.Varanda A.S., Santos M., Soares A.R., Vitorino R., Oliveira P., Oliveira C., et al. Human cells adapt to translational errors by modulating protein synthesis rate and protein turnover. RNA Biol. 2020;17:135–149. doi: 10.1080/15476286.2019.1670039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Kachuk C., Faulkner M., Liu F., Doucette A.A. Automated SDS depletion for mass spectrometry of intact membrane proteins though transmembrane electrophoresis. J. Proteome Res. 2016;15:2634–2642. doi: 10.1021/acs.jproteome.6b00199. [DOI] [PubMed] [Google Scholar]
- 46.Kim M.S., Zhong J., Pandey A. Common errors in mass spectrometry-based analysis of post-translational modifications. Proteomics. 2016;16:700–714. doi: 10.1002/pmic.201500355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Lenčo J., Khalikova M.A., Švec F. Dissolving peptides in 0.1% formic acid brings risk of artificial formylation. J. Proteome Res. 2020;19:993–999. doi: 10.1021/acs.jproteome.9b00823. [DOI] [PubMed] [Google Scholar]
- 48.Akashi H. Synonymous codon usage in Drosophila melanogaster: natural selection and translational accuracy. Genetics. 1994;136:927–935. doi: 10.1093/genetics/136.3.927. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Sun M., Zhang J. Preferred synonymous codons are translated more accurately: proteomic evidence, among-species variation, and mechanistic basis. Sci. Adv. 2022;8 doi: 10.1126/sciadv.abl9812. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Bezerra A.R., Oliveira C., Correia I., Guimarães A.R., Sousa G., Carvalho M.J., et al. The role of non-standard translation in Candida albicans pathogenesis. FEMS Yeast Res. 2021;21:foab032. doi: 10.1093/femsyr/foab032. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Weil T., Santamaría R., Lee W., Rung J., Tocci N., Abbey D., et al. Adaptive mistranslation accelerates the evolution of fluconazole resistance and induces major genomic and gene expression alterations in Candida albicans. mSphere. 2017;2 doi: 10.1128/mSphere.00167-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Amador-García A., Zapico I., Borrajo A., Malmström J., Monteoliva L., Gil C. Extending the proteomic characterization of Candida albicans exposed to stress and apoptotic inducers through data-independent acquisition mass spectrometry. mSystems. 2021;6 doi: 10.1128/mSystems.00946-21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Song N., Zhou X., Li D., Li X., Liu W. A proteomic landscape of Candida albicans in the stepwise evolution to fluconazole resistance. Antimicrob. Agents Chemother. 2022;66 doi: 10.1128/aac.02105-21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Mühlhausen S., Schmitt H.D., Plessmann U., Mienkus P., Sternisek P., Perl T., et al. Proteogenomics analysis of CUG codon translation in the human pathogen Candida albicans. BMC Biol. 2021;19:258. doi: 10.1186/s12915-021-01197-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Kulak N.A., Pichler G., Paron I., Nagaraj N., Mann M. Minimal, encapsulated proteomic-sample processing applied to copy-number estimation in eukaryotic cells. Nat. Methods. 2014;11:319–324. doi: 10.1038/nmeth.2834. [DOI] [PubMed] [Google Scholar]
- 56.Mostovenko E., Hassan C., Rattke J., Deelder A.M., van Veelen P.A., Palmblad M. Comparison of peptide and protein fractionation methods in proteomics. EuPA Open Proteomics. 2013;1:30–37. [Google Scholar]
- 57.den Ridder M., Knibbe E., van den Brandeler W., Daran-Lapujade P., Pabst M. A systematic evaluation of yeast sample preparation protocols for spectral identifications, proteome coverage and post-isolation modifications. J. Proteomics. 2022;261 doi: 10.1016/j.jprot.2022.104576. [DOI] [PubMed] [Google Scholar]
- 58.Halder A., Verma A., Biswas D., Srivastava S. Recent advances in mass-spectrometry based proteomics software, tools and databases. Drug Discov. Today Technol. 2021;39:69–79. doi: 10.1016/j.ddtec.2021.06.007. [DOI] [PubMed] [Google Scholar]
- 59.Verheggen K., Raeder H., Berven F.S., Martens L., Barsnes H., Vaudel M. Anatomy and evolution of database search engines-a central component of mass spectrometry based proteomic workflows. Mass Spectrom. Rev. 2020;39:292–306. doi: 10.1002/mas.21543. [DOI] [PubMed] [Google Scholar]
- 60.Shao W., Lam H. Tandem mass spectral libraries of peptides and their roles in proteomics research. Mass Spectrom. Rev. 2017;36:634–648. doi: 10.1002/mas.21512. [DOI] [PubMed] [Google Scholar]
- 61.Vitorino R., Guedes S., Trindade F., Correia I., Moura G., Carvalho P., et al. De novo sequencing of proteins by mass spectrometry. Expert Rev. Proteomics. 2020;17:595–607. doi: 10.1080/14789450.2020.1831387. [DOI] [PubMed] [Google Scholar]
- 62.Knudsen G.M., Chalkley R.J. The effect of using an inappropriate protein database for proteomic data analysis. PLoS One. 2011;6 doi: 10.1371/journal.pone.0020873. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Han X., He L., Xin L., Shan B., Ma B. PeaksPTM: mass spectrometry-based identification of peptides with unspecified modifications. J. Proteome Res. 2011;10:2930–2936. doi: 10.1021/pr200153k. [DOI] [PubMed] [Google Scholar]
- 64.Fang Y., Robinson D.P., Foster L.J. Quantitative analysis of proteome coverage and recovery rates for upstream fractionation methods in proteomics. J. Proteome Res. 2010;9:1902–1912. doi: 10.1021/pr901063t. [DOI] [PubMed] [Google Scholar]
- 65.Shevchenko A., Tomas H., Havlis J., Olsen J.V., Mann M. In-gel digestion for mass spectrometric characterization of proteins and proteomes. Nat. Protoc. 2006;1:2856–2860. doi: 10.1038/nprot.2006.468. [DOI] [PubMed] [Google Scholar]
- 66.Deng L., Handler D.C.L., Multari D.H., Haynes P.A. Comparison of protein and peptide fractionation approaches in protein identification and quantification from Saccharomyces cerevisiae. J. Chromatogr. B Analyt. Technol. Biomed. Life Sci. 2021;1162 doi: 10.1016/j.jchromb.2020.122453. [DOI] [PubMed] [Google Scholar]
- 67.Ji Q.Q., Fang Z.P., Ye Q., Ruan Z.R., Zhou X.L., Wang E.D. C-Terminal domain of leucyl-tRNA synthetase from pathogenic Candida albicans recognizes both tRNASer and tRNALeu. J. Biol. Chem. 2016;291:3613–3625. doi: 10.1074/jbc.M115.699777. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Zhou X.L., Fang Z.P., Ruan Z.R., Wang M., Liu R.J., Tan M., et al. Aminoacylation and translational quality control strategy employed by leucyl-tRNA synthetase from a human pathogen with genetic code ambiguity. Nucleic Acids Res. 2013;41:9825–9838. doi: 10.1093/nar/gkt741. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Michalski A., Cox J., Mann M. More than 100,000 detectable peptide species elute in single shotgun proteomics runs but the majority is inaccessible to data-dependent LC-MS/MS. J. Proteome Res. 2011;10:1785–1793. doi: 10.1021/pr101060v. [DOI] [PubMed] [Google Scholar]
- 70.Fernández-Costa C., Martínez-Bartolomé S., McClatchy D.B., Saviola A.J., Yu N.K., Yates J.R. Impact of the identification strategy on the reproducibility of the DDA and DIA results. J. Proteome Res. 2020;19:3153–3161. doi: 10.1021/acs.jproteome.0c00153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Chen C., Hou J., Tanner J.J., Cheng J. Bioinformatics methods for mass spectrometry-based proteomics data analysis. Int. J. Mol. Sci. 2020;21:2873. doi: 10.3390/ijms21082873. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Perez-Riverol Y., Bai J., Bandla C., García-Seisdedos D., Hewapathirana S., Kamatchinathan S., et al. The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res. 2022;50:D543–D552. doi: 10.1093/nar/gkab1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE (72) partner repository with the dataset identifier PXD047025 and 10.6019/PXD047025. R scripts of the pipeline used for data processing are available at the GitHub repository: https://github.com/andreia-reis/proteogenomic_pipeline_calbicans and at Zenodo: https://doi.org/10.5281/zenodo.10651622.








