Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jun 9;37(7):1562–1570. doi: 10.1021/jasms.6c00012

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Xi Xing †,‡,§, Wenyun Lu †,‡, Xi Li †,‡,§, Jimmy S Pratas †,‡, Anna M Oschmann †,‡, Joshua D Rabinowitz †,‡,§,∥,*
PMCID: PMC13329995  PMID: 42261582

Abstract

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.


graphic file with name js6c00012_0008.jpg


graphic file with name js6c00012_0006.jpg

1. Introduction

The discovery of unknown metabolites represents a critical frontier in metabolomics, offering insights into novel biochemical pathways, enzyme functions, and disease mechanisms. − Yeasts are one of the most extensively studied eukaryotic model systems in biology. Saccharomyces cerevisiae is both a classical model for cellular and metabolic research and an important industrial organism for baking, brewing, and biofuel production. Nonmodel yeasts have distinct metabolic properties from S. cerevisiae and hold potential for enhanced production of biofuels and chemical feedstocks. , For example, Rhodotorula toruloides, an oleaginous yeast, naturally accumulates copious lipids and has been engineered to produce fatty acid ethyl esters, triacetic acid lactone, and other valuable chemicals. − Similarly, Issatchenkia orientalis has attracted attention for its remarkable tolerance to acidic and harsh conditions, making it useful for organic acid production under low-pH fermentation. − These species’ robust metabolism, resilience to harsh conditions, and ability to utilize diverse carbon sources make them attractive for industrial applications. To capitalize fully on the biotechnology potential of these yeasts, it is important to identify their full metabolic capacity, which likely goes beyond textbook reactions.

Liquid chromatography–mass spectrometry (LC-MS) has emerged as a cornerstone technology for metabolite profiling, offering high sensitivity, broad dynamic range, and the ability to separate and detect diverse chemical species. − In untargeted analysis, there are two primary objectives: complete peak annotation/quantitation and unknown discovery. Complete peak annotation seeks to comprehensively interpret all detected features (peaks). In contrast, unknown discovery focuses on solving the structures of a set of prioritized peaks. Both tasks involve a shared challenge: Untargeted LC-MS experiments typically detect tens of thousands of features, the majority of which are redundant “sibling” peaks arising from natural isotopes, adducts, in-source fragments, and related phenomena or are instrumental artifacts, with only a small fraction of peaks assignable to “parent” metabolite molecular ions [M+H]+ or [M–H]–. , Due to the diversity of redundancies and artifacts and the wide range of molecular formulas that can be assigned to a given mass (even when measured at high resolution), unambiguous assignment of whether a peak is a metabolite molecular ion remains challenging, as does molecular formula assignment.

Moving from formulas to structures presents an even greater challenge. MS2 spectral library matching is a widely used strategy for structural identification but is more effective for proteins and peptides than small molecules, whose structures do not follow the defined template of peptides and for which MS2 spectral variation (entropy) across isomeric structures is often low. Furthermore, MS2 library searching is inherently restricted to compounds already reported in existing databases, limiting its utility for the unknown discovery.

Stable isotope tracing–based metabolite annotation and identification has emerged as a complementary strategy. − In particular, the mass isotopomer distribution (MID) derived from isotope tracing provides information orthogonal to MS2, enhancing confidence in metabolite characterization. Recently, Gao et al. developed an isotopologue similarity networking strategy (IsoNet) to deduce unknown metabolic reactions in living cells and mice. Similarly, Secilmis et al. measures pairwise distances between mass isotopologue distributions (MIDs) to identify unknown metabolites related to known compounds. These strategies rely on partial 13C labeling that differs across metabolites. In yeast grown on uniformly 13C-glucose, nearly all extracted metabolites become fully labeled, eliminating such information. The near-complete 13C labeling, however, enables the carbon number of an unknown peak to be determined, with greater confidence than based on natural isotope abundances alone, which are susceptible to mass-spectral interferences, ion suppression effects, and low signal-to-noise. Mahieu et al. used 13C credentialing to identify biologically relevant peaks with confident 13C incorporation. Similarly, Wang et al. introduced a dual13C-glucose and 15N-NH3 labeling to assign formulas based on the HMDB database match. Inferred atom counts were not used, however, for distinguishing whether unknown peaks could be assigned sensible formulas.

For a given m/z value, an exhaustive formula generator can produce many plausible molecular formulas (often >100 within 1 ppm mass error). If carbon numbers can be precisely inferred and fixed, however, then a unique formula can often be assigned. Moreover, we find that the combination of m/z and carbon count can effectively improve the discrimination between true metabolites and artifacts based on whether a sensible formula exists. Thus, 13C-labeling-informed formula assignment substantially narrows the pool of candidate features for novel metabolite discovery, complements biological peak filtering, and aids in prioritization of peaks for future structural elucidation. Applying this approach to three bioenergy-relevant yeast species revealed high-abundance biological peaks with sensible formulas that were not found in YMDB or HMDB, pointing to previously uncharacterized yeast metabolites.

2. Experiment

2.1. Cell Culture and Metabolite Extraction

Wild-type S. cerevisiae (DBY11096, MATa derivative of S288C), R. toruloides (IFO0880), and I. orientalis (SD108) were cultured in a shaker at 30 °C and 250 rpm in medium containing 20 g L–1 glucose and 6.7 g L–1 yeast nitrogen base (YNB) without amino acids (pH 5; Sigma, Y0626). Cells were cultured under two conditions, with glucose provided either unlabeled or uniformly 13C-labeled in both the overnight seed culture and the fresh medium. For metabolite extraction, 2.4 mL of culture (OD600 = 0.6–0.8) was vacuum-filtered through nylon membranes (0.45 μm pore size; GVS Magna, 1213776) on a fritted-glass support. The membrane containing the cell was immediately immersed in 1.5 mL of prechilled extraction solvent (40:40:20 acetonitrile:methanol:water, with 0.5% formic acid, v/v; −20 °C) in a Petri dish and incubated for ∼1 min on dry ice–cooled wet ice. Extracts were then neutralized by adding 132 μL of 15% (w/v) NH4HCO3, transferred to an Eppendorf tube, and centrifuged at 16,000g for 10 min at 4 °C. The supernatant was collected and stored in −80 °C before LC–MS analysis. Three biological replicates were prepared for each condition (unlabeled and uniformly 13C-labeled) for all three strains.

2.2. LC-MS Analysis

The LC-MS analysis was performed on a Thermo Fisher Scientific Vanquish UHPLC system coupled with an Exploris 480 orbitrap mass spectrometer. LC separation was achieved using a Waters XBridge BEH Amide column (2.1 × 150 mm, 2.5 μm particle size) with a 25 min gradient. Solvent A is 95:5 water:acetonitrile with 20 mM ammonium hydroxide and 20 mM ammonium acetate, pH 9.4. Solvent B is acetonitrile. The gradients are: 0 min, 90% B; 2 min, 90% B; 3 min, 75%; 7 min, 75% B; 8 min, 70%, 9 min, 70% B; 10 min, 50% B; 12 min, 50% B; 13 min, 25% B; 14 min, 25% B; 16 min, 0% B; 20.5 min, 0% B; 21 min, 90% B; 25 min, 90% B. Other LC parameters are: flow rate 150 μL/min, column temperature 25 °C, injection volume 5 μL. Mass spectrometry parameters are: full scan MS1 scan range m/z 70–1000, spray voltage 2800 V (negative mode), sheath gas 35 (Arb), aux gas 10 (Arb), sweep gas 0.5 (Arb), ion transfer tube temperature 300 °C, vaporizer temperature 35 °C, internal mass calibration on, and RF lens 60%. LC-MS data were generated from both unlabeled and 13C-labeled yeast cultures across three yeast species.

Targeted MS2 were performed for selected peaks of interest using a parallel reaction monitoring (PRM) approach. Samples were analyzed with a full scan, followed by targeted MS2 scans using an inclusion list in the same LC-MS run. Full scan parameters were: resolution 60,000, range m/z 70–1000, AGC target 1e7, ITmax 100 ms. MS/MS parameters were: isolation window 1.5 m/z, collision energies 15, 20, 30 eV, resolution 15,000, AGC target 1e6, ITmax 100 ms, RT window 3 min.

3. Methods, Results, and Discussion

We designed a computational pipeline to identify mass spectrometry features enriched for biologically interesting unknown metabolites with assignable formulas, thereby narrowing the pool from thousands of features to a few hundred priority ones. MS1-based analysis begins with peak picking and progresses to peak annotation, artifact removal, and molecular formula assignment (Figure a–f). This is followed by targeted MS2 experiments on the selected peaks, MS2 spectral cleaning and structure prediction/validation (Figure g–i). Each step is described sequentially below.

1.

1

Computational pipeline for yeast metabolite discovery. (a) LC–MS data were generated from unlabeled and 13C-labeled yeast cultures across three yeast species. Peak picking from unlabeled data sets (three biological replicates) was performed using El-Maven. (b) Peak quality was assessed based on extracted ion chromatograms (EICs) using a convolutional neural network (CNN) model. (c) Biological peaks were identified by comparing unlabeled and 13C-labeled data sets (lack of labeling is not biological). (d) Co-eluting MS1 features were analyzed to annotate and flag common redundant peaks. (e) Carbon number was determined based on mass shift in 13C-labeled samples. (f) Molecular formulas were generated using a custom formula generator constrained by elemental composition and by the identified carbon number, filtered by machine learning–based formula plausibility assessment. (g) Targeted MS2 spectra were acquired and annotated. (h) MS2 spectra were cleaned based on EIC correlation between MS1 and MS2 signals. (I) For each prioritized peak, the top 5 structural candidates were generated by DeepMet. Verification against authentic standards via retention time and MS2 spectral matching represents an ongoing work.

3.1. Peak Picking and Quality Control

A variety of peak-detection software − and peak curation algorithms exist. − In this study, peak picking from raw LC–MS data of cells grown in unlabeled media (three biological replicates) was performed for S. cerevisiae, R. toruloides, and I. orientalis in parallel, using the El-Maven software package (v0.12.0; mass resolution = 5 ppm, time domain resolution = 10 scans, minimum intensity = 5e4). Each data set generated an untargeted peak list containing approximately ∼5000 features, which served as the starting point of the pipeline (Figure a). Peaks with matching m/z and retention time (RT) values within predefined tolerance windows (10 ppm for m/z and 0.1 min for RT) were considered identical features. A customized peak quality filter was then applied to remove bad peaks with low-quality extracted ion chromatograms (EICs) (Figure b). The filter employed a pretrained deep learning model using a similar architecture and training data as prior work (Supporting Information Figure S1). For each peak, the three replicate EICs were evaluated. The peak was retained if at least two out of three EICs were classified as “good” in peak shape. This filter removed approximately 10% of peaks due to low quality.

3.2. High-Abundance Biological Peaks

We next focused on peaks with high signal intensities (peak height >1 × 105 ion counts per second), as such peaks are more likely to yield high-quality MS/MS spectra and to reflect biological metabolites present at substantial concentration, reducing the total number of peaks by half. In addition, peaks were required to exhibit clear 13C incorporation from the sole carbon source, 13C-glucose, as defined by at least a 10-fold reduction in the mean signal intensity of the unlabeled peak in the 13C-labeled samples compared to the unlabeled controls (Figure c). Three biological replicates were used for both 12C and 13C conditions to minimize the effect of culture-to-culture variance. This 13C incorporation criterion serves as a preliminary filter to deprioritize features that do not show strong evidence of 13C incorporation. Features that pass this filter are retained in subsequent analyses. Collectively, these steps together are conceptually similar to the credentialing approach used in ref. . The combined procedures above reduce the total MS1 features from 6011→606, 5302→1021 and 4250→818 for S. cerevisiae, R. toruloides, and I. orientalis, respectively.

3.3. Detection of Redundant Peaks Such as Adducts and Isotopes

For the remaining peaks, we performed several analyses in parallel, including redundant peak detection (Figure d), carbon number determination (Figure e), and formula assignment (Figure f). Ringing artifacts (satellite peaks surrounding highly abundant signals in Orbitrap data), natural-abundance isotopic peaks, co-eluting adducts, in-source fragments, dimers, multiply charged species, or a combination of the above were detected and deprioritized. Importantly, all these peaks coelute with the monoisotopic molecular ion, facilitating their identification. A rule-based pipeline based on pairwise peak relationships was used to identify these redundancies. This pipeline is conceptually similar to that used in ref. . In-source fragment peaks were identified using all-ion fragmentation (AIF) data acquired at three collision energies (0, 5, and 10 eV) with three replicates per condition. Peaks showing a statistically significant increase of at least 1.5 fold in intensity at either 5 or 10 eV collision energy were annotated as in-source fragments. For detecting isotopic, adducts and dimers, a list of commonly observed m/z differences, along with a range of relevant intensity ratios, was defined (Supporting Information Table S1). A limitation is that certain mass differences can correspond either to redundant features originating from the same metabolite or to distinct metabolite ions. For example, the [M+Na–2H]− adduct exhibits a characteristic mass difference of 21.982 Da relative to the precursor ion. However, metabolites differing by +2 C and −2 H have a similar m/z difference of 21.984 Da and such a peak with +2 C and −2 H could be misannotated as a sodium adduct. Similarly, isotopes such as 34S, 37Cl, 30Si, and 41K each exhibit characteristic mass differences of approximately 1.997 Da, which makes their isotope-based annotation less reliable. A second limitation is that only common adducts included in the rule-based list are detected, leaving many unknown redundant features unannotated. Inaccuracy and incompleteness continue to pose challenges for effective redundant peak removal, highlighting the need for complementary approaches.

3.4. Carbon Number Determination from 13C-Labeling Data

For singly charged 13C-labeled peaks, the observed m/z shift between unlabeled and labeled samples corresponds to the number of incorporated carbons (n × 1.00335 Da) (Figure ). Using three replicates of both unlabeled and labeled data sets, a pattern-matching algorithm compares the experimental isotopic distributions with theoretical patterns across all possible carbon counts (n = 1···N), from which the number of labeled carbon atoms is deduced. To minimize interference from neighboring peaks, the algorithm focuses only on the signals combining M + 0 and M + 1 positions in the unlabeled samples, and M + n – 1 and M + n positions in the 13C-labeled samples. While the primary signal is expected to shift from M + 0 in the unlabeled to M + n in the labeled condition, the M + 1 signal reflects the natural 13C isotope abundance, and the M + n – 1 signal arises from tracer impurity (∼1%), facilitating carbon-number inference. For each candidate carbon number n, a theoretical isotopic pattern was generated and compared with the experimental pattern, both normalized to unity, and a similarity score (Pearson correlation coefficient) reflects the likelihood that n is the correct carbon count. For each peak, the algorithm systematically enumerates all possible carbon counts up to the theoretically maximum allowable value and outputs n with the highest similarity score as the predicted carbon count.

2.

2

Inference of carbon number from 13C-labeling data. (a) Mass isotopologue distributions (MIDs) in unlabeled (blue bars, 3 replicates) and labeled (red bars, 3 replicates) samples. This example shows a carbon shift of 7. The inset displays scores for all possible carbon numbers. The correct carbon shift corresponds to the highest score. (b) Labeling-pattern similarity score for a given presumed carbon number n is calculated by comparing the experimental data with simulated distributions and finding the Pearson correlation.

3.5. Formula Assignment

A customized formula generator tool was developed to deduce molecular formulas from m/z values and carbon counts for each peak of interest in yeast. Similar to existing tools, , this generator exhaustively enumerates all possible elemental compositions within user-defined elemental constraints. In this work, only elements C, H, N, O, S, and P are considered, with reasonable constraintsmost notably the fixed carbon numberto balance completeness while avoiding an excessive number of unlikely formulas. In addition, a classification model is employed to evaluate the plausibility of each candidate formula (formula plausibility classifier). For training the classifier, molecular formulas from HMDB were used as positive examples, while an approximately equal amount of computer-generated “fake” formulas were used as negative examples. Fake formulas were generated by applying the formula generator to decoy m/z values obtained by perturbing the m/z of genuine metabolites within a 10–100 ppm window. This perturbation deviated from the correct m/z while preserving a mass defect distribution similar to that of the known compounds, and a random formula returned by the generator (if any) was selected. Model predictors included the atom counts for each element (C, H, N, O, S, and P), elemental ratios relative to hydrogen, the mass defect, and the m/z value. Four different classification models were tested: fine tree, boosted tree, bagged tree, and naïve Bayes. The top-performing formula plausibility classifier turned out to be the bagged tree model, which achieved an accuracy of 97%.

Using all unique formulas of known metabolites in YMDB (∼1500 formulas with m/z < 1000, excluding inorganics and those containing unusual elements), we simulated the impact of incorporating the carbon-number constraint and the machine learning–based plausibility classifier, which stepwise improve the accuracy of formula assignment (Supporting Information Figure S2). With the correct carbon number provided, more than 80% of these m/z values receive a unique molecular formula assignment, compared to <5% when no carbon constraint is applied. In contrast, when the carbon number is incorrectly input (perturbed by ±1 in simulation), most m/z values return no valid formula. Together, these results demonstrate that the customized formula generator is effective at recognizing metabolite [M–H]- ions based on m/z and carbon count. This helps to deprioritize many ambiguous MS featuresthose not clearly identified as common adducts yet lacking a plausible formula assignment. The underlying logic is that if no valid formula can be assigned to a given m/z peak, the likely explanations are (1) the presence of other elements (e.g., Na, K, B, Si, F, and Mg, which means that the peak is likely an adduct); (2) elementary atom counts beyond the bounds; (3) multiply charged or isotopic variants; (4) high m/z error; or (5) incorrect carbon number input. Most of these reasons are appropriate grounds to deprioritize the peak as unlikely to be a genuine novel metabolite ion.

To further evaluate whether true unknown metabolites can be effectively distinguished from redundant peaks using the formula generator, we assessed 1021 biological peaks detected in R. toruloides. As shown in Figure , these peaks were classified into four categories: (1) redundant peaks identified as isotopes, adducts, or in-source fragments based on MS1 signatures; (2) putative metabolites with molecular formula matches in YMDB; (3) putative metabolites without formula matches in YMDB, but with matches in HMDB; and (4) unknowns without HMDB or YMDB matches, representing all remaining peaks (note that the unknowns may match formula in larger chemical databases such as PubChem). Without applying carbon number constraints, the formula generator typically produced more than 10 candidate formulas for most peaks, failing to distinguish between redundant peaks and metabolites, and rarely assigning unique formulas even for the simplest known metabolites. Incorporating carbon count substantially improved performance, reducing the number of candidate formulas and enhancing group separation. When further combined with the formula plausibility classifier, group separation further improved: Most redundant peaksdominated by adducts, isotopes, or other nonbiological featuresfailed to yield plausible formulas, whereas approximately 80% of YMDB peaks were assigned unique formulas. The proportion of peaks with ambiguous assignments (2–5 formula candidates) was also markedly reduced compared with carbon-number restriction alone. Most unknowns returned no plausible formula, suggesting that the unknowns are likely dominated by artifacts or redundant peaks of unknown origin rather than unknown genuine metabolites (Figure ). Similar results are obtained for the other two yeast species (Supporting Information Figures S3 and S4). Table below shows peak categorization for three yeast species.

3.

3

Evaluation of the formula generator for assigning unique molecular formulas to unannotated peaks. A total of 1021 biologically relevant peaks in R. toruloides were classified into four groups: (1) artifacts peaks (isotopes, adducts, or in-source fragments based on co-eluting MS1 peaks); (2) putative metabolites with molecular formula matches in YMDB; (3) putative metabolites without formula matches in YMDB but with matches in HMDB; and (4) unknowns, representing all remaining high-abundance biological peaks. The bar plots show the distribution of peaks in each category according to the number of candidate formulas assigned by the formula generator under three conditions: (a) no carbon-number restriction, (b) with carbon-number restriction, and (c) with both carbon-number restriction and plausibility classification. The combined constraints substantially improve the ability to assign realistic, unique formulas. Note that true metabolites (with formula matches in YMDB and HMDB) most commonly have a single formula match, while redundancies and unknowns most commonly have zero formula matches.

1. Categorization of High-Abundance Biological Peaks Across the Three Yeast Species, Based on Redundancy Filtering and Matches Returned from Ymdb and Hmdb Database Formula Searches.

yeast species S. cerevisiae R. toruloides I. orientalis
total biological peaks 606 1021 818
redundant peaks 249 388 287
YMDB formula match 142 177 147
HMDB formula match 104 214 166
unknowns 111 242 218

3.6. Structure Identification by Targeted MS2

As illustrated in Figure g, targeted MS2 spectra were acquired using parallel reaction monitoring for all biological peaks excluding the redundancy peaks. Although a narrow window of 1.5 Da was used to isolate the precursor ion, contaminants from co-eluting ions in that mass range are still common, resulting in “chimeric” spectra that often require spectral deconvolution or complementary computational methods. Here, we employed an MS2 cleaning procedure to remove fragments arising from contaminant peaks that are not well separated from the precursor ion (Figure h). PRM-type MS2 experiment yields chromatographic peaks for each fragment m/z, allowing for peak shape comparison. The cleaning is based on extracted ion chromatogram (EIC) correlation between MS1 and MS2 signals, using the Pearson correlation coefficient. A correlation score threshold of 0.8 was used, with MS2 fragments below this cutoff excluded. For each precursor ion with an assignable formula, the formulas of the corresponding fragment peaks were inferred using the same formula generator, applying elemental constraints derived from the precursor formula, thereby providing insights into potential substructural units.

By matching the cleaned MS2 spectra against the MS2 library (MetaboAnalyst 6.0, MS2 Spectra Reference Databases) and using the cosine similarity score, we validated 50 peaks in R. toruloides, and 55 peaks in I. orientalis as matching known metabolites in YMDB, and 19 and 17 additional peaks, respectively, as matching known metabolites in HMDB. Figure shows a few representative examples of peaks with HMDB match confirmed by MS2 library matching, which are not reported previously as yeast metabolites (not in YMDB), including ophthalmic acid and 5-methoxy-3-indoleacetic acid in R. toruloides, as well as N-Acetyl Histidine and Fatty Acid 18:3 in I. orientalis.

4.

4

Example peaks not found in YMDB but found in HMDB and confirmed by MS2 library matching after the MS2 spectral cleaning.

Structural identification of unknown metabolites is assisted by fragment interpretation of cleaned MS2 spectra from the previous steps (Figure g,h) and involves the use of a language model–based AI tool DeepMet (Figure i). DeepMet generates plausible structures by learning patterns from known metabolites and scoring candidates based on how frequently they are predicted. The SMILES of the top five scoring structural candidates are generated, and peaks with plausible formulas cross-validated by DeepMet predictions are prioritized. Additional efforts to resolve isomeric structures and validate candidates predicted by DeepMet using authentic standards or chemical synthesis remain ongoing and will be reported in the future.

3.7. Accuracy of C Number Determination

The success of assigning unique formulas to true metabolites in this untargeted yeast analysis relies on the assumption that full 13C incorporation occurs for all of the metabolites. However, is this assumption valid? If not, what fraction of metabolites represent exceptional cases? Another important question concerns the accuracy of carbon number determination and the rate of false formula assignment. To address these questions, we examined each peak group in more detail.

For peaks with a YMDB formula match, we compared the reported formulas with those generated using the customized formula generator with known carbon shift and found only ∼5% mismatches, representing a low error rate. The discrepancies arise either because the YMDB formula is incorrectresulting from an accidental m/z matchor because the carbon number is wrong. As an example of a wrong match to YMDB, an I. orientalis-specific peak at m/z = 204.0513 is matched in YMDB to C8H15NOS2 with <5 ppm error. However, the carbon shift pattern clearly indicates 7 carbons, and the formula generator returns a unique formula of C7H11NO6. Further inspection of the MS1 signature confirms the absence of sulfur, and the C7 compound is ultimately found in HMDB and verified by MS2 as succinyl-serine. As an example of wrong carbon number determination, a peak at m/z = 328.2858 is matched in YMDB to C19H39NO3, but the carbon shift indicates 12 carbons, resulting in no formula being returned. A close manual inspection of the labeling pattern shows that there is indeed labeling at M+19, which aligns with the YMDB formula. This assignment is further supported by the natural-abundance 13C ratio, which indicates a carbon number range of 18–21 (Supporting Information Figure S5). There are also a few cases where formula-carbon number mismatch arises from incompletely labeled peaks, which we define as peaks with “carbon defects”a difference between the actual total carbon number and the observed 13C shift. These cases account for only a small fraction (∼1%) of the peaks with a YMDB match. For example, a compound with the formula C5H9NO4 shows a carbon shift of 4, indicating a carbon defect of 1. Similarly, a compound with the formula C6H8O5 exhibits a carbon shift of 5, also corresponding to a carbon defect of 1. The unlabeled carbon in these cases has been confirmed to originate from the formic acid in the extraction buffer, i.e., formic acid can react with serine to generate a new compound of formyl-serine, which is an abiotic product rather than an endogenous metabolite, thereby accounting for the presence of the unlabeled carbon atom.

For peaks matching to an HMDB formula but not the YMDB formula, we also compared the reported formulas with those generated using the customized formula generator constrained by the known carbon shift. A substantially higher mismatch rate (∼25%) was observed, and more than half of these cases were puzzling as no plausible formula was returned by the generator. In such instances, it remains unclear whether the discrepancy arises from an incidental m/z match in HMDB or from an incorrect determination of the carbon number. Since HMDB represents a much larger formula pool, the likelihood of incidental m/z matches is correspondingly higher. Interestingly, a group of seven compounds with a carbon defect of 10 were found, all from R. toruloides. They share similar molecular formulas and retention times, and each contains the minimal repeating unit C10H13NO (Supporting Information Table S2). The carbon source for these compounds has not yet been identified but may be derived from trace vitamins present in the growth medium (the glucose but not vitamins in the media is 13C-labeled). So far, only carbon defects of 1, 2, and 10 have been observed and confirmed, representing only a small fraction of all peaks. This is encouraging, as it supports our main assumption that yeast metabolites are mainly fully labeled by 13C-glucose.

3.8. Peaks Likely Corresponding to Novel Metabolites

For peaks without matches in either YMDB or HMDB, approximately one-third yielded at least one plausible novel molecular formula. This subset represents a valuable pool of potential true unknown metabolites. Figure a shows a Venn diagram summarizing the overlap of unique novel formulas among the three species, with low overlap suggesting that the novel metabolites are largely species-specific. Inspection of these putative formulas shows that the majority contain more than 10 carbons and can often be readily recognized as lipids or long-chain fatty acid–related molecules, especially those from the oleaginous yeast R. toruloides. There are fewer smaller compounds (≤10 carbons), but these are particularly structurally and metabolically intriguing.

5.

5

Novel formulas largely reflect nonmodel yeast metabolites with >15 carbon atoms. (a) Overlap of unique novel formulas identified across three yeast species. (b) Number of novel metabolites as a function of carbon count.

The complete list of putative novel metabolites with formulas is provided in the Supporting Information, containing 89, 64, and 32 features for R. toruloides, I. orientalis, and S. cerevisiae, respectively. There are no matching MS2 spectra in existing libraries for these putatively novel metabolites. We instead work to infer their structures directly from the observed fragment signatures. We are also testing an alternative strategy that combines in silico MS/MS prediction tools with LLM-predicted structural candidates.

4. Conclusion

Imposing a carbon-number constraint derived from uniform 13C-labeling data, together with an appropriate formula generator incorporating elemental constraints and a plausibility filter, enables unique molecular formula assignment for most true metabolites and thereby effectively separates them from the vast majority of other unknown mass-spectral features. Building on this insight, we developed an untargeted analysis pipeline tailored for the discovery of unknown metabolites in bioenergy-relevant yeasts, substantially reducing the pool of high-abundance biological unknowns to about one hundred priority candidates. This analysis suggests that the number of true unknown metabolites may be even smaller than estimates based on high-quality credentialing and redundant peak removal. We anticipate that this work will facilitate future unknown metabolite discovery in yeast and other fully labeled systems by focusing efforts on a high-confidence set of candidate peaks and formulas.

Supplementary Material

js6c00012_si_001.xlsx (34.3KB, xlsx)
js6c00012_si_002.pdf (399.1KB, pdf)

Acknowledgments

This work was funded by Department of Energy (DOE) DESC0018260 supporting J.D.R.; the DOE Center for Advanced Bioenergy and Bioproducts Innovation (U.S. Department of Energy, Office of Science, Biological, and Environmental Research Program under Award Number DE-SC0018420); any opinions, findings, and conclusions or recommendations are those of the author(s) and do not necessarily reflect the views of the U.S. Department of Energy. W.L. is supported by NIH grant R50CA211437.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jasms.6c00012.

The authors declare no competing financial interest.

Published as part of Journal of the American Society for Mass Spectrometry special issue “Biemann: Embracing the Unknown”.

References

  1. Agongo J., Grady S. F., Cho K., Patti G. J., Bythell B. J., Arnatt C. K., Edwards J. L.. Discovery and Identification of Three Homocysteine Metabolites by Chemical Derivatization and Mass Spectrometry Fragmentation. Anal. Chem. 2024;96(29):11639–11643. doi: 10.1021/acs.analchem.4c01706. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Giera M., Yanes O., Siuzdak G.. Metabolite discovery: Biochemistry’s scientific driver. Cell Metab. 2022;34(1):21–34. doi: 10.1016/j.cmet.2021.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Lin C., Tian Q., Guo S., Xie D., Cai Y., Wang Z., Chu H., Qiu S., Tang S., Zhang A.. Metabolomics for Clinical Biomarker Discovery and Therapeutic Target Identification. Molecules. 2024;29(10):2198. doi: 10.3390/molecules29102198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Patti G. J., Yanes O., Siuzdak G.. Metabolomics: the apogee of the omics trilogy. Nat. Rev. Mol. Cell Biol. 2012;13(4):263–269. doi: 10.1038/nrm3314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Prosser G. A., Larrouy-Maumus G., de Carvalho L. P.. Metabolomic strategies for the identification of new enzyme functions and metabolic pathways. EMBO Rep. 2014;15(6):657–669. doi: 10.15252/embr.201338283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Sindelar M., Patti G. J.. Chemical Discovery in the Era of Metabolomics. J. Am. Chem. Soc. 2020;142(20):9097–9105. doi: 10.1021/jacs.9b13198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Lacerda M. P., Oh E. J., Eckert C.. The Model System Saccharomyces cerevisiae Versus Emerging Non-Model Yeasts for the Production of Biofuels. Life (Basel) 2020;10(11):299. doi: 10.3390/life10110299. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Navarrete C., Martínez J. L.. Non-conventional yeasts as superior production platforms for sustainable fermentation based bio-manufacturing processes. AIMS Bioengineering. 2020;7(4):289–305. doi: 10.3934/bioeng.2020024. [DOI] [Google Scholar]
  9. Jagtap S. S., Deewan A., Liu J.-J., Walukiewicz H. E., Yun E. J., Jin Y.-S., Rao C. V.. Integrating transcriptomic and metabolomic analysis of the oleaginous yeast Rhodosporidium toruloides IFO0880 during growth under different carbon sources. Appl. Microbiol. Biotechnol. 2021;105(19):7411–7425. doi: 10.1007/s00253-021-11549-8. [DOI] [PubMed] [Google Scholar]
  10. Li X., Weilandt D. R., Keber F. C., Subramanian A. M., Loynes S. R., Rao C. V., Shen Y., Wühr M., Rabinowitz J. D.. Lipid accumulation in nitrogen and phosphorus-limited yeast is caused by less growth-related dilution. Metab. Eng. 2026;93:60–72. doi: 10.1016/j.ymben.2025.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Mooney E. J., Suthers P. F., Schroeder W. L., Dinh H. V., Li X., Shen Y., Xiao T., Call C. M., Baron H., Subramanian A. M.. et al. Metabolic flux and resource balance in the oleaginous yeast Rhodotorula toruloides . Metab. Eng. 2026;94:169–181. doi: 10.1016/j.ymben.2025.11.012. [DOI] [PubMed] [Google Scholar]
  12. Dubinkina V., Bhogale S., Hsieh P.-H., Dibaeinia P., Nambiar A., Maslov S., Yoshikuni Y., Sinha S.. A transcriptomic atlas of acute stress response to low pH in multiple Issatchenkia orientalis strains. Microbiol. Spectrum. 2024;12(1):e0253623. doi: 10.1128/spectrum.02536-23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Suthers P. F., Dinh H. V., Fatma Z., Shen Y., Chan S. H. J., Rabinowitz J. D., Zhao H., Maranas C. D.. Genome-scale metabolic reconstruction of the non-model yeast Issatchenkia orientalis SD108 and its application to organic acids production. Metab. Eng. Commun. 2020;11:e00148. doi: 10.1016/j.mec.2020.e00148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Tran V. G., Tan S.-I., Xu H., Weilandt D. R., Li X., Bhagwat S. S., Zhu Z., Guest J. S., Rabinowitz J. D., Zhao H.. Decompartmentalization of the yeast mitochondrial metabolism to improve chemical production in Issatchenkia orientalis . Nat. Commun. 2025;16(1):7110. doi: 10.1038/s41467-025-62304-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Lu W., Bennett B. D., Rabinowitz J. D.. Analytical strategies for LC–MS-based targeted metabolomics. J. Chromatogr. B. 2008;871(2):236–242. doi: 10.1016/j.jchromb.2008.04.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Patti G. J.. Separation strategies for untargeted metabolomics. J. Sep Sci. 2011;34(24):3460–3469. doi: 10.1002/jssc.201100532. [DOI] [PubMed] [Google Scholar]
  17. Zhou B., Xiao J. F., Tuli L., Ressom H. W.. LC-MS-based metabolomics. Mol. Biosyst. 2011;8(2):470–481. doi: 10.1039/C1MB05350G. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Mahieu N. G., Patti G. J.. Systems-Level Annotation of a Metabolomics Data Set Reduces 25 000 Features to Fewer than 1000 Unique Metabolites. Anal. Chem. 2017;89(19):10397–10406. doi: 10.1021/acs.analchem.7b02380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Novoa-Del-Toro E. M., Witting M.. Navigating common pitfalls in metabolite identification and metabolomics bioinformatics. Metabolomics. 2024;20(5):103. doi: 10.1007/s11306-024-02167-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Li Y., Kind T., Folz J., Vaniya A., Mehta S. S., Fiehn O.. Spectral entropy outperforms MS/MS dot product similarity for small-molecule compound identification. Nat. Methods. 2021;18(12):1524–1531. doi: 10.1038/s41592-021-01331-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Cano P. M., Jamin E. L., Tadrist S., Bourdaud’hui P., Péan M., Debrauwer L., Oswald I. P., Delaforge M., Puel O.. New Untargeted Metabolic Profiling Combining Mass Spectrometry and Isotopic Labeling: Application on Aspergillus fumigatus Grown on Wheat. Anal. Chem. 2013;85(17):8412–8420. doi: 10.1021/ac401872f. [DOI] [PubMed] [Google Scholar]
  22. Bueschl C., Kluger B., Koutnik A., Adam G., Wiesenberger G., Maschietto V., Marocco A., Strauss J., Bödi S., Thallinger G.. et al. A novel stable isotope labeling assisted workflow for improved untargeted LC–HRMS based metabolomics research. Metabolomics. 2014;10:754–769. doi: 10.1007/s11306-013-0611-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Dong Y., Feldberg L., Aharoni A., Heinig U.. Metabolite Annotation through Stable Isotope Labeling. TrAC, Trends Anal. Chem. 2024;181:118037. doi: 10.1016/j.trac.2024.118037. [DOI] [Google Scholar]
  24. Gao Y., Luo M., Wang H., Zhou Z., Yin Y., Wang R., Xing B., Yang X., Cai Y., Zhu Z.-J.. Charting unknown metabolic reactions by mass spectrometry-resolved stable-isotope tracing metabolomics. Nat. Commun. 2025;16(1):5059. doi: 10.1038/s41467-025-60258-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Secilmis, D. ; Begzati, A. ; Grankvist, N. ; Roci, I. ; Watrous, J. ; Majithia, A. R. ; Smith, G. I. ; Klein, S. ; Jain, M. ; Nilsson, R. . Isotope tracing-based metabolite identification for mass spectrometry metabolomics bioRxiv 2025. 10.1101/2025.04.07.647691. [DOI]
  26. Mahieu N. G., Huang X., Chen Y. Jr., Patti G. J.. Credentialing Features: A Platform to Benchmark and Optimize Untargeted Metabolomic Methods. Anal. Chem. 2014;86(19):9583–9589. doi: 10.1021/ac503092d. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Wang L., Xing X., Chen L., Yang L., Su X., Rabitz H., Lu W., Rabinowitz J. D.. Peak Annotation and Verification Engine for Untargeted LC–MS Metabolomics. Anal. Chem. 2019;91(3):1838–1846. doi: 10.1021/acs.analchem.8b03132. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kind T., Fiehn O.. Metabolomic database annotations via query of elemental compositions: mass accuracy is insufficient even at less than 1 ppm. BMC Bioinf. 2006;7:234. doi: 10.1186/1471-2105-7-234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Agrawal S., Kumar S., Sehgal R., George S., Gupta R., Poddar S., Jha A., Pathak S.. El-MAVEN: A Fast, Robust, and User-Friendly Mass Spectrometry Data Processing Engine for Metabolomics. Methods Mol. Biol. 2019;1978:301–321. doi: 10.1007/978-1-4939-9236-2_19. [DOI] [PubMed] [Google Scholar]
  30. Clasquin M. F., Melamud E., Rabinowitz J. D.. LC-MS data processing with MAVEN: a metabolomic analysis and visualization engine. Curr. Protoc. Bioinformatics. 2012;Chapter 14:Unit14.11. doi: 10.1002/0471250953.bi1411s37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Schmid R., Heuckeroth S., Korf A., Smirnov A., Myers O., Dyrlund T. S., Bushuiev R., Murray K. J., Hoffmann N., Lu M.. et al. Integrative analysis of multimodal mass spectrometry data in MZmine 3. Nat. Biotechnol. 2023;41(4):447–449. doi: 10.1038/s41587-023-01690-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Smith C. A., Want E. J., O’Maille G., Abagyan R., Siuzdak G.. XCMS: Processing Mass Spectrometry Data for Metabolite Profiling Using Nonlinear Peak Alignment, Matching, and Identification. Anal. Chem. 2006;78(3):779–787. doi: 10.1021/ac051437y. [DOI] [PubMed] [Google Scholar]
  33. Gloaguen Y., Kirwan J. A., Beule D.. Deep Learning-Assisted Peak Curation for Large-Scale LC-MS Metabolomics. Anal. Chem. 2022;94(12):4930–4937. doi: 10.1021/acs.analchem.1c02220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Guo J., Shen S., Xing S., Chen Y., Chen F., Porter E. M., Yu H., Huan T.. EVA: Evaluation of Metabolic Feature Fidelity Using a Deep Learning Model Trained With Over 25000 Extracted Ion Chromatograms. Anal. Chem. 2021;93(36):12181–12186. doi: 10.1021/acs.analchem.1c01309. [DOI] [PubMed] [Google Scholar]
  35. Stancliffe E., Patti G. J.. PeakDetective: A Semisupervised Deep Learning-Based Approach for Peak Curation in Untargeted Metabolomics. Anal. Chem. 2023;95(25):9397–9403. doi: 10.1021/acs.analchem.3c00764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Patiny L., Borel A.. ChemCalc: A Building Block for Tomorrow’s Chemical Infrastructure. J. Chem. Inf. Model. 2013;53(5):1223–1228. doi: 10.1021/ci300563h. [DOI] [PubMed] [Google Scholar]
  37. Meija J.. Mathematical tools in analytical mass spectrometry. Anal Bioanal Chem. 2006;385(3):486–499. doi: 10.1007/s00216-006-0298-4. [DOI] [PubMed] [Google Scholar]
  38. Stancliffe E., Schwaiger-Haber M., Sindelar M., Patti G. J.. DecoID improves identification rates in metabolomics through database-assisted MS/MS deconvolution. Nat. Methods. 2021;18(7):779–787. doi: 10.1038/s41592-021-01195-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Chen L., Lu W., Wang L., Xing X., Chen Z., Teng X., Zeng X., Muscarella A. D., Shen Y., Cowan A.. et al. Metabolite discovery through global annotation of untargeted metabolomics data. Nat. Methods. 2021;18(11):1377–1385. doi: 10.1038/s41592-021-01303-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Pang Z., Lu Y., Zhou G., Hui F., Xu L., Viau C., Spigelman A. F., MacDonald P. E., Wishart D. S., Li S., Xia J.. MetaboAnalyst 6.0: towards a unified platform for metabolomics data processing, analysis and interpretation. Nucleic Acids Res. 2024;52(W1):W398–W406. doi: 10.1093/nar/gkae253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Qiang H., Wang F., Lu W., Xing X., Kim H., Mérette S. A. M., Ayres L. B., Oler E., AbuSalim J. E., Roichman A.. et al. Language model-guided anticipation and discovery of mammalian metabolites. Nature. 2026;651(8104):211–220. doi: 10.1038/s41586-025-09969-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Lu W., Xing X., Wang L., Chen L., Zhang S., McReynolds M. R., Rabinowitz J. D.. Improved Annotation of Untargeted Metabolomics Data through Buffer Modifications That Shift Adduct Mass and Intensity. Anal. Chem. 2020;92(17):11573–11581. doi: 10.1021/acs.analchem.0c00985. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

js6c00012_si_001.xlsx (34.3KB, xlsx)
js6c00012_si_002.pdf (399.1KB, pdf)

Articles from Journal of the American Society for Mass Spectrometry are provided here courtesy of American Chemical Society

RESOURCES