Skip to main content
NPJ Science of Food logoLink to NPJ Science of Food
. 2026 Mar 18;10:147. doi: 10.1038/s41538-026-00802-x

Tracing the origins of molecular signals in food through integrative metabolomics and chemical databases

Alejandro Mendoza Cantu 1, Julia M Gauglitz 1, Wout Bittremieux 1,✉
PMCID: PMC13149677  PMID: 41851115

Abstract

Foods contain thousands of chemical constituents beyond macronutrients, including bioactive metabolites, processing by-products, and contaminants that remain poorly characterized. The Periodic Table of Food Initiative (PTFI) is establishing a standardized global reference for food composition using untargeted mass spectrometry. We analyzed the first PTFI release ( ~ 24,000 molecular features across 500 foods) by linking annotated and unannotated signals to curated databases of pharmaceuticals, agrochemicals, food contact chemicals, and natural products. Annotated compounds revealed characteristic chemical patterns across food groups, while unannotated features exposed xenobiotic signatures and potential contamination pathways. A taxonomy-aware search identified unexpected natural products, such as biochanin A and tiliroside produced by Canada thistle. Together, these analyses show how agricultural practices, environmental exposures, and processing shape food chemistry and highlight the value of food metabolomics for advancing a One Health understanding of the molecular connections between the environment, food systems, and human health.

Subject terms: Biochemistry, Chemistry, Computational biology and bioinformatics, Environmental sciences

Introduction

Food is one of the most direct and complex interfaces between humans and their environment. It provides essential nutrients, mediates chemical exposures, and reflects agricultural and processing practices. Understanding the molecular composition of food is therefore important not only for health and nutrition, but also for assessing environmental and industrial influences on the food supply.

While foods are typically described in terms of macronutrients—such as protein, fat, and carbohydrates—this represents only one view of their full chemical complexity. In reality, foods contain thousands of additional compounds, including bioactive molecules, processing by-products, and potential contaminants, many of which remain poorly characterized1.

Despite advances in analytical chemistry, comprehensive molecular profiles are still unavailable for most foods1. This lack of information limits our ability to assess dietary exposures, evaluate food quality, and understand the broader health implications of food composition. Better molecular-level data could support more accurate nutritional recommendations, improved food safety monitoring, and the discovery of beneficial natural compounds. Existing food composition databases2–4 are often limited in scope, focusing on industrially processed or widely commercialized foods, with minimal representation of regionally important, traditional, or under-studied items.

Untargeted liquid chromatography mass spectrometry (LC-MS)-based metabolomics allows for broad detection of both known and unknown metabolites, making it especially valuable for exploring poorly characterized foods. Untargeted metabolomics initiatives like the Global FoodOmics Project5 and the Food Metabolome Repository6 yield insights into food by providing publicly available raw MS data, revealing the chemical complexity in food beyond compositional analysis.

To complement these ongoing efforts to understand what is in our food, the Periodic Table of Food Initiative (PTFI) was established7. The PTFI is a global research effort to develop standardized protocols for untargeted metabolomics analysis of foods, with the goal of building a comprehensive, publicly accessible reference of food composition. It combines high-resolution LC-MS profiling with structured metadata based on ontologies like FoodOn8, enabling scalable and interpretable foodomics analyses. As an initial release, the PTFI has made data for 500 commonly consumed foods available. These data include quantitative measurements for thousands of molecular features, covering both known reference compounds and a large set of unknown signals. As is typical in untargeted metabolomics, a substantial portion of the detected features remain unannotated, presenting an opportunity to explore their origins using external chemical databases.

In this study, we analyzed the PTFI dataset by matching both annotated and unannotated molecular features against curated databases of natural products, pharmaceuticals, agrochemicals, and food contact chemicals. This enabled us to examine the likely origins of chemical signals in food, identify potential contaminants, and detect unexpected sources of known natural products. By tracing these molecular linkages between biological, environmental, and industrial domains, our analysis contributes to a One Health understanding of how interconnected systems shape the chemical composition of the global food supply. These examples illustrate how knowledge bases can help interpret complex metabolomics signals.

Results & discussion

The PTFI untargeted LC-MS dataset includes 500 globally relevant foods like grains, fruits, vegetables, meats, and dairy collected from various U.S. locations. Plant foods are the most represented group, spanning 70% of the data, followed by animal food products (15%), and the remaining percentage corresponding to algal and fungal products. The data consist of 24,721 unique metabolic features, which were annotated against PTFI in-house reference libraries and subsequently split into structurally annotated and non-structurally annotated features for analysis, allowing us to apply a two-pronged analysis strategy (Fig. 1a). First, we cross-referenced the set of structure-annotated compounds with curated knowledge bases to infer likely functional classes and potential exogenous sources (Fig. 1b). Of the 900 structure-annotated compounds, a subset matched external databases of drugs, agrochemicals, and food contact chemicals. Furthermore, we performed a taxonomy-aware screen to identify putative novel producers of known natural products.

Fig. 1. Analysis workflow and categorization of molecular features in the PTFI dataset.

Fig. 1

a Schematic of the data analysis workflow. Structure-annotated molecules were matched by InChI key to curated databases of pharmaceuticals, agrochemicals, and food contact chemicals to infer potential sources. Formula-annotated features were screened by molecular formula against natural product databases to identify formulas not found in those references; these were classified as putative xenobiotics. b Sankey diagram illustrating the categorization of 24,721 molecular features detected across 500 foods, subdivided by chemical characteristic or source of annotation.

Second, for non-structurally annotated features, molecular formulas were provided to enable downstream matching. We used formula-level filtering to prioritize unannotated features that are unlikely to be biogenic and may represent xenobiotics (Fig. 1b). Of the 23,821 features with only formula annotations, 22,634 match a natural product molecular formula, indicating that they may have potential pharmacological, agrochemical, or contaminant origins. In addition, 1,187 features were absent from the natural product reference databases and flagged as potential xenobiotics. This xenobiotic subset was further prioritized by screening for formulas containing fluorine.

To obtain a global view of the chemical diversity captured in the PTFI dataset, we visualized all molecular formula-annotated features using a Van Krevelen diagram (Fig. 2). Plotting hydrogen-to-carbon (H/C) versus oxygen-to-carbon (O/C) ratios revealed a structured chemical space spanning multiple major metabolite superclasses, consistent with the compositional complexity of foods. A prominent high-density region centered around H/C ~ 1.5 and O/C ~ 0.4 was dominated by lipids and lipid-like molecules, reflecting the prevalence of saturated and partially unsaturated aliphatic chains typical of fatty acids, glycerolipids, and related compounds across many food matrices. In contrast, phenylpropanoids and polyketides, including flavonoids and other polyphenolic metabolites, occupied a broader region of the diagram characterized by lower H/C ratios, consistent with increased aromaticity and unsaturation. Beyond these well-defined clusters, a substantial fraction of features lacking structural annotations was distributed broadly across the Van Krevelen space. This widespread dispersion indicates that unannotated molecular formulas are not confined to a single chemical class but instead span diverse regions associated with lipids, oxygenated metabolites, and aromatic compounds. Together, these patterns underscore both the chemical richness of the food metabolome and the extent of its remaining uncharacterized molecular diversity.

Fig. 2. Van Krevelen representation of molecular features detected across 500 foods in the PTFI dataset.

Fig. 2

Hydrogen-to-carbon (H/C) and oxygen-to-carbon (O/C) ratios derived from molecular formulas reveal distinct clustering by chemical superclass, assigned by ClassyFire. Structurally annotated lipids and lipid-like molecules (blue) occupy regions of high H/C and relatively low O/C typical of aliphatic compounds, whereas phenylpropanoids and polyketides (orange) span a broader chemical space reflecting aromatic and polyphenolic structures. Gray points denote features annotated at the molecular formula level only, highlighting the extensive and chemically diverse uncharacterized component of the food metabolome.

External database matching identifies contaminants, bioactives, and therapeutic agents in foods

To obtain a high-level view of chemical classes with potential regulatory or toxicological relevance, we summed signal intensities for compounds matched to DrugBank, FCCdb, and the PubChem subset within each FoodOn level 1 category (Fig. 3a). Plant food products showed the highest aggregate signal for DrugBank-matched compounds, consistent with the breadth of phytochemicals that have direct pharmacological activity or overlap structurally with the pharmacologically relevant space9. Animal food products exhibited the next-highest DrugBank signal, which may reflect endogenous bioactive metabolites, feed-derived exposures, veterinary drugs carryover, or processing-related carryover.

Fig. 3. Analysis of annotated metabolites matched to external drug, contact chemical, and agrochemical knowledge bases.

Fig. 3

a Heatmap showing the mean intensity of features from the four principal PTFI ontologies (plant, animal, algal, fungus) across three databases: DrugBank, FCCdb, and Agrochemical PubChem. b–d Bar plots showing the signal intensity of selected compounds matched to external databases across their up to top 10 food sources for (b) the piscicide rotenone, (c) the bioactive compound phlorizin, and (d) the contact chemical pulegone. e Bar plot showing the signal intensity of rosmarinic acid in animal products.

The FCCdb-associated molecular signal was consistently high across animal, plant, and algal food products, in contrast to the much lower agrochemical matches observed across all categories. This pattern reflects both the fundamental difference in what these databases capture and their relative sizes: FCCdb contains 35 matched compounds in our dataset compared to only nine agrochemical matches. The top FCCdb matches include linoleic acid in plant products, hexadecanoic acid in animal products, and azelaic acid across multiple food categories; all naturally occurring, abundant fatty acids which are common components of food matrices10–12. While FCCdb was designed to catalog food contact chemicals and additives, it does not distinguish between natural food constituents and industrially produced chemicals.

In contrast, agrochemical matches were dominated by naturally occurring compounds with pesticidal properties (piperine, rotenone), agrochemical precursors (4-hydroxybenzoic acid), and plant growth regulators (indoleacetic acid, folic acid). The agrochemical signal was lowest in animal products, which is consistent with the expectation that most agrochemical compounds originate from plant sources due to agricultural activity or environmental contamination, rather than being intrinsic to animal food matrices.

Next, we highlight four compounds which matched to the selected chemical databases to illustrate how database-informed comparison reveals diverse possible routes of chemical entry into foods. We selected rotenone for its prominence among agrochemical matches and known health implications, phlorizin for its unexpected distribution beyond documented sources, pulegone for its dual nature as both a natural product and regulated contaminant, and rosmarinic acid as an example of plant-derived bioactives in animal products. From these cases we hypothesize intentional application, environmental contamination, and processing-related transfer.

Rotenone, a naturally occurring isoflavonoid found in non-edible parts of some plant species, has insecticidal and piscicidal applications13, and was among the most prominent agrochemical matches in the dataset. The top 10 food sources for this compound included both aquatic and terrestrial products (Fig. 3b), indicating its broad presence. Sugar kelp showed the highest mean normalized abundance, which aligns with prior exploration of rotenone as an algal crop protective agent13. In terrestrial contexts, rotenone was prevalent in vegetables and herbs/spices, consistent with its continued use on certain crops despite regulatory restrictions in some regions14. Considering that all foods in this study were procured within the United States, particularly notable was the detection of rotenone in organic oat flour; an unexpected finding, as it is a prohibited substance under U.S. National Organic Program standards15. Possible explanations include environmental drift from nearby conventional farms; persistence of residues in soil from prior applications; regulated use of fishery and wildlife management agencies to control invasive fish species in lakes, streams, and reservoirs16; or non-compliance. The documented association of rotenone exposure with Parkinson’s disease17 underscores the relevance of such observations, once validated, for food safety monitoring.

Phlorizin is a dihydrochalcone primarily reported in apples and other members of the Malus genus with reported antidiabetic, antioxidant, and anti-inflammatory activities18. While two kinds of apples are among the top foods where phlorizin was detected, the highest concentrations were actually found in anise hyssop (Fig. 3c), possibly playing an unnamed role in its use in traditional medicine19. Additionally, phlorizin was also detected in several other non-apple plant products such as allspice, pecan nut, and a number of different berries. This finding suggests either a broader biosynthetic distribution than currently described in the literature or transfer during multi-ingredient processing.

Pulegone is a cyclic monoterpenic compound primarily found in plants of the Lamiaceae family, such as mint (Mentha spp.) and thyme (Thymus spp.)20. Some regulatory bodies have defined it as a hazardous substance; for example, it is listed as a carcinogen under California’s Proposition 65. While the PTFI dataset does not include mint and thyme, it contains several other Lamiaceae herbs (basil, lemon balm, anise hyssop, and beautyberry). However, our analysis revealed that pulegone was detected exclusively in dairy-based products, specifically three types of cheese, labneh, and sour cream (Fig. 3d). Two plausible sources could explain this observation. First, pulegone’s antimicrobial properties make it an occasional component in cheese coatings designed to extend shelf life, from which it could migrate into the food21,22. A second possibility is carryover from animal feed, as it has been reported that aromatic compounds from essential oils in cattle diets, derived from pulegone-producing plants, can be transferred to milk and concentrated in cheese23. This latter pathway would also explain its presence in products not typically coated, such as labneh and sour cream.

As a final case study motivated by the search for bioactive compounds in animal sources, we present the finding of rosmarinic acid (Fig. 3e). This compound was only found in two animal products in the dataset: beef rib and sockeye salmon. This is a bioactive compound known for its beneficial effects in treating ailments like cancer and diabetes24. Its detection in fish and beef meat might be explained by its intentional addition as a natural antioxidant, as it is a common additive used to prevent lipid oxidation in meat products25. Furthermore, these are not just spurious amounts introduced as trace environmental contaminants. The levels of rosmarinic acid detected in beef (signal intensity 7,932) are higher than in some plants, such as purslane leaf (signal intensity 5,890), although there are also exceptionally high abundances observed in other plants, such as holy basil (signal intensity >5,000,000).

Together, these case studies highlight two complementary strengths of our analysis: the ability to flag potential contamination events that merit regulatory or analytical follow-up, and the capacity to uncover unexplored sources of bioactive molecules.

Presence of DrugBank features across food types

Building on the previous examples of agrochemical contamination, unexpected bioactive distributions, and food contact chemical transfer, we extended our analysis to systematically examine DrugBank therapeutic categories across food types. We selected 20 groups of interest from DrugBank’s therapeutic categories, which classify drugs based on their use in treatment and therapy. The analysis revealed plant foods dominating all categories, as expected given their phytochemical diversity (Fig. 4a). However, certain categories showed unexpected animal product signals. Anticarcinogenic agents, anticoagulants, and phytoestrogens all appeared in animal foods despite not being endemic to these matrices, pointing to potential carry-over via the food chain. Cross-referencing the annotated molecules with DrugBank therapeutic categories enabled identification of potential drug residues and bioactive compound sources in the food supply. To illustrate these transfer mechanisms, we examined three compounds that exemplify feed-to-animal pathways: 4-hydroxycoumarin (anticoagulant), daidzein (phytoestrogen), and salicylic acid (anti-inflammatory agent).

Fig. 4. Distribution of DrugBank therapeutic categories and veterinary-approved drugs across foods.

Fig. 4

a Heatmap showing the log-transformed mean intensity of compounds grouped by therapeutic category across FoodOn level 1 food groups. b–d Violin plot showing the abundance of (b) 4-hydroxycoumarin, (c) daidzein, and (d) salicylic acid across food categories.

4-hydroxycoumarin, a microbially produced anticoagulant26 as well as found as a product of fermentation and in some plant biosynthetic pathways27,28, was most commonly detected in algal food products (Fig. 4b). The elevated coumarin presence may be attributed to their photoprotective role, as coumarins have been shown to serve as UV-protective compounds in macroalgae, accumulating in cell walls and vacuolar membranes to absorb harmful radiation29. Furthermore, it appears in 97% of animal food samples (63 of 65) and 97% of plant samples (194 of 200), with comparable intensities across both food categories. Two animal products which show the highest intensity are butter and (house) cricket powder, demonstrating its presence across diverse animal-derived matrices. The presence of 4-hydroxycoumarin in animal products is notable because it is not approved as a veterinary drug, ruling out pharmaceutical administration. This nearly ubiquitous presence at similar abundance levels in both animal and plant products suggests transfer from plant-based feeds or other dietary sources, or possibly carryover during mass spectral data collection.

Daidzein, classified under DrugBank’s phytoestrogen category, showed accumulation in animal products despite its primary association with soy-based plant sources. While present in hen eggs, it is substantially accumulated in (house) cricket powder (Fig. 4c), possibly due to the fact that insects are known to efficiently absorb and metabolize dietary phenolic compounds from plant-based feeds30. Furthermore, daidzein has also been used as a supplement to improve egg quality, as demonstrated in laying hens where dietary daidzein supplementation (30 mg/kg) showed reduced cholesterol deposition in egg yolk through modulation of lipid metabolism31.

Salicylic acid, which matched DrugBank’s anti-inflammatory and analgesic categories, was detected in dairy products including skyr and sour cream (Fig. 4d). While this compound is a known anti-inflammatory agent, its presence in food is noteworthy for individuals with salicylate intolerance. The detected residues (skyr: signal intensity 1,533; sour cream: signal intensity 2,295) are comparatively low, representing only 2.5–5.0% of the levels found in salicylate-rich plant sources such as lemon balm leaf (signal intensity 60,000) or hemp seed (signal intensity 18,600). While one potential source is its natural occurrence in animal feed like alfalfa, the transfer of salicylic acid from feed to milk in ruminants is considered very limited and unlikely32. Given the relatively low concentrations, the origin of these residues remains uncertain and could include multiple pathways: natural occurrence in feed with minimal transfer, veterinary treatments such as aspirin administration, or biocide residues from processing or transfer of salicylates from fermented silage, as it has been found that salicylates may be present in silage fed to dairy cattle33.

Importantly, as this analysis is limited to compounds annotated through the PTFI reference library of authentic standards, the detection of pharmacologically relevant compounds in animal products may represent only a subset of signatures from carry-over compounds and transfer mechanisms.

Discovery of underreported bioactive natural product producers

To identify unexpected biological sources of bioactive natural products, we compared the genus of each compound in the PTFI dataset that matched pharmaceutical databases (DrugBank or PubChem Drugs) against genus-level occurrence records in the aggregated LOTUS, COCONUT, and SuperNatural databases. This procedure yielded 776 compound–genus pairs for which the producing genus was absent from all reference sources. To focus on specialized metabolites rather than ubiquitous compounds, we excluded any metabolite detected in more than 10% of all samples, retaining 84 novel producer candidates (plant genus) and 243 compound–genus pairs.

To contextualize the observation of known compounds in novel food matrices, each compound was assigned a chemical class using ClassyFire34, with the most common sample name used for identification (Fig. 5a). The results highlight several species as particularly rich in unreported compounds, such as carrot (Daucus carota) and soybean (Glycine max), with 30 and 20 compounds, respectively. While soybean shows a high total count of unreported molecules, the majority of these molecules are in one class, prenol lipids. In contrast, carrot (Daucus carota), Canada thistle (Cirsium arvense), purslane (Portulaca oleracea), and common mallow leaf (Malva sylvestris) present a more diverse novel compound content, notably containing a significant proportion of flavonoids that are not commonly reported for these foods (Fig. 5a).

Fig. 5. Analysis of novel plant natural product producers.

Fig. 5

a Top 10 species by number of unique unreported metabolites discovered, categorized according to chemical classes determined by ClassyFire. b–e Top putative novel sources for (b) biochanin A, (c) chrysin, (d) isoliquiritigenin, and (e) tiliroside, showcasing the intensity against the most abundant known producers of these compounds.

After mapping the overall diversity of the unreported screened compounds, we focused on specific examples to highlight our findings. The compounds discussed below (Fig. 5b–e) were selected both for their known bioactivity and therapeutic properties35–38, and because they represent discoveries where a novel producer showed a remarkably high abundance, often surpassing all known sources in the database.

For instance, biochanin A is an isoflavone normally found in chickpeas with diverse therapeutic applications, such as neuroprotective and anticancer properties35. Notably, anise hyssop (Agastache foeniculum) was not only identified as a novel source but also emerged as the top producer of this metabolite, with an intensity far exceeding all other foods analyzed (Fig. 5b).

Similarly, we observed a compelling pattern for isoliquiritigenin (Fig. 5d), a chalcone with antitumoral and antioxidant activities36. The two highest concentrations were found in the pulp of two passion fruit (Passiflora edulis) variants. This consistency across samples from the same botanical family strengthens the evidence that these are genuine biosynthetic findings rather than experimental artifacts.

A similar trend was observed for chrysin, a flavonoid typically associated with honey and bee pollen, where hepatoprotective and antidiabetic properties have been reported37. In the PTFI data, the highest intensities were detected in two parts of the same plant, the flower and leaf of Canada thistle (Cirsium arvense), with the concentration in the flower being considerably higher than in the leaf (Fig. 5c).

Finally, a comparable trend was observed for tiliroside, a dietary flavonoid with documented antiproliferative and antimicrobial activities38. While tiliroside is commonly associated with strawberries, Canada thistle (Cirsium arvense) emerged as the top producer in our dataset, with additional detections in peach (Prunus persica) and common mallow leaf (Malva sylvestris) (Fig. 5e). The consistency of Canada thistle as a high producer across multiple bioactive compounds and high count of unreported discoveries suggests that this species may represent an underexplored source of therapeutic natural products.

Molecular formula annotated features reveal putative xenobiotics

The unannotated portion of the dataset comprised ~24,000 molecular features for which only elemental formulas were available. We screened these formulas against the combined reference space of LOTUS, COCONUT, and SuperNatural to identify candidates unlikely to originate from known natural products (Fig. 1). A total of 1187 features (4.8%) were absent from the natural product formula set and were therefore classified as putative xenobiotics. To further enrich for likely anthropogenic compounds, we applied an elemental filter targeting formulas containing fluorine. Fluorine-containing natural products are exceptionally rare and thus strong candidates for anthropogenic xenobiotics in food matrices, as they are prevalent in synthetic agrochemicals, industrial additives, and persistent pollutants39.

To characterize molecular features whose elemental formulas did not match entries in natural product reference databases, we examined their degree of unsaturation, quantified as double bond equivalents (DBE), together with their heteroatom composition (Fig. 6). These analyses provide insight into the chemical nature of the putative xenobiotic feature pool beyond structural annotation. The DBE distribution of non-natural formulas exhibited a peak at DBE = 0 (Fig. 6a), indicating fully saturated molecular structures. Fluorinated features were particularly concentrated in this low-DBE region, consistent with the physicochemical properties of highly saturated fluorinated compounds, including perfluoroalkyl substances and related chemistries40. Analysis of heteroatom content further supported a xenobiotic-associated chemical signature (Fig. 6b). Among the 1,187 formulas classified as non-natural, 213 contained halogens, with fluorine being the most prevalent (126 features), followed by chlorine (77 features) and iodine (5 features). In addition, sulfur (125 features) and phosphorus (102 features) were frequently observed. The enrichment of these heteroatoms is consistent with chemical motifs commonly found in agrochemicals, pharmaceuticals, and industrial processing agents41,42, reinforcing the interpretation that this subset of features likely reflects environmental or anthropogenic contributions to the food metabolome.

Fig. 6. Degree of unsaturation and heteroatom composition of non-natural molecular formulas.

Fig. 6

a Distribution of double bond equivalents (DBE) for molecular formulas classified as non-natural (gray), with fluorinated features highlighted (blue). b Heteroatom distribution among non-natural features, showing frequent occurrence of halogens (F, Cl, I) as well as sulfur and phosphorus, consistent with chemical signatures of agrochemicals and other anthropogenic compounds.

The distribution of fluorine-containing xenobiotics across FoodOn level 1 categories revealed distinct xenobiotic fingerprints for major food groups (Fig. 7a). Within the animal category, subcategory-level clustering uncovered profiles for specific product types (Fig. 7b). For example, there is a distinct grouping of dairy products characterized by a unique combination of fluorinated formulas. At the level of molecular formula annotation it is not possible to group these features into chemical categories, though several potential sources likely contribute to these signals. The fluorinated signals observed could originate from veterinary pharmaceuticals and pesticide residues used in dairy production. Additionally, a class of fluorinated compounds that have reported health consequences and which are now ubiquitous in our environment are per- and polyfluoroalkyl substances (PFAS), which may be a contributing factor to these detections. While there have been reports of PFAS contamination in dairy products43, the distinct fluorine patterns in these dairy products indicate that comprehensive screening may be necessary to fully assess xenobiotic exposure from dairy consumption.

Fig. 7. Fluorine-containing features reveal the presence of potential xenobiotic unknowns.

Fig. 7

a Clustered heatmap of top 10 fluorine-containing putative xenobiotics for each FoodOn level 1 category. b Clustered heatmap of top 10 fluorine-containing putative xenobiotics for each FoodOn level 2 animal product categories. A distinct group of dairy foods is indicated by the red box. c Clustered heatmap of top 10 fluorine-containing putative xenobiotics within dairy products. Hard cheeses are marked with an asterisk (*).

A deeper analysis demonstrated that even within this specific set of dairy samples, xenobiotic profiles can be highly differentiated (Fig. 7c). Hard, aged cheeses share a core set of three recurring fluorinated features (C33H35FN2O5, C21H23ClFNO2, C23H27FN4O2) that were absent or low-abundance in fresh cheeses, suggesting that aging processes, protective coatings, or packaging materials used for hard cheeses may introduce specific fluorinated contaminants not found in fresh dairy products.

One prominent feature in the dairy cluster had the formula C8F17KO3S, potentially corresponding to the potassium salt of perfluorooctanesulfonic acid (PFOS), a well-documented PFAS contaminant in milk43. This feature was consistently detected in milk but appeared at higher intensities in aged cheeses. The increase may be a concentration effect driven by the higher protein and fat content of cheese; the amphiphilic nature of PFOS causes it to bind to fats and proteins44, which become significantly more concentrated as water is removed during maturation. Inspection of the compositional PTFI dataset, which accompanies the metabolomics data, shows that a cheese can greatly increase the protein content from 3.37 g/100 g in milk to over 22.90 g/100 g in Gouda45,46.

In conclusion, by matching the first release of the PTFI’s food metabolomics data with multiple external chemical and natural product knowledge bases, we sought to understand the origin of molecular signals in food data. Our complementary analyses highlight the interplay of agricultural practices, environmental exposures, biological biosynthesis, and industrial processing in shaping the molecular composition of foods. This systems-level perspective aligns with the One Health concept, which emphasizes the interdependence of human, animal, and environmental health. By tracing chemical linkages across these domains, food metabolomics can contribute to identifying routes of exposure, monitoring contaminants, and uncovering beneficial compounds that connect agricultural and public health outcomes. Together, these findings demonstrate how curated chemical occurrence information can be used to identify contamination pathways and discover new food sources of bioactive compounds with potential nutritional or therapeutic value, supporting One Health-oriented efforts to build safer and more sustainable food systems.

Limitations of the study

This study has several important limitations that should be considered when interpreting the results. Most fundamentally, the analyses are constrained by the availability of MS1-only LC-MS data, as MS/MS data were not available in the current public release of the PTFI dataset. While a subset of compounds was annotated using accurate mass and retention time matching to authentic reference standards, the majority of detected features were annotated at the molecular formula level rather than at the structural level. As a result, structural isomers sharing the same elemental composition cannot be distinguished, and definitive chemical identities cannot be assigned without additional fragmentation or orthogonal analytical evidence. Consequently, grouping features into specific chemical subclasses based solely on formula information is not possible. For example, fluorine-containing signals observed in dairy products could plausibly correspond to veterinary pharmaceuticals, fluorinated agrochemicals, food contact materials, or PFAS. Discriminating among these possibilities requires structural information that cannot be resolved from MS1 data alone. All formula-level interpretations should therefore be regarded as putative.

Interpretation of the data further depends on the coverage and biases of reference databases used for contextualization, including LOTUS, COCONUT, SuperNatural, and DrugBank. These resources vary in scope, curation depth, and chemical focus, and none provide comprehensive or unbiased coverage of the chemical space present in foods. As a result, classification of features as “natural products,” “xenobiotics,” or “novel” reflects current database and literature coverage rather than definitive biological absence or presence. In particular, the identification of “novel producers” should be interpreted cautiously, as a compound may appear unreported for a given genus simply because that food source has not previously been studied using untargeted metabolomics or has been underrepresented in the literature.

The study also relies on a single, standardized extraction protocol applied uniformly across all 500 foods. This approach was chosen to maximize comparability and scalability; however, no universal extraction method can efficiently recover all chemical classes present across highly diverse food matrices. Certain compound classes may therefore be preferentially extracted or underrepresented, introducing systematic bias into the detected chemical profiles.

Quantitative interpretation represents an additional limitation. The reported signal intensities are semi-quantitative and reflect relative abundances rather than absolute concentrations. Comparisons of the same molecular feature across different foods are valid within the dataset, but direct comparisons between different compounds should be interpreted with caution, owing to compound-specific differences in extraction efficiency, ionization behavior, and matrix effects. Statements regarding compound abundance therefore describe relative patterns rather than absolute exposure estimates.

Finally, while the dataset includes a diverse collection of 500 foods, it represents only a subset of global food biodiversity. Geographic origin, agricultural practices, food processing, seasonal variation, and supply-chain factors may all influence chemical composition and are not comprehensively captured here. Accordingly, the results should be interpreted as illustrative of broad chemical trends rather than as an exhaustive characterization of the global food supply.

Taken together, these limitations underscore that the present study is exploratory and hypothesis-generating. Its primary contribution lies in demonstrating how large-scale untargeted metabolomics data can be leveraged to reveal global patterns in food-associated chemical space. Future studies incorporating MS/MS data, targeted validation, and quantitative measurements will be essential to refine and validate the molecular interpretations presented here.

Methods

Data acquisition

We analyzed the first public release of the PTFI dataset, which includes untargeted metabolomics measurements for 500 commonly consumed food items7. Each sample is accompanied by structured metadata describing its source organism, production method, and food classification, the latter organized hierarchically using the FoodOn ontology8.

Sample preparation and data collection followed the standardized untargeted metabolomics protocol developed by the PTFI47,48. Briefly, foods were lyophilized and homogenized to create uniform composite specimens. Aliquots of 55 ± 5 mg were extracted with 80% methanol:water, followed by lipid removal. This extraction method was selected based on its effectiveness across different food types, allowing efficient capture of a diverse set of chemical classes. A custom internal retention time standard mixture of 33 compounds non-endogenous to food was added to enable chromatographic alignment across laboratories. Extracts were analyzed by reverse-phase liquid chromatography coupled to high-resolution mass spectrometry. For the 500 foods dataset analyzed in this study, data were acquired using an Agilent 6560 ion mobility quadrupole time-of-flight (IM-qTOF) mass spectrometer, resulting in 24,721 metabolic features defined by their mass-to-charge ratio (m/z) and chromatographic retention time index. All intensity values were normalized within each injection to the median intensity of the 33 internal standards.

Feature detection and molecular formula assignment

Feature detection and molecular formula assignment were performed by the PTFI team following the PTFI Metabolomics v1 data processing pipeline49. Briefly, chromatographic feature detection was conducted using XCMS50 with the centWave algorithm51. Retention times were converted to retention indices using a second-order polynomial regression fitted to internal retention time standard signals with a Tukey M-estimator. Redundant signals arising from isotopes and common adducts were removed using IsoSpecPy52,53.

Molecular formula assignment was performed against a consensus reference library constructed from compounds and molecular formulas curated from FooDB and additional chemical databases, subject to defined elemental constraints. Common adduct forms ([M + H]+, [M + H-H2O]+, [M+Na]+, [M + NH4]+, [M + K]+ for positive mode; and [M-H]-, [M-H-CHOO]-, [M-H-H2O]-, [M-Cl]-, [M+Na-2H]-, [M + K-2H]- for negative mode) were included in the molecular search space. Library inclusion required observation of a given signal in two independent laboratories, and isotopic pattern verification was used to reduce formula ambiguity. Experimental features were matched to library entries using a composite scoring scheme incorporating mass difference, retention index difference, isotopic pattern similarity, adduct preference, and monoisotopic feature intensity, applying a 15 ppm mass tolerance.

Two distinct annotation confidence levels can be assigned in accordance with Metabolomics Standards Initiative (MSI) guidelines (Supplementary Data S1)54. First, 900 features were structurally annotated by matching experimental signals to authentic reference standards using accurate mass, retention index, and isotopic envelope agreement. These annotations rely on two orthogonal analytical properties and therefore correspond to MSI Level 1 (confirmed identification). Note that retention time was not provided in the PTFI dataset for these compounds. All remaining features ( ~ 23,800) were annotated at the molecular formula level only, based on accurate mass and isotopic pattern matching without confirmation by reference standards or MS/MS fragmentation data. These annotations do not distinguish between structural isomers and are therefore designated as MSI Level 3 (putative compound / formula-level annotation). These formula-level annotations form the basis for downstream chemical space, elemental composition, and database membership analyses but are not interpreted as definitive structural identifications.

All subsequent data analysis was conducted within the PTFI online runtime environment hosted by the American Heart Association’s Precision Medicine Platform. The data were originally provided as a wide-format abundance matrix with features as rows and samples as columns. We converted this into a long-format table, where each row corresponds to a metabolite–sample measurement. No raw MS data were directly available, and MS/MS data were not available for this PTFI release.

Database integration and matching strategy

To interpret the chemical origin of observed features, we cross-referenced annotated compounds with DrugBank (v6.0)55 for pharmaceuticals, the Food Contact Chemicals Database (FCCdb; v5.0)56 for packaging-related substances, and a filtered subset of PubChem (v2025)57 entries corresponding to agrochemicals (e.g. pesticides and fungicides, based on the “Agrochemicals” table of contents category) and pharmaceuticals (based on Medical Subject Headings (MeSH) classifications). 93 compounds matched to DrugBank, corresponding to 400 unique therapeutic category terms. To focus on biologically relevant functions, we excluded non-therapeutic classifications such as chemical descriptors and industrial uses. From the remaining terms, we selected 20 broad therapeutic categories (e.g., Enzyme Inhibitors, Antioxidants) that capture diverse pharmacological activities. To assess potential biological origins of known compounds, we further integrated three natural product databases: LOTUS (NPOC2021, February 2021)58, COCONUT (v2.0)59, and SuperNatural (v3.0)60. Matching was performed using InChI keys for annotated compounds and molecular formulas for unannotated features.

Compound origin analyses

To identify potential novel sources of known natural products, compounds were matched to food samples based on genus-level taxonomy. The genus was extracted from each sample’s organism name metadata field and compared against the genus-level organism annotations in the three natural product databases (LOTUS, COCONUT, SuperNatural). A compound was flagged as a potential novel discovery if the plant genus of the food sample was not among the known producers listed in any of the external databases. To focus on rare or specialized metabolites, compounds detected in more than 10% of all plant samples were excluded. The final output was a list of compound–genus pairs not previously reported in available natural product literature.

We used ClassyFire (version 1.0)34, a hierarchical chemical taxonomy system that assigns compounds to categories based on their structural features, to classify compounds for the novel producer analysis. Each compound was classified using its InChI.

H/C and O/C ratios were calculated from molecular formulas for all features. Compounds with structural annotations were classified into chemical superclasses using ClassyFire (version 1.0)34. The five most abundant superclasses were visualized in the Van Krevelen diagram (Fig. 2).

For features that remained unannotated, i.e. those for which only the molecular formula was known, a formula-based screen was used to identify putative xenobiotics. Molecular formulas for 23,821 such features were compared to the union of formulas from the three natural product databases (LOTUS, COCONUT, SuperNatural). Any formula not found in this reference set was classified as a putative xenobiotic. To further prioritize likely contaminants, we filtered for formulas containing fluorine (F), as this element is uncommon in biosynthetic products but frequently observed in industrial and agricultural compounds. DBEs were calculated from molecular formulas using the standard formula DBE = C − (H + X)/2 + N/2 + 1, where X represents halogens (F, Cl, Br, I); formulas containing elements incompatible with this calculation were excluded. Heteroatom counts (S, P, F, Cl, Br, I) were extracted directly from parsed formulas. The resulting list provided a set of candidate xenobiotic features for further analysis and interpretation.

Supplementary information

Supplemental Information (1.1MB, xlsx)

Acknowledgements

This work was conducted using data provided through the PTFI 2024 Data Challenge. We thank the PTFI team and the American Heart Association’s Precision Medicine Platform for enabling access to the dataset and analysis environment. This work was supported by the Research Foundation – Flanders (FWO G0AGQ24N).

Author contributions

A.M.C.: conceptualization, methodology, software, investigation, data curation, writing - original draft, writing - review & editing, visualization. J.M.G.: conceptualization, methodology, writing - original draft, writing - review & editing, supervision, funding acquisition. W.B.: conceptualization, methodology, writing - original draft, writing - review & editing, supervision, funding acquisition.

Data availability

The PTFI dataset analyzed in this study was accessed through the American Heart Association Precision Medicine Platform at https://pmp.heart.org/ptfi-data-resource, with version v2_2024_08. PTFI data and analysis tools are provided under the CC-BY-NC-4.0 license. External compound databases were retrieved from their corresponding web resources: LOTUS (vNPOC2021; https://lotus.naturalproducts.net/), COCONUT (v2.0; https://coconut.naturalproducts.net/), SuperNatural (v3.0; http://bioinf-applied.charite.de/supernatural_3/), DrugBank (v6.0; https://www.drugbank.ca/), Food Contact Chemicals Database (v5.0; https://foodpackagingforum.org/resources/databases/fccdb) and PubChem (v2025; https://pubchem.ncbi.nlm.nih.gov/).

Code availability

All analyses were conducted in Python 3.12 within the online PTFI data analysis platform (https://pmp.heart.org/). Data processing was performed using the Pandas (v2.2.3)61,62 and Scikit-Learn (v1.7.0)63 libraries. Figures were generated using Matplotlib (v3.9.3)64 and Seaborn (v0.13.2)65. Analysis code, including notebooks for data processing, database matching, and figure generation, is available as open source at https://github.com/AlejandroMC28/ptfi_externaldb.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors jointly supervised this work: Julia M. Gauglitz, Wout Bittremieux.

Supplementary information

The online version contains supplementary material available at 10.1038/s41538-026-00802-x.

References

  • 1.Barabási, A.-L., Menichetti, G. & Loscalzo, J. The unmapped chemical complexity of our diet. Nat. Food1, 33–37 (2019). [Google Scholar]
  • 2.FooDB. FR-0015. https://foodb.ca/.
  • 3.USDA FoodData Central. https://fdc.nal.usda.gov/.
  • 4.Food composition data | EFSA. https://www.efsa.europa.eu/en/data-report/food-composition-data (2021).
  • 5.Gauglitz, J. M. et al. Untargeted mass spectrometry-based metabolomics approach unveils molecular changes in raw and processed foods and beverages. Food Chem302, 125290 (2020). [DOI] [PubMed] [Google Scholar]
  • 6.Sakurai, N. et al. The Thing Metabolome Repository family (XMRs): comparable untargeted metabolome databases for analyzing sample-specific unknown metabolites. Nucleic Acids Res51, D660–D677 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Jarvis, A. et al. Periodic Table of Food Initiative for generating biomolecular knowledge of edible biodiversity. Nat. Food5, 189–193 (2024). [DOI] [PubMed] [Google Scholar]
  • 8.Dooley, D. M. et al. FoodOn: a harmonized food ontology to increase global food traceability, quality control and data integration. Npj Sci. Food2, 23 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chihomvu, P., Ganesan, A., Gibbons, S., Woollard, K. & Hayes, M. A. Phytochemicals in Drug Discovery—A Confluence of Tradition and Innovation. Int. J. Mol. Sci.25, 8792 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Showing Compound Azelaic acid (FDB012192) - FooDB. https://foodb.ca/compounds/FDB012192.
  • 11.Showing Compound alpha-Linoleic acid (FDB006287) - FooDB. https://foodb.ca/compounds/FDB006287.
  • 12.Showing Compound Hexadecanoic acid (FDB011679) - FooDB. https://foodb.ca/compounds/FDB011679.
  • 13.El-Sayed, W. M. M. et al. Environmental influence on rotenone performance as an algal crop protective agent to prevent pond crashes for biofuel production. Algal Res33, 277–283 (2018). [Google Scholar]
  • 14.Gupta, R. C. & Milatovic, D. Insecticides. in Biomarkers in Toxicology 389–407 (Elsevier, 2014). 10.1016/B978-0-12-404630-6.00023-3.
  • 15.Federal Register: National Organic Program; Amendments to the National List of Allowed and Prohibited Substances (Crops, Livestock and Handling).
  • 16.Finlayson, B. J. Planning and Standard Operating Procedures for the Use of Rotenone in Fish Management: Rotenone SOP Manual. (American Fisheries Society, Bethesda, Maryland, USA, 2018).
  • 17.Van Laar, A. D. et al. Transient exposure to rotenone causes degeneration and progressive parkinsonian motor deficits, neuroinflammation, and synucleinopathy. Npj Park. Dis.9, 121 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Ehrenkranz, J. R. L., Lewis, N. G., Ronald Kahn, C. & Roth, J. Phlorizin: a review. Diabetes Metab. Res. Rev.21, 31–38 (2005). [DOI] [PubMed] [Google Scholar]
  • 19.Duda, M. M., Varban, D. I., Muntean, S., Moldovan, C. & Olar, M. Use of species Agastache foeniculum (Pursh) Kuntze. Hop Med. Plants21, 52–54 (2014). [Google Scholar]
  • 20.Soleimani, M., Arzani, A., Arzani, V. & Roberts, T. H. Phenolic compounds and antimicrobial properties of mint and thyme. J. Herb. Med.36, 100604 (2022). [Google Scholar]
  • 21.Shahdadi, F. et al. Mentha longifolia Essential Oil and Pulegone in Edible Coatings of Alginate and Chitosan: Effects on Pathogenic Bacteria in Lactic Cheese. Molecules28, 4554 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Voigt, V., Franke, H. & Lachenmeier, D. W. Risk Assessment of Pulegone in Foods Based on Benchmark Dose–Response Modeling. Foods13, 2906 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Faehnrich, B., Chizzola, R., Schabauer, A., Pracser, N. & Duerrschmid, K. Volatiles in dairy products after supplementation of essential oils in the diet of cows and influence on taste of cheese. Eur. Food Res. Technol.243, 1783–1797 (2017). [Google Scholar]
  • 24.Singh, D., Kumari, K. & Ahmed, S. Natural herbal products for cancer therapy. in Understanding Cancer 257–268 (Elsevier, 2022). 10.1016/B978-0-323-99883-3.00010-X.
  • 25.Lorenzo, J. M. et al. Preservation of meat products with natural antioxidants from rosemary. IOP Conf. Ser. Earth Environ. Sci.854, 012053 (2021). [Google Scholar]
  • 26.Lin, Y., Shen, X., Yuan, Q. & Yan, Y. Microbial biosynthesis of the anticoagulant precursor 4-hydroxycoumarin. Nat. Commun.4, 2603 (2013). [DOI] [PubMed] [Google Scholar]
  • 27.Bye, A. & King, H. K. The biosynthesis of 4-hydroxycoumarin and dicoumarol by Aspergillus fumigatus Fresenius. Biochem. J.117, 237–245 (1970). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Liu, B., Raeth, T., Beuerle, T. & Beerhues, L. A novel 4-hydroxycoumarin biosynthetic pathway. Plant Mol. Biol.72, 17–25 (2010). [DOI] [PubMed] [Google Scholar]
  • 29.Perez-Rodriguez, E., Aguilera, J. & Figueroa, F. L. Tissular localization of coumarins in the green alga Dasycladus vermicularis (Scopoli) Krasser: a photoprotective role?. J. Exp. Bot.54, 1093–1100 (2003). [DOI] [PubMed] [Google Scholar]
  • 30.Torres-Castillo, J. A. & Olazarán-Santibáñez, F. E. Insects as source of phenolic and antioxidant entomochemicals in the food industry. Front. Nutr.10, 1133342 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Liu, J. et al. Effects of quercetin and daidzein on egg quality, lipid metabolism, and cecal short-chain fatty acids in layers. Front. Vet. Sci.10, 1301542 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Houben, K. et al. Salicylic acid residues in products of animal origin. Food Risk Assess Eur.2, FR-0015 (2024). [Google Scholar]
  • 33.Xu, D. et al. The bacterial community and metabolome dynamics and their interactions modulate fermentation process of whole crop corn silage prepared with or without inoculants. Microb. Biotechnol.14, 561–576 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Djoumbou Feunang, Y. et al. ClassyFire: automated chemical classification with a comprehensive, computable taxonomy. J. Cheminformatics8, 61 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Feng, Z.-J. & Lai, W.-F. Chemical and Biological Properties of Biochanin A and Its Pharmaceutical Applications. Pharmaceutics15, 1105 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Chen, Z., Ding, W., Yang, X., Lu, T. & Liu, Y. Isoliquiritigenin, a potential therapeutic agent for treatment of inflammation-associated diseases. J. Ethnopharmacol.318, 117059 (2024). [DOI] [PubMed] [Google Scholar]
  • 37.Naz, S. et al. Chrysin: Pharmacological and therapeutic properties. Life Sci235, 116797 (2019). [DOI] [PubMed] [Google Scholar]
  • 38.Grochowski, D. M., Locatelli, M., Granica, S., Cacciagrano, F. & Tomczyk, M. A Review on the Dietary Flavonoid Tiliroside. Compr. Rev. Food Sci. Food Saf.17, 1395–1421 (2018). [DOI] [PubMed] [Google Scholar]
  • 39.Petkowski, J. J., Seager, S. & Bains, W. Reasons why life on Earth rarely makes fluorine-containing compounds and their implications for the search for life beyond Earth. Sci. Rep.14, 15575 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Evich, M. G. et al. Per- and polyfluoroalkyl substances in the environment. Science375, eabg9065 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Wahab, S. et al. Advancement and New Trends in Analysis of Pesticide Residues in Food: A Comprehensive Review. Plants11, 1106 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kariyanna, B. et al. Comprehensive insights into pesticide residue dynamics: unraveling impact and management. Chem. Biol. Technol. Agric.11, 182 (2024). [Google Scholar]
  • 43.Curci, D., Sundaram, T. S., Ghidini, S. & Arioli, F. What We Know About per- and Polyfluoroalkyl Contamination Levels in Milk. A Review from the Last Decade. Foods14, 2274 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Bonato, T., Pal, T., Benna, C. & Di Maria, F. Contamination of the terrestrial food chain by per- and polyfluoroalkyl substances (PFAS) and related human health risks: A systematic review. Sci. Total Environ.961, 178337 (2025). [DOI] [PubMed] [Google Scholar]
  • 45.Food Detail - gouda cheese - PTFI. https://ptfi.markerlab.com/detail/food/GGB100362.
  • 46.Food Detail - cow milk (vitamin D, pasteurized, homogenized) - PTFI. https://ptfi.markerlab.com/detail/food/GGB100035.
  • 47.Odenkirk, M. T. et al. A standardized nontargeted metabolomics method for cross-laboratory comparison of food profiles. Food Chem493, 145934 (2025). [DOI] [PubMed] [Google Scholar]
  • 48.Odenkirk, Melanie T. Nontargeted Metabolomics (Reverse Phase) - Analysis of Foods Standard Operating Procedure (SOP). https://hdl.handle.net/10568/179764 (2025).
  • 49.Chien, et al. Metabolomics: V1 Data Processing Pipeline. (2025).
  • 50.Smith, C. A., Want, E. J., O’Maille, G., Abagyan, R. & Siuzdak, G. XCMS: Processing Mass Spectrometry Data for Metabolite Profiling Using Nonlinear Peak Alignment, Matching, and Identification. Anal. Chem.78, 779–787 (2006). [DOI] [PubMed] [Google Scholar]
  • 51.Tautenhahn, R., Böttcher, C. & Neumann, S. Highly sensitive feature detection for high resolution LC/MS. BMC Bioinformatics9, 504 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Łącki, M. K., Valkenborg, D. & Startek, M. P. IsoSpec2: Ultrafast Fine Structure Calculator. Anal. Chem.92, 9472–9475 (2020). [DOI] [PubMed] [Google Scholar]
  • 53.Łącki, M. K., Startek, M., Valkenborg, D. & Gambin, A. IsoSpec: Hyperfast Fine Structure Calculator. Anal. Chem.89, 3272–3277 (2017). [DOI] [PubMed] [Google Scholar]
  • 54.Sumner, L. W. et al. Proposed minimum reporting standards for chemical analysis: Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI). Metabolomics3, 211–221 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Knox, C. et al. DrugBank 6.0: the DrugBank Knowledgebase for 2024. Nucleic Acids Res52, D1265–D1275 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Groh, K., Geueke, B. & Muncke, J. FCCdb: Food Contact Chemicals database. Version 5.0. Zenodo 10.5281/ZENODO.4296944 (2020).
  • 57.Kim, S. et al. PubChem 2025 update. Nucleic Acids Res53, D1516–D1525 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Rutz, A. et al. The LOTUS initiative for open knowledge management in natural products research. eLife11, e70780 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Chandrasekhar, V. et al. COCONUT 2.0: a comprehensive overhaul and curation of the collection of open natural products database. Nucleic Acids Res53, D634–D643 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Gallo, K. et al. SuperNatural 3.0—a database of natural products and natural product-based derivatives. Nucleic Acids Res51, D654–D659 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.The pandas development team. pandas-dev/pandas: Pandas. Zenodo 10.5281/ZENODO.13819579 (2024).
  • 62.McKinney, W. Data Structures for Statistical Computing in Python. in 56–61 (Austin, Texas, 2010). 10.25080/Majora-92bf1922-00a.
  • 63.Pedregosa, F. et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
  • 64.Hunter, J. D. Matplotlib: A 2D Graphics Environment. Comput. Sci. Eng.9, 90–95 (2007). [Google Scholar]
  • 65.Waskom, M. seaborn: statistical data visualization. J. Open Source Softw.6, 3021 (2021). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Information (1.1MB, xlsx)

Data Availability Statement

The PTFI dataset analyzed in this study was accessed through the American Heart Association Precision Medicine Platform at https://pmp.heart.org/ptfi-data-resource, with version v2_2024_08. PTFI data and analysis tools are provided under the CC-BY-NC-4.0 license. External compound databases were retrieved from their corresponding web resources: LOTUS (vNPOC2021; https://lotus.naturalproducts.net/), COCONUT (v2.0; https://coconut.naturalproducts.net/), SuperNatural (v3.0; http://bioinf-applied.charite.de/supernatural_3/), DrugBank (v6.0; https://www.drugbank.ca/), Food Contact Chemicals Database (v5.0; https://foodpackagingforum.org/resources/databases/fccdb) and PubChem (v2025; https://pubchem.ncbi.nlm.nih.gov/).

All analyses were conducted in Python 3.12 within the online PTFI data analysis platform (https://pmp.heart.org/). Data processing was performed using the Pandas (v2.2.3)61,62 and Scikit-Learn (v1.7.0)63 libraries. Figures were generated using Matplotlib (v3.9.3)64 and Seaborn (v0.13.2)65. Analysis code, including notebooks for data processing, database matching, and figure generation, is available as open source at https://github.com/AlejandroMC28/ptfi_externaldb.


Articles from NPJ Science of Food are provided here courtesy of Nature Publishing Group

RESOURCES