Skip to main content
EPA Author Manuscripts logoLink to EPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Feb 9.
Published in final edited form as: J Breath Res. 2021 Mar 18;15(2):025001. doi: 10.1088/1752-7163/abdb03

Using the US EPA CompTox Chemicals Dashboard to interpret targeted and non-targeted GC–MS analyses from human breath and other biological media

Joachim D Pleil 1,*, Charles N Lowe 2, M Ariel Geer Wallace 3, Antony J Williams 2
PMCID: PMC12883259  NIHMSID: NIHMS2137398  PMID: 33734097

Abstract

The U.S. EPA CompTox Chemicals Dashboard is a freely available web-based application providing access to chemistry, toxicity, and exposure data for ~900 000 chemicals. Data, search functionality, and prediction models within the Dashboard can help identify chemicals found in environmental analyses and human biomonitoring. It was designed to deliver data generated to support computational toxicology to reduce chemical testing on animals and provide access to new approach methodologies including prediction models. The inclusion of mass and formula-based searches, together with relevant ranking approaches, allows for the identification and prioritization of exogenous (environmental) chemicals from high resolution mass spectrometry in need of further evaluation. The Dashboard includes chemicals that can be detected by liquid chromatography, gas chromatography–mass spectrometry (GC–MS) and direct-MS analyses, and chemical lists have been added that highlight breath-borne volatile and semi-volatile organic compounds. The Dashboard can be searched using various chemical identifiers (e.g. chemical synonyms, CASRN and InChIKeys), chemical formula, MS-ready formulae monoisotopic mass, consumer product categories and assays/genes associated with high-throughput screening data. An integrated search at a chemical level performs searches against PubMed to identify relevant published literature. This article describes specific procedures using the Dashboard as a first-stop tool for exploring both targeted and non-targeted results from GC–MS analyses of chemicals found in breath, exhaled breath condensate, and associated aerosols.

Keywords: non-targeted analysis, compound identification, chemical dashboard, InChIKeys, mass and formula search

1. Introduction

Biomarkers research relies on analysis of compounds in biological media (i.e. serum, urine, biological tissue, expired breath, saliva, etc) and subsequent interpretation of relative amounts. One challenge is to unambiguously identify chemicals within complex media; this is almost always subject to some level of error as there are often hundreds to thousands of components in a sample, and many are not easily separated by chromatographic techniques, resulting in incomplete chromatographic separation and complex mass spectra. To simplify the mixture, analytical experiments are often designed to optimize detection of a smaller set of compounds of interest (targeted analysis). As the list of potentially important compounds is expanded, or if the analyses are non-targeted (widened to include all detectable analytes in a sample), this task becomes almost impossible without the application of advanced informatics platforms. Researchers can go through a variety of external web-based resources to come up with information to help identify constituents in their samples, or they could stay within the proprietary software of their instrumentation, primarily vendor-provided software. However, searching among resources on the internet is time consuming, and one may encounter inaccurate information. Staying within vendor/instrument-provided software may be overly restrictive, and thus may not provide a sufficiently broad search useful for comparing among other analytical platforms. Although modern analytical systems can provide suggestions as to the chemical group, volatility, or expected formula of a constituent based on either chromatography and/or mass spectrometry, potential chemical candidates for a specific feature may range from a few to hundreds of possibilities. As such, additional information provided by an appropriate informatics structure is crucial for reliable and unambiguous chemical identification. As will be demonstrated, the U.S. EPA CompTox Chemicals Dashboard web-based application is a powerful tool for facilitating chemical identifications in non-targeted analyses of complex biological matrices (Williams et al 2017, Pleil and Isaacs 2016 and Pleil et al 2016).

In this tutorial article, we discuss how the Dashboard can be used to answer questions that often arise when dealing with the interpretation of new mass spectrometry data. The focus of this article is on gas chromatography–mass spectrometry (GC–MS) questions arising from breath analyses, but the approaches described could be applied to other techniques and biological media as well. For background, we first introduce some of the underlying concepts that are featured herein.

1.1. Targeted approach

Traditional biomarkers research implements a targeted compound approach generally referred to as ‘bottom-up’ analysis wherein specific chemicals known to be in the environment, or where specific metabolites known to be a response to stressors, are deliberately sought in environmental or biological media. The analyst uses quality control criteria to confirm the identity of each detected compound, such as by analyzing authentic chemical standards to check retention times, confirm exact masses, and match isotopic patterns of the spectra for each compound on a particular instrument. Although the chemical standards are generally very pure, the signal for the detected compound may include overlapping signals from the total chemical space of real-world samples that confound identification. Despite the a priori knowledge gained from standards, unambiguous identification can therefore become a challenge in complex samples. Post-processing and subject matter expertise are required.

Some examples of targeted breath analyses include the article by Wallace et al (1987) that focused on 20 volatile organic compounds (VOCs) in a crosssectional study of environmental exposures and breath in three cities. Pleil et al (2000) focused on a series of n-alkanes and single-ring aromatics from the breath of military fuel systems workers. Dweik et al (2011) developed guidance for exhaled nitric oxide as a disease diagnostic, and Lomonaco et al (2018) measured carbonyl compounds in breath to indicate health state. Miekisch et al (2004) explored a variety of targeted metabolic chemicals in breath, especially acetone, isoprene, ethane and sulfur compounds. Gaude et al (2019) proposed the study of exogenous breath VOCs as probes to assess the activity of metabolic enzymes. The targeted analysis concept employed in these studies is useful for exploring and documenting the potential effects of known chemicals on the public, developing regulations, and non-invasively assessing health state. The caution is that a targeted compound may be obscured or confounded in a complex matrix requiring additional confirmation for identification.

1.2. Non-targeted approach

Non-targeted mass-spectrometry-based approaches have been traditionally applied in proteomics and metabolomics investigations (Denkert et al 2006, Patti et al 2012, Cajka and Fiehn 2016, Bloszies and Fiehn 2018, Seerangaiyan et al 2019). More recently, non-targeted analysis has been used to discover as many compounds as possible in environmental and biological media for assessing responses to environmental exposures and for screening disease states (Pleil and Stiegel 2013, Cariou et al 2016, Wallace et al 2017, 2019a, Newton et al 2018, Sobus et al 2018, Pleil et al 2019). Non-targeted and suspect chemical screening approaches have been implemented using liquid chromatography–high resolution mass spectrometry (LC–HRMS) for environmental measures in dust, water, and airborne particles, and for liquid biological media, including blood, exhaled breath condensate and urine (Bouatra et al 2013, Psychogios et al 2011, Ulrich et al 2019). Suspect chemical screening involves the use of chemical libraries for identification without a specific chemical target list. In non-targeted analysis, all chemical features (detected ions with a unique m/z ratios at unique retention times) in the chromatogram are considered unknown until external knowledge is applied. This approach has become more useful as analytical instrumentation, and the ancillary diagnostic software systems have improved. The most important advances have occurred in the realm of GC and LC coupled with HRMS that can be used to detect thousands of chemical features representative of many compounds (Cariou et al 2016, Newton et al 2018, Gómez-Ramos et al 2019, Singh et al 2020). Generally, instrument manufacturers provide proprietary software that helps to identify unknowns, but the search results are limited to the chemicals present in the vendor’s software and the scoring metrics for identification are often not understood by users. Thus, the software may propose identifications that are unlikely in the particular sample type. As such, post-processing and subject matter expertise, often involving the development of automated workflows, are even more important for non-targeted analysis than for targeted analysis (Siren et al 2019, Chao et al 2020). The advent of breath metabolomics and applied biometric data processing solutions have also improved non-targeted analysis in breath research (Rattray et al 2014, Smolinska et al 2014).

The non-targeted concept is a complement to targeted analyses. In fact, a well-designed non-targeted study could ultimately lead to additional specific probative chemicals for application in simpler targeted methods as recently implemented for a study of firefighters (Wallace et al 2019a). The caution is that even the most sensitive and specific instrumentation can provide ambiguous results. A combination of database searches, literature assessment, and subject matter expertise all need to be applied to categorize and ultimately gain confidence in specific identifications.

Although sample-to-sample pattern shifts in unknown or tentatively identified biological chemicals are a valuable tool, ultimately researchers need to correctly identify as many compounds as possible to better develop the biochemistry underlying the human response to environmental and disease stressors. As described above, the challenges are much greater in non-targeted analyses; however, even the best targeted approaches may also require additional scrutiny. In either case, a comprehensive cheminformatics approach is necessary.

1.3. Breath applications

In breath gas analysis, GC–MS is a prominent analytical methodology for VOCs and other related species (Amman et al 2004, Pleil and Wallace 2018). Until recently, the analysis of certain compounds within the breath volatilome had not been directly supported by the Dashboard. This gap has been addressed by adding a series of compound lists to the Dashboard that can be used specifically to narrow hit lists of chemicals to those previously identified in the volatilome and to support structural identification using LC–MS or GC–MS. Most recently, the Dashboard has been augmented to include other compounds found in breath, blood and urine that are less volatile, but still detectable in GC–MS analyses. This was prompted by recent work in our laboratories in which we sought to use the Dashboard to identify compounds in breath aerosol samples (Pleil et al 2019, Wallace et al 2019b, Wallace and Pleil 2018).

This article specifically focuses on the use of the Dashboard as an investigative tool for breath VOCs and other relatively small organic compounds found in the headspace of liquid media, including blood, urine and condensed phase breath. We pose some common questions arising in GC–MS data interpretation and show how to use the Dashboard to help obtain answers. The focus of this tutorial is on standard electron impact (EI) fragmentation spectra (MS1), obtained using both single Da nominal mass and high-resolution mass spectrometry. We also provide some additional resources for focusing on breath-specific literature and a ‘crowdsourcing’ tool to help categorize the compounds in the list.

2. U.S. EPA CompTox Chemicals Dashboard

The U.S. EPA CompTox Chemicals Dashboard is an online database that contains comprehensive chemical data such as properties, structures, hazards, exposures, and bioactivity for ~900 000 chemicals. The website URL is https://comptox.epa.gov/dashboard, and the Dashboard homepage is shown in figure 1. In many ways the dashboard can be considered to be a rather small resource compared to other popular online databases, specifically PubChem with ~110 million chemicals and ChemSpider with ~100 million chemicals. However, there is a distinct difference between the quality of data in these resources as well as some of the available functionality. There have been a number of reports of data quality in public domain databases (e.g. https://pubmed.ncbi.nlm.nih.gov/21871970/, www.sciencedirect.com/science/article/abs/pii/S135964461200075X) but the dashboard data differ in that the data, which is a far smaller and focused data set to support EPA research, has a dedicated team of curators, has been developed over almost two decades of ongoing curation (www.sciencedirect.com/science/article/pii/S2468111319300234) and only grows by a few thousand records per year. This should be compared with PubChem and ChemSpider that can expand by many millions of records a year without the level of curation provided for the underlying Dashboard data. It should be noted that both PubChem and ChemSpider are open to public depositions while this is not possible through the dashboard which is very rigorously controlled and only a small group of EPA employees and curators have data registration and curation privileges. The impact of the effort put into data quality, especially in terms of compound identification by mass spectrometry techniques, has been reported previously (https://link.springer.com/article/10.1007/s00216-016-0139-z and www.ncbi.nlm.nih.gov/pmc/articles/PMC7345619/).

Figure 1.

Figure 1.

Homepage of the Dashboard. The top menu (in the blue header) allows navigation to multiple types of searches, lists of chemicals, assays, real time predictions and downloadable files. The search box allows for searching based on chemical identifiers (e.g. CAS Registry Number, chemical name, InChIKey) or, with selection of the appropriate tab, searching for chemicals by product use or category as well as by assay or gene associated with bioactivity data. The scrolling text ‘Latest News’ contains announcements for the Dashboard.

The applications of the dashboard to host and deliver data have been discussed in a number of articles and the reader is referred to these for details (e.g. https://jcheminf.biomedcentral.com/articles/10.1186/s13321-017-0247-6, https://doi.org/10.14573/altex.1811292). A number of searches are available and include the access to bioactivity data (https://doi.org/10.1007/s00216-018-1526-4 and https://doi.org/10.1021/acs.chemrestox.0c00264) and consumer product data (https://doi.org/10.1038/sdata.2018.125). In relation to this report the mass and formula-based search options available on the Dashboard can be used to aid in identification of chemicals detected in LC–MS and GC–MS analyses (Williams et al 2017, Newton et al 2018, McEachran et al 2017, McEachran et al 2020, Pleil and Williams 2019). This will be our starting point as we consider topics relevant to human breath research.

In the following section, a series of questions involving common problems in GC–MS will be introduced. These issues include how to identify chemicals using integer mass via source ranking approaches and how to perform literature searches to ensure tentative identifications have previously appeared in breath research related articles. We will conclude with a summary of the current state of the Dashboard and future improvements that will further enhance breath research.

2.1. Topics/questions

The following questions may arise when analyzing breath samples using GC–MS, especially in situations wherein there are multiple unknown peaks or ‘features’, or when an apparent identification from the instrument software is suspect. Briefly, these questions are addressed in detail in the text:

  • There are some tentatively identified compounds (from a library search) in my chromatogram; how likely are they to be found in breath?

  • There are a series of compounds in the chromatogram with the same apparent integer mass; how can I tell them apart and ultimately assign tentative identifications (i.e. identified by names, CASRN, etc)?

  • Is there a way to search for multiple related compounds at once to avoid repeating similar efforts?

  • We have some new compounds we want to look for in breath sample EI spectra; how can I get some major characteristic ions to extract from a non-targeted chromatogram?

  • My instrument is real-time (no retention time information) and is operating in soft chemical ionization mode (e.g. PTR or SIFT); can I search the database for M + 1, or other adducts to help get possible identifications?

  • How can we determine whether certain compounds derived from consumer products are found in breath, or inversely, what compounds we could look for in breath to identify a particular type of consumer product?

  • The chromatogram has some unknown features with tentative identification embedded among some known calibrated compounds; how can I assign tentative identifications with the help of retention time order?

  • How can I quickly query the literature to see how often a particular compound is mentioned for breath, exhaled breath condensate, or aerosols with respect to a broader generic search?

  • We have identified some interesting compounds in a chromatogram; how can we quickly determine how often these compounds were found in breath, blood, and urine, how likely they are to be exogenous, exogenous metabolites, or endogenous and what exposure sources or disease states to which they are related?

Each of these questions is addressed separately (in order) in the ensuing text in the section labelled ‘Individual questions and related Comptox Dashboard procedures’.

Throughout the development of this article, we found that some issues are currently beyond the scope of the tools in the Dashboard and are listed here:

  • Is there a quick way to obtain information about bacterial and other cellular based off-gas products from in vitro analyses?

  • There are a series of features in my breath chromatogram that have a common prominent base peak ion fragment (such as 43, 57, 58, 71, 72, 85, and 86, etc). These base peak ions represent known fragments of aldehydes, amines, thiols, hydrocarbons and ketones that are probably not the molecular ion; how can I assign structure or identification?

  • We have a list of ions from EI spectra for some particular features; can we directly query the Volatilome list with these ions?

Each of these are excellent questions and the authors are developing new procedures for the Dashboard to expand the scope. The current activities are described in the section labelled ‘Future Developments’.

3. Individual questions and related Comptox Dashboard procedures

  • There are some tentatively identified compounds (from a library search) in my chromatogram; how likely are they to be found in breath?

This is the most common question when developing new methods. New compounds or unknown features in a breath sample could have multiple sources, including contamination from the environment, or could simply be misidentified by the particular search algorithms. The Dashboard has three chemical lists available, which are subsets of the total Dashboard content and assembled from volatilome research articles and data assembled in breath research laboratories. Chemical lists can be accessed from the main page of the Dashboard by clicking ‘Lists’ in the top bar and selecting ‘Lists of Chemicals.’ The chemical lists can then be searched from this page, and the query ‘Volatilome’ can be entered to filter the available chemical lists. The three Volatilome lists presently included in the Dashboard (https://comptox.epa.gov/dashboard/chemical_lists/?search=volatilome) are shown in table 1. For the purpose of this article, we are focusing on the first Volatilome list (VOLATILOME: human breath).

Table 1.

Names and descriptions of the three Volatilome lists included on the Dashboard. The number of chemicals in each list reflects those present at the time of writing, and these numbers may increase as additional volatilome data are obtained and submitted.

List acronym List name List description Number of chemicals
VOLATILOME VOLATILOME: human breath A subset of compounds detected in human breath 1177
VOLATILOME2 VOLATILOME polar, semi-volatile, and condensed phase organic compounds found in human blood, condensed breath and urine VOLATILOME subset of compounds detected in human biological media including blood, dried blood spots (DBS), urine, exhaled breath condensate (EBC), and exhaled breath aerosols (EBA)  133
VOLATILSALIVA VOLATILOME: saliva This list is a subset of compounds detected in saliva  307

A simple example of how the Volatilome list can be used to tentatively identify compounds in GC–MS breath analysis is the NIST library search of a spectrum containing 91 and 92 Da as primary ions, with only a few other fragment ions. Although the ratios of 91/92 are different in the spectra for these two compounds, in relatively complex samples, the NIST library may select either toluene or 1,3,5-cycloheptatriene as the top hit during spectral matching. The spectra for both compounds are shown in figure 2.

Figure 2.

Figure 2.

NIST library spectra comparison of toluene and 1,3,5-cycloheptatriene from a quadrupole mass spectrometer with 70 eV EI ionization. Note that both of these compounds have the chemical formula C7H8 and cannot be distinguished based on mass, but have different ratios of 91/92 m/z.

It is possible that both toluene and 1,3,5-cycloheptatriene are exogenous compounds that could be present in breath as a result of recent exposures. However, a query of the Volatilome: human breath list shows that toluene, but not 1,3,5-cycloheptatriene, is present in the list. 1,3,5-Cycloheptatriene is present in the Dashboard and is included in several other compound lists (11 in total at the time of writing). Toluene is associated with >80 individual chemical lists, from the >220 chemical lists represented in the Dashboard, but it is the presence of the chemical in the Volatilome list that increases the likelihood that the compound detected in breath is toluene, regardless of the ranking of the NIST search results.

  • There are a series of compounds in the chromatogram with the same apparent integer mass; how can I tell them apart and ultimately assign tentative identifications (i.e. identified by names, CASRN, etc)?

A common issue in exhaled breath analysis, especially for MS instruments with single Da resolution, is the identification of multiple compounds that have the same integer mass. Volatile compounds with integer mass below about 150 Da are extremely common in biological gas samples with most falling into the CHNOPS category, i.e. composed of carbon, hydrogen, nitrogen, oxygen, phosphorus and/or sulfur. In addition to these elements, many exogenous VOCs of interest may also contain halogens, especially fluorine, chlorine or bromine (F, Cl, Br), as well as silicon (Si) and, on rare occasions, some metals. In general, the integer mass of such compounds is defined by the fragment with the largest m/z ratio in the mass spectrum, within the context of common isotope patterns.

In order to demonstrate the utility of the Dashboard for chemical identification, two examples are investigated for integer masses 86 and 92 Da. Using the ‘Advanced Search’ page on the Dashboard, searching for the mass 86 ± 0.2 Da without adducts yields 307 chemicals, 24 of which are represented as multi-component chemicals (e.g. salts and hydrates) and filtered out of the display, leaving 283 single component compounds. Distinguishing single component chemicals from multiple component chemical substances is based on a cheminformatics processing approach that standardizes chemical substances to produce ‘MS-ready’ chemical structures (McEachran et al 2018). This standardization separates multi-component chemicals into individual components that are then neutralized, with stereocenters removed and with isotopic labeling removed. As a result of the processing of the chemicals into MS-Ready forms, linkages are retained to all multicomponent forms, stereo-forms, and isotopically labeled forms. For example, the MS-Ready form for nicotine maps to a total of 37 related chemicals (https://comptox.epa.gov/dashboard/dsstoxdb/ms_ready_mixture?cid=28128).

The filtering of multi-component chemicals is simply a way to display those single components that match the input mass searched and is based on the ‘MS-ready’ mass but the actual chemical may be contained within a multicomponent chemical. Similarly, an advanced search for 92 ± 0.2 Da results in 215 chemicals of which 186 are single component features. Figure 3 shows an example of a series of chemical tiles from the 92 Da search that can be further scrolled down to look at all 186 chemicals. It is important to note that integer masses must be searched using the ‘Advanced Search’ page and not the ‘Batch Search’ page. Under ‘Advanced Search’ the mass error can be searched using Daltons, while under ‘Batch Search’ compounds can only be searched according to monoisotopic mass but not integer mass (between 1 and 10 ppm mass error in Batch Search). This is a limitation of using low resolution mass spectrometry data for non-targeted analyses.

Figure 3.

Figure 3.

A subset of the results from the advanced search of 92 ± 0.2 Da querying the entire Dashboard. These results can be scrolled to find all 186 chemicals that fit the search criteria.

In the search results reported in figure 3, this advanced search is based on the total content of the Dashboard; this is not particularly useful when we want to limit the search to only those chemicals likely to be found in breath. These kinds of searches can be modified by navigating to the ‘Lists’ button in the top navigation bar, selecting ‘Lists of Chemicals,’ and searching for the Volatilome list that represents the compounds found in breath. Here, all chemicals within that list can be viewed. After selecting the ‘Mono. Mass’ feature in the left center drop down menu, click on the ‘hot key’ toggle on the upper righthand side of the search results menu to switch from the graphic tiles (as shown in figure 3 above) to the list representation, which displays all chemicals in the Volatilome list in order of increasing monoisotopic mass. By scrolling down to 86 or 92 Da monoisotopic mass ranges, the number of chemicals that fall within these integer mass ranges can be viewed. Figure 4 shows a subset of the results for the 86 Da portion of the list.

Figure 4.

Figure 4.

Extracted entries of a subset of 86 Da integer mass constituents of the Volatilome list, ordered by monoisotopic mass. These entries include ketones, aldehydes, and hydrocarbons.

In the Volatilome list of the candidates with integer mass 86 Da, there are only three distinct monoisotopic masses: 86.036779, 86.073165, and 86.10955, representing various carbonyls and hydrocarbons. Certainly, EI spectra contain fragments that can further help differentiate these compounds. Tentatively identified chemicals could also be further queried against the NIST database or one of the myriad online databases such as MassBank (Horai et al 2010) and Mass-Bank of North America (MoNa; https://mona.fiehnlab.ucdavis.edu/) databases that host experimental mass spectra. If there are representative spectra for specific chemicals in these open databases, they would be visible on the external links page for a particular chemical, for example for pentanal at https://comptox.epa.gov/dashboard/dsstoxdb/results?search=pentanal#links. For this example, however, we focus only on using the monoisotopic masses to obtain an initial list of chemicals.

Similarly, the candidates from the Volatilome list are represented by five entries with 92 Da, shown in figure 5. The chlorinated compounds could be quickly identified by taking a closer look at the isotope patterns, as chlorinated compound spectra contain an M + 2 molecular ion peak with a 3:1 peak height ratio in comparison to the M+ molecular ion peak. Toluene has been found in breath samples in exposure scenarios, for example in breath samples collected from firefighters after exposure to controlled structure burns during training exercises (Pleil et al 2014, Wallace et al 2019c). These additional clues from the mass spectra and knowledge of the origins of the sample (i.e. endogenous or exogenous/exposure) can help to narrow down the list of potential compounds in breath samples using the Volatilome list.

Figure 5.

Figure 5.

Extracted entries of all 92 Da integer mass constituents of the Volatilome list, ordered by monoisotopic mass.

  • Is there a way to search for multiple related compounds at once to avoid repeating similar efforts?

Finding compounds related to a particular list can be accomplished by using the batch search feature available on the Dashboard (https://comptox.epa.gov/dashboard/dsstoxdb/batch_search), as shown in figure 6. Inputs include, but are not limited to, chemical identifiers (i.e. chemical name, CASRN, and InChIKeys (or any combination of these)), exact formula(e), or monoisotopic mass. To perform a search these can be entered, by copy-paste for example, into the box located directly under ‘Enter Identifiers to Search’. At present, a maximum of 5000 identifiers are possible; due the number of hits based on a mass or formula search, no more than 200 identifiers are recommended. Identifiers are generally a one-to-one match so 5000 identifiers will produce 5000 or less results. However a single formula can give hundreds of hits and slow down the search dramatically. 5000 masses or formulae input could result in many tens of thousands of hits. Following a search the user can then click on the ‘Display All Chemicals’ button to ensure that each input value has a match in the database or can click the ‘Download Chemical Data’ button. This will produce a list of selectable options available to download into specific download formats, including comma and tab-separated formats as well as Microsoft Excel spreadsheets. Data types of interest to mass spectrometrists that can be included in the download file include MS-ready SMILES, molecular formula, and monoisotopic mass. The inputs that are present in the available lists on the Dashboard, including the Volatilome lists, can easily be determined. For example, searching for ‘C6H4Cl2’ using the molecular formula(e) input type and selecting all Volatilome lists for download yields three chemicals (ortho-, meta-, and para-dichlorobenzene) with that formula, which are all present in the Volatilome: human breath list but not in the other two Volatilome lists (see table 1). The batch search offers an easy path to source information from the databases underlying the Dashboard, including physicochemical predictions (Mansouri et al 2018), the number of consumer products associated with a particular chemical (Dionisio et al 2018), whether or not a chemical has been examined in bioactivity screens and many other data types.

Figure 6.

Figure 6.

The batch search feature of the Dashboard. The entry box on the top right accepts any of the input types listed on the top left side. Selecting the ‘Download Chemical Data’ button produces the downloadable options shown on the bottom half of the screen.

  • We have some new compounds we want to look for in breath sample EI spectra; how can I get some major characteristic ions to extract from a non-targeted chromatogram?

From the perspective of using the Dashboard without invoking analytical instrument software, the first step would be to explore if these compounds have ever been reported in breath. The Volatilome lists can be queried using the search box on an individual list page (i.e. search by CASRN or name) or by pasting a list of identifiers into the batch search box and checking whether those chemicals are in any of the Volatilome lists as explained above. If the compounds of interest are present in the Volatilome lists, then a user can use their own copy of the NIST database to look up the spectra or could navigate to online data sources such as MassBank, MoNa or the NIST webbook via the ‘External Links’. The mass spectra can then be visually inspected using those available online resources and the best representative ions that distinguish the compound of interest can be selected. For example, nonanoic acid is listed in two of the three Volatilome lists (VOLATILOME: Human Breath and VOLATILOME: Saliva), and there are links to numerous online databases as shown in figure 7. Some of these online databases can be used to view the mass spectrum for nonanoic acid in order to select representative ions.

Figure 7.

Figure 7.

Hyperlinks to public analytical resources for nonanoic acid, including the EI mass spectrum on the NIST Webbook.

It is possible that suspect compounds may not be reported in the Dashboard. If new compounds are identified in breath, the authors encourage researchers to forward this information for inclusion in the next revision of the Dashboard by either searching for a chemical and submitting a comment for that chemical to be associated with the Volatilome lists or by suggesting a new chemical to be added by commenting through the contact page (https://comptox.epa.gov/dashboard/contact_us).

  • My instrument is real-time (no retention time information) and is operating in soft chemical ionization mode (e.g. PTR or SIFT); can I search the database for M + 1, or other adducts to help get possible identifications?

The advanced search functionality allows for the searching of relevant (pre-populated) adducts as shown in figure 8. The user can either search for a specific adduct from the list (note that the default selection is no adducts) or search for all possible adducts using the toggle switch.

Figure 8.

Figure 8.

An adduct selection dropdown menu for inclusion in searches is available in the advanced search. The user can select either a single adduct from the drop-down menu or include all potential adducts in the search by clicking the ‘All Adducts’ toggle switch.

As an example of the utility of adduct searching, consider a suspect chemical screen that returns an exact mass of 222.22 g mol−1. A mass-based search for 222.22 ± 5 ppm produces no hits (except for a three component magnesium salt), while a search for all adducts shows 20 hits, which in this case are all for sodiated adducts. These adducts are all potentially related to C13-chain alkylamines (https://comptox.epa.gov/dashboard/dsstoxdb/multiple_results?error_ppm=5&mass_adducts=−22.989229&input_type=ms_ready_monoisotopic_mass&inputs=222.22). While this is not definitive in terms of identifying a specific chemical, it is highly suggestive of the MS-response being for a sodium adduct form of a chemical with formula C13H29NNa.

  • How can we determine whether certain compounds derived from consumer products are found in breath, or inversely, what compounds we could look for in breath to identify a particular type of consumer product?

The Chemicals and Products Database (CPDat) is a database of consumer product composition and functional use (Dionisio et al 2018), including information on 15 000 consumer products. The data are available as a list on the Dashboard (https://comptox.epa.gov/dashboard/chemical_lists/CPDAT) and are available as metadata to use when searching using mass or formula to identify a particular chemical. For example, performing a Batch Search for the MS-ready formula C6H10O and selecting ‘Display All Chemicals’ produces a list of over 200 chemicals, as shown in figure 9. Selecting ‘CPDAT’ in the two drop down menus on the left side of the menu reveals that only six of the >200 chemicals have CPDAT product counts with cyclohexanone being the most prominent of these chemicals with an occurrence in ~1500 products and the second most prominent chemical having a CPDat count of five products. These metadata are important for the identification of known unknowns and candidate ranking as explained earlier. To identify whether any of these chemicals are present in any of the Volatilome lists, the user would select the subset of compounds present in CPDat (by clicking on the check circle at the end of the row or at the top right corner of each tile) and select the ‘Send to Batch Search’ button on the left-hand side. This brings up the Batch Search page, in which the DSS Tox Substance IDs for the selected chemicals of interest will be populated. The user would select ‘Download Chemical Data’ and under the ‘Presence in Lists’ column, check the boxes for the three Volatilome lists (List: VOLATILOME polar, semi-volatile, and condensed phase organic compounds found in human blood, condensed breath and urine, List: VOLATILOME: Human Breath, and List: VOLATILOME: Saliva). The data can be downloaded by choosing an output format (Excel, TSV, CSV, or SDF) and selecting the ‘Download’ button. The batch search reveals that three of the compounds with the MS-ready formula C6H10O (cyclohexanone, 4-methylpent-3-en-2-one, and 5-hexen-2-one) are present in the VOLATILOME: Human Breath list.

Figure 9.

Figure 9.

A table view of the results set based on a search for the MS-ready formula C6H10O. The resulting hit list is sorted based on descending count of frequency of occurrence in CPDat. The results show that cyclohexanone is the most frequently occurring chemical with that formula and demonstrates the value of CPDat counts as useful metadata for candidate ranking.

  • The chromatogram has some unknown features with tentative identification embedded among some known calibrated compounds; how can I assign tentative identifications with the help of retention time order?

If a candidate feature has been tentatively identified among a set of known compounds within the same chromatogram, the Dashboard can be utilized to obtain retention time indices to support the tentative identification. To do this, search each of the known compounds using the search feature previously discussed to reach each chemical’s page. From there, locate the ‘Links’ tab on the left-hand side and click it. The ‘Links’ page includes links to numerous websites with additional information for the searched chemical. Where available, the ‘NIST Kovats Index values’ link navigates to NIST’s webpage. For a particular chemical, the index values associated with specific GC–MS methods and associated literature references are listed. The relevant retention time indices can be harvested from the NIST page for each chemical of interest. The investigator can then approximate where different tentatively identified features should elute and test the indices against the different tentative identifications. For example, for nonanoic acid (https://webbook.nist.gov/cgi/cbook.cgi?ID=C112-05-0&Mask=2000#Gas-Chrom), dozens of rows of information are available regarding the retention indices under different experimental conditions (i.e. active phases, column dimensions, temperature ramps, etc). Links to references for each entry are included. It should be noted that Kovats Index values are not available for all chemicals contained in the Dashboard and are limited to data contained in the NIST webbook.

  • How can I quickly query the literature to see how often a particular compound is mentioned for breath, exhaled breath condensate, or aerosols with respect to a broader generic search?

The Dashboard has an ‘abstract sifter’ function available under the ‘PubMed Abstract sifter’ subtab of the ‘Literature’ tab for each individual chemical’s Dashboard page. This capability, described in detail elsewhere (Baker et al 2017, Williams et al 2017), searches PubMed using the application programming interface. For compounds of interest such as cholesterol, malondialdehyde, heptanal, hexanal, acetaldehyde, etc that have been detected in breath, it is possible to identify articles of interest referenced in PubMed. First, navigate from the Dashboard’s home page and to a chemical of interest’s page as discussed previously. Using the subtabs on the left side of the page, select ‘Literature’ and then ‘PubMed Abstract Sifter’. This page shows a list of potential queries of interest and a pre-populated search box containing the CASRN and chemical name. Selecting a specific query populates the query box with an appropriate query string. For example, choosing the query ‘Dust and Exposure’ will populate the search box with (’66-25-1’ OR ‘hexanal’ OR ‘n-hexanal’) AND Dust AND Environmental Exposure for the chemical hexanal. However, the search can be optimized for more focused searches. For example, on the right side insert key words (separate by ‘AND’ for the particular topic). For this example, use either: ‘66-25-1’ OR ‘hexanal’; ‘66-25-1’ OR ‘hexanal’ AND ‘breath’; ‘66-25-1’ OR ‘hexanal’ AND ‘breath condensate’; ‘66-25-1’ OR ‘hexanal’ AND ‘breath aerosol’ in separate searches. The results not only include the total number of PubMed abstracts that match the search criteria, but also provide links to each article. For example, a search for ‘111-71-7’ OR ‘heptanal’ AND ‘breath’ results in 26 hits from a query against over 30 million abstracts in PubMed at the time of writing. The resulting hits can then be additionally ‘sifted’ by entering terms of interest into the boxes directly below ‘to find articles quickly, enter terms to sift abstracts,’ as shown in figure 10. In this case, the abstracts associated with the 26 articles are retrieved and the sifting terms ‘breath’, ‘cancer’ and ‘biomarkers’ are entered. The resulting frequency counts for each of the additional terms are included in the three color-coded columns. The color coding is also used to mark up the relevant terms in the associated abstracts. When the user identifies an article of interest, they can click on the PubMed identifier highlighted in blue to navigate to the relevant abstract.

Figure 10.

Figure 10.

A filtered search against 30 million abstracts in PubMed using the CASRN and name for heptanal and the additional term breath. The abstracts for 26 articles were retrieved and sifted using the terms breath, cancer and biomarkers, as indicated using the color coding. The numbers are the frequency counts for the individual sifting terms in the abstracts. The color coding is used to mark up the relevant terms in the associated abstracts. The user can read an associated article of interest by clicking on the associated PubMedID.

  • We have identified some interesting compounds in a chromatogram; how can we quickly determine how often these compounds were found in breath, blood, and urine, how likely they are to be exogenous, exogenous metabolites, or endogenous and what exposure sources or disease states to which they are related?

These types of searches are possible in the abstract sifter module of the Dashboard using search strings such as those shown in the previous section. However, this is a more difficult procedure because query terms such as ‘exogenous’ or ‘endogenous’ might miss many abstracts as these conditions are often taken out of context. Furthermore, determining exposure sources and disease relationships would require separate searches for each term and post-search combination and duplication would be additional work.

It is possible to use the list functionality to query segregated lists of interest if they are constructed appropriately. For example, if lists were split into relevant body fluids then presence in those lists can be helpful indicators. There are currently three lists for blood (BLOODEXPOSOME, HUMANBLOOD, and VOLATILOME2; and two lists for breath (VOLATILOME and VOLATILOME2; so, this utility is currently limited. We have embarked on an alternative path to address this question and have developed a stand-alone spreadsheet of the Volatilome list with additional columns for a series of descriptors. This spreadsheet could be enhanced using a ‘crowdsourcing’ approach by engaging the breath community wherein members who have measured a particular compound can provide their input as to which biological media it has been identified in by placing a vote in the appropriate column. The breath research community will ultimately benefit from this aggregate experience. This ‘Breath Research Exchange Forum’ (https://docs.google.com/spreadsheets/d/1wspixeki0oI-ew8vx09YRDwlvYBABiVk8ABrre4U5r4/edit?usp=sharing) is already online and we are actively encouraging engagement.

Figure 11 shows an example for the first few compounds in the Volatilome list.

Figure 11.

Figure 11.

The ‘Breath Research Exchange Forum’ spreadsheet containing chemicals that have been found in samples via GC–MS. The spreadsheet includes both chemical identifiers and indicators for the type of sample in which the chemical was identified.

4. Future developments

In this article, we have demonstrated how to use the Dashboard as a common, instrument independent, starting point for investigating unknown GC–MS features for a series of specific questions arising from non-targeted GC–MS analyses. We realize that new questions or better search methods will continually arise. We are happy to entertain questions from the readership to continually improve the value of the Dashboard.

At present, we are working on Dashboard projects to address the following three questions:

  • Is there a quick way to obtain information about bacterial and other cellular based off-gas products from in vitro analyses?

We have become aware of separate efforts in developing databases for in vitro off-gas products for cell samples, bacteria cultures, tissue samples and microbioreactors. We hope to incorporate these as a separate list, and ultimately as linked features for the Dashboard. This question can already be addressed specifically using some of the tools outlined above, especially the ‘batch search’ coupled with the ‘abstract sifter’.

  • There are a series of features in my breath chromatogram that have a common prominent base peak ion fragment (such as 43, 57, 58, 71, 72, 85, and 86, etc). These base peak ions represent known fragments of aldehydes, amines, thiols, hydrocarbons and ketones that are probably not the molecular ion; how can I assign structure or identification?

This question represents one of the more vexing problems in non-targeted GC-MS analyses of breath samples using EI ionization. Most alcohols, ketones, aldehydes and hydrocarbons, as well as less common compounds such as amines and thiols, tend to fragment with common ions such as C3H7 (43 Da), C2H3O (43 Da), C4H9 (57 Da), C3H6O (58 Da), C5H11 (71 Da), C4H8O (72 Da), C3H8S (76 Da), C5H9O (85 Da), C5H12N (86 Da) and C5H10O (86 Da). Larger compounds or cyclic compounds can also have common fragments, including 78, 84, 91, 92, 105, 106, 112 Da, etc.

Generally, these fragments can occur within and between features in a confined part of the chromatogram, making it difficult to assign molecular structures, or even determine if they are from different compounds. Standard library searches from instrument software can often provide too many guesses to be useful.

We are now working on a Dashboard-based triage mechanism to narrow down these possibilities by developing search algorithms that allow for the direct searching of spectral data (using X,Y input data to represent fragment ions) against in silico fragmentation data generated using CFM–ID computational fragmentation approaches (Allen et al 2016). We are now studying such applications to real world samples in our laboratories.

  • We have a list of ions from EI spectra for some particular features; can we directly query the Volatilome list with these ions?

This question is analogous to the issue immediately above, but now the searches could be focused on individual major ions. This is more difficult in that co-eluting compounds need to be addressed with some form of learning approach. There are proprietary (instrument based) search algorithms that already have such capabilities, but for now, we are still relegated to whole spectra searches. In the future, we hope to adapt some of this technology for the Dashboard.

5. Summary discussions

The U.S. EPA Comptox Chemicals Dashboard is a useful tool for exploring all manner of non-targeted analytical results. Herein, we have focused on typical questions that could arise from GC–MS analyses of samples for which we have no targeted preconceptions. These revolve around the raw data from the instrument that provide retention times, mass fragments, polarities and peak shapes (based on column choice), and tentative identifications from instrument software, etc that make it difficult to narrow down the true identities, or even the class of compounds present in our samples.

The value of the Dashboard is the overall consolidation of analytical information in a one-stop framework that is an open-sourced, regularly updated and curated resource. Furthermore, the Dashboard provides curated lists of chemicals with some common themes. Herein we focused on the Volatilome list which contains compounds found in human breath (Volatilome: human breath) as a subset of the total Dashboard as a starting point for GC–MS. The list named Volatilome2 (see table 1) contains compounds found in GC–MS analyses of other biological media, including blood, breath condensate, urine, and breath aerosols. However, we are not restricted to these Volatilome lists; it is a simple task to query the total database, or to investigate other lists wherein these compounds may also appear.

This tutorial has been designed in the format of ‘frequently asked questions’ to allow the researcher to skip quickly to a section that addresses a specific question or concept. We encourage the readership to ask us other questions that we may have missed, and we intend to update this tutorial as necessary with application notes in the Journal of Breath Research, or directly from reader queries. We will also add new compounds to the Volatilome lists and encourage the readership to submit compounds from their own non-targeted studies for inclusion. New updates go live at 6–12 month intervals. The contact for updates to the Dashboard is Dr Antony Williams: email williams.antony@epa.gov .

Acknowledgments

The authors are grateful for helpful discussions and advice from Jon Sobus and James McCord from the U.S. Environmental Protection Agency, North Carolina, USA; Wolfram Miekisch from Rostock University Hospital, Germany; Ben de Lacy Costello and Norman Ratcliffe from University of West England, Bristol, UK.

This article was reviewed in accordance with the policies of the Office of Research and Development, U.S. Environmental Protection Agency, and approved for publication. Mention of trade names or commercial products does not constitute endorsement or recommendation for use.

References

  1. Allen F, Pon A, Greiner R and Wishart D 2016. Computational prediction of electron ionization mass spectra to assist in GC/MS compound identification Anal. Chem 88 7689–97 [DOI] [PubMed] [Google Scholar]
  2. Amann A, Poupart G, Telser S, Ledochowski M, Schmid A and Mechtcheriakov S 2004. Applications of breath gas analysis in medicine Int. J. Mass. Spectrom 239 227–33 [Google Scholar]
  3. Baker N, Knudsen T and Williams A 2017. Abstract sifter: a comprehensive front-end system to PubMed F1000Res. 6 1–10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Bloszies CS and Fiehn O 2018. Using untargeted metabolomics for detecting exposome compounds Curr. Opin. Toxicol 8 87–92 [Google Scholar]
  5. Bouatra S. et al. The human urine metabolome. PLoS One. 2013;8:e73076. doi: 10.1371/journal.pone.0073076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Cajka T and Fiehn O 2016. Toward merging untargeted and targeted methods in mass spectrometry-based metabolomics and lipidomics Anal. Chem 88 524–45 [DOI] [PubMed] [Google Scholar]
  7. Cariou R, Omer E, Léon A, Dervilly-Pinel G and Le Bizec B 2016. Screening halogenated environmental contaminants in biota based on isotopic pattern and mass defect provided by high resolution mass spectrometry profiling Anal. Chim. Acta 936 130–8 [DOI] [PubMed] [Google Scholar]
  8. Chao A et al. 2020. In silico MS/MS spectra for identifying unknowns: a critical examination using CFM-ID algorithms and ENTACT mixture samples Anal. Bioanal. Chem 412 1303–15 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Denkert C, Budczies J, Kind T, Weichert W, Tablack P, Sehouli J, Niesporek S, Könsgen D, Dietel M and Fiehn O 2006. Mass spectrometry-based metabolic profiling reveals different metabolite patterns in invasive ovarian carcinomas and ovarian borderline tumors Cancer Res. 66 10795–804 [DOI] [PubMed] [Google Scholar]
  10. Dionisio KL, Phillips K, Price PS, Grulke CM, Williams AJ, Biryol D, Hong T and Isaacs KK 2018. The Chemical and Products Database, a resource for exposure-relevant data on chemicals in consumer products Sci. Data 5 180125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Dweik RA, Boggs PB, Erzurum SC, Irvin CG, Leigh MW, Lundberg JO, Olin AC, Plummer AL and Taylor DR 2011. American Thoracic Society Committee on Interpretation of Exhaled Nitric Oxide Levels (FENO) for Clinical Applications. An official ATS clinical practice guideline: interpretation of exhaled nitric oxide levels (FENO) for clinical applications Am. J. Respir. Crit. Care Med 184 602–15 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Gaude E, Nakhleh MK, Patassini S, Boschmans J, Allsworth M, Boyle B and van der Schee MP 2019. Targeted breath analysis: exogenous volatile organic compounds (EVOC) as metabolic pathway-specific probes J. Breath Res 13 032001. [DOI] [PubMed] [Google Scholar]
  13. Gómez-Ramos MM, Ucles S, Ferrer C, Fernández-Alba AR and Hernando MD 2019. Exploration of environmental contaminants in honeybees using GC-TOF-MS and GC-Orbitrap-MS Sci. Total Environ 647 232–44 [DOI] [PubMed] [Google Scholar]
  14. Horai H et al. 2010. MassBank: a public repository for sharing mass spectral data for life sciences J. Mass Spectrom 45 703–14 [DOI] [PubMed] [Google Scholar]
  15. Lomonaco TO, Romani AN, Ghimenti SI, Biagini DE, Bellagambi FG, Onor MA, Salvo PI, Fuoco RO and Di Francesco F A 2018. Determination of carbonyl compounds in exhaled breath by on-sorbent derivatization coupled with thermal desorption and gas chromatography-tandem mass spectrometry J. Breath Res 12 046004. [DOI] [PubMed] [Google Scholar]
  16. Mansouri K, Grulke CM, Judson RS and Williams AJ 2018. OPERA models for predicting physicochemical properties and environmental fate endpoints J. Cheminform 10 10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. McEachran AD, Chao A, Al-Ghoul H, Lowe C, Grulke C, Sobus JR and Williams AJ 2020. Revisiting five years of CASMI contests with EPA identification tools Metabolites 10 260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. McEachran AD, Mansouri K, Grulke C, Schymanski EL, Ruttkies C and Williams AJ 2018. ‘MS-Ready’ structures for non-targeted high-resolution mass spectrometry screening studies J. Cheminform 10 45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. McEachran AD, Sobus JR and Williams AJ 2017. Identifying known unknowns using the US EPA’s CompTox Chemistry Dashboard Anal. Bioanal. Chem 409 1729–35 [DOI] [PubMed] [Google Scholar]
  20. Miekisch W, Schubert JK and Noeldge-Schomburg GF 2004. Diagnostic potential of breath analysis—focus on volatile organic compounds Clin. Chim. Acta 347 25–39 [DOI] [PubMed] [Google Scholar]
  21. Newton SR, McMahen RL, Sobus JR, Mansouri K, Williams AJ, McEachran AD and Strynar MJ 2018. Suspect screening and non-targeted analysis of drinking water using point-of-use filters Environ. Pollut 234 297–306 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Patti GJ, Yanes O and Siuzdak G 2012. Metabolomics: the apogee of the omics trilogy Nat. Rev. Mol. Cell Biol 13 263–9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Pleil JD and Isaacs KK 2016. High-resolution mass spectrometry: basic principles for using exact mass and mass defect for discovery analysis of organic molecules in blood, breath, urine and environmental media J. Breath Res 10 012001. [DOI] [PubMed] [Google Scholar]
  24. Pleil JD, Smith LB and Zelnick SD 2000. Personal exposure to JP-8 jet fuel vapors and exhaust at air force bases Environ. Health Perspect 108 183–92 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Pleil JD and Stiegel MA 2013. The evolution of environmental exposure science: using breath-borne biomarkers for ‘discovery’ of the human exposome Anal. Chem 85 9985–90 [DOI] [PubMed] [Google Scholar]
  26. Pleil JD, Stiegel MA and Kent FW 2014. Exploratory breath analyses for assessing toxic dermal exposures of firefighters during suppression of structural burns J. Breath Res 8 037107. [DOI] [PubMed] [Google Scholar]
  27. Pleil JD and Wallace MAG 2018. New breath related topics: sample collection for exhaled breath condensate and aerosol, development of real-time medical alerts, measurement of artificial atmospheres, and analysis of legalized cannabis product J. Breath Res 12 039001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Pleil JD, Wallace MAG and Mccord J 2016. Beyond mono-isotopic accurate mass spectrometry: ancillary techniques for identifying unknown features in non-targeted discovery analysis J. Breath Res 13 012001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Pleil JD, Wallace MAG, Mccord J, Madden MC, Strynar MJ, Sobus JR and Ferguson G 2019. How do cancer-sniffing dogs sort biological samples? Exploring case-control samples with non-targeted LC-Orbitrap, GC-MS, and immunochemistry methods J. Breath Res 14 016006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Pleil JD and Williams A 2019. Centralized resource for chemicals from the human volatilome in an interactive open-sourced database J. Breath Res 13 040201. [DOI] [PubMed] [Google Scholar]
  31. Psychogios N. et al. The human serum metabolome. PLoS One. 2011;6:e16957. doi: 10.1371/journal.pone.0016957. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Rattray NJ, Hamrang Z, Trivedi DK, Goodacre R and Fowler SJ 2014. Taking your breath away: metabolomics breathes life in to personalized medicine Trends Biotechnol. 32 538–48 [DOI] [PubMed] [Google Scholar]
  33. Seerangaiyan K, Maruthamuthu M, van Winkelhoff AJ and Winkel EG 2019. Untargeted metabolomics of the bacterial tongue coating of intra-oral halitosis patients J. Breath Res 13 046010. [DOI] [PubMed] [Google Scholar]
  34. Singh RR, Chao A, Phillips KA, Xia XR, Shea D, Sobus JR, Schymanski EL and Ulrich EM 2020. Expanded coverage of non-targeted LC-HRMS using atmospheric pressure chemical ionization: a case study with ENTACT mixtures Anal. Bioanal. Chem 412 4931–9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Siren K, Fischer U and Vestner J 2019. Automated supervised learning pipeline for non-targeted GC-MS data analysis Anal. Chim. Acta 1 100005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Smolinska A, Hauschild AC, Fijten RR, Dallinga JW, Baumbach J and Van Schooten FJ 2014. Current breathomics—a review on data pre-processing techniques and machine learning in metabolomics breath analysis J. Breath Res 8 027105. [DOI] [PubMed] [Google Scholar]
  37. Sobus JR et al. 2018. Integrating tools for non-targeted analysis research and chemical safety evaluations at the US EPA J. Expo. Sci. Environ. Epidemiol 28 411–26 [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Ulrich EM, Sobus JR, Grulke CM, Richard AM, Newton SR, Strynar MJ, Mansouri K and Williams AJ 2019. EPA’s non-targeted analysis collaborative trial (ENTACT): genesis, design, and initial findings Anal. Bioanal. Chem 411 853–66 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Wallace LA, Pellizzari ED, Hartwell TD, Sparacino C, Whitmore R, Sheldon L, Zelon H and Perritt R 1987. The TEAM (Total exposure assessment methodology) study: personal exposures to toxic substances in air, drinking water, and breath of 400 residents of New Jersey, North Carolina, and North Dakota Environ. Res 43 290–307 [DOI] [PubMed] [Google Scholar]
  40. Wallace MAG and Pleil JD 2018. Evolution of clinical and environmental health applications of exhaled breath research: review of methods and instrumentation for gas-phase, condensate, and aerosols Anal. Chim. Acta 1024 18–38 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Wallace MAG, Pleil JD and Madden MC 2019b. Identifying organic compounds in exhaled breath aerosol: non-invasive sampling from respirator surfaces and disposable hospital masks J. Aerosol. Sci 137 105444. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Wallace MAG, Pleil JD, Mentese S, Oliver KD, Whitaker DA and Fent KW 2017. Calibration and performance of synchronous SIM/scan mode for simultaneous targeted and discovery (non-targeted) analysis of exhaled breath samples from firefighters J. Chromatogr. A 1516 114–24 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Wallace MAG, Pleil JD, Oliver KD, Whitaker DA, Mentese S, Fent KW and Horn GP 2019a. Non-targeted GC-MS analysis of exhaled breath samples: exploring human biomarkers of exogenous exposure and endogenous response from professional firefighting activity J. Toxicol. Environ. Health A 82 244–60 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Wallace MAG, Pleil JD, Oliver KD, Whitaker DA, Mentese S, Fent KW and Horn GP 2019c. Targeted GC-MS analysis of firefighters’ exhaled breath: exploring biomarker response at the individual level J. Occup. Environ. Hyg 16 355–66 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Williams AJ. et al. The CompTox Chemistry Dashboard: a community data resource for environmental chemistry. J. Cheminform. 2017;9:61. doi: 10.1186/s13321-017-0247-6. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES