Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2026 Mar 31;9:399. doi: 10.1038/s41746-026-02527-3

Rapid and noninvasive artificial intelligence-assisted diagnostic method for oral squamous cell carcinoma

Yilan Sun 1,2,3,4,5,6,#, Xin Hu 1,2,3,4,5,6,#, Jing Han 1,2,3,4,5,6, Yujue Wang 1,2,3,4,5,6, Jiacheng Luo 1,2,3,4,5,6, Jiayi Yu 7, Yixiang Duan 8, Xu Wang 1,2,3,4,5,6,✉, Jiannan Liu 1,2,3,4,5,6,✉
PMCID: PMC13201731  PMID: 41917329

Abstract

Oral squamous cell carcinoma (OSCC) remains the most common head and neck malignancy, for which early detection is critical yet challenging with current invasive methods. This study aimed to establish a comprehensive diagnostic framework for OSCC by integrating proton transfer reaction-time-of-flight mass spectrometry (PTR-TOF-MS) breath analysis and metagenomic sequencing with artificial intelligence (AI). Exhaled breath and saliva samples were collected from participants in a discovery cohort (n = 222) and an external validation cohort (n = 83). Samples were analyzed using PTR-TOF-MS and metagenomic sequencing, and multimodal diagnostic models were constructed and trained on the discovery cohort data. We identified OSCC-specific biomarkers, including methanethiol and Fusobacterium nucleatum, and developed an interactive online platform (https://bio.futurecnn.com/) enabling real-time predictions and biomarker interpretability. The AI-driven diagnostic model achieved excellent accuracy (ROC-AUC: 0.92) in distinguishing OSCC patients from healthy controls in the external set. This approach offers a practical, noninvasive solution for OSCC screening and establishes an adaptable framework for other breath-based diagnostics.

graphic file with name 41746_2026_2527_Figa_HTML.jpg

Subject terms: Biomarkers, Cancer, Computational biology and bioinformatics, Oncology

Introduction

Oral squamous cell carcinoma (OSCC) is the predominant subtype of malignant tumors in the head and neck region and accounts for more than 90% of malignancies in this area1,2. Globally, the number of OSCC cases exceeds 300,000 annually and continues to increase steadily, posing significant public health concerns3. When OSCC is detected at early stages (I-II), early and precise diagnosis substantially influences prognosis, with survival rates exceeding 60%. However, survival rate decreases to less than 50% once local spread or lymph node metastasis occurs. This is often accompanied by severe impairments in speech, swallowing, and mastication, significantly reducing quality of life and increasing healthcare burden4. Conventional invasive methods for OSCC diagnosis predominantly rely on histopathological biopsy, supplemented by visual examination and imaging techniques such as computed tomography, magnetic resonance imaging, and positron emission tomography5–7. However, biopsy is inherently invasive and carries risks, and its diagnostic yield is critically dependent on accurate lesion targeting, often leading to sampling errors and missed diagnoses of small, early, or occult tumors. Imaging modalities, while valuable for staging, have limited sensitivity for detecting early mucosal changes and may involve drawbacks like radiation exposure. On the one hand, this reliance on physician experience and sampling accuracy often overlook small or hidden lesions, making early detection challenging8. On the other hand, although noninvasive methods such as saliva-based assays are promising, they are hindered by confounding factors such as oral hygiene and inflammatory conditions, leading to reduced specificity9,10. Thus, rapid, noninvasive, and accurate diagnostic methods for OSCC are urgently needed.

Breath analysis has emerged as a novel noninvasive diagnostic approach, distinguished by its real-time analysis capability, patient-friendly sampling process, and reliable repeatability11. This method leverages the principle that volatile organic compounds (VOCs) in exhaled breath reflect metabolic alterations in tumor tissues, enabling minute-scale rapid detection without invasive procedures12. This intrinsic advantage stems from the direct gas-phase analysis of breath samples, bypassing complex pretreatment steps required in liquid or tissue-based assays. Proton transfer reaction-time-of-flight mass t spectrometry (PTR-TOF-MS) represents a technological breakthrough in breathomics13. Its ultrahigh sensitivity (ppt-level detection), near-instantaneous response time, and capacity to resolve hundreds of VOCs in a single exhalation make it uniquely suited for clinical applications14,15. Notably, PTR-TOF-MS requires minimal sample preparation given that breath is directly ionized via proton transfer reactions, eliminating chromatographic separation delays. This design ensures low interbatch variability, which is critical for large-scale screening. Although PTR-TOF-MS has demonstrated success in lung cancer detection16, challenges remain in establishing standardized breath collection protocols that control for pre-analytical variables (e.g., fasting) and environmental contamination, in decoding high-dimensional VOC signatures amid confounding background signals, and in validating specific biomarkers across diverse populations. Addressing these gaps requires synergistic innovations in both analytical instrumentation and data science.

Artificial intelligence (AI), particularly machine learning techniques, offers powerful solutions for managing complex, high-dimensional datasets typical of breathomics and multiomics analyses17. AI algorithms can efficiently reduce dimensionality, select critical biomarkers, and integrate heterogeneous datasets, significantly increasing diagnostic precision and interpretability18–23. By leveraging advanced computational tools, AI-based models increase biomarker detection accuracy, facilitate real-time clinical decision-making, and enable personalized diagnostic strategies. However, a critical gap remains in developing an integrated diagnostic platform to specifically address the challenges of early, non-invasive OSCC detection, including robust biomarker validation across diverse cohorts and real-time clinical translation.

In this study, we aimed to integrate PTR-TOF-MS breath analysis with AI-driven multiomics strategies to establish a novel, rapid, and noninvasive OSCC diagnostic platform. We constructed a comprehensive pipeline involving standardized breath sampling, robust VOC detection via PTR-TOF-MS, feature extraction using advanced machine learning algorithms, and integration with clinical and microbial data through multiomics analyses. Furthermore, we developed an interactive, open-access online platform to facilitate real-time data interpretation and clinical application. This integrated breath-based platform represents a significant step towards practical, non-invasive point-of-care screening and early detection of OSCC, with the potential to improve patient outcomes and reduce healthcare burdens.

Results

Cohort Profiling and PTR-TOF-MS Method Development

A total of 222 participants were enrolled in the discovery cohort (OSCC:HC = 111:111) and 83 participants in the external validation cohort (OSCC:HC = 53:30), forming a well-balanced case–control population with an independent validation set. Detailed demographic and clinical characteristics are presented in Supplementary Fig. 1. The OSCC group included both newly diagnosed and recurrent cases, with all samples collected prior to any treatment. The median age was 57 (range: 19-89), and 63.5% were male. Clinical staging indicated that 38.74% of OSCC patients were diagnosed at early stages (I–II), while the remainder presented with advanced disease. The anatomical sites of OSCC occurrence included the tongue (45.0%), gums (22.0%), cheek/buccal mucosa (17.4%), floor of mouth (6.4%), palate (3.7%), oropharynx (4.6%), and jawbone/alveolar ridge (0.9%) (Supplementary Fig. 2A). Participants were geographically diverse, originating from 19 provinces across China, reflecting broad regional representation and enhancing the generalizability of the cohort (Supplementary Fig. 2B).

Through our established PTR-TOF-MS platform, we identified a total of 212 high-confidence VOC peaks in breath samples. Analytical robustness was ensured by the inclusion of 13 internal standards with standardized calibration curves (Fig. 1A).

Fig. 1. Construction of PTR-TOF-MS Analytical Method and Cohort Analysis.

Fig. 1

A Schematic of the experimental workflow and PTR-TOF-MS method development (Created with BioRender.com). B Correlation analysis of clinical factors associated with OSCC; * indicates statistical significance, and the Y-axis shows the Spearman correlation coefficient (rho). C Odds ratios (ORs) for risk factors associated with OSCC tumor classification (OR > 1: risk factor; OR < 1: protective factor). D Odds ratios (ORs) for risk factors associated with OSCC tumor stage (OR > 1: risk factor; OR < 1: protective factor).

Multivariate logistic regression analysis revealed significant associations between OSCC incidence and multiple covariates, including rural residence, canker sore, marital status, family history, betel quid usage, alcohol intake, occupational exposure, and smoking intensity (Fig. 1B and Supplementary Table 1). Subsequent subgroup analyses identified marital status as the strongest discriminator of tumor classification and disease stage (Fig. 1C, D; Supplementary Tables 2-3). Importantly, these findings reflect univariable associations and are substantially influenced by the underlying age distribution. Nearly all unmarried participants in our cohort were younger individuals, who inherently have a lower risk of OSCC. Therefore, the observed association likely represents confounding by age and socioeconomic factors, rather than a biological or causal effect. Age itself was not significantly associated with tumor classification in the univariable framework. Occupational exposure was associated with tumor classification but not with cancer stage. However, oral ulceration showed significant associations with both tumor classification and cancer stage.

In summary, we established a rigorously defined cohort and a standardized PTR-TOF-MS platform for accurate detection and characterization of OSCC-associated volatile biomarkers.

PTR-TOF-MS Breath Analysis Enables High-Accuracy OSCC Diagnosis

Having established a standardized PTR-TOF-MS analytical platform and a well-defined cohort, we next sought to leverage this analytical platform to evaluate its diagnostic performance. To evaluate whether PTR-TOF-MS analysis can distinguish human exhaled breath samples from ambient air, as well as differentiate OSCC patients from healthy controls (HC), OPLS-DA was performed. A significant separation was noted between ambient air samples and human breath samples (R2Y = 0.563, Q2 = 0.402, RMSEE = 0.227; Fig. 2A). Furthermore, a clear distinction was observed between OSCC and HC breath samples (R2Y = 0.749, Q2 = 0.464, RMSEE = 0.255; Fig. 2B). Among human breath samples, a total of 212 VOC compounds were detected using PTR-TOF-MS, with hydrocarbons constituting the largest category, accounting for 29.9% of total (Fig. 2C). The top 10 metabolites by relative abundance included acetamide, ethylbenzene, styrene, butanal, 2-propenal, diethyl phthalate, 2-acetylfuran, methanamide, methyl methacrylate and carbon disulfide (Fig. 2D). PERMANOVA analysis revealed several clinical factors significantly correlated with PTR metabolite abundance, notably dental calculus, cancer status, and exercise (Fig. 2E), indicating that clinical covariates significantly modulate VOC patterns and support breath-based diagnostics.

Fig. 2. PTR-TOF-MS Analysis of Oral Exhaled Breath: Diagnostic Model Differentiates Disease and Healthy Groups.

Fig. 2

A OPLS-DA score plot of exhaled metabolites between the human exhaled breath samples and ambient air, indicating distinct metabolic profiles. B OPLS-DA score plot of exhaled metabolites between the OSCC and healthy control (HC) groups. C Distribution of metabolite classes detected using PTR-TOF-MS. D Abundance distribution of the top 10 metabolites detected using PTR. E PERMANOVA R2 values showing the influence of clinical factors on PTR metabolite abundance. F Schematic diagram of diagnostic model construction using PTR data. G Performance comparison of the five machine learning models. H Hyperparameter optimization trajectory for 600 iterations of 5-fold cross-validation showing the highest model accuracy. I Parallel coordinate plot visualizing optimal parameter combinations and corresponding accuracy; each line represents one training run. J Average relative importance of each input parameter on model accuracy across all 600 cross-validation runs. K ROC curve of the LGBM model in the training set. L Confusion matrix of the LGBM model in the training set. M ROC curve of the LGBM model in the testing set. N Confusion matrix of the LGBM model in the testing set.

An AI-driven diagnostic model was constructed using PTR-derived VOC data. The data were split into training (80%) and testing (20%) sets (Fig. 2F). The open-source AutoML framework AutoGluon (by Amazon AWS) was initially utilized to automatically tune parameters and compare multiple models24. Among the tested models, the five best-performing approaches represented different machine learning paradigms (Fig. 2F). Among these candidates, LGBM outperformed others, with an accuracy of 0.893, precision of 0.893, recall of 0.893, F-beta of 0.892, and ROC-AUC of 0.960 (Fig. 2G and Supplementary Table 4). The superior performance of LGBM was consistent across multiple evaluation metrics, reflecting its ability to capture complex nonlinear interactions and feature dependencies inherent in high-dimensional VOCs datasets. To further refine the model, hyperparameter optimization of LGBM was performed using Optuna25 with 600 iterations of 5-fold cross-validation (Fig. 2H). The optimized trajectory demonstrated a progressive improvement in accuracy and resulted in the highest model performance. The corresponding parallel coordinate plot visualizes the optimal parameter combinations and their relationships with classification accuracy, where each polyline represents one training run (Fig. 2I). Furthermore, the relative importance of all the input hyperparameters on accuracy was quantified across 600 iterations, highlighting “min_child_samples” as the most influential parameter (Fig. 2J). The detailed optimal parameter combinations are provided in the Supplementary Table 5.

Model robustness was validated by analyzing its ability to discriminate between OSCC and HC samples in both training and testing datasets. In the training set, the LGBM model demonstrated exceptional performance, achieving an AUC of 0.97 and correctly diagnosing 100% of the OSCC cases (Fig. 2K, L). Similarly, the model demonstrated high diagnostic accuracy in the testing dataset, with an AUC of 0.96 and 94.7% correct diagnosis of OSCC patients (Fig. 2M, N). Through rigorous AutoML selection and hyperparameter tuning, the final LGBM model exhibited superior and consistent performance in both discovery and validation phases, establishing a reliable foundation for breath-based OSCC diagnosis.

Salivary Microbiome Signatures Enable Accurate OSCC Diagnosis across Clinical Subgroups

To further improve the robustness of the diagnostic model, we performed salivary metagenomic sequencing. This complementary approach aimed to capture differences in the microbiome between OSCC patients and healthy controls (HCs). The α-diversity, represented by the Simpson index, revealed significantly greater microbial richness in the HC group than in the OSCC group (Wilcoxon, p = 0.0092; Fig. 3A). β-diversity analysis using principal coordinate analysis (PCoA) based on Bray‒Curtis distances revealed distinct differences in microbial composition between the two groups (ANOVA, p < 0.001; Fig. 3B). Among all the salivary samples, the nine most abundant microbial genera identified were Prevotella, Neisseria, Streptococcus, Veillonella, Porphyromonas, Capnocytophaga, Haemophilus, Alloprevotella and Fusobacterium (Supplementary Fig. 3A). Specifically, the abundance of the genus Fusobacterium and Neisseria significantly increased in OSCC samples compared with that in HC samples, whereas the abundance of Prevotella decreased (Fig. 3C). These distinct microbial signatures suggest their potential utility as diagnostic biomarkers for OSCC. We subsequently developed diagnostic models based on salivary microbiome data, splitting the data into training (80%) and testing (20%) sets. Five machine learning models, namely, Logistic Regression, Gradient Boosting Classifier, Extreme Gradient Boosting, Light Gradient Boosting Machine, and Random Forest were evaluated. Across the five-performance metrics, the random forest model demonstrated superior diagnostic performance, with ROC-AUC scores of 0.962 and 0.924 based on training and testing sets, respectively (Fig. 3D, E and Supplementary Table 6).

Fig. 3. Salivary Metagenomic Analysis: Diagnostic Model Differentiates OSCC and Healthy Groups.

Fig. 3

A Comparison of the Simpson diversity index (α-diversity) between the OSCC and HC groups; higher values indicate greater microbial diversity. B PCoA plot of β-diversity (Bray-Curtis distance) showing differences in microbial distribution between the OSCC and HC groups. C Percent change in the relative abundance of the top 10 microbial genera between the OSCC and HC groups. D Comparison of five machine learning models across five performance metrics. E ROC curves of the random forest model in the training and testing sets. F Accuracy of the diagnostic model stratified by cancer stage in the validation set. G PERMANOVA R2 values for assessing the influence of clinical information on microbial abundance. H Schematic of the integrated multiomics fusion model combining PTR-TOF-MS-detected metabolites and metagenomic data (Created with BioRender.com). I ROC curve of the fusion model in the external validation set.

Furthermore, the diagnostic accuracy stratified by cancer stage in the validation set demonstrated clear differences across subgroups (Fig. 3F). For primary early-stage cancers, the model achieved accuracies of 75.0% (3/4) for stage I and 83.3% (25/30) for stage II. Robust performance was also observed in advanced-stage cancers, with accuracies of 91.7% (22/24) for stage III, 92.9% (13/14) for stage IVA, and 100% (7/7) for stage IVB, although the accuracy declined to 53.3% (8/15) for stage IVC. In recurrent cases, the model achieved 100% accuracy for recurrent stage II (rII, 3/3) and recurrent stage III (rIII, 4/4), whereas the single recurrent stage I case (rI) was misclassified (0/1). Stratification by cancer type also confirmed the robustness of the model. The diagnostic accuracy reached 100% (6/6) in healthy controls, 90.1% (100/111) in untreated relapse cases, 89.5% (34/38) in primary OSCC, and 82.1% (55/67) in post-chemotherapy/radiotherapy or neoadjuvant-treated patients. Multivariate PERMANOVA analysis further indicated that cancer stage (R² = 0.1556, p < 0.001) and cancer type (R² = 0.0613, p < 0.001) were the strongest explanatory factors for microbial abundance variation, whereas smoking frequency (R² = 0.0269, p < 0.01), cigarettes per day (R² = 0.0263, p < 0.01), denture wearing (R² = 0.0139, p < 0.05), toothbrushing frequency (R² = 0.0123, p < 0.05), and rural residence (R² = 0.0133, p < 0.05) also contributed significantly, though with lower explanatory power (Fig. 3G, Supplementary Fig. 3B and Supplementary Table 7). These data demonstrated excellent sensitivity for early-stage cancers (I–II) and consistent accuracy across primary, recurrent, and posttreatment cases.

By aligning microbial features with exhaled VOCs biomarkers, we aimed to enable rapid and highly precise diagnosis with complementary information from both data sources, thereby increasing diagnostic accuracy and enhancing mechanistic interpretability. Building upon the individual strengths of VOCs- and microbiome-based classifiers, we developed a multiomics fusion model through systematic feature integration. Three design principles guided this process: (1) ensuring complementarity between exhaled metabolites and microbial pathways, (2) minimizing redundancy through orthogonal feature selection, and (3) preserving biological interpretability for downstream mechanistic exploration (Fig. 3H). To validate the diagnostic performance of the fusion model, we analyzed another external validation cohort of 30 HCs and 53 patients with OSCC (N = 30 + 53). The fusion model achieved ROC-AUC values of 0.80 in the training cohort (Supplementary Fig. 4A), 1.00 in the internal validation cohort, and 0.923 in the external validation cohort (Fig. 3I).

To further ensure that the fused feature space was stable and not vulnerable to overfitting, we assessed cross-cohort feature distribution stability using the Population Stability Index (PSI)26. Most features exhibited PSI values within the 0.1–0.25 range (Supplementary Fig. 4B), indicating only moderate distribution shift between the discovery and external validation cohorts, consistent with realistic population heterogeneity rather than model drift. A small subset of features demonstrated higher PSI values (Supplementary Fig. 4C), but single-feature decile comparisons (Supplementary Fig. 4D) and PCA visualization (Supplementary Fig. 4E) confirmed that the two cohorts remained partially overlapping in the multi-omics latent space. These results demonstrate that the fusion model is learning biologically meaningful patterns rather than memorizing cohort-specific noise.

These results demonstrate that the multiomics fusion model achieves superior diagnostic accuracy and generalizability.

Feature Selection Offers Mechanistic Insights into Diagnostic Power

To elucidate the biological basis of the multiomics fusion model, we analyzed the individual contributions of VOC and microbiome features. In the established cohort, feature selection and model interpretability analysis were performed using both metabolomics and microbiome datasets to identify specific biomarkers.

For the VOCs dataset, we applied three complementary approaches: statistical testing, regularization-based selection, and model-based filtering to identify robust biomarkers. Specifically, Mann-Whitney U tests with false discovery rate correction were used to detect statistically significant metabolites, while LASSO-Boruta regression was applied to retain features with non-zero coefficients. In parallel, RF and LGBM classifiers were trained, and Shapley additive explanations (SHAP) values were computed to quantify feature importance. The intersection of significant features across these methods defined a final set of 12 VOCs biomarkers (Fig. 4A). Among them, methanethiol showed the largest fold change in OSCC patients, while dimethylthioformamide was more abundant in healthy controls (Fig. 4B).

Fig. 4. Multiomics Feature Selection and Model Interpretability.

Fig. 4

A Venn diagram of 12 shared differentially abundant metabolites identified using Random Forest, LGBM, Boruta, and Mann-Whitney U tests. B showing the fold change in 12 metabolites between the OSCC and HC groups (positive values indicate higher abundance in OSCC); Maaslin2 was used to analyze differential microbial abundance. C Venn diagram showing overlap between significant microbial species identified by Maaslin2 and the top 50 species identified by SHAP analysis. D Coefficients and SHAP values of the 17 overlapping microbial species between the OSCC and HC groups. E Standard curve for methanethiol quantification using PTR-TOF-MS. F Methanethiol abundance comparison between the OSCC and HC groups based on the standard curve. G Comparison of the relative abundance of five Fusobacterium nucleatum subspecies between the OSCC and HC groups based on metagenomic sequencing.

For the salivary microbiome, differential abundance was assessed using two independent strategies: the MaAsLin2 regression framework and SHAP analysis derived from machine learning classifiers. These complementary analyses identified 17 microbial species that consistently distinguished OSCC patients from HCs (Fig. 4C). Among these, Fusobacterium nucleatum (F.n) and Neisseria were significantly enriched in OSCC samples, whereas Porphyromonas gingivalis was more abundant in HCs (Fig. 4D and Supplementary Fig. 5A). The concordance between VOCs- and microbiome-derived biomarkers supports the biological plausibility of the fusion model’s decision-making process.

To systematically explore how these biomarkers interact, we performed network-based integrative correlation analysis linking differential metabolites, microbial taxa, and functional pathways. Methanethiol clustered with Fusobacterium nucleatum, Veillonella dispar, and Peptostreptococcus stomatis through sulfur metabolism pathways, whereas metabolites such as benzyl alcohol and methyl methacrylate were associated with Streptococcus salivarius through alternative metabolic routes (Supplementary Fig. 5B). These results demonstrate that methanethiol is not an isolated biomarker but functions within a broader interaction network connecting specific microbes, metabolic pathways and volatile compounds.

Given methanethiol’s central role in this network, we conducted quantitative analysis using PTR-TOF-MS with certified calibration gases. Six-point calibration curves demonstrated high linearity (R² = 0.996), and results confirmed a striking difference in methanethiol concentrations between groups: 44.6 ± 7.2 ppb in OSCC versus 12.2 ± 4.1 ppb in healthy controls, representing a 4- to 5-fold increase (Fig. 4E, F). In parallel, strain-level metagenomic profiling revealed that among five subspecies of Fusobacterium nucleatum, F.nucleatum subspecies polymorphum (F.np) showed the most significant enrichment in OSCC patients (Fig. 4G). This finding highlights a potential microbial driver of methanethiol accumulation in the oral environment, offering mechanistic insights into OSCC-associated metabolic reprogramming and pointing to potential targets for further investigation.

Source Analysis Identify F.n Producing Methanethiol through Sulfur Metabolism in OSCC

Following the identification of specific biomarkers and the establishment of diagnostic models, we investigated the biological interpretability and origins of these markers. Metabolic pathway analysis revealed that F.n shares significant sulfur metabolism pathways with methanethiol, specifically through the L-methionine biosynthesis IV and S-adenosyl-L-methionine salvage I pathways (Fig. 5A). Mediation analyses revealed significant relationships, indicating a deep biological linkage between F.n and methanethiol production via both pathways (Fig. 5B, C).

Fig. 5. Source Analysis of Key Biomarkers: Methanethiol and Fusobacterium nucleatum.

Fig. 5

A A MetaCyc pathway map showing the Top 20 differentially expressed pathways (including P. stomatis, F. nucleatum, and S. salivarius). B Mediation analysis among F.nucleatum, L-methionine biosynthesis IV, and methanethiol. C Mediation analysis among F.nucleatum, S-adenosyl-L-methionine salvage I, and methanethiol. D Schematic of co-culture experiments involving F.nucleatum and oral squamous cell carcinoma cell lines (Created with BioRender.com). E Methanethiol production under different multiplicities of infection (MOI = 1, 10, 50, 100) in the MOC1 cell line; Methanethiol production after 24 h of co-culture of F.nucleatum with oral squamous cell carcinoma cell lines. F Time course of methanethiol production at 0, 6, 12, 18, and 24 h in F.nucleatum and HN30 co-cultures. G, H Targeted LC-MS analysis of one-carbon metabolism in saliva. I Hypothetical model linking F.nucleatum/methanethiol with OSCC pathogenesis (Created with BioRender.com).

Gas-phase methanethiol concentrations were quantified by PTR-TOF-MS (Fig. 5D). All data represent the mean of three independent biological replicates. To elucidate the production phenotype, we first performed an MOI-optimization experiment using the MOC1 cell line by co-culturing F.n with cancer cells at MOI values of 1, 10, 50, and 100 (Fig. 5E), MOI 10 yielded the highest methanethiol signal, and was therefore selected for all subsequent assays. Using the optimized MOI of 10:1, we next conducted co-culture experiments with three OSCC cell lines (HN6, HN30, and CAL33) to assess methanethiol production differences across tumor cell backgrounds. Methanethiol levels in the F.n-OSCC co-culture group were significantly higher than those in F.n alone, OSCC cells alone, and blank controls (Fig. 5E), confirming that the interaction between F.n and OSCC cells promotes methanethiol generation. To further evaluate temporal production dynamics, HN30 cells were selected for a time-course co-culture experiment (0, 6, 12, 18, and 24 h). As shown in Fig. 5F, methanethiol concentrations increased progressively across the co-culture period. To determine whether methanethiol elevation was due to interaction-driven metabolic activation rather than bacterial proliferation, we quantified F.n colony-forming units (CFU) before and after co-culture using MOC1 matrices. As shown in Supplementary Fig. 6A, co-culture did not promote bacterial growth; instead, viable F.n counts decreased relative to the initial inoculum. This reduction is likely attributable to the use of DMEM-based culture medium, which does not support the anaerobic growth requirements of F.n, along with minor cell loss during differential centrifugation. These results confirm that methanethiol increase cannot be explained by bacterial overgrowth and is instead driven by interaction-associated metabolic responses27.

S-adenosylhomocysteine and S-adenosylmethionine serve as precursors within the shared sulfur metabolism pathways of F.n and methanethiol production28,29. Targeted LC-MS analysis of one-carbon metabolism in saliva samples from the external validation set revealed significantly elevated levels of S-adenosylhomocysteine (Fig. 5G) and S-adenosylmethionine (Fig. 5H) in OSCC patients compared with those in HCs, further validating their biological linkage. Collectively, these multilevel analyses support that F.n drives methanethiol production in OSCC through sulfur metabolic reprogramming, laying the mechanistic foundation for its diagnostic utility (Fig. 5I).

Our Open-Access Multiomics Platform Facilitates Clinically Actionable OSCC Diagnosis

Following comprehensive cohort sampling, multiomics integration, and diagnostic model development, we constructed and publicly launched an integrated online OSCC multiomics database (accessible at https://bio.futurecnn.com/). Developed using Streamlit, the database employs a layered application architecture comprising a data layer, a logic layer, and a user interface (UI) layer (Fig. 6A, B). The main functional modules of the website are as follows: (1) Home: Features external resource links and comprehensive introductions to research content and project details. (2) Search: Allows queries based on phenotype information and relevant dataset exploration. (3) Analysis: Provides embedded tools for principal component analysis (PCA), correlation analysis, and feature selection on integrated datasets. (4) Resource: Enables users to access multiomics data for external analysis. (5) Model Prediction: Facilitates disease diagnosis through loading user-specific volatile metabolomics or metagenomics data and using an integrated multiomics diagnostic model. (6) Tutorials and Q&A: Provides a step-by-step guidance module that introduces the functionality of each section and includes corresponding example datasets to facilitate practical use and reproducibility, while also incorporating a large language model to provide interactive and intelligent responses to user queries (Fig. 6C). Specifically, the embedded, trained multiomics fusion diagnostic model provides immediate online predictions. The model prediction module enables real-time detection using validated biomarkers (e.g., methanethiol and F.n abundance) with SHAP-based interpretability (Fig. 6D).

Fig. 6. Schematic and Functional Modules of the OSCC Multiomics Database.

Fig. 6

A Overview of the OSCC multiomics database integrating metagenomics, LC-MS, and PTR-TOF-MS data; website: https://bio.futurecnn.com/ (Created with BioRender.com). B Code architecture of the OSCC data platform. C, D Database module showcase: the Search, Analysis, Resource, Model Prediction, and Tutorials and Q&A modules with example outputs from the diagnostic model.

In summary, this study establishes a rapid, noninvasive AI-assisted diagnostic framework for OSCC, integrating breath VOC analysis, salivary metagenomics, and mechanistic biomarker validation into a publicly accessible multiomics platform and to demonstrate high accuracy and clinical translatability.

Discussion

In this study, we developed an integrated breathomic and microbiome analysis approach for the rapid, noninvasive, and AI-assisted diagnosis of OSCC. By combining a standardized PTR-TOF-MS workflow with machine learning, we achieved high diagnostic accuracy and identified robust biomarkers, notably methanethiol and Fusobacterium nucleatum (F.n). Beyond the diagnostic performance, these findings provide new biological insights into the metabolic reprogramming associated with oral carcinogenesis.

The integration of PTR-TOF-MS breathomics with AI-driven computational modeling advances the field of cancer diagnostics in several important ways. Breath analysis has long been considered promising for noninvasive cancer detection but has struggled with reproducibility, confounding factors, and lack of biomarker specificity30. Our study addresses these limitations by (i) optimizing sampling protocols to enrich oral cavity-derived VOCs while minimizing alveolar contamination, (ii) applying robust computational pipelines to reduce noise and improve interpretability, and (iii) integrating VOCs signatures with salivary microbial data to enhance specificity. Importantly, methanethiol emerged as a reproducible marker of OSCC, aligning with prior evidence that sulfur metabolism dysregulation is a hallmark of oral cancer31,32. The parallel enrichment of F.n further reinforces the mechanistic link between microbial activity and volatile metabolite production, adding a novel layer of interpretability to breath-based diagnostics.

Our novel platform provides a rapid, noninvasive screening tool capable of real-time analysis, complementing rather than replacing the gold-standard histopathological examination. This approach offers distinct advantages for early detection and initial risk stratification, particularly in settings where invasive procedures are impractical or as a first-line triage before confirmatory biopsy. By minimizing processing delays and sampling variability inherent in traditional workflows33, these refinements collectively address longstanding challenges in breath diagnostics34, such as intersample variability and confounding factors that have hindered the clinical adoption of earlier methods35.

The top 10 VOCs identified in our cohort fall into chemical classes that have been repeatedly associated with malignancies of the aerodigestive tract. Consistent with previous breathomics studies, aldehydes (e.g., acetaldehyde, propanal) are linked to oxidative stress and lipid peroxidation, and have been reported in lung, esophageal, prostate, and oral cancers36,37. Ketones such as acetone reflect metabolic reprogramming and have been observed in OSCC and colorectal cancer36. Aromatic hydrocarbons (ethylbenzene, styrene, toluene) can arise from membrane lipid degradation or environmental exposure and have been detected in lung, laryngeal, and oral cancers38. Alcohols (ethanol) and VSCs such as methanethiol and methyl mercaptan may originate from endogenous or microbial metabolism. Notably, ethanol36 and methanethiol32 have been previously proposed as OSCC or HNSCC breath biomarkers.

Our findings establish a standardized analytical framework for OSCC screening and provide a transferable paradigm for developing noninvasive, breath-based screening tools across oncology. This is especially relevant for malignancies where current screening options are invasive, inaccessible, or lack sensitivity for early-stage disease39,40. Future implementation studies should focus on validating this technology in real-world screening programs, emphasizing protocol standardization across diverse populations and integration into point-of-care or community health settings alongside established diagnostic pathways.

Our study provides compelling evidence that F.n contributes to OSCC pathogenesis not only through direct tumor–microbe interactions but also through the reshaping of metabolic pathways. Although numerous studies have established the association of F.n with OSCC progression41, primarily focusing on inflammation42, immune evasion43, and direct cellular interactions44, our findings reveal a novel and complementary mechanism for F.n-driven metabolic reprogramming. Previous studies have established F.n as a key oral pathobiont enriched in OSCC, with roles in chronic inflammation through activation of NF-κB and IL-6/STAT3 signaling42,45,46, immune evasion via modulation of T cells and natural killer cell responses47, and direct cellular interactions that promote epithelial-to-mesenchymal transition (EMT) and tumor invasiveness48,49. We revealed that F.n influences key sulfur metabolic pathways, particularly those involving L-methionine biosynthesis and S-adenosylmethionine recycling, which are directly linked to methanethiol production. This relationship was statistically validated using mediation analysis and further confirmed by in vitro experiments demonstrating the ability of F.n to induce methanethiol release from OSCC cells. This discovery expands our understanding of tumor–microenvironment interactions and provides a metabolic basis for the observed breath signatures in OSCC patients.

Although the co-culture experiments demonstrated consistent methanethiol elevation across multiple OSCC cell lines, several limitations should be acknowledged. The model relies on a single optimized MOI and a simplified DMEM-based microtube system that does not fully recapitulate the complex oral microenvironment. CFU quantification showed that F.n did not proliferate, indicating that methanethiol elevation reflects interaction–associated metabolic activation rather than bacterial overgrowth. However, the current system does not directly visualize physical association or quantify metabolic flux. Future work incorporating isotope-labeled tracing, improved anaerobic culture systems, and microenvironment-mimetic co-culture platforms will be required to define the causal pathways underlying this interaction.

This study has several limitations. First, our cohort was restricted to histologically confirmed OSCC and did not include precancerous or dysplastic lesions. As such, the utility of methanethiol and F.n as early-warning biomarkers remains to be validated. Second, the study population was geographically limited, raising the possibility of population-specific microbial or metabolic patterns; broader validation across diverse cohorts will be essential. Third, while our in vitro co-culture experiments support the causal role of F.n in methanethiol production, they cannot fully recapitulate the complexity of the in vivo tumor microenvironment. Fourth, while PTR-TOF-MS is a highly sensitive technique ideal for high-throughput, direct analysis, its lack of chromatographic separation presents a challenge for resolving structural isomers. Consequently, identification relies heavily on comparisons with existing databases and literature. To achieve definitive characterization, integration with chromatographic techniques like GC-MS is often necessary. Finally, breath and saliva samples were collected under controlled research conditions, and protocol adaptation for real-world clinical settings remains a challenge. Importantly, these limitations do not undermine our findings, but highlight the need for prospective, multicenter, longitudinal studies to confirm the generalizability and predictive power of our diagnostic framework.

This work demonstrates how integrating breathomics, microbiome profiling, and AI-driven modeling can generate a noninvasive, accurate, and mechanistically informed diagnostic tool for OSCC. The study exemplifies a broader principle that volatile metabolites produced through tumor-microbe interactions can serve as accessible and biologically interpretable biomarkers.

Methods

Study Cohort

A total of 305 individuals were screened for eligibility between July 2024 and May 2025. Of these, 281 met inclusion criteria, and 222 consented to participate (111 histopathologically confirmed OSCC patients and 111 healthy controls), resulting in a final analytical cohort of 222 individuals. In addition, an independent external validation cohort comprising 53 OSCC patients and 30 controls was recruited for confirmatory testing. Eligibility for OSCC patients required: (1) a primary or recurrent OSCC diagnosis, (2) no prior oncological treatment before sample collection, and (3) good periodontal health or completion of periodontal treatment. Healthy controls were required to have: (1) no history of malignancy, (2) no detectable oral lesions, and (3) no recent oral surgical procedures. Exclusion criteria for both groups included: (1) other active malignancies, (2) severe systemic diseases (e.g., COPD, cirrhosis), (3) antibiotic or immunosuppressant use within the last 3 months, and (4) inability to provide informed consent. Demographic and clinical variables, including age, sex, smoking, alcohol use, betel quid chewing, oral hygiene status, and family history, were collected for all participants.

To evaluate the generalizability and robustness of the diagnostic model, an independent external validation cohort of 83 individuals (53 OSCC and 30 healthy controls) was recruited. This validation cohort differed from the primary cohort in two key aspects: (1) Sampling period: all samples in the external cohort were collected at least 3 months later than those in the primary cohort, reducing temporal and environmental overlap. (2) Clinical source: patients in the external cohort were recruited from the Pudong branch of Shanghai Ninth People’s Hospital, whereas the primary cohort was collected at the main Huangpu campus. Although both cohorts were derived from the same hospital system, the use of distinct clinical sites introduces institutional heterogeneity, which enhances the model’s evaluation by testing its performance across real-world variations in patient demographics and clinical practices.

This study was conducted in accordance with the Declaration of Helsinki50. Ethical approval was obtained from the Ethics Committee of the Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine (Approval No: SH9H-2025-T92-2; Clinical trial number: not applicable). Written informed consent was obtained from all participants prior to their inclusion in the study. Additionally, consent for the publication of any potentially identifiable images or data included in this article was also obtained from the participants.

Exhaled Breath Collection

Breath samples were collected using 50 mL polyvinyl fluoride (PVF) sampling bags (ALIBEN Science & Technology, China), following a standardized protocol. Before each collection, PVF bags were flushed five times with 99.999% pure nitrogen (Shanghai Wetry Standard Gas Analysis Technology Co., LTD) to eliminate contaminants.

Prior to sampling, participants fasted for ≥10 hours and refrained from consuming odoriferous foods (e.g., garlic, onions), alcohol, or tobacco for 24 h. Sampling was conducted in a controlled environment (temperature: 22 ± 2 °C; humidity: 50 ± 5%). After 3 minutes of normal breathing to stabilize respiratory patterns, subjects performed closed-mouth nasal breathing for 10 minutes to allow oral air incubation. Participants exhaled 50 mL of oral breath through a sterile, disposable plastic mouthpiece (single-use, medical-grade polypropylene) into PVF sampling bags. Triplicate samples were collected per participant, with concurrent ambient air samples obtained as background controls (1-2 m from subject at equivalent height). Immediately following collection, samples were transported to the laboratory at ambient temperature for batch processing within 2 h to minimize VOCs degradation51.

PTR-TOF-MS Analysis

Breath samples were analyzed using a PTR-TOF-MS 2500 instrument (ALIBEN Science & Technology, Sichuan, China), based on proton transfer reaction principles. Exhaled gases were introduced into the instrument through a heated PEEK capillary (2.10 mm i.d., 1.50 m length; 70°C), preventing condensation of humidified breath. The ion source was also maintained at 70°C. In the drift tube, VOCs were ionized via proton transfer with H3O+ ions, and detected by time-of-flight mass spectrometry, generating intensity and m/z data. After analysis, the system was purged with 99.999% pure nitrogen for 2 minutes before the next sample. Key operating conditions are provided in Supplementary Table 8.

PTR-TOF-MS Data Preprocessing

All analyses were performed in Python (v3.10.14), integrating multidimensional data into structured feature matrices for subsequent predictive modeling. The preprocessing workflow for PTR-TOF-MS raw spectra focused on peak detection, quality control, normalization, and batch correction.

Raw Data Parsing and Peak Extraction: Custom algorithms were used to extract mass-to-charge ratio (m/z), retention time (RT), and peak intensity information from raw spectra. Processing steps included peak detection, baseline correction, and alignment across samples, followed by calculation of integrated peak areas.

Quality Control and Feature Filtering: (1) Low-reproducibility filtering: Features detected in fewer than 80% of samples were excluded. (2) Background removal: Potential matrix interference peaks were identified by comparing feature intensities between experimental samples and blank controls. Features with high blank/sample ratios were discarded.

Data Normalization: To minimize technical variation across samples, Total Ion Current (TIC) normalization was applied. Peak areas for each metabolite were divided by the sum of all peak areas within the same sample, thereby correcting for variations in sample loading and ionization efficiency. This procedure was implemented using sklearn.preprocessing.Normalizer with parameter norm= 'l1'.

Batch Effect Correction: For datasets spanning multiple experimental batches, systematic variation was corrected using the ComBat algorithm. Normalized intensity matrices were combined with batch labels as covariates to adjust for batch-associated artifacts.

Compound Annotation

Characteristic m/z features were selected based on significant differences between breath and ambient air using statistical methods. Raw spectra were visualized to confirm peak shape and quality, excluding inconsistent or noisy features. Refined m/z features were cross-referenced with GLOVOCs52, the Human Breathomics Database (https://hbdb.cmdm.tw/)53, and Metabolomics Workbench54, as well as relevant literature, for compound assignment.

Complementary GC-MS Identification of VOCs

To aid compound identification, pooled breath samples from healthy participants were analyzed by GC-MS. VOCs were pre-concentrated by SPME (DVB/CAR/PDMS fiber, 30 min at 37°C) and desorbed into a Thermo Scientific TRACE 1300 GC-TSQ 8000 MS system using a VF-624 column. GC oven and MS conditions are described in Supplementary Table 9. Spectra were processed with AMDIS (v2.66), and tentative identifications were made by spectral matching against the NIST library. GC-MS results were compared to PTR-TOF-MS features for annotation validation.

VOCs Biomarker Selection from PTR-TOF-MS Data

To identify robust volatile biomarkers, four complementary approaches were applied: Mann–Whitney U test, LASSO-Boruta, Random Forest (RF), and LightGBM (LGBM). (1) Mann-Whitney U test with FDR correction retained features with q < 0.05. (2) LASSO logistic regression with five-fold CV selected non-zero coefficient features, further confirmed by Boruta. (3) Model-based selection, RF and LGBM models were trained, and SHAP values quantified feature contributions; the top 30% features were retained. (4) The intersection of U test, LASSO-Boruta, and SHAP-filtered RF/LGBM features defined the final VOCs biomarker set.

Biomarker Quantification

Methanethiol was quantitatively analyzed using a standardized calibration protocol. Certified reference gas (1.10 × 10-6 mol/mol in nitrogen, 10 MPa cylinder) was obtained from Shanghai Wetry Standard Gas Analysis Technology Co., Ltd (Shanghai, China). Six-point calibration curves (0, 10, 20, 30, 40, and 50 ppb) were generated using a portable gas dilution system (HT-702C, ALIBEN Science & Technology, Sichuan, China) with demonstrated linearity (R2 = 0.996).

Saliva Collection

Following exhaled breath collection, saliva samples were obtained from participants after completion of all breath analysis procedures. After maintaining a ≥ 10-hour fasting period with avoidance of oral interventions (including brushing, mouthwash use, or eating/drinking), unstimulated whole saliva was collected between 6:00 and 10:00 AM using the established Navazesh method (1992)55. Participants rinsed with 30 mL physiological saline for 1 min, then expectorated into sterile tubes. Samples were aliquoted into six 5 mL cryovials, flash-frozen in liquid nitrogen (5–10 min), stored on dry ice during transport, and maintained at -80 °C until analysis.

Metagenomic Sequencing

DNA was extracted from saliva samples using the QIAamp DNA Mini Kit (Qiagen) with sterile deionized water as negative controls. Library preparation was performed with the NEBNext Ultra II DNA Library Prep Kit (New England Biolabs), followed by 150 bp paired-end sequencing on a NovaSeq XPlus platform (Illumina) with ≥6 Gb raw data per sample.

Quality-filtered reads were analyzed using HUMAnN3 for taxonomic profiling and MetaPhlAn4 for functional annotation, with alignment to integrated reference databases (UniRef90, Chocophlan).

Metagenomic Data Processing

Standardized bioinformatic analyses were performed on metagenomic data. Quality control was conducted using Kneaddata (v0.10.0) to remove host-derived sequences, ensuring that downstream analyses focused exclusively on microbial sequences. Taxonomic annotation was carried out with MetaPhlAn (v4.0.6) for species- and strain-level classification. Functional pathway abundance was predicted using HUMAnN (v3.9), with metabolic pathway abundance calculated based on the MetaCyc database.

Microbiome Diversity and Statistical Analysis

All microbiome analyses were performed in R (v4.1.0). Community structure differences were assessed using PERMANOVA (permutational multivariate analysis of variance) with the adonis function in the vegan package, applying 999 permutations. α-diversity was calculated using the Simpson index (vegan package), while β-diversity was assessed with principal coordinate analysis (PCoA) based on Bray-Curtis distances.

Metabolomic OPLS-DA Analysis

Metabolomic data were normalized to relative abundance (metabolite peak area as a percentage of the total sample peak area) to minimize inter-sample variation in total metabolite levels. Orthogonal partial least squares discriminant analysis (OPLS-DA) was then performed using the ropls package (v1.30.0), with metabolite abundance matrices as predictors (X) and sample group information as the response variable (Y). Orthogonal signal correction was applied to separate predictive variation from noise, improving group discrimination. Model validity was assessed through 1000 permutation tests.

Feature Importance Analysis Based on Shapley Values

The iml package in R was used to compute Shapley values, quantifying the contribution of each microbial genus to model predictions. Shapley values were calculated per sample and aggregated to evaluate the overall importance of each genus across predictions.

Differential Microbial Analysis (MaAsLin2)

Differential microbial abundance was analyzed using the MaAsLin2 package. Fixed effects included disease status, while random effects included height, weight, sex, and age. A minimum detection threshold of 0.001 was applied, and genera with FDR-adjusted q-values < 0.05 were considered statistically significant.

Mediation Analysis

The mediation package in R was applied with 1000 bootstrap resamples to calculate the average causal mediation effect (ACME), average direct effect (ADE), and mediation proportion. Statistical significance was determined using p-values and confidence intervals.

Statistical Analysis

All statistical analyses were performed using SPSS 25.0 and R (v4.1.0). Continuous variables were tested for normality using the Shapiro–Wilk test and expressed as mean ± standard deviation (normal distribution) or median with interquartile range (non-normal distribution). Between-group comparisons were made using independent t-tests (normal distribution) or Wilcoxon rank-sum tests (non-normal distribution). Odds ratios (ORs) were calculated using binary logistic regression. Multiple comparisons were corrected using the Benjamini–Hochberg FDR method, with statistical significance defined as p < 0.05 and FDR < 0.05.

Microbiota–Function–Metabolite Network Construction

Based on differentially identified microbial genera, functional pathways, and metabolites, Spearman correlation analysis was performed to calculate associations: (1) between differential genera and pathways, and (2) between pathways and metabolites. Only significant positive associations (r > 0.3, FDR < 0.05) were retained. Networks were visualized using Cytoscape (v3.9). Node size was scaled according to the absolute log2 fold change ( |log2FC| ), edge width was scaled to correlation strength (r), and network topology was optimized using a force-directed layout.

Diagnostic Model Construction (Single-Omics)

For each omics dataset (PTR-TOF-MS VOCs and salivary metagenomics), individual diagnostic models were constructed using the AutoML framework AutoGluon (v1.2.0). Datasets were split into training (80%) and validation (20%) sets using stratified random sampling to preserve class balance. AutoGluon automatically handled data preprocessing (missing value imputation, normalization, and feature formatting), followed by large-scale model exploration across numerous candidate algorithms. From this pool, the top five performing models were compared, and the best-performing model was selected for further optimization.

Logistic Regression, a linear baseline classifier assuming proportional log-odds relationships between features and outcomes; Random Forest (RF), an ensemble of decision trees using bootstrap aggregation, which improves robustness but may sacrifice fine-grained optimization; Gradient Boosting Classifier (GBC) and Extreme Gradient Boosting (XGBoost), tree-based boosting algorithms that sequentially minimize residual errors, generally improving predictive power for structured tabular data; Light Gradient Boosting Machine (LGBM), a gradient boosting framework that uses histogram-based binning and leaf-wise growth, providing faster training, efficient memory use.

To refine performance, hyperparameter tuning was conducted with Optuna (v3.2.0) using 600 iterations of 10-fold cross-validation. Optimization trajectories, parallel coordinate visualizations of parameter, accuracy combinations, and averaged feature importance across all parameter sets confirmed the stability of the tuning process.

Final model performance was evaluated on the held-out validation set using multiple metrics, including accuracy, precision, recall, F beta, and ROC-AUC. ROC curves and confusion matrices were generated to visualize discriminative ability. AutoGluon’s leaderboard function summarized comparative results across all models, ensuring transparent selection of the most robust diagnostic classifier.

Multi-Omics Fusion Model Construction

To achieve maximal diagnostic performance, we constructed a multi-omics fusion framework that integrates PTR-TOF-MS derived VOC features with salivary microbiome profiles. Prior to model development, the combined feature matrix underwent preprocessing to address potential inter-omics heterogeneity and batch effects. Specifically, proportion-standardized inputs were processed using PSI (Population Stability Index) assessment, following which ComBat harmonization and log1p transformation were applied to reduce batch-related biases arising from experimental conditions and microbial abundance distributions.

For model development, we designed an innovative three-level stacked ensemble architecture tailored for high-dimensional biomedical classification tasks. (1) Level 0 (Base Learners): Three complementary algorithmic families: LGBM, XGBoost, and logistic regression were trained on each omics modality. To enhance robustness, bagging was performed for each model type, generating 10 bootstrapped learners, yielding a diverse ensemble of base predictors. (2) Level 1 (Intermediate Integration Layer): Predictions from all Level 0 learners were concatenated with the original omics features to form an expanded meta-feature matrix that captures both raw biological signatures and model-derived representations. (3) Level 2 (Meta-Learner): A multi-layer perceptron (MLP) implemented in PyTorch served as the meta-learning module. The network contained two hidden layers (64 and 32 units) with ReLU activation. Hyperparameters, including learning rate (0.001–0.01) and L2 regularization (1e-4 to 1e-2), were optimized via Optuna, using 600 iterations of 5-fold cross-validation targeting maximal ROC-AUC.

Model training followed a 5-fold cross-validation strategy for internal tuning, after which the final optimized framework was evaluated on the held-out internal test set and on an independent external validation cohort. ROC curves were generated to assess discriminative power across discovery, test, and external validation phases.

LC-MS/MS Analysis of One-Carbon and Sulfur Metabolites

Targeted LC-MS/MS analysis of one-carbon metabolism-related intermediates was performed on saliva samples from the external validation cohort, consisting of 17 patients with OSCC and 17 HCs. Liquid chromatography-tandem mass spectrometry (LC-MS/MS) was employed to quantify one-carbon metabolism-related intermediates.

  1. Chemicals and reagents: HPLC-grade acetonitrile (ACN) and methanol (MeOH) were purchased from Merck (Darmstadt, Germany). MilliQ water (Millipore, Bradford, USA) was used in all experiments. Ammonium acetate was purchased from Sigma-Aldrich. All of the standards were purchased from Sigma-Aldrich/ISOREAG. The stock solutions of standards were prepared at the concentration of 1 mg/mL in MeOH. All stock solutions were stored at -20°C. The stock solutions were diluted with MeOH to working solutions before analysis.

  2. Sample preparation and extraction: Two milliliters of the sample were subjected to vacuum freeze-drying, then reconstitute with 250 μL of 70% methanol/water. 10 μL internal standard mixed solution (250 ng/mL) was added into the extract as internal standards (IS) for the quantification. Then the extract was vortexed for 3 min, then kept in a refrigerator at -20 °C for 30 min, centrifuge at 12000 r/min for 10 min at 4 °C, 150 μL of the supernatant was collected. The supernatant was again centrifuged at 12000 r/min for 5 min at 4 °C, 100 μL of the supernatant was transferred for further LC-MS analysis.

  3. UPLC Conditions: The sample extracts were analyzed using an LC-ESI-MS/MS system (UPLC, ExionLC AD, Sciex, Singapore; MS, QTRAP® 6500+ System, Sciex, Singapore). The analytical conditions were as follows, HPLC: column, Thermo Scientific™ Accucore™ 150 Amide hilic (100 mm×2.1 mm i.d., 2.6 µm); solvent system, water with 10 mM ammonium acetate (A), acetonitrile (B); The gradient was started at 95% B (0–1 min), increased to 50% B (1–6 min), 50% B (6–7.5 min), finally ramped back to 95% B (7.6–10 min); flow rate, 0.66 mL/min; temperature, 40 °C; injection volume: 2 μL.

  4. ESI-MS/MS Conditions: AB 6500 + QTRAP® LC-MS/MS System, equipped with an ESI Turbo Ion-Spray interface, operating in positive ion modes and controlled by Analyst 1.6 software (AB Sciex). The ESI source operation parameters were as follows: ion source, turbo spray; source temperature 550°C; ion spray voltage (IS) 5500 V; curtain gas (CUR) were set at 35.0 psi; DP and CE for individual MRM transitions was done with further DP and CE optimization. A specific set of MRM transitions were monitored for each period according to the metabolites eluted within this period.

Cell Lines and Bacteria Cultures

The human HNSCC cell lines HN6, HN30 and CAL33 were all cultured in high glucose DMEM medium (BasalMedia, China) supplemented with 10% fetal bovine serum (FBS; Cellmax, China) and 1% penicillin-streptomycin (P/S; NCM Biotech, China). The human oral keratinocytes (HOK) were maintained in defined keratinocyte serum-free basal medium (DK-SFM; Thermo Fisher Scientific, USA). The murine oral squamous cell carcinoma line MOC1 was cultured in growth medium composed of IMDM:F12 (2:1 ratio) (BasalMedia, China) supplemented with 10% FBS, 5 mg/L insulin (Selleck, China), 40 μg/L hydrocortisone (YuanYe, China), 5 μg/L epidermal growth factor (EGF) (Anweisci, China), and 1% P/S. All cells were incubated in a humidified atmosphere with 5% CO2 at 37°C.

Fusobacterium nucleatum (ATCC 10953) were cultured anaerobically in Fluid Thioglycollate Medium (FTG; MingzhouBio, China) at 37°C (90% N2, 5% CO2, 5% H2).

Direct Cell Co-culture

HOK cells and OSCC tumor cell lines (including HN6, HN30, CAL33 and MOC1) were trypsinized from 100 mm culture dishes, then centrifuged at 1200 rpm for 3 min and resuspended in high glucose DMEM basal medium without serum, penicillin and streptomycin. Following cell counting, the suspension was adjusted to a density of 3 × 106 viable cells per tube. F.n was monitored by measuring the optical density at 600 nm (OD value), when reaching plateau phase, F.n was centrifuged at 4500 rpm for 5 minutes and resuspended in basal medium mentioned above. Cells resuspended in centrifuge tubes were stimulated by live F.n at gradient MOI (1:1, 10, 50, 100) for 24 h, or stimulated at MOI = 1:10 for 0, 6, 12, 18, 24 h. All co-culture, VOC quantification, and CFU assays were performed with three independent biological replicates unless otherwise specified.

Sample Preparation and Analysis

For optimal bacterial-cancer cell interaction, the F.n suspension was carefully mixed with cancer cells by performing repeated pipetting in 15 mL centrifuge tubes. The tube openings were sealed with Parafilm® and co-cultured under specified experimental conditions as previously described. For analysis, the probe of the analytical instrument was aseptically punctured through the Parafilm® to maintain a gas-tight microenvironment that prevents atmospheric exchange between co-culture-generated volatiles and the external environment during measurement.

Bacterial CFU Enumeration

For effective separation of cancer cells and bacteria from direct co-culture systems, we employed differential centrifugation. Specifically, the co-culture was first centrifuged at 300 ×g for 5 min at 4 °C to pellet the cancer cells. The bacterial-containing supernatant was then collected and further centrifuged at 8000 × g for 10 min at 4 °C to isolate the bacterial fraction. Then bacterial suspension of F.n was subjected to optical density measurement at 600 nm, with the absorbance values then converted to colony-forming unit (CFU) counts using established calibration curves in 0 h and 24 h.

Database Construction and Online Platform Development

An integrated online platform was developed using Streamlit (v1.34.0) to enable multi-omics data visualization, interactive analysis, and diagnostic prediction. The platform adopts a three-tier architecture.

  1. Data layer: Multi-omics datasets, including VOC mass spectrometry data (m/z, retention time, peak intensity), metagenomic abundance tables, clinical phenotype information, and analysis outputs (e.g., differential features, model parameters), were stored in a structured SQLite database. Data I/O was handled through Python’s sqlite3 library, supporting batch imports in .csv format and query responses within <1 s.

  2. Logic layer: Core computational functions were implemented in Python. Data preprocessing used pandas and scikit-learn for cleaning, standardization, and imputation. Statistical analyses applied scipy and statsmodels, while machine learning leveraged pre-trained model weights for diagnostic prediction and shap for interpretability. A large language model interface was integrated via the Qwen 2.5-7B-Instruct API to support intelligent Q&A functionality.

  3. User interface layer: Built on Streamlit’s modular design, the UI incorporated st.pyplot() for visualization, st.dataframe() for tabular display, and st.file_uploader for user data input. Responsive design ensured compatibility across major browsers (Chrome, Firefox, Safari).

The platform includes seven functional modules: (1) Home for project overview and external links, (2) Search for phenotype-based queries, (3) Analysis offering PCA, correlation, and feature selection tools, (4) Resource for data access, (5) Diagnostic Prediction enabling user-uploaded VOC or metagenomic data to generate online diagnostic predictions with SHAP-based interpretability, (6) Q&A integrating a large language model for interactive question answering, and (7) Tutorial, a comprehensive guidance module that explains the usage of each platform function and provides matched example datasets for step-by-step practice.

Navigation was implemented via a fixed sidebar for module switching. Each page provided collapsible instruction panels and code examples to guide user interaction. Real-time input validation (e.g., format checks on uploaded files) ensured smooth operation and user-friendly experience.

Supplementary information

Supplementary Data 1 (35.2KB, xlsx)

Acknowledgements

This project was supported by the National Natural Science Foundation of China Outstanding Youth Fund (Grant No. 62322114) and the Medical Engineering Cross Foundation of Shanghai Jiao Tong University (Grant No. YG2023LC06) and the National Natural Science Foundation of China (Grant No. 82272815).

Author contributions

Y.S.: Conceptualization; methodology; software; data curation; investigation; validation; project administration; resources; visualization; writing—original draft; writing—review and editing; supervision. X.H.: Methodology; data curation; validation; visualization; writing—original draft; writing—review and editing. J.H.: Methodology; data curation; formal analysis. Y.W.: Methodology; visualization. J.L.(Luo): Resources; writing—review and editing. J.Y.: Visualization, writing—review and editing. Y.D.: Supervision; writing—review and editing. X.W.: Supervision; methodology; funding acquisition; writing—review and editing; project administration. J.L.(Liu): Supervision; methodology; funding acquisition; writing—review and editing; project administration.

Data availability

The original data generated and analyzed in this study are included in the article and Supplementary Material. Additional information or raw data can be obtained from the corresponding author upon reasonable request. The diagnostic model developed in this study has been fully open-sourced and is publicly accessible at: https://github.com/SunYilan/biomodel.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yilan Sun, Xin Hu.

Contributor Information

Xu Wang, Email: wangx312016@sh9hospital.org.cn.

Jiannan Liu, Email: liujiannan@sh9hospital.org.cn.

Supplementary information

The online version contains supplementary material available at 10.1038/s41746-026-02527-3.

References

  • 1.Tan, Y. et al. Oral Squamous Cell Carcinomas: State of the Field and Emerging Directions. Int J. Oral. Sci.15, 44 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Bray, F. et al. Global Cancer Statistics 2022: Globocan Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. Ca: A Cancer J. Clinicians74, 229–263 (2024). [DOI] [PubMed] [Google Scholar]
  • 3.Silveira, F. M., Schuch, L. F. & Bologna-Molina, R. Classificatory Updates in Verrucous and Cuniculatum Carcinomas: Insights From the 5Th Edition of Who-Iarc Head and Neck Tumor Classification. World J. Clin. Oncol.15, 464–467 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ken, K. et al. Doublet Chemotherapy, Triplet Chemotherapy, Or Doublet Chemotherapy Combined with Radiotherapy as Neoadjuvant Treatment for Locally Advanced Oesophageal Cancer (Jcog1109 Next): A Randomised, Controlled, Open-Label, Phase 3 Trial. Lancet404, 55–66 (2024). [DOI] [PubMed] [Google Scholar]
  • 5.Ahmed, A. et al. Mutation Detection in Saliva From Oral Cancer Patients. Oral. Oncol.151, 106717 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lennon, A. M. et al. Feasibility of Blood Testing Combined with Pet-Ct to Screen for Cancer and Guide Intervention. Science369, eabb9601 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Feng, X. et al. Cancer stage compared with mortality as end points in randomized clinical trials of cancer screening: a systematic review and meta-analysis. Jama-J. Am. Med. Assoc.331, 1910–1917 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Yan, B. et al. Genai synthesis of histopathological images from raman imaging for intraoperative tongue squamous cell carcinoma assessment. Int. J. Oral. Sci.17, 12 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Balakittnen, J. et al. A novel saliva-based mirna profile to diagnose and predict oral cancer. Int. J. Oral. Sci.16, 14 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Song, M., Bai, H., Zhang, P., Zhou, X. & Ying, B. Promising applications of human-derived saliva biomarker testing in clinical diagnostics. Int. J. Oral. Sci.15, 2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhou, M. et al. Exhaled breath and urinary volatile organic compounds (Vocs) for cancer diagnoses, and microbial-related Voc metabolic pathway analysis: a systematic review and meta-analysis. Int. J. Surg.110, 1755–1769 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Li, L. et al. Qualitative and quantitative transformer-cnn algorithm models for the screening of exhale biomarkers of early lung cancer patients. Anal. Chem.97, 6651–6660 (2025). [DOI] [PubMed] [Google Scholar]
  • 13.Li, X. et al. Volatile organic compounds in exhaled breath: a promising approach for accurate differentiation of lung adenocarcinoma and squamous cell carcinoma. J. Breath Res. 18 (2024). [DOI] [PubMed]
  • 14.Roquencourt, C., Grassin-Delyle, S. & Thevenot, E. A. Ptairms: real-time processing and analysis of Ptr-Tof-Ms data for biomarker discovery in exhaled breath. Bioinformatics38, 1930–1937 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Roquencourt, C., Lamy, E., Bardin, E., Devillier, P., & Grassin-Delyle, S. A benchmark study of data normalisation methods for Ptr-Tof-Ms exhaled breath metabolomics. J. Breath Res. 18 (2023). [DOI] [PubMed]
  • 16.Wang, H. et al. A combined screening study for evaluating the potential of exhaled acetone, isoprene, and nitric oxide as biomarkers of lung cancer. Rsc Adv.13, 31835–31843 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zhang, K. et al. A generalist vision-language foundation model for diverse biomedical tasks. Nat. Med.30, 3129–3141 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Beck, A. G. et al. Recent developments in machine learning for mass spectrometry. Acs Meas. Sci. Au.4, 233–246 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Brinkman, P. et al. Fulfilling the promise of breathomics: considerations for the discovery and validation of exhaled volatile biomarkers. Am. J. Resp. Crit. Care210, 1079–1090 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Deng, F. et al. A novel accurate peak extraction algorithm of mass spectrometry based on iterative adaptive curve fitting. J. Am. Soc. Mass Spectr.35, 2900–2909 (2024). [DOI] [PubMed] [Google Scholar]
  • 21.Garg, M. et al. Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank. Nat. Genet.56, 1821–1831 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Melnikov, A. D., Tsentalovich, Y. P. & Yanshole, V. V. Deep learning for the precise peak detection in high-resolution Lc-Ms Data. Anal. Chem.92, 588–592 (2020). [DOI] [PubMed] [Google Scholar]
  • 23.Sun, Y. et al. Integrative plasma and fecal metabolomics identify functional metabolites in adenoma-colorectal cancer progression and as early diagnostic biomarkers. Cancer Cell42, 1386–1400 (2024). [DOI] [PubMed] [Google Scholar]
  • 24.He, X., Zhao, K. & Chu, X. Automl: a survey of the state-of-the-art. Knowl. -Based Syst.212, 106622 (2020). [Google Scholar]
  • 25.Pinichka, C., Chotpantarat, S., Cho, K. H. & Siriwong, W. Comparative analysis of swat and swat coupled with xgboost model using optuna hyperparameter optimization for nutrient simulation: a case study in the upper Nan River Basin, Thailand. J. Environ. Manag.388, 126053 (2025). [DOI] [PubMed] [Google Scholar]
  • 26.Lu, S., Song, W., Pfob, A. & Gibbons, C. Assessing the representativeness of large medical data using population stability index. Bmc Med. Res. Methodol.25, 44 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Gu, Y. et al. Influence of the densities and nutritional components of bacterial colonies on the culture-enriched gut bacterial community structure. Amb. Express11, 78 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Hara, T., Sakanaka, A., Lamont, R. J., Amano, A. & Kuboniwa, M. Interspecies metabolite transfer fuels the methionine metabolism of fusobacterium nucleatum to stimulate volatile methyl mercaptan production. Msystems9, e76423 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Alon-Maimon, T., Mandelboim, O. & Bachrach, G. Fusobacterium Nucleatum and Cancer. Periodontol 200089, 166–180 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Vassilenko, V., Moura, P. C., & Raposo, M. Diagnosis of carcinogenic pathologies through breath biomarkers: present and future trends. Biomedicines11, (2023). [DOI] [PMC free article] [PubMed]
  • 31.Philipp, T. M., Scheller, A. S., Krafczyk, N., Klotz, L., & Steinbrenner, H. Methanethiol: a scent mark of dysregulated sulfur metabolism in cancer. Antioxidants-Basel12, (2023). [DOI] [PMC free article] [PubMed]
  • 32.Kwon, I. et al. Detection of Volatile Sulfur Compounds (Vscs) in Exhaled Breath as a Potential Diagnostic Method for Oral Squamous Cell Carcinoma. BMC Oral. Health22, 268 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Henderson, B. et al. The Peppermint Breath Test Benchmark for Ptr-Ms and Sift-Ms. J. Breath Res. 15 (2021). [DOI] [PubMed]
  • 34.Desai, K. M. et al. Screening of oral potentially malignant disorders and oral cancer using deep learning models. Sci. Rep. -UK15, 17949 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Golsanamlu, Z., Jin, H., Soleymani, J. & Jouyban, A. Application of high-resolution analytical techniques in breathomics studies. Microchem. J.211, 113073 (2025). [Google Scholar]
  • 36.Borden, S. A. et al. Characterizing volatile organic compound profiles in oral cancer using multiple sample collection approaches by Gc-Ims and Td-Gc-Ms. Sci. Rep. -Uk15, 37014 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Mentel, S. et al. Prediction of oral squamous cell carcinoma based on machine learning of breath samples: a prospective controlled study. Bmc Oral Health 21 (2021). [DOI] [PMC free article] [PubMed]
  • 38.Bouza, M., Gonzalez-Soto, J., Pereiro, R., de Vicente, J. C. & Sanz-Medel, A. Exhaled breath and oral cavity Vocs as potential biomarkers in oral cancer patients. J. Breath. Res.11, 16015 (2017). [DOI] [PubMed] [Google Scholar]
  • 39.Jia, Z. et al. Advanced strategy for cancer detection based on volatile organic compounds in breath. J. Nanobiotechnol.23, 468 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Le, T. & Priefer, R. Detection technologies of volatile organic compounds in the breath for cancer diagnoses. Talanta265, 124767 (2023). [DOI] [PubMed] [Google Scholar]
  • 41.Fitzsimonds, Z. R., Rodriguez-Hernandez, C. J., Bagaitkar, J. & Lamont, R. J. From beyond the pale to the pale riders: the emerging association of bacteria with oral cancer. J. Dent. Res.99, 604–612 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Xiang, Z. et al. Fusobacterium nucleatum Exacerbates Colitis via Stat3 activation induced by acetyl-Coa accumulation. Gut Microbes17, 2489070 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Sun, J. et al. F.nucleatum facilitates oral squamous cell carcinoma progression via Glut1-driven lactate production. Ebiomedicine88, 104444 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Cavallucci, V. et al. Proinflammatory and cancer-promoting pathobiont fusobacterium nucleatum directly targets colorectal cancer stem cells. Biomolecules 12 (2022). [DOI] [PMC free article] [PubMed]
  • 45.Nie, F. et al. The role of Cxcl2-mediated crosstalk between tumor cells and macrophages in fusobacterium nucleatum-promoted oral squamous cell carcinoma progression. Cell Death Dis.15, 277 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Wang, Y. et al. Study of the Inflammatory Activating Process in the Early Stage of Fusobacterium Nucleatum Infected Pdlscs. Int. J. Oral. Sci.15, 8 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Sun, J., Chen, F. & Wu, G. Potential effects of gut microbiota on host cancers: focus on immunity, DNA damage, cellular pathways, and anticancer therapy. Isme J.17, 1535–1551 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Zhu, H. et al. Fusobacterium Nucleatum Promotes Tumor Progression in Kras P.G12D-Mutant Colorectal Cancer by Binding to Dhx15. Nat. Commun.15, 1688 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Dan, W., Xiong, C., Zhou, G., Chen, J. & Pan, F. Gut microbiota as a mediator of cancer development and management: from colitis to colitis-associated dysplasia and carcinoma. BBA Rev. Cancer1880, 189381 (2025). [DOI] [PubMed] [Google Scholar]
  • 50.Association, W. M World medical association declaration of Helsinki: ethical principles for medical research involving human subjects. Jama-J. Am. Med. Assoc.310, 2191–2194 (2013). [DOI] [PubMed] [Google Scholar]
  • 51.Czippelova, B. et al. Impact of Breath Sample Collection Method and Length of Storage of Breath Samples in Tedlar Bags On the Level of Selected Volatiles Assessed Using Gas Chromatography-Ion Mobility Spectrometry (Gc-Ims). J. Breath Res.18 (2024). [DOI] [PubMed]
  • 52.Yáñez-Serrano, A. M. et al. Glovocs - master compound assignment guide for proton transfer reaction mass spectrometry users. Atmos. Environ.10, 689–698 (2021). [Google Scholar]
  • 53.Kuo, T. et al. Human Breathomics Database. Database-Oxford 2020 (2020). [DOI] [PMC free article] [PubMed]
  • 54.He, Y. et al. Plasma metabolomics dataset of race-walking athletes illuminating systemic metabolic reaction of exercise. Sci. Data12, 448 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Navazesh, M., Mulligan, R. A., Kipnis, V., Denny, P. A. & Denny, P. C. Comparison of whole saliva flow rates and mucin concentrations in healthy caucasian young and aged adults. J. Dent. Res.71, 1275–1278 (1992). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Data 1 (35.2KB, xlsx)

Data Availability Statement

The original data generated and analyzed in this study are included in the article and Supplementary Material. Additional information or raw data can be obtained from the corresponding author upon reasonable request. The diagnostic model developed in this study has been fully open-sourced and is publicly accessible at: https://github.com/SunYilan/biomodel.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES