Skip to main content
Cancer & Metabolism logoLink to Cancer & Metabolism
. 2026 Apr 1;14:8. doi: 10.1186/s40170-026-00427-4

Leveraging untargeted metabolomics in combination with machine learning to uncover novel insights into bladder cancer

Abu Hena Mostafa Kamal 1,2,#, Vasanta Putluri 2,#, Tanmay Gandhi 2,4,#, Chandra Shekar R Ambati 2, Chandra Sekhar Amara 1, Karthik Reddy Kami Reddy 1, Meredith Lauren Spradlin 3, Amrit Koirala 1,4,11, Sachin B Jorvekar 1, Sandra L Grimm 1,2,4,11, Dexue Fu 5, Krishna Parsawar 6, Felice de Jong 7, Chris Beecher 7, Subrata Sen 8, Seth P Lerner 9, M Minhaj Siddiqui 5, Yair Lotan 10, Livia S Eberlin 3, Arun Sreekumar 1, Cristian Coarfa 1,2,4,11,, Nagireddy Putluri 1,2,
PMCID: PMC13040883  PMID: 41923172

Abstract

Background

Untargeted metabolomics has emerged as a powerful approach to uncover metabolic dysregulation associated with cancer progression. When integrated with a machine learning strategy it facilitates the discovery of key metabolic pathways and predictive biomarkers with high diagnostic and prognostic value.

Methods

In this study, we employed liquid chromatography coupled to high-resolution Tribrid Orbitrap mass spectrometry to perform comprehensive metabolic profiling of bladder cancer (BLCA) as well as predict invasiveness of the disease.

Results

By leveraging both in-house retention time-based MS/MS spectral libraries and commercial databases, we robustly identify over 2000 metabolites. In addition, this platform allows identification of novel pathways highlighting metabolic vulnerabilities in BLCA. The application of machine learning algorithms and advanced computational modeling uncovered metabolic signatures that differentiate BLCA from adjacent normal/benign samples and distinguish muscle-invasive from non-muscle-invasive bladder cancer. Our integrative analytical pipeline addresses key challenges in metabolomics-including high dimensionality, metabolite annotation, and biological variability-through feature selection and predictive modeling. We identify candidate metabolic markers with strong potential for early detection and characterize invasiveness of the disease and identify potential therapeutic target pathways.

Conclusions

This work highlights the power of combining untargeted metabolomics with machine learning to map the metabolic landscape of BLCA and to accelerate the development of precision diagnostics and future therapeutic strategies.

Supplementary Information

The online version contains supplementary material available at 10.1186/s40170-026-00427-4.

Keywords: Untargeted metabolomics, Orbitrap IQ-X, Bladder cancer, Machine learning

Introduction

Metabolomics has emerged as a powerful tool for elucidating metabolic alterations underlying disease processes, offering unique insights into pathophysiology and biomarker discovery. Unlike genomics or proteomics, metabolomics captures real-time dynamic snapshots of cellular activities, reflecting the complex interplay of genetic, environmental, and lifestyle factors. By profiling small-molecule metabolites in biological systems, this approach enables the identification of dysregulated pathways that contribute to disease progression and may serve as novel therapeutic targets [1].

Among metabolomics strategies, untargeted metabolomics has gained importance because it enables comprehensive detection of a wide range of metabolites without prior assumptions. This capacity is especially valuable in cancer research, where it facilitates the discovery of cancer-specific metabolic signatures and novel metabolic markers [2]. The integration of high-throughput liquid chromatography-mass spectrometry (LC-MS) has significantly enhanced untargeted metabolomics, enabling simultaneous detection of a broad spectrum of metabolites in biological matrices [3].

Despite its advantages, untargeted LC-MS metabolomics presents several challenges, including complexities in data interpretation, variability in sample preparation, and difficulties in metabolite identification. The extensive datasets generated necessitate sophisticated computational tools for robust analysis. Additionally, inconsistencies in sample processing can introduce biases, affecting reproducibility. Metabolite identification remains a significant challenge due to the structural diversity of metabolites and the limited availability of comprehensive spectral libraries. Furthermore, LC-MS-based methodologies are subject to limitations such as ion suppression effects and restricted dynamic range, which may impact the accuracy of quantification. Additionally, the structural diversity of metabolites and the limited scope of spectral libraries hinder the comprehensive identification of metabolites in biological matrices [46].

Bladder cancer (BLCA) is a common and clinically heterogeneous malignancy, representing a significant global health burden [7, 8]. While progress has been made in diagnosis and therapeutic approaches, outcomes remain variable, with five year survival of 71% for localized disease, 39% for regional, and 8% for distant disease [9, 10]; particularly due to a lack of molecular stratification tools. A deeper understanding of metabolic reprogramming that accompanies BLCA development and progression could inform effective diagnostic, prognostics, and therapeutic implications [11, 12].

Although previous studies have uncovered metabolic alterations in bladder cancers, many have relied on targeted metabolomic approaches or lacked the resolution and depth necessary to characterize the metabolic landscape [1118]. There remains a critical need for comprehensive, untargeted studies that can capture the dynamic and diverse metabolic aspects of BLCA metabolism and progression.

In this study, we applied an advanced LC coupled to high-resolution MS to identify global metabolomics alterations in BLCA and to characterize the aggressiveness of the disease. By combining cutting-edge analytical instrumentation with robust machine learning (ML) approaches, we aim to identify metabolic signatures that enhance our understanding of BLCA biology and support the advancement of precision medicine, improving diagnosis, risk stratification, and clinical outcomes of BLCA.

Materials and methods

Clinical samples

For this study, human bladder cancer (BLCA) and adjacent normal/benign tissues were obtained in a de-identified manner from the Cooperative Human Tissue Network (CHTN), the University of Maryland Baltimore (UMB), and the University of Texas Southwestern Medical Center (UTSW). All samples were collected in de-identified manner under approved IRB protocols and stored at −140 °C until further analysis. The clinical information used for the study was provided in Supplementary Table 1.

Sample preparation and metabolite extraction

Approximately 10 mg of tissue was homogenized in 300 µL of a 1:1 (v/v) methanol-water solution on ice (4°C). The lysate was transferred to a 2 mL Eppendorf tube, and 900 µL of methanol-acetonitrile (1:1, v/v) was added (3-fold volume relative to the lysate). The mixture was vortexed for 5 minutes, then incubated at −20 °C for 20 minutes to precipitate proteins and other insoluble components. After incubation, the samples were centrifuged at 15,000 rpm (4 °C, 10 min). The supernatant was transferred into 2 mL tubes and dried under a speed vacuum (GeneVac, SP Scientific). The dried samples were reconstituted in 100 µL of methanol-water (1:1, v/v), followed by vortexing (5 min) and water bath sonication (5 min) to ensure complete solubilization. The samples were then centrifuged again (15,000 rpm, 4 °C, 5 min) to remove any particulates. For quality control (QC), 10 µL aliquots from each sample were pooled and used for this study.

Liquid chromatography-mass spectrometry

The study was conducted utilizing the Thermo Scientific Vanquish UHPLC system (Thermo Fisher Scientific), which was equipped with a Vanquish Horizon Binary Pump H, a Vanquish Column Compartment H, and a Vanquish Split sampler HT. Metabolites were separated using Hydrophilic Interaction Liquid Chromatography (HILIC) with an ACQUITY UPLC BEH Amide Column (130 Å, 1.7 µm, 2.1 mm X 150 mm) from Waters Corporation, as well as Reverse Phase (RP) chromatography utilizing an ACQUITY UPLC HSS T3 Column (100 Å, 1.8 µm, 2.1 mm X 150 mm), both in electro spray ionization (ESI) positive and negative modes [19]. The Thermo Scientific Orbitrap IQ-X Tribrid mass spectrometer from Thermo Fisher Scientific is operated using IQ-X Tune instrumental control software. Real-time data is collected through Thermo Xcalibur Software, utilizing ultra-high purity nitrogen and helium gas. This instrument offers high-resolution accurate mass (HRAM) analysis, coupled with quadrupole isolation and fragmentation of precursor ions achieved through high-energy collision-energy dissociation (HCD) through Orbitrap. This fragmentation process generates spectra containing fragment ions in both ESI positive and negative modes for comprehensive detection. Data acquisition occurs in two modes: full scan MS mode and MS2 mode via data-dependent acquisition (DDA).

Quality controls

The injection sequence in LC-MS is carefully planned to ensure system readiness, assess quality, monitor control parameters, and normalize analytes to the reference matrix. A detailed sequence table is prepared, outlining sample details, acquisition method, and sample locations for the LC autosampler. Before initiating the analysis of biological samples, a series of blank samples are run to condition the system and remove any residual matrix from the column and LC system, minimizing the risk of analyte carryover. Following this steps, pool quality control (QC) samples are injected to optimize column performance in preparation for analyzing biological samples. Pool QC samples are strategically injected at the beginning and end of the sample batch and at intervals of the samples throughout the acquisition process. This ensures consistent monitoring of system performance and quality across the entire batch to prevent any potential bias caused by temporal drift in instrument performance. Biological samples are randomized before analysis, ensuring accurate comparison of metabolite abundances during subsequent data analysis.

Generation of the In-house retention time-based MS/MS spectral library

An in-house spectral library was built using reference standards of 630 metabolite compounds from IROA Technologies. The compounds were prepared according to the manufacturer’s protocol and analyzed using a Thermo IQ-X Orbitrap Tribrid mass spectrometer coupled to Vanquish Horizon LC. Data acquisition was performed in both ESI positive and negative modes using four LC methods: HILIC-Positive, HILIC-Negative, RP-Positive, and RP-Negative. The LC system was coupled to the mass spectrometer with an optimized gradient and mobile phase composition for each method. Mass spectra were acquired in DDA mode with high-resolution full MS and MS/MS scans. The acquired spectra were processed using MLS Discovery software.

After review, based on metabolite peaks detected using MLS Discovery software, 518 unique metabolites out of 630 were qualified for inclusion in the MS/MS spectral library. Compound identification was confirmed, and an in-house curated spectral database was created as NIST file (.msp), then converted to Thermo Compound Discoverer (CD) format (.db) for use as a plug-in the library for database search. Identified spectra were annotated with retention time, accurate mass, and MS/MS fragmentation patterns (Supplementary Figure 1). Additionally, adducts were also considered when generating the spectral database.

Metabolite identification and statistical analysis

Compound Discoverer (CD), version 3.3, serves as an effective tool for qualitative and semi-quantitative analysis in untargeted metabolomics. It leverages accurate mass data, isotope pattern matching, fragment matching, and mass spectral library searches to identify the structural composition of small molecules, including chromatogram alignment. CD boasts a user-friendly and customizable node-based processing workflow tailored to handle Xcalibur RAW files. This workflow seamlessly processes raw data into single result files while facilitating statistical analysis, thus streamlining the metabolomics data processing pipeline. Processing raw files through CD involves several sequential steps. Initially, spectra were extracted from the raw data, followed by retention time (RT) alignment with a tolerance of 0.2 minutes, peak detection tolerance of 1.5 minutes and a mass tolerance of 5 ppm. In a subsequent stage, missing values were imputed using CD automatic workflow with a Fill Gaps node for further data processing. For metabolite identification, CD utilizes m/z information from multiple MS/MS Spectral databases; therefore, naming of the metabolites may vary between databases. This applies to the in-house reference library as well as public spectral libraries such as mzCloud and the NIST Mass Spectral Library. These processes utilized QC samples and compound identification to ensure analytical consistency and accuracy. The identified compounds were annotated utilizing various online and offline databases, including an in-house RT-based library, online mzCloud, NIST 2020, and mass list from Human Metabolome Database (HMDB). To enhance identification accuracy, the mzLogic algorithm utilized all available fragmentation scans (full MSn depth) for an compound [20], scoring potential matches. In cases where mzCloud search results did not yield any matches for an unknown compound, the mzLogic algorithm provided ranking scores based on identification. Adducts were taken into account during the identification process. Additionally, the algorithm provided ranking scores for different database search results when an unknown compound had available data-dependent MS2 scans and similarity results from spectral libraries search.

Metabolites were matched against the in-house RT-based MS/MS spectral library, with higher priority given to multiple matches over single matches. Metabolites were categorized as “A” when they matched not only the “retention time and MS/MS spectra in our in-house IROA-based reference library” but were also spectral matched with either the “mzCloud MS/MS spectral libraries” or “NIST MS/MS” spectral libraries. Features were categorized as “B” when they matched either the “mzCloud MS/MS spectral libraries” or “NIST MS/MS” spectral libraries. The features were categorized as a “C” when they matched only the accurate mass from a curated mass list (HMDB database). The names of metabolites are not assigned exclusively from our in-house RT-based spectral library. During annotation, CD prioritizes matches from integrated databases such as mzCloud and ChemSpider when assigning metabolite names. As a result, even when a metabolite is confirmed using our in-house retention time and MS/MS reference library, the reported name may reflect the preferred match selected by the software from mzCloud or ChemSpider rather than the alternative name (e. g., isomeric name/synonym name/common name/IUPAC name) used in our in-house library for category “A”. Therefore, the metabolite names in our in-house RT-based spectral library may differ from those reported by mzCloud and ChemSpider, and the final reported name typically follows the mzCloud or ChemSpider database entry.

In cases where metabolites matched with the same score, the lowest coefficient of variation (CV) of pool QC was used to determine the best match. High mass accuracy was maintained as the standard. After filtering the data, the drug metabolites were further filtered using the DrugBank Database 2023 [21] (DRUGBANK; https://go.drugbank.com/) to exclude the known drugs, since the current study focused on solely metabolomics.

The peak areas from four different methods including RP-Positive, RP-Negative, HILIC-Positive, and HILIC-Negative, were normalized using a method-wise median interquartile range (IQR) normalization method as previously described [19]. Altered metabolites were identified using a t-test followed by the Benjamini-Hochberg False Discovery Rate (FDR) test [22], considering an FDR of less than 0.25 as indicative of differentially altered metabolites [12, 16, 19, 2327]. For the analysis, BLCA pathological stages Ta and T1 were classified as non–muscle invasive bladder cancer (NMIBC), whereas stage T2 and higher (>T2) were classified as muscle invasive bladder cancer (MIBC). After filtering the data, we used an R package-based comprehensive pipeline (https://github.com/CoarfaBCM/runModac) to help normalize the data and perform the required statistical analysis, including ensuring data quality.

Pathway enrichment analysis

The metabolites with significant alterations between adjacent normal/benign and BLCA as well as MIBC and NMIBC patient tissues were selected for the metabolic pathway analysis. Differentially expressed metabolites were mapped to associated genes using the HMDB database [12, 16, 28]. These identified genes were then used for pathway enrichment analysis through over-representation analysis [2931] (ORA, refer to supplementary table 29) using the KEGG [32], Reactome [33], Hallmark [34], and GO [35] databases.

Machine learning models

For supervised machine learning analyses, normalized untargeted metabolomic profiles were used. Machine learning (ML) models were derived using k-nearest neighbor (KNN), Random Forest (RF), and Support Vector Machines (SVM) [31, 36, 37]. A cross-validation approach was used, with data split into 80% training for model building, then 20% for model testing; the R package caret was used to implement the cross-validation approach [38]. Performance of classification was quantified using the Area Under the Receiver Operating Characteristic curve (AUROC). The training/testing split was conducted over 100 cross-validation iterations, and the median value and distribution of AUROC were collected. The importance of individual features (metabolites) was determined using the Interpretable Machine Learning R package [38, 39], and features were further filtered for those informative in least 70% of iterations. For the best performing model for each machine learning problem, the top 20 informative features were reported, together with their direction of association with the outcome, positive or negative. Lastly, models were derived using top 2–20, 25, 30, 40, 50, 75, and 100 most informative features to determine a minimum complement of features sufficient to classify the outcomes.

Results

We developed a robust untargeted metabolic profiling platform using high-throughput liquid chromatography-mass spectrometry (LC-MS) with the Thermo Scientific IQ-X Orbitrap. This platform was meticulously designed to capture a comprehensive metabolic fingerprint of biological samples by incorporating four distinct separation and detection methods: reverse-phase (RP) in both positive and negative ion modes (RP-positive, RP-negative) and hydrophilic interaction liquid chromatography (HILIC) in both positive and negative ion modes (HILIC-positive, HILIC-negative). We developed an in-house retention time (RT)-based spectral library using standard compounds to characterize metabolites via LC-MS method. This in-house spectral library was integrated with external databases such as mzCloud and NIST to enhance metabolite identification. This multi-modal approach enabled a broad coverage of metabolites, encompassing both polar and non-polar compounds, thereby enhancing the detection sensitivity and overall analytical depth (Fig. 1).

Fig. 1.

Fig. 1

Schematic diagram represents the entire workflow used for the study

Following data acquisition, the mass spectra were processed using the Compound Discoverer (version: 3.3) software, which facilitated the identification of metabolites by searching against several mass spectral libraries. These libraries included an online mzCloud, NIST 2020, an in-house RT-based library, and a mass list from Human Metabolome Database (HMDB). The comprehensive use of these libraries ensured highly accurate measurement in metabolite identification, allowing for extensive coverage of the metabolome. A rigorous data filtering and quality assurance process was implemented to ensure the accuracy and reliability of the identified metabolites. The data processing workflow included filtering and MS/MS-based identification (Category A-B), emphasizing library matches and mass accuracy. Metabolites that did not identified a definitive match in the spectral libraries were further subjected to metabolites annotation from Category C (Mass List). When multiple library matches had the same score, the metabolite with the lowest CV among the QC samples was selected as the best match. The above filtering criteria allowed for the inclusion of metabolites in all levels that, although not fully characterized by spectral libraries, were likely to be true positives due to their consistent detection across replicates. As shown in Supplementary Figure 2, the QC samples shows robust reproducibility, with correlation coefficients > 75% for peak areas (Supplementary Figure 2A) and their corresponding histogram (Supplementary Figure 2C). After normalization, the QC data show even higher correlations (>94%) with consistent signal distribution and stability across injections (Supplementary Figure 2B and 2D).

Assessment of the untargeted metabolomics platform

LC-MS analysis was performed to profile the global metabolome, including adjacent normal/benign tissues, and BLCA patient tissues (Supplementary Table 1). We employed a liquid-liquid extraction method using acetonitrile and methanol to extract metabolites from tissue samples of both normal and BLCA patients. This procedure enabled the extraction of both polar and non-polar metabolites. The extracted samples were then analyzed using high-throughput LC-MS. Data analysis was performed using CD software (Thermo Fisher Scientific), utilizing multiple spectral libraries. Additionally, data acquisition was carried out using four developed LC-MS methods, and the mass spectra were searched against various spectral libraries, including in-house library and mass lists. Subsequently, the data were filtered using established criteria including different categories [4042], and statistical analysis was conducted to identify the metabolites that were altered between the normal and BLCA patient tissue samples (Fig. 1 and Supplementary Figure 1). In each of the four analytical methods, we ran eight pooled quality control (QC) samples, distributed across the sample acquisition. Subsequent analysis revealed excellent correlations among the QC samples across all methods, as demonstrated by correlation plots (Supplementary Figure 2). These strong correlations indicate the high reproducibility and consistency of our metabolomic analyses, ensuring the reliability of the data obtained for the subsequent comparative studies between benign and cancerous tissue samples. After filtering the metabolites, we identified 2590 metabolites/compounds across all analytical methods (Supplementary Figure 3 and 4).

Altered metabolites and their associated pathways

When comparing between adjacent normal/benign (n = 29) and BLCA (n = 65) samples, a total of 1475 altered metabolites (FDR < 0.25) were identified across all analytical methods. These metabolites were detected across 4 methods and categorized as “A-C” based on the nature of the identifications (Fig. 2A and Supplementary Data 1). Out of 1475 metabolites, 812 metabolites were decreased and 663 were increased in BLCA tumor samples compared to adjacent benign/normal samples (Fig. 2B).

Fig. 2.

Fig. 2

Identifying altered metabolites and key pathways in bladder cancer (BLCA) compared to adjacent normal/benign. A) Heatmap showing differentially expressed metabolites (DEM) in BLCA (n = 65) compared to adjacent normal/benign (n = 29) tissues (FDR <0.25). The color scale (z-score) represents relative abundance of metabolites. B) Volcano plot illustrating the significance and fold changes of metabolites in BLCA (n = 65) compared to adjacent normal/benign (n = 29) tissue samples. C) Dot plot showing the top 10 significantly enriched hallmark pathways (refer to supplementary table 2) obtained using DEMs between BLCA over adjacent normal/benign. DEMs (FDR < 0.25) were mapped to genes and used for pathway analysis. The number of genes from pathways and their significance is represented

Next, we mapped the differentially expressed metabolites (DEMs; 1475) to genes using HMDB, then we performed over-representation analysis against the pathway collections Hallmark, KEGG, Reactome, and Gene Ontology (GO) to understand the potential functional impact of DEMs in BLCA. In Hallmark pathways, fatty acid and xenobiotic metabolism showed the most significant enrichment, indicating major lipid metabolism alterations in BLCA. Other notable pathways include glycolysis and oxidative phosphorylation, reflecting significant metabolic reprogramming. Adipogenesis and bile acid metabolism were also enriched (Fig. 2C and Supplementary Table 2). A top enriched KEGG pathway (Supplementary Figure 5A and Supplementary Table 3) was glycolysis and gluconeogenesis, consistent with the Warburg effect; cytochrome P450-mediated xenobiotic metabolism and glycerophospholipid pathways were also prominently enriched. Analysis of Reactome pathways (Supplementary Figure 5B and Supplementary Table 4) highlighted significant changes in lipid metabolism, biological oxidations, and amino acid metabolism. Gene Ontology analysis (Supplementary Figure 5C and Supplementary Table 5) showed strong enrichment for organic acid metabolism and other critical processes like oxidation-reduction and lipid metabolism. These results emphasize the extensive metabolic alterations in BLCA, particularly in carbohydrate, lipid, and amino acid metabolism pathways. The identification of these enriched pathways provides valuable insights into the metabolic vulnerabilities of BLCA, which could be exploited for the development of targeted therapies.

In this study, BLCA pathological stage Ta-T1 tumors were considered as NMIBC, and T2–T4 tumors were considered as MIBC for the analysis. When comparing MIBC (n = 41) over NMIBC (n = 10) samples, a total of 16 altered metabolites (FDR < 0.25) were identified across all analytical methods. In the RP-Positive method, 5 altered metabolites were identified, whereas 5 in RP-Negative method, 2 in HILIC-Positive method, and 4 in HILIC-Negative method (Fig. 3A). Out of 16 metabolites, 3 metabolites were decreased whereas 13 metabolites were increased in MIBC samples compared to NMIBC samples (Fig. 3B). Next, we mapped the differential metabolites (DEMs; 16) to genes using HMDB, then we performed over-representation analysis using the Hallmark, KEGG, Reactome, and GO pathway compendia to understand the key pathways that are impacted in MIBC. Among the Hallmark pathways (Fig. 3C and Supplementary Table 6), xenobiotic metabolism exhibited significant enrichment. Other prominently enriched pathways include reactive oxygen species (ROS) signaling, glycolysis, and bile acid metabolism, highlighting key aspects of metabolic reprogramming. Additionally, apoptosis and hypoxia were also enriched, indicating broader alterations in lipid metabolism (Fig. 3C). KEGG pathway analysis (Supplementary Figure 6A and Supplementary Table 7) revealed enrichment of glutathione metabolism as well as other key pathways including cytochrome P450-mediated xenobiotic metabolism. Cytochrome P450 enzymes play a crucial role in xenobiotic metabolism by detoxifying carcinogens in BLCA as shown in our earlier publication [11]. However, their altered expression in tumors can lead to impaired detoxification, contributing to increased toxicity or drug resistance. Dysregulated P450 activity also affects hormone metabolism, influencing cancer progression and therapeutic outcomes [4345]. Reactome pathway analysis (Supplementary Figure 6B and Supplementary Table 8) highlighted significant enrichment of glutathione conjugation. GO analysis (Supplementary Figure 6C and Supplementary Table 9) revealed strong enrichment in xenobiotic metabolism and glutathione metabolism. The findings reveal significant metabolic changes in MIBC, particularly in pathways related to xenobiotic metabolism, lipid metabolism, glycan biosynthesis. The enrichment of glutathione metabolism across both KEGG and Reactome analyses emphasizes its crucial role in detoxification and maintaining redox homeostasis. These altered pathways provide a deeper understanding of the metabolic vulnerabilities in BLCA, offering potential targets for therapeutic intervention.

Fig. 3 .

Fig. 3

Metabolic and pathway alterations in MIBC compared to NMIBC. A) Heatmap showing differentially expressed, metabolites (DEM) in MIBC (n = 41) compared to NMIBC (n = 10) patients (FDR <0.25). The color scale (z-score) indicates relative metabolite abundance: upregulated metabolites are shown in yellow, and downregulated metabolites are shown in blue. B) Volcano plot illustrating the significance and fold changes (Log2) of metabolites in MIBC (n = 41) and NMIBC (n = 10). C) Dot plot showing the top 10 significantly enriched hallmark pathways (refer to supplementary table 6) obtained from DEMs between MIBC (n = 41) and NMIBC (n = 10). DEMs were mapped to genes and used for pathway analysis

Detection of biomarker candidates for BLCA and MIBC using machine learning

To complement the analysis of differentially expressed metabolites analysis, we explored the potential of machine learning applications to our untargeted metabolomics BLCA dataset. A challenging problem in the field is determining metabolic markers of MIBC vs NMIBC. A total of 2590 metabolites were used as input features, and the samples were categorized into MIBC (n = 41) and NMIBC (n = 10). We applied three machine learning classification models; all models demonstrated robust performance, achieving AUROC values of 0.750 for KNN, 0.766 for Random Forest (RF) and 0.813 for Support Vector Machine (SVM) in distinguishing MIBC from NMIBC (Fig. 4A). Based on its superior performance, the SVM model was selected to determine metabolic markers. Informative features were identified in at least 70% of the cross-validation iterations, and feature importance was computed using Interpretable Machine Learning (IML); as few as 8 metabolites sufficed to classify effectively MIBC vs NMIBC (Fig. 4B). The top 20 informative metabolites, ranked by feature importance, are shown in Fig. 4C, and their corresponding fold changes are displayed in Fig. 4D. Notably, 3-hydroxyglutaric acid, 3-methylglutaric acid, mesaconic acid, and 2-methylacetoacetic acid are decreased, and guanine is increased in patients with MIBC (Figs. 4C and 4D). Next, we investigated BLCA over adjacent normal/benign tissue samples; whereas this analysis is less challenging than the MIBC vs NMIBC, we wanted to determine the potential for machine learning to determine and rank informative metabolic markers for BLCA. A total of 2590 metabolites were used as input features, with samples classified as adjacent normal/benign (n = 29) or BLCA (n = 65). Comparing BLCA to adjacent normal/benign samples, all three models used demonstrated expected robust performance, each achieving an AUROC above 0.90-specifically, 0.938 for KNN, 0.938 for RF, and 0.923 for SVM (Supplementary Figure 7A). Based on its superior performance, the KNN model was selected to determine metabolic markers for BLCA. Informative features identified in at least 70% of the cross-validation iterations were ranked based on feature importance as few as 20 metabolites sufficed to effectively classify BLCA over adjacent normal/benign tissues (Supplementary Figure 7B). The top 20 metabolites, ranked by feature importance, are shown in Supplementary Figure 7C, along with their corresponding fold changes in Supplementary Figure 7D. Adipoylglycine, S-Adenosylmethionine, 2’-O-Methylguanosine, Glutaminylproline, trans-2-Hexacosenoic acid, N-propyl-L-arginine, N-Acetylmuramate, O-Succinyl-L-homoserine, 1-Methylguanosine, and 1,1-Dimethylurea were increased while 4-Thiouridine decreased in BLCA patients compared to adjacent normal/benign (Supplementary Figure 7C and 7D).

Fig. 4.

Fig. 4

Predicting metabolic markers in MIBC over NMIBC using parsimonious machine learning models. A) Three machine learning classification methods: k-nearest neighbor (KNN), Random Forest (RF), and Support Vector Machines (SVM) linear were used to predict MIBC (n = 41) over NMIBC (n = 10); One hundred cross-validation iterations were performed, using 80% of the data as training and 20% as testing. The median area under the receiver operating characteristic curve (AUROC) of each model is listed for each method. B) Minimal feature classifiers were identified by selecting metabolite features informative in ≥70% of 100 iterations. SVM classification was performed with an increasing number of features to predict MIBC over NMIBC. Each green dot represents the median AUROC for a given number of metabolites. C) The top 20 informative metabolites are plotted and sorted in decreasing order by important features in MIBC over NMIBC. D) Lollipop plot shows the fold changes of the top 20 metabolites in MIBC over NMIBC, and data derived from Figure 3B. Detailed statistical analyses, including significance values for these metabolites, are presented in Supplementary Data 2

Discussion

Metabolites are central regulators of cellular functions and play critical roles in the initiation and progression of a wide range of diseases, including cancers [46, 47]. The emergence of metabolomics technologies has enabled systematic capture of dynamic metabolic changes under pathological conditions, providing transformative insights into disease mechanisms. Over the past years, the metabolomics field has advanced rapidly, driven by innovations in analytical platforms, acquisition strategies and computational analysis methods [12, 1418].

Accurately characterizing cellular metabolism is particularly crucial in cancer metabolism, where metabolic reprogramming is now recognized as a hallmark of tumor progression and therapeutic resistance [4850]. Several analytical platforms are available to interrogate cancer metabolism including Seahorse [12], Nuclear Magnetic Resonance (NMR) [51], and MS-based approaches [12, 14, 16, 17]. While each of the available methods has its merits, they also have distinct limitations. NMR is non-destructive but suffers from relatively low sensitivity and limited dynamic range [52]. MS-based platforms utilizing Electro Spray Ionization (ESI), Matrix-Assisted Laser Desorption/Ionization (MALDI), Desorption Electrospray Ionization (DESI) sources have been widely applied to metabolomics [5356]. Gas Chromatography (GC)-MS is known for strong separation capacity for volatile compounds, but it often require labor intensive derivatization [57, 58].

In contrast, liquid chromatography coupled with a high-resolution Orbitrap-based instrument-offers high sensitivity, a broad dynamic range, and the ability to detect thousands of metabolites in a single analysis. In this study, we developed and applied a high-throughput untargeted metabolomics platform using high-resolution Orbitrap IQ-X mass spectrometry coupled with liquid chromatography to profile the metabolome in BLCA. By integrating both RP and HILIC in both positive and negative ionization modes, our platform captures the full spectrum of polar and non-polar metabolites, enabling comprehensive metabolic coverage. The Tribrid Orbitrap system used in our study integrates a quadrupole, linear ion trap, and Orbitrap analyzer, enabling ultra-high resolution, sub-ppm mass accuracy, and rapid scan speeds [59]. Advanced acquisition modes such as Data Dependent Acquision (DDA) enhance confidence of metabolites identification, reduce FDR, and enable precise measurement of metabolites.

Our previous studies have employed both targeted and untargeted LC-MS/TOF strategies to investigate the metabolic landscape of BLCA [11]. This effort includes analysis of tumor tissue and serum samples using LC-MS. This study reveals key dysregulated pathways including xenobiotic metabolism [11], amino acid metabolism [60], and fatty acid oxidation across BLCA stages as well as alterations in oxidative phosphorylation, amino acid metabolism and xenobiotic metabolism in BLCA [12, 16, 17, 61, 62]. Collectively, these findings highlight critical metabolic vulnerabilities that may be exploited therapeutically. A subset of the metabolites, either from in-house retention time (RT)-based spectral library or from the NIST spectral database, may represent structurally distinct isoforms. However, annotation of the metabolites, derived from mzCloud or ChemSpider, were mapped to corresponding genes, and subsequent pathway analyses were performed as part of the current study.

Despite these advances, the full scope of metabolic rewiring in BLCA remains incompletely characterized, owing to biological heterogeneity and technical limitations of existing platforms. The analytical depth of our LC–Tribrid MS platform overcomes many of these challenges, enabling comprehensive metabolic profiling of BLCA tumors relative to adjacent normal/benign tissues. In addition, our platform was able to identify distinct metabolic signatures in aggressive MIBC patients. This approach not only validates previous findings, including the well-documented elevation of xenobiotic metabolism in BLCA, but also provides a deeper understanding of the broader metabolic framework underlying the disease.

In parallel, the complexity and volume of data generated by the LC-MS platform demand advanced analytical approaches. Machine Learning (ML) has emerged as a powerful tool to extract meaningful patterns, enable disease prediction, and facilitate patient stratification [31, 63, 64]. ML with metabolomics has been applied to diseases such as cardiovascular [65, 66], gastric cancer [63], lung adenocarcinoma [67], and other diseases [6870]. In the context of BLCA, we leverage this synergy to identify novel metabolic markers and gain mechanistic insights into tumor biology.

In this study, we employed three ML models namely k-nearest neighbors (KNN), random forest (RF), and support vector machine (SVM). Among the various models, KNN slightly outperformed the others in adjacent normal/benign vs BLCA. Using a KNN machine learning approach, we identified several metabolic markers. Among the key metabolites altered in BLCA patients were adipoylglycine, 2′-O-methylguanosine, S-adenosylmethionine (SAM), and 4-thiouridine. The accumulation of methylated nucleosides such as 2’-O-methylguanosine and the elevated presence of S-adenosylmethionine suggest increased methylation activity, which is commonly associated with epigenetic dysregulation in tumor cells. Additionally, the increased in amino acid derivatives like adipoylglycine may reflect perturbations in nitrogen metabolism, further supporting the notion of reprogrammed metabolic pathways that contribute to cancer progression [28, 7174]. SAM, a key molecules in cellular metabolism, regulates crucial DNA and protein methylation for gene expression and epigenetics modulation [7577]. In BLCA, altered SAM metabolism contributes to unusual methylation and activation [78, 79]. These disruptions associate SAM to metabolic and epigenetic reprogramming that promotes tumor progression. Additionally, 2′-O-Methylguanosine is an endogenous RNA methylated nucleoside that is altered in BLCA in our study, which may reflect RNA methylation dysregulation including cell proliferation [80, 81].

Of note, among the various models tested, SVM exhibited slightly superior performance in distinguishing NMIBC from MIBC. Using SVM model, we also identified several metabolic markers metabolites associated with MIBC. An example, 3-Hydroxyglutaric acid, 3-Methylglutaric acid and 2-Methylacetoacetic acid were significantly decreased, while guanine was increased. The reduction of the organic acids may indicate distresses in branched-chain amino acid catabolism and related pathways, which are often linked to energy homeostasis and cellular stress responses [82]. Those changes could reflect a shift toward alternative metabolic routes that support tumor growth and survival. To the best of our knowledge, the suppression of 3-hydroxyglutaric acid, 3-methylglutaric acid, and 2-methylacetoacetic acid has not been explored in BLCA. We hypothesize that this reduction may be associated with cancer-related metabolic reprogramming, whereby tumor cells divert carbon flux away from oxidative catabolism toward anabolic biosynthesis and redox homeostasis to sustain proliferation and survival [83, 84]. Additionally, the pronounced increase in guanine intend to enhanced purine metabolism, consistent with the high nucleotide demand of rapidly proliferating cancer cells [8587]. Dysregulated purine biosynthesis has been implicated in tumor progression and may contribute to chemoresistance, highlighting its potential as a therapeutic target [85, 88, 89]. In contrast, the elevated guanine levels are consistent with accelerated purine salvage and heightened nucleotide turnover, metabolic programs known to sustain rapid DNA replication, support RNA biogenesis, and enable continuous proliferation in high-grade malignancies [86, 90]. Enhanced purine metabolism also contributes to redox buffering and epigenetic regulation through SAM and folate-linked one-carbon pathways, aligning with known oncogenic metabolic dependencies in BLCA [85, 88, 91]. Our study has a few limitations, the sample size in our study was relatively small which requires validation with a larger number of samples. Additionally, our platform also detected metabolites with high mass accuracy (Category C) which may require further validation in future studies.

Taken together, our LC-MS-based untargeted metabolomics platform offers exceptional analytical depth, resolution, and versatility for investigating cancer metabolism. When coupled with advanced computational analyses, such as machine learning approaches, it provides a powerful framework for discovering metabolic markers in BLCA. This integrative approach could greatly enhance our understanding of tumor metabolism and support early detection, prognosis, and future personalized therapy in BLCA.

Conclusion

This study presents the successful development and validation of a robust untargeted metabolomics platform leveraging the Thermo IQ-X Orbitrap Tribrid Mass Spectrometer coupled with the UHPLC system. We have applied these advanced platforms to BLCA tissue samples, this platform enabled high-resolution metabolic profiling and revealed distinct metabolic signatures between tumor and adjacent normal/benign tissues, as well as between patients with muscle-invasive bladder cancer (MIBC) and non-muscle-invasive bladder cancer (NMIBC). Our analysis identified novel dysregulated pathways-including xenobiotic metabolism, glutathione metabolism, amino acid metabolism, lipid metabolism, and glycolysis-highlighting potential metabolic vulnerabilities and offered critical insights into BLCA biology and therapeutic targeting. Furthermore, utilizing machine learning approaches holds promise for predicting key metabolic markers for both BLCA diagnosis and their stratification.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Data (1-8) (476.5KB, xlsx)
Supplementary Tables (1-9) (241.1KB, xlsx)

Abbreviations

BLCA

Bladder cancer

ML

Machine learning

LC

Liquid chromatography

MS

Mass spectrometry

HILIC

Hydrophilic interaction liquid chromatography

RP

Reverse phase

DDA

Data dependent acquisition

HCD

High energy collisional dissociation

MIBC

Muscle-invasive bladder cancer

NMIBC

Non-muscle-invasive bladder cancer

QC

Quality control

KNN

k-nearest neighbor

RF

Random forest

SVM

Support vector machine

CD

Compound discoverer

KEGG

Kyoto encyclopedia of genes and genomes

GO

Gene ontology

Author contributions

AHMK contributed development and validation of the methods; analyzed data, and takes full responsibility for the finished work and/or the conduct of the study, had access to the entire data set, and conducted manuscript writing and preparing the figures. VP contributed acquisition of the standard compounds, contributed to method developemnt and validation. TG developed the bioinformatics analysis pipeline, and conducted bioinformatics and machine learning data analysis. CRA contributed on LC methods and provided technical input. KRKR and CSA contributed to the collection of clinical samples, clinical data and edited the manuscript. AK contributed data analysis, and edited the manuscript. SBJ contributed to annotation of metabolites and edit the manuscript. SLG contributed data analysis, and edited the manuscript. KP contributed to development of the methods. FDJ and CB provided support and standard compounds for developing in-house library. SS and SPL edited the manuscript. MMS, YL, MLS, and LSE provided clinical samples and DF provided technical input on sample storage and shipment. CC provided project conception, supervised the development and application of bioinformatics and machine learning pipelines, provided intellectual input, contributed to manuscript writing, and controlled the decision to publish. AS provided project conception, and provided supervision. NP provided project conception, contributed to the development and validation of the methods, provided intellectual input, had access to the entire data set, provided supervision and manuscript writing, and controlled the decision to publish.

Funding

This research was supported by NIH S10OD032218 (ASK), NIH/NCI R01CA282282 (NP), DOD W81XWH-21-1-0613 (CA200996) (NP), and P20CA284971 (NP and ASK) and was partially funded by NIH grant nos. P42ES027725 (NP and CC), NIH/NCI R37CA289419 (NP), P30ES030285 (NP and CC), NIH R01HD112886 (NP), NIH R01GM141366 (NP), R01 NS137577 (NP), NIH U01CA214263 (NP, SS) and P30DK144025 (AHMK, VP and NP). This work received support from UTSW Simmons Comprehensive Cancer Center’s Tissue Management Shared Resource and was supported by the National Cancer Institute of the NIH under award no. P30CA142543. This work was supported in part by Cancer Prevention & Research Institute of Texas Proteomics & Metabolomics Core Facility Support Award CPRIT RP210227 (NP, CC and ASK) and by NCI Cancer Center Support Grant (P30CA125123) to the Metabolomics Core Shared Resource and Quantitative Science Shared Resource. TG, SLG, and CC were partially supported by CPRIT RP200504 (CC), NIH/NCI U54 CA274321 (NP and CC) and R21 DE032344 (CC). We thank the Cooperative Human Tissue Network (CHTN) for providing the bladder normal and cancer tissues, supported by University BioRepository & Precision Pathology Center (Duke BRPC) (P30CA014236), and the National Cancer Institute’s CHTN (RRID: SCR_004446), supported at Duke University by UM1CA239755. We thank to IROA Teachnologies for gifting the reference standard used for the study.

Data availability

Raw and processed metabolomics data from samples are available at NIH Metabolomics workbench (Study ID: ST003990). All the raw data values used for the study are in supplementary data (1–8).

Declarations

Ethics approval and consent to participate

All human tissue samples were collected from tumor bank from UTSW, UMB and de-identified samples were analysed under protocols approved by the Institutional Review Board: protocol H-35808 (BCM) from UTSW, UMB. De-identified samples from CHTN were obtained by Dr. Eberlin (under MTA) and trasnsfered to Dr. Putluri lab (under MTA) for metabolomics analysis.

Consent for publication

All the authors have read and approved the submission of the current version of the manuscript.

Competing interests

AS is Scientific Advisor to Karkinos Health Care Pvt Ltd., India, and is an unpaid visiting faculty to Sri Sathya Sai Institute for Higher Learning, India.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Abu Hena Mostafa Kamal, Vasanta Putluri and Tanmay Gandhi contributed equally to this work.

Contributor Information

Cristian Coarfa, Email: coarfa@bcm.edu.

Nagireddy Putluri, Email: putluri@bcm.edu.

References

  • 1.Wishart DS. Emerging applications of metabolomics in drug discovery and precision medicine. Nat Rev Drug Discov. 2016;15(7):473–84. [DOI] [PubMed] [Google Scholar]
  • 2.Patti GJ, Yanes O, Siuzdak G. Innovation: Metabolomics: the apogee of the omics trilogy. Nat Rev Mol Cell Biol. 2012;13(4):263–69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Beger RD, Dunn W, Schmidt MA, Gross SS, Kirwan JA, Cascante M, et al. Metabolomics enables precision medicine: “A white paper, community Perspective”. Metabolomics. 2016;12(10):149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Cui L, Lu H, Lee YH. Challenges and emergent solutions for LC-MS/MS based untargeted metabolomics in diseases. Mass Spectrom Rev. 2018;37(6):772–92. [DOI] [PubMed] [Google Scholar]
  • 5.Benton HP, Ivanisevic J, Mahieu NG, Kurczy ME, Johnson CH, Franco L, et al. Autonomous metabolomics for rapid metabolite identification in global profiling. Anal Chem. 2015;87(2):884–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Calderon-Santiago M, Priego-Capote F, Luque de Castro MD. Enhanced detection and identification in metabolomics by use of LC-MS/MS untargeted analysis in combination with gas-phase fractionation. Anal Chem. 2014;86(15):7558–65. [DOI] [PubMed] [Google Scholar]
  • 7.Bray F, Ferlay J, Soerjomataram I, Siegel RL, Torre LA, Jemal A. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2018;68(6):394–424. [DOI] [PubMed] [Google Scholar]
  • 8.Amara CS, Kami Reddy KR, Yuntao Y, Chan YS, Piyarathna DWB, Dobrolecki LE, et al. The IL6/JAK/STAT3 signaling axis is a therapeutic vulnerability in SMARCB1-deficient bladder cancer. Nat Commun. 2024;15(1):1373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Siegel RL, Kratzer TB, Giaquinto AN, Sung H, Jemal A. Cancer statistics, 2025. CA Cancer J Clin. 2025;75(1):10–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Survival Rates for Bladder Cancer [https://www.cancer.org/cancer/types/bladder-cancer/detection-diagnosis-staging/survival-rates.html].
  • 11.Putluri N, Shojaie A, Vasu VT, Vareed SK, Nalluri S, Putluri V, et al. Metabolomic profiling reveals potential markers and bioprocesses altered in bladder cancer progression. Cancer Res. 2011;71(24):7376–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kami Reddy KR, Piyarathna DWB, Park JH, Putluri V, Amara CS, Kamal AHM, et al. Mitochondrial reprogramming by activating OXPHOS via glutamine metabolism in African American patients with bladder cancer. JCI Insight. 2024;9(17). [DOI] [PMC free article] [PubMed]
  • 13.Kamat AM, Hahn NM, Efstathiou JA, Lerner SP, Malmstrom PU, Choi W, et al. Bladder cancer. Lancet. 2016;388(10061):2796–810. [DOI] [PubMed] [Google Scholar]
  • 14.Sreekumar A, Poisson LM, Rajendiran TM, Khan AP, Cao Q, Yu J, et al. Metabolomic profiles delineate potential role for sarcosine in prostate cancer progression. Nature. 2009;457(7231):910–14. [DOI] [PMC free article] [PubMed] [Google Scholar] [Research Misconduct Found]
  • 15.Vantaku V, Putluri V, Bader DA, Maity S, Ma J, Arnold JM, et al. Epigenetic loss of AOX1 expression via EZH2 leads to metabolic deregulations and promotes bladder cancer progression. Oncogene. 2020;39(40):6265–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Vantaku V, Dong J, Ambati CR, Perera D, Donepudi SR, Amara CS, et al. multi-omics integration anal robustly predicts high-grade patient survival identifies CPT1B Effect Fat Acid Metab Bladder Cancer. Clin Cancer Res. 2019;25(12):3689–701. [DOI] [PMC free article] [PubMed]
  • 17.Vantaku V, Donepudi SR, Piyarathna DWB, Amara CS, Ambati CR, Tang W, et al. Large-scale profiling of serum metabolites in African American and European American patients with bladder cancer reveals metabolic pathways associated with patient survival. Cancer. 2019;125(6):921–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Amara CS, Vantaku V, Lotan Y, Putluri N. Recent advances in the metabolomic study of bladder cancer. Expert Rev Proteomics. 2019;16(4):315–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Li Y, Yadollahi P, Essien FN, Putluri V, Ambati CSR, Kami Reddy KR, et al. Tobacco smoke exposure is a driver of altered oxidative stress response and immunity in head and neck cancer. J Transl Med. 2025;23(1):403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Cooper B, Yang R. An assessment of AcquireX and Compound Discoverer software 3.3 for non-targeted metabolomics. Sci Rep. 2024;14(1):4841. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wishart DS, Feunang YD, Guo AC, Lo EJ, Marcu A, Grant JR, et al. DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic Acids Res. 2018;46(D1):D1074–82. [DOI] [PMC free article] [PubMed]
  • 22.Benjamini Y, Hochberg Y. Controlling the false discovery Rate: a Practical and powerful approach to multiple testing. J R Stat Soc: Ser B (Methodological). 1995;57(1):289–300. [Google Scholar]
  • 23.Thaiparambil J, Dong J, Grimm SL, Perera D, Ambati CSR, Putluri V, et al. Integrative metabolomics and transcriptomics analysis reveals novel therapeutic vulnerabilities in lung cancer. Cancer Med. 2023;12(1):584–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Murthy D, Dutta D, Attri KS, Samanta T, Yang S, Jung KH, et al. CD24 negativity reprograms mitochondrial metabolism to PPARalpha and NF-kappaB-driven fatty acid beta-oxidation in triple-negative breast cancer. Cancer Lett. 2024;587:216724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kaushik AK, Vareed SK, Basu S, Putluri V, Putluri N, Panzitt K, et al. Metabolomic profiling identifies biochemical pathways associated with castration-resistant prostate cancer. J Proteome Res. 2014;13(2):1088–100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kannan V, Srimadh Bhagavatham SK, Dandamudi RB, Kunchala H, Challa S, Almansour AI, et al. Integrated clinical and metabolomic analysis identifies molecular signatures, biomarkers, and therapeutic targets in primary angle closure glaucoma. Front Mol Biosci. 2024;11:1421030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Braun T, Feng R, Amir A, Levhar N, Shacham H, Mao R, et al. Diet-omics in the study of Urban and rural crohn disease evolution (SOURCE) cohort. Nat Commun. 2024;15(1):3764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Wishart DS, Guo A, Oler E, Wang F, Anjum A, Peters H, et al. HMDB 5.0: the Human metabolome database for 2022. Nucleic Acids Res. 2022;50(D1):D622–31. [DOI] [PMC free article] [PubMed]
  • 29.Wieder C, Frainay C, Poupin N, Rodriguez-Mier P, Vinson F, Cooke J, et al. Pathway analysis in metabolomics: recommendations for the use of over-representation analysis. PLoS Comput Biol. 2021;17(9):e1009105. [DOI] [PMC free article] [PubMed]
  • 30.Lu Y, Pang Z, Xia J. Comprehensive investigation of pathway enrichment methods for functional interpretation of LC-MS global metabolomics data. Brief Bioinform. 2023;24(1). [DOI] [PMC free article] [PubMed]
  • 31.Powell H, Coarfa C, Ruiz-Echartea E, Grimm SL, Najjar O, Yu B, et al. Differences in prediagnostic serum metabolomic and lipidomic profiles between cirrhosis patients with and without incident hepatocellular carcinoma. J Hepatocell Carcinoma. 2024;11:1699–712. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kanehisa M, Furumichi M, Tanabe M, Sato Y, Morishima K. KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res. 2017;45(D1):D353–61. [DOI] [PMC free article] [PubMed]
  • 33.Croft D, O’Kelly G, Wu G, Haw R, Gillespie M, Matthews L, et al. Reactome: a database of reactions, pathways and biological processes. Nucleic Acids Res. 2011;39(Database issue):D691–697. [DOI] [PMC free article] [PubMed]
  • 34.Liberzon A, Birger C, Thorvaldsdottir H, Ghandi M, Mesirov JP, Tamayo P. The molecular signatures database (MSigDB) hallmark gene set collection. Cell Syst. 2015;1(6):417–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Harris MA, Clark J, Ireland A, Lomax J, Ashburner M, Foulger R, et al. The Gene Ontology (GO) database and informatics resource. Nucleic Acids Res. 2004;32(Database issue):D258–261. [DOI] [PMC free article] [PubMed]
  • 36.Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. [Google Scholar]
  • 37.Cortes C, Vapnik V. Support-vector networks. Mach Learn. 1995;20(3):273–97. [Google Scholar]
  • 38.Kuhn M. Building predictive models in R using the caret package. J Stat Softw. 2008;28(5):1–26.27774042 [Google Scholar]
  • 39.Molnar C, Casalicchio G. Bischl B: iml: an R package for Interpretable machine learning. J Educ Chang Open Source Softw. 2018;3(26).
  • 40.Zhou Z, Luo M, Zhang H, Yin Y, Cai Y, Zhu ZJ. Metabolite annotation from knowns to unknowns through knowledge-guided multi-layer metabolic networking. Nat Commun. 2022;13(1):6656. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Reisdorph NA, Walmsley S, Reisdorph R. A Perspective and framework for developing sample type specific databases for LC/MS-Based clinical metabolomics. Metabolites. 2019;10(1). [DOI] [PMC free article] [PubMed]
  • 42.Chaleckis R, Meister I, Zhang P, Wheelock CE. Challenges, progress and promises of metabolite annotation for LC-MS-based metabolomics. Curr Opin Biotechnol. 2019;55:44–50. [DOI] [PubMed] [Google Scholar]
  • 43.Rodriguez-Antona C. Ingelman-sundberg M: cytochrome P450 pharmacogenetics and cancer. Oncogene. 2006;25(11):1679–91. [DOI] [PubMed] [Google Scholar]
  • 44.Nebert DW, Dalton TP. The role of cytochrome P450 enzymes in endogenous signalling pathways and environmental carcinogenesis. Nat Rev Cancer. 2006;6(12):947–60. [DOI] [PubMed] [Google Scholar]
  • 45.Guengerich FP. Cytochrome P450 enzymes in the generation of commercial products. Nat Rev Drug Discov. 2002;1(5):359–66. [DOI] [PubMed] [Google Scholar]
  • 46.Mrowiec K, Debik J, Jelonek K, Kurczyk A, Ponge L, Wilk A, et al. Profiling of serum metabolome of breast cancer: multi-cancer features discriminate between healthy women and patients with breast cancer. Front Oncol. 2024;14:1377373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Kosmides AK, Kamisoglu K, Calvano SE, Corbett SA, Androulakis IP. Metabolomic fingerprinting: challenges and opportunities. Crit Rev Biomed Eng. 2013;41(3):205–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Xiao Z, Dai Z, Locasale JW. Metabolic landscape of the tumor microenvironment at single cell resolution. Nat Commun. 2019;10(1):3763. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Schiliro C, Firestein BL. Mechanisms of metabolic reprogramming in cancer Cells supporting Enhanced growth and proliferation. Cells. 2021;10(5). [DOI] [PMC free article] [PubMed]
  • 50.Phan LM, Yeung SC, Lee MH. Cancer metabolic reprogramming: importance, main features, and potentials for precise targeted anti-cancer therapies. Cancer Biol Med. 2014;11(1):1–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Li T, Deng P. Nuclear magnetic resonance technique in tumor metabolism. Genes Dis. 2017;4(1):28–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Nagana Gowda GA, Raftery D. Can NMR solve some significant challenges in metabolomics? J Magn Reson. 2015;260:144–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.King ME, Lin M, Spradlin M, Eberlin LS. Advances and Emerging Medical applications of direct mass spectrometry technologies for tissue analysis. Annu Rev Anal Chem (Palo Alto Calif). 2023;16(1):1–25. [DOI] [PubMed] [Google Scholar]
  • 54.Godfrey TM, Shanneik Y, Zhang W, Tran T, Verbeeck N, Patterson NH, et al. Integrating ambient ionization mass spectrometry imaging and spatial transcriptomics on the same cancer tissues to identify RNA-Metabolite correlations. Angew Chem Int Ed Engl. 2025;e202502028. [DOI] [PubMed]
  • 55.Gatmaitan AN, Weaver K, Badal SP, Bird HE, Huhn A, Hussey C, et al. Development and application of the modular MasSpec Pen System for clinical opioid screening. Anal Chem 2025. [DOI] [PubMed]
  • 56.Garza KY, King ME, Nagi C, DeHoog RJ, Zhang J, Sans M, et al. Intraoperative Evaluation of breast tissues during breast cancer operations using the MasSpec Pen. JAMA Netw Open. 2024;7(3):e242684. [DOI] [PMC free article] [PubMed]
  • 57.Koek MM, Jellema RH, van der Greef J, Tas AC, Hankemeier T. Quantitative metabolomics based on gas chromatography mass spectrometry: status and perspectives. Metabolomics. 2011;7(3):307–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Fiehn O. Metabolomics by gas chromatography-mass spectrometry: combined targeted and untargeted profiling. Curr Protoc Mol Biol. 2016;114: 30 34 31–30 34 32. [DOI] [PMC free article] [PubMed]
  • 59.He Y, Shishkova E, Peters-Clarke TM, Brademan DR, Westphall MS, Bergen D, et al. Evaluation of the Orbitrap ascend Tribrid mass spectrometer for shotgun Proteomics. Anal Chem. 2023;95(28):10655–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.von Rundstedt Fc, Rajapakshe K, Ma J, Arnold JM, Gohlke J, Putluri V, et al. Integrative pathway analysis of metabolic signature in bladder cancer: a linkage to the cancer genome atlas project and prediction of survival. J Urol. 2016;195(6):1911–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Razavi S, Khan A, Fu DX, Mayer D, McConkey D, Putluri N, et al. Metabolic landscape in bladder cancer. Curr. Opin. Oncol. 2025;37(3):259–66. [DOI] [PubMed] [Google Scholar]
  • 62.Vantaku V, Amara CS, Piyarathna DWB, Donepudi SR, Ambati CR, Putluri V, et al. DNA methylation patterns in bladder tumors of African American patients point to distinct alterations in xenobiotic metabolism. Carcinogenesis. 2019;40(11):1332–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Chen Y, Wang B, Zhao Y, Shao X, Wang M, Ma F, et al. Metabolomic machine learning predictor for diagnosis and prognosis of gastric cancer. Nat Commun. 2024;15(1):1657. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Liebal UW, Phan ANT, Sudhakar M, Raman K, Blank LM. Machine learning applications for mass spectrometry-based metabolomics. Metabolites. 2020;10(6). [DOI] [PMC free article] [PubMed]
  • 65.Zhou X, Sun X, Zhao H, Xie F, Li B, Zhang J. Biomarker identification and risk assessment of cardiovascular disease based on untargeted metabolomics and machine learning. Sci Rep. 2024;14(1):25755. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Moskaleva NE, Shestakova KM, Kukharenko AV, Markin PA, Kozhevnikova MV, Korobkova EO, et al. Target metabolome profiling-based machine learning as a diagnostic approach for cardiovascular diseases in adults. Metabolites. 2022;12(12). [DOI] [PMC free article] [PubMed]
  • 67.Huang L, Wang L, Hu X, Chen S, Tao Y, Su H, et al. Machine learning of serum metabolic patterns encodes early-stage lung adenocarcinoma. Nat Commun. 2020;11(1):3556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Zulqarnain F, Rhoads SF, Syed S. Machine and deep learning in inflammatory bowel disease. Curr Opin Gastroenterol. 2023;39(4):294–300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Uddin S, Khan A, Hossain ME, Moni MA. Comparing different supervised machine learning algorithms for disease prediction. BMC Med Inf Decis Mak. 2019;19(1):281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Summers HD. Practical machine learning for disease diagnosis. Cell Rep Methods. 2021;1(6):100103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Chang Q, Chen P, Yin J, Liang G, Dai Y, Guan Y, et al. Discovery and validation of bladder cancer related excreted nucleosides biomarkers by dilution approach in cell culture supernatant and urine using UHPLC-MS/MS. J Proteomics. 2023;270:104737. [DOI] [PubMed] [Google Scholar]
  • 72.Enokida H, Nakagawa M. Epigenetics in bladder cancer. Int J Clin Oncol. 2008;13(4):298–307. [DOI] [PubMed] [Google Scholar]
  • 73.Luo Y, Yao Y, Wu P, Zi X, Sun N, He J. The potential role of N(7)-methylguanosine (m7G) in cancer. J Hematol Oncol. 2022;15(1):63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Spratlin JL, Serkova NJ, Eckhardt SG. Clinical applications of metabolomics in oncology: a review. Clin Cancer Res. 2009;15(2):431–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Song Z, Uriarte S, Sahoo R, Chen T, Barve S, Hill D, et al. S-adenosylmethionine (SAMe) modulates interleukin-10 and interleukin-6, but not TNF, production via the adenosine (A2) receptor. Biochim Biophys Acta. 2005;1743(3):205–13. [DOI] [PubMed] [Google Scholar]
  • 76.Fukumoto K, Ito K, Saer B, Taylor G, Ye S, Yamano M, et al. Excess S-adenosylmethionine inhibits methylation via catabolism to adenine. Commun Biol. 2022;5(1):313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Fernandez-Ramos D, Lopitz-Otsoa F, Lu SC, Mato JM. S-Adenosylmethionine: a multifaceted regulator in cancer pathogenesis and therapy. Cancers (Basel). 2025;17(3). [DOI] [PMC free article] [PubMed]
  • 78.Loras A, Segovia C, Ruiz-Cerda JL. Epigenomic and Metabolomic Integration reveals dynamic metabolic regulation in bladder cancer. Cancers (Basel). 2021;13(11). [DOI] [PMC free article] [PubMed]
  • 79.Erichsen L, Ghanjati F, Beermann A, Poyet C, Hermanns T, Schulz WA, et al. Aberrant methylated key genes of methyl group metabolism within the molecular etiology of urothelial carcinogenesis. Sci Rep. 2018;8(1):3477. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zhou X, Zhu H, Luo C, Yan Z, Zheng G, Zou X, et al. The role of RNA modification in urological cancers: mechanisms and clinical potential. Discov Oncol. 2023;14(1):235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Wu H, Chen S, Li X, Li Y, Shi H, Qing Y, et al. RNA modifications in cancer. MedComm (2020). 2025;6(1):e70042. [DOI] [PMC free article] [PubMed]
  • 82.Richard E, Gallego-Villar L, Rivera-Barahona A, Oyarzabal A, Perez B, Rodriguez-Pombo P, et al. Altered redox homeostasis in Branched-Chain amino acid disorders, organic acidurias, and homocystinuria. Oxid Med Cell Longev. 2018;2018:1246069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Sivanand S, Vander Heiden MG. Emerging roles for Branched-Chain amino acid metabolism in cancer. Cancer Cell. 2020;37(2):147–56. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Hattori A, Tsunoda M, Konuma T, Kobayashi M, Nagy T, Glushka J, et al. Cancer progression by reprogrammed BCAA metabolism in myeloid leukaemia. Nature. 2017;545(7655):500–04. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Yin J, Ren W, Huang X, Deng J, Li T, Yin Y. Potential mechanisms connecting purine metabolism and cancer Therapy. Front Immunol. 2018;9:1697. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Tran DH, Kim D, Kesavan R, Brown H, Dey T, Soflaee MH, et al. De Novo and salvage purine synthesis pathways across tissues and tumors. Cell. 2024;187(14):3602–18 e 3620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Huang Z, Xie N, Illes P, Di Virgilio F, Ulrich H, Semyanov A, et al. From purines to purinergic signalling: molecular functions and human diseases. Signal Transduct Target Ther. 2021;6(1):162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Yu J, Jin C, Su C, Moon D, Sun MA, Zhang H, et al. Resilience and vulnerabilities of tumor cells under purine shortage stress. Clin Cancer Res. 2025;31(20):4345–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Lu Z, Li J, Liu Y, Li H, Sun Y, Geng R, et al. Purine metabolism in tumorigenesis and its clinical implications. Semin Oncol. 2025;52(6):152409. [DOI] [PubMed] [Google Scholar]
  • 90.Diehl FF, Miettinen TP, Elbashir R, Nabel CS, Darnell AM, Do BT, et al. Nucleotide imbalance decouples cell growth from cell proliferation. Nat Cell Biol. 2022;24(8):1252–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Rosenzweig A, Blenis J, Gomes AP. Beyond the Warburg Effect: how Do cancer cells regulate One-carbon metabolism? Front. Cell Dev. Biol. 2018;6:90. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Data (1-8) (476.5KB, xlsx)
Supplementary Tables (1-9) (241.1KB, xlsx)

Data Availability Statement

Raw and processed metabolomics data from samples are available at NIH Metabolomics workbench (Study ID: ST003990). All the raw data values used for the study are in supplementary data (1–8).


Articles from Cancer & Metabolism are provided here courtesy of BMC

RESOURCES