Abstract
Metabolomics and artificial intelligence (AI) form a synergistic partnership. Metabolomics generates large datasets comprising hundreds to thousands of metabolites with complex relationships. AI, aiming to mimic human intelligence through computational modeling, possesses extraordinary capabilities for big data analysis. In this review, we provide a recent overview of the methodologies and applications of AI in metabolomics studies in the context of systems biology and human health. We first introduce the AI concept, history, and key algorithms for machine learning and deep learning, summarizing their strengths and weaknesses. We then discuss studies that have successfully used AI across different aspects of metabolomic analysis, including analytical detection, data preprocessing, biomarker discovery, predictive modeling, and multi-omics data integration. Lastly, we discuss the existing challenges and future perspectives in this rapidly evolving field. Despite limitations and challenges, the combination of metabolomics and AI holds great promises for revolutionary advancements in enhancing human health.
Keywords: Artificial Intelligence, Metabolomics, Machine Learning, Deep Learning, Systems Biology, Disease Diagnosis, Precision Medicine, Drug Discovery
1. Introduction
Metabolomics, broadly including lipidomics, is a powerful approach to study the metabolome, compositions and interactions of small molecule metabolites within a biological system [1, 2]. Metabolome impacts or is impacted by virtually every cellular process, and altered metabolism is one of the hallmarks for many diseases, nutrition intervention, environmental exposure, etc. [3] The metabolites of interest in metabolomics include sugars, amino acids, lipids, and other organic compounds essential for cellular functions [4]. In terms of metabolic reactions, they can be broadly categorized into two major types: catabolism, breaking down molecules to yield energy; and anabolism, synthesis of compounds for cellular functions [5]. In addition, metabolism participates in deactivation, detoxification, and elimination of exogenous substances [6]. To study the whole metabolome, metabolomics measures thousands of metabolite compositions and abundances in parallel, which has emerged as a highly useful approach to examine metabolic pathways, biochemical reactions, and physiological transformations happening within biological systems [7–9].
Metabolomics often results in large datasets. The two most common analytical technologies in metabolomics are nuclear magnetic resonance (NMR) spectroscopy and mass spectrometry (MS), and MS is often coupled to liquid or gas chromatography (LC-MS, GC-MS) [10]. These technologies can detect hundreds to thousands of metabolites and/or metabolic features from a typical biological sample, paving the way for large-scale exploration of a wide range of metabolome [11]. Based on the large datasets, metabolomic studies can comprehensively analyze the metabolic profiles of biological systems, providing insights into the biochemical processes underpinning internal and/or external stimuli [11–14].
Metabolomics has shown great potential in a wide variety of applications. For example, metabolomics was used to identify metabolic signatures associated with specific diseases, and the potential metabolite biomarkers were used for disease diagnosis, prognosis, and monitoring [8, 15–17]. Roche group and Scalbert group reviewed the applications of metabolomics to improve human health, covering various diseases and human nutrition sciences [14, 18]. Armitage and Ciborowski, as well as Xie et al., discussed how metabolomics was used in cancer research, including the identification of cancer biomarkers and metabolic pathways associated with cancer mechanisms [19, 20]. Botas et al. reviewed metabolomics studies in neurodegenerative diseases, such as Alzheimer’s and Parkinson’s diseases [21]. In addition, metabolomics can discover new drug candidates and study drug metabolism, toxicology, efficacy, and potential side effects [22–25]. Kell et al. reviewed the applications of metabolomics in targeted drug discovery [22]. Robertson et al. discussed the challenges and opportunities of metabolomics in drug metabolism and pharmaceutical development [24]. Furthermore, Chen et al. reviewed metabolomics in studying the metabolic products of human gut microbiome [25]. We also published a comprehensive review of recent studies investigating xenobiotic exposures of wide interest and their effects on the gut microbiome and metabolome [26]. Notably, metabolomics offers a functional assessment of cellular activity, and from a systems biology point of view, it can be integrated with various omics domains (such as genes and proteins) to provide a “big picture” of a biological system [27–29]. Both metabolomics itself and this complex interplay require sophisticated data analytical approaches.
Artificial intelligence (AI) makes uses of machines and/or software to perform tasks that typically require human intelligence, such as understanding natural languages, recognizing patterns, making decisions, and solving problems [30–32]. When using complex datasets, AI has demonstrated remarkable proficiency in recognizing hidden patterns, discerning informative features amidst noisy backgrounds, and delivering accurate predictions [33, 34]. These capabilities make it exceptionally well suited for metabolomic data analysis. Indeed, innovative AI algorithms have been developed to process raw mass-spectrometry data, detect chromatogram peaks, identify metabolites, discover biomarkers, build predictive models, construct pathways, and other tasks that challenge traditional statistical approaches [35–37]. In addition, AI shows significant promises in integrative multi-omics analysis, thereby advancing the field of systems biology.
Current literatures have comprehensive reviews of metabolomics techniques and applications [3, 24, 38–42], as well as development and applications of AI in fundamental and applied sciences [43–47]. In this article, we present a systematic review of AI techniques and applications in metabolomic studies in the framework of systems biology. We first introduce the AI concept, history, and landscape of the AI algorithms that are integral to big data analysis. We then present a comprehensive review of AI applications across different aspects of metabolomics studies, highlighting the unique characteristics of metabolomics data and important considerations for effective analysis and appropriate interpretation. We further expand the discussion to underscore the roles of AI in integrative multi-omics analysis. Finally, we discuss the existing challenges and future perspectives in this rapidly evolving field, outlining potential directions for AI-driven metabolomics.
2. AI History, Concepts, and Techniques in Big Data Analysis
Since November 2022, many of us have been greatly impressed by Chat Generative Pre-trained Transformer (ChatGPT), which uses natural language processing (NLP) to create humanlike conversational dialogues [48–50]. ChatGPT is a typical example of AI to simulate and replicate human intelligence, and in fact, AI can do much more.
As shown in Figure 1, the origin of AI can be tracked back to 1950s when the field of computer science just emerged. In 1956, the term “artificial intelligence” was formally introduced at the historic Dartmouth Conference [51]. Early AI systems were used for symbolic reasoning, problem-solving, mathematics, and logical thinking [52]. Interestingly, one of the first AI applications, the Dendral project (1965), was developed to analyze MS data to identify organic compounds [53]. Between 1970s-1980s, the expert systems, as one of the first practical applications of AI, were created by Edward Feigenbaum to simulate decisions of human experts [54, 55]. For example, MYCIN (1976) and Prospector (1983) were well-known expert systems used for medical diagnosis and analyzing geological data in mineral exploration, respectively [35, 56]. Since 1990s, AI techniques, mainly including machine learning (ML) and neural networks (NNs), have gained wide applications in data analysis and pattern recognition. For instance, ML was used for sequence analysis and identifying genetic markers in genomics, and clustering algorithms were employed in chemistry for compound classification [57–59]. Cartwright discussed the impact of ML in chemistry, covering applications in molecular modeling, property prediction, and drug discovery [60]. From 2010s to now, the explosion of big data and advances in computing power facilitated the use of deep learning (DL) techniques. Deep NNs were employed in image analysis for tasks like medical imaging and astronomy [61, 62]. In climate science, AI is able to analyze large datasets to model and predict complex climate patterns [63, 64]. In addition, AI has been used to develop and control robots for data collection in extreme environments, such as deep-sea exploration and planetary exploration [65]. AI-assisted robotic systems have automated high-throughput experimentation in biology and chemistry [66, 67]. Further, AI was used to optimize quantum algorithms and extract valuable information from scientific literature using NLP, which is highly useful in precise and high-throughput drug discovery [68].
Figure 1.

History of artificial intelligence and its connections to metabolomics.
Currently, AI continues to rapidly grow and diversify, with applications in various disciplines, from biology and chemistry to astronomy, climate science, and beyond [69]. Petersen et al. reviewed the use of AI in drug discovery, biomarker identification, and personalized medicine [70]. Lai et al. reviewed how AI and ML are used for analyzing genomics data, including gene expression, DNA sequencing, and variant calling [71]. Impressively, AI has outperformed human intelligence in tasks that demand processing vast volumes of data, discerning intricate relationships, and making rapid decisions, where human cognitive processes may be comparatively slower or more prone to errors [72]. As the technology advances, it is expected that in the near future AI will bring revolutionary improvements in many sectors including banking, transportation, healthcare, etc.
2.1. Machine Learning (ML)
As shown in Figure 1, ML is a subfield of AI that focuses on developing algorithms and models that enable computers to learn from the data and improve their performance without manual intervention [73]. There are 3 major types of ML algorithms, supervised learning, unsupervised learning, and reinforcement learning (Table 1).
Table 1.
Major types of ML algorithms
| Supervised Learning |
| Linear classifiers: linear regression, logistic regression (LogReg), mixed-effect regression, and other generalized linear models, partial least squares (PLS) regression, partial least squares discriminant analysis (PLS-DA), regularized regression |
| Non-linear classifiers: K-nearest neighbors (KNN), decision trees, random forest (RF), gradient boosting, support vector machine (SVM), naive Bayes, artificial neural network (ANN) |
| Unsupervised Learning |
| Clustering: K-means clustering, hierarchical clustering, self-organizing map (SOM) |
| Dimension reduction: principal component analysis (PCA), Independent Component Analysis (ICA), t-Distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP) |
| Reinforcement Learning |
| Q-learning, Deep Q-Networks (DQN), Policy Gradients, Actor-Critic, Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradients (DDPG), Trust Region Policy Optimization (TRpO) |
In supervised learning, the objective is to learn the mapping from input features to output labels [74, 75]. Once a model is trained, it can make predictions on new, unseen data. Supervised learning has been widely used in metabolomics studies. For example, Catharino’s group combined MS with supervised ML analyses to discover biomarkers predictive of body weight [76]. Tiedt et al. used supervised ML and found a set of four metabolites that accurately distinguished between patients with ischemic stroke and those with stroke mimics [77]. Notably, ensemble learning can be employed to improve predictions from multiple models, further enhancing the overall performance and robustness [78].
Unsupervised learning does not use output labels to guide the learning process. The objective is to unveil inherent structures in the data, solely based on the input features [79]. Clustering is a common unsupervised learning task, which aims to partition data into groups based on their similarities. The discovered clustering patterns can aid in identifying novel sample subtypes, exploring correlated features, detecting anomalies, and revealing other distribution characteristics worthy of further examination [80]. Another unsupervised learning application is dimensionality reduction, which transforms data from a high dimensional space onto a low dimensional space while preserving the most meaningful information of the original data [81]. Data in the reduced dimensions can be easily visualized to inspect structures, and the transformed features have lower redundancy and higher information content compared to the untransformed ones [82]. Dimension reduction algorithms and clustering algorithms are often utilized together to explore and visualize structures in a metabolomic dataset, such as grouping metabolites with similar concentration profiles [83, 84], discovering sample subsets with similar metabolic profiles [85], identifying novel compounds [86], and detecting unexpected patterns [87].
In reinforcement learning, the task is to make sequential decisions that are robust to stochasticity, uncertainty, partial observability, and other challenges in a complex environment [88]. The algorithm optimizes a model via interactions with the environment in trials and errors, during which desired behaviors are rewarded and undesired ones are penalized [89]. While reinforcement learning has achieved great successes in fields like robotics, game playing, and autonomous systems, its application in systems biology is still in its early stages [90]. The demand for substantial amounts of data and extensive training makes it computationally intensive and impractical in biomedical scenarios with limited data and high sample complexity [91]. Therefore, simplification of complex real-world biological mechanisms into more manageable systems is frequently necessary. For instance, Hu et al. selectively incorporated the influxes/outfluxes of six bile acids across six organs as parameters in a differential equation and employed reinforcement learning to approximate the body’s adjustment, treating homeostasis as a reward signal [92]. This model successfully predicted the disease progression of primary sclerosing cholangitis.
2.2. Deep Learning (DL)
As shown in Figure 1, DL is a subset of ML algorithms that use multiple layers of interconnected artificial neurons to simulate the way how the human brain works [93]. Compared to traditional ANNs where neurons are organized in a few fully-connected feedforward layers, DL models feature diverse types of neurons arranged in deep hierarchical structures [94]. This design allows DL models to automatically capture complex information within the data, and it has been used for supervised, unsupervised, or reinforcement learning tasks that are challenging for traditional ML approaches [95]. Leveraging these powerful and versatile DL technologies, myriads of models have been built to analyze and integrate genomic, transcriptomic, and epigenetic data [96, 97], as well as to develop novel drugs [98].
Recently, these DL models have also begun to find applications in metabolomics. For example, a convolutional neural network (CNN) can autonomously construct informative features from spatiotemporally distributed signals, eliminating the need for time-consuming manual feature engineering [99]. Kim et al. designed a CNN to process 1H-13C NMR spectra, achieving precise detection of peaks in complex mixtures [100]. ChromAlignNet uses the recurrent neural network (RNN) technology, which allows information to be memorized and passed over multiple layers to form feedback loops [101], and it demonstrated superior accuracy over traditional methods to align GC-MS chromatograms [102]. Choudhury et al. built kinetic models for dynamical metabolism analysis [103] using a generative adversarial network (GAN) that enables a competition between two NNs, one trained on real samples and the other trained on synthesized artificial samples, to enhance model performance [104].
Attention mechanisms can significantly improve the performance of a DL model by selectively focusing on important elements in the input [105]. A groundbreaking development in DL architecture was the introduction of a model solely based on attention, known as Transformer [106]. The Transformer model utilizes a self-attention mechanism, allowing it to weigh different parts of the input sequence, which enables it to effectively capture long-range dependencies and surpass traditional RNNs in various tasks [107]. In fact, ChatGPT, as mentioned above, harnesses the Transformer architecture to comprehend and generate human-like texts [108]. Using the Transformer framework, Shrivastava et al. developed a MassGenie tool to explore millions of fragmentation patterns in silico, which showed an unprecedented ability to identify de novo small molecules from MS data [109].
3. AI in Metabolomics
3.1. Overview
Applying AI models on metabolomics data suits well because of large datasets and complex patterns among them. The integration of AI in metabolomics has greatly improved data analysis, biological interpretation, and predictive statistical modeling [110, 111]. As a result, AI-assisted metabolomics has made significant discoveries in diverse applications including disease diagnosis [112], precision medicine [113], drug discovery [114], etc.
Figure 2 depicts the general metabolomics workflow. Metabolomics is a comparative science, and it needs samples from two or more groups under different endogenous and/or exogenous conditions [10]. Various kinds of biological samples have been used, such as blood, urine, tissue, and cells. These samples are subject to untargeted (broad screening) and/or targeted (only metabolites of interest) metabolomics analysis. NMR and LC/GC-MS are common analytical technologies to measure a wide range of metabolites, generating large datasets. These analytical platforms require different sample preparation protocols to extract metabolites and improve detection [115–117]; for example, chemical derivatization is often performed to volatilize the compounds prior to GC-MS detection [118]. After collection of the raw data, data will be extracted, which includes steps of background correction, peak picking, peak deconvolution, peak integration, etc. In addition, metabolites are identified by comparing their experimental data to chemical standards, reference libraries or databases, and their concentrations can be measured in targeted measurement [119]. Prior to statistical analysis, data pretreatment involves baseline correction, binning, scaling, and/or normalization. In statistical modeling, both univariate and multivariate statistics can be used to discover significant metabolites. These metabolites are then mapped to metabolic pathways and integrated with other omics data to reconstruct the molecular networks. An optimal set of metabolites and/or genes/transcripts/proteins can be obtained to accurately distinguish samples under different conditions [120]. Further, internal and external cross validations are required to ensure study rigor in metabolomics studies [121]. In the following sections, we will discuss in details how AI can be applied to different steps in metabolomics.
Figure 2.

A general workflow in metabolomics.
3.2. AI in Analytical Detection
LC-MS and GC-MS data contain abundances of not only identified metabolites, but also those without annotations [122]. To identify these unknows, AI is highly efficient to compare the mass spectra and retention times (RTs) from the experiments with authentic chemical standards as well as those present in existing databases [123]. For example, Jang et al. constructed a standalone software named AI-SIDA (AI-screener of illicit drugs and analogues), which provided a very powerful platform for screening unknown erectile dysfunction (ED) drugs and analogues [124]. The OngLai algorithm was developed to detect homologous series using the RDKit [125]. Entropy of AI was used for the classification of stilbenoid compounds [126]. In addition, AI algorithms have demonstrated exceptional capabilities to predict compound structures, even in cases where no prior reference exists [127]. AlphaFold2 (AF2) was developed by DeepMind to predict 3D structures of chemical compounds with an atomic-level accuracy. These multifaceted AI approaches for metabolite identification not only accelerate the annotation process but also are promising to discover novel compounds that have never been studied before [128].
AI algorithms were also utilized to analyze complex spectral data. For example, Zhang et al. developed the DeepSpectra model to learn patterns from raw spectroscopic data and improve the quantitative performance [129]. The results showed that DeepSpectra generated improved quantitation results than conventional linear and nonlinear calibration approaches, through minimizing artifacts, retaining useful patterns, and providing better repeatability and accuracy.
3.3. AI in Data Extraction and Preprocessing
The primary goal of data extraction in metabolomics is to convert complex, high-dimensional raw data into a format that can be used for downstream analysis and interpretation [111, 130]. CNNs have been proven particularly effective at these tasks, such as processing, selecting, and integrating chromatographic peaks with strong local correlations, outperforming other methods by an average increase of 6% to 24% in classification accuracy [131, 132]. In addition, CNN exhibits less dependency on the data preprocessing methods compared to other chemometric techniques [133], as well as the ability to deconvolute overlapping peaks in complex chromatograms [134]. Serrano group introduced a novel ANN model derived from PUNNs [135]. This model significantly improved peak selection, which makes it effective in quantifying analytes with overlapped chromatographic peaks from a single detector instrument [135].
The use of RT, along with the m/z ratio, can enhance metabolite identification in untargeted metabolic profiling [136, 137]. The accuracy of peak picking is high when the RT of a metabolite is known; however, RT prediction is challenging when the metabolite is unknown. To address this challenge, many ML methods have been developed. Fiehn’s group integrated five different ML algorithms into the Retip R package, and they observed a significant decrease (68%) in the number of candidate structures when conducting a search for all isomers in mouse plasma samples [138]. In the initial phase of RT prediction, structural information is encoded into a vector format. This can be achieved using techniques such as quantitative structure–retention relationship (QSRR) models, or molecular fingerprints [139–141]. Liapikos et al. developed a model using four ML regression algorithms (Bayesian Ridge Regression, Extreme Gradient Boosting Regression, and Support Vector Regression (SVR) using both linear and non-linear kernels), in which the molecular descriptors were employed to represent the physical, chemical, and structural properties of the molecules in suitable mathematical structures [142]. In 2019, Bouwmeester et al. carried out a comprehensive comparison of various ML methods for predicting RTs in LC [143]. 151 features were extracted from the simplified molecular input line entry system (SMILES) notation to train seven models, incorporating both linear and nonlinear approaches. Stand-alone models, such as ANN and SVM, demonstrated high accuracy in the RT prediction task. In addition, the results highlighted the superior performance of ensemble methods that combined multiple ML tools.
Metabolomics data may have systematic errors, including batch differences, longitudinal drifts, instrument-to-instrument variation, etc. To minimize these technical variances, Kehoe et al. developed a sparse SVM model to address batch effects that are a common challenge in metabolomics studies with many samples [144]. Shen et al. developed the SVR method to reduce unwanted intra- and inter-batch variations and integrate multiple data batches in large-scale metabolomics studies, which significantly improved the performance of downstream multivariate statistical analyses in terms of classification and prediction accuracy [145]. The ability of ML methods to handle non-linear data makes them suitable for capturing complex patterns and removing systematic errors in the data. Fan et al. performed systematic error removal using the random forest (SERRF) approach. They showed that SERRF outperformed 15 other commonly used normalization methods, significantly reducing the unwanted systematic variation and revealing the biological variance of interest [146].
Class imbalance, in which each group has a very different sample size, may introduce severe biases into model fitting and performance assessment. To address this issue, oversampling techniques such as Synthetic Minority Over-sampling Technique (SMOTE) [147], as well as undersampling techniques such as random subsampling [148] and bootstrapping [149] are commonly used. These techniques are employed to balance classes in the training data but not in the validation and testing data, such that the resulted performance reflects the original sample size ratios. Gui compared over-sampling (OS), under-sampling (US), and SMOTE, and as a result, SMOTE technique was shown to have the minimum value of Cost [150].
3.4. AI in Statistical Modeling
Given preprocessed data, statistical modeling aims to uncover complicated metabolic interactions, identify novel biomarkers, and improve our understanding of various biological processes [151–153]. To select suitable AI algorithms, adjust the configuration of a model effectively, and design new algorithms, a profound understanding of the characteristics inherent in metabolomics data is imperative. Key aspects include:
Curse of dimensionality [154]: While metabolomic experiments can profile a vast array of metabolites often numbering in the hundreds to thousands, the high costs often limit the number of samples examined. This results in datasets with the number of variables/metabolites much larger than the number of samples, which is prone to cause data overfitting.
Complex relationship between metabolites [155]: As different metabolites are jointly involved in a multitude of biochemical reactions and pathways, their concentrations may show various levels of correlation.
Large variations [156]: Metabolite levels can change rapidly in response to intrinsic and extrinsic stimuli. The biological variation among individuals or experimental replicates may overshadow subtle metabolic changes of interest. In addition, covariates, such as age, gender, and demographic features, also affect metabolite levels. These covariates shall also be considered together with the metabolites in feature selection and predictive modeling.
Ambiguity in metabolic causality: Metabolic changes could either cause or result from phenotypic changes. Given the diversity of metabolites involved, it is challenging to ascertain which ones should be modeled as dependent variables and which as independent variables.
A variety of strategies have been developed to address these challenges at each stage of the analysis process, mainly including feature extraction and selection, as well as modeling and evaluation.
3.4.1. Feature Extraction and Selection
Feature extraction and feature selection both aim to identify the most relevant and insightful features while reducing the data dimensionality [157]. Feature extraction accomplishes this goal by transforming the original features into a smaller set of artificial features, typically using the unsupervised learning algorithms described in Section 2.1. However, the interpretation of the transformed artificial features can be challenging. In contrast, feature selection employs supervised ML algorithms to pinpoint metabolites associated with the phenotype/condition of interest [158].
Figure 3 and Table 2 summarize four major categories of feature selection strategies. Filter methods, encompassing various univariate and multivariate statistical tests, are the most commonly used strategies due to their simplicity and efficiency [20, 159, 160]. However, the selected features may not be the best predictors to distinguish samples of different phenotypes. Wrapper methods address this issue by adding an evaluation step that uses the selected metabolites to build a predictive model and choosing the best-performing set. For example, Gaul et al. combined a recursive feature elimination (RFE) step with an SVM evaluator, which generated a 16-metabolite predictive model to diagnose early-stage serous epithelial ovarian cancer [161]. Embedded methods, such as decision tree, RF, and least absolute shrinkage and selection operator (LASSO) [162], simultaneously perform feature selection and model training. As an example, LASSO was applied to construct significant lipid profiles for different molecular subtypes of breast cancer [163]. When these methods individually cannot produce satisfactory results, hybrid techniques to combine them could be used. To predict renal cell carcinoma status, Bifarin et al. first used filter techniques to narrow down the feature set from 7,147 NMR and MS features to 128 based on t-test p-values, fold changes, and correlation tests [164]. They then utilized wrapper approaches, specifically RF-RFE and PLS, to further refine this selection to the top 20 features. Ultimately, a vote-based approach was used to identify a highly effective panel of seven metabolites [164].
Figure 3.

Illustration of the four categories of feature selection strategies.
Table 2.
Comparison of the four categories of feature selection strategies
| Filter | Wrapper | Embedded | Hybrid | |
|---|---|---|---|---|
| Advantages | Computationally efficient; independent of model. | Tailored to model performance; can capture feature interactions. | Efficient as part of model training; directly linked to model’s predictive power; reduces overfitting risk. | Balances efficiency and accuracy; flexible. |
| Limitations | May miss important interactions; not tailored to specific model performance. | Computationally intensive; risk of overfitting; dependent on specific model. | Tied to specific model; can be complex; interpretability may vary. | computationally intensive; requires careful balance and tuning. |
| Best Use in Metabolomics | Feature reduction before modeling in very large data sets with high dimensionality. | When computational resources allow and model performance is critical, it may produce smaller, well-defined feature sets. | Suited for datasets where model integration is beneficial and /or when specific models are preferred. | When a single method is insufficient and/or when a balanced approach is needed. |
| Examples | t-test, correlation test, linear regression, ANOVA, etc. | Forward selection, backward elimination, RFE, etc. | LASSO, decision trees, etc. | Combines filter, wrapper, and embedded methods. |
3.4.2. Predictive Modeling
Table 3 compares commonly used ML algorithms with specific examples of their applications in metabolomic modeling. AI offers significant advantages over traditional statistical methods for modeling complex, non-linear relationships between features and responses. For targeted metabolomics studies where the datasets are relatively small, or when interpretability is crucial, methods that provide clear insights into data relationships, such as decision trees or traditional statistical regression, are preferable. However, for untargeted metabolomics studies where the datasets are large, or when predictive performance is paramount and interpretability is less critical, complex “black box” models like DL may be more suitable. Iterative testing is crucial for selecting the optimal model. Initially, a variety of models are tested, from simple ones like linear regression to more complex types such as RF, SVM, and DL, to establish baseline performance. Based on these initial results, model selection will be refined by tuning hyperparameters and re-evaluating performance. This structured approach ensures that the chosen method optimally addresses the study objectives, such as classifying samples based on phenotypes or conditions, or predicting quantitative outcomes.
Table 3.
Comparison of ML algorithms for predictive modeling
| Algorithm | Advantages | Limitations | Examples |
|---|---|---|---|
| PLS-DA | Supporting modeling multiple metabolites as features or responses. | Assumptions about data distribution. | Fabritiis et al. (2023) [166] |
| RF | Providing feature importance scores, robust to overfitting. | Can be biased with categorical variables, large trees can be complex. | Chen et al. (2021) [167] |
| SVM | Flexible with kernel choice. | Choice of kernel critical, computationally intensive for large datasets. | Kang et al. (2022) [168] |
| KNN | Intuitive, simple, effective in data with clear local structures. | Sensitive to k and distance metric, struggles with high-dimensional data. | Wang et al. (2022) [169] |
| CNN | Excellent for capturing spatial hierarchies, robust to translation and rotation. | Requiring large training datasets, computationally intensive. | Kim et al. (2022) [100] |
| RNN | Effective for modeling time-dependent data, captures temporal dependencies. | Requiring large training data, prone to vanishing gradient problem, complex to train. | Li et al. (2019) [102] |
| GAN | Capable of generating realistic new samples, enhancing data diversity. | May produce unrealistic samples, challenging to train. | Choudhury et al. (2022) [103] |
The need for large training datasets has historically limited the applications of DL in metabolomics [165]. Recently, however, Choudhury et al. employed GAN to overcome this limitation by generating synthetic data that closely mimic real metabolite profiles [103]. By augmenting a small sample set with a large synthetic dataset, the Reconstruction of Kinetic Models using Deep Learning (REKINDLE) model was trained. This approach allows REKINDLE to explore a vast parameter space, ensuring that sufficient diversity is incorporated in the learned model.
4. AI-assisted Metabolomics in Systems Biology
Metabolomics is one of the omics fields in systems biology. Systems biology is an interdisciplinary area to understand the big picture of a biological system, often at the levels of organism, tissue, and cell [170–174]. As shown in Figure 4, a fundamental process in biology involves the expression of genetic information (genotype, encoded in DNAs) as observable characteristics (phenotype) [175, 176]. This transition is governed by various molecular and cellular processes. To initiate it, genes are expressed in a multi-step process that includes transcription and translation, synthesis of RNAs from DNAs and proteins from RNA sequences, respectively [177]. Although not all genes are always expressed, the resulting proteome is crucial in determining phenotypes. For example, the proteasome is a highly sophisticated protease complex that is responsible for the degradation of proteins, and it plays an important role in regulating various cellular metabolism [178, 179]. The proteasome is extensively involved in protein turnover, regulation of specific metabolic pathways, and control of key metabolic processes [180]. At the phenotype level, metabolome provides significant insights into how genetic information and environmental factors (co-)influence the composition and abundance of metabolites (small molecules) in cells, tissues, and organisms [181–185]. With combined input from genome, transcriptome, proteome, and metabolome, systems biology studies the interactions among these components and aims to gain a comprehensive understanding of a biological system.
Figure 4.

Systems biology: from genotype to phenotype in a biological system.
4.1. Metabolic Network and Pathway Analysis
AI can assist in the construction and analysis of metabolic networks and pathways [186–188]. AI algorithms were used to provide a comprehensive understanding of how metabolic pathways were regulated and perturbed in different conditions [189]; therefore, it gained a deeper understanding of the complex interactions and networks of metabolites in biological systems [151, 152]. ML algorithms, especially DL techniques, are applied to metabolic pathway and network analysis to address the complexity and high-dimensional nature of the data [111, 190]. Jin et al. showed that graph DL was a novel approach to process network structures and analyze large, multidimensional, and complex biological data [188].
Elovici group used a ML model to map metabolic pathways onto the metabolite correlation networks and compute network properties for each pathway [191]. Using this RF model, they predicted previously unidentified pathways and successfully validated the predictions in vivo [191]. Their approach involved associating metabolites with pathways and using measured metabolite correlations to calculate feature vectors for each pathway. These feature vectors were generated using a combination of statistical, graph-related, and correlation network metrics. The RF model was then trained to classify the activity of 169 organism-related active pathways from the MetaCyc database [192], along with 85 non-active pathways and 85 randomly selected metabolite combinations. Notably, this approach is limited to identified metabolites and predefined pathways during the training. In another study by Hosseini and colleagues, pathway activity was weighed based on the likelihood of metabolites being connected to the pathway [193]. A generative model was developed to establish links between pathway activity probabilities, metabolites, and measured spectral masses. This approach emphasized metabolites that were unique to a specific pathway, resulting in novel predictions compared to standard enrichment analysis.
Metabolic flux analysis (MFA) is also broadly included in the big field of metabolomics, and MFA provides a dynamic description of metabolic networks in a biological system [194, 195]. MFA tracks the footprint of isotopically labeled substrates, such as 13C6-glucose, in metabolic pathways to examine the impacts of internal or external stimuli [196]. Using the ANN approaches, Antoniewicz et al. predicted the fluxes of 13C-labeled metabolites in mammalian gluconeogenesis [197]. Compared to traditional methods that require many samples (>200), ANN algorithms, including multiple linear regression (MLR), PCR, and PLS, performed much better for flux prediction using the mass isotopomer data.
4.2. Integrative Analysis of Multi-omics Data
Integrating metabolomics data with other omics data presents unique challenges due to the inherent differences in metabolites compared to genes, proteins, and other molecular components. Table 4 presents the data characteristics in different omics that need to be accounted for in integrative analysis. First, harmonizing data scales is challenging because the data units vary by omics types and instruments [198]. For example, metabolites are measured in relative abundance (e.g., counts per second, cps, in MS) or absolute quantitation (mM, μM, or ng/mL), while genomic sequences are reported as A/T/C/G characters, and gene expression abundance is reported as continuous values or discrete counts. In addition, the diversity of chemical structures and the presence of isomers and modifications introduce uncertainties in integrated analysis [199], which are higher in metabolomics than other omics as most genetic variants and transcripts can be accurately mapped to unique genomic regions [200]. The sparsity of metabolomics data, where many metabolites have dramatically different concentrations, complicates the process of finding meaningful associations with other omics data [201]. Furthermore, metabolites often participate in multiple pathways, making it difficult to link changes in metabolite levels to specific genetic or protein-level changes due to the pathway crosstalk. High biological variation of metabolites, influenced by genetics, diet, and environmental factors, adds another layer of complexity [202].
Table 4.
Characteristics of various omics data in systems biology
| Feature | Metabolomic Data | Genomic Data | Transcriptomic Data | Proteomic Data | Microbiome Data |
|---|---|---|---|---|---|
| Nature | Full set of small molecule metabolites in a biological sample | Entire DNA sequence of an organism, including all genes | Full set of RNA transcripts from the genome at a specific time | Complete range of proteins produced in a system | Genetic material of all microbes in a specific environment |
| Stability | Highly dynamic; changes rapidly in response to internal and external factors | Relatively stable; changes only due to mutations | Dynamic; varies with time, conditions, and tissue type | Dynamic; influenced by post-translational modifications and conditions | Dynamic; varies with environmental factors and host interactions |
| Dynamics | Shows real-time metabolic activity and state | Static representation of potential cellular capabilities | Indicates active gene expression at a given moment | Reflects the proteins currently present and functional | Reflects the composition and functions of microbial communities |
| Information | Biochemical activity and metabolic state | Genetic potential and hereditary information | Gene expression levels and activity | Functional state at the protein level | Microbial diversity, community structure, and potential functions |
| Applications | Disease diagnosis, drug metabolism, nutrition | Hereditary diseases, genetic predispositions | Disease mechanisms, biomarker discovery | Disease mechanisms, drug targets, biomarkers | Human health, disease states, environmental studies |
| Challenges | Diversity of metabolites, sensitivity of detection | Data size, interpretation of variants | Temporal and spatial complexity, data volume | Protein modifications, complexity of interactions | Diversity of species, functional interpretation |
Therefore, it is critically important to choose appropriate data integration methods that accommodate the unique characteristics of metabolomics data. Currently, new AI methods and tools are facilitating the integration of metabolomics with other omics to extract meaningful insights [203]. For example, ML and network-based approaches can help identify connections between genes, enzymes, and metabolites, with the potential of discovering previously unknown metabolic pathways [152, 186]. In addition, ML methods have been used to analyze multi-omics data, resulting in groundbreaking discoveries in cancer molecular biology [204]. These AI approaches can lead to a deeper understanding of biological mechanisms, and in the case of diseases, the discovery of novel therapeutic targets and more effective treatments.
The AI integration of multi-omics mainly uses approaches of sequential integration and simultaneous integration. Sequential integration adopts a step-wise approach, where omic datasets are initially analyzed individually or in specific combinations, followed by the integration of their findings in subsequent stages. The sequential integration method allows for many types of data analyses, even in cases where omic measurements for the same samples are unavailable. It is particularly suitable for bulk datasets, as demonstrated by a previous study that successfully integrated ATAC-seq and RNA-seq data to identify critical genes and regulatory pathways associated with the neuroprotection of S-adenosylmethionine (SAM) against perioperative neurocognitive disorder (PND) [205].
Simultaneous integration involves the parallel analysis of multiple omic layers, allowing the analysis of all available omic data at once within a single modeling step. While this approach requires that data originate from the same biological samples or individuals, it is a powerful method for deriving valuable insights into cellular functions by capturing the interplay between different omic datasets. These integration methods can be broadly categorized into concatenation, model-based, and transformation-based approaches [206].
Concatenation-based integration methods involve the direct merging of data matrices from individual omics domains, creating a comprehensive multi-omic dataset. This integrated data is then used as the input for various AI techniques. Concatenation-based integration can also incorporate unsupervised methods, as demonstrated in a study on muscle-invasive bladder cancer (MIBC) [207].
Model-based integration methods involve a multi-stage process where individual models are first developed for different omics data types, and then these models are combined into a joint model, allowing for the integration of heterogeneous data sources. Model-based integration methods can be further categorized as either supervised or unsupervised. As an illustrative example, Multi-omics Supervised Autoencoder (MOSAE) was applied to the TCGA Pan-Cancer dataset and to effectively predict four distinct clinical outcomes [208].
Finally, transformation-based integration methods begin by converting individual omics datasets into graphs or kernel matrices, which are then combined to create a unified representation and capture the relationships between diverse omics data. The versatility of these methods is a significant advantage, allowing them to integrate various omic types effectively. For instance, the fast multiple kernel learning for dimensionality reduction (fMKL-DR) method employs kernel matrix transformation and an SVM classifier to effectively stratify samples [209]. Weighted-nearest neighbor (WNN) method is an example of transformation-based unsupervised algorithm that effectively integrates single-cell multimodal datasets [210].
4.3. Biological and Functional Integration
AI algorithms were used to incorporate biological and functional information. Because of the intercorrelation between metabolites and biological variations, it is often challenging to identify reproducible and generalizable biomarkers. A RF model was developed to derive a feature importance score based on the existing biological knowledge [211]. This updated RF model was used in recognizing features, and it provided more discriminating power. In addition, metabolites were ranked using prior knowledge, and this rank was successfully used to guide feature selection and predictive modeling [212, 213].
AI methods were also used to improve biological interpretation. For example, local interpretable model-agnostic explanations (LIME) [214] and SHapley Additive exPlanations (SHAP) analysis [215] were used to interpret significant variables within the disease context. In addition, strategies, such as simplifying complex models and prioritizing functionally important features, were employed to enhance interpretability. Further, unsupervised learning methods combined with visualization techniques, such as PCA followed by t-SNE plot, as well as clustering results displayed in dendrograms are useful to uncover hidden structures in the data.
5. AI-assisted Metabolomics in Health
5.1. Disease Biomarker Discovery
AI has found many applications in disease diagnosis and classification in the field of metabolomics. Many diseases induce altered metabolism, resulting in measurable changes in the levels of metabolites. As an example, it has been well known that cancer cells preferentially use glycolysis as their source of energy, a phenomenon known as the Warburg effect [216]. In subjects with congestive heart failure, cardiomyocytes resume the fetal metabolic program that favors glucose versus fatty acids [217]. AI algorithms have the capacity to identify distinct metabolite patterns from different groups, and these patterns can subsequently serve as valuable tools for disease diagnosis and progression monitoring [218, 219]. Table 5 summarizes the recent applications of AI-assisted metabolomics in human health. Turi et al. and Jendoubi et al. reviewed ML approaches for analyzing metabolic profiles, focusing on biomarker discovery and disease diagnosis [126, 220]. In particular, AI techniques were used for identifying metabolite biomarkers for various diseases, such as cancer, cardiovascular diseases (CVD), and Alzheimer’s disease (AD) [221–223].
Table 5.
Summary of AI-assisted Metabolomics in Health
| AI Algorithm | Strength | Biomarker | Analytical Platform | Sample | Application | Data Source | Reference |
|---|---|---|---|---|---|---|---|
| SVM-RFE | Effective in high-dimensional spaces | 16 diagnostic metabolites | UPLC-MS | Serum | Epithelial ovarian cancer | N/A | Gaul et al. (2015) [161] |
| LogReg | N/A | 63 metabolites | RP-HPLC-QQQ/MS | N/A | Bladder cancer | N/A | Kouznetsova et al. (2019) [253] |
| FCBF | N/A | 5 biomarkers | Targeted LC/MS | Lung tumor tissue | Lung cancer | N/A | Xie et al. (2021) [20] |
| ML-based LDA | N/A | 10 ratios of metabolites, 2 ratios of metabolites | LC-MS/MS | Tumor tissue | Cancer | N/A | Wallace et al. (2020) [254] |
| DL, RF, SVM, RPART, LDA, PAM, GBM | N/A | 20 metabolites, 8 pathways | GC-TOF/MS | Tumor tissue | Breast cancer | http://www.ncbi.nlm.nih.gov/geo. | Alakwaa et al. (2018) [255] |
| OPTICS | Hierarchical, easily interpretable data structure | Imaging | N/A | Tumor tissue | Cancer | http://www.ebi.ac.uk/metabolights/MTBLS415 | Inglese et al. (2017) [224] |
| Multiple LogReg, ADTree | Higher accuracy | 31 metabolites | CE-TOF-MS, LC-QQQ-MS | Saliva | Breast cancer | N/A | Murata et al. (2019) [256] |
| RF | N/A | 5 biomarkers | LTQ-XL Orbitrap Discovery | Plasma | Weight gain | http://dx.doi.org/10.21227/k446-fp12 | Dias-Audibert et al. (2020) [76] |
| RF | N/A | 17 discriminatory metabolites | LC-MS/MS | Stool | Cirrhosis | metabolomic data were deposited at GNPS | Oh et al. (2020) [257] |
| LogReg | N/A | 68 metabolites | RPLC-MS | Plasma | Human pregnancy metabolome | https://doi.org/10.21228/M81H58 | Liang et al. (2020) [258] |
| RF, LDA, LogReg, KNN, naive Bayes, SVM | N/A | 4 metabolites | RP/UPLC-MS/MS, HILIC/HPLC-MS/MS, LC-Q-Exactive high resolution MS | Serum | Ischemic stroke | N/A | Tiedt et al. (2020) [77] |
| RFE, PLS, RF, KNN | Optimization of hyperparameters | 10 metabolites | LC-MS | Urine | Renal cell carcinoma | http://dx.doi.org/10.21228/M8P97V | Bifarin et al. (2021) [164] |
| MUVR | Simultaneously identify minimal-optimal and all-relevant variables | 13 metabolites | LC-MS | Serum | Gout and asymptomatic hyperuricemia | N/A | Shen etal. (2021) [259] |
| LogReg, RF, SVM | Effective in handling high-dimensional data | 26 metabolites and lipids | HILIC-LC-MS, RP-LC-MS | Serum | Rheumatoid arthritis | N/A | Luan etal. (2021) [260] |
| RF | Highly conserved, predictive chemical feature subsets for taxonomic identity | 34 metabolites | LC-MS qTOF MS | Gut bacteria | Human gut microbiota | N/A | Han etal. (2021) [261] |
| LogReg | N/A | 5 purine metabolites | UPLC-Q/TOF MS, UHPLC-TQ MS/MS | Serum | Coronary artery disease | N/A | Jung etal. (2021) [262] |
| LDA, SVM, RF, LogReg | N/A | 58 metabolites | GC-MS, UHPLC-TQ MS/MS | Serum | Diabetic kidney disease | N/A | Liu etal. (2021) [263] |
| SVM, RF | Easily build non-linear classifiers; effectively reduce the high-variance | N/A | GC-MS, LC-MS | Plasma | Major depressive disorder | N/A | Liu etal. (2016) [264] |
| ANN, O2PLS, PCA, PLS | Nonlinear MSA | 11 compounds | HPLC-HRMS | Plant extracts | Anti-inflammation | N/A | Chagas-Paula et al. (2015) [265] |
| RF | Built-in feature selection, Variable importance scores are useful for biomarker prioritization | 22 serum proteins, 7 metabolites | UPLC-MS/MS | Serum | COVID-19 | https://www.iprox.org/(IPX0002106000andIPX0002171000) | Shen et al. (2020) [266] |
| Gradient Boosting, RF | N/A | Top 20 ion features | LC/Q-TOF, LC-MS/MS | Nasopharyngeal | Influenza | N/A | Hogan et al. (2021) [267] |
| AdaBoost, GDB, RF, XRF, PLS, SVM | N/A | 19 discriminant biomarkers | HESI-Q Exactive Orbitrat-MS | Plasma | COVID-19 | https://10.5281/zenodo.4329381 | Delafiori et al. (2021) [268] |
| RF | N/A | N/A | GC-MS | Pericarp tissue | Tomato | N/A | Toubiana et al. [191] |
| MCR-ALS, PCA, k-Means | N/A | N/A | GC-MS | Saffron | Saffron extract | N/A | Aliakbarzadeh et al. (2016) [269] |
| EDNN | superior to the conventional methods in regression, based on RMSEs | N/A | NMR | Fish muscle | Fish | N/A | Asakura et al. (2018) [270] |
AI was used to improve cancer diagnosis and treatment. Inglese et al. developed a novel 3D MS imaging technology, which is a powerful tool to investigate the chemical and biological interactions occurring in tumor tissues [224]. Inspection of the 3D data derived from DL and cluster analysis is a straightforward and robust approach to identify the presence of tumor cells, which is extremely difficult by visual inspection [224]. In addition, ML algorithms have been applied to analyze images to detect various types of cancer at an early stage [225]. Bebas et al. used classification methods, including SVM, KNN, RF, and DL, and the best results were achieved using the SVM classifier with the highest specificity and sensitivity [226]. Takats, Nicholson, and Glen groups made use of a deep unsupervised NN, specifically parametric t-SNE, and they projected a 3D-desorption electrospray ionization (DESI)-MS dataset from a human colorectal adenocarcinoma biopsy onto a 2-dimensional manifold, which enabled the identification of clusters that may not be discernible with linear methods [224]. Further, AI not only improves the accuracy of diagnosis but also allows for quicker and more effective treatments. Elemento group reviewed the recent progress in the applications of AI in oncology, which highlighted the potential of AI to be used in clinical settings [227].
AI is also very useful in clinical decision-making and management for other diseases, such as CVD and strokes [228, 229]. Diverse AI methodologies, including robotics, ML, and NLP, have been employed for the examination of CVD [228]. Chang group compared the strengths and limitations of both traditional and AI approaches in CVD biomarker discovery, and AI algorithms were found to have high potential in risk stratification and early disease detection [230]. In addition, a number of advanced ML models have been developed to detect biomarkers of coronary artery disease, coronary artery stenosis, myocardial ischemia, and stroke lesion outcomes [231–235]. Juarez-Orozco et al. used ML to identify predictors for recognizing patients who were likely to have myocardial ischemia and an elevated risk of Major Adverse Cardiovascular Events (MACE) [236].
AI, particularly ML, has been used for extracting dependable predictors and automatically classifying diverse dementia subtypes, such as mild cognitive impairment (MCI) that is likely to progress to AD [223]. Li group reviewed AI methods that have been successfully used for the diagnosis and prognosis of AD at the MCI or preclinical stages [237]. Battista et al. summarized ML methods that can automatically classify patients with AD, even in the early stages of the condition, and these AI approaches can identify an optimal combination of neuropsychological predictors [238].
5.2. Personalized Medicine & Drug Discovery
AI can assist in the development of personalized medicine by analyzing an individual’s metabolome [239]. Personalized medicine aims to optimize medical treatments for individual patients based on their unique characteristics, including their age, sex, genetics, lifestyle, and metabolism. AI-driven robotic metabolomics is able to analyze large and complex datasets, offering insights into complex gene-protein-metabolite interactions for targeted drug discovery and disease treatments [240]. Particularly, AI-driven robotic metabolomics automates labor-intensive tasks such as sample preparation and data acquisition, which increases data robustness and accuracy [240–242]. This approach is not only cost-effective, reducing the likelihood of ineffective treatments, but also empowers patients by involving them more directly in their healthcare decisions, providing specific health status information and potential risks [243]. Azad et al. suggested that ML/AI could be integrated with metabolomics for automated clinical diagnosis, treatment, and prognosis [244]. As a result, AI-assisted metabolomics can help generate personalized recommendations for medication, diet, and lifestyle modifications, which has the potential to significantly improve patient outcomes and reduce healthcare costs [245].
In addition, metabolomics can examine metabolic effects of potential drugs, based on which AI can predict how drugs interact with metabolic pathways and prioritize drug candidates for further investigation [246]. David et al. provided a guideline for the practice of AI in drug discovery [247]. AI-driven analysis has been used to determine the most effective drug dosages, safety, and efficacy. Martinelli reviewed the ML models, such as elastic net regression and RF, which have been used to study metabolism and provide invaluable insights to determine attrition rates and optimize drug efficacy [248]. Similarly, ML algorithms have been applied to expedite the process of identifying potential drug candidates and predicting their effectiveness. Tanoli et al. described the use of supervised ML models for three levels of predictions related to drug repurposing [249]. ML has also been used to predict the specific effects of drugs on cancer [250]. In a study conducted by McDermott and Garnett groups, the responses of 1,001 human cancer cell lines were tested on 265 anticancer drugs, establishing a comprehensive pharmaceutical landscape in cancer [251]. Varshney group showed that AI can address many disadvantages during the classical drug development, such as too many candidates and being slow in determining drug side effects, through reducing human intervention and bridging the translational gap between drugs and diseases [252].
6. Challenges and Future Perspectives
Although it is of increasing interest to use AI in metabolomics, there are challenges to analyze metabolomics data using AI approaches. As mentioned above, one of the main challenges is curse of dimensionality. Training a reliable AI model demands a substantial number of samples. In statistics, a common guideline suggests that around 20 samples are needed for each parameter to be estimated. AI models often comprise a considerable number of parameters. For instance, in a RF model with 100 decision trees each containing an average of 10 nodes, there are 1,000 parameters to be estimated (as the splitting cutoff at each node is a parameter). Training such a model would necessitate 20,000 samples. However, in many applications, the number of samples falls short of this criterion, or may even be smaller than the number of parameters, leading to a situation known as the curse of dimensionality. Under this curse, the model is inevitably prone to overfitting the data, capturing noise and intricacies specific to the training data but failing to generalize well to new, unseen data. To mitigate the curse of dimensionality, it is essential that the number of model parameters is kept lower than the number of samples. In cases where increasing the sample size is not feasible, another approach is to reduce the complexity of the model. This can be achieved by, for example, decreasing the number of trees in a RF model or reducing the number of neurons in an ANN model. In addition, it is also recommended to reduce the number of input features. To further combat overfitting, robust validation methods and diverse, well-annotated datasets are essential, ensuring that AI models are trained to generalize effectively. Stringent internal cross validation should be embedded in AI approaches, and when possible, external multi-center cross validation using a large sample set should be performed to validate the significant metabolites established by AI.
The complexity of AI models can lead to interpretation challenges, making it difficult to understand the rationale behind specific predictions or recommendations. Developing more transparent AI models can help understand how decisions are made, fostering trust and practical utility in clinical settings [271]. Another significant challenge is the requirement for extensive computational resources and expertise in AI, which may not be readily available in all healthcare settings. Improving access to computational resources and fostering interdisciplinary collaborations among AI experts, bioinformaticians, and clinicians can facilitate the adoption of AI in diverse healthcare environments. In addition, the dynamic nature of AI models, while beneficial in rapidly adapting to new data, can also lead to model drift over time, necessitating continuous monitoring and updating. This weakness highlights the need for a balanced approach in integrating AI into metabolomics to ensure that the benefits of advanced data analysis are harmonized with practical applications. Regularly updating AI models and employing adaptive learning algorithms can help prevent model drift, maintaining the relevance and accuracy of the analysis over time.
Further, the transition of AI-assisted metabolomics to routine clinical practices faces hurdles, including technological and scientific challenges, practicality, regulatory compliance, and integration into existing healthcare frameworks. The resulted AI models need to have friendly interfaces for end-users such as clinicians, ideally as plug-ins in electronic health record software that can be directly used in clinical decision-making. Data security and patient privacy should also be protected, similar to other omics data. By tackling these challenges, AI-assisted metabolomics can be more effectively used to enhance patient care and treatment outcomes.
7. Conclusions
In this review, we reviewed recent applications of AI in the field of omics, with a special focus on metabolomics. Metabolomics studies the metabolome in systems biology, and it produces large datasets examining hundreds to thousands of metabolites. AI was designed to work with large datasets, and it has the potential to bring “human intelligence” to metabolomics, but with computers and software packages. Recently, various AI algorithms have been developed and used in metabolomics to improve data collection, data extraction, and statistical modeling. AI algorithms can efficiently and accurately identify metabolites, detect changes in their concentrations, and determine their biological significance. Based on previous successful applications, it is expected that AI-assisted metabolomics has the potential to revolutionize disease prevention, diagnosis, and treatment.
Highlights.
We provide a recent overview of the methodologies and applications of artificial intelligence (AI) in metabolomics.
We introduce metabolomics and AI algorithms that are commonly used in big data analyses.
Recent applications of AI-assisted metabolomics in the context of systems biology and human health are summarized.
Challenges and future perspectives for the development and applications of AI approaches in metabolomics are discussed.
Acknowledgments
This work was supported by the NIH 1R01ES030197, R01LM013438, and T32DK137525.
Abbreviations:
- AI
artificial intelligence
- ML
machine learning
- DL
deep learning
- MS
mass spectrometry
- NMR
nuclear magnetic resonance
- LC
liquid chromatography
- GC
gas chromatography
- CNNs
convolutional neural networks
- RNNs
recurrent neural networks
- GANs
generative adversarial networks
- NLP
natural language processing
- HRMS
high-resolution mass spectrometry
- QSRR
quantitative structure–retention relationship
- NN
neural networks
- ANN
artificial neural networks
- PCA
principal component analysis
- ICA
independent component analysis
- UMAP
uniform manifold approximation and projection
- t-SNE
t-distributed stochastic neighbor embedding
- SVM
support vector machines
- SVR
support vector regression
- RF
random forest
- SOM
self-organizing maps
- KNN
k-nearest neighbors
- DQN
deep q-networks
- PPO
proximal policy optimization
- DDPG
deep deterministic policy gradients
- TRPO
trust region policy optimization
- SMILES
simplified molecular input line entry system
- MFA
metabolic flux analysis
- MLR
multiple linear regression
- PLS
partial least square
- AD
Alzheimer’s Disease
- MCI
mild cognitive impairment
- PLS-DA
partial least squares regression–discriminant analysis
- LogReg
logistic regression
- ADTree
alternative decision tree
- MSA
multivariate statistical analysis
- XRF
extreme random forest
- GDB
gradient tree boosting
- ADA
ADA tree boosting
- EDNN
ensemble deep neural network
- VAEs
variational autoencoders
- ANOVA
analysis of variance
- REKINDLE
reconstruction of kinetic models using deep learning
- ChatGPT
chat generative pre-trained transformer
- SAM
S-adenosylmethionine
- PND
perioperative neurocognitive disorder
- MIBC
muscle-invasive bladder cancer
- MOSAE
multi-omics supervised autoencoder
- WNN
weighted-nearest neighbor
- fMKL-DR
fast multiple kernel learning for dimensionality reduction
- DESI
desorption electrospray ionization
- CVD
cardiovascular disease
- MACE
major adverse cardiovascular events
- RT
retention time
- LIME
local interpretable model-agnostic explanations
- SHAP
SHapley Additive exPlanations
- AF2
AlphaFold2
- ED
erectile dysfunction
- SMOTE
synthetic minority over-sampling technique
- OS
over-sampling
- US
under-sampling
- SERRF
systematic error removal using random forest
- LASSO
least absolute shrinkage and selection operator
- MUVR
multivariate methods with unbiased variable selection in R
- FCBF
fast correlation-based filter
- AdaBoost
adaptive boosting
- RFE
recursive feature elimination
- MCR-ALS
multivariate curve resolution-alternating least squares
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
The authors declare no competing conflicts of interest.
Declaration of interests
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- 1.Muthubharathi BC, Gowripriya T, and Balamurugan K, Metabolomics: Small molecules that matter more. Molecular omics, 2021. 17(2): p. 210–229. [DOI] [PubMed] [Google Scholar]
- 2.Gu H, et al. , Principal component directed partial least squares analysis for combining NMR and MS data in metabolomics: application to the detection of breast cancer. Analytica chimica acta, 2011. 686(1–2): p. 57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Putri SP, et al. , Current metabolomics: practical applications. Journal of bioscience and bioengineering, 2013. 115(6): p. 579–589. [DOI] [PubMed] [Google Scholar]
- 4.Yang PL, Metabolomics and Lipidomics: Yet more ways your health is influenced by fat, in Viral pathogenesis. 2016, Elsevier. p. 181–198. [Google Scholar]
- 5.Chandel NS, Basics of metabolic reactions. Cold Spring Harbor Perspectives in Biology, 2021. 13(8). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Vermeulen N, ROLE OF METABOLISM IN. Cytochromes P450: metabolic and toxicological aspects, 1996: p. 29. [Google Scholar]
- 7.Schilling CH, et al. , Metabolic pathway analysis: basic concepts and scientific applications in the post-genomic era. Biotechnology progress, 1999. 15(3): p. 296–303. [DOI] [PubMed] [Google Scholar]
- 8.Johnson CH, et al. , Bioinformatics: the next frontier of metabolomics. Analytical Chemistry, 2015. 87(1): p. 147–156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wishart DS, Metabolomics for investigating physiological and pathophysiological processes. Physiological reviews, 2019. 99(4): p. 1819–1875. [DOI] [PubMed] [Google Scholar]
- 10.Zhang A, et al. , Modern analytical techniques in metabolomics analysis. Analyst, 2012. 137(2): p. 293–300. [DOI] [PubMed] [Google Scholar]
- 11.Qiu S, et al. , Small molecule metabolites: discovery of biomarkers and therapeutic targets. Signal Transduct Target Ther, 2023. 8(1): p. 132. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Pearce EL, et al. , Fueling immunity: insights into metabolism and lymphocyte function. Science, 2013. 342(6155): p. 1242454.24115444 [Google Scholar]
- 13.Zhang J, et al. , Metabolomics study of esophageal adenocarcinoma. The Journal of thoracic and cardiovascular surgery, 2011. 141(2): p. 469–475. e4. [DOI] [PubMed] [Google Scholar]
- 14.Gibney MJ, et al. , Metabolomics in human nutrition: opportunities and challenges–. The American journal of clinical nutrition, 2005. 82(3): p. 497–503. [DOI] [PubMed] [Google Scholar]
- 15.Gowda GN, et al. , Metabolomics-based methods for early disease diagnostics. Expert review of molecular diagnostics, 2008. 8(5): p. 617–633. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Monteiro M, et al. , Metabolomics analysis for biomarker discovery: advances and challenges. Current medicinal chemistry, 2013. 20(2): p. 257–271. [DOI] [PubMed] [Google Scholar]
- 17.Xu R, et al. , Integrated models of blood protein and metabolite enhance the diagnostic accuracy for Non-Small Cell Lung Cancer. Biomarker Research, 2023. 11(1): p. 71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Primrose S, et al. , Metabolomics and human nutrition. British Journal of Nutrition, 2011. 105(8): p. 1277–1283. [DOI] [PubMed] [Google Scholar]
- 19.Armitage EG and Ciborowski M, Applications of metabolomics in cancer studies. Metabolomics: From fundamentals to clinical applications, 2017: p. 209–234. [DOI] [PubMed] [Google Scholar]
- 20.Xie Y, et al. , Early lung cancer diagnostic biomarker discovery by machine learning methods. Translational oncology, 2021. 14(1): p. 100907. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Botas A, et al. , Metabolomics of neurodegenerative diseases. International review of neurobiology, 2015. 122: p. 53–80. [DOI] [PubMed] [Google Scholar]
- 22.Kell DB and Goodacre R, Metabolomics and systems pharmacology: why and how to model the human metabolic network for drug discovery. Drug Discovery Today, 2014. 19(2): p. 171–182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Robertson DG, Watkins PB, and Reily MD, Metabolomics in toxicology: preclinical and clinical applications. Toxicological sciences, 2011. 120(suppl_1): p. S146–S170. [DOI] [PubMed] [Google Scholar]
- 24.Robertson D and Frevert U, Metabolomics in drug discovery and development. Clinical Pharmacology & Therapeutics, 2013. 94(5): p. 559–561. [DOI] [PubMed] [Google Scholar]
- 25.Chen MX, et al. , Metabolome analysis for investigating host-gut microbiota interactions. Journal of the Formosan Medical Association, 2019. 118: p. S10–S22. [DOI] [PubMed] [Google Scholar]
- 26.Jin Y, et al. , Recent Review on Selected Xenobiotics and Their Impacts on Gut Microbiome and Metabolome. Trends in analytical chemistry: TRAC, 2023. 166: p. 117155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Ahmed M, et al. , Preclinical and Clinical Applications of Metabolomics and Proteomics in Glioblastoma Research. International Journal of Molecular Sciences, 2022. 24(1): p. 348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Saito K and Matsuda F, Metabolomics for functional genomics, systems biology, and biotechnology. Annual review of plant biology, 2010. 61: p. 463–489. [DOI] [PubMed] [Google Scholar]
- 29.Pinu FR, et al. , Systems biology and multi-omics integration: viewpoints from the metabolomics research community. Metabolites, 2019. 9(4): p. 76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Wang P, On defining artificial intelligence. Journal of Artificial General Intelligence, 2019. 10(2): p. 1–37. [Google Scholar]
- 31.Lucci S, Kopec D, and Musa SM, Artificial intelligence in the 21st century. 2022: Mercury learning and information. [Google Scholar]
- 32.Tien JM, Internet of things, real-time decision making, and artificial intelligence. Annals of Data Science, 2017. 4: p. 149–178. [Google Scholar]
- 33.Barnes EA, et al. , Indicator patterns of forced change learned by an artificial neural network. Journal of Advances in Modeling Earth Systems, 2020. 12(9): p. e2020MS002195. [Google Scholar]
- 34.Li RC, Asch SM, and Shah NH, Developing a delivery science for artificial intelligence in healthcare. NPJ digital medicine, 2020. 3(1): p. 107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Fisher PF, et al. , Artificial intelligence and expert systems in geodata processing. Progress in physical geography, 1988. 12(3): p. 371–388. [Google Scholar]
- 36.Greener JG, et al. , A guide to machine learning for biologists. Nat Rev Mol Cell Biol, 2022. 23(1): p. 40–55. [DOI] [PubMed] [Google Scholar]
- 37.Alajaji SA, et al. , Generative Adversarial Networks in Digital Histopathology: Current Applications, Limitations, Ethical Considerations, and Future Directions. Mod Pathol, 2023: p. 100369. [DOI] [PubMed] [Google Scholar]
- 38.Dunn WB and Ellis DI, Metabolomics: current analytical platforms and methodologies. TrAC Trends in Analytical Chemistry, 2005. 24(4): p. 285–294. [Google Scholar]
- 39.Spratlin JL, Serkova NJ, and Eckhardt SG, Clinical applications of metabolomics in oncology: a review. Clinical cancer research, 2009. 15(2): p. 431–440. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Wishart DS, Applications of metabolomics in drug discovery and development. Drugs in R & D, 2008. 9: p. 307–322. [DOI] [PubMed] [Google Scholar]
- 41.Zeng L, et al. , Comprehensive scRNA-seq Model Reveals Artery Endothelial Cell Heterogeneity and Metabolic Preference in Human Vascular Disease. Interdisciplinary Sciences: Computational Life Sciences, 2023: p. 1–19. [DOI] [PubMed] [Google Scholar]
- 42.Liu J, Li B, and Xiong J-H, MDAS: An integrated system for metabonomic data analysis. Interdisciplinary Sciences: Computational Life Sciences, 2009. 1: p. 61–71. [DOI] [PubMed] [Google Scholar]
- 43.Bhargava C and Sharma PK, Artificial intelligence: fundamentals and applications. 2021: CRC Press. [Google Scholar]
- 44.Agah A, Medical applications of artificial intelligence. 2013: CRC Press. [Google Scholar]
- 45.Pannu A, Artificial intelligence and its application in different areas. Artificial Intelligence, 2015. 4(10): p. 79–84. [Google Scholar]
- 46.Das S, et al. , Applications of artificial intelligence in machine learning: review and prospect. International Journal of Computer Applications, 2015. 115(9). [Google Scholar]
- 47.Górriz JM, et al. , Artificial intelligence within the interplay between natural and artificial computation: Advances in data science, trends and applications. Neurocomputing, 2020. 410: p. 237–270. [Google Scholar]
- 48.Hai HN, ChatGPT: The Evolution of Natural Language Processing. Authorea Preprints, 2023. [Google Scholar]
- 49.Ray PP, ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 2023. [Google Scholar]
- 50.Kalla D and Smith N, Study and Analysis of Chat GPT and its Impact on Different Fields of Study. International Journal of Innovative Science and Research Technology, 2023. 8(3). [Google Scholar]
- 51.Ertel W, Introduction to artificial intelligence. 2018: Springer. [Google Scholar]
- 52.Newell A, Intellectual issues in the history of artificial intelligence. Artificial Intelligence: Critical Concepts, 1982: p. 25–70. [Google Scholar]
- 53.Buchanan BG, et al. , Applications of artificial intelligence for chemical inference. 22. Automatic rule formation in mass spectrometry by means of the meta-DENDRAL program. Journal of the American Chemical Society, 1976. 98(20): p. 6168–6178. [Google Scholar]
- 54.Feigenbaum EA, Expert systems in the 1980s. State of the art report on machine intelligence. Maidenhead: Pergamon-Infotech, 1981. 23. [Google Scholar]
- 55.El-Najdawi M and Stylianou AC, Expert support systems: integrating AI technologies. Communications of the ACM, 1993. 36(12): p. 55–ff. [Google Scholar]
- 56.Duda RO and Shortliffe EH, Expert systems research. Science, 1983. 220(4594): p. 261–268. [DOI] [PubMed] [Google Scholar]
- 57.Libbrecht MW and Noble WS, Machine learning applications in genetics and genomics. Nature Reviews Genetics, 2015. 16(6): p. 321–332. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Leung MK, et al. , Machine learning in genomic medicine: a review of computational problems and data sets. Proceedings of the IEEE, 2015. 104(1): p. 176–197. [Google Scholar]
- 59.Watson DS, Interpretable machine learning for genomics. Human genetics, 2022. 141(9): p. 1499–1513. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Cartwright HM, Machine Learning in Chemistry. 2020: Royal Society of Chemistry. [Google Scholar]
- 61.Meher SK and Panda G, Deep learning in astronomy: a tutorial perspective. The European Physical Journal Special Topics, 2021. 230: p. 2285–2317. [Google Scholar]
- 62.Shen D, Wu G, and Suk H-I, Deep learning in medical image analysis. Annual review of biomedical engineering, 2017. 19: p. 221–248. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Camps-Valls G, et al. , Deep learning for the Earth Sciences: A comprehensive approach to remote sensing, climate science and geosciences. 2021: John Wiley & Sons. [Google Scholar]
- 64.Kurth T, et al. Exascale deep learning for climate analytics. in SC18: International conference for high performance computing, networking, storage and analysis. 2018. IEEE. [Google Scholar]
- 65.Aguzzi J, et al. , Developing technological synergies between deep-sea and space research. Elem Sci Anth, 2022. 10(1): p. 00064. [Google Scholar]
- 66.Dar YL, High-throughput experimentation: A powerful enabling technology for the chemicals and materials industry. Macromolecular rapid communications, 2004. 25(1): p. 34–47. [Google Scholar]
- 67.Eyke NS, Koscher BA, and Jensen KF, Toward machine learning-enhanced high-throughput experimentation. Trends in Chemistry, 2021. 3(2): p. 120–132. [Google Scholar]
- 68.Gupta R, et al. , Artificial intelligence to deep learning: machine intelligence approach for drug discovery. Molecular diversity, 2021. 25: p. 1315–1360. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Wang H, et al. , Scientific discovery in the age of artificial intelligence. Nature, 2023. 620(7972): p. 47–60. [DOI] [PubMed] [Google Scholar]
- 70.Petersen EV, et al. , The extracellular matrix-derived biomarkers for diagnosis, prognosis, and personalized therapy of malignant tumors. Frontiers in Oncology, 2020. 10: p. 575569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Lai K, et al. , Artificial intelligence and machine learning in bioinformatics. Encyclopedia of Bioinformatics and Computational Biology: ABC of Bioinformatics, 2018. 1(3). [Google Scholar]
- 72.Xu Y, et al. , Artificial intelligence: A powerful paradigm for scientific research. Innovation (Camb), 2021. 2(4): p. 100179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Raschka S, Patterson J, and Nolet C, Machine learning in python: Main developments and technology trends in data science, machine learning, and artificial intelligence. Information, 2020. 11(4): p. 193. [Google Scholar]
- 74.Kotsiantis SB, Zaharakis I, and Pintelas P, Supervised machine learning: A review of classification techniques. Emerging artificial intelligence applications in computer engineering, 2007. 160(1): p. 3–24. [Google Scholar]
- 75.Cunningham P, Cord M, and Delany SJ, Supervised learning, in Machine learning techniques for multimedia: case studies on organization and retrieval. 2008, Springer. p. 21–49. [Google Scholar]
- 76.Dias-Audibert FL, et al. , Combining machine learning and metabolomics to identify weight gain biomarkers. Frontiers in bioengineering and biotechnology, 2020. 8: p. 6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Tiedt S, et al. , Circulating metabolites differentiate acute ischemic stroke from stroke mimics. Annals of neurology, 2020. 88(4): p. 736–746. [DOI] [PubMed] [Google Scholar]
- 78.Sen P, et al. , Deep learning meets metabolomics: a methodological perspective. Brief Bioinform, 2021. 22(2): p. 1531–1542. [DOI] [PubMed] [Google Scholar]
- 79.Dayan P, Sahani M, and Deback G, Unsupervised learning. The MIT encyclopedia of the cognitive sciences, 1999: p. 857–859. [Google Scholar]
- 80.Han J, Kamber M , Data Mining: Concepts and Techniques. 2nd ed. 2006. [Google Scholar]
- 81.Hurtik P, Molek V, and Perfilieva I, Novel dimensionality reduction approach for unsupervised learning on small datasets. Pattern Recognition, 2020. 103: p. 107291. [Google Scholar]
- 82.Lin Y, et al. , The individual identification method of wireless device based on dimensionality reduction and machine learning. The journal of supercomputing, 2019. 75(6): p. 3010–3027. [Google Scholar]
- 83.Meinicke P, et al. , Metabolite-based clustering and visualization of mass spectrometry data using one-dimensional self-organizing maps. Algorithms for Molecular Biology, 2008. 3: p. 1–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Goodwin CR, et al. , Phenotypic mapping of metabolic profiles using self-organizing maps of high-dimensional mass spectrometry data. Analytical chemistry, 2014. 86(13): p. 6563–6571. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Ceusters N, et al. , Hierarchical clustering reveals unique features in the diel dynamics of metabolites in the CAM orchid Phalaenopsis. Journal of experimental botany, 2019. 70(12): p. 3269–3281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Rawlinson C, et al. , Hierarchical clustering of MS/MS spectra from the firefly metabolome identifies new lucibufagin compounds. Scientific Reports, 2020. 10(1): p. 6043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Chen X, et al. , Robust principal component analysis for accurate outlier sample detection in RNA-Seq data. BMC bioinformatics, 2020. 21: p. 1–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Li H, Liao X, and Carin L, Multi-task Reinforcement Learning in Partially Observable Stochastic Environments. Journal of Machine Learning Research, 2009. 10(5). [Google Scholar]
- 89.Moerland TM, Broekens DJ, Plaat A, Jonker CM Model-based Reinforcement Learning: A Survey. 2023. DOI: 10.1561/2200000086. [DOI] [Google Scholar]
- 90.Neftci EO and Averbeck BB, Reinforcement learning in artificial and biological systems. Nature Machine Intelligence, 2019. 1(3): p. 133–143. [Google Scholar]
- 91.Chen X, et al. , Reinforcement learning for selective key applications in power systems: Recent advances and future challenges. IEEE Transactions on Smart Grid, 2022. 13(4): p. 2935–2958. [Google Scholar]
- 92.Hu C, et al. REMEDI: REinforcement learning-driven adaptive MEtabolism modeling of primary sclerosing cholangitis DIsease progression. in Machine Learning for Health (ML4H). 2023. PMLR. [Google Scholar]
- 93.Kim H, Deep learning, in Artificial Intelligence for 6G. 2022, Springer. p. 247–303. [Google Scholar]
- 94.Zou J, Han Y, and So S-S, Overview of artificial neural networks. Artificial neural networks: methods and applications, 2009: p. 14–22. [Google Scholar]
- 95.Montesinos López OA, Montesinos López A, and Crossa J, Fundamentals of artificial neural networks and deep learning, in Multivariate statistical machine learning methods for genomic prediction. 2022, Springer. p. 379–425. [PubMed] [Google Scholar]
- 96.Avsec Ž, et al. , Effective gene expression prediction from sequence by integrating long-range interactions. Nature methods, 2021. 18(10): p. 1196–1203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Chandrashekar PB, et al. , DeepCORE: An interpretable multi-view deep neural network model to detect co-operative regulatory elements. Computational and Structural Biotechnology Journal, 2024. 23: p. 679–687. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Askr H, et al. , Deep learning in drug discovery: an integrative review and future challenges. Artificial Intelligence Review, 2023. 56(7): p. 5975–6037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Li Z, et al. , A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems, 2021. [DOI] [PubMed] [Google Scholar]
- 100.Kim HW, et al. , SMART-Miner: a convolutional neural network-based metabolite identification from 1H-13C HSQC spectra. Magnetic Resonance in Chemistry, 2022. 60(11): p. 1070–1075. [DOI] [PubMed] [Google Scholar]
- 101.Medsker LR and Jain L, Recurrent neural networks. Design and Applications, 2001. 5(64–67): p. 2. [Google Scholar]
- 102.Li M and Wang XR, Peak alignment of gas chromatography–mass spectrometry data with deep learning. Journal of Chromatography A, 2019. 1604: p. 460476. [DOI] [PubMed] [Google Scholar]
- 103.Choudhury S, et al. , Reconstructing kinetic models for dynamical studies of metabolism using generative adversarial networks. Nature Machine Intelligence, 2022. 4(8): p. 710–719. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Creswell A, et al. , Generative adversarial networks: An overview. IEEE signal processing magazine, 2018. 35(1): p. 53–65. [Google Scholar]
- 105.Niu Z, Zhong G, and Yu H, A review on the attention mechanism of deep learning. Neurocomputing, 2021. 452: p. 48–62. [Google Scholar]
- 106.Ekman M, Learning Deep Learning: Theory and Practice of Neural Networks, Computer Vision, Natural Language Processing, and Transformers Using TensorFlow. 2021: Addison-Wesley Professional. [Google Scholar]
- 107.Alharthi AG and Alzahrani SM, Do it the transformer way: A comprehensive review of brain and vision transformers for autism spectrum disorder diagnosis and classification. Comput Biol Med, 2023. 167: p. 107667. [DOI] [PubMed] [Google Scholar]
- 108.Schramowski P, et al. , Large pre-trained language models contain human-like biases of what is right and wrong to do. Nature Machine Intelligence, 2022. 4(3): p. 258–268. [Google Scholar]
- 109.Shrivastava AD, et al. , MassGenie: A transformer-based deep learning method for identifying small molecules from their mass spectra. Biomolecules, 2021. 11(12): p. 1793. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Odenkirk MT, Reif DM, and Baker ES, Multiomic big data analysis challenges: increasing confidence in the interpretation of artificial intelligence assessments. Analytical Chemistry, 2021. 93(22): p. 7763–7773. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Liebal UW, et al. , Machine learning applications for mass spectrometry-based metabolomics. Metabolites, 2020. 10(6): p. 243. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Galal A, Talal M, and Moustafa A, Applications of machine learning in metabolomics: Disease modeling and classification. Frontiers in genetics, 2022. 13: p. 1017340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Liao J, et al. , Artificial intelligence assists precision medicine in cancer treatment. Frontiers in Oncology, 2023. 12: p. 998222. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Yang X, et al. , Concepts of artificial intelligence for computer-assisted drug discovery. Chemical reviews, 2019. 119(18): p. 10520–10594. [DOI] [PubMed] [Google Scholar]
- 115.Dettmer K, Aronov PA, and Hammock BD, Mass spectrometry-based metabolomics. Mass spectrometry reviews, 2007. 26(1): p. 51–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Wishart D, Metabolomics: the principles and potential applications to transplantation. American journal of transplantation, 2005. 5(12): p. 2814–2820. [DOI] [PubMed] [Google Scholar]
- 117.Want EJ, et al. , From exogenous to endogenous: the inevitable imprint of mass spectrometry in metabolomics. Journal of proteome research, 2007. 6(2): p. 459–468. [DOI] [PubMed] [Google Scholar]
- 118.Halket JM, et al. , Chemical derivatization and mass spectral libraries in metabolic profiling by GC/MS and LC/MS/MS. Journal of experimental botany, 2005. 56(410): p. 219–243. [DOI] [PubMed] [Google Scholar]
- 119.Ribbenstedt A, Ziarrusta H, and Benskin JP, Development, characterization and comparisons of targeted and non-targeted metabolomics methods. PLoS One, 2018. 13(11): p. e0207082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Chen Y, Li E-M, and Xu L-Y, Guide to metabolomics analysis: a bioinformatics workflow. Metabolites, 2022. 12(4): p. 357. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Masutin V, Kersch C, and Schmitz-Spanke S, A systematic review: metabolomics-based identification of altered metabolites and pathways in the skin caused by internal and external factors. Experimental Dermatology, 2022. 31(5): p. 700–714. [DOI] [PubMed] [Google Scholar]
- 122.Petrick LM and Shomron N, AI/ML-driven advances in untargeted metabolomics and exposomics for biomedical applications. Cell Reports Physical Science, 2022. 3(7). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Giese SH, et al. , Retention time prediction using neural networks increases identifications in crosslinking mass spectrometry. Nature communications, 2021. 12(1): p. 3237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Jang I, et al. , LC–MS/MS software for screening unknown erectile dysfunction drugs and analogues: Artificial neural network classification, peak-count scoring, simple similarity search, and hybrid similarity search algorithms. Analytical chemistry, 2019. 91(14): p. 9119–9128. [DOI] [PubMed] [Google Scholar]
- 125.Lai A, et al. , An algorithm to classify homologous series within compound datasets. Journal of Cheminformatics, 2022. 14(1): p. 1–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126.Turi KN, et al. , A review of metabolomics approaches and their application in identifying causal pathways of childhood asthma. Journal of Allergy and Clinical Immunology, 2018. 141(4): p. 1191–1201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Yang Z, et al. , AlphaFold2 and its applications in the fields of biology and medicine. Signal Transduction and Targeted Therapy, 2023. 8(1): p. 115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Frusciante L, et al. , Artificial Intelligence Approaches in Drug Discovery: Towards the Laboratory of the Future. Current Topics in Medicinal Chemistry, 2022. 22(26): p. 2176–2189. [DOI] [PubMed] [Google Scholar]
- 129.Zhang X, et al. , DeepSpectra: An end-to-end deep learning approach for quantitative spectral analysis. Analytica chimica acta, 2019. 1058: p. 48–57. [DOI] [PubMed] [Google Scholar]
- 130.Boccard J, Veuthey JL, and Rudaz S, Knowledge discovery in metabolomics: an overview of MS data handling. Journal of separation science, 2010. 33(3): p. 290–304. [DOI] [PubMed] [Google Scholar]
- 131.Su K-M, et al. , Intelligent geochemical interpretation of mass chromatograms: Based on convolution neural network. Petroleum Science, 2023. [Google Scholar]
- 132.Acquarelli J, et al. , Convolutional neural networks for vibrational spectroscopic data analysis. Analytica chimica acta, 2017. 954: p. 22–31. [DOI] [PubMed] [Google Scholar]
- 133.Bjerrum EJ, Glahder M, and Skov T, Data augmentation of spectral data for convolutional neural network (CNN) based deep chemometrics. arXiv preprint arXiv:1710.01927, 2017. [Google Scholar]
- 134.Kensert A, et al. , Convolutional neural network for automated peak detection in reversed-phase liquid chromatography. Journal of Chromatography A, 2022. 1672: p. 463005. [DOI] [PubMed] [Google Scholar]
- 135.Hervás C, et al. , Improving the quantification of highly overlapping chromatographic peaks by using product unit neural networks modeled by an evolutionary algorithm. Journal of chemical information and modeling, 2005. 45(4): p. 894–903. [DOI] [PubMed] [Google Scholar]
- 136.Zeki ÖC, et al. , Integration of GC–MS and LC–MS for untargeted metabolomics profiling. Journal of Pharmaceutical and Biomedical Analysis, 2020. 190: p. 113509. [DOI] [PubMed] [Google Scholar]
- 137.Choi E, et al. , Machine learning liquid chromatography retention time prediction model augments the dansylation strategy for metabolite analysis of urine samples. Journal of Chromatography A, 2023. 1705: p. 464167. [DOI] [PubMed] [Google Scholar]
- 138.Bonini P, et al. , Retip: retention time prediction for compound annotation in untargeted metabolomics. Analytical chemistry, 2020. 92(11): p. 7515–7522. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 139.Wolfer AM, et al. , UPLC–MS retention time prediction: a machine learning approach to metabolite identification in untargeted profiling. Metabolomics, 2016. 12(1): p. 8. [Google Scholar]
- 140.Creek DJ, et al. , Toward global metabolomics analysis with hydrophilic interaction liquid chromatography–mass spectrometry: improved metabolite identification by retention time prediction. Analytical chemistry, 2011. 83(22): p. 8703–8710. [DOI] [PubMed] [Google Scholar]
- 141.Berry MW, Mohamed A, and Yap BW, Supervised and unsupervised learning for data science. 2019: Springer. [Google Scholar]
- 142.Liapikos T, et al. , Quantitative structure retention relationship (QSRR) modelling for Analytes’ retention prediction in LC-HRMS by applying different Machine Learning algorithms and evaluating their performance. Journal of Chromatography B, 2022. 1191: p. 123132. [DOI] [PubMed] [Google Scholar]
- 143.Bouwmeester R, Martens L, and Degroeve S, Comprehensive and empirical evaluation of machine learning algorithms for small molecule LC retention time prediction. Analytical chemistry, 2019. 91(5): p. 3694–3703. [DOI] [PubMed] [Google Scholar]
- 144.Kehoe ER, et al. , Biomarker selection and a prospective metabolite-based machine learning diagnostic for lyme disease. Sci Rep, 2022. 12(1): p. 1478. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 145.Shen X, et al. , Normalization and integration of large-scale metabolomics data using support vector regression. Metabolomics, 2016. 12: p. 1–12. [Google Scholar]
- 146.Fan S, et al. , Systematic Error Removal Using Random Forest for Normalizing Large-Scale Untargeted Lipidomics Data. Anal Chem, 2019. 91(5): p. 3590–3596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 147.Maw M, Haw SC, and Ho CK, Utilizing data sampling techniques on algorithmic fairness for customer churn prediction with data imbalance problems. F1000Res, 2021. 10: p. 988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 148.Khuvis J, et al. , The impact of diagnostic stewardship interventions on Clostridiodes difficile test ordering practices and results. Clin Biochem, 2023. 117: p. 23–29. [DOI] [PubMed] [Google Scholar]
- 149.Santos-Perez MI, et al. , A cross-sectional study of psychotropic drug use in the elderly: Consuming patterns, risk factors and potentially inappropriate use. Eur J Hosp Pharm, 2021. 28(2): p. 88–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 150.Gui C, Analysis of imbalanced data set problem: The case of churn prediction for telecommunication. Artif. Intell. Res, 2017. 6(2): p. 93-. [Google Scholar]
- 151.Biswas N and Chakrabarti S, Artificial intelligence (AI)-based systems biology approaches in multi-omics data analysis of cancer. Frontiers in Oncology, 2020. 10: p. 588221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 152.Helmy M, Smith D, and Selvarajoo K, Systems biology approaches integrated with artificial intelligence for optimized metabolic engineering. Metabolic Engineering Communications, 2020. 11: p. e00149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 153.Edison AS, et al. , NMR: unique strengths that enhance modern metabolomics research. Analytical chemistry, 2020. 93(1): p. 478–499. [DOI] [PubMed] [Google Scholar]
- 154.Berisha V, et al. , Digital medicine and the curse of dimensionality. NPJ Digit Med, 2021. 4(1): p. 153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 155.Cuperlovic-Culf M, Machine learning methods for analysis of metabolic data and metabolic pathway modeling. Metabolites, 2018. 8(1): p. 4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 156.Miyazawa T, et al. , Artificial intelligence in food science and nutrition: a narrative review. Nutrition Reviews, 2022. 80(12): p. 2288–2300. [DOI] [PubMed] [Google Scholar]
- 157.Grissa D, et al. , Feature selection methods for early predictive biomarker discovery using untargeted metabolomic data. Frontiers in molecular biosciences, 2016. 3: p. 30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 158.Barberis E, et al. , Precision Medicine Approaches with Metabolomics and Artificial Intelligence. Int J Mol Sci, 2022. 23(19). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 159.He Z, Liu Z, and Gong L, Biomarker identification and pathway analysis of rheumatoid arthritis based on metabolomics in combination with ingenuity pathway analysis. Proteomics, 2021. 21(11–12): p. 2100037. [DOI] [PubMed] [Google Scholar]
- 160.Mehmood T, Sæbø S, and Liland KH, Comparison of variable selection methods in partial least squares regression. Journal of Chemometrics, 2020. 34(6): p. e3226. [Google Scholar]
- 161.Gaul DA, et al. , Highly-accurate metabolomic detection of early-stage ovarian cancer. Scientific reports, 2015. 5(1): p. 16351. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 162.Vidyasagar M, Identifying predictive features in drug response using machine learning: opportunities and challenges. Annu Rev Pharmacol Toxicol, 2015. 55: p. 15–34. [DOI] [PubMed] [Google Scholar]
- 163.Santoro AL, et al. , In situ DESI-MSI lipidomic profiles of breast cancer molecular subtypes and precursor lesions. Cancer research, 2020. 80(6): p. 1246–1257. [DOI] [PubMed] [Google Scholar]
- 164.Bifarin OO, et al. , Machine learning-enabled renal cell carcinoma status prediction using multiplatform urine-based metabolomics. Journal of proteome research, 2021. 20(7): p. 3629–3641. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 165.Zampieri G, et al. , Machine and deep learning meet genome-scale metabolic modeling. PLoS computational biology, 2019. 15(7): p. e1007084. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 166.De Fabritiis S, et al. , Targeted metabolomics detects a putatively diagnostic signature in plasma and dried blood spots from head and neck paraganglioma patients. Oncogenesis, 2023. 12(1): p. 10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 167.Chen N, et al. , Using random forest to detect multiple inherited metabolic diseases simultaneously based on GC-MS urinary metabolomics. Talanta, 2021. 235: p. 122720. [DOI] [PubMed] [Google Scholar]
- 168.Kang Y, Vijay S, and Gujral TS, Deep neural network modeling identifies biomarkers of response to immune-checkpoint therapy. Iscience, 2022. 25(5). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 169.Wang H, et al. , Machine learning of plasma metabolome identifies biomarker panels for metabolic syndrome: findings from the China Suboptimal Health Cohort. Cardiovascular Diabetology, 2022. 21(1): p. 288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 170.Kitano H, Systems biology: toward system-level understanding of biological systems. Foundations of systems biology, 2001: p. 1–36. [Google Scholar]
- 171.Tong W, Analyzing the biology on the system level. Genomics, Proteomics & Bioinformatics, 2004. 2(1): p. 6–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 172.Kitano H, Systems biology: a brief overview. science, 2002. 295(5560): p. 1662–1664. [DOI] [PubMed] [Google Scholar]
- 173.Veenstra TD, Systems biology and multi-omics. Proteomics, 2021. 21(3–4): p. e2000306. [DOI] [PubMed] [Google Scholar]
- 174.Kaushik AC, et al. , Cheminformatics and bioinformatics at the interface with systems biology: bridging chemistry and medicine. Vol. 24. 2023: Royal Society of Chemistry. [Google Scholar]
- 175.Mahner M and Kary M, What exactly are genomes, genotypes and phenotypes? And what about phenomes? Journal of theoretical biology, 1997. 186(1): p. 55–63. [DOI] [PubMed] [Google Scholar]
- 176.Orgogozo V, Morizot B, and Martin A, The differential view of genotype–phenotype relationships. Frontiers in genetics, 2015: p. 179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 177.Strachan T and Read A, Human molecular genetics. 2018: Garland Science. [Google Scholar]
- 178.Konstantinova IM, Tsimokha AS, and Mittenberg AG, Role of proteasomes in cellular regulation. International review of cell and molecular biology, 2008. 267: p. 59–124. [DOI] [PubMed] [Google Scholar]
- 179.Lecker SH, Goldberg AL, and Mitch WE, Protein degradation by the ubiquitin–proteasome pathway in normal and disease states. Journal of the American society of nephrology, 2006. 17(7): p. 1807–1819. [DOI] [PubMed] [Google Scholar]
- 180.Mishra R, et al. , Proteasome-mediated proteostasis: Novel medicinal and pharmacological strategies for diseases. Medicinal Research Reviews, 2018. 38(6): p. 1916–1973. [DOI] [PubMed] [Google Scholar]
- 181.Zampieri M and Sauer U, Metabolomics-driven understanding of genotype-phenotype relations in model organisms. Current Opinion in Systems Biology, 2017. 6: p. 28–36. [Google Scholar]
- 182.Martins MC, et al. , The Contribution of Metabolomics to Systems Biology: Current Applications Bridging Genotype and Phenotype in Plant Science, in Advances in Plant Omics and Systems Biology Approaches. 2022, Springer. p. 91–105. [DOI] [PubMed] [Google Scholar]
- 183.Aardema MJ and MacGregor JT, Toxicology and genetic toxicology in the new era of “toxicogenomics”: impact of “-omics” technologies. Toxicogenomics, 2003: p. 171–193. [DOI] [PubMed] [Google Scholar]
- 184.Nicholson JK and Wilson ID, Understanding’global’systems biology: metabonomics and the continuum of metabolism. Nature Reviews Drug Discovery, 2003. 2(8): p. 668–676. [DOI] [PubMed] [Google Scholar]
- 185.Bundy JG, Davey MP, and Viant MR, Environmental metabolomics: a critical review and future perspectives. Metabolomics, 2009. 5: p. 3–21. [Google Scholar]
- 186.Jang WD, et al. , Applications of artificial intelligence to enzyme and pathway design for metabolic engineering. Current Opinion in Biotechnology, 2022. 73: p. 101–107. [DOI] [PubMed] [Google Scholar]
- 187.Kim GB, et al. , Machine learning applications in systems metabolic engineering. Current opinion in biotechnology, 2020. 64: p. 1–9. [DOI] [PubMed] [Google Scholar]
- 188.Jin S, et al. , Application of deep learning methods in biological networks. Briefings in bioinformatics, 2021. 22(2): p. 1902–1917. [DOI] [PubMed] [Google Scholar]
- 189.Dasgupta A, Chowdhury N, and De RK, Metabolic pathway engineering: Perspectives and applications. Computer Methods and Programs in Biomedicine, 2020. 192: p. 105436. [DOI] [PubMed] [Google Scholar]
- 190.Sen P, et al. , Deep learning meets metabolomics: A methodological perspective. Briefings in Bioinformatics, 2021. 22(2): p. 1531–1542. [DOI] [PubMed] [Google Scholar]
- 191.Toubiana D, et al. , Combined network analysis and machine learning allows the prediction of metabolic pathways from tomato metabolomics data. Communications biology, 2019. 2(1): p. 214. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 192.Karp PD, et al. , The metacyc database. Nucleic acids research, 2002. 30(1): p. 59–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 193.Hosseini R, et al. , Pathway-activity likelihood analysis and metabolite annotation for untargeted metabolomics using probabilistic modeling. Metabolites, 2020. 10(5): p. 183. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 194.Fontaine N, et al. , Flux prediction using artificial neural network (ANN) for the upper part of glycolysis. Plos one, 2019. 14(5): p. e0216178–e0216178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 195.Shi X, et al. , Comprehensive isotopic targeted mass spectrometry: reliable metabolic flux analysis with broad coverage. Analytical chemistry, 2020. 92(17): p. 11728–11738. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 196.Antoniewicz MR, A guide to 13C metabolic flux analysis for the cancer biologist. Experimental & molecular medicine, 2018. 50(4): p. 1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 197.Antoniewicz MR, Stephanopoulos G, and Kelleher JK, Evaluation of regression models in metabolic physiology: predicting fluxes from isotopic data without knowledge of the pathway. Metabolomics, 2006. 2: p. 41–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 198.Misra BB, et al. , Integrated omics: tools, advances and future approaches. Journal of molecular endocrinology, 2019. 62(1): p. R21–R45. [DOI] [PubMed] [Google Scholar]
- 199.D. TM, Textbook of biochemistry: with clinical correlations. 2011, USA: John Wiley & Sons; Hoboken, NJ. [Google Scholar]
- 200.Di Minno A, et al. , Challenges in Metabolomics-Based Tests, Biomarkers Revealed by Metabolomic Analysis, and the Promise of the Application of Metabolomics in Precision Medicine. Int J Mol Sci, 2022. 23(9). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 201.Jendoubi T, Approaches to Integrating Metabolomics and Multi-Omics Data: A Primer. Metabolites, 2021. 11(3). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 202.Picard M, et al. , Integration strategies of multi-omics data for machine learning analysis. Comput Struct Biotechnol J, 2021. 19: p. 3735–3746. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 203.Picard M, et al. , Integration strategies of multi-omics data for machine learning analysis. Computational and Structural Biotechnology Journal, 2021. 19: p. 3735–3746. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 204.Cai Z, et al. , Machine learning for multi-omics data integration in cancer. Iscience, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 205.Xu F, et al. , Integration of ATAC-Seq and RNA-Seq identifies key genes and pathways involved in the neuroprotection of S-adenosylmethionine against perioperative neurocognitive disorder. Comput Struct Biotechnol J, 2023. 21: p. 1942–1954. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 206.Reel PS, et al. , Using machine learning approaches for multi-omics data analysis: A review. Biotechnol Adv, 2021. 49: p. 107739. [DOI] [PubMed] [Google Scholar]
- 207.Mo Q, et al. , Integrative multi-omics analysis of muscle-invasive bladder cancer identifies prognostic biomarkers for frontline chemotherapy and immunotherapy. Commun Biol, 2020. 3(1): p. 784. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 208.Tan K, et al. , A multi-omics supervised autoencoder for pan-cancer clinical outcome endpoints prediction. BMC Med Inform Decis Mak, 2020. 20(Suppl 3): p. 129. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 209.Giang TT, Nguyen TP, and Tran DH, Stratifying patients using fast multiple kernel learning framework: case studies of Alzheimer’s disease and cancers. BMC Med Inform Decis Mak, 2020. 20(1): p. 108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 210.Hao Y, et al. , Integrated analysis of multimodal single-cell data. Cell, 2021. 184(13): p. 3573–3587 e29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 211.Guan X, Runger G, and Liu L, Dynamic incorporation of prior knowledge from multiple domains in biomarker discovery. BMC Bioinformatics, 2020. 21(Suppl 2): p. 77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 212.Melo C, et al. , A Machine Learning Application Based in Random Forest for Integrating Mass Spectrometry-Based Metabolomic Data: A Simple Screening Method for Patients With Zika Virus. Front Bioeng Biotechnol, 2018. 6: p. 31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 213.Dias-Audibert FL, et al. , Combining Machine Learning and Metabolomics to Identify Weight Gain Biomarkers. Front Bioeng Biotechnol, 2020. 8: p. 6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 214.Ribeiro MTS, S.; Guestrin C Why Should I Trust You? Explaining the Predictions of Any Classifier. . in 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. San Francisco, CA, USA. [Google Scholar]
- 215.Lundberg SM, et al. , Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nat Biomed Eng, 2018. 2(10): p. 749–760. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 216.Liberti MV and Locasale JW, The Warburg Effect: How Does it Benefit Cancer Cells? Trends Biochem Sci, 2016. 41(3): p. 211–218. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 217.Bae J, Paltzer WG, and Mahmoud AI, The Role of Metabolism in Heart Failure and Regeneration. Front Cardiovasc Med, 2021. 8: p. 702920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 218.Gebregiworgis T and Powers R, Application of NMR metabolomics to search for human disease biomarkers. Combinatorial chemistry & high throughput screening, 2012. 15(8): p. 595–610. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 219.Young SP and Wallace GR, Metabolomic analysis of human disease and its application to the eye. Journal of ocular biology, diseases, and informatics, 2009. 2: p. 235–242. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 220.Jendoubi T, Approaches to integrating metabolomics and multi-omics data: a primer. Metabolites, 2021. 11(3): p. 184. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 221.Fiehn O and Spranger J, Use of metabolomics to discover metabolic patterns associated with human diseases, in Metabolic Profiling: Its Role in Biomarker Discovery and Gene Function Analysis. 2003, Springer. p. 199–215. [Google Scholar]
- 222.Shah SH, Kraus WE, and Newgard CB, Metabolomic profiling for the identification of novel biomarkers and mechanisms related to common cardiovascular diseases: form and function. Circulation, 2012. 126(9): p. 1110–1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 223.Fabrizio C, et al. , Artificial intelligence for Alzheimer’s disease: promise or challenge? Diagnostics, 2021. 11(8): p. 1473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 224.Inglese P, et al. , Deep learning and 3D-DESI imaging reveal the hidden metabolic heterogeneity of cancer. Chemical science, 2017. 8(5): p. 3500–3511. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 225.Sengodan P, et al. , Early detection and classification of malignant lung nodules from CT images: An optimal ensemble learning. Expert Systems with Applications, 2023. 229: p. 120361. [Google Scholar]
- 226.Bębas E, et al. , Machine-learning-based classification of the histological subtype of non-small-cell lung cancer using MRI texture analysis. Biomedical Signal Processing and Control, 2021. 66: p. 102446. [Google Scholar]
- 227.Bhinder B, et al. , Artificial intelligence in cancer research and precision medicine. Cancer discovery, 2021. 11(4): p. 900–915. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 228.Tran BX, et al. , The current research landscape of the application of artificial intelligence in managing cerebrovascular and heart diseases: A bibliometric and content analysis. International journal of environmental research and public health, 2019. 16(15): p. 2699. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 229.Dave D, et al. , Explainable ai meets healthcare: A study on heart disease dataset. arXiv preprint arXiv:2011.03195, 2020. [Google Scholar]
- 230.Faizal ASM, et al. , A review of risk prediction models in cardiovascular disease: conventional approach vs. artificial intelligent approach. Computer methods and programs in biomedicine, 2021. 207: p. 106190. [DOI] [PubMed] [Google Scholar]
- 231.Bom MJ, et al. , Predictive value of targeted proteomics for coronary plaque morphology in patients with suspected coronary artery disease. EBioMedicine, 2019. 39: p. 109–117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 232.Alawieh A, et al. , Using machine learning to optimize selection of elderly patients for endovascular thrombectomy. Journal of NeuroInterventional Surgery, 2019. 11(8): p. 847–851. [DOI] [PubMed] [Google Scholar]
- 233.Awwalu J, et al. , Artificial intelligence in personalized medicine application of AI algorithms in solving personalized medicine problems. International Journal of Computer Theory and Engineering, 2015. 7(6): p. 439. [Google Scholar]
- 234.Maier O, et al. , Extra tree forests for sub-acute ischemic stroke lesion segmentation in MR sequences. Journal of neuroscience methods, 2015. 240: p. 89–100. [DOI] [PubMed] [Google Scholar]
- 235.Lucas C, et al. , Learning to predict ischemic stroke growth on acute CT perfusion data by interpolating low-dimensional shape representations. Frontiers in neurology, 2018. 9: p. 989. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 236.Juarez-Orozco LE, et al. , Machine learning in the integration of simple variables for identifying patients with myocardial ischemia. Journal of Nuclear Cardiology, 2020. 27: p. 147–155. [DOI] [PubMed] [Google Scholar]
- 237.Liu X, et al. , Use of multimodality imaging and artificial intelligence for diagnosis and prognosis of early stages of Alzheimer’s disease. Translational Research, 2018. 194: p. 56–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 238.Battista P, et al. , Artificial intelligence and neuropsychological measures: The case of Alzheimer’s disease. Neuroscience & Biobehavioral Reviews, 2020. 114: p. 211–228. [DOI] [PubMed] [Google Scholar]
- 239.Singh S, et al. , Unveiling the future of metabolic medicine: omics technologies driving personalized solutions for precision treatment of metabolic disorders. Biochemical and Biophysical Research Communications, 2023. [DOI] [PubMed] [Google Scholar]
- 240.Pun FW, Ozerov IV, and Zhavoronkov A, AI-powered therapeutic target discovery. Trends in Pharmacological Sciences, 2023. [DOI] [PubMed] [Google Scholar]
- 241.Rao RV, AI for Drug Discovery: Transforming Pharmaceutical Research.
- 242.Stasevych M and Zvarych V, Innovative robotic technologies and artificial intelligence in pharmacy and medicine: paving the way for the future of health care—a review. Big Data and Cognitive Computing, 2023. 7(3): p. 147. [Google Scholar]
- 243.Zielinski JM, et al. , High Throughput Multi-Omics Approaches for Clinical Trial Evaluation and Drug Discovery. Front Immunol, 2021. 12: p. 590742. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 244.Azad RK and Shulaev V, Metabolomics technology and bioinformatics for precision medicine. Briefings in bioinformatics, 2019. 20(6): p. 1957–1971. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 245.Sahu M, et al. , Artificial intelligence and machine learning in precision medicine: A paradigm shift in big data analysis. Progress in Molecular Biology and Translational Science, 2022. 190(1): p. 57–100. [DOI] [PubMed] [Google Scholar]
- 246.You Y, et al. , Artificial intelligence in cancer target identification and drug discovery. Signal Transduction and Targeted Therapy, 2022. 7(1): p. 156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 247.David L, et al. , Molecular representations in AI-driven drug discovery: a review and practical guide. Journal of Cheminformatics, 2020. 12(1): p. 1–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 248.Martinelli DD, Machine learning for metabolomics research in drug discovery. Intelligence-Based Medicine, 2023: p. 100101. [Google Scholar]
- 249.Tanoli Z, Vähä-Koskela M, and Aittokallio T, Artificial intelligence, machine learning, and drug repurposing in cancer. Expert opinion on drug discovery, 2021. 16(9): p. 977–989. [DOI] [PubMed] [Google Scholar]
- 250.Vamathevan J, et al. , Applications of machine learning in drug discovery and development. Nature reviews Drug discovery, 2019. 18(6): p. 463–477. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 251.Iorio F, et al. , A landscape of pharmacogenomic interactions in cancer. Cell, 2016. 166(3): p. 740–754. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 252.Chakravarty K, et al. , Driving success in personalized medicine through AI-enabled computational modeling. Drug Discovery Today, 2021. 26(6): p. 1459–1465. [DOI] [PubMed] [Google Scholar]
- 253.Kouznetsova VL, et al. , Recognition of early and late stages of bladder cancer using metabolites and machine learning. Metabolomics, 2019. 15: p. 1–15. [DOI] [PubMed] [Google Scholar]
- 254.Wallace PW, et al. , Metabolomics, machine learning and immunohistochemistry to predict succinate dehydrogenase mutational status in phaeochromocytomas and paragangliomas. The Journal of pathology, 2020. 251(4): p. 378–387. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 255.Alakwaa FM, Chaudhary K, and Garmire LX, Deep learning accurately predicts estrogen receptor status in breast cancer metabolomics data. Journal of proteome research, 2018. 17(1): p. 337–347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 256.Murata T, et al. , Salivary metabolomics with alternative decision tree-based machine learning methods for breast cancer discrimination. Breast cancer research and treatment, 2019. 177: p. 591–601. [DOI] [PubMed] [Google Scholar]
- 257.Oh TG, et al. , A universal gut-microbiome-derived signature predicts cirrhosis. Cell metabolism, 2020. 32(5): p. 878–888. e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 258.Liang L, et al. , Metabolic dynamics and prediction of gestational age and time to delivery in pregnant women. Cell, 2020. 181(7): p. 1680–1692. e15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 259.Shen X, et al. , Serum metabolomics identifies dysregulated pathways and potential metabolic biomarkers for hyperuricemia and gout. Arthritis & Rheumatology, 2021. 73(9): p. 1738–1748. [DOI] [PubMed] [Google Scholar]
- 260.Luan H, et al. , Serum metabolomic and lipidomic profiling identifies diagnostic biomarkers for seropositive and seronegative rheumatoid arthritis patients. Journal of translational medicine, 2021. 19(1): p. 1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 261.Han S, et al. , A metabolomics pipeline for the mechanistic interrogation of the gut microbiome. Nature, 2021. 595(7867): p. 415–420. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 262.Jung S, et al. , Purine metabolite-based machine learning models for risk prediction, prognosis, and diagnosis of coronary artery disease. Biomedicine & Pharmacotherapy, 2021. 139: p. 111621. [DOI] [PubMed] [Google Scholar]
- 263.Liu S, et al. , Serum integrative omics reveals the landscape of human diabetic kidney disease. Molecular Metabolism, 2021. 54: p. 101367. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 264.Liu Y, et al. , Metabolomic biosignature differentiates melancholic depressive patients from healthy controls. BMC genomics, 2016. 17: p. 1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 265.Chagas-Paula DA, et al. , Prediction of anti-inflammatory plants and discovery of their biomarkers by machine learning algorithms and metabolomic studies. Planta Medica, 2015. 81(06): p. 450–458. [DOI] [PubMed] [Google Scholar]
- 266.Shen B, et al. , Proteomic and metabolomic characterization of COVID-19 patient sera. Cell, 2020. 182(1): p. 59–72. e15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 267.Hogan CA, et al. , Nasopharyngeal metabolomics and machine learning approach for the diagnosis of influenza. EBioMedicine, 2021. 71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 268.Delafiori J, et al. , Covid-19 automated diagnosis and risk assessment through metabolomics and machine learning. Analytical Chemistry, 2021. 93(4): p. 2471–2479. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 269.Aliakbarzadeh G, Sereshti H, and Parastar H, Pattern recognition analysis of chromatographic fingerprints of Crocus sativus L. secondary metabolites towards source identification and quality control. Analytical and Bioanalytical Chemistry, 2016. 408: p. 3295–3307. [DOI] [PubMed] [Google Scholar]
- 270.Asakura T, Date Y, and Kikuchi J, Application of ensemble deep neural network to metabolomics studies. Analytica Chimica Acta, 2018. 1037: p. 230–236. [DOI] [PubMed] [Google Scholar]
- 271.Linardatos P, Papastefanopoulos V, and Kotsiantis S, Explainable ai: A review of machine learning interpretability methods. Entropy, 2020. 23(1): p. 18. [DOI] [PMC free article] [PubMed] [Google Scholar]
