Abstract
Personalized medicine is essential for delivering efficacious care to children, yet pediatric-specific research and treatment guidelines lag behind those for adults. Multi-omics integration may offer clinicians biologically grounded patient stratification beyond traditional clinical biomarkers. While this approach holds promise for advancing precision medicine, it remains incompletely validated, and significant barriers to routine clinical implementation persist. In this review, we discuss practical considerations for integrating proteomics and metabolomics in pediatric multi-omics research. We describe popular code- and web-based tools, and recent examples of multi-omics studies in pediatric cohorts, concluding with current challenges and future outlooks for integrative multi-omics studies.
Subject terms: Biomarkers, Computational biology and bioinformatics
Introduction
Although clinical trials are widely regarded as the standard approach for determining treatment efficacy, they do not prioritize the vast heterogeneity of molecular phenotypes, disease mechanisms, and responses to treatments1,2. This heterogeneity can limit the generalizability of clinical trial outcomes, potentially resulting in suboptimal patient care3. Furthermore, pediatric populations face an additional challenge in that some treatment guidelines, such as those for septic shock or fluid resuscitation in trauma, are extrapolated from adult studies rather than derived from pediatric-specific research4–6. One solution to the challenges posed by heterogeneity is personalized medicine, an approach that tailors treatments and prevention strategies to each individual’s biology7.
To transition from a population-based to a personalized approach, the research community has turned to omics technologies to provide comprehensive molecular datasets at various biological levels to achieve a deeper molecular understanding of disease heterogeneity, mechanisms, and other sources of variability8. However, single-omics approaches do not account for interconnectivity between biological layers (e.g., transcriptomic, proteomic, and metabolomic layers, amongst many others). A multi-omics approach can provide a more holistic view, linking the biological layers to provide insights into trajectories over time that are not readily captured by single-omics analyses. For clinicians, multi-omics profiling offers a framework to understand diseases at a systems level rather than through isolated biomarkers. This approach has the potential to increase the accuracy of patient and risk stratification to inform targeted therapeutic decisions9. Through complex data analyses enabled by machine learning (ML) models, the research community aims to better characterize molecular processes, disease states, and predict outcome trajectories in individual patients, representing a promising but still incompletely validated step toward personalized medicine10,11.
Pediatric-specific considerations
Pediatric physiology shows marked differences from adult physiology and varies across developmental stages, making the clinical relevance of the majority of omics studies based on adult cohorts unclear12. For example, there are significant differences in the pediatric metabolome between children who are 6 months old, 2–3 years old, and young adolescents13,14. Furthermore, the lack of pediatric-specific studies has led to pediatric treatment guidelines (e.g., septic shock, trauma) being based on adult guidelines4–6. Consequently, there is a need for pediatric-specific research, and omics studies in particular, to define age- and disease-specific molecular profiles, with the aim of improving the precision of clinical decision-making.
Pediatric studies also require extra considerations regarding biospecimen sampling, as ethical and safety constraints differ from adult studies. For example, safe blood draw volumes are reduced in pediatric patients, potentially constraining the data acquisition methods that can be employed15. Additionally, obtaining informed consent is more complex, typically requiring parental/guardian consent in parallel with the child’s assent, adding layers of ethical and logistical complexity16. It is necessary to accommodate these and other challenges to achieve the pediatric-specific multi-omics profiling required to better understand disease mechanisms, delineate sub-phenotypes, and identify treatment responsiveness.
Why proteomics and metabolomics?
Ideally, multi-omics research would include all available modalities (genomics, transcriptomics, proteomics, metabolomics, etc.) along with comprehensive clinical data from healthcare settings and personal wearables to provide a holistic, systems biology view of a patient’s phenotype. While currently impractical, its development and routine clinical implementation remain an important goal for the future of personalized medical care17. In this review, we focus on a tractable starting point for multi-omics integration: the combination of two modalities, proteomics and metabolomics. Proteins are major executors of cellular processes, playing functional roles in reaction catalysis, signal transduction, and cellular structure. Therefore, a sample’s proteome can provide substantial insights into functional capabilities. Metabolites, the complement of small molecules in a sample that are less than 1.5 kDa, can be endogenous (produced by the organism) or exogenous (derived from the environment) and are substantially derived from, or associated with, proteins18. As such, the metabolome is widely considered the closest omics layer to an organism’s immediate phenotype19. The integration of proteomics with metabolomics allows researchers to connect the main functional machinery of the cell (the proteome) to the outcomes (metabolome) driven (at least in part) by that machinery, to provide insights into how organisms are reacting to their environment.
Goals of this review
While single-omics workflows are well-established, multi-omics integration remains a challenging task. Therefore, the goals of this review are to 1) establish foundational terminology and a conceptual framework for multi-omics integration, 2) critically evaluate available code- and web-based integration tools, 3) characterize the current state of integrated proteomic-metabolomic studies in pediatric cohorts, and 4) identify the key methodological and infrastructural gaps that must be addressed to advance the field towards robust, clinically translatable multi-omics profiling in pediatric health and disease. Figure 1 depicts a generalized multi-omics workflow, from idea conceptualization through to implementation. It references the other figures and serves as a rough outline of this review’s organization.
Fig. 1. Overview of proteomics-metabolomics integration.

Depiction of a multi-omics workflow from study conceptualization to implementation, illustrating the various steps and processes along the way. This figure roughly tracks the outline of this review and references other figures within this review. MS: Mass spectrometry, NMR Nuclear magnetic resonance, DDA Data-dependent acquisition, DIA Data-independent acquisition. Created in BioRender. Matusa, A. (2026) https://BioRender.com/k68h1nt.
Proteomics overview
Proteomics is the study of the entire protein complement in a biological sample and is crucial for investigating molecular profiles or disease mechanisms20,21. Contemporary classifications of proteomics focus on the analytical goal and context, such as quantitative proteomics (the focus of this review), post-translational modification (PTM) proteomics, functional proteomics (which characterizes protein activity, interactions, and regulatory dynamics in addition to abundance), spatial proteomics, and single-cell proteomics22–29. Broadly speaking, proteomics may be targeted, where a pre-defined panel of proteins is quantified, or untargeted, where protein quantification is unbiased30.
The two main analytical approaches to proteomics are mass spectrometry- and affinity-based methods (Fig. 2). Mass spectrometry (MS) is typically performed after initial processing by liquid chromatography (LC) and remains a popular method for performing both targeted and untargeted proteomics31,32. MS detects molecules based on their mass-to-charge ratio, where proteins are first enzymatically digested into peptides, ionized, and fragmented. Peptides may be analyzed in data-dependent acquisition (DDA) mode, where the most abundant ions are selected for fragmentation, or data-independent acquisition (DIA) mode, where all ions within defined mass windows are systematically fragmented, regardless of abundance20. DIA generates more complex spectra than DDA, often requiring more advanced algorithms for interpretation; however, it offers improved reproducibility, quantitative depth, and data completeness, and is becoming increasingly favored in biomarker discovery studies20. Peptide identification can be achieved through database searching, matching observed spectra against theoretical spectra (generated by in silico digestion of a reference protein database), or through spectral library matching against previously observed, empirically validated spectra33. When these approaches do not yield a match (e.g., for novel, unannotated peptides), de novo sequencing can help identify peptide sequences directly from fragmentation spectra33. Protein levels are then inferred by aggregating peptide-level evidence using methods such as summing or median-aggregating peptide intensities across a given protein34.
Fig. 2. Summary of proteomics and metabolomics technologies.

A LC-MS/MS workflow for obtaining proteomic datasets. B Proximity extension assay-based workflow for obtaining proteomic datasets. C LC-MS/MS workflow for obtaining metabolomic datasets. D NMR-based workflow for obtaining metabolomic datasets. LC Liquid chromatography; MS Mass spectrometry; NMR Nuclear magnetic resonance. Created in BioRender. Matusa, A. (2026) https://BioRender.com/6nkdyoh.
Affinity-based methods use antibodies or aptamers to facilitate targeted approaches with exceptional sensitivity35. Unlike ionization-based methods, affinity-based methods measure proteins indirectly by quantifying molecules that detect proteins of interest. SomaScan is a popular aptamer-based (short, single-stranded nucleotide polymer) technology that uses microarrays or next-generation sequencing (NGS) to quantify aptamers bound to target proteins36,37. Olink’s proximity extension assay (PEA) employs a pair of oligonucleotide-tagged antibodies and uses PCR or NGS to quantify target proteins with high specificity, a wide dynamic range, and very low cross-reactivity32,38–40. However, a limitation of affinity-based methods is that they are strictly targeted approaches that produce relative, affinity-based measurements. Additionally, they are susceptible to off-target binding, epitope/conformational blindness, and protein-altering variants (e.g., PTMs), all of which disconnect the observed signal from the protein’s true abundance41. For these reasons, MS is widely regarded as the reference approach for untargeted, discovery-based proteomics; however, in targeted settings, affinity platforms can offer superior sensitivity, better detection of low-abundance proteins, and higher throughput, with lower sample-volume requirements at a massively multiplexed scale42,43.
An implicit assumption made in quantitative proteomics is that quantity approximates function. While this assumption is reasonable for exploratory work, it is worth noting that protein function is heavily mediated through proteoforms rather than abundance alone. Proteoforms refer to the repertoire of molecular forms that a protein product, from a single gene, may take44. Proteoforms arise by various mechanisms, including genetic variants, alternative splicing, PTMs (phosphorylation, glycosylation, etc.), and proteolytic processing, to name a few45. This brings about a limitation of quantitative proteomics that researchers must note: quantitative proteomics often collapses proteoforms of a gene into one value, which may mask meaningful biological variation45,46. Depending on the research goals, investigators may choose alternative routes with better proteoform resolution such as top-down MS, which analyzes undigested proteins47.
Metabolomics overview
Metabolomics is the quantitative study of all metabolites (molecules <1.5 kDa) in a biological sample and is the layer considered the most reflective of an organism’s immediate phenotype18,19,48. Major metabolite classes include amino acids, carbohydrates, lipids, biogenic amines, organic acids, and steroids49. As with proteomics, metabolomic approaches may be targeted or untargeted, with targeted approaches trading higher sensitivity at the cost of discovery potential50,51. Common separation techniques in metabolomics workflows include liquid chromatography (LC), gas chromatography (GC), and capillary electrophoresis (CE)49.
The major metabolomic technologies are MS and nuclear magnetic resonance (NMR) spectroscopy (Fig. 2), with MS being the predominant approach in metabolomic studies52. Technological advancements over the years have increased the sensitivity, specificity, and metabolite coverage of modern MS instrumentation. The workflow resembles MS-based proteomics, where ionized metabolites are detected by their mass-to-charge ratio and identified by comparison to MS spectra databases. However, compared to protein databases such as UniProt [v2026_02], metabolite databases such as HMDB [v5.0] or METLIN [v1.0.0r00] are much less comprehensive53–55. Therefore, a large proportion of the metabolites detectable by MS remains unidentified and undocumented in databases, highlighting a significant gap in the field56.
NMR spectroscopy is an alternative, complementary approach in metabolomics studies. Although less sensitive than MS-based approaches, NMR is a highly reproducible, non-selective, and non-destructive approach that can accurately identify metabolites, even in complex mixtures57. NMR relies on the detection of dipole moments of various atomic nuclei, the generation of unique spectra, and alignment to reference databases. In metabolomics, the most commonly used approaches utilize 1H-NMR and 13C-NMR, and may be used for targeted or untargeted analyses58.
Integrating omics: multi-omics
In their simplest forms, omics datasets are matrices that quantify the molecules detected in a particular biological layer for a given number of samples. These data vary in format, complexity, size, noise, and distribution between different omics layers. Therefore, integrating multiple omics layers is often not as simple as combining two matrices into one; instead, integration typically requires sophisticated computational methods such as ML approaches59. Despite the difficulties inherent to processing them, integrating multiple omics layers can provide valuable information. Multi-omics analyses allow researchers to detect any patterns in the data that are obscured in single omic datasets60. The value of multi-omics analyses stems from the notion that biological layers are highly interconnected, with diseases often perturbing multiple layers. This is the central concept of “systems biology”, where biological organisms are viewed as dynamic and highly interconnected systems61. Changes in one layer, such as the proteome, can directly affect the status of another layer, such as the metabolome, with each layer providing complementary information that the other may lack. Thus, a multi-omics approach may provide a deeper mechanistic insight compared to a single-omics approach.
Multi-omics integration can also increase the predictive power of models. Using omics data as prognostic tools has immense clinical value as they provide a high-resolution view of a patient’s molecular phenotype. Omics profiling can be used to stratify patients, identify disease sub-phenotypes, and ultimately guide personalized treatment decisions62. ML models can use a patient’s omics dataset to make predictions such as likelihood of death or responsiveness to a particular treatment. However, ML models must be trained on a large volume of high-quality data to be useful. In principle, a training dataset with more information derived from multiple omics layers could, in certain cases, increase a predictive model’s performance63.
Preprocessing
It is worth noting that integration of multiple omics datasets requires preprocessed datasets and that preprocessing is performed on each omics layer separately before integration. Preprocessing omics datasets involves imputing missing values, dealing with outliers, correcting batch effects, and transformations such as scaling, log transformation, and normalization (Fig. 3). Unlike metabolomics, which directly quantifies metabolites, proteomics workflows usually quantify a proxy (e.g., an MS approach quantifying peptides), thereby warranting special considerations when it comes to preprocessing. For example, because peptides are often shared across multiple homologous proteins, assigning a measured peptide back to a single protein is inherently ambiguous, a challenge known as the protein inference problem64,65. Resolving this ambiguity requires a roll-up strategy to summarize peptide-level measurements into a single protein-level value, with methods ranging from simple summation to more robust statistical summarization approaches34.
Fig. 3. Preprocessing procedures for omics datasets.

Red squares indicate missing values; yellow squares indicate outliers; the various shades of blue represent the heterogeneous datapoints of any omics set. Created in BioRender. Matusa, A. (2026) https://BioRender.com/ccdgjjz.
Confidence in these peptide and protein identifications is often established upstream through a target-decoy search strategy, which estimates a false discovery rate (FDR) before quantitative values are carried into downstream preprocessing66. The normalization approach (e.g., median- or quantile-based scaling) can also affect quantification accuracy, with performance shown to vary across label-free proteomics datasets67. Batch effects (e.g., from instrument drift or different reagent lots) are similarly addressed at the protein level (e.g., via ComBat [sva v3.60.0]68), which recent benchmarking found to improve robustness across multi-batch studies69. Missingness in this context is also frequently non-random, for example, peptides present below the instrument’s detection limit are systematically missing rather than missing by chance; this necessitates imputation methods suited to this missing-not-at-random pattern (e.g., a hybrid left-censored imputation approach)70. In multi-omics integration, some frameworks like MOFA [v1.22.0] are built to handle missing data; however others, such as Support Vector Machines (SVM), are sensitive to missing data or outliers in the training dataset, thereby underscoring the importance of dataset preprocessing before integration71,72.
Integrating targeted and untargeted platforms introduces particular mismatches that must be addressed before downstream analyses. Consider, for example, the pairing of targeted affinity (Olink, SomaScan) proteomics with untargeted LC-MS metabolomics. These platforms differ along several axes. First, dynamic range and detection: affinity panels are bounded to their designated analytes, are calibrated to their abundance ranges, and report relative values (e.g., NPX, RFU), whereas untargeted LC-MS may capture more analytes at a narrower range, where low-abundance features may be lost due to ion suppression; therefore, each platform is blind to different portions of the entire abundance spectrum41–43,73. Additionally, targeted panels can produce near-complete datasets, while untargeted LC-MS data are sparse46,70,74. Second, coverage: preselected panels of a few hundred to a few thousand proteins contrast with thousands of (largely unannotated) metabolite features, resulting in asymmetric block sizes and unequal interpretability32,56. Depending on the downstream analyses, asymmetric block sizes can result in the masking of signals from smaller blocks, though this may be remedied by employing methods with built-in functionality to work with uneven block sizes (e.g., MOFA)72. Third, normalization: affinity data often arrive pre-normalized, while untargeted metabolomics often requires additional steps, and the two datasets occupy non-comparable numerical scales (e.g., relative NPX versus raw analyte intensities)39,40,75. The above points emphasize that preprocessing is not a one-size-fits-all process; rather, the choice of platforms and the corresponding preprocessing steps are design-specific and crucial for the validity of downstream analyses.
Terminology and classifications: early, intermediate, and late integration
Multi-omics integration is far from a one-size-fits-all approach and includes integration categories and subcategories. This review focuses on the simultaneous integration of different omics layers from a single sample cohort, termed meta-dimensional, vertical integration76–78. This review will focus on three subcategories within meta-dimensional, vertical integration: concatenation-based (early), transformation-based (intermediate), and model-based (late) integration (Fig. 4). Unfortunately, there is currently no consensus for subcategorization, and the literature describes other categories, including mixed, middle, or hierarchical integration77,79–81. As multi-omics integrative studies become more commonplace, the development of standardized terminology may help eliminate this confusion.
Fig. 4. Comparison of early, intermediate, and late integration strategies.

Early integration shows concatenation of preprocessed omics datasets, intermediate integration shows concatenation of sub-representations of omics datasets, while late integration shows concatenation of results. Created in BioRender. Matusa, A. (2026) https://BioRender.com/2l3a0wu.
Concatenation-based integration, also known as early integration, involves the combination of the different omics datasets (matrices) into one large matrix before analyses are performed59. A major benefit of this approach is the simplicity of combining the two preprocessed datasets and analyzing them as one. However, a major disadvantage of this approach is that it exacerbates already-existing integration challenges through the creation of a more complex, highly dimensional matrix80. This concatenated (combined) matrix makes subsequent machine learning difficult, aggravating model overfitting, which can result in less generalizable conclusions82.
In transformation-based (intermediate) integration, the data from each separate omics are first transformed into an intermediate form that shares a common, lower dimensional format, and then combined80. While this ameliorates the dimensionality and noise issues experienced in early integration, drawbacks include the loss of potentially important information during the transformation process, specifically complementary information between the omics layers. Today, multi-omics integration methods gravitate towards intermediate approaches83,84.
In model-based (late) integration, each omics dataset is analyzed separately, and the results are combined at the end into a final model; e.g., a final prediction model for drug responsiveness79. The strengths of this approach include the availability of established analysis tools, and that late integration circumvents the challenges associated with the early combination of multi-omics datasets of different formats. However, a limitation of this approach is that separate omics analyses preclude inter-omic interactions80. Additionally, hybrid frameworks such as CustOmics [v0.1.1] utilize a mixture of intermediate and late integration approaches and are becoming more commonplace in multi-omics research85.
Machine learning terminology
ML serves specific analytical roles in multi-omics research, including clustering samples into biologically meaningful groups, identifying disease sub-phenotypes, reducing high-dimensional data, and generating predictions such as disease severity or treatment responsiveness. The following sections are focused on explaining several ML-associated terms to better define their roles in multi-omics research and to serve as context for the section on pediatric applications and studies.
Machine learning classifiers
A classifier is an ML model that allocates a sample to a particular class. In omics research, a classifier may utilize a patient’s molecular profile (an omics dataset) to classify the patient into a particular group, such as healthy vs. disease. To do this, the model must first be trained on a training dataset, where the model learns the complex patterns in the dataset and associates those patterns with predefined outcomes (Fig. 5). The model then identifies combinations of variables (features) that can reliably predict an outcome. The underlying premise here is that only a small proportion of the dataset, e.g., a subset of the proteome, is relevant for classification purposes and the classifier learns which proteins are relevant. Machine learning occurs through repeated cycles of testing feature importances, known as weights, predicting outcomes, computing error, and re-adjusting weights to minimize the computed error (Fig. 5).
Fig. 5. Conceptual visualization of machine learning.

A A simplified view of how a classifier is trained. Blue squares represent features, green squares—positive outcome, red squares—negative outcome, purple squares—testing dataset not used in training. Dials represent feature weights. B Comparing SL and DL algorithm structures. SL Shallow learning, DL Deep learning. Created in BioRender. Matusa, A. (2026) https://BioRender.com/6av7dz6.
Shallow vs. deep learning, supervised vs. unsupervised learning
ML algorithms are broadly divided into shallow learning (SL) and deep learning (DL). SL algorithms are relatively straightforward, handle features in a defined way, are computationally efficient, and are well established in omics research86. DL algorithms, commonly referred to as neural networks, are multi-layer models with hidden layer architecture that learn feature relationships in a more computationally intensive and less interpretable manner86. ML can also be divided into supervised or unsupervised approaches. In supervised learning, the training dataset contains each sample’s label (outcome), whereas in unsupervised learning, no outcome is provided87. Therefore, unsupervised algorithms are primarily used for clustering or dimensionality reduction, and are useful for identifying sub-phenotypes within a cohort.
Dimensionality reduction
Omics datasets have high dimensionality, where the number of features in the dataset determines the number of dimensions. The well-known curse of dimensionality arises in datasets where the number of dimensions, p, is much, much greater than the number of samples, n. The problem with high dimensionality is that it risks model over-fitting where the model ends up learning patterns from noise rather than informative features, resulting in poor predictive performance and poor generalizability outside the training dataset59,88. To ameliorate this problem, dimensionality reduction (DR) techniques are employed prior to performing analyses on omics datasets. DR techniques can be subclassified into two approaches: feature selection and feature extraction80. Feature selection refers to the retention of informative, original features, while feature extraction derives new features that represent a subset of the original features.
DR techniques can be linear or non-linear. Principal Component Analysis (PCA) is the most popular linear technique, and popular non-linear techniques include t-distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP)89. t-SNE is a slower, unsupervised DR approach, while UMAP, which may be either supervised or unsupervised, is faster and more computationally efficient59,90. Finally, dimensionality reduction also plays a key role in outlier detection and cluster discovery91,92.
Model evaluation and metrics
An important step in ML analyses is model evaluation, i.e., determining how well a model does what it was trained to do. Models in omics research can be evaluated in two parts: by biological relevance and by predictive performance. In the context of biological relevance, a model’s performance can also be assessed based on the published literature, validation from single-omics analyses, or expertise in the field. Metrics used to evaluate predictive performance include accuracy, F1 score (combines precision and recall), Area Under the Receiver Operating Characteristic Curve (ROC-AUC), C-index, precision-recall curves, calibration curves, and decision-curve analyses79. However, not all metrics are equally appropriate, especially when models deal with imbalanced datasets, e.g., a model predicting mortality in a cohort where only 10% of the patients died. Precision-recall curves are more informative in these cases as they focus on model performance on the minority class, e.g., the 10% of patients who died93. Calibration curves assess whether predicted probabilities (e.g., a predicted 10% risk of death) reflect true outcome frequencies (the actual frequency of death) and can prove crucial if the model outputs are used to inform clinical decision-making94.
While evaluation metrics assess a model’s ability to discriminate between outcomes, discrimination alone does not guarantee clinical usefulness. A model with strong discriminative performance (e.g., high ROC-AUC) may still provide no clinical benefit if its predictions do not translate to better clinical decision-making. Decision-curve analysis addresses this by quantifying net benefit across a range of decision thresholds that a clinician might implicitly set (e.g., when a clinician chooses to escalate treatment if a model predicts a mortality risk of >30%)95. In other words, decision-curve analysis bridges discriminative performance with a model’s practical utility at the bedside.
Interpretability and explainability
Interpretability and explainability are two additional concepts that are increasingly prioritized in translational research, particularly when models inform clinical decision-making79. The notion of interpretability remains an active area of research in the ML community96,97. Broadly speaking, interpretability refers to understanding the model’s inner workings, i.e., how the model makes decisions. Interpretability is straightforward in shallow learning algorithms but much harder in deep learning (“black box”) approaches, since they employ hidden layer architectures. Post-hoc interpretations can be as simple as investigating computed variable importance scores to determine which features were most vital to the model’s performance97.
That said, it is important to guard against overinterpreting model outputs. Features deemed important by a model must be independently validated to ensure their importance is not due to random effects97. To guard against overinterpretation, one can test the same interpretation method (e.g., feature importance scores) on multiple models and assess the similarity of the interpretation results. Additionally, one must guard against a phenomenon known as feature selection instability, where feature importance rankings vary substantially across different data splits of the same dataset98. When feature rankings differ substantially across splits, the biological interpretation of the model should be treated with caution. Independent validation with an external cohort remains the most reliable safeguard against unstable feature selection99.
The concept of explainability refers to the justification of results derived from a “black box” model. Given the hidden layer architecture of neural networks, additional frameworks are often required to interpret these complex models. Two frameworks that are commonly utilized for this purpose are SHAP [v0.52.0] (Shapley Additive exPlanations) and LIME [v0.2.0.1] (Local Interpretable Model-agnostic Explanations)100,101. Compared to LIME, SHAP is more mathematically robust and consistent, albeit at the cost of being more computationally demanding102,103. Regardless of the approach, interpretability and explainability of ML models are considered crucial to biomedical research and the implementation of predictive models in clinical settings.
Pitfalls in machine learning analyses
Despite the advantages of ML approaches, several pitfalls such as data leakage, improper validation, and batch confounding can undermine model credibility unless they are carefully addressed. Data leakage occurs when external information influences model development, artificially inflating performance results104. An example of data leakage in omics research might be using a list of differentially expressed proteins as the input for a classifier rather than a raw dataset. Next, the validation of model results is crucial. If the same dataset is used for both training and testing (internal validation), performing non-nested cross-validation produces optimistically-biased performance estimates105. Therefore, nested cross-validation, which better isolates the testing dataset from the training dataset, improves model credibility. The majority of published omics studies utilize internal validation, while external validation, which is generally considered the reference standard, remains nearly absent106. To improve credibility in omics prediction studies and facilitate clinical adoption of models, external validation must be prioritized.
Finally, site or batch structure can also confound model training, where there is a mismatch between the samples in the training dataset and the samples in the testing dataset107. This pitfall is particularly important when assessing pediatric cohorts, which are often small, single-site, and imbalanced. Awareness of these limitations is essential when designing workflows and performing ML analyses in pediatric omics research.
When multi-omics integration might help, and when it might not
Multi-omics integration can improve biological insights when the axes of variation underlying a condition are not fully captured by a single omics layer. The axes of variation underlying a condition refer to the primary, independent, and typically continuous dimensions (principal components) that explain the greatest amount of heterogeneity within that condition. Rather than seeing a condition as a single entity, axes of variation represent how individual cases differ across key clinical dimensions. Tools like MOFA (further described in the Tools section) provide insights for each omics layer to aid in assessing sources of variation. In a study of 200 chronic lymphocytic leukemia patients, MOFA revealed certain factors (e.g., factors 1 and 2) that were shared across multiple biological layers while others (e.g., factor 4) were layer-specific72. Focusing on proteomics and metabolomics, their integration is most valuable when there is mechanistic coupling between them and the disease or condition under study. If the dominant axis of variation is confined to a single layer, e.g., the proteome, integrated analyses of an additional layer, e.g., the metabolome, are unlikely to yield additional insights. Applying a variance-partitioning approach such as MOFA to characterize the main sources of variation therefore serves as an informative starting point for determining whether proteo-metabolomic integration will improve biological insights.
Multi-omics integration may not always improve results; rather, integration can sometimes hinder predictive performance. A multi-omics analysis across 14 TCGA cancer datasets demonstrated that models incorporating one or two omics layers outperformed those combining four or five, attributable to the introduction of additional noise and the exacerbation of the curse of dimensionality108. With these limitations in mind, the failure modes of multi-omics integration are examined below.
Several failure modes and statistical pitfalls are particularly prevalent in multi-omics studies. Batch effects (e.g., non-biological variations due to technical factors such as reagent lots or processing dates) are widespread and, when confounded with biological outcomes, can produce misleading conclusions109; in a proteomics study of ovarian cancer, for instance, sample processing day was a primary driver of the reported results110. Data leakage occurs when training and testing conditions do not reflect real-world implementation, resulting in over-inflated model performance. For example, a multi-omics model trained to predict trauma severity using samples collected 24 h post-admission effectively incorporates outcome information, whereas its intended clinical use requires predictions at presentation104. Overfitting, a pervasive concern in high-dimensional analyses, is further exacerbated in multi-omics settings and is closely linked to inadequate validation. As described in the previous section, both internal validation and external validation are optimal to mitigate bias and ensure generalizability99,111.
Collectively, multi-omics integration presents both advantages and limitations. Whether integration will be beneficial for a given study cannot be assumed a priori, particularly in the absence of prior biological knowledge. For this reason, it is advisable to assess the contribution of each omics layer (using a tool such as MOFA) and to conduct parallel single-omics analyses as a basis for comparison. Although studies demonstrating superior performance of multi-omics models over single-omics counterparts exist, there is limited evidence of their clinical translation. To our knowledge, no pediatric-specific studies have demonstrated that multi-omics models improve prediction in real-world clinical settings. This can be attributed to several factors, including the novelty of the field, the scarcity of implemented studies, and the methodological reality that integration of multi-omics layers does not necessarily improve predictive performance over the most informative single omics layer. Furthermore, prospective validation against established clinical panels remains largely unreported, especially in pediatric cohorts112,113. A goal of this review is to encourage and facilitate pediatric-specific multi-omics studies that may eventually enhance clinical decision-making and patient outcomes. Toward this goal, the following section provides a brief overview of the most widely used code- and web-based tools for integrative multi-omics analyses.
Biological integration approaches
While the integration methods described in the previous sections can identify statistical associations between and across proteomic and metabolomic features, these associations do not inherently establish biological mechanisms. Interpreting computational integration results in the context of known biology strengthens both interpretability and translational relevance. Enzyme-metabolite relationships can be established by reconciling multi-omics integration results with known enzyme-substrate associations annotated in pathway databases (e.g., KEGG [v119.0]114) and concordant or discordant changes between enzyme-metabolite pairs can then flag specific dysregulated reaction steps or pathways115.
Pathway-constrained approaches map significant proteins and metabolites onto curated biological pathways (e.g., KEGG114), while network biology approaches may build protein-metabolite interaction networks (based on validated interactions) to identify key hubs and modules that point to the overarching biological themes at play116. In contrast, data-driven approaches such as co-expression module detection (e.g., the weighted gene co-expression network analysis algorithm117) build networks without relying on known interactions, offering a complementary route to identifying potential biological themes that curated networks may miss118. As another application, integrated proteo-metabolomic data can be used to constrain genome-scale metabolic network models (e.g., via flux balance analysis) to estimate plausible pathway flux distributions and hint at underlying mechanisms119. While flux balance analyses using static multi-omics datasets serve to generate mechanistic hypotheses, they remain model-based estimates rather than validated flux measurements, as establishing true flux quantification requires dedicated isotope-tracing experiments120. Collectively, these approaches move multi-omics research from computational associations toward the interpretation of biological themes and disease mechanisms at play, thereby strengthening translational relevance.
Tools and approaches to integrating omics
Code-based tools
While code-based tools are the most diverse and customizable tools for multi-omics integration, they also require coding expertise to employ. The main advantages of code-based tools are their customizability and transparency when disseminating results. Various web-based tools employ code-based tools “under the hood” albeit with limitations in the user’s ability to customize parameters. The most commonly used code-based tools are R [v4.5.2] and Python [v3.14]121,122. For multi-omics integration, R-based tools are sufficient for almost all aspects of the workflow, including ML analyses employing shallow learning algorithms. Python becomes advantageous when the workflow requires deep learning algorithms. For both R and Python, numerous omics-specific packages and libraries (respectively) are available to streamline workflows and reduce the amount of coding and tweaking required. A community-maintained collection of multi-omics integration tools with respective publications is available on GitHub: https://github.com/mikelove/awesome-multi-omics. Several popular tools are described below, with Table 1 summarizing their features, strengths, and limitations, with reference to the terminology introduced in the previous sections. Figure 6 is included to provide a practical, general guide to choosing the appropriate tool for multi-omics integration. All software versions reported in this review were verified on 4 August 2026.
Table 1.
Summary of popular multi-omics integration tools
| Tool | Interface/language | Integration stage | Learning style | Missing data handling | Interpretability | Typical use cases | Strengths and limitations |
|---|---|---|---|---|---|---|---|
| mixOmics (DIABLO) | R | Intermediate |
Supervised; Shallow learning |
Limited; assumes mostly complete data | High | Biomarker discovery and classification across omics layers (e.g., responders vs. non-responders) |
Strengths: interpretable latent components; built-in tools for model evaluation, parameter tuning, and imputation. Limitations: poor handling of substantial missingness; requires class labels. |
| MOFA+ | R and Python | Intermediate |
Unsupervised; Shallow learning |
Yes; tolerates missing values and entire missing modalities | High | Hypothesis generation, sub-phenotyping, identifying shared and omics-specific sources of variation |
Strengths: probabilistic framework; outputs variance explained per factor and per omics layer; downstream-friendly results. Limitations: unsupervised only; factors may not align with phenotype of interest. |
| SNFtool | R | Intermediate (network fusion) |
Unsupervised; Shallow learning |
Not built-in | Moderate | Patient stratification and clustering using sample-similarity networks |
Strengths: captures both shared and complementary information across layers; pairs well with WGCNA and Cytoscape for downstream module detection. Limitations: clustering only; no built-in classification or supervised modeling. |
| OmiEmbed | Python | Intermediate |
Supervised; Deep learning |
Not built-in | Low | Multi-task prediction: classification, regression (e.g., disease severity), and survival analysis |
Strengths: strong predictive performance across multiple task types; structured latent space preserves biological signal. Limitations: limited interpretability; computationally heavier than shallow methods. |
| CustOmics | Python | Hybrid |
Supervised; Deep learning |
Yes; via phase 1 learning stage | Low | Phenotype prediction with two-phase learning of per-omic and joint representations |
Strengths: flexible architecture; learns per-omic sub-representations before joint fine-tuning. Limitations: limited interpretability; computationally heavier than shallow methods. |
| MOGONET | Python | Intermediate |
Supervised; Deep learning |
Not built-in | Low | Patient classification and biomarker identification using sample-sample graphs |
Strengths: leverages graph structure across samples; competitive classification performance. Limitations: low interpretability; requires graph-construction choices that affect outputs. |
| OmicsAnalyst | Web-based (no-code); R package available | Multiple (early, intermediate, late depending on workflow) |
Supervised and unsupervised, depending on analysis; Shallow learning |
Depends on selected method | High | End-to-end exploratory analysis: preprocessing, dimensionality reduction, clustering, classification, pathway enrichment |
Strengths: integrates DIABLO and MOFA under one interface; cross-omics correlation and phenotype-association analyses; rich visualizations. Limitations: limited parameter customization compared to native R/Python implementations. |
| OmicsNet | Web-based (no-code); R package available | Late (knowledge-driven network mapping) | Not applicable (network/topology) | Not applicable | High | Network construction, topology analysis, and pathway-level interpretation across omics layers |
Strengths: integrates molecule–molecule and pathway-level relationships; strong visualization. Limitations: not designed for predictive modeling or sample classification. |
| PaintOmics4 | Web-based (no-code) | Late (pathway-level mapping) | Not applicable (visualization, enrichment) | Not applicable | High | Pathway-centric integration and visualization using KEGG, Reactome, and similar databases | Strengths: intuitive pathway visualization; broad omics platform support.Limitations: not predictive; complementary to (rather than substitute for) statistical or ML-based tools. |
DIABLO Data Integration Analysis for Biomarker discovery using Latent cOmponents, MOFA Multi-Omics Factor Analysis, SNF Similarity Network Fusion, WGCNA weighted gene co-expression network analysis, ML machine learning. Bolding: “Strengths:” and “Limitations:” are bolded to improve readability.
Fig. 6. Decision tree to guide selecting presented multi-omics integration tools.

This decision tree is meant to be a guide, not absolute, and overlaps are not shown (for example, OmicsAnalyst contains both MOFA and DIABLO). Tools are stratified by interface, then programming language, analytical goals, methodological requirements, supervision, integration architecture, and output types. Note that MOFA+ appears in both R and Python branches, reflecting its availability in both languages.
R
R is the most widely used programming language for biological research with a vast collection of packages tailored to specific fields, including omics research. Single-omics research using R is well established, with packages for virtually all aspects of the workflow. Additionally, a variety of shallow learning ML algorithms are readily available for implementation, rendering R as a “one-stop-shop” for omics research. Some of the popular tools available in R for performing integrative analyses on multi-omics data are listed below.
mixOmics [v6.36.0] is an R package capable of analyzing single and multi-omics datasets123. It has a variety of integrative multi-omics analyses, but it is most widely known for DIABLO (Data Integration Analysis for Biomarker discovery using Latent cOmponents), a supervised, multivariate, multi-omics integration method124. DIABLO learns correlated feature sets across omics layers that effectively discriminate predefined classes, e.g., treatment responders vs. non-responders. DIABLO is also referred to as multiblock sparse partial least squares-discriminant analysis (multiblock sPLS-DA) and is an intermediate integration method that builds latent components by maximizing covariances between datasets. A limitation of this tool is that, since DIABLO assumes mostly complete data, it does not handle datasets with a substantial amount of missing data well. In addition to integrative analysis, mixOmics has tools for evaluating model performance, tuning parameters, and even data preprocessing tools for missing value imputation.
MOFA + (Multi-Omics Factor Analysis) is an R/Python package that integrates multi-omics data through a matrix factorization approach125. MOFA is an unsupervised, probabilistic latent factor model that integrates multiple omics layers to identify both shared and omics-specific sources of variation. Conceptually, MOFA is a multi-view, probabilistic extension of PCA. MOFA outputs a set of factors and their association with each sample, as well as the loadings (the biological molecules) that make up each factor. It outputs the variance explained per factor and per omics layer, aiding interpretation and factor prioritization. The loadings specify the features that comprise each factor and their importance to the factor. Practically speaking, the loadings matrices contain the data, e.g., the proteins (in the case of a proteomics dataset) that comprise the factor, allowing for downstream analyses and biological interpretation using common methods such as Gene Set Enrichment Analysis (GSEA) or Over-Representation Analysis (ORA)126,127. Additionally, the factors may be used for downstream clustering, sub-phenotyping, or association with clinical symptoms. Unlike DIABLO, MOFA is explicitly modeled to handle noise, missing values, or even missing modalities, so dealing with missing values is not necessary. However, a limitation is that MOFA is strictly unsupervised, meaning its factors may not align with phenotypes or comparison groups of interest.
SNFtool [v2.3.1] (Similarity Network Fusion tool) is an R package that takes a network-based approach to integrating multi-omics data128. In network-based approaches, molecules (proteins, metabolites, etc.) or even patients are treated as nodes, and associations between nodes (correlation, co-expression patterns) are represented as edges connecting nodes. Bundles of closely connected nodes form modules, which are biologically meaningful units often enriched for common pathways or shared patterns8. These modular structures are especially useful for identifying coordinated biological responses to injury or infection. SNF is a well-known, unsupervised technique that constructs sample-similarity networks separately for each omics layer, subsequently fusing all layers into a single network that captures both shared and complementary information across multi-omics layers60. SNFtool enables this integration and downstream clustering. Tools like weighted gene co-expression network analysis (WGCNA) [v1.74] and Cytoscape [v3.10.4] can be used post-integration for module detection and visualization117,129,130.
Python
Python is an alternative language for omics analyses. While less popular than R, Python is preferred for deep learning and includes a variety of omics-specific algorithms, some of which are described below131,132.
OmiEmbed [v0.2.1] is a supervised, intermediate integration, deep learning framework for analyzing multi-omics data133. It is an upgrade from the original OmiVAE [v0.1] and uses a Variational AutoEncoder (VAE) to process the data. Variational autoencoders are deep learning models that learn flexible, non-linear mappings of data into latent spaces. In principle, an autoencoder takes an input source, encodes it into a latent representation, and then decodes the representation into a product that resembles the original input. Variational autoencoders contain a probabilistic framework and learn a continuous, structured latent space, thereby outperforming standard autoencoders in preserving meaningful biological structures85. OmiEmbed is open-source, implemented only in Python, and shows superior performance in classification tasks, regression tasks (predicting a value such as disease severity), and survival prediction133. Given that this is a DL approach, interpretability is lower than that of SL methods, although latent embeddings can be inspected to provide some level of interpretability.
CustOmics is a supervised integration method that also employs an Autoencoder (AE) to analyze multi-omics data85. CustOmics uses two learning phases, where the first learns sub-representations of each omics dataset independently, while in the second, the sub-representations are concatenated, followed by a learning stage to fine tune the model. Therefore, CustOmics can be viewed as a hybrid rather than an intermediate integration approach. Autoencoder weights and post-hoc feature importance can be extracted, but as with other DL models, interpretability is lower than in SL models.
MOGONET [updated 31 March 2021] (Multi-Omics Graph cOnvolutional NETworks) is a supervised, intermediate integration method that adopts a graphical approach to integrating multi-omics data134. Specifically, MOGONET uses graph neural networks to integrate multiple omics layers by constructing sample-sample graphs for each omics layer and learning representations that are optimized for phenotype classification. The learned representations for each omics are then combined to make final predictions, again, with low model interpretability.
Web-based, user-friendly tools
A variety of web-based programs are available to perform integrative analyses of multi-omics data. These programs often employ the same code-based tools in user-friendly, graphical user interfaces. These programs exclusively employ SL approaches, and to our knowledge, there are currently no user-friendly apps/webpages that employ DL approaches to analyze multi-omics data. A few user-friendly tools that perform integrative analyses on multi-omics datasets are discussed below.
OmicsAnalyst [v2.0] is a web-based, no-code tool capable of a variety of integrative analyses for multi-omics datasets135. OmicsAnalyst comes from a suite of web-based tools for analyzing omics data, including OmicsNet [v2.0], ProteoAnalyst [updated 13 May 2026], ExpressAnalyst [updated 14 April 2026], and the popular MetaboAnalyst [v6.0]135,136. OmicsAnalystR [updated 4 August 2026], ExpressAnalystR [v1.0.0], MetaboAnalystR [v4.2], and OmicsNetR [v1.0.0] are also available as R packages. OmicsAnalyst is designed to be an end-to-end statistical and machine-learning-driven integration tool. It supports data preprocessing, differential analysis, dimensionality reduction, clustering, supervised classification, and pathway enrichment. It has tools like DIABLO and MOFA built-in, supports cross-omics correlation and phenotype association analyses, and produces a variety of graphical and network-based visualizations.
OmicsNet is a web-based network integration and visualization platform that emphasizes molecule-molecule and pathway-level relationships135. It integrates omics data onto biological networks (protein-protein, gene-metabolite, pathway networks) to enable network construction, topology analysis, and visual interpretation across omics layers.
PaintOmics4 [v1.0.0] is a web-based multi-omics pathway analysis and visualization platform137. Supporting a variety of omics platforms, PaintOmics4 maps data onto biological pathways from well-known databases (e.g., KEGG, Reactome [v97]), enabling pathway-centric integration114,138. This tool allows users to visualize coordinated changes across omics layers, perform pathway enrichment, and explore regulatory activity within pathways. PaintOmics4 is primarily designed for biological interpretation and visualization, rather than predictive modeling or sample classification, thereby rendering it a complementary tool to statistical or ML-based multi-omics analyses.
Examples of multi-omics studies in pediatric cohorts
Multi-omics approaches have been used to investigate a variety of diseases, with different combinations of omics layers having been employed to study both pediatric and adult populations. This section focuses on multi-omics, specifically proteomic and metabolomic studies of pediatric populations from the past 5 years.
Search strategy and retrieved studies
A search and screen of PubMed and Web of Science (WoS) databases (on April 24, 2026) yielded 28 studies integrating proteomics and metabolomics in pediatric samples from the past 5 years (Table 2). Our search queries were:
Table 2.
Studies integrating proteomics and metabolomics in pediatric samples from the past 5 years
| Year/Ref | Disease/Topic, Longitudinal (LG) vs Cross-sectional (CS) | ~Total Patients | Age Range | Sample Type(s) | Proteomics (approach and approx. coverage) | Metabolomics (approach and approx. coverage) | Integration Stage | Main Analyses |
|---|---|---|---|---|---|---|---|---|
| 2026144 | Tuberculosis (CS) | 237 | 4 (2–8) | Plasma |
MS (unspecified) Untargeted ~850 proteins |
MS (unspecified) Targeted ~100 metabolites |
Early; Intermediate; Late | classifier training (early), DIABLO (intermediate), multiGSEA pathway (late) integration (aggregated p-values) |
| 2026167 | Acute illness leading to death (CS) | 3101 | <2 | Plasma (proteomics), serum (metabolomics) |
SomaScan Targeted ~6,400 proteins |
MS, unspecified Targeted ~170 metabolites |
Late | Built a multi-omics classifier to predict mortality risk |
| 2026168 | Chronic Rhinosinusitis with Nasal Polyps (CS) | 39 | ~10.5 (2.1) | Nasal secretions |
LC-MS/MS Untargeted ~5,600 proteins |
LC-MS Untargeted |
Late | Correlation analysis of significant metabolites and proteins |
| 2026169 | Mycoplasma pneumoniae infection (CS) | 67 | 3–15 | Bronchoalveolar lavage fluid | LC-MS/MS | UPLC-MS | Late | Correlation analyses between proteins and metabolites, built a multi-omics classifier using significant proteins and metabolites. |
| 2025170 | Cerebral palsy (LG) | 91 | ~2–7 | Plasma |
LC-MS/MS Untargeted ~670 proteins |
GC-MS Untargeted ~350 metabolites |
Late | Joint pathway enrichment in MetaboAnalyst |
| 2025147 | Antrochoanal Polyps; Chronic Rhinosinusitis with Nasal Polyps (CS) | 18 | 6–14 | Nasal tissue |
LC-HRMS Untargeted ~4,900 proteins |
UHPLC-MS/MS Untargeted |
Late | Pathway enrichment, correlation analyses between proteins and metabolites. |
| 2025171 | Infantile epileptic spasms syndrome (CS) | 20 | <2 | Cerebrospinal fluid |
iTRAQ Untargeted |
GC-MS Untargeted |
Late | Correlation analyses of top proteins and metabolites |
| 2025172 | Mycoplasma pneumoniae pneumonia (CS) | 170 | ~6 (3–8) | Plasma |
Olink PEA Targeted ~2,900 proteins |
LC-MS/MS Targeted & Untargeted |
Late | Pathway analysis overlap, correlation analyses of proteins and metabolites |
| 2025173 | Cow’s milk allergy (LG) | 40 | <13 months | Fecal, saliva |
LC-MS/MS; Olink PEA Targeted & Untargeted |
UPLC-HRMS ~170 metabolites |
Early; Late | Classifier training via early and late integration methods |
| 2025174 | Autism (LG) | 138 | 7–10 | Umbilical cord blood plasma |
LC-MS/MS Untargeted ~500 proteins |
LC-MS/MS Targeted & Untargeted ~1,800 metabolites |
Late | Separate omics analyses, separate ML models; late integration at results/interpretation stage |
| 2025175 | Barth syndrome (CS) | 5* | 5mo– 15 yrs* | Left ventricle heart tissue |
UPLC-MS/MS Untargeted ~1,400 proteins |
UHPLC-MS/MS Semi-targeted ~120 metabolites |
Late | PCA clustering, pathway annotation |
| 2025140 | Obesity (CS) | 1041 | 6–11 | Plasma (proteomics), serum & urine (metabolomics) |
Luminex assay Targeted ~36 proteins |
LC-MS/MS (serum), 1H-NMR (urine) Targeted ~221 metabolites |
Intermediate | MOFA (latent factors) |
| 2025176 | Crohn’s disease (CS) | 58 | 9–18 | Fecal |
LC-MS/MS Untargeted ~140 proteins |
HILIC-HRMS Targeted ~200 metabolites |
Early | Integration to build a classifier for active disease vs. remission classification; correlation analysis between proteins and metabolites |
| 2025142 | Obesity and metabolic dysfunction (CS) | 863 | 7.8 (1.4) | Plasma (proteomics), serum (metabolomics) |
Luminex assay Targeted ~35 proteins |
LC-MS/MS Targeted ~100 metabolites |
Intermediate | Multi-omics similarity network fusion (via SNFtool) |
| 2025139 | Rhabdomyosarcoma (CS) | 53 | 1–16 | Plasma |
LC-MS/MS Untargeted ~500 proteins |
LC-MS/MS Untargeted ~250 metabolites |
Intermediate | DIABLO, MOFA |
| 2025141 | Telomere length (CS) | 1001 | 6–11 | Plasma (proteomics), serum & urine (metabolomics) |
Luminex assay Targeted ~36 proteins |
LC-MS/MS (serum), 1H-NMR (urine) Targeted ~221 metabolites |
Intermediate | Latent components analyses via mixOmics package |
| 2025177 | Autism (CS) | 60 | ~7.3 (2.93) | Fecal |
LC-MS/MS Untargeted |
LC-MS/MS Untargeted |
Late | Pathway overlap, post analysis correlations |
| 2024178 | Idiopathic nephrotic syndrome (LG) | 18 | ~ 6–15 | Saliva |
LC-MS/MS Untargeted ~4,300 proteins |
LC-MS/MS Targeted ~1,500 metabolites |
Late | Protein-metabolite correlations, pathway enrichment overlap |
| 2024143 | Metabolic dysfunction-associated fatty liver disease (LG) | 420 | 7.19 (1.11) | Plasma (proteomics), serum (metabolomics) |
Luminex assay Targeted ~36 proteins |
LC-MS/MS Targeted ~180 metabolites |
Early; Intermediate; Late | Latent clustering (LUCID framework) at multiple integration stages; joint pathway analyses |
| 2023149 | RSV Pneumonia (CS) | 162 | 2.17 (2.33) | Serum |
LC-MS/MS Untargeted ~500 proteins |
UPLC-MS/MS Untargeted ~2,700 metabolites |
Late | Separate omics analyses, combining results at interpretation stage |
| 2023179 | Autism (CS) | 20 | 2-6 | Plasma |
SWATH-MS Untargeted ~850 proteins |
HPLC-MS Untargeted ~3,400 metabolites |
Late | Joint pathway analysis in MetaboAnalyst |
| 2023180 | Nephrotic syndrome (LG) | 15 | 2–14 | Plasma |
LC-MS/MS Untargeted ~1,100 proteins |
NMR ~45 metabolites |
Late | Joint Pathway Analysis in MetaboAnalystR |
| 2023181 |
Community-acquired pneumonia (CS) |
179 | ~ 1–8 | Serum |
LC-MS/MS Untargeted ~500 proteins |
UPLC-MS/MS Untargeted ~2,700 metabolites |
Late | Build classifier, pathway analyses, correlation analyses |
| 2023182 | Lesch-Nyhan syndrome (CS) | 3* | 4mo– 4 yr* | Red blood cells |
LC-MS/MS Untargeted ~900 proteins |
UHPLC-MS Targeted ~190 metabolites |
Late | Network analyses and pathway analyses performed in OmicsNet |
| 2023183 | General child development and aging (CS) | 1173 | 5–12 | Plasma (proteomics), serum & urine (metabolomics) |
Luminex assay Targeted ~36 proteins |
LC-MS/MS (serum), 1H-NMR (urine) Targeted ~221 metabolites |
Early | Prediction model training via concatenated proteomic and metabolomic datasets |
| 2023184 | Endocrine disrupting chemicals (LG) | 143 | 6–11 | Plasma (proteomics), serum & urine (metabolomics) |
Luminex assay Targeted ~36 proteins |
LC-MS/MS Targeted ~221 metabolites |
Early | Multi-omics network construction to depict exposure-to-molecular feature associations |
| 2022185 | Leukemia (LG) | 8 | 3–15 | Bone marrow interstitial fluid and peripheral blood plasma |
LC-MS/MS Untargeted |
LC-MS Untargeted ~160 metabolites |
Early | UMAP and fuzzy-c clustering to identify clusters |
| 2022186 | Exposome (LG) | 1301 | 6–11 | Plasma (proteomics), serum & urine (metabolomics) |
Luminex assays Targeted ~36 proteins |
LC-MS/MS Targeted ~221 metabolites |
Late | Associations of omics layers with exposome, combined network visualizations |
LC Liquid chromatography, UPLC ultra performance LC, UHPLC ultra high performance LC, GC gas chromatography, MS mass spectrometry, MS/MS tandem MS, PEA proximity extension assay, WGCNA weighted gene co-expression network analysis, MOFA Multi-Omics Factor Analysis, DIABLO Data Integration Analysis for Biomarker discovery using Latent cOmponents, PCA principal component analysis, iTRAQ isobaric Tags for Relative and Absolute Quantitation, ML machine learning.
Bolded terms are tools that are described in this review. Due to high study and reporting heterogeneity, numbers are largely approximations and should not be regarded as absolute. Additionally, study features are not reported if the original study did not clearly specify (e.g., number of proteins quantified). Age range is reported in years unless otherwise specified; reported as range: X-Y, mean/SD: X(Y), or median/IQR: X(Y-Z) depending on the study. *: only pediatric patients and age ranges reported; these two denoted studies used adult controls. LG longitudinal, CS cross-sectional.
Web of Science: ((TI = (Integrat*) AND TI = (omic*)) OR (TI = (multi-omic*) OR TI = (multiomic*))) AND (TI = (pediatric*) OR TI = (paediatric*) OR TI = (children) OR AB = (pediatric*) OR AB = (paediatric*) OR AB = (children))
PubMed: (integrat*[Title] AND omic*[Title]) OR (multi-omic*[Title] OR multiomic*[Title]) AND (pediatric*[Title/Abstract] OR children[Title/Abstract])
Results were subsequently filtered by publication date (the last 5 years) and to exclude reviews, resulting in 203 PubMed and 273 WoS studies. Covidence was used to input (n = 476), screen for duplicates (n = 158), and extract information from these studies, resulting in 318 studies for screening. After title/abstract screening, 284 studies were excluded, leaving 34 for full-text screening. Six studies were subsequently excluded and 28 studies were selected for data extraction and presentation in Table 2.
Inclusion/exclusion criteria
Studies were included if they were primary articles from the past 5 years, focused on pediatric cohorts, and performed both proteomics and metabolomics on patient samples. Studies that reported proteomics, metabolomics and additional omics analyses (e.g., transcriptomics) were included. Data extracted from the 28 included studies were: publication year, disease/topic studied, total number of participants, participant age range, sample type(s), analytical approaches for proteomics and metabolomics, multi-omics integration stage, and main analyses.
Studies were excluded if the omics analyses were not pediatric-focused, did not perform both proteomics and metabolomics, were performed outside the 5-year window, or were not primary articles. Of the 28 studies included, all but two performed omics analyses on pediatric cohorts exclusively. The remaining two, while still pediatric-focused, used either the parents or healthy adults as control comparators. Of note, a limitation of our search strategy is that restricting terms like “integration” or “multi-omics” to study titles may exclude pediatric studies that performed data integration but used alternative terminology, potentially underestimating the breadth of the field.
Observed patterns in pediatric multi-omics studies
Taken together, the 28 studies listed in Table 2 reveal a field shaped as much by the practical constraints of pediatric research as by methodological maturity. The dominant integration strategy was late integration, in which separate omics analyses were carried out and conclusions compared or correlated post hoc. A shift toward intermediate integration, which captures cross-omic covariance rather than simple agreement between separate analyses, was evident in the most recent studies (2024–2026), consistent with a trend for newly developed integration tools to employ the intermediate integration approach83. Future analyses should aim to continue incorporating intermediate integrative approaches (alongside late integration) in multi-omics pediatric research, as intermediate approaches preserve inter-omics interactions that late integration approaches omit. Nevertheless, late integration (e.g., at the pathway level) retains value for biological interpretation and for corroborating conclusions: when two independently analyzed omics layers converge on the same finding, confidence in that conclusion is inherently strengthened.
Six studies from Table 2 employed intermediate integration, with the most commonly used tools being MOFA and DIABLO (both described in detail in the Tools section), applied across three studies. Osama et al.139 was notable for applying both MOFA and DIABLO to plasma samples from 53 pediatric rhabdomyosarcoma patients and for identifying co-varying proteomic and metabolomic signatures associated with disease. While it was the only intermediate integration study in a disease-focused, small-cohort setting, the authors demonstrated that these methods can be deployed outside of large consortium contexts.
Two HELIX cohort studies also employed latent factor approaches. Wang et al.140 applied MOFA to characterize shared proteomic and metabolomic variation in obesity (n = 1041), and Wang et al.141 used mixOmics latent component analysis to investigate multi-omics correlates of telomere length (n = 1001). Stratakis et al.142 used SNFtool for similarity network fusion to identify metabolic sub-phenotypes in obesity and metabolic dysfunction (n = 863), Goodrich et al.143 applied the LUCID framework to MAFLD (n = 420) using intermediate integration as one stage within a broader multi-stage analytical design, as did Mousavian et al.144 who employed DIABLO within a similar, multi-stage pipeline. Notably, three of the six intermediate integration studies drew from the HELIX cohort—a large pediatric initiative linking environmental exposures to multi-omics profiles—underscoring that the observed shift toward methodological advancement in pediatric multi-omics was partially driven by a single well-resourced consortium rather than broad field-wide adoption.
Most integrated proteo-metabolomic studies skew towards respiratory (RSV pneumonia, community-acquired pneumonia, tuberculosis), metabolic (MAFLD, metabolic dysfunction, obesity), and neurological (autism, infantile epileptic spasms, cerebral palsy) conditions. Plasma or serum were the most common samples collected, most likely due to their accessibility, the whole-body sampling nature of blood, as well as pediatric-specific sampling constraints15,145,146. As biofluids, blood/plasma/serum samples reflect systemic, whole-body signals rather than signals localized to the disease-affected tissue146. Ultimately, whether plasma sampling is optimal depends on a variety of factors such as the research goals, the spread (or localization) of the signal/disease, and any accessibility barriers. In the 28 extracted studies, explicit justification for sample type selection (beyond accessibility) was rarely provided; disease-site sampling (e.g., nasal tissue for rhinosinusitis147) was rare, meaning that studies skew toward systemic results rather than focal signals.
Most (15/28) cohorts are small to moderate in size (n < 100), while 6/28 were large-cohort studies (n > 800). All but 1 of the 6 large-cohort studies drew from the same sample pool (the HELIX study) containing methylomics, transcriptomics, serum metabolomics, urinary metabolomics, and plasma proteomics on up to 1301 children148. The studies with the smallest cohorts almost exclusively utilized late integration, possibly reflecting insufficient data to train latent factor models and the need to avoid curse-of-dimensionality problems. The age range of these cohorts skews towards younger children and infants, with middle adolescents (14–17) being underrepresented. While this may reflect disease epidemiology (e.g., some conditions are more prevalent in early childhood), it underscores the current lack of proteomic and metabolomic analyses of this age group.
LC-MS/MS was the most frequently used method for assessing proteomic and metabolomic samples, with targeted proteomic approaches such as Olink, SomaScan, and Luminex being mostly confined to the large-cohort studies. Metabolomics was almost exclusively assessed using LC-MS variants (UPLC-, UHPLC-, HILIC-MS), with few reports utilizing alternative approaches such as GC-MS and NMR. The platform monoculture (LC-MS) for proteomics and metabolomics analyses means that the field’s coverage reflects not only the strengths of this technology, but also potential blind spots.
Beyond integration strategy and platform choice, the included studies also varied in study design, analytical depth, and cross-platform validation. Study designs were mainly cross-sectional (19/28), with nine employing a longitudinal approach. As most designs capture only a single timepoint, few studies track molecular changes across development, which highlights a notable gap given the significant physiological and developmental changes in pediatric populations. Reported feature counts and analytical approaches were inconsistent across studies and not always provided, so the analytical depth values summarized in Table 2 should be interpreted as approximate. Proteomic analytical depth varied substantially, ranging from ~35 to ~6400 proteins, with untargeted approaches more common than targeted (~18:10 untargeted:targeted, of studies that reported a designation). Metabolomic coverage showed a similar spread, ranging from ~45 to ~3400 metabolites, with targeted and untargeted approaches also fairly evenly represented across studies (~12:14 untargeted:targeted, of studies that reported a designation). Studies did not cross-validate proteomic or metabolomic findings using an independent analytical platform, nor did they replicate full results in independent pediatric cohorts. For example, the study by Huang et al.149 on RSV pneumonia had a dedicated validation cohort but only validated two metabolites using the same analytical platform. Therefore, cross-platform and cross-cohort reproducibility of proteo-metabolomic findings in pediatric populations remains largely unestablished.
Another trend evident in these studies was the use of multi-omics classifiers as a recurring output. Several studies built multi-omics predictive classifiers that were trained using supervised learning and featured some form of dimensionality reduction technique (e.g., feature selection by ranked variable importance scores; reference sections on Supervised learning and Dimensionality reduction). The classifiers were commonly used to discriminate between disease/healthy states, classify subtypes, or distinguish between active and remission states (e.g., Table 2, Crohn’s disease). A lack of true external validation was noted, with only a few studies from Table 2 featuring dedicated internal validation cohorts. As a result, these models carry a risk of over-estimating performance and generalizability beyond the original cohort, although we recognize the difficulty in performing external validation, especially in rare diseases with small cohorts. Nevertheless, validation using independent data remains crucial to the correction of biases and generalizability of results111.
Collectively, the reviewed studies demonstrate that pediatric proteo-metabolomic integration is an emerging but structurally constrained field. Progress has been made across a range of clinically relevant conditions (respiratory, metabolic, and neurological) and the methodological toolkit introduced in the preceding sections (intermediate integration methods, supervised classifiers, dimensionality reduction) is gradually being adopted. The field’s most significant structural limitations include cohort sizes, which were predominantly small and single-site, underrepresentation of adolescent subjects, minimal evidence of external validation, and the concentration on a single large cohort (HELIX) rather than multiple cohorts across independent groups. Addressing these gaps will require not only the continued development of the computational tools and integration frameworks described in this review, but also dedicated investment in pediatric cohort infrastructure, specifically longitudinal, age-stratified, multi-site studies that can generate the sample sizes and diversity necessary to build and validate robust multi-omics models of pediatric health and disease.
Challenges and future outlooks
Multi-omics integration is an exciting next step toward providing a more comprehensive view of molecular profiles by incorporating the interconnectedness between biological layers. To normalize multi-omics workflows, we need to overcome the current barriers to performing these studies and their clinical implementation. The first problem is the coding barrier: many multi-omics research tools require coding experience, especially when employing deep learning algorithms, hindering research groups with reduced coding abilities but high expertise in other areas. User-friendly tools like OmicsAnalyst and OmicsNet represent early progress toward no-code workflows, though DL-capable equivalents remain absent. Agentic AI coding tools have the potential to further lower the coding barrier; however, their application in multi-omics workflows remains largely unvalidated. Human verification of code correctness and output validity is essential, given the known limitations of LLM-based tools, including hallucination and gaps in domain-specific knowledge150.
From a technical perspective, the high dimensionality of omics datasets in conjunction with the use of DL algorithms has created a demand for powerful computational tools. The need for substantial computational power to accomplish these analyses can be challenging for research groups lacking this infrastructure. The creation of more computationally efficient DL algorithms and improvements in cloud-based computing (both in capabilities and in reduction of cost) will facilitate complex analyses and reduce both effort and time151. Additionally, the black-box nature of DL approaches, i.e., their limited interpretability and explainability, poses a challenge. Although neural networks are extremely powerful and are better at prediction and classification tasks compared to shallow learning algorithms, their hidden layer architecture renders their interpretability and explainability well below that of SL algorithms.
Focusing on clinical implementation, multi-omics workflows need to become more streamlined from compound detection to biological interpretation. MetaboAnalyst [v6.0] takes a step in this direction by offering a platform that performs end-to-end analyses on metabolomics datasets, from processing raw MS spectra to biological interpretation136. However, streamlined analyses alone are insufficient. A candidate biomarker must first clear several hurdles. First, analytical validation, showing that the assay measures the intended analyte accurately and reproducibly152. Second, clinical validation, showing that the measurement associates with the condition or outcome in independent samples/cohorts152. Third, clinical utility, showing that use of the biomarker test actually improves patient outcomes and/or helps deliver cost-effective care152. The assay must also be authorized for clinical use under the applicable regulatory framework (e.g., FDA clearance or approval in the United States, Health Canada licensing, or conformity assessment under the EU IVDR), and involved laboratories must hold appropriate certifications (e.g., CLIA/CAP, ISO 15189)153,154. Together, these validation, regulatory, and accreditation requirements represent important barriers to routine clinical implementation of multi-omics-derived biomarkers.
Finally, there is a need to expand databases and repositories. This is particularly problematic for metabolomics, since most metabolites that are detected remain unidentified and metabolomics databases are largely incomplete52. Furthermore, molecular sub-phenotypes remain undefined for many diseases, especially multi-omics sub-phenotypes. Improving databases and expanding repositories of disease-specific sub-phenotypes will facilitate personalized medicine approaches, should such profiles be validated, allowing treatment to be tailored to each patient’s molecular phenotype and treatment tolerances. An example of this principle is emerging in pediatric septic shock, where transcriptomic-based endotypes A and B show differing responses to steroid treatment155. The establishment of age- and developmental-specific molecular profiles will aid in the creation of pediatric-specific guidelines and facilitate personalized approaches.
Expanding these resources also depends on adherence to the FAIR principles, which emphasize that scientific data should be Findable, Accessible, Interoperable, and Reusable156. In practice, this means depositing proteomic data in repositories such as PRIDE or MassIVE, and metabolomic data in MetaboLights or Metabolomics Workbench157–159. That said, deposition is only informative when accompanied by sufficient metadata, and reporting frameworks already exist for this purpose: the Minimum Information About a Proteomics Experiment (MIAPE) for proteomics, and the Metabolomics Standards Initiative (MSI) minimum reporting standards for metabolomics160,161. The same discipline applies to code, where dependency management (e.g., renv, which tracks exact R package versions) together with version-controlled, publicly shared repositories such as GitHub allows published code to be re-run by others162. These practices ultimately improve reproducibility and facilitate future meta-analyses and external validation.
Another structural challenge specific to pediatric multi-omics is the absence of a large, multi-site, longitudinal cohort infrastructure analogous to adult biobanks163,164. Addressing this will require coordinated investment across pediatric research networks, not only to generate the sample sizes needed for robust integration and model validation, but also to capture age-stratified data, which will be necessary to distinguish developmental variation from true disease signals across the full pediatric age/developmental spectrum.
Emerging analytical modalities may further refine multi-omics integration. Single-cell proteomics enables protein profiling at single-cell resolution, capturing potential cellular heterogeneity that is otherwise masked by bulk measurements24. Similarly, spatial proteomics and spatial metabolomics preserve the tissue context of molecular signals, allowing protein and metabolite distributions to be mapped directly onto tissue architecture25,165. Incorporating these spatially- and cellularly-resolved layers into pediatric multi-omics frameworks may help disentangle developmental and disease-specific signals that are otherwise averaged out in bulk analyses; however, these approaches require samples matched to the tissue(s) of interest, which are more invasive and may therefore be more difficult to obtain from pediatric patients166.
Conclusion
The goal of this review was to provide a structured framework for understanding and applying proteomics-metabolomics integration in pediatric research, spanning foundational concepts, available tools, and the current state of the literature. Our review of the 28 pediatric proteo-metabolomic studies published in the past 5 years reveals a field that is progressing but structurally constrained. Late integration remains the dominant approach, though there is evidence of a shift toward intermediate integration methods such as MOFA and DIABLO, particularly in larger-cohort settings. This shift is encouraging, as intermediate approaches capture cross-omic covariance that late integration inherently misses. However, several pitfalls and shortcomings must be kept in mind, such as overfitting, data leakage, batch effects, and a lack of external validation. The methodological advances observed are largely concentrated within a single consortium (HELIX) rather than distributed across independent groups, suggesting that this assessment of progress may overstate its breadth. Adolescent populations remain systematically underrepresented, and the computational toolkit, although maturing, has not yet been matched by the cohort infrastructure needed to fully leverage it. Moving forward, progress will require coordinated investment in large, longitudinal, multi-site pediatric cohorts that are age-stratified across the full developmental spectrum, alongside continued standardization of integration workflows and rigorous validation practices. Ultimately, the convergence of methodological maturity, cohort infrastructure, and clinical collaboration will determine whether multi-omics profiling can transition from a research tool to a validated contributor to personalized medicine in pediatric health and disease.
Author contributions
A.M.M.: conceptualization, writing—original draft, writing—editing and revisions, figure preparation. D.T.: writing—editing and revisions. M.D.: writing—editing and revisions, supervision. D.B.O.: writing—editing and revisions. DDF: conceptualization, writing—editing and revisions, supervision. All authors read and approved the final manuscript.
Peer review
Peer review information
Communications Medicine thanks all reviewer(s) for their contribution to the peer review of this work.
Funding
D.D.F. is funded by a GSK Chair in Clinical Pharmacology (Western University), an AMOSO Innovation Grant and a Donor Legacy Gift provided via the Children’s Health Foundation (https://childhealth.ca/ accessed on 24 December 2025).
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Frieden, T. R. Evidence for health decision making—beyond randomized, controlled trials. N. Engl. J. Med.377, 465–475 (2017). [DOI] [PubMed] [Google Scholar]
- 2.Agarwal, A., Rochwerg, B. & Sevransky, J. E. 21st century evidence: randomized controlled trials versus systematic reviews and meta-analyses. Crit. Care Med49, 2001–2002 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ioannidis, J. P. A. Why most published research findings are false. PLoS Med.2, e124 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Matheny Antommaria, A. H., Kelleher, M. & Peterson, R. J. Quality of evidence and strength of recommendations in American Academy of Pediatrics’ Guidelines. Pediatrics155, e2024067836 (2025). [DOI] [PubMed]
- 5.Morgan, K. M., Gaines, B. A. & Leeper, C. M. Pediatric trauma resuscitation practices. Curr. Trauma Rep.8, 160–171 (2022). [Google Scholar]
- 6.Schlapbach, L. J. et al. International consensus criteria for pediatric sepsis and septic shock. JAMA331, 665–674 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.National Research Council Committee. Toward Precision Medicine: Building a Knowledge Network for Biomedical Research and a New Taxonomy of Disease. in A Framework for Developing a New Taxonomy of Disease (National Academies Press (US), National Academy of Sciences, 2011). [PubMed]
- 8.Hasin, Y., Seldin, M. & Lusis, A. Multi-omics approaches to disease. Genome Biol.18, 83 (2017). This review discusses the major omics layers, integration strategies, and how they can be combined to study disease mechanisms, which are topics central to our work here. [DOI] [PMC free article] [PubMed]
- 9.Chandra Sekar, P. K. & Veerabathiran, R. Systems immunology meets clinical translation: Multi-omic approaches to predict therapy response in cancer and autoimmune disease. Clin. Immunol. Commun.9, 12–22 (2026). [Google Scholar]
- 10.Acharya, D. & Mukhopadhyay, A. A comprehensive review of machine learning techniques for multi-omics data integration: challenges and applications in precision oncology. Brief. Funct. Genom.23, 549–560 (2024). [DOI] [PubMed] [Google Scholar]
- 11.Reel, P. S., Reel, S., Pearson, E., Trucco, E. & Jefferson, E. Using machine learning approaches for multi-omics data analysis: A review. Biotechnol. Adv.49, 107739 (2021). [DOI] [PubMed] [Google Scholar]
- 12.Verhelst, S. et al. The promise of omics approaches for pediatric drug development. in Essentials of Translational Pediatric Drug Development (eds. Gasthuys, E., Allegaert, K., Dossche, L. & Turner, M.) 257–280 (Academic Press, 2024).
- 13.Chiu, C.-Y. et al. Metabolomics reveals dynamic metabolic changes associated with age in early childhood. PLOS ONE11, e0149823 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Francis, E. C. et al. Metabolomic profiles in childhood and adolescence are associated with fetal overnutrition. Metabolites12, 265 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Peplow, C. et al. Blood draws up to 3% of blood volume in clinical trials are safe in children. Acta Paediatr.108, 940–944 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Katz, A. L. et al. Informed consent in decision-making in pediatric practice. Pediatrics138, e20161485 (2016). [DOI] [PubMed]
- 17.Denny, J. C. & Collins, F. S. Precision medicine in 2030—seven ways to transform healthcare. Cell184, 1415–1419 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Wishart, D. S. et al. HMDB: a knowledgebase for the human metabolome. Nucleic Acids Res.37, D603–D610 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Hasanzad, M. et al. Precision medicine journey through the omics approach. J. Diab. Metab. Disord.21, 881–888 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Guo, T., Steen, J. A. & Mann, M. Mass-spectrometry-based proteomics: from single cells to clinical applications. Nature638, 901–911 (2025). This review summarizes the current state of MS-based proteomics, covering a variety of technical topics such as sample preparation, instrumentation, data-acquisition strategies, and extending towards clinical applications. Therefore, it serves as an important reference in our proteomics section. [DOI] [PubMed] [Google Scholar]
- 21.Fan, S. & Zeng, S. Plasma proteomics in pediatric patients with sepsis–hopes and challenges. Clin. Proteomics22, 10 (2025). [DOI] [PMC free article] [PubMed]
- 22.Basu, A. A. & Zhang, X. Quantitative proteomics and applications in covalent ligand discovery. Front. Chem. Biol.3, 1352676 (2024). [DOI] [PMC free article] [PubMed]
- 23.Gheybi, E., Hosseinzadeh, P., Tayebi-Khorrami, V., Rostami, M. & Soukhtanloo, M. Proteomics in decoding cancer: A review. Clin. Chim. Acta574, 120302 (2025). [DOI] [PubMed] [Google Scholar]
- 24.Momenzadeh, A. & Meyer, J. G. Single-cell proteomics using mass spectrometry. Cell Genom.5, 100973 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Wu, M. et al. Spatial proteomics: unveiling the multidimensional landscape of protein localization in human diseases. Proteome Sci.22, 7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Zhang, N., Wu, J. & Zheng, Q. Chemical proteomics approaches for protein post-translational modification studies. Biochim. Biophys. Acta (BBA) - Proteins Proteom.1872, 141017 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Cheng, X. Understanding signal transduction through functional proteomics. Expert Rev. Proteom.2, 103–116 (2005). [DOI] [PubMed] [Google Scholar]
- 28.Figeys, D. Functional proteomics: mapping protein-protein interactions and pathways. Curr. Opin. Mol. Ther.4, 210–215 (2002). [PubMed] [Google Scholar]
- 29.Fraser, D. D. et al. Functional mass spectrometry indicates anti-protease and complement activity increase with COVID-19 severity. Exp. Biol. Med.250, 10308 (2025). [DOI] [PMC free article] [PubMed]
- 30.Goetze, S. et al. Simultaneous targeted and discovery-driven clinical proteotyping using hybrid-PRM/DIA. Clin. Proteom.21, 26 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Aebersold, R. & Mann, M. Mass spectrometry-based proteomics. Nature422, 198–207 (2003). [DOI] [PubMed] [Google Scholar]
- 32.Suhre, K., McCarthy, M. I. & Schwenk, J. M. Genetics meets proteomics: perspectives for large population-based studies. Nat. Rev. Genet.22, 19–37 (2021). This review covers affinity-based proteomics platforms such as SomaScan and Olink, and how protein measurements behave at population scale. It serves as an important reference discussing the trade-offs between affinity and mass spectrometry approaches. [DOI] [PubMed] [Google Scholar]
- 33.Lou, R. & Shui, W. Acquisition and analysis of DIA-based proteomic data: a comprehensive survey in 2023. Mol. Cell. Proteom.23, 100712 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Sticker, A., Goeminne, L., Martens, L. & Clement, L. Robust summarization and inference in proteome-wide label-free quantification. Mol. Cell Proteom.19, 1209–1219 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Shuken, S. R. An introduction to mass spectrometry-based proteomics. J. Proteome Res.22, 2151–2171 (2023). [DOI] [PubMed] [Google Scholar]
- 36.Gold, L. et al. Aptamer-based multiplexed proteomic technology for biomarker discovery. PLoS ONE5, e15004 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Gold, L., Walker, J. J., Wilcox, S. K. & Williams, S. Advances in human proteomics at high scale with the SOMAscan proteomics platform. N. Biotechnol.29, 543–549 (2012). [DOI] [PubMed] [Google Scholar]
- 38.Martínez-Moreno, J.M., Llamas-Urbano, A., Barbarroja, N. & Pérez-Sánchez, C. Proteomics by qPCR using the proximity extension assay (PEA). In Methods in Molecular Biology, Vol. 2929, 129–142 (Springer US, 2025). [DOI] [PubMed]
- 39.Wik, L. et al. Proximity extension assay in combination with next-generation sequencing for high-throughput proteome-wide analysis. Mol. Cell. Proteom.20, 100168 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Assarsson, E. et al. Homogenous 96-Plex PEA immunoassay exhibiting high sensitivity, specificity, and excellent scalability. PLoS ONE9, e95192 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Singh, B., Tzoulaki, I. & Mayr, M. Precision medicine requires precision proteomics: discordance between proteomic and clinical assays in UK Biobank. Cardiovasc. Res.121, 2293–2295 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Rooney, M. R. et al. Correlations within and between highly multiplexed proteomic assays of human plasma. Clin. Chem.71, 677–687 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Smith, J. G. & Gerszten, R. E. Emerging affinity-based proteomic technologies for large-scale plasma profiling in cardiovascular disease. Circulation135, 1651–1664 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Smith, L. M. & Kelleher, N. L. Proteoform: a single term describing protein complexity. Nat. Methods10, 186–187 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Aebersold, R. et al. How many human proteoforms are there? Nat. Chem. Biol.14, 206–214 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Sissala, N. et al. Comparative evaluation of Olink Explore 3072 and mass spectrometry with peptide fractionation for plasma proteomics. Commun. Chem.8, 327 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Roberts, D. S. et al. Top-down proteomics. Nat. Rev. Methods Prim.4, 38 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Bauermeister, A., Mannochio-Russo, H., Costa-Lotufo, L. V., Jarmusch, A. K. & Dorrestein, P. C. Mass spectrometry-based metabolomics in microbiome investigations. Nat. Rev. Microbiol.20, 143–160 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Becker, S., Kortz, L., Helmschrodt, C., Thiery, J. & Ceglarek, U. LC–MS-based metabolomics in the clinical laboratory. J. Chromatogr. B883-884, 68–75 (2012). [DOI] [PubMed] [Google Scholar]
- 50.Dudley, E., Yousef, M., Wang, Y. & Griffiths, W. J. Targeted metabolomics and mass spectrometry. Adv. Protein Chem. Struct. Biol.80, 45–83 (2010). [DOI] [PubMed] [Google Scholar]
- 51.Johnson, C. H., Ivanisevic, J. & Siuzdak, G. Metabolomics: beyond biomarkers and towards mechanisms. Nat. Rev. Mol. Cell Biol.17, 451–459 (2016). This review argues that metabolomics should move beyond biomarker lists towards mechanistic insight and frames the metabolome as the omics layer closest to phenotype. It provides the conceptual justification for including metabolomics in a multi-omics design rather than treating it as a biomarker screen. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Sahu, D., Matusa, A. M., DiBattista, A., Urquhart, B. L. & Fraser, D. D. Mass spectrometry-based metabolomics in pediatric health and disease. Metabolites16, 49 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.UniProt Consortium UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res.53, D609–D617 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Smith, C. A. et al. METLIN: a metabolite mass spectral database. Ther. Drug Monit.27, 747–751 (2005). [DOI] [PubMed] [Google Scholar]
- 55.Wishart, D. S. et al. HMDB 5.0: the Human Metabolome Database for 2022. Nucleic Acids Res50, D622–d631 (2022). This is the Human Metabolome Database 5.0 paper, describing the reference resource used for metabolite annotation. Many mass spectrometry-based metabolomics studies depend on this database, and its incompleteness underlies our argument that most metabolites remain unidentified. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Gika, H.G., Theodoridis, G. & Wilson, I.D. Metabolic Profiling: A Perspective on the Current Status, Challenges, and Future Directions. in Methods in Molecular Biology, Vol. 2891, 1–14 (Springer US, 2025). [DOI] [PubMed]
- 57.Emwas, A. H. et al. NMR spectroscopy for metabolomics research. Metabolites9, 123 (2019). [DOI] [PMC free article] [PubMed]
- 58.Raja, G., Jung, Y., Jung, S. H. & Kim, T.-J. 1H-NMR-based metabolomics for cancer targeting and metabolic engineering—a review. Process Biochem.99, 112–122 (2020). [Google Scholar]
- 59.Correa-Aguila, R., Alonso-Pupo, N. & Hernandez-Rodriguez, E. W. Multi-omics data integration approaches for precision oncology. Mol. Omics18, 469–479 (2022). [DOI] [PubMed] [Google Scholar]
- 60.Sanches, P. H. G., De Melo, N. C., Porcari, A. M. & De Carvalho, L. M. Integrating molecular perspectives: strategies for comprehensive multi-omics integrative data analysis and machine learning applications in transcriptomics, proteomics, and metabolomics. Biology13, 848 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Elgar, G. Editorial: What is Systems Biology?. Brief. Funct. Genom. Proteom.7, 237–238 (2008). [DOI] [PubMed] [Google Scholar]
- 62.Zhu, B. et al. Integrating clinical and multiple omics data for prognostic assessment across human cancers. Sci. Rep.7, 16954 (2017). [DOI] [PMC free article] [PubMed]
- 63.Abdelhamid, S. S. et al. Multi-omic admission-based prognostic biomarkers identified by machine learning algorithms predict patient recovery and 30-day survival in trauma patients. Metabolites12, 774 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Alves, P. et al. Advancement in protein inference from shotgun proteomics using peptide detectability. Pac. Symp. Biocomput12, 409–420 (2007). [PubMed] [Google Scholar]
- 65.Schork, K., Turewicz, M., Uszkoreit, J., Rahnenführer, J. & Eisenacher, M. Characterization of peptide-protein relationships in protein ambiguity groups via bipartite graphs. PLOS ONE17, e0276401 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Lin, A., See, D., Fondrie, W. E., Keich, U. & Noble, W. S. Target-decoy false discovery rate estimation using Crema. Proteomics24, 2300084 (2024). [DOI] [PubMed] [Google Scholar]
- 67.Välikangas, T., Suomi, T. & Elo, L. L. A systematic evaluation of normalization methods in quantitative label-free proteomics. Brief. Bioinform19, 1–11 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Johnson, W. E., Li, C. & Rabinovic, A. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics8, 118–127 (2007). [DOI] [PubMed] [Google Scholar]
- 69.Chen, Q. et al. Protein-level batch-effect correction enhances robustness in MS-based proteomics. Nat. Commun.16, 9735 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Gardner, M. L. & Freitas, M. A. Multiple imputation approaches applied to the missing value problem in bottom-up proteomics. Int. J. Mol. Sci.22, 9650 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kanamori, T., Fujiwara, S. & Takeda, A. Breakdown point of robust support vector machines. Entropy19, 83 (2017). [Google Scholar]
- 72.Argelaguet, R. et al. Multi-Omics Factor Analysis—a framework for unsupervised integration of multi-omics data sets. Mol. Syst. Biol.14, e8124 (2018). This is the original MOFA paper which outlines the tool and its functionality. MOFA is a central tool for multi-omics integration, and is mentioned many times throughout this manuscript. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Chamberlain, C. A., Rubio, V. Y. & Garrett, T. J. Impact of matrix effects and ionization efficiency in non-quantitative untargeted metabolomics. Metabolomics15, 135 (2019). [DOI] [PubMed] [Google Scholar]
- 74.Do, K. T. et al. Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies. Metabolomics14, 128 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Dunn, W. B. et al. Procedures for large-scale metabolic profiling of serum and plasma using gas chromatography and liquid chromatography coupled to mass spectrometry. Nat. Protoc.6, 1060–1083 (2011). [DOI] [PubMed] [Google Scholar]
- 76.Baião, A. R. et al. A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches. Brief. Bioinform.26, bbaf355 (2025). [DOI] [PMC free article] [PubMed]
- 77.Ritchie, M. D., Holzinger, E. R., Li, R., Pendergrass, S. A. & Kim, D. Methods of integrating data to uncover genotype-phenotype interactions. Nat. Rev. Genet.16, 85–97 (2015). This paper introduced the formal taxonomy of data integration strategies, distinguishing meta-dimensional from multi-staged analyses. It is the origin of the concatenation- (early), transformation- (intermediate), and model-based (late) integration vocabulary used throughout this manuscript. [DOI] [PubMed] [Google Scholar]
- 78.Sathyanarayanan, A. et al. A comparative study of multi-omics integration tools for cancer driver gene identification and tumour subtyping. Brief. Bioinform.21, 1920–1936 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Abdelaziz, E. H., Ismail, R., Mabrouk, M. S. & Amin, E. Multi-omics data integration and analysis pipeline for precision medicine: systematic review. Comput. Biol. Chem.113, 108254 (2024). [DOI] [PubMed] [Google Scholar]
- 80.Picard, M., Scott-Boyer, M. P., Bodein, A., Perin, O. & Droit, A. Integration strategies of multi-omics data for machine learning analysis. Comput. Struct. Biotechnol. J.19, 3735–3746 (2021). This review systematically compares early, intermediate, mixed, and late integration strategies for machine learning analyses, along with their respective dimensionality and information loss penalties. These concepts are central to our section “Integrating Omics: Multi-omics.” [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Zitnik, M. et al. Machine learning for integrating data in biology and medicine: principles, practice, and opportunities. Inf. Fusion50, 71–91 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Berisha, V. et al. Digital medicine and the curse of dimensionality. npj Digit. Med.4, 153 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Athieniti, E. & Spyrou, G. M. A guide to multi-omics data collection and integration for translational medicine. Comput. Struct. Biotechnol. J.21, 134–149 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Belharar, F.Z., Belharar, O., Retal, S., Kharmoum, N. & Ziti, S. Comparative Analysis of Multi-Omic Integration Approaches for Tumor Subtype Classification: a Systematic Literature Review. 189–199 (Springer Nature Switzerland, 2026).
- 85.Benkirane, H., Pradat, Y., Michiels, S. & Cournède, P.-H. CustOmics: a versatile deep-learning based strategy for multi-omics integration. PLOS Comput. Biol.19, e1010921 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Akinci, T.C., Topsakal, O. & Akbas, M.I. Machine learning methods from shallow learning to deep learning. In Shallow Learning vs. Deep Learning: A Practical Guide for Machine Learning Solutions (eds. Ertuğrul, Ö.F., Guerrero, J.M. & Yilmaz, M.) 1–28 (Springer Nature Switzerland, 2024).
- 87.Babu, M. & Snyder, M. Multi-omics profiling for health. Mol. Cell. Proteom.22, 100561 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Kernbach, J.M. & Staartjes, V.E. Foundations of Machine Learning-based Clinical Prediction Modeling: Part II—Generalization and Overfitting. 15–21 (Springer International Publishing, 2022). [DOI] [PubMed]
- 89.Li, J. & Wang, Y. nPCA: a linear dimensionality reduction method using a multilayer perceptron. Front. Genet.14, 1290447 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Nanga, S. et al. Review of dimension reduction methods. J. Data Anal. Inf. Process.09, 189–231 (2021). [Google Scholar]
- 91.Meng, C. et al. Dimension reduction techniques for the integrative analysis of multi-omics data. Brief. Bioinform.17, 628–641 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Cantini, L. et al. Benchmarking joint multi-omics dimensionality reduction approaches for the study of cancer. Nat. Commun.12, 124 (2021). [DOI] [PMC free article] [PubMed]
- 93.Saito, T. & Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE10, e0118432 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Steyerberg, E. W. et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology21, 128–138 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Vickers, A. J. & Elkin, E. B. Decision curve analysis: a novel method for evaluating prediction models. Med. Decis. Mak.26, 565–574 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Barredo Arrieta, A. et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion58, 82–115 (2020). [Google Scholar]
- 97.Sidak, D., Schwarzerová, J., Weckwerth, W. & Waldherr, S. Interpretable machine learning methods for predictions in systems biology from omics data. Front. Mol. Biosci.9, 926623 (2022). [DOI] [PMC free article] [PubMed]
- 98.Khaire, U. M. & Dhanalakshmi, R. Stability of feature selection algorithm: a review. J. King Saud. Univ. Comput. Inf. Sci.34, 1060–1073 (2022). [Google Scholar]
- 99.Broadhurst, D. I. & Kell, D. B. Statistical strategies for avoiding false discoveries in metabolomics and related experiments. Metabolomics2, 171–196 (2006). [Google Scholar]
- 100.Lundberg, S.M. & Lee, S.-I. A unified approach to interpreting model predictions. In: Proc. 31st International Conference on Neural Information Processing Systems 4768–4777 (Curran Associates Inc., 2017).
- 101.Ribeiro, M.T., Singh, S. & Guestrin, C. “Why should i trust you?”: explaining the predictions of any classifier. In: Proc. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 1135–1144 (ACM, 2016).
- 102.Gramegna, A. & Giudici, P. SHAP and LIME: an evaluation of discriminative power in credit risk. Front. Artif. Intell.4, 752558 (2021). [DOI] [PMC free article] [PubMed]
- 103.Hermosilla, P., Berríos, S. & Allende-Cid, H. Explainable AI for forensic analysis: a comparative study of SHAP and LIME in intrusion detection models. Appl. Sci.15, 7329 (2025). [Google Scholar]
- 104.Kapoor, S. & Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns4, 100804 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Varma, S. & Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinform.7, 91 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Slatni, R. et al. Machine learning for immune biomarkers in severe mental illness: a systematic review. Neurosci. Appl.5, 107003 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Soneson, C., Gerster, S. & Delorenzi, M. Batch effect confounding leads to strong bias in performance estimates obtained by cross-validation. PLoS One9, e100335 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Li, Y., Herold, T., Mansmann, U. & Hornung, R. Does combining numerous data types in multi-omics data improve or hinder performance in survival prediction? Insights from a large-scale benchmark study. BMC Med. Inform. Decis. Mak.24, 244 (2024). [DOI] [PMC free article] [PubMed]
- 109.Leek, J. T. et al. Tackling the widespread and critical impact of batch effects in high-throughput data. Nat. Rev. Genet11, 733–739 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Baggerly, K. A., Morris, J. S. & Coombes, K. R. Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments. Bioinformatics20, 777–785 (2004). [DOI] [PubMed] [Google Scholar]
- 111.Jelizarow, M., Guillemot, V., Tenenhaus, A., Strimmer, K. & Boulesteix, A. L. Over-optimism in bioinformatics: an illustration. Bioinformatics26, 1990–1998 (2010). [DOI] [PubMed] [Google Scholar]
- 112.Catalano, M. et al. Navigating cancer complexity: integrative multi-omics methodologies for clinical insights. Clin. Med. Insights Oncol.19, 11795549251384582 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Jiang, Z., Zhang, H., Gao, Y. & Sun, Y. Multi-omics strategies for biomarker discovery and application in personalized oncology. Mol. Biomed.6, 115 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Ogata, H. et al. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res.27, 29–34 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Papaioannou, N. et al. Multi-omics analysis reveals that co-exposure to phthalates and metals disturbs urea cycle and choline metabolism. Environ. Res.192, 110041 (2021). [DOI] [PubMed] [Google Scholar]
- 116.Kozlova, A. et al. PMconv: how to compare proteomes and metabolomes? Int. J. Mol. Sci.27, 5086 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117.Langfelder, P. & Horvath, S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinform.9, 559 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Huang, Y. et al. Integrated proteomics and metabolomics network analysis across different delivery modes in human pregnancy: a pilot study. BMC Preg. Childbirth24, 868 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 119.Qiu, S. et al. Proteome trade-off between primary and secondary metabolism shapes acid stress-induced bacterial exopolysaccharide production. Metab. Eng.91, 254–266 (2025). [DOI] [PubMed] [Google Scholar]
- 120.Antoniewicz, M. R. A guide to 13C metabolic flux analysis for the cancer biologist. Exp. Mol. Med.50, 1–13 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Python Software Foundation. Python. Vol. 3 (Python Software Foundation, 2025).
- 122.R Core Team. R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, 2025).
- 123.Rohart, F., Gautier, B., Singh, A. & Lê Cao, K.-A. mixOmics: an R package for ‘omics feature selection and multiple data integration. PLOS Comput. Biol.13, e1005752 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Singh, A. et al. DIABLO: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinformatics35, 3055–3062 (2019). This is the original DIABLO paper, which outlines the tool and its features. DIABLO is the flagship algorithm in the mixOmics toolset and proves vital for integrating multi-omics datasets in a supervised manner. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 125.Argelaguet, R. et al. MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol.21, 111 (2020). [DOI] [PMC free article] [PubMed]
- 126.Subramanian, A. et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc. Natl. Acad. Sci.102, 15545–15550 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Khatri, P., Sirota, M. & Butte, A. J. Ten years of pathway analysis: current approaches and outstanding challenges. PLoS Comput. Biol.8, e1002375 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Wang, B. et al. Similarity network fusion for aggregating data types on a genomic scale. Nat. Methods11, 333–337 (2014). This is the original SNF paper, outlining the tool and its functionality. SNF is a powerful approach to multi-omics integration, providing clustering functionality through a network approach. [DOI] [PubMed] [Google Scholar]
- 129.Shannon, P. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res.13, 2498–2504 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Ono, K. et al. Cytoscape Web: bringing network biology to the browser. Nucleic Acids Res.53, W203–W212 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Raschka, S., Patterson, J. & Nolet, C. Machine learning in Python: main developments and technology trends in data science, machine learning, and artificial intelligence. Information11, 193 (2020). [Google Scholar]
- 132.Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
- 133.Zhang, X., Xing, Y., Sun, K. & Guo, Y. OmiEmbed: a unified multi-task deep learning framework for multi-omics data. Cancers13, 3047 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Wang, T. et al. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat. Commun.12, 3445 (2021). [DOI] [PMC free article] [PubMed]
- 135.Ewald, J. D. et al. Web-based multi-omics integration using the Analyst software suite. Nat. Protoc.19, 1467–1497 (2024). [DOI] [PubMed] [Google Scholar]
- 136.Pang, Z. et al. MetaboAnalyst 6.0: towards a unified platform for metabolomics data processing, analysis and interpretation. Nucleic Acids Res.52, W398–W406 (2024). This is the original MetaboAnalyst 6.0 paper, describing an end-to-end web platform that takes raw spectra through statistics to pathway interpretation. MetaboAnalyst is a widely used no-code omics platform, and is our main example of no-code omics tools. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 137.Liu, T. et al. PaintOmics 4: new tools for the integrative analysis of multi-omics datasets supported by multiple pathway databases. Nucleic Acids Res.50, W551–W559 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 138.Milacic, M. et al. The Reactome Pathway Knowledgebase 2024. Nucleic Acids Res.52, D672–D678 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 139.Osama, A. et al. Integrative multi-omics profiling of rhabdomyosarcoma subtypes reveals distinct molecular pathways and biomarker signatures. Cells14, 1115 (2025). This study applies both MOFA and DIABLO to plasma from children with rhabdomyosarcoma, recovering co-varying proteomic and metabolomic disease signatures. It is a direct link between our pediatric focus and the modern integration tools discussed earlier in the manuscript. [DOI] [PMC free article] [PubMed]
- 140.Wang, C. et al. Meet-in-the-middle meets multi-omics identifying molecular signatures of environmental drivers of childhood overweight. Environ. Int202, 109630 (2025). [DOI] [PubMed] [Google Scholar]
- 141.Wang, C. et al. The multi-omics signatures of telomere length in childhood. BMC Genom.26, 75 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 142.Stratakis, N. et al. Multi-omics architecture of childhood obesity and metabolic dysfunction uncovers biological pathways and prenatal determinants. Nat. Commun.16, 654 (2025). This study uses similarity network fusion on HELIX data to resolve childhood obesity into molecular subphenotypes. It is a direct link between our pediatric focus and SNF, another modern tool for multi-omics integration. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Goodrich, J. A. et al. Integrating Multi-Omics with environmental data for precision health: a novel analytic framework and case study on prenatal mercury-induced childhood fatty liver disease. Environ. Int.190, 108930 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 144.Mousavian, Z. et al. A multi-omics study reveals pathway-level insights and predictive biomarkers in pediatric TB. Clin. Proteom. 23, 31 (2026). [DOI] [PMC free article] [PubMed]
- 145.Shraim, R. et al. A method for comparing proteins measured in serum and plasma by Olink proximity extension assay. Mol. Cell. Proteom.24, 101000 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 146.Surinova, S. et al. On the development of plasma protein biomarkers. J. Proteome Res.10, 5–16 (2011). [DOI] [PubMed] [Google Scholar]
- 147.Chen, Y. C., Wang, X., Pan, Y. W., Teng, Y. S. & Pan, H. G. Distinct and shared molecular mechanisms in pediatric antrochoanal polyps and chronic rhinosinusitis with nasal polyps: a proteomic and metabolomic integrative analysis. J. Inflamm. Res.18, 4435–4447 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 148.Maitre, L. et al. Human Early Life Exposome (HELIX) study: a European population-based exposome cohort. BMJ Open8, e021311 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 149.Huang, X. et al. Multi-omics analysis reveals underlying host responses in pediatric respiratory syncytial virus pneumonia. iScience26, 106329 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 150.Kuehl, M. et al. BioContextAI is a community hub for agentic biomedical systems. Nat. Biotechnol.43, 1755–1757 (2025). [DOI] [PubMed] [Google Scholar]
- 151.Mienye, I. D. & Swart, T. G. A comprehensive review of deep learning: architectures, recent advances, and applications. Information15, 755 (2024). [Google Scholar]
- 152.Füzéry, A. K., Levin, J., Chan, M. M. & Chan, D. W. Translation of proteomic biomarkers into FDA approved cancer diagnostics: issues and challenges. Clin. Proteom.10, 13 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 153.Han, C.-L. et al. Lessons learned: establishing a CLIA-equivalent laboratory for targeted mass spectrometry assays—navigating the transition from research to clinical practice. Clin. Proteom.21, 12 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 154.Kahles, A. et al. Regulation of laboratory-developed tests and in-house in vitro diagnostic medical devices in the United States and the European Union—a comparative overview. ESMO Open10, 105909 (2025). [DOI] [PMC free article] [PubMed]
- 155.Wong, H. R., Hart, K. W., Lindsell, C. J. & Sweeney, T. E. External corroboration that corticosteroids may be harmful to septic shock endotype A patients. Crit. Care Med.49, e98–e101 (2021). This study discusses external corroboration that supports the notion that corticosteroids may be harmful in patients with pediatric septic shock endotype A. It is strong evidence that omics-defined pediatric endotypes can change treatment decisions, which is a clinical endpoint this review argues multi-omics should reach. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 156.Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data3, 160018 (2016). This paper defines the FAIR principles: findable, accessible, interoperable, and reusable. This is a central reference for our discussion on data deposition and reproducibility, and a precondition for the meta-analyses and external validation the field currently lacks. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 157.Perez-Riverol, Y. Proteomic repository data submission, dissemination, and reuse: key messages. Expert Rev. Proteom.19, 297–310 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 158.Sud, M. et al. Metabolomics Workbench: an international repository for metabolomics data and metadata, metabolite standards, protocols, tutorials and training, and analysis tools. Nucleic Acids Res.44, D463–D470 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 159.Yurekten, O. et al. MetaboLights: open data repository for metabolomics. Nucleic Acids Res.52, D640–D646 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 160.Sumner, L. W. et al. Proposed minimum reporting standards for chemical analysis. Metabolomics3, 211–221 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 161.Taylor, C. F. et al. The minimum information about a proteomics experiment (MIAPE). Nat. Biotechnol.25, 887–893 (2007). [DOI] [PubMed] [Google Scholar]
- 162.Siraji, M. A. & Rahman, M. Primer on reproducible research in R: enhancing transparency and scientific rigor. Clocks Sleep.6, 1–10 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 163.McBain, K. et al. A scoping review of adult NCD-relevant phenotypes measured in today’s large child cohort studies. Pediatr. Res.98, 2058–2072 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 164.Wadhwa, L. Landscape of pediatric biobanking: challenges and current efforts. Biopreserv. Biobank19, 119–123 (2021). [DOI] [PubMed] [Google Scholar]
- 165.Min, X. et al. Spatially resolved metabolomics: from metabolite mapping to function visualising. Clin. Transl. Med.14, e70031 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 166.Weiser, D. A. et al. Progress toward liquid biopsies in pediatric solid tumors. Cancer Metastasis Rev.38, 553–571 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 167.Espinosa, C. A. et al. Multiomics characterization of acute child illness and mortality in Africa and South Asia. Nat. Commun.17, 5171 (2026). [DOI] [PMC free article] [PubMed]
- 168.Jia, C. et al. Multi-omics reveal the potential associations of streptococcus, 13’-hydroxy-alpha-tocopherol and glutathione metabolism in children with chronic rhinosinusitis with nasal polyps. J. Inflamm. Res.19, 567582 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 169.Li, G. et al. Multi-omics and machine learning-based profiling of severity signatures in Mycoplasma pneumoniae infection in children. iScience29, 114861 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 170.Chen, Z. et al. Integrated metabolomics and proteomics analysis in children with cerebral palsy exposed to botulinum toxin-A. Pediatr. Res.98, 2300–2310 (2025). [DOI] [PubMed] [Google Scholar]
- 171.Chen, J. et al. Integrated analysis of proteomics and metabolomics in infantile epileptic spasms syndrome. Sci. Rep.15, 4457 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 172.Tang, Y. et al. Multi-omics analysis of human plasma reveals reprogramming of tryptophan metabolism associated with inflammation in Mycoplasma pneumoniae pneumonia in children. J. Infect.91, 106525 (2025). [DOI] [PubMed]
- 173.Hendrickx, D. M. et al. A multi-omics machine learning classifier for outgrowth of cow’s milk allergy in children. Mol. Omics21, 343–352 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 174.Noone, A. et al. Longitudinal multi-omics analysis of umbilical cord blood and childhood serum in autism. Mol Psychiatry31, 701-713 (2025). [DOI] [PubMed]
- 175.Schomakers, B. V. et al. Integrated multi-omics mapping of mitochondrial dysfunction and substrate preference in Barth syndrome cardiac tissue. EMBO Mol. Med.17, 3227–3246 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 176.Koopman, N. et al. Integrated multi-omics of feces, plasma and urine can describe and differentiate pediatric active Crohn’s Disease from remission. Commun. Med.5, 281 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 177.Osama, A. et al. Integrative multi-omics analysis of autism spectrum disorder reveals unique microbial macromolecules interactions. J. Adv. Res.77, 265–279 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 178.Ye, Q. et al. Comprehensive mapping of saliva by multiomics in children with idiopathic nephrotic syndrome. Nephrology29, 565–578 (2024). [DOI] [PubMed] [Google Scholar]
- 179.Tang, X. et al. A study of genetic heterogeneity in autism spectrum disorders based on plasma proteomic and metabolomic analysis: multiomics study of autism heterogeneity. MedComm4, e380 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 180.Bhayana, S. et al. Multiomics analysis of plasma proteomics and metabolomics of steroid resistance in childhood nephrotic syndrome using a “patient-specific” approach. Kidney Int. Rep.8, 1239–1254 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 181.Wang, Y. et al. Serum-integrated omics reveal the host response landscape for severe pediatric community-acquired pneumonia. Crit. Care27, 79 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 182.Reisz, J. A. et al. Red blood cells from individuals with Lesch-Nyhan syndrome: multi-omics insights into a novel S162N mutation causing hypoxanthine-guanine phosphoribosyltransferase deficiency. Antioxidants12, 1699 (2023). [DOI] [PMC free article] [PubMed]
- 183.Robinson, O. et al. Associations of four biological age markers with child development: a multi-omic analysis in the European HELIX cohort. Elife12, e85104 (2023). [DOI] [PMC free article] [PubMed]
- 184.Fabbri, L. et al. Childhood exposure to non-persistent endocrine disrupting chemicals and multi-omic profiles: a panel study. Environ. Int173, 107856 (2023). [DOI] [PubMed] [Google Scholar]
- 185.Nierves, L. et al. Multi-omic profiling of the leukemic microenvironment shows bone marrow interstitial fluid is distinct from peripheral blood plasma. Exp. Hematol. Oncol.11, 56 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 186.Maitre, L. et al. Multi-omics signatures of the human early life exposome. Nat. Commun.13, 7024 (2022). This study reports multi-omics signatures of the early-life exposome across approximately 1300 children in the HELIX cohort, linking exposures to several omics layers, including the proteome and metabolome. It is the main pediatric multi-omics dataset from which most of the pediatric proteo-metabolomic work derives. [DOI] [PMC free article] [PubMed] [Google Scholar]
