Abstract
Liquid chromatography coupled to high-resolution mass spectrometry (LC–HRMS) is a widely used analytical technique for characterizing the chemical composition of organic samples. Due to its high sensitivity and ability to detect thousands of chemical features in a single run, untargeted LC–HRMS experiments generate highly complex and data-rich datasets that typically require advanced computational methods, including machine learning, for meaningful interpretation. While traditional machine learning approaches have been applied to LC–HRMS data, their performance remains limited for complex tasks. Deep learning has demonstrated improved performance, but both machine and deep learning are often constrained by the complexity and scarcity of labeled LC–HRMS data. Foundation models present a promising new horizon for LC–HRMS data analysis, given their ability to learn transferable representations from large-scale unlabeled data and adapt efficiently to downstream tasks with limited labeled samples. Recent studies have shown that foundation models can outperform conventional machine learning approaches in chemical annotation and molecular property prediction. We envision that foundation models for LC–HRMS data will benefit from the expansion of curated sample repositories and spectral libraries, developing privacy-preserving training strategies, enabling simultaneous modeling of multiple LC–HRMS data types, and improving model explainability.


Introduction
High-resolution mass spectrometry (HRMS) is a widely used technique in analytical chemistry that enables the detection and partial characterization of a vast diversity of molecules in complex mixtures. Coupling liquid chromatography with HRMS (LC–HRMS) has become a dominant analytical method, enabling the analysis of a wide range of compounds with diverse physicochemical properties. Given the widespread adoption of LC–HRMS across multiple scientific fields and the challenges and recent advances in LC–HRMS data analysis (see the next sections), this manuscript focuses on LC–HRMS.
LC–HRMS data consist of a series of mass spectra collected at specific time points during chromatographic separation (i.e., retention time). These spectra contain multiple ionized molecules and/or fragment ions, represented by their mass-to-charge ratios (m/z) and relative intensities. A key advantage of HRMS is its versatility to acquire mass spectra using various acquisition modes, including full-scan MS1, MS2 data-dependent acquisition (DDA), and MS2 data-independent acquisition (DIA). MS1 spectra provide m/z and intensity values for ions produced during primary ionization, while DDA and DIA generate tandem mass spectra (known as MS2 or MS/MS) containing fragment ions derived from precursors selected from previous MS1 scans. The selection of precursor ions differs between the DDA and DIA. In DDA, precursor ions are individually selected for fragmentation based on predefined criteria, such as intensity thresholds, and are subsequently fragmented. In contrast, DIA fragments all ions within predefined m/z isolation windows simultaneously without prior selection of individual precursors. In untargeted analyses, which aim to comprehensively profile a wide range of molecules without predefined targets, LC–HRMS acquisition typically combines MS1 with DDA and/or DIA. As a result, each injection generates an LC–HRMS data file containing thousands of spectra acquired across multiple acquisition modes (MS1, DDA, and DIA).
The large number of spectra typically acquired during untargeted LC–HRMS experiments limits the feasibility of manual data interpretation. Consequently, automated data analysis pipelines, often incorporating machine learning, are used to process these complex datasets and translate them into meaningful chemical information. Several machine learning-based methods have been developed for HRMS data analysis, including methods that classify MS1 peaks as sample’s compounds or instrumental noise, , calculate similarity between different MS2 spectra, , predict molecular properties directly from MS2 spectra, identify compound analogues, and perform de novo structure generation from MS2 data. , Despite these advances, modeling MS1 and MS2 data is a complex task and presents several challenges (see next section), which can lead to suboptimal model performance. Although recent deep learning approaches have demonstrated improved performance over traditional machine learning methods in tasks such as classification and structure annotation, they often depend on labeled datasets (e.g., MS2 spectra labeled with compound names or other descriptors). This dependency is a major limitation, as the majority of spectra generated in untargeted LC–HRMS experiments remain unlabeled.
Foundation models represent a promising approach to address both the scarcity of labeled datasets and the complexity of HRMS data. , Foundation models are often pretrained using self-supervised learning, allowing them to leverage large amounts of unlabeled data to learn meaningful spectral representations by creating labels based on the data itself. After this pretraining step, foundation models can be adapted to a wide range of downstream tasks, including compound classification and compound annotation, by training on a labeled dataset. The continued training of an already pretrained (foundation) model is called fine-tuning. As a result of the learned spectral representations, foundation models typically require substantially less labeled data during fine-tuning for downstream tasks. In other fields, such as natural language processing, computer vision, and molecular sciences, foundation models have already been successfully developed and deployed. A well-known example is ChatGPT, which was trained on a vast amount of text and can now be applied to diverse tasks, including text generation, code writing, and language translation. Therefore, analogous to large language models such as ChatGPT, foundation models based on HRMS data can learn latent representations correlated with chemical properties directly from spectral data, without requiring explicit annotations of molecular structures or physicochemical properties as input during pretraining.
In this perspective, we explore the potential of foundation models for LC–HRMS data analysis. As untargeted HRMS experiments often rely on the analysis of both MS1 and MS2 data, we will discuss key challenges associated with both of these data types together with the most relevant solutions proposed by the scientific community. We focus on small-molecule HRMS data, where predicting molecular properties directly from raw mass spectra remains more challenging compared to proteins or peptides. In the discussion, we include some solutions proposed in proteomics, as they hold value as potentially transferable approaches to the small-molecule domain. Lastly, we provide an outlook on future developments and key areas that will enable further advances in foundation models for LC–HRMS.
Challenges Associated with LC–HRMS Data
Challenges Associated with MS1 Modeling
Machine learning modeling of MS1 data has predominantly followed two strategies: modeling of MS1 data represented as feature tables using traditional machine learning algorithms or modeling MS1 more directly through raw or learned representations of the underlying spectral data. Both approaches face inherent limitations. MS1 data are unstructured, sparse, and noisy, with signal intensities that only scale linearly with molecular abundance within a limited dynamic range and are further confounded by ion suppression, matrix effects, and detector saturation in addition to substantial technical and biological variability. Moreover, the data are high-dimensional with complex dependencies across retention time and m/z dimensions. These characteristics violate key assumptions of traditional machine learning, such as fixed-length, well-conditioned feature spaces, stable signal-to-noise ratios, and sufficient sample-to-feature ratios, making the model highly susceptible to overfitting. As a consequence, extensive preprocessing is often required to impose structure, which can inadvertently discard informative signals or introduce additional biases (see below), while modeling raw or alternatively processed data relies on assumptions about signal regularity and noise that often do not hold. These challenges underscore the critical need for frameworks that can robustly capture MS1 complexity while minimizing the level of preprocessing or restrictive assumptions.
MS1 data preprocessing remains a central challenge. Its primary objective is to structure raw spectra and enhance data quality through algorithms that detect mass features and align them across samples, while minimizing noise. “Featurization” structures mass spectra into feature tables by arranging detected mass features as columns and samples as rows. , Currently, this process is based on well-defined rule-based algorithms, including binning, construction of extracted ion chromatograms, peak detection, and alignment, to group elution profiles across samples based on similarities in m/z and retention time. Featurization is highly sensitive to user-defined parameters and may underrepresent the underlying data structure. Consequently, resulting feature tables can lack reproducibility, misclassify noise as chemical features (false positives), and fail to capture true features (false negatives). Importantly, featurization does not eliminate the intrinsic challenges of MS1 data. Feature tables inherit key limitations, including the high dimensionality of chemical features relative to sample size, the complex and context-dependent relationship between ion intensity and molecular abundance, and the influence of technical and biological variability. Consequently, feature tables still need preprocessing to become suitable for machine learning modeling, a step that could be avoided by directly analyzing the mass spectra.
Deep learning approaches aim to overcome some of these limitations by modeling raw or minimally processed MS1 data directly. − Convolutional neural networks (CNNs) are among the most commonly used architectures in this context. For example, Kantz and colleagues proposed a CNN-based strategy applied after featurization. Rather than modeling raw MS1 spectra or conventional feature tables, they implemented an image-based representation of predetected peak groups (features represented as m/z-retention time windows) and trained a CNN to classify each feature as a true peak or noise. Although innovative, this approach still depends on precomputed feature tables. Melnikov and colleagues extended this concept by reducing the reliance on predefined feature tables. They developed a CNN that operates directly on raw MS1 data to classify signals as true peaks, noise, or uncertain. Their method first applies dynamic binning to identify regions of interest (ROIs), defined as clusters of m/z centroids with limited mass deviation and approximately Gaussian intensity profiles, which are then classified by CNN. More recently, Deng and colleagues proposed a multimodule framework that processes raw MS1 data through sequential components: an initial module prepares candidate regions, followed by CNN-based modules for feature extraction and final classification.
Other studies developed in the proteomics community bypassed explicit featurization altogether. Iravani and colleagues trained a CNN to classify samples directly from MS1 spectra and applied layer-wise relevance propagation to interpret model decisions, enabling the identification of spectral regions contributing most strongly to classification. Similarly, Xu and colleagues implemented a transformer architecture combined with a feed-forward neural network to perform sample classification from raw MS1 data. Despite these advances, the generalizability of deep learning models for MS1 analysis remains unclear. Most approaches are evaluated on limited datasets, and the absence of comprehensive ground truth makes it difficult to determine whether such models robustly capture chemical signals across experimental settings. Moreover, these methods are typically developed in a supervised setting; incorporating self-supervised learning strategies could provide a promising alternative to address similar challenges.
Testing the generalizability and performance of deep learning models for untargeted MS1 analysis is inherently difficult. In untargeted LC–HRMS studies, a comprehensive ground truth is unavailable because the full set of true chemical signals present in a sample is unknown. As a result, it is not possible to rigorously assess whether a model accurately captures all relevant features or generalizes beyond the datasets on which it was trained. To address this limitation, some studies have relied on synthetic (artificially generated) datasets for model development and evaluation. , While synthetic data provide controlled conditions in which the true signal composition is known, they fail to fully replicate the complexity of the experimental MS1 data. In particular, important sources of technical and biological variability, such as interday and intraday instrument fluctuations, matrix effects, and random electronic noise, are not adequately represented. These limitations highlight the need for alternative validation strategies and more realistic benchmarking frameworks, which become even more critical when extending modeling efforts beyond MS1 to more structurally informative data types, such as MS2 spectra.
Challenges Associated with MS2 Modeling
MS2 spectra consist of multiple signals arising from the fragmentation of a precursor ion under the defined instrumental conditions. The composition and intensity of these fragment ions vary widely, depending on the molecular structure, fragmentation method, collision energy, and ionization mode. Moreover, MS2 spectra are intrinsically sparse, containing relatively few peaks distributed across a broad and variable m/z range. Therefore, when modeling MS2 data, two major challenges emerge: first, representing MS2 spectra in a form suitable for machine learning is nontrivial, as their variable length, sparsity, and intensity variability complicate the extraction of chemically meaningful features. Second, the limited availability of MS2 spectra from known chemical structures restricts the amount of labeled data for supervised learning, constraining model training and generalization.
Learning chemically meaningful representations from mass spectra remains a core challenge in developing robust machine learning algorithms. Mass spectra represent two-dimensional data with variable length, which makes it inherently difficult to model. To make spectra compatible with machine learning architectures, preprocessing steps such as binning or subsampling peaks are commonly used to create vectors of uniform size and to remove noise. Binning (as mentioned for MS1), for instance, partitions spectra into fixed-width intervals to generate uniform representations and has been used in several machine learning-based models such as Spec2Mol, MS2DeepScore, and ChemEmbed. While these strategies simplify spectral processing and reduce complexity, they often lead to sparse, noise-sensitive, and even incorrect representations, limiting their ability to encode chemically relevant information. Alternative representations of MS2 spectra using molecular fingerprints have also been explored. These molecular fingerprints are fixed-length numerical vectors representing the likelihood of certain chemical substructures or functional groups being present in the spectrum based on predictions. Machine learning-based models, such as CSI:FingerID and MSNovelist, integrate this approach to rank candidate molecules from databases or for de novo structure generation, respectively. Similarly, MS-BART also implements molecular fingerprints as a core spectral representation, enabling scalable pretraining across millions of molecules–fingerprint pairs while preserving chemical meaning.
Beyond binning and fingerprint-based preprocessing, alternative strategies have been proposed to better represent the complex structures of MS2 spectra. A recent study by de Jonge and coworkers compared traditional binned inputs with set-based and graph-based representations on a regression task of predicting the quantitative estimate of drug-likeness (QED) of a molecule from its mass spectra. Inspired by the work of Boulougouri et al., set-based representations of mass spectra involved viewing MS2 spectra as sets of intensity-m/z pairs of unequal numbers across different spectra. On the other hand, in graph-based representations, each peak in a mass spectrum is represented by a vertex (or node) connected to neighboring peaks through edges. In this representation, the peak intensities were used as vertex attributes, while the m/z differences were coded as edge attributes. The results from this study showed that machine learning models trained on set- and graph-based representations of MS2 spectra performed substantially better on the QED task, indicating that representing spectra as sets of intensity-m/z values or graphs can more effectively capture molecular properties related to chemical structures encoded in the mass spectra than binned representations of the same spectra. However, it remains unclear whether these alternative representations of mass spectra result in significant improvements in machine learning algorithms for more complex tasks, such as molecular structure prediction, as this has not yet been systematically tested.
An additional challenge for modeling MS2 data is the limited availability of MS2 spectra from known chemical structures to be used in supervised learning, which constrains both model training and generalization. For example, chemical annotation (assigning chemical structures to unknown mass spectra) has been the focus of many recent ML methods for MS2 data, yet their performance remains insufficiently reliable. This limited performance is partly caused by the scarcity of labeled spectra, as the number of organic compounds with available MS2 spectra in spectral libraries is several orders of magnitude smaller than the total number of known compounds. For instance, Stravs et al. claimed that even by aggregating major spectral libraries, such as NIST, MoNA, MassBank, and GNPS, experimental MS2 spectra exist for only ∼60,000 unique molecules, a number of molecules estimated to be an order of magnitude lower than typically required to train robust chemical annotation models. To address this limitation, the scientific community has explored unsupervised machine learning algorithms. Two notable examples are Spec2Vec and MS2LDA. , Spec2Vec adapts techniques from natural language processing, including word embedding algorithms and co-occurrence-based vectorization, to represent fragment ions in mathematical space that the model can use to predict structural similarity from MS2 spectra. This method was proven to correlate better with structural similarity than traditional spectral similarity metrics, such as cosine-based scores. However, as Spec2Vec was trained on only ∼95,000 MS2 spectra from ∼13,000 unique molecules, its ability to generalize to chemically diverse or underrepresented compound classes may be limited. This contrasts with MS2LDA, which, by identifying co-occurring fragment peaks and neutral losses (termed Mass2Motifs), can capture structural features shared across chemically distinct molecules, potentially revealing relationships that are not evident from spectral similarity alone. , This method may significantly increase the annotation coverage of unknown spectra by providing substructural information for molecules absent from spectral libraries. Nonetheless, the interpretability and utility of the discovered motifs remain dependent on comparisons with manually curated data and expert validation.
Foundation Models Based on Mass Spectrometry Data
Overcoming the limitations in modeling MS1 and MS2 data requires approaches capable of learning from large-scale data, including unlabeled spectra. Techniques such as self-supervised learning and contrastive learning can extract meaningful representations from the unlabeled HRMS data. These techniques provide the basis for large-scale strategies such as foundation models, making them a particularly promising direction for improving MS1 and MS2 data analysis. Foundation models are built on transfer learning principles, where knowledge learned during pretraining is reused and refined for downstream tasks. Compared to conventional transfer learning approaches, foundation models scale both the size of training data and the diversity of downstream applications through task-agnostic pretraining on large unlabeled datasets. Currently, many foundation models rely on transformer neural network architectures that process segments of mass spectra and learn relationships between spectral features using self-supervised learning strategies. , A commonly used self-supervised learning strategy is masking, in which parts of the input data are intentionally removed, and the model is trained to reconstruct them from the remaining data. This forces the model to learn meaningful relationships within the data, such as correlations between fragment ions (m/z) in MS2, trends in chromatographic retention times, or patterns in MS1 m/z ions (and intensity) profiles. These learned relationships are encoded as mathematical representations in the model that can later be fine-tuned for specific applications. This paradigm is particularly valuable for LC–HRMS, where most spectra generated in untargeted experiments lack structural annotations. Because foundation models do not require labeled spectra during pretraining, they can exploit the full scale of LC–HRMS data available in raw data repositories, such as GNPS/MassIVE, MetaboLights, Metabolomics Workbench, and NORMAN, which contain billions of unlabeled mass spectra. Once trained, these models can serve as a general foundation for multiple downstream tasks, including chemical annotation and molecular property prediction.
Current Foundation Models for Small Molecules
Several foundational models have been developed for the analysis of small molecules using HRMS data. These include DreaMS, MSBERT, LSM1-MS2, and PRISM, all of which were introduced within the past five years. These models differ substantially in training data scale, learning strategies, input representations, and the availability of code and datasets. Together, they illustrate both a paradigm shift toward self-supervised learning in mass spectrometry data analysis and the rapid evolution of the field.
DreaMS and MSBERT are widely recognized open-source foundation models for small-molecule MS2 data that have demonstrated state-of-the-art performance across multiple tasks, such as spectral similarity scoring, molecular fingerprint prediction, and molecular property prediction. , DreaMS and MSBERT share similar transformer-based architectures but differ in the size of the training dataset and training strategy. DreaMS was trained on datasets of varying size and curation quality, with the GeMS-A10 dataset (comprising over 24 million well-curated spectra from GNPS/MassIVE) yielding the best performance. Its training relied primarily on masked spectrum reconstruction and successive fine-tuning via contrastive learning. Similarly, MSBERT combined masked reconstruction with contrastive learning but used differently masked versions of the same spectrum during training. Although this strategy improved performance relative to traditional machine learning-based models, such as Spec2Vec, the relatively small training dataset (∼164,000 Orbitrap spectra) may limit scalability and generalization across chemically diverse compound classes. Other foundation models for small molecules include LSM1-MS2 , and PRISM. These commercially developed models were trained on substantially larger datasets and have been reported to achieve state-of-the-art performance in tasks such as spectral similarity prediction, molecular property prediction, and de novo structure generation. However, limited details regarding data curation and restricted access to the final training datasets and source code constrain reproducibility and broader adoption.
Although current foundation models differ substantially in the size and curation of their training datasets, most are trained primarily on DDA data and target similar applications, such as compound annotation and molecular property prediction (Table ). However, the foundation model framework is not restricted to DDA data and can be extended to other types of data, including MS1 and DIA. Furthermore, it is also possible to adapt foundation models from other fields, such as audio and speech processing, to learn from mass spectrometry data. An example is the “Foundation Model to Assess Cancer Tissue” (FACT). Unlike other small molecule foundation models, FACT was developed by adapting an existing audio-language foundation model to use Rapid Evaporative Ionization MS data (REIMS) as input. REIMS is an ambient ionization mass spectrometry technique that generates MS1 spectra from metabolites present in surgical aerosols produced by the thermal evaporation of biological tissue during electrosurgical procedures. This model exploits similarities between REIMS spectral patterns, such as short-time frequency variations and sudden fluctuations in intensities, with audio Mel-spectrograms, commonly used in speech processing and machine learning tasks. Accordingly, using REIMS data, this model was fine-tuned to distinguish cancerous from healthy tissue. The development of FACT illustrates that foundation models can be adapted beyond other spectral types to alternative MS modalities.
1. Recently Developed Foundation Models for Small Molecules Based on HRMS Data .
| Tool name (year) | Origin of the training dataset | Data type | Size training dataset | Fine-tunning tasks | Open source model/data | Ref. |
|---|---|---|---|---|---|---|
| DreaMS (2025) | GNPS/MassIVE | DDA Retention orders | 24 million spectra (GeMS-A10) | 1) spectral similarity 2) prediction of molecular fingerprints 3) Fluorine detection | Yes/Yes | |
| MSBERT (2024) | GNPS/MassIVE | DDA | 163,953 spectra | 1) spectral similarity 2) analogous compounds searching | Yes/Yes | |
| LSM1-MS2 (2024) | Metabolomics Workbench and in-house data | DDA | 100 million spectra | 1) spectral similarity 2) prediction of chemical properties 3) de novo structure generation | No/No | |
| PRISM (2024) | GNPS/MassIVE, MetaboLights, Metabolomics Workbench, and in-house data | DDA | 1.2 billion spectra | 1) spectral similarity 2) prediction of chemical properties | No/No |
DDA: data-dependent acquisition; DIA: data-independent acquisition; REIMS: rapid evaporative ionization mass spectrometry; DreaMS: deep representations empowering the annotation of mass spectra; MSBERT: bidirectional encoder representations from transformers for mass spectrometry; PRISM: pretrained representations informed by spectral masking.
Current foundation models for small molecules present several advantages over those of traditional machine learning approaches. For example, they show superior performance than commonly implemented methods in tasks such as spectral similarity. ,− In addition, once fine-tuned, these models can handle more complex tasks, such as detecting fluorinated compounds from mass spectra and de novo molecular generation. Another key benefit is data efficiency: fine-tuning can require up to 50% less data than training a machine learning model from scratch. However, the development, adaptation, and implementation of foundation models for small molecules remain limited due to lack of expertise, data, and computational requirements. Fine-tuning of foundation models demands not only expertise in machine learning, chemistry, and programming but also access to advanced computational infrastructure, such as high-performance computing clusters, which may not be available to all laboratories. Furthermore, their practical applicability for routine use presents several challenges. For example, even the best-in-class foundation models reliably annotate only a low percentage of unknown spectra in a sample. Next to this, current models are developed only for positive-mode spectra, which is a significant practical limitation for routine metabolomics and environmental analysis. Likewise, applying these models to spectra generated by low-resolution instruments (e.g., quadrupole or ion trap) is a significant barrier to routine adoption.
From a machine learning perspective, there are also important caveats regarding their training, benchmarking, and applicability. Foundation models are difficult to benchmark and can be biased by large-scale pretraining. Generally, these models are evaluated on the basis of their performance on fine-tuning tasks. However, this approach complicates their benchmarking, as performance can vary depending on the types of fine-tuning tasks and testing data, neither of which is standardized across studies. As shown in the development of LMS1-MS2, the performance of the foundation model in analogous compound searching ranged from outperforming traditional machine learning models to being outperformed by cosine similarity, depending on the testing dataset. In addition, not all of the foundation models developed for small molecules were tested on the same fine-tuning tasks. While MSBERT outperforms DreaMS in compound annotation tasks, it remains unclear whether this advantage generalizes to all chemical classes and other untested tasks, such as fluorine detection. Furthermore, the quality of the data used for the pretraining is also critical for the performance of foundation models, as this directly reflects how well the model can learn a correct representation of the data.
Additionally, potential biases in the training data can propagate to downstream fine-tuning tasks. This phenomenon has been well documented in the language domain, where foundation models, particularly large language models, exhibit the phenomenon of “toxic degeneration”. In such cases, biases present in the training data are amplified during fine-tuning, resulting in the generation of non-normative or harmful text. In the LC–HRMS field, such biases may arise from imbalances in factors related to both analytical conditions and sample composition in the pretraining data. These include, for example, sample matrices (e.g., food or water), instrument types (e.g., Orbitrap or TOF), and acquisition settings. As a result, model performance is likely to vary depending on how closely these factors in the test data align with those represented during pretraining.
In the small-molecule domain, no foundation models have yet been developed specifically for DIA- or ESI-based MS1 data. As a result, tasks associated with these acquisition modes, such as peak detection and integration, DIA-based chemical annotation, and retention time prediction, are not currently supported within a foundation-model framework. This can be seen as a form of large-pretraining bias toward a specific type of mass spectra. The direct implementation of existing foundation models designed for DDA data for other types of MS data, while theoretically possible, would likely show suboptimal performance without modifications to the model architecture (e.g., masking approach). For example, when testing the DreaMS model for REIMS data, as part of the FACT model adaptation, a near baseline model performance was achieved, comparable to that of traditional machine learning techniques such as Principal Component Analysis/Linear Discriminant Analysis.
Outlook on the Future of LC–HRMS Foundation Models
Foundation models for LC–HRMS have demonstrated considerable potential; however, their development remains in its early stage. We envision that future progress in this field will be driven by four key areas currently underdeveloped in existing foundation models for small molecules: (i) harmonization and centralization of curated datasets, (ii) implementation of federated learning strategies, (iii) integration of multiple HRMS data types, and (iv) improved model explainability. Below, we outline these areas and describe how they may accelerate progress in this field.
Harmonization and Centralization of Curated Datasets
The performance of foundation models is critically dependent on both the quantity and the quality of the training data. As demonstrated by Bushuiev et al., models trained on carefully curated, high-quality spectra can outperform those trained on substantially larger but noisier datasets. Key quality criteria include instrument m/z accuracy, maximum m/z range, and polarity, among others, as well as reliable metadata annotations.
Currently, researchers rely on large public repositories such as GNPS/MassIVE and MetaboLights to assemble sufficiently large datasets for pretraining. While these repositories provide invaluable access to raw LC–HRMS data, data and metadata curation are not standardized across platforms. Variability in file formats, spectra quality, instrument types, and acquisition settings, including collision energies and polarity, necessitates extensive preprocessing before data can be used for large-scale model pretraining. This process is often labor-intensive and requires ad hoc filtering and assumptions that could be avoided by standardizing curation of the repositories. We envision that the implementation of standardized, automated curation pipelines for untargeted LC–HRMS datasets (similar to the periodic data cleaning and metadata harmonization done in GNPS for spectral libraries using the MatchMS package) will represent a major advance. Establishing community-agreed quality control metrics, harmonized metadata schemas (e.g., consistent reporting of instrument type, acquisition parameters, etc.), and automated preprocessing workflows would facilitate large-scale model pretraining while improving transparency and reproducibility.
Beyond benefiting foundation models, such harmonization would strengthen benchmarking efforts, enable fair model comparisons, and accelerate the development and validation of next-generation machine learning approaches for LC–HRMS. As shown in other fields, such as genomics and pathology, the performance and robustness of foundation models can be addressed by standardizing the testing dataset and fine-tuning tasks. The performance on fine-tuning tasks can be evaluated via well-known metrics like area under the curve. The robustness of the foundation models can be evaluated by studying the organization of the embedding space via clustering techniques. This was shown in the development of DreaMS, where the authors showed that the embeddings cluster as a function of chemical properties. We anticipate that the definition and standardization of benchmarking tasks and data would lead to the development of novel benchmarking metrics, as previously shown in the field of pathology, where a specific metric was developed to measure the degree of representation between the embedding space and meaningful experimental features.
Enabling Sensitive Data Use with Federated Learning
Foundation model development is intertwined with data availability. These models benefit not only from open-access datasets but also from proprietary in-house data, as demonstrated by current foundation models for small molecules (Table ). However, because of privacy regulations or the sensitive nature of certain information, data and metadata cannot always be shared with the broader scientific community. This lack of accessibility can limit model reproducibility and reduce the amount of data available for pretraining.
Federated learning offers a solution to this challenge as it consists of a decentralized and privacy-preserving training strategy. In federated learning, a shared model is trained on local data across multiple devices without the transfer of sensitive raw data to a central server. Instead, only model hyperparameters are shared with a central server, which aggregates these updates, often using algorithms like federated averaging, to create a better global model. − This approach enables collaborative AI training while keeping the data secure and localized.
As demonstrated by Chen and colleagues, the federated learning strategy can be applied to untargeted MS data. We envision that foundation models will greatly benefit from federated strategies because they can lead to a substantial increase in the effective size of the training dataset. This increase can improve the performance of foundation models that can then be fine-tuned for specific tasks. Although literature on federated foundation models remains limited, recent research has explored their development for text data. While these federated models did not yet surpass their centralized counterparts, the concept holds significant promise for LC–HRMS applications, as it could enable learning from private in-house data, for example, from laboratories performing routine analyses for regulatory purposes that would otherwise be inaccessible.
Modeling of Other Types of HRMS Data
Current foundation models do not fully exploit the diversity of information available in LC–HRMS experiments. Most models developed for small-molecule analysis rely primarily on DDA data and incorporate little or no information from other acquisition modes.
For example, DreaMS is specifically designed for DDA spectra and only includes precursor m/z values and retention time information (encoded as retention orders) derived from MS1 scans. As a result, it is not readily applicable to untargeted MS1-based workflows or DIA data. Conversely, models such as FACT are adapted for REIMS data and are not intended for MS2-driven tasks. Therefore, an important next step is the development of dedicated foundation models for other major HRMS data types such as ESI MS1 full-scan data and DIA data. These acquisition modes capture distinct, but complementary, information. MS1 data encode chromatographic peak shapes, isotope patterns, and adduct distributions, which are central to feature detection and quantification. DIA data provide untargeted fragmentation of precursor ions across the chromatographic peak, which not only enables compound annotation at low concentrations but also captures chemically relevant intensity patterns analogous to those of MS1 features. We envision that building foundation models directly on these data types will lead to many advantages, from learning data-type-specific representations directly from raw spectra to enabling downstream tasks in untargeted workflows (e.g., peak detection, deconvolution, compound classification, and improved annotation). For example, current feature detection algorithms rely entirely on hand-crafted rules and fixed, user-defined thresholds to make decisions that are inherently dataset- and instrument-dependent. The parameters governing these decisions, like m/z tolerance, peak width, signal-to-noise cutoffs, and shape criteria, must be manually tuned by the user. These parameters do not generalize well across instruments, acquisition settings, or sample types, as extensively documented in comparative studies. − A reframing of this problem in a machine learning setting would involve training a model to learn these decision boundaries directly from data: inferring which combinations of m/z deviation patterns, intensity profiles, chromatographic shapes, and noise characteristics constitute a true chemical feature rather than encoding them as fixed rules. A foundation model operating directly on raw spectral data could, in principle, learn to perform the full detection pipeline end-to-end. Evidence that this is achievable comes from the proteomics field. CASCADIA is a supervised transformer-based model that takes as input temporally adjacent raw MS1 and MS2 scans, represented as four-tuple peaks (m/z, intensity, time offset, and acquisition mode flag). It processes the input jointly through a 2D positional encoder, linear projection, and transformer layers to compute embeddings, without any rule-based preprocessing such as precursor detection or feature alignment. Thus, by treating the raw spectral data as a sequence of peaks spanning both acquisition modes, the model internally learns which signals across MS1 and MS2 are chemically meaningful, effectively performing a learned equivalent of what rule-based algorithms approximate through fixed thresholds and hand-crafted criteria. Although designed for a specific supervised task, we argue that the CASCADIA architecture is highly transferable to self-supervised learning settings and could therefore serve as a basis for future foundation models. We envision future foundation models for LC–HRMS could replace fixed algorithmic thresholds with representations learned from the data and potentially incorporate self-supervised strategies such as Joint Embedding Predictive Architectures, in which masking operates in embedding space rather than on raw data, enabling richer spectral representations without requiring labeled datasets.
Furthermore, no foundational models have been developed that systematically integrate multiple types of HRMS data within a unified framework. Incorporating multiple HRMS data types during pretraining would substantially broaden the scope of downstream applications. Future multimodal foundation models could be fine-tuned not only for MS2-based annotation tasks but also for MS1- and DIA-related tasks. Additionally, expanding the diversity of input data during pretraining could enable the model to learn richer representations of chromatographic and spectral patterns, supporting additional fine-tuning tasks and likely improving performance.
Improving Model Explainability
Foundation models, and deep learning models in general, would benefit substantially from improved explainability. These models are often regarded as “black boxes” because it is difficult to determine which input features contribute to a specific prediction and how they influence the final output. This lack of transparency can limit user trust, hinder validation, and slow adoption in high-stakes scientific applications.
To address this challenge, the research community has developed a range of explainable artificial intelligence (XAI) techniques. These approaches can be model-agnostic, such as SHAP (Shapley Additive Explanations), which estimates the contribution of individual input features to a prediction, or they can be inherently integrated into the model architecture. In the context of HRMS and chemical analysis, explainability is critical. Identifying which regions of a mass spectrum, such as specific m/z values, fragmentation patterns, or isotopic profiles, drive the prediction of a molecular property or chemical structure would significantly enhance confidence in the model’s outputs. Moreover, interpretable models could facilitate error analysis, reveal systematic biases, and potentially uncover novel structure–spectra relationships. Therefore, improving explainability is not merely a matter of transparency but a key requirement for the responsible and scientifically rigorous deployment of foundation models in mass spectrometry.
We envision future research on foundation models for LC–HRMS data not only focusing on performance but also on explainability. Examples of explainable foundation model architectures include prototype- or attention-based frameworks. These frameworks provide interpretable intermediate representations linked to specific training examples or feature regions, and they have already been explored in other domains. For instance, Huang and colleagues developed a foundation model for therapeutic drug selection that incorporates an explainable network to justify its predictions. Understanding why a model classifies a drug as therapeutic is particularly important, as such insights can guide further optimization and experimental validation.
Conclusion
Foundation models represent a promising new paradigm for LC–HRMS data analysis. Through self-supervised training, these models can learn meaningful representations from large volumes of unlabeled spectral data that have been underexploited. This capability is particularly valuable in mass spectrometry, where labeled datasets remain limited, yet raw, unlabeled spectra generation continues to expand rapidly. Once pretrained, they can be adapted to diverse fine-tuning tasks, reflecting the wide range of LC–HRMS applications, from sample classification to chemical annotation and prediction of molecular properties.
The future development and successful deployment of LC–HRMS foundation models will be dependent on several key factors. Improvements in data quality, standardization, and accessibility are essential to ensuring robust and transferable representations. A more comprehensive exploitation of available information, including metadata and different data types, is equally important to fully capturing the complexity of LC–HRMS experiments. Furthermore, the development of explainable and transparent model architectures is critical to foster trust, enable validation, and support regulatory acceptance. Collaborative efforts among researchers, instrument vendors, data repositories, and regulatory bodies will be crucial for establishing benchmarks, promoting interoperability, and validating models across applications. Such initiatives could advance LC–HRMS beyond task-specific solutions toward more generalizable and robust AI-driven analytical frameworks.
Acknowledgments
Financial support is acknowledged from the Dutch Ministry of Agriculture, Fisheries, Food Security and Nature through the Food Safety Knowledge Development Program (project KB-54-000-007) granted to Wageningen Food Safety Research (part of Wageningen University and Research).
#.
A.J.C. and F.P.-G. contributed equally to this manuscript and indicate joint first authorship. All authors contributed to the conceptualization of this perspective paper and its overall structure. A.J.C., F.P.-G., L.M.B., and D.K. carried out most of the literature review. A.J.C. and F.P.-G. led the writing of the original draft and carried out the final revisions of the manuscript, with contributions from all other authors. L.M.B., B.V., and D.K. contributed substantially to drafting and revising the manuscript. A.J.C. and F.P.-G. managed the project administration. M.A., B.V., and M.H.B. contributed to the review and editing of the first draft and final versions. B.V. led the funding acquisition for the project, overall project supervision, and provided critical review and editing of the first draft and final versions. All authors reviewed the manuscript and have given approval to the final version of the manuscript.
The authors declare no competing financial interest.
References
- Guo J., Shen S., Xing S., Chen Y., Chen F., Porter E. M., Yu H., Huan T.. EVA: Evaluation of Metabolic Feature Fidelity Using a Deep Learning Model Trained with over 25000 Extracted Ion Chromatograms. Anal. Chem. 2021;93(36):12181–12186. doi: 10.1021/acs.analchem.1c01309. [DOI] [PubMed] [Google Scholar]
- Gloaguen Y., Kirwan J. A., Beule D.. Deep Learning-Assisted Peak Curation for Large-Scale LC-MS Metabolomics. Anal. Chem. 2022;94(12):4930–4937. doi: 10.1021/acs.analchem.1c02220. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huber F., Ridder L., Verhoeven S., Spaaks J. H., Diblen F., Rogers S., van der Hooft J. J. J.. Spec2Vec: Improved Mass Spectral Similarity Scoring through Learning of Structural Relationships. PLoS Comput. Biol. 2021;17(2):e1008724. doi: 10.1371/journal.pcbi.1008724. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huber F., van der Burg S., van der Hooft J. J. J., Ridder L.. MS2DeepScore: A Novel Deep Learning Similarity Measure to Compare Tandem Mass Spectra. J. Cheminform. 2021;13(1):84. doi: 10.1186/s13321-021-00558-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Voronov, G. ; Lightheart, R. ; Frandsen, A. ; Bargh, B. ; Haynes, S. E. ; Spencer, E. ; Schoenhardt, K. E. ; Davidson, C. ; Schaum, A. ; Macherla, V. R. ; et al. MS2Prop: A Machine Learning Model That Directly Generates de Novo Predictions of Drug-Likeness of Natural Products from Unannotated MS/MS Spectra. bioRxiv 2022. [Google Scholar]
- de Jonge N. F., Louwen J. J. R., Chekmeneva E., Camuzeaux S., Vermeir F. J., Jansen R. S., Huber F., van der Hooft. MS2Query: reliable and scalable MS2 mass spectra-based analogue search. Nat. Commun. 2023;14(1):1752. doi: 10.1038/s41467-023-37446-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stravs M. A., Dührkop K., Böcker S., Zamboni N.. MSNovelist: De Novo Structure Generation from Mass Spectra. Nat. Methods. 2022;19(7):865–870. doi: 10.1038/s41592-022-01486-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Litsa E. E., Chenthamarakshan V., Das P., Kavraki L. E.. An End-to-End Deep Learning Framework for Translating Mass Spectra to de-Novo Molecules. Commun. Chem. 2023;6(1):132. doi: 10.1038/s42004-023-00932-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bittremieux W., Noble W. S.. Self-Supervised Learning from Small-Molecule Mass Spectrometry Data. Nat. Biotechnol. 2026;44:538–539. doi: 10.1038/s41587-025-02677-x. [DOI] [PubMed] [Google Scholar]
- Bushuiev R., Bushuiev A., Samusevich R., Brungs C., Sivic J., Pluskal T.. Self-Supervised Learning of Molecular Representations from Millions of Tandem Mass Spectra Using DreaMS. Nat. Biotechnol. 2026;44:630–640. doi: 10.1038/s41587-025-02663-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paaß, G. ; Giesselbach, S. . Foundation Models for Natural Language Processing, Pre-trained Language Models Integrating Media; Springer, 2023. DOI: 10.1007/978-3-031-23190-2. [DOI] [Google Scholar]
- Singh A.. Self-Supervised Learning of Molecular Representations. Nat. Methods. 2025;22(7):1395–1395. doi: 10.1038/s41592-025-02757-5. [DOI] [PubMed] [Google Scholar]
- Renner G., Reuschenbach M.. Critical Review on Data Processing Algorithms in Non-Target Screening: Challenges and Opportunities to Improve Result Comparability. Anal. Bioanal. Chem. 2023;415:4111–4123. doi: 10.1007/s00216-023-04776-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tautenhahn R., Bottcher C., Neumann S.. Highly Sensitive Feature Detection for High Resolution LC/MS. BMC Bioinf. 2008;9:504. doi: 10.1186/1471-2105-9-504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Myers O. D., Sumner S. J., Li S., Barnes S., Du X.. Detailed Investigation and Comparison of the XCMS and MZmine 2 Chromatogram Construction and Chromatographic Peak Detection Methods for Preprocessing Mass Spectrometry Metabolomics Data. Anal. Chem. 2017;89(17):8689–8695. doi: 10.1021/acs.analchem.7b01069. [DOI] [PubMed] [Google Scholar]
- Melnikov A. D., Tsentalovich Y. P., Yanshole V. V.. Deep Learning for the Precise Peak Detection in High-Resolution LC-MS Data. Anal. Chem. 2020;92(1):588–592. doi: 10.1021/acs.analchem.9b04811. [DOI] [PubMed] [Google Scholar]
- Kantz E. D., Tiwari S., Watrous J. D., Cheng S., Jain M.. Deep Neural Networks for Classification of LC-MS Spectral Peaks. Anal. Chem. 2019;91(19):12407–12413. doi: 10.1021/acs.analchem.9b02983. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Risum A. B., Bro R.. Using Deep Learning to Evaluate Peaks in Chromatographic Data. Talanta. 2019;204:255–260. doi: 10.1016/j.talanta.2019.05.053. [DOI] [PubMed] [Google Scholar]
- Iravani S., Conrad T. O. F.. An Interpretable Deep Learning Approach for Biomarker Detection in LC-MS Proteomics Data. IEEE/ACM Trans. Comput. Biol. Bioinform. 2023;20(1):151–161. doi: 10.1109/TCBB.2022.3141656. [DOI] [PubMed] [Google Scholar]
- Deng Y., Yao Y., Wang Y., Yu T., Cai W., Zhou D., Yin F., Liu W., Liu Y., Xie C., Guan J., Hu Y., Huang P., Li W.. An End-to-End Deep Learning Method for Mass Spectrometry Data Analysis to Reveal Disease-Specific Metabolic Profiles. Nat. Commun. 2024;15(1):7136. doi: 10.1038/s41467-024-51433-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu W., Zhang L., Qian X., Sun N., Tu X., Zhou D., Zheng X., Chen J., Xie Z., He T., Qu S., Wang Y., Yang K., Su K., Feng S., Ju B.. A Deep Learning Framework for Hepatocellular Carcinoma Diagnosis Using MS1 Data. Sci. Rep. 2024;14(1):26705. doi: 10.1038/s41598-024-77494-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Han, Y. ; Wang, P. ; Yu, K. ; Chen, X. ; Chen, L. . MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation. arXiv.2025. [Google Scholar]
- de Jonge, N. ; van der Hooft, J. J. J. ; Probst, D. . To Bin or Not to Bin: Alternative Representations of Mass Spectra. arXiv.2025. [Google Scholar]
- Faizan-Khan M., Giné R., Badia J. M., Pérez-Ribera M., Junza A., Vinaixa M., Sales-Pardo M., Guimerà R., Yanes O.. ChemEmbed: A Deep Learning Framework for Metabolite Identification Using Enhanced MS/MS Data and Multidimensional Molecular Embeddings. Briefings Bioinf. 2026;27:bbag054. doi: 10.1093/bib/bbag054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dührkop K., Shen H., Meusel M., Rousu J., Böcker S.. Searching Molecular Structure Databases with Tandem Mass Spectra Using CSI: FingerID. Proc. Natl. Acad. Sci. U. S. A. 2015;112(41):12580–12585. doi: 10.1073/pnas.1509788112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boulougouri M., Vandergheynst P., Probst D.. Molecular Set Representation Learning. Nat. Mach. Intell. 2024;6(7):754–763. doi: 10.1038/s42256-024-00856-0. [DOI] [Google Scholar]
- van der Hooft J. J. J., Wandy J., Barrett M. P., Burgess K. E. V., Rogers S.. Topic Modeling for Untargeted Substructure Exploration in Metabolomics. Proc. Natl. Acad. Sci. U. S. A. 2016;113(48):13738–13743. doi: 10.1073/pnas.1608041113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Torres-Ortega, L. R. ; Dietrich, J. ; Wandy, J. ; Mol, H. ; van der Hooft, J. J. J. . Large-Scale Discovery and Annotation of Hidden Substructure Patterns in Mass Spectrometry Profiles. bioRxiv 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vaswani, A. ; Shazeer, N. ; Parmar, N. ; Uszkoreit, J. ; Jones, L. ; Gomez, A. N. ; Kaiser, L. ; Polosukhin, I. . Attention Is All You Need. arXiv. 2023. [Google Scholar]
- Farahmand M., Jamzad A., Fooladgar F., Connolly L., Kaufmann M., Ren K. Y. M., Rudan J., McKay D., Fichtinger G., Mousavi P. F.. Foundation Model for Assessing Cancer Tissue Margins with Mass Spectrometry. Int. J. Comput. Assist. Radiol. Surg. 2025;20(6):1097–1104. doi: 10.1007/s11548-025-03355-8. [DOI] [PubMed] [Google Scholar]
- Zhang H., Yang Q., Xie T., Wang Y., Zhang Z., Lu H.. MSBERT: Embedding Tandem Mass Spectra into Chemically Rational Space by Mask Learning and Contrastive Learning. Anal. Chem. 2024;96(42):16599–16608. doi: 10.1021/acs.analchem.4c02426. [DOI] [PubMed] [Google Scholar]
- Asher, G. ; Shah, D. ; Caudy, A. A. ; Ferro, L. ; Amar, L. ; Costa, A. S. H. ; Patton, T. ; O’Connor, N. ; Campbell, J. M. ; Geremia, J. . LSM-MS2: A Foundation Model Bridging Spectral Identification and Biological Interpretation. arXiv 2025. [Google Scholar]
- Asher, G. ; Cadosh Delmar, M. ; Campbell, J. M. ; Geremia, J. ; Kassis, T. . LSM1-MS2: A Foundation Model for MS/MS, Encompassing Chemical Property Predictions, Search and de Novo Generation. ChemRxiv 2024. [Google Scholar]
- () Healey, D. ; Domingo-Fernández, D. ; Taylor, J. ; Krettler, C. ; Lightheart, R. ; Park, T. ; Kind, T. ; Allen, A. ; Colluru, V. . PRISM: a Foundation Model For life’s Chemistry https://enveda.com/prism-a-foundation-model-for-lifes-chemistry/.
- Schramowski P., Turan C., Andersen N., Rothkopf C. A., Kersting K.. Large Pre-Trained Language Models Contain Human-like Biases of What Is Right and Wrong to Do. Nat. Mach. Intell. 2022;4(3):258–268. doi: 10.1038/s42256-022-00458-8. [DOI] [Google Scholar]
- de Jonge N. F., Hecht H., Strobel M., Wang M., van der Hooft J. J. J., Huber F.. Reproducible MS/MS Library Cleaning Pipeline in Matchms. J. Cheminform. 2024;16:88. doi: 10.1186/s13321-024-00878-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kömen, J. ; de Jong, E. D. ; Hense, J. ; Marienwald, H. ; Dippel, J. ; Naumann, P. ; Marcus, E. ; Ruff, L. ; Alber, M. ; Teuwen, J. ; Klauschen, F. ; Müller, K.-R. . Towards Robust Foundation Models for Digital Pathology. arXiv 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Feng H., Wu L., Zhao B., Huff C., Zhang J., Wu J., Lin L., Wei P., Wu C.. Benchmarking DNA Foundation Models for Genomic and Genetic Tasks. Nat. Commun. 2025;16(1):10780. doi: 10.1038/s41467-025-65823-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McMahan, H. B. ; Moore, E. ; Ramage, D. ; Hampson, S. ; Arcas, B. A. Y. . Communication-Efficient Learning of Deep Networks from Decentralized Data. arXiv. 2023. [Google Scholar]
- Hanser T.. Federated Learning for Molecular Discovery. Curr. Opin. Struct. Biol. 2023;79:102545. doi: 10.1016/j.sbi.2023.102545. [DOI] [PubMed] [Google Scholar]
- Chen J., Yang Q., Dai Q., Chang S., Tian J., Cong S., Ji H.. FederEI: Federated Library Matching Framework for Electron Ionization Mass Spectrum Based Compound Identification. Anal. Chem. 2024;96(40):15840–15845. doi: 10.1021/acs.analchem.4c02313. [DOI] [PubMed] [Google Scholar]
- Fendor Z., van der Velden B. H. M., Wang X., Carnoli A. J., Mutlu O., HürriyetoĞlu A.. Federated Learning in Food Research. J. Agric Food Res. 2025;23:102238. doi: 10.1016/j.jafr.2025.102238. [DOI] [Google Scholar]
- Tian Y., Wan Y., Lyu L., Yao D., Jin H., Sun L.. FedBERT: When Federated Learning Meets Pre-Training. ACM Trans. Intell. Syst. Technol. 2022;13(4):1–26. doi: 10.1145/3510033. [DOI] [Google Scholar]
- Aigensberger M., Bueschl C., Castillo-Lopez E., Ricci S., Rivera-Chacon R., Zebeli Q., Berthiller F., Schwartz-Zimmermann H. E.. Modular Comparison of Untargeted Metabolomics Processing Steps. Anal. Chim. Acta. 2025;1336:343491. doi: 10.1016/j.aca.2024.343491. [DOI] [PubMed] [Google Scholar]
- Guo J., Huan T.. Mechanistic Understanding of the Discrepancies between Common Peak Picking Algorithms in Liquid Chromatography-Mass Spectrometry-Based Metabolomics. Anal. Chem. 2023;95(14):5894–5902. doi: 10.1021/acs.analchem.2c04887. [DOI] [PubMed] [Google Scholar]
- Sadia M., Boudguiyer Y., Helmus R., Seijo M., Praetorius A., Samanipour S.. A Stochastic Approach for Parameter Optimization of Feature Detection Algorithms for Non-Target Screening in Mass Spectrometry. Anal. Bioanal. Chem. 2025;417(27):6033–6047. doi: 10.1007/s00216-024-05425-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sanders J., Wen B., Rudnick P. A., Johnson R. S., Wu C. C., Riffle M., Oh S., MacCoss M. J., Noble W. S.. A Transformer Model for de Novo Sequencing of Data-Independent Acquisition Mass Spectrometry Data. Nat. Methods. 2025;22(7):1447–1453. doi: 10.1038/s41592-025-02718-y. [DOI] [PubMed] [Google Scholar]
- Assran, M. ; Duval, Q. ; Misra, I. ; Bojanowski, P. ; Vincent, P. ; Rabbat, M. ; Le Cun, Y. ; Ballas, N. . Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. In 2023 IEEE/CVF Conference On Computer Vision And Pattern Recognition (CVPR); IEEE, 2023. 10.1109/CVPR52729.2023.01499. [DOI] [Google Scholar]
- van der Velden B. H. M., Kuijf H. J., Gilhuijs K. G. A., Viergever M. A.. Explainable Artificial Intelligence (XAI) in Deep Learning-Based Medical Image Analysis. Med. Image Anal. 2022;79:102470. doi: 10.1016/j.media.2022.102470. [DOI] [PubMed] [Google Scholar]
- Lundberg, S. ; Lee, S.-I. . A Unified Approach to Interpreting Model Predictions. arXiv 2017. [Google Scholar]
- Chen, C. ; Li, O. ; Tao, C. ; Barnett, A. J. ; Su, J. ; Rudin, C. . This Looks Like That: Deep Learning for Interpretable Image Recognition. arXiv 2019. [Google Scholar]
- Huang K., Chandak P., Wang Q., Havaldar S., Vaid A., Leskovec J., Nadkarni G. N., Glicksberg B. S., Gehlenborg N., Zitnik M.. A Foundation Model for Clinician-Centered Drug Repurposing. Nat. Med. 2024;30(12):3601–3613. doi: 10.1038/s41591-024-03233-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bessadok A., Grisoni F.. An Explainable Foundation Model for Drug Repurposing. Nat. Med. 2024;30(12):3422–3423. doi: 10.1038/s41591-024-03333-8. [DOI] [PubMed] [Google Scholar]
