Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 Sep 26;27(5):bbag507. doi: 10.1093/bib/bbag507

MCOD: a memory-constrained deep learning framework for robust outlier detection in quantitative proteomics

Jinze Huang 1, Huanyue Liao 2, Bo Meng 3, Guangkui Fan 4, Dong An 5, Xinhua Dai 6,✉, Xiang Fang 7,✉, Yang Zhao 8,✉
PMCID: PMC13615539  PMID: 42799700

Abstract

Ensuring robust quality control (QC) remains a major challenge in quantitative proteomics, particularly in detecting and managing outliers. Deep learning offers powerful representational capacity for ultra-high-dimensional data but often suffers from overfitting in small-sample scenarios. To address this, we propose memory-constrained outlier detection (MCOD), a deep anomaly detection framework that directly processes MaxQuant outputs and achieves competitive recall at high precision levels, which suggests a reduced risk of overlooking true outliers. MCOD integrates two innovations: (i) a memory-constrained module (MC module) that mitigates over-representation of samples via prototype-based regularization, and (ii) an adaptive steady-aware regulator that dynamically adjusts the per-sample loss weights in the MC module according to the estimated overfitting risk. Across two simulation settings based on a human cervical cancer cell line (HeLa) proteomics dataset and two real-world cancer proteomics datasets, MCOD consistently outperformed 18 statistical, machine learning, and deep learning baselines, achieving superior area under the receiver operating characteristic curve and area under the precision-recall curve scores. Functional enrichment analyses on the two real-world datasets showed that MCOD performed favorably compared to the three domain-specific models. Furthermore, feature-level visualization provided insights into the rationale behind the model’s anomaly assignments. Collectively, MCOD establishes a robust and scalable framework for data QC in quantitative proteomics.

Keywords: outlier detection, memory constraint, proteomics, quality control, deep learning

Introduction

Proteomics, an important branch of life sciences fundamentally enabled by mass spectrometry, plays a crucial role in elucidating fundamental biological mechanisms, disease research, drug discovery, and biotechnological innovation [1, 2]. Consequently, this technique imposes stringent demands on data quality, necessitating rigorous quality control (QC) protocols for every individual sample [3, 4]. In response, researchers have developed extensive QC tools to process the resulting data, making this an essential step in proteomics data analysis [5, 6].

Data-driven QC methods have been widely adopted in proteomics due to their ease of use and flexible software algorithms [7–9]. For mass spectrometry data, researchers have developed comprehensive QC processes by incorporating substantial prior knowledge, such as distance measures and missing rates. For example, Mass speCtrometry Proteomics QC (MaCProQC) [9] processes raw files by analyzing chromatographic profiles (precursor and fragment ions) and ion charge states, whereas Proteomics Quality Control (PTXQC) [10] evaluates MaxQuant results across 24 metrics, including ProteinGroups, Evidence, and tandem mass spectrometry (MS/MS) scans. Although the aforementioned methods consider important evaluation metrics in the field, existing prior knowledge cannot fully capture the characteristics of outliers, leading to shortcomings in monitoring low-quality (outlier) data and instrument status [11]. Therefore, in recent years, researchers have increasingly turned to dedicated packages designed for protein-centric QC, the most representative of which are Ensemble Methods for Outlier Detection (EnsMOD) [12] and protein abundance outlier detection (PROTRIDER) [13]. These tools directly evaluate sample-wise completeness and reproducibility, enabling robust filtering of low-quality samples.

With advances in artificial intelligence, key advantages have emerged, such as scalable analysis and high-precision outlier detection, particularly in high-dimensional proteomics applications [14]. Representative models including Isolation Forest (IForest) [15], Local Outlier Factor (LOF) [16], k-nearest neighbors (KNN) [17], Connectivity-Based Outlier Factor (COF) [18], Angle-Based Outlier Detection (ABOD) [19], Unsupervised Outlier Detection Using Empirical Cumulative Distribution Functions (ECOD) [20], and skewness density ratio outlier factor (SDROF) [21] demonstrate these capabilities. However, proteomics datasets exhibit ultra-high dimensionality and small sample sizes, with feature counts ranging from thousands to tens of thousands per sample [22]. This dimensionality exacerbates the curse of dimensionality in traditional machine learning models, generating spurious correlations that prevent the effective capture of intricate feature interactions [23]. Consequently, their anomaly detection performance becomes significantly compromised when processing proteomic data. In comparison, deep learning models enable dimensionality reduction by mapping high-dimensional data into low-dimensional latent spaces, facilitating efficient representation learning. By capturing complex correlations within these compressed embeddings, they achieve robust anomaly detection. Methods such as Deep One-Class Classification (DeepSVDD) [24], AutoEncoder (AE) [25], Variational AutoEncoder (VAE) [26], Deep Isolation Forest (DIF) [27], and Single-Objective Generative Adversarial Active Learning (SOGAAL) [28] have outperformed conventional approaches across diverse applications. However, these data-hungry models require substantial training datasets [29]. When applied to proteomics studies involving only a few hundred samples, these methods are prone to overfitting, often misclassifying true outliers as normal and compromising data quality.

To address these limitations, we developed memory-constrained outlier detection (MCOD)—a novel deep learning framework for robust outlier detection in quantitative proteomics. MCOD is designed to mitigate overfitting through two key innovations. First, during the training of a standard AE, the anomaly scores of individual samples are recorded, and their feature representations are stored for constructing memory features. A memory-constrained (MC) module is designed so that for each sample, the stored memory feature is used to constrain its current semantic representation, thereby reducing the risk of overlooking true outliers at high precision. Second, an adaptive steady-aware regulator is proposed: to further focus on overfitted samples, MCOD extracts the variance, first-order difference magnitude, and global trend from historical anomaly scores to construct a 3D stability representation for each sample, from which a dynamic adjustment factor is derived. This factor works in concert with the MC module to achieve sample-wise adaptive regularization—stronger constraints are imposed on samples with high overfitting risk, while stable samples are allowed more relaxed learning. In this way, precise selective regularization is achieved at the individual sample level. Figure 1 provides an intuitive illustration of the model’s design philosophy.

Figure 1.

A schematic comparison of conventional and MCOD decision boundaries for outlier detection. MCOD uses a constrain the direction mechanism to shift the learned boundary toward the ideal boundary (red), improving separation between inliers, shown as circles, and outliers, shown as triangles.

The comparison of previous model and MCOD model.

To demonstrate the effectiveness of MCOD, we conducted comprehensive evaluations using three distinct quantitative proteomics datasets. Our results show that MCOD achieves robust outlier detection and leads to meaningful improvements in downstream analyses. We also provide a feature-level visualization to identify proteins contributing most to outlier assignments. Overall, MCOD offers a comprehensive framework for proteomics QC, delivering high accuracy in outlier detection while supporting more reproducible downstream analyses.

Materials and methods

Overview of memory-constrained outlier detection method

MCOD is a deep anomaly detection framework which accepts and processes the protein and peptide quantification results generated by MaxQuant directly. MCOD leverages representation dynamics within feature spaces through three core modules: (i) the Basic module, which serves as the backbone network to map protein features to anomaly score predictions; (ii) the MC module, which mitigates overfitting by constraining the excessive representation of overfitted samples; and (iii) the adaptive steady-aware regulation module, which assesses the fitting stability of individual samples and derives an adjustment factor from this stability, enabling dynamic per-sample weight modulation within the MC module. The workflow of the model is presented in Fig. 2.

Figure 2.

A three-module schematic of the MCOD architecture. Module B1 uses an autoencoder to reconstruct input data. Module B2 applies memory constraints and momentum-updated sample features. Module B3 evaluates the stability of historical anomaly scores and generates an adaptive regulator that dynamically adjusts model training. Fire and snowflake symbols indicate trainable and frozen parameters, respectively.

The overview of MCOD model. B1 illustrates the basic architecture of MCOD; mass spectral data are first processed by MaxQuant software and formatted into a tabular structure; the processed data are normalized, missing values are imputed as preprocessing, and the data are then fed into an encoder to obtain a latent vector, while a decoder is used to reconstruct the input; B2: it measures the similarity loss of different samples between the current state and memory feature using the regulator derived from the B3 module, and updates the memory feature of each sample through a momentum update approach; B3: it collects the anomaly scores of each sample over a specified number of iterations, evaluates the stability of the model’s representation with respect to each sample, and generates a regulator based on this stability to dynamically adjust the loss in B1.

Basic module

We introduce a baseline solution based on an AE to address proteomic data outlier detection. The baseline model consists of a feature encoder Inline graphic with parameters θ and a decoder Inline graphic with parameters ϕ. The whole network Inline graphic maps the input space directly to latent vector via encoder and re-maps to the output space by decoder, which is defined as: Inline graphic. The parameters θ and ϕ are optimized by a reconstruction loss, which is extensively used in AE-based modeling tasks. Specifically, MCOD minimizes the mean squared reconstruction loss:

graphic file with name DmEquation1.gif (2.1)

where N is the sample size, Inline graphic is the input feature vector of the Inline graphic sample, and Inline graphic is the model’s reconstructed output.

Memory-constrained module

The characteristic scarcity of proteomics samples coupled with ultra-high dimensionality predisposes models to rapid overfitting during training, necessitating the preservation of early-stage pre-overfitting representations to constrain sample-specific overlearning. Therefore, we introduce the MC module, which constructs a comprehensive state training archive, continuously capturing each sample’s transition from non-steady to steady representation states throughout model training. This dynamic recording provides regularization constraints, specifically enabling targeted optimization toward pre-overfitting representations. The proposed module involves memory initialization, memory updating, and a memory-constrained loss.

Memory initialization. For each sample, the encoder’s output Inline graphic, obtained from early training epochs well before reaching the total scheduled number of iterations, is stored in the memory-based feature pool. Specifically, the feature vectors are initialized as the representations, i.e.

graphic file with name DmEquation2.gif (2.2)

where Inline graphic denotes the latent representation of the Inline graphic samples, obtained by performing inference with the model parameters frozen at the k-th epoch, and K indicates the total number of epochs used for memory initialization.

Memory updating. During training, the representation memory of the n-th sample, Inline graphic, is stored and updated via a momentum-based mechanism. Since the initial representations are not necessarily optimal for all samples, the feature vectors corresponding to each mini-batch are progressively refined using representations transferred across epochs, following the update protocol:

graphic file with name DmEquation3.gif (2.3)

where m is the momentum updating factor, the initial memory representation for the Inline graphic sample is computed by Equation (2.2), and Inline graphic denotes the latent semantic representation obtained by freezing the encoder weights at the end of the i-th epoch after memory initialization.

Memory-constrained loss. The goal of this module is to encourage the learned feature representations of samples to stay as close as possible to their corresponding memory features. The distance function is implemented through an inner product, which is bounded above and thus conducive to optimization:

graphic file with name DmEquation4.gif (2.4)

where N is the sample size, Inline graphic denotes the inner product operation of normalized features, and τ is a temperature value that is set as 0.5 in this model.

Adaptive steady-aware regulator module

Since the MC module aims to regularize the distance between the semantic representations produced by the feature encoder Inline graphic and the memory features in order to mitigate sample overfitting, the state and quality of the memory features themselves become critical to the constraint’s effectiveness. To prevent the subsequent constraint from introducing negative feedback due to unreliable memory features, we introduce the adaptive steady-aware regulator (ASAR) module that leverages historical anomaly score analysis to quantify the stability of sample-wise reconstruction. The module comprises two stages: (a) steady measurement, which extracts multi-scale stability features from historical anomaly scores; and (b) regulator calculation, which converts the measured stability into a dynamic adjustment factor.

Steady measurement. This module first identifies samples with severe instability by analyzing three key characteristics derived from each sample’s historical outlier scores: first-order difference magnitude, variance, and global trend. The overall flowchart is shown in Fig. 3.

Figure 3.

A flowchart illustrating extraction of a stability vector from the historical anomaly scores of an individual sample. Scores within a sliding window are used to calculate the first-order difference, standard error, and linear-regression slope, which are combined to represent sample stability.

The extraction process of stability vector; the stability vector is constructed by first calculating the first-order difference magnitude and variance for each sample; the mean values of these two metrics are then computed and combined with the slope obtained from linear regression to form a comprehensive representation of sample stability.

Specifically, starting from epoch e, at each subsequent epoch, the model weights are temporarily frozen to perform full-sample inference and obtain temporary anomaly scores for M consecutive checkpoints. These temporary anomaly scores across sequential checkpoints are recorded into a history information pool. For sample n, let its temporary anomaly scores be denoted as Inline graphic. As the finest indicator, the first-order difference magnitude is computed as the average absolute first-order difference over the M temporary anomaly scores, capturing the instantaneous rate of change of the anomaly score within the evaluation window. The formula is given as follows:

graphic file with name DmEquation5.gif (2.5)

The standard deviation of the n-th sample, Inline graphic, quantifies the dispersion of its anomaly scores. The average variance across M epochs serves as a key metric for assessing representation stability during model training. The global trend Inline graphic is represented by the slope of a least-squares linear fit, which quantifies the temporal evolution rate of the anomaly scores within the cohort and serves as a critical indicator of representation stability. Finally, for each sample, we construct a 3D stability vector:

graphic file with name DmEquation6.gif (2.6)

where Inline graphic indexes the evaluation window. It quantifies the instantaneous rate of representation change for the Inline graphic sample. Higher values indicate accelerated manifold drift. The pseudocode of the module is presented in Algorithm 1.

Algorithm 1.

Pseudo-code of ASAR (B3 module).

Input: Proteomic sample i: Inline graphic
Output: Regulator Inline graphic  
  1. Initialization: 3D stability vector of sample i:Inline graphic; Anomaly score pool with length M of sample i: Inline graphic; the matrix of regulator: Inline graphic

  2. for  epoch in E  do

  3.  /* Standard AE model training and loss optimization*/

  4.  Inline graphic

  5.  Loss (AS,P)

  6.  /* At the e-th epoch end, the ASAR module is triggered. */

  7.  if  epoch≥ e  then

  8.   /* The standard AE is set to evaluation mode with its weights frozen. */

  9.   Freeze (AE)

  10.  /* Put temporary anomaly scores into queue of ASPool. */

  11.  Inline graphic

  12.  if  epoch-e==M  then

  13.  /*Calculating the 3D stability vector of sample i  */

  14.  for  Inline graphic  in  Inline graphic  do

  15.   Inline graphic

  16.   Inline graphic

  17.   Inline graphic

  18.   Inline graphic

  19.  end

  20.  /* 3D stability vectors are split into two clusters using the K-means algorithm. */

  21.  Inline graphic

  22.  /* Calculating the L2 distance of samples between cluster centers */

  23.  Inline graphic

  24. End

  25. /* Set the AE weights back to unfreeze. */

  26. Unfreeze (AE)

  27. end

  28. return  Inline graphic

Regulator calculation. After clustering, the stable cluster is identified as the one with the lower average centroid value. The stability vectors are first processed by taking the absolute value of each component and then normalizing to [0,1] using Min-Max scaling, ensuring that all features are non-negative and comparable before clustering. Then, the Euclidean distance from each sample to the stable cluster centroid is then computed as:

graphic file with name DmEquation7.gif (2.7)

where Inline graphic is the distance measurement function of two vectors, Inline graphic is the stability vector representation of Inline graphic sample, and Inline graphic is the centroid of stable cluster, which is calculated by K-Means. Note that G is set to two in this work, based on the following considerations. After adequate training, the representations of the vast majority of samples tend to stabilize, and only a small fraction exhibit drastic fluctuations. In this regime, partitioning the stability states into two groups is sufficient to characterize varying degrees of overfitting risk, rendering a finer-grained division unnecessary. This design avoids introducing excessive hyperparameters and the ambiguity of cluster assignment, while preserving the effectiveness of selective regularization and simultaneously reducing computational overhead and tuning difficulty.

Total loss function

Based on the foregoing discussion, Equation (2.4) with regulator can be written as:

graphic file with name DmEquation8.gif (2.8)

where the metric Inline graphic is computed by Equation (2.7), and to preserve its monotonicity during stability computation, we apply an exponential transformation to obtain the weighting factor Inline graphic. Thus, the overall loss is defined as:

graphic file with name DmEquation9.gif (2.9)

where Inline graphic is the reconstruction loss of AE model, Inline graphic is the memory-constrained loss, and Inline graphic is an indicator function that equals 1, when the training epoch is greater than or equal to e, and 0 otherwise.

Experiments details

Tool implementation

The MCOD framework was implemented entirely in Python (supporting versions 3 via the six library, utilizing Numerical Python (NumPy), Scientific Computing in Python (SciPy), scikit-learn, and Python Outlier Detection (PyOD)). Deployment occurred on an Ubuntu 24.04 server equipped with an Intel Xeon E5-2683 processor. This implementation fully incorporates core benefits characteristic of PyOD, including runtime (JIT (just-in-time)) compilation, broad operating system support, a consistent programmatic interface, comprehensive guides, and practical sample code. Access to the MCOD source code is provided on its GitHub repository (https://github.com/whisperH/MCOD) under the Apache 2.0 license. Installation procedures and operational instructions are detailed within supplementary documentation hosted on the repository.

Baseline models

Since proteomics data typically lacks standardized ground-truth annotations, our evaluation is confined to unsupervised algorithms. Specifically, we conduct a comparative analysis between MCOD and 18 other established unsupervised methods in the field of proteomic outlier detection. These methods include Gaussian Mixture Modeling [30], IForest [15], LOF [16], KNN [17], COF [18], ABOD [19], ECOD [20], Clustering-Based Local Outlier Factor (CBLOF) [31], FeatureBagging [32], Locally Selective Combination of Parallel Outlier Ensembles (LSCP) [33], Histogram-based Outlier Score (HBOS) [34], One-Class Support Vector Machines (OCSVM) [35], and Principal Component Analysis (PCA) [36] represent classical statistical learning approaches for outlier detection that remain widely used. Meanwhile, numerous deep learning models, including DeepSVDD [24], AE [30], VAE [26], DIF [27], and SOGAAL [28], are also considered in this study. All baseline methods were sourced from the PyOD library [37], a comprehensive Python library for outlier detection, and they were implemented with the same dataset partitioning and preprocessing procedures.

Datasets and evaluation protocol

Datasets description

To evaluate the effectiveness and applicability of MCOD, we used three large-scale datasets from label-free proteomics analysis via data-dependent acquisition MS (DDA-MS). Missing values in all three datasets were imputed using the global minimum value, a widely used approach for such cases [38]. Because label-free proteomics data are relative rather than absolute, and unsupervised anomaly detection relies on the internal structure of the data, all models were directly applied to the entire dataset; for the simulated HeLa data, anomaly scores were evaluated against the known simulated labels. The details of the datasets used are shown in Table 1.

Table 1.

Details of datasets used.

Datasets Samples Dimensions Anomaly
HeLa-simulation 325 6448 29
HCC 206 9238 –
LUAD 206 11246 –

The first dataset comprised simulated outliers based on a HeLa cell line DDA-MS dataset (data collection procedure is detailed in our earlier publication [39]). Following raw data acquisition, the dataset was refined by using protein identification numbers, correlation coefficients, and existing outlier detection methods to eliminate potential outliers [40]. After that, the 296 high-quality samples were randomly divided into 10 cohorts, each with about 29 samples. In each of the 10 experimental iterations, one cohort was designated as the outlier group. To evaluate model performance more objectively, we employed two different strategies for generating outlier samples. The first is random feature shuffling on non-imputed data, in which we perform systematic data corruption by gradually increasing the shuffle ratio from 1% up to 40% in 1% increments per step. The second is Gaussian noise injection, where the fraction of features to be perturbed is randomly sampled from a uniform distribution Inline graphic. For each selected feature, its standard deviation across samples is computed; Gaussian noise with a mean of 0 and a standard deviation of 0.5σ is added to the original feature values to construct the anomalous data [41]. These approaches allow for a systematic assessment of the robustness of different models under varying degrees of feature changes. The entire process was repeated 10 times.

We further validated our framework using clinically annotated proteomic cohorts from the Chinese Human Proteome Project, encompassing hepatocellular carcinoma (HCC, 103 patients) and lung adenocarcinoma (LUAD, 103 patients) [39, 42]. All samples were processed using standardized Filter Aided Sample Preparation (FASP) digestion coupled with DDA-MS acquisition (6 fractions for HCC; 10 for LUAD). The raw mass spectrometry data comprised 618 MS runs (HCC) and 2060 runs (LUAD), processed through MaxQuant (v2.0.3.0) [43]. Experimental methodologies—detailing tissue procurement, liquid chromatography-tandem mass spectrometry (LC–MS/MS) configurations, and QC protocols—followed established procedures in primary publications. All data are publicly accessible via iProX under accessions IPX0000937000 and IPX0001804000 [44].

Evaluation protocol

For the HeLa dataset, each batch contains precisely 29 outliers (10% prevalence) with known simulated labels. We therefore used the area under the receiver operating characteristic curve (AUC-ROC) and the area under the precision-recall curve (AUC-PR) as evaluation metrics [45]. The AUC-ROC metric reflects the model’s ability to discriminate between normal and anomalous samples. It is calculated as the area under the ROC curve, with values ranging from 0 to 1—higher values indicate better classification performance. On the other hand, AUC-PR evaluates the model’s effectiveness in identifying anomaly samples, particularly under class imbalance. Similarly, higher AUC-PR values correspond to better detection performance, making it especially informative in scenarios where one class is rare.

Results

Memory-constrained outlier detection evaluation

In the MCOD evaluation, we compared MCOD against 18 representative outlier detection methods, covering traditional machine learning models (e.g. KNN, LOF, IForest), ensemble learning models (e.g. LSCP, FeatureBagging), and deep learning models (e.g. standard AE, VAE, Deep SVDD). All models were assessed under two distinct synthetic anomaly generation strategies mentioned in section Datasets and Evaluation Protocol. As summarized in Fig. 4, MCOD consistently achieved the highest AUC-ROC and AUC-PR across all settings.

Figure 4.

Line and bar charts comparing MCOD with deep learning and conventional machine learning methods using AUC-ROC and AUC-PR. The upper and middle panels show performance across increasing feature-shuffle ratios, and the lower panels show results after Gaussian-noise injection. MCOD generally maintains higher performance across increasing levels of data perturbation.

Comparative performance between MCOD and 18 other models in terms of AUC-ROC and AUC-PR on feature shuffled and Gaussian noise injection simulated dataset; (a, b, c, d) shows results on feature-shuffled data, while (e, f) shows results on Gaussian noise injection data; (a) and (b) depict the performance comparison between MCOD and deep learning models on AUC-ROC and AUC-PR, respectively; (c) and (d) show the comparison between MCOD and other machine learning models on AUC-ROC and AUC-PR; the x-axis represents the proportion of randomly shuffled features, while the y-axis indicates the model’s performance measured by AUC-ROC and AUC-PR under each corruption level; each perturbation level was tested over 10 independent runs, with results averaged to ensure robustness; (e) and (f) display the performance of all models on the Gaussian noise-injected simulated dataset.

Specifically, on the feature-shuffled simulated dataset, AE-based mode ls (including AE and VAE) achieved better performance than DeepSVDD, SOGAAL, and DIF (Fig. 4(a) and (b), Supplementary Tables Sa1 and Sa2). MCOD follows the same reconstruction-based approach as AE and VAE, and benefits from the AE architecture. However, the AE and VAE models significantly underperform MCOD in terms of AUC-PR. This result demonstrates that MCOD effectively enhances the detection capability for anomalies in small-sample scenarios (Fig. 4(b)). Compared to statistical learning models, MCOD also achieves optimal performance (Fig. 4(c) and Fig. 4(d), Supplementary Tables Sb1 and Sb2). Among the top five models ranked by AUC-PR, the leading methods besides MCOD are FeatureBagging, LSCP, LOF, and PCA. FeatureBagging is an ensemble method based on LOF, while LSCP integrates both LOF and KNN. Given that LOF is specifically designed for detecting local outliers, the strong performance of MCOD suggests its effectiveness in identifying such anomalies. Furthermore, since both MCOD and PCA incorporate dimensionality reduction capabilities, MCOD is likewise able to preserve critical information in proteomics data.

On the Gaussian noise injection dataset, MCOD again achieves the best performance; three of the top five models—PCA, AE, and VAE—are all based on dimensionality reduction, indicating that for ultra-high-dimensional proteomics data, reducing the feature space can effectively mitigate the impact of noise (Fig. 4(e) and Fig. 4(f)). This observation carries several implications for anomaly detection in ultra-high-dimensional proteomics data. The strong performance of decomposition-based models can be attributed to their ability to project data onto a lower-dimensional manifold where the essential biological signal is retained while high-frequency noise, including the injected Gaussian perturbations, is largely discarded. Building on this principle, the MC module in MCOD stores stable historical representations for each sample and penalizes large deviations from these memory features. This mechanism effectively prevents the encoder from drifting toward noise-driven representations during training, thereby further enhancing detection robustness.

Effectiveness of each module and hyperparameters

To evaluate the contribution of each component in the proposed MCOD method, we conducted baseline model comparisons, a standard AE with early stopping [46], an ablation study, and a sensitivity analysis of three key hyperparameters: the start of epochs for memory feature initialization, the window size used for stability assessment, and the number of training epochs. All experiments were performed under the most challenging setting: a feature shuffle ratio ranging from 1% to 5%, with the anomaly ratio held constant, and each experiment was repeated 10 times. For the early stopping strategy, we randomly sampled 20% of the training data as a validation set to determine whether to stop training. The patience was set to 5, meaning that training was terminated if the reconstruction loss on the validation set did not decrease for five consecutive epochs. The performance was assessed using both AUC-ROC and AUC-PR metrics. The effectiveness of each module is shown in Fig. 5. As shown in Fig. 5, the introduction of the MC module leads to significant improvements in AUC-PR, with performance gains ranging from 0.057 to 0.215 across perturbation ratios from 1% to 5%, further demonstrating the proposed model’s strong capability and reliability in detecting anomalies. Even compared with an AE that employs an early stopping strategy, MCOD still achieves improvements in AUC-PR ranging from 0.037 to 0.573. A possible reason is that early stopping requires a portion of the training data to be held out as a validation set (20% in our setting), which reduces the number of samples available for model learning. In proteomics, where samples are extremely scarce, such a sacrifice can noticeably degrade model performance. In contrast, MCOD fully utilizes all training samples by integrating an MC module and ASAR to prevent overfitting without the need for a separate validation set. Furthermore, to demonstrate that the performance gain of the ASAR module is not due to random perturbation, we performed paired two-tailed t-tests to assess whether the differences significantly deviated from zero, followed by multiple testing correction, using the SciPy (1.10.1) and statsmodels (0.13.2) packages. The results are shown in Table 2.

Figure 5.

Bar charts comparing mean AUC-PR for AE, AE with early stopping, MCOD with the memory-constraint module, and the complete MCOD model across feature-shuffle ratios of 1%-5%. The MCOD variants show higher AUC-PR than the AE baselines, with asterisks marking statistically significant differences.

Effectiveness of each module on the HeLa simulated dataset; this figure presents the mean and standard deviation of AUC-PR for different variants of the baseline: AE and AE with early stopping (ES) strategy and the MCOD model, evaluated under the most challenging setting with a feature shuffle ratio from 1% to 5%; “MCOD (MC)” denotes the model incorporating only the MC module, without the adaptive steady-aware regulation module; “MCOD (ASAR-MC)” refers to the complete proposed model with both components; “*” denotes a significant difference between MCOD (MC) and MCOD (ASAR-MC).

Table 2.

Results of paired t-tests with FDR correction across noise levels.

Shuffle ratio Mean ΔAUC t-statistic P-value (uncorrected) Cohen’s d P-value (FDR) FDR significant
0.01 0.00189 1.42563 .18772 0.45082 .31287 False
0.02 0.00055 0.13439 .89605 0.04249 .89605 False
0.03 0.01421 3.21169 .01062 1.01562 .03033 True
0.04 0.01127 3.12941 .01213 0.98961 .03033 True
0.05 0.00241 0.80553 .44127 0.254732 .55159 False

As shown in Table 2, only at shuffle ratios of 3% and 4% did the improvement reach significance with P < .05 (0.01062 and 0.01213, respectively). After Benjamini–Hochberg false discovery rate (FDR) correction, these two points remained significant, indicating that the regulator does not generate random gains under arbitrary noise conditions but instead functions stably within a moderate noise range. Furthermore, within the 2% to 4% range, all mean AUC differences were positive, and Cohen’s d effect sizes ranged from 0.04 to 1.02. This directional consistency further rules out the possibility of random fluctuation. When the shuffle ratio is extremely low (1% and 2%), the baseline model itself can already fit the data accurately, and the steady-state constraint of the regulator brings no additional benefit. In general, this comparison suggests that the memory mechanism in MCOD is more sample-efficient and provides a more effective regularization strategy than early stopping in small-sample proteomics applications.

Meanwhile, we also present a sensitivity analysis of the two key parameters: the number of epochs for memory feature initialization, and the window size (i.e. the length of the temporary anomaly score sequence) used for assessing sample stability. As shown in Fig. 6(a), the AUC-ROC metric of the model with a fixed window size initially increases as the memory initialization is delayed, peaking at 40 epochs before subsequently declining. In contrast, in Fig. 6(b), the AUC-PR metric exhibits an overall decreasing trend. Nevertheless, both AUC-ROC and AUC-PR values consistently outperform those of the baseline model.

Figure 6.

Bar charts showing the effects of memory-feature initialization epoch and window-size strategy on MCOD performance. The upper panels compare AUC-ROC and AUC-PR across different initialization epochs with a fixed window size of 10, while the lower panels compare fixed and dynamic window strategies across different window sizes.

The impact of hyperparameters on the MCOD model; the bars and error bars represent the mean and standard deviation of the performance over 10 repeated simulations under a 1% shuffle ratio, respectively; (a) and (b) show the influence of different memory feature initialization epochs on AUC-ROC and AUC-PR, respectively, under a fixed window size of 10; (c) and (d) illustrate the effects of different window sizes on AUC-ROC and AUC-PR under both fixed and dynamic window size strategies.

To investigate how window size affects model performance, we consider two strategies: Dynamic Window Size and Fixed Window Size. Here, “Dynamic window size” refers to initializing the memory feature at a specified iteration, after which the number of historical observations increases until the window size is reached. “Fixed window size” indicates that initialization occurs only once the number of historical observations meets the predefined window size, and this size is maintained throughout training. As shown in Fig. 6(c) and (d), the dynamic strategy yields relatively stable performance across different window sizes. In contrast, under the fixed window size strategy, both AUC-ROC and AUC-PR gradually decline as the window size increases. However, this behavior is consistent with the trends observed in Fig. 6(a) and Fig. 6(b), suggesting that the memory initialization timing—rather than the window size itself—is the fundamental factor influencing performance.

MCOD was originally designed to mitigate overfitting in ultra-high-dimensional, small-sample proteomics data. Therefore, the number of training epochs serves as a natural test indicator for evaluating this capability. To investigate the impact of extended training, we performed an extreme test by increasing the number of epochs from the default 150 to 200, 300, and 500, respectively. As shown in Fig. 7, the performance exhibits only a slight decline across these extended training regimes, demonstrating the effectiveness of MCOD’s design in resisting overfitting even under prolonged optimization.

Figure 7.

Bar charts showing mean AUC-ROC and AUC-PR for MCOD trained for different numbers of epochs, from the default 150 to 500 epochs. Model performance remains relatively stable as the number of training epochs increases.

Effect of the number of training epochs on model performance; the default setting for MCOD is 150 epochs; E denotes the training epoch number.

Individual feature visualization

Kernel density estimation (KDE) is a non-parametric method for estimating the probability density function of a random variable. It provides a smooth representation of the data distribution, avoiding binning artifacts inherent in histograms and revealing finer details of the underlying distribution. We therefore used KDE to visualize the distribution of normal and abnormal samples in terms of protein anomaly scores before and after reconstruction. Specifically, we first standardized the input data, separated the features of normal and abnormal samples along with their reconstructed counterparts, and then visualized the protein features of abnormal and normal samples, respectively. Standardization is performed using scikit-learn (version 1.3.2), while KDE and visualization are implemented with Seaborn (version 0.12.1). The protein feature with the highest anomaly contribution in abnormal samples is shown in Fig. 8. The reconstructed distributions of both the baseline model and the MCOD model exhibit substantial overlap with the distribution of normal samples. In both cases, the reconstructed values are primarily concentrated within the range of −2 to 1, similar to the training data distribution (Fig. 8(a) and (c)). In terms of fitting abnormal samples, the reconstruction distribution of the baseline model closely resembles that of the abnormal samples. In contrast, while the MCOD model’s reconstruction partially overlaps with the abnormal sample distribution, it remains overall closer to the distribution of normal samples, as shown in Fig. 8(b). This indicates that the MCOD model demonstrates superior reconstruction efficacy compared to the baseline model.

Figure 8.

KDE plots showing distributions of the Q9UIJ7 protein feature in the HeLa dataset. Normal and abnormal samples before and after reconstruction are compared with the training-data distribution. Panels a1 and b1 show MCOD results, whereas panels a2 and b2 show results from the baseline model.

KDE of Q9UIJ7 protein feature of HeLa dataset; the left/right figure shows the distribution of normal/abnormal sample before and after reconstruction, both are compared with the distribution of training data; the x-axis represents the values of each sample for feature Q9UIJ7, which have been standardized for better visualization; the y-axis represents the probability density function value, indicating the relative likelihood of data points occurring within a unit interval around a given value; subfigures (a) and (b) display the results of the MCOD model, while (c) and (d) show those of the baseline model.

Assessment of memory-constrained outlier detection on real-world proteomics data

To further validate the practical utility of MCOD, we evaluated its performance on a real-world proteomics dataset without any simulated corruption. Here, the MCOD model was evaluated on HCC and LUAD datasets to assess its performance in QC. It was comprehensively benchmarked against the Villena AE, SEAOP (a Python toolkit designed for robust outlier detection in quantitative proteomics) [40], EnsMOD [12], and PROTRIDER [13], which are the state-of-the-art QC models in proteomic analysis. As HCC and LUAD datasets are oncology datasets, samples analyzed by MCOD were divided into two groups—non-tumor-adjacent tissues (NATs) and tumors—to minimize the potential confounding effects of pathological conditions, whereas the other models were implemented following the protocols specified in the original publications.

The comparison of computation cost

As shown in Table 3, PROTRIDER achieves the fastest runtime, followed by single-model approaches such as the vanilla AE and MCOD, and then by ensemble models like SEAOP and EnsMOD, which integrate multiple base learners. PROTRIDER’s speed advantage stems from its VAE-based architecture with only a single hidden layer, resulting in a small parameter count and consequently the lowest computational cost. MCOD exhibits a higher runtime than the vanilla AE because, after a certain number of iterations, it needs to compute the stability vector and perform clustering, which introduces additional overhead compared to the vanilla AE. The longer runtime of EnsMOD primarily stems from its internal clustering procedure, which exhaustively searches from 1 to 20 cluster centers to obtain more accurate performance estimates; we adopted its default configuration without any modification. However, MCOD’s runtime remains substantially lower than that of the two ensemble models.

Table 3.

Comparison of computational efficiency.

Model name Running time (s)
HCC LUAD
EnsMOD 1977.90 3088.44
PROTRIDER 16.46 17.73
SEAOP 96.53 161.22
Vallina AE 22.63 37.27
MCOD (ours) 28.73 28.48

Note: Because SEAOP and EnsMOD do not support GPU acceleration, all experiments were conducted with GPU disabled to maintain a fair comparison across all methods. Bold values indicate the best performance.

Effect of outlier removal in hepatocellular carcinoma and lung adenocarcinoma data analysis

We performed unsupervised hierarchical clustering of tumor and NAT samples in the LUAD and HCC datasets before and after outlier removal by each of the four models (MCOD, SEAOP, PROTRIDER, and EnsMOD). Clustering of tumor and NATs was conducted using Euclidean distance and average linkage. The outliers identified by each method are annotated in the resulting dendrograms. As shown in Fig. 9, MCOD achieved superior clustering performance compared with the other three models. Notably, in the LUAD dataset, only one cancer sample was incorrectly assigned, and in the HCC dataset, MCOD also produced more accurate clustering of normal samples than the other three methods.

Figure 9.

Hierarchical clustering dendrograms of tumor and normal-adjacent-tissue samples from the LUAD and HCC datasets before and after outlier removal using EnsMOD, SEAOP, PROTRIDER, and MCOD. Red and blue annotation bars indicate tumor and NAT samples, respectively, with the MCOD-processed data showing clearer separation between the two sample classes.

Unsupervised clustering of tumor and NAT samples in the LUAD (a) and HCC (b) datasets, before and after the exclusion of outliers; we employed the Euclidean distance and average linkage method for clustering; in the clustering dendrograms, the key regions of divergence are highlighted with curly braces. The red bar represents the samples with tumor label and blue bar represents the samples with NAT label.

Functional enrichment analysis

To gain biological insight into the impact of outlier removal, we performed differential expression analysis on the HCC data with and without outlier removal, followed by Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses for upregulated and downregulated proteins, as shown in Fig. 10. Enrichment analysis was performed using the clusterProfiler package (version 4.0) in R. Differentially expressed proteins were defined using a significance threshold of Inline graphic and a log2 fold-change cutoff of Inline graphic. Specifically, we focused on a curated set of well-established cancer-related pathways, including proliferation-related processes (e.g. DNA replication) and other cancer-associated processes (e.g., DNA repair, transcription initiation, and elongation), as well as immune-related biological processes commonly dysregulated in tumors (e.g. humoral immune response, immune effector processes). MCOD markedly increased the average enrichment strength (−log10 adjusted P-value) of upregulated cancer pathways in the KEGG analysis, outperforming not only the raw data but also SEAOP, EnsMOD, and PROTRIDER. This indicates that MCOD selectively sharpens the signals of activated oncogenic programs that are central to tumor maintenance. In the GO analysis and for downregulated KEGG pathways, MCOD performed comparably to the existing methods (EnsMOD, SEAOP). Importantly, no artificial enrichment of pathways unrelated to cancer biology was observed after MCOD-based outlier removal. Clustering analysis further confirmed that after removing MCOD-identified outliers, tumor and normal samples formed purer, separate clusters (Fig. 10). While removing outliers improves statistical power, we acknowledge that in heterogeneous diseases such as HCC and LUAD, extreme samples may represent genuine biological subtypes; therefore, we interpret this step as a technical QC measure to enhance data robustness rather than a strict biological filtering criterion. Together, these independent enrichment and clustering analyses demonstrate that MCOD enhances the biological coherence of downstream proteomic analyses by selectively recovering genuine cancer-associated signals rather than simply increasing the number of detected features.

Figure 10.

Comparative GO and KEGG enrichment plots for upregulated and downregulated proteins in the HCC dataset. The panels compare enrichment patterns obtained from the raw data with those obtained after outlier removal using four different detection methods, showing changes in the enriched biological processes and pathways across preprocessing strategies.

The GO and KEGG results of upregulated and downregulated proteins on the HCC data without outlier removal and processed by four outlier detection methods.

Discussion

In this study, we present MCOD, a deep learning-based anomaly detection framework designed for quantitative proteomics data. MCOD first introduces an MC module that employs a momentum-based update algorithm to refresh the semantic representations of samples within the memory unit. By incorporating similarity metrics, we constrain the gradient descent updates in the backbone network, thereby mitigating overfitting in deep learning models applied to high-dimensional proteomic data with limited sample sizes. Considering that individual overfitted samples in the later stages of training substantially influence model updates, we propose a novel approach, aiming to enhance MCOD’s sensitivity toward such samples. This is achieved by constructing a comprehensive representation for each sample based on three features derived from its historical anomaly scores. Building on this diagnostic outcome, we implement a regulator that adaptively weights the memory constraint for each sample according to its degree of overfitting. The effectiveness of the proposed framework has been thoroughly validated on both simulated and real experimental datasets.

In practical applications, the performance of the MCOD model was evaluated against a range of traditional methods—including statistical approaches, machine learning models, deep learning architectures, and ensemble techniques—using a simulated HeLa dataset and two real-world datasets (HCC and LUAD). The results clearly demonstrate the superiority of MCOD. On the simulated HeLa dataset, MCOD achieved significantly higher AUC-PR values across data subsets of varying difficulty levels, outperforming not only leading deep learning models such as DIF, DeepSVDD, AE, VAE, and SOGAAL, but also surpassing ensemble and other machine learning methods by a considerable margin. MCOD demonstrated substantially better performance than some domain methods including SEAOP, EnsMOD, and PROTRIDER across the HCC and LUAD datasets. The advantages were evident across multiple lines of evidence—including the number of significantly enriched GO and KEGG terms and their biological relevance—thus underscoring its practical utility for identifying novel potential outliers.

When integrated with existing model results, MCOD enables more comprehensive QC. At high precision levels, it achieves competitive recall, suggesting a reduced risk of overlooking true outliers. Overall, MCOD demonstrates consistent superiority over contemporary deep learning and machine learning methods, particularly when applied to complex proteomic datasets. Furthermore, MCOD utilizes KDE to analyze the distributional divergence between normal and abnormal samples at the individual protein level, providing guidance for anomaly localization. Extending MCOD to phosphoproteomics is a promising future direction. Tailored developments—including phosphosite-aware anomaly scoring, motif-oriented memory banks, and heterogeneity-adaptive stability evaluation—could further optimize phosphorylation-site-specific outlier detection and we plan to systematically validate these strategies on phosphoproteomics benchmarks.

Overall, we have developed a deep learning model for quantitative proteomics outlier detection, exhibiting notable superiority over proteomic outlier detection algorithms. However, there are still some issues to be addressed, such as developing more interpretable methods within MCOD to rapidly identify which specific proteins contribute to sample anomalies—an advancement that would greatly benefit the optimization of data acquisition protocols and instrument parameter adjustments in future experiments.

Key Points

  • The high-dimensionality and limited sample size typical of proteomics data make it highly susceptible to overfitting in deep learning models. Effectively mitigating this overfitting is therefore a critical challenge for developing robust, deep learning-based quality control algorithms.

  • We developed memory-constrained outlier detection (MCOD), a deep learning-based algorithm specifically designed for outlier detection in proteomics. By integrating memory-constrained module and adaptive steady-aware regulation into the AutoEncoder framework, MCOD enables accurate identification of anomalous samples with minimal manual intervention.

  • Experimental results on quantitative proteomics datasets demonstrate that MCOD effectively detects technical outliers, providing insights into the rationale behind the model’s anomaly assignments.

Supplementary Material

Supplementary_material_bbag507

Acknowledgments

This work was supported by National Key R&D Program of China (32503260, 2022YFF0608404, 2022YFF0705001, 2022FY101202), Plan for Leading Talents of Science and Technology Innovation (No. WR2202), National Natural Science Foundation of China (No. 21927812), and Research Project of the National Institute of Metrology (AKYZD2111).

Contributor Information

Jinze Huang, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Huanyue Liao, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Bo Meng, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Guangkui Fan, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Dong An, College of Information and Electrical Engineering, China Agricultural University, No. 17 Tsinghua East Road, Haidian District, Beijing 100083, China.

Xinhua Dai, China National Institute of Standardization, No. 9 Madian East Road, Haidian District, Beijing 100191, China.

Xiang Fang, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Yang Zhao, Technology Innovation Center of Mass Spectrometry for State Market Regulation, Center for Advanced Measurement Science, National Institute of Metrology, 18, Beisanhuandonglu, Chaoyang District, 100029, Beijing, China.

Author contributions

Jinze Huang (directed and designed research, developed the corresponding algorithmic workflow and software), Xiang Fang (directed and designed research, reviewed the manuscript and provided suggestions for revisions), Yang Zhao (directed and designed research, provided guidance and suggestions for improvements to the analysis workflow and code), Bo Meng (provided guidance and suggestions for improvements to the analysis workflow and code), Dong An (provided guidance and suggestions for improvements to the analysis workflow and code), Guangkui Fan (visualized the analysis results and wrote the manuscript), Huanyue Liao (visualized the analysis results and wrote the manuscript), and Xinhua Dai (reviewed the manuscript and provided suggestions for revisions)

Conflicts of interest

The authors declare no competing interests.

Funding

None declared.

Data availability

All datasets utilized in this research are publicly accessible to promote transparency and allow other researchers to replicate our findings. All data have been deposited on GitHub: https://github.com/whisperH/MCOD. Specifically, HeLa cell line DDA-MS dataset is supported by public research: Jiang, Y.; Sun, A.; Zhao, Y.; Ying, W.; Sun, H.; Yang, X.; Xing, B.; Sun, W.; Ren, L.; Hu, B.; et al. Proteomics identifies new therapeutic targets of early-stage hepatocellular carcinoma. Nature  2019, 567 (7747), 257–261. DOI: 10.1038/s41586-019-0987-8. Meanwhile, HCC datasets and LUAD dataset (IPX0000937000 and IPX0001804000) are available from iProX in 2021: connecting proteomics data sharing with big data.

Code availability

The source code and trained logs for MCOD are freely available under the Apache 2.0 license and can be found at: https://github.com/whisperH/MCOD.

References

  • 1. Zhao  Y, Xue  Q, Wang  M  et al. Evolution of mass spectrometry instruments and techniques for blood proteomics. J Proteome Res  2023;22:1009–23. 10.1021/acs.jproteome.3c00102 [DOI] [PubMed] [Google Scholar]
  • 2. Giudice  G, Petsalaki  E. Proteomics and phosphoproteomics in precision medicine: applications and challenges. Brief Bioinform  2019;20:767–77. 10.1093/bib/bbx141 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Dai  C, Pfeuffer  J, Wang  H  et al. Quantms: a cloud-based pipeline for quantitative proteomics enables the reanalysis of public proteomics data. Nat Methods  2024;21:1603–7. 10.1038/s41592-024-02343-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Yan  G-J, Mei  X, Yang  J  et al. A novel real-time cell electronic analysis technology for the authentication and quality control of natural medicines. Chin Chem Lett  2013;24:1140–4. 10.1016/j.cclet.2013.07.016 [DOI] [Google Scholar]
  • 5. Tsantilas  KA, Merrihew  GE, Robbins  JE  et al. A framework for quality control in quantitative proteomics. J Proteome Res  2024;23:4392–408. 10.1021/acs.jproteome.4c00363 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Hsiao  Y, Zhang  H, Li  GX  et al. Analysis and visualization of quantitative proteomics data using FragPipe-analyst. J Proteome Res  2024;23:4303–15. 10.1021/acs.jproteome.4c00294 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Mann  M, Kumar  C, Zeng  WF  et al. Artificial intelligence for proteomics and biomarker discovery. Cell Syst  2021;12:759–70. 10.1016/j.cels.2021.06.006 [DOI] [PubMed] [Google Scholar]
  • 8. Huang  L, Zhang  D, Jiao  L  et al. A new quality control method for lateral flow assay. Chin Chem Lett  2018;29:1853–6. 10.1016/j.cclet.2018.11.028 [DOI] [Google Scholar]
  • 9. Rozanova  S, Uszkoreit  J, Schork  K  et al. Quality control-a stepchild in quantitative proteomics: a case study for the human CSF proteome. Biomolecules  2023;13:491. 10.3390/biom13030491 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Bielow  C, Mastrobuoni  G, Kempa  S. Proteomics quality control: quality control software for MaxQuant results. J Proteome Res  2016;15:777–87. 10.1021/acs.jproteome.5b00780 [DOI] [PubMed] [Google Scholar]
  • 11. Ahmed  F, Mahmood  T, Riaz  M  et al. Comprehensive review of high-dimensional monitoring methods: trends, insights, and interconnections. Qual Technol Quant Manag  2024;22:727–51. 10.1080/16843703.2024.2395745 [DOI] [Google Scholar]
  • 12. Manes  NP, Song  J, Nita-Lazar  A. EnsMOD: a software program for omics sample outlier detection. J Comput Biol  2023;30:726–35. 10.1089/cmb.2022.0243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Klaproth-Andrade  D, Scheller  IF, Tsitsiridis  G  et al. PROTRIDER: protein abundance outlier detection from mass spectrometry-based proteomics data with a conditional autoencoder. Bioinformatics  2025;41:btaf628. 10.1093/bioinformatics/btaf628 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Mikuni  V, Nachman  B. High-dimensional and permutation invariant anomaly detection. SciPost Phys  2024;16:062. 10.21468/SciPostPhys.16.3.062 [DOI] [Google Scholar]
  • 15. Liu  FT, Ting  KM, Zhou  Z-H. Isolation forest. In: Giannotti F, Gunopulos D, Turini F et al. (eds.), Eighth IEEE International Conference on Data Mining, pp. 413–22. New York, United States: IEEE Computer Society, 2008.
  • 16. Breunig  MM, Kriegel  H-P, Ng  RT  et al.  LOF: identifying density-based local outliers. ACM SIGMOD Rec  2000;29:93–104. 10.1145/335191.335388 [DOI] [Google Scholar]
  • 17. Ramaswamy  S, Rastogi  R, Shim  K. Efficient algorithms for mining outliers from large data sets. In: Chen W, Naughton JF, Bernstein PA (eds.), Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, pp. 427–38. New York, NY, USA: Association for Computing Machinery, 2000.
  • 18. Tang  J, Chen  Z, Fu  AW-C  et al. Enhancing effectiveness of outlier detections for low density patterns. In: Chen M-S, Yu PS, Liu B (eds.), Advances in Knowledge Discovery and Data Mining, pp. 535–48. Berlin, Heidelberg: Springer Berlin Heidelberg, 2002. [Google Scholar]
  • 19. Kriegel  H-P, Schubert  M, Zimek  A. Angle-based outlier detection in high-dimensional data. In: Li Y, Liu B, Sarawagi S (eds.), Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 444–52. New York, NY, USA: Association for Computing Machinery, 2008. [Google Scholar]
  • 20. Li  Z, Zhao  Y, Hu  X  et al. ECOD: unsupervised outlier detection using empirical cumulative distribution functions. IEEE Trans Knowl Data Eng  2023;35:12181–93. 10.1109/TKDE.2022.3159580 [DOI] [Google Scholar]
  • 21. Zhang  Z, Wang  K, Dong  J  et al. SDROF: outlier detection algorithm based on relative skewness density ratio outlier factor. Appl Intell  2024;55:67. 10.1007/s10489-024-06092-8 [DOI] [Google Scholar]
  • 22. Lualdi  M, Fasano  M. Statistical analysis of proteomics data: a review on feature selection. J Proteome  2019;198:18–26. 10.1016/j.jprot.2018.12.004 [DOI] [PubMed] [Google Scholar]
  • 23. Feldner-Busztin  D, Nisantzis  PF, Edmunds  SJ  et al. Dealing with dimensionality: the application of machine learning to multi-omics data. Bioinformatics  2023;39:btad021. 10.1093/bioinformatics/btad021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Ruff  L, Vandermeulen  R, Goernitz  N  et al. Deep one-class classification. In: Jennifer  D, Andreas  K (eds.), Proceedings of the 35th International Conference on Machine Learning, pp. 4393–402. Proceedings of Machine Learning Research: PMLR, 2018. [Google Scholar]
  • 25. Abhaya  A, Patra  BK. An efficient method for autoencoder based outlier detection. Expert Syst Appl  2023;213:118904. 10.1016/j.eswa.2022.118904 [DOI] [Google Scholar]
  • 26. Kingma  DP, Welling  M. Auto-encoding variational bayes. In: Bengio Y, LeCun Y (eds.), 2nd International Conference on Learning Representations, pp. 14. Banff, AB, Canada: ICLR, 2014. [Google Scholar]
  • 27. Xu  H, Pang  G, Wang  Y  et al. Deep isolation forest for anomaly detection. IEEE Trans Knowl Data Eng  2023;35:12591–604. 10.1109/TKDE.2023.3270293 [DOI] [Google Scholar]
  • 28. Liu  Y, Li  Z, Zhou  C  et al. Generative adversarial active learning for unsupervised outlier detection. IEEE Trans Knowl Data Eng  2019;32:1–1528. 10.1109/TKDE.2019.2905606 [DOI] [Google Scholar]
  • 29. Boehm  KM, Khosravi  P, Vanguri  R  et al. Harnessing multimodal data integration to advance precision oncology. Nat Rev Cancer  2022;22:114–26. 10.1038/s41568-021-00408-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Aggarwal  CC. An Introduction to Outlier Analysis. Outlier Analysis. Cham: Springer International Publishing, 2017. 1–34. [Google Scholar]
  • 31. He  Z, Xu  X, Deng  S. Discovering cluster-based local outliers. Pattern Recognit Lett  2003;24:1641–50. 10.1016/S0167-8655(03)00003-5 [DOI] [Google Scholar]
  • 32. Lazarevic  A, Kumar  V. Feature bagging for outlier detection. In: Grossman R, Bayardo R, Bennett KP (eds.), Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining, pp. 157–66. New York, NY, USA: Association for Computing Machinery, 2005. [Google Scholar]
  • 33. Zhao  Y, Nasrullah  Z, Hryniewicki  MK  et al. LSCP: locally selective combination in parallel outlier ensembles. In: Berger-Wolf T, Chawla N (eds.), Proceedings of the 2019 SIAM International Conference on Data Mining (SDM), pp. 585–93. Philadelphia, PA, USA: Society for Industrial and Applied Mathematics, 2019.
  • 34. Goldstein  M, Dengel  A. Histogram-based outlier score (HBOS): A fast unsupervised anomaly detection algorithm. In: KI-2012: Poster and Demo Track, Vol. 1, pp. 59–63. Germany: DFKI (German Research Center for Artificial Intelligence), 2012.
  • 35. Scholkopf  B, Platt  JC, Shawe-Taylor  J  et al. Estimating the support of a high-dimensional distribution. Neural Comput  2001;13:1443–71. 10.1162/089976601750264965 [DOI] [PubMed] [Google Scholar]
  • 36. Shyu  M-L, Chen  S-C, Sarinnapakorn  K  et al. Principal component-based anomaly detection scheme. In: Young  LT, Ohsuga  S, Liau  C-J  et al. (eds.), Foundations and Novel Approaches in Data Mining, pp. 311–29. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006. [Google Scholar]
  • 37. Zhao  Y, Nasrullah  Z, Li  Z. PyoD: a python toolbox for scalable outlier detection. J Mach Learn Res  2019;20:1–7. [Google Scholar]
  • 38. Wang  S, Li  W, Hu  L  et al. NAguideR: performing and prioritizing missing value imputations for consistent bottom-up proteomic analyses. Nucleic Acids Res  2020;48:e83. 10.1093/nar/gkaa498 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Jiang  Y, Sun  A, Zhao  Y  et al. Proteomics identifies new therapeutic targets of early-stage hepatocellular carcinoma. Nature  2019;567:257–61. 10.1038/s41586-019-0987-8 [DOI] [PubMed] [Google Scholar]
  • 40. Huang  J, Zhao  Y, Meng  B  et al. SEAOP: a statistical ensemble approach for outlier detection in quantitative proteomics data. Brief Bioinform  2024;25:bbae129. 10.1093/bib/bbae129 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Wen  F, Su  Y, Liu  D  et al. Automated sparse feature selection in high-dimensional proteomics data via 1-bit compressed sensing and K-medoids clustering. BMC Bioinformatics  2025;26:165. 10.1186/s12859-025-06193-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Xu  JY, Zhang  C, Wang  X  et al. Integrative proteomic characterization of human lung adenocarcinoma. Cell  2020;182:e217. 10.1016/j.cell.2020.05.043 [DOI] [PubMed] [Google Scholar]
  • 43. Cox  J, Mann  M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat Biotechnol  2008;26:1367–72. 10.1038/nbt.1511 [DOI] [PubMed] [Google Scholar]
  • 44. Chen  T, Ma  J, Liu  Y  et al. iProX in 2021: connecting proteomics data sharing with big data. Nucleic Acids Res  2022;50:D1522–7. 10.1093/nar/gkab1081 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Sofaer  HR, Hoeting  JA, Jarnevich  CS  et al. The area under the precision-recall curve as a performance metric for rare binary events. Methods Ecol Evol  2019;10:565–77. 10.1111/2041-210X.13140 [DOI] [Google Scholar]
  • 46. Yao  Y, Rosasco  L, Caponnetto  A. On early stopping in gradient descent learning. Constr Approx  2007;26:289–315. 10.1007/s00365-006-0663-2 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary_material_bbag507

Data Availability Statement

All datasets utilized in this research are publicly accessible to promote transparency and allow other researchers to replicate our findings. All data have been deposited on GitHub: https://github.com/whisperH/MCOD. Specifically, HeLa cell line DDA-MS dataset is supported by public research: Jiang, Y.; Sun, A.; Zhao, Y.; Ying, W.; Sun, H.; Yang, X.; Xing, B.; Sun, W.; Ren, L.; Hu, B.; et al. Proteomics identifies new therapeutic targets of early-stage hepatocellular carcinoma. Nature  2019, 567 (7747), 257–261. DOI: 10.1038/s41586-019-0987-8. Meanwhile, HCC datasets and LUAD dataset (IPX0000937000 and IPX0001804000) are available from iProX in 2021: connecting proteomics data sharing with big data.


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES