Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jun 9;25(7):3384–3398. doi: 10.1021/acs.jproteome.5c01235

ProteoForge: An Imputation-Aware Framework for Differential Proteoform Discovery in Bottom-Up Proteomics

Enes K Ergin †,‡,*, Agustina Conrrero ‡, Kirsty M Ferguson †,‡, Philipp F Lange †,‡,*
PMCID: PMC13340428  PMID: 42261773

Abstract

The human genome contains approximately 20,000 protein-coding genes. However, millions of diverse protein variants, called proteoforms, exist. Despite originating from the same gene, proteoforms often have distinct biological roles. In bottom-up proteomics, the aggregation of peptide measurements into protein-level quantities often obscures this information. Existing methods for differential proteoform discovery are limited by their handling of missing data, which can introduce a significant bias. To address this, we developed ProteoForge, which builds on an imputation-aware statistical model to identify and group covarying peptides into quantitatively differential proteoforms (dPFs). Benchmarking against existing methods demonstrated that ProteoForge provides high accuracy and stability in data sets with high rates of missing values, complex experimental designs, or varying signal strengths. Application of ProteoForge to proteomics data from lung cancer cells under hypoxia revealed extensive proteoform-level regulation hidden by a standard protein-level analysis.

Keywords: Proteomics, Bioinformatics, Proteoform, Differential Discovery, Quantification, Hypoxia


graphic file with name pr5c01235_0008.jpg


graphic file with name pr5c01235_0006.jpg

1. Introduction

The human genome contains approximately 20,000 protein-coding genes. However, contrary to a simplified view of the central dogma, millions of diverse protein forms exist. , Made through mechanisms such as alternative splicing and post-translational modifications (PTMs), , proteoforms originating from the same gene are not functionally redundant and can exhibit different or even antagonistic roles. − This limits the suitability of gene or RNA-level analysis alone in explaining biological phenomena and underscores the need for analysis methods that encompass proteoform diversity. ,

Bottom-up proteomics is a standard method for large-scale protein analysis where proteins are digested into smaller peptides before measurement. This process introduces the ″protein inference problem″: a single peptide sequence can map to multiple proteins, isoforms, or proteoforms, making it difficult to trace back to its specific origin. − To manage this ambiguity, peptide measurements are often aggregated into ″protein groups″ following the Occam’s-razor principle, which simplifies data analysis but obscures the quantitative behavior of individual proteoforms. − Top-down proteomics analyzes intact proteoforms; however, its application remains limited by technical constraints. As an alternative, several computational methods have been developed to infer proteoform-level changes from bottom-up data. − Established methods such as PeCorA (Peptide Correlation Analysis) detect quantitative changes in single modified peptides, while COPF (Correlation-based functional ProteoForm assessment) and the hierarchical clustering approach of Lukasse et al. (2014) identify coregulated peptide groups that may define a proteoform. ,, AlphaQuant propagates quantitative confidence hierarchically from fragment ions to genes using counting statistics, and MSstatsWeightedSummary resolves shared-peptide abundances via convex optimization. ,

The existing tools address different facets of the proteoform problem and address missing values to different degrees, making robust differential proteoform discovery from typical sparse bottom-up proteomics data a largely unsolved problem. AlphaQuant and MSstatsWeightedSummary focus on resolving known shared peptides or propagating quantification confidence, while PeCorA and COPF target de novo detection of peptide-level quantitative discordances. MSstats addresses censored missing values through an accelerated failure time (AFT) model at the summarization stage, MSstatsWeightedSummary resolves shared peptide abundances via convex optimization but currently ignores missing feature intensities, AlphaQuant sidesteps intensity-based imputation through counting statistics, and OLS-based methods like PeCorA treat imputed values as equally reliable as measured intensities. A framework is therefore needed that robustly identifies differentially abundant as well as condition-specific proteoforms avoiding information loss during protein inference while mitigating imputation bias.

Here, we present ProteoForge, a computational framework that identifies discordant peptide-level quantitative patterns de novo from standard, imputed bottom-up proteomics data sets. ProteoForge combines a Huber M-estimation, as originally proposed for PeCorA, with a Robust Linear Model (RLM) to provide ‘imputation-awareness’. Unlike OLS, which treats imputed values as equally reliable as measured intensities, RLM iteratively down-weights observations with large residuals, reducing the influence of imputed data points on parameter estimates. ProteoForge further integrates this robust RLM-linked M-estimation with a peptide-level interaction model originally proposed by PeCorA, coupled with hierarchical clustering to provide de novo proteoform assembly.

We evaluate ProteoForge against COPF and PeCorA frameworks, which share the primary objective of identifying de novo peptide discordance in standard processed data sets. Using both simulated and publicly available proteomics data sets, we compare ProteoForge’s performance and demonstrate its ability to uncover biological insights obscured by conventional protein-level quantification.

2. Materials and Methods

2.1. ProteoForge Framework

The core workflow of ProteoForge involves four modules: (1) data processing and normalization, (2) identification of significantly discordant peptides, (3) clustering of quantitatively similar peptides, and (4) construction of differential proteoforms (dPFs) (Figure A).

1.

1

Overview of the ProteoForge framework. (A) Schematic representation of the four primary stages of the ProteoForge analytical workflow: (1) Data Processing, (2) Identifying Discordant Peptides, (3) Clustering Peptides, and (4) Building Proteoforms. (B) Density plots demonstrating the impact of different normalization stages on the peptide intensity distributions for four conditions (control, cond-1, cond-2, cond-3). (C) Examples for three proteins (Protein 1, Protein 2, and Protein 3) illustrating the identification of Significantly Discordant Peptides. Plots show the Control Adjusted Intensity (Y-axis) for individual peptides (X-axis, ordered by location) across the conditions. Error bars represent the 95% confidence interval (CI) around the mean of replicates. The weighted linear model used to identify significant peptide-condition interaction; orange stars indicate significantly discordant peptides (p < 0.001). Examples of Peptide Clustering (three heatmaps) and Proteoform Building (right text). Heatmaps show the scaled Euclidean distance between all peptide pairs (color bar, 0.0 to 1.0) for each protein. The resulting Canonical peptide groups, final differentially expressed Proteoforms (dPFs), and nongrouped (Single peptide PTMs) are listed to the right.

2.1.1. Module 1: Data Processing and Normalization

This module selects fully quantified peptides from proteins with at least four peptides. Prior imputation is required if there are missing values. The data is normalized to highlight changes relative to the control, standard, or normal samples. Raw nonlog intensities are transformed into a logarithmic scale. Each sample undergoes z-score normalization: z i = (y i – μ)/σ, where y i is the log-transformed value, μ is the mean, and σ is the standard deviation, to ensure a mean of zero and a standard deviation of one. To establish a baseline for each peptide relative to its control average, the average z-scored intensity of each peptide across control samples (C p) is calculated and subtracted from the corresponding peptide’s z-scored intensity in each sample: A p,i = z p,i – C p. This facilitates more precise comparisons between experimental conditions and highlights potential changes in peptide abundance (Figure B). For example, in Section normoxia 72 h served as the control. For experimental designs without a designated control, ProteoForge supports grand-mean centering, where the peptide-level mean across all conditions is subtracted, enabling application to multicondition factorial or time-course designs. For example, in Section Comparison 3, three hypoxia conditions were compared without a control.

2.1.2. Module 2: Imputation-Aware Identification of Significantly Discordant Peptides

A linear model with weights is used to identify peptides that are significantly discordant within proteins across conditions. The process iteratively compares each peptide with its sibling peptides, fitting a linear model that includes an interaction term between these groups. Formally, for a given protein group, let y ijk denote the standardized log-intensity of peptide i in condition j and replicate k. The model is specified as y ijk = βPi + βCj + βPi:Cj + ε ijk , where βPi is the baseline peptide abundance, βCj is the global condition effect, βPi:Cj is the peptide-condition interaction of interest, and ε ijk is the residual error. While there is a variety of linear models to choose from (details in Note S1), the main model we chose to highlight ProteoForge’s application and benchmarks is RLM. The RLM uses the Huber loss function, which transitions from quadratic (L2) loss for small residuals to linear (L1) loss beyond a tuning constant c (default = 1.345), thereby bounding the influence of any single observation. Model parameters are estimated via Iteratively Reweighted Least Squares (IRLS): at each iteration, observations with residuals exceeding c receive reduced weights, converging to parameter estimates that are resistant to outlying data points, including those arising from imputed values. We distinguish between a technical outlier, a data point generating a large residual (e.g., from faulty imputation or instrument noise) that is automatically down-weighted by the Huber loss function, and a discordant peptide, which exhibits a statistically significant peptide-condition interaction term ( βPi:Cj ) that persists despite robust down-weighting, indicating a systematic biological deviation from the parent protein’s consensus trajectory.

The statistical significance of all interaction terms was determined using a Wald test, with the specific test statistic and its underlying distribution tailored to the linear model used. P-values are adjusted in two stages to control the False Discovery Rate. Stage 1 applies a Bonferroni correction within each protein, penalizing each raw p-value by the number of peptide-level tests for that protein. This prevents peptide-rich proteins from producing spuriously significant results through test volume alone. Stage 2 takes the minimum Bonferroni-adjusted p-value per protein and applies the Benjamini–Hochberg procedure globally across all proteins. , This hierarchical correction is more conservative than a single global adjustment, trading marginal statistical power for robust hit selection. The two-stage procedure is the default; users can alternatively apply protein-only correction or a single global Benjamini–Hochberg adjustment, with empirical false-positive rates for each configuration reported in Note S2. Note S2 (Figure ) presents null-calibration results across K = 50 null simulation runs comparing six correction strategies, confirming that the two-stage procedure maintains the peptide-level false-positive rate well below the nominal α.

3.

3

ProteoForge resolves proteoform heterogeneity to identify specific PTMs and potential proteoform-like events. ProteoForge was applied to a lung cancer cell hypoxia data set to identify and characterize peptide-level regulation. (A) Stacked bar charts show the distribution of protein count and the corresponding total peptide count (48 and 72 h) based on the number of significantly discordant peptides identified (adj. p < 0.001): no single significant peptide, multiple significant peptides, or single significant peptide. (B) Stacked bar charts further classify proteins into functional categories: no significant peptides, single peptide PTM, with dPF, or with single PTM and dPF. (C) Bar plots detail the UniProt annotations for the significant peptides from PTM and dPF proteins (48 h and 72 h). An inset pie chart shows the proportion of these peptides that match a known biological feature with phosphorylation, ubiquitination, and acetylation being predominant modifications. (D,F) Quantitative profiles for peptides from glyoxalase I (GLO1) and fucose-1-phosphate guanylyltransferase (FPGT) are shown across four experimental conditions. Error bars: 95% confidence intervals. (E,G) Schematics for GLO1 and FPGT mapping the locations of the quantified peptides and known protein features.

2.1.3. Module 3: Clustering of Quantitatively Similar Peptides

Peptides are first grouped by calculating a distance matrix based on the Euclidean distance between their z-score-normalized median-adjusted intensity profiles across all samples, ensuring that magnitude differences between high- and low-abundance peptides do not dominate clustering. This distance matrix was used for hierarchical clustering with the Ward linkage method. While the module provides several functions, the default (hybrid_outlier_cut) automatically determines the optimal cluster number. Modules 2 and 3 apply distinct statistical criteria to the same data: the significance of the peptide-condition interaction (Module 2) and the Euclidean distance between median intensity profiles (Module 3), preventing circular inference. Module 4 intersects these independent assessments to declare a dPF.

2.1.4. Module 4: Construction of Differential Proteoforms (dPFs)

To call a set of peptides a differential proteoform (dPF), significantly discordant peptides must belong to a cluster containing at least two peptides (Figure C). For proteins without such clusters, all peptides are assigned to the nonsignificant representation or canonical (dPF_0). When a cluster with significant peptides is identified, its peptides form a new differential proteoform (dPF_1), while the remaining peptides stay in dPF_0. If additional significant clusters are detected, they are numbered sequentially (dPF_2, dPF_3, etc.), as illustrated in Figure C. Due to insufficient evidence, significantly discordant peptides in singleton clusters are excluded from both groups, denoted as dPF_-1. These cases, potentially representing single peptide PTMs, are reported separately for further analysis. Finally, a peptide-to-dPF mapping table is generated. This table uses the naming convention “proteinacc_dPF_N″ to represent each proteoform, where N is the cluster, and dPF_ denotes the nonsignificant representation, excluding significantly discordant peptides. The output preserves protein accessions from the upstream search engine (e.g., DIA-NN, MaxQuant) and appends dPF labels (e.g., P12345_dPF_0, P12345_dPF_1), maintaining traceability to the initial database search. No reinference of protein groups is performed. Figure S9 illustrates the full Module 4 assembly logic on a representative 18-peptide protein, demonstrating how cluster-consistent peptides that individually fall below the significance threshold are recovered through the grouping step.

2.2. Benchmarking Performance of ProteoForge against PeCorA and COPF

ProteoForge’s performance was benchmarked against COPF and PeCorA using both a publicly available mass spectrometry data set with in-silico modifications and a series of controlled simulations.

2.2.1. SWATH-MS Interlab Data Set

A benchmark data set was recreated from the COPF manuscript, which originated from a SWATH-MS interlaboratory study of HEK293 cell lysates across 21 replicate runs performed on days 1, 3, and 5. These acquisition days (1, 3, and 5) represent separate dates of data collection within the reproducibility study, not distinct biological treatments. Because the unmodified data set contains residual technical variation between acquisition days (Note S16), the fully simulated data sets in section isolate method performance from such confounds. After filtering, a ground truth data set was established by introducing two types of in silico changes. First, to simulate differential expression, peptide intensities for samples from days 3 and 5 were multiplied by random factors ranging from 1 to 6. Second, to create artificial proteoforms, a subset of peptides from 1000 randomly selected proteins had their day 5 intensities modified by a random factor between 0.01 and 0.9. The number of modified peptides per protein was varied to be one, two, a random fraction between 2% and 50%, or 50% (rounded up per protein (e.g., perturbing three out of five peptides)). The first three data sets (one, two, or random) directly mirrored the COPF benchmark.

2.2.2. Simulated Scenarios

A base data set was first established to serve as the foundation for four distinct simulation scenarios. Data was generated for 500 proteins across 10 replicates in 3 conditions, with each protein containing between 5 and 50 peptides. Simulation began by sampling mean protein abundances from a normal distribution (mean = 20, standard deviation = 2) and protein-level coefficients of variation (CVs) from a log-normal distribution. The number of peptides per protein was determined using a beta-binomial distribution (α = 0.5, β = 3). These choices reflect empirically characterized properties of bottom-up proteomics data sets: protein abundances span several orders of magnitude and approximate a log-normal distribution on the raw intensity scale, while tryptic peptide counts per protein follow an overdispersed, right-skewed distribution well captured by the beta-binomial model. Peptide-level abundances were subsequently generated, incorporating random noise and a small fraction of outliers (0.01% of peptides) to simulate experimental data.

  • 1.

    Simulation 1: evaluate the impact of data imputation across different peptide perturbation levels (Figure S10). First, a perturbation layer was modeled by varying the number of altered peptides per protein across four categories: two, random (2–50%), 50%, and >50%. For each perturbation level, two data versions were generated: a complete version and an imputed version. For the imputed version, of the 200 proteins designated for imputation testing, 100 received sporadic random dropouts in up to 35% of replicates (imputed via kNN) and 100 had all replicates within one experimental condition set to missing. This simulated a missing-not-at-random (MNAR) scenario (imputed using a downshifted normal distribution).

  • 2.

    Simulation 2: systematically test the effect of data missingness. Twenty-five distinct data sets were created combining different levels of protein and peptide missingness (ranging from 0% to 80%).

  • 3.

    Simulation 3: investigate the magnitude of perturbation by applying log2 fold changes (six different intervals, from (0.1, 0.25) to (1.25, 1.50)).

  • 4.

    Simulation 4: examine multicondition complexity. Data sets were generated with two to six conditions, testing the effects of peptide perturbation overlap and directionality.

Simulation 2 uses the ‘random’ peptide perturbation level as its base, while Simulations 3 and 4 use imputed data and random peptide perturbation, followed by modification of other parameters to establish the experiments.

To validate the statistical calibration of the two-stage correction procedure (section ), a null simulation was generated by retaining the data-generating process of Simulation 1 (500 proteins, Beta-Binomial peptide counts, 10 replicates across 3 conditions) but applying zero biological perturbations. K = 50 independent null data sets were generated under complete-data conditions. Any peptide flagged as significant in these data sets constitutes a false positive. A parallel MNAR-null simulation applied the Simulation 2 missingness pattern (condition-complete missingness with downshifted normal imputation) to the same zero-perturbation data sets (K = 50 runs), isolating the effect of structured missingness on Type I error (Notes S2 and S3).

2.2.3. Performance Metrics

Performance was quantified using receiver operating characteristic (ROC) curves, the area under the curve (AUC), and the Matthews correlation coefficient (MCC), using standard formulas. , These metrics were calculated from a confusion matrix (true/false positives/negatives) by comparing method outputs against the in silico ground truth across a range of significance thresholds. The MCC (−1 to +1) was selected for its balanced, class-imbalance-robust performance: +1 for perfect prediction, 0 for random prediction, and −1 for inverse prediction. The benchmark is formulated as a binary classification: for each peptide within a protein, discordant vs concordant (module 2) and correctly grouped vs incorrectly grouped (module 3).

All AUROC and MCC values are accompanied by 95% confidence intervals from B = 1000 stratified bootstrap iterations. Pairwise DeLong tests evaluate AUROC differences between methods; for metrics without a closed-form variance estimator (Best MCC, matched-FPR sensitivity, precision), pairwise paired permutation tests (n = 10,000) were used. All p-values are Benjamini–Hochberg corrected and reported as q-values. To complement threshold-optimized comparisons, a matched operating point analysis evaluated bootstrapped 95% CIs for Sensitivity, Specificity, Precision, and MCC at a fixed False-Positive Rate (FPR) of 5%, corresponding to the nominal α = 0.05 significance level. The Adjusted Rand Index (ARI) was calculated for the peptide grouping benchmark using ‘sklearn.metrics.adjusted_rand_score’ to provide a clustering-specific evaluation metric alongside MCC.

2.3. Analysis of the Lung Cancer Cell Hypoxia Data Set

Quantitative proteomics data were sourced from Tomin et al., who used bottom-up proteomics to investigate the response of lung cancer cells to hypoxia (PRIDE repository, PXD062503). Our analysis used the precursor-level quantitative data (report.parquet) generated by DIA-NN (version 1.9.2) as the input.

2.3.1. Data Preprocessing

Modified peptide-level data were first calculated by summing the precursor intensities for each peptide. The initial data set consisted of 16 samples (four biological replicates per group, across two time points (48 h vs 72 h) and two conditions (Hypoxia vs Normoxia)). Data were filtered: (i) to include only peptides quantified in at least four biological replicates in one or more experimental groups and (ii) to retain only proteins identified by four or more peptides. This resulted in a final matrix of 95,369 peptides across 7161 proteins.

Missing values were handled using a two-step imputation strategy. First, peptides with values completely missing across all replicates of a condition were imputed using values drawn from a normal distribution derived from the 0.1 percentile of observed intensities, shifted down by a magnitude of two. Second, all remaining sparse missing values were imputed using a k-Nearest Neighbors (kNN) approach. Finally, the complete data matrix was normalized using the ProteoForge control-based protocol, with the ‘Normoxia (21% O2)72 h’ samples serving as the control group.

2.3.2. ProteoForge Application

Three comparisons were made: (1) hypoxia vs normoxia at 48 h, (2) hypoxia vs normoxia at 72 h, and (3) a time-course comparing hypoxia at 48 h, hypoxia at 72 h, and normoxia at 72 h.

For all analyses, a robust linear model (RLM) was used. A two-step multiple testing correction was applied, using a Bonferroni correction at the protein level followed by a global Benjamini–Hochberg FDR correction (fdr_bh). Peptides with an adjusted p-value below 0.001 were classified as significantly discordant. Peptide clusters were generated by using the hybrid_outlier_cut method, and these clusters were used to construct differential proteoforms (dPFs).

Peptides within dPFs were annotated using a custom utility (annotate.py) module, integrating feature annotations from UniProt (May 2024), protease-substrate information from MEROPS, and post-translational modification data from iPTMnet. All significant peptides were manually cross-referenced with the TopFind database , to identify potential neo-N- and C-termini.

2.3.3. Downstream Analysis

Four different global protein-level quantification (or ‘summarization’) strategies were compared: (1) Original, the protein-level output from DIANN; (2) mean­(pep), based on averaging top peptide intensities per protein; (3) mean­(all), based on averaging all identified peptide intensities per protein; and (4) dPFs, based on averaging peptide intensities within each ProteoForge-identified dPF group. Simple averaging-based summarization was chosen to isolate the direct effect of peptide grouping resolution on quantitative variance. Model-based tools such as MSqRob or MSstats apply their own normalization, variance shrinkage, and error modeling, which would confound the comparison by introducing variables beyond peptide selection. All four strategies were applied to the 7161 proteins quantified by ProteoForge. The DIA-NN protein-level output (‘Original’) was subset to this protein pool, ensuring differences in statistical power originate from peptide selection and grouping.

These strategies were compared using QuEStVar, , which tests for both equivalence and difference using log2 fold-change boundaries of [-0.25, 0.25] and [-0.5, 0.5], respectively, at a global FDR of 0.05. The upregulated, downregulated, and equivalent proteins from each strategy were subjected to pathway analysis using g:Profiler (version e113_eg59_p19_6be52918, database updated May 23, 2025; accessed November 28, 2025), with all identified proteins as the background and an FDR threshold of 0.001.

2.4. Technical Details

ProteoForge was built and tested by using Python (version 3.13.2). Numpy (v.1.24.3) and Pandas (v.1.5.2) facilitated data manipulation and analysis. The statsmodels (v.0.14.1) package was used to build and test the linear models, whereas scikit-learn (v.1.5.2) was used for performance metrics and kNN imputation. SciPy (v.1.10.1) served as the backbone in the clustering analyses. All visualizations were generated using Matplotlib (version 3.7.2) and seaborn (version 0.11.2). Benchmarking against COPF and PeCorA utilized custom R (v4.4.2) scripts interfacing with CCProfiler (v0.99.1) , for COPF and PeCorA (v0.0.90). The complete analysis codes, associated data sets, computational notebooks, and their rendered outputs are permanently archived on GitHub and can be accessed via https://github.com/LangeLab/ProteoForge_Analysis/.

3. Results

3.1. ProteoForge: A Four-Module Framework to Identify Differential Proteoforms

Identifying distinct proteoforms from bottom-up proteomics data is challenging; peptides are often shared between proteoforms and existing methods are frequently limited by their handling of missing data. We therefore developed ProteoForge, a framework that leverages peptide-level quantitative data to identify and assemble theoretical peptide groups, termed quantitatively differential proteoforms (dPFs). ProteoForge identifies ‘discordant peptides’ that show a consistent quantitative pattern that diverges from the parent protein’s trajectory across all conditions and replicates. This is in contrast to an outlier, which is a single aberrant measurement that does not replicate across conditions. It then defines dPFs as a set of at least two peptides from a protein with similar behavior patterns. Importantly, this framework is an imputation-aware statistical model, minimizing bias from missing values.

ProteoForge has four key modules. First, it normalizes peptide intensity data (Figure A.1). Next, a robust or weighted linear model identifies peptides with significantly discordant quantitative patterns (Figure A.2). The third module clusters peptides with similar quantitative profiles (Figure A.3). Finally, dPFs are built by combining statistical and clustering data: a cluster of a minimum of two peptides, with at least one being a significant discordant peptide, is assigned as a dPF. Significant singletons are flagged, and all other peptides not in a dPF are assigned to form a canonical proteoform (dPF_0) (Figure A.4).

3.2. ProteoForge Demonstrates Robust Identification and Peptide Grouping

We set out to assess ProteoForge’s performance in two key areas, compared to two alternative proteoform discovery methods, PeCorA and COPF: (i) quantitative assessment of significantly discordant peptide identification (referred to here as ‘peptide identification’) and (ii) grouping of peptides with similar variance and intensity profiles within parent proteins to infer potential proteoforms (referred to here as ‘peptide grouping’). For peptide grouping, only COPF was used for comparison as PeCorA does not provide this functionality. To model both complete data and data with imputed missing values, we used real data with known perturbations as well as simulated data sets.

First, we used the COPF benchmarking approach of Bludau et al. (2021) and a published HEK293 data set, to generate four distinct in-silico data sets of increasing perturbation complexity (referred to as “the SWATH-MS Interlab benchmark”). For 1000 randomly selected proteins, either one, two, a random number (2%–50%), or half (50%) of peptides were perturbed per protein.

For peptide identification, at one-peptide, two-peptide, and random conditions ProteoForge achieved the highest MCC scores among tested methods. However, when half of peptides were perturbed, ProteoForge drastically reduced in prediction performance to 0.06 MCC (Figure A). For the half peptide condition, COPF performed the best up to 0.14 MCC. The trends observed in Figure A were supported by Receiver Operating Characteristic (ROC) curve analysis, where ProteoForge maintained a high area under the curve (AUC 0.79–0.94) across all scenarios except 50% perturbation condition (Figure S1A). ProteoForge’s MCC performance curve formed a broad plateau across a wide range of p-values (Figure S1B); this demonstrates that, while peak performance is achieved at stringent thresholds (-log10­(p) > 4), a lower stringency threshold, such as -log10­(p) = 3, still captures a substantial fraction of the maximum predictive accuracy (Figure S1B). ProteoForge therefore provides a robust strategy for peptide identification that balances high precision with broad discovery, avoiding true proteoform loss due to overly conservative filtering.

2.

2

ProteoForge demonstrates more robust identification and grouping of peptides across diverse benchmark conditions than PeCorA and COPF. The performance of ProteoForge, PeCorA, and COPF were evaluated using peptide identification (detecting discordant peptides) and peptide grouping (clustering co varying peptides). Unless otherwise specified, the method is indicated by colour and shape (ProteoForge: triangle, red, PeCorA: square, yellow, COPF: circle, teal). Performance is reported as the Matthews Correlation Coefficient (MCC). PeCorA was excluded from peptide grouping analyses. (A) Performance on a real-world benchmark SWATH-MS Interlab dataset. Bar heights represent the mean MCC score. Overlaid points (triangles, squares, circles) indicate the MCC score achieved at a Benjamini-Hochberg adjusted p-value of 0.001. Error bars: 95% confidence interval calculated from 1000 bootstrap samples. (B) Simulation 1: MCC comparing complete data (grey) and imputed data with random number of perturbed peptides (coloured). (C) Simulation 2: MCC scores on imputed simulated data across a range of increasing missingness levels, from 0% (complete) to 80%. (D) Simulation 3: Method performance as a function of the mean log2-fold change (Perturbation Magnitude) for perturbed peptides in a simulation. (E) Simulation 4: Performance on simulated data with an increasing number of experimental conditions. Method is encoded by colour, the data overlap is indicated by linestyle (solid: True, dashed: False), and the direction of peptide change is represented by shape (square: same direction, circle: random direction).

Next, we evaluated the performance of ProteoForge and COPF in peptide grouping. The same data set was used with three perturbation scenarios (2%, 2–50%, or 50% peptides perturbed) (Figure A). Across all conditions, ProteoForge maintained a stable and high mean MCC of approximately 0.45, whereas COPF’s scores were consistently around 0.27. With increased perturbations up to half, ProteForge’s performance reduced to 0.35. This performance difference was also reflected in a consistently higher AUC for ProteoForge (0.85) compared to COPF (∼0.77) (Figure S2A). The similar broad plateau observed for peptide identification was also prevalent for peptide grouping, but this time reaching its optimal MCC score within a clear range of stringent adjusted p-values (−log10­(p)≈ 4–6). COPF’s peak performance occurred at a less stringent threshold (−log10­(p)≈ 1) (Figure S2B). ProteoForge’s performance reduction at 50% perturbation is mitigated by better cluster accuracy. This performance decline is consistent with the theoretical breakdown point of the Huber M-estimator (Note S4, Section S2.6).

To complement this real-data benchmark containing technical variation (Note S16), we created four fully simulated data scenarios, free from confounding batch effects (Figure B–E). This allowed us to evaluate all ProteoForge, PeCorA, and COPF across data quality, experimental design, and pattern changes. 500 proteins were used, with 250 proteins perturbed, to investigate the impact of (i) data imputation in the context of different peptide perturbation levels (Simulation 1), (ii) the level of data missingness (Simulation 2), (iii) the magnitude of peptide perturbations (Simulation 3), and (iv) multicondition complexity including peptide perturbation overlap and directionality (Simulation 4).

First, we assessed the effect of data imputation. For data with a random number (2–50%) of perturbed peptides per protein, ProteoForge’s peptide identification performance was largely unaffected by imputed data; the MCC score decreased by only 0.005 (from 0.574 to 0.569). In contrast, PeCorA’s MCC score decreased by 0.158 (Figure B). This stability was consistent across different peptide perturbation scenarios, where ProteoForge maintained high AUC values, while PeCorA’s scores decreased with data imputation (Figure S3A).

For peptide grouping, ProteoForge achieved higher MCC scores than COPF (Figure B). In scenarios where individual peptide identification was challenging (such as ’>50% Peptides’ perturbed condition), ProteoForge maintained a high AUC (∼0.91), demonstrating effective peptide clustering (Figure S4A).

Simulation 2 assessed the impact of increasing data missingness. For peptide identification, ProteoForge’s MCC score declined modestly (0.57 to 0.46) as the percentage of missing values increased (0% to 60%); however, increasing from 60% to 80% missingness resulted in a sharp MCC decline (Figure C and Note S1). In contrast, PeCorA’s accuracy was highly dependent on data completeness and degraded sharply as the level of missingness increased (Figure C). Heatmaps of performance at a p-value of 0.001 confirmed that ProteoForge maintains high accuracy across combined protein- and peptide-level missingness up to 60%. In contrast, PeCorA’s performance is primarily impacted by protein-level missingness (Figure S5A). For peptide grouping, ProteoForge’s MCC scores also decreased with higher missingness but were consistently higher than those of COPF up to 60% (Figures C and S5B).

Next, for Simulation 3, we evaluated the impact of peptide perturbation magnitude (incremental log2 fold-changes from 0.1 to 1.50) on method performance, using imputed data with random peptide perturbations (Figures D and S6). As expected, the performance of all methods improved with a stronger biological signal; however, ProteoForge achieved higher MCC scores than the other methods across the entire tested range (Figure D). For peptide identification, ProteoForge’s MCC increased to 0.91 at the highest perturbation magnitude, while PeCorA reached 0.52. A similar trend was observed for peptide grouping, where ProteoForge’s MCC reached 0.91 while COPF’s remained below 0.29 (Figures D and S6).

Finally, Simulation 4 assessed the impact of increasing experimental complexity by varying the number of conditions from two to six, again using imputed data with random peptide perturbations. ProteoForge’s performance remained high and stable for both peptide identification and grouping as the number of conditions increased (Figure E). This stability was maintained across various complex scenarios, including those with random directional changes and overlapping perturbations between conditions (Figure S7).

Null-simulation analyses confirmed that ProteoForge’s two-stage correction controls false-positive rates at both the peptide and proteoform-grouping levels, including under MNAR conditions, and that its sensitivity and clustering accuracy exceed competing methods across all tested scenarios (Notes S2–S4).

3.3. ProteoForge Identifies Alternative Hypoxia-Induced dPFs in Lung Cancer Cells

Having evaluated ProteoForge’s performance under a variety of simulated scenarios, we tested its ability to bring new biological insights to a published proteomics data set. To assess the influence of hypoxia on the glyoxalase pathway, Tomin et al. performed bottom-up proteomics on H358 lung cancer cells exposed to normal (21% O2) or hypoxic (1% O2) conditions for 48 or 72 h. Following data preparation and processing, the data set contained 95,369 peptides across 7161 proteins. ProteoForge ran for both 48 and 72 h conditions separately to determine hypoxia vs normoxia proteoforms at both time points. At 48 h, ProteoForge found ∼55% (3920) of proteins had no significantly discordant peptides, ∼30% (2185) had a single significantly discordant peptide, and the remaining ∼15% (1056) had multiple significantly discordant peptides (Figure A). At 72 h, ProteoForge found ∼29% (2109) of proteins had no significantly discordant peptides, ∼35% (2477) had a single significantly discordant peptide, and the remaining ∼36% (2575) had multiple significantly discordant peptides (Figure A).

Peptides were then clustered to build proteoforms. At 48 h, ∼7% (522) of proteins had at least one additional dPF, while ∼33% (2385) of proteins contained significantly discordant peptides in singleton clusters predicted to be PTMs (named ‘single peptide PTMs’). Furthermore, ∼5% (334) of proteins were found to have both a single peptide PTM and at least one additional dPF (Figure B). At 72 h, ∼17% (1255) of proteins had at least one additional dPF, while ∼39% (2818) of proteins contained at least one ‘single peptide PTMs’. ∼14% (979) of proteins were found to have both a single peptide PTM and at least one additional dPF (Figure B).

To identify the molecular basis for discordant peptide behavior, we first examined single peptide PTMs across two time points, where 48 h had 2802 peptides identified as potential PTMs across 2719 proteins and at 72 h 3945 peptides across 3797 proteins. It is important to note that a single protein could have multiple single peptide PTMs belonging to distinct clusters. For these peptides, we cross-referenced their genomic locations with annotations in a custom database. This revealed that the majority of peptides across both time points (75% for 48 h and 73% for 72 h) matched a known biological feature, while the remaining had no corresponding annotation (Figure C). The most frequent feature type was a modified residue (MOD_RES), with phosphorylation being the most prevalent (∼83% of the modified residues) followed by acetylation (∼10%) and methylation (∼5–6%) (Figure C).

The enzyme glyoxalase I (GLO1; Q04760) exemplifies the additional information provided by ProteoForge. While Tomin et al. observed increased levels of the GLO1 product (S-lactoylglutathione) under hypoxia, they reported no correlation between hypoxia and GLO1 protein abundance. ProteoForge analysis resolved this discrepancy. At 48 h, we identified a single discordant peptide (peptide 1) that remained abundant under hypoxia despite low levels in normoxia (Figure D). This peptide harbors known phosphorylation (T35) and mutagenesis (Q34) sites, suggesting a condition-specific PTM (Figure E). By 72 h, this response changed into a multipeptide dPF (peptides 1, 2, and 8) characterized by elevated abundance in hypoxia (Figure D). Peptides 1 and 2 share similar annotations, containing mutagenized sites (Q34 and S45) and phosphorylation sites (T35 and S45) (Figure E). Although peptide 8 carries distinct modifications (ubiquitination/acetylation at K148) compared to peptides 1 and 2, ProteoForge grouped them into a single dPF based on their synchronized upregulation under hypoxic stress (Figure E). This demonstrates that GLO1 undergoes complex, multisite regulation that standard protein-level averaging obscures.

Fucose-1-phosphate guanylyltransferase (FPGT; Q14772) is an enzyme involved in the salvage pathway of l-fucose metabolism and provides another example of ProteoForge’s utility. At 48 h, no discordant peptides were found, indicating a stable protein status. Conversely, the 72 h comparison yielded a two peptide dPF, signifying peptide heterogeneity (Figure F). The dPF was defined by two discordant peptides (Peptides 1 and 2 spanning amino acid positions 60–95) that showed a reduction in hypoxia relative to normoxia. Notable annotations were limited to ubiquitination at residue K73; however, these two peptides likely behave distinctly due to isoforms that span the rest of the protein. FPGT Isoform 5 is missing amino acids 150–403, while Isoform 6 generates a modified version of the protein from amino acids 128–607, a product of alternative splicing (Figure G). ProteoForge therefore identified a two-peptide dPF spanning amino acids 60–95 whose quantitative behavior aligns with known isoform boundaries, resolving condition-dependent proteoform heterogeneity that standard protein inference does not capture.

Collectively, these analyses demonstrate ProteoForge’s ability to resolve complex quantitative peptide signals into proteoform patterns and successfully distinguish variations linked to known annotations, such as PTM- or isoform-driven variance, from other unannotated, condition-dependent peptide behaviors.

3.4. Proteoform Deconvolution Improves Quantitative Reproducibility and Enhances Statistical Power for Differential Analysis

We next sought to examine the impact of proteoform-level resolution on a typical quantitative global protein analysis. Four different protein-level quantification strategies were applied to Tomin et al.’s data set: DIANN’s protein group quantification (‘Original’), the mean of the top three most intense peptides (‘mean­(top)’), the mean of all peptides (‘mean­(all)’), and the mean of peptides within each ProteoForge-defined dPF (‘dPFs’).

First, we assessed quantitative reproducibility, measured as the median coefficient of variation (CV = σ/μ) across the four biological replicates within each experimental group, calculated independently for each protein or dPF (Figure A). A lower CV indicates higher reproducibility. The ‘Original’ and ‘mean­(top)’ methods showed the highest variability, with median CVs ranging from 7.7% to 9.8%. In contrast, the ‘mean­(all)’ method consistently yielded the most reproducible results, with the lowest CVs across all conditions (ranging from 6.0% to 7.7%), likely as averaging all peptides has a variance-smoothing effect. While not surpassing ‘mean­(all)’, the ‘dPFs’ method showed strong reproducibility, with median CVs ranging from 7.0% to 8.1%, and a robust foundation for subsequent quantitative differential analyses.

4.

4

Proteoform deconvolution improves quantitative reproducibility and enhances statistical power for differential analysis. The performance of four protein-level quantification strategies was compared using the lung cancer hypoxia data set: original (DIANN output), mean­(top) (mean of top 3 peptides), mean­(all) (mean of all peptides), and dPFs (mean of peptides within each ProteoForge-defined proteoform). (A) Bar chart comparing the median coefficient of variation (CV) across biological replicates for each method and condition. Error bars: 95% confidence intervals. (B) Line graph showing the cumulative variance explained by principal components for each quantification strategy. (C) Stacked bar charts show the percentage of proteins classified into different regulatory states (equivalent, downregulated, upregulated, etc.) following differential analysis at 48 and 72 h. The dPFs method identifies a substantially larger fraction of regulated proteins compared with the other three methods, demonstrating greater statistical power.

Next, we examined the data variance structure with a principal component analysis (PCA) (Figure B). The ‘dPFs’ method captures the most variance with the fewest components, explaining more than 80% of the cumulative variance with just three principal components. This indicates that resolving proteins into proteoforms distinguishes experimental groups more effectively. This is further supported by the improved separation of groups in the PCA score plot (Figure S8). dPFs inherently increase feature dimensionality relative to standard protein-level summaries (e.g., splitting one protein group into dPF_0 and dPF_1 doubles its feature count). The improved PCA separation therefore reflects a combined effect of resolving conflicting biological signals within protein groups and the expanded feature space.

This improved separation translates into greater statistical power for differential analyses (Figure C). At 72 h, the ‘dPFs’ method classified a combined 23% of proteins as either upregulated (12%) or downregulated (11%), a substantial increase compared to the ‘mean­(all)’ (4%), ‘mean­(top)’ (2%), and ‘Original’ (2%) methods. By providing a clearer biological signal for a larger portion of the proteome, the ‘dPFs’ method also reduces proteins with an ambiguous ″Unexplained″ status, maximizing the biological insights derived from the data set.

Finally, to understand the functional consequences of proteoform-level changes, we performed a global pathway enrichment analysis on quantitative data from the ‘Original’ or ‘dPFs’ method. While both methods identified a similar total number of enriched terms per statistical status query, their distribution was strikingly different (Figure A). The ‘Original’ method predominantly classified pathways as ″Equivalent″, particularly at the 72 h time point, where almost no “Elevated” terms (those enriched in the upregulated set) were found. In contrast, ‘dPFs’ uncovered a substantial number of both “Elevated” and “Reduced” terms (those enriched in the downregulated set) at both 48 and 72 h. This trend was consistent across all GO categories. For instance, at 72 h, the ‘dPFs’ method identified 87 “Elevated” and 181 “Reduced” GO Biological Process (GO:BP) terms, whereas the ‘Original’ method found no GO:BP terms to be “Reduced” or “Elevated” (Figure B).

5.

5

dPF-based analysis reveals dynamic pathway regulation masked by conventional protein quantification. Pathway enrichment analysis was performed on regulated protein sets derived from the original (DIANN) and dPFs (ProteoForge) quantification methods at 48 and 72 h of hypoxia. (A) Alluvial plot showing the number and enrichment status (elevated, reduced, equivalent) of enriched Gene Ontology (GO) terms. (B) Heatmap detailing the number of enriched terms by GO category (BP: Biological Process, CC: Cellular Component, MF: Molecular Function), confirming the trend observed in (A). (C–H) Dot plots show the enrichment status and significance for selected GO term groups. Dot color indicates the regulatory status (red: elevated, blue: reduced, orange: equivalent), and dot size corresponds to the statistical significance (−log10 p-value). The dPFs method correctly identifies the elevation of canonical hypoxia pathways in (C) ″Response to Stimulus″ and reveals complex patterns in (D) ″Cellular Organization”, (E) ″Cell Cycle”, (F) ″Metabolism″, (G) ″Signaling″, and (H) ″RNA Processing″.

Further analysis of the enrichment status and significance of several GO terms found distinct differences between the ‘dPFs’ and ‘Original’ methods. For example, in the ″Response to Stimulus″ category, ‘dPFs’ identified a distinct, time-dependent elevation of canonical hypoxia pathways, specifically ″cellular response to hypoxia″ and ″response to decreased oxygen levels″ at 72 h (Figure C). The ‘Original’ method did not detect this response. A similarly striking contrast appeared in ″Cellular Organization″ and ″Cell Cycle″ (Figure D,E). While the ‘Original’ method showed broad equivalence at the protein level, differential proteoform-level analysis showed reduction of cell cycle-associated ‘dPFs’ and differential abundance of Cellular Organization-associated dPFs at 72 h (Figure D,E). A notable divergence was also observed in ″Metabolism″: while at the 48 h time points both methods suggested equivalence for metabolic processes, at 72 h, ‘dPFs’ lost any enrichment for these terms, suggesting the protein-level signal may not reflect proteoform behavior (Figure F). Another interesting category showed the “Signaling” terms to be generally elevated in “dPFs”, while not enriched in or equivalent in Original, indicating, while the enrichment significance is lower for these terms, the proteoform-level resolution is expanding the insights (Figure G). Finally, provided are some terms within the RNA processing category where both methods at both time points agree to be equivalent, which shows the main housekeeping processes did not change meaningfully under hypoxic conditions and did not result in distinct proteoform profiles (Figure H). Together, these results highlight that proteoform-level resolution enhances detection of pathway changes and uncovers distinct functional signatures that are averaged out in standard protein-level assessments.

4. Discussion

In this study, we present ProteoForge, a framework that detects and groups peptides with distinct quantitative patterns into differential proteoforms (dPFs). The primary goal of ProteoForge is to move beyond conventional protein-level quantification and provide a more granular view of peptide-level dynamics, enabling the identification of candidate proteoforms for further biological validation. Importantly, ProteoForge shows robust discordance detection with unsupervised grouping at high rates of missing data, something not previously demonstrated in a unified workflow. ProteoForge navigates missing data using an imputation-aware robust statistical model; while operating on imputed matrices, imputation bias is mitigated at the model-fitting stage through robust M estimation.

Comprehensive benchmarking on both real-world and simulated data sets demonstrated that ProteoForge provides stable performance across a wide range of analytical challenges. Its accuracy was largely unaffected by high rates of missing and imputed values, a condition which substantially degrades the performance of other methods. Notably, our analysis identified a critical stability threshold: ProteoForge maintains consistent identification and grouping performance up to 60% missingness. While all methods eventually succumb to data sparsity, ProteoForge’s conservative handling of outliers prevents the rapid performance collapse observed in standard models. Furthermore, ProteoForge achieved higher accuracy across all tested signal strengths and maintained its performance as experimental complexity increased through more conditions or random directional changes (Figure ).

This distinction between outliers and discordant peptides is central to ProteoForge’s design: transient measurement errors are separated from reproducible, condition-dependent biological signals at the peptide level, enabling detection of genuine proteoform differences rather than technical noise. Tsai et al. (2020) proposed explicit prefiltering of low-quality peptides prior to protein quantification. ProteoForge’s Robust Linear Model (RLM) provides a quality filter that operates during model fitting, down-weighting aberrant observations without discarding them, and is therefore complementary to prefiltering approaches.

The robust performance of ProteoForge can also be attributed to the RLM with Huber M-estimation. Benchmarking (Note S1) revealed that RLM offers a favorable balance between performance and usability. By effectively treating the artificial variance introduced by imputation as ‘outliers’, RLM automatically down-weights their influence without requiring manual specification of complex weight matrices. This enables the incorporation of measurement variance and accounts for the uncertainty of imputed values without the fragility of standard weighting schemes. This also explains the performance increase after imputing completely missing conditions; introducing low values creates a strong perturbation signal that the algorithm correctly identifies.

ProteoForge’s contribution is best understood by comparison with existing tools. PeCorA employs the same peptide-condition interaction model but fits it with OLS, which treats all observations equally and produces biased estimates when imputed values introduce non-Gaussian error. In addition, it does not group peptides into proteoforms; discordant peptides are removed to improve the accuracy of parent protein quantification. While this strategy is valid for focus solely on protein-level abundance, ProteoForge operates on the principle that these peptides are not errors but sources of novel biological information.

COPF groups peptides by pairwise correlation on raw intensities but does not incorporate a robust estimation step, making it sensitive to imputation artifacts. Module 3 shares conceptual similarities with the hierarchical clustering approach of Lukasse et al. (2014); however, ProteoForge clusters peptides based on median-adjusted intensity profiles derived from the RLM-normalized data rather than raw correlation matrices and integrates grouping with the statistical significance output of Module 2 to form biologically interpretable dPFs. ProteoForge uses the peptide grouping strategy proposed by Plubell et al. (2022), generating 1-to-N quantities per protein group based on covarying peptide behavior across conditions. Note S3 compares ProteoForge’s performance across multiple imputation strategies under MNAR-null calibration. ProteoForge maintained strict type I error control regardless of the imputation method, while PeCorA and COPF exhibited reduced statistical power, confirming that RLM’s residual-based weighting generalizes across imputation strategies.

A potential concern with any method operating on imputed data is that downshifted MNAR imputation may create artificial contrast between conditions. Our MNAR-null simulation demonstrated that ProteoForge’s RLM does not convert imputation artifacts into false biological signals: corrected FPR remained at 0.0000 across all tested thresholds (Note S3). This robustness generalized across seven imputation strategies spanning left-censored, local, global, and hybrid approaches, with null FPR ranging from 0.027 to 0.035 at α = 0.05 (Note S3).

The resolution of peptide signals into dPFs has broader implications for proteomics, particularly in addressing the protein inference problem, the challenge of accurately assigning peptides to their originating proteins. , The presence of a dPF can help resolve ambiguity for shared peptides; if a peptide common to multiple proteins exhibits a discordant pattern, then it provides evidence that it originates from a specific, regulated proteoform. This concept complements other advanced methods that utilize weighted networks or peptide-level evidence to enhance confidence in protein identification and quantification. , For instance, the recently introduced AlphaQuant framework leverages a hierarchical tree structure to propagate error models from the fragment ion level up to the gene, offering high sensitivity by using counting statistics rather than imputation for missing values. Similarly, MSstatsWeightedSummary addresses the challenge of quantifying proteins with known shared peptides by estimating contribution weights via convex optimization, though it currently ignores missing feature intensities. Protein inference methods such as Bayesian proteoform modeling and single-isoform inference operate upstream of ProteoForge, resolving which proteins or isoforms are present from the shared peptide evidence. In contrast, ProteoForge operates on established protein-to-peptide mappings and identifies condition-dependent quantitative discordances within those groups. When upstream protein inference distributes peptides from the same proteoform across multiple group labels, users should verify or consolidate the group assignments before analysis.

Using a published proteomics data set in lung cancer cells, we demonstrated ProteoForge’s ability to successfully distinguish hypoxia-induced variations linked to known annotations, such as the condition-specific PTM in GLO1 and isoform-driven differences in FPGT. This provides a clear rationale to prioritize specific proteoform candidates for functional validation. Furthermore, by retaining peptide-level variance, the analysis revealed a directional reorganization of cell cycle and signaling machinery (Figure E,H). The resulting divergence between ProteoForge (‘dPF’) and the standard analysis (‘Original’) for the ‘Metabolism’ and ‘Signaling’ pathway enrichment (Figure F,H) highlights the importance of retaining proteoform-level data for accurate biological conclusions.

A limitation of ProteoForge’s robust estimation framework is its behavior when the fraction of perturbed peptides approaches or exceeds 50% (Note S4 and Table S1). The Huber M-estimator’s breakdown point dictates that when half or more of the observations deviate from the consensus, the robust fit can no longer reliably distinguish the discordant group from the concordant baseline. Biologically, this arises during isoform switching events or high-occupancy post-translational modifications. One should therefore interpret results cautiously for proteins where the number of flagged discordant peptides exceeds 40% of the total. However, grouping-level metrics still remain significantly above COPF (Note S4 and Table S5), suggesting the proteoform structure is still recoverable. In this scenario, we recommend complementary analysis with methods that do not rely on a consensus assumption (e.g., correlation-based approaches such as COPF).

We acknowledge that interpreting peptide-level information still has limitations and resolution of similar abundance patterns remains a challenge for the field. , Technical artifacts from sample preparation, such as missed cleavages, can also confound the dPF detection. The development of centralized, comprehensive proteoform databases would greatly aid high-throughput annotation of identified dPFs and reduce the need for manual verification. ,, While our data (Figure S2) confirmed that Euclidean distance recovers simulated ground truth groupings across all tested missingness levels, alternative distance metrics such as correlation-based or Mahalanobis distance also warrant future evaluation.

Future work will focus on refining the peptide grouping logic to incorporate context-aware, adaptive weighting. RLM serves as a robust default; however, weighting schemes incorporating biological priors, such as limit of detection or local data density, could further enhance sensitivity. This would allow more flexible models such as WLS or GLM to surpass current baselines in complex cases where data-driven weights can be reliably calculated.

Supplementary Material

pr5c01235_si_001.pdf (14.9MB, pdf)

Acknowledgments

We thank all Lange lab members for their valuable feedback upon testing the framework in various projects and experimental settings.

Glossary

Abbreviations

AUC

Area Under the Curve

BH

Benjamini–Hochberg

COPF

Correlation-based functional Proteoform

CV

Coefficient of Variation

dPFs

Quantitatively Differential Proteoforms

FDR

False Discovery Rate

kNN

k-Nearest Neighbors

LFC

Log Fold Change

MCC

Matthews Correlation Coefficient

PeCorA

Peptide Correlation Analysis

PTM

Post-Translational Modifications

QuEStVar

Quantitative Exploration of Stability and Variability

ROC

Receiver Operating Characteristic

WLS

Weighted Least Squares

The raw mass spectrometry data analyzed in this study were obtained from two public repositories: the SWATH-MS interlaboratory data set (PXD004886) and the lung cancer hypoxia data set (PXD062503), both available through the ProteomeXchange Consortium via PRIDE. The exact input files derived from these data sets, all analysis scripts, computational notebooks, and their rendered outputs are archived on Zenodo (10.5281/zenodo.17795845) and GitHub (https://github.com/LangeLab/ProteoForge_Analysis/) under a CC BY-NC 4.0 license.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jproteome.5c01235.

  • Performance of discordant peptide identification and grouping on the SWATH-MS benchmark; the framework’s resilience to data imputation; performance stability across varying levels of missingness, perturbation magnitudes, and experimental complexities; improved separation of experimental groups in PCA using dPF-based quantification; the Module 4 inference logic for dPF assembly on a representative protein; the simulation framework for peptide perturbation patterns and missingness modeling; evaluation and justification for selecting linear estimators for imputation-aware peptide analysis; empirical FDR calibration and p-value framework comparison across ProteoForge, PeCorA, and COPF; imputation robustness through MNAR-null calibration and sensitivity to alternative imputation strategies; and formal benchmarking methodology including bootstrap inference, matched operating points, and clustering quality metrics (PDF)

This work was partially supported by grants from the Canadian Institutes of Health Research (CIHR, PJT-169190), Natural Sciences and Engineering Research Council of Canada (NSERC, RGPIN-2018-05645), the Michael Cuccione Foundation, and the BC Children’s Hospital Foundation (to P.F.L.). P.F.L was supported by the Canada Research Chairs program (CRC-RS 950-230867, P.F.L.), the Michael Smith Foundation for Health Research Scholar program (16442, P.F.L.), and University of British Columbia (E.K.E).

The authors declare no competing financial interest.

This paper was published ASAP on June 9, 2026. The Figure captions for Figures 1, 2, and 3 have been corrected due to a production error. The corrected paper was reposted on June 11, 2026.

Published as part of Journal of Proteome Research special issue “One Health Powered by Proteomics”.

References

  1. Smith L. M., Agar J. N., Chamot-Rooke J., Danis P. O., Ge Y., Loo J. A., Paša-Tolić L., Tsybin Y. O., Kelleher N. L., Proteomics C., for T.-D.. The Human Proteoform Project: Defining the Human Proteome. Sci. Adv. 2021;7(46):eabk0734. doi: 10.1126/sciadv.abk0734. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Smith L. M., Kelleher N. L.. Consortium for Top Down Proteomics. Proteoform: A Single Term Describing Protein Complexity. Nat. Methods. 2013;10(3):186–187. doi: 10.1038/nmeth.2369. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Aebersold R., Agar J. N., Amster I. J., Baker M. S., Bertozzi C. R., Boja E. S., Costello C. E., Cravatt B. F., Fenselau C., Garcia B. A., Ge Y., Gunawardena J., Hendrickson R. C., Hergenrother P. J., Huber C. G., Ivanov A. R., Jensen O. N., Jewett M. C., Kelleher N. L., Kiessling L. L., Krogan N. J., Larsen M. R., Loo J. A., Ogorzalek Loo R. R., Lundberg E., MacCoss M. J., Mallick P., Mootha V. K., Mrksich M., Muir T. W., Patrie S. M., Pesavento J. J., Pitteri S. J., Rodriguez H., Saghatelian A., Sandoval W., Schlüter H., Sechi S., Slavoff S. A., Smith L. M., Snyder M. P., Thomas P. M., Uhlén M., Van Eyk J. E., Vidal M., Walt D. R., White F. M., Williams E. R., Wohlschlager T., Wysocki V. H., Yates N. A., Young N. L., Zhang B.. How Many Human Proteoforms Are There? Nat. Chem. Biol. 2018;14(3):206–214. doi: 10.1038/nchembio.2576. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Carbonara K., Andonovski M., Coorssen J. R.. Proteomes Are of Proteoforms: Embracing the Complexity. Proteomes. 2021;9(3):38. doi: 10.3390/proteomes9030038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Faria M., Félix D., Domingues R., Bugalho M. J., Matos P., Silva A. L.. Antagonistic Effects of RAC1 and Tumor-Related RAC1b on NIS Expression in Thyroid. J. Mol. Endocrinol. 2019;63(4):309–320. doi: 10.1530/JME-19-0195. [DOI] [PubMed] [Google Scholar]
  6. Degan S. E., Gelman I. H.. Emerging Roles for AKT Isoform Preference in Cancer Progression Pathways. Mol. Cancer Res. 2021;19(8):1251–1257. doi: 10.1158/1541-7786.MCR-20-1066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. McCool E. N., Xu T., Chen W., Beller N. C., Nolan S. M., Hummon A. B., Liu X., Sun L.. Deep Top-down Proteomics Revealed Significant Proteoform-Level Differences between Metastatic and Nonmetastatic Colorectal Cancer Cells. Sci. Adv. 2022;8(51):eabq6348. doi: 10.1126/sciadv.abq6348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Adams L. M., DeHart C. J., Drown B. S., Anderson L. C., Bocik W., Boja E. S., Hiltke T. M., Hendrickson C. L., Rodriguez H., Caldwell M., Vafabakhsh R., Kelleher N. L.. Mapping the KRAS Proteoform Landscape in Colorectal Cancer Identifies Truncated KRAS4B That Decreases MAPK Signaling. J. Biol. Chem. 2023;299(1):102768. doi: 10.1016/j.jbc.2022.102768. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Chetta M., Basile A., Tarsitano M., Rivieccio M., Oro M., Capitanio N., Bukvic N., Priolo M., Rosati A.. The Target Therapy Hyperbole: “KRAS (p.G12C)”the Simplification of a Complex Biological Problem. Cancers. 2024;16(13):2389. doi: 10.3390/cancers16132389. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Jiang Y., Rex D. A. B., Schuster D., Neely B. A., Rosano G. L., Volkmar N., Momenzadeh A., Peters-Clarke T. M., Egbert S. B., Kreimer S., Doud E. H., Crook O. M., Yadav A. K., Vanuopadath M., Hegeman A. D., Mayta M. L., Duboff A. G., Riley N. M., Moritz R. L., Meyer J. G.. Comprehensive Overview of Bottom-up Proteomics Using Mass Spectrometry. ACS Meas. Sci. Au. 2024;4(4):338–417. doi: 10.1021/acsmeasuresciau.3c00068. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Ivanov M. V., Solovyeva E. M., Bubis J. A., Gorshkov M. V.. Improving the Protein Inference from Bottom-up Proteomic Data Using Identifications from MS1 Spectra. J. Am. Soc. Mass Spectrom. 2021;32(5):1258–1262. doi: 10.1021/jasms.1c00061. [DOI] [PubMed] [Google Scholar]
  12. Nesvizhskii A. I., Aebersold R.. Interpretation of Shotgun Proteomic Data: The Protein Inference Problem. Mol. Cell. Proteomics. 2005;4(10):1419–1440. doi: 10.1074/mcp.R500012-MCP200. [DOI] [PubMed] [Google Scholar]
  13. He Z., Huang T., Liu X., Zhu P., Teng B., Deng S.. Protein Inference: A Protein Quantification Perspective. Comput. Biol. Chem. 2016;63:21–29. doi: 10.1016/j.compbiolchem.2016.02.006. [DOI] [PubMed] [Google Scholar]
  14. Nesvizhskii A. I., Keller A., Kolker E., Aebersold R. A.. Statistical Model for Identifying Proteins by Tandem Mass Spectrometry. Anal. Chem. 2003;75(17):4646–4658. doi: 10.1021/ac0341261. [DOI] [PubMed] [Google Scholar]
  15. Plubell D. L., Käll L., Webb-Robertson B.-J., Bramer L. M., Ives A., Kelleher N. L., Smith L. M., Montine T. J., Wu C. C., MacCoss M. J.. Putting Humpty Dumpty Back Together Again: What Does Protein Quantification Mean in Bottom-up Proteomics? J. Proteome Res. 2022;21(4):891–898. doi: 10.1021/acs.jproteome.1c00894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Saltzman A. B., Leng M., Bhatt B., Singh P., Chan D. W., Dobrolecki L., Chandrasekaran H., Choi J. M., Jain A., Jung S. Y., Lewis M. T., Ellis M. J., Malovannaya A.. gpGrouper: A Peptide Grouping Algorithm for Gene-Centric Inference and Quantitation of Bottom-up Proteomics Data. Mol. Cell. Proteomics. 2018;17(11):2270–2283. doi: 10.1074/mcp.TIR118.000850. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Schork K., Turewicz M., Uszkoreit J., Rahnenführer J., Eisenacher M.. Characterization of Peptide-Protein Relationships in Protein Ambiguity Groups via Bipartite Graphs. PLoS One. 2022;17(10):e0276401. doi: 10.1371/journal.pone.0276401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Brown K. A., Melby J. A., Roberts D. S., Ge Y.. Top-down Proteomics: Challenges, Innovations, and Applications in Basic and Clinical Research. Expert Rev. Proteomics. 2020;17(10):719–733. doi: 10.1080/14789450.2020.1855982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Dou Y., Liu Y., Yi X., Olsen L. K., Zhu H., Gao Q., Zhou H., Zhang B.. SEPepQuant Enhances the Detection of Possible Isoform Regulations in Shotgun Proteomics. Nat. Commun. 2023;14(1):5809. doi: 10.1038/s41467-023-41558-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Ammar C., Thielert M., Weiss C. A. M., Rodriguez E. H., Strauss M. T., Rosenberger F. A., Zeng W.-F., Mann M.. Tree-Based Quantification Infers Proteoform Regulation in Bottom-up Proteomics Data. bioRxiv. 2025:641844. doi: 10.1101/2025.03.06.641844. [DOI] [Google Scholar]
  21. Dermit M., Peters-Clarke T. M., Shishkova E., Meyer J. G.. Peptide Correlation Analysis (Pecora) Reveals Differential Proteoform Regulation. J. Proteome Res. 2021;20(4):1972–1980. doi: 10.1021/acs.jproteome.0c00602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Bludau I., Frank M., Dörig C., Cai Y., Heusel M., Rosenberger G., Picotti P., Collins B. C., Röst H., Aebersold R.. Systematic Detection of Functional Proteoform Groups from Bottom-up Proteomic Datasets. Nat. Commun. 2021;12(1):3810. doi: 10.1038/s41467-021-24030-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Staniak M., Huang T., Figueroa-Navedo A. M., Kohler D., Choi M., Hinkle T., Kleinheinz T., Blake R., Rose C. M., Xu Y., Jean Beltran P. M., Xue L., Bogdan M., Vitek O.. Relative Quantification of Proteins and Post-Translational Modifications in Proteomic Experiments with Shared Peptides: A Weight-Based Approach. Bioinformatics. 2025;41(3):btaf046. doi: 10.1093/bioinformatics/btaf046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Lukasse P. N. J., America A. H. P.. Protein Inference Using Peptide Quantification Patterns. J. Proteome Res. 2014;13(7):3191–3199. doi: 10.1021/pr401072g. [DOI] [PubMed] [Google Scholar]
  25. Hu J. X., Zhao H., Zhou H. H.. False Discovery Rate Control with Groups. J. Am. Stat. Assoc. 2010;105(491):1215–1227. doi: 10.1198/jasa.2010.tm09329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Heller R., Bogomolov M., Benjamini Y.. Deciding Whether Follow-up Studies Have Replicated Findings in a Preliminary Large-Scale Omics Study. Proc. Natl. Acad. Sci. U. S. A. 2014;111(46):16262–16267. doi: 10.1073/pnas.1314814111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Kriegeskorte N., Simmons W. K., Bellgowan P. S. F., Baker C. I.. Circular Analysis in Systems Neuroscience: The Dangers of Double Dipping. Nat. Neurosci. 2009;12(5):535–540. doi: 10.1038/nn.2303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Collins B. C., Hunter C. L., Liu Y., Schilling B., Rosenberger G., Bader S. L., Chan D. W., Gibson B. W., Gingras A.-C., Held J. M., Hirayama-Kurogi M., Hou G., Krisp C., Larsen B., Lin L., Liu S., Molloy M. P., Moritz R. L., Ohtsuki S., Schlapbach R., Selevsek N., Thomas S. N., Tzeng S.-C., Zhang H., Aebersold R.. Multi-Laboratory Assessment of Reproducibility, Qualitative and Quantitative Performance of SWATH-Mass Spectrometry. Nat. Commun. 2017;8(1):291. doi: 10.1038/s41467-017-00249-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Zubarev R. A.. The Challenge of the Proteome Dynamic Range and Its Implications for In-Depth Proteomics. Proteomics. 2013;13(5):723–726. doi: 10.1002/pmic.201200451. [DOI] [PubMed] [Google Scholar]
  30. Pham T. V., Piersma S. R., Warmoes M., Jimenez C. R.. On the Beta-Binomial Model for Analysis of Spectral Count Data in Label-Free Tandem Mass Spectrometry-Based Proteomics. Bioinformatics. 2010;26(3):363–369. doi: 10.1093/bioinformatics/btp677. [DOI] [PubMed] [Google Scholar]
  31. Tyanova S., Cox J.. Perseus: A Bioinformatics Platform for Integrative Analysis of Proteomics Data in Cancer Research. Methods Mol. Biol. 2018;1711:133–148. doi: 10.1007/978-1-4939-7493-1_7. [DOI] [PubMed] [Google Scholar]
  32. Chicco D., Jurman G.. The Matthews Correlation Coefficient (MCC) Should Replace the ROC AUC as the Standard Metric for Assessing Binary Classification. BioData Min. 2023;16(1):4. doi: 10.1186/s13040-023-00322-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Chicco D., Jurman G.. The Advantages of the Matthews Correlation Coefficient (MCC) over F1 Score and Accuracy in Binary Classification Evaluation. BMC Genomics. 2020;21(1):6. doi: 10.1186/s12864-019-6413-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. DeLong E. R., DeLong D. M., Clarke-Pearson D. L.. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics. 1988;44(3):837–845. doi: 10.2307/2531595. [DOI] [PubMed] [Google Scholar]
  35. Tomin T., Honeder S. E., Liesinger L., Gremel D., Retzl B., Lindenmann J., Brcic L., Schittmayer M., Birner-Gruenberger R.. Increased Antioxidative Defense and Reduced Advanced Glycation End-Product Formation by Metabolic Adaptation in Non-Small-Cell-Lung-Cancer Patients. Nat. Commun. 2025;16(1):5157. doi: 10.1038/s41467-025-60326-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Benjamini Y., Hochberg Y.. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J. R. Stat. Soc. Ser. B. 1995;57(1):289–300. doi: 10.1111/j.2517-6161.1995.tb02031.x. [DOI] [Google Scholar]
  37. The UniProt Consortium. UniProt: The Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 2025;53(D1):D609–D617. doi: 10.1093/nar/gkae1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Rawlings N. D., Barrett A. J., Thomas P. D., Huang X., Bateman A., Finn R. D.. The MEROPS Database of Proteolytic Enzymes, Their Substrates and Inhibitors in 2017 and a Comparison with Peptidases in the PANTHER Database. Nucleic Acids Res. 2018;46(D1):D624–D632. doi: 10.1093/nar/gkx1134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Huang H., Arighi C. N., Ross K. E., Ren J., Li G., Chen S.-C., Wang Q., Cowart J., Vijay-Shanker K., Wu C. H.. iPTMnet: An Integrated Resource for Protein Post-Translational Modification Network Discovery. Nucleic Acids Res. 2018;46(D1):D542–D550. doi: 10.1093/nar/gkx1104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Lange P. F., Overall C. M.. TopFIND, a Knowledgebase Linking Protein Termini with Function. Nat. Methods. 2011;8(9):703–704. doi: 10.1038/nmeth.1669. [DOI] [PubMed] [Google Scholar]
  41. Fortelny N., Yang S., Pavlidis P., Lange P. F., Overall C. M.. Proteome TopFIND 3.0 with TopFINDer and PathFINDer: Database and Analysis Tools for the Association of Protein Termini to Pre- and Post-Translational Events. Nucleic Acids Res. 2015;43(D1):D290–D297. doi: 10.1093/nar/gku1012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Goeminne L. J. E., Gevaert K., Clement L.. Peptide-Level Robust Ridge Regression Improves Estimation, Sensitivity, and Specificity in Data-Dependent Quantitative Label-Free Shotgun Proteomics. Mol. Cell. Proteomics. 2016;15(2):657–668. doi: 10.1074/mcp.M115.055897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Choi M., Chang C.-Y., Clough T., Broudy D., Killeen T., MacLean B., Vitek O.. MSstats: An R Package for Statistical Analysis of Quantitative Mass Spectrometry-Based Proteomic Experiments. Bioinformatics. 2014;30(17):2524–2526. doi: 10.1093/bioinformatics/btu305. [DOI] [PubMed] [Google Scholar]
  44. Ergin, E. K. ; Myung, J. J. K. ; Lange, P. F. . Snapshot of the Quantitative Protein Stability Analysis in Cancer Cell Lines Using QuEStVar; Zenodo, 2024. 10.5281/zenodo.10694635. [DOI] [Google Scholar]
  45. Ergin E. K., Myung J. J. K., Lange P. F.. Statistical Testing for Protein Equivalence Identifies Core Functional Modules Conserved across 360 Cancer Cell Lines and Presents a General Approach to Investigating Biological Systems. J. Proteome Res. 2024;23(6):2169–2185. doi: 10.1021/acs.jproteome.4c00131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Kolberg L., Raudvere U., Kuzmin I., Adler P., Vilo J., Peterson H. G.. Profiler–Interoperable Web Service for Functional Enrichment Analysis and Gene Identifier Mapping (2023 Update) Nucleic Acids Res. 2023;51(W1):W207–W212. doi: 10.1093/nar/gkad347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Harris C. R., Millman K. J., van der Walt S. J., Gommers R., Virtanen P., Cournapeau D., Wieser E., Taylor J., Berg S., Smith N. J., Kern R., Picus M., Hoyer S., van Kerkwijk M. H., Brett M., Haldane A., Del Río J. F., Wiebe M., Peterson P., Gérard-Marchant P., Sheppard K., Reddy T., Weckesser W., Abbasi H., Gohlke C., Oliphant T. E.. Array Programming with NumPy. Nature. 2020;585(7825):357–362. doi: 10.1038/s41586-020-2649-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. McKinney, W. Data Structures for Statistical Computing in Python. In Proceedings of the 9th Python in Science Conference; SciPy: Austin, Texas, 2010; pp 56–61. DOI: 10.25080/Majora-92bf1922-00a. [DOI] [Google Scholar]
  49. Seabold, S. ; Perktold, J. . Statsmodels: Econometric and Statistical Modeling with Python. In Proceedings of the 9th Python in Science Conference; SciPy: Austin, Texas, 2010; pp 92–96. 10.25080/Majora-92bf1922-011. [DOI] [Google Scholar]
  50. Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., Vanderplas J., Passos A., Cournapeau D., Brucher M., Perrot M., Duchesnay E. ´.. Scikit-Learn: Machine Learning in Python. J. Mach. Learn. Res. 2011;12:2825. doi: 10.5555/1953048.2078195. [DOI] [Google Scholar]
  51. Virtanen P., Gommers R., Oliphant T. E., Haberland M., Reddy T., Cournapeau D., Burovski E., Peterson P., Weckesser W., Bright J., van der Walt S. J., Brett M., Wilson J., Millman K. J., Mayorov N., Nelson A. R. J., Jones E., Kern R., Larson E., Carey C. J., Polat I. ˙., Feng Y., Moore E. W., VanderPlas J., Laxalde D., Perktold J., Cimrman R., Henriksen I., Quintero E. A., Harris C. R., Archibald A. M., Ribeiro A. H., Pedregosa F., van Mulbregt P.. et al. SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nat. Methods. 2020;17(3):261–272. doi: 10.1038/s41592-019-0686-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Hunter J. D.. Matplotlib: A 2D Graphics Environment. Comput. Sci. Eng. 2007;9(3):90–95. doi: 10.1109/MCSE.2007.55. [DOI] [Google Scholar]
  53. Waskom M.. Seaborn: Statistical Data Visualization. J. Open Source Softw. 2021;6(60):3021. doi: 10.21105/joss.03021. [DOI] [Google Scholar]
  54. Heusel M., Bludau I., Rosenberger G., Hafen R., Frank M., Banaei-Esfahani A., van Drogen A., Collins B. C., Gstaiger M., Aebersold R.. Complex-Centric Proteome Profiling by SEC-SWATH-MS. Mol. Syst. Biol. 2019;15(1):e8438. doi: 10.15252/msb.20188438. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Tsai T.-H., Choi M., Banfai B., Liu Y., MacLean B. X., Dunkley T., Vitek O.. Selection of Features with Consistent Profiles Improves Relative Protein Quantification in Mass Spectrometry Experiments. Mol. Cell. Proteomics. 2020;19(6):944–959. doi: 10.1074/mcp.RA119.001792. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Audain E., Uszkoreit J., Sachsenberg T., Pfeuffer J., Liang X., Hermjakob H., Sanchez A., Eisenacher M., Reinert K., Tabb D. L., Kohlbacher O., Perez-Riverol Y.. In-Depth Analysis of Protein Inference Algorithms Using Multiple Search Engines and Well-Defined Metrics. J. Proteomics. 2017;150:170–182. doi: 10.1016/j.jprot.2016.08.002. [DOI] [PubMed] [Google Scholar]
  57. Webb-Robertson B.-J. M., Matzke M. M., Datta S., Payne S. H., Kang J., Bramer L. M., Nicora C. D., Shukla A. K., Metz T. O., Rodland K. D., Smith R. D., Tardiff M. F., McDermott J. E., Pounds J. G., Waters K. M.. Bayesian Proteoform Modeling Improves Protein Quantification of Global Proteomic Measurements. Mol. Cell. Proteomics. 2014;13(12):3639–3646. doi: 10.1074/mcp.M113.030932. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Bollon J., Shortreed M. R., Jeffery E., Jordan B. T., Miller R., Cavalli A., Smith L. M., Dewey C. N., Sheynkman G. M., Tiberi S.. IsoBayes: A Bayesian Approach for Single-Isoform Proteomics Inference. Bioinformatics. 2025;41(8):btaf450. doi: 10.1093/bioinformatics/btaf450. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Bamberger C., Martínez-Bartolomé S., Montgomery M., Pankow S., Hulleman J. D., Kelly J. W., Yates J. R.. Deducing the Presence of Proteins and Proteoforms in Quantitative Proteomics. Nat. Commun. 2018;9(1):2320. doi: 10.1038/s41467-018-04411-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Po A., Eyers C. E.. Top-down Proteomics and the Challenges of True Proteoform Characterization. J. Proteome Res. 2023;22(12):3663–3675. doi: 10.1021/acs.jproteome.3c00416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Zahn-Zabal M., Michel P.-A., Gateau A., Nikitin F., Schaeffer M., Audot E., Gaudet P., Duek P. D., Teixeira D., Rech de Laval V., Samarasinghe K., Bairoch A., Lane L.. The neXtProt Knowledgebase in 2020: Data, Tools and Usability Improvements. Nucleic Acids Res. 2019;48(D1):D328–D334. doi: 10.1093/nar/gkz995. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

pr5c01235_si_001.pdf (14.9MB, pdf)

Data Availability Statement

The raw mass spectrometry data analyzed in this study were obtained from two public repositories: the SWATH-MS interlaboratory data set (PXD004886) and the lung cancer hypoxia data set (PXD062503), both available through the ProteomeXchange Consortium via PRIDE. The exact input files derived from these data sets, all analysis scripts, computational notebooks, and their rendered outputs are archived on Zenodo (10.5281/zenodo.17795845) and GitHub (https://github.com/LangeLab/ProteoForge_Analysis/) under a CC BY-NC 4.0 license.


Articles from Journal of Proteome Research are provided here courtesy of American Chemical Society

RESOURCES