Abstract
Analysis of intact proteins by mass spectrometry enables direct quantitation of the specific proteoforms present in a sample and is an increasingly important tool for biopharmaceutical and academic research. Interpreting and quantifying intact protein species from mass spectra typically involves many challenges including mass deconvolution and peak processing as well as determining optimal spectral averaging parameters and matching masses to theoretical proteoforms. Each of these steps can present informatic hurdles, as parameters often need to be tailored specifically to the data sets. To reduce intact mass deconvolution data analysis burdens, we built upon the widely used “sliding window” mass deconvolution technique with several additional concepts. First, we found that how spectra are averaged and the overlap in spectral windows can be tuned to favor either sensitivity or speed. A multiple window averaging approach was found to be the most effective way to increase mass detection and yielded a >2-fold increase in the number of masses detected. We also developed a targeted feature-finding routine that boosted sensitivity by >2-fold, decreased coefficient of variation across replicates by 50%, and increased the quality of mass elution profiles through 3-fold more detected time points. Lastly, we furthered existing approaches for annotating detected masses with potential proteoforms through spectral fitting for possible proteoform family modifications and network viewing. These proteoform annotation approaches ultimately produced a more accurate way of finding related, but previously unknown proteoforms from intact mass-only data. Together, these quantitation workflow improvements advance the information obtainable from intact protein mass spectrometry analyses.
Graphical Abstract

INTRODUCTION
Proteins often contain a myriad of post-translational modifications (PTMs) and sequence variations from alternative splicing that can have dramatic impacts on their biochemical function. Precise analysis of the different proteoforms present in a biological system is a key part of understanding disease. Additionally, proteoform analysis is critical in the development of biotherapeutics and other protein products, where minute proteoform changes can have sizable impacts on product efficacy and safety. Mass spectrometry analysis is the dominant analytical approach for defining and quantifying proteoforms in a sample. However, accurate and reliable data analysis remains a significant challenge.
Intact protein mass spectrometry analysis relies on accurate deconvolution of raw mass spectra to yield proteoform masses, which ideally are also combined in the time domain to produce mass features and elution profiles. Over the past decade, several approaches have been developed for this step of intact protein feature detection. ProMex1 identifies isotope envelope signals across multiple spectra, performing charge assignment across these isotope elution profiles and grouping charge states into final mass features. FLASHDeconv2 utilizes a log transformation-based deconvolution approach on individual spectra, with subsequent binning of masses in the time domain to produce mass feature elution profiles. Multiple approaches have also relied on spectral averaging or summing to increase the signal-to-noise ratio of proteoform charge states before mass deconvolution. The “sliding window” approach present in Thermo Fisher Scientific software such as Protein Deconvolution, BioPharma Finder, and Proteome Discoverer involves a spectral window of a fixed size that “moves” along the chromatogram to average spectra before deconvolution using ReSPECT or Xtract.3,4 Protein Metrics’ Intact software5 similarly features sliding window spectral summing, as well as a time slicing approach that divides a chromatogram into nonoverlapping time regions for summing. The Bayesian deconvolution program UniDec6 has likewise recently added modules for mass feature detection across chromatograms, including the UniChrom7 module with averaging of sliding windows or time sections. Other commercial software tools such as Agilent’s MassHunter and Bruker’s BioPharma Compass also feature workflows that combine spectral deconvolution across multiple spectra to produce mass features.
After mass deconvolution, the mass features can then be annotated with potential proteoforms. One approach for generating these possible proteoforms is using UniProt XML files that contain isoform sequences and annotated PTMs, cleavages, and SNPs. Another approach is to utilize sequences from FASTA files or manual user input and then rely on users to manually add specific PTMs of interest. In contrast to fragmentation search engines that typically consider PTMs applied to different sites of the protein, intact mass-only searches can disregard PTM location, instead searching for only a single candidate proteoform mass represented by the PTM unlocalized across the sequence.
Once candidate proteoforms are generated, these possible masses can be used to annotate detected mass features. The most straightforward approach for intact mass searching involves simply matching the masses of detected features to the calculated masses of the proteoform candidates. The Smith group has developed an alternative proteoform family network-based approach, available in the software Proteoform Suite.8 The proteoform family approach involves comparing the mass differences between detected features and identifying those differences that correspond to a PTM or amino acid changes.9 These mass differences are then represented as edges connecting proteoform nodes, resulting in a network data structure in which individual clusters represent proteoform families originating from the same gene. In Proteoform Suite, the resulting network can be exported for visualization in network viewers such as Cytoscape, resulting in a visual format that can be used with shape and color mapping to easily identify quantitative trends and suggest biochemical PTM pathways.8 To the best of our knowledge, the proteoform family approach has not been applied to sliding window deconvolution feature detection.
Here, we describe the refinement of our intact mass analysis and quantitation workflow in ProSight Native, initially introduced for rapid deconvolution for high-throughput screening10 and now optimized for more conventional LC-MS applications. First, we compare four different spectral averaging modes, the choice of which drastically impacts the total runtime and sensitivity of mass deconvolution. We additionally investigate a targeted routine that uses a deconvolution-free approach for mass feature finding and find that this further increases the sensitivity, improves the feature quality, and reduces variation across replicate injections. These identified mass components are then annotated using a combination of intact search routines to identify potential proteoform identifications. By introducing a theoretical proteoform spectra and a proteoform family network approach, we find that this integrated strategy can sharply increase the overall precision and usefulness of intact mass searches.
EXPERIMENTAL SECTION
Cell Culture and Sample Preparation.
Primary IMR90 fibroblasts with an ER:ras fusion construct for inducing cellular senescence were cultured as previously mentioned,11 including culture in Dulbecco’s Modified Eagle Medium (Sigma-Aldrich) and 10% fetal bovine serum and treatment of cells with 4-hydroxytamoxifen (4-OHT) for 7 days to induce cellular senescence. Following cell culture, cell pellets were enriched for mitochondria as previously described.12 Briefly, cells were homogenized in 250 mM sucrose, 10 mM Tris-HCl, 0.1 mM EGTA, 1 mM DTT, and 10 mM sodium butyrate using a Teflon homogenizer. Nuclei were removed by centrifugation at 2,500g for 15 min, followed by centrifugation of the supernatants at 15,000g for 10 min. The resulting mitochondrial pellets were washed with 250 mM sucrose and 10 mM Tris-HCl, resuspended in 1% SDS, and heated for 10 min at 95 °C. Resulting samples were separated by molecular weight with a GELFREE 8100 device with an 8% gel column (Expedeon, Harston, Cambridgeshire, UK). The fraction corresponding to the 0–30 kDa range was collected for LC-MS analysis.
LC-MS Analysis of Mitochondrial Extracts.
Untreated and 4-OHT treated samples were each analyzed with four replicate LC-MS injections. Samples were analyzed with 90 min nanoRPLC gradient on a Dionex Ultimate 3000 Rapid Separations LCnano system (Thermo Fisher Scientific, San Jose, CA) as described previously.12 The intact quantitative data was acquired on an LTQ-Velos Orbitrap Elite (Thermo Fisher) at 120,000 FTMS resolution. Fragmentation data for intact mass search validation were acquired by targeting the most abundant charge states for HCD fragmentation using 25 NCE.
Sliding Window Deconvolution.
Sliding window deconvolution was performed by ProSight Native (Proteinaceous, Inc., Evanston, IL). The input to this workflow is mass spectrometry files, with direct support for Thermo .RAW files and support for most other vendor formats via an integrated mzML conversion wizard based on ProteoWizard. The first step of this workflow involves averaging of profile (preferred for low resolution spectra) or centroid spectra (preferred for faster runtimes, and used in this work), implemented as four different averaging approaches (Figure 2A). For the peaks averaging mode, chromatographic peaks were detected by identifying major peaks and valleys within the total ion chromatograms. For each peak, we produced a single spectrum from averaged scans within the time range. For other averaging modes, various window sizes (either nonoverlapping, overlapping or multiple windows) and offsets were used as described elsewhere in the text. After spectral averaging, mass deconvolution was determined using a heavily modified version of the THRASH algorithm,13 considering masses between 5,000 to 25 000 Da and charge states between 2 and 50, with a signal-to-noise threshold of 3. Masses were combined using a “top-down” approach that proceeds through raw mass results of decreasing intensity, for each mass binning other masses within 50 ppm until no masses remain. Peak detection was performed on these binned masses by splitting the masses into separate peaks when the gap between adjacent time points reached 1 min. Mass components from individual files were grouped across files using a similar “top-down” mass binning approach described above (with a 50 ppm tolerance), also applying a time tolerance of 1 min between files. Final mass values were computed as the intensity-weighted average across all deconvolution results for the mass. To ensure that only high-quality features were preserved, we required final mass components to be detected in at least three spectral windows, contain at least three consecutive charge states detected, and have a peak area in at least one data file of at least 0.01% relative to the most abundant mass component. This sliding window deconvolution routine was compared to BioPharma Finder 5.2 (Thermo Fisher Scientific) using the “Sliding Windows” mode with a spectrum width of 0.5 min and offset of 20%, Xtract deconvolution with a mass range of 5,000 to 25 000 Da, and a minimum of three detected spectral intervals per mass component. A full table of settings is provided in Table S1. BioPharma Finder comparisons were compared without the 0.01% relative peak area threshold and without targeted feature finding.
Figure 2.

Spectral averaging strategy has a strong impact on the runtime, sensitivity, and quality of deconvolution results. (A) Sliding window averaging modes present in ProSight Native. “Peaks” only averages regions that correspond to a detected peak in the total ion chromatogram. “Time segments” divides the chromatogram into equal-length segments for averaging. “Overlapping windows” likewise uses equal-length segments, but each is offset so that the windows overlap. “Multiple windows” uses different sets of overlapping windows. (B) Overlapping windows and multiple windows averaging modes increase the number of detected mass components (top) in both untreated and 4-hydroxytamoxifen-treated replicates. This increase in sensitivity comes at the cost of runtime, with the peaks and time Segments approach completing much more rapidly (bottom). (C) Adding multiple sets of averaging windows increases the signal-to-noise and quality of deconvolution results.
Targeted Feature Detection.
We used the averagine model14 to estimate the chemical formula for each mass component, followed by the Mercury7 algorithm15 to compute isotopic distributions, with each simulated isotope envelope limited to the most abundant seven peaks. Mass component start and end times were used along with 1 min padding to produce mass targets for subsequent quantitation. For each mass target, averaged spectral windows within the appropriate time range were compared to theoretical isotope distributions using the isotope fitting algorithm previously described.16–18 This isotope fitter compares theoretical versus experimental and outputs a score from zero to one that represents how well the isotope envelopes match in the m/z error and intensity distribution shape. To estimate a false discovery rate (FDR) at various isotope fit score thresholds, we used a decoy distribution of randomly generated masses. Masses were generated using a pseudorandom number generator within the 5,000 to 25 000 Da mass range used for THRASH deconvolution. Based on this decoy analysis, we only considered isotope fitting results with a score of 0.5 and an average m/z error less than 20 ppm. After isotope fitting, these results were used in combination with the deconvolution results to produce merged elution profiles. If both deconvolution and isotope fitting results were found for the same spectral window, the more intense of the two was used as the elution profile time point.
Intact Mass Searching.
We implemented three intact mass search routines. For the naïve search, candidate proteoforms were computed by starting with a FASTA file composed of reviewed human mitochondrial isoforms in UniProt (downloaded 2/9/2023). For each isoform, we considered N-terminally cleaved methionine sequences as well as variable modifications of acetyl, phosphorylation, methylation, formyl, and oxidation. Modification mass deltas were added to candidate sequences up to a maximum of three modifications per candidate proteoform. For the naïve search, the masses of these candidate proteoforms were directly compared to each detected mass component using a 30 ppm mass tolerance.
For the family finder search routine, proteoform root nodes were created from unmodified sequences and N-terminal acetylated sequences. Similar to the naïve search, these root nodes were compared to each detected mass component using a 30 ppm mass tolerance. If a root node matched a mass component, variable modifications were added to create new proteoform candidate nodes. These new candidate nodes were again compared to detected mass components, and this iterative process was repeated until the number of modifications per proteoform reached four or no new matches to mass components could be found.
RESULTS AND DISCUSSION
Ultrafast Sliding Window Deconvolution.
We implemented a new workflow in ProSight Native that combines rapid mass deconvolution with chromatographic feature finding routines and proteoform identification (Figure 1). This workflow requires only minimal input from users with mass spectrometry data files, protein sequences, and a parameter set. All intact mass detection and annotation analyses are then completed as a single contiguous workflow. The first step involves various averaging modes (described below) that control how a mass spectrometry data file is divided into spectral windows for averaging and deconvolution. This choice of averaging mode allows users to prioritize rapid runtimes versus maximal sensitivity for diverse applications. For mass deconvolution, we prioritized existing deconvolution algorithms that can readily scale to high-throughput screening applications. These deconvolution algorithms are tailored to provide millisecond analysis times. Additional speedups for the sliding window were achieved through parallelization of analysis to take advantage of multicore computing and extensive software profiling to identify processing inefficiencies. Low-resolution data are deconvolved using the kDecon algorithm for charge state deconvolution,10 and high-resolution spectra are deconvolved using a modified THRASH algorithm for isotopic envelope deconvolution.13 Once averaged spectra have been deconvolved, masses are binned based on a user-specified tolerance (in ppm or Daltons), before splitting into individual mass components based on chromatographic peak profiles. Depending on the length of analysis, averaging mode, and deconvolution parameters such as minimum signal-to-noise, data files were typically completed in between ~5 and 30 s. More details of the sliding window algorithm are provided in the Experimental Section.
Figure 1.

Workflow for rapid sliding window deconvolution. Sliding window averaging (first panel, with example averaging windows colored) can be used to balance the sensitivity with runtime. Averaged spectra are processed using rapid mass deconvolution algorithms (second panel) that can process whole data files in seconds. Individual mass results are combined to produce final mass features and elution profiles for quantitation (third panel). Final mass components are annotated using proteoform family networks, which provides a visual tool for organizing putative proteoform identifications and related forms for a protein.
Choice of Averaging Strategy Impacts Speed, Sensitivity, and Feature Quality.
Spectral averaging is an often crucial data processing step for increasing the signal-to-noise ratio of proteoform charge states and isotope peaks. Higher signal-to-noise ratios increase the likelihood of correct mass assignments by the deconvolution. Differences in spectral averaging can significantly affect the number of averaged spectra that must be deconvolved. Overall, the averaging strategy is a major determinant of the number and quality of detected mass features as well as the runtime of the workflow.
To assess the impact of spectral averaging on results and runtime, we compared four different methods of averaging scans using ProSight Native (Figure 2A). To test these averaging modes and investigate other improvements to sliding window deconvolution, we utilized a top-down proteomics data set focused on quantifying the differences in mitochondrial proteoforms upon chemical treatment to induce cellular senescence. We focused on replicate injections from this data set that could be used to filter out spurious mass components and optimize consistency across replicates. The first averaging mode we implemented was the “peaks” approach, which determines time ranges for averaging based either on user-selected regions or from major detected peaks from the base peak chromatogram (BPC) or total ion chromatogram (TIC). The automated peak detection of the Peaks workflow parallels the “auto peak detection” option in BioPharma Finder. With the automated peak finding focused on major peaks and valleys, we identified 58 ± 4 peaks per file (arithmetic mean ± standard deviation) for control samples (hereinafter “untreated”) and 77 ± 3 peaks per file for senescence-induced samples (hereinafter “treated”) (Figure S1). These peaks had corresponding full width at half-maximum (FWHM) values of 0.73 ± 0.01 min and 0.60 ± 0.02 min for untreated and treated samples, respectively (Figure S1). When these peaks were used to determine averaged spectra for sliding window deconvolution, we identified 466 masses from untreated samples and 541 from treated (Figure 2B, top). From these totals, 244 (45%) and 173 (37%) masses were found in all four untreated and treated replicates, respectively. This peaks approach required 40 ± 3 s to process all eight data files, averaged across three trials (Figure 2B, bottom).
A second “time segments” mode (Figure 2A) involved dividing the chromatograms into equal-length time regions for averaging. In contrast to the peaks approach, this mode did not require major peaks to be visible from the total ion chromatogram. Using these averaged spectra as the basis for sliding window deconvolution, we observed a close relationship between the retention time length used for each averaging segment and both the number of mass components detected across untreated and treated conditions and the total runtime (Figure S2). Using a time segment of 6 s that yielded ≥10 averaged spectra across typical peaks, we found a total of 653 and 632 mass components for untreated and treated samples, respectively, which was a 28% increase in mass components compared to the peaks approach. The fraction of masses found in all four replicate injections decreased slightly to 215 (33%) and 226 (36%) masses for untreated and treated, respectively. Analyzing these eight files using the “time segments” approach completed in 62 ± 8 s, a ~ 50% increase versus peaks (Figure 2C).
Next, we implemented an averaging approach that similarly used averaging windows of equal time length but shifted them across the chromatogram with a time offset. This approach, “overlapping windows” (Figure 2A), parallels the main “sliding window” approach in Thermo Fisher Scientific commercial implementations. In contrast to the time segments approach, we found that averaging time lengths near the average FWHM for the overlapping windows mode produced the highest number of detected components across untreated and treated replicates (Figure S3), with higher window time offsets greatly decreasing the number of components as the sampling rate across the peaks decreased. At a window time length of 0.5 min and offset of 0.1 min that balanced this observation with runtime (Figure S4), we detected a total of 1332 mass components combined across untreated and treated replicates (Figure S3). This corresponded to 1041 and 961 masses across untreated and treated injections, respectively (Figure 2B). The fraction of masses detected in all replicate injections was similar to time segments: 1041 (34%) and 961 (36%) for untreated and treated injections, respectively. The “overlapping windows” approach further increased the runtime to 122 ± 19 s for the eight data files (Figure 2C), which was twice the runtime of the time segments approach.
Having observed that window time lengths and offsets can drastically impact both the speed and the number of detected components, we reasoned that combining multiple sets of averaging window sizes might further increase the number of components detected. Using pairs of window time lengths and corresponding time offset lengths of 20%, we found that a “multiple window” mode with two time lengths had a large impact on the number of detected components across untreated and treated conditions (Figure S5) and total runtime (Figure S6). Compared to the optimal overlapping windows window time of 0.5 min with a 0.1 min offset, adding a second averaging window of 0.2 min with a 0.04 min offset resulted in 1328 and 1204 total mass components in untreated and treated replicates, respectively, an ~25% increase over a single overlapping window. Similar to single windows, 432 (33%) and 415 (34%) of these masses were in all untreated and treated replicates, respectively. This increase in sensitivity was accompanied by a 3-fold increase in runtime for a total of 303 ± 7 s across three trials (Figure 2C).
To further compare these averaging approaches not only on the basis of runtime and number of detected components, but also feature quality, we simulated the elution profiles and spectra for 18 masses ranging from 6–23 kDa and “spiked” these data in silico into the data file for an untreated sample injection. Analyzing these simulated mass spikes using the four averaging modes above, only 15/18 simulated components were detected using the peaks approach, while the other approaches detected all 18 (Figure S7). We also observed a stark difference in the quality of the elution profiles. As an example, one simulated mass component contained a total of 28 spectra, corresponding to a 62 s peak in the total ion chromatogram. While all averaging modes correctly identified the mass in each corresponding averaged spectrum, the differences in spectral sampling rates greatly affected the number of elution profile time points. Because the peaks approach averages once across a chromatographic peak, the elution profile only contained a single time point (Figure S8). The time segments approach with the chosen 6 s time width resulted in eight time points across the peak (Figure S8). While the overlapping windows approach further increased the number of time points to 12, using multiple windows produced an elution profile with 28 time points, the maximum number of points expected from the total ion chromatogram (Figure S8). We also observed an increase in the quality of individual spectral deconvolution results with the multiple windows approach detecting more charge states and resulting in higher signal-to-noise spectra than the overlapping windows approach (Figure 2C). We likewise noted an improvement in the elution profiles of nonsimulated detected masses from this untreated file, which was evident when comparing TICs to “total mass chromatograms” computed from detected mass components (Figure S9). While the total mass chromatogram from the peaks approach included only 50 time points from the 3048 total time points in the TIC (2% of spectra), the number of time points was increased to 1547 for the multiple windows approach (51% of spectra). This increase in spectral coverage across the file resulted in enhanced peak resolution for multiple windows.
Dramatic Speedup Versus Existing Software Solutions.
We next sought to compare these sliding window improvements to the popular implementation in BioPharma Finder. The entire senescence top-down proteomics data set was analyzed in a head-to-head comparison. The sliding window analysis was completed after 80.5 min in BioPharma Finder and 97 s in ProSight Native, using the overlapping windows mode and analogous averaging and mass deconvolution settings (Figure S10). To identify the main causes of this difference in runtime, we performed a side-by-side comparison of THRASH (used by ProSight Native) versus Xtract (used by BioPharma Finder), deconvolving each individual spectrum (n = 2937) for a single untreated replicate injection. In this comparison, THRASH had a runtime of 8 ± 5 ms per spectrum compared to 534 ± 257 ms for Xtract (Figure S10). This comparison suggests that differences in deconvolution algorithm are a major determinant in the difference in runtime between the two software tools.
For mass detection, the two software tools were in agreement, with 1051/1067 (99%) of the masses from BioPharma Finder also detected in ProSight Native. ProSight Native detected an additional 1632 masses. Of these masses, 302 were found in three replicates, and 132 were found in all four replicates for untreated or treated samples. These additional masses detected by ProSight Native represented mainly minor, lower abundance components, with 0.1 ± 0.2% abundance by peak area compared to 0.7 ± 3.3% abundance for masses shared with BioPharma Finder (Figure S10). Despite being lower in abundance, these additional masses contained 7 ± 7 spectral observation time points per mass and 3 ± 2 charge states (Figure S10). One example mass component (12382.698 ± 0.006 Da) detected by ProSight Native but not BioPharma Finder was present with consistent elution profiles across all eight total injections and contained seven consecutive charge states (Figure S11). Overall, these data suggest that ProSight Native can detect major, higher abundance protein species that would be found by BioPharma Finder but in a fraction of the runtime. Furthermore, ProSight Native offers substantially improved detection of lower abundance proteoforms.
Targeted Feature Finding Increases Consistency Across Replicates.
While mass deconvolution is an effective strategy for intact mass analysis, variations in the shape of isotope distributions and low signal-to-noise ratios can result in masses being missed. Such missing values can decrease the quality of elution profiles by reducing the number of time points and can increase the variance across replicates. To address this shortcoming of sliding window deconvolution, we introduced a new targeted feature finding routine that aims to fill in these gaps from missing values.
To implement this targeted feature finding strategy, we added a step that directly searches spectra for masses that have already been identified in other spectra, (Figure 3a). Before producing final elution profiles for each mass component, we computed theoretical spectra for each mass component, using chemical formulas estimated using the averagine model to generate theoretical isotope distributions. For each mass, we reanalyzed averaged spectral windows, using the appropriate start and end times from mass deconvolution with an extra 15 s “extension time” on either end to allow the peak boundaries to be extended. We then compared theoretical isotope envelopes to averaged spectra using an isotope fitting score previously described.16–18 This resulted in isotope fit scores between 0 and 1 for each isotope envelope with 1 representing a perfect match. To identify an appropriate isotope fit score threshold, we compared scores based on detected masses to scores from a decoy distribution of 1000 randomly generated masses (Figure S12). A threshold score of 0.5 corresponded to an approximately 8% false discovery rate. The summed intensity values from matching isotope distributions with a score above the 0.5 threshold were used to fill in elution profile gaps due to missed deconvolution.
Figure 3.

Targeted feature finding improves the consistency across replicates. (A) After sliding window deconvolution to detect initial mass components, theoretical isotope distributions are created for each mass. Averaged spectra that did not contain the mass are reanalyzed using an isotope fitting strategy, resulting in new time points to consider for quantitation. (B) Example of targeted feature finding improving detection of an inconsistently detected mass. While untargeted feature detection (top, orange) failed to detect the mass across all replicates, targeted feature finding (bottom, turquoise) detected the mass in all replicate injections. The number of time points detected for mass components was also increased with targeted feature detection, resulting in smoother elution profiles.
Using this targeted feature finding routine, we detected an additional 731 and 837 masses in untreated and treated injections, respectively, compared to the untargeted feature finding using the multiple windows mode (Figure S13). We also observed a substantial decrease in variation across replicates. For untargeted feature finding average coefficient of variation (CV%) values were 27 and 23% for untreated and treated, respectively. Using targeted feature finding, variation was improved, with ~50% decreases to CV% values yielding 18 and 15% for untreated and treated (Figure S13). Targeted feature finding also increased the number of time points across elution profiles, with an increase from 10 ± 11 to 33 ± 39 time points for untreated and 10 ± 11 to 33 ± 36 time points for treated injections (Figure S13). As an example, one mass component at 6,529.5 Da was initially sporadically detected across untreated and treated replicate injections, but after, targeted feature finding was successfully detected in all replicates (Figure 3b). Each file additionally showed a noticeable improvement in the quality of elution profiles (Figure 3b), with between 13 and 21 time points across peaks for untargeted versus 75–101 time points for targeted feature finding. These increases in the number of spectral observations resulted in fewer final mass features that were filtered out due to an insufficient number of observations or poor quality.
We were also interested in the effect of targeted feature finding on abundant mass components that were more readily detected by deconvolution. The most abundant species in our top-down proteomics data set was a mass component of 10835.864 ± 0.006 Da that represented between 37 and 46% of the total deconvolved intensity per file in untreated and treated injections. For this mass component, we observed differences in elution profile quality across different averaging modes and with targeted feature finding (Figure S14). For untreated and treated replicates, the number of time points across the peak increased from 1.2 ± 0.4 and 1.3 ± 0.5 time points for the peaks averaging approach to 33 ± 1 and 44 ± 1 time points using multiple windows (Figure S14). Using targeted feature finding and the same multiple windows averaging windows, the number of time points across the peak was further increased to 39 ± 3 and 65 ± 2 time points for untreated and treated replicates, respectively (Figure S14). For this abundant mass component, the increase in detected time points was most evident at the front and tail of the peaks, where the signal was less intense and thus more likely to be missed by deconvolution (Figure S15). Targeted feature finding also improved the detection of related masses within the 10835.9 Da proteoform family (Figure S16, further characterized below). These improvements in the elution profiles of this abundant mass component and related forms suggest that targeted feature finding may similarly improve data quality in studies of purified or enriched proteins and proteoform families of interest.
Annotating Mass Components Using Complementary Intact Mass Search Routines.
Once masses have been detected and quantified, the identification step is crucial in providing putative proteoform identifications. While intact mass-only searches cannot provide information about the localization of PTMs and sequence variants, they are a useful first-pass search that provides possible proteoforms that can be validated by additional top-down fragmentation or peptide mapping experiments. In intact mass proteomics applications, the number of false positives can easily be very high, increasing the number of erroneous assignments that must be navigated. We aimed to develop an intact mass searching routine that addressed this issue with false positives, while also allowing search with multiple variable modifications of interest.
We implemented three aspects of this intact mass search routine (Figure S17). First, the “naïve search” involves simple mass comparisons between experimentally observed mass components and theoretical proteoform masses. Candidate proteoforms were created by adding variable modifications to a set of protein sequences, up to a threshold number of maximum modifications per proteoform. For the mitochondrial senescence data set, we implemented this naïve search using all reviewed human mitochondrial isoform sequences (n = 1321) and added acetyl, oxidation, methyl, formyl, and phosphorylation, up to a maximum of three modifications per proteoform.
As a second strategy (Figure S17), we instead utilized theoretical isotope distributions as the basis for comparing the detected mass components to the theoretical proteoforms. This provided a complementary approach to the naïve search, allowing candidate proteoforms to be filtered in the m/z domain in addition to merely by deconvolved mass. Candidate proteoforms were populated identically with the naïve Search. For each candidate proteoform, its chemical formula was used to compute theoretical isotope distributions, which were compared to the most abundant averaged spectrum for each mass component.
The third family finder strategy (Figure S17) focused on the biochemical relationships among detected proteoform masses. Because multiply modified proteoforms often are detected alongside proteoforms with fewer modifications, we reasoned that this observation could be utilized to filter out false positive identifications. The Family Finder approach begins with a base proteoform “root” candidate that is initially considered. If the proteoform root matches experimental proteoform masses, then a new proteoform “leaf” candidate with a modification is considered. This new leaf candidate is compared to detected proteoform masses, and this iterative process of proteoform leaf consideration is continued until no more matches can be found, or a threshold maximum number of modifications per proteoform is reached. We implemented the family finder approach for our mitochondrial senescence data set using both unmodified isoform sequences and those containing a single N-terminal acetyl.
First, we compared these intact search strategies based on how many proteoform candidates matched to each mass component. Assuming that each mass component corresponded to a single correct proteoform (disregarding PTM localization), the number of candidates per mass is expected to be an indicator of false positives. Using a mass tolerance of 30 ppm (corresponding to 0.15 to 0.75 Da mass error for the 5 to 25 kDa mass range considered) resulted in 3.1 ± 2.6 matched proteoform candidates per mass component using the naïve search (Figure S18), with a maximum of 24 matched proteoforms. This value was decreased to 2.4 ± 2.5 and 2.3 ± 1.7 candidate proteoforms per mass component for the family finder and isotope fitting strategies, respectively. When the family finder was combined with the isotope fitting strategies, and only matching candidate proteoforms that passed both approaches were considered, the number of matched proteoforms per mass component was further reduced to 1.5 ± 1.1. We observed a similar improvement to algorithm runtime. While the naïve search completed in 88 s, the runtime was decreased to only 2 s by reducing the number of proteoform candidates considered during the search with the family finder approach. Adding the isotope fitting step after the family finder approach increased the total runtime by only 1 s.
To further investigate these intact mass search improvements, we focused on the abundant 10835.9 Da mass component described above. Using naïve search with a 30 ppm tolerance (corresponding to ±0.33 Da), this mass component matched ten possible candidate proteoforms (Table S2). These candidates represented four different protein accessions with between one and three acetyl, oxidation, formyl, and methyl modifications. When we used the isotope fitting approach to further filter these candidates by the fits between their experimental and theoretical isotope distributions, one candidate proteoform could be excluded due to an undetectable isotope fit (Table S2). Using the family finder approach to filter candidate proteoforms based on those that also had related forms detected, only two candidates remained: an acetylated form of P61604 (10 kDa heat shock protein) and an acetylated and methylated form of P56134 (ATP synthase subunit f). To decipher between these two proteoforms, we collected additional fragmentation data on the 10,835 kDa mass component. Using targeted top-down fragmentation tools in ProSight Native, this mass could be confidently assigned to acetylated 10 kDa heat shock protein, with 46% of residues explained by matching fragments (Figure S19).
In addition to providing a filter for reducing the number of possible candidate proteoforms for a mass, the family finder approach is also useful as a tool for the network visualization of proteoform families. Using proteoform family network visualization tools in ProSight Native, we generated a network of putative proteoform family identifications from our top-down proteomics data set (Figure 4A). Using the targeted feature finding approach described above, 1272 unique putative proteoform identification nodes were organized into 29 families containing multiple proteoform nodes and 1153 singleton families. One useful aspect of network visualization is that nodes and families can be colored based on various properties, allowing researchers to easily identify trends across the data set. Coloring nodes based on their abundance in treated versus untreated (Figure 4A) revealed that many proteoform families were up- or downregulated in their entirety, suggesting control of expression at the protein-level of these entire proteoform families. Others show mixed abundance patterns within the family, raising the possibility of proteoform-specific expression patterns for these proteins. Visualizing quantitative trends can also be useful for corroborating intact mass proteoform assignments, as families with proteoforms that are coexpressed across different treatment conditions can provide support that they are indeed biochemically related.
Figure 4.

Annotating mass components using proteoform family networks. (A) Proteoform family network comprised of 1272 putative proteoform identifications organized into 1182 families based on biochemical relationships between PTMs. Proteoform nodes are colored based on their fold-change abundance in treated versus untreated files, with node size corresponding to the maximum abundance observed for the proteoform. (B) The 10 kDa heat shock proteoform family, anchored by MS2 fragmentation evidence for the acetylated form, includes eight proteoforms with a variety of modifications. All proteoforms are more abundant in the treated files.
The 10 kDa heat shock protein proteoform family (Figure 4B), anchored by fragmentation evidence for the acetylated form, included seven additional putative forms. These represented a range of phosphorylation, acetylation, oxidation, methylation, and formylation modifications. All proteoforms in the family were more abundant in the treated injections, with five only detected in treated injections and the others 1.7 to 5.1-fold more abundant in treated injections.
CONCLUSIONS
As intact mass analyses become more routine and increase in scale and throughput, practitioners must balance data processing sensitivity with speed to make these experiments feasible. We present several aspects of sliding window deconvolution that can be tuned to sharply impact the number and quality of detected mass features as well as the analysis runtime. We present four complementary spectral averaging modes, each useful in different research contexts. In high-throughput screening applications that require the shortest processing times, our peaks approach provided ultrafast intact mass detection of the most abundant masses. The peaks approach was also useful for quick assessments such as “walk-up” instrument usage and for optimizing deconvolution parameters before beginning more comprehensive and long-running analyses. Researchers wishing to characterize their samples more exhaustively or produce the highest quality elution profiles, irrespective of runtime, are best suited to the multiple windows approach with multiple averaging widths. The time segments and overlapping windows approaches balance runtime with sensitivity and feature quality and are ideal averaging modes for routine daily analyses. In addition to these averaging optimizations, we further show how targeted isotope fitting can complement mass deconvolution by increasing the sensitivity of detection, improving elution profiles, and reducing variation across replicates. Following feature finding, annotation of detected masses via proteoform family networks can significantly limit the search space and reduce false positives for intact mass-only searches while also providing a useful data visualization tool for generating biological insights. Future advances in rapid and reliable intact mass detection and annotation will be an invaluable asset to the field, helping these experiments become more commonplace and powerful across biopharma and academia.
Supplementary Material
ACKNOWLEDGMENTS
This work was supported by the National Institute of General Medical Sciences of the National Institutes of Health by an SBIR grant R44GM130262 award to Proteinaceous.
Footnotes
Supporting Information
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.analchem.3c02345.
Settings for BioPharma Finder; illustration of chromatogram peaks for sliding window definition; effect of averaging parameters on sensitivity and runtime for time segments, overlapping windows, and multiple windows modes; effect of averaging mode on simulated mass detection and elution profiles; comparisons between ProSight Native and BioPharma Finder; estimation of targeted feature finding false discovery rate; comparisons between untargeted and targeted feature detection; elution profile comparisons for the 10 kDa heat shock protein; targeted feature finding for the 10 kDa heat shock protein and its proteoform family; illustration of three intact mass search strategies; comparison of excess proteoform matches across intact mass search strategies; proteoform candidates for the 10835 Da mass component; and fragmentation evidence for identification of the 10 kDa heat shock protein (PDF)
Complete contact information is available at: https://pubs.acs.org/10.1021/acs.analchem.3c02345
The authors declare the following competing financial interest(s): All authors declare a competing financial interest due to their involvement in the commercialization of proteomics software.
Contributor Information
Matthew T. Robey, Proteinaceous, Inc., Evanston, Illinois 60201, United States; Northwestern University, Evanston, Illinois 60208, United States
Daisha Utley, Proteinaceous, Inc., Evanston, Illinois 60201, United States.
Joseph B. Greer, Proteinaceous, Inc., Evanston, Illinois 60201, United States; Northwestern University, Evanston, Illinois 60208, United States
Ryan T. Fellers, Proteinaceous, Inc., Evanston, Illinois 60201, United States; Northwestern University, Evanston, Illinois 60208, United States
Neil L. Kelleher, Proteinaceous, Inc., Evanston, Illinois 60201, United States; Northwestern University, Evanston, Illinois 60208, United States
Kenneth R. Durbin, Proteinaceous, Inc., Evanston, Illinois 60201, United States; Northwestern University, Evanston, Illinois 60208, United States
REFERENCES
- (1).Park J; Piehowski PD; Wilkins C; Zhou M; Mendoza J; Fujimoto GM; Gibbons BC; Shaw JB; Shen Y; Shukla AK; Moore RJ; Liu T; Petyuk VA; Tolic N; Pasa-Tolic L; Smith RD; Payne SH; Kim S Nat. Methods 2017, 14 (9), 909–914. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (2).Jeong K; Kim J; Gaikwad M; Hidayah SN; Heikaus L; Schluter H; Kohlbacher O Cell Syst. 2020, 10 (2), 213–218. [DOI] [PubMed] [Google Scholar]
- (3).Millan-Martin S; Carillo S; Fussl F; Sutton J; Gazis P; Cook K; Scheffler K; Bones J Eur. J. Pharm. Biopharm. 2021, 158, 83–95. [DOI] [PubMed] [Google Scholar]
- (4).Bailey AO; Han G; Phung W; Gazis P; Sutton J; Josephs JL; Sandoval W MAbs 2018, 10 (8), 1214–1225. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (5).Bern M; Caval T; Kil YJ; Tang W; Becker C; Carlson E; Kletter D; Sen KI; Galy N; Hagemans D; Franc V; Heck AJ R. J. Proteome Res 2018, 17 (3), 1216–1226. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (6).Marty MT; Baldwin AJ; Marklund EG; Hochberg GKA; Benesch JLP; Robinson CV Anal. Chem 2015, 87 (8), 4370–4376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (7).Kostelic MM; Marty MT; Sun L; Liu X Methods Mol. Biol 2022, 2500, 159–180. [DOI] [PubMed] [Google Scholar]
- (8).Cesnik AJ; Shortreed MR; Schaffer LV; Knoener RA; Frey BL; Scalf M; Solntsev SK; Dai Y; Gasch AP; Smith LM J. Proteome Res 2018, 17 (1), 568–578. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (9).Shortreed MR; Frey BL; Scalf M; Knoener RA; Cesnik AJ; Smith LM J. Proteome Res 2016, 15 (4), 1213–1221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (10).Pu F; Knizner KT; Robey MT; Radosevich AJ; Ugrin SA; Elsen NL; Durbin KR; Williams JD J. Am. Soc. Mass Spectrom 2022, 33 (12), 2338–2341. [DOI] [PubMed] [Google Scholar]
- (11).Tran JC; Zamdborg L; Ahlf DR; Lee JE; Catherman AD; Durbin KR; Tipton JD; Vellaichamy A; Kellie JF; Li M; Wu C; Sweet SMM; Early BP; Siuti N; LeDuc RD; Compton PD; Thomas PM; Kelleher NL Nature 2011, 480 (7376), 254–258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (12).Catherman AD; Durbin KR; Ahlf DR; Early BP; Fellers RT; Tran JC; Thomas PM; Kelleher NL Mol. Cell. Proteomics 2013, 12 (12), 3465–3473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (13).Horn DM; Zubarev RA; McLafferty FW J. Am. Soc. Mass Spectrom 2000, 11 (4), 320–332. [DOI] [PubMed] [Google Scholar]
- (14).Senko MW; Beu SC; McLaffertycor FW J. Am. Soc. Mass Spectrom 1995, 6 (4), 229–233. [DOI] [PubMed] [Google Scholar]
- (15).Rockwood AL; Haimi PJ Am. Soc. Mass Spectrom 2006, 17 (3), 415–419. [DOI] [PubMed] [Google Scholar]
- (16).Fornelli L; Srzentic K; Huguet R; Mullen C; Sharma S; Zabrouskov V; Fellers RT; Durbin KR; Compton PD; Kelleher NL Anal. Chem 2018, 90 (14), 8421–8429. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (17).Seckler HDS; Fornelli L; Mutharasan RK; Thaxton CS; Fellers R; Daviglus M; Sniderman A; Rader D; Kelleher NL; Lloyd-Jones DM; Compton PD; Wilkins JTJ Proteome Res. 2018, 17 (6), 2156–2164. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (18).Melani RD; Gerbasi VR; Anderson LC; Sikora JW; Toby TK; Hutton JE; Butcher DS; Negrão F; Seckler HS; Srzentić K; Fornelli L; Camarillo JM; LeDuc RD; Cesnik AJ; Lundberg E; Greer JB; Fellers RT; Robey MT; DeHart CJ; Forte E; Hendrickson CL; Abbatiello SE; Thomas PM; Kokaji AI; Levitsky J; Kelleher NL Science 2022, 375 (6579), 411–418. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
