Abstract
The crucial roles of proteoforms in biological processes and disease mechanisms have been increasingly recognized. However, the rate at which new proteoforms are being discovered using top-down proteomics has far outpaced the rate at which the functional significance of different proteoforms can be determined. Because of the close connection between protein folding and protein function, protein folding stability measurements on proteoforms have the potential to identify functionally significant proteoforms of a given protein. While a number of mass spectrometry-based proteomics methods for making protein folding stability measurements on the proteomic scale have been reported over the past decade, none have been interfaced with top-down proteomics. Described here is a top-down (TD) stability of proteins from the rates of oxidation (SPROX) approach for making proteoform specific folding stability measurements. This approach is validated using a mixture of three model proteins with well-characterized protein folding behavior by conventional SPROX as well as other more conventional biophysical techniques. The method is also used to evaluate the relative folding stabilities of the <30 kDa protein fraction isolated from an MCF-7 cell lysate. The relative folding stabilities of 150 proteoforms from 83 proteins were successfully characterized in the cell lysate analysis using the TD-SPROX approach.
Graphical Abstract

INTRODUCTION
The rapid development of top-down proteomics (TDP) in recent years has led to the discovery of increasingly large numbers of proteoforms (i.e., the different forms of proteins resulting from splice variants, sequence variations, different posttranslational modifications (PTMs) and combinations thereof) across diverse proteomes.1,2 While the number of protein-coding genes in a proteome is relatively easy to estimate from genomic sequencing data, the number of proteoforms that exist in a proteome is more difficult to determine.2,3 Recent estimates of the number of proteoforms in the human proteome suggest that the number is at least an order of magnitude more than the number of protein-encoding genes.2 Concomitant with proteoform discovery efforts is the realization that proteins can have proteoform-specific functions and dysfunctions with important biological consequences.2,4–6 However, deciphering the functional significance of the myriad of specific proteoforms has been a formidable challenge.2,7 This has been in large part due to challenges associated with obtaining purified or synthetic proteoforms, which are often needed for functional studies.8,9
Because of the close connection between protein folding stability and protein function, protein folding stability measurements would be useful for studying proteoform function/dysfunction.10,11 Recent years have seen the creation of a new toolbox of mass spectrometry-based proteomics methods for making folding stability measurements of protein folding on complex proteins mixtures. These protein stability profiling strategies include the Cellular Thermal Shift Assay (CETSA),12 Thermal Proteome Profiling (TPP),13 several limited proteolysis strategies (e.g., Drug Affinity Responsive Target Stability (DARTS),14 Limited Proteolysis (LiP),15 Pulse Proteolysis (PP)16), several covalent labeling strategies (e.g., a fast photochemical oxidation strategy termed, FPOP,17 the Stability of Proteins from Rates of Oxidation (SPROX),18,19 plus a lysine modification strategy termed, Covalent Protein Painting (CPP)20). However, all the mass spectrometry-based methods in the existing toolbox for making protein folding stability measurements on the proteomic scale currently rely on bottom-up proteomics (BUP). Therefore, the above methods all have the same protein inference problem that plagues all BUP strategies when it comes to the identification of proteoforms.21,22
In contrast to BUP, TDP approaches are well-suited for the identification of proteoforms. The direct analysis of intact proteins in TDP methods eliminates the protein inference problem created by protease digestion in BUP analyses. Recent progress in TDP has afforded the identification of proteoforms with higher and higher molecular weights and in samples of increasing complexity.23–25 Unfortunately, the quantitation strategies employed in TDP have not been as robust and generally applicable as in BUP.26 For example, despite several recent applications of quantitative TDP analyses using isobaric mass tags (albeit not with protein folding stability workflows) showing some promising results,27,28 the use of isobaric mass tags in TDP is far from routine and widespread. Problems with protein precipitation during labeling and significant sample heterogeneity at the intact protein level, particularly in larger proteoforms,29 make isobaric mass tagging workflows in TDP challenging. This has made it nontrivial to incorporate TDP workflows into protein folding stability methods such as SPROX and TPP, which so far have relied almost exclusively on quantitative BUP strategies using isobaric mass tags.13,18 In these BUP protocols, proteins undergo covalent labeling or precipitation under gradient denaturant conditions before being enzymatically digested into peptides. These peptides are then reacted with isobaric mass tags to quantify the extent of labeling (SPROX) or precipitation (TPP) as a function of denaturant to generate protein unfolding curves.
Reported here is a top-down (TD)-SPROX workflow exploiting a so-called “one-pot” SPROX strategy, which we have previously described in BUP SPROX studies using isobaric mass tags for quantitation.30,31 The TD-SPROX protocol described in this work utilizes a label-free quantitation strategy where the relative ion signals of the wild-type and oxidized versions of detected proteoforms are used to evaluate their chemical-denaturant induced equilibrium unfolding transitions. The method was validated with a mixture of three model proteins with well-established chemical-denaturant induced equilibrium unfolding properties and SPROX behaviors. The assay was also applied to the proteins (and proteoforms) in an MCF-7 cell lysate, uncovering relative folding stabilities of 150 proteoforms from 83 proteins. Furthermore, TD-SPROX addresses the inference challenges in traditional BUP SPROX workflows where PTMs and the reporter methionine are often not on the same peptide. Our data revealed protein folding stability changes due to PTMs including phosphorylation and acetylation that can be directly linked to proteoform functions.
MATERIALS AND METHODS
TD-SPROX Protocol.
The TD-SPROX approach developed in this work is based on the one-pot SPROX protocol we previously interfaced with a quantitative BUP readout utilizing isobaric mass tags.30,31 In this protocol, aliquots of the protein sample were diluted into a series with 20 mM sodium phosphate buffers (pH 7.4) containing 150 mM NaCl and increasing concentrations of guanidinium chloride (GdmCl). The protein in each denaturant-containing buffer was equilibrated and reacted with hydrogen peroxide under the same reaction conditions. The reaction conditions (i.e., hydrogen peroxide concentration and reaction time) were tuned such that the oxidation reaction of an unprotected (i.e., solvent exposed) methionine residue will proceed half-lives for the pseudo-first order reaction. The hydrogen peroxide concentration and reaction time used in this work were 1 M and 3 min. Under these reaction conditions, the primary sites of protein modification (i.e., oxidation) are methionine residues that are exposed to solvent, either in the native structure or as a result of protein unfolding. After the protein oxidation reaction in each denaturant-containing buffer is quenched with tris(2-carboxyethyl)phosphine (TCEP), the resulting protein samples are combined into a single sample (i.e., “one-pot”). More detailed information about the protein and denaturant concentrations used in the “one-pot” and control samples generated in this work is included in the Detailed Materials and Methods section in the Supporting Information.
TDP Data Acquisition and Analysis.
The “one-pot” and control samples generated in the TD-SPROX analyses of the 3-component model protein mixture and the proteins in an MCF-7 cell lysate were desalted using a modified methanol/chloroform/water precipitation protocol described previously.32 The MCF-7 cell lysate samples were also fractioned/extracted for proteins <30 kDa by passively eluting proteins from polyacrylamide gels as intact species (PEPPI) as previously described.32,33 The resulting samples were subjected to triplicate LC-MS/MS analyses using a Vanquish Neo UHPLC chromatographic system (Thermo Fisher Scientific) coupled inline to an EasySpray source and an Orbitrap Eclipse mass spectrometer (Thermo Fisher Scientific) operating in intact protein mode with 2 mTorr of pressure in the ion routing multiple (IRM). Detailed chromatographic conditions and data acquisition parameters are included in the Detailed Materials and Methods section in the Supporting Information. All mass spectrometry.raw data files were uploaded to MassIVE (repository number MSV000095546).
The LC-MS/MS data generated on a control MCF-7 cell lysate sample that was not treated with hydrogen peroxide or chemical denaturant was searched on TDportal v4.1.0 (https://portal.nrtdp.northwestern.edu/) against a proteoform database (Taxon 9606 – June 2020) that contained 2.4 million proteoform entries (Figure S1). The results were reported with a 1% context-dependent false discovery rate (FDR) assigned at the protein, isoform, and proteoform levels.
Proteoform Finder (available in ProSight Native v1.0.24208, Proteinaceous, Inc.)32 was used to generate the ion signal intensities of the wild-type and oxidized forms of the three model proteins in the three “one-pot” and control samples generated in the model protein mixture analysis (Figure S2). In each case, the ion signals for the three most intense charge states were extracted and used for quantitation. Proteoforms identified in the MCF-7 cell lysate sample were analyzed in the same way and focused on the top 50% most abundant proteoforms (as determined by their total ion signal intensity in the control sample that saw neither GdmCl nor hydrogen peroxide).
Transition Midpoint Determination.
The Proteoform Finder generated ion signal intensities of each wild-type proteoform and its methionine oxidation products were used to generate a value for each one-pot and control sample using eq 1.
| (1) |
In eq 1, is the mass difference between the wild-type proteoforms and their oxidation products (e.g., 0 Da, 16 Da, 32 Da, 48 Da, etc.), is the intensity detected at value corresponding to proteoform species associated with . The values were used to calculate a SPROX transition midpoint value for each proteoform using eq 2.
| (2) |
In eq 2, A is the GdmCl concentration at the pre-transition point; B is the GdmCl concentration in the post-transition point. In the 3-component model protein mixture analysis, A was 0 M and B was 5 M; In the MCF-7 cell lysate analysis, A was 0.5 M and B was 2.5 M. The value is the value of the one-pot sample; is the value at the pre-transition point A; is the value at the post-transition point B (Figure 1).
Figure 1.

Workflow of TD-SPROX. Protein samples were diluted into gradient denaturant buffers and reacted with under consistent conditions. Following cleanup, samples were subjected to top-down proteomics analysis. Transition midpoint values were determined based on values of the “one-pot” and control samples. Abbreviations: Den = Denaturant, PFR = Proteoform, AUC = Area Under the Curve, OP = One-Pot.
Protein Folding Free Energy Calculation.
In some cases, the values determined above were used in eq 3 to calculate protein folding free energies.
| (3) |
In eq 3, which has been previously described,34 R is the gas constant, T is the temperature in Kelvin, is the average pseudo-first-order rate constant for the oxidation of the unprotected methionine residue in the protein, t is the time of oxidation in seconds, is is the folding free energy in the absence of denaturant. The TD-SPROX protocol employed in this work cannot generate -values. Thus, the -values used in the value calculations in this work were assigned based on values previously reported on the same protein in similar chemical denaturant-induced equilibrium unfolding experiments.
RESULTS AND DISCUSSION
Three-Component Model Protein Mixture Analysis.
The TD-SPROX protocol outlined in Figure 1 was initially established and tested using a model system composed of three proteins (ubiquitin, RNase A, and BCA-II) with well-established chemical denaturant induced equilibrium unfolding and SPROX behaviors.34 Shown in Figure 2 are the ion signals recorded for one of the charge states used in the RNase A analysis. Typical ion signals used to generate the values for ubiquitin and BCA-II are shown in Figure S3. Summarized in Table S1A are the wild-type and oxidized protein ion intensities obtained from Proteoform Finder. Summarized in Table S1B are the calculated value obtained for the three “one-pot” and control samples. These value were used to generate values and ultimately values for the three proteins (see Table 1), which were in reasonable agreement with values obtained from more conventional method (CD and/or conventional SPROX). We note that the different values obtained in the TD-SPROX and SPROX-ESI experiments on ubiquitin and RNase A are due to the different reaction conditions ( concentration and reaction time) used in each experiment.
Figure 2.

Mass spectrum showing the 16+ charge state of the RNase A, which contain 4 methionine residues, under various experimental conditions: Control-1 , Control-2 , One-pot (protein samples diluted across gradient GdmCl concentrations at 0.5 M increments from 0 to 5 M, treated with , and combined as the “One-pot” sample), and Control-3 .
Table 1.
Thermodynamic Parameters for the Folding Reactions of Ubiquitin, RNase A, and BCA-II
| protein | technique | transition midpoint [GdmCl] (M) | -value (kcal/(mol·M)) | (kcal/mol) |
|---|---|---|---|---|
| ubiquitin | TD-SPROX | 4.0 ± 0.1a | 2.1b | −8.2 ± 0.1a |
| SPROX-ESI | 3.2b | 2.1 ± 0.1b | −8.0 ± 0.3b | |
| CDc | 4.0d | 2.1 ± 0.2d | −8.5 ± 0.3d | |
| RNase A | TD-SPROX | 2.8 ± 0.1a | 2.0b | −6.5 ± 0.2a |
| SPROX-ESI | 2.2b | 2.0 ± 0.2b | −5.3 ± 0.5b | |
| CDc | 2.9e | 2.6e | −7.5e | |
| BCA-II | TD-SPROX | 2.3 ± 0.3a | 3.7b | −9.0 ± 1.0a |
| SPROX-MALDI | 2.1 ± 0.1b | 3.7 ± 0.7b | −8.2 ± 1.2b |
As is the case with conventional chemical denaturant-induced equilibrium unfolding experiments, the calculation of meaningful values from transition midpoint values requires the protein be well-modeled by a two-state (folded/unfolded) process and an appropriate -value be known. The values in Table 1 were all in reasonable agreement with values obtained from CD and conventional SPROX experiments, suggesting that the TD-SPROX protocol described here can generate reasonably accurate values. Moreover, the values can, in turn, be used to calculate reasonably accurate and precise values for well-behaved proteins in the chemical-denaturant induced equilibrium unfolding experiment, provided an appropriate -value is known, both of which are the case for ubiquitin and RNase A. BCA-II is not a two-state folding protein,35 therefore the calculated values are not meaningful. However, it is noteworthy that both the midpoint and values determined by TD-SPROX and SPROX-MALDI techniques were in good agreement. We also note that while the values calculated for BCA-II are not meaningful, it has been shown that value determined for this protein by SPROX can still yield useful information about perturbations to the protein’s folding properties.34
MCF-7 Cell Lysate Analysis.
The proteins in an MCF-7 cell lysate were also subjected to the TD-SPROX protocol outlined in Figure 1. The denaturant concentration range used in this experiment was abbreviated (compared to the range used for the model protein mixture) in order to focus on the transition region expected for most proteins, which is 1–2 M GdmCl based on the results of previous SPROX studies on proteins in an MCF-7 cell lysate.36 This approach has been shown to maximize the sensitivity of the one-pot SPROX experiment to detect changes in protein folding stability.37 In total, 1029 proteoforms were successfully identified in the TDP analysis (1% FDR), which only focused on the <30 kDa protein fraction, from the control sample that was not treated with GdmCl or hydrogen peroxide (Table S2).
A total of 802 proteoforms (~80%) contained at least one methionine residue (Table S2), which is required for the evaluation of values from TD-SPROX data. The data analysis here focused on the 454 most abundant proteoforms, based on their ion signals in the control sample that was not treated with GdmCl or hydrogen peroxide. The data showed good technical reproducibility (Figure S4). Unfortunately, the less abundant proteoforms did not yield high quality data sets (see below). The extracted ion intensities and values generated for each of the 454 most abundant proteoforms in the “one-pot” and control samples are summarized in Table S3A and Table S3B.
The values generated for the 454 most abundant proteoforms were filtered (see Identification of “High-Quality” SPROX Data sets in MCF-7 Cell Lysate Analysis in Supporting Information) to identify 150 proteoforms with “high-quality” TD-SPROX data sets (Table S3C). Not surprisingly, these 150 proteoforms were ones with relatively high WT-proteoform intensities in the untreated control sample (Figure S5). Higher signal intensities provide more reliable quantification because they are more easily distinguished from noise and measured more accurately and precisely than lower signal intensities. Conversely, lower signal intensities can be obscured by noise, resulting in poorer quantification. Additionally, for these high-quality SPROX data sets, the R-squared values observed for the TDP replicates were above 0.95 (Figure S6), further confirming the high-quality of these 150 data sets.
Inspection of the high-quality TD-SPROX data sets on the 150 proteoforms revealed that the values for most proteoforms were in the denaturant concentration range used in the one-pot protocol (i.e., 0.5 to 2.5 M GdmCl). Figure 3 shows an example of the TD-SPROX result for heat shock protein beta-1 (HSPB1/HSP27), with additional representative examples provided in the Supporting Information (Figures S8–S10). However, some proteoforms did have values outside of this denaturant concentration range. For example, 27 proteoforms (Table S3C) had nearly complete methionine oxidation in all the hydrogen peroxide treated samples except the Control-2 sample, which is consistent with a . A total of 5 proteoforms (Table S3C) had consistently low levels of methionine oxidation in all the hydrogen peroxide treated samples with exception of the 5 M GdmCl control sample, which is consistent with the values of these proteins being greater than 2.5 M GdmCl. Also, there were 23 proteoforms from 14 proteins (Table S3C) in which the methionine oxidation reaction was nearly complete in all the samples treated with hydrogen peroxide, indicating that these methionines were solvent-exposed in the protein’s native state. Interestingly, these proteoforms were either significantly truncated (EEF1A1, EEF1A2, KRT18, KRT8) or did have solvent-exposed methionine(s), as evidenced by X-ray crystallographic data or AlphaFold prediction (Figure S7). Among the 150 proteoforms with high-quality SPROX data, four proteins had previously reported chemical denaturant-induced equilibrium unfolding data on purified constructs using traditional techniques (Table 2). In the case of all four proteins, the TD-SPROX data were in reasonable agreement with that obtained from more conventional denaturation experiments (Table 2).
Figure 3.

TD-SPROX result for wild type heat shock protein--1 (HSPB1) from MCF-7 cell lysate. (A) Mass spectrum showing the +23 charge state of wt-HSPB1, which contains one methionine residue. Shown in (A) are data from the different experimental conditions: Control-1 , Control-2 , Control-3 , One-pot (protein samples diluted across gradient GdmCl concentrations at 0.25 M increments from 0 to 2.5 M, treated with , and combined as the “One-pot” sample), Control-4 , and Control-5 . (B) Single-scan mass spectrum showing the charge-state distributions of wt-HSPB1. (C) Sequence coverage map of wt-HSPB1 generated using TDViewer.
Table 2.
Summary of the Folding/Unfolding Transition Midpoints in [GdmCl] (M) of Selected Proteins from MCF-7 Lysate
Proteoform-Specific Folding Stability Measurements.
Included in the high-quality TD-SPROX data were 8 proteins with a number of proteoforms ranging from 2 to 21 (Table S4). A total of 5 proteins (EEF1A1, KRT8, PARK7, RKIP, and HSPB1) had at least one differentially stabilized proteoform. The proteoforms from the other 3 proteins (TPI, UBB, and TXN) were similarly stabilized (TPI) or had values that were >2.5 M GdmCl (UBB, and TXN).
The one protein that had multiple proteoforms with similar stabilities was triosephosphate isomerase (TPI, Figure S8), which is known to be a dimer of TIM barrels. Moreover, dimerization is critical for its function, enhancing catalytic turnover by rigidifying each monomer.38 The values of nine TPI proteoforms with different PTMs, point mutations, and multiple isoforms were evaluated, and all were consistently within experimental error of each other (Table S4). An analysis of the TPI homodimer’s crystal structure revealed that each monomer contains two methionine residues (M14, M82), both located at the dimer interface (Figure 4A). Intriguingly, these methionines are crucial for dimer formation and stability: each methionine is interlocked with the opposite monomer, enhancing structural integrity. Notably, mutation at either M14 or M82 significantly disrupts the dimer and reduces the protein’s thermal stability.39,40 The TD-SPROX results in this work suggest that the identified PTMs and mutations, which are located away from the dimer interface and active site (Figure 4A, PDB 1HTI), do not impact TPI’s folding stability and dimerization. These findings align with prior native top-down analyses of other TPI proteoforms, which also all had similar dimerization properties (i.e., the PTMs did not affect dimerization).41,42
Figure 4.

Schematic representation of proposed protein structures/functions (A) not impacted and (B) impacted by observed PTMs. (A) The crystal structure of the wild-type triosephosphate isomerase dimer (PDB 1HTI). Methionine residues (M14, M82) are highlighted in yellow, and residues with PTMs identified in this study are shown in pink. One monomer is distinguished in gray, and the other is depicted in a darker shade. (B) Schematic representation of HSPB1 oligomers showing large oligomers with buried methionine residue M169 (PDB 6DV5). Upon phosphorylation, these large oligomers are known to disassemble and expose the methionine residue (PDB 2N3J), which explains their oxidation in TD-SPROX. Structures were created using PyMOL (version 2.5.8).
Two proteins, elongation factor 1-alpha 1 (EEF1A1) and keratin, type II cytoskeletal 8 (KRT8), included proteoforms with significant truncations (Table S4). The TD-SPROX data collected on these truncated proteoforms revealed that the methionine residues in these proteoforms were either solvent exposed or the proteoform was unfolded at a low GdmCl concentration (e.g., ~ 0.5 M).
One protein with proteoform-specific stabilities was phosphatidylethanolamine-binding protein 1 (PEBP1/RKIP, Figure S9). RKIP is a signaling modulator that regulates key signal transduction cascades in mammalian cells, and it is known to inhibit Raf kinase by binding directly to Raf-1, preventing its phosphorylation and activation.43 The function of RKIP is known to be modulated by its own phosphorylation. For instance, phosphorylation at Ser153 in RKIP by protein kinase C (PKC) is known to disrupt the binding of RKIP to Raf-1, thereby inactivating its role as a Raf-1 inhibitor.44 Additionally, RKIP is phosphorylated by CDK5 at Thr42, a modification that facilitates the release of Raf-1 and activates the ERK pathway.45 In TD-SPROX experiments, the sole methionine (M92) on RKIP-pT42/pT6 was solvent-accessible, while in wild-type RKIP the methionine was globally protected and yielded a transition midpoint of 1.64 M GdmCl (Table S4). X-ray crystallography of RKIP indicates M92 is located in a half-pocket (Figure S7G, PDB 2L7W). Although high resolution structural data is not available on the RKIP/Raf-1 complex, our TD-SPROX data suggest that M92 is globally protected in the RKIP/Raf-1 complex.
For Parkinson disease protein 7 (PARK7), a multifunctional protein that plays a role in maintaining cellular functions,46 the three detected PARK7 proteoforms yielded transition midpoints ranging from 0.5 to 1.7 M GdmCl, indicating differential folding stabilities (Table S4). Notably, among the three PARK7 proteoforms characterized here, the two proteoforms with phosphorylation at Tyr67 were more stable than the nonphosphorylated proteoform (Table S4). This result suggests that phosphorylation at Tyr67 may confer a novel function on PARK7, warranting further investigation in future studies.
The high-quality TD-SPROX data on the heat shock protein beta-1 (HSPB1/HSP27, Figure 3, Figure S10) proteoforms can be categorized into four groups (Table S4). Group 1 contains nonphosphorylated proteoforms that all exhibit similar folding stabilities. Group 2 consists of truncated proteoforms, which are less stable than those in Group 1. Group 3 includes singly phosphorylated proteoforms, which are also less stable than those in Group 1. Proteoforms in Group 4, which have two or three phosphorylation sites, feature methionine residues that appear to be exposed. Notably, HSP27 normally assembles into large oligomers.47 Studies have shown that a single serine-to-aspartic acid phosphomimetic mutation in HSP27 slightly increases the presence of smaller oligomers, though larger oligomeric structures still predominate; with two or three such mutations, smaller oligomers become significantly more prevalent; and with three mutations, there is an almost complete inability to form large oligomeric structures.48,49 The TD-SPROX data on the proteoforms in Group 1 are consistent with the formation of large oligomeric structures, where the sole methionine residue (M169) is buried and protected (Figure 4B, PDB 6DV5). The similar transition midpoint values of the proteoforms in Group 1 suggest that PTMs such as acetylation and methylation do not alter the folding stability of these large oligomers. Conversely, in Group 4, which consists of multiply phosphorylated proteoforms, the TD-SPROX data is consistent with disassembly of the large oligomers into primarily dimer and tetramers (Figure 4B, PDB 2N3J),48,49 causing the methionine residue to become exposed and more vulnerable to oxidation (Figure 4B). The TD-SPROX data on the singly phosphorylated proteoforms in Group 3 are consistent with a mixture of both large and smaller oligomers, where the measured transition midpoint represents a weighted average of the respective midpoints of the large and smaller oligomers (~1.4 and 0 M GdmCl, respectively). The ~1.2 M transition midpoint measured for Group 3 proteoforms implies a 85/15 ratio of large/small oligomer, consistent with findings previously reported in solution.48
The proteoform specific measurements on the above proteins would not have been possible using traditional BUP SPROX workflows. This is because the TDP readout addresses the protein inference challenge in the traditional BUP SPROX workflow. For example, proteins such as PARK7, PEBP1, TPI and HSPB1 were all effectively assayed in a previously reported one-pot SPROX analysis of the proteins in an MCF-7 cell lysate using a BUP readout (Table S5).31 However, none of the methionine-containing peptides detected for these proteins in this previously reported BUP SPROX analysis covered any of the phosphosites in the phosphorylated proteoforms detected in the TD-SPROX experiment described here. Furthermore, in the case of proteoforms containing multiple PTMs, it is also impossible to assign a given methionine-containing peptide to one proteoform over another using a BUP readout.
In an earlier study we evaluated the folding stability of proteins in an MCF-7 cell using a quantitative BUP SPROX protocol involving isobaric tags for relative and absolute quantitation (iTRAQ).36 The iTRAQ-SPROX protocol in that work enabled the evaluation of transition midpoints from 850 proteins using nearly 2000 methionine-containing peptides. Included in these 850 proteins were several proteins from Table S4 (e.g., TPI, HSPB1/Hsp27, RKIP/PEBP1, and TXN). Not surprisingly, in the three cases in which transition midpoints were successfully measured in both the iTRAQ- and TD-SPROX experiments, the BUP iTRAQ-SPROX midpoints were in close agreement with that of the most abundant proteoform for each protein (Table S6), which was the wild-type, or near wild-type proteoform (i.e., Proteoforms 1375, 1078, and 1088). However, lost in the BUP iTRAQ-SRPOX experiment is any proteoform-specific stability information, which is not only important but also functionally relevant in the case of HSPB1/Hsp27 and RKIP/PEBP1.
CONCLUSION
The TD-SPROX protocol developed here enabled the evaluation of proteoform-specific folding stabilities. In cases where protein folding -values are known and the chemical denaturant-induced equilibrium unfolding properties of the test protein are well modeled by a two-state (folded/unfolded) process, the values determined in the TD-SPROX experiment can be used to calculate reasonably accurate and precise folding free energies. In other cases, TD-SPROX derived values offer valuable insights into the relative folding stabilities of different proteoforms. For example, transition midpoints for proteoforms of several proteins (e.g., HSPB1, PARK7, and RKIP) varied by up to 1.5 M GdmCl. Such relative folding stability measurements on the different proteoforms from the same gene can help identify functionally significant proteoforms. Notably, the differentially stabilized proteoforms of HSPB1 detected here can be linked to functionally important PTMs that have been previously documented. Consequently, we expect the TD-SPROX protocol described here to be a useful tool for screening proteoform function/dysfunction.
Not surprisingly, the TD-SPROX protocol is most easily applied to more abundant proteoforms. Future enhancements, such as incorporating an readout50 or targeted top-down mass spectrometry approaches using immunoenrichment,51 could help increase proteoform signal intensities and ultimately lead to better outcomes and/or broader coverage. Additionally, while the mass spectral data generated in this work did not resolve PTM isomers or distinguish between different oxidized methionine residues, future applications of the tandem mass spectrometry (e.g., proteoform reaction monitoring32) could potentially resolve these isomers.
Supplementary Material
ACKNOWLEDGMENTS
This work was supported in part by the National Institute of General Medical Sciences of the National Institutes of Health under Award Numbers R01GM134716 (to M.C.F.) and P41GM108569 to N.L.K.
Footnotes
Complete contact information is available at: https://pubs.acs.org/10.1021/acs.analchem.4c04469
Supporting Information
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.analchem.4c04469.
Detailed materials and methods; TDportal configuration; Proteoform Finder configuration; selected ion signals utilized for generating values in 3-component model proteins; technical reproducibility of signal intensity for each sample across three technical replicates; distribution of wide type proteoform ion signal intensities in the untreated control sample for proteoforms that meet the “high-quality” criteria and those that do not; reproducibility of values across three top-down replicates for the 150 “high-quality” proteoforms; protein structure with methionine residues highlighted; TD-SPROX result for triosephosphate isomerase from MCF-7 cell lysate; TD-SPROX result for phosphatidylethanolamine-binding protein 1 from MCF-7 cell lysate; TD-SPROX result for phosphorylated heat shock protein- from MCF-7 cell lysate; Peptides identified in the bottom-up SPROX data from Bailey et al; example of bottom up SPROX data from Bailey et al; Comparison of BUP SPROX data in SI ref 52 to TD SPROX data generated in this work (the TD-SPROX proteoform listed in the table is the most abundant proteoform in each protein) (PDF)
TD-SPROX analysis of 3-component model protein mixture (XLSX)
1029 assayed proteoforms in the TD-SPROX proteomics analysis of MCF-7 lysate (XLSX)
Summary of the extracted ion intensities and values generated for methionine-containing proteoforms in the MCF-7 lysate (XLSX)
The authors declare the following competing financial interest(s): N.L.K. is involved in entrepreneurial activities in top-down proteomics and a shareholder of Proteinaceous Inc., which commercializes the software used for the analysis of top-down mass spectrometry data. The other authors declare no further competing interests.
Contributor Information
You Zou, Department of Chemistry, Duke University, Durham, North Carolina 27708, United States.
Che-Fan Huang, Department of Chemistry, Northwestern University, Evanston, Illinois 60208, United States; Department of Molecular Biosciences and Feinberg School of Medicine, Northwestern University, Evanston, Illinois 60208, United States.
Grace R. Sturrock, Department of Chemistry, Duke University, Durham, North Carolina 27708, United States
Neil L. Kelleher, Department of Chemistry, Northwestern University, Evanston, Illinois 60208, United States; Department of Molecular Biosciences and Feinberg School of Medicine, Northwestern University, Evanston, Illinois 60208, United States.
Michael C. Fitzgerald, Department of Chemistry, Duke University, Durham, North Carolina 27708, United States.
REFERENCES
- (1).Smith LM; Kelleher NL Nat. Methods 2013, 10 (3), 186–187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (2).Aebersold R; Agar JN; Amster IJ; Baker MS; Bertozzi CR; Boja ES; Costello CE; Cravatt BF; Fenselau C; Garcia BA; et al. Nat. Chem. Biol 2018, 14 (3), 206–214. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (3).Burnum-Johnson KE; Conrads TP; Drake RR; Herr AE; Iyengar R; Kelly RT; Lundberg E; MacCoss MJ; Naba A; Nolan GP Mol. Cell. Proteomics 2022, 21 (7), 100254. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (4).Schmitt ND; Agar JN J. Mass Spectrom 2017, 52 (7), 480–491. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (5).Ntai I; Fornelli L; DeHart CJ; Hutton JE; Doubleday PF; LeDuc RD; Van Nispen AJ; Fellers RT; Whiteley G; Boja ES; Rodriguez H; Kelleher NL Proc. Natl. Acad. Sci. U. S. A 2018, 115 (16), 4140–4145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (6).Gregorich ZR; Peng Y; Cai W; Jin Y; Wei L; Chen AJ; McKiernan SH; Aiken JM; Moss RL; Diffee GM; Ge YJ Proteome Res. 2016, 15 (8), 2706–2716. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (7).Kang J; Seshadri M; Cupp-Sutton KA; Wu S Front. Anal. Sci 2023, 3, 1186623. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (8).Melby JA; Roberts DS; Larson EJ; Brown KA; Bayne EF; Jin S; Ge YJ Am. Soc. Mass Spectrom 2021, 32 (6), 1278–1294. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (9).Regnier FE; Kim J Anal. Chem 2018, 90 (1), 361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (10).Geer MA; Fitzgerald MC Annu. Rev. Anal. Chem 2014, 7 (1), 209–228. [DOI] [PubMed] [Google Scholar]
- (11).Shoichet BK; Baase WA; Kuroki R; Matthews BW Proc. Natl. Acad. Sci. U. S. A 1995, 92 (2), 452–456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (12).Molina DM; Jafari R; Ignatushchenko M; Seki T; Larsson EA; Dan C; Sreekumar L; Cao Y; Nordlund P Science 2013, 341 (6141), 84–87. [DOI] [PubMed] [Google Scholar]
- (13).Savitski MM; Reinhard FBM; Franken H; Werner T; Savitski MF; Eberhard D; Molina DM; Jafari R; Dovega RB; Klaeger S; Kuster B; Nordlund P; Bantscheff M; Drewes G Science 2014, 346 (6205), 1255784. [DOI] [PubMed] [Google Scholar]
- (14).Lomenick B; Hao R; Jonai N; Chin RM; Aghajan M; Warburton S; Wang J; Wu RP; Gomez F; Loo JA; Wohlschlegel JA; Vondriska TM; Pelletier J; Herschman HR; Clardy J; Clarke CF; Huang J Proc. Natl. Acad. Sci. U. S. A 2009, 106 (51), 21984–21989. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (15).Feng Y; De Franceschi G; Kahraman A; Soste M; Melnik A; Boersema PJ; De Laureto PP; Nikolaev Y; Oliveira AP; Picotti P Nat. Biotechnol 2014, 32 (10), 1036–1044. [DOI] [PubMed] [Google Scholar]
- (16).Liu P-F; Kihara D; Park CJ Mol. Biol 2011, 408 (1), 147–162. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (17).Espino JA; Mali VS; Jones LM Anal. Chem 2015, 87 (15), 7971–7978. [DOI] [PubMed] [Google Scholar]
- (18).West GM; Tucker CL; Xu T; Park SK; Han X; Yates JR; Fitzgerald MC Proc. Natl. Acad. Sci. U. S. A 2010, 107 (20), 9078–9082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (19).DeArmond PD; Xu Y; Strickland EC; Daniels KG; Fitzgerald MC J. Proteome Res 2011, 10 (11), 4948–4958. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (20).Bamberger C; Pankow S; Martínez-Bartolomé S; Ma M; Diedrich J; Rissman RA; Yates JR J. Proteome Res 2021, 20 (5), 2762–2771. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (21).Nesvizhskii AI; Aebersold R Mol. Cell. Proteomics 2005, 4 (10), 1419–1440. [DOI] [PubMed] [Google Scholar]
- (22).Huang C-F; Su P; Fisher TD; Levitsky J; Kelleher NL; Forte E Front. Transplant 2023, 2, 1286881. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (23).Cai W; Tucholski T; Chen B; Alpert AJ; McIlwain S; Kohmoto T; Jin S; Ge Y Anal. Chem 2017, 89 (10), 5467–5475. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (24).Li Y; Compton PD; Tran JC; Ntai I; Kelleher NL PROTEOMICS 2014, 14 (10), 1158–1164. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (25).Tran JC; Zamdborg L; Ahlf DR; Lee JE; Catherman AD; Durbin KR; Tipton JD; Vellaichamy A; Kellie JF; Li M; et al. Nature 2011, 480 (7376), 254–258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (26).Cupp-Sutton KA; Wu S Mol. Omics 2020, 16 (2), 91–99. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (27).Hung C-W; Tholey A Anal. Chem 2012, 84 (1), 161–170. [DOI] [PubMed] [Google Scholar]
- (28).Yu D; Wang Z; Cupp-Sutton KA; Guo Y; Kou Q; Smith K; Liu X; Wu SJ Am. Soc. Mass Spectrom 2021, 32 (6), 1336–1344. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (29).Winkels K; Koudelka T; Tholey AJ Proteome Res. 2021, 20 (9), 4495–4506. [DOI] [PubMed] [Google Scholar]
- (30).Cabrera A; Wiebelhaus N; Quan B; Ma R; Meng H; Fitzgerald MC J. Am. Soc. Mass Spectrom 2020, 31 (2), 217–226. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (31).Bailey MA; Tang Y; Park H-J; Fitzgerald MC J. Am. Soc. Mass Spectrom 2023, 34 (3), 383–393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (32).Huang C-F; Kline JT; Negrão F; Robey MT; Toby TK; Durbin KR; Fellers RT; Friedewald JJ; Levitsky J; Abecassis MMI; Melani RD; Kelleher NL; Fornelli L Anal. Chem 2024, 96 (8), 3578–3586. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (33).Takemori A; Butcher DS; Harman VM; Brownridge P; Shima K; Higo D; Ishizaki J; Hasegawa H; Suzuki J; Yamashita M; Loo JA; Loo RRO; Beynon RJ; Anderson LC; Takemori NJ Proteome Res. 2020, 19 (9), 3779–3791. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (34).West GM; Tang L; Fitzgerald MC Anal. Chem 2008, 80 (11), 4175–4185. [DOI] [PubMed] [Google Scholar]
- (35).Henkens RW; Kitchell BB; Lottich SC; Stein PJ; Williams TJ Biochemistry 1982, 21 (23), 5918–5923. [DOI] [PubMed] [Google Scholar]
- (36).Ogburn RN; Jin L; Meng H; Fitzgerald MC J. Proteome Res 2017, 16 (11), 4073–4085. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (37).Xu Y; West GM; Abdelmessih M; Troutman MD; Everley RA ACS Chem. Biol 2021, 16 (8), 1445–1455. [DOI] [PubMed] [Google Scholar]
- (38).Schliebs W; Thanki N; Jaenicke R; Wierenga RK Biochemistry 1997, 36 (32), 9655–9662. [DOI] [PubMed] [Google Scholar]
- (39).Mainfroid V; Terpstra P; Beauregard M; Frere J-M; Mande SC; Hol WG; Martial JA; Goraj KJ Mol. Biol 1996, 257 (2), 441–456. [DOI] [PubMed] [Google Scholar]
- (40).Roland BP; Zeccola AM; Larsen SB; Amrich CG; Talsma AD; Stuchul KA; Heroux A; Levitan ES; VanDemark AP; Palladino MJ PLoS Genet. 2016, 12 (3), No. e1005941. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (41).Yates J; Gomes F; Durbin K; Schauer K; Nwachukwu J; Russo R; Njeri J; Saviola A; McClatchy D; Diedrich J Res. Sq 2023, DOI: 10.21203/rs.3.rs-3097806/v1. [DOI] [Google Scholar]
- (42).Schachner LF; Soye BD; Ro S; Kenney GE; Ives AN; Su T; Goo YA; Jewett MC; Rosenzweig AC; Kelleher NL ACS Chem. Biol 2022, 17 (10), 2769–2780. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (43).Yeung K; Seitz T; Li S; Janosch P; McFerran B; Kaiser C; Fee F; Katsanakis KD; Rose DW; Mischak H; et al. Nature 1999, 401 (6749), 173–177. [DOI] [PubMed] [Google Scholar]
- (44).Corbit KC; Trakul N; Eves EM; Diaz B; Marshall M; Rosner MR J. Biol. Chem 2003, 278 (15), 13061–13068. [DOI] [PubMed] [Google Scholar]
- (45).Wen Z; Shu Y; Gao C; Wang X; Qi G; Zhang P; Li M; Shi J; Tian B Neurobiol. Aging 2014, 35 (12), 2870–2880. [DOI] [PubMed] [Google Scholar]
- (46).Clements CM; McNally RS; Conti BJ; Mak TW; Ting JP-Y Proc. Natl. Acad. Sci. U. S. A 2006, 103 (41), 15091–15096. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (47).Sun Y; MacRae TH Cell. Mol. Life Sci. CMLS 2005, 62, 2460–2476. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (48).Rogalla T; Ehrnsperger M; Preville X; Kotlyarov A; Lutsch G; Ducasse C; Paul C; Wieske M; Arrigo A-P; Buchner J; et al. J. Biol. Chem 1999, 274 (27), 18947–18956. [DOI] [PubMed] [Google Scholar]
- (49).Jovcevski B; Kelly MA; Rote AP; Berg T; Gastall HY; Benesch JL; Aquilina JA; Ecroyd H Chem. Biol 2015, 22 (2), 186–195. [DOI] [PubMed] [Google Scholar]
- (50).Kafader JO; Melani RD; Durbin KR; Ikwuagwu B; Early BP; Fellers RT; Beu SC; Zabrouskov V; Makarov AA; Maze JT; et al. Nat. Methods 2020, 17 (4), 391–394. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (51).Drown BS; Gupta R; McGee JP; Hollas MAR; Hergenrother PJ; Kafader JO; Kelleher NL Anal. Chem 2024, 96 (11), 4455–4462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (52).Dai SY; Fitzgerald MC J. Am. Soc. Mass Spectrom 2006, 17 (11), 1535–1542. [DOI] [PubMed] [Google Scholar]
- (53).Ahmad F; Bigelow CC J. Protein Chem 1986, 5, 355–367. [Google Scholar]
- (54).Weber C; Kraemer S; Drechsler M; Lue H; Koenen RR; Kapurniotu A; Zernecke A; Bernhagen J Proc. Natl. Acad. Sci. U. S. A 2008, 105 (42), 16278–16283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (55).Wang MZ; Shetty JT; Howard BA; Campa MJ; Patz EF; Fitzgerald MC Anal. Chem 2004, 76 (15), 4343–4348. [DOI] [PubMed] [Google Scholar]
- (56).Main ER; Fulton KF; Jackson SE J. Mol. Biol 1999, 291 (2), 429–444. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
