Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Apr 29.
Published in final edited form as: Anal Chem. 2024 Dec 16;96(52):20481–20490. doi: 10.1021/acs.analchem.4c04466

Improving glycoproteomic analysis workflow by systematic evaluation of glycopeptide enrichment, quantification, mass spectrometry approach, and data analysis strategies

Zhenyu Sun 1, T Mamie Lih 1, Jongmin Woo 1, Liyuan Jiao 1, Yingwei Hu 1, Yuefan Wang 1, Hongyi Liu 1, Hui Zhang 1,*
PMCID: PMC12039365  NIHMSID: NIHMS2067566  PMID: 39679613

Abstract

Glycosylation is one of the most prevalent and crucial protein modifications. Quantitative site-specific characterization of glycosylation usually requires sophisticated intact glycopeptide analysis using glycoproteomics. Recent efforts have focused on the interrogation of intact glycopeptide analysis using tandem mass spectrometry. However, a systematic evaluation of quantitative glycoproteomic workflow is still lacking. This study compared different strategies for glycopeptide enrichment alongside glycopeptide quantitation, mass spectrometry strategies, and data analysis strategies, providing a comprehensive assessment of their efficacy. The ZIC-HILIC enrichment method demonstrated superior in identification capacity, representing a 26% improvement compared to the MAX enrichment method. Quantification using TMT provided high precision and throughput, with an average CV of 8%. Through systematic evaluation, this study established that ZIC-HILIC enrichment method, quantification with TMT, and stepped collision energies of 25, 35, and 45 using tandem mass spectrometry are optimal workflow for HCD fragmentation, significantly enhancing the analysis of intact glycopeptides. Precise energy adjustment is crucial for the identification of certain glycans. Our results indicate that glycan dissociation is highly sensitive to collision energy, with higher energies potentially causing excessive fragmentation or reducing detection. TMTpro-labeled intact glycopeptides were analyzed using three different software tools, pGlyco 3.0, MS-PyCloud, and MSFragger-Glyco to investigate the identification and quantification of each analysis tool. By applying optimal settings, 5,514 unique intact glycopeptides were identified in luminal and basal patient-derived xenograft (PDX) samples, highlighting distinct glycosylation profiles that may influence tumor behavior and patient response to therapy. These findings offer valuable insights for glycoproteomic analyses and clinical applications.


Glycosylation is one of the most prevalent and crucial protein modifications PTMs and plays pivotal roles in various biological processes, including protein folding, trafficking, and stability 1,2. The synthesis of glycans is template-independent, resulting in micro-and macro-heterogeneity, posing significant challenges to glycosylation study 3,4. Mass spectrometry (MS) has emerged as a powerful tool for protein glycosylation analysis of intact glycopeptides for site-specific characterization of micro- and macro-heterogeneity 5,6. Recently, a range of MS-based methods for comprehensive intact N-glycopeptide analysis have been developed 7,8. MS analysis can simultaneously provide information on glycosites and the types of glycans for intact glycopeptides. Additionally, quantitative analysis of intact glycopeptides can be performed, providing protein glycosylation information. Currently, the workflow for analyzing intact glycopeptides involves enzymatic digestion of proteins, followed by enrichment of glycopeptides, and subsequent analysis using LC-MS/MS analysis 9. In peptide mixtures resulting from trypsin or other proteolytic enzyme treatments, intact glycopeptides typically represent in low abundance (approximately 1% of total peptide content) 10. Consequently, their enrichment of intact glycopeptides before MS analysis is crucial for complex biological or clinical samples. This step enhances the signal intensity of intact glycopeptides and minimizes interference from non-glycosylated peptides from the complex samples. Various enrichment strategies for intact glycopeptides have been developed, significantly improving their coverage in different types of samples 11. Hydrophilic interaction liquid chromatography (HILIC) and mixed-mode strong anion exchange (MAX) enrichment strategies were the most used in glycoproteomics, which can separate intact glycopeptides via the hydrophilic properties of glycans 12,13. There are several HILIC materials, and ZIC-HILIC, based on zwitterionic stationary phases, is a common type of HILIC. ZIC-HILIC can selectively enrich glycopeptides based on their hydrophilicity and charge state, offering higher enrichment specificity and sensitivity compared to traditional HILIC methods 14. Currently, a systematic evaluation of MAX and ZIC-HILIC enrichment methods are still lacking.

Even though excellent enrichment methods can mitigate the impact of non-glycosylated peptides on intact glycopeptides analysis, the micro-heterogeneity of intact glycopeptides still makes their quantitative analysis challenging. In the era of precision medicine, the demand for MS analysis of large-scale clinical samples is steadily increasing 15. This emphasizes the need to develop robust and highly reproducible quantitative methods for analyzing large cohorts of clinical samples. The MS-based quantification glycoproteomic technique consists of label based and label-free methods. Sample preparation for the label-free method is more convenient and flexible. However, the abundance of some intact glycopeptides carrying certain glycans is low, which are difficult to identify or quantify. Missing values is one of the main problems associated with label-free glycoproteomics, significantly undermining the quantification reliability and downstream data analysis 16. Label-based methods, such as Tandem mass tagging (TMT) strategy, have the advantage of allowing multiple samples analysis simultaneously, which can be used for investigating multiple samples at the same time. For intact glycopeptides, glycans dissociate more easily than reporter groups because glycosidic bonds absorb most of the collisional energy, which limits the generation of reporter ions 17. Nevertheless, systematically comparing the quantitative abilities of labeled and label-free methods for intact glycopeptides serves as a fundamental cornerstone for subsequent large-scale glycoproteomic investigations based on clinical samples.

After sample preparation, interrogating glycopeptides using tandem MS to produce efficient fragments for successful peptide and glycan determination remains a formidable task. Unlike most other protein modifications, which involve a fixed mass shift, glycosylation is complex due to the varied composition of glycans 18. Optimal MS/MS acquisition can generate comprehensive fragments, both the glycan and the peptide, for each intact glycopeptide, where collision energy plays critical role that can significantly influence the identification of glycan types due to variations in oxonium ion fragmentation 19.

The combination of various strategies for glycoproteomic using glycopeptide enrichment, mass spectrometry techniques, quantification, and data analysis strategies have not been systematically evaluated in previous research. In this study, we conducted a systematic investigation into glycoproteomic workflow. Firstly, Label-free and TMT labeling strategies were utilized for intact N-glycopeptide quantification. A total of 2,924 unique intact N-glycopeptides can be quantified after being labeled using the TMT approach. Compared to label-free method, TMT labeling for its high precision with the CV averaging 8% and multiplex analysis capabilities. Subsequently, we compared the capacity of ZIC-HILIC and MAX enrichment methods, revealing that the ZIC- HILIC method demonstrated superior performance in enriching intact N-glycopeptides, which can identify 3,004 GPSMs from a single injection of patient-derived xenograft (PDX) sample. Furthermore, we systematically evaluated different higher-energy collisional dissociation (HCD) collision energies, determining that the stepped collision energy HCD (sceHCD) settings of 25, 35, and 45 was the optimal for both the identification and quantification of intact N-glycopeptides. Of note, our findings indicated a relationship between dissociation of glycans and applied collision energy, emphasizing the necessity of tailored fragmentation energies for comprehensive glycoform analysis. TMTpro-labeled intact glycopeptides were analyzed using three different software tools, pGlyco 3.0, MS-PyCloud and MSFragger-Glyco, to investigate the identification of quantification of different glycoproteomic analysis tools for glycopeptide identification and quantification. We applied the aforementioned optimal settings to breast cancer PDX samples, a total of 5,514 unique intact N-glycopeptides were identified from luminal and basal subtypes of the PDX samples, revealing distinct glycosylation profiles. These unique profiles may play a significant role in influencing tumor behavior and patient response to therapy. The significance of this study lies in providing effective technical routes and methodological strategies for optimizing MS analysis of glycosylation, laying an important groundwork for further exploration of the value of glycoproteomics in clinical applications.

EXPERIMENTAL SECTION

Materials and Chemicals.

Trypsin was from Promega (Madison, WI); Lys-C endopeptidase (Wako Chemicals); Sep-Pak C18 1 cm3 Vac Cartridge was from Waters (Milford, MA); ZIC-HILIC columns from Sigma aldrich (St. Louis, MO), MAX columns from Waters (Milford, MA); Other chemicals such as urea, a, acetonitrile (ACN), trifluoroacetic acid (TFA), Iodoacetamide (IAA), iodoacetamide, triethylammonium acetate were purchased from Sigma Aldrich (St. Louis, MO).

Protein extract and digestion.

The PDX tumors used in this study were from previously established basal (WHIM2) and luminal (WHIM16) breast cancers. For protein extraction, 400 μL of 8M urea lysis buffer were added to 100 mg tissue for protein extract. The concentrations of proteins were determined using BCA assay. Samples were incubated in 5 mM dithiothreitol (DTT) for 1 h at room temperature and then alkylated by 10 mM iodoacetamide (IAA) for 45 min at room temperature in the dark. Proteins were then diluted to 2M urea with gentle shaking for 2 h by Lys-C digestion at an enzyme-to-protein ratio of 1:50 (wt/wt). And then trypsin was added at an enzyme-to-protein ratio of 1:50, digested overnight at room temperature, ensuring complete digestion for accurate glycopeptide analysis. Digestion was terminated using 50% formic acid (FA) by adjusting solution pH to about 2–3 20.

Data availability.

The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD054874 21.

RESULTS AND DISCUSSION

Glycoproteomic workflow for analyzing intact glycopeptides involves enzymatic digestion of proteins, followed by enrichment of glycopeptides, and subsequent data acquisition using LC-MS/MS and data analysis (Scheme 1). In this work, we will evaluate each step for setting up the optimal workflow for clinical samples.

Scheme 1.

Scheme 1

The general workflow of glycoproteomics analysis

Quantification strategies for glycopeptides analysis

For quantification strategies, both label-free and label-based methods could be used. Isobaric labeling with TMT enables proteome-wide relative quantification for multiple samples simultaneously in one LC-MS/MS analysis 22, which could reduce MS analysis time. 18plex TMTpro labeling was included in our workflow, a total of 18 individual samples can be labeled and analyzed in one LC-MS/MS analysis. In this study, the label-free strategy, we identified 5253 glycopeptide spectral matches (GPSMs) and 3042 unique glycopeptides from one LC-MS/MS injection. Following TMT labeling, 2,702 GPSMs and 1,992 unique glycopeptides were identified (Figure 1a). The overlap between label-free and TMT labeling glycopeptide is 1,193 (Figure 1b). We also calculated TMT labeling efficiency and achieved an efficiency of 99%.

Figure 1. The comparison between label-free and TMTpro labeling intact N- glycopeptides.

Figure 1

(a) The identification number of GPSMs and unique intact N-glycopeptides from label-free and TMTpro labeling glycopeptides. (b) The overlap between label-free and TMTpro labeling strategies. (c) The distribution of identified N-glycan types from unlabeled and TMTpro labeling strategies. (d) The CV distribution between label-free and TMTpro labeling strategies.

We further compared the N-glycan types between TMT and label-free methods. The number of fucose and high mannose glycans was similar, but 7% more sialic acids were identified after TMTpro labeling, while glycans without sialic acids were increased in label-free method (Figure 1c). TMTpro labeling converts the primary amines of the lysine residues into moieties containing dimethyl piperidine groups, increasing proton affinity and the number of possible charge states during ionization. Consequently, more sialic acids could be detected compared to unlabeled glycopeptides, as sialic acids contain carboxyl groups, which have low intensity in MS. Furthermore, among the identified sialylated glycopeptides, more than 7% of N-glycans contained two sialic acids and more than 1.6% N-glycans contained three sialic acids (Figure S1). Moreover, we compared the coefficient of variation (CV) distributions between label-free and TMTpro to assess the quantitative precision of each method (Figure 1d). We observed that the average CV was 21% for label-free method and 8% for TMTpro labeling method, indicating higher quantitative reproducibility and reduced variability with TMTpro labeling. In addition, the TMTpro labeling approach maintained high reproducibility, with the CV for quantified glycopeptides remaining below 15% in more than 98% glycopeptides. Thus, we selected the TMTpro labeling strategy to decrease variations among replicates. In addition, stable isotope labeling boosted the intensity of intact N-glycopeptides with relatively lower abundance in the samples, such as those containing sialic acids.

Enrichment strategies for Intact glycopeptides

Hydrophilic interaction liquid chromatography (HILIC) and mixed-mode strong anion exchange (MAX) enrichment strategies were the most used in glycoproteomics, which can separate intact glycopeptides via the hydrophilic properties of glycans 15,23. ZIC-HILIC, based on zwitterionic stationary phases, is a common type of HILIC. ZIC-HILIC can selectively enrich glycopeptides based on their hydrophilicity and charge state, offering higher enrichment specificity and sensitivity compared to traditional HILIC methods 24,25. In our previous studies, we employed a sequential enrichment method to simultaneously analyze glycosylation and phosphorylation modifications 15,2627. Peptides were labeled with TMTpro. After being labeled, phosphorylation enrichment using IMAC was employed at first since our previous research has already optimized the workflow 28. The flowthrough peptides that were enriched by phosphorylation enrichment were used for glycopeptide enrichment. The ZIC-HILIC method showed superior enrichment efficacy for intact glycopeptide identifications compared to the MAX approach, identifying 880, 1,271, and 816 GPSMs through IMAC, ZIC-HILIC and MAX enrichment methods, respectively. These GPSMs represented 572 (IMAC), 881 (ZIC-HILIC), and 582 (MAX) unique intact glycopeptides were identified, respectively (Figure 2a). From these, a total of 2,151 GPSMs were identified through combination of IMAC with ZIC-HILIC enrichment from flow-through enrichment while 1,696 GPSMs were identified through combination of IMAC with MAX. The enrichment specificity was determined by the ratio of oxonium ion-containing MS2 spectra to total MS2 spectra. The ratios were 45% for HILIC enrichment and 6% for MAX enrichment, respectively. We analyzed the overlap of unique intact glycopeptides identified by ZIC-HILIC and MAX enrichment methods to assess their comparative performance. The overlapping identification between these two methods was only 374 unique glycopeptides, which indicated that more unique glycopeptides were identified by using ZIC-HILIC method (Figure 2b). In terms of glycan types, the ZIC-HILIC column preferred to enrich high mannose N-glycans and sialic acid containing glycans (Figure 2c). Both methods showed similar efficiency for enrichment of fucose and other N-glycan types, consistent with the previous study 12. Furthermore, the MAX method identified 10% more N-glycans than the ZIC-HILIC method in the mass range between 1800–2300, while ZIC-HILIC method identified 4% more N-glycans with mass weight over 2300, which indicated ZIC-HILIC preferred to enrich glycans with larger mass weight (Figure 2d). An analysis of direct glycopeptide enrichment, without prior phosphopeptide enrichment, yielded consistent findings, demonstrating that the ZIC-HILIC strategy outperformed the MAX enrichment approach (Figure S2). Therefore, we opted to utilize the ZIC-HILIC strategy for the enrichment and analysis of glycopeptides, based on its demonstrated superiority in identifying a wider range of N-glycan types.

Figure 2. The characteristics of glycopeptides identified by ZIC-HILIC and MAX methods.

Figure 2

(a) The numbers of glycopeptide identifications using ZIC-HILIC or MAX. (b) The overlapped and uniquely identified intact N-glycopeptides using ZIC-HILIC and MAX. (c) The distribution of N-glycan types identified using ZIC-HILIC and MAX methods. (d) The distribution of N- glycan mass weight using ZIC-HILIC and MAX methods.

Optimal dissociation conditions for enhancing Identification of various glycan types in Intact glycopeptide analysis

To characterize the site-specific N-glycosylation of proteins from complex samples, it is critical to select an appropriate dissociation method. HCD has been used with success in large-scale glycoproteomic experiments 29. In this study, we created nine methods to systematically evaluate different HCD collision energy for their performance in characterizing intact N- glycopeptides (Table S1). The average number of GPSMs from each method is summarized in Figure 3a. Stepped collision energy HCD (sceHCD) methods clearly outperform single HCD methods, since much higher number of intact N- glycopeptides were identified using sceHCD (Figure 1a), especially when using sceHCD of 25,35, and 45. A total of 5,305 GPSMs and 2,924 unique N-glycopeptides were identified using sceHCD of 25, 35, and 45. Therefore, sceHCD of 25,35, and 45 showed the optimal dissociation condition, allowing the identification of various glycan types, thereby improving the reliability of glycopeptide profiling. Moreover, as shown in Figure 3b, high mannose N-glycans emerged as the most prevalent type across various collision energies. Intriguingly, a direct correlation was observed between increasing collision energy and the identification of sialic acids, with an energy setting of 25_35_45 being optimal for their identification. Raising the collision energy from 20 to 35 resulted in an increase in the proportion of sialic acid from 3.15% to 8.51%, showing a linear relationship between collision energy and sialic acid fragmentation in this range. However, the highest energy level of 45, only a small number of sialic acids were detected (0.3%). The identification of intact N-glycopeptides is influenced by both the peptide sequence and the glycan moiety 19. We selected a peptide sequence, LLNINPNK of lysosome-associated membrane glycoprotein 1 (LAPM1) as an example to further demonstrate the impact of collision energy on different glycan types (Figure S3). With an energy setting of 35, a total of 12 salic acids were identified, which was the highest number among the nine methods. By examing the intensities of the representative oxonium ions from all the MS2 spectrum, we found the intensity of NeuAc and NeuGc ions were correlated with the identification of sialic acid (Figure 3c). The intensities of these oxonium ions progressively increased with the rise in normalized collision energy (NCE) levels. When the collision energy exceeds 40, the intensity of HexNAc, NeuAc and NeuGc significantly decreased. However, the oxonium ions of fucose decreased when the energy was increased to 45. Under collision energy ranging from 20 to 35, the ability to detect fucose remains consistent. Increasing the collision energy to 40 resulted in a significant reduction in the number of fucose identifications, although the intensity of related oxonium ion remains notably high. The results show that glycan dissociation is sensitive to the collision energies used, suggesting that higher energies may cause excessive fragmentation or reduce their detection. These findings highlighted the necessity of selecting specific fragmentation energies for targeted glycoform analysis. In summary, precise energy adjustment was crucial for accurate identification of glycans, particular for sialic acids. By setting appropriate collision energy would enhance the overall efficacy and specificity of glycoproteomic investigations. Since TMT labeling needs additional collision energies to form reporter ions, we also set different collision energy conditions for MS analysis (Table S2) 16. The sceHCD setting at 25, 35, and 45 was still the best condition for the TMT-labeled intact N-glycopeptide analysis (Figure S4). Therefore, sceHCD setting of 25, 35, and 45 was selected for the later analysis.

Figure 3. The characterization of label-free intact N- glycopeptides using different collision energy methods.

Figure 3

(a) The number of GPSMs and unique intact N-glycopeptides identified from different collision energies. (b) The glycan type distribution from different collision energy methods, “SA” represents sialic acids, “Fuc” represents fucose, and “HM” represents high mannose. (c) The MS2 intensity of oxonium ions from different collision energy methods.

Evaluation of software tools for intact glycopeptide analysis

Compared to proteomics analysis, intact glycopeptide data analysis presents significant challenges. An intact glycopeptide comprises both a glycan and a peptide backbone, leading to the generation of multiple ion types during mass spectrometry fragmentation. This substantially increases the complexity and information density of the spectra. Unlike other protein PTM, glycosylation does not add a fixed mass. With over 1,000 glycan structures reported in glycan databases 30, the diversity of glycans adds complexity to intact glycopeptide analysis.

Recently, advancements have been made in MS data analysis for intact glycopeptide analysis. Three different software tools, MS-PyCloud, pGlyco 3.0, and MSFragger-Glyco, were used to analyze ZIC-HILIC enriched and TMTpro-labeled intact glycopeptides in this study 3133. The analysis identified 2,701, 3,095, and 3,704 GPSMs using pGlyco 3.0, MS-PyCloud, and MSFragger-Glyco, respectively. The number of unique intact glycopeptides identified was 1,992 for pGlyco 3.0, 2,444 for MS-PyCloud, and 2,093 for MSFragger-Glyco (Figure 4a). The number of identified glycoproteins, glycosites, and site-specific glycans is summarized in Figure S5. A total of 403, 560, and 432 glycoproteins were identified using pGlyco 3.0, MS-PyCloud, and MSFragger-Glyco, respectively. For N-glycans, 119, 262, and 291 were identified using pGlyco 3.0, MS-PyCloud, and MSFragger, respectively. Additionally, 617, 798, and 705 N-glycosites were identified with pGlyco 3.0, MS-PyCloud, and MSFragger-Glyco, respectively. Across these three software tools, a total of 593 glycans, 1163 peptide backbones, and 626 proteins were identified. All three analyses used the same N-glycan database, and the glycan types identified are summarized in Figure 4b. High-mannose glycans were the most prevalent in all three datasets. Notably, 23.9% of the intact glycopeptides identified by MS-PyCloud contained sialic acids, compared to 12.8% and 19.6% in the results from pGlyco 3.0 and MSFragger-Glyco, respectively. We also summarized the peptides and proteins identified by the three software tools, which illustrates the distribution of glycans, peptides, and proteins across different software combinations, highlighting both overlaps and unique identifications in each (Figure S6). Each software tool identified unique components, with the overlap between MS-PyCloud and MSFragger-Glyco being the most, 181 glycans, 501 peptides, and 349 proteins were commonly identified by both tools. Each software provides reliable results and users can select the most suitable software based on their sample type and specific needs for glycans for data analysis. pGlyco 3.0 was selected for further data analysis in this study.

Figure 4. The characterization of TMTpro-labeled intact N- glycopeptides using different software tools.

Figure 4

(a) The number of GPSMs and unique intact N-glycopeptides identified from different software. (b) The glycan type distribution from different software, “SA” represents sialic acids, “Fuc” represents fucose, and “HM” represents high mannose.

Investigating protein glycosylation in luminal and basal breast cancer subtypes using PDX

Glycosylation of cell surface proteins plays an important role in the regulation of apoptosis 34. Aberrant glycosylation remodeling and metabolism are associated with the epithelial mesenchymal transition (EMT) and metastasis in breast cancer 35. The epithelium of the lactiferous ducts in the breast is comprised of luminal epithelial cells and underlying basal myoepithelial cells 36. We applied the optimal method to investigate the differences in N-glycosylation of luminal and basal subtypes using PDX mouse models for better understanding the function of N-glycosylation in breast cancer.

Proteins extracted from the PDX samples were digested to peptides, followed by TMTpro labeling and offline high-pH research phase separation to separate peptides into 12 fractions to increase the depth of identification. Each fraction underwent phosphorylation enrichment prior to the enriched of intact N-glycopeptides using the follow-through as aforementioned.

A total of 5,514 unique intact N-glycopeptides (FDR < 0.01), mapping to 686 glycoproteins were quantified from basal and luminal samples. Principal component analysis (PCA) showed a clear separation between luminal and basal subtypes samples based on the abundances of intact glycopeptides (Figure 5a). There were notable differences in the expression of intact N-glycopeptides between the two subtypes (Figure 5b). The majority of the intact N-glycopeptides were from the proteins in the extracellular matrix, cell surface, lysosome, and plasma membrane regions (Figure 5c). Differential analysis between basal and luminal subtypes revealed a total of 711 intact glycopeptides mapping to 179 proteins with significantly upregulation of at least 2-fold changes (p-value < 0.05), and 1299 intact glycopeptides mapping to 242 proteins were downregulated more than 2-fold changes (p-value < 0.05) in luminal subtype compared to basal subtype (Figure 5d and Table S3). The differential expression of intact N-glycopeptides was statistically analyzed by glycan types, revealing a marked decrease in the expression of fucose, high mannose and sialic acids in luminal samples (Figure 5e). EpCAM (CD326), which plays diverse roles in cell adhesion and proliferation and is known to be overexpressed in primary breast carcinomas (PBCs) 37. A total of 12 intact N-glycopeptides of EpCAM was found with reduced expression in luminal samples. Among these 12 intact N-glycopeptides, 4 were exclusively fucosylated, while the remainder were both fucosylated and sialylated (Figure 5f). The decreased expression of sialylated glycopeptides in luminal samples may align with findings reported in previous studies 38,39. These results validated the successful application of our optimized workflow for intact N-glycopeptide analysis, demonstrating its potential for large-scale clinical sample analysis. Taken together, our established workflow not only enhanced understanding of N-glycosylation differences between breast cancer subtypes but also provided strong backing for future research and treatment approaches.

Figure 5. Analysis to altered N-glycosylation in luminal and basal breast cancer subtypes of PDX.

Figure 5

(a) Principal component analysis between luminal and basal breast cancer subtypes; (b) Heatmap shows the intact N- glycopeptide abundance and N-glycans between luminal and basal breast cancer subtypes; (c) GO analysis of identified intact N-glycopeptides in terms of cellular component; (d) Differential analysis between luminal and basal breast cancer subtypes. Significantly altered intact glycopeptides were defined as > 2-fold changes with a p-value < 0.05. (e) The distribution of altered N-glycans between luminal and basal breast cancer subtypes; (f) The distribution of downregulated intact N-glycopeptides of EpCAM in luminal subtype compared to basal subtype.

Conclusions

In conclusion, we selected the TMTpro labeling strategy for quantification due to its high accuracy and high throughput, with the CV averaging 8% and the ability to simultaneously quantify 18 samples in a single run. The ZIC-HILIC enrichment method is superior in terms of identification capacity and unbiased detection of N-glycan types. A total of 2,151 GPSMs were identified, representing a 26% improvement compared to the MAX method. We systematically optimized the HCD fragmentation process to improve the identification of intact N-glycopeptides, focusing on glycoform identification and identification capacity. Precise energy adjustment is crucial for the accurate identification of glycans. Our results indicate that glycan dissociation is highly sensitive to collision energy, with higher energies potentially causing excessive fragmentation or reducing detection. The optimal HCD collision energy was set as stepping energies of 25, 35, and 45 for subsequent analyses. By comparing label-free and TMTpro labeling quantitative methods, we found that the TMTpro labeling quantification strategy had higher reproducibility. TMTpro-labeled intact glycopeptides were analyzed using three different software tools, pGlyco 3.0, MS-PyCloud and MSFragger-Glyco, respectively. The number of unique intact glycopeptides identified was 1,992 for pGlyco, 2,444 for MS-PyCloud, and 2,093 for MSFragger-Glyco. Additionally, the TMTpro labeling strategy reduced MS analysis time, enabling more robust analysis of large-scale clinical samples. Using the optimal method, we investigated differences in N-glycosylation between luminal and basal subtypes of breast cancer PDX mouse models, identifying a total of 5,514 unique intact N-glycopeptides. Further analysis suggested potential distinctions in N-glycosylation patterns associated with breast cancer subtypes. Our established optimal intact N-glycopeptide analysis workflow has provided a crucial reference for future glycoproteomic studies on clinical samples, laying the foundation for further clinical applications.

Supplementary Material

Supporting Information
Supplementary Tables

ACKNOWLEDGMENT

This work was supported by National Institutes of Health, National Cancer Institute, The Clinical Proteomic Tumor Analysis Consortium (CPTAC, U24CA271079), the Early Detection Research Network (EDRN, U2CCA271895), and Pancreatic Cancer Detection Consortium (PCDC, U01CA274514).

Footnotes

Notes

The authors declare no competing financial interest

ASSOCIATED CONTENT

Supporting Information

Additional experimental details, including sepPak C18 desalting of digested peptides; TMTpro labeling of peptides; enrichment of glycopeptides by ZIC-HILIC enrichment; enrichment of glycopeptides by MAX enrichment; basic reverse phase fractionation; LC-MS/MS analysis; data analysis. Supplementary tables, including Table S1S3; Supplementary figures, including the identification number of salic acid-containing intact glycopeptides from unlabeled and TMTpro labeled intact glycopeptides; Comparative performance of ZIC-HILIC and MAX Enrichment strategies for direct glycopeptide enrichment; The N-glycans of LLNINPNK.T (LAPM1) from different collision energies for mass spectrometry analysis. Numbers in the X-axis represent glycans containing the numbers of Hex (H), HexNAc (N), A (NeuAc), F (fucose), respectively; The characterization of TMT pro labeling intact N- glycopeptides using different collision energies; The number of glycoproteins, glycosites, and N-Glycans identified from different software; Distribution of glycans, peptide backbones, and proteins in TMTpro-labeled intact glycopeptides analyzed with different software tools.

REFERENCES

  • (1).Gemmer M; Chaillet ML; van Loenhout J; Arenas RC; Vismpas D; Gröllers-Mulderij M; Koh FA; Albanese P; Scheltema RA; Howes SC; Kotecha A; Fedry J; Förster F Visualization of translation and protein biogenesis at the ER membrane. Nature 2023, 614, 160–167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (2).Stadlmann J; Taubenschmid J; Wenzel D; Gattinger A; Durnberger G; Dusberger F; Elling U; Mach L; Mechtler K; Penninger JM Comparative glycoproteomics of stem cells identifies new players in ricin toxicity. Nature 2017, 549, 538–542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (3).Higel F; Seidl A; Sörgel F; Friess W N-glycosylation heterogeneity and the influence on structure, function and pharmacokinetics of monoclonal antibodies and Fc fusion proteins. Eur. J. Pharm. Biopharm. 2016, 100, 94–100. [DOI] [PubMed] [Google Scholar]
  • (4).Gudelja I; Lame G; Pezer M Immunoglobulin G glycosylation in aging and diseases. Cell. Immunol. 2018, 333, 65–79. [DOI] [PubMed] [Google Scholar]
  • (5).Sun S; Shah P; Eshghi ST; Yang W; Trikannad N; Yang S; Chen L; Aiyetan P; Hoti N; Zhang Z; Chan DW; Zhang H Comprehensive analysis of protein glycosylation by solid-phase extraction of N-linked glycans and glycosite-containing peptides. Nat. Biotechnol. 2016, 34, 84–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (6).de Haan N; Yang S; Cipollo J; Wuhrer M Glycomics studies using sialic acid derivatization and mass spectrometry. Nature Reviews Chemistry 2020, 4, 229–242. [DOI] [PubMed] [Google Scholar]
  • (7).Sun ZY; Fu B; Wang GL; Zhang L; Xu RF; Zhang Y; Lu HJ High-throughput site-specific N-glycoproteomics reveals glyco-signatures for liver disease diagnosis. Natl. Sci. Rev. 2023, 10, nwac059. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (8).Sun ZY; Ji GH; Wang GL; Wei L; Zhang Y; Lu HJ One step carboxyl group isotopic labeling for quantitative analysis of intact glycopeptides by mass spectrometry. Chem. Commun. 2021, 57, 4154–4157. [DOI] [PubMed] [Google Scholar]
  • (9).Fang P; Ji YL; Oellerich T; Urlaub H; Pan KT Strategies for Proteome-Wide Quantification of Glycosylation Macro- and Micro-Heterogeneity. Int. J. Mol. Sci. 2022, 23, 1609. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (10).Xiao HP; Chen WX; Smeekens JM; Wu RH An enrichment method based on synergistic and reversible covalent interactions for large-scale analysis of glycoproteins. Nat. Commun. 2018, 9, 1692. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (11).Suttapitugsakul S; Sun FX; Wu RH Recent Advances in Glycoproteomic Analysis by Mass Spectrometry. Anal. Chem. 2020, 92, 267–291. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (12).Dong WB; Chen L; Jia L; Chen ZX; Shen JC; Li PF; Sun SS Maximal performance of intact glycopeptide enrichment using sequential HILIC and MAX columns. Anal. Bioanal. Chem. 2023, 415, 6431–6439. [DOI] [PubMed] [Google Scholar]
  • (13).Yang WM; Shah P; Hu YW; Eshghi ST; Sun SS; Liu Y; Zhang H Comparison of Enrichment Methods for Intact N- and O-Linked Glycopeptides Using Strong Anion Exchange and Hydrophilic Interaction Liquid Chromatography. Anal. Chem. 2017, 89, 11193–11197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (14).Xia CS; Jiao FL; Gao FY; Wang HP; Lv YY; Shen YH; Zhang YJ; Qian XH Two-Dimensional MoS-Based Zwitterionic Hydrophilic Interaction Liquid Chromatography Material for the Specific Enrichment of Glycopeptides. Anal. Chem. 2018, 90, 6651–6659. [DOI] [PubMed] [Google Scholar]
  • (15).Cao L; et al. Proteogenomic characterization of pancreatic ductal adenocarcinoma. Cell 2021, 184, 5031–5052 e5026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (16).Fang P; Ji Y; Silbern I; Doebele C; Ninov M; Lenz C; Oellerich T; Pan KT; Urlaub H A streamlined pipeline for multiplexed quantitative site-specific N-glycoproteomics. Nat. Commun. 2020, 11, 5268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (17).Lee HJ; Cha HJ; Lim JS; Lee SH; Song SY; Kim H; Hancock WS; Yoo JS; Paik YK Abundance-Ratio-Based Semiquantitative Analysis of Site-Specific N-Linked Glycopeptides Present in the Plasma of Hepatocellular Carcinoma Patients. J. Proteome Res. 2014, 13, 2328–2338. [DOI] [PubMed] [Google Scholar]
  • (18).Fang Z; Qin HQ; Mao JW; Wang ZY; Zhang N; Wang Y; Liu LY; Nie YZ; Dong MM; Ye ML Glyco-Decipher enables glycan database-independent peptide matching and in-depth characterization of site-specific N-glycosylation. Nat. Commun. 2022, 13, 1900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (19).Halim A; Westerlind U; Pett C; Schorlemer M; Rüetschi U; Brinkmalm G; Sihlbom C; Lengqvist J; Larson G; Nilsson J Assignment of Saccharide Identities through Analysis of Oxonium Ion Fragmentation Profiles in LC MS/MS of Glycopeptides. J. Proteome Res. 2014, 13, 6024–6032. [DOI] [PubMed] [Google Scholar]
  • (20).Mertins P; Tang LC; Krug K; Clark DJ; Gritsenko MA; Chen LJ; Clauser KR; Clauss TR; Shah P; Gillette MA; et al. Reproducible workflow for multiplexed deep-scale proteome and phosphoproteome analysis of tumor tissues by liquid chromatography-mass spectrometry. Nature Protocols 2018, 13, 1632–1661. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (21).Perez-Riverol Y; Bai JW; Bandla C; García-Seisdedos D; Hewapathirana S; Kamatchinathan S; Kundu DJ; Prakash A; Frericks-Zipper A; Eisenacher M; Walzer M; Wang SB; Brazma A; Vizcaíno JA The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res. 2022, 50, D543–D552. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (22).Li J; Van Vranken JG; Pontano Vaites L; Schweppe DK; Huttlin EL; Etienne C; Nandhikonda P; Viner R; Robitaille AM; Thompson AH; Kuhn K; Pike I; Bomgarden RD; Rogers JC; Gygi SP; Paulo JA TMTpro reagents: a set of isobaric labeling mass tags enables simultaneous proteome-wide measurements across 16 samples. Nat. Methods 2020, 17, 399–404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (23).Qing GY; Yan JY; He XN; Li XL; Liang XM Recent advances in hydrophilic interaction liquid interaction chromatography materials for glycopeptide enrichment and glycan separation. Trac-Trends in Analytical Chemistry 2020, 124,115570. [Google Scholar]
  • (24).Neue K; Mormann M; Peter-Katalinic J; Pohlentz G Elucidation of Glycoprotein Structures by Unspecific Proteolysis and Direct nanoESI Mass Spectrometric Analysis of ZIC-HILIC-Enriched Glycopeptides. J. Proteome Res. 2011, 10, 2248–2260. [DOI] [PubMed] [Google Scholar]
  • (25).Yeh CH; Chen SH; Li DT; Lin HP; Huang HJ; Chang CI; Shih WL; Chern CL; Shi FK; Hsu JL Magnetic bead-based hydrophilic interaction liquid chromatography for glycopeptide enrichments. J. Chromatogr. A 2012, 1224, 70–78. [DOI] [PubMed] [Google Scholar]
  • (26).Hu Y; Shah P; Clark DJ; Ao M; Zhang H Reanalysis of Global Proteomic and Phosphoproteomic Data Identified a Large Number of Glycopeptides. Anal. Chem. 2018, 90, 8065–8071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (27).Lih TM; Cho KC; Schnaubelt M; Hu Y; Zhang H Integrated glycoproteomic characterization of clear cell renal cell carcinoma. Cell Rep 2023, 42, 112409. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (28).Cho KC; Chen LJ; Hu YW; Schnaubelt M; Zhang H Developing Workflow for Simultaneous Analyses of Phosphopeptides and Glycopeptides. ACS Chem. Biol. 2019, 14, 58–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (29).Riley NM; Malaker SA; Driessen MD; Bertozzi CR Optimal Dissociation Methods Differ for N- and O-Glycopeptides. J. Proteome Res. 2020, 19, 3286–3301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (30).Tiemeyer M; Aoki K; Paulson J; Cummings RD; York WS; Karlsson NG; Lisacek F; Packer NH; Campbell MP; Aoki NP; et al. GlyTouCan: an accessible glycan structure repository. Glycobiology 2017, 27, 915–919. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (31).Zeng WF; Cao WQ; Liu MQ; He SM; Yang PY Precise, fast and comprehensive analysis of intact glycopeptides and modified glycans with pGlyco3. Nat. Methods 2021, 18, 1515–1523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (32).Polasky DA; Yu FC; Teo GC; Nesvizhskii AI Fast and comprehensive and glycoproteomics analysis with MSFragger-Glyco. Nat. Methods 2020, 17, 1125–1136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (33).Eshghi ST; Shah P; Yang WM; Li XD; Zhang H GPQuest: A Spectral Library Matching Algorithm for Site-Specific Assignment of Tandem Mass Spectra to Intact N-glycopeptides. Anal. Chem. 2015, 87, 5181–5188. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (34).Hamouda H; Kaup M; Ullah M; Berger M; Sandig V; Tauber R; Blanchard V Rapid Analysis of Cell Surface N-Glycosylation from Living Cells Using Mass Spectrometry. J. Proteome Res. 2014, 13, 6144–6151. [DOI] [PubMed] [Google Scholar]
  • (35).Yu R; Longo J; van Leeuwen JE; Zhang CJ; Branchard E; Elbaz M; Cescon DW; Drake RR; Dennis JW; Penn LZ Mevalonate Pathway Inhibition Slows Breast Cancer Metastasis via Reduced glycosylation Abundance and Branching. Cancer Res. 2021, 81, 2625–2635. [DOI] [PubMed] [Google Scholar]
  • (36).Gudjonsson T; Adriance MC; Sternlicht MD; Petersen OW; Bissell MJ Myoepithelial cells: Their origin and function in breast morphogenesis and neoplasia. J. Mammary Gland Biol. Neoplasia 2005, 10, 261–272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (37).Cimino A; Halushka M; Illei P; Wu X; Sukumar S; Argani P Epithelial cell adhesion molecule (EpCAM) is overexpressed in breast cancer metastases. Breast Cancer Res. Treat. 2010, 123, 701–708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (38).Sethi MK; Kim H; Park CK; Baker MS; Paik YK; Packer NH; Hancock WS; Fanayan S; Thaysen-Andersen M In-depth N-glycome profiling of paired colorectal cancer and non-tumorigenic tissues reveals cancer-, stage- and EGFR-specific protein N-glycosylation. Glycobiology 2015, 25, 1064–1078. [DOI] [PubMed] [Google Scholar]
  • (39).Liu X; Gao J; Sun Y; Zhang D; Liu T; Yan Q; Yang X Mutation of N-linked glycosylation in EpCAM affected cell adhesion in breast cancer cells. Biol. Chem. 2017, 398, 1119–1126 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting Information
Supplementary Tables

Data Availability Statement

The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD054874 21.

RESOURCES