Abstract
Given the significant role of glycosylation in modulating protein structure and activity, glycoproteomics is gaining increased interest from the broad scientific community. Large-scale site-resolved N-glycoproteomics relies on automated MS2-based searches, but candidate glycopeptide assignments can remain ambiguous when isomeric or isobaric glycan structures, adducts, chemical modifications, in-source fragments, or incomplete MS2 evidence support more than one plausible interpretation. Here, we systematically categorize common challenges and misassignments in glycoproteomics and present a post-search validation workflow using Skyline software to identify and correct these. The workflow matches search-engine-derived candidate assignments to LC–MS/MS evidence for correct precursor monoisotope assignment, retention time behavior, and glycosite context. The workflow is demonstrated with Byonic-derived glycopeptide candidate lists and converts automated search results into curated, verifiable, site-resolved N-glycopeptide features for downstream quantification and reporting. We applied the workflow to data from 52 human serum samples, and reviewed 3,071 candidate N-glycopeptide IDs. From these, 1,722 MS2 candidate IDs were refuted as inconsistent with chromatographic and/or precursor-level evidence. Curation added 320 glycopeptide features, comprising 152 MS1-supported composition-level assignments and 168 additional LC-resolved isomer features, yielding a final curated feature set of 1,436 N-glycopeptides across the serum N-glycoproteome. Together, these results show that reviewing the raw LC-MS/MS data associated with search-engine results improves both the accuracy and comprehensiveness of detectable N-glycopeptides, supporting more transparent and reliable reporting. The curated dataset provides a resource for future method development, benchmarking and machine–learning efforts directed at automated glycopeptide validation.
Keywords: N-Glycoproteomics, Retention time monitoring, LC-MS/MS, Human serum, Glycopeptide validation


1. Introduction
Glycoproteomics is a rapidly evolving field at the intersection of proteomics and glycobiology, focused on elucidating how glycans modulate protein function, stability, and interactions within biological systems. Understanding glycosylation patterns is essential because they influence a wide range of biological processes, such as immune response, cell signaling, and disease progression, making glycans indispensable for elucidating mechanisms underlying various pathologies. Mass spectrometry (MS) technologies have advanced significantly over the past decades, enabling deeper insights into the structure and function of glycoproteins through improved sensitivity, resolution, and throughput. − Higher-resolution MS has enhanced our ability to detect fine differences in glycan compositions, while advanced tandem MS (MS2) strategies provide valuable fragment information for glycan and peptide identification. These methodological innovations have established glycoproteomics as a powerful tool for characterizing complex biological samples and discovering disease-associated biomarkers. ,
Despite these technological advances, the inherent complexity and heterogeneity of glycan structures pose considerable challenges in glycoproteomic data processing. Glycosylation can vary widely in terms of monosaccharide composition, branching, and glycosidic linkages, creating numerous potential isomeric structures and complicating both glycan identification and the interpretation of results. − Additional measurement artifacts can also arise from sample preparation and adduct formation. , These challenges have motivated diverse computational strategies for automated N-glycopeptide candidate generation based on MS2 data. One of the first and still most common approaches are database-searches, where Byonic provides flexible glycopeptide searching with user-defined protein and glycan databases, pGlyco2 introduced separate glycan-, peptide-, and glycopeptide-level quality control, and pGlyco3 uses a glycan-first, site-aware assignment strategy. Other candidate-generation strategies include open or glycan-mass-offset searching with MSFragger-Glyco, glycan-database-independent peptide matching with Glyco-Decipher, de novo glycan sequencing with PEAKS GlycanFinder, and structure-resolved interpretation of N-glycans with StrucGP.
Beyond candidate generation based on MS2, the reliability of site-resolved N-glycopeptide assignments depends on how automated candidate matches are evaluated against the complementary evidence. Evidence-aware refinement can draw on precursor-level MS1 features (accurate mass, isotopic pattern matching), detailed fragment ion support, chromatographic behavior, and glycosite-level context. These are crucial to differentiate isomeric and isobaric structures, glycopeptide modifications, and possible measurement artifacts. Selected workflows include, assign, and/or report such possibilities, although their implementation and downstream disambiguation differ. ,,, GlycReSoft illustrates this by linking MS2 identifications to MS1 features and incorporating fragmentation, relative-retention-time, and site-specific glycome-network models to support and, in defined cases, revise ambiguous N-glycopeptide assignments. Once accepted, candidate N-glycopeptide assignments become the molecular features quantified in downstream glycoproteomics. Multiple glycopeptide quantification tools have been developed so far using both precursor and fragment ion information. Among these is the highly specialized SugarQuant-Glycobinder pipeline, which performs TMT-based quantification at the MS3 level. The flexible pGlycoQuant can import search glycopeptide results from multiple sources (Byonic, MSFragger-Glyco, pGlyco) and perform match between runs label-free quantification using MS1 and MS2 data. LacyTools performs high-throughput, targeted integration and quality control based on MS1 data. Complementary post-identification tools then help convert quantified feature tables into interpretable data sets: GlycoDash focuses on visually assisted curation, normalization, and reporting of quantified glycopeptide data, whereas StrucGAP extends the analysis to structural feature extraction, functional annotation, visualization, and regulatory or pathway-level data mining. ,
Community evaluations and comparative software studies show that glycopeptide search results can vary substantially across informatics workflows or users and that incorrect or incomplete candidate assignments may remain in reported results and literature when not evaluated against orthogonal evidence. , This underscores the need for reproducible post-search quality control that makes the evidence behind reported assignments verifiable before quantitative reporting and downstream interpretation.
The present work addresses this need through a Skyline-based workflow for post-search validation and curation of site-resolved N-glycoproteomics data. Rather than introducing a competing search engine, the workflow organizes automated MS2-based candidate glycopeptide assignments together with LC-MS/MS evidence for precursor isotope agreement, retention time behavior, fragment-ion support, glycosite context, and sample-level quantitative consistency. Although demonstrated here with Byonic-derived candidate lists, the validation logic is not search-engine specific in principle, provided that their assignments and supporting spectra can be represented in Skyline-compatible formats. We first summarize common challenges in large-scale N-glycoproteomics data analysis, then describe the Skyline-based integration of automated MS2-based searches with MS1-level validation and quantification. Finally, we demonstrate the workflow on 52 human serum samples, confirming 1,436 N-glycopeptides and providing a curated resource for future multi-tool comparisons and benchmarking.
2. Common Challenges in Glycoproteomic Data Analysis
2.1. Isobaric and Isomeric Monosaccharide Combinations
A major challenge in MS2-based glycopeptide identification arises from the fact that there are multiple combinations of monosaccharides that can converge to (partly) overlapping masses within their isotopic envelopes (Table ). ,,− MS2 spectra may not discriminate between isobaric glycans that share a common peptide backbone, due to the large overlap of glycan B- and Y-ions. This results in false positive annotations of certain glycoforms while the correct ones are missed (false negatives). This issue becomes especially problematic for overlapping species and lower intensity ions, for which the isotopic envelope cannot be accurately annotated. A common example is the substitution of two fucoses (F2; 292.1158 Da) for one sialic acid (S1; 291.0954 Da), generating a 1.0204 Da offset. The overlap in the isotopic envelopes can cause MS2-based software tools to, for example, label a glycopeptide as carrying a glycan with composition H5N4F2S1 instead of H5N4S2 (H: hexose, N: N-acetylhexosamine, F: fucose, S: N-acetyl neuraminic acid) and vice versa (Supporting Information, Figure S1). A second representative example is the incorrect assignment of H8N7F1 instead of H6N5S3 (H2N2F1 = S3 + 3.0361 Da; Supporting Information, Figure S2), illustrating that even larger neutral-mass differences can still result in overlapping isotopic envelopes and lead to misidentifications.
1. Common Glycan Misassignments in LC–MS/MS Glycoproteomics .

Presence is specific; it is valid only for NeuAc-free compositions.
Absence is specific; presence does not exclude in multi-NeuGc glycans.
Misassignments originating from isobaric/isomeric monosaccharide combinations, ammonia adducts, carbamidomethylation, and oxidation are summarized. Each ambiguous assignment is paired with an alternative composition, monoisotopic mass, exact mass difference, RT shift, and a diagnostic fragment, supporting correction. Ambiguous and alternative compositions are interchangeable. Symbols follow SNFG system. H, hexose; N, N-acetylhexosamine; F, fucose; S, N-acetylneuraminic acid; G, N-glycolylneuraminic acid; KDN, 2-keto-3-deoxy-nononic acid; GlcA, Glucuronic acid. “A” (in black box) = Ammonia adduct (+17.02655 Da); “C” (in black box) = CAM (+57.0215 Da); “O” (in black box) = Oxidation (+15.9949 Da).
An important analytical measure for correct glycan composition assignment is the elution behavior of glycopeptides in reversed-phase LC. Glycopeptide retention is primarily driven by the peptide sequence, with minor contributions from glycoform variations. A key exception is the sialic acid content. Under standard low-pH reversed-phase conditions (e.g., C18 with 0.1% formic acid; pH∼ 2.5), sialic acid residues are largely protonated and introduce an increased retention as compared to their non-sialylated counterparts, possibly through changes in hydrophobicity, conformation, and/or secondary interactions with the LC system. In practice, and for the LC configuration used here, glycoforms with the same number of sialic acids co-elute within narrow retention time (RT) windows, and each additional sialic acid shifts to later RT (Figure ). Monitoring this behavior, in combination with high-resolution MS1-based isotopic envelope matching, enables a rapid detection of misidentifications. Changes in pH, stationary phase (e.g., porous graphitized carbon) and ion-pairing reagents (e.g., TFA) will alter the elution behavior of glycopeptides, for which the RT-based rules described here do not directly apply.
1.

N-glycopeptides follow a clustered retention time behavior dependent on their sialic acid content. Extracted ion chromatograms (EICs) of N-glycopeptide glycoforms from the haptoglobin peptide VVLHPNYSQVDIGLIK in human serum, analyzed by reversed-phase LC-MS/MS, demonstrating how the number of sialic acids in the glycan composition affects retention time. Each trace corresponds to a specific glycoform composition (e.g., H6N5F1S3, where H denotes hexose, N denotes N-acetylhexosamine, F denotes fucose, and S denotes N-acetylneuraminic acid). Notably, glycoforms such as H6N5S3 and H7N6S4 exhibit double peaks, indicating chromatographically resolved isomers that can be quantified individually. Glycan symbols follow SNFG formatting and structures are inferred from literature, supported by MS2 data. Exact linkages are not determined.
Also, isomeric glycans pose a challenge in glycoproteomics, sharing the same monosaccharide composition yet with differences in linkages and branching positions. Although MS2 fragmentation can provide partial insights into glycan structure based on diagnostic monosaccharide ions, or specific ion ratios indicative of monosaccharide identity or linkages, , isomer resolution remains incomplete in many standard bottom-up workflows. Importantly, isomeric glycans are often (partially) separated by LC, e.g., H6N5S3 resolves into two peaks (H6N5S3_a, H6N5S3_b) with distinct RTs (Figure ), highlighting the need for careful identification and individual quantification. Automated workflows frequently overlook these chromatographically separated isomers, leading to an incomplete characterization.
2.2. Ammonia Adducts
Ammonia adducts are a common artifact in bottom-up proteomics, leading to a +17.026 Da shift in mass. In glycopeptide analyses, such adducts confound data interpretation because the extra mass mimics the presence of specific monosaccharides (Table ). For example, the mass offset between F1+ammonia and H1 is 1.0318 Da, while the difference between H1F1 and S1+ammonia is only 0.0112 Da (Supporting Information, Figure S3). These ammonia-adduct-related ambiguities may also combine with other recurring mass overlaps, further complicating glycopeptide assignment (Supporting Information, Figure S4). To overcome ammonia-adduct-induced misidentifications, correct monoisotope selection and RT monitoring are again critical.
2.3. Carbamidomethylation
Carbamidomethylation (CAM) is a chemical modification commonly introduced in proteomics during sample preparation to prevent (re)formation of cysteine bridges, resulting in an addition of +57.0215 Da. Unintended side reactions can cause overalkylation, affecting not only cysteine but also lysine, histidine, glutamic acid, methionine, and N-termini of peptides, leading to incorrect glycan assignments by mimicking natural glycan compositions (Table ). On the one hand, this results in similar situations as described for the glycan compositions and ammonia adducts, where S1+CAM can be misidentified as N1F1, despite a mass difference of 1.0204 Da (Supporting Information, Figure S5). On the other hand, it introduces potential isomers of affected glycopeptides, in particular, when F1+CAM is misidentified as N1 (Supporting Information, Figure S6) or when H1+CAM is mistaken for N1+peptide oxidation. In these cases, evaluation of the presence of the Y1 ion (peptide+N1) in the MS2 data is crucial for accurate assignment, as it confirms the (un)modified peptide mass (Supporting Information, Figure S7).
2.4. In-Source Fragmentation
In-source fragmentation is a common phenomenon in mass spectrometric analysis of peptides with labile modifications, such as glycans. In particular, glycans containing sialic acids and oligomannose-type N-glycans are prone to in-source degradation, resulting in the detection of a truncated glycan variant at the same RT apex as its precursor. , A representative example is the apparent detection of a monosialylated glycoform after in-source loss of one terminal NeuAc from a disialylated precursor; for instance, an H5N4S1-like ion may be observed at the same RT apex as that of the corresponding H5N4S2 glycopeptide. On the basis of MS2 identification and precursor mass evidence alone, these in-source fragments can easily be misidentified as unique glycoforms. Only critical evaluation of the elution behavior can tell in-source fragments apart from their natural counterparts.
2.5. Missing GlycoformsFalse Negatives
Also, the occurrence of missing glycoforms remains a significant challenge in glycoproteomics. First, a glycan may be missing from the input database, meaning that even if it is present in the sample, the search engine will not identify it. Second, some glycopeptides are more susceptible to collision-induced dissociation (CID) fragmentation, particularly those with multiple sialic acids, leading to overfragmented MS2 spectra with low intensity or missing glycopeptide-specific Y-ions that prevent confident glycoform identification. Finally, instrumental limitations in fragmentation time or sensitivity may cause certain glycopeptides to remain undetected. These challenges become more apparent in large-scale glycoproteomics studies, where MS2 coverage might be inconsistent between runs, while glycopeptide presence is often shared between samples. Furthermore, the more complex a glycoproteomics sample is, the higher the chance that low-abundance glycoforms are missed entirely.
MS1-based approaches incorporating RT parameters have proven effective for glycopeptide identification, addressing limitations in MS2-centric workflows. Notably, GlycReSoft was the first tool to integrate RT patterns as a validation feature within its MS2 search pipeline. Complementing this, Skyline allows the quantification of all detected glycoforms across data sets, independent of MS2 identification in individual runs. Importantly, it facilitates MS1-level evaluation (accurate mass, isotope envelope, RT) of glycoforms not directly observed in MS2 spectra but inferred from biosynthetic logic and the identified glycan repertoire.
3. Post-Search Validation and Curation of N-Glycoproteomics Data
Here we provide the steps essential to overcome the limitations of MS2-based glycopeptide identification software packages for the processing of large LC-MS/MS N-glycoproteomics data sets, encompassing multiple samples. While full implementation currently relies on manual data curation, automated solutions are emerging. , Our approach combines automated MS2-based glycopeptide identification tools with critical MS1-level validation using chromatographic profiles, facilitated by the open-source software Skyline. ,
Given its widespread use in the field, this workflow is based on output from Byonic, , following data acquisition of human serum samples on a Thermo Orbitrap instrument and a standard glycoproteomics sample preparation protocol (Supporting Information, Sections S1–S3). Although different instruments and search engines may result in different levels of artifacts and pitfalls, the principles outlined here are broadly applicable, possibly with minor adaptations, across platforms and identification tools (Figure ).
2.

Glycopeptide data analysis workflow. Raw mass spectrometry (MS) data acquired from glycopeptide digest mixtures are first pre-processed to correct precursor isotope assignments. Initial protein identification is performed using a broad protein database (e.g., the human proteome). Subsequently, glycopeptide identification is carried out against a reduced protein database derived from the initial search, along with a curated list of glycosylation modifications. High-confidence peptide-spectrum matches (PSMs) should exhibit both glycan-specific diagnostic ions (e.g., oxonium and Y-ions) and peptide backbone fragment ions (e.g., b- and y-ions). All peptide identifications are exported in standardized formats (e.g., .mzIdentML), including MS/MS spectrum metadata, peptide assignments, and confidence scores. The standardized output and raw MS data are then imported into Skyline to generate a spectral library and extracted ion chromatograms for all proposed glycopeptides. Final data validation is based on both fragment level (i.e., detection of glycan- and peptide-specific ions) and precursor level data metrics (i.e., mass error, isotope correction, isotopic dot product, and glycoform retention time). Detailed post-search validation and curation decisions are summarized in Supporting Information, Figure S8.
3.1. Monoisotopic Ion Selection by Preprocessing in Monocle
Accurate identification of the precursor’s monoisotopic ion is essential for determining the exact mass of glycopeptides and avoiding false positives (Table ). This becomes challenging for low-abundance species in the higher m/z range, where low intensity monoisotopic ions may be obscured by noise. Several glycoproteomics tools, including Byonic, offer monoisotopic correction. ,,,, However, pre-processing with the free software Monocle yields a higher rate of correct N-glycopeptide monoisotope assignments than those reported by acquisition software or Byonic (Supporting Information, Figure S9 and Tables S1 and S2). Monocle enhances signal-to-noise ratios by summing and averaging multiple precursor scans around the fragmentation event. It supports Thermo raw MS files or universal formats (.mzXML, .mzML), and exports corrected .mzXML files compatible with Byonic and other search engines.
3.2. Fine-Tuning Protein and Glycan Database for MS2-level Searches
Standard glycan databases provided for MS2-based glycopeptide identification often do not suffice because 1) not all N-glycans are included, resulting in false negatives, and 2) unnecessarily expanding the search space results in the higher risk of false positive assignments. Furthermore, the diverse set of glycans that can be attached at any queried glycosylation site challenges automated identification, since each glycan increases the search space exponentially. Thus, both the protein and glycan databases should be tuned for a specific experiment, based on the literature background and/or experimental information. Both the proteome and the glycome , components of the human serum glycoproteome have been extensively characterized. For the glycopeptide identification in this study, a subset of 4608 human proteins validated by MS and reported in PeptideAtlas for human plasma were considered and combined with a database of 67 N-glycans previously reported in human plasma, , 159 N-glycans reported in breast cancer tumors, and 3 additional N-glycans found during data validation (see sections below).
For biological samples that have not been characterized as extensively, we recommend 1) the protein database should be based on an initial identification search across the full proteome, excluding glycan modifications. This also allows monitoring of digestion efficiency and the co-occurrence of (artificial) protein modifications, providing the input for the glycopeptide-specific search settings. Ideally, this step is performed using a full (glyco)proteome sample, e.g., the flow-through in case a specific glycopeptide enrichment is performed, or a deglycosylated fraction of the unenriched samples. 2) To obtain an accurate and reduced glycan database, ideally a parallel released glycan (glycomics) analysis is performed from the same matrix as used for glycoproteomics analysis. Further extension of the glycan database may be based on the availability of literature reports on similar samples, and/or initial oxonium ion screening for specific glycan substructures.
Upon definition of the protein and glycan databases, the isotope-corrected MS files (3.1) can be searched for glycopeptide identifications based on the MS2 data. For the currently used Byonic search parameters, see Supporting Information, Section S4.2.
3.3. Generating Spectral Libraries and Extracted Ion Chromatograms in Skyline
To validate the output of the MS2-based search engine, transfer identifications across measurements, and eventually quantify all identified glycopeptides across large sample sets, the identified glycopeptides should be imported into Skyline via spectral libraries, after which extracted ion chromatograms (EICs) are computed for each data file. To generate spectral libraries in Skyline, search engine result files should minimally contain the glycopeptides identified, their corresponding peptide spectrum matches (PSMs) and RTs, and the score assigned to these PSMs, in a universal format such as mzIdentML (a.k.a. mzID). Proposed glycopeptides are filtered based on a user-defined certainty threshold (e.g., 5% false discovery rate at the PSM level) after which Skyline generates EICs based on calculated m/z and isotopic distributions for the relevant charge states per glycopeptide. Some glycan modifications, reported in unimod, are stored by default in Skyline. However, custom glycan modifications can also be added. It is recommended to add new PTMs by describing their elemental composition, so isotopologue distributions can be calculated; e.g. H5N4S2 = C84H136N6O61. Glycan-specific fragment losses can be defined to calculate glycopeptide Y-ions; e.g. H5N4S2 - C76H123N5O56 (H5N3S2) to specify the Y1-ion. Additionally, structure-specific glycan B-ions (oxonium ions) can be defined, e.g. C20H34N1O14 representing the H1N1F1 oxonium ion indicative for antenna fucosylation. Although all modifications can be added manually through Skyline’s user interface, this becomes tedious for hundreds of glycan modifications. Therefore, we provide an R script (i.e. RNotebook_SkylineGlycoproteomicsTemplate.Rmd) and glycan database to streamline Skyline template modification (Supporting Information of the Panorama Public submission: https://panoramaweb.org/Navigating_HumanSerum_NGlycoproteomics.url).
EICs are generated in Skyline for multiple isotopologues calculated per precursor ion. We recommend using predicted relative intensities to capture the most intense isotopologues using an intensity threshold of 20% from the base peak. EIC generation is guided by the glycopeptide PSMs using an extraction window around the time these MS/MS events were triggered (e.g. ±5 min; highly dependent on LC settings; Figure A). Restricting the extraction windows for generating EICs facilitates the selection of the correct chromatographic peaks and reduces file sizes. However, incorrect glycopeptide identifications can result in incorrect extraction windows. Therefore, we recommend using extraction windows that include not only the RT cluster corresponding to the annotated glycopeptide but also the clusters of the same peptide portion carrying glycans with one sialic acid more or less.
3.

Glycopeptide data inspection and validation. (A) Aligned extracted ion chromatograms (EICs) of the isotopologues from the triply protonated ion of the haptoglobin (HPT) glycopeptide VVLHP N [H6N5F1S3]YSQVDIGLIK. Sample #5 lacked a peptide-spectrum match (PSM) for this glycopeptide, but its EIC was inferred using retention time information from other samples (e.g., Sample #45). (B) Examples of incorrect glycoform assignments for the glycopeptide VVLHPNYSQVDIGLIK carrying the glycan H5N4S2. These misassignments result from (1) incorrect isotope assignment, (2) incorrect isotope assignment combined with off-target carbamidomethyl (CAM) alkylation, and (3) formation of an ammonium (NH3) adduct. (C) Mass spectra at the chromatographic apex of the incorrectly assigned glycoforms. Retention time and precursor-level information aid in diagnosing misassignments that are not apparent from fragment ion data alone. All cases show elution times inconsistent with the expected glycoform elution cluster. This is especially critical in case (3), where it is the sole discriminant of the incorrect assignment (Supporting Information, Figure S3).
Importantly, EICs are also generated in sample files for which no PSMs were defined, based on MS/MS events in other samples (Figure A). This ensures uniform quantification of glycopeptides across large sample sets. While the automated, RT-restricted EIC generation is an important step in large-scale glycoproteomics, validation of the glycopeptide assignments is crucial for accurate glycosite profiling as it allows the detection of wrong assignments and the inclusion of LC-separated isomers and missing glycoforms.
3.4. Glycopeptide Validation and Quantification
Once the PSM-guided EICs have been generated, data validation should be performed by checking both precursor and fragment level data (Supporting Information, Figure S8). Glycopeptide inspection starts by evaluating the EICs that were created based on their guiding PSMs (Figure A). The steps described below are often run iteratively with incremental improvements:
-
1)
EIC peak RT matching between samples. Consistent peak integration requires using the same retention-time window across samples, which Skyline facilitates through visualization of RT plots (Figure A). When elution patterns shift between samples, indexed retention time libraries , (based either on endogenous (glyco)peptides or on an added indexed retention time (iRT) mix prior to measurement) facilitate alignment in the RT domain. All 27 iRT values used here were empirically indexed and not predicted based on peptide and glycan composition, since this is still notoriously challenging for glycopeptides. They included highly abundant endogenous human plasma glycopeptides and were used to align the retention time of all other targets across 52 LC-MS/MS measurements (Supporting Information, Table S5). After fitting the glycopeptides to a linear regression, the majority of validated glycopeptides were within ±2.5 min of the predicted retention time. After recalibration, this iRT library can be used for other LC-MS/MS platforms and data sets.
-
2
) Comparing recorded isotopologue ratios to theoretical values derived from the glycopeptide’s elemental composition. This comparison, described in Skyline as the isotopic dot product (idotp) ideally approaches 1, while values between 0.8 and 0.9 may indicate integration interferences or incorrect glycopeptide assignment and should be critically assessed. Additionally, precursor ion ppm errors of the integrated EIC peaks as indicated on top of the peaks should be within the acceptable boundaries, i.e. within 95% confidence intervals of the measuring LC-MS/MS platform. Peak selections that do not match these MS1 quality parameters should be removed from the data, or corrected to matching annotations (Table ).
-
3)
Diagnosing wrong isotope envelope assignments by both idotp and precursor scan visualization. Precursor scans can be visualized by selecting any MS1 scan within the integration window (Figure B and C). If the idotp is below 0.9, and MS1 inspection reveals there is no co-eluting species that induces integration interference and, consequently, deviates the mass error and idotp, then the glycopeptide assignment is incorrect. Depending on the direction of the offset and elution rules, an incorrect assignment may represent another glycoform of the same site and can be corrected (Table ). A reason for the misidentification might be that the glycan composition was missing from the initial glycan database (see Section 3.2) and should in that case be flagged for inclusion in a reanalysis with the MS2-level search engine (see Section 3.4, point 8).
-
4)
PSM Inspection to Confirm Glycopeptide Identifications. Importantly, the PSM acquisition time has to be aligned with the precursor-ion integration window, ideally close to the peak apex. An ideal glycopeptide match allows for full peptide sequence and glycan composition coverage. However, the fragmentation quality is highly dependent on peptide sequence, precursor charge state, and glycan composition and hardly ever results in full coverage. , Therefore, for large-scale studies, we emphasize the importance of a subset of fragments to confirm at least one glycoform per peptide portion in at least one of the samples. Reliable glycopeptide PSMs should contain, as a minimum, fragment ions that allow confirmation of peptide and glycan elemental composition. First, if at least one of the B-ions representing N1 and H1N1 is observed, fragmentation of a glycopeptide ion can be assumed and validation can proceed. For glycan compositions proposing sialic acids, the presence of S1 and S1-water oxonium ions is a strict requirement. Next, the detection of Y-ions Y0, Y1, and Y2 confirms the mass of the proposed peptide sequence. The remaining mass should match the proposed glycan mass. Next, peptide backbone fragments should provide sequence information for at least one glycoform per glycosylation site. Because of the lability of glycans, b- and y-ions are generally low-abundance, limiting peptide sequence coverage in low-intensity glycoforms. The evidence of one glycoform per peptide portion can be used to support the detection of other glycoforms that show chromatographic behavior consistent with their glycan composition (see Section 3.4, point 6). In a multi-sample study, identifications can be transferred between runs based on RT, accurate precursor mass, and matching idotp. N-Glycosylation sites are inferred based on the sequence motif N-X-S/T where X is any amino acid that is not proline. If a glycopeptide sequence carries two or more of these motifs, then the glycan composition(s) assigned to one or more sites should be determined from electron-mediated dissociation (EThcD in our case) that generates additional site-localizing c- and z-ions carrying the full or partial glycan. Further glycan structure refinement is possible with the detection of specific B- and Y ions. For example, for N-glycopeptides containing fucoses, a Y1+F1 fragment should be present in the case of N-glycan core fucosylation and/or the B-ion H1N1F1 in the case of antenna fucosylation. The Y-ion H1N3 confirms the presence of bisecting GlcNAc. Importantly, annotation of diagnostic ions should be on the monoisotopic peaks with at least one more isotopologue co-detected.
-
5)
Annotating glycopeptide isomers. Distinct glycoform isomers present at the same peptide portion are often observed as multiple EIC peaks from the same precursor m/z, and positively match the MS1 data validation described above. Routine MS2-based glycopeptide identification does not necessarily distinguish chromatographically separated isomers, particularly when MS1 elution profiles are not considered. In Skyline, glycopeptide isomers may be annotated by making a copy of the glycopeptide and altering the name of the glycan modification to represent multiple isomers. For each peptide–glycan composition, suffixes _a, _b, _c, and so forth are assigned in increasing elution order across the RT-aligned dataset, e.g., H6N5S3_a for the first-eluting feature and H6N5S3_b for the second (Figure ). These labels are retained across runs and are not renumbered when individual features are absent in a specific run (NA downstream). When an established isomer pair is only partially resolved in an occasional run, it is approximately apportioned between the fixed suffix-specific feature by splitting the signal at the saddle point. If reproducible isomer separation is not established across the dataset, the composition is retained as a single unsuffixed feature. The suffixes denote chromatographic elution order only and do not imply linkage, branching, or other structural identity.
-
6)
RT cluster evaluation. Different glycoforms on the same peptide portion elute in tight (1-2 min) clusters, under the condition that they carry the same number of sialic acids, where a higher degree of sialylation results in a significant jump in RT (Figure ). If the EIC of a glycoform with two sialic acids is present in the cluster representing one sialic acid, the assignment is incorrect (Figure and Figure B). Furthermore, elution profile analysis is essential for identifying glycopeptide candidates arising from in-source fragments and adducts. These share RT apexes with their precursors and cluster incorrectly when e.g. missing a sialic acid. Skyline facilitates RT cluster inspection by displaying overlay EICs of multiple glycoforms with annotations and by offering retention time comparisons in a “live report” via the Document Grid, which updates dynamically during peak curation for interactive data validation.
-
7)
Expanding the glycan repertoire. Steps 1 to 6 are essential to obtain high quality identifications based on the MS2 search results. Next, tools like GlyConnect Compozitor can predict missing glycan compositions using biosynthetic logic. For this, the curated set of glycans per glycosylation site can be uploaded, after which so called “missing nodes” will be provided connecting all glycoforms with one monosaccharide increment. Inferred novel glycopeptides may be manually introduced in Skyline by defining new glycan modifications (see Section 3.3), re-importing the raw MS data to generate the missing EICs, and validation using MS1-level criteria. Although PSMs are initially absent for these glycoforms, their identity can be supported by accurate mass, matching isotope patterns and consistent elution behavior with related glycoforms.
-
8)
Perform a new MS 2 -level analysis based on all insights into curated glycoforms. This is an iterative process that can be repeated for complete or subsets of the data.
-
9)
Export of quantitative glycoproteomics data. Once the integration windows of all accepted glycoforms have been curated, the integrated peak areas of the precursor ions are exported in tabulated format for downstream analysis. In addition to peak area, other peak characteristicssuch as retention time, iRT offset, idotp, precursor mass errorare used for data filtering, either with free open-source software or custom processing scripts. An example of such processing is available in the Supporting Information section of the following Panorama Public submission: https://panoramaweb.org/Navigating_HumanSerum_NGlycoproteomics.url. Subsequently, all precursor areas corresponding to the same glycopeptide are aggregated to yield the total glycopeptide area. Each glycoform area is then normalized to the summed area of all glycoforms sharing the same peptide sequence. The resulting ratios (relative abundances of individual glycoform per glycosite) are used to compare across experimental conditions or patient groups. Additionally, N-glycan traits are calculated per site, representing e.g., proportions of oligomannose, complex, or hybrid glycans, the degree of sialylation, the degree of fucosylation, and the presence of bisecting GlcNAc, providing relevant insights into the regulation of glycosylation pathways and the presence of specific glycan epitopes. ,
4. Results from a Human Serum N-Glycoproteomics Dataset
Serum samples were obtained from Masaryk Memorial Cancer Institute from a cohort comprising 32 female patients with breast cancer and 20 female healthy controls, with a mean age of 49.6 years (SD 11.7). Details on samples, ethics and sample preparation and mass spectrometry acquisition procedures are provided in the Supporting Information, Sections S1–S3.
4.1. N-Glycopeptide ID Validation Resulted in a Highly Curated Dataset
We applied our thorough N-glycoproteomics data processing to a human serum dataset comprising samples from 52 individuals. In total, 1,756,567 MS/MS spectra were acquired, for which 249,322 PSMs were found in the MS2-based search using a 5% FDR (Figure A). Prior to further validation, a PSM selection was performed by excluding non-glycosylated and decoy matches, as well as PSMs containing >1 assigned N-glycan site. Furthermore, we restricted our validation to 58 high-confidence glycoprotein identifications (Supporting Information, Section S4.1). These filters retained 186,203 PSMs, which were collapsed into 3,071 candidate N-glycopeptide IDs for Skyline-based inspection. Our validation method identified 1,722 MS2 candidate IDs as assignments inconsistent with chromatographic and/or precursor m/z-level evidence (Figure B, Table ). Furthermore, 93 oxidized-peptide forms and 162 alternative digestion products (fully tryptic versus missed cleavage) were identified as redundant peptide-forms, rather than distinct glycopeptides. Beyond the MS2-supported identifications, glycan-repertoire expansion added 152 MS1-supported composition-level assignments, and 168 LC-resolved isomer features, of which 164 were associated with MS2-level evidence and 4 with MS1-level evidence. Together, these additions increased the curated dataset by 320 features, yielding a final set of 1,436 glycopeptide features spanning 148 N-glycosylation sites on 58 glycoproteins (average number of features per site: 9.7; range: 1–52). The total list of MS2-based glycopeptide IDs contains 3,093 entries (Figure B), including the 3,071 candidates from the Monocle-based search, 19 glycopeptide entries supported only by the initial search workflow (Supporting Information, Section S4.1) and three entries retained after manual protein reassignment from IGHA2 to IGHA1 (Supporting Information, Section S4).
4.

N-glycoproteomics accuracy gains and N-glycopeptide structural diversity across 52 human serum samples. (A) Funnel summarizing acquisition and curation from 1,756,567 MS/MS spectra to 1,436 glycopeptide IDs across 148 glycosites on 58 glycoproteins. (B) Stacked bar chart summarizing validation outcomes for glycopeptide IDs in the curated Skyline dataset. Category assignments are based on the final Skyline Peptide Note classification used for curation (Supporting Information, Section S4). MS2 misannotated IDs: misannotated identifications assigned by the MS2 search engine but manually refuted; MS2 redundant oxidation: additional accepted MS2 IDs corresponding to oxidized peptide forms; MS2 redundant digestion: alternative digestion forms (fully tryptic or missed cleavage) not retained as the dominant (most intense) glycopeptide; MS2 confirmed: MS2 IDs manually confirmed; MS1 search: glycopeptides identified via MS1-based searching within RT intervals; MS2-based extra isomers: manually detected additional isomers derived from MS2 IDs; MS1-based extra isomers: manually detected additional isomers derived from MS1 IDs. (C) Glycopeptide diversity across the curated dataset was binned into five structural traits, and the relative frequency is displayed in pie charts. (a) Glycan type: complex-type versus hybrid/oligomannose. (b) Sialylation: S0–S5 correspond to the number of sialic acids per glycan. (c) Antennarity: A1–A4 denote mono- to tetra-antennary complex glycans, while “other” groups hybrid and oligomannose species. (d) Bisection: B1 indicates the presence of a bisecting GlcNAc, B0 its absence. (e) Fucosylation: F0–F5 reflect the total count of fucose residues.
Characterization of the curated set of 1,436 N-glycopeptides revealed substantial structural heterogeneity across the human serum N-glycoproteome, consistent with the heterogeneity reported in blood glycomics studies (Figure C). Structural trait frequencies showed that complex-type glycans dominated the dataset (96%), with hybrid and oligomannose glycans each contributing 2%. Within the complex-type subset, diantennary (A2; 48%), triantennary (A3; 31%), and tetraantennary (A4; 14%) structures predominated. Consistent with the highly sialylated nature of serum glycoproteins, the majority of glycans carried one to three NeuAc residues (S1; 26%, S2; 36%, S3; 21%), while asialo (S0; 12%) and tetrasialylated (S4; 6%) forms occurred less frequently. Core and/or antennary fucosylation was observed in 45% of all glycoforms, with monofucosylated species (F1; 35%) predominating over multifucosylated variants (F2; 7%, F3; 1%, F4; 0.6%, F5; 0.3%). A bisecting GlcNAc was detected in 9% of the structures.
The 152 MS1-supported glycopeptide additions following glycan repertoire expansion (see Section 3.4, point7) closely resembled the glycan-class distribution of the complete curated data set: 145 (95%) were complex-type, four were oligomannose, and three were hybrid (Supporting Information, Figure S10 and Table S4). The additions spanned 70 glycosites on 40 proteins but were concentrated at a subset of sites, with the largest expansions observed for IGHA2 N92 (10), ORM1 N103 (9), ATRN N416 and IGHA2 N205 (7 each), IGHG3 N227 (6), and IGHG4 N177 and SERPING1 N238 (5 each).
It is important to notice that a substantial fraction of the original MS2-based glycopeptide candidates were refuted during subsequent validation (Figure B), leading us to examine how refuted assignments were distributed across the serum glycoproteome and which orthogonal criteria most frequently disqualify them (Figure D). The fraction of misannotated glycopeptides ranged between 0% and 85% both across proteins (Figure A) and individual glycosylation sites (Figure B). Further evaluation of two representative proteins with multiple glycosylation sites (kininogen-1; P01042 and alpha-1-acid glycoprotein 1; P02763) partitioned site-resolved outcomes into MS2-confirmed identifications, misannotated MS2 candidates, redundant peptide-form variants (oxidation/digestion), MS1-supported additions, and chromatographically resolved isomers detected during curation (Figure C). Within each protein, refuted candidates were concentrated at a subset of glycosites rather than uniformly distributed across sites, ranging from 15% to 55% per site in kininogen-1 and 45% to 70% per site in alpha-1-acid glycoprotein 1. Across both proteins, the largest fraction of the misannotated putative MS2-based IDs could be assigned to isobaric/isomeric monosaccharide combinations (42%; Table ; Figure D), followed by ammonia adduct artifacts (23%), peptide modification artifacts (overalkylation, 11%; oxidation, 2%), while the remaining 22% were categorized as unexplained. Among the 175 refuted candidate assignments examined for these two proteins, refutation was supported by both RT and MS1 isotope-envelope evidence in 117 cases (67%), by RT evidence alone in 35 cases (20%), by MS1 evidence alone in 21 cases (12%), and by MS2 diagnostic fragments alone in 2 cases (1%) (Figure D). When both RT and MS1 evidence supported refutation, the two criteria were considered jointly with no fixed hierarchy between them.
5.

Evidence-based filtering of MS2 assignments quantifies the misannotation burden across serum glycoproteins and glycosylation sites. (A) Protein-level misannotation plotted as fraction of the total glycopeptide MS2 candidates per protein (refuted MS2 candidates/all MS2 candidates). Bubble size scales with the number of unique N-glycosites detected per protein, and bubble color denotes the fraction of confirmed N-glycosites (≥1 MS2 ID per site). (B) Glycosite-level misannotation plotted as fraction of the total glycopeptide MS2 candidates per site (refuted MS2 candidates/all MS2 candidates). Squares denote the four highlighted glycosites from P02763 alpha-1-acid glycoprotein 1 and triangles denote the four highlighted glycosites from P01042 kininogen-1. (C) Site-resolved validation outcomes for N-glycosites of alpha-1-acid glycoprotein 1 (P02763) and kininogen-1 (P01042). MS2 misannotated IDs: misannotated identifications assigned by the MS2 search engine but manually refuted; MS2 redun. digestion: alternative digestion forms (fully tryptic or missed cleavage) not retained as the dominant (most intense) glycopeptide; MS2 redun. peptide oxidation: additional accepted MS2 IDs corresponding to oxidized peptide forms; MS2 confirmed: MS2 IDs manually confirmed; MS1 search: glycopeptides identified via MS1-based searching within RT intervals; MS2 extra isomers: manually detected additional chromatographic isomers derived from MS2 IDs; MS1 extra isomers: manually detected additional chromatographic isomers derived from MS1 IDs. (D) Evidence types supporting the refutation of MS2-derived misannotated IDs for the two representative proteins, stratified by error class. RT+MS1 denotes cases in which both retention-time behavior and MS1 isotopic-envelope inspection supported refutation. RT, MS1, and MS2 denote cases in which the respective evidence type alone supported refutation. Error classes are defined as follows: Overalkylation; Ammonia adducts; Oxidation, oxidized peptide variants misannotated as distinct glycopeptide IDs; isobaric/isomeric combinations, ambiguous monosaccharide compositions yielding overlapping precursor masses; and unexplained, isolated low support IDs that could not be confidently assigned to a specific error class, but can be refuted by RT behavior, MS1 isotope-envelope consistency, or MS2 diagnostic fragments. The two proteins exemplify misannotations from all categories in different proportions, illustrating how error profiles may differ among glycosites.
4.2. Site-Resolved Glycan Profiles of 58 Human Serum N-Glycoproteins across Individuals
Following validation of MS2 search outputs (Figures and ), we resolved the curated serum N-glycoproteome at the level of individual glycosylation sites to complement the global structural-trait summary (Figure C) and to summarize cohort-wide glycoform distributions (Figure ; Supporting Information, Table S3). Importantly, while a total of 137 N-glycans were detected across the entire dataset, glycoform richness varied markedly across sites and proteins. For example, while IgM (P01871) site N46 and N209 are characterized by the presence of complex diantennary glycans with a high degree of bisection, site N441 on the same protein exclusively carries oligomannose type glycans. Glycans carrying bisecting GlcNAcs are largely absent for acute phase plasma glycoproteins such as haptoglobin (P00738) and alpha-1-acid glycoprotein 1 (P02763), which are in turn characterized by multi-antennary glycans carrying high degrees of sialylation. Importantly, most identified glycopeptides could be quantified in all 52 serum samples, although low abundant glycoforms were excluded from quantification in a fraction of the samples, due to the limit of detection (Figure ). Also, relatively high abundant glycoforms were not always detected across all samples, for example the fucosylated glycans on haptoglobin (P00738) N211, indicating that these could be related to specific (patho)physiological conditions of the sample donors. Overall, our curated serum N-glycoproteomics dataset shows a wide variety of glycoforms of which subsets are specific for individual proteins and glycosylation sites. Site-specific relative quantification of glycan features allows comparison of the N-glycoproteome across individuals.
6.

Site-resolved bubble heatmap of glycoform abundance and prevalence across 52 human sera. Rows list curated N-glycosylation sites (protein | glycosite position) annotated with the number of glycoforms detected per site (n). Columns denote glycan compositions (e.g., H7N6F1S4, where H, N, F, and S denote: hexose, N-acetylhexosamine, fucose, and N-acetylneuraminic acid, respectively) with each composition assigned to structural traits indicated by the headers above the matrix (glycan type: complex (C); oligomannose (OM); hybrid (Hy); sialylation (S0–S4); fucosylation (F0–F2); antennarity (A2–A4 for complex glycans; A0 for non-complex); bisection (B0/B1)). Of note, some of the compositions represent multiple chromatographically resolved isomers that were separately quantified, see Supporting Information, Table S3. Bubble color encodes the median within-site relative abundance (%) of each glycoform, computed across samples in which that glycoform was detected, and bubble size encodes the fraction of 52 samples in which the glycoform was detected.
5. Discussion
In large-scale site-resolved N-glycoproteomics, LC-MS/MS data interpretation is complicated by glycan isomers, isobaric species, adducts, chemical modifications, and in-source fragments that generate ambiguous spectral assignments. Consequently, automatically generated candidate lists often retain false-positive glycopeptide assignments that are difficult to recognize from the fragment-ion evidence alone. To address this issue, we comprehensively outline common misassignments in glycoproteomics and present a Skyline-based post-search validation workflow for systematic evaluation of candidate assignments using precursor isotope-envelope evidence, retention-time behavior, and glycosite context, thereby converting search-engine-derived candidate lists into a curated, verifiable, site-resolved N-glycopeptide feature set.
In a case-study dataset analyzing the serum N-glycoproteome of 52 individuals, we observed these assignment ambiguities and artifacts at scale. To provide a practical framework for their interpretation, Table consolidates the recurring pitfall classes observed in our dataset and reported in the literature, together with orthogonal cues that can help disambiguate them. Using our Skyline-based workflow to curate the candidate assignments, we refuted 1,722 of 3,071 MS2 glycopeptide candidates (56%), highlighting that MS2 scoring alone is insufficient in complex glycopeptide search spaces. Importantly, refutation rates were highly glycosite-dependent (0–85%), showing that global FDR thresholds do not translate into uniform confidence at the site level. This argues for glycosite-level confidence control that integrates retention-time constraints and MS1 precursor verification. For two representative proteins spanning eight glycosites, retention-time consistency contributed to 87% of refutations, while MS1 isotopic-envelope evidence supported 79%, indicating that orthogonal LC–MS constraints explained most refutations in this focused subset without relying exclusively on manual MS2 reinterpretation. Beyond removing misannotated candidates, curation added 152 MS1-supported composition-level glycopeptides and 168 additional LC-resolved isomer features, not identified in the MS2-based search (false negatives). Together, this filtering and expansion yielded high-quality glycopeptide identification and quantification across 52 samples, capturing serum glycosylation features consistent with prior blood N-glycome reports, , but now resolved at the protein level. These results position post-search validation as an assignment-level curation layer that determines which candidate glycopeptides are retained, refuted, corrected, split into LC-resolved isomers, or expanded before downstream quantification and biological interpretation.
The need for post-MS2-search validation is recognized in the field and has been recently implemented in glycoproteomics informatics tools. Yet, no consensus has been reached on the scope of this validation and the level at which it is implemented. E.g., GlycReSoft represents an advanced automated evidence-aware refinement strategy, integrating fragmentation modeling, MS1 feature mapping, relative-retention-time-based assignment revision, and site-specific glycome-network smoothing to support, expand, or revise N-glycopeptide assignments. Protein Prospector, on the other hand, provides FDR-controlled Batch-Tag/Search Compare results to initiate MS-Filter searches for additional glycoforms of confidently identified peptide backbones using Y-ion evidence and glycan scoring. Search Compare also reports mean retention time and retention-time spread for identified glycoforms, making chromatographic consistency available for user inspection. StrucGAP is primarily a downstream data-mining platform, but its preprocessing layer is relevant here as a complementary implementation of postidentification quality control. It supports standardized inputs from multiple search engines and combines data-quality filters with glycan plausibility annotation, including GlyTouCan mapping, KEGG-based biosynthetic annotation, and rule-based flags for improbable structures. Together, these examples show that state-of-the-art glycoproteomics informatics increasingly incorporates orthogonal evidence into automated scoring, filtering, or postidentification quality control.
Although the present case study was performed using Byonic output from Thermo Orbitrap Fusion data, the post-search validation strategy is not restricted to this search engine or instrument data. To build a Skyline spectral library and to guide EIC generation, the glycopeptide search output should contain sufficient PSM level metadata. These include the peptide sequence, glycan composition, glycosylation site assignment where applicable, precursor charge state and m/z, retention time, MS/MS spectrum metadata, and a search engine confidence metric such as a score, posterior error probability, or q-value. Standardized formats such as mzIdentML are directly applicable, although tabular or tool-specific outputs may also be used after an in-house conversion.
Accordingly, outputs from, e.g., pGlyco3, MSFragger-Glyco, and Glyco-Decipher can be evaluated with our workflow. The 56% refutation fraction observed here is specific to the present data set, Byonic/Monocle workflow, search space, confidence thresholds, instrument platform, and curation criteria; it should not be interpreted as an intrinsic error rate of Byonic or of glycopeptide search engines in general. Although our final workflow already used Monocle-based monoisotope correction and a 5 ppm precursor tolerance, differences among search engines in precursor-isotope handling, scoring, and quality-control procedures may change the number and types of candidates requiring downstream review, particularly for cases resulting from monoisotope misassignments. , Further refinement of precursor mass tolerance, confidence thresholds, and matrix-appropriate databases could reduce unsupported candidates but may also exclude low-abundance true identifications and cannot resolve ambiguities absent from the search score, including isobaric compositions, adducts, in-source fragments, and RT-inconsistent assignments. We did not perform a matched multi-engine or parameter sweep reanalysis of our data and can therefore not estimate the direction and magnitude of any change in the refutation fraction. Importantly, our recommendations regarding protein and glycan database tuning apply most directly to database search workflows, where the search space is defined by the researcher. Open, de novo, and database-independent approaches may reduce or alter the reliance on predefined glycan lists, but ambiguity in candidate interpretation persists, making post-search evidence review equally essential for these strategies.
Our workflow should be viewed as a post-FDR, assignment-level evidence-review layer rather than a replacement for statistical FDR control. Target–decoy or score-based procedures estimate the expected false-discovery proportion among accepted matches within the candidate space and scoring model used by the search workflow. While important as a first step, such estimates often fail to capture RT/MS1-based artifacts, including in-source fragmentation or isobaric overlaps, making post-hoc orthogonal validation irreplaceable. In Skyline, this complementary review is explicit and verifiable, and can be applied to an entire dataset or specific proteins of interest.
In practice, this assignment-level review is especially relevant for 1) matching glycoforms across large sample sets, 2) validating glycoproteins, glycosites, and glycoforms that appear to discriminate between biologically relevant conditions or originate from industry-produced therapeutic proteins, 3) preserving and separately quantifying LC-resolved glycopeptide isomers, and 4) achieving broad glycoform coverage at individual sites.
Preserving LC-resolved glycopeptide isomers is relevant not only for analytical coverage but also for biological interpretation. Distinct peaks sharing the same peptide sequence and glycan composition may correspond to linkage or branching glycan isomers that arise from alternative glycan biosynthesis pathways. Changes in the relative abundance of these features may reveal site-specific glycan remodeling that is not apparent from the composition-level measurements alone. In our data, suffixes such as _a and _b denote the elution order only and should not be interpreted as linkage or branching assignments. Identification of the glycan structure can be supported by, e.g., MS2 fragments, exoglycosidase treatment, and external reference standards. Such standards would also be valuable for benchmarking our validation approach but were not included here due to the limited availability of well-defined glycopeptide materials spanning multiple peptides and glycoforms.
A limitation of our work is that the manual data validation is time-consuming and requires specialized expertise. Assignment-level review of the 3,071 MS2-derived candidate IDs was estimated at approximately 2 min per candidate, corresponding to approximately 102 person-hours. MS1 repertoire expansion involved screening 4,824 additional protein–peptide–glycan hypotheses, as defined in Section 3.4, at approximately 1 min per hypothesis, corresponding to approximately 80 person-hours. Final run-by-run inspection and, where necessary, correction of peak-integration boundaries for the 1,436 retained features across 52 LC–MS/MS runs was estimated at approximately 5 s per feature–run pair, corresponding to approximately 104 person-hours. Together, these principal hands-on validation and integration stages required approximately 290 person-hours for one experienced operator. This estimate excludes one-time Skyline library and template construction and elapsed computer-processing time for data import, chromatogram extraction, and chromatogram recalculation.
Yet, the effort is front-loaded for a particular biological matrix and an LC method. Once a curated glycopeptide library and empirical retention time behavior have been established, these resources can be reused for subsequent cohorts acquired under comparable conditions. Our deposited Skyline documents provided in https://panoramaweb.org/Navigating_HumanSerum_NGlycoproteomics.url contain all validated glycopeptides indexed against 27 abundant human serum N-glycopeptides and can serve as templates for new serum N-glycoproteomics studies. The empirical iRT values support targeted EIC extraction and consistent peak integration across the runs. Immediate opportunities for automation include rule-based flagging of incorrect monoisotope assignments, low isotopic dot product, systematic precursor-mass offsets, retention-time-cluster outliers, coeluting in-source fragments, recurrent adduct- or overalkylation-related mass shifts, and chromatographically resolved isomers. Such rules could prioritize ambiguous cases for expert inspection rather than attempt to replace manual review entirely. Automated cross-run glycopeptide identification and quantification and integration of partially resolved glycopeptide-isomer peaks nevertheless remain challenging and represent a specific target for future workflow development.
In conclusion, our work highlights and quantifies recurrent sources of glycopeptide misannotation in MS2-based candidate assignments and provides a Skyline-based strategy to recognize and address them through post-search validation. The integrated workflow improves the accuracy and comprehensiveness of glycopeptide identification and quantification, resulting in a well-curated dataset of serum N-glycopeptides. Beyond providing protein- and site-specific N-glycosylation profiles in human serum, this dataset offers a resource for future method development, search-engine benchmarking, and machine-learning efforts aimed at automated glycopeptide validation, thereby supporting more transparent and reliable site-resolved N-glycopeptide reporting.
Supplementary Material
Acknowledgments
This work was supported by the project SALVAGE (OP JAC; reg. no. CZ.02.01.01/00/22_008/0004644), co-funded by the European Union and by the State Budget of the Czech Republic. This work was further supported by MH CZ–DRO (MMCI, 00209805) and by the project BBMRI.CZ (LM2023033). J.C.R.E. and N.d.H. were for this work supported by the European Union, projects HORIZON-EIC-101161509 (GLUCOTYPES) and HORIZON-EIC-101161281 (Bugs4Urate). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them. We would like to thank Manfred Wuhrer for providing his valuable input on the early manuscript.
Data of this study is publicly available as a Panorama Public submission https://panoramaweb.org/Navigating_HumanSerum_NGlycoproteomics.url with the associated ProteomeXchange ID PXD077279. The Supporting Information section of the Panorama repository contains the curated Skyline documents, associated raw LC–MS/MS data, the RNotebook_SkylineGlycoproteomicsTemplate.Rmd notebook for Skyline template generation together with the files required for its execution, and the PostSearchValidation_Glycoproteomics_Reproducibility_Workflow.Rmd notebook for processing Skyline exports and reproducing the reported analyses together with the files required for its execution.
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jacsau.6c00875.
Supporting materials and methods, additional experimental details on human serum samples and ethics approval, sample preparation, mass spectrometry data acquisition, MS2-based glycoproteomics database searches, and glycopeptide validation in Skyline, representative annotated MS2 spectra illustrating common glycan misassignments and their manual correction in LC–MS/MS glycoproteomics, decision-oriented flowchart summarizing the post-search validation and curation workflow for glycopeptide assignments, comparison of monoisotopic peak assignment strategies based on overlap with validated glycopeptide candidates, structural-trait distribution of the 152 MS1-supported composition-level additions, and summary tables of glycopeptide identification acceptance rates and monoisotopic peak assignment offsets across isotope assignment strategies (PDF)
Site-resolved curated N-glycopeptide table across 52 human serum samples (XLSX)
Feature-level annotations for the 320 features added during curation, comprising 152 MS1-supported composition-level assignments and 168 additional LC-resolved isomer features (XLSX)
Twenty-seven endogenous human plasma N-glycopeptides used as indexed retention-time standards for cross-run alignment (XLSX)
Adam P. Urminsky and Juan C. Rojas E. contributed equally to this work. The manuscript was written through contributions of all authors. All authors have given approval to the final version of the manuscript.
The authors declare no competing financial interest.
References
- Bagdonaite I., Malaker S. A., Polasky D. A., Riley N. M., Schjoldager K., Vakhrushev S. Y., Halim A., Aoki-Kinoshita K. F., Nesvizhskii A. I., Bertozzi C. R., Wandall H. H., Parker B. L., Thaysen-Andersen M., Scott N. E.. Glycoproteomics. Nature Reviews Methods Primers. 2022;2:48. doi: 10.1038/s43586-022-00128-4. [DOI] [Google Scholar]
- Reily C., Stewart T. J., Renfrow M. B., Novak J.. Glycosylation in health and disease. Nature Reviews Nephrology. 2019;15:346–366. doi: 10.1038/s41581-019-0129-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Suttapitugsakul S., Sun F., Wu R.. Recent Advances in Glycoproteomic Analysis by Mass Spectrometry. Analytical Chemistry. 2020;92:267–291. doi: 10.1021/acs.analchem.9b04651. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Werner T., Fahrner M., Schilling O.. Advancements in mass spectrometry-based proteomics: a new era in pathology research and diagnostics. Pathologie. 2024;45:56–62. doi: 10.1007/s00292-024-01390-x. [DOI] [PubMed] [Google Scholar]
- Wenk D., Zuo C., Kislinger T., Sepiashvili L.. Recent developments in mass-spectrometry-based targeted proteomics of clinical cancer biomarkers. Clinical Proteomics. 2024;21:6. doi: 10.1186/s12014-024-09452-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hu H., Khatri K., Klein J., Leymarie N., Zaia J.. A review of methods for interpretation of glycopeptide tandem mass spectral data. Glycoconjugate Journal. 2016;33:285–296. doi: 10.1007/s10719-015-9633-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pinho S. S., Reis C. A.. Glycosylation in cancer: mechanisms and clinical implications. Nature Reviews Cancer. 2015;15:540–555. doi: 10.1038/nrc3982. [DOI] [PubMed] [Google Scholar]
- He K., Baniasad M., Kwon H., Caval T., Xu G., Lebrilla C., Hommes D. W., Bertozzi C.. Decoding the glycoproteome: a new frontier for biomarker discovery in cancer. Journal of Hematology & Oncology. 2024;17:12. doi: 10.1186/s13045-024-01532-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang L., Luo S., Zhang B.. Glycan analysis of therapeutic glycoproteins. mAbs. 2016;8:205–215. doi: 10.1080/19420862.2015.1117719. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Polasky D. A., Geiszler D. J., Yu F., Nesvizhskii A. I.. Multiattribute Glycan Identification and FDR Control for Glycoproteomics. Molecular & Cellular Proteomics. 2022;21:100205. doi: 10.1016/j.mcpro.2022.100205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lu H., Zhang Y., Yang P.. Advancements in mass spectrometry-based glycoproteomics and glycomics. National Science Review. 2016;3:345–364. doi: 10.1093/nsr/nww019. [DOI] [Google Scholar]
- Darula Z., Medzihradszky K. F.. Carbamidomethylation Side Reactions May Lead to Glycan Misassignments in Glycopeptide Analysis. Anal. Chem. 2015;87:6297–6302. doi: 10.1021/acs.analchem.5b01121. [DOI] [PubMed] [Google Scholar]
- Klein J. A., Zaia J.. Relative Retention Time Estimation Improves N-Glycopeptide Identifications by LC–MS/MS. J. Proteome Res. 2020;19:2113–2121. doi: 10.1021/acs.jproteome.0c00051. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zeng W.-F., Cao W.-Q., Liu M.-Q., He S.-M., Yang P.-Y.. Precise, fast and comprehensive analysis of intact glycopeptides and modified glycans with pGlyco3. Nat. Methods. 2021;18:1515–1523. doi: 10.1038/s41592-021-01306-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bern M., Kil Y. J., Becker C.. Byonic: Advanced Peptide and Protein Identification Software. Current Protocols in Bioinformatics. 2012;40:13.20.1. doi: 10.1002/0471250953.bi1320s40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu M.-Q.. et al. pGlyco 2.0 Enables Precision N-Glycoproteomics with Comprehensive Quality Control and One-Step Mass Spectrometry for Intact Glycopeptide Identification. Nat. Commun. 2017;8:438. doi: 10.1038/s41467-017-00535-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Polasky D. A., Yu F., Teo G. C., Nesvizhskii A. I.. Fast and comprehensive N- and O-glycoproteomics analysis with MSFragger-Glyco. Nat. Methods. 2020;17:1125–1132. doi: 10.1038/s41592-020-0967-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fang Z., Qin H., Mao J., Wang Z., Zhang N., Wang Y., Liu L., Nie Y., Dong M., Ye M.. Glyco-Decipher enables glycan database-independent peptide matching and in-depth characterization of site-specific N-glycosylation. Nat. Commun. 2022;13:1900. doi: 10.1038/s41467-022-29530-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sun W., Zhang Q., Zhang X., Tran N. H., Ziaur Rahman M., Chen Z., Peng C., Ma J., Li M., Xin L., Shan B.. Glycopeptide database search and de novo sequencing with PEAKS GlycanFinder enable highly sensitive glycoproteomics. Nat. Commun. 2023;14:4046. doi: 10.1038/s41467-023-39699-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shen J.. et al. StrucGP: de novo structural sequencing of site-specific N-glycan on glycoproteins using a modularization strategy. Nat. Methods. 2021;18:921–929. doi: 10.1038/s41592-021-01209-0. [DOI] [PubMed] [Google Scholar]
- Chalkley R. J., Baker P. R.. Improving the Depth and Reliability of Glycopeptide Identification Using Protein Prospector. Molecular & Cellular Proteomics. 2025;24:100903. doi: 10.1016/j.mcpro.2025.100903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Klein J. A., Carvalho L., Zaia J.. Expanding N-glycopeptide identifications by modeling fragmentation, elution, and glycome connectivity. Nat. Commun. 2024;15:6168. doi: 10.1038/s41467-024-50338-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fang P., Ji Y., Silbern I., Doebele C., Ninov M., Lenz C., Oellerich T., Pan K.-T., Urlaub H.. A Streamlined Pipeline for Multiplexed Quantitative Site-Specific N-Glycoproteomics. Nat. Commun. 2020;11:5268. doi: 10.1038/s41467-020-19052-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kong S., Gong P., Zeng W.-F., Jiang B., Hou X., Zhang Y., Zhao H., Liu M., Yan G., Zhou X., Qiao X., Wu M., Yang P., Liu C., Cao W.. pGlycoQuant with a Deep Residual Network for Quantitative Glycoproteomics at Intact Glycopeptide Level. Nat. Commun. 2022;13:7539. doi: 10.1038/s41467-022-35172-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jansen B. C., Falck D., de Haan N., Hipgrave Ederveen A. L., Razdorov G., Lauc G., Wuhrer M.. LaCyTools: A Targeted Liquid Chromatography–Mass Spectrometry Data Processing Package for Relative Quantitation of Glycopeptides. J. Proteome Res. 2016;15:2198–2210. doi: 10.1021/acs.jproteome.6b00171. [DOI] [PubMed] [Google Scholar]
- Pongracz T., Gijze S., Hipgrave Ederveen A. L., Derks R. J. E., Falck D.. GlycoDash: automated, visually assisted curation of glycoproteomics datasets for large sample numbers. Anal. Bioanal. Chem. 2025;417:2003–2014. doi: 10.1007/s00216-025-05794-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang M., Wu Y., Zhang Z., Xu Y., Lei T., Wang X., Jin Z., Hou K., Cai Y., Sun S.. StrucGAP: a modular, streamlined and traceable data mining platform for structural and site-specific glycoproteomics. Nat. Commun. 2026;17:2579. doi: 10.1038/s41467-026-70560-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kawahara R.. et al. Community evaluation of glycoproteomics informatics solutions reveals high-performance search strategies for serum glycopeptide analysis. Nat. Methods. 2021;18:1304–1316. doi: 10.1038/s41592-021-01309-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hogan R. A., Pepi L. E., Riley N. M., Chalkley R. J.. Comparative analysis of glycoproteomic software using a tailored glycan database. Anal. Bioanal. Chem. 2025;417:1985–2001. doi: 10.1007/s00216-025-05780-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- MacLean B., Tomazela D. M., Shulman N., Chambers M., Finney G. L., Frewen B., Kern R., Tabb D. L., Liebler D. C., MacCoss M. J.. Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics. 2010;26:966–968. doi: 10.1093/bioinformatics/btq054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shu Q., Li M., Shu L., An Z., Wang J., Lv H., Yang M., Cai T., Hu T., Fu Y., Yang F.. Large-scale Identification of N-linked Intact Glycopeptides in Human Serum using HILIC Enrichment and Spectral Library Search. Molecular & Cellular Proteomics. 2020;19:672–689. doi: 10.1074/mcp.RA119.001791. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu S.-W., Pu T.-H., Viner R., Khoo K.-H.. Novel LC-MS2 Product Dependent Parallel Data Acquisition Function and Data Analysis Workflow for Sequencing and Identification of Intact Glycopeptides. Analytical Chemistry. 2014;86:5478–5486. doi: 10.1021/ac500945m. [DOI] [PubMed] [Google Scholar]
- Zuniga-Banuelos F. J., Hoffmann M., Reichl U., Rapp E.. New Avenues for Human Blood Plasma Biomarker Discovery via Improved In-Depth Analysis of the Low-Abundant N-Glycoproteome. Engineering. 2026;57:23–42. doi: 10.1016/j.eng.2024.11.039. [DOI] [Google Scholar]
- Varki A.. et al. Symbol Nomenclature for Graphical Representations of Glycans. Glycobiology. 2015;25:1323–1324. doi: 10.1093/glycob/cwv091. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Palmisano G., Larsen M. R., Packer N. H., Thaysen-Andersen M.. LC-MS based detection of glycoprotein sialylation – part II: Structural analysis. RSC Adv. 2013;3:22706–22726. doi: 10.1039/c3ra42969e. [DOI] [Google Scholar]
- Lettow M., Greis K., Mucha E., Lambeth T. R., Yaman M., Kontodimas V., Manz C., Hoffmann W., Meijer G., Julian R. R., von Helden G., Marianski M., Pagel K.. Decoding the Fucose Migration Product during Mass-Spectrometric Analysis of Blood Group Epitopes. Angew. Chem., Int. Ed. 2023;62:e202302883. doi: 10.1002/anie.202302883. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maliepaard J. C., Damen J. M. A., Boons G.-J. P., Reiding K. R.. Glycoproteomics-Compatible MS/MS-Based Quantification of Glycopeptide Isomers. Analytical Chemistry. 2023;95:9605–9614. doi: 10.1021/acs.analchem.3c01319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Halim A., Westerlind U., Pett C., Schorlemer M., Rüetschi U., Brinkmalm G., Sihlbom C., Lengqvist J., Larson G., Nilsson J.. Assignment of Saccharide Identities through Analysis of Oxonium Ion Fragmentation Profiles in LC–MS/MS of Glycopeptides. J. Proteome Res. 2014;13:6024–6032. doi: 10.1021/pr500898r. [DOI] [PubMed] [Google Scholar]
- Kruve A., Kaupmees K.. Adduct Formation in ESI/MS by Mobile Phase Additives. Journal of the American Society for Mass Spectrometry. 2017;28:887–894. doi: 10.1007/s13361-017-1626-y. [DOI] [PubMed] [Google Scholar]
- Müller T., Winter D.. Systematic Evaluation of Protein Reduction and Alkylation Reveals Massive Unspecific Side Effects by Iodine-containing Reagents. Molecular & Cellular Proteomics. 2017;16:1173–1187. doi: 10.1074/mcp.M116.064048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Medzihradszky K. F., Kaasik K., Chalkley R. J.. Characterizing sialic acid variants at the glycopeptide level. Analytical Chemistry. 2015;87:3064–3071. doi: 10.1021/ac504725r. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dong Q., Yan X., Liang Y., Stein S. E.. In-Depth Characterization and Spectral Library Building of Glycopeptides in the Tryptic Digest of a Monoclonal Antibody Using 1D and 2D LC–MS/MS. J. Proteome Res. 2016;15:1472–1486. doi: 10.1021/acs.jproteome.5b01046. [DOI] [PubMed] [Google Scholar]
- Desaire H.. Glycopeptide analysis, recent developments and applications. Molecular & Cellular Proteomics. 2013;12:893–901. doi: 10.1074/mcp.R112.026567. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Choo M. S., Wan C., Rudd P. M., Nguyen-Khuong T.. GlycopeptideGraphMS: Improved Glycopeptide Detection and Identification by Exploiting Graph Theoretical Patterns in Mass and Retention Time. Analytical Chemistry. 2019;91:7236–7244. doi: 10.1021/acs.analchem.9b00594. [DOI] [PubMed] [Google Scholar]
- Alocci D., Mariethoz J., Gastaldello A., Gasteiger E., Karlsson N. G., Kolarich D., Packer N. H., Lisacek F.. GlyConnect: Glycoproteomics Goes Visual, Interactive, and Analytical. J. Proteome Res. 2019;18:664–677. doi: 10.1021/acs.jproteome.8b00766. [DOI] [PubMed] [Google Scholar]
- Pino L. K., Searle B. C., Bollinger J. G., Nunn B., MacLean B. X., MacCoss M. J.. The Skyline ecosystem: Informatics for quantitative mass spectrometry proteomics. Mass Spectrom. Rev. 2020;39:229–244. doi: 10.1002/mas.21540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rad R., Li J., Mintseris J., O’Connell J., Gygi S. P., Schweppe D. K.. Improved Monoisotopic Mass Estimation for Deeper Proteome Coverage. J. Proteome Res. 2021;20:591–598. doi: 10.1021/acs.jproteome.0c00563. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Deutsch E. W., Omenn G. S., Sun Z., Maes M., Pernemalm M., Palaniappan K. K., Letunica N., Vandenbrouck Y., Brun V., Tao S.-c., Yu X., Geyer P. E., Ignjatovic V., Moritz R. L., Schwenk J. M.. Advances and Utility of the Human Plasma Proteome. J. Proteome Res. 2021;20:5241–5263. doi: 10.1021/acs.jproteome.1c00657. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clerc F., Reiding K. R., Jansen B. C., Kammeijer G. S., Bondt A., Wuhrer M.. Human Plasma Protein N-Glycosylation. Glycoconjugate J. 2016;33:309–343. doi: 10.1007/s10719-015-9626-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jager S., Zeller M., Pashkova A., Schulte D., Damoc E., Reiding K. R., Makarov A. A., Heck A. J. R.. In-depth plasma N-glycoproteome profiling using narrow-window data-independent acquisition on the Orbitrap Astral mass spectrometer. Nat. Commun. 2025;16:2497. doi: 10.1038/s41467-025-57916-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Geyer P. E.. et al. The Circulating ProteomeTechnological Developments, Current Challenges, and Future Trends. J. Proteome Res. 2024;23:5279–5295. doi: 10.1021/acs.jproteome.4c00586. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lageveen-Kammeijer G. S. M., de Haan N., Mohaupt P., Wagt S., Filius M., Nouta J., Falck D., Wuhrer M.. Highly sensitive CE-ESI-MS analysis of N-glycans from complex biological samples. Nat. Commun. 2019;10:2137. doi: 10.1038/s41467-019-09910-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benesova I., Nenutil R., Urminsky A., Lattova E., Uhrik L., Grell P., Kokas F. Z., Halamkova J., Zdrahal Z., Vojtesek B., Novotny M. V., Hernychova L.. N-glycan profiling of tissue samples to aid breast cancer subtyping. Sci. Rep. 2024;14:320. doi: 10.1038/s41598-023-51021-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chau, T. H. ; Chernykh, A. ; Ugonotti, J. ; Parker, B. L. ; Kawahara, R. ; Thaysen-Andersen, M. In Serum/Plasma Proteomics: Methods and Protocols; Greening, D. W. , Simpson, R. J. , Eds.; Humana: 2023; Methods in Molecular Biology, Vol. 2628; pp 235–263. [DOI] [PubMed] [Google Scholar]
- Robin T., Mariethoz J., Lisacek F.. Examining and Fine-tuning the Selection of Glycan Compositions with GlyConnect Compozitor. Molecular & Cellular Proteomics. 2020;19:1602–1618. doi: 10.1074/mcp.RA120.002041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- White M. E. H., Sinn L. R., Jones D. M., Flynn H. R., Messner C. B.. et al. Oxonium ion scanning mass spectrometry for large-scale plasma glycoproteomics. Nature Biomedical Engineering. 2024;8:233–247. doi: 10.1038/s41551-023-01067-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jones A. R.. et al. The mzIdentML Data Standard for Mass Spectrometry-Based Proteomics Results. Molecular & Cellular Proteomics. 2012;11:M111.014381-1. doi: 10.1074/mcp.M111.014381. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Creasy D. M., Cottrell J. S.. Unimod: Protein modifications for mass spectrometry. PROTEOMICS. 2004;4:1534–1536. doi: 10.1002/pmic.200300744. [DOI] [PubMed] [Google Scholar]
- Parker S. J., Rost H., Rosenberger G., Collins B. C., Malmström L., Amodei D., Venkatraman V., Raedschelders K., Van Eyk J. E., Aebersold R.. Identification of a Set of Conserved Eukaryotic Internal Retention Time Standards for Data-independent Acquisition Mass Spectrometry. Mol. Cell. Proteomics. 2015;14:2800–2813. doi: 10.1074/mcp.O114.042267. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Escher C., Reiter L., MacLean B., Ossola R., Herzog F., Chilton J., MacCoss M. J., Rinner O.. Using iRT, a Normalized Retention Time for More Targeted Measurement of Peptides. Proteomics. 2012;12:1111–1121. doi: 10.1002/pmic.201100463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mukherjee S., Jankevics A., Busch F., Lubeck M., Zou Y., Kruppa G., Heck A. J. R., Scheltema R. A., Reiding K. R.. Oxonium Ion–Guided Optimization of Ion Mobility–Assisted Glycoproteomics on the timsTOF Pro. Molecular & Cellular Proteomics. 2023;22:100486. doi: 10.1016/j.mcpro.2022.100486. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hinneburg H., Stavenhagen K., Schweiger-Hufnagel U., Pengelley S., Jabs W., Seeberger P. H., Silva D. V., Wuhrer M., Kolarich D.. The Art of Destruction: Optimizing Collision Energies in Quadrupole-Time of Flight (Q-TOF) Instruments for Glycopeptide-Based Glycoproteomics. Journal of the American Society for Mass Spectrometry. 2016;27:507–519. doi: 10.1007/s13361-015-1308-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hevér H., Nagy K., Xue A., Sugár S., Komka K., Vékey K., Drahos L., Révész Á.. Diversity Matters: Optimal Collision Energies for Tandem Mass Spectrometric Analysis of a Large Set of N-Glycopeptides. J. Proteome Res. 2022;21:2743–2753. doi: 10.1021/acs.jproteome.2c00519. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bladergroen M. R., Reiding K. R., Hipgrave Ederveen A. L., Vreeker G. C. M., Clerc F., Holst S., Bondt A., Wuhrer M., van der Burgt Y. E. M.. Automation of High-Throughput Mass Spectrometry-Based Plasma N-Glycome Analysis with Linkage-Specific Sialic Acid Esterification. J. Proteome Res. 2015;14:4080–4086. doi: 10.1021/acs.jproteome.5b00538. [DOI] [PubMed] [Google Scholar]
- Fu B., Wang G., Li C., Li Y., Liu X., Lu H., Zhang Y.. GlyTrait: A Versatile Bioinformatics Tool for Glycomics Analysis. J. Proteome Res. 2025;24:5484–5497. doi: 10.1021/acs.jproteome.5c00320. [DOI] [PubMed] [Google Scholar]
- Varki A.. Sialic acids in human health and disease. Trends in Molecular Medicine. 2008;14:351–360. doi: 10.1016/j.molmed.2008.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pongracz T., Mayboroda O. A., Wuhrer M.. The Human Blood N-Glycome: Unraveling Disease Glycosylation Patterns. JACS Au. 2024;4:1696–1708. doi: 10.1021/jacsau.4c00043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chu C. S., Niñonuevo M. R., Clowers B. H., Perkins P. D., An H. J., Yin H., Killeen K., Miyamoto S., Grimm R., Lebrilla C. B.. Profile of native N-linked glycan structures from human serum using high performance liquid chromatography on a microfluidic chip and time-of-flight mass spectrometry. Proteomics. 2009;9:1939–1951. doi: 10.1002/pmic.200800249. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data of this study is publicly available as a Panorama Public submission https://panoramaweb.org/Navigating_HumanSerum_NGlycoproteomics.url with the associated ProteomeXchange ID PXD077279. The Supporting Information section of the Panorama repository contains the curated Skyline documents, associated raw LC–MS/MS data, the RNotebook_SkylineGlycoproteomicsTemplate.Rmd notebook for Skyline template generation together with the files required for its execution, and the PostSearchValidation_Glycoproteomics_Reproducibility_Workflow.Rmd notebook for processing Skyline exports and reproducing the reported analyses together with the files required for its execution.
