Abstract
The post-translational modification (PTM) of proteins by O-linked β-N-acetyl-D-glucosamine (O-GlcNAcylation) is widely found across the proteome and regulates diverse cellular processes, from transcription and translation to signal transduction and metabolism. However, most functional studies to date have focused on individual modifications, overlooking other simultaneous O-GlcNAcylation events that work together to coordinate cellular activities. Here we describe networking of O-GlcNAc transferase interactors and substrates (NOTISE), a systems-level approach that monitors O-GlcNAcylation rapidly and comprehensively across the proteome to reveal important functional and regulatory relationships. The NOTISE method integrates affinity purification–mass spectrometry and site-specific chemoproteomic technologies with network generation to connect putative upstream regulators and downstream targets of O-GlcNAcylation. The resulting data-rich networks identify critical conserved activities of O-GlcNAcylation and tissue-specific functions. This holistic and unbiased approach provides a broadly applicable framework to catalyze investigations into the functional roles of coordinated, multisubstrate PTMs in specific cellular and physiological contexts.

Subject terms: Glycobiology, Chemical tools, Carbohydrates
Functional studies of O-GlcNAcylation have often focused on individual modifications. Now, a systems-level approach has identified simultaneous O-GlcNAcylation events that coordinate cellular activities and tissue-specific functions.
Main
Nearly all eukaryotic cellular processes are regulated by post-translational modifications (PTMs), which enable the dynamic coordination of protein function, localization, stability and interactions in response to diverse stimuli. The PTM O-linked β-N-acetyl-D-glucosamine (O-GlcNAc) is essential for cell survival and couples the metabolic status of the cell to a wide range of critical functions, including transcription, translation, metabolism, protein homeostasis and signal transduction1. Addition of O-GlcNAc to serine and threonine residues (known as O-GlcNAcylation) occurs on thousands of intracellular proteins in all multicellular organisms. O-GlcNAcylation is induced in response to normal physiological stimuli such as extracellular signaling, nutrient availability and cellular stress2. Aberrant O-GlcNAcylation has a role in many human diseases, including diabetes, cancer and neurodegenerative disorders, making the modification an attractive therapeutic target3,4. Despite this, the vast majority of identified O-GlcNAc sites remain functionally uncharacterized and how O-GlcNAcylation is coordinated on a global level across the cell is poorly understood.
Cycling of O-GlcNAc occurs on different timescales and across multiple subcellular compartments to dynamically regulate protein activity, localization and stability, while influencing long-term genetic reprogramming through transcription factor recruitment and epigenetic modifications. For example, cancer cells reprogram gene expression and metabolic flux by altering the O-GlcNAcylation status of c-Myc5, phosphofructokinase6 and Tet methylcytosine dioxygenases7, suggesting that multiple O-GlcNAcylation events collectively influence the aggressiveness and metastatic potential of growing tumors. In neurons, excitatory and inhibitory stimuli modulate the O-GlcNAcylation status of numerous proteins from the synapse to the nucleus, thereby contributing to synaptic remodeling8, neuronal survival9 and mitochondrial bioenergetics10. These and other findings provide strong evidence for high-level coordination of O-GlcNAcylation within the cell. However, virtually all studies to date have characterized individual proteins one at a time, neglecting the multitude of simultaneous O-GlcNAcylation events that collectively contribute to various concurrent cellular responses. Furthermore, because of a lack of suitable methods, systematic, quantitative comparisons of O-GlcNAcylation across the proteome in different cell types, states or tissues have not been readily accessible. Thus, holistic methods are greatly needed to determine how this broadly acting PTM orchestrates cellular outcomes through its many protein targets.
Dynamic protein assemblies and proximity-induced effects are fundamental mechanisms for the regulation of cellular processes such as signal transduction, transcription and chromatin remodeling11. PTM writers and erasers are often brought in proximity to their substrates through specific, stable protein–protein interactions (PPIs). For example, the binding of cyclic AMP-dependent protein kinase (PKA) to A-kinase-anchoring proteins serves to compartmentalize PKA activity within cells12, while ubiquitination and degradation of specific proteins are controlled by the association of E3 ubiquitin ligases with E2 conjugating enzymes13. Notably, O-GlcNAc transferase (OGT), the sole enzyme responsible for O-GlcNAcylation, forms stable protein complexes through its N-terminal tetratricopeptide repeat domains, suggesting that OGT may also be targeted to certain substrates14–16. Therefore, we sought to use stable OGT–protein interactions across the cell as a discovery framework17 to decipher the diverse specificity, cellular functions and dynamic regulation of O-GlcNAcylation.
Here, we describe the networking of OGT interactors and substrates (NOTISE) approach for the comprehensive, systems-level analysis of O-GlcNAcylation. To demonstrate the approach, we developed robust, proteomics-based methods to identify both OGT-interacting proteins and O-GlcNAc-modified substrates, which enabled us to determine the molecular components of O-GlcNAc activity for a given biological state (Fig. 1a,b). These matched elements were then combined using global PPI datasets to produce large functional O-GlcNAcylation networks that were partitioned into process-specific subnetworks by unsupervised clustering (Fig. 1c). Analysis of the resulting communities revealed functional outcomes of O-GlcNAcylation for that biological state, providing insights unattainable from the individual interactome or O-GlcNAcome datasets alone. Furthermore, incorporation of bioinformatics data regarding protein domains, amino acid motifs and other PTMs into the O-GlcNAcylation networks created a large, information-rich resource that connected specific glycosites to potential upstream and downstream regulatory proteins and pathways. By exploiting the unique connectivity provided by NOTISE, we identified many highly interconnected OGT interactors including BAP1 and found that this protein can both positively and negatively regulate the O-GlcNAcylation of specific sites on proximal substrates within the network. Lastly, we expanded our NOTISE approach in vivo (iNOTISE) through the generation of a valuable mouse model and organ-specific networks that revealed conserved, core activities of O-GlcNAcylation and remarkable tissue-specific functions. Together, our studies provide a powerful experimental platform, global framework and information-rich resource for understanding the specificity, regulation and functions of O-GlcNAcylation with unparalleled depth and scale.
Fig. 1. Overview of NOTISE, a systems-level approach to reveal functional and regulatory relationships for O-GlcNAcylation.

a, OGT interactors are isolated by TAP of Dox-inducible, dual-tagged OGT using SILAC for quantification and identified by MS. IP, immunoprecipitation. b, OGT substrates are biotinylated by selective CL of the O-GlcNAc moiety and CuAAC using a chemically cleavable biotin tag (1; Fig. 2e). Tryptic digestion, affinity purification of the biotinylated peptides, tag cleavage and MS2 analysis leads to protein identification and mapping of O-GlcNAc events to localized sites or small peptide regions. R = H for serine or CH3 for threonine. c, Networks using context-specific datasets generated in a and b are constructed by connecting interactors and substrates (nodes) through edges derived from established PPI databases, followed by community-based partitioning, which assembles functionally related subclusters.
Results
Identification of context-specific, native OGT interactors
To implement NOTISE, we first required comprehensive, robust and general methods to identify context-specific OGT-interacting proteins. Although all substrates must interact at least transiently with OGT, we aimed to determine proteins in stable complexes that could potentially target OGT to its substrates. Therefore, we combined tandem affinity purification–mass spectrometry (TAP–MS) with stable isotope labeling by amino acids in cell culture (SILAC) and spectral counting to generate quantifiable, statistically robust OGT interactomes.
Full-length OGT with dual C-terminal FLAG and HA epitope tags (OGT–FH; Fig. 1a) was stably expressed in HEK293T cells under the tetracycline-responsive element (TRE) promoter using lentiviral vectors. This system allowed for uniform, controllable levels of OGT–FH (Extended Data Fig. 1a) through the addition of doxycycline (Dox). OGT–FH and its associated proteins were then subjected to two sequential rounds of immunoprecipitation (Figs. 1a and 2a). Several known OGT interactors were pulled down and detected by western blotting, including retinoblastoma-binding protein 5 (RBBP5), nuclear pore glycoprotein p62 (NUP62), WD repeat-containing protein 5 (WDR5) and host cell factor 1 (HCFC1), validating our TAP procedure (Fig. 2b). For quantitative OGT interactome analyses, we compared proteins immunoprecipitated from Dox-treated, OGT–FH-expressing cells to those immunoprecipitated from vehicle-treated, non-OGT–FH-expressing cells in triplicate. In total, we identified 519 high-confidence OGT interactors from HEK293T cells, with hits defined as proteins with a corrected P value < 0.05 from either SILAC enrichment ratio analysis or spectral counting (Fig. 2c and Supplementary Table 1). The dataset was highly robust and reproducible, with all interactors observed in at least two replicates and nearly all interactors (498 proteins, 96%) found in all three independent replicates (Fig. 2d). A substantial number of the hits (323 proteins, 62%) were previously undiscovered OGT interactors when compared to the widely used, highly curated PPI databases, BioGRID and IntAct. Notably, we identified high-confidence interactors associated with seminal advances in the O-GlcNAc field in addition to many previously unknown interactors (Supplementary Note 1).
Extended Data Fig. 1. Methods for OGT expression and O-GlcNAc site identification.

a, Western blotting of inducible OGT-FH expression by HEK293T cells. Cells were treated with the indicated doxycycline (DOX) concentrations for 24 h. After lysis, lysates were probed for OGT-FH expression or O-GlcNAc levels by western blotting with an anti-HA or anti-O-GlcNAc (RL-2) antibody. b, Synthesis of biotin-Dmpt-alkyne 1. (a) 1,3-Dimethylbarbituric acid (DBA, 1.3 eq), EDC·HCl (1.3 eq), TEA (1.5 eq), DMAP (0.1 eq), CH2Cl2, 0 °C to RT, 16 h, 92% yield. (b) Propargylamine (neat), 50 °C, 4 h, 28% yield. Detailed synthetic protocols are provided in the Methods section. c, Comparative stability of biotin-Dde-alkyne and biotin-Dmpt-alkyne. Biotin-Dde-alkyne or biotin-Dmpt-alkyne 1 was reacted with azido-Cy3 and immobilized onto high-capacity Neutravidin resin (top). After overnight incubation in phosphate-buffered saline, the dye conjugated to the Dde linker showed leaching into the buffer at higher temperatures necessary for tryptic digestion (middle; arrow). Both linkers showed efficient release of the immobilized dye after treatment with NH2OH (bottom). d, Chemical structure and exact mass of oxonium diagnostic ion produced by O-GlcNAcylated peptides after chemoenzymatic labeling, linker functionalization, and Dmpt cleavage. e, Instrument workflow for O-GlcNAc site identification. Each MS1 parent ion selected by the instrument is first subjected to data-dependent HCD fragmentation with Orbitrap detection (ddMS2-OT-HCD) to potentially generate 300.1 m/z diagnostic ion from the tagged O-GlcNAc moiety (1). Presence of the diagnostic ion in the HCD MS2 spectrum triggers either ETD or EThcD (in separate runs) fragmentation on the same precursor (2), also with Orbitrap detection (ddMS2-OT-ETD or ddMS2-OT-EThcD). f-g, Annotated MS2 spectrum of the HCFC1 peptide containing a GlcNAc moiety at T1525 fragmented by (f) HCD and (g) ETD. Note that the O-GlcNAc modification is only able to be definitively localized to a specific threonine by ETD. h, Pie chart of the overlap between all unique, localized O-GlcNAc sites identified by ETD and EThcD fragmentation methods in HEK293T cells. i, Annotated MS2 spectrum of the SPEN peptide containing a GlcNAc moiety at S2620 fragmented by ETD. j, Annotated MS2 spectrum of the HIPK1 peptide containing a GlcNAc moiety at T1001 fragmented by EThcD. For f-g and i-j, differences between calculated and observed m/z values for peptide fragment ions are presented as parts per million (ppm) error.
Fig. 2. Generation of context-specific OGT interactor and substrate datasets.

a, Silver-stained gel profiles of cell lysates and purified OGT interactors after TAP using HEK293T cells stably expressing Dox-inducible OGT–FH. b, Validation of the TAP method by western blotting for known OGT interactors. c, Plot of the −log10P of the limma-moderated t-test (two-sided) of SILAC ratios versus the inverted beta-binomial test (one-sided) of spectral counts for all proteins identified in the OGT TAP–MS experiments from HEK293T cells. Proteins with P < 0.05 in either test (after adjusting for multiple comparisons using the Benjamini–Hochberg method) are considered statistically significant interactors. Proteins with P > 0.05 (dashed lines) in both tests are indicated with gray circles and not counted toward the total number of interactors. Known interactors from the BioGRID and IntAct databases are denoted in red and previously unidentified interactors are denoted in blue. d, Venn diagram showing overlap of identified OGT interactors across three independent replicates. e, Chemical structure and hydroxylamine-mediated release of cleavable biotin tag 1. R, labeled O-GlcNAcylated peptide. f, Quantification of O-GlcNAc events (Methods and Extended Data Fig. 2) identified in HEK293T cells by HCD, ETD and EThcD fragmentation. The number of localized events mapped to an individual amino acid (dark and light purple) is compared to the number of unlocalized events found in a region but not localized to an individual amino acid (gray). Peptides with a single S or T residue (dark purple), which by definition have an inherent localization probability of 100%, represent only a tiny fraction of the localized events. g, Distribution of localization probabilities of O-GlcNAc events identified from HEK293T cells. h, Annotation of localized O-GlcNAc sites identified on HCFC1, with sites not annotated in O-GlcNAc database version 2.0 shown in red. i,j, Frequency plot (i) and enrichment plot (j) of amino acid residues surrounding localized O-GlcNAc sites generated using WebLogo. A slight preference for small hydrophobic residues near the modification site was observed, including proline at the −3 and −2 positions as reported previously22, as well as higher frequencies of both proximal and distal serine and threonine residues.
Proteome-wide mapping of protein O-GlcNAcylation sites
We next generated a comprehensive database of context-specific OGT substrates by expanding on our chemoenzymatic labeling (CL) approach18,19. CL uses a mutant galactosyltransferase (GalT-Y289L) to append the nonnatural monosaccharide β-N-azidoacetyl-D-galactosamine (GalNAz) onto O-GlcNAc-modified proteins as a ‘clickable’ handle for downstream functionalization and enrichment (Supplementary Note 2 and Methods). To enhance the capabilities of CL further, we developed a new chemically cleavable tag to improve the capture-and-release of modified targets. Biotin-Dmpt-alkyne 1 (Fig. 2e and Extended Data Fig. 1b) incorporates a N,N′-dimethylpyrimidinetrione (Dmpt) group that exhibits higher chemical and thermal stability than the previous generation 1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl (Dde) scaffold, allowing for peptide level enrichment after trypsin digestion (Extended Data Fig. 1c)20. Upon Dmpt cleavage, the resulting primary amine aids in glycopeptide fragmentation through electron transfer dissociation (ETD) by increasing the peptide charge density. Moreover, the tagged peptides produce a signature ion of 300.1 m/z during higher-energy collisional dissociation (HCD) (Extended Data Fig. 1d), which enables the rapid, definitive identification of O-GlcNAc-modified peptides20.
To accelerate O-GlcNAc peptide identification, we developed an efficient MS method that capitalized on this unique signature ion and targeted the tagged glycopeptides for sequencing and glycosite mapping (Extended Data Fig. 1e). Detection of the 300.1 m/z reporter ion by HCD was used to trigger fragmentation of the precursor ion by either ETD or ETD with HCD supplemental activation (EThcD), thereby avoiding the time-intensive fragmentation of unmodified peptides and leading to faster instrument cycle times. HCD alone identified the most O-GlcNAcylated events, followed by HCD-triggered EThcD and then HCD-triggered ETD (Fig. 2f). As expected19,20, HCD fragmentation led to neutral loss of the glycan moiety, preventing exact glycosite localization. Importantly, we found that HCD-triggered ETD and EThcD each identified unique, nonoverlapping sets of glycosites (Extended Data Fig. 1f–j), underscoring the power of using both ETD and EThcD in combination to maximize O-GlcNAc event identification and site localization.
To facilitate counting, we developed a computational program to calculate the total number of nonredundant O-GlcNAcylation events in the most accurate and parsimonious way possible. Unlike previous methods, this approach eliminates redundant modifications assigned by MS analytical software packages to solve the common problem of duplicate site counting because of missed cleavages or different peptide charge states (Extended Data Fig. 2 and Methods). To describe this nonredundant list of modification events, we adopt the terminology ‘sites’ and ‘regions’ herein, where ‘sites’ refers to modifications localized onto an individual amino acid with a probability greater than random chance and ‘regions’ refers to modifications localized to a short peptide but not a single site. Because regions cover a span of multiple amino acids, they can contain more than one modification. For counting purposes, each unique O-GlcNAcylation event can occur in either a site or region but not both. Notably, this general quantification strategy can be directly applied to other PTMs such as phosphorylation.
Extended Data Fig. 2. Workflow to identify O-GlcNAc sites and regions.

Method to process peptide spectrum matches (PSMs) and compile a minimized list of O-GlcNAc localized sites and unlocalized regions. O-GlcNAc sites are only considered localized if the localization probability (L) is greater than random chance. The threshold for nonrandom probability that a site is modified (P) is calculated as the number of sites in the PSM (m) divided by 1 + m. a, For a singly modified PSM, localization probabilities are calculated for each site based on MSn fragmentation and ranked in descending order. The highest ranked probability is compared to P. If L > P as in PSM1, then the site is considered localized. If L < P as in PSM2 and PSM3, then probabilities are summed in descending order until the sum is greater than P. This new unlocalized region is then compared to other sites within the list to see if it is better explained by another PSM. For example, PSM2 is not added to the list because the putative region is better explained by a localized site in PSM1. PSM3 is added to the list as an unlocalized region because the putative region is not better explained by another PSM. b, For a PSM modified by m sites, the highest m localization probabilities are summed and compared against m*P. For PSM4, one site is already included from PSM1; thus, this site is not added to the list as a new entry. The other site better explains the region from PSM3 and is included to the list as a replacement for the PSM3 region. For PSM5, the workflow generates a region of three sites, two of which are better explained by other PSMs. Therefore, PSM5 only adds a single unlocalized region to the overall list. For PSMs where multiple amino acid residues have equal localization probability such as HCD spectra that lack localization, an unlocalized region incorporating all sites is added to the total list unless the site is better explained by another PSM. Together, this workflow generates a list containing the least number of unique O-GlcNAc events that explains all PSM data from the experiment.
Together, our analysis of HEK293T cells yielded 1,764 unique O-GlcNAc modification events (Supplementary Table 2) across 646 proteins (Supplementary Table 3), with a false discovery rate of 1%. The improved MS protocols produced highly efficient peptide fragmentation without loss of the glycan moiety, allowing us to localize the modification to a single serine or threonine residue (site) with >99% probability for most events (Fig. 2g). For example, we definitively assigned 66 of the 72 total O-GlcNAc modifications we detected on HCFC1 as localized sites, including six sites that were not previously annotated in the O-GlcNAc database version 2.0 (Fig. 2h)21. The frequency (Fig. 2i) and enrichment analysis (Fig. 2j) of all localized sites using WebLogo revealed no consensus sequence for OGT as expected22. Interestingly, 80% of the OGT interactors were not found to be O-GlcNAc modified in the HEK293T samples, resulting in two largely distinct datasets.
Generation of functionally partitioned O-GlcNAc networks
We next connected our large datasets of OGT interactors and substrates (or ‘nodes’) through defined, physical interactions. Because we hypothesized that proximity-induced effects involving OGT, its interactors and its substrates may regulate OGT activity, known PPIs between OGT interactor–interactor pairs or OGT interactor–substrate pairs were obtained from the BioGRID or IntAct databases and used as undirected edges between interactor and substrate nodes. This combination of protein nodes and PPI edges resulted in a proteome-wide O-GlcNAc network that was visualized using Cytoscape. To partition the network into individual clusters for further analysis, we used the community clustering algorithm GLay23. This algorithm scores edge betweenness for node groups to extract local communities from a larger network and partition them into functionally relevant clusters. Notably, the GLay algorithm is inherently agnostic to protein identity, which provides a functionally unbiased approach to network partitioning. Partitioning of the complete network revealed ten, highly intraconnected communities, five of which had more than ten proteins (Fig. 3a and Supplementary Data 2). To assess whether the communities contained functionally related nodes, we then performed a Gene Ontology (GO) analysis using biological process terms. Grouping of the significant GO terms using ClueGO revealed distinct cellular functions within the five major subnetworks: (1) post-transcriptional regulation; (2) chromatin organization and remodeling; (3) nuclear and organelle transport and cytoskeleton organization; (4) RNA/DNA synthesis and catabolism; and (5) mRNA splicing and translational initiation (full annotation of all clusters in Extended Data Fig. 3).
Fig. 3. NOTISE reveals diverse potential regulatory roles for the O-GlcNAc modification.

a, O-GlcNAc functional network from HEK293T cells with communities annotated on the basis of enriched biological process GO terms (Methods and Extended Data Fig. 3). b, Hierarchical clustering of enriched biological process GO terms from initial interactor (Int) and substrate (Sub) datasets and individual GLay communities demonstrates efficient partitioning of biological functions. Significant GO terms were identified using g:Profiler (P < 0.01 by two-sided Fisher’s exact test after correction for multiple comparisons using the default g:SCS method; Methods and Supplementary Data 8). c, Annotation of subcluster 2 highlighting four protein complexes (NSL, SET1/COMPASS, PR-DUB and MLL) involved in epigenetic regulation of gene expression.
Extended Data Fig. 3. Enriched pathways of O-GlcNAc functional networks in HEK293T cells.

Pie charts of the percentage of genes per GO term for significantly enriched GO biological process terms in each annotated community. Pie charts were generated by ClueGO. For clarity, only terms with more than 5% of included genes are labeled.
To evaluate the effectiveness of our GLay partitioning approach, we compared the biological process GO terms that were significantly enriched (P < 0.01) in the interactor or substrate datasets alone, in the combined dataset and within each partitioned community. Hierarchical clustering based on coenrichment demonstrated that the GO terms were partitioned well into individual clusters, with relatively few terms distributed across multiple clusters (Fig. 3b). Moreover, nearly all GO terms from the original datasets were found to be enriched within the final groupings, suggesting no notable loss of functional data. Similar results were obtained through analysis of other terms (Extended Data Fig. 4a–f and Supplementary Table 4). NOTISE also uncovered functional annotations within the individual clusters that were absent from analyses of the interactor, substrate or combined datasets before GLay partitioning (Extended Data Fig. 4g and Supplementary Table 5). Notable examples from each cluster are described in Supplementary Note 3. In total, there were 145 biological process, 20 molecular function and 40 cellular component GO terms, as well as nine KEGG, 66 Reactome, nine WikiPathways and 43 CORUM complexes that emerged only after network clustering (Supplementary Table 5).
Extended Data Fig. 4. Partitioning and novel pathway analyses of O-GlcNAc functional networks in HEK293T cells.

a–f, Hierarchical clustering of significant functional annotations from OGT interactors (Int), substrates (Sub), and network clusters. Significant functional annotations were identified using g:Profiler, P < 0.01 by two-sided Fisher’s exact test after correction for multiple comparisons using the default g:SCS method (see Methods and Supplementary Data 8). g, Hierarchical clustering of 145 GO Biological Process terms that were newly enriched after network clustering. GO term from each cluster with the lowest P value indicated on the right. Significant GO terms were identified as described for a–f above.
The NOTISE networks and annotations provide critical insights into the central functions of O-GlcNAc for a given cell type or state. For example, cluster 2 (chromatin organization and remodeling) was enriched for 44 of the 82 CORUM complexes in the network (Extended Data Fig. 4f and Supplementary Table 4), including the histone H3K4 methyltransferase complexes SET1/COMPASS and mixed lineage leukemia (MLL), the Polycomb repressive deubiquitinase (PR-DUB) complex and the nonspecific lethal (NSL) histone acetyltransferase complex (Fig. 3c). These complexes suggest functional roles for O-GlcNAcylation in the regulation of three major epigenetic histone marks. Consistent with these findings, recent studies have shown that OGT and O-GlcNAcylation contribute to SET1-mediated histone methylation24, MLL stability25 and NSL-mediated histone acetylation26. O-GlcNAcylation was also found to coordinate transcriptional repression by the PR-DUB complex in Drosophila27 and our NOTISE network indicates that this activity is likely an important, conserved function in HEK293T and other mammalian cells. Moreover, the wealth of site-specific and PPI information in other proteins surrounding the canonical members of these complexes within our network suggests that O-GlcNAc may have broader roles in coordinating the overall activities of these histone-modifying complexes.
Cluster 5 revealed previously uncharacterized roles for O-GlcNAcylation in RNA splicing. Consistent with this function, pharmacological inhibition of OGT has been reported to globally alter detained intron splicing28. Notably, our network contains several splicing factors shown to be differentially phosphorylated upon OGT inhibition, thereby providing possible molecular mechanisms for this intron-splicing activity. For example, we detected OGT interactions with known intron retention mediators, including protein arginine N-methyltransferase 5, numerous heterogeneous ribonucleoprotein particles (hnRNPs) and serine/arginine-rich splicing factors SRSF1, SRSF3 and SRSF5 (Extended Data Fig. 5a), and O-GlcNAcylation of the phosphoprotein SRSF11. Beyond its role in intron splicing, we found evidence that O-GlcNAcylation may alter core and auxiliary spliceosome components to regulate the splicing machinery more broadly (Supplementary Table 6). For example, the E complex components U1 small nuclear polypeptide A (SNRPA) and splicing factor 1 (SF1), which are involved in the initial recognition of splicing sites and commit pre-mRNA to the splicing pathway, act as both OGT interactors and substrates (Extended Data Fig. 5b and Supplementary Note 4). Lastly, multiple OGT interactors and substrates comprise the Prp19/Nineteen complex, a core component of the catalytic spliceosome C complex that is required for intron removal (Extended Data Fig. 5c). These data exemplify how our holistic network-based approach can be used to provide deeper insights into established functions and discover new activities for O-GlcNAcylation.
Extended Data Fig. 5. Potential roles of OGT in RNA processing from O-GlcNAc functional networks.

a, Schematic of identified OGT interactors and substrates involved in intron retention and pre-mRNA stabilization. b, Schematic of identified OGT interactors and substrates within the spliceosome early complex, including SNRPA within the U1 small nuclear ribonucleoprotein and SF1. Ub = ubiquitination, P = phosphorylation, Me = methylation. c, Schematic of identified OGT interactors and substrates within the Prp19/Nineteen complex of the catalytic spliceosome complex. Structure from PDB 5YZG.
NOTISE analysis of PTM crosstalk and PPI modulation
As crosstalk between O-GlcNAcylation and other PTMs provides exquisite control over protein activity2,29, we incorporated PTM databases into our network structure. Integration with the PhosphoSitePlus database showed a striking number of proximal PTMs (Fig. 4a and Supplementary Table 7). Roughly 25% of the 1,101 localized O-GlcNAc sites identified within canonical Swiss-Prot protein sequences directly overlapped with annotated phosphorylation sites and nearly 65% of the sites resided within ten amino acid residues of a phosphorylation site. For example, we identified a well-documented regulatory O-GlcNAcylation site (T717) within signal transduction and activator of transcription 3 (STAT3) (Fig. 4b) that negatively modulates the phosphorylation of neighboring sites critical for STAT3 transcriptional activity30. We also detected O-GlcNAcylation directly at the S727 phosphorylation site, which may provide further, unexplored mechanisms for regulating STAT3 function. Beyond phosphorylation, hundreds of O-GlcNAc sites were found within ten amino acid residues of acetylation, methylation or ubiquitination sites, indicating that PTM crosstalk with O-GlcNAc may be an understudied, general feature of cell signaling and protein regulation.
Fig. 4. Integration of existing datasets into NOTISE-generated networks uncovers potential molecular mechanisms of O-GlcNAc regulation.

a, Quantification of PTM sites within the PhosphoSitePlus database that overlap with (black) or are within ten amino acids of (red) O-GlcNAc events from the functional network. b, Schematic of potential STAT3 regulation by PTM crosstalk between known phosphorylation sites (P) and localized O-GlcNAc sites. c, Schematic of potential SPEN regulation by PPI disruption through O-GlcNAcylation of receptor-binding motifs. d, Ranking of potential OGT hub proteins within the network (Supplementary Table 9). The top ten hub proteins are shown in red. e, Extracted first-neighbor network of potential OGT hub protein RBM25 and connected OGT substrates. f, Validation of selective OGT decoupling from the PR-DUB complex through BAP1 silencing and western blotting of PR-DUB and SET1 complex components. g,h, Site-specific quantification of network-wide (g) and BAP1-adjacent (h) O-GlcNAc site occupancies upon knockout of the OGT interactor BAP1. Only sites with significantly changing occupancies are shown. The log2 fold changes for individual sites are reported as insets on each node and the node outline color represents the average log2 fold change of all sites on the protein. i,j, Quantification of edge distances from BAP1 for all nodes, individual changing sites and proteins with changing sites, displayed as a histogram (i) or table (j) comparing neighbors to more distant interactors. The red bars indicate the increased number of changing O-GlcNAc events and substrates that are one edge away from BAP1 in the network. Differences between the proportions of nodes that are neighboring (one edge away) versus distant (two or more edges away) in each dataset were determined by two-sided Fisher’s exact test and exact P values for j are reported when possible.
Another major mechanism by which O-GlcNAcylation alters protein function is by modifying critical allosteric or protein–protein interfaces to promote or inhibit PPIs6,8. Because proteome-wide information regarding the exact amino acid residues found at PPIs is scarce and poorly annotated, we integrated protein domain, motif, region and amino acid repeat data from the UniProtKB database into the network (Supplementary Table 8) and examined whether the identified O-GlcNAc modifications occurred within protein regions implicated in PPIs. This approach yielded numerous O-GlcNAc events that overlapped with canonical binding motifs and protein-specific regions identified as protein–protein interfaces. For example, we found seven localized O-GlcNAcylation sites within the nuclear receptor interaction domain of Msx2-interacting protein SHARP (SPEN), including two sites, T2619 and S2620, that directly overlap with a conserved IxxVI binding motif (Fig. 4c and Extended Data Fig. 1i)31. Thus, O-GlcNAcylation of SPEN may alter its association with nuclear receptors such as the metabolic regulator peroxisome proliferator-activated receptor-δ. Our analysis also uncovered O-GlcNAcylation of the autophagy regulator serine/threonine protein kinase ULK1 at S409 in a region that mediates its interaction with the GABA type A receptor-associated protein, which promotes ULK1 activity and autophagosome formation upon starvation32. Lastly, we identified a single O-GlcNAc site on homeodomain-interacting protein kinase 1 (HIPK1) at T1001 within a region required for tumor suppressor p53 (TP53) binding that has been suggested to enhance TP53 activation and limit tumor cell growth (Extended Data Fig. 1j)33. These examples underscore the ability of NOTISE to reveal unique, readily testable hypotheses regarding mechanisms of O-GlcNAc function. We created a Cytoscape file integrating the above information from UniProtKB and PhosphoSitePlus (Supplementary Data 2) to allow for live interrogation, contextualization and visualization of these networks by the broader scientific community.
NOTISE identifies hubs for selective O-GlcNAc modulation
OGT interactors attached to multiple substrates within our networks (or ‘network hubs’) may direct OGT activity toward a specific subset of proteins by molecular proximity. However, the prevalence of such network hubs and the cellular mechanisms of OGT regulation remain largely unknown. To search for hubs, we developed a program to rank OGT interactors on the basis of their relative connectivity. Interactors were scored positively (+1) for each connection or edge to an OGT substrate and penalized (−0.5) for each connection to another OGT interactor. This facilitated the identification of interactors that were highly connected to multiple substrates, without overselecting for members of large protein complexes (Fig. 4d and Supplementary Table 9). Among the top hits were DPY30 and Set1/Ash2 histone methyltransferase complex subunit ASH2, which are core scaffolding components of the SET1/COMPASS and MLL complexes (Supplementary Table 9 and Supplementary Note 5)34. Other highly ranked interactors suggested multiple roles for targeted O-GlcNAcylation in the post-transcriptional regulation of protein expression. These hits included the RNA exon junction complex component, RNA-binding protein 25 (RBM25), which engages spliceosome components and OGT substrates PRPF19, PRPF40A and SRRM2 (Fig. 4e), and the RNA-induced silencing complex component, endoribonuclease DICER1.
The existence of OGT interactor hubs could serve to decouple specific substrates and enable independent regulation of their O-GlcNAcylation status. Despite the discovery of many therapeutically important OGT substrates such as tau and ɑ-synuclein2, the selective modulation of individual glycosites on endogenous proteins remains an important challenge. Here, we examined the protein deubiquitinase BAP1, a non-O-GlcNAcylated interactor and component of the PR-DUB complex (Fig. 3c), as a potential target to decouple OGT activity. BAP1 was connected to several OGT substrates in the network, including forkhead box protein K1 (FOXK1) and additional sex combs-like protein 3 (ASXL3). Consistent with the notion of OGT network hubs, small interfering RNA (siRNA) knockdown of BAP1 in HEK293T cells disrupted the association of OGT with FOXK1 of the PR-DUB complex but not with components of the SET1/COMPASS complex, such as SET1, RBBP5 and WDR5 (Fig. 4f). These results indicate that the OGT–BAP1 complex may assemble and function independently from the OGT–SET1/COMPASS complex.
We next investigated whether OGT activity toward substrates proximal to BAP1 in the network was decoupled from other substrates. For this, we knocked out BAP1 in HEK293T cells using CRISPR–Cas9 and evaluated site-specific O-GlcNAc levels (or site occupancy; Supplementary Note 6) across the proteome compared to wild-type (WT) cells (Extended Data Fig. 6a). Unlike quantitative methods that compare only changes in the modified peptide abundance between different conditions, site occupancy measurements normalize the modified peptide abundance to the total protein level for each condition, revealing critical differences in O-GlcNAc stoichiometry at specific sites and eliminating changes merely because of protein expression. To quantify O-GlcNAc occupancies, we incorporated tandem mass tag (TMT)–synchronous precursor selection (SPS)–MS3 into our HCD-triggered ETD and EThcD fragmentation workflow (Extended Data Fig. 6b–f and Methods). A total of 1,087 modifications across 365 unique proteins were identified and quantified over three biological replicates (Supplementary Table 10). To account for changes in protein expression between conditions, we performed quantitative proteomics in tandem, quantifying 6,485 proteins across all conditions, and calculated the protein expression-corrected fold change (that is, relative site occupancy) for 935 modifications (86.0% of the total) across 290 proteins.
Extended Data Fig. 6. Workflow and characterization of O-GlcNAc site quantification in wild-type and BAP1 knockout HEK293T cells.

a, Validation of BAP1 knockout in HEK293T cells by western blotting. b, Instrument workflow for O-GlcNAc site identification and quantification. Each MS1 parent ion selected by the instrument is first subjected to data-dependent HCD fragmentation with Orbitrap detection (ddMS2-OT-HCD) to potentially generate diagnostic ions from the tagged O-GlcNAc moiety. Scans containing any of the diagnostic ions trigger three additional MS scans: (1) synchronous precursor selection of MS2 ions followed by HCD fragmentation with Orbitrap detection (SPS, ddMS3-OT-HCD) for isobaric tag quantification, (2) ETD fragmentation of the MS1 parent ion with the ion trap (ddMS2-IT-ETD) for site localization, and (3) EThcD fragmentation of the MS1 parent ion with the ion trap (ddMS2-IT-EThcD) for site localization. c, Representative MS scan of diagnostic ions and chemical structures produced by TMT-labeled chemoenzymatic linker fragment at 529.2930, 547.3036, and 732.3720 m/z. d, Pie chart of localized and quantified O-GlcNAc sites found in BAP1 WT vs. KO experiment based on ion selection method. e, Bar graph of localized O-GlcNAc sites produced by the three fragmentation methods used. f, Pie chart showing the overlap between unique, localized O-GlcNAc sites identified by ETD and EThcD fragmentation methods in the BAP1 WT vs. KO quantitative glycomics experiment.
Only a small fraction of the quantified modifications (6.8%; 60 events across 27 proteins) across the O-GlcNAcome underwent significant changes upon BAP1 knockout (Fig. 4g,h; full annotated Cytoscape file in Supplementary Data 3) and these changes were not randomly distributed throughout the network. Notably, the frequency of changing modifications and proteins was significantly enriched at nodes directly connected to BAP1 (Fig. 4i,j; events: P < 0.0001, proteins: P = 0.0172, according to a two-sided Fisher’s exact test). For example, ASXL3 underwent an 85.9% decrease in O-GlcNAcylation at S1031, although total ASXL3 protein levels were not quantifiable by MS or western blotting analysis. Interestingly, we also observed an increase in O-GlcNAcylation of the BAP1 interactor ankyrin repeat domain-containing protein 17 (ANKRD17) at S1820 upon BAP1 knockout and both increases (four events) and decreases (three events) at distinct locations on the BAP1 interactor HCFC1. Changes in O-GlcNAc occupancy on HCFC1 also covaried in specific regions, with increases clustered at neighboring residues T515/S518 and T858/T861 and decreases at S1238/T1239/T1243. Conversely, occupancy on the direct BAP1 interactor FOXK1 was unchanged, both overall and at individual quantified events. As FOXK1 dissociates from OGT upon BAP1 knockout (Fig. 4f), a stable complex involving OGT, BAP1 and FOXK1 is not required for FOXK1 O-GlcNAcylation, suggesting that only some substrates may be subject to proximity-directed OGT regulation. Nonetheless, the highly localized changes observed in the network provide direct experimental support that stable OGT interactions with hub proteins can target OGT activity toward a subset of substrates. Moreover, they suggest an expanded, more nuanced form of OGT regulation than proximity-induced glycosylation, highlighting that OGT interactors can promote or even block O-GlcNAcylation at specific locations. Overall, these quantitative proteomic studies demonstrate that O-GlcNAcylation can be selectively modulated on a small subset of OGT substrates and glycosites through manipulation of NOTISE-identified network hubs.
Defining conserved and tissue-specific O-GlcNAc networks
Having established methods to identify functional O-GlcNAc networks in cells, we next expanded our approach in vivo (iNOTISE). CRISPR–Cas9 technology was used to generate a mouse model in which the endogenous Ogt gene was replaced with OGT–FH (Fig. 5a and Extended Data Fig. 7a). The OGT–FH genotype status did not affect breeding (Extended Data Fig. 7b) or cause any detectable phenotype alterations. Using western blotting, we confirmed that the expression levels of tagged and untagged OGT (Fig. 5b) and the overall O-GlcNAcylation levels (Extended Data Fig. 7c) were equivalent in major organs of OGT–FH and WT adult mice, indicating that addition of the dual tags did not affect OGT expression levels or activity in vivo.
Fig. 5. Development and application of O-GlcNAc functional networks in vivo (iNOTISE).

a, Schematic of genome-editing strategy to introduce the FLAG–HA dual tag onto chromosomal OGT. b, Validation that OGT–FH expression levels are equivalent to native OGT across multiple tissues by western blotting. c, Plot of the −log10P of the two-sided limma-moderated t-test of SILAC ratios versus the one-sided inverted beta-binomial test of spectral counts for all proteins identified in the OGT TAP–MS experiments from mouse brains and livers. Proteins with P > 0.05 (dashed lines) in both tests (after adjusting for multiple comparisons using the Benjamini–Hochberg method) are indicated by gray circles and not counted toward the total number of interactors; statistically significant interactors (P < 0.05 in either test) are indicated by purple (brain) or green (liver) circles. d, Number of O-GlcNAc sites (with their average localization probability in each tissue) and regions identified from murine brain and liver (Extended Data Fig. 2). e, O-GlcNAc functional networks from the brain (purple) and liver (green) with communities annotated on the basis of significantly enriched biological process GO terms (Methods and Extended Data Fig. 8). f,g, Hierarchical clustering of enriched biological process GO terms from initial interactor and substrate datasets and individual GLay communities within the brain (f) or liver (g) networks. Significant GO terms were identified using g:Profiler (P < 0.01 by two-sided Fisher’s exact test after correction for multiple comparisons using the default g:SCS method; Methods and Supplementary Data 9). h, Hierarchical clustering of enriched biological process GO terms from brain and liver networks confirms the presence of both shared and tissue-specific functions for O-GlcNAcylation. Significant GO terms were identified as described for f and g. i, Venn diagram of identified OGT interactors and substrates and overall proteomes35 from murine brain and liver.
Extended Data Fig. 7. Characterization of OGT-FH mouse model.

a, Genotyping strategy for OGT-FH mouse model with representative agarose gel. A region containing the C-terminus of the Ogt gene was amplified by PCR, and the PCR product was treated with BamHI. Wild-type animals yield a single band at 2.5 kb, and OGT-FH animals yield two bands at 1.9 kb and 0.6 kb. b, Punnett square of Ogtfh gene inheritance outcomes during the derivation of OGT-FH mouse model. Data analyzed by chi-square (χ2) test with 3 degrees of freedom (df) assuming 25% of each possible genotype. c, Analysis of global O-GlcNAc levels in brain and liver tissue from WT and OGT-FH mice by western blot. The RL-2 clonal antibody was used to measure global O-GlcNAc levels. d, Workflow to identify in vivo OGT interactors. Organs were harvested and lysed in nondenaturing lysis buffer. Lysates were subjected to tandem FLAG and HA co-immunoprecipitation, and eluted proteins were subjected to in solution digestion. Tryptic peptides were tagged by dimethyl labeling using CH2O or CD2O and NaBH3CN. For samples labeled with CD2O, peptides gained 4 or 8 Da based on whether the terminal residue is Arg or Lys, respectively. Samples were then combined 1:1 and subjected to LC-MS/MS identification and quantification. e, Pie chart of 1,543 O-GlcNAc sites localized using ETD or EThcD fragmentation in murine brain and liver tissues.
We then applied iNOTISE to this OGT–FH mouse model. OGT and its interacting partners were immunoprecipitated from the brains and livers of OGT–FH and WT animals. After TAP, proteins were proteolytically digested and the resulting peptides were isotopically labeled by reductive amination using light (H2CO) or heavy (D2CO) formaldehyde before mixing and downstream identification (Extended Data Fig. 7d). MS analysis revealed 993 and 643 high-confidence interactors (n = 3 per tissue type; Fig. 5c and Supplementary Tables 11 and 12) and 2,219 and 1,150 O-GlcNAc events (n = 2 per tissue type; Fig. 5d and Supplementary Tables 13 and 14) across 810 and 492 proteins (Supplementary Tables 15 and 16) from the brain and liver, respectively. Together, our datasets revealed 2,785 nonredundant O-GlcNAc events, consisting of 1,641 localized sites and 1,114 unlocalized events across 1,061 regions. In this extensive, multitissue dataset, most O-GlcNAc modifications were uniquely identified by a single MS fragmentation method (Extended Data Fig. 7e), providing further evidence that combining both ETD and EThcD fragmentation is critical for the comprehensive mapping of O-GlcNAc sites.
Overall, iNOTISE identified 11 communities in the brain and 14 communities in the liver (Fig. 5e; complete annotation of all clusters in Extended Data Fig. 8a,b; results of enrichment analyses (gene, pathway, etc.) in Supplementary Table 17). Similar to the HEK293T networks, GO terms were well partitioned into individual clusters in the liver (Fig. 5f) and brain (Fig. 5g) networks. Hierarchical clustering of all significantly enriched biological process GO and other terms from these organs revealed both conserved and tissue-specific functions for O-GlcNAcylation (Fig. 5h and Extended Data Fig. 9a). For example, chromatin organization, RNA processing and cytoskeletal organization were essential functions identified in the brain, liver and HEK293T networks that likely represent conserved OGT functions, while neuronal development and synaptic vesicle exocytosis were found only in the brain networks and endocytosis, lipid metabolism, and ATP synthesis were unique to the liver (Supplementary Note 7). Interestingly, the overlap between protein substrates in each dataset (39.2%) was close to the reported overlap between total brain and liver proteomes (50.8%)35, whereas the interactor datasets were more distinct (only 27.7% overlap; Fig. 5i), suggesting that differences in OGT–protein interactions may be primarily responsible for the tissue-specific functions of O-GlcNAcylation. Moreover, among the top 25 OGT network hubs identified from each dataset (Fig. 6a, Supplementary Tables 18 and 19 and Supplementary Note 8), only one protein (karyopherin subunit α3) overlapped between the brain and liver, further underscoring the tissue-specific differences and importance of stable OGT interactions in driving those differences.
Extended Data Fig. 8. Analysis of O-GlcNAc functional networks from murine brain and liver tissues.

a, b, Pie charts of the percentage of genes per GO term for significantly enriched GO biological process terms for each annotated community in the (a) brain and (b) liver O-GlcNAc networks. Pie charts were generated by ClueGO. For clarity, only terms with more than 5% of included genes are labeled.
Extended Data Fig. 9. Analysis of O-GlcNAc functional networks in murine brain and liver tissues.

a, Hierarchical clustering of enriched functional annotations from brain and liver networks identifies both shared and tissue-specific functions for O-GlcNAcylation (see Fig. 5h). Significant functional annotations were identified using g:Profiler, P < 0.01 by two-sided Fisher’s exact test after correction for multiple comparisons using the default g:SCS method (see Methods and Supplementary Data 9). b, Liver network with color-coded nodes pertaining to the selected biological functions listed on the right allows for their rapid visualization throughout the network. c, Brain network with color-coded nodes pertaining to long-term potentiation (LTP) and learning and memory. d, Validation of the Ppfia1-OGT interaction by co-immunoprecipitation of Ppfia1 with OGT-FH from adult murine cortical neurons and western blotting.
Fig. 6. NOTISE-generated networks reveal regulatory roles for O-GlcNAcylation in neuronal actin dynamics and synaptic organization.

a, Ranking of network hubs within the brain and liver networks with the top ten hub proteins for each set shown in red (Supplementary Tables 18 and 19). b, Quantification of PTM sites within the PhosphoSitePlus database that overlap with (black) or are within ten amino acids (purple, brain; green, liver) of O-GlcNAc events from the brain and liver networks. c, Frequency of specific protein domains in OGT interactors and substrates from the brain and liver datasets. PH, pleckstrin homology; SH3, Src homology 3; PDZ, PSD95/Dlg1/ZO1 homology; RRM, RNA recognition motif (Supplementary Table 22). int., interactors; sub., substrates. d, Workflow to identify potential O-GlcNAc activities from functional-level networks on the basis of cluster annotation, identification of specific OGT-associated protein complexes and domain-specific and site-specific information. Cyfip2 structure predicted by AlphaFold. e, Extracted network of Ppfia (red) proteins and neighboring interactors (solid node) and substrates (outlined node) within the brain network. f–h, Protein blots (f) and quantification (g,h) of O-GlcNAcylated Ppfia1 in Thiamet G (TMG)-treated, KCl-stimulated and resting primary mouse embryonic cortical neurons, as detected by CL with UDP-GalNAz (n = 4 biological replicates)18,19. Exact P values are reported and were calculated using a paired two-sided Student’s t-test. i,j, Schematic of the highly conserved putative C-terminal O-GlcNAcylation event across Ppfia proteins (i) and its complex with the Grip1 PDZ6 domain (j) (structure from PDB 1N7F). O-GlcNAc regions were identified in Ppfia1 and Ppfia3, whereas exact O-GlcNAc sites were identified in Ppfia2 and Ppfia4.
Integration of our datasets with the UniProtKB and PhosphoSitePlus databases as described above (Supplementary Data 4 and 5 and Extended Data Fig. 9b,c) uncovered a wealth of potential PTM crosstalk within the brain and liver proteomes, with 1,928 O-GlcNAc events overlapping with or adjacent to other PTMs within the brain network alone and multiple sites of potential importance to neurological disorders and cancer (Fig. 6b, Supplementary Tables 20 and 21 and Supplementary Note 9). Annotated structural domains of interest were also readily identified and quantified throughout the networks, providing additional functional information (Fig. 6c, Supplementary Table 22 and Supplementary Note 9). For example, protein kinase domains were a recurring feature of many brain and liver network nodes, revealing unexplored instances of potential O-GlcNAc regulation of kinase function and specificity29.
iNOTISE as a discovery platform for novel O-GlcNAc functions
The iNOTISE-generated networks connect multiple levels of information to facilitate hypothesis-driven investigations into regulatory roles for O-GlcNAcylation. Although O-GlcNAc has been previously implicated in cytoskeletal remodeling36, our networks suggest expanded roles for O-GlcNAcylation in cytoskeletal dynamics and reorganization, particularly in the Arp2/3 complex and WAVE regulatory complex (WRC) (Fig. 6d and Supplementary Note 10). For instance, O-GlcNAcylation of WRC member Cyfip2 occurred within a region that mediates interactions with Nck-associated protein 1 (Nckap1) and fragile X mental retardation 1 (Fmr1), thus potentially affecting complex assembly and stimulation-dependent actin polymerization. Interestingly, WRC member Abl interactor 2 (Abi2) was identified as both an OGT substrate and a highly ranked network hub, suggesting that it may regulate OGT activity toward numerous substrates involved in actin remodeling and synapse formation.
Although global changes in O-GlcNAcylation are known to affect synapse maturation37, few studies have characterized the OGT interactors and substrates involved in synaptic structure and function. The iNOTISE brain network contained 40 interactors and 28 substrates directly implicated in learning, memory and long-term potentiation (Extended Data Fig. 9c), as well as 144 OGT interactors and 93 OGT substrates associated with synaptic function (Supplementary Table 23). For example, iNOTISE revealed that all members of the Ppfia (liprin-α) family of synaptic scaffolding proteins are O-GlcNAcylated and both Ppfia1 and Ppfia3 were identified as putative OGT interactors (Fig. 6e). Consistent with these findings, we found that Ppfia1 coimmunoprecipitated with OGT from cortical neurons cultured from OGT–FH mice (Extended Data Fig. 9d). We also confirmed that Ppfia1 was O-GlcNAc modified (Fig. 6f) in primary cortical neurons and found that this modification was actively cycled by both OGA inhibition by Thiamet G (Fig. 6g) and KCl-induced neuronal depolarization (Fig. 6h). Ppfia1 glycosylation occurred within a C-terminal peptide that is highly conserved across the protein family (Fig. 6i) and is recognized by the PDZ6 domain of glutamate receptor-interacting protein 1 (Grip1) (Fig. 6j), a protein critical for AMPA receptor clustering in dendrites38. These and other interesting network links between OGT and proteins containing PDZ domains, the third most frequent domain in the iNOTISE brain network (Fig. 6c, Supplementary Table 22 and Supplementary Data 4), reveal unexplored molecular mechanisms for O-GlcNAc regulation of synapse organization, maturation and plasticity.
Discussion
O-GlcNAcylation and other PTMs act simultaneously on multiple protein targets, generating many context-specific proteoforms that coordinate biological activity on a global scale. Comprehensive, function-driven approaches that report on the full scope of PTMs within cells represent the next level of PTM understanding and will provide insights into their myriad functions and dynamic regulation. Here, we developed NOTISE, a broad platform for examining and contextualizing the O-GlcNAc modification across the proteome (Supplementary Data 2–5). By organizing comprehensive OGT interactome and O-GlcNAcome datasets from paired cells and tissues into network communities, NOTISE uncovered both established and new functions for the modification related to histone modification, intracellular transport, transcription and translation, synaptic organization and carbohydrate metabolism. Within these functions, NOTISE also generated a plethora of protein-specific and site-specific data for driving mechanistic hypotheses and functional validation studies. Thus, NOTISE delivers a powerful, systems-level approach to discover the central functions and mechanisms underlying O-GlcNAc activity—a long-standing challenge in the field given the large number of O-GlcNAcylated proteins and the complexity of OGT regulation.
NOTISE also reveals important connections between the upstream regulators and downstream effectors of O-GlcNAc, including potential network ‘hubs’ that direct OGT to specific target substrates. We found that the stable OGT interactor BAP1 can either enhance or block the O-GlcNAcylation of specific sites. Thus, a hub protein can serve not only to anchor OGT toward a substrate but also precisely position it and thereby direct its activity toward certain sites or regions of the protein, providing an elegant mechanism to control the broad specificity of OGT. The observed positive and negative regulation of O-GlcNAcylation by OGT interactors has important implications for proximity-based modalities aimed at enhancing site-specific O-GlcNAcylation4. In addition, our results suggest that it may be possible to selectively modulate the O-GlcNAcylation status of a small subset of OGT substrates or glycosites through manipulation of NOTISE-identified network hubs. Our approach might also guide efforts to modify specific proteins with PTMs39, while confirming that the engineered modifications occur at physiologic sites and stoichiometries.
Lastly, the iNOTISE approach using our dual-tagged OGT mouse model enables comprehensive, comparative studies of O-GlcNAc in complex tissues, illuminating both conserved and tissue-specific functions for O-GlcNAcylation. We identified a core, shared subset of O-GlcNAc activities centered on the regulation of transcription and translation through epigenetic modifications, transcription factors and intron splicing, which highlight important, conserved roles of O-GlcNAc spanning the entire central dogma of molecular biology. In addition to highly conserved functions, iNOTISE revealed diverse tissue-specific functions for O-GlcNAc such as the regulation of synaptic structure and function. The networks also indicated potential roles for O-GlcNAcylation in presynaptic and postsynaptic signaling through multiple distinct substrates, which may reveal new neuronal functions for O-GlcNAcylation and additional mechanisms to explain its established effects on long-term potentiation and long-term depression40.
By applying the tools and methods developed herein to other PTMs, the NOTISE approach can be readily expanded for in-depth multiomic analyses. The advancement of global interactome datasets will undoubtedly improve both node connectivity within the networks and resulting functional annotations. In the future, other resources such as cell-type-specific expression levels, protein disorder and structural prediction, germline or disease-specific protein mutations and small-molecule ligandability could be readily integrated into the networks to further enable molecular hypothesis generation. We envision that networks could be constructed for diverse biological contexts across individual tissues, developmental stages and disease progression. Thus, the global functional annotation of O-GlcNAcylation and other PTMs will greatly advance our understanding of the spatiotemporal, dynamic regulation of proteins and PTM coordination across the proteome.
Methods
Chemical synthesis
General procedures and analyses
Unless otherwise indicated, all chemical synthesis reagents were obtained from MilliporeSigma (Merck) and used as provided. Reactions were performed in flame-dried glassware under Ar using freshly dried solvents passed through an activated alumina column under Ar. Preparative high-performance liquid chromatography (HPLC) was conducted using a 1260 Infinity II LC system (Agilent) with a custom Zorbax Eclipse XDB-C18 PrepHT column (Agilent, 21.2 × 250 mm, 5 μm) with a linear gradient of 5–50% CH3CN with 0.1% formic acid over 70–85 min with a flow rate of 6 ml min−1. 1H and 13C nuclear magnetic resonance (NMR) experiments were recorded on a Bruker 400-MHz spectrometer at the Caltech NMR Facility. Spectra are reported in parts per million (δ) relative to CDCl3 (7.26 ppm). Data are reported as the chemical shift (ppm), multiplicity (s, singlet; d, doublet; t, triplet; q, quartet; p, pentet; m, multiplet; b, broad), coupling constant (Hz) and integration. Mass spectra were obtained using a Waters LCT Premier XE electrospray ionization (ESI) time-of-flight MS instrument at the Caltech Multiuser MS Laboratory.
5-(3-(2-(2-(2-(2-(N-biotinyl)aminoethoxy)ethoxy)ethoxy)ethoxy)-1-hydroxypropylidene)-1,3-dimethylpyrimidine-2,4,6-trione (biotin-Dmpt-OH, 2)
Biotin-PEG4-CO2H (BroadPharm, 400 mg, 0.814 mmol) and N,N′-dimethylbarbituric acid (165 mg, 1.06 mmol, 1.3 equivalents) were suspended in 10 ml of dry dichloromethane (DCM) under Ar and cooled to 0 °C. Separately, 1-(3-dimethylaminopropyl)-3-ethylcarbodiimide hydrochloride (EDC·HCl, 203 mg, 1.06 mmol, 1.3 equivalents), DMAP (9.9 mg, 0.0810 mmol, 0.1 equivalents) and TEA (170 μL, 1.22 mmol, 1.5 equivalents) were dissolved in 5 mL of dry DCM. The coupling solution was added to the suspension dropwise. The reaction was allowed to warm to room temperature and stirred overnight. The solution was then concentrated to a yellow oil and redissolved in a minimal volume of 5% aqueous CH3CN. The product was then purified by preparative HPLC. Fractions were combined, concentrated and lyophilized to afford the product as a pale-yellow powder (469 mg, 92% yield). 1H NMR (500 MHz, CDCl3): δ 7.33 (bs, 0.2H, OH), 6.73 (bt, J = 5.4 Hz, 0.8H, OH), 6.28 (s, 1H, NH), 5.43 (s, 1H, NH), 4.51 (t, J = 6.5 Hz, 1H, biotin-CH), 4.32 (dd, J = 7.6, 4.5 Hz, 1H, biotin-CH), 3.86 (t, J = 6.3 Hz, 2H, OCH2), 3.68-3.59 (m, 12H, OCH2), 3.55 (t, J = 5.1 Hz, 2H, CH2), 3.48-3.27 (m, 10H, OCH2, CH3), 3.15 (dt, J = 12.0, 6.0 Hz, 1H, biotin-CH), 2.91 (dd, J = 12.7, 4.9 Hz, 1H, biotin-CH), 2.73 (d, J = 12.7 Hz, 1H, biotin-CH), 2.22 (td, J = 7.1, 3.4 Hz, 2H, biotin-CH2), 1.82-1.57 (m, 4H, biotin-CH2) and 1.44 (p, J = 7.5 Hz, 2H, biotin-CH2). 13C NMR (100 MHz, CDCl3): 196.88, 173.41, 169.85, 164.02, 150.45, 96.02, 77.36, 70.68, 70.57, 70.53, 70.34, 70.22, 70.08, 66.91, 61.91, 60.33, 55.55, 40.68, 39.29, 37.26, 35.96, 28.21 and 25.65. ESI high-resolution MS (m/z): [M + H]+ calculated for C27H44N5O10S+: 630.2803, found: 630.2791.
5-(3-(2-(2-(2-(2-(N-biotinyl)aminoethoxy)ethoxy)ethoxy)ethoxy)-1-(N-propargyl)aminopropylidene)-1,3-dimethylpyrimidine-2,4,6-trione (biotin-Dmpt-alkyne, 1)
Compound 2 (50.0 mg, 0.0794 mmol) was suspended in propargylamine (2.00 ml, 31.2 mmol) and stirred at 50 °C for 4 h. The reaction was then concentrated, and the residue was diluted in DCM. The solution was extracted with H2O, and the aqueous layer was back-extracted twice with DCM. The organic layers were combined, dried with MgSO4 and concentrated. The residue was redissolved in a minimal volume of 5% aqueous CH3CN and then purified by preparative HPLC. Fractions were combined, concentrated and lyophilized to afford the product as a pale-yellow powder (14.8 mg, 28% yield). 1H NMR (400 MHz, CDCl3): δ 12.87 (bt, J = 5.5 Hz, 1H, NH), 6.64 (bt, J = 5.6 Hz, 1H, NH), 6.20 (s, 1H, NH), 5.28 (s, 1H, NH), 4.50 (ddt, J = 7.5, 5.0, 1.2 Hz, 1H, biotin-CH), 4.45 (dd, J = 5.5, 2.6 Hz, 2H, NCH2), 4.31 (ddd, J = 7.8, 4.6, 1.4 Hz, 1H, biotin-CH), 3.86 (t, J = 5.5 Hz, 2H, OCH2), 3.67-3.53 (m, 14H, OCH2), 3.51-3.40 (m, 4H, OCH2), 3.33 (s, 3H, NCH3), 3.29 (s, 3H, NCH3), 3.14 (td, J = 7.4, 4.6 Hz, 1H, biotin-CH), 2.90 (dd, J = 12.8, 4.9 Hz, 1H, biotin-CH), 2.74 (d, J = 12.8 Hz, 1H, biotin-CH), 2.41 (t, J = 2.6 Hz, 1H, CCH), 2.22 (td, J = 7.3, 2.2 Hz, 2H, biotin-CH2), 1.78-1.58 (m, 4H, biotin-CH2) and 1.44 (q, J = 7.6 Hz, 2H, biotin-CH2). 13C NMR (100 MHz, CDCl3): δ 175.65, 173.30, 166.84, 163.79, 162.63, 151.35, 90.31, 77.86, 77.36, 73.87, 70.64, 70.62, 70.56, 70.52, 70.25, 70.12, 69.91, 61.88, 60.28, 55.56, 40.68, 39.28, 36.02, 33.97, 31.14, 28.25, 28.22, 28.14, 27.91 and 25.67. ESI high-resolution MS (m/z): [M + H]+ calculated for C30H47N6O9S+: 667.3120, found: 667.3133.
Antibodies
Primary antibodies were obtained from Cell Signaling Technology and secondary antibodies were obtained from Thermo Fisher Scientific unless otherwise specified. For western blotting, the following primary antibodies were used at 1:1,000 dilution unless otherwise specified: anti-OGT (Proteintech, 11576-2-AP), anti-HA (C29F4), anti-Nup62 (BD Biosciences, 610498), anti-HCFC1 (Bethyl Laboratories, A301-399A-M), anti-SET1A (61702S), anti-RBBP5 (13171S), anti-BAP1 (rabbit, Bethyl Laboratories, A302-243A-M), anti-WDR5 (13105S), anti-FoxK1 (Bethyl Laboratories, A301-728A-M), anti-FoxK2 (12008S), anti-liprin-α1 (Proteintech, 14175-1-AP), anti-BAP1 (mouse, Santa Cruz Biotechnology, sc-28383), anti-O-GlcNAc (RL-2, Abcam, ab2739, 1:500) or anti-α-tubulin (MilliporeSigma, T9026, 1:3,000 or Thermo Fisher Scientific, 62204, 1:1,000). The following secondary antibodies were used at 1:10,000 dilution unless otherwise specified: anti-mouse AlexaFluor 680 conjugate (A21057), anti-rabbit AlexaFluor 680 (A21109), anti-rabbit AlexaFluor 790 (A11369) or anti-mouse IgG DyLight 800 (A11357).
Data analysis
All data analysis was performed using the Python language (Python Software Foundation, Python Language Reference, version 3.7.3; available at http://www.python.org) with all Anaconda packages installed (Anaconda Software Distribution, version 4.8.3; available at https://anaconda.com) unless otherwise noted.
Statistical analyses and reproducibility
All MS data were filtered to a 1% false discovery rate on PSMs using the Percolator or target decoy algorithms in Proteome Discoverer (Thermo Fisher Scientific, version 2.4.0.305). Enrichment ratio analysis and spectral counting were analyzed by the limma-moderated t-test and inverted beta-binomial tests, respectively. Distances between nodes within the network were analyzed by a two-sided Fisher’s exact test using Prism (GraphPad, version 9.3.0). When quantified, western blotting of two conditions was analyzed using a paired two-sided Student’s t-test and multiple-hypothesis-testing correction was performed with the Holm–Bonferroni method. A P value (adjusted or corrected when appropriate) of less than 0.05 was considered significant unless otherwise noted. Further details of statistical analyses, including corrections for multiple-hypothesis testing and network pathway enrichment, are provided in subsequent sections. In cases where experimental data were used to validate a model system (for example, OGT induction by Dox), at least two independent experiments were conducted during the development of the model system with similar results. Genotyping of OGT–FH mice was performed for every mouse in the colony (for example, Extended Data Fig. 7a) and, thus, before all experiments on mouse tissue. In cases where experimental data were used to verify interactions (for example, OGT coimmunoprecipitations), at least two independent experiments were conducted with western blotting during optimization followed by identical experiments using proteomics as the readout for interactors. For O-GlcNAcomics experiments, two biological replicates were split into separate LC–MS2 runs (Extended Data Fig. 1e). In experiments with biological stimuli (for example, O-GlcNAc levels or OGT tagging in cells and tissue), n = 3 independent experiments were conducted with similar results unless otherwise noted.
Generation of HEK293T O-GlcNAc network
Generation of the OGT–FH cell line
Lentiviral plasmid construction
All primers were purchased from Integrated DNA Technologies and all molecular biology supplies were purchased from New England Biolabs, unless otherwise indicated. The OGT lentiviral vector was produced by first cloning OGT with C-terminal FLAG and HA tags into the pCMV6 entry vector (Origene) using the NEBuilder HiFi DNA assembly and Q5 site-directed mutagenesis kits according to the manufacturer’s protocols. Both the tagged form of OGT (OGT–FH) and pTRIPZ vector were linearized by PCR amplification using the Q5 HotStart high-fidelity 2× master mix. Fragments were gel-purified using the Zymoclean gel DNA recovery kit (Zymo Research) and assembled with the NEBuilder HiFi DNA assembly kit according to the manufacturer’s protocol. The ligated plasmid was transformed into stable competent cells (New England Biolabs) and individual clones were picked and sequenced (Laragen) to ensure proper incorporation and no recombination. Bacterial cultures were stored as 50% glycerol stocks at −80 °C and 100-ml maxiprep cultures of the plasmid-expressing bacteria were grown and purified by ZymoPure plasmid kits (Zymo Research). Each purification of the pTRIPZ-OGT–FH plasmid was resolved by 1% agarose gel to ensure no recombination occurred.
Cell culture and lentivirus production
The parental HEK293T cell line was obtained from American Type Culture Collection. All cell culture reagents were purchased from Thermo Fisher Scientific unless otherwise indicated. HEK293T cells were cultured in DMEM supplemented with high glucose, GlutaMAX, 10% FBS and 1× penicillin–streptomycin (P/S) (complete DMEM). Transfection and cell culture reagents were purchased from Thermo Fisher Scientific unless otherwise indicated. The lentiviral packaging plasmids pMD2.G and psPAX2 (D. Trono, Addgene, 12259 and 12260) were obtained from Addgene. Four 15-cm culture plates containing HEK293T cells at 80% confluency were each transfected with 8 μg pTRIPZ-OGT–FH, 6 μg psPAX2 and 2 μg pMD2.G using Lipofectamine 3000 at a 2:1 lipid-to-DNA ratio in OptiMEM according to the manufacturer’s protocol. The medium was replaced at 6 h after transfection with complete DMEM containing 2% FBS. Medium was isolated at 48 h and 72 h after transfection and lentiviral particles were concentrated using 15-ml Amicon concentrator tubes (100-kDa molecular weight cutoff (MWCO); MilliporeSigma). Lentiviral concentrate was used immediately.
Cell line generation
To a six-well plate containing HEK293T cells at 80% confluency in complete DMEM with 8 μg ml−1 hexadimethrine bromide, 0–50 μl of lentiviral concentrate was added. Medium was replaced with complete DMEM after 24 h. After 48 h, cells were passaged 1:4 into fresh six-well plates with complete DMEM containing 2 μg ml−1 puromycin (Thermo Fisher Scientific). After selection for 2 weeks, surviving cells were then split for clonal dilution in 96-well plates. After 2–3 weeks, single clones were isolated and expanded. Clones were then tested for induction of OGT–FH expression by incubating with complete DMEM containing 0.1 μg ml−1 Dox hyclate for 24 h followed by western blotting. A single clone for the HEK293T-iOGT–FH (OGT–FH) cell line was then chosen and used for subsequent experiments.
Identification of high-confidence OGT interactors (HEK293T cells)
SILAC prelabeling
To perform quantitative proteomics, OGT–FH cells were cultured for at least six passages in SILAC medium. SILAC medium was produced using SILAC DMEM Flex medium supplemented with high glucose, GlutaMAX, 10% dialyzed FBS, 1× P/S and either 0.80 mM L-lysine and 0.40 mM L-arginine (MilliporeSigma; light, K0, R0) or 0.80 mM [U-13C6,15N2] L-lysine and 0.40 mM [U-13C6] L-arginine (Cambridge Isotope Laboratories, heavy, K8, R6).
TAP
Two 15-cm plates of heavy-labeled HEK293T-iOGT–FH cells at 60% confluency were treated with freshly made 0.04 μg ml−1 Dox hyclate in H2O for 24 h. As a control, light-labeled HEK293T-iOGT–FH cells were vehicle treated for the same time period. The cells were collected with a scraper and homogenized on ice in 10 ml of 50 mM Tris-HCl pH 7.4, 150 mM NaCl, 0.5% NP-40 and 1× cOmplete EDTA-free protease inhibitor cocktail (MilliporeSigma) using 30–40 strokes with a 20-ml Potter-Elvehjem tissue grinder (Wheaton) followed by centrifugation at 15,000g for 10 min at 4 °C. The protein concentration of the clarified supernatant was measured using the Pierce BCA protein assay kit (Thermo Fisher Scientific). For each replicate, approximately 15 mg of cell lysate from heavy-labeled and light-labeled cells were subjected to immunoprecipitation with a 200-μl settled volume of anti-FLAG M2 affinity agarose gel (MilliporeSigma) overnight at 4 °C. In all cases, 1% of the sample (approximately 150 μg) was saved as input for total protein staining and western blotting. The beads were washed three times 5 min in wash buffer (50 mM Tris-HCl pH 7.4, 150 mM NaCl and 0.1% NP-40) and eluted with 150 μg ml−1 of 3×FLAG peptide (MilliporeSigma) in wash buffer, followed by immunoprecipitation with a 100-μl settled volume of anti-HA agarose beads (Thermo Fisher Scientific) for 6 h at 4 °C. The beads were washed with wash buffer and eluted with 3 M NaSCN. The eluates were acetone precipitated, redissolved in 1× SDS loading buffer (50 mM Tris-Cl pH 6.8, 100 mM dithiothreitol (DTT), 2% SDS, 0.1% bromophenol blue and 10% glycerol) and subjected to in-gel digestion.
In-gel digestion
The eluates from heavy-labeled and light-labeled samples were mixed immediately before loading onto a NuPAGE 4–12% Bis–Tris protein gel (Thermo Fisher Scientific). The gel was stained and visualized using Imperial protein staining reagent (Thermo Fisher Scientific) and destained in doubly deionized H2O (ddH2O). The gel lane was cut into 30–40 gel slices. Each gel piece was reduced with 10 mM DTT for 1 h at 37 °C and then alkylated with 50 mM iodoacetamide for 45 min at room temperature protected from light. The slices were dried and incubated with 20 ng μl−1 trypsin protease (Pierce MS grade, Thermo Fisher Scientific) solution in 100 mM NH4HCO3 pH 8.0 and 1 mM CaCl2 overnight at 37 °C. Peptides were recovered using sequential washes with 100 mM NH4HCO3, 1:1 (v/v) 100 mM NH4HCO3 and CH3CN and 5% formic acid. The recovered peptides were dried and desalted with a C18 ZipTip pipette tip (MilliporeSigma). Finally, the desalted peptides were resuspended in 0.2% formic acid.
LC–MS2
Samples were subjected to LC–MS2 analysis on a nanoflow LC system, EASY-nLC II (Thermo Fisher Scientific), coupled to a LTQ Orbitrap Elite MS instrument (Thermo Fisher Scientific). Solvent A consisted of 97.8% H2O, 2% CH3CN and 0.2% formic acid and solvent B consisted of 19.8% H2O, 80% CH3CN and 0.2% formic acid. Samples were directly loaded onto a 16-cm analytical HPLC column (inner diameter: 50 µm) packed in-house with ReproSil-Pur C18AQ 3-µm resin (pore size: 120 Å; Dr. Maisch). The column was heated to 60 °C during separation. The peptides were separated with a 60-min gradient at a flow rate of 220 nl min−1 with the following gradient: 2–30% solvent B (60 min), 30–100% B (1 min) and 100% B (9 min). Eluted peptides were then ionized using a Nanospray Flex ion source (Thermo Fisher Scientific) and introduced into the MS instrument. The LTQ Orbitrap Elite was operated in data-dependent mode, automatically alternating between a full scan (400–1,600 m/z, 120,000 resolution) in the orbitrap and subsequent MS2 scans of the 20 most abundant peaks in the linear ion trap (Top20 method).
MS data analysis and identification of high-confidence OGT interactors
All raw files were searched using Proteome Discoverer version 2.4.0.305 (Thermo Fisher Scientific) with the Byonic (Protein Metrics) search node version 3.7.4. Search parameter files are provided in Supplementary Data 7 (‘Interactomics_SILAC_ProcessingWorkflow.pdProcessingWF’ and ‘Interactomics_All_ConsensusWorkflow.pdConsensusWF’). Peak lists were searched against the species-specific complete UniProtKB databases and the UniprotKB/Swiss-Prot databases (Homo sapiens, retrieved January 25, 2020) supplemented with a database of frequently observed contaminants (245 sequences). The following parameters were set: MS1 tolerance of 5 ppm, MS2 tolerance of 10 ppm or 0.5 Da for orbitrap and ion trap spectra, respectively, minimum peptide length of 6 aa, a maximum of three missed cleavages, carbamidomethylation of cysteine residues as a fixed modification, acetylation of the protein N terminus and oxidation of methionine (common 2) as variable modifications and K0/8 and R0/6 as the heavy/light (H/L) SILAC isotopes. Data were filtered to a 1% false discovery rate on PSMs using the Percolator algorithm in Proteome Discoverer. Both SILAC ratios and H/L spectral counts were extracted from the Proteome Discoverer output for statistical analysis. We chose to combine both SILAC ratios and H/L spectral counting to provide a more robust statistical accounting of interactors in an experimental design, where true interactors would be expected to have a SILAC ratio of infinity (that is, the interactor would only be present in the OGT pulldown sample and not the sham control). Here, the use of spectral counting is necessary to perform meaningful statistical tests under ideal conditions. We chose to also include significant SILAC ratios because, in practice, many proteins are sticky and are in fact present in both the OGT and the sham pulldowns. In this scenario, where large variation in spectral counts is expected, statistical testing on SILAC ratios provides a much more robust and quantitative comparison. This approach has been explored and discussed in depth elsewhere41. To analyze SILAC ratios, the limma-moderated t-test42 (package: limma version 3.42.2) was used and multiple-hypothesis-testing correction was performed with the Benjamini–Hochberg method by default. To analyze count data, the inverted beta-binomial test43 (package, ibb version 13.06) was used. Multiple-hypothesis-testing correction for the inverted beta-binomial test was performed using the Benjamini–Hochberg method. For both statistical tests, the R language (R Development Core Team, version 3.6.1; available at https://www.R-project.org) was applied using the rpy2 (Laurent Gautier, version 2.9.4) Python package. ‘High-confidence’ OGT interactors were defined as hits that generated an adjusted P value of <0.05 from either enrichment ratio analysis or spectral counting.
Western blotting
OGT–FH cell lysates were prepared as described above. Cell lysates (~30 µg) were diluted with 4× SDS–PAGE buffer (200 mM Tris-HCl pH 6.8, 400 mM DTT, 8% SDS, 0.4% bromophenol blue and 40% glycerol), resolved on a NuPAGE 4–12% Bis–Tris protein gel (Thermo Fisher Scientific) and transferred to Immobilion-FL PVDF membrane (MilliporeSigma). Blots were blocked for 1 h at room temperature with Odyssey blocking buffer (Tris-buffered saline, TBS; LI-COR Biosciences) before probing overnight at 4 °C with the appropriate primary antibody. Blots were rinsed three times with TBS and 0.1% Tween-20 (TBST) and then probed for 1 h at room temperature with the appropriate secondary antibody in blocking buffer. Blots were washed three times for 5 min with TBST and then imaged using an Odyssey Infrared Imaging System (LI-COR Biosciences). Images were processed using ImageStudio version 5.2 (LI-COR Biosciences).
Identification of O-GlcNAc modification events (HEK293T cells)
Sample preparation
Two 15-cm plates of OGT–FH cells at 60% confluency were lysed in 5 ml of 1% SDS, 50 mM Tris-HCl pH 7.4, 150 mM NaCl, 1× cOmplete EDTA-free protease inhibitor cocktail and 10 μM Thiamet G and probe sonicated to shear DNA. The protein concentration was measured as described above.
CL and digestion
OGT–FH cell lysate (5 mg) was diluted to 2.5 mg ml−1 with 1% SDS and 20 mM HEPES pH 7.6 in a 15-ml conical tube. Samples were then reduced using 20 mM DTT (final concentration) for 45 min at 60 °C, cooled to room temperature and alkylated with 100 mM iodoacetamide (final concentration) for 45 min at room temperature protected from light. Protein was then precipitated by the addition of 3 volumes of CH3OH, 1 volume of CHCl3 and 2.25 volumes of H2O. The precipitated protein was pelleted by centrifugation at 7,068g for 35 min at 4 °C. Solvents were removed and the pellet was washed twice with three volumes of CH3OH before redissolving at 5 mg ml−1 with 1 ml of 1% SDS and 20 mM HEPES pH 7.9 at 95 °C for 5 min. Samples were then diluted with 2 ml of 2.5× GalT labeling buffer (50 mM HEPES pH 7.9, 125 mM NaCl and 5% IGEPAL CA-630), 1.2 ml of ddH2O and 275 μl of 100 mM MnCl2. The mixture was chilled on ice for 5 min and then 250 μl of 0.1 mM UDP-GalNAz in 10 mM HEPES pH 7.9, 250 μl of 2 mg ml−1 GalT-Y289L (ref. 19), 50 μL of 500 kU per ml PNGase F (New England Biolabs, 5 U per μg of protein) and 62.5 μl of lambda protein phosphatase (New England Biolabs, 5 U per μg of protein) were added with gentle mixing after each addition. Samples were rotated end-over-end overnight at 4 °C and were then precipitated as described above. The protein pellet was redissolved at 5 mg ml−1 in 1 ml of 1% SDS and 20 mM HEPES pH 7.6 at 95 °C for 5 min. The resulting solution was diluted with 2.85 ml of ddH2O and 0.15 ml of 200 mM HEPES pH 7.6. In a separate tube, a 5× copper-catalyzed azide–alkyne cycloaddition (CuAAC) mix was prepared by sequential addition of 885 μl of ddH2O, 5 μl of 100 mM BTTAA (Click Chemistry Tools) in DMSO, 100 μl of 50 mM CuSO4 (freshly prepared) and 10 μl of 50 mM 1 in DMSO. The 5× mix was then added directly to the protein sample and mixed gently before adding 100 μl of 100 mM TCEP (freshly prepared). The CuAAC reaction was allowed to proceed with end-over-end rotation for 1 h at room temperature before the proteins were precipitated as described above. The pellet was resuspended in 5 ml of 50 mM HEPES pH 7.6 and 10 mM EDTA to chelate residual Cu2+. To the resuspension, 200 μg of trypsin protease (1:25 w/w) was added and the sample was rotated end-over-end for 20 h at 37 °C. Two 15-ml Amicon concentrator tubes (10-kDa MWCO) were rinsed twice with 5 ml of 50% CH3OH and twice with 5 ml of ddH2O by centrifugation at 4,000g for 5 min. Tryptic solutions were then centrifuged in the rinsed concentrator tubes at 4,000g for 20 min. The remaining residue was rinsed twice with 2.5 ml of H2O by centrifugation and the flowthrough fractions were combined and diluted twofold with PBS.
O-GlcNAc peptide enrichment
Before use, a 900-μl spin filter (Thermo Fisher Scientific) was prerinsed twice with 1:1 CH3OH and ddH2O. Then, 500 μl of a 50% slurry of high-capacity Neutravidin agarose (Thermo Fisher Scientific) was added and rinsed twice with 0.5 ml of PBS to remove storage buffer. The washed agarose resin was then added to the filtrate sample and rotated end-over-end for 1 h at room temperature. The resin was pelleted by centrifugation at 500g for 5 min and then transferred to a fresh, prerinsed 900-μl spin filter. Beads were rinsed by centrifugation with the following solutions: five times with 0.5 ml of PBS, five times with 0.5 ml of 1 M NaCl in PBS and five times with 0.5 ml of PBS. Captured peptides were then eluted twice with 0.75 ml of 2% v/v NH2OH with end-over-end rotation for 1.5 h at 37 °C. Elution fractions were combined and concentrated to dryness with vacuum centrifugation before desalting as described for affinity-purified peptides.
LC–MS2
Chemoenzymatically labeled, enriched and digested samples were analyzed on a nanoflow LC system, EASY-nLC 1000 (Thermo Fisher Scientific), coupled to an Orbitrap Fusion Tribrid MS instrument (Thermo Fisher Scientific) equipped with a Nanospray Flex ion source. Sample loading and LC separation were performed identically to the Q Exactive HF Orbitrap LC method outlined for the HEK293T OGT interactor samples. Full spectra were acquired over 350–1,800 m/z in the orbitrap operating at 120,000 resolution at 200 m/z. Automatic gain control (AGC) was set to accumulate 50,000 ions, with a maximum injection time of 50 ms. Data-dependent MS2 analysis was performed using a top-speed approach (cycle time of 5 s) with multiple fragmentation methods (see below). Dynamic exclusion was set to exclude features after one time for 15 s with exclude isotopes turned on. The normalized collision energy was optimized at 28% for HCD fragmentation. The intensity threshold for fragmentation was set to 25,000. HCD–MS2 fragmentation spectra were collected in the orbitrap operating at 30,000 resolution at 200 m/z. Either ETD–MS2 or EThcD–MS2 fragmentation was then performed (on separate instrument runs) for precursor MS1 ions whose corresponding HCD–MS2 spectra contained a fragment of mass 300.1303 m/z (15-ppm tolerance, within the top 30 ions). AGC was set to 50,000 with a maximum injection time set at 500 ms for the orbitrap operating at 30,000 resolution. ETD–MS2 reaction time was charge-dependent and supplemental activation energy for EThcD–MS2 was set to 20%.
MS data analysis and identification of O-GlcNAc modifications
MS data analysis was conducted as described above for OGT interactor samples with some modifications. The Swiss-Prot searches were used specifically to match O-GlcNAc sites with known PTMs in the PhosphoSitePlus database. The chemoenzymatically labeled O-GlcNAc residue was added as a variable modification (common 2). Localization of O-GlcNAc sites was performed using the ptmRS node. For HCD spectra, equal localization probabilities were assigned to every serine/threonine because of neutral loss of the glycan. Search parameter files are provided in Supplementary Data 7 (‘Glycomics_All_ProcessingWorkflow.pdProcessingWF’ and ‘Glycomics_All_ConsensusWorkflow.pdConsensusWF’). Output files from Proteome Discoverer were then subjected to the ‘sites and regions’ workflow and algorithm. This workflow provides a valuable means to count the total number of events identified in the most parsimonious way possible.
Previous efforts to enumerate the total number of O-GlcNAcylation events have been inconsistent and at times inaccurate. This variability has arisen in part because of different approaches by which MS analytical software handle the counting of ‘nonredundant’ events. For example, in Proteome Discoverer, a PeptideGroups file is created where PSMs are combined into peptide ‘groups’ that account for different charge states. However, related peptides resulting from missed proteolytic cleavage events or amino acid oxidation are grouped separately, resulting in a potential overcounting of PTM events. Similarly, MS fragmentation by HCD and ET(hc)D across different MS runs will generate both unlocalized and localized PSMs for the same O-GlcNAc event that can also lead to overcounting. Our ‘sites and regions’ algorithm counts such redundant events only once. As a direct comparison, we would have reported 3,144 O-GlcNAcylation events across 3,012 unique O-GlcNAcylated peptides in our combined mouse brain and liver dataset using the PeptideGroups file from Proteome Discoverer instead of the more accurate 2,785 sites and regions obtained using our ‘sites and regions’ algorithm. The general theoretical workflow for calling sites and regions is described in detail in Extended Data Fig. 2. Original Python code is available in Supplementary Data 6 and in Github (https://github.com/michael-sweredoski/O-GlcNAc-Sites-And-Regions)44. The algorithm yields two files, maxparcon.txt (number of sites with maximum parsimony; modified to Supplementary Table 2 for 293T cells) and bestms2.txt (best spectra for every possible site; Supplementary Table 24 for 293T cells).
Once properly enumerated, frequency and enrichment plots for O-GlcNAc recognition sequences were generated using WebLogo45 (version 2.8.2; available at https://weblogo.berkeley.edu/). Bar graphs quantifying O-GlcNAc events were generated in GraphPad Prism (version 9.3.0). Identified O-GlcNAc sites on HCFC1 were compared to the O-GlcNAc database21 (version 2.0; available at https://www.oglcnac.mcw.edu/, accessed December 18, 2024).
Generation of O-GlcNAc network (HEK293T cells)
Network assembly and partitioning
The input for the network analysis included (1) OGT interactors and substrates identified as above and (2) known interactions from PPI databases. Two PPI databases were used: BioGRID46 (H. sapiens-specific dataset, downloaded February 25, 2020) and IntAct47 (human–human PPIs, downloaded February 25, 2020 and filtered for only multivalidated interactions). Interactions between an OGT interactor and other interactors or substrates were filtered to generate the edges of the O-GlcNAc network. Edges from individual OGT interactor ‘baits’ with more than 30 connections to other protein ‘prey’ were removed to avoid biasing during community clustering. Examples of these proteins include ribosomal proteins, hnRNPs and heat-shock proteins. Overall, these eliminated edges essentially coincided with interactions from proteins highly represented within the CRAPome48, implying a high rate of nonspecific interactions. However, edges generated in the reverse direction (for example, from an interactor/substrate bait to a ribosomal protein prey) were retained. The nodes and edges were then used to build a network of OGT interactors and substrates, which was visualized and analyzed using Cytoscape (version 3.7.2)49. To partition the network, we used the community cluster GLay algorithm23 in the Cytoscape plugin clusterMaker2 (version 1.3.1)50, which is a community clustering algorithm that implements the Girvan–Newman fast greedy algorithm51.
Network functional analyses
Information from the UniProt Knowledgebase52 and PhosphoSitePlus53 database used for the integrated Cytoscape tables (Supplementary Data 2–5) was retrieved on April 22, 2020 and April 28, 2020, respectively. The protein names from the whole network or from each cluster were submitted for statistical overrepresentation testing using g:Profiler54. Only GO55 terms with a P value < 0.01 with the g:SCS significance threshold were used and electronic GO annotations were excluded. The complete datasets along with all submission parameters are provided in Supplementary Data 8. To group GO terms and appropriately name clusters, the Cytoscape plugin ClueGO (version 2.5.6)56 was used. For each cluster, the proteins were submitted to ClueGO analysis with GO biological processes (downloaded February 17, 2020), again with all electronic annotations excluded. GO term fusion was used and only pathways with a P value < 0.01 were considered significant. Network specificity was set to medium with all dependent parameters left as default. GO term grouping was used. Enrichment was also measured using the same workflow for KEGG pathways57, Reactome pathways58, WikiPathways59 and CORUM protein complexes60. Hierarchical clustering and generation of figures were done with Seaborn using the average linkage method.
BAP1 network hub studies
Manipulation of BAP1 expression levels
Coimmunoprecipitation of OGT interactors upon BAP1 siRNA knockdown
BAP1 RNA knockdown was performed using a set of three unique 27-mer siRNA duplexes targeting BAP1 (Origene, SR305435) using siTrans 1.0 (Origene). Briefly, OGT–FH cells in six-well plates at 70% confluency were transfected with 20 nmol of siRNA mixture according to the manufacturer’s protocol. Cells were allowed to grow for 3 days before harvesting as described for cell lysates. Knockdown was validated by western blotting. Coimmunoprecipitation experiments were carried out in the same manner as described for TAP–MS sample preparation.
CRISPR–Cas9-mediated BAP1 knockout
BAP1 CRISPR–Cas9 knockout Plasmid (h; sc-400232), BAP1 HDR Plasmid (h; sc-400232-HDR) and control CRISPR–Cas9 plasmid (sc-418922) were purchased from Santa Cruz Biotechnology (Dallas, TX) and prepared per the manufacturer’s instructions. HEK293T cells were maintained as described above and plated in a 12-well plate. At 60–70% confluency, cells were then transfected with 1 µg of BAP1-knockout and HDR plasmids (0.5 µg each) or the control plasmid using Lipofectamine 3000 (Thermo Fisher Scientific) at a 2:1 lipid-to-DNA ratio in OptiMEM according to the manufacturer’s protocol. Once confluent, the cells were split 1:1 into medium containing 10 μg ml−1 puromycin (Thermo Fisher Scientific) or normal medium for the control. Cells were then selected for 2 weeks, replacing the medium or splitting as necessary. BAP1 knockout was confirmed by western blotting as described above.
Quantification of global O-GlcNAc site occupancies
Sample preparation for O-GlcNAcomics analysis
BAP1 knockout and HEK293T WT cells were grown in three 10-cm plates, lysed and subjected to CL and enrichment as described above without modification. Dried peptides after enrichment and elution were resuspended in 50 μl of 100 mM TEAB pH 8.5 and labeled using six tags (n = 3 for WT and n = 3 for BAP1 knockout) from the TMT10plex isobaric labeling kit (Thermo Fisher Scientific) per the manufacturer’s instructions. Labeled peptides were then combined, desalted and dried as before and then resuspended in 10 μl of 0.2% formic acid.
LC–MS2 to quantify O-GlcNAc modifications
Quantitative O-GlcNAcomics samples were analyzed on a nanoflow LC system, EASY-nLC 1000, coupled to an Orbitrap Fusion Tribrid MS instrument equipped with a Nanospray Flex ion source. For each sample, ~4 μg of peptides were loaded onto a 25-cm Aurora column (inner diameter: 75 µm) packed in-house with 1.6-µm C18 resin (Ion Opticks). Solvents A and B were the same as described above for other MS samples. Peptides were separated over 135 min at a flow rate of 350 nl min−1 with the following gradient: 2–6% solvent B (7.5 min), 6–25% B (82.5 min), 25–40% B (30 min), 40–100% B (1 min) and 100% B (14 min). MS1 spectra were acquired at 120,000 resolution with a scan range of 350–1,800 m/z. AGC was set to accumulate 50,000 ions, with a maximum injection time of 50 ms. Data-dependent MS2 analysis, either top speed (5 s) or top 10, was then performed, in which features were filtered for monoisotopic peaks with a charge state of 3–8 and a minimum intensity of 25,000, with dynamic exclusion set to exclude features after one time for 15 s with a 10-ppm mass tolerance and exclude isotopes turned on. HCD–MS2 fragmentation was performed with normalized collision energy of 28% after quadrupole isolation of features using an isolation window of 1.6 m/z, an AGC target of 100,000 and a maximum injection time of 100 ms. MS2 scans were acquired at 30,000 resolution in centroid mode with the scan range set to automatic. Detection of at least one fragment with 732.3726, 547.3037 or 529.2931 m/z in the top 30 MS2 ions (15-ppm tolerance) triggered three additional data-dependent scans (ddMSn) as shown in Extended Data Fig. 6b:
ddMS3-OT–HCD: SPS–MS3 fragmentation by HCD within the orbitrap on the top five HCD–MS2 precursor ions using an isolation window of 1.3 m/z, 65% normalized collision energy, resolution of 50,000, AGC target of 50,000, an automatic scan range and a maximum injection time of 250 ms, along with precursor selection range, ion exclusion and isobaric tag loss exclusion;
ddMS2-IT–ETD: MS2 fragmentation by ETD within the ion trap on the MS1 precursor ion using an isolation window of 1.6 m/z, calibrated charge-dependent ETD parameters, rapid scan mode, automatic scan and mass range, an AGC target of 50,000 and a maximum injection time of 100 ms;
ddMS2-IT–EThcD: MS2 fragmentation by EThcD within the ion trap on the MS1 precursor ion using the same parameters as ddMS2-IT–ETD except with 20% supplemental activation energy.
Sample preparation for quantitative proteomics
Samples of BAP1 knockout or WT total protein (100 μg, n = 3 for each) were removed from the lysates used for CL experiments and precipitated using the CH3OH–CHCl3–H2O method as described above except that the samples were centrifuged at 21,130g for 5 min to pellet proteins. The pellet was resuspended in 100 μl of 100 mM TEAB pH 8.5 and 3 μg of trypsin protease (1:33 w/w enzyme and protein) was then added. The samples were digested for 20 h at 37 °C and then labeled with the TMTsixplex isobaric labeling kit (Thermo Fisher Scientific) per the manufacturer’s instructions. After labeling, peptides from all groups were combined and concentrated to dryness by vacuum centrifuge. Peptides were then resuspended in 100 μl of 10 mM NH4OH and subjected to offline high-pH, reversed-phase fractionation using an Agilent Extend-C18 (2.1 × 150 mm, 5 μm) on an Agilent 1100 HPLC operating at 0.2 ml min−1. Solvent A consisted of 10 mM NH4OH and solvent B consisted of 10 mM NH4OH in 90% CH3CN. The combined peptide sample (600 μg) was injected onto the column and separated using the following gradient: 1% solvent B (4 min), 1–30% B (50 min), 30–60% B (4 min), 60–70% B (2 min) and 70–90% B (5 min). A total of 64 1-min fractions were collected in 96 deep-well plates, concatenated to 16 samples and dried by vacuum centrifugation. Lastly, each fraction was resuspended in 10 µl of 0.2% formic acid by bath sonication.
LC–MS2 to quantify total proteome
Samples were subjected to LC–MS2 on a nanoflow LC system, EASY-nLC 1000, coupled to an Orbitrap Eclipse Tribrid MS instrument. From each resuspended fraction, 3 μl was loaded onto a monolithic column (Capillary EX-Nano MonoCap C18 high-resolution 2000, 0.1 × 2,000 mm; Merck) fitted with a silica coated PicoTip emitter (New Objective, FS360-20-10-D). Solvents A and B were the same as described above. Peptides were separated over 180 min at a flow rate of 500 nl min−1 with the following gradient: 2–6% solvent B (10 min), 6–40% B (140 min), 40–98% B (1 min) and 98% B (29 min). MS1 spectra were acquired in the orbitrap at 120,000 resolution with a scan range of 400–1,500 m/z, an AGC target of 400,000 and the maximum injection time set automatically. Features were filtered for monoisotopic peaks with a charge state of 2–5 and a minimum intensity of 5,000, with dynamic exclusion set to exclude features after one time for 60 s with a 10-ppm mass tolerance. Data-dependent MS2 analysis was performed using a top-speed approach (cycle time of 5 s). Collision-induced dissociation fragmentation was performed with normalized collision energy of 35% and an activation time of 10 ms with an activation Q of 0.25 after quadrupole isolation of features using an isolation window of 0.7 m/z. Scans were acquired in the ion trap using automatic scan range, a maximum injection time of 250 ms and rapid scan mode. HCD–SPS–MS3 scans were also performed in the orbitrap on the top ten MS2 precursors with an isolation window of 0.7 m/z, 55% normalized collision energy, resolution of 60,000, AGC target of 150,000, scan range of 100–500 m/z and a maximum injection time of 118 ms along with precursor selection range, ion exclusion and isobaric tag loss exclusion.
MS data analysis and mapping to OGT–FH network
Data analysis for O-GlcNAcomic and total proteome samples was conducted as described above with some modifications. Swiss-Prot-only searches were used to facilitate matching of O-GlcNAcylated peptides to their parent protein expression. TMT labeling of peptide N termini and lysines was set as fixed modifications. The quantitative proteomics data were filtered to a 1% false discovery rate on PSMs using the target decoy algorithm in Proteome Discoverer. Search parameter files are provided in Supplementary Data 7 (files beginning with ‘QuantitativeGlycomics’ or ‘QuantitativeProteomics’ for the quantitative O-GlcNAcomics and quantitative proteomics, respectively). The maxparcon.txt and bestms2.txt output data from the ‘sites and regions’ algorithm are provided in Supplementary Tables 10 and 25, respectively. Quantitative data were mapped to the existing OGT–FH network in 293T cells (Supplementary Data 4). Total-protein-normalized changes in O-GlcNAcylation events (that is, changes in individual site occupancies) were reported as log2 fold change (part 5 of ‘Sites and Regions Analysis.ipynb’; https://github.com/michael-sweredoski/O-GlcNAc-Sites-And-Regions).
Generation of in vivo O-GlcNAc networks
Generation of the OGT–FH genetic mouse model
General information
All mouse procedures were performed in accordance with protocols approved and guidelines set by the Caltech Institutional Animal Care and Use Committee (protocol IA21-1435). All mice were obtained from Jackson Laboratory or Charles River. The OGT–FH mouse model was maintained on a C57BL/6J background. Animals were housed with a 12-h dark–light cycle at ambient temperature and humidity in Caltech facilities.
Design and validation of the single guide RNA (sgRNA) targeting sequence
sgRNA candidates were designed using the CRISPR design (www.genome-engineering.org) and CHOPCHOP programs (https://chopchop.cbu.uib.no/). Genomic DNA was isolated from the tail tip of a WT C57BL/6J mouse using the DNEasy blood and tissue kit (Qiagen). The genomic C-terminal region of Ogt was amplified using the Q5 HotStart high-fidelity 2× master mix (New England Biolabs). The pCAG-EGxxFP plasmid (Addgene, 50716; M. Ikawa) was cut with BamHI and SalI and the construct was assembled using the NEBuilder HiFi assembly kit. sgRNA candidates were cloned into the pX330 plasmid (Addgene, 42230; F. Zhang) according to the published protocol. The two plasmids were then transfected into HEK293T cells at 80% confluency in a six-well plate using Lipofectamine 3000 at a 2:1 lipid-to-DNA ratio according to the manufacturer’s protocol. After 48 h, cells were observed using an LSM 710 confocal microscope (Carl Zeiss) for GFP+ cells, indicating active sgRNA. The most active sgRNA found within 20 bp of the Ogt stop codon was chosen for in vivo genomic editing (5′-CCTGAATAAAGACTGCGCAC-3′).
Preparation of materials for zygote injection
The sgRNA was amplified from the pX330 plasmid and then transcribed and purified using the MEGAshortscript T7 transcription and cleanup kits (Thermo Fisher Scientific), respectively. For homology-directed recombination, a single-stranded oligodeoxynucleotide (ssODN) was synthesized as an Ultramer oligonucleotide (Integrated DNA Technologies) containing an MluI cut site, FLAG tag, BamHI cut site, HA tag and stop codon flanked on either side by 60 nt homologous to the genomic region surrounding the insertion site. The protospacer-adjacent motif was mutated to prevent further nuclease activity after homology-directed recombination.
Zygote injection and validation of knock-in
Zygotes from C57BL/6N mice were produced, collected, cultured and implanted as previously described61. Microinjection of embryos was performed using an inverted microscope (Carl Zeiss) equipped with a micromanipulator (Leica Microsystems), CellTram (Eppendorf) and FemtoJet (Eppendorf). Injections were carried out as previously described62. Solutions containing 2.5 ng μl−1 sgRNA, 10 ng μl−1 ssODN and 5 ng μl−1 Cas9 mRNA (System Biosciences) in 10 mM Tris-HCl pH 8.0 and 0.1 mM EDTA were used for injection into the pronucleus of fertilized mouse zygotes. Following injection, the zygotes were cultured to the two-cell stage and implanted into pseudopregnant foster mothers at 0.5 days after coitum (up to 30 two-cell embryos per recipient). Approximately 19.5 days after implantation, the pups were delivered; then, 3 weeks after birth, the pups were tailed and separated by gender. Offspring were genotyped by restriction enzyme digestion using BamHI after purification of the genomic DNA from tail tips using the DNEasy blood and tissue kit and PCR amplification of the genomic C-terminal region of Ogt. A single heterozygous female was obtained containing the correct on-target inserted sequence. The line was then backcrossed for at least six generations with the C57BL/6J strain before experiments. Subsequent homozygous mice (OGT–FH mice) obtained from this founder exhibited no abnormalities in growth, behavior or breeding and were indistinguishable from WT and heterozygous mice. Western blotting was performed as described above to verify expression of tagged OGT and compare overall O-GlcNAc levels.
Identification of high-confidence OGT interactors (brain and liver tissue)
Sample preparation
To prepare brain and liver tissue lysates, 2-month-old OGT–FH and WT mouse forebrains and livers were freshly dissected and rinsed with ice cold PBS. One brain hemisphere or liver lobe was used per experiment. Lysate from each of OGT–FH and WT organs was isolated and homogenized on ice in 10 ml of 50 mM Tris-HCl pH 7.4, 150 mM NaCl, 0.5% NP-40 and 1× cOmplete EDTA-free protease inhibitor cocktail using 30–40 strokes with a Potter-Elvehjem homogenizer followed by centrifugation at 15,000g for 10 min at 4 °C.
TAP
Lysates (15 mg) from each mouse were subjected to immunoprecipitation using the same procedure outlined above.
Filter-assisted sample preparation and dimethyl labeling
Bead eluates were loaded onto 0.5-ml Amicon Ultra centrifugal filters (10-kDa MWCO; MilliporeSigma) and buffer-exchanged into 100 mM TEAB pH 8.5. The samples were then reduced with 10 mM DTT for 1 h at 37 °C and alkylated with 50 mM iodoacetamide for 45 min at room temperature protected from light. The proteins were digested with 50 ng μl−1 trypsin in 100 mM TEAB at 37 °C overnight. The resulting peptides were recovered by washing the centrifugal filter with 100 mM TEAB three times for a final volume of 100 μl. To each peptide mixture, 4 μl of 4% v/v CH2O (WT) or CD2O (OGT–FH) was added followed by 4 μl of 0.6 M NaBH3CN. Samples were then rotated end-over-end for 1 h at room temperature and quenched with 16 μl of 1% v/v NH3 followed by 8 μl of formic acid. The differentially labeled samples from OGT–FH and WT mice were mixed and desalted by an Agilent 1100 Series HPLC system with a MicroTrap cartridge (Michrom, inner diameter: 1 × 8 mm, bed volume: 5.0 μl, capacity: 20 μg, sample volume: 1–1,000 μl). Solvent A consisted of 99.8% H2O and 0.2% formic acid and solvent B consisted of 99.8% CH3CN, and 0.2% formic acid. The gradient was as follows: 0% solvent B (10 min), 0–85% B (2 min), 85% B (5 min), 85–90% B (1 min), 90% B (5 min), 90–100% B (1 min) and 100% B (6 min). Eluted peptide fractions were collected in 1-ml fractions, lyophilized and resuspended in 0.2% formic acid for LC–MS2 analysis.
LC–MS2
Affinity-purified samples from brain tissue were subjected to LC–MS2 analysis on a nanoflow LC system, EASY-nLC 1200 (Thermo Fisher Scientific), coupled to a Q Exactive HF Orbitrap MS instrument (Thermo Fisher Scientific) equipped with a Nanospray Flex ion source. Samples were directly loaded onto a 20-cm PicoFrit column (inner diameter: 50 µm; New Objective) packed in-house with ReproSil-Pur C18AQ 1.9-µm resin (pore size: 120 Å; Dr. Maisch). The column was heated to 60 °C during separation. The peptides were separated with a 120-min gradient at a flow rate of 220 nl min−1 with the following gradient: 2–6% solvent B (7.5 min), 6–25% B (82.5 min), 25–40% B (30 min) and 100% B (9 min). Solvents A and B were the same as described above. The Q Exactive HF Orbitrap was operated in data-dependent mode. Full scan resolution was set to 60,000 at 200 m/z. Full scan target was 3 × 106 with a maximum injection time of 15 ms. Mass range was set to 300–1,650 m/z. For data-dependent MS2 scans, the loop count was 12, target value was set at 1 × 105 and intensity threshold was kept at 1 × 105. Isolation width was set at 1.2 m/z and a fixed first mass of 100 was used. Normalized collision energy was set at 28%. Peptide match was set to off and isotope exclusion was on.
MS data analysis and identification of high-confidence OGT interactors
MS data analysis and interactor identification were carried out as described above for HEK293T interactome experiments with some modifications. Peak lists were searched against the species-specific complete UniProtKB databases and the UniprotKB/Swiss-Prot databases (Mus musculus, retrieved January 21, 2020) supplemented with a database of frequently observed contaminants (245 sequences). Dimethyl labeling of peptide N termini and lysines were set as fixed modifications. Search parameter files are provided in Supplementary Data 7 (‘Interactomics_Dimethyl_ProcessingWorkflow.pdProcessingWF’ and ‘Interactomics_All_ConsensusWorkflow.pdConsensusWF’). Both dimethyl ratios and H/L spectral counts were extracted from the PD output for statistical analysis as described above. The maxparcon.txt and bestms2.txt output data from the sites and regions algorithm are provided in Supplementary Tables 13, 14, 26 and 27, respectively. Dimethyl ratios and spectral counts were used to identify high-confidence interactors as described above for SILAC-labeled HEK293T proteins. High-confidence OGT interactors were defined as hits that generated an adjusted P value < 0.05 from either enrichment ratio analysis or spectral counting.
Identification of O-GlcNAc modification events (brain and liver tissue)
Sample preparation
Lysates from brain and liver tissue were prepared as described above for HEK293T interactor studies and 5 mg of lysate was precipitated and redissolved as described above for HEK293T sample preparation and CL.
CL and digestion
For both brain and liver samples, 5 mg of total protein was labeled, enriched and digested as described above for HEK293T CL with no additional modifications.
LC–MS2
LC–MS2 detection for brain and liver O-GlcNAcomic samples were conducted as described above for HEK293T samples with no additional modifications.
MS data analysis and identification of O-GlcNAc modification sites
MS data analysis and interactor identification were carried out as described above for HEK293T site identification experiments (including the search parameter files in Supplementary Data 7) with some modifications. Peak lists were searched against the species-specific complete UniProtKB databases and the UniProtKB/Swiss-Prot databases (M. musculus, retrieved January 21, 2020) supplemented with a database of frequently observed contaminants (245 sequences).
Generation of O-GlcNAc network (brain and liver tissue)
Network assembly and partitioning
Network assembly and partitioning was performed as described above for the HEK293T network with no additional modifications. Human interactors were used as a proxy for mouse interactors because of the relative lack of data on protein–protein interactors in mice in BioGRID and IntAct.
Network functional analyses
Network analyses were carried out as described above for the HEK293T network with no additional modifications. The complete datasets along with all submission parameters are provided in Supplementary Data 9 (brain and liver).
Functional characterization experiments in neurons
Neuronal culture, stimulation and lysis
Primary mouse cortical neurons were isolated as previously described8, plated on poly(D-lysine)-coated plates, and cultured in Neurobasal medium with 1% P/S, 2 mM GlutaMAX Supplement and 1× B-27 Plus. Half of the medium was changed every 2–4 days. After 20 days in vitro, neuronal activity was silenced using 10 μM tetrodotoxin (Tocris Biosciences) and 100 μM D-AP5 (Tocris Biosciences) and a subset was also treated with 50 μM Thiamet G. The following day, silenced neurons were depolarized with 50 mM KCl or vehicle for 2 h and subsequently lysed with 2% SDS and 100 mM HEPES pH 7.9 containing cOmplete protease inhibitor cocktail and 100 μM Thiamet G. Protein concentrations were measured using the BCA assay. For the immunoprecipitation experiments, silenced and depolarized neurons were prepared as described and lysed with 1% Triton X-100 in TBS pH 7.6 containing cOmplete protease inhibitor cocktail, 100 μM Thiamet G and 0.25 U per μl benzonase nuclease (Santa Cruz Biotechnology).
CL of neuronal lysates
O-GlcNAcylated proteins from cell lysates (150 μg) were labeled as previously described63. Briefly, proteins were precipitated by the CH3OH–CHCl3–H2O method described above and pellets were then resolubilized with 40 µl of 1% SDS and 20 mM HEPES pH 7.9 for 5 min at 95 °C. ddH2O (49 µl), 5.5 mM MnCl2 (11 μl) and 80 µl of 2.5× GalT labeling buffer were added and the solution was vortexed gently before adding 10 µl of 1 mg ml−1 GalT-Y289L and 10 µl of 0.5 mM UDP-GalNAz in 10 mM HEPES pH 7.9. The reaction was then rotated end-over-end at 4 °C overnight. Control experiments were carried out in parallel in the absence of UDP-GalNAz. The following day, samples were alkylated with 12.5 mM iodoacetamide for 1 h at room temperature protected from light. Proteins were then precipitated as before and resolubilized by boiling in 199.6 μl of 1% SDS and TBS pH 7.6. After allowing the redissolved proteins to cool to room temperature, 0.4 μl of 50 mM WS DBCO-Biotin (Click Chemistry Tools) was added. The solution was rotated end-over-end for 1 h in the dark before being precipitated as described above. Here, 10% of the reaction volume was taken as input. The remaining aliquot was then incubated with streptavidin magnetic beads (Thermo Fisher) for 1.5 h in the dark. Beads were then washed five times with 0.5 ml of low-salt buffer (100 mM Na2HPO4, 150 mM NaCl, 0.1% SDS, 1% Triton X-100 and 0.5% sodium deoxycholate) and five times with 1 ml of high-salt buffer (100 mM Na2HPO4, 500 mM NaCl and 0.2% Triton X-100). Biotinylated proteins were eluted by boiling the resin in 50 mM Tris-HCl pH 6.8, 2.5% SDS, 100 mM DTT, 10% glycerol and 2 mM biotin for 15 min with intermittent vortexing. To calculate O-GlcNAcylation stoichiometry, the ratios of eluents to inputs, weighted by proportion of the total protein in each lane, were calculated as previously described19.
Coimmunoprecipitation of OGT–FH from neuronal lysate
Cortical neurons isolated from OGT–FH mice were grown, silenced, treated with KCl or vehicle and lysed after 21 days in vitro as described above. From each condition, 1 mg of protein was diluted to 1 ml with TBS containing 100 μM Thiamet G and 1× cOmplete protease inhibitor cocktail. Here, 15 μg of protein was saved as input. The remaining lysate was incubated with anti-FLAG magnetic beads (MilliporeSigma) at 4 °C overnight with end-over-end rotation. The following day, beads were washed three times with 0.1% Triton X-100 in TBS (pH 7.9) and eluted by boiling in 2% SDS and TBS pH 8.0. Inputs and eluents were run on SDS–PAGE and western-blotted for HA and liprin-α1 as described above.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Online content
Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41589-025-02108-7.
Supplementary information
Supplementary Notes 1–10 and titles and legends for Supplementary Data 1–10.
Supplementary Tables 1–29.
O-GlcNAc network of OGT interactors and substrates from HEK293T OGT–FH cells.
O-GlcNAc network of OGT interactors and substrates from wild-type and BAP1-knockout HEK293T cells with O-GlcNAc site quantification.
O-GlcNAc network of OGT interactors and substrates from forebrain tissue of OGT–FH mice.
O-GlcNAc network of OGT interactors and substrates from liver tissue of OGT–FH mice.
Original Python code to perform sites and region analysis, network generation and network hub ranking.
Search parameter files for MS data using Proteome Discoverer with Byonic search node.
Complete dataset and submission parameters for g:Profiler analysis of proteins in the HEK293T network.
Complete dataset and submission parameters for g:Profiler analysis of proteins in the OGT–FH forebrain and liver networks.
Source data
Numerical source data.
Uncropped gel and blot images.
Acknowledgements
We thank S. Pease and staff at the Caltech Genetically Engineered Mouse Services Core for help with generating the OGT–FH mouse line. We also thank D. VanderVelde and the Caltech Liquid NMR Facility, M. Shahgholi and the Caltech Multiuser MS Lab and the Dow Next-Generation Fund for supporting the MS and NMR characterization of biotin-Dmpt-alkyne and its precursors. This project was supported by the National Institutes of Health (R01AG060540 to L.C.H.-W.; T32GM008042, T32GM007616 and F30AG055314 to J.W.T.), UCLA-Caltech Medical Scientist Training Program (J.W.T.), National Science Foundation (GRFP DGE-1144469 to M.E.G.) and Department of Defense (NDSEG to E.H.J.).
Extended data
Author contributions
Conceptualization, M.E.G., J.W.T., Y.X. and L.C.H.-W. Methodology development, M.E.G., J.W.T., Y.X. and M.J.S. Software, J.W.T., M.J.S. and P.C. Investigation, M.E.G., J.W.T., Y.X., R.B.A., E.H.J., Y.K., A.L.S., H.A., T.D.K., A.M. and B.L. Writing—original draft, M.E.G., J.W.T., Y.X. and L.C.H.-W. Writing—review and editing, M.E.G., J.W.T. and L.C.H.-W. Visualization, M.E.G., J.W.T. and L.C.H.-W. Supervision, M.E.G., S.D.G. and L.C.H.-W. Funding acquisition, L.C.H.-W.
Peer review
Peer review information
Nature Chemical Biology thanks Junfeng Ma and the other, anonymous reviewer(s) for their contribution to the peer review of this work.
Data availability
All experimental data that support the findings of this study are available in the Supplementary Information or through ProteomeXchange (PXD035902). Information on source data for figures can be found in Supplementary Tables 28 and 29 in Supplementary Data 1. Source data are provided with this paper.
Code availability
All custom scripts used for data analysis in this study are provided in Supplementary Data 6. All code was also deposited along with supporting documentation in GitHub (https://github.com/michael-sweredoski/O-GlcNAc-Sites-And-Regions)44. Any additional materials, including code used to generate all figures and supplementary files in the paper, are available from the corresponding authors.
Competing interests
S.D.G. is the founder, president and chief technology officer of Proteas Health. The other authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Matthew E. Griffin, John W. Thompson.
Contributor Information
Matthew E. Griffin, Email: griffin@uci.edu
Linda C. Hsieh-Wilson, Email: lhw@caltech.edu
Extended data
is available for this paper at 10.1038/s41589-025-02108-7.
Supplementary information
The online version contains supplementary material available at 10.1038/s41589-025-02108-7.
References
- 1.Zachara, N. et al. The O-GlcNAc modification. In Essentials of Glycobiology (eds Varki, A. et al.) (Cold Spring Harbor Laboratory Press, 2022). [PubMed]
- 2.Yang, X. & Qian, K. Protein O-GlcNAcylation: emerging mechanisms and functions. Nat. Rev. Mol. Cell Biol.18, 452–465 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Fehl, C. & Hanover, J. A. Tools, tactics and objectives to interrogate cellular roles of O-GlcNAc in disease. Nat. Chem. Biol.18, 8–17 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Cheng, S. S., Mody, A. C. & Woo, C. M. Opportunities for therapeutic modulation of O-GlcNAc. Chem. Rev.124, 12918–13019 (2024). [DOI] [PubMed] [Google Scholar]
- 5.Chou, T. Y., Hart, G. W. & Dang, C. V. c-Myc is glycosylated at threonine 58, a known phosphorylation site and a mutational hot spot in lymphomas. J. Biol. Chem.270, 18961–18965 (1995). [DOI] [PubMed] [Google Scholar]
- 6.Yi, W. et al. Phosphofructokinase 1 glycosylation regulates cell growth and metabolism. Science337, 975–980 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chen, Q., Chen, Y., Bian, C., Fujiki, R. & Yu, X. TET2 promotes histone O-GlcNAcylation during gene transcription. Nature493, 561–564 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Rexach, J. E. et al. Dynamic O-GlcNAc modification regulates CREB-mediated gene expression and memory formation. Nat. Chem. Biol.8, 253–261 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wang, A. C., Jensen, E. H., Rexach, J. E., Vinters, H. V. & Hsieh-Wilson, L. C. Loss of O-GlcNAc glycosylation in forebrain excitatory neurons induces neurodegeneration. Proc. Natl Acad. Sci. USA113, 15120–15125 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Yu, S. B. et al. Neuronal activity-driven O-GlcNAcylation promotes mitochondrial plasticity. Dev. Cell59, 2143–2157 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Pawson, T. & Nash, P. Protein–protein interactions define specificity in signal transduction. Genes Dev.14, 1027–1047 (2000). [PubMed] [Google Scholar]
- 12.Wong, W. & Scott, J. D. AKAP signalling complexes: focal points in space and time. Nat. Rev. Mol. Cell Biol.5, 959–970 (2004). [DOI] [PubMed] [Google Scholar]
- 13.Cruz Walma, D. A., Chen, Z., Bullock, A. N. & Yamada, K. M. Ubiquitin ligases: guardians of mammalian development. Nat. Rev. Mol. Cell Biol.23, 350–367 (2022). [DOI] [PubMed] [Google Scholar]
- 14.Stephen, H. M., Praissman, J. L. & Wells, L. Generation of an interactome for the tetratricopeptide repeat domain of O-GlcNAc transferase indicates a role for the enzyme in intellectual disability. J. Proteome Res.20, 1229–1242 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Hu, C. W., Wang, K. & Jiang, J. The non-catalytic domains of O-GlcNAc cycling enzymes present new opportunities for function-specific control. Curr. Opin. Chem. Biol.81, 102476 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Potter, S. C. et al. Dissecting OGT’s TPR domain to identify determinants of cellular function. Proc. Natl Acad. Sci. USA121, e2401729121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Wang, S. et al. Protein–protein interaction networks as miners of biological discovery. Proteomics22, e2100190 (2022). [DOI] [PubMed] [Google Scholar]
- 18.Clark, P. M. et al. Direct in-gel fluorescence detection and cellular imaging of O-GlcNAc-modified proteins. J. Am. Chem. Soc.130, 11576–11577 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Thompson, J. W., Griffin, M. E. & Hsieh-Wilson, L. C. Methods for the detection, study, and dynamic profiling of O-GlcNAc glycosylation. Methods Enzymol.598, 101–135 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Griffin, M. E. et al. Comprehensive mapping of O-GlcNAc modification sites using a chemically cleavable tag. Mol. Biosyst.12, 1756–1759 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wulff-Fuentes, E. et al. The human O-GlcNAcome database and meta-analysis. Sci. Data8, 25 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Trinidad, J. C. et al. Global identification and characterization of both O-GlcNAcylation and phosphorylation at the murine synapse. Mol. Cell. Proteomics11, 215–229 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Su, G., Kuchinsky, A., Morris, J. H., States, D. J. & Meng, F. GLay: community structure analysis of biological networks. Bioinformatics26, 3135–3137 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Deplus, R. et al. TET2 and TET3 regulate GlcNAcylation and H3K4 methylation through OGT and SET1/COMPASS. EMBO J.32, 645–655 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Ding, X. et al. Mixed lineage leukemia 5 (MLL5) protein stability is cooperatively regulated by O-GlcNac transferase (OGT) and ubiquitin specific protease 7 (USP7). PLoS ONE10, e0145023 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Wu, D.-L. et al. O-Linked N-acetylglucosamine transferase 1 regulates global histone H4 acetylation via stabilization of the non-specific lethal protein NSL3. J. Biol. Chem.292, 10014–10025 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Sinclair, D. A. R. et al. DrosophilaO-GlcNAc transferase (OGT) is encoded by the Polycomb group (PcG) gene, super sex combs (sxc). Proc. Natl Acad. Sci. USA106, 13427–13432 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Tan, Z.-W. et al. O-GlcNAc regulates gene expression by controlling detained intron splicing. Nucleic Acids Res.48, 5656–5669 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Hart, G. W., Slawson, C., Ramirez-Correa, G. & Lagerlof, O. Cross talk between O-GlcNAcylation and phosphorylation: roles in signaling, transcription, and chronic disease. Ann. Rev. Biochem.80, 825–858 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.White, C. W. et al. Age-related loss of neural stem cell O-GlcNAc promotes a glial fate switch through STAT3 activation. Proc. Natl Acad. Sci. USA117, 22214–22224 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Shi, Y. et al. Sharp, an inducible cofactor that integrates nuclear receptor repression and activation. Genes Dev.15, 1140–1151 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Joachim, J. et al. Activation of ULK kinase and autophagy by GABARAP trafficking from the centrosome is regulated by WAC and GM130. Mol. Cell60, 899–913 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Kondo, S. et al. Characterization of cells and gene-targeted mice deficient for the p53-binding kinase homeodomain-interacting protein kinase 1 (HIPK1). Proc. Natl Acad. Sci. USA100, 5431–5436 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Tremblay, V. et al. Molecular basis for DPY-30 association to COMPASS-like and NURF complexes. Structure22, 1821–1830 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Huttlin, E. L. et al. A tissue-specific atlas of mouse protein phosphorylation and expression. Cell143, 1174–1189 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Chatham, J. C., Zhang, J. & Wende, A. R. Role of O-linked N-acetylglucosamine protein modification in cellular (patho)physiology. Physiol. Rev.101, 427–493 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Hwang, H. & Rhim, H. Functional significance of O-GlcNAc modification in regulating neuronal properties. Pharmacol. Res.129, 295–307 (2018). [DOI] [PubMed] [Google Scholar]
- 38.Im, Y. J. et al. Crystal structure of GRIP1 PDZ6–peptide complex reveals the structural basis for class II PDZ target recognition and PDZ domain-mediated multimerization. J. Biol. Chem.278, 8501–8507 (2003). [DOI] [PubMed] [Google Scholar]
- 39.Liu, X. & Ciulli, A. Proximity-based modalities for biology and medicine. ACS Cent. Sci.9, 1269–1284 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Han, S. et al. Modulation of synaptic transmission through O-GlcNAcylation. Mol. Brain17, 1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Parker, S. J., Halligan, B. D. & Greene, A. S. Quantitative analysis of SILAC data sets using spectral counting. Proteomics10, 1408–1415 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Ritchie, M. E. et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res.43, e47 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Pham, T. V., Piersma, S. R., Warmoes, M. & Jimenez, C. R. On the beta-binomial model for analysis of spectral count data in label-free tandem mass spectrometry-based proteomics. Bioinformatics26, 363–369 (2010). [DOI] [PubMed] [Google Scholar]
- 44.Thompson, J. & Sweredoski, M. J. michael-sweredoski/O-GlcNAc-Sites-And-Regions. Zenodo10.5281/zenodo.17240013 (2025).
- 45.Crooks, G. E., Hon, G., Chandonia, J. M. & Brenner, S. E. WebLogo: a sequence logo generator. Genome Res.14, 1188–1190 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Stark, C. et al. BioGRID: a general repository for interaction datasets. Nucleic Acids Res.34, D535–D539 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Orchard, S. et al. The MIntAct project—IntAct as a common curation platform for 11 molecular interaction databases. Nucleic Acids Res.42, D358–D363 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Mellacheruvu, D. et al. The CRAPome: a contaminant repository for affinity purification–mass spectrometry data. Nat. Methods10, 730–736 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Shannon, P. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res.13, 2498–2504 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Morris, J. H. et al. clusterMaker: a multi-algorithm clustering plugin for Cytoscape. BMC Bioinformatics12, 436 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Newman, M. E. & Girvan, M. Finding and evaluating community structure in networks. Phys. Rev. E69, 026113 (2004). [DOI] [PubMed] [Google Scholar]
- 52.UniProt Consortium UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res.47, D506–D515 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Hornbeck, P. V. et al. PhosphoSitePlus, 2014: mutations, PTMs and recalibrations. Nucleic Acids Res.43, D512–D520 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Raudvere, U. et al. g:Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res.47, W191–W198 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Ashburner, M. et al. Gene ontology: tool for the unification of biology. Nat. Genet.25, 25–29 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Bindea, G. et al. ClueGO: a Cytoscape plug-in to decipher functionally grouped gene ontology and pathway annotation networks. Bioinformatics25, 1091–1093 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Kanehisa, M. & Goto, S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res.28, 27–30 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Jassal, B. et al. The Reactome pathway knowledgebase. Nucleic Acids Res.48, D498–D503 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Slenter, D. N. et al. WikiPathways: a multifaceted pathway database bridging metabolomics to other omics research. Nucleic Acids Res.46, D661–D667 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Giurgiu, M. et al. CORUM: the comprehensive resource of mammalian protein complexes—2019. Nucleic Acids Res.47, D559–D563 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Yang, H., Wang, H. & Jaenisch, R. Generating genetically modified mice using CRISPR/Cas-mediated genome engineering. Nat. Protoc.9, 1956–1968 (2014). [DOI] [PubMed] [Google Scholar]
- 62.Behringer, R., Gertsenstein, M., Nagy, K. V. & Nagy, A. Manipulating the Mouse Embryo: A Laboratory Manual 4th edn (Cold Spring Harbor Press, 2014).
- 63.Darabedian, N., Thompson, J., Chuh, K. N., Hsieh-Wilson, L. C. & Pratt, M. R. Optimization of chemoenzymatic mass-tagging by strain-promoted cycloaddition (SPAAC) for the determination of O-GlcNAc stoichiometry by western blotting. Biochemistry57, 5769–5774 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Notes 1–10 and titles and legends for Supplementary Data 1–10.
Supplementary Tables 1–29.
O-GlcNAc network of OGT interactors and substrates from HEK293T OGT–FH cells.
O-GlcNAc network of OGT interactors and substrates from wild-type and BAP1-knockout HEK293T cells with O-GlcNAc site quantification.
O-GlcNAc network of OGT interactors and substrates from forebrain tissue of OGT–FH mice.
O-GlcNAc network of OGT interactors and substrates from liver tissue of OGT–FH mice.
Original Python code to perform sites and region analysis, network generation and network hub ranking.
Search parameter files for MS data using Proteome Discoverer with Byonic search node.
Complete dataset and submission parameters for g:Profiler analysis of proteins in the HEK293T network.
Complete dataset and submission parameters for g:Profiler analysis of proteins in the OGT–FH forebrain and liver networks.
Numerical source data.
Uncropped gel and blot images.
Data Availability Statement
All experimental data that support the findings of this study are available in the Supplementary Information or through ProteomeXchange (PXD035902). Information on source data for figures can be found in Supplementary Tables 28 and 29 in Supplementary Data 1. Source data are provided with this paper.
All custom scripts used for data analysis in this study are provided in Supplementary Data 6. All code was also deposited along with supporting documentation in GitHub (https://github.com/michael-sweredoski/O-GlcNAc-Sites-And-Regions)44. Any additional materials, including code used to generate all figures and supplementary files in the paper, are available from the corresponding authors.
