Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 4.
Published in final edited form as: Nat Chem Biol. 2025 Apr 4;21(8):1205–1213. doi: 10.1038/s41589-025-01866-8

Native top-down proteomics enables discovery in endocrine resistant breast cancer

Fabio P Gomes 1, Kenneth R Durbin 2, Kevin Schauer 3, Jerome C Nwachukwu 4, Robin R Kobylski 4,5, Jacqline W Njeri 4,5, Ciaran P Seath 6, Anthony J Saviola 7, Daniel B McClatchy 8, Jolene K Diedrich 8, Patrick T Garrett 8, Alexandra B Papa 4,9, Ianis Ciolacu 4,9, Neil L Kelleher 2,10, Kendall W Nettles 4,5,*, John R Yates III 8,*
PMCID: PMC12307125  NIHMSID: NIHMS2072218  PMID: 40186031

Abstract

Oligomerization of proteoforms produces functional protein complexes. Characterization of these assemblies within cells is critical to understanding the molecular mechanisms involved in disease and to designing effective drugs. Here, we present a native top-down proteomics (nTDP) strategy to identify protein assemblies (≤70 kDa) in breast cancer cells and in cells that overexpress epidermal growth factor receptor (EGFR), which serves as a resistance model of estrogen receptor-α (ER) targeted therapies. This nTDP approach identified ~104 complexoforms from 17 protein complexes, which revealed several molecular features of the breast cancer proteome, including EGFR-induced dissociation of nuclear transport factor 2 (NUTF2) assemblies that modulate ER activity. We found that the K4 and K55 posttranslational modification sites discovered with nTDP differentially impact the effects of NUTF2 on the inhibition of the ER signaling pathway. The characterization of endogenous proteoform-proteoform/ligand interactions revealed the molecular diversity of complexoforms and their role in breast cancer growth.

Graphical Abstract

graphic file with name nihms-2072218-f0016.jpg

Introduction

Nearly all critical functions in cells, including those associated with malignancies like breast tumors, are driven by proteins in complexes that often assemble via non-covalent interactions of monomeric subunits1. Individual proteins exist as a collection of proteoforms due to differential splicing, sequence variations, posttranslational modifications (PTMs), or mutations that can influence the formation, stability, and activity of functional protein assemblies2. The ability to identify protein complexes and to characterize their monomeric proteoform arrangements within the intracellular space improves our understanding of how biological processes are regulated. Mutated genes in cancer cells can alter protein sequences, thus creating new proteoforms that are difficult to characterize. Here, we use the term “complexoform” to define a protein complex formed by monomeric proteoform arrangements of one or more gene products3.

Bottom-up proteomics is the most common approach to study proteins from complex biological mixtures such as cellular extracts. This mass spectrometry (MS)-based method uses the analysis of peptides released from digested proteins to characterize proteins4. However, bottom-up proteomics is limited in its ability to elucidate the proteoforms that arise from individual proteins. The characterization of ligands and cofactors bound to protein complexes is also a significant challenge, as fragile non-covalent interactions are lost during the protein digestion step4,5. And because the presence of proteins is inferred from digested peptides, mapping combinatorial PTM patterns is difficult6,7. Native mass spectrometry (native MS) has emerged as a powerful complement to biophysical techniques (e.g., X-ray crystallography) for elucidating the structural biology of protein complexes8. Native MS provides compositional information on the architecture of intact macromolecular assemblies, whereas top-down proteomics (TDP) enables in-depth characterization of intact proteoforms4. These two methods have been recently combined in a single MS method “native top-down proteomics (nTDP)”, which provides detailed molecular information about protein assemblies in a single experiment813. Despite its structural elucidation power, nTDP has primarily been used on purified protein complexes, with very few applications to simple14,15 or complex mixtures9,11,13,16. This may be attributed to the notorious difficulty separating intact protein assemblies in complex biological mixtures, but it may also be due to a lack of bioinformatics tools that can effectively handle large-scale nTDP datasets. The low intracellular abundances of many proteoform assemblies in complex biological mixtures further impedes the ability to study them17,18.

Here, we developed and validated a nTDP approach for discovery of complexoforms from cell extracts (Extended Data Fig. 1) using estrogen receptor-alpha (ER)-positive MCF-7 breast cancer cells and MCF-7 cells overexpressing the Epidermal Growth Factor Receptor (EGFR) gene19. Growth factor signaling is one of the major mechanisms of resistance to ER-targeted therapy in breast cancer. EGFR amplification and mutations in downstream signaling components are clinically associated with tamoxifen resistance in patients19 and cell lines20 through poorly understood mechanisms. We used offline native size-exclusion chromatography (nSEC) to reduce sample complexity and enhance the detection of extremely low abundance endogenous protein assemblies in the mass spectrometer (Extended Data Fig. 1a). nSEC fractions were “injected-infused” at low nano flow rate and ionized by nano electrospray ionization (nano ESI) (Extended Data Figs. 1bc). The resulting protein ions were fractionated in the gas-phase with Field Asymmetric Ion Mobility Spectrometry (FAIMS) prior to entering the MS analyzer (Extended Data Figs. 1cd), which allowed us to separate complexoform ions with similar mass-to-charge ratios, but different charge states. This critical step further reduced sample complexity and facilitated MS analysis. Multistage tandem MS, a strategy that measures the intact mass of protein assemblies before ejecting and identifying monomeric proteoforms for subsequent dissociation and sequencing, was used to characterize the protein complexes and their constituent proteoforms (Extended Data Fig. 1e). We started our semi-automated data analysis with ProSight Native software21 (Extended Data Figs. 1fg) and incorporated manual curation with the aid of additional software to enable a more accurate and comprehensive interpretation of the extremely complex MS raw data. We identified a total of 104 complexoforms from 17 protein complexes in the breast cancer cells including proteoform-metal complexes, proteoform-proteoform-metal complexes, and proteoform-proteoform complexes. We characterized complexoforms of triosephosphate isomerase (TPI), macrophage migration inhibitory factor (MIF), and superoxide dismutase [Cu-Zn] (SOD1). These proteins are known for their roles in cancer. We found that EGFR induced dissociation of nuclear transport factor 2 (NUTF2) dimers. NUTF2 dimer functions as a nuclear pore shuttle protein responsible for importing RAN. Overexpression of NUTF2 inhibited breast cancer cell growth, which could be reversed or worsened by different mutations at sites of PTMs. Changes in ER DNA binding and interacting proteins further support ER as an indirect target of NUTF2 transcriptional and growth regulation, contributing to inhibition of breast cancer growth. This work demonstrates the development and application of nTDP to complex biological mixtures from breast cancer cells to discover proteoforms and their complexoforms that regulate breast cancer growth and treatment resistance.

Results

nTDP Workflow

To develop a general approach for identifying complexoforms from cellular extracts, we optimized all aspects of our nTDP workflow. For instance, the reproducibility and robustness of the nano ESI-FAIMS-MSn were established using three purified protein-protein and protein-metal assemblies including carbonic anhydrase II (CA II, ~29.1kDa), streptavidin (SA, ~53kDa), and avidin (AV, ~67kDa). All these protein systems showed a low charge state distribution (Supplementary Figs. 13), which confirmed that our nano ESI-FAIMS-MSn approach can maintain, isolate, and fragment protein-protein/metal assemblies to yield fingerprint spectra that can be used to elucidate monomeric proteoform arrangements in protein assemblies. Additionally, we confirmed the gas-phase fractionation efficiency of the method for protein complexes up to 70kDa from base peak ion mobiligram of each protein complex in a mixture (Supplementary Fig. 4). Importantly, the intact protein assemblies were identified and characterized in a single MS experiment. This experiment comprises three critical steps: 1) intact protein assemblies are ionized, fractionated in the gas-phase by FAIMS, and analyzed in their near native states (MS1); 2) component subunits are expelled from the intact assemblies (MS2); and 3) each subunit is fragmented for proteoform identification (MS3) (Extended Data Fig. 1e).

We chose MCF-7 and MCF-7-EGFR breast cancer cells for discovery of complexoforms. EGFR amplification and/or activation of downstream signaling components are clinically associated with tamoxifen resistance19. The effects of amplification can be reproduced in the EGFR-overexpressing cells, which grow 25% faster than the parental MCF-7 cells, and are resistant to breast cancer therapies such as tamoxifen and other Selective Estrogen Receptor Modulators (SERMs) (Extended Data Fig. 2)20. The isolation of complexoforms from breast cancer cells required extensive optimization, including fractionation and sample concentration with offline nSEC. We then employed our nano ESI-FAIMS-MSn approach for native top-down characterization of complexoforms from MCF-7 and MCF-7-EGFR breast cancer extracts. FAIMS further reduced the complexity of these cellular extracts by fractionating them in gas-phase. The MS data analysis required a combination of bioinformatics tools and manual curation due to the high complexity of the nTDP datasets, followed by the application of a set of criteria for reliable identification of intracellular proteoforms and their complexoforms (described further in Methods).

nTDP Analysis of MCF-7 and MCF-7-EGFR cellular extracts

We identified and characterized three distinct groups of protein assemblies: 1) protein-metal assemblies, 2) protein-protein-metal assemblies, and 3) protein-protein assemblies. The components of these protein assemblies were identified, and the stoichiometry of the assemblies defined. A total of 104 complexoforms from 17 protein assemblies were identified with high confidence from the breast cancer cell extracts. Most complexes were identified in the nSEC fractions 4 and 5 (Supplementary Fig. 5). By matching theoretical and observed masses (MS1 spectrum only) within ±1 Da, ~13 protein assemblies were identified with confidence. We also considered 3 protein assemblies (formed by truncated proteoforms) that were identified based only on their observed masses (MS1 spectrum only). Complexoforms containing truncated proteoforms (an irreversible PTM that can significantly affect the function and fate of protein complexes22) were identified. Novel heterodimeric assemblies of TPI, heterotrimeric assemblies of MIF, homodimeric assemblies of SOD1, and heterodimeric assemblies of NUFT2 were identified. And heteromeric and homomeric assemblies that ranged from dimers to hexamers were also identified. A complete list of identified proteins/proteoforms and their complexes/complexoforms is provided in Supplementary Data 1. The characterization of four protein complexes in MCF-7 cells including three protein-protein and one protein-protein-metal assemblies serve as examples of our methodology’s efficiency.

We initially interrogated the dimeric structure of TPI, an important glycolytic enzyme involved in breast cancer. Studies suggest that TPI may be associated with drug resistance, tumor progression, and metastasis. TPI could be regarded as a potential therapeutic target and a biomarker in various cancers23,24. We observed TPI dimeric structures with molecular masses of ~53kDa in both MCF-7 and MCF-7-EGFR cells (nSEC - Fraction 4). Fig. 1a shows the native MS1 spectrum of intact TPI dimers at the mass/range (m/z) of 3318 – 3539 with two charge states (16+ and 15+). Closer examination of the MS1 spectrum revealed that dimeric TPI complexes are formed by truncated and phosphorylated monomeric proteoform arrangements (Fig. 1b)25. Following HCD activation and release from dimeric TPI, monomeric proteoforms (11+ charge state, [MS2]) were isolated using the quadrupole mass filter and fragmented by HCD in the MS3 (Fig. 1a). The MS3 spectrum revealed a monomeric proteoform with two deamidations (N15 and N71). It yielded isotopically resolved sequence ions that were mapped to the sequence of TPI and displayed in a graphical fragmentation map (Figs. 1cd).

Fig. 1: nTDP spectra of dimeric TPI complexoforms from MCF-7 cell extracts.

Fig. 1:

a. nTDP spectra of TPI.

b. Expanded MS1 spectrum from a. (+15 charge state, TPI complexoforms, Pi, phosphorylation) for detailed structural information of TPI assemblies.

c. Fragmentation statistics.

d. TPI fragmentation map derived from the MS3 spectrum (deamination [N15 and N71] shown in gray).

TPI has been used to study biological catalysis; however, it has been suggested that endogenous proteolysis and deamidations of this enzyme can negatively affect its catalytic activity26. Although these deamidations and their location have been previously reported27, top-down analysis provides definitive confirmation of the location of these modifications. The HCD fragmentation resulted in backbone cleavages that yield diagnostic ions that allowed us to unambiguously pinpoint the deamidation sites and confirm previous observations28. This high level of characterization makes it possible to define the specific proteoform associated with a particular pathological state. We found TPI complexoforms that were formed by mutated (E104D) forms of TPI, and we also found that E104D monomeric subunits of TPI can be phosphorylated, deamidated, or acetylated (Supplementary data 1). Although mutations in E104D can induce endogenous proteolysis due to loss of rigidity of the 3-dimensional structure of TPI, the biological implications of PTMs on mutated TPI remains to be elucidated. In humans, TPI deficiency has been associated with neurological diseases, cardiomyopathy, and mutations in E104D result in premature death29.

Next, we identified MIF as highly expressed in both MCF-7 and MCF-7-EGFR cells (nSEC Fractions 5 and 6), a versatile cytokine with biological relevance in several cancers, autoimmune diseases, and inflammation30,31. As illustrated in Fig. 2a, the native MS1 spectrum provided a complete overview of MIF complexoforms. We used a wide quadrupole isolation window to isolate and fragment multiple peaks and characterized the peaks via diagnostic product ions of the constitutive proteins, which enabled the characterization of proteoforms and determined how they were arranged in the trimeric assemblies. For instance, as shown in Fig. 2a, we isolated the 12+ charge states (MS1) found in the m/z range ~3080 – 3100 from MCF-7 cells, which corresponds to MIF trimeric structures (~37 kDa) and confirms that MIF’s native-like assemblies are retained during the MS1 measurements. Figs. 2bc illustrate the MS2 spectrum of the monomers (6+ charge state) ejected from the 12+ charge states (MS1) of the MIF trimmers using HCD activation and their respective fragmentation maps along with sequence coverage, identification confidence scores (P-Score), and mass accuracy. When trimeric structures were disassembled, truncated, unmodified, nitrosylated, and acetylated monomeric proteoforms were revealed (Figs. 2bc). We also identified MIF complexoforms that were formed by interactions of phosphorylated monomeric subunits (Supplementary Data 1). Five MIF (homo and hetero) trimeric structures were formed by different monomeric MIF proteoform arrangements (Fig. 2a), while MIF truncated forms not related to MIF assemblies were also observed (Fig. 2b). The truncated monomeric proteoforms were not characterized, but the three monomeric proteoforms that form four different trimeric structures are shown in Fig. 2c. As illustrated in Fig. 2c, the high sequence coverage accomplished on either side of the modified residues with HCD allowed us to unequivocally assign C80 and K77 as nitrosylated and acetylated residues, respectively. Although S-nitrosylation at C81 has been previously reported32, top-down analysis provides clear evidence of the presence and location of this PTM and its association with a protein complex. This demonstrates that we can discover complexoform oligomeric structures with different proteoform components from cancer cell extracts.

Fig. 2: nTDP spectra of MIF complexoforms from MCF-7 cell extracts.

Fig. 2:

a. MS1 analysis of trimeric MIF complexoforms.

b. MS2 analysis of MIF subunits released from MIF complexoforms (MS1 spectrum).

c. Fragmentation maps and statistics of MIF monomeric subunits (unmodified, nitrosylated and acetylated forms).

We then characterized the metalloenzyme SOD1, an important antioxidant that is essential for oncogene-driven mammary tumor formation and for the conversion of superoxide into oxygen and hydrogen peroxide33. We observed SOD1 dimeric structures with molecular masses of ~32kDa in both MCF-7 and MCF-7-EGFR cells (nSEC - Fraction 5). Non-covalent bindings of Cu2+ and Zn2+ ions were observed for both the complex (11+ charge state, MS1) and the ejected proteoform subunits from the complex (6+ charge state, MS2). The ejected subunits were subsequently fragmented in the MS3 stage (Figs. 3ad). This MS1−3 cycle confirms that electrostatic interactions between the backbone side chains and that metal ions are maintained during HCD activations to yield binding site information. The fragmentation maps of SOD1 with apo fragment ions (non-metal) and holo fragment ions (metal) are shown in Figs. 3bd. Holo fragment ions were identified with mass shifts consistent with the binding regions of Cu2+ and Zn2+ ions, as previously observed in purified proteins (Figs. 3cd)34. Holo fragment ions were manually verified to increase confidence in the localization of metal cofactors35. The TDValidator module of ProSight Native software was used to generate theoretical isotopic distributions of the detected holo fragment ions and to overlay them on the raw spectrum (Figs. 3ef, Extended Data Fig. 3). This extended analysis verified cofactor binding regions and demonstrates that structural information about SOD1 was effectively extracted from the complex cellular mixtures of breast cancer cells. The SOD1 proteoform also carried an N-terminal acetylation, initial methionine cleavage, and a disulfide bond (C57-C111) (Figs. 3be). We also found methylated SOD1 complexoforms. However, the proteoforms of these complexes were not localized to a specific residue, and they were characterized with Confidence level 2A36. The importance of protein methylation has been observed in numerous cellular and physiological processes37, but the biological relevance of methyl groups on the SOD1’s primary and three-dimensional structures in breast cancer cells are unknown.

Fig. 3: nTDP spectra of the SOD1 complexoforms in MCF-7 cells.

Fig. 3:

a. MS1 and MS2 spectra of SOD1.

b. SOD1 modifications and metal bindings, shown in c–e.

c–d. Fragmentation maps and ststistics of SOD1 with c) apo and d) holo/apo fragment ions that were generated using the software ProSight Lite.

e. Fragmentation map of SOD1 monomeric proteoform covering regions of the protein where the metal ions are bound (generated using the software TDValidator).

f. Isotopic distribution of the diagnostic ion b144 that was generated using the software TDValidator.

Finally, we characterized NUFT2 in both MCF-7 and MCF-7-EGFR cells (nSEC - Fraction 5). NUTF2 is a dimeric protein that mediates the nuclear import of Ran, a small GTPase that directs nucleocytoplasmic trafficking of cargo proteins with a nuclear localization sequence (NLS), and other proteins38,39. Our nTDP strategy enabled identification of 8 proteoforms of NUTF2 with different combinations of PTMs that were uniquely distributed among 26 different complexoforms (Extended Data Figs. 45 and Supplementary data 1). The latter included 15 complexoforms with SOD1 and/or D-dopachrome decarboxylase (DDT), forming NUTF2-SOD1-DDT, NUTF2-SOD1, and NUTF2-DDT assemblies (Extended Data Fig. 5 and Supplementary data 1), demonstrating the remarkable diversity of related complexoforms identifiable with our advances in nTDP.

We observed a substantial decrease of the endogenous NUTF2 dimeric complexoforms in MCF-7-EGFR when compared to MCF-7 cells (Fig. 4a). The NUTF2 heterodimer with three acetylations was unique to MCF-7-EGFR cells (Fig. 4a, complexoform 8). Following HCD activation for MS2, we observed numerous monomeric NUTF2 proteoforms including acetylated, methylated, and dimethylated forms (Fig. 4b). Initial methionine cleavage was observed in all NUTF2 proteoforms. We also observed differences in the expression of the monomeric NUTF2 proteoforms in the two cells (Extended Data Fig. 6), although the overall expression was similar (Supplementary Fig. 6). Importantly, we observed a peak that may have arisen from either of two monomeric products, K4 acetylated-K55 methylated-K63 acetylated form, or K4 acetylated-K55 dimethylated-K63 acetylated in both cells (Fig. 4b, Extended Data Fig. 6). However, there was no molecular ion evidence in the MS3 stage for these proteoforms and identifications were based on the mass shift of the MS2 spectra. We were also unable to unambiguously distinguish isomeric NUFT2 proteoforms that carry either a dimethyl group in a single residue or 2 methyl groups in two different residues (Extended Data Fig. 6). While acetylations in NUTF2 subunits are known, the methylations, dimethylations, and multiple PTMs in a single NUTF2 monomeric subunit have not been previously reported, nor has the existence of multiple heterodimer and monomeric species. The discovery and characterization of a remarkable ensemble of NUTF2 complexoforms (Fig. 4, Extended Data Figs. 45), and the many other identified complexoforms (Figs. 13, Supplementary Data 1) demonstrates that nTDP can be used for discovery in complex cellular environments.

Fig. 4. nTDP spectra of NUTF2 in MCF-7 and MCF-7-EGFR cells.

Fig. 4.

a. MS1 spectrum of dimeric NUTF2 assemblies in MCF-7 and MCF-7-EGFR cells. Peaks with modifications in dimeric complexoforms are numbered and found in both cell types, except complexoform 8, found only in MCF-7-EGFR. Ac, acetyl; Me, methyl; Me2, dimethyl.

b. MS2 spectrum of NUTF2 monomeric proteoforms in MCF-7 or MCF-7-EGFR cells with the indicated modifications.

NUTF2 mediates crosstalk between the EGFR and ER pathways

Co-precipitation of NUTF2 with two different affinity tags verified that EGFR reduced the number of NUTF2 dimers (Fig. 5a and Extended Data Fig. 7). The NUTF2 PTM sites that we identified are conserved in vertebrates, with variable levels of conservation in other species (Supplementary Fig. 7). Crystal structures of NUTF2 protein complexes indicate that K4 is near the nucleoporin binding site (Fig. 5b). K55 and K63 are near the Ran binding site, with K55 participating in water mediated contacts between Ran and NUTF2 (Figs. 5cd)40,41, leading us to mutate K4 and K55. We engineered MCF-7NUTF2 cells to stably express epitope tagged HA-NUTF2 proteins since MCF-7 cells were not viable after transduction with short hairpin RNAs to silence the NUTF2 gene. Exogenous NUTF2 levels were less than the endogenous protein in these stable cells (Supplementary Fig. 8a), and the mutations did not affect HA-NUTF2 protein levels or nuclear localization in the stable cell lines (Supplementary Figs. 8bc).

Fig. 5: NUTF2 Modulates ER signaling and inhibits growth.

Fig. 5:

a. Dimerization of NUTF2 was assayed by transfection of MCF7 or MCF7-EGFR cells with the indicated tagged expression vectors, and the complex purified with streptavidin beads for Western blot analyses. Representative of two replicates, shown with additional controls in Extended Data Fig. 7.

b–d. Crystal structure of one subunit of the NUTF2 dimer interacting with RanGDP and a model of the FxFG nucleoporin peptide, based on PDBs 1GYB and 5BXQ.

e. Stable MCF-7 cells were cultured with vehicle or 1 μM 4OHT for 5 days and assayed for cell number. N = 2 biological experiments with 2–4 wells each. *Significantly different from WT NUTF2 by 1-way ANOVA.

f. NUTF2 downregulates GREB1 gene expression. Stable MCF-7 cells in complete medium were treated for 24 h with vehicle or 1 μM 4OHT. GREB1 mRNA levels were compared by qPCR. N = 2.

NUTF2 expression in the MCF-7 cells suppressed cell proliferation (Fig. 5e). The K4Q mutation rescued this phenotype, while K55 NUTF2 was more growth suppressive, suggesting that both the nucleoporin and RAN binding sites regulate cell growth (Fig. 5e). NUTF2 downregulated the estrogen-induced proliferative gene, GREB1, and the K4Q mutation also rescued this phenotype in the absence of the active tamoxifen metabolite, 4-hydroxytamoxifen (4OHT) (Fig. 5f). The differential effects of the PTM mutants on breast cancer growth showed that the nTDP-identified NUTF2 proteoforms drove distinct biological outcomes.

To further investigate how NUTF2 modulates ER signaling and the response to 4OHT, we identified cellular DNA binding sites that comprise the ER cistromes. As a control, MCF-7 cells were starved of estradiol for 3 days and then treated with estradiol for 1hr. Using Cleavage Under Targets and Release Using Nuclease (CUT&RUN), we identified 1582 DNA binding sites for ER, including enrichment for the canonical ER binding site (Extended Data Figs. 8ab). In MCF-7 cells (with empty vector transduction) and MCF-7NUTF2 cells, ER and NUTF2 cistromes were essentially non-overlapping (Fig. 6a), but there were complex interactions between NUTF2 expression and tamoxifen 4OHT treatment on ER binding patterns (as illustrated in Extended Data Fig. 8c). ER showed 1,353 differential binding sites between vehicle treated MCF-7 vs MCF-7NUTF2 (Fig. 6b). Treatment with 4OHT induced some overlapping and some distinct ER binding sites between the MCF-7 and MCF-7NUTF2 cells (Fig. 6b). NUTF2 expression also altered ER binding sites in the vehicle (Extended Data Fig. 8d) and at 4OHT-sensitive DNA binding sites (Fig. 6c). 4OHT increased NUTF2 occupancy at >10% of its binding sites (Extended Data Fig. 8e). Further analysis of the 4OHT-sensitive NUTF2 sites using de novo motif discovery tools revealed a common motif (Extended Data Fig. 8f, Supplementary Table 1) for consensus binding sites of Forkhead box (FOX) proteins (Supplementary Tables 23), some of which function as pioneer factors for ER42,43. Only 132 consensus ER motif half-sites (i.e., “1-AGGTCA”) were found at 4OHT-sensitive NUTF2 sites, suggesting that NUTF2 modulates ER chromatin occupancy and activity through an indirect mechanism.

Fig. 6. NUTF2-ER crosstalk.

Fig. 6.

a. CUT&RUN identified ER and NUTF2 binding sites in MCF-7 cells expressing empty vector or NUTF2, n =2. b. Differential occupancy analysis of ER binding sites identified by CUT&RUN. MCF-7 cells expressing empty vector or NUTF2 were steroid-deprived and then treated with vehicle or 1 μM 4OHT for 1 h, n = 2.

c. ER CUT&RUN binding profile at 4OHT-sensitive sites in MCF-7 cells (empty vector) or MCF-7-NUTF2 cells.

d. Volcano plot of RNA-seq data showing the effects of NUTF2 overexpression on gene expression in MCF-7 cells, n = 3.

e. Potential upstream regulators of NUTF2 target genes were identified based on published ChIP-seq data sets using LISA (http://www.lisa.cistrome.org).

ER-APEX2 gene fusion was transduced into MCF-7 stable cells expressing empty vector, wild type NUTF2, or NUTF2-K4Q. After 24 h, cells were treated with H2O2 for 1 minute to induce biotin labeling of ER proximal proteins. The ER interactome (left panel) was defined relative to non-transfected cells (bead control), including only proteins significantly enriched relative to both bead and GFP-APEX2 transfected controls. See also Supplementary Data 2.

g. ER-APEX2 was transduced into MCF7 cells expressing NUTF2-WT (left) or NUTF2-K4Q (right) and treated as in f. Data are shown relative to MCF7 cells expressing empty vector (left) or NUTF2-WT (right).

To understand the molecular basis of NUTF2 growth inhibition, we completed RNA-seq of MCF-7 and MCF-7NUTF2 cells (Fig. 6d). Gene set enrichment analyses suggest that among the 600+ differentially expressed genes, NUTF2 upregulated RNA catabolic processes, components of the NuRD complex such as HDAC1 and SPEN1, transcriptional repressor complex, and other chromatin regulatory factors that could indirectly impact ER activity (Fig. 6d, Supplementary Fig. 9, Supplementary Table 4). The apoptotic gene ERFFI1 was also upregulated, as was PDK2, the suppressor of aerobic respiration. Genes encoding mitochondrial proteins including those that mediate oxidative phosphorylation were downregulated, such as the antiapoptotic gene, BCL2, and respiratory chain gene MT-CO1 (Fig. 6d, Supplementary Fig. 9). The oncogenic RAS family genes RAB20 and RAP2A, five PARP family genes, and several genes involved in the interferon signaling pathway (including IRF9) were also downregulated (Fig. 6d, Supplementary Fig. 9). Pathway enrichment analysis integrating NUTF2 cistrome (i.e., CUT&RUN), and functions (i.e., RNA-seq), revealed over-representation of many cancer- and EGFR-related pathways (Extended Data Fig. 9), reinforcing the identified role of NUTF2 as an effector of EGFR signaling. Additionally, we noted that there were 51 genes that are regulated by both NUTF2 and estradiol in quiescent MCF-7 cells, along with genes that were regulated in opposite directions, highlighting the broader role of NUTF2 in regulating gene expression (Supplementary Fig. 10). This role was supported by an analysis of ChIP-seq data sets, which identified ER binding sites among the top five upstream regulatory elements of both upregulated and downregulated genes, but other transcription factors and coregulators were also identified as upstream regulators (Fig. 6e). This suggests that NUTF2 may be regulating cell growth through several different transcriptional pathways and additional roles as part of the nuclear pore complex and its known regulation of RAN activity.

Biotin proximity proteomics with transfection of an ERα-APEX2 fusion enabled rapid H2O2 induction of biotinylation and MS identification of streptavidin purified ER proximal proteins. This revealed an ER interactome that contained components of histone readers, writers, and erasures (Fig. 6f, Supplementary Table 5, Supplementary Data 2), which was modulated by NUTF2 expression. These included increased labeling of JunB and decreased labeling of the HSPBP1 tumor suppressor and CSDE1, an upstream regulator of RAS and WNT signaling (Fig. 6g). Comparison of NUTF2-WT versus -K4Q expressing cells showed an altered ER-interactome, including CNOT9 and additional increased labeling of JunB, along with decreased labeling for other proteins, including the DNMT3A methyltransferase (Fig. 6g, Supplementary Data 2).

Discussion

nTPD has advanced the structural biology field by complementing traditional biophysical techniques with unique insights into the dynamics, interactions, and structure of protein complexes. However, it has been largely limited to analysis of purified proteins due to the challenge of isolating intact protein assemblies from complex biological mixtures such as cellular extracts. Another barrier is the lack of bioinformatics tools for accurate, comprehensive, and high-throughput analysis of highly complex nTDP datasets from cellular extracts, especially when the goal of the experiment is to elucidate the composition of intact complexoforms.

Here, we demonstrated that a combination of advances enabled us to isolate and characterize 104 complexoforms from 17 protein complexes from breast cancer cell extracts using nTDP. These advances included offline liquid fractionation to reduce sample complexity and concentrate the extremely low abundant intracellular protein complexes. We were able to further reduce sample complexity with additional online gas-phase fractionation with FAIMS, which makes large scale, sensitive, and efficient characterization of complexoforms from cells possible, particularly when the goal of the study is to achieve extensive protein characterization rather than in-depth proteomic coverage. These improvements were critical through the process of identifying complexoforms, metal cofactors, the ensemble of component monomeric proteoforms, site-specific identification of PTMs, and other proteoform variants (e.g., truncations and mutations). Our nTPD workflow also benefitted from extensive optimization of the MS strategy, including HCD tuning for controlled release of proteoform subunits (MS2) and fragmentation of proteoform subunits (MS3). While ProSight Native software provided a semi-automated analysis of nTDP datasets (MS1–M3 spectra)21, the complexity of breast cancer cell extracts required manual curation with the aid of the software ProSight Lite and TDValidator to enable discovery and characterization of protein assemblies35. With the single gene complexoforms of TPI, SOD1, and MIF, we were able to resolve oligomeric states, PTMs, and metal binding regions with high confidence. NUTF2 revealed a more complex set of assemblies, including complexes with SOD1 and/or DDT. The NUTF2 proteoforms included sets of PTMs that were found in different monomers, homodimers, and heterodimers. The power of nTDP for discovery is further demonstrated by our findings that the EGFR model of endocrine resistant breast cancer was associated with monomeric NUTF2 compared to the higher levels of dimeric NUTF2 in the hormone sensitive MCF-7 parental cells, and there were differences in proteoform PTMs in the EGFR model as well. Remarkably, inhibition of breast cancer cell growth by NUTF2 expression was blocked by mutation of one PTM site and enhanced by a different PTM site mutation, demonstrating roles for distinct biological outcomes for different proteoforms.

Growth factor signaling is one of the major drivers of endocrine resistant breast cancer, including EGFR, Her2, FGFR family, and their downstream signaling pathways19. Mechanisms of resistance include circumvention of a requirement for ER for growth or through coopting ER signaling, such as through PTMs known to drive ER constitutive activity. Here, NUTF2 inhibited breast cancer growth through ER independent mechanisms, but also through altered ER DNA binding and protein interaction networks, while tamoxifen influenced NUTF2 DNA localization. Rather than direct interactions, we suggest that NUTF2 and ER can alter the chromatin states independently, which in turn modulates cellular responses to the other signal. NUTF2 is poorly studied outside of its role in nuclear translocation. Mutations in the dimer interface reduce NUTF2 interaction with the nuclear pore and RAN, but our work suggests a role in transcriptional regulation.

Much of the activity in the nTDP field has been focused on the characterization of highly purified individual protein complexes via direct infusion experiments using mass analyzers that can handle large protein ions (e.g., Orbitrap with extended mass range44, Time of Flight45, and Fourier-Transform Ion Cyclotron Resonance8). Recently, a native proteomics study using a conventional Orbitrap platform was reported, but it was limited to the characterization of protein complexes < 30kDa11. As demonstrated here, our innovative nTDP platform can be used in discovery mode for characterization of complexoforms up to 70 kDa in human breast cancer cells. Although there are many protein complexes larger than 70 kDa in the breast cancer proteome, we were limited to protein assemblies ≤70 kDa due to mass range limitations of our conventional Orbitrap mass spectrometer. The overall consistency of benchmark proteins across triplicate measurements reflected the reliability of our nTDP approach. Efficient fractionation and concentration of low abundance protein complexes allowed us to identify the proteoforms that constitute different protein assemblies. Obtaining satisfactory fractionation was particularly important to this nTDP workflow, not only to reduce ion suppression effects and signal superposition, but also to avoid precursor overlap during quadrupole isolation of ions due to our large quadrupole isolation windows. In a single MS experiment, our nTDP approach provided three different levels of molecular information: 1) the intact masses of protein assemblies including non-covalent cofactors were obtained under near physiological conditions; 2) endogenous proteoform subunits and non-covalent cofactors were released from their respective protein assemblies using controlled HCD; and 3) this controlled disassembly of proteoform subunits and non-covalent cofactors allowed us to obtain detailed primary sequence information and to localize non-covalent cofactors after their quadrupole isolation and HCD fragmentation.

Our platform represents a paradigm shift in the study of protein complexes, with implications for the fields of cancer biology, discovery proteomics, drug discovery, and structural biology. It opens a new avenue for large-scale native top-down proteomic characterization of endogenous complexoforms. Our nTDP workflow is uniquely positioned to reveal intracellular molecular pathways and illuminate functional PTM differences in the expression landscapes of proteoforms and their complexoforms. By preserving fragile non-covalent proteoform-proteoform and proteoform-ligand interactions, our integrated platform captures, with very high sensitivity and accuracy, the full spectrum of protein assemblies within cells. This led us to the discovery of numerous endogenous protein assemblies that could be further explored to inform important cellular processes and as potential therapeutic targets or cancer biomarkers.

Methods

Overall development, optimization, and data interpretation of the nTDP workflow.

Following the analysis of a standard protein complex mixture, whole cell lysates of MCF-7 and MCF-7-EGFR were fractionated with nSEC (Supplementary Fig. 5). Recently, online nSEC-MS has been effectively used to separate protein complexes from small molecule non-volatile buffer components prior to MS analysis47. SEC has been previously used under denaturing conditions for fractionation and concentration of intact proteins48. Our nSEC separation strategy was developed using protein standards (bovine serum albumin [BSA], ~66kDa; ovalbumin [OV], ~44kDa; and myoglobin [M], ~17kDa) and a chromatographic column with a small pore size (300 Å) to facilitate the separation of medium and low molecular weight protein complexes (Supplementary Figs. 5ab). One of the major problems in the nTDP field is the notoriously low abundance of compexoforms within cells. Here, we addressed this issue by injecting each sample into the chromatographic system 16 times. The separation window and signal intensity of the nSEC method were found to be highly reproducible, which allowed enrichment of extremely low abundant protein assemblies from breast tumor cells (Supplementary Figs. 5ab). nSEC used a mobile phase compatible with native MS; thus, the fractions were ready for direct nano ESI infusion FAIMS-MSn analysis to maintain the integrity of proteoform assemblies. After concentration of each fraction (16 replicates) with molecular weight cutoff filters, Native-PAGE was used to visualize the nSEC fractions from each sample to confirm the presence of protein complexes (Supplementary Figs. 5cd). These results show that the complexity and composition of the two samples are similar and that protein complexes remain intact. Four-fractions of MCF-7 (3 – 6) and four fractions of MCF-7-EGFR (3 – 6) cells with identical retention times were analyzed (Supplementary Figs. 5cd). Although we collected fractions from ~9–30 min, we only analyzed fractions from ~11 – 14 min for each sample, which included molecules that are relevant for the scope of this work. Each concentrated nSEC fraction was then “injected-infused (~90 min)” using a chromatographic system, ionized by nano ESI, and fractionated using specific optimized voltage settings on the FAIMS, which allowed us to separate complexoform ions with similar mass-to-charge ratios, but different charge states. The overall patterns were preserved across triplicate nano ESI-FAIMS-MSn analyses.

We employed FAIMS in a multistage MS approach for native top-down characterization of complexoforms in MCF-7 and MCF-7-EGFR breast cancer cells. We pushed our Orbitrap Fusion Lumos without extended mass range (EMR) capabilities to its limit to attain optimum transmission and detection of sub-70kDa intact biomolecular complexes. CA II, SA, and AV were used to develop and evaluate the reproducibility and robustness of our nTDP method because these proteins are examples of protein-metal and protein-protein assemblies that can be found in breast cancer cells, and they span a suitable molecular weight range (~30–70kDa) for the Orbitrap Fusion Lumos (determined through meticulous optimization of key MS parameters). In this nano ESI-FAIMS-MSn method, high pressure (14mTorr) in the ion-route multipole (IRM) was necessary to ensure optimal collisional cooling of the intact protein ions from the noncovalent assemblies, which enabled them to be effectively transmitted, trapped, and ejected to the detector49. At higher pressures (e.g., 16mTorr), we noticed collisional gas instability in the mass spectrometer. A high resolving power in the MS3 stage was necessary to resolve overlapping isotope patterns, thereby characterizing the constituent proteoforms in protein assemblies. Tandem spectral quality also benefited from the high number of μscans and long ion injection times in the MS2 and MS3 stages. Importantly, all these instrument parameters are directly associated with scan time. Since our approach is based on direct infusion and gas-phase fractionation, MS acquisition parameters can be set at higher values than are possible for online liquid-phase separation strategies without affecting data collection. For example, in capillary zone electrophoresis-mass spectrometry (CZE-MS) approaches, scan time must be adjusted according to the short elution time of each protein peak, as spectral quality is driven by real-time peak area of protein precursor ions. Additionally, we evaluated the gas-phase fractionation efficiency of the method by analyzing the standard protein complex mixture, with CVs ranging from −45 to −30 V with 5 V stepping. Our multistage MS approach employed a three-tiered strategy that determined the intact masses of non-covalent protein assemblies (MS1). Higher-energy collisional dissociation (HCD) was optimized for both the disassembly of intact proteoform subunits (MS2) and fragmentation of proteoform subunits (MS3).

Proteoforms and their complexoforms were identified using the software ProSight Native21. This bioinformatic tool integrates stoichiometry calculations and spectral deconvolutions with top-down database searches for the determination of molecular masses. Although ProSight Native is proficient in defining the full composition of intracellular proteoforms and their complexoforms, the characterization of these biomolecules is performed in a semi high-throughput manner. For instance, protein assemblies were individually searched. The software “ProSight Lite” was also used to assist with the characterization of the monomeric proteoform subunits and to visualize their sequences. Briefly, ProSight Lite matches fragment ions with proteoform sequences50 and complements the software “ProSight Native”. Other software, such as QualBrowser and Freestyle were used to assist with deconvolution and deisotoping. To obtain fragmentation maps, the MS3 spectra were deconvoluted using either QualBrowser or Freestyle, and the fragment ions were manually curated using the software “ProSight Lite”. Although a large-scale identification of intracellular proteoforms and their complexoforms was possible in this manner, our identification strategy required manual supervision and interpretation.

Monomeric proteoform subunits of the protein assemblies were considered identified when the P-Score was ≤1 × 10−05 with a minimum of 4 matching fragment ions (Supplementary Data 1). Protein-metal assemblies were putatively considered identified when the difference between observed and theoretical masses was ≤ 2 Da and the number of non-covalent ligands bound to the complex was ≤ 4. To expand the list of identifications, the mass shifts from observed masses of the protein assemblies were matched with the mass shifts of metal cofactors and/or PTMs. Protein assemblies were considered identified with confidence when they were observed in at least two of the three replicates with a minimum of two charge states.

Protein complex standard preparation.

Ammonium acetate solution 7.5M (A2706) and carbonic anhydrase II (C2522) were purchased from Sigma-Aldrich. Avidin (PI21121) and streptavidin (434301) were purchased from Thermo Fisher Scientific. Proteins were diluted using 100 mM ammonium acetate and were desalted and concentrated (~1–30μM) using 10kDa molecular weight cutoff filters (Millipore). The desalting procedure was repeated 10 times (12,000 × g for 5 min at 4°C) to minimize salt effects.

Cell lines.

MCF7-EGFR cells were derived by stably transducing MCF7 cells (ATCC HTB-22) with a lentivirus expressing wildtype human EGFR20. MCF7 cells were also stably transduced with lentiviruses expressing wildtype human NUTF2 with a C-terminal DYKDDDDK (FLAG) or hemagglutinin (HA) epitope tag, or the empty lentiviral vector as a control. The HA epitope-tagged NUTF2-K4M, -K4Q, and -K55R lentiviral vectors were generated from the wildtype NUTF2 vector using the Q5® Site-Directed Mutagenesis Kit (New England BioLabs, E0554S). All cells were cultured in Dulbecco’s modified Eagle medium (DMEM) (Thermo Fisher Scientific, 11995073) supplemented with 10% fetal bovine serum (Sigma-Aldrich, F0926), GlutaMAX (Thermo Fisher Scientific, 35050061), non-essential amino acids (Thermo Fisher Scientific; 11140050), penicillin, streptomycin, and neomycin (Thermo Fisher Scientific, 15640055) and maintained at 37°C in a 5% CO2 incubator.

Protein extraction.

Cells were cultured in 15 cm dishes for 3 days, placed on ice for 5 min, aspirated and rinsed with chilled 1X phosphate-buffered saline (PBS), lifted into 2 ml per dish of PBS, collected in 15 mL conical tubes, and then centrifuged (500 × g for 3 min). The supernatant was carefully removed, and the pellet was stored at −80°C. The cell pellet (equivalent to ~100μL) was resuspended in 500μL of ice-cold 200 mM ammonium acetate pH 6.9 with protease and phosphatase inhibitors (Thermo Fisher Scientific). Proteins were mechanically extracted from MCF-7 cell lines using a Dounce homogenizer (20 strokes). Cellular debris and intact organelles were pelleted by centrifugation (17,000 × g for 15 min at 4 °C) and the supernatant, which contained proteins, was filtered with a 0.45 μm syringe filter, transferred to a clean microcentrifuge tube, and quantified using a BCA assay kit (Thermo Fisher Scientific). The lysate was immediately fractionated.

nSEC conditions.

A standard protein mixture was used to develop the nSEC strategy and to monitor the performance of the method. This mixture was composed of 3 proteins (bovine serum albumin (BSA), [Sigma-Aldrich, A3059], ovalbumin (OV), [Sigma-Aldrich, A2512], and myoglobin (M), [Sigma-Aldrich, M1882]). Five milligrams of each standard protein was placed in a single vial and dissolved in 500μL of ultrapure water. The protein mixture was then filtered using a 0.45 μm syringe filter and 20μL (~600 μg of total protein) was injected (triplicate measurements) into an Agilent 1200 HPLC system. Approximately 200μg of total protein (cell lysate) was injected (16 replicates) using the HPLC system. The mobile phase was 200 mM ammonium acetate pH 6.9. Samples were fractionated in the cold room (~7°C) on a PolyHYDROXYETHYL A column (200 × 9.4 mm; 3 μm; 300 Å [PolyLC Inc]) using isocratic elution at a flow rate of 0.5 mL/min. Fractions were collected from ~9–30 min at 1 min intervals. A 20μL injection loop was used for the protein standard mixture and a 200μL injection loop was used for the cell lysates. The elution profile was monitored by UV absorbance at 280 nm using the Agilent ChemStation Software. Each nSEC fraction was then concentrated/desalted using a 10kDa molecular weight cutoff filter, quantified using a BCA assay kit, and diluted with 100 mM ammonium acetate to approximately ~1μg (total protein) in 18μL for subsequent FAIMS-MSn analysis.

One-dimensional native polyacrylamide gel rlectrophoresis (Native-PAGE).

Ten microliter aliquots from each SEC fraction were added to each lane of a 10–20% tris-glycine polyacrylamide gel (Thermo Fisher Scientific). Five microliters of NativeMark Unstained Protein Standard (Thermo Fisher Scientific) was used for molecular weight estimation and the gel was electrophoresed at 225 V for ~37 min. Coomassie blue (Invitrogen) was used for visualization.

Nano injection-infusion.

The FAIMS Pro device was placed between an Easy-nLC 1200 system and Orbitrap Fusion Lumos mass spectrometer (Thermo Fisher Scientific). Standard protein complexes and nSEC fractions were infused into the mass spectrometer on a fused silica capillary (~5 cm × 75 μm i.d., with a 2 μm pulled tip), using the Easy-nLC platform in direct infusion mode at a constant flow rate of 150nL/min. Ultrapure water was used as a carrier fluid. The auto sampler temperature was set to 4°C. The “injection-infusion” volume was 18μL (~1μg of total protein) for each nSEC fraction, which was sufficient to acquire data for 90 min for each fraction. All fractions were “injected-infused” in triplicate.

FAIMS conditions.

The FAIMS Pro device was connected to an Orbitrap Fusion Lumos mass spectrometer. FAIMS was set to the standard resolution mode (inner and outer electrodes at 100 °C), with supplementary FAIMS gas flow (nitrogen) of 5L/min. The FAIMS Pro front plate was set to 250 V and the dispersion voltage was −5 kV. FAIMS was operated in static mode. Compensation voltage (CV) scans for the optimization of the transmission of each standard protein complex were individually performed using a scan range of −60 V to −10 V at a rate of −10 V per scan (0.2 seconds each CV transmission). Subsequent FAIMS analyses were performed according to the optimal CV transmissions for the protein complex mixture (−45 V to −30 V at a rate of −5 V per scan (4 minutes for each CV transmission) and fractions (−40 V to −30 V at a rate of −5 V per scan [30 minutes for each CV transmission]).

MS conditions.

MS measurements were performed on an Orbitrap Fusion Lumos mass spectrometer. Before native top-down experiments were run, a full set of MS calibrations was performed in positive mode in “standard pressure” with the ion-route multipole (IRM) pressure at 8mTorr. The mass spectrometer was then changed to “intact protein mode” and the pressure in the IRM was set to 14mTorr for calibration of the optics in high mass range. The instrument was operated in “intact protein mode”, with a pressure of 14mTorr in the IRM. All calibrations were performed using the Pierce FlexMix Calibration Solution (Pierce). The mass spectrometer was operated in positive mode and spray voltages ranged from 1300 and 2500 V. The mass range was set to “high m/z”. The inlet capillary temperature was 200°C and in-source fragmentation voltage was set to 25–50 V. The automatic gain control (AGC) was set to accumulate 4E5–1E6 ions (full MS) and 4.5E5–1E6 ions (MS2) in a maximum injection time of 25–200 milliseconds (full MS) and 50–1000 milliseconds (MS2). An AGC value of 1E6 ions in 2500–3000 milliseconds was set for the MS3 stage. Full MS scan and MS2 spectra were acquired with a resolving power of 7.5–60k and MS3 spectra were acquired using a resolving power of 240k. The full MS and MS2 spectra were acquired by averaging 3–5 μscans, while the MS3 spectra were acquired by averaging 7–10 μscans. MS2 was set to 20 m/z isolation in the quadrupole, and monomeric subunits were ejected from the protein assemblies using normalized higher-energy collisional dissociation (HCD) with a collision energy of 15–20%. In the MS3 stage, monomeric proteoforms were isolated by the quadrupole (isolation width of 50 m/z) and fragmented using 35% HCD. Precursor ions were dynamically excluded for 6000 seconds after being selected (10–20 m/z) and “exclude isotopes” was enabled. The MS3 spectra were acquired in data-dependent mode (top 3–5).

MS data analysis.

Raw files were processed using ProSight Native v1.0.24275 (Proteinaceous, Inc., Evanston, IL)21,35. Protein complex masses were determined from MS1 spectra by charge state deconvolution using the kDecon algorithm in ProSight Native. Subunit masses from MS2 spectra and fragment ions from MS3 spectra were determined using the THRASH algorithm in ProSight Native. Subunit identifications were made via ProSight Annotated Proteoform searches against the full human SwissProt UniProt database containing annotated modifications. Searches were performed using precursor mass tolerance windows of 100–10 Da and a fragment ion mass tolerance of 10 ppm. A wide mass tolerance window (e.g., 100 Da) was used to capture modifications to a protein sequence that are not included in the database. Additional complex and subunit analyses were performed by manual analysis in QualBrowser 3.1, Freestyle software 1.8 SP2 (Thermo Fisher Scientific), and the ProSight Lite and TDValidator modules of ProSight Native. To obtain fragment ions for ProSight Lite analysis, the MS3 spectral data were deconvoluted using the QualBrowser or Freestyle. ProSight Lite, with mass tolerance of either 20 or 10 ppm, was utilized to support the assignment of product ions from deconvoluted MS3 spectra and to produce graphic interpretation of the MS3 spectra, with b/y ion bond cleavages indicated in blue. TDValidator was used to confirm the location of cofactor ligands and PTMs. Theoretical fragment ion isotopic distributions containing metal cofactors and PTMs were produced using the BRAIN algorithm (calibration of 10.00 ppm) in TDValidator. The intensity of the protein assemblies was determined by averaging the spectra of the target species for 90 minutes in each of the three replicates. The averaged intensity of each species among the three replicates was calculated. For the monomeric proteoforms, the intensity of each species was obtained by deconvoluting the spectrum of the target species to zero-charge mass spectrum and averaging them in each of the three replicates using the software TDValidator.

Transactivation assays.

MCF7 cells in 96-well plates containing phenol red-free DMEM plus 10% charcoal-stripped FBS, 1X GlutaMAX, and 1X nonessential amino acids, were co-transfected with 50 ng per well of 3x estrogen-response element-driven luciferase (3xERE-Luc), mouse mammary tumor virus promoter-driven luciferase (MMTV-Luc), or 5x NF-κB response element-driven luciferase (5xκBRE-Luc) reporter plasmid; 10 ng per well of empty vector or human NUTF2 expression vectors; and 5 ng per well empty vector, wildtype human ER expression vector, using FuGENE® HD transfection reagent (Promega, cat no. ME2311). The next day, the medium was replaced with phenol red-free DMEM plus 10% complete FBS, 1X GlutaMAX, and 1X nonessential amino acids (complete medium) and treated with vehicle, 1 μM 4-hydroxytamoxifen (4OHT), 100 nM Dexamethasone (Dex), or 20 ng/ml Tumor necrosis factor-alpha (TNFα). After 24 h, luciferase activity was measured using the Britelite plus reporter gene assay system (PerkinElmer, cat no. 6066761).

Western blot.

Whole cell lysates: Cells were lysed directly in 4% SDS or 2X sample buffer (Bio-Rad, cat no. 1610737), denatured at 95°C for 10 min, analyzed by SDS-PAGE, and transferred to PVDF membranes. Soluble fractions: Cells were place on ice for 5 min, lysed in M-PER (Thermo Fisher Scientific, cat no. 78501) on ice for 15 minutes, and centrifuged at 13,000 rpm for 15 min at 4°C. Supernatants were carefully transferred to fresh tubes, mixed with sample buffer (ThermoFisher Scientific, cat no. J61337-AC), denatured at 95°C for 10 min, analyzed by SDS-PAGE, and transferred to PVDF membranes. The membrane was incubated for 1 h in blocking solution i.e., 5% bovine serum albumin in 1X Tris-buffered saline plus 0.1% Tween-20 (TBS-T) and probed overnight at 4°C with anti-HA (F-7) mouse IgG2a monoclonal antibody (mAb) (Santa Cruz Biotechnology, cat no. sc-7392X, 1:2,000) or anti-HA (C29F4) rabbit IgG mAb (Cell Signaling Tech. cat no. 3724, 1:1,000), anti-FLAG (D6W5B) rabbit IgG mAb (Cell Signaling Technology, cat no. 14793, 1:1,000), anti-NUTF2 (5A3) mouse IgG2a mAb (Cell Signaling Technology, cat no. 3053S, 1:1,000), anti-Ran (A-7) mouse IgG2b mAb (Santa Cruz Biotechnology, cat no. sc-271376, 1:100), and/or anti-β-actin (C4) mouse IgG1 mAb (Santa Cruz Biotechnology, cat no. sc-47778, 1:500). The next day, the membrane was washed 4X with 1X TBS-T for 5 minutes per wash, and incubated in the dark for 1 hour at room temperature with blocking solution containing the following fluorescent secondary antibodies: Anti-mouse IgG polyclonal antibody (pAb) DyLight 800 (Cell Signaling Technology, cat no. 5257S, 1:15,000), anti-mouse IgG pAb DyLight 680 (Thermo Fisher Scientific, cat no. 35519, 1:10,000), anti-mouse IgG2b pAb IRDye® 800CW (LI-COR Biosciences, cat no. 926–32352, 1:15,000), and anti-rabbit IgG pAb IRDye® 800CW (LI-COR biosciences, cat no. 926–32211, 1:10,000). The membranes were washed again with TBS-T and rinsed twice with 1X TBS. The membranes were then imaged using the Odyssey® M imaging system (LI-COR Biosciences).

Immunofluorescence.

Cells in complete medium were cultured on 8-well Millicell® EZ slides (Millipore Sigma, cat no. PEZGS0816). After 3 days, the cells were placed on ice for 5 minutes and fixed in 4% formaldehyde for 15–20 minutes. After a 15-minute permeabilization with 1X PBS plus 0.1% Triton-X100, the cells were rinsed with cold PBS, and incubated in blocking buffer (Rockland Immunochemicals, cat no. MB-070) for 1 hour at room temperature. The cells were then probed with anti-ER (D6R2W) rabbit IgG mAb (Cell Signaling Technology, cat no. 13258, 1:100) overnight at 4°C. The next day, the cells were washed 4 times with 1X TBS-T, for 5 minutes per wash, and probed for 1 hour in the dark at room temperature with anti-rabbit IgG Alexa Fluor® 647 (Thermo Fisher Sci. cat no. A21245, 1:500). Remaining in the dark, the cells were washed with 1X TBS-T four times, probed with anti-HA (C29F4) rabbit IgG mAb (Cell Signaling Tech. cat no. 3724S, 1:800) overnight at 4°C, washed with TBS-T, and probed for 1 hour at room temperature with anti-rabbit IgG Alexa Fluor® 488 (Cell Signaling Tech. cat no. 4412S, 1:1,000). The cells were washed again with TBS-T and rinsed twice with 1X TBS. Coverslips were carefully mounted on the slides using EverBrite Hardset Mounting Medium with DAPI (Biotium, cat no. 23004), and cured for 24 hours at room temperature. The slides were imaged on a Zeiss LSM 880 confocal microscope and adjusted for presentation using ImageJ/Fiji software version 1.53t.

Proliferation assay.

Cell proliferation assays were performed in 384-well format20. Cells were treated for 5 days with indicated compounds and cell number assayed using the CellTiter-Glo® luminescent cell viability assay (Promega, cat no. G7573) and a standard curve. Luminescence was measured using an Envision plate reader (PerkinElmer).

Quantitative RT-PCR (qPCR).

Cells were cultured for 3 days in 6-well plates containing 2 ml per well of complete medium. After ligand treatment, total RNA was extracted using the Rneasy® mini kit (QIAGEN, cat no. 74106), and reverse transcribed using the High-Capacity RNA-to-cDNA Kit (Thermo Fisher Scientific, cat no. 4387406). Real time PCR was performed on a LightCycler 480 system (Roche Diagnostics), using the TaqMan Fast Advanced Master Mix (Thermo Fisher Scientific, cat no. 4444557) and TaqMan gene expression assays (Thermo Fisher Scientific, cat no. 4331182) for GREB1 (Hs00536409_m1) and GAPDH (Hs02758991_g1). Relative mRNA levels were normalized to GAPDH using the ΔΔCT method51.

CUT&RUN.

MCF-7 cells stably transduced with empty vector or HA-tagged, wild type NUTF2, were placed in phenol red-free DMEM supplemented with 10% csFBS for 3 days and treated with DMSO (vehicle) or 1 μM 4-hydroxytamoxifen (4OHT) for 1 hour. CUT&RUN was performed using the CUTANA kit (EpiCypher, cat no. 14–1048) as previously described52. Specifically, 500,000 cells per reaction were immobilized on concanavalin A beads, permeabilized, and nutated overnight at 4°C with normal rabbit IgG (EpiCypher, cat no. 13–0042, 1:50), anti-ER rabbit IgG pAb (EpiCypher, cat no. 13–2011, 1:100), or anti-HA tag (C29F4) rabbit IgG mAb (Cell Signaling Tech. cat no. 3724, 1:800). The next day, after washing the cells to remove unbound antibody, the cells were incubated with protein A/G fused to micrococcal nuclease (pAG-MNase) for 10 min. The cells were washed again to remove unbound pAG-MNase. Target DNA was cleaved and released by adding CaCl2 and nutating at 4°C for 2 h. The nuclease digest was stopped by incubation with the EDTA-containing Stop Buffer for 10 min at 37°C. Released DNA fragments were purified using SPRIselect beads, eluted with TE buffer, and used to construct DNA libraries, which were then sequenced on the Illumina NovaSeq X series platform to a depth of >10 million reads per sample. Sequences were analyzed using the nf-core cutandrun pipeline (v3.2.2, doi: 10.5281/zenodo.5653535)53, including alignment to the human genome (GRCh38) with Bowtie254, sorting and indexing with SAMtools55, normalization by counts per million mapped reads, and peak calling using MACS2 with a q-value < 0.0156. Reads that mapped to mitochondrial DNA were excluded from further analysis. Differential occupancy analysis was performed using DiffBind (v3.10.1)57. Blacklisted peaks were removed from all samples, then the cognate IgG or empty vector control peaksets were used as “greylists” which were also removed from the ER or NUTF2 peak sets, respectively. The samples were normalized based on sequencing depth and then compared using DESeq2. The criteria for modulation were p < 0.1 for ERα-binding sites and FDR < 0.1 for NUTF2 sites. ER and NUTF2 binding profiles were computed using the dba.plotProfile function in DiffBind and plotted using the plotHeatmap function in deepTools (v3.5.5). De novo DNA motif discovery at ER or NUTF2 binding sites resized to ±50 bp from the center, was performed using rGADEM58. We constructed a position weight matrix (PWM) for the discovered motif and searched the JASPAR database for human transcription factor binding sites with similar PWMs using TFBSTools59,60. We also used MEME-ChIP v5.5.7 (https://meme-suite.org/meme/tools/meme-chip)61 to search the full-sized ±200 bp NUTF2 sites. We further evaluated the candidates by comparing their expression levels observed by RNA-seq in the stable MCF-7 cells.

RNA-seq.

Total RNA was isolated from 3 separate cell passages as biological replicates using the RNeasy kit (QIAGEN) according to the manufacturer’s instructions. RNA was quantified using the Qubit 2.0 Fluorometer (Invitrogen) and run on the Agilent 4200 TapeStation RNA tape (Agilent Technologies) for quality assessment. Only samples with RNA integrity number ≥ 9.5 were used. Messenger RNA was selectively isolated using the NEBNext poly(A) mRNA magnetic isolation module (Cat. #: E7490, NEB) according to the manufacturer’s instructions. Libraries were prepared with NEBNext Ultra II Directional RNA kit (Cat. # E7760, NEB) according to the manufacturer’s instructions. Sequencing data were uploaded to the Galaxy web platform (v2.2.1)62, aligned to the human reference genome hg38 (female) using Hisat2 (v2.2.1)63 and counted using featureCounts in the Subread (v2.0.1) package. Raw and processed RNA-seq datasets were deposited in the gene expression omnibus (GEO), with accession no. GSE230939 and GSE231397. Differential expression analysis was performed using DESeq2 (v1.36.0)64. Gene set and pathway enrichment analysis was performed using the Search Tool for the Retrieval of Interacting Genes/Protein (STRING) database v12.065. Upstream transcriptional regulators were predicted from public ChIP-seq datasets using epigenetic landscape in silico analysis (LISA; http://www.lisa.cistrome.org)66. Gene set and pathway enrichment analysis incorporating CUT&RUN and RNA-seq data was performed using Cistrome-GO67.

Strep-Tactin pull-down.

MCF7 and MCF7-EGFR cells in 10 cm dishes containing 10 ml of phenol red-free DMEM + 10% FBS were co-transfected with 9.0 μg of a plasmid mixture containing 4.5 μg each of empty vector, NUTF2-FLAG, and/or NUTF2-HA fused to an N-terminal Strep-tag®II epitope. After 48 h, the cells placed on ice for 5 min, lifted into 1 ml of 1X PBS, and transferred into microfuge tubes. 100 μl aliquots were transferred to separate tubes (Input). The samples were centrifuged at 1,500 rpm for 3 min at 4°C, and the supernatant was removed. Cells were lysed on ice for 15 min in 500 μl 1X RIPA buffer + protease inhibitor cocktail (PIC), centrifuged at 13,000 rpm for 15 min at 4°C, and transferred to fresh tubes. Input samples were saved at −80°C. For each precipitation, 40 μl of MagStrep “type3” XT magnetic beads (IBA Life Sci. cat no. 2–4090-002 or NC0776437) was placed in a fresh tube, wash twice beads with 500 μl of 1X Buffer W (100 mM Tris/HCl pH 8.0, 150 mM NaCl, 1 mM EDTA), and rotated with 500 μl of cell lysate overnight at 4°C. The next day, the beads were washed three times with 500 μl of 1X Buffer W and resuspended in 35 μl of 2X sample buffer (Bio-Rad, cat no. 1610737EDU). Input samples were lysed directly in 50 μl sample buffer + PIC. The samples were denatured at 95°C for 10 min and analyzed by fluorescent western blot using anti-FLAG (Cell Signaling Technology, cat no. 14793, 1:1,000), anti-HA (Cell Signaling Technology, cat no. 3724, 1:1,000) rabbit IgG, and anti-β-Actin mouse IgG1 (Santa Cruz Biotechnology, cat no. sc-47778, 1:1,000) primary antibodies, and anti-rabbit IgG DyLight 800 (ThermoFisher Scientific, cat no. SA510036, 1:5,000) and anti-mouse IgG DyLight 680 (ThermoFisher Scientific, cat no. 35519, 1:10,000) secondary antibodies.

APEX2 proximity proteomics.

MCF7 cells stably transduced with empty vector, NUTF2-WT, or NUTF2-K4Q, were seeded in 15 cm dishes. After 24 h, the medium was replaced with 15 ml of phenol red-free DMEM + 10% FBS. The cells were then transfected with 12 μg of GFP– or ERα–APEX2 expression vector using 36 μl of ViaFect transfection reagent (Promega, cat no. E4981). The next day, cells were labeled as previously described68, with some modifications. Specifically, the medium was replaced with 12 ml of phenol red-free DMEM +10% FBS +1.5 mM biotinyl tyramide (BT) (Millipore Sigma, cat no. SML2135). After 1 h, the medium was quickly replaced with 10 ml of Activation buffer (1X PBS + 0.5 mM MgCl2 + 1 mM CaCl2, with or without 0.5 mM H2O2) for ~1 min. Labeling was stopped with three 3 ml washes in Stop buffer (1X PBS + 0.5 mM MgCl2 + 1 mM CaCl2 + 5 mM Trolox + 10 mM sodium ascorbate + 10 mM sodium azide). The cells were lifted into 1 ml of stop buffer and transferred to microfuge tubes. The samples were centrifuged at 1,500 rpm for 3 min at 4°C, and dry cell pellets were stored at −80°C.

To verify ERα–APEX2 fusion protein expression and activity, transfected cells were lysed in 30 μl 2X sample buffer (Bio-Rad, cat no. 1610737EDU) + protease inhibitor cocktail (PIC), denatured at 95°C for 10 min, and analyzed by fluorescent western blot with Streptavidin DyLight 800 (ThermoFisher Scientific, cat no. 21851, 1:15,000), anti-ER (D6R2W) rabbit IgG (Cell Signaling Tech, cat no. 13258S, 1:1,000) and anti-β-Actin mouse IgG1 (Santa Cruz Biotechnology, cat no. sc-47778, 1:500) primary antibodies, and anti-rabbit IR Dye 800 CW (LI-COR, cat no. 926–32211, 1:10,000) and anti-mouse IgG DyLight 680 (ThermoFisher Scientific, cat no. 35519, 1:10,000) secondary antibodies.

The cells were washed twice with cold DPBS (4 °C) and resuspended in 1 mL of 1X RIPA with 1X protease inhibitor cocktail. The lysed cells were incubated on ice for 15 minutes and then sonicated (35%, 3 × 10s with 30s rest). The lysate was then centrifuged at 15×1000g for 20 mins at 4 °C and the supernatant collected. The concentration of the cell lysate was measured by BCA assay. 600 μg of each lysate was then incubated with pre-washed streptavidin beads (60 μL, Cytiva, Streptavidin Mag Sepharose) at 4 °C for 16 h. Samples were then washed 3X with 1% SDS in DPBS, 3X with 1 M NaCl in DPBS, and 3X with 10% EtOH in DPBS. Finally, samples are washed 3X with 50 mM Ammonium Bicarbonate buffer.

The beads were resuspended in 500 μL 3 M urea in PBS and 25 μL of 200 mM DTT in 25 mM NH4HCO3 was added. The beads were incubated at 55°C for 30 min. Subsequently, 30 μL 500 mM iodoacetamide in 25 mM ammonium bicarbonate was added and incubated for 30 min at room temperature in the dark. The supernatant was removed and the beads were washed with 3 × 0.5 mL DPBS and 6 × 0.5 mL ammonium bicarbonate (50 mM). The beads were resuspended in 0.5 mL ammonium bicarbonate (50 mM) and transferred to a new protein LoBind tube. The beads were resuspended in 40 μL ammonium bicarbonate (50 mM), 1.2 μL trypsin (1 mg/mL in 50 mM acetic acid) was added and the beads incubated overnight with end-over-end rotation at 37 °C. After 16 hours, an additional 0.8 μL trypsin was added and the beads were incubated for an additional 1 hour at 37 °C. The tubes were centrifuged at 1000g for 1 min and placed on a magnetic rack. The digested peptides were removed, and each biological replicate was split into two technical replicates (6 total for each condition).

Peptide digests were acidified with TFA to 0.1% (v:v) and desalted using 2 μg capacity ZipTips (Millipore, Billerica, MA) according to manufacturer instructions. After drying under vacuum, peptides were resolubilized in 20μL of 0.1% formic acid (FA) to a final concentration of 100ng/μL. Samples were analyzed on a nanoElute (plug-in V2.1.60.0; Bruker, Germany) coupled to a Bruker TimsTOF Pro 2 mass spectrometer (Bremen, Germany), equipped with a CaptiveSpray source and a 20μm zero dead volume (ZDV) Sprayer. Peptides (corresponding to 400ng) were separated on a reversed-phase C18 column (10cm X 75μm X 1.9μm, Bruker PepSep Ten Series). The column temperature was maintained at 50°C using an integrated Bruker Column Toaster (Bremen, Germany). The column was equilibrated using 4 column volumes before loading samples in 100% buffer A (99.9% Fisher Optima® LC/MS water, 0.1% FA), with both steps performed at 800 bar. Samples were separated at 500nl/min using a linear gradient from 3% to 30% buffer B (99.9% Fisher Optima® LC/MS acetonitrile, 0.1% FA) over 17.90 min before ramping to 95% buffer B (0.5min) and sustained at 95% buffer B for 2.4min (total separation method time 20.7min). The Bruker TimsTOF Pro 2 was operated in DIA-PASEF mode using Tims Control v. 5.0.2. Settings for the MS method were as follows: Mass Range 100 to 1700m/z, 1/K0 Start 0.6 V.·/cm2 End 1.4 V·s/cm2, TIMS Ramp and accumulation time 75ms, Capillary Voltage 1700V, Dry Gas 3 l/min, Dry Temp 200°C, DIA-PASEF settings: 18 MS/MS scans (50m/z windons, 0.21 1/K0 windows, total cycle time 0.74), mass range 300 to 1200, and CID collision energy 20eV (at 0.60, 1/K0) to 65eV (at 1.60, 1/K0). The analysis was performed at The Herbert Wertheim UF Scripps Institute for Biomedical Innovation & Technology, Mass Spectrometry and Proteomics Core Facility (RRID:SCR_023576).

Data was processed via DIANN 1.8.1. Parameters were set as follows: trypsin/P digestion, 3 missed cleavages, 3 max. variable modifications, N-term M excision, Ox(M), Ac(N-term) and carbamidomethylation. Peptide length range was 7–30, precursor charge range 1–4, m/z range 300– 1800, and fragment ion range 200–1800. Mass accuracy and MS accuracy were both set to 10. The following settings on the algorithm were checked: “Use isotopologues”, “MBR”, “No shared spectra”, “Heuristic protein inference”. Precursor FDR was set to 1%. A spectral library generated via DIANN from all known human proteins (In-Silico spectral library) was used. Resulting matrix.pg file was opened in Perseus (v2.0.7.0). Intensities were input as “main”, the rest of the descriptors were categorical. Data was then transformed (Log base 2). Normalization was performed via median subtraction. Following this process, a volcano plot was generated utilizing a two tailed t-test for statistical significance. Resulting volcano plots were plotted in GraphPad Prism 9.

Extended Data

Extended Data Fig. 1. Schematic of the nTDP workflow for identification and characterization of proteoforms and their complexoforms within cells.

Extended Data Fig. 1.

a. Soluble proteins are extracted from cells and subject to off-line fractionation with native size-exclusion chromatography (nSEC) to separate intact protein assemblies according to their molecular sizes. Created in BioRender. Nwachukwu, J. (2025) https://BioRender.com/s78v974.

b. Equivalent nSEC fractions from multiple runs are collected, combined, concentrated, and “injected-infused” into the mass-spectrometer using a nano liquid-chromatographic system.

c. Nano electrospray ionization (nano ESI) enables gas-phase ionization of the intact protein assemblies.

d. Field Asymmetric Ion Mobility Spectrometry (FAIMS) separates intact protein ions according to their differential mobility in high and low electric fields.

e. Mass spectrometry (MS) uses a three-stage MS approach. 1) MS1 scan identifies intact protein assemblies in their native states; 2) MS2 scan identifies constituent subunits that are ejected from protein assemblies; and 3) MS3 scan fragments each subunit for proteoform identification. While low energy HCD is used to release subunits from protein assemblies, high energy HCD is required to fragment subunits and obtain structural information of protein primary sequences.

f. Complex-down analysis with ProSight Native software and manual curation defines the full composition of protein assemblies.

g. The crystal structure of the NUTF2 dimer is shown as spheres and colored by monomer, showing complexoform-specific acetylation and methylation of lysine side chains.

Extended Data Fig. 2. MCF-7-EGFR cells are resistant to SERMs and grow faster.

Extended Data Fig. 2.

a. MCF-7 cells were transduced with GFP or GFP/EGFR expression plasmids and sorted for EGFR and side scatter by FACS. The MCF7-EGFR were separated into 4 bins.

b. The sorted MCF-7 cells were analyzed by FACS.

c. The bin with the highest EGFR expression was processed for Western blot along with MFC-7-GFP control cells. Representative of multiple experiments.

d. Expression of EGFR enabled growth inhibition by the EGFR inhibitor, lapatinib. Mean + SEM of n = 16 wells per condition. Representative of 3 experiments with similar results.

e. MCF-7-EGFR Cells were treated with the indicated compounds for 5 days and assayed for cell number. Data is normalized to DMSO vehicle and the starting cell number. Mean +SEM of 4 wells per condition. Representative of multiple experiments.

f. MCF-7 and MCF-7-EGFR cells were grown for 5 days and assayed for proliferation. Day 5 EGFR cells grew significantly faster than the MCF-7 (Student’s two-tailed T test, p = 9 × 10−24). Day 5 was 48 wells per cell type. Representative of two biological replicates.

Extended Data Fig. 3. Fragmentation map of SOD1 with apo and holo fragment ions that was generated using TDValidator software.

Extended Data Fig. 3.

Holo fragment ions are marked in blue and apo fragment ions are marked in purple where backbone cleavages occurred upon ion dissociation.

Extended Data Fig. 4. Intensity of the NUTF2 complexoforms and proteoforms.

Extended Data Fig. 4.

Data are averaged from three replicates.

Extended Data Fig. 5. Protein-NUTF2 interactions in MCF-7 and MCF-7-EGFR cells.

Extended Data Fig. 5.

a. Native MS spectra of the SOD1-NUTF2 and DDT-NUTF2 interactions in MCF-7 and MCF-7-EGFR cells.

b. MS2 spectrum of the SOD1, DDT, NUTF2 species in MCF-7 cells.

c. MS2 spectrum of the SOD1, DDT, NUTF2 monomeric subunits in MCF-7-EGFR cells.

Extended Data Fig. 6. NUTF2 proteoforms in MCF-7 and MCF-7-EGFR cells. MS2 spectra were generated by TDValidator.

Extended Data Fig. 6.

a. Overlayed isotopic distribution of the NUFT2 proteoforms in MCF-7 cells.

b. Overlayed isotopic distribution of the NUFT2 proteoforms in MCF-7-EGFR cells.

Extended Data Fig. 7. Reduced NUTF2 dimerization in EGFR-expressing MCF-7 cells.

Extended Data Fig. 7.

a–b. MCF7 or MCF7-EGFR cells were transduced with the indicated tagged vectors. S2, Strep-Tactin II tag, HA, hemagglutinin tag. Two biological replicates are shown.

a. NUTF2 complexes were isolated with streptavidin beads and characterized with Western blots.

b. Whole cell extracts were probed for the indicated epitopes.

Extended Data Fig. 8. Chromatin occupancy by ER and NUTF2.

Extended Data Fig. 8.

a. ER binding profile in steroid-deprived MCF-7 cells stimulated with vehicle or 10 nM 17β-estradiol (E2) for 1 h and analyzed by CUT&RUN.

b. The consensus ER motif (top), and the motif enriched at ER sites in MCF-7 cells (bottom) are almost identical.

c. CUT&RUN for ER in stable MCF-7 cells treated with vehicle or 1 μM 4OHT for 1 h. Genome browser tracks show ER occupancy in the boxed regions.

d. ER binding profile at some NUTF2-sensitive sites.

e. NUTF2 binding profile at 4OHT-sensitive sites identified by CUT&RUN for the HA epitope tag.

f. The enriched motif at 4OHT-sensitive NUTF2 sites.

N = 3 biological replicates.

Extended Data Fig. 9. Pathway Enrichment incorporating NUTF2-binding sites and target genes identified by RNA-seq.

Extended Data Fig. 9.

Over-represented KEGG pathways were calculated with Cistrome-GO using the mean hypergeometric test with FDR < 0.2.

Supplementary Material

Source data for Table
Supp Data Tables 1-2
SI Doc

Acknowledgements:

The authors thank Dr. Claire Delahunty (Yates Lab), Rosa Viner, Weijing Liu, and John E.P. Syka (Thermo Fisher Scientific) for helpful discussions. Next-generation sequencing was performed at the Genomics Core at the Herbert Wertheim UF Scripps Institute for Biomedical Innovation & Technology. This work was supported by the National Institutes of Health 5R01 AG075862-02, 1R01 HL165168-01, 1R01 AG077046-02, JRY; 5R44 GM130262, KRD; 5P41 GM108569-08 NLK; 1R01CA275142 and R01CA220284, KWN.

Footnotes

Competing Interests: KRD and NLK are involved in the commercialization of top-down proteomics software including ProSight Native. The other authors have no competing interests to disclose.

Data Availability

The mass spectra Raw data files, database, and fragment maps that support the findings of this study is publicly available online at https://massive.ucsd.edu under the accession number MSV000094241. RNA-seq datasets are deposited at GEO GSE230939 and GSE231397. A Life Sciences Reporting Summary and source data accompany this paper.

References

  • 1.Ryan CJ, Kennedy S, Bajrami I, Matallanas D & Lord CJ A Compendium of Co-regulated Protein Complexes in Breast Cancer Reveals Collateral Loss Events. Cell Syst 5, 399–409.e395 (2017). 10.1016/j.cels.2017.09.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Smith LM, Kelleher NL & Proteomics TC f. T. D. Proteoform: a single term describing protein complexity. Nat Meth 10, 186–187 (2013). 10.1038/nmeth.2369 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Jensen MH, Morris EJ, Tran H, Nash MA & Tan C Stochastic ordering of complexoform protein assembly by genetic circuits. PLoS Comput Biol 16, e1007997 (2020). 10.1371/journal.pcbi.1007997 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Gomes FP & Yates JR 3rd. Recent trends of capillary electrophoresis-mass spectrometry in proteomics research. Mass Spectrom Rev 38, 445–460 (2019). 10.1002/mas.21599 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zhou M et al. Higher-order structural characterisation of native proteins and complexes by top-down mass spectrometry. Chem. Sci 11, 12918–12936 (2020). 10.1039/D0SC04392C [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Zhai Z et al. Characterization of Complex Proteoform Mixtures by Online Nanoflow Ion-Exchange Chromatography-Native Mass Spectrometry. Analytical Chemistry 96, 8880–8885 (2024). 10.1021/acs.analchem.4c01760 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Brown KA, Melby JA, Roberts DS & Ge Y Top-down proteomics: challenges, innovations, and applications in basic and clinical research. Expert Review of Proteomics 17, 719–733 (2020). 10.1080/14789450.2020.1855982 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Li H, Nguyen HH, Loo RRO, Campuzano IDG & Loo JA An integrated native mass spectrometry and top-down proteomics method that connects sequence to structure and function of macromolecular complexes. Nature Chemistry (2018). 10.1038/nchem.2908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Skinner OS et al. Top-down characterization of endogenous protein complexes with native proteomics. Nature Chemical Biology, nchembio.2515 (2017). 10.1038/nchembio.2515 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Gault J et al. Combining native and ‘omics’ mass spectrometry to identify endogenous ligands bound to membrane proteins. Nat Meth 17, 505–508 (2020). 10.1038/s41592-020-0821-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Shen X et al. Native Proteomics in Discovery Mode Using Size-Exclusion Chromatography–Capillary Zone Electrophoresis–Tandem Mass Spectrometry. Analytical Chemistry 90, 10095–10099 (2018). 10.1021/acs.analchem.8b02725 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Li H, Wongkongkathep P, Van Orden SL, Ogorzalek Loo RR & Loo JA Revealing Ligand Binding Sites and Quantifying Subunit Variants of Noncovalent Protein Complexes in a Single Native Top-Down FTICR MS Experiment. J. Am. Soc. Mass Spectrom 25, 2060–2068 (2014). 10.1007/s13361-014-0928-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chapman EA et al. Native Top-Down Mass Spectrometry for Characterizing Sarcomeric Proteins Directly from Cardiac Tissue Lysate. J. Am. Soc. Mass Spectrom 35, 738–745 (2024). 10.1021/jasms.3c00430 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Jooß K et al. Separation and Characterization of Endogenous Nucleosomes by Native Capillary Zone Electrophoresis-Top-Down Mass Spectrometry. Anal Chem 93, 5151–5160 (2021). 10.1021/acs.analchem.0c04975 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Mehaffey MR, Xia Q & Brodbelt JS Uniting Native Capillary Electrophoresis and Multistage Ultraviolet Photodissociation Mass Spectrometry for Online Separation and Characterization of Escherichia coli Ribosomal Proteins and Protein Complexes. Analytical Chemistry 92, 15202–15211 (2020). 10.1021/acs.analchem.0c03784 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Chapman EA et al. Structure and dynamics of endogenous cardiac troponin complex in human heart tissue captured by native nanoproteomics. Nat Commun 14, 8400 (2023). 10.1038/s41467-023-43321-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Vimer S, Ben-Nissan G & Sharon M Direct characterization of overproduced proteins by native mass spectrometry. Nature Protocols 15, 236–265 (2020). 10.1038/s41596-019-0233-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Rogawski R & Sharon M Characterizing Endogenous Protein Complexes with Biological Mass Spectrometry. Chemical Reviews 122, 7386–7414 (2022). 10.1021/acs.chemrev.1c00217 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Razavi P et al. The Genomic Landscape of Endocrine-Resistant Advanced Breast Cancers. Cancer Cell 34, 427–438.e426 (2018). https://doi.org/ 10.1016/j.ccell.2018.08.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Min J et al. Dual-mechanism estrogen receptor inhibitors. Proc Natl Acad Sci U S A 118 (2021). 10.1073/pnas.2101657118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Durbin KR et al. ProSight Native: Defining Protein Complex Composition from Native Top-Down Mass Spectrometry Data. Journal of Proteome Research 22, 2660–2668 (2023). 10.1021/acs.jproteome.3c00171 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Chen D, Geis-Asteggiante L, Gomes FP, Ostrand-Rosenberg S & Fenselau C Top-Down Proteomic Characterization of Truncated Proteoforms. J Proteome Res 18, 4013–4019 (2019). 10.1021/acs.jproteome.9b00487 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Pekel G & Ari F Therapeutic Targeting of Cancer Metabolism with Triosephosphate Isomerase. Chem Biodivers 17, e2000012 (2020). 10.1002/cbdv.202000012 [DOI] [PubMed] [Google Scholar]
  • 24.Stein BD et al. LKB1-Dependent Regulation of TPI1 Creates a Divergent Metabolic Liability between Human and Mouse Lung Adenocarcinoma. Cancer Discov 13, 1002–1025 (2023). 10.1158/2159-8290.Cd-22-0805 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Schachner LF et al. Revving an Engine of Human Metabolism: Activity Enhancement of Triosephosphate Isomerase via Hemi-Phosphorylation. ACS Chem Biol 17, 2769–2780 (2022). 10.1021/acschembio.2c00324 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kulkarni YS et al. Enzyme Architecture: Modeling the Operation of a Hydrophobic Clamp in Catalysis by Triosephosphate Isomerase. J. Am. Chem. Soc 139, 10514–10525 (2017). 10.1021/jacs.7b05576 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yüksel KU & Gracy RW In vitro deamidation of human triosephosphate isomerase. Arch Biochem Biophys 248, 452–459 (1986). 10.1016/0003-9861(86)90498-4 [DOI] [PubMed] [Google Scholar]
  • 28.Ugur I, Marion A, Aviyente V & Monard G Why does Asn71 deamidate faster than Asn15 in the enzyme triosephosphate isomerase? Answers from microsecond molecular dynamics simulation and QM/MM free energy calculations. Biochemistry 54, 1429–1439 (2015). 10.1021/bi5008047 [DOI] [PubMed] [Google Scholar]
  • 29.De La Mora-De La Mora I et al. The E104D mutation increases the susceptibility of human triosephosphate isomerase to proteolysis. Asymmetric cleavage of the two monomers of the homodimeric enzyme. Biochim Biophys Acta 1834, 2702–2711 (2013). 10.1016/j.bbapap.2013.08.012 [DOI] [PubMed] [Google Scholar]
  • 30.Schindler L, Dickerhof N, Hampton MB & Bernhagen J Post-translational regulation of macrophage migration inhibitory factor: Basis for functional fine-tuning. Redox Biology 15, 135–142 (2018). https://doi.org/ 10.1016/j.redox.2017.11.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Xiao Z et al. Structure-activity relationships for binding of 4-substituted triazole-phenols to macrophage migration inhibitory factor (MIF). European Journal of Medicinal Chemistry 186, 111849 (2020). https://doi.org/ 10.1016/j.ejmech.2019.111849 [DOI] [PubMed] [Google Scholar]
  • 32.Luedike P et al. Cardioprotection through S-nitros(yl)ation of macrophage migration inhibitory factor. Circulation 125, 1880–1889 (2012). 10.1161/circulationaha.111.069104 [DOI] [PubMed] [Google Scholar]
  • 33.Gomez ML, Shah N, Kenny TC, Jenkins EC & Germain D SOD1 is essential for oncogene-driven mammary tumor formation but dispensable for normal development and proliferation. Oncogene 38, 5751–5765 (2019). 10.1038/s41388-019-0839-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Li H et al. Structural Characterization of Native Proteins and Protein Complexes by Electron Ionization Dissociation-Mass Spectrometry. Analytical Chemistry 89, 2731–2738 (2017). 10.1021/acs.analchem.6b02377 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Schachner LF et al. Reassembling protein complexes after controlled disassembly by top-down mass spectrometry in native mode. Int J Mass Spectrom 465 (2021). 10.1016/j.ijms.2021.116591 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Smith LM et al. A five-level classification system for proteoform identifications. Nat Meth 16, 939–940 (2019). 10.1038/s41592-019-0573-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Małecki JM, Davydova E & Falnes P Protein methylation in mitochondria. J Biol Chem 298, 101791 (2022). 10.1016/j.jbc.2022.101791 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Steggerda SM, Black BE & Paschal BM Monoclonal antibodies to NTF2 inhibit nuclear protein import by preventing nuclear translocation of the GTPase Ran. Mol Biol Cell 11, 703–719 (2000). 10.1091/mbc.11.2.703 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Lui K & Huang Y RanGTPase: A Key Regulator of Nucleocytoplasmic Trafficking. Mol Cell Pharmacol 1, 148–156 (2009). 10.4255/mcpharmacol.09.19 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Stewart M, Kent HM & McCoy AJ Structural basis for molecular recognition between nuclear transport factor 2 (NTF2) and the GDP-bound form of the Ras-family GTPase Ran. J Mol Biol 277, 635–646 (1998). 10.1006/jmbi.1997.1602 [DOI] [PubMed] [Google Scholar]
  • 41.Bayliss R et al. Structural basis for the interaction between NTF2 and nucleoporin FxFG repeats. Embo j 21, 2843–2853 (2002). 10.1093/emboj/cdf305 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Jozwik KM & Carroll JS Pioneer factors in hormone-dependent cancers. Nat Rev Cancer 12, 381–385 (2012). 10.1038/nrc3263 [DOI] [PubMed] [Google Scholar]
  • 43.Hurtado A, Holmes KA, Ross-Innes CS, Schmidt D & Carroll JS FOXA1 is a key determinant of estrogen receptor function and endocrine response. Nat Genet 43, 27–33 (2011). 10.1038/ng.730 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.VanAernum ZL et al. Surface-Induced Dissociation of Noncovalent Protein Complexes in an Extended Mass Range Orbitrap Mass Spectrometer. Analytical Chemistry 91, 3611–3618 (2019). 10.1021/acs.analchem.8b05605 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Larson EJ et al. High-Throughput Multi-attribute Analysis of Antibody-Drug Conjugates Enabled by Trapped Ion Mobility Spectrometry and Top-Down Mass Spectrometry. Analytical Chemistry 93, 10013–10021 (2021). 10.1021/acs.analchem.1c00150 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Fornelli L et al. Accurate Sequence Analysis of a Monoclonal Antibody by Top-Down and Middle-Down Orbitrap Mass Spectrometry Applying Multiple Ion Activation Techniques. Analytical Chemistry 90, 8421–8429 (2018). 10.1021/acs.analchem.8b00984 [DOI] [PMC free article] [PubMed] [Google Scholar]

Methods-only references

  • 47.VanAernum ZL et al. Rapid online buffer exchange for screening of proteins, protein complexes and cell lysates by native mass spectrometry. Nat Protoc 15, 1132–1157 (2020). 10.1038/s41596-019-0281-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Tucholski T et al. A Top-Down Proteomics Platform Coupling Serial Size Exclusion Chromatography and Fourier Transform Ion Cyclotron Resonance Mass Spectrometry. Analytical Chemistry 91, 3835–3844 (2019). 10.1021/acs.analchem.8b04082 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Chernushevich IV & Thomson BA Collisional Cooling of Large Ions in Electrospray Mass Spectrometry. Analytical Chemistry 76, 1754–1760 (2004). 10.1021/ac035406j [DOI] [PubMed] [Google Scholar]
  • 50.DeHart CJ, Fellers RT, Fornelli L, Kelleher NL & Thomas PM Bioinformatics Analysis of Top-Down Mass Spectrometry Data with ProSight Lite. Methods Mol Biol 1558, 381–394 (2017). 10.1007/978-1-4939-6783-4_18 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Bookout AL, Cummins CL, Mangelsdorf DJ, Pesola JM & Kramer MF High-throughput real-time quantitative reverse transcription PCR. Curr Protoc Mol Biol Chapter 15, Unit 15.18 (2006). 10.1002/0471142727.mb1508s73 [DOI] [PubMed] [Google Scholar]
  • 52.Meers MP, Bryson TD, Henikoff JG & Henikoff S Improved CUT&RUN chromatin profiling tools. Elife 8 (2019). 10.7554/eLife.46314 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Ewels PA et al. The nf-core framework for community-curated bioinformatics pipelines. Nat Biotechnol 38, 276–278 (2020). 10.1038/s41587-020-0439-x [DOI] [PubMed] [Google Scholar]
  • 54.Langmead B & Salzberg SL Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357–359 (2012). 10.1038/nmeth.1923 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Li H et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079 (2009). 10.1093/bioinformatics/btp352 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Zhang Y et al. Model-based analysis of ChIP-Seq (MACS). Genome Biol 9, R137 (2008). 10.1186/gb-2008-9-9-r137 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ross-Innes CS et al. Differential oestrogen receptor binding is associated with clinical outcome in breast cancer. Nature 481, 389–393 (2012). 10.1038/nature10730 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Jayaram N, Usvyat D & AC RM Evaluating tools for transcription factor binding site prediction. BMC Bioinformatics 17, 547 (2016). 10.1186/s12859-016-1298-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Tan G & Lenhard B TFBSTools: an R/bioconductor package for transcription factor binding site analysis. Bioinformatics 32, 1555–1556 (2016). 10.1093/bioinformatics/btw024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Rauluseviciute I et al. JASPAR 2024: 20th anniversary of the open-access database of transcription factor binding profiles. Nucleic Acids Res 52, D174–d182 (2024). 10.1093/nar/gkad1059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Machanick P & Bailey TL MEME-ChIP: motif analysis of large DNA datasets. Bioinformatics 27, 1696–1697 (2011). 10.1093/bioinformatics/btr189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2022 update. Nucleic Acids Res 50, W345–w351 (2022). 10.1093/nar/gkac247 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Kim D, Langmead B & Salzberg SL HISAT: a fast spliced aligner with low memory requirements. Nat Meth 12, 357–360 (2015). 10.1038/nmeth.3317 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Love MI, Huber W & Anders S Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 15, 550 (2014). 10.1186/s13059-014-0550-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Szklarczyk D et al. STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Research 47, D607–D613 (2018). 10.1093/nar/gky1131 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Qin Q et al. Lisa: inferring transcriptional regulators through integrative modeling of public chromatin accessibility and ChIP-seq data. Genome Biology 21, 32 (2020). 10.1186/s13059-020-1934-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Li S et al. Cistrome-GO: a web server for functional enrichment analysis of transcription factor ChIP-seq peaks. Nucleic Acids Res 47, W206–w211 (2019). 10.1093/nar/gkz332 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Tan B et al. An Optimized Protocol for Proximity Biotinylation in Confluent Epithelial Cell Cultures Using the Peroxidase APEX2. STAR Protoc 1, 100074 (2020). 10.1016/j.xpro.2020.100074 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Source data for Table
Supp Data Tables 1-2
SI Doc

Data Availability Statement

The mass spectra Raw data files, database, and fragment maps that support the findings of this study is publicly available online at https://massive.ucsd.edu under the accession number MSV000094241. RNA-seq datasets are deposited at GEO GSE230939 and GSE231397. A Life Sciences Reporting Summary and source data accompany this paper.

RESOURCES