Abstract
Precise characterization of proteoforms within the protein corona is essential for developing safer and more effective nanomedicines for diagnostic and therapeutic applications. Here, we advance the characterization of proteoforms within the protein corona by integrating mass spectrometry (MS)-based top-down proteomics (TDP) and bottom-up proteomics (BUP). TDP of the protein corona formed on polystyrene nanoparticles (PSNPs) identifies over 5000 proteoforms of 400 genes from breast cancer plasma samples, creating one of the largest human plasma TDP datasets and the most comprehensive proteoform database of the protein corona. Combining TDP and BUP improves the characterization quality for about 35% of identified proteoforms containing mass shifts, producing a more precise proteoform landscape of protein corona. It also enables the discovery and precise characterization of potential proteoform biomarkers of breast cancer. The approach will eventually enhance our understanding of the protein corona, offer valuable insights into nanoparticle–biosystem interactions, and advance proteoform biomarker discovery.
Subject terms: Nanoparticles, Nanomedicine, Nanobiotechnology
Characterisation of proteoforms within the protein corona is needed to advance diagnostic and therapeutic applications. Here, the authors combining mass spectrometry-based top-down and bottom-up proteomics to produce a more precise proteoform landscape of the protein corona.
Introduction
Nanomedicine employs biocompatible nanoparticles (NPs) in various applications, e.g., targeted drug delivery, imaging, and diagnostics1–8. When NPs are used in these biomedical applications, the biomolecules, predominantly proteins, in the complex biological systems/fluids rapidly interact with NPs, forming a protein corona on their surfaces9. This protein corona influences the interaction between NPs and cells/tissues, determines the biological identity of NPs, and affects the pharmacokinetics and overall efficacy of nanomedicine10. Accurate characterization of the composition and structure of the protein corona is crucial for advancing the field of nanomedicine11.
Mass spectrometry (MS)-based bottom-up proteomics (BUP) has been widely used to profile the proteins within the corona11–13. While BUP offers high peptide-fragmentation coverage and enables the accurate localization of post-translational modifications (PTMs), it cannot accurately identify proteoforms—various forms of protein molecules derived from the same gene—including their combinatorial PTMs. This limitation is due to the enzymatic digestion step, peptide loss, and the peptide-to-protein inference problem in BUP, making the identification of full proteoforms impossible14,15. Proteoforms stemming from sequence variations and PTMs have divergent biological functions16–19 and substantially influence disease progression20–23. For example, human serum albumin (HSA) proteoforms with distinct PTMs exhibit distinct binding interactions with NPs, thereby affecting NP–cell interactions24. Therefore, accurate measurement of proteoforms within the protein corona is vital for advancing our understanding of NP–cell dynamics and discovering proteoform biomarkers.
We established a top-down proteomics (TDP) approach to characterize proteoforms in protein corona by measuring intact proteins without enzymatic digestion in 202425. The TDP approach enabled the identification of full proteoforms and was further enhanced by employing diverse separation techniques and NPs, thereby significantly improving proteoform identification26–29. However, the TDP approach still has some challenges, for example, the backbone cleavage coverage of proteoforms is typically limited by the commonly used collision-based gas-phase fragmentation techniques, making precise localization of PTMs challenging30.
The combination of BUP and TDP approaches combines their complementary strengths for accurate characterization of proteoforms, i.e., precise localization of PTMs, which cannot be achieved by either technique alone (Fig. 1). The TDP approach offers information on proteoform diversity and PTM patterns, and the BUP approach provides high backbone cleavage coverage of peptides, enabling accurate localization and validation of PTMs. In this work, we investigate the synergistic potential of this integrated approach by measuring protein coronas of polystyrene NPs (PSNPs). The PTM-TBA (top-down and bottom-up MS and annotations) software pipeline31 was employed to combine the proteoform-level details from TDP and the peptide-level data from BUP. Our findings indicate that the combined strategy produces unprecedentedly accurate localization of modifications on proteoforms in protein corona. The strategy of combining BUP and TDP may enhance the discovery of proteoform biomarkers and advance our understanding of nanoparticle–biosystem interactions.
Fig. 1. Schematic illustration of enhanced proteoform characterization achieved by integrating top-down and bottom-up proteomics using the PTM-TBA pipeline.
The figure was created using BioRender.com.
Results and discussion
Polystyrene nanoparticles (PSNPs; 75 µL, 25 mg/mL) were incubated with 364 µL of 55% human plasma to generate protein corona-coated NPs following our previously reported protocol25. After isolation and washing, the resulting corona samples were characterized by cryo-transmission electron microscopy (cryo-TEM), dynamic light scattering (DLS), and zeta potential analysis (Supplementary Fig. 1). Verification of minimal aggregation and removal of excess unbound proteins is important to ensure reliable evaluation of protein–NP interactions32. Cryo-TEM analysis revealed well-dispersed PSNPs surrounded by a uniform corona layer, indicating successful and homogeneous protein corona formation (Supplementary Fig. 1a). Consistent with corona adsorption, DLS measurements showed an increase in particle size after plasma incubation, whereas zeta potential measurements demonstrated a reduction in the negative surface charge following corona formation (Supplementary Fig. 1b). These results agree with previous reports describing protein corona formation on NP surfaces33–39, supporting the robustness and reproducibility of the preparation workflow.
The protein corona samples were further analyzed by MS-based TDP and BUP. We have two main aims in this study. The first aim is to align the BUP and TDP data for the protein corona to improve the quality of proteoform characterization (i.e., PTM localization). The second aim is to determine the potential proteoform biomarkers of breast cancer in human plasma by integrating the protein corona-based sample preparation, MS-based TDP, and the combination of BUP and TDP to improve proteoform characterization.
Improving the proteoform characterization by integrating BUP and TDP data of protein corona
To systematically evaluate how the integration of BUP and TDP improves the confidence and depth of proteoform characterization in protein corona studies, extensive BUP and TDP datasets were generated. In this pilot study, protein coronas were prepared from human plasma samples collected from 12 individuals, including three healthy donors, five patients with grade I breast cancer, and four patients with grade II breast cancer. All plasma samples were obtained from age-matched individuals without known comorbidities. Incorporating plasma samples from individuals with different health conditions increased proteomic diversity within the dataset, thereby improving proteoform and peptide coverage in the resulting protein corona analyses13. In addition, the proteoform composition of the coronas may capture interindividual biological variability associated with disease state and personalized physiological differences40.
For this cohort of 12 human plasma samples, 1/3 of each protein corona sample was analyzed by BUP, and 2/3 of each sample was measured by TDP (Fig. 2). For BUP, we first used reversed-phase liquid chromatography (RPLC)-MS/MS to quickly check each sample and analyzed the data using the MaxQuant database search41. Then we employed two-dimensional (2D) RPLC-MS/MS (high-pH and low-pH) to boost the proteome coverage by analyzing the pooled leftover protein corona digests from the 12 human plasma samples. To maximize the PTM information from this study, we employed MSFragger open search42,43 to analyze the data for protein and peptide identifications. For comparison purposes, we also analyzed the data using Proteome Discoverer (2.2) using SEQUEST HT for a regular database search. To improve coverage of phosphorylated peptides in the BUP dataset, we further performed a phosphopeptide enrichment experiment using a TiO2-based approach as described in the literature44. We used a larger cohort of breast cancer human plasma samples to cover as many phosphopeptides as possible, including 9 healthy control samples, 10 grade I breast cancer samples, 10 grade II breast cancer samples, and 10 metastatic breast cancer samples. The human plasma samples were restricted to age-matched individuals with no known comorbidities. After enrichment and RPLC-MS/MS, MSFragger open search was further used to analyze the data to produce a phosphopeptide-enriched peptide dataset. All the peptide datasets from MSFragger open search were used to create a PTM library for matching with the proteoform data from TDP by a bioinformatics tool, PTM-TBA31 (Fig. 2). For producing a large proteoform dataset of protein corona, we employed two different measurement approaches, capillary zone electrophoresis-tandem mass spectrometry (CZE-MS/MS) and RPLC-MS/MS (Fig. 2), to analyze the cohort of 12 human plasma samples, because these two approaches have been well documented for complementary peptide/proteoform identification from complex proteomes25,45–51. In this study, CZE-MS/MS and RPLC-MS/MS analyses were performed in technical triplicate for each protein corona sample. The proteoform data from TDP and the PTM library from BUP were matched using the PTM-TBA tool to improve the PTM localization on proteoforms in the protein corona (Fig. 2).
Fig. 2. The overall design of the workflow integrating TDP and BUP to improve the proteoform characterization in protein corona and discover proteoform biomarkers of breast cancer using the protein corona-based sample preparation and TDP by analyzing a spectrum of human plasma samples covering a range of breast cancer stages.
The PTM-TBA pipeline31 was used to combine BUP and TDP data to improve proteoform characterization. The figure was created using BioRender.com.
By combining CZE-MS/MS and RPLC-MS/MS analyses of protein coronas generated from the 12 human plasma samples, we identified 3503 proteoforms corresponding to 344 genes (Fig. 3a), representing the most comprehensive PSNP protein corona proteoform dataset reported to date. Using a 2D-RPLC-MS/MS workflow together with MSFragger open search, we further identified 4570 protein groups, 45,790 peptides, and 23,632 modified peptides, including glycosylated, phosphorylated, acetylated, oxidized, and deamidated peptides. The number of protein identifications obtained in this study is more than 150% higher than previously reported PSNP protein corona datasets52. The large number of modified peptides enabled the construction of a PTM library for PSNP protein coronas. An additional database search using Proteome Discoverer 2.2 (SEQUEST HT) identified 4504 protein groups, 35,543 peptides, and 3933 modified peptides (Fig. 3d). The smaller number of PTM-containing peptides compared to the MSFragger open search is mainly due to the limited set of PTMs included in the PD search, including oxidation, acetylation, methylation, succinylation, and phosphorylation. Phosphopeptide enrichment experiments further identified 709 phosphopeptides corresponding to 22,270 peptide-spectrum matches (PSMs) from the MSFragger open search. All BUP and TDP datasets are provided in Source data, including Sheets 2 and 3 (TDP) and Sheets 6, 7, 11, and 12 (BUP).
Fig. 3. Proteoform and protein identifications obtained from different analytical platforms and sample groups.
a Bar graphs summarize the numbers of proteoforms and proteoform families identified from human plasma samples by RPLC-MS/MS and CZE-MS/MS. For the RPLC and CZE datasets, bars indicate the average number of proteoforms identified from technical triplicates, and error bars represent standard deviations across the three replicate analyses. The “Individual” category corresponds to the average number of nonredundant proteoforms identified per sample after combining RPLC and CZE results, with error bars showing variation among the 12 plasma samples. The “Total” category represents the overall nonredundant proteoform dataset generated from all samples, separation methods, and technical replicates combined. Data are shown as mean ± S.D. b Violin plots showing molecular weight distributions of proteoforms identified by TDP and proteins identified by BUP. BUP results were generated from pooled plasma samples analyzed by 2D-RPLC-MS/MS with MSFragger open search, whereas TDP results were obtained from the combined CZE-MS/MS and RPLC-MS/MS analyses of the 12 plasma samples. The central dashed line represents the median, while the upper and lower dashed lines correspond to the 75th and 25th percentiles, respectively. c Overview of BUP identifications, including the numbers of protein groups and peptide groups identified per sample. Error bars represent standard deviations from triplicate LC-MS/MS analyses. Protein and peptide identifications were generated using MaxQuant. Data are presented as mean ± S.D. d Total nonredundant protein, peptide, and PTM-containing peptide identifications obtained from 50 offline high-pH RPLC fractions used for the generation of the PTM library. Equal amounts of peptides from all plasma samples were pooled, fractionated by high-pH RPLC, and analyzed by low-pH RPLC-MS/MS. Each fraction was analyzed once. Results are shown for both the MSFragger open-search workflow and Proteome Discoverer 2.2 (PD2.2) using the SEQUEST HT search engine.
The protein mass from BUP is up to 600 kDa, and the TDP data only covers proteoforms smaller than 30 kDa (Fig. 3b), which represents another technical challenge of TDP regarding large proteoform identification. The signal-to-noise ratios of proteoforms are reduced exponentially as the proteoform mass increases due to the much wider charge state distributions of larger proteoforms53. This fact makes the identification of large proteoforms (i.e., >30 kDa) a well-recognized challenge in TDP30. Significant improvement in the TDP workflow is needed to achieve global characterization of large proteoforms in complex samples54,55.
We need to point out that the BUP data for individual protein corona samples from RPLC-MS/MS and MaxQuant database search showed an average of 390 protein groups and 2645 peptide groups per sample (Fig. 3c), totaling 588 unique proteins and 4899 unique peptides across all samples, which agrees well with the literature data regarding the number of protein identifications from PSNP protein corona by RPLC-MS/MS12.
We next integrated the BUP and TDP datasets using PTM-TBA to improve proteoform characterization in the protein corona, particularly for PTM annotation and localization. Unless otherwise noted, the MSFragger open-search BUP datasets, including phosphopeptide enrichment data, were used for the integration. Using this workflow, BUP-derived PTM and mass-shift information were matched with TDP mass-shift data for hundreds of proteoforms, including 480 proteoforms from the CZE-MS/MS dataset (36.6% of the 1312 proteoforms containing mass shifts) and 338 proteoforms from the RPLC-MS/MS dataset (35.3% of the 958 proteoforms with mass shifts) (Fig. 4). The matched proteoform and peptide information are provided in Source data, Sheets 8–10. The integrated BUP-TDP strategy enabled confirmation or annotation of several common PTMs on proteoforms, including oxidation, combined oxidation events, deamidation, acetylation, phosphorylation, and lysine (K) deletion (Fig. 4). For example, 9 proteoforms with a 79–81-Da mass shift were identified in the CZE-MS/MS TDP dataset, suggesting phosphorylation, and five of these phosphorylated proteoforms were confirmed through integration with the BUP data. However, many proteoform mass shifts could not be matched to PTMs identified in the BUP datasets. This limitation is mainly due to two factors: some mass shifts likely arise from combinations of multiple PTMs, which are not currently supported by the present PTM-TBA workflow, and the existing BUP-derived PTM library does not yet cover all PTMs observed in TDP. Further expansion of the PTM library, particularly for important modifications such as phosphorylation and glycosylation, will improve proteoform annotation.
Fig. 4. Distributions of matched mass shifts between BUP and TDP datasets using the PTM-TBA software.

Integrated analysis of BUP and TDP datasets for a the RPLC-MS/MS-based TDP dataset and b the CZE-MS/MS-based TDP dataset. MSFragger open-search BUP results, including datasets generated with phosphopeptide enrichment, were used for PTM matching. Selected commonly identified PTMs are annotated in the plots.
Figure 5 presents four representative examples, showing how integration of BUP and TDP improves proteoform characterization. Detailed annotations, including mass shifts, PTM assignments, and peptide-level evidence, are provided in Source data, Sheet 10_Proteoform Examples. One proteoform derived from myosin-9 (MYH9) showed a + 79.96 Da mass shift that was assigned to serine phosphorylation at residue 1943 through integrated TDP-BUP analysis supported by matching phosphopeptides identified in the bottom-up dataset (Fig. 5d). In another example, a major HDL-associated protein, apolipoprotein A-I (APOA1), displayed a − 128.06 Da mass shift, which was identified as lysine deletion at residue 262 based on the BUP data (Fig. 5c). Another APOA1 proteoform exhibited a + 42 Da mass shift that was confirmed by the integrated analysis as lysine acetylation at residue 250 (Fig. 5a). In addition, a proteoform from apolipoprotein F (APOF) containing a + 48.07 Da mass shift was characterized as triple oxidation at residues 227, 229, and 232 by combining TDP and BUP results (Fig. 5b). Supplementary Figs. 2–4 provide additional examples illustrating improved PTM assignment and localization for proteoforms of transthyretin (TTR) and APOA1. All PTM-TBA results are provided in Source data, Sheets 8–10.
Fig. 5. Representative examples demonstrating enhanced proteoform characterization through integration of BUP and TDP datasets using the PTM-TBA software.
Four representative cases are presented in which combined peptide- and proteoform-level information enabled confident assignment and localization of PTMs or sequence variations on intact proteoforms. Panels show examples of a acetylation, b oxidation, c lysine deletion, and d phosphorylation. The peptide data shown in (d) were obtained from the PD2.2 search workflow, whereas peptide information for all other panels was derived from the MSFragger open-search analysis.
Discovering potential proteoform biomarkers of breast cancer
To accomplish the second main aim of this project, we further studied the proteoform profile differences in the protein corona of PSNPs for human plasma samples from healthy controls and breast cancer patients. To enhance the robustness and accuracy of the protein corona findings, the human plasma samples were restricted to age-matched individuals with no known comorbidities. This selection strategy was employed to minimize biological variability and ensure that the resulting data reflected the specific pathological states under investigation.
The study cohort consisted of 12 human plasma samples, including 3 healthy controls, 5 Grade I breast cancer samples, and 4 Grade II breast cancer samples. Proteoform overlap among the three groups was relatively low (Supplementary Fig. 5), with only 848 of the 3503 identified proteoforms (24%) shared across all conditions. Label-free quantification enabled comparisons of proteoform abundance among healthy, Grade I, and Grade II samples. Differential expression analysis identified 115 proteoforms from 23 genes in the RPLC-MS/MS dataset (Supplementary Figs. 6 and 7) and 31 proteoforms from 10 genes in the CZE-MS/MS dataset (Supplementary Fig. 8). These differentially expressed proteoforms clearly distinguished the disease groups, highlighting the potential of TDP-based protein corona analysis for disease classification. Integration of TDP and BUP data further improved characterization of several differentially expressed proteoforms (Supplementary Fig. 7). One representative example was an apolipoprotein C-II (APOC2) proteoform that was substantially enriched in Grade II samples relative to healthy and Grade I groups. TDP analysis detected a + 16 Da mass shift, consistent with oxidation, while the integrated BUP data localized the modification to methionine oxidation. Methionine oxidation is commonly associated with oxidative stress56, and the increased abundance of this oxidized APOC2 proteoform in Grade II samples may reflect the elevated oxidative environment associated with advanced cancer. Another example involved an apolipoprotein B-100 (APOB) proteoform that was abundant in healthy samples but reduced in both cancer groups. This proteoform contained a + 31.98 Da mass shift that was assigned as dihydroxylation and localized to a specific APOB region using the BUP data. The altered abundance of these PTM-defined proteoforms demonstrates the potential of protein corona proteoform profiling for distinguishing healthy and disease states.
The data from the cohort of 12 human plasma samples fully demonstrates the potential of the proteoform profile of the PSNP protein corona for distinguishing breast cancer and healthy control human plasma samples. The results also motivated us to study a much larger cohort of human plasma samples to determine potential proteoform biomarkers of breast cancer. We analyzed 39 human plasma samples, including healthy controls (9 samples), Grade I (10 samples), Grade II (10 samples), and metastatic breast cancer (10 samples), using the PSNP protein corona approach and RPLC-MS/MS-based TDP. Each protein corona sample was analyzed by RPLC-MS/MS in technical triplicate. In total, 117 MS raw files were acquired. The TDP data were analyzed using the TopPIC suite to identify and quantify proteoforms in the protein corona. In total, 2292 proteoforms from 166 proteins were identified from the 39 human plasma samples. Label-free quantification (LFQ) of proteoforms across the four types of human plasma samples was performed, and the data were further processed using the Perseus software54 to determine the differentially expressed proteoforms among healthy control, Grade I, Grade II, and metastatic breast cancer stages. In total, 118 differentially expressed proteoforms were determined. The proteoforms with statistically significant abundance differences among the four breast cancer conditions can separate the four breast cancer stages well, and each stage has a distinct profile of the differentially expressed proteoforms (Fig. 6a). The control and Grade I were clustered together, and the Grade II and metastatic stages were grouped together according to the abundance profiles of the differentially expressed proteoforms.
Fig. 6. Summary of quantitative TDP data of the cohort of 39 human plasma samples covering healthy control (9 samples), Grade I (G1, 10 samples), Grade II (G2, 10 samples), and metastatic (Met, 10 samples) breast cancer.
a The heatmap and hierarchical clustering analysis show the clear separation of different sample conditions based on the expression profiles of differentially expressed proteoforms across the four biological conditions. The proteoform intensity data were normalized by z-score. b Volcano plot shows the differentially expressed proteoforms between the healthy control and metastatic breast cancer. The Perseus software54 was used to perform the clustering analysis and generate the volcano plot. A two-sided Student’s t test was used to determine the p-values for comparing these two sample groups. To determine the differentially expressed proteoforms, these parameters were applied in the Perseus software: FDR 0.05 and s0 0.1. c The sequence and fragmentation pattern of one differentially expressed phosphorylated proteoform. For these proteoforms, the BUP and TDP data matched by the PTM-TBA software31 regarding the site of phosphorylation.
We then focused on the healthy control and metastatic breast cancer samples to determine the differentially expressed proteoforms after breast cancer metastasis compared to the control. As shown in the volcano plot in Fig. 6b, 89 proteoforms from 12 proteins have statistically significant abundance differences between the two biological conditions. The list of differentially expressed proteoforms is listed in Source data, sheet 13. The 12 proteins include apolipoprotein (A1, C1, C2, and C3), fibrinogen (alpha, beta, and gamma chains), thymosin beta-4, inter-alpha-trypsin inhibitor heavy chain H4, myosin-9, alpha-2-HS-glycoprotein, and kininogen-1. We need to highlight that inter-alpha-trypsin inhibitor heavy chain H4, apolipoprotein, myosin-9, and kininogen-1 are prognostic biomarkers of cancer, and alpha-2-HS-glycoprotein is cancer-related, according to the information from The Human Protein Atlas (https://www.proteinatlas.org/). Here, we identified specific proteoforms of these genes in human plasma samples as potential proteoform biomarkers of breast cancer. Interestingly, all five differentially expressed apolipoprotein proteoforms are much more abundant in metastatic breast cancer samples than in controls. All five differentially expressed proteoforms from inter-alpha-trypsin inhibitor heavy chain H4 have significantly higher abundance in healthy controls compared to metastatic breast cancer. For the differentially expressed proteoforms from the fibrinogen alpha chain and beta chain (74 in total), nearly 96% of them are more abundant in controls. However, we determined three fibrinogen proteoforms (two from the alpha chain and one from the beta chain) that have much higher abundance in metastatic breast cancer (Fig. 6b). The opposite expression profiles of different proteoforms from the same gene between control and metastatic breast cancer highlight the potentially divergent biological functions of these proteoforms. The results also demonstrate the value of MS-based TDP in advancing our understanding of the molecular mechanisms of disease at the proteoform level.
Of the 89 differentially expressed proteoforms between control and metastatic breast cancer, 33 contain mass shifts, and in many cases, we do not know the exact location of the mass shift within the proteoform sequence due to the limited backbone cleavage coverage of the HCD method used. We utilized the PTM-TBA31 tool to match the PTM library built in this work to the differentially expressed proteoforms to improve the proteoform characterization, Fig. 2. Mass shifts on eleven differentially expressed proteoforms were matched with the PTM library, including four proteoforms with oxidation (+16 Da mass shift) and two proteoforms with phosphorylation (+80 Da mass shift). The matched differentially expressed proteoforms are listed in Source data, Sheet 14. The sequence and fragmentation pattern of one matched phosphorylated proteoform of fibrinogen alpha chain is shown in Fig. 6c. The PTM library helped confirm the phosphorylation site on the highlighted serine residue. The results here demonstrate that combining BUP and TDP data is valuable for improving the characterization quality of potential proteoform biomarkers for diseases, e.g., cancer. It is worth pointing out that by combining all the TDP data from the two cohorts of human plasma samples (the 12 and 39 samples), we identified 5121 proteoforms in total from the PSNP protein coronas of 51 human plasma samples, and those proteoforms are derived from over 400 genes. The proteoform dataset represents the most comprehensive proteoform database of PSNP protein corona and the largest TDP dataset of human plasma produced in a single study.
This study presents an integrated BUP and TDP strategy for detailed characterization of the proteoform landscape within the protein corona. The combined workflow substantially improves proteoform characterization, particularly for PTM identification and localization. By leveraging the complementary strengths of TDP for intact proteoform analysis and BUP for peptide-level PTM assignment, the approach provides a higher level of confidence and molecular detail than either method alone. In addition, the PTM-TBA pipeline facilitates integration of the two datasets and supports more accurate PTM annotation and localization in complex biological samples. Our results demonstrate that PTMs play an important role in protein–NP interactions and emphasize the value of proteoform-level analysis in nanomedicine research. The workflow also enables improved characterization of disease-associated proteoform diversity and PTM patterns. Differences observed between healthy and breast cancer plasma samples further support the potential of proteoform profiling for biomarker discovery and personalized nanomedicine applications. Overall, this integrated characterization strategy provides deeper insight into NP–biosystem interactions, NP biodistribution, and proteoform-based biomarker discovery, which may contribute to the future development of safer and more effective nanomedicines.
We also need to note that the quality and comprehensiveness of the PTM library are vital for improving the characterization quality of proteoforms from TDP analysis of protein corona. In our future study, more efforts will be made to improve the confidence of PTM localization in the PTM library and to improve the coverage of common PTMs, e.g., phosphorylation and glycosylation. For our BUP and TDP combination approach, the sample-specific PTM library will be ideal, but the biological condition-specific (i.e., breast cancer) PTM library, or even an organism-specific PTM library, will still be useful for proteoform characterization.
Methods
Chemicals and materials
TPCK-treated trypsin, ammonium bicarbonate (ABC), dithiothreitol (DTT), iodoacetamide (IAA), acrylamide, 3-(trimethoxysilyl) propyl methacrylate, and Amicon Ultra centrifugal filter units (0.5 mL, 10 kDa molecular weight cutoff) were obtained from Sigma-Aldrich. LC/MS-grade water, acetonitrile (ACN), HPLC-grade acetic acid, and fused silica capillaries (75 µm ID, 360 µm OD, and 50 µm ID, 360 µm OD) were purchased from Fisher Scientific. Plain polystyrene nanoparticles (PSNPs, ~100 nm) were acquired from Polysciences.
Human plasma samples
Two independent batches of human plasma samples were used in this study. The first batch included 12 plasma samples collected from healthy female donors and female patients with primary breast cancer at different disease stages and was provided by the Van Andel Institute. Sample collection was conducted under approval from the Van Andel Institutional Review Board following established standard operating procedures. The second batch consisted of 39 plasma samples from healthy female donors and patients with primary or metastatic breast cancer and was obtained from Accio Biobank Online. The use of all plasma samples in this study was approved by the Michigan State University Biomedical and Health Institutional Review Board (Study ID: STUDY00010556).
To improve the robustness of the protein corona analysis and reduce biological variability, both cohorts were limited to age-matched individuals without known comorbidities. This selection criterion was intended to better capture proteomic differences associated with the disease conditions under investigation.
Preparation of protein corona
Two human plasma cohorts containing 12 and 39 samples, respectively, were included in this study. Protein coronas were generated from all plasma samples using a modified nanoparticle incubation workflow12. Briefly, 75 µL of PSNPs (25 mg/mL) were mixed with 364 µL of 55% human plasma prepared by diluting 200 µL of plasma with 164 µL of PBS. The mixtures were incubated at 37 °C for 1 h with gentle agitation to allow corona formation. After incubation, nanoparticle–protein complexes were collected by centrifugation at 14,000 × g for 20 min to remove excess unbound proteins. The pellets were washed twice with 500 µL of cold PBS and finally resuspended in 500 µL of cold PBS. Each sample was then separated into two portions for downstream analyses, with approximately two-thirds (334 µL) used for TDP experiments and the remaining one-third (166 µL) reserved for BUP sample preparation and LC-MS/MS analysis.
NP characterization
The hydrodynamic size and surface charge of the PSNPs were measured before and after protein corona formation using a Zetasizer Nano Series instrument (Malvern Instruments) equipped for dynamic light scattering (DLS) and zeta potential analysis. DLS measurements were performed at room temperature (~25 °C) using a 632 nm He–Ne laser. Multiple replicate measurements were collected for each sample to ensure reproducibility and reliable evaluation of changes in particle size distribution and surface charge following corona formation.
For cryo-EM analysis, 10 nm BSA-coated gold nanoparticles were mixed with corona-coated PSNPs at an optimized ratio of 1:4.3. A 5 µL aliquot of the mixture was deposited onto glow-discharged holey carbon grids (C-Flat R2/2, Protochips, Inc.), followed by blotting and vitrification in liquid ethane using a Vitrobot Mark IV system (Thermo Fisher Scientific, Hillsboro, OR, USA). Cryo-EM imaging was carried out on a Titan Krios 300 kV microscope equipped with a Falcon 2 direct electron detector and phase plate (Thermo Fisher Scientific). Images were collected at a nominal magnification of 75,000× with a calibrated pixel size of 1.075 Å. Defocus values ranged from −2.0 to −3.0 µm, and low-dose imaging conditions were applied to minimize electron beam-induced damage.
Cryo-electron tomography
Single-axis cryo-electron tomography datasets were acquired using a Titan Krios transmission electron microscope (Thermo Fisher Scientific) operated at 300 kV and equipped with a Falcon 2 direct electron detector. Tomographic tilt series were recorded using FEI Batch Tomography software at a magnification of 59,000× across a tilt range from −60° to +60° with 2° increments. The resulting images had a nominal pixel size of 1.375 Å. Defocus settings were maintained between −2 and −3 µm, and the cumulative electron dose for each tomogram was ~80 e⁻/Ų.
Sample preparation for BUP
Protein corona-coated PSNP pellets prepared from each plasma sample (~70 µg total protein per sample) were processed for bottom-up proteomics analysis. Pellets were resuspended in 35 µL of 100 mM ABC (pH 8.0), containing 8 M urea, and incubated at 37 °C for 30 min for protein denaturation. Reduction was performed by adding 5 µL of 70 mM DTT followed by incubation at 37 °C for 30 min. Proteins were subsequently alkylated with 12.5 µL of 70 mM IAA for 20 min at room temperature in the dark, and excess IAA was quenched with 1 µL of 70 mM DTT. Samples were diluted fourfold with 100 mM ABC prior to overnight digestion at 37 °C using 1.5 µg of trypsin. Proteolysis was terminated by adjusting the final formic acid concentration to 0.6% (v/v). Peptides were desalted using Sep-Pak C18 SPE cartridges (Waters, Milford, MA), dried in a vacuum concentrator, and reconstituted in 70 µL of 100 mM ABC buffer (pH 8.0).
Phosphoproteomics sample preparation
Peptide samples generated from the 39 plasma-derived coronas were combined before phosphopeptide enrichment. After desalting and lyophilization, ~250 µg of pooled peptides were enriched using a TiO₂-based immobilized metal affinity chromatography (IMAC) workflow adapted from a previously reported method44. Dried peptides were dissolved in loading buffer containing 70% acetonitrile, 5% TFA, and 1 M lactic acid, then incubated with pre-equilibrated TiO₂ beads (5 µm) using a 4:1 bead-to-peptide mass ratio for 30 min with gentle mixing. To improve phosphopeptide recovery, enrichment was carried out in two consecutive binding steps. After the first incubation, the supernatant was transferred to a fresh batch of TiO₂ beads and incubated again under the same conditions. Following enrichment, the beads were washed sequentially with loading buffer, washing buffer I (80% acetonitrile, 1% TFA), and washing buffer II (30% acetonitrile, 0.5% TFA) to remove non-specifically bound peptides. Bound phosphopeptides were then eluted using two solutions containing 4% ammonium hydroxide with or without 50% acetonitrile. The collected eluates were pooled, vacuum dried, desalted with Sep-Pak C18 cartridges, dried again, and reconstituted in 30 µL of 0.1% formic acid prior to RPLC-MS/MS analysis.
Sample preparation for TDP
Intact protein corona samples were prepared following a workflow adapted from our recent publication27. Briefly, protein corona-coated PSNP pellets generated from each plasma sample in both cohorts (12 and 39 plasma samples) were treated under elution conditions to dissociate corona proteins from the nanoparticle surface. The pellets were incubated in 1% (w/v) SDS elution buffer at 60 °C for 3 h with continuous agitation. After elution, the samples were centrifuged at 19,000 × g for 20 min at 15 °C to pellet the nanoparticles, and the supernatant containing the released proteins was collected. A second centrifugation step under identical conditions was performed to ensure complete removal of residual nanoparticles.
SDS removal and buffer exchange were carried out using Amicon Ultra centrifugal filters (10 kDa MWCO). Before use, the filter units were conditioned with 40 µL of 100 mM ABC buffer (pH 8.0) followed by centrifugation at 14,000 × g for 10 min. Approximately 200 µg of protein sample was then loaded onto the filter and centrifuged for 20 min. To remove SDS, the samples were washed twice with 200 µL of 8 M urea prepared in 100 mM ABC, with centrifugation at 14,000 × g for 20 min after each wash. Residual urea was subsequently removed by three additional buffer exchange steps using 200 µL of 100 mM ABC buffer. Protein concentrations after buffer exchange were measured using a BCA assay (Fisher Scientific) according to the manufacturer’s instructions. Samples were stored at 4 °C overnight prior to MS analysis. The final protein solutions consisted of 60 µL of 100 mM ABC buffer at a protein concentration of ~2.4 mg/mL and were used for both CZE-MS/MS and RPLC-MS/MS experiments.
CZE-MS/MS for TDP of protein corona for the batch of 12 plasma samples
Fused silica capillaries (50 µm ID, 360 µm OD) were coated with linear polyacrylamide (LPA) according to previously reported methods55,57. Following coating, one end of each separation capillary was chemically etched with hydrofluoric acid to reduce the outer diameter to ~100 µm58.
CE-MS/MS analysis was carried out using an ECE-001 autosampler (CMP Scientific, Brooklyn, NY, USA) coupled to a Q Exactive HF mass spectrometer (Thermo Fisher Scientific) through an in-house electrokinetically pumped sheath-flow nanospray interface47,59. Separations were performed using a 1 m LPA-coated fused silica capillary (50 µm ID, 360 µm OD). The background electrolyte consisted of 5% acetic acid (pH 2.4), while the sheath liquid contained 0.2% formic acid and 10% methanol (v/v). Samples were pressure injected into the capillary, and the injection volume (~100 nL) was estimated using Poiseuille’s equation based on a pressure of 5 psi applied for 20 s. A dynamic pH junction strategy was employed to improve sample stacking and separation efficiency. During electrophoresis, 30 kV was applied at the capillary inlet, and an electrospray voltage of 2.0–2.2 kV was supplied through the sheath buffer reservoir. Nanospray emitters were fabricated from borosilicate glass capillaries (1.0 mm OD, 0.75 mm ID, 10 cm length) using a Sutter P-1000 micropipette puller, producing emitter tip openings of ~25–35 µm. Mass spectrometric analysis was performed on the Q Exactive HF in data-dependent acquisition mode. Full MS spectra were acquired over an m/z range of 600–2000 at a resolution of 120,000 (m/z 200) using 3 microscans. The AGC target was set to 3 × 10⁶ with a maximum injection time of 100 ms. Precursors with charge states greater than +5 and intensities above 1 × 10⁴ were isolated using a 2 m/z window and fragmented by HCD with a normalized collision energy of 20%. MS/MS scans were acquired at a resolution of 60,000 with an AGC target of 1 × 10⁶ and a single microscan. Dynamic exclusion was enabled for 30 s with a ± 10 ppm tolerance, and isotopic peaks were excluded from fragmentation.
RPLC-MS/MS for TDP of protein corona for the batch of 12 plasma samples
RPLC-MS/MS experiments were performed using an EASY-nLC 1200 system (Thermo Fisher Scientific). For each analysis, 1 µL of protein corona sample (0.3 mg/mL) was injected onto a custom-packed Bio-C4 capillary column (75 µm ID × 360 µm OD, 20 cm length) packed with 3 µm, 300 Å particles (Sepax Technologies). Protein separation was conducted at 400 nL/min using a binary mobile phase system consisting of 2% ACN with 0.1% FA in water as solvent A and 80% ACN with 0.1% FA as solvent B.
The LC gradient increased from 20 to 100% solvent B over 60 min, followed by a hold at 100% B from 60 to 80 min. An additional 30 min was used for column equilibration and sample loading, resulting in a total runtime of approximately 135 min per analysis. All samples were analyzed in technical triplicate.
Mass spectrometric data were acquired on a Q Exactive HF instrument (Thermo Fisher Scientific) operating in data-dependent acquisition mode. Full MS scans were collected across an m/z range of 600–2000 at a resolution of 120,000 (m/z 200) using 3 microscans. The AGC target for MS scans was set to 1 × 10⁶ with a maximum injection time of 100 ms. Selected precursor ions with charge states greater than +3 and signal intensities above 5 × 10⁴ were isolated using a 2 m/z isolation window and fragmented by HCD with a normalized collision energy of 20%. MS/MS spectra were acquired at a resolution of 120,000 with 3 microscans, an AGC target of 1 × 10⁵, and a maximum injection time of 200 ms. Dynamic exclusion was applied for 30 s using a ± 10 ppm mass tolerance, and isotopic precursor ions were excluded from fragmentation.
RPLC-MS/MS for TDP of protein corona for the batch of 39 plasma samples
RPLC-MS/MS analysis was performed using an EASY-nLC 1200 system (Thermo Fisher Scientific). For each injection, 2 µL of protein corona sample (0.2 mg/mL) was loaded onto a custom-packed Bio-C4 capillary column (100 µm ID × 360 µm OD, 17 cm length) packed with 3 µm, 300 Å particles (Sepax Technologies). Chromatographic separation was carried out at 400 nL/min using solvent A composed of 2% ACN and 0.1% FA in water and solvent B consisting of 80% ACN with 0.1% FA.
A 70 min gradient was used for separation, with mobile phase B increased from 15% to 100% over 60 min, followed by a 10 min hold at 100% B. Each run also included ~30 min for sample loading and column equilibration, resulting in a total runtime of ~100 min per sample. All measurements were performed in triplicate.
MS data were acquired on a Q Exactive HF mass spectrometer (Thermo Fisher Scientific) operated in DDA mode. Full MS scans were collected from m/z 600–2000 at a resolution of 120,000 (m/z 200) with 3 microscans. The AGC target was set to 1 × 10⁶ with a maximum injection time of 100 ms. Precursor ions with charge states above +3 and intensities greater than 5 × 10⁴ were isolated using a 4 m/z isolation window and fragmented by HCD at an NCE of 20%. MS/MS spectra were acquired at a resolution of 120,000 using 3 microscans, with an AGC target of 1 × 10⁵ and a maximum injection time of 200 ms. Dynamic exclusion was enabled for 30 s with a ± 10 ppm mass tolerance, and isotopic precursor ions were excluded from MS/MS selection.
Offline high-pH RPLC peptide fractionation of peptide samples
For offline high-pH fractionation, ~600 µg of pooled peptides generated from the 12 plasma samples were separated using a 1260 Infinity II HPLC system (Agilent Technologies, Santa Clara, CA). Fractionation was performed on a Poroshell 120 EC-C18 column (2.7 µm, 3 × 150 mm) operated at a flow rate of 0.2 mL/min. Mobile phase A consisted of 5 mM ABC in water, and mobile phase B consisted of 5 mM ABC in 80% ACN. Both solvents were adjusted to pH 9 with ammonium hydroxide.
Peptide separation was carried out using a 65 min gradient. Fractions were collected during a linear gradient from 5% to 55% solvent B over 50 min, followed by a ramp from 55% to 100% B over 10 min and a final 5 min hold at 100% B. In total, 50 fractions were collected. All fractions were dried in a SpeedVac concentrator and stored at −20 °C before LC-MS/MS analysis using the workflow described in the “RPLC-MS for BUP of protein corona” section.
RPLC-MS for BUP of protein corona
Peptide LC-MS/MS analysis was carried out using an EASY-nLC 1200 system (Thermo Fisher Scientific) coupled to a home-packed C18 capillary column (75 µm ID × 360 µm OD, 23 cm length) packed with ReproSil-Pur 120 C18-AQ particles (1.9 µm, 120 Å). For each run, 1 µL of protein corona digest (0.4 mg/mL) was injected and separated at a flow rate of 200 nL/min. Mobile phase A consisted of 2% ACN with 0.1% FA in water, while mobile phase B contained 80% ACN with 0.1% FA.
Peptides were separated using a 115 min gradient, increasing from 8% to 55% solvent B over 85 min, followed by a ramp to 100% B at 95 min and a hold at 100% B until 115 min. An additional 30 min was used for sample loading and column equilibration, resulting in a total runtime of ~145 min per sample. All samples were analyzed in triplicate.
MS analysis was performed on a Q Exactive HF mass spectrometer (Thermo Fisher Scientific) operated in DDA mode. Full MS scans were acquired over an m/z range of 350–1800 at a resolution of 60,000 (m/z 400). Precursor ions with charge states greater than +1 and intensities above 1 × 10⁵ were selected for HCD fragmentation using a normalized collision energy of 28%. MS/MS spectra were acquired in the Orbitrap at a resolution of 60,000 using a single microscan. Maximum injection times were set to 50 ms for MS scans and 100 ms for MS/MS scans. A 2 m/z isolation window was used, dynamic exclusion was set to 30 s, and isotopic peaks were excluded from precursor selection.
BUP database search and data analysis
BUP datasets were processed using three different software platforms: FragPipe (v22.0)42,43, Proteome Discoverer 2.2 (PD2.2) with the SEQUEST HT search engine, and MaxQuant (v1.5.5.1)41. FragPipe was primarily employed for MSFragger open searches to generate a PTM library for matching TDP-derived proteoform mass shifts. The FragPipe workflow was applied to both the 2D-RPLC-MS/MS datasets and the phosphopeptide enrichment datasets. PD2.2 was used for conventional database searching to compare against the MSFragger open-search results. MaxQuant analysis was performed only for the individual BUP datasets from the 12 plasma samples to obtain protein and peptide identifications. For MaxQuant analysis, raw MS files were searched against the UniProt human proteome database (UP000005640; 82,733 entries; December 29, 2023 release). A reversed decoy database was included for FDR estimation. Trypsin was specified as the digestion enzyme with up to two missed cleavages allowed. Carbamidomethylation of cysteine residues was set as a fixed modification, whereas methionine oxidation, deamidation of asparagine/glutamine, and protein N-terminal acetylation were included as variable modifications. Peptide-spectrum matches (PSMs) and protein groups were filtered to a 1% FDR. All remaining parameters were kept at default settings.
FragPipe open searches were also conducted against the UniProt human proteome database (UP000005640; 82,733 entries; December 29, 2023 release). A combined target-decoy database containing 165,466 entries was generated within the FragPipe workflow for FDR estimation. Variable modifications included methionine oxidation (+15.9949 Da), N-terminal acetylation (+42.0106 Da), and phosphorylation on serine, threonine, and tyrosine residues (+79.96633 Da), while carbamidomethylated cysteine was specified as a fixed modification. Searches allowed up to three variable modifications per peptide, and precursor mass shifts ranging from −150 to +500 Da were permitted according to the default open-search settings. The “report mass shift as variable modification” option was disabled. PSMs, peptides, and proteins were each filtered using a 1% FDR threshold. Identified PSM and protein lists are provided in Source Data.
For PD2.2 analysis using the SEQUEST HT algorithm, MS data were searched against the same UniProt human proteome database. Search parameters included a precursor mass tolerance of 20 ppm and a fragment mass tolerance of 0.05 Da, with trypsin specified as the digestion enzyme. Variable modifications included methionine oxidation, deamidation of asparagine and glutamine, phosphorylation of serine, threonine, and tyrosine residues, and protein N-terminal acetylation. Carbamidomethylation of cysteine residues was set as a fixed modification. A target-decoy strategy was used for FDR estimation, and peptide and PSM identifications were filtered to 1% FDR. The complete list of identified proteins is available in Source Data.
TDP database search and data analysis
Proteoform identification and quantification were performed using the TopPIC software suite (v1.7.8)60. Raw MS files were first converted into centroided mzML format using MSConvert (v3.0)61. Monoisotopic mass assignment, spectral deconvolution, and proteoform feature extraction were subsequently carried out using TopFD (v1.7.8)62. The resulting processed spectra and extracted features were exported as msalign and txt files, respectively.
Proteoform database searching was conducted with TopPIC (v1.7.8)60 using a custom protein database containing 7576 sequences compiled from proteins identified in the present BUP datasets, together with proteins reported in previous literature studies. Search permitted one unexpected mass shift per proteoform, with precursor and fragment mass tolerances both set to 10 ppm. Unknown mass shifts were limited to a maximum of 500 Da. False discovery rates were estimated using a target-decoy strategy, and both spectrum-level and proteoform-level identifications were filtered to 1% FDR. Proteoform-spectrum matches (PrSMs) were grouped using a 2.2 Da mass tolerance. Proteoforms identified across multiple fractions from the same sample were merged, and redundant identifications were removed by grouping proteoforms originating from the same protein with precursor mass differences below 2.2 Da. Within each group, the proteoform with the best E-value was retained.
To generate the combined proteoform dataset from both RPLC-MS/MS and CZE-MS/MS analyses and across the two plasma cohorts, proteoforms were considered identical if they originated from the same gene, shared the same sequence, and differed in mass by less than 3 Da. These criteria were used to construct the final nonredundant proteoform list.
Label-free quantification of proteoforms was performed using TopDiff (v1.7.8) with default parameters. Quantified proteoforms were further analyzed in Perseus software (v2.0.10.0)54 for differential expression analysis across disease groups. Samples were annotated and grouped according to disease condition. To improve data quality, only proteoforms containing at least 10 valid intensity values per group in the 39-sample cohort and at least 4 valid values per group in the 12-sample cohort were retained. Intensity values were log₂ transformed prior to statistical analysis. Differentially expressed proteoforms were identified using one-way ANOVA with permutation-based FDR correction (FDR = 0.05, S0 = 0.1). Significant proteoforms were selected for downstream analysis and visualized using z-score normalized heatmaps. For the 39-sample cohort, additional comparison between healthy controls and metastatic breast cancer samples was performed using a t test in Perseus. Differentially expressed proteoforms identified from this analysis are listed in the Source Data. Volcano plots were generated using an FDR threshold of 0.05 and S0 value of 0.1.
Matching the BUP and TDP data regarding mass shifts and PTMs
The PTM-TBA workflow (v1.1)31 was used to cross-validate mass shifts and PTM assignments between proteoforms identified by top-down MS and peptides identified from bottom-up MS datasets. Peptide-spectrum match (PSM) information generated from the MSFragger open-search analysis was used to construct the PTM library for matching. Within the PTM-TBA framework, a proteoform mass shift (m1) localized between positions a1 and b1 was considered matched to a peptide mass shift (m2) localized between positions a2 and b2 when the localization regions overlapped ([a1, b1] ⋂ [a2, b2] ≠ Ø) and the mass difference satisfied either |m1 − m2 | <e or 1.00235 − e <|m1 − m2 | <1.00235 + e, where e = 0.1 Da. The 1.00235 Da allowance accounts for a common deconvolution-related precursor mass error observed in top-down MS datasets. For each mass shift identified at the proteoform level, PTM-TBA searched for peptide identifications containing corresponding PTMs or mass shifts. When multiple peptide matches were available, the peptide with the lowest E-value was selected. PTM assignments identified at both the peptide and proteoform levels were subsequently compared for validation.
Because the MSFragger open-search workflow does not provide PTM localization probabilities, additional manual validation steps were performed for PTM-containing peptides matched to proteoforms. Fragmentation spectra were manually inspected to confirm sufficient sequence coverage and adequate fragment ion support for both peptide identification and PTM localization. PTM site assignments on proteoforms were further supported by overlapping localization regions between peptide-level and proteoform-level identifications, improving confidence in PTM localization.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Source data
Author contributions
S.A.S. performed the sample preparation and MS experiments and made the draft of the manuscript. K.L. performed data analysis to combine the TDP and BUP datasets for better proteoform characterization. Y.Y. helped with all the data analysis. R.T.N. helped with the sample preparation. S.G., A.A.S., H.V. and F.F. helped with the protein corona preparation, MS experiments, and/or data analysis. X.L., M.M. and L.S. proposed the study, obtained funding to support the study, oversaw the project, and edited the manuscript. All authors made comments on the manuscript.
Peer review
Peer review information
Nature Communications thanks anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.
Funding
The authors thank the support from the National Institute of General Medical Sciences (NIGMS) through grant R35GM153479 (to L.S.), the National Cancer Institute (NCI) through the grant R01CA247863 (to L.S. and X.L.), MSU College of Human Medicine and Henry Ford Jean P. Schultz Endowed Biomedical Research Funding (to M.M.), and the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) through the grant DK131417 (to M.M.).
Data availability
The MS RAW files about TDP measurement and the corresponding identification lists generated in this study have been deposited in the ProteomeXchange Consortium via the PRIDE63 with the dataset identifier PXD077545. Source data is available for Figs. 3–6, and Supplementary Fig. 2–8 in the associated source data file. Source data are provided with this paper.
Competing interests
M.M. discloses that (1) he is a co-founder and director of the Academic Parity Movement (www.paritymovement.org), a non-profit organization dedicated to addressing academic discrimination, violence and incivility; (2) he is a co-founder of and shareholder in Targets’ Tip, AlbuDerm, and XProteome Inc.; and (3) he receives royalties/honoraria for his published books, plenary lectures and licensed patents. A.A.S. is a co-founder of XProteome Inc. X.L. has a project contract with Bioinformatics Solutions Inc., a company that develops software for MS data processing. The remaining authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Xiaowen Liu, Email: xwliu@tulane.edu.
Morteza Mahmoudi, Email: mahmou22@msu.edu.
Liangliang Sun, Email: lsun@chemistry.msu.edu.
Supplementary information
The online version contains supplementary material available at 10.1038/s41467-026-74306-3.
References
- 1.Riehemann, K. et al. Nanomedicine—challenge and perspectives. Angew. Chem. Int. Ed.48, 872–897 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Bhatia, S. N., Chen, X., Dobrovolskaia, M. A. & Lammers, T. Cancer nanomedicine. Nat. Rev. Cancer22, 550–556 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Pelaz, B. et al. Diverse applications of nanomedicine. ACS Nano11, 2313–2381 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hajipour, M. J. et al. Antibacterial properties of nanoparticles. Trends Biotechnol.30, 499–511 (2012). [DOI] [PubMed] [Google Scholar]
- 5.Patra, J. K. et al. Nano based drug delivery systems: recent developments and future prospects. J. Nanobiotechnol.16, 71 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Mitchell, M. J. et al. Engineering precision nanoparticles for drug delivery. Nat. Rev. Drug Discov.20, 101–124 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Khoee, S. & Sadeghi, A. An NIR-triggered drug release and highly efficient photodynamic therapy from PCL/PNIPAm/porphyrin modified graphene oxide nanoparticles with the Janus morphology. RSC Adv.9, 39780–39792 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Attia, M. F., Anton, N., Wallyn, J., Omran, Z. & Vandamme, T. F. An overview of active and passive targeting strategies to improve the nanocarriers efficiency to tumour sites. J. Pharm. Pharmacol.71, 1185–1198 (2019). [DOI] [PubMed] [Google Scholar]
- 9.Monopoli, M. P., Aberg, C., Salvati, A. & Dawson, K. A. Biomolecular coronas provide the biological identity of nanosized materials. Nat. Nanotechnol.7, 779–786 (2012). [DOI] [PubMed] [Google Scholar]
- 10.Liu, K. et al. Multiomics analysis of naturally efficacious lipid nanoparticle coronas reveals high-density lipoprotein is necessary for their function. Nat. Commun.14, 4007 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Mahmoudi, M., Landry, M. P., Moore, A. & Coreas, R. The protein corona from nanomedicine to environmental science. Nat. Rev. Mater.8, 422–438 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Ashkarran, A. A. et al. Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities. Nat. Commun.13, 6610 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Blume, J. E. et al. Rapid, deep and precise profiling of the plasma proteome with multi-nanoparticle protein corona. Nat. Commun.11, 3662 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Nesvizhskii, A. I. & Aebersold, R. Interpretation of shotgun proteomic data. Mol. Cell. Proteom.4, 1419–1440 (2005). [DOI] [PubMed] [Google Scholar]
- 15.Chick, J. M. et al. A mass-tolerant database search identifies a large proportion of unassigned spectra in shotgun proteomics as modified peptides. Nat. Biotechnol.33, 743–749 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Smith, L. M. & Kelleher, N. L. Proteoforms as the next proteomics currency. Science359, 1106–1107 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Yang, X. et al. Widespread expansion of protein interaction capabilities by alternative splicing. Cell164, 805–817 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McCool, E. N. et al. Deep top-down proteomics revealed significant proteoform-level differences between metastatic and nonmetastatic colorectal cancer cells. Sci. Adv. 8, eabq6348 (2022). [DOI] [PMC free article] [PubMed]
- 19.Adams, L. M. et al. Mapping the KRAS proteoform landscape in colorectal cancer identifies truncated KRAS4B that decreases MAPK signaling. J. Biol. Chem.299, 102768 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Smith, L. M. et al. The Human Proteoform Project: defining the human proteome. Sci. Adv. 7, eabk0734 (2021). [DOI] [PMC free article] [PubMed]
- 21.Forgrave, L. M. et al. Truncated TDP-43 proteoforms diagnostic of frontotemporal dementia with TDP-43. Pathol. Alzheimer’s. Dement.20, 103–111 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Schmitt, N. D. & Agar, J. N. Parsing disease-relevant protein modifications from epiphenomena: perspective on the structural basis of SOD1-mediated ALS. J. Mass Spectrom.52, 480–491 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Tucholski, T. et al. Distinct hypertrophic cardiomyopathy genotypes result in convergent sarcomeric proteoform profiles revealed by top-down proteomics. Proc. Natl. Acad. Sci. USA117, 24691–24700 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Treuel, L. et al. Impact of protein modification on the protein corona on nanoparticles and nanoparticle–cell interactions. ACS Nano8, 503–513 (2014). [DOI] [PubMed] [Google Scholar]
- 25.Sadeghi, S. A. et al. Mass spectrometry-based top-down proteomics in nanomedicine: proteoform-specific measurement of protein corona. ACS Nano18, 26024–26036 (2024). [DOI] [PMC free article] [PubMed]
- 26.Tabatabaeian Nimavard, R., Sadeghi, S. A., Mahmoudi, M., Zhu, G. & Sun, L. Top-down proteomic profiling of protein corona by high-throughput capillary isoelectric focusing-mass spectrometry. J. Am. Soc. Mass Spectrom.36, 778–786 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Sadeghi, S. A. et al. Mass spectrometry-based top-down proteomics for proteoform profiling of protein coronas. Nat. Protoc. 21, 1092–1125 (2026). [DOI] [PMC free article] [PubMed]
- 28.Zhu, G., Sadeghi, S. A., Mahmoudi, M. & Sun, L. Deciphering nanoparticle protein coronas by capillary isoelectric focusing-mass spectrometry-based top-down proteomics. Chem. Commun.60, 11528–11531 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Huang, C.-F. et al. Deep profiling of plasma proteoforms with engineered nanoparticles for top-down proteomics. J. Proteome Res.23, 4694–4703 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Roberts, D. S. et al. Top-down proteomics. Nat. Rev. Methods Prim.4, 38 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Chen, W., Ding, Z., Zang, Y. & Liu, X. Characterization of proteoform post-translational modifications by top-down and bottom-up mass spectrometry in conjunction with annotations. J. Proteome Res.22, 3178–3189 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Sheibani, S. et al. Nanoscale characterization of the biomolecular corona by cryo-electron microscopy, cryo-electron tomography, and image simulation. Nat. Commun.12, 573 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Cedervall, T. et al. Understanding the nanoparticle–protein corona using methods to quantify exchange rates and affinities of proteins for nanoparticles. Proc. Natl. Acad. Sci. USA104, 2050–2055 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Lundqvist, M. et al. Nanoparticle size and surface properties determine the protein corona with possible implications for biological impacts. Proc. Natl. Acad. Sci. USA105, 14265–14270 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Ke, P. C., Lin, S., Parak, W. J., Davis, T. P. & Caruso, F. A decade of the protein corona. ACS Nano11, 11773–11776 (2017). [DOI] [PubMed] [Google Scholar]
- 36.Li, S., Cortez-Jugo, C., Ju, Y. & Caruso, F. Approaching two decades: biomolecular coronas and bio–nano interactions. ACS Nano18, 33257–33263 (2024). [DOI] [PubMed] [Google Scholar]
- 37.Tonigold, M. et al. Pre-adsorption of antibodies enables targeting of nanocarriers despite a biomolecular corona. Nat. Nanotechnol.13, 862–869 (2018). [DOI] [PubMed] [Google Scholar]
- 38.Tenzer, S. et al. Rapid formation of plasma protein corona critically affects nanoparticle pathophysiology. Nat. Nanotechnol.8, 772–781 (2013). [DOI] [PubMed] [Google Scholar]
- 39.Schöttler, S. et al. Protein adsorption is required for stealth effect of poly(ethylene glycol)- and poly(phosphoester)-coated nanocarriers. Nat. Nanotechnol.11, 372–377 (2016). [DOI] [PubMed] [Google Scholar]
- 40.Caracciolo, G. et al. Disease-specific protein corona sensor arrays may have disease detection capacity. Nanoscale Horiz.4, 1063–1076 (2019). [Google Scholar]
- 41.Cox, J. & Mann, M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat. Biotechnol.26, 1367–1372 (2008). [DOI] [PubMed] [Google Scholar]
- 42.Kong, A. T., Leprevost, F. V., Avtonomov, D. M., Mellacheruvu, D. & Nesvizhskii, A. I. MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry-based proteomics. Nat. Methods14, 513–520 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Chang, H.-Y. et al. Crystal-C: a computational tool for refinement of open search results. J. Proteome Res.19, 2511–2515 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Li, Y. et al. Ionic liquid-assisted protein extraction method for plant phosphoproteome analysis. Talanta213, 120848 (2020). [DOI] [PubMed] [Google Scholar]
- 45.Xu, T., Wang, Q., Wang, Q. & Sun, L. Mass spectrometry-intensive top-down proteomics: an update on technology advancements and biomedical applications. Anal. Methods16, 4664–4682 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Fang, F. et al. Quantitative proteomics reveals the dynamic proteome landscape of zebrafish embryos during the maternal-to-zygotic transition. iScience27, 109944 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Chen, D. et al. Recent advances (2019–2021) of capillary electrophoresis-mass spectrometry for multilevel proteomics. Mass Spectrom. Rev.42, 617–642 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Ludwig, K. R., Sun, L., Zhu, G., Dovichi, N. J. & Hummon, A. B. Over 2300 phosphorylated peptide identifications with single-shot capillary zone electrophoresis-tandem mass spectrometry in a 100 min separation. Anal. Chem.87, 9532–9537 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Zhu, G., Sun, L., Yan, X. & Dovichi, N. J. Single-shot proteomics using capillary zone electrophoresis-electrospray ionization-tandem mass spectrometry with production of more than 1250 Escherichia coli peptide identifications in a 50 min separation. Anal. Chem.85, 2569–2573 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Sun, L. et al. Over 10,000 peptide identifications from the HeLa proteome by using single-shot capillary zone electrophoresis combined with tandem mass spectrometry. Angew. Chem. Int. Ed. Engl.53, 13931–13933 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Wang, Q., Wang, Q., Zhu, G. & Sun, L. Capillary electrophoresis-mass spectrometry for top-down proteomics. Annu. Rev. Anal. Chem.18, 125–147 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Hajipour, M. J. et al. An Overview of Nanoparticle Protein Corona Literature. Small19, 2301838 (2023). [DOI] [PMC free article] [PubMed]
- 53.Compton, P. D., Zamdborg, L., Thomas, P. M. & Kelleher, N. L. On the scalability and requirements of whole protein mass spectrometry. Anal. Chem.83, 6868–6874 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Tyanova, S. et al. The Perseus computational platform for comprehensive analysis of (prote)omics data. Nat. Methods13, 731–740 (2016). [DOI] [PubMed] [Google Scholar]
- 55.Zhu, G., Sun, L. & Dovichi, N. J. Thermally-initiated free radical polymerization for reproducible production of stable linear polyacrylamide coated capillaries, and their application to proteomic analysis using capillary zone electrophoresis-mass spectrometry. Talanta146, 839–843 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Suzuki, S. et al. Methionine sulfoxides in serum proteins as potential clinical biomarkers of oxidative stress. Sci. Rep.6, 38299 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Chen, D., Shen, X. & Sun, L. Capillary zone electrophoresis–mass spectrometry with microliter-scale loading capacity, 140 min separation window and high peak capacity for bottom-up proteomics. Analyst142, 2118–2127 (2017). [DOI] [PubMed] [Google Scholar]
- 58.Sun, L. et al. Ultrasensitive and fast bottom-up analysis of femtogram amounts of complex proteome digests. Angew. Chem. Int. Ed. Engl.52, 13661–13664 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Jorgenson, J. W. & Lukacs, K. D. Capillary zone electrophoresis. Science222, 266–272 (1983). [DOI] [PubMed] [Google Scholar]
- 60.Kou, Q., Xun, L. & Liu, X. TopPIC: a software tool for top-down mass spectrometry-based proteoform identification and characterization. Bioinformatics32, 3495–3497 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Kessner, D., Chambers, M., Burke, R., Agus, D. & Mallick, P. ProteoWizard: open source software for rapid proteomics tools development. Bioinformatics24, 2534–2536 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Basharat, A. R., Zang, Y., Sun, L. & Liu, X. TopFD: a proteoform feature detection tool for top–down proteomics. Anal. Chem.95, 8189–8196 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Perez-Riverol, Y. et al. The PRIDE database at 20 years: 2025 update. Nucleic Acids Res.53, D543–D553 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The MS RAW files about TDP measurement and the corresponding identification lists generated in this study have been deposited in the ProteomeXchange Consortium via the PRIDE63 with the dataset identifier PXD077545. Source data is available for Figs. 3–6, and Supplementary Fig. 2–8 in the associated source data file. Source data are provided with this paper.





