Skip to main content
Science Advances logoLink to Science Advances
. 2025 Aug 8;11(32):eadt8974. doi: 10.1126/sciadv.adt8974

Non-native entanglement protein misfolding observed in all-atom simulations and supported by experimental structural ensembles

Quyen V Vu 1,2,, Ian Sitarik 2,, Yang Jiang 2,, Yingzi Xia 3,, Piyoosh Sharma 3, Divya Yadav 3, Hyebin Song 4,5, Mai Suan Li 1,6, Stephen D Fried 3,7,*, Edward P O’Brien 2,4,8,*
PMCID: PMC12333692  PMID: 40779622

Abstract

Several mechanisms are known to cause monomeric protein misfolding. Coarse-grained simulations have predicted an additional mechanism exists involving off-pathway, noncovalent lasso entanglements, which are long-lived kinetic traps and structurally resemble the native state. Here, we examine whether such misfolded states occur in long-timescale, all-atom folding simulations of ubiquitin and λ-repressor. We find that these entangled misfolded states are populated in higher-resolution models. However, because of the small size of ubiquitin and λ-repressor, these states are short-lived. In contrast, coarse-grained simulations of a larger protein, IspE, predict that it populates long-lived misfolded states. Using an Arrhenius extrapolation applied to all-atom simulations, we estimate that these IspE misfolded states have lifetimes similar to the native state while remaining soluble. We further show that these misfolded states are consistent with the structural changes inferred from limited proteolysis and cross-linking mass spectrometry experiments. Our results indicate that misfolded states composed of non-native entanglements can persist for long timescales in both all-atom simulations and experiments.


All-atom simulations populate a predicted class of misfolding, and ensembles are proposed based on structural mass spectrometry.

INTRODUCTION

Various factors are known to cause misfolding in monomeric proteins (Table 1). Recently, high-throughput, coarse-grained simulations of protein synthesis and folding of the Escherichia coli proteome have suggested that there exists an additional widespread mechanism of misfolding (13): Proteins can populate off-pathway misfolded states that involve a change in noncovalent lasso entanglements (1, 2). This type of entanglement is defined by the presence of two structural components: a loop formed by a protein backbone segment closed by a noncovalent native contact and another segment of the protein that is threaded through this loop (46) and, in some cases, wrapped around it multiple times (Fig. 1, A and B). Between 49 and 71% of native structures contain a noncovalent lasso entanglement depending on the organism (7).

Table 1. Mechanisms of protein misfolding.

Protein misfolding can result from factors inherent to the amino acid sequence or from external conditions or processes, see references (5661).

Intrinsic to the primary structure Extrinsic to the primary structure
Proline isomerization Transcription & translation errors
Intrachain domain swapping mRNA splicing errors
Out-of-register β strands Perturbation of temperature, oxidative stress, pH, osmotic pressure, protein concentration, chaperones, processing enzymes
Incorrect helix packing Soluble & insoluble aggregation
Mispacking of side chains Genetic mutations
Alternative secondary or tertiary structure Erroneous complex formation
Backtracking (frustration) Improper/incomplete degradation
Incorrect disulfide bond formation Cofactor and ligand concentrations
Posttranslational modifications
Interchain domain swapping

Fig. 1. Visualizing changes of entanglement through various representations.

Fig. 1.

(A) Illustration of the two geometric elements that compose an entanglement: The closed loop is colored in red, and the threading segment is in blue. The loop is closed by a noncovalent contact between two residues (yellow). Such entanglements occur naturally in the native states of some proteins. In other cases, non-native formation of these entanglements can occur. (B) Three-dimensional (3D) structure of a misfolded entangled state from the protein d-alanine—d-alanine ligase B (DDLB; the closed loop and crossing section of the threading segment of their entangled regions are colored in red and blue, respectively) taken from previously reported coarse-grained simulations (1). Q and G correspond respectively to the fraction of native contacts formed in the structure and the fraction of native contacts that exhibit a change in entanglement. (C) Flattened two-dimensional (2D) structure representation of the entangled state of DDLB. S, β sheet; A, α helix. (D) 3D structure of the crystal structure of DDLB. All native contacts are formed in the crystal structure ( Q = 1), and there is no entanglement change (hence, G=0 ). (E) Flattened 2D structure representation of the native state of DDLB. The coloring scheme of the relevant elements is the same in all panels.

The misfolded states observed in the coarse-grained simulations involved either the gain of a non-native entanglement (i.e., the formation of an entanglement that is not present in the native ensemble; table S1 and fig. S1) or the loss of a native entanglement (i.e., an entanglement present in the native state fails to form; table S1 and fig. S1) (13). This predicted class of misfolding offers an explanation for the decades-old observations that nonfunctional (or less functional) protein molecules can persist for long timescales without aggregating and, in the presence of chaperones, neither get refolded nor degraded (810).

The coarse-grained simulations, in which individual residues are represented by single interaction sites, suggest that these misfolded states are often long-lived because, to correctly fold, they need to change their entanglement state: For example, an entangled state would need to disentangle, which is energetically costly as correctly folded portions of the protein would need to unfold (see folded gray segments in Fig. 1, B and C). Moreover, these states are more likely to bypass cellular quality control mechanisms and remain long-lived in vivo (810) because, in some cases, they are similar in size and structure to the native state and do not expose that much more hydrophobic surface (compare Fig. 1, B and D).

Two criticisms of these findings are that they are based on a model with limited spatial resolution and use force field approximations that have the potential to affect the results. It could be the case, for example, that these states are never populated in a more detailed, transferable all-atom model. Furthermore, the coarse-grained force field previously used was “structure-based,” meaning that the native state is encoded to be favored over other states (11). Both approximations have the potential to affect the accuracy of the earlier findings.

Here, we examine whether an all-atom model of the protein folding process—based on a transferable force field—also exhibits these self-entangled states, as observed previously in coarse-grained models; and if they do, we test whether those states are off-pathway long-lived traps with properties similar to the native state. We show that this type of misfolding occurs in all-atom simulations and that for typically sized proteins, these misfolded states can be very long-lived, soluble kinetically trapped states and then combine information from mass spectrometry and simulations to propose experimentally informed misfolded ensembles.

RESULTS

Ubiquitin and λ-repressor exhibit short-lived entangled states

To address these questions, we analyze previously reported all-atom protein folding trajectories of ubiquitin (12) and the N-terminal domain (NTD) of λ-repressor (13) starting from their native and unfolded states (Fig. 2, A and B). Ubiquitin is a small protein (76 residues) found in eukaryotic organisms and regulates a range of processes including protein degradation and the cell cycle. Ubiquitin folds on the millisecond timescale under typical conditions in vitro (14) and in 3 ms in molecular dynamics simulations (12). The NTD of λ-repressor in the published trajectories (13) is 80 residues in length, binds DNA, and folds on the microsecond timescale (13, 15). Noncovalent lasso entanglements are not present in the native structure of either of these proteins.

Fig. 2. Misfolded gains of entanglement are observed in all-atom protein folding simulations.

Fig. 2.

(A) Ubiquitin’s native structure and (B) λ-repressor’s native structure have no entanglement present. (C) Upper panel: ubiquitin’s fraction of native contacts Q versus time from all-atom simulations. The blue dashed line represents Q=0.6 . Lower panel: ubiquitin’s degree of entanglement G versus time for the same trajectory. (D) Same as (C) but for λ-repressor. (E) 3D structure of a ubiquitin entangled state observed at the time point labeled (i) in (C), with entanglement elements colored as in Fig. 1A. (F) Same as (E) but for λ-repressor and from point labeled (ii) in (D). (G) Flattened secondary structure representation for the entangled state in (E). (H) Same as (G) but for λ-repressor.

We analyzed whether non-native entanglements are populated during the folding trajectories of ubiquitin and λ-repressor totaling 8 ms and 643 μs of simulation time, respectively. To characterize the folding events in the simulations, we used the fraction of native contacts ( Q ) and the order parameter G (defined in Eq. 4, which measures the fraction of native contacts that exhibit a change in entanglement). We observed several non-native entangled states (denoted by nonzero G values, shown in Fig. 2, C and D) during folding transitions, which are indicated by shifts from low to high Q values. Thus, protein self-entangled states are observed in all-atom models. Next, we asked whether these entangled states are off-pathway. We never observe in these trajectories a direct transition from entangled clusters (clusters 1 and 4 for ubiquitin and clusters 2 and 4 for λ-repressor from structural clustering; fig. S2) to the native cluster (clusters 6 and 5, respectively, for ubiquitin and λ-repressor). Hence, these misfolded states are off-pathway—they must unfold to reach the native state (as indicated in fig. S2, these entangled structures always transition to clusters with lower Q and G ). These entangled states clearly are short-lived, lasting just a few nanoseconds according to Fig. 2 (C and D). This is seemingly at odds with the coarse-grained model prediction that these can be long-lived states (2). However, these all-atom simulations were performed at high temperatures (390 K for ubiquitin and 350 K for λ-repressor), near their melting temperatures in the all-atom force field, and configurational transitions are accelerated. Therefore, these entangled states may be able to rapidly disentangle because of these high temperatures.

To test whether these entangled states are near-native, kinetic traps at physiological temperatures, we performed unrestrained molecular dynamics simulations at 310 K (37°C). We selected entangled structures that had at least 60% of their native contacts formed as starting structures, resulting in 21 and 12 structures, respectively, for ubiquitin and λ-repressor. Structurally, these entangled structures are qualitatively similar to each other (table S2); 20 of the 21 ubiquitin structures have closed loops located toward the C terminus, and the threading segment is composed of the N-terminal portion. For λ-repressor, the loop forms toward the C-terminal end and the N terminus threads through it in all 12 structures. Representative structures for these entangled states are shown in Fig. 2 (E and F). Three independent trajectories were started from each of these conformations, and the simulations were run for 700 ns or until the entanglement was lost (i.e., a value of G=0 was reached).

For the 63 trajectories (= 21 × 3) of ubiquitin, 71% (45 of 63) of trajectories persist for 700 ns (an example in fig. S3A and detailed in table S2). On average then, the time to disentangle is ~2.1 μs (estimated using Eq. 6, setting t=700 ns). For λ-repressor, 17% of trajectories (6 of 36) persist in an entangled state for 700 ns (an example is shown in fig. S3B and table S2). The average time to disentangle for λ-repressor is ~390 ns, 50 times faster than for ubiquitin entanglements.

The radii of gyration of these entangled states are comparable to the native state (the differences are less than 10%; table S3), indicating that the entangled and native states are of similar sizes. These entangled states are estimated to have solubility similar to the native state (table S3) according to a model that accounts for the chemical properties of the exposed protein surface (1). Next, we analyzed the secondary structure of proteins using Stride software (16) and found that the non-native entangled states contain up to 83% of the secondary structure found in the native state (table S2). Thus, the size, solubility, and secondary structure analyses indicate that these misfolded states are similar to the native state.

Relative to their protein folding timescales, these entangled states are not long-lived kinetic traps even at 310 K. On the basis of the previous coarse-grained simulation results, this is to be expected (1), because ubiquitin and the NTD of λ-repressor are small single-domain proteins. They are representative of the “model proteins” that have been traditionally studied by biophysicists and demonstrate the ability to rapidly and reversibly unfold and refold (17) but are not representative of the complexity of proteomes (18) [for example, the median protein length among eukaryotes is 361 residues and that among archaea is 247 residues (19)]. In such small proteins, entanglements comprise a large proportion of the total protein structure present—and hence, it is easier to disentangle since most of the protein structure is misfolded and less stable. For example, for the entanglements reported in table S2, up to 78% of ubiquitin’s primary structure is involved in the entanglement (i.e., the minimum distance along the primary sequence required to maintain the entanglement; Fig. 2, G and H), and up to 84% of λ-repressor’s primary structure is involved in the entanglement. In contrast, when a protein of median length misfolds, a larger proportion of the already folded structure will need to unfold to permit disentanglement (Fig. 1B). Unfolding the correctly folded portions of these proteins is energetically more costly, and hence, misfolded states involving non-native entanglements in typically sized proteins are more likely to be long-lived kinetic traps (2).

Misfolded states in a typically sized protein, IspE, are long-lived states in all-atom simulations

Unrestrained all-atom simulations are incapable of folding such large proteins on tractable simulation timescales. Therefore, to test the aforementioned “size-effect” hypothesis, we calculated the lifetime of misfolded states of a larger protein, E. coli’s 4-diphosphocytidyl-2-C-methyl-d-erythritol kinase (283 residues; gene ispE), within an aqueous, all-atom simulation model. This protein was chosen because it was previously identified to exhibit entangled misfolded states in coarse-grained simulations (2, 11). To create models of misfolded IspE for all-atom simulations, we (i) performed protein folding simulations in a coarse-grained model; (ii) carried out structural clustering; (iii) scored the structures within each cluster for their consistency with limited proteolysis mass spectrometry (LiP-MS) and cross-linking mass spectrometry (XL-MS) data (described in the next section and Materials and Methods); (iv) chose, within each misfolded cluster, the single structure that had the highest score; and (v) backmapped those structures to an all-atom representation. These backmapped structures were then used as starting conformations for the aqueous, all-atom temperature jump simulations that were used to estimate their lifetimes at 298 K using an Arrhenius analysis. Five distinct misfolded clusters are observed in the simulations of this enzyme (step ii; Fig. 3A), and their backmapped all-atom representative structures (step v) are visualized in Fig. 3B. An Arrhenius analysis [i.e., performing unfolding simulations at high temperatures (600, 650, 700, 750, and 800 K) and extrapolating the unfolding time to 298 K] is necessary because all-atom simulations cannot reach the timescales associated with unfolding at low temperatures. The resulting unfolding rates exhibited super-Arrhenius behavior (figs. S5 and S6) (20) and were fit to Eq. 11 to estimate the unfolding rate at 298 K. The lifetimes were obtained as the inverse of the extrapolated unfolding rates.

Fig. 3. IspE misfolded states persist for long periods while remaining soluble.

Fig. 3.

(A) −ln(P) surface and structural clusters identified from clustering analysis. (B) Structures selected from each cluster for all-atom temperature jump simulations. Q and G values of the structure are reported below the structure. Entanglement elements colored as in Fig. 1A. Interactive visualizations of these structures are available at https://obrien-lab.github.io/All_Atom_Misfolding_Exp_Ensemble/. Additional structures from each cluster can be found in data S2. (C) Apparent unfolding rates and estimated lifetimes of the structures in (B), derived from Arrhenius extrapolation. (D) Relative aggregation propensities of misfolded clusters compared to the native state. The aggregation propensity of cluster 4 is not significantly different from that of the native cluster, which is indicated by the blue circle. Error bars in (C) and (D) represent the 95% confidence intervals.

To determine whether these misfolded structures are long-lived, we compared their estimated lifetimes to the lifetime of the native cluster in the simulation model. (There are no experimentally reported lifetimes of native IspE; hence, we had to use the lifetime of native IspE within the model.) We find that the native cluster is estimated to have a lifetime of 3 hours (Fig. 3C), corresponding to an unfolding rate of 9.1 × 10−5 s−1. In contrast, structures from clusters 2 to 4 persist for days to months (Fig. 3, B and C), while less structured clusters 0 and 1 last only seconds to minutes. We emphasize that because of the large uncertainties in these lifetime estimates (see 95% confidence intervals in Fig. 3C), only qualitative conclusions can be drawn that are relative to the estimated lifetime of the native cluster. We conclude that in all-atom simulations, misfolded clusters 2 to 4 are likely to be kinetic traps with lifetimes of a similar order of magnitude to the native cluster.

IspE misfolds for over an hour in refolding experiments

We next experimentally tested whether IspE populates misfolded states. To address this, we performed LiP-MS (2124) and XL-MS (2527) on IspE after denaturing and refolding, and compared it to a native IspE sample that was never denatured to detect structural changes such as those associated with changes in entanglement (2, 28). Refolded samples were prepared by first chemically denaturing IspE and then refolding it by dilution from high to low denaturant concentration. Then, the sample was incubated at room temperature for 1 hour before being exposed either to proteinase K (PK) in the LiP-MS experiment or to the cross-linking agent in the XL-MS experiment. We then detect differential signal changes between native (untreated) samples and refolded (treated) samples. We define a signal as significantly different if it exhibits more than a twofold difference in the mean abundance between the refolded and native samples ( log2FC(fold change)1 ) and an adjusted P value, false discovery rate corrected—applied separately to the XL-MS and LiP-MS datasets—using the Benjamini-Hochberg method (29), of less than or equal to 0.05. We observe that there are changes in the proteolytic susceptibility and cross-linking patterns between chemically renatured IspE 1 hour after refolding conditions have been established and a native reference that was not denatured. Specifically, we identified 13 PK cleavage sites from LiP-MS (Table 2 and data S1, sheet labeled “EXP-LiPMS”) and 20 cross-links from XL-MS (Table 3 and data S1, sheet labeled “EXP-XLMS”) experiments that exhibited significant differences in abundance between refolded and native samples. These results suggest that in the case of LiP-MS, the solvent exposure of these regions along the primary structure of IspE was significantly altered in the refolded sample, and in the case of XL-MS, structural changes in the protein led to altered cross-linking propensity between various pairs of residues. We therefore conclude that IspE can populate soluble misfolded conformations over long timescales, qualitatively consistent with our simulation predictions.

Table 2. PK cut sites exhibiting significant differences in abundance between refolded and native samples identified by LiP-MS.

PK site log2FC * Adj.P value PK site log2FC * Adj.P value
R69 1.78 0.028 F207 2.60 0.000
M75 4.14 0.007 A254 2.15 0.016
A78 2.17 0.000 R255 1.53 0.033
I159 2.14 0.002 L258 1.90 0.009
L202 1.46 0.007 Q260 2.47 0.018
K204 1.74 0.038 A261 2.67 0.004
E206 1.82 0.003

*FC: fold change defined as the ratio of the mean refold-sample abundance to native-sample abundance across five replicates each. †Adj. P value: Benjamini-Hochberg adjusted P value.

Table 3. Cross-links exhibiting significant differences in abundance between refolded and native samples identified by XL-MS.

Cross-link log2FC * Adj.P value Cross-link log2FC * Adj.P value
S7-Y16 1.54 0.050 K96-K196 5.33 0.000
K10-K196 7.24 0.003 K96-S208 5.33 0.001
Y16-K196 2.72 0.003 T161-K204 2.19 0.042
K76-T86 4.53 0.000 K196-S198 1.72 0.000
K76-T161 11.93 0.008 K196-K204 3.90 0.000
K76-K196 1.85 0.000 K196-S208 6.83 0.000
K76-S208 5.55 0.014 K196-K271 2.93 0.000
T77-S208 4.74 0.016 K196-S276 1.34 0.042
S81-T86 1.20 0.002 S198-S208 3.87 0.024
S81-K96 1.88 0.000 T201-S208 1.60 0.001

*FC: fold change defined as the ratio of the mean refold-sample abundance to native-sample abundance across five replicates each. † Adj. P value: Benjamini-Hochberg adjusted P value.

Misfolded structural ensembles of IspE that are consistent with LiP-MS and XL-MS

With 33 unique experimental signals reporting on changes in structure in IspE relative to the native state, we can examine which misfolded states seen in our simulations are consistent with these signals. To do this, we take a two-stage approach: First, we simulated the refolding of IspE in our coarse-grained model, and second, we scored the resulting misfolded subpopulations for their overall consistency with the mass spectrometric data (see Materials and Methods). From the simulations, we structurally clustered the resulting conformations and found one cluster corresponding to the native ensemble and five distinct misfolded clusters (Fig. 3A) that vary in their degree of native-structure formation (Q) and changes in entanglement status (G). Scoring each misfolded cluster in terms of its relative change in solvent accessibility for LiP-MS results and solvent-accessible surface distance (SASD) [as calculated with JWalk (30)] for XL-MS results, we find that cluster 3 is the highest scoring with 22 of 33 experimental signals consistent with the structural changes observed in this cluster. (For LiP-MS and XL-MS, respectively, 7 of 13 and 15 of 20 are consistent; data S1 labeled “Comparison-LiPMS” and “Comparison-XLMS”). In comparison, cluster 0 is consistent with 14 of the signals, cluster 2 with 11, and cluster 1 with 9. Cluster 4, which is the most similar to the native cluster, is consistent with only four of the signals.

Across the five misfolded clusters, 29 unique experimental signals can be explained by the structural changes observed in this combination of clusters. One LiP-MS and three XL-MS signals cannot be explained by our methodology. Thus, 88% of the experimental signals can be explained by this ensemble of five misfolded clusters. For interactive visualization, we provide structures from each misfolded cluster in PDB format in data S2.

Most IspE misfolded states are likely to stay soluble

Although most of these misfolded states are long-lived, we next examine whether they are likely to remain soluble or aggregate. To address this, we estimated the relative aggregation propensity of these misfolded clusters compared to the native cluster by incorporating information on the solvent accessibility of aggregation-prone regions (Eq. 14). We first used the AMYLPRED2 server (31) to identify regions in IspE that are prone to aggregation when exposed to solvent. We then calculated the solvent-accessible surface area (SASA) of these regions from our simulated trajectories for each misfolded and native cluster. Our analysis indicates that the aggregation propensity of these misfolded clusters is very similar to that of the native cluster (Fig. 3D), with differences of less than 5%—except for cluster 3, which exhibits a 12.15% relative increase to the native state. Although these small differences are statistically significant (Benjamini-Hochberg adjusted P value ≤ 0.05) for all clusters except cluster 4, the small effect sizes (<5% increase) suggest that most of these misfolded ensembles would remain soluble on timescales similar to the native ensemble.

DISCUSSION

Coarse-grained simulations predicted the existence of an additional widespread mechanism of protein misfolding. Motivated by the possibility this may be an artifact of those simulations’ resolution and structure-based force field, in this study, we examined whether protein folding modeled at all-atom resolution also exhibits this type of misfolding. We have demonstrated that ubiquitin and λ-repressor populate these entangled misfolded states in all-atom folding simulations. Thus, protein folding involving a change of entanglement status occurs regardless of model resolution, lending support to the validity of the earlier conclusions drawn from coarse-grained modeling.

Because of the small size of ubiquitin and λ-repressor, the entangled states they populate are short-lived, persisting for just a few microseconds. For larger proteins, all-atom simulations indicate that these misfolded states can persist for much longer periods—on timescales similar to the lifetime of the native state while remaining soluble. This relationship between protein size and misfolded lifetimes arises from the fact that a more stabilizing native structure can form around the localized misfolded elements in larger proteins as compared to smaller proteins, creating a larger energy barrier to unfold improperly folded portions.

Experimentally, we observed that IspE can populate misfolded structures that form after dilution from the denaturant and remain misfolded for over an hour. This result is consistent with our simulations of IspE (Fig. 3, A to C), giving us more confidence in our broader conclusions. When cross-referencing the misfolded signals from LiP-MS and XL-MS experiments with the misfolded clusters from the IspE coarse-grained simulations, we find that almost 90% of the experimental signals can be explained by the simulated structural changes. Thus, the structures we have identified in Fig. 3B and data S2 represent an experimentally informed misfolded ensemble involving soluble, long-lived states.

Misfolding involving changes of entanglement status can arise from either a gain of a non-native entanglement or a loss of a native entanglement. In ubiquitin and λ-repressor, which contain no noncovalent lasso entanglements in their native state, it is only possible to misfold through the formation of a non-native entanglement (gain of entanglement), which we observed during their folding. In contrast, IspE has two unique representative, noncovalent lasso entanglements present in its native structure (table S5). In this system, we observed misfolded structures involving either gains of entanglement, losses of entanglement, or combinations thereof. Both classes of misfolding are manifestations of topological frustration. In the one case, a protein backbone segment is improperly positioned outside a backbone loop, corresponding to a loss of entanglement; in the other case, it is improperly positioned within a backbone loop, corresponding to a gain of entanglement. The consequence of both of these types of topological traps is that they both must backtrack (32) out of these states, meaning that already natively structured portions of the protein must unfold to allow proper folding. Small proteins that misfold in this way have additional folding pathways that bypass backtracking that include threading of the loop in cases in which the native entanglement failed to form or unthreading of a loop in misfolded states involving formation of non-native entanglements (1, 33). However, for typically sized proteins, such as IspE, backtracking will often be the dominant pathway because substantial portions of the structure can still properly fold.

These changes of entanglement that lead to misfolded, long-lived kinetic traps expand the mechanisms of topological frustration that give rise to the kinetic partitioning mechanism for understanding biomolecular self-assembly in proteins and nucleic acids (3436). In the kinetic partitioning model, frustration is inherent to the nature of biopolymers, because interactions at local and global length scales can be in energetic conflict, giving rise to different structural preferences. Particular packing arrangements, for example, of hydrophobic residues or secondary structural elements might be energetically favored locally but disfavored in the presence of tertiary interactions (34). As a consequence, near-native misfolded states can be populated, especially for larger proteins, which will tend to exhibit greater frustration. The current study adds to the structural mechanisms by which such frustration can arise, specifically through the formation of native-like states that either fail to form their native entanglement or form a non-native entanglement.

Proteins containing entanglements in their native state are widespread across different organisms. The proportion of globular proteins with them ranges from 49% in humans to 71% in E. coli (7). Perhaps expectedly then, this type of misfolding has been found to be common based on high-throughput coarse-grained simulations of the E. coli proteome (2). Furthermore, because structure equals function, such misfolding has been implicated in influencing ensemble-averaged functioning of proteins like specific activity (1) and, in some cases, in bypassing the chaperone machinery in cells depending on how native-like in structure are these misfolded states (3). We hypothesize that this understudied class of protein misfolding will offer insights into protein function and homeostasis, including new causal mechanisms for disease and aging that can open up new therapeutic targets.

MATERIALS AND METHODS

All-atom simulations of ubiquitin and λ-repressor

Initial entangled conformations of the protein were placed at the center of a rectangular periodic box with the minimum distance between the box edge and protein atoms of 1.5 nm, the system was solvated with the TIP3P water model (37), and NaCl was added to neutralize and mimic relevant salt concentrations [ubiquitin: 0 mM; λ-repressor: 50 mM; the salt concentration was added according to original studies (12, 13)]. Energy minimization was then carried out using the steepest descent algorithm (38, 39) to minimize steric conflict. The systems were then gradually heated from 2 to 310 K over 1 ns. After that, the system was equilibrated for 1 ns in the NVT ensemble and 1 ns in the NPT ensemble with all heavy atoms restrained using a harmonic potential with a force constant of 1000 kJ/mol per square nanometer before it went into production run in the NPT ensemble. The particle mesh Ewald method (40) was used to calculate long-range electrostatic interactions beyond 12 Å, and the Lennard-Jones interactions were calculated with a cutoff distance of 12 Å and applied, smoothly switching the forces to zero between 10 and 12 Å. Temperature and pressure were maintained at 310 K and 1 atm using a Nose-Hoover thermostat (41, 42) and Parrinello-Rahman barostat (43), respectively. The LINCS algorithm (44) was used to constrain all bonds involving hydrogen atoms, and the integration time step was set to 2 fs. Simulations were performed using GROMACS 2018 (45) with the CHARMM36m force field (46). This procedure was used in all the new simulations carried out in this study.

Characterizing changes in entanglement

Calculating changes in entanglement has been reported elsewhere (13). Here, we briefly describe our procedure. To detect noncovalent lasso entanglements, we use linking numbers, which require at least one closed loop. We define this loop as being composed of the backbone trace connecting residues 𝑖 and 𝑗 that form a noncovalent contact in the given protein conformation. The looped portion (red segment in Fig. 1A) is identified if the Cα coordinates of residues 𝑖 and 𝑗 (yellow spheres in Fig. 1A) are closer than 8 Å and ij10 residues. Outside this loop is an N-terminal segment composed of residues 1 through 𝑖 − 1, and there is a C-terminal segment composed of residues 𝑗 + 1 through 𝑁. These two segments represent open curves, whose entanglement through the closed loop was characterized by linking numbers denoted 𝑔N and 𝑔C. For a given structure of an N-residue protein, with a contact present at residues (𝑖, 𝑗), the average coordinates R𝑙 and the gradient dR𝑙 of point 𝑙 on the curves were calculated as

{Rl=12(rl+rl+1)dRl=rl+1rl (1)

where r𝑙 is the coordinates of the Cα atom in residue 𝑙. The linking numbers 𝑔N(𝑖, 𝑗) and 𝑔C(𝑖, 𝑗) of N and C termini, respectively, were calculated as

{gN(i,j)=14πm=6i5n=ij1RmRnRmRn3(dRm×dRn)gC(i,j)=14πm=ij1n=j+4N6RmRnRmRn3(dRm×dRn) (2)

where we excluded the first five residues on the N-terminal curve, the last five residues on the C-terminal curve, and four residues before and after the native contact to eliminate the error introduced by both the high flexibility and contiguity of the termini and trivial entanglements in the local structure. It is worth noting that partial linking values between 0.5 and 0.7 may not signify real entanglements and, thus, those have to be checked manually (47). The above integrations yield two noninteger values, and the total linking number for a contact (𝑖, 𝑗) is estimated as the sum of N- and C-terminal linking numbers

g(i,j)=round[gN(i,j)]+round[gC(i,j)] (3)

The use of the round function in Eq. 3 serves to discretize the computed linking numbers, which can result in noninteger values because of numerical integration between a closed loop and an open segment. By applying the round function, we map such values to discrete integers (for example, g = 0.6 → 1), effectively capturing the presence or absence of an entanglement in a robust and consistent way. This discretization allows us to perform meaningful comparisons between the entanglement states of different structures. For instance, even if both a conformation in the simulated trajectory and the native structure are entangled, their computed linking values may slightly differ (for example, g = 0.6 versus g = 0.7), yet both indicate the same qualitative topological feature (i.e., one event of terminal threading). The round function ensures that such topologies are treated identically in our analysis.

Comparing the absolute value of the total linking number for a contact (𝑖, 𝑗) at a given conformation to that of a reference state (i.e., native state) allows us to ascertain a gain or loss of linking between the backbone trace loop and the terminal open curves as well as any switches in chirality. There are six changes in linking cases we should consider (table S1 and fig. S1) when using this approach to quantify entanglement.

We defined the G metric, the degree of entanglement, as a time-dependent order parameter that reflects the extent of the topological entanglement changes in a given structure compared to the native structure and is calculated as

G(t)=1M(i,j)Θ[(i,j)NCg(i,j,t)gnative(i,j)] (4)

where (𝑖, 𝑗) is one of the native contacts in the native crystal structure; NC is the set of native contacts formed in the current structure at time 𝑡; 𝑔(𝑖, 𝑗, 𝑡) and 𝑔native(𝑖, 𝑗) are, respectively, the total linking number of the contact (𝑖, 𝑗) at time 𝑡 and native structures estimated using Eq. 3; M is the total number of native contacts within the native structure; and the selection function Θ equals 1 when the condition is true and equals 0 when it is false. The larger G is, the greater the number of native contacts that have changed their entanglement status relative to the reference state.

Disentanglement time constant from unrestrained simulations at physiological temperature

The disentangled time constant can be estimated through the connection with the portion of disentangled trajectories

Pdis(t)=1etτ (5)
τ=tln[1Pdis(t)] (6)

where Pdis(t) is the portion of disentangled trajectories at time t. τ is a time constant that entanglement disappears. We use this equation to estimate the average lifetime of entangled states of ubiquitin and λ-repressor by setting t = 700 ns.

Secondary structure similarity definition

The secondary structure similarity in table S2 is defined as the fraction of residues that are in the correct secondary element at the current structure compared to the native structure; we only count residues in the α helix or β sheet. The secondary structural elements of protein are assigned by the STRIDE program (16).

Calculation of solubility of the entangled structures

The solubility of the entangled structures in table S3 is defined as the percent soluble protein estimated via the insolubility propensities of structure 𝑖 ( χisol ), a fully disordered structure, which has no secondary and tertiary structure elements ( χdisorderedinsol ), and the minimum propensity values of all structures [ min(χisol)]

fisol=χdisorderedinsolχiinsolχdisorderedinsolmin(χiinsol) (7)

The insolubility was estimated by considering the aggregation propensity ( χagg ) and degradation propensity ( χdeg ) minus the 70-kDa heat shock protein (Hsp70) binding propensity ( χHsp70 ) that is considered to prevent the misfolding protein from aggregation

χinsol=χagg+χdegχHsp70 (8)

For a given entangled structure, the aggregation, degradation, and Hsp70 binding propensities were estimated as the relative change of the SASA of the aggregation-prone, degradation-prone, and Hsp70 binding regions of the entangled state 𝑖 against the SASA of the native structure

χiT=SASAiTSASAnativeT1 (9)

where T can be aggregation, degradation, or Hsp70 binding propensities. The aggregation-prone region was predicted by the AmylPred2 server (31). The degradation-prone region was defined as the hydrophobic residues (Ile, Val, Phe, Cys, Met, Ala, Gly, and Trp). The Hsp70 binding region was predicted by using ChaperISM (48).

Coarse-grained protein folding simulations of IspE protein and structural clustering

We use a previously published Go̅-based coarse-grain model (1) in which each amino acid is represented by a single interaction site placed at the Cα atom with a specific van der Waals radius for each amino acid. Initial conformations for refolding simulations were generated by heating the native state of the protein to 1000 K for 15 ns. The final conformation from heating was then temperature quenched at 310 K to initialize refolding. All simulations were carried out using a Langevin thermostat at a temperature of 310 K, with a time step of 15 fs and a friction coefficient of 0.050 ps−1. All simulations were carried out using OpenMM version 7.7 (49). Trajectories were saved after every 5000 steps; a total of 56 independent trajectories of 1.5 μs was performed. The last 500 ns from each trajectory was used for structural clustering on the basis of two order parameters that capture the nativeness of the structures (fraction of native contacts, Q) and the changes in self-entanglement of the protein (fraction of native contact change in entanglement, G). The fraction of native contacts at each conformation was defined as the number of native contacts formed divided by the total number of native contacts in the crystal structure. Two residues (𝑖, 𝑗) were considered to form a native contact if the distance between their Cα atoms was less than 1.2 times their distance in the crystal structure, and ij>3 . The fraction of native contact change in entanglement (G) was calculated as described in Eq. 4. Using the K-mean++ clustering algorithm (50), 100 microstates were grouped from the coarse-grained simulations. These microstates are then coarse grained into structural states using the PCCA+ algorithm (51). To determine a reasonable number of clusters for the PCCA+ analysis, we used a timescale separation criterion based on identifying a large gap in the implied timescale (ITS) spectrum derived from eigenvalues of the transition probability matrix (52). Specifically, we calculated the ITSs, sorted them in descending order, and computed the ITS ratio, defined by the ratio ITSi/ITSi+1 . A large ratio indicates that the corresponding dynamical processes are kinetically well separated, while a ratio close to 1 suggests that the processes can be grouped together. All clustering analyses were performed by using the PyEmma package (53).

Structure selection for Arrhenius analyses

The last 500-ns simulation structures were labeled using the cluster ID and the experimental signals they produce. Here, we identify a simulated structure capable of generating an experimental signal only if it demonstrates a SASA greater than the 97.5th percentile or less than the 2.5th percentile of all structures in the native ensemble for a LiP-MS signal or exhibits a cross-linking propensity score outside the corresponding percentiles for an XL-MS signal. These groups represent the misfolded structural ensemble that can generate the observed LiP-MS and/or XL-MS signals.

For each of the clusters 0 to 4, we choose a single structure that can generate the greatest number of experimental signals. If there is a tie (multiple structures generate the same number of consistent experiment signals), preference is given to the structure with the highest microstate probability, followed by the highest Q and G values.

In silico temperature jump simulations and Arrhenius analysis of 4-diphosphocytidyl-2-C-methyl-d-erythritol kinase

To estimate the lifetime of the representative structure of each metastable state, we calculate the time it takes for this structure state to unfold. It is not feasible to trace the unfolding process of such large protein in conventional simulation; hence, we performed simulation at high temperatures, 600, 650, 700, 750, and 800 K, and then used Arrhenius analysis to extrapolate the unfolding at 298 K. The simulation protocol was similar to that used for ubiquitin and λ-repressor, with the exception that the salt concentration was set at 150 mM NaCl. For the production runs, at each temperature, we performed 50 independent trajectories in the NVT ensemble and then got the list of first-passage times that the structure is unfolded. To account for fluctuation, a trajectory was considered to reach the unfolded state if Q0.3 for 1 ns. The starting structures are highly native-like (Q > 0.7; see Fig. 3B in the main text), so there is an initial delay time during which native contacts are partially lost, but the system has not yet reached the unfolding threshold. This results in a delay time in the survival probability curves, during which the fraction of unfolded conformations remains zero (fig. S4). Once sufficient disruption of native contacts occurs, the survival probability begins to decay exponentially. The survival probability was then fitted to a single-exponential function to find the fitting parameters (t0 and k)

SU(t)={1,0t<t0ek(tt0),tt0 (10)

where t0 is the delay time of entanglement (fig. S4). The apparent unfolding rate at temperature T is kapp(T)=1t0+1k.

The apparent unfolding rates are found to have a super-Arrhenius behavior (figs. S5 and S6) (20, 54) and are then fitted as the function of T

ln(kapp)=aT2+bT+c (11)

Here a, b, and c are the free fitting parameters. We then extrapolated to obtain the apparent unfolding rate at the target temperature of 298 K. The extrapolated unfolding time in simulation is τsim=1kapp.

The uncertainties were estimated using bootstrapped resampling with 104 iterations. For each iteration, at each temperature, we randomly selected with replacement 50 samples from the list of mean first-passage times of the original data at that temperature and then repeated the fitting procedure on the resampled data. Resampled datasets were discarded if the coefficient of determination R2<0.9.

Limited proteolysis experiments

To perform limited proteolysis, 2 μl of PK (from Tritirachium album) stock solutions (prepared at a concentration of 0.1 mg/ml in native buffer with 20% glycerol and then stored at −20°C for future use) was added to either native or refolded samples, which were prepared on a 200-μl scale. The reaction mixtures were rapidly mixed using a pipette [the enzyme:substrate ratio is 1:100 (w/w)]. Samples were incubated for 1 min before being transferred to a preheated mineral oil bath at 105°C to quench the PK activity. The samples were left in the oil bath for 5 min. After quenching, the proteolyzed samples were transferred to a microcentrifuge tube containing 152 mg of urea to reach a final concentration of 8 M. The limited proteolysis experiments were conducted in five replicates for both native and refolded IspE samples, resulting in a total of 10 samples.

Cross-linking experiments

Two hundred microliters of native samples was prepared, consisting of 20 mM Hepes, 100 mM NaCl, 2 mM MgCl2, 1 mM tris(2-carboxyethyl)phosphine (TCEP), and 0.1% glycerol at pH 7.4, with IspE protein (0.1 mg/ml). Cross-linking was initiated by adding 2 μl of DSBU-d0 (100 mM stock in dimethyl sulfoxide; Thermo Fisher Scientific, A35459) to a final concentration of 1 mM. Simultaneously, refolded IspE samples were cross-linked with 2 μl of DSBU-d12 [deuterated disuccinimidyl dibutyric urea (DSBU); 100 mM stock in dimethyl sulfoxide]. DSBU-d12 was prepared as described in a separate publication (55).

The samples were then mixed and incubated at room temperature for 1 hour using a Roto-Mini rotator (Alkali Scientific RS3024) at 10 rpm. The reactions were quenched by adding 1 M tris-HCl (pH 7.5) to a final concentration of 20 mM, followed by 30-min incubation at room temperature, rotating at 10 rpm on the Roto-Mini. After quenching, the native and refolded samples were combined and transferred to a new microfuge tube containing 304 mg of urea to achieve a final concentration of 8 M. Cross-linking experiments were conducted in five replicates for both native and refolded IspE samples; after mixing the pairs, a total of five replicate samples (400 μl) was prepared.

Modifying matched and nonaccessible cross-link score function

To computationally estimate the cross-linking propensity between residues 𝑖 and 𝑗, we first use the JWalk algorithm (30) to calculate the SASD between their Cα atoms. Subsequently, the cross-linking propensity is assessed using a customized version of the matched and nonaccessible cross-link (MNXL) score function (30). The original function was designed for pairs of lysines cross-linked by the dextran sulfate sodium linker (with a length of 11.4 Å); the MNXL score function needs adaptation for our XL-MS experiment where we use the DSBU cross-linker molecule (has a length of 12.5 Å) (26) and permit nonlysine residues to participate in cross-linking. To account for the differences, we have adjusted the parameters of the MNXL scoring function

MNXL score(SASDi,j)={N(SASDi,j,μi,j,σi,j2),SASDi,jthresholdi,j0,otherwise (12)

where N(SASDi,j,μi,j,σi,j2) represents the normal distribution

N(SASDi,j,μi,j,σi,j2)=12πσi,je12(SASDi,jμi,jσi,j)2 (13)

μi,j , σi,j2 , and thresholdi,j denote the mean, variance, and the SASD cutoff threshold for a specific pair of residues (i,j) involved in cross-linking, respectively. If the SASD of the pair of residues is less than the cutoff threshold, the propensity is determined by a normal distribution with the mean and variance specifically tailored for the pair of residues (table S6). If the SASD exceeds the threshold or if at least one residue in the pair is buried and inaccessible, then the propensity is assigned a value of 0.

Comparison of misfolded and native clusters for consistency between simulations and experimental data

LiP-MS and XL-MS detected 13 and 20 structural signals, respectively, exhibiting significant differences between refolded and native samples. To assess the consistency between the misfolded ensembles from our simulations and experimental data, we examined whether the SASA of a region within five residues around the PK cut site identified from the experiment and the cross-linking propensity (MNXL score, Eq. 12) of pairs of residues from misfolded ensembles are statistically significantly different from the native ensemble from our simulations. The P values were computed using numerical permutation tests with 105 iterations with the null hypothesis stating that there is no difference in the ensemble mean of SASA or cross-linking propensity between misfolded clusters and the native cluster. The P values were then adjusted using the Benjamini-Hochberg method (29). If adj.Pvalue0.05 , we reject the null hypothesis and accept the alternative hypothesis: The ensemble means of SASA or cross-linking propensity are significantly different between misfolded and native clusters. To mitigate the risk of underestimating the confidence interval of the ensemble mean (thus potentially overestimating the number of consistent signals) because of data correlation, we subsampled our data every 50 frames (as suggested by block averaging).

Calculation of the relative aggregation propensity of the IspE misfolded clusters

The aggregation propensity of the misfolded clusters of IspE relative to the native cluster is quantified by the difference in the SASA of aggregation-prone regions between these two clusters. This propensity of misfolded cluster i relative to the native state is calculated using the following equation

χi=SASAiSASAnativeSASAnative×100% (14)

The aggregation-prone region was predicted by the AmylPred2 server (31). The uncertainties were estimated from bootstrap resampling with 105 iterations.

Uncertainty estimation

The uncertainties, represented as 95% confidence intervals, were estimated using bootstrapped resampling with 105 iterations (except for the Arrhenius extrapolation, where 104 iterations were used because of computational expensive). For a sample of N elements, in each iteration, we randomly selected N values from the original data with replacement. The mean was calculated from these resampled values, and the distribution of these means was used to estimate the 95% confidence interval.

Permutation test

To evaluate whether the sample means ( xA and xB ) of two groups, A (with size nA) and B (with size nB), are significantly different, we performed a permutation test with 105 iterations. The null hypothesis states that there is no difference between the sample means of the two groups. First, we calculated the absolute observed mean difference: Tobs=xAxB . Next, we pooled all the data from both groups and randomly selected, without replacement, nA values for group A and nB values for group B to form two new resampled groups. We then computed the absolute mean difference for these resampled groups. This process was repeated for 105 iterations, generating a distribution of mean differences under the null hypothesis.

The two-tailed P value was calculated as the proportion of times the absolute resampled mean difference was equal or more extreme than the absolute observed difference ( Tobs ). For multiple test comparison, we applied Benjamini-Hochberg correction to P values. If the adj.Pvalue0.05 , we reject the null hypothesis and accept the alternative hypothesis suggesting a statistically significant difference between the sample means of the two groups.

Supporting methods for LiP-MS and XL-MS experiments

Additional methods related to LiP-MS and XL-MS experiments can be found in the Supplementary Materials.

Acknowledgments

We thank D. E. Shaw’s group for providing access to molecular dynamics trajectories for ubiquitin and λ-repressor. This research was supported in part by the TASK Supercomputer Center in Gdansk, Poland, and PLGrid Infrastructure in Poland (Grant ID: plgribliq), as well as Roar supercomputer in the US.

Funding: This work was supported by the following: National Science Centre, Poland, grant 2019/35/B/ST4/02086 (to M.S.L.), National Science Foundation grant MCB-2045844 (to S.D.F.), National Institutes of Health grant DP2-GM140926 (to S.D.F.), National Science Foundation grant MCB-1553291 (to E.P.O.), and National Institutes of Health grant R35-GM124818 (to E.P.O.).

Author contributions: Conceptualization: S.D.F. and E.P.O. Methodology: Q.V.V., I.S., Y.J., Y.X., H.S., M.S.L., S.D.F., and E.P.O. Investigation: Q.V.V., I.S., Y.J., Y.X., P.S., D.Y., M.S.L., S.D.F., and E.P.O. Resources: Q.V.V., Y.J., Y.X., P.S., M.S.L., S.D.F., and E.P.O. Data curation: Q.V.V., Y.J., and Y.X. Validation: Q.V.V., I.S., Y.J., and Y.X. Formal analysis: Q.V.V., I.S., Y.J., and Y.X. Software: Q.V.V., I.S., and Y.J. Visualization: Q.V.V., I.S., and Y.J. Funding acquisition: M.S.L., S.D.F., and E.P.O. Project administration: S.D.F. and E.P.O. Supervision: S.D.F. and E.P.O. Writing—original draft: Q.V.V. and E.P.O. Writing—review and editing: Q.V.V., I.S., Y.J., Y.X., P.S., D.Y., H.S., M.S.L., S.D.F., and E.P.O.

Competing interests: The authors declare that they have no competing interests.

Data and materials availability: All data needed to evaluate the conclusions in the paper are present in the paper and/or the Supplementary Materials. LiP-MS and XL-MS data and comparison between misfolded clusters and the native cluster from simulations based on experimental signals are provided in data S1. Misfolded structures of IspE are provided in data S2. The computer code for coarse-grained simulations and entanglement analysis is available at https://github.com/obrien-lab/cg_simtk_protein_folding. Additional analyses and data are available at https://github.com/obrien-lab/All_Atom_Misfolding_Exp_Ensemble or https://zenodo.org/records/15412020. Interactive visualizations of structures in Fig. 3B are available at https://obrien-lab.github.io/All_Atom_Misfolding_Exp_Ensemble/. Software used in experimental data analysis is available at https://github.com/FriedLabJHU/Refoldability-Tools. The mass spectrometry data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifiers PXD055795 for LiP-MS and PXD066083 for XL-MS.

Supplementary Materials

The PDF file includes:

Supplementary Text

Figs. S1 to S7

Tables S1 to S6

Legends for data S1 and S2

References

sciadv.adt8974_sm.pdf (2.2MB, pdf)

Other Supplementary Material for this manuscript includes the following:

Data S1 and S2

REFERENCES AND NOTES

  • 1.Jiang Y., Neti S. S., Sitarik I., Pradhan P., To P., Xia Y., Fried S. D., Booker S. J., O’Brien E. P., How synonymous mutations alter enzyme structure and function over long timescales. Nat. Chem. 15, 308–318 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Nissley D. A., Jiang Y., Trovato F., Sitarik I., Narayan K. B., To P., Xia Y., Fried S. D., O’Brien E. P., Universal protein misfolding intermediates can bypass the proteostasis network and remain soluble and less functional. Nat. Commun. 13, 3081 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Halder R., Nissley D. A., Sitarik I., Jiang Y., Rao Y., Vu Q. V., Li M. S., Pritchard J., O’Brien E. P., How soluble misfolded proteins bypass chaperones at the molecular level. Nat. Commun. 14, 3689 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Baiesi M., Orlandini E., Trovato A., Seno F., Linking in domain-swapped protein dimers. Sci. Rep. 6, 33872 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Baiesi M., Orlandini E., Seno F., Trovato A., Exploring the correlation between the folding rates of proteins and the entanglement of their native states. J. Phys. A. Math. Theor. 50, 504001 (2017). [Google Scholar]
  • 6.Baiesi M., Orlandini E., Seno F., Trovato A., Sequence and structural patterns detected in entangled proteins reveal the importance of co-translational folding. Sci. Rep. 9, 8426 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Rana V., Sitarik I., Petucci J., Jiang Y., Song H., O’Brien E. P., Non-covalent lasso entanglements in folded proteins: Prevalence, functional implications, and evolutionary significance. J. Mol. Biol. 436, 168459 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Komar A. A., Lesnik T., Reiss C., Synonymous codon substitutions affect ribosome traffic and protein folding during in vitro translation. FEBS Lett. 462, 387–391 (1999). [DOI] [PubMed] [Google Scholar]
  • 9.Zhou M., Guo J., Cha J., Chae M., Chen S., Barral J. M., Sachs M. S., Liu Y., Non-optimal codon usage affects expression, structure and function of clock protein FRQ. Nature 494, 111–115 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Zhou M., Wang T., Fu J., Xiao G., Liu Y., Nonoptimal codon usage influences protein structure in intrinsically disordered regions. Mol. Microbiol. 97, 974–987 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Nissley D. A., Vu Q. V., Trovato F., Ahmed N., Jiang Y., Li M. S., O’Brien E. P., Electrostatic interactions govern extreme nascent protein ejection times from ribosomes and can delay ribosome recycling. J. Am. Chem. Soc. 142, 6103–6110 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Piana S., Lindorff-Larsen K., Shaw D. E., Atomic-level description of ubiquitin folding. Proc. Natl. Acad. Sci. U.S.A. 110, 5915–5920 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Lindorff-Larsen K., Piana S., Dror R. O., Shaw D. E., How fast-folding proteins fold. Science 334, 517–520 (2011). [DOI] [PubMed] [Google Scholar]
  • 14.Sivaraman T., Arrington C. B., Robertson A. D., Kinetics of unfolding and folding from amide hydrogen exchange in native ubiquitin. Nat. Struct. Biol. 8, 331–333 (2001). [DOI] [PubMed] [Google Scholar]
  • 15.Yang W. Y., Gruebele M., Folding at the speed limit. Nature 423, 193–197 (2003). [DOI] [PubMed] [Google Scholar]
  • 16.Frishman D., Argos P., Knowledge-based protein secondary structure assignment. Proteins Struc. Func. Genet. 23, 566–579 (1995). [DOI] [PubMed] [Google Scholar]
  • 17.Maxwell K. L., Wildes D., Zarrine-Afsar A., De Los Rios M. A., Brown A. G., Friel C. T., Hedberg L., Horng J.-C., Bona D., Miller E. J., Vallée-Bélisle A., Main E. R. G., Bemporad F., Qiu L., Teilum K., Vu N.-D., Edwards A. M., Ruczinski I., Poulsen F. M., Kragelund B. B., Michnick S. W., Chiti F., Bai Y., Hagen S. J., Serrano L., Oliveberg M., Raleigh D. P., Wittung-Stafshede P., Radford S. E., Jackson S. E., Sosnick T. R., Marqusee S., Davidson A. R., Plaxco K. W., Protein folding: Defining a “standard” set of experimental conditions and a preliminary kinetic data set of two-state proteins. Protein Sci. 14, 602–616 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Braselmann E., Chaney J. L., Clark P. L., Folding the proteome. Trends Biochem. Sci. 38, 337–344 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Brocchieri L., Karlin S., Protein length in eukaryotic and prokaryotic proteomes. Nucleic Acids Res. 33, 3390–3400 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sarkar D., Kang P., Nielsen S. O., Qin Z., Non-arrhenius reaction-diffusion kinetics for protein inactivation over a large temperature range. ACS Nano 13, 8669–8679 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Feng Y., De Franceschi G., Kahraman A., Soste M., Melnik A., Boersema P. J., de Laureto P. P., Nikolaev Y., Oliveira A. P., Picotti P., Global analysis of protein structural changes in complex proteomes. Nat. Biotechnol. 32, 1036–1044 (2014). [DOI] [PubMed] [Google Scholar]
  • 22.P. To, Whitehead B., Tarbox H. E., Fried S. D., Nonrefoldability is pervasive across the E. coli proteome. J. Am. Chem. Soc. 143, 11435–11448 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Schopper S., Kahraman A., Leuenberger P., Feng Y., Piazza I., Müller O., Boersema P. J., Picotti P., Measuring protein structural changes on a proteome-wide scale using limited proteolysis-coupled mass spectrometry. Nat. Protoc. 12, 2391–2410 (2017). [DOI] [PubMed] [Google Scholar]
  • 24.Malinovska L., Cappelletti V., Kohler D., Piazza I., Tsai T.-H., Pepelnjak M., Stalder P., Dörig C., Sesterhenn F., Elsässer F., Kralickova L., Beaton N., Reiter L., de Souza N., Vitek O., Picotti P., Proteome-wide structural changes measured with limited proteolysis-mass spectrometry: An advanced protocol for high-throughput applications. Nat. Protoc. 18, 659–682 (2023). [DOI] [PubMed] [Google Scholar]
  • 25.O’Reilly F. J., Rappsilber J., Cross-linking mass spectrometry: Methods and applications in structural, molecular and systems biology. Nat. Struct. Mol. Biol. 25, 1000–1008 (2018). [DOI] [PubMed] [Google Scholar]
  • 26.Pan D., Brockmeyer A., Mueller F., Musacchio A., Bange T., Simplified protocol for cross-linking mass spectrometry using the MS-cleavable cross-linker DSBU with efficient cross-link identification. Anal. Chem. 90, 10990–10999 (2018). [DOI] [PubMed] [Google Scholar]
  • 27.Herzog F., Kahraman A., Boehringer D., Mak R., Bracher A., Walzthoeni T., Leitner A., Beck M., Hartl F.-U., Ban N., Malmström L., Aebersold R., Structural probing of a protein phosphatase 2A network by chemical cross-linking and mass spectrometry. Science 337, 1348–1352 (2012). [DOI] [PubMed] [Google Scholar]
  • 28.To P., Xia Y., Lee S. O., Devlin T., Fleming K. G., Fried S. D., A proteome-wide map of chaperone-assisted protein refolding in a cytosol-like milieu. Proc. Natl. Acad. Sci. U.S.A. 119, e2210536119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Benjamini Y., Hochberg Y., Controlling the false discovery rate: A practical and powerful approach to multiple testing. J. R. Stat. Soc. Series B Stat. Methodol. 57, 289–300 (1995). [Google Scholar]
  • 30.Bullock J. M. A., Schwab J., Thalassinos K., Topf M., The importance of non-accessible crosslinks and solvent accessible surface distance in modeling proteins with restraints from crosslinking mass spectrometry. Mol. Cell. Proteomics 15, 2491–2500 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Tsolis A. C., Papandreou N. C., Iconomidou V. A., Hamodrakas S. J., A consensus method for the prediction of “Aggregation-Prone” peptides in globular proteins. PLOS ONE 8, e54175 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Sułkowska J. I., Noel J. K., Onuchic J. N., Energy landscape of knotted protein folding. Proc. Natl. Acad. Sci. U.S.A. 109, 17783–17788 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Salicari L., Baiesi M., Orlandini E., Trovato A., Folding kinetics of an entangled protein. PLOS Comput. Biol. 19, e1011107 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Guo Z., Thirumalai D., Kinetics of protein folding: Nucleation mechanism, time scales, and pathways. Biopolymers 36, 83–102 (1995). [Google Scholar]
  • 35.Thirumalai D., Woodson S. A., Kinetics of folding of proteins and RNA. Acc. Chem. Res. 29, 433–439 (1996). [Google Scholar]
  • 36.Thirumalai D., Klimov D. K., Woodson S. A., Kinetic partitioning mechanism as a unifying theme in the folding of biomolecules. Theor. Chem. Acta 96, 14–22 (1997). [Google Scholar]
  • 37.Jorgensen W. L., Chandrasekhar J., Madura J. D., Impey R. W., Klein M. L., Comparison of simple potential functions for simulating liquid water. J. Chem. Phys. 79, 926–935 (1983). [Google Scholar]
  • 38.Bryson A. E., Denham W. F., A steepest-ascent method for solving optimum programming problems. J. Appl. Mech. 29, 247–257 (1962). [Google Scholar]
  • 39.Haug E. J., Arora J. S., Matsui K., A steepest-descent method for optimization of mechanical systems. J. Optim. Theory Appl. 19, 401–424 (1976). [Google Scholar]
  • 40.Darden T., York D., Pedersen L., Particle mesh Ewald: An N·log(N) method for Ewald sums in large systems. J. Chem. Phys. 98, 10089–10092 (1993). [Google Scholar]
  • 41.Nosé S., A unified formulation of the constant temperature molecular dynamics methods. J. Chem. Phys. 81, 511–519 (1984). [Google Scholar]
  • 42.Hoover W. G., Canonical dynamics: Equilibrium phase-space distributions. Phys. Rev. A. (Coll Park) 31, 1695–1697 (1985). [DOI] [PubMed] [Google Scholar]
  • 43.Parrinello M., Rahman A., Polymorphic transitions in single crystals: A new molecular dynamics method. J. Appl. Phys. 52, 7182–7190 (1981). [Google Scholar]
  • 44.Hess B., Bekker H., Berendsen H. J. C., Fraaije J. G. E. M., LINCS: A linear constraint solver for molecular simulations. J. Comput. Chem. 18, 1463–1472 (1997). [Google Scholar]
  • 45.Abraham M. J., Murtola T., Schulz R., Páll S., Smith J. C., Hess B., Lindah E., Gromacs: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 1–2, 19–25 (2015). [Google Scholar]
  • 46.Huang J., Rauscher S., Nawrocki G., Ran T., Feig M., De Groot B. L., Grubmüller H., MacKerell A. D., CHARMM36m: An improved force field for folded and intrinsically disordered proteins. Nat. Methods 14, 71–73 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Niemyska W., Millett K. C., Sulkowska J. I., GLN: A method to reveal unique properties of lasso type topology in proteins. Sci. Rep. 10, 15186 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Gutierres M. B. B., Bonorino C. B. C., Rigo M. M., ChaperISM: Improved chaperone binding prediction using position-independent scoring matrices. Bioinformatics 36, 735–741 (2020). [DOI] [PubMed] [Google Scholar]
  • 49.Eastman P., Swails J., Chodera J. D., McGibbon R. T., Zhao Y., Beauchamp K. A., Wang L. P., Simmonett A. C., Harrigan M. P., Stern C. D., Wiewiora R. P., Brooks B. R., Pande V. S., OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Comput. Biol. 13, e1005659 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.D. Arthur, S. Vassilvitskii, “K-means++: The advantages of careful seeding,” in Proceedings of the Annual ACM-SIAM Symposium on Discrete Algorithms (Society for Industrial and Applied Mathematics, 2007), pp. 1027–1035.
  • 51.Röblitz S., Weber M., Fuzzy spectral clustering by PCCA+: Application to Markov state models and data classification. Adv. Data Anal. Classif. 7, 147–179 (2013). [Google Scholar]
  • 52.Prinz J.-H., Wu H., Sarich M., Keller B., Senne M., Held M., Chodera J. D., Schütte C., Noé F., Markov models of molecular kinetics: Generation and validation. J. Chem. Phys. 134, 174105 (2011). [DOI] [PubMed] [Google Scholar]
  • 53.Scherer M. K., Trendelkamp-Schroer B., Paul F., Pérez-Hernández G., Hoffmann M., Plattner N., Wehmeyer C., Prinz J. H., Noé F., PyEMMA 2: A software package for estimation, validation, and analysis of markov models. J. Chem. Theory Comput. 11, 5525–5542 (2015). [DOI] [PubMed] [Google Scholar]
  • 54.Onuchic J. N., Luthey-Schulten Z., Wolynes P. G., THEORY OF PROTEIN FOLDING: The energy landscape perspective. Annu. Rev. Phys. Chem. 48, 545–600 (1997). [DOI] [PubMed] [Google Scholar]
  • 55.Jiang Y., Xia Y., Sitarik I., Sharma P., Song H., Fried S. D., O’Brien E. P., Protein misfolding involving entanglements providesa structural explanation for the origin of stretched-exponential refolding kinetics. Sci. Adv. 11, eads7379 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Gosavi S., Chavez L. L., Jennings P. A., Onuchic J. N., Topological frustration and the folding of interleukin-1β. J. Mol. Biol. 357, 986–996 (2006). [DOI] [PubMed] [Google Scholar]
  • 57.Sherman M. Y., Qian S. B., Less is more: Improving proteostasis by translation slow down. Trends Biochem. Sci. 38, 585–591 (2013). [DOI] [PubMed] [Google Scholar]
  • 58.Kiefhaber T., Kohler H. H., Schmid F. X., Kinetic coupling between protein folding and prolyl isomerization. I. Theoretical models. J. Mol. Biol. 224, 217–229 (1992). [DOI] [PubMed] [Google Scholar]
  • 59.E. A. Kikis, “The intrinsic and extrinsic factors that contribute to proteostasis decline and pathological protein misfolding” in Advances in Protein Chemistry and Structural Biology (Elsevier Ltd., 2019), vol. 118. [DOI] [PubMed]
  • 60.Salahuddin P., Siddiqi M. K., Khan S., Abdelhameed A. S., Khan R. H., Mechanisms of protein misfolding: Novel therapeutic approaches to protein-misfolding diseases. J. Mol. Struct. 1123, 311–326 (2016). [Google Scholar]
  • 61.Lafita A., Tian P., Best R. B., Bateman A., Tandem domain swapping: Determinants of multidomain protein misfolding. Curr. Opin. Struct. Biol. 58, 97–104 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Yu F., Haynes S. E., Teo G. C., Avtonomov D. M., Polasky D. A., Nesvizhskii A. I., Fast quantitative analysis of timsTOF PASEF data with MSFragger and IonQuant. Mol. Cell. Proteomics 19, 1575–1585 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Manriquez-Sandoval E., Brewer J., Lule G., Lopez S., Fried S. D., FLiPPR: A processor for limited proteolysis (LiP) mass spectrometry data sets built on FragPipe. J. Proteome Res. 23, 2332–2342 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Mendes M. L., Fischer L., Chen Z. A., Barbon M., O’Reilly F. J., Giese S. H., Bohlke-Schneider M., Belsom A., Dau T., Combe C. W., Graham M., Eisele M. R., Baumeister W., Speck C., Rappsilber J., An integrated workflow for crosslinking mass spectrometry. Mol. Syst. Biol. 15, e8994 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Text

Figs. S1 to S7

Tables S1 to S6

Legends for data S1 and S2

References

sciadv.adt8974_sm.pdf (2.2MB, pdf)

Data S1 and S2


Articles from Science Advances are provided here courtesy of American Association for the Advancement of Science

RESOURCES