Abstract
AlphaFold has revolutionized protein structure prediction by accurately creating 3D structures from just the amino acid sequence. However, even with extensive research validating its overall accuracy, a key question remains: Can AlphaFold predict the conformation of individual amino acid residue side chains within a folded protein? This is important for the field of molecular modeling, particularly when predicting the effects of mutations on protein stability and ligand binding. AlphaFold generates a set of atomic coordinates not just for the mutated side chain but also for potential rearrangements across the entire protein structure. In this study we investigate the ability of ColabFold, an online implementation of AlphaFold2 (AF2), to predict the conformations of residue side chains in folded proteins. We find that over a set of 10 benchmark proteins, the side chain conformation prediction error of ColabFold is on average for dihedral angles, and increases to for dihedral angles. The prediction error is smaller for non-polar side chains and is somewhat improved using structural templates. ColabFold demonstrates a bias towards the most prevalent rotamer states in the protein data bank (PDB), potentially limiting its ability to capture rare side chain conformations effectively. As an application of AlphaFold to explore the structural consequences of strongly cooperative mutations on side chain rearrangements, we employ a Potts sequence-based statistical energy model to perform large scale mutational scans of two proteins ABL1 and PIM1 kinase, searching for the most strongly cooperative mutational pairs, and then use ColabFold to predict the structural signatures of this cooperativity on the interacting side chains. Our results demonstrate that integration of the sequence-based Potts model with AlphaFold into a single pipeline provides a new tool that can be used to explore the fundamental relationship between protein mutations, cooperative changes in structure, and fitness.
Introduction
To perform their functions in living organisms, most proteins must fold into a unique three-dimensional structure, or more generally, into ensembles of structures which constitute distinct functional states. These ensembles taken together characterize the protein conformational landscape. Starting from the famous experiments of Anfinsen et al. (1), not only is the question of how proteins reach their biologically active ensembles of conformations not fully answered, but also determining the three-dimensional folded structures of proteins remains one of the most challenging problems. Several experimental techniques, such as nuclear magnetic resonance (NMR) (2), X-ray crystallography (3), and cryo-electron microscopy (cryo-EM) (4), can be used to determine protein structures; however, the structures of only ( structures) of the entire protein universe have been determined. To this end, AlphaFold2 (AF2) (5), a machine-learning-based model to predict protein structures developed by DeepMind, has been groundbreaking for predicting protein structures and interactions (6). It took DeepMind only one and a half years to release structures of more than 200 million proteins predicted by AF2, i.e., structures of almost all the known proteins on the planet. Notably the 2024 Nobel Prize in Chemistry was awarded to two of the developers of AlphaFold for this achievement (https://www.nobelprize.org/prizes/chemistry/2024/summary/).
While the AF2 algorithm represents a significant advancement in protein structure prediction, one of the unresolved questions is: how accurate are these predicted structures for various computational modeling applications? Free Energy Perturbation (FEP) simulations are used to estimate the thermodynamic effects of mutations on protein stability by estimating the folding free energy difference between wild-type and mutant structures (7–13). The accuracy of these calculations depends on the accurate representation of the three-dimensional structure of the residue side chains in the folded protein (14–17). While current computational tools for predicting side chain conformations are widely used, they have not yet incorporated machine learning methods. AF2 offers an approach for side chain conformational (rotamer state) predictions by generating detailed atomic coordinates for the entire protein which captures both the local backbone and side chain structural changes at the mutation site, and potentially larger rearrangements in the protein structure.
First, it should be noted that the structure predictions by AF2 are much more accurate than by other prediction algorithms, as demonstrated at the 14th Critical Assessment of Structure Prediction (CASP14) competition (18). Second, it has recently been reported that the prediction accuracy by AF2 is often within the range of the accuracy of experimentally-derived structures of the same protein (19,20). In particular, using the root mean square deviation of atoms, as a measure of the accuracy, the authors found that the average root mean square deviation of atoms between different experimental structures (e.g. crystallized in a different space group) of the same protein is , whereas it is between experimental structures and the structures predicted by the AF2 (19). Moreover, it was shown that the AF2 models for nine small, monomeric NMR protein structures, not used in the training of AF2, have accuracies that are comparable to the experimental NMR models deposited in the PDB (20). Despite this relatively high accuracy, it was found that sometimes predictions are incompatible with experimental structure. For example, using another measure of the accuracy of the AF2 predictions, the local distance difference test (LDDT) scores, which evaluate local distance differences of all atoms in a model by comparing the inter-atomic distances of one structure to the same distances in the other structure (21) and correlate well with the confidence level of the prediction, which is quantified as the predicted-LDDT (or pLDDT), it was noted (19) that of the side chains in moderate (for 70 < pLDDT < 90) to high (for pLDDT > 90) confidence residues of predicted structures had different conformations than in the experimental structure, and about 1/3 of these (7%) were incompatible with the experimental data. These findings are in general agreement with a previous study by DeepMind (22), in which 7% (for pLDDT > 90) to 30% (for 70 < pLDDT < 90) of side chains had a angle deviation of at least 40° from experimental structures. It was pointed out (19) that the lack of explicitly accounting for ligands, ions, covalent modifications or environmental conditions in AF2 predictions might be the reasons for some of these discrepancies.
In the present study, we performed a detailed analysis of side chain conformations, including all angles, for ten different proteins predicted by ColabFold (23) to determine (i) for which amino acid residues are side chain conformations most difficult to predict correctly; (ii) whether the addition of structural templates improves the accuracy of the prediction; (iii) if there are some other reasons in addition to those mentioned above that may cause these discrepancies. Here, we consider a side chain prediction is correct if it is within +/− 40° from the experimental value (17).
Methods
ColabFold (23), a fast and user-friendly implementation of AF2, was employed to predict the structures of ten different proteins, including, six proteins [staph nuclease (PDB ID: 1EY0 (24)), T4 lysozymes (PDB IDs: 2LZM (25) and 1L63 (26)), ABL-1 kinase (PDB ID: 2V7A (27)), PIM-1 kinase (PDB ID: 1XWS (28)), and antifreeze protein (PDB ID: 3VN3 (29))], two sheet proteins [choristoneura fumiferana (spruce budworm) antifreeze protein (PDB ID: 1L0S (30)) and unknown function protein (BF3112) from bacteroides fragilis NCTC 9343 (PDB ID: 3MSW (31))], and two helical proteins [myoglobin (PDB ID: 3V2V (32)) and the transmembrane protein TMHC2_E (PDB ID: 6B87 (33))]. By replacing AF2’s homology search with the 40 to 60-fold faster, Many-against-Many sequence searching (MMseqs2 (34,35)), ColabFold significantly accelerates predictions of protein structures and complexes. The structure predictions on ColabFold were performed with- and without structural templates. The proteins selected in this study are from diverse families, and differ from each other by size and biological function. Predicted structures have been validated by comparing them to experimental PDB structures.
ColabFold allows users to customize the protein structure prediction. There are several important parameters, including particularly the depth of the multiple sequence alignment (MSA), and the use of structural templates, that can be changed. Supplying a structural template, pushes ColabFold prediction to resemble the provided structure, although the extent to which the prediction is biased by the template and the accuracy of the prediction depends on the MSA depth, as well. In this study, we used the target proteins themselves as the structural templates, which enabled us to investigate the contributions from the structural template and the MSA in the prediction accuracy.
Results and Discussion
Application of structural templates for side-chain structure prediction and performance by residue type
We have performed structure predictions by ColabFold with and without structural templates for ten different proteins to elucidate the prediction accuracy of side chain conformations in the folded protein: (a) when only the MSA and the sequence of the target protein are provided; and (b) to determine how the prediction accuracy changes with the addition of a structural template.
The results illustrated in Figure 1 show that the prediction accuracy for side chain conformations with and without templates, on average, is highest for angles and decreases with the increase of index (the average errors for angles are 12% and 17% with and without templates, respectively, and increase to 47% and 50% with and without templates on average, respectively, for angles. The angles are an exception because of their existence only in two amino acids – Arginine and Lysine). Moreover, the prediction accuracy of depends on the depth of the MSA and the secondary structure. ColabFold tends to predict side chain dihedral angles for proteins with structures more accurately than ones for proteins with -helical or -strand only structures.
Figure 1.
The percentages of incorrectly predicted (panel A), (panel B), (panel C), and (panel D) angles, in predicted structures with (blue bars) and without (orange bars) templates, determined by the formula: Error , where and is number of predicted angles, values of which differ from their experimental counterparts by more than corresponds to the number of all angles.
The utilization of structural templates in protein structure prediction by ColabFold results in an enhancement in the accuracy of the predicted rotamer states of side chains in the ten benchmark proteins ( averaged over all side chain torsions) (Figure 1). The addition of templates has a significant effect on the conformations (improvement in the accuracy of on average using templates) but very little effect on the accuracy of predictions. There is one protein (2LZM) among ten studied proteins, for which the and angles exhibit an opposite trend; the side chain prediction accuracy is better without the structural template (Figure 1). The accuracy of the ColabFold structure prediction strongly depends on the MSA. A diverse and deep MSA, with thousands of sequences in the alignment, helps ColabFold to identify coevolutionary signals, and when these coevolutionary signals are strong ColabFold tends to ignore template structures. 2LZM has a deep MSA (> 3,500 sequences), and might be one of the reasons for the inconsistency. Secondly, the depth of the MSA and the inclusion of a structural template mainly affect the accuracy of the backbone conformations of predicted structures, whereas the accuracy of side chain conformations of some amino acids does not always rely on these parameters. The factors playing a role in inaccurate prediction of side chain conformations are investigated below in this section and following sections.
Our results in Figure 1 illustrate that incorporating a structural template has a significantly greater effect improving the accuracy of proteins assembled from sheets (e.g., 1.8 times in 1L0S, and 2.2 times in 3MSW), which suggests a potential advantage of template-based approaches for -sheet prediction. -sheets, characterized by tightly packed, antiparallel strands connected by hydrogen bonds, can exhibit significant structural variability, which is usually not the case for -helical proteins. Moreover, shallow MSAs are another reason for the noticeable improvement when template structures are used; in particular, 1L0S and 3MSW have only and sequences, respectively, in the alignments.
In Figure 2, we decompose the overall performance of ColabFold for predicting side chain dihedral angles in folded proteins by residue type. The percentages of incorrectly predicted angles for the most polar amino acids are higher than for non-polar amino acids in both with and without template predicted structures. It should be noted that ColabFold with and without templates, predicts correctly all angles for two hydrophobic aromatic amino acids tryptophan and tyrosine (the total number of TYR + TRP residues in the ten target proteins is 93). Furthermore, Figure 2 reveals a general trend of lower prediction errors for angles compared to and (as in Figure 1, angles are an exception due to their existence only in two amino acids) across all residue types. This is related to the inherent greater flexibility of angles closer to the side chain terminus compared to the more constrained which are subject to restriction due to steric hindrance between the side-chain atoms and the main chain. In all cases the AF2 predictions with templates have lower error than without templates. Moreover, Figure 2 highlights a significant improvement ( error reduction) in the prediction of the dihedral angles for Cysteine (Cys) residues upon template addition. This observation suggests that AF2 might benefit from modifications of the objective function specifically focused on accurately predicting the rotamer state for Cys residues (36). It is important to mention that only 21 cysteine residues are observed in the ten studied proteins (five of them do not have cysteine residues at all). All errors in dihedral angles are found in 1L0S, which contains the most cysteines (eight residues). The structure of 1L0S reveals a left-handed -helical structure consisting of 15 amino acid loops. Each helix loop has three sides ( strands) and Ser/Thr/Cys stacking within the interior of the protein at the corners of the helix strengthened by four interstrand disulfide bonds formed by cysteines (30). All four interstrand disulfide bonds and eight dihedral angles of cysteines are correctly predicted in ColabFold predicted structure with template. The ColabFold predicted structure without template reveals only two disulfide bonds formed by cysteine residue pairs, although these pairs are not the same as in the 1L0S PDB structure and ColabFold predicted structure with template; plus, four dihedral angles of cysteines are incorrectly predicted. This finding indicates that not only side chains but also backbone of 1L0S is predicted poorly by ColabFold without template (see sub-section “Correlation between pLDDT, backbone and side-chain angles” below). Notable improvement was also observed in the prediction of angles for Methionine (Met).
Figure 2.
The same as in Fig. 1 but for amino acids.
AlphaFold2 bias towards dominant rotamer states: Implications for rare conformations
Our analysis of AF2’s rotamer state predictions reveals a potential bias towards the most frequently observed rotamer conformations in the PDB (dominant states). This bias suggests the difficulties that AF2 may have predicting rare rotamer states, even when these rare conformations are present in the experimental crystal structure provided as a template. By “rare,” we refer to conformations observed in a very small percentage (< 1%) of protein data bank entries, according to the Schrödinger knowledge base (Schrödinger suite release 2023–2) that contains the rotamer state populations for each amino acid based on an analysis of the PDB(37,38). More precisely, out of a total 1453 side chain rotamer states, predicted in this work, 8.2% of the experimental side chain conformations are rare. For visual illustration, we superposed experimental and AF2 predicted structures of staph nuclease (1EY0) (Figure 3A) and T4 lysozyme (2LZM) (Figure 3B), and selected two residues (LYS-49 in 1EY0 and LYS-16 in 2LZM) with rare and PDB dominant rotamer states in experimental and predicted structures, respectively. As shown in Figure 3, side chains of both residues have enough room to adopt various conformations, however, there are some differences. In particular, in 1EY0, LYS-49 is in a loop in both the experimental and AF2 predicted structures, which have a large deviation from each other [backbone root mean square distance (RMSD) at residue 49 is ], that permits the formation of a salt bridge between the side chains of LYS-49 (orange) and GLU-52 in the AF2 predicted structure, not observed in the PDB structure. In 2LZM, LYS-16 is embedded in different secondary structures (-sheet in experimental structure and loop in AF2 predicted structure), however, these secondary structures are only slightly deviated from each other (backbone RMSD at residue 16 is ), and hence, probably do not play a role in the discrepancy between the predicted and experimental rotamer states of LYS-16. It should be noted that and of errors for angles and for rotamer states, respectively, in all ten studied proteins come from the loop regions. Also, rotamer states of only of loop regions are incorrectly predicted by AF2.
Figure 3.
Superposed experimental and AF2 predicted structures along with RMSDs of atoms between these two structures for staph nuclease (panel A) and T4 lysozyme (panel B).
Using the Schrödinger database we have assigned the closest rotamer state in the PDB-derived database for each residue over the entire sequence for the AF2 predicted and the experimentally determined protein structure and plotted those values in the X and Y axis of Figure 4, respectively. Because there are a large number of points in Figure 4 which sometimes overlap and also differences in the distribution of rotamer state populations for amino acids in the Schrödinger knowledge base (Table S1 shows the percentages of five most dominant rotamer states of each amino acid), in Supporting Material we additionally (i) plotted the results separately for each amino acid (Figure S1); (ii) calculated the ratios between the PDB derived database percentages for each residue of ten experimental and AF2 predicted structures (Figure S2); (iii) computed probability distribution functions (PDFs) of rotamer states of each amino acid obtained from ten experimental and AF2 predicted structures along the populations of rotamer states in protein data base (Figure S3); and (iv) computed how frequently each amino acid in experimental and AF2 predicted structures adopt more dominant rotamer state from protein data base (Figure S4).
Figure 4.
Comparison of rotamer states of each amino acid [non-polar (triangles), polar uncharged (circles), polar charged (squares)] of experimental and AF2 predicted (with template) structures in terms of the probability (in %) of the closest rotamer state in the protein database.
The results show that rotamer states of most of the amino acids, except for Glu, Phe, and Tyr, in the AF2 predicted structures tend to adopt the dominant conformations from the protein data base. That is, the majority of triangles, circles, and squares in Figure 4 are located below the diagonal indicating that AF2 frequently predicts the most populated states whereas rotamer states of these amino acids obtained from experimental structures of the ten studied proteins are either in less populated states or rarely observed in the PDB database. It is not surprising that the most easily visible residues, in Figure 4, showing a bias towards the most frequently observed conformations in protein databases in AF2 predicted structures are Leu, Ile and Val, which have much higher percentages of dominant rotamer states than others (see Table S1). Further, as illustrated in Figure S4, (i) the rotamer states of polar amino acids Arg, Asn, Asp, Gln, and Lys in AF2 predicted structures exhibit a strong inclination towards more dominant rotamer states in protein data base than their counterparts in experimental structures; (ii) Glu, Phe and Tyr show opposite results, i.e., rotamer states of these amino acids in experimental structures exhibit a strong tendency to adopt more dominant rotamer states in protein data base than their counterparts in AF2 predicted structures; (iii) the rotamer states of Pro and Trp in experimental and AF2 predicted structures exhibit an equal propensity towards the dominant rotamer states in protein data base; (iv) the rotamer states of rest of the amino acids in AF2 predicted structures are more inclined towards the dominant rotamer states in protein data base than their counterparts in experimental structures, although in majority cases rotamer states of these amino acids in AF2 and experimental structures do not differ from each other and fall into the same category defined by protein data base. These results suggest training bias in the AF2 model that corresponds to a potential limitation in AF2’s ability to handle low population rotamer states when predicting the protein structure.
It is well known that “rare” rotamer states might represent local minima in the protein’s energy landscape (16). Hence, without MD refinement that accounts for various interactions, such as hydrogen bonding, steric clashes, and solvation effects, capturing these subtle energy minima can be difficult for the AF2 algorithm (39). This also suggests that AF2’s internal sampling procedures might not be sufficiently comprehensive to explore the full conformational space, especially for regions with high flexibility or the presence of rare rotamers.
Spatial clustering of rotamer prediction errors adjacent to Alanine and Glycine residues
One of the most intriguing observations from our analysis of multiple ColabFold-predicted PDB structures is the spatial clustering of rotamer prediction errors. We observed a frequent trend where prediction errors tend to occur next to Alanine (Ala) or Glycine (Gly) residues in the protein sequence. In particular, of all errors in the range of 40° - 100°, found in predicted structures for angles occur in residues next to Ala or Gly. Moreover, the error rate increases to for errors .
One possible explanation lies in backbone flexibility. Both Ala and Gly lack bulky side chains, leading to a more flexible backbone compared to other residues. This inherent flexibility can introduce conformational heterogeneity, particularly in neighboring peptide planes.
To explore this problem, we computed root mean square distances between the residues of experimental and AF2 predicted structures for ten studied proteins, and found that of the side chain prediction errors adjacent to Gly and Ala residues might be associated with a lack of packing constraints due to the small size of Ala and Gly, and the remainder () might be associated with backbone prediction errors, respectively. For visual illustration of both mechanisms, we selected one of the proteins (2LZM), which exhibits large errors (> 40°) in rotamer states at eight residues next to Ala and Gly. Two curves are depicted in Figure 5: the blue curve corresponds to errors between angles obtained from experimental and AF2 predicted structures, and the orange curve corresponds to backbone RMSDs. Because Ala and Gly do not have angles, the blue curve looks like a piecewise function, in which the “breaks” correspond to Ala and Gly residues. Based on the behavior of the RMSD over the sequence, the threshold at was determined indicating which mechanism may cause a large error in rotamer state. In other words, if the backbone RMSD , a large discrepancy in rotamer state might be caused by backbone prediction errors, and if , a large error in rotamer state might be induced by lack of packing constraints. All eight residues are numbered with red () and black colors depending which mechanism may cause the large discrepancies in angles. Also, two inserts are included in Figure 5 illustrating the structural differences caused by these two mechanisms. The insert for Lys-48 shows the discrepancy between the backbones of experimental and AF2 predicted structures, which may consequently originate the difference in rotamer states, however, due to a large free space in the neighborhood of Lys-48 the tendency of AF2 to adopt the more dominant rotamer states in protein data base could easily cause this discrepancy in angles. The rotamer state of Lys-48 in the experimental structure has a frequency of 0.9% in the PDB database, while the AF2-predicted rotamer state has a frequency of 5.6%. The insert for Lys-147 demonstrates that there is a large free space in the vicinity of Lys-147 allowing side chains to adopt different conformations. While the small size of Ala may be the reason for the discrepancy in angles in Lys-147, it is possible its side chain could adopt different conformations with having other neighbor residues, as well. The main reason for the discrepancy may again be due to AF2’s bias for dominant rotamer states, because the rotamer state of Lys-147 in the experimental structure has frequency very close to zero (very rare) in the protein data base, while it is 5.6% in AF 2 predicted structure.
Figure 5.
Absolute differences between values (blue curve) and RMSDs (orange curve) of experimental and AF2 predicted structures with template for 2LZM. Red color numbers indicate large errors in rotamer states that might be caused by backbone prediction errors, black color numbers indicate large errors in rotamer states that might be induced by lack of packing constraints. Inserts with experimental (gold) and AF2 predicted (cyan) structures illustrate the differences between side chain conformations, in particular, for Lys-48 and Lys-147, caused by these two mechanisms. Purple dashed line represents the threshold for angles (40°) and RMSD ().
To understand whether discrepancies between backbones of experimental and AF2 predicted structures are correlated with errors of side chain rotamer states, we computed (i) the Pearson correlation coefficients between the RMSD and values; (ii) average RMSD values for all Ala and Gly residues, separately, and for entire sequences without Ala and Gly residues; and also, for all residues with large rotamer state errors (see Table S2). The results show that (i) there is either very weak or no correlation between RMSD and values; (ii) average RMSD values are changing (“oscillating”) from protein to protein. However, as was expected, average RMSD values for residues with large rotamer state errors are greater than average RMSD values for Ala, Gly and for rest of sequences in most proteins. It should be noted that large discrepancies between backbones of experimental and AF2 predicted structures are mainly found in loops and ends (only two proteins, 3 V 2 V and 6B87, exhibit opposite results) (see Table S2). Therefore, if loops contain Ala and Gly residues, average backbone error at Ala and Gly residues may increase, which is the reason of significantly large average RMSD values for Ala and Gly in proteins 3MSW and 2V7A, respectively.
Based on these results, the driving force causing the discrepancies between rotamer states of residues of experimental and AF2 predicted structures is the inclination of AF2 to adopt more dominant rotamer states from protein data base. The lack of packing constraints caused by a small size of Ala and Gly or increased backbone flexibility can play a role in the discrepancies between rotamer states, although it may not be as significant as the former.
Correlation between pLDDT, backbone and side-chain dihedral angles
As stated in the Introduction, the accuracy of the AF2 predictions is evaluated by the LDDT scores that correlate well with the confidence level of the prediction (pLDDT). It has been shown that a pLDDT score greater than 90 is taken as the benchmark for very high confidence and a pLDDT score greater than 70 is considered to be a moderate-to-high confidence structure and corresponds to a correct backbone prediction, and a pLDDT < 50 indicates unreliable predictions (19). Moreover, it was observed that AF2 can predict side-chains ( angles) with high accuracy when the backbone prediction is high confidence (pLDDT > 90) (22). It is of interest to know how the pLDDT scores are correlated with the prediction accuracy of backbone and side-chain angles. Among the ten studied proteins we selected three proteins which illustrated different behavior. In particular, (i) the predicted structures of staph nuclease (PDB ID: 1EY0) with and without template have very high pLDDT (> 90) along the entire sequence except for one loop region (43–53, 62 < pLDDT < 85) (Fig. S5A); (ii) the predicted structures of T4 lysozyme (PDB ID: 1L63) with and without template exhibit very high (>90) pLDDT score along almost entire sequence (Fig. S5B); (iii) the predicted structure of choristoneura fumiferana (spruce budworm) antifreeze protein (PDB ID: 1L0S) with template shows very high (>90) pLDDT score along the entire sequence, whereas the predicted structure without template illustrates moderate and below moderate confidence (Fig. S5C). Figure S5 shows moderate negative (for 1EY0 and 1L0S) and weak negative (for 1L63) correlations between pLDDT and backbone angles and very weak negative correlation between pLDDT and angles for all three proteins. There is no correlation observed between pLDDT and angles, consequently no correlation between and angles.
Prediction accuracy of buried vs solvent exposed side chains
To elucidate how the solvent accessibility of each residue changes in the AF2 predicted structures compared with the experimental PDB structures, we computed the solvent accessible surface areas (SASAs), which is considered to be an important factor in protein folding and stability, of each residue in all ten studied proteins. Figure S6 illustrates the probability distribution functions of percentages of solvent accessibility areas of each amino acid computed from experimental and predicted structures (with and without template). As was expected, most of the non-polar residues are buried in protein except for proline, which despite its non-polar nature is not hydrophobic (40), consequently it is more solvent exposed than other non-polar residues. The differences between the results of experimental and both predicted structures are not significant, which indicate that side-chain conformations of non-polar residues are predicted correctly in terms of solvent accessibility.
The residues with uncharged polar and charged polar side chains, in general, are more exposed to the solvent than non-polar residues (Figure S6), however, side chains of some polar residues either are buried in protein (cysteine and tyrosine) or exhibiting no preference to be solvent exposed (histidine). Although the cysteine side chain is polar, it is considered hydrophobic based on the observation that it is often found in the interior of proteins largely due to its ability to form disulfide bonds. Moreover, the tyrosine is sometimes considered as a polar amino acid because its side chain has OH group, however, it also has a nonpolar benzene ring, and tyrosine is generally classified as a hydrophobic amino acid. Although histidine is a polar, positively charged amino acid, its physical properties depend very much on pH. In particular, the imidazole group of histidine is the only amino acid side chain affected by the range of physiological pH values; i.e., at pH 5.0 the group is positively charged, polar, and hydrophilic, whereas at pH 7.4 it is neutral, apolar, and hydrophobic (41). Therefore, a strong dependence of physical properties of histidine on pH might be a reason for broad distribution of percentages of SASA.
As in non-polar residues, the differences between the results of experimental and both predicted structures with and without templates for polar uncharged and polar charged residues, in general, are not significant.
It should be noted that PDFs for some polar uncharged (threonine and serine) and polar charged (aspartic acid and glutamic acid) amino acids do not exhibit strong propensity to be solvent exposed. This behavior of amino acids is in harmony with the hydropathy index defined by Kyte and Doolittle (42). Based on this work, (i) amino acids with a hydropathy index equal to or more than 1.8 are defined as hydrophobic; (ii) amino acids with a hydropathy index equal to or less than −3.3 are defined as hydrophilic; (iii) amino acids with a hydropathy index less than 1.8 and more than −3.3 are defined as neutral. Hydropathy indices for threonine and serine are −0.7 and −0.8, respectively, which indicate that both amino acids are neutral, consequently, it is not surprising that they do not exhibit strong propensity to be solvent exposed. Hydropathy index for both aspartic acid and glutamic acid is −3.5, which indicates that both amino acids are hydrophilic, however, their closeness to the threshold value (−3.3) might be a reason for a weak propensity to be solvent exposed.
Structural signatures of strongly cooperative double mutations are discovered by combining Potts sequence co-variation models with AlphaFold
There is a long history of using co-evolutionary information encoded in protein MSAs to probe the relationship between protein structure, function, and fitness (43–48). Potts Hamiltonian models provide a simple but powerful framework to capture the sequence co-variation patterns in MSAs using a machine learning approach, which share some characteristics with transformers, the machine learning algorithm employed by ColabFold (23). Implicitly, ColabFold is exploiting information about correlated mutations in its predictions of the sequence dependence of the most stable three-dimensional structure(s) of a protein. Furthermore, being able to predict non-additive effects of pairs of mutations is important for understanding how mutations alter protein stability, protein-protein interactions, ligand binding affinity, and many other sequence-structure properties that impact protein function. We have used Potts models extensively to study the fitness and conformational free energy landscapes of kinase family proteins (43,49–52), and to investigate the effects of mutations on the acquisition of drug resistance in HIV protein targets (43,53–56). The non-linear effects of pairs of cooperative mutations on protein fitness can be estimated at scale for large numbers of mutations using the Potts model for mutational scanning on many sequences in a protein family, whereby Potts double mutant cooperativity is calculated for all pairs of mutations in each sequence. Using Potts model mutational scanning as a preprocessing step to select mutation pairs that are predicted to be strongly coupled, followed by AlphaFold structure predictions of the corresponding sequences, is a way to identify changes in the rotamer states of side chains that may provide a possible structural explanation for the strong cooperativity indicated by the Potts model sequence-based calculations. In this section we illustrate how such modeling might proceed.
The Potts model can be used to estimate the cooperative effect of a pair of mutations on protein fitness through the calculation of :
| [Eq.1] |
The first term in square brackets is the relative log probability of the double mutation at positions and relative to the reference residues , while the second and third terms in square brackets are the corresponding log relative probabilities of the two single mutations. The field contributions to each of the terms cancel when combined in the formula. then depends on the four couplings between the residues pairs and at positions and . Importantly, is independent of the gauge of the Potts statistical energy function. The non-linear cooperativity is a measure of how much more favorable (or unfavorable) the fitness of the double mutation is than the sum of the effects of the individual mutations on the fitness. See the Appendix for a detailed description of the Potts Hamiltonian Model and the derivation of [Eq.1].
We have used a Potts model of Kinase Family proteins to compute a histogram of values for all pairs of mutations in ABL1 and PIM1, the results are shown in Figure 6 as histogram plots. We found that less than 1% of the total double mutations (0.78% for ABL 1 and 0.84% for PIM1) have the values greater than 1.5, indicating that only a small fraction of mutation pairs are strongly cooperative. Among these, the 10 most positively cooperative mutation pairs appearing in the tail of the distribution are labelled in Figure 6.
Figure 6.

Histogram Plots for double mutant cycle values for ABL1 and PIM1. The annotated mutant pairs represent the top 10 most positively cooperative double mutant pairs, with black indicating pairs that show the direct cooperative effect, while pink representing those whose cooperativity can be explained in an alternate way (see text).
The simplest structural signature of mutational cooperativity involving mutations at positions and in the sequence, is for one of the mutated residues to perturb the rotamer state of the other mutation. In other words, the rotamer state of residue changes depending on whether the residue at position is or . To test this, three AlphaFold structure predictions are required for each mutation pair: , and . We have performed this analysis for the ten mutation pairs that are predicted by the Potts model to be most strongly cooperative for both ABL1 and PIM1. For 5 of the 10 most cooperative mutation pairs predicted by the Potts model for ABL1 and PIM1, the AlphaFold structural predictions show the most direct structural signature of double mutant cooperativity (indicated by the black residue labels in the histogram, Fig. 6); AlphaFold predicts that the rotamer state of one mutant residue perturbs the rotamer state of the other mutant residue. These five cooperative mutation pairs which serve as an example, can mostly be explained as resulting from the maintenance of a charge-charge interaction involving the double mutant residues (e.g. ABL1: K247D/Y257K, and PIM1: L43D/Y53K), or by charge swapping (e.g. PIM1: K169S/T204K).
The cooperativity of several of the remaining 5 out of 10 pairs can also be accounted for (represented by pink in the histogram). In the case of the most cooperative pair K247E/E255R (ABL1), we observe the single mutant K247E in the wild-type sequence causes a rotamer change in the wild-type residue E255, supporting a cooperative effect without reference to the interactions between the double-mutant side chain rotamer states described above and listed in Table 1. It is also possible for mutation pairs to be cooperative without any rotamer state change, such as in K419T/S485Q (ABL1). The single mutant AF2 predictions show electrostatically repulsive groups in proximity while the wildtype and double mutant have electrostatically attractive groups in proximity. Even though the rotamers are locked in place by their environment, the attractive pairings of the double mutant stabilize the protein and improve folding. Others of the remaining pairs can be similarly rationalized, for instance in the double mutant W423Y/L452H (ABL1) the hydrophobically attractive pair WL is replaced by the electrostatically attractive mutant pair YH, without a change in the rotamer states. It is also possible that AF2 predicted rotamer states are incorrect in some cases, falsely predicting no rotamer state changes pairs that are cooperative. Experimental structures would be needed to test this. For the pair V422Q/P484F (ABL1), although no experimental structure of the double mutant is available, an experimental structure is available for a different kinase protein with the QF residue pair at position corresponding to 422/484 in the MSA, tuberculosis PknI kinase, and it shows a different rotamer state than the AF2 prediction of the QF double mutant, suggesting a possible AF2 rotamer misprediction. The remaining 5 cooperative double mutant pairs for PIM1 exhibit similar behavior to those of ABL1.
Table 1.
Mutant pairs with the double mutant rotamer states differs from the single mutant rotamer states with differences highlighted in red
| Protein | Cooperative Mutations | Mutant | Residue | RMSD () | ||||
|---|---|---|---|---|---|---|---|---|
| ABL1 (PDB ID: 2V7A) | H295S/Y353D | H295S | SER-295 | 173 | 0.25 | |||
| H295S/Y353D | SER-295 | 175 | ||||||
| ASP-353 | −74 | −17 | 1.87 | |||||
| Y353D | ASP-353 | −173 | 71 | |||||
| E258H/T267E | E258H | HIS-258 | −171 | −77 | 1.02 | |||
| E258H/T267E | HIS-258 | −176 | −135 | |||||
| GLU-267 | −63 | −168 | 163 | 0.26 | ||||
| T267E | GLU-267 | −71 | −177 | 163 | ||||
| E316R/K378F | E316R | ARG-316 | −176 | 75 | 72 | −165 | 1.79 | |
| E316R/K378F | ARG-316 | −175 | 169 | −72 | −101 | |||
| PHE-378 | −56 | −102 | 0.06 | |||||
| K378F | PHE-378 | −58 | −103 | |||||
| K247D/Y257K | K247D | ASP-247 | −174 | −7 | 1.81 | |||
| K247D/Y257K | ASP-247 | −69 | −23 | |||||
| LYS-257 | −62 | −176 | 175 | −176 | 0.19 | |||
| Y257K | LYS-257 | −61 | −178 | 178 | −178 | |||
| E409N/S420D | E409N | ASN-409 | 61 | −66 | 2.53 | |||
| E409N/S420D | ASN-409 | −162 | 1 | |||||
| ASP-420 | −62 | −76 | 0.39 | |||||
| S420D | ASP-420 | −68 | −89 | |||||
| PIM1 (PDB ID: 1XWS) | K169S/T204K | K169S | SER-169 | 72 | 0.06 | |||
| K169S/T204K | SER-169 | 70 | ||||||
| LYS-204 | −66 | −177 | −176 | −177 | 4.03 | |||
| T204K | LYS-204 | −176 | 77 | −178 | −177 | |||
| S54R/R122D | S54R | ARG-54 | −173 | 179 | −171 | −113 | 3.18 | |
| S54R/R122D | ARG-54 | −168 | 168 | −88 | −177 | |||
| ASP-122 | −172 | 18 | 0.04 | |||||
| R122D | ASP-122 | −171 | 19 | |||||
| D176R/K183C | D176R | ARG-176 | −169 | 167 | −164 | −125 | 4.59 | |
| D176R/K183C | ARG-176 | −58 | 174 | −172 | −86 | |||
| CYS-183 | 59 | 0.18 | ||||||
| K183C | CYS-183 | 61 | ||||||
| L43D/Y53K | L43D | ASP-43 | −174 | 14 | 1.80 | |||
| L43D/Y53K | ASP-43 | −71 | −23 | |||||
| LYS-53 | −63 | −175 | 175 | −175 | 0.32 | |||
| Y53K | LYS-53 | −62 | −178 | −179 | 179 | |||
| V225Q/P279F | V225Q | GLN-225 | −77 | 89 | 24 | 1.45 | ||
| V225Q/P279F | GLN-225 | −66 | 157 | −9 | ||||
| PHE-279 | −81 | −8 | 0.47 | |||||
| P279F | PHE-279 | −82 | −23 |
In this section we have illustrated how the Potts sequence-based energy model can be used to perform mutational scans of a protein to search for the most strongly cooperative mutational pairs, and AlphaFold can then be used to detect the structural changes associated with the very strong cooperativity. Further developing these ideas, it should be possible to integrate the Potts model and AlphaFold into a single pipeline which can search for mutational pairs that have the strongest effects on protein fitness, build structural models of these mutant proteins and their assemblies, and use these models to guide experiments designed to probe the role this cooperativity plays in the protein’s function and dysfunction.
Conclusions
We have analyzed the side chain rotamer states predicted by ColabFold for every amino acid residue in ten different proteins; a total of 1453 side chain rotamer state predictions. ColabFold is an online implementation of AlphaFold2. Unlike the striking accuracy (the root mean square deviation of atoms averaged over the ten studied proteins is less than ) observed between backbones of experimental PDB and ColabFold predicted structures, discrepancies between side-chain angles of experimental and predicted structures are significantly larger. Using a deviation of more than +/− 40 degrees from the experimental value as a definition of side chain error, we find that over a set of 10 benchmark proteins, the rotamer state prediction error of ColabFold is on average for dihedral angles, and increases to for dihedral angles. The prediction error is smaller for non-polar side chains and is improved using structural templates. When structural templates are employed, the largest improvement is for dihedrals ( improvement compared to without templates).
ColabFold demonstrates a bias towards the most prevalent rotamer states in the protein data bank, potentially limiting its ability to capture rare side chain conformations effectively. The rotamer state prediction error rate increases near Alanine and Glycine residues, which surprisingly appears to be caused by a bias for ColabFold to predict the most populated rotamer states in the protein data bank, while the lack of packing constraints adjacent to Alanine and Glycine residues or the increased backbone flexibility near these residues may play a less significant role. We hope these observations can lead to approaches to improve the training of AlphaFold to better account for rotamer states of residues that are only rarely observed in the PDB.
As a first application of AlphaFold to explore the structural consequences of strongly cooperative mutations on side chain rearrangements, we employ a Potts sequence-based energy model to perform large scale mutational scans of two proteins ABL1 and PIM1 kinase, searching for the most strongly cooperative mutational pairs, and then use ColabFold to predict the structural signatures of this cooperativity on the interacting side-chains. The predictions of the sequence-based Potts statistical energy model, and structure based AlphaFold are largely consistent, in that for more than 50% of the examples, the structure of one side chain mutant significantly perturbs the structure of the second side chain (either wild type or mutant), leading to increased stability of the folded protein according to the Potts sequence-based statistical energy model; in the remaining cases the increased stability can be rationalized without requiring a rotamer state change. Our results demonstrate that the integration of the sequence-based Potts model with AlphaFold into a single pipeline provides a new tool that can be used to explore the fundamental relationship between protein mutations, cooperative changes in structure, and fitness.
Supplementary Material
Acknowledgement
We thank Dr. Sompriya Chatterjee for performing the RMSD calculations and providing the scripts used to perform these calculations. This work was supported in part by the National Institutes of Health grant NIH R35 GM132090.
Appendix: Introduction to the Potts model Hamiltonian and derivation of equation 1 in the main text
The cooperative mutation effects can be modeled using statistical mechanics and Boltzmann networks by learning the mutational patterns in the protein families using Multiple Sequence Alignments (MSA) (57). Our sequence-based double mutant cooperativity calculation is based on the MSA used to derive the protein Kinase Potts Hamiltonian model which describes the coevolutionary interactions between all possible residue pairs in the sequence with all possible amino acid combinations at each pair of positions. Analogous to the Hamiltonian of the magnetic spin system in the Condensed Matter Physics, it is defined by,
| [Eq. 1] |
where and are the fields and couplings, also called single-site and pairwise statistical energy terms, respectively (58). The coupling term gives the strength of the interaction between the residues and at positions i and j, respectively whereas represents the field for residue Si at position i in the sequence . The indices i and j represent the amino acid positions in the aligned sequence of length . Each position can have 20 possible states with one gap character. The statistical energy E of a sequence of length gives the probability of that sequence to be observed in the MSA relative to all other sequences following the Boltzmann distribution.
where is the partition function that takes sum over all possible sequences and ensures the normalization.
The model is parametrized such that it exactly reproduces the following MSA marginals, the single-site residue frequencies and joint frequencies of every pair of residues in the , also referred to as univariate and bivariate residue frequencies, respectively, and are defined by:
| [Eq. 2] |
Here, is the MSA depth (number of sequences in the MSA), and is Kronecker delta function where if the residue is observed at position i in the sequence , and otherwise.
The univariate and bivariate marginals from the model is difficult to estimate directly as it requires computing partition function which takes sums over sequences which is complex. There are several methods for inferring the model parameters such that the model distribution reproduces the empirical univariate and bivariate marginals of the MSA. The protein Kinase Potts Hamiltonian Model implements Markov Chain Monte Carlo (MCMC) based Mi3-GPU method for such purpose. Through the MCMC sampling, it generates a synthetic MSA from which marginals can be calculated. With an initial guess for the couplings, the model generates synthetic MSA from which marginals are calculated and used those values to iteratively update the coupling parameter until the empirical marginals of the MSA dataset are recovered. The resulting Hamiltonian can generate the sequences whose univariate and bivariate marginals match with empirical MSA marginals. The detailed explanation can be studied in (59).
The Potts model can be used to estimate the cooperative effect of a pair of mutations on protein fitness. This double mutant cooperativity is defined by subtracting the additive effects of two single mutations at position i and j from the double mutation effect. Mathematically, it is defined by:
| [Eq. 3] |
The first term in right hand side gives the change in energy when both residues i and j are mutated while the two terms inside parenthesis contribute energy change due to single mutation at position and , respectively. These terms can be further defined as:
| [Eq. 4] |
| [Eq. 5] |
Substituting [Eq.4] and [Eq. 5] in [Eq. 3], we get:
| [Eq. 6] |
where the first term in right hand side represents the relative log probability of the double mutation at positions i and j relative to the reference residues , while the second and third terms are the corresponding log relative probabilities of the two single mutations. After using [Eq.1] in [Eq.6], the field contributions to each of the terms cancel when combined in the formula, and [Eq.6] can be written:
| [Eq. 7] |
then depends on the four couplings between the residues pairs and at positions and . Importantly, is independent of the gauge of the Potts statistical energy function. The non-linear cooperativity is a measure of how much more favorable (or unfavorable) the fitness of the double mutation is than the sum of the effects of the individual mutations on the fitness.
Footnotes
Declaration of Interests
The authors declare no competing interests.
References
- 1.Anfinsen C. B., Haber E., Sela M., and White F. H.. 1961. The kinetics of formation of native ribonuclease during oxidation of the reduced polypeptide chain. Proc. Natl. Acad. Sci. U.S.A. 47:1309–1314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Wüthrich K. 1990. Protein structure determination in solution by NMR spectroscopy. J. Biol. Chem. 265:22059–22062. [PubMed] [Google Scholar]
- 3.Shi Y. 2014. A glimpse of structural biology through X-ray crystallography. Cell 159:995–1014. [DOI] [PubMed] [Google Scholar]
- 4.Earl L. A., Falconieri V., Milne J. L., and Subramaniam S.. 2017. Cryo-EM: beyond the microscope. Curr. Opin. Struct. Biol. 46:71–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Zidek A., Potapenko A., et al. 2021. Highly accurate protein structure prediction with AlphaFold. Nature 596:583–589. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Gao M., and Skolnick J.. 2024. Improved deep learning prediction of antigen-antibody interactions. Proc. Natl. Acad. Sci. U.S.A. 121:e2410529121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Steinbrecher T., Abel R., Clark A., and Friesner R.. 2017. Free energy perturbation calculations of the thermodynamics of protein side-chain mutations. J. Mol. Biol. 429:923–929. [DOI] [PubMed] [Google Scholar]
- 8.Pucci F., Bernaerts K. V., Kwasigroch J. M., and Rooman M.. 2018. Quantification of biases in predictions of protein stability changes upon mutations. Bioinformatics 34:3659–3665. [DOI] [PubMed] [Google Scholar]
- 9.de Oliveira C., Yu H. S., Chen W., Abel R., and Wang L.. 2019. Rigorous free energy perturbation approach to estimating relative binding affinities between ligands with multiple protonation and tautomeric states. J. Chem. Theory Comput. 15:424–435. [DOI] [PubMed] [Google Scholar]
- 10.Duan J., Lupyan D., and Wang L.. 2020. Improving the accuracy of protein thermostability predictions for single point mutations. Biophys. J. 119:115–127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Coskun D., Chen W., Clark A. J., Lu C., Harder E. D., Wang L., Friesner R. A., and Miller E. B.. 2022. Reliable and accurate prediction of single-residue pKa values through free energy perturbation calculations. J. Chem. Theory Comput. 18:7193–7204. [DOI] [PubMed] [Google Scholar]
- 12.Sergeeva A. P., Katsamba P. S., Liao J., Sampson J. M., Bahna F., Mannepalli S., Morano N. C., Shapiro L., Friesner R. A., and Honig B.. 2023. Free energy perturbation calculations of mutation effects on SARS-CoV-2 RBD::ACE2 binding affinity. J. Mol. Biol. 435:168187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Thakur A., Gizzio J., and Levy R. M.. 2024. Potts Hamiltonian models and molecular dynamics free energy simulations for predicting the impact of mutations on protein kinase stability. J. Phys. Chem. B 128:1656–1667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Dunbrack R. L. Jr., and Karplus M.. 1993. Backbone-dependent rotamer library for proteins: Application to side-chain prediction. J. Mol. Biol. 230:543–574. [DOI] [PubMed] [Google Scholar]
- 15.Dunbrack R. L. Jr., and Karplus M.. 1994. Conformational analysis of the backbone-dependent rotamer preferences of protein side chains. Nat. Struct. Biol. 1:334–340. [DOI] [PubMed] [Google Scholar]
- 16.Dunbrack R. L. Jr. 2002. Rotamer libraries in the 21st century. Curr. Opin. Struct. Biol. 12:431–440. [DOI] [PubMed] [Google Scholar]
- 17.Shapovalov M.S., and Dunbrack R. L. Jr. 2011. A smoothed backbone-dependent rotamer library for proteins derived from adaptive kernel density estimates and regressions. Structure 19:844–858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Kryshtafovych A., Schwede T., Torf M., Fidelis K., and Moult J.. 2021. Critical assessment of methods of protein structure prediction (CASP) - Round XIV. Proteins 89:1607–1617. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Terwilliger T. C., Liebschner D., Croll T. I., Williams C. J., McCoy A. J., Poon B. K., Afonine P. V., Oeffner R. D., Richardson J. S., Read R. J., et al. 2024. AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. Nat. Methods 21:110–116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Li E. H., Spaman L. E., Tejero R., Huang Y. J., Ramelot T. A., Fraga K. J., Prestegard J. H., Kennedy M. A., and Montelione G. T.. 2023. Blind assessment of monomeric AlphaFold2 protein structure models with experimental NMR data. J. Magn. Reson. 352:107481. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Mariani V., Biasini M., Barbato A., and Schwede T.. 2013. IDDT: a local superposition – free score for comparing protein structures and models using distance difference tests. Bioinformatics 29:2722–2728. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Tunyasuvunakool K., Adler J., Wu Z., Green T., Zielinski M., Zidek A., Bridgland A., Cowie A., Meyer C., Laydon A., et al. 2021. Highly accurate protein structure prediction for the human proteome. Nature 596:590–596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Mirdita M., Schutze K., Moriwaki Y., Heo L., Ovchinnikov S., and Steinegger M.. 2022. ColabFold: making protein folding accessible to all. Nat. Methods 19:679–682. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Chen J., Lu Z., Sakon J., and Stites W. E.. 2000. Increasing the thermostability of staphylococalnuclease: implications for the origin of protein thermostability. J. Mol. Biol. 303:125–130. [DOI] [PubMed] [Google Scholar]
- 25.Weaver L. H., and Matthews B. W.. 1987. Structure of bacteriophage T4 lysozyme refined at 1.7 Å resolution. J. Mol. Biol. 193:189–199. [DOI] [PubMed] [Google Scholar]
- 26.Nicholson H., Anderson D. E., Dao-pin S., and Matthews B. W.. 1991. Analysis of the interaction between charged side chains and the -helix dipole using designed thermostable mutants of phage T4 lysozyme. Biochemistry 30:9816–9828. [DOI] [PubMed] [Google Scholar]
- 27.Modugno M., Casale E., Soncini C., Rosettani P., Colombo R., Lupi R., Rusconi L., Fancelli D., Carpinelli P., Cameron A. D., et al. 2007. Crystal structure of the T315I Abl mutant in complex with the aurora kinases inhibitor PHA-739358. Cancer Res. 67:7987–7990. [DOI] [PubMed] [Google Scholar]
- 28.Lori C., Lantella A., Pasquo A., Alexander L. T., Knapp S., Chiaraluce R., and Consalvi V.. 2013. Effect of single amino acid substitution observed in cancer on Pim-1 kinase thermodynamic stability and structure. PLOS One 8:e64824. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Kondo H., Hanada Y., Sugimoto H., Hoshino T., Garnham C. P., Davies P. L., and Tsuda S.. 2012. Ice-binding site of snow mold fungus antifreeze protein deviates from structural regularity and high conservation. Proc. Natl. Acad. Sci. U.S.A. 109:9360–9365. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Leinala E. K., Davies P. L., and Jia Z.. 2002. Crystal structure of β-helical antifreeze protein points to a general ice binding model. Structure 10:619–627. [DOI] [PubMed] [Google Scholar]
- 31.Xu Q., Biancalana M., Grant J. C., Chiu H. J., Jaroszewski L., Knuth M. W., Lesley S. A., Godzik A., Elsliger M. A., Deacon A. M., et al. 2019. Structures of single-layer β-shette proteins evolved from β-hairpin repeats. Protein Sci. 28:1676–1689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Yi J., Thomas L. M., and Richter-Addo G. B.. 2012. Distal pocket control of nitrite binding in myoglobin. Angew. Chem. Int. Ed. Engl. 51:3625–3627. [DOI] [PubMed] [Google Scholar]
- 33.Lu P., Min D., Dimaio F., Wei K. Y., Vahey M. D., Boyken S. E., Chen Z., Fallas J. A., Ueda G., Sheffler W., et al. 2018. Accurate computational design of multipass transmembrane proteins. Science 359:1042–1046. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Steinegger M., and Soding J.. 2017. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35:1026–1028. [DOI] [PubMed] [Google Scholar]
- 35.Mirdita M., Steinegger M., and Sording J.. 2019. MMserqs2 desktop and local web server app for fast, interactive sequence searches. Bioinformatics 35:2856–2858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Wiedemann C., Kumar A., Lang A., and Ohlenschlager O.. 2020. Cysteines and disulfide bonds as structure-forming units: Insights from different domains of life and the potential for characterization by NMR. Front. Chem. 8:280. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Beard H., Cholleti A., Pearlman D., Sherman W., and Lovin K. A.. 2013. Applying physics-based scoring to calculate free energies of binding for single amino acid mutations in protein-protein complexes. PLOS One 8:e82849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Salam N. K., Adzhigirey M., Sherman W., and Pearlma D. A.. 2014. Stricture-based approach to the prediction of disulfide bonds in proteins. PEDS 27:365–374. [DOI] [PubMed] [Google Scholar]
- 39.Towse C-L., Rysavy S. J., Vulovic I. M., and Daggett V.. 2016. New dynamic rotamer libraries: Data-driven analysis of side-chain conformational propensities. Structure 24:187–199. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Pommie C., Levadoux S., Sabatier R., Lefranc G., and Lefranc M-P.. 2004. IMGT standardized criteria for statistical analysis of immunoglobulin V-REGOIN amino acid properties. J. Mol. Recognit. 17:17–32. [DOI] [PubMed] [Google Scholar]
- 41.Rotzschke O., Lau J. M., Hofstatter M., Falk K., and Strominger J. L.. 2002. A pH-sensitive histidine residue as control element for ligand release from HLA-DR molecules. Proc. Natl. Acad. Sci. U.S.A. 99:16946–16950. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Kyte J., and Doolittle R. F.. 1982. A simple method for displaying the hydropathic character of a protein. J. Mol. Biol. 157:105–132. [DOI] [PubMed] [Google Scholar]
- 43.Levy R. M., Haldane A., and Flynn W. F.. 2017. Potts Hamiltonian models of protein covariation, free energy landscapes, and evolutionary fitness. Curr. Opin. Struct. Biol. 43:55–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.de Juan D., Pazos F., and Valencia A.. 2013. Emerging methods in protein co-evolution. Nat. Rev. Genet. 14:249–261. [DOI] [PubMed] [Google Scholar]
- 45.Stein R. R., Marks D. S., and Sander C.. 2015. Inferring pairwise interactions from biological data using maximum-entropy probability models. PLOS Comput. Biol. 11:e1004182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Serohijos A. W., and Shakhnovich E. I.. 2014. Merging molecular mechanism and evolution: theory and computation at the interface of biophysics and evolutionary population genetics. Curr. Opin. Struct. Biol. 26:84–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Neuwald A. F. 2016. Gleaning structural and functional information from correlations in protein multiple sequence alignments. Curr. Opin. Struct. Biol. 38:1–8. [DOI] [PubMed] [Google Scholar]
- 48.Dinan J. C., McCormick J. W., and Reynolds K. A.. 2024. Engineering proteins using statistical models of coevolutionary sequence information. Cold Spring Harb. Perspect. Biol. 16:a041463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Haldane A., Flynn W. F., He P., Vijayan R. S. K., and Levy R. M.. 2016. Structural propensities of kinase family proteins from a Potts model of residue co-variation. Prot. Sci. 25:1378–1384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Haldane A., Flynn W. F., He P., and Levy R. M.. 2018. Co-Evolutionary landscape of kinase family proteins, subsequence probabilities, and functional motifs. Biophys. J. 114:21–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Gizzio J., Thakur A., Haldane A., and Levy R. M.. 2022. Evolutionary divergence in the conformational landscapes of tyrosine vs serine/threonine kinases. eLife 11:e83368. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Gizzio J., Thakur A., Haldane A., Post C. B., and Levy R. M.. 2024. Evolutionary sequence and structural basis for the district conformational landscapes of Tyr and Ser/Thr kinases. Nat. Commun. 15:6545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Flynn W. F., Haldane A., Torbett B. E., and Levy R. M.. 2017. Inference of epistatic effects leading to entrenchment and drug resistance in HIV-1 protease. Mol. Biol. Evol. 34:1291. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Biswas A., Haldane A., Arnold E., and Levy R. M.. 2019. Epistasis and entrenchment of drug resistance in HIV-1 subtype B. eLife 8:e50524. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Li M., Oliveira Passos D., Shan Z., Smith S. J., Sun Q., Biswas A., Choudhuri I., Strutzenberg T. S., Haldane A., Deng N., et al. 2023. Mechanisms of HIV-1 integrase resistance to dolutegravir and potent inhibition of drug-resistant variants. Sci. Adv. 9:eadg5953. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Biswas A., Choudhuri I., Arnold E., Lyumkis D., Haldane A., and Levy R. M.. 2024. Kinetic coevolutionary models predict the temporal emergence of HIV-1 resistance mutations under drug selection pressure. Proc. Natl. Acad. Sci. U.S.A. 121:e2316662121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Lapades A., Giraud B., and Jarzynski C.. 2012. Using sequence alignments to predict protein structure and stability with high accuracy. arXiv, arXiv:1207.2484. [Google Scholar]
- 58.Levy R. M., Haldane A., and Flynn W. F.. 2017. Potts Hamiltonian models of protein co-variation, free energy landscapes, and evolutionary fitness. Curr. Opin. Struct. Biol. 43:55–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Haldane A., and Levy R. M.. 2021. Mi3-GPU: MCMC-based inverse Ising inference on GPUs for protein covariation analysis. Comput. Phys. Commun. 260:107312. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.





