Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 Jan 1.
Published in final edited form as: Methods Mol Biol. 2022;2340:41–50. doi: 10.1007/978-1-0716-1546-1_3

Using Surface Hydrophobicity Together with Empirical Potentials to Identify Protein-Protein Binding Sites. Application to the Interactions of E-cadherins

Robert L Jernigan 1,*, Pranav Khade 1, Ambuj Kumar 1, Andrzej Kloczkowski 2,3
PMCID: PMC9131873  NIHMSID: NIHMS1803779  PMID: 35167069

Abstract

Studying the interactions within protein structures can inform about the details of how proteins of various types interact and aggregate. Empirical contact potentials have proven to be extremely important in the evaluation of individual modeled protein structures, but have found few applications to protein-protein interactions. In part this is caused by a lack of properly formulated potentials with a proper reference state. Since the comparisons are made between different bound structures, the proper reference state should take into account other contacts. Therefore a preferred reference state should be defined with respect to a given residue type interacting with an average residue instead of interacting with solvent as typically is used in derivation of statistical contact potentials. Here, a two stage procedure for generating and evaluating interacting protein pairs is described, and an example of E-cadherin interactions is shown.

Introduction

Uncovering the underlying geometric and chemical complementarity for protein-protein binding has a long history, with a wide variety of approaches having been pursued. In the early days, regions of high surface shape complementarity (Shoichet et al., 1992) were used to identify the best molecular geometry for docking (Kuntz et al., 1982). Nicholls et al. (1991) even proposed surface curvature as a useful tool for locating binding sites. Also considered were regions of high density in hydrogen bond sites (Danziger & Dean, 1989a, 1989b), and those areas with favorable electrostatic properties (Goodford, 1985) were used to optimize the chemical complementarity as well as to search selected surfaces for potential binding sites. Studies using molecular dynamics simulations (Brooks et al., 1988) were also applied to understand better the principal characteristics of molecular associations (Janin & Chothia, 1990). And, there are many newer papers on this subject (Jones & Thornton, 1996; Keskin et al., 2008 & 2016).

Protein-protein interactions play a critical role in protein function. Completion of many genomes is being followed rapidly by major efforts to identify interacting protein pairs experimentally in order to decipher problem with broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. Experimental identification of residues in protein-protein interaction surfaces must come either from determination of the structure of protein-protein complexes or from functional assays, such as yeast two hybrid experiments. The rapidly growing experimental data of protein structures (although the number of solved protein-protein complexes is still limited) and on protein-protein interactions, as well as improved methods of protein structure prediction led in the last 20 years to the development of computational methods for identifying and for prediction of amino acid residues that participate in protein-protein interactions. Integration of information from various sources (sequential information, structural information, interactome information, pathway databases), identification of conserved residues, and of co-evolution couplings, as well as literature text mining helps to identify protein-protein interactions and to better understand molecular mechanisms of protein-protein interactions. This accumulating information is being assembled in several public databases such as DIP (Salwinski, 2004), BIOGRID (Oughtred et al., 2019), and STRING (Szklarczyk, 2019).

Hydrophobicity

Hydrophobicity is important for protein-protein interactions and binding of other molecules. Usually these occur in clusters and the binding sites are found to be located at one of the strongest hydrophobic clusters on the surface of the protein acceptor. (Young et al., 1994) found that in 25 of 38 cases, the binding was to the strongest hydrophobic cluster and in the remaining cases it was located in one of the top 6 cases. Also, it was reported that the most hydrophobic cluster coinciding with more than one-third of the surface buried by the bound ligand. The remaining 13 cases bound to one of the top 6 hydrophobic clusters. These results suggest that, at surface hydrophobicity can be used to identify those regions on a protein's surface most likely to bind any ligand (Young, Jernigan & Covell, 1994).

Why is this? It is primarily because the hydrophobic interacting pairs are the strongest types of interactions, so matching hydrophobic residues is the most important for gaining strength (low energy). The interactions between two proteins have been shown to follow the usual contact interactions for interactions between amino acids coming from interactions with solvent (Keskin et al., 1998). And, just as for contacts between amino acids within a single protein, the specificity for the interactions comes from the interactions between polar and charged residues. The work showed that the reference state for these potentials is forming an interactions from the amino acid being solvent exposed, which holds for both intermolecular interactions and intramolecular interactions. The relative values of this potential function are shown in Fig. 1. Principally these reflect hydrophobicity values. Another hydrophobicity-based scale was developed by Chris Dobson’s group for use to identify the aggregation tendencies of each amino acid (Pawar et al., 2005).

Figure 1.

Figure 1.

Contact potential for pairs of amino acids in different classes. H is for hydrophobic residues, P for polar residues, A for negatively charged amino acids and B for positively charged amino acids. It can be clearly seen that the strongest interactions by far are for pairs of hydrophobic amino acids in contact in structures.

Both of these studies point out the importance of hydrophobicity for binding strength between proteins. The importance of hydrophobicity of interacting residues for protein contact potentials (CPs) was revealed by the detailed statistical analysis of 29 different contact potentials published in the literature (Pokarowski et al., 2005). It has been shown that all analyzed CPs can be divided into two groups, regardless of their having completely different derivation origin. Most of these knowledge-based statistical potentials could be well approximated by appropriate combinations of one-body components. The CP matrices of the first class can be approximated with a correlation of order 0.9 by the formula eij = hi + hj, 1 ≤ i, j ≤ 20, where the residue-type dependent factor h is highly correlated with the frequency of occurrence of a given amino acid type inside proteins. Potentials belonging to the second class can be approximated with a correlation of 0.9 by the formula eij = c0hihj + qiqj, where c0 is a constant, h is highly correlated with the Kyte–Doolittle hydrophobicity scale, and a new, less dominant, residue-type dependent factor q is correlated (~0.9) with amino acid isoelectric points pI.

Cadherins

Cadherins are a class of type-1 transmembrane proteins that play important role in cell adhesion. Their name comes from the phrase "calcium-dependent adhesion". They play a critical role in the formation of adherens junctions that bind cells together. Cadherins mediate homotypic adhesion between cells. Another type of proteins named integrins mediate adhesion between the cell and the extracellular matrix. By regulating contact formation and stability, cadherins play a crucial role in development: cell adhesion, sorting, and tissue morphogenesis. Cadherins are also involved in signaling in cancer, and play an important role as a therapeutic target in oncological studies.

The external domain of a cadherin that is located outside of a cell is composed of multiple repeats of the same protein chain where each repeat can bind calcium. Calcium binding makes the whole chain rigid enough to connect it with a chain from another cell. The role of cadherins in molecular biology and medicine has been reviewed in (Van Roy, 2013).

Cadherin Interactions

One of the most important cases of protein-protein interactions is between two identical cadherins, whose interactions control cell-cell interactions (Manibog et al., 2016). It serves as an interesting example of protein-protein interactions.

Method for evaluating various interacting conformations between two rigid protein structures.

This is a two stage procedure: 1) identify the strong potential hydrophobic patches and 2) rotate and translate pairs of hydrophobic patches with respect to each other to achieve the most favorable non-hydrophobic interactions. This first step is premised on the fact that hydrophobic interactions are quite non-specific and are often able to rearrange side chains to accommodate one another and maintain strong interactions. The second step is critical for obtaining the optimal interaction position. For this purpose the interactions that evaluate the specificity should be used, which use a reference state of exposure to an average residue.

There are many of these empirical contact potentials, which have proven themselves to be the most effective type of potentials for evaluating predicted protein structures at the CASP competitions (Feng, 2010). One of the earliest was that of Miyazawa and Jernigan from ‘1985, which was subsequently refined by using more structures in 1996 (Miyazawa & Jernigan, 1996). This latter potential is the one we choose to describe here.

Stage 1 – Identify the strongest hydrophobic patches. Use the values from Table 1 for all possible patches separately. Then, choose a set of the strongest ones for considering in Stage 2.

Table 1.

Contact potential between residues types and an average residue (Miyazawa & Jernigan, 1996). These values can also be used as a hydrophobicity scale and have a similarity to other hydrophobicity scales.

C 3.57
M 3.92
F 4.76
I 4.42
L 4.81
V 3.89
W 3.81
Y 3.41
A 2.57
G 2.19
T 2.29
S 1.98
N 1.92
Q 2.00
D 1.84
E 1.79
H 2.56
R 2.11
K 1.52
P 2.09

Stage 2 – Evaluate interactions between all pairs of interacting non-hydrophobic residues outside the hydrophobic patches identified in Stage 1 by using the contact potentials with a reference state of solvent exposure (Table 2). These values were not published in the 1996 paper (Miyazawa & Jernigan, 1996) but have been computed here according to the equation Eij = eij + err – eir - ejr. Since placing intermolecular hydrophobic patches together will simultaneously cause polar residues on the outside fringe to come into contact, it is important to evaluate these interactions with a different reference state of interacting with another amino acid. These values are provided in Table 2. And, at the same time, similar evaluations of the hydrophobic residue pairs should be carried out.

Table 2.

Specificity contact potentials for pairs of residue types. Eight residue types (C, M, F, I, L, V, W, Y) with the highest values of hydrophobicity in Table 1 were excluded. Because of the symmetry of the potentials only the upper half is shown.

A G T S N Q D E H R K P
A 1.07 0.84 0.93 0.88 0.92 1.08 0.94 1.18 1.15 1.22 1.13 1.12
G 0.27 0.53 0.43 0.38 0.67 0.41 0.83 0.77 0.69 0.65 0.64
T 0.59 0.39 0.34 0.53 0.3 0.41 0.60 0.61 0.59 0.71
S 0.32 0.28 0.58 0.11 0.31 0.55 0.53 0.49 0.68
N 0.05 0.23 −0.07 0.15 0.45 0.38 0.20 0.59
Q 0.61 0.36 0.45 0.76 0.43 0.33 0.60
D 0.28 0.52 0.09 −0.39 −0.39 0.67
E 0.68 0.31 −0.32 −0.46 0.79
H 0.28 0.66 0.86 0.67
R 0.76 1.11 0.71
K 0.97 0.83
P 0.76

Some Simple Questions about Protein Aggregation

Protein aggregation is a widespread problem, from protein syntheses where inclusion bodies are improperly folded aggregates to amyloid deposits in the brain. The characteristics of amino acid interactions are important considerations for this problem. In general, hydrophobic interactions are the strongest and provide substantial stability for protein binding. The remaining residue types provide specificity to favor specific cases out of the ones having hydrophobic interactions.

In an earlier study we observed that the most hydrophobic patches on protein surfaces are often the sites for not only proteins to interact with one another but also for small molecule ligand binding. There were 38 co-molecule crystal structures considered where we observed that a small set of the most hydrophobic patches were the actual binding site. This also holds clearly for cadherin, as can be seen in Fig. 3 for E cadherin dimers. These were static structures, and there is also some possibility that the protein’s dynamics exposes some essential hydrophobic binding sites.

Figure 3.

Figure 3.

High Hydrophobicity (95% highest cases shown in magenta) of all the tessellations in the nrPDB) tessellation patches for the interface of the cadherin trans-dimer (PDB: 3Q2V); The interaction between the monomers is stabilized by this large magenta patch shown in the center and further supported by a geometric fit by tryptophan 2 near the high hydrophobicity patch. See Fig. 2 for a closer view of the binding hydrophobic patch at the interaction site.

The characteristics of the sequences of globular proteins are to have a balance of hydrophobic residues for strength of interactions together with a proper balance polar and charged residues. The polar and charged residues impart specificity to yield a more limited set of specific favorable structures. The polar non-charged residues are especially important since they can form hydrogen bonds.

Many multimeric proteins have strong hydrophobic patches at their interfaces. Some other structures such as hemoglobin have more unusual distributions of hydrophobic patches on their surfaces (Fig. 4). Clearly the hemes bind at hydrophobic patches but the remaining hydrophobic patch distribution is somewhat puzzling. Does avoiding these relatively small hydrophobic patches help to guide oxygen along the more polar surface to reach its binding site? Or could it be guiding the oxygen to entries of some of the important pathways to the hemes (Lin & Song, 2011)?

Figure 4.

Figure 4.

Tetrameric hemoglobin showing its hydrophobic surface patches in magenta, with the hemes shown as sticks. The four subunitsare shown in yellow, beige, gray and turquoise.

There are significant advantages to investigating protein aggregation with coarse-grained models such as is proposed here. What has been discussed here is appropriate for globular proteins where the dynamics has been ignored. Further studies should include treatment of the dynamics. Also, there are other types of proteins where aggregation is important, such as membrane proteins and fibrous proteins. In addition there is the important problem of aggregation within the intrinsically disordered proteins that is an important aspect of their structures. The general sequence characterization of disordered proteins is having a high fraction of charged amino acids as their hallmark, and this has implications of the amino acid composition in general. The remaining hydrophobic residues are likely to aggregate, with some the polar determining more precisely how these interact. This is an undeveloped research area. Would it be possible to predict which protein sequences would be most likely to aggregate? And, even to develop molecular models of possible interacting structures, no matter how evanescent they may be? Could sequentially close hydrophobic residues induce partial intramolecular folding establishing a balance between that and intermolecular aggregation? All these problems need further studies to shed more light on the molecular mechanisms of protein aggregation.

Figure 2.

Figure 2.

E-Cadherin domain showing the residues in dimeric interactions when attached to an identical domain from an opposing cell. The structure shown is one chain the PDB structure 3q2v. The closely interacting residues are shown in color, with the hydrophobic residues colored magenta, and the interacting residues on the edge of the hydrophobic patch in yellow. In addition tryptophan (strongly hydrophobic in red) interacts strongly by fitting into a pocket ono the opposing chain. The residues in yellow are mostly polar but in this case also include several prolines, which could also have been considered to be part of the hydrophobic patch.

Acknowledgments:

We acknowledge support from NSF grant DBI 1661391, and NIH grant R01 GM127701.

REFERENCES

  1. Brooks CL Ill, Karplus M, Pettit M 8. 1988. Proteins: A theoretical perspective of dynamics, structure, and thermodynamics. New York: J.Wiley [Google Scholar]
  2. Capon OJ. Ward RHR. 1991. The CD4-gpl20 interaction and AJDS pathogenesis. Annu Rev Immunol 9:649–678. [DOI] [PubMed] [Google Scholar]
  3. Danziger DJ, Dean PM. 1989a. Automated site-directed drug design: A general algorithm for knowledge acquisition about hydrogen-bonding regions at protein surfaces. Proc Roy Soc Lond B 236:101–113. [DOI] [PubMed] [Google Scholar]
  4. Danziger DJ, Dean PM. 1989b. Automated site-directed drug design: The prediction and observation of ligand point positions at hydrogen-bonding regions on protein surfaces. Proc Roy Soc Lond B 236:115–124. [DOI] [PubMed] [Google Scholar]
  5. Feng Y, Kloczkowski A, Jernigan RL. 2010. Potentials’R’Us web -server for protein energy estimations with coarse-grained knowledge-based potentials. BMC Bioinformatics 11:92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Goodford P 1985. A computational procedure for determining energetically favorable binding sites on biologically important macromolecules. J Med Chem 28:849–857. [DOI] [PubMed] [Google Scholar]
  7. Janin J, Chothia C. 1990. The structure of protein-protein recognition sites. J Biol Chem 26.S:16027–16030. [PubMed] [Google Scholar]
  8. Jones S, Thornton JM. 1996. Principles of protein-protein interactions. Proc Natl Acad Sci USA, 93:13–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Keskin O, Bahar I, Badretdinov AY, Ptitsyn OB, Jernigan RL. 1998. Empirical solvent-mediated potentials hold for both intra-molecular and inter-molecular inter-residue interactions. Prot Sci 1998; 7: 2578–2586. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Keskin O , Gursoy A, Ma B, Nussinov R. 2008. Principles of Protein–Protein Interactions: What are the Preferred Ways For Proteins To Interact? Chem Rev 108:1225–1244. [DOI] [PubMed] [Google Scholar]
  11. Keskin O, Tuncbag N, Gursoy A. 2016. Predicting Protein–Protein Interactions from the Molecular to the Proteome Level, Chem Rev 116:4884–4909. [DOI] [PubMed] [Google Scholar]
  12. Kuntz ID, Blaney JM, Oatley SJ, Langridge R, Ferrin TE. 1982. A geometric approach to macromolecular-ligand interactions. J Mo/ Biol 161:269–288. [DOI] [PubMed] [Google Scholar]
  13. Lin T-L, Song G. 2011. Efficient mapping of ligand migration channel networks in dynamic protein. Proteins 79:2475–2490. [DOI] [PubMed] [Google Scholar]
  14. Manibog K, Sankar K, Kim SA, Zhang Y, Jernigan RL 2016. Sivasankar S. Molecular determinants of cadherin ideal bond formation: Conformation-dependent unbinding on a multidimensional landscape. Proc Natl Acad Sci U S A. 2016;113,E5711–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Miyazawa S, Jernigan RL. 1996. Residue-residue potentials with a favorable contact pair term and an unfavorable high packing density term for simulation and threading. J Mol Biol 256:623–644. [DOI] [PubMed] [Google Scholar]
  16. Nicholls A, Sharp KA, Honig 1991. Protein folding and association: In- sights from the interfacial and thermodynamic properties of hydrocarbons. Protein 8:281–296. [DOI] [PubMed] [Google Scholar]
  17. Oughtred R, Stark C, Breitkreutz BJ, Rust J, Boucher L, Chang C, Kolas N, O'Donnell L, Leung G, McAdam R, Zhang F, Dolma S, Willems A, Coulombe-Huntington J, Chatr-Aryamontri A, Dolinski K, Tyers M. 2019. The BioGRID interaction database: 2019 update. Nucleic Acids Res 47:D529–D541 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Pawar AP, duBay KF, Zurdo J, Chiti F,Vendruscolo M, Dobson CM. 2005. Prediction of “Aggregation-prone” and”Aggregation-susceptible” Regions in Proteins Associated with Neurodegenerative Diseases J Mol Biol 350: 379–392. [DOI] [PubMed] [Google Scholar]
  19. Pokarowski P, Kloczkowski A, Jernigan RL, Kothari NS, Pokarowska M, Kolinski A. Inferring ideal amino acid interaction forms from statistical protein contact potentials, Proteins, 2005, 59, 49–57 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Salwinski L; Miller CS; Smith AJ; Pettit FK; Bowie JU; Eisenberg D. 2004. "The Database of Interacting Proteins: 2004 update". Nucleic Acids Res. 32 :D449–D451. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Shoichet BK, Bodian DL, Kuntz ID. 1992. Molecular docking using shape descriptors. J Comput Chem /3(2):1–18. [Google Scholar]
  22. Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ, Mering CV. 2019. STRING v11: protein-protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Res 47:D607–D613. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Young L, Jernigan RL, Covell DG. 1994. A role for surface hydrophobicity in protein-protein recognition Prot Sci 1994;3:717–729 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Van Roy F (Ed.). The Molecular Biology of Cadherins, Progress in Molecular Biology and Translational Science Vol 116, (Academic Press, Oxford, UK: ). [DOI] [PubMed] [Google Scholar]

RESOURCES