Abstract
Vertebrate classical cadherins mediate selective calcium-dependent cell adhesion by mechanisms now understood at the atomic level. However, structures and adhesion mechanisms of cadherins from invertebrates, which are highly divergent yet function in similar roles, remain unknown. Here we present crystal structures of three- and four-tandem extracellular cadherin (EC) domain segments from Drosophila N-cadherin (DN-cadherin), each including the predicted N-terminal EC1 domain (denoted EC1’) of the mature protein. While the linker regions for the EC1’-EC2’ and EC3’-EC4’ pairs display binding of three Ca2+ ions similar to that of vertebrate cadherins, domains EC2’ and EC3’ are joined in a “kinked” orientation by a previously uncharacterized Ca2+-free linker. Biophysical analysis demonstrates that a construct containing the predicted N-terminal nine EC domains of DN-cadherin forms homodimers with affinity similar to vertebrate classical cadherins, whereas deleting the ninth EC domain ablates dimerization. These results suggest that, unlike their vertebrate counterparts, invertebrate cadherins may utilize multiple EC domains to form intercellular adhesive bonds. Sequence analysis reveals that similar Ca2+-free linkers are widely distributed in the ectodomains of both vertebrate and invertebrate cadherins.
Cell-cell adhesion is a distinguishing feature of metazoan species essential to the development and maintenance of solid tissues (1). In vertebrates, calcium-dependent cell adhesion is mediated primarily by members of the cadherin superfamily (2). Cadherins are defined as proteins containing “extracellular cadherin” (EC) domains (3–6), protein modules of ∼110 amino acids, which adopt a β-sandwich fold with a Greek key topology similar to that of immunoglobulin (Ig) domains. The best characterized cadherins are vertebrate classical cadherins, a family of proteins which share similar domain structures, each consisting of an ectodomain with five tandem EC domains, a single transmembrane region, and a conserved cytoplasmic tail (7). The connections between each set of successive EC domains are rigidified by the stereotyped binding of three Ca2+ ions (8, 9). Classical cadherins have been shown to function in intercellular adhesion by binding through their ectodomains, which are in turn linked to the actin cytoskeleton through associations of the cytoplasmic domains with catenin adaptor proteins (reviewed in ref. 10 and 11).
The cadherin superfamily is also broadly represented in invertebrates. Analysis of the Drosophila genome has revealed 17 genes that encode proteins containing EC-like domains (12). Three of these molecules, DN-cadherin encoded by CadN, DE-cadherin encoded by Shg, and DN-cad2 encoded by CadN2, which appears likely to be a partial duplication product of the CadN gene, contain catenin binding sites in their cytoplasmic regions, and have been shown to interact with the Drosophila β-catenin homolog armadillo (12–14). DN- and DE-cadherins serve cell adhesion and tissue patterning functions analogous to their vertebrate counterparts (13, 14). Also, like vertebrate classical cadherins, overexpression of DN-cadherin or DE-cadherin in otherwise nonadhesive cells induces Ca2+-dependent cell aggregation (13, 15). Although DN- and DE- cadherins perform biological roles roughly orthologous to those assumed by classical cadherins in vertebrate species, their ectodomains differ markedly from their vertebrate counterparts both in size and sequence features. Compared to vertebrate counterparts, DN- and DE-cadherins include a larger number of EC domains, some of which are highly diverged from vertebrate counterparts, and a membrane-proximal region consisting of EGF-like and laminin G domains. While the precise number of EC domains in DN-cadherin is not known with certainty because various prediction methods currently available identify different numbers ranging from 9 to 16, it is unclear how as many as 16 EC domains per cadherin could be arranged at intercellular junctions, because intercellular distances in both vertebrate and invertebrate tissues are similar (20–30 nm) (16, 17) and can be spanned by only five EC domains per molecule in vertebrate species.
Here we report crystal structures of DN-cadherin ectodomain regions corresponding to the predicted N-terminal four EC domains in the mature protein. While the linker regions between domains 1 and 2 and domains 3 and 4 display binding of three Ca2+ ions similar to that of vertebrate cadherins, domains 2 and 3 are joined in a “kinked” orientation by a Ca2+-free linker previously uncharacterized in cadherins. The orientations of domains 2 and 3 defined by this Ca2+-free interdomain linker are similar in all three crystal structures. The DN-cadherin fragments containing the predicted N-terminal four EC domains are monomeric both in crystals and in solution, whereas a larger construct that includes the N-terminal nine EC domains forms homodimers with a dissociation constant of ∼0.35 μM. These data, taken together, suggest that in contrast to vertebrate classical and T-cadherins, DN-cadherin and related cadherins function in intercellular adhesion through binding interfaces that are not localized to their distal N-termini. Rather, it is more likely that DN-cadherin and related cadherins form a globular structure with adhesive interfaces involving several EC domains, perhaps thematically similar to the arrangement of multiple Ig domains required for Dscam binding (18). Finally, based on the unique Ca2+-free interdomain linkage found in the DN-cadherin crystal structures, we present bioinformatic analyses of the entire cadherin superfamily that reveal the widespread presence of similar Ca2+-free linkers between successive EC domains in a large number of cadherins.
Results
DN-Cadherin EC Domains and N Terminus.
Prior to this work, the number of EC domains and domain boundaries for DN-cadherin have not been known with certainty, as sequence analyses have predicted a composition of 9 to 16 EC domains and different boundaries for each domain (Fig. S1). Moreover, the precise N terminus of the mature DN-cadherin has not been determined experimentally.
To determine the precise N terminus of the mature ectodomain as well as the EC domain boundaries of DN-cadherin, we first used the XPXF/W motif, a marker of the beginning of an EC domain (5), to divide the sequence of the extracellular region of DN-cadherin. Using this approach, we identified 19 segments of approximately 110 amino acids that can be considered as putative EC domains in the extracellular region of DN-cadherin. We next used PSI-BLAST to search for a set of DN-cadherin related proteins that include ectodomains of similar length with at least 60% sequence identity and a cytoplasmic region with a consensus β-catenin interaction motif. Within this set of DN-cadherin homologs, the first three EC domains have on average 35% sequence identity between corresponding EC domains, whereas for the following EC domains, the pair wise identity between corresponding domains is about 60% or higher (Fig. S2). Furin-like protease prediction algorithms implemented in Predict Protein (19) and Signal Pro P (20) place a prodomain cleavage site just C-terminal to the third putative EC domain in all of these related molecules. The observed pattern that the sequence identity is low (< 35%) N-terminal, and high (> 60%) C-terminal, to this predicted site is consistent with the patterns of sequence conservation found for vertebrate cadherins, and gives confidence that this site demarcates the connection between the prodomain and mature ectodomain in DN-cadherin. We thus designated DN-cadherin EC domains following the furin site EC1’ to EC16’, with the N-terminal boundary of EC1’ corresponding to the putative furin recognition site at position 434. The EC domains identified in the region preceding the cleavage site were assigned negative number identifiers (EC-3’, EC-2’, and EC-1’) to indicate their likely absence from the mature protein. The hypothesis that DN-cadherin proprotein is cleaved by a subtilisin-like convertase is supported by the work of Iwai and colleagues, which identified two distinct DN-cadherin species in the lysates of embryos and S2 cells transfected with a DN-cadherin encoding cDNA construct (14). The approximate molecular mass of the dominant species is 300 kD, consistent with the calculated mass of DN-cadherin regions C-terminal to the predicted furin cleavage site, while the second species is approximately 330 kD, consistent with the molecular mass of the uncleaved form (14).
Structures of DN-Cadherin Ectodomain Regions Reveal a Unique Architecture.
To gain insight into the structural basis for DN-cadherin adhesive functions, we sought to obtain structures of DN-cadherin ectodomain regions. We initially focused on ectodomain fragments containing the four predicted N-terminal EC domains in the mature protein, partly because prior structural and functional studies on vertebrate classical- and T-cadherins showed that cell adhesion is mediated through the N-terminal EC domains (8, 9, 21–23). We determined crystal structures of two DN-cadherin ectodomain fragments: one comprising domains EC1’-EC3’ in two crystal forms (I and II), both to 2.5 Å resolution, and the other comprising domains EC1’-EC4’ to 2.7 Å resolution. The structure of DN-cadherin EC1’-EC3’ in crystal form I was solved by the single-wavelength anomalous diffraction method using anomalous signal from zinc (Fig. S3A), and was used as the search model to determine the remaining two structures by molecular replacement. Data and refinement statistics are listed in Table S1.
The overall architecture of DN-cadherin EC1’-EC4’ fragment does not resemble any of the previously determined structures of vertebrate cadherin ectodomains, which adopt an elongated curved structure (8, 9) (Fig. 1). Instead, DN-cadherin EC1’-EC4’ fragment adopts a V-shaped structure imparted by a prominent “kink” between domains EC2’ and EC3’ with a ∼80° angle between the long axes of these two domains. In all three structures, each EC domain presents a seven-stranded β-sandwich fold with a Greek key topology seen in other structures of cadherin ectodomain fragments published to date (Fig. 1A). The linker regions between the EC1’ and EC2’ domains, and that between the EC3’ and EC4’ domain pairs contain three calcium ions bound in a way similar to that seen in classical cadherins (Fig. 1 A and B): Specifically, the three calcium ions are coordinated by a DXNDX (Asp-X-Asn-Asp-X) linker between the two successive EC domains, a DR/YE motif and a single glutamate (E) residue contributed from the prelinker domain, a DXD (Asp-X-Asp) motif and a single aspartate (D) residue contributed from the postlinker domain. The EC1’-EC2’ interdomain linkage in the EC1’-EC3’ structure determined in crystal form I is an exception in that the conserved calcium coordination patterns were disrupted by zinc ions present at an excessive concentration in the crystallization solution (Fig. S3A). As a result, the EC1’-EC2’ interdomain loop is rearranged such that EC1’ is rotated almost 180° along the long axis of EC2’ relative to the other two structures (Fig. S3B).
Fig. 1.
Structure of DN-cadherin EC1’-EC4’ ectodomain fragment. (A) The DN-cadherin EC1’-EC4’ ectodomain region adopts a V-shaped structure, in clear contrast to the overall elongated curved structure of mouse N-cadherins shown in (B). Calcium binding in the DN-cadherin EC1’-2’ and EC3’-4’ interdomain linker regions is very similar to that of classical cadherins as shown in the close-up views of the EC1’-2’ interdomain linker region in DN-cadherin and the EC1-2 interdomain linker region in mouse N-cadherin. Side chains or backbone atoms of residues involved in calcium coordination are shown as sticks, calcium ions are shown as green spheres. Dashed lines indicate coordinating interactions.
Notably, domains EC2’ and EC3’ are connected by a short Gly-Gly loop instead of a linker with bound Ca2+ ions, which joins most other successive EC domain pairs characterized to date. The EC2’-EC3’ interdomain interface facilitated by this unique Ca2+-free linker is very similar in all three structures (Fig. 2A), despite their presentation in distinct crystal lattices and large variations in crystallization conditions. Indeed, the kink between domains EC2’ and EC3’ is so rigid that the EC2’-EC3’ fragments from all four distinct chains (one in EC1’-EC3’ crystal form I, two in EC1’-EC3’ crystal form II, and one in EC1’-EC4’) can be superposed with pair wise root mean square deviation (rmsd) values of 1.1 Å or less for all 210 Cα atoms. This interdomain interface buries a surface area of 994 Å2, and is formed mainly by the A-strand, the Pro109-Leu110-Pro111 bulge immediately N-terminal to the A-strand, a portion of the G-strand from the EC2’ domain, and the loop region preceding the A-strand from the EC3’ domain (Fig. 2B). Interactions at this interface are mostly van der Waals contacts and nonsequence specific hydrogen bonds with main chain carbonyl oxygens, except for one salt bridge interaction between Arg209 from EC3’ and Ser 203 from EC2’ (Fig. 2B). Although the DN-cadherin EC2’-EC3’ interdomain interface, which is similar in all three crystal structures reported here, may represent a stable interface, we cannot exclude the possibility that the linker region could be flexible and the two domains may adopt different orientations under other conditions.
Fig. 2.
EC2’-3’ interface facilitated by a Ca2+-free interdomain linker. (A) Superposition of all three crystal structures: EC1’-EC3’ construct in crystal form I shown in yellow, EC1’-EC3’ construct in crystal form II shown in salmon, and EC1’-EC4’ construct shown in orange. (B) Stereoview of the EC2’-EC3’ interface observed in all three crystal structures. For clarity, the EC2’ domain is colored orange, and EC3’ domain is colored yellow.
DN-Cadherin Ectodomain Requires Multiple EC Domains for Adhesive Dimerization.
No distinct dimer interfaces were observed between molecules related by either crystallographic or noncrystallographic symmetry in the three crystal structures we report here. Consistent with the absence of sequence features required for strand-swap dimerization in vertebrate classical cadherins (5, 24), the DN-cadherin EC1’ domain does not engage in strand-swap dimerization with its N-terminal strand entirely integrated into the main body of its own protomer. In agreement with the crystallographic observations, our sedimentation equilibrium analytical ultracentrifugation (AUC) results show that both EC1’-EC3’ and EC1’-EC4’ fragments are monomers in solution, either in the presence or absence of calcium (Table 1).
Table 1.
Binding analysis of DN-cadherin ectodomain constructs by equilibrium AUC
| Constructs | Oligomeric states |
| EC1’-3’ | monomer |
| EC1’-4’ | monomer |
| EC1’-8’ | monomer |
| EC1’-9’ | dimer* |
| EC1’-10’ | dimer† |
| EC2’-9’ | aggregates |
Dimerizaton affinites obtained from three independent measurements (μM):
*0.36 ± 0.03; 0.14 ± 0.03; 0.028 ± 0.008
†5.4 ± 0.2; 4.1 ± 0.1; 0.14 ± 0.02
To determine the ectodomain regions responsible for adhesive functions of DN-cadherin, we produced a series of domain deletion constructs and carried out binding analysis using sedimentation equilibrium AUC. Our AUC measurements show that an ectodomain construct comprising the predicted EC1’-EC10’ domains forms a tight dimer in the presence of Ca2+ with low micromolar dimerization affinity (Table 1). A shorter construct containing EC1’-EC9’ domains also forms a tight dimer. However, deleting the EC9’ domain leads to a monomer (Table 1), in agreement with results from cell aggregation studies reported by Yonekura, et al., which demonstrated that the highly related DN-cad2, which aligns with the C-terminal region of DN-cadherin, but contains nine fewer EC domains at the N terminus, cannot induce cell aggregation in vitro (25). On the other hand, a DN-cadherin construct containing EC2’-EC9’ domains appears to form soluble aggregates in AUC experiments.
Taken together, these results suggest that DN-cadherin requires EC1’-EC9’ domains for adhesive dimerization, in clear contrast to vertebrate classical cadherins, which trans-dimerize through an interface entirely confined to the EC1 domain, or T-cadherin, which uses both EC1 and EC2 domains for adhesive dimerization. Consistently, the sequence determinants of the strand-swapped interface of vertebrate classical cadherins and the X-interface of vertebrate T-cadherin are absent in the DN-cadherin EC1’ and EC2’ domains.
Sequence Analysis Reveals the Widespread Presence of Ca2+-Free Linkers in the Cadherin Superfamily.
It is apparent from the sequence alignment of all DN-cadherin EC domains (Fig. 3) that the Ca2+-free linker between EC2’ and EC3’ is correlated with the absence of most of the calcium-binding motifs. Specifically, the interdomain DXNDX motif, the E and DR/YE motifs in EC2’, and the DXD motif in EC3’ are all missing. Notably, the same combination of Ca2+-binding motifs is missing in all of the DN-cadherin homologs identified in other arthropods (26) (Fig. S4), suggesting the functional importance of these missing motifs.
Fig. 3.
Structure based sequence alignment of type I and type II classical cadherin domains with DN-cadherin domains. Secondary structure elements are indicated above the alignment. Red boxes above the alignment denote the five canonical Ca2+-binding motifs or residues, and the red arrow on top of each box indicates which of the two interdomain Ca2+-binding sites the motif belongs to. Conserved Ca2+-binding residues are highlighted in red; D/E to N substitutions are highlighted in yellow; the black boxes indicate the missing Ca2+-binding elements. The highly conserved XPXF/W motif that marks the beginning of each EC domain is highlighted in blue. Conserved hydrophobic residues that constitute the core of the EC domains are represented in blue.
To determine whether the Ca2+-free linker found in DN-cadherin is common to other cadherins, we searched the entire cadherin superfamily for similar instances. For each of the 20,310 pairs of consecutive EC domains we determined whether the combination of Ca2+-binding motifs corresponding to the interdomain linker is either completely or partially present or missing (see Table S2 for the contribution from the different Ca2+-binding motifs to each Ca2+-binding site). Among the 20,310 interdomain linkers, we found 2,504 linkers distributed in 936 proteins with different combinations of missing Ca2+-binding residues. The analysis of the domains directly preceding and following these 2,504 linkers did not reveal any sequence commonalities other than the missing Ca2+-binding motifs. In particular, the residues that constitute the EC2’-EC3’ interface of DN-cadherin are not conserved in Ca2+-free linkers of different cadherins. Furthermore, because no residues with oxygen- or nitrogen-containing side chains that would normally coordinate metal ions are conserved in position for the Ca2+-free linkers, it seems unlikely that these linkers bind an ion other than calcium under physiological conditions.
In contrast to the DN-cadherin EC2’-EC3’ linker where almost all of the Ca2+-binding motifs are absent, we found many instances where only some of the motifs are missing. Among the 2,504 linkers with missing Ca2+-binding motifs, we found 366 cases where the missing Ca2+-binding residues mostly correspond to one Ca2+ site and the remaining residues may be sufficient to coordinate the two other Ca2+ ions. Whether this sequence based definition of partial Ca2+-free linkers indeed corresponds to distinct linker structures where only one or two Ca2+ ions are bound, however, remains to be determined. Of note, in addition to the EC2’-EC3’ Ca2+-free linker revealed by our crystal structures, DN-cadherin contains a second linker with missing Ca2+-binding sites between EC7’ and EC8’. Because the missing motifs in this linker mostly correspond to residues that coordinate one of the three Ca2+ ions (Fig. 3), we categorized it as a partial Ca2+-free linker. Importantly, the same combination of missing motifs is also conserved in the EC7’-EC8’ regions of the related arthropod N-cadherins (Fig. S4). However, as observed in the domains flanking Ca2+-free linkers, no other sequence feature, apart from the missing Ca2+-binding residues, appears to distinguish these domains from regular domains.
DE-cadherin and DN-cad2 are closely related to DN-cadherin, and contain eight and seven EC domains in the mature proteins, respectively. In fact, the sequence identity between EC domains of DN-cadherin and those of either DE-cadherin or DN-cad2 and their homologs in other arthropods suggests that DE-cadherin and DN-cad2 are evolutionarily related to different parts of the DN-cadherin ectodomain: the first six EC domains of DE-cadherin correspond to the EC6’-EC11’ fragment in DN-cadherin, and the seven EC domains of DN-cad2 correspond to the EC10’-EC16’ fragment of DN-cadherin (Fig. 4). Because the CadN2 gene is likely the product of a partial gene duplication of the CadN gene (12), the homology between DN-cad2 and DN-cadherin EC10’-EC16’ correlates with that between the gene structures. On the other hand, the gene structure of Shotgun is not homologous to that of CadN, and DE-cadherin may have evolved from a different origin. Our analysis shows that while DE-cadherin, which dimerizes homophilically (13), contains a Ca2+-free linker between EC2 and EC3, DN-cad2, which does not aggregate cells (25), has none. Remarkably, the pattern of Ca2+-free linkers in the evolutionarily related fragments is well conserved: the putative partial Ca2+-free linker between EC7’ and EC8’ of DN-cadherin corresponds to the Ca2+-free linker between EC2 and EC3 of DE-cadherin; DN-cad2 preserves all of the Ca2+-binding linkers as does the DN-cadherin EC10’-EC16’ fragment (Fig. 4).
Fig. 4.
Relationship between DN-cadherin, DE-cadherin, and DN-cad2. (A) Matrices of pair wise sequence identity between EC domains of DN-cadherin and DE-cadherin (top) and between EC domains of DN-cadherins and DN-cad2 (bottom). The arrows indicate the presence of Ca2+-free linkers as revealed by the DN-cadherin crystal structure reported here (in green) or predicted from the sequence (in blue). The homologous sequences used in the analysis are indicated in the inset. (B) Mapping of correspondence between the three proteins, with the known regions required for binding functions indicated.
The mapping of the Ca2+-free and partial Ca2+-free linkers for the major cadherin subfamilies is presented in Fig. 5. While we did not find any missing Ca2+-binding sites in the shorter members of the cadherin superfamily such as type I and type II classical cadherins, desmocollins, desmogleins, and protocadherins, we found many instances of Ca2+-free and partial Ca2+-free linkers in nonclassical cadherins containing a large number of EC domains, including DN- and DE-related cadherins, Fat, Dachsous, and Flamingo/CELSR cadherins. Furthermore, we observed that within a given subfamily, the distribution and pattern of the linker types are generally conserved across species. Remarkably, these Ca2+-free linkers are also present in the majority of the choanoflagellate Monosiga brevicollis (M. brevicollis) cadherins (27) (Table S3). Of note, the largest M. brevicollis cadherin, MBCDH18 (27) with 58 EC domains, contains as many as 19 Ca2+-free linkers (Fig. 5). On the whole superfamily scale, the larger the cadherins, the more Ca2+-free and partial Ca2+-free linkers they contain (Fig. 6A). Ca2+-free and partial Ca2+-free linkers mostly appear in cadherins containing more than six EC domains, and we observe three or more of these new linkers only in cadherins with 15 EC domains or more. As shown in Fig. 6B for the whole cadherin superfamily, the majority (93%) of consecutive segments uninterrupted by a Ca2+-free or a partial Ca2+-free linker contain six EC domains or less.
Fig. 5.
A broad search for Ca2+-free and partial Ca2+-free linkers in the major cadherin subfamilies. The mapping of Ca2+-free and partial linkers in the major subfamilies is given. The domain numbering is based on the mature proteins. Red arrows indicate Ca2+-free linkers in the extracellular domains, and gray arrows indicate partial Ca2+-free linkers. For DN-cadherin, only the EC domains in the putative mature protein are shown. In Fats 1–3 and Fat-like, the striped domains labeled ΔEC correspond to the absence of an EC domain that is found in Fat-4 and Fat-like.
Fig. 6.
Length distribution of ectodomains and ectodomain segments. (A) The number of proteins is given by the number of EC domains in the whole ectodomain. The colors of each bar indicate the number of Ca2+-free linkers found in the protein. Segments containing eight EC domains and more are shown magnified. (B) The number of segments that are connected by full Ca2+-binding sites is given by the number of EC domains in the segment.
Discussion
Although the molecular mechanisms of vertebrate classical cadherin mediated cell adhesion have been understood at the atomic level for some time (28), little is known about the structures and adhesion mechanisms of the vast majority of cadherins including invertebrate cadherins. The structural and biophysical results presented here provide important understanding of the ectodomain architecture of DN-cadherin and suggests a unique adhesive mechanism. Our bioinformatic analysis shows that Ca2+-free linkers similar to the one between domains EC2’ and EC3’ of DN-cadherin revealed by our crystal structures are likely to be present in a large number of nonclassical cadherins and may play important structural roles in adhesive functions of these cadherins.
DN-Cadherin Ectodomain Architecture Suggests a Unique Adhesive Mechanism.
The structures of DN-cadherin ectodomain fragments we report here reveal a V-shaped architecture, in clear contrast to the elongated curved structure of vertebrate classical cadherins (8, 9). The “V” shape of DN-cadherin EC1’-EC4’ is imparted by a kink between domains EC2’ and EC3’ that are connected by a unique Ca2+-free linker, and is thematically similar to the U-shaped structure of Ig1-Ig4 fragments from several members of the immunoglobulin cell adhesion molecule (IgCAM) family, including hemolin, axonin, Dscam, and neurofascin (18, 29–31). In these IgCAMs, the first four Ig domains form a U-shaped structure with significant intramolecular contacts between Ig1 and Ig4 domains that bury ∼1,400 Å2 of surface area. By contrast, DN-cadherin forms a V-shaped structure, as the Ca2+-bound EC1’-EC2’ and EC3’-EC4’ interdomain regions orient domains EC1’ and EC4’ such that they do not make any intramolecular interdomain contacts. In vertebrate classical cadherins and T-cadherin, calcium-binding interdomain linkers are critical structural elements that determine interdomain orientation and facilitate strand swapping and X-dimer formation (8, 9, 21). The unique Ca2+-free linker seen between the EC2’ and EC3’ domains in all three structures presented here provides new clues about the ectodomain architecture of DN-cadherin and other related molecules. The similarity of the EC2’-EC3’ interdomain interface facilitated by this Ca2+-free linker in all three structures suggests that it may be a stable structural feature required for the proper tertiary folding and adhesive functions of DN-cadherin.
The results from our AUC measurements suggest that the DN-cadherin EC1'-EC9' region includes the adhesive interface, in agreement with results from a previous study, which demonstrated that DN-cad2, a highly related but truncated protein with nine fewer EC domains at the N terminus, is not required for R7 target selection and cannot induce cell aggregation in vitro (25). The observation that a larger ectodomain region with a greater number of EC domains is required for DN-cadherin dimerization than is evident for vertebrate classical cadherins, along with our structural data, suggest that the adhesive mechanism utilized by DN-cadherin is likely to be completely different from that of vertebrate classical cadherins, which bridge the intercellular space by an elongated arrangement of EC repeats (8, 9). Our structures reveal and AUC measurements support that the EC1’ domain of DN-cadherin does not engage in strand-swap dimerization, however, we cannot exclude the possibility that domains EC8’ and EC9’ might use an interface similar to the X-dimer adhesive interface of vertebrate T-cadherin (21).
Our sequence analysis reveals that the apparently adhesive EC1’-EC9’ fragment of DN-cadherin also contains a second putative partial Ca2+-free linker between domains EC7’ and EC8’, and may therefore mediate intercellular adhesion through a more globular structure reminiscent of Dscams. Dscams, which require domains Ig1-Ig7 for homophilic binding, contain two interdomain “bends”: one between domains Ig2 and Ig3, and the other between domains Ig5 and Ig6 (18). These two bends enable Dscams to adopt an S-shaped structure, which presents a fixed binding surface comprised of three variable domains, Ig2, Ig3, and Ig7. This conformation is thought to allow each pair of the three variable domains to match in an antiparallel fashion to confer homophilic binding specificity to each of the 19,008 possible Dscam isoforms with different ectodomains. Whether DN-cadherin also uses a similar surface comprised of multiple EC domains for adhesive interactions remains to be investigated. Interestingly, the CadN gene contains three mutually exclusive exons (MEs) to produce 12 possible alternatively spliced variants (15, 32). Alternative splicing at one of these three MEs generates two isoforms with different amino acid compositions in a region encompassing the EC7’-EC8’ interdomain linker (32). We note that this variable region corresponds to a region within the EC2’-EC3’ segment that constitutes the EC2’-EC3’ interface observed in our crystal structures (Figs. 2 and 3). An intriguing possibility arises that, by analogy to the EC2’-EC3’ interface, the predicted partial Ca2+-free linker between domains EC7’ and EC8’ generates different EC7’-EC8’ interfaces in different splice isoforms. Thus, each isoform can adopt distinct ectodomain architecture and exhibit differential homophilic binding properties. Understanding of the detailed molecular basis for DN-cadherin homophilic binding, however, must await structural and functional studies of larger regions of ectodomains in different isoforms.
As described above, our sequence analysis shows that the N-terminal six EC domains of DE-cadherin, corresponding to its minimal binding fragment in cell aggregation assays (33), are evolutionarily related to a region that overlaps with the binding fragment of DN-cadherin and conserves the partial Ca2+-free linker present in that fragment (Fig. 4). While it seems unlikely that DN- and DE-cadherins dimerize through the same mechanism given their different sizes and numbers of Ca2+-free linkers, our study suggests that Ca2+-free linkers may play an important and similar role in adhesive binding by these cadherins. Future structural and functional studies will be required to elucidate the detailed molecular mechanisms for homophilic binding by these cadherins.
Ca2+-Free Linkers may Represent Functionally Important Structural Elements in a Large Number of Cadherin Ectodomains.
To date, cadherin ectodomains have mostly been perceived as extended, rigid structures, as they appear to be in vertebrate classical cadherins. The present study shows that Ca2+-free linkers are present in many cadherin families such as Fat, Dachsous, Flamingo, DN- and DE-cadherins, and can impart complex ectodomain architectures to these proteins. While it is not clear whether the Ca2+-free and partial Ca2+-free linkers correspond to stable bends or points of higher flexibility in all of these cases, they are likely to interrupt the rigid linear arrangements of cadherin ectodomains and enable them to fold into more globular structures. The specific combination and placement of these Ca2+-free and partial Ca2+-free linkers, highly conserved across species, could allow these large cadherins to fold into a specific, more globular shape. Moreover, because there appears to be no conserved feature other than the absence of Ca2+-binding motifs among the domains flanking the Ca2+-free linkers, it is possible that different Ca2+-free linkers might adopt different conformations suited for the assumed functions of different cadherins. That the vast majority (more than 90%) of the stretches of consecutive EC domains connected by full Ca2+-binding linkers contain six or fewer EC domains (Fig. 6) suggests that nonclassical cadherins containing Ca2+-free linkers have a molecular architecture involving extended structural regions seen in vertebrate classical cadherins that may fold back on each other in still undetermined ways. Notably, among the few exceptions (less than 1% of the stretches containing 10 EC domains or more), we found cadherin-23, whose 27 EC domains remain “uninterrupted” by missing or partial Ca2+-binding sites. In good agreement with our hypothesis, this cadherin does not function in a regular intercellular space, but instead associates with protocadherin-15 (11 EC domains) to form the tip link of the inner ear’s stereocilia (34), where a large distance (150 to 200 nm) must be spanned by extended conformations of both cadherins.
Of note, the choanoflagellate M. brevicollis genome contains 23 cadherin genes (27). Among the 19 M. brevicollis cadherins that contain two or more EC domains, 13 are predicted to contain Ca2+-free linkers (Table S3). The widespread presence of Ca2+-free linkers in a premetazoan species suggests that the earliest function of cadherins, for example, potentially in binding bacterial prey for recognition or capture, may have required structural complexities imparted by these Ca2+-free interdomain linkages.
Future studies will be required to characterize the specific structural and functional roles of Ca2+-free linkers in cadherins, and to link the patterns of different linker types to the specific functions of individual cadherins.
Materials and Methods
Expression and Purification of Recombinant DN-Cadherin Proteins.
The EC1’-EC3’ construct (residues 439–753) and EC1’-EC4’ construct (residues 434–851) were subcloned into the pSMT3 vector (Invitrogen) to produce the N-terminal His-tagged SUMO fusion. Overexpression in E. coli was induced with IPTG. Cells were sonicated in lysis buffer (20 mM Tris pH 8.0, 150 mM NaCl, 4 mM CaCl2, 100 ng/mL Leupeptin, and 1 mM PMSF) followed by centrifugation at 25,000 × g for 1.5 h at 4 °C. The resulting supernatant was then applied to Ni2+-charged agarose resin (Qiagen), and washed extensively with lysis buffer. Subsequently, ULP1 was applied to the protein-bound resin to a concentration of 0.1 mg/mL. The mixture was then incubated at 4 °C overnight with stirring. DN-cadherin proteins were then washed off the SUMO-bound resin in small volumes of lysis buffer and then dialyzed into a low salt (50 mM NaCl) buffer and passed over a MonoQ10/100 column (GE Healthcare). The flow-through was pooled, concentrated, and applied to a Superdex 75 column (GE Healthcare).
DN-cadherin EC1’-EC10’, EC1’-EC9’, EC1’-EC8’, and EC2’-EC9’ proteins were heterologously expressed in N-acetylglucosaminyltransferase I (GnTI)-deficient HEK293 cell lines as secreted proteins with a PTPα signal sequence and N-terminal hexahistidine tag. Conditioned media were harvested two days after transient transfection using Polyethylenimine buffered to pH 8.0 with Tris, and brought to high salt concentration before application to sepharose resin (GE Healthcare) charged with nickel sulfate. The histidine-tagged protein was eluted with imidazole in the following buffer: 10 mM Tris pH 8.0, 500 mM NaCl, 4 mM CaCl2. Digestion with Precission protease (GE Healthcare) was then followed by further purification on a MonoQ 10/100 global column (GE Healthcare) and Superdex 200 column (GE Healthcare).
Crystallization, Data Collection, and Structure Determination.
Initial factorial-based crystallization screens were conducted using a Mosquito robotic crystallization system (TTPlabtech) using 0.2 μL drop volumes to screen numerous commercially prepared crystallization reagents. DN-cadherin EC1’-EC3’ crystallized in two major forms. The first crystal form was optimized in 4% isopropanol, 0.1M zinc acetate, 0.1M sodium cacodylate, pH 6.5. These crystals belong to space group P6222 with unit cell dimensions of a = b = 96.3 Å, c = 148.2 Å, and has one molecule in the asymmetric unit. Crystals in the second crystal form grew in 34% PEG 3350, 0.1M lithium sulfate, 0.1M Tris, pH 8.5, and are in space group C2 with unit cell dimensions of a = 106.0 Å, b = 113.8 Å, c = 87.8 Å, β = 123.7° and two molecules in the asymmetric unit. A single crystal form was obtained for DN-cadherin EC1’-EC4’ in 9% PEG 3350, 0.05M L-proline, 0.1M Hepes, pH 7.5. These crystals are in space group P21212 with unit cell dimension of a = 81.5 Å, b = 126.7 Å, c = 62.0 Å and contain one molecule in the asymmetric unit. All crystals were cryoprotected in mother liquor containing 35% glycerol and flash-frozen in liquid nitrogen for diffraction data collection.
All X-ray diffraction data were collected on single crystals at 100 K at the X4A and X4C beamlines of the National Synchrotron Light Source, Brookhaven National Laboratory. Data were processed using the HKL package (35). The structure of a DN-cadherin EC1’-EC3’ construct in crystal form I (space group P6222) was determined by single-wavelength anomalous diffraction method. Experimental phases were obtained using zinc sites located with the program SOLVE (36). Initial experimental phases were improved by solvent flattening using RESOLVE (36). The resultant density-modified experimental map was used to manually build the model using COOT (37), and iterative refinement was carried out using CNS (38). The structures of EC1’-EC3’ in crystal form II (space group C2) and EC1’-EC4’ were solved by molecular replacement with PHASER (39) using the structure of EC1’-EC3’ in crystal form I as the search model. Manual rebuilding was done with COOT (37), and refinement was performed using REFMAC (40) implemented in the CCP4 program suite (41) and CNS (38). The statistics of data collection and refinement are summarized in Table S1. All molecular graphics figures were generated with the program Pymol (DeLano Scientific, LLC). Coordinates will be deposited in the Protein Data Bank prior to publication.
Analytical Ultracentrifugation.
Sedimentation equilibrium measurements were performed using a Beckman XL-A/I analytical ultracentrifuge (Beckman-Coulter), equipped with 12mm six-cell centerpieces with sapphire or quartz windows. The proteins were dialyzed in 10 mM Tris, pH 8.0, 150 mM NaCl, 3 mM CaCl2 at 4 °C. All proteins were run at 25 °C for 20 h at 7,000 rpm after which four scans at one-hour intervals were collected. Speed was then increased to 9,000 rpm for 10 h, then to 11,000 rpm for 10 h, and lastly to 13,000 rpm for 10 h, with four hourly scans taken after each period., Each protein was simultaneously run at three different concentrations, 0.70, 0.49, and 0.26 mg/mL for detection using UV 280 nm and interference at 660 nm, and 0.13, 0.092, and 0.049 mg/mL for UV detection at 230 nm. Solvent density and protein v-bar were calculated using the program SednTerp (Alliance Protein Laboratories). For calculation of KDs and apparent molecular weights, all data were used in a global fit, using the program HeteroAnalysis version 1.1.44, obtained from University of Connecticut (www.biotech.uconn.edu/auf).
Sequence Analysis.
All cadherin domain hits from the SMART dataset (42) were aligned one by one to a reference alignment of cadherin domains using Muscle (43). The calcium-binding motifs were found by searching for each motif in the region that aligned to that motif in the reference alignment. Specifically, we searched for DXD, DXE, and DXNDX at the appropriate positions as indicated by the reference alignment. Each of the three motifs was scored as being present, absent, or partial [if only one of the two D/E (Asp/Glu) residues was there], allowing for D/E substitutions. Consecutive pairs of EC domains were taken to determine whether each individual motif from the two domains contributing to the interdomain Ca2+-binding site were present, absent, or partially present. The linker was indicated as a full calcium-binding linker if all three motifs were fully present, or if two were fully and one partially present. The linker was marked as a partial calcium-free linker if two motifs were fully present and one was missing. In all other cases, if only one motif or no motifs were fully present, the linker was marked a calcium-free linker.
Supplementary Material
Acknowledgments.
We acknowledge Dr. Chi-Hon Lee at the National Institute of Health for kindly providing us with the DN-cadherin cDNA. X-ray data were collected at the X4A and X4C beamlines of the National Synchrotron Light Source, Brookhaven National Laboratory; the beamlines are operated by the New York Structural Biology Center. This work was supported by grants R01 GM062270-07 (to L.S.) from the National Institutes of Health and MCB-0918535 (to B.H.) from the National Science Foundation. M.A.W. was supported by a fellowship F30 NS061400 and a training Grant T32 EY13933-07 from the National Institute of Health. K.F. was supported by a training Grant T32 GM082797 from the National Institute of Health.
Footnotes
The authors declare no conflict of interest.
See Author Summary on page 659.
This article contains supporting information online at www.pnas.org/lookup/suppl/doi:10.1073/pnas.1117538108/-/DCSupplemental.
Data deposition: The atomic coordinates have been deposited in the Protein Data Bank, www.pdb.org [PDB ID codes 3UBF (EC1’-EC3’ crystal form I), 3UBG (EC1’-EC3’ crystal form II), and 3UBH (EC1’-4’)].
References
- 1.Gumbiner BM. Regulation of cadherin-mediated adhesion in morphogenesis. Nat Rev Mol Cell Biol. 2005;6:622–634. doi: 10.1038/nrm1699. [DOI] [PubMed] [Google Scholar]
- 2.Takeichi M. Morphogenetic roles of classic cadherins. Curr Opin Cell Biol. 1995;7:619–627. doi: 10.1016/0955-0674(95)80102-2. [DOI] [PubMed] [Google Scholar]
- 3.Nollet F, Kools P, van Roy F. Phylogenetic analysis of the cadherin superfamily allows identification of six major subfamilies besides several solitary members. J Mol Biol. 2000;299:551–572. doi: 10.1006/jmbi.2000.3777. [DOI] [PubMed] [Google Scholar]
- 4.Overduin M, et al. Solution structure of the epithelial cadherin domain responsible for selective cell adhesion. Science. 1995;267:386–389. doi: 10.1126/science.7824937. [DOI] [PubMed] [Google Scholar]
- 5.Posy S, Shapiro L, Honig B. Sequence and structural determinants of strand swapping in cadherin domains: do all cadherins bind through the same adhesive interface? J Mol Biol. 2008;378:954–968. doi: 10.1016/j.jmb.2008.02.063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Shapiro L, et al. Structural basis of cell-cell adhesion by cadherins. Nature. 1995;374:327–337. doi: 10.1038/374327a0. [DOI] [PubMed] [Google Scholar]
- 7.Takeichi M, et al. Identification of a gene family of cadherin cell adhesion molecules. Cell Differ Dev. 1988;25(Suppl):91–94. doi: 10.1016/0922-3371(88)90104-9. [DOI] [PubMed] [Google Scholar]
- 8.Boggon TJ, et al. C-cadherin ectodomain structure and implications for cell adhesion mechanisms. Science. 2002;296:1308–1313. doi: 10.1126/science.1071559. [DOI] [PubMed] [Google Scholar]
- 9.Harrison OJ, et al. The extracellular architecture of adherens junctions revealed by crystal structures of type I cadherins. Structure. 2011;19:244–256. doi: 10.1016/j.str.2010.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Takeichi M. Cadherins: a molecular family important in selective cell-cell adhesion. Annu Rev Biochem. 1990;59:237–252. doi: 10.1146/annurev.bi.59.070190.001321. [DOI] [PubMed] [Google Scholar]
- 11.Nelson WJ. Regulation of cell-cell adhesion by the cadherin-catenin complex. Biochem Soc Trans. 2008;36:149–155. doi: 10.1042/BST0360149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hill E, et al. Cadherin superfamily proteins in Caenorhabditis elegans and Drosophila melanogaster. J Mol Biol. 2001;305:1011–1024. doi: 10.1006/jmbi.2000.4361. [DOI] [PubMed] [Google Scholar]
- 13.Oda H, et al. A Drosophila homolog of cadherin associated with armadillo and essential for embryonic cell-cell adhesion. Dev Biol. 1994;165:716–726. doi: 10.1006/dbio.1994.1287. [DOI] [PubMed] [Google Scholar]
- 14.Iwai Y, et al. Axon patterning requires DN-cadherin, a novel neuronal adhesion receptor, in the Drosophila embryonic CNS. Neuron. 1997;19:77–89. doi: 10.1016/s0896-6273(00)80349-9. [DOI] [PubMed] [Google Scholar]
- 15.Yonekura S, et al. The variable transmembrane domain of Drosophila N-cadherin regulates adhesive activity. Mol Cell Biol. 2006;26:6598–6608. doi: 10.1128/MCB.00241-06. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Farquhar MG, Palade GE. Junctional complexes in various epithelia. J Cell Biol. 1963;17:375–412. doi: 10.1083/jcb.17.2.375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Ho KL. Intercellular septate-like junction of neoplastic cells in myxopapillary ependymoma of the filum terminale. Acta Neuropathol. 1990;79:432–437. doi: 10.1007/BF00308720. [DOI] [PubMed] [Google Scholar]
- 18.Sawaya MR, et al. A double S shape provides the structural basis for the extraordinary binding specificity of Dscam isoforms. Cell. 2008;134:1007–1018. doi: 10.1016/j.cell.2008.07.042. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Rost B, Yachdav G, Liu J. The PredictProtein server. Nucleic Acids Res. 2004;32(Web Server issue):W321–W326. doi: 10.1093/nar/gkh377. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Duckert P, Brunak S, Blom N. Prediction of proprotein convertase cleavage sites. Protein Eng Des Sel. 2004;17:107–112. doi: 10.1093/protein/gzh013. [DOI] [PubMed] [Google Scholar]
- 21.Ciatto C, et al. T-cadherin structures reveal a novel adhesive binding mechanism. Nat Struct Mol Biol. 2010;17:339–347. doi: 10.1038/nsmb.1781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Shan WS, et al. Functional cis-heterodimers of N- and R-cadherins. J Cell Biol. 2000;148:579–590. doi: 10.1083/jcb.148.3.579. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Troyanovsky RB, Sokolov E, Troyanovsky SM. Adhesive and lateral E-cadherin dimers are mediated by the same interface. Mol Cell Biol. 2003;23:7965–7972. doi: 10.1128/MCB.23.22.7965-7972.2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Vendome J, et al. Molecular design principles underlying beta-strand swapping in the adhesive dimerization of cadherins. Nat Struct Mol Biol. 2011;18:693–700. doi: 10.1038/nsmb.2051. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Yonekura S, et al. Adhesive but not signaling activity of Drosophila N-cadherin is essential for target selection of photoreceptor afferents. Dev Biol. 2007;304:759–770. doi: 10.1016/j.ydbio.2007.01.030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Oda H, Tagawa K, Akiyama-Oda Y. Diversification of epithelial adherens junctions with independent reductive changes in cadherin form: identification of potential molecular synapomorphies among bilaterians. Evol Dev. 2005;7:376–389. doi: 10.1111/j.1525-142X.2005.05043.x. [DOI] [PubMed] [Google Scholar]
- 27.Abedin M, King N. The premetazoan ancestry of cadherins. Science. 2008;319:946–948. doi: 10.1126/science.1151084. [DOI] [PubMed] [Google Scholar]
- 28.Shapiro L, Weis WI. Structure and biochemistry of cadherins and catenins. Cold Spring Harbor Perspectives in Biology. 2009;1:a003053. doi: 10.1101/cshperspect.a003053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Freigang J, et al. The crystal structure of the ligand binding module of axonin-1/TAG-1 suggests a zipper mechanism for neural cell adhesion. Cell. 2000;101:425–433. doi: 10.1016/s0092-8674(00)80852-1. [DOI] [PubMed] [Google Scholar]
- 30.Liu H, Focia PJ, He X. Homophilic adhesion mechanism of neurofascin, a member of the L1 family of neural cell adhesion molecules. J Biol Chem. 2011;286:797–805. doi: 10.1074/jbc.M110.180281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Su XD, et al. Crystal structure of hemolin: a horseshoe shape with implications for homophilic adhesion. Science. 1998;281:991–995. doi: 10.1126/science.281.5379.991. [DOI] [PubMed] [Google Scholar]
- 32.Hsu SN, et al. Conserved alternative splicing and expression patterns of arthropod N-cadherin. PLoS Genet. 2009;5:e1000441. doi: 10.1371/journal.pgen.1000441. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Haruta T, et al. The proximal half of the Drosophila E-cadherin extracellular region is dispensable for many cadherin-dependent events but required for ventral furrow formation. Genes Cells. 2010;15:193–208. doi: 10.1111/j.1365-2443.2010.01389.x. [DOI] [PubMed] [Google Scholar]
- 34.Kazmierczak P, et al. Cadherin 23 and protocadherin 15 interact to form tip-link filaments in sensory hair cells. Nature. 2007;449:87–91. doi: 10.1038/nature06091. [DOI] [PubMed] [Google Scholar]
- 35.Otwinowski Z, Minor W. Processing of X-ray diffraction data collected in oscillation mode. Method Enzymol. 1997;276:307–326. doi: 10.1016/S0076-6879(97)76066-X. [DOI] [PubMed] [Google Scholar]
- 36.Terwilliger TC. SOLVE and RESOLVE: automated structure solution and density modification. Methods Enzymol. 2003;374:22–37. doi: 10.1016/S0076-6879(03)74002-6. [DOI] [PubMed] [Google Scholar]
- 37.Emsley P, et al. Features and development of COOT. Acta Crystallogr D. 66:486–501. doi: 10.1107/S0907444910007493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Brunger AT. Version 1.2 of the Crystallography and NMR system. Nat Protoc. 2007;2:2728–2733. doi: 10.1038/nprot.2007.406. [DOI] [PubMed] [Google Scholar]
- 39.McCoy AJ, et al. Phaser crystallographic software. J Appl Crystallogr. 2007;40:658–674. doi: 10.1107/S0021889807021206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Murshudov GN, Vagin AA, Dodson EJ. Refinement of macromolecular structures by the maximum-likelihood method. Acta Crystallogr D. 1997;53:240–255. doi: 10.1107/S0907444996012255. [DOI] [PubMed] [Google Scholar]
- 41.Collaborative Computational Project, N. The CCP4 suite: programs for protein crystallography. Acta Crystallogr D. 1994;50:760–763. doi: 10.1107/S0907444994003112. [DOI] [PubMed] [Google Scholar]
- 42. Letunic I, Doerks T, Bork P. SMART 6: recent updates and new developments. Nucleic Acids Res. 2009;37:D229–D232. doi: 10.1093/nar/gkn808. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004;32:1792–1797. doi: 10.1093/nar/gkh340. [DOI] [PMC free article] [PubMed] [Google Scholar]







