Abstract
The structures of the N-terminal domains of two integrases of closely related but not identical asn tDNA-associated genomic islands, Yersinia HPI (high pathogenicity island; encoding siderophore yersiniabactin biosynthesis and transport) and an Erwinia carotovora genomic island with yet unknown function, HAI7, have been resolved. Both integrases utilize a novel four-stranded β-sheet DNA-binding motif, in contrast to the known proteins that bind their DNA targets by means of three-stranded β-sheets. Moreover, the β-sheets in IntHPI and IntHAI7 are longer than those in other integrases, and the structured helical N terminus is positioned perpendicularly to the large C-terminal helix. These differences strongly support the proposal that the integrases of the genomic islands make up a distinct evolutionary branch of the site-specific recombinases that utilize a unique DNA-binding mechanism.
Introduction
Genomic islands, together with temperate phages, integrative plasmids, transposons, and integrative conjugative elements, make up the group of mobile genetic elements that play an important role in bacterial quantum leap evolution and adaptation (1–3). Genomic islands that carry clustered genes encoding vital functions supply bacteria with additional capabilities to withstand and overcome host defenses and to improve fitness. Genomic islands are integrative elements that are not able to self-transfer and replicate. Typically, they are composed of functional and recombination modules. The recombination module consists of a tyrosine family integrase and two attachment sites involved in recombination. The integrase promotes attP× attB site-specific DNA recombination of the genomic islands into highly conserved tRNA-encoding genes (attB recombination targets) of the host genome and subsequent excision (4, 5).
It has been demonstrated that the N-terminal domain of phage integrases is responsible for specific recognition of the arm-type site sequence of the attachment sites, a step that is essential for activity of the catalytic C-terminal domain responsible for the strand exchange. Three-dimensional structures of the N-terminal domains of two prokaryotic integrases, namely bacteriophage λ integrase and Tn916 transposon integrase, have been determined. Although they do not share significant sequence homology, both adopt similar structures and recognize the arm-type DNA site by inserting their N-terminal domain into a major groove of DNA (6–9). The N-terminal domain, consisting of a three-stranded antiparallel β-sheet, is proposed to be a new DNA-binding motif whose residue composition and position within the major DNA groove varied to alter specificity (6). Nevertheless, the genomic islands are evolutionarily divergent from phages and other mobile elements and represent a distinct mobile genetic element class. Moreover, island-encoded integrases are not closely related to phage integrases, as was expected previously (3, 10).
Four closely related genomic islands, Yersinia HPI (high pathogenicity island; encoding siderophore yersiniabactin biosynthesis and transport) (11), Ecoc54N or the pks island (encoding the cytotoxic polyketide colibactin in uropathogenic Escherichia coli CFT073) (12), and two genomic islands with yet unknown functions in Erwinia carotovora, HAI7 and HAI13, contain highly similar but not identical integrases that recognize asn tDNA genes as their attB integration sites (13). In contrast to the highly conserved bacterial attB attachment site, the element-encoded attP sites and integrases are subjected to sequence fluctuations. Parallel evolution of the corresponding integrase with its cognate attP sites results in inability of the evolved integrase to support recombination of the heterologous attP and to substitute heterologous integrase that has the same bacterial target. In this work, we present new insights into the evolution of two site-specific genomic island integrases through the three-dimensional structures of their N-terminal domains that utilize a novel four-stranded β-sheet DNA-binding motif.
EXPERIMENTAL PROCEDURES
Protein Expression and Purification
The construct of the arm-type binding domain of HPI integrase comprised residues 1–80, whereas a corresponding domain of the HAI7 integrase covered residues 1–100. Both constructs were cloned into the pET21b vector (Novagen). Native proteins were expressed in the BL21(DE3) strain grown on standard LB medium, whereas for the selenomethionine labeling, the B834(DE3) strain was used in a defined modified M9 medium with methionine being replaced by selenomethionine. For all preparations, the medium was inoculated with freshly transformed overnight cultures, grown at 37 °C until A600 = 0.6–0.7 was reached, and induced with 1 mm isopropyl β-d-thiogalactopyranoside. Afterward, the temperature was lowered to 27 °C, and the cells were grown for ∼16 h. After harvesting by centrifugation, cells were resuspended in lysis buffer (300 mm NaCl, 50 mm Na2HPO4, and 10 mm imidazole), broken by sonication, and centrifuged to remove cell debris. Clear supernatant was applied on nickel-nitrilotriacetic acid resin pre-equilibrated with lysis buffer. The column was then washed with lysis buffer supplemented with 20 mm imidazole, and the proteins were eluted with lysis buffer containing 250 mm imidazole. Eluted proteins were applied to a Superdex S-75 preparative gel filtration column in crystallization buffer (50 mm NaCl and 10 mm Tris, pH 7.6). Purity of the proteins was confirmed by SDS-PAGE and mass spectroscopy.
Electrophoretic Mobility Shift Assay Experiments
The ability of recombinant IntHPI80 and IntHAI7100 to bind putative arm-type DNA was evaluated by electrophoretic mobility shift assays using FAMTM-labeled probes generated by annealing two DNA oligonucleotides for each probe: attpHPI_DR_for (5′-FAM-CGCAACTATTGGTGGTCATTATGGTGGTCTTGACAGAAATACAATTCAGGTTTAGCTTAATTA-3′) and attpHPI_DR_rev (5′-TAATTAAGCTAAACCTGAATTGTATTTCTGTCAAGACCACCATAATGACCACCAATAGTTGCG-3′) for the arm-binding site of attPHPI and attpHAI7_DR_for (5′-FAM-AATAATGTTGGTAGAAGTGTTGGTATAAAAAACACTTAAAAAAACAAACTCAATGATAAC-3′) and attpHAI7_DR_rev (5′-GTTATCATTGAGTTTGTTTTTTTAAGTGTTTTTTATACCAACACTTCTACCAACATTATT-3′) for the arm-binding site of attPHAI7. Approximately 50 fmol of the probe was incubated at 28 °C with IntHPI80 and IntHAI7100 as appropriate in 10 μl of binding buffer (20 mm Hepes, pH 7.6, 1 mm EDTA, 10 mm (NH4)2SO4, 1 mm dithiothreitol, 0.2% (w/v) Tween 20, 30 mm KCl, and 250 ng/μl each poly[d(I-C)] and poly[d(A-T)]). After 1 h, the samples were applied to a 6% (29:1) acrylamide/bisacrylamide gel in 0.25× Tris borate/EDTA buffer and electrophoresed at 10 V/cm. Gels were visualized with a Fujifilm FLA-3000 fluorescent scanner.
Crystallization
Crystallization was performed with a sitting drop vapor diffusion method by mixing equal volumes of the proteins and reservoir solutions, using varied starting concentrations of the proteins. Crystals of the HPI integrase domain appeared typically after 24 h. The crystals chosen for the native data set originated from the condition containing 0.1 m trisodium citrate, pH 5.6, and 35% (v/v) tert-butyl alcohol. The selenomethionine crystals were harvested from a drop containing 0.2 m trisodium citrate, 0.1 m sodium cacodylate, pH 6.5, and 30% (v/v) isopropyl alcohol with cryoprotection with 2-methyl-2,4-pentanediol. Crystals of the HAI7 domain initially grew in 0.1 m MES,2 pH 6.0, and 3.2 m ammonium sulfate, and crystals optimized for data collection were obtained from a drop containing 0.1 m MES, pH 6.3, and 2.8 m ammonium sulfate.
Data Collection and Structure Determination
Crystals were plunge-frozen in the cryoprotectant solution containing 30% 2-methyl-2,4-pentanediol in the mother liquor. The diffraction data were measured on the PXII beamline of the Swiss Light Source (Villigen, Switzerland). Data were indexed, integrated, and scaled with the XDS package (14). The native data set from the IntHPI80 crystals was collected at 0.9873 Å, and native crystals were diffracted up to 1.3 Å. The selenomethionine derivative was measured at 0.9796 Å for the peak, 0.9796 Å for the inflection point, 0.972 Å for the high remote, and 0.9875 Å for the low remote data sets; all data sets were collected up to 2.0 Å. Two selenomethionine crystals were measured, and corresponding data sets were merged to obtain higher redundancy. Data sets were of high quality and showed strong anomalous signals. Anomalous scatterers were found using SHELXD software (15). Interestingly, two sites, instead of the expected one site, were found. Map analysis showed that the selenium atom is very close to a special position and that the selenomethionine side chain assumes different conformations in each asymmetric unit, thus locally breaking the crystallographic symmetry. (This is also the case in the native data for the sulfur atom.) Two initial atom positions were refined using the autoSHARP software package (16). The resulting phases were improved by the DM program (17) and used for automated model building with ARP/wARP software (18). The resulting model of ∼80% completeness was inspected and finished manually with the Xfit program (19). Restrained refinement enforced by Refmac5 software was then performed using the native data, followed by the addition of water molecules by ARP/wARP (20). Data collection, phasing, and refinement statistics are presented in Table 1. Most of the model has a clear and well interpretable electron density with the exception of a few solvent-exposed side chains. These parts were omitted in the final model. The R-factor of the presented structure of the HPI integrase domain is 20.8%, and Rfree is 23.5%.
TABLE 1.
Data collection and refinement statistics for HPI and HAI7
| HPI |
HAI7 |
|||||
|---|---|---|---|---|---|---|
| Native | SeMet peak | Inflection | High remote | Low remote | Native | |
| Data collection | ||||||
| Space group | P41212 | P41212 | P65 | |||
| Cell constants (Å) | ||||||
| a | 48.75 | 52.56 | ||||
| b | 78.75 | 52.53 | ||||
| c | 74.20 | 79.47 | ||||
| Resolution range (Å) | 20–1.3 | 50–2.0 | 6–1.6 | |||
| Wavelength (Å) | 0.9873 | 0.9796 | 0.9799 | 0.9720 | 0.9875 | 1.0015 |
| Observed reflections | 391,656 | 196,899 | 196,596 | 196,716 | 196,899 | 183,210 |
| Unique reflections | 19,463 | 11,516 | 11,505 | 11,510 | 11,516 | 15,089 |
| Whole resolution range | ||||||
| Completeness (%) | 85 | 98.4 | 98.3 | 98.4 | 98.4 | 91.6 |
| Rmerge | 3.6 | 1.7 | 1.7 | 2.6 | 2.8 | 2.7 |
| I/σ(I) | 29.07 | 45.32 | 43.69 | 46.44 | 45.32 | 30.67 |
| Last resolution shell | ||||||
| Resolution range (Å) | 1.3–1.4 | 2.0–2.1 | 1.6–1.7 | |||
| Completeness (%) | 65.5 | 96.6 | 96.3 | 96.8 | 96.6 | 76 |
| Rmerge | 10.3 | 2.9 | 2.9 | 2.6 | 2.8 | 10.7 |
| I/σ(I) | 8.98 | 29.85 | 28.42 | 31.62 | 29.85 | 9.31 |
| Phasing | ||||||
| No. of sites found/present | 2/2 (each Se site as 50% occupancy) | |||||
| Phasing power anomalous/isomorphous | 4.2/1.622 | |||||
| FOMa | 0.65 | |||||
| Refinement | ||||||
| No. of reflections | 18,404 | 14,113 | ||||
| Resolution (Å) | 10-1.3 | 6-1.6 | ||||
| R-factor (%) | 20.8 | 14.76 | ||||
| Rfree(%) | 23.5 | 21.21 | ||||
| Average B (Å2) | 16.7 | 23.3 | ||||
| r.m.s. bond length (Å) | 0.008 | 0.024 | ||||
| r.m.s. angles | 1.28° | 2.051° | ||||
| Content of asymmetric unit | ||||||
| No. of protein molecules | 1 | 1 | ||||
| No. of protein residues/atoms | 77/612 | 95/866 | ||||
| No. of solvent atoms | 125 | 130 | ||||
| No. of heavy atom sites identified | 2 | |||||
| Ramachandran statistics | ||||||
| Most favored regions (no./%) | 59/91 | 81/95 | ||||
| Additionally allowed regions (no./%) | 6/9 | 4/5 | ||||
| Generously allowed regions (no./%) | 0/0 | 0/0 | ||||
| Disallowed regions (no./%) | 0/0 | 0/0 | ||||
a FOM, figure of merit; r.m.s., root mean square.
Data for the HAI7 integrase domain crystal, collected to 1.6 Å, were integrated, scaled, and merged by XDS and XSCALE programs. The structure was determined by molecular replacement using the Molrep program from the CCP4 suite with the structure of the previously solved IntHPI80 used as a probe. The model was then refined by Refmac5 and rebuilt by XtalView/Xfit and by a subsequent Refmac5 refinement. Water molecules were added by the ARP/wARP program. Side chains of Lys7, Lys16, Lys46, Lys47, Arg49, Lys90, Gln93, Lys95, and Arg96 had no interpretable electron densities and therefore were omitted in the model. The final R-factor of the structure of the HAI7 integrase domain is 14.7%, and Rfree is 21.2%.
RESULTS
N-terminal Domains of IntHPI and IntHAI7 Bind Short Direct Repeats in Heterologous attP Sites
Previously, we reconstituted the attP site of HPI and defined several copies of imperfect direct repeats located in the regions that can be designated “arms” by analogy with integrases of temperate phages (5). To determine whether an N-terminal part of IntHPI indeed specifically recognizes attPHPI DNA, IntHPI80 was tested in the electrophoretic mobility shift assay for its ability to bind the attPHPI fragment containing direct repeats. This IntHPI80 protein (spanning Met1–Asn80) was able to bind only its cognate attPHPI but not attPHAI7 fragment (Fig. 1). Similarly, the IntHAI7100 protein (Met1–Arg100), representing the N-terminal domain of IntHAI7, was able to form retarded bands with only attPHAI7 but not attPHPI DNA (Fig. 1).
FIGURE 1.
Electrophoretic mobility shift assay of different purified integrase N-terminal domains: IntHAI7-(1–100) and IntHPI-(1–80) with attPHAI7 (A) and attPHPI (B) arm DNAs.
Overall Structures
The structure of the arm-type binding domain of the HPI integrase (residues 1–80) comprises a four-stranded antiparallel β-sheet preceded and followed by α-helices (Fig. 2A). The N-terminal α-helix H1 covers Asp5–Thr10 and is positioned perpendicularly to the C-terminal larger helix H2. Helix H1 is followed by a 7-residue-long loop connecting it to the first β-strand B1 (Phe18–Ser23). β-Strands B1, B2 (Leu26–Lys31), B3 (Gly34–Ile44), and B4 (Lys47–Ala55) are connected by short turns comprising 2 amino acids. β-Strands are parallel in the following pairs: B1 with B2 and B3 with B4. The first two strands (B1 and B2) form angle of ∼45° with respect to strands B3 and B4. The β-sheet is connected to α-helix H2 (Leu61–Ala76) by the 5-amino acid loop. This helix is positioned almost parallel to strands B3 and B4. The β-sheet, together with helix H2, forms the L-shaped hydrophobic core of the protein.
FIGURE 2.
Crystal structures of arm-type binding domains of HPI (A) and HAI7 (B) integrases.
The structure of the HPI integrase domain shows a high resemblance to that of the arm-type binding domain of the HAI7 integrase (Fig. 2B). The positioning of helices and β-strands is identical, and the root mean square deviation of superimposed structures (aligned taking into account C-α atoms) is 0.81 Å. The largest difference in the main chain tracing is localized in the loop between strands B3 and B4, with the IntHAI7100 main chain bending toward helix H1. This segment shows also a difference in its primary structure with the large residues (Glu45 and Lys46) in IntHAI7100 exchanged for small ones in IntHPI80 (serine and glycine, respectively) (Fig. 3). The crystallized IntHAI7 construct is longer than IntHPI (spanning residues 1–100). The residues, located C-terminally to helix H2, are visible in the crystal structure, but this region is largely unstructured, with only residues 81–83, 86–88, and 92–95 forming very short helices. The electron density is, however, well defined, and the polypeptide chain is arranged antiparallel to helix H2.
FIGURE 3.
Alignment of structures of IntHPI80 (cyan) and IntHAI7100 (orange). The inset shows a magnification of the loop between strands B3 and B4.
A comparison of the amino acid sequences of the arm-type binding domains of IntHPI and IntHAI7 is shown in Fig. 6. Conserved residues can be divided into several groups. The amino acids of the N-terminal helix H1 and the following loop form a large cluster of conserved residues (8 being identical and 4 being similar). Several of these residues stabilize the “vertical position” of helix H1 through formation of hydrogen bonds. For example, Arg65 binds the carbonyl group of Leu3, and Leu3 also seems to create a hydrogen connection with the conserved Asp22. Additionally, in IntHPI, the hydroxyl group of Tyr56 binds the carboxyl group of the conserved Asp5 from the top of helix H1.
FIGURE 6.
Alignment of IntHPI80 with IntHAI7100 and IntHAI13100. Identical residues are colored dark blue; similar residues are in light blue. The strands and helices are indicated above the sequences. In the case of the last 20 residues, the helices that are taken from the structure of the HAI7 integrase only are presented as a dotted line. Residues pointing toward the cavity are indicated in red.
The interactions that connect helix H2 to strand B1 seem to be of special importance because all residues responsible for this interaction are conserved. Asp22, Arg65, and Arg68 (exclusively in IntHPI) create a network of hydrogen bonds between one another. Arg68 in IntHAI7, although conserved, does not contribute to the interactions. Another group of well conserved residues are those present in or in the proximity of the turns: Gly25 in the turn between strands B1 and B2; Gly34 and Ser35 in the proximity of the turn between strands B2 and B3; and Gly54, Pro57, and Ala58 in the turn between strand B4 and helix H2.
The IntHAI7100 structure gives additional insight into the positioning of the C-terminal part of the protein. A dense network of hydrogen bonds formed by Lys73–Ser82 caps the C terminus of helix H1 and thus fixes the region with respect to this helix. The following part of the C terminus does not form hydrogen interactions with the rest of the protein, with the exceptions of the carbonyl oxygen of Asn89 bonding Arg67 and the side chain of Asp63 interacting with NH of Ile92 (Fig. 4). It should be mentioned here that, as for proteins of such relatively small sizes (10–12 kDa), the arm-type binding domains of integrases and especially IntHPI possess a rich H-bonding network.
FIGURE 4.
Interactions of the C terminus (shown in backbone representation) with the main body of IntHAI7100 (shown in ribbon representation). Hydrogen bonds are colored green.
Comparison with Structures of Three-stranded β-Sheet DNA-binding Proteins
Several structures of arm-type binding domains of integrases have been solved (6, 21, 22). The alignment of the arm-type binding domain of HPI with the λ integrase and the GCC box-binding protein is shown in Fig. 5; the root mean square factors are 1.51 and 1.48 Å for the backbone atoms, respectively. These proteins bind their DNA recognition sites by means of the three-stranded β-sheet. Our structures differ from these published structures by a prominent extra β-strand, which is an integral part of the β-sheet. Furthermore, all β-strands in the structure of the arm-type binding domains of HPI and HAI7 integrases are significantly longer than in any other arm-type DNA-binding domains, with as many as 11 residues projecting from the β-sheet in the direction of the potential DNA-binding cavity (with the Tn916 transposon integrase having 9 of them and the λ integrase only 7) (Fig. 6).
FIGURE 5.
Structure of IntHPI80 (blue) best fitted to structures of N-terminal domains of the λ integrase (green) and the GCC box-binding protein (red).
The other distinguishing feature of our integrase domain is its structured helical N terminus positioned perpendicularly to the large C-terminal helix. Among all known integrases, a small N-terminal helix is present only in a newly described structure of the λ integrase in complex with DNA (9). The positioning of this helix is, however, different from that in the HPI and HAI integrases, with the N terminus of the λ integrase pointing away from the main body of protein (Fig. 5). Because the structure of the N terminus of the λ integrase proved to be significantly different in the free versus DNA-bound state, it cannot be excluded that the N-terminal helix of HPI and HAI7 integrases could change its conformation upon binding to the attachment site.
It is still unconfirmed whether the mode of DNA binding is conserved between the integrases possessing three-stranded β-sheets and the HPI and HAI7 integrases. The space created by a concave structure beneath the β-sheet is, however, large enough to fit the major groove of DNA. Furthermore, the charge distribution in IntHPI80 and IntHAI7100 is similar to that in other arm-type binding proteins, with positive charges gathered on the concave surface of the β-sheet (Fig. 7).
FIGURE 7.
Comparison of DNA-binding cavities of N-terminal domains of HPI (A) and HAI7 (B) integrases. A model of DNA binding is shown in a backbone representation.
Comparison of DNA-binding Cavities of N-terminal Domains of HAI7 and HPI Integrases
The largest observed differences between our solved structures appear, as expected, on the putative DNA-binding interface. The majority of the residues conserved between IntHPI80 and IntHAI7100 present in the β-sheet have their side chains directed toward the core of the interface between the β-sheet and helix H2 (Leu28, Val30, Trp38, Leu40, and Tyr42). Three conserved residues are facing the putative DNA-binding cavity, namely Lys19, Leu29, and Arg43. Other residues pointing in the same direction differ substantially between the HPI and HAI7 integrases, as well as (although to a lesser extent) between the HAI13 and HAI7 integrases.
Comparison of the putative DNA-binding cavities of IntHPI80 and IntHAI7100 shows several differences. First, the proximal border of the cavity (composed of loop L1, parts of β-strands B1 and B2, and the turn between them) in IntHPI80 bears a strong positive charge, introduced by Lys16, Lys19, and Lys31 (Fig. 7). The situation differs in the case of IntHAI7100, which not only lacks Lys31 (substituted with histidine, which remains buried) but also presents the negative charge from Glu17 protruding from the surface of the cavity. An exact position of the side chain of Lys16, which is present in both proteins, could not be detected in the IntHAI7100 structure, probably due to its high mobility.
The central part of the cavity (formed by the middle parts of strands B3 and B4) in IntHPI80 is dominated by the presence of the positive charge of Lys41, next to the negative charge supplied by Glu48. In IntHAI7100, none of these residues is present (Lys41 being substituted with Ser, and Glu48 with Gln). However, Leu50 protrudes into the cavity approximately in the same place as the charged residues in IntHPI80. The distal rim of the cavity in IntHAI7100 (formed by the C-terminal end of strand B3, the N-terminal part of strand B4, and the loop between them) seems larger and bulkier than that in IntHPI80.
Putative Residues Interacting with DNA
The IntHPI and IntHAI7 DNA-binding interfaces bear too many differences to unambiguously link specific residues with the DNA base recognition. Therefore, we used a related integrase, IntHAI13, which shows a strong sequential resemblance to IntHAI7, to obtain more insight into the residues' functions. A model of the IntHAI13 DNA-binding cavity was constructed based on the structure of HAI7 with surface-exposed residues mutated in the Swiss-PdbViewer (Fig. 8). Comparison of the crystal and modeled structures shows that differences group in a cluster in the central part of the cavity (residues mutated: H27T, R39Q, S41G, and L50V). This results in a slight change of the surface of the cavity, mostly induced by the substitution of Leu50 with a smaller Val side chain. Furthermore, the hydrogen bond pattern between the protein and the target DNA fragment can be changed because a histidine, which can serve both as a donor and an acceptor, is substituted with a donor residue, tyrosine. The reverse is the case for the R39Q substitution, where a double donor is exchanged for the donor + acceptor residue. Another significant change in the DNA-binding interface is the substitution of a negatively charged Glu17 for Val.
FIGURE 8.
Model of the DNA-binding cavity of the N-terminal domain of the HAI13 integrase. A, side view of the DNA-binding cavity in a stick representation. Mutated residues are colored magenta and labeled. B, surface of the cavity colored by an electrostatic potential. Protruding residues are labeled.
The sequences of DNA recognized by the inspected integrases, supplemented by the sequences bound by HAI13, are presented in Fig. 9. Each of the integrases binds to two attachment sites: the right (attR) and the left (attL). Furthermore, each of them contains two binding sites (P1 and P2). Alignment of all 12 binding sites shows that base 4 (indicated in blue in Fig. 9) differs among all integrases. Base 9 (indicated in green) remains the same in the attachment site of both HAI integrases, whereas it is changed in IntHPI. Therefore, differences in the structure of the binding cavity between IntHAI7 and IntHAI13 can be attributed to the recognition of base 4.
FIGURE 9.
Attachment sites of HPI, HAI7, and HAI13 integrases. Bases that are different for each integrase are indicated in blue. Bases that are identical in HAI7 and HAI13, although different from those in HPI, are indicated in green.
Given these two major spatial localizations of differences, it is impossible to safely predict which area is responsible for the differentiation between the IntHAI7 and IntHAI13 binding sites. However, the residues interacting with base 9 are most probably localized in the Glu48/Lys41 region.
DISCUSSION
Two closely related but still non-identical asn tDNA-associated integrases of the four genomic islands give us a unique opportunity to superimpose differences in their amino acid structure and the cognate co-evolved attP recombination sites to address specific protein-DNA recognition. The structures of the arm-type binding domains of both IntHPI and IntHAI7 integrases resemble each other and comprise a four-stranded antiparallel β-sheet preceded and followed by α-helices. The main difference can be localized in the loop between strands B3 and B4, with the HAI7 main chain bending toward helix H1.
Comparison with the structures of the arm-type binding domains of other integrases demonstrates that the arm-type binding domains of IntHPI and IntHAI7 differ by the presence of an extra β-strand. In contrast, all known proteins bind their DNA recognition sites by means of three-stranded β-sheets (6, 21, 22). Additionally, all β-strands in the structures of the arm-type binding domains of HPI and HAI7 integrases are significantly longer than those in three-stranded arm-type binding domains, which enables mimicking the shape of the major groove of a longer DNA fragment. HPI and HAI7 attachment sites comprise eight nucleotides, the same number found as DNA-interacting for the λ integrase (9). In the λ integrase, however, only six nucleotides are recognized through the interaction of the β-sheets with the major groove, whereas the remaining two are bound by the N-terminal helix inserting into the minor groove (9). Therefore, it can be speculated that the long β-sheets in IntHPI and IntHAI7 are sufficient for recognition without a contribution of the N-terminal helix. The other distinguishing feature of the IntHPI and IntHAI7 N-terminal domains is their structured helical N terminus positioned perpendicularly to the large C-terminal helix. These differences support the proposal that integrases of the genomic islands form a distinct evolutionary branch of the site-specific recombinases different from that of phage integrases (3). On the other hand, integrases of the genomic islands are bidirectional recombinases that efficiently support both integrative and excisive recombination and do not require an additional recombination directionality factor (RDF), in contrast to the phage integrases. For example, the excisive activity of HPI by its RDF is supported <10-fold (23), in contrast to the highly efficient RDF of bacteriophage λ. Moreover, not all asn tDNA-associated islands possess an RDF.
Evolutionary integrases of the temperate phages might represent a further step in function speciation to increase the efficiency of prophage rescue under unfavorable conditions. In contrast, genomic islands, being unable to replicate and to transfer, are long term passengers of the bacterial chromosome with the dominating integrative function. In addition, the inefficient RDF is not clustered with the cognate integrase, as in the case of temperate phages. This supports the proposal of their independent and consecutive acquisition by the integrative modules of genomic islands. Likewise, differences in the structures of IntHPI and IntHAI7 might be more suited for the recognition of both attP and recombinant attL/attR sites to support efficiently both types of site-specific recombination. Until now, it also had not been demonstrated whether the mode of DNA binding is conserved between the integrases possessing three-stranded β-sheets and IntHPI and IntHAI7.
Our resolved structures of the arm-type binding domains of the two asn tDNA-associated integrases show the largest observed differences between two structures on putative DNA-binding interfaces. Although 3 conserved residues are facing the putative DNA-binding cavity (Lys19, Leu29, and Arg43), all other residues pointing in the same direction differ substantially between the HPI and HAI7 integrases, as well as (although to a lesser extent) between IntHAI13 and IntHAI7. Such β-sheet sequence-specific DNA binding is very rare and has been demonstrated only for three proteins that form complexes with DNA, namely the bacteriophage λ integrase, the Tn916 transposon integrase, and the ethylene-responsive factor AtERF1 from Arabidopsis thaliana.
At the moment, it is impossible to demonstrate what DNA nucleotides are actually interacting with the amino acid residues in the DNA-binding interface. However, co-crystallization of the arm-type binding domains with the cognate DNA arms might be the next step to solve this problem.
The work was supported by Deutsche Forschungsgemeinschaft Grants RA 991/1-2 and RA 991/1-3.
The atomic coordinates and structure factors (codes 3JTZ and 3JU0) have been deposited in the Protein Data Bank, Research Collaboratory for Structural Bioinformatics, Rutgers University, New Brunswick, NJ (http://www.rcsb.org/).
- MES
- 4-morpholineethanesulfonic acid
- RDF
- recombination directionality factor.
REFERENCES
- 1.Dobrindt U., Hochhut B., Hentschel U., Hacker J. (2004) Nat. Rev. Microbiol. 2, 414–424 [DOI] [PubMed] [Google Scholar]
- 2.Hacker J., Blum-Oehler G., Mühldorfer I., Tschäpe H. (1997) Mol. Microbiol. 23, 1089–1097 [DOI] [PubMed] [Google Scholar]
- 3.Boyd E. F., Almagro-Moreno S., Parent M. A. (2009) Trends Microbiol. 17, 47–53 [DOI] [PubMed] [Google Scholar]
- 4.Williams K. P. (2002) Nucleic Acids Res. 30, 866–875 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Rakin A., Noelting C., Schropp P., Heesemann J. (2001) Mol. Microbiol. 39, 407–415 [DOI] [PubMed] [Google Scholar]
- 6.Wojciak J. M., Sarkar D., Landy A., Clubb R. T. (2002) Proc. Natl. Acad. Sci. U.S.A. 99, 3434–3439 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Biswas T., Aihara H., Radman-Livaja M., Filman D., Landy A., Ellenberger T. (2005) Nature 435, 1059–1066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Connolly K. M., Wojciak J. M., Clubb R. T. (1998) Nat. Struct. Biol. 5, 546–550 [DOI] [PubMed] [Google Scholar]
- 9.Fadeev E. A., Sam M. D., Clubb R. T. (2009) J. Mol. Biol. 388, 682–690 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Balding C., Bromley S. A., Pickup R. W., Saunders J. R. (2005) Environ. Microbiol. 7, 1558–1567 [DOI] [PubMed] [Google Scholar]
- 11.Rakin A., Schubert S., Pelludat C., Brem D., Heesemann J. (1999) in The High-Pathogenicity Island of Yersinia (Kaper J. B., H. J., eds) ASM Press, Washington, DC [Google Scholar]
- 12.Nougayrède J. P., Homburg S., Taieb F., Boury M., Brzuszkiewicz E., Gottschalk G., Buchrieser C., Hacker J., Dobrindt U., Oswald E. (2006) Science 313, 848–851 [DOI] [PubMed] [Google Scholar]
- 13.Bell K. S., Sebaihia M., Pritchard L., Holden M. T., Hyman L. J., Holeva M. C., Thomson N. R., Bentley S. D., Churcher L. J., Mungall K., Atkin R., Bason N., Brooks K., Chillingworth T., Clark K., Doggett J., Fraser A., Hance Z., Hauser H., Jagels K., Moule S., Norbertczak H., Ormond D., Price C., Quail M. A., Sanders M., Walker D., Whitehead S., Salmond G. P., Birch P. R., Parkhill J., Toth I. K. (2004) Proc. Natl. Acad. Sci. U.S.A. 101, 11105–11110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Kabsch W. (1993) J. Appl. Crystallogr. 26, 795–800 [Google Scholar]
- 15.Schneider T. R., Sheldrick G. M. (2002) Acta Crystallogr. D Biol. Crystallogr. 58, 1772–1779 [DOI] [PubMed] [Google Scholar]
- 16.de La Fortelle E., Bricogne G. (1997) Methods Enzymol. 276, 472–494 [DOI] [PubMed] [Google Scholar]
- 17.CCP4, Collaborative Computational Project Number 4 (1994) Acta Crystallogr. D Biol. Crystallogr. 50, 760–763 [DOI] [PubMed] [Google Scholar]
- 18.Perrakis A., Morris R., Lamzin V. S. (1999) Nat. Struct. Biol. 6, 458–463 [DOI] [PubMed] [Google Scholar]
- 19.McRee D. E. (1999) J. Struct. Biol. 125, 156–165 [DOI] [PubMed] [Google Scholar]
- 20.Lamzin V. S., Wilson K. S. (1993) Acta Crystallogr. D Biol. Crystallogr. 49, 129–147 [DOI] [PubMed] [Google Scholar]
- 21.Allen M. D., Yamasaki K., Ohme-Takagi M., Tateno M., Suzuki M. (1998) EMBO J. 17, 5484–5496 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wojciak J. M., Connolly K. M., Clubb R. T. (1999) Nat. Struct. Biol. 6, 366–373 [DOI] [PubMed] [Google Scholar]
- 23.Antonenka U., Nölting C., Heesemann J., Rakin A. (2006) Int. J. Med. Microbiol. 296, 341–352 [DOI] [PubMed] [Google Scholar]









