Abstract
Insertion sequences (IS) are mobile genetic elements that are distributed in many prokaryotes. In particular, in the genomes of the symbiotic nitrogen-fixing bacteria collectively known as rhizobia, IS are fairly abundant in plasmids or chromosomal islands that carry the genes needed for symbiosis. Here, we report an analysis of the distribution and genetic conservation of the IS found in the genome of Rhizobium etli CFN42 in a collection of 87 Rhizobium strains belonging to populations with different geographical origins. We used PCR to generate presence/absence profiles of the 39 IS found in R. etli CFN42 and evaluated whether the IS were located in consistent genomic contexts. We found that the IS from the symbiotic plasmid were frequently present in the analyzed strains, whereas the chromosomal IS were observed less frequently. We then examined the evolutionary dynamics of these strains based on a population genetic analysis of two chromosomal housekeeping genes (glyA and dnaB) and three symbiotic sequences (nodC and the two IS elements). Our results indicate that the IS contained within the symbiotic plasmid have a higher degree of genomic context conservation, lower nucleotide diversity and genetic differentiation, and fewer recombination events than the chromosomal housekeeping genes. These results suggest that the R. etli populations diverged recently in Mexico, that the symbiotic plasmid also had a recent origin, and that the IS elements have undergone a process of cyclic infection and expansion.
Insertion sequences (IS) are the smallest transposable elements found in prokaryotes (usually less than 3 kb in size). They encode a transposase and may also encode small hypothetical proteins (4, 9). IS are distinguished by their ability to move within prokaryotic replicons, including both the chromosome and plasmids, and to copy themselves into various genomic sites. In this manner, IS elements can inactivate or alter the expression of adjacent genes (4). When IS occur in two or more identical copies within a genome, they can participate in various types of genetic rearrangements (e.g., duplications, inversions, and deletions), suggesting that IS may play an important role in the evolution of their hosts by promoting genomic plasticity (34). Due to these evolutionary dynamics, the diversity and distribution of IS elements differ greatly between taxa and even within strains of the same species (27).
Various theories have been proposed to explain the evolution of IS elements in laboratory model strains and environmental bacterial populations (8, 18, 25, 29). Two main hypotheses seek to explain how these elements are maintained over the long term in their host genomes. The first proposes that they occasionally generate beneficial mutations and therefore may represent a selective advantage to their hosts (34). The second suggests that IS elements are genomic parasites that are maintained by their high rate of transposition and might be disseminated among different bacterial lineages by horizontal gene transfer (HGT). Data supporting the second hypothesis have shown that some IS elements may transpose at high rates upon entering a new host (42). After the initial infection, however, purifying selection may continuously remove these elements from the genome. Thus, IS may undergo an infection-expansion-extinction cycle that allows them to remain in different bacterial populations within the gene pool (42). These two hypotheses are not contradictory, and the evolutionary dynamics and distribution of IS may differ greatly depending on several factors, including (most notably) the rate of transposition and HGT, as well as selective pressures, population size, and the host's habitat (6, 18, 21, 25, 27, 29).
In the nitrogen-fixing symbiotic bacteria of the genera Rhizobium, Sinorhizobium, Mesorhizobium, Bradyrhizobium (of the alphaproteobacteria), Cupriavidus, and Burkholderia (of the betaproteobacteria), IS are particularly abundant in symbiotic plasmids (pSym) and symbiotic chromosomal islands (SI) (2, 12, 14, 19, 20, 43). SI and pSym include most of the genes needed to establish symbiosis in the roots of leguminous plants through nodule formation and nitrogen fixation (11). It is generally believed that these elements entered the rhizobial genomes through HGT (39, 40). Comparative genomic analyses have shown that both pSym and SI are highly variable, with the exception of a common set of genes encoding factors critical to nitrogen fixation (nif) and nodulation (nod) (5, 14). SI and pSym have been found to have lower GC contents and different codon usages than the corresponding chromosomal and nonsymbiotic plasmid sequences, suggesting that they were recently acquired by HGT.
Some of these symbiotic elements, such as in pSym of Rhizobium etli CFN42 and the SI of Mesorhizobium loti, are conjugative and mobile (30, 32). Genomic analysis of R. etli CFN42 revealed the presence of 39 IS belonging to different families (14); they were found in the chromosome (11 IS); the 371-kb symbiotic plasmid (13 IS); the smaller 192-kb conjugative plasmid, p42a (13 IS); and two other plasmids, p42b and p42c (2 IS). Interestingly, this particular strain shows no evidence of IS disrupting open reading frames (ORFs) or having transpositional activity. However, another 42 incomplete IS may be found in the chromosome, pSym, and the conjugative plasmid; these incomplete sequences are truncated or contain stop codons in their coding sequences.
Here, we focused on the dynamics and distribution of IS in different populations of the nitrogen-fixing symbiont R. etli. Since the maintenance of IS in bacterial species might depend on their transpositional activities and horizontal transfer rates, the identification of IS in the same genomic contexts across different strains of the same species could provide new insights into their persistence and divergence over short evolutionary periods. To examine the evolutionary dynamics of IS in natural populations of R. etli, we characterized the distributions, genomic contexts, and sequence variations of IS in isolates of R. etli from three populations with different origins, as well as in some other Rhizobium species. More specifically, we used PCR to generate presence/absence profiles of the 39 IS found in R. etli CFN42 in a collection of 87 strains representing different geographical sites and a gradient of domestication of the bacterial host, the common bean (Phaseolus vulgaris). We also evaluated whether the IS were conserved in the same genomic context relative to their position in R. etli CFN42 and determined the nucleotide sequences of two IS found in most of the isolates. Several population genetic tests applied to these IS, another pSym gene (nodC), and two chromosomal housekeeping genes (dnaB and glyA) suggest that these two IS elements have been inherited vertically and represent recent components of the R. etli gene pool. Finally, the present study strongly suggests that symbiotic plasmids have a recent origin within the R. etli populations.
MATERIALS AND METHODS
Bacterial strains, growth conditions, and DNA extraction.
The 87 Rhizobium strains used in this study correspond to three different collections (Table 1). Two of the strain collections are from Mexico. The first collection (36) was derived from two plots in a traditional milpa system of native bean landraces (San Miguel Acuexcomac, Puebla, Mexico); the collection consists of 30 strains that were the dominant strains for several years. The second collection (13) comes from the Michoacan-Guanajuato area, which is the reported center of origin of bean domestication (24); the 30 strains in this collection include a bean plant domestication gradient from wild, nondomesticated P. vulgaris to milpa landrace cultivars, as well as wild bean plants that are probably the descendants of cultivated plants. The third collection (33) represents R. etli from Spain and includes 27 strains obtained from 21 soil samples collected along the Guadalquivir River Valley. They are believed to represent either original native rhizobia or a subsample of New World rhizobia that traveled along with the original bean seeds (Table 1). Some strains from the Puebla and Spanish collections were previously characterized as Rhizobium gallicum, Rhizobium giardini, Rhizobium fredii, and Rhizobium leguminosarum (Table 1).
TABLE 1.
Rhizobium isolates and reference strains used in this study
| Origin | No. of strains | Species | Identifier(s) | Reference |
|---|---|---|---|---|
| Rhizobium isolates | ||||
| Puebla, Mexico | 30 | R. etli | 1009, 1006, 4771, 1004, 2737, 4815, 4803, 4777, 950, 2704, 4813, 4794, 951, 994, 4837, 4804, 954, 4877, 4795, 2730, 4810 | 34 |
| R. gallicum | 4868, 4872, 4770, 4845, 988, 2751, 2735, 992, 2703 | |||
| Guanajuato, Mexico | 30 | R. etli | Domesticated: 6854, 6861, 6857, 6832, 6868, 6867, 6840, 6846, 6850, 6851, 6833, 6845, 6862 | 12 |
| Wild: 6779, 6778, 6760, 6794, 6805, 6795, 6766, 6784, 6797, 6798, 6776, 6768, 6763 | ||||
| Weedy: 6815, 6824, 6823, 6813 | ||||
| Valle Guadalquivir, Spain | 27 | R. etli | 21PR-1, 16C-1, 21NJ-2, 8C-3, 8NJ-2, 14PR-2, 17NJ-2, 16NJ-2, 14C-1, 4PR-2, 17C-2, 6PR-1, 6C-1, 4C-2, GR10, GR62, GR14, GR87, GR56 | 31 |
| R. gallicum | GR18, GR45, GR60, GR42 | |||
| R. giardinni | GR93, GR03 | |||
| R. fredii | GR64 | |||
| R. leguminosarum | GR84 | |||
| Rhizobium reference strains | ||||
| Mexico | 1 | R. etli | CFN42 | 13 |
| Costa Rica | 1 | R. etli | CIAT652 | 10 |
The various Rhizobium strains were grown in peptone-yeast (PY) medium for 24 to 48 h. Genomic DNA was extracted with a GenomicPrep DNA Isolation Kit (Amersham Biosciences), and the DNA concentration was determined with a DyNA Quant 200 fluorometer (Hoefer). Two completely sequenced R. etli strains, CFN42 and CIAT652, were used as references strains.
PCR amplification and DNA sequencing.
We first localized each of the 39 selected IS within the CFN42 genome, thereby “anchoring” our examination of whether these IS conserved their genomic locations (i.e., their synteny) across the wild R. etli isolates from different geographical regions. These 39 IS elements represent 11 different families, and some of these families have identical or nearly identical copy numbers: 6 and 4 copies for 2 different elements of the IS66 family, 5 copies for IS630, 2 for IS1111, and 2 for IS21. We then designed specific PCR primers spanning the immediate neighborhoods of the 39 studied IS. These primers allowed us to test for conservation (i.e., amplification of a fragment of a size similar to that predicted in CFN42). In cases where we found two contiguous IS elements, we designed primers in the neighboring genes of both IS. There were three possible results: (i) no PCR product, indicating that there were no identical priming sites in the wild isolate (this was not an absolute positive or negative result for the presence of the IS); (ii) a small DNA fragment equivalent to the distance between the two sequences near the insertion site but lacking the IS (a negative result for the IS); and (iii) a large DNA fragment of a size similar to that in CFN42 (a positive result for the IS). The PCRs were performed in a DNA thermal cycler (Gene Amp 9700; Applied Biosystems) in a 25-μl reaction mixture containing approximately 10 ng genomic template DNA, 2 mM MgCl2, 1 mM deoxynucleoside triphosphates (dNTPs), 5 pmol of each primer, and 2 U of Taq polymerase (AltaEnzymes). The mixtures were subjected to 5 min of denaturation at 94°C followed by 30 cycles of 1 min at 94°C, 1 to 3 min at 58 to 62°C, and 3 min at 72°C.
For a comparison of evolutionary dynamics, we performed direct sequencing on the following: two IS elements, ISRel4 and ISRel2 from pSym, which were present in the highest proportion of test samples (see Results and Fig. 2); nodC, which is a pSym gene encoding an N-acetylglucosaminyltransferase that participates in the nodulation process; and two chromosomal housekeeping genes, glyA (serine hydroxymethyltransferase) and dnaB (replicative DNA helicase), which have been proposed to serve as predictors of genome relatedness because their sequence divergence rates reflect the overall rate of genome divergence (44). Internal PCR primers were designed for the last three genes to obtain partial sequences for each gene. The sequencing reaction mixtures contained approximately 10 ng genomic template DNA, 1 mM MgCl2, 1 mM dNTPs, 5 pmol of each primer, and 1 U of rTth Polymerase XL (Applied Biosystems). The mixtures were subjected to 5 min of denaturation at 92°C, followed by 30 cycles of 30 s at 92°C and 1 to 6 min at 58 to 62°C. The PCR products were purified with an Exo/SAP kit (Affymetrix) and sequenced using a Dye-terminator cycle-sequencing kit (Perkin Elmer Applied Biosystems). The sequencing reactions were run on an ABI 3730 sequencer (Applied Biosystems).
Sequence analysis, alignments, and phylogenetic reconstruction.
Sequence quality analysis, assembly, and comparison were performed using SeqScape software V2.5 (Applied Biosystems). For IS characterization, we compared all of the putative transposases in the complete annotated genome sequence of R. etli CFN42 to the nr Database of NCBI and the Insertion Sequence Database (35) with BLASTp (1).
To define a complete and (most likely) functional IS, we used the same criteria applied by the authors of the Insertion Sequence Database as follows: (i) similar organizations of the genes, (ii) sequence similarity of the transposases and inverted repeats, and (iii) presence of direct repeats (14). Gene alignments for the assessed genes and IS elements were done using the MUSCLE program (7). For each alignment, the best substitution model was determined using FindModel (31). Phylogenetic reconstruction was performed using the maximum-likelihood method implemented in PHYML (16). The analysis was carried out with a nonparametric bootstrap analysis of 100 replicates for each alignment. The complete chromosome and plasmid sequences of R. etli CFN42 and CIAT652 were compared with the Mummer program (23). The genomic contexts of the IS in the two strains were examined using Perl programs. The rate of synonymous and nonsynonymous substitutions, the dn/ds ratio, was evaluated using SNAP (22).
Population genetic parameter estimation.
DNASP version 4 (26) was used to assess the following parameters: the average nucleotide divergence between populations (Dxy), the average nucleotide diversity within populations, the average nucleotide divergence per site (π), the average nucleotide diversity at synonymous (πs) and nonsynonymous (πns) sites, gene flow estimates (Fst and Nm), genetic differentiation estimates (Ks, Kst, Z, and Snn), the numbers of shared haplotypes and fixed and shared polymorphisms, and Tajima's D neutrality test.
Recombination analysis.
Recombination tests were performed with RDP3 (28) and SplitsTree4 (17) for the five genes studied in each of the collections. The RDP3 software was applied to detect and analyze recombination signals using eight different methods (RDP, GENECONV, Bootscan, MaxChi, Chimaera, SiScan, 3SEQ, and LARD), with 100 permutation steps, a Bonferroni correction for multiple tests, and a P value threshold of 0.05. Recombination events and breakpoints were accepted if they were detected by two or more methods that used different approximations to detect recombination. The recombination events were then confirmed independently for each recombinant strain, as suggested in the literature (RDP website user manual [http://darwin.uvigo.es/rdp/rdp.html]) (28). Specific parameter modifications were used for each gene alignment depending on the number of sequences in the analysis. Split decomposition and neighbor net analyses were performed with SplitsTree4 software. Both phylogenetic networks represent incompatibilities within and between data sets by the use of splits. For the neighbor net analysis, we applied a split filter based on the network weight and a threshold value of 95% (3).
Nucleotide sequence accession numbers.
The relevant sequence data have been deposited in the GenBank database under accession numbers GUO84443 to GUO84796.
RESULTS
Differentiation of R. etli populations.
To study the evolutionary dynamics of the IS of R. etli, we first asked if the three studied collections of R. etli strains, which were obtained from the root nodules of P. vulgaris plants located in different areas, were geographically structured. The phylogenies constructed based on the chromosomal housekeeping genes, glyA and dnaB, were relatively similar to each other, with a few differences in the groupings of some strains (data not shown). In order to obtain a phylogenetic reconstruction that more accurately represented the evolutionary relationships among the 87 strains contained within the three collections, we concatenated the sequence alignments (2,238 bp) of both genes. We obtained three large R. etli clusters that broadly corresponded to the three different geographical areas from which the strains were collected. The Puebla collection formed one group with some intermingled strains from the Michoacan-Guanajuato and Spanish collections (Fig. 1). The second group consisted of the Michoacan-Guanajuato strains with three intermingled sequences belonging to the Spanish collection. The third R. etli group represented the Spanish collection with a single intermingled sequence from the Puebla collection. The Michoacan-Guanajuato and Spanish collections were more closely related to one another than to the Puebla collection (Fig. 1). A distant fourth group, which was later used as the outgroup, contained several R. gallicum strains from Puebla and Spain. Since this phylogenetic reconstruction clearly separated the three different collections according to their geographic origins, we considered them to be different populations in the subsequent analyses.
FIG. 1.
Phylogenetic reconstruction representing the evolutionary relationships across the three collections. The circles inside the branches indicate bootstrap support of >70%. The external circles indicate intermingled R. etli strains from one of the other collections. The arrowheads indicate other intermingled species (4872, GR42, and GR60 are R. gallicum; GR93 is R. giardinii; and GR84 is R. leguminosarum). The underlined strains represent recombinant strains for dnaB or glyA between the three populations and the R. gallicum strain.
IS profiles of the R. etli populations.
We next examined which IS from R. etli CFN42 were also maintained in the same genomic context among the R. etli isolates from the three different populations. Our PCR-based profiling (see Materials and Methods) yielded either positive or negative results for 24 of the 39 IS and showed marked differences in the conservation of the various IS among the tested populations (Fig. 2). ISRel4 and ISRel2 (from pSym of R. etli CFN42) were the most common IS, present in more than 50% of the individual strains. Sequencing experiments showed that ISRel9 (also from pSym) in most cases was actually a hypothetical protein (hyp304) not related to any IS element and similar in size to ISRel9. Only five isolates (all from Puebla) had an intact ISRel9 in the expected genomic context; these elements had an average identity of 100%. ISRel5 and ISRel12 of pSym were relatively frequent, as they were present in 25% of the tested strains. The other seven IS from pSym were not common. The IS from p42a were found in the Puebla and Michoacan-Guanajuato populations, but not in the collection from Spain, while the chromosomal IS were found only in the Puebla strains. Interestingly, the CFN42 reference strain was found to be close to the Michoacan-Guanajuato population (Fig. 1).
FIG. 2.
Presence/absence profile of the IS elements present in the chromosome and conjugative and symbiotic plasmids across the three populations. The bottom line represents the CFN42 control strain. The red bars indicate positive confirmation of the IS element (a DNA fragment of similar size to that in CFN42), the blue bars indicate a negative result for the presence of the IS element (a DNA fragment equivalent to the distance between the two sequences near the insertion site but lacking the IS), no color indicates no PCR product, and the yellow bars indicate DNA fragments that were bigger than expected.
DNA sequence divergence of ISRel4 and ISRel2.
The observation that some IS occur in the same genomic context suggests that these IS may have been present since the arrival of pSym in the populations, as it is far less likely that independent transposition events could have inserted the IS into the same genomic sites. It is important to mention that ISRel4 and ISRel2 belong to IS families, ISAs1 and IS481, respectively, that do not have specific target sequences for transposition. To investigate the evolutionary dynamics of the two IS elements, ISRel4 and ISRel2, that show conservation of their genomic contexts, we looked for differences in the degree of diversification and selection pressures among these IS elements, as well as in other symbiotic and chromosomal housekeeping genes (see Table S1 in the supplemental material). The DNA sequence conservation was very high for ISRel4 and ISRel2, with average identities of 98.7% and 99.4%, respectively. This is surprisingly high compared to the sequence identities for the housekeeping genes glyA and dnaB, which were 75.3% and 76.4%, respectively. Given that the diversification of chromosomal genes could differ from that on the symbiotic plasmid, we sequenced the nodC gene, which encodes an N-acetylglucosaminyltransferase involved in the synthesis of sugar backbones for the production of lipooligosaccharides (critical for nodulation signaling). This was done in order to have another gene to trace the diversification differences between pSym and the chromosome. The average nucleotide identity among the 44 analyzed nodC sequences was 96.7%. The nucleotide diversity (π) values for the pSym sequences (ISAs1, IS481, and nodC) were very low compared with those of glyA and dnaB (Table 2). The total average π values of dnaB and glyA were 0.066 and 0.088, respectively, whereas those for ISRel4, ISRel2, and nodC were 0.002, 0.002, and 0.014, respectively (Table 2). Since the number of polymorphic sites and the π values for ISRel4, ISRel2, and nodC were very low, we hypothesize that these sequences may have entered the gene pools of the three populations relatively recently.
TABLE 2.
DNA divergence in R. etli IS elements, plasmid genes, and chromosomal genes
| Gene | No. of polymorphic sitesa | dn/ds | Θb | Nucleotide diversity (π) | Nucleotide diversity at synonymous sites (πs) | Nucleotide diversity at nonsynonymous sites (πns) | Tajima's D |
|---|---|---|---|---|---|---|---|
| dnaB | 263 (23.6) | 0.019 | 0.047 | 0.066 | 0.237 | 0.009 | 1.39 |
| glyA | 278 (24.7) | 0.029 | 0.070 | 0.088 | 0.122 | 0.075 | 0.84 |
| nodC | 19 (3.3) | 0.355 | 0.007 | 0.014 | 0.028 | 0.009 | 0.72 |
| ISRel2 | 6 (0.6) | 0.267 | 0.001 | 0.002 | 0.001 | 0.002 | 2.72 |
| ISRel4 | 13 (1.3) | 0.325 | 0.002 | 0.002 | 0.003 | 0.007 | 1.50 |
The percentages of polymorphic sites per sequence length are shown in parentheses.
Population mutation rate (per bp).
Recent origins of ISRel4, ISRel2, and pSym.
To evaluate the hypothesis that ISRel4 and ISRel2 may have originated recently, we measured their average nucleotide diversities at synonymous sites (πs) and nonsynonymous sites (πns) and compared these values with those for nodC, glyA, and dnaB. The πs value reflects the age of an allele in the genetic pool; genes with lower πs values are generally considered to have a recent common ancestor. Our results revealed that the πs values were low for ISRel4, ISRel2, and nodC but high for dnaB and glyA, whereas the πns values were low for all five sequences (Table 2). Similar results were obtained when these πs and πns values were measured within each population (see Table S2 in the supplemental material). These findings indicate that both the IS elements and nodC might have recently entered the gene pools of the three populations.
Overall, the dn/ds ratio was lower for the chromosomal genes than for the pSym genes. To test what this might mean in terms of pSym evolution, we compared the shared haplotypes for chromosomal and pSym genes among the three populations. We did not find any shared dnaB and glyA (i.e., chromosomal) haplotypes between the populations. In contrast, we identified a number of shared haplotypes for the pSym genes (nodC, ISRel2, and ISRel4) (data not shown). These results suggest either that pSym has a relatively recent origin or that there is high gene flow among pSym sequences; in any case, it seems that the pSym sequences have a different evolutionary history than the chromosomal genes. To further test the possible recent origin of pSym sequences in the gene pools of the three populations, we compared the mean genetic divergence values between populations (Dxy) to the π values describing the mean genetic divergence within the populations. Higher values of Dxy indicate increasing time since population divergence. Table 3 shows that the Dxy values were higher than the mean within-group divergence (π) for the chromosomal genes, potentially reflecting phylogenetic differentiation among the three populations. In the case of the pSym sequences, however, the Dxy and π values were very low and did not show clear phylogenetic differentiation between the populations. This further supports the idea that the pSym sequences have a relatively recent origin in the gene pools of the three populations and have followed a different evolutionary history than the chromosomal genes.
TABLE 3.
Comparison of mean genetic divergence between populations (Dxy) and mean genetic divergence (π) within each population
| Gene and origina | Dxy | π |
|---|---|---|
| dnaB | ||
| Pue-Gto | 0.072 | 0.056-0.052 |
| Pue-Spain | 0.068 | 0.056-0.064 |
| Gto-Spain | 0.073 | 0.052-0.064 |
| glyA | ||
| Pue-Gto | 0.097 | 0.080-0.066 |
| Pue-Spain | 0.100 | 0.080-0.077 |
| Gto-Spain | 0.089 | 0.066-0.077 |
| ISRel2 | ||
| Pue-Gto | 0.003 | 0.002-0.002 |
| Pue-Spain | 0.002 | 0.002-0.001 |
| Gto-Spain | 0.002 | 0.002-0.001 |
| ISRel4 | ||
| Pue-Gto | 0.002 | 0.001-0.003 |
| Pue-Spain | 0.002 | 0.001-0.002 |
| Gto-Spain | 0.002 | 0.003-0.002 |
| nodC | ||
| Pue-Gto | 0.003 | 0.001-0.005 |
| Pue-Spain | 0.026 | 0.001-0.008 |
| Gto-Spain | 0.025 | 0.005-0.008 |
Pue, Puebla; Gto, Guanajuato.
We applied additional genetic-differentiation tests (Ks*, Kst, Z, Z*, and Snn) to the pSym and chromosomal sequences to test whether the two IS and nodC have a lower level of genetic differentiation than the chromosomal sequences, which would further support a recent origin within the three populations. The results of the genetic-differentiation tests for the two housekeeping genes were highly significant (P < 0.001), supporting the idea of genetic differentiation among the three populations (Table 4; see Table S3 in the supplemental material). In contrast, the three pSym sequences had much lower levels of genetic differentiation across the three populations. For example, comparison of the pSym sequences from the two Mexican populations (those from Puebla and Guanajuato-Michoacan) yielded Kst values that were either nonsignificant or just barely significant at a P value of <0.05 (Table 4). Although the chromosomal housekeeping genes showed evidence of genetic differentiation, we did not find evidence of fixed polymorphisms; this may indicate that the three populations also diverged recently. Given that the pSym sequences had low πs values, their Dxy and π values were not well differentiated and had several shared alleles, and the genetic-differentiation tests were not significant, it is not surprising that these sequences also lacked fixed polymorphisms.
TABLE 4.
Genetic differentiation and gene flow tests in R. etli IS elements, plasmid genes, and chromosomal genes
| Gene and origina | Kstb | Fst | Nm |
|---|---|---|---|
| dnaB | |||
| Pue-Gto | 0.146*** | 0.252 | 1.48 |
| Pue-Spain | 0.064*** | 0.118 | 3.72 |
| Gto-Spain | 0.118*** | 0.208 | 1.90 |
| glyA | |||
| Pue-Gto | 0.144*** | 0.249 | 1.51 |
| Pue-Spain | 0.142*** | 0.246 | 1.54 |
| Gto-Spain | 0.129*** | 0.226 | 1.71 |
| ISRel2 | |||
| Pue-Gto | 0.007 | 0.025 | 19.75 |
| Pue-Spain | 0.652** | 0.845 | 0.09 |
| Gto-Spain | 0.607** | 0.750 | 0.17 |
| ISRel4 | |||
| Pue-Gto | 0.056 | 0.217 | 1.80 |
| Pue-Spain | 0.310** | 0.634 | 0.29 |
| Gto-Spain | 0.156** | 0.270 | 1.35 |
| nodC | |||
| Pue-Gto | 0.184* | 0.374 | 0.84 |
| Pue-Spain | 0.118* | 0.258 | 1.44 |
| Gto-Spain | 0.225** | 0.357 | 0.90 |
Pue, Puebla; Gto, Guanajuato.
The asterisks represent the probabilities obtained by a permutation test with 1,000 replicates (*, 0.01 < P < 0.05; **, 0.001 < P < 0.01; ***, P < 0.001).
Next, we analyzed whether there was some degree of gene flow across the three populations. If an Fst value above 0.25 and an Nm value of >1 were taken to represent a significant between-population gene flow (37), most of the analyzed genes could be considered to have evidence of gene flow. The pSym sequences often had higher Fst values than the chromosomal genes (Table 4). A comparison between the Mexican and Spanish populations also yielded values indicative of gene flow. Given the geographical distance between these populations, it seems conceivable that these values could reflect a recent divergence of the three populations and a recent origin for the pSym genes.
There is published evidence indicating that genetic exchange occurred within R. etli populations from a single field; this process involved both plasmid and chromosomal loci (36). Accordingly, we hypothesized that the numbers of recombination events would be different between the chromosomal and pSym genes if their divergence times were clearly distinct. To assess this, we used the RDP3 software package to implement eight different recombination tests to search for recombination events across the three populations (see Table S4 in the supplemental material). The chromosomal genes, dnaB and glyA, showed evidence of four and two intragenic recombination events involving 10 and 5 strains, respectively. In contrast, only one pSym sequence (ISRel4) showed a recombination event, even though the low sequence divergence made it harder to detect recombination. Interestingly, some of the same recombination events were found in strains of both R. etli and R. gallicum; moreover, four recombination events found between populations could explain some intermingled strains in the phylogeny of the three populations (Fig. 1; see Table S4 in the supplemental material).
Next, we used a phylogenetic-network approach, consisting of split decomposition and neighbor net analyses, to test whether the inconsistencies in the phylogenetic reconstructions could be due to recombination events. The split decomposition analysis of the two chromosomal genes showed a partition in the data between sequences from the R. etli and R. gallicum species. In the case of glyA, we found some splits within the R. etli strains; these could come from the detected recombination events. ISRel4 was the only pSym sequence for which we obtained a split that clearly represented the unique recombination event we detected in the R. etli and R. gallicum strains. The neighbor net analysis, which is more sensitive to the conflicting sequence data that may represent recombination events, showed that the two chromosomal genes had several splits, some of which were related to the detected recombination events (see Fig. S1 in the supplemental material). ISRel4 yielded the same split we had found in the split decomposition analysis. The other two pSym sequences failed to reveal splits with either analysis; this was consistent with the lack of any recombination events detected by our RDP3 analysis of these sequences. The relative lack of splits and recombination events in the pSym sequences supports our hypothesis that these genes (and probably the symbiotic plasmid itself) originated relatively recently in the three populations.
Another possibility that could explain the high DNA conservation of the pSym sequences and the conserved genomic contexts of ISRel2 and ISRel4 is that selection pressures may inhibit variation at these sequences. However, Tajima's D neutrality tests showed that most of the observed nucleotide substitution patterns in the chromosomal and pSym sequences from the three populations were under a neutral equilibrium model (i.e., all statistical results were nonsignificant) (Table 2; see Table S2 in the supplemental material). The neutral equilibrium model was rejected only for nodC. However, this could be due to the effect of a decrease in population size and/or balancing selection.
IS elements in R. etli symbiotic plasmids.
We recently sequenced another R. etli strain, the Costa Rican strain CIAT652, which has one chromosome, three plasmids, and 22 IS elements distributed mainly in the chromosome and the symbiotic plasmid (15). Comparative analysis of the CFN42 and CIAT652 symbiotic plasmids showed that their sequences were highly conserved, with a nucleotide identity value of 99% (Fig. 3). In contrast, the identities of the chromosome and other plasmid sequences ranged from 90 to 95%. The two pSyms differed mainly in the presence/absence of certain genomic regions, the diversity of IS families, and the presence/absence of the IS elements in a given genomic context. This indicates that the major evolutionary events modifying these two symbiotic plasmids were the gains/losses of genomic regions and the losses/translocations of IS elements. In contrast, the chromosomes and other plasmids did not contain IS in the same genomic contexts. Comparison of the IS elements in the chromosomes and plasmids indicated that the symbiotic plasmids may have originated externally and harbored IS that thereafter transposed to the chromosomes. Both CFN42 and CIAT652 harbored IS elements that were 100% identical in sequence between the chromosome and pSym. If these plasmids have a recent origin (as suggested by our population genetic analyses), this would seem to indicate that the IS had recently moved to the chromosome in both strains, in the manner of an infection-expansion-extinction cycle (42).
FIG. 3.
Dot plot graphic of pSym from R. etli CFN42 and CIAT652. The diagonal red lines represent DNA regions that were aligned between the two replicons. The blue and red points represent small DNA regions, in many cases repeated sequences, that were aligned in different regions of each replicon.
DISCUSSION
In order to examine the evolutionary dynamics of IS elements in relation to the evolutionary history of the chromosomal and symbiotic plasmid genes, we compared the distributions and genetic diversity of these sequences in three R. etli populations. First, we used the concatenated alignment of the chromosomal glyA and dnaB genes to perform a phylogenetic reconstruction and found that the three tested R. etli populations were almost completely differentiated according to their geographic origins. Moreover, they were clearly differentiated from the R. gallicum strains, which formed a fourth clade. The close phylogenetic relationship between the Michoacan-Guanajuato and Spanish populations was highlighted by the individual phylogenies of dnaB and glyA, especially the one based on glyA, where several strains appeared to be intermingled in both populations. This pattern might be explained by recombination events, which were assessed using the RDP3 software (see Table S4 in the supplemental material). However, another possibility is that the intermingled strains from Guanajuato could be closely related to migrant strains isolated from the Spanish population. The phylogeny made with the concatenated alignment of dnaB and glyA was used as the reference for mapping the IS distributions.
Comparison of the IS presence/absence profiles of the R. etli populations to that of the CFN42 reference strain showed that the most conserved IS were on the symbiotic plasmid. Because the majority of the IS elements did not maintain their genomic contexts, the presence/absence profiles were inconsistent with the housekeeping gene phylogeny; some strains that were closely related in the phylogeny had very different IS profiles. For example, most of the Puebla strains had IS profiles more similar to that of the model strain, CFN42, which belonged to the Guanajuato-Michoacan clade. These data demonstrate that the distribution of IS elements between and within populations is strain specific, in concordance with the HGT and transpositional dynamics of IS (4).
We analyzed two of the IS elements, ISRe12 and ISRel4 (both from pSym), in detail, as they were the most conserved, in terms of their genomic contexts, across the three R. etli populations. This pattern may indicate that some selective advantage is conferred by the presence of these IS elements, or it could be a consequence of the apparently recent origin of the pSym plasmid. Other studies on the evolutionary dynamics of IS elements have offered different explanations for the maintenance of such elements in populations or strains of the same species. For example, a report on IS in strains of Helicobacter pylori from different geographical origins suggested that IS are ancient components of the H. pylori gene pool and that they evolve at approximately the same rate as normal chromosomal genes (18). A paper on IS within a single population of the hyperthermophilic archaeon Pyrococcus suggested that the high frequencies of some IS were due to genetic drift (8). Neither of these explanations seems to be applicable to the IS in R. etli examined here. Instead, we think that the IS have undergone rapid turnover in the population and that the contextual and sequence conservation of ISRel2 and ISRel4 should be viewed within the evolutionary history of pSym.
Due to the transpositional nature of IS elements, they are not expected to be found at the same genomic site in populations of different geographical origin. However, here, we observed such conservation for ISRel2 and ISRel4. One explanation for our observation is that that the IS might have been frequently found in the same genomic context as a consequence of the small population size of the R. etli collections analyzed in this work. The probability of finding several strains with the same IS in the same location is lower for large populations than for small populations (38). In our case, however, this may not be the best explanation, given the characteristics of the three populations (see Materials and Methods) and the geographical distances separating the collection sites. Another possibility is that these IS elements were recently acquired by the studied R. etli strains. Comparisons among these IS elements, nodC, and the chromosomal genes suggested that the pSym sequences recently entered the gene pools of the three R. etli populations. Population genetic analyses of the pSym sequences showed that they had the following features: (i) low π values, (ii) Dxy and π values that imply no clear differentiation, (iii) several different shared alleles, and (iv) nonsignificant results from their genetic differentiation tests. In contrast, opposing results were obtained for the chromosomal genes, glyA and dnaB. Thus, the differences noted in our population genetic analyses suggest that the chromosome and pSym have different evolutionary histories in the studied R. etli populations. These findings, along with the high sequence conservation found in the pSym sequences of CFN42 and CIAT652, support the proposition that pSym could have a recent origin in R. etli (14). However, the pSym sequences from CFN42 and CIAT652 have clear differences in terms of the presence/absence of certain genomic regions (some as large as 50 kb) and IS elements; notably, the gains and losses of IS elements appear to be the main events differentiating these two pSyms.
A prior analysis of the complete genomes of R. etli CFN42 and R. etli CIAT652 demonstrated that the IS do not disrupt genes (with one exception) and that these strains appear to harbor the same pSym plasmid (15). Horizontal transfer of this plasmid to R. etli CFN42 and R. etli CIAT652 could explain the asymmetric distribution of IS, which were found only in these plasmids and in the chromosome. Some IS elements probably moved from the plasmids to the chromosome, but they do not appear to have moved to other plasmids. The asymmetrical distribution of IS in the chromosome and plasmids could be due to the sizes of the different replicons in the R. etli CFN42 genome (41). In the chromosome of CFN42, there are six probable recent transposition events of IS elements (100% identity between the chromosomal IS and another IS in the pSym or conjugative plasmid). This implies a transposition rate of 1 IS per 730 kb, which is greater than the size of any other replicon of CFN42 (14), so the asymmetrical distribution could be the product of the chromosome size.
The complete genomes of the two strains of R. etli, CFN42 and CIAT652, revealed the presence of different IS family members in different copy numbers. It has been proposed that when new IS elements enter the genome, they actively transpose to different genomic positions (42). This hypothesis may help explain the behavior of some of the IS in R. etli CFN42. For instance, members of the IS66 family, which is a very common IS family in rhizobial species (27), were found in several copies in the pSym plasmids of CFN42 and CIAT652, and IS66 copies with 100% nucleotide sequence identity could also be found in the chromosome. The most plausible explanation for this finding is that the IS originally entered the genome via plasmids and were then transposed to the chromosome.
The present work provides useful new insights into differences of the distribution and genomic context maintenance of IS elements in natural populations of R. etli. Although the pSym in these populations seems to be of relatively recent origin, the harbored IS appear to be active elements that may participate in the genomic plasticity of these organisms. Our comparisons of the two genomes of R. etli and across the natural populations provide further evidence that different IS can rapidly expand when they arrive in a new host genome, as recently proposed by Wagner (43).
Supplementary Material
Acknowledgments
We thank Santiago Castillo for critical reading of the mnuscript. We thank Susana Brom (Universidad Nacional Autónoma de México) for providing the Spanish strains, Antonio Cruz for providing those from Puebla and Michoacan-Guanajuato, and J. Espíritu for technical and computational assistance.
This work was supported by grants from CONACyT (grant U4633), PAPIIT-UNAM (grant IN223005), and the Natural Environment Research Council. L.L. was a recipient of a CONACyT scholarship.
L.L. was responsible for data analysis and manuscript preparation. L.L. and V.G. were responsible for the experimental design. I.H.-G. was responsible for bioinformatics analysis. P.B. and R.I.S. were responsible for DNA sequencing. L.L., V.S., J.P.W.Y., G.D., and V.G. were responsible for discussion of the data.
Footnotes
Published ahead of print on 30 July 2010.
Supplemental material for this article may be found at http://aem.asm.org/.
REFERENCES
- 1.Altschul, S., T. Madden, A. Schaffer, J. Zhang, Z. Zhang, W. Millar, and D. Lipman. 1997. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 25:3389-3402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Amadou, C., G. Pascal, S. Mangenot, M. Glew, C. Bontemps, D. Capela, S. Carrere, S. Cruveiller, C. Dossat, A. Lajus, M. Marchetti, V. Poinsot, B. Rouy, Z. Servin, M. Saad, C. Schenowitz, V. Barbe, J. Batut, C. Medigue, and C. Masson-Boivin. 2008. Genome sequence of the beta-rhizobium Cupriavidus taiwanensis and comparative genomics of rhizobia. Genome Res. 18:1472-1483. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Bailly, X., I. Olivieri, S. De Mita, J. C. Cleyet-Marel, and G. Bena. 2006. Recombination and selection shape the molecular diversity pattern of nitrogen-fixing Sinorhizobium sp. associated to Medicago. Mol. Ecol. 15:2719-2734. [DOI] [PubMed] [Google Scholar]
- 4.Chandler, M., and J. Mahillon. 2002. Insertion sequences revisited, p. 305-366. In L. Craig et al. (ed.), Mobile DNA II. ASM Press, Washington, DC.
- 5.Crossman, L. C., S. Castillo-Ramírez, C. McAnnula, L. Lozano, G. S. Vernikos, J. L. Acosta, Z. F. Ghazoui, I. Hernández-González, G. Meakin, A. W. Walker, M. F. Hynes, J. P. Young, J. A. Downie, D. Romero, A. W. Johnston, G. Dávila, J. Parkhill, and V. González. 2008. A common genomic framework for a diverse assembly of plasmids in the symbiotic nitrogen fixing bacteria. PloS ONE 3:e2567. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Doolittle, W. F., and C. Sapienza. 1980. Selfish genes, the phenotype paradigm and genome evolution. Nature 284:601-603. [DOI] [PubMed] [Google Scholar]
- 7.Edgar, R. 2004. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinform. 5:113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Escobar-Páramo, P., S. Ghosh, and J. DiRuggiero. 2005. Evidence for genetic drift in the diversification of a geographically isolated population of the hyperthermophilic archaeon Pyrococcus. Mol. Biol. Evol. 22:2297-2303. [DOI] [PubMed] [Google Scholar]
- 9.Filée, J., P. Siguier, and M. Chandler. 2007. Insertion sequence diversity in Archea. Microbiol. Mol. Biol. Rev. 71:121-157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Flores, M., L. Morales, M. Avila, P. Bustos, V. González, D. Garcia, Y. Mora, X. Guo, J. Collado-Vides, D. Piñero, G. Davila, J. Mora, and R. Palacios. 2005. Diversification of DNA sequences in the symbiotic genome of Rhizobium etli. J. Bacteriol. 187:7185-7192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Freiberg, C., R. Fellay, A. Bairoch, W. J. Broughton, A. Rosenthal, and X. Perret. 1997. Molecular basis of symbiosis between Rhizobium and legumes. Nature 387:394-401. [DOI] [PubMed] [Google Scholar]
- 12.Galibert, F., T. M. Finan, S. R. Long, A. Puhler, P. Abola, F. Ampe, F. Barloy-Hubler, M. J. Barnett, A. Becker, P. Boistard, G. Bothe, M. Boutry, L. Bowser, J. Buhrmester, E. Cadieu, D. Capela, P. Chain, A. Cowie, R. W. Davis, S. Dreano, N. A. Federspiel, R. F. Fisher, S. Gloux, T. Godrie, A. Goffeau, B. Golding, J. Gouzy, M. Gurjal, I. Hernandez-Lucas, A. Hong, L. Huizar, R. W. Hyman, T. Jones, D. Kahn, M. L. Kahn, S. Kalman, D. H. Keating, E. Kiss, C. Komp, V. Lelaure, D. Masuy, C. Palm, M. C. Peck, T. M. Pohl, D. Portetelle, B. Purnelle, U. Ramsperger, R. Surzycki, P. Thebault, M. Vandenbol, F. J. Vorholter, S. Weidner, D. H. Wells, K. Wong, K. C. Yeh, and J. Batut. 2001. The composite genome of the legume symbiont Sinorhizobium meliloti. Science 293:668-672. [DOI] [PubMed] [Google Scholar]
- 13.Gasca, J. 2004. Genética de poblaciones de Rhizobium etli asociado a Phaseolus vulgaris en dos localidades del bajío mexicano. B.S. thesis. Istituto de Ecología, Universidad Nacional Autónoma de México, Mexico DF, Mexico.
- 14.González, V., R. Santamaría, P. Bustos, I. Hernandez-González, A. Medrano-Soto, G. Moreno-Hagelsieb, S. Janga, M. Ramirez, V. Jiménez-Jacinto, J. Collado-Vides, and G. Davila. 2006. The partitioned Rhizobium etli genome: genetic and metabolic redundancy in seven interacting replicons. Proc. Natl. Acad. Sci. U. S. A. 103:3834-3839. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.González, V., J. L. Acosta, R. Santamaría, P. Bustos, J. L. Fernandez, I. Hernandez-González, R. Diza, M. Flores, R. Palacios, J. Mora, and G. Davila. 2010. Conserved symbiotic plasmid DNA sequences in the multireplicon pangenomic structure of Rhizobium etli. Appl. Environ. Microbiol. 76:1604-1614. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Guindon, S., and O. Gascuel. 2002. Efficient biased estimation of evolutionary distances when substitution rates vary across sites. Mol. Biol. Evol. 19:534-543. [DOI] [PubMed] [Google Scholar]
- 17.Huson, D. H., and D. Bryant. 2006. Application of phylogenetic networks in evolutionary studies. Mol. Biol. Evol. 23:254-267. [DOI] [PubMed] [Google Scholar]
- 18.Kalia, A., A. K. Mukhopadhyay, G. Dailide, Y. Ito, T. Azuma, B. C. Wong, and D. E. Berg. 2004. Evolutionary dynamics of insertion sequences in Helicobacter pylori. J. Bacteriol. 186:7508-7520. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Kaneko, T., Y. Nakamura, S. Sato, E. Asamizu, T. Kato, S. Sasamoto, A. Watanabe, K. Idesawa, A. Ishikawa, K. Kawashima, T. Kimura, Y. Kishida, C. Kiyokawa, M. Kohara, M. Matsumoto, A. Matsuno, Y. Mochizuki, S. Nakayama, N. Nakazaki, S. Shimpo, M. Sugimoto, C. Takeuchi, M. Yamada, and S. Tabata. 2000. Complete genome structure of the nitrogen-fixing symbiotic bacterium Mesorhizobium loti. DNA Res. 7:331-338. [DOI] [PubMed] [Google Scholar]
- 20.Kaneko, T., Y. Nakamura, S. Sato, K. Minamisawa, T. Uchiumi, S. Sasamoto, A. Watanabe, K. Idesawa, M. Iriguchi, K. Kawashima, M. Kohara, M. Matsumoto, S. Shimpo, H. Tsuruoka, T. Wada, M. Yamada, and S. Tabata. 2002. Complete genomic sequence of nitrogen-fixing symbiotic bacterium Bradyrhizobium japonicum USDA110. DNA Res. 9:189-197. [DOI] [PubMed] [Google Scholar]
- 21.Kidwell, M. G., and D. R. Lisch. 2002. Transposable elements as sources of genomic variation, p. 305-366. In L. Craig et al. (ed.), Mobile DNA II. ASM Press, Washington, DC.
- 22.Korber, B. 2000. HIV signature and sequence variation analysis, p. 55-72. In A. G. Rodrigo and G. H. Learn (ed.), Computational analysis of HIV molecular sequences. Kluwer Academic Publishers, Dordrecht, Netherlands.
- 23.Kurtz, S., A. Phullippy, A. Delcher, M. Smoot, M. Shumway, C. Antonescu, and S. Salzberg. 2004. Versatile and open software for comparing large genomes. Genome Biol. 5:R12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Kwak, M., and P. Gepts. 2009. Structure of genetic diversity in the two major gene pools of common bean (Phaseolus vulgaris L., Fabaceae). Theor. Appl. Genet. 118:979-992. [DOI] [PubMed] [Google Scholar]
- 25.Lenski, R., C. Winkwirth, and M. Riley. 2003. Rates of DNA sequence evolution in experimental populations of Escherichia coli during 20,000 generations, J. Mol. Evol. 56:498-508. [DOI] [PubMed] [Google Scholar]
- 26.Librado, P., and J. Rozas. 2009. DnaSP v5: a software for comprehensive analysis of DNA polymorphism data. Bioinformatics 25:1451-1452. [DOI] [PubMed] [Google Scholar]
- 27.Mahillon, J., C. Léonard, and M. Chandler. 1999. IS elements as constituents of bacterial genomes. Res. Microbiol. 150:675-687. [DOI] [PubMed] [Google Scholar]
- 28.Martin, D. P., C. Williamson, and D. Posada. 2005. RDP2: recombination detection and analysis from sequence alignments. Bioinformatics 21:260-262. [DOI] [PubMed] [Google Scholar]
- 29.Papadopoulos, D., D. Schneider, J. Meier-Eis, W. Arber, R. E. Lenski, M. Blot, et al. 1999. Genomic evolution during a 10,000-generation experiement with bacteria. Proc. Natl. Acad. Sci. U. S. A. 96:3807-3812. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Pérez-Mendoza, D., A. Dominguez-Ferreras, S. Munoz, M. J. Soto, J. Olivares, S. Brom, L. Girard, J. A. Herrera-Cervera, and J. Sanjuan. 2004. Identification of functional mob regions in Rhizobium etli: evidence for self-transmissibility of the symbiotic plasmid pRetCFN42d. J. Bacteriol. 186:5753-5761. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Posada, D., and K. A. Crandall. 1998. MODELTEST: testing the model of DNA substitution. Bioinformatics 14:817-818. [DOI] [PubMed] [Google Scholar]
- 32.Ramsay, J. P., J. T. Sullivan, G. S. Stuart, I. L. Lamont, and C. W. Ronson. 2006. Excision and transfer of the Mesorhizobium loti R7A symbiosis island requires an integrase IntS, a novel recombination directionality factor RdfS, and a putative relaxase RlxS. Mol. Microbiol. 62:723-734. [DOI] [PubMed] [Google Scholar]
- 33.Rodríguez-Navarro, D., A. Buendia, M. Camacho, M. Lucas, and C. Santamaria. 2000. Characterization of Rhizobium spp. Bean isolates from south-west Spain. Soil Biol. Biochem. 32:1601-1613. [Google Scholar]
- 34.Schneider, D., and R. Lenski. 2004. Dynamics of insertion sequence elements during experimental evolution of bacteria. Res. Microbiol. 155:319-327. [DOI] [PubMed] [Google Scholar]
- 35.Siguier, P., J. Filée, and M. Chandler. 2006. Insertion sequences in prokaryotic genomes. Curr. Opin. Microbiol. 9:1-6. [DOI] [PubMed] [Google Scholar]
- 36.Silva, C., P. Vinuesa, L. Eguiarte, E. Martínez-Romero, and V. Souza. 2003. Rhizobium etli and Rhizobium gallicum nodulate common bean (Phaseolus vulgaris) in a traditionally managed milpa plot in Mexico: population genetics and biogeographic implications. Appl. Environ. Microbiol. 69:884-893. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Silva, C., P. Vinuesa, L. Eguiarte, V. Souza, and E. Martínez-Romero. 2005. Evolutionary genetics and biogeographic structure of Rhizobium gallicum sensu lato, a widely distributed bacterial symbiont of diverse legumes. Mol. Ecol. 14:4033-4050. [DOI] [PubMed] [Google Scholar]
- 38.Slatkin, M. 1985. Genetic differentiation of transposable elements under mutation and unbiased gene conversion. Genetics 110:145-158. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Sullivan, J. T., and C. W. Ronson. 1998. Evolution of rhizobia by acquisition of a 500-kb symbiosis island that integrates into a phe-tRNA gene. Proc. Natl. Acad. Sci. U. S. A. 95:5145-5149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Sullivan, J. T., J. R. Trzebiatowski, R. W. Cruickshank, J. Gouzy, S. D. Brown, R. M. Elliot, D. J. Fleetwood, N. G. McCallum, U. Rossbach, G. S. Stuart, J. E. Weaver, R. J. Webby, F. J. De Bruijn, and C. W. Ronson. 2002. Comparative sequence analysis of the symbiosis island of Mesorhizobium loti strain R7A. J. Bacteriol. 184:3086-3095. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Touchon, M., and E. Rocha. 2007. Causes of insertion sequences abundance in prokaryotic genomes. Mol. Biol. Evol. 4:969-981. [DOI] [PubMed] [Google Scholar]
- 42.Wagner, A. 2006. Periodic extinctions of transposable elements in bacterial lineales: evidence from intragenomic variation in multiple genomes. Mol. Biol. Evol. 23:723-733. [DOI] [PubMed] [Google Scholar]
- 43.Young, J. P., L. C. Crossman, A. W. Johnston, N. R. Thomson, Z. F. Ghazoui, K. H. Hull, M. Wexler, A. R. Curson, J. D. Todd, P. S. Poole, T. H. Mauchline, A. K. East, M. A. Quail, C. Churcher, C. Arrowsmith, I. Cherevach, T. Chillingworth, K. Clarke, A. Cronin, P. Davis, A. Fraser, Z. Hance, H. Hauser, K. Jagels, S. Moule, K. Mungall, H. Norbertczak, E. Rabbinowitsch, M. Sanders, M. Simmonds, S. Whitehead, and J. Parkhill. 2006. The genome of Rhizobium leguminosarum has recognizable core and accessory components. Genome Biol. 7:R34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Zeigler, D. R. 2003. Gene sequences useful for predicting relatedness of whole genomes in bacteria. Int. J. Syst. Evol. Microbiol. 53:1893-1900. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



