Abstract
The spontaneous deamination of cytosine produces uracil mispaired with guanine in DNA, which will produce a mutation, unless repaired. In all domains of life, uracil-DNA glycosylases (UDGs) are responsible for the elimination of uracil from DNA. Thus, UDGs contribute to the integrity of the genetic information and their loss results in mutator phenotypes. We are interested in understanding the role of UDG genes in the evolutionary variation of the rate and the spectrum of spontaneous mutations. To this end, we determined the presence or absence of the five main UDG families in more than 1,000 completely sequenced genomes and analyzed their patterns of gene loss and gain in eubacterial lineages. We observe nonindependent patterns of gene loss and gain between UDG families in Eubacteria, suggesting extensive functional overlap in an evolutionary timescale. Given that UDGs prevent transitions at G:C sites, we expected the loss of UDG genes to bias the mutational spectrum toward a lower equilibrium G + C content. To test this hypothesis, we used phylogenetically independent contrasts to compare the G + C content at intergenic and 4-fold redundant sites between lineages where UDG genes have been lost and their sister clades. None of the main UDG families present in Eubacteria was associated with a higher G + C content at intergenic or 4-fold redundant sites. We discuss the reasons of this negative result and report several features of the evolution of the UDG superfamily with implications for their functional study. uracil-DNA glycosylase, mutation rate evolution, mutational bias, GC content, DNA repair, mutator gene.
INTRODUCTION
Thanks to advances in high-throughput sequencing, we can estimate the spontaneous mutation rate of many eukaryotes with startling precision. Whole-genome sequencing of mutation accumulation lines has allowed the mutational spectrum for several model organisms to be characterized and compared in great detail (fig. 1; Denver et al. 2009, Keightley et al. 2009, Ossowski et al. 2010). Knowledge of the mutation rate to different base substitutions is critical to the detection of positive selection, the quantification of selective coefficients, and the estimation of effective population sizes.
FIG. 1.
Conditional spontaneous base substitution rates per base per generation for Drosophila melanogaster (white, Keightley et al. 2009), Caenorhabditis elegans (pattern, Denver et al. 2009), and Arabidopsis thaliana (black, Ossowski et al. 2010). Horizontal lines indicate total mutation rates for each species with their standard errors in grey. Total mutation rates are equal to the average of the sum of mutation rates at A:T sites and the sum of mutation rates at G:C sites, weighted by the proportion of each kind of sites in the genome. Data for D. melanogaster and for C. elegans were derived from the original publications assuming that the proportion of A:T sites surveyed is the same as the genomic proportion of A + T. Data for D. melanogaster were not corrected for a small downward bias described in Keightley et al. (2009).
Three main factors have been proposed to explain the diversity of mutation rates among lineages: generation time, metabolic rate, and quality of DNA-repair machinery (Baer et al. 2007). Despite the fact that the organisms represented in figure 1 have different numbers of germ-line cell divisions per generation (8.5 in Caenorhabditis elegans, Wilkins 1992; 36 in Drosophila melanogaster, Drost and Lee 1995; and ∼50 in Arabidopsis thaliana, Hoffman et al. 2004), the mutation rates per generation of most kinds of base substitutions are remarkably similar among them, although significant interspecific differences exist for the mutation spectrum. Therefore, in addition to generation time, other factors such as the presence and quality of specific pathways of DNA modification and repair must be playing an important role in the diversity of mutation rates and spectra. For example, the relatively high rate of transitions at G:C sites in A. thaliana can be largely, if not completely, explained by the levels of cytosine methylation in this species (Ossowski et al. 2010).
In parallel to the estimation of spontaneous mutation rates in several species, others and we are following a phylogenomic approach, pioneered by Eisen and Hanawalt (1999), to characterize the DNA-repair machinery present in as many species as possible, one pathway at a time (Denver et al. 2003, Lin et al. 2007, Lucas-Lledó and Lynch 2009). One aim of this effort is to gain insight into the mechanisms underlying differences in mutation rates among species. Here, we focus on the uracil excision repair pathway, which removes uracils from DNA. Uracils appear in DNA from two different sources. During DNA replication, polymerases introduce uracil in front of adenine with a frequency similar to the ratio of the concentrations of dUTP and dTTP (Tye et al. 1978). Properly paired with adenine, uracil is not mutagenic, although it can disrupt the binding sites of transcription factors and other DNA-binding proteins (Ivarie 1987). Uracils can also appear in DNA, mispaired with guanine, by deamination of cytosine, which is one of the most frequent spontaneous mutagenic reactions (Lindahl and Nyberg 1974). Half the progeny of a DNA molecule with an U:G mismatch will carry a mutation. The enzymes that specifically target uracil and start the uracil excision repair pathway are called uracil-DNA glycosylases (UDGs). They catalyze the excision of uracil, producing an abasic site and free uracil.
We hypothesize that the mutation rate, the transition–transversion ratio, and the mutational bias toward A + T (Lynch 2010) should be increased in lineages where UDG activity has been evolutionarily lost. To test this hypothesis, we first performed a presence/absence analysis of the five main families of UDGs in completely sequenced genomes of Archaea, Eubacteria, and Eukarya, and then we tested the correlation between the presence or absence of each family with the G + C content in intergenic and 4-fold redundant sites of 779 eubacterial genomes. The rationale behind this test is that intergenic and 4-fold redundant sites may reflect mutational biases more faithfully than more functionally constrained sites. We also study the evolutionary interaction between the four UDG families widely spread among Eubacteria and show that, in a large temporal scale, almost any of them can compensate for the absence of the other.
The five UDG families considered in this study are usually numbered in order of discovery (Pearl 2000). Family 1 is present in Eubacteria, Eukarya and in some eukaryotic viruses. Members of family 1 are, by far, the most efficient UDGs (Slupphaug et al. 1995). Family 2 UDGs are also found in Eubacteria and Eukarya, where they seem to have complementary functions to those of family 1. For example, family 2 UDGs target uracil (and also thymine in Eukaryotes) only when mismatched with guanine (Neddermann and Jiricny 1994, Gallinari and Jiricny 1996), also target ethenocytosine (Saparbaev and Laval 1998), are upregulated during stationary phase in Escherichia coli (Mokkapati et al. 2001) and prefer CpG sites in human cells (Abu and Waters 2003). Family 3 was originally described in animals and characterized by the ability to excise 5-hydroxymethyluracil and uracil from single-stranded DNA (Haushalter et al. 1999, Boorstein et al. 2001), although family 3 UDGs were later shown to be active on double-stranded DNA and present in a few eubacterial species (Pettersen et al. 2007). Family 4 is present in Eubacteria, Archaea, and some phages. It is thermostable and functionally similar to family 1, as far as it has been characterized (Koulis et al. 1996, Sandigursky and Franklin 1999, Sandigursky and Franklin 2000, Sandigursky et al. 2001). Family 5 is also present in Eubacteria and Archaea but absent in Eukarya. It is also thermostable. A family 5 UDG from Mycobacterium tuberculosis complements a strain of E. coli deficient in its family 1 (Srinath et al. 2007). The five families have a common evolutionary origin, although their homology is hardly recognized at the sequence level, even in the two most conserved motifs of the catalytic pocket (Aravind and Koonin 2000). Other enzymes with UDG activity have evolved in some lineages but have been excluded from this study, such as the Archaea-specific UDGs from the helix-hairpin-helix family (Chung et al. 2003), and the MBD4 homologs found in mammals (Hendrich et al. 1999). We also excluded viral UDGs from our analysis.
MATERIALS AND METHODS
Phylogenies
Using queries from the five UDG families identified in the literature, we performed PSI-Blast searches (Altschul et al. 1997) against the nonredundant protein database in the National Center for Biotechnology Information (NCBI) server until convergence. Results for families 2, 4, and 5 overlapped, and they were combined in a single file. An alignment with Muscle (Edgar 2004) and a preliminary neighbor joining tree was built for these three families, and family 2 was then defined to include a clade (family 2a) with all proteins previously annotated as family 2 UDGs, and a sister clade (family 2b) of only prokaryotic sequences that did not have any functional annotation and that were closer to family 2 proteins than to either family 4 or 5. All the remaining proteins in the alignment, including families 4 and 5, and other more distantly related sequences, were kept together and labeled “extended family 4”. The resulting four alignments (family 1, family 2, family 3, and extended family 4) were manually curated with BioEdit (Hall 1999). We removed from the alignments the less reliable sites with Gblocks (Talavera and Castresana 2007), and we used ProtTest (Abascal et al. 2005) to determine the most likely models of amino acid evolution, which were then used to infer four maximum likelihood (ML) phylogenies with RAxML (Stamatakis 2006). An approximation to bootstrap sampling was used to determine the branch support (Stamatakis et al. 2008).
To analyze the patterns of evolutionary loss of UDG genes in Eubacteria, we needed a phylogenetic tree of the eubacterial species with completely sequenced genomes. To this end, we downloaded the aligned sequences of their small ribosomal subunit from the RDP II (Cole et al. 2007) or the Greengenes databases (DeSantis et al. 2006). To merge sequences from the two databases, we aligned their profiles with muscle, and then we manually checked the alignment with BioEdit (Hall 1999). We also trimmed the ends and used Gblocks to purge from the alignment the unreliably aligned regions. The final alignment consisted of 832 sequences and was 1,437 nucleotides long. We ran 240 replicates of the search of the ML tree with RAxML, using the GTRMIX model of molecular evolution. We kept the 100 most likely trees to represent the uncertainty about the true phylogeny. We used Archaea as outgroup.
Protein Domain Analysis
We searched for additional protein domains from the Pfam database (Finn et al. 2006) in all full-length UDGs found in the PSI-Blast searches, using HMMER3 (Eddy 1998). We disregarded all domain predictions with an E value higher than 0.001.
Presence–Absence Data
Given the sensitivity of the PSI-Blast searches, we consider unlikely that any homologous protein with UDG activity present in the databases was missed. However, some UDGs could be missing from the protein databases if their genes were not annotated. To fill this gap, we undertook exhaustive TBlastN searches of all UDG families in completely sequenced genomes with or without a definitive assembly. We scanned over 1,000 genomes through either the microbial genomic database or the whole-genome shotgun database at NCBI (http://www.ncbi.nlm.nih.gov, Last accessed December 8, 2010), or through the species-specific databases at the Broad Institute (http://www.broadinstitute.org/, Last accessed December 8, 2010), the Genome Center at Washington University (http://genome.wustl.edu, Last accessed December 8, 2010), or the DOE Joint Genome Institute (http://www.jgi.doe.gov/, Last accessed December 8, 2010), depending on where each genome was available. To search for a UDG family in a genome, we used as queries all the sequences of that UDG family already present in our alignment that belonged to the lowest taxonomic rank of that genome's classification in which the presence of that protein family was already known.
A presence was tentatively called whenever at least one of the queries retrieved a hit with an E value lower than 0.001. Best hits with E values higher than were manually checked, and they were determined to be false positives if any of the following were true: Homology was limited to a portion of the protein lower than 70%; an additional domain in the query, other than UDG, was causing the hit; or the best reciprocal hit on the “nr” protein database was not a UDG. In addition, we excluded several hits suspected to be caused by bacterial contamination of eukaryotic contigs. Given that all UDG families analyzed share a common origin (Aravind and Koonin 2000), we also corrected a few cases of cross-detection between families. To avoid false negatives, we excluded from the analysis several eukaryotic genomes supposed to be in an assembly stage but with sequencing coverage depths lower than 2.
Test of Correlated Evolution among UDG Families
We tested whether the probabilities of loss and gain of a UDG family were independent of the presence of other UDG families in Eubacteria. We used the program BayesTraits (Pagel and Meade 2006), which implements the comparison of dependent and independent models of evolution of pairs of binary traits developed by Pagel (1994). We ran Bayesian analyses with reversible jumps between models to compare dependent and independent models of evolution between families 1 and 2, 1 and 4, and 2 and 4. Because the presence of family 3 in Eubacteria is incidental (see below), we did not include it in the analyses. The use of more than one phylogeny to represent its uncertainty was not feasible due to extremely low rates of acceptance of parameter value change proposals in the Markov chains, resulting in a very inefficient exploration of the parameter space. We therefore used the most likely Eubacterial phylogeny inferred before, with branch lengths in units of substitutions per site. Because we expected both gains and losses of UDG genes to happen at a slower rate than the fixation of base substitutions, we used exponential priors of mean 1.0 for all the parameters. The Markov chains were sampled every 1,000 iterations, during 65 million iterations after a burn in period of 500,000 iterations. Under these settings, no correlation between successive samples of parameter values was observed.
Test of Correlation between Presence–Absence of UDG Families and G + C Content
We used a python script (R.M.) to scan all GenBank contigs composing the complete genomes analyzed. In those contigs with gene predictions, the script determined the G + C content in coding, intergenic, and 4-fold redundant sites, as well as the overall G + C content. Plasmids were excluded from the analysis. We used the aotf routine of the Phylocom 4.1 package (Webb et al. 2008) to test whether the presence of families 1, 2, or extended family 4 was correlated with the G + C content at 4-fold redundant sites among 779 eubacterial species. The best of the phylogenies found in the ML searches was used, with untransformed branch lengths, and the logarithmic transformation of the G + C proportion was used because it removed any correlation between the absolute value of the standardized contrasts and their standard deviations (Garland et al. 1992) as tested with Mesquite (Maddison WP and Maddison DR 2010, Midford et al. 2010). A first run of aotf with presence–absence data set as a binary trait determined what nodes of the tree were informative about the relationship between G + C content and presence–absence of UDG. A second run, calculated the standardized contrasts for all the nodes as suggested in the Phylocom manual. A t-test was performed with the standardized contrasts of the informative nodes.
The presence of any UDG family is expected to increase the mutation equilibrium G + C content. However, the G + C content at 4-fold redundant sites in Eubacteria is also affected by codon usage bias (Sharp et al. 2005). Therefore, we also tested the effect of UDG families on the G + C content at intergenic sites. Unfortunately, the G + C content at intergenic sites did not fit the assumed Brownian model of evolution along the most likely Eubacterial phylogeny. We had to use the seventh most likely phylogeny, with branch lengths transformed as suggested by Pagel (1992). Both intergenic and coding G + C content fitted the Brownian model along this phylogeny.
Lastly, we wanted to control for other factors affecting the genomic G + C content, such as optimal growth temperature (Musto et al. 2006). Because optimal growth temperatures are not readily available for all the species included in this analysis, we assumed that the regression of intergenic G + C content on the G + C content at coding sites captured most of the variation due to genome-wide factors. The regression was calculated through the origin using the phylogenetically independent contrasts of all nodes in the tree. Sign tests and t-tests were applied to the residuals of the nodes previously determined to be informative about the effect of each UDG family.
RESULTS
Evolution of the UDG Families
Because they are present in all domains of life (fig. 2), UDGs must have appeared very early in evolution. Family 1 is the most conserved one. We failed to recover the likely monophyly of all eukaryotic family 1 UDGs with our ML phylogenetic analysis (fig. 2). Specifically, some highly divergent protozoan UDGs are grouped together with eubacterial UDGs with high bootstrap values, probably due to long branch attraction. Most family 1 UDGs are encoded by single-copy genes. The high degree of conservation of the catalytic motifs in this family suggests that all its members have the same function. Family 1 UDG domains are very rarely associated with other domains in the same protein. The main exceptions to this are the Actin domain in at least two fungal genera of the order Onygenales and the ART domain in three genera of Actinobacteria and Planctomycetes (supplementary table S2, Supplementary Material online). None of these domains are known to participate in DNA repair.
FIG. 2.
ML phylogenies of the main UDG families, and the sequence logos of their most conserved homologous motif. The branch color indicates eubacterial (black), archaeal (green), or eukaryotic (red) sequences. A putative family 6, not functionally described in the literature, is indicated with a question mark.
All proteins functionally characterized as family 2 UDGs are located in figure 2 in the clade labeled “2a.” Another clade, “2b,” is composed of proteins clearly more similar to family 2 UDGs than to any other UDG families, although none of them has been functionally characterized, to the best of our knowledge. Family 2 has been previously described only in Eubacteria and Eukarya, and we identify for the first time two Archaeal orthologs of subfamily 2a (accession numbers EET89853 and EEZ92451) and several others in subfamily 2b. Archaeal and eukaryotic sequences in subfamily 2b are scattered in the clade in positions that do not match their taxonomy. Although some of them may be artifactual or due to lateral gene transfer, we consider it very likely that the last universal common ancestor had at least one and probably two family 2 UDGs.
The ML phylogeny of family 3 also reveals the existence of two main clades: the cannonical family 3 UDGs, almost exclusively animal, and a eubacterial clade with some representatives from Bacteroidetes, Chlorobi, Actinobacteria, and Firmicutes. Among the animal family 3 UDGs, we found some nonanimal proteins: three from the moss Physcomitrella patens, and some from nonpathogenic eubacterial species, that belong to Proteobacteria, Planctomycetes, or Verrucomicrobia.
UDGs from families 4 and 5 can be reliably aligned together with many other sequences of uncertain function. Figure 2 depicts the ML phylogeny of all these proteins. Family 5 and another clade of UDG-like proteins (putative family 6) are supported by bootstrap values of 66% and 69%, respectively. However, family 4 UDGs form a highly polytomic ensemble together with unannotated proteins, many of which do not conserve the catalytic motifs. To simplify, we considered families 4 and 5, putative family 6, and all other related proteins as part of the “extended family 4” in all subsequent analyses.
Two family 4 UDGs have been independently acquired by the eukaryotes Micromonas sp. (Chlorophyta) and Paulinella chromatophora (Rhizaria) through their primary photosynthetic organelles. The annotation of many family 4 UDGs as “Phage SPO1 polymerase-related protein” is an unfortunate consequence of the existence of a polymerase in Bacillus phage SPO1 fused to a UDG domain related to family 4. Interestingly, plasmids in Yersinia pestis and in Salmonella enterica subsp. enterica serovar Typhi str. CT18, but not in other serovars of the same subspecies, also encode a polymerase III alpha subunit fused to a family 4 UDG domain (supplementary table S1, Supplementary Material online).
Supplementary table S1, Supplementary Material online shows what UDG families are present or absent in more than 1,000 completely sequenced genomes. Only 26 genomes have been observed without any UDG family, 14 of which are from Archaea (out of 63), 10 in Bacteria (out of 839), and 2 in Eukarya (out of 281). Most organisms with only one UDG family rely on either family 1, family 4, or family 5. The mosquitoes from the family Culicidae are the only known organisms that seem to rely only on family 3 UDGs to remove the uracil from their genomes. And family 2 is the only UDG family apparently available in the red alga Cyanidioschyzon merolae and in the parasitic alveolate Perkinsus marinus.
Comparative Functional Review
All UDG families share specificity for uracil in U:G mispairs and probably for uracil paired with adenine (see supplementary table S3, Supplementary Material online). All families, and especially families 2 and 3, have additional substrates, which may have become their primary targets, such as ethenocytosine or T:G mispairs in many family 2 UDGs.
A convenient measure of catalytic efficiency, to establish comparisons across species, would be the ratio between the maximum reaction rate, , and the Michaelis constant, (substrate concentration at which the reaction rate is half the maximum). Unfortunately, many studies do not report but only and the turnover, (maximum number of enzymatic reactions catalyzed per unit of time). Figure 3 represents the distributions of several measures of and for family 1 UDGs of four different species. Most experiments attribute a lower toward U:A (higher affinity) to E. coli's family 1 UDG than to its human ortholog (t-test P value = 0.012). Supplementary table S4, Supplementary Material online, displays the catalytic properties of the family 1 UDGs represented in figure 3 and also of some members of families 2, 3, and 4.
FIG. 3.

Summary of values of the Michaelis constant () and the rate of uracil excision (; see text for details) from U:A pairs by family 1 UDGs of different species found in the literature (see raw data and references in the Supplementary table S4, Supplementary Material online). Boxes are interquartile ranges; horizontal lines within the boxes are medians; whiskers extend to the most extreme data points that are no more than 1.5 times the interquartile range away from the box; data points further than 1.5 times the interquartile range from the box are drawn as empty circles. Horizontal lines without a box represent single values.
Table 1 shows how UDG deficiencies affect mutation rates across species. The inactivation of family 1 UDGs causes on average a 5.4 ± 2.3-fold increase in the overall mutation rate in Eubacteria, a 2.9 ±1.3-fold increase in mammals, and an 18.7 ± 20.1-fold increase in yeast, where family 1 is the only UDG family available. At G:C sites, the deficiency of family 1 UDG causes an average increase in mutation rate by a factor of 21.4 ± 12.1 in E. coli, 6.9 ± 0.3 in mammals, and 21.74 in yeast. When two or more UDG families are present in a genome, one of them (either family 1 or family 4) prevents most of the mutations, whereas the others contribute marginally to replication fidelity (see table 1). However, the role of the secondary UDGs in preventing mutations is increased in the absence of the main UDG.
Table 1.
Inflation of the Spontaneous Mutation Rate due to Deficiency of UDGs. Values are Ratios of Mutation Rates of UDG-Deficient and Wild-Type Strains, Estimated from Fluctuation Analysis, or Ratios of the Frequency of Mutants, If the Former Was not Available. K.O. Indicates What UDG Family or Families, among Those Present in the Wild Type, Have Been Knocked Out in Each Study. UDG Deficient Strains Are not Known to Be Defective in any other DNA-Repair Pathway
| Species | UDG Families |
Mutation Rate Inflation |
|||
| Wild Type | K.O. | G:C Sites | All Sites | Reference | |
| Escherichia coli | 1, 2 | 1? | 30.00 | 3.73 | Duncan and Weiss (1982) |
| 1 | — | 4.45 | Venkatesh et al. (2003) | ||
| 1 | — | 1.76 | Nordman and Wright (2008) | ||
| 1 | 12.86 | — | Lutsenko and Bhagwat (1999) | ||
| 2 | 1.00 | — | Lutsenko and Bhagwat (1999) | ||
| 1, 2 | 17.62 | — | Lutsenko and Bhagwat (1999) | ||
| Salmonella typhimurium | 1, 2 | 1 | — | 6.00 | Lind and Andersson (2008) |
| Helicobacter pylori | 1, 4 | 1 | — | 4.00 | Huang et al. (2006) |
| Mycobacterium smegmatis | 1, 2, 4, 5 | 1 | — | 9.20 | Kurthkoti et al. (2008) |
| 1 | — | 6.87 | Venkatesh et al. (2003) | ||
| Pseudomonas aeruginosa | 1 | 1 | — | 7.00 | Venkatesh et al. (2003) |
| Streptococcus pneumoniae | 1, 4 | ? a | 4.60 | 2.65 | Chen and Lacks (1991)b |
| Thermus thermophilusc | 4, 5 | 4 | 12.20 | — | Sakai et al. (2008) |
| 5 | 2.90 | — | Sakai et al. (2008) | ||
| 4, 5 | 55.90 | — | Sakai et al. (2008) | ||
| Saccharomyces cerevisiae | 1 | 1 | — | 41.57 | Burgers and Klein (1986) |
| 1 | 21.74 | 10.81 | Impellizzeri et al. (1991) | ||
| 1 | — | 3.80 | Guillet et al. (2006) | ||
| 1 | — | 2.04d | Chatterjee and Singh (2001) | ||
| Mus musculus | 1, 2, 3 | 1 | 7.11 | 3.79 | An et al. (2005) |
| 1 | — | 1.36 | Nilsen et al. (2000) | ||
| 3 | 3.00 | 2.07 | An et al. (2005) | ||
| 1, 3 | 18.11 | 21.16 | An et al. (2005) | ||
| Homo sapiens | 1, 2, 3 | 1e | 6.67 | 3.57 | Radany et al. (2000) |
Mutants selected with UDG activity lower than 0.5% of wild type.
We used Drake's method A0 (Drake 1991) to estimate the mutation rates.
Results for its optimal temperature, 70°C.
In mitochondria, average from two different assays.
UDG inhibited using the Ugi protein.
Correlated Rates of Gain and Loss of UDG Families
The patchy taxonomic distribution of some UDG families suggested that UDG genes have been evolutionarily lost and gained several times in Eubacterian lineages. Therefore, we used the patterns of gene gain and loss to study the functional relationships among UDG families along a large timescale. First, we combined the presence–absence data with a ML phylogeny of Eubacteria, built from the alignment of the 16S ribosomal subunit of 779 species with completely sequenced genomes, in which Archaea was used as an outgroup. With this data set, we performed Bayesian analyses to fit two kinds of models of the evolution of the genomic UDG composition: Independent models in which the probabilities of gain and loss of a family are unaffected by the presence of another family; and dependent models, where the presence of one family changes the probability of losing or gaining another one. Each specific model is defined by the number of parameters used to represent the rates of loss and gain of each family and by what rates are set to be equal. Due to computational limitations, only pairwise comparisons among families can be performed with currently available software. Specifically, we performed all pairwise comparisons among families 1, 2, and the extended family 4. The Markov chains were allowed to jump between dependent and independent models to estimate the posterior probability of dependent evolution between families. We found strong evidence for the dependent evolution of families 1 and 2 (Bayes factor, or ratio between posterior and prior odds of dependent evolution, 52.4) and of family 1 and extended family 4 (Bayes factor >157). In contrast, family 2 and the extended family 4 were shown to be lost or gained in Eubacterial lineages independently (Bayes factor, 0.01). The use of alternative, less likely, phylogenies gave similar results (data not shown).
Dependent evolution between UDG families imply that the presence of one family in a genome affects the chance of another family to be lost or gained. To determine exactly what rates are modified by the presence or absence of each family, we also compared the posterior probabilities of specific models of dependent evolution, taken from the frequency with which the Markov chain visited them, with their prior probabilities. Following Pagel and Meade (2006), we determined that the prior odds of any specific dependence is 4.11 because among all the models randomly sampled by the Markov chain, there are this many more models where a specific rate depends on the presence of the alternative family than models where the rate is equal under the presence or absence of the alternative family. Using this information, we determine that only two rates showed strong evidence of dependent evolution: The rate of loss of family 1 is dependent on the presence of family 2 (Bayes factor, 12.2), and the rate of loss of the extended family 4 is dependent on the presence of family 1 (Bayes factor >15,800).
Figure 4 represents the posterior distributions of the rates of loss of family 1 and of extended family 4 in the presence and the absence of families 2 and 1, respectively. It can be seen that family 1 and extended family 4 are very rarely lost in the absence of either family 2 or family 1, whereas the presence of family 2 increases the rate of loss of family 1, and the presence of family 1 increases the rate of loss of the extended family 4. To put it another way, the loss of family 2 enhances the preservation of family 1 (but not of family 4), and the loss of family 1 enhances the preservation of family 4.
FIG. 4.
Posterior distributions of the rates of loss of UDG family 1 and of extended family 4, in the absence or presence of families 2 or 1. The meanings of boxes, whiskers, and empty circles are the same as in figure 3.
Effect of UDG Families in the Evolution of G + C Content
The genomic G + C content evolves by mutation, drift, and natural selection. The removal of uracil from DNA by UDGs prevents G:C → A:T transitions, and therefore the presence of UDGs is expected not only to reduce the total mutation rate (see table 1) but also to balance the mutational spectrum in favor of a higher G + C composition. We hypothesized that lineages where a UDG family has been lost, the mutational spectrum will have changed in favor of a lower G + C content. To test this hypothesis, we calculated the G + C composition in intergenic and 4-fold redundant sites of 779 completely sequenced and annotated eubacterial genomes, and we used their ML phylogeny to calculate independent contrasts of G + C composition in all nodes. We then used the contrasts at nodes whose daughter branches differed in the presence of a UDG family to test the effect of the absence of that UDG family in the G + C composition. In all, 40 nodes were informative for family 1, 83 nodes for family 2, 18 nodes for family 3, and 55 for the extended family 4. We also combined the presence of family 1 and extended family 4 to test whether the absence of both of them affected the G + C composition, but only 9 nodes were informative because most eubacterial genomes have either one or the other. Within each family, the standardized independent contrasts at informative nodes represent the difference in G + C content between a lineage with that family and a lineage without it. Overall, a positive average was expected. None of the t-tests reported a significant departure from zero at the 0.05 level for the G + C content in 4-fold redundant sites (table 2). Sign tests were also performed, with similar results.
Table 2.
Average of Independent Contrasts of (logarithmic) G + C Content Between Lineages with or without a UDG Family, Their Standard Deviation (SD), Sample Size (N), and P value of a t-Test
| Family | Mean | SD | N | t-Test |
| 1 | – 0.0197 | 0.5180 | 40 | 0.8114 |
| 2 | 0.0993 | 0.9717 | 83 | 0.3547 |
| 3 | 0.0973 | 0.2826 | 18 | 0.1625 |
| 4 | 0.0454 | 0.4206 | 55 | 0.4265 |
| 1 or 4 | 0.3491 | 0.5000 | 9 | 0.5999 |
To avoid the confounding effect of codon usage bias, intergenic G + C content was also analyzed, using a suboptimal phylogeny that allowed the assumption of a Brownian model of evolution (see Methods). The t-tests yielded nonsignificant results for all families, indicating that intergenic G + C content is not different between species with and without a UDG family (data not shown). To account for the effect of optimal growth temperature and other factors affecting the genome-wide G + C content, we repeated the analysis using the residuals of the regression of intergenic G + C content on G + C content at coding sites. Again, none of the UDG families was associated with a significantly higher intergenic G + C content relative to the G + C content at coding sites.
DISCUSSION
We now have the ability to detect small differences in the genomic mutation rate and the mutational spectra among species, and we can aim to explain those differences. One of the types of base substitutions contributing to the higher mutation rate in D. melanogaster relative to C. elegans are G:C → A:T transitions (fig. 1). It is tempting to hypothesize that the lack of the main UDG family in the former (supplementary table S1, Supplementary Material online) could be at least part of the explanation, although there are other types of substitution with significantly different rates between the two species.
We showed that the presence of none of the UDG families was associated with a significantly higher G + C content at intergenic or 4-fold redundant sites (table 2) in contrast to what was expected from the known effect of UDG enzymes on the mutational spectrum. There are several reasons for this negative result. First, G + C content is the net outcome of mutation and selection effects. Selection for optimal codon usage (Sharp et al. 2005) and the optimal growth temperature (Musto et al. 2006) affect the G + C content at 4-fold redundant sites and limit our ability to detect the mutational effects. Second, a change in the mutational spectrum, such as the one expected after the loss of a UDG family, does not affect the G + C composition immediately, but through potentially long periods of time, whereas the new mutation selection–drift equilibrium is reached. Therefore, recent losses of UDG genes may have had a strong effect in the mutation rate and spectrum, but little effect in the G + C composition of intergenic and 4-fold redundant sites. Indeed, all complete UDG losses have happened relatively recently in lineages that have not diversified yet (supplementary table S1, Supplementary Material online) and may never diversify, unless they recover the ability to excise uracil. And third, given the scarcity of species where all UDG families have been lost, it is difficult to rule out the possibility of the loss of one UDG family being rapidly compensated by other UDG families still present in the lineage.
Despite the failure to detect a relationship between the absence of UDG families and the mutational spectrum of eubacterial lineages, we set the bases for more detailed studies in the future. We gathered important information about the functions and the evolution of the UDG superfamily. Our phylogenomic analysis underlines the essentiality of the UDG activity in all domains of life and suggests that the UDG superfamily may have been the main provider of UDG activity since the last universal common ancestor. The only obvious exception to this hypothesis are some families of Euryarchaeota, whose genomes are devoid of all UDG families. A different family of enzymes with UDG activity seems to have taken over the ancestral function of the UDG superfamily in those archaeas (Chung et al. 2003). In addition, archaeal DNA polymerases have the ability to stall replication 4–6 nucleotides before the presence of a uracil residue (Greagg et al. 1999), providing them with another line of defense against uracil.
Overall, either family 1 or families 4 or 5 seem to be essential, given the few examples of species without them (supplementary table S1, Supplementary Material online). The vast majority of thermophiles have either family 4 or 5, rather than (or sometimes in addition to) family 1, although families 4 and 5 are not restricted by any means to thermophiles. Our analysis of rates of gene loss and gain also supports the idea of a functionally equivalent role of families 1 and 4 or 5 (fig. 4). Results from UDG-deficiency experiments where more than one family was sequentially knocked out (table 1) show that either family 1 (E. coli, and Mus musculus) or family 4 (Thermus thermophilus) act as the main uracil remover, whereas UDGs from families 2, 3, or 5 have a secondary role in mutation avoidance in the presence of the former. It is likely that the difference between the main and the secondary UDGs is the ability of the former to interact with the replication machinery and to function during DNA synthesis in a highly processive fashion (Kavli et al. 2002, Dionne and Bell 2005, Kiyonari et al. 2008). The fusion of UDG domains and polymerase domains in phages and plasmids (supplementary table S2, Supplementary Material online) also underlines the tendency of some UDGs to become associated with DNA replication.
Although all UDG families share a common origin (Aravind and Koonin 2000) and uracil as a common substrate (supplementary table S3, Supplementary Material online), it has been argued that the main biological role of some members of families 2 and 3 is different from the excision of uracil from DNA in species where family 1 is also present (Saparbaev and Laval 1998, Lutsenko and Bhagwat 1999, Masaoka et al. 2003). Although that might well be the case, our finding that a few species lost all UDGs but those from families 2 or 3 strongly suggests that these UDG families can become the main providers of UDG activity, at least over an evolutionary timescale. The finding that the presence of family 2 significantly increases the chance of evolutionarily loss of family 1 in Eubacteria (fig. 4) strongly supports the idea that family 2 UDGs can complement the absence of family 1.
It is sometimes assumed that the ability of an enzyme to repair a damaged base must be adaptive, even when that damage is very rare (O'Brien 2006). Thus, it was thought that the main biological role of family 3 UDGs was the removal 5-hydroxymethyluracil (Boorstein et al. 1989), before its UDG activity was discovered (Boorstein et al. 2001); or that ethenocytosine was the main substrate of the family 2 UDG MUG from E. coli (Saparbaev and Laval 1998), even though ethenocytosine is not detected in E. coli DNA (Mokkapati et al. 2001). Given the large number of substrates of some UDG families (supplementary table S3, Supplementary Material online), we consider it unlikely that the specificity for all of them has been actively acquired by positive selection. Instead, we consider that a broad substrate specificity may be the result of relaxed selection, on the bases of two likely assumptions: First, that most neutral mutations expand the substrate specificity, rather than narrowing it, and second, that the lower the frequency of a substrate, the less efficient natural selection is to promote the specificity for that substrate. Under this scenario, natural selection can promote and keep a narrow substrate specificity of a DNA repair enzyme only if the substrate is frequent enough in DNA. When more than one enzyme repairs the same damage, as in the case of the UDG superfamily, and one of them ends up repairing most of the damage, as in the case of families 1 (An et al. 2005) and 4 (Sakai et al. 2008), the rest of the enzymes may evolve a broader substrate specificity by random drift. Along this line of thinking, it is commonly thought that the reason why UDG enzymes able to recognize thymine can only act on double-stranded substrates is the strong negative selection against the acquisition of thymine-DNA glycosylase activity in single-stranded DNA or in properly paired contexts. In addition, it has been shown that single substitutions on the human family 1 UDG gene can produce cytosine- and thymine-DNA glycosylases with mutator phenotypes (Kavli et al. 1996).
Our functional review is consistent with the idea that species with larger effective population sizes, where natural selection is more efficient, should have more efficient DNA-repair enzymes, and lower spontaneous mutation rates, in consequence. Thus, family 1 UDGs seem to have higher affinity for its most frequent substrate, U:A (fig. 3), and to prevent more mutations (table 1) in E. coli than in mammals. However, the data available up to now are not fully conclusive due to difficulties in the comparison of catalytic efficiency among species. First, data are scarce for most species and very noisy in some of them. Experimental conditions, such as the saline concentration (Slupphaug et al. 1995, Masaoka et al. 2003) or the presence of Mg2+ and other cofactors (Kavli et al. 2002), affect the catalytic properties of these enzymes and are not homogeneous across the reports (supplementary table S4, Supplementary Material online). UDGs from all families, including highly processive members of family 1, have been reported to be inhibited, to different extents, by their products, the abasic site, and the free uracil (Domena et al. 1988, Waters and Swann 1998, Sandigursky and Franklin 2000, Sung and Mosbaugh 2000, Sandigursky et al. 2001, Srinath et al. 2007, Nilsen et al. 2001, Kavli et al. 2002, Hardeland et al. 2003, O'Neill et al. 2003, Peña Diaz et al. 2004) (although, see as well Kavli et al. 2002, Sartori et al. 2002). This implies that measures of or could be biased if the reaction times chosen to determine the amount of products and the enzyme and substrate concentrations were not chosen carefully (Waters and Swann 1998, Abu and Waters 2003). Another problem with this comparative approach is that differences in the ability to process a substrate could be attributed to divergence in the preferred substrate, rather than to the efficacy of selection. This is likely to have happened in family 2 UDGs, whose mammalian members are specialized in the excision of thymine from T:G mispairs originated from deamination of methylated cytosine in CpG contexts (Waters and Swann 1998), whereas family 2 UDGs from species without DNA methylation may not even recognize thymine (Hardeland et al. 2003).
With respect to the number of mutations prevented by UDGs, it has already been pointed out that the role of an enzyme depends on what other enzymes are encoded in the same genome. Thus, the fact that family 1 prevents on average more mutations in Eubacteria than in mammals (table 1) may be due to the presence of family 3 in the latter.
In summary, we provide a broad evolutionary report of the UDG superfamily with direct implications in the interpretation of their functions. We expect this and future analysis of DNA-repair genes to help explain the observed diversity of mutation rates and spectra among species.
SUPPLEMENTARY MATERIAL
Supplementary tables S1–S4 are available at Molecular Biology and Evolution online (http://www.mbe.oxfordjournals.org/).
Acknowledgments
We thank Adam Eyre-Walker and an anonymous reviewer for helpful comments and suggestions. This work was funded by the National Institute of Health grant number GM36827 to M.L. and W. Kelly Thomas.
References
- Abascal F, Zardoya R, Posada D. ProtTest: selection of best-fit models of protein evolution. Bioinformatics. 2005;21:2104–2105. doi: 10.1093/bioinformatics/bti263. [DOI] [PubMed] [Google Scholar]
- Abu M, Waters TR. The main role of human thymine-DNA glycosylase is removal of thymine produced by deamination of 5-methylcytosine and not removal of ethenocytosine. J Biol Chem. 2003;278:8739–8744. doi: 10.1074/jbc.M211084200. [DOI] [PubMed] [Google Scholar]
- Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997;25:3389–3402. doi: 10.1093/nar/25.17.3389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- An Q, Robins P, Lindahl T, Barnes D. C → T mutagenesis and γ-radiation sensitivity due to deficiency in the Smug1 and Ung DNA glycosylases. EMBO J. 2005;24:2205–2213. doi: 10.1038/sj.emboj.7600689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aravind L, Koonin EV. The α/β uracil DNA glycosylases: a common origin with diverse fates. Genome Biol. 2000 doi: 10.1186/gb-2000-1-4-research0007. 1:research0007.1–research0007.8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Baer CF, Miyamoto MM, Denver DR. Mutation rate variation in multicellular eukaryotes: causes and consequences. Nat Rev Genet. 2007;8:619–631. doi: 10.1038/nrg2158. [DOI] [PubMed] [Google Scholar]
- Boorstein RJ, Chiu LN, Teebor G. Phylogenetic evidence of a role for 5-hydroxymethyluracil-DNA glycosylase in the maintenance of 5-methylcytosine in DNA. Nucleic Acids Res. 1989;17:7653–7661. doi: 10.1093/nar/17.19.7653. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boorstein RJ, Cummings A, Jr., Marenstein DR, Chan MK, Ma Y, Neubert TA, Brown SM, Teebor GW. Definitive identification of mammalian 5-hydroxymethyluracil DNA N-glycosylase activity as SMUG1. J Biol Chem. 2001;276:41991–41997. doi: 10.1074/jbc.M106953200. [DOI] [PubMed] [Google Scholar]
- Burgers PMJ, Klein MB. Selection by genetic transformation of a Saccharomyces cerevisiae mutant defective for the nuclear uracil-DNA-glycosylase. J Bacteriol. 1986;166:905–913. doi: 10.1128/jb.166.3.905-913.1986. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chatterjee A, Singh KK. Uracil-DNA glycosylase-deficient yeast exhibit a mitochondrial mutator phenotype. Nucleic Acids Res. 2001;29:4935–4940. doi: 10.1093/nar/29.24.4935. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen JD, Lacks SA. Role of uracil-DNA glycosylase in mutation avoidance by. Streptococcus pneumoniae. J Bacteriol. 1991;173:283–290. doi: 10.1128/jb.173.1.283-290.1991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chung JH, Im EK, Park HY, Kwon JH, Lee S, Oh J, Hwang KC, Lee JH, Jang Y. A novel uracil-DNA glycosylase family related to the helix-hairpin-helix DNA glycosylase superfamily. Nucleic Acids Res. 2003;31:2045–2055. doi: 10.1093/nar/gkg319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cole JR, Chai B, Farris RJ, Wang Q, Kulam-Syed-Mohideen AS, McGarrell DM, Bandela AM, Cardenas E, Garrity GM, Tiedje JM. The ribosomal database project (RDP-II): introducing myRDP space and quality controlled public data. Nucleic Acids Res. 2007;35:D169–D172. doi: 10.1093/nar/gkl889. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Denver DR, Dolan PC, Wilhelm LJ, et al. (11 co-authors) A genome-wide view of Caenorhabditis elegans base-substitution mutation processes. Proc Natl Acad Sci U S A. 2009;106:16310–16314. doi: 10.1073/pnas.0904895106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Denver DR, Swenson SL, Lynch M. An evolutionary analysis of the helix-hairpin-helix superfamily of DNA repair glycosylases. Mol Biol Evol. 2003;20:1603–1611. doi: 10.1093/molbev/msg177. [DOI] [PubMed] [Google Scholar]
- DeSantis TZ, Hugenholtz P, Larsen N, Rojas M, Brodie EL, Keller K, Huber T, Dalevi D, Hu P, Andersen GL. Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB. Appl Environ Microbiol. 2006;72:5069–5072. doi: 10.1128/AEM.03006-05. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dionne I, Bell SD. Characterization of an archaeal family 4 uracil DNA glycosylase and its interaction with PCNA and chromatin proteins. Biochem J. 2005;387:859–863. doi: 10.1042/BJ20041661. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Domena JD, Timmer RT, Dicharry SA, Mosbaugh DW. Purification and properties of mitochondrial uracil-DNA glycosylase from rat liver. Biochemistry. 1988;27:6742–6751. doi: 10.1021/bi00418a015. [DOI] [PubMed] [Google Scholar]
- Drake JW. A constant rate of spontaneous mutation in DNA-based microbes. Proc Natl Acad Sci U S A. 1991;88:7160–7164. doi: 10.1073/pnas.88.16.7160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Drost JB, Lee WR. Biological basis of germline mutation: comparisons of spontaneous germline mutation rates among Drosophila, mouse, and human. Environ Mol Mutagen. 1995;25((Suppl. 26)):48–64. doi: 10.1002/em.2850250609. [DOI] [PubMed] [Google Scholar]
- Duncan BK, Weiss B. Specific mutator effects of ung (uracil-DNA glycosylase) mutations in. Escherichia coli. J Bacteriol. 1982;151:750–755. doi: 10.1128/jb.151.2.750-755.1982. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eddy SR. Profile hidden Markov models. Bioinformatics. 1998;14:755–763. doi: 10.1093/bioinformatics/14.9.755. [DOI] [PubMed] [Google Scholar]
- Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004;32:1792–1797. doi: 10.1093/nar/gkh340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eisen JA, Hanawalt PC. A phylogenomic study of DNA repair genes, proteins, and processes. Mutat Res. 1999;435:171–213. doi: 10.1016/s0921-8777(99)00050-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Finn RD, Mistry J, Schuster-Böckler B, et al. (13 co-authors) Pfam: clans, web tools and services. Nucleic Acids Res. 2006;34:D247–D251. doi: 10.1093/nar/gkj149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gallinari P, Jiricny J. A new class of uracil-DNA glycosylases related to human thymine-DNA glycosylase. Nature. 1996;383:735–738. doi: 10.1038/383735a0. [DOI] [PubMed] [Google Scholar]
- Garland T, Jr., Harvey PH, Ives AR. Procedures for the analysis of comparative data using phylogenetically independent contrasts. Syst Biol. 1992;41:18–32. [Google Scholar]
- Greagg MA, Fogg MJ, Panayotou G, Evans SJ, Connolly BA, Pearl LH. A read-ahead function in archaeal DNA polymerases detects promutagenic template-strand uracil. Proc Natl Acad Sci U S A. 1999;96:9045–9050. doi: 10.1073/pnas.96.16.9045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Guillet M, Van Der Kemp PA, Boiteux S. dUTPase activity is critical to maintain genetic stability in Saccharomyces cerevisiae. Nucleic Acids Res. 2006;34:2056–2066. doi: 10.1093/nar/gkl139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hall TA. BioEdit: a user-friendly biological sequence alignment editor and analysis program for Windows 95/98/NT. Nucl Acids Symp Ser. 1999;41:95–98. [Google Scholar]
- Hardeland U, Bentele M, Jiricny J, Schär P. The versatile thymine DNA-glycosylase: a comparative characterization of the human, Drosophila and fission yeast orthologs. Nucleic Acids Res. 2003;31:2261–2271. doi: 10.1093/nar/gkg344. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haushalter KA, Stukenberg PT, Kirschner MW, Verdine GL. Identification of a new uracil-DNA glycosylase family by expression cloning using synthetic inhibitors. Curr Biol. 1999;9:174–185. doi: 10.1016/s0960-9822(99)80087-6. [DOI] [PubMed] [Google Scholar]
- Hendrich B, Hardeland U, Ng HH, Jiricny J, Bird A. The thymine glycosylase MBD4 can bind to the product of deamination at methylated CpG sites. Nature. 1999;401:301–304. doi: 10.1038/45843. [DOI] [PubMed] [Google Scholar]
- Hoffman PD, Leonard JM, Lindberg GE, Bollmann SR, Hays JB. Rapid accumulation of mutations during seed-to-seed propagation of mismatch-repair-defective. Arabidopsis. Genes Dev. 2004;18:2676–2685. doi: 10.1101/gad.1217204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huang S, Kang J, Blaser MJ. Antimutator role of the DNA glycosylase mutY gene in. Helicobacter pylori. J Bacteriol. 2006;188:6224–6234. doi: 10.1128/JB.00477-06. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Impellizzeri KJ, Anderson B, Burgers PMJ. The spectrum of spontaneous mutations in a Saccharomyces cerevisiae uracil-DNA-glycosylase mutant limits the function of this enzyme to cytosine deamination repair. J Bacteriol. 1991;173:6807–6810. doi: 10.1128/jb.173.21.6807-6810.1991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ivarie R. Thymine methyls and DNA—protein interactions. Nucleic Acids Res. 1987;15:9975–9983. doi: 10.1093/nar/15.23.9975. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kavli B, Slupphaug G, Mol CD, Arvai AS, Petersen SB, Tainer JA, Krokan HE. Excision of cytosine and thymine from DNA by mutants of human uracil-DNA glycosylase. EMBO J. 1996;15:3442–3447. [PMC free article] [PubMed] [Google Scholar]
- Kavli B, Sundheim O, Akbari M, Otterlei M, Nilsen H, Skorpen F, Aas PA, Hagen L, Krokan HE, Slupphaug G. hUNG2 is the major repair enzyme for removal of uracil from U: A matches, U: G mismatches, and U in single-stranded DNA, with hSMUG1 as a broad specificity backup. J Biol Chem. 2002;277:39926–39936. doi: 10.1074/jbc.M207107200. [DOI] [PubMed] [Google Scholar]
- Keightley PD, Trivedi U, Thomson M, Oliver F, Kumar S, Blaxter ML. Analysis of the genome sequences of three Drosophila melanogaster spontaneous mutation accumulation lines. Genome Res. 2009;19:1195–1201. doi: 10.1101/gr.091231.109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kiyonari S, Uchimura M, Shirai T, Ishino Y. Physical and functional interactions between uracil-DNA glycosylase and proliferating cell nuclear antigen from the euryarchaeon. Pyrococcus furiosus. J Biol Chem. 2008;283:24185–24193. doi: 10.1074/jbc.M802837200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koulis A, Cowan DA, Pearl LH, Savva R. Uracil-DNA glycosylase activities in hyperthermophilic micro-organisms. FEMS Microbiol Lett. 1996;143:267–271. doi: 10.1111/j.1574-6968.1996.tb08491.x. [DOI] [PubMed] [Google Scholar]
- Kurthkoti K, Kumar P, Jain R, Varshney U. Important role of the nucleotide excision repair pathway in Mycobacterium smegmatis in conferring protection against commonly encountered DNA-damaging agents. Microbiology. 2008;154:2776–2785. doi: 10.1099/mic.0.2008/019638-0. [DOI] [PubMed] [Google Scholar]
- Lin Z, Nei M, Ma H. The origins and early evolution of DNA mismatch repair genes–multiple horizontal gene transfers and co-evolution. Nucleic Acids Res. 2007;35:7591–7603. doi: 10.1093/nar/gkm921. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lind PA, Andersson DI. Whole-genome mutational biases in bacteria. Proc Natl Acad Sci U S A. 2008;105:17878–17883. doi: 10.1073/pnas.0804445105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lindahl T, Nyberg B. Heat-induced deamination of cytosine residues in deoxyribonucleic acid. Biochemistry. 1974;13:3405–3410. doi: 10.1021/bi00713a035. [DOI] [PubMed] [Google Scholar]
- Lucas-Lledó JI, Lynch M. Evolution of mutation rates: phylogenomic analysis of the photolyase/cryptochrome family. Mol Biol Evol. 2009;26:1143–1153. doi: 10.1093/molbev/msp029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lutsenko E, Bhagwat A. The role of the Escherichia coli Mug protein in the removal of uracil and 3, N4-ethenocytosine from DNA. J Biol Chem. 1999;274:31034–31038. doi: 10.1074/jbc.274.43.31034. [DOI] [PubMed] [Google Scholar]
- Lynch M. Rate, molecular spectrum, and consequences of human mutation. Proc Natl Acad Sci U S A. 2010;107:961–968. doi: 10.1073/pnas.0912629107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maddison WP, Maddison DR. 2010. Mesquite: a modular system for evolutionary analysis. Version 2.74 [cited 2010 Dec 8]. Available from: http://mesquiteproject.org. [Google Scholar]
- Masaoka A, Matsubara M, Hasegawa R, Tanaka T, Kurisu S, Terato H, Ohyama Y, Karino N, Matsuda A, Ide H. Mammalian 5-formyluracil-DNA glycosylase. 2. Role of SMUG1 uracil-DNA glycosylase in repair of 5-formyluracil and other oxidized and deaminated base lesions. Biochemistry. 2003;42:5003–5012. doi: 10.1021/bi0273213. [DOI] [PubMed] [Google Scholar]
- Midford PE, Jr., Garland T, Maddison WP. 2010. PDAP package of Mesquite. Version 1.15 [cited 2010 Dec 8] Available from: http://mesquiteproject.org/pdap_mesquite. [Google Scholar]
- Mokkapati SK, Fernández de Henestrosa AR, Bhagwat AS. Escherichia coli DNA glycosylase Mug: a growth-regulated enzyme required for mutation avoidance in stationary-phase cells. Mol Microbiol. 2001;41:1101–1111. doi: 10.1046/j.1365-2958.2001.02559.x. [DOI] [PubMed] [Google Scholar]
- Musto H, Naya H, Zavala A, Romero H, Alvarez-Valín Bernardi G. Genomic GC level, optimal growth temperature, and genome size in prokaryotes. Biochem Biophys Res Commun. 2006;347:1–3. doi: 10.1016/j.bbrc.2006.06.054. [DOI] [PubMed] [Google Scholar]
- Neddermann P, Jiricny J. Efficient removal of uracil from G·C mispairs by the mismatch-specific thymine DNA glycosylase from HeLa cells. Proc Natl Acad Sci U S A. 1994;91:1642–1646. doi: 10.1073/pnas.91.5.1642. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nilsen H, Haushalter KA, Robins P, Barnes DE, Verdine GL, Lindahl T. Excision of deaminated cytosine from the vertebrate genome: role of the SMUG1 uracil-DNA glycosylase. EMBO J. 2001;20:4278–4286. doi: 10.1093/emboj/20.15.4278. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nilsen H, Rosewell I, Robins P, Skjelbred CF, Andersen S, Slupphaug G, Daly G, Krokan HE, Lindahl T, Barnes DE. Uracil-DNA glycosylase (UNG)-deficient mice reveal a primary role of the enzyme during DNA replication. Mol Cell. 2000;5:1059–1065. doi: 10.1016/s1097-2765(00)80271-3. [DOI] [PubMed] [Google Scholar]
- Nordman J, Wright A. The relationship between dNTP pool levels and mutagenesis in an Escherichia coli NDP kinase mutant. Proc Natl Acad Sci U S A. 2008;105:10197–10202. doi: 10.1073/pnas.0802816105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- O'Brien PJ. Catalytic promiscuity and the divergent evolution of DNA repair enzyme. Chem Rev. 2006;106:720–752. doi: 10.1021/cr040481v. [DOI] [PubMed] [Google Scholar]
- O'Neill RJ, Vorob'eva OV, Shahbakhti H, Zmuda E, Bhagwat AS, Baldwin GS. Mismatch uracil glycosylase from Escherichia coli. A general mismatch or a specific DNA glycosylase? J Biol Chem. 2003;278:20526–20532. doi: 10.1074/jbc.M210860200. [DOI] [PubMed] [Google Scholar]
- Ossowski S, Schneeberger K, Lucas-LLedó JI, Warthmann N, Clark RM, Shaw RG, Weigel D, Lynch M. The rate and molecular spectrum of spontaneous mutations in. Arabidopsis thaliana. Science. 2010;327:92–94. doi: 10.1126/science.1180677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pagel M. Detecting correlated evolution on phylogenies: a general method for the comparative analysis of discrete characters. Proc R Soc Lond B. 1994;255:37–45. [Google Scholar]
- Pagel M, Meade A. Bayesian analysis of correlated evolution of discrete characters by reversible-jump Markov chain Monte Carlo. Am Nat. 2006;167:808–825. doi: 10.1086/503444. [DOI] [PubMed] [Google Scholar]
- Pagel MD. A method for the analysis of comparative data. J Theor Biol. 1992;156:431–442. [Google Scholar]
- Pearl LH. Structure and function in the uracil-dna glycosylase superfamily. Mutat Res. 2000;460:165–181. doi: 10.1016/s0921-8777(00)00025-2. [DOI] [PubMed] [Google Scholar]
- Peña Diaz J, Akbari M, Sundheim O, Farez-Vidal ME, Andersen S, Sneve R, Gonzalez-Pacanowska D, Krokan HE, Slupphaug G. Trypanosoma cruzi contains a single detectable uracil-DNA glycosylase and repairs uracil exclusively via short patch base excision repair. J Mol Biol. 2004;342:787–799. doi: 10.1016/j.jmb.2004.07.043. [DOI] [PubMed] [Google Scholar]
- Pettersen HS, Sundheim O, Gilljam KM, Slupphaug G, Krokan HE, Kavli B. Uracil-DNA glycosylases SMUG1 and UNG2 coordinate the initial steps of base excision repair by distinct mechanisms. Nucleic Acids Res. 2007;35:3879–3892. doi: 10.1093/nar/gkm372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Radany EH, Dornfeld KJ, Sanderson RJ, Savage MK, Majumdar A, Seidman MM, Mosbaugh DW. Increased spontaneous mutation frequency in human cells expressing the phage PBS2-encoded inhibitor of uracil-DNA glycosylase. Mutat Res. 2000;461:41–58. doi: 10.1016/s0921-8777(00)00040-9. [DOI] [PubMed] [Google Scholar]
- Sakai T, Tokishita SI, Mochizuki K, Motomiya A, Yamagata H, Ohta T. Mutagenesis of uracil-DNA glycosylase deficient mutants of the extremely thermophilic eubacterium. Thermus thermophilus. DNA Repair. 2008;7:663–669. doi: 10.1016/j.dnarep.2008.01.006. [DOI] [PubMed] [Google Scholar]
- Sandigursky M, Faje A, Franklin WA. Characterization of the full length uracil-DNA glycosylase in the extreme thermophile. Thermotoga maritima. Mutat Res. 2001;485:187–195. doi: 10.1016/s0921-8777(00)00083-5. [DOI] [PubMed] [Google Scholar]
- Sandigursky M, Franklin WA. Thermostable uracil-DNA glycosylase from Thermotoga maritima, a member of a novel class of DNA repair enzymes. Curr Biol. 1999;9:531–534. doi: 10.1016/s0960-9822(99)80237-1. [DOI] [PubMed] [Google Scholar]
- Sandigursky M, Franklin WA. Uracil-DNA glycosylase in the extreme thermophile. Archaeoglobus fulgidus. J Biol Chem. 2000;275:19146–19149. doi: 10.1074/jbc.M001995200. [DOI] [PubMed] [Google Scholar]
- Saparbaev M, Laval J. 3, N4-ethenocytosine, a highly mutagenic adduct, is a primary substrate for Escherichia coli double-stranded uracil-DNA glycosylase and human mismatch-specific thymine-DNA glycosylase. Proc Natl Acad Sci U S A. 1998;95:8508–8513. doi: 10.1073/pnas.95.15.8508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sartori AA, Fitz-Gibbon S, Yang H, Miller JH, Jiricny J. A novel uracil-DNA glycosylase with broad substrate specificity and an anusual active site. EMBO J. 2002;21:3182–3191. doi: 10.1093/emboj/cdf309. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharp PM, Bailes E, Grocock RJ, Peden JF, Sockett RE. Variation in the strength of selected codon usage bias among bacteria. Nucleic Acids Res. 2005;33:1141–1153. doi: 10.1093/nar/gki242. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Slupphaug G, Eftedal I, Kavli B, Bharati S, Helle NM, Haug T, Levine DW, Krokan HE. 1995. Properties of a recombinant human uracil-DNA. [DOI] [PubMed] [Google Scholar]
- glycosylase from the UNG gene and evidence that UNG encodes the major uracil-DNA glycosylase. Biochemistry 34:128–138. [DOI] [PubMed] [Google Scholar]
- Srinath T, Bharti SK, Varshney U. Substrate specificities and functional characterization of a thermo-tolerant uracil DNA glycosylase (UdgB) from. Mycobacterium tuberculosis. DNA Repair. 2007;6:1517–1528. doi: 10.1016/j.dnarep.2007.05.001. [DOI] [PubMed] [Google Scholar]
- Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22:2688–2690. doi: 10.1093/bioinformatics/btl446. [DOI] [PubMed] [Google Scholar]
- Stamatakis A, Hoover P, Rougemont J. A rapid bootstrap algorithm for the RAxML web servers. Syst Biol. 2008;57:758–771. doi: 10.1080/10635150802429642. [DOI] [PubMed] [Google Scholar]
- Sung JS, Mosbaugh DW. Escherichia coli double-strand uracil-DNA glycosylase: involvement in uracil-mediated DNA base excision repair and stimulation of activity by endonuclease IV. Biochemistry. 2000;39:10224–10235. doi: 10.1021/bi0007066. [DOI] [PubMed] [Google Scholar]
- Talavera G, Castresana J. Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments. Syst Biol. 2007;56:564–577. doi: 10.1080/10635150701472164. [DOI] [PubMed] [Google Scholar]
- Tye BK, Chien J, Lehman IR, Duncan BK, Warner HR. Uracil incorporation: a source of pulse-labeled DNA fragments in the replication of the Escherichia coli chromosome. Proc Natl Acad Sci U S A. 1978;75:233–237. doi: 10.1073/pnas.75.1.233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Venkatesh J, Kumar P, Krishna PSM, Manjunath R, Varshney U. Importance of uracil DNA glycosylase in Pseudomonas aeruginosa and Mycobacterium smegmatis, G + C-rich bacteria, in mutation prevention, tolerance to acidified nitrite, and endurance in mouse macrophages. J Biol Chem. 2003;278:24350–24358. doi: 10.1074/jbc.M302121200. [DOI] [PubMed] [Google Scholar]
- Waters TR, Swann PF. Kinetics of the action of thymine DNA glycosylase. J Biol Chem. 1998;273:20007–20014. doi: 10.1074/jbc.273.32.20007. [DOI] [PubMed] [Google Scholar]
- Webb CO, Ackerly DD, Kembel SW. Phylocom: software for the analysis of phylogenetic community structure and trait evolution. Bioinformatics. 2008;24:2098–2100. doi: 10.1093/bioinformatics/btn358. [DOI] [PubMed] [Google Scholar]
- Wilkins AS, editor. Genetic analysis of animal development. New York: Wiley-Liss; 1992. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



