Abstract
Halohydrin dehalogenases are very rare enzymes that are naturally involved in the mineralization of halogenated xenobiotics. Due to their catalytic potential and promiscuity, many biocatalytic reactions have been described that have led to several interesting and industrially important applications. Nevertheless, only a few of these enzymes have been made available through recombinant techniques; hence, it is of general interest to expand the repertoire of these enzymes so as to enable novel biocatalytic applications. After the identification of specific sequence motifs, 37 novel enzyme sequences were readily identified in public sequence databases. All enzymes that could be heterologously expressed also catalyzed typical halohydrin dehalogenase reactions. Phylogenetic inference for enzymes of the halohydrin dehalogenase enzyme family confirmed that all enzymes form a distinct monophyletic clade within the short-chain dehydrogenase/reductase superfamily. In addition, the majority of novel enzymes are substantially different from previously known phylogenetic subtypes. Consequently, four additional phylogenetic subtypes were defined, greatly expanding the halohydrin dehalogenase enzyme family. We show that the enormous wealth of environmental and genome sequences present in public databases can be tapped for in silico identification of very rare but biotechnologically important biocatalysts. Our findings help to readily identify halohydrin dehalogenases in ever-growing sequence databases and, as a consequence, make even more members of this interesting enzyme family available to the scientific and industrial community.
INTRODUCTION
Halohydrin dehalogenases (also called haloalcohol dehalogenases, haloalcohol/halohydrin epoxidases, or hydrogen-halide lyases; EC 4.5.1.−) (HHDHs) are biotechnologically relevant enzymes that catalyze the reversible dehalogenation of β-haloalcohols with epoxide formation (1–3). Besides being useful for the production of enantiopure haloalcohols (4–6) and epoxides (5, 7–9), these enzymes can also be applied in the formation of novel carbon-carbon, carbon-nitrogen, or carbon-oxygen bonds. Due to the promiscuous epoxide ring-opening activity of these enzymes, cyanide, azide, and nitrite, for example, are accepted as nucleophiles in the ring-opening reaction, leading to a diverse range of products (10). Examples of important HHDH applications are the production of optically pure C3 or C4 fine-chemical precursors (5, 8, 11, 12), including the multiton-scale production of enantiopure (R)-4-cyano-3-hydroxybutyrate esters for statin drugs (13), and the production of chiral tertiary alcohols (14–17), for which conventional organic synthesis is rather challenging (Fig. 1) (18, 19).
FIG 1.
Examples of HHDH-catalyzed reactions include the preparation of (optically pure) haloalcohols and epoxides (5) (A) as well as synthetic routes toward statin side chain precursors (13) (B) and tertiary alcohols (15, 17) (C).
Despite their designated potential as biocatalysts, few HHDHs have thus far been made available to the scientific and industrial community. Since the initial discovery of bacterial enzymes with HHDH activities more than 45 years ago, a couple of bacterial species have been reported to possess HHDH activity, but only a very few HHDH enzymes have been purified and characterized biochemically (recently reviewed in detail [2, 3]). Of these, only six HHDH genes have been cloned and expressed recombinantly: hheA from Corynebacterium sp. strain N-1074 (20), hheA2 from Arthrobacter sp. strain AD2 (21), hheB from Corynebacterium sp. strain N-1074 (20), hheB2 from Mycobacterium sp. strain GP1 (21), and two identical hheC sequences from Agrobacterium radiobacter AD1 (21) and Rhizobium sp. strain NHG3 (22).
All of the cloned HHDHs belong to the short-chain dehydrogenase/reductase (SDR) superfamily and exhibit several major features of this diverse enzyme class (21, 23, 24). For example, all known HHDHs make use of a catalytic triad and share the commonly found homomultimeric quaternary assembly, and both crystallized HHDHs, namely, HheA2 (25) and HheC (26), possess a tertiary structure similar to those of other Rossmann fold proteins. Aside from these overall similarities to SDR enzymes, HHDHs can be distinguished from them by a combination of mechanistic and sequence/structure characteristics. The concerted activity of Ser-Tyr-Lys in classical SDR enzymes abstracts a proton from the substrate's hydroxyl group, and an enzyme-bound NAD(P)+ cofactor is responsible for hydride abstraction. In contrast, HHDHs possess a catalytic triad composed of Ser-Tyr-Arg, and instead of a cofactor-binding site, a spacious anion binding pocket is present in the structures of HheA2 and HheC (25, 26). As a consequence, all known HHDH sequences form only a minute but well-defined fraction within the SDR superfamily, which comprises more than 163,000 SDR enzymes that can be retrieved from UniProt (27).
Based on activity profiles and sequence identities, the available HHDH enzymes have been classified into three different phylogenetic subgroups: types A, B, and C (21). HHDHs within each type share more than 97% sequence identity, while the level of identity between enzymes of different types is below 33%. Due to these high sequence identities within each of the three subtypes, only three different HHDH enzymes are currently available for biotechnological exploitation.
Although these few known HHDHs have already given rise to many interesting applications, it is of great interest to increase the number of functionally diverse HHDH enzymes (2, 3). For this purpose, a viable approach employs rational and random protein engineering strategies that can yield drastically improved and functionally diverse enzyme variants (28). Such strategies have been successfully applied to HheA2 (29–31) and HheC (13, 32–34), addressing specific drawbacks of the respective parental enzymes and yielding HHDH mutants with sometimes substantial improvements with regard to target reactions. Nevertheless, parental sequences will always govern the overall accessible sequence space in every protein engineering study. For example, HheC has been shown to be extraordinarily tolerant, with mutations at 153 of its total 254 residues; a maximum of 42 simultaneous substitutions per variant have been described (13). However, these heavily engineered enzyme variants are still rather similar to parental HheC; for example, HheC-2360 (35) shows more than 85% sequence identity. Clearly, novel sequences would be a valuable addition to the functional diversity of the HHDH enzyme toolbox. Further, novel HHDH enzymes might already exhibit activities or characteristics that are difficult to engineer or even unlikely to be accessible by laboratory evolution.
Here we report the identification of novel HHDHs in publicly available sequence databases by making use of specific sequence motifs that allow for the unambiguous discrimination of true HHDH sequences from the vast number of other SDR sequences.
MATERIALS AND METHODS
Database mining.
The sequences of HheA, HheB, and HheC (Table 1) were used as queries for blastp searches (36) of the nonredundant (nr) database of GenBank (release 195) (37). For each query, 20,000 sequences were retrieved and were used together with HheA, HheB, and HheC for the construction of multiple sequence alignments (MSAs) with MAFFT (FFT-NS-2) (38). In the resulting alignments, sequences were dismissed if they possessed the typical Ser-Tyr-Lys catalytic residues of SDR enzymes which aligned with the Ser-Tyr-Arg catalytic triads of HheA/HheA2 (S134-Y148-R151), HheB/HheB2 (S127-Y139-R143), and HheC (S132-Y145-R149).
TABLE 1.
Sources and accession numbers of previously known and novel HHDHs
| HHDHa | Organism or source | Protein accession no. |
|---|---|---|
| HheA* | Corynebacterium sp. | BAA14361 |
| HheA2* | Arthrobacter sp. strain AD2 | AAK92100 |
| HheA3 | Parvibaculum lavamentivorans DS-1 | ABS64560 |
| HheA4 | Arthrobacter sp. strain JBH1 | AFI98638 |
| HheA5 | Tistrella mobilis KA081020-065 | AFK51877 |
| HheB* | Corynebacterium sp. | BAA14362 |
| HheB2* | Mycobacterium sp. strain GP1 | AAK73175 |
| HheB3 | Marine metagenome (Ralstonia)b | EBL02020 |
| HheB4 | Marine metagenome (Shewanella)b | EBP61646 |
| HheB5 | Marine metagenome (Burkholderia)b | ECR06649 |
| HheB6 | Marine metagenome (Sorangium)b | EDB56284 |
| HheB7 | Marine metagenome (Bradyrhizobium)b | EDD65701 |
| HheC* | Agrobacterium tumefaciens | AAK92099 |
| HheD | Dechloromonas aromatica RCB | AAZ44846 |
| HheD2 | Gammaproteobacterium strain HTCC2207 | EAS46473 |
| HheD3 | Methylibium petroleiphilum PM1 | ABM93639 |
| HheD4 | Marine metagenome (Haliangium)b | ECY18578 |
| HheD5 | Thauera sp. strain MZ1T | YP_002355872 |
| HheE | Marine metagenome (Acaryochloris)b | EBP63112 |
| HheE2 | Marine metagenome (Sorangium)b | ECW41905 |
| HheE3 | Marine metagenome (Burkholderia)b | EDF62577 |
| HheE4 | Marine metagenome (Catenulispora)b | EDH34310 |
| HheE5 | Gammaproteobacterium strain IMCC3088 | EGG28524 |
| HheF | Uncultured bacterium | BAH89601 |
| HheG | Ilumatobacter coccineus YM16-304 | BAN03849 |
Previously known HHDHs are marked with asterisks.
Taxonomic classification according to the NBC Web server.
After the removal of those putative SDR sequences, the pool of remaining sequences was realigned using MAFFT (FFT-NS-2). Afterwards, only sequences with a catalytic triad of Ser-Tyr-Arg, also present in known HHDH enzymes, were selected for the construction of a new MAFFT (L-INS-i) (39) MSA. From the latter alignment, sequences were considered to be putative novel HHDHs only if they possessed an aromatic Phe or Tyr that aligned with F12, Y27, or F12 from HheA, HheB, or HheC, respectively.
All putative novel HHDH sequences were used as queries in subsequent search routines, which consisted of blastp searches, construction of MSAs, and their inspection for the presence of the typical HHDH catalytic triad in combination with the conserved aromatic residue close to the N terminus as outlined above. This strategy was continued until no additional putative HHDH sequences could be identified in the nr database with the protocol specified.
Similarly, the GenBank collection of nonredundant sequences from environmental sources (env_nr [release 195]) was searched for putative novel HHDH sequences by using the known and putative novel HHDH sequences as blastp queries. Again, as many as 20,000 sequences per query were retrieved with blastp and an expect threshold (E) of 10. Before MSA construction for the identification of conserved HHDH sites as outlined above, env_nr sequences shorter than 180 residues (80% of the HheE5 sequence) were removed from the environmental sequence pool. Each complete nucleotide record for each HHDH of metagenomic origin was taxonomically classified via the Naïve Bayes Classification (NBC) tool Web server (40).
Gene synthesis and cloning.
Prior to gene synthesis, the respective nucleotide records of all putative novel HHDHs were first inspected for the presence of an alternative ATG start codon that was also preceded by a putative Shine-Dalgarno (SD) sequence downstream of the annotated transcription start. If the shorter protein sequences also contained the conserved aromatic Phe or Tyr residue (see above), then these corrected protein sequences were used as a basis for gene synthesis, activity tests, and phylogenetic analysis.
Synthetic genes were ordered from Life Technologies (Darmstadt, Germany) after back translation of the curated protein sequences and codon optimization for Escherichia coli with GeneOptimizer (41). The synthetic genes were excised from the plasmids received by restriction with NdeI and either HindIII or XhoI, followed by ligation with T4 DNA ligase (all DNA-modifying enzymes were from New England BioLabs, Frankfurt, Germany) into the linearized expression vector pET-28a (Merck, Darmstadt, Germany). After transformation into E. coli DH5α (Life Technologies), recombinant plasmid DNA was prepared with the NucleoSpin Plasmid kit (Macherey-Nagel, Düren, Germany) and was sent to GATC Biotech (Constance, Germany) for sequencing.
Expression and activity assays.
Expression plasmids containing the different HHDH genes were transformed into either E. coli BL21(DE3) (Life Technologies) or E. coli C43(DE3) (Lucigen Corporation, Middleton, WI, USA) for heterologous protein expression. Each 20 ml of Terrific broth (TB) medium containing 50 mg liter−1 kanamycin and 0.2 mM isopropyl-β-d-thiogalactopyranoside (IPTG) was inoculated with 10% (vol/vol) of an E. coli preculture and was incubated at 20, 30, or 37°C for 7 to 24 h. Afterwards, cultures were centrifuged, and pellets were stored at −20°C. For reference, the known enzymes HheA2, HheB2, and HheC (21) were recombinantly expressed in E. coli TOP10 cells (Life Technologies) from the pBAD vector (Life Technologies) in TB medium with 100 mg liter−1 ampicillin and 0.02% (wt/vol) l-arabinose. To prepare cell-free extracts (CFEs), cell pellets were resuspended in 1.2 ml of 25 mM Tris·SO4 buffer, pH 7.5, and were disrupted by sonication and centrifuged. SDS-PAGE and Western blot analyses of CFEs were performed to analyze the heterologous expression of the different HHDHs. For SDS-PAGE, about 10 μg of total protein present in CFEs was separated in 12% polyacrylamide gels for 45 min at 200 V. In Western blot analyses, His-tagged proteins were detected by using horseradish peroxidase (HRP)-conjugated Ni-nitrilotriacetic acid (NTA) (Qiagen, Hilden, Germany) and the Pierce ECL Western blotting substrate (Thermo Fisher Scientific, Rockford, IL, USA) according to the manufacturers' instructions.
After centrifugation, 200 μl of each CFE was added to 600 μl of 25 mM Tris·SO4 buffer, pH 7.5, containing one of the substrates 1,3-dichloro-2-propanol, 2-chlorophenylethanol, and 1,3-dibromo-2-propanol at a final concentration of 10 mM for activity assays. Reaction mixtures were incubated at 30°C, and 50-μl samples were taken after 5, 20, and 60 min of incubation to monitor dehalogenase activity by using the halide release assay as described elsewhere (35). In addition, after 60 min, the remaining 650 μl of the reaction mixture was extracted once with 600 μl methyl tert-butyl ether containing 0.1% (vol/vol) dodecane as an internal standard. Organic extracts were dried over magnesium sulfate and were analyzed on a GC-2010 gas chromatograph (GC) (Shimadzu, Duisburg, Germany) equipped with a Supreme-5ms column (CS Chromatography Service, Langerwehe, Germany).
Reaction mixtures containing 1,3-dichloro-2-propanol or 1,3-dibromo-2-propanol as the substrate were analyzed by GC with a temperature gradient starting at 40°C for 1 min, with heating first at 10°C min−1 to 120°C and then at 20°C min−1 to 300°C. For reaction mixtures containing 2-chlorophenylethanol as the substrate, the temperature gradient started at 80°C for 1 min, with heating first at 10°C min−1 to 160°C and then at 20°C min−1 to 300°C. Substrates eluted at 6.6 min (1,3-dichloro-2-propanol), 8.6 min (2-chlorophenylethanol), and 9.4 min (1,3-dibromo-2-propanol), whereas corresponding products were detected at 3.7 min (epichlorohydrin), 5.8 min (styrene oxide), and 4.8 min (epibromohydrin), respectively. Product formation was monitored based on relative peak areas and was quantified with the help of standard curves. All chemicals were obtained from Sigma-Aldrich (Steinheim, Germany) at the highest purity available.
HHDH fingerprinting.
Sequence logos were generated with WebLogo (42) on the basis of a MAFFT MSA either with all HHDHs or with 718 homologous SDR sequences (see below). For HheC, the precomputed BLAST link (BLINK) at the National Center for Biotechnology Information website was used to collect sequences with an E cutoff of 100 from the UniProtKB/Swiss-Prot (43) and PDB (44) databases. The T-X4-(F/Y)-X-G or S-X12-Y-X3-R pattern, or combinations of these, were used as seeds in PHI-BLAST (45) searches of the GenBank nr database (release 202). Here HheA, HheB, HheC, HheD, HheE, HheF, and HheG were used as queries with an E of 10, and the resulting alignments were inspected for sequences with both conserved HHDH sequence features as outlined above. From the PHI-BLAST results of the combined-pattern searches, PSI-BLAST searches (36) were initiated after all HHDHs identified were selected for building the position-specific scoring matrix (PSSM).
Phylogenetic analysis.
MAFFT (L-INS-i) and webPRANK (46) MSAs with all previously known and novel HHDH sequences were used for minimum evolution and maximum likelihood tree reconstruction. Minimum evolution trees were generated by FastME (version 2.07) (47) with nearest neighbor interchange (NNI) and subtree pruning and regrafting (SPR) tree topology refinement from a PROTDIST distance matrix (program included in the PHYLIP package, version 3.695) (48) using the JTT substitution matrix. For maximum likelihood trees, PhyML (version 3.1) (49) was used with the WAG substitution matrix, NNI and SPR tree refinement, and 5 random starting trees. Trees were rooted with the help of SDR enzymes DHRS4 (UniProt ID Q9BTZ2) and FabG (UniProt ID P55336), which were included in MSAs and tree reconstruction as an outgroup. Phylogenetic trees were assessed for their reliabilities by bootstrap analyses with 100 replicas for PhyML trees or 1,000 replicas for FastME trees.
For phylogenetic placement of HHDH sequences with respect to other SDR sequences, 100 of the most homologous SDR sequences were collected from BLINK results for a diverse selection of HHDHs (HheA, HheA3, HheA5, HheB, HheB5, HheC, HheD, HheD2, HheD5, HheE, HheE2, HheE3, HheE5, HheF, and HheG). In total, 718 SDR sequences that possessed a Ser-Tyr-Lys catalytic triad were used for the construction of a FastME minimum evolution tree based on a MAFFT MSA as outlined above, and the resulting tree was assessed for its reliability with 100 bootstrap replicas.
RESULTS
Sequence characteristics of HHDHs.
First, in order to identify novel HHDH enzymes, all known HHDH sequences were inspected for distinctive residues that distinguish HHDHs from other SDR enzymes. As deduced from a previous ClustalW MSA and confirmed by mutational studies (21), all known HHDHs possess a catalytic triad of Ser-Tyr-Arg that aligns with the catalytic residues Ser-Tyr-Lys present in SDR sequences. Thus, the presence of Arg in the HHDH catalytic triad can be used as an initial criterion to distinguish putative novel HHDHs from other SDR sequences. For the certain discrimination of novel putative HHDH enzymes, however, an additional identification criterion was identified from the alignment of all known HHDHs as well as structural data (see below).
The crystal structures of HheA2 (25) and HheC (26) show that both HHDHs possess a spacious anion binding pocket, which is formed in part by residues that align with residues of a Gly-rich motif responsible for nucleotide cofactor binding in Rossmann fold enzymes such as SDR enzymes (23, 24). Specifically, the large residue F12 in both HheA2 and HheC is essential for the formation of the HHDH anion binding pocket and replaces a central small Gly or Ala in the T-G-X3-(G/A)-X-G nucleotide binding motif of classical SDR enzymes (24–26). In the first ClustalW alignments of all known HHDHs (21), residue R7 of both known B-type enzymes aligned with residue F12 of the other HHDHs. The overall alignment quality, however, was low in this region. Later, an improved alignment of all known HHDHs was published which also incorporated structural information and showed that now Y27 of both B-type enzymes aligned with F12 of the other HHDHs (26). Thus, in all known HHDH enzymes with sequence identities as low as 33%, the large aromatic amino acid Phe or Tyr disturbs the commonly observed Gly-rich cofactor binding motif of SDR enzymes; this feature, therefore, might indicate sequences with HHDH activity. Instead of a structural alignment algorithm, we used MAFFT (50) because of its high computational efficiency and accuracy in the alignment of thousands of individual sequences (51, 52). Furthermore, MAFFT also correctly aligns F12 with Y27 as well as aligning the catalytic triad residues of the known HHDH enzymes (see Fig. S1 in the supplemental material).
In conclusion, we propose that HHDH enzymes can be distinguished from other SDR enzymes by the presence of the HHDH catalytic triad and a conserved aromatic Phe or Tyr which replaces the central small Gly or Ala in the T-G-X3-(G/A)-X-G motif of classical SDR enzymes.
Database mining.
In order to identify novel HHDH sequences, blastp searches were initiated to collect homologous sequences from GenBank which could then be assessed for the presence of both conserved HHDH sequence features.
Initially, homologous sequences were collected by blastp from the GenBank nr protein sequence databases with the sequences of HheA, HheB, and HheC as queries. Afterwards, a MAFFT (FFT-NS-2) alignment was used to identify the large majority of putative SDR sequences with catalytic Ser-Tyr-Lys residues and to exclude them from the sequence pool. Then, after realignment with MAFFT (FFT-NS-2), sequences that lacked the catalytic Ser-Tyr-Arg triad of known HHDH enzymes were removed. This much smaller sequence set was then effectively aligned by MAFFT (L-INS-i) to identify putative novel HHDH sequences that also possess the conserved aromatic Phe or Tyr present in known HHDH enzymes. This entire process was iterated until no further putative novel HHDH sequence could be identified (Fig. 2).
FIG 2.
Flow scheme for the in silico identification of novel HHDHs. Homologous protein sequences were collected using blastp and were aligned by MAFFT to identify novel HHDH sequences by sequentially removing sequences that possessed the typical Ser-Tyr-Lys catalytic residues of SDR enzymes (“S-Y-K”), that lacked the conserved HHDH catalytic triad of Ser-Tyr-Arg (no “S-Y-R”), and that did not possess the specific aromatic Phe or Tyr (no “F/Y”). Through iteration, all of the novel HHDH sequences (Table 1) were identified.
From the GenBank collection of nonredundant (nr) sequences, 35,448 unique sequences were obtained by using HheA, HheB, and HheC as blastp queries. Of these, 23 sequences contained the catalytic triad of known HHDHs, but only 9 sequences also possessed the aromatic Phe or Tyr specific for known HHDH enzymes. Sequences not considered to be putative novel HHDHs were, for example, too short to contain the conserved aromatic Phe or Tyr of known HHDHs or possessed a variation of the Gly-rich T-G-X3-(G/A)-X-G motif required for nucleotide binding in SDR enzymes.
The nine putative novel HHDH sequences originated from Parvibaculum lavamentivorans DS-1 (HheA3), Arthrobacter sp. strain JBH1 (HheA4), Tistrella mobilis KA081020-065 (HheA5), Dechloromonas aromatica RCB (HheD), the marine gammaproteobacterium strain HTCC2207 (HheD2), Methylibium petroleiphilum PM1 (HheD3), Thauera sp. strain MZ1T (HheD5), the gammaproteobacterium strain IMCC3088 (HheE5), and an uncultured bacterium (HheF) (Table 1). By using these putative novel HHDH sequences as queries for subsequent blastp searches, another 37,469 unique sequences were collected from the nr database. These contained one additional putative novel HHDH from Ilumatobacter coccineus YM16-304 (HheG) (Table 1) with both the conserved HHDH catalytic triad and aromatic Phe or Tyr. By using the HheG sequence as a query for a subsequent blastp search, another 2,035 unique sequences were retrieved, but no additional sequence that contained the correct Ser-Tyr-Arg catalytic triad of known HHDHs in combination with the HHDH-specific aromatic Phe or Tyr residue could be identified.
No information on the activities of the putative novel HHDH sequences obtained can be retrieved from associated GenBank records. Except for HheA4, sequence identities to known HHDHs range between 32% and 48%. Although sequence HheA4 from Arthrobacter sp. strain JBH1 is annotated as “3-oxoacyl-acyl-carrier-protein,” it very likely represents a true HHDH enzyme, since it is identical to HheA in 242 of 244 residues. For two further sequences, additional information from the NCBI website can be retrieved which indicates that these enzymes might be HHDHs. Sequences HheA3 from P. lavamentivorans and HheA5 from T. mobilis exhibit only low sequence identities of 33% and 38% to HheA and HheA2, respectively. Nevertheless, HheA3 belongs to the “haloalcohol dehalogenase, classical (c) SDRs” cluster cd05361 of the Conserved Domain Database (53). This cluster is part of the Rossmann fold NAD(P)+-binding protein superfamily, and the HheA3 sequence is the threshold setting representative for this cluster, which also comprises the sequences of HheA, HheA2, and HheC, as well as associated crystal structures. Due to its high sequence identity to HheA, HheA4 is also on the specific hit list of an RPS-BLAST search (36, 54) for the conserved domain of cluster cd05361. In contrast, HheA5 from T. mobilis is not found among the specific hits of an RPS-BLAST search but has been annotated as “halohydrin epoxidase A,” most likely as a result of the submitting authors' gene annotation algorithms (55). The remaining sequences are mostly annotated as “SDR enzyme” or “oxidoreductase” but never as HHDH enzymes.
Since no further HHDH sequences could be retrieved from the nr database, the GenBank collection of nonredundant sequences from environmental sources (env_nr) was also surveyed for the presence of putative novel HHDH sequences. All known and putative novel HHDH sequences were used as blastp queries to collect 27,907 unique environmental sequences, but the alignment of these sequences was more challenging than before.
The majority of env_nr sequences were derived from shotgun sequencing projects of environmental samples such as the Global Ocean Sampling (GOS) studies (56, 57). Usually, these short sequencing reads were assembled into larger contigs, which could cause some of the annotated open reading frames to extend beyond the sequenced and assembled boundaries. As a consequence, the env_nr database contains a higher portion of either N- or C-terminally truncated proteins than the nr database, a feature that is also indicated by a 60% shorter average length of GOS protein sequences (56). Apparently, during MSA construction, this higher portion of truncated proteins caused misalignment of the catalytic triad residues for the known HHDH sequences, which were always included as an indicator of alignment accuracy and reliability.
To circumvent this, sequences shorter than 180 residues—a value corresponding to 80% of the sequence length of the shortest putative novel HHDH, HheE5— were removed from the original environmental sequence set. With this measure, the catalytic triad residues of the known HHDHs could be correctly aligned, and 40 sequences with an HHDH Ser-Tyr-Arg catalytic triad were identified from the remaining 19,068 sequences. Of these, 10 complete sequences (HheB3 through HheB7, HheD4, and HheE through HheE4) (Table 1) also possessed the HHDH-specific aromatic Phe or Tyr residue. According to the taxonomic classification of the NBC Web server, each of the metagenomic nucleotide contigs that harbored a novel HHDH originated from a bacterial strain, and eight of these were proteobacteria. Sequence identities with the known HHDH enzymes ranged from 37% to 60%, and again, no GenBank record indicated HHDH activity, because all these sequences were annotated as “hypothetical protein.”
In summary, a total of 20 novel HHDH sequences were initially identified from more than 100,000 protein sequences present in the GenBank nr and env_nr databases on the basis of HHDH-specific features (Table 1).
Experimental verification of HHDH activity.
In order to investigate if the sequences identified indeed represent enzymes with true HHDH activity, codon-optimized synthetic genes were ordered for heterologous expression of the respective proteins in E. coli.
Prior to gene synthesis, all coding sequences were inspected for alternative translation start signals, since 9 of the 20 sequences were annotated with start codons different from the commonly observed ATG (see Fig. S2 in the supplemental material). First, the sequences were inspected for the presence of an alternative standard ATG start codon downstream of the annotated translation start. Then, and only if the shorter alternative gene product also contained the conserved aromatic Phe or Tyr residue of known HHDHs, this curated sequence was used for further analysis. In addition, sequences immediately upstream of each curated ATG start codon were inspected for overall similarity to the TAAGGAGGTGA SD sequence of E. coli, required as a ribosomal binding site. In all nine of the coding sequences with a nonstandard start codon, an intact shortened gene product, now starting from a standard ATG start codon, could be identified (see Fig. S2). In each case, curated ATG start codons were preceded by sequences with at least weak homology to the E. coli SD sequence.
Except for HheA4, which likely possesses HHDH activity, as evidenced by its exceptionally high sequence identity with HheA, synthetic genes coding for the putative HHDHs were ordered and were cloned into the well-established pET-28a expression vector. Heterologous expression of soluble enzyme was optimized for each HHDH by varying parameters such as the expression host [E. coli BL21(DE3) or E. coli C43(DE3)] and expression temperature (20, 30, or 37°C). The optimization of expression conditions (see Table S1 in the supplemental material) yielded visible bands of soluble enzyme for most HHDHs after Coomassie staining of polyacrylamide gels; such bands were not present in empty-vector controls (see Fig. S3 in the supplemental material). Since cloning into pET-28a resulted in the N-terminal addition of a His tag to the HHDH, the same bands were also detected in Western blots by making use of a His tag-specific Ni-NTA horseradish peroxidase conjugate (see Fig. S3).
To assess whether these recombinant enzymes possess true HHDH activity, a colorimetric activity assay for halide release from chloroalcohols 1,3-dichloro-2-propanol and 2-chlorophenylethanol, as well as from the bromoalcohol 1,3-dibromo-2-propanol, was performed with extracts of cells containing the recombinantly expressed enzymes (35). Afterwards, gas chromatography analysis was used to confirm the formation of the corresponding epoxides. By using this approach, halide release as well as epoxide formation could be detected with at least one of the three substrates for all but one of the novel HHDHs. Only HheB3 exhibited no activity in these assays. However, this enzyme was also barely expressed in soluble form in E. coli as judged by SDS-PAGE and Western blot analysis (see Fig. S3 in the supplemental material). Figure 3 summarizes the specific activities of the different HHDH-containing CFEs as calculated from the amounts of epoxide formed. CFEs containing the known HHDHs HheA2, HheB2, and HheC were included in the measurements as positive controls.
FIG 3.

Specific activities of CFEs from recombinant HHDH expression. Activities are shown for the formation of epoxides (epichlorohydrin and styrene oxide) from chloroalcohols (1,3-dichloro-2-propanol [filled bars] and 2-chlorophenylethanol [open bars], respectively) after the subtraction of background activities from empty-vector controls. For reference, activity data are also shown for CFEs from the expression of previously known HHDHs (marked with asterisks).
For conversions with 1,3-dibromo-2-propanol, rather high background activities were observed in empty-vector controls. Such high rates have already been reported in the past for the uncatalyzed formation of epoxide from this bromoalcohol (7, 58). In contrast, control reactions with chloroalcohols 1,3-dichloro-2-propanol and 2-chlorophenylethanol showed only marginal epoxide formation (<5%). Therefore, only results for the conversion of both chloroalcohols are given in Fig. 3. Nonetheless, again with the exception of HheB3, all HHDHs showed significantly higher epoxide formation in the conversion of 1,3-dibromo-2-propanol than in control reactions (data not shown).
It should be noted that the specific activities reported are based on the total protein of the CFEs used and thus are not comparable between individual enzymes, since these values are largely affected by differences in protein expression. Similarly, conversion results from independent expression cultures for each HHDH differed in absolute values due to differences in total protein content. Nevertheless, repetition of activity measurements for at least two different CFEs per HHDH confirmed the representativeness of the activity data obtained. Furthermore, the reaction conditions applied were likely not optimal for each enzyme, since the pH and temperature optima of the individual HHDHs are not yet known. Instead, reaction conditions (pH 7.5 and 30°C) that have been used consistently in the past to report HHDH activities were chosen. Therefore, higher activities might be observed if the reaction conditions for each individual enzyme were optimized. Nevertheless, true HHDH activity could be confirmed for 18 out of the 19 HHDHs tested.
Fingerprinting for exclusive recovery of HHDH sequences.
After we confirmed that our approach specifically identifies sequences that exhibit true HHDH activity, we tried to optimize our search routine further so as to identify even more distantly related sequences and, at the same time, simplify the overall procedure.
In our initial blastp searches for novel HHDH sequences, the large majority of results were dominated by SDR sequences due to the large number of SDR enzyme sequences available (27) and the overall high similarity of HHDHs with SDR enzymes (21, 26). As a consequence, distantly related HHDH sequences might not have been included in the 20,000 sequences retrieved per blastp query. To circumvent this, PHI-BLAST can be used to detect more distantly related sequences that match a user-defined pattern (45). In order to recognize such a pattern, the MAFFT alignment of all known and novel HHDH sequences was inspected for the presence of conserved motifs (see Fig. S1 in the supplemental material).
As expected from the outlined identification protocol, all HHDHs necessarily possessed the Ser-Tyr-Arg catalytic triad, with Tyr consistently separated from Arg by three amino acids (Fig. 4B). This is in agreement with the pattern for enzymes of the SDR superfamily, since the corresponding catalytic Tyr and Lys residues are also separated by three amino acids in more than 86% of the classified SDR enzymes that belong to either the classical, extended, or intermediate subfamily (24, 27). Surprisingly, Ser always precedes Tyr by 12 amino acids in HHDH enzymes, whereas the position of this upstream catalytic residue seems to be less conserved in SDR enzymes. For example, 184 of the 201 sequences most similar to HheC from UniProt/Swiss-Prot or PDB (E, >100) possess a catalytic Tyr residue that is separated by three amino acids from Lys and are thus likely classical, extended, or intermediate SDR enzymes. Interestingly, for 18 of these 184 SDR sequences, an upstream catalytic Ser is not separated from Tyr by 12 amino acids; instead, this distance between Ser and Tyr varies by as many as 2 residues. Therefore, in contrast to SDR enzymes, all HHDHs identified seem to possess a catalytic triad that fits the pattern S-X12-Y-X3-R (Fig. 4B) despite overall low sequence identities. Consequently, this catalytic triad motif can serve as a seed pattern for PHI-BLAST searches.
FIG 4.
Partial MAFFT alignment of previously known (asterisked) and novel HHDH sequences. Excerpts from the complete alignment (see Fig. S1 in the supplemental material) show sequences around the conserved Phe or Tyr residue (shaded) (A) or around the Ser-Tyr-Arg catalytic triad residues (shaded) (B) in HHDHs, as well as sequences around the corresponding residues in two homologous, experimentally verified SDR enzymes, FabG and DHRS4. The sequence logos given above each of the alignment excerpts schematize the amino acid distributions observed in all HHDHs or in 718 homologous SDR sequences.
In addition to the HHDH catalytic triad pattern, the conserved aromatic Phe or Tyr residue was utilized to infer another HHDH-specific pattern. As deduced from the sequence logos in Fig. 4A, the position of either aromatic HHDH residue corresponds to the central Gly in the T-(A/G)-X3-G-X-G motif. The latter motif constitutes a variation of the commonly observed nucleotide binding motif, which was observed for a selection of 718 SDR enzymes most homologous to HHDHs. The sequence logo for the corresponding HHDH enzymes revealed the pattern T-X4-(F/Y)-X-G, which is present only in HHDH sequences, instead of the SDR motif (Fig. 4A).
With either pattern as a seed for PHI-BLAST searches, sequences of each phylogenetic HHDH type (see below), namely, HheA, HheB, HheC, HheD, HheE, HheF, and HheG, were used to query the nr database. As anticipated, these PHI-BLAST searches allowed for a much deeper look into sequence space, since it was possible to retrieve sequences with an E value of 10 within the first few hundred hits. In contrast, E values of previous blastp searches did not exceed 0.01 within the first 20,000 results. Within these results, SDR enzymes were again included, but more importantly, so too were all novel HHDHs. Thus, each pattern alone is sufficient to recover the available HHDH sequences independently of the query sequence. Rather than assessing tens of thousands of sequences, it is possible to reduce the sequence pool to only a few hundred candidates without compromising the quality of the result by using either of the two patterns. This measure is far less demanding on the (computational) alignment efforts and might also generate higher-quality MSAs due to the exclusion of unrelated (contaminating) sequences.
To our surprise, when we performed these PHI-BLAST searches, 17 additional sequences that represent putative novel HHDHs were identified in the nr database (Table 2). All these sequences (HheA6 through HheA9, HheD6 through HheD18) possessed the conserved aromatic Phe or Tyr residue (Fig. 4A) together with the catalytic Ser-Tyr-Arg triad (Fig. 4B) and originated from different alpha-, beta-, and gammaproteobacteria. Apparently, these sequences had been included in the updated GenBank release and were now successfully recovered by our optimized PHI-BLAST queries.
TABLE 2.
Sources and accession numbers of putative novel HHDHs
| HHDH | Organism | Protein accession no. |
|---|---|---|
| HheA6 | Bacterium strain Ec32 | CDO61292 |
| HheA7 | Sneathiella glossodoripedis | WP_025899379 |
| HheA8 | Alphaproteobacterium strain Mf 1.05b.01 | WP_029639308 |
| HheA9 | Alphaproteobacterium strain MA2 | GAK44072 |
| HheD6 | Marinobacter nanhaiticus D15-8W | ENO15189 |
| HheD7 | Thauera sp. strain 27 | ENO82779 |
| HheD8 | Thauera aminoaromatica S2 | ENO87252 |
| HheD9 | Thauera phenylacetica B4P | ENO98837 |
| HheD10 | Limnohabitans sp. strain Rim28 | WP_019427705 |
| HheD11 | Thiothrix disciformis | WP_020394200 |
| HheD12 | Pseudomonas pelagia | WP_022962804 |
| HheD13 | Betaproteobacterium strain MOLA814 | ESS13801 |
| HheD14 | Gammaproteobacterium strain MOLA455 | ETN91936 |
| HheD15 | “Candidatus Competibacter denitrificans” | CDI00977 |
| HheD16 | Methylibium sp. strain T29 | EWS52496 |
| HheD17 | Curvibacter gracilis | WP_027476209 |
| HheD18 | Curvibacter lanceolatus | WP_031254602 |
Sequences HheA6 to HheA9 and HheD6 through HheD18 matched the other experimentally verified A- and D-type HHDHs with identities between 32% and 64% and between 59% and 99%, respectively. As seen for the other novel sequences, none of the associated sequence records indicated any relation to HHDH enzymes. Although enzymatic activities have not been verified for these putative novel HHDH sequences, all 17 of them very likely exhibit typical HHDH activity, since their levels of identity to the sequences of experimentally verified enzymes are well within the range observed for the novel HHDHs initially discovered in this study. Currently, these putative novel HHDH sequences are under investigation for their HHDH activities. No further HHDH sequence could be identified by increasing the E cutoff value (up to 1,000) or by using any of the other HHDH sequences as a query.
Although each pattern alone greatly reduces the total number of sequences to be assessed for the second required HHDH sequence feature, SDR sequences were always included in all of the results. For the specific retrieval of HHDH sequences only, it is possible to combine the two patterns so as to identify sequences that must agree in both of the sequence characteristics mentioned above. Since the two patterns are separated by 93 to 131 residues in the experimentally verified HHDHs, they can be combined in the pattern T-X4-(F/Y)-X-G-X93–131-S-X12-Y-X3-R. Now, by using the combined pattern with HheA, HheB, HheC, HheD, HheE, HheF, or HheG as the PHI-BLAST query, only the initially identified (Table 1) and putative novel (Table 2) HHDHs are retrieved as relevant hits (E, <0.001). From the results recovered with the combined-pattern searches, other sequences had E values of at least 0.54 to 10, depending on the query sequence, and were often annotated as large membrane proteins (>1,000 amino acids). Because these other sequences greatly exceeded the conventional lengths observed for SDR enzymes (<350 amino acids), they were not considered to represent enzymes with HHDH activity. Varying the gap length between zero and 250 residues or using any of the other HHDH sequences as the query did not result in the retrieval of any additional sequence with both HHDH sequence features within the boundaries of relevant results.
Despite the apparent impact that positional sequence information has on homology search results, subsequent PSI-BLAST searches did not result in the identification of any further HHDH sequences, although all HHDH sequences identified were included during PSSM construction. Querying the updated env_nr database with separate or combined seed patterns did not result in the recovery of additional putative novel HHDH sequences.
In summary, the use of the restrictive HHDH-specific sequence pattern T-X4-(F/Y)-X-G or S-X12-Y-X3-R, as well as their combination, simplified the overall identification routine. In addition to the 20 HHDHs identified earlier, another 17 putative novel HHDH sequences, which likely possess HHDH activities, were identified (Table 2).
Phylogenetic classification of the HHDH enzyme family.
Since we suspected that our diverse HHDH sequences might be challenging for reliable phylogenetic inference, we decided to use the FastME minimum evolution tree-building algorithm as well as the PhyML maximum likelihood tree-building algorithm, both of which ranked highest in a recent benchmark of phylogenetic tree-building methods (59). Since phylogenetic tree reconstruction can also suffer from substantial bias caused by the underlying MSA, two different algorithms, MAFFT (50) and PRANK +F (60), were used for MSA construction, since both outperformed other algorithms in recent benchmark studies (51, 52, 59). Because it has been suggested that gaps in MSAs carry substantial phylogenetic signals (52, 61), none of the resulting MSAs were curated from unaligned columns.
Overall, all phylogenetic trees generated were very similar independently of the underlying MSA or tree reconstruction algorithm, but corresponding bootstrap values were higher for the PhyML phylogram on the basis of the PRANK +F MSA (Fig. 5) than for the other trees (see Fig. S4 in the supplemental material). In all trees, the previously known enzymes HheA, HheA2, and HheC clustered at a major clade together with novel HHDHs HheA3 through HheA9. This clade of A- and C-type enzymes diverged early from the major clade of remaining HHDHs. In this second major clade, the known HheB and HheB2 enzymes were grouped together with five novel B-type HHDHs of metagenomic origin (HheB3 through HheB7).
FIG 5.
Phylogram of previously known (asterisked) and novel HHDH enzymes. The PhyML tree shown was constructed on the basis of a PRANK +F MSA that included the experimentally verified, homologous SDR enzymes DHRS4 and FabG as an outgroup for rooting (percentages give bootstrap supports at indicated nodes).
In contrast, the remaining novel enzymes formed distinct additional phylogenetic branches, expanding the previous classification of known HHDHs into A-, B-, and C-type enzymes (21). As a consequence, we propose to cluster the novel HHDHs in a total of four additional phylogenetic clades encompassing members of the D-, E-, F-, and G-type enzymes. For example, the 5 experimentally verified and 13 putative novel D-type enzymes clustered together in one clade with shared ancestry with the B-type enzymes. Another clade was formed by enzymes HheE through HheE5, which diverged earlier from the lineage of B- and D-type enzymes. Except for the clade of E-type enzymes, minor differences in topology and bootstrap support were observed between the four trees with regard to each clade's individual members. Thus, it was difficult to conclusively analyze their true internal phylogenetic relation. Nevertheless for the A-, B-, C-, D-, and E-type enzymes, the phylogenetic grouping outlined was observed consistently in all four trees with high bootstrap support.
The branch points of HheF and HheG, however, differed depending on the MSA or tree-building algorithm. In both FastME trees, for example, HheF branched after the segregation of D-type HHDHs but prior to the branch point of B-type enzymes with reasonable bootstrap support (54% and 76%). The PhyML trees, on the other hand, indicated the branch point of HheF before the segregation of B- and D-type enzymes with overall higher bootstrap confidence (86% and 90%). Consistently for all trees, low bootstrap confidence was observed for the placement of HheG (36% to 55%). In both of the trees built on the PRANK +F MSA, HheG branched prior to any other HHDH clade. In contrast, the FastME/MAFFT tree specified the HheG branch point prior to the major clade of A- and C-type enzymes, while HheG diverged within the second major clade of B- through F-type enzymes in the PhyML/MAFFT tree (see Fig. S4 in the supplemental material). Despite these inconsistencies, since the overall placement of HheF and HheG did not occur within any other clade, we decided that HheF and HheG should become archetypes of their own phylogenetic subtype, keeping in mind that, with the data at hand, their actual branch point cannot be determined with absolute certainty.
Despite these minor variations and uncertainties, the overall classification into six different phylogenetic HHDH subtypes was observed for the majority of bootstrap trees independently of the MSA or tree-building algorithm. This suggests that, in principle, the HHDH enzyme family is reliably represented by the phylogram in Fig. 5.
Phylogenetic placement of HHDHs in relation to homologous SDR sequences.
To elucidate the phylogenetic origins of HHDH enzymes, minimum evolution trees including all HHDH sequences identified, as well as a broad selection of SDR sequences, were constructed.
In addition to all HHDHs identified (Tables 1 and 2), 718 unique homologous sequences, which originated from organisms in all three domains of life and which possessed the typical Ser-Tyr-Lys catalytic residues found in SDR enzymes, were used for FastME tree building. Besides sequences with only putative SDR activity, a well-studied human SDR enzyme, the dehydrogenase/reductase SDR family member 4 (DHRS4) (62–64), and experimentally verified 3-ketoacyl-acyl carrier protein reductases from Vibrio harveyi (FabG) (65) and Burkholderia pseudomallei (66) were included.
The resulting tree revealed that HHDH enzymes form a monophyletic clade that does not contain any SDR sequences (see Fig. S5 in the supplemental material). Bootstrap analysis confirmed the proposed branching for the majority of HHDHs (HheA through HheF) for 95% of the bootstrap replicates. Again, as observed for the phylogenetic classification of the HHDH enzyme family (see above), the branching of HheG showed only low bootstrap support (42%), but still, the observed monophyletic clustering of HHDHs was confirmed for the majority of replica trees.
Proposed nomenclature.
As a consequence of the growing number of HHDH sequences, we propose to adopt a general naming scheme for genes and enzymes of the HHDH enzyme family on the basis of their clustering with any of the phylogenetic subtypes A through G and, if necessary, additional phylogenetic subtypes. Then the numbering of enzymes within each subtype shall be consecutive according to the time of submission to public sequence databases such as GenBank. Throughout this study, this nomenclature has already been implemented (Table 1 and 2), which should help to avoid future conflicts or inconsistencies. A repository of available HHDH enzyme sequences will be maintained online and will be updated regularly (http://tiny.cc/hhdhs).
DISCUSSION
The exponentially growing deposition of sequence information in public databases currently outpaces efforts to characterize novel biocatalysts biochemically (67). Moreover, the annotation of a large portion of sequence records lacks proper information on their true activity or, in the worst case, includes no functional information at all (68). For this reason, the accuracy of automated protein function annotation methods is crucially important (69). To aid in the in silico assignment of enzyme functions, specific enzyme family fingerprints can facilitate the identification of novel biocatalysts. In order to expand the current short list of HHDH biocatalysts, we extracted HHDH-specific sequence motifs that allowed us, for the first time, to discern true HHDHs from the large majority of homologous SDR sequences.
First, the consensus motif S-X12-Y-X3-R represents the typical HHDH catalytic triad present in the 5 previously known and 37 novel HHDHs. This consensus motif is more precise than the less-specific pattern S-X7–17-Y-X3-R used in the past without any reports on specific enzyme sequences or their HHDH activities (70). Additionally, motif T-X4-(F/Y)-X-G was deduced, which specifies conserved residues in the nucleophile binding pocket architecture of HHDH enzymes that align with residues of the SDR nucleotide binding motif. The combination of these two features is highly effective in distinguishing HHDHs from SDR enzymes, while the use of the HHDH catalytic triad pattern alone in PHI-BLAST searches also retrieved sequences that possessed only the typical T-G-X3-(G/A)-X-G nucleotide binding motif found in SDR enzymes.
Surprisingly, among the sequences with an HHDH catalytic triad, no amino acid other than the central Phe/Tyr or Gly/Ala typical for HHDHs or SDRs, respectively, was observed for the nucleophile binding motif. This limited diversity was unexpected, since Fox and coworkers observed at least Leu or Ile as nondetrimental substitutions at position F12 in HheC mutants (13). In wild-type HHDHs, the aromatic rings of Phe or Tyr might provide an evolutionary benefit for the stabilization of cleaved halide anions and the positioning of incoming nucleophiles during epoxide ring formation and opening, respectively. Due to the still relatively small number of HHDH sequences identified, however, variations of both motifs might be present in hitherto undescribed enzymes with HHDH activities.
With the combination of both restrictive HHDH-specific sequence patterns, it was possible to precisely identify true HHDH sequences present in public databases and to distinguish them from the vast number of similar SDR sequences. These analyses represent a critically important effort to functionally characterize the enormous number of predicted genes that have unknown or incorrectly annotated functions (68, 71). Especially sequence data of metagenomic origin can be a viable source of novel biocatalysts, but our results reflect the fact that annotated sequences should be utilized only after careful inspection. Eight of the nine HHDH sequences from environmental DNA sources were apparently annotated with likely incorrect start codons. Evidently, the automated gene annotation algorithms now in use require further development and optimization—especially for challenging environmental sequence data that lack host-specific translation initiation signals, such as 16S rRNA sequences.
Of all novel (Table 1) and putative novel (Table 2) HHDHs, only the coding sequence of HheA4 appeared to be part of a degradative operon, since it was flanked by an epoxide hydrolase gene upstream and a glycerol kinase gene downstream. The concerted activity of all three enzymes allows for the utilization of compounds such as 3-chloro-1,2-propanediol via glycidol and eventually glycerol—a pathway that has been proposed for HHDH-containing Arthrobacter sp. strain AD2 and Agrobacterium radiobacter AD1 (72). For the latter strain, the epoxide hydrolase gene echA was cloned (73), and the resulting enzyme exhibited only 31% sequence identity with the epoxide hydrolase from Arthrobacter sp. strain JBH1, a value similar to the identity observed for the respective HHDHs (33%).
Regarding the substrate specificity of each novel HHDH toward the small aliphatic 1,3-dichloro-2-propanol in comparison to the larger aromatic 2-chlorophenylethanol, a first conclusion might be drawn according to Fig. 3. Most HHDHs converted both chloroalcohols tested, but enzymes of the D and E types exhibited a preference for the smaller aliphatic substrate over the aromatic substrate. In contrast, a larger portion of A- and B-type HHDHs showed a significantly higher relative activity in the conversion of the aromatic 2-chlorophenylethanol than in that of the aliphatic 1,3-dichloro-2-propanol. Only for HheB4, HheB6, and HheF are the overall activities observed on both substrates too low to support any conclusion on their substrate preferences. For HheB4 and HheB6, only very small amounts of soluble protein were obtained upon heterologous expression in E. coli, explaining their low activities (see Fig. S3 in the supplemental material). However, for HheF, a significant band of soluble protein can be observed on SDS-PAGE. Despite its rather low activity on 1,3-dichloro-2-propanol, HheF exhibited high activity in reactions using 1,3-dibromo-2-propanol (data not shown). Hence, HheF seems to rather prefer bromosubstituted to chlorosubstituted haloalcohol substrates.
Instead of the traditional classification of HHDHs based on their biochemical and sequence similarities (21), classification of HHDH family enzymes was determined by phylogenetic methods for this study. This approach is especially advantageous in light of the exponential growth of public databases, in which more and more members of this enzyme class will certainly become available over time. Moreover, this objective means of classification has been used, for example, for the unrelated but similarly diverse haloalkane dehalogenase enzyme family (74). While our phylogenetic classification efforts were consistent with the previous HHDH classification (21), several novel HHDHs did not cluster with previous subtypes and therefore had to be grouped into four additional enzyme subtypes. Due to the limited number of HHDH enzymes, however, the phylogenetic relationships of the HHDH enzyme family might be affected by future discoveries of additional family members.
Previously, due to the low sequence identities generally observed, the possibility that HHDH enzymes segregated very early from SDR enzymes was discussed (23, 70). Indeed, our phylogenetic analyses suggest that HHDH enzymes are only distantly related to highly homologous SDR enzymes. As measured by branch lengths from the root of SDR enzymes, the phylogenetic distance for any HHDH enzyme is greater than 0.7 amino acid exchange per residue (see Fig. S5 in the supplemental material). To put this distance into perspective, a recent analysis of DHRS4 and its paralog, the dehydrogenase/reductase SDR family member 2 (DHRS2), concluded that these two human genes diverged from each other before the formation of the mammalian clade (75). Here the phylogenetic distance is only 0.2 amino acid exchange per residue since the event of divergence of the two SDR enzymes (not shown). Hence, the large phylogenetic distances to any SDR homolog that we have observed for all known and newly identified HHDHs indicate that previous assumptions correctly reflect the evolutionary path of HHDH enzymes.
Conclusion.
Public sequence databases hold an immense treasure of biotechnologically relevant enzymes. In the past, many biotechnologically important enzymes have been successfully identified through in silico enzyme discovery. In contrast to enzyme classes that can be found in many different organisms, all of the few HHDH enzymes known previously have been found in species obtained after microbial enrichment techniques (2). HHDHs constitute only a minute fraction of the SDR enzyme superfamily, and on average, only 1 HHDH enzyme sequence can be found among >106 of GenBank's nonredundant protein sequences. Despite these facts, we could show that it is still feasible to extract such rare enzymes from sequence databases after thorough sequence analysis and identification of exclusive sequence motifs. With the help of the conserved sequence fingerprints T-X4-(F/Y)-X-G and S-X12-Y-X3-R, the process of identifying novel HHDH enzymes in the future will be highly effective and much faster than the time-consuming classical microbiology and molecular biology approaches. Especially as an answer to the implications of advancing sequencing technologies, effective in silico methods for the discovery of novel enzymes will become more and more important. Ultimately, we are convinced that the novel halohydrin dehalogenases will enable interesting applications enriching the synthetic chemist's toolbox.
Supplementary Material
ACKNOWLEDGMENTS
We thank Shiva Saraeian for technical assistance. We thank D. B. Janssen for critical reading and suggestions on the manuscript.
This work was financially supported by the German Research Foundation (DFG) within the national Excellence Initiative funding scheme to promote science and research at German universities. Additionally, the work was supported by the German Federal Ministry for Economic Affairs and Energy (BMWi).
M.S., J.K., R.W., and A.S. have filed a patent application on the identification of novel halohydrin dehalogenases from public databases and their use. R.W. is an employee of the company Enzymicals AG interested in the commercialization of biocatalysts.
Footnotes
Published ahead of print 19 September 2014
Supplemental material for this article may be found at http://dx.doi.org/10.1128/AEM.01985-14.
REFERENCES
- 1.Janssen DB, Majeric-Elenkov M, Hasnaoui G, Hauer B, Lutje Spelberg JH. 2006. Enantioselective formation and ring-opening of epoxides catalysed by halohydrin dehalogenases. Biochem. Soc. Trans. 34:291–295. 10.1042/BST20060291. [DOI] [PubMed] [Google Scholar]
- 2.Schallmey M, Floor RJ, Szymanski W, Janssen DB. 2012. Halohydrin dehalogenases, p 143–155 In Carreira EM, Yamamoto H. (ed), Comprehensive chirality. Elsevier, Amsterdam, The Netherlands. [Google Scholar]
- 3.You Z-Y, Liu Z-Q, Zheng Y-G. 2013. Properties and biotechnological applications of halohydrin dehalogenases: current state and future perspectives. Appl. Microbiol. Biotechnol. 97:9–21. 10.1007/s00253-012-4523-0. [DOI] [PubMed] [Google Scholar]
- 4.Nakamura T, Nagasawa T, Yu F, Watanabe I, Yamada H. 1992. Resolution and some properties of enzymes involved in enantioselective transformation of 1,3-dichloro-2-propanol to (R)-3-chloro-1,2-propanediol by Corynebacterium sp. strain N-1074. J. Bacteriol. 174:7613–7619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Lutje Spelberg JH, van Hylckama Vlieg JET, Bosma T, Kellogg RM, Janssen DB. 1999. A tandem enzyme reaction to produce optically active halohydrins, epoxides and diols. Tetrahedron Asymmetry 10:2863–2870. 10.1016/S0957-4166(99)00308-0. [DOI] [Google Scholar]
- 6.Haak RM, Tarabiono C, Janssen DB, Minnaard AJ, de Vries JG, Feringa BL. 2007. Synthesis of enantiopure chloroalcohols by enzymatic kinetic resolution. Org. Biomol. Chem. 5:318–323. 10.1039/b613937j. [DOI] [PubMed] [Google Scholar]
- 7.Haak RM, Berthiol F, Jerphagnon T, Gayet AJA, Tarabiono C, Postema CP, Ritleng V, Pfeffer M, Janssen DB, Minnaard AJ, Feringa BL, de Vries JG. 2008. Dynamic kinetic resolution of racemic β-haloalcohols: direct access to enantioenriched epoxides. J. Am. Chem. Soc. 130:13508–13509. 10.1021/ja805128x. [DOI] [PubMed] [Google Scholar]
- 8.Jin H-X, Hu Z-C, Liu Z-Q, Zheng Y-G. 2012. Nitrite-mediated synthesis of chiral epichlorohydrin using halohydrin dehalogenase from Agrobacterium radiobacter AD1. Biotechnol. Appl. Biochem. 59:170–177. 10.1002/bab.1004. [DOI] [PubMed] [Google Scholar]
- 9.Majeric Elenkov M, Primozic I, Hrenar T, Smolko A, Dokli I, Salopek-Sondi B, Tang L. 2012. Catalytic activity of halohydrin dehalogenases towards spiroepoxides. Org. Biomol. Chem. 10:5063–5072. 10.1039/c2ob25470k. [DOI] [PubMed] [Google Scholar]
- 10.Hasnaoui-Dijoux G, Majeric Elenkov M, Lutje Spelberg JH, Hauer B, Janssen DB. 2008. Catalytic promiscuity of halohydrin dehalogenase and its application in enantioselective epoxide ring opening. Chembiochem 9:1048–1051. 10.1002/cbic.200700734. [DOI] [PubMed] [Google Scholar]
- 11.Nakamura T, Nagasawa T, Yu F, Watanabe I, Yamada H. 1994. A new enzymatic synthesis of (R)-γ-chloro-β-hydroxybutyronitrile. Tetrahedron 50:11821–11826. 10.1016/S0040-4020(01)89297-8. [DOI] [Google Scholar]
- 12.Lutje Spelberg JH, Tang L, Kellogg RM, Janssen DB. 2004. Enzymatic dynamic kinetic resolution of epihalohydrins. Tetrahedron Asymmetry 15:1095–1102. 10.1016/j.tetasy.2004.02.009. [DOI] [Google Scholar]
- 13.Fox RJ, Davis SC, Mundorff EC, Newman LM, Gavrilovic V, Ma SK, Chung LM, Ching C, Tam S, Muley S, Grate J, Gruber J, Whitman JC, Sheldon RA, Huisman GW. 2007. Improving catalytic function by ProSAR-driven enzyme evolution. Nat. Biotechnol. 25:338–344. 10.1038/nbt1286. [DOI] [PubMed] [Google Scholar]
- 14.Majeric Elenkov M, Hauer B, Janssen DB. 2006. Enantioselective ring opening of epoxides with cyanide catalysed by halohydrin dehalogenases: a new approach to non-racemic β-hydroxy nitriles. Adv. Synth. Catal. 348:579–585. 10.1002/adsc.200505333. [DOI] [Google Scholar]
- 15.Majeric Elenkov M, Hoeffken HW, Tang L, Hauer B, Janssen DB. 2007. Enzyme-catalyzed nucleophilic ring opening of epoxides for the preparation of enantiopure tertiary alcohols. Adv. Synth. Catal. 349:2279–2285. 10.1002/adsc.200700146. [DOI] [Google Scholar]
- 16.Fuchs M, Simeo Y, Ueberbacher BT, Mautner B, Netscher T, Faber K. 2009. Enantiocomplementary chemoenzymatic asymmetric synthesis of (R)- and (S)-chromanemethanol. Eur. J. Org. Chem. 2009:833–840. 10.1002/ejoc.200800950. [DOI] [Google Scholar]
- 17.Molinaro C, Guilbault A-A, Kosjek B. 2010. Resolution of 2,2-disubstituted epoxides via biocatalytic azidolysis. Org. Lett. 12:3772–3775. 10.1021/ol101406k. [DOI] [PubMed] [Google Scholar]
- 18.Cozzi PG, Hilgraf R, Zimmermann N. 2007. Enantioselective catalytic formation of quaternary stereogenic centers. Eur. J. Org. Chem. 2007:5969–5994. 10.1002/ejoc.200700318. [DOI] [Google Scholar]
- 19.Kourist R, Bornscheuer UT. 2011. Biocatalytic synthesis of optically active tertiary alcohols. Appl. Microbiol. Biotechnol. 91:505–517. 10.1007/s00253-011-3418-9. [DOI] [PubMed] [Google Scholar]
- 20.Yu F, Nakamura T, Mizunashi W, Watanabe I. 1994. Cloning of two halohydrin hydrogen-halide-lyase genes of Corynebacterium sp. strain N-1074 and structural comparison of the genes and gene products. Biosci. Biotechnol. Biochem. 58:1451–1457. 10.1271/bbb.58.1451. [DOI] [PubMed] [Google Scholar]
- 21.van Hylckama Vlieg JET, Tang L, Lutje Spelberg JH, Smilda T, Poelarends GJ, Bosma T, van Merode AEJ, Fraaije MW, Janssen DB. 2001. Halohydrin dehalogenases are structurally and mechanistically related to short-chain dehydrogenases/reductases. J. Bacteriol. 183:5058–5066. 10.1128/JB.183.17.5058-5066.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Higgins TP, Hope SJ, Effendi AJ, Dawson S, Dancer BN. 2005. Biochemical and molecular characterisation of the 2,3-dichloro-1-propanol dehalogenase and stereospecific haloalkanoic dehalogenases from a versatile Agrobacterium sp. Biodegradation 16:485–492. 10.1007/s10532-004-5670-5. [DOI] [PubMed] [Google Scholar]
- 23.de Jong RM, Dijkstra BW. 2003. Structure and mechanism of bacterial dehalogenases: different ways to cleave a carbon-halogen bond. Curr. Opin. Struct. Biol. 13:722–730. 10.1016/j.sbi.2003.10.009. [DOI] [PubMed] [Google Scholar]
- 24.Kavanagh KL, Jörnvall H, Persson B, Oppermann U. 2008. Medium- and short-chain dehydrogenase/reductase gene and protein families. Cell. Mol. Life Sci. 65:3895–3906. 10.1007/s00018-008-8588-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.de Jong RM, Kalk KH, Tang L, Janssen DB, Dijkstra BW. 2006. The X-ray structure of the haloalcohol dehalogenase HheA from Arthrobacter sp. strain AD2: insight into enantioselectivity and halide binding in the haloalcohol dehalogenase family. J. Bacteriol. 188:4051–4056. 10.1128/JB.01866-05. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.de Jong RM, Tiesinga JJW, Rozeboom HJ, Kalk KH, Tang L, Janssen DB, Dijkstra BW. 2003. Structure and mechanism of a bacterial haloalcohol dehalogenase: a new variation of the short-chain dehydrogenase/reductase fold without an NAD(P)H binding site. EMBO J. 22:4933–4944. 10.1093/emboj/cdg479. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Persson B, Kallberg Y. 2013. Classification and nomenclature of the superfamily of short-chain dehydrogenases/reductases (SDRs). Chem. Biol. Interact. 202:111–115. 10.1016/j.cbi.2012.11.009. [DOI] [PubMed] [Google Scholar]
- 28.Bornscheuer UT, Huisman GW, Kazlauskas RJ, Lutz S, Moore JC, Robins K. 2012. Engineering the third wave of biocatalysis. Nature 485:185–194. 10.1038/nature11117. [DOI] [PubMed] [Google Scholar]
- 29.Tang L, Jiang R, Zheng K, Zhu X. 2011. Enhancing the recombinant protein expression of halohydrin dehalogenase HheA in Escherichia coli by applying a codon optimization strategy. Enzyme Microb. Technol. 49:395–401. 10.1016/j.enzmictec.2011.06.021. [DOI] [PubMed] [Google Scholar]
- 30.Tang L, Gao H, Zhu X, Wang X, Zhou M, Jiang R. 2012. Construction of “small-intelligent” focused mutagenesis libraries using well-designed combinatorial degenerate primers. Biotechniques 52:149–158. [DOI] [PubMed] [Google Scholar]
- 31.Tang L, Zhu X, Zheng H, Jiang R, Elenkov MM. 2012. Key residues for controlling enantioselectivity of halohydrin dehalogenase from Arthrobacter sp. strain AD2, revealed by structure-guided directed evolution. Appl. Environ. Microbiol. 78:2631–2637. 10.1128/AEM.06586-11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Tang L, van Hylckama Vlieg JET, Lutje Spelberg JH, Fraaije MW, Janssen DB. 2002. Improved stability of halohydrin dehalogenase from Agrobacterium radiobacter AD1 by replacement of cysteine residues. Enzyme Microb. Technol. 30:251–258. 10.1016/S0141-0229(01)00488-4. [DOI] [Google Scholar]
- 33.Tang L, Torres Pazmino DE, Fraaije MW, de Jong RM, Dijkstra BW, Janssen DB. 2005. Improved catalytic properties of halohydrin dehalogenase by modification of the halide-binding site. Biochemistry 44:6609–6618. 10.1021/bi047613z. [DOI] [PubMed] [Google Scholar]
- 34.Tang L, Li Y, Wang X. 2010. A high-throughput colorimetric assay for screening halohydrin dehalogenase saturation mutagenesis libraries. J. Biotechnol. 147:164–168. 10.1016/j.jbiotec.2010.04.002. [DOI] [PubMed] [Google Scholar]
- 35.Schallmey M, Floor RJ, Hauer B, Breuer M, Jekel PA, Wijma HJ, Dijkstra BW, Janssen DB. 2013. Biocatalytic and structural properties of a highly engineered halohydrin dehalogenase. Chembiochem 14:870–881. 10.1002/cbic.201300005. [DOI] [PubMed] [Google Scholar]
- 36.Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ. 1997. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 25:3389–3402. 10.1093/nar/25.17.3389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW. 2013. GenBank. Nucleic Acids Res. 41:D36–D42. 10.1093/nar/gks1195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Katoh K, Misawa K, Kuma K, Miyata T. 2002. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 30:3059–3066. 10.1093/nar/gkf436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Katoh K, Kuma K, Toh H, Miyata T. 2005. MAFFT version 5: improvement in accuracy of multiple sequence alignment. Nucleic Acids Res. 33:511–518. 10.1093/nar/gki198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Rosen GL, Reichenberger ER, Rosenfeld AM. 2011. NBC: the Naïve Bayes Classification tool webserver for taxonomic classification of metagenomic reads. Bioinformatics 27:127–129. 10.1093/bioinformatics/btq619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Raab D, Graf M, Notka F, Schödl T, Wagner R. 2010. The GeneOptimizer algorithm: using a sliding window approach to cope with the vast sequence space in multiparameter DNA sequence optimization. Syst. Synth. Biol. 4:215–225. 10.1007/s11693-010-9062-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Crooks GE, Hon G, Chandonia J-M, Brenner SE. 2004. WebLogo: a sequence logo generator. Genome Res. 14:1188–1190. 10.1101/gr.849004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.UniProt Consortium. 2013. Update on activities at the Universal Protein Resource (UniProt) in 2013. Nucleic Acids Res. 41:D43–D47. 10.1093/nar/gks1068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, Shindyalov IN, Bourne PE. 2000. The Protein Data Bank. Nucleic Acids Res. 28:235–242. 10.1093/nar/28.1.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zhang Z, Miller W, Schäffer AA, Madden TL, Lipman DJ, Koonin EV, Altschul SF. 1998. Protein sequence similarity searches using patterns as seeds. Nucleic Acids Res. 26:3986–3990. 10.1093/nar/26.17.3986. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Löytynoja A, Goldman N. 2010. webPRANK: a phylogeny-aware multiple sequence aligner with interactive alignment browser. BMC Bioinformatics 11:579. 10.1186/1471-2105-11-579. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Desper R, Gascuel O. 2002. Fast and accurate phylogeny reconstruction algorithms based on the minimum-evolution principle. J. Comput. Biol. 9:687–705. 10.1089/106652702761034136. [DOI] [PubMed] [Google Scholar]
- 48.Felsenstein J. 1989. PHYLIP—Phylogeny Inference Package (version 3.2). Cladistics 5:164–166. [Google Scholar]
- 49.Guindon S, Dufayard J-F, Lefort V, Anisimova M, Hordijk W, Gascuel O. 2010. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Syst. Biol. 59:307–321. 10.1093/sysbio/syq010. [DOI] [PubMed] [Google Scholar]
- 50.Katoh K, Standley DM. 2013. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol. 30:772–780. 10.1093/molbev/mst010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Dessimoz C, Gil M. 2010. Phylogenetic assessment of alignments reveals neglected tree signal in gaps. Genome Biol. 11:R37. 10.1186/gb-2010-11-4-r37. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Thompson JD, Linard B, Lecompte O, Poch O. 2011. A comprehensive benchmark study of multiple sequence alignment methods: current challenges and future perspectives. PLoS One 6:e18093. 10.1371/journal.pone.0018093. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Marchler-Bauer A, Zheng C, Chitsaz F, Derbyshire MK, Geer LY, Geer RC, Gonzales NR, Gwadz M, Hurwitz DI, Lanczycki CJ, Lu F, Lu S, Marchler GH, Song JS, Thanki N, Yamashita RA, Zhang D, Bryant SH. 2013. CDD: conserved domains and protein three-dimensional structure. Nucleic Acids Res. 41:D348–D352. 10.1093/nar/gks1243. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Marchler-Bauer A, Bryant SH. 2004. CD-Search: protein domain annotations on the fly. Nucleic Acids Res. 32:W327–W331. 10.1093/nar/gkh454. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Xu Y, Kersten RD, Nam S-J, Lu L, Al-Suwailem AM, Zheng H, Fenical W, Dorrestein PC, Moore BS, Qian P-Y. 2012. Bacterial biosynthesis and maturation of the didemnin anti-cancer agents. J. Am. Chem. Soc. 134:8625–8632. 10.1021/ja301735a. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Yooseph S, Sutton G, Rusch DB, Halpern AL, Williamson SJ, Remington K, Eisen JA, Heidelberg KB, Manning G, Li W, Jaroszewski L, Cieplak P, Miller CS, Li H, Mashiyama ST, Joachimiak MP, van Belle C, Chandonia J-M, Soergel DA, Zhai Y, Natarajan K, Lee S, Raphael BJ, Bafna V, Friedman R, Brenner SE, Godzik A, Eisenberg D, Dixon JE, Taylor SS, Strausberg RL, Frazier M, Venter JC. 2007. The Sorcerer II global ocean sampling expedition: expanding the universe of protein families. PLoS Biol. 5:e16. 10.1371/journal.pbio.0050016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Rusch DB, Halpern AL, Sutton G, Heidelberg KB, Williamson S, Yooseph S, Wu D, Eisen JA, Hoffman JM, Remington K, Beeson K, Tran B, Smith H, Baden-Tillson H, Stewart C, Thorpe J, Freeman J, Andrews-Pfannkoch C, Venter JE, Li K, Kravitz S, Heidelberg JF, Utterback T, Rogers Y-H, Falcon LI, Souza V, Bonilla-Rosso G, Eguiarte LE, Karl DM, Sathyendranath S, Platt T, Bermingham E, Gallardo V, Tamayo-Castillo G, Ferrari MR, Strausberg RL, Nealson K, Friedman R, Frazier M, Venter JC. 2007. The Sorcerer II global ocean sampling expedition: Northwest Atlantic through Eastern Tropical Pacific. PLoS Biol. 5:e77. 10.1371/journal.pbio.0050077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Tang L, Lutje Spelberg JH, Fraaije MW, Janssen DB. 2003. Kinetic mechanism and enantioselectivity of halohydrin dehalogenase from Agrobacterium radiobacter. Biochemistry 42:5378–5386. 10.1021/bi0273361. [DOI] [PubMed] [Google Scholar]
- 59.Gonnet GH. 2012. Surprising results on phylogenetic tree building methods based on molecular sequences. BMC Bioinformatics 13:148. 10.1186/1471-2105-13-148. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Löytynoja A, Goldman N. 2008. Phylogeny-aware gap placement prevents errors in sequence alignment and evolutionary analysis. Science 320:1632–1635. 10.1126/science.1158395. [DOI] [PubMed] [Google Scholar]
- 61.Golubchik T, Wise MJ, Easteal S, Jermiin LS. 2007. Mind the gaps: evidence of bias in estimates of multiple sequence alignments. Mol. Biol. Evol. 24:2433–2442. 10.1093/molbev/msm176. [DOI] [PubMed] [Google Scholar]
- 62.Matsunaga T, Endo S, Maeda S, Ishikura S, Tajima K, Tanaka N, Nakamura KT, Imamura Y, Hara A. 2008. Characterization of human DHRS4: an inducible short-chain dehydrogenase/reductase enzyme with 3β-hydroxysteroid dehydrogenase activity. Arch. Biochem. Biophys. 477:339–347. 10.1016/j.abb.2008.06.002. [DOI] [PubMed] [Google Scholar]
- 63.Yan Y, Song X, Liu G, Su Z, Du Y, Sui X, Chang X, Huang D. 2012. Human NRDRB1, an alternatively spliced isoform of NADP(H)-dependent retinol dehydrogenase/reductase enhanced enzymatic activity of benzil. Cell. Physiol. Biochem. 30:1371–1382. 10.1159/000343326. [DOI] [PubMed] [Google Scholar]
- 64.Song X-H, Liang B, Liu G-F, Li R, Xie J-P, Du K, Huang D-Y. 2007. Expression of a novel alternatively spliced variant of NADP(H)-dependent retinol dehydrogenase/reductase with deletion of exon 3 in cervical squamous carcinoma. Int. J. Cancer 120:1618–1626. 10.1002/ijc.22306. [DOI] [PubMed] [Google Scholar]
- 65.Shen Z, Byers DM. 1996. Isolation of Vibrio harveyi acyl carrier protein and the fabG, acpP, and fabF genes involved in fatty acid biosynthesis. J. Bacteriol. 178:571–573. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Baugh L, Gallagher LA, Patrapuvich R, Clifton MC, Gardberg AS, Edwards TE, Armour B, Begley DW, Dieterich SH, Dranow DM, Abendroth J, Fairman JW, Fox D, III, Staker BL, Phan I, Gillespie A, Choi R, Nakazawa-Hewitt S, Nguyen MT, Napuli A, Barrett L, Buchko GW, Stacy R, Myler PJ, Stewart LJ, Manoil C, Van Voorhis WC. 2013. Combining functional and structural genomics to sample the essential Burkholderia structome. PLoS One 8:e53851. 10.1371/journal.pone.0053851. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Gerlt JA, Allen KN, Almo SC, Armstrong RN, Babbitt PC, Cronan JE, Dunaway-Mariano D, Imker HJ, Jacobson MP, Minor W, Poulter CD, Raushel FM, Sali A, Shoichet BK, Sweedler JV. 2011. The enzyme function initiative. Biochemistry 50:9950–9962. 10.1021/bi201312u. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Schnoes AM, Brown SD, Dodevski I, Babbitt PC. 2009. Annotation error in public databases: misannotation of molecular function in enzyme superfamilies. PLoS Comput. Biol. 5:e1000605. 10.1371/journal.pcbi.1000605. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Radivojac P, Clark WT, Oron TR, Schnoes AM, Wittkop T, Sokolov A, Graim K, Funk C, Verspoor K, Ben-Hur A, Pandey G, Yunes JM, Talwalkar AS, Repo S, Souza ML, Piovesan D, Casadio R, Wang Z, Cheng J, Fang H, Gough J, Koskinen P, Törönen P, Nokso-Koivisto J, Holm L, Cozzetto D, Buchan DWA, Bryson K, Jones DT, Limaye B, Inamdar H, Datta A, Manjari SK, Joshi R, Chitale M, Kihara D, Lisewski AM, Erdin S, Venner E, Lichtarge O, Rentzsch R, Yang H, Romero AE, Bhat P, Paccanaro A, Hamp T, Kaßner R, Seemayer S, Vicedo E, Schaefer C, Achten D, Auer F, Boehm A, Braun T, et al. 2013. A large-scale evaluation of computational protein function prediction. Nat. Methods 10:221–227. 10.1038/nmeth.2340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Janssen DB, Dinkla IJT, Poelarends GJ, Terpstra P. 2005. Bacterial degradation of xenobiotic compounds: evolution and distribution of novel enzyme activities. Environ. Microbiol. 7:1868–1882. 10.1111/j.1462-2920.2005.00966.x. [DOI] [PubMed] [Google Scholar]
- 71.Galperin MY, Koonin EV. 2010. From complete genome sequence to “complete” understanding? Trends Biotechnol. 28:398–406. 10.1016/j.tibtech.2010.05.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Van den Wijngaard AJ, Janssen DB, Witholt B. 1989. Degradation of epichlorohydrin and halohydrins by bacterial cultures isolated from freshwater sediment. J. Gen. Microbiol. 135:2199–2208. [Google Scholar]
- 73.Rink R, Fennema M, Smids M, Dehmel U, Janssen DB. 1997. Primary structure and catalytic mechanism of the epoxide hydrolase from Agrobacterium radiobacter AD1. J. Biol. Chem. 272:14650–14657. 10.1074/jbc.272.23.14650. [DOI] [PubMed] [Google Scholar]
- 74.Chovancova E, Kosinski J, Bujnicki JM, Damborsky J. 2007. Phylogenetic analysis of haloalkane dehalogenases. Proteins 67:305–316. 10.1002/prot.21313. [DOI] [PubMed] [Google Scholar]
- 75.Gabrielli F, Tofanelli S. 2012. Molecular and functional evolution of human DHRS2 and DHRS4 duplicated genes. Gene 511:461–469. 10.1016/j.gene.2012.09.013. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




