Skip to main content
Molecular Biology and Evolution logoLink to Molecular Biology and Evolution
. 2025 Aug 19;42(9):msaf203. doi: 10.1093/molbev/msaf203

The Evolutionary History and Modern Diversity of Triterpenoid Cyclases

Hanon Solomon McShea 1,2, Robb A Viens 3, Babatunde O Olagunju 4, José-Luis Giner 5, Paula V Welander 6,✉,b
Editor: Fabia Ursula Battistuzzi
PMCID: PMC12414744  PMID: 40827364

Abstract

Cyclic terpenoids are a class of lipid compounds containing immense structural and functional diversity, with many cyclic triterpenoids acting as regulators of the physical properties and spatial organization of lipid membranes. Cyclic terpenoids are also readily preserved as terpane fossils, such as steranes and hopanes, forming a rich record of the evolution of life on Earth. Formation of the multiple ring structure of all cyclic terpenoids is catalyzed by terpenoid cyclase enzymes, among which are whole clades of proteins—many from environmental metagenomes and uncultured organisms—whose substrates and products are completely unknown. We investigate the function of these divergent cyclases through biochemical assays, and the evolutionary processes that produced them by testing and applying a variety of evolutionary models. We find deep divergence between the diterpenoid cyclases and triterpenoid cyclases, with other clades branching between the two, rooting the triterpenoid cyclase subtree between squalene-hopene cyclases and sterol cyclases. Through a simple test of evolutionary rate shifts, we find an elevated evolutionary rate in the enzyme active site on the squalene-hopene cyclase stem, potentially indicative of positive selection. Finally, by testing the activity of divergent cyclases for a variety of substrates, we find a group of early branching sterol cyclases from bacteria that synthesize arborinols, two of which produce the molecular precursor to a Permian “orphan biomarker.” Together, our data present an evolutionary framework for triterpenoid cyclases that can inform both the biochemical potential of these proteins and their products’ occurrence in the geological record.

Keywords: protein evolution, geological biomarkers, triterpenoid cyclases, sterols, hopanoids

Introduction

In 1971, “geohopanes” were discovered in vast quantities in organic-rich sediments, ancient rocks, and fuel reservoirs (Ourisson and Albrecht 1992). In the following years, hopanoid molecules from diverse bacteria were identified as the biological source of these molecular fossils (Förster et al. 1973; Ourisson and Rohmer 1992). These findings precipitated a new era of research on the taxonomic distribution and physiological function of cyclic triterpenoids (hopanoids, sterols, and similar molecules shown in Fig. 1a) and the triterpenoid cyclase enzymes that synthesize them, motivated by questions of what the molecular fossil record could reveal about the evolution of life on Earth, and of Earth itself (e.g. Runnegar 1991; Ourisson and Nakatani 1994; Schaeffer et al. 1994; Fischer 2008). This body of research has uncovered diverse triterpane fossils, as well as an increasingly deep understanding of cyclic terpenoid structural diversity (Hayashi et al. 2007; Kontnik et al. 2008; Moosmann et al. 2020; Rudolf et al. 2021), taxonomic distribution (Pearson et al. 2003; Ricci et al. 2014; Villanueva et al. 2014; Wei et al. 2016; Takishita et al. 2017; Mayer et al. 2021), biosynthesis (Welander et al. 2010; Banta et al. 2015; Pan et al. 2015; Lee et al. 2018, 2023; Pollier et al. 2019; Brown et al. 2023), and function (Flesch and Rohmer 1987; Welander et al. 2009, 2012; Welander and Summons 2012; Schmerk et al. 2015; Brenac et al. 2019; Gudde et al. 2019; Rivas-Marin et al. 2019; Zhai et al. 2024). Unsurprisingly, the catalogues of triterpane fossils and modern triterpenoids are not isomorphic—there are many “orphan biomarkers” in the geological record with no known biological source (Brocks and Pearson 2005), and conversely, there are many organisms and environments that have genes for triterpenoid cyclases, but no known biosynthetic products (Pearson et al. 2007). To reconcile such orphan biomarkers and mystery cyclases, we analyzed the modern diversity and evolutionary history of the terpenoid cyclase protein superfamily, of which triterpenoid cyclases are one protein family.

Fig. 1.

Fig. 1.

Reactions performed by Type II terpenoid cyclases, including on a) C30 substrates in the case of triterpenoid cyclases (with annotation showing important chemical differences among products), b) C20 substrates in the case of diterpenoid cyclases, and c) C15, C25, C35, and C40 substrates for other Type II terpenoid cyclases. Other reactions exist for some substrates; for example, in addition to lanosterol cyclases, there are oxidosqualene cyclases that synthesize parkeol, cycloartenol, or amyrins. d) The crystal structure of squalene-hopene cyclase from Alicyclobacillus acidocaldarius (Wendt et al. 1999). SHC, squalene-hopene cyclase; STC, tetrahymanol cyclase; OSC, oxidosqualene cyclase; OIC, oxidosqualene-isoarborinol cyclase; BjCPS, Bradyrhizobium japonicum copalyl diphosphate synthase; Bra4, brasilicardin cyclase; MstE, merosterolic acid synthase; DMS, drimenol synthase; AtoE, atolypene synthase; SqhC, sporulenol synthase; Lon15, terpene cyclase (longestin).

Terpenoid cyclases rearrange carbon–carbon bonds to form polycyclic molecules through an exquisitely precise carbocation cascade (Hoshino and Sato 2002; Eschenmoser and Arigoni 2005; Abe et al. 2014). We refer here specifically to the Type II terpenoid cyclase protein family, which protonates substrates with a general acid, and is distinct from the Type I terpenoid cyclase family, which adopts a different fold and employs a different, metal-dependent mechanism to cyclize C10, C15, and C20 substrates (Christianson 2017). Type II terpenoid cyclases typically have a dumbbell shape composed of two homologous structural domains known as the β and γ domains. The enzyme active site is between the two domains, accessible by a substrate channel through the γ domain, while the β domain has a catalytic aspartic acid, and a water tunnel needed to replenish it. Some Type II cyclases have additional domains or lack the γ domain. All domains share the alpha–alpha toroid fold, which is a torus of 12 alpha helices arranged in two concentric rings of six (Fig. 1). As described by the structural classification of proteins—extended 2.08 (Fox et al. 2014; Chandonia et al. 2022), terpenoid cyclases and protein prenyltransferases make up one of the six superfamilies in the alpha–alpha toroid fold, the others being glycosidases and polysaccharide lyases (Syrén et al. 2016), an enzyme involved in the synthesis of lanthipeptide antibiotics (Zhang et al. 2012), and eukaryotic α2-macroglobulins and complement proteins (Cao et al. 2010).

Type II terpenoid cyclases include triterpenoid (C30 substrate) and diterpenoid (C20 substrate) cyclases, which share a catalytic mechanism and βγ domain architecture (Christianson 2017). Triterpenoid cyclases include oxidosqualene cyclases, which cyclize an epoxidized substrate to form tetracyclic sterols such as lanosterol, parkeol, cycloartenol, arborinols, and amyrins (Xu et al. 2004). These lipid alcohols are often modified to form functional membrane components such as cholesterol. Squalene–hopene cyclases and tetrahymanol cyclases cyclize squalene to form pentacyclic hopenes and tetrahymanol, respectively. Unmodified tetrahymanol functions as a membrane lipid (Conner et al. 1971), and hopanoids most often undergo polyfunctionalization at the “tail” alkene to become amphipathic membrane components as well (Bradley et al. 2010; Belin et al. 2018). Diterpenoid cyclases cyclize geranylgeranyl pyrophosphate (GGPP) (C20) to form bicyclic copalyl (Ikeda et al. 2007; Morrone et al. 2009; Smanski et al. 2011), halimadienyl (Nakano et al. 2005), terpentedienyl (Dairi et al. 2001; Stowell et al. 2022), and kolavenyl (Nakano et al. 2015) diphosphates, among other molecules. Other diterpenoid cyclases work on epoxidized GGPP, forming tricyclic molecules such as brasilicardin (Hayashi et al. 2008), phenalinolactone (Dürr et al. 2006), and tiancilactone (Dong et al. 2018) precursors. As shown in Fig. 1, the protein superfamily includes the above and also C15 (Pan et al. 2022; Vo et al. 2022), C25 (Kim et al. 2019), C35 (Kontnik et al. 2008; Sato, Hoshino et al. 2011; Sato, Yoshida et al. 2011), and C40 (Hayashi et al. 2007; Ozaki et al. 2018) cyclases, which work on a variety of alkenyl, epoxidized, and otherwise functionalized isoprenoid substrates.

Products of triterpenoid cyclases often undergo modification by other enzymes and go on to regulate membrane homeostasis and dynamics, among other functions (Bloch 1983). The best-studied example is cholesterol in the animal cell membrane, where it buffers membrane flexibility against temperature changes (Crockett 1998; Mouritsen and Zuckermann 2004) and organizes the membrane into liquid-ordered microdomains (Xu and London 2000; Levental et al. 2020). There is evidence for similar function for hopanoids in bacteria (López and Kolter 2010; Sáenz 2010; Sáenz et al. 2012, 2015). Cyclic diterpenoids, on the other hand, are commonly modified with solubilizing functional groups after cyclization, and function in cell–cell and interspecific communication. Meanwhile, diterpenoids cyclized by Type I cyclases are often insoluble and are involved in other functions such as plant wound healing in the form of resins. Other well-studied diterpenoids include the gibberellin hormones made by plants and their symbionts to regulate tissue growth (Nett et al. 2022). Cyclic diterpenoids are also involved in interspecies conflict: between microbes in the form of antibiotics such as terpentecin (Tamamura et al. 1985), platencin (Jayasuriya et al. 2007; Wang et al. 2007), platensimycin (Wang et al. 2006), and phenalinolactone (Gebhardt et al. 2011), antifungals such as viguiepinol and oxaloterpins (Bi and Yu 2016), and between humans and their pathogens in the case of tuberculosinol from Mycobacterium tuberculosis (Mann et al. 2009) and brasilicardin A from Nocardia brasiliensis (Shigemori et al. 1998; Usui et al. 2006).

The evolutionary processes that generated the structural and functional diversity of terpenoids and of triterpenoids in particular have long been a subject of great scientific interest (Bloch 1983; Ourisson 1989; Ourisson and Nakatani 1994). Fischer and Pearson (2007) presented several alternative models for relationships among triterpenoid cyclases under parsimonious evolution of three functional traits, which they posit evolve independently and slowly, if not irreversibly: substrate (squalene vs. oxidosqualene), reaction favorability (number of anti-Markovnikov carbocations propagated during the reaction), and product stereochemistry (whether the first three rings are chair–chair–chair vs. chair–boat–chair). In this study, we adopt their rigorous evolutionary approach and update it with four new sources of information, in pursuit of the same biological question: in what order, and through what processes, did the diversity of polycyclic membrane regulators arise?

We first apply a probabilistic method (maximum likelihood) of estimating relationships among cyclases, using an explicit model of protein sequence evolution. This allows us to relax Fischer and Pearson's (2007) assumptions about the independence of the traits and the (slow) speed at which they evolve. Therefore, we can observe the distribution of the selected cyclase traits across a tree generated under unrelated assumptions (amino acid exchangeabilities and rate distributions) rather than using these trait assumptions to build the tree. We also revise the list of traits based on functional studies of cyclases over the past 18 years. Our use of probabilistic phylogenetic methods follows that of previous workers (Desmond and Gribaldo 2009; Frickey and Kannenberg 2009; Gold et al. 2017; Santana-Molina et al. 2020), and we find that our analyses, where they overlap, recapitulate theirs. Second, we analyze the astounding cyclase diversity since sequenced, with special attention to divergent sequences from environmental metagenomes and metagenome-assembled genomes as a potential source of cyclases that synthesize orphan biomarkers or novel natural products. Previous environmental sequencing efforts focused on triterpenoid cyclases found entire clades comprised solely of metagenomic sequences (Pearson et al. 2007; Pearson and Rusch 2009), all of which remain uncharacterized. Third, we broaden the analysis to include the entire terpenoid cyclase protein superfamily, allowing us to root the triterpenoid cyclase protein family tree, again independently of assumptions about the manner in which functional traits evolve. Fourth and finally, using the evolutionary framework provided by an explicit model of protein evolution, sequence diversity, and the alpha–alpha toroid fold, we identify cyclases that branch at key transitions in the superfamily's evolution. We test the function of these cyclases via heterologous expression in Escherichia coli and find two surprising enzymes. One is a cyclase that synthesizes isoarborinol, an orphan biomarker from the Paleozoic. The other is a group of cyclases within the squalene–hopene cyclase clade which cyclize oxidosqualene (the native sterol cyclase substrate) but not squalene (the native squalene–hopene cyclase substrate). In addition to other lines of evidence, the substrate specificity of these cyclases suggests that alkenyl substrates may not be ancestral to triterpenoid cyclases, and were gained relatively late in the squalene–hopene cyclase lineage.

Results and Discussion

Phylogeny

We began our analyses of cyclase evolution by estimating the phylogeny of all proteins known to share the cyclase fold, including triterpenoid, diterpenoid, and meroterpenoid cyclases, and cyclase homologs of unknown function from environmental metagenomes (Fig. 2). To do so, we retrieved cyclase homologs from public databases using BLAST and Pfam-based methods (see Material and Methods). Combined, these searches resulted in 44,209 unique sequences after removing duplicates and filtering for length. Briefly, all sequences in this database were aligned, alignments were subset to 1,139 sequences, and phylogenetic trees were estimated under maximum likelihood.

Fig. 2.

Fig. 2.

Maximum-likelihood phylogeny of terpenoid cyclases estimated under the EX_EHO model with five free rate categories. Stars indicate cyclases from environmental metagenomes; squares indicate cyclases of known function. Ultrafast bootstrap support is shown for deep splits, where black circles indicate bootstrap ≥ 93. Purple clades are composed of proteins without a γ domain (β-only or βx) as annotated; pink clades are composed of two-domain (βγ or αβγ) proteins.

We tested the robustness of tree topology to uncertainty in alignment, model choice, and history of domain duplication and loss. Figure 2 shows the cyclase phylogeny estimated under EX_EHO + R5, which applies separate evolutionary models to each site depending on its probability of being on the protein's surface or buried, and depending on whether it is in an alpha helix, a loop, or another structural element (Le and Gascuel 2010). This model also calculates the probability of each site being in one of five evolutionary rate categories. While the distribution of these categories differed slightly between the β and γ domains, we found that the likelihood-maximizing partitioning scheme for the concatenation of domains was a single partition. The topology estimated under EX_EHO + R5 is robust to omission of γ domains, to alignment masking, to larger and smaller subsets of the sequence database, and to model misspecification, as the same deep splits were recovered under runner-up models (supplementary table S2, Supplementary Material online). However, ultrafast bootstrap support improved dramatically under better-fitting models. Thus, although protein sequences contain limited evolutionary information, and can lose information as they evolve due to site saturation caused by repeated substitution (Xia et al. 2003), the phylogeny presented here represents a strong hypothesis given the sequence information available.

The domain structure of all cyclases in this tree is either β, βγ, βx, or αβγ. We noticed that β and γ domains crystal structures do not align well, despite clear secondary structural and topological homology (supplementary table S1, Supplementary Material online). Similarly, diterpenoid γ domains and triterpenoid γ domains do not align well, due either to divergence or because they arose from independent duplications of the β domain. If so, the diterpenoid and triterpenoid cyclase γ domains would still be homologous but could have arisen from β domains with different sequences, in different organisms, resulting in different trajectories through fitness landscapes and resultingly independent evolutionary constraints, rates, and substitution processes. For this reason, we used structural alignments of protein crystals to create custom hidden Markov models (HMMs) for the β domain, the diterpenoid cyclase γ domain, and the triterpenoid cyclase γ domain. Our triterpenoid γ domain HMM did not pick up the diterpenoid γ domain, and vice versa. We estimated cyclase phylogeny with the two γ domains aligned separately and recovered the same topology as when γ domains were aligned together, or omitted.

Major Clades

The overall tree topology is consistent with previously published triterpenoid and diterpenoid subtrees (Desmond and Gribaldo 2009; Frickey and Kannenberg 2009; Gold et al. 2017; Santana-Molina et al. 2020), and with a recent tree of all terpenoid cyclases (Hoshino and Villanueva 2023). Squalene-hopene cyclase homologs (SHCs) and oxidosqualene cyclase homologs (OSCs) form large monophyletic clades, with squalene-tetrahymanol cyclase homologs (STCs) from microbial eukaryotes and sporulenol synthase homologs (SqhCs) from Bacillota forming small groups sister to SHC, and bacterial OSC groups branching earliest among the OSC subclades.

Sister to this mostly triterpenoid group is another pair of clades, one of which is a small group of epoxy-polyprenyl cyclases. These enzymes take a substrate that is epoxidized like oxidosqualene but has all head–tail joined isoprene units rather than the central tail–tail linkage of squalene and oxidosqualene (Fig. 1). The epoxy-polyprenyl cyclases include Lon15, a meroterpenoid octaprenyl (C40) cyclase from Actinomycete Streptomyces argenteolus (Hayashi et al. 2007) and AtoE, a unique C25 cyclase from Actinomycete Amycolatopsis tolypomycina NRRL B-24205 (Kim et al. 2019). This clade is quite taxonomically restricted, with most sequences coming from Actinomycetes and a handful from Chloroflexota and Clostridia. Surprisingly, those diterpenoid cyclases that utilize an epoxidized substrate—Bra4 and PlaT2, epoxy-diterpenoid cyclases from Actinomycetes Nocardia brasiliensis IFM 0406 (Hayashi et al. 2008) and Streptomyces sp. Tü 6071 (Dürr et al. 2006), respectively—also fall in this clade rather than with other diterpenoid cyclases. It seems that the number of substrate isoprenoid chain carbons is a polyphyletic trait in the cyclase tree, as demonstrated by these diterpenoid cyclases and MstE (which takes a substrate with a dihydroxybenzoate-functionalized C20 isoprenoid chain) and by the triterpenoid cyclases, which are paraphyletic with respect to C35 SqhC. The position of SqhC also demonstrates the relative lability of isoprenoid chain linkage (all head–tail vs. tail–tail) as a trait.

Sister to the epoxy-polyprenyl cyclase group is a clade of unknown function referred to in the literature (e.g. Santana-Molina et al. 2020) as “SHC-like.” This clade contains Gracilicutes sequences from the PVC group (Planctomycetes, Verrucomicrobia, and Chlamydiae), the FCB group (Fibrobacterota, Chlorobiota, and Bacteroidota), Spirochaetota, and Acidobacteriota. The function of these proteins, including those from cultured and genetically tractable organisms such as Gemmata obscuriglobus, has not been demonstrated, and their molecular structure has not been determined. However, we found that for both groups, the γ domains are detectable by the custom “triterpenoid” γ domain HMM we generated but not by our “diterpenoid” γ domain HMM. This could be due to functional divergence of the γ domains in these two halves of the tree (see “evolutionary implications” below).

On the tree's central long branch are three subgroups, each containing a single characterized enzyme. One is composed of monodomain (β-only) meroterpenoid cyclase homologs, including a single characterized enzyme from Cyanobacterium Scytonema sp. PCC 10023, which cyclizes a quinone precursor with a functionalized benzyl headgroup and a C20 isoprenoid chain (Moosmann et al. 2020). Interestingly, this cyclase, called MstE, uses a DxD motif rather than the canonical DxDD motif to protonate an alkenyl substrate. Other sequences in this clade are from diverse bacterial taxa, including other Cyanobacteria as well as Alphaproteobacteria, Betaproteobacteria, Deltaproteobacteria, Gammaproteobacteria, Acidobacteriota, Myxococcota, Planctomycetes, Bacillota, Bacteroidota, Actinomycetota, and Chloroflexi, spanning the Gracilicutes and Terrabacteria divisions, although not recapitulating their branching order. There are also several archaeal sequences, including from cultured representatives such as Methanosarcina acetivorans C2A. The other subgroups are composed of drimenyl pyrophosphate synthases (DMS), recently discovered enzymes that cyclize a C15 substrate. One (“DMS-2”) contains the two-domain DMS characterized from Actinomycete Streptomyces showdoensis (Pan et al. 2022), along with homologs from other bacteria, while the other (“DMS-1”) contains the DMS characterized from Bacteroidete Aquimarina spongiae (Vo et al. 2022), which has only one alpha–alpha toroid domain, fused to a haloacid dehalogenase domain. This clade also contains homologs from Actinomycetota, FCB group bacteria, and fungi.

The final major clade consists of diterpenoid cyclases. All characterized proteins in this group take C20 GGPP as a substrate. The two most basal clades consist of only bacterial proteins, including characterized bacterial cyclases that make ent-copalyl pyrophosphate (Jayasuriya et al. 2007; Morrone et al. 2009; Smanski et al. 2011), a gibberellin or antibiotic precursor, and the cyclase from Kitasatospora griseola that synthesizes the terpentecin precursor (Dairi et al. 2001). This is sister to a pair of clades containing fungal ent-copalyl pyrophosphate cyclases (Quin et al. 2014) and plant ent-copalyl pyrophosphate cyclases (Jia et al. 2022), respectively. The plant clade has many paraphyletic groups of prokaryotic homologs on its stem, including the Mycobacterium tuberculosis cyclase that synthesizes the tuberculosinyl adenosine precursor (Young et al. 2015) and homologs from the Asgardarchaeota that produce the same molecule (McShea et al. 2025).

Functional Analysis of Putative Metagenomic Cyclases

Cyclases from environmental metagenomes are distributed throughout the tree, especially in bacteria-dominated clades (Fig. 3). They fall in the SHC and OSC crown groups, as well as on the stems of these clades, in small groups or as singleton branches. To determine if these metagenomic cyclases are capable of cyclizing squalene or oxidosqualene, we expressed a select set of cyclases in two strains of E. coli, engineered to synthesize squalene or oxidosqualene, respectively, and measured production of polycyclic lipids by gas chromatography–mass spectrometry (GC–MS). We identified 14 triterpenoid cyclase homologs from environmental metagenomes which either branched on the OSC or SHC stem lineage rather than within the crown group or branched within the crown in monophyletic subgroups with no cultured representatives. Many of these cyclases have substitutions at sites of known importance for cyclization of squalene or oxidosqualene.

Fig. 3.

Fig. 3.

Phylogenetic positions of triterpenoid cyclases from environmental metagenomes heterologously expressed in E. coli. Metagenomic cyclases are marked with stars, while previously characterized enzymes from cultured organisms are marked with squares, and both are annotated with product profiles. a) Oxidosqualene cyclases. b) Squalene-hopene cyclases. c) Substrates and products of cyclases from environmental metagenomes.

Many triterpenoid cyclases were active in the heterologous expression system. Most enzymes from early branching OSC subclades produce lanosterol or cycloartenol (Table 1; structures shown in Fig. 3c), consistent with their phylogenetic proximity to cyclases from cultured bacteria known to make those products (Fig. 3a). None of the cyclases from early branching bacterial OSC groups had activity for squalene, which requires the greater active site acidity of an SHC (Xu et al. 2004).

Table 1.

Substrates and products of triterpenoid cyclases from environmental metagenomes heterologously expressed in E. coli.

Cyclase gene ID Environmental metagenome Clade Substrate Major products Strain
Ga0079300_100027797 Deep subsurface shale carbon reservoir, OH, USA OSC Oxidosqualene Isoarborinol HM1
Ga0123336_1000111412 Glacier valley Borup Fiord, NU, Canada OSC Oxidosqualene Lanosterol HM3
Ga0114968_100454421 Lake Montjoie, QC, Canada OSC Oxidosqualene Isoarborinol HM4
Ga0055584_1000802711 Pelagic marine sediment, Helgoland, North Sea OSC Oxidosqualene Eudoraenol HM6
Ga0099828_100017293 Vadose zone soil and rhizosphere, Eel River Critical Zone Observatory, CA, USA OSC Oxidosqualene Cycloartenol HM23
Ga0126314_100002927 Serpentine soil from UC McLaughlin Reserve, CA, USA OSC Oxidosqualene Cycloartenol HM31
Ga0068707_100022523 Anoxygenic and chlorotrophic microbial mat, Yellowstone National Park, WY, USA OSC Oxidosqualene Lanosterol, parkeol HM33
Ga0129299_10203701 Hot spring microbial mat, Owens Valley, CA, USA SHC Squalene Hop-17(21)-ene RAV2
Ga0068707_10120735 Anoxygenic and chlorotrophic microbial mat, Yellowstone National Park, WY, USA SHC Squalene Hop-17(21)-ene RAV3
Ga0066848_100168651 Oceanic oxygen minimum zone, Eastern Pacific Ocean SHC Squalene Hop-17(21)-ene RAV7
JGI11876J14442_100020056 Hot spring microbial mat, Elkhorn Slough, CA, USA SHC Squalene Hop-17(21)-ene RAV10
BBAY77_1256936218 Macroalgal surface, Botany Bay, NSW, Australia SHC Squalene Hop-17(21)-ene, hop-22(29)-ene RAV13
Abe_100035313 Black smoker hydrothermal plume, Abe, Lau Basin, Pacific Ocean SHC Oxidosqualene/squalene Hopanol/no product HM163
JGI24120J20309_100072311 Serpentinite rock and fluid subsurface biosphere, McLaughlin Reserve, CA, USA SHC Oxidosqualene/squalene Hopanol/no product HM164

Products were determined by comparison to published GC–MS spectra (see supplementary figs. S1 and S2, Supplementary Material online), and in the case of isoarborinol, confirmed by NMR (supplementary fig. S3 and table S3, Supplementary Material online).

One cyclase from a marine metagenome produces eudoraenol, a recently discovered arborane triterpenoid. It is the second eudoraenol-synthesizing enzyme to be reported after the discovery of a eudoraenol cyclase in Eudoraea adriatica (Banta et al. 2017), a marine heterotroph from phylum Bacteroidetes (Alain et al. 2008). The new eudoraenol cyclase has the same substitutions shown by Banta et al. (2017) to be sufficient for eudoraenol production—tryptophan to serine at position 230 (Homo sapiens OSC numbering), histidine to tyrosine at position 232, tyrosine to valine at position 521, and asparagine to tyrosine at position 697. In the enzyme active site, these residues are clustered near the substrate tail where the fifth ring forms in arborinols.

Closely related to the eudoraenol cyclases are two cyclases from deep subsurface and inland lake ecosystems which we found synthesize isoarborinol (Fig. 4). These cyclases have the eudoraenol cyclase substitutions at positions 230, 232, and 697 and a unique cysteine at position 503. The subsurface cyclase contig is predicted by the Whokaryote classifier (Pronk and Medema 2022) to be prokaryotic, making it the first reported bacterial source of isoarborinol. While the lake cyclase contig is too short for classification, its phylogenetic proximity to the subsurface and other bacterial cyclases make it likely to be prokaryotic as well. Arborinol is named for the tree from which it was first isolated (Vorbrüggen et al. 1963; Kennard et al. 1965), but the arborane fossil record stretches back to the Permian (Hauke et al. 1995), well before the evolution of angiosperms, leading to the hypothesis (Ourisson et al. 1982) that microbial sources of arborinols, and specifically isoarborinol, must exist. The formation of isoarborinol by these two cyclases identifies a modern and most likely bacterial source for this orphan biomarker. Bacterial arborinol cyclases form a monophyletic clade within one of the bacterial OSC clades. Thus, it is unclear whether bacterial arborinol cylcases arose before or after the radiation of eukaryotic OSCs, which show a pattern of vertical inheritance (i.e. major clades of plant, fungal, and animal cyclases largely recapitulate organismal phylogeny). The bacterial arborinol cyclases are thus a reasonable, but not definitive, source for Paleozoic and Proterozoic arborane fossils.

Fig. 4.

Fig. 4.

Biosynthesis of isoarborinol by a triterpenoid cyclase from a deep subsurface shale metagenome. a) The reaction performed by isoarborinol cyclase. b) GC–MS total ion chromatogram of total lipid extract from E. coli strain HM1 expressing the oxidosqualene biosynthesis pathway and Ga0079300_100027797, derivatized to trimethylsilyl (TMS) ethers. c) Mass spectrum of peak at 36.0 min, identified by comparison to Vorbrüggen et al. (1963) and by nuclear magnetic resonance (NMR), see supplementary fig. S3 and table S3, Supplementary Material online.

Most enzymes from early branching SHC subclades (Fig. 3b) that we tested produce typical hopene isomers hop-17(21)-ene and/or hop-22(29)-ene from a squalene substrate (Table 1). Hopene-producing enzymes also all had activity for oxidosqualene, typically producing a single hopanol isomer (hopan-22(29)-ol), consistent with previous work, which has shown squalene cyclases to be nonspecific (Rohmer et al. 1980; Abe and Rohmer 1994; Seitz et al. 2012; Hammer et al. 2013). However, a small group of cyclases within the SHC crown had activity for only oxidosqualene and no activity for squalene, producing only hopanols. This is surprising given that all other cyclases in the SHC clade have some activity for squalene, including sporulenol cyclases, which are functional for both β-curcumene and squalene, and tetrahymanol cyclases. These oxidosqualene cyclases in the SHC clade do not lack any residues needed to activate squalene. Indeed, they have the electrophilic DxD[D] motif thought to be necessary to protonate an alkenyl substrate. It is possible that these enzymes are in fact true SHCs but cannot cyclize squalene in the E. coli cell under the conditions tested, perhaps due to minor protein misfolding that weakens the electrophilicity of the active site, retaining activity for the more reactive substrate but abolishing it for the more recalcitrant one. It is also possible that these hopanol-producing enzymes are true OSCs in the functional sense, despite being SHCs in the phylogenetic sense, and that substrate functionalization is a relatively labile trait in this group.

Adaptation in the Terpenoid Cyclase Superfamily

What evolutionary dynamics drove the diversification of terpenoid cyclases? Did the diversity of substrates and products in this superfamily arise from neutral drift and divergence, or did natural selection play a role? To assess the relative influences of drift and selection, we employed a quantitative test that interprets positive selection as an elevated rate of amino acid mutations in functional regions of the protein structure, along a given branch in a phylogenetic tree (similar to Ritchie et al. 2021; Maddamsetti and Grant 2022). This approach applies the theory developed for protein-coding nucleotide sequences (e.g. the branch-site test, McDonald and Kreitman 1991; Yang and Dos Reis 2011; Messer and Petrov 2013; Gharib and Robinson-Rechavi 2013) to protein structures, which allows the test to work on deeply divergent groups where nucleotide substitutions have saturated but structure has evolved more slowly.

Along all the major internal branches of the terpenoid cyclase phylogeny, there are two where a functional domain had a significantly elevated proportion of amino acid substitutions over background (Fig. 5). One is the stem of the triterpenoid cyclases which take an alkenyl substrate (that is, squalene–hopene, sporulenol, and tetrahymanol cyclases), where the active site of the protein has an unusually high substitution density. The second is the stem of the group of unknown function, where the SHC-type dimer contact surface has unusually high substitution density. We interpret clustering of substitutions in a functional region as a potential signal of protein adaptation, in which case these may be two instances of positive selective pressure and/or functional divergence in the evolutionary history of triterpenoid cyclase homologs. Adaptation could be to new environmental conditions (e.g. to a new host organism after horizontal gene transfer), to new function, or both. That only two instances of rate shifts in functional regions were detectable means either that adaptation on other branches occurred at a molecular length scale undetectable when focusing on functional regions (e.g. single-residue substitutions), at a temporal scale undetectable when focusing on long branches (e.g. rapid evolution within crown groups), or that triterpenoid cyclase divergence was otherwise dominated by drift.

Fig. 5.

Fig. 5.

Windows of rapid evolution in functionally important domains of triterpenoid cyclases. a) Functional regions shown on the squalene-hopene cyclase structure (PDB ID 1SQC). Regions found to have an evolutionary rate elevated above background are marked with a rhombus (⧫) and a star (★). b) Branches with a rapidly evolving functional region, marked with symbols corresponding to panel (a). As in Fig. 2, purple clades are composed of proteins without a γ domain (β-only or βX); pink clades are composed of two-domain (βγ or βγX) proteins.

Evolutionary Implications

To interrogate domain evolution in the cyclase family, we turned to protein structure, which evolves more slowly than sequence composition (Caetano-Anollés et al. 2009; Ingles-Prieto et al. 2013). We applied structural data to estimate relationships in three ways: we used structures to better align sequences before estimating phylogeny under a model of sequence evolution (Fig. 6a), we measured relationships as simple distances between structural coordinates (Fig. 6b), and we measured relationships using the local structural alphabet implemented in Foldtree (Moi et al. 2023) (Fig. 6c).

Fig. 6.

Fig. 6.

Relationships between cyclase domains estimated using three structure-informed methods. a) Phylogeny estimated using standard sequence-based method after aligning sequences using structure. b) Neighbor-joined dendrogram of domains using root mean squared deviation (RMSD) as distances. c) Structure relationships estimated using the Foldtree algorithm. Horizontal stripes are for visual aid, to show the position of each group across trees.

The structure-aligned sequence tree (Fig. 6a) and the RMSD dendrogram (Fig. 6b) show a deep split between β domains and γ domains. The β and γ clades have the same internal branching order in the structure-aligned sequence tree, consistent with a single duplication event at the root followed by parallel divergence, with the only other domain rearrangements being the loss of the γ domain in monodomain cyclase MstE (PDB ID 6SBB), gain of the α domain in the ancestor of plant cyclases, and loss of the γ domain later within that clade. The RMSD dendrogram also supports this scenario, and differs only in the relationships among β domains, placing the β domains of bacterial diterpenoid cyclases sister to those of the triterpenoid cyclases rather than those of the plant cyclases. The Foldtree estimate (Fig. 6c) predicts a much more complex scenario, with a deep split between plant β domains and the rest of the structures. After the split with plant βs, the first branches are triterpenoid and sesquiterpenoid γ domains, then triterpenoid and sesquiterpenoid βs, then diterpenoid βs and γs. This is consistent with our hypothesis of triterpenoid and diterpenoid β domains arising from independent duplication events but suggests a surprising γ domain ancestor for all nonplant β domains. Given that Foldtree does not yet measure statistical support, and given the agreement between the structure-aligned sequence tree, the RMSD dendrogram, and our large sequence-based phylogeny (Fig. 2), we tentatively reject the Foldtree estimate and the independent duplications hypothesis. Most evidence supports a single evolution of the terpenoid cyclase γ domain. However, this scenario should be revisited as structural phylogenetics (Malik et al. 2020; Moi et al. 2023; Garg and Hochberg 2025) becomes more sophisticated (Mutti et al. 2025).

Given the possible selection on oligomerization and the active site revealed through our structural analyses of cyclases, we can posit and assess evolutionary scenarios focused on related traits. The longest internal branch of the cyclase phylogeny is between the diterpenoid cyclases and all other proteins (Fig. 7a). Rooting the tree on this branch, in agreement with the root position in the β and γ domain subtrees (Fig. 6), places all-bacterial clades basal to the eukaryotic clades. This is consistent with separate transfer events of diterpenoid cyclases from bacteria to plants and to fungi, likely at different times and involving different donor organisms and proteins. For this and most possible superfamily roots, the triterpenoid cyclase root is between the squalene-hopene (and STC and SqhC) and oxidosqualene cyclases. This OSC-SHC root position does not further elucidate the biosynthetic substrate and product of the triterpenoid cyclase ancestor (i.e. whether it produced hopanoids, sterols, or a different molecule). Given that squalene-hopene cyclase was likely present in the common ancestor of Gracilicutes, one of the major divisions of bacteria (Santana-Molina et al. 2020), the OSC-SHC ancestor is likely quite ancient, although hopanes and steranes do not appear in the fossil record until 1.6 billion years ago (Brocks et al. 2005; French et al. 2015).

Fig. 7.

Fig. 7.

Distribution of terpenoid cyclase traits and hypotheses for superfamily root position. a) Cyclase traits mapped across the unrooted tree, which has the same topology and internal branch lengths as the tree shown in Fig. 2. Indicated traits include the number of rings in the final product, the structural conformation of the substrate (CBC: chair–boat–chair or CCC: chair–chair–chair); whether the substrate is epoxidized, and the number of domains (β = purple half-oval, γ = red half-oval, α (present in the plant subclade of diterpenoid cyclases) = green circle, haloacid dehalogenase domain = orange half-oval). a–d) Alternate hypotheses for the root position of the cyclase tree with substrate and domain structure mapped. For simplicity, STC and SqhC are not shown, as their domain architecture and substrate functionalization is identical to SHC. OSC, oxidosqualene cyclases; SHC, squalene-hopene cyclases; MstE, meroterpenoid cyclases; DTC, diterpenoid cyclases; DMS, drimenyl synthases; Unk, homologs of unknown function; oPPC, epoxy-polyprenyl cyclases.

However, rooting on the diterpenoid cyclases implies either three independent evolutions of the γ domain (Fig. 7b), or two losses of the γ domain on the MstE and DMS-1 stems (Fig. 7c). The first scenario is not supported by the analysis of domain evolution discussed above (Fig. 6). The second scenario is consistent with the domain trees but seems biochemically unlikely given that both domains are necessary for function in modern two-domain enzymes for substrate positioning and water exclusion (Wendt 2005; Abe 2007; Fischer and Pearson 2007). The unusual cyclase DMS-1 may have had an easier path, with the novel haloacid dehalogenase-like domain (Vo et al. 2022) replacing the γ domain, or inserted between β and γ, in one mutational event, never requiring an evolutionary intermediate with a water-exposed active site. On the other hand, the monodomain cyclase merosterolic acid synthase (MstE) excludes water with a mobile loop that caps the active site when substrate is bound (Moosmann et al. 2020). This loop is significantly shorter in the closest related two-domain cyclase, drimenyl diphosphate synthase (Pan et al. 2022), suggesting that unless loop extension and domain loss occurred simultaneously, the active site of the MstE ancestor would have been exposed to water, making catalysis impossible. However, domain loss is not unknown in terpenoid cyclases (Oldfield and Lin 2012). Biochemical investigation of basal or ancestral monodomain and didomain cyclases, as well as of, e.g. synthetic constructs of monodomain cyclases with shortened mobile loops, are necessary to establish the relative probabilities of domain loss and duplication in this superfamily.

If we then take domain loss in an MstE ancestor to be less likely than domain loss in a DMS-1 ancestor, a monodomain cyclase root (Fig. 7d) becomes worth investigating. Rooting the structure-informed domain tree and the RMSD domain dendrogram on the monodomain cyclase (PDB ID 6SBB) places sesquiterpenoid and triterpenoid β domains as the next most basal, followed by remaining β domains, from which a monophyletic γ domain clade emerges. Although this position is in all cases far from the tree midpoint, it remains plausible given the evolutionary rarity of domain loss (Bridgham et al. 2009), and balances that rarity with the evidence for a single origin of the γ domain.

Within the triterpenoid cyclases, the oxidosqualene cyclase clade contains enzymes that form different numbers of rings (four for sterol cyclases and five for arborinol cyclases). The squalene-hopene cyclases, which form five rings, have a stem group comprised of enzymes that form four (sporulenol cyclases) or five (squalene-tetrahymanol cyclase) rings, respectively. Although we have found an outgroup with fewer rings (three for Bra4 and PlaT2) as predicted by Fischer and Pearson (2007), this group also contains an enzyme that forms many more (eight for Lon15). A gradualist model of evolution, whereby additional rings evolve irreversibly and one-at-a-time, does not fit the data for this trait; rather, the number of rings is relatively labile. This is not surprising since anti-Markovnikov (secondary) carbocations are no longer considered to be cyclization intermediates (Hess 2002; Tantillo 2010; Hess and Smentek 2013), removing the physical argument for a teleology of ring number.

Finally, substrate functionalization could have followed two evolutionary paths in the updated phylogeny. Fischer and Pearson (2007) hypothesized that the use of an epoxidized substrate evolved from an alkenyl ancestor once, on the oxidosqualene cyclase stem after the divergence of OSC and SHC. In the updated phylogeny, because epoxy-polyprenyl cyclases also used an epoxidized substrate, epoxide use could have evolved twice independently, once in the ancestor of each group. Because alkenyl-substrate enzymes can also cyclize epoxidized substrates, but epoxy-substrate enzymes cannot cyclize alkenyl substrates, evolutionary transitions from alkenyl to epoxy substrates seem in general more likely than the reverse. This is consistent with our discovery of enzymes within the squalene–hopene cyclase crown that only work on an epoxidized substrate (Table 1). Alternately, use of an epoxidized substrate may have evolved only once, in the common ancestor of oxidosqualene cyclases and epoxy-polyprenyl cyclases, and subsequently been lost in the squalene–hopene cyclase/squalene-tetrahymanol cyclase/sporulenol cyclase ancestor. Because epoxidized substrate biosynthesis requires molecular oxygen (Jahnke 1986; Jahnke and Nichols 1986), this scenario is less likely if the OSC-SHC-oPPC ancestor predates the radiation of Gracilicutes (Santana-Molina et al. 2020) and if Gracilicutes predate the rise of molecular oxygen around 2.4 billion years ago (Luo et al. 2016; Gumsley et al. 2017; Martinez-Gutierrez et al. 2023; Ostrander et al. 2024; Wang and Luo 2025). However, it bears further investigation, e.g. via molecular clock and duplication-transfer-loss analysis, given the uncertainties in both antecedents. The relative probabilities of the two scenarios for substrate evolution will also be better-informed when the substrate of the “unknown” clade is discovered.

Conclusions

Here, we present an updated phylogenetic and functional analysis of triterpenoid cyclases. One result is that the triterpenoid cyclase active site has an elevated evolutionary rate on the SHC-STC-SqhC stem lineage. We hypothesize that this may be due to adaptation to an alkenyl substrate from an epoxy ancestor of all triterpenoid and epoxy-polyprenyl cyclases, if this ancestor evolved after the great oxidation event. Further, we found possible reversions to an epoxidized substrate in the SHC clade. This experimental result, as well as the discovery of a novel bacterial oxidosqualene-isoarborinol cyclase, shows that metagenomic cyclases from uncultured sources are a significant source of diversity and can help inform our understanding of the biochemical potential and evolutionary history of these fascinating enzymes. Future work focused on the functional and structural characterization of cyclases, particularly those from the uncharacterized clade sister to the epoxy-polyprenyl cyclases, will provide invaluable insight that will help constrain the evolutionary scenarios proposed in this study.

Materials and Methods

Gene Synthesis and Molecular Cloning

Cyclase DNA sequences were codon-optimized for expression in E. coli and synthesized through the Department of Energy Joint Genome Institute (DOE JGI) DNA Synthesis Science Program. These genes were obtained in the IPTG-inducible plasmid pSRKGm-lacUV5-rbs5 (Banta et al. 2017) in E. coli TOP10. Plasmid DNA was isolated using the GeneJET Plasmid Miniprep Kit (ThermoFisher Scientific) and sequenced to confirm promoter and gene sequences by ELIM Biopharm (Hayward, CA) using the following primers: 5′-AATGCAGCTGGCACGACAGG-3′ (forward) and 5′-CCAGGGTTTTCCCAGTCAC-3′ (reverse), purchased from Integrated DNA technologies (Coralville, IA).

Bacterial Culture and Heterologous Expression

E. coli DH10B expression strains were transformed by electroporation with the cyclase-bearing pSRKGm plasmid (Khan et al. 2008) and two other plasmids: pJBEI2997 (Addgene plasmid #351515) (Peralta-Yahya et al. 2011) encoding the MEV pathway and a pTrc99a derivative (Amann et al. 1988) encoding squalene synthase (and, in some cases, squalene monooxygenase). Gentamicin (15 μg/mL), carbenicillin (100 μg/mL), and/or chloramphenicol (20 μg/mL) were added to media as appropriate to select for maintenance of each plasmid. Plasmids are described in detail in supplementary table S4, Supplementary Material online.

Strains were cultured in 50 mL TYGPN broth at 37 °C while shaking at 225 rpm. Cultures were induced with 500 μM isopropyl β-D-1-thiogalactopyranoside (IPTG) at an OD600 of ∼0.6, and then incubated an additional 48 h at 30 °C while shaking at 225 rpm. Cells were harvested by centrifugation at 4500 × g for 10 min at 4 °C. Cell pellets were stored at −20 °C until lipid extraction. In every experiment, a negative control strain with an empty pSRKgm plasmid and relevant positive controls (strains with pSRKGm containing a triterpenoid cyclase of known function) were grown and extracted alongside experimental strains.

Lipid Extraction

Cell pellets were extracted using a modified Bligh–Dyer method (Bligh and Dyer 1959; Welander et al. 2012): cells were resuspended in 2 mL of deionized water and transferred to a solvent-washed Teflon centrifuge tube containing 5 mL of methanol and 2.5 mL of dichloromethane (DCM), vortexed for 30 s, and sonicated for 1 h in a water bath sonicator. 10 mL of deionized water and 10 mL of DCM were then added, and samples were vortexed and incubated overnight at −20 °C. Samples were centrifuged for 10 min at 2,800 × g, and the organic layer was transferred to a baked glass vial and evaporated at 40 °C under a gentle stream of N2 to yield the total lipid extract (TLE). TLE was stored at −20 °C, and derivatized to trimethylsilyl ethers with 1:1 (v:v) bis(trimethylsilyl)trifluoroacetamide (BSTFA):pyridine for 1 h at 70 °C before analysis by GC–MS. The alcohol-soluble fractions of some TLEs were further purified by silicon column chromatography before derivatization, using a solvent series of hexanes, 8:2 hexanes:DCM, DCM, 1:1 DCM:ethyl acetate, and ethyl acetate (Summons et al. 2013).

GC–MS Analysis

Lipids were separated with an Agilent 7890B Series GC equipped with two Agilent DB-17HT columns (30 m × 0.25 mm i.d. × 0.15 μm film thickness) in tandem with helium as the carrier gas at a constant flow of 1.1 mL/min and programmed as follows: 100 °C for 2 min, then 12 °C/min to 250 °C and held for 10 min, then 10 °C/min to 330 °C and held for 17.5 min. 2 μL of each sample was injected in splitless mode at 250 °C. The GC was coupled to an Agilent 5977A Series MSD with the ion source at 230 °C and operated at 70 eV in electron ionization mode scanning from 50 to 850 Da in 0.5 s. All lipids except isoarborinol were identified based on their retention time and comparison with published spectra. Isoarborinol was identified as a likely sterol by GC–MS, and its structure was confirmed by nuclear magnetic resonance spectroscopy (NMR).

NMR Analysis

NMR experiments were performed following Banta et al. (2017). TLE saponification was accomplished by heating with 10% (vol/vol) sodium hydroxide (NaOH)/MeOH at reflux for 16 h. The reaction mixture was partitioned between water and hexane/ethyl acetate (EtOAc) 2:1; the organic layers were filtered through neutral alumina and concentrated to dryness with a stream of nitrogen. The saponified lipids were fractionated by preparative TLC on glass-backed plates (10 cm in length) coated with a 0.25-mm layer of silica gel 60 F254 using hexane/EtOAc 4:1 as the developing solvent. The triterpenol fraction was further fractionated by reversed-phase HPLC with a system consisting of a Waters 6000A pump, Waters 410 differential refractometer, and two Altex Ultrasphere ODS 5-μm 10 × 250 mm columns in series using a flow rate of 3 mL/min MeOH. After evaporation of the HPLC solvent, the triterpenols were characterized by NMR using a Bruker Avance III HD with an Ascend 800 MHz magnet and a 5-mm TCI cryoprobe at 30 °C using deuterated chloroform (CDCl3) as the solvent. Calibration was by the residual solvent signal (1H: 7.26 ppm; 13C 77.0 ppm).

Sequence Search and Curation

Cyclase homologs were retrieved from the Joint Genome Institute Integrated Microbial Genomes and Microbiomes database (JGI IMG; https://img.jgi.doe.gov) and the National Center for Biological Informatics nonredundant sequence database (https://www.ncbi.nlm.nih.gov/genbank) using the set of query sequences listed in supplementary table S5, Supplementary Material online. Protein–protein BLAST algorithms were used in both cases, with an expect threshold of 0.05 and a word size of 5 for NCBI, and an e-value cutoff of 1e−5 for JGI. Additionally, we downloaded all hits for Pfams PF13243 (squalene-hopene cyclase C-terminal domain) and PF13249 (squalene-hopene cyclase N-terminal domain) from the European Bioinformatics Institute (EBI; https://www.ebi.ac.uk/interpro/entry/pfam/#table). Combined, these searches resulted in an initial set of 44,209 unique sequences. BLAST results were filtered by length dependent on query sequence as follows: 600 to 800 amino acids (aa) for triterpenoid cyclases, 550 to 700 aa for epoxy-polyprenyl cyclases and members of the “unknown function” clade, 700 to 1,000 aa for plant and fungal diterpenoid cyclases, 450 to 600 aa for bacterial diterpenoid cyclases, 550 to 650 aa for DMS1, 450 to 550 aa for DMS2, and 300 to 400 aa for MstE. These length ranges were determined based on length distribution of results and length range of functionally characterized representatives of these groups. Sequences from EBI were filtered to a length range of 350 (the length of the smallest known Type II cyclase; Moosmann et al. 2020) to 900 (the length of the largest known Type II domain-bearing cyclase; Mitsuhashi et al. 2017) amino acids. Remaining sequences were subsetted with CD-hit (command cd-hit -i in.fasta -o out.fasta -c 0.90 -n 5 -M 6000 -d 0 -T 8) (Fu et al. 2012) and aligned with the MAFFT v7.402 fftns2 algorithm (Rozewicki et al. 2019), after which sequences lacking an aligned DxD motif were removed (except for oPPC homologs with an [E/D]SA[E/N] motif; Dürr et al. 2006; Hayashi et al. 2008) to yield a database of 13,680 sequences.

Phylogenetic Estimation

All sequences in the database described above were aligned using the MAFFT linsi algorithm (Rozewicki et al. 2019) implemented in Magus (Smirnov and Warnow 2021). The alignment was trimmed using trimAl with a gap threshold of 0.1 (command trimal -in in.aln -out out.aln -gt 0.1), and subsetted with a maximum sequence identity of 50% (command trimal -in in.aln -out out.aln -maxidentity 0.5) (Capella-Gutiérrez et al. 2009). Model selection and maximum likelihood phylogeny estimation were performed with IQ-TREE 2.2.0 (Bui et al. 2020), with the model EX_EHO + R5 (Le and Gascuel 2010) found to be the best model by Bayesian Information Criterion and Akaike Information Criterion (Kalyaanamoorthy et al. 2017). For all phylogeny estimation, we performed ≤3,000 ultrafast bootstrap replicates (Hoang et al. 2018). All input commands are reported in supplementary table S2, Supplementary Material online. Resulting phylogenetic trees were visualized in Dendroscope (Huson and Scornavacca 2012), Figtree (Rambaut 2009), or iTOL (Letunic and Bork 2021).

Analysis of Adaptation

Using the phylogeny estimated under the model EX_EHO + R5 in IQ-TREE, ancestral protein sequences were reconstructed using IQ-TREE. For each of the branches between major clades (highlighted in Fig. 5), most-likely ancestral sequences at the branches’ terminal nodes were modeled using ColabFold (Mirdita et al. 2022). We then used Fisher's exact test (2-tail, multiple-test corrected) to measure whether the amino acid substitutions that occurred along that branch (differences between the two ancestral proteins) are statistically clustered in functional regions of the protein. Six functional regions were tested: the protein's N terminus, the C terminus, the helix that acts as a membrane anchor, the dimer contact surface, the tunnel from the surface to the active site, and the active site itself (precise region definitions can be found in supplementary table S7, Supplementary Material online). Substrate tunnel residues were calculated with Caver (Pavelka et al. 2016) and active site was defined as residues within 5 Å of substrate molecule.

Supplementary Material

msaf203_Supplementary_Data

Acknowledgments

The authors would like to thank the students of ESS 143: Molecular Geomicrobiology Lab for hard work and curiosity in initial experiments characterizing the lipid products of several metagenomic triterpenoid cyclases. The authors would like to thank members of the Welander lab, M. Horst, B. Kapili, C.D. Stark, and B. Barros-McShea for critically examining, and thereby improving, this work. We would also like to thank Stanford Research Computing for access to and assistance with the Sherlock High-Performance Computing Environment, where we performed phylogeny estimation and protein structure prediction, and the Stanford Geomicrobiology Shared Laboratories Core Facility (RRID:SCR_025000), where we performed portions of the experimental work.

Contributor Information

Hanon Solomon McShea, Department of Earth System Science, Stanford University, Stanford, CA, USA; Department of Microbiology and Immunology, University of California, San Francisco, San Francisco, CA, USA.

Robb A Viens, Department of Earth System Science, Stanford University, Stanford, CA, USA.

Babatunde O Olagunju, Department of Chemistry, SUNY-ESF, Syracuse, NY, USA.

José-Luis Giner, Department of Chemistry, SUNY-ESF, Syracuse, NY, USA.

Paula V Welander, Department of Earth System Science, Stanford University, Stanford, CA, USA.

Supplementary Material

Supplementary material is available at Molecular Biology and Evolution online.

Funding

H.S.M. and P.V.W. were supported by National Science Foundation (NSF) Grant EAR-1752564. H.S.M. was supported by the NSF Graduate Research Fellowship, the Stanford Enhancing Diversity in Graduate Education (EDGE) Fellowship, and the Stanford School of Sustainability McGee and Levorsen Graduate Research Grant. R.A.V. was supported by the Stanford Earth Summer Undergraduate Research (SESUR) Program. B.O.O. and J.L.G. were supported (in part) by National Institutes of Health (NIH) Grant R15GM143714; the 800 MHz NMR was acquired through NIH S10 OD012254.

Data Availability

The data supporting the findings in this study are available in the main text and supplementary information. The terpenoid cyclase alignment and phylogeny (shown in Figs. 2, 3, and 5) are available as supplementary files.

References

  1. Abe  I. Enzymatic synthesis of cyclic triterpenes. Nat Prod Rep.  2007:24(6):1311–1331. 10.1039/b616857b. [DOI] [PubMed] [Google Scholar]
  2. Abe  I. The oxidosqualene cyclases: one substrate, diverse products. In: Osbourn  A, Goss  R, Carter  G, editors. Natural products: discourse, diversity, and design. John Wiley & Sons, Inc.; 2014. p. 293–316. [Google Scholar]
  3. Abe  I, Rohmer  M. Enzymic cyclization of 2,3-dihydrosqualene and squalene 2,3-epoxide by squalene cyclases: from pentacyclic to tetracyclic triterpenes. J Chem Soc Perkin Trans. 1994:1(7):783–791. 10.1039/P19940000783. [DOI] [Google Scholar]
  4. Alain  K, Intertaglia  L, Catala  P, Lebaron  P. Eudoraea adriatica gen. nov., sp. nov., a novel marine bacterium of the family Flavobacteriaceae. Int J Syst Evol Microbiol.  2008:58(10):2275–2281. 10.1099/ijs.0.65446-0. [DOI] [PubMed] [Google Scholar]
  5. Amann  E, Ochs  B, Abel  K-J. Tightly regulated tac promoter vectors useful for the expression of unfused and fused proteins in Escherichia coli. Gene. 1988:69(2):301–315. 10.1016/0378-1119(88)90440-4. [DOI] [PubMed] [Google Scholar]
  6. Banta  AB, Wei  JH, Gill  CCC, Giner  J-L, Welander  PV. Synthesis of arborane triterpenols by a bacterial oxidosqualene cyclase. Proc Natl Acad Sci U S A.  2017:114(2):245–250. 10.1073/pnas.1617231114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Banta  AB, Wei  JH, Welander  PV. A distinct pathway for tetrahymanol synthesis in bacteria. Proc Natl Acad Sci U S A.  2015:112(44):13478–13483. 10.1073/pnas.1511482112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Belin  BJ, Busset  N, Giraud  E, Molinaro  A, Silipo  A, Newman  DK. Hopanoid lipids: from membranes to plant-bacteria interactions. Nat Rev Microbiol.  2018:16(5):304–315. 10.1038/nrmicro.2017.173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Bi  Y, Yu  Z. Diterpenoids from Streptomyces sp. SN194 and their antifungal activity against Botrytis cinerea. J Agric Food Chem.  2016:64(45):8525–8529. 10.1021/acs.jafc.6b03645. [DOI] [PubMed] [Google Scholar]
  10. Bligh  EG, Dyer  WJ. A rapid method of total lipid extraction and purification. Can J Biochem Physiol.  1959:37(1):911–917. 10.1139/o59-099. [DOI] [PubMed] [Google Scholar]
  11. Bloch  KE. Sterol structure and membrane function. Crit Rev Biochem Mol Biol.  1983:14(1):47–92. 10.3109/10409238309102790. [DOI] [Google Scholar]
  12. Bradley  AS, Pearson  A, Sáenz  JP, Marx  CJ. Adenosylhopane: the first intermediate in hopanoid side chain biosynthesis. Org Geochem.  2010:41(10):1075–1081. 10.1016/j.orggeochem.2010.07.003. [DOI] [Google Scholar]
  13. Brenac  L, Baidoo  EEK, Keasling  JD, Budin  I. Distinct functional roles for hopanoid composition in the chemical tolerance of Zymomonas mobilis. Mol Microbiol.  2019:112(5):1564–1575. 10.1111/mmi.14380. [DOI] [PubMed] [Google Scholar]
  14. Bridgham  JT, Ortlund  EA, Thornton  JW. An epistatic ratchet constrains the direction of glucocorticoid receptor evolution. Nature. 2009:461(7263):515–519. 10.1038/nature08249. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Brocks  JJ, Love  GD, Summons  RE, Knoll  AH, Logan  GA, Bowden  SA. Biomarker evidence for green and purple sulphur bacteria in a stratified palaeoproterozoic sea. Nature. 2005:437(7060):866–870. 10.1038/nature04068. [DOI] [PubMed] [Google Scholar]
  16. Brocks  JJ, Pearson  A. Building the biomarker tree of life. Rev Mineral Geochem.  2005:59(1):233–258. 10.2138/rmg.2005.59.10. [DOI] [Google Scholar]
  17. Brown  MO, Olagunju  BO, Giner  J-L, Welander  PV. Sterol methyltransferases in uncultured bacteria complicate eukaryotic biomarker interpretations. Nat Commun.  2023:14(1):Article 1. 10.1038/s41467-023-37552-3. [DOI] [Google Scholar]
  18. Bui  MQ, Schmidt  HA, Chernomor  O, Schrempf  D, Woodhams  MD, von Haeseler  A, Lanfear  R. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol.  2020:37(5):1530–1534. 10.1093/molbev/msaa015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Caetano-Anollés  G, Wang  M, Caetano-Anollés  D, Mittenthal  JE. The origin, evolution and structure of the protein world. Biochem J.  2009:417(3):621–637. 10.1042/BJ20082063. [DOI] [PubMed] [Google Scholar]
  20. Cao  R, Zhang  Y, Mann  FM, Huang  C, Mukkamala  D, Hudock  MP, Mead  ME, Prisic  S, Wang  K, Lin  F-Y, et al.  Diterpene cyclases and the nature of the isoprene fold. Proteins Struct Funct Bioinformatics. 2010:78(11):2417–2432. 10.1002/prot.22751. [DOI] [Google Scholar]
  21. Capella-Gutiérrez  S, Silla-Martínez  JM, Gabaldón  T. Trimal: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009:25(15):1972–1973. 10.1093/bioinformatics/btp348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Chandonia  J-M, Guan  L, Lin  S, Yu  C, Fox  NK, Brenner  SE. SCOPe: improvements to the structural classification of proteins—extended database to facilitate variant interpretation and machine learning. Nucleic Acids Res.  2022:50(D1):D553–D559. 10.1093/nar/gkab1054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Christianson  DW. Structural and chemical biology of terpenoid cyclases. Chem Rev.  2017:117(17):11570–11648. 10.1021/acs.chemrev.7b00287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Conner  RL, Mellory  FB, Landrey  JR, Ferguson  KA, Kaneshiro  ES, Ray  E. Ergosterol replacement of tetrahymanol in Tetrahymena membranes. Biochem Biophys Res Commun.  1971:44(4):995–1000. 10.1016/0006-291X(71)90810-2. [DOI] [PubMed] [Google Scholar]
  25. Crockett  EL. Cholesterol function in plasma membranes from ectotherms: membrane-specific roles in adaptation to temperature. Am Zool.  1998:38(2):291–304. 10.1093/icb/38.2.291. [DOI] [Google Scholar]
  26. Dairi  T, Hamano  Y, Kuzuyama  T, Itoh  N, Furihata  K, Seto  H. Eubacterial diterpene cyclase genes essential for production of the isoprenoid antibiotic terpentecin. J Bacteriol.  2001:183(20):6085–6094. 10.1128/JB.183.20.6085-6094.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Desmond  E, Gribaldo  S. Phylogenomics of sterol synthesis: insights into the origin, evolution, and diversity of a key eukaryotic feature. Genome Biol Evol.  2009:1(0):364–381. 10.1093/gbe/evp036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Dong  L-B, Rudolf  JD, Deng  M-R, Yan  X, Shen  B. Discovery of the tiancilactone antibiotics by genome mining of atypical bacterial type II diterpene synthases. ChemBioChem. 2018a:19(16):1727–1733. 10.1002/cbic.201800285. [DOI] [Google Scholar]
  29. Dürr  C, Schnell  H-J, Luzhetskyy  A, Murillo  R, Weber  M, Welzel  K, Vente  A, Bechthold  A. Biosynthesis of the terpene phenalinolactone in Streptomyces sp. Tü6071: analysis of the gene cluster and generation of derivatives. Chem Biol.  2006:13(4):365–377. 10.1016/j.chembiol.2006.01.011. [DOI] [PubMed] [Google Scholar]
  30. Eschenmoser  A, Arigoni  D. Revisited after 50 years: the ‘stereochemical interpretation of the biogenetic isoprene rule for the triterpenes’. Helv Chim Acta.  2005:88(12):3011–3050. 10.1002/hlca.200590245. [DOI] [Google Scholar]
  31. Fischer  WW. Life before the rise of oxygen. Nature. 2008:455(7216):1051–1052. 10.1038/4551051a. [DOI] [PubMed] [Google Scholar]
  32. Fischer  WW, Pearson  A. Hypotheses for the origin and early evolution of triterpenoid cyclases. Geobiology. 2007:5(1):19–34. 10.1111/j.1472-4669.2007.00096.x. [DOI] [PubMed] [Google Scholar]
  33. Flesch  G, Rohmer  M. Growth inhibition of hopanoid synthesizing bacteria by squalene cyclase inhibitors. Arch Microbiol.  1987:147(1):100–104. 10.1007/BF00492912. [DOI] [Google Scholar]
  34. Förster  HJ, Biemann  K, Haigh  WG, Tattrie  NH, Colvin  JR. The structure of novel C35 pentacyclic terpenes from Acetobacter xylinum. Biochem J.  1973:135(1):133–143. 10.1042/bj1350133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Fox  NK, Brenner  SE, Chandonia  J-M. SCOPe: structural classification of proteins—extended, integrating SCOP and ASTRAL data and classification of new structures. Nucleic Acids Res.  2014:42(D1):D304–D309. 10.1093/nar/gkt1240. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. French  KL, Hallmann  C, Hope  JM, Schoon  PL, Zumberge  JA, Hoshino  Y, Peters  CA, George  SC, Love  GD, Brocks  JJ, et al.  Reappraisal of hydrocarbon biomarkers in Archean rocks. Proc Natl Acad Sci U S A.  2015:112(19):5915–5920. 10.1073/pnas.1419563112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Frickey  T, Kannenberg  E. Phylogenetic analysis of the triterpene cyclase protein family in prokaryotes and eukaryotes suggests bidirectional lateral gene transfer. Environ Microbiol.  2009:11(5):1224–1241. 10.1111/j.1462-2920.2008.01851.x. [DOI] [PubMed] [Google Scholar]
  38. Fu  L, Niu  B, Zhu  Z, Wu  S, Li  W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012:28(23):3150–3152. 10.1093/bioinformatics/bts565. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Garg  SG, Hochberg  GKA. A general substitution matrix for structural phylogenetics. Mol Biol Evol.  2025:42(6):msaf124. 10.1093/molbev/msaf124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Gebhardt  K, Meyer  SW, Schinko  J, Bringmann  G, Zeeck  A, Fiedler  H-P. Phenalinolactones A–D, terpenoglycoside antibiotics from Streptomyces sp. Tü 6071. J Antibiot (Tokyo).  2011:64(3):229–232. 10.1038/ja.2010.165. [DOI] [PubMed] [Google Scholar]
  41. Gharib  WH, Robinson-Rechavi  M. The branch-site test of positive selection is surprisingly robust but lacks power under synonymous substitution saturation and variation in GC. Mol Biol Evol.  2013:30(7):1675–1686. 10.1093/molbev/mst062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Gold  DA, Caron  A, Fournier  GP, Summons  RE. Paleoproterozoic sterol biosynthesis and the rise of oxygen. Nature. 2017:543(7645):420–423. 10.1038/nature21412. [DOI] [PubMed] [Google Scholar]
  43. Gudde  LR, Hulce  M, Largen  AH, Franke  JD. Sterol synthesis is essential for viability in the planctomycete bacterium Gemmata obscuriglobus. FEMS Microbiol Lett. 2019:366(3):fnz019. 10.1093/femsle/fnz019. [DOI] [PubMed] [Google Scholar]
  44. Gumsley  AP, Chamberlain  KR, Bleeker  W, Söderlund  U, de Kock  MO, Larsson  ER, Bekker  A. Timing and tempo of the great oxidation event. Proc Natl Acad Sci U S A.  2017:114(8):1811–1816. 10.1073/pnas.1608824114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Hammer  SC, Syrén  P-O, Seitz  M, Nestl  BM, Hauer  B. Squalene hopene cyclases: highly promiscuous and evolvable catalysts for stereoselective CC and CX bond formation. Curr Opin Chem Biol.  2013:17(2):293–300. 10.1016/j.cbpa.2013.01.016. [DOI] [PubMed] [Google Scholar]
  46. Hauke  V, Adam  P, Trendel  JM, Albrecht  P, Schwark  L, Vliex  M, Hagemann  H, Püttmann  W. Isoarborinol through geological times: evidence for its presence in the permian and triassic. Org Geochem.  1995:23(1):91–93. 10.1016/0146-6380(95)00002-V. [DOI] [Google Scholar]
  47. Hayashi  Y, Matsuura  N, Toshima  H, Itoh  N, Ishikawa  J, Mikami  Y, Dairi  T. Cloning of the gene cluster responsible for the biosynthesis of brasilicardin A, a unique diterpenoid. J Antibiot (Tokyo).  2008:61(3):164–174. 10.1038/ja.2008.126. [DOI] [PubMed] [Google Scholar]
  48. Hayashi  Y, Onaka  H, Itoh  N, Seto  H, Dairi  T. Cloning of the gene cluster responsible for biosynthesis of KS-505a (longestin), a unique tetraterpenoid. Biosci Biotechnol Biochem.  2007:71(12):3072–3081. 10.1271/bbb.70477. [DOI] [PubMed] [Google Scholar]
  49. Hess  BA. Concomitant C-ring expansion and D-ring formation in lanosterol biosynthesis from squalene without violation of Markovnikov's rule. J Am Chem Soc.  2002:124(35):10286–10287. 10.1021/ja026850r. [DOI] [PubMed] [Google Scholar]
  50. Hess  BA, Smentek  L. The concerted nature of the cyclization of squalene oxide to the protosterol cation. Angew Chem Int Ed. 2013:52(42):11029–11033. 10.1002/anie.201302886. [DOI] [Google Scholar]
  51. Hoang  DT, Chernomor  O, von Haeseler  A, Minh  BQ, Vinh  LS. UFBoot2: improving the ultrafast bootstrap approximation. Mol Biol Evol.  2018:35(2):518–522. 10.1093/molbev/msx281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Hoshino  T, Sato  T. Squalene—hopene cyclase: catalytic mechanism and substrate recognition. Chem Commun. 2002:(4):291–301. 10.1039/B108995C. [DOI] [Google Scholar]
  53. Hoshino  Y, Villanueva  L. Four billion years of microbial terpenome evolution. FEMS Microbiol Rev.  2023:47(2):fuad008. 10.1093/femsre/fuad008. [DOI] [PubMed] [Google Scholar]
  54. Huson  DH, Scornavacca  C. Dendroscope 3: an interactive tool for rooted phylogenetic trees and networks. Syst Biol.  2012:61(6):1061–1067. 10.1093/sysbio/sys062. [DOI] [PubMed] [Google Scholar]
  55. Ikeda  C, Hayashi  Y, Itoh  N, Seto  H, Dairi  T. Functional analysis of eubacterial ent -copalyl diphosphate synthase and pimara-9(11),15-diene synthase with unique primary sequences. J Biochem.  2007:141(1):37–45. 10.1093/jb/mvm004. [DOI] [PubMed] [Google Scholar]
  56. Ingles-Prieto  A, Ibarra-Molero  B, Delgado-Delgado  A, Perez-Jimenez  R, Fernandez  JM, Gaucher  EA, Sanchez-Ruiz  JM, Gavira  JA. Conservation of protein structure over four billion years. Structure. 2013:21(9):1690–1697. 10.1016/j.str.2013.06.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Jahnke  LL. The effects of low oxygen on the synthesis of unsaturated fatty acids and sterols: implications for the evolution of eukaryotes. Orig Life Evol Biosph.  1986:16(3-4):317–318. 10.1007/BF02422046. [DOI] [Google Scholar]
  58. Jahnke  LL, Nichols  PD. Methyl sterol and cyclopropane fatty acid composition of Methylococcus capsulatus grown at low oxygen tensions. J Bacteriol.  1986:167(1):238–242. 10.1128/jb.167.1.238-242.1986. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Jayasuriya  H, Herath  KB, Zhang  C, Zink  DL, Basilio  A, Genilloud  O, Diez  MT, Vicente  F, Gonzalez  I, Salazar  O, et al.  Isolation and structure of platencin: a FabH and FabF dual inhibitor with potent broad-Spectrum antibiotic activity. Angew Chem Int Ed. 2007:46(25):4684–4688. 10.1002/anie.200701058. [DOI] [Google Scholar]
  60. Jia  Q, Brown  R, Köllner  TG, Fu  J, Chen  X, Wong  GK-S, Gershenzon  J, Peters  RJ, Chen  F. Origin and early evolution of the plant terpene synthase family. Proc Natl Acad Sci U S A.  2022:119(15):e2100361119. 10.1073/pnas.2100361119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Kalyaanamoorthy  S, Minh  BQ, Wong  TKF, von Haeseler  A, Jermiin  LS. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods.  2017:14(6):587–589. 10.1038/nmeth.4285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Kennard  O, Riva di Sanseverino  L, Vorbrüggen  H, Djerassi  C. The complete structure of the triterpene arborinol. Tetrahedron Lett.  1965:6(39):3433–3438. 10.1016/S0040-4039(01)89324-2. [DOI] [Google Scholar]
  63. Khan  SR, Gaines  J, Roop  RM, Farrand  SK. Broad-host-range expression vectors with tightly regulated promoters and their use to examine the influence of TraR and TraM expression on ti plasmid quorum sensing. Appl Environ Microbiol.  2008:74(16):5053–5062. 10.1128/AEM.01098-08. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Kim  S-H, Lu  W, Ahmadi  MK, Montiel  D, Ternei  MA, Brady  SF. Atolypenes, tricyclic bacterial sesterterpenes discovered using a multiplexed in vitro cas9-TAR gene cluster refactoring approach. ACS Synth Biol.  2019:8(1):109–118. 10.1021/acssynbio.8b00361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Kontnik  R, Bosak  T, Butcher  RA, Brocks  JJ, Losick  R, Clardy  J, Pearson  A. Sporulenes, heptaprenyl metabolites from Bacillus subtilis spores. Org Lett.  2008:10(16):3551–3554. 10.1021/ol801314k. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Le  SQ, Gascuel  O. Accounting for solvent accessibility and secondary structure in protein phylogenetics is clearly beneficial. Syst Biol.  2010:59(3):277–287. 10.1093/sysbio/syq002. [DOI] [PubMed] [Google Scholar]
  67. Lee  AK, Banta  AB, Wei  JH, Kiemle  DJ, Feng  J, Giner  J-L, Welander  PV. C-4 sterol demethylation enzymes distinguish bacterial and eukaryotic sterol synthesis. Proc Natl Acad Sci U S A.  2018:115(23):5884–5889. 10.1073/pnas.1802930115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Lee  AK, Wei  JH, Welander  PV. De novo cholesterol biosynthesis in bacteria. Nat Commun.  2023:14(1):Article 1. 10.1038/s41467-023-38638-8. [DOI] [Google Scholar]
  69. Letunic  I, Bork  P. Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res.  2021:49(W1):W293–W296. 10.1093/nar/gkab301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Levental  I, Levental  KR, Heberle  FA. Lipid rafts: controversies resolved, mysteries remain. Trends Cell Biol.  2020:30(5):341–353. 10.1016/j.tcb.2020.01.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. López  D, Kolter  R. Functional microdomains in bacterial membranes. Genes Dev.  2010:24(17):1893–1902. 10.1101/gad.1945010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Luo  G, Ono  S, Beukes  NJ, Wang  DT, Xie  S, Summons  RE. Rapid oxygenation of Earth's atmosphere 2.33 billion years ago. Sci Adv.  2016:2(5):e1600134. 10.1126/sciadv.1600134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Maddamsetti  R, Grant  NA. Discovery of positive and purifying selection in metagenomic time series of hypermutator microbial populations. PLoS Genet.  2022:18(8):e1010324. 10.1371/journal.pgen.1010324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Malik  AJ, Poole  AM, Allison  JR. Structural phylogenetics with confidence. Mol Biol Evol.  2020:37(9):2711–2726. 10.1093/molbev/msaa100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Mann  FM, Xu  M, Chen  X, Fulton  DB, Russell  DG, Peters  RJ. Edaxadiene: a new bioactive diterpene from Mycobacterium tuberculosis. J Am Chem Soc.  2009:131(48):17526–17527. 10.1021/ja9019287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Martinez-Gutierrez  CA, Uyeda  JC, Aylward  FO. A timeline of bacterial and archaeal diversification in the ocean. eLife. 2023:12:RP88268. 10.7554/eLife.88268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Mayer  MH, Parenteau  MN, Kempher  ML, Madigan  MT, Jahnke  LL, Welander  PV. Anaerobic 3-methylhopanoid production by an acidophilic photosynthetic purple bacterium. Arch Microbiol.  2021:203(10):6041–6052. 10.1007/s00203-021-02561-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. McDonald  JH, Kreitman  M. Adaptive protein evolution at the adh locus in Drosophila. Nature. 1991:351(6328):652–654. 10.1038/351652a0. [DOI] [PubMed] [Google Scholar]
  79. McShea  H, Anda  VD, Brocks  JJ, Welander  PV, Baker  BJ. Archaeal lineages related to eukaryotes encode functional diterpenoid cyclases (p. 2025.02.07.637177). bioRxiv 2025.02.07.637177. 10.1101/2025.02.07.637177, 8 February 2025, preprint: not peer reviewed. [DOI]
  80. Messer  PW, Petrov  DA. Frequent adaptation and the McDonald–Kreitman test. PNAS. 2013:110(21):8615–8620. 10.1073/pnas.1220835110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Mirdita  M, Schütze  K, Moriwaki  Y, Heo  L, Ovchinnikov  S, Steinegger  M. ColabFold: making protein folding accessible to all. Nat Methods.  2022:19(6):679–682. 10.1038/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Mitsuhashi  T, Okada  M, Abe  I. Identification of chimeric αβγ diterpene synthases possessing both type II terpene cyclase and prenyltransferase activities. ChemBioChem. 2017:18(21):2104–2109. 10.1002/cbic.201700445. [DOI] [PubMed] [Google Scholar]
  83. Moi  D, Bernard  C, Steinegger  M, Nevers  Y, Langleib  M, Dessimoz  C. Structural phylogenetics unravels the evolutionary diversification of communication systems in gram-positive bacteria and their viruses. bioRxiv 2023.09.19.558401. 10.1101/2023.09.19.558401., 23 September 2023, preprint: not peer reviewed. [DOI]
  84. Moosmann  P, Ecker  F, Leopold-Messer  S, Cahn  JKB, Dieterich  CL, Groll  M, Piel  J. A monodomain class II terpene cyclase assembles complex isoprenoid scaffolds. Nat Chem.  2020:12(10):968–972. 10.1038/s41557-020-0515-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Morrone  D, Chambers  J, Lowry  L, Kim  G, Anterola  A, Bender  K, Peters  RJ. Gibberellin biosynthesis in bacteria: separate ent-copalyl diphosphate and ent-kaurene synthases in Bradyrhizobium japonicum. FEBS Lett.  2009:583(2):475–480. 10.1016/j.febslet.2008.12.052. [DOI] [PubMed] [Google Scholar]
  86. Mouritsen  OG, Zuckermann  MJ. What's so special about cholesterol?  Lipids. 2004:39(11):1101–1113. 10.1007/s11745-004-1336-x. [DOI] [PubMed] [Google Scholar]
  87. Mutti  G, Ocaña-Pallarès  E, Gabaldón  T. Newly developed structure-based methods do not outperform standard sequence-based methods for large-scale phylogenomics. Mol Biol Evol.  2025:42(7):msaf149. 10.1093/molbev/msaf149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  88. Nakano  C, Okamura  T, Sato  T, Dairi  T, Hoshino  T. Mycobacterium tuberculosis H37Rv3377c encodes the diterpene cyclase for producing the halimane skeleton. Chem Commun. 2005:8(8):1016–1018. 10.1039/B415346D. [DOI] [Google Scholar]
  89. Nakano  C, Oshima  M, Kurashima  N, Hoshino  T. Identification of a new diterpene biosynthetic gene cluster that produces O-methylkolavelool in Herpetosiphon aurantiacus. ChemBioChem. 2015:16(5):772–781. 10.1002/cbic.201402652. [DOI] [PubMed] [Google Scholar]
  90. Nett  RS, Bender  KS, Peters  RJ. Production of the plant hormone gibberellin by rhizobia increases host legume nodule size. ISME J.  2022:16(7):1809–1817. 10.1038/s41396-022-01236-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. Oldfield  E, Lin  F-Y. Terpene biosynthesis: modularity rules. Angew Chem Int Ed. 2012:51(5):1124–1137. 10.1002/anie.201103110. [DOI] [Google Scholar]
  92. Ostrander  CM, Heard  AW, Shu  Y, Bekker  A, Poulton  SW, Olesen  KP, Nielsen  SG. Onset of coupled atmosphere–ocean oxygenation 2.3 billion years ago. Nature. 2024:631(8020):335–339. 10.1038/s41586-024-07551-5. [DOI] [PubMed] [Google Scholar]
  93. Ourisson  G. The evolution of terpenes to sterols. Pure Appl Chem. 1989:61(3):345–348. 10.1351/pac198961030345. [DOI] [Google Scholar]
  94. Ourisson  G, Albrecht  P. Hopanoids. 1. Geohopanoids: the most abundant natural products on earth?  Acc Chem Res.  1992:25(9):398–402. 10.1021/ar00021a003. [DOI] [Google Scholar]
  95. Ourisson  G, Albrecht  P, Rohmer  M. Predictive microbial biochemistry—from molecular fossils to procaryotic membranes. Trends Biochem Sci.  1982:7(7):236–239. 10.1016/0968-0004(82)90028-7. [DOI] [Google Scholar]
  96. Ourisson  G, Nakatani  Y. The terpenoid theory of the origin of cellular life: the evolution of terpenoids to cholesterol. Chem Biol. 1994:1(1):11–23. 10.1016/1074-5521(94)90036-1. [DOI] [PubMed] [Google Scholar]
  97. Ourisson  G, Rohmer  M. Hopanoids. 2. Biohopanoids: a novel class of bacterial lipids. Acc Chem Res.  1992:25(9):403–408. 10.1021/ar00021a004. [DOI] [Google Scholar]
  98. Ozaki  T, Shinde  SS, Gao  L, Okuizumi  R, Liu  C, Ogasawara  Y, Lei  X, Dairi  T, Minami  A, Oikawa  H. Enzymatic formation of a skipped methyl-substituted octaprenyl Side chain of longestin (KS-505a): involvement of homo-IPP as a common extender unit. Angew Chem Int Ed. 2018:57(22):6629–6632. 10.1002/anie.201802116. [DOI] [Google Scholar]
  99. Pan  J-J, Solbiati  JO, Ramamoorthy  G, Hillerich  BS, Seidel  RD, Cronan  JE, Almo  SC, Poulter  CD. Biosynthesis of squalene from farnesyl diphosphate in bacteria: three steps catalyzed by three enzymes. ACS Cent Sci.  2015:1(2):77–82. 10.1021/acscentsci.5b00115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  100. Pan  X, Du  W, Zhang  X, Lin  X, Li  F-R, Yang  Q, Wang  H, Rudolf  JD, Zhang  B, Dong  L-B. Discovery, structure, and mechanism of a class II sesquiterpene cyclase. J Am Chem Soc.  2022:144(48):22067–22074. 10.1021/jacs.2c09412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  101. Pavelka  A, Sebestova  E, Kozlikova  B, Brezovsky  J, Sochor  J, Damborsky  J. CAVER: algorithms for analyzing dynamics of tunnels in macromolecules. IEEE/ACM Trans Comput Biol Bioinform.  2016:13(3):505–517. 10.1109/TCBB.2015.2459680. [DOI] [PubMed] [Google Scholar]
  102. Pearson  A, Budin  M, Brocks  JJ. Phylogenetic and biochemical evidence for sterol synthesis in the bacterium Gemmata obscuriglobus. Proc Natl Acad Sci U S A.  2003:100(26):15352–15357. 10.1073/pnas.2536559100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  103. Pearson  A, Flood Page  SR, Jorgenson  TL, Fischer  WW, Higgins  MB. Novel hopanoid cyclases from the environment. Environ Microbiol.  2007:9(9):2175–2188. 10.1111/j.1462-2920.2007.01331.x. [DOI] [PubMed] [Google Scholar]
  104. Pearson  A, Rusch  DB. Distribution of microbial terpenoid lipid cyclases in the global ocean metagenome. ISME J.  2009:3(3):352–363. 10.1038/ismej.2008.116. [DOI] [PubMed] [Google Scholar]
  105. Peralta-Yahya  PP, Ouellet  M, Chan  R, Mukhopadhyay  A, Keasling  JD, Lee  TS. Identification and microbial production of a terpene-based advanced biofuel. Nat Commun.  2011:2(1):Article 1. 10.1038/ncomms1494. [DOI] [Google Scholar]
  106. Pollier  J, Vancaester  E, Kuzhiumparambil  U, Vickers  CE, Vandepoele  K, Goossens  A, Fabris  M. A widespread alternative squalene epoxidase participates in eukaryote steroid biosynthesis. Nat Microbiol.  2019:4(2):226–233. 10.1038/s41564-018-0305-5. [DOI] [PubMed] [Google Scholar]
  107. Pronk  LJU, Medema  MH. Whokaryote: distinguishing eukaryotic and prokaryotic contigs in metagenomes based on gene structure. Microb Genom.  2022:8(5):000823. 10.1099/mgen.0.000823. [DOI] [Google Scholar]
  108. Quin  MB, Flynn  CM, Schmidt-Dannert  C. Traversing the fungal terpenome. Nat Prod Rep.  2014:31(10):1449–1473. 10.1039/C4NP00075G. [DOI] [PMC free article] [PubMed] [Google Scholar]
  109. Rambaut  A. FigTree, version 1.4.3; 2009. http://tree.bio.ed.ac.uk/software/Figtree/
  110. Ricci  JN, Coleman  ML, Welander  PV, Sessions  AL, Summons  RE, Spear  JR, Newman  DK. Diverse capacity for 2-methylhopanoid production correlates with a specific ecological niche. ISME J.  2014:8(3):675–684. 10.1038/ismej.2013.191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  111. Ritchie  AM, Stark  TL, Liberles  DA. Inferring the number and position of changes in selective regime in a non-equilibrium mutation-selection framework. BMC Ecol Evol.  2021:21(1):39. 10.1186/s12862-021-01770-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  112. Rivas-Marin  E, Stettner  S, Gottshall  EY, Santana-Molina  C, Helling  M, Basile  F, Ward  NL, Devos  DP. Essentiality of sterol synthesis genes in the planctomycete bacterium Gemmata obscuriglobus. Nat Commun.  2019:10(1):2916. 10.1038/s41467-019-10983-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  113. Rohmer  M, Anding  C, Ourisson  G. Non-specific biosynthesis of hopane triterpenes by a cell-free system from Acetobacter pasteurianum. Eur J Biochem.  1980:112(3):541–547. 10.1111/j.1432-1033.1980.tb06117.x. [DOI] [PubMed] [Google Scholar]
  114. Rozewicki  J, Li  S, Amada  KM, Standley  DM, Katoh  K. MAFFT-DASH: integrated protein sequence and structural alignment. Nucleic Acids Res.  2019:47(W1):W5–W10. 10.1093/nar/gkz342. [DOI] [PMC free article] [PubMed] [Google Scholar]
  115. Rudolf  JD, Alsup  TA, Xu  B, Li  Z. Bacterial terpenome. Nat Prod Rep.  2021:38(5):905–980. 10.1039/D0NP00066C. [DOI] [PMC free article] [PubMed] [Google Scholar]
  116. Runnegar  B. Precambrian oxygen levels estimated from the biochemistry and physiology of early eukaryotes. Palaeogeogr Palaeoclimatol Palaeoecol.  1991:97(1-2):97–111. 10.1016/0031-0182(91)90186-U. [DOI] [Google Scholar]
  117. Sáenz  JP. Hopanoid enrichment in a detergent resistant membrane fraction of Crocosphaera watsonii: implications for bacterial lipid raft formation. Org Geochem.  2010:41(8):853–856. 10.1016/j.orggeochem.2010.05.005. [DOI] [Google Scholar]
  118. Sáenz  JP, Grosser  D, Bradley  AS, Lagny  TJ, Lavrynenko  O, Broda  M, Simons  K. Hopanoids as functional analogues of cholesterol in bacterial membranes. Proc Natl Acad Sci U S A.  2015:112(38):11971–11976. 10.1073/pnas.1515607112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  119. Sáenz  JP, Sezgin  E, Schwille  P, Simons  K. Functional convergence of hopanoids and sterols in membrane ordering. Proc Natl Acad Sci U S A.  2012:109(35):14236–14240. 10.1073/pnas.1212141109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  120. Santana-Molina  C, Rivas-Marin  E, Rojas  AM, Devos  DP. Origin and evolution of polycyclic triterpene synthesis. Mol Biol Evol.  2020:37(7):1925–1941. 10.1093/molbev/msaa054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  121. Sato  T, Hoshino  H, Yoshida  S, Nakajima  M, Hoshino  T. Bifunctional triterpene/sesquarterpene cyclase: tetraprenyl-β-curcumene cyclase is also squalene cyclase in Bacillus megaterium. J Am Chem Soc.  2011a:133(44):17540–17543. 10.1021/ja2060319. [DOI] [PubMed] [Google Scholar]
  122. Sato  T, Yoshida  S, Hoshino  H, Tanno  M, Nakajima  M, Hoshino  T. Sesquarterpenes (C35 terpenes) biosynthesized via the cyclization of a linear C35 isoprenoid by a tetraprenyl-β-curcumene synthase and a tetraprenyl-β-curcumene cyclase: identification of a new terpene cyclase. J Am Chem Soc.  2011b:133(25):9734–9737. 10.1021/ja203779h. [DOI] [PubMed] [Google Scholar]
  123. Schaeffer  P, Poinsot  J, Hauke  V, Adam  P, Wehrung  P, Trendel  J-M, Albrecht  P, Dessort  D, Connan  J. Novel optically active hydrocarbons in sediments: evidence for an extensive biological cyclization of higher regular polyprenols. Angew Chem Int Ed. 1994:33(11):1166–1169. 10.1002/anie.199411661. [DOI] [Google Scholar]
  124. Schmerk  CL, Welander  PV, Hamad  MA, Bain  KL, Bernards  MA, Summons  RE, Valvano  MA. Elucidation of the Burkholderia cenocepacia hopanoid biosynthesis pathway uncovers functions for conserved proteins in hopanoid-producing bacteria. Environ Microbiol.  2015:17(3):735–750. 10.1111/1462-2920.12509. [DOI] [PubMed] [Google Scholar]
  125. Seitz  M, Klebensberger  J, Siebenhaller  S, Breuer  M, Siedenburg  G, Jendrossek  D, Hauer  B. Substrate specificity of a novel squalene–hopene cyclase from Zymomonas mobilis. J Mol Catal B Enzym.  2012:84:72–77. 10.1016/j.molcatb.2012.02.007. [DOI] [Google Scholar]
  126. Shigemori  H, Komaki  H, Yazawa  K, Mikami  Y, Nemoto  A, Tanaka  Y, Sasaki  T, In  Y, Ishida  T, Kobayashi  J. Brasilicardin A. A novel tricyclic metabolite with potent immunosuppressive activity from actinomycete Nocardia brasiliensis. J Org Chem.  1998:63(20):6900–6904. 10.1021/jo9807114. [DOI] [PubMed] [Google Scholar]
  127. Smanski  MJ, Yu  Z, Casper  J, Lin  S, Peterson  RM, Chen  Y, Wendt-Pienkowski  E, Rajski  SR, Shen  B. Dedicated ent-kaurene and ent-atiserene synthases for platensimycin and platencin biosynthesis. Proc Natl Acad Sci U S A.  2011:108(33):13498–13503. 10.1073/pnas.1106919108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  128. Smirnov  V, Warnow  T. MAGUS: multiple sequence alignment using graph clUStering. Bioinformatics. 2021:37(12):1666–1672. 10.1093/bioinformatics/btaa992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  129. Stowell  EA, Ehrenberger  MA, Lin  Y-L, Chang  C-Y, Rudolf  JD. Structure-guided product determination of the bacterial type II diterpene synthase Tpn2. Commun Chem.  2022:5(1):Article 1. 10.1038/s42004-022-00765-6. [DOI] [Google Scholar]
  130. Summons  RE, Bird  LR, Gillespie  AL, Pruss  SB, Roberts  M, Sessions  AL. Lipid biomarkers in ooids from different locations and ages: evidence for a common bacterial flora. Geobiology. 2013:11(5):420–436. 10.1111/gbi.12047. [DOI] [PubMed] [Google Scholar]
  131. Syrén  P-O, Henche  S, Eichler  A, Nestl  BM, Hauer  B. Squalene-hopene cyclases—evolution, dynamics and catalytic scope. Curr Opin Struct Biol.  2016:41:73–82. 10.1016/j.sbi.2016.05.019. [DOI] [PubMed] [Google Scholar]
  132. Takishita  K, Chikaraishi  Y, Tanifuji  G, Ohkouchi  N, Hashimoto  T, Fujikura  K, Roger  AJ. Microbial eukaryotes that lack sterols. J Eukaryot Microbiol.  2017:64(1):897–900. 10.1111/jeu.12426. [DOI] [PubMed] [Google Scholar]
  133. Tamamura  T, Sawa  T, Isshiki  K, Masuda  T, Homma  Y, Iinuma  H, Naganawa  H, Hamada  M, Takeuchi  T, Umezawa  H. Isolation and characterization of terpentecin, a new antitumor antibiotic. J Antibiot (Tokyo).  1985:38(12):1664–1669. 10.7164/antibiotics.38.1664. [DOI] [PubMed] [Google Scholar]
  134. Tantillo  DJ. The carbocation continuum in terpene biosynthesis—where are the secondary cations?  Chem Soc Rev.  2010:39(8):2847–2854. 10.1039/B917107J. [DOI] [PubMed] [Google Scholar]
  135. Usui  T, Nagumo  Y, Watanabe  A, Kubota  T, Komatsu  K, Kobayashi  J, Osada  H. Brasilicardin A, a natural immunosuppressant, targets amino acid transport system L. Chem Biol.  2006:13(11):1153–1160. 10.1016/j.chembiol.2006.09.006. [DOI] [PubMed] [Google Scholar]
  136. Villanueva  L, Rijpstra  WIC, Schouten  S, Damsté  JSS. Genetic biomarkers of the sterol-biosynthetic pathway in microalgae. Environ Microbiol Rep.  2014:6(1):35–44. 10.1111/1758-2229.12106. [DOI] [PubMed] [Google Scholar]
  137. Vo  NNQ, Nomura  Y, Kinugasa  K, Takagi  H, Takahashi  S. Identification and characterization of bifunctional drimenol synthases of marine bacterial origin. ACS Chem Biol.  2022:17(5):1226–1238. 10.1021/acschembio.2c00163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  138. Vorbrüggen  H, Pakrashi  SC, Djerassi  C. Terpenoide, LIV. Arborinol, ein neuer triterpen-typus. Justus Liebigs Ann Chem.  1963:668(1):57–76. 10.1002/jlac.19636680107. [DOI] [Google Scholar]
  139. Wang  J, Kodali  S, Lee  SH, Galgoci  A, Painter  R, Dorso  K, Racine  F, Motyl  M, Hernandez  L, Tinney  E, et al.  Discovery of platencin, a dual FabF and FabH inhibitor with in vivo antibiotic properties. Proc Natl Acad Sci U S A.  2007:104(18):7612–7616. 10.1073/pnas.0700746104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  140. Wang  J, Soisson  SM, Young  K, Shoop  W, Kodali  S, Galgoci  A, Painter  R, Parthasarathy  G, Tang  YS, Cummings  R, et al.  Platensimycin is a selective FabF inhibitor with potent antibiotic properties. Nature. 2006:441(7091):358. 10.1038/nature04784. [DOI] [PubMed] [Google Scholar]
  141. Wang  S, Luo  H. Dating the bacterial tree of life based on ancient symbiosis. Syst Biol.  2025:syae071. 10.1093/sysbio/syae071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  142. Wei  JH, Yin  X, Welander  PV. Sterol synthesis in diverse bacteria. Front Microbiol.  2016:7:1–19. 10.3389/fmicb.2016.00990. [DOI] [PMC free article] [PubMed] [Google Scholar]
  143. Welander  PV, Coleman  ML, Sessions  AL, Summons  RE, Newman  DK. Identification of a methylase required for 2-methylhopanoid production and implications for the interpretation of sedimentary hopanes. Proc Natl Acad Sci U S A.  2010:107(19):8537–8542. 10.1073/pnas.0912949107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  144. Welander  PV, Doughty  DM, Wu  C-H, Mehay  S, Summons  RE, Newman  DK. Identification and characterization of Rhodopseudomonas palustris TIE-1 hopanoid biosynthesis mutants: hopanoid biosynthesis in R. palustris TIE-1. Geobiology. 2012:10(2):163–177. 10.1111/j.1472-4669.2011.00314.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  145. Welander  PV, Hunter  RC, Zhang  L, Sessions  AL, Summons  RE, Newman  DK. Hopanoids play a role in membrane integrity and pH homeostasis in Rhodopseudomonas palustris TIE-1. J Bacteriol.  2009:191(19):6145–6156. 10.1128/JB.00460-09. [DOI] [PMC free article] [PubMed] [Google Scholar]
  146. Welander  PV, Summons  RE. Discovery, taxonomic distribution, and phenotypic characterization of a gene required for 3-methylhopanoid production. Proc Natl Acad Sci U S A.  2012:109(32):12905–12910. 10.1073/pnas.1208255109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  147. Wendt  KU. Enzyme mechanisms for triterpene cyclization: new pieces of the puzzle. Angew Chem Int Ed. 2005:44(26):3966–3971. 10.1002/anie.200500804. [DOI] [Google Scholar]
  148. Wendt  KU, Lenhart  A, Schulz  GE. The structure of the membrane protein squalene-hopene cyclase at 2.0 Å resolution. J Mol Biol.  1999:286(1):175–187. 10.1006/jmbi.1998.2470. [DOI] [PubMed] [Google Scholar]
  149. Xia  X, Xie  Z, Salemi  M, Chen  L, Wang  Y. An index of substitution saturation and its application. Mol Phylogenet Evol.  2003:26(1):1–7. 10.1016/S1055-7903(02)00326-3. [DOI] [PubMed] [Google Scholar]
  150. Xu  R, Fazio  GC, Matsuda  SPT. On the origins of triterpenoid skeletal diversity. Phytochemistry. 2004:65(3):261–291. 10.1016/j.phytochem.2003.11.014. [DOI] [PubMed] [Google Scholar]
  151. Xu  X, London  E. The effect of sterol structure on membrane lipid domains reveals how cholesterol can induce lipid domain formation. Biochemistry. 2000:39(5):843–849. 10.1021/bi992543v. [DOI] [PubMed] [Google Scholar]
  152. Yang  Z, Dos Reis  M. Statistical properties of the branch-site test of positive selection. Mol Biol Evol.  2011:28(3):1217–1228. 10.1093/molbev/msq303. [DOI] [PubMed] [Google Scholar]
  153. Young  DC, Layre  E, Pan  S-J, Tapley  A, Adamson  J, Seshadri  C, Wu  Z, Buter  J, Minnaard  AJ, Coscolla  M, et al.  In vivo biosynthesis of terpene nucleosides provides unique chemical markers of Mycobacterium tuberculosis infection. Chem Biol.  2015:22(4):516–526. 10.1016/j.chembiol.2015.03.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  154. Zhai  L, Bonds  AC, Smith  CA, Oo  H, Chou  JC-C, Welander  PV, Dassama  LM. Novel sterol binding domains in bacteria. eLife. 2024:12:RP90696. 10.7554/eLife.90696. [DOI] [PMC free article] [PubMed] [Google Scholar]
  155. Zhang  Q, Yu  Y, Vélasquez  JE, van der Donk  WA. Evolution of lanthipeptide synthetases. Proc Natl Acad Sci U S A.  2012:109(45):18361–18366. 10.1073/pnas.1210393109. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

msaf203_Supplementary_Data

Data Availability Statement

The data supporting the findings in this study are available in the main text and supplementary information. The terpenoid cyclase alignment and phylogeny (shown in Figs. 2, 3, and 5) are available as supplementary files.


Articles from Molecular Biology and Evolution are provided here courtesy of Oxford University Press

RESOURCES