Skip to main content
Engineering Microbiology logoLink to Engineering Microbiology
. 2023 Mar 29;3(3):100085. doi: 10.1016/j.engmic.2023.100085

Recent advances in the direct cloning of large natural product biosynthetic gene clusters

Jiaying Wan 1, Nan Ma 1, Hua Yuan 1,
PMCID: PMC11611023  PMID: 39628928

Abstract

Large-scale genome-mining analyses have revealed that microbes potentially harbor a huge reservoir of uncharacterized natural product (NP) biosynthetic gene clusters (BGCs), and this has spurred a renaissance of novel drug discovery. However, the majority of these BGCs are often poorly or not at all expressed in their native hosts under laboratory conditions, and thus are regarded as silent/orphan BGCs. Currently, connecting silent BGCs to their corresponding NPs quickly and on a large scale is particularly challenging because of the lack of universal strategies and enabling technologies. Generally, the heterologous host-based genome mining strategy is believed to be a suitable alternative to the native host-based approach for prioritization of the vast and ever-increasing number of uncharacterized BGCs. In the last ten years, a variety of methods have been reported for the direct cloning of BGCs of interest, which is the first and rate-limiting step in the heterologous expression strategy. Essentially, each method requires that the following three issues be resolved: 1) how to prepare genomic DNA; 2) how to digest the bilateral boundaries for release of the target BGC; and 3) how to assemble the BGC and the capture vector. Here, we summarize recent reports regarding how to directly capture a BGC of interest and briefly discuss the advantages and disadvantages of each method, with an emphasis on the notion that direct cloning is very beneficial for accelerating genome mining research and large-scale drug discovery.

Keywords: Natural product, Silent BGCs, Genome mining, Direct cloning, Heterologous expression

Graphical abstract

Image, graphical abstract

1. Introduction

The discovery of the Nobel Prize molecule penicillin produced by the microbe Penicillium chrysogenum marked the beginning of the antibiotic era, and it is regarded as a most valuable contribution to modern medicine. Since then, natural products (NPs) isolated from microbes have had a particularly large impact on human medicine, animal health, and plant crop protection. For example, numerous small molecule drugs used in the clinic are either original or modified NPs; 67% of anti-infective agents and 83% of anticancer drugs used in the past 40 years are of natural origin [1]. Most classes of antibiotics used today were identified in the 1940s to 1960s, mainly isolated through the Waksman platform-guided massive screening of microbial fermentation broths. This approach, which was very effective and fruitful in the initial screening programs, led to the creation of the Golden Age of antibiotics [2]. However, the frequency of discovery of new antibiotics from randomly screened microbes has decreased by over six orders of magnitude, from 10−1 to 10−7, with screens gradually becoming more and more unproductive due to the high rate of re-isolation or rediscovery of known compounds [3,4]. Eventually, novel antibiotic scaffolds became harder to identify using this bioactivity-guided screening strategy, causing fewer new antibiotic drugs to enter the market [5]. At the same time, the use and misuse of antibiotics led to the inevitable rise of antimicrobial resistance, which has become a looming global crisis [5,6]. Therefore, an alternative revolutionary strategy for novel drug discovery is required to meet the perpetual need for new antibiotics.

Biosynthetic studies of numerous microbial NPs have revealed that the genes responsible for biosynthesis, self-resistance, regulation, and transport are generally physically clustered in the genome forming a biosynthetic gene cluster (BGC) [7,8]. This evolved characteristic strongly facilitates horizontal gene transfer among microbes, which can lead to multiple microbial species producing the same NP [9]. Significantly, the complete genome sequencing of the first two Streptomyces chromosomes (from Streptomyces coelicolor and Streptomyces avermitilis) clearly established that, even though derived from two of the most well-studied bacterial species, these chromosomes contain numerous predicted BGCs with as yet unknown functions under standard culturing conditions [10], [11], [12]. Since then, with the advent of fast and relatively inexpensive genome sequencing technologies, microbial BGCs have been uncovered at an unprecedented rate. Currently, however, only a very small fraction of BGCs (about 2500 as of February 2023) have been characterized [13], and the vast majority are not connected to any known compounds. Because most of the annotated BGCs are poorly or not at all expressed in their native hosts under conventional culture conditions, they are designated as silent/orphan BGCs. These BGCs are believed to be a great resource for novel drug discovery. For example, through bioinformatics analyses of the product diversity of orphan assembly-line polyketide synthases (PKSs), Nivina et al. concluded that more than one-half of all orphan assembly lines could likely produce novel polyketide structures [14].

Currently, translating silent BGCs to their encoded products through genome mining quickly and on a large scale is particularly challenging because of the lack of universal strategies and enabling technologies. To date, a rich variety of methods have been reported for the activation of BGCs, which mainly include native host-based and heterologous host-based strategies. For genome mining in natural producers, the diverse methods mainly fall into the following two groups [15], [16], [17], [18], [19], [20], [21], [22]: 1) BGC-targeted (pathway-specific) approaches, such as the overexpression of positive regulator genes and/or the deletion of negative regulator genes, and the insertion of a strong promoter upstream of a target BGC; and 2) non-targeted (pleiotropic) approaches, such as ribosome engineering, media optimization, co-culturing, and the addition of elicitors. One of the advantages of this native host-based strategy is the maintenance of the integrity of the target BGC and the precursor supply. The genetic manipulation system developed in natural producers can facilitate the rewiring of the transcriptional regulatory circuit of the BGC of interest. However, when taking into consideration that developing genetic manipulation tools is varied from strain to strain, time-consuming, and often fruitless, the native host-based strategy is not generally applicable for accessing the NP-producing potential of most microbes. Moreover, this strategy is completely unsuitable for the exploitation of some genetically intractable strains and uncultivated microbes, which also possess a vast array of NP BGCs as revealed by metagenomics analyses [23], [24], [25], [26].

The heterologous host-based strategy, on the other hand, is more flexible and can be applied to all BGCs from both cultivated microbes and those yet to be cultivated in the laboratory. The steps of this strategy mainly include: 1) the cloning (or refactoring) of BGCs from microbial genomic DNA; 2) heterologous expression in a genetically tractable host (such as S. coelicolor and Escherichia coli); and 3) chemical analyses and structure determination of the heterologously expressed NPs. Therefore, two types of techniques are the key drivers of bioactive NP discovery using the heterologous host-based strategy. The first one includes the various methods developed for the cloning of a BGC of interest, and the second is the use of modern analytical techniques, such as high-performance liquid chromatography (HPLC), high-resolution mass spectrometry (HRMS), tandem mass spectrometry (MS/MS), and nuclear magnetic resonance (NMR), which have greatly facilitated analyses and structure elucidation of organic chemical compounds [27]. To date, PCR amplification or complete de novo DNA synthesis of a large BGC (e.g., >30 kb) is still impractical for most research laboratories. Therefore, the cloning of large BGCs is still regarded as a rate-limiting step, although various cloning methods have been developed [28], [29], [30], such as library construction, PCR-based bottom-up assembly, and direct cloning. Currently, the construction of genomic libraries (e.g., cosmid/fosmid, BAC, PAC, FAC) is used to capture all classes of NP BGCs [31], [32], [33], [34], [35]. However, library construction requires extensive screening (1000–2000 clones), making it laborious and time-consuming. Moreover, due to its untargeted nature, the library-based cloning method is also very inefficient in prioritizing BGCs from a vast number of uncharacterized BGCs. The bottom-up assembly methods frequently need to be performed in combination with other assembly methods (e.g., type IIS restriction endonuclease [36], Gibson assembly [37]), but the assembly efficiency is severely limited by the length, amount of repetitive sequences, and GC content of target BGCs [37]. In addition, random mutations can be easily introduced during the multiple PCR amplifications and the assembly process itself.

In stark contrast, direct cloning offers many advantages over other cloning and assembly methods [38]. In direct cloning BGCs of interest are directly captured from genomic DNA; thus, this approach can both circumvent the library construction and screening steps and minimize the introduction of random mutations by PCR amplification. Thus, direct cloning represents the most promising strategy for the BGC prioritization-based large-scale discovery of novel bioactive NPs. All the direct cloning methods essentially include the following three parts: 1) the preparation of microbial genomic DNA; 2) the digestion of genomic DNA for release of the target BGC; and 3) the assembly of the BGC and the capture plasmid. In this Review, we briefly summarize recent advances in the direct cloning of BGCs by dissecting the above three parts of each method. In the future, we believe that complete de novo DNA synthesis (combined with bottom-up refactoring) of BGCs will rapidly facilitate NP genome mining. Currently, however, we believe that direct cloning is very important for large-scale novel drug discovery to combat the growing problem of antimicrobial resistance.

2. In vivo DNA circularization between the target BGC and a capture plasmid

2.1. Saccharomyces cerevisiae transformation-associated recombination-based method

Homologous DNA molecules transformed into yeast are highly recombinogenic, and by taking advantage of this feature, in 1996, Larionov et al. developed a revolutionary method called transformation-associated recombination (TAR), which can be harnessed to selectively clone large consecutive DNA fragments from gently isolated human genomic DNA [39,40]. Fourteen years later in 2010, Kim et al. adapted this TAR cloning method and designed a BAC-based S. cerevisiae/E. coli/Streptomyces shuttle capture vector (pTARa) [26]. This pTARa vector enables direct cloning of large BGCs in S. cerevisiae, maintenance and manipulation of the cloned BGC in E. coli, and heterologous expression in Streptomyces hosts via integrative conjugation. In this report, it was demonstrated that pTARa could be used to both directly clone and reassemble BGCs of interest. The authors first employed pTARa to directly clone the 56-kb colibactin BGC from the sequenced Citrobacter koseri genome (Fig. 1a). Specifically, a capture plasmid was first constructed with one homology arm located upstream of the colibactin BGC and another homology arm located downstream. Then, the capture plasmid was linearized by restriction enzyme digestion and co-transformed with the C. koseri genomic DNA (without any enzymatic digestion) into S. cerevisiae spheroplasts for in vivo DNA circularization between the target BGC and the linear capture plasmid.

Fig. 1.

Fig. 1

Examples of in vivo DNA circularization between targeted BGC and a capture plasmid. (a) In the S. cerevisiae TAR-based method, a capture plasmid with two homology arms (red and green solid boxes) is linearized via restriction enzyme (RE) digestion and co-transformed with the microbial genomic DNA (with/without RE digestion) into S. cerevisiae spheroplasts for in vivo DNA circularization. (b) In the LLHR method, the PCR products with two designed homology arms are obtained by the amplification of a universal receiver vector, and co-transformed with the microbial genomic DNA (with RE digestion) into an engineered E. coli strain, GB05-dir. (c) In the CAPTURE method, the capture vector is split into two fragments with a loxP site at the ends and obtained by the PCR amplification of designed universal receiver vectors. The target BGC is released by the Cas12a digestion of purified genomic DNA and pre-assembled in vitro with the two capture vector fragments via the T4 polymerase exo + fill-in DNA assembly approach. Finally, the pre-assembled linearized products are electrotransformed into an engineered E. coli strain for in vivo DNA circularization.

In the same report, these authors also applied pTARa to create complete BGCs by stitching overlapping cosmid clones. Similarly, for each reassembly experiment, a unique pathway-specific capture plasmid with homology arms was constructed. Meanwhile, the overlapping cosmids required for reassembly of a complete BGC were first digested with the restriction enzyme DraI (which recognizes the AT-rich hexamer, TTTAAA, and mainly digests the cosmid backbone), and then co-transformed with the linearized capture plasmid into competent S. cerevisiae cells for the generation of a complete BGC. They used this strategy to successfully reassemble three complete BGCs (a 39-kb PKS gene cluster, an 89-kb non-ribosomal peptide synthetase [NRPS] gene cluster, and a 90-kb friulimicin BGC) from soil-derived environmental DNA cosmid libraries.

Later in 2014, Yamanaka et al. also designed and developed a S. cerevisiae-E. coli-actinobacterial chromosome integrative capture vector, pCAP01, around the SuperCos1 cosmid [41]. One of the main differences between this vector and pTARa is that the pCAP01 vector is equipped with the pUC ori element and thus functions at multiple copies in E. coli. In this report, the authors used pCAP01 to successfully TAR clone a 73-kb genomic region containing the taromycin (tar) BGC from Saccharomonospora sp. CNQ490, which shows sequence similarity with the BGC of the clinically approved antibiotic daptomycin. Specifically, the tar pathway-specific capture plasmid with 1-kb homologous DNA arms was first constructed. Then, this construct was linearized and co-transformed with the restriction enzyme XbaI-digested genomic DNA into S. cerevisiae VL6–48 spheroplasts. Initially, this directly cloned tar cluster was not efficiently expressed in the model host S. coelicolor M1146. However, through subsequent regulatory gene remodeling the tar cluster was successfully activated, leading to the production of taromycin A, which has notable structural differences from daptomycin in three amino acid residues and the lipid side chain. Remarkably, by performing NcoI restriction mapping analysis of two different E. coli clones the authors observed that the 81.8-kb pCAP01-tar plasmid could be stably propagated in E. coli. Therefore, they suggested that the pCAP01 vector can be used to directly TAR clone most NP BGCs.

Since the authors observed that nonhomologous end joining led to high levels of self-recircularization of the capture plasmid, which greatly reduced the efficiency of capture rates to below 2%, in 2015 they constructed another capture vector, pCAP03, with the counter-selectable marker gene URA3 placed under the strong promoter of the Schizosaccharomyces pombe ADH1 gene (pADH1) by ligating it into the SpeI and KpnI restriction sites of pCAP01 [42,43]. Furthermore, based on the knowledge that pADH1 can tolerate an insertion of up to 130 bp of DNA between the TATA box and the transcription initiation site, the authors then designed a ready-to-use pCAP03 vector (RTU-pCAP03) for BGC capture. Specifically, a 144-bp DNA fragment with the following three features is synthesized: 1) two 18-bp overlapping fragments used for Gibson assembly with pADH1 and URA3 of the pCAP03 vector; 2) two 50-bp homologous DNA arms for recombination with BGC boundaries; and 3) an 8-bp PmeI restriction site between the two homologous DNA arms for linearization. Subsequently, they used RTU-pCAP03 to successfully clone the ∼22-kb thiolactomycin and 33-kb thiotetroamide BGCs with high efficiency (66.7% and 20%, respectively). Therefore, the RTU-pCAP03-based method can eliminate the need for PCR for capture plasmid construction and also greatly enhance the capture efficiency.

In 2019, Li et al. constructed pCL01 by replacing the pUC ori element of pCAP01 with the copy-control element from the plasmid pCC1BAC. The resulting construct pCL01 facilitates the capture of larger BGCs in a single-copy form, and the addition of a copy-control inducer to this construct can increase the number of cloned BGCs to 10–20 copies per E. coli cell [44]. The authors then used pCL01 to successfully capture the 5-oxomilbemycin BGC and enhanced the production of 5-oxomilbemycin using an advanced multiplex site-specific genome engineering strategy [44].

Notably, all the above TAR-cloned BGCs were introduced into Streptomyces hosts for heterologous expression or overexpression. To extend the number of heterologous hosts, the Moore group, in collaboration with other groups, has developed a series of TAR cloning vectors, such as the yeast/E. coli shuttle-Bacillus subtilis chromosome integrative capture vector pCAPB02 [45], and the yeast/broad-host-range Gram-negative host expression vector pCAP05 [46]. Therefore, a series of TAR cloning vectors supporting the production of NPs in diverse hosts have been developed and optimized.

2.2. Rac prophage RecET-based method

In 2012, Fu et al. developed a strategy called LLHR (linear plus linear homologous recombination) involving homologous recombination between two linear DNA molecules using the full-length Rac prophage protein RecE and its partner RecT [38,47]. The main steps of this strategy are as follows: 1) designing pathway-specific PCR primers and generation of a linear capture vector with two homology arms, each >30 bp; 2) restriction enzyme digestion of purified genomic DNA to release the BGC of interest; and 3) co-electrotransformation of the linear capture vector plus digested genomic DNA into an engineered E. coli strain GB05-dir (with RecET-Redγ-RecA under the control of the PBAD promoter integrated in the chromosome) and validation of positive clones (Fig. 1b). They applied this platform to directly clone nine of the ten PKS-NRPS gene clusters (each 10–52 kb in length) from the Photorhabdus luminescens genome. Subsequently, through heterologous expression in E. coli, they successfully identified the metabolites luminmycin A (encoded by the plu1881plu1877 gene cluster) and luminmide A/B (encoded by plu3263) and proposed their biosynthetic pathways accordingly.

Of note, because of the high background levels of capture vector recircularization, the tenth and largest gene cluster, plu2670 (∼52 kb), could not be directly cloned using LLHR alone. The authors then tried an alternative two-step, double recombination ‘fishing’ strategy by combining LLHR with Redαβ-mediated linear plus circular homologous recombination (LCHR). With this two-step cloning method, they finally obtained the 52-kb plu2670 gene cluster (6/21 correct clones).

Later, in 2018, the same group improved the above RecET-based direct cloning method and developed the ExoCET platform (Exonuclease Combined with RecET recombination), which is based on the notion that RecET-based direct cloning efficiencies can be enhanced by annealing the linear capture vector and the targeted BGC in vitro prior to transformation into E. coli [48]. The authors used T4 polymerase as the in vitro 3′ exonuclease to generate 5′ sticky ends, which greatly facilitated annealing between the linear capture vector and its target BGC, thus increasing the probability of the paired DNA molecules simultaneously entering one E. coli cell. Using this optimized platform, they successfully cloned much larger DNA fragments (>50 kb) from bacterial and mammalian genomes with high efficiency, including the salinomycin BGC (∼106 kb) from both EcoRV- and Cas9-digested genomic DNA preparations. Very recently, this ExoCET platform was used to directly clone a 142-kb pseudorabies virus genome into a BAC vector [49]. The resulting infectious viral BAC was stably propagated in E. coli, thus facilitating studies of the biology of this virus and the development of vaccines.

Notably, in 2019, Song et al. used ExoCET to successfully join 12 overlapping PCR products with a linearized capture plasmid to rebuild an artificial 79-kb spinosad BGC in one round with a success rate of 58% [50]. In 2020, the same group developed a RedEx method (by combining Redαβ-mediated LCHR, ccdB counterselection, and exonuclease-mediated in vitro annealing) for seamless DNA insertion and deletion in large multimodular PKS gene clusters [51].

A number of platforms have been developed by exploiting RecET- and Redαβ-mediated homologous recombination, and these platforms have facilitated the direct cloning of DNA sequences from diverse and complex sources, the stitching of overlapping DNA fragments, and the seamless DNA mutation.

2.3. Cre-lox-based method

In 2021, Enghiad et al. reported a highly efficient direct cloning method named Cas12a-assisted precise targeted cloning using in vivo Cre-lox recombination (CAPTURE) [52]. The main steps of this method are as follows: 1) Cas12a-sgRNA-mediated digestion of purified genomic DNA; 2) PCR amplification of two pathway-specific capture plasmid fragments that each carry a loxP site at their ends; 3) joining the digested genomic DNA and the two capture plasmid fragments via the T4 polymerase exo + fill-in DNA assembly approach; and 4) transformation of the in vitro pre-assembled linearized products into an engineered E. coli strain (that expresses the Cre recombinase and the Redγ protein) for intramolecular DNA circularization (Fig. 1c). The authors applied this method to successfully clone a total of 43 uncharacterized NP BGCs (10–113 kb in size) from 14 Streptomyces and three Bacillus species without any failure (close to 100% cloning efficiency). Subsequent heterologous expression of the 43 clusters revealed seven BGCs with positive HPLC peaks, and 15 previously uncharacterized NPs were identified from five of these seven BGCs. Among these compounds were bipentaromycins, which showed strong antimicrobial activity toward both Gram-positive and Gram-negative bacteria. Therefore, the authors suggested that using this method to directly capture BGCs is inexpensive, rapid, robust, and highly efficient, and thus suitable for large-scale discovery of novel NPs.

3. In vitro DNA circularization between the target BGC and a capture plasmid

3.1. Gibson assembly-based method

In 2015, Jiang et al. reported a method, Cas9-assisted targeting of chromosome segments (CATCH), that allows the cleavage of a user-defined DNA region in vitro from intact bacterial chromosomes embedded in agarose plugs, which can be subsequently ligated with a capture plasmid through Gibson assembly [53,54]. The main steps are as follows: 1) agarose plug preparation and in-gel cell lysis; 2) sgRNA preparation and in-gel Cas9 digestion; 3) capture plasmid construction and Gibson assembly-mediated circularization; and 4) electrotransformation and validation of positive clones (Fig. 2a). As a proof of principle, the authors first demonstrated that their method could enable one-step targeted cloning of long E. coli genomic regions of up to 100 kb with good efficiency. Then, to demonstrate the application of CATCH, they used it to successfully clone three known NP BGCs, namely the 78-kb bacillaene BGC from B. subtilis, the 36-kb jadomycin BGC from Streptomyces venezuelae, and the 32-kb chlortetracycline BGC from Streptomyces aureofaciens.

Fig. 2.

Fig. 2

Examples of in vitro DNA circularization between targeted BGC and a capture plasmid. (a) In the CATCH or CAT-FISHING method, the microbial cells are first embedded and lysed in agarose plugs. The BGC of interest is released by Cas9/12a digestion, and the plugs are liquified by agarase. Then, the target BGC in the liquefied DNA mixture is used for ligation into a capture plasmid through Gibson assembly/DNA ligation. Finally, the circularization products are electrotransformed into E. coli cells. (b) In the in vitro λ packaging-assisted method, the purified genomic DNA is first dephosphorylated and digested by Cas9, and then ligated into the EcoRV-linearized and dephosphorylated universal vector pJTU2554 by T4 DNA ligase. The ligation products are used for the subsequent in vitro λ packaging and infection of E. coli EPI300.

Compared with traditional in-solution digestion of purified genomic DNA, the advantage of in-gel Cas9 cleavage is that the chromosomal DNA is well protected from mechanical shearing by the agarose gel matrix, which greatly reduces background DNA fragments and thus can lead to generation of fewer falsely ligated clones. In addition, the utility of Gibson assembly-mediated in vitro circularization saves time because the assembly products can be directly transformed into E. coli or other bacteria. However, when handling DNA segments longer than 100 kb, the CATCH method appears to be less efficient. For example, the authors obtained only one colony containg a 150-kb E. coli genomic fragment from three independent trials (51 colonies in total). However, when considering that most of the known microbial BGCs are less than 100 kb in length [13,55], this method is likely sufficient for the direct cloning of most BGCs of interest.

3.2. λ packaging-based method

In 2019, Tao et al. designed an elegant in vitro λ packaging-based method for one-step targeted cloning of NP pathways [56]. The main steps of this method are as follows: 1) genomic DNA isolation and dephosphorylation; 2) sgRNA preparation and Cas9 digestion; 3) T4 DNA ligase-mediated ligation between the digested genomic DNA mixture and the EcoRV-linearized and dephosphorylated universal vector pJTU2554 (an integrative Streptomyces cosmid [57]); and 4) in vitro λ packaging of the ligation products and infection of E. coli EPI300 (Fig. 2b). The authors employed this method to directly obtain the 27.4-kb Tü3010 (stu) and the 40.7-kb sisomicin (sis) BGCs with high cloning efficiencies of 18% and 54%, respectively. Because the λ phage has a packaging size limit of 37.4–50.4 kb (78–105% of the wild-type genome size) the cloning efficiencies for the stu and sis BGCs were also believed to test the lower and upper packaging limits, respectively. Therefore, one advantage of this method is the high λ phage packaging and infection efficiency, which will facilitate subsequent positive clone verification. Another is that the EcoRV-linearized vector is used for blunt-end ligation and thus can be for all pathways without the introduction of homologous DNA sequences for recombination. Hence, in combination with the CRISPR/Cas9 system for specific release of the target BGC, λ packaging system-assisted infection of E. coli cells can directly and quickly capture BGCs of ∼30–40 kb with high efficiency.

3.3. DNA ligase-based method

In 2022, Liang et al. reported an in vitro platform for directly capturing large BGCs, named CAT-FISHING (CRISPR/Cas12a-mediated fast direct BGC cloning) [55]. In this platform, the method for release of targeted BGCs from genomic DNA is similar to that of the CATCH method, but the main difference is using Cas12a instead of Cas9 for in-gel digestion (Fig. 2a). Unlike Cas9, Cas12a generates sticky ends on the target BGC and thus inspired the authors to develop a DNA ligase-based method. The capture plasmid is likewise linearized by Cas12a so as to create complementary sticky ends with the released target BGC. Then, the in-gel digested genomic DNA is ligated into the linearized capture plasmid using E. coli DNA ligase. Finally, the ligation products are electrotransformed into E. coli for validation of positive clones.

Using this CAT-FISHING method, the authors successfully cloned eight BGCs (with GC contents ranging from 66% to 76% and lengths from 41 kb to 145 kb) from different actinomycetes with very good efficiency (ranging from 8% to 55%). Furthermore, when combined with isolation and purification of the target BGC DNA fragments using pulsed field gel electrophoresis (PFGE), the capture efficiency was greatly increased to ∼70%. The high efficiency is mainly attributed to the greatly reduced number of genomic DNA fragments in the in-gel digested products, which reduces the interference with the ligation reaction.

4. Conclusions and outlook

Rapid and cost-effective sequencing technologies have led to the exponential accumulation of NP BGCs present in microbial genomes, the vast majority of which are uncharacterized, and thus have spurred a renaissance of novel drug discovery [58]. Because the rate of chemical decoding of silent BGCs cannot keep up with their identification, prioritization of BGCs for further activation is required. Although automated and high-throughput biofoundries have been successfully developed for discovery of new terpenoids and RiPPs (ribosomally synthesized and post-translationally modified peptides) [59,60], the lengths of the BGCs for such classes of NPs are relatively small and easily refactored. Currently, however, we believe that direct cloning is suitable for targeting prioritized BGCs from the vast number of microbial BGCs that have accumulated, and that it can facilitate large-scale discovery of novel bioactive NPs. As discussed above, the traditional preparation of high-quality and high-molecular-weight genomic DNA (especially >100 kb) may be technically challenging for most research laboratories [29]. Instead, the preparation and manipulation of genomic DNA in agarose gel plugs can avoid mechanical shearing and is easy to perform even though the cost is a bit high. CRISPR/Cas-mediated digestion can be designed at near-arbitrary DNA sites and thus is the most precise and quickest way to release a target BGC [53,55]. To date, the biggest BGC cloned using direct cloning methods is the 145-kb candicidin BGC (GC content 75%; PKS) from Streptomyces albus J1074, which was cloned by CAT-FISHING [55]. The direct cloning method is more convenient and more efficient than BAC library construction, which uses an endonuclease for partial digestion of genomic DNA and requires further purification of DNA fragments (e.g., 75–145 kb) by PFGE. Notably, using BAC vectors Hashimoto et al. have successfully cloned the largest PKS gene cluster (for the biosynthesis of quinolidomicin, >215 kb) [35]. TAR cloning has enabled the direct cloning of a chromosomal DNA fragment of up to almost 300 kb in length from complex mammalian genomes, and thus so far it is undoubtedly the most reliable method for selective isolation of very large chromosomal regions [61]. However, the majority of BGCs of interest often contain large PKS and NRPS genes, whose intrinsically repetitive and high GC-content sequences probably lead to unexpected inter- and intra-molecular recombination and thus the formation of unwanted BGCs (which are difficult to detect without sequencing). In summary, although different research laboratories may leverage different methods to acquire targeted BGCs, coordinated efforts from academia and industry are necessarily required to provide more drug lead compounds for use in the fight against the growing problem of antimicrobial resistance.

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

This work was financially supported by the National Key Research and Development Program of China (2021YFC2100600), the National Natural Science Foundation of China (31870034), and the Science and Technology Commission of Shanghai Municipality (20ZR1469100).

References

  • 1.Newman D.J., Cragg G.M. Natural products as sources of new drugs over the nearly four decades from 01/1981 to 09/2019. J. Nat. Prod. 2020;83:770–803. doi: 10.1021/acs.jnatprod.9b01285. [DOI] [PubMed] [Google Scholar]
  • 2.Baltz R.H. Natural product drug discovery in the genomic era: realities, conjectures, misconceptions, and opportunities. J. Ind. Microbiol. Biotechnol. 2019;46:281–299. doi: 10.1007/s10295-018-2115-4. [DOI] [PubMed] [Google Scholar]
  • 3.Baltz R.H. Marcel Faber Roundtable: is our antibiotic pipeline unproductive because of starvation, constipation or lack of inspiration? J. Ind. Microbiol. Biotechnol. 2006;33:507–513. doi: 10.1007/s10295-005-0077-9. [DOI] [PubMed] [Google Scholar]
  • 4.Pye C.R., Bertin M.J., Lokey R.S., et al. Retrospective analysis of natural products provides insights for future discovery trends. Proc. Natl. Acad. Sci. USA. 2017;114:5601–5606. doi: 10.1073/pnas.1614680114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Fischbach M.A., Walsh C.T. Antibiotics for emerging pathogens. Science. 2009;325:1089–1093. doi: 10.1126/science.1176667. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Brown E.D., Wright G.D. Antibacterial drug discovery in the resistance era. Nature. 2016;529:336–343. doi: 10.1038/nature17042. [DOI] [PubMed] [Google Scholar]
  • 7.Malpartida F., Hopwood D.A. Molecular cloning of the whole biosynthetic pathway of a Streptomyces antibiotic and its expression in a heterologous host. Nature. 1984;309:462–464. doi: 10.1038/309462a0. [DOI] [PubMed] [Google Scholar]
  • 8.Hopwood D.A., Malpartida F., Kieser H.M., et al. Production of 'hybrid' antibiotics by genetic engineering. Nature. 1985;314:642–644. doi: 10.1038/314642a0. [DOI] [PubMed] [Google Scholar]
  • 9.Fischbach M.A., Walsh C.T., Clardy J. The evolution of gene collectives: how natural selection drives chemical innovation. Proc. Natl. Acad. Sci. USA. 2008;105:4601–4608. doi: 10.1073/pnas.0709132105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Hopwood D.A. The Streptomyces genome–be prepared! Nat. Biotechnol. 2003;21:505–506. doi: 10.1038/nbt0503-505. [DOI] [PubMed] [Google Scholar]
  • 11.Bentley S.D., Chater K.F., AM C.T., et al. Complete genome sequence of the model actinomycete Streptomyces coelicolor A3(2) Nature. 2002;417:141–147. doi: 10.1038/417141a. [DOI] [PubMed] [Google Scholar]
  • 12.Omura S., Ikeda H., Ishikawa J., et al. Genome sequence of an industrial microorganism Streptomyces avermitilis: deducing the ability of producing secondary metabolites. Proc. Natl. Acad. Sci. USA. 2001;98:12215–12220. doi: 10.1073/pnas.211433198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Kautsar S.A., Blin K., Shaw S., et al. MIBiG 2.0: a repository for biosynthetic gene clusters of known function. Nucl. Acid. Res. 2020;48:D454–D458. doi: 10.1093/nar/gkz882. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Nivina A., Yuet K.P., Hsu J., et al. Evolution and diversity of assembly-line polyketide synthases. Chem. Rev. 2019;119:12524–12547. doi: 10.1021/acs.chemrev.9b00525. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Laureti L., Song L., Huang S., et al. Identification of a bioactive 51-membered macrolide complex by activation of a silent polyketide synthase in Streptomyces ambofaciens. Proc. Natl. Acad. Sci. USA. 2011;108:6258–6263. doi: 10.1073/pnas.1019077108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhang M.M., Wong F.T., Wang Y., et al. CRISPR–Cas9 strategy for activation of silent Streptomyces biosynthetic gene clusters. Nat. Chem. Biol. 2017;13:607–609. doi: 10.1038/nchembio.2341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zhang Q., Ren J.W., Wang W., et al. A versatile transcription-translation in one approach for activation of cryptic biosynthetic gene clusters. ACS Chem. Biol. 2020;15:2551–2557. doi: 10.1021/acschembio.0c00581. [DOI] [PubMed] [Google Scholar]
  • 18.Bode H.B., Bethe B., Höfs R., et al. Big effects from small changes: possible ways to explore nature's chemical diversity. Chembiochem. 2002;3:619–627. doi: 10.1002/1439-7633(20020703)3:7<619::AID-CBIC619>3.0.CO;2-9. [DOI] [PubMed] [Google Scholar]
  • 19.Lautru S., Deeth R.J., Bailey L.M., et al. Discovery of a new peptide natural product by Streptomyces coelicolor genome mining. Nat. Chem. Biol. 2005;1:265–269. doi: 10.1038/nchembio731. [DOI] [PubMed] [Google Scholar]
  • 20.Wu C., Zacchetti B., Ram A.F.J., et al. Expanding the chemical space for natural products by Aspergillus-Streptomyces co-cultivation and biotransformation. Sci. Rep. 2015;5:10868. doi: 10.1038/srep10868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Seyedsayamdost M.R. High-throughput platform for the discovery of elicitors of silent bacterial gene clusters. Proc. Natl. Acad. Sci. USA. 2014;111:7266–7271. doi: 10.1073/pnas.1400019111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Ochi K. Insights into microbial cryptic gene activation and strain improvement: principle, application and technical aspects. J. Antibiot. 2017;70:25–40. doi: 10.1038/ja.2016.82. [DOI] [PubMed] [Google Scholar]
  • 23.Brady S.F., Chao C.J., Handelsman J., et al. Cloning and heterologous expression of a natural product biosynthetic gene cluster from eDNA. Org. Lett. 2001;3:1981–1984. doi: 10.1021/ol015949k. [DOI] [PubMed] [Google Scholar]
  • 24.Brady S.F., Chao C.J., Clardy J. New natural product families from an environmental DNA (eDNA) gene cluster. J. Am. Chem. Soc. 2002;124:9968–9969. doi: 10.1021/ja0268985. [DOI] [PubMed] [Google Scholar]
  • 25.Brady S.F., Simmons L., Kim J.H., et al. Metagenomic approaches to natural products from free-living and symbiotic organisms. Nat. Prod. Rep. 2009;26:1488–1503. doi: 10.1039/b817078a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kim J.H., Feng Z., Bauer J.D., et al. Cloning large natural product gene clusters from the environment: piecing environmental DNA gene clusters back together with TAR. Biopolymers. 2010;93:833–844. doi: 10.1002/bip.21450. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Atanasov A.G., Zotchev S.B., Dirsch V.M., et al. Natural products in drug discovery: advances and opportunities. Nat. Rev. Drug Discov. 2021;20:200–216. doi: 10.1038/s41573-020-00114-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Zhang J.J., Tang X., Moore B.S. Genetic platforms for heterologous expression of microbial natural products. Nat. Prod. Rep. 2019;36:1313–1332. doi: 10.1039/c9np00025a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wang W., Zheng G., Lu Y. Recent advances in strategies for the cloning of natural product biosynthetic gene clusters. Front. Bioeng. Biotechnol. 2021;9 doi: 10.3389/fbioe.2021.692797. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Li L., Jiang W., Lu Y. New strategies and approaches for engineering biosynthetic gene clusters of microbial natural products. Biotechnol. Adv. 2017;35:936–949. doi: 10.1016/j.biotechadv.2017.03.007. [DOI] [PubMed] [Google Scholar]
  • 31.Miao V., Coëffet-LeGal M.F., Brian P., et al. Daptomycin biosynthesis in Streptomyces roseosporus: cloning and analysis of the gene cluster and revision of peptide stereochemistry. Microbiology. 2005;151:1507–1523. doi: 10.1099/mic.0.27757-0. [DOI] [PubMed] [Google Scholar]
  • 32.Yu Y., Bai L., Minagawa K., et al. Gene cluster responsible for validamycin biosynthesis in Streptomyces hygroscopicus subsp. jinggangensis 5008. Appl. Environ. Microbiol. 2005;71:5066–5076. doi: 10.1128/AEM.71.9.5066-5076.2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Bok J.W., Ye R., Clevenger K.D., et al. Fungal artificial chromosomes for mining of the fungal secondary metabolome. BMC Genom. 2015;16:343. doi: 10.1186/s12864-015-1561-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Jones A.C., Gust B., Kulik A., et al. Phage p1-derived artificial chromosomes facilitate heterologous expression of the FK506 gene cluster. PLoS ONE. 2013;8:e69319. doi: 10.1371/journal.pone.0069319. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hashimoto T., Hashimoto J., Kozone I., et al. Biosynthesis of quinolidomicin, the largest known macrolide of terrestrial origin: identification and heterologous expression of a biosynthetic gene cluster over 200 kb. Org. Lett. 2018;20:7996–7999. doi: 10.1021/acs.orglett.8b03570. [DOI] [PubMed] [Google Scholar]
  • 36.Chen W.H., Qin Z.J., Wang J., et al. The MASTER (methylation-assisted tailorable ends rational) ligation method for seamless DNA assembly. Nucl. Acid. Res. 2013;41:e93. doi: 10.1093/nar/gkt122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Li L., Zhao Y., Ruan L., et al. A stepwise increase in pristinamycin II biosynthesis by Streptomyces pristinaespiralis through combinatorial metabolic engineering. Metab. Eng. 2015;29:12–25. doi: 10.1016/j.ymben.2015.02.001. [DOI] [PubMed] [Google Scholar]
  • 38.Cobb R.E., Zhao H. Direct cloning of large genomic sequences. Nat. Biotechnol. 2012;30:405–406. doi: 10.1038/nbt.2207. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Larionov V., Kouprina N., Graves J., et al. Specific cloning of human DNA as yeast artificial chromosomes by transformation-associated recombination. Proc. Natl. Acad. Sci. USA. 1996;93:491–496. doi: 10.1073/pnas.93.1.491. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Larionov V., Kouprina N., Graves J., et al. Highly selective isolation of human DNAs from rodent-human hybrid cells as circular yeast artificial chromosomes by transformation-associated recombination cloning. Proc. Natl. Acad. Sci. USA. 1996;93:13925–13930. doi: 10.1073/pnas.93.24.13925. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Yamanaka K., Reynolds K.A., Kersten R.D., et al. Direct cloning and refactoring of a silent lipopeptide biosynthetic gene cluster yields the antibiotic taromycin A. Proc. Natl. Acad. Sci. USA. 2014;111:1957–1962. doi: 10.1073/pnas.1319584111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Tang X., Li J., Millán-Aguiñaga N., et al. Identification of thiotetronic acid antibiotic biosynthetic pathways by target-directed genome mining. ACS Chem. Biol. 2015;10:2841–2849. doi: 10.1021/acschembio.5b00658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Zhang J.J., Yamanaka K., Tang X., et al. Direct cloning and heterologous expression of natural product biosynthetic gene clusters by transformation-associated recombination. Methods Enzymol. 2019;621:87–110. doi: 10.1016/bs.mie.2019.02.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Li L., Wei K., Liu X., et al. aMSGE: advanced multiplex site-specific genome engineering with orthogonal modular recombinases in actinomycetes. Metab. Eng. 2019;52:153–167. doi: 10.1016/j.ymben.2018.12.001. [DOI] [PubMed] [Google Scholar]
  • 45.Li Y., Li Z., Yamanaka K., et al. Directed natural product biosynthesis gene cluster capture and expression in the model bacterium Bacillus subtilis. Sci. Rep. 2015;5:9383. doi: 10.1038/srep09383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Zhang J.J., Tang X., Zhang M., et al. Broad-host-range expression reveals native and host regulatory elements that influence heterologous antibiotic production in Gram-negative bacteria. MBio. 2017;8:e01217–e01291. doi: 10.1128/mBio.01291-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Fu J., Bian X., Hu S., et al. Full-length RecE enhances linear-linear homologous recombination and facilitates direct cloning for bioprospecting. Nat. Biotechnol. 2012;30:440–446. doi: 10.1038/nbt.2183. [DOI] [PubMed] [Google Scholar]
  • 48.Wang H., Li Z., Jia R., et al. ExoCET: exonuclease in vitro assembly combined with RecET recombination for highly efficient direct DNA cloning from complex genomes. Nucleic Acids Res. 2018;46:e28. doi: 10.1093/nar/gkx1249. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Yuan H., Zheng Y., Yan X., et al. Direct cloning of a herpesvirus genome for rapid generation of infectious BAC clones. J. Adv. Res. 2023;43:97–107. doi: 10.1016/j.jare.2022.02.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Song C., Luan J., Cui Q., et al. Enhanced Heterologous Spinosad Production from a 79-kb Synthetic Multioperon Assembly. ACS Synth. Biol. 2019;8:137–147. doi: 10.1021/acssynbio.8b00402. [DOI] [PubMed] [Google Scholar]
  • 51.Song C., Luan J., Li R., et al. RedEx: a method for seamless DNA insertion and deletion in large multimodular polyketide synthase gene clusters. Nucl. Acid. Res. 2020;48:e130. doi: 10.1093/nar/gkaa956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Enghiad B., Huang C., Guo F., et al. Cas12a-assisted precise targeted cloning using in vivo Cre-lox recombination. Nat. Commun. 2021;12:1171. doi: 10.1038/s41467-021-21275-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Jiang W., Zhao X., Gabrieli T., et al. Cas9-Assisted Targeting of CHromosome segments CATCH enables one-step targeted cloning of large gene clusters. Nat. Commun. 2015;6:8101. doi: 10.1038/ncomms9101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Jiang W., Zhu T.F. Targeted isolation and cloning of 100-kb microbial genomic sequences by Cas9-assisted targeting of chromosome segments. Nat. Protoc. 2016;11:960–975. doi: 10.1038/nprot.2016.055. [DOI] [PubMed] [Google Scholar]
  • 55.Liang M., Liu L., Xu F., et al. Activating cryptic biosynthetic gene cluster through a CRISPR-Cas12a-mediated direct cloning approach. Nucl. Acid. Res. 2022;50:3581–3592. doi: 10.1093/nar/gkac181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Tao W., Chen L., Zhao C., et al. In vitro packaging mediated one-step targeted cloning of natural product pathway. ACS Synth. Biol. 2019;8:1991–1997. doi: 10.1021/acssynbio.9b00248. [DOI] [PubMed] [Google Scholar]
  • 57.Li L., Xu Z., Xu X., et al. The mildiomycin biosynthesis: initial steps for sequential generation of 5-hydroxymethylcytidine 5′-monophosphate and 5-hydroxymethylcytosine in Streptoverticillium rimofaciens ZJU5119. Chembiochem. 2008;9:1286–1294. doi: 10.1002/cbic.200800008. [DOI] [PubMed] [Google Scholar]
  • 58.Cimermancic P., Medema M.H., Claesen J., et al. Insights into secondary metabolism from a global analysis of prokaryotic biosynthetic gene clusters. Cell. 2014;158:412–421. doi: 10.1016/j.cell.2014.06.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Ayikpoe R.S., Shi C., Battiste A.J., et al. A scalable platform to discover antimicrobials of ribosomal origin. Nat. Commun. 2022;13(1):6135. doi: 10.1038/s41467-022-33890-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Yuan Y., Cheng S., Bian G., et al. Efficient exploration of terpenoid biosynthetic gene clusters in filamentous fungi. Nat. Catal. 2022;5:277–287. [Google Scholar]
  • 61.Kouprina N., Noskov V.N., Larionov V. Selective isolation of large segments from individual microbial genomes and environmental DNA samples using transformation-associated recombination cloning in yeast. Nat. Protoc. 2020;15:734–749. doi: 10.1038/s41596-019-0280-1. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Engineering Microbiology are provided here courtesy of Shandong University

RESOURCES