Abstract
Accurate, economical and high-throughput gene and genome synthesis is essential to the development of synthetic biology and biotechnology. New large scale gene synthesis methods harnessing the power of DNA microchips have recently been demonstrated. Yet, the technology is still compromised by a high occurrence of errors in the synthesized products. These errors still require substantial effort to correct. To solve this bottleneck, novel approaches based on new chemistry, enzymology or next generation sequencing have emerged. This review discusses these new trends and promising strategies of error-filtration, error correction and error-prevention in de novo gene and genome synthesis. Continued innovation in error correction technologies will enable affordable and large scale gene and genome synthesis in the near future.
Keywords: Gene synthesis, error correction, synthetic biology
Gene synthesis
With the rise of synthetic biology, the era of creating new functional genes, genetic networks and whole genomes is upon us. Heralding the dawn of this new era are the rapid technological breakthroughs allowing for on demand synthesis of DNA of any sequence orlength. Recent breakthroughs have resulted in synthesis and assembly of an entire bacterial genome and creation of a new cell controlled by this transplanted synthetic genome[1]. The demand for synthetic genes and genomes will likely continue to increase as the scope of their applications expands.
Gene synthesis is typically accomplished by enzymatic assembly of chemically synthesized overlapping oligonucleotides which span the entire length of the gene construct (reviewed in [2, 3]). The resulting products unavoidably contain errors such as deletions, insertions, or base substitutions, due largely to mistakes in chemical oligonucleotide synthesis and to a lesser extent, to the subsequent enzymatic gene assembly processes. Cloning and sequencing of multiple clones is normally required in order to identify the clones with the correct sequence. Improvement in fidelity of gene synthesis is essential to continued progress in the development of this technology, since a substantial fraction of the overall cost of gene synthesis goes to cloning and sequencing, which is required for selecting and confirming a correct sequence.
Nature has evolved sophisticated error-correction mechanisms to ensure that DNA replication proceeds with high fidelity [4]. The error rates in prokaryotic and eukaryotic replication machineries range from 10−7 to 10−8 thanks to various proofreading and mismatch repair mechanisms [5, 6]. In contrast, current gene synthesis process has a typical error rate of 10−2 to 10−3, or 1–10 errors per kilo base-pairs (kbp) synthesized [7–10].
Given an error rate (P), the probability of a synthetic DNA sequence being error-free, (1-P)N, decreases exponentially as its length (N) increases. The number of clones that need to be sequenced in order to have 95% confidence of obtaining one perfect clone, Ln(1-0.95)/Ln(1-(1-P)N), can be dramatically reduced with a 10-fold reduction in error rate [8]. The presence of a variety of error types and error sources from the series of DNA synthesis and gene assembly steps allows establishment of error-control procedures on multiple levels. This article reviews the current methods for error-filtration, error correction and error-prevention. We also discuss new developments and new directions which may potentially enable error-free synthesis for oligonucleotides, genes, and genomes in the near future.
Error-removal from synthetic oligonucleotides
The dominant source of errors in synthetic DNA comes from chemical synthesis of oligonucleotides. Standard solid-phase oligonucleotide synthesis uses the classical phosphoramidite chemistry[11, 12] which adds each nucleotide monomers to the 5’end of the elongating DNA chain in a four-step cycle: i) Deprotection: an acid is used to remove the protecting demethoxytrityl (DMT) group from the 5’-end of the growing oligonucleotide chain and generate a reactive 5’-OH group; ii) Coupling: the 5’-OH group generated from the deprotection step reacts with an activated monomer created by adding the desired phosphoramidite and an appropriate activator (i.e. tetrazole) simultaneously. Tetrazole, a weak acid, protonates the trivalent phosphorus on the 3’-end of the monomer. This results in a slow displacement of the secondary amine and formation of a highly reactive tetrazolide that then immediately couples with the OH group; iii) Capping: uncoupled 5’-OH groups are blocked by an acylating capping reagent, which is delivered along with a nucleophilic catalyst, to minimize deletion products; and iv) Oxidation: the unstable phosphite triester internucleotide linkages are oxidized to a more stable pentavalent phosphotriester [13].
The most frequent type of errors in oligonucleotide synthesis happens when a new phosphoramidite monomer fails to couple to the elongating chain, which results in a typical stepwise coupling efficiency of 98.5% – 99.5% [14–16]. Uncoupled chains will be terminated from growth by acetylation and result in truncated oligonucleotides. However, failures in acetylation or deprotection do happen with a frequency as high as 0.5% per position, which leads to deletion errors in the final synthetic DNA. Insertions also occur when DMT is cleaved by excess activator and can reach 0.4% per base [15].
Post synthesis, the purity of the synthesized oligonucleotide pool can be improved by size exclusion purification using high performance liquid chromatography (HPLC) [17] or polyacrylamide gel electrophoresis (PAGE) [15]. Hydrophobic purification cartridges can also be used for purification of the oligonucleotide pool before the hydrophobic DMT blocking group is removed (trityl-on). Full-length oligonucleotides with a hydrophobic DMT terminus can be readily separated from prematurely terminated sequences lacking the blocking group. With these methods, more than 90% of the impurities (mostly insertions/deletions and truncations) can be eliminated before assembly and, as a result, the error rate in the final product can be reduced by several fold [18–20]. Although such methods are relatively laborious, the effort can be justified if gene-construction oligos of the same length are pooled and purified together [19]. The drawback of size exclusion purification methods is that they are generally ineffective against base-substitutions or single-base insertions/deletions, especially for long oligos. In fact, single-base deletions are the most frequently observed error type in assembled DNA constructs [8] and they cannot be removed effectively by size exclusion purification methods.
Besides post-synthesis purification, a fundamental approach to increase the accuracy of chemical DNA synthesis is to develop more efficient synthesis chemistry. Along this line, an alternative two-step DNA synthesis methodutilizes a peroxy anion as the nucleophile to simultaneously remove a 5’-carbonate and oxidize the internucleotide phosphite trimester [21]. The removal of the 5’-protecting group with peroxy anion under mildly basic conditions is considered essentially irreversible and quantitative and therefore has the potential to completely eliminate depurination and reduce mutation frequencies in cloned, synthetic DNA. To date, there has been no report on the wide spread use of the new two-step DNA synthesis method. Whether it will become the next new standard remains unclear.
Error-removal from chip-synthesized oligo pools
Oligonucleotide pools synthesized from microarrays have recently been used as an economical source for large scale gene synthesis. However, oligos synthesized on planar surfaces tend to be more prone to errors. Depurination of purine bases seems to be of major concern[22]. Due to prolonged exposure of deprotecting/detritylation agents, adenine and guanine bases often undergo degradation by hydrolysis, leaving behind only the ribose sugar backbone. The presence of the 5’-OH-presenting sugar allows the oligonucleotide chain to elongate further; however, during the final side-group removal step (typically using ammonium hydroxide) these apurinic bases are cleaved and thereby result in truncated products.
DNA synthesis on planar surfaces often uses a modified version of the four-step phosphoramidite chemistry where a certain step is gated in order to provide spatial control of the individual oligonucleotides being synthesized. This is typically the coupling step for inkjet printing-based synthesis (Agilent, Protogene), or the deblocking step for light-based (LC Sciences, Invitrogen, Affymetrix) and redox reaction-based platforms (Combimatrix, Oxamer). In any case, truncated products could arise from misalignment of printed droplets or from partial deblocking as a result of poor light-source registration and improper sequestering of redox ions. Erroneous products caused by such “edge effects” could be mitigated by employing patterned substrates for synthesis [23, 24]. Studies using the inkjet chip synthesis platformdetermined that the error rate can be reduced from 1 error in ~200 bases to 1 error in ~600 bases by using patterned silica features on a plastic chip [25].
In addition to size exclusion purifications, strategies using the hybridization-selection principle have been used to reduce errors in microarray synthesized oligo pools. Error-containing oligos can be removed by stringent hybridization selections using short complementary oligos (selection oligos) immobilized on beads [7]. Gene construction oligos with errors form imperfect matches with the selection oligos and can be washed away under stringent washing conditions whereas those without errors can be retained and enriched. This strategy may be useful for cleaning up large pools of microarray-derived oligos but may not be convenient or economical for purifying small numbers of oligos due to the burden of synthesizing complementary selection oligos for all gene-construction oligos.
Without a separate pre-purification step, simply increasing hybridization stringency during gene assembly reaction helps prevent incorporation of erroneous oligos into the final assembly products [26]. This is because perfect hybridization among error-free sequences creates better templates for the polymerase or the ligase used in gene assembly. Ligation-based chain assembly (LCA) methods tend to benefit more from this effect as no gap is allowed between oligos. Polymerase-based cyclic assembly (PCA) methods allow the use of longer oligos and, as a result, the middle portion of the oligos are not subject to hybridization selection and tend to carry over more errors into assembled gene products [7, 26–28]. As it takes fewer long oligos to assemble a gene, there is an advantage of using long oligos provided that the sequence quality is satisfactory. High quality long oligos are available from commercial sources such as Integrated DNA Technologies (“Ultramers”, as long as 200 bp, Coralvill, IA, USA). Agilent Technologies (Santa Clara, CA, USA) has also recently started producing oligonucleotide libraries composed of >100bp sequences synthesized by their SurePrint microarray platform [22].
The rapid development of next-generation sequencing (NGS) technology has made it possible to sequence large pools of oligo sequences at affordable costs. This has triggered the temptation of selecting sequence-verified oligos as input for gene assembly. A proof-of-concept experiment has been performed which demonstrated that the so-called “megacloning” method can reduce error rates by a factor of 500 compared to the starting oligonucleotide pool generated by microarray [29]. In principle, with future development in platform automation, millions of oligos can be sequenced and sorted in a single megacloner run, which will potentially enable gene construction up to megabases [29].
Error-removal from synthetic genes
Despite exhaustive purification, errors that remained in synthetic oligos will be carried over during the assembly process and accumulated in downstream gene constructs. Additional errors may also be introduced by polymerase elongation or miss-hybridization among oligos [30]. Picking an error-free synthetic gene sequence often requires expensive and time consuming cloning and sequencing steps and some good luck. Besides cloning and sequencing, expression assays or functional screens can also be used to select desired products against errors which can cause reading frame shift and/or loss of protein functions [19, 31–34]. Nevertheless, this method is only selective to protein-coding regions or sequences encoding functional elements and is not effective for spotting silent or conservative mutations.
Longer DNA constructs (5-50kb) are usually assembled step-wise to achieve desired quality and efficiency. Short segments (<1kb) are first synthesized as building blocks, and their sequences verified individually by cloning and sequencing before being joined together into longer segments. Considering the error frequencies in starting synthetic oligonucleotides and the efficiency of cloning and sequencing, many current error-reducing strategies target this intermediate assembly stage in order to achieve maximum error-elimination efficiency through simple, robust and ideally automatable procedures [35, 36].
Most of the current error-removal techniques make use of DNA mismatch recognition agents. In contrast to mutation detection and correction in vivo [6, 37], synthetic DNA does not have chemical labels to distinguish between the “correct” and the “mutant” strands. Moreover, the PCR-mediated assembly process will copy the existing errors into the complementary strand. Therefore the likelihood of errors existing as mismatches in the assembled polynucleotide constructs is very small. It is then necessary to first randomly re-associate the polynucleotides through a denaturation and re-hybridization process during which erroneous bases form mismatches with the corresponding correct bases in the reverse-complementary strand. The bulging mismatch sites are then recognized and removed by using mismatch-binding proteins or mismatch-cleavage enzymes (Fig. 1).
Figure 1.
Summary of general schemes of mismatch-based error correction in synthetic DNA constructs. Assembly products are heat-denatured and then re-annealed to allow correct (blue lines) and mutant (red lines) strands to randomly re-hybridize and form mismatches. The mismatches are then removed by various error correction methods using either mismatch-binding proteins or mismatch-cleaving enzymes to yield products with correct sequences.
Error-removal using MutS mismatch-binding protein
The MutS protein is part of the bacterial MutHLS mismatch repair mechinary [38]. It detects and binds to a variety of mispaired bases and small single-strand loops in vivo. Methods have been developed that use MutS as an error-removal agent for gene synthesis. In one strategy, MutS protein from the thermophilic bacteria Thermus aquaticus was used to directly filter out full-length heteroduplexes [8]. After denaturation and reannealing, mismatch containing heteroduplexes are recognized and bound by the Taq MutS protein, which can be separated from the unbound homoduplexes (mostly error-free) by a gel-mobility shift assay (Fig. 2a). This method was reported to reduce error rate to 1 error per 3.8kb with one round of filtration, which represented several folds of improvement from raw synthesis products. Repeating the procedure one more time reportedly reduced the error rate to 1 error per 10 kb, which is good enough to ensure finding a correct sequence with a single round of cloning and sequencing. A major limitation with removing full-length heteroduplexes is that it requires a substantial fraction of the sequences to be error-free in order to survive the MutS binding filtration.
Figure 2.
Schematic of error removal strategies using MutS mismatch-binding protein. Assembled gene segments are denatured and re–annealed to form heteroduplexes where errors are revealed as mismatches between correct (blue) and mutant (red) sequences. a) MutS proteins recognize and bind to mismatch structures. Full-length heteroduplexes with MutS bound can be separated from homodimers without MutS by gel-shift assay. b) Consensus Shuffling. Re-annealed constructs are fragmented by type IIS restriction endonucleases. Fragments containing mismatches are filtered away by MutS binding column. Error-free fragments that pass through the column are collected and reassembled into full-length genes by assembly PCR.
For error-rich sequences, an alternative MutS-based error-correction strategy is called “consensus shuffling” [39, 40] (Fig. 2b). Re-annealed polynucleotide products are first cleaved by a mixture of restriction endonucleases and the resulting overlapping short fragments are then subjected to MutS column filtration. Short, mismatch-containing pieces are captured by immobilized Taq MutS, while error-free fragments are eluted and re-assembled into full-length constructs by assembly PCR . Two iterations of consensus shuffling improve the error rate 3.5- to 4.3-fold to 1 error per 3.5kb [39].
Consensus shuffling is a mechanism that resembles DNA shuffling where several iterations can be made until the consensus sequences are concentrated enough to far outnumber other species. Consensus shuffling provides several advantages over direct MutS filtration of full-length sequences: i) it is more effective on longer DNA constructs where perfect hybridization after re-assortment is rare; ii) it tolerates more errors in the starting material; iii) the removal of the entire DNA heteroduplex is not necessary. Importantly, fragmentation is ideally performed with type IIS restriction endonucleases that cut outside of the recognition sites. In this way, distinct ends can be generated by the same restriction endonuclease so as to minimize nonspecific hybridization of cut fragments after error-filtration. However, despite the merit of the idea, consensus shuffling requires extraneous use of endonucleases and the reaction conditions may still need to be improved to achieve better outcome [39]
Error-correction using mismatch-cleaving enzymes
Mismatch-cleaving enzymes here refer to a group of mismatch-specific endonucleases that recognize and cleave at or near mismatch sites in DNA heteroduplexes. The family embraces a variety of members including but not limited to: i) resolvases, such as phage T4 endonuclease VII [41], T7 endonuclease I [42] and E.coli endonuclease V [43]; ii) mismatch repair endonuclease MutH [38]; iii) single-strand specific nucleases, such as S1 nuclease from Aspergillus oryzae, P1 nuclease from Penicillium citrinum, mung bean nuclease, and CEL nuclease from celery [44, 45]. The ability of these endonucleases to cleave heteroduplex DNA at the mismatch sites has lead to their extensive applications as probes for mutation and polymorphism detection [44, 46–50]. The mismatch-cleaving activity also makes them very useful in error removal for gene synthesis and several different assays were developed utilizing this activity.
Resolvases
T7 endonuclase I was used in an early report:after assembly and amplification of a target DNA, the purified product was denatured, re-annealed and treated with T7 endonuclease I. Afterwards, the pool was run on a gel and the remaining intact band was excised and extracted, while the cleaved fragments were discarded [51]. This way of enriching error-free sequences seems to work for high-quality synthetic products where correct sequences out number mutants.
The efficacy of various resolvases was systematically explored to recognize and cleave single-base mismatches using a different error-removal assay, which is applicable for low-quality sequences. Three resolvases, T4 endonuclease VII, T7 endonuclease I and E.coli endonuclease V were tested respectively in the process of ligation-based gene synthesis and functional cloning of bacterial chloramphenicol-acetyltransferase (cat) gene. After treatment with T4 endo VII and E.coli endo V, the fraction of “functional clones” increased over 10-fold, and sequence analysis revealed an over 4-fold reduction of error-rate. T7 endonuclease I was not fully characterized in this setting [32]. Instead of using gel purification, it was demonstrated that following mismatch-cleavage, an additional 3’to 5’ exonuclease treatment to remove single-stranded overhangs was essential for error removal, either by adding E.coli exonuclease I or using the intrinsic 3’to 5’ exonuclease activity of proof-reading DNA polymerases (Fig. 3). The processed DNA fragments were reassembled into full-length products by assembly PCR.
Figure 3.
Schematic diagram showing the principal of error removal using mismatch-cleaving enzymes. The re-annealed heteroduplexes are cleaved in both strands by mismatch-specific enzymes 2–5bp downstream of the mismatches, generating stagger ends that are immediately dissociated at the reaction temperature. The single-stranded overhangs which contain mismatch base(s) are degraded by an exonuclease in the reaction mixture. Full-length genes with fewer errors are recovered by overlap assembly PCR[32].
Another method, called circular assembly amplification, combines exonuclease and endonuclease treatments to improve gene synthesis quality [52]. In this method, the overlapping and complementary oligonucleotides are first ligated into circular molecules under stringent conditions. An exonuclease is added to degrade any un-circularized polynucleotides. The circular DNA constructs are then treated with E. coli endonuclease V, which linearizes and degrades the mismatch-containing circles in concert with the exonuclease. By using the above three tiers of selections, the method is able to reduce error-rate by at least 7-fold compared to the conventional PCA assembly protocol [52].
A problem with phage resolvase-based mismatch cleavage is preferential cleavage of certain types of mis-matches. For instance, T7 Endonuclease I has been reported to fail to achieve quantitative cleavage for selected mis-pairs [32]. In addition, the degree of specificity of T4-endonuclease VII has been found to be highly dependent on the length of the substrate and the sequence surrounding the mismatch sites [53, 54]. This enzymology limitation is intrinsic to these resolvases because they naturally function to recognized holiday junctions rather than mismatches in a heteroduplex [55]. Their potency and efficacy of mismatch cleavage are therefore compromised by both low sensitivity and high undesired background. Future work to optimize their substrate specificity or use enzyme combinations may potentially produce better results.
MutHLS
E. coli proteins MutH, MutL and MutS constitute a bacterial mismatch repair mechanism: MutS detects and binds to mispaired bases and small single-strand loops, MutL couples the MutH endonuclease to the MutS bound mutant sites, leading to incision of the unmethylated strand of a hemi-methylated dsDNA. The error containing staggered ends are trimmed by a helicase and an exonuclease, and the resulting gap is filled by a DNA polymerase and a ligase [5]. MutHLS can be used to remove polymerase-produced mutant sequences in PCR products [38]. Treating re-annealed PCR products with MutHLS allows MutS to bind to the mismatches and create incisions on both unmethylated strands at d(GATC) in the vicinity of the mismatch sites. The cleaved heteroduplexes are then removed by gel electrophoresis. Mismatch removal using E. coli MutHLS reduces the error rate by one order of magnitude and that is effective to G-T, A-C, G-G, A-A and small insertion/deletion mis-pairs, while less effective for other type of mismatches [38].
Single-strand specific nucleases
Single-strand specific nucleases are a family of multifunctional enzymes that are present ubiquitously and display a wide variety of characteristics. Although more than 30 single-strand nucleases have been isolated from various sources, only a few have been thoroughly characterized. The most widely used single-strand nucleases are almost exclusively extracellular glycoproteins, including S1 and P1 nucleases from fungi, the mung bean nuclease and the CEL nuclease from plant. These nucleases vary drastically in the degree of specificity towards different types of mismatches. For instance, S1 and P1 nucleases appear to be strongly specific to AT-rich regions [55]; S1 nuclease seems incapable of recognizing single base mismatches [56]; and quantitative single nucleotide mismatch cleavage by mung bean nuclease occurs only at pH 6–6.5 [57]. Moreover, contradictory results have been reported regarding the use of S1, P1 and mung bean nucleases in detecting different types of single nucleotide polymorphism (SNP) in heteroduplex DNA [57]. The incomplete capability of these single strand nucleases to identify all possible types of mismatches greatly discourages their general use in synthetic gene error correction.
CEL endonuclease is an ortholog of the S1 nucleases isolated from celery [45, 58]. The uniqueness of CEL endonuclease is that it is the first eukaryotic endonuclease known to cleave DNA with high specificity towards all types of base mismatches and DNA distortions at neutral pH. Following the discovery of CEL nuclease, the techniques for high-throughput mutation detection have significantly improved, and CEL has become the most commonly used enzyme in TILLING (targeting-induced local lesions in genomes) [59–62], polymorphism analysis [63, 64] and disease diagnosis [58, 63–68].
CEL endonuclease is a mannosyl glycoprotein that nicks a DNA strand at the 3’end of the base-substitution mismatch and DNA distortion. Prolonged incubation and a high enzyme to substrate ratio facilities double-stranded cleavage by CEL at opposite phosphodiester bonds 3’ to the mismatch, generating i) two single-base 3’ overhangs in the case of base-substitutions or ii) a short stagger end in the case of insertion/deletion [45, 58]. The commercialization of CEL endonuclease as Surveyor™ nuclease by Transgenomic, Inc (Omaha, NE, USA) has led to its extensive applications in mutation detection as a simple yet cost-effective method, resulting in the discovery of many novel mutations associated with human genetic disorders, including BRCA1 [45, 58, 69], EGFR [65], HCDC4 [66], p53 [67], mitochondrial [63, 64] and kidney-related [68] genes, etc. The broad substrate specificity and low undesired activity of CEL nucleases make them the most promising candidates for error-correction in synthetic gene assemblies.
Recent studies employing on-chip gene synthesis technology explored the effectiveness of Surveyor™ nuclease in error correction of microchip synthesized genes [25, 70] (Fig. 4). The incorporation of a 60-minute digestion step with Surveyor™ nuclease followed by PCR assembly with Phusion DNA polymerase was able to reduce the error rate of the synthetic gene products from 1 error per 526 bp to 1 error per 3,883 bp. Two iterations of enzymatic error cleavage reactions further reduced the error rate to 1 error per 8,701bp, representing an over 16-fold error reduction [70]. Selective sub-pool amplification of chip-synthesized oligonucleotides achieve an error-rate of 1 error per 7,017bp after enzymatic error correction using ErrASE kit, a commercially available CEL-based enzyme cocktail that corrects mismatches [71]in a similar fashion to [25].
Figure 4.
Schematic diagram of error correction strategy using CEL mismatch-specific endonuclease. Multiple CEL error-correction cycles maybe integrated into a gene synthesis process. Each cycle consists of four steps: 1) Re-annealing of assembled gene constructs to present erroneous bases as mismatches; 2) CEL nuclease cleavage on both strands at the 3’ side of the mismatches; 3) Exonuclease trimming of single-stranded mismatch overhangs by added exonuclease or the 3’->5’ exonuclease activity of the proof-reading PCR enzyme; and 4) Reassembly and amplification of the processed fragments by assembly PCR. The final products are used for down-stream applications such as cloning and sequencing.
The combined actions of mismatch cleavage by CEL nuclease and proof-reading by Phusion DNA polymerase makes the procedure particularly robust and suitable for correcting error-rich synthetic sequences. Re-annealing of low-quality synthetic sequences often yield little error-free duplexes that can survive the mismatch-binding filtrationor mismatch-cleavage. Rather than discarding all error-containing heteroduplexes, CEL nuclease makes double-stranded cuts next to the mismatches. The mismatch bases are then chewed away by the exonuclease activity of the proof-reading DNA polymerase. The resulting error-free fragments are reassembled into full-length gene constructs by assembly PCR (Fig. 4). This method enables error elimination at the single-base level, while salvaging a majority of the correct portions of the sequences. More importantly, the “inspection-correction-reassemble” cycle can be iterated multiple times until products with desired purity are obtained.
Minimizing errors in genome synthesis
Identifying and correcting errors in synthetic genomes appears to be costly and painful due to the length of the sequences and difficulty of construct assembly. The key then is to use error-free starting materials (i.e. sequence-confirmed intermediate sequences) and avoid the introduction of errors during the genome assembly process. Single-step assembly using yeast homologous recombination could produce long, high quality DNA constructs with a very low error-rate (0.054%), comparable to many in vitro error-removal protocols as discussed previously. A microbial genome over 1 Mb has successfully been assembled using this strategy [1, 72]. This technique has also been applied to rapid assembly of metabolic pathways with multiple parts [73]. In addition to using overlapping dsDNA fragments as starting material, yeast recombination mechanism has recently been successfully applied to direct assembly of single-strand oligonucleotides into double-stranded gene constructs [74].
Conclusions and perspectives
Over the past half century, de novo chemical DNA synthesis and enzymatic gene and genome assembly techniques have been continuously improved. The longest published synthetic DNA has extended from less than 1kb to over 1Mb, which encoded functional genes, genetic pathways, and viral or bacterial genomes [1, 2, 75, 76]. Nevertheless, the gap between the expanding demand and the capability to deliver synthetic genes and genomes in accurate, economical and high-throughput fashion, is still wide. To close the gap, current efforts in automation and miniaturization of gene synthesis need to seamlessly incorporate convenient and effective error prevention and error removal strategies. The methods discussed in this review reflect current thinking and practice in error correction. There is still plenty of room for improvements and innovations. For example, the new high-fidelity DNA synthesis chemistry has yet to be put into general use, which can reduce error generation from the beginning. The error-correction enzymes can be further engineered to be more effective on correcting single-base mutations. The cloning and sequencing process used to select correct sequences should be eliminated or simplified and automated using advanced NGS technology. Ultimately, what the field is looking for is a simple, effective and automatable method that ensures maximum quality of DNA synthesis at minimum cost. Concerted efforts combining more efficient DNA synthesis chemistry, more effective error-correction enzymes, and more powerful NGS sequencing technology seems to be the way to go. With continued innovation, this goal is expected to be accomplished in the near future, which will significantly benefit the synthetic biology and biotechnology fields.
Acknowledgments
This work was supported by an NIH grant (R01HG005862) and by grant from the Chinese Academy of Sciences. JT was a Beckman Young Investigator and a recipient of The Hartwell Foundation Individual Biomedical Research Award.
Footnotes
Conflict of Interest
The authors declare no conflict of interest.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Gibson DG, et al. Creation of a bacterial cell controlled by a chemically synthesized genome. Science. 2010;329:52–56. doi: 10.1126/science.1190719. [DOI] [PubMed] [Google Scholar]
- 2.Tian J, et al. Advancing high-throughput gene synthesis technology. Molecular bioSystems. 2009;5:714–722. doi: 10.1039/b822268c. [DOI] [PubMed] [Google Scholar]
- 3.Czar MJ, et al. Gene synthesis demystified. Trends Biotechnol. 2009;27:63–72. doi: 10.1016/j.tibtech.2008.10.007. [DOI] [PubMed] [Google Scholar]
- 4.Kunkel TA, Erie DA. DNA mismatch repair. Annu Rev Biochem. 2005;74:681–710. doi: 10.1146/annurev.biochem.74.082803.133243. [DOI] [PubMed] [Google Scholar]
- 5.Schofield MJ, Hsieh P. DNA mismatch repair: molecular mechanisms and biological function. Annu Rev Microbiol. 2003;57:579–608. doi: 10.1146/annurev.micro.57.030502.090847. [DOI] [PubMed] [Google Scholar]
- 6.Li GM. Mechanisms and functions of DNA mismatch repair. Cell Res. 2008;18:85–98. doi: 10.1038/cr.2007.115. [DOI] [PubMed] [Google Scholar]
- 7.Tian J, et al. Accurate multiplex gene synthesis from programmable DNA microchips. Nature. 2004;432:1050–1054. doi: 10.1038/nature03151. [DOI] [PubMed] [Google Scholar]
- 8.Carr PA, et al. Protein-mediated error correction for de novo DNA synthesis. Nucleic acids research. 2004;32:e162. doi: 10.1093/nar/gnh160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Hoover DM, Lubkowski J. DNAWorks: an automated method for designing oligonucleotides for PCR-based gene synthesis. Nucleic Acids Res. 2002;30:e43. doi: 10.1093/nar/30.10.e43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xiong AS, et al. A simple, rapid, high-fidelity and cost-effective PCR-based two-step DNA synthesis method for long gene sequences. Nucleic Acids Res. 2004;32:e98. doi: 10.1093/nar/gnh094. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Caruthers MH. Gene synthesis machines: DNA chemistry and its uses. Science. 1985;230:281–285. doi: 10.1126/science.3863253. [DOI] [PubMed] [Google Scholar]
- 12.Beaucage SL, Caruthers MH. Deoxynucleoside phosphoramidites - a new class of key intermediates for deoxypolynucleotide synthesis. Tetrahedron Letters. 1981;22:1859–1862. [Google Scholar]
- 13.Caruthers MH, et al. Chemical synthesis of deoxyoligonucleotides by the phosphoramidite method. Methods Enzymol. 1987;154:287–313. doi: 10.1016/0076-6879(87)54081-2. [DOI] [PubMed] [Google Scholar]
- 14.Eckstein F. Oligonucleotides and analogues : a practical approach. IRL Press; 1991. [Google Scholar]
- 15.Ellington A, Pollard JD., Jr Introduction to the synthesis and purification of oligonucleotides. Curr Protoc Nucleic Acid Chem. 2001 doi: 10.1002/0471142700.nca03cs00. Appendix 3, Appendix 3C. [DOI] [PubMed] [Google Scholar]
- 16.Zhou X, et al. Microfluidic PicoArray synthesis of oligodeoxynucleotides and simultaneous assembling of multiple DNA sequences. Nucleic Acids Res. 2004;32:5409–5417. doi: 10.1093/nar/gkh879. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Andrus A, Kuimelis RG. Analysis and purification of synthetic nucleic acids using HPLC. Curr Protoc Nucleic Acid Chem. 2001;Chapter 10(Unit 10):15. doi: 10.1002/0471142700.nc1005s01. [DOI] [PubMed] [Google Scholar]
- 18.Cello J, et al. Chemical synthesis of poliovirus cDNA: generation of infectious virus in the absence of natural template. Science. 2002;297:1016–1018. doi: 10.1126/science.1072266. [DOI] [PubMed] [Google Scholar]
- 19.Smith HO, et al. Generating a synthetic genome by whole genome assembly: phiX174 bacteriophage from synthetic oligonucleotides. Proc Natl Acad Sci U S A. 2003;100:15440–15445. doi: 10.1073/pnas.2237126100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Xiong AS, et al. PCR-based accurate synthesis of long DNA sequences. Nature Protocols. 2006;1:791–797. doi: 10.1038/nprot.2006.103. [DOI] [PubMed] [Google Scholar]
- 21.Sierzchala AB, et al. Solid-phase oligodeoxynucleotide synthesis: a two-step cycle using peroxy anion deprotection. J Am Chem Soc. 2003;125:13427–13441. doi: 10.1021/ja030376n. [DOI] [PubMed] [Google Scholar]
- 22.LeProust EM, et al. Synthesis of high-quality libraries of long (150mer) oligonucleotides by a novel depurination controlled process. Nucleic Acids Research. 2010;38:2522–2540. doi: 10.1093/nar/gkq163. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Saaem I, et al. In situ Synthesis of DNA Microarray on Functionalized Cyclic Olefin Copolymer Substrate. Acs Appl Mater Inter. 2010;2:491–497. doi: 10.1021/am900884b. [DOI] [PubMed] [Google Scholar]
- 24.Chow BY, et al. Photoelectrochemical synthesis of DNA microarrays. Proceedings of the National Academy of Sciences. 2009;106:15219–15224. doi: 10.1073/pnas.0813011106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Quan J, et al. Parallel on-chip gene synthesis and application to optimization of protein expression. Nat Biotechnol. 2011;29:449–452. doi: 10.1038/nbt.1847. [DOI] [PubMed] [Google Scholar]
- 26.Borovkov AY, et al. High-quality gene assembly directly from unpurified mixtures of microarray-synthesized oligonucleotides. Nucleic acids research. 2010;38:e180. doi: 10.1093/nar/gkq677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Smith HO, et al. Generating a synthetic genome by whole genome assembly: phi X174 bacteriophage from synthetic oligonucleotides. P Natl Acad Sci USA. 2003;100:15440–15445. doi: 10.1073/pnas.2237126100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gao X, et al. Thermodynamically balanced inside - out (TBIO) PCR–based gene synthesis: a novel method of primer design for high–fidelity assembly of longer gene sequences. Nucleic Acids Research. 2003;31:e143. doi: 10.1093/nar/gng143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Matzas M, et al. High-fidelity gene synthesis by retrieval of sequence-verified DNA identified using high-throughput pyrosequencing. Nat Biotechnol. 2010;28:1291–1294. doi: 10.1038/nbt.1710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Cline J, et al. PCR fidelity of pfu DNA polymerase and other thermostable DNA polymerases. Nucleic Acids Res. 1996;24:3546–3551. doi: 10.1093/nar/24.18.3546. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Stemmer WP, et al. Single-step assembly of a gene and entire plasmid from large numbers of oligodeoxyribonucleotides. Gene. 1995;164:49–53. doi: 10.1016/0378-1119(95)00511-4. [DOI] [PubMed] [Google Scholar]
- 32.Fuhrmann M, et al. Removal of mismatched bases from synthetic genes by enzymatic mismatch cleavage. Nucleic Acids Research. 2005;33 doi: 10.1093/nar/gni058. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Cox JC, et al. Protein fabrication automation. Protein Sci. 2007;16:379–390. doi: 10.1110/ps.062591607. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Kim H, et al. A Fluorescence Selection Method for Accurate Large-Gene Synthesis. ChemBioChem. 2010;11:2448–2452. doi: 10.1002/cbic.201000368. [DOI] [PubMed] [Google Scholar]
- 35.Kodumal SJ, et al. Total synthesis of long DNA sequences: Synthesis of a contiguous 32-kb polyketide synthase gene cluster. Proceedings of the National Academy of Sciences of the United States of America. 2004;101:15573–15578. doi: 10.1073/pnas.0406911101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Reisinger SJ, et al. Total synthesis of multi-kilobase DNA sequences from oligonucleotides. Nat Protocols. 2007;1:2596–2603. doi: 10.1038/nprot.2006.426. [DOI] [PubMed] [Google Scholar]
- 37.Modrich P. Mechanisms and biological effects of mismatch repair. Annu Rev Genet. 1991;25:229–253. doi: 10.1146/annurev.ge.25.120191.001305. [DOI] [PubMed] [Google Scholar]
- 38.Smith J, Modrich P. Removal of polymerase-produced mutant sequences from PCR products. Proc Natl Acad Sci U S A. 1997;94:6847–6850. doi: 10.1073/pnas.94.13.6847. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Binkowski BF, et al. Correcting errors in synthetic DNA through consensus shuffling. Nucleic acids research. 2005;33:e55. doi: 10.1093/nar/gni053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Binkowski BF, et al. Correcting errors in synthetic DNA through consensus shuffling. Nucleic Acids Research. 2005;33:e55. doi: 10.1093/nar/gni053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Kemper B, Garabett M. Studies on T4-head maturation. 1. Purification and characterization of gene-49-controlled endonuclease. European journal of biochemistry / FEBS. 1981;115:123–131. [PubMed] [Google Scholar]
- 42.West SC. Enzymes and molecular mechanisms of genetic recombination. Annu Rev Biochem. 1992;61:603–640. doi: 10.1146/annurev.bi.61.070192.003131. [DOI] [PubMed] [Google Scholar]
- 43.Yao M, Kow YW. Cleavage of insertion/deletion mismatches, flap and pseudo-Y DNA structures by deoxyinosine 3'-endonuclease from Escherichia coli. J Biol Chem. 1996;271:30672–30676. doi: 10.1074/jbc.271.48.30672. [DOI] [PubMed] [Google Scholar]
- 44.Desai NA, Shankar V. Single-strand-specific nucleases. FEMS Microbiol Rev. 2003;26:457–491. doi: 10.1111/j.1574-6976.2003.tb00626.x. [DOI] [PubMed] [Google Scholar]
- 45.Yang B, et al. Purification, cloning, and characterization of the CEL I nuclease. Biochemistry. 2000;39:3533–3541. doi: 10.1021/bi992376z. [DOI] [PubMed] [Google Scholar]
- 46.Youil R, et al. Enzymatic mutation detection (EMD) of novel mutations (R565X and R1523X) in the FBN1 gene of patients with Marfan syndrome using T4 endonuclease VII. Human Mutation. 2000;16:92–93. doi: 10.1002/1098-1004(200007)16:1<92::AID-HUMU24>3.0.CO;2-1. [DOI] [PubMed] [Google Scholar]
- 47.Babon JJ, et al. The use of resolvases T4 endonuclease VII and T7 endonuclease I in mutation detection. Molecular Biotechnology. 2003;23:73–81. doi: 10.1385/MB:23:1:73. [DOI] [PubMed] [Google Scholar]
- 48.Qiu P, et al. Mutation detection using Surveyor nuclease. Biotechniques. 2004;36:702–707. doi: 10.2144/04364PF01. [DOI] [PubMed] [Google Scholar]
- 49.Pimkin M, et al. Recombinant nucleases CEL I from celery and SP I from spinach for mutation detection. BMC Biotechnol. 2007;7:29. doi: 10.1186/1472-6750-7-29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Triques K, et al. Mutation detection using ENDO1: application to disease diagnostics in humans and TILLING and Eco-TILLING in plants. BMC Mol Biol. 2008;9:42. doi: 10.1186/1471-2199-9-42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Young L, Dong Q. Two-step total gene synthesis method. Nucleic acids research. 2004;32:e59. doi: 10.1093/nar/gnh058. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Bang D, Church GM. Gene synthesis by circular assembly amplification. Nat Methods. 2008;5:37–39. doi: 10.1038/nmeth1136. [DOI] [PubMed] [Google Scholar]
- 53.Babon JJ, et al. Mutation detection using fluorescent enzyme mismatch cleavage with T4 endonuclease VII. Electrophoresis. 1999;20:1162–1170. doi: 10.1002/(SICI)1522-2683(19990101)20:6<1162::AID-ELPS1162>3.0.CO;2-Y. [DOI] [PubMed] [Google Scholar]
- 54.Norberg T, et al. Enzymatic mutation detection method evaluated for detection of p53 mutations in cDNA from breast cancers. Clin Chem. 2001;47:821–828. [PubMed] [Google Scholar]
- 55.Yeung AT, et al. Enzymatic mutation detection technologies. Biotechniques. 2005;38:749–758. doi: 10.2144/05385RV01. [DOI] [PubMed] [Google Scholar]
- 56.Silber JR, Loeb LA. S1 nuclease does not cleave DNA at single-base mismatches. Biochim Biophys Acta. 1981;656:256–264. doi: 10.1016/0005-2787(81)90094-0. [DOI] [PubMed] [Google Scholar]
- 57.Till BJ, et al. Mismatch cleavage by single-strand specific nucleases. Nucleic Acids Res. 2004;32:2632–2641. doi: 10.1093/nar/gkh599. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Oleykowski CA, et al. Mutation detection using a novel plant endonuclease. Nucleic Acids Research. 1998;26:4597–4602. doi: 10.1093/nar/26.20.4597. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Till BJ, et al. Large-scale discovery of induced point mutations with high-throughput TILLING. Genome Res. 2003;13:524–530. doi: 10.1101/gr.977903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Colbert T, et al. High-throughput screening for induced point mutations. Plant Physiol. 2001;126:480–484. doi: 10.1104/pp.126.2.480. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Comai L, et al. Efficient discovery of DNA polymorphisms in natural populations by Ecotilling. Plant J. 2004;37:778–786. doi: 10.1111/j.0960-7412.2003.01999.x. [DOI] [PubMed] [Google Scholar]
- 62.Perry JA, et al. A TILLING reverse genetics tool and a web-accessible collection of mutants of the legume Lotus japonicus. Plant Physiol. 2003;131:866–871. doi: 10.1104/pp.102.017384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Bannwarth S, et al. Rapid identification of unknown heteroplasmic mutations across the entire human mitochondrial genome with mismatch-specific Surveyor Nuclease. Nature Protocols. 2006;1:2037–2047. doi: 10.1038/nprot.2006.318. [DOI] [PubMed] [Google Scholar]
- 64.Bannwarth S, et al. Rapid identification of mitochondrial DNA (mtDNA) mutations in neuromuscular disorders by using surveyor strategy. Mitochondrion. 2008;8:136–145. doi: 10.1016/j.mito.2007.10.008. [DOI] [PubMed] [Google Scholar]
- 65.Janne PA, et al. A rapid and sensitive enzymatic method for epidermal growth factor receptor mutation screening. Clin Cancer Res. 2006;12:751–758. doi: 10.1158/1078-0432.CCR-05-2047. [DOI] [PubMed] [Google Scholar]
- 66.Nowak D, et al. Mutation analysis of hCDC4 in AML cells identifies a new intronic polymorphism. Int J Med Sci. 2006;3:148–151. doi: 10.7150/ijms.3.148. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Mitani N, et al. Surveyor nuclease-based detection of p53 gene mutations in haematological malignancy. Ann Clin Biochem. 2007;44:557–559. doi: 10.1258/000456307782268174. [DOI] [PubMed] [Google Scholar]
- 68.Voskarides K, Deltas C. Screening for mutations in kidney-related genes using SURVEYOR nuclease for cleavage at heteroduplex mismatches. J Mol Diagn. 2009;11:311–318. doi: 10.2353/jmoldx.2009.080144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Kulinski J, et al. CEL I enzymatic mutation detection assay. Biotechniques. 2000;29:44–46. 48. doi: 10.2144/00291bm07. [DOI] [PubMed] [Google Scholar]
- 70.Saaem I, et al. Error correction of microchip synthesized genes using Surveyor nuclease. Nucleic acids research. 2011 doi: 10.1093/nar/gkr887. (In print) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kosuri S, et al. Scalable gene synthesis by selective amplification of DNA pools from high-fidelity microchips. Nat Biotechnol. 2010;28:1295–1299. doi: 10.1038/nbt.1716. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Gibson DG, et al. Complete chemical synthesis, assembly, and cloning of a Mycoplasma genitalium genome. Science. 2008;319:1215–1220. doi: 10.1126/science.1151721. [DOI] [PubMed] [Google Scholar]
- 73.Shao Z, Zhao H. DNA assembler, an in vivo genetic method for rapid construction of biochemical pathways. Nucleic Acids Res. 2009;37:e16. doi: 10.1093/nar/gkn991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Gibson DG. Synthesis of DNA fragments in yeast by one-step assembly of overlapping oligonucleotides. Nucleic Acids Res. 2009;37:6984–6990. doi: 10.1093/nar/gkp687. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Wimmer E, et al. Synthetic viruses: a new opportunity to understand and prevent viral disease. Nat Biotechnol. 2009;27:1163–1172. doi: 10.1038/nbt.1593. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Carr PA, Church GM. Genome engineering. Nat Biotechnol. 2009;27:1151–1162. doi: 10.1038/nbt.1590. [DOI] [PubMed] [Google Scholar]




