Skip to main content
Nucleic Acids Research logoLink to Nucleic Acids Research
. 2026 Aug 10;54(15):gkag775. doi: 10.1093/nar/gkag775

A self-iterative orthogonal base-editing platform enables multiplex N-to-N diversification and genome-scale functional screening in Escherichia coli

Xiangrui Fan 1,c, Liya Liang 2,✉,c, Hongle Wang 3, Guangning Liu 4, Huiping Tan 5, Fa Zhang 6, Lin Liu 7, Rongming Liu 8,9,
PMCID: PMC13454841  PMID: 42573069

Abstract

Base editing enables precise genome modification without double-strand breaks but remains limited by narrow editing windows, DNA repair pathway biases, and restricted nucleotide diversity. Here, we report MUTATOR, a MUlTiplexAble and self-iTerative ORthogonal base-editing platform that enables N-to-N diversification in Escherichia coli. MUTATOR combines CWBE and ABE with iterative editing on two complementary DNA strands, thereby overcoming endogenous DNA repair constraints and expanding A-to-N and C-to-N editing outcomes across both strands. This strategy substantially expands accessible nucleotide outcomes, codon variants, and amino-acid diversity within existing editing windows relative to conventional editors. Using four gRNAs, MUTATOR facilitated four-site editing of ompR, generating 84 distinct amino-acid combinations and 252 codon combinations, with the synonymous OmpR_P160P variant increasing isobutanol production by up to 56.2%. We further applied MUTATOR to a 151-gene library encompassing transcriptional regulators, translation factors, DNA repair proteins, ribosomal components, and NAD(P)H-associated metabolic genes, identifying single and combinatorial mutations that markedly enhanced cell growth and ethanol utilization when ethanol was used as the sole carbon source. Together, these results establish MUTATOR as a broadly applicable platform for genome-wide diversification, functional dissection, and rapid engineering of industrial microbial chassis.

Graphical Abstract

Graphical Abstract.

For image description, please refer to the figure legend and surrounding text.

Introduction

Genetic modifications are essential for phenotypic diversification and functional optimization in living systems, underpinning both natural evolution and the engineering of microbial systems for industrial biomanufacturing [1]. Advanced genome engineering platforms such as Multiplex Automated Genome Engineering (MAGE) [2, 3], CRISPR-enabled trackable genome engineering (CREATE) [4], and CRISPR–Cas9 and homology-directed repair-assisted saturation editing (CHASE) [5] have greatly facilitated the functional interrogation of genetic elements and accelerated microbial strain improvement, enabling the generation of large collections of amino-acid-level variants in vivo.

Despite these advances, most genome engineering strategies continue to rely on programmable DNA double-strand breaks (DSBs) to stimulate homologous recombination [6, 7]. DSB-dependent editing imposes intrinsic limitations on multi-locus genome engineering [6, 8, 9]. Excessive DSBs provoke strong DNA-damage responses across species. In bacteria, they trigger SOS-mediated apoptosis-like death [1012], while in eukaryotes they activate DNA-damage checkpoints and apoptosis [13]. These responses markedly reduce cell viability and limit the scalability of iterative genome diversification. Moreover, high-throughput CRISPR–Cas-mediated gene editing often requires the introduction of synonymous mutations at protospacer adjacent motif (PAM) or seed regions to prevent re-cutting [4, 5], and even these “silent” substitutions have been shown to affect gene expression and phenotype, further complicating functional interpretation in DSB-based systems and interfering with perturbation designs at nucleotide and codon resolution [1416].

DSB-free base editors (BEs) [17] were developed as an alternative to DSB-dependent recombination, offering a means to install precise nucleotide substitutions without triggering DNA-damage responses. However, canonical adenine base editor (ABE) and cytosine base editor (CBE) remain largely restricted to transition mutations such as C-to-T and A-to-G, granting access to only a narrow portion of the mutational landscape [1824]. Engineering efforts have produced transversion-type editors, including the C-to-G base editor (CGBE), the adenine Y-base editor (AYBE, Y = C or T), and the adenine X-base editor (AXBE, X = any nucleotide) [2529], which operate through apurinic/apyrimidinic (AP) intermediates that must be resolved by translesion synthesis (TLS) polymerases [30, 31]. In bacteria, engineering of TLS-associated repair pathways has also enabled controllable C-to-A and C-to-G editing outcomes, highlighting the importance of AP-site processing in shaping bacterial base-editing spectra [26, 32]. Recent studies have also developed bacterial dual-base editors and combinatorial multi-base editing systems that further expand editable nucleotide outcomes and multiplex editing capabilities [3335]. However, these existing approaches remain constrained by predefined editing chemistries and/or endogenous repair pathways and do not provide a unified strategy for accessing all four nucleotide states at the same genomic position. As a result, these processes yield strongly skewed and incomplete mutational spectra that fall well short of saturating the accessible outcome space [36]. Such mechanistic constraints fundamentally limit the ability of existing transversion-type editors to generate balanced and comprehensive C-to-N or A-to-N mutational outcomes in Escherichia coli (E. coli).

More critically, current base editors cannot achieve N-to-N using a single system. Since existing base editors are intrinsically base-specific, with ABEs acting only on adenines and CBEs acting only on cytosines, no single editor can access all four nucleotide states at the same genomic position (Fig. 1A). In addition, the narrow, position-specific editing windows of most systems further limit which nucleotides within a target region can be edited. Collectively, these mechanistic constraints leave substantial fractions of the local sequence space inaccessible [36], preventing combinatorial saturation even within short programmable windows [37], and hindering unbiased reconstruction of sequence–function landscapes required for engineering multigenic traits.

Figure 1.

For image description, please refer to the figure legend and surrounding text.

Characterization of base editors in E. coli. (A) Overview of MUTATOR-enabled N-to-N mutational space and its applications in multi-site and genome-scale editing. Blue arrows represent substitutions achievable by both MUTATOR and existing editors, whereas pink arrows represent mutations unique to MUTATOR. (B) Editing efficiencies of CBE at target positions 4–8 immediately after recovery (0 h) and following 16 h of continuous induction. (C) Editing efficiencies of ABE at positions 3–8 under the same induction conditions as in panel (B). (D) Editing efficiencies of CWBEmini, CWBE-v1, and CWBE-v2 at positions 4–8. (E) Base substitution frequencies (C-to-A versus C-to-T) mediated by CWBEmini at positions 4–8. Each bar represents n = 3 biological replicates, and error bars denote the mean ± standard deviation.

In this study, we report MUTATOR, a multiplexable and self-iterative orthogonal base-editing platform that enables N-to-N diversification in E. coli (Fig. 1A). MUTATOR combines a C-to-W base editor (CWBE, W = A or T) and an ABE with iterative editing on both DNA strands, thereby overcoming endogenous repair-driven outcome biases and expanding both A-to-N and C-to-N mutation spectra. This design expands the mutational outcome space within defined editing windows, thereby increasing accessible nucleotide outcomes, codon variants, and amino-acid diversity compared with conventional base editors. These capabilities markedly broaden the accessible sequence–function landscape, providing a versatile framework for scalable genome diversification, high-resolution functional interrogation, and accelerated evolution of industrial microbial chassis.

Materials and methods

Strains and plasmids used in the MUTATOR method

Escherichia coli DH5α was used as the cloning host. Escherichia coliMG1655 was employed for genome editing. Strains were cultivated in LB broth at 30℃ or 37℃ .When appropriate, antibiotics were supplemented at the following concentrations: chloramphenicol (34 μg/ml), kanamycin (50 μg/ml), and ampicillin (100 μg/ml). Expression of the base editor was induced with 10 μg/l anhydrotetracycline when required.

Spacer replacement in gRNA plasmids was performed using a pre-constructed counter-selection plasmid pDYBE-nfsI-scaffold, whose design is shown in Supplementary Fig. S1. A complete list of bacterial strains and plasmids is provided in Supplementary Tables S1 and S2, and all primers and oligonucleotides are listed in Supplementary Table S3.

Evaluation of base editing efficiency in E. coli

To assess editing efficiency, ABE, CBE, or CWBEmini plasmids either individually or in combination were transformed into E. coli MG1655. Spacer sequences of gRNAs used in this study are listed in Supplementary Table S4.

Electrocompetent cells were prepared as follows. Overnight cultures were inoculated into 50 ml of fresh LB medium containing 0.2% arabinose, 34 μg/ml chloramphenicol, and 50 μg/ml kanamycin and incubated at 30°C and 200 rpm. When OD600 reached 0.5–0.6, the cultures were chilled on ice for 15–30 min. Then, the cells were washed twice with ice-cold 10% glycerol at 50% of the initial culture volume. After two rounds of centrifugation and supernatant removal, cells were resuspended to 1% of the initial volume, and 50 µl aliquots were distributed into electroporation cuvettes (0.1 cm gap) and electroporated with 100 ng of the respective editing plasmids at 1800 V. Immediately after electroporation, cells were resuspended in 950 µl of SOB medium and recovered at 30°C with shaking at 200 rpm for 2 h.

After recovery, cultures were either directly plated onto LB agar supplemented with the corresponding antibiotic and anhydrotetracycline (10 μg/l) or inoculated into 10 ml of LB medium containing the same selection agents and incubated at 30°C and 200 rpm for 16 h prior to plating.

Editing efficiency was evaluated by two methods depending on the editor combinations: (i) colony color screening on MacConkey agar plates, or (ii) polymerase chain reaction (PCR) amplification of the genomic regions flanking the target sites using specific primers, followed by next-generation sequencing (NGS) to quantify editing rates.

Construction of the OmpR library

EC41 strain [7] harboring both ABE and CWBEmini plasmids was cultured until OD600 reached 0.5–0.6. Cells were chilled on ice for 15–30 min, then washed twice with ice-cold 10% (v/v) glycerol, using 50% of the initial culture volume for each wash. After two rounds of centrifugation and supernatant removal, the cells were resuspended in ice-cold 10% glycerol to 1% of the original culture volume. Aliquots of 200 µl were dispensed into 0.2-cm-gap electroporation cuvettes.

Library plasmids were added to the electrocompetent cells, and the mixtures were incubated on ice for 5 min. Electroporation was performed at 2500 V. Immediately afterward, cells were recovered in SOB medium and incubated at 30°C with shaking at 200 rpm for 2 h.

Recovered cultures were transferred into LB medium (10 ml) containing 50 μg/ml kanamycin, 100 μg/ml carbenicillin, 50 μg/ml spectinomycin, and 10 μg/ml anhydrotetracycline. After 16 h incubation, cells were harvested by centrifugation and resuspended in fresh LB or selective media for downstream target screening.

Screening assay and fermentation validation of OmpR variants

Following base editing using the library strains, cultures were incubated overnight at 37°C to promote plasmid curing and eliminate ABE and CWBEmini editing plasmids.

Cured strains were rendered electrocompetent and subsequently transformed with the isobutanol biosensor plasmid pSelect-1 [7]. After overnight growth, cultures were diluted to an OD600 of 0.1, and selection was initiated by adding tetracycline at a final concentration of 14 μg/l. High-throughput screening was performed in 48-well plates using EZ Rich medium at 30°C for 12 h to identify variants that responded to increased isobutanol production.

Isobutanol production was assessed following a two-stage fermentation protocol. Strains were cultured in M9 minimal medium supplemented with 30 g/l glucose, 0.5 g/l NaCl, 17.1 g/l Na2HPO4·12H2O, 3 g/l KH2PO4, 2 g/l NH4Cl, 246 mg/l MgSO4·7H2O, 14.7 mg/l CaCl2·2H2O, and 2.78 mg/l FeSO4·7H2O, with the addition of 10 g/l yeast extract. The growth was performed in 50 ml conical tubes containing 25 ml of media. Cultures were first incubated at 30°C with shaking at 200 rpm for 6 h, followed by a second 6-h incubation at 30°C and 150 rpm. Fermentation was terminated at the 12-h timepoint, and isobutanol titers were subsequently measured.

Glucose and isobutanol were analyzed by HPLC (Agilent 1260 Infinity equipped with a refractive index detector, Agilent Technologies, Santa Clara, USA) with a Bio-Rad Aminex HPX-87H column at 65°C. The mobile phase was 5 mM H2SO4 at a flow rate of 0.6 ml/min. All the samples were centrifuged at 15 800 × g for 6 min, and then filtered through a 0.22-μm filter before analysis.

Library construction for ethanol-utilization strain selection

Using plasmid pDYBE-nfsI-scaffold as the template, the vector fragment was amplified with primers bb-library-new-F/bb-library-new-R under the following reaction conditions: 98°C for 30 s; then 30 cycles of 98°C for 10 s, 61°C for 5 s, and 72°C for 15 s; followed by 72°C for 1 min. Subsequently, gel electrophoresis was performed to confirm the correctness of the bands, and gel extraction was carried out.

According to the designed library plasmid, the gRNA region was amplified using the In-library-new-F/In-library-new-R primer pair. After that, gel electrophoresis was conducted to verify the correct bands, and gel extraction was performed using the Hipure Gel Pure DNA Mini Kit (Magen Biotechnology, Guangzhou, China). The resulting fragment served as the insert.

Finally, ligation was performed using the Golden Gate method, and the resulting plasmids were library plasmids. For the Golden Gate reaction, an assembly premix was prepared containing 0.05 pmol of vector, a three-fold excess of insert, NEB Golden Gate Enzyme Mix, and T4 DNA Ligase Buffer (10×). The reaction program was as follows: 37°C for 1 h; 60°C for 5 min. After ligation was completed, 10 µl of the ligation product was transformed into E. coli DH5α competent cells and spread on plates containing 50 μg/ml spectinomycin to estimate the transformation efficiency and library coverage.

Selection assay for ethanol-utilizing strains

The constructed library plasmids were introduced into the ES64 strain (Supplementary Table S1). After editing using the MUTATOR, the pCDF-adhEMut plasmid [38] was transformed into the strains thereby enabling E. coli to grow using ethanol as the sole carbon source. Selection assays were performed under different ethanol concentrations.

When transferring the bacterial culture to the ethanol screening medium, the initial OD600 of the bacterial culture in the screening medium was controlled to be 0.1. Specifically, after measuring the OD600 of the bacterial culture, the appropriate volume of the bacterial culture was calculated. The culture was centrifuged at 5000 rpm for 1 min. Then, the supernatant was removed, and the bacterial pellet was resuspended in sterile ddH2O before being transferred to the screening medium. The culture was incubated in a shaker at 30°C with 200 rpm, and the OD600 was measured every 24 h to observe the growth trend of the strains. After selection, the library showing a clear growth advantage (higher OD600 relative to the control) at high ethanol stress was harvested by centrifugation. The resulting cell pellets were subjected to genomic DNA extraction for subsequent mutation profiling and enrichment analysis.

Construction of high-enrichment gRNA designs

gRNA libraries before and after ethanol selection were extracted and subjected to NGS to identify variants with high enrichment scores. gRNA designs that ranked highest in both enrichment magnitude and post-selection read abundance were selected for reconstruction.

For single-gRNA plasmid reconstruction, spacer replacement was performed using the pDYBE-nfsI-scaffold vector as the backbone. Primers carrying homology arms flanking the spacer region were used to amplify the entire plasmid. PCR products were purified (HiPure Gel Pure DNA Mini Kit, Magen Biotechnology, Guangzhou, China) and assembled using Gibson Assembly. The assembly mixture was electroporated into E. coli MG1655, recovered in SOB medium, and plated on LB agar supplemented with 0.4% (w/v) metronidazole and 50 μg/ml spectinomycin. More than 90% of the resulting colonies contained correctly assembled gRNA plasmids.

Dual-gRNA plasmids were constructed by Golden Gate Assembly using the corresponding reconstructed single-gRNA plasmids as templates. For each dual-gRNA design, the first and second gRNA cassettes were amplified with primer pairs IN-dual_gRNA-F/IN-dual_gRNA-R and bb-dual_gRNA-F/bb-dual_gRNA-R, respectively. PCR products were gel-purified and assembled with the destination vector using NEB Golden Gate Enzyme Mix (BsmB I-v2). The reaction was incubated at 37°C for 1 h followed by 60°C for 5 min. A total of 10 µl of the reaction mixture was transformed into E. coli DH5α and plated on spectinomycin-containing LB agar for selection.

High-throughput multi-locus genotype tracing using OE-PCR coupled with digital PCR (dPCR)

We designed multiplex primer sets using the primer design tool [39]. The target strain was cultured overnight to obtain a high-density bacterial suspension. The culture density was adjusted based on OD600 measurements to ensure that each dPCR reaction chamber contained fewer than ∼5 × 104 cells.

PCR system components: 5 µl of 5× PerfeCTa Multiplex qPCR SuperMix; 5 µl of Probes luciferin sodium salt solution (10 μM); 1 µl of diluted bacterial suspension; 2 µl of primers (10 μM); and ddH2O was added to make up the total volume of the system to 25 µl.

Reaction program: 40°C for 5 min (droplet formation); 95°C for 10 min; then 30 cycles of 95°C for 10 s, 55°C for 5 s, and 72°C for 60 s; followed by 72°C for 1 min; and finally 40°C for 30 min to complete the program.

Droplet recovery: The instrument recovery program “Sapphire Protocol Droplet Recover” was used to recover the droplets. Droplets from each reaction chamber were transferred to microtubes and allowed to stand for 2–3 min to enable phase separation. After the droplets and oil were separated into layers, all the upper aqueous phase was gently transferred to a new EP tube, and the lower oil phase was discarded. Then, 60 µl of TE buffer or deionized water was added, followed by 210 µl of chloroform. The mixture was vortexed for ~1 min and centrifuged at 15 500 × g for 10 min. The upper aqueous phase (~20–30 µl) was transferred to a new EP tube, and gel electrophoresis was subsequently performed. The band position was determined according to the marker position, and gel extraction was carried out. Gel-purified amplicons were re-amplified with barcoded primers and submitted for NGS to resolve multi-locus genotypes.

High-throughput DNA sequencing and data analysis

Genomic DNA was extracted from mixed E. coli populations subjected to base editing. Target loci were amplified using barcode-indexed primers (Supplementary Table S3), and PCR products were verified by electrophoresis on 2% agarose gels. Gel-purified amplicons were submitted to GENEWIZ for paired-end sequencing (2 × 150 bp) using the Illumina HiSeq platform.

Raw sequencing data were processed using custom Python scripts. Paired-end reads were merged using VSEARCH, followed by quality filtering with a minimum Phred score threshold of Q30. Demultiplexing was performed based on barcode sequences. Downstream analyses were customized according to the experimental requirements. The full analysis pipeline and source code are available at https://github.com/Liurmlab/NGS.git.

High-throughput DNA sequencing and data analysis were performed. The sequencing data were classified according to the barcodes at both ends, and the generated sequences were ranked by read counts in the specified editing regions. Enrichment values represent the log2 ratio of the proportion of reads relative to total reads post-selection to that pre-selection for each mutant.

The enrichment score (Ej) was calculated as follows:

graphic file with name TM0001.gif

where Xj is the proportion of mutant j pre-selection in the deep-sequencing measurement, and Yj is the proportion of mutant j post-selection in the deep-sequencing measurement.

Reconstruction of the top enriched mutants

Since some mutations involve changes in synonymous codons, the direct construction using the traditional CRISPR–Cas system may interfere with the functional interpretation of target mutations due to the need to introduce additional synonymous mutations in the PAM region. To avoid this issue, we adopted a two-step strain construction strategy. First, an ampicillin resistance cassette (AmpR) was inserted within or adjacent to the target gene as a replacement site. In the second step, a gRNA was designed to target the AmpR fragment, and homologous recombination using donor templates containing the desired mutations enabled precise introduction of the selected variants into the target locus without introducing additional synonymous PAM-disrupting changes.

Validation of top enriched mutants from ethanol selection

To validate the functional contribution of the top enriched variants identified from ethanol selection, representative synonymous, nonsynonymous, and combinatorial mutations were individually reconstructed and subjected to growth assays in ethanol-based minimal medium.

For liquid-culture validation, reconstructed mutants were inoculated into M9 minimal medium supplemented with 10 or 40 g/l ethanol, with an initial OD600 = 0.2, and incubated at 30°C and 200 rpm. Cell growth was monitored by measuring OD600 every 24 h.

For plate-based ethanol tolerance assays, reconstructed mutants were serially diluted and spotted onto M9 minimal medium plates supplemented with 40 g/l ethanol. Liquid cultures were normalized to an OD600 of 4 prior to dilution. Serial dilutions (10−4–10−7) were prepared, and 5-µl aliquots of each dilution were spotted onto agar plates, followed by incubation at 30°C.

Results

Optimizing the base editor for high-efficiency and multiplex mutagenesis in E. coli

To establish an efficient foundation for high-efficiency genome editing, we first optimized the performance of SpCas9-NG-based CBE and ABE constructed with the high-activity deaminases evoFERNY and TadA-8e_V106W [21, 22, 40]. Prolonged editing duration markedly improved editing efficiencies at the edges of the canonical editing windows for each editor, with positions 4–8 for CBE and 3–8 for ABE showing the greatest gains, while the central positions remained consistently high regardless of induction duration (Fig. 1B and C). Additional time-course experiments performed in the absence of aTc revealed a similar temporal increase in editing efficiency, indicating that prolonged exposure time substantially contributes to the accumulation of editing outcomes (Supplementary Fig. S2). These results define the maximal editable window achievable under optimized editing conditions in E. coli.

To characterize C-to-N editing performance, we next explored the development of the CGBE by designing three variants, including a uracil-DNA glycosylase (UNG)-free design, CWBEmini (evoFERNY-nCas9), and two UNG-containing variants, CWBE-v1 (UNG-evoFERNY-nCas9) and CWBE-v2 (evoFERNY-nCas9-UNG) (Supplementary Fig. S3). Among all variants, CWBEmini exhibited the highest overall activity, with strong editing at positions 5–7 and markedly reduced activity at positions 4 and 8 (Fig. 1D). In E. coli, CGBE activity was dominated by C-to-A and C-to-T substitutions, whereas C-to-G products were detected only at very low frequencies, which differs from what has been reported in eukaryotic systems (Fig. 1E). Incorporating UNG, either at the N- or C-terminus, did not alter these outcomes, indicating that enhanced uracil excision does not substantially influence the repair trajectory in bacteria. The editing result analysis further revealed that C-to-T conversions dominated across most positions, particularly at positions 4 and 7, where they accounted for 68.0% and 58.6% of edits (Fig. 1E), respectively. In comparison, C-to-A conversions were less frequent at these positions, representing 31.4% and 40.2% of edits (Fig. 1E). At positions 5, 6, and 8, the editing outcomes were more balanced, with C-to-A and C-to-T frequencies falling within a similar range (~45%–55%) (Fig. 1E). These results establish CWBEmini as the most effective C-to-T/A editor in E. coli. This bacteria-specific outcome bias highlights fundamental differences between bacterial and eukaryotic AP/TLS pathways and underscores the necessity of tailoring C-to-N base editors for prokaryotic systems.

We further evaluated a panel of AXBE-based editors, including AXBE-alkA, AYBE, AXBE, and their corresponding dCas9-containing variants (dAXBE-alkA and dAYBE), which were designed to expand the range of adenine-editing outcomes [27, 28, 41, 42] (Supplementary Fig. S4A). However, the results demonstrated that the editing outcomes for these AXBE and AYBE variants largely mirrored those of ABE, predominantly resulting in A-to-G substitutions, with minimal detection of other base conversions (Supplementary Fig. S4B and C). For instance, all the AXBE variants, as well as ABE, showed a strong preference for converting adenine (A) to guanine (G), with frequencies exceeding 90%. Minor proportions of other base changes, such as A-to-T and A-to-C conversions, were observed but accounted for <5% of the total edits (Supplementary Fig. S4B and C). These results, consistent with the limited AP-site processing observed for CWBE variants, suggest that bacterial AP/TLS pathways do not efficiently support diversified adenine transversions.

Given that AXBE variants did not provide advantages or novel outcomes beyond ABE, all subsequent analyses focused on ABE, CBE, and CWBEmini. Together, these results identify the most effective adenine and cytosine editors in E. coli and define the mechanistic constraints that limit editing diversity. These insights established the basis for designing the MUTATOR platform in the following sections.

Design of a self-iterative base editing system for C-to-N and A-to-N

Modified editors such as CWBEmini and AXBE were unable to generate comprehensive C-to-N or A-to-N outcomes in E. coli, as their edits remained strongly restricted to C-to-T/A or A-to-G substitutions. To expand mutational breadth beyond these intrinsic substrate biases and the resulting outcome limitations, we developed MUTATOR, a self-iterative platform capable of driving sequential and complementary base conversions at the same genomic position. Accordingly, we designed a configuration in which CWBEmini and ABE operate iteratively, allowing one round of base conversion to generate the appropriate substrate for the next (Fig. 2A).

Figure 2.

For image description, please refer to the figure legend and surrounding text.

Orthogonal editing systems enable C-to-N and A-to-N mutagenesis in E. coli. (A) Schematic illustration of orthogonal C-to-N (left) and A-to-N (right) editing using CWBE and ABE. Phase I: targeting by individual editors; Phase II: co-targeting with orthogonal systems; Phase III: diverse base conversion outcomes (C-to-A/T/G and A-to-G/T/C). (B) Comparison of cytosine mutation outcomes between CWBEmini alone and the orthogonal MUTATOR system (CWBEmini + ABE). Left: base substitution frequencies; Right: representative mutation spectra. (C) Comparison of adenine mutation outcomes between ABE alone and MUTATOR (CWBEmini + ABE). Left: base substitution frequencies; Right: representative mutation spectra. Each bar represents n = 3 biological replicates, and error bars denote the mean ± standard deviation.

For C-to-N diversification, CWBEmini was used to generate C-to-A/T alleles, followed by ABE-mediated A-to-G editing at the same locus. This strategy requires that the system recognize both the wild-type C-containing target and the edited A-containing allele, so we implemented two gRNA configurations: a matched-pair (two-gRNA) design, in which one gRNA is fully matched to the wild-type sequence and the second carries a single A mismatch to recognize the edited allele (Supplementary Fig. S5A), and a single-gRNA design that is fully matched to the wild-type locus and relies on Cas9’s mismatch tolerance to target edited alleles (Supplementary Fig. S5B). Surprisingly, two-gRNA and single-gRNA configurations yielded nearly identical editing efficiencies (Fig. 2B and Supplementary Fig. S5C), demonstrating substantial gRNA–DNA mismatch tolerance. NGS confirmed that the sequential action of CWBEmini followed by ABE markedly increased both the diversity and the overall frequency of C-to-N outcomes compared with CWBEmini alone (Fig. 2B). Furthermore, even after CWBEmini converted the native CAA motif into alleles such as CGG or TAA, the original gRNA continued to guide CWBEmini or ABE despite carrying one or two mismatches (Supplementary Fig. S6A and B). This level of mismatch tolerance eliminates the need to redesign gRNAs after each editing step and allows sequential rounds of editing to proceed using the same gRNA.

We next examined whether this platform could support A-to-N diversification. ABE-mediated A-to-G substitution forms a deamination endpoint on a single DNA strand, because the resulting G cannot be further edited within standard base-editing windows. To bypass this limitation, we implemented a dual-strand design in which ABE targets adenine on one strand while CWBEmini simultaneously edits the complementary cytosine on the opposite strand (Fig. 2A). Co-expression of both editors and two gRNAs targeting complementary strands did not impair growth or editing activity (Supplementary Fig. S7A and B), confirming that dual-strand targeting is compatible with efficient base editing. Similar to the C-to-N system, both editors remained active even when gRNA–DNA mismatches were introduced by earlier editing events, generating overlapping and sequential edits within the same genomic window (Fig. 2C). High-throughput sequencing revealed that non-G products accounted for ~20% of total A-editing outcomes (Fig. 2C), representing a substantial increase relative to ABE alone, which generates nearly exclusive A-to-G substitutions.

Collectively, these results establish a unified self-iterative base-editing platform in which CWBEmini and ABE function orthogonally through their complementary substrate specificities and coordinated activities on opposite DNA strands. By exploiting one round of base conversion to create substrates for subsequent editing events, this platform circumvents the intrinsic base-specificity limitations of conventional editors and mitigates editing outcome biases imposed by endogenous repair pathways in E. coli. As a result, the accessible single-nucleotide mutational space is expanded far beyond what is achievable with any individual base editor, providing a robust foundation for comprehensive base diversification, high-throughput mutagenesis, and directed evolution in bacteria.

Orthogonal dual-editor platform enables multi-locus N-to-N mutagenesis

Having established iterative C-to-N and A-to-N pathways at single loci, we next applied MUTATOR to achieve N-to-N diversification in E. coli (Fig. 3A). In this configuration, CWBEmini and ABE act on complementary DNA strands and sequentially generate substrates for one another, allowing all four nucleotide states to be accessed at the same genomic position without requiring redesign of gRNAs after intermediate editing events.

Figure 3.

For image description, please refer to the figure legend and surrounding text.

N-to-N substitutions within the editing window mediated by MUTATOR. (A) Substitution profiles obtained at the galK-206 site using the CAGG editing window. For CWBEmini (left), the two gRNAs were tested independently, and the resulting single-gRNA editing outcomes were summed to generate the overall substitution profile. For MUTATOR (right), the two gRNAs were used simultaneously, enabling coordinated editing on complementary strands and thereby producing combined substitution outcomes. (B) Codon and amino acid diversity achievable by CWBEmini and MUTATOR within the editing window at the galK-206 site. The circular plot shows the number of codons and resulting amino acids generated by base editing. Orange indicates substitutions specific to MUTATOR, blue denotes CWBEmini-specific substitutions, and green represents shared substitutions. (C) Substitution profiles obtained at the galK-20 site using the ACAA editing window. For CWBEmini (left), the two gRNAs were tested independently, and the resulting single-gRNA editing outcomes were summed to generate the overall substitution profile. For MUTATOR (right), the two gRNAs were used simultaneously, enabling coordinated editing on complementary strands and thereby producing combined substitution outcomes. (D) Codon and amino acid diversity achievable by CWBEmini and MUTATOR within the editing window at the galK-20 site. The circular plot shows the number of codons and resulting amino acids generated by base editing. Orange indicates substitutions specific to MUTATOR, blue denotes CWBEmini-specific substitutions, and green represents shared substitutions.

To assess N-to-N mutagenesis under genomic contexts, we selected two genomic positions within galK, one in the internal coding region and the other near the 3′ terminus, and designed pairs of gRNAs that shared the same programmed editing window for each site. MUTATOR was benchmarked against CWBEmini using matched gRNA architectures. At the galK-206 site, MUTATOR generated 71 distinct single-nucleotide substitutions across the 5–8 nt window, representing a 1.9-fold improvement in variant diversity relative to CWBEmini (Fig. 3A). At the amino-acid level, MUTATOR produced 31 codons encoding 17 amino acids plus stop codons, substantially exceeding the diversity obtained with CWBEmini (Fig. 3B).

At the less favorable galK-20 locus, MUTATOR generated 41 nucleotide-based mutations, whereas CWBEmini produced only 17, corresponding to a 141% improvement (Fig. 3C). Consistently, MUTATOR yielded 24 codons and 13 amino-acid variants, compared with 11 codons and 8 amino-acid variants from CWBEmini (Fig. 3D). Although a few low-frequency variants appeared unique to CWBEmini, nearly all of these were present within the MUTATOR mutational space but fell below the detection threshold due to the substantially broader diversity generated by MUTATOR (Supplementary Fig. S8A and B).

We further evaluated potential off-target editing by performing targeted deep sequencing of the top predicted off-target sites identified using Cas-OFFinder [43] (three to five mismatches) for the gRNA pairs targeting the galK-206 and galK-20 loci (Supplementary Fig. S9 and Supplementary Table S5). Across all analyzed sites, the vast majority of reads corresponded to wild-type sequences. Although a small number of low-frequency variant reads were detected, these variants occurred sporadically and lacked reproducible editing signatures or consistent mutation patterns across biological replicates (Supplementary Fig. S9). These observations suggest that MUTATOR maintains high target specificity under the conditions tested.

Collectively, these results demonstrate that the orthogonal, self-iterative design of MUTATOR markedly expands both the depth and breadth of base- and amino-acid-level diversity within defined genomic windows. By integrating complementary deaminase activities across both DNA strands, MUTATOR enables N-to-N mutagenesis in E. coli and provides a scalable platform for comprehensive sequence diversification.

MUTATOR-enabled multi-site functional diversification of OmpR

To evaluate whether the MUTATOR platform can be extended across separate chromosomal regions, we prepared a schematic in which distinct genomic loci are assigned independent programmable editing windows (Fig. 4A) and applied this strategy to diversify the global transcriptional regulator OmpR. Previous structural and functional studies have shown that specific residues within OmpR modulate DNA-binding dynamics and regulatory output, thereby influencing complex phenotypes including isobutanol tolerance and production [7]. Guided by these findings, we selected two functional regions encompassing Met159, Pro160, Leu180, and Ser181, and each segment was targeted with two gRNAs (Fig. 4B).

Figure 4.

For image description, please refer to the figure legend and surrounding text.

MUTATOR-enabled multi-site diversification of OmpR for improved isobutanol production. (A) Schematic overview of MUTATOR-enabled multi-site editing. Multiple target sites are programmed for coordinated diversification, enabling simultaneous editing across distinct loci. (B) Structural visualization of key residues in the OmpR protein targeted for mutagenesis. Critical sites Met159, Pro160, Leu180, and Ser181 are highlighted. (C) Codon usage and amino acid diversity generated by MUTATOR in two critical regions of ompR using four gRNAs. The heatmap shows the number of codons resulting in different amino acid substitutions at each targeted position. (D) Combinatorial mutation landscape of ompR library. Variants fixed at position 180 as Leu are shown together with combinatorial substitutions at other positions, with read counts indicated by color intensity. (E) Enrichment analysis of the ompR mutation library. Enrichment values represent the log2 ratio of post-selection to pre-selection read count for each variant. High-performing variants are highlighted. (F) Reconstruction and fermentation validation of top-performing variants identified in (E). Each bar represents n = 3 biological replicates, and error bars denote the mean ± standard deviation.

High-throughput profiling of the edited libraries revealed extensive amino-acid diversification across all four programmed editing windows. At the single-site level, MUTATOR generated broad substitution spectra at each residue (Fig. 4C). Pro160 exhibited the greatest mutational plasticity, producing 16 amino-acid identities encoded by one to five distinct codons. Even the least variable positions, such as Met159 and Ser181, each yielded 10 or more amino-acid substitutions, demonstrating that MUTATOR effectively accesses deep mutational space at both targeted regions. These results confirm that each programmed site can be diversified into a broad range of codon and amino-acid states, substantially surpassing the capabilities of conventional single-editor systems.

At the combinatorial level, MUTATOR extensively sampled multi-residue sequence space spanning all four targeted positions. By quantifying joint amino-acid distributions, we reconstructed the combinatorial mutation landscape across Met159, Pro160, Leu180, and Ser181 (Fig. 4D). When fixing Leu180 as a representative example, MUTATOR produced numerous multi-residue configurations distributed across a wide span of read depths. The resulting density patterns illustrate nonuniform yet extensive exploration of multi-site sequence space, reflecting positional interdependencies and editor-specific activity preferences. In total, MUTATOR generated 84 amino-acid combinations and 252 codon variants across these four residues.

Biosensor-based screening under sub-inhibitory tetracycline selection enriched diverse OmpR variants (Supplementary Fig. S10). After NGS validation (Fig. 4E), representative variants spanning synonymous substitutions, single-residue changes accompanied by synonymous variation, double-residue substitutions, and combinatorial variants were reconstructed. Fermentation assays demonstrated that all reconstructed strains exhibited significantly elevated isobutanol titers relative to the control strain (Fig. 4F). The strongest improvement was observed for the synonymous variant OmpR_P160P, which increased isobutanol production to 3.92 g/l from 2.51 g/l in the control strain, corresponding to a 56.2% increase. This finding is consistent with prior evidence that synonymous changes modulate translation kinetics, co-translational folding and regulatory protein function in E. coli [14, 44, 45]. These results demonstrate that MUTATOR enables high-resolution diversification of multiple functional residues within a key transcriptional regulator, providing access to mutational configurations not reachable through traditional genome engineering approaches.

Genome-scale identification of ethanol-utilization mutant using the mutator platform

Ethanol is not a native carbon source for E. coli, and efficient ethanol assimilation typically relies on introducing ethanol-metabolizing pathways and modifying key regulators such as CRP to improve ethanol tolerance [35, 46, 47]. These modifications enable initial growth on ethanol but remain insufficient for industrially relevant conditions, where high ethanol concentrations and strong redox perturbations severely limit cellular performance. Since ethanol can be biologically produced from C1 feedstocks including methanol and syngas, improving microbial ethanol utilization has substantial value for C1-based biomanufacturing [48]. Thus, further enhancement of ethanol assimilation is expected to require identifying beneficial mutations in genes spanning regulatory, metabolic, and stress-response functions, whether as single-site changes or as specific combinatorial variants, rather than relying on alterations within a single regulatory pathway. To define the physiological limits of the parental strain, we evaluated its growth across a gradient of ethanol concentrations. The strain exhibited the highest biomass formation at 10 g/l ethanol but showed sharply reduced growth at 20 g/l and near-complete inhibition at 40–50 g/l (Supplementary Fig. S11). These measurements guided our selection of 20 and 40 g/l as moderate and severe stress conditions for subsequent genetic screens.

The multi-locus editing capability of the MUTATOR platform is highly suited for this challenge, as it enables semi-rational construction of combinatorial mutation libraries targeting key regulatory nodes (Fig. 5A). To evaluate whether MUTATOR could uncover mutations that enhance ethanol tolerance and utilization, we assembled a comprehensive mutation library comprising 151 genes associated with global regulation, transcriptional control, translation regulation, DNA repair, ribosomal function, and NAD(P)H metabolism (Supplementary Table S6). A total of 4597 gRNAs were partitioned into five libraries (Supplementary Table S6). Among these, 76 transcriptional regulators were consolidated into the S1 core library and served as the backbone for generating multiplex perturbation libraries, in which gRNAs from the remaining libraries were incorporated in either single or paired configurations. This platform ensured representation of both single-site and multi-site perturbations, enabling systematic interrogation of regulatory and metabolic networks.

Figure 5.

For image description, please refer to the figure legend and surrounding text.

Genome-scale identification of ethanol-utilization mutants using the MUTATOR platform. (A) Design–Build–Test–Learn (DBTL) workflow for MUTATOR-driven combinatorial genome engineering. (B) Growth of the five gRNA libraries under 20 and 40 g/l ethanol stress. S123 denotes a library containing gRNAs from S1, S2, and S3 (similarly for S124, S125, S134, and S145). (C) NGS-based enrichment analysis of gRNA designs after ethanol selection. Yellow bars indicate enriched single-gRNA designs, and blue bars indicate enriched dual-gRNA designs. (D) Secondary screening of a focused MUTATOR library, showing OD600 values after cultivation under 40 g/l ethanol. Strains exhibiting improved growth were selected for mutation tracing and further reconstruction. (E) Structural mapping of representative enriched mutations in CRP, RpoB, SoxR, and RecB. Highlighted residues correspond to sites identified by MUTATOR as beneficial under ethanol stress. (F) Growth validation of reconstructed variants in shake-flask cultures containing 40 g/l ethanol. Data represent n = 3 biological replicates, with error bars indicating mean ± standard deviation.

Ethanol selection was performed at 20 and 40 g/l to impose moderate and high metabolic stress (Fig. 5B). Substantial differences in biomass accumulation were observed among the libraries. Library S134 and S145 consistently exhibited the strongest growth under both ethanol concentrations, suggesting that beneficial mutations were enriched in these two pools (Fig. 5B). NGS confirmed significant enrichment of single and double gRNA designs (Fig. 5C). In S134, single-gRNA designs targeting rpoA (5.55-fold enrichment) and crp (5.20-fold enrichment), which originated from different library groups, together with the dual-gRNA combination targeting soxR and recB (4.80-fold enrichment), showed clear enrichment after selection (Supplementary Tables S7 and S8). In S145, gRNAs targeting hns, pntA, and tufA, each derived from distinct library groups, exhibited enrichment levels of 4.49-, 4.35-, and 3.85-fold, respectively, indicating that mutations in these loci contributed to the improved ethanol-associated phenotypes (Supplementary Table S9).

To further refine the mutation landscape, we selected gRNAs that ranked the highest in both enrichment fold and post-selection read abundance, and assembled a focused mutation library using the MUTATOR platform. Secondary screening was then performed using 20 and 40 g/l ethanol (Fig. 5D). For single gRNA designs, we analyzed genomic mutational profiles at targeted sites and compared mutation abundances before and after selection to identify variants that conferred growth advantages (Supplementary Fig. S12 and Supplementary Tables S10–12). For double gRNA designs, we established a high-throughput multi-locus tracing workflow that couples large-scale Overlap Extension PCR (OE-PCR) with a digital PCR platform (Supplementary Fig. S13). This approach enables distal genomic loci to be assembled into a single amplicon from each colony, allowing efficient detection of co-occurring mutations and precise tracking of multi-site editing events at single-cell resolution (Supplementary Table S13).

Structural mapping of enriched substitutions revealed clear mechanistic patterns (Fig. 5E and Supplementary Fig. S14). Enriched CRP substitutions including L188L, K189K, V184, G185, and M190 were located near the α-helix DNA-binding domain, suggesting altered DNA-binding dynamics or changes in allosteric coupling. Enriched RpoB variants such as A676A, D675D, N677X, and A679V localized to the β-subunit clamp-channel region involved in transcription elongation and stress-responsive gene regulation. For SoxR, the co-occurring substitutions L102P and D103N are located immediately adjacent to the Dithiothreitol (DTT) redox cofactor, indicating that simultaneous perturbation of this local structural motif may modify redox sensing and signal transduction. For the cross-locus combination involving SoxR_L102P and RecB_V777I, the SoxR substitution lies near the DTT center whereas RecB_V777I maps to a helicase–nuclease interface that is essential for DNA end processing under genotoxic stress. The joint localization patterns of these combinatorial mutations suggest coordinated modulation of redox-responsive regulation and DNA repair capacity during ethanol challenge.

To validate the functional relevance of enriched mutations, representative variants including synonymous, nonsynonymous, and combinatorial mutations were reconstructed and evaluated in shake-flask cultures in M9 media with 10 and 40 g/l ethanol (Fig. 5F and Supplementary Fig. S15). Under 10 g/l ethanol, RpoB_A679V, Crp_L188L and K189K, and SoxR_L102P and D103N markedly improved biomass accumulation relative to the non-targeting control (Supplementary Fig. S15). These mutations increased specific growth rates by >70%, with Crp_L188L and K189K reaching an OD600 of 1.55 at 24 h and exhibiting a 79% increase in growth rate compared with the control. Preliminary tests confirmed that 10 g/l ethanol does not inhibit growth of the wild-type strain, suggesting that these mutations accelerate ethanol assimilation rather than relieve toxicity. At 40 g/l ethanol, all 11 reconstructed high-enrichment variants outperformed the control strain (Fig. 5F). After 72 h, the OD600 of high-enrichment variants increased 2.1- to 5.5-fold relative to WT. The synonymous RpoB_A676A mutation, a GCC-to-GCA codon substitution, demonstrates that even silent mutations can reshape global transcriptional and metabolic activity in ways that alter ethanol utilization.

The combinatorial mutation SoxR_L102P and RecB_V777I produced a substantial synergistic effect (Fig. 5F). After 72 h, its OD600 reached 0.895, representing a 2.1-fold improvement compared with the single-site SoxR_L102P mutation and a 2.8-fold improvement compared with the control. These results indicate that combinatorial mutations can provide additional benefits over single-site perturbations and highlight the importance of multi-locus editing for reshaping stress response networks. Consistent with liquid-culture assays, spot tests on M9 minimal plates supplemented with 40 g/l ethanol (Supplementary Fig. S16) further verified the superior stress tolerance of reconstructed mutants. Serial dilution spotting showed markedly enhanced survival compared with the non-targeting control, reinforcing the robustness of the mutations identified through the MUTATOR screening pipeline. Together, these findings demonstrate that the MUTATOR platform can leverage efficient multi-site editing to construct semi-rational combinatorial libraries targeting global regulators, enabling systematic rewiring of metabolic and transcriptional networks. This capability allows the identification of mutations that enhance both ethanol tolerance and ethanol utilization in E. coli, providing a powerful strategy for engineering strains capable of metabolizing ethanol derived from C1-based feedstocks.

Discussion

In this study, we developed MUTATOR, a high-throughput platform for multi-locus genome engineering in E. coli that enables diverse base-substitution patterns without requiring homologous repair templates. By integrating self-iterative editing dynamics with an orthogonal dual-editor configuration, MUTATOR achieves complementary A-to-N and C-to-N conversions that together provide access to all four nucleotide states at the same genomic position. This capability bypasses the substrate restrictions and repair-dependent outcome biases that fundamentally constrain existing E. coli base editors. As a result, MUTATOR establishes a generalizable platform for efficient, template-free base editing across multiple genomic loci and markedly expands the accessible mutational space in E. coli.

Building on this expanded outcome space, MUTATOR overcomes long-standing mechanistic limitations of transversion-type base editors in E. coli [4951]. Whereas CGBE, AXBE, and related systems can generate biased subsets of C-to-N or A-to-N substitutions in eukaryotic cells [52, 53], their activity in E. coli is constrained by asymmetric AP-site processing and inefficient TLS engagement, resulting in highly skewed mutational spectra and limited saturation potential [26]. These mechanistic constraints prevent existing editors from sampling balanced transversion outcomes or accessing all four nucleotide states. Consequently, no single bacterial base editor can mediate interconversion across nucleotide classes or provide full N-to-N accessibility at an individual genomic position. Recent bacterial dual-base editors and multi-base editing systems have expanded simultaneous or combinatorial editing capabilities, but they still rely on predefined editor chemistries and do not provide unified access to all four nucleotide states at the same genomic position [34, 35]. In contrast, the self-iterative configuration combining CWBEmini with ABE bypasses these repair bottlenecks by leveraging sequential and strand-complementary editing. This strategy yields enriched noncanonical outcomes, including C-to-G and A-to-C/T substitutions, that are effectively inaccessible in current E. coli base-editing systems and yields more uniform C-to-N profiles while simultaneously expanding the accessible A-to-N mutational space. Thus, MUTATOR enables a qualitatively distinct mutational regime, expanding the reach of base editing beyond the mechanistic boundaries imposed by endogenous repair pathways.

Compared with high-throughput genome engineering strategies such as MAGE [3, 6] and CREATE [4, 7], MUTATOR offers higher efficiencies, concurrent multi-locus diversification, and simpler design requirements. These features make the system well suited for large-scale strain engineering. To demonstrate its utility, we applied MUTATOR to the ompR regulatory gene, which influences isobutanol biosynthesis [7]. Four gRNAs targeting two functional regions generated extensive sequence diversity, resulting in 252 codon variants and 84 amino acid combinations across four residues. Screening identified a synonymous P160 codon change that improved isobutanol production by 56.2%. Such beneficial variants are rarely detected by conventional genome editing methods. These approaches often depend on introducing synonymous mutations to disrupt the PAM and prevent re-cutting. However, synonymous changes can alter codon usage or local mRNA structure, which may unintentionally affect translation kinetics and obscure the functional contribution of the targeted residue [14, 15].

To evaluate MUTATOR’s ability to dissect complex multigenic phenotypes, we applied the system to a genome-scale screen targeting 151 genes spanning regulatory, metabolic, and stress-response pathways under ethanol selection. Ethanol is an increasingly important C1-derived feedstock for biomanufacturing, yet its assimilation imposes strong redox and transcriptional burdens, making it an informative model for testing multi-locus adaptive rewiring. Multi-locus perturbation libraries revealed strong enrichment of both single-site and combinatorial variants, indicating that adaptive ethanol utilization requires distributed rewiring across transcriptional control, redox management, and DNA-repair networks. Beneficial substitutions in CRP, RpoA, and RpoB clustered near domains associated with allosteric regulation or transcriptional elongation. Synergistic combinations such as the SoxR_L102P and D103N double-site variant further supported by cross-locus interactions involving SoxR and RecB, reveal non-additive effects that are inaccessible to single-locus approaches. These findings establish MUTATOR as an effective platform for systematically uncovering epistatic and cross-network interactions that shape microbial stress responses and carbon assimilation.

MUTATOR defines a conceptually distinct paradigm for bacterial genome engineering. It bridges the gap between classical single-base editors, which are limited by narrow mutational spectra, and genome-scale recombineering platforms that require extensive design and template synthesis. The ability to install, iterate, and combine substitutions across multiple loci without template redesign positions MUTATOR as a powerful tool for functional genomics, directed evolution, and the rapid optimization of industrial microbial chassis. Beyond coding sequences, MUTATOR could also be applied to regulatory elements such as promoters, 5′ untranslated regions, and ribosome binding sites. Previous BETTER-based strategies used cytosine base editing to generate combinatorial regulatory libraries for expression tuning [54]. Because MUTATOR expands nucleotide outcome diversity beyond CBE-only systems, it may enable finer and broader diversification of regulatory sequences for synthetic biology and metabolic engineering applications.

Although extending the system to organisms with more restrictive repair environments may require rebalancing deaminase activity or modulating mismatch tolerance, the modularity of the architecture readily supports such adaptations. Furthermore, emerging transversion-type editors such as TBE and gGBE could in principle be incorporated into this framework [29, 55], offering additional routes to diversify mutational spectra as these tools mature. Notably, MUTATOR expands mutational outcome diversity without broadening the physical editing window of the underlying base editors. Future integration with editors possessing expanded editing windows or altered activity profiles may further increase the number of editable codons accessible per gRNA. Coupling MUTATOR with machine-learning-guided gRNA design [5658], biosensor-coupled selection, and continuous evolution schemes may further enhance its capacity for autonomous navigation of high-dimensional fitness landscapes. In addition, the current implementation uses SpCas9-NG, which expands targeting scope beyond canonical NGG PAMs but still imposes PAM-dependent constraints. Incorporating PAM-relaxed or near-PAMless Cas variants, such as SpRY [59], may further expand the targetable codon space of MUTATOR, although potential trade-offs in editing efficiency and specificity will need to be carefully evaluated.

Together, these results demonstrate that MUTATOR enables comprehensive and scalable exploration of sequence–function landscapes in E. coli, offering a versatile platform for engineering emergent phenotypes such as stress tolerance, metabolic rewiring, and carbon utilization. The modular architecture of MUTATOR may facilitate future adaptation to additional prokaryotic chassis following empirical optimization and validation. As synthetic biology increasingly demands precise, high-throughput, and multi-locus genome-editing capabilities, MUTATOR provides a robust technological foundation for the next generation of microbial strain development.

Supplementary Material

gkag775_Supplemental_File

Acknowledgements

We thank Dr Dongdong Yang from University of Colorado Boulder for generously providing the FERNY-CBE4max plasmid. We thank Prof. Hui Wu from Dalian University of Technology for generously providing the pCDF-adhEMut plasmid. We thank Core Facilities of School of Bioengineering, Dalian University of Technology for their assistance with digital PCR. 

Author contributions: Conceptualization: R.L., L. Liang, and X.F. Methodology: R.L., L. Liang, X.F., and H.W. Investigation: R.L., L. Liang, X.F., H.W., G.L., and L.L. Visualization: H.T., F.Z., and H.W. Supervision: R.L. and L. Liang. Writing—original draft: X.F., H.W., and R.L. Writing—review & editing: R.L. and L. Liang.

Contributor Information

Xiangrui Fan, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Liya Liang, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Hongle Wang, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Guangning Liu, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Huiping Tan, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Fa Zhang, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Lin Liu, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China.

Rongming Liu, MOE Key Laboratory of Bio-Intelligent Manufacturing, School of Bioengineering, Dalian University of Technology, Dalian 116024, China; Ningbo Institute of Dalian University of Technology, Ningbo 315016, China.

Supplementary data

Supplementary data is available at NAR online.

Conflict of interest

None declared.

Funding

This work was supported by the National Key R&D Program of China (2023YFC3402300 to R.L.), the National Natural Science Foundation of China (22278058 to R.L., 22578048 to L.L., and 22208044 to L.L.), Scientific Research Innovation Capability Support Project for Young Faculty (SRICSPYF-ZY2025107 to R.L.), the Xingliao Talent Plan project (XLYC2203075 to R.L.), the Natural Science Foundation of Liaoning Province (2024-MSBA-09 to L.L. and 2025JH2/101330156 to R.L.), the Science and Technology Innovation Foundation of Dalian (2023JJ12SN030 to R.L.), and Fundamental Research Funds for the Central Universities (DUT24YG131 to R.L. and DUT25LAB105 to L.L.). Funding to pay the Open Access publication charges for this article was provided by the National Key R&D Program of China, and Scientific Research Innovation Capability Support Project for Young Faculty.

Data availability

The targeted amplicon sequencing data generated in this study have been deposited in the NCBI Sequence Read Archive (SRA) under the accession number PRJNA1450559. All other relevant data are avilable in the main text or the source data. All custom codes are available at https://github.com/Liurmlab/NGS and https://doi.org/10.5281/zenodo.18993788.

References

  • 1. Lawson  CE, Harcombe  WR, Hatzenpichler  R  et al.  Common principles and best practices for engineering microbiomes. Nat Rev Microbiol. 2019;17:725–41. 10.1038/s41579-019-0260-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Warner  JR, Reeder  PJ, Karimpour-Fard  A  et al.  Rapid profiling of a microbial genome using mixtures of barcoded oligonucleotides. Nat Biotechnol. 2010;28:817–25. 10.1038/nbt.1662. [DOI] [PubMed] [Google Scholar]
  • 3. Wang  HH, Isaacs  FJ, Carr  PA  et al.  Programming cells by multiplex genome engineering and accelerated evolution. Nature. 2009;460:894–8. 10.1038/nature08187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Garst  A, Bassalo  M, Pines  G  et al.  Genome-wide mapping of mutations at single-nucleotide resolution for protein, metabolic and genome engineering. Nat Biotechnol. 2017;34:1126–8. 10.1038/nbt.3724. [DOI] [PubMed] [Google Scholar]
  • 5. Deng  L, Zhou  YL, Cai  Z  et al.  Massively parallel CRISPR-assisted homologous recombination enables saturation editing of full-length endogenous genes in yeast. Sci Adv. 2024;10:adh9382. 10.1126/sciadv.adh9382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Ronda  C, Pedersen  LE, Sommer  MO  et al.  CRMAGE: CRISPR Optimized MAGE Recombineering. Sci Rep. 2016;6:19452. 10.1038/srep19452. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Liu  R, Liang  L, Freed  EF  et al.  Engineering regulatory networks for complex phenotypes in E. coli. Nat Commun. 2020;11:4050. 10.1038/s41467-020-17844-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Jiang  Y, Chen  B, Duan  C  et al.  Multigene editing in the Escherichia coli genome via the CRISPR–Cas9 system. Appl Environ Microbiol. 2015;81:2506–14. 10.1128/AEM.03445-14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Ao  X, Yao  Y, Li  T  et al.  A multiplex genome editing method for Escherichia coli based on CRISPR–Cas12a. Front Microbiol. 2018;9:2307. 10.3389/fmicb.2018.02307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Erental  A, Kalderon  Z, Saada  A  et al.  Apoptosis-like death, an extreme SOS response in Escherichia coli. mBio. 2014;5:e01426–14. 10.1128/mBio.01426-14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Kogoma  T. Stable DNA replication: interplay between DNA replication, homologous recombination, and transcription. Microbiol Mol Biol Rev. 1997;61:212–38. 10.1128/mmbr.61.2.212-238.1997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Gao  B, Liang  L, Su  L  et al.  Structural basis for regulation of SOS response in bacteria. Proc Natl Acad Sci USA. 2023;120:e2217493120. 10.1073/pnas.2217493120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Huang  M, Zhou  Z, Elledge  SJ. The DNA replication and damage checkpoint pathways induce transcription by inhibition of the Crt1 repressor. Cell. 1998;94:595–605. 10.1016/S0092-8674(00)81598-8. [DOI] [PubMed] [Google Scholar]
  • 14. Yang  DD, Rusch  LM, Widney  KA  et al.  Synonymous edits in the Escherichia coli genome have substantial and condition-dependent effects on fitness. Proc Natl Acad Sci USA. 2024;121:e2316834121. 10.1073/pnas.2316834121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Shen  X, Song  S, Li  C  et al.  Synonymous mutations in representative yeast genes are mostly strongly non-neutral. Nature. 2022;606:725–31. 10.1038/s41586-022-04823-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Robertson  WE, Funke  LFH, de la Torre  D  et al.  Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science. 2021;372:1057–62. 10.1126/science.abf8848. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Anzalone  AV, Koblan  LW, Liu  DR. Genome editing with CRISPR–Cas nucleases, base editors, transposases and prime editors. Nat Biotechnol. 2020;38:824–44. 10.1038/s41587-020-0561-9. [DOI] [PubMed] [Google Scholar]
  • 18. Komor  AC, Kim  YB, Packer  MS  et al.  Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature. 2016;533:420–4. 10.1038/nature17946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Gaudelli  NM, Komor  AC, Rees  HA  et al.  Programmable base editing of A.T to G.C in genomic DNA without DNA cleavage. Nature. 2017;551:464–71. 10.1038/nature24644. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Doman  JL, Raguram  A, Newby  GA  et al.  Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat Biotechnol. 2020;38:620–8. 10.1038/s41587-020-0539-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Neugebauer  ME, Hsu  A, Arbab  M  et al.  Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nat Biotechnol. 2023;41:673–85. 10.1038/s41587-023-01704-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Thuronyi  BW, Koblan  LW, Levy  JM  et al.  Continuous evolution of base editors with expanded target compatibility and improved activity. Nat Biotechnol. 2019;37:1070–9. 10.1038/s41587-019-0263-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Tong  Y, Whitford  CM, Robertsen  HL  et al.  Highly efficient DSB-free base editing for streptomycetes with CRISPR-BEST. Proc Natl Acad Sci USA. 2019;116:20366–75. 10.1073/pnas.1909903116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Banno  S, Nishida  K, Arazoe  T  et al.  Deaminase-mediated multiplex genome editing in Escherichia coli. Nat Microbiol. 2018;3:423–9. 10.1038/s41564-018-0137-6. [DOI] [PubMed] [Google Scholar]
  • 25. Koblan  LW, Arbab  M, Shen  MW  et al.  Efficient C•G-to-G•C base editors developed using CRISPRi screens, target-library analysis, and machine learning. Nat Biotechnol. 2021;39:1414–25. 10.1038/s41587-021-01079-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Zhao  D, Li  J, Li  S  et al.  Glycosylase base editors enable C-to-A and C-to-G base changes. Nat Biotechnol. 2021;39:35–40. 10.1038/s41587-020-00784-6. [DOI] [PubMed] [Google Scholar]
  • 27. Tong  H, Wang  X, Liu  Y  et al.  Programmable A-to-Y base editing by fusing an adenine base editor with an N-methylpurine DNA glycosylase. Nat Biotechnol. 2023;41:1080–4. 10.1038/s41587-023-01848-7. [DOI] [PubMed] [Google Scholar]
  • 28. Chen  L, Hong  M, Luan  C  et al.  Adenine transversion editors enable precise, efficient A•T-to-C•G base editing in mammalian cells and embryos. Nat Biotechnol. 2024;42:638–50. 10.1038/s41587-024-02217-4. [DOI] [PubMed] [Google Scholar]
  • 29. Tong  H, Liu  N, Wei  Y  et al.  Programmable deaminase-free base editors for G-to-Y conversion by engineered glycosylase. Natl Sci Rev. 2023;10:nwad143. 10.1093/nsr/nwad143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Choi  JY, Lim  S, Kim  EJ  et al.  Translesion synthesis across abasic lesions by human B-family and Y-family DNA polymerases α, δ, η, ι, κ, and REV1. J Mol Biol. 2010;404:751–72. 10.1016/j.jmb.2010.09.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Haracska  L, Unk  I, Johnson  RE  et al.  Roles of yeast DNA polymerases delta and zeta and of Rev1 in the bypass of abasic sites. Genes Dev. 2001;15:839–44. 10.1101/gad.875201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Koblan  LW, Doman  JL, Wilson  C  et al.  Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat Biotechnol. 2018;36:843–6. 10.1038/s41587-018-0059-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Wang  Y, Zhao  D, Sun  L  et al.  Engineering of the translesion DNA synthesis pathway enables controllable C-to-G and C-to-A base editing in Corynebacterium glutamicum. ACS Synth Biol. 2022;11:3368–78. 10.1021/acssynbio.2c00265. [DOI] [PubMed] [Google Scholar]
  • 34. Shelake  RM, Pramanik  D, Kim  JY. Improved dual base editor systems (iACBEs) for simultaneous conversion of adenine and cytosine in the bacterium Escherichia coli. mBio. 2023;14:e0229622. 10.1128/mbio.02296-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Wu  Y, Wang  J, Zhang  Z  et al.  Triple base editor catalyzes saturation mutation of adenine, cytidine, and guanine. Nucleic Acids Res. 2026;54:gkaf1423. 10.1093/nar/gkaf1423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Chen  L, Park  JE, Paa  P  et al.  Programmable C:G to G:C genome editing with CRISPR–Cas9-directed base excision repair proteins. Nat Commun. 2021;12:1384. 10.1038/s41467-021-21626-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Yuan  YM, Liao  XH, Li  S  et al.  Base editor-mediated large-scale screening of functional mutations in bacteria for industrial phenotypes. Sci China Life Sci. 2024;67:1051–60. 10.1007/s11427-023-2366-1. [DOI] [PubMed] [Google Scholar]
  • 38. Lu  J, Wang  Y, Xu  M  et al.  Efficient biosynthesis of 3-hydroxypropionic acid from ethanol in metabolically engineered Escherichia coli. Bioresour Technol. 2022;363:127907. 10.1016/j.biortech.2022.127907. [DOI] [PubMed] [Google Scholar]
  • 39. Zeitoun  RI, Garst  AD, Degen  GD  et al.  Multiplexed tracking of combinatorial genomic mutations in engineered cell populations. Nat Biotechnol. 2015;33:631–7. 10.1038/nbt.3177. [DOI] [PubMed] [Google Scholar]
  • 40. Nishimasu  H, Shi  X, Ishiguro  S  et al.  Engineered CRISPR–Cas9 nuclease with expanded targeting space. Science. 2018;361:1259–62. 10.1126/science.aat7290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Labahn  J, Schärer  OD, Long  A  et al.  Structural basis for the excision repair of alkylation-damaged DNA. Cell. 1996;86:321–9. 10.1016/S0092-8674(00)80151-0. [DOI] [PubMed] [Google Scholar]
  • 42. Vincent  MS, Uphoff  S. Cellular heterogeneity in DNA alkylation repair increases population genetic plasticity. Nucleic Acids Res. 2021;50:D211–21. 10.1093/nar/gkab808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Bae  S, Park  J, Kim  JS. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics. 2014;30:1473–5. 10.1093/bioinformatics/btu048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Grantham  R, Gautier  C, Gouy  M  et al.  Codon catalog usage is a genome strategy modulated for gene expressivity. Nucleic Acids Res. 1981;9:43–74. 10.1093/nar/9.1.43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Subramaniam  AR, Deloughery  A, Bradshaw  N  et al.  A serine sensor for multicellularity in a bacterium. eLife. 2013;2:e01501. 10.7554/eLife.01501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Liang  H, Ma  X, Ning  W  et al.  Constructing an ethanol utilization pathway in Escherichia coli to produce acetyl-CoA derived compounds. Metab Eng. 2021;65:52–65. 10.1016/j.ymben.2021.03.008. [DOI] [PubMed] [Google Scholar]
  • 47. Chong  H, Huang  L, Yeow  J  et al.  Improving ethanol tolerance of Escherichia coli by rewiring its global regulator cAMP receptor protein (CRP). PLoS One. 2013;8:e57628. 10.1371/journal.pone.0057628. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Sanford  PA, Woolston  BM. Synthetic or natural? Metabolic engineering for assimilation and valorization of methanol. Curr Opin Biotechnol. 2022;74:164–70. 10.1016/j.copbio.2021.11.010. [DOI] [PubMed] [Google Scholar]
  • 49. He  Y, Zhou  X, Chang  C  et al.  Protein language models-assisted optimization of a uracil-N-glycosylase variant enables programmable T-to-G and T-to-C base editing. Mol Cell. 2024;84:1406–1421.e8. 10.1016/j.molcel.2024.02.015. [DOI] [PubMed] [Google Scholar]
  • 50. Li  Y, Li  S, Li  C  et al.  Engineering a plant A-to-K base editor with improved performance by fusion with a transactivation module. Plant Communications. 2023;4:100667. 10.1016/j.xplc.2023.100667. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Huang  ME, Qin  Y, Shang  Y  et al.  C-to-G editing generates double-strand breaks causing deletion, transversion and translocation. Nat Cell Biol. 2024;26:294–304. 10.1038/s41556-024-01263-8. [DOI] [PubMed] [Google Scholar]
  • 52. Chen  L, Hong  M, Luan  C  et al.  Adenine transversion editors enable precise, efficient A•T-to-C•G base editing in mammalian cells and embryos. Nat Biotechnol. 2024;42:638–50. 10.1038/s41587-024-02217-4. [DOI] [PubMed] [Google Scholar]
  • 53. Wang  Y, Zhao  D, Sun  L  et al.  Engineering of the translesion DNA synthesis pathway enables controllable C-to-G and C-to-A base editing in Corynebacterium glutamicum. ACS Synth Biol. 2022;11:3120–33. 10.1021/acssynbio.2c00287. [DOI] [PubMed] [Google Scholar]
  • 54. Wang  Y, Cheng  H, Liu  Y  et al.  In-situ generation of large numbers of genetic combinations for metabolic reprogramming via CRISPR-guided base editing. Nat Commun. 2021;12:678. 10.1038/s41467-021-21003-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Ye  L, Zhao  D, Li  J  et al.  Glycosylase-based base editors for efficient T-to-G and C-to-G editing in mammalian cells. Nat Biotechnol. 2024;42:1538–47. 10.1038/s41587-024-02267-5. [DOI] [PubMed] [Google Scholar]
  • 56. Guo  J, Wang  T, Guan  C  et al.  Improved sgRNA design in bacteria via genome-wide activity profiling. Nucleic Acids Res. 2018;46:e87–. 10.1093/nar/gky411. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Liu  RM, Liang  LL, Freed  E  et al.  Synthetic chimeric nucleases function for efficient genome editing. Nat Commun. 2019;10:5524. 10.1038/s41467-019-13419-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Liu  X, Tan  H, Wang  J  et al.  SELECT: high-precision genome editing strategy via integration of CRISPR–Cas and DNA damage response for cross-species applications. Nucleic Acids Res. 2025;53:gkae1156. 10.1093/nar/gkae1156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Christie  KA, Guo  JA, Silverstein  RA  et al.  Precise DNA cleavage using CRISPR-SpRYgests. Nat Biotechnol. 2023;41:409–16. 10.1038/s41587-022-01492-y. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

gkag775_Supplemental_File

Data Availability Statement

The targeted amplicon sequencing data generated in this study have been deposited in the NCBI Sequence Read Archive (SRA) under the accession number PRJNA1450559. All other relevant data are avilable in the main text or the source data. All custom codes are available at https://github.com/Liurmlab/NGS and https://doi.org/10.5281/zenodo.18993788.


Articles from Nucleic Acids Research are provided here courtesy of Oxford University Press

RESOURCES