Abstract
Cytosine base editors (CBEs) show promise for multiplex gene knockout applications, but impure edits, indels and off-targets still frequently occur. We describe here QBEmax, which exhibits high efficiency, low indel and off-targets and high product purity with up to 99.8% of edits comprised of C-to-T. Through molecular dynamic modeling, QBEmax presents as a compact and stable base editor that shields protected bases from undesired repair processes.
Subject terms: Genetic engineering, Targeted gene repair
A base editor designed using circular permutation reduces off-target effects while enhancing purity.
Main
The two most common classes of base editors are modular fusion proteins comprised of a deaminase and DNA-binding protein to perform targeted and precise C·G-to-T·A (cytosine base editors (CBEs)) or A·T-to-G·C (adenine base editors (ABEs)) base conversions with high efficiency1,2. In addition to the use of base editors to correct disease-causing genetic mutations, base editors have shown promise as an alternative to nucleases for multiplex gene knockout applications. CBEs can precisely edit arginine, glutamine or tryptophan codons to generate premature stop codons; because this process does not undergo the formation of DNA double-strand breaks, base editing for multiplex gene knockouts minimizes the risk of genomic translocations, cell toxicity and DNA chromothripsis3,4. However, unintended indels and impure base editing byproducts (for example, C-to-G and C-to-A) are still frequently observed when using CBEs. The formation of impure edits could be detrimental and act as a missense mutation in a gene otherwise targeted for knockout. In this regard, an ideal CBE for gene knockout should exhibit properties such as high efficiency, low indel formation, low off-target edits, an expanded editing window to reach more potential bases and, importantly, high product purity for safety purposes.
Because canonical base editors are comprised of different proteins fused together end-to-end, each protein’s orientation may not be best to balance catalytic processes including nontarget strand association/dissociation by the deaminase, target base deamination, base protection/exposure and, ultimately, endogenous cellular DNA repair enzymes. Previous engineering studies have demonstrated that either the use of an inlaid deaminase domain or a circularly permuted Cas protein could affect the on-target or off-target editing efficiencies, product purity or editing windows of base editors5–9. However, the combination of these parameters has not been optimized using any one approach. To obtain a more ideal base editor, we hypothesized that we could treat the deaminase and Cas9 protein together as one complex comprised of different domains. Therefore, by shuffling the orientation of domains from both the deaminase and Cas9 together, which combines concepts of circularly permuted proteins and inlaid deaminases, we hoped to identify a CBE that maximizes the beneficial properties of using CBEs for gene knockout.
We first designed four base editor orientations in which a deaminase was internally embedded within a circularly permuted Cas9(D10A) protein10 (Fig. 1a) based on Cas9 positions previously found to be amenable for circular permutation or for inlaying a small peptide10,11. These architectures are hereby designated as Q base editors (QBE1 to QBE4) to reflect the reconstitution process of circularizing domains from the Cas9 protein and inserting a deaminase internally to generate a new start codon (circle with an internal cut resembling the letter ‘Q’). To characterize the properties of different QBE architectures, six different cytidine deaminases (rAPOBEC1 (ref. 1), hA3A, mini-Sdd3, mini-Sdd6, Sdd7 and mini-Sdd9 (ref. 12)) were evaluated.
Fig. 1. Design and characterization of QBEmax.
a, Schematic representations of the BE4max, QBE1, QBE2, QBE3 (QBEmax), QBE4 editors. Numbers under and above nCas9(D10A) represent amino acid positions in reference to wild-type SpCas9. b, Frequencies of C-to-T or C-to-R conversions (left y axis) and indels (right y axis) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax; values and error bars represent the means and s.e.m. for three independent biological replicates. c, Gray scale heat map showing average cytosine base editing frequencies by mini-Sdd9-BE4max and mini-Sdd9-QBEmax at each protospacer position across 17 endogenous sites tested. Numbers below the heat map indicate protospacer positions. d, AlphaFold3-predicted structures of mini-Sdd9-BE4max and mini-Sdd9-QBEmax bind to a sgRNA and DNA target (site 1). Blue, Cas9 domains; orange, sgRNA; purple, dsDNA target; pink, mini-Sdd9. e–g, Average editing frequencies of C-to-T conversions (e), indels (f) and ratio of base edit-to-indel (g) induced by the mini-Sdd9-BE4max or mini-Sdd9-QBEmax editors across 17 endogenous sites; each dot represents the mean for three independent biological replicates for a specified target site, the violin plot shows the base editing frequency distribution with medians and quartiles, and significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t-test, n = 17. h, Percent of edited reads with C-to-T conversions among edited events at each C1 to C16 base position, cumulated across 17 endogenous sites; values and error bars represent the means and s.e.m. for three independent biological replicates. CMV pro, enhanced cytomegalovirus promoter; DEA, deaminase; NLS, nuclear localization signals; bGH, bovine growth hormone polyadenylation signal.
Plasmids encoding corresponding CBEs were transfected in HEK293T cells and compared with the canonical BE4max architecture13 at two endogenous genomic sites. Deep sequencing revealed that QBE3-based editors showed comparable or higher editing frequencies for four of the six deaminases evaluated; in contrast, QBE1-based, QBE2-based and QBE4-based editors exhibited lower editing frequencies (Supplementary Fig. 1a,b). Compared to the BE4max counterparts, the QBE3-based editors showed substantially lower indels, with an average indel reduction of 60.4% for site 1 and 62.6% for site 3 (Supplementary Fig. 1c,d). We next measured the edit-to-indel ratios and found that six of the seven editors at site 1, and five of the seven editors at site 3 showed substantially higher edit/indel ratios (Supplementary Fig. 1e,f). We then analyzed the editing window and product purities of QBE editors. We observed higher editing efficiencies at PAM-proximal Cs for QBE3 editors and greatly improved product purities compared to BE4max editors (Supplementary Fig. 2a,b). Based on these, we hereby refer to QBE3 as QBEmax (Fig. 1a).
From initial evaluations, we found that mini-Sdd9-based QBE editors exhibited superior editing properties in terms of editing activity, indel formation and product purity. Using mini-Sdd9, we designed seven additional QBEs (QBE5–QBE11; Supplementary Fig. 3a). We found that only one editor, QBE6 demonstrated similar performance to QBEmax in terms of editing efficiency and purity; however, its editing window appeared narrow, so we hereby designate it as QBEn (Supplementary Fig. 3b,c). Because of mini-Sdd9-QBEmax’s overall superior performance and relatively wide editing window, which is desired for expanding the targeting scope of gene knockout applications with base editors, we selected it for further study.
It was reported that fusing the deaminase to the N terminus of a circularly permuted Cas9 (referred to as CP-BE)5 could broaden the editing window, and inlaying the deaminase within the Cas9 protein (referred to as inlaid-BE)6–9 could affect the editing efficiency or off-target efficiency. We next compared mini-Sdd9-QBEmax with CP-BEs and inlaid-BEs with Cas9 permutation or deaminases inlaid at positions used in the QBE1–QBE11 editors. We found that all CP-BEs and inlaid-BEs induced more impure products than QBEmax. Notably, QBEmax (which uses positions 1,031 and 1,244 for modular assembly) outperforms individual CP-1031 and inlaid-1244 CBEs when comparing the combination of editing frequencies, indel formation and product purities (Extended Data Fig. 1), suggesting that the modularly designed QBEmax architecture enhances desired properties from each functional domain.
Extended Data Fig. 1. Comparison of QBEmax, BE4max, CP-BEs and inlaid-BEs at site 1 and site 3.
a, Three-dimensional plots of editing efficiencies (y axis), indels (z axis) and product purities (x axis) induced by BE4max, QBEmax, CP-BEs and inlaid-BEs at site 1 (left) and site 3 (right). A golden cube in the top right corner highlights ideal base editors that exhibit high editing efficiency, high product purity and low indel formation simultaneously (thresholds for efficiency, purity and indel are set to be >70%, >98% and <3%, respectively). Each point represents the mean for three independent biological replicates; the target sites are shown above the plots with targeted Cs in red and PAM in blue. b, Values used for plotting the graph in a.
To further profile editing properties of mini-Sdd9-QBEmax, we compared mini-Sdd9-QBEmax with mini-Sdd9-BE4max across 17 endogenous genomic sites in HEK293T cells. We found that, in contrast to mini-Sdd9-BE4max, which biases editing at the PAM-distal region, mini-Sdd9-QBEmax showed a wider editing window as far as C16 (Fig. 1b and Supplementary Fig. 4). Aggregate analyses revealed a ‘forward-shifted’ and wider editing window for mini-Sdd9-QBEmax with target Cs between 4 and 14 being favored (Fig. 1c). We used AlphaFold3 (ref. 14) to predict a mini-Sdd9-QBEmax–sgRNA–target DNA ternary structure and compared it with that of the corresponding BE4max architecture (Fig. 1d). We found that these structures revealed the deaminase in mini-Sdd9-QBEmax being more closely associated to PAM-proximal Cs, while in BE4max, more closely associated to PAM-distal Cs, which is consistent with experimental results from genomic edits.
The average editing frequencies across 17 genomic sites for mini-Sdd9-QBEmax and mini-Sdd9-BE4max were 52.4 ± 2.4% and 54.5 ± 2.2%, respectively (Fig. 1e). Mini-Sdd9-QBEmax induced lower indels at 16 of 17 sites tested, with average indel frequencies decreasing by 56.5% from 2.8 ± 0.3% to 1.2 ± 0.2% (Fig. 1f and Supplementary Fig. 4), which also substantially increases average edit-to-indel ratios (Fig. 1g). We next evaluated cytosine base editing product purities, which is calculated as the proportion of ‘C’ edited to ‘T’ as opposed to ‘G’ or ‘A’, for each position within the protospacer and aggregated all sites together. Importantly, mini-Sdd9-QBEmax exhibited superior product purities (99.4% ± 0.4%) at all positions within a C1–C16 editing window compared to that of mini-Sdd9-BE4max (95.5% ± 2.8%; Fig. 1b,h and Supplementary Fig. 4). To test the versatility of the QBEmax, we evaluated mini-Sdd9-QBEmax in additional mammalian cell lines, including A549, HeLa and HCT116. QBEmax exhibited higher or comparable editing efficiencies, improved product purities and decreased indels in all cell lines tested (Extended Data Fig. 2). These results highlight QBEmax in achieving efficient and precise base edits at a flexible editing window with minimal indel and byproducts.
Extended Data Fig. 2. Performance of QBEmax in other mammalian cell lines.
a–c, Desired editing efficiencies (a), indels (b), and percent of edited reads (c) with C-to-T conversions induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax editors at site 1 with C7 in red used to calculate the efficiency and product purity. Values and error bars represent the means and s.e.m. for three independent biological replicates.
Chimeric antigen receptor-T cell (CAR-T) therapy has demonstrated success as a cancer immunotherapy for hematological malignancies. Many clinical trials and research studies have found that multiplex knockout of genes related to immune rejection and graft-versus-host disease (GvHD) would further benefit CAR-T therapy in terms of both durability and potency. Encouraged by the performance of mini-Sdd9-QBEmax, we next sought to perform multiplex gene knockout to simultaneously edit genes that could compromise the efficacy of CAR-Ts. Five genes, PD-1 (ref. 15), CISH16, Fas17, TGFBR2 (truncating off the endodomain)18–20 and TRAC21, were selected, which all previously demonstrated potential in improving CAR-T performance when knocked out or downregulated.
We first identified all possible SpCas9 protospacers with an NGG PAM and target C located within codons encoding tryptophan (W), arginine (R) or glutamine (Q), so that a cytosine base edit would generate a stop codon (TAA, TAG or TGA). We obtained 46, 21, 16, 30 and 4 protospacers for PD-1, CISH, Fas, TGFBR2 and TRAC, respectively (Supplementary Table 1). We next filtered for potential off-target sites (mismatches ≤ 3) and ultimately selected 16 targets for PD-1, 9 for CISH, 10 for Fas, 10 for TGFBR2 and 3 for TRAC. Lastly, we included one additional target for CISH, PD-1 and TRAC, which disrupts a splice site to perform gene knockout as reported previously21,22 (Fig. 2a).
Fig. 2. Multiplex base editing gene knockout of CAR-T targets.
a, Schematic representations showing PD-1, CISH, Fas, TGFBR2 and TRAC genes. Light blue boxes indicate exons of genes, and short red lines represent the position of selected protospacers. b–d, Average editing frequencies of desired editing efficiencies (b), indels (c) and ratio of base edits to indels (d) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax editors for PD-1 (n = 17), CISH (n = 10), Fas (n = 10), TGFBR2 (n = 10) and TRAC (n = 4) genes; each dot represents the mean for three independent biological replicates for a specified target site and the violin plot shows base editing frequency distribution with medians and quartiles. e, Percent of edited reads with C-to-T or C-to-R conversions at target genes indicated with values and error bars representing the means and s.e.m. for three independent biological replicates across all target sites for PD-1 (n = 17), CISH (n = 10), Fas (n = 10), TGFBR2 (n = 10) and TRAC (n = 4) genes; significances are indicated between BE4max and QBEmax by exact P value using two-tailed Student’s t-test. f, Schematic representations of potential editing outcomes induced by C-to-T, C-to-G and C-to-A conversions of tryptophan (W), arginine (R) and glutamine (Q). g–i, Desired editing efficiencies (g), indels (h) and percent of edited reads with C-to-T or C-to-R conversions at target genes indicated (i) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax editors during multiplexed base editing of five genes; values and error bars represent the means and s.e.m., respectively, for four independent biological replicates. j, Single-cell colony analysis of multiplex base editing distributions in unsorted and sorted cell populations. k, Schematic representation of the experimental design for the R-loop assay. l, Frequencies of C-to-T conversions at the dSaCas9-induced R-loop sites; values and error bars represent the means and s.e.m. for four independent biological replicates. m, Number of C-to-U RNA variants induced by the editors indicated; values and error bars represent the means and s.e.m. for two (Cas9(D10A)) or four (BE4max and QBEmax) independent biological replicates, significances are indicated by exact P value using one-way ANOVA Tukey’s multiple comparisons.
HEK293T cells were transfected with mini-Sdd9-BE4max or mini-Sdd9-QBEmax together with each sgRNA plasmid. We then analyzed desired editing efficiencies (calculated as percent C-to-T for stop codon creation), indel frequencies, desired edit-to-indel ratios and product purities. We found that mini-Sdd9-QBEmax achieved comparable or slightly higher average desired editing at the target base for all sites aggregated for each of the five genes (Fig. 2b). Average indel frequencies for all sites aggregated decreased by 76%, 75%, 59% and 71% for PD-1, CISH, Fas and TGFBR2, respectively (Fig. 2c). The average indel frequency at TRAC was 0.8 ± 0.15% for mini-Sdd9-BE4max and 1.0 ± 0.34% for mini-Sdd9-QBEmax due to one outlier at TRAC-site 3 whereby mini-Sdd9-BE4max and mini-Sdd9-QBEmax exhibited 1.4 ± 0.18% and 2.9 ± 0.22%, respectively. Cumulatively, the desired edit-to-indel ratios induced by mini-Sdd9-QBEmax were 3.26, 3.99, 2.02, 9.91 or 1.99-fold higher than that of mini-Sdd9-BE4max (Fig. 2d). Importantly, average product purities at the target cytosine base for stop codon creation by mini-Sdd9-QBEmax and mini-Sdd9-BE4max were 99.7% versus 95.7% for PD-1, 99.7% versus 97.7% for CISH, 99.7% versus 96.5% for Fas, 99.8% versus 98.0% for TGFBR2 and 99.5% versus 95.8% for TRAC (Fig. 2e). This increase in product purity minimizes the formation of missense mutations from imprecise C-to-G or C-to-A edits (Fig. 2f). When analyzed individually, mini-Sdd9-QBEmax induced lower indels at 47 of the 51 target sites and higher product purities at 45 of the 51 target sites (Extended Data Figs. 3–7).
Extended Data Fig. 3. Evaluation of BE4max and QBEmax for base editing-mediated knockout of CISH.
a, Frequencies of C-to-T or C-to-R conversions (left y axis) and indels (right y axis) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax for ten sites targeting CISH knockout. b, Ratio of base edits to indels induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax; significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t test. c, Percent of edited reads with C-to-T or C-to-R conversions at the target sites indicated. Values and error bars represent the means and s.e.m. for three independent biological replicates.
Extended Data Fig. 7. Evaluation of BE4max and QBEmax for base editing-mediated knockout of TRAC.
a, Frequencies of C-to-T or C-to-R conversions (plot on the left y axis) and indels (plot on the right y axis) induced by the mini-Sdd9-BE4max and mini-Sdd9-QBEmax for four sites targeting TRAC knockout. b, Ratio of base edits to indels induced by the mini-Sdd9-BE4max and the mini-Sdd9-QBEmax; significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t test. c, Percent of edited reads with C-to-T or C-to-R conversions at the target sites indicated. Values and error bars represent the means and s.e.m. for three independent biological replicates.
For each gene, we identified one ideal guide and next edited PD-1, CISH, Fas, TGFBR2 and TRAC simultaneously for multiplex gene knockout. We transformed plasmids for each of the five sgRNAs together with QBEmax or BE4max editors into HEK293T cells. We observed that mini-Sdd9-QBEmax exhibited comparable or higher editing compared to mini-Sdd9-BE4max across all five sites in HEK293T cells in the absence of any selection pressure (Fig. 2g). Notably, mini-Sdd9-QBEmax achieved lower indel formations (Fig. 2h) and exhibited superior product purity (Fig. 2i) at all five genes. To validate that all five base edits occurred in a single cell, we sequenced 48 and 112 QBEmax-transfected single-cell colonies arising from unsorted or sorted cell populations, respectively. We found that 21 (43.8%) and 100 (89.3%) cell colonies exhibited all five genes edited, respectively, demonstrating successful multiplex base editing by QBEmax in a single cell (Fig. 2j). To further evaluate the potential of QBEmax, we co-electroporated QBEmax and all five sgRNA plasmids into an immortal Jurkat T cell line. In these T cells, QBEmax also exhibited superior editing efficiencies, improved product purities and decreased indels, which is similar to its performance in HEK293T cells and further supports the versatility of QBEmax for safe and robust base editing in clinical applications (Extended Data Fig. 8).
Extended Data Fig. 8. Multiplex base editing with QBEmax and BE4max in immortal Jurkat cells.
a–c, Desired editing efficiencies (a), indels (b) and percent of edited reads (c) with C-to-T conversions at target genes induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax editors during multiplexed base editing of five genes; values and error bars represent the means and s.e.m. for three independent biological replicates.
Because DNA off-targets are a major concern for base editing therapeutic applications, we next evaluated Cas-independent DNA off-target effects of mini-Sdd9-QBEmax using the orthogonal R-loop assay23–25. We cotransfected a dead-SaCas9 (dSaCas9) and sgRNA to induce the formation of an orthogonal R-loop simultaneously with the multiplexed gene knockout strategy (Fig. 2k). We evaluated five orthogonal sites and deep sequencing at each orthogonal R-loop showed that mini-Sdd9-QBEmax induced lower Cas-independent off-target editing at all five R-loop sites compared to that of mini-Sdd9-BE4max (Fig. 2l). We next evaluated the RNA off-target effects of mini-Sdd9-QBEmax. We conducted whole transcriptome sequencing and analyzed the number of C-to-U variants in QBEmax, BE4max and nCas9 (D10A) treated samples together with a sgRNA plasmid targeting CISH. We found that QBEmax induced substantially lower RNA off-target edits on transcriptome-wide RNA transcripts compared to that of the BE4max without compromising DNA on-target editing (Fig. 2m and Supplementary Fig. 5). The robust desired editing efficiencies, minimized indels, high product purities and decreased DNA and RNA off-target effects portray QBEmax as an ideal base editor for multiplex gene knockout applications.
We next sought to probe the molecular basis by which mini-Sdd9-QBEmax embodies its desired properties. We performed molecular dynamic (MD) simulation analyses based on the AlphaFold3-predicted ternary structures of mini-Sdd9-QBEmax or BE4max with a sgRNA and target DNA (Fig. 1d). With these models, all-atom MD simulations of approximately 300 ns were performed (Supplementary Fig. 6a,b). We first investigated the conformational stability of these two systems by projecting their free energy landscapes onto corresponding root mean square deviation (RMSD) and radius of gyration (Rg) components. We found that both the RMSD and Rg of mini-Sdd9-QBEmax were lower than that of mini-Sdd9-BE4max, and only one single stable energy state was observed (Fig. 3a,b and Supplementary Fig. 6b). This suggests that the QBEmax architecture better treats the deaminase and Cas protein as one complex so that each domain is oriented compactly within itself. At the minimum energy state, while the deaminase was predicted to be associated with the nontarget strand in both systems (Supplementary Fig. 6c,d), the RMSD of mini-Sdd9-QBEmax system was lower throughout the 300 ns MD process (Supplementary Fig. 6b). When analyzing individual amino acids, we observed that the root mean square fluctuation (RMSF) of the linkers connecting the deaminase to the Cas protein were lower in the QBEmax system (Fig. 3c,d) compared to the BE4max architecture. We speculate a more compact QBEmax architecture that limits the deaminase from sporadically swinging in space, thereby contributing to lower Cas-independent DNA off-target editing, lower indel formation and higher product purity.
Fig. 3. Molecular dynamic modeling reveals QBEmax as a more compact and protective editor.
a,b, The free energy landscape against RMSD and Rg for mini-Sdd9-BE4max (a) and mini-Sdd9-QBEmax (b) during a 300 ns MS simulation. c,d, RMSF plot for mini-Sdd9-BE4max (c) and mini-Sdd9-QBEmax (d) in the MD simulation, systems equilibrated after 150 ns; schematic representations of editors are shown above the plots. e, SASA analysis of Cs within the editing window of site 1 in predicted mini-Sdd9-BE4max and mini-Sdd9-QBEmax ternary structures; each replicate represents the SASA by a 1.0 nm probe during a 1 ns time scale; n = 150 ns following system equilibration, boxes and lines represent the interquartile range (IQR) and median, respectively, and whiskers represent 1.5× IQR. f, Snapshots showing exposed Cs in the editing window. Blue, Cas9 and UGI; yellow, linker; green, mini-Sdd9 deaminase.
During cytosine base editing, intermediate uracil cleavage by endogenous uracil DNA glycosylase (UNG) drives the formation of indels and imprecise C-to-G or C-to-A edits1,26. We envisioned that an ideal base editor adopts a compact and protective conformation for the exposed R-loop so that the intermediate uracil base is not excised before cellular mismatch repair resolving a permanent C-to-T conversion. To evaluate R-loop exposure, we performed solvent accessibility analyses using a 1.0 nm probe and found that the solvent-accessible surface area (SASA) was increased for C3 and substantially increased for C7 and C8 in the mini-Sdd9-BE4max compared to the mini-Sdd9-QBEmax architectures (Fig. 3e,f), suggesting that these residues are accessible by UNG and ultimately form indels and byproducts. We also speculated that the distance of the UGI to the ssDNA target bases may affect product purity. We evaluated the position and distance of the two UGI domains to the ssDNA target bases in the QBEmax, BE4max or individual CP-1031 and inlaid-1244 CBEs based on molecular dynamic modeling data. We observed that indeed both UGI domains in QBEmax exhibited a relatively shorter distance to the target bases in the ssDNA R-loop region, which suggests an inverse relationship between UGI positioning to product purities and indel formation, as others previously have also identified5 (Supplementary Fig. 6e).
Based on these results, we propose a model for QBEmax base editing. In cytosine base editing using a canonical BE4max architecture, indel formation and impure C-to-G or C-to-A base edits arise from uracil excision and abasic site formation. Because QBEmax exhibits a more compact architecture, limits deaminase swinging and shields the Cas9-induced R-loop, base editing intermediates are protected from cellular UNG excision before Cas9 detaching from the target DNA and subsequent mismatch repair. Therefore, a protective and compact base editor conformation reduces unintended effects driven by DNA repair processes and further promotes desired base editing events (Extended Data Fig. 9).
Extended Data Fig. 9. Model showing base editing outcomes for BE4max and QBEmax.
UNG, uracil DNA glycosylase; AP lyase, apurinic or apyrimidinic site lyase; NHEJ, nonhomologous end joining.
Taken together, we designed and identified a base editor architecture, QBEmax, which achieves high efficiency on-target editing while decreasing low indel formation, exhibits high edit product purities and minimizes DNA off-targets. The development of QBEmax serves as a promising base editor architecture for developing more efficient and precise base edits toward the use of base editing in multiplex therapeutic applications such as CAR-T immunotherapies. The efficient delivery of QBEmax in vivo will further expand on its use and analysis of desired base editing properties. Advances in protein prediction and MD further help shed light on the molecular basis of genome editors and would greatly aid in future developments of new editing technologies.
Methods
Plasmid construction
The construct fragments were PCR amplified using 2× Phanta Max Master Mix (Vazyme Biotech) and 2× KOD one Master Mix (TOYOBO Life Sciences) and cloned into the pCMV24 backbone, using Uniclone One Step Seamless Cloning Kit (Genesand). Plasmids for HEK293T cell line transfection were extracted and purified using EndoFree Plasmid Kits (Qiangen) or FastPure EndoFree Plasmid Mini Kit DC203 (Vazyme Biotech). The amino acid sequences of the representative vectors used are provided in the Supplementary Note. Primers used for amplicon sequencing were synthesized by the Beijing Genomics Institute and are listed in Supplementary Table 2.
Cell culture and transfection
HEK293T, A549, HeLa and HCT116 cells are cultured in DMEM (Gibco) supplemented with 10% (vol/vol) FBS (Gibco) and 1% (vol/vol) penicillin–streptomycin (Gibco); Jurkat, Clone E6-1 cells are cultured in Roswell Park Memorial Institute (RPMI, Gibco) 1640 medium supplemented with 10% (vol/vol) FBS (Gibco) in a humidified incubator at 37 °C with 5% CO2. All the cells were routinely tested for Mycoplasma contamination with a mycoplasma detection kit (TransGen Biotech). For HEK293T cells transfection, 6 × 104 cells per well were seeded into 48-well poly-d-lysine-coated plates (Corning) in the absence of antibiotic. After 16–24 h, plasmids were transfected into the cells; for single target base editing, cells were transfected with 1 μl jetPRIME transfection reagent (Polyplus), 375 ng editor plus 125 ng sgRNA plasmids per well, for simultaneous editing of the five CAR-T relevant genes and the R-loop assays, 200 ng of editor plasmid and 66 ng of sgRNA plasmid for each gene was transfected, together with 200 ng of dSaCas9 plasmid and 66 ng of an orthogonal R-loop inducing sgRNA plasmid, at a 60–80% cell confluency. Cells were washed with PBS and followed by DNA extraction 72 h after transfection. To isolate single-cell colonies, transfected cells with or without sorting were diluted and plated into 96-well plates at a density of 0.9 cells per well. The cell colonies were cultured for 2 to 3 weeks and then transferred into 48-well plates to grow for another 5 days before cell lysis, DNA extraction and sequencing of all five targeted genomic loci. Effective editing in a single-cell colony for any individual gene is benchmarked as having an editing efficiency surpassing 50% at the targeted base when analyzed by CRISPResso2 (ref. 27). For the sorted cell population, sgRNA plasmids and a mini-Sdd9-QBEmax-P2A-mScarlet plasmid were transfected and the top 5% of mScarlet-positive cells were collected and diluted into single-cell colonies. For A549, HeLa and HCT116 cells, 5 × 104, 2 × 105 or 4 × 105 cells were seeded into 12-well poly-d-lysine-coated plates (Corning), respectively, in the absence of antibiotic. After 16–24 h, 375 ng of mini-Sdd9-BE4max-P2A-mScarlet or mini-Sdd9-QBEmax-P2A-mScarlet plasmids were transfected into cells together with 125 ng sgRNA plasmid while using 2 μl of jetPRIME transfection reagent (Polyplus). After 48 h, cells were resuspended for fluorescence-activated cell sorting. For Jurkat cells, 1,000 ng of mini-Sdd9-BE4max-P2A-mScarlet or mini-Sdd9-QBEmax-P2A-mScarlet editor plasmids and 250 ng of sgRNA plasmid for each of the five genes were cotransfected into 4 × 105 cells using Entranster-E (Engreen Biosystem) and the 4D-Nucleofector (Lonza Biosciences) with program DS167. mScarlet-positive cells were collected for DNA extraction.
Fluorescence-activated cell sorting
Cells were collected and resuspended in a culture medium supplemented with 2% FBS in a 0.5 ml volume. Cells were sorted on a FACSAria III (BD Biosciences) cytometer after gating for the singlet-cell population by the mScarlet signal. For RNA off-target analysis, both the fluorescence-positive and negative cells were collected for further paired analysis. A representative flow cytometry gating strategy can be found in Supplementary Fig. 7.
Genomic DNA isolation from mammalian cell culture
Genomic DNA extraction was performed by the addition of 100 μl freshly prepared lysis buffer (10 mM Tris–HCl (pH8.0), 0.05% SDS and 25 μg ml−1 proteinase K (Thermo Fisher Scientific) directly into the 48-well culture plate after cells were washed once with 1× Dulbecco’s PBS (Thermo Fisher Scientific). The mixture was incubated at 37 °C for 60 min and then treated at 80 °C in the thermocycler for 20 min.
Transcriptome-wide RNA off-target analysis
To profile the transcriptome-wide RNA off-target effects of the QBEmax and BE4max, 2 × 105 HEK293T cells were seeded into 12-well poly-d-lysine-coated plates (Corning) in the absence of antibiotics. After 24 h, 750 ng of mini-Sdd9-BE4max-P2A-mScarlet, mini-Sdd9-QBEmax-P2A-mScarlet or Cas9(D10A)-P2A-mScarlet plasmids were transfected into cells together with 250 ng sgRNA plasmid targeting the CISH-site 1 using 2 μl of jetPRIME transfection reagent (Polyplus). After 48 h, cells were collected for cell sorting. For each treatment, both the mScarlet-positive and negative cells were collected for paired analyses. RNA samples were sequenced using an MGI T7 (2 × 150 PE) platform at the JMDNA, at a depth of ~45 million reads per sample. Raw reads were cleaned by fastp28. The reads were mapped to the human reference genome (hg38) by Hisat2 (ref. 29) software (version 2.2.1). After removing duplication, variants were identified using strelka30 (version 2.9.10). Finally, C-to-U edits in the transcribed strand were considered for downstream analyses.
Amplicon deep sequencing and data analysis
Two rounds of PCR were used to amplify a DNA fragment spanning the target site. In the first round PCR, the target region was amplified from genomic DNA with site-specific primers using the DNA template. In the second round, both forward and reverse barcodes were added to the ends of the PCR products for library construction. Equal amounts of PCR product were pooled and purified with a FastPure Gel DNA Extraction Mini Kit (Vazyme Biotech. Inc.) and quantified with a Qubit 4 (Thermo Fisher Scientific). The purified products were sequenced using the Illumina Miseq platform or the BGI G99 platform, and the sequences around the target regions were examined to analyze the editing outcomes. Sequences of the NGS primers and the corresponding amplicons are listed in Supplementary Table 2. Analysis of the base editing outcomes was performed as described previously31.
MD simulations and analyses
Initial protein structure models of mini-Sdd9-BE4max and mini-Sdd9-QBEmax ternary complex were predicted using AlphaFold3 (ref. 14), with the optimal model_0 selected as the starting structure for MD simulations. MD simulations were conducted using GROMACS32 (version 2023.1, CUDA) with the Amber99BSC1 force field and the TIP3P water model33,34. In each simulation system, the complex was positioned at the center of a cubic box with a distance of 15 Å maintained between the surface and the edges of the box. Na⁺ and Cl− ions were added to the system to achieve charge neutrality with an ion concentration set at 0.15 M. During the simulation, the particle mesh Ewald method35 was employed to calculate electrostatic interactions and a cut-off distance of 14 Å was used for analysis of short-range electrostatic and van der Waals forces. The LINCS algorithm36 was applied to constrain bonds involving hydrogen atoms, and periodic boundary conditions were specified in all three dimensions. To optimize the system, energy minimization was performed using the steepest descent method, continuing until a maximum of 2,500 steps was reached or the maximum force <1,000 kJ mol−1 nm−1. Next, 100 ps of NVT equilibration and 100 ps of NPT equilibration were conducted. During the equilibration process, the V-rescale temperature and Parrinell–Rahman pressure coupling methods were employed, with the system’s temperature and pressure set to 310 K and 1.0 bar, respectively. After system equilibration, a 300 ns MD simulation was performed.
RMSD and RMSF data were generated using the gmx rmsd and gmx rmsf commands in GROMACS. The SASA data were obtained through the gmx sasa command, with a probe diameter of 1.0 nm, and snapshots of the SASA state were collected every 1 ns within the stable region. All data were statistically analyzed and visualized using R packages ggplot2 (3.5.1) and tidyverse (2.0.0). The free energy landscape data were generated using the gmx sham command with the parameter -tsham set to 310 K. The data were plotted along the two projection components of RMSD and Rg. Visualization of this data was performed using the matplotlib (3.8.2) package in Python37.
Statistical analysis
GraphPad Prism 9 software was used to analyze the data. All numerical values are presented as means ± s.e.m. Significant differences between controls and treatments were tested using the Student’s t-test or one-way ANOVA Tukey’s multiple comparisons. P < 0.05 was considered statistically significant, and P < 0.01 was considered statistically extremely significant.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Online content
Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41587-025-02641-9.
Supplementary information
Supplementary Figs. 1–7, Supplementary Tables 1–3 and Supplementary Note (sequences of mini-Sdd9-BE4max, mini-Sdd9-QBEmax and mini-Sdd9-QBEn).
sgRNA list of the 5 CAR-T genes.
Acknowledgements
This work was supported by the National Key R&D Program (2023YFF1001600), Beijing Municipal Science & Technology Commission (Z241100009024035) and Beijing Rural Revitalization Agricultural Science and Technology (project NY2401010024). We thank Y. Li at Qi Biodesign for assistance in RNA-sequencing raw data analysis and B. Zhou at Xin Hua Hospital, Shanghai Jiao Tong University School of Medicine for assistance in cell colony isolation. We thank the Flow Cytometry Facility at the National Center for Protein Sciences Beijing, particularly C. Han and Y. Wang, for their assistance. The Jurkat, Clone E6-1 cells were kindly provided by Cell Bank, Chinese Academy of Sciences.
Extended data
Extended Data Fig. 4. Evaluation of BE4max and QBEmax for base editing-mediated knockout of Fas.
a, Frequencies of C-to-T or C-to-R conversions (left y axis) and indels (right y axis) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax for ten sites targeting Fas knockout. b, Ratio of base edits to indels induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax; significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t test. c, Percent of edited reads with C-to-T or C-to-R conversions at the target sites indicated. Values and error bars represent the means and s.e.m. for three independent biological replicates.
Extended Data Fig. 5. Evaluation of BE4max and QBEmax for base editing-mediated knockout of PD-1.
a, Frequencies of C-to-T or C-to-R conversions (left y axis) and indels (right y axis) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax for 17 sites targeting PD-1 knockout. b, Ratio of base edits to indels induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax; significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t test. c, Percent of edited reads with C-to-T or C-to-R conversions at the target sites indicated. Values and error bars represent the means and s.e.m. for three independent biological replicates.
Extended Data Fig. 6. Evaluation of BE4max and QBEmax for base editing-mediated knockout of TGFBR2.
a, Frequencies of C-to-T or C-to-R conversions (left y axis) and indels (right y axis) induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax for ten sites targeting TGFBR2 knockout. b, Ratio of base edits to indels induced by mini-Sdd9-BE4max and mini-Sdd9-QBEmax; significances are indicated between the mini-Sdd9-BE4max and mini-Sdd9-QBEmax by exact P value using two-tailed Student’s t test. c, Percent of edited reads with C-to-T or C-to-R conversions at the target sites indicated. Values and error bars represent the means and s.e.m. for three independent biological replicates.
Author contributions
K.T.Z. and J.H. conceptualized the study. M.G., J.H., M.H., G.L. and Quan Gao assembled the vectors and conducted experiments in HEK293T cells. M.H., L.G. and G.L. performed amplicon sequencing. J.H. and M.G. collected and analyzed amplicon sequencing data. Qiang Gao and H.J. performed the MD analysis. Qiang Gao and Z.W. wrote scripts and processed the raw amplicon sequencing data. J.H., M.G., M.H. and K.T.Z. prepared the figures. J.H. and K.T.Z. wrote the manuscript with input from all authors.
Peer review
Peer review information
Nature Biotechnology thanks the anonymous reviewers for their contribution to the peer review of this work.
Data availability
The deep amplicon sequencing data have been deposited in the NCBI BioProject database under accession code PRJNA1147008 (ref. 38). Plasmids encoding QBEmax and QBEn are available at Addgene. All other data are available in the main paper or Supplementary Information.
Competing interests
The authors have submitted a patent application based on the results reported in this paper. K.T.Z. is the founder and holds equity at Qi Biodesign. J.H., M.G., Qiang Gao, H.J., M.H., Z.W., L.G., G.L. and Quan Gao are employees of Qi Biodesign.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data
is available for this paper at 10.1038/s41587-025-02641-9.
Supplementary information
The online version contains supplementary material available at 10.1038/s41587-025-02641-9.
References
- 1.Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A. & Liu, D. R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature533, 420–424 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Gaudelli, N. M. et al. Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage. Nature551, 464–471 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Webber, B. R. et al. Highly efficient multiplex human T cell engineering without double-strand breaks using Cas9 base editors. Nat. Commun.10, 5222 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Kosicki, M., Tomberg, K. & Bradley, A. Repair of double-strand breaks induced by CRISPR–Cas9 leads to large deletions and complex rearrangements. Nat. Biotechnol.36, 765–771 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Huang, T. P. et al. Circularly permuted and PAM-modified Cas9 variants broaden the targeting scope of base editors. Nat. Biotechnol.37, 626–631 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Jiang, L. et al. Internally inlaid SaCas9 base editors enable window specific base editing. Theranostics12, 4767–4778 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chu, S. H. et al. Rationally designed base editors for precise editing of the sickle cell disease mutation. CRISPR J.4, 169–177 (2021). [DOI] [PubMed] [Google Scholar]
- 8.Nguyen Tran, M. T. et al. Engineering domain-inlaid SaCas9 adenine base editors with reduced RNA off-targets and increased on-target DNA editing. Nat. Commun.11, 4871 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wang, Y., Zhou, L., Liu, N. & Yao, S. BE-PIGS: a base-editing tool with deaminases inlaid into Cas9 PI domain significantly expanded the editing scope. Signal Transduct. Target Ther.4, 36 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Oakes, B. L. et al. Profiling of engineering hotspots identifies an allosteric CRISPR–Cas9 switch. Nat. Biotechnol.34, 646–651 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Oakes, B. L. et al. CRISPR–Cas9 circular permutants as programmable scaffolds for genome modification. Cell176, 254–267 e216 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Huang, J. et al. Discovery of deaminase functions by structure-based protein clustering. Cell186, 3182–3195 e3114 (2023). [DOI] [PubMed] [Google Scholar]
- 13.Koblan, L. W. et al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol.36, 843–846 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.McGowan, E. et al. PD-1 disrupted CAR-T cells in the treatment of solid tumors: promises and challenges. Biomed. Pharmacother.121, 109625 (2020). [DOI] [PubMed] [Google Scholar]
- 16.Zhu, H. et al. Metabolic reprograming via deletion of CISH in human iPSC-derived NK cells promotes in vivo persistence and enhances anti-tumor activity. Cell Stem Cell27, 224–237 e226 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Yamamoto, T. N. et al. T cells genetically engineered to overcome death signaling enhance adoptive cancer immunotherapy. J. Clin. Invest.129, 1551–1565 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Rouce, R. H. et al. The TGF-β/SMAD pathway is an important mechanism for NK cell immune evasion in childhood B-acute lymphoblastic leukemia. Leukemia30, 800–811 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Bollard, C. M. et al. Adapting a transforming growth factor β-related tumor protection strategy to enhance antitumor immunity. Blood99, 3179–3187 (2002). [DOI] [PubMed] [Google Scholar]
- 20.Bollard, C. M. et al. Tumor-specific T-cells engineered to overcome tumor immune evasion induce clinical responses in patients with relapsed Hodgkin lymphoma. J. Clin. Oncol.36, 1128–1139 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Diorio, C. et al. Cytosine base editing enables quadruple-edited allogeneic CART cells for T-ALL. Blood140, 619–629 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kluesner, M. G. et al. CRISPR–Cas9 cytidine and adenosine base editing of splice-sites mediates highly-efficient disruption of proteins in primary and immortalized cells. Nat. Commun.12, 2437 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Doman, J. L., Raguram, A., Newby, G. A. & Liu, D. R. Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat. Biotechnol.38, 620–628 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol.38, 883–891 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Jin, S. et al. Rationally designed APOBEC3B cytosine base editors with improved specificity. Mol. Cell79, 728–740 e726 (2020). [DOI] [PubMed] [Google Scholar]
- 26.Komor, A. C. et al. Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity. Sci. Adv.3, eaao4774 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol.37, 224–226 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Chen, S. Ultrafast one‐pass FASTQ data preprocessing, quality control, and deduplication using fastp. iMeta2, e107 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol.37, 907–915 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat. Methods15, 591–594 (2018). [DOI] [PubMed] [Google Scholar]
- 31.Jin, S., Lin, Q., Gao, Q. & Gao, C. Optimized prime editing in monocot plants using PlantPegDesigner and engineered plant prime editors (ePPEs). Nat. Protoc.18, 831–853 (2023). [DOI] [PubMed] [Google Scholar]
- 32.Abraham, M. J. et al. GROMACS: high performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX1–2, 19–25 (2015). [Google Scholar]
- 33.Ivani, I. et al. Parmbsc1: a refined force field for DNA simulations. Nat. Methods13, 55–58 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Nayar, D., Agarwal, M. & Chakravarty, C. Comparison of tetrahedral order, liquid state anomalies, and hydration behavior of mTIP3P and TIP4P water models. J. Chem. Theory Comput.7, 3354–3367 (2011). [DOI] [PubMed] [Google Scholar]
- 35.Darden, T., York, D. & Pedersen, L. Particle mesh Ewald: an N⋅log(N) method for Ewald sums in large systems. J. Chem. Phys.98, 10089–10092 (1993). [Google Scholar]
- 36.Hess, B., Bekker, H., Berendsen, H. J. C. & Fraaije, J. G. E. M. LINCS: a linear constraint solver for molecular simulations. J. Comput. Chem.18, 1463–1472 (1997). [Google Scholar]
- 37.Hunter, J. D. Matplotlib: a 2D graphics environment. Comput. Sci. Eng.9, 90–95 (2007). [Google Scholar]
- 38.Qi Biodesign, Inc. Homo sapiens (human): NGS raw sequencing data. BioProjectwww.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1147008 (2025).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Figs. 1–7, Supplementary Tables 1–3 and Supplementary Note (sequences of mini-Sdd9-BE4max, mini-Sdd9-QBEmax and mini-Sdd9-QBEn).
sgRNA list of the 5 CAR-T genes.
Data Availability Statement
The deep amplicon sequencing data have been deposited in the NCBI BioProject database under accession code PRJNA1147008 (ref. 38). Plasmids encoding QBEmax and QBEn are available at Addgene. All other data are available in the main paper or Supplementary Information.












