Abstract
Advancing the performance of programmable genome editing nucleases remains a key challenge in expanding their research and therapeutic applications. Here, we introduce a scalable deep learning–guided protein engineering framework for improving nuclease activity without requiring experimental training data. As a demonstration, we apply this strategy to SpuFz1, a compact Fanzor nuclease of eukaryotic origin, identifying and validating beneficial mutations that produces a multi-mutant variant with an 11.6-fold increase in editing efficiency. In parallel, we use comparative sequence analysis to design and experimentally validate a 75-nt ultrashort ωRNA scaffold, reducing guide RNA length by 79% while maintaining activity. Integration of these optimized components yields enFanzor, a compact genome editing system that achieves editing efficiencies up to 81.9% in mammalian cells, with strong editing performance in both human hematopoietic stem and progenitor cells (HSPCs) and mouse embryos. The outperforming variant developed through this strategy also supports robust CBE and ABE activity. Notably, the shortened ωRNA not only improves nuclease editing specificity but also leads to a substantial increase in base editing efficiency. Together, this work demonstrates the power of combining AI-guided protein optimization with rational RNA design, and establishes a generalizable strategy for engineering next-generation genome editing tools.
Subject terms: Synthetic biology, Molecular engineering, Protein design
How to quickly and systematically advance the activity of programmable RNA-guided endonucleases remains a key challenge to be overcome. Here the authors combined deep learning-guided protein optimization with rational RNA design, generating enFanzor system which supports robust genome editing.
Introduction
The success of CRISPR-Cas systems has catalyzed a revolution in genome editing, enabling precise and programmable manipulation across a wide range of organisms and applications1–4. However, most CRISPR-Cas nucleases originate from prokaryotes5, and their large molecular size and immunogenicity pose persistent challenges for in vivo delivery and therapeutic translation6–9. These limitations have motivated the search for compact, programmable genome editors that are more delivery-compatible and less immunogenic.
The OMEGA (obligate mobile element-guided Activity) system has recently emerged as an alternative class of RNA-guided nucleases, comprising proteins such as TnpB, IscB, and Fanzor (Fz)10–12. These systems are thought to represent evolutionary predecessors of CRISPR-Cas enzymes13,14. Among them, Fanzor stands out as the only known OMEGA nuclease of eukaryotic origin11,15,16. Fz proteins are highly compact (~600 amino acids), lack trans-cleavage activity, and exhibit potentially low immunogenicity17. Notably, compared to larger nucleases such as SpCas9 (1368 amino acids) and SaCas9 (1053 amino acids), the small size of Fanzor enables efficient packaging into a single adeno-associated virus (AAV) vector, which is constrained to ~4.7 kb for effective delivery, thereby enhancing their suitability for in vivo genome editing applications. However, their native activity in mammalian cells remains insufficient for practical use11, necessitating further optimization.
Traditional protein engineering strategies, such as X > R/K scanning mutagenesis, aim to enhance nuclease activity by increasing surface positive charge and improving nucleic acid binding. This method has shown partial success in other compact nucleases like Cas12i18,19 and IscB20,21, but yielded only modest improvements in SpuFanzor1 (SpuFz1)11, suggesting that charge-focused mutagenesis alone is insufficient to fully exploit its editing potential. To overcome these limitations, we develop a scalable, deep learning–based protein engineering framework capable of predicting high-activity variants of genome editing nucleases without the need for experimental training data. As a proof of concept, we apply this strategy to SpuFz1, using our Fanzor-Fitness Predictor to identify and prioritize single mutations in silico. Combinatorial assembly of experimentally verified top-ranked single substitutions yields a multi-mutant variant, SpuFz1-v3, with an 11.6-fold increase in indel formation relative to wild-type. These results demonstrate the utility of data-efficient AI-driven design in overcoming the limitations of conventional engineering approaches.
Beyond protein engineering, the guide RNA component of the Fanzor system also presents a bottleneck to practical application. Previous designs aimed to truncate the ωRNA scaffold for improved delivery and synthesis11, but editing efficiency remained variable across contexts. To systematically address this, we conduct a comparative analysis of naturally occurring ωRNA variants, followed by engineering design. This leads to the identification of a 75-nt ultrashort scaffold, gst28.m2, which retains high editing efficiency while reducing RNA length by 79% compared with the previously engineered 275ext_75 scaffold (350 nt)11. Its compact size supports cost-effective synthesis and enhances compatibility with delivery vectors. Intriguingly, GUIDE-seq analysis reveals that the engineering of ωRNA simultaneously improves editing specificity.
The system also supports both ABE and CBE modalities and, unexpectedly, the use of an ultrashort ωRNA substantially enhances base editing efficiency, further highlighting its versatility and design potential.
We combine SpuFz1-v3 with gst28.m2 to create enFanzor, a compact and efficient eukaryotic genome editing system. enFanzor demonstrates robust activity in HEK293T cells, outperforms the engineered miniature enIscB at most tested sites, and shows strong editing performance in both human hematopoietic stem and progenitor cells (HSPCs) and mouse embryos.
In summary, we present a generalizable AI-guided framework for optimizing genome editing nucleases, which we validate by substantially enhancing the activity of the compact Fanzor nuclease SpuFz1. In parallel, we experimentally identify and optimize an ultrashort ωRNA scaffold that further boosts editing efficiency and specificity. Together, these advances establish a deep learning-guided protein engineering framework, complemented by systematic guide RNA discovery and optimization, enabling the development of compact and programmable genome editors and expanding the toolkit for precise and scalable manipulation of gene function in physiologically relevant systems.
Results
AI-guided engineering of Fanzor to enhance genome editing activity
RNA-guided nuclease systems have evolved diverse mechanisms and structural features, with recent analyses suggesting that TnpB—a central component of the OMEGA system—represents a direct evolutionary precursor to Cas12 nucleases in CRISPR systems10. Within this evolutionary context, Fanzor stands out as the only well-characterized eukaryotic member of the OMEGA family, representing a lineage that bridges prokaryotic and eukaryotic genome editing systems11. OMEGA nucleases such as TnpB and Fanzor possess minimalist architectures, retaining essential catalytic domains (e.g., RuvC) while lacking many accessory features present in Cas nucleases. This results in significantly smaller protein sizes, which improves compatibility with delivery platforms. For example, AsCas12a is 1307 amino acids in length22, while TnpB12 and SpuFz111 are 408 and 638 amino acids, respectively (Fig. 1a). Some analyses suggest that RNA-guided nuclease systems have evolved with increasing protein complexity alongside progressive simplification of the guide RNA23,24. From this perspective, the Fanzor system occupies an evolutionarily intermediate position, balancing protein size and RNA length. Its compact protein and concise guide RNA offer dual advantages: ease of protein delivery and compatibility with cost-effective, chemically synthesized guide RNAs. Nevertheless, despite recent engineering efforts, the editing efficiency of the Fanzor system remains suboptimal, posing a major barrier to its practical application as a genome-editing tool.
Fig. 1. Development of the Fanzor-Fitness Predictor for SpuFz1 engineering.

a Simplified loci and domain/lobe organization of AsCas12a, ISDra2 TnpB, and SpuFz1, based on information from refs. 45,46. Protein lengths are shown in amino acids; elements are color-coded and not drawn to scale. b Schematic overview of the Fanzor-Fitness Predictor workflow. The left panel illustrates the dataset construction process for model training: multiple sequence alignment (MSA) data of the protein is generated, flattened, and encoded prior to input. The middle panel depicts the architecture and training workflow of the deep learning model, which consists of an encoder–decoder structure. The right panel outlines the inference phase, highlighting the computational logic used to predict the fitness scores of protein mutants. c Heatmap showing the predicted fitness landscape of single amino acid substitutions across the SpuFz1 protein. Each row corresponds to a predictable amino acid position, and each column represents a substituting amino acid residue. The color intensity indicates the predicted fitness change relative to the wild-type residue at that position. To enhance the visualization of beneficial mutations, all fitness values below zero were set to zero. Black blocks indicate the wild-type amino acid at each respective site. d Schematic of the experimental workflow for detecting Fanzor cleavage activity using the EGxxFP reporter system in HEK293T cells. NLS nuclear localization signal, DSB DNA double-strand break. e Flow cytometry scatter plots showing representative results from two SpuFz1 protein variants, analyzed using the EGxxFP reporter system. f Cleavage activities of the top 98 single-point mutants predicted by the Fanzor-Fitness Predictor, as measured by the fluorescent reporter assay. The horizontal dashed line denotes the mean editing efficiency of SpuFz1-v2, used here as a reference. Data represent mean ± s.d., n = 3 biological replicates. g Editing efficiencies of combinatorial variants generated by combining the EGxxFP-identified top five single-point mutations. Genome editing of 32 combinatorial variants at the B2M locus was assessed 72 h post-transfection by amplicon deep sequencing. The horizontal line indicates the mean editing efficiency of wild-type SpuFz1, shown as a reference. Data represent mean ± s.d., n = 3 biological replicates.
Current engineering strategies for gene-editing nucleases primarily include deep mutational scanning (DMS)25, X > R/K charge-scanning, and structure-guided rational design. While these methods differ in complexity and throughput, there is typically a trade-off between screening workload and the breadth of protein sequence space explored—more targeted approaches tend to evaluate fewer variants. In the case of SpuFz1, engineering efforts using X > R/K scanning have already been conducted. However, the resulting improvements in gene-editing efficiency remain limited. For example, in HEK293T cells, the highest average indel rate reached only 18.4% across 12 endogenous target sites11. This outcome suggests that restricting exploration to a relatively small subset of variants is insufficient to achieve substantial performance enhancement, highlighting the need for broader and more systematic engineering approaches.
To more efficiently explore the mutational landscape of SpuFz1, we developed a deep learning–driven zero-shot prediction model for protein engineering. This approach enables rapid evaluation of a large number of unscreened variants prior to experimental screening and characterization, effectively simulating a virtual deep mutational scanning and reducing the randomness inherent in conventional experimental approaches. The strategy begins with the construction of a high-quality multiple sequence alignment (MSA) for the target protein, followed by training a variational autoencoder (VAE)–based mutation effect prediction model using the DeepSequence framework26. DeepSequence has demonstrated strong performance and robustness across diverse benchmark datasets27. Compared to large protein language models (lLMs) such as ESM28, DeepSequence has substantially lower computational requirements, which facilitates repeated local execution in iterative protein engineering workflows. During training, the model learns coordinated residue variation patterns from the MSA, capturing evolutionary constraints in a low-dimensional latent space. In the inference phase, it performs high-throughput fitness prediction for all possible amino acid substitutions and calculates the evolutionary index (Eᵥ) relative to the wild-type sequence. Variants with high predicted activity are then prioritized for experimental validation (Fig. 1b).
Although DeepSequence has been successfully applied to various protein fitness prediction tasks—such as enhancing enzyme activity and binding affinity27—its applicability to RNA-guided nuclease systems remains largely unexplored. To evaluate the suitability of this approach for engineering OMEGA effectors, we first validated the predictive performance of DeepSequence using the OgeuIscB system, a recently reported prokaryotic OMEGA nuclease10 that was optimized via an X > R charge-scanning strategy to function efficiently in mammalian cells. We constructed a mutation effect prediction model for the OgeuIscB protein and performed full model inference (Supplementary Fig. 1a, b). Predicted fitness scores were then compared with experimentally measured indel activities reported in previous X > R mutagenesis studies20. The predicted scores showed a strong concordance with the experimentally measured editing activities, with a Spearman correlation coefficient of 0.68 and a Pearson correlation coefficient of 0.66 (Supplementary Fig. 1a), indicating that the model reliably captures overall mutational effects. In addition, ranking-based evaluations demonstrated robust performance in prioritizing high-activity variants: Precision@k analysis showed a high fraction of beneficial mutations among top-ranked predictions, while NDCG@k values remained close to 0.9 across a wide range of k, reflecting consistent agreement between predicted rankings and experimental outcomes. This result provides supporting evidence for the predictive and practical power of deep learning–based approaches in the engineering of RNA-guided nucleases, and highlights their potential as a viable alternative to conventional mutagenesis strategies for optimizing OMEGA effector proteins.
After validating the reliability of our approach, we applied the deep learning–based prediction strategy to the Fanzor system. We trained the Fanzor-Fitness Predictor and used it to infer the fitness landscape of saturation mutations spanning the majority of the SpuFz1 protein sequence—covering 392 amino acid positions from residues 151 to 616, which represent 61.44% of the total sequence (Fig. 1c). Notably, the model identified several novel mutational hotspots predicted to enhance SpuFz1 activity, which differed from those previously revealed by the X > R/K scanning strategy (Supplementary Fig. 2). This divergence indicates that our deep learning–based method can overcome the inherent limitations of traditional charge-scanning approaches and capture a broader set of potentially beneficial mutations across the Fanzor protein sequence space. These findings provide a valuable resource for guiding the design of next-generation Fanzor variants with enhanced genome-editing performance.
Subsequently, we ranked all 7448 predicted single-point mutations by their fitness improvement scores for SpuFz1 and selected the top 98 variants for experimental validation (excluding two mutations at position 513 already present in SpuFz1-v2 from the original top 100 list). To evaluate editing activity, we established a three-color fluorescent reporter system: red fluorescent protein (mCherry) was co-expressed with the editor, while blue fluorescent protein (BFP) was co-expressed with the ωRNA. The BFP construct also carried two partially overlapping EGFP coding sequences separated by the editing target region—an arrangement termed EGxxFP29. When double-strand breaks (DSBs) are introduced at the target site, homologous recombination between the two EGFP fragments restores the open reading frame, leading to the expression of green fluorescence. We assessed the functional impact of each predicted mutation by measuring the percentage of EGFP-positive cells within the mCherry⁺/BFP⁺ double-positive population (Fig. 1d, e). Based on this assay, 67 out of 98 variants (68.37%) showed enhanced editing activity relative to SpuFz1-v2, confirming the effectiveness and high predictive success rate of our deep learning–based strategy for Fanzor protein engineering (Fig. 1f).
To further improve SpuFz1 activity, we selected the top five experimentally validated single mutations (one per site, selecting the best variant at each) and performed full combinatorial assembly with the previously reported SpuFz1-v2. These combinations were evaluated using a minimal 81-nt ωRNA scaffold—comprising a 52-nt truncated ωRNA and a 5′-terminal 29-nt MS2 stem-loop—along with a 20-nt guide targeting the endogenous B2M locus in HEK293T cells. Among the tested combinations, the variant carrying C318K, C373L, E425K, and A469I—designated SpuFz1-v3—demonstrated the highest indel activity at the B2M locus, achieving an average editing efficiency of 35.52 ± 0.59%, representing an 11.59-fold improvement over wild-type SpuFz1 (3.07 ± 0.12%) and a 3.51-fold increase over SpuFz1-v2 (10.13 ± 0.46%) (Fig. 1g). Collectively, these results demonstrate that the deep learning-guided Fanzor-Fitness Predictor enables high-accuracy identification of activity-enhancing mutations. Furthermore, combining top-performing single-point mutations resulted in the generation of SpuFz1-v3, a variant with significantly improved editing efficiency at endogenous genomic loci.
Based on the resolved structure of the SpuFz1–ωRNA–target DNA complex (PDB: 8gkh), we performed structure-based analyses of high-activity variants identified by the DeepSequence model (Supplementary Fig. 3). Several mutations are predicted to enhance editing activity, potentially by strengthening protein–nucleic acid interactions. For instance, C318, located in a loop region of the WED domain, does not contact nucleic acids in the wild-type structure; however, substitution with lysine or arginine is predicted to introduce a positively charged side chain that may enable formation of a salt bridge with A75 in the ωRNA scaffold (Supplementary Fig. 3b). Similarly, the E425K mutation in the RuvC domain may reinforce polar interactions with the DNA–RNA heteroduplex, potentially stabilizing the catalytic complex (Supplementary Fig. 3c). In contrast, mutations such as W473M, A469I, and C373L are less likely to act primarily through direct nucleic acid binding and may instead modulate local conformational flexibility within the RuvC domain. In the wild-type structure, W473 forms a π–π stacking interaction with F416, contributing to stabilization of two adjacent helices (Supplementary Fig. 3d). Replacement of W473 with methionine would be expected to weaken this aromatic interaction, potentially increasing interhelical flexibility, which may facilitate contacts with the DNA–RNA heteroduplex or ωRNA scaffold and enhance cleavage activity. Consistent with this hypothesis, enforcing stronger interhelical constraints at W473 and F416 through engineered putative noncovalent or covalent interactions led to reduced SpuFz1 cleavage activity (Supplementary Fig. 3e). These analyses suggest that deep learning-assisted protein engineering can introduce or enhance structural interactions and may enable exploration of sequence variants beyond those readily identified through structure-guided approaches.
Editing efficiency evaluation and further optimization of SpuFz1 and its mutants
To objectively assess the impact of our protein engineering strategy, we selected 12 endogenous genomic target sites in HEK293T cells and directly compared the editing activities of SpuFz1-v3, wild-type SpuFz1, and the previously reported engineered variant SpuFz1-v2. Across all tested loci, SpuFz1-v3—generated via deep learning-guided engineering—demonstrated consistently superior editing performance relative to both SpuFz1-v2 and the wild-type protein. Specifically, SpuFz1-v3 exhibited editing efficiencies up to 11.59-fold higher than the wildtype, with an average improvement of 7.39-fold (Fig. 2a). Compared to SpuFz1-v2, which was previously optimized using X > R/K charge-scanning and represented the best-performing variant to date, SpuFz1-v3 achieved up to 4.57-fold and 2.58-fold average higher editing efficiency (Fig. 2a). These results further validate the effectiveness of our Fanzor-Fitness Predictor in enhancing SpuFz1 activity. Importantly, these results demonstrate that deep learning–based approaches can uncover potent variants by characterizing only a limited number of candidates, surpassing traditional X > R/K scanning strategies. The latter explores only a narrow subset of the mutational landscape and rely primarily on modulating electrostatic interactions between the nuclease and target DNA or guide RNA—thereby underutilizing the broader mutational potential inherent in the protein sequence space.
Fig. 2. Engineering strategies to enhance SpuFz1 activity.

a Comparison of indel editing efficiencies at 12 endogenous genomic loci in HEK293T cells for wild-type SpuFz1, SpuFz1-v2, and the optimized SpuFz1-v3. Editing efficiencies were quantified by amplicon-based deep sequencing. Each point represents an individual biological replicate. Fold changes above each bar denote the mean efficiency of SpuFz1-v3 relative to SpuFz1 at each locus. Data represent mean ± standard deviation (s.d.), n = 3 independent biological replicates. b Schematic diagrams of SpuFz1 expression constructs with additional fusion domains, including the HMG-D DNA-binding domain from the Drosophila melanogaster high-mobility group protein family, nuclear localization signal (NLS), and the SARD linker. c Comparison of indel editing efficiencies at endogenous loci for SpuFz1-v3 with HMG-D domain fused at either the C-terminus. Editing was assessed by amplicon deep sequencing. Data represent mean ± s.d., n = 3 biological replicates. d Comparison of indel editing efficiencies for SpuFz1-v3 with C-terminal fusion of either an NLS or a SARD linker at endogenous loci in HEK293T cells. Editing efficiencies were determined by deep sequencing. Data represent mean ± s.d., n = 3 biological replicates.
Previous studies have reported that fusing additional domains—such as DNA 5′ exonuclease domains (e.g., T5 exonuclease, 291 amino acids)20 and DNA minor groove–binding domains (e.g., HMG-D, 111 amino acids)21,30—can enhance the editing efficiency of certain CRISPR and OMEGA effector proteins. To investigate whether such strategies could further improve the activity of SpuFz1-v3, we evaluated the effects of fusing T5E and HMG-D domains to either the N- or C-terminus of SpuFz1-v3 (Fig. 2b and Supplementary Fig. 4a). We found that T5E fusion at either terminus reduced editing efficiency, suggesting that augmenting SpuFz1-v3 with an additional nuclease domain is not a viable strategy for enhancing its function (Supplementary Fig. 4b). In contrast, the impact of HMG-D fusion varied depending on the fusion orientation. N-terminal fusion of HMG-D led to no improvement or a slight decrease in editing activity (Supplementary Fig. 4b). However, when HMG-D was fused to the C-terminus, a modest enhancement in editing efficiency was observed at multiple target sites (Fig. 2c and Supplementary Fig. 4b). Specifically, editing efficiency at the CA2 site increased from 33.95 ± 0.43% to 40.36 ± 4.06%, at the HPRT site from 25.33 ± 2.03% to 33.69 ± 1.69%, and at the IFNG site from 20.44 ± 0.15% to 29.12 ± 0.47% (Fig. 2c). While the addition of exogenous domains increases the overall molecular size of the editor, these results suggest that C-terminal fusion with DNA-binding domains may represent a promising and site-specific optimization strategy for improving SpuFz1-v3 activity—particularly when delivery constraints are relaxed.
Previous studies have demonstrated that fusing a basic nuclear localization signal (NLS) can enhance the nuclear import of nucleases, thereby increasing the efficiency of nuclease-mediated genome editing6,31. In this study, we evaluated two SpuFz1-v3 constructs with different NLS configurations. The first followed our previous design, incorporating two NLS sequences at the C-terminus32. The second, inspired by another study33, employed a SRAD linker (serine–arginine–alanine–aspartic acid) containing basic residues to improve the connection between the NLS and the Fanzor protein, with the goal of enhancing nuclear localization and, consequently, editing efficiency. Cell-based editing assays revealed that the addition of a supplementary NLS improved gene-editing efficiency at most tested sites, with the SRAD-linked construct exhibiting slightly greater enhancements overall (Fig. 2d and Supplementary Fig. 4b). Notably, improvements were more pronounced at the CA2, CXCR4, and HPRT3 loci. Although SpuFz1 contains a short N-terminal peptide with intrinsic nuclear localization function, these findings highlight that both the presence and positioning of NLS sequences—and the nature of the linker used—can meaningfully influence editing performance, providing a straightforward yet effective strategy to enhance the functional utility of Fanzor nucleases.
ωRNA scaffold engineering to generate a compact size with robust activity
Previous studies have shown that, when co-expressed with wild-type SpuFz1 and tested at the B2M target site, a 75-nt canonical ωRNA scaffold with a 5′ 275-nt extension (referred to as 275ext_75) exhibits comparable insertion/deletion (indel) activity to a 75-nt scaffold bearing a 5′ 29-nt MS2 stem-loop extension (MS_75)11. Similarly, when paired with SpuFz1-v2, a previously designed 81-nt ωRNA scaffold—comprising a 52-nt truncated ωRNA with a 5′ 29-nt MS2 stem-loop (MS_52)—showed editing efficiency similar to that of MS_75 at the same locus11. However, a direct comparison of these three ωRNA scaffolds across different SpuFz1 variants had not been systematically performed. Here, we evaluated and compared the editing efficiencies of MS_75, 275ext_75, and MS_52 scaffolds in HEK293T cells when used with both SpuFz1-v2 and our developed SpuFz1-v3. Across all conditions, SpuFz1-v3 consistently outperformed SpuFz1-v2, regardless of the ωRNA scaffold used. Notably, 275ext_75 exhibited the highest editing efficiency in both editor contexts (Fig. 3a). Specifically, at the B2M locus, the mean indel efficiencies of MS_52, MS_75, and 275ext_75, when paired with SpuFz1-v2 were 13.41, 15.57, and 41.38%, respectively (Fig. 3a). When paired with SpuFz1-v3, these values increased to 34.16, 43.00, and 62.93%, respectively (Fig. 3a). While prior studies reported no significant difference between 275ext_75 and MS_75 when used with wild-type SpuFz1, our data demonstrate that 275ext_75 consistently outperforms the other scaffolds when paired with engineered variants. In particular, its mean editing efficiency was 2.66-fold higher than MS_75 and 3.09-fold higher than MS_52 with SpuFz1-v2, and 1.46-fold and 1.84-fold higher, respectively, with SpuFz1-v3 (Fig. 3a).
Fig. 3. Optimization and evaluation of ωRNA scaffold variants for SpuFz1.

a Comparison of indel editing efficiencies at the endogenous B2M locus in HEK293T cells using SpuFz1-v2 and SpuFz1-v3, in combination with ωRNA scaffolds containing various 5′ protective sequences appended to a 52-nt core scaffold. Editing efficiency was quantified by amplicon-based deep sequencing. Data represent the mean ± standard deviation (s.d.), n = 3 independent biological replicates. b Sequence similarity analysis of all annotated ghost scaffold sequences from the SpuFz1 host genome. c Comparison of SpuFz1-v3–mediated editing efficiency at the B2M locus using the canonical ωRNA scaffold, 5’ 275-nt extension canonical ωRNA (275ext_75) scaffold and representative ghost ωRNA scaffolds. Editing activity was assessed by amplicon sequencing. Data represent mean ± s.d., n = 3 biological replicates. P values from two-tailed Student’s t-tests are indicated above the corresponding data points. d Predicted secondary structures of the canonical ωRNA scaffold (left) and gst28 ωRNA (middle). The right panel illustrates the sequence modifications made to gst28 ωRNA, with matching nucleotides color-coded to reflect corresponding structural regions. e Editing efficiencies of modified gst28 ωRNA variants at the endogenous B2M locus. The horizontal line represents the mean editing efficiency of unmodified wild-type gst28 ωRNA, used as a reference. Editing was quantified by amplicon-based deep sequencing. Data represent mean ± s.d., n = 3 independent biological replicates. Statistical significance was assessed using a two-tailed unpaired Student’s t-test. f Predicted complex structures of SpuFz1-v3 protein bound to either the 275ext_75 ωRNA scaffold (350 nt) or the gst28.m2 ωRNA scaffold (75 nt). In the structural models, ωRNA is shown in blue, and the SpuFz1-v3 protein is shown in light gray. Structural predictions were generated using the AlphaFold3 web server.
These results indicate that the performance of different ωRNA scaffold variants is context-dependent and varies with the specific SpuFz1 editor variant used. Notably, the absence of a significant difference between 275ext_75 and MS_75 when used with wild-type SpuFz1, and between MS_75 and MS_52 when used with SpuFz1-v2, does not imply that MS_52 and 275ext_75 will behave similarly under SpuFz1-v2 conditions. Our findings clearly demonstrate that, to maximize editing efficiency, 275ext_75 should be prioritized. When paired with either SpuFz1-v2 or SpuFz1-v3, 275ext_75 significantly outperforms MS_52, the compact scaffold previously proposed as a final optimized version11 (Fig. 3a). However, this performance gain comes with a trade-off: at 350 nucleotides, 275ext_75 is 4.32 times the length of MS_52 (81 nucleotides), posing considerable challenges for chemical synthesis and in vivo delivery. These findings underscore a persistent tension in guide RNA design: the need to balance maximal editing efficiency with minimal scaffold length—a conflict that remains unresolved and central to future optimization efforts.
To improve the delivery efficiency, in vivo stability, and chemical synthesis feasibility of the SpuFz1 gene-editing system, miniaturization of the ωRNA scaffold has emerged as a key engineering strategy. A streamlined ωRNA not only facilitates packaging into viral vectors but also enhances expression efficiency and supports diverse chemical modifications. However, the current 350-nt ωRNA scaffold presents a significant barrier to synthesis, as efficient RNA synthesis methods typically limit transcript length to ~100 nucleotides. To address this, we first explored replacing the original 275-nt 5′extension with shorter sequences. Specifically, we integrated 15 compact structured RNA motifs—ranging mainly from 10 to 20 nucleotides and derived from previously reported engineered pegRNAs34—into the 5′ end of a 52-nt truncated canonical ωRNA scaffold. These designs aimed to minimize total RNA length while maintaining or improving functional activity. However, screening results showed that none of the mini structural motifs significantly enhanced ωRNA activity compared to the 5′ 275-nt extended canonical ωRNA scaffold (275ext_75)—regardless of whether the scaffold was paired with SpuFz1-v2 or SpuFz1-v3 (Fig. 3a). These findings suggest that shortening the 5′extension via structural motif substitution alone is insufficient to replicate the functional enhancement provided by the original 275-nt 5′extension.
In addition, we observed that the indel efficiency of the 275ext_75 ωRNA scaffold was significantly higher than that of the 275-nt extension +52-nt trimmed canonical scaffold (275ext_52). Specifically, editing efficiency reached 41.38 vs. 21.33% when combined with SpuFz1-v2—a 1.94-fold increase—and 62.93 vs. 50.00% when combined with SpuFz1-v3—a 1.26-fold increase (Fig. 3a). Under conditions where the same 5′ MS extension was applied, the performance difference between the 75-nt canonical scaffold and the 52-nt trimmed scaffold was context-dependent. In line with previous findings, when paired with SpuFz1-v2, the difference between MS_75 and MS_52 was not statistically significant (Fig. 3a). However, when combined with SpuFz1-v3, MS_75 exhibited a 1.26-fold higher efficiency than MS_52 (43.00 vs. 34.16%) (Fig. 3a). These results suggest that when the same 5′ sequence extension is applied, the longer canonical ωRNA scaffold tends to outperform its trimmed counterpart in the context of SpuFz1-v3, but not necessarily with SpuFz1-v2. This implies that editor-scaffold compatibility plays a critical role in determining editing efficiency, and that truncation of the ωRNA scaffold may reduce performance when paired with more active editor variants.
On the other hand, previous studies have reported no significant difference in editing efficiency between the 275ext_75 scaffold and several 75-nt ghost ωRNA scaffolds when used with wild-type SpuFz111. However, further investigation into these ghost ωRNA scaffolds has remained limited. Notably, the SpuFz1 host genome harbors a large number of such ghost ωRNA sequences, with over 100 annotated to date11. These ghost scaffolds may represent naturally evolved variants, potentially offering a rich source of structural diversity for ωRNA engineering—prompting our interest in their systematic characterization. We extracted all annotated ghost backbone sequences from the SpuFz1 host genome and conducted a sequence similarity analysis (Fig. 3b and Supplementary Data 1). Excluding ghost1, whose sequence has already been characterized in prior studies11, we collated 126 unique sequences (designated gst1–gst126), ranging in length from 74 to 76 nucleotides, with the majority being 75 nt in length. A phylogenetic dendrogram based on sequence similarity revealed clear clustering patterns: approximately one-third of the sequences grouped with the canonical and ghost1 scaffolds, about half formed a distinct second clade, and the remainder fell into several smaller, more divergent clusters (Fig. 3b). We selected representative sequences from each major clade for synthesis and constructed corresponding expression plasmids. These constructs were transfected into HEK293T cells to evaluate their editing performance. Our results showed that ghost scaffolds with high sequence similarity to ghost1 and the canonical ωRNA (e.g., gst2, gst9, gst25, gst28, gst38, gst43, and gst122) generally exhibited higher editing efficiency than more distantly related sequences (gst7, gst8, gst16, gst73, and gst98) (Fig. 3c). Among all tested variants, gst28 demonstrated the highest mean editing efficiency and was therefore selected as the basis for further engineering (Fig. 3c). Notably, gst28 achieved the comparable editing activity to the 350-nt 275ext_75 scaffold (p = 0.2623) in this assay, highlighting its potential as a compact and efficient ωRNA scaffold for SpuFz1-based genome editing (Fig. 3c).
Building on the predicted secondary structure of gst28 ωRNA, we explored several engineering strategies aimed at enhancing its functionality (Fig. 3d). These included increasing stem length to improve structural stability and introducing base substitutions to disrupt poly-T stretches, thereby enhancing transcription efficiency (Fig. 3d). Through these modifications, we successfully identified an optimized variant, gst28.m2, which exhibited superior editing efficiency (Fig. 3e). In summary, by systematically screening natural ωRNA backbone variants and applying rational structural engineering, we developed gst28.m2—a compact yet highly effective ωRNA scaffold. At only 75 nucleotides, gst28.m2 is 79% shorter than the previously most efficient 275ext_75 scaffold (350 nt), significantly improving its suitability for chemical synthesis, viral packaging, and in vivo delivery, and thereby expanding the practical utility of the SpuFz1 genome editing system (Fig. 3f).
Characterization and comparison of enFanzor
By engineering both components of the SpuFz1 gene-editing system, we developed a more active protein variant, SpuFz1-v3, and a compact yet efficient ωRNA scaffold, gst28.m2, without compromising editing performance. We combined these two elements to construct an enhanced Fanzor editing system, termed enFanzor, with significantly improved functionality in mammalian cells. To further validate the performance of the gst28.m2 scaffold, we systematically compared its editing efficiency to that of the 350-nt 275ext_75 ωRNA scaffold when paired with SpuFz1-v3 across multiple endogenous loci in HEK293T cells. Overall, enFanzor exhibited comparable activity to the SpuFz1-v3 + 350-nt ωRNA combination, despite the substantial reduction in RNA length (Fig. 4a). These results further underscore the efficacy and practical advantages of the enFanzor system and highlight its potential for scalable and efficient genome editing applications in mammalian systems.
Fig. 4. Evaluation of enFanzor performance across multiple endogenous loci and in comparison with other genome-editing systems.

a Comparison of indel editing efficiencies at 12 endogenous genomic loci in HEK293T cells using SpuFz1-v3 in combination with either the gst28.m2 ωRNA scaffold or the 275-extension canonical ωRNA scaffold (275ext_75). Editing efficiencies were quantified by amplicon deep sequencing. Data represent mean ± standard deviation (s.d.), n = 3 independent biological replicates. b Schematic illustration of TAM/PAM-matched sites used for head-to-head comparisons between SpuFz1-v3 and SpCas9 (top), and between SpuFz1-v3 and enIscB (bottom). Colored nucleotide blocks indicate the respective TAM/PAM recognition sequences for each editor. c Comparison of indel editing efficiencies at 20 endogenous loci in HEK293T cells using enFanzor (SpuFz1-v3 + gst28.m2) and SpCas9. Data were presented as mean ± s.d., n = 3 biological replicates. d Summary dot plots showing the activity of enFanzor and SpCas9 in HEK293T cells. Summary dot plots on the right show overall performance across all tested sites. All data were presented as means ± s.d., n = 60. e Comparison of indel editing efficiencies at 24 endogenous loci using enFanzor and enIscB in HEK293T cells. Data are presented as mean ± s.d., n = 3 biological replicates. f Summary dot plots on the right show overall performance across all tested sites. Data were presented as mean ± s.d., n = 72. g–k Sequences of off-target sites identified by GUIDE-seq for identification of the impact of SpuFz1 protein engineering and ωRNA scaffold engineering on the specificity of the Fanzor editing system. Data were shown for five groups of Fanzor systems targeting the CA2 site: g SpuFz1-v3 paired with ωRNA scaffold in 275ext_75 format, h SpuFz1-v3 paired with ωRNA scaffold in MS_52 format, i SpuFz1-v3 paired with ωRNA scaffold in engineered gst28.m2 format, j SpuFz1-v2 paired with ωRNA scaffold in gst28.m2 format, k wild-type SpuFz1 paired with ωRNA scaffold backbone in gst28.m2 format.
We next conducted a head-to-head comparison of the enFanzor system with other established genome-editing platforms, including the widely used SpCas96 system and the recently reported enhanced enIscB20 system. Notably, while SpCas9 and enIscB recognize PAM/TAM sequences at the 3′ end of the spacer, enFanzor targets a TAM at the 5′ end, facilitating the design of shared test sites for direct performance comparison (Fig. 4b). We designed 20 target sites across 12 human genes that met the sequence requirements of two systems—TAM (CATA) for enFanzor and PAM (NGG) for SpCas9—and evaluated editing efficiencies across these loci in parallel (Fig. 4b). Experimental results demonstrated that enFanzor achieved robust editing activity at multiple sites. In particular, at HRNF2.site1, HRNF2.site2, IFNG.site1, VEGFA.site1, and HEK2.site2, enFanzor’s performance was comparable to or even exceeded that of SpCas9 (Fig. 4c, d). While SpCas9 generally outperformed enFanzor at most tested loci, the ability of enFanzor to deliver efficient editing at specific sites, combined with its substantially smaller molecular size, highlights its potential as a compact and versatile genome-editing platform. These findings underscore enFanzor’s promise for targeted genome modification, particularly in applications where payload size and delivery constraints are critical.
We further compared the editing efficiency of enFanzor with that of the previously reported miniaturized OMEGA system editor enIscB20 (496 amino acids) across 24 target sites in 13 human genes. These sites were selected based on compatibility with both systems’ TAM recognition requirements—CATA for enFanzor and NWRRNA for enIscB (Fig. 4b). Strikingly, at the 24 shared TAM-compatible sites, enFanzor outperformed enIscB at 22 loci, demonstrating a clear advantage in editing efficiency (Fig. 4e, f). In particular, at CA2.site2, DYRK1A.site4, and EMX1.site2, enIscB exhibited only weak editing activity (5.97 ± 1.19%), whereas enFanzor achieved robust editing efficiencies (73.19 ± 5.68%), representing a more than 12-fold increase (Fig. 4e). These results demonstrate that within co-targetable regions, enFanzor is markedly more effective at inducing gene modifications compared to enIscB. This highlights enFanzor’s strong potential as a compact, high-performance genome-editing platform.
In addition, we evaluated the editing specificity of enFanzor across multiple genomic loci in HEK293T cells using both targeted deep sequencing and GUIDE-seq analyses. Across representative target sites, including B2M, CA2, and CXCR4, enFanzor exhibited minimal off-target editing at loci predicted by Cas-OFFinder35, with off-target events detected at only a small fraction of the examined sites and at substantially lower frequencies than on-target editing (Supplementary Fig. 5a). Moreover, direct GUIDE-seq36 comparisons at the same target site revealed that enFanzor generated fewer detectable off-target sites than another engineered OMEGA nuclease, enIscB, under matched conditions (Supplementary Fig. 5b–e).
Furthermore, based on the same genomic target site (CA2), we systematically compared the effects of SpuFz1 protein engineering and ωRNA scaffold engineering on Fanzor editing specificity using five parallel GUIDE-seq assays (Fig. 4g–k). When the SpuFz1-v3 protein was used, different ωRNA scaffolds showed distinct off-target profiles. Both the 275ext_75 and MS-52 scaffolds exhibited frequent off-target events, and in some cases, the most abundant off-target reads exceeded the corresponding on-target reads. In contrast, the engineered gst28.m2 ωRNA scaffold reduced both the number and abundance of off-target sites while maintaining higher levels of on-target editing reads.
Under the same gst28.m2 ωRNA scaffold conditions, further comparison of different protein variants showed that SpuFz1-v2 and wild-type SpuFz1 produced only very low levels of on-target editing, with no off-target signals detected. This observation was accompanied by their overall lower editing activity, which is below the effective detection threshold of GUIDE-seq37. The above results indicate that our engineering of both protein and RNA acted synergistically to improve the editing efficiency and specificity of the system. Together, these results indicate that enFanzor displays a favorable specificity profile across multiple loci and detection platforms, supporting its suitability for precise genome-editing applications.
Establishing a compact base editing platform based on Fanzor and ultrashort ωRNA optimization
Although IscB and Cas12f family proteins have been successfully engineered into compact base editors20,38, no compact base editor based on Fanzor proteins has been reported to date. To explore the feasibility of Fanzor as a base editing platform, we established a reporter system containing an EGxxFP fluorescent reporter gene with either a single target site or two identical TAM-flanked target sites spaced 15 bp apart (Fig. 5a), enabling the assessment of both double-strand and single-strand DNA cleavage activities of various Fanzor mutants.
Fig. 5. Development of Fanzor-derived compact base editors and enhancement of editing efficiency using optimized ωRNA.

a Schematic of reporter constructs used to evaluate double-strand break (DSB) or single-strand nick activities of Fanzor variants. The reporter contains either a single target site or two identical TAM-flanked targets spaced 15 bp apart. Restoration of the EGFP signal indicates cleavage at the designed target site(s). b Quantification of EGFP+ cells following transfection with SpuFz.v3 variants carrying different catalytic residue mutations. Data were presented as mean ± s.d., n = 3 biological replicates. c, d Schematic representation of Fz-ABE (top) and Fz-CBE (bottom) designs. Fz-ABE consists of a catalytically dead SpuFz1-v3 (dSpuFz1-v3) fused to TadA8e; Fz-CBE consists of the same protein scaffold fused to hAPOBEC3A(N57Q) and UGI. e A-to-G editing efficiencies at two endogenous loci (CA2 and CXCR4) using 275ext_75, MS_52 or optimized gst28.m2 ωRNA scaffold in Fz-ABE. Data were presented as mean ± s.d., n = 3 biological replicates. f C-to-T editing efficiencies at two endogenous loci (B2M and VEGFA) using Fz-CBE with 275ext_75, MS_52 or gst28.m2 ωRNA scaffold. Sequence context of endogenous target sites used for base editing assays are annotated above. TAM highlighted in green and target nucleotides in black. Data were presented as mean ± s.d., n = 3 biological replicates.
Based on the SpuFz-v3, we selected four residues (E541, D383, N385, and D606) predicted to be catalytically essential11 and individually substituted them with alanine. All four single-point mutants lost both double- and single-strand DNA cleavage activities, resulting in catalytically dead variants—a feature reminiscent of Cas12 family proteins (Fig. 5b).
We next systematically generated all 15 possible combinatorial mutants of these four residues based on SpuFz-v3 and fused each variant with the effector modules of either ABE (TadA8e39) or CBE (human APOBEC3A + UGI40), yielding a series of Fz-ABE and Fz-CBE base editors(Fig. 5c, d). We evaluated their editing activities at two endogenous ABE target sites and two endogenous CBE target sites. Among them, the D383A–E541A–D606A triple mutant exhibited the highest A-to-G editing efficiency at both ABE loci and showed the highest and second-highest C-to-T editing efficiencies at the two CBE loci, respectively (Supplementary Fig. 6a–d). Therefore, we selected this dead SpuFz-v3 variant for further characterization.
Notably, although the 275ext_75 version of ωRNA induced robust DNA cleavage at several endogenous loci, its base editing efficiency remained limited to ~5–10% at the same target sites. In addition, the previously reported 81-nt MS_52 scaffold, which was generated through structure-guided truncation, also exhibited similarly low base editing efficiency at these loci. But strikingly, when we replaced the ωRNA with the optimized ultrashort version (gst28.m2), base editing efficiency improved substantially at all tested loci. For example, at the CA2 locus, A-to-G editing at the A4 position increased from an average of 7.06% with 275ext ωRNA to 33.04% with gst28.m2—a 4.68-fold enhancement. Similarly, editing efficiencies increased from 13.09 to 38.96% at CXCR4 A6, from 14.22 to 33.78% at B2M C6, and from 4.29 to 14.09% at VEGFA C4 (Fig. 5e, f).
These findings not only demonstrate the effectiveness of our engineered ultrashort ωRNA in enhancing nuclease editing outcomes but also highlight its potential as a critical component for the development of next-generation compact base editors.
Genome editing application of enFanzor in human primary cells and mouse embryos
To date, most research on the Fanzor system has been limited to HEK293T cells, with few published reports demonstrating its effectiveness in human primary cells or mouse embryos. To address this gap, we evaluated the genome-editing capabilities of the enFanzor system in human hematopoietic stem and progenitor cells (HSPCs) and mouse embryos. HSPCs present a technical challenge for plasmid-based delivery due to their low transfection efficiency. However, they are amenable to in vitro electroporation, a method we have previously used for genome editing in both preclinical41 and clinical settings42. In this study, we established an RNA electroporation–based enFanzor delivery strategy, combining SpuFz1-v3 mRNA with a synthetic guide RNA based on the gst28.m2 scaffold. To enhance guide RNA stability, chemical modifications consisting of 2′-O-methyl and 3′ phosphorothioate (MS) groups were added to the first and last three nucleotides at both ends of the gst28.m2 ωRNA43. We electroporated purified SpuFz1 mRNA and modified ωRNA targeting the BCL11A enhancer into CD34⁺ HSPCs derived from healthy donors, and compared editing outcomes between wild-type SpuFz1 and SpuFz1-v3 (Fig. 6a, b). The results revealed that SpuFz1-v3, when combined with chemically modified gst28.m2 ωRNA, achieved a high editing efficiency of 49.24 ± 1.72%, whereas the wild-type SpuFz1 system (paired with the canonical ωRNA scaffold) showed minimal editing activity (1.31 ± 0.21%) (Fig. 6c). These findings demonstrate that the enFanzor system is capable of efficient genome editing in primary human HSPCs via mRNA electroporation, significantly expanding its potential for therapeutic genome engineering.
Fig. 6. In vivo and ex vivo genome editing using enFanzor in human HSPCs and mouse embryos.

a Schematic representation of the human BCL11A locus. The target sequence and the TAM motif are indicated by gray and green boxes, respectively. b Experimental workflow for disrupting the +58 enhancer site of BCL11A in human CD34⁺ HSPCs. Purified SpuFz1-v3 mRNA and chemically modified gst28.m2 ωRNA were delivered into cells via electroporation. c Quantification of genome editing efficiency at the BCL11A enhancer +58 site in CD34⁺ HSPCs, measured by amplicon-based deep sequencing. Data are represented as mean ± s.d., n = 3 independent biological replicates. Statistical significance was assessed using a two-tailed unpaired Student’s t-test. d Schematic of mouse Tyr exon 1 with four ωRNA target sites used for efficiency screening. The target sequence and TAM motif are shown in gray and green, respectively. e Indel frequencies in F₀ mice following cytoplasmic co-injection of SpuFz1 mRNA variants and their respective guide RNAs, measured by deep sequencing. Three groups were analyzed: SpuFz1-v2 and 275ext_75 (n = 28), SpuFz1-v3 and 275ext_75 (n = 18), and SpuFz1-v3 and gst28.m2 (n = 12). Data were presented as mean ± s.d. f Representative coat color phenotypes of F₀ pups following cytoplasmic injection of SpuFz1-v3 mRNA and gst28.m2 ωRNA. Images were captured on postnatal day 10. g Summary of microinjection results for the generation of an albino mouse model via enFanzor-mediated editing.
We further evaluated the in vivo genome-editing potential of the Fanzor system in mouse embryos by microinjecting mRNA and guide RNA to generate a phenotypic disease model. Disruption of the tyrosinase (Tyr) gene leads to albinism, making it an ideal target for phenotypic assessment of editing efficiency44. We designed four guide RNAs targeting exon 1 of the Tyr gene and co-transfected them into mouse N2a cells along with a SpuFz1-v3 expression plasmid to identify the most effective guide, Tyr-ωRNA2, for subsequent experiments (Fig. 6d and Supplementary Fig. 7a). We then compared the editing efficiencies of three configurations: SpuFz1-v2 mRNA with a 275-nt extended canonical ωRNA scaffold, SpuFz1-v3 mRNA with the same scaffold, and SpuFz1-v3 mRNA combined with a chemically synthesized and modified gst28.m2 ωRNA scaffold. All mRNAs were produced by in vitro transcription, and the gst28.m2 ωRNA contained 2′-O-methyl and 3′ phosphorothioate modifications to enhance stability. The constructs were microinjected into mouse embryos, and editing outcomes were evaluated by deep sequencing and phenotypic analysis based on coat color. The SpuFz1-v3 and gst28.m2 group exhibited the highest average editing efficiency (91.37 ± 6.03%), followed by the SpuFz1-v3 and 275ext_75 group (85.95 ± 24.94%) and the SpuFz1-v2 and 275ext_75 group (47.09 ± 38.56%) (Fig. 6e). These molecular results were consistent with the observed phenotypes: all 12 mice in the SpuFz1-v3 & gst28.m2 group displayed an albino phenotype, with 11 fully albino and one partially albino, compared to albino rates of 39.3% (11/28) and 83.3% (15/18) in the SpuFz1-v2 and 275ext_75 and SpuFz1-v3 and 275ext_75 groups, respectively (Fig. 6f, g and Supplementary Fig. 7b–d). These findings demonstrate that the enFanzor system, combining the engineered SpuFz1-v3 protein with the optimized gst28.m2 ωRNA scaffold, achieves robust and efficient genome editing in mouse embryos, further validating its potential as a compact and effective gene-editing platform for in vivo applications.
In summary, our engineered enFanzor system achieved up to 50.17% insertion/deletion efficiency in unselected human hematopoietic stem and progenitor cells (HSPCs) and demonstrated 100% targeting efficiency in generating albino offspring through mouse embryo editing. These results underscore the potential of enFanzor as a versatile genome-editing platform for both therapeutic applications in humans and the generation of precise mouse models for human disease research.
Discussion
In this study, we present a scalable deep learning–based framework for engineering genome editing nucleases, demonstrated through the optimization of SpuFz1, a compact eukaryotic Fanzor nuclease. Our trained model, the Fanzor-Fitness Predictor, enables zero-shot prediction of functional mutations without requiring experimental screening data, allowing efficient navigation of the protein sequence space. By eliminating the need for labeled datasets or structure-based input, this strategy offers a data-efficient and rapid path to functional protein enhancement, particularly suited for emerging or data-scarce genome editing systems.
Using this framework, we identified and assembled high-performing variants into SpuFz1-v3, which exhibited markedly enhanced editing efficiency in mammalian cells. Although Fanzors were selected as a validation system due to their compact size and evolutionary distinction, the framework itself is designed to be adaptable to other nuclease families, including Cas variants, base editors, and synthetic enzymes. In this study, we further applied the same strategy to OgeuIscB, where model predictions showed strong consistency with previously published experimental mutational data, providing orthogonal validation of our approach. Additionally, in independent efforts, we have applied the same pipeline to the engineering of an identified IscB-based genome editor, further suggesting the broader applicability of this framework beyond the systems examined here.
To complement protein optimization, we performed comparative analysis and structural refinement of ωRNA scaffolds, resulting in gst28.m2—a highly compact 75-nt guide RNA that maintains efficient editing activity. This represents a 79% reduction in length compared to the previously most efficient 350-nt version, significantly improving its suitability for chemical synthesis and in vivo delivery. The integration of SpuFz1-v3 with gst28.m2 yielded the enFanzor system, which demonstrated strong genome-editing performance in human primary cells and mouse embryos.
Conventional optimization methods, such as charge-scanning or deep mutational scanning, often require labor-intensive experimentation and may miss complex epistatic interactions. Our AI-guided approach overcomes these limitations by enabling systematic exploration of sequence-function landscapes at scale. The success of SpuFz1-v3 and its validation across distinct nucleases highlights the generalizability of machine learning–based protein engineering strategies in genome editing.
Although our analysis revealed low off-target activity, more comprehensive profiling may be required in future therapeutic contexts. Similarly, expanding the TAM recognition profile of SpuFz1 may improve target site accessibility and warrants further study. In addition, the model evaluation is primarily focused on single-mutation effects, and extending the framework to robustly capture strong epistasis in multi-mutation regimes—potentially through incorporation of site–site interaction modeling or limited double-mutation calibration—represents an important direction for future work.
In summary, this work introduces a generalizable AI-driven strategy for the enhancement of genome editing nucleases, offering an efficient and scalable alternative to conventional protein engineering. Through the design of SpuFz1-v3 and the compact ωRNA gst28.m2, we constructed enFanzor, a compact and effective genome editing platform. More broadly, this study demonstrates how deep learning can accelerate the functional refinement of genome editors and lays a foundation for methodologically driven improvements across the genome engineering field.
Methods
Ethics statement
All animal procedures were approved by the Institutional Animal Care and Use Committee (IACUC) of the Center for Excellence in Molecular Cell Science, Chinese Academy of Sciences. Peripheral blood mobilized human CD34+ HSPCs from anonymous healthy donors were obtained from the First Affiliated Hospital of Zhejiang University School of Medicine (FAHZU), approved by the Medical Ethics Committee (MEC) of FAHZU.
Deep learning–driven engineering of Fanzor proteins
To identify homologous sequences of Fanzor proteins, the HMMER suite was employed to query the UniRef100 database using the SpuFanzor1 gene as a reference. Homology search was performed using the jackhmmer.sh script with the following parameters: jackhmmer -N 5 --incT $scaled_bitscore --incdomT $scaled_bitscore -T $scaled_bitscore --domT $scaled_bitscore --popen 0.02 --pextend 0.4 --mx BLOSUM62 -A $alignmentfile --noali --notextw --cpu 32 $query $seqdb. Here, $scaled_bitscore represents a bitscore value scaled by the length of the amino acid sequence. The resulting STO alignment file was converted into A2M format using a Python script. Model training was conducted using the DeepSequence variational autoencoder (VAE) framework, deployed on a remote server equipped with five NVIDIA 2080Ti GPUs. Given that previous studies have demonstrated the critical role of multiple sequence alignment (MSA) quality in model performance, a range of scaled bitscore thresholds (from 0.01 to 0.2) was tested for optimal MSA construction (Supplementary Data 2). Based on the total number of aligned amino acid positions and the diversity of retrieved sequences, a threshold of bitscore = 0.04 was selected. The final a2m file was then used to train the Fanzor-Fitness Predictor models within the DeepSequence VAE architecture. Model training was carried out in parallel across five GPUs, with each model trained for 300,000 epochs. In total, five independent models were trained. During inference, the mean prediction score across all five models was used to evaluate variant fitness (Supplementary Data 3). The top 98 single-point mutants with the highest predicted mean fitness values were selected for plasmid construction and functional validation in human cells using a fluorescent reporter assay.
Plasmid construction
Expression plasmids encoding Fanzor variants were constructed using the Seamless Cloning Kit (Beyotime, D7010M). Specifically, the SpuFz1 coding sequence was codon-optimized for human expression and synthesized by Dynegene (Shanghai). The optimized sequence was then inserted into the ABE8e expression vector (Addgene plasmid #138489) via seamless cloning. Designed amino acid mutations were subsequently introduced using mutation-specific PCR primers (Tsingke), as detailed in Supplementary Data 4. Plasmids expressing ωRNAs were constructed based on the backbone of lentiGuide-Puro (Addgene plasmid #52963). ωRNA sequences were amplified by PCR and inserted via seamless cloning. For the construction of guide RNA expression vectors targeting Cas9 and enIscB, forward and reverse primers encoding each guide RNA were annealed at 98 °C for 2 min, followed by ligation into expression plasmids containing Type IIS restriction sites using T4 DNA ligase (Takara). The sequences corresponding to all editing targets are listed in Supplementary Data 5. All plasmid constructs were prepared using endotoxin-free plasmid extraction kits (TIANGEN, DP118) following the manufacturer’s instructions. Final constructs were validated by Sanger sequencing to confirm sequence accuracy.
Mammalian cell culture, transfection, and fluorescence-activated cell sorting (FACS) analysis
HEK293T cells were obtained from ATCC (CRL-3216), and Neuro-2a (N2A) cells were obtained from the National Collection of Authenticated Cell Cultures (SCSP-5035). (https://www.cellbank.org.cn/). HEK293T cells were cultured in Dulbecco’s Modified Eagle Medium (DMEM, Gibco C11995500BT) supplemented with 10% fetal bovine serum (FBS, ExCell Bio FSP500). N2a cells were maintained in minimum essential medium (MEM, Gibco 11095080) supplemented with 10% FBS and 1% non-essential amino acids (NEAA, Gibco 11140050). All cell lines were incubated at 37 °C in a humidified atmosphere containing 5% CO₂. For transfection, ~30,000 cells were seeded into each well of a 48-well plate 18–24 h prior to transfection. Transfection was performed when the cell confluence reached ~70%. To screen predicted protein variants, a fluorescent reporter system (EGxxFP) was employed to sensitively detect Fanzor-mediated cleavage activity in HEK293T cells. This reporter includes blue fluorescent protein (BFP) and a split enhanced green fluorescent protein (EGFP) that can be reconstituted via the single-strand annealing (SSA) repair pathway following RNA-guided DNA cleavage. HEK293T cells were co-transfected with plasmids expressing SpuFanzor1 variants (fused with mCherry) and the EGxxFP reporter, along with the corresponding target-specific ωRNA. A total of 500 ng of plasmid DNA (1:1 weight ratio of editor to reporter plasmids) was delivered using polyetherimide (PEI) at a 1:3 (DNA:PEI) ratio. After 72 h, mCherry, BFP, and EGFP fluorescence signals were analyzed using a BD LSR Fortessa flow cytometer. To evaluate editing efficiency at endogenous genomic loci, 250 ng of editor plasmid and 250 ng of ωRNA plasmid were co-transfected into mammalian cells using the same PEI-based transfection protocol. Transfection mixtures were prepared in Opti-MEM and added dropwise to the culture medium. In experiments requiring analysis of sorted cell populations, mCherry-positive cells were isolated using a BD FACSAria II cell sorter. After ~72 h of expression, approximately the top 20% of cells with the highest mCherry fluorescence were sorted, and ~20,000 cells were collected. Cell pellets were lysed in 10 µL of lysis buffer, and lysates were incubated in a PCR thermal cycler. A volume of 1 µL of the lysate was used as a PCR template to determine genome-editing efficiency.
Detection of gene-editing frequency
To assess genome editing efficiency at endogenous loci, target genomic regions were amplified from cell lysates using high-fidelity KOD-plus-neo DNA polymerase (TOYOBO, KOD-401). In the first round of PCR, primers containing unique index sequences were used to enable subsequent sample demultiplexing during data processing. Gene-specific primers were listed in Supplementary Data 5. PCR products from the first round were pooled and purified using the QIAquick PCR Purification Kit (QIAGEN). A total of 20 ng of purified DNA was used as input for the second round of PCR, which was performed for ten cycles to further amplify the target regions and introduce sequencing adapters. The resulting pooled libraries were indexed to allow differentiation of DNA products derived from each sample within the sequencing run. Final PCR products were purified and subjected to 150-bp paired-end sequencing on an Illumina MiSeq or BGI sequencing platform. Raw sequencing data were downloaded and demultiplexed using the corresponding index combinations for each sample. Genome editing efficiency was subsequently quantified using CRISPResso2 (version 2.2.7) with default parameters.
Ghost ωRNA analysis
The Fanzor ghost loci in S. punctatus were obtained from datasets reported by Feng Zhang et al. After removing duplicate and incomplete sequences, the nucleotide identity of over 100 ghost ωRNA scaffold sequences was analyzed. Sequence alignment was performed using the NGPhylogeny.fr platform, followed by sequence similarity–based clustering. One or more representative sequences were selected from each cluster. These representative ωRNA sequences were chemically synthesized and cloned into ωRNA expression vectors for functional characterization in gene-editing assays. To gain additional insights into the structural properties of these RNAs, secondary structure prediction was performed using the RNA folding tool provided by UNAfold (http://www.unafold.org/mfold/applications/rna-folding-form.php). The resulting predicted structures were visualized using RNA2Drawer (https://rna2drawer.app/) to inform and guide scaffold optimization and engineering.
Guide RNA-dependent off-target analysis
To evaluate the gRNA-dependent off-target potential of the enFanzor system, Cas-OFFinder, part of the CRISPR RGEN Tools suite (http://www.rgenome.net/cas-offinder/), was used to predict putative off-target sites. As the Fanzor system lacks a defined TAM selection interface in Cas-OFFinder, the reverse-complementary 24-nt query sequence, comprising a 4-nt TAM motif (CATA) followed by the 20-nt target sequence, was input. The PAM type was set to SpRY Cas9 from Streptococcus pyogenes (5′-NNN-3′) to allow flexible base recognition. The maximum number of mismatches was restricted to four, while all other parameters were kept at default settings. Off-target prediction was conducted for three representative target sites: B2M, CA2, and CXCR4. Candidate sites with mismatches occurring within the TAM sequence (CATA) were excluded from downstream analysis. For each target site, ten potential off-target sites were manually selected based on the lowest number of mismatches, in ascending order. All predicted sequences and corresponding PCR primer sets are listed in Supplementary Data 6.
For GUIDE-seq analysis, HEK293T cells were electroporated with 600 ng of SpuFz1 variant expression plasmid, 600 ng of the corresponding ωRNA plasmid, and 50 pmol of double-stranded oligodeoxynucleotide (dsODN). Electroporation was performed using a Lonza 4D-Nucleofector system with program CM-130 in 20 μL of Nucleofector Solution (Lonza, Cat# V4XC-2032), following the manufacturer’s instructions. Seventy-two hours after electroporation, cells were harvested, and genomic DNA was extracted. GUIDE-seq libraries were generated by PCR amplification of dsODN-tagged genomic fragments and subsequently subjected to next-generation sequencing. Sequencing data were analyzed using the recommended parameters of the GUIDE-seq analysis pipeline (https://github.com/tsailabSJ/guideseq).
In vitro transcription of Fanzor mRNA and ωRNA
For in vitro transcription of Fanzor mRNA, transcription templates corresponding to different Fanzor variants were amplified by PCR using KOD-Plus-Neo polymerase (TOYOBO) from their respective expression plasmids. PCR products were purified using the Universal DNA Purification Kit (TIANGEN, Cat# DP214) and subsequently transcribed with the mMESSAGE mMACHINE T7 ULTRA Transcription Kit (Invitrogen, Cat# AM1345), following the manufacturer’s protocol. For the transcription of 350-nt ωRNAs based on the full-length scaffold, ωRNA templates containing a T7 promoter were first amplified from plasmids using specific primers, and in vitro transcription was performed using the MEGAshortscript T7 Kit (Life Technologies). Following transcription, all mRNA and ωRNA products were purified using the MEGAclear Transcription Clean-Up Kit (Invitrogen, Cat# AM1908), resuspended in pre-heated RNase-free water (95°C), and stored at −80 °C until use. The resulting Fanzor mRNAs were used for electroporation into human hematopoietic stem and progenitor cells (HSPCs) and microinjection into mouse embryos, while the in vitro–transcribed ωRNAs were utilized specifically for cytoplasmic injection into mouse embryos.
mRNA electroporation into CD34⁺ HSPCs
Electroporation was performed using the Lonza 4D-Nucleofector system with 20 μL Nucleocuvette Strips (V4XP-3032), following the manufacturer’s protocol. Chemically modified synthetic guide RNAs—with 2′-O-methyl and 3′-phosphorothioate modifications at the first and last three nucleotides—were ordered from GenScript. CD34⁺ HSPCs (5 × 10⁴ cells) were thawed 24 h prior to electroporation. For each electroporation reaction, 1.0 μg of Fanzor mRNA and 300 pmol of ωRNA (full-length chemically modified guide) were mixed prior to electroporation. Following electroporation, cells were resuspended in X-VIVO medium supplemented with cytokines, and the medium was replaced with erythroid differentiation medium (EDM) after 24 h. Cells were maintained in EDM for 96 h post-electroporation, after which editing frequencies were quantified.
Mouse housing and treatment
All animal procedures were approved by the Institutional Animal Care and Use Committee (IACUC) of the Center for Excellence in Molecular Cell Science, Chinese Academy of Sciences. Mice were maintained in individually ventilated cages (IVCs) under a certified specific pathogen-free (SPF) facility. All animals were housed with their littermates under a 12-h light/12-h dark (LD) cycle, at a room temperature of 20–26 °C and relative humidity of 30–70%. Zygotes were collected from eight-week-old female B6D2F1 mice (C57BL/6 J♀ × DBA2♂) that had been mated with ten- to twenty-week-old B6D2F1 males. Prior to zygote collection, all mice were euthanized using carbon dioxide (CO₂) in accordance with institutional ethical guidelines. Ten- to 14-week-old ICR females were used as pseudopregnant foster mothers for embryo transfer.
Microinjection of murine zygotes
Eight-week-old B6D2F1 female mice were superovulated by intraperitoneal injection of six international units (IU) of pregnant mare’s serum gonadotropin (PMSG), followed 48 h later by injection of 6 IU human chorionic gonadotropin (hCG). The females were then mated with B6D2F1 males for a 12-h mating period. Zygotes were collected from the oviducts of plug-positive females at 24 h post-hCG injection, following treatment with hyaluronidase (Sigma, Cat# H3884) to remove cumulus cells. For microinjection, a solution containing Fanzor mRNA (20 ng/μL) and sgRNA (20 ng/μL) was prepared in RNase-free water, centrifuged at 13,400×g for 10 min at 4 °C, and injected into the cytoplasm of zygotes. Injections were performed in HCZB medium supplemented with 5 μg/mL cytochalasin B (Sigma, Cat# C6762), using a micromanipulator (Olympus) and a FemtoJet microinjector (Eppendorf). Following injection, zygotes were cultured in AA-KSOM medium (Millipore, Cat# MR-106-D) at 37 °C with 5% CO₂ in air for 24 h, until reaching the two-cell stage. The resulting embryos were then transferred into the oviducts of pseudopregnant ICR females at 0.5 days post-copulation. About 2 weeks after birth, the offspring were examined for coat color phenotypes, and tail biopsies were collected for genomic DNA extraction. Editing efficiency at the target locus was quantified by amplicon-based next-generation sequencing (NGS).
Statistics and reproducibility
All experiments were performed with at least three independent replicates unless otherwise stated. Data were presented as mean ± standard deviation (s.d.), as indicated in the figure legends. Statistical analyses were conducted using appropriate methods as described in the corresponding figure legends, and exact p values are provided where applicable. No statistical method was used to predetermine sample size. No data were excluded from the analyses. The experiments were not randomized. The investigators were not blinded to allocation during experiments and outcome assessment.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Description of Additional Supplementary Files
Source data
Acknowledgements
We thank the Instruments Sharing Platform of the School of Life Sciences, East China Normal University. We also thank Ying Zhang from the Flow Cytometry Core Facility of the School of Life Sciences at East China Normal University for technical assistance.
Author contributions
S.C. and J.Y.L. conceived and designed the project, planned and analyzed the experiments, and co-led the research. S.C. performed ωRNA and protein engineering, endogenous site cleavage, and off-target analyses. J.Y.L. trained the deep learning model for protein engineering, conducted bioinformatics analyses, and supervised the entire project. J.Y.L. and S.C. conducted FACS experiments. S.H. assisted with experimental procedures. Animal experiments were performed by T.X., D.C., and N.C., and supervised by J.J. and J.S.L. Y.W. contributed to project supervision. Y.W. and J.S.L. provided critical input in manuscript reviewing and editing. All authors contributed to and approved the final manuscript.
Peer review
Peer review information
Nature Communications thanks Han Xiao and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.
Funding
This work was supported by the National Key R&D Program of China 2023YFC3403400 & 2024YFA1803300 (Y.W.), 2024YFC3408100 (S.C.), 2024YFC3407900 (J.L.), the National Natural Science Foundation of China 32300667 (J.L.), 32371535 (S.C.), 82450107 (S.C.), the project of Shanghai Municipal Science and Technology Commission 23HC1400400 (Y.W.), the National Program for Support of Top-Notch Young Professionals (Y.W.).
Data availability
All data supporting the findings of this study are available within the Article, Supplementary Information, Source Data and Supplementary Data files. Source data are provided with this paper. The raw sequencing data generated in this study have been deposited in the NCBI Sequence Read Archive under BioProject accession number PRJNA1440988. Source data are provided with this paper.
Code availability
The deep learning model for protein engineering was implemented using the publicly available DeepSequence framework, which is available at the GitHub repository https://github.com/debbiemarkslab/DeepSequence under the MIT License. There are no access restrictions for this publicly available code.
Competing interests
S.C., Y.W., and J.L. have submitted a patent application to the China National Intellectual Property Administration (CNIPA) pertaining to Fanzor variants and ωRNA variants described in this work (application number: 202610067574.3). Y.W. and J.L. are employees of YolTech Therapeutics, Shanghai, China. The remaining authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Yuxuan Wu, Email: yxwu@bio.ecnu.edu.cn.
Jiaoyang Liao, Email: jyliao@bio.ecnu.edu.cn.
Supplementary information
The online version contains supplementary material available at 10.1038/s41467-026-74624-6.
References
- 1.Doudna, J. A. & Charpentier, E. The new frontier of genome engineering with CRISPR-Cas9. Science346, 1258096 (2014). [DOI] [PubMed] [Google Scholar]
- 2.Hsu, P. D., Lander, E. S. & Zhang, F. Development and applications of CRISPR-Cas9 for genome engineering. Cell157, 1262–1278 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Barrangou, R. & Doudna, J. A. Applications of CRISPR technologies in research and beyond. Nat. Biotechnol.34, 933–941 (2016). [DOI] [PubMed] [Google Scholar]
- 4.Barrangou, R. & Horvath, P. A decade of discovery: CRISPR functions and applications. Nat. Microbiol.2, 1–9 (2017). [DOI] [PubMed] [Google Scholar]
- 5.Koonin, E. V. & Makarova, K. S. Origins and evolution of CRISPR-Cas systems. Philos. Trans. R. Soc. B374, 20180087 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Cong, L. et al. Multiplex genome engineering using CRISPR/Cas systems. Science339, 819–823 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ran, F. A. et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature520, 186–191 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Charlesworth, C. T. et al. Identification of preexisting adaptive immunity to Cas9 proteins in humans. Nat. Med.25, 249–254 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Chew, W. L. Immunity to CRISPR Cas9 and Cas12a therapeutics. Wiley Interdiscip. Rev. Syst. Biol. Med.10, e1408 (2018). [DOI] [PubMed] [Google Scholar]
- 10.Altae-Tran, H. et al. The widespread IS200/IS605 transposon family encodes diverse programmable RNA-guided endonucleases. Science374, 57–65 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Saito, M. et al. Fanzor is a eukaryotic programmable RNA-guided endonuclease. Nature620, 660–668 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Karvelis, T. et al. Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature599, 692–696 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Kato, K. et al. Structure of the IscB–ωRNA ribonucleoprotein complex, the likely ancestor of CRISPR-Cas9. Nat. Commun.13, 6719 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Makarova, K. S. et al. Evolutionary classification of CRISPR–Cas systems: a burst of class 2 and derived variants. Nat. Rev. Microbiol.18, 67–83 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Yoon, P. H. et al. Eukaryotic RNA-guided endonucleases evolved from a unique clade of bacterial enzymes. Nucleic Acids Res.51, 12414–12427 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Jiang, K. et al. Programmable RNA-guided DNA endonucleases are widespread in eukaryotes and their viruses. Sci. Adv.9, eadk0171 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Jiang, K., Gootenberg, J. S. & Abudayyeh, O. O. Fanzors, a family of eukaryotic RNA-guided DNA endonucleases. FEBS Lett.599, 1089–1093 (2025). [DOI] [PubMed] [Google Scholar]
- 18.McGaw, C. et al. Engineered Cas12i2 is a versatile high-efficiency platform for therapeutic genome editing. Nat. Commun.13, 2833 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhang, H. et al. An engineered xCas12i with high activity, high specificity, and broad PAM range. Protein Cell14, 540–545 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Han, D. et al. Development of miniature base editors using engineered IscB nickase. Nat. Methods20, 1029–1036 (2023). [DOI] [PubMed] [Google Scholar]
- 21.Xue, N. et al. Engineering IscB to develop highly efficient miniature editing tools in mammalian cells and embryos. Mol. Cell84, 3128–3140. e3124 (2024). [DOI] [PubMed] [Google Scholar]
- 22.Zetsche, B. et al. Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell163, 759–771 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Sun, A. et al. The compact Casπ (Cas12l)‘bracelet’provides a unique structural platform for DNA manipulation. Cell Res.33, 229–244 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Liu, Z.-X. et al. Hydrolytic endonucleolytic ribozyme (HYER) is programmable for sequence-specific DNA cleavage. Science383, eadh4859 (2024). [DOI] [PubMed] [Google Scholar]
- 25.Hino, T. et al. An AsCas12f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell186, 4920–4935. e4923 (2023). [DOI] [PubMed] [Google Scholar]
- 26.Riesselman, A. J., Ingraham, J. B. & Marks, D. S. Deep generative models of genetic variation capture the effects of mutations. Nat. Methods15, 816–822 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Hsu, C., Nisonoff, H., Fannjiang, C. & Listgarten, J. Learning protein fitness models from evolutionary and assay-labeled data. Nat. Biotechnol.40, 1114–1122 (2022). [DOI] [PubMed] [Google Scholar]
- 28.Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science379, 1123–1130 (2023). [DOI] [PubMed] [Google Scholar]
- 29.Mashiko, D. et al. Generation of mutant mice by pronuclear injection of circular plasmid expressing Cas9 and single guided RNA. Sci. Rep.3, 3355 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Yin, S. et al. Engineering of efficiency-enhanced Cas9 and base editors with improved gene therapy efficacies. Mol. Ther.31, 744–759 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Shui, S., Wang, S. & Liu, J. Systematic investigation of the effects of multiple SV40 nuclear localization signal fusion on the genome editing activity of purified SpCas9. Bioengineering9, 83 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Wu, Y. et al. Highly efficient therapeutic gene editing of human hematopoietic stem cells. Nat. Med.25, 776–783 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Marquart, K. F. et al. Effective genome editing with an enhanced ISDra2 TnpB system and deep learning-predicted ωRNAs. Nat. Methods21, 2084–2093 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Nelson, J. W. et al. Engineered pegRNAs improve prime editing efficiency. Nat. Biotechnol.40, 402–410 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Bae, S., Park, J. & Kim, J.-S. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics30, 1473–1475 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Tsai, S. Q. et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat. Biotechnol.33, 187–197 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Malinin, N. L. et al. Defining genome-wide CRISPR–Cas genome-editing nuclease activity with GUIDE-seq. Nat. Protoc.16, 5592–5615 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Kim, D. Y. et al. Hypercompact adenine base editors based on a Cas12f variant guided by engineered RNA. Nat. Chem. Biol.18, 1005–1013 (2022). [DOI] [PubMed] [Google Scholar]
- 39.Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol.38, 883–891 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Gehrke, J. M. et al. An APOBEC3A-Cas9 base editor with minimized bystander and off-target activities. Nat. Biotechnol.36, 977–982 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Liao, J. et al. Therapeutic adenine base editing of human hematopoietic stem cells. Nat. Commun.14, 207 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Fu, B. et al. CRISPR-Cas9-mediated gene editing of the BCL11A enhancer for pediatric beta(0)/beta(0) transfusion-dependent beta-thalassemia. Nat. Med.28, 1573–1580 (2022). [DOI] [PubMed] [Google Scholar]
- 43.Hendel, A. et al. Chemically modified guide RNAs enhance CRISPR-Cas genome editing in human primary cells. Nat. Biotechnol.33, 985–989 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Wang, H. et al. One-step generation of mice carrying mutations in multiple genes by CRISPR/Cas-mediated genome engineering. Cell153, 910–918 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Wang, F., Ma, S., Zhang, S., Ji, Q. & Hu, C. CRISPR beyond: harnessing compact RNA-guided endonucleases for enhanced genome editing. Sci. China Life Sci.67, 2563–2574 (2024). [DOI] [PubMed] [Google Scholar]
- 46.Yang, H. & Patel, D. J. Fanzors: Striking expansion of RNA-guided endonucleases to eukaryotes. Cell Res.34, 99–100 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Description of Additional Supplementary Files
Data Availability Statement
All data supporting the findings of this study are available within the Article, Supplementary Information, Source Data and Supplementary Data files. Source data are provided with this paper. The raw sequencing data generated in this study have been deposited in the NCBI Sequence Read Archive under BioProject accession number PRJNA1440988. Source data are provided with this paper.
The deep learning model for protein engineering was implemented using the publicly available DeepSequence framework, which is available at the GitHub repository https://github.com/debbiemarkslab/DeepSequence under the MIT License. There are no access restrictions for this publicly available code.
