Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 Feb 20:2026.02.20.706353. [Version 1] doi: 10.64898/2026.02.20.706353

Mechanistic machine learning enables interpretable and generalizable prediction of prime editing outcomes

Alvin Hsu 1,2,3,15, Peter J Chen 1,2,3,15, Angus H Li 1,2,3,15, Colin F Hemez 1,2,3, Xin D Gao 1,2,3, Markus Terrey 4, Charlie Nelson 4, Vijay Selvam 4, Ana Cristian 1,2,3, Amber N McElroy 5, Benjamin J Steinbeck 5, Gandhar K Mahadeshwar 1,2,3, Smriti Pandey 1,2,3, Zachary Barsdale 1,2,3,6, Paul Z Chen 1,2,3,7, Alexander A Sousa 1,2,3, Holt A Sakai 1,2,3, Rachel A Silverstein 8,9,10, Ilias Morad, Ryan K Krueger 11, Max W Shen 1,2,3,12, Benjamin P Kleinstiver 8,9,13, Cathleen M Lutz 4,14, Jakub Tolar 5, Bruce R Blazar 5, Mark J Osborn 5, David R Liu 1,2,3,16,17,*
PMCID: PMC12934823  PMID: 41757025

Abstract

Although prime editing (PE) can effect virtually any specified local change to genomic DNA in living systems, its efficient application currently requires extensive optimization of prime editing guide RNA (pegRNA) sequences. We present OptiPrime, a machine learning model of PE efficiency based on our current understanding of the mechanism of prime editing. OptiPrime achieves state-of-the-art accuracy on PE efficiency prediction and also enables prediction of nicking guide RNA (PE3) and dual pegRNA (twinPE) outcomes. We validated that OptiPrime has learned the determinants of mammalian mismatch repair (MMR), and is therefore well suited for nominating MMR-evasive silent edits that improve PE efficiency. We demonstrate the utility of OptiPrime in a variety of prospective therapeutic contexts, including in primary human and mouse cells. Finally, we show how OptiPrime can be used to achieve highly streamlined and efficient in vivo correction of a pathogenic mutation in the brain of a mouse model of KIF1A-associated neurological disorder.

Introduction

Precision gene-editing technologies have revolutionized research in the life sciences, as well as the treatment of both heritable and non-heritable diseases in the clinic1-48. Prime editing (PE) is a versatile gene-editing method that uses a programmable nickase and a prime editing guide RNA (pegRNA) to nick a target DNA sequence at a specified location, reverse transcribe the edited sequence from the pegRNA onto the nicked target DNA strand, and guide the cell through DNA repair processes to replace the original DNA sequence and make the resulting edit permanent on both DNA strands49 (Figure 1A). PE can install virtually any substitution, small insertion, small deletion, or combination thereof in the genome of a living cell without requiring double-stranded breaks (DSBs) or a donor DNA template (Figure 1B). For example, PE has been used clinically to correct a 2-bp deletion in NCF1, resulting in the effective treatment of chronic granulomatous disease (CGD) in multiple patients50.

Figure 1. Design of prime editing (PE) strategies in a high-dimensional search space.

Figure 1.

(A) Prime editors write genetic information encoded by a prime editing guide RNA (pegRNA) directly into the genome of a cell with a reverse transcriptase (RT). After 3′-flap resolution, DNA repair processes, including mismatch repair (MMR), can resolve the heteroduplex into either the unedited sequence or the edited sequence. See also Figure S1. (B) Prime editors can effect small substitutions, insertions, deletions, and combinations thereof. (C) The search space of PE strategies is vast. Designing a pegRNA requires choosing (1) a PE protospacer, (2) MMR-evasive silent edits, (3) RT template (RTT) length, and (4) primer binding site (PBS) length.

Although the basic mechanistic steps of PE are well understood (Supplementary Figure 1), achieving efficient prime editing, especially in challenging cell types, can require evaluating hundreds or thousands of combinations of pegRNA design parameters in order to identify a pegRNA that maximizes PE efficiency51-56 (Figure 1C). Each pegRNA contains a spacer that directs the prime editor to its genomic target site, a reverse transcriptase template (RTT) that encodes the edit to be installed, and a primer-binding site (PBS) that hybridizes to the DNA primer to initiate reverse transcription. All three of these components are critical determinants of prime editing outcomes, and many options are often plausible for each. In addition, we and others previously showed that evading cellular mismatch repair (MMR) substantially increases PE efficiency, and therefore incorporating additional silent or benign mutations—of which there are typically many possibilities—can substantially improve PE outcomes by causing PE intermediates to be poor substrates for MMR57,58.

The vast number (typically, thousands) of potential pegRNAs that could support a given prime edit of interest, together with the high variability of PE efficiencies among these possible pegRNAs, is daunting for those who wish to use PE for both research and therapeutic purposes. To address this challenge, we and others have sought to develop a machine learning (ML) model that computationally identifies pegRNA designs with the highest potential to support efficient PE. The ability of such a model to decrease the number of experiments required to identify a high-performing PE strategy is dependent on the accuracy of the model. Currently published ML models for PE efficiency include PRIDICT54,56, DeepPrime55,59, and OPED60. Although these models achieve high accuracy on test datasets that were sampled from the large datasets on which they were trained, a substantial portion of each of these datasets consists of poor-performing pegRNAs that an experienced researcher would typically avoid during an initial screen. Consequently, although these models can reject low-efficiency pegRNAs in prospective design contexts, they struggle to accurately predict optimal pegRNAs amid a set of well-performing ones (vide infra). Moreover, these existing models are not based on the biological mechanism of PE. As a result, they may not take advantage of mechanistic insights that suggest key nodes that strongly determine PE outcomes, and their lack of correspondence with the molecular processes involved in PE makes it difficult to glean biological insights from their trained state and from their output.

Here, we describe the development of OptiPrime, an ML model that generates predictions of PE efficiency by directly incorporating knowledge about the biological mechanism of PE into its mathematical structure. By design, OptiPrime uses separate ML models to predict “pseudo-rates” of biochemical steps involved in the PE mechanism. Each pseudo-rate model only uses features that are biologically relevant to each process to make its prediction, ensuring a clean separation of tasks. The system of differential equations described by these rates is then integrated through time to obtain a final output prediction of editing efficiency. This approach contrasts with those used by PRIDICT54,56 and DeepPrime55,59, which incorporate all sequence features and representations directly into a single black-box predictor for PE efficiency. To train OptiPrime, we performed matched pegRNA–target site screens in both MMR-deficient HEK293T cells and MMR-competent HeLa cells, resulting in a dataset of 74,769 PE efficiencies across 1,290 unique target sites that enabled us to systematically study both MMR-dependent and MMR-independent sequence determinants of editing efficiency. While traditional ML models usually do not directly include different experimental conditions as covariates and instead rely on additional model fine-tuning with new data to make predictions in new contexts, the mechanistic structure of OptiPrime enabled its joint training on a dataset of 297,962 PE efficiencies collected across 40 experimental contexts in our laboratory and others. We demonstrate, through benchmarking and ablation studies, that OptiPrime offers best-in-class accuracy on PE efficiency prediction and that its performance is dependent on its mechanism-based mathematical structure.

Moreover, we show that the measurable effects of individual mechanistic steps in the PE process are consistent with their corresponding pseudo-rate constants in OptiPrime, that these pseudo-rates can be used to train a model of PE3 efficiency, and that OptiPrime can make useful predictions about twinPE (a dual-flap, two-pegRNA application of PE61), even though the model was never trained on twinPE data. Finally, we demonstrate the ability of OptiPrime to streamline the development of therapeutic PE applications in several prospective contexts, including the rapid development of a PE strategy to correct a pathogenic mutation in Kif1a that resulted in >40% in vivo prime editing efficiency in the bulk brain cortex of mice. Collectively, these findings demonstrate that OptiPrime augments the potential of PE by greatly accelerating the identification of efficient editing strategies, and that mechanistic ML models can provide insights into the dynamics of complex biomolecular processes, enabling modular, high-performance prediction of their outcomes in living systems.

Results

High-throughput prime editing library screens recapitulate structural and cellular determinants of mismatch repair

In a previous CRISPRi screen of DNA repair and DNA metabolism genes, we discovered that MMR is a major bottleneck for PE efficiency57. This insight led us to develop the PE4 and PE5 editing strategies, in which a dominant negative variant of the MMR protein MLH1 (MLH1dn) is co-delivered with the editing components to suppress MMR and favor other repair pathways that lead to desired PE outcomes57. Consistent with these results, we also found that the use of pegRNAs that install additional silent or benign mutations near the desired edit could substantially increase PE efficiencies by causing heteroduplex PE intermediates to contain sufficient mismatches between the edited and unedited strands to evade recognition by MMR proteins, which preferentially engage single-nucleotide or small mismatched regions57.

The MMR pathway is initiated upon mismatch binding by the MutSα or MutSβ complexes62, which together recognize small mismatches and small insertion–deletion loops63,64. To comprehensively evaluate the sequence determinants of binding by MutSα and MutSβ, we designed “Lib-MMR,” a set of 10,000 pegRNAs designed to target 200 randomly selected exonic sites in the human genome with diverse edit types, including single-base substitutions (1,592), contiguous substitutions of 2–5 bases (2,792), non-contiguous substitutions of 2–5 bases (2,780), deletions of ≤10 bases (1,000), insertions of ≤10 bases (1,000), protospacer-adjacent motif (PAM) edits with deletions (400), and PAM edits with insertions (400). Lib-MMR also included 36 positive control pegRNAs that we previously used to edit endogenous genomic sites (Table 1)57.

Table 1. Design of Lib-MMR.

Prime edit class # members
Single SNP 1,592
Contiguous SNPs 2,792
Non-contiguous SNPs 2,780
Deletions (<10 bp) 1,000
Insertions (<10 bp) 1,000
PAM SNP + deletion 400
PAM SNP + insertion 400
Endogenous control sites 36
Total 10,000

Lib-MMR was designed to encompass a diverse set of prime edits in order to probe the sequence determinants of mammalian mismatch repair.

Although Lib-MMR contains diverse edit types and is therefore useful for assessing the determinants of binding by the MutSα/MutSβ complexes, we also sought to generate a dataset that would include many therapeutically relevant edits. Therefore, we designed “Lib-CV”, which comprises 10,406 pegRNAs that correct 944 known pathogenic variants in protein-coding regions from the ClinVar database65. In addition to a pegRNA that directly corrects each variant, we designed up to nine pegRNAs that also included benign “silent edits” that would not alter the corrected protein coding sequence (Table 2). Importantly, we designed the pegRNAs in Lib-MMR and Lib-CV with empirically determined heuristics to support efficient PE52 (Methods) so that an ML model that performs well on the resulting dataset will be required to distinguish between several candidates that perform well, rather than between rare candidates that perform well amidst many poor-performing alternatives.

Table 2. Design of Lib-CV.

Prime edit class # members
Correction to reference allele:
 SNP reversion 859
 Insertion reversion (deletion) 6
 Deletion reversion (insertion) 71
 Mixed reversion 6
Correction with silent edits:
 SNP reversion 8,600
 Insertion reversion (deletion) 56
 Deletion reversion (insertion) 723
 Mixed reversion 62
Total 10,406

Design of Lib-CV. Lib-CV was designed to investigate how silent edits affect prime editing in the context of correcting pathogenic mutations in ClinVar.

To assay PE efficiencies in a high-throughput manner, we packaged Lib-MMR and Lib-CV into lentiviral libraries of pegRNA–target site pairs. We transduced these libraries into cells at low multiplicity of infection (<0.3) to ensure that most transduced cells were only infected with a single lentivirus, selected for cells with integrated library members, then transfected these cells with PE construct plasmids to induce editing (Figure 2A). To thoroughly interrogate the effects of MMR on PE outcomes, we performed the screens with PE2 and PE4 in both HEK293T cells, which are partially MMR-deficient, and HeLa cells, which are MMR-proficient. As expected, we found that PE4 consistently outperformed PE2 across Lib-MMR. This effect was especially pronounced in HeLa cells, in which we observed a median 4.3-fold increase in editing rates with PE4 versus PE2, compared to only a median 1.5-fold increase in HEK293T cells (Figure 2B, 2C, Supplementary Figure 2A), consistent with previous observations that HEK293T cells are MMR-deficient. We chose to use the ratio of PE4 editing efficiency to PE2 editing efficiency (PE4:PE2) as a proxy for the propensity of a given edit to be reverted to the unedited sequence by MMR. We found that PE4:PE2 correlated more strongly with PE2 in HeLa cells (Spearman correlation ρ = −0.793) than in HEK293T cells (ρ = −0.390), suggesting that while MMR is the primary determinant of PE2 efficiency in MMR-proficient cell types, other factors determine PE2 efficiency in MMR-deficient cell types (Supplementary Figure 2B).

Figure 2. High-throughput lentiviral screens recapitulate the sequence determinants of mammalian mismatch repair (MMR).

Figure 2.

(A) Paired pegRNA–target site lentiviral libraries enable high-throughput evaluation of PE2 and PE4 outcomes in HEK293T and HeLa cells. (B) A scatter plot comparing editing efficiencies across Lib-MMR between PE2 (x-axis) and PE4 (y-axis) in HeLa cells. The gray line plots the equation y = x. PE4 consistently outperforms PE2 in HeLa cells. Brighter colors correspond to increased point density. See also Supplementary Figure 2A. (C–H) PE4:PE2 represents the ratio of PE4 editing efficiency to PE2 editing efficiency across library members in Lib-MMR. All PE4:PE2 values are plotted on a logarithmic scale. PE4:PE2 values below 0.25 (HeLa) or 0.5 (HEK293T) are clipped to each respective minimum value. PE4:PE2 values above 1024 (HeLa) or 4 (HEK293T) are clipped to each respective maximum value. (C) Beeswarm plots comparing PE4:PE2 between HeLa cells and HEK293T cells. PE4:PE2 is higher in HeLa cells than in HEK293T cells. (D) A scatter plot comparing PE4:PE2 across library members between cell types. PE4:PE2 correlates well between HEK293T and HeLa cells. Brighter colors correspond to increased point density. (E) Beeswarm plots of PE4:PE2 in HeLa cells for all twelve possible single-base substitutions, sorted by median PE4:PE2. See also Supplementary Figure 2E. (F–H) Beeswarm plots comparing PE4:PE2 in HeLa cells across varying lengths of (F) contiguous substitutions, (G) pure insertions, and (H) pure deletions. See also Supplementary Figure 2F-H. (I) A scatter plot comparing PE4:PE2 editing efficiency ratio of each edit to the ratio of PE2 editing efficiency with and without silent edits across Lib-CV in HeLa cells. Improvements in PE2 efficiency from including silent edits are positively correlated with PE4:PE2 efficiency ratios of the reference edit without any silent edits in HeLa cells. The red line is a total least squares regression line. Brighter colors correspond to increased point density.

We found that while PE4:PE2 values were generally lower in HEK293T cells than in HeLa cells as expected given the partial MMR deficiency of HEK293T cells, mean PE4:PE2 values correlated well between the cell types (ρ = 0.733), suggesting that trends in binding by MutSα and MutSβ are conserved (Figure 2D). Moreover, the PE4:PE2 ratios for single-base substitutions in Lib-MMR were highly consistent and linearly correlated with those obtained at endogenous genomic loci from our previous work (Pearson correlation r = 0.800, Figure S2C), with G-to-C and A-to-G edits resulting in the lowest median PE4:PE2 ratios, and C-to-T edits showing the highest median PE4:PE2 ratios57 (Figure 2E, Supplementary Figure 2D). This ranking is consistent with the known propensity of the corresponding mismatched DNA intermediates to be repaired in vitro by HeLa cell extracts66. Although both A-to-G and C-to-T edits create mismatched T:G intermediates, these edits qualitatively show very different PE4:PE2 ratios, suggesting that T:G mismatches may be asymmetrically repaired by MMR or preferentially processed by an orthogonal repair pathway.

Additionally, we observed that edits consisting of longer contiguous substitutions resulted in decreasing PE4:PE2 values, with Spearman ρ values of −0.484 (HeLa) and −0.427 (HEK293T) (Figure 2F, Supplementary Figure 2E). These data are consistent with structural observations that MutSα preferentially binds small mismatched regions rather than larger stretches of mismatched nucleotides63. Similarly, PE4:PE2 for both pure insertions and pure deletions decreased as the number of nucleotides inserted or removed increased, with deletion edits producing much higher PE4:PE2 than the corresponding insertion edits of the same length (Figure 2G, 2H, Supplementary Figure 2G, 2H). with strong negative rank-correlations for both pure insertions (HeLa ρ = −0.638; HEK293T ρ = −0.624) and pure deletions (HeLa ρ = −0.632; HEK293T ρ = −0.615). These data together suggest that known structural features of MMR proteins may underlie differences between PE2 and PE4 across a broad range of editing contexts.

In Lib-CV, we observed that including silent edits increased PE2 efficiency in an MMR-dependent fashion. The average log-fold-increase in PE2 efficiency at each site had positive, linear correlations with log-PE4:PE2 ratios of the simple reversion edit (HeLa r = 0.539; HEK293T r = 0.457, Figure 2I, Supplementary Figure 2I). Collectively, these data strongly reinforce our previous findings that including silent or benign mutations improves PE efficiency by shielding heteroduplex PE intermediates from engagement by cellular MMR machinery57. Moreover, by performing these screens in both HEK293T and HeLa cells, we confirmed that the bottleneck for efficient PE changes depending on the specific context in which the edit is performed, despite the common mechanism by which these edits are thought to occur.

OptiPrime is a mechanism-based ML model of prime editing that enables high-performance prediction of prime editing efficiency and mismatch repair effects

A challenging aspect of modeling PE efficiency with ML is that any one of the many individual steps of the PE process can be a strong determinant of editing outcomes. For example, MMR on the unproductive strand results in the target site’s reversion to the original sequence, restarting the PE process from the beginning. However, a pegRNA that fails to engage the prime editor would not allow the target site to reach the heteroduplex state to be corrected by MMR. While it is theoretically possible for traditional ML models to learn the intricacies of these dynamics given sufficient training data, we hypothesized that an ML paradigm in which the model’s mathematical structure is based on a researcher-defined multi-step reaction mechanism would be particularly well suited for training a model of PE efficiency. Although no biophysical parameters are directly provided, we hypothesized that imposing this inductive bias on the structure of the model would allow for better model generalization through incorporation of domain knowledge. Moreover, this approach may allow relationships between key PE intermediates to be learned during training in ways that may accurately reflect effective rates between these intermediates in cells. Likewise, incorporating details regarding variations in experimental procedure is difficult with traditional ML models. While prior works on PE efficiency prediction have addressed this issue by tuning a separate model on each experimental context55,56, this solution narrows the applicability of the resulting models because it is unclear to what extent each individual model has learned rules intrinsic to PE itself that are general across many contexts.

Inspired in part by prior work in physics-informed neural networks, neural ordinary differential equations, and virtual whole-cell models67-71, we created OptiPrime, an ML model of PE efficiency based on our current understanding of the PE mechanism (Figure 3A, Supplementary Figure 1). In OptiPrime, a combination of deep neural networks and linear regression models are used to model the effective rates of steps in our hypothesized mechanism of PE (“pseudo-rates”). These rates form a system of ordinary differential equations that can be integrated through time to evaluate the editing at the terminal timepoint for a given editing experiment (Figure 3B-C). While it would be experimentally impractical to create an ML training dataset with direct measurements of ground-truth rates in cells for many prime edits, we reasoned that pseudo-rates predicted using only handpicked salient features would represent their respective distinct steps in the PE mechanism. For a detailed discussion of the OptiPrime architecture and the considerations made in its design, see Supplementary Text 1.

Figure 3. OptiPrime is a mechanism-based machine learning model of prime editing efficiency.

Figure 3.

(A) The mechanism of prime editing used by OptiPrime as an inductive bias. Numbers (1–6) in circles represent the different genomic states that are used in the OptiPrime mechanism. Colored rates represent those which are predicted by machine learning models (left), while gray rates represent those that are kept constant (for a given cell type). (B) The differential equation that governs how the concentration of each state (ci) evolves over time. Predicted and constant rates are placed in the generator matrix at off-diagonal positions. Φi terms are the negative sum of all rates in the row, which maintains mass balance. (C) An overview of the procedure used to generate predictions with OptiPrime.

In brief, we divided our model into mechanistic steps that are mediated by the prime editor or those that only involve DNA repair machinery (Figure 3A). We modeled the former with hand-crafted features that were selected for their relevance to either DNA binding and nicking (kon,PE) or flap synthesis (ksyn). We then explicitly modeled the MMR pathway by including a reversible step for heteroduplex binding by the MutSα/MutSβ complexes. We represented MMR-independent heteroduplex repair as well as the off-rate of the MutSα/MutSβ complexes with the “HetFormer”, a neural network architecture we developed for this work inspired by the EvoFormer used by AlphaFold72,73. We pre-trained HetFormer on 64 million simulated heteroduplexes (Supplementary Figure 3, Supplementary Text 1). We trained OptiPrime on a dataset of 297,962 PE experiments that combines the data from our paired pegRNA–target site screens with prime editing data collected by other laboratories in previous studies54-56 to maximize the likelihood that the resulting model learns aspects of the PE mechanism that are general across all tested conditions.

To estimate model performances on unseen data, we split our dataset five ways, stratified by protospacer sequence (to avoid information leakage across splits). For each split, we trained models on four of the data partitions and evaluated on the other partition (5-fold cross-validation). Under these conditions, we found that OptiPrime consistently achieved high predictive performance on all training datasets, including those collected in this study (mean r = 0.693, ρ = 0.745, Figure 4A). We also compared OptiPrime to the previously reported DeepPrime55 and PRIDICT54,56 models (Figure 4A). We found that OptiPrime obtained significantly higher cross-validation rank correlations than DeepPrime-FT (mean r = 0.325, ρ = 0.399), PRIDICT2.0-HEK (mean r = 0.426, ρ = 0.538), and PRIDICT2.0-K562 (mean r = 0.516, ρ = 0.590) on the datasets collected in this work (all Steiger’s p < 10−323, Figure 4A).

Figure 4. OptiPrime accurately predicts prime editing efficiencies across diverse experimental contexts.

Figure 4.

(A) Spearman ρ values on different 5-fold cross-validation splits for the data collected in this study. The results for the final version of OptiPrime are in blue, those for ablation studies are shown in light blue, those for the PRIDICT2.0 models are shown in green, and those for the DeepPrime models are shown in purple. “OP without time-evolution” refers to replacing the mechanistic time-evolution layer in OptiPrime with a linear model that combines the outputs of the individual rate models. “OP without HF pretraining” refers to training OptiPrime with randomly initialized HetFormer weights, rather than from weights obtained from pretraining on 64,000,000 synthetically generated edits. “OP with scrambled mechanism” refers to training a version of OptiPrime in which the ML model that predicts the rate of each mechanistic step has been assigned to steps randomly, rather than based on the mechanism of prime editing. “PRIDICT2.0 (HEK)” and “PRIDICT2.0 (K562)” refer to the latest publicly available models trained on each cell type, respectively. DeepPrime-FT uses the HEK293T PE2max-e, HEK293T PE4max-e, HeLa PE2max, and HEK293T PE4max-e (since there is no DeepPrime-FT version for HeLa PE4max-e) models to predict values for HEK293T PE2max, HEK293T PE4max, HeLa PE2max, and HeLa PE4max, respectively. (B, C) Scatter plot of test set predictions with OptiPrime for the data collected in this study in HeLa cells with (B) PE2 and (C) PE4. (D) Comparison of rank correlations between OptiPrime, PRIDICT2.0, and DeepPrime-FT on data collected at the endogenous ATP1A3 locus for correcting four pathogenic mutations. (E) Scatter plots of ATP1A3 correction with OptiPrime scores. Colors represent the spacers, as denoted in ref. 47. (F) Beeswarm plots of predicted MMR in HeLa cells across the test sets collected in this work. As expected, predicted kMMR decreases when changing from a PE2 context to a PE4 context. (G) A scatter plot of observed PE4:PE2 editing efficiency ratio (PE4:PE2) in HeLa cells vs. predicted koff,MutS across the test sets collected in this work. As expected, decreased koff,MutS corresponds to an increased binding affinity for the MutSα/MutSβ complexes, resulting in increased MMR correction (as quantified by PE4:PE2). (H) A scatter plot comparing observed PE4:PE2 in HeLa cells to the ratio of the corresponding predictions by OptiPrime. Predicted PE4:PE2 correlates well (r = 0.582, ρ = 0.668), indicating that OptiPrime has learned the sequence determinants of MMR correction propensity.

To determine the design features that contribute to OptiPrime’s predictive performance, we also performed an ablation study, in which several versions of OptiPrime were re-trained with each component systematically removed. Overall, we found that the results achieved by OptiPrime were dependent on all of its key design features—removing any component resulted in significantly worse cross-validation performance on all of our datasets, including the mechanistic time-evolution layer and the pretraining of the HetFormer (all Steiger’s p ≤ 2×10−12, Figure 4A). Moreover, “scrambling” the reactions by re-assigning the rate models of the mechanistic graph to randomly chosen reaction edges also significantly deteriorated the model performance (all Steiger’s p < 10−323, Figure 4A), indicating that OptiPrime’s high performance is derived from the correspondence between its architecture and the underlying PE mechanism. Collectively, these findings indicate that the above design choices each contribute to OptiPrime’s high accuracy.

With these cross-validation results in hand, we computed predictions with the final version of OptiPrime on held-out test sets sampled from each dataset (Figures 4B, 4C, Supplementary Figure 4). The cross-validation results translated well to the test set across all four experimental conditions tested in this work (mean r = 0.723, ρ = 0.775). Lastly, we compared OptiPrime’s predictions to those generated by PRIDICT2.0 and DeepPrime on several test sets of prime edits from arrayed experiments that we performed during the course of our work on correcting mutations in ATP1A3 that cause alternating hemiplegia of childhood47 (Figure 4D, 4E). We observed high correlations between OptiPrime and empirical editing for D801N (ρ = 0.604) and L839P (ρ = 0.599) and modest correlations for E815K (ρ = 0.326) and G947R (ρ = 0.196). We noted that these relatively poor performances could be primarily attributed to bad predictions with protospacers E815K-NGG3 and G947R-NGG2, which both supported efficient PE in spite of their poor predicted protospacer scores by the CRISPR nuclease sgRNA model Rule Set 374. Indeed, removing E815K-NGG3 and G947R-NGG2 improved OptiPrime’s rank correlation on the respective test sets to ρ = 0.786 and ρ = 0.787, respectively (Figure 4E). In comparison, neither PRIDICT2.0 nor DeepPrime-FT were able to achieve rank-correlations above 0.43 on any of the four test sets, and both models generated predictions with negative rank-correlations at one site, suggesting that following the predictions of PRIDICT2.0 or DeepPrime-FT can sometimes produce results that would be on average results worse than selecting at random from the candidate pegRNAs we evaluated. Taken together, these results show that OptiPrime models PE efficiency with an accuracy that can exceed that of previously published models.

Another potential benefit of OptiPrime’s mechanistic architecture is the possibility that the resulting learned pseudo-rates may provide insights into specific steps of the PE process. Unlike non-mechanistic ML models, in which outcomes cannot be disentangled from their influence by individual model features, the mechanistic nature of OptiPrime enabled us to investigate in greater detail how input features influence each pseudo-rate. For example, we found that predicted kMMR, which takes MLH1dn use as an input, decreases by a median of 14-fold under PE4 conditions compared to under PE2 conditions, consistent with our understanding that the use of MLH1dn in PE4 inhibits MMR (Figure 4F).

To further probe OptiPrime’s ability to evaluate the role of MMR across diverse edits, we investigated the relationships between PE4:PE2 and the three pseudo-rates predicted by HetFormer: krep,u, krep,e, and koff,MutS. In our proposed PE mechanism, we consider binding of MutSα/MutSβ complexes to an editing intermediate to be necessary for that edit to be reverted to the unedited sequence by MMR machinery. In OptiPrime, we assume that the on-rate is governed by diffusion and that all variation in binding affinity can be attributed to the off-rate. We found a significant negative rank-correlation between predicted koff,MutS rates and empirical PE4:PE2 (HeLa ρ = −0.42, p = 1.5×10−104) across our test set, consistent with our hypothesis that tighter binding by the MutSα/MutSβ complexes (i.e., a lower off-rate) results in increased reversion by MMR (i.e., increased PE4:PE2) (Figure 4G). Moreover, the ratio of predicted PE4 efficiencies to predicted PE2 efficiencies with OptiPrime correlates well with observed PE4:PE2 values (r = 0.582, ρ = 0.668), indicating that the HetFormer has learned the sequence determinants salient to MMR in the context of PE (Figure 4H) and that OptiPrime is well-suited for selecting MMR-evading silent edits. We speculate that this ability arises from the fact that OptiPrime was jointly trained on both PE2 and PE4 simultaneously, rather than through fine-tuning on PE4 data.

Pseudo-rates learned by OptiPrime enable PE3 and twinPE predictions

We next asked whether OptiPrime pseudo-rates could be used to predict outcomes of PE variants such as adding a nicking guide RNA (‘PE3’49) or using two pegRNAs to generate complementary edited strands (‘TwinPE’61), despite a lack of training data on these variants. In PE3, a nicking sgRNA (nsgRNA) is used to drive DNA repair towards incorporation of the edit by nicking the DNA strand complementary to the PE protospacer (Figure 5A). Analogous to the relationship between PE2 and PE4, we previously found that including MLH1dn could improve PE3 efficiencies and reduce indel formation, a strategy we designated PE557. We assembled a dataset of 940 PE3 (or PE5) efficiencies paired with their corresponding PE2 (or PE4) efficiencies with 159 prime edits collected at endogenous sites in human cells. Although some PE strategy design tools allow users to nominate nsgRNAs for PE3, these tools give recommendations solely on the basis of predicted Cas9 nuclease efficiencies for indel generation51,54,56. While creating double-strand DNA breaks by Cas9 nuclease is a prerequisite for indel formation in these contexts, indel sequence outcomes are shaped by various end-joining pathways (i.e., c-NHEJ and MMEJ). This reasoning is consistent with our observation that neither DeepSpCas9 (mean ρ = −0.086, all p > 0.05)75 nor Rule Set 3 (mean ρ = −0.084, all p > 0.05)74 scores provided meaningful predictive PE3 editing efficiency signal on any test dataset (Figure 5B), although we note a modest correlation between DeepSpCas9 (mean ρ = 0.241) or Rule Set 3 (mean ρ = 0.221) scores and PE3 indel efficiencies. These data suggest that nuclease scores are not useful for predicting optimal nsgRNAs in the context of PE3, motivating the development of bespoke models of PE3 efficiency.

Figure 5. OptiPrime rates enable prediction of PE3 and twinPE outcomes.

Figure 5.

(A) In PE3, a complementary strand-nicking sgRNA (nsgRNA) drives editing towards the desired product. (B) Test set correlations of PE3 editing data taken from our work correcting three pathogenic ATP1A3 mutations47. (C) Beeswarm plots of SHAP values for PE3 editing predictions on test sets. (D) TwinPE installs longer edits by synthesizing complementary flaps with two pegRNAs that target complementary strands. We predict twinPE efficiencies with OptiPrime by assuming that 3’-flap synthesis is rate-limiting, and that synthesis rates of each flap is independent of the other. (E) Scatter plots showing zero-shot prediction correlations of twinPE efficiencies on four datasets of twinPE efficiencies taken from our previous work61,77.

Transfer learning is a ML approach in which a model trained on one task is repurposed for a different, typically related, task. A common strategy for transfer learning is to fix the base model’s weights and use its outputs or computational intermediates as model inputs for the new task. Although the detailed mechanisms by which PE3 improves editing efficiency are not fully understood, we hypothesized that the pseudo-rates learned by OptiPrime could serve as mechanistically relevant features for predicting PE3 efficiencies. We trained gradient-boosted tree models to predict PE3 editing efficiencies and indel frequencies based on these pseudo-rates, predicted OptiPrime PE2 (or PE4) efficiencies, as well as other hand-extracted features (Methods) and found that the resulting model was able to achieve good predictive correlation for each PE3 test dataset collected during the course of prime editing experiments correcting pathogenic mutations in ATP1A3 (mean ρ = 0.534, Figure 5B). Consistent with our observation that nuclease efficiency scores alone do not predict PE3 nsgRNA efficiency, we found that removing OptiPrime-predicted pseudo-rates and editing efficiencies as input features abrogated predictive power on our test sets (mean ρ = 0.025, all p > 0.05, Figure 5B). Consistent with these findings, when we quantified the contribution of each input feature to our PE3 model’s test set prediction performance with the widely used SHAP explanation technique76, we found that features taken from OptiPrime were among the most important factors influencing its outcomes—kon,PE, OptiPrime-predicted editing efficiency, and koff,MutS were all among the top five predictive features (Figure 5C). We also found that including only the final predicted PE2 (or PE4) efficiency from OptiPrime could still achieve some predictive power, albeit lower than when pseudo-rates are included (mean ρ = 0.464, Figure 5C). Taken together, these results further indicate that integrating inferred mechanistic dynamics into models of PE efficiency with nicking guide RNAs (PE3) is crucial for accurate performance.

We also hypothesized that OptiPrime could predict the outcomes of twinPE, in which two pegRNAs targeting complementary genomic DNA strands are used to generate two 3′ flaps that anneal to each other, obviating the need for second-strand DNA synthesis to generate stable PE products (Figure 5D). While a traditional ML model trained on PE data would be unsuitable for such a task because predictions about the relevant flap synthesis steps cannot be disentangled from those for irrelevant DNA repair steps, mechanistic ML could enable us to rewire pseudo-rate models to generalize to processes such as twinPE with different but related mechanisms compared to those used to generate the base model.

Although it is not possible to model DNA repair in twinPE using OptiPrime because twinPE does not create an unresolved heteroduplex intermediate as traditional PE does, we assumed that 3′-flap installation efficiency is the rate-limiting step in the twinPE mechanism. We thus chose to predict twinPE efficiencies by taking 3′-flap synthesis pseudo-rates from OptiPrime for each individual pegRNA, ignoring the pseudo-rates for all DNA repair steps downstream of flap synthesis (Figure 5D). We collected a dataset of 667 twinPE outcomes across four endogenous loci from our previous reports disclosing twinPE61 and eePASSIGE77, a technique that uses PE to insert landing sites for gene integration with the Bxb1 recombinase. We found that this procedure achieves moderate predictive performance across the four twinPE datasets (mean ρ = 0.412, Figure 5E) despite never having been trained on any twinPE data, demonstrating that OptiPrime has learned features inherent to the steps in the PE mechanism that can then be disentangled and used separately for new mechanisms. Collectively, these data demonstrate that mechanistic ML models can facilitate downstream fine-tuning or rewiring to generate predictions on new, mechanistically related processes.

OptiPrime accelerates the development of therapeutic prime editing strategies

To establish the utility of OptiPrime in finding efficient PE strategies in various therapeutic contexts, we designed pegRNAs capable of effecting several therapeutically relevant edits, including correction of four pathogenic mutations in the CFTR gene and one pathogenic mutation in the COL7A1 gene, orthogonalization of the IL2 receptor, and correction of a pathogenic Kif1a mutation in a mouse model of KIF1A-associated neurological disorder (KAND) (Methods). We tested these pegRNAs in various cell types, including in human primary T cells and in embryonic fibroblasts from the Kif1algdg model mouse. Finally, we used the results from an OptiPrime-streamlined optimization campaign in an in vivo study to demonstrate how OptiPrime can greatly accelerate the development of PE strategies to correct pathogenic mutations that cause rare diseases.

Cystic fibrosis (CF) is caused by mutations in the CF transmembrane regulator (CFTR) gene that encodes the CFTR protein, an anion transporter of Cl and HCO3−82. We recently reported the development of a PE strategy for correcting CFTR p.F508del, a mutation found in 85% of people with CF12. Our efforts towards developing this strategy involved several cycles of iterative pegRNA optimization over the course of several years, resulting in the identification of “SE2”, an epegRNA encoding several silent edits that corrects p.F508del with 11% efficiency in HEK293T cell lines homozygous for CFTR p.F508del when used in a PE2 strategy (without a nicking guide RNA). Assaying the eight designs with the highest OptiPrime-predicted editing efficiencies enabled us to identify an epegRNA capable of correcting this mutation with 22% efficiency in a PE2 strategy, a 2.0-fold increase over that achieved by SE2 (Figure 6A, Supplementary Figure 5A). In contrast, none of the 16 top-scoring designs proposed by either PRIDICT2.0 or DeepPrime was able to achieve over 1% editing (Figure 6A).

Figure 6. OptiPrime streamlines identification of efficient therapeutic prime editing strategies in vitro.

Figure 6.

(A) Correction of the CFTR p.F508del mutation with OptiPrime, PRIDICT2.0, and DeepPrime in HEK293T cells. Points denote individual experimental editing efficiencies (N = 3) for each ranked pegRNA on the x-axis, while lines indicate the maximum editing efficiency achieved after evaluating the top x-ranked pegRNAs. (B) Correction of the COL7A1 p.R185X mutation with OptiPrime, PRIDICT2.0, and DeepPrime in fibroblasts derived from patients with recessive dystrophic epidermolysis bullosa. (C) Installation of the orthogonal IL2 receptor (IL2RB H134D Y135F) in primary human T cells. “293T opt.” denotes editing efficiency with a previously reported pegRNA without silent edits optimized by hand in HEK293T cells57. (D) PASSIGE is a strategy for the insertion of large DNA cargo into mammalian cells61,77. First, PE is used to install an attP site into a genomic locus. Then, a DNA donor containing the corresponding attB site is recombined into the genome with a large serine recombinase, such as Bxb1. (E) Installation of a gene-sized cargo into intron 1 of CFTR in HEK293T cells with eePASSIGE with either OptiPrime-nominated pegRNA pairs or expert-designed pegRNA pairs. Points denote individual experimental editing efficiencies (N = 3) for each ranked pegRNA pair on the x-axis, while lines indicate the maximum editing efficiency achieved after evaluating the top x-ranked pegRNA pairs.

Although the CFTR p.F508del mutation is present in the majority of people with CF, more than 2,000 CFTR variants have been identified in the human population, over 700 of which are verified to cause CF82. We used OptiPrime to score and select pegRNAs that are capable of correcting either the p.W1282X or p.G542X nonsense mutations, which are not responsive to small-molecule CFTR modulators such as Trikafta. For both mutations, assaying the eight top OptiPrime-predicted epegRNAs resulted in higher editing than assaying the same number of epegRNAs scored by PRIDICT2.0 or DeepPrime (Supplementary Figure 5B-E). We noted that, despite allowing OptiPrime to evaluate silent edits proximal to p.W1282X, the optimal strategy proposed by OptiPrime did not have any silent edits, indicating that the higher editing efficiency of the OptiPrime-nominated design is due to an RTT+PBS length combination that was more optimal than those nominated by PRIDICT2.0 or DeepPrime, rather than an advantage conferred by silent edit designs. Moreover, the PE-suitable NGG PAM closest to CFTR p.G542X would place the corrective T-to-G edit 45 bp downstream of the PE nick. Although OptiPrime, PRIDICT2.0, and DeepPrime do not allow design of edits this distance away, OptiPrime alone was able to identify a protospacer that places the c.1624G>T mutation 4 bp downstream from the PE nick with an NGA PAM that is compatible with a Cas9-VRQR variant83. For the sake of comparison, we assessed the RTT+PBS length combinations output by PRIDICT2.0 and DeepPrime by replacing the NGA PAM used by the OptiPrime-selected protospacer with NGG.

Next, we tested how the pegRNAs from OptiPrime performed in primary human cell types, which were not used to generate any of OptiPrime’s training data. Recessive dystrophic epidermolysis bullosa (RDEB) is caused by mutations in the COL7A1 gene that result in impaired function of type VII collagen, leading to breakdown of the dermal–epidermal junction and a severe blistering phenotype84. Using OptiPrime, PRIDICT2.0, and DeepPrime, we chemically synthesized the top three epegRNAs as evaluated by each model and electroporated them into RDEB patient-derived fibroblasts with in vitro-transcribed mRNA encoding PEmax. We found that the best of the three OptiPrime-selected pegRNAs achieved 37% correction, while none of the six pegRNAs selected by PRIDICT2.0 and DeepPrime were able to effect above 5% correction (Figure 6B). We note that the DeepPrime-nominated pegRNAs showed high levels of indels, which we speculate may arise from its selection of pegRNAs with short RTT homology lengths. Notably, the levels of editing from OptiPrime-generated pegRNAs were also higher than those we previously achieved with adenine base editing16, demonstrating that OptiPrime is capable of designing PE strategies as efficient as those using base editing in difficult cell types without additional optimization.

We also chose to revisit the installation of the IL2RB H134D Y135F variant that enables selective stimulation of T cells with a corresponding orthogonal IL-2 cytokine85. We electroporated primary human T cells with PEmax-encoding mRNA and chemically synthesized pegRNAs. In this assay, we compared three OptiPrime designs that include silent edits to three designs without silent edits, two of which were designed with PRIDICT2.0 (which does not generate silent edits) and one of which we previously identified as the top performer from a manual assessment of 24 pegRNAs in HEK293T cells57. We did not include any comparisons with DeepPrime, since it is limited to making ≤3-bp changes. We found that the best of the three OptiPrime-designed pegRNAs achieved 7.0% editing in a PE2 context with synthetic pegRNAs. In comparison, the best pegRNA without any silent edits, which was optimized beforehand in HEK293T cells, achieved only 1.1% editing, and both of the PRIDICT2.0-generated pegRNAs yielded ≤0.1% editing (Figure 6C). Taken together, these results demonstrate that the silent edit combinations identified by OptiPrime can improve editing efficiency at previously challenging targets and in cell types beyond those used to generate OptiPrime’s training data, including primary human T cells.

Given OptiPrime’s ability to predict twinPE efficiencies, we designed twinPE epegRNAs capable of installing the attP landing site for Bxb1 recombinase in intron 1 of CFTR, as one step towards a potential mutation-agnostic therapy for people with CF caused by rare mutations (Figure 6D). We performed eePASSIGE77 integration experiments comparing one-pot integration efficiencies of eight epegRNA pairs nominated by OptiPrime against eight epegRNA pairs designed by a human PE expert with assistance with CRISPick74,86,87. We found that the top OptiPrime-nominated epegRNA pair yielded 10% integration efficiency of a multi-kilobase plasmid cargo, compared to 8.0% achieved by the best-performing human expert-designed epegRNA pair (Figure 6E).

Finally, we tested if OptiPrime could streamline efforts to address a longstanding need in medicine—the development of N=1 treatments for de novo genetic mutations that result in disease88. While high-throughput sequencing allows us to diagnose the genetic causes of some diseases rapidly, the development of a corresponding genome editing treatment has until recently been too slow to rescue patient phenotypes for most urgent genetic diseases. In a study led by Musunuru, Ahrens-Nicklas, and coworkers, we recently reported the development of the first personalized gene-editing treatment of a rare genetic disease, in which an infant was dosed with a bespoke base editor developed specifically to correct the infant’s unique mutation in CPS1 at only seven months of age48. Developing this treatment so rapidly was possible due to the relatively small design space of base editing strategies; such a timeline would be more difficult for PE, given the far larger design space of pegRNAs.

We envision a pipeline in which detection of a pathogenic mutation is followed by streamlined generation of a small number of promising candidate corrective pegRNAs with OptiPrime. These candidates are then tested directly in patient-derived cells to rapidly identify efficient editing strategies (Figure 7A). As a proof-of-concept for this strategy, we used OptiPrime to design pegRNAs to correct the “leg dragger” mutation in Kif1algdg mice. This spontaneous missense mutation results in the Kif1a p.L181F protein change and causes spastic paraplegia in homozygous Kif1algdg similar to the defects observed in patients with pathogenic mutations in KIF1A89,90.

Figure 7. OptiPrime accelerates the development of corrective strategies for pathogenic mutations in vivo.

Figure 7.

(A) An experimental outline for rapid development of therapeutic prime editing strategies with OptiPrime. Rather than generating a cell line bearing the mutation of interest, we propose directly editing patient-derived cells through electroporation of in vitro-transcribed mRNA and synthetic pegRNA/nsgRNA constructs or viral transduction. (B) Kif1a genotypes relevant to this study. Kif1a+ refers to the genotype from the GRCm38 reference assembly. Kif1algdg is caused by a C-to-T transition mutation resulting in the p.L181F protein change. The editing strategy used by the OP-5 family of pegRNAs changes a different codon position than the pathogenic mutation and includes five non-coding silent edits. (C–F) Lead optimization of a corrective strategy for Kif1algdg/+ in mouse embryonic fibroblasts through electroporation of PE mRNA, synthetic pegRNAs, and synthetic nsgRNAs. (C) OP-5 was identified in an initial assessment of eight pegRNAs encoding different silent edit strategies. (D) OP-5.1 was identified in a follow-up experiment investigating eight RTT+PBS lengths for OP-5. (E) Assaying 15 nsgRNAs identified the +65 and +114 nsgRNA sequences as promising candidates for further assessment. (F) Assessing engineered and evolved PE protein sequences identified PE6b as our lead candidate. (G) In vivo correction strategy of Kif1algdg. Dual AAV9 encoding split PE6b, the in vitro-optimized pegRNA (converted to an epegRNA), and the in vitro-optimized nsgRNA were delivered through intracerebroventricular injection into P0 Kif1algdg/+ mice. (H) In vivo editing of Kif1a genomic DNA in mouse cortex and spinal cord. Genomic DNA from both bulk mouse tissue and GFP-positive nuclei (to enrich for transduced cells) was evaluated for PE efficiency 4 weeks post-injection.

We designed pegRNAs to correct the Kif1algdg mutation with OptiPrime and tested these pegRNAs in mouse embryonic fibroblasts (MEFs) derived from Kif1algdg mice by electroporating cells with in vitro-transcribed PE mRNA and chemically synthesized pegRNAs and nsgRNAs. While such a campaign would typically be cost-prohibitive due to the difficulty of synthesizing long RNA molecules, we hypothesized that the accuracy of OptiPrime would reduce the number of pegRNAs required to find an initial lead.

First, we evaluated the top eight silent edit combinations nominated by OptiPrime with only one RTT+PBS combination each (OP-1–8), which identified OP-5 as the pegRNA encoding the silent edit combination that yielded the highest editing efficiency of 10% (Figure 7B, 7C). In our second optimization assay, we evaluated seven additional RTT+PBS combinations with the same silent edit strategy as OP-5, revealing OP-5.4 as the lead pegRNA candidate that yielded 22% correction in heterozygous Kif1algdg/+ MEFs without any nsgRNA (Figure 7D). In total, these pegRNA optimization experiments required only 15 total pegRNAs—far fewer than the hundreds required for each edit in previous therapeutic PE studies12,47. With OP-5.4 in hand, we evaluated 15 nsgRNAs and identified several that enabled increases in editing efficiency exceeding two-fold. Although comprehensive evaluation of candidates in homozygous Kif1algdg/lgdg MEFs was not possible due to low cell availability, we performed small-scale investigations in this system. We observed that some nsgRNA candidates induced high levels of indel formation in this setting, whereas nsgRNAs +65 and +114 maintained indel rates below 5%. We reasoned that these data, while limited, may more accurately reflect editing outcomes in the phenotypically relevant homozygous Kif1algdg/lgdg model system. We therefore advanced both nsgRNAs +65 and +114, which in larger-scale characterization in heterozygous MEFs conferred 2.5- and 1.5-fold improvements in editing efficiency over PE2, yielding 39% and 23% editing, respectively. Although nsgRNA +65 yielded higher editing efficiency, it also generated more indel byproducts compared to nsgRNA +114 (Figure 7E, Supplementary Figure 6A). Finally, we assayed both of these PE strategies with evolved PE6 prime editor variants91, which revealed that PE6b offers an additional 1.1-fold improvement over PEmax (Figure 7F), now achieving 64% (nsgRNA +65) and 51% (nsgRNA +114) editing efficiencies in heterozygous Kif1algdg/+ MEFs.

These results, obtained from only four experiments requiring a total of 4 weeks, demonstrate how OptiPrime can greatly accelerate the development of new, efficient therapeutic prime editing strategies. We packaged dual adeno-associated virus 9 (AAV9)92 with an epegRNA version of OP-5.4, either nsgRNA +65 or +114, and PE6b (Figure 7G). These AAV, along with an AAV encoding an EGFP–KASH transduction marker, were delivered via intracerebroventricular injection into heterozygous Kif1algdg/+ mice at birth (P0). After four weeks, mice treated with the AAV encoding both nsgRNAs showed above 40% average bulk brain cortex editing, with >70% editing efficiency among transduced (GFP+) cells (Figure 7H, Supplementary Figure 6B-C). Although we found lower editing levels in some brain regions, consistent with previous studies47, we found that over 50% of transcripts were edited in bulk tissue, suggesting that AAV9 transduction and prime editing occurs preferentially in cells that express Kif1a. Collectively, these findings show that OptiPrime can be used to rapidly generate therapeutic PE strategies to efficiently correct pathogenic mutations in vivo, including those that arise from de novo mutations that would often require N=1 treatments.

Discussion

The versatility of PE offers immense potential as a tool for research and as a way to precisely tailor the genome of living systems for therapeutic and other applications. In this work, we demonstrated that OptiPrime, a mechanistic ML model, is able to predict PE efficiencies with high enough accuracy to greatly streamline an otherwise laborious pegRNA optimization process. Moreover, we showed that the mechanistic formulation underlying OptiPrime confers several advantages over traditional ML models, including the ability to use transfer learning and mechanism rewiring to predict outcomes of PE3 and twinPE, editing mechanisms not seen during training.

While OptiPrime was trained on many experiments in diverse cell types, we note that high-performing pegRNAs optimized using its HeLa cell parameters tend to also maintain their performance in HEK293T cells, while the converse is not true. Therefore, we recommend that researchers interested in designing pegRNAs with OptiPrime use its default settings for typical applications. This recommendation is supported by our encouraging in vitro results in primary human T cells, patient-derived fibroblasts, and mouse embryonic fibroblasts, as well as in vivo mouse brain editing, all contexts in which the pegRNAs generated using OptiPrime’s default parameters performed well.

OptiPrime performs less accurately in cases in which Cas9 nuclease scores fail to correlate with PE efficiency (Figure 4G). Similarly, during our exploratory data analysis of PE3 data, we showed that Cas9 nuclease scores fail to predict nsgRNA performance (Figure 5C). In future studies, we hope to further improve the accuracy of OptiPrime by replacing Rule Set 3 scores with PE-specific heuristics for protospacer efficiency. Moreover, we speculate that further studies elucidating the DNA repair mechanisms governing PE3, PE3b, and twinPE can lead to improved performance when selecting optimal nsgRNAs and twinPE pegRNA pairs.

Similar to other models that predict PE efficiency based on pegRNA sequences, OptiPrime was primarily trained on data from high-throughput PE experiments performed on a library of synthetic reporters randomly integrated into the genome. Information about chromatin context is therefore not incorporated into OptiPrime. Researchers interested in augmenting OptiPrime predictions with chromatin state information can use models such as ePRIDICT, which predicts relative editing efficiency at genomic loci in K562 cells without considering pegRNA sequences. Future studies may integrate chromatin state information into OptiPrime directly, although the cell type dependence of chromatin state may limit the generality of such an approach.

We found that empirical editing efficiencies cluster based on pegRNA similarity. While a deep exploration of this phenomenon to further increase the efficiency of navigating OptiPrime-nominated pegRNA space is the subject of ongoing efforts, and comparisons in this work were made simply by taking the top ranked outputs from OptiPrime, we expect that future work may integrate OptiPrime with iterative pegRNA optimization using techniques such as Bayesian optimization to further streamline the development of PE strategies by minimizing the number of pegRNAs tested to achieve efficient prime editing.

Methods

Culture conditions for immortalized cell lines

HEK293T and HeLa cells were cultured in Dulbecco’s Modified Eagle Medium (DMEM) plus GlutaMAX (Thermo Fisher) supplemented with 10% fetal bovine serum (FBS) (Thermo Fisher). Both cell types were passaged every 2–3 days, maintained below 80% confluency, cultured at 37 °C with 5% CO2, and tested negative for mycoplasma.

Isolation of and culture conditions for primary human fibroblasts

With IRB approval, informed patient consent, and in accordance the Declaration of Helsinki Principles, a punch biopsy was obtained from a patient with recessive dystrophic epidermolysis bullosa (RDEB). The tissue was minced and primary fibroblasts were expanded and subsequently propagated in Alpha MEM and 1× final concentrations of non-essential amino acids, penicillin, streptomycin, GlutaMAX, and anti-oxidant supplement (all from Thermo Fisher Scientific), 20% fetal bovine serum (Atlas Biologicals, Fort Collins, CO), 0.5 ng/mL epidermal growth factor, and 10 ng/mL epidermal growth factor (each from Sigma-Aldrich, St Louis, MO).

Isolation of and culture conditions for primary human T cells

Blood units were obtained from unrelated donors (n = 3) without identifying information from Memorial Blood Centers (Saint Paul, MN). Peripheral blood mononuclear cells (PBMCs) were isolated from buffy coats using Lymphoprep (STEMCELL Technologies) and SepMate tubes (STEMCELL Technologies) and cryopreserved in CryoStor CS10 Cell Cryopreservation Media (STEMCELL Technologies). T-cells were isolated from PBMCs using the EasySep Human T Cell Isolation Kit (STEMCELL Technologies) and activated with Dynabeads Human T-Expander CD3/CD28 at a 3:1 bead to cell ratio in complete T-cell media comprised of: X-VIVO 15 (Lonza), 10% Human Ab serum (Valley Biomedical), 1× GlutaMAX, 1× penicillin/streptomycin (Life Technologies), 1× N-acetyl-l-cysteine (Sigma-Aldrich) and 300 IU/mL interleukin-2 (PeproTech).

Isolation of and culture conditions for mouse embryonic fibroblasts

To obtain mouse embryonic fibroblasts (MEFs), mice embryos were collected from timed matings of heterozygous JR #16894 C3.Cg-Kif1algdg/GrsrJ x JR #664 C57BL/6J between embryonic day (E) E12.5 and E14.5. Embryos were dissected out of uterine horns, washed in 70% ethanol, and placed in a 10 cm petri dish with PBS. Embryos were separated from the placenta, head, liver, and heart and blood clots were washed out. Embryo cells were minced with a razor blade and dissociated with 2 mL of trypsin-EDTA and incubated at 37 °C for 10 minutes with shaking. Cells were resuspended with an additional 4 mL of DMEM (Thermo Fisher) supplemented with 10% embryonic stem cell FBS (Thermo Fisher) and vigorously resuspended. Any residual debris was removed by centrifugation and cells were plated in a T75 flask. MEFs were maintained in DMEM with GlutaMAX (Thermo Fisher), PSN (Thermo Fisher), and 10% embryonic stem cell fetal bovine serum (Thermo Fisher) at 37 °C in 5% CO2. Media was changed after 48 h. Four days post-isolation, cells were washed with PBS, trypsinized (0.05% Trypsin-EDTA), and frozen at passage P1 in cell culture media supplemented with 10% DMSO, 5.0×106 cells/mL per cryo-vial. Cryo-preserved P1 MEFs were thawed and used for in vitro gene editing experiments.

Paired pegRNA–target site library cloning

A lentiviral transfer plasmid for library cloning (pPC1535) was designed to contain a human U6 promoter–2xBsmBI and a PuroR–T2A–BFP marker under expression from an EF1α promoter (pEF1α). An oligonucleotide library of paired pegRNAs (containing a 2×BsmBI sequence instead of the sgRNA scaffold) and target sites (Twist Biosciences) was PCR amplified with Kapa HiFi Master Mix DNA Polymerase (Roche) and purified with the QIAquick PCR Purification Kit (QIAGEN). Backbone plasmid (pPC1535) was digested with BsmBI-v2 (New England Biolabs), treated with Quick CIP (New England Biolabs), then gel purified on a 1% agarose gel and purified with 0.75× SPRIselect beads (Beckman Coulter). Digested pPC1535 and amplified oligonucleotide libraries were then cloned with Gibson assembly with NEBuilder HiFi DNA Assembly (New England Biolabs) and purified with 1× SPRIselect beads (Beckman Coulter). Next, assembled plasmids were electroporated into 10-beta electrocompetent E. coli (New England Biolabs) and plated on LB agar plates with 50 μg/mL carbenicillin. After cells were grown for ~16 h at 37 °C, colonies were scraped and plasmid DNA was isolated with the Plasmid Plus Midi Kit with endotoxin removal (QIAGEN).

For the second cloning step, in which the pegRNA scaffold was incorporated into the U6-driven pegRNAs, midiprepped plasmid was digested with BsmBI-v2 (New England Biolabs) and gel purified on a 1% agarose gel. Next, phosphorylated oligos encoding the BlpI-modified SpCas9 F+E sgRNA scaffold were ligated to digested plasmid with T4 DNA ligase (New England Biolabs) in the presence of Esp3I (ThermoFisher), then purified with 1× SPRIselect (Beckman Coulter). Assembled plasmids were electroporated into 10-beta electrocompetent E. coli (New England Biolabs) and plated on LB agar plates with 50 μg/mL carbenicillin. After cells were grown for ~16 h at 37 °C, colonies were scraped and plasmid DNA was isolated with the Plasmid Plus Midi Kit with endotoxin removal (QIAGEN).

Lentivirus production, functional titering, and transduction

The day before transfection, 10×106 HEK293T cells were seeded in 15-cm dishes with DMEM (ThermoFisher) supplemented with 10% FBS (Millipore Sigma). When cells reached 60% confluency (~16 h after seeding), cells were transfected with 13.3 μg lentiviral transfer plasmid, 6.7 μg pMD2.G (Addgene #12259), and 10 μg psPAX2 (Addgene #12260) with 120 μL Lipofectamine 2000 (ThermoFisher) according to the manufacturer’s protocol.

6 h post transfection, media was exchanged with fresh DMEM (ThermoFisher) supplemented with 10% FBS (Millipore Sigma). 48 h after transfection, viral supernatant was centrifuged at 3000× g for 15 min to remove cellular debris, filtered through a 0.45 mm PVDF filter (Corning), aliquoted, and stored at −80 °C. 2×106 HEK293T or HeLa cells were “spinfected” with 0, 10, 20, 50, 100, or 200 μL of a thawed aliquot of virus in the presence of 10 μg/mL polybrene (Sigma Aldrich) in 6-well plates at 1000× g for 120 min, then transferred to an incubator at 37 °C with 5% CO2. 24 h post-infection, cells were trypsinized, resuspended in PBS, and fixed with 4% w/v paraformaldehyde (Sigma Aldrich). Fixed cells were strained into a round bottom test tube with a 35 mm strainer cap (Falcon) and BFP expression was quantified with a CytoFlex LX cytometer to calculate the MOI.

To make the final cell libraries, cells were transduced with a volume of virus that would result in ~0.2 MOI with 1,000× coverage of each library member post-selection. 1 day after infection, cells were treated with 1 μg/mL puromycin (ThermoFisher) to select for cells with integrated library members 2 days after infection, an additional 2 μg/mL puromycin was added to cells.

Paired pegRNA–target site library screens

All screens were performed in biological duplicate and with PE2 and PE4 conditions on cells that had undergone at least 6 days total of puromycin selection. A single replicate of untreated cells was always sequenced for background subtraction of editing efficiency. The day before transfection, 10×106 HEK293T cells or 3×106 HeLa cells were seeded in DMEM (ThermoFisher) supplemented with 10% FBS (Millipore Sigma) into 15 cm dishes. Cells were transfected at 70% confluency (~16 h after seeding).

Each 15-cm dish of HEK293T cells was transfected with 6.7 μg of pCMV–PEmax and 3.3 μg of either pEF1α–RFP (PE2) or pEF1α–hMLH1dn (PE4) with 100 μL of Lipofectamine 2000 (ThermoFisher) according to manufacturer's protocols. Each 15 cm dish of HeLa cells was transfected with 25 μg of pCMV–PEmax–P2A–BSD and 25 μg of either pEF1α–RFP (PE2) or pEF1α–hMLH1dn (PE4) with 140 μL of TransIT-HeLa reagent (Mirus Bio; no HeLaMONSTER reagent used) according to manufacturer's protocols. 24 h following transfection of HeLa cells, cells were treated with 10 μg/mL blasticidin S (ThermoFisher) to select for expression of PEmax. 48 h following transfection of HeLa cells, the media was replaced with DMEM (ThermoFisher) supplemented with 10% FBS (Millipore Sigma), 1× Penicillin/Streptomycin (ThermoFisher), 10 μg/mL blasticidin S (ThermoFisher), and 1 μg/mL puromycin. 72 h following transfection of HEK293T cells or 120 h following transfection of HeLa cells, cells were trypsinized, pelleted, and stored at −80 °C.

High-throughput sequencing of pooled screens

Genomic DNA was extracted from all screen cells with the NucleoSpin Blood XL Maxi kit (Machery-Nagel). The entirety of the genomic DNA from each screen condition was used in the initial round of PCR (PCR1) to amplify the region containing the gRNA and target site. Each 100 mL PCR1 was performed with 5 μg of genomic DNA as template, 1 mM of each of Lib-FWD and Lib-REV, and 50 mL of NEBNext Ultra II Q5 Master Mix (New England Biolabs) on a BioRad C1000 thermal cycler with the following cycling parameters:

98 °C for 30 s,
22 cycles of
 98 °C for 10 s,
 65 °C for 75 s,
65 °C for 5 min.

The reactions were characterized by Agilent TapeStation. 1 mL of pooled PCRs1 from each condition were gel purified (QIAGEN) and further purified with 0.8× SPRISelect. A following PCR step (PCR2) enabled indexing of the samples through addition of barcodes and binding to the flow cell through addition of the P5/P7 Illumina sequences.

For each screen condition, 4×50 μL PCRs2 were performed. Each PCR2 was performed with 10 ng of purified PCR1 amplicon was used as template, 600 nM each of a PE-FWD primer and a PE-REV primer, and 25 μL of Kapa HiFi Master Mix DNA Polymerase (Roche) on a BioRad C1000 thermal cycler with the following cycling parameters:

98 °C for 3 min,
8 cycles of
 98 °C for 20 s,
 65 °C for 15 s,
 72 °C for 15 s,
72 °C for 1 min.

The reactions were characterized by Agilent TapeStation. Reactions were combined and purified with 0.8× SPRISelect, then normalized for concentration prior to pooling and sequencing. Libraries were sequenced on an Illumina NovaSeq 6000 with the S1 Reagent Kit v1.5 with 50 cycles for R1, 8 cycles for I1, 8 cycles for I2, and 150 cycles for R2.

Processing of pegRNA–target site library screen sequencing data

Sequencing data from the NovaSeq 6000 were demultiplexed into fastq.gz files for each individual screen based on the index reads with bcl2fastq2. In these data, R1 reads cover the pegRNA spacer, while R2 reads contain the sequence outcome resulting from PE as well as a barcode specifying pegRNA identity. Within each screen, R2 reads were demultiplexed into individual fastq.gz files based on the barcode sequence in R2. If the sequenced barcode did not match the sequenced spacer (from R1), then the read was discarded. These individual R2 fastq.gz files were then analyzed individually with CRISPResso2 93 to obtain editing and indel frequencies for each pegRNA. For CRISPResso analysis, R2 read sequences were aligned to a reference sequence in HDR mode with a quality cutoff of Q30. For each amplicon, the CRISPResso2 quantification window was positioned to include the entire sequence between the nick site and homology end with 10 additional bp of buffer on either side. Frequencies for these edits were quantified as:

(number of HDR-aligned reads)(number of reads aligned to any amplicon).

Frequencies for indels were quantified as:

(number of indel-containig reads)(number of reads aligning to any amplicon).

Preparation of mRNA

Prime editor mRNA was generated using in vitro transcription (IVT). Briefly, the prime editor transcript, containing a 5′ untranslated region (UTR), Kozak sequence, prime editor open reading frame, and 3′ UTR were PCR amplified from a template plasmid containing an inactive T7 (dT7) promoter. PCR primers repaired this dT7 promoter and also installed a 119 nt poly(A) tail. The purified PCR double-stranded DNA amplicon was used as an IVT template using the HiScribe T7 high-yield RNA synthesis kit (NEB). IVT was performed following the manufacturer’s optional protocol to include CleanCap reagent AG (Trilink) and substitute N1-methylpseudouridine-5′-triphosphate (Trilink) for uridine triphosphate, after which IVT template was degraded with DNase I (NEB). A Monarch Spin RNA Cleanup Kit (NEB) was used to purify mRNA from complete IVT reactions, and mRNA transcripts were reconstituted in nuclease-free water.

Preparation of AAV

Transfer vectors were cloned in a v1em twin prime editor AAV architecture92 adapted for the PE6b editor variant. Briefly, synthetic gene fragments (IDT) with appropriate restriction recognition sites and containing an nsgRNA cassette driven by a mouse U6 promoter and an epegRNA cassette driven by a human U6 promoter were installed into the C-terminal transfer plasmid via restriction digestion and T4 ligation (NEB). Plasmids were transformed into chemically competent NEB Stable cells (NEB), which were plated on LB agar plates with 50 μg/mL carbenicillin and grown at 30 °C overnight. Single colonies were picked and grown in 2xYT media with 50 μg/mL carbenicillin for 20 hours, and plasmids were isolated using a ZymoPURE II Plasmid Maxiprep Kit (Zymo Research).

Recombinant AAVs (AAV9) were produced in suspension HEK293T cells, using F17 media (ThermoFisher). Cell suspensions were incubated at 37 °C, 8% CO2, 80 RPM. 24 hours before transfection, cells were seeded in 500–1000 mL of media at ~1×106 cells/mL. The day after, cells (~2×106 cells/mL) were transfected with pHelper, pRepCap and pTransgene (2:1:1 ratio, 2 μg DNA per million cells) using Transporter 5 transfection reagent (Polysciences) with a 2:1 reagent:DNA ratio. Three days post-transfection, cells were pelleted via centrifugation at 2000 RPM for 12 minutes in Nalgene conical bottles. The supernatant was discarded, and cell pellets were stored at −20 °C until purification. Each pellet, corresponding to 500 mL of cell culture, was resuspended in 14 mL of 500 mM NaCl, 40 mM Tris-base, 10 mM MgCl2, with Salt Active Nuclease (ArcticZymes) at 100 U/mL. Afterwards, the lysate was clarified at 5000 RCF for 20 minutes and loaded onto a density step gradient containing OptiPrep (Cosmo Bio) at 60%, 40%, 25%, and 15% at a volume of 6, 6, 8, and 5 mL respectively in OptiSeal tubes (Beckman Coulter). The step gradients were spun in a Beckman Type 70ti rotor (Beckman Coulter) in a Sorvall WX+ ultracentrifuge (ThermoFisher) at 67,000 RPM for 1 hour and 15 minutes at 18 °C. Afterwards, ~4.5 mL of the 40–60% interface was extracted using a 16-gauge needle, filtered through a 0.22 μm PES filter, buffer exchanged with 100K MWCO protein concentrators (ThermoFisher) into PBS containing 0.001% Pluronic F-68, and concentrated down to a volume of 200–1000 μL. The concentrated virus was filtered through a 0.22 μm PES filter and stored at 4 °C.

Titers of purified AAVs were measured using an AAVpro Titration Kit (for Real Time PCR) Ver. 2 (Takara) with 1:50,000 final dilutions of AAV stocks as input, following the manufacturer’s protocol. Viral stocks were concentrated to a minimum of 2×1013 vg/mL using Pierce PES 100K MWCO Concentrators (ThermoFisher) pre-equilibrated with 0.001% F-68 in PBS (ThermoFisher) and sterile filtered using 0.22 μm GV Durapore centrifugal filter units (Millipore). Concentrated viral titers were confirmed using an AAVpro Titration Kit (for Real Time PCR) Ver.2 (Takara) with 1:50,000 final dilutions of AAV stocks as input, following the manufacturer’s protocol. Samples were sent to Plasmidsaurus for verification of AAV genome integrity via Oxford Nanopore long-read sequencing.

Transfection and genomic DNA processing of HEK293T cells

The day before transfection, approximately 16,000 HEK293T cells were plated into each well of a 96-well plate (Corning) in DMEM (ThermoFisher) supplemented with 10% FBS (Millipore Sigma). When cells reached 70% confluency (~18 h after plating), cells were transfected with variable amounts of plasmid DNA and 0.5 μL of Lipofectamine 2000 (ThermoFisher) diluted in Opti-MEM I (ThermoFisher) according to manufacturer’s protocols. For a typical 96-well transfection using standard prime editor systems, the following plasmid amounts were transfected:

  • PE2: 200 ng prime editor and 50 ng pegRNA

  • PE3: 200 ng prime editor, 50 ng pegRNA, and 15 ng nsgRNA

  • PE4: 200 ng prime editor, 100 ng MLH1dn, and 50 ng pegRNA

  • PE5: 200 ng prime editor, 100 ng MLH1dn, 50 ng pegRNA, and 15 ng nsgRNA

  • eePASSIGE: 100 ng prime editor, 100 ng Bxb1, 150 ng donor, 15 ng pegRNA 1, 15 ng pegRNA 2

Following transfection, cells were incubated at 37 °C with 5% CO2 for 72 h before lysis and genomic DNA harvest.

Genomic DNA was isolated from HEK293T cells following a custom lysis protocol and paramagnetic bead extraction. To lyse cells, growth medium was carefully removed from cell culture plates and 100 μL of lysis buffer (100 mM Tris-HCl pH 8.0, 200 mM NaCl, 5 mM EDTA, 0.05% SDS, 4.0 mg/mL proteinase K (QIAGEN), and 12.5 mM dithiothreitol) was added to each well of a 96-well plate. Plates with lysis buffer were incubated at 55 °C for 12–20 h with shaking. To extract genomic DNA, one volume of lysate was thoroughly mixed with one volume of Ampure XP paramagnetic beads (Beckman Coulter) and incubated for 5 min, separated on plate magnet, washed with 70% ethanol three times with bead resuspension, and eluted in 50 μL of nuclease-free water.

Electroporation of primary human fibroblasts

On the day of electroporation, the fibroblasts were trypsinized (final concentration of 0.05% in 1 mM EDTA/PBS) and enumerated using a Countess Cell Counter (Thermo Fisher Scientific). 100,000 cells were mixed with 1 μg of PEmax-encoding mRNA, 250 ng GFP-encoding mRNA, and 90 pmol of pegRNA. Cells were electroporated with the Neon NxT Electroporation System (Thermo Fisher Scientific) using a 10 μL tip using one 20 ms pulse of 1700 V. Cells then were plated in a 24 well plate in complete media and harvested four days later for genomic DNA isolation with the Monarch Spin gDNA Extraction Kit (New England Biolabs).

Electroporation of primary human T cells

48 h post stimulation, T cell activation beads were removed and the cells were replated for ~2 hours in T-cell media. Cells were harvested, counted, and 300,000 T cells were mixed with 1 μg of PEmax-encoding mRNA and 90 pmol of pegRNA. Cells were electroporated with the Neon NxT Electroporation System (Thermo Fisher Scientific) with 10 μL tips and Buffer T and E10 using three 10 ms pulses of 1400 V. Cells were then plated in a 48 well plate with 600 μL of complete T-cell media and then harvested at 72 hours by centrifugation and resuspended in QuickExtract DNA Extraction Solution (Biosearch Technologies). Plates with lysis buffer were incubated at 65 °C for 10 m with shaking.

Electroporation of mouse embryonic fibroblasts

Mouse embryonic fibroblasts were trypsinized at 70% confluency, counted, and washed with PBS. Per electroporation, 200,000 cells were resuspended into 20 μL of SE nucleofector solution with SE supplement solution (Lonza) containing 1 μg of prime editor mRNA, 90 pmol of synthetic pegRNA (IDT), and 60 pmol synthetic nsgRNA (Synthego) (if applicable). Cells were then electroporated with a 4D-Nucleofector with X Unit (Lonza) using program CM-130 and the SE Cell Line 4D-Nucleofector X Kit S (Lonza). Cells were allowed to recover with addition of 80 μL of prewarmed DMEM for 10 minutes then plated in a 24-well plate with 500 μL DMEM and collected after 72 hours by addition of lysis buffer.

Mouse strains

The leg dragger mutation (Kif1a L181F) arose as a spontaneous mutation on a subline of C3.NB-H2 mice (The Jackson Laboratory, JR #16894, MGI:5752896) and referred to as C3.Cg-Kif1algdg/GrsrJ mice. The lgdg mutation is maintained on a hybrid genetic background with C57BL/6J (The Jackson Laboratory, JR #664, B6J;C3-Kif1algdg/Lutzy). For the generation of B6J;C3-Kif1algdg mice, heterozygous C3.Cg-Kif1algdg/GrsrJ mice (JR #16894) were intercrossed with C57BL/6J mice (JR #664). Mice were genotyped at birth, wean age (~postnatal P28), and triple confirmed when reaching study end (e.g., necropsy for tissue collection). The genotyping primers are listed below and are applicable regardless of the genetic background. Furthermore, B6J;C3-Kif1algdg embryos (small piece of the tail, day of MEF isolation) and MEFs (small aliquot of cells during passaging and cryo-preservation) were genotyped using the same genotyping assay. The Jackson Laboratory Animal Care and Use Committee approved all mouse protocols.

For animal experiments, simplified genotype abbreviations, such as Kif1a+/+ (B6J;C3-Kif1a+/+) and Kif1algdg/+ (B6J;C3-Kif1algdg/+) mice, were used throughout.

Genotyping Primers (lgdg mutation):

  • Common Forward: 5′-ATGGAGATTTACTGTGAGCGAG-3′

  • Common Reverse: 5′-GGTCCTGGATGTCATTGTAGG-3′

  • WT Probe (HEX): 5′-TGTGGAGGACCTCTCCAAGCT-3′

  • lgdg mutant Probe (FAM): 5′-TGTGGAGGACTTCTCCAAGCT-3′

Intracerebroventricular injection

For neonatal ICV injection (targeting of brain ventricles located below the cortex), heterozygous C3.Cg-Kif1algdg/GrsrJ (JR #16894) males were crossed with C57BL/6J (JR #664) females. C57BL/6J females were singly housed 10–12 days post mating set up. The week of birth, females were checked daily for the birth of the litter. On the day of birth (postnatal day P0), the litter of B6J;C3-Kif1algdg mice was injected, and operators were blinded to genotype and the treatment (test articles AAV9-OP5.4-+65 and AAV9-OP5.4-+114). For gene editing studies, mice were bilaterally injected with either prime editor (total dose: 1.1×1011 vg/mouse, consisting of 5×1010 vg each N- and C-terminal vector and 1×1010 vg EGFP-KASH vector, 2.5μl per hemisphere). Uninjected B6J;C3-Kif1algdg mice were used as a control group for each prime editor. Post-injection, B6J;C3-Kif1algdg pups were monitored daily for survival.

Necropsy and tissue collection

Mice were euthanized via CO2 asphyxiation 28 days after birth. Brain tissues were harvested and dissected into cortex, brainstem, hippocampus, and cerebellum. Samples were frozen on dry ice and stored at −80 °C until further analysis.

Nuclear isolation, sorting, nucleic acid extraction, and cDNA synthesis

Frozen brain tissues were resuspended in 2 mL of Nuclei Extraction Buffer (Miltenyi Biotec) supplemented with 0.2 U/μL Murine RNase Inhibitor (NEB) and transferred to gentleMACS C Tubes (Miltenyi Biotec). Samples were homogenized using the default program “4C_nuclei_1” on a gentleMACS Octo Dissociator (Miltenyi Biotec). The nuclei suspension was washed with an additional 2 mL of Nuclei Extraction Buffer, filtered through a 100 μm MACS SmartStrainer (Miltenyi Biotec) into a 5 mL centrifuge tube, and centrifuged at 500×g at 4 °C for 5 minutes. The supernatant was aspirated, and the pelleted nuclei were resuspended in 4 mL of ice-cold nuclei suspension buffer consisting of PBS supplemented with 100 μg/μL albumin (NEB), 3.33 μM Vybrant DyeCycle Ruby Stain (Thermo Fisher), and 0.2 U/μL RNase Inhibitor, Murine (NEB). Samples were centrifuged again at 500×g at 4 °C for 5 minutes, and the supernatant was aspirated. The pellet was resuspended in 1 mL of ice-cold nuclei suspension buffer and filtered through a 35 μL cell strainer into a 5 mL Falcon tube (Corning). Nuclei were sorted using a Sony SH800 Cell Sorter (Sony Biotechnology) with a 100 μm chip in the Broad Institute Technology Space. Singlet DyeCycle Ruby-positive events were collected and back-gated to establish forward and side scatter area gates for nuclei. Nuclei were then sorted based on GFP fluorescence to yield “bulk” (containing all nuclei irrespective of GFP expression) and “GFP-positive” populations. For each population, 25,000 nuclei per tissue sample were sorted into 540 μL of Buffer RLT Plus from the AllPrep DNA/RNA Mini kit (QIAGEN) supplemented with 74 mM DTT. DNA and RNA were purified according to the manufacturer’s protocol. The SuperScript IV First-Strand Synthesis System with ezDNase was used with kit-supplied random hexamers to generate cDNA from 9 μL of RNA input, following the manufacturer’s protocol. Synthesis reactions were cleaned up using SPRIselect (Beckman Coulter) at a 1.8x bead ratio and eluted in 35 μL of nuclease-free water. For sequencing analysis, 10 μL of purified cDNA was used as input for PCR1.

High-throughput sequencing of genomic loci

To assess gene editing, loci were amplified from genomic DNA samples with two rounds of PCR and deep sequenced. Briefly, an initial PCR step (PCR1) amplified the genomic sequence of interest with synthetic primers (IDT) containing Illumina forward and reverse I5/I7 adapters. Each 20 μL PCR1 reaction was performed with 500 nM of each primer, 0.8 to 1.0 μL of genomic DNA, 1× SYBR Green I, and 10 μL Q5 High-Fidelity 2× Master Mix (New England Biolabs) on a CFX96 Touch Real-Time PCR detection System with the following thermocycling conditions:

98 °C for 3 min,
30 cycles of
 98 °C for 20 s,
 65 °C for 15 s,
 72 °C for 15 s,
72 °C for 2 min

PCR1 reactions were monitored with SYBR Green I fluorescence to avoid over-amplification.

The subsequent PCR step (PCR2) added unique i7 and i5 Illumina barcode combinations to both ends of the PCR1 DNA fragment to enable demultiplexing. Each PCR2 was performed with 10 ng of purified PCR1 amplicon was used as template, 600 nM each of a PE-FWD primer and a PE-REV primer, and 25 μL of Kapa HiFi Master Mix DNA Polymerase (Roche) on a BioRad C1000 thermal cycler with the following cycling parameters:

98 °C for 3 min,
8 cycles of
 98 °C for 20 s,
 65 °C for 15 s,
 72 °C for 15 s,
72 °C for 1 min

PCR2 products were pooled by common amplicons, separated by electrophoresis on a 1% agarose gel, purified using the QIAquick Gel Extraction Kit (QIAGEN), and eluted in nuclease-free water. DNA amplicon libraries were quantified with a Qubit 3.0 Fluorimeter (ThermoFisher), then normalized, pooled, and sequenced using the Illumina MiSeq Reagent Kit v2 with 280–300 single-read cycles.

Quantification of editing in amplicon sequencing data

Sequencing reads were demultiplexed with bcl2fastq2. Amplicon sequences were aligned to a reference and edited sequence with CRISPResso293 in HDR mode using the parameters-q 30 --discard_indel_reads. For each amplicon, the CRISPResso2 quantification window was positioned to include the entire sequence between pegRNA- and (where applicable) sgRNA-directed Cas9 cut sites as well as an additional 10 bp beyond both cut sites. All prime editing efficiencies describe percentage of

(number of HDR-aligned reads)(number of reads aligned to any amplicon).

All prime editing indel rates describe percentage of

(number of aligned reads discarded)(number of reads aligning to any amplicon).

Design of paired pegRNA–target site libraries

For both Lib-MMR and Lib-CV, we designed pegRNAs that we hypothesized would give high editing efficiencies while matching the criteria in Tables 1 and 2. For pegRNAs that encode substitutions, we used an initial RTT homology length of 9+L, where L denotes the number of substituted bases. For pegRNAs that encode insertions or deletions, we used an initial RTT homology length of 19+L, where L is the absolute length difference between the unedited and edited sequences. These initial RTT lengths were then altered (if necessary) to find the closest RTT length such that the first base of the RTT would not be a cytosine. All pegRNAs used PBS length 13, since we hypothesized that this PBS length would provide sufficient signal to learn MMR preferences while not occupying excessive library space.

Model implementation and training

OptiPrime and all related models were implemented in Python 3.11 with jax (v.0.4.30) and flax (v.0.10.2). Stochastic gradient descent was performed with optax (v.0.2.4) using the AdamW optimizer with a linear warm-up ramp over 25 epochs to an initial learning rate of 0.001, then exponentially decaying by a factor of 0.99 each epoch for 75 epochs. Model hyperparameters were optimized manually with 5-fold stratified cross-validation. For detailed modeling considerations, readers are encouraged to read Supplemental Text 1.

We used code provided by each respective initial report to evaluate retrospective predictions made by PRIDICT2.0 and DeepPrime. In each case, we chose the model version most appropriate for the cellular system being studied, as listed in the caption for the respective figures. We validated that our code produced the correct outputs by comparing to outputs from each model’s respective webserver.

PE3 models were implemented in Python 3.11 with xgboost (v.3.0.0). In addition to outputs from OptiPrime, the following hand-crafted PE3-specific features were implemented and used:

  • nick_offset: Distance (bp) between pegRNA and nsgRNA nicks

  • hom_offset: Distance (bp) between end of RTT homology arm and nsgRNA nick

  • PE3b: Whether the nsgRNA spacer only matches its target after flap synthesis

  • pe_type: Whether PE3 or PE5 is used

  • cell_type: Categorical encoding of the cell type used

Selection of pegRNAs for prospective use cases

To best simulate prospective use cases, we entered the mutation that we wanted to correct into OptiPrime, PRIDICT2.0, or DeepPrime. For all experiments with OptiPrime, we used our default settings (HeLa cells; PE2max; epegRNAs) and took the top pegRNAs as ordered by silent edit combination. For experiments with PRIDICT2.0, we selected outputs as ranked by HEK score for experiments performed in HEK293T cells and as ranked by K562 score for all other experiments. For experiments with DeepPrime, we selected outputs as ranked by HEK293T_PE2max-e for experiments performed in HEK293T cells and as ranked by A549_PE2max-e for all other experiments, since the original authors only performed screens in A549 cells with both PE2max and epegRNAs.

Quantification and statistical analysis

Correlation coefficients were calculated in Python with the scipy.stats package (v.1.14.1). Unless otherwise noted, r denotes a Pearson (linear) correlation and ρ denotes a Spearman rank correlation. Accordingly, we compute and report r when assessing quantitative predictive accuracy between correlates and ρ when assessing the ability to rank-order outcomes. P-values reported for both types of correlation coefficient represent the two-tailed significance of observing a correlation of the given magnitude under Student’s t-distribution with n − 2 degrees of freedom. We used the two-sided Steiger’s Z-test94 to compute P-values of differences between correlations on identical datasets to compare model performances. The number of independent biological replicates and technical replicates for each experiment are described in either the figure legends or the Methods section. Unless otherwise specified, all error bars represent standard deviation.

Supplementary Material

Supplement 1
media-1.pdf (4.2MB, pdf)

Acknowledgements

Funding was provided by US NIH grants U01AI142756 (D.R.L.), RM1HG009490 (D.R.L), R01EB022376 (D.R.L.), R35GM118062 (D.R.L.), R01HL147324 (M.J.O.), P01CA065493 (M.J.O.), R01AR063070 (M.J.O.), DP2CA281401 (B.P.K.), and P01HL142494 (B.P.K.), the Howard Hughes Medical Institute, the Bill & Melinda Gates Foundation. This work was supported by the NIH IGNITE Program (R61NS133266-02 to M.T. and C.M.L.). The authors acknowledge funding from NSF Graduate Research Fellowships to A.H., P.J.C., A.C., A.A.S., and M.W.S., and a Natural Sciences and Engineering Research Council of Canada (NSERC) Postgraduate Scholarship-Doctoral (PGS D – 567791) to R.A.S.

The authors thank Peyton Randolph, Jordan Doman, and Żaneta Matuszek for helpful discussions and John Harvey and Benjamin Deverman for technical assistance. We thank the Jackson Laboratory (JAX), all RDTC (Rare Disease Translational Center) and JCPG (JAX Center for Precision Genetics) members for their operational support.

Data availability

Raw high-throughput sequencing data have been deposited at the NCBI Sequence Read Archive (SRA) under accession PRJNA1314411 and are publicly available as of the date of publication. This paper analyzes existing, publicly available high-throughput sequencing data, accessible on the SRA (PRJNA735408, PRJNA1055086, PRJNA1211588). Plasmids generated in this study will be available from Addgene as of the date of publication. Requests for other resources should be directed to David R. Liu (drliu@fas.harvard.edu).

Code availability

All original code relevant to training and modeling will be publicly available as of the date of publication under a non-commercial license. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

References

  • 1.Gao P., Lyu Q., Ghanam A.R., Lazzarotto C.R., Newby G.A., Zhang W., Choi M., Slivano O.J., Holden K., Walker J.A., et al. (2021). Prime editing in mice reveals the essentiality of a single base in driving tissue-specific gene expression. Genome Biology 22, 83. 10.1186/s13059-021-02304-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Gao X., Tao Y., Lamas V., Huang M., Yeh W.-H., Pan B., Hu Y.-J., Hu J.H., Thompson D.B., Shu Y., et al. (2018). Treatment of autosomal dominant hearing loss by in vivo delivery of genome editing agents. Nature 553, 217–221. 10.1038/nature25164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Geurts M.H., de Poel E., Pleguezuelos-Manzano C., Oka R., Carrillo L., Andersson-Rolf A., Boretto M., Brunsveld J.E., van Boxtel R., Beekman J.M., and Clevers H. (2021). Evaluating CRISPR-based prime editing for cancer modeling and CFTR repair in organoids. Life Science Alliance 4, e202000940. 10.26508/lsa.202000940. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Lin J., Liu X., Lu Z., Huang S., Wu S., Yu W., Liu Y., Zheng X., Huang X., Sun Q., Qiao Y., and Liu Z. (2021). Modeling a cataract disorder in mice with prime editing. Molecular Therapy Nucleic Acids 25, 494–501. 10.1016/j.omtn.2021.06.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Liu Y., Li X., He S., Huang S., Li C., Chen Y., Liu Z., Huang X., and Wang X. (2020). Efficient generation of mouse models with the prime editing system. Cell Discovery 6, 27. 10.1038/s41421-020-0165-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Musunuru K., Chadwick A.C., Mizoguchi T., Garcia S.P., DeNizio J.E., Reiss C.W., Wang K., Iyer S., Dutta C., Clendaniel V., et al. (2021). In vivo CRISPR base editing of PCSK9 durably lowers cholesterol in primates. Nature 593, 429–434. 10.1038/s41586-021-03534-y. [DOI] [PubMed] [Google Scholar]
  • 7.Newby G.A., and Liu D.R. (2021). In vivo somatic cell base editing and prime editing. Molecular Therapy 29, 3107–3124. 10.1016/j.ymthe.2021.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Park S.-J., Jeong T.Y., Shin S.K., Yoon D.E., Lim S.-Y., Kim S.P., Choi J., Lee H., Hong J.-I., Ahn J., Seong J.K., and Kim K. (2021). Targeted mutagenesis in mouse cells and embryos using an enhanced prime editor. Genome Biology 22, 170. 10.1186/s13059-021-02389-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Rothgangl T., Dennis M.K., Lin P.J.C., Oka R., Witzigmann D., Villiger L., Qi W., Hruzova M., Kissling L., Lenggenhager D., et al. (2021). In vivo adenine base editing of PCSK9 in macaques reduces LDL cholesterol levels. Nature Biotechnology 39, 949–957. 10.1038/s41587-021-00933-4. [DOI] [Google Scholar]
  • 10.Schene I.F., Joore I.P., Oka R., Mokry M., van Vugt A.H.M., van Boxtel R., van der Doef H.P.J., van der Laan L.J.W., Verstegen M.M.A., van Hasselt P.M., Nieuwenhuis E.E.S., and Fuchs S.A. (2020). Prime editing for functional repair in patient-derived disease models. Nature Communications 11, 5352. 10.1038/s41467-020-19136-7. [DOI] [Google Scholar]
  • 11.She K., Liu Y., Zhao Q., Jin X., Yang Y., Su J., Li R., Song L., Xiao J., Yao S., et al. (2023). Dual-AAV split prime editor corrects the mutation and phenotype in mice with inherited retinal degeneration. Signal Transduction and Targeted Therapy 8, 57. 10.1038/s41392-022-01234-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Sousa A.A., Hemez C., Lei L., Traore S., Kulhankova K., Newby G.A., Doman J.L., Oye K., Pandey S., Karp P.H., McCray P.B., and Liu D.R. (2024). Systematic optimization of prime editing for the efficient functional correction of CFTR F508del in human airway epithelial cells. Nature Biomedical Engineering. 10.1038/s41551-024-01233-3. [DOI] [Google Scholar]
  • 13.Everette K.A., Newby G.A., Levine R.M., Mayberry K., Jang Y., Mayuranathan T., Nimmagadda N., Dempsey E., Li Y., Bhoopalan S.V., et al. (2023). Ex vivo prime editing of patient haematopoietic stem cells rescues sickle-cell disease phenotypes after engraftment in mice. Nature Biomedical Engineering 7, 616–628. 10.1038/s41551-023-01026-0. [DOI] [Google Scholar]
  • 14.Yeh W.-H., Chiang H., Rees H.A., Edge A.S.B., and Liu D.R. (2018). In vivo base editing of post-mitotic sensory cells. Nature Communications 9, 2184. 10.1038/s41467-018-04580-3. [DOI] [Google Scholar]
  • 15.Lee H.K., Willi M., Smith H.E., Miller S.M., Liu D.R., Liu C., and Hennighausen L. (2019). Simultaneous targeting of linked loci in mouse embryos using base editing. Scientific Reports 9, 1662. 10.1038/s41598-018-33533-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Osborn M.J., Newby G.A., McElroy A.N., Knipping F., Nielsen S.C., Riddle M.J., Xia L., Chen W., Eide C.R., Webber B.R., et al. (2020). Base Editor Correction of COL7A1 in Recessive Dystrophic Epidermolysis Bullosa Patient-Derived Fibroblasts and iPSCs. Journal of Investigative Dermatology 140, 338–347.e335. 10.1016/j.jid.2019.07.701. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Song C.-Q., Jiang T., Richter M., Rhym L.H., Koblan L.W., Zafra M.P., Schatoff E.M., Doman J.L., Cao Y., Dow L.E., et al. (2020). Adenine base editing in an adult mouse model of tyrosinaemia. Nature Biomedical Engineering 4, 125–130. 10.1038/s41551-019-0357-8. [DOI] [Google Scholar]
  • 18.Yeh W.-H., Shubina-Oleinik O., Levy J.M., Pan B., Newby G.A., Wornow M., Burt R., Chen J.C., Holt J.R., and Liu D.R. (2020). In vivo base editing restores sensory transduction and transiently improves auditory function in a mouse model of recessive deafness. Science Translational Medicine 12, eaay9101. doi: 10.1126/scitranslmed.aay9101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hanna R.E., Hegde M., Fagre C.R., DeWeirdt P.C., Sangree A.K., Szegletes Z., Griffith A., Feeley M.N., Sanson K.R., Baidi Y., et al. (2021). Massively parallel assessment of human variants with base editor screens. Cell 184, 1064–1080.e1020. 10.1016/j.cell.2021.01.012. [DOI] [PubMed] [Google Scholar]
  • 20.Koblan L.W., Erdos M.R., Wilson C., Cabral W.A., Levy J.M., Xiong Z.-M., Tavarez U.L., Davison L.M., Gete Y.G., Mao X., et al. (2021). In vivo base editing rescues Hutchinson–Gilford progeria syndrome in mice. Nature 589, 608–614. 10.1038/s41586-020-03086-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Krishnamurthy S., Traore S., Cooney A.L., Brommel C.M., Kulhankova K., Sinn Patrick L., Newby Gregory A., Liu David R., and McCray Paul B. Jr (2021). Functional correction of CFTR mutations in human airway epithelial cells using adenine base editors. Nucleic Acids Research 49, 10558–10572. 10.1093/nar/gkab788. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Newby G.A., Yen J.S., Woodard K.J., Mayuranathan T., Lazzarotto C.R., Li Y., Sheppard-Tillman H., Porter S.N., Yao Y., Mayberry K., et al. (2021). Base editing of haematopoietic stem cells rescues sickle cell disease in mice. Nature 595, 295–302. 10.1038/s41586-021-03609-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Suh S., Choi E.H., Leinonen H., Foik A.T., Newby G.A., Yeh W.-H., Dong Z., Kiser P.D., Lyon D.C., Liu D.R., and Palczewski K. (2021). Restoration of visual function in adult mice with an inherited retinal disease via adenine base editing. Nature Biomedical Engineering 5, 169–178. 10.1038/s41551-020-00632-6. [DOI] [Google Scholar]
  • 24.Choi E.H., Suh S., Foik A.T., Leinonen H., Newby G.A., Gao X.D., Banskota S., Hoang T., Du S.W., Dong Z., et al. (2022). In vivo base editing rescues cone photoreceptors in a mouse model of early-onset inherited retinal degeneration. Nature Communications 13, 1830. 10.1038/s41467-022-29490-3. [DOI] [Google Scholar]
  • 25.Knipping F., Newby G.A., Eide C.R., McElroy A.N., Nielsen S.C., Smith K., Fang Y., Cornu T.I., Costa C., Gutierrez-Guerrero A., et al. (2022). Disruption of HIV-1 co-receptors CCR5 and CXCR4 in primary human T cells and hematopoietic stem and progenitor cells using base editing. Molecular Therapy 30, 130–144. 10.1016/j.ymthe.2021.10.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Mouri K., Guo M.H., de Boer C.G., Lissner M.M., Harten I.A., Newby G.A., DeBerg H.A., Platt W.F., Gentili M., Liu D.R., et al. (2022). Prioritization of autoimmune disease-associated genetic variants that perturb regulatory element activity in T cells. Nature Genetics 54, 603–612. 10.1038/s41588-022-01056-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Sheriff A., Guri I., Zebrowska P., Llopis-Hernandez V., Brooks I.R., Tekkela S., Subramaniam K., Gebrezgabher R., Naso G., Petrova A., et al. (2022). ABE8e adenine base editor precisely and efficiently corrects a recurrent COL7A1 nonsense mutation. Scientific Reports 12, 19643. 10.1038/s41598-022-24184-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Suh S., Choi E.H., Raguram A., Liu D.R., and Palczewski K. (2022). Precision genome editing in the eye. Proceedings of the National Academy of Sciences 119, e2210104119. doi: 10.1073/pnas.2210104119. [DOI] [Google Scholar]
  • 29.Arbab M., Matuszek Z., Kray K.M., Du A., Newby G.A., Blatnik A.J., Raguram A., Richter M.F., Zhao K.T., Levy J.M., et al. (2023). Base editing rescue of spinal muscular atrophy in cells and in mice. Science 380, eadg6518. doi: 10.1126/science.adg6518. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Cooney A.L., Brommel C.M., Traore S., Newby G.A., Liu D.R., McCray P.B., and Sinn P.L. (2023). Reciprocal mutations of lung-tropic AAV capsids lead to improved transduction properties. Frontiers in Genome Editing 5. 10.3389/fgeed.2023.1271813. [DOI] [Google Scholar]
  • 31.Joynt A.T., Kavanagh E.W., Newby G.A., Mitchell S., Eastman A.C., Paul K.C., Bowling A.D., Osorio D.L., Merlo C.A., Patel S.U., et al. (2023). Protospacer modification improves base editing of a canonical splice site variant and recovery of CFTR function in human airway epithelial cells. Molecular Therapy - Nucleic Acids 33, 335–350. 10.1016/j.omtn.2023.06.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kabra M., Shahi P.K., Wang Y., Sinha D., Spillane A., Newby G.A., Saxena S., Tong Y., Chang Y., Abdeen A.A., et al. (2023). Nonviral base editing of KCNJ13 mutation preserves vision in a model of inherited retinal channelopathy. The Journal of Clinical Investigation 133. 10.1172/JCI171356. [DOI] [Google Scholar]
  • 33.Li C., Georgakopoulou A., Newby G.A., Chen P.J., Everette K.A., Paschoudi K., Vlachaki E., Gil S., Anderson A.K., Koob T., et al. (2023). In vivo HSC prime editing rescues sickle cell disease in a mouse model. Blood 141, 2085–2099. 10.1182/blood.2022018252. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Li H., Ma T., Remsberg J.R., Won S.J., DeMeester K.E., Njomen E., Ogasawara D., Zhao K.T., Huang T.P., Lu B., et al. (2023). Assigning functionality to cysteines by base editing of cancer dependency genes. Nature Chemical Biology 19, 1320–1330. 10.1038/s41589-023-01428-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Martin-Rufino J.D., Castano N., Pang M., Grody E.I., Joubran S., Caulier A., Wahlster L., Li T., Qiu X., Riera-Escandell A.M., et al. (2023). Massively parallel base editing to map variant effects in human hematopoiesis. Cell 186, 2456–2474.e2424. 10.1016/j.cell.2023.03.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Mayuranathan T., Newby G.A., Feng R., Yao Y., Mayberry K.D., Lazzarotto C.R., Li Y., Levine R.M., Nimmagadda N., Dempsey E., et al. (2023). Potent and uniform fetal hemoglobin induction via base editing. Nature Genetics 55, 1210–1220. 10.1038/s41588-023-01434-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.McAuley G.E., Yiu G., Chang P.C., Newby G.A., Campo-Fernandez B., Fitz-Gibbon S.T., Wu X., Kang S.-H.L., Garibay A., Butler J., et al. (2023). Human T cell generation is restored in CD3δ severe combined immunodeficiency through adenine base editing. Cell 186, 1398–1416.e1323. 10.1016/j.cell.2023.02.027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Peters C.W., Hanlon K.S., Ivanchenko M.V., Zinn E., Linarte E.F., Li Y., Levy J.M., Liu D.R., Kleinstiver B.P., Indzhykulian A.A., and Corey D.P. (2023). Rescue of hearing by adenine base editing in a humanized mouse model of Usher syndrome type 1F. Molecular Therapy 31, 2439–2453. 10.1016/j.ymthe.2023.06.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Reichart D., Newby G.A., Wakimoto H., Lun M., Gorham J.M., Curran J.J., Raguram A., DeLaughter D.M., Conner D.A., Marsiglia J.D.C., et al. (2023). Efficient in vivo genome editing prevents hypertrophic cardiomyopathy in mice. Nature Medicine 29, 412–421. 10.1038/s41591-022-02190-7. [DOI] [Google Scholar]
  • 40.Tao Y., Lamas V., Du W., Zhu W., Li Y., Whittaker M.N., Zuris J.A., Thompson D.B., Rameshbabu A.P., Shu Y., et al. (2023). Treatment of monogenic and digenic dominant genetic hearing loss by CRISPR-Cas9 ribonucleoprotein delivery in vivo. Nature Communications 14, 4928. 10.1038/s41467-023-40476-7. [DOI] [Google Scholar]
  • 41.Du S.W., Newby G.A., Salom D., Gao F., Menezes C.R., Suh S., Choi E.H., Chen P.Z., Liu D.R., and Palczewski K. (2024). In vivo photoreceptor base editing ameliorates rhodopsin-E150K autosomal-recessive retinitis pigmentosa in mice. Proceedings of the National Academy of Sciences 121, e2416827121. 10.1073/pnas.2416827121. [DOI] [Google Scholar]
  • 42.Ely Z.A., Mathey-Andrews N., Naranjo S., Gould S.I., Mercer K.L., Newby G.A., Cabana C.M., Rideout W.M., Jaramillo G.C., Khirallah J.M., et al. (2024). A prime editor mouse to model a broad spectrum of somatic mutations in vivo. Nature Biotechnology 42, 424–436. 10.1038/s41587-023-01783-y. [DOI] [Google Scholar]
  • 43.Gould S.I., Wuest A.N., Dong K., Johnson G.A., Hsu A., Narendra V.K., Atwa O., Levine S.S., Liu D.R., and Sánchez Rivera F.J. (2024). High-throughput evaluation of genetic variants with prime editing sensor libraries. Nature Biotechnology. 10.1038/s41587-024-02172-9. [DOI] [Google Scholar]
  • 44.Kennedy P.H., Alborzian Deh Sheikh A., Balakar M., Jones A.C., Olive M.E., Hegde M., Matias M.I., Pirete N., Burt R., Levy J., et al. (2024). Post-translational modification-centric base editor screens to assess phosphorylation site functionality in high throughput. Nature Methods 21, 1033–1043. 10.1038/s41592-024-02256-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Li C., Georgakopoulou A., Paschoudi K., Anderson A.K., Huang L., Gil S., Giannaki M., Vlachaki E., Newby G.A., Liu D.R., et al. (2024). Introducing a hemoglobin G-Makassar variant in HSCs by in vivo base editing treats sickle cell disease in mice. Molecular Therapy 32, 4353–4371. 10.1016/j.ymthe.2024.10.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Schmidt R., Ward C.C., Dajani R., Armour-Garb Z., Ota M., Allain V., Hernandez R., Layeghi M., Xing G., Goudy L., et al. (2024). Base-editing mutagenesis maps alleles to tune human T cell functions. Nature 625, 805–812. 10.1038/s41586-023-06835-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Sousa A.A., Terrey M., Sakai H.A., Simmons C.Q., Arystarkhova E., Morsci N.S., Anderson L.C., Xie J., Suri-Payer F., Laux L.C., et al. In vivo prime editing rescues alternating hemiplegia of childhood in mice. Cell. 10.1016/j.cell.2025.06.038. [DOI] [Google Scholar]
  • 48.Musunuru K., Grandinette Sarah A., Wang X., Hudson Taylor R., Briseno K., Berry Anne M., Hacker Julia L., Hsu A., Silverstein Rachel A., Hille Logan T., et al. (2025). Patient-Specific In Vivo Gene Editing to Treat a Rare Genetic Disease. New England Journal of Medicine 392, 2235–2243. 10.1056/NEJMoa2504747. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Anzalone A.V., Randolph P.B., Davis J.R., Sousa A.A., Koblan L.W., Levy J.M., Chen P.J., Wilson C., Newby G.A., Raguram A., and Liu D.R. (2019). Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157. 10.1038/s41586-019-1711-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Prime Medicine Announces Breakthrough Clinical Data Showing Rapid Restoration of DHR Positivity After Single Infusion of PM359, an Investigational Prime Editor for Chronic Granulomatous Disease. (2025). Prime Medicine. May/19/2025. https://investors.primemedicine.com/news-releases/news-release-details/prime-medicine-announces-breakthrough-clinical-data-showing. [Google Scholar]
  • 51.Hsu J.Y., Grünewald J., Szalay R., Shih J., Anzalone A.V., Lam K.C., Shen M.W., Petri K., Liu D.R., Joung J.K., and Pinello L. (2021). PrimeDesign software for rapid and simplified design of prime editing guide RNAs. Nature Communications 12, 1034. 10.1038/s41467-021-21337-7. [DOI] [Google Scholar]
  • 52.Doman J.L., Sousa A.A., Randolph P.B., Chen P.J., and Liu D.R. (2022). Designing and executing prime editing experiments in mammalian cells. Nature Protocols 17, 2431–2468. 10.1038/s41596-022-00724-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Chen P.J., and Liu D.R. (2023). Prime editing for precise and highly versatile genome manipulation. Nature Reviews Genetics 24, 161–177. 10.1038/s41576-022-00541-1. [DOI] [Google Scholar]
  • 54.Mathis N., Allam A., Kissling L., Marquart K.F., Schmidheini L., Solari C., Balázs Z., Krauthammer M., and Schwank G. (2023). Predicting prime editing efficiency and product purity by deep learning. Nature Biotechnology 41, 1151–1159. 10.1038/s41587-022-01613-7. [DOI] [Google Scholar]
  • 55.Yu G., Kim H.K., Park J., Kwak H., Cheong Y., Kim D., Kim J., Kim J., and Kim H.H. (2023). Prediction of efficiencies for diverse prime editing systems in multiple cell types. Cell 186, 2256–2272.e2223. 10.1016/j.cell.2023.03.034. [DOI] [PubMed] [Google Scholar]
  • 56.Mathis N., Allam A., Tálas A., Kissling L., Benvenuto E., Schmidheini L., Schep R., Damodharan T., Balázs Z., Janjuha S., et al. (2024). Machine learning prediction of prime editing efficiency across diverse chromatin contexts. Nature Biotechnology. 10.1038/s41587-024-02268-2. [DOI] [Google Scholar]
  • 57.Chen P.J., Hussmann J.A., Yan J., Knipping F., Ravisankar P., Chen P.-F., Chen C., Nelson J.W., Newby G.A., Sahin M., et al. (2021). Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635–5652.e5629. 10.1016/j.cell.2021.09.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Ferreira da Silva J., Oliveira G.P., Arasa-Verge E.A., Kagiou C., Moretton A., Timelthaler G., Jiricny J., and Loizou J.I. (2022). Prime editing efficiency and fidelity are enhanced in the absence of mismatch repair. Nature Communications 13, 760. 10.1038/s41467-022-28442-1. [DOI] [Google Scholar]
  • 59.Kim H.K., Yu G., Park J., Min S., Lee S., Yoon S., and Kim H.H. (2021). Predicting the efficiency of prime editing guide RNAs in human cells. Nature Biotechnology 39, 198–206. 10.1038/s41587-020-0677-y. [DOI] [Google Scholar]
  • 60.Liu F., Huang S., Hu J., Chen X., Song Z., Dong J., Liu Y., Huang X., Wang S., Wang X., and Shu W. (2023). Design of prime-editing guide RNAs with deep transfer learning. Nature Machine Intelligence 5, 1261–1274. 10.1038/s42256-023-00739-w. [DOI] [Google Scholar]
  • 61.Anzalone A.V., Gao X.D., Podracky C.J., Nelson A.T., Koblan L.W., Raguram A., Levy J.M., Mercer J.A.M., and Liu D.R. (2022). Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing. Nature Biotechnology 40, 731–740. 10.1038/s41587-021-01133-w. [DOI] [Google Scholar]
  • 62.Fishel R., and Kolodner R.D. (1995). Identification of mismatch repair genes and their role in the development of cancer. Current Opinion in Genetics & Development 5, 382–395. 10.1016/0959-437X(95)80055-7. [DOI] [PubMed] [Google Scholar]
  • 63.Warren J.J., Pohlhaus T.J., Changela A., Iyer R.R., Modrich P.L., and Beese Lorena S. (2007). Structure of the Human MutSα DNA Lesion Recognition Complex. Molecular Cell 26, 579–592. 10.1016/j.molcel.2007.04.018. [DOI] [PubMed] [Google Scholar]
  • 64.Gupta S., Gellert M., and Yang W. (2012). Mechanism of mismatch recognition revealed by human MutSβ bound to unpaired DNA loops. Nature Structural & Molecular Biology 19, 72–78. 10.1038/nsmb.2175. [DOI] [Google Scholar]
  • 65.Landrum M.J., Lee J.M., Benson M., Brown G., Chao C., Chitipiralla S., Gu B., Hart J., Hoffman D., Hoover J., et al. (2015). ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Research 44, D862–D868. 10.1093/nar/gkv1222. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Thomas D.C., Roberts J.D., and Kunkel T.A. (1991). Heteroduplex repair in extracts of human HeLa cells. Journal of Biological Chemistry 266, 3744–3751. 10.1016/S0021-9258(19)67858-0. [DOI] [PubMed] [Google Scholar]
  • 67.Raissi M., Yazdani A., and Karniadakis G.E. (2020). Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations. Science 367, 1026–1030. doi: 10.1126/science.aaw4741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Chen Z., King W.C., Hwang A., Gerstein M., and Zhang J. (2022). DeepVelo: Single-cell transcriptomic deep velocity field learning with neural ordinary differential equations. Science Advances 8, eabq3745. doi: 10.1126/sciadv.abq3745. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Chen R.T.Q., Rubanova Y., Bettencourt J., and Duvenaud D.K. (2018). Neural Ordinary Differential Equations. [Google Scholar]
  • 70.Resasco D.C., Gao F., Morgan F., Novak I.L., Schaff J.C., and Slepchenko B.M. (2012). Virtual Cell: computational tools for modeling in cell biology. WIREs Systems Biology and Medicine 4, 129–140. 10.1002/wsbm.165. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Schaff J., Fink C.C., Slepchenko B., Carson J.H., and Loew L.M. (1997). A general computational framework for modeling cellular structure and function. Biophysical Journal 73, 1135–1146. 10.1016/S0006-3495(97)78146-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A.J., Bambrick J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.DeWeirdt P.C., McGee A.V., Zheng F., Nwolah I., Hegde M., and Doench J.G. (2022). Accounting for small variations in the tracrRNA sequence improves sgRNA activity predictions for CRISPR screening. Nature Communications 13, 5255. 10.1038/s41467-022-33024-2. [DOI] [Google Scholar]
  • 75.Kim H.K., Kim Y., Lee S., Min S., Bae J.Y., Choi J.W., Park J., Jung D., Yoon S., and Kim H.H. SpCas9 activity prediction by DeepSpCas9, a deep learning–based model with high generalization performance. Science Advances 5, eaax9249. 10.1126/sciadv.aax9249. [DOI] [Google Scholar]
  • 76.Lundberg Scott M., S.-I.L. (2017). A Unified Approach to Interpreting Model Predictions. Proceedings of the 31st Conference on Neural Information Processing Systems. [Google Scholar]
  • 77.Pandey S., Gao X.D., Krasnow N.A., McElroy A., Tao Y.A., Duby J.E., Steinbeck B.J., McCreary J., Pierce S.E., Tolar J., et al. (2024). Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing. Nature Biomedical Engineering. 10.1038/s41551-024-01227-1. [DOI] [Google Scholar]
  • 78.Sherry S.T., Ward M.H., Kholodov M., Baker J., Phan L., Smigielski E.M., and Sirotkin K. (2001). dbSNP: the NCBI database of genetic variation. Nucleic Acids Research 29, 308–311. 10.1093/nar/29.1.308. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Perez G., Barber Galt P., Benet-Pages A., Casper J., Clawson H., Diekhans M., Fischer C., Gonzalez Jairo N., Hinrichs Angie S., Lee Christopher M., et al. (2025). The UCSC Genome Browser database: 2025 update. Nucleic Acids Research 53, D1243–D1249. 10.1093/nar/gkae974. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Walton R.T., Hsu J.Y., Joung J.K., and Kleinstiver B.P. (2021). Scalable characterization of the PAM requirements of CRISPR–Cas enzymes using HT-PAMDA. Nature Protocols 16, 1511–1547. 10.1038/s41596-020-00465-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Walton R.T., Christie K.A., Whittaker M.N., and Kleinstiver B.P. (2020). Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368, 290–296. doi: 10.1126/science.aba8853. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Marshall B. (2022). Cystic Fibrosis Foundation Patient Registry: 2021 Annual Data Report. Cystic Fibrosis Foundation. [Google Scholar]
  • 83.Kleinstiver B.P., Pattanayak V., Prew M.S., Tsai S.Q., Nguyen N.T., Zheng Z., and Joung J.K. (2016). High-fidelity CRISPR–Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529, 490–495. 10.1038/nature16526. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Has C., Nyström A., Saeidian A.H., Bruckner-Tuderman L., and Uitto J. (2018). Epidermolysis bullosa: Molecular pathology of connective tissue components in the cutaneous basement membrane zone. Matrix Biology 71-72, 313–329. 10.1016/j.matbio.2018.04.001. [DOI] [PubMed] [Google Scholar]
  • 85.Sockolosky J.T., Trotta E., Parisi G., Picton L., Su L.L., Le A.C., Chhabra A., Silveria S.L., George B.M., King I.C., et al. (2018). Selective targeting of engineered T cells using orthogonal IL-2 cytokine-receptor complexes. Science 359, 1037–1042. doi: 10.1126/science.aar3246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Doench J.G., Fusi N., Sullender M., Hegde M., Vaimberg E.W., Donovan K.F., Smith I., Tothova Z., Wilen C., Orchard R., et al. (2016). Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nature Biotechnology 34, 184–191. 10.1038/nbt.3437. [DOI] [Google Scholar]
  • 87.Sanson K.R., Hanna R.E., Hegde M., Donovan K.F., Strand C., Sullender M.E., Vaimberg E.W., Goodale A., Root D.E., Piccioni F., and Doench J.G. (2018). Optimized libraries for CRISPR-Cas9 genetic screens with multiple modalities. Nature Communications 9, 5416. 10.1038/s41467-018-07901-8. [DOI] [Google Scholar]
  • 88.Urnov F.D. (2021). Imagine CRISPR cures. Molecular Therapy 29, 3103–3106. 10.1016/j.ymthe.2021.10.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Esmaeeli Nieh S., Madou M.R.Z., Sirajuddin M., Fregeau B., McKnight D., Lexa K., Strober J., Spaeth C., Hallinan B.E., Smaoui N., et al. (2015). De novo mutations in KIF1A cause progressive encephalopathy and brain atrophy. Annals of Clinical and Translational Neurology 2, 623–635. 10.1002/acn3.198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Harris B.S.F., E. H.; Reinholdt L. G.; Berry M. L.; Bergstrom D. E.; Schroeder D. G.; Cox G. A. (2016). Two viable spontanenous mutation in Kif1a. National Institutes of Health. [Google Scholar]
  • 91.Doman J.L., Pandey S., Neugebauer M.E., An M., Davis J.R., Randolph P.B., McElroy A., Gao X.D., Raguram A., Richter M.F., et al. (2023). Phage-assisted evolution and protein engineering yield compact, efficient prime editors. Cell 186, 3983–4002.e3926. 10.1016/j.cell.2023.07.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Davis J.R., Banskota S., Levy J.M., Newby G.A., Wang X., Anzalone A.V., Nelson A.T., Chen P.J., Hennes A.D., An M., et al. (2024). Efficient prime editing in mouse brain, liver and heart with dual AAVs. Nature Biotechnology 42, 253–264. 10.1038/s41587-023-01758-z. [DOI] [Google Scholar]
  • 93.Clement K., Rees H., Canver M.C., Gehrke J.M., Farouni R., Hsu J.Y., Cole M.A., Liu D.R., Joung J.K., Bauer D.E., and Pinello L. (2019). CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nature Biotechnology 37, 224–226. 10.1038/s41587-019-0032-3. [DOI] [Google Scholar]
  • 94.Steiger J.H. (1980). Tests for comparing elements of a correlation matrix. 87, 245–251. 10.1037/0033-2909.87.2.245. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
media-1.pdf (4.2MB, pdf)

Data Availability Statement

Raw high-throughput sequencing data have been deposited at the NCBI Sequence Read Archive (SRA) under accession PRJNA1314411 and are publicly available as of the date of publication. This paper analyzes existing, publicly available high-throughput sequencing data, accessible on the SRA (PRJNA735408, PRJNA1055086, PRJNA1211588). Plasmids generated in this study will be available from Addgene as of the date of publication. Requests for other resources should be directed to David R. Liu (drliu@fas.harvard.edu).

All original code relevant to training and modeling will be publicly available as of the date of publication under a non-commercial license. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.


Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES