Abstract
Retrons are bacterial immune systems that use reverse transcribed DNA as a detector of phage infection. They are also increasingly deployed as a component of biotechnology. For genome editing, for instance, retrons are modified so that the reverse transcribed DNA (RT-DNA) encodes an editing donor. Retrons are commonly found in bacterial genomes; thousands of unique retrons have now been predicted bioinformatically. However, only a small number have been characterized experimentally. Here, we add substantially to the corpus of experimentally studied retrons. We synthesized >100 previously untested retrons to identify the natural sequence of RT-DNA they produce, quantify their RT-DNA production, and test the relative efficacy of editing using retron-derived donors to edit bacterial genomes, phage genomes, and human genomes. We add 62 new empirically determined, natural RT-DNAs, which are not predictable from the retron sequence alone. We report a large diversity in RT-DNA production and editing rates across retrons, finding that top performing editors are drawn from a subset of the retron phylogeny and outperform those used in previous studies, reaching precise editing rates of up to 40% in human cells.
INTRODUCTION
The signature of a retron – copious single-stranded DNA of a uniform length produced spontaneously by bacteria – was first observed in 19841. Five years later, it was shown that this DNA was the product of an endogenous reverse transcriptase (RT), the first RT discovered in prokaryotes2. The template for the reverse transcribed DNA (RT-DNA) of a retron is a short, highly structured RNA, which is produced from the same operon as the retron RT. The RT recognizes this structured, noncoding RNA (ncRNA) and partially reverse transcribes it into RT-DNA, polymerizing from a critical 2’ hydroxyl of the ncRNA3,4. The result is a transcriptional/reverse-transcriptional cascade that amplifies a single locus in the bacterial genome up to hundreds of copies of RT-DNA per cell5. This phenomenon of RT-DNA production was observed from 16 natural retrons over the next thirty years, without any conclusive evidence for a cellular function6.
More recently, a cellular role for retrons has become clear. Retrons confer phage resistance to bacteria7–10. Full mechanistic details are still emerging, but the general model is that retron RT-DNA is a sensor of phage infection. Either directly or indirectly, the phage modifies or degrades the RT-DNA, which releases an accessory retron protein that acts as a cellular toxin to remove the infected cell and spare the bacterial population8,9,11. Individual retrons have been shown to confer specific resistance to particular phages8,12. This specificity is due, at least in part, to the unique sequence of RT-DNA produced by the retron8. More than a thousand retrons have now been identified in genomic and metagenomic databases by homology to known retron RTs10. However, the sequence of the ncRNA has remained difficult to identify bioinformatically and the exact RT-DNA produced by each retron is, thus far, impossible to predict.
Retrons, like many other bacterial immune systems, have also proven valuable as a component of biotechnology. Specifically, the faithful production of abundant single-stranded DNA in cells by retrons has been used to produce templates for genome engineering, transcriptional receipts for molecular recorders, and transcription factor decoys13–17. Yet, to date, these technologies have been built using only a small number of the early-discovered retrons.
It is critical that we survey the diversity of retrons, both to inform investigations of bacterial immunity and to improve retron-based technologies. Of the thousand plus retrons identified10, only eighteen have actually been shown to produce RT-DNA6,8,9, only one has been used for bacterial recombineering13,16,18, and only eight have been tested for human genome editing18–20. Here, we address the lack of experimental testing across the diversity of predicted retrons by undertaking a census of RT-DNA production and editing in multiple cellular contexts. We provide validation of noncoding RNA sequence and RT-DNA production for retrons across the phylogeny, and report sequences of the unpredictable RT-DNA that are critical to inform new studies of function and technology. We identify the first subsets of retrons that appear to lack RT-DNA production, which will also inform future work in the field. We also find that particular subtypes of retrons produce higher amounts of RT-DNA and better editing performance.
RESULTS
In order to expand the corpus of experimentally validated retrons, we synthesized 163 never-before-tested retron RTs and ncRNAs distributed across various experiment types to enable quantification of RT-DNA production, as well as editing in bacterial, phage, and human genomes. These retrons were drawn from a set of 1,928 recently published retrons10 that were bioinformatically predicted based on protein homology to known retron RTs. While the RT component was previously annotated, the ncRNA component of retrons is difficult to predict across diverse retrons. We identified and annotated the cognate ncRNAs of each RT by simulating RNA folding in regions within 500bp of the RT and comparing the overall secondary structure to the consensus ncRNA structures. This consists of conserved characteristics such as an inverted repeat (the a1/a2 region of a retron ncRNA) >8bp, a priming guanine near the beginning of the ncRNA, and multiple hairpins in close proximity to a retron RT. Retrons have been recently shown to engage in phage defense that requires one or more additional accessory proteins, multiple of which have been shown to be cellular toxins8,9,12. The accessory proteins and phage defense mechanisms are extremely diverse among retrons10. For this work, we chose to focus on a census of the core reverse transcription machinery and use in biotechnology and thus excluded accessory genes from our analysis. However, it is important to note that in some cases the retron RT is fused to accessory domains and, in these instances, we did not truncate the protein.
RT-DNA production by a diverse set of retrons
For testing RT-DNA production, we synthesized 98 new retron RTs and ncRNAs, and one known standard, retron-Eco1 (ec86). These retrons were chosen to give a broad representation of the RT phylogeny (Fig 1a), including retrons from each RT clade (Fig 1b). Retron ncRNAs are frequently found immediately upstream of the retron RT. However, in this natural arrangement, putative ribosomal binding sites can be contained within the ncRNA, which makes standardization impossible. Therefore, we chose to invert the architecture and create operons driven by a T7/lac promoter, followed by a strong RBS, the RT coding region, and finally the ncRNA (Fig 1c). In this arrangement, no additional surrounding sequence other than the inverted repeat a1/a2 regions are present with the ncRNA, so in cases where an RT-DNA is produced, we can be sure that we have rigorously identified the ncRNA.
Figure 1.

RT-DNA production by a diverse set of retrons. a. Phylogenic analysis of retron RTs10. b. Number of retrons synthesized by clade (percentage of clade synthesized indicated in gray above each bar). Bar o indicates an orphan, clade-less retron. c. Schematic of RT-DNA quantification experiment. Bottom left shows the architecture of retrons for expression, right shows strategy for unbiased sequencing of RT-DNA. d. PAGE analysis of all retrons tested (composite of multiple gels, uncropped gels in Supplemental Figure 7) and proportion of retrons with detectable RT-DNA (bottom right donut). e. Quantification of RT-DNA production by density, relative to retron-Eco1 (marked in blue) arranged by order (terminal number) in the RT phylogeny shown in 1a. f. RT-DNA production by clade (one-way ANOVA, effect of clade P=0.0003). Bar o indicates an orphan, clade-less retron. Plot to the right splits production of the top clades by retron subtype. Bars show mean ±SEM. g. RT-DNA sequencing coverage plotted onto folded ncRNA for retron Mestre-1673. Orange indicates the bases that are reverse transcribed. h. Plots of RT-DNA sequencing coverage by ncRNA position for retron Mestre-1673. Conditions include pA or pC tailing of the RT-DNA for sequencing prep and inclusion or exclusion of DBR1 prior to sequencing prep. Dashed lines indicate the initiation and termination points, defined from the pA/DBR1 condition (initiation: coverage increases to ≥15% of maximum; termination: coverage decreases to ≤50% of maximum). Additional statistical details in Supplemental Table 2.
We transformed each of these 99 retrons (98 new plus one reference standard) into BL21-AI E. coli and induced expression of the synthetic operon for five hours. We prepared non-genomic nucleotides by midiprep and visualized these nucleotides on a TBE/Urea gel to quantify the presence of a characteristic RT-DNA band. Next, we prepared any resulting RT-DNA for sequencing. Because the retron RT-DNA sequences were unknown to us, we adapted an unbiased sequencing approach that we have previously used to examine Eco1 variant libraries12. Briefly, we added a single poly-nucleotide handle to the 5’ end of RT-DNA using TdT, used a complementary poly-nucleotide anchored primer and Klenow to synthesize a strand complementary to the RT-DNA, and then ligated adapters at the 3’ end of the now double-stranded RT-DNA (Fig 1c). The resulting molecules were indexed and sequenced on an Illumina machine.
We found that 62 of the 98 newly tested retrons produced RT-DNA that were detectable on a TBE/Urea gel, which varied in both length and abundance (Fig 1d). We quantified the intensity of these bands relative to a standard, retron-Eco1. When these relative production values are plotted by phylogenetic terminal, we observed that retrons in certain regions of the phylogeny produce more RT-DNA than others (Fig 1e). When examined by RT clade, we see greatest production from clades 2, 9, and 10; almost no production from clades 3 and 4; and intermediate production from the rest (Fig 1f). This can be further broken down by retron subtype – which is based on features of the accessory proteins rather than the RT itself – within the high producing clades. Here, we see that within clades 2 and 9, the highest producing retrons are confined to certain subtypes. This subtype dependence is also evident in some lower producing clades (Supplementary Fig 1).
We next used the sequencing data to determine the sequence of RT-DNA from retrons that produced it. We aligned RT-DNA sequences determined from the blind prep of single-stranded DNA to each retron ncRNA. An example of sequencing coverage is shown for Mestre-1673, plotting coverage onto the secondary structure of the ncRNA (Fig 1g). Each RT-DNA was prepped in parallel using four similar approaches, varying two parameters. First, we varied the nucleotide that we used to extend the sequence (poly adenine, pA, or poly cytosine, pC) for unbiased sequencing as TdT polymerase has nucleotide preferences that could affect the efficiency of the prep based on the terminating nucleotide of the RT-DNA. Second, we varied whether we pretreated the nucleotides with a debranching enzyme (DBR1) prior to sequencing prep. The characteristic 2’−5’ RNA to DNA linkage created when retron RTs polymerize from a structured RNA inhibits our sequencing preparation. For some retrons, this branched molecule is long lasting and DBR1 is required for efficient preparation. However, other retrons are debranched in vivo by nucleases that cleave in the RT-DNA. In these cases, DBR1 is not required for efficient sequencing preparation. By including both conditions, we can estimate whether a given retron remains branched or is debranched in vivo.
Figure 1h shows each of these four conditions for Mestre-1673, with sequencing coverage plotted over linear ncRNA position. In the pA/DBR1 condition (also shown in Figure 1g), we see even coverage of the RT-DNA region of the ncRNA. We define the initiation point (dashed line i) as the point where coverage increases to ≥15% of maximum coverage and termination (dashed line t) as the point where coverage decreases to ≤50% of maximum coverage. These initiation and termination points are propagated to the other conditions. In this case the pC/DBR1 condition closely matches the pA, but with fewer reads, which is typical of all RT-DNA sequencing. The conditions below, which lack DBR1 addition during sequencing prep, have substantially fewer reads. This is consistent with a retron that remains 2’−5’ branched in vivo. These coverage maps also illustrate a pattern that occurs frequently among retrons. In the absence of DBR1, the 5’ end of the RT-DNA begins further from the branched initiation point. This likely represents a fraction of RT-DNA that was debranched in vivo by cleavage of the RT-DNA at a position 3’ to the branch point or partial nuclease degradation from the free 5’ end of a debranched fraction.
Characteristics of RT-DNA production across retrons
We generated RT-DNA coverage maps for 54 retrons. Sequencing coverage (RT-DNA) plotted onto retron ncRNA secondary structures is shown for a subset of retrons in Figure 2a. Coverage and sequencing data for all retrons is provided in accompanying supplemental data (Supplemental Data 1). When quantified across retrons, the initiation point is typically quite stringent, with RT-DNA sequencing coverage rising from 15% of maximum to 85% of maximum within 3 nucleotides for >65% of retrons tested (Fig 2b). In contrast, termination tended to occur over a wider window, suggesting more variability in the termination point or some degradation of the 3’ end after termination (Fig 2c). The branching nucleotide of the ncRNA from which the RT-DNA is built has been found to be most frequently a guanine often preceded by TTA, which we also see in our set of functional retrons (Supplemental Fig 2). We find that the first reverse transcribed base is also most frequently a guanine and that the termination point of RT-DNA typically precedes an AT rich region of the ncRNA (Fig 2d). The length of the ncRNA varies across retrons, and is correlated with the length of the RT-DNA (Fig 2e). We can use the ratio of DBR1+ sequencing reads to DBR1- sequencing reads to infer which retrons remain branched in vivo. Retrons with similar numbers of reads regardless of the presence or absence of DBR1 in the sequencing prep are likely unbranched already in vivo, whereas retrons with many more reads in the DBR1+ condition likely remain branched in vivo (Fig 2f). Finally, this RT-DNA sequencing data can be harnessed to estimate the relative fidelity of the retrons by quantifying the errors per base in the RT-DNA with respect to the retron’s ncRNA reference. Critically, these are not absolute error rates of the RT as this includes errors introduced by sequencing prep or Illumina sequencing. We find that most retrons are of a similar fidelity, with just a few retrons making ~5x greater errors (Fig 2g).
Figure 2.

Characteristics of RT-DNA production across retrons. a. Examples of RT-DNA sequencing coverage plotted on folded ncRNA for a subset of sequenced retrons (additional data in Supplemental Data 1). b. Strictness of initiation. For all retrons with at least 100 sequencing reads, number of nucleotides at the RT initiation point between 15% an 85% of maximum coverage. c. Strictness of termination. For all retrons with at least 100 sequencing reads, number of nucleotides at the RT termination point from 85% to 15% of maximum coverage. d. Probability logos showing overrepresented (positive) and underrepresented (negative) nucleotides at the initiation (left) and termination (right) points. Red line indicates P<0.05. e. RT-DNA length versus ncRNA length (Pearson Correlation, r2=0.6586, P<0.0001). f. Ratio of the sequencing coverage in the +DBR1 condition versus the –DBR1 condition for each retron. Retrons with a high ratio are more likely to exist in a 2’−5’ branched form with their msr, which requires DBR1 for efficient sequencing, whereas retrons with ratios around or under 1 are equally sequenced in the absence or presence of DBR1 indication that they were not 2’−5’ branched. g. Errors per base versus the ncRNA reference for retrons with at least 1,000 sequenced bases, calculated using Levenshtein Distance. Errors include fidelity of the RT, but also errors introduced by the sequencing prep and sequencing. Additional statistical details in Supplemental Table 2.
Surveying retrons for bacterial and phage recombineering
The ability of retrons to continuously produce a ssDNA containing a precise target mutation has been recently exploited for recombineering purposes in both bacteria and phages13,16,18,21–23. Retron-based recombineering relies on the reverse transcription of a modified RT-DNA that contains an editing donor (Fig 3a). The host single-stranded binding protein (SSB) will bind the resulting RT-Donor to promote its interaction with a single-stranded annealing protein (SSAP, e.g. CspRecT), which leads to installation of the edit in the lagging strand during chromosome replication24. However, only retron-Eco1 has been extensively characterized and optimized for use in prokaryotic genome engineering.
Figure 3.

Bacterial editing across retrons. a. Schematic of modifications to the retron architecture for encoding a recombineering donor. b. Precise editing rate across retrons for bacterial genome recombineering. Points show mean ±SEM. c. RT-DNA production versus editing rate (r2=0.7477, P<0.0001). d. Predicted secondary structure of Kva1 ncRNA with sequencing data in orange. e. Precise editing rate across retrons for phage genome recombineering. Points show mean ±SEM. f. Bacterial versus phage editing rate (r2=0.2184, P=0.0505). Additional statistical details in Supplemental Table 2.
To survey the ability of retrons to support recombineering, a batch of 29 candidates were selected from our retron library, prioritizing higher RT-DNA producers across the phylogenic tree, and including retron-Eco1 as a reference. The ncRNAs of these retrons were modified to contain a 90 bp template donor to make a precise single nucleotide mutation in rpoB gene. The template location was placed within the RT-DNA region as determined by our sequencing data and in a centrally located hairpin stem. After analyzing recombineering efficiencies, we found that 8 retrons yield higher editing rates than retron-Eco1 (Fig 3b). We find that RT-DNA production from the wildtype ncRNA and editing using a modified DNA are strongly correlated, consistent with previous work showing that the abundance of recombineering donor is a limiting reagent in editing (Fig 3c). However, the most effective retron for bacterial editing deviates from this correlation. Strikingly, a retron from Klebsiella variicola, referred from now on as retron-Kva1 (Mestre-208, clade 1), shows a 10-fold increase editing efficiency compared to Eco1 despite producing only a moderate amount of RT-DNA. Thus, retron-Kva1 represents a particular case that will require further characterization to understand whether another specific feature, such as its predicted non-canonical ncRNA structure, could explain its high recombineering abilities (Fig 3d).
18 out of the 29 retron candidates were also assessed for their ability to support phage genome recombineering. These retrons were drawn from a subset of high RT-DNA producers across clades, modified to edit phage lambda xis gene (stop codon TGA>TAA) using a 70 bp RT-DNA donor. In this case, 4 retrons, including retron-Kva1, show higher recombineering rates than wild-type retron-Eco1 (Fig 3f). We found no strong correlation between RT-DNA production and lambda recombineering rates for this set of retrons (Supplemental Fig 3); the relative editing of bacterial and phage genomes by individual retrons is only weakly correlated (Fig 3f). This could indicate that the biological mechanism of bacterial and phage recombineering differ and that phage biology, RT-DNA structure or additional factors could impact editing rates. However, care should be taken with this interpretation, given that the set of retrons tested for phage editing was selected based on RT-DNA production and genome editing in bacteria.
Human precise editing by a diverse set of retrons
To extend the retron census beyond function in bacteria, we tested a diverse set of retrons for their ability to precisely edit the genomes of cultured human cells. Analogous to bacterial and phage editing, a modified RT-DNA encodes an editing donor. In eukaryotic editing, however, this donor is used to precisely repair a genomic site after a targeted double strand break by a Cas9 nuclease. For clarity, we term this combination of a retron with Cas9 for precise editing as an editron.
We synthesized a set of 136 never-before-tested editrons for human editing and four editrons that have been previously tested as internal references. For human editing of the EMX1 locus, a human codon-optimized RT was driven by a constitutive CAG promoter, while a modified ncRNA fused to a CRISPR sgRNA was driven by a Pol III promoter (Fig 4a). We modified each ncRNA to encode an editing donor in the RT-DNA, which is designed to insert a retron-specific 10 base barcode. The donor also contains a recoding of the PAM nucleotides to protect the ncRNA/sgRNA/RT plasmid and prevent further cutting of the genome after precise editing. Based on previous findings that the length of the retron ncRNA’s a1/a2 region can affect editing rates18, we increased this region of the ncRNA up to 16 bases for any editrons where the endogenous a1/a2 was shorter than 16 bases. We used an H1 promoter to express the ncRNA/sgRNA for most editrons, but used a U6 for a small subset where the H1 promoter was incompatible with synthesis of the retron components. In a parallel experiment, we found no difference between H1 and U6 promoters for retron-Eco1 (Supplementary Fig 4a). Given the potential for an innate immune reaction to a 2’5’ branched RNA/DNA hybrid, we also tested for activation of the cGAS-STING response with expression of two wildtype retrons (retron-Eco1 and Mestre-1036) and found no detectable induction of the cGAS-STING response by either retrons in cultured human cells (Supplementary Fig 4b).
Figure 4.

Human precise editing by a diverse set of retrons. a. Schematic of synthesized retron architecture for editing human genomes. b. Number of retrons synthesized by clade (percentage of clade synthesized indicated in gray above each bar). Bar o indicates an orphan, clade-less retron. c. Schematic of experimental workflow to test retrons in batches of 12. d. Proportion of retrons that enable precise editing at any efficiency. e. Demultiplexed editing percentage of functional editors, showing a range of efficiencies both above and below the prior reference retron-Eco1. f. Demultiplexed editing by retron. Points are mean ±SEM. g. Demultiplexed editing by clade (one-way ANOVA, effect of clade, P=0.2197). Bars are mean ±SEM. Bar o indicates an orphan, clade-less retron. h. Demultiplexed editing for retron Mestre-814 with the editing donor placed in different stem positions of the msd (one-way ANOVA, effect of donor location, P=0.3774). i. Demultiplexed editing for retrons tested with long and short versions of the msd stem (two-way ANOVA, effect of stem length, P=0.0017). Additional statistical details in Supplemental Table 2.
As in the bacterial work, we chose the human editrons to sample the phylogenic diversity of retrons, representing each RT clade. Of the 99 total retrons quantified for RT-DNA production and 140 total retrons quantified for human editing, 72 were tested in both conditions. We quantified editing in multiplexed pools of twelve editrons at a time to facilitate testing, using the unique 10 base pair insertion to demultiplex the edits and extrapolate the relative editing rate for each individual member of a pool. Specifically, we quantified the total precise editing rate for a given pool, then assigned a subset of that editing to each member of the pool based on the relative number of barcodes for that retron in the cell genomes and the relative number of plasmids for that retron in the particular pool (Fig 4c). Each editron appeared in at least two separate pools to mitigate effects of specific cross-reactivity within a pool.
The editron plasmids were transfected into HEK293T cells containing Cas9 pre-integrated into their genomes. Cells were collected three days after transfection. We generated two amplicons per sample for targeted Illumina sequencing: one which amplified the genome (to quantify editing) and one which amplified the plasmid (to account for differences in the relative abundance of retrons plasmids in the pool). We tested retrons in cells with either constitutively expressing or doxycycline inducible versions of the Cas9 and found similar results in both (Supplementary Fig 4c), and therefore present a merged dataset.
We found that the majority of plasmids (101/140) yielded edited genomes containing their precise barcode at some rate (Fig 4d). Of these editrons, we find that 58 of the 101 support higher rates of precise editing than the previous standard retron, retron-Eco1 (Fig 4e). As in the case of RT-DNA production, we find that the functional retrons are not randomly distributed across the RT phylogeny, but rather occur in hotspots (Fig 4f). In fact, the best editrons were drawn from the same clades (1, 2, 6, 9, 10) and subtypes as the highest RT-DNA producing retrons (Fig 4g, Supplementary Fig 4d). For one retron where we empirically determined the RT-DNA, we tested the placement of the donor at various positions either inside or outside the reverse transcribed region and found, unsurprisingly, that a donor in the core of the RT-DNA had a tendency to support higher levels of editing, consistent with the use of the RT-DNA as the editing donor and not the plasmid itself (Fig 4h). Additionally, we tested the effect of donor placement within the RT-DNA stem for three retrons and found that shortening the stem length that flanks the donor yielded higher rates of editing (Fig 4i).
We next individually tested a set of ten retrons that showed higher rates of editing than retron-Eco1 in the multiplexed experiments. Interestingly, during initial individual testing, we found that only a few of the retrons outperformed retron-Eco1, but we also found high variability in the rates of indel formation across the retrons tested that were roughly related to editing rate (Supplementary Fig 5a–d). Given that all retrons in the multiplexed experiment used the same gRNA, we wondered if potential ‘sharing’ of gRNAs between retron constructs might have obscured a negative effect on Cas9/gRNA activity from fusing the retron ncRNA to the gRNA in some retrons, leading to variability in editing and indel rates when tested individually. To test this, we split the ncRNA and gRNA for those ten retrons, driving the ncRNA with an U6 promoter and the gRNA with an H1 promoter on the same plasmid. We found that this arrangement increased editing rates for many of the retrons and led to equal rates of indel formation across all retrons (Supplementary Fig 5e–h). Thus, while a fused ncRNA/gRNA is sometimes functional for editrons, a split ncRNA/gRNA architecture is more reliable to prevent interference with the gRNA. None of the retrons tested edited better with a fused architecture.
Using the split ncRNA/gRNA, we found that six of the ten new retrons significantly outperformed retron-Eco1 (Fig 5a,b). We also checked fidelity of the retron-derived insertions. We found that edited genomes carried around ~0.003 excess errors per base compared to wild-type genomes from the same sample (Fig 5c). The slight differences in error rate across retrons were not significant. Looking deeper into the error type, we found that base substitutions accounted for essentially all the errors with no increase in deletions or insertions in edited reads versus wild-type reads (Supplementary Fig 5i–j). This pattern is consistent with rare polymerase errors during reverse transcription of the donor.
Figure 5.

Additional editron testing. a. Editron architecture using a split ncRNA/gRNA driven by U6 and H1 promoters, respectively. b. Precise editing for a set of individually tested retrons, each inserting a ten base barcode at the EMX1 locus. Three biological replicates are shown as open circles for each retron. Complete constructs are shown in purple, while constructs lacking a reverse transciptase are shown in black, retron-Eco1 (1550) marked in blue (one-way ANOVA effect of retron +RT P<0.0001; Dunnett’s corrected follow-up testing versus 1550: 412 P=0.0021; 482 P=0.0005; 1429, 6, 1467, and 1531 P≤0.0001). c. Errors per base on edited reads beyond the errors on wild-type reads. Open circles are biological replicates, each compared to wild-type reads from the same sample, closed circles are the mean (one-way ANOVA effect of retron P=0.2647). d. Precise editing for three retrons across three sites with a three base edit at each site, bars are mean ±SEM of three biological replicates (two-way ANOVA effect of site P=0.0029, retron P=0.3, and interaction P=0.4931). e. Retron Mestre-1531 precise editing (left axis, purple bars) and indel percentage (right axis, grey bars) at EMX1 with modified ncRNA structure and the addition of pharmacology. Bars are mean ±SEM of three biological replicates (one-way ANOVA effect of condition for precise editing P=0.0002; Dunnett’s corrected follow-up testing versus unmodified/no drug: modified structure P=0.7584, AZD7648/PolQi2 P=0.0003). f. Retron Mestre-1531 precise editing (left axis, purple bars) and indel percentage (right axis, grey bars) at HEK4 with modified ncRNA structure and the addition of pharmacology. Bars are mean ±SEM of three biological replicates (one-way ANOVA effect of condition for precise editing P<0.0001; Dunnett’s corrected follow-up testing versus unmodified/no drug: modified structure P=0.0165, modified structure and AZD7648/PolQi2 P<0.0001). Additional statistical details in Supplemental Table 2.
We next selected the top three performing retrons and tested them across three genomic sites. We found a small difference in editing rates across sites, as expected, but no effect of the retron or interaction between the retron and site, suggesting to us that this difference is primarily dictated by Cas9/gRNA activity differences across sites rather than an effect of retron type (Fig 5d). Finally, we tested structural modifications and the addition of edit-enhancing pharmacology on our top performing retron across two sites. For the structural modification, we elongated the a1/a2 region of Mestre-1531 to 23bp and shortened the msd donor stem to 9bp. This structural change had no significant effect on editing at the EMX1 site, but small positive effect at the HEK4 site (Fig 5e,f). We next added pharmacology that has previously been shown to enhance precise editing25: the DNA-dependent protein kinase (DNA-PK) inhibitor AZ7648, and the DNA polymerase theta (PolƟ) inhibitor PolQi2. This drug cocktail improved precise editing at both sites, reaching ~30% at EMX1 and 40% at HEK4, while simultaneously reducing the rate of indel formation (Fig 5e,f).
DISCUSSION
Retrons are notable for their innate immune function against phages and their applications in biotechnology, particularly for genome engineering. Despite their diversity, prior studies have only explored a limited range of retrons. Here, we conduct a comprehensive survey of these bacterial systems, focusing on RT-DNA production, bacterial and phage recombineering, and precise human genome editing. This work complements and advances earlier work on retron systems, including an expansion of validated retron RTs, ncRNAs and RT-DNAs. Some of the validated ncRNAs are derived from retron subtypes with no previously predicted ncRNAs (e.g. clade 8, type XI).
The data regarding RT-DNA position within the ncRNA, which can currently only be determined empirically, are particularly valuable. As the number of empirically determined RT-DNA sequences increases, it may become possible to predict future sequences without experimentation. We observed significant variability in RT-DNA production among different retrons, with retrons from certain clades consistently producing more RT-DNA. Another interesting finding is the lack of RT-DNA production in clades 3 and 4. Perhaps the ncRNA from these retrons is noncanonical or positioned far from the retron operon. Alternatively, these retrons may require additional host factors for RT-DNA production.
Our study also sheds light on reverse transcription initiation and termination within retrons. We confirm the presence of a preceding TTA consensus in functional retrons and identify a prevalent AT-rich region at the site of termination area. It is also important to note that some of retrons were visualized as double bands on the gel, and different branched/debranched conformations in sequencing. It has been described that retrons undergo varying maturation processes, which may explain these results2,26–28.
This research not only provides biological insights but also expands the genome editing toolkit for both prokaryotic and eukaryotic cells. As in previous work18, we found a strong correlation between RT-DNA production and bacterial editing rates. When editing lambda phage genomes, this correlation was lost. Speculatively, this could be due to anti-retron systems in the lambda genome. Recent work has suggested the presence of an anti-retron protein in phage T5 that reduces the amount of RT-DNA inside the cell29.
We also show that retron-Kva1 significantly outperforms the previous standard retron-Eco1 for both bacterial and phage recombineering, becoming an ideal candidate for further characterization and optimization. Although retron-Kva1 was functional for editing in human cells, it was not among the very best retrons. It also substantially deviated from the correlation between RT-DNA production and editing rate in bacteria, yielding our highest editing rates despite a moderate amount of RT-DNA production. There may be interactions between Kva1 and the bacterial host biology or recombineering machinery that explain these discrepancies, but which we do not yet understand.
We also characterized the ability of phylogenetically diverse retrons to edit mammalian cells, an application which has been limited by a small number of validated retrons. Our work identifies 58 retrons that edit mammalian cells at greater efficiencies than the previous biotechnological standard of retron-Eco1. Interestingly, the top retrons for bacterial RT-DNA production, bacterial recombineering, and human editors are all derived from an overlapping set of RT clades. However, bacterial RT-DNA production is not correlated with human editing at the level of individual retrons (Supplementary Fig 6). This difference could reflect RT expression differences in human versus bacterial cells, as well as reverse transcription efficiency. One general lesson we take from this census is that an efficient pipeline for mammalian technology optimization can employ an initial survey of part diversity in bacterial cells for a property of interest to identify particularly promising clades/subtypes of that part, but then the ideal individual component still needs to be empirically determined in the mammalian cell. We also demonstrate a new method to test editing of retrons in multiplexed pools which harnesses the unique RT-DNA produced by each retron. This enables a more efficient and trackable screening platform for genome engineering.
Overall, our study is the most extensive characterization of retrons to date, validating a larger set of retrons and identifying retrons with better RT-DNA production, bacterial and phage recombineering, and human editing.
METHODS
Biological replicates were taken from distinct samples, not the same sample measured repeatedly.
Plasmids
For bacterial expression, retron ncRNA-RT pairs were synthesized into high-copy pET21 backbones with T7/lac inducible promoters (Twist). Codons were optimized only if necessary for synthesis. Constructs for recombineering were cloned from these plasmids to add an rpoB editing donor (S512P) or lambda xis editing donor (stop codon TGA to TAA). pORTMAGE-Ec1 was generated previously30.
All human vectors are derivatives of pSCL.273, itself a derivative of pCAGGS31. pCAGGS was modified by replacing the MCS and rb_glob_polyA sequence with an IDT gblock containing inverted BbsI restriction sites and a SpCas9 tracrRNA, using Gibson Assembly. The resulting plasmid, pSCL.273, contains an SV40 ori for plasmid maintenance in HEK293T cells. The strong CAG promoter is followed by the BbsI sites and SpCas9 tracrRNA. BbsI-mediated digestion of pSCL.273 yields a backbone for single or library cloning of plasmids with inserts that contain (retron RT - pol III promoter - modified retron ncRNA), by Gibson Assembly or Golden Gate cloning.
Retrons for human editing were synthesized into pSCL.273 (Twist). The RT is driven by the CAG promoter. The ncRNA/gRNA cassette was placed under a H1/U6 promoter, with the choice of promoter not impacting editing rates (Supplementary Fig 4a). The donor location loop was determined by viewing RNAFold structures. Typically, the donor was placed at the first unpaired bases of the predicted RT-DNA stem. The reverse transcribed repair template was slightly asymmetric (49 bp of genome site homology upstream of the Cas9 cut site; 71 bp of genome site homology downstream of the cut site) and was complementary to the target strand; in practice, this means that after reverse transcription, the repair template RT-DNA is complementary to the non-target strand, as recommended in previous studies32. The repair template for human editing carried three distinct mutations to the EMX1 locus: the first inserts a 10-bp sequence at the Cas9 cut site, with a unique sequence generated for each retron plasmid. The second changes a G>A after the 10-bp insert for the gRNA. The third recodes the Cas9 PAM (NGG → NTT). The gRNA is 20 bp and targets EMX1. For individual validation, new plasmids were cloned where the ncRNA was driven by a U6 promoter and the gRNA was driven by a H1 promoter. For each editron in this set, a corresponding version was made without an RT. The same barcoded donor was used as in the pool. To test single-base recodes at different sites, the same donor and guide were used as previously described to create a three base precise edit18. The last set of plasmids cloned created modifications to Mestre-1531 ncRNA including elongating the a1/2 to 23bp and shortening the msd donor stem to 9bp.
Retron accession information as well as all plasmids and ncRNA sequences are listed in Supplemental Tables 1 and 6.
Bacterial Strains and Growth Conditions
The E. coli strains used in this study were DH5α (New England Biolabs) for cloning, bSLS.11418 for RT-DNA production, bMS.34615 for bacterial and phage retron recombineering assays. bSLS.114 was constructed from BL21-AI cells using lambda-red replacement to remove the retron-Eco1 locus. bMS.346 was generated from E. coli MG1655 by inactivating the exoI and recJ genes with early stop codons. Bacterial cultures were grown in LB (supplemented with 0.1 mM MnCl2 and 5 mM MgCl2 (MMB) for phage assays), shaking at 37 °C with appropriate inducers and antibiotics. Inducers and antibiotics were used at the following working concentrations: 1mM m-toluic acid (Sigma-Aldrich), 1 mM IPTG (GoldBio), 2 mg/ml L-arabinose (GoldBio), 35 μg/ml kanamycin (GoldBio), and 100 μg/ml carbenicillin (GoldBio).
RT-DNA expression and gel analysis.
RT-DNA expression and analysis was performed as previously described12. Briefly, retron plasmids were transformed into bSLS.114 for expression. A starter culture from a single clone was grown overnight in 3 ml LB plus antibiotic. After 16h, the culture was diluted (1:100) into 25 ml LB. This culture was allowed to reach OD ~0.5 (approximately 2 hours) and then induced with 1M IPTG and 200 ug/ml l-arabinose. After 5 hours, OD600 was measured and bacteria were harvested for RT-DNA analysis.
RT-DNA was recovered using a Qiagen Plasmid Plus Midi kit, eluted into a volume of 150 ul. Volume RT-DNA prep was adjusted based on bacterial OD, measured at the point of collection, to normalize the input prior to loading into Novex TBE-Urea gels (15% Invitrogen). The gels were run (45 minutes at 200 V) in pre-heated (>75°C) TBE running buffer. Gels were stained with SYBR Gold (ThermoFisher) and then imaged on a Gel Doc Imager (BioRad). To quantify the amount of RT-DNA production relative to retron-Eco1, a retron-Eco1 expressing strain was included every batch of culture grown, and the resulting prep of was always run on the same gel as the experimental retron for quantification. The density of the strongest band of each retron was quantified with ImageJ software.
Multiplexed RT-DNA Sequencing
Following Qiagen midipreps, mixes of up to 11 different retrons (including Eco1) were subjected to a DBR1 or sham-DBR1 (water) treatment in the following reaction: 79 ul of RT-DNA prep split evenly depending on the number of retrons in the mix, 10 ul of DBR1 (50 ng/uL), 1 ul of RNaseH (NEB), 10 ul rCutSmart Buffer (NEB). The mixes of retrons were organized according to their relative production to retron-Eco1, with the highest producing retrons grouped together and lowest producing retrons grouped together. This minimized the skew in sequencing among the retrons in a pool.
These reactions were cleaned up with ssDNA/RNA clean & concentrator kit (Zymo Research) and RT-DNA was prepped for sequencing by taking the resulting material and extending the 3′ end in parallel with two types of nucleotides: dCTP or dATP, using terminal deoxynucleotidyl transferase (TdT) (NEB). This reaction was carried out in 1× TdT buffer, with 60 units of TdT and 125 M dATP for 60 s, or 125 M dCTP for 4 minutes at room temperature with the aim of adding ~25 adenosines adenosines or ~6 cytosines before inactivating the TdT at 70°C for 5 min. Next, a primer a polynucleotide, anchored primer was used to create a complementary strand to the TdT extended products using 15 units of Klenow Fragment (3′→5′ exo-) (NEB) in 1× NEB2, 1 mM dNTP and 50 nM of primer containing an Illumina adapter sequence, nine thymines (for the pA extended version) or six guanines (for the pC extended version), and a non-thymine (V) or non-guanine (H) anchor. This product was cleaned up using Qiagen PCR cleanup kit and eluted in 10 ul water. Finally, Illumina adapters were ligated on at the 3′ end of the complementary strand using 1× TA Ligase Master Mix (NEB). All products were indexed and sequenced on an Illumina MiSeq instrument. Previous data from Palka et al. 202212 was included in the sequencing analysis for retron-Sen2 (Mestre-116), retron-Eco6 (Mestre-488), retron-Eco2 (Mestre-1036), and retron-Eco4 (Mestre-64).
Bacterial Recombineering and Analysis.
To edit bacterial genomes, the different retron cassettes encoded in a pET-21 (+) plasmid and the CspRecT and mutLE32K in the plasmid pORTMAGE-Ec1 were transformed into the bMS.346 strain by electroporation. Experiments were conducted in 500uL cultures in a deep 96-well plate. 3 individual colonies of every retron tested were grown in LB with kanamycin and carbenicillin for 16 h at 37°C. A 1:1000 dilution of these cultures was grown into LB with 1 mM IPTG, 0.2% arabinose and 1 mM m-toluic acid for 16 h with shaking at 37°C. A volume of 25 μl of culture was collected, mixed with 25 μl of water and incubated at 95°C for 10 min. A volume of 1 μl of this boiled culture was used as a template in 25-μl reactions with primers flanking the edit site, which additionally contained adapters for Illumina sequencing preparation. These amplicons were indexed and sequenced on an Illumina MiSeq instrument and processed with custom Python software to quantify the percentage of precisely edited genomes.
Lambda Propagation and Plaque Assays
Lambda was propagated from an ATCC stock (#97538) into a 2mL culture of E. coli bMS.346 strain at 37°C at OD600 0.25 in MMB medium until culture collapse. The culture was then centrifuged and the supernatant was filtered to remove bacterial remnants. Lysate titer was determined using the full plate plaque assay method as described by Kropinski et al. 200933.
Lambda Recombineering and analysis
The diverse retron cassettes with modified ncRNAs to contain a lambda donor to edit xis gene, was co-expressed with CspRecT and mutL E32K from the plasmid pORTMAGE-Ec1. Experiments were conducted in 500uL cultures in a deep 96-well plate. 3 individual cultures were grown for every retron tested for 16h at 37°C. A 1:100 dilution of each culture was induced for 2 h at 37°C. The OD600 of each culture was measured to approximate cell density and cultures were diluted to OD600 0.25. A volume of pre-titered lambda was added to the culture to reach a multiplicity of infection (MOI) of 0.1. The infected culture was grown overnight for 16 h, before being centrifuged for 10 min at 4000 rpm to remove the cells. For amplicon-based sequencing, a volume of 25 μl of the medium containing the phage was collected, mixed with 25 μl of water and incubated at 95°C for 10 min. A volume of 1 μl of this boiled culture was used as a template in 25-μl reactions with primers flanking the edit site, which additionally contained adapters for Illumina sequencing preparation. Sequencing and editing rates were analyzed as described previously for bacterial recombineering.
Human Editing Pool Design
Editrons were pooled in groups of 12. All editrons appear in at least two different pools (Supplementary Table 1). Individual plasmid cultures were started from glycerol stocks in 500 uL of LB Media +Carb for ~6 hours of growth at 37 degrees. Then, OD600 was measured using SpectraMax platereader to determine how much of each culture to add to a mixed culture for equal distribution. Mixed cultures were grown overnight at 37 degrees and Midiprepped according to manufacturer’s protocol (Qiagen).
Human Cell Culture
For pooled experiments, two Cas9-expressing HEK293T cell lines were used. The first expresses Cas9 from a piggybac integrated, TRE3G driven, doxycycline-inducible (1 μg/ml) cassette, which we have previously described18. The second expresses Cas9 constitutively from a CBh promoter in the AAVS1 Safe Harbor locus (GeneCopoeia #SL502). All HEK cells were cultured in DMEM +GlutaMax supplement (Thermofisher #10566016).
For pooled experiments, T25 cultures were transiently transfected with 12.76 ug of pooled plasmid per T25 using Lipofectamine 3000. Individual follow up experiments were conducted in 6-well and were transiently transfected with 7.32 ug of plasmid using Lipofectamine 3000. All cultures were passaged and cultured for an total of 72h. In the inducible Cas9 line, doxycycline was refreshed at passaging. Three days after transfection, cells were collected for sequencing analysis. To prepare samples for sequencing, cell pellets were collected, and gDNA was extracted using a QIAamp DNA mini kit according to the manufacturer’s instructions. DNA was eluted in 100–200 μl of ultra-pure, nuclease-free water. AZD7648 (1 μM, Selleck Chemicals S8843) and PolQi2 (3 μM, MedChem Express HY-150279) were added immediately before transfection and replenished during passaging.
Human Sample Preparation and Analysis
For pooled samples, 0.5 μl of the gDNA was used as template in 25-μl PCR reactions with primer pairs to amplify the locus of interest and a PCR reaction with primer pairs to amplify the plasmid region, both of which also contained adapters for Illumina sequencing preparation. Lastly, the amplicons were indexed and sequenced on an Illumina MiSeq instrument and processed with custom Python software to quantify the percentage of on-target precise genomic edits normalized to the representation of each plasmid in the pool. Individual experiments were analyzed with a separate custom script to quantify on target precise genome edits. Levenshetein fuzzysearch was adjusted from 1 to 0 based on barcode editing or SNP editing respectively. For fidelity analysis, reads were compared to either wildtype or edited sequences in the region of the edit (depending on the outcome assigned) using Python’s difflib to quantify the frequency of each error type. Excess errors were calculated as the errors per base for edited reads minus the errors per base for wildtype reads in the same sample.
STING Activation Assay
Wild-type versions of Mestre-1036 (retron-Eco2) and Mestre-1550 (retron-Eco1) were integrated semi-randomly with the PiggyBac transposase and Lipofectamine 3000 (Thermofisher) into the HEK STING reporter line (Invivogen) prior to the experiment. HEK STING reporter line was either treated with 2’3’-cGAMP (Invivogen) or had retron expression induced with doxycycline (1 μg/mL). 2’3’-cGAMP was transfected into the cells with Lipofectamine 2000 (Thermofisher) at concentrations ranging from 0 to 30 μg/mL. After 48 hours of 2’3’-cGAMP treatment or retron expression, 20 μL of supernatant incubated with 100 μL QuantiBlue solution (Invivogen) (1 μL QuantiBlue reagents: 1 μL QuantiBlue buffer: 98 μL of water) for 4 hours at 37 C, and then OD655 of the reaction was measured on a plate reader. Fold change is relative to a mock-transfection control with no 2’3’-cGAMP.
Supplementary Material
Supplementary Figure 3. Related to Figure 3. RT-DNA production versus with phage lambda recombineering rates (Pearson Correlation (r2=0.001926, P=0.8766).
Supplementary Figure 2. Related to Figure 2. Probability logo of ncRNA nucleotides adjacent to the branching guanosine for retrons that produce RT-DNA.
Supplementary Figure 1. Related to Figure 1. RT-DNA production relative to Eco1 by clade, separated by subtype. Bars show mean ±SEM.
Supplementary Figure 4. Related to Figure 4. a. Demultiplexed editing percentage by retron-Eco1 (Mestre-1550) comparing the use of an H1 and U6 promoter to drive ncRNA/sgRNA (unpaired T-test, P=0.6256). Bars are mean ±SEM. b. Fold-change in STING response after 48 hours of retron induction or 2’−3’ cGAMP (0.003 to 30 ug/mL) relative to a mock transfection control with 0 ug/mL 2’−3’ cGAMP. Open circles are biological replicates Each biological replicate is the mean of three technical replicates. c. Comparison of demultiplexed editing rates of editrons in the constitutive and inducible Cas9 cell lines (Pearson Correlation, r2=0.4196, P<0.0001). d. Average demultiplexed editing rates across retron subtypes. Bars are mean ±SEM.
Supplementary Figure 6. Related to Figure 4. Bacterial RT-DNA production compared with human demultiplexed editing.
Supplementary Figure 5. Related to Figure 5. a. Editron architecture with fused ncRNA/gRNA. b. Precise editing for eleven retrons, including retron-Eco1 in the fused architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. c. Indel percentage for eleven retrons, including retron-Eco1 in the fused architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. d. Indel versus precise editing percentage for the fused architecture, mean ±SEM for three biological replicates. e. Editron architecture with split ncRNA and gRNA. f. Precise editing for eleven retrons, including retron-Eco1 in the split architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. Note that this data is replotted in Figure 5b. g. Indel percentage for eleven retrons, including retron-Eco1 in the split architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. d. Indel versus precise editing percentage for the split architecture, mean ±SEM for three biological replicates. i. Substitutions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. j. Deletions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. In cases where there are fewer than three points, there were not enough errors to quantify the frequency. k. Insertions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. In cases where there are fewer than three points, there were not enough errors to quantify the frequency.
Acknowledgements
Work was supported by funding from the National Science Foundation (MCB 2137692), the National Institute of Biomedical Imaging and Bioengineering (R21EB031393), the National Institute of General Medical Sciences (1DP2GM140917), and research support from Retronix Bio. S.L.S. is a Chan Zuckerberg Biohub – San Francisco Investigator and acknowledges additional funding support from the L.K. Whittier Foundation and the Pew Biomedical Scholars Program. A.G.-D. was supported by the California Institute of Regenerative Medicine (CIRM) scholar program. S.C.L. was supported by a Berkeley Fellowship for Graduate Study. R.F.F. was supported by a UCSF Discovery Fellowship. K.D.C. was supported by a National Science Foundation Graduate Research Fellowship and a UCSF Discovery Fellowship. We would like to thank Alex Pico and the Gladstone Bioinformatics Core for assistance with data database management as well as Karen Zhang and David Wen for comments on the manuscript.
Footnotes
Code Availability
Custom code to process or analyze data from this study will be made available on GitHub prior to peer-reviewed publication.
Competing Interests
S.L.S. is a co-founder of Retronix Bio and Sprint Synthesis. A.G-D., S.C.L., and S.L.S. are named inventors on patent applications related to the technologies described in this work that are assigned to the Gladstone Institutes and the University of California, San Francisco.
Data Availability
All data supporting the findings of this study are available within the article and its supplementary information, or will be made available from the authors upon request. Sequencing data associated with this study is be available on NCBI SRA (PRJNA1047666).
References:
- 1.Yee T, Furuichi T, Inouye S & Inouye M Multicopy single-stranded DNA isolated from a gram-negative bacterium, Myxococcus xanthus. Cell 38, 203–209, doi: 10.1016/0092-8674(84)90541-5 (1984). [DOI] [PubMed] [Google Scholar]
- 2.Inouye S, Hsu MY, Eagle S & Inouye M Reverse transcriptase associated with the biosynthesis of the branched RNA-linked msDNA in Myxococcus xanthus. Cell 56, 709–717, doi: 10.1016/0092-8674(89)90593-x (1989). [DOI] [PubMed] [Google Scholar]
- 3.Hsu MY, Eagle SG, Inouye M & Inouye S Cell-free synthesis of the branched RNA-linked msDNA from retron-Ec67 of Escherichia coli. The Journal of biological chemistry 267, 13823–13829 (1992). [PubMed] [Google Scholar]
- 4.Shimamoto T, Inouye M & Inouye S The formation of the 2’,5’-phosphodiester linkage in the cDNA priming reaction by bacterial reverse transcriptase in a cell-free system. The Journal of biological chemistry 270, 581–588, doi: 10.1074/jbc.270.2.581 (1995). [DOI] [PubMed] [Google Scholar]
- 5.Shimamoto T, Kawanishi H, Tsuchiya T, Inouye S & Inouye M In vitro synthesis of multicopy single-stranded DNA, using separate primer and template RNAs, by Escherichia coli reverse transcriptase. Journal of bacteriology 180, 2999–3002, doi: 10.1128/jb.180.11.2999-3002.1998 (1998). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Simon AJ, Ellington AD & Finkelstein IJ Retrons and their applications in genome engineering. Nucleic Acids Res 47, 11007–11019, doi: 10.1093/nar/gkz865 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Gao L et al. Diverse enzymatic activities mediate antiviral immunity in prokaryotes. Science 369, 1077–1084, doi: 10.1126/science.aba0372 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Millman A et al. Bacterial Retrons Function In Anti-Phage Defense. Cell 183, 1551–1561.e1512, doi: 10.1016/j.cell.2020.09.065 (2020). [DOI] [PubMed] [Google Scholar]
- 9.Bobonis J et al. Bacterial retrons encode phage-defending tripartite toxin-antitoxin systems. Nature 609, 144–150, doi: 10.1038/s41586-022-05091-4 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Mestre MR, González-Delgado A, Gutiérrez-Rus LI, Martínez-Abarca F & Toro N Systematic prediction of genes functionally associated with bacterial retrons and classification of the encoded tripartite systems. Nucleic Acids Res 48, 12632–12647, doi: 10.1093/nar/gkaa1149 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wang Y et al. Cryo-EM structures of Escherichia coli Ec86 retron complexes reveal architecture and defence mechanism. Nature microbiology 7, 1480–1489, doi: 10.1038/s41564-022-01197-7 (2022). [DOI] [PubMed] [Google Scholar]
- 12.Palka C, Fishman CB, Bhattarai-Kline S, Myers SA & Shipman SL Retron reverse transcriptase termination and phage defense are dependent on host RNase H1. Nucleic Acids Res 50, 3490–3504, doi: 10.1093/nar/gkac177 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Farzadfard F & Lu TK Genomically encoded analog memory with precise in vivo DNA writing in living cell populations. Science 346, 1256272–1256272, doi: 10.1126/science.1256272 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Sharon E et al. Functional Genetic Variants Revealed by Massively Parallel Precise Genome Editing. Cell 175, 544–557.e516, doi: 10.1016/j.cell.2018.08.057 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Bhattarai-Kline S et al. Recording gene expression order in DNA by CRISPR addition of retron barcodes. Nature 608, 217–225, doi: 10.1038/s41586-022-04994-6 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Schubert MG et al. High-throughput functional variant screens via in vivo production of single-stranded DNA. PNAS 118, e2018181118, doi: 10.1073/pnas.2018181118 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lee G & Kim J Engineered retrons generate genome-independent protein-binding DNA for cellular control. bioRxiv, 2023.2009.2027.556556, doi: 10.1101/2023.09.27.556556 (2023). [DOI] [Google Scholar]
- 18.Lopez SC, Crawford KD, Lear SK, Bhattarai-Kline S & Shipman SL Precise genome editing across kingdoms of life using retron-derived DNA. Nat Chem Biol 18, 199–206, doi: 10.1038/s41589-021-00927-y (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Kong X et al. Precise genome editing without exogenous donor DNA via retron editing system in human cells. Protein & cell 12, 899–902, doi: 10.1007/s13238-021-00862-7 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Zhao B, Chen SA, Lee J & Fraser HB Bacterial Retrons Enable Precise Gene Editing in Human Cells. The CRISPR journal 5, 31–39, doi: 10.1089/crispr.2021.0065 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Liu W et al. Retron-mediated multiplex genome editing and continuous evolution in Escherichia coli. Nucleic Acids Res 51, 8293–8307, doi: 10.1093/nar/gkad607 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Fishman CB et al. Continuous Multiplexed Phage Genome Editing Using Recombitrons. bioRxiv, 2023.2003.2024.534024, doi: 10.1101/2023.03.24.534024 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.González-Delgado A, Lopez SC, Rojas-Montero M, Fishman CB & Shipman SL Simultaneous multi-site editing of individual genomes using retron arrays. bioRxiv, doi: 10.1101/2023.07.17.549397 (2023). application related to the technologies described in this work. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Mosberg JA, Lajoie MJ & Church GM Lambda red recombineering in Escherichia coli occurs through a fully single-stranded intermediate. Genetics 186, 791–799, doi: 10.1534/genetics.110.120782 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Wimberger S et al. Simultaneous inhibition of DNA-PK and Polϴ improves integration efficiency and precision of genome editing. Nature communications 14, 4761, doi: 10.1038/s41467-023-40344-4 (2023). S.DC., P.I., M.B., T.M., S.R., O.E., E.B.C., J.V.F., S.Š., P.A., A.T.G. and M.M. are presently or were previously employed by AstraZeneca and may be AstraZeneca shareholders. M.K.S., M.R.S. and TM are presently employed by Promega Corporation. S.W., N.A., S.Š. and M.M. are listed as co-inventors in an AstraZeneca patent (WO2023052508A2) related to this work. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Lampson BC, Inouye M & Inouye S Retrons, msDNA, and the bacterial genome. Cytogenet Genome Res 110, 491–499, doi: 10.1159/000084982 (2005). [DOI] [PubMed] [Google Scholar]
- 27.Kim K, Jeong D & Lim D A mutational study of the site-specific cleavage of EC83, a multicopy single-stranded DNA (msDNA): nucleotides at the msDNA stem are important for its cleavage. Journal of bacteriology 179, 6518–6521, doi: 10.1128/jb.179.20.6518-6521.1997 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Lease RA & Yee T Early events in the synthesis of the multicopy single-stranded DNA-RNA branched copolymer of Myxococcus xanthus. The Journal of biological chemistry 266, 14497–14503 (1991). [PubMed] [Google Scholar]
- 29.Azam AH et al. Viruses encode tRNA and anti-retron to evade bacterial immunity. bioRxiv, 2023.2003.2015.532788, doi: 10.1101/2023.03.15.532788 (2023). [DOI] [Google Scholar]
- 30.Wannier TM et al. Improved bacterial recombineering by parallelized protein discovery. Proceedings of the National Academy of Sciences of the United States of America 117, 13689–13698, doi: 10.1073/pnas.2001588117 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Niwa H, Yamamura K & Miyazaki J Efficient selection for high-expression transfectants with a novel eukaryotic vector. Gene 108, 193–199, doi: 10.1016/0378-1119(91)90434-d (1991). [DOI] [PubMed] [Google Scholar]
- 32.Richardson CD, Ray GJ, DeWitt MA, Curie GL & Corn JE Enhancing homology-directed genome editing by catalytically active and inactive CRISPR-Cas9 using asymmetric donor DNA. Nature biotechnology 34, 339–344, doi: 10.1038/nbt.3481 (2016). [DOI] [PubMed] [Google Scholar]
- 33.Kropinski AM, Mazzocco A, Waddell TE, Lingohr E & Johnson RP Enumeration of bacteriophages by double agar overlay plaque assay. Methods in molecular biology (Clifton, N.J.) 501, 69–76, doi: 10.1007/978-1-60327-164-6_7 (2009). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Figure 3. Related to Figure 3. RT-DNA production versus with phage lambda recombineering rates (Pearson Correlation (r2=0.001926, P=0.8766).
Supplementary Figure 2. Related to Figure 2. Probability logo of ncRNA nucleotides adjacent to the branching guanosine for retrons that produce RT-DNA.
Supplementary Figure 1. Related to Figure 1. RT-DNA production relative to Eco1 by clade, separated by subtype. Bars show mean ±SEM.
Supplementary Figure 4. Related to Figure 4. a. Demultiplexed editing percentage by retron-Eco1 (Mestre-1550) comparing the use of an H1 and U6 promoter to drive ncRNA/sgRNA (unpaired T-test, P=0.6256). Bars are mean ±SEM. b. Fold-change in STING response after 48 hours of retron induction or 2’−3’ cGAMP (0.003 to 30 ug/mL) relative to a mock transfection control with 0 ug/mL 2’−3’ cGAMP. Open circles are biological replicates Each biological replicate is the mean of three technical replicates. c. Comparison of demultiplexed editing rates of editrons in the constitutive and inducible Cas9 cell lines (Pearson Correlation, r2=0.4196, P<0.0001). d. Average demultiplexed editing rates across retron subtypes. Bars are mean ±SEM.
Supplementary Figure 6. Related to Figure 4. Bacterial RT-DNA production compared with human demultiplexed editing.
Supplementary Figure 5. Related to Figure 5. a. Editron architecture with fused ncRNA/gRNA. b. Precise editing for eleven retrons, including retron-Eco1 in the fused architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. c. Indel percentage for eleven retrons, including retron-Eco1 in the fused architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. d. Indel versus precise editing percentage for the fused architecture, mean ±SEM for three biological replicates. e. Editron architecture with split ncRNA and gRNA. f. Precise editing for eleven retrons, including retron-Eco1 in the split architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. Note that this data is replotted in Figure 5b. g. Indel percentage for eleven retrons, including retron-Eco1 in the split architecture (1550, in blue). Bars are mean ±SEM for three biological replicates. d. Indel versus precise editing percentage for the split architecture, mean ±SEM for three biological replicates. i. Substitutions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. j. Deletions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. In cases where there are fewer than three points, there were not enough errors to quantify the frequency. k. Insertions per base on edited reads beyond the errors on wild-type reads. Open circles are three biological replicates, each compared to wild-type reads from the same sample. In cases where there are fewer than three points, there were not enough errors to quantify the frequency.
Data Availability Statement
All data supporting the findings of this study are available within the article and its supplementary information, or will be made available from the authors upon request. Sequencing data associated with this study is be available on NCBI SRA (PRJNA1047666).
