Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2025 Dec 23;122(52):e2516306122. doi: 10.1073/pnas.2516306122

Anellovirus protein encoded by ORF2/3 functions as the viral replication initiation protein

Nicole Boisvert a,1, Stephanie Thurmond a,1, Carmen Elenberger a, Patricio Jeraldo a, Cato Prince a, Nolan Sutherland a, José Melo a, Ken Tsheowang a, Cameron Dodier a, Maciej Nogalski a, Dinesh Verma a, Geoffrey Parsons a,2, Joseph Cabral a,2
PMCID: PMC12772153  PMID: 41433061

Significance

Anelloviruses (ANVs) are the most prevalent eukaryotic viruses in humans, yet their molecular biology remains poorly understood due to the lack of experimental systems. Here, we identify a previously uncharacterized protein we have termed Rip (Replication initiation protein) as necessary and sufficient to initiate viral genome replication. We demonstrate that Rip initiates replication from noncoding regions of both alpha- and betatorquevirus genomes, and interacts with host recombination machinery, suggesting a recombination-dependent replication mechanism. This work provides experimental evidence of the molecular mechanism of ANV replication and the function of the ORF2/3-encoded protein. These findings represent a major advance in the understanding of this near-ubiquitous human virus family and enable future studies into their biology, evolution, and therapeutic potential.

Keywords: host-virus interaction, anellovirus, virology, human virome, viral replication

Abstract

Anelloviridae is a family of single-stranded DNA viruses that are thought to be nonpathogenic and commensal. Despite their ubiquitous presence in human populations, little is known about the anellovirus mechanism of replication in host cells. We identified the protein coded by ORF2/3 as necessary and sufficient to initiate replication from the minimal origin of replication for viruses of both the Beta- and Alphatorquevirus genera. Supporting this observation, we identified components of the polymerase alpha and BTR (Bloom’s syndrome helicase (BLM), topoisomerase IIIα, RMI1, and RMI2) complexes as interacting with the viral replication initiation protein (Rip) during DNA replication, suggesting a recombination-dependent mechanism of replication that uses host cell machinery to mediate dissolution of replication intermediates. Furthermore, we mapped a 92-bp minimal origin of replication sequence for the Betatorquevirus genus composed of an adenine and thymine (AT)-rich stretch and a portion of the guanine and cytosine (GC)-rich region. Altogether, this study provides insight into the mechanism by which anelloviruses manipulate host cell machinery to facilitate viral genome replication and represents a significant step forward in understanding the complex processes underlying anellovirus replication and persistent infection of these important commensal viruses.


Viruses of the family Anelloviridae are the dominant eukaryotic virus in the healthy human virome and are most abundant in the blood and bone marrow (15). Anelloviruses (ANVs) are nonenveloped, negative-sense, circular, single-stranded DNA (ssDNA) viruses that are nearly ubiquitous in human populations (2, 4). Anelloviruses establish persistent infections from childhood, remaining detectable throughout life in greater than 90% of the global adult population (69). T cells are suspected to be the most likely reservoir of persistent ANV in humans; (10) however, one study reported that granulocytes have the highest density of ANV genomes (11). The only known ANV pathogen is the chicken anemia virus (CAV), a gyrovirus which infects young chickens and induces anemia. CAV infects T cell precursors in the thymic cortex, CD8+ splenic lymphocytes, and hemocytoblasts in the bone marrow (12). Remarkably, with the exception of CAV, all other known ANVs are thought to be nonpathogenic and have evolved to coexist with their diverse vertebrate host species as commensal viruses (6, 13, 14). ANVs are not known to infect organisms outside of vertebrates.

Identification of Torque teno virus (TTV), the first ANV described, was reported in 1997 (15). Since then, ANVs have been detected in a wide range of vertebrate species, including humans and other primates. Human ANVs are classified largely into three genera based on genome size: Alpha-, Beta-, and Gammatorqueviruses, corresponding to TTV (3.7 to 3.9 kb), Torque teno mini virus (TTMV) (2.8 to 2.9 kb), and Torque teno midi virus (TTMDV) (3.2 kb), respectively. More recently, another genus of human ANVs, Hetorquevirus, was recognized (16, 17). Anelloviruses are the most prevalent eukaryotic virus found in humans and are highly genetically diverse, with thousands of unique capsid sequences identified across various tissue types (2, 6, 13, 17). Additionally, it has been suggested that anelloviruses may benefit human health by shaping immunity during early development (1, 1820). A recent study reported that TTV infections induce an exhausted TTV-specific CD8+ T cell response which may mediate an imprinting of the immune system toward NKG2A+ T cells (19). It has been proposed that an immune system biased toward NKG2A+ CD8+ T cells is associated with protection against disease severity, mortality, and autoimmune/postacute chronic disease (19, 21). Importantly, elevated ANV levels in the blood are correlated with immunosuppression; thus, ANV levels in the blood have been proposed to serve as immune markers to inform clinical outcomes (18, 22).

Despite their near-universal presence in humans, there is limited understanding of the molecular virology of these ubiquitous yet enigmatic viruses. Viruses of the Cressdnaviricota phylum are circular Rep-encoding single-stranded DNA (CRESS DNA) viruses that replicate via a rolling-circle mechanism initiated by virus-coded Rep proteins of the HUH endonuclease superfamily (2325). Owing to the absence of a recognizable Rep protein coding sequence, Anelloviridae is the only family of eukaryotic circular ssDNA viruses not included in this phylum (23). It has been proposed that the ANV origin of replication (ORI) is at a stem-loop structure with an octanucleotide motif that partially resembles the conserved nonanucleotide motif of CRESS DNA virus ORIs (14, 17, 26). However, there is no experimental evidence supporting this hypothesis. The protein coded by open reading frame 1 (ORF1) has been identified as the capsid protein (Cap) for ANVs, and it has also been suggested that the ORF1 protein may play a role in replication of the viral genome (17, 2729). Only recently has the ORF1 protein been experimentally demonstrated to form an ANV-like particle with icosahedral symmetry that is composed of 60 jelly roll domain–containing protein subunits (30), and another recent study identified two distinct nuclear localization signals on an ANV ORF1 that interact with importins to facilitate distinct subnuclear localization (31). Previous reports identified potential rolling circle replication (RCR) motifs associated with viral HUH Rep proteins as present in the ANV Cap (29, 32, 33). Recently, structural modeling of a TTMV Cap suggests the RCR motifs are spatially arranged in a manner that bears no resemblance to the HUH Rep proteins and that the highly conserved RNA-recognition-motif (RRM) fold does not form (28). In addition, these putative motifs are not well conserved in ANVs (25). The remaining ANV proteins lack homology to other known proteins, viral or otherwise; (14, 34) thus, their functions are largely unknown. The protein coded by ORF2 contains a conserved N-term W-x7-H-x3-C-x-C-x5-H motif which may suggest it has phosphatase activity; however, this activity has not been experimentally verified (28, 29, 33, 35). Studying the function of ANV proteins has been challenging due to a lack of an in vitro cell culture system; (36) although, recent advances in ANV virion production in MOLT-4 cells and the development of an ANV-based gene therapy vector system have paved the way for studying ANV protein functions (37, 38).

In this study, we report on the molecular mechanism of replication of human ANVs. Using a series of constructs to express viral gene products, we identified the protein coded by ORF2/3 as the human ANV replication initiation protein (Rip). Rip proteins from TTMV and TTV were necessary and sufficient to initiate replication from DNA constructs containing the entire noncoding regions (NCR) of either the Beta- or Alphatorquevirus genera, respectively. Further supporting this observation, we used immunoprecipitation-mass spectrometry (IP-MS) under multiple conditions to demonstrate that components of the polymerase alpha (POLα) complex and the BTR complex, which resolves DNA replication and recombination intermediates, (39) interact with ANV Rip during viral DNA replication, suggesting that ANVs may employ a recombination-dependent replication (RDR) mechanism for copying viral genomes. Additionally, we mapped the betatorquevirus ORI to a 92-bp sequence that immediately follows the protein coding sequences and contains an AT-rich stretch and a portion of the GC-rich region.

Results

nrVL4619 Transcript mRNA3 Accumulation Is Elevated at Early Times During Wildtype Virion Production.

The genomes of TTMVs are circular ssDNA that contain two distinct regions: the NCR and the protein coding sequences (Fig. 1A). The NCR contains the viral promoter and all of the cis elements required for replication and packaging of viral genomes (38). The TTMV NCR comprises a GC-rich region with 2 to 5 highly conserved hairpin structures followed by conserved domain 1 (CD1). Downstream of CD1 is a TATA box and a hyperconserved domain (HCD) which is conserved across the human ANV genera and serves as the transcriptional start site (Fig. 1B). The proposed replication loop resides in conserved domain 2 (CD2) (17, 26). An intron is also contained within CD2 (Fig. 1 A and B). When this intron is spliced out, the noncanonical Kozak sequence (GCCGAAGATG) for ORF2 is formed. The TTMV coding region contains three known ORFs: ORF2, ORF1, and ORF3, in order from 5’ to 3’ (Fig. 1A). Alternative splicing of the full transcript results in 8 detectable mRNA isoforms that code for up to 6 proteins (Fig. 1B). The ORF2 and ORF1 proteins are coded by mRNA1; the ORF2/2 and ORF1/1 proteins are coded by mRNA2; the ORF2/3 and ORF1/2 proteins are coded by mRNA3 (Fig. 1B).

Fig. 1.

Fig. 1.

Relative viral mRNA transcript abundance and viral protein expression over time. (A) (Top) Circular representation of TTMV genome. (Bottom) A linearized representation of the TTMV genome annotated to represent key features of the coding and NCR. The linearized sequence is represented as having the contiguous nrVL4619 NCR at the 5’ end of the genome and ending with the nrVL4619 polyadenylation signal at the 3’ end. Sequence features of the NCR include the GC-rich region, which contains 3 to 5 highly conserved hairpin structures, CD1, the core promoter region (containing the canonical TATA box and HCD), CD2, and the NCR intron. Open reading frames are indicated with labeled bars (ORF1, ORF2 which has a gap, ORF3). Not to scale. (B) The eight detectable mRNA species of TTMV, which are generated via alternative splicing. Open reading frames are indicated by an ORF number and black arrow. (C and D) MOLT4 cells were transfected with nrVL4619 tandem genome plasmid, and the relative abundance of the mRNA transcripts (B) was determined at specific timepoints during the viral life cycle using short-read RNA-seq. Relative expression of viral mRNAs was analyzed by normalization to (C) all viral reads or to (D) host + viral reads. Each bar represents one of three replicates per timepoint. (E and F) Immunoblots showing the detectable viral proteins expressed in MOLT-4 cells over 4 d (lanes representing 24, 28, 72, and 96 h for each condition) using either WT Tandem or SRR constructs. Antibodies against (E) the C-term domain of ORF1 or (F) the N-term domain of ORF2 were used to detect viral proteins. GAPDH was used as a loading control.

DNA virus life cycles tend to follow a cascade of events, with early viral gene product functions biasing toward manipulation of host cell innate defenses and initiating viral genome replication. Viral gene products that are associated with packaging, virion assembly, and egress are usually expressed to higher levels following the onset of genomic replication (40). We used RNA sequencing (RNA-seq) and immunoblotting with antibodies specific for either the C-terminal (C-Term) of ORF1 proteins or the N-terminal (N-term) of ORF2 proteins to investigate the kinetics of viral gene expression. Variants of mRNAs 1 to 3 in which the intron contained in CD2 is not spliced out were detected by RNA-seq (Fig. 1 BD). As the Kozak sequence for ORF2 does not form if this intron is not spliced out, it is not clear whether these mRNA species are immature transcripts or if they are variants that bias toward translation from the ORF1 start codon.

We previously reported the development of an ANV particle production system that uses plasmids containing two tandem ANV genomes to express viral proteins (37). Interestingly, cells transfected with a plasmid containing tandem genomes of TTMV-nrVL4619 identified from human tissue samples using the AnelloScope platform (2, 37), here referred to as nrVL4619, accumulated increasing amounts of the “long” mRNA isoforms over time (Fig. 1 C and D). Another set of minor mRNA species, mRNA2 prime (mRNA2p) and mRNA2p long, were detected (Fig. 1 BD). However, their functions are not yet known. Over a 72-h time course of viral replication and virion production, mRNA3 appears to be the most abundant transcript, particularly at early time points (Fig. 1 C and D). The overall number of viral transcripts detected increased over time, peaking at 48 h posttransfection (hpt) (Fig. 1D). As with the RNA-seq, we used nrVL4619 tandem genome constructs to monitor the expression of viral proteins at 24, 48, 72, and 96 h post transfection (Fig. 1 E and F). Viral proteins translated from ORF1 and ORF1/1 accumulated over the course of 4 d with little to no detectable protein on day 1 and increased to days 3-4 (Fig. 1E). The levels of proteins translated from ORF2/3 and ORF2/2 peaked earlier, between days 2 and 3, while the protein coded by ORF2 peaked from days 3 to 4 (Fig. 1F). We recently reported on the development of a self-replicating rescue plasmid (SRR) designed for expression of viral proteins to trans-rescue replication and packaging of ANV-based gene therapy vectors in MOLT-4 cells (38). The SRR expresses the simian virus 40 (SV40) large T antigen (LT) to drive replication of the plasmid through an SV40 ORI (SI Appendix, Fig. S1). This enables expression of a protein of interest from an upstream cassette in MOLT-4 cells, which are refractory to expression from exogenous DNA. Compared to the nrVL4619 tandem genome construct, ANV protein expression from the SRR demonstrated accelerated expression kinetics and higher protein levels, with high levels of the protein coded by ORF2/3 occurring 1-d posttransfection followed by peak accumulation of ORF2 and ORF2/2 proteins by 2 d posttransfection (Fig. 1F). We currently do not have an antibody reagent capable of detecting the protein coded by ORF1/2. Expression of ORF1 and ORF1/1 proteins were also greatly enhanced from the SRR compared to a tandem genome plasmid. The early accumulation of mRNA3 and the ORF2/2 and ORF2/3 proteins suggests they may play a role in early events of the ANV life cycle, such as innate immune modulation and viral DNA replication.

The Protein Coded by nrVL4619 ORF2/3 Is Necessary and Sufficient to Drive ANV Reporter Construct Replication.

We next investigated the potential for proteins coded by nrVL4619 to drive replication of a reporter construct. Based on RNA-seq and protein data, we hypothesized that proteins translated from mRNA3 may play a role early in the ANV life cycle, including replication of the viral genome. To evaluate this hypothesis, we generated three constructs based on the SRR expression plasmid that contain the sequences of either the mRNA1, mRNA2, or mRNA3 transcripts (SI Appendix, Fig. S1B). Each one of these transcripts can translate two proteins, one from the ORF2 frame and one from the ORF1 frame. We used these expression constructs to drive replication of a reporter construct that has the entire nrVL4619 NCR driving an enhanced green fluorescent protein (eGFP) cassette. Using a Southern blot assay with a set of probes targeting the eGFP coding sequence, we detected DpnI-resistant DNA in cells transfected with the reporter plasmid and either the nrVL4619 SRR or the SRR encoding mRNA3 (Fig. 2A). DpnI targets the methylated adenine within GATC sequences of DNA that were replicated in Escherichia coli, leading to digestion of the input DNA. In contrast, DNA replicated in eukaryotic cells remains resistant to DpnI (41, 42). Cells transfected with SRRs encoding mRNA1 and mRNA2 did not produce DpnI-resistant DNA bands (Fig. 2A). As mRNA3 can produce both ORF2/3 and ORF1/2 proteins, we used SRRs that encoded either protein (SI Appendix, Fig. S1B) to detect replicated DNA via a Southern blot assay. As previously observed, the mRNA3 expression construct produced a DpnI-resistant band of plasmid size (Fig. 2B). Likewise, an SRR expressing a codon-optimized ORF2/3 produced a DpnI-resistant band of plasmid size. In contrast, the SRR expressing a codon-optimized ORF1/2 did not produce a DpnI-resistant band. These observations suggest that the protein coded by ORF2/3 is both necessary and sufficient to initiate replication of ANV NCR containing DNA making it the Rip protein for nrVL4619.

Fig. 2.

Fig. 2.

The ORF2/3 coded protein is necessary and sufficient for initiation of viral DNA replication. (A) Southern blot probing for eGFP DNA sequence in Day 3 lysates of MOLT-4 cells transfected with the nrVL4619 Full-NCR eGFP reporter and one of four separate SRR constructs expressing either the full nrVL4619 coding sequence (nrVL4619 SRR) or each of the three canonical TTMV mRNAs. Each condition was treated with a single-cut restriction enzyme to linearize the nrVL4619 Full-NCR eGFP sequence with (+) or without (−) DpnI. (B) Southern blot probing eGFP DNA sequences as in (C), but for MOLT-4 cells transfected with the nrVL4619 Full-NCR eGFP reporter alone or with SRR constructs expressing mRNA3, codon-optimized ORF1/2, or codon-optimized ORF2/3. Each condition was treated with a single-cut restriction enzyme to linearize the nrVL4619 Full-NCR eGFP sequence with (+) or without (−) DpnI. (C) Immunoblot probing for the protein encoded by ORF2/3 when expressed by constructs expressing the ORF2/3 coding sequence in either an SRR format (+ SRR::ORF2/3) or in non-SRR formats with a CMV promoter driving expression of a codon-optimized ORF2/3 coding sequence (CMV-ORF2/3 and ORF2/3-3xFLAG). GAPDH was probed as a loading control. (D) Southern blot probing for eGFP sequences when nrVL4619 Full-NCR eGFP was transfected into HEK293 cells alone or with ORF2/3 expression constructs: SRR::nrVL4619 (full ANV coding sequence), SRR::EMPTY (expresses no ANV viral proteins), CMV-ORF2/3, or CMV-ORF2/3-3xFLAG. (E) Southern blot probing for eGFP sequences when TTMV-LY2-Full NCR-eGFP was transfected into MOLT-4 cells alone or with expression construct coding the full TTMV-LY2 coding sequence (+ TTMV-LY2 SRR) or the protein encoded by ORF2/3 of TTMV-LY2 (+ SRR::TTMV-LY2-ORF2/3). (F) Southern blot probing for eGFP sequences when TTV-16-Full NCR-eGFP was transfected into MOLT-4 cells alone or with expression construct coding the full TTV-16 coding sequence (+ TTV-16 SRR) or the protein encoded by ORF2/3 of TTV-16 (+ SRR::TTV-16-ORF2/3).

Because the SV40 LT is a viral replication protein, we could not rule out that viral reporter replication observed using an SRR-based construct was driven by LT. While the SRRs for mRNA1, mRNA2, and ORF1/2 did not result in replicated reporter plasmid as detected by Southern blot, we set out to determine whether the protein coded by ORF2/3 could replicate a reporter construct in the absence of LT. We generated two CMV-driven ORF2/3 expression plasmids that do not produce LT, with and without a 3xFLAG tag at the ORF2/3 C-term. Both the ORF2/3 and ORF2/3-3xFLAG constructs (SI Appendix, Fig. S1B) produce the ORF2/3 protein as detected by immunoblot (Fig. 2C). While the ORF2/3 expression construct produced similar or slightly less protein than the SRR-ORF2/3 construct, the ORF2/3-3xFLAG construct produced a greater quantity of protein, possibly through stabilizing the protein (Fig. 2C). We used these expression constructs to transfect HEK293 cells that lack the LT protein. The ORF2/3 expression construct did not produce detectable DpnI-resistant DNA by Southern blot; however, the ORF2/3-3xFLAG construct did produce detectable DpnI-resistant reporter DNA in the absence of LT (Fig. 2D). Interestingly, the truncation of the last 69 amino acids of the C-term of ORF2/3 through introduction of a stop codon in the ORF3 unique sequence also resulted in an increased accumulation of protein (SI Appendix, Fig. S2A), yet the SRR expressing the ORF2/3 truncation was not able to drive replication of a nrVL4619 NCR-containing reporter (SI Appendix, Fig. S2B). Taken together, these experiments demonstrate that the nrVL4619 ORF2/3 protein acts as the viral Rip and is both necessary and sufficient to drive replication of DNA from the viral NCR. In addition, the inability of the ORF2/3 truncation proteins to drive replication of DNA from the viral NCR suggests the C-term domain may be important for its replication function.

The Protein Coded by ORF2/3 Is the Rip for TTMV-LY2 and TTV-16.

Following identification of the nrVL4619 ORF2/3 protein as the Rip responsible for driving replication from the viral NCR, we set out to test whether the protein coded by ORF2/3 was also the Rip for other betatorqueviruses and for alphatorqueviruses. To test this, we first made reporter constructs that have an eGFP cassette driven by the NCR of a betatorquevirus identified from human pleural effusion, TTMV-LY2 (GenBank: JX134045.1) (43) or a similar reporter construct with the entire NCR of TTV-16 (GenBank: AB017613.1) (44), an alphatorquevirus (SI Appendix, Fig. S2C). The NCR of alphatorqueviruses is approximately twice the size of the NCR of betatorqueviruses and contains two GC-rich stretches. We transfected the TTMV-LY2 reporter into MOLT-4 either alone, with a TTMV-LY2 SRR encoding all viral proteins, or a TTMV-LY2-ORF2/3 SRR that expresses only the ORF2/3 protein. DpnI-resistant reporter DNA was detected in both the TTMV-LY2 SRR and the TTMV-LY2-ORF2/3 SRR conditions (Fig. 2E). As observed with TTMV-LY2 constructs, the ORF2/3 of TTV-16 was capable of producing DpnI-resistant DNA from the construct containing the TTV-16 NCR (Fig. 2F). Our results indicated that the proteins coded by ORF2/3 were sufficient to drive DNA replication from the TTMV-LY2 and TTV-16 NCRs. These observations argue that the ORF2/3 protein serves as the Rip for human ANVs and possibly for ANVs infecting other vertebrates.

Influence of Rip on Promoter Activity of nrVL4619.

When driving replication of TTMV-LY2 and TTV-16 reporter constructs Rips, we observed that expression of eGFP from both the TTMV-LY2 and TTV-16 NCR-driven reporter cassettes were dependent on the presence of viral proteins (SI Appendix, Fig. S2 D and E). This dependence suggested that either ORF2/3 can activate the promoter or that the promoter may activate through DNA replication. We aimed to determine how ANV proteins influence viral promoter expression. Constructs with deleted GC (ΔGC) or both GC and CD1 regions (ΔGC-ΔCD1) were designed to remove potential regulatory elements upstream of the TATA box (SI Appendix, Fig. S3A). MOLT-4 cells were transfected with nrVL4619 Full-NCR eGFP, ΔGC, or ΔGC-ΔCD1, with or without nrVL4619 mRNA3, ORF1/2, or ORF2/3 expression constructs. DNA was analyzed via Southern blot. As expected, nrVL4619 Full-NCR eGFP replicated in the presence of mRNA3 or ORF2/3, while ΔGC and ΔGC-ΔCD1 constructs showed no DpnI-resistant DNA, suggesting disrupted replication initiation (SI Appendix, Fig. S3 B and C). As determined by flow cytometry, eGFP positivity was 50 to 70% for cells transfected with nrVL4619 Full-NCR eGFP and ΔGC but <10% for ΔGC-ΔCD1 (SI Appendix, Fig. S3D). MFI analysis showed that both nrVL4619 Full-NCR eGFP and ΔGC constructs had elevated GFP levels with viral proteins, though ΔGC-ΔCD1 showed no increase (SI Appendix, Fig. S3E). These results suggest that viral proteins enhance promoter expression without DNA replication, and the CD1 region is crucial for promoter activity. Interestingly, the results also suggest that GC region elements may be essential for replication from the ANV NCR.

The nrVL4619 Origin of Replication Maps to a 92 bp Sequence Just Downstream of the Coding Sequences.

While there is no experimental evidence to support it, it has been suggested that the ORI for ANVs lies within the CD2 region of the ANV NCR (14, 17, 26). The proposed octanucleotide motif bears some resemblance to the nonanucleotide CRESS DNA virus ORI motif and has been speculated to be part of a hairpin structure that may be cleaved by the previously undescribed replication protein(s), so it has also been referred to as the replication loop. However, the ANV Rip has no homology to the HUH Reps of CRESS DNA viruses. Additionally, we observed that a ΔGC was incapable of replicating DNA in the presence of ANV Rip (SI Appendix, Fig. S3B), suggesting that elements in the GC region may play a role in replication. To evaluate the replication loop hypothesis, we designed a series of constructs based on the nrVL4619 Full-NCR eGFP reporter which have increasingly large truncations from the 3’ end of the NCR (Δ-1-5) (Fig. 3A). Surprisingly, constructs Δ-1-4 produced DpnI-resistant bands as did the nrVL4619 Full-NCR eGFP control when transfected into MOLT-4 cells with the nrVL4619 SRR (Fig. 3B). The replication appeared to be less efficient when portions of the NCR were removed, consistent with results from truncated porcine circovirus ORIs (45). Consistent with our observations from the ΔGC reporter, construct Δ-5, which also removes the GC-rich region, did not to produce DpnI-resistant DNA, indicating a failure to initiate replication. We next designed a series of constructs with truncations from the 5’ end of the contiguous nrVL4619 NCR (Fig. 3C). The smallest truncation, Δ-6, removes a 59nt sequence of DNA that lies between the ANV polyadenylation (polyA) signal and the GC-rich region. Unlike the nrVL4619 Full-NCR eGFP construct, all 5’ truncation constructs, Δ-6-9, failed to produce DpnI-resistant DNA when transfected with the nrVL4619 SRR (Fig. 3D). Together, these results argue that the ANV ORI resides in the 5’ end of the contiguous NCR rather than in the CD2 region.

Fig. 3.

Fig. 3.

The nrVL4619 minimal origin of replication maps to a 92 bp sequence. (A) Linear representations of deletion mutants of nrVL4619 Full-NCR eGFP from the 3’ end of the NCR toward the 5’ targeting key sequence features (Δ-1-5). Targeted regions that were deleted are outlined in red and their respective nomenclature is denoted within each red box. (B) Southern blot probing for eGFP DNA sequences when the deletion mutants Δ-1-5 were cotransfected with an expression construct with the full ANV coding sequence (+ nrVL4619 SRR). (C) Linear representations of deletion mutants of nrVL4619 Full-NCR eGFP from the 5’ end of the NCR toward the 3’ targeting key sequence features (Δ-6-9). Targeted regions that were deleted are outlined in red and their respective nomenclature is denoted within each red box. (D) Southern blot probing for eGFP DNA sequences when the deletion mutants Δ-6-9 were cotransfected with an expression construct with the full ANV coding sequence (+ nrVL4619 SRR). (E) Linear representations of deletion mutants of the Δ-4 construct. The AT-rich near the GC-rich domain is depicted in blue. Deleted sequences are denoted by red Xs. (F) Southern blot probing for eGFP DNA sequences when Δ-4, ΔHairpins, or AT-rich only constructs were cotransfected with the full ANV coding sequence (+ nrVL4619 SRR) in MOLT-4 cells. (G) A linear representation of the nrVL4619 full NCR depicting the 92 bp minimal origin of replication sequence.

The Δ-4 construct contained the minimal amount of nrVL4619 NCR that was sufficient to replicate DNA with nrVL4619 proteins (Fig. 3B). This region contains several notable genetic elements which include a set of three highly conserved hairpin repeats in the GC-rich region and a 27nt AT-rich region (7% GC) immediately upstream of the GC-rich region, a common feature of ORIs for prokaryotes, eukaryotes, and DNA viruses (4650). To map the nrVL4619 minimal ORI, we made further modifications to the Δ-4 construct by deleting the conserved hairpins (ΔHairpins) and by deleting the hairpins and the sequence upstream of the AT-rich region (AT-rich only), leaving the AT-rich region along with the succeeding 34nt of GC-rich region upstream of the hairpins (Fig. 3E). The nrVL4619 Full-NCR eGFP produced DpnI-resistant DNA only in the presence of the nrVL4619 SRR, as expected (Fig. 3F). Surprisingly, like the Δ-4 construct, the ΔHairpins construct produced DpnI-resistant DNA in the presence of the nrVL4619 SRR, indicating that the highly conserved hairpins are not required for initiating replication from the ANV Ori. Loss of the sequence preceding the AT-rich region abolished DNA replication. The 92 bp of nrVL4619 sequence that comprises the ΔHairpins construct has very high homology to the corresponding sequence of TTMV-LY2 (SI Appendix, Fig. S3F). These observations indicate the 92 bp sequence immediately following the ANV polyA signal is sufficient to function as the ORI for betatorqueviruses (Fig. 3G).

nrVL4619 Rip Interacts with the DNA Polymerase α Complex and the BTR Complex.

As the nrVL4619 Rip has no homology to known proteins, including the Reps of CRESS DNA viruses, it is unknown how it might facilitate replication from the ANV ORI. To investigate how Rip might initiate replication, we performed IP-MS on MOLT-4 cells expressing a 3xFLAG-tagged nrVL4619 Rip from an SRR. The Rip sequence was codon optimized to prevent expression from the ORF1 start codon. Rip-3xFLAG was immunoprecipitated with either a FLAG antibody (Ab) or an Ab targeting the ORF2 N-term domain. Gene ontology (GO) analysis revealed that proteins enriched in both the FLAG and ORF2 IP-MS were heavily involved in resolution of recombination intermediates and DNA synthesis (Fig. 4A). Of particular interest, all four subunits of the DNA POLα complex (51) (POLA1 catalytic subunit; POLA2 regulatory subunit; primase subunits PRIM1 and PRIM2) and the core components of the BTR complex (5254) (BLM RecQ helicase [BLM]; Topoisomerase 3A [TOP3A]; RMI1; RMI2) were detected in both pulldown schemes (Fig. 4 B and C). The POLα complex is the polymerase complex that is responsible for the priming of both the leading and lagging strands and initiation of DNA replication at ORIs (55, 56). The BTR complex, also known as the BLM dissolvasome, regulates homologous recombination and facilitates dissolution of Holliday junctions (54, 57). The BTR complex also plays a significant role in the alternative lengthening of telomeres (ALT) pathway and C-circle formation that are hallmarks of many cancers (5860).

Fig. 4.

Fig. 4.

Rip IP-MS is enriched with proteins involved in DNA replication initiation and homologous recombination. (A) Overlapping hits between the FLAG and ORF2 IP/MS datasets with log2 FC > 2 & P < 0.05 [−log10(FDR)] were analyzed using the Go Ontology Biological Process algorithm. (B and C) ORF2 and FLAG IP/MS hits from Rip-expressing MOLT-4 cells are represented in volcano plots. Log2 FC represents the ratio of peptide counts from Rip+ lysates to peptide counts from Rip- lysates. The ORF2/3 bait protein is highlighted in blue, and proteins involved in the BTR and Polymerase α Complex are highlighted in red. (DF) Volcano plots of Rip IP/MS hits from HEK293 (D), 293T (E), and MOLT-4 (F) cells. (G) Table showing BTR and Polymerase α Complex hits in 5 different IP designs. (+) denotes enrichment of log2 FC > 2 with statistical significance P < 0.05 [−log10(FDR)]; (–) denotes lack of enrichment in the indicated IP design.

The initial IP-MS experiment in MOLT-4 cells used constructs that expressed the LT to facilitate expression of proteins from plasmids. Additionally, lysates were not treated with nucleases so that we might capture indirect interactions mediated by nucleic acid. As a follow up experiment, we performed IP-MS in MOLT4 cells transfected with a nrVL4619 tandem genome plasmid and in HEK293 and HEK293T cells transfected with a plasmid in which the CMV promoter drives expression of a codon-optimized 3xFLAG tagged Rip (non-SRR). While LT is present in the HEK293T experimental condition because it is expressed by the cells, LT is absent in the MOLT4 and HEK293 experimental conditions. All three experimental schemes incorporated DNase and RNase treatments to examine more direct interactions. Immunoprecipitation was performed using an antibody specific to the ORF3 domain of Rip. Components of the POLα complex (POLA1, POLA2, and PRIM1) and BTR complex (TOP3A and RMI1) were detected as interactors of nrVL4619 Rip in all three cell lines, including HEK293 and MOLT-4 cells in which no LT was present (Fig. 4 DG). In addition to POLα and BTR components, minichromosome maintenance (MCM) proteins 2-7 were detected in IP-MS from all three cell lines (SI Appendix, Fig. S4 AC), but the enrichment was most significant in MOLT-4 and 293T cells lines. MCM proteins 2-7 are the six subunits of the MCM helicase complex which melts DNA strands at origins of replication (6163). These IP-MS experiments support our observation that the nrVL4619 ORF2/3 protein serves as the ANV Rip. Furthermore, the interaction with the BTR complex suggests that ANVs may use a recombination-related mechanism during replication, a model that is supported by our recent report that tandem ANV genomes resolve to single unit genomes through a recombination mediated mechanism during replication (38).

Discussion

ANVs are exceptionally successful vertebrate viruses, and the near-ubiquitous presence in humans as the dominant feature of the eukaryotic human virome makes it all the more surprising as to how little is known about their replication. In this study, we demonstrated that the ANV protein coded by ORF2/3 of both the Betatorquevirus and Alphatorquevirus genera, referred to as ANV Rip, is necessary and sufficient to initiate replication from cis elements located in the NCR of the ANV genome. Furthermore, we identified components of both the POLα complex and the BTR complex as interacting with ANV Rip in multiple cell types, arguing that the ANV replication mechanism employs host cell machinery-mediated recombination.

While it has been assumed that ANVs would replicate using a form of RCR, no definitive evidence has been found for the presence of a Rep-like protein to facilitate initiation of RCR from the ANV genome. It had been speculated in the literature that the ORF1 protein, demonstrated to be the ANV capsid protein, contained RCR-related motifs that might resemble those found in Rep proteins (29, 32, 33). However, both sequence conservation analysis and predictive structural models called this hypothesis into question (25, 28). Indeed, we found that the ORF1 protein is dispensable for driving replication from viral genetic elements. ANV Rip has no known homology to the HUH Rep proteins of CRESS DNA viruses or to any other known protein, suggesting ANVs may employ a mechanism for the replication of their circular ssDNA genomes that has not been previously described.

We mapped the minimal nrVL4619 ORI to a 92 bp sequence of DNA that bears no resemblance to the replication loops of CRESS DNA viruses. The ANV ORI has an AT-rich stretch of DNA that is similar to the AT-rich stretches found at the ORIs of prokaryotes, eukaryotes, and dsDNA viruses (4750, 56). The SV40 ORI is embedded within its own promoter and has a 17 bp AT-rich stretch (47, 64). Herpes Simplex Virus (HSV) also has AT-rich stretches at its ORI (48). AT-rich sequences are often found at origins of DNA replication as the fewer hydrogen bonds between adenine and thymine facilitate the unwinding of the double-stranded helix of DNA (47, 49). Unwinding of DNA at the ORI allows replication machinery to access the individual strands to prime and initiate replication. Interestingly, there is no evidence of a helicase domain or activity for ANV Rip. However, the detection of all six subunits of the MCM helicase complex argues that ANV Rip recruits the MCM helicase complex to melt open the ANV ORI. It is also unclear whether the ANV ORI initiates replication in a unidirectional or bidirectional manner. While RCR generally proceeds in a unidirectional manner, the ANV ORI may more closely resemble bidirectional cellular ORIs or the SV40 ORI than the replication loops of CRESS DNA viruses. It is also unknown whether the ANV ORI undergoes nicking, as observed for CRESS DNA virus RCR processes, but the presence of the MCM helicase complex proteins suggest ANV ORIs undergo a melting process to open the viral origin and initiate replication fork formation. In order to fully understand the mechanism ANV Rip uses to initiate viral replication, it will be important to experimentally determine whether ANV Rip directly binds to the viral origin of replication.

It is important to note that the data described in this study was gathered using a dsDNA version of the viral NCR and ORI sequence. While ANV genomes are circular ssDNA in the virion (38), upon entry to the host cell nucleus, the genome is converted to dsDNA which allows for gene expression. The process through which the ANV genome is converted from ssDNA to dsDNA is currently unknown. Thus far, there is no evidence that ANV Rip is packaged in viral particles, and its ability to promote replication from a ssDNA origin remains to be tested. Unlike parvoviruses, ANVs do not have a free end of DNA to form a complementary structure and prime for ssDNA to dsDNA conversion. It is possible that ANVs package a small oligonucleotide that complements the ssDNA genome to prime for second strand synthesis as has been demonstrated for porcine circoviruses (65). Further studies will be required to understand this step in the ANV lifecycle.

The detection of BTR complex components interacting with the ANV Rip offers some intriguing possibilities for modes of replication. The ALT pathway requires localization of the BTR complex to telomeres (58). The ALT pathway is a recombination-dependent form of telomere replication that is a hallmark of several forms of cancer, but it is also known to play a role in the replication and maintenance of some DNA virus genomes (6669). Many DNA viruses employ recombination strategies in the replication of their DNA (70, 71), and ANVs are well documented to have recombination hotspots within their genomes (2, 17). Supporting the use of recombination in ANV replication, we recently reported that replicating tandem ANV genomes resolve single unit genomes through a mechanism seemingly regulated by homologous recombination (38). It is of note that the BTR complex is required for the formation of extratelomeric ssDNA C-circles (58, 60). It is possible the BTR complex plays a role in the formation of the genomic ssDNA in a manner that resembles the formation of C-circles. It is also interesting to note that TOP3A was detected in all IP-MS conditions. TOP3A has an isoform that localizes to the mitochondria and facilitates dissolution of hemicatenated replication products after D-loop-mediated replication of the circular mitochondrial genome; (72) thus, it is possible TOP3A may play a role in the dissolution of ANV replication intermediates. While additional experimentation is required to elucidate the role of TOP3A and the BTR complex in ANV replication, it is clear that host cell recombination machinery plays an important role in this process.

To gain insight into the nrVL4619 Rip, we used AlphaFold to model its structure (73). The models suggest that the conserved ANV Rip may be largely unstructured in its ORF3 domain, while the N terminus ORF2 domain has helical structures that coordinate a Zn2+ ion (SI Appendix, Fig. S5A). This zinc finger-like domain may target the ANV origin, possibly dimerized with another ORF2-domain protein, or affect host or viral protein function (74). ORF2 and ORF2/2, provided by constructs expressing mRNA1 and mRNA2, respectively (Fig. 2A), were insufficient to initiate replication but contain the same N-term domain as Rip. The model predicts possible interaction between N-term ORF2 domains of two Rip proteins, both coordinating Zn2+ (SI Appendix, Fig. S5 BD). This raises the possibility that interactions between ORF2 domain-containing proteins modulate replication processes, but the functions of other ORF2 domain-containing proteins remain to be explored. The C-terminus ORF3 domain is largely unstructured, except for a single helix. Intrinsically disordered regions (IDRs) are often associated with liquid–liquid phase separation (LLPS) (75), and many viral proteins with IDRs form compartments through LLPS (76). This may aid the virus in establishing a replication compartment to sequester replication and packaging proteins while excluding restriction factors. Folding predictions on the TTMV-LY2 ORF2/3 protein showed a structure similar to that of the nrVL4619 Rip (SI Appendix, Fig. S5 EG). Further studies on ANV ORF2 protein structure and interaction are needed to validate these models.

We observed that adding a 3xFLAG tag to or truncating the last 69 amino acids of the Rip C-term, resulted in enhanced accumulation of protein as detected by immunoblot. This suggests that the C-term of Rip may play a role in protein stability and turnover. Furthermore, truncating the last 69 amino acids of Rip resulted in a loss of reporter DNA replication, suggesting the C-term domain of Rip is important for its function in replication. This is supported by the observation that the other ORF2 N-term domain–containing proteins, ORF2 and ORF2/2, were not required for replication. Additional studies will be needed to understand the properties of the domains within Rip and the roles they play in targeting the ANV ORI, recruiting host cell machinery, homo- and heterodimerization, and protein stability.

Methods and Materials

Cell Culture.

MOLT-4 cells were sourced from the National Cancer Institute. These cells were cultured at 37 °C with 5% CO2 in suspension culture using a complete growth medium (Gibco’s RPMI 1640 with 10% fetal bovine serum (FBS), supplemented with 1 mM sodium pyruvate, 0.1% Pluronic F-68, and 2 mM L-glutamine) while shaking at 100 rpm with over 85% relative humidity (RH). HEK293 and HEK293T cells were obtained from the American Type Culture Collection (Manassas, VA) and maintained in Dulbecco’s Modified Eagle’s Medium (DMEM; Corning, Corning NY) supplemented with 10% (v/v) FBS in humidified incubators at 5% CO2 and 37 °C. Cells were regularly passaged by treatment with trypsin-EDTA (0.05%; Corning) and transferred to fresh media.

Plasmid Construction.

For expanded plasmid construction methods, see SI Appendix. A plasmid containing two tandem copies of the TTMV-nrVL4619 genome was constructed by synthesizing a single copy flanked by BsaI sites, circularizing it in vitro, and ligating it to another linearized copy. Specialized SRR plasmids were designed to express various forms of nrVL4619 mRNAs and ORFs using a human EF1a promoter and EMCV IRES2, and cloned into a pUC57-Kan vector. Variants including truncations and codon-optimized ORFs for TTV-16 and TTMV-LY2 were also created using GenScript seamless cloning. Codon-optimized ORF2/3 and ORF2/3-3xFLAG sequences were cloned into a CMV-driven pcDNA3.1(+) vector for expression studies. Reporter plasmids incorporating full or truncated NCRs from nrVL4619, TTV-16, and TTMV-LY2 followed by eGFP and SV40 poly(A) signal were synthesized to study replication origins.

Transfections.

MOLT-4 cells were electroporated in 2S buffer (5 mM KCl, 15 mM MgCl2, 15 mM HEPES, 150 mM Na2HPO4, pH 7.2, 50 mM sodium succinate) using a NEPA electroporator. Cell pellets (1 × 10^7 cells) were resuspended in 500 µL 2S buffer with 50 µg of expression and eGFP plasmids for five cuvettes with identical DNA content. After electroporation, 300 µL of prewarmed medium was added to each cuvette, gently mixed to dissociate clumps, and transferred to 25 mL of prewarmed media in a flask. Incubate at 37 °C with shaking at 125 rpm for 3 d. Prior to harvesting, image cells in brightfield and GFP channels at 10× magnification using an EVOS M5000 microscope (Thermo Fisher, AMF5000SV). HEK293 and HEK293T cells were transfected using PEIpro (Polyplus) per the manufacturer’s protocol. Incubate at 37 °C for the indicated time.

RNA-seq.

Total RNA was extracted from transfected cells using Trizol (Thermo Fisher Cat# 15596026) per the manufacturer’s protocol followed by rRNA depletion. NEBNext Ultra II Directional RNA Library Prep For Illumina (NEB Cat# E7760L) was used to prepare the sequencing library from the rRNA-depleted RNA samples per the manufacturer’s protocol. Samples were sequenced with an Illumina NextSeq. Sequencing reads were processed using nf-core/rnaseq v3.17.0 (DOI: 10.5281/zenodo.1400710) of the nf-core collection of workflows (77), utilizing reproducible software environments from the Bioconda (78) and Biocontainers (79) projects. We used the “star_salmon” workflow and chose fastp v0.23.4 for read quality filtering, while otherwise retaining the default settings and software versions used in nf-core/rnaseq v3.17.0. For the genome reference, we used the ENSEMBL v113 Homo sapiens genome, together with the genome sequence of nrVL4619, plus their respective annotations. We used the resulting alignments form the RNA-seq pipeline to quantify the host and nrVL4619 transcript isoforms using StringTie v2.2.3 (80), in units of Transcripts per Million Mapped (TPM), then extracted the nrVL4619 transcripts using samtools v1.20 (81) to quantify the relative TPM values of the nrVL4619 transcripts exclusively.

Immunoblotting.

Cells were harvested 3 d postelectroporation, washed with PBS, and lysed in 0.5% Triton-X, 300 mM NaCl, and 50 mM Tris pH 8.0. Genomic DNA was degraded using mSAN, and proteins were denatured with LI-COR protein sample loading buffer and reducing agent at 95 °C for 10 min. Lysates were run on BOLT SDS-PAGE gels (12% for ORF2 and 4 to 12% for ORF1). Proteins were transferred to nitrocellulose membranes using the Bio-Rad Trans-Blot Turbo Transfer system. Membranes were blocked for 1 h with LI-COR Intercept Blocking buffer, followed by overnight primary antibody incubation (1:1,000) at 4 °C. GAPDH detection used an antibody from Cell Signaling (Cat#97166S), while ORF1 and ORF2 used GenScript antibodies. Membranes were then incubated with IRDye® 800CW and 680CW goat IgG secondary antibodies (1:10,000) from LI-COR for 2 h and imaged using the Bio-Rad ChemiDoc MP imaging system.

Southern Assay.

For Southern Blot analysis, transfected MOLT-4 cells were lysed with the DNeasy Blood and Tissue Kit (QIAgen #69504) and proteinase K. The DNA was linearized overnight with a restriction endonuclease specific to the viral DNA. DpnI (NEB #R0176) was used to digest input plasmid DNA. DNA was electrophoresed on a 1.0% agarose gel, depurinated, denatured, and transferred overnight onto a Hybond-N+ membrane (Cytiva #RPN203B). The membrane was prehybridized in ULTRAhyb buffer (Thermo Fisher #AM8670) and probed overnight with biotin-labeled oligos specific to viral DNA. Probes were generated with the BioPrime™ Array CGH Genomic Labeling System (Thermo Fisher #18095011). Membranes were blocked, incubated with IRDye800CW Streptavidin (LI-COR #926-32230), and imaged using the Bio-Rad ChemiDoc MP Imaging System. A mixed ladder (Φ+λ) composed of ΦX174 DNA-HaeIII Digest (NEB N3026S) and Lambda DNA HindIII Digest (NEB 3012S) was used as a marker to determine DNA band size.

Flow Cytometry.

The entire culture of cells was harvested by pipetting into a 50 mL conical tube at 3 d post transfection. 300 μL of cells were transferred to 96-well round bottom plates. A Cytek Guava clow cytometer was used to measure %GFP and mean fluorescence intensity over 10,000 events.

IP/MS.

For immunoprecipitation, HEK293, HEK293T, and MOLT-4 cells were cotransfected with constructs expressing nrVL4619 ORF2/3 or ORF2/3-FLAG and the nrVL4619 NCR to generate Bait (+) lysates. Bait (−) lysates were generated by transfecting cells with the NCR construct alone. Cells were harvested 2 d posttransfection, washed, and lysed at 4C for 30 min in nondenaturing buffer (0.5% Triton-X, 300 mM NaCl, 50 mM Tris pH 8.0, protease/phosphatase inhibitor) with or without mSAN and RNase A. mSAN/RNase A-treated lysates were incubated for 1 h at 37C. Cell debris was pelleted, and the clarified supernatants were subjected to IP using bead-conjugated antibodies targeting the C-term FLAG tag, the ORF2 N-term domain, or the ORF3 C-term domain of ORF2/3. ORF2 and ORF2/3 antibodies were conjugated to M-270 Epoxy beads (Thermo Fisher #14311D) for FLAG IPs, and ANTI-FLAG® M2 Magnetic Beads (Sigma # M8823) were used for ORF2/3 IPs. Cell lysates were incubated with antibody-bead slurries overnight at 4C, beads were pelleted, washed twice with cold lysis buffer, and proteins were eluted using 4X Protein Sample Loading buffer (LI-COR). Beads were incubated in elution buffer at 65C for 20 min, separated from the eluate, and stored at −80 °C until quantitative proteomic sample preparation and analysis (performed at IQ Proteomics, Cambridge, MA). Fold change values represent the ratio of Bait+ signal to Bait- signal, and P-values were calculated using the two-sample t test.

AlphaFold Server.

Protein sequence from nrVL4619 and TTMV-LY2 were input into the AlphaFold server (https://alphafoldserver.com/). pLDDT is a per-atom confidence estimate on a 0-100 scale where a higher value indicates higher confidence. This information is subject to AlphaFold Server Output Terms of Use found at alphafoldserver.com/output-terms.

Supplementary Material

Appendix 01 (PDF)

Acknowledgments

We would like to thank Brian Luque, Chris McNulty, and Kelly Morgan for facilitating this manuscript.

Author contributions

N.B., S.T., D.V., and J.C. designed research; N.B., S.T., C.E., N.S., J.M., K.T., and C.D. performed research; C.P. and M.N. contributed new reagents/analytic tools; N.B., S.T., C.E., P.J., C.P., N.S., K.T., C.D., G.P., and J.C. analyzed data; D.V. preliminary research and concepts; G.P. and J.C. project supervision; and J.C. wrote the paper.

Competing interests

All authors were employees of Ring Therapeutics during their contributions to this manuscript.

Footnotes

This article is a PNAS Direct Submission. A.V. is a guest editor invited by the Editorial Board.

Contributor Information

Geoffrey Parsons, Email: gparsons@ringtx.com.

Joseph Cabral, Email: cabral.joseph@gmail.com.

Data, Materials, and Software Availability

The data from the RNA-sequencing of anellovirus transcripts have been deposited in the Sequence Read Archive at the National Center for Biotechnology Information under BioProject PRJNA1366609 (82), the genome reference for nrVL4619 and the IP-mass spectrometry data are both deposited at Figshare (https://doi.org/10.6084/m9.figshare.30698825) (83). All other data are included in the manuscript and/or SI Appendix.

Supporting Information

References

  • 1.Freer G., et al. , The virome and its major component, anellovirus, a convoluted system molding human immune defenses and possibly affecting the development of asthma and respiratory diseases in childhood. Front. Microbiol. 9, 686 (2018), 10.3389/fmicb.2018.00686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Arze C. A., et al. , Global genome analysis reveals a vast and dynamic anellovirus landscape within the human virome. Cell Host Microbe 29, 1305–1315.e6 (2021), 10.1016/j.chom.2021.07.001. [DOI] [PubMed] [Google Scholar]
  • 3.Okamoto H., et al. , TT virus mRNAs detected in the bone marrow cells from an infected individual. Biochem. Biophys. Res. Commun. 279, 700–707 (2000), 10.1006/bbrc.2000.4012. [DOI] [PubMed] [Google Scholar]
  • 4.Kaczorowska J., et al. , Diversity and long-term dynamics of human blood anelloviruses. J. Virol. 96, e00109-22 (2022), 10.1128/jvi.00109-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Cebriá-Mendoza M., et al. , Deep viral blood metagenomics reveals extensive anellovirus diversity in healthy humans. Sci. Rep. 11, 6921 (2021), 10.1038/s41598-021-86427-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Taylor L. J., Keeler E. L., Bushman F. D., Collman R. G., The enigmatic roles of Anelloviridae and Redondoviridae in humans. Curr. Opin. Virol. 55, 101248 (2022), 10.1016/j.coviro.2022.101248. [DOI] [PubMed] [Google Scholar]
  • 7.Xin X., et al. , Mother-to-infant vertical transmission of transfusion transmitted virus in South China. J. Perinat. Med. 32, 404–406 (2004), 10.1515/JPM.2004.136. [DOI] [PubMed] [Google Scholar]
  • 8.Kyathanahalli C., Snedden M., Hirsch E., Human anelloviruses: Prevalence and clinical significance during pregnancy. Front. Virol. 1, 782886 (2021), 10.3389/fviro.2021.782886. [DOI] [Google Scholar]
  • 9.Morrica A., et al. , TT virus: Evidence for transplacental transmission. J. Infect. Dis. 181, 803–804 (2000), 10.1086/315296. [DOI] [PubMed] [Google Scholar]
  • 10.Brundin P. M. A., Landgren B.-M., Fjällström P., Johansson A. F., Nalvarte I., Blood hormones and torque teno virus in peripheral blood mononuclear cells. Heliyon 6, e05535 (2020), 10.1016/j.heliyon.2020.e05535. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kosulin K., et al. , Post-transplant replication of torque teno virus in granulocytes. Front. Microbiol. 9, 2956 (2018), 10.3389/fmicb.2018.02956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Miller M. M., Jarosinski K. W., Schat K. A., Positive and negative regulation of chicken anemia virus transcription. J. Virol. 79, 2859–2868 (2005), 10.1128/JVI.79.5.2859-2868.2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Spezia P.G., et al. , TTV and other anelloviruses: The astonishingly wide spread of a viral infection. Asp. Mol. Med. 1 (2023), 10.1016/j.amolm.2023.100006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kaczorowska J., van der Hoek L., Human anelloviruses: Diverse, omnipresent and commensal members of the virome. FEMS Microbiol. Rev. 44, 305–313 (2020), 10.1093/femsre/fuaa007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Nishizawa T., et al. , A novel DNA virus (TTV) associated with elevated transaminase levels in posttransfusion hepatitis of unknown etiology. Biochem. Biophys. Res. Commun. 241, 92–97 (1997), 10.1006/bbrc.1997.7765. [DOI] [PubMed] [Google Scholar]
  • 16.Varsani A., et al. , Taxonomic update for mammalian anelloviruses (family Anelloviridae). Arch. Virol. 166, 2943–2953 (2021), 10.1007/s00705-021-05192-x. [DOI] [PubMed] [Google Scholar]
  • 17.Modha S., Hughes J., Orton R. J., Lytras S., Expanding the genomic diversity of human anelloviruses. Virus Evol. 11, veaf002 (2025), 10.1093/ve/veaf002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Gore E. J., Gard L., Niesters H. G. M., Van Leer Buter C. C., Understanding torquetenovirus (TTV) as an immune marker. Front. Med. 10, 1168400 (2023), 10.3389/fmed.2023.1168400. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Vietzen H., et al. , Torque teno viruses exhaust and imprint the human immune system via the HLA-E/NKG2A axis. Front. Immunol. 15, 1447980 (2024), 10.3389/fimmu.2024.1447980. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Timmerman A. L., Schönert A. L. M., van der Hoek L., Anelloviruses versus human immunity: How do we control these viruses? FEMS Microbiol. Rev. 48, fuae005 (2024), 10.1093/femsre/fuae005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chen D. G., et al. , Integrative systems biology reveals NKG2A-biased immune responses correlate with protection in infectious disease, autoimmune disease, and cancer. Cell Rep. 43, 113872 (2024), 10.1016/j.celrep.2024.113872. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mouton W., et al. , Torque teno virus viral load as a marker of immune function in allogeneic haematopoietic stem cell transplantation recipients. Viruses 12, 1292 (2020), 10.3390/v12111292. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Krupovic M., et al. , Cressdnaviricota: A virus phylum unifying seven families of rep-encoding viruses with single-stranded, circular DNA genomes. J. Virol. 94, e00582-20 (2020), 10.1128/jvi.00582-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhao L., Rosario K., Breitbart M., Duffy S., “Chapter three—Eukaryotic circular rep-encoding single-stranded dna (cress dna) viruses: Ubiquitous viruses with small genomes and a diverse host range” in Advances in Virus Research, Kielian M., Mettenleiter T. C., Roossinck M. J., Eds. (Academic Press, 2019), pp. 71–133, 10.1016/bs.aivir.2018.10.001. [DOI] [PubMed] [Google Scholar]
  • 25.Rosario K., Duffy S., Breitbart M., A field guide to eukaryotic circular single-stranded DNA viruses: Insights gained from metagenomics. Arch. Virol. 157, 1851–1871 (2012), 10.1007/s00705-012-1391-y. [DOI] [PubMed] [Google Scholar]
  • 26.de Villiers E.-M., Borkosky S. S., Kimmel R., Gunst K., Fei J.-W., The diversity of torque teno viruses: In vitro replication leads to the formation of additional replication-competent subviral molecules. J. Virol. 85, 7284–7295 (2011), 10.1128/JVI.02472-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Takahashi K., Ohta Y., Mishiro S., Partial ~2.4-kb sequences of TT virus (TTV) genome from eight Japanese isolates: Diagnostic and phylogenetic implications. Hepatol. Res. 12, 111–120 (1998), 10.1016/S1386-6346(98)00042-4. [DOI] [Google Scholar]
  • 28.Butkovic A., et al. , Evolution of anelloviruses from a circovirus-like ancestor through gradual augmentation of the jelly-roll capsid protein. Virus Evol. 9, vead035 (2023), 10.1093/ve/vead035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Hijikata M., Takahashi K., Mishiro S., Complete circular DNA genome of a TT virus variant (isolate name SANBAN) and 44 partial ORF2 sequences implicating a great degree of diversity beyond genotypes. Virology 260, 17–22 (1999), 10.1006/viro.1999.9797. [DOI] [PubMed] [Google Scholar]
  • 30.Liou S., et al. , Structure of anellovirus-like particles reveal a mechanism for immune evasion. Nat. Commun. 15, 7219 (2024), 10.1038/s41467-024-51064-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Petersen G. F., et al. , Nuclear trafficking of Anelloviridae capsid protein ORF1 reflects modular evolution of subcellular targeting signals. Virus Evol. 11, veaf069 (2025), 10.1093/ve/veaf069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Huang Y. W., Ni Y. Y., Dryman B. A., Meng X. J., Multiple infection of porcine torque teno virus in a single pig and characterization of the full-length genomic sequences of four U.S. prototype PTTV strains: Implication for genotyping of PTTV. Virology 396, 289–297 (2010), 10.1016/j.virol.2009.10.031. [DOI] [PubMed] [Google Scholar]
  • 33.Biagini P., et al. , Genetic analysis of full-length genomes and subgenomic sequences of TT virus-like mini virus human isolates. J. Gen. Virol. 82, 379–383 (2001), 10.1099/0022-1317-82-2-379. [DOI] [PubMed] [Google Scholar]
  • 34.Webb B., Rakibuzzaman A., Ramamoorthy S., Torque teno viruses in health and disease. Virus Res. 285, 198013 (2020), 10.1016/j.virusres.2020.198013. [DOI] [PubMed] [Google Scholar]
  • 35.Fisher M., et al. , Discovery and comparative genomic analysis of a novel equine anellovirus, representing the first complete Mutorquevirus genome. Sci. Rep. 13, 3703 (2023), 10.1038/s41598-023-30875-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Focosi D., Antonelli G., Pistello M., Maggi F., Torquetenovirus: The human virome from bench to bedside. Clin. Microbiol. Infect. 22, 589–593 (2016), 10.1016/j.cmi.2016.04.007. [DOI] [PubMed] [Google Scholar]
  • 37.Nawandar D. M., et al. , Human anelloviruses produced by recombinant expression of synthetic genomes. bioRxiv [Preprint] (2022). 10.1101/2022.04.28.489885 (Accessed 21 February 2025). [DOI]
  • 38.Prince C., et al. , A novel gene delivery platform based on a commensal human anellovirus demonstrates transduction in multiple tissue types. Mol. Ther. Methods Clin. Dev. 33, 101597 (2025), 10.1016/j.omtm.2025.101597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Manthei K. A., Keck J. L., The BLM dissolvasome in DNA replication and repair. Cell. Mol. Life Sci. 70, 4067–4084 (2013), 10.1007/s00018-013-1325-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Schmid M., Speiseder T., Dobner T., Gonzalez R. A., DNA virus replication compartments. J. Virol. 88, 1404–1420 (2014), 10.1128/jvi.02046-13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Rao B. S., Martin R. G., DpnI assay for DNA replication in animal cells: Enzyme-resistant material can result from factors not related to replication. Nucleic Acids Res. 16, 4171 (1988), 10.1093/nar/16.9.4171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wilson V. G., Cell culture assay for transient replication of human and animal papillomaviruses. Curr. Protoc. Microbiol. 24, 14B.1.1–14B.1.18 (2012), 10.1002/9780471729259.mc14b01s24. [DOI] [PubMed] [Google Scholar]
  • 43.Galmès J., et al. , Potential implication of new torque teno mini viruses in parapneumonic empyema in children. Eur. Respir. J. 42, 470–479 (2013), 10.1183/09031936.00107212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Okamoto H., et al. , The entire nucleotide sequence of a TT virus isolate from the United States (TUS01): Comparison with reported isolates and phylogenetic analysis. Virology 259, 437–448 (1999), 10.1006/viro.1999.9769. [DOI] [PubMed] [Google Scholar]
  • 45.Mankertz A., Persson F., Mankertz J., Blaess G., Buhk H. J., Mapping and characterization of the origin of DNA replication of porcine circovirus. J. Virol. 71, 2562–2566 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Deyerle K. L., Sajjadi F. G., Subramani S., Analysis of origin of DNA replication of human papovavirus BK. J. Virol. 63, 356–365 (1989), 10.1128/JVI.63.1.356-365.1989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Hertz G. Z., Young M. R., Mertz J. E., The A+T-rich sequence of the simian virus 40 origin is essential for replication and is involved in bending of the viral DNA. J. Virol. 61, 2322–2325 (1987), 10.1128/JVI.61.7.2322-2325.1987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Olsson M., et al. , Stepwise evolution of the Herpes Simplex Virus origin binding protein and origin of replication. J. Biol. Chem. 284, 16246–16255 (2009), 10.1074/jbc.M807551200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Rajewska M., Wegrzyn K., Konieczny I., AT-rich region and repeated sequences–The essential elements of replication origins of bacterial replicons. FEMS Microbiol. Rev. 36, 408–434 (2012), 10.1111/j.1574-6976.2011.00300.x. [DOI] [PubMed] [Google Scholar]
  • 50.Yella V. R., Vanaja A., Kulandaivelu U., Kumar A., Delving into eukaryotic origins of replication using DNA structural features. ACS Omega 5, 13601–13611 (2020), 10.1021/acsomega.0c00441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Yuan Z., Georgescu R., Li H., O’Donnell M. E., Molecular choreography of primer synthesis by the eukaryotic Pol α-primase. Nat. Commun. 14, 3697 (2023), 10.1038/s41467-023-39441-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Xu D., et al. , RMI, a new OB-fold complex essential for Bloom syndrome protein to maintain genome stability. Genes Dev. 22, 2843–2855 (2008), 10.1101/gad.1708608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Pradhan A., Singh T. R., Ali A. M., Wahengbam K., Meetei A. R., Monopolar spindle 1 (MPS1) protein-dependent phosphorylation of RecQ-mediated genome instability protein 2 (RMI2) at serine 112 is essential for BLM-Topo III α-RMI1-RMI2 (BTR) protein complex function upon spindle assembly checkpoint (SAC) activation during mitosis. J. Biol. Chem. 288, 33500–33508 (2013), 10.1074/jbc.M113.470823. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Shorrocks A.-M. K., et al. , The bloom syndrome complex senses RPA-coated single-stranded DNA to restart stalled replication forks. Nat. Commun. 12, 585 (2021), 10.1038/s41467-020-20818-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.D’Urso G., Grallert B., Nurse P., DNA polymerase alpha, a component of the replication initiation complex, is essential for the checkpoint coupling S phase to mitosis in fission yeast. J. Cell Sci. 108, 3109–3118 (1995), 10.1242/jcs.108.9.3109. [DOI] [PubMed] [Google Scholar]
  • 56.Aria V., Yeeles J. T. P., Mechanism of bidirectional leading-strand synthesis establishment at eukaryotic DNA replication origins. Mol. Cell 73, 199–211.e10 (2019), 10.1016/j.molcel.2018.10.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.West S. C., et al. , Resolution of recombination intermediates: Mechanisms and regulation. Cold Spring Harb. Symp. Quant. Biol. 80, 103–109 (2015), 10.1101/sqb.2015.80.027649. [DOI] [PubMed] [Google Scholar]
  • 58.Loe T. K., et al. , Telomere length heterogeneity in ALT cells is maintained by PML-dependent localization of the BTR complex to telomeres. Genes Dev. 34, 650–662 (2020), 10.1101/gad.333963.119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Sobinoff A. P., et al. , BLM and SLX4 play opposing roles in recombination-dependent replication at human telomeres. EMBO J. 36, 2907–2919 (2017), 10.15252/embj.201796889. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Chen Y.-Y., et al. , The C-circle biomarker is secreted by alternative-lengthening-of-telomeres positive cancer cells inside exosomes and provides a blood-based diagnostic for ALT activity. Cancers (Basel) 13, 5369 (2021), 10.3390/cancers13215369. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Lei M., Tye B. K., Initiating DNA synthesis: From recruiting to activating the MCM complex. J. Cell Sci. 114, 1447–1454 (2001), 10.1242/jcs.114.8.1447. [DOI] [PubMed] [Google Scholar]
  • 62.Bell S. D., Botchan M. R., The minichromosome maintenance replicative helicase. Cold Spring Harb. Perspect. Biol. 5, a012807 (2013), 10.1101/cshperspect.a012807. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Henrikus S. S., et al. , Unwinding of a eukaryotic origin of replication visualized by cryo-EM. Nat. Struct. Mol. Biol. 31, 1265–1276 (2024), 10.1038/s41594-024-01280-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Parsons R., Anderson M. E., Tegtmeyer P., Three domains in the simian virus 40 core origin orchestrate the binding, melting, and DNA helicase activities of T antigen. J. Virol. 64, 509–518 (1990). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Cheung A. K., Specific functions of the Rep and Rep׳ proteins of porcine circovirus during copy-release and rolling-circle DNA replication. Virology 481, 43–50 (2015), 10.1016/j.virol.2015.01.004. [DOI] [PubMed] [Google Scholar]
  • 66.Deng Z., Wang Z., Lieberman P. M., Telomeres and viruses: Common themes of genome maintenance. Front. Oncol. 2, 201 (2012), 10.3389/fonc.2012.00201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Lippert T. P., et al. , Oncogenic herpesvirus KSHV triggers hallmarks of alternative lengthening of telomeres. Nat. Commun. 12, 512 (2021), 10.1038/s41467-020-20819-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Kamranvar S. A., Chen X., Masucci M. G., Telomere dysfunction and activation of alternative lengthening of telomeres in B-lymphocytes infected by Epstein-Barr virus. Oncogene 32, 5522–5530 (2013), 10.1038/onc.2013.189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Deng Z., et al. , HSV-1 remodels host telomeres to facilitate viral replication. Cell Rep. 9, 2263–2278 (2014), 10.1016/j.celrep.2014.11.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Pérez-Losada M., Arenas M., Galán J. C., Palero F., González-Candelas F., Recombination in viruses: Mechanisms, methods of study, and evolutionary consequences. Infect. Genet. Evol. 30, 296–307 (2015), 10.1016/j.meegid.2014.12.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Lo Piano A., Martínez-Jiménez M. I., Zecchi L., Ayora S., Recombination-dependent concatemeric viral DNA replication. Virus Res. 160, 1–14 (2011), 10.1016/j.virusres.2011.06.009. [DOI] [PubMed] [Google Scholar]
  • 72.Nicholls T. J., et al. , Topoisomerase 3α is required for decatenation and segregation of human mtDNA. Mol. Cell 69, 9–23.e6 (2018), 10.1016/j.molcel.2017.11.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Abramson J., et al. , Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024), 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Li X., et al. , Structures and biological functions of zinc finger proteins and their roles in hepatocellular carcinoma. Biomark. Res. 10, 2 (2022), 10.1186/s40364-021-00345-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Martin E. W., Holehouse A. S., Intrinsically disordered protein regions and phase separation: Sequence determinants of assembly or lack thereof. Emerg. Top. Life Sci. 4, 307–329 (2020), 10.1042/ETLS20190164. [DOI] [PubMed] [Google Scholar]
  • 76.Brocca S., Grandori R., Longhi S., Uversky V., Liquid-liquid phase separation by intrinsically disordered protein regions of viruses: Roles in viral life cycle and control of virus-host interactions. Int. J. Mol. Sci. 21, 9045 (2020), 10.3390/ijms21239045. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Ewels P. A., et al. , The nf-core framework for community-curated bioinformatics pipelines. Nat. Biotechnol. 38, 276–278 (2020), 10.1038/s41587-020-0439-x. [DOI] [PubMed] [Google Scholar]
  • 78.Grüning B., et al. ; Bioconda Team, Bioconda: Sustainable and comprehensive software distribution for the life sciences. Nat. Methods 15, 475–476 (2018), 10.1038/s41592-018-0046-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.da Veiga Leprevost F., et al. , BioContainers: An open-source and community-driven framework for software standardization. Bioinformatics 33, 2580–2582 (2017), 10.1093/bioinformatics/btx192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Pertea M., et al. , StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat. Biotechnol. 33, 290–295 (2015), 10.1038/nbt.3122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Danecek P., et al. , Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021), 10.1093/gigascience/giab008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Boisvert N., et al. , Transcriptome sequencing of an anellovirus-like synthetic construct. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1366609. Deposited 19 November 2025.
  • 83.Boisvert N., et al. , Supplemental data for “Anellovirus protein encoded by ORF2/3 functions as the viral replication initiation protein.” Figshare Datasets. 10.6084/m9.figshare.30698825.v1. Deposited 9 December 2025. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Data Availability Statement

The data from the RNA-sequencing of anellovirus transcripts have been deposited in the Sequence Read Archive at the National Center for Biotechnology Information under BioProject PRJNA1366609 (82), the genome reference for nrVL4619 and the IP-mass spectrometry data are both deposited at Figshare (https://doi.org/10.6084/m9.figshare.30698825) (83). All other data are included in the manuscript and/or SI Appendix.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES