Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2025 Aug 21;122(34):e2513552122. doi: 10.1073/pnas.2513552122

Definition of the components required for selective packaging of coronavirus genomic RNA

Janice D Pata a,b, Taína K Stevens c, Lili Kuo d, Paul S Masters b,d,1
PMCID: PMC12403154  PMID: 40838885

Significance

Coronaviruses selectively package their genomic RNA into assembling virions, excluding all other viral and host RNA species. This high selectivity is critical for efficient replication in vivo. In the prototype Betacoronavirus mouse hepatitis virus packaging specificity is due to an RNA packaging signal (PS) embedded in the viral genome. We define three distinct viral structural protein regions governing PS recognition: two separate domains of the nucleocapsid protein, as well as the endodomain of the membrane protein. We demonstrate that the nucleocapsid protein carboxy-terminal domain is the principal driver of packaging selectivity through specifically binding to PS RNA. This work draws attention to coronavirus protein domains and RNA structures that represent attractive targets for the development of antiviral therapeutics.

Keywords: RNA virus genome packaging, packaging signal, coronavirus, nucleocapsid protein, membrane protein

Abstract

Most cytoplasmic RNA viruses have evolved mechanisms to identify their genomic RNA (gRNA) as the sole RNA species to be packaged into assembled virions. Coronaviruses exhibit highly selective packaging of their gRNA, in spite of the presence of a large excess of subgenomic viral and host RNA in infected cells. Failure to accomplish this selectivity does not impair virion assembly but renders the virus vulnerable to host innate immunity. In the prototype coronavirus mouse hepatitis virus (MHV) packaging selectivity is brought about by a packaging signal (PS), an RNA structure found exclusively in gRNA and absent from subgenomic RNAs. However, it is not well resolved how virion structural proteins participate in the recognition of the PS. Previous studies with virions have shown that the nucleocapsid (N) protein carboxy-terminal RNA-binding domain (CTD) and the carboxy-terminal tail (domain N3) both have roles in PS recognition. Separately, work with a virus-like particle system has implicated the viral membrane (M) protein as the main factor in this process. Here, we pinpoint key amino acids in the CTD and domain N3 that govern selective gRNA packaging into virions, and we provide genetic evidence that the M protein endodomain also plays a required role. We show that the N protein CTD is the primary determinant of PS recognition, and we localize residues in the CTD structure that are critical for specific and nonspecific RNA binding.


Genome packaging represents a fundamental problem to be addressed by most viruses in the later stages of their cellular replication cycle. In particular, cytoplasmic RNA viruses must somehow distinguish their own genetic material, amid a vast array of other viral and host RNA species, as the correct substrate for incorporation into assembling virions. Many RNA virus families have evolved means of accomplishing this selectivity through the recognition of unique sequences or structures contained within the viral genome. Coronaviruses are a family of enveloped viruses with exceptionally large positive-strand RNA genomes. Among the genera in this family that infect mammalian hosts, the Betacoronaviruses have received heightened attention because they include three highly virulent human pathogens: severe acute respiratory syndrome coronavirus (SARS-CoV), which emerged in the SARS outbreak of 2002–2003; Middle East respiratory syndrome coronavirus (MERS-CoV), the etiologic agent of a frequently fatal disease that is confined to the Arabian peninsula; and SARS-CoV-2, the cause of the COVID-19 pandemic (1). Two other endemic human coronaviruses (HCoVs), HCoV-OC43 and HCoV-HKU1, also fall within the Betacoronavirus genus and are generally associated with milder upper respiratory tract infections.

Coronavirus infection entails the synthesis of numerous RNA species (Fig. 1A). Genomic RNA (gRNA), serving as the first mRNA of infection, is translated via ribosomal frameshifting into two huge polyproteins that are autoproteolytically processed into 16 nonstructural proteins (nsp1-nsp16), which constitute the viral replicase–transcriptase complex. This RNA-synthetic machinery, proceeding through negative-strand RNA intermediates, both replicates gRNA and transcribes a 3′-nested set of subgenomic RNAs (sgRNAs) that are the mRNAs for all of the downstream-encoded genes. Despite the abundant production of sgRNA, only gRNA ultimately becomes packaged into virions. For the prototype coronavirus mouse hepatitis virus (MHV) this occurs by means of a packaging signal (PS) embedded in the coding region for nsp15 (2). The MHV PS was initially localized through comparison of packaged and unpackaged defective interfering RNAs, which are extensively deleted gRNA derivatives that replicate at the expense of helper virus (35). Subsequently, a model of the MHV PS (Fig. 1A), supported by in vitro structural probing and phylogenetic analysis, presented it as a 95-nt bulged stem containing four repeated modules centered around an internal loop (6). Significantly, this particular PS is found only in members of the Embecovirus subgenus of the Betacoronaviruses (which includes MHV, HCoV-OC43, and HCoV-HKU1). Members of the Sarbecovirus subgenus (which includes SARS-CoV and SARS-CoV-2) possess neither a homolog of the MHV PS nor the nsp15 surface loop that it encodes (7, 8). Thus, Sarbecoviruses must harbor a different PS at a different genomic site (9, 10).

Fig. 1.

Fig. 1.

Elements of MHV gRNA packaging. (A) Top, coronavirus RNA species. The 5′ end of the MHV genome (gRNA) comprises the genes rep1a and rep1b, the protein products of which are processed into the components of the viral replicase–transcriptase complex (nsp1–nsp16). A set of subgenomic RNAs (sgRNA2–sgRNA7), transcribed during infection, serve as mRNAs for the downstream-encoded viral structural proteins, spike (S), envelope (E), membrane (M), and nucleocapsid (N) proteins, as well as various accessory proteins [NS2a, HE (hemagglutinin-esterase), NS4, NS5a]. Left, structure of the wild-type MHV gRNA PS (wt PS), which is embedded in the coding region for nsp15 (6). Genomic nucleotide coordinates are numbered; the four repeat units are boxed. The PS structure was disrupted by the creation of 20 coding-silent mutations (nucleotides highlighted in green) to generate the silPS mutant (11). (B) Northern blot of RNA isolated from highly purified virions of wild-type MHV and the silPS mutant; also shown is Alb424, a mutant in which the SARS-CoV 3′ UTR was substituted for that of MHV (12). Viral RNA was detected by a probe corresponding to the N gene, which is contained in gRNA and all sgRNAs. (C) Schematic of wild-type MHV N protein showing the amino-terminal RNA-binding domain (NTD or N1b), the carboxy-terminal RNA-binding domain (CTD or N2b), and domain N3, the carboxy terminus of the molecule. The CTD is also the N dimerization domain. Flanking these domains are three unstructured regions, N1a, N2a, and spacer B; numbering indicates amino-acid residues. Below the wild type are two previously characterized N protein mutants that are packaging-defective despite having a wild-type PS. S2b contains the SARS-CoV N protein CTD (blue) replacing that of MHV (13); CCA5 harbors a set of four clustered charged-to-alanine mutations (red) in the distal portion of domain N3 (14).

We previously investigated the authentic role of the PS in the viral genome by the construction of an MHV mutant, named silPS, containing 20 coding-silent mutations that totally disrupted PS structure while leaving the nsp15 protein unaltered (11). This completely abolished packaging selectivity: the silPS mutant incorporated large amounts of all sgRNAs, as well as gRNA, into purified virions (Fig. 1B). Surprisingly, however, the silPS mutant was fully competent to assemble into virions. It had the same plaque size and growth kinetics as the wild type, and it was only marginally less fit than the wild type, as shown by growth competition. Other types of PS manipulation, by mutation or total deletion, generated the same phenotype as silPS (11, 15). This revealed that for coronaviruses there is an important distinction between packaging and assembly, two terms that are often used interchangeably. The expendable nature of the PS in tissue culture was somewhat paradoxical until it was demonstrated that inclusion of the silPS mutations into a highly neurovirulent strain of MHV dramatically attenuated viral infection in the mouse host (16). It was further shown that this attenuation depended upon type-I interferon signaling, suggesting that innate immunity is triggered by a pathogen-associated molecular pattern in sgRNA that is absent or hidden in gRNA.

How viral proteins participate in selective genome packaging is not well resolved. The most salient candidate for a role in PS recognition is the viral nucleocapsid (N) protein, which wraps the viral genome into a “beads-on-a-string” arrangement within the virion (1719). The N protein has two separate structural domains that bind to RNA [designated the amino-terminal domain (NTD) and the carboxy-terminal domain (CTD)] and a domain at the actual carboxy terminus (N3), which interacts with the viral membrane protein (M) (20) (Fig. 1C). The CTD also functions as the dimerization domain of the N molecule. The NTD, CTD, and N3 are flanked by three unstructured regions. We previously constructed mutants substituting the SARS-CoV NTD or CTD in the MHV N protein, based on the assumption that SARS-CoV N protein would not have evolved to recognize the heterologous MHV PS (13). One of these, the SARS-CoV CTD substitution mutant, had the same packaging-defective phenotype as the silPS mutant, even though it maintained a wild-type PS. Remarkably, the NTD substitution mutant was not affected in selective gRNA packaging. In a subsequent study we found that, distinct from the CTD, the domain N3 is also critical for PS recognition, independent of its role in assembly with the M protein. Certain N3 mutants that were fully assembly-competent were found to exhibit a phenotype identical to that of silPS, despite having a wild-type PS (14). These results pointed to two N domains controlling the packaging process (Fig. 1C). However, there also exists strong evidence that M protein plays a part in packaging. Immunoprecipitation studies showed that in infected cell lysates N protein was bound to all sgRNAs as well as to gRNA, but anti-M antibody coimmunoprecipitated only the fraction of N protein that was bound to gRNA (21, 22). Moreover, work with a virus-like particle system indicated that M protein, in the absence of N, could bring about the selective packaging of an expressed nonviral RNA only if it contained the PS (23).

In the present study, we pinpoint critical amino acids in the CTD and domain N3 of N protein that govern selective gRNA packaging into virions. Additionally, we provide genetic evidence for a role for M protein as a necessary participant in the packaging process. Finally, we show by in vitro RNA-binding assays that the N protein CTD is the primary determinant of PS recognition, and we localize residues in the CTD structure that participate in specific and nonspecific RNA binding.

Results

MHV Virions Package Cellular 7SL RNA in Addition to Viral RNA.

To assay gRNA packaging, we first stringently purified virions by a procedure that included two cycles of equilibrium centrifugation on glycerol-tartrate gradients (SI Appendix, Materials and Methods). Virions were normalized by protein concentration and evaluated by SDS-PAGE, as shown for overloaded amounts of wild-type MHV and the previously characterized packaging-defective N3 mutant CCA5 (SI Appendix, Fig. S1A). This established that our virion preparations were highly pure, predominantly containing the three major structural proteins N, M, and spike (S), as well as minor amounts of the envelope (E) protein and the I protein, which is translated from the +1 reading frame of the N gene (24, 25). Specifically, we did not detect contaminating cellular proteins that might have had the potential to bind RNA. It should be noted that for wild-type and mutant MHV, M protein bands in SDS-PAGE appear heterogeneous due to varying levels of oligosaccharide processing of the O-glycosylated ectodomain (26); the N-glycosylated S protein likewise exhibits a degree of heterogeneity.

Northern blot analysis remains the definitive method to determine the RNA content of purified coronavirus virions (10, 11). However, this technique is cumbersome and can be hampered by variability of transfer of the exceptionally large coronavirus gRNA from gels to filters. To have a complementary assay, we used native agarose gel electrophoresis in the presence of ethidium bromide for further examination of RNA species that are packaged by MHV virions (SI Appendix, Fig. S1B). RNA size markers in these gels were provided by purified total RNA from infected and mock-infected cells, tRNA, and synthesized sgRNA7. We also included a set of DNA markers as a standard reference, although the sizes of dsDNA species cannot be compared to those of ssRNA. Native gels allowed the direct visualization of gRNA and could be used to more quickly screen constructed mutants. As expected, we observed that purified virions of wild-type MHV and the CCA5 mutant contained equivalent amounts of gRNA. As well, the more abundant sgRNA species, sgRNA5–sgRNA7, were readily detected in CCA5 virions but not in wild-type virions. In contrast to the discrete bands seen for sgRNAs and cellular rRNAs, under nondenaturing conditions gRNA migrated as a broad band, likely due to the extensive ensemble of secondary structures this ~31 kb RNA can assume (27, 28). We cannot rule out that some degradation of gRNA occurred, but we think this unlikely because the sgRNA bands were sharply defined in native gels, and we did not observe degradation of gRNA or sgRNA in Northern blots.

Unexpectedly, we noticed that both wild-type and CCA5 virions packaged a small cellular RNA. We tentatively identified this as 7SL RNA because it corresponded to a prominent small RNA in uninfected and infected cells (SI Appendix, Fig. S1B) and because 7SL RNA is the most abundant nonribosomal RNA in eukaryotic cells (29). To confirm the identity of this packaged species, and to clearly differentiate it from the similarly sized 7SK RNA, RNA isolated from purified virions was reverse-transcribed (RT) with a random primer, and the resulting cDNA was amplified by PCR with diagnostic primers for 7SL RNA or 7SK RNA (SI Appendix, Fig. S1C). Control RNA from mock-infected cells yielded amplicons of the predicted sizes for both targets (299 bp for 7SL RNA, 325 bp for 7SK RNA). By contrast, only a robust product corresponding to 7SL RNA was obtained with wild-type and CCA5 virion RNA. For CCA5 virion RNA a faint 7SK band was also seen, likely because there is general low-level incorporation of cellular RNA by packaging-defective mutants (16). The identities of all RT-PCR amplicons were verified by DNA sequencing. This result showed that, despite its stringent packaging specificity, MHV incorporates a substantial amount of 7SL RNA into virions, similar to HIV and all other retroviruses that have been examined (30, 31). Notably, however, retroviruses also package both primer and other tRNAs, whereas we did not detect tRNA in MHV virions.

Mutational Analysis of the Role of the N Protein CTD in PS Recognition.

To more precisely define the components that govern selective packaging of gRNA we carried out mutational analyses of the relevant MHV structural protein domains. In earlier work, we had discovered that the CTD of N protein is a major determinant in the packaging process through the creation of chimeric viruses in which either of the two RNA-binding domains of the MHV N protein was replaced by its counterpart from the SARS-CoV N protein (13). Substitution of the SARS-CoV N CTD, but not the NTD, totally eliminated packaging competence, despite the presence of a fully intact PS. To explore whether a minimal set of point mutations could decouple specific and nonspecific RNA binding by the CTD, we adopted two different strategies. In the first, we replaced single residues or pairs of residues of the MHV N CTD by the corresponding aligned amino acids from the SARS-CoV N CTD. We focused on MHV N residues 295–328, based on work with previously generated partial chimeras, which indicated that critical differences between the MHV and SARS-CoV CTDs might map to this region (13). By this scheme, mutants Sm1–Sm5 were produced (Fig. 2A). These and all other mutants described in the present study were generated by the reverse-genetic method of targeted RNA recombination (32, 33). At least two isolates of each mutant were obtained, and the sequences of the entire N gene and the upstream region encoding the M protein endodomain were determined to confirm the presence of the expected mutations and the absence of extraneous or compensating mutations. Additionally, the central portion of the nsp15 coding region of each characterized mutant was sequenced to verify that the wild-type PS was intact. All five constructed viruses exhibited growth comparable to the wild type. However, when RNA from purified virions was analyzed, mutants Sm1–Sm5 were each found to have a wild-type packaging phenotype or showed only marginally less specificity for packaging gRNA than did wild-type MHV. These results suggested that it is unlikely there is any crucial individual SARS-CoV N CTD residue (or pair of residues) essential to maintaining or abolishing packaging competence when substituted in the MHV N CTD.

Fig. 2.

Fig. 2.

Effect of CTD mutations on selective gRNA packaging. (A) Schematic of N protein, with an expansion showing the wild-type sequence of the MHV CTD and the locations of mutations constructed by two separate strategies. Residues highlighted in green were mutated to their aligned SARS-CoV CTD counterparts to create mutants Sm1–Sm5. Residues in blue are basic amino acids that were changed to alanine to create mutants NB1–NB8; red circles indicate sets of NB mutations that were lethal. (B) SDS-PAGE analysis of purified virions of wild-type MHV and mutants CCA5, NB4, NB6, and NB8 (5 µg protein each). Virion structural proteins indicated are S0, uncleaved spike protein; S1 and S2, cleaved spike protein; N, nucleocapsid protein; M, membrane protein. Std, broad-range protein standard (New England Biolabs). (C) Northern blots of RNA isolated from purified virions of wild-type MHV and mutants CCA5, NB4, NB6, and NB8 detected with a probe corresponding to the 5′ half of the N gene; gRNA and the more prominent sgRNAs are indicated. Of note, the CCA5 mutant produces (and packages) less sgRNA4 because, unlike all the other viruses in the present study, it lacks an engineered mutation upstream of the gene 4 transcription regulatory signal that upregulates synthesis of sgRNA4 (32). (D) Native agarose gels of RNA isolated from purified virions of wild-type MHV and mutants CCA5, NB4, NB6, and NB8; gRNA and the more prominent sgRNAs are indicated. RNA standards for comparison were in vitro synthesized sgRNA7 and mock-infected cellular RNA; a DNA standard (1 kb Plus DNA ladder; New England Biolabs) was included for reference. In panels B and D, duplicate samples shown for wild-type and CCA5 are from two independent virion purifications of each.

A second, more productive strategy prioritized mutation of surface clusters of basic amino acids, based on their positions in the aligned structure of the SARS-CoV N CTD (PDB ID 2JW8) (34). We hypothesized that changing such residues to alanines might disrupt interactions with RNA or else with acidic residues in domain N3. Our expectation was that groups of lysines and arginines that were conserved between the MHV and SARS-CoV N CTDs were likely to be required for general nonspecific RNA binding, and their alteration would result in temperature-sensitivity or lethality. Conversely, lysine and arginine clusters that were found only in MHV and other Embecoviruses might be involved in specific recognition of the PS, but their elimination would not affect virus assembly or growth. Eight mutants, NB1–NB8 (Fig. 2A), were constructed according to this plan and sorted into three classes. Mutants NB2, NB3, and NB5 were deemed lethal, as they were not obtained after five independent targeted recombination trials, all of which had strong wild-type controls. Mutants NB1 and NB7 were very defective, producing small plaques and displaying slow growth. Although we could select faster-growing variants of NB1 and NB7 containing second-site CTD mutations, these pseudorevertants were still markedly defective compared to the wild type and were thus not pursued further.

By contrast, the three remaining mutants, NB4 (K257A, K261A), NB6 (K279A, K303A), and NB8 (K345A, K370A), all exhibited growth yields as robust as the wild type, indicating they were fully assembly-competent. They were therefore chosen for more detailed analysis. Virions of the NB4, NB6, NB8 mutants were isolated, along with the wild type and the domain N3 mutant CCA5. Equal amounts of virions, normalized by protein concentration, were evaluated by SDS-PAGE and determined to be highly pure (Fig. 2B). RNA isolated from equal amounts of virions of the CTD mutants, wild type, and CCA5 was analyzed by Northern blot (Fig. 2C) and by native agarose gel electrophoresis (Fig. 2D), which revealed that all of these constructed CTD mutants were impaired in RNA packaging selectivity to the same extent as CCA5. Each of them packaged substantial quantities of sgRNAs in proportion to the relative abundance of these species in infected cells. These results point to a distinct set of positively charged residues in the MHV N CTD that are critical for distinguishing gRNA from all other available RNA species.

Mutational Analysis of the Role of N Protein Domain N3 in PS Recognition.

We previously discovered that N3, the 46-residue carboxy terminus of N protein, plays a role in gRNA packaging that is independent from its role in virion assembly. This was determined through screening a set of N3 mutants that had been created for other purposes (14). Based on comparison of those mutants that were packaging-competent versus those that were packaging-defective, the critical region of domain N3 was mapped to a short segment adjacent to the carboxy terminus (SI Appendix, Fig. S2A). The clustered charged-to-alanine mutant CCA5 (D446A, E449A, D450A, D451A) was found to be the minimally altered packaging-defective virus among this set. To learn if the lesion in CCA5 could be localized further, we constructed two viruses, C1 (D446A) and C2 (E449A, D450A), each harboring a subset of the CCA5 mutations. The C1 and C2 mutants, like CCA5, exhibited a wild-type growth phenotype, and we produced highly purified virion preparations of the two, as well as of CCA5 and a wild-type control (SI Appendix, Fig. S2B). As observed previously (14), the N protein of CCA5 migrated slightly faster than wild-type N protein. Analysis by Northern blot (SI Appendix, Fig. S2C) and native agarose gel electrophoresis (SI Appendix, Fig. S2D) of RNA purified from virions demonstrated that the C1 and C2 mutants each had a wild-type packaging phenotype, in contrast to CCA5. These results suggest that there is no key acidic amino acid among the four that are mutated in CCA5. To abolish packaging competence it appears that multiple negative charges within a small interval of N3 must be ablated.

We additionally constructed two N3 mutants containing mutations in a pair of amino acids that are targets of phosphorylation (35), which would add to the distribution of negative charges in the N3 domain. Notably, these two residues, S424 and T428, are conserved in all Embecovirus N proteins, suggesting a possible involvement in the temporal coordination of packaging during infection. In the first mutant, C3, phosphorylation was knocked out by alanine substitution; in the second, C4, phosphorylation was mimicked by substitution of serine and threonine with aspartate and glutamate, respectively (SI Appendix, Fig. S2A). Neither set of mutations impaired viral growth. Purified virions of both these mutants (SI Appendix, Fig. S2B) showed no alteration in packaging specificity (SI Appendix, Fig. S2 C and D), indicating that phosphorylation of domain N3 does not affect this function of N protein. To reinforce these findings, we confirmed that the C1–C4 and CCA5 mutants all contained the unaltered wild-type PS sequence in the nsp15 coding region of their gRNA.

Genetic Evidence for a Role for the M Protein in PS Recognition.

Although our work to this point appeared to solely implicate the N protein in recognition of the PS, a series of compelling studies from the Makino laboratory has suggested that the M protein is the principal MHV component that interacts with the PS (2123). This prompted us to attempt a genetic search that would either support or rule out this possibility. The coronavirus M protein is a dimer, with each monomer anchored in the virion envelope by three transmembrane domains. The greater part of the molecule, its carboxy-terminal endodomain, projects into the interior of the virion (3638), where it contacts domain N3 of the N protein (Fig. 3A). Most M protein interactions with the nucleocapsid have been localized to the distal portion of the endodomain (3941).

Fig. 3.

Fig. 3.

Effect of M protein endodomain mutations on selective gRNA packaging. (A) Schematics of the MHV N and M proteins; shown for M are the amino-terminal ectodomain (ecto), three transmembrane domains (Tm1–Tm3), and the carboxy-terminal endodomain (endo). The region of N protein domain N3 highlighted in red marks the major determinant of interaction with the M protein in virion assembly (42, 43); the adjacent region highlighted in green is the locus of the CCA5 mutations that abolish selective gRNA packaging. The region of M highlighted in red marks the principal segment of the M endodomain that interacts with N3 in virion assembly (40, 41). The expansion shows the wild-type M endodomain sequence and the locations of mutations constructed by two separate strategies. MC1–MC5 were designed as clustered charged-to-alanine mutants; red circles indicate sets of mutations that were lethal. Residues highlighted in blue are basic amino acids unique to Embecoviruses that were individually changed to alanine to create mutants MT1–MT4. (B) SDS-PAGE analysis of purified virions of wild-type MHV and mutants MT4C and MT2 (5 µg protein each); purified virions of MT4 are shown in Fig. 2B. Virion proteins and molecular mass standards are indicated as in Fig. 2B. (C) Northern blots of RNA isolated from purified virions of wild-type MHV and mutants CCA5, NB4, MT2, MT4, and MT4C, an independently constructed isolate of the MT4 mutant; CTD mutant Sm4 was included as an additional wild-type-like control. Viral RNA was detected with a probe corresponding to the 5′ half of the N gene; gRNA and the more prominent sgRNAs are indicated. (D) Native agarose gels of RNA isolated from purified virions of wild-type MHV and mutants MT2, MT4, and MT4C; gRNA and the more prominent sgRNAs are indicated. RNA and DNA standards are as in Fig. 2D. (E) Antiparallel association of the carboxy termini of the N and M proteins suggested by the separate locations of amino acid residues critical for assembly interactions (shown in red) (4043) or packaging interactions (shown in green) (14).

As a first approach to manipulating the M endodomain we carried out clustered charged-to-alanine mutagenesis, in which two or more charged amino acids appearing within a sliding window of five residues were all replaced with alanine (44). However, five mutants constructed according to this scheme, MC1–MC5 (Fig. 3A), uniformly turned out to be lethal. This outcome was consistent with previous findings that the M protein is much less tolerant to mutational changes than the N protein (45, 46). We therefore next adopted a strategy similar to that taken with the N protein CTD. In the design of mutants MT1–MT4 (Fig. 3A) we targeted only single basic residues in the M endodomain that are conserved among the Embecoviruses but are not present in the other four subgenera of the Betacoronaviruses. All four of these mutants were viable. One of them, MT3, had severely impaired growth and consequently was not examined further. On the other hand, mutants MT1, MT2, and MT4 were readily isolated and displayed robust growth phenotypes, equivalent to that of the wild type. Because its mutation fell closest to the carboxy terminus of the endodomain, mutant MT4 was chosen for more complete analysis.

The protein composition of highly purified virions of MT4 and other M mutants was indistinguishable from that of the wild type (Figs. 2B and 3B). This, together with the strong growth of these viruses, established that their particular mutations did not cause any detectable defects in virion assembly. Analysis of MT4 virion RNA by Northern blot (Fig. 3C) and native agarose gel electrophoresis (Fig. 3D) revealed that this mutant was impaired in selective gRNA packaging to an equal degree as the N protein mutants NB4 and CCA5. Since it was surprising that this lesion could be caused by merely a single point mutation, we constructed an entirely independent isolate, designated MT4C, harboring the same K217A mutation. Purified virions of mutant MT4C exhibited the same RNA packaging defect as MT4 (Fig. 3 C and D). Contrary to this, analysis of a different mutant, MT2, showed that such a defect was not a general property of M endodomain mutants. As with all other mutants in this study, we sequenced MT4, MT4C, and MT2 to verify that, other than for the engineered mutations, the entire M and N genes and the region of nsp15 encompassing the PS were identical to the wild type. These results clearly show that the carboxy terminus of M protein participates in recognition of the PS, and this function is distinct and separable from its role in virion assembly. The key N–M interaction essential for MHV virion assembly was previously mapped to R227 in the tail of the M protein (40, 41) and D440–D441 in domain N3 (42, 43). An analogous interaction, also between a single basic residue in the M tail and multiple acidic residues in N3, now appears to be essential for selective packaging of the genome (Fig. 3E). Moreover, the spacing of the two types of interaction suggests that the carboxy termini of the two proteins align in an antiparallel configuration within virions.

Collectively, our mutational analyses therefore revealed that three individual domains found in the two most abundant coronavirus structural proteins – the CTD and N3 domains of N protein and the endodomain of M protein—play critical roles in selective packaging of gRNA. Small localized sets of mutations, or even a single mutation, can abolish the ability of these proteins to participate in recognition of the wild-type PS. As a further check on the integrity of members of each mutant class, we confirmed that during infection the NB4, CCA5, and MT4 viruses synthesized the same canonical set of sgRNAs in the same relative amounts as wild-type MHV (Fig. 4A). Additionally, each mutant aberrantly packaged sgRNAs in proportion to the frequencies of these species in infected cells. Mutant NB4 was found to produce a slightly smaller sgRNA2 owing to a 215-nt deletion in HE; since HE is a pseudogene in MHV-A59 this was deemed to not be relevant. We also measured single-step growth for this set of mutants in comparison to that of the wild type (Fig. 4B). All of the viruses displayed identical kinetics during the initial phase of growth, although maximal titers of NB4 reached a plateau some four- to fivefold below those attained by CCA5, MT4, and the wild type. This showed that the packaging-negative phenotype of the mutants cannot be due to a general instability of the N or M protein but, rather, it results from a specific loss of the ability to discriminate between unique properties of the genomic PS and more general features of RNA.

Fig. 4.

Fig. 4.

RNA synthesis and viral replication of N and M protein mutants. (A) Northern blots: Left, total RNA isolated from 17Cl1 cells that were mock-infected or infected with wild-type MHV or mutants CCA5, NB4, Sm4, or MT4; Right, RNA isolated from highly purified virions of the same set of viruses. Viral RNA was detected with probes corresponding to the 3′ half of the N gene plus the 3′ UTR (infected cells) or the 5′ half of the N gene (virions); gRNA and all sgRNAs are indicated. The virion Northern blot shown in the Right panel was generated in an experiment separate from the one shown in Fig. 3C. (B) Growth kinetics of PS recognition mutants in tissue culture. Confluent monolayers of 17Cl1 cells were infected with wild-type, CCA5, NB4, or MT4 viruses at a multiplicity of 5.0 PFU per cell. At the indicated times postinfection, aliquots of medium were removed and infectious titers were determined by plaque assay on mouse L2 cells. Each data point is the mean (±SD) of three biological replicates.

Selective RNA Binding by the N Protein CTD.

To acquire biochemical evidence in support of our in vivo findings, we generated His-tagged constructs of the wild-type N protein CTD as well as versions containing the mutations of the packaging-defective mutants NB4, NB6, and NB8 (SI Appendix, Fig. S3A). Wild-type and mutant CTD proteins were expressed in Escherichia coli, purified by affinity chromatography (SI Appendix, Fig. S3B), and used to perform electrophoretic mobility shift assays (EMSAs). The substrates for these RNA-binding assays were fluorescent-labeled 45-nt RNA oligomers (PS-45 and silPS-45) (SI Appendix, Fig. S3C) identical to the apical stem-loop of the wild-type genomic PS or that of the packaging-defective silPS mutant, respectively (Fig. 1A). Increasing concentrations of CTD protein were incubated with each probe and then separated by electrophoresis in TBE gels (SI Appendix, Fig. S3 D and E). CTD was seen to bind to either probe with an affinity in the micromolar range, similar to what had previously been observed with the CTDs of SARS-CoV and SARS-CoV-2 binding to short nonspecific RNA probes (47, 48). We also consistently noted that unbound PS-45 had a slightly higher electrophoretic mobility than unbound silPS-45, likely because the wild-type probe is more highly structured. However, no clear difference was detectable between CTD binding to PS-45 vs. silPS-45. This apparent lack of specificity was unchanged if binding reactions were incubated at 37 °C instead of room temperature or if gels were run at 4 °C. Moreover, each of the CTD mutants behaved in the same way as the wild type.

In marked contrast, when excess tRNA was included as a nonspecific competitor RNA in EMSAs, pronounced differences between CTD binding to the two substrates became evident. Under these conditions, increasing amounts of wild-type CTD formed two complexes with PS-45 (Fig. 5 A and E). The faster-migrating one, designated complex 1, was very sharply defined and predominated at lower CTD concentrations. The second, slower-migrating complex 2, was more broadly dispersed and did not become substantially populated until the highest concentration of CTD. At variance with this, binding of CTD to silPS-45 produced only faint amounts of a diffuse band in the position of complex 1, while most bound probe migrated as complex 2 (Fig. 5 B and F). To further investigate these results, the same competitive RNA-binding experiments were performed with the packaging-defective CTD mutants. Unlike the wild-type CTD, the binding of CTD-NB4 to PS-45 produced complex 2 as the major product and only small amounts of a dispersed complex 1 band (Fig. 5 C and G), a pattern highly similar to the binding of wild-type CTD to silPS-45. In the same manner, binding of CTD-NB4 to silPS-45 almost exclusively produced complex 2 (Fig. 5 D and H). EMSAs carried out with the mutants CTD-NB6 (SI Appendix, Fig. S4) and CTD-NB8 (SI Appendix, Fig. S5) yielded essentially identical results to those with CTD-NB4.

Fig. 5.

Fig. 5.

Competitive RNA binding by expressed wild-type N protein CTD and mutant CTD-NB4. (AD) Representative EMSAs of wild-type CTD (A and B) and mutant CTD-NB4 (C and D) binding to PS-45 or silPS-45 in the presence of 100-fold molar excess of unlabeled competitor tRNA. RNA was incubated with 2-fold increasing concentrations of protein, from 0.313 to 10.0 µM, and complexes were resolved by TBE gel electrophoresis. A reaction without tRNA is included as the last lane of each gel. The positions of unbound probe and complexes 1 and 2 are marked beside each gel. (EH) Quantification of RNA–protein complexes in EMSAs. In each lane, the amount of probe bound in complex 1 or complex 2 was normalized as the fraction of total bound probe for the lane. Each data point is the mean (±SD) of four (wild-type CTD) or three (mutant CTD) replicate experiments.

On this basis, we concluded that complex 1 represents packaging-specific RNA binding by the CTD. Notably, there was a direct correspondence between the RNA-binding and genetic results. Formation of complex 1, like selective gRNA packaging in virions, could be abolished by two independent routes—either by mutating the PS or else by mutating key residues in the CTD that do not diminish overall nonspecific RNA binding. This congruence strongly suggests that the CTD is the principal determinant of PS recognition.

Discussion

In the present study we have shown that genome packaging in the prototype coronavirus MHV is governed by four separate molecular components: the RNA PS that resides in the coding region of nsp15, both the CTD and domain N3 of the N protein, and the endodomain of the M protein. Appropriately targeted impairment of any one of these components abrogates the ability of the virus to selectively package gRNA, resulting in indiscriminate inclusion of abundant sgRNAs into assembled virions. Importantly, our work highlights that for coronaviruses there is a critical distinction between gRNA packaging and virion assembly. All of the mutants characterized in this and previous studies (11, 13, 14) showed no deficiencies in virion assembly even though they failed to exclusively package gRNA.

A surprising finding that emerged from our work was that considerable amounts of a highly transcribed RNA polymerase III product, 7SL RNA, are incorporated into both wild-type and mutant virions (Figs. 2 and 3 and SI Appendix, Figs. S1 and S2). 7SL RNA is the RNA component of the signal recognition particle (SRP), which guides cotranslational protein insertion into membranes. Multiple double-stranded substructures within 7SL RNA may be akin to the stem-loop structures throughout coronavirus gRNA shown to favor nonspecific binding by N protein (19). It is not yet known whether 7SL RNA incorporation is a universal property of coronaviruses, but it has also been reported in unpurified viral preparations of SARS-CoV-2 (49). Due to its 5′-triphosphate terminus, free 7SL RNA, not shielded by SRP proteins, is a substrate for the innate immune sensor RIG-I (50). Consequently, it is possible that N protein binding serves to sequester free 7SL RNA, thereby countering activation of host antiviral pathways.

Specific and nonspecific RNA binding by N protein are integral to virion packaging and assembly, respectively. We were previously able to attribute to the N protein CTD a major role in gRNA packaging (13), and this has now been brought into finer focus. The NB4, NB6, and NB8 mutants identify positively charged residues in the CTD that are key to specific RNA binding in PS recognition but negligibly affect nonspecific RNA binding, which is essential for assembly (Fig. 2). To complement these genetic results, we carried out RNA-binding assays that revealed the formation of a unique complex between wild-type CTD and PS RNA in the presence of excess competitor tRNA (Fig. 5). Notably, tRNA was an apt competitor RNA in these binding assays because it is not packaged into MHV virions (SI Appendix, Fig. S1B). We were thus able to demonstrate selective binding of an authentic coronavirus PS RNA by a purified N protein domain, building on earlier studies that explored PS RNA binding by N protein in lysates from MHV-infected cells (51, 52). The specific binding of SARS-CoV-2 N protein CTD to a candidate PS RNA substrate was recently shown (53), but the actual identity of the SARS-CoV-2 PS is still unresolved. While we have yet to determine the exact stoichiometry of the complexes seen in competitive EMSAs, our results potentially indicate that each CTD dimer harbors both a specific and a nonspecific RNA-binding site. The combination of these in vitro results with our genetic results strongly argues for the assignment of the N protein CTD as the primary factor in PS recognition.

The structure of the CTD offers further insight into distinctions between specific and nonspecific RNA binding. At the outset of this study, there were no available structures of the CTD of MHV or any other Embecovirus, but high-confidence structural predictions recently became accessible through the introduction of AlphaFold 3 (54). The modeled wild-type MHV N CTD structure is highly similar to those of other Betacoronaviruses, as well as of Alpha- and Gammacoronaviruses, despite the relatively low amino-acid sequence conservation of this domain across the coronavirus family. As described for the earliest determined N CTD structures (34, 55, 56), the MHV N CTD is a domain-swapped dimer having a rectangular slab shape (Fig. 6A). The two β-strands (β1 and β2) from each monomer are joined in a four-stranded antiparallel β-sheet that is stabilized by internal hydrogen bonding and by hydrophobic interactions with the longest α-helix (α6). The dimer slab has two faces, one formed by the β-sheet and the other consisting of the six α-helices. A striking difference in the MHV N CTD is that the loops between the β-strands in each monomer are much larger than those in the SARS-CoV N CTD (Fig. 6B). This disparity holds for the N CTDs of all Embecoviruses, compared to those of the other four subgenera of the Betacoronaviruses (SI Appendix, Figs. S6 and S7). For SARS-CoV-2, the β1-β2 hairpin loop has been shown by hydrogen–deuterium exchange to be the only flexible portion of the otherwise compact CTD structure (57). The substantially larger loop in MHV would be expected to have even greater flexibility, which may have consequences for binding to RNA or to other parts of the N protein. Significantly, residues in the β1–β2 hairpin loop are proposed to directly participate in RNA binding, as inferred from a cocrystal structure of the SARS-CoV-2 CTD with GTP (53).

Fig. 6.

Fig. 6.

Structure of the MHV CTD. (A) Model of the MHV N protein CTD generated with AlphaFold 3 (54). Ribbon diagrams show the α-helical and β-sheet faces of the dimer slab, with the individual monomers depicted in green and turquoise. The amino and carboxy termini and the β1–β2 hairpin loop are labeled on each monomer. Residues that were mutated in the packaging-defective assembly-competent mutants are shown in yellow for NB4 (K257, K261), orange for NB6 (K279, K303), and magenta for NB8 (K345, K370). (B) Structural alignment of the N protein CTDs of MHV, SARS-CoV-1 (PDB ID 2CJR) (56), and SARS-CoV-2 (PDB ID 6YUN) (48). (C) Electrostatic surface potential of the N protein CTDs of wild-type MHV and the mutants NB4, NB6, and NB8. Surfaces are colored by electrostatic potential from –4 kT/e (red) to +4 kT/e (blue).

As for all other CTD structures, the α face of the MHV CTD exhibits a markedly electropositive surface brought about by the clustering of lysines and arginines. This strongly basic groove is postulated to be the site of nonspecific RNA binding (20, 56). In the three CTD mutants that are packaging-defective but fully assembly-competent (NB4, NB6, and NB8), all of the mutated amino-acid residues are located on the edges of the CTD slab structure, peripheral to the basic groove, not within it (Fig. 6A). These mutations do not detectably alter the structure or the electrostatic surface potential of the CTD (Fig. 6C), suggesting that specific RNA binding is governed by residues outside of the region that accommodates general RNA binding. The key residues identified by the NB4, NB6, and NB8 mutants are all lysines, the side chains of which are capable of both electrostatic and nonpolar interactions with RNA. By contrast, the mutations in the lethally altered CTD mutants (NB2, NB3, and NB5) are located more centrally on the α face (SI Appendix, Fig. S8A). This suggests that these mutants were not viable due to impairment of general nonspecific RNA binding, although only for NB2 is there a dramatic disruption of folding and electrostatic surface potential (SI Appendix, Fig. S8B). However, the inferences drawn here from predicted N CTD structures must be considered tentative until actual structures of N protein complexed with PS RNA can be obtained.

While the preponderance of the evidence points to the N protein CTD as the principal driver of PS recognition, our genetic results make clear that there exist ancillary, but essential, roles for both domain N3 of the N protein and the carboxy terminus of the M protein endodomain. These findings indicate that the mechanism of specific packaging of gRNA is more intricate than merely an initial N–PS binding event serving to nucleate cooperative assembly of N throughout the remainder of the genome. Indeed, it now appears that initiation of general N–RNA binding must occur at multiple sites. Recent cryoelectron tomographic imaging of SARS-CoV-2 virions has overturned the long-held notion that coronavirus nucleocapsids have helical symmetry analogous to that of most negative-strand RNA viral nucleocapsids (17, 18). These studies model the nucleocapsid as a series of viral nucleoprotein (vRNP) complexes separated by loops of unbound RNA arranged in a beads-on-a-string configuration. Bolstered by detailed in vitro reconstitution analyses (19), there are now estimated to be roughly 35 to 40 vRNPs per virion, each containing 10 to 12 molecules of N, which are stabilized by a variety of N–N interactions.

In earlier ultrastructural studies of SARS-CoV and MHV virions, the contacts between the N and M proteins were visualized as thread-like connections (58, 59). Based on previous mapping of essential assembly interactions, these connections correspond to the carboxy termini of the N and M molecules (4043). We have now found that essential gRNA packaging interactions likewise map to nearby residues of the same regions. N protein domain N3 is largely disordered (20, 57), and the first coronavirus M protein structures reveal that the extreme carboxy terminus of M protein is also disordered (3638). The spatial alignment of the two separate types of interaction in these two unstructured segments (Fig. 3E) suggests that the roles of N3 and the M endodomain in packaging and assembly are somehow interrelated. Alternatively, the relationship suggested by Fig. 3E could be coincidental; N3–M packaging interactions may come into play at a different time than N3–M assembly interactions. Moreover, a further complication arises from the observation that tetramerization of N protein can be mediated by N3–N3 interactions between N dimers (57, 60). Taken together, our results are in accord with previous findings that M protein is a necessary participant in the packaging process (21, 22). However, they do not support M as the sole protein determinant of packaging fidelity (23). In particular, if M protein were capable of binding to the PS independently from N protein, then our CTD mutants would have no packaging defect. Nevertheless, the exact role of M remains enigmatic. PS mutants, such as silPS and others (11, 15), are fully competent in virion assembly and therefore must retain the ability to assemble into beads-on-a-string vRNPs. This begs the question of what is unique about the particular vRNP containing the PS and how its contact with M protein triggers the incorporation of the genomic nucleocapsid into budding virions, while excluding all other RNA species that can be bound by N protein. Experimental structural resolution of this M endodomain–vRNP complex would help to illuminate this issue.

Selective gRNA packaging is a characteristic of every coronavirus for which it has been carefully examined. However, such examination can be hindered, either because some coronaviruses are intractable to propagation in sufficient quantities for thorough virion purification or else because of the hazards involved in manipulations of large amounts of highly pathogenic coronaviruses. Notably, it was only recently that selective gRNA packaging was rigorously demonstrated for SARS-CoV-2 (10). MHV and other Embecoviruses remain the only coronaviruses for which there is a clearly defined PS (2). In the alphacoronavirus transmissible gastroenteritis virus the PS has been shown to fall within the 5′-most 598 nucleotides of the genome, but it could not be further delimited (61). For SARS-CoV-2 there are disparate PS candidates. Based on bioinformatic analyses, the PS of this and multiple other coronaviruses was proposed to consist of repeating stem-loop structures in the 5′ UTR of the genome (8), an assignment partially supported by the construction of an engineered defective interfering RNA (62). At variance with this, a virus-like particle system that comprehensively tiled RNA segments across the entire SARS-CoV-2 genome localized a possible PS to a 1.1-kb region encompassing parts of nsp15 and nsp16, some 21 kb from the 5′ end of the genome (9). Differing from both of these results, dissection of a defective interfering particle efficiently packaged by SARS-CoV-2 virions mapped the PS to a 1.4-kb genomic segment crossing from nsp12 into nsp13, 16 kb from the 5′ end (10).

The acute in vivo attenuation of MHV PS mutants makes it very likely that selective gRNA packaging is an essential property common to all coronaviruses (16). Further investigation of packaging in other members of this family is therefore warranted. There are several reasons to expect that the protein determinants of selective gRNA packaging defined in the present study will generalize to all coronaviruses, even though the identity and genomic location of the PS can vary. Coronaviruses share the same mechanism of RNA synthesis, and they preserve a common organization of the unique replicase–transcriptase machinery that produces viral gRNA and sgRNAs. Coronavirus structural proteins have a uniform architecture, and they maintain the same modes of protein–protein interaction in assembled virions. The protein determinants of selective gRNA packaging thus represent attractive targets for small-molecule antiviral compounds. Additionally, as RNA structures become increasingly recognized as druggable entities, the PSs of the most virulent coronaviruses could be developed for therapeutic intervention (63, 64).

Materials and Methods

Detailed and referenced descriptions of methods used in this study, including growth of cells and viruses, generation of MHV mutants, plasmid construction, virion purification, virion RNA purification and analysis, viral growth kinetics, protein expression and purification, EMSA, and protein structural prediction are provided in the SI Appendix. In brief, MHV mutants were created by targeted RNA recombination, as described previously (32, 33). Virions were purified by a procedure involving two cycles of equilibrium ultracentrifugation on preformed glycerol-tartrate gradients, as described previously (65). Virion RNA was analyzed by Northern blotting (11, 13, 14) and by native agarose gel electrophoresis. Expressed recombinant CTD proteins were purified using HisTrap-FF Ni2+ Sepharose columns (Cytiva).

Supplementary Material

Appendix 01 (PDF)

Acknowledgments

We thank the Wadsworth Center’s Applied Genomics Technology Core Facility for DNA sequencing and the Media and Tissue Culture Core Facility for media preparation. We are grateful to Jon Paczkowski for helpful discussions and advice. This work was supported by the Wadsworth Center and by NIH (National Institute of Allergy and Infectious Diseases) Grant R01 AI064603 to P.S.M.

Author contributions

J.D.P. and P.S.M. designed research; J.D.P., T.K.S., L.K., and P.S.M. performed research; J.D.P., T.K.S., and P.S.M. analyzed data; and J.D.P., L.K., and P.S.M. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

Data, Materials, and Software Availability

All data from this study are included in the article and the SI Appendix. Unique materials produced for this study are available from the corresponding author upon request.

Supporting Information

References

  • 1.V’kovski P., Kratzel A., Steiner S., Stalder H., Thiel V., Coronavirus biology and replication: Implications for SARS-CoV-2. Nat. Rev. Microbiol. 19, 155–170 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Masters P. S., Coronavirus genomic RNA packaging. Virology 537, 198–207 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Makino S., Yokomori K., Lai M. M., Analysis of efficiently packaged defective interfering RNAs of murine coronavirus: Localization of a possible RNA-packaging signal. J. Virol. 64, 6045–6053 (1990). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.van der Most R. G., Bredenbeek P. J., Spaan W. J., A domain at the 3′ end of the polymerase gene is essential for encapsidation of coronavirus defective interfering RNAs. J. Virol. 65, 3219–3226 (1991). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Fosmire J. A., Hwang K., Makino S., Identification and characterization of a coronavirus packaging signal. J. Virol. 66, 3522–3530 (1992). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chen S. C., et al. , New structure model for the packaging signal in the genome of group IIa coronaviruses. J. Virol. 81, 6771–6774 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Joseph J. S., et al. , Crystal structure of a monomeric form of severe acute respiratory syndrome coronavirus endonuclease nsp15 suggests a role for hexamerization as an allosteric switch. J. Virol. 81, 6700–6708 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Chen S. C., Olsthoorn R. C. L., Yu C. H., Structural phylogenetic analysis reveals lineage-specific RNA repetitive structural motifs in all coronaviruses and associated variations in SARS-CoV-2. Virus. Evol. 7, veab021 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Syed A. M., et al. , Rapid assessment of SARS-CoV-2-evolved variants using virus-like particles. Science 374, 1626–1632 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Terasaki K., Narayanan K., Makino S., Identification of a 1.4-kb-long sequence located in the nsp12 and nsp13 coding regions of SARS-CoV-2 genomic RNA that mediates efficient viral RNA packaging. J. Virol. 97, e0065923 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kuo L., Masters P. S., Functional analysis of the murine coronavirus genomic RNA packaging signal. J. Virol. 87, 5182–5192 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Goebel S. J., Taylor J., Masters P. S., The 3′ cis-acting genomic replication element of the severe acute respiratory syndrome coronavirus can function in the murine coronavirus genome. J. Virol. 78, 7846–7851 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Kuo L., Koetzner C. A., Hurst K. R., Masters P. S., Recognition of the murine coronavirus genomic RNA packaging signal depends on the second RNA-binding domain of the nucleocapsid protein. J. Virol. 88, 4451–4465 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kuo L., Koetzner C. A., Masters P. S., A key role for the carboxy-terminal tail of the murine coronavirus nucleocapsid protein in coordination of genome packaging. Virology 494, 100–107 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Athmer J., et al. , In situ tagged nsp15 reveals interactions with coronavirus replication/transcription complex-associated proteins. mBio 8, e02320-16 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Athmer J., et al. , Selective packaging in murine coronavirus promotes virulence by limiting type i interferon responses. mBio 9, e00272-18 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Yao H., et al. , Molecular architecture of the SARS-CoV-2 virus. Cell 183, 730–738.e13 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Klein S., et al. , SARS-CoV-2 structure and replication characterized by in situ cryo-electron tomography. Nat. Commun. 11, 5885 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Carlson C. R., et al. , Reconstitution of the SARS-CoV-2 ribonucleosome provides insights into genomic RNA packaging and regulation by phosphorylation. J. Biol. Chem. 298, 102560 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chang C. K., Hou M. H., Chang C. F., Hsiao C. D., Huang T. H., The SARS coronavirus nucleocapsid protein—Forms and functions. Antivir. Res. 103, 39–50 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Narayanan K., Maeda A., Maeda J., Makino S., Characterization of the coronavirus M protein and nucleocapsid interaction in infected cells. J. Virol. 74, 8127–8134 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Narayanan K., Makino S., Cooperation of an RNA packaging signal and a viral envelope protein in coronavirus RNA packaging. J. Virol. 75, 9059–9067 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Narayanan K., Chen C. J., Maeda J., Makino S., Nucleocapsid-independent specific viral RNA packaging via viral envelope protein and viral RNA signal. J. Virol. 77, 2922–2927 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Senanayake S. D., Hofmann M. A., Maki J. L., Brian D. A., The nucleocapsid protein gene of bovine coronavirus is bicistronic. J. Virol. 66, 5277–5283 (1992). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Fischer F., Peng D., Hingley S. T., Weiss S. R., Masters P. S., The internal open reading frame within the nucleocapsid gene of mouse hepatitis virus encodes a structural protein that is not essential for viral replication. J. Virol. 71, 996–1003 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.de Haan C. A., et al. , Structural requirements for O-glycosylation of the mouse hepatitis virus membrane protein. J. Biol. Chem. 273, 29905–29914 (1998). [DOI] [PubMed] [Google Scholar]
  • 27.Ziv O., et al. , The short- and long-range RNA-RNA interactome of SARS-CoV-2. Mol. Cell. 80, 1067–1077.e5 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Lan T. C. T., et al. , Secondary structural ensembles of the SARS-CoV-2 RNA genome in infected cells. Nat. Commun. 13, 1128 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Boivin V., et al. , Simultaneous sequencing of coding and noncoding RNA reveals a human transcriptome dominated by a small number of highly expressed noncoding genes. RNA 24, 950–965 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Keene S. E., Telesnitsky A., Cis-acting determinants of 7SL RNA packaging by HIV-1. J. Virol. 86, 7934–7942 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Eckwahl M. J., Telesnitsky A., Wolin S. L., Host RNA packaging by retroviruses: A newly synthesized story. mBio 7, e02025-15 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kuo L., Godeke G. J., Raamsman M. J., Masters P. S., Rottier P. J., Retargeting of coronavirus by substitution of the spike glycoprotein ectodomain: Crossing the host cell species barrier. J. Virol. 74, 1393–1406 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Goebel S. J., Hsue B., Dombrowski T. F., Masters P. S., Characterization of the RNA components of a putative molecular switch in the 3′ untranslated region of the murine coronavirus genome. J. Virol. 78, 669–682 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Takeda M., et al. , Solution structure of the C-terminal dimerization domain of SARS coronavirus nucleocapsid protein solved by the SAIL-NMR method. J. Mol. Biol. 380, 608–622 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.White T. C., Yi Z., Hogue B. G., Identification of mouse hepatitis coronavirus A59 nucleocapsid protein phosphorylation sites. Virus Res. 126, 139–148 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Dolan K. A., et al. , Structure of SARS-CoV-2 M protein in lipid nanodiscs. eLife 11, e81702 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Zhang Z., et al. , Structure of SARS-CoV-2 membrane protein essential for virus assembly. Nat. Commun. 13, 4399 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wang X., Yang Y., Sun Z., Zhou X., Crystal structure of the membrane (M) protein from a bat betacoronavirus. PNAS Nexus 2, pgad021 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Escors D., Ortego J., Laude H., Enjuanes L., The membrane M protein carboxy terminus binds to transmissible gastroenteritis coronavirus core and contributes to core stability. J. Virol. 75, 1312–1324 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Kuo L., Masters P. S., Genetic evidence for a structural interaction between the carboxy termini of the membrane and nucleocapsid proteins of mouse hepatitis virus. J. Virol. 76, 4987–4999 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Verma S., Lopez L. A., Bednar V., Hogue B. G., Importance of the penultimate positive charge in mouse hepatitis coronavirus A59 membrane protein. J. Virol. 81, 5339–5348 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Hurst K. R., et al. , A major determinant for membrane protein interaction localizes to the carboxy-terminal domain of the mouse coronavirus nucleocapsid protein. J. Virol. 79, 13285–13297 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Verma S., Bednar V., Blount A., Hogue B. G., Identification of functionally important negatively charged residues in the carboxy end of mouse hepatitis coronavirus A59 nucleocapsid protein. J. Virol. 80, 4344–4355 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wertman K. F., Drubin D. G., Botstein D., Systematic mutational analysis of the yeast ACT1 gene. Genetics 132, 337–350 (1992). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Arndt A. L., Larson B. J., Hogue B. G., A conserved domain in the coronavirus membrane protein tail is important for virus assembly. J. Virol. 84, 11418–11428 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Kuo L., Hurst-Hess K. R., Koetzner C. A., Masters P. S., Analyses of coronavirus assembly interactions with interspecies membrane and nucleocapsid protein chimeras. J. Virol. 90, 4357–4368 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Chang C. K., et al. , Multiple nucleic acid binding sites and intrinsic disorder of severe acute respiratory syndrome coronavirus nucleocapsid protein: Implications for ribonucleocapsid protein packaging. J. Virol. 83, 2255–2264 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Zinzula L., et al. , High-resolution structure and biophysical characterization of the nucleocapsid phosphoprotein dimerization domain from the Covid-19 severe acute respiratory syndrome coronavirus 2. Biochem. Biophys. Res. Commun. 538, 54–62 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Peña N., et al. , Profiling selective packaging of host RNA and viral RNA modification in SARS-CoV-2 viral preparations. Front. Cell Dev. Biol. 10, 768356 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Zhang Y., et al. , 5-methylcytosine (m5C) RNA modification controls the innate immune response to virus infection by regulating type I interferons. Proc. Natl. Acad. Sci. U.S.A. 119, e2123338119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Molenkamp R., Spaan W. J., Identification of a specific interaction between the coronavirus mouse hepatitis virus A59 nucleocapsid protein and packaging signal. Virology 239, 78–86 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Cologna R., Spagnolo J. F., Hogue B. G., Identification of nucleocapsid binding sites within coronavirus-defective genomes. Virology 277, 235–249 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Ciges-Tomas J. R., Franco M. L., Vilar M., Identification of a guanine-specific pocket in the protein N of SARS-CoV-2. Commun. Biol. 5, 711 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Abramson J., et al. , Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Jayaram H., et al. , X-ray structures of the N- and C-terminal domains of a coronavirus nucleocapsid protein: Implications for nucleocapsid formation. J. Virol. 80, 6612–6620 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Chen C. Y., et al. , Structure of the SARS coronavirus nucleocapsid protein RNA-binding dimerization domain suggests a mechanism for helical packaging of viral RNA. J. Mol. Biol. 368, 1075–1086 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ye Q., West A. M. V., Silletti S., Corbett K. D., Architecture and self-assembly of the SARS-CoV-2 nucleocapsid protein. Protein Sci. 29, 1890–1901 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Bárcena M., et al. , Cryo-electron tomography of mouse hepatitis virus: Insights into the structure of the coronavirion. Proc. Natl. Acad. Sci. U.S.A. 106, 582–587 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Neuman B. W., et al. , Supramolecular architecture of severe acute respiratory syndrome coronavirus revealed by electron cryomicroscopy. J. Virol. 80, 7918–7928 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Lo Y. S., et al. , Oligomerization of the carboxyl terminal domain of the human coronavirus 229E nucleocapsid protein. FEBS Lett. 587, 120–127 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Morales L., et al. , Transmissible gastroenteritis coronavirus genome packaging signal is located at the 5′ end of the genome and promotes viral RNA incorporation into virions in a replication-independent process. J. Virol. 87, 11579–11590 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Chaturvedi S., et al. , Identification of a therapeutic interfering particle—A single-dose SARS-CoV-2 antiviral intervention with a high barrier to resistance. Cell 184, 6022–6036.e18 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Disney M. D., The druggable transcriptome project: From chemical probes to precision medicines. Biochemistry 64, 1647–1661 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Veenbaas S. D., Koehn J. T., Irving P. S., Lama N. N., Weeks K. M., Ligand-binding pockets in RNA and where to find them. Proc. Natl. Acad. Sci. U.S.A. 122, e2422346122 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Ye R., Montalto-Morrison C., Masters P. S., Genetic analysis of determinants for spike glycoprotein assembly into murine coronavirus virions: Distinct roles for charge-rich and cysteine-rich regions of the endodomain. J. Virol. 78, 9904–9917 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Data Availability Statement

All data from this study are included in the article and the SI Appendix. Unique materials produced for this study are available from the corresponding author upon request.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES