Skip to main content
mSphere logoLink to mSphere
. 2025 Jan 23;10(2):e00709-24. doi: 10.1128/msphere.00709-24

PRRSV-2 variant classification: a dynamic nomenclature for enhanced monitoring and surveillance

Kimberly VanderWaal 1,✉, Nakarin Pamornchainavakul 1, Mariana Kikuti 1, Jianqiang Zhang 2, Michael Zeller 2, Giovani Trevisan 2, Stephanie Rossow 3, Mark Schwartz 1, Daniel C L Linhares 2, Derald J Holtkamp 2, João Paulo Herrera da Silva 1, Cesar A Corzo 1, Julia P Baker 1, Tavis K Anderson 4, Dennis N Makau 1,2, Igor A D Paploski 1
Editor: Michael J Imperiale5
PMCID: PMC11852939  PMID: 39846734

ABSTRACT

Existing genetic classification systems for porcine reproductive and respiratory syndrome virus type 2 (PRRSV-2), such as restriction fragment length polymorphisms and sub-lineages, are unreliable indicators of close genetic relatedness or lack sufficient resolution for epidemiological monitoring routinely conducted by veterinarians. Here, we outline a fine-scale classification system for PRRSV-2 genetic variants in the United States. Based on >25,000 U.S. open reading frame 5 (ORF5) sequences, sub-lineages were divided into genetic variants using a clustering algorithm. Through classifying new sequences every 3 months and systematically identifying new variants across 8 years, we demonstrated that prospective implementation of the variant classification system produced robust, reproducible results across time and can dynamically accommodate new genetic diversity arising from virus evolution. From 2015 to 2023, 118 variants were identified, with ~48 active variants per year, of which 26 were common (detected >50 times). Mean within-variant genetic distance was 2.4% (max: 4.8%). The mean distance to the closest related variant was 4.9%. A routinely updated webtool (https://stemma.shinyapps.io/PRRSLoom-variants/) was developed and is publicly available for end users to assign newly generated sequences to a variant ID. This classification system relies on U.S. sequences from 2015 onward; further efforts are required to extend this system to older or international sequences. Finally, we demonstrate how variant classification can better discriminate between previous and new strains on a farm, determine possible sources of new introductions into a farm/system, and track emerging variants regionally. Adoption of this classification system will enhance PRRSV-2 epidemiological monitoring, research, and communication, and improve industry responses to emerging genetic variants.

IMPORTANCE

The development and implementation of a fine-scale classification system for PRRSV-2 genetic variants represent a significant advancement for monitoring PRRSV-2 occurrence in the swine industry. Based on systematically applied criteria for variant identification using national-scale sequence data, this system addresses the shortcomings of existing classification methods by offering higher resolution and adaptability to capture emerging variants. This system provides a stable and reproducible method for classifying PRRSV-2 variants, facilitated by a freely available and regularly updated webtool for use by veterinarians and diagnostic labs. Although currently based on U.S. PRRSV-2 ORF5 sequences, this system can be expanded to include sequences from other countries, paving the way for a standardized global classification system. By enabling accurate and improved discrimination of PRRSV-2 genetic variants, this classification system significantly enhances the ability to monitor, research, and respond to PRRSV-2 outbreaks, ultimately supporting better management and control strategies in the swine industry.

KEYWORDS: PRRS, molecular epidemiology, evolution, machine learning, nomenclature, lineages, clades, clustering, genetic variants

INTRODUCTION

In the United States, porcine reproductive and respiratory syndrome virus type 2 (PRRSV-2) circulates within 30%–50% of swine breeding farms in any given year (1, 2), causing both reproductive and respiratory impacts that result in >$600 million USD of productivity losses annually (3). These economic losses make PRRSV the most important virus affecting swine in the United States. Classified as the species Betaarterivirus americense (formerly Betaaterivirus suid 2) in the family Arteriviridae and order Nidovirales, PRRSV-2 is a rapidly evolving RNA virus characterized by enormous genetic and antigenic variability in the United States and globally (4–6). Control of this virus is hindered by routine emergence of novel, sometimes more virulent genetic variants (7–9), which result in recurrent epidemic waves of viral spread in the industry (5, 10).

PRRSV-2 is also one of the most sequenced viruses in the world (11), largely because sequencing is used by animal health professionals as a tool for routine monitoring of virus circulation within and between farms. While phylogenetic analysis is still the gold standard for interpretation of sequence data, practitioners and field epidemiologists often find it faster and more convenient to have a name in which they can refer to a given genetic variant as part of everyday communication and outbreak investigations. Currently, the naming method used by the industry to discriminate between sequences is restriction fragment length polymorphism (RFLP) typing (12), sometimes in combination with an additional label corresponding to phylogenetic lineage (4, 5). However, lineages and sub-lineages are large and diverse and hence are too coarse for on-farm disease monitoring, and using RFLP types to refer to PRRSV-2 viruses often leads to misleading conclusions (e.g., viruses assigned to the same RFLP type often are not genetically similar and vice versa) (13–15). For example, RFLP 1-4-4, which is one of the most abundantly reported RFLP types in the United States today, occurs in seven different lineages (16).

In a previous work, VanderWaal et al. (15) evaluated and compared 140 approaches for fine-scale classification of open reading frame 5 (ORF5) sequences. Three approaches were found to be robust and reproducible across trees built with different methods and data and thus could form the foundation for fine-scale classification of PRRSV-2 below the sub-lineage level. However, previous work did not explore the performance of PRRSV-2 variant classification on a rolling basis, and it is necessary to validate the performance and associated procedures for fine-scale classification that accommodates expanding genetic diversity on a prospective basis.

Taking insights and needs of practitioners and diagnosticians alongside a rigorous comparison of alternative approaches for classifying PRRSV-2 (15), the purpose of this paper is to introduce a new fine-scale genetic classification system for PRRSV-2 that is tailored to meet the needs of animal health professionals. Specifically, we outline criteria used for defining PRRSV-2 genetic variants, establish and test procedures for prospective implementation of the system, and assess the adaptability of the classification system to accommodate expanding genetic diversity at national scales. We also introduce a machine-learning webtool that can be used to identify the variant to which newly generated sequences belong and introduce naming conventions for PRRSV-2 variants. Finally, we report the results of a survey conducted with field practitioners on their motivations for submitting samples for sequencing and demonstrate how variant classification can enhance the utility of sequence data for the purposes of epidemiological monitoring and surveillance.

RESULTS

Variant classification

Utilizing sequences from the United States from 2015 to 2023, 25,403 PRRSV-2 ORF5 sequences were analyzed on a rolling quarterly basis to simulate prospective application of the variant classification system. Each quarter, groups of closely related sequences were identified in phylogenetic trees using a clustering algorithm and defined as a variant if the group (i) had five or more sequences, (ii) showed robust support of their shared ancestry in the ORF5 phylogeny (bootstrap value >85), and (iii) was >2% different from the nearest named variant. Any sequences belonging to clades that did not meet these requirements were labeled as “unclassified.” In total, the fine-scale classification system identified 118 genetic variants, 37 of which were common (detected >50 times) and 19 were rare (detected <10 times). Of the total sequences, 89.7% belonged to common variants, while 1.3% belonged to rare variants. The median number of sequences per variant was 25.5, with an interquartile range (IQR) of 11.25–68.75 sequences. The average within-variant genetic distance was 2.4% (IQR: 1.6%–3.2%, max: 4.8%). The mean distance to the closest related variant was 4.9% (IQR: 2.5%–5.6%). The distribution of variants on a phylogenetic tree is shown in Fig. 1. Variant nomenclature incorporates the sub-lineage to which the variant belonged, followed by an integer (i.e., 1A.3 and 1H.3 are the third variants identified within sub-lineages 1A and 1H, respectively). For contemporary sequences (2015 onward), several variants were a one-to-one correspondence with vaccine-like sequences, namely, variant 5A.1 (Ingelvac PRRS MLV; Boehringer Ingelheim Animal Health, Duluth, GA), 8A.1 (Ingelvac PRRS ATP; Boehringer Ingelheim Animal Health), 8C.1 (Fostera PRRS; Zoetis, Parsipanny, NH), 1D.2 (Prevacent PRRS; Elanco, Greenfield, IN), 7.1 (PrimePac PRRS; Merck, Rahway, NJ), and 1F.1 (PRRSGard, Pharmgate Animal Health, Wilmington, NC).

Fig 1.

Phylogenetic tree depicts hierarchical clustering of diverse taxa with labeled groups, branch lengths indicating genetic distances, and color-coded clades corresponding to distinct classifications or evolutionary relationships.

Phylogenetic tree (36 months) at the last timepoint (December 2023). Tip colors indicate variant. Color bars indicate sub-lineage.

Across 21 quarterly data sets (each containing 36 months of data, 2015–2023), genetic diversity within a variant did not show an increasing trend through time (Fig. S1a). The median bootstrap for the basal node of each clade was 100 (IQR: 95–100). Clade purity was calculated for each variant as the proportion of sequences in a phylogenetic clade that was assigned to the same variant ID. Clade purity was consistently high across quarters, with a median of 100% (IQR: 99%–100%, mean: 89.9%; Fig. S1b), indicating that variants formed compact groups and were not inter-mixed across the phylogenetic tree. Initially, ~36% of sequences could not be reliably grouped into a well-supported clade that met the variant criteria and were thus considered unclassified, but this value reduced and stabilized to ~11% in 2020 and 2021 and ~7% in 2022 and 2023. Lineage 1A accounted for 94.6% of unclassified sequences. Lineage 1A has a lower genetic diversity than other sub-lineages due to its more recent emergence approximately 10 years ago (5), and classification for this sub-lineage improved as clades became more diverged through time.

The median number of active variants per year was 48 (i.e., variants that are detected at least once during a calendar year; see Fig. 4). This compares to 65 RFLP types and 112 lineage + RFLP types, which are methods currently employed for fine-scale PRRSV classification. The median number of “common” active variants was 26, 25, and 39 for variants, RFLP types, and lineage + RFLPs, respectively. Thus, the new classification system does not result in a greater number of IDs than the industry currently is accustomed to with RFLP types.

There was a median of 19 new variants per year, but only four new common variants (those that would eventually be detected >50 times), demonstrating that variant classification is able to scale up to accommodate newly emerging PRRSV diversity (Fig. 2). In contrast, there were no new common RFLP types across the study period. Most newly identified variants were created from sequences that were previously unclassified and not from splits of existing variants. In total, 0.9% of sequences were re-named (i.e., a result of splitting a variant) at some point during the 21 quarters assessed here.

Fig 2.

Graphs depict trends in the number of IDs, including active and new variants, RFLP-types, and combined Lin+RFLP data from 2019 to 2023, depicting annual counts and thresholds for high activity levels exceeding 50 IDs per year.

Number of variants per year. Yearly number of active (blue) and new (red) variants for each classification method, with those that reach at least 50 sequences considered “common.” Solid lines show the number per year. Dashed lines show the median number across years.

Using a subset of data, we also constructed time-scaled phylogenetic trees to contextualize variant emergence and divergence on a time frame that is interpretable for epidemiological investigations of within- and between-farm transmission. We found that sequences belonging to the same variant typically descended from a common ancestor that existed ~2.3 years previously (median: 2.3 years, IQR: 1.5–3.5 years), which can be interpreted as that all sequences belonging to a single variant were part of chains of transmission originating around 2–3 years previously. This gives a time frame for which to search for epidemiological connections among cases. Divergence time from the closest relatives was 3.9 years (median: 2.9 years, IQR: 3.1–4.6 years). Clade purity in time-scaled Bayesian trees was high, with a median of 100% and an interquartile range from 90% to 100%.

Tools for assigning variant IDs to new sequences

Across the 21 quarters, the quarterly updated assignment algorithm had accuracies ranging between 97.1% and 99.9%, groupwise accuracies between 95.1% and 99.8%, mean precision (positive predictive value) between 96.6% and 99.9%, and mean recall between 95.1% and 99.8%. The percentage of sequences that were undetermined (i.e., probability of assignment was <0.25) ranged from 0.2% to 8.1%, with a median of 4.4%. The most up-to-date model is accessible via an RShiny webtool (https://stemma.shinyapps.io/PRRSLoom-variants/), and the trained model is available on GitHub in both R and Python (https://github.com/kvanderwaal/prrsv2_classification). Whether using the webtool or the R/Python code, the user uploads ORF5 sequences in fasta format, which are then realigned to the PRRSV-2 prototype sequence VR2332 (GenBank accession number EF536003). Sequences are not saved or retained by the webtool in any way. The tool then estimates the probability that the sequence belongs to each defined variant (Fig. 3). For each sequence, outputs include the assignment probability for the variant ID with the highest (top) and second highest probability. A final assignment is also given, with sequences that could not be assigned to any variant with >0.25 probability listed as “undetermined.” It is possible that the variant with the highest probability is not substantially greater than the second highest probability, which may indicate potential misclassification. If the highest probability is more than double the second highest probability, then the assigned variant ID can be interpreted with greater confidence. While this paper reports results up to December 2023, the PRRSLoom-variant webtool and the trained model have been updated since then to reflect recent data, and this will be maintained on a quarterly basis.

Fig 3.

Interface depicts PRRSLoom-Variants tool for uploading FASTA sequences, assigning most likely variants with probabilities, and depicting match confidence for uploaded sequences with options to download predictions and read assignment details.

PRRS-Loom webtool interface. (a) Sequences can either be uploaded as a FASTA file or copy-pasted into the window. The webtool then estimates the probability that each uploaded sequence belongs to each defined variant. For each sequence, outputs include the assignment probability for the variant ID with the highest (top) and second highest probability (b and c, respectively). (d) A final assignment is also given, with sequences that could not be assigned to any variant with >0.25 probability listed as “undetermined.”

Survey on use of PRRSV-2 sequence data by animal health professionals

In a survey administered by the American Association of Swine Veterinarians, swine practitioners (n = 92) were asked to rank the primary motivations for which they submitted samples for sequencing. The motivations that were consistently highly ranked included (i) “anticipate and track the spread of novel and emerging variants,” (ii) “discriminate between previous and new wild-type strains on the same farm,” and (iii) “determine possible source of introduction” (Fig. 4).

Fig 4.

Bar chart depicts various objectives for within-farm monitoring, regional spread analysis, and immunological studies, including vaccine differentiation, source tracing, variant tracking, herd evolution documentation, and immune phenotype classification.

Results of survey where animal health professionals were asked to rank their top four reasons for submitting samples for sequencing from a list of 10 options. Bars represent the number of respondents that selected each answer, with color shading representing rank (with 1 being high).

To better visualize how the new classification system tracks the spread of emerging variants, which was the most highly ranked motivation for sequencing, we selected variants that had <25 sequences at the time of naming and >200 sequences by the end of the study period. Five variants met these criteria (Fig. 5). We also considered the 1H.18 variant due to interest in this variant at the time of writing (17). There was a median of 10.5 different RFLP types per variant. Of the common RFLP types (n > 50), none were exclusive to any of the emerging variants. Indeed, these common RFLPs were all found in ≥3 of the six emerging variants and across a median of 33 variants overall. These insights show the benefits of using variant classification as opposed to RFLP typing for identifying and tracking emerging strains of the virus.

Fig 5.

Line graph depicts increasing sequences for emerging variants 1A.2, 1C.2, 1H.18, 1C.5, 1H.11, 1A.29, and others from 2018 to 2023. Phylogenetic trees depict evolutionary relationships and variant spread across years.

Cumulative number of sequences per emerging variant over time (top). Phylogenetic trees (bottom) from September of each year, with emerging variants colored and all other sequences shown in gray.

To discriminate between previous and new wild-type strains on the same farm and to determine possible source of introduction, which were the second and third most highly ranked motivations for sequencing, we partnered with production system veterinarians and applied the new variant classifications to sequences collected from their farms. System 1 shared 28 PRRSV-2 ORF5 sequences from 12 farms, with a particular interest in 13 sequences from farm 5, which was a sow farm (Fig. 6). This farm experienced four PRRSV circulation events: variant 1B.8 in 2015–2016, 1C.3 in 2016–2018, 1A.13 in 2019, and 1C.5 in 2020–2024. Of note, RFLP types or lineages were not able to discriminate between new introductions on farm 5. Either a new introduction did not receive a unique label (as in 2020, where sub-lineages failed to discern a new 1C virus, despite having <91% nucleotide identity with the previous 1C virus on the farm) or multiple sequences that were part of the same circulation event received different labels (three different RFLP types among the five 1C.3 sequences, despite having >98% nucleotide identity). This limitation of RFLP types is more thoroughly quantified in VanderWaal et al. (15), wherein 43% of on-farm circulation events attributable to a single variant had multiple associated RFLP types.

Fig 6.

Phylogenetic trees depict relationships among RFLP and variant classifications. Heatmap depicts nucleotide identity percentages for ORF5 among sequences from various farms, with values ranging from 84.2 to 100. Scale bars indicate 0.02 substitutions.

Phylogenetic tree of ORF5 sequences for system 1 reconstructed by Bayesian inference (mrBayes v.3.2 [18]), rooted on the only sequence not belonging to lineage 1 (sub-lineage 8A), with tip color indicating (a) RFLP type and (b) variant. (c) ORF5 nucleotide identity, with redder colors representing higher identity.

System 2, which is a large production system operating in five states, shared 1,095 ORF5 sequences from 2014 to 2022. In the phylogenetic tree in Fig. 7, the majority of the sequences belonged to sub-lineages 1C, 1H, and 1E. Lineages and sub-lineage do not provide sufficient resolution to distinguish new and already circulating wild-type viruses on farms and also fail to provide a distinguishing label for one clade that contains recombinant viruses (e.g., sequences in this clade share the same recombination profiles, suggesting that they descended from the same recombinant ancestor). The vast majority of sequences belong to two RFLP types (1-8-4 and 1-4-4), which do not cluster together genetically. Furthermore, the RFLP 1-4-4 sequences are not related to the recently emerged outbreak variant bearing the same RFLP type (the so-called novel L1C-1-4-4 variant, which is referred to as 1C.5 in the new variant system [7]). In contrast, the variant classifications are aligned with clades that are visually well differentiated and also provide a unique variant ID for the recombinant clade. Thus, the improved labeling of closely related sequences in the variant system enhances a practitioner’s ability to track spread between farms, detect the introduction of new variants into a farm or flow, and narrow the possible sources of introduction.

Fig 7.

Phylogenetic trees depict lineage, variant, and RFLP classifications with recombinant regions indicated. Color-coded clusters represent diverse groups, and scale bars depict 0.08 substitutions per site, highlighting evolutionary relationships.

Maximum-likelihood tree of PRRSV-2 ORF5 sequences for system 2, colored by (a) lineage, (b) variant, and (c) RFLP type. The red bar indicates a clade that includes sequences that share the same recombination profiles, as determined by recombination screening conducted with RDP4 software (19).

As a final example, we zoom in on 25 sequences collected during a single year from 15 farms. In Fig. 8, these farms are plotted on a landscape, with dots colored according to the variant ID or the RFLP type of the sequences found from those farms. The farm of interest (circled in red) has a sequence classified as 1H.26 with an RFLP type of 1-4-4. In total, there were six farms where 1H.26 was detected in this area, and the phylogenetic tree (Fig. 8a) and homology table (Fig. 8c) demonstrate that these sequences were closely related with genetic distances mostly <2.5%. There were also six farms with RFLP 1-4-4, but they were not the same six farms nor were the 1-4-4 sequences closely related (Fig. 8b). Compared to variant classification, investigations based on RFLP-typing would result in the wrong farms being identified, thus leading to erroneous conclusions about epidemiological links and avenues of transmission between these farms.

Fig 8.

Maps depict spatial distribution of genetic variants and RFLP types across six farms with arrows indicating spread. Heatmap depicts genetic distances in percentages. Phylogenetic trees depict evolutionary relationships with 0.04 scale bars.

(a) Map and phylogenetic trees of 25 sequences collected in a single year from 15 farms in a small geographic area, colored according to (a) variant ID and (b) RFLP type. A farm of interest is circled in red, and green arrows are used to indicate other farms where the same variant ID (1H.26) or RFLP type (1-4-4) was detected. (c) A genetic homology table with redder colors indicating higher relatedness.

DISCUSSION

While the shortcomings of existing PRRSV-2 ORF5-based classification systems have been apparent as early as 2011 (20), recent advances in computational power and the creation of national-scale sequence databases have created the opportunity to finally address these issues. Here, we outline a fine-scale classification system (below sub-lineage level) for PRRSV-2 in the United States that is expandable to new genetic diversity that emerges as consequence of virus evolution. We lay out procedures for quarterly updating of the classification system and for assignment of newly generated sequences via a centrally maintained machine-learning model which facilitates a unified naming scheme across the United States. The level of granularity represented by genetic variants was tailored to meet the needs of animal health professionals, who primarily reported using sequence data for epidemiological monitoring. In this paper and in VanderWaal et al. (15), we demonstrate that as compared to RFLP typing, variant classifications more reliably group viruses based on relatedness in the ORF5 gene and provide better discrimination between unrelated viruses. This facilitates on-farm monitoring, detection of new introductions to a farm or production system, and tracking of regional and between-farm spread.

Our fine-scale classification system is an extension of the lineage and sub-lineage classifications first proposed in 2010 (21) and refined in the past 5 years (4, 5, 10). Lineages represent the broadest classification, with genetic distances typically <11% within a lineage based on ORF5. Sub-lineages typically have genetic distances of <8.5% and are made up of numerous genetic variants. Sequences belonging to the same variant typically have an average genetic distance of 2%–3% but can sometimes be as much as almost 5% different. Our intent is not to replace lineages, as we do believe that these larger classifications are useful for explorations of phenotype as well as tracking the macro-evolutionary dynamics of PRRSV-2. Therefore, we incorporate lineage into the IDs utilized in the variant classification system to provide a general zip code of the variant within the larger genetic diversity of PRRSV-2.

A major advantage of a unified variant classification system within the United States is to facilitate communication among animal health professionals, diagnostic laboratories, and researchers. In the past, tracking of emerging variants was typically accomplished either by using RFLP types or by calculating the genetic distance to “anchor” sequences established for a particular strain. RFLP typing, with its associated limitations and inaccuracies, can generate confusion in the field about which farms are part of an outbreak, both missing farms that should be included (i.e., a closely related virus with a different RFLP type) and sparking false alarms (i.e., a distantly related virus with the same RFLP type), as happened in the early days of the emergence of the 1C.5 variant (7). Calculating the genetic distance between sequences is a viable alternative for determining relatedness but requires someone to set anchor sequences for a particular outbreak (which is usually only done for variants of heightened concern), and importantly, requires a several step process of sharing sequences (which are often considered confidential), aligning them and calculating distances in bioinformatic software. These steps are not required if the variant ID of the respective sequences is already assigned by diagnostic labs as part of their reports to clients.

However, an important caveat is that variant classification is not based on immunological or virulence variability of the virus (i.e., the phenotype), and most variants will not have been fully characterized from a phenotypic standpoint even when a whole genome is available. Thus, variant classification is not designed to provide information on the clinical manifestations of the virus in a herd, which is influenced by a myriad of factors in addition to the virus itself (e.g., co-infections, host genetics, and immunological history) (22–25). Variant classification also does not directly translate to immunological cross-protection. Although viruses labeled as the same variant are more genetically homologous on ORF5, cross-protection is not simply a function of genetic distance between viruses (26, 27). Whole-genome data are required for phenotypic investigations, and variant classifications may be too fine scale to expect major phenotypic differences among closely related variants. However, if we made variants less granular, we would lose their utility for epidemiological investigations, such as determining possible sources of introduction and tracking regional spread.

Sequences belonging to the same variant have a relatively recent common ancestor (2–3 years), minimizing the time period during which recombination may occur and making it more likely that relatedness based on ORF5 reflects relatedness on the whole genome. That being said, recombination can occur. If a recombination event occurs that leaves a large number of descendants, these are often discernible as divergent clades in ORF5 phylogenies. This can be observed, for example, in the recombinant clade present in system 2 (all sequences share a similar recombination profile, Fig. 7). When the recombination breakpoint is outside of ORF5 (28), the diverging trajectory of the recombinant group of viruses is sometimes discernible on ORF5-based trees, as is the case of variant 1C.5, even though whole-genome sequencing would be required to fully characterize this (29). That being said, viruses resulting from recombination events that do not spawn large numbers of descendants (i.e., their descendants are not numerous enough to form identifiable clades) would likely not receive distinct IDs in the variant classification system, which is a limitation.

Variant classification also facilitates the generation, organization, and findability of additional information or research related to a particular variant, such as regional incidence trends, production impacts, associated whole genomes and genetic markers of virulence therein (30), and whether the group includes recombinant viruses. This, combined with the ease of cross-communication across diagnostic labs, researchers, and the field, could lay the foundation for additional research on PRRSV epidemiology, virology, and immunology.

Limitations to the proposed variant classification include the coverage of our sequence data set. This system was based on U.S. PRRSV-2 sequences from 2015 to the present. Given that our data set covers >55% of U.S. swine production and spans 27 states (including all major pig producing regions, notably, the Midwest, Eastern, and Great Plains states), we believe that our data are reasonably representative of PRRSV-2 diversity circulating in the major pork-producing regions in the country since 2015. Thus, we urge potential users of the webtool to be cognizant of the year of sequence collection and origin (country) of any sequences they may upload. While earlier sequences can be input into the webtool, they are likely to be predicted as unclassified, given that diversity present in previous decades was not represented in our data set. Similarly, it is not meant to encompass genetic diversity from other countries. Our analysis of time-scaled trees suggests that sequences belonging to the same variant typically evolved from a common ancestor that existed 1.5–3.5 years previously. This short timescale and the rapid evolutionary rate of the virus (5) support the idea that variants circulating in the United States are likely distinct from variants in other countries, except perhaps in cases where there is more frequent transboundary movement (such as Canada). The system could be expanded to PRRSV-2 in other countries by incorporating their data into our quarterly updates to identify and name new variants. In the absence of sharing of larger databases, five representative sequences from a particular clade (either in the United States or elsewhere) that is currently unclassified and for which a variant ID is desired can be submitted to our system. Alternatively to using our platform, we suggest that other countries could adopt our criteria for defining a variant so that this term can be used more consistently across continents.

While having an improved naming scheme for PRRSV-2 genetic variants will not solve PRRS in the United States, a classification system for field-based epidemiological monitoring is needed and has been requested by practitioners for many years. In this paper, we outlined the definition of genetic variants, the procedures for systemic identification of variants in a nationwide sequence data set, and a validated workflow for routinely updating variant classification on a quarterly basis. The latter will ensure that the nomenclature system can dynamically adapt to evolving PRRSV-2 diversity across time and space. Variant classification will facilitate communication about outbreaks, tracking of emerging and endemic variants across time and space, as well as provide a framework to more rigorously analyze the genetic basis of variability in phenotype or production impacts. Finally, this work was conducted with iterative feedback from a working group of veterinarians, researchers, and diagnosticians. This close engagement with stakeholders and end users has been crucial for the operationalization and adoption of the variant classification, ensuring that it is tailored to the needs of animal health professionals utilizing sequence data for disease management in the field.

MATERIALS AND METHODS

Data source and pre-processing

Sequence data were obtained from the Morrison Swine Health Monitoring Project (MSHMP), which is a voluntary initiative operated by University of Minnesota that monitors PRRS occurrence in the United States. MSHMP was initiated in 2011 and currently collects sequence data for farms belonging to 37 production systems, accounting for >55% of the U.S. pig population (2). Participating production systems share PRRSV ORF5 sequences that are generated as part of routine monitoring and outbreak investigations in breeding, gilt developing units, growing, and finishing herds (31). Sequences are generally obtained either directly from each MSHMP participant or from the main veterinary diagnostic laboratory where participants submit their diagnostic samples. Metadata for each sequence include farm name, date, and farm type of origin (e.g., breeding or growing herd).

A total of 16,260 sequences were available from 1 October 2015 to 30 June 2021. These sequences were used to establish the rolling procedures for updating the classification system across time. An additional 9,143 sequences were available from 1 July 2021 to 31 December 2023. Based on VanderWaal et al., sequence data sets that lack duplicated (100% nucleotide identity) sequences produced the most consistent variant classifications (15). Therefore, sequences with 100% identity were de-duplicated before phylogenetic tree building. Duplicated sequences were retained for calculations of the frequency and mean genetic distances of variants. Sequences with ≥4 ambiguous nucleotides (0.5% of ORF5) or with gaps greater than 24 positions were removed from the data set (4). To assess how the system would function if utilized prospectively, we initiated the classification system in 2018 with 36 months of data (2015–2018), then added new data every 3 months up to December 2023. Thus, each quarter of data included the previous 36 months of sequence data, with a median of 8,620 sequences per set.

Tree building

For each data set of 36 months, sequences were aligned to a consensus reference sequence based on all previous data with --6merpair, --keeplength, and --addfragments options of the MAFFT algorithm (32, 33). All tree building, unless otherwise specified, utilized the maximum likelihood method performed using IQ-TREE2 with 1,000 ultrafast bootstraps (34). As described in VanderWaal et al. (15), strict majority-rule consensus trees were constructed (clades with bootstrap support of <50 were collapsed), with the general time-reversible substitution model with empirical base frequencies and gamma plus invariant site heterogeneity (GTR + F + I + G4). The ggtree package in R was used for all tree visualizations, with trees re-rooted on lineage 5, which contains the PRRSV-2 prototype virus (VR2332, GenBank accession number EF536003) (21, 35).

Variant classification: initialization

Initial time step

Utilizing the first tree of 36 months (9,783 sequences, 1 October 2015–30 September 2018), we applied a tree-based clustering approach to the phylogeny using the average-clade method in the TreeCluster package available in Python (36); clusters of genetically related sequences in the trees were referred to as “variants.” Briefly, this method identifies monophyletic clades where the average pairwise patristic distance between sequences within the clade is <7%. This threshold was selected based on a rigorous comparison of thresholds performed by VanderWaal et al. (15). In that analysis, the observed average pairwise distance between sequences belonging to the same variant was 2.3% (15).

Additional steps were applied based on preliminary results showing that some variants defined on the first tree of 36 months did not consistently group together in subsequent trees. First, post-processing of TreeCluster outputs was performed to merge clusters with low support (supplemental text). Second, additional criteria for defining a variant were that the group must have (i) five or more sequences, (ii) robust support of their shared ancestry in the ORF5 phylogeny (bootstrap value >85 in the tree), and (iii) that the genetic distance to the nearest named variant must be >2%. Any sequences belonging to clades that did not meet these requirements were labeled as unclassified. Nomenclature for variants was the sub-lineage to which the variant belonged, followed by an integer (i.e., 1A.3 and 1H.3 are the third variants identified within sub-lineages 1A and 1H, respectively). To align with the five sub-groups within sub-lineage 1C delineated by reference 4, we identified the variants that corresponded to those groups and utilize the same IDs and continue numbering onward from 1C.6.

Variant classification: updating

With each new quarter, new sequences from the most recent 3 months (median: 663 sequences) were assigned to variants using the assignment algorithm trained at the end of the previous quarter (see next section). A new tree of 36 month was constructed, and variant IDs were annotated to the tree. The tree was systematically examined for new variants using TreeCluster as well as for splits in existing variants (see supplemental text for details). Splits were systematically considered for variants where the 95th percentile of pairwise genetic distances was ≥5% (based on sequences from the previous 12 months to better capture recent genetic divergence). A new variant was only created if a clade met the following conditions: (i) it consisted of five or more sequences; (ii) it had robust support of their shared ancestry in the ORF5 phylogeny (bootstrap value >85 in the tree); and (iii) the genetic distance to the nearest named variant was >2%. If the creation of the new variant was due to a split in an existing variant, a new variant was created only if the minimum and median genetic distance between the new and original variant were >3% and >5%, respectively. These high thresholds were set to minimize the number of sequences being re-named as a result of variant splitting, as per the request of diagnostic laboratories. New variants receive names in the same manner as described above (i.e., if 1A.3 is split, one daughter group retains the name 1A.3, and the other receives the next integer in the series, e.g., 1A.8).

Algorithm for assignment of new sequences

For prospective application of any classification system, it is desirable to be able to assign new sequences to variants without performing computationally heavy analysis. We thus trained a random forest machine learning algorithm to assign new sequences to the appropriate variant ID (15, 37). Up to 120 sequences per variant (approximately 10 per quarter) were randomly selected from the initial time step to build a training data set, which was then appended quarterly with new sequence data. Using the training data set for each quarter, a random forest algorithm was fitted using the “caret” package in R using 10-fold cross-validation and auto-tuning of the “mtry” hyper-parameter (38). In parallel, we also trained a random forest in Python for Python end users (supplemental text). Model performance on the training set was assessed using 10-fold cross-validation (i.e., performance evaluated on 10% of observations that were left out of 10 iterative random forest runs). We report the overall accuracy (percentage of sequences correctly classified by the algorithm) of the training data set. We also calculated the mean groupwise precision (a.k.a. positive predictive value), recall (a.k.a. sensitivity), and accuracy (i.e., percentage of sequences correctly classified per variant was first calculated, and then a mean of these groupwise accuracies was reported).

Outputs from the trained algorithms include the probabilities of the first, second, and third most likely variant ID for a given sequence, with the highest probability ID being assigned to the sequence for downstream analyses of predictive performance. In some cases, the highest probability ID was quite low, indicating that the model had poor confidence in the assignment. Therefore, sequences with assignment probabilities of <0.25 were considered undetermined and not considered in calculations of model accuracy. The proportion undetermined was tracked and reported. More stringent thresholds do not markedly improve model accuracy but resulted in a higher percentage of undetermined sequences (15). The training data set and algorithm are updated each quarter to include sequences (up to 120) from new variants as well as additional recent sequences from existing variants (up to 60). Older sequences and variants are not removed from the assignment algorithm, in order for the model to retain the ability to predict on older sequences from 2015 to the present. Only sequences with assignment probabilities of >0.4 were included in the training data set.

An RShiny webtool was developed and is updated quarterly (https://stemma.shinyapps.io/PRRSLoom-variants/) so that end users can assign new sequences to variants. The updated algorithms are also available as R and Python scripts so that they can be used in command lines or ported to external applications maintained by diagnostic laboratories or other groups (https://github.com/kvanderwaal/prrsv2_classification). This ensures that all potential end users will obtain the same variant classifications regardless of the platform.

Genetic characterization and phylogenetic properties of variants

At the final timepoint, genetic characterization of variants produced by each approach included (i) the number of variants identified; (ii) the number of common variants (n > 50 sequences belonging to the variant); (iii) the median size (sequences per variant) and interquartile range; (iv) the percentage of sequences belonging to common variants; (v) the percentage of sequences belonging to rare variants (n < 10 sequences); (vi) the median bootstrap value and interquartile range of the ancestral node; (vii) the mean genetic distance (raw p-distance) within a variant; and finally, (viii) calculation of the genetic distance to the most closely related cluster for each variant. The cutoffs for rare and common variants (10 and 50, respectively) were chosen based on the distribution of the data. Three years post-detection, these cutoffs of resulted in 15% of the variants being classified as rare and 30% classified as common, with over 60% of sequences belonging to common variants. Based on preliminary analysis of the data set utilized in VanderWaal et al. (15), common variants were found in a median of three U.S. states (IQR: two to five states), whereas other variants were found in a median of only one state (IQR: one to two states).

Across all timepoints, we also evaluated the mean genetic distance through time to better understand how the mean within-variant distance may expand as a result of ongoing evolution, as well as clade purity over time to assess the tendency of sequences belonging to the same variant to remain grouped together in the tree over time. Clade purity was calculated as the proportion of sequences in a phylogenetic clade that was assigned to the same variant ID (see supplemental text for details).

We also calculated the number of new variants detected per year and number of active variants per year. The number of new variants per year was based on the calendar year of the earliest sequence belonging to a variant, and active variants per year included all variants whose earliest and latest detected sequences occurred before, during, or after the considered calendar year. For comparison purposes, these values were compared to RFLP types and lineage + RFLP types.

Using a subset of data, we also constructed time-scaled phylogenetic trees (supplemental methods) to contextualize the timeline of variant emergence and divergence on a time frame that is interpretable for epidemiological investigations of within- and between-farm transmission. Briefly, for each variant in the time-scaled trees, we extracted the time to the most recent common ancestor, which was used to calculate clade age, divergence time from the most closely related variant, and clade purity.

Working group and practitioner survey

A working group was established in March 2021 that included representatives from major swine-oriented diagnostic labs (University of Minnesota, Iowa State University, South Dakota State University, and Ohio Animal Disease Diagnostic Lab), swine disease monitoring programs that serve as national repositories of PRRSV sequences (Morrison Swine Health Monitoring Program [31], Swine Disease Reporting System [39, 40], and USDA-Agricultural Research Service’s Swine Pathogen Database [41]), and swine veterinarians and production systems. This group was involved iteratively in the development of the new classification system (Fig. S2).

The working group developed and administered a survey to swine health professionals in the United States to better understand how they use genetic sequence data. In this survey, practitioners were asked to rank the top four reasons for which they submitted samples for sequencing (out of 10 options related to within-farm monitoring, between-farm or regional spread, or immunological/phenotype considerations). This survey was distributed in April 2022 by the American Association of Swine Veterinarians. Based on the results of this survey and with input from the working group, the final classification system was tailored to meet the needs of practitioners by explicitly addressing the primary motivations for sequencing.

ACKNOWLEDGMENTS

The authors thank members of the American Association of Swine Veterinarians porcine reproductive and respiratory syndrome virus nomenclature working group, including Andreia Arruda, Srijita Chandra, Eric Nelson, Tom Petznick, Melanie Prarat, Mark Schwartz, Donna Drebes, Mark Wagner, Jessica Seate, Gustavo Silva, Joel Sparks, and Paul Yeske. The authors also thank the Morrison Swine Health Monitoring Project (MSHMP) participants and the Veterinary Diagnostic Laboratory, University of Minnesota, for sharing its porcine reproductive and respiratory syndrome virus 2 genetic sequences.

Mention of trade names or commercial products in this paper is solely for the purpose of providing specific information and does not imply recommendation or endorsement by the U.S. Department of Agriculture (USDA). The findings and conclusions in this publication are those of the authors and should not be construed to represent any official USDA or U.S. Government determination or policy. USDA is an equal opportunity provider and employer.

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication. This study was funded by the joint NIFA-NSF-NIH Ecology and Evolution of Infectious Disease award 2019-67015-29918, the Intramural Research Program of the U.S. Department of Agriculture, National Institute of Food and Agriculture, Data Science for Food and Agricultural Systems Program, grant number 2023-67021-40018, a grant from the American Association of Swine Veterinarians, and the U.S. Department of Agriculture, Agricultural Research Service project 5030-32000-231-000-D. MSHMP was funded by the Swine health Information Center (www.swinehealth.org, project #23-079).

Author order of the first three and last two authors was based on effort. The remaining authors were listed in reverse alphabetical order. K.V. designed the research, performed the analysis, and wrote the paper. N.P., M.K., and I.P. performed some analyses, contributed to interpretation, and edited the paper. N.P. and D.N.M. developed analytical pipelines and the RShiny app. D.C.L.L., G.T., J.Z., T.K.A., M.Z., C.A.C., and D.J.H. contributed to research design, results interpretation, and manuscript editing.

Contributor Information

Kimberly VanderWaal, Email: kvw@umn.edu.

Michael J. Imperiale, University of Michigan, Ann Arbor, Michigan, USA

DATA AVAILABILITY

Reference sequences for the findings of this paper may be available upon reasonable request to the corresponding author. The data are not publicly available as they are part of diagnostic data from third parties (companies and veterinarians submitting samples for diagnosis).

SUPPLEMENTAL MATERIAL

The following material is available online at https://doi.org/10.1128/msphere.00709-24.

Supplemental material. msphere.00709-24-s0001.pdf.

Supplemental text and figures.

DOI: 10.1128/msphere.00709-24.SuF1

ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.

REFERENCES

  • 1. Perez AM, Linhares DCL, Arruda AG, VanderWaal K, Machado G, Vilalta C, Sanhueza JM, Torrison J, Torremorell M, Corzo CA. 2019. Individual or common good? Voluntary data sharing to inform disease surveillance systems in food animals. Front Vet Sci 6:194. doi: 10.3389/fvets.2019.00194 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. MSHMP . 2024. PRRS cumulative incidence. Morrison Swine Health Monitoring Project. Available from: https://mshmp.umn.edu/reports#Charts [Google Scholar]
  • 3. Holtkamp DJ, Kliebenstein JB, Zimmerman JJ, Neumann E, Rotto H, oder TK, Wang C, Yeske P, Mowrer CL, Haley C. 2012. Economic impact of porcine reproductive and respiratory syndrome virus on U.S. Pork producers. Iowa State University Animal Industry Report 9 [Google Scholar]
  • 4. Yim-Im W, Anderson TK, Paploski IAD, VanderWaal K, Gauger P, Krueger K, Shi M, Main R, Zhang J. 2023. Refining PRRSV-2 genetic classification based on global ORF5 sequences and investigation of their geographic distributions and temporal changes. Microbiol Spectr 11:e02916-23. doi: 10.1128/spectrum.02916-23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Paploski IAD, Pamornchainavakul N, Makau DN, Rovira A, Corzo CA, Schroeder DC, Cheeran MCJ, Doeschl-Wilson A, Kao RR, Lycett S, VanderWaal K. 2021. Phylogenetic structure and sequential dominance of sub-lineages of PRRSV Type-2 lineage 1 in the United States. Vaccines (Basel) 9:608. doi: 10.3390/vaccines9060608 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Meng XJ. 2000. Heterogeneity of porcine reproductive and respiratory syndrome virus: implications for current vaccine efficacy and future vaccine development. Vet Microbiol 74:309–329. doi: 10.1016/s0378-1135(00)00196-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Kikuti M, Paploski IAD, Pamornchainavakul N, Picasso-Risso C, Schwartz M, Yeske P, Leuwerke B, Bruner L, Murray D, Roggow BD, Thomas P, Feldmann L, Allerson M, Hensch M, Bauman T, Sexton B, Rovira A, VanderWaal K, Corzo CA. 2021. Emergence of a new lineage 1C variant of porcine reproductive and respiratory syndrome virus 2 in the United States. Front Vet Sci 8:752938. doi: 10.3389/fvets.2021.752938 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. van Geelen AGM, Anderson TK, Lager KM, Das PB, Otis NJ, Montiel NA, Miller LC, Kulshreshtha V, Buckley AC, Brockmeier SL, Zhang J, Gauger PC, Harmon KM, Faaberg KS. 2018. Porcine reproductive and respiratory disease virus: evolution and recombination yields distinct ORF5 RFLP 1-7-4 viruses with individual pathogenicity. Virology (Auckl) 513:168–179. doi: 10.1016/j.virol.2017.10.002 [DOI] [PubMed] [Google Scholar]
  • 9. Rawal G, Almeida MN, Gauger PC, Zimmerman JJ, Ye F, Rademacher CJ, Armenta Leyva B, Munguia-Ramirez B, Tarasiuk G, Schumacher LL, Aljets EK, Thomas JT, Zhu JH, Trexel JB, Zhang J. 2023. In vivo and In vitro characterization of the recently emergent PRRSV 1-4-4 L1C variant (L1C.5) in comparison with other PRRSV-2 lineage 1 isolates. Viruses 15:2233. doi: 10.3390/v15112233 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Paploski IAD, Corzo C, Rovira A, Murtaugh MP, Sanhueza JM, Vilalta C, Schroeder DC, VanderWaal K. 2019. Temporal dynamics of co-circulating lineages of porcine reproductive and respiratory syndrome virus. Front Microbiol 10:2486. doi: 10.3389/fmicb.2019.02486 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. VanderWaal K. 2024. Surprising but true: PRRSV one of the most sequenced viruses in the world. National Hog Farmer. Available from: https://www.nationalhogfarmer.com/livestock-management/surprising-but-true-prrsv-one-of-the-most-sequenced-viruses-in-the-world [Google Scholar]
  • 12. Wesley RD, Mengeling WL, Lager KM, Clouser DF, Landgraf JG, Frey ML. 1998. Differentiation of a porcine reproductive and respiratory syndrome virus vaccine strain from North American field strains by restriction fragment length polymorphism analysis of ORF 5. J Vet Diagn Invest 10:140–144. doi: 10.1177/104063879801000204 [DOI] [PubMed] [Google Scholar]
  • 13. Murtaugh MP, Stadejek T, Abrahante JE, Lam TTY, Leung FC-C. 2010. The ever-expanding diversity of porcine reproductive and respiratory syndrome virus. Virus Res 154:18–30. doi: 10.1016/j.virusres.2010.08.015 [DOI] [PubMed] [Google Scholar]
  • 14. Cha SH, Chang CC, Yoon KJ. 2004. Instability of the restriction fragment length polymorphism pattern of open reading frame 5 of porcine reproductive and respiratory syndrome virus during sequential pig-to-pig passages. J Clin Microbiol 42:4462–4467. doi: 10.1128/JCM.42.10.4462-4467.2004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. VanderWaal K, Pamornchainavakul N, Kikuti M, Linhares DCL, Trevisan G, Zhang J, Anderson TK, Zeller M, Rossow S, Holtkamp DJ, Makau DN, Corzo CA, Paploski IAD. 2024. Phylogenetic-based methods for fine-scale classification of PRRSV-2 ORF5 sequences: a comparison of their robustness and reproducibility. Front Virol 4:1433931. doi: 10.3389/fviro.2024.1433931 [DOI] [Google Scholar]
  • 16. SDRS . 2024. PRRSV genotyping dashboard: RFLP and lineages. Swine Disease Reporting System. Available from: https://fieldepi.org/domestic-swine-disease-monitoring-program [Google Scholar]
  • 17. Kikuti M, Paploski IAD, Pamornchainavakul N, Mellini M, Yue X, Vadnais S, Baker J, Silva J, Wagner M, Kurt J, Schwartz M, Rossow S, Corzo C, VanderWaal K. 2024. Monitoring the detection of PRRSV variant 1H.18. Morrison Swine Health Monitoring Project Science Page. Available from: https://mshmp.umn.edu/sites/mnshmp.umn.edu/files/2024-05/SHMP%202023l24.43%20%5BMonitoring%20PRRS%20Variant%5D2.pdf [Google Scholar]
  • 18. Ronquist F, Teslenko M, van der Mark P, Ayres DL, Darling A, Höhna S, Larget B, Liu L, Suchard MA, Huelsenbeck JP. 2012. MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space. Syst Biol 61:539–542. doi: 10.1093/sysbio/sys029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Martin DP, Murrell B, Golden M, Khoosal A, Muhire B. 2015. RDP4: detection and analysis of recombination patterns in virus genomes. Virus Evol 1:vev003. doi: 10.1093/ve/vev003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Murtaugh M. 2012. Use and interpretation of sequencing in PRRSV control programs
  • 21. Shi M, Lam TT-Y, Hon C-C, Murtaugh MP, Davies PR, Hui RK-H, Li J, Wong LT-W, Yip C-W, Jiang J-W, Leung FC-C. 2010. Phylogeny-based evolutionary, demographical, and geographical dissection of North American type 2 porcine reproductive and respiratory syndrome viruses. J Virol 84:8700–8711. doi: 10.1128/JVI.02551-09 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Niederwerder MC, Jaing CJ, Thissen JB, Cino-Ozuna AG, McLoughlin KS, Rowland RRR. 2016. Microbiome associations in pigs with the best and worst clinical outcomes following co-infection with porcine reproductive and respiratory syndrome virus (PRRSV) and porcine circovirus type 2 (PCV2). Vet Microbiol 188:1–11. doi: 10.1016/j.vetmic.2016.03.008 [DOI] [PubMed] [Google Scholar]
  • 23. Lough G, Rashidi H, Kyriazakis I, Dekkers JCM, Hess A, Hess M, Deeb N, Kause A, Lunney JK, Rowland RRR, Mulder HA, Doeschl-Wilson A. 2017. Use of multi-trait and random regression models to identify genetic variation in tolerance to porcine reproductive and respiratory syndrome virus. Genet Sel Evol 49:37. doi: 10.1186/s12711-017-0312-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Stoian AMM, Rowland RRR. 2019. Challenges for porcine reproductive and respiratory syndrome (PRRS) vaccine design: reviewing virus glycoprotein interactions with CD163 and targets of virus neutralization. Vet Sci 6:9. doi: 10.3390/vetsci6010009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Popescu LN, Trible BR, Chen N, Rowland RRR. 2017. GP5 of porcine reproductive and respiratory syndrome virus (PRRSV) as a target for homologous and broadly neutralizing antibodies. Vet Microbiol 209:90–96. doi: 10.1016/j.vetmic.2017.04.016 [DOI] [PubMed] [Google Scholar]
  • 26. Makau DN, Prieto C, Martínez-Lobo FJ, Paploski IAD, VanderWaal K. 2023. Predicting antigenic distance from genetic data for PRRSV-type 1: applications of machine learning. Microbiol Spectr 11:e04085-22. doi: 10.1128/spectrum.04085-22 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Kim WI, Kim JJ, Cha SH, Wu WH, Cooper V, Evans R, Choi EJ, Yoon KJ. 2013. Significance of genetic variation of PRRSV ORF5 in virus neutralization and molecular determinants corresponding to cross neutralization among PRRS viruses. Vet Microbiol 162:10–22. doi: 10.1016/j.vetmic.2012.08.005 [DOI] [PubMed] [Google Scholar]
  • 28. Pamornchainavakul N, Kikuti M, Paploski IAD, Makau DN, Rovira A, Corzo CA, VanderWaal K. 2022. Measuring how recombination re-shapes the evolutionary history of PRRSV-2: a genome-based phylodynamic analysis of the emergence of a novel PRRSV-2 variant. Front Vet Sci 9:846904. doi: 10.3389/fvets.2022.846904 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. VanderWaal K, Baker J, Pamornchainavakul N, Corzo C, Holtkamp D, Rovira A, Paploski I. PRRSV sub-types: what’s new in the lineage system? AASV Pre-Conference Seminars, Perry, Iowa, USA. doi: 10.54846/am2023/s2-11 [DOI] [Google Scholar]
  • 30. Ruedas-Torres I, Rodríguez-Gómez IM, Sánchez-Carvajal JM, Larenas-Muñoz F, Pallarés FJ, Carrasco L, Gómez-Laguna J. 2021. The jigsaw of PRRSV virulence. Vet Microbiol 260:109168. doi: 10.1016/j.vetmic.2021.109168 [DOI] [PubMed] [Google Scholar]
  • 31. Kikuti M, Sanhueza J, Vilalta C, Paploski IAD, VanderWaal K, Corzo CA. 2021. Porcine reproductive and respiratory syndrome virus 2 (PRRSV-2) genetic diversity and occurrence of wild type and vaccine-like strains in the United States swine industry. PLoS One 16:e0259531. doi: 10.1371/journal.pone.0259531 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Young C, Meng S, Moshiri N. 2022. An evaluation of phylogenetic workflows in viral molecular epidemiology. Viruses 14:774. doi: 10.3390/v14040774 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Katoh K, Standley DM. 2013. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30:772–780. doi: 10.1093/molbev/mst010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, Lanfear R. 2020. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic Era. Mol Biol Evol 37:1530–1534. doi: 10.1093/molbev/msaa015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Yu GC, Smith DK, Zhu HC, Guan Y, Lam TTY. 2017. GGTREE: an R package for visualization and annotation of phylogenetic trees with their covariates and other associated data. Methods Ecol Evol 8:28–36. doi: 10.1111/2041-210X.12628 [DOI] [Google Scholar]
  • 36. Balaban M, Moshiri N, Mai U, Jia X, Mirarab S. 2019. TreeCluster: Clustering biological sequences using phylogenetic trees. PLoS One 14:e0221068. doi: 10.1371/journal.pone.0221068 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. O’Toole Á, Scher E, Underwood A, Jackson B, Hill V, McCrone JT, Colquhoun R, Ruis C, Abu-Dahab K, Taylor B, Yeats C, du Plessis L, Maloney D, Medd N, Attwood SW, Aanensen DM, Holmes EC, Pybus OG, Rambaut A. 2021. Assignment of epidemiological lineages in an emerging pandemic using the pangolin tool. Virus Evol 7:veab064. doi: 10.1093/ve/veab064 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Kuhn M. 2008. Building predictive models in R using the caret package. J Stat Softw 28:26. doi: 10.18637/jss.v028.i05 [DOI] [Google Scholar]
  • 39. Trevisan G, Linhares LCM, Crim B, Dubey P, Schwartz KJ, Burrough ER, Main RG, Sundberg P, Thurn M, Lages PTF, Corzo CA, Torrison J, Henningson J, Herrman E, Hanzlicek GA, Raghavan R, Marthaler D, Greseth J, Clement T, Christopher-Hennings J, Linhares DCL. 2019. Macroepidemiological aspects of porcine reproductive and respiratory syndrome virus detection by major United States veterinary diagnostic laboratories over time, age group, and specimen. PLoS One 14:e0223544. doi: 10.1371/journal.pone.0223544 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Trevisan G, Sharma A, Gauger P, Harmon KM, Zhang JQ, Main R, Zeller M, Linhares LCM, Linhares DCL. 2021. PRRSV2 genetic diversity defined by RFLP patterns in the United States from 2007 to 2019. J Vet Diagn Invest 33:920–931. doi: 10.1177/10406387211027221 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Anderson TK, Inderski B, Diel DG, Hause BM, Porter EG, Clement T, Nelson EA, Bai J, Christopher-Hennings J, Gauger PC, Zhang J, Harmon KM, Main R, Lager KM, Faaberg KS. 2021. The United States Swine Pathogen Database: integrating veterinary diagnostic laboratory sequence data to monitor emerging pathogens of swine. Database (Oxford) 2021:baab078. doi: 10.1093/database/baab078 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental material. msphere.00709-24-s0001.pdf.

Supplemental text and figures.

DOI: 10.1128/msphere.00709-24.SuF1

Data Availability Statement

Reference sequences for the findings of this paper may be available upon reasonable request to the corresponding author. The data are not publicly available as they are part of diagnostic data from third parties (companies and veterinarians submitting samples for diagnosis).


Articles from mSphere are provided here courtesy of American Society for Microbiology (ASM)

RESOURCES