Abstract
Motivation
Porcine Reproductive and Respiratory Syndrome Virus (PRRSV) is a rapidly evolving RNA virus causing significant economic losses, posing a formidable challenge to vaccine efficacy due to its high mutational variability and immune escape. As the viral mutants evolve, their ability to sustain in population is driven by a range of host biology factors such as receptor binding, fusion, and uncoating. Existing tools that predict viral fitness and escape propensities rely heavily on extensive, up-to-date sequence data and lack integration of biochemical host interactions, limiting mechanistic understanding of the mutational landscape. We introduce Esca, a sequence-only toolchain framework that identifies immune escape-prone residues by exhaustively scanning each residue position for all amino acid substitutions using a Bayesian Variational Autoencoder (VAE) trained on protein language model embeddings. We demonstrate Esca on the GP5(ORF5) glycoprotein of PRRSV (EscaPRRS-ORF5) by training on ESM-2 embeddings of 32 146 GP5 sequences (2015–2022) spanning 140 sub-lineages.
Results
Despite being trained only on GP5 sequence data, EscaPRRS-ORF5 recovered 85.7% of the surface-exposed receptor binding interfaces as escape-prone regions. We use a mutation-sensitive fitness scoring scheme that goes beyond Hamming distances, to predict antibody escape tendencies, supporting surveillance of (re) emerging PRRSV variants. We do not claim that ORF5 alone captures PRRSV evolution or serves as a surveillance endpoint; rather, Esca offers a scalable path toward whole-genome, structure-aware surveillance.
Availability and implementation
EscaPRRS-ORF5 is freely available at https://doi.org/10.6084/m9.figshare.32661033 with an interactive Colab notebook at https://colab.research.google.com/drive/1TEgzAhPwvNAZ01VXeJbIFibfri2jnDA5? usp=sharing.
1 Introduction
With over $1.2 billion annual losses in the United States, Porcine Reproductive and Respiratory Virus (PRRSV) poses a significant economic challenge for the swine industry, with implications to food security and human health. Vaccine design efforts for PRRSV have consistently been challenged by the virus’s complex and tightly controlled infection mechanisms. Traditional approaches that identify antiviral targets via directed evolution are often insufficient to keep pace with its extensive plasticity (Baker et al. 2025). Thus, semi-rational designs aware of host-antigen interactions and dynamic, disordered regions of viral proteins are suggested. The emergence of protein language model-based tools (Cheng et al. 2023, Ito et al. 2025), which implicitly capture molecular interaction landscapes, enables predictions of epidemiological fitness propensities of viral variants.
In metazoan cells, antibodies neutralize pathogens by blocking viral entry, propagation, and immune evasion, often by binding to Receptor Binding Domains (RBDs) and interfering with host-virus interactions (Chan et al. 2025). However, viruses evade these defenses through mutations in surface glycoproteins that alter charge distribution and structural complementarity. Consequently, immune escape undermines vaccine efficacy, as evolving viral surfaces diminish antibody binding to designed immunogens (Li et al. 2025).
Contemporary PRRSV is classified into Betaarterivirus europense (PRRSV-1) and Betaarterivirus americanese (PRRSV-2), colloquially referred to as the European PRRSV and American PRRSV respectively (ICTV Report), reflecting substantial genetic divergence. Although PRRSV-1 is present in North America, it accounts for ≤1% of annual detections since 2015 based on Swine Disease Reporting System data (Trevisan et al. 2020), and remains associated with lower clinical impact. In contrast, PRRSV-2 comprises > 99% of PRRSV in the U.S and is responsible for significant disease burden and mortality, including the emergence “highly pathogenic” (HP) strains affecting southeast Asia and China (Tong et al. 2007). The standard control strategy for PRRSV is to use modified live vaccines that contain an attenuated virus. In the United States there are currently 6 licensed PRRSV-2 vaccine products available (Rawal et al. 2022, Li et al. 2024). While PRRSV-1 is a potential future concern, there are no currently approved vaccines available, due to the low detection rate and lesser clinical signs presented by this species. Accordingly, this study focusses on ORF5-linked fitness scoring of PRRSV-2, given its endemic persistence and high pathogenicity.
PRRSV enters host cells through a multistep process mediated by interactions between viral glycoproteins and receptors on porcine alveolar macrophages (PAMs) (Jiang et al. 2025). Its envelope contains two main glycoprotein complexes: the GP5-M heterodimer, responsible for receptor binding, and the GP2-GP3-GP4 trimer, which stabilizes the virion and enhances infectivity (Gao and Wen 2025). Viral entry occurs through receptor mediated endocytosis, involving binding, clustering of receptor-antigen complexes and membrane reorganization (Wei et al. 2020). Cryo-electron microscopy shows PRRSV virions as 55 nm spherical particles with small ectodomains (2 nm) of GP5-M proteins.
Of the ten open reading frames (ORFs), ORF1a encodes the polyprotein 1a, which is processed into 14 non-structural proteins and four proteinases, while ORF2-ORF7 encode structural proteins including GP2-GP5, membrane proteins and the nucleocapsid (Meng 2000). ORF5 is the most studied region of PRRSV and is seen to be the most variable region that harbors antibody neutralizing epitopes (Jian et al. 2025, VanderWaal et al. 2025, Cotaquispe Nalvarte et al. 2026). This 200-residue GP5 glycoprotein is known to affect neutralization of antibody targets, impairing cross-protection by immune responses raised against different strains (Wei et al. 2012). While GP5 solely does not govern viral tropism, available literature (Gao and Wen 2025) on porcine receptor interactions to PRRSV-2 indicated the GP5 mediates weak attachment to Heparansulfate that possibly positions the virus for subsequent high affinity binding with Sialoadhesin where the proteoglycans attach to the M-GP5 complex. Additionally, CD163 receptor is widely reported to be involved in uncoating viral particles to activate the infector pathway. Nontarget cells expressing both Sialoadhesin and CD163 are found more susceptible than those that express only CD163 (Van Gorp et al. 2008). In-vitro studies demonstrate that the expression of CD163 alone is sufficient for infection and replication as CD163 is solely responsible for disassembly of the viral envelope and capsid, causing conformational changes in GP2-GP4 that destabilizes the membrane-nucleocapsid interface (Xu et al. 2022, Zhu et al. 2023). Complete knockout of CD163 in porcine gestational pigs proved resistant to PRRSV infection with no fetal pathology while WT fetuses were up to 92% PRRSV positive. This is because CD163 blocks the infection pathway in maternal macrophages that normally carry the virus across the placenta (Prather et al. 2017). Several studies report improved resistance across all PRRSV genotypes in CD163 knockout pigs (Yang et al. 2018, Guo et al. 2019). Hence, a high affinity binding of one or more envelope proteins of PRRSV to porcine receptors is indicated to govern the overall tropism of PRRSV-2.
Our previous work (Dey et al. 2024) on PRRSV epitope atlas serves as a roadmap for structural characterization of host-pathogen interactions beyond the CD163 receptor. Given the immunodominance and abundant availability of GP5 sequences over other viral glycoproteins, we analyze all possible single point mutants that potentially obscure neutralizing epitopes, alter local charge distribution and sterically block antibodies to confer viral escape. To this end, we have developed a structurally guided escape-variant predictor CTRL-V and demonstrated its efficacy in predicting mutations observed in SARS-CoV-2 variants of concern starting with the parental sequence for which the receptor binding domain was co-crystallized with the human ACE2 protein and commercial Ly-CoV1404 antibodies (Teoh et al. 2025).
Anticipating escape tendencies using Machine Learning (ML) based computational tools is viewed as a promising approach to identify and predict pathogenic variants (Chen et al. 2021, Teoh et al. 2025). The EVE model (Frazer et al. 2021) uses a Bayesian Variational Autoencoder (VAE) on the homologous protein sequences, represented as one-hot encodings to capture evolutionary constraints that maintain protein fitness. Thus, variants with a lower modeled probability are predicted as less fit (deleterious). EVEscape (Thadani et al. 2023) extends these fitness predictions by integrating them with structure and biophysical constraints to score escape propensities in SARS-CoV-2 spike protein. The integration of protein language models into genomic surveillance pipelines has proven useful for real-time monitoring and characterization of emerging variants(Lytras et al. 2025). This work (Esca) fundamentally expands the scope and power of EVEscape scoring capabilities to predict critical future variants by replacing the one hot amino acid encoding with state-of-the-art ESM-2 (https://huggingface.co/facebook/esm2_t33_650M_UR50D) residue embeddings thus providing a structural basis of interpreting these variants through receptor interactions. ESM residue embeddings carry pre-trained evolutionary patterns from millions of protein sequences and even capture spatial relationships between amino acid pairs. This allows us to accurately assess mutational effects to pinpoint functionally important loci for vaccine targets without the need for explicit sequence alignments and experimentally resolved structures.
Here, we demonstrate Esca (Escape propensity prediction) for PRRSV-2 variants (EscaPRRS), a structure-aware fitness and escape prediction model, trained on ESM-2 embeddings using a Bayesian VAE. We demonstrate EscaPRRS’s fidelity for a robust understanding of GP5 glycoprotein, by training on 32 146 PRRSV-2 GP5 protein sequences (encoded by ORF5 genes) from the United States from 2015 to 2022, referred to as EscaPRRS-ORF5 hereon. Although trained on local sequences, ESM embeddings capture structural and evolutionary context, generalizing the validity of EscaPRRS-ORF5 scores for all (5291) publicly available PRRSV-2 GP5 sequences. We demonstrate the biological relevance of EscaPRRS-ORF5 antibody escape scores for PRRSV GP5 mutants, providing insights to competitive relative binding affinities at interfacial hotspots for seven different porcine receptor proteins. These receptors include: CD151, CD209, CD163, Sialoadhesin, Heparansulfate sulfotransferase, Vimentin, and MYH-9, identified through our epitope atlas mapping endeavor (Dey et al. 2024). EscaPRRS-ORF5 residue level escape score predictions accurately capture evolutionary patterns and pinpointed 85.7% of the surface exposed receptor-binding interfaces as escape-prone regions, while preserving seasonal and lineage-specific patterns arising from population sampling. In recognizing the limitations inherent in conventional fitness models that rely solely on sequence similarity patterns, we underscore the value of integrating protein language model embeddings to capture both evolutionary and structural signals that affect the fitness landscape of complete sequences of emerging variants.
2 Methodology: Esca scoring framework
Esca combines evolutionary information with local structural and physiochemical changes to score antibody escape propensity across the complete mutational landscape of any viral protein. Figure 1 illustrates the structure-function basis of EscaPRRS antibody escape propensity scoring at the host-receptor-antibody interface, highlighting its ability to complement experimental efforts that typically assay only a limited subset of PRRSV GP5 variants against available monoclonal antibodies.
Figure 1.

Antibody escape scoring for a single point PRRSV mutant with EscaPRRS-ORF5. The current implementation employs a multiplicative combination of the individual component scores. The genetic construct of GP5 is indicated as ORF5. Illustrations adapted from NIAID NIH BioArt Source #465 and #652.
Sequence distributions under natural selection have been modeled (Pugh et al. 2025) to obey a Boltzmann distribution profile, where the log-likelihood of a sequence is proportional to its evolutionary fitness, providing an unbiased zero-shot proxy for mutational tolerance. In EscaPRRS, fitness is estimated using a Bayesian VAE framework adapted from EVE. Instead of one-hot encodings, input sequences are represented using ESM-2 embeddings, which capture contextual and evolutionary features. These embeddings are processed through a probabilistic latent space, where the encoder is trained to minimize the evidence lower bound (ELBO) with a Kullback–Leibler (KL) divergence penalty to enforce consistency with the prior distribution.
To further enhance biological realism, EscaPRRS incorporates a 1D convolutional layer in its output, capturing local dependencies across sequence positions. This improves upon the positionally independent one-hot encoding scheme used in EVE (Fig. S1, available as supplementary data at Bioinformatics online), enabling better generalization to unseen sequences with conserved structural motifs. The sensitivity of ESM-2 embeddings to single-point mutations supports downstream analyses, such as ELBO-based evolutionary indices and identification of functionally important residues at virus–host interfaces.
Antibody escape is governed by three complementary requirements: (i) the variant must remain evolutionarily viable, (ii) the mutation site must be structurally accessible to antibodies, and (iii) the mutation must induce sufficient physicochemical change to disrupt binding. Accordingly, EscaPRRS models escape propensity using three corresponding components: fitness, surface accessibility, and chemical dissimilarity.
Surface accessibility is quantified using the weighted contact number (WCN) (Lin et al. 2008), defined as the sum of inverse squared distances between all residue pairs in a protein structure. Distances are computed between centers of mass of residue side chains, ensuring that nearby residues contribute more strongly. Lower local packing density corresponds to higher surface exposure, making WCN an effective proxy for identifying residues accessible to antibody binding.
Chemical dissimilarity is computed using a charge–hydrophobicity metric, following the EVEscape (Thadani et al. 2023) paradigm, based on differences in Eisenberg–Weiss hydrophobicity and residue charge. This captures the extent to which a mutation perturbs the local physicochemical environment at potential antibody binding sites.
These three components are transformed into probability distributions using Boltzmann-like scaling (SoftMax normalization) with temperature parameters controlling sensitivity. The resulting probabilities—fitness (P1) accessibility (P2), and dissimilarity (P3) are combined multiplicatively to yield the overall escape score.
This multiplicative formulation represents a joint likelihood in which the three components act as complementary and conditionally independent requirements for immune escape. It enforces an AND-like interaction, ensuring that mutations must be simultaneously viable, exposed, and antigenically novel to achieve high escape propensity. Consequently, a low probability in any single component proportionally suppresses the overall score, resulting in robust, biologically consistent, and interpretable predictions across complex mutational landscapes.
As shown in Fig. 2, the overall escape score is calculated as a product of fitness, accessibility and dissimilarity scores with suitable logistic temperature scaling factors , and referred to as escape factors as shown in Equation 1.
Figure 2.

The EscaPRRS-ORF5 (generally Esca) escape propensity scoring of emerging sequences incorporating fitness, surface accessibility and charge dissimilarity. ESM-2 embeddings from PRRSV GP5 sequences are trained on a Bayesian Autoencoder that captures variational inference from sequence distributions as in EVE30, to estimate evolutionary fitness of all single point mutants from the wild type. These fitness calculations are enriched with surface avidity and dissimilarity metrics to predict relative escape propensities. The single point mutation effects are aggregated to score emerging variants and sequence constructs for escape propensities.
| (1) |
| (2) |
Where i and denotes the temperature (scaling) parameter associated with each escape component i, (1 = fitness, 2 = accessibility, and 3 = dissimilarity) which modulates the sensitivity of the calculated probability to the corresponding scorei that represents the raw score calculated for that component. The denominator ensures the probabilities are normalized to sum to unity. Details on training dataset curation and autoencoder training are provided in the Supplementary Methods S1, available as supplementary data at Bioinformatics online.
3 Results and discussion
3.1. Complete fitness landscape for PRRSV GP5 mutants using EscaPRRS-ORF5
A major component of an escape prediction framework is to demarcate relative fitness of circulating variants as soon as they emerge, to aid parallel efforts on escape-proof vaccine design, as demonstrated in our CTRL-V (Teoh et al. 2025) effort. In other RNA viruses like in HIV and SARS-Cov-2, the replicative fitness of a variant in the sequence population can be compromised in high antibody escape variants but its fitness is reestablished by means of structural and functional changes at the secondary residues not directly linked to primary escape mutation sites (Lynch et al. 2015, Teruel et al. 2025). Thus, capturing the fitness landscape beyond phylogenetic context is important to identify critical structural domains conferring immune escape and those enabling fitness.
ESM-2 captures evolutionary patterns by storing coevolutionary statistics of interacting sequence parameters and pairwise residue dependencies, implicitly learning local motifs and their separation distances (Zhang et al. 2024). The Bayesian Variational Autoencoder trained on ESM-2 embeddings for 32 146 PRRSV-2 sequences under a fixed hyperparameter configuration (described in Table S1, available as supplementary data at Bioinformatics online) was used to calculate the departure of the resultant sequence distribution for all single point mutants from the reference sequence. This is used to calculate the fitness as the log likelihood of every mutant, visualized in Fig. 3, panel(a). The fitness values are then parsed into a two-component Gaussian Mixture Model (GMM) to classify into low and high ORF5-linked fitness. The posterior fitness probabilities are computed with quantified uncertainty values for downstream analyses.
Figure 3.

(a) EscaPRRS-ORF5 fitness calculation for all possible single point mutants to a consensus PRSSV2-GP5 reference sequence captures specific residue position and choices that compromise the evolutionary fitness of the resulting single point mutant. This is calculated from the latent space trained on 32 146 sequences. (b) IQ-TREE Maximum likelihood phylogenetic tree of the sequences used in this study annotated with nine lineages of PRRSV-2 and the percentage of sequences within each lineage. (c) Population occurrence trends in PRRSV-2 lineages publicly reported by VanderWaal et al. (d) t-SNE visualization of ESM-2 embeddings derived from 32 146 protein sequences in the training dataset color coded based on DBSCAN (Density-based spatial clustering of applications with noise) cluster numbers (e) UMAP (Uniform Manifold Approximation and Projection) visualization of the latent space learned by the VAE on EscaPRRS-ORF5 capturing smoothened sequence variability. (f) 2-component Gaussian Mixture Model (GMM) fit to the evolutionary index distribution to allow probabilistic separation between ORF5 fitness classes.
It is important for a mutation to be fit over its evolutionary counterparts to survive in time and for its surface to be available for binding to the host receptor, and the antibodies. From our dataset, we observed 9 sequences consistently re-emerging 100–870 times from 2015 to 2022, implying superior replicative fitness, transmission potential and immune evasion under shared selective pressures. This is supported by low gamma shape parameter of the maximum likelihood phylogenetic tree () (Fig. 3, panel (b), Supplementary Methods S2, available as supplementary data at Bioinformatics online), showing a prevalence of high ORF5-fitness variants that have nearly zero evolution rates, with a few hotspots driving rapid evolution under selection. We can also see a higher number of high ORF5-fitness variants over time publicly reported by VanderWaal et al. (2025) (Fig. 3, panel (c)).
Thus, capturing the fitness landscape beyond phylogenetic context is important to identify critical structural domains conferring immune escape and those enabling fitness. ESM-2 captures evolutionary patterns by storing coevolutionary statistics of interacting sequence parameters and pairwise residue dependencies, implicitly learning local motifs and their separation distances (Zhang et al. 2024). The Bayesian Variational Autoencoder trained on ESM-2 embeddings for 32 146 PRRSV-2 sequences under a fixed hyperparameter configuration (described in Supplementary Table S1, available as supplementary data at Bioinformatics online) was used to calculate the departure of the resultant sequence distribution for all single point mutants from the reference sequence. This is used to calculate the fitness as the log likelihood of every mutant, visualized in Fig. 3, panel(a). The fitness values are then parsed into a two-component Gaussian Mixture Model (GMM) to classify into low and high ORF5-linked fitness. The posterior fitness probabilities are computed with quantified uncertainty values for downstream analyses.
It is important for a mutation to be fit over its evolutionary counterparts to survive in time and for its surface to be available for binding to the host receptor, and the antibodies. From our dataset, we observed 9 sequences consistently re-emerge 100–870 times from 2015 to 2021, implying superior replicative fitness, transmission potential and immune evasion under shared selective pressures. This is supported by low gamma shape parameter of the maximum likelihood phylogenetic tree () (Fig. 3, panel (b), Supplementary Methods S2, available as supplementary data at Bioinformatics online), showing a prevalence of high ORF5-fitness variants that have nearly zero evolution rates, with a few hotspots driving rapid evolution under selection. We can also see a higher number of high ORF5-fitness variants over time publicly reported by VanderWaal et al. (2025) (Fig. 3, panel (c)).
3.2. EscaPRRS-ORF5 for antigenic escape landscape
While phylogeny can influence apparent virulence of co-evolving strains, evolutionary classification is only a partial proxy for virulence and genetically homologous variants do not directly translate to immunological cross protection (Geoghegan and Holmes 2018). To this end, antibody escape propensity predictions offer a powerful refinement of the antigenic dimension to identify variants of concern (VOCs), forecast mutations with high escape potential from monoclonal and polyclonal sera, and monitor the combinatorial fitness landscape (Starr et al. 2021).
For evolutionary fitter mutants, immune escape is conferred by the change in the local structure and charge distribution that render the interfacial surface incompatible/unstable for antibody binding. As shown in Equation 2, the EscaPRRS protocol integrates these features to calculate the antibody escape propensity score using a combined charged-hydrophobicity metric and the surface accessibility. Figure 4, panel (a) shows the standardized distribution of each of the three escape conferring factors for the complete mutational landscape of PRRSV-GP5.
Figure 4.

(a) Comparison of standardized distribution of ORF5-fitness, surface accessibility score, combined charge-hydrophobicity metric for the complete single point mutational landscape for PRRSV-GP5. (b) Overlay of AlphaFold3 predicted high confidence structures (pTM > 0.5) for GP5 glycoprotein of PRRSV-2 across different lineages (L1A-L1F, L5A, L8C) observed in the training dataset, indicating a structurally conserved backbone (Relative RMSD < 1.5 Å) used to calculate weighted contact number for each residue position. (c) Ablation analysis of EscaPRRS-ORF5 showing how per-factor temperature scaling affects the self-consistency of the combined EscaPRRS score. (d) Line plot of mean and variable EscaPRRS scores along the sequence length of PRRSV-GP5. (e) Heatmap of calculated EscaPRRS-ORF5 antibody escape propensities for all possible single point mutations from the reference PRRSV-GP5 sequence. Red regions indicate high escape propensity while blue regions indicate lower escape propensity.
We use the AlphaFold3 (Abramson et al. 2024) generated backbone structure of the consensus PRRSV-GP5 sequence to calculate the weighted contact number for each residue as a scalable surface accessibility metric, owing to the highly conserved backbone structure across different PRRSV lineage sequences, as point mutations do not drastically influence the global fold of the glycoprotein, seen in Fig. 4, panel (b).
The escape conferring factors including evolutionary fitness, surface accessibility and charge dissimilarity are subjected to ablation analysis, which helps adjust the inherent trade-off between evolutionary plausibility and immune evasion, which is uniquely distributed for each viral protein based on its structural rigidity and evolutionary plasticity. Youssef et al. (2024) demonstrated hyperparameter searches to optimize escape factors by varying the temperature across a range of values and reported the resulting AUROC values for different viral protein families including flu, HIV, and SARS-Cov-2. Higher AUROC values corresponded to the efficacy of the escape factor setting in the model to predict true escape mutations while avoiding false positives.
In our setting, because no experimental escape labels were available for PRRSV GP5, ablation was performed in an unsupervised manner, and the reported AUROC values quantify self-consistency with a baseline EscaPRRS ranking, rather than accuracy against ground truth. Our ablation results, as seen in Fig. 4, panel (c) revealed that structural accessibility reflecting solvent exposure of residues, is the most critical factor in EscaPRRS-ORF5 performance. Adjusting the temperature scaling for accessibility resulted in a 38.71% performance decrease, underscoring its strong sensitivity. In contrast, the fitness feature (mutational tolerance) showed a 30.75% increased performance, and dissimilarity (biochemical divergence from wild type) had the smallest effect with 13.66% improvement. Furthermore, the R2 score decreased most significantly to 0.1717 upon removal of the accessibility feature, confirming its essential role in PRRSV GP5 mutational escape prediction. The minimal impact observed when removing fitness (R2 = 0.9006) and dissimilarity (R2 = 0.9042) suggests that structural features predominantly govern escape tendencies, indicating that immune escape mutations are primarily located at structurally accessible sites, likely antibody epitopes, rather than being driven by selection for enhanced receptor binding or altered host mimicry. This is analogous to hypermutability of CDR3 surface exposed loops of antibodies which are constantly under selection pressure to bind to a rapidly evolving antigen (Chowdhury et al. 2018, Boorla et al. 2023).
Inspecting unsupervised temperature-response ablation curves, the fitness plateaus beyond a scaling of 10, while accessibility feature drops sharply after 0.5 and dissimilarity peaks around 2–5. Thus, pushing temperature scaling factors beyond these regions yields diminishing returns and can induce numerical saturation. Since the default escape factors (in EVEscape) were initially set to 1, 1, and 2, we performed a targeted adjustment of temperatures in EscaPRRS-ORF5 based on this ablation analysis with an unsupervised objective that favors stability and entropy for a balanced scoring of mutational escape in EscaPRRS-ORF5. Figure 4, panel (d) shows the average EscaPRRS-ORF5 scores for each residue position in the PRRSV-GP5 sequence, with panel (e) showing the complete mutational landscape for all single point mutants colored based on EscaPRRS-ORF5 scores calculated at the adjusted escape scaling factors (10,2,1.5 for fitness, accessibility, and dissimilarity respectively). Some of the low escape (blue-hued) regions are reported in the literature as GP5 residues prone to antibody neutralization, including N51, N29, N32, N44 (Vu et al. 2011). Positions 37 through 44 are reported to induce neutralizing antibody in addition to V102 (Han et al. 2019).
3.3. Immune escape propensities at top receptor binding hotspots
The normalized mutation-level EscaPRRS-ORF5 scores grouped by residue positions were used to calculate site average antibody binding escape propensities. Our previous work (Dey et al. 2024) annotated binding interface residues of all glycoproteins, membrane envelope proteins and non-structural proteins to seven different porcine receptors consistent across different docking platforms including GRAMMDock (Singh et al. 2024), HADDOCK3 (Giulini et al. 2025) and ClusPro (Kozakov et al. 2017) to characterize the epitope atlas of PRRSV-2.
These docked complexes of seven receptors with GP5, namely CD151, CD209, CD163, Sialoadhesin, Heparansulfate sulfotransferase, and MYH-9 indicated high EscaPRRS-ORF5 escape propensity regions interface with critical binding hotspots with porcine receptors, as seen in Fig. 5.
Figure 5.

3D structure of PRRSV GP5 annotated with residues known to interact with seven porcine receptor proteins (yellow: CD151, cyan: CD209, purple: CD163, magenta: vimentin, peach: Sialoadhesin, green: Heparansulfate sulfotransferase, violet: MYH-9). GP5 Residue colors indicate EscaPRRS-ORF5 scores where red indicates higher escape propensity and blue corresponds to lower escape propensity.
This underscores the biological fidelity of EscaPRRS as the model is not trained in explicit information on receptor or antibody binding. This correlation between interfacial maximum escape propensity sites and HADDOCK binding affinity trends with the seven receptors highlights potential antibody target sites. A comparison of binding affinity scores with the maximum site average scores with the interaction residues indicated a Pearson Correlation Coefficient (PCC) of 0.74 with a P value of 0.0569 (Fig. 6). We observed that high escape propensity residues recover 85.7% of the interfacial residues of GP5 with the seven porcine receptors considered. This result supports the hypothesis that immune evasion by steric effects, rather than improved receptor usage is the dominant selection pressure driving the evolution and emergence of pathogenic variants.
Figure 6.

(a) Structure free binding affinity calculation using TwinPeaks (Dey et al. 2024) for GP5 complexes with seven porcine receptors across 13 different lineages of PRRSV-2 classified based on ORF5. (b) Correlation between maximum EscaPRRS-ORF5 antibody escape scores at the interfacial residues with HADDDOCK3 binding affinities for the reference GP5 sequence-receptor complexes. Residue-wise scores at the interfacial residues are provided in Table S2, available as supplementary data at Bioinformatics online. (c) Sequence-only binding affinity prediction for six unique sequences belonging to different ESM-2 embedding space t-SNE clusters on seven porcine receptors reveal similar trends with GP5 showing a strong binding affinity for CD151 and weaker binding to Heparansulfate sulfotransferase, as described in the introduction. Binding affinity scores for individual GP5-sequence receptor pairs are given in Table S3, available as supplementary data at Bioinformatics online.
For such viral-host interactions, scaling binding affinity predictions are challenging as most tools often rely on structure predicting and geometry governed docking, wherein longer sequence length results in low confidence/disordered structures. To overcome this, we use our sequence-only binding affinity predictor tool, Seq2Bind (TwinPeaks module; tracker) (Dey and Chowdhury 2025, Ma et al. 2025) that maps protein language embeddings of the protein sequences to a latent space was used. We learnt and predicted GP5 variant affinity score () trends and interfacial hotspot residues with all seven porcine receptors. A more negative binding affinity score () is correlated with stronger binding between the viral protein and the porcine receptor. The ability to score binding affinities and antibody escape propensities at scale with Seq2Bind (Dey and Chowdhury 2025, Ma et al. 2025) offers a promising opportunity to incorporate EscaPRRS for concerted testing and surveillance of emerging/re-emerging strains and design constructs for immunogenicity.
It is important to note that antibody escape and receptor binding are independent events and are dictated by a range of host biology factors and non-modeled events like population-level effects, glycosylation, and host cell dynamics. This calls for integrative metabolic modeling in host systems incorporating viral entry, enzyme activity, metabolite transport, and molecular recognition, as outlined in our review (Noor et al. 2025).
3.4. Incorporating population and mutational dynamics for fitness prediction of variant sequences
Given the complete mutational landscape, fitness prediction of complete sequences involves aggregating the effects of single point mutations. We infer that mere aggregation of the escape scores for the single point mutations that govern the fitness landscape to score a full sequence, as employed in EVEscape (Thadani et al. 2023), resulted in overdependency of the score on the number of mutations, which dilutes the effective predictive performance of the VAE in annotating fitness of any sequence from the latent space (Fig. 7, panel a). While traditional one hot encoding cannot represent a full sequence in a single embedding dimension, the use of protein language model embeddings allows a motif-informed representation of any sequence which can be conveyed and projected into the latent space for fitness predictions by ELBO difference relative to the reference sequence which captures the periodic and seasonal trends in variant fitness while agreeing with the sub-lineage occurrence trends reported in the PRRSLoom-Variants (n.d.) server (https://stemma.shinyapps.io/PRRSLoom-variants/) (Fig. 7, panel (b) and (c)). It is important to note that the lineage classification and RFLP typing are based solely on ORF5 and therefore do not fully explain the occurrence dynamics of variants in the population.
Figure 7.

(a) Sigmoid sum of EscaPRRS-ORF5 scores of single point mutation in a sequence relative to reference calculated for all sequences in the local dataset over time, indicating the loss in capturing mutational variability beyond the sequence Hamming distance. (b) EscaPRRS-ORF5 fitness values calculated for each sequence from the ESM-2 embedding space learnt by the Bayesian VAE. (c) EscaPRRS-ORF5 fitness values colored based on population occurrence trends for the sub-lineages reported in PRRS-Loom webserver. Sub lineage wise temporal trends for the variants in the training dataset are provided in Supplementary file S3, available as supplementary data at Bioinformatics online. (d) Distribution of sequence edit distances versus fitness differences between two sequences classified into different lineages sampled across 4.18 billion inter-lineage pairs. (e) Density of mutation hotspots that drive lineage jumps for all sequences that have inter-lineage edit distances less than or equal to ten. (f) Lineage jump hotspots shown in (e) visualized on Alphafold-3 generated structure of PRRSV-GP5.
Since protein language model embeddings naturally capture higher order epistasis (Tang et al. 2026), our approach offers a significant leap over traditional methods that penalize fitness based on individual mutational effects to score full sequences for their epidemiologic fitness (Read et al. 2012) in highlighting the dynamic fitness landscapes in emerging/re-emerging variants within a sub-lineage.
ORF5-based variant classification (VanderWaal et al. 2025, Zeller et al. 2024) of PRRSV-2 is quite powerful to pinpoint dominating lineages when inferred with sampling intensities to annotate variants of concern (VOCs). However, phylogenic analysis efforts do not detail on charting the mutations that shift a sequence from a lineage of lower reported pathogenicity to that of a VOC. To address this, we performed inter-lineage analysis (Fig. 7, panel (d)) by sampling 4.18 billion random pairs of sequences, each belonging to a different lineage, and comparing their sequence edit distances and fitness changes. This pinpoints mutation regions that drive lineage jumps by a small number of mutations between the two sequences. After filtering inter-lineage pairs where the sequences differ by less than or equal to 10 mutations, we identified hotspot regions that drive lineage jumps (Tewawong et al. 2017), as visualized in Fig. 7, panel (e) and (f). This included a Lysine at position 59, a Glutamine at position 58, and an Isoleucine at position 121 as the top three spots seen to be mutated to potentially shift to the next closest lineages (LIB, LIC) from L1A.
Thus, EscaPRRS-ORF5 gives a tractable static fitness score for each genotype, to highlight high fitness variants, which is inherently tied to a competitive replication in a population. When new sequences appear through mutation, fitness is counteracted by population growth, and the compensatory evolution requires a dynamic tracking. To address this, we apply a basic logistic population growth ODE model on EscaPRRS-ORF5 fitness for all re-emerging sequences in our dataset to follow multiple genotypes and their frequencies, as detailed in Supplementary File S2, available as supplementary data at Bioinformatics online. This formulation shall be enhanced with experimental data on growth and clearance rates for each lineage for reliable quantitative modeling of time-varying selection regimes (Read et al. 2012).
4 Conclusion
With rapidly evolving PRRSV infection patterns that challenge the current vaccine design paradigms, there is an urgent need to score and monitor emerging and re-emerging strains to enable timely interventions and strengthen preparedness for effective protection. While lineage classifications purely based on sequence similarity and phylogeny can help detect the spread and introduction of newer strains, fitness prediction is crucial to flag specific mutants and clades with higher propensity to immune escape and transmission. To our knowledge, there are no structure-aware tools available to predict and inform the epidemiological fitness for PRRSV mutants. We present EscaPRRS to score point mutations, complete variant sequences, and design constructs. A fundamental version of EscaPRRS is demonstrated in this work with PRRSV ORF5 sequence data, where we encode ESM-2 protein language model embeddings to a latent space to infer fitness of point mutants and sequences of the GP5 glycoprotein encoded by ORF5. ESM-2 is chosen for its open-source capabilities for downstream embedding processing, and its versatility in representing sequence motifs and non-additive mutational patterns without being trained on explicit structural features, which allows capturing subtle local mutational differences without being smoothened by the global fold information. Consequently, the current version of EscaPRRS-ORF5 is limited to GP5 sequences of length 200 (as usually observed in PRRSV-2), other protein language models including ESM-c can be leveraged to handle the effect of insertions and deletions.
ORF5 constitutes less than 5% of the whole PRRSV-2 genome and is widely characterized owing to rapid mutational variability and its role in antibody binding. To account for the structural and biochemical changes at the mutation sites that evade antibody binding, position-wise escape propensity scores are predicted for the complete mutational landscape (Supplementary Methods S1, available as supplementary data at Bioinformatics online), which can be extended to the complete genome to help identify critical antibody escape hotspots. This calls for a complete genome based PRRSV-2 classification, which when combined with EscaPRRS will provide scope for mapping potential hotspots for recombination, and domain-wise receptor binding epitopes.
The key takeaway of this work is not to argue that ORF5 alone provides a complete representation of PRRSV evolutionary dynamics, nor to suggest that an ORF5-centered model constitutes an endpoint of genomic surveillance. Rather, we establish a structurally informed, evolution-aware Esca toolchain that converts lineage-resolved sequence variation into mechanistic fitness and escape insight, forming a reusable and extensible framework that can integrate metagenomic coupling datasets (Hodges et al. 2025) and epitope–receptor interactomes to enable scalable, whole-genome, structure-aware surveillance of PRRSV-2 (or other viruses).
Overall, EscaPRRS-ORF5 allows surveillance of PRRSV-2 GP5 variants for fitness and identification of hotspots for receptor binding and antibody escape without being explicitly trained on any antibody or receptor information. EscaPRRS reduces the reliance on deep sequence alignments to assess the evolutionary patterns of emerging/re-emerging variants and escape vulnerability of design constructs that are antigenically similar to the emerging sequences, without heavy dependence on three-dimensional structures of antibody-antigen complexes. As a result, PRRSV-2 GP5 variants of concern can be flagged in advance as they emerge, to aid comprehensive experimental assays. Subsequent goals include incorporating EscaPRRS for an inverse design toolchain for constructing prototypic libraries of de novo antibodies or immunogens, as demonstrated in our other work (Teoh et al. 2025).
Supplementary Material
Contributor Information
Ratul Chowdhury, Department of Chemical and Biological Engineering, Iowa State University, Ames, IA 50011, United States.
Vaishnavey SR, Department of Chemical and Biological Engineering, Iowa State University, Ames, IA 50011, United States.
Supantha Dey, Department of Chemical and Biological Engineering, Iowa State University, Ames, IA 50011, United States.
Sakib Ferdous, Department of Chemical and Biological Engineering, Iowa State University, Ames, IA 50011, United States.
Riza Danurdoro, Department of Chemical and Biological Engineering, Iowa State University, Ames, IA 50011, United States.
Michael A Zeller, Department of Veterinary Diagnostic and Production Animal Medicine, College of Veterinary Medicine, Iowa State University, Ames, IA 50011, United States.
Author contributions
Vaishnavey SR (Conceptualization [Equal], Data curation [Equal], Formal analysis [Lead], Investigation [Lead], Methodology [Equal], Software [Lead], Validation [Equal], Visualization [Lead], Writing—original draft [Lead], Writing—review & editing [Equal]), Supantha Dey (Formal analysis [Equal], Investigation [Equal], Writing—review & editing [Equal]), Sakib Ferdous (Methodology [Equal], Writing—review & editing [Equal]), Riza Danurdoro (Software [Lead], Writing—review & editing [Equal]), and Michael A. Zeller (Data curation [Lead], Resources [Lead], Validation [Equal], Writing—review & editing [Equal]), Ratul Chowdhury (Conceptualization [Lead], Funding acquisition [Lead], Methodology [Lead], Project administration [Lead], Supervision [Lead], Writing—review & editing [Equal])
Supplementary material
Supplementary material is available at Bioinformatics online.
Conflicts of interest
None declared.
Funding
This work is partially supported by the Iowa State University Startup Grant (Building a World of Difference Faculty Fellowship), Iowa State University Vice President of Research Seed Grant (Vaccines, Diagnostics, and Immunotherapeutics) for PRRSV study, and NSF 22–599, EPSCoR RII Track-1, Award Number 2242763 to R.C. V.S.R. acknowledges Dr. Monica H. Lamm, Professor, Department of Chemical and Biological Engineering, Iowa State University for reading and providing feedback on this manuscript.
Data availability
While the training and test sequences used in this work are of proprietary nature, the PRRSView (Zeller et al. 2022) server shall be referred for a phylogenetic overview of PRRSV-2 ORF5 GP5 sequences, including BLAST, RFLP scores, lineage and sub-lineage classification tools.
Code availability
EscaPRRS is freely accessible at https://agrivax.studio suite of tools and within the StructF tab. Scripts to calculate EscaPRRS-ORF5 fitness scores are available at the following Google colab link: https://colab.research.google.com/drive/1TEgzAhPwvNAZ01VXeJbIFibfri2jnDA5? usp=sharing. Training and inference scripts for Esca are available in our Hugging face repository at https://huggingface.co/chowdhury-lab/escaprrs_demo for a generalized workflow which can be implemented for any viral protein sequence data. Additionally, EscaPRRS scores calculated as aggregated sigmoid sum escape scores of point mutations can be accessed at the following Google Colab link: https://colab.research.google.com/drive/1qs48k65aDkxuxO0rYcCDxtZPHjYlVMrO? usp=sharing.
References
- Abramson J, Adler J, Dunger J et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024;630:493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Baker JP, Rovira A, VanderWaal K. Repeat offenders: PRRSV-2 clinical re-breaks from a whole genome perspective. Vet Microbiol 2025;302:110411. 10.1016/J.VETMIC.2025.110411. [DOI] [PubMed] [Google Scholar]
- Boorla VS, Chowdhury R, Ramasubramanian R et al. De novo design and Rosetta-based assessment of high-affinity antibody variable regions (Fv) against the SARS-CoV-2 spike receptor binding domain (RBD). Proteins 2023;91:196–208. 10.1002/PROT.26422. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chan AC, Martyn GD, Carter PJ. Fifty years of monoclonals: the past, present and future of antibody therapeutics. Nat Rev Immunol 2025; 25:745–65. 10.1038/s41577-025-01207-9. [DOI] [PubMed] [Google Scholar]
- Chen C, Boorla VS, Banerjee D et al. Computational prediction of the effect of amino acid changes on the binding affinity between SARS-CoV-2 spike RBD and human ACE2. Proc Natl Acad Sci USA 2021;118:e2106480118. 10.1073/pnas.2106480118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cheng J, Novati G, Pan J et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 2023;381:eadg7492. 10.1126/SCIENCE.ADG7492. [DOI] [PubMed] [Google Scholar]
- Chowdhury R, Allan MF, Maranas CD. OptMAVEn-2.0: de novo design of variable antibody regions against targeted antigen epitopes. Antibodies 2018;7:23. 10.3390/ANTIB7030023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cotaquispe Nalvarte R, Legua Barrios M, Escajadillo Luján P et al. Caracterización Molecular Del Gen ORF5 (GP5) Del Agente Causal Del Síndrome Respiratorio y Reproductivo Porcino (PRRS) Detectado En Granjas Porcinas de Lima, Perú. Rev Argent Microbiol 2026;58:23–30. 10.1016/J.RAM.2025.07.005. [DOI] [PubMed] [Google Scholar]
- Dey S, Bruner J, Brown M et al. Identification and biophysical characterization of epitope atlas of porcine reproductive and respiratory syndrome virus. Comput Struct Biotechnol J 2024;23:3348–57. 10.1016/J.CSBJ.2024.08.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dey S, Chowdhury R. 2025. Twin peaks: dual-head architecture for structure-free prediction of protein-protein binding affinity and mutation effects. https://arxiv.org/pdf/2509.22950, 18 December 2025, preprint: not peer reviewed.
- Frazer J, Notin P, Dias M et al. Disease variant prediction with deep generative models of evolutionary data. Nature 2021;599:91–5. 10.1038/s41586-021-04043-8. [DOI] [PubMed] [Google Scholar]
- Gao F, Wen G. Strategies and scheming: the war between PRRSV and host cells. Virol J 2025;22:191–21. 10.1186/S12985-025-02685-Y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Geoghegan JL, Holmes EC. The phylogenomics of evolving virus virulence. Nat Rev Genet 2018;19:756–69. 10.1038/s41576-018-0055-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giulini M, Reys V, Teixeira JM et al. "HADDOCK3: a modular and versatile platform for integrative modeling of biomolecular complexes." J Chem Inf Model 2025;65:7315–24. 10.1021/acs.jcim.5c00969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gorp HV, Van Breedam W, Delputte PL et al. Sialoadhesin and CD163 join forces during entry of the porcine reproductive and respiratory syndrome virus. J Gen Virol 2008;89:2943–53. 10.1099/VIR.0.2008/005009-0. [DOI] [PubMed] [Google Scholar]
- Guo C, Wang M, Zhu Z et al. Highly efficient generation of pigs harboring a partial deletion of the CD163 SRCR5 domain, which are fully resistant to porcine reproductive and respiratory syndrome virus 2 infection. Front Immunol 2019;10:1846. 10.3389/FIMMU.2019.01846. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Han G, Xu H, Wang K et al. Emergence of two different recombinant PRRSV strains with low neutralizing antibody susceptibility in China. Sci Rep 2019;9:2490. 10.1038/s41598-019-39059-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hodges AL, Walker LR, Everding T et al. Metagenomic detection and genome assembly of novel PRRSV-2 strain using oxford nanopore flongle flow cell. J Anim Sci 2025;103:skae395. 10.1093/jas/skae395. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ito J, Strange A, Liu W et al. A protein language model for exploring viral fitness landscapes. Nat Commun 2025;16:4236. 10.1038/s41467-025-59422-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jian Y, Lu C, Shi Y et al. Genetic evolution analysis of PRRSV ORF5 gene in five provinces of Northern China in 2024. BMC Vet Res 2025;21:242. 10.1186/S12917-025-04679-Y/FIGURES/5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiang B, Wei R, Wang X et al. Single-Cell RNA sequencing reveals immune response dynamics and infection mechanisms in PRRSV-Infected porcine alveolar macrophages. Int J Biol Macromol 2025;334:148719. 10.1016/J.IJBIOMAC.2025.148719. [DOI] [PubMed] [Google Scholar]
- Kozakov D, Hall DR, Xia B et al. The ClusPro web server for protein–protein docking. Nat Protoc 2017 12:2, 2017;12:255–78. 10.1038/nprot.2016.169. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li J, Miller LC, Sang Y. Current status of vaccines for porcine reproductive and respiratory syndrome: interferon response, immunological overview, and future prospects. Vaccines (Basel) 2024;12:606. 10.3390/VACCINES12060606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li Y, Yang S, Qian J et al. Molecular characteristics of the immune escape of coronavirus PEDV under the pressure of vaccine immunity. J Virol 2025;99:e0219324. 10.1128/JVI.02193-24/FORMAT/EPUB. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lin C-P, Huang S-W, Lai Y-L et al. Deriving protein dynamical properties from weighted protein contact number. Proteins 2008;72:929–35. 10.1002/prot.21983. [DOI] [PubMed] [Google Scholar]
- Lynch RM, Wong P, Tran L et al. HIV-1 fitness cost associated with escape from the VRC01 class of CD4 binding site neutralizing antibodies. J Virol 2015;89:4201–13. 10.1128/JVI.03608-14; WGROUP: STRING: PUBLICATION. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lytras S, Lamb KD, Ito J et al. Pathogen genomic surveillance and the AI revolution. J Virol 2025;99:e0160124. 10.1128/JVI.01601-24/FORMAT/EPUB. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma X, Dey S, Sr V et al. Seq2bind webserver for binding site prediction from sequences using fine-tuned protein language models. NAR Genom Bioinform 2025;7:lqaf154. 10.1093/nargab/lqaf154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meng XJ. Heterogeneity of porcine reproductive and respiratory syndrome virus: implications for current vaccine efficacy and future vaccine development. Vet Microbiol 2000;74:309–29. 10.1016/S0378-1135(00)00196-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noor MS, Ferdous S, Salehi R et al. Next-generation metabolic models informed by biomolecular simulations. Curr Opin Biotechnol 2025;92:103259. 10.1016/J.COPBIO.2025.103259. [DOI] [PubMed] [Google Scholar]
- Prather RS, Wells KD, Whitworth KM et al. Knockout of maternal CD163 protects fetuses from infection with porcine reproductive and respiratory syndrome virus (PRRSV). Sci Rep 2017;7:13371–5. 10.1038/S41598-017-13794-2; SUBJMETA. [DOI] [PMC free article] [PubMed] [Google Scholar]
- ‘PRRSLoom-Variants’. https://stemma.shinyapps.io/PRRSLoom-variants/, n.d. (29 August 2025, date last accessed).
- Pugh CW, Nuñez-Valencia PG, Dias M et al. From likelihood to fitness: improving variant effect prediction in protein and genome language models. Advances in Neural Information Processing Systems 2026;38:130835–66. [Google Scholar]
- Rawal G, Yim-Im W, Chamba F et al. Development and validation of a reverse transcription real-time PCR assay for specific detection of PRRSGard vaccine-like virus. Transbound Emerg Dis 2022;69:1212–26. 10.1111/TBED.14084. [DOI] [PubMed] [Google Scholar]
- Read EL, Tovo-Dwyer AA, Chakraborty AK. Stochastic effects are important in intrahost HIV evolution even when viral loads are high. Proc Natl Acad Sci USA 2012;109:19727–32. 10.1073/pnas.1206940109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Singh A, Copeland MM, Kundrotas PJ et al. GRAMM web server for protein docking. Methods Mol Biol 2024;2714:101–12. 10.1007/978-1-0716-3441-7_5. [DOI] [PubMed] [Google Scholar]
- Starr TN, Greaney AJ, Addetia A et al. Prospective mapping of viral mutations that escape antibodies used to treat COVID. Science 2021;371:850–4. 10.1126/science.abf9302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tang M, Cromie GA, Kabir A et al. Predicting epistasis across proteins by structural logic. Proc Natl Acad Sci USA 2026;123:e2516291123. 10.1073/pnas.2516291123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teoh YC, Noor MS, Aghakhani S et al. Viral escape-inspired framework for structure-guided dual bait protein biosensor design. PLoS Comput Biol 2025;21:e1012964. 10.1371/JOURNAL.PCBI.1012964. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teruel NFB, Crown M, Rajsbaum R et al. Comprehensive analysis of SARS-CoV-2 spike evolution: epitope classification and immune escape prediction. Virus Evol 2025;11:veaf027. 10.1093/VE/VEAF027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tewawong N, Suntronwong N, Korkong S et al. Evidence for influenza B virus lineage shifts and reassortants circulating in Thailand in 2014. Infect Genet Evol 2017;47:35–40. 10.1016/j.meegid.2016.11.010. [DOI] [PubMed] [Google Scholar]
- Thadani NN, Gurev S, Notin P et al. Learning from prepandemic data to forecast viral escape. Nature 2023;622:818–25. 10.1038/S41586-023-06617-0; SUBJMETA=114,2161,2397,250,631; KWRD=COMPUTATIONAL+MODELS, IMMUNE+EVASION. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tong G-Z, Zhou Y-J, Hao X-F et al. Highly pathogenic porcine reproductive and respiratory syndrome, China. Emerg Infect Dis 2007;13:1434–6. 10.3201/EID1309.070399. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Trevisan G, Linhares LCM, Crim B et al. Prediction of seasonal patterns of porcine reproductive and respiratory syndrome virus RNA detection in the U.S. Swine industry. J Vet Diagn Invest 2020;32:394–400. 10.1177/1040638720912406. [DOI] [PMC free article] [PubMed] [Google Scholar]
- VanderWaal K, Pamornchainavakul N, Kikuti M et al. PRRSV-2 variant classification: a dynamic nomenclature for enhanced monitoring and surveillance. mSphere 2025;10:e0070924. 10.1128/MSPHERE.00709-24/FORMAT/EPUB. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vu HLX, Kwon B, Yoon K-J et al. Immune evasion of porcine reproductive and respiratory syndrome virus through glycan shielding involves both glycoprotein 5 as well as glycoprotein 3. J Virol 2011;85:5555–64. 10.1128/JVI.00189-11; CTYPE: STRING: JOURNAL. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wei X, Li R, Qiao S et al. Porcine reproductive and respiratory syndrome virus utilizes viral apoptotic mimicry as an alternative pathway to infect host cells. J Virol 2020;94:10–1128. 10.1128/JVI.00709-20 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wei Z, Lin T, Sun L et al. N-Linked glycosylation of GP5 of porcine reproductive and respiratory syndrome virus is critically important for virus replication In vivo. J Virol 2012;86:9941–51. 10.1128/jvi.07067-11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu Y, Ye M, Sun S et al. CD163-Expressing porcine macrophages support NADC30-like and NADC34-like PRRSV infections. Viruses 2022;14:2056. 10.3390/V14092056/S1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang H, Zhang J, Zhang X et al. CD163 knockout pigs are fully resistant to highly pathogenic porcine reproductive and respiratory syndrome virus. Antiviral Res 2018;151:63–70. 10.1016/J.ANTIVIRAL.2018.01.004. [DOI] [PubMed] [Google Scholar]
- Youssef N, Gurev S, Ghantous F et al. Protein design for evaluating vaccines against future viral variation. BioRxiv, 10.1101/2023.10.08.561389, 7 March 2024, preprint: not peer reviewed. [DOI]
- Zeller MA, Saxena A, Trevisan G et al. PRRSView: An analytical platform for the assessment of PRRSV ORF5 genetic sequences. In: Proceedings of the 13th ACM international conference on bioinformatics, computational biology and health informatics, 2022. 10.1145/3535508.3545105 [DOI]
- Zeller MA, Chang J, Trevisan G et al. Rapid PRRSV-2 ORF5-based lineage classification using Nextclade. Front Vet Sci 2024a;11:1419340. 10.3389/FVETS.2024.1419340/BIBTEX. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z, Wayment-Steele HK, Brixi G et al. Protein language models learn evolutionary statistics of interacting sequence motifs. Proc Natl Acad Sci USA 2024b;121:e2406285121. 10.1073/pnas.2406285121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhu J, He X, Bernard D et al. Identification of new compounds against PRRSV infection by directly targeting CD163. J Virol 2023;97:e0005423. 10.1128/JVI.00054-23/FORMAT/EPUB. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
While the training and test sequences used in this work are of proprietary nature, the PRRSView (Zeller et al. 2022) server shall be referred for a phylogenetic overview of PRRSV-2 ORF5 GP5 sequences, including BLAST, RFLP scores, lineage and sub-lineage classification tools.
