Abstract
Motivation
Sequence repositories have few well-annotated virus mature peptide sequences. Therefore post-translational proteolytic processing of polyproteins into mature peptides (MPs) has been performed in silico, with a new computational method, for over 200 species in 5 pathogenic virus families (Caliciviridae, Coronaviridae, Flaviviridae, Picornaviridae and Togaviridae).
Results
Using pairwise alignment with reference sequences, MPs have been annotated and their sequences made available for search, analysis and download. At publication the method had produced 156 216 sequences, a large portion of the protein sequences now available in https://www.viprbrc.org. It represents a new and comprehensive mature peptide collection.
Availability and implementation
The data are available at the Virus Pathogen Resource https://www.viprbrc.org, and the software at https://github.com/VirusBRC/vipr_mat_peptide.
1 Introduction
Viruses from the families Caliciviridae, Coronaviridae, Flaviviridae, Picornaviridae and Togaviridae are responsible for many human and animal diseases. They are positive-sense, non-segmented, single-stranded RNA viruses. Their genomes encode multifunctional polyproteins (PPs) (Yost and Marcotrigiano, 2013), which are typically annotated as a single feature in GenBank (Benson et al., 2005).
These PPs are proteolytically cleaved in the host into ‘mature peptides’ (MPs) (Lindenbach et al., 2013). Cleavage is performed by the virus-encoded proteases and by host furin proteases, such as at pr-M (Yu et al., 2008). The MPs control virus replication, transmission, pathogenicity and host responses, thus determining patient outcomes.
Yet MP sequences are highly underrepresented in public database genomes, where only 2–20% of PPs are annotated. In 2016, when our work began on this effort, GenBank had annotated 20.4% of its Flaviviridae PPs with a MP. Four other viral families had even lower coverage at 1.9% for Coronaviridae, and up to 14.3% for Togaviridae. At the time of submission, GenBank coverage had improved to between 4% and 50%. Our method currently now produces MP annotations for >99% these genomes.
To address this issue, we developed a method to annotate viral genomes, using a set of highly curated reference sequences as a blueprint. This method is straightforward in that it calculates and transfers curated mature protein annotation positions from a reference to a target sequence, and the predicted output content is gathered by data processing pipelines to be published to a web application. It is flexible enough that new species can be added (Sun et al., 2017). We developed a modular package to input a query, identify the closest intra-species reference within the NCBI taxonomy tree, then align and generate annotations. Finally, we processed all available genome data in five families, and the content was made openly available at the Virus Pathogen Resource (ViPR) www.viprbrc.org.
2 Materials and methods
Viruses were selected as targets for data enrichment. We focused on species that had (i) nucleotide records encoding PPs, (ii) a RefSeq representative for that species and (iii) valid cleavage sites for all proteases, vetted against known hydrolytic specificities, for example cutting to the right of Arginine (^R; P1’=Arg). This check ensured that the published cleavage sites reflected the known protease specificities.
Two hundred and nineteen reference genomes were gathered from RefSeq (O’Leary et al., 2016). Inspection verified that all the records had complete PPs and all mature protein annotations had starts, stops and no gaps. Records with cut discrepancies or premature terminations were not used. After validation, a collection of 219 reference sequences was assembled into a folder, inclusive of all relevant target species and strains.
The method is invoked by calling a perl script, for example as ./vipr_mat_peptide.pl -d ./-i NC_001477_test.gb ≫ out.txt 2≫ err.txt as described further in code comments under #USAGE. Dependencies to BioPerl (Stajich et al., 2002) and other packages are listed. First, a query genome is aligned by ClustalW (Thompson et al., 1994) to a reference sequence, known cleavage sites are used and the locations of MP ‘starts and stops’ are calculated in the new nucleotide strand. Cuts are made. The method is non-codon-aware such as Virulign (Libin et al., 2019), but if the alignment at the cut site presents any gaps within four residues of either side of the cut, the sequence is discarded. Protein is virtually translated for each product. Methods choose the closest in-species virus reference sequence from among the curated RefSeq folder, based on taxonomically-nearest taxon_ID and not the blast score. If no reference match is made at the species level or below, annotations are halted for that sequence and the script proceeds to the next genome in batch. Names for products are transferred from the canonical Reference sequences, which are all available in the vipr-mat-peptide/refseq folder from github and are reviewed with each data release. Perl modules were built to handle sequence analysis and extraction of features, including Annotate_Align.pm, Annotate_def.pm, Annotate_Math.pm, Annotate_Util.pm, to go along with the invoked executable script, vipr_mat_peptide.pl. The software is openly available for inspection and use at https://github.com/VirusBRC/vipr_mat_peptide.
3 Results
MP sequences of 156 216 were predicted synoptically across five virus families. A major outcome is the annotation and storage of a complete set of MP sequences, in the publicly accessible ‘ViPR’ (Pickett et al., 2012). As new genomes become available, MPs are automatically added.
We describe the results using Zika as a model. Flaviviruses encode a single open reading frame. The MPs for Zika virus (ZIKV) include capsid, membrane, envelope and non-structural proteins, as shown in Figure 1.
Fig. 1.

Example ZIKV processing. The polyprotein of Zika virus is post-translationally cleaved by viral and host proteases into: capsid (C), intracellular capsid (Ci), premembrane (pr), membrane (M), envelope (E), and nonstructural proteins (NS1–NS5). Numbering corresponds to the first amino acid of each mature peptide in the reference genome NC_012532
Very few MPs for ZIKV have been directly experimentally identified and sequenced to date. Results are based on the first complete ZIKV genome and our study of other Flavivirus proteases (Enfissi et al., 2016); we previously reported results for ZIKV (Sun et al., 2017). Idiosyncracies in the ZIKV N-terminus have been noted (Theys, 2017) and these were accommodated by allowing the start in the GenBank file to be used so that artificial changes to coding sequence are not introduced.
The method produced a large amount of completely new content. For comparison, the new Flavivirus data here includes 145 487 proteins from 13 894 complete genomes, a > 10x multiple. In contrast, GenBank has only 60k proteins from 47k genomes. Table 1 illustrates the current data summary for all the applicable virus families.
Table 1.
Results: data summary for 156 216 MP annotations, using input genomes with complete PPs only
| Family | Genomes | MP/PP | Proteins | MP | |
|---|---|---|---|---|---|
| 1 | Caliciviridae | 1878 | 6 | 11 768 | 11 268 |
| 2 | Coronaviridae | 2875 | 11 | 47 378 | 17 250 |
| 3 | Flaviviridae | 13 939 | 12–14 | 145 893 | 83 634 |
| 4 | Picornaviridae | 5456 | 12 | 52 799 | 32 736 |
| 5 | Togaviridae | 1888 | 8 | 16 373 | 11 328 |
| Total | 26 036 | 274 211 | 156 216 |
Note: We report that 56% of all ViPR protein data are computationally derived.
Example row (top): 1878 Caliciviridae genomes with 6 MP/PP yield 11 268 new MPs. Total protein numbers are also included.
4 Discussion
Computational biology pipelines were developed to predict viral PP protease cleavage sites, produce MPs and store them. The tools were used to produce comprehensive, consistent and continually updated annotations for publicly available sequences, which enhance the available body of tools used for the research, diagnosis and treatment of viral infectious diseases.
Funding
This work was supported by the US National Institutes of Health, National Institute of Allergy and Infectious Diseases (NIH/DHHS) [HHSN272201400028C].
Conflict of Interest: none declared.
Contributor Information
Christopher N Larsen, Vecna Technologies, Inc., Greenbelt, MD 20770.
Guangyu Sun, Vecna Technologies, Inc., Greenbelt, MD 20770.
Xiaomei Li, Northrop Grumman, Rockville, MD 20850, USA.
Sam Zaremba, Northrop Grumman, Rockville, MD 20850, USA.
Hongtao Zhao, Northrop Grumman, Rockville, MD 20850, USA.
Sherry He, Northrop Grumman, Rockville, MD 20850, USA.
Liwei Zhou, Northrop Grumman, Rockville, MD 20850, USA.
Sanjeev Kumar, Northrop Grumman, Rockville, MD 20850, USA.
Vince Desborough, Northrop Grumman, Rockville, MD 20850, USA.
Edward B Klem, Northrop Grumman, Rockville, MD 20850, USA.
References
- Benson D.A. et al. (2005) GenBank. Nucleic Acids Res., 33, D34–D38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Enfissi A. et al. (2016) Zika virus genome from the Americas. Lancet, 387, 227–228. [DOI] [PubMed] [Google Scholar]
- Libin P.J. et al. (2019) VIRULIGN: fast codon-correct alignment and annotation of viral genomes. Bioinformatics, 35, 1763–1765. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lindenbach B.D. et al. (2013) Flaviviridae: the viruses and their replication. In: Knipe D.M., Howley P.M. (eds.) Flaviviridae. Fields Virology, 6th edn. Lippincott-Raven Publishers, Philadelphia, pp. 712–746. [Google Scholar]
- O’Leary N.A. et al. (2016) Reference sequence (RefSeq) database at NCBI. Nucleic Acids Res., 44, D733–D745. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pickett B.E. et al. (2012) ViPR: an open bioinformatics database and analysis resource for virology research. Nucleic Acids Res., 40, D593–D598. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stajich J.E. et al. (2002) The Bioperl toolkit: Perl modules for the life sciences. Genome Res., 12, 1611–1618. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sun G. et al. (2017) Comprehensive annotation of mature peptides and genotypes for Zika virus. PLoS One, 12, e0170462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Theys K. et al. (2017) Zika genomics urgently need standardized and curated reference sequences. PLoS Pathogens, 13, e1006528. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thompson J.D. et al. (1994) CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res., 22, 4673–4680. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yost S.A., Marcotrigiano J. (2013) Viral precursor polyproteins: keys of regulation from replication to maturation. Curr. Opin. Virol., 3, 137–142. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu I.M. et al. (2008) Structure of the immature dengue virus at low pH primes proteolytic maturation. Science, 319, 1834–1837. [DOI] [PubMed] [Google Scholar]
