ABSTRACT
Metal-binding proteins, including proteases, play critical roles in the replication and pathogenesis of RNA viruses. This study conducted a comprehensive analysis to elucidate the landscape of metalloproteins and identify metal-binding proteases within RNA viruses. The analysis revealed the presence of predicted metal-binding proteins in 19 significant RNA viral families, encompassing 45 viral genera and 375 viral species. The predicted metalloproteome primarily consisted of magnesium (47%) and zinc (40%) binding proteins, with manganese (11%) binding proteins also identified. Through advanced bioinformatics tools, a data set of 905 carefully selected metalloproteins underwent rigorous functional characterization. The prevalence of polyproteins (37%), RNA polymerases (22%), and proteases (10%) was observed. Furthermore, metal ion binding was achieved in viral structural, non-structural, and accessory proteins such as integrase and methyltransferase. The in-depth investigation focused on protease domains, identifying a refined set of 456 non-redundant protease domains. Among them, 78 protease domains were determined to possess metal-binding properties, with zinc and magnesium binding emphasized. The findings provide significant insights into metal-binding proteases’ distribution and evolutionary patterns, particularly in major human RNA viral proteases. Sequence similarity network analysis highlighted the presence of different classes of peptidases in viral families, such as zinc-binding peptidase C30 in the coronaviridae family, including human coronavirus proteases, and peptidase C16 in all genera of coronaviruses. This comprehensive analysis sheds light on the immense potential of metal-binding proteases as therapeutic targets. Continued exploration of metal-binding proteomes will further enhance our understanding of metal-dependent biological processes and facilitate the development of innovative antiviral strategies.
IMPORTANCE
Metal-binding proteins are pivotal components with diverse functions in organisms, including viruses. Despite their significance, many metalloproteins in viruses remain uncharacterized, posing challenges to understanding viral systems. This study addresses this knowledge gap by identifying and analyzing metal-binding proteins and proteases in RNA viruses. The findings emphasize the prevalence of these proteins as essential functional classes within viruses and shed light on the role of metal ions and metalloproteins in viral replication and pathogenesis. Moreover, this research serves as a crucial foundation for further investigations in this field, offering the potential for developing innovative antiviral strategies. Additionally, the study enhances our understanding of the distribution and evolutionary patterns of metal-binding proteases in major human viruses. Continually exploring metal-binding proteomes across diverse viruses will deepen our knowledge of metal-dependent biological processes and provide valuable insights for combating viral infections, including respiratory viruses and other life-threatening diseases.
KEYWORDS: bioinformatics, metalloproteins, viral functional proteins, metal-binding proteases, RNA viruses, zinc ion
INTRODUCTION
Interestingly, a significant portion of an organism’s proteome is metal-binding proteins (MBP) or metalloproteins. These proteins have the ability to bind with at least one metal ion and play critical roles in various functions, including catalytic and structural roles (1). Previous research has demonstrated the involvement of metal ions and their major-binding proteins in the pathogenesis of several viral infections (2). They are a possible aim for new therapeutic medications to cure life-threatening illnesses due to their position in the life cycle of a range of pathogenic organisms. Despite their significant importance, it seems that only a small fraction of metal-binding protein sequences have been uncovered thus far. Metalloproteomes are dynamic in nature and likely to exceed current estimates (3). The analysis of protein sequences using bioinformatics techniques for identifying metalloproteins has steadily increased (4). However, a vast metalloproteome is still waiting to be discovered (5). On the other hand, proteases, a crucial type of enzyme and a promising drug target, play significant roles in catalyzing various biochemical processes in all living organisms, including viruses (6). Some of the most notable human viruses that have proteases which are utilized in the development of drugs with wide therapeutic applications include herpesvirus, hepatitis C virus, alphavirus, picornavirus, adenovirus, flavivirus, and human immunodeficiency virus (7). These viruses encode one or more proteases that mediate the viral assembly and replication by processing viral polyproteins into functional proteins and also help to antagonize the host immune system (8). However, not all viruses encode proteases; only those viruses that use a polyprotein expression strategy encode the protease [i.e., mainly (+ss) RNA and some DNA viruses] (9, 10). Proteases may use different catalytic mechanisms to perform the same reaction. Based on the catalytic mechanism and canonical catalytic residues, they are classified into seven groups, i.e., serine protease, cysteine protease, aspartic protease, glutamic protease, metalloprotease, asparagine protease, and threonine protease. They can detect and cleave various substrate sequences with diverse specificities (11). A class of proteases called “metalloprotease” require a bound divalent metal cation at their active site; it acts as a catalyst for hydrolysis peptide bonds. However, other than metalloproteases, other types of proteases have also been reported to interact with metal ions (6). Metals can also play other important roles in metalloproteins, such as providing structural support or regulation. The zinc ion (Zn2+) is metalloproteases’ most common metal ion. Other transition metals, such as Co2+ and Mn2+, have also been detected in their active sites (12–14). Exploring the metal-binding abilities of proteases presents an exciting opportunity to gain a deeper understanding of the diverse functions that metals can have in proteases while also revealing new metal-binding groups in viruses. This study specifically aims to identify metal-binding proteins in RNA viruses and locate potential metal-binding proteases. Established techniques and resources for detecting the metalloproteome are utilized, with a thorough comparative evaluation being conducted (15–17). This study sheds light on the diverse roles that metal and metal-binding proteins play in viral pathogenesis and disease progression. Additionally, the study delves into the characteristics of metal-binding proteases found in RNA viruses that infect humans and other animals. It also examines important features of human RNA viral metal-binding protease, including conserved domain information, as well as various structural and evolutionary aspects through various bioinformatics tools and approaches. Ultimately, this study estimates the number of existing or new metal-binding groups of RNA viral proteins. Moreover, this study may offer new possibilities for antiviral treatments by exploring the potential of metal-binding viral proteases as therapeutic interventions.
RESULTS
Identification of probable metal-binding proteins of RNA viruses
In this study, an investigation was conducted on a data set comprising 18,852 RefSeq RNA viral proteins. These proteins belong to about 110 RNA virus families and contain the ds (double-stranded), ss (single-stranded), and reverse transciptase (RT) RNA genomes. The aim was to identify potential metal-binding proteins within this data set using various computational methods. BLASTP and hmmer hits were filtered based on their E-value (<0.0001), and the MebiPred tool results were filtered based on a predefined score (>0.5). These filtering criteria helped narrow down the proteins for further analysis. Figure 1 shows the number of viral genera, viral species, total proteins, and metalloproteins in the respective viral family.
Fig 1.
The distribution of predicted metalloproteins within viral families, based on the total available RNA RefSeq viral proteins, viral genus, species, and the number of available proteins in each family. The inner histogram highlights the predicted metal-binding proteins. The data reveal a correlation between the number of available proteins in the RefSeq database and the presence of metal-binding proteins within viral families. Viral families with a higher number of available proteins exhibit a greater abundance of metal-binding proteins.
The results obtained from each method were analyzed and compared to identify overlaps and unique predictions. The UniProt-based method predicted a total of 1,443 metal-binding proteins, while the PDBeChem-based method predicted 1,818. The MetalPDB-based method yielded 1,425 predictions, and the Hmmer/Pfam-based method predicted the highest number of metal-binding proteins, with 7,018 predictions. Interestingly, the MeBiPred tool predicted a significantly higher number of metal-binding proteins, reaching 11,371 predictions (Fig. 2). These results demonstrate that different computational methods can provide varying predictions for metal-binding proteins within the protein data set. Notably, annotation-based methods utilizing UniProt, MetalPDB, and PDBeChem protein databases predict a lower number of proteins compared to Pfam and MeBiPred, which predicts metalloproteins based solely on sequence-derived features. After manually comparing the results from various approaches, we identified common or overlapping proteins in at least four of them. They have a higher probability of being metal binding. However, further analysis and experimental validation would be necessary to confirm the accuracy and relevance of these predictions.
Fig 2.
The figure presents the diverse landscape of metalloproteomes in RNA viruses, as predicted by different bioinformatics methods. (A) The total numbers of predicted RNA viral metalloproteomes and metal-binding proteases are shown. Notably, annotation-based methods utilizing UniProt, MetalPDB, and PDBeChem protein databases predict fewer proteins than Pfam and MeBiPred, which predicts metalloproteins based solely on sequence-derived features. (B) The panel illustrates the metal-wise distribution of predicted metalloproteins, showcasing the varying numbers for each metal as predicted by the different methods. (C) Similarly, the panel highlights the metal-binding proteases predicted by each method, further emphasizing the distinct predictions made by each approach. A total of 905 metalloproteins were carefully selected and subjected to a rigorous selection process, resulting in the shortlisting of 78 proteases with a high probability of being metal-binding proteases. These findings underscore the complexity of predicting metalloproteins and metal-binding proteases while providing valuable insights into their abundance and distribution within RNA viruses.
Distribution of predicted metalloproteome of RNA viruses
For further analysis, a total of 905 metal-binding proteins were shortlisted (Data Set S1). This predicted metalloproteome is mainly composed of magnesium (47%) and zinc (40%) binding proteins, followed by manganese (11%) binding proteins. However, toxic metal-binding proteins for mercury, cadmium, and nickel are also found (Fig. 3A). The analysis of predicted metal-binding proteins reveals their presence in 19 significant RNA viral families encompassing 45 viral genera and 375 viral species. These metal-binding proteins are predominantly found in RNA viruses capable of infecting a wide range of hosts, including mammals, aves, insects, arachnids, bacteria, and plants (tracheophytes) (Fig. 3B). Moreover, a closer examination of the distribution of these metalloproteins across viral families highlights that viral families known to commonly infect humans, such as Coronaviridae, Flaviviridae, Retroviridae, Togaviridae, and Picornaviridae, possess a higher abundance of metal-binding proteins (Fig. 3C). In addition, the identified metal-binding proteins within these viral families exhibit a preference for binding with zinc and magnesium ions, as depicted in Fig. 3C. It is worth mentioning that the predicted metalloproteome might exhibit variations among different viral families owing to the specific criteria used for selection. Nonetheless, these findings shed light on the potential presence of metal-binding proteins within the RefSeq RNA viral protein data set and provide a basis for future investigations into their functional roles and implications.
Fig 3.
The figure presents the distribution of 905 selected metal-binding proteins derived from RNA viruses. (A) The donut chart illustrates the proportion of metal-binding proteins categorized by their respective metals. Zinc, magnesium, and manganese emerged as the predominant constituents of the predicted metalloproteins. These metal-binding proteins originate from 375 viral species, infecting a diverse range of hosts, including mammals, insects, bacteria, and plants. (B) The panel highlights the prevalence of zinc-binding proteins in viruses infecting mammals, birds, insects, arachnids, and actinoptergians. Conversely, viruses infecting plant and bacterial species predominantly exhibit magnesium binding capability. Notably, toxic metal-binding proteins were exclusively found in viruses primarily infecting mammals. The family-wise distribution of metal-binding proteins is depicted in panel (C), revealing a higher abundance of metal-binding proteins in viral families that predominantly infect humans.
Functional annotation
The functional characterization of the shortlisted metalloproteins (905) using high-throughput bioinformatics tools revealed several significant findings (Fig. 4). Most of the predicted metalloproteins were identified as polyproteins, comprising an impressive 37% of the total. Additionally, RNA replication enzyme polymerases accounted for 22% of the metalloproteins, highlighting their crucial role in metal-ion binding processes. Proteases, constituting 10% of the metalloproteins, were also identified, suggesting their involvement in metal-dependent catalytic activities. Notably, viral structural proteins, including Gag polyproteins, capsid, and nucleocapsid proteins, exhibited an affinity for metal ions, indicating their importance in metal coordination within viral mechanisms and host-virus interactions. Furthermore, in addition to the previously mentioned functional categories, it is worth noting that we also categorized certain essential viral proteins as accessory proteins, as they play crucial roles in virus survival and replication. These accessory proteins, comprising a significant portion of the predicted metalloproteins (16%), include integral components like integrase and methyltransferase, which have been predicted to possess metal-binding capabilities. However, it is important to note that the proportions mentioned above provide a general overview of the presence of metalloproteins in various processes within viruses. Obtaining an accurate proportion for each functional class can be challenging due to the presence of polyproteins and multifunctional proteins. For instance, proteins like NS5 of the Zika virus have been found to exhibit both helicase and protease activities, blurring the boundaries between distinct functional classes. These multifunctional proteins further emphasize the complexity and versatility of metalloprotein functions within viral systems. Therefore, while the proportions mentioned earlier provide valuable insights, it is crucial to consider the presence of polyproteins and multifunctional proteins when assessing the exact distribution and functional diversity of metalloproteins in viral processes.
Fig 4.
Functional annotations of selected metal-binding proteins highlight the diverse functional classes of viral proteins with metal-binding properties. The figure showcases the prevalence of specific functional categories among the predicted metalloproteins. Polyproteins constitute a substantial portion (37%), followed by RNA replication enzyme polymerases (22%) and proteases (10%). Additionally, accessory proteins comprise a significant proportion (16%) of the predicted metal-binding proteins. Major structural proteins such as capsid, nucleocapsid, and Gag polyproteins also demonstrate metal-binding capabilities. These findings shed light on the importance of metal-binding properties in various functional classes of viral proteins and further emphasize their significance in viral replication, pathogenesis, and potentially as therapeutic targets.
Extracting metal-binding non-redundant protease domains
In our investigation, we identified a total of 461 metalloproteins that exhibit the presence of protease domains. Since many of these proteins were found within polyproteins, we specifically extracted the protease domains from these polyproteins, resulting in a final set of 796 individual proteases. To ensure the accuracy of our data set, we checked for redundancy among the protease domains and eliminated any redundant protein sequences. This process resulted in 456 non-redundant protease domains. It is important to note that not all extracted protease domains necessarily exhibit metal-binding properties, which can lead to false positives in the data. To address this, we re-evaluated these proteins by employing the same approaches mentioned earlier for identifying metal-binding proteins. This rigorous re-searching process allows us to refine our data set and provide a more reliable assessment of metal-binding protease domains. Interestingly, despite the previous filtering criteria being successfully passed by all the predicted proteins, we observed distinct numbers of metal-binding proteins when employing these prediction methods again. As illustrated in Fig. 2, the Uniprot-based method predicted 319 metal-binding proteins, MetalPDB predicted 171, the PDBeChem-based method predicted 235, and the Hmmer/Pfam-based method predicted 434 proteins. These contrasting results further highlight the importance of considering multiple prediction methods to understand the diverse landscape of metalloproteins comprehensively. In our further analysis, we focused on a subset of 78 protease domains that were predicted as metal binding by at least four of the applied approaches (Fig. 2). We specifically chose proteins that exhibited overlapping predictions for a single metal across all four methods. This selection criterion was crucial because different prediction methods may assign different metal-binding properties to the same protein. For example, a protein identified as zinc binding by one method might be classified as magnesium binding by another method.
Distribution of metal-binding proteases
In this study, we initially explored the binding efficiency of different metals to viral proteins. However, after employing various filtering methods, the focus narrowed down to zinc- and magnesium-binding proteins as potential metal-binding proteases. Interestingly, it was revealed that all the predicted proteases are classified as zinc-binding proteins, with some also having the ability to associate with magnesium ions (as depicted in Fig. 5). Four virus families were identified among the predicted viral metal-binding proteases: Coronaviridae, Picornviridae, Retroviridae, and Flaviviridae. Notably, these families include significant human viruses such as coronaviruses, hepatitis C virus, HIV, enteroviruses, and rhinoviruses. The proteases associated with these viruses also exhibit the efficiency to bind both zinc and magnesium ions, suggesting the importance of these metals in the functional mechanisms of these viral proteases.
Fig 5.
The figure illustrates the distribution of predicted metal-binding proteases across four virus families, with a specific focus on human viral proteases represented by the outer red ring. Zinc-binding proteases are predominantly found in coronaviruses and retroviruses, while picornaviruses and flaviviruses primarily exhibit magnesium-binding proteases. These predicted proteases are associated with significant human viral pathogens and demonstrate the potential to bind with zinc and magnesium metal ions.
Sequence similarity network analysis of predicted metal-binding proteases
Predicted metal-binding proteases belong to four different virus families (Fig. 5) and five MEROPS (18) peptidase families [peptidase C30, peptidase C16, peptidase S29, peptidase C03, and peptidase A02 (Data Set S2)]. MEROPS is a manually curated database and information resource of peptidases (proteases) that includes all information about protease classification, their inhibitors, and substrates (18). The similarity between every two protein sequences was calculated, and the Sequence Similarity Network (SSN) was constructed [using a lower alignment score (19)]. Initially, seven clusters were formed (Fig. 6A). All the nodes (metal-binding proteases) are zinc-binding proteins (circle shape), except for peptidase C30 and C16. The other three peptidase family proteases also have binding sites for magnesium (hexagonal node shape). Peptidase C30 and C16 family proteases are distributed in the Coronaviridae viral family, generating two distinct clusters for each peptidase type. In comparison, peptidase C3 and A2 family proteases are found in the Picornaviridae and Retroviridae, forming two clusters for each family. These family groups are further separated by their viral genus. Members of the peptidase S29 family belong to the Flaviviridae viral family. These family groups are further isolated by their viral genus. When the alignment score rises to 30 (sequence similarity above 30%; Fig. 6B) and 50 (sequence similarity above 50%; Fig. 6C), the proteases of peptidase C30, C16, and A2 generate more clusters. The viral genus accounts for the majority of the grouping. Following these arrangements, substantially similar sequences are grouped. Red highlighted circles show human viral proteins, and edges with sequence similarity of more than 70% are also highlighted in brown. In the coronaviridae family (Fig. 6C), the putative zinc-binding peptidase C30 seen in alpha and betacoronavirus proteases and human coronavirus proteases is more related to bat and rat viral proteases. In peptidase C16, a similar trend was identified for the human betacoronavirus protease.
Fig 6.
The weighted undirected sequence similarity network of a total of 78 metal-binding proteases of different RNA Viruses (denoted by nodes). Edges denote the similarity between the connected nodes. The viral family is denoted by the node’s color, while shapes denote the zinc (circle) and magnesium (hexagon) binding proteins; on the other hand, respective hosts are also demonstrated by the node’s outline color. Predicted metal-binding proteases seem to be differentiated by the catalytic MEROPS peptidase family of peptidases (peptidase C30, etc.). SSN shows edges when sequence similarity exceeds to above 30% (B) and above 50% (C), while the brown color of edges shows highly similar sequences (more than 75% similar)). In panel (B) of the figure, edges are shown when the sequence similarity between proteases exceeds 30%, while in panel (C), edges are displayed for sequence similarity above 50%. The brown color of the edges signifies highly similar sequences with a similarity level exceeding 75%. Different peptidase families distribute in varying viral families. Similar protease sequences clustered together and sub-clustering occur on the basis of different genera in the same family. It was also observed that in the case of coronaviridae, a different class of peptidase can also be found in the same family, but they are making two separate clusters. Different similarity filters show that human viral proteases have similarities with rat and bat viral proteases. This sequence similarity network analysis provides valuable insights into the relationships and similarities among the predicted metal-binding proteases, revealing clusters of proteases within specific peptidase families.
The peptidase C16 domain, on the other hand, is found in alpha, beta, and gamma coronaviruses. There is just one cluster of peptidases S29 proteases in the Flaviviridae family, and it has been shown that zinc- and magnesium-binding proteins are present. Red highlighted circles represent human viral proteins, and the highlighted edge shows sequence similarity above 70%. The Flaviviridae family has just one cluster of S29 proteases, and zinc- and magnesium-binding human flavivirus proteases have been found to be closely linked to other primate viral proteins. The A2 and C3 proteases form two clusters for each associated family, and each cluster represents a different viral genus. The human retrovirus and picornavirus protease appears to be associated with the other non-human primate viral protease.
Phylogenetic analysis
For all putative metal-binding proteases, a phylogenetic tree was built. Based on the diverse functional protease families, six groups were formed. In addition, sub-clustering can also be recognized based on their distinct viral genus. A remarkable finding was observed (Fig. 7) that a protease of the rhinovirus-B genus (a member of the Picornaviridae family) belongs to the peptidase C3 family, but when we look at the tree, it is associated with the peptidase C30 group. All C3 proteases are magnesium-binding proteins, but this protease has zinc-binding sites identical to C30 proteases. Although the predicted metal-binding proteases (from C3, A2, and S29 families) share similar clade origins, apparent diversification is observed by analyzing viral hosts belonging to these proteases. It was also observed that lentivirus (mainly HIV) proteases are in the same subset but have distinct zinc- and magnesium-binding proteins. Viruses containing potential metal-binding proteases from three peptidase families (C3, A2, and S29) have predominantly been identified to infect primates. In contrast, other viruses containing metal-binding peptidase C30 and C16 have been reported to infect a broader range of hosts. Metal-binding proteases, on the other hand, show a limited host range. In the case of peptidases C16 and C30, human viral proteases are closely associated with the Chiroptera and Rodentia hosts.
Fig 7.
Phylogeny of the 78 predicted metal-binding proteases. Proteases of family A02, S29, and C03 are clustered together, while peptidases C16 and C30 form diverse sub-cluster based on viral genera. The clustering of proteases is primarily based on the catalytical family of the proteases; the viral genera can see further sub-clustering. Magnesium-binding ability can be found in some proteases; most of these proteases are associated with human and non-human primates.
Metal-binding assessment
In order to assess the metal-binding capabilities of chosen viral proteases, which had limited three-dimensional structure data available, structural modeling was conducted on a total of 78 proteases. The modeled structures were then analyzed to determine their predicted capacity for binding to Mg2+/Zn2+ metal ions. It was found that the predictor was able to confidently transplant these respective predicted metal ions into the modeled structure. After analyzing 78 proteases, the predictor found that a significant number of them bound to their predicted metal ion with a high degree of similarity. Data Set S2 showed that approximately 33 proteases had a similarity of 70% or greater, while 41 proteases bound to their respective metal ion with 50%–70% similarity. Additionally, two proteases were predicted to have 40%–50% similarity. Only two proteases did not predict to bind with their respective metals. These proteins are proteases (A2) of primate T-lymphotropic virus 1 and human immunodeficiency virus 2. These viral proteins were further validated by the similar available structure in the Protein Data Bank (PDB) database and filled by the respective metal utilizing AlphaFill. These results indicate a high degree of confidence in the predicted metal-binding capabilities of the proteases analyzed.
DISCUSSION
The varying role of metal and metal-binding proteins in numerous dynamic disease processes makes metalloproteins a very intriguing research subject. With the exponential utilization of metalloproteins as therapeutic candidates, various bioinformatics algorithms for predicting metal-binding proteins using multiple computational tools and techniques have also increased dramatically in recent years (16, 20). Various metalloproteins have been identified by applying one of the methods based on the sequence homology with previously identified metalloproteins (3, 16, 17, 21, 22). By applying this method, we used previously reported approaches to identify RNA viral metal-binding proteins (19). Surprisingly, each method yields different significant results, which might be attributed to the increased diversity of metal-binding sites and the exponential accumulation of experimental data. It should be noted that in silico prediction of metal-binding proteins is difficult since different bioinformatics approaches have varied benefits and limitations in capturing essential characteristics of metalloproteins (16). However, the present resource complexity associated with the necessary experimental labor renders identifying an organism’s whole collection of metal-binding proteins impossible. As a result, bioinformatics approaches may provide valuable support (23), although variation in result outputs of each applied approach indicates the higher requirement of a bioinformatics method or tool to predict metalloproteins of an organism. To increase our prediction accuracy, we shortlisted those proteases which are predicted as metal binder by at least four applied approaches.
In addition, we perform various analyses on these predicted metal-binding protease domains, offering vital insights into their distribution, function, and evolution. However, each approach predicts different types of metalloproteomes, but only the zinc- and magnesium-binding proteases are shortlisted, using high accuracy. Zinc is essential not only for the proteins and enzymes of humans and other living organisms but also for viruses. Several viral proteins are reported as zinc-dependent metalloenzymes and zinc fingers (24, 25). On the other hand, magnesium is also the most common divalent cation present in living cells (26) and can bind with several viral proteins (27). Zinc and magnesium deficiencies have been reported in patients with the recently emerged COVID-19 virus, indicating a possible role of these metal ions in disease progression (24, 26). In our study, we identified zinc- and magnesium-binding viral proteases belonging to different RNA virus families and classes of proteases. Although these viral proteases perform multiple catalytic activities during the virus’s life cycle, their dependence on these metals for catalysis still needs to be determined. On the other hand, in these metal-binding proteases, these metals may also have different potential roles, such as regulatory or structural. In-depth insights are still needed to discover the potential roles of zinc and magnesium in metal-binding proteases and to explore other metal-binding features. Some human viral proteases related to important human viruses, such as coronaviruses, hepatitis C virus, HIV, enteroviruses, and rhinoviruses, are also identified as potential metal-binding proteases. Because of their essentiality in viral survival and pathogenesis, they are the best possible drug targets and are being widely investigated for their therapeutic potential in most human viral diseases. Several inhibitors have already become successful examples of inhibiting HIV and HCV (hepatitis C virus) viral proteases (11) to halt disease progression. However, the metal-binding characteristics of these viral proteases still need to be fully explored.
In addition to aiding in the catalytic activity of proteases, metals can also help in protein-protein and protein-ligand interactions. In the search for new antiviral drugs, these metal-based inhibitors could prove to be a breakthrough invention. Using metal ionophores and chelators to target metalloproteins has shown promise in fighting viral diseases. FDA-approved drugs like disulfiram, chloroquine, dolutegravir, and baloxavir marboxil have been effective in inhibiting viral diseases by either blocking the viral protein activity by altering metal ions concentration or forming stable complexes with metals to hinder viral replication (1, 28–31). Baloxavir marboxil (Xofluza) is another successful example of a metal-chelating drug that has recently been expanded in its approval by the FDA (www.fda.gov) for the post-exposure prevention of influenza in both pediatric and adult patients (31). This study highlights the metal-binding potential within numerous viral proteases, underscoring their potential suitability as promising candidates for drug targets and designing metal-based inhibiting strategies. However, various metal-based inhibitors against other viral protein targets, such as RNA-dependent RNA polymerase, are also being analyzed (32, 33). Metals are often associated with the active sites of proteins and enzymes and can regulate the functional activity of proteins by coordinating with ligands in a three-dimensional configuration. This property prompted the development of metal-based inhibitors with promising new therapeutic possibilities and applications (34). Some metals exhibit different oxidation states and can interact with a few specific amino acids of the proteins. Deep insights into the three-dimensional structure and metal-binding pockets are major steps to understanding metal-protein interactions.
In addition, we also focused on the evolutionary aspects of these predicted metal-binding proteases. Few previous studies have been done on the evolutionary possibilities of viral proteases, but they still lack focus on metal-binding features. Some previous studies reported the divergent nature of viral proteases even among closely related viruses, but on the other hand, they have been emphasized for conservation in terms of maintaining catalytic sequence specificity (35). The same trend is also visible in our study showing the SSN, with catalytically distinct proteases clustering separately. Although genus-wise clustering is also visible, the primary clustering is based on the catalytic class of proteases, which may be due to similar catalytic sequence sites. It is also observed that the metal-binding domains of human peptidase C30 and peptidase C16 are identical (above 70%) to the proteases of rat and bat viruses. It also followed the evolutionary theory previously described in the coronavirus (36), which suggested that the gene pools of the rat and bat are responsible for the evolution of the coronavirus. Our other observation by phylogeny tree suggests that metal-binding proteases belonging to the peptidase C16 groups are more diverse than those of C30, whereas proteases of flaviviruses, picornaviruses, and retroviruses may have a common ancestor. Strikingly, proteases of human rhinoviruses (peptidase C03) are showing clustering with proteases of coronaviruses (peptidase C30). Along with rhinoviruses, coronaviruses, and influenza viruses are the major cold-causing viruses in humans (37). These findings will motivate the study of the evolutionary relationship between rhinovirus and coronavirus. In addition, phylogeny clustering also indicates a close association of human coronavirus proteases with rat and bat viral proteases. Another observation noted in the metal-binding propensity is that most of the proteases that bind to magnesium are associated with human and non-human primate viruses. Currently, this type of study focusing on various features, including the metal-binding ability of viral proteases, is rare.
Our comprehensive analysis also involving structural modeling and assessment of selected viral proteases has provided robust confidence in predicting their capacity for metal binding. The utilization of the AI-based “Alphafold2” (38) and “AlphaFill” (39) approaches has effectively enabled the incorporation of predicted Mg2+/Zn2+ metal ions into the modeled structures, enhancing the precision of predictions. Notably, our findings were also cross-referenced with available three-dimensional viral protease structures retrieved from the PDB database (40). This thorough comparison unveiled compelling instances where the presence of metal ions aligned precisely with our predictions. Noteworthy examples include the NS3 protease of Hepacivirus (PDB id: 2xcf), which exhibits both zinc and magnesium ions, human enterovirus B (PDB id: 3q3x) displaying magnesium ion attributes, and specific coronaviruses such as SARS Co-V-2 (PDB id: 6w9c), and avian infectious bronchitis virus (PDB id: 4x2z), both featuring zinc-binding characteristics within their respective structures. Furthermore, by utilizing AlphaFill (39), we also transplanted metal ions into those available three-dimensional structures that lacked them, revealing consistent metal-binding potential across these viral proteases (Data Set S2). Therefore, this work may be beneficial in providing hope for new and efficient antiviral therapeutic possibilities.
MATERIALS AND METHODS
Data retrieval
The complete reference protein sequences (RefSeq) of all the available RNA viruses were extracted from the Virus Variation database of the National Centre for Biotechnology Information (NCBI) (https://www.ncbi.nlm.nih.gov/labs/virus/vssi/#/); date of data retrieval: 5 May 2022). NCBI’s Virus Variation database is the value-added resource of curated viral sequences assembled using GeneBank and other NCBI repositories (41). All the retrieved protein sequences were subjected to sequence-wide identification of the metal-binding patterns using different high-throughput computational approaches (15, 17, 23, 42). All the methodology steps used in this study are shown in Fig. S1.
Mining of potential metal-binding proteins
Identifying metal-binding proteins is challenging, whether done experimentally or through in silico methods (43). Here, we employed the sequence-homology method utilized by some protein databases like UniProtKB (44), Pfam (45), MetalPDB (46), and PDBeChem (47), in addition to a standalone tool called MeBiPred, which is a sequence alignment-free method (43), to predict metal-binding proteins. The UniProt Knowledgebase (UniProtKB) is a widely used protein database (44) that contains all protein sequences derived from genome sequencing projects and annotated with reliable, functional annotations. The Pfam database is a resource that includes all protein families and domains with various annotations and multiple sequence alignments generated by using hidden Markov models (45). The MetalPDB is a metal-specific database that contains all the information on diverse metal-binding sites found in three-dimensional biological structures gathered from protein databases (PDB) (46). The PDBeChem database is a chemical component dictionary service provided by the European Molecular Biology Laboratory’s European Bioinformatics Institute (https://www.ebi.ac.uk/pdbe-srv/pdbechem), which was also utilized to access the information of proteins that are bound to specific metals as ligand. The standalone tool MebiPred is the machine learning-based method that provides over 80% accuracy in recognizing metal-binding proteins using sequence-derived features (43). Based on the available literature, the 10 metals [zinc (Zn2+), iron (Fe3+/2+), magnesium (Mg2+), manganese (Mn2+), copper (Cu3+/2+), cadmium (Cd2+), mercury (Hg2+), cobalt (Co2+), molybdenum (Mo2+), and nickel (Ni2+)] may play an essential role in viral pathogenesis and have been shortlisted for the current study (1, 2). All the above-mentioned databases (UniProt, MetalPDB, Pfam, and PDBeChem) were used to access the data of already known metal-binding protein sequences and converted into local databases by using standalone BLASTP (48) and the HMMER search tool (http://hmmer.org). HMMER is a search tool that can be used with profile databases (such as Pfam) for searching sequence homologs. The keywords for the advanced databases search are “zinc-binding,” “magnesium-binding,” and so on. The idea is to find the sequence similarity between the already known metal-binding proteins and query proteins. It is well known that highly similar protein sequences show similar structural and functional features (21). All the downloaded reference sequences were used as a query sequence against the prepared local databases. Hits were shortlisted based on E-value (i.e., less than 1 × 10−4). All the downloaded reference sequences were also used as input to the Standalone tool MebiPred to predict metal-binding proteins using a threshold value of >0.5. Outputs of all the approaches were compared manually, and proteins predicted as possible metal-binding proteins by at least four methods were considered high-confidence putative “metal-binding proteins.”
Functional classification
We used the bioinformatics tool InterProScan and the NCBI’s Conserved Domains Database (NCBI-CDD) to detect functional and conserved domains in the predicted MBPs. InterProScan is a tool that recognizes the different protein signatures by scanning the various secondary protein databases (such as Pfam, Prosite, and PANTHER) and assigns specific functions to a query protein (49). The NCBI-CDD is a database containing the conserved domain footprints of proteins exported from various other resources like Entrez Protein, PubMed, and NCBI Biosystems (50). After conducting a functional characterization, we shortlisted those metal-binding proteins that have the potential to act as proteases. Additionally, we have also found many polyproteins present in the data, from which only protease domains were extracted and utilized for further study.
Redundancy settlements
CD-Hit is a widely used program that clusters the biological sequences on user-defined similarity threshold values to remove redundancy from the data set (51). We perform similarity-based clustering using the CD-Hit tool with a cut-off of >99% to remove the redundancy from the shortlisted data. The sequences with >99% similarity were removed manually to ensure the non-redundancy of the sequences.
Searching for sequence similarity to extract metalloprotease domains
After the segmentation of polyproteins, many false positives may be present in the data; to ensure the accuracy of the identification process, shortlisted proteases are again subjected to a search for metal-binding protease domains [repeated the step 2 (methodology sub-section 2)]. The metal-binding protease domain selection was based on defined selection criteria (i.e., E-value less than 1 × 10−4 for sequence similarity and >0.5 scores for MeBiPred). All the results were evaluated, and proteases that were again common in predicted outputs of at least four methods were shortlisted. Shortlisted proteins are considered to perform other significant analyses and have a high probability of containing metalloprotease domains.
Proteins sequence similarity network analysis
A weighted undirected graph (SSN) was generated for these predicted metal-binding proteases using the web server EFI-EST (52). An alignment score was calculated for each edge obtained from all-against-all BLASTP to create SSN. The magnitude of this alignment score is nearly identical to the negative logarithm of the E-value. Edges represent the alignment scores between the two linked protein sequences; highly similar proteins (nodes) have short edge lengths and are located near each other. We defined a sequence identity threshold and paired the nodes (MBP) when the sequence identity surpassed the threshold value. This threshold was calculated by comparing the networks built with an incremental range of threshold values. This constructed network is visualized and edited using Cytoscape (53).
Inference and visualization of phylogenetic relationships
A phylogenetic tree helps to understand evolutionary patterns by showing the common ancestry of different organisms or species that descended from a common ancestor, or it can provide information about evolutionary events, biological diversity, and hierarchical classification (54). To build the clustergram of putative metal-binding protease domains of RNA viruses, the MAFFT tool (a multiple alignment program) was utilized to perform the multiple-sequence alignment with default parameters (55). A phylogenetic tree was built to determine the evolutionary connection of these aligned sequences using MEGAX with the maximum likelihood approach at default settings and bootstrap value 500 (56). To comprehend the evolutionary links among each protein better, we annotated each protein with various related information in the tree and displayed it using the iTol: Interactive Tree Of Life tool (https://itol.embl.de) (57).
Metal-binding assessment
For validation purposes, we modeled selected viral proteases to assess their propensity as metal-binding proteins. This includes the use of ColabFold (58), which is an implementation of AlphaFold2 (38) optimized for notebook-based use. Through this method, we generated accurate, high-quality structural models for metal-binding proteases. Next, we employed the AlphaFill (39) approach. It uses both sequence and structural similarity to intuitively incorporate small molecules and ions from experimentally established structures into predicted protein models.
ACKNOWLEDGMENTS
The authors acknowledge the Indian Council of Medical Research, Government of India, for providing the Senior Research Fellowship to Himisha Dixit (ICMR: BMI/11(77)/2020), Institution of Eminence grant by the University of Delhi to Shailender Kumar Verma (Ref. No./IoE/2021/12/FRP), and the Central University of Himachal Pradesh for providing the laboratory space and computational facilities.
Contributor Information
Shailender Kumar Verma, Email: sverma@es.du.ac.in.
Colin R. Parrish, Cornell University Baker Institute for Animal Health, Ithaca, New York, USA
SUPPLEMENTAL MATERIAL
The following material is available online at https://doi.org/10.1128/jvi.01399-23.
Information of all RNA metalloproteins, functional annotations, segmentation of polyproteins, total protease, total metal binding proteins, and metal binding proteases.
Details of all the predicted metal-binding proteases, metal-binding assessment analysis, and other supporting information regarding the evolutionary analysis.
Overall workflow to predict metal-binding proteins and metal-binding proteases within RNA viruses.
ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.
REFERENCES
- 1. Chen AY, Adamek RN, Dick BL, Credille CV, Morrison CN, Cohen SM. 2019. Targeting metalloenzymes for therapeutic intervention. Chem Rev 119:7719. doi: 10.1021/acs.chemrev.9b00322 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Chaturvedi UC, Shrivastava R. 2005. Interaction of viral proteins with metal ions: role in maintaining the structure and functions of viruses. FEMS Immunol Med Microbiol 43:105–114. doi: 10.1016/j.femsim.2004.11.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Maret W. 2010. Metalloproteomics, metalloproteomes, and the annotation of metalloproteins. Metallomics 2:117–125. doi: 10.1039/b915804a [DOI] [PubMed] [Google Scholar]
- 4. Andreini C, Bertini I, Rosato A. 2004. A hint to search for Metalloproteins in gene banks. Bioinformatics 20:1373–1380. doi: 10.1093/bioinformatics/bth095 [DOI] [PubMed] [Google Scholar]
- 5. Cvetkovic A, Menon AL, Thorgersen MP, Scott JW, Poole FL, Jenney FE, Lancaster WA, Praissman JL, Shanmukh S, Vaccaro BJ, Trauger SA, Kalisiak E, Apon JV, Siuzdak G, Yannone SM, Tainer JA, Adams MWW. 2010. Microbial Metalloproteomes are largely Uncharacterized. Nature 466:779–782. doi: 10.1038/nature09265 [DOI] [PubMed] [Google Scholar]
- 6. Hoppert M. 2011. Metalloenzymes. In Finkl CW, Fairbridge RW (ed), Encyclopedia of earth sciences series [Google Scholar]
- 7. Sharma A, Gupta SP. 2017. Fundamentals of viruses and their proteases. doi: 10.1016/B978-0-12-809712-0.00001-0 [DOI]
- 8. Tsu BV, Beierschmitt C, Ryan AP, Agarwal R, Mitchell PS, Daugherty MD. 2021. Diverse viral proteases activate the nlrp1 inflammasome. Elife 10:e60609. doi: 10.7554/eLife.60609 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Konvalinka J, Kräusslich HG, Müller B. 2015. Retroviral proteases and their roles in virion maturation. Virology:403–417. doi: 10.1016/j.virol.2015.03.021 [DOI] [PubMed] [Google Scholar]
- 10. Majerová T, Novotný P. 2021. Precursors of viral proteases as distinct drug targets. Viruses 13:1981. doi: 10.3390/v13101981 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Kurt Yilmaz N, Swanstrom R, Schiffer CA. 2016. Improving viral protease inhibitors to counter drug resistance. Trends Microbiol 24:547–557. doi: 10.1016/j.tim.2016.03.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Ward OP. 2011. Proteases, p 604–615. In Comprehensive biotechnology. Elsevier. [Google Scholar]
- 13. Nagase H. 2001. Metalloproteases. CP Protein Science 24. doi: 10.1002/0471140864.ps2104s24 [DOI] [PubMed] [Google Scholar]
- 14. Rao MB, Tanksale AM, Ghatge MS, Deshpande VV. 1998. Molecular and biotechnological aspects of microbial proteases. Microbiol Mol Biol Rev 62:597–635. doi: 10.1128/MMBR.62.3.597-635.1998 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Sharma A, Sharma D, Verma SK. 2019. In silico identification of copper-binding proteins of Xanthomonas translucens pv. undulosa for their probable role in plant-pathogen interactions. Physiol Mol Plant Pathol 106:187–195. doi: 10.1016/j.pmpp.2019.02.005 [DOI] [Google Scholar]
- 16. Zhang Y, Zheng J. 2020. Bioinformatics of metalloproteins and metalloproteomes. Molecules 25:3366. doi: 10.3390/molecules25153366 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Dixit H, Upadhyay V, Kulharia M, Verma SK. 2023. The putative metal-binding proteome of the coronaviridae family. Metallomics 15:mfad001. doi: 10.1093/mtomcs/mfad001 [DOI] [PubMed] [Google Scholar]
- 18. Rawlings ND, Barrett AJ, Thomas PD, Huang X, Bateman A, Finn RD. 2018. The MEROPS database of proteolytic enzymes, their substrates and inhibitors in 2017 and a comparison with peptidases in the PANTHER database. Nucleic Acids Res 46:D624–D632. doi: 10.1093/nar/gkx1134 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Dixit H, Kulharia M, Verma SK. 2023. Metalloproteome of human-infective RNA viruses: a study towards understanding the role of metal ions in virology. Pathog Dis 81:ftad020. doi: 10.1093/femspd/ftad020 [DOI] [PubMed] [Google Scholar]
- 20. Andreini C, Arnesano F, Rosato A. 2022. The zinc proteome of SARS-CoV-2. Metallomics 14:mfac047. doi: 10.1093/mtomcs/mfac047 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Andreini C, Putignano V, Rosato A, Banci L. 2018. The human iron-proteome. Metallomics 10:1223–1231. doi: 10.1039/c8mt00146d [DOI] [PubMed] [Google Scholar]
- 22. Kushwah AS, Dixit H, Upadhyay V, Yadav S, Verma SK, Prasad R. 2023. Elucidating the zinc-binding proteome of Fusarium oxysporum f. sp. lycopersici with particular emphasis on zinc-binding effector proteins. Arch Microbiol 205:298. doi: 10.1007/s00203-023-03638-1 [DOI] [PubMed] [Google Scholar]
- 23. Andreini C, Bertini I, Rosato A. 2009. Metalloproteomes: a bioinformatic approach. Acc Chem Res 42:1471–1479. doi: 10.1021/ar900015x [DOI] [PubMed] [Google Scholar]
- 24. Doboszewska U, Wlaź P, Nowak G, Młyniec K. 2020. Targeting zinc metalloenzymes in coronavirus disease 2019. Br J Pharmacol 177:4887–4898. doi: 10.1111/bph.15199 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Lei J, Kusov Y, Hilgenfeld R. 2018. Nsp3 of coronaviruses: structures and functions of a large multi-domain protein. Antiviral Res 149:58–74. doi: 10.1016/j.antiviral.2017.11.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Dominguez LJ, Veronese N, Guerrero-Romero F, Barbagallo M. 2021. Magnesium in infectious diseases in older people. Nutrients 13. doi: 10.3390/nu13010180 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Viswanathan T, Misra A, Chan SH, Qi S, Dai N, Arya S, Martinez-Sobrido L, Gupta YK. 2021. A metal ion orients SARS-CoV-2 mRNA to ensure accurate 2′-O methylation of its first nucleotide. Nat Commun 12:3287. doi: 10.1038/s41467-021-23594-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Enoki Y, Kishi N, Sakamoto K, Uchiyama E, Hayashi Y, Suzuki N, Ito M, Taguchi K, Yokoyama Y, Kizu J, Matsumoto K. 2021. Multivalent cation and polycation polymer preparations influence pharmacokinetics of dolutegravir via chelation-type drug interactions. Drug Metab Pharmacokinet 37:100371. doi: 10.1016/j.dmpk.2020.11.006 [DOI] [PubMed] [Google Scholar]
- 29. Xue J, Moyer A, Peng B, Wu J, Hannafon BN, Ding W-Q. 2014. Chloroquine is a zinc ionophore. PLoS One 9:e109180. doi: 10.1371/journal.pone.0109180 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Lin MH, Moses DC, Hsieh CH, Cheng SC, Chen YH, Sun CY, Chou CY. 2018. Disulfiram can inhibit MERS and SARS coronavirus papain-like proteases via different modes. Antiviral Res 150:155–163. doi: 10.1016/j.antiviral.2017.12.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Noshi T, Kitano M, Taniguchi K, Yamamoto A, Omoto S, Baba K, Hashimoto T, Ishida K, Kushima Y, Hattori K, et al. 2018. In vitro characterization of baloxavir acid, a first-in-class cap-dependent endonuclease inhibitor of the influenza virus polymerase PA subunit. Antiviral Res 160:109–117. doi: 10.1016/j.antiviral.2018.10.008 [DOI] [PubMed] [Google Scholar]
- 32. Ioannou K, Vlasiou MC. 2022. Metal-based complexes against SARS-CoV-2. BioMetals 35:639–652. doi: 10.1007/s10534-022-00386-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Pathania S, Rawal RK, Singh PK. 2022. RdRp (RNA-dependent RNA polymerase): a key target providing anti-virals for the management of various viral diseases. J Mol Struct 1250:131756. doi: 10.1016/j.molstruc.2021.131756 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Sodhi RK. 2019. Metal complexes in medicine: an overview and update from drug design perspective. Cancer Ther Oncol Int J 14. doi: 10.19080/CTOIJ.2019.14.555883 [DOI] [Google Scholar]
- 35. Tsu BV, Fay EJ, Nguyen KT, Corley MR, Hosuru B, Dominguez VA, Daugherty MD. 2021. Running with scissors: evolutionary conflicts between viral proteases and the host immune system. Front Immunol 12:769543. doi: 10.3389/fimmu.2021.769543 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Cui J, Li F, Shi ZL. 2019. Origin and evolution of pathogenic coronaviruses. Nat Rev Microbiol 17:181–192. doi: 10.1038/s41579-018-0118-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Lewis-Rogers N, Seger J, Adler FR. 2017. Human rhinovirus diversity and evolution: how strange the change from major to minor. J Virol 91:e01659-16. doi: 10.1128/JVI.01659-16 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, et al. 2021. Highly accurate protein structure prediction with AlphaFold. Nature 596:583–589. doi: 10.1038/s41586-021-03819-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Hekkelman ML, de Vries I, Joosten RP, Perrakis A. 2023. Alphafill: enriching AlphaFold models with ligands and cofactors. Nat Methods 20:205–213. doi: 10.1038/s41592-022-01685-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, Shindyalov IN, Bourne PE. 2000. The protein data bank. Nucleic Acids Res 28:235–242. doi: 10.1093/nar/28.1.235 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Hatcher EL, Zhdanov SA, Bao Y, Blinkova O, Nawrocki EP, Ostapchuck Y, Schäffer AA, Brister JR. 2017. Virus variation resource-improved response to emergent viral outbreaks. Nucleic Acids Res 45:D482–D490. doi: 10.1093/nar/gkw1065 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Piovesan D, Profiti G, Martelli PL, Casadio R. 2012. The human "magnesome": detecting magnesium binding sites on human proteins. BMC Bioinformatics 13:S10. doi: 10.1186/1471-2105-13-S14-S10 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Aptekmann AA, Buongiorno J, Giovannelli D, Glamoclija M, Ferreiro DU, Bromberg V. 2022. mebipred: identifying metal-binding potential in protein sequence. Bioinformatics 38:3532–3540. doi: 10.1093/bioinformatics/btac358 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Bateman A, Martin MJ, Orchard S, Magrane M, Agivetova R, Ahmad S, Alpi E, Bowler-Barnett EH, Britto R, Bursteinas B, et al. 2021. Uniprot: The universal protein knowledgebase in 2021. Nucleic Acids Res 49:D480–D489. doi: 10.1093/nar/gkaa1100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, Tosatto SCE, Paladin L, Raj S, Richardson LJ, et al. 2021. Pfam: the protein families database in 2021. Nucleic Acids Res 49:D412–D419. doi: 10.1093/nar/gkaa913 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Putignano V, Rosato A, Banci L, Andreini C. 2018. MetalPDB in 2018: a database of metal sites in biological macromolecular structures. Nucleic Acids Res 46:D459–D464. doi: 10.1093/nar/gkx989 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Varadi M, Anyango S, Armstrong D, Berrisford J, Choudhary P, Deshpande M, Nadzirin N, Nair SS, Pravda L, Tanweer A, et al. 2022. PDBe-KB: collaboratively defining the biological context of structural data. Nucleic Acids Res 50:D534–D542. doi: 10.1093/nar/gkab988 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ. 1997. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res 25:3389–3402. doi: 10.1093/nar/25.17.3389 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Blum M, Chang HY, Chuguransky S, Grego T, Kandasaamy S, Mitchell A, Nuka G, Paysan-Lafosse T, Qureshi M, Raj S, et al. 2021. The InterPro protein families and domains database: 20 years on. Nucleic Acids Res 49:D344–D354. doi: 10.1093/nar/gkaa977 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Marchler-Bauer A, Bo Y, Han L, He J, Lanczycki CJ, Lu S, Chitsaz F, Derbyshire MK, Geer RC, Gonzales NR, et al. 2017. CDD/SPARCLE: functional classification of proteins via subfamily domain architectures. Nucleic Acids Res 45:D200–D203. doi: 10.1093/nar/gkw1129 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Fu L, Niu B, Zhu Z, Wu S, Li W. 2012. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics 28:3150–3152. doi: 10.1093/bioinformatics/bts565 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Gerlt JA, Bouvier JT, Davidson DB, Imker HJ, Sadkhin B, Slater DR, Whalen KL. 2015. Enzyme function initiative-enzyme similarity tool (EFI-EST): a web tool for generating protein sequence similarity networks. Biochim Biophys Acta 1854:1019–1037. doi: 10.1016/j.bbapap.2015.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, Amin N, Schwikowski B, Ideker T. 2003. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res 13:2498–2504. doi: 10.1101/gr.1239303 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Baum DA, Smith SD, Donovan SSS. 2005. The tree-thinking challenge. Science 310:979–980. doi: 10.1126/science.1117727 [DOI] [PubMed] [Google Scholar]
- 55. Katoh K, Rozewicki J, Yamada KD. 2019. MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Brief Bioinform 20:1160–1166. doi: 10.1093/bib/bbx108 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Kumar S, Stecher G, Li M, Knyaz C, Tamura K. 2018. MEGA X: molecular evolutionary genetics analysis across computing platforms. Mol Biol Evol 35:1547–1549. doi: 10.1093/molbev/msy096 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Letunic I, Bork P. 2021. Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotatio. Nucleic Acids Res 49:W293–W296. doi: 10.1093/nar/gkab301 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Mirdita M, Schütze K, Moriwaki Y, Heo L, Ovchinnikov S, Steinegger M. 2022. ColabFold: making protein folding accessible to all. Nat Methods 19:679–682. doi: 10.1038/s41592-022-01488-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Information of all RNA metalloproteins, functional annotations, segmentation of polyproteins, total protease, total metal binding proteins, and metal binding proteases.
Details of all the predicted metal-binding proteases, metal-binding assessment analysis, and other supporting information regarding the evolutionary analysis.
Overall workflow to predict metal-binding proteins and metal-binding proteases within RNA viruses.







