Summary
Emerging evidence suggests that amino acid homorepeats (HRs) in proteins (HRPs) contribute to protein interactability. What is the role of HRs in HPIs? We found that pathogens engage physiologically important human HRPs, thereby affecting diverse host physiological processes. From the pathogen standpoint, (1) eukaryotic pathogens engage more HRPs but with host-sparse HRs (HR types that are rare in the host), leading to disparate and discriminate interactions, (2) prokaryotic pathogens engage less HRPs but with host-abundant non-polar HRs via host-protein proxies, bringing about discriminate or promiscuous interactions, and (3) viral pathogens engage more HRPs with host-abundant polar uncharged HRs, affecting promiscuous interactions using host-partner HR tract mimicry. To propel further research, we introduce a resource Hi-PHI (http://hiphi.iisertirupati.ac.in/), cataloging critical information about human and pathogen HRPs and HRs. We propose mechanisms to (1) repurpose drugs targeting human HRPs engaged by pathogens for treating different infections, and (2) exploit HRs and their flanks as targets for pathogen-targeted anti-infectives.
Subject areas: computational bioinformatics, microbiology, molecular biology
Graphical abstract

Highlights
-
•
Pathogens engage physiologically important human proteins with homorepeats (HRP)
-
•
Eukaryotic, prokaryotic and viral pathogens have distinct modes of HR and HRP engagement
-
•
Homorepeats are involved in human-pathogen protein interactions (HPIs)
-
•
Hi-PHI database details HRPs in HPIs
Computational bioinformatics; Microbiology; Molecular biology
Introduction
Invasion of host by pathogens involves engaging host machinery effectively for the propagation and/or physiological functions of the pathogens, thereby triggering pathogenic response. This includes abrogating host physiological functions and/or hijacking host system. Pathogen proteins aid host invasion by (1) molecular mimicry to sequester host proteins and/or (2) de novo pathological interactions to abrogate host physiological functions (e.g., immune response), and/or hijack host machinery for pathophysiological outcomes (e.g., viral replication).1,2 Pathogen proteins are known to interact with host proteins using protein functional units such as (1) protein domains, which are self-folding units that can function independently3,4,5,6 (Figures S1A and S1B), and (2) sequence motifs, which are small stretches of consensus residues that are associated with molecular functions7,8 (Figures S1C and S1D).
Emerging evidence suggests that stretches of identical amino acids in proteins, referred to as homorepeats (HRs), can act as interaction modules and contribute to high interactability, leading to multi-functionality of proteins with HRs (HRPs).9,10,11 Importantly, human HRPs are associated with core processes, such as gene expression regulation, development, and signaling in humans,9,10 making them attractive targets for pathogens. On the other hand, the increased interaction potential conferred by HRs might be pivotal for pathogen proteins to better engage human proteins. How are human and pathogen HRPs and HRs engaged in human-pathogen interactions? In this study, we elucidate the role of HRPs and HRs in human-pathogen interactions by assembling, integrating, and analyzing diverse large-scale publicly available datasets pertaining to human-pathogen interactions, human functional and regulatory interactions, spatial gene expression, phenotypic screens, drug-target networks, and priority pathogens identified by the World Health Organization (WHO) (Figure 1A).
Figure 1.
Modalities of HPIs
(A) Illustration of the framework adopted in this study to understand the role of HRPs and HRs in human-pathogen pathophysiology.
(B) Attributes of the human-pathogen protein interaction (HPI) network assembled in this study.
(C) Taxonomic classification of human unicellular and viral pathogens represented in the HPI network. The numbers in parentheses indicate the number of species in each taxonomic class.
(D) Classification of human and pathogen proteins based on their participation in HPIs and the presence of HRs; n denotes the number of proteins in each class.
Results
We assembled a comprehensive human-pathogen protein interaction (HPI) network, which comprised 19,535 interactions between 5,295 human proteins and 3,407 pathogen proteins, pertaining to a total of 286 human cellular (prokaryotic and eukaryotic) and viral pathogens spanning different taxa (Figures 1B and 1C). We classified proteins with identical amino acid runs (length ≥5) as HRPs, as previously defined.9,12 Based on their participation in HPIs and the presence of HRs, we classified the human proteins into HPI proteins with HRs (hHPI-HRPs) and without HRs (hHPI-NonHRPs), and those that do not participate in HPIs with HRs (hNonHPI-HRPs) and without HRs (hNonHPI-NonHRPs) (Figure 1D). We drew comparisons for different attributes of hHPI-HRPs with hHPI-NonHRPs and hNonHPI-HRPs spanning phenotypes, molecular networks, and gene expression. Based on taxonomy and presence of HRs, pathogen proteins participating in the HPIs (pHPI) were categorized into those that belonged to (1) eukaryotic pathogens (epHPI-NonHRPs and epHPI-HRPs), (2) prokaryotic pathogens (ppHPI-NonHRPs and ppHPI-HRPs), and (3) viral pathogens (vpHPI-NonHRPs and vpHPI-HRPs). To understand the taxa-specific attributes of pathogen HRPs, comparisons were drawn between (1) the HRPs and NonHRPs belonging to each of the three taxa, and (2) the pathogen HRs across the three taxa (Figure 1D).
Physiologically important human HRPs are engaged by pathogens
Are human HRPs engaged in HPIs? Human HRPs, especially those engaged by prokaryotic and viral pathogens, are enriched among human proteins that participate in HPI (Figures 2A and 2B). Importantly, we found that hHPI-HRPs engaged by prokaryotic and viral pathogens are predominantly enriched for regulatory processes such as chromatin organisation, transcription and post-transcriptional regulation (Figure 2C). Additionally, hHPI-HRP partners of prokaryotes and viruses are also enriched for development and differentiation, while those of the prokaryotes are enriched for signaling. We could not find any enriched biological processes for the hHPI-HRPs engaged by eukaryotic pathogens. This can be due to (1) fewer hHRPs identified thus far to be interacting with eukaryotic pathogens, and/or (2) the hHRPs interacting with eukaryotic pathogen proteins participating in diverse biological processes. Contrarily, hHPI-NonHRPs are predominantly enriched for post-translational regulation, cell death, and signaling across the three taxa of pathogens (Figure 2C). Jaccard similarity indices across different biological processes show that hHPI-NonHRP interacting partners of prokaryotic and viral pathogen proteins are more shared across biological processes than the hHPI-HRPs (Figure 2D). This suggests that prokaryotic and viral pathogens target hHPI-HRPs that participate in or regulate specific processes.
Figure 2.
Physiological importance of human proteins engaged by pathogens
(A) Enrichment of HRPs in human proteins that are engaged by pathogens. The histogram in gray represents the random expectation, and blue arrow denotes the actual observation of the number of hHPI-HRPs.
(B) Enrichment of human HRPs that are engaged by different taxa of pathogens. The dotted red line denotes the significance estimate, and the dotted grey line highlighted by the red asterisk represents the threshold for statistical significance. Z scores and p values were estimated using permutation testing.
(C) Bubble plot indicating the proportion of significantly enriched (FDR < 0.05) gene ontology biological process (GO-BP) terms in each of the manually classified major biological processes. The size of the bubble indicates the proportion of enriched GO-BP terms in each broad category for a given protein class in each taxon.
(D) Proportion of shared proteins across different biological processes, estimated using Jaccard similarity index (JSI); n denotes the number of pairs in each class. p value was estimated using Wilcoxon rank-sum test.
(E–J) Bar plots and boxplots.
(E, H, I, and J) Bar plots showing the proportion of human (E) essential proteins, (H) transcription factors (TFs), (I) RNA-binding proteins (RBPs), and (J) proteins that participate in more than one human phase separated condensates engaged by different pathogens in HPIs.
(F and G) Boxplots representing the (F) total number of tissues in which the different classes of human proteins are expressed and (G) number of protein-protein interactions in humans.
p values were computed using Fisher’s exact test (E, H, I, and J) or Wilcoxon rank-sum test (F and G); n denotes the total number of proteins in each class.
What is the physiological relevance of hHPI-HRPs? Physiological importance can be assessed through various attributes such as gene essentiality, tissue-wide expression, centrality in human protein-protein interaction network, and/or impact on global regulation. We found that hHPI-HRPs engaged by prokaryotic and viral pathogens (1) are preponderantly essential, (2) have broader tissue distribution in humans, (3) have higher number of protein-protein interactions, and (4) are enriched for human regulatory HRPs involved in transcriptional and post-transcriptional regulation (Figures 2E–2I). Strikingly, hHPI-HRPs engaged by viral pathogens predominantly participate in multiple biological condensates, which represent both higher-order assemblies resulting from multiple interactions and dynamic regulatory units (Figure 2J). The attributes that represent physiological importance are lot more pronounced in hHPI-HRPs than in hNonHPI-HRPs (Figure S2). Collectively, these findings imply that functionally important human HRPs are engaged by pathogens in HPIs, affecting a diverse range of host biological processes, including global regulation across tissues.
Sequestration of hHRPs by pathogens might significantly impact human physiology
How important are hHPI-HRPs for human physiology? To mimic the impact of sequestration of human proteins by pathogen proteins on human protein interaction network, we performed network dismantling. For this, we removed each protein (node) in the human protein-protein interaction network, corresponding to a virtual knockout of that protein, and then assessed the impact of node removal on the network topology by estimating the (1) link density—average global connectivity of the network,13 and (2) assortativity—preferential association of proteins with similar interaction potential14 (Figures 3A–3C). Removal of hHPI-HRPs that interact with prokaryotic and viral pathogens causes (1) significant reduction in the link density, resulting in sparsely connected networks and (2) significant increase in the assortativity, leading to an increase in disjoint modules, compared with removal of hHPI-NonHRPs (Figure 3D). While hHPI-HRPs participate in a limited number of processes (Figure 2C), they exhibit high interactability and inter-connectivity in human protein interaction networks (Figure 3C), implying that they might connect pathways or protein complexes involved in closely related biological processes. These observations suggest that the hHPI-HRPs engaged by prokaryotic and viral pathogen proteins bring about more interactability at an individual protein level and denser connectivity at the network level, facilitating inter-connectedness across modules affecting fundamental processes.
Figure 3.
Impact of human HRPs engaged by pathogens on human protein-protein interaction network and on different human cellular phenotypes
(A–C) Schema illustrating the network dismantling approach (A), and estimation of the network topological attributes, link density (B)13 and assortativity (C),14 upon node (protein) removal from the human protein-protein interaction network. The numbers in the nodes represent the number of connections (degree) for each node.
(D) Boxplots representing the link density (top) and assortativity (bottom) of human proteins that are engaged in HPIs. p value was computed using Wilcoxon rank-sum test.
(E and F) Bar plots showing the proportion of proteins that affect (E) cell differentiation and cell proliferation during neural induction of human iPSCs, and (F) survival of neurons in oxidative stress, assessed using CRISPRa and CRISPRi. n denotes the total number of proteins in each class. p value was estimated using Fisher’s exact test.
To decipher the plausible phenotypic impact when hHPI-HRPs are engaged by pathogens, we investigated publicly available genome-wide CRISPRi/a-based phenotypic screens assessing (1) cell fate determination, (2) cell proliferation, and (3) cell survival during stress. Compared with HPI-NonHRPs and NonHPI-HRPs, removal of hHPI-HRPs tends to significantly influence the neural differentiation of human induced pluripotent stem cells (iPSCs) but not cell proliferation (Figures 3E and S2H). This suggests that the hHPI-HRPs are key for neuronal differentiation. While the impact of removal of hHPI-HRPs is comparable to that of hHPI-NonHRPs for neuronal survival under oxidative stress (Figure S2I), there is a significant change in the survival of neurons when hHPI-HRPs are activated using CRISPRa (Figure 3F). This finding suggests that upregulation of hHPI-HRPs influences the stress response behavior of neurons in a dosage-dependent manner. Thus, engaging such proteins by pathogens would have greater impact on host physiology.
Pathogens of different taxa show diverse modes of engaging pHRPs
How do pathogens employ their HRPs to engage human proteins? Analysis of the proportion of pathogen HRPs (pHRPs) of different taxa (ep, eukaryotic pathogen; pp, prokaryotic pathogen; vp, viral pathogen) revealed that higher proportion of proteins of eukaryotic pathogens that participate in HPIs are HRPs, compared to those of prokaryotes and viruses (Figure 4A). Although more in number, epHPI-HRPs predominantly engage in solitary interactions with human proteins, bringing about one-to-one or discriminate interactions (Figures 4B and 4C). Contrarily, ppHPI-HRPs bring about discriminate as well as one-to-few/many i.e., promiscuous interactions with human proteins (Figures 4B and 4C), vpHPI-HRPs predominantly engage in promiscuous interactions with human proteins (Figures 4B and 4C). These findings capture considerable diversity in pHRP deployment in terms of the number of pHRPs and their mode of engagement of human proteins across different taxa of pathogens.
Figure 4.
Modes of HRP and HR engagement by different taxa of human pathogens
(A and B) Bar plots showing the proportion of pHRPs (A) belonging to pathogens of different taxa (ep, eukaryotic pathogens; pp, prokaryotic pathogens; vp, viral pathogens) that participate in HPIs, with n denoting the number of proteins in each class, and (B) that engage single or multiple human proteins in HPIs. p value was estimated using Fisher’s exact test.
(C) Illustration of different modes by which pHPI-HRPs engage human proteins in HPIs.
(D) Top: Heatmap showing fold change of different amino acid HR types of different taxa of pathogens that are engaged in HPI over the total HRs engaged in HPI. Bottom: Bar plot showing the distribution of HR types in the human proteome.
(E) Bar plot showing the proportion of interactions of pHRPs and hHRPs with similar amino acid HR types; n denotes the number of unique interactions in each taxon. p value was estimated using Fisher’s exact test.
(F) Boxplot depicting the total number of human interactors of viral HRPs with at least one human HRP interactor having similar amino acid HR type (similar) and viral HRPs with human HRP interactors with different amino acid HR types (different); n denotes the number of unique viral HRPs in each class. p value was estimated using Wilcoxon rank-sum test.
(G) Illustration of the different modes of HRP and amino acid HR type engagement by human pathogens of different taxa.
Amino acid HR type might influence diverse modes of interactions of pHRPs across different taxa
What is the impact of the nature of amino acid HRs on the variability in the mode of interactions of pHRPs belonging to different taxa of pathogens? To address this, we examined the proportion of different amino acid pHRPs across the three different taxa in HPIs. We found that epHPI-HRPs show a higher tendency for the presence of polar uncharged HRs such as polySer, polyThr, polyAsn, and polyGln (Figure 4D). Strikingly, the human proteome is depleted for these HR types except polySer (Figure 4D). Furthermore, about 75% of epHPI-HRPs that interact with human HRPs lack HRs with similar physicochemical properties (Figure 4E). For example, the epHPI-HRP transcription factor PDR1 (polyAsn-containing HRP) of Saccharomyces cerevisiae S288C interacts with human mediator of RNA polymerase II transcription subunit 15 MED15 (polyPro- and polyGln-containing hHPI-HRP). Binding of human MED15 to S. cerevisiae PDR1 promotes transcriptional activation of ATP-dependent drug efflux pumps, inducing multi-drug resistance.15 This implies that eukaryotic pathogens employ HRs that are less frequent in the human proteome and engage the human proteins using “host-sparse” HRs, probably leading to de novo interactions. This type of HR engagement could lead to disparate discriminate interactions (Figure 4G).
Conversely, ppHPI-HRPs show an overrepresentation of non-polar HRs such as polyGly, polyAla, and polyLeu (Figure 4D). Strikingly, such amino acid HR types are predominant in the human proteome. In the human protein interaction network, more than 50% of interactions of non-polar HRs containing hHPI-HRPs are with other non-polar human HRPs (Figure S3A). Furthermore, the grand average of hydropathy (GRAVY) score of human proteins that interact with non-polar HRs containing hHPI-HRPs is high, indicating higher hydrophobicity (Figure S3B). Akin to these observations, in the HPIs, about 60% of hHPI-HRPs with ppHPI-HRPs have amino acid HR types of similar physicochemical nature (Figure 4E). This could indicate that pHPI-HRPs with non-polar HRs can potentially mimic human HRPs with non-polar HRs and hijack their neighborhood by hydrophobic interactions. For instance, (1) the cell division protein FtsZ (polyGly-containing ppHPI-HRP) in Francisella tularensis subsp. tularensis SCHU S4 interacts with human proteins involved in cell adhesion and cytoskeleton proteins, including protocadherin-16 DCHS1 (polyLeu-containing hHPI-HRP), phospholipid scramblase 1 (PLSCR1), and moesin (MSN), as well as transcription factors involved in immune response such as TNFAIP3-interacting protein 1 (TNIP1),16 and (2) inner membrane permease of ABC transporter fhuB of Yersinia pestis (polyLeu-containing ppHPI-HRP) interacts with human transcription factors; immune response and cell proliferation proteins, including NF-κB p105 subunit (NFKB1; polyGly-containing hHPI-HRP); cyclic AMP-dependent transcription factor ATF-6α (ATF6); and thioredoxin reductase 1 (TXNRD1)16 (Figure S3C). NF-κB signaling is critical for the effective host response to infection. Binding of Y. pestis fhuB to human NF-κB is associated with the impairment of T cell and antigen-presenting cell activation and differentiation, leading to rapid apoptosis of infected macrophages, thereby contributing to evasion of host immune response.16,17 This implies that HRPs of prokaryotic pathogens engage human HRPs by (1) mimicking the host partner tract or (2) hijacking the host protein neighborhood through host-protein proxy, specifically non-polar HRs containing HRPs, via hydrophobic interactions, facilitating promiscuous interactions (Figures 4G and S3).
Alternatively, vpHPI-HRPs engage HRs such as polyGly, polyAla, polyPro, and polyArg to engage host proteins (Figure 4D); these HRs, except polyArg, are also predominant in the human proteome. The interactions between vpHPI-HRPs and human HRPs containing both similar (mainly non-polar HRs) as well as different HR types are comparable (Figure 4E). Interestingly, compared to those that interact with different HR types containing human HRPs, vpHPI-HRPs that interact with similar HR tract containing human HRPs have a higher number of host interactions, probably contributing to promiscuous interactions (Figure 4F). For instance, polyPro-containing UL27 vpHPI-HRP of human herpesvirus 5 (strain Merlin) interacts with multiple non-polar HR-containing human proteins—26S proteasome non-ATPase regulatory subunit 3 (PSMD3; polyAla and polyPro HRs), E3 ubiquitin-protein ligase UBR5 (UBR5; polyAla, polyLeu, and polySer HRs), WD repeat-containing protein 26 (WDR26; polyGly and polySer HRs), protein disulfide isomerase A4 (PDIA4; polyGlu and polyLeu HRs), E1A-binding protein p400 (EP400; polyGlu, polyPro, and polyGln HRs), and transformation/transcription domain-associated protein (TRRAP; polyPro HR).18 The interaction of human herpesvirus 5 UL27 with components of the human proteasome (PSMD3, UBR5) facilitates viral RNA transcription, subsequently leading to efficient viral replication.18,19 This implies that vpHPI-HRPs show both disparate interactions that lead to de novo interactions as well as interactions using host-partner tract mimicry, with the latter bringing about promiscuous interactions (Figure 4G). Collectively, these findings highlight that pathogens belonging to different taxa show distinct patterns in employing different amino acid HR types to engage human proteins.
Human and pathogen HRs constitute HPI interfaces
The findings thus far highlight the different modes of HRP and HR engagement in HPIs. Do HRs bring about HPIs? To examine the role of HRs in facilitating HPIs, we obtained the protein structures of hHRPs, pHRPs, and their interacting partners from AlphaFold Protein Structure Database20 (human, eukaryotic, and prokaryotic pathogens) and Viro3D21 (viral pathogens). This extensive search indicated 1,074 human proteins and 1,161 pathogen proteins involved in 3,265 HRP-mediated interactions. By selecting structures that had a reliable confidence score (average pLDDT [predicted local distance difference test] score ≥50) and appropriate sequence length (50–1,000 residues), we performed blind molecular docking of 1,501 HPIs between 652 human proteins and 784 pathogen proteins by using ZDOCK 3.0.2/2.3.222 (Figure S4A and Table S3). Subsequently, we identified the interfaces of each of the docked HPIs and extracted information of those interfaces where HRs are involved in the interaction with a reliable confidence score (average pLDDT score ≥50 of residues involved in HR-mediated interactions). We found that hHRs are involved in the interface of protein interactions with 8 eukaryotic proteins, 33 prokaryotic proteins, and 72 viral proteins. On the other hand, 9 HPIs involved epHRs and 2 involved ppHRs. We could not find any reliable vpHRs at the interaction interface (Figure S4B and Table S4). The following instances illustrate the potential role(s) of HR-mediated interactions in molecular pathogenesis.
PolyPro-containing human vasodilator-stimulated phosphoprotein (VASP) is engaged in the interaction with major actin (ACT1) of the eukaryotic pathogen Dictyostelium discoidum (Figure 5A). ACT1-VASP-profilin complex formation has been shown to be essential for actin polymerization in both humans and D. discoidum. The interaction of human VASP with D. discoidum ACT1 could accelerate pathogen actin polymerization, leading to the growth and division of the pathogen.23,24 Similarly, polyAsp of Plasmodium falciparum 3D7 nucleosome assembly protein (NAP) is engaged in interaction with human histone H3.1 (H3C1) (Figure 5B). This interaction could aid hijacking of the host transcriptional regulation.25 Importantly, polyAsp is a host-sparse HR, which is engaged by the pathogen to interact with the human partner, suggesting a discriminate disparate interaction.
Figure 5.
Human and pathogen HRs at the interface of hHRP and pHRP interactions
(A–E) Representative structures of interaction interfaces involving hHRs (A, C, and E) and pHRs (B and D). Docked structures of (A) hHPI-HRP (VASP; AlphaFold ID: AF-P50552-F1-v6) with epHPI-NonHRP (major actin; AlphaFold ID: AF-P07830-F1-v6), (B) hHPI-NonHRP (histone 3.1; AlphaFold ID: AF-P68431-F1-v6) with epHPI-HRP (nucleosome assembly protein; AlphaFold ID: AF-Q8I2W3-F1-v6), (C) hHPI-HRP (FCER1G; AlphaFold ID: AF-P30273-F1-v6) with ppHPI-NonHRP (putative ankyrin repeat protein CBU_0781; AlphaFold ID: AF-Q83DF6-F1-v6), (D) hHPI-HRP (NFKB1; AlphaFold ID: AF-P19838-F1-v6) with ppHPI-HRP (FxsA; AlphaFold ID: AF-Q8D1E6-F1-v6), and (E) hHPI-HRP (TBP; AlphaFold ID: AF-P20226-F1-v6) with vpHPI-HRP (E2; GenBank ID: CAA52472.1 for Viro3D structures21). These docked structures show how HRs in either humans or pathogens bring about HPIs; insets show the interaction interface between HR and other residues. The dotted lines in red show atomic distances between the interacting residues.
In the case of human and prokaryotic protein interactions, the non-polar HR polyLeu of human high-affinity immunoglobulin epsilon receptor subunit gamma (FCER1G) is engaged in interaction with putative ankyrin repeat protein CBU_0781 of Coxiella burnetii RSA 493 (Figure 5C). The direct binding of the pathogen protein to a key component of the human immune system (Fc receptor) may have cascading effects that interfere with phagocytosis and cytokine generation.26 The non-polar HR polyLeu of Yersinia pestis FsxA protein is involved in interaction with the human NFKB1, which has a polyGly HR (Figure 5D), representing host-partner tract mimicry and/or host-protein proxy. As highlighted above, attenuation of the NF-κB pathway could impede activation and differentiation of T cells and antigen-presenting cells, affecting accelerated apoptosis of infected macrophages.16,17
The polyGln HR of human TATA-box-binding protein (TBP) is engaged in the interaction with the regulatory protein E2 of human papillomavirus 3 (HPV-3), which has a polyGln HR, highlighting host-partner tract mimicry (Figure 5E). The HPV-3 E2 interaction with the key human general transcription factor TBP could aid in the transcriptional activity of the double-stranded DNA virus. This interaction not only facilitates the transcription of the genetic material of the virus but also abrogates the global transcription of the host cell, making it a double whammy that contributes to a rapid and persistent infection.27 Strikingly, the extent of engagement of hHRs and pHRs in the interaction interfaces varied from a single residue in the HRs to the entire HR region, with about one-third of the hHRs showing ≥75% of hHR length being engaged in the interaction interface with pathogen proteins (Figure S4C). These findings clearly illustrate that HRs participate in interaction interfaces of HPIs, facilitating human-pathogen protein interactions.
Discussion
System-level and molecular-level studies so far have elucidated the role of protein domains and sequence motifs in host-pathogen interactions.6,8,28 However, little is known about how amino acid repeats, an emerging class of interaction modules, are engaged in host-pathogen interactions. Here, we present a comprehensive system-level analysis and show how different taxa of pathogens engage both human and pathogen proteins with HRs distinctly in HPIs. We found that (1) HRPs are engaged by both human and pathogen proteins during host-pathogen interactions, (2) pathogen proteins engage physiologically relevant human HRPs in HPIs whose depletion might have severe consequences to host physiology, and (3) pathogen proteins across different taxa engage pHRPs of diverse HR types differently in HPIs, bringing about distinct modes of engagement of host proteins.
Notably, some of the priority pathogens identified by the WHO (based on parameters such as mortality, incidence, non-fatal health burden, transmissibility, preventability and treatability), namely Mycobacterium tuberculosis, Escherichia coli, Candida parapsilosis and Candida tropicalis, engage a high number of pHRPs in HPIs and/or have high HRP content in their proteomes (Figure S4). Similarly, some of the biodefense pathogens classified by the National Institute of Allergy and Infectious Diseases (NIAID)—such as M. tuberculosis, E. coli, Bacillus anthracis and Y. pestis—owing to their ability to cause diseases of high public health concern—engage a relatively high number of pHRPs in HPIs (Figure S4). Strikingly, pathogens causing life-threatening diseases such as malaria (Plasmodium falciparum, which caused nearly 250 million cases and 600,000 deaths worldwide in 2022) and tuberculosis (M. tuberculosis that caused ∼1.3 million deaths in 2022), which also result from multi-drug resistance, engage more pHRPs in HPIs. These findings imply that high-risk human pathogens show a higher number of HRPs in their proteomes and/or employ more numbers of pHRPs in HPIs. From a viral pathogen perspective—(1) viruses such as human herpesvirus 5 and Epstein-Barr virus that infect multiple organs and organ systems, and/or (2) epidemic- and pandemic-causing life-threatening viruses such as Ebola and SARS coronavirus—have high HRPs in their proteomes and/or engage more HRPs in HPIs (Figure S4). Furthermore, in addition to the presence of HRs, repeat-associated variation can also contribute to pathogenicity. This is best exemplified by the interaction of polyHis HR of Knob-associated histidine-rich protein (KAHRP) in P. falciparum with β-spectrin of human-infected erythrocytes in a repeat length-dependent manner, affecting cytoadhesion and, hence, the pathogenesis.29,30,31 Interestingly, about 26% of human HRPs that are engaged in HPIs (304 out of 1,174) and about 18% mappable hHRs at the interaction interface of HPIs (13 out of 72) show repeat-associated variation.32 Importantly, tandem repeat variability in the genomes of microbes has been shown to contribute to rapid adaptability against host immune responses, either through rapid modulation of gene-function or via virulence factors,33,34,35,36 and that in humans has been associated with various diseases.37,38,39,40,41,42,43 In this context, delineating the effects of repeat variation arising from instability of tandem repeats in coding genes of both humans and pathogens is paramount to decipher their roles in incidence of/susceptibility to infections, extent and type(s) of symptom manifestation, duration of infection, possibility of co-infections, and response/resistance to drug(s). These facets emphasize the need for immediate and thorough investigations to delineate the impact of different amino acid HRs and repeat-associated variation and tandem repeats in the genome, as a whole, on pathogenicity and virulence.
From the pharmacological standpoint, reconstruction of a human-drug interaction network of hHPI-HRPs revealed that many HRPs are already being targeted to alleviate certain infectious diseases (Figure 6A). For instance, human NF-κB1 with polyGly HR is targeted in HIV, herpes simplex virus, and SARS CoV-2 infections.44,45 In the HPI network, human NF-κB1 interacts with 8 different bacterial and viral pathogens, amounting to about 172 HPIs. Human epidermal growth factor receptor (EGFR) is targeted in several skin and soft tissue infections.46 In the HPI network, human EGFR, which contains polyLeu HR, interacts with 8 different bacterial and viral pathogens, with 19 pathogen protein interactions. Human plasminogen (PLMN), which contains polyLeu HR, is targeted by drugs for conjunctivitis47 and interacts with 6 eukaryotic and bacterial pathogens, participating in a total of 21 HPIs. Tranexamic acid, which competitively binds to human PLMN48,49,50 and aids in its targeted removal, has the potential to be repurposed as a drug for infections that involve engagement of human PLMN. Notable examples for drugs targeting HRPs that have already been repurposed for infectious diseases are (1) aspirin and ibuprofen, from inflammation to anti-bacterial and anti-fungal drugs,51,52 (2) thalidomide, from anti-emetic to leprosy,52,53 and (3) the broad-spectrum anti-viral ribavirin to COVID-19.52,54,55 These instances highlight that some of the existing drugs that target the important hHRPs can be repurposed for treating and/or alleviating symptoms of infections caused by pathogens, even from different taxa, offering a cost-effective and low-risk strategy.56
Figure 6.
Therapeutic intervention strategies for targeting human and/or pathogen HRPs or HRs that participate in HPI
(A) Tripartite network with connections between FDA-approved drugs/molecules under clinical trial administered during infections (left nodes), drug target hHRPs participating in HPI (middle nodes) and pathogens that engage the hHRPs identified in this study (right nodes). The color of the drug nodes denotes the drug type based on function, while that of the pathogen nodes on the right denotes the taxonomic class of the pathogen. The number in parentheses for hHRP nodes indicates the total number of pathogen protein partners of each hHRP. The conditions for which the drugs are being currently administered are provided on the left side of the drugs.
(B) Schema representing the approach that involves HR- and their flank-aided PROTAC-based protein targeting and degradation.
(C) Top: Pairwise alignment of the N-terminal and C-terminal flanking residues of pathogen HRs with corresponding flanks of the same amino acid HR type in human proteome. Bottom: Histograms showing the sequence identity of HR-associated flanks of pathogen and human proteins with same amino acid HR type. The pie chart shows the total pairs of pathogen and human flanking sequences that qualify (lighter color in the pie) for the analysis (at least 10 residues from pathogen and human flanking sequences should be aligned). The gray part of the pie denotes pathogen HRs and the flanks, with <10 residues aligning with the regions in human proteome, are ideal for designing species-specific PROTAC molecule without any cross-reactivity with human proteins. n denotes the total number of pathogens and human HR-flanking pairs that were analyzed for the study.
Abbreviations: PK3CG, phosphatidylinositol 4,5-bisphosphate 3-kinase catalytic subunit gamma isoform; SIR1, NAD-dependent protein deacetylase sirtuin-1; PK3CB, phosphatidylinositol 4,5-bisphosphate 3-kinase catalytic subunit beta isoform; P4K2A, phosphatidylinositol 4-kinase type 2-alpha; GPX1, glutathione peroxidase 1; NU6M, NADH-ubiquinone oxidoreductase chain 6; CP2C8, cytochrome P450 2C8; NFKB1, nuclear factor NF-κB p105 subunit; NFKB2, nuclear factor NF-κB p100 subunit; TRBM, thrombomodulin; PE2R2, prostaglandin E2 receptor EP2 subtype; PDE4D, 3′,5′-cyclic-AMP phosphodiesterase 4D; 5NTC, cytosolic purine 5′-nucleotidase; PLMN, plasminogen; ESR1, estrogen receptor; ANDR, androgen receptor; ADA2C, alpha-2C adrenergic receptor; DCAM, S-adenosylmethionine decarboxylase proenzyme; COMT, catechol O-methyltransferase; TOP2B, DNA topoisomerase 2-beta; EGFR, epidermal growth factor receptor; SCN9A, sodium channel protein type 9 subunit alpha; QCR6, cytochrome b-c1 complex subunit 6, mitochondrial; CY1, cytochrome c1, heme protein, mitochondrial; CSF1R, macrophage colony-stimulating factor 1 receptor; DDR1, epithelial discoidin domain-containing receptor 1; PGFRB, platelet-derived growth factor receptor beta; EPHA2, ephrin type-A receptor 2.
Additionally, some of the human HRPs can be also be targeted for host-directed therapies by using alternate approaches. One approach to achieve this is to employ a targeted protein degradation (TPD) method, using proteolysis-targeting chimeras (PROTACs),57,58 which utilize the ubiquitin-proteasome system to degrade the target protein (Figure 6B). Owing to the impediments resulting from the presence of broad active sites, shallow pockets, and/or “smooth” surfaces that offer few binding sites, proteins cannot be targeted using small molecule antibiotics. TPD is an effective strategy to modulate such difficult proteins.57 For instance, degraders are available for hHRPs involved in HPIs such as androgen receptor (ANDR), estrogen receptor (ESR1), and EGFR, which are in clinical and pre-clinical trials.57 We propose that human HRs along with their flanks can be designated targets to abrogate human-pathogen interactions. Especially in case of host-partner tract mimicry of human HRs by pathogens, HR-based targeting can be used as a potential therapeutic intervention strategy. Similarly, pathogen-directed therapies involving TPD, especially for eukaryotic and viral pathogens that bring about disparate interactions with human proteins by engaging host-sparse HRs, can also be developed, with minimal impact on host system. For instance, the degrader DGY-08-097 inhibits and degrades the hepatitis C virus (HCV) NS3/4A protease, thereby acting as an anti-viral agent, even against viral strains that show resistance to the conventional enzymatic inhibitor telaprevir.58,59 We posit that the HRs with flanking sequences could be effective targets for PROTAC-mediated TPD. In this regard, we observed that the sequences flanking the repeats on either side of pathogen proteins show very less similarity with human proteins (Figure 6C). This makes such unique sequences encompassing the HRs and the flanks ideal candidates to be targeted by PROTACs, with high specificity and low cross-reactivity with host proteins. To propel research on the role of HRs in pathogenicity and virulence of human infectious agents, we have developed a database, named Hi-PHI (Homorepeats in Pathogen-Human Interactions; http://hiphi.iisertirupati.ac.in/), which catalogs all the data generated and analyzed in this study (Figure S5). Furthermore, Hi-PHI provides the percentage sequence identity of pathogen HRs, along with their flanks, with the HRs and flanks in the entire human proteome (Figure 6C) to facilitate designing pathogen-directed targeting approaches. Notably, large pharmacological screens such as those identifying epitopes for widely used antibodies for disease treatment or to predict cross-reactivity60,61,62,63 often do not include HRs, possibly due to the inherent difficulties associated with studying such low-complexity sequences. We anticipate that the results and the datasets provided here will aid in designing studies that target human HRs and/or pathogen HRs to abrogate host-pathogen interactions, thereby attenuating infections.
Limitations of the study
The findings in this study are contingent to the large-scale datasets integrated and analyzed, which might influence some of the interpretations drawn here. This includes the number of HPIs, the density of interactions per pathogen identified thus far, and the availability of reliable protein structural data. Furthermore, through system-level analyses, we obtained the general trends, and every trend reported might not be applicable to each protein in every species across the different taxa studied here. While the network dismantling and phenotypic screen analyses provide important insights, they essentially do not present an exhaustive picture of the impact of engagement of hHRPs by pathogens on human physiology. Although this study provides a comprehensive account of the different modes of HRP engagement in HPIs and potential roles of HRs, targeted experimental investigations are required to elucidate the functional roles of both hHRs and pHRs.
Resource availability
Lead contact
Further information and requests for resources should be direct to and will be fulfilled by the lead contact, Sreenivas Chavali (schavali@iisertirupati.ac.in).
Materials availability
This study did not generate any new unique reagents.
Data and code availability
All data are available in the main text or the supplemental information. All data generated and analyzed in this study have been deposited as a resource, Hi-PHI (Homorepeats in Pathogen-Human Interactions), accessible at http://hiphi.iisertirupati.ac.in/. The HPIs and all codes used in this study have been deposited in Mendeley Data: https://doi.org/10.17632/nzk4swk7xy.1. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Acknowledgments
We thank P.L. Chavali, A. Dhayalan, and A.D. Allu for their feedback on the manuscript. This study was supported by IISER Tirupati core funding (to V.S., A.A., R.V.K., and S.C.); Prime Minister’s Research Fellowship, Government of India (to A.K.S. and N.R.); Ramalingaswami Re-entry Fellowship BT/RLF/Re-entry/05/2018 from Department of Biotechnology, Government of India (to R.V.K. and S.C.); INSPIRE Fellowship from Department of Science and Technology, Government of India (to V.S.); Junior/Senior Research Fellowships from Department of Biotechnology, Government of India (to K.S.K.); and Core Research Grant CRG/2023/004691 from Anusandhan National Research Foundation, Government of India (to A.K.S. and S.C.).
Author contributions
Conceptualization, S.C.; methodology, S.C. and A.K.S.; formal analysis, A.K.S., N.R., V.S., A.A., and K.S.K.; visualization, A.K.S., N.R., V.S., R.V.K., and S.C.; writing – original draft: S.C. and A.K.S.; writing – review & editing, S.C., A.K.S., N.R., V.S., A.A., K.S.K., and R.V.K.; supervision, S.C.
Declaration of interests
The authors declare that they have no competing interests.
STAR★Methods
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Deposited data | ||
| Human-pathogen protein interactions | This paper | http://hiphi.iisertirupati.ac.in/ |
| Human protein-drug interactions | This paper | http://hiphi.iisertirupati.ac.in/ |
| Percentage identity (PID) of pathogen and human HR flanks | This paper | http://hiphi.iisertirupati.ac.in/ |
| Software and algorithms | ||
| Adobe illustrator | Adobe | https://www.adobe.com/in/products/illustrator.html |
| Cytoscape 3.0 | Shannon P et al.64 | https://cytoscape.org/ |
| R | R Foundation for Statistical Computing | https://www.r-project.org/ |
| R: ggplot2 package | R Foundation for Statistical Computing | https://ggplot2.tidyverse.org |
| R: Biostrings package | R Foundation for Statistical Computing | https://bioconductor.org/packages/Biostrings |
| R: fuzzySim package | R Foundation for Statistical Computing | https://doi.org/10.1111/2041-210X.12372 |
| R: reshape2 package | R Foundation for Statistical Computing | http://www.jstatsoft.org/v21/i12/ |
| R: splitstackshape package | R Foundation for Statistical Computing | https://CRAN.R-project.org/package=splitstackshape |
| R: stringr package | R Foundation for Statistical Computing | https://CRAN.R-project.org/package=stringr |
| R: dplyr package | R Foundation for Statistical Computing | https://CRAN.R-project.org/package=dplyr |
| Python | Python Software Foundation | https://www.python.org/ |
| Python: pandas library | Python Software Foundation | https://pandas.pydata.org/ |
| Python: numpy library | Python Software Foundation | https://numpy.org/ |
| Python: Biopython library | Python Software Foundation | https://biopython.org/ |
| Python: openpyxl library | Python Software Foundation | https://openpyxl.readthedocs.io/en/stable/ |
| Python: networkx library | Python Software Foundation | https://networkx.org/ |
Method details
Assembly of human-pathogen interactome
Experimentally determined human-pathogen protein interactions (HPI) were assembled from HPIDB 3.0,65,66 MorCVD,67 HoPaCI-DB,68 BIOGRID69 and a high-throughput study for human-SARS-CoV-2 interactions70 [Table S1]. These were mapped to subset one strain per pathogen species (except for viruses), retaining the strain with the maximum number of interactions with human proteins. We have removed duplicate interactions and have retained only unique interactions between a pathogen and a human protein. The source article(s) for each HPI has been provided in the Hi-PHI database. The human and pathogen protein sequences were retrieved from UniProt. Pathogen taxonomy information was extracted from NCBI Taxonomy browser.71 An in-house Perl script was used to identify homorepeats (HR; amino acid type, length and location in the protein) in the human and pathogen proteins. A protein was considered as an HRP if it had at least one HR stretch with a minimum length of 5 residues, as previously defined.9,12
Gene ontology enrichment analysis
Assessment of enriched Gene Ontology-Biological Processes (GO-BP) was done by DAVID server using an FDR cut-off of < 0.05.72,73 The enriched GO-BP terms were then manually organized into major biological processes. We computed the extent of overlapping genes between any two enriched processes using Jaccard Similarity Index, which represents the ratio of the number of genes common to both biological processes to the sum of all genes in the two processes, as previously described.12
Assembly of human proteome-level datasets
Dataset for human (i) essential genes/proteins, (ii) tissue-wide expression of proteins, (iii) physical protein-protein interaction network, (iv) protein complexes, (v) transcription factors and RNA binding proteins and (vi) phase-separated condensates were obtained from previously published study12 [Table S2]. Human drug-target network for hHPI proteins was reconstructed using data obtained from the drug-target network47 and STITCH database.74 Drug-target interactions from STITCH with binding data were considered here. The dataset was curated for drugs that target hHPI proteins involved in diseases caused by cellular and viral pathogens. The curated dataset comprised of 251 human proteins which were targets of 179 drugs. The functional classification of drugs based on their mode of action was obtained from DrugBank.75 We obtained the variable human HRs from a large-scale study,32 which mapped HR-associated variability using 125,748 human exomes.
Assembly of pathogen-related datasets
We compiled the disease caused by each pathogen and the number of affected organs and organ systems through extensive literature search. The information for viral pathogens was curated from Human Virus Database (HVD).76 We obtained a list of 20 bacterial (in 2024) and 19 fungal (in 2022) priority pathogens identified by the World Health Organization (WHO; https://www.who.int/). This categorization of priority pathogens is based on parameters such as mortality, incidence, non-fatal health burden, transmissibility, preventability and treatability. We also obtained the list of biodefense pathogens, those that cause diseases of high public health concern, from National Institute of Allergy and Infectious Diseases (NIAID; https://www.niaid.nih.gov/).
Network analysis
We computed the topological properties of human-pathogen protein interaction network and human protein-protein interaction network using Cytoscape 3.064 and in-house Python and R scripts. Using an in-house written python script, we performed network dismantling of human protein-protein interaction network, by removing one node (protein) at a time. This led to the generation of 17,144 networks, each of which lacked one human protein. To assess the impact of a protein on protein interaction network, we estimated the network topological properties such as link density and assortativity for the network resulting from the removal of that protein and compared the network properties of the real network.
Assembly of large-scale phenotypic screens
We assembled (i) CRISPRi-mediated genome-wide phenotypic screen comprising of 18,489 human coding genes involved in cell proliferation and differentiation during neural induction of human iPSCs to form neural stem cells from Wu et al.,77 and (ii) CRISPRa-mediated phenotypic screen comprising of 20,115 human genes and CRISPRi-mediated phenotypic screen comprising of 20,071 human genes involved in survival of neurons under oxidative stress (without anti-oxidant treatment) from Tian et al.78
Molecular docking of human-pathogen protein interactions
We identified 5,073 HPIs where atleast one of the interacting partner is an HRP. For the selected human-pathogen protein interactions, structural models were retrieved from the AlphaFold Protein Structure Database20 using UniProt accession identifiers and Viro3D21 using Genbank ID for viruses. This resulted in a dataset comprising of 758 hHPI-HRPs, 316 hHPI-NonHRPs, 125 epHPI-HRPs, 138 epHPI-NonHRPs, 34 ppHPI-HRPs, 681 ppHPI-NonHRPs, 33 vpHPI-HRPs, 150 vpHPI-NonHRPs involved in 3,265 HPIs. Each structure was processed to remove low-confidence terminal regions by applying a sliding-window trimming approach based on per-residue predicted local distance difference test (pLDDT) scores. Terminal residues with pLDDT values below 50 were excluded. We calculated both the sequence length and average pLDDT score before and after trimming to assess structural integrity. Only trimmed structures with an average pLDDT≥50 and a sequence length between 50 and 1000 residues were considered for this study. The curated structures, which included 1,501 human-pathogen protein interactions between 479 hHPI-HRPs, 173 hHPI-NonHRPs, 80 epHPI-HRPs, 124 epHPI-NonHRPs, 25 ppHPI-HRPs, 437 ppHPI-NonHRPs, 18 vpHPI-HRP and 100 vpHPI-NonHRPs, were used for protein-protein docking. We performed blind molecular docking using ZDOCK 3.0.2/2.3.222, generating 100 poses per complex. The top-scoring conformation, based on the lowest predicted energy, was selected for further analysis [Table S3]. The interaction interface was mapped using PDBParser from Biopython for structural parsing. We used the interaction criterion for any atom-atom distance cut-off as ≤5Å (default) to predict the interaction interface, using Biopython NeighborSearch module. We then mapped this to the HR residue information to identify HRs that participate in the interaction interface. We picked up reliable docked structures of HRs in interaction interfaces using a cut-off of average pLDDT≥50 of the residues involved in HR-mediated interactions [Table S4].
Estimation of percentage sequence identity
We retrieved 20 residue flanking the N-terminus and C-terminus of each HR in the human and pathogen proteins. This resulted in 5,412 N-terminal and 5,963 C-terminal flank sequences of hHRs and 342 N-terminal and 332 C-terminal flank sequences of pHRs. We disregarded flanks with <20 residues. For each flanking sequence of a pathogen HR, we computed pairwise percentage sequence identity using an in-house written R script (defined as the ratio of number of aligned residues to total length of the alignment) with corresponding flank of all human proteins with the same amino acid HR type. For instance, pair-wise percentage sequence identity of the N-terminal flank of a pathogen polyGly HR was estimated with the N-terminal flank of all human polyGly HRs. We could not perform pair-wise alignments of flanks with <10 aligning residues from either pathogen or human proteins. This led to 117,846 (out of 129,372) and 122,978 (out of 135,914) pairwise alignments of N-terminal and C-terminal flanks, respectively.
Quantification and statistical analysis
Statistical significance for the differences in the distribution of discrete variables was assessed using Fisher’s Exact test and that of continuous variables was estimated using the non-parametric Wilcoxon rank sum test. Enrichment of proteins in different classes (for instance, human HRPs in HPI) was examined using permutation tests by performing 10,000 randomizations. In each permutation, the protein of interest (e.g., human proteins that participate in HPI) was replaced with a random protein and the number of random proteins that overlapped with a specific class of proteins (e.g., human HRPs) was noted for all the randomizations. We estimated Z score and P value from the distribution of overlapping proteins from the random expectations. Z score represents the magnitude of deviation of the actual observation compared to the mean of random distribution in terms of the number of standard deviations. P values for enrichment were estimated as the ratios of the number of the randomly observed proteins ≥ to the number of actually observed proteins to the total number of randomized samples (10,000). All statistical analyses were performed using R. The details pertaining to the statistical tests undertaken are provided in the figure legends, while the number of datapoints analyzed and the statistical significance estimates are provided in the figures.
Published: March 5, 2026
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.isci.2026.115249.
Contributor Information
Rajashekar Varma Kadumuri, Email: rajashekar@labs.iisertirupati.ac.in.
Sreenivas Chavali, Email: schavali@iisertirupati.ac.in.
Supplemental information
References
- 1.Speranza E. Understanding virus-host interactions in tissues. Nat. Microbiol. 2023;8:1397–1407. doi: 10.1038/s41564-023-01434-7. [DOI] [PubMed] [Google Scholar]
- 2.Davey N.E., Travé G., Gibson T.J. How viruses hijack cell regulation. Trends Biochem. Sci. 2011;36:159–169. doi: 10.1016/j.tibs.2010.10.002. [DOI] [PubMed] [Google Scholar]
- 3.Karabadzhak A.G., Petti L.M., Barrera F.N., Edwards A.P.B., Moya-Rodríguez A., Polikanov Y.S., Freites J.A., Tobias D.J., Engelman D.M., DiMaio D. Two transmembrane dimers of the bovine papillomavirus E5 oncoprotein clamp the PDGF beta receptor in an active dimeric conformation. Proc. Natl. Acad. Sci. USA. 2017;114:E7262–E7271. doi: 10.1073/pnas.1705622114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Goldstein D.J., Andresson T., Sparkowski J.J., Schlegel R. The BPV-1 E5 protein, the 16 kDa membrane pore-forming protein and the PDGF receptor exist in a complex that is dependent on hydrophobic transmembrane interactions. EMBO J. 1992;11:4851–4859. doi: 10.1002/j.1460-2075.1992.tb05591.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Scarth J.A., Patterson M.R., Morgan E.L., Macdonald A. The human papillomavirus oncoproteins: a review of the host pathways targeted on the road to transformation. J. Gen. Virol. 2021;102 doi: 10.1099/jgv.0.001540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ashrafi G.H., Haghshenas M., Marchetti B., Campo M.S. E5 protein of human papillomavirus 16 downregulates HLA class I and interacts with the heavy chain via its first hydrophobic domain. Int. J. Cancer. 2006;119:2105–2112. doi: 10.1002/ijc.22089. [DOI] [PubMed] [Google Scholar]
- 7.Machesky L.M., Insall R.H., Volkman L.E. WASP homology sequences in baculoviruses. Trends Cell Biol. 2001;11:286–287. doi: 10.1016/s0962-8924(01)02009-8. [DOI] [PubMed] [Google Scholar]
- 8.Pery E., Rajendran K.S., Brazier A.J., Gabuzda D. Regulation of APOBEC3 proteins by a novel YXXL motif in human immunodeficiency virus type 1 Vif and simian immunodeficiency virus SIVagm Vif. J. Virol. 2009;83:2374–2381. doi: 10.1128/JVI.01898-08. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Chavali S., Chavali P.L., Chalancon G., de Groot N.S., Gemayel R., Latysheva N.S., Ing-Simmons E., Verstrepen K.J., Balaji S., Babu M.M. Constraints and consequences of the emergence of amino acid repeats in eukaryotic proteins. Nat. Struct. Mol. Biol. 2017;24:765–777. doi: 10.1038/nsmb.3441. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Chavali S., Singh A.K., Santhanam B., Babu M.M. Amino acid homorepeats in proteins. Nat. Rev. Chem. 2020;4:420–434. doi: 10.1038/s41570-020-0204-1. [DOI] [PubMed] [Google Scholar]
- 11.Gemayel R., Chavali S., Pougach K., Legendre M., Zhu B., Boeynaems S., van der Zande E., Gevaert K., Rousseau F., Schymkowitz J., et al. Variable Glutamine-Rich Repeats Modulate Transcription Factor Activity. Mol. Cell. 2015;59:615–627. doi: 10.1016/j.molcel.2015.07.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Singh A.K., Amar I., Ramadasan H., Kappagantula K.S., Chavali S. Proteins with amino acid repeats constitute a rapidly evolvable and human-specific essentialome. Cell Rep. 2023;42 doi: 10.1016/j.celrep.2023.112811. [DOI] [PubMed] [Google Scholar]
- 13.Bedru H.D., Yu S., Xiao X., Zhang D., Wan L., Guo H., Xia F. Big networks: A survey. Comput. Sci. Rev. 2020;37 doi: 10.1016/j.cosrev.2020.100247. [DOI] [Google Scholar]
- 14.Foster D.V., Foster J.G., Grassberger P., Paczuski M. Clustering drives assortativity and community structure in ensembles of networks. Phys. Rev. E - Stat. Nonlinear Soft Matter Phys. 2011;84 doi: 10.1103/PhysRevE.84.066117. [DOI] [PubMed] [Google Scholar]
- 15.Thakur J.K., Arthanari H., Yang F., Pan S.J., Fan X., Breger J., Frueh D.P., Gulshan K., Li D.K., Mylonakis E., et al. A nuclear receptor-like pathway regulating multidrug resistance in fungi. Nature. 2008;452:604–609. doi: 10.1038/nature06836. [DOI] [PubMed] [Google Scholar]
- 16.Dyer M.D., Neff C., Dufford M., Rivera C.G., Shattuck D., Bassaganya-Riera J., Murali T.M., Sobral B.W. The human-bacterial pathogen protein interaction networks of Bacillus anthracis, Francisella tularensis, and Yersinia pestis. PLoS One. 2010;5 doi: 10.1371/journal.pone.0012089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Zhang Y., Bliska J.B. Role of Toll-like receptor signaling in the apoptotic response of macrophages to Yersinia infection. Infect. Immun. 2003;71:1513–1519. doi: 10.1128/IAI.71.3.1513-1519.2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Reitsma J.M., Savaryn J.P., Faust K., Sato H., Halligan B.D., Terhune S.S. Antiviral inhibition targeting the HCMV kinase pUL97 requires pUL27-dependent degradation of Tip60 acetyltransferase and cell-cycle arrest. Cell Host Microbe. 2011;9:103–114. doi: 10.1016/j.chom.2011.01.006. [DOI] [PubMed] [Google Scholar]
- 19.Tran K., Mahr J.A., Spector D.H. Proteasome subunits relocalize during human cytomegalovirus infection, and proteasome activity is necessary for efficient viral gene transcription. J. Virol. 2010;84:3079–3093. doi: 10.1128/JVI.02236-09. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Fleming J., Magana P., Nair S., Tsenkov M., Bertoni D., Pidruchna I., Lima Afonso M.Q., Midlik A., Paramval U., Žídek A., et al. AlphaFold Protein Structure Database and 3D-Beacons: New Data and Capabilities. J. Mol. Biol. 2025;437 doi: 10.1016/j.jmb.2025.168967. [DOI] [PubMed] [Google Scholar]
- 21.Litvin U., Lytras S., Jack A., Robertson D.L., Hughes J., Grove J. Viro3D: a comprehensive database of virus protein structure predictions. Mol. Syst. Biol. 2025;21:1599–1617. doi: 10.1038/s44320-025-00147-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Pierce B.G., Hourai Y., Weng Z. Accelerating protein docking in ZDOCK using an advanced 3D convolution library. PLoS One. 2011;6 doi: 10.1371/journal.pone.0024657. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ferron F., Rebowski G., Lee S.H., Dominguez R. Structural basis for the recruitment of profilin-actin complexes during filament elongation by Ena/VASP. EMBO J. 2007;26:4597–4606. doi: 10.1038/sj.emboj.7601874. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Baek K., Liu X., Ferron F., Shu S., Korn E.D., Dominguez R. Modulation of actin structure and function by phosphorylation of Tyr-53 and profilin binding. Proc. Natl. Acad. Sci. USA. 2008;105:11748–11753. doi: 10.1073/pnas.0805852105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Gill J., Kumar A., Yogavel M., Belrhali H., Jain S.K., Rug M., Brown M., Maier A.G., Sharma A. Structure, localization and histone binding properties of nuclear-associated nucleosome assembly protein from Plasmodium falciparum. Malar. J. 2010;9:90. doi: 10.1186/1475-2875-9-90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Wallqvist A., Wang H., Zavaljevski N., Memišević V., Kwon K., Pieper R., Rajagopala S.V., Reifman J. Mechanisms of action of Coxiella burnetii effectors inferred from host-pathogen protein interactions. PLoS One. 2017;12 doi: 10.1371/journal.pone.0188071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Muller M., Jacob Y., Jones L., Weiss A., Brino L., Chantier T., Lotteau V., Favre M., Demeret C. Large scale genotype comparison of human papillomavirus E2-host interaction networks provides new insights for e2 molecular functions. PLoS Pathog. 2012;8 doi: 10.1371/journal.ppat.1002761. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Hagai T., Azia A., Babu M.M., Andino R. Use of host-like peptide motifs in viral proteins is a prevalent strategy in host-virus interactions. Cell Rep. 2014;7:1729–1739. doi: 10.1016/j.celrep.2014.04.052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Rug M., Prescott S.W., Fernandez K.M., Cooke B.M., Cowman A.F. The role of KAHRP domains in knob formation and cytoadherence of P falciparum-infected human erythrocytes. Blood. 2006;108:370–378. doi: 10.1182/blood-2005-11-4624. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Cutts E.E., Laasch N., Reiter D.M., Trenker R., Slater L.M., Stansfeld P.J., Vakonakis I. Structural analysis of P. falciparum KAHRP and PfEMP1 complexes with host erythrocyte spectrin suggests a model for cytoadherent knob protrusions. PLoS Pathog. 2017;13 doi: 10.1371/journal.ppat.1006552. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Davies H.M., Nofal S.D., McLaughlin E.J., Osborne A.R. Repetitive sequences in malaria parasite proteins. FEMS Microbiol. Rev. 2017;41:923–940. doi: 10.1093/femsre/fux046. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Mier P., Andrade-Navarro M.A., Morett E. Homorepeat variability within the human population. NAR Genom. Bioinf. 2024;6 doi: 10.1093/nargab/lqae053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Zhou K., Aertsen A., Michiels C.W. The role of variable DNA tandem repeats in bacterial adaptation. FEMS Microbiol. Rev. 2014;38:119–141. doi: 10.1111/1574-6976.12036. [DOI] [PubMed] [Google Scholar]
- 34.Gemayel R., Cho J., Boeynaems S., Verstrepen K.J. Beyond junk-variable tandem repeats as facilitators of rapid evolution of regulatory and coding sequences. Genes. 2012;3:461–480. doi: 10.3390/genes3030461. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Verstrepen K.J., Reynolds T.B., Fink G.R. Origins of variation in the fungal cell surface. Nat. Rev. Microbiol. 2004;2:533–540. doi: 10.1038/nrmicro927. [DOI] [PubMed] [Google Scholar]
- 36.Hoyer L.L. The ALS gene family of Candida albicans. Trends Microbiol. 2001;9:176–180. doi: 10.1016/s0966-842x(01)01984-9. [DOI] [PubMed] [Google Scholar]
- 37.Tanudisastro H.A., Deveson I.W., Dashnow H., MacArthur D.G. Sequencing and characterizing short tandem repeats in the human genome. Nat. Rev. Genet. 2024;25:460–475. doi: 10.1038/s41576-024-00692-3. [DOI] [PubMed] [Google Scholar]
- 38.Xiao X., Zhang C.Y., Zhang Z., Hu Z., Li M., Li T. Revisiting tandem repeats in psychiatric disorders from perspectives of genetics, physiology, and brain evolution. Mol. Psychiatr. 2022;27:466–475. doi: 10.1038/s41380-021-01329-1. [DOI] [PubMed] [Google Scholar]
- 39.Malik I., Kelley C.P., Wang E.T., Todd P.K. Molecular mechanisms underlying nucleotide repeat expansion disorders. Nat. Rev. Mol. Cell Biol. 2021;22:589–607. doi: 10.1038/s41580-021-00382-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Wheeler V.C., Dion V. Modifiers of CAG/CTG Repeat Instability: Insights from Mammalian Models. J. Huntingtons Dis. 2021;10:123–148. doi: 10.3233/JHD-200426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Deshmukh A.L., Porro A., Mohiuddin M., Lanni S., Panigrahi G.B., Caron M.C., Masson J.Y., Sartori A.A., Pearson C.E. FAN1, a DNA Repair Nuclease, as a Modifier of Repeat Expansion Disorders. J. Huntingtons Dis. 2021;10:95–122. doi: 10.3233/JHD-200448. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Hannan A.J. Tandem repeats mediating genetic plasticity in health and disease. Nat. Rev. Genet. 2018;19:286–298. doi: 10.1038/nrg.2017.115. [DOI] [PubMed] [Google Scholar]
- 43.Hannan A.J. Tandem repeat polymorphisms: modulators of disease susceptibility and candidates for 'missing heritability'. Trends Genet. 2010;26:59–65. doi: 10.1016/j.tig.2009.11.008. [DOI] [PubMed] [Google Scholar]
- 44.Hariharan A., Hakeem A.R., Radhakrishnan S., Reddy M.S., Rela M. The Role and Therapeutic Potential of NF-kappa-B Pathway in Severe COVID-19 Patients. Inflammopharmacology. 2021;29:91–100. doi: 10.1007/s10787-020-00773-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Shankar S., Singh G., Srivastava R.K. Chemoprevention by resveratrol: molecular mechanisms and therapeutic potential. Front. Biosci. 2007;12:4839–4854. doi: 10.2741/2432. [DOI] [PubMed] [Google Scholar]
- 46.Pfalzgraff A., Brandenburg K., Weindl G. Antimicrobial Peptides and Their Therapeutic Potential for Bacterial Skin Infections and Wounds. Front. Pharmacol. 2018;9:281. doi: 10.3389/fphar.2018.00281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Yildirim M.A., Goh K.I., Cusick M.E., Barabási A.L., Vidal M. Drug-target network. Nat. Biotechnol. 2007;25:1119–1126. doi: 10.1038/nbt1338. [DOI] [PubMed] [Google Scholar]
- 48.Dunn C.J., Goa K.L. Tranexamic acid: a review of its use in surgery and other indications. Drugs. 1999;57:1005–1032. doi: 10.2165/00003495-199957060-00017. [DOI] [PubMed] [Google Scholar]
- 49.Jansen A.J., Andreica S., Claeys M., D'Haese J., Camu F., Jochmans K. Use of tranexamic acid for an effective blood conservation strategy after total knee arthroplasty. Br. J. Anaesth. 1999;83:596–601. doi: 10.1093/bja/83.4.596. [DOI] [PubMed] [Google Scholar]
- 50.Hanson A.J., Quinn M.T. Effect of fibrin sealant composition on human neutrophil chemotaxis. J. Biomed. Mater. Res. 2002;61:474–481. doi: 10.1002/jbm.10196. [DOI] [PubMed] [Google Scholar]
- 51.Miró-Canturri A., Ayerbe-Algaba R., Smani Y. Drug Repurposing for the Treatment of Bacterial and Fungal Infections. Front. Microbiol. 2019;10:41. doi: 10.3389/fmicb.2019.00041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Kulkarni V.S., Alagarsamy V., Solomon V.R., Jose P.A., Murugesan S. Drug Repurposing: An Effective Tool in Modern Drug Discovery. Russ. J. Bioorg. Chem. 2023;49:157–166. doi: 10.1134/S1068162023020139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Schein C.H. Repurposing approved drugs on the pathway to novel therapies. Med. Res. Rev. 2020;40:586–605. doi: 10.1002/med.21627. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Jean S.S., Hsueh P.R. Old and re-purposed drugs for the treatment of COVID-19. Expert Rev. Anti Infect. Ther. 2020;18:843–847. doi: 10.1080/14787210.2020.1771181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Ng Y.L., Salim C.K., Chu J.J.H. Drug repurposing for COVID-19: Approaches, challenges and promising candidates. Pharmacol. Ther. 2021;228 doi: 10.1016/j.pharmthera.2021.107930. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Pushpakom S., Iorio F., Eyers P.A., Escott K.J., Hopper S., Wells A., Doig A., Guilliams T., Latimer J., McNamee C., et al. Drug repurposing: progress, challenges and recommendations. Nat. Rev. Drug Discov. 2019;18:41–58. doi: 10.1038/nrd.2018.168. [DOI] [PubMed] [Google Scholar]
- 57.Békés M., Langley D.R., Crews C.M. PROTAC targeted protein degraders: the past is prologue. Nat. Rev. Drug Discov. 2022;21:181–200. doi: 10.1038/s41573-021-00371-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Liu Z., Hu M., Yang Y., Du C., Zhou H., Liu C., Chen Y., Fan L., Ma H., Gong Y., Xie Y. An overview of PROTACs: a promising drug discovery paradigm. Mol. Biomed. 2022;3:46. doi: 10.1186/s43556-022-00112-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.de Wispelaere M., Du G., Donovan K.A., Zhang T., Eleuteri N.A., Yuan J.C., Kalabathula J., Nowak R.P., Fischer E.S., Gray N.S., Yang P.L. Small molecule degraders of the hepatitis C virus protease reduce susceptibility to resistance mutations. Nat. Commun. 2019;10:3468. doi: 10.1038/s41467-019-11429-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Matsumoto K., Harada S.Y., Yoshida S.Y., Narumi R., Mitani T.T., Yada S., Sato A., Morii E., Shimizu Y., Ueda H.R. DECODE enables high-throughput mapping of antibody epitopes at single amino acid resolution. PLoS Biol. 2025;23 doi: 10.1371/journal.pbio.3002707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Rockberg J., Löfblom J., Hjelm B., Uhlén M., Ståhl S. Epitope mapping of antibodies using bacterial surface display. Nat. Methods. 2008;5:1039–1045. doi: 10.1038/nmeth.1272. [DOI] [PubMed] [Google Scholar]
- 62.Moreira G.M.S.G., Fühner V., Hust M. Epitope Mapping by Phage Display. Methods Mol. Biol. 2018;1701:497–518. doi: 10.1007/978-1-4939-7447-4_28. [DOI] [PubMed] [Google Scholar]
- 63.Reyes S.G., Kuruma Y., Fujimi M., Yamazaki M., Eto S., Nishikawa S., Tamaki S., Kobayashi A., Mizuuchi R., Rothschild L., et al. PURE mRNA display and cDNA display provide rapid detection of core epitope motif via high-throughput sequencing. Biotechnol. Bioeng. 2021;118:1736–1749. doi: 10.1002/bit.27696. [DOI] [PubMed] [Google Scholar]
- 64.Shannon P., Markiel A., Ozier O., Baliga N.S., Wang J.T., Ramage D., Amin N., Schwikowski B., Ideker T. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13:2498–2504. doi: 10.1101/gr.1239303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Ammari M.G., Gresham C.R., McCarthy F.M., Nanduri B. HPIDB 2.0: a curated database for host-pathogen interactions. Database. 2016;2016 doi: 10.1093/database/baw103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Kumar R., Nanduri B. HPIDB--a unified resource for host-pathogen interactions. BMC Bioinf. 2010;11 doi: 10.1186/1471-2105-11-S6-S16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Singh N., Bhatia V., Singh S., Bhatnagar S. MorCVD: A Unified Database for Host-Pathogen Protein-Protein Interactions of Cardiovascular Diseases Related to Microbes. Sci. Rep. 2019;9:4039. doi: 10.1038/s41598-019-40704-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Bleves S., Dunger I., Walter M.C., Frangoulidis D., Kastenmüller G., Voulhoux R., Ruepp A. HoPaCI-DB: host-Pseudomonas and Coxiella interaction database. Nucleic Acids Res. 2014;42:D671–D676. doi: 10.1093/nar/gkt925. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Oughtred R., Rust J., Chang C., Breitkreutz B.J., Stark C., Willems A., Boucher L., Leung G., Kolas N., Zhang F., et al. The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Sci. 2021;30:187–200. doi: 10.1002/pro.3978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Zhou Y., Liu Y., Gupta S., Paramo M.I., Hou Y., Mao C., Luo Y., Judd J., Wierbowski S., Bertolotti M., et al. A comprehensive SARS-CoV-2-human protein-protein interactome reveals COVID-19 pathobiology and potential host therapeutic targets. Nat. Biotechnol. 2023;41:128–139. doi: 10.1038/s41587-022-01474-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Sayers E.W., Bolton E.E., Brister J.R., Canese K., Chan J., Comeau D.C., Connor R., Funk K., Kelly C., Kim S., et al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 2022;50:D20–D26. doi: 10.1093/nar/gkab1112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Huang D.W., Sherman B.T., Lempicki R.A. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nat. Protoc. 2009;4:44–57. doi: 10.1038/nprot.2008.211. [DOI] [PubMed] [Google Scholar]
- 73.Sherman B.T., Hao M., Qiu J., Jiao X., Baseler M.W., Lane H.C., Imamichi T., Chang W. DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update) Nucleic Acids Res. 2022;50:W216–W221. doi: 10.1093/nar/gkac194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Kuhn M., von Mering C., Campillos M., Jensen L.J., Bork P. STITCH: interaction networks of chemicals and proteins. Nucleic Acids Res. 2008;36:D684–D688. doi: 10.1093/nar/gkm795. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Knox C., Wilson M., Klinger C.M., Franklin M., Oler E., Wilson A., Pon A., Cox J., Chin N.E.L., Strawbridge S.A., et al. DrugBank 6.0: the DrugBank Knowledgebase for 2024. Nucleic Acids Res. 2024;52:D1265–D1275. doi: 10.1093/nar/gkad976. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Ye S., Lu C., Qiu Y., Zheng H., Ge X., Wu A., Xia Z., Jiang T., Zhu H., Peng Y. An atlas of human viruses provides new insights into diversity and tissue tropism of human viruses. Bioinformatics. 2022;38:3087–3093. doi: 10.1093/bioinformatics/btac275. [DOI] [PubMed] [Google Scholar]
- 77.Wu D., Poddar A., Ninou E., Hwang E., Cole M.A., Liu S.J., Horlbeck M.A., Chen J., Replogle J.M., Carosso G.A., et al. Dual genome-wide coding and lncRNA screens in neural induction of induced pluripotent stem cells. Cell Genom. 2022;2 doi: 10.1016/j.xgen.2022.100177. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Tian R., Abarientos A., Hong J., Hashemi S.H., Yan R., Dräger N., Leng K., Nalls M.A., Singleton A.B., Xu K., et al. Genome-wide CRISPRi/a screens in human neurons link lysosomal failure to ferroptosis. Nat. Neurosci. 2021;24:1020–1034. doi: 10.1038/s41593-021-00862-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data are available in the main text or the supplemental information. All data generated and analyzed in this study have been deposited as a resource, Hi-PHI (Homorepeats in Pathogen-Human Interactions), accessible at http://hiphi.iisertirupati.ac.in/. The HPIs and all codes used in this study have been deposited in Mendeley Data: https://doi.org/10.17632/nzk4swk7xy.1. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.






