Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Aug 29;17:10318. doi: 10.1038/s41467-026-77295-5

Widespread genomic islands are hotspots of genome variations and mosaicism in giant viruses

Benjamin Minch 1, Mohammad Moniruzzaman 1,✉
PMCID: PMC13620141  PMID: 42805991

Abstract

Giant viruses in the phylum Nucleocytoviricota possess exceptionally large and mosaic genomes, yet the mechanisms underlying their remarkable plasticity remain poorly understood. Genomic islands are dynamic genomic regions that are major drivers of diversification and adaptation in bacteria. However, their contribution to giant virus evolution remains largely unexplored. Here, we characterize the genomic island landscape of giant viruses using 369 high-quality genomes spanning cultured isolates and long-read metagenome-assembled genomes. We identify 307 genomic islands across >50% of the genomes, demonstrating that these regions are pervasive across Nucleocytoviricota. These genomic islands are frequently associated with genomic hypervariability and enriched in genes involved in host interaction, particularly surface adhesion proteins, suggesting roles in host adaptation during the virus-host arms race. Comparative analyses further reveal these islands as hotspots of genome diversification, exhibiting frequent gain/loss and rearrangement even among highly similar genomes. Notably, many genomic islands are enriched in bacterial homologs, and several exhibit striking synteny with genomic regions recovered from co-occurring bacterial genomes, supporting large-scale genetic exchange between bacteria and giant viruses. Together, these findings identify genomic islands as pervasive and dynamic drivers of giant virus genome evolution, providing a framework for genome plasticity, mosaicism, and adaptive potential of giant viruses.

Subject terms: Microbial ecology, Viral evolution


Researchers analyzed 369 giant virus genomes and identified widespread genomic islands, regions of rapid gene turnover that help these viruses adapt to hosts and even exchange genes with bacteria, revealing new insight into viral genome evolution.

Introduction

Giant viruses in the phylum Nucleocytoviricota (NCVs) are a unique group of eukaryote-infecting viruses with both large capsid (up to 2 µm) and genome sizes (up to 2.5 Mbp)1–4. These viruses have redefined what was previously the limits of virus functional potential, as they have the ability to encode many genes previously unknown in the virosphere3,5. There are currently six identified orders of NCVs – Imitervirales, Algavirales, Pimascovirales, Asfuvirales, ‘Pandoravirales’, and Chitovirales6. These viruses are widespread in the global oceans, lakes, and sediments and have a seemingly broad host range, consisting mainly of single-celled eukaryotes or protists4,7. NCVs infecting protists have the potential to have a significant impact on planetary biogeochemical cycles due to both the large biomass of protists on the planet and the known potential of NCVs to modulate host metabolism and carbon fates5,8,9.

While their ecological significance has been well established, many questions remain regarding the large size of NCV genomes, as they are stand-outs in the virosphere10,11. Most viruses exist as largely successful replicators due to their small genome size, allowing for quick and efficient replication12,13. Even viruses infecting complex organisms such as mammals have quite streamlined genomic resources available to them14,15. This is not the case for NCVs, as their genomes are often larger than 200 kilobase pairs (kbp) and contain many proteins that do not seem immediately useful for viral replication3,16. This gigantism of NCV genomes has evolved multiple times throughout their evolutionary history, and it is thought that the pattern of gene gain and loss follows an “accordion model”17,18. These gene gains are thought to be primarily due to gene transfer from other viruses (viral HGT), gene duplications, and originations19.

On top of having large genomes, NCV genomes are often seen as mosaic, comprising many genes from diverse cellular lineages as well as a large number of ORFan genes2,20,21. The large amount of genetic material from cellular sources within NCV genomes initially led researchers to hypothesize that these viruses derived from an ancient cellular ancestor, an idea that has since lost favor11,22,23. While we know that NCVs acquire many genes through horizontal gene transfer (HGT), we still lack a clear understanding of the genomic constraints and ecological interactions facilitating these acquisitions.

From extensive research on HGT in bacteria, we know that gene exchange can occur on the level of individual genes as well as in large chunks in the form of genomic islands24–26. Genomic islands are well characterized in bacteria as large segments of DNA often harboring many horizontally transferred genes24,25. These regions are also usually distinct in nucleotide composition signatures compared to the surrounding genome and are enriched in genes involved in pathogenesis, symbiosis, antibiotic resistance, or response to different environmental factors24,26. Genomic islands often play a crucial role in the evolution and adaptability of bacteria as they introduce genetic diversity and improve fitness in changing environments25. These regions are typically fast-evolving, with higher mutation rates and greater variability than the surrounding genome. In this context, hypervariability refers to the fact that rapid gain or loss of genes and a fast-evolving gene compendium make islands within closely related bacterial strains slightly different from each other24,27.

Since genomic islands (10–200 kbp)28 are often as large or larger than many viral genomes, we often don’t associate them with viral populations. The discovery of NCVs, however, calls for a reevaluation of the role of genomic islands in viruses, as these viruses have sufficiently large genomes to contain island regions. Since their discovery in 2003, there have been a few examples of putative hypervariable regions found within a few cultured NCV genomes29–32, as well as metagenomic islands (regions of the genome with anomalous metagenomic read recruitment) in uncultured genomes33, suggesting that regions with features characteristic of genomic islands exist in NCVs. However, their distribution across NCVs, functional characteristics, and ecological significance in NCVs has largely remained elusive. Part of this is due to the majority of NCV genomes coming from short-read, metagenomic binning of NCVs3,4,34, which is a barrier against assembling NCV genomes with high contiguity. However, recent developments in long-read sequencing represent a significant advance in reducing this barrier, given the possibility that this technology can uncover complete or near-complete genomes from metagenomic datasets.

Given that there is little information on the origin and functional composition of these islands in NCVs, a method is necessary to identify them that is function-agnostic but still rooted in the fundamental nature of these elements, such as deviated nucleotide composition. Motivated by this, we developed a tool to identify these regions in NCVs based on compositional deviations from the surrounding genomic context. Here, we apply this tool on a curated set of long-read and cultured representative NCV genomes to uncover the genomic island landscape in NCVs35,36. We demonstrate that genomic islands are widespread among all orders of Nucleocytoviricota and are enriched in surface adhesion, replication, and metabolic genes, providing evidence of their use as host-adaptation mechanisms. Additionally, by leveraging long-read genomic data, we show hypervariability in these regions among closely related strains, suggesting that genomic islands are concentrated regions of gene-content variability in giant viruses. We also demonstrate possible HGT of these islands across strains. This likely facilitates their host adaptation. Furthermore, we show that many genomic islands harbor genes highly identical to bacterial homologs, indicating a strong bacterial contribution to their gene content. Notably, in several cases, these islands exhibit remarkable synteny and gene content similarity to genomic regions of bacteria that co-occur within the same environmental sample, providing evidence of large-scale transfer of genomic regions between bacteria and NCVs, shaping the genomic island landscape in NCVs. Taken together, these findings position genomic islands as key mediators through which NCVs acquire adaptive traits, shaping their ecological interactions, genome evolution, and host-virus arms race. The results also lay the groundwork for future research into the ecological roles of genomic islands in NCVs and how they may be shaping the ecology and evolution of these viruses on a global scale.

Results and discussion

Genomic islands are widespread across NCV diversity

Utilizing our pipeline and 369 high-quality NCV genomes, we identified a total of 307 genomic islands with flanking direct repeats across 187 unique genomes (51%). Islands ranged from 5 to 135 kbp in length, and 68 genomes had more than one genomic island. Genomic island regions occupied an average of 15% of the genome of the virus that had them, with the highest percentage being 44% (Fig. 1). The distribution of islands across taxonomic orders varied substantially, with the majority of identified islands and the highest density of islands per kilobase (kb) concentrated within the Imitervirales order.

Fig. 1. Distribution of genomic islands in giant viruses.

Fig. 1

A phylogenetic tree using a concatenated alignment of DEAD/SNF2-like helicase (SFII), DNA polymerase B (polB), Packaging ATPase (A32), and Poxvirus Late Transcription Factor VLTF3 (VLTF3) was constructed using our 369 reference genomes across 5 orders of Nucleocytoviricota. Rings around the tree are as follows: (1) Total genome size, (2) the total number of islands identified, (3) Islands per kbp, normalized to genome size, (4) total length of island regions, (5) percent of the genome that island regions took up, and (6) island bacterial enrichment which is calculated as the difference between the proportion of genes having best hits to bacteria inside and outside the genomic islands in a genome. Negative values were rounded to 0.

Despite the high absolute number of genomic islands observed within large orders such as the Imitervirales, this distribution largely mirrors the underlying taxon sampling density. All 5 phylogenetic orders searched contained genomes with genomic islands, and there did not appear to be statistically significant phylogenetic enrichment of genomic islands within a particular order or family of NCV (Fig. S2). However, there tended to be more genomic islands per genome as genome size increased (R^2 = 0.44) (Fig. S3). Genomic islands also tended to have slightly higher GC % than the rest of the genome (by around 2%), with no noticeable trend in differences in codon bias or gene length (Fig. S4). These findings align with what has been found in bacteria, as bacterial genomic islands tend to have GC differences ranging from 2 to 10%37,38.

To determine whether the mechanisms of genomic island integration have signals of evolutionary conservation across viral lineages, we evaluated the sequence similarity of the flanking direct repeats (Fig. S5). Pairwise Levenshtein distances showed a significant taxonomic signal within both the order and family levels, with a mean order similarity of 0.443 vs 0.329 between orders (Mann-Whitney U, p < 0.001) (Fig. S5a). A similar pattern was observed on the family level as well (Fig. S5b); however, similarities were highly variable across families, ranging from 0.556 in IM_08 to 0.321 in AG_02 (Fig. S5c). These findings suggest that repeat regions can be taxonomically conserved, although this is not always the case.

Using the San Francisco Estuary short-read data for read recruitment, 53 of the 72 islands (74%) identified in the San Francisco Estuary NCV genomes showed deviations in the read recruitment (+/−2 SD) of island regions (Fig. S1b). This finding supports the possibility that these islands can either be hypervariable regions within a viral population or widely distributed across many distinct viral populations, as has been shown elsewhere32,33,39.

Functional diversity and phylogenetic differences in genomic island proteins

A total of 12,691 proteins were found to be encoded within the genomic island regions, representing 6894 unique protein clusters when clustered into orthologous groups. Of these, 42 proteins were significantly enriched in the islands of a particular NCV order (Fig. 2a). Many of the identified enriched proteins contained domains of unknown function according to PFAM, but further embedding-based homology searches (see methods) showed they were related to viral membranes and structure, highlighting unique mechanisms and proteins that may be used in capsid assembly and attachment between NCV orders40.

Fig. 2. Functional differences in genomic islands across Nucleocytoviricota orders.

Fig. 2

a A Fisher exact test (one-sided) with Bonferroni correction was performed (FDR p-value cutoff of .05), resulting in 42 proteins enriched in islands of a particular order. The heatmap shows the log10 gene counts (log10 + 1 to remove errors), and the barplot shows the odds ratio when comparing that order to all the other orders. b Proteins of interest associated with viral structure, replication, metabolism, environmental response, and interaction were selected based on previous studies3. The proportion of islands containing these functions for each Nucleocytoviricota family is shown as bubbles in the plot.

Of particular interest, certain NCV hallmark genes, such as the major capsid protein (MCP) and VLTF3 were found inside some genomic islands. To investigate whether these markers show evidence of cross-family or cross-order transfer from other NCVs, a phylogenetic tree was constructed for VLTF3 markers found within genomic islands (Fig. S6a) and MCPs coming from both inside and outside islands in the same genome (Fig. S6b). While most of the VLTF3 and MCP genes clustered within their appropriate clade (IM_01), some showed significant phylogenetic dispersion, as 2 VLTF3s clustered with the IM_07 family (Fig. S6a) and 2 MCPs within genomic islands clustered within the PV_05 and AG_01 families (Fig. S6b), crossing the order boundary. While the evolutionary history of MCP is complicated due to the presence of many paralogs, the results from both MCP and VLTF3 phylogeny provide some evidence that key hallmark genes within NCV can be transferred through genomic islands across phylogenetically distant genomes, having potential implications for host range and adaptation. A search for entire genomic islands that were shared across NCV orders, however, did not yield any such islands.

Looking at key functional differences between NCV orders, Imitervirales genomic islands tend to encode more genes on average related to DNA repair and environmental response, such as MutS, rhodopsins, superoxide dismutase, and photolyase, while photosynthesis genes were found nearly exclusively in Algavirales islands, consistent with previous genome-wide findings3,41 (Fig. 2b). Across all orders, islands encoded high proportions of genes with likely roles in host interaction, with the most common types being surface adhesion proteins, methyltransferases, and glycosyltransferases (Fig. 2b)42,43. Interestingly, multiple studies have shown that bacterial genomic islands frequently harbor genes involved in similar functions44–46, and recent studies on hypervariable regions within the genomes of some prasinoviruses reported these functions as well32. These suggest some functional parallel between bacterial and NCV genomic islands, which might point to similar adaptive functions that are mediated through the genomic islands.

Clustering the genomic islands based on the presence/absence of unique orthogroups (see methods) revealed nine clusters. (Fig. S7). Of these nine clusters, seven formed tight groups, likely due to the close phylogenetic similarity of the genomes from which the islands originated. The largest cluster was characterized by the sporadic presence of numerous orthogroups, highlighting the high diversity of proteins found in most genomic islands. Further evidence for this high diversity comes from the fact that of the 12,691 island proteins, 5,390 are singletons (42.5%). This is consistent with previous observations in diverse bacterial genomic islands that harbor large amounts of hypothetical proteins47,48.

Genomic islands in NCVs are enriched in surface adhesion proteins

The island proteins were mainly categorized as metabolic, interaction, or replication proteins (Fig. 3). Of the 369 total islands, 75% had at least one protein involved in host interaction, 71% had a replication protein, 58% encoded a gene with possible metabolic functions, and 16% had all three types of genes encoded on the island (Fig. 3a, b).

Fig. 3. Genomic island functional characteristics.

Fig. 3

a An upset plot showing the number of islands encoding different types of functions related to structure, adaptation, metabolism, replication, or host interaction. An island was placed into a category if it encoded at least one gene in that category. Islands entirely composed of genes of unknown function were placed in the unknown category. b Representative islands for each island type were mapped with lovis4u, highlighting the genes fitting into defined categories. The location of insertion of these genomic islands is highlighted in blue in the small genome plots under each respective island. The full genomic context for these islands can be found in Fig. S23. c A chi-square test (two-sided) with Yates’ correction was run comparing all of the proteins inside the genomic islands to the proteins found in giant virus genomes outside the island regions. Enriched proteins (p < 0.05) in both groups are shown here with their odds ratios and categorization. d Representative enriched surface adhesion proteins were folded using AlphaFold3. The different surface adhesion domains are highlighted. Proteins were aligned to the closest reference within the Seqhub platform. Alignments are available on figshare. TM-scores and RMSD are displayed for each alignment. More information about the closest reference bacterial genomes can be found on figshare.

To determine the functional specialization of the NCV genomic islands, we performed a functional enrichment analysis comparing genes within genomic islands vs those encoded in the rest of the genome. This revealed that genomic island regions were significantly enriched in various surface adhesion-related proteins (Ig-like domain, YadA, Collagen helix repeat, F5/8, Chaperone of endosialidase), as well as a gene related to DNA mobility with an HMG-box domain (Fig. 3c).

Further phylogenetic investigation into the enriched surface adhesion proteins demonstrated that they seemed to be of bacterial origin, clustering within clades of similar proteins found in bacterial isolate genomes (Fig. S8) as well as having high protein-level alignment with annotated proteins from bacterial genomes (Fig. 3d). Many of these enriched surface adhesion proteins have known adhesion roles within bacteria and phage communities, such as the YadA domain having a role in adhesion of bacteriophage tail fibers49, as well as the cell-to-cell adhesion role of Ig domains, being the most widely used class of adhesion protein among viruses50,51. Other proteins (F5/8 and FG-GAP) have been previously identified in genomic island regions of phages and NCVs infecting the marine alga Ostreococcus32,52, highlighting the potential unique role these islands play in host adaptation through the exchange and capture of surface adhesion proteins.

While surface adhesion has not been an area of in-depth study within NCVs, it is undoubtedly an important step in successful infection, as most NCVs enter the host cell through phagocytosis or receptor-mediated endocytosis, both of which rely on some form of surface adhesion10,53. One example of a well-studied surface adhesion protein in an NCV is the Mimivirus fibrils, which bind mannose and N-acetylglucosamine to mimic bacteria when binding to the surface of the host54,55. Iridoviruses also encode IG-like domains, similar to those found to be enriched in our study, that have a predicted role in cell adhesion and evasion of the host immune system56. Within the realm of eukaryotic viruses, surface adhesion proteins are a key area of host adaptation, and studies have shown that switching of surface proteins is often indicative of host switching57,58. Specific combinations of surface adhesion proteins have also been shown to be important for host entry, as is shown in HIV having both an IG-like domain, along with co-receptors to bypass multiple different cellular “gates”59.

Our findings of enrichment of surface adhesion proteins within genomic island regions of NCVs are intriguing, as they show that these “keys” can be potentially shared throughout virus populations to aid the binding and opening of diverse cellular “locks”. Further phylogenetic investigation shows this to be a possibility, as surface protein phylogeny (FG-gap, Peptidase-S78, F5/8 domain) often shows incongruence with the phylogeny of the host genomes (Fig. S9a–c). This acquisition could allow NCVs to broaden their host range or counter changes in surface recognition proteins from their current host as part of the host-virus arms race, increasing viral fitness.

Genomic islands can be widely shared throughout the NCV population

Leveraging the long-read sequenced giant viral populations of San Francisco Estuary, we evaluated the population-level dynamics in genomic islands across closely related NCVs. Specifically, we identified a conserved island shared across 56 NCV genomes (order Algavirales) and contigs (Fig. 4a). Mapping short reads from the same dataset to both the shared island region and the rest of the genomes revealed that the islands were, on average, ~4x more abundant than the genomes they were derived from, providing further evidence for the exchange of these islands across related NCV populations (Fig. 4c).

Fig. 4. Finding shared genomic islands in the San Francisco Estuary.

Fig. 4

a An all-vs-all BLASTn search was performed using reference genomic islands identified in this study and additional giant virus contigs found in the San Francisco Estuary metagenome36. The resulting connections at 70% minimum identity and 50% coverage were displayed as a Cytoscape network colored by giant virus order. A large connected node of shared Algavirales islands is highlighted by the red circle. b A phylogeny of the selected giant viruses with shared island regions was performed using the same paralog of the major capsid protein (MCP) found outside of the island region. Synteny-based hierarchical clustering of the genomic islands was also performed (Fig. S11). Lines between islands and MCPs coming from the same genome were drawn to show phylogenetic incongruence. c Short reads from the San Francisco Estuary dataset were mapped to both the island regions and the other genomic regions of the selected group (red) at 95% identity. The resulting barplot shows the average RPKM of the island and the rest of the genome, with error bars representing +/− 2 SEM (i = 57). d The same phylogenetic tree as in (b) with overlaid genomic maps of selected contigs showing the annotated conserved island region.

Further phylogenetic analyses showed that the 56 contigs and genomes sharing this island region formed at least two closely related but distinct clades within the order Algavirales (Fig. 4d and Fig. S10a). In addition, average nucleotide identity (ANI) within and between clades supported the phylogenetic patterns. Specifically, both clades had significantly higher within-clade ANI (means of 76.3% and 77.5%) when compared to the ANI between clades (mean 71.7%) (Fig. S10b). Although formal genus-level distinctions cannot be assigned to these clades solely based on ANI differences, their phylogenetic and nucleotide-level divergence indicates that they represent clearly differentiated lineages within Algavirales. Across these genomes, the islands shared a broadly conserved gene repertoire, with many orthogroups present in both clades (Fig. S11). However, the organization of genes (synteny) within the island did not strictly mirror viral phylogeny - in several cases, closely related genomes harbored islands with markedly different gene order, while showing greater structural similarity to islands found in genomes from the alternate clade (Fig. 4b and Fig. S12). This discordance between viral phylogeny and island architecture suggests that the evolutionary history of these regions is partly decoupled from that of the surrounding genome. One parsimonious explanation is horizontal transfer of the island region between related viruses, potentially during co-infection of the same host cell, although recombination and structural reshuffling within the island could also generate the observed mosaic patterns.

The shared islands contained many virion module proteins (minor and major capsid) as well as proteins putatively involved in viral replication and evasion of host defenses, such as methyltransferases and restriction enzymes (Fig. 4d)60,61. The widespread distribution of a viral surface antigen protein in these islands also gives evidence to the hypothesis that genomic islands can be beneficial to closely related populations through the sharing of cellular “keys” for entry and successful infection.

Genomic islands are a driver of hypervariability among closely related NCVs

While we showed that some genomic islands can be shared widely across a NCV population, these islands are not always totally identical, as genes can be added, deleted, or shuffled around within the island region, creating a hypervariable region in the genome that differs between even closely related NCVs32,52. To investigate this phenomenon, we examined several groups of closely related NCVs from diverse environments. Using our genomic island detection framework, we identified conserved island loci that act as modular regions of genomic variability. In these loci, islands share similar boundaries and portions of core gene content, but differ substantially in gene composition, organization, and sequence identity.

A clear example of this pattern occurs within prasinoviruses. Building on recent observations of hypervariable genomic regions in these viruses32, we expanded the analysis by identifying 12 additional closely related genomes from the San Francisco Estuary that contain islands located at the same genomic loci. Although these islands share conserved flanking regions and portions of core gene content, their internal organization is highly variable, particularly in genes encoding surface adhesion proteins (Fig. 5a and Fig. S13). Overall, this analysis demonstrated that these island/hypervariable regions persist over ocean basins (previous prasinoviruses were isolated from the North Sea, Mediterranean Sea, and Pacific Ocean near Hawaii) with only minor changes to potentially adapt to different host regimes32. The conservation of this island region provides strong evidence that this island and those similar to it are critical to viral fitness and host adaptation24.

Fig. 5. Genomic hypervariability across similar genomes.

Fig. 5

a A phylogenetic tree consisting of a concatenated alignment of SFII, PolB, TFIIB, Topoisomerase 2, A32 ATPase, and VLTF3 was constructed using reference prasinoviruses from Thomy et al. (2026)32 and newly discovered prasinoviruses from the San Francisco Estuary data in this study. The hypervariable region identified in Thomy et al. (2026) is highlighted in yellow along with gene functional annotations and phylogenetic groups. b,c Complete genome plots of two pairs of similar viruses with different genomic islands are shown with their island regions highlighted in yellow. Vertical gray bars represent synteny of 90% identity or greater at the amino acid level. All plots were made using lovis4u. A figure showing synteny lines can be found in Fig. S13.

This phenomenon of hypervariability within similar genomic islands is not unique to prasinoviruses, as we identified a similar case in closely related Mimivirus strains from different aquatic basins (Fig. S14). For four different Mimivirus strains, we identified genomic islands (~55 kbp each) with a well-conserved gene order structure in the first half of the island. However, this conservation falls apart in a hypervariable region of the island made up mostly of surface adhesion proteins, such as the collagen helix repeats, adhesin autotransporters previously identified to be enriched in island regions (Fig. 3c), and the YadA head domain. No two islands shared perfect synteny in these regions, and even similar proteins show differences in sequence and length.

In addition to hypervariability within islands, we also identified variability in terms of the presence or absence of genomic islands in closely related genomes (Fig. 5bc). Comparing two closely related Chlorella viruses (CVR-1 and OR0704) isolated from different water basins (Germany and Oregon) revealed a genomic island present in OR0704 that was missing in CVR-1 (Fig. 5b). This island region contained many surface adhesion glycoproteins, common in Chlorella viruses, as well as a chaperone of endosialidase, which also has a putative surface adhesion role62,63.

Previous research on variability in chlorovirus genomes identified multiple “gene gangs”, conserved clusters of collinear genes, among various chlorovirus species64. Since our study identified genomic islands in various chloroviruses, we leveraged this information to investigate the hypotheses of the study, particularly whether these gene gangs are conserved due to the proximity of genes that facilitate their collective transfer64. The genomic island identified in chlorovirus PBCV-1 contained two gene gangs associated with transcription and DNA synthesis (Fig. S15), together with multiple genes putatively involved in facilitating horizontal gene transfer. This co-localization is consistent with the possibility that genomic islands act as modular exchange regions through which linked gene sets can move between viral lineages.

Another occurrence of hypervariability was also observed within two genomic islands from similar NCVs within the same ecosystem (Lake Biwa, Japan). One of these NCVs (LBGVMAG_0006) contained a large 76 kb genomic island that represents a source of syntenic disruption from LBGVMAG_0143 (Fig. 5c). This genomic island contained a surface adhesion protein as well as other genes involved in replication (polymerases) and host adaptation through glycosyltransferases43,65. Interestingly, this genomic island also contained a MutS gene, a commonly found DNA mismatch repair gene found in NCVs3. Upon further inspection of the data from the initial study, both of these viruses were predicted to occupy the same niche layer of the lake (epilimnion), but the island-containing LBGVMAG_0006 was estimated to be 11.7x more abundant than the similar LBGVMAG_014335. This finding may suggest that the island region is critical to efficient viral infection and reproduction in this ecosystem, and the DNA repair capacity against UV damage or modulation of host adhesion might play a role in the success of this virus strain over the strain without this island66. This example highlights the possible ecological advantages genomic islands likely confer to diverse NCVs.

Overall, these four examples demonstrate that hypervariability in genomic islands containing surface adhesion proteins may be a widespread phenomenon within NCVs and may be a key aspect of their environmental fitness in adapting to changing hosts or host-defense mechanisms. This hypervariability is likely driven by a combination of gene gain and loss as well as gene transfer between similar NCVs19.

Some genomic islands have bacterial origins

Given that genomic islands are widespread in NCVs and potentially aid in the virus-host arms race, we sought to identify the potential origins of these islands. It is well known that NCVs can exchange genes with their eukaryotic hosts as well as other viruses19. Interestingly, many NCVs are known to harbor genes of bacterial origin, but the driver or mechanism of transfer to and from bacteria remains unclear, given that bacteria are not hosts of NCVs. In our study, we noticed that island regions were sometimes enriched in bacterial, rather than eukaryotic genes (Fig. 1).

To further investigate the potential for gene transfer between NCV genomic islands and other cellular organisms, we performed a protein clustering between genomic islands of NCVs from the San Francisco estuary and other cellular contigs identified within the same metagenome data (Fig. 6a). This clustering analysis revealed 68 proteins clustering both within the NCV islands and cellular contigs, with the majority clustering with bacterial contigs (42) (Fig. 6a). These proteins came from both Imitervirales and Algavirales genomes in similar proportions (6.7 and 8.2%) and represent putative recent transfer events given the high similarity (>80%) to the cellular hits67,68. Annotations of the clustered proteins revealed that the most frequent transfer events included the FG-GAP domain surface adhesion protein found in both Algavirales and Imitervirales genomes (Fig. S16). Other surface adhesion proteins were also among the most transferred, including 2 chaperones of endosialidase and a Chlorovirus glycoprotein (Fig. S16). As described earlier, many of these surface adhesion proteins were also found to be phylogenetically related to proteins from bacterial isolates (Fig. S8).

Fig. 6. Evidence for the bacterial nature of certain giant virus genomic islands.

Fig. 6

a Proteins from genomic islands from the San Francisco Estuary giant virus genomes were clustered with bacterial, eukaryotic, archaeal, and phage contigs from the same metagenome using mmseqs. The resulting alluvial plot shows the number of proteins from each giant virus order that clustered with proteins from a cellular or phage contig (minimum 80% identity and 20% coverage). The proportion of total proteins that clustered with other (cellular) contigs in each order is shown in the barplot on top. b Total island regions from the San Francisco Estuary were also clustered with bacterial contigs from the same metagenome using mmseqs, resulting in 6 matches. The above genome maps highlight regions of synteny as well as giant virus hallmark genes (MCP, VLTF3), bacterial single-copy genes (SCGs), and genes involved in the toxin/antitoxin system. Bacterial contigs were classified based on consensus of their SCGs in anvi’o, and the lowest level of classification is given here. Zoomed-in synteny plots for selected islands with Pfam annotations can be found in Fig. S17.

In addition to evidence of many individual genes within genomic islands being of bacterial origin, we also found evidence of partial or near-entire genomic islands being shared with bacterial genomes (Fig. 6b and Fig. S17). Within the SF estuary genomic islands, we identified six NCV islands that were also present (> 25% length coverage) on well-established bacterial contigs with taxonomic assignments. All of these contigs with established taxonomy could also be binned into bacterial MAGs (Fig. S18). These contigs often contained multiple bacterial single-copy core genes used for taxonomy, as well as bacterial toxin/antitoxin systems (Fig. 6b). To ensure these bacterial contigs were not assembly chimeras, nanopore long reads were mapped to the contigs, and reads were identified in each example that spanned both the island region and the rest of the contig (Fig. S19). Of interest, there was even one read that spanned a bacterial region that harbored both an NCV major capsid protein (MCP) within the island and multiple bacterial genes, such as a gene with the LTXXQ motif, ADP ribosylglycohydrolase, and S-adenosyl-L-homocysteine hydrolase (Fig. S19). The same was done for NCV contigs as well (Fig. S20). Bacterial contigs were classified into many different phyla, demonstrating that this phenomenon isn’t limited to specific phyla or orders. The finding of NCV hallmark genes such as MCP within these islands raises questions about the potential origins of these genes as well as the direction of transfer. Further in-depth evolutionary studies are needed to fully resolve these questions.

While NCVs are thought to infect predominantly eukaryotic organisms4,7, gene exchange with other cellular organisms, such as bacteria, has been previously reported. In an overview of Chlorovirus gene transfer, Filee (2009)69 found that 4–7% of Chlorovirus genes were bacterial in origin. Similar studies found chunks of genes within Mimivirus that seem to be transferred from bacteria near the termini of the genome or located near transposable elements20,21,70. There are also reports of three adjacent Mimivirus genes being potentially acquired from a Clostridium bacterium20. Our findings importantly suggest that not only do genomic islands in NCVs routinely harbor bacterial genes, but some islands feature large segments of DNA shared with bacteria. However, it is important to emphasize that while gene transfer from bacteria to giant viruses through the genomic islands seems like a plausible mechanism of how giant viruses acquire bacterial genes, it is not definitively possible to establish the directionality of gene transfer without further, in-depth phylogenetic studies that incorporate a large number of bacterial and giant virus genomes with high contiguity and wider taxon sampling. Our finding of a bona fide giant virus gene (major capsid protein) on a bacterial contig exemplifies this problem - at least in this particular case, gene transfer from giant virus to bacteria seems most plausible given MCP is a bona fide viral gene. Together, our results provide further evidence that gene exchange between NCVs and bacteria occurs and suggest that genomic islands may facilitate the transfer of large segments of bacterial DNA into NCV genomes (or vice versa), contributing to their genomic mosaicism.

Potential contexts for gene exchange between NCVs and bacteria

The logical question remains as to how bacterial genes are transferred to NCVs that infect Eukaryotic organisms. Filee et al. (2007)21 observed that two conditions would need to be met for gene exchange to occur:(i) an ‘ecological niche bringing viral and bacterial DNA into close contact ‘, and (ii) a mechanism to drive gene acquisition. We will explore both of these conditions in light of our findings.

The first condition is quite easy to satisfy, as the hosts of many NCVs are protists - complex, often single-celled, eukaryotic organisms that regularly host bacterial symbionts or phagocytose bacteria as a food source71,72. These processes could create conditions where infecting viruses are introduced to foreign bacterial DNA inside the protist host, making it plausible for their integration into the genome. Of the six identified bacterial contigs containing NCV islands in our dataset, one of these was binned into a genome that was predicted to be an endosymbiote using Symclatron (Myxococcota sp.)73. While the others were not predicted symbiotes, they include genera such as Luminiphilus that are known to be preferentially grazed by protists due to their large size and lack of defenses74.

The second condition involves identifying a mechanism by which NCVs could acquire genetic material from a bacterium (or vice versa). It was noted by Filee et al. (2007)21 that recombination-primed replication could be an attractive mechanism for gene transfer, as poxviruses and some Algavirales members undergo high levels of homologous recombination during replication75,76.

Another potential mechanism for the exchange of genomic islands between NCVs would be through mobile genetic elements (MGEs), genetic material that can move around a genome77. NCVs have previously been shown to contain a diverse array of MGEs, with many having a putative prokaryotic origin78. We found that diverse MGEs are also prevalent within NCV genomic islands across all identified orders (Fig. S21a), as 117 (38%) islands have at least one MGE (Fig. S21b). Among the most common are restriction-modification systems (R-M), restriction enzymes, and HNH endonucleases (Fig. S21b). These R-M systems are of particular interest as they have been proposed to be another potential candidate for promoting recombination in NCVs, as is seen in herpesvirus recombination79. HNH endonucleases have also been previously identified in Algaviruses and Mimiviruses, often associated with other genes of bacterial origin21,80. Some islands also appear to contain transposons within them, characterized by diverse transposases found within these regions (Fig. S22). While this finding offers a promising putative mechanism, further mechanistic and culture-based studies are needed to confirm these hypotheses.

Our study establishes genomic islands as a widespread, dynamic engine of genome plasticity across Nucleocytoviricota. By analyzing 369 high-quality genomes, we demonstrate that these islands are broadly distributed across all major phylogenetic orders and frequently display signatures of hypervariability among closely related viral genomes and enrichment in host-interaction genes. Crucially, our structural and synteny analyses reveal multiple mechanisms of genetic exchange, including the acquisition of large bacterial segments and the lateral transfer of entire islands between co-occurring viral strains. These patterns suggest that infected protist hosts could function as critical ecological hubs where giant viruses, eukaryotic hosts, and endosymbiotic bacteria actively exchange genetic material. Ultimately, the rapid gain, loss, and remodeling of these genomic modules allow NCVs to adapt to shifting host defenses and contribute to the evolutionary trajectory of giant viruses. Overall, genomic islands in giant viruses provide a new model for how gene flow drives viral innovation and diversification in the biosphere.

Methods

Gathering high-quality NCV genomes

To identify genomic islands across the diversity of NCVs, we obtained NCV genomes that were either complete or near-complete (based on the presence of key hallmark genes) from various sources. These included 59 genomes from cultured viruses (only those with an N50 of >= 100 kb) listed in the giant virus genome database (GVDB)6 and 146 published nanopore long-read genomes from Lake Biwa, Japan35. In addition to these, we identified an additional 164 single-contig nanopore long-read genomes from published data from the south bay of the San Francisco Estuary36 using BEREN (v1.0) in contig mode81.

All of these genomes were selected based on the presence of at least 5 Nucleocytoviricota marker genes (out of 9 total searched for: MCP, SFII, RNAPS, RNAPL, PolB, TFIIB, TopoII, A32, VLTF3)6, identified using the ncldv_markersearch tool6. Among the 5 marker genes, genomes had to contain a DEAD/SNF2-like helicase (SFII), DNA polymerase B (polB), Packaging ATPase (A32), and Poxvirus Late Transcription Factor VLTF3 (VLTF3), as these are common genes used in establishing the phylogeny and identity of NCVs and yielded the highest tree confidence scores in previous phylogenetic reconstructions4,6. This resulted in a total of 369 high-quality or near-complete genomes for downstream analysis. A concatenated phylogeny of these 4 marker genes was created using MAFFT (v5)82 for alignment, trimAl (v1.4) with ‘-gt 0.1’83 for alignment trimming, and IQTREE (v2.0.3)(LG + F + R10 model)84. The resulting tree was visualized using Anvi’o (v8)85 (Fig. 1).

Identification of genomic islands

To identify the genomic islands in the genomes of NCVs, we developed a Python tool (IslaMine (v1.0.0)) that uses deviations in tetranucleotide frequency (TNF) and flanking direct repeats to identify these regions, common features of genomic islands found in diverse bacteria and eukaryotes24,86. Each genome is first divided into non-overlapping chunks of 5 kbp from which the TNF profile is computed and compared to the TNF of the full genome using a Pearson correlation. The resulting correlation values are converted to a cumulative sum and first derivative to highlight compositional shifts. A sliding window variance is then calculated across the derivative signal to identify regions of local instability. Peaks in the variance signal are then detected using adaptive thresholds based on genome size (90th percentile of the variance), with adjacent or overlapping peaks being merged. An example output can be found in Fig. S1a.

Each region of high variance that is predicted through this pipeline is then scanned for the presence of flanking direct repeats (minimum length of 10 bp), +/-500 bp on either side of the region. Islands and the flanking regions (+/-5 kb) were scanned for tRNAs using tRNAscanSE (v2.0)87. The IslaMine tool is publicly available on GitHub (https://github.com/BenMinch/IslaMine).

Validation of the tool using known NCV genomic islands

In separate studies, several NCV genomes were reported to contain putative genomic islands31,32. We attempted to independently recover these genomic islands using our Python tool, IslaMine. Our tool successfully recovered all 21 of the genomic islands present in these genomes, with the identified islands often covering > 90% of the reported island regions (Supplemental Table 1). This validates IslaMine’s ability to identify NCV genomic islands and indicates that island regions previously identified in NCVs share common compositional and structural features (altered tetranucleotide frequency and flanking repeats), which possibly represent generalizable hallmarks of genomic islands in NCV genomes.

Genomic island read recruitment

To identify genomic islands that also had heterogeneous read recruitment, Illumina short-read data from the San Francisco Estuary36 were mapped to all identified genomes from that region using minimap2 (v2.17)(-pa 95)88. Resulting alignments were converted to BAM files with samtools (v1.10)89 and visualized using a custom Python script to look at coverage across the genome (Fig. S1b).

Functional analysis of genomic islands

For each genomic island, open reading frames (ORFs) were predicted using prodigal-gv (v2.10)90,91 and annotated using Annomazing (v1.0) (a HMMER-based annotation pipeline92)34 with the Pfam (v36)93 and GVOG6 databases (e-value cutoff of 1e-5). Resulting annotations were then categorized based on their Pfam annotation (Structure, Environment Response, Metabolism, Replication, Interaction, or Unknown) (Supplementary Table 2). Proteins with unknown Pfam domains (DUFs) were additionally annotated using Gaia, which utilizes protein embeddings94. Mobile genetic elements within the genomic islands were identified using REBASE (v31.07)95, ISfinder (v2.0)96, and Pfam databases93.

To determine genes enriched in islands of specific phylogenetic orders, a Fisher exact test with Bonferroni correction was performed utilizing an FDR-corrected p-value of <0.05 for significance. For testing enriched genes in island regions, a chi-square test with Yates’ correction was performed utilizing an FDR-corrected p-value of <0.05 for significance. Maps of genomic islands were created using lovis4u97 (Figs. 3–5).

Proteins encoded within and outside of genomic islands were independently clustered into orthogroups using MMseqs2 (v13.45111)(--min-seq-id 0.25 -c 0.5 –cov-mode 1 -s 7.5)98. Genomic islands were subsequently clustered functionally based on the presence or absence of orthogroups containing at least five members. To this end, hierarchical clustering was performed using Jaccard distances and the Ward.D2 method implemented in the vegan (v2.7-5) package99,100 in R (v4.5.2). The optimal number of clusters was determined by minimizing the Davies-Bouldin index, and the resulting clusters were visualized using pheatmap (v1.0.13) in R101.

For certain surface adhesion proteins, phylogenetic trees were built alongside 100 reference sequences from the top 100 hits to the Open Genomes Database (based on embedding distances)102 within Seqhub94. Alignments were made using MAFFT82 and trimmed using trimAl83. Trees were made using IQTREE84 with 1000 bootstraps and displayed using iTOL (v7)103. Structures of certain island-encoded proteins were predicted using AlphaFold3104, and alignments to the closest reference (Open Genomes Database based on embedding distance) were performed using the Seqhub platform94.

Locating shared genomic islands from the San Francisco Estuary

To locate similar or shared genomic islands in the co-occurring NCVs in the San Francisco Estuary environmental dataset, we first created a BLAST (v2.9.0)105 database of all the genomic islands found from genomes recovered from this location. Additional NCV contigs were recovered from the San Francisco nanopore data utilizing BEREN in ‘contig’ mode81.

BLASTn was then run with a minimum percent identity of 70% and a minimum coverage of 50% of the subject to locate similar islands in these NCV contigs. This modified pipeline was used as these islands were often on smaller contigs, not representing full genomes. This analysis revealed a cluster of 56 contigs and genomes that shared a similar genomic island. The BLAST network was visualized using Cytoscape (v2.1)106 (Fig. 4).

To further characterize the distributions of the islands across these contigs, a phylogenetic tree was created for each contig/genome sharing an island using MCP as the phylogenetic marker. Effort was made to use the same MCP paralog for each contig, as NCVs can have multiple MCP copies. This was determined by similar flanking genes and gene neighborhood. MCPs were first aligned with MAFFT82, then trimmed with trimal ‘-gt 0.1’83, and the final tree was constructed using IQTREE84. This phylogeny was compared with hierarchical clustering of the gene synteny in the island regions to determine whether gene transfer or the accumulation of mutations was most likely the cause of island heterogeneity107. Additionally, contigs containing shared islands were compared to each other using pyani (v0.2.12)108.

To assess island similarity, island proteins were clustered into orthologous groups using proteinortho (v6.3.6) in default mode109. The synteny of orthologous groups within islands was compared using Levenshtein distance, treating each protein like a word in a sentence. The resulting synteny dissimilarity was plotted in R.

Finally, to quantify differences in coverage of this widely shared island to the individual NCV contigs and genomes housing the island, short reads from the San Francisco dataset were mapped at 95% identity to both the genomic islands and the rest of the genome individually using minimap288. The raw read counts were normalized using RPKM.

Genomic hypervariability among similar strains

Using a comparative phylogenetic framework, we identified and analyzed genomic islands (identified using Islamine) from prasinoviruses from the San Francisco Estuary to evaluate the prevalence of genomic hypervariability in prasinoviruses. We utilized 19 prasinoviruses from a previous study by Thomy et al. 32, as well as 13 new prasinoviruses from the San Francisco Estuary, and created a phylogenetic tree using a concatenated alignment of SFII, PolB, TFIIB, Topoisomerase 2, A32 ATPase, and VLTF3. The alignment was constructed using MAFFT, trimmed with trimAl82,83 using ‘-gt 0.1 ‘, and the tree was built using IQTREE84. Maps for all the genomic islands in the prasinoviruses were created using lovis4u (v0.2.0)97 (Fig. 5a).

Additional pairs of similar Mimivirus genomes with unique genomic islands were identified through clustering all genomes used in the initial analysis at 97% identity using mmseqs98. Genomes clustering together were then aligned using synteny plots from lovis4u97 and manually inspected for unique genomic islands (Fig. S14).

Analysis of the potential origins of genomic islands in NCVs

To identify possible transfer of genomic regions across genomes of co-occurring NCVs and cellular organisms in the same environment, we leveraged the metagenomic data from the San Francisco Estuary. First, all nanopore-assembled contigs were filtered to a minimum length of 10 kbp. Then these contigs were classified to the kingdom level using a consensus approach of Tiara (v1.0.3)110 and mmseqs taxonomy using the UniRef50 database (Release 2026_02)98,111. Bacterial contigs were further filtered to only include those encoding a bacterial hallmark gene (bacterial SCG)112. The proteins in the NCV genomic islands were clustered with a database of proteins from other contigs in the dataset using mmseqs linclust with a minimum sequence identity of 80% and a minimum coverage of 20%. Only island proteins that clustered with a protein on a contig with known taxonomy were further illustrated in the alluvial plot using ggplot2 (v4.0.0) (Fig. 6b)113.

To obtain further evidence of island transfer between NCVs and bacteria, we sought to identify possible cases of entire or near-entire island regions in bacterial contigs. First, all bacterial contigs>100 kbp were obtained from our previous analysis of HGT in the San Francisco Estuary bacterial contigs. The NCV islands were then aligned to these regions using BLASTn with a minimum alignment length of 25% of the island and an e-value cutoff of 1e-20105. The resulting bacterial contigs were assigned taxonomy using single-copy genes in Anvi’o (anvi-estimate-taxonomy)85. Contigs that were found to have these islands were then binned with other contigs in the San Francisco long-read metagenome using Semibin2 (v2.3.0)114, and the symbiotic status of the bins was identified using symclatron (v0.10.10)73 after manual bin refinement in Anvi’o85. Genome maps were made using Proksee115.

To ensure that the island regions on the bacterial contigs were not erroneous due to chimeric assemblies, the nanopore long reads from the San Francisco Estuary were mapped back to the bacterial contigs using minimap2 to identify reads that spanned both the island regions and the flanking genomic regions88. The alignment results were then visualized using pysam and a custom script.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

41467_2026_77295_MOESM2_ESM.docx (9.2KB, docx)

Description of Additional Supplementary Files

Reporting Summary (1.6MB, pdf)

Acknowledgements

This work was supported by the Rosenstiel School of Marine, Atmospheric, and Earth Sciences, University of Miami. We would also like to thank Dr. Torben Nielsen and Dr. Lauren Lui (Lawrence Berkeley National Laboratory, Berkeley, California, USA) for their provision of the San Francisco Estuary metagenomic dataset and for help in editing the final manuscript.

Author contributions

B.M. and M.M. conceived the study, performed the data analysis, and wrote the manuscript.

Peer review

Peer review information

Nature Communications thanks Rodrigo Rodrigues and the other anonymous reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Funding

This work was supported by National Science Foundation grant OCE-2346438 to M.M. The funding bodies had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Data availability

All genomes used in the analysis as well as raw files for phylogenetic trees, alignments, genomic islands, and protein annotations have been deposited on FigShare (https://figshare.com/projects/_b_Widespread_genomic_islands_in_giant_viruses_shape_genome_plasticity_and_mosaicism_b_/273924).

Code availability

The code for recovering genomic islands is publicly available on GitHub (https://github.com/BenMinch/IslaMine)116.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s41467-026-77295-5.

References

  • 1.Abrahão, J. et al. Tailed giant Tupanvirus possesses the most complete translational apparatus of the known virosphere. Nat. Commun.9, 749 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Philippe, N. et al. Pandoraviruses: amoeba viruses with genomes up to 2.5 mb reaching that of parasitic eukaryotes. Science341, 281–286 (2013). [DOI] [PubMed] [Google Scholar]
  • 3.Moniruzzaman, M., Martinez-Gutierrez, C. A., Weinheimer, A. R. & Aylward, F. O. Dynamic genome evolution and complex virocell metabolism of globally distributed giant viruses. Nat. Commun.11, 1710 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Schulz, F. et al. Giant virus diversity and host interactions through global metagenomics. Nature578, 432–436 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Moniruzzaman, M. et al. Virologs, viral mimicry, and virocell metabolism: the expanding scale of cellular functions encoded in the complex genomes of giant viruses. FEMS Microbiol. Rev.47, fuad053 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Aylward, F. O., Moniruzzaman, M., Ha, A. D. & Koonin, E. V. A phylogenomic framework for charting the diversity and evolution of giant viruses. PLOS Biol.19, e3001430 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sun, T.-W. et al. Host range and coding potential of eukaryotic giant viruses. Viruses12, 1337 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Bar-On, Y. M. & Milo, R. The biomass composition of the oceans: a blueprint of our blue planet. Cell179, 1451–1454 (2019). [DOI] [PubMed] [Google Scholar]
  • 9.Christien, P. et al. Coccolithovirus facilitation of carbon export in the North Atlantic. Nat. Microbiol.3, 537–547 (2018). [DOI] [PubMed] [Google Scholar]
  • 10.Brandes, N. & Linial, M. Giant viruses—big surprises. Viruses11, 404 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Claverie, J.-M. & Abergel, C. Giant viruses: The difficult breaking of multiple epistemological barriers. Stud. Hist. Philos. Sci. Part C. Stud. Hist. Philos. Biol. Biomed. Sci.59, 89–99 (2016). [DOI] [PubMed] [Google Scholar]
  • 12.DiMaio, D. Viruses, masters at downsizing. Cell Host Microbe11, 560–561 (2012). [DOI] [PubMed] [Google Scholar]
  • 13.Belshaw, R., Pybus, O. G. & Rambaut, A. The evolution of genome compression and genomic novelty in RNA viruses. Genome Res.17, 1496–1504 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Phan, T. G. et al. Small viral genomes in unexplained cases of human encephalitis, diarrhea, and in untreated sewage. Virology482, 98–104 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Mollentze, N., Babayan, S. A. & Streicker, D. G. Identifying and prioritizing potential human-infecting viruses from their genome sequences. PLoS Biol.19, e3001390 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Brahim Belhaouari, D. et al. Metabolic arsenal of giant viruses: Host hijack or self-use? eLife11, e78674 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Filée, J. Genomic comparison of closely related giant viruses supports an accordion-like model of evolution. Front. Microbiol.6, 593 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Sun, T.-W. & Ku, C. Unraveling gene content variation across eukaryotic giant viruses based on network analyses and host associations. Virus Evol.7, veab081 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wu, J. et al. Gene transfer among viruses substantially contributes to gene gain of giant viruses. Mol. Biol. Evol.41, msae161 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Moreira, D. & Brochier-Armanet, C. Giant viruses, giant chimeras: the multiple evolutionary histories of mimivirus genes. BMC Evol. Biol.8, 12 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Filée, J., Siguier, P. & Chandler, M. I am what I eat and I eat what I am: acquisition of bacterial genes by giant viruses. Trends Genet.23, 10–15 (2007). [DOI] [PubMed] [Google Scholar]
  • 22.Claverie, J.-M. & Abergel, C. Open questions about giant viruses. Adv. Virus Res.85, 25–56 (2013). [DOI] [PubMed] [Google Scholar]
  • 23.Guglielmini, J., Woo, A. C., Krupovic, M., Forterre, P. & Gaia, M. Diversification of giant and large eukaryotic dsDNA viruses predated the origin of modern eukaryotes. Proc. Natl. Acad. Sci. USA116, 19585–19592 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Juhas, M. et al. Genomic islands: tools of bacterial horizontal gene transfer and evolution. Fems Microbiol. Rev.33, 376–393 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hacker, J. & Carniel, E. Ecological fitness, genomic islands and bacterial pathogenicity. EMBO Rep.2, 376–381 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Busby, B., Kristensen, D. M. & Koonin, E. V. Contribution of phage-derived genomic islands to the virulence of facultative bacterial pathogens. Environ. Microbiol.15, 307–312 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.López-Pérez, M., Martin-Cuadrado, A.-B. & Rodriguez-Valera, F. Homologous recombination is involved in the diversity of replacement flexible genomic islands in aquatic prokaryotes. Front. Genet. 5, 147 (2014). [DOI] [PMC free article] [PubMed]
  • 28.Hacker, J. & Kaper, J. B. Pathogenicity islands and the evolution of microbes. Annu. Rev. Microbiol.54, 641–679 (2000). [DOI] [PubMed] [Google Scholar]
  • 29.Fischer, M. G., Allen, M. J., Wilson, W. H. & Suttle, C. A. Giant virus with a remarkable complement of genes infects marine zooplankton. Proc. Natl. Acad. Sci.107, 19508–19513 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Filée, J. & Chandler, M. Convergent mechanisms of genome evolution of large and giant DNA viruses. Res. Microbiol.159, 325–331 (2008). [DOI] [PubMed] [Google Scholar]
  • 31.Mukherjee, I. et al. Cultivation, genomics, and giant viruses of a ubiquitous and heterotrophic freshwater cryptomonad. ISME J.19, wraf271 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Thomy, J. et al. Genomic analysis of Ostreococcus tauri-infecting viruses reveals a hypervariable region associated with host–virus interactions. Virus Evol.12, veaf096 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Pagarete, A. et al. Dip in the gene pool: metagenomic survey of natural coccolithovirus communities. Virology466–467, 129–137 (2014). [DOI] [PubMed] [Google Scholar]
  • 34.Minch, B. & Moniruzzaman, M. Expansion of the genomic and functional diversity of global ocean giant viruses. npj Viruses3, 32 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Zhang, L., Meng, L., Fang, Y., Ogata, H. & Okazaki, Y. Spatiotemporal dynamics of giant viruses within a deep freshwater lake reveal a distinct dark-water community. ISME J.18, wrae182 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Lui, L. M. & Nielsen, T. N. Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses. mSystems9, e00242–24 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.da Silva Filho, A. C. et al. Comparative analysis of genomic island prediction tools. Front. Genet.9, 619 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Che, D., Hasan, M. S. & Chen, B. Identifying pathogenicity islands in bacterial pathogenomics using computational approaches. Pathogens3, 36–56 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Schön, M. E. et al. Strain-level diversity of giant viruses infecting chlorarachniophyte algae in the subtropical North Pacific. ISME J. 20, p.wrag093 (2026). [DOI] [PMC free article] [PubMed]
  • 40.Talbert, P. B., Henikoff, S. & Armache, K.-J. Giant variations in giant virus genome packaging. Trends Biochem. Sci.48, 1071–1082 (2023). [DOI] [PubMed] [Google Scholar]
  • 41.Ha, A. D., Moniruzzaman, M. & Aylward, F. O. High transcriptional activity and diverse functional repertoires of hundreds of giant viruses in a coastal marine system. mSystems6, e0029321 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Jeudy, S. et al. The DNA methylation landscape of giant viruses. Nat. Commun.11, 2657 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Speciale, I. et al. The astounding world of glycans from giant viruses. Chem. Rev.122, 15717–15766 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Franzin, F. M. & Sircili, M. P. Locus of enterocyte effacement: a pathogenicity island involved in the virulence of enteropathogenic and enterohemorrhagic Escherichia coli subjected to a complex network of gene regulation. BioMed. Res. Int.2015, 534738 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Moon, B. Y. et al. Mobilization of Genomic Islands of Staphylococcus aureus by Temperate Bacteriophage. PLoS ONE11, e0151409 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.De Maayer, P. & Cowan, D. A. Comparative genomic analysis of the flagellin glycosylation island of the Gram-positive thermophile Geobacillus. BMC Genomics17, 913 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Haas, D. et al. Synteruptor: mining genomic islands for non-classical specialized metabolite gene clusters. NAR Genomics Bioinforma.6, lqae069 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.McCarlie, S. J., Boucher, C. E. & Bragg, R. R. Genomic islands identified in highly resistant Serratia sp. HRI: a pathway to discover new disinfectant resistance elements. Microorganisms11, 515 (2023). [DOI] [PMC free article] [PubMed]
  • 49.Xiang, Y. et al. Crystallographic insights into the autocatalytic assembly mechanism of a bacteriophage tail spike. Mol. Cell34, 375–386 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Aricescu, A. R. & Jones, E. Y. Immunoglobulin superfamily cell adhesion molecules: zippers and signals. Curr. Opin. Cell Biol.19, 543–550 (2007). [DOI] [PubMed] [Google Scholar]
  • 51.Fraser, J. S., Yu, Z., Maxwell, K. L. & Davidson, A. R. Ig-like domains on bacteriophages: a tale of promiscuity and deceit. J. Mol. Biol.359, 496–507 (2006). [DOI] [PubMed] [Google Scholar]
  • 52.Aldaihani, R. & Heath, L. S. Connecting genomic islands across prokaryotic and phage genomes via protein families. Sci. Rep.13, 344 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Bosmon, T., Abergel, C. & Claverie, J.-M. 20 years of research on giant viruses. npj Viruses3, 1–11 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.de Aquino, I. L. M. et al. Diversity of surface fibril patterns in mimivirus isolates. J. Virol.97, e01824-22 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Rodrigues, R. A. L. et al. Mimivirus fibrils are important for viral attachment to the microbial world by a diverse glycoside interaction repertoire. J. Virol.89, 11812–11819 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Song, W. J. et al. Functional genomics analysis of singapore grouper iridovirus: complete sequence determination and proteomic analysis. J. Virol.78, 12576–12590 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Bhella, D. The role of cellular adhesion molecules in virus attachment and entry. Philos. Trans. R. Soc. B Biol. Sci.370, 20140035 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Maginnis, M. S. Virus–receptor interactions: the key to cellular invasion. J. Mol. Biol. 430, 2590–2611 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Prabakaran, P., Dimitrov, A. S., Fouts, T. R. & Dimitrov, D. S. Structure and function of the hiv envelope glycoprotein as entry mediator, vaccine immunogen, and target for inhibitors. Adv. Pharmacol. San. Diego Calif.55, 33–97 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Li, J. & Zhang, L. The emerging role of m5C modification in viral infection. Virology610, 110606 (2025). [DOI] [PubMed] [Google Scholar]
  • 61.Jivaji, A. M. et al. Giant endogenous viral elements in the genome of the model protist Euglena gracilis reveal past interactions with giant viruses. J. Virol.99, e00713-25 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Choi, J., Shin, J.-H., An, H. J., Oh, M. J. & Kim, S.-R. Analysis of secretome and N-glycosylation of Chlorella species. Algal Res. 59, 102466 (2021). [Google Scholar]
  • 63.Schwarzer, D., Stummeyer, K., Gerardy-Schahn, R. & Mühlenhoff, M. Characterization of a novel intramolecular chaperone domain conserved in endosialidases and other bacteriophage tail spike and fiber proteins*. J. Biol. Chem.282, 2821–2831 (2007). [DOI] [PubMed] [Google Scholar]
  • 64.Seitzer, P. et al. Gene gangs of the chloroviruses: conserved clusters of collinear monocistronic genes. Viruses10, 576 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Pyle, J. D., Lund, S. R., O’Toole, K. H. & Saleh, L. Virus-encoded glycosyltransferases hypermodify DNA with diverse glycans. Cell Rep.43, 114631 (2024). [DOI] [PubMed] [Google Scholar]
  • 66.Furuta, M. et al. Chlorella virus PBCV-1 encodes a homolog of the bacteriophage T4 UV damage repair gene denV. Appl. Environ. Microbiol.63, 1551–1556 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Jaramillo, V. D. A., Sukno, S. A. & Thon, M. R. Identification of horizontally transferred genes in the genus Colletotrichum reveals a steady tempo of bacterial to fungal gene transfer. BMC Genomics16, 2 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Li, X. et al. A novel strategy for detecting recent horizontal gene transfer and its application to rhizobium strains. Front. Microbiol. 9, 973 (2018). [DOI] [PMC free article] [PubMed]
  • 69.Filée, J. Lateral gene transfer, lineage-specific gene expansion and the evolution of Nucleo Cytoplasmic Large DNA viruses. J. Invertebr. Pathol.101, 169–171 (2009). [DOI] [PubMed] [Google Scholar]
  • 70.Boyer, M. et al. Mimivirus shows dramatic genome reduction after intraamoebal culture. Proc. Natl. Acad. Sci. USA108, 10296–10301 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Zhang, B., Xiao, L., Lyu, L., Zhao, F. & Miao, M. Exploring the landscape of symbiotic diversity and distribution in unicellular ciliated protists. Microbiome12, 96 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Weisse, T. et al. Functional ecology of aquatic phagotrophic protists – Concepts, limitations, and perspectives. Eur. J. Protistol.55, 50–74 (2016). [DOI] [PubMed] [Google Scholar]
  • 73.Villada, J. C. et al. A genomic catalog of Earth’s bacterial and archaeal symbionts. 2025.05.29.656868 Preprint at 10.1101/2025.05.29.656868 (2025). [DOI] [PubMed]
  • 74.Ferrera, I., Gasol, J. M., Sebastián, M., Hojerová, E. & Koblízek, M. Comparison of growth rates of aerobic anoxygenic phototrophic bacteria and other bacterioplankton groups in coastal Mediterranean waters. Appl. Environ. Microbiol.77, 7451–7458 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Tessman, I. Genetic recombination of the DNA plant virus PBCV1 in a Chlorella-like alga. Virology145, 319–322 (1985). [DOI] [PubMed] [Google Scholar]
  • 76.Esposito, J. J. et al. Genome sequence diversity and clues to the evolution of variola (Smallpox) virus. Science313, 807–812 (2006). [DOI] [PubMed] [Google Scholar]
  • 77.Frost, L. S., Leplae, R., Summers, A. O. & Toussaint, A. Mobile genetic elements: the agents of open source evolution. Nat. Rev. Microbiol.3, 722–732 (2005). [DOI] [PubMed] [Google Scholar]
  • 78.Filée, J. Giant viruses and their mobile genetic elements: the molecular symbiosis hypothesis. Curr. Opin. Virol.33, 81–88 (2018). [DOI] [PubMed] [Google Scholar]
  • 79.Liao, Y., Bajwa, K., Reddy, S. M. & Lupiani, B. Methods for the manipulation of herpesvirus genome and the application to Marek’s disease virus research. Microorganisms9, 1260 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Van Etten, J. L. Unusual life style of giant chlorella viruses. Annu. Rev. Genet.37, 153–195 (2003). [DOI] [PubMed] [Google Scholar]
  • 81.Minch, B. & Moniruzzaman, M. BEREN: a bioinformatic tool for recovering giant viruses, polinton-like viruses, and virophages in metagenomic data. Bioinforma. Adv.5, vbaf284 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Katoh, K., Misawa, K., Kuma, K. & Miyata, T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 30, 3059–3066 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Capella-Gutiérrez, S., Silla-Martínez, J. M. & Gabaldón, T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics25, 1972–1973 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Nguyen, L.-T., Schmidt, H. A., von Haeseler, A. & Minh, B. Q. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol. Biol. Evol.32, 268–274 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Eren, A. M. et al. Community-led, integrated, reproducible multi-omics with anvi’o. Nat. Microbiol.6, 3–6 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Clasen, F. J., Pierneef, R. E., Slippers, B. & Reva, O. EuGI: a novel resource for studying genomic islands to facilitate horizontal gene transfer detection in eukaryotes. BMC Genomics19, 323 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Chan, P. P., Lin, B. Y., Mak, A. J. & Lowe, T. M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res. 49, 9077–9096 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics34, 3094–3100 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics25, 2078–2079 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Hyatt, D. et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinforma.11, 119 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Camargo, A. P. et al. Identification of mobile genetic elements with geNomad. Nat. Biotechnol. 1–10 10.1038/s41587-023-01953-y (2023). [DOI] [PMC free article] [PubMed]
  • 92.Eddy, S. R. Accelerated profile HMM searches. PLOS Comput. Biol.7, e1002195 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Mistry, J. et al. Pfam: the protein families database in 2021. Nucleic Acids Res. 49, D412–D419 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Jha, N. et al. Gaia: an AI-enabled genomic context–aware platform for protein sequence annotation. Sci. Adv.11, eadv5109 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Roberts, R. J., Vincze, T., Posfai, J. & Macelis, D. REBASE: a database for DNA restriction and modification: enzymes, genes and genomes. Nucleic Acids Res.51, D629–D630 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Siguier, P., Perochon, J., Lestrade, L., Mahillon, J. & Chandler, M. ISfinder: the reference centre for bacterial insertion sequences. Nucleic Acids Res. 34, D32–D36 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Egorov, A. A. & Atkinson, G. C. LoVis4u: a locus visualization tool for comparative genomics and coverage profiles. NAR Genomics Bioinforma.7, lqaf009 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Steinegger, M. & Söding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol.35, 1026–1028 (2017). [DOI] [PubMed] [Google Scholar]
  • 99.Xiao, J., Lu, J. & Li, X. Davies Bouldin Index-based hierarchical initialization K-means. Intell. Data Anal.21, 1327–1338 (2017). [Google Scholar]
  • 100.Dixon, P. VEGAN, a package of R functions for community ecology. J. Veg. Sci.14, 927–930 (2003). [Google Scholar]
  • 101.Kolde, R. pheatmap: Pretty Heatmaps. 1.0.13 10.32614/CRAN.package.pheatmap (2010). [DOI]
  • 102.Cornman, A. et al. The OMG dataset: an open metagenomic corpus for mixed-modality genomic language modeling. 2024.08.14.607850 Preprint at 10.1101/2024.08.14.607850 (2024). [DOI]
  • 103.Letunic, I. & Bork, P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res.49, W293–W296 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinforma.10, 421 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Shannon, P. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res.13, 2498–2504 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Ravenhall, M., Škunca, N., Lassalle, F. & Dessimoz, C. Inferring horizontal gene transfer. PLoS Comput. Biol.11, e1004095 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Pritchard, L., Glover, R. H., Humphris, S., Elphinstone, J. G. & Toth, I. K. Genomics and taxonomy in diagnostics for food security: soft-rotting enterobacterial plant pathogens. Anal. Methods8, 12–24 (2015). [Google Scholar]
  • 109.Lechner, M. et al. Proteinortho: detection of (Co-)orthologs in large-scale analysis. BMC Bioinforma.12, 124 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Karlicki, M., Antonowicz, S. & Karnkowska, A. Tiara: deep learning-based classification system for eukaryotic sequences. Bioinformatics38, 344–350 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Suzek, B. E., Huang, H., McGarvey, P., Mazumder, R. & Wu, C. H. UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics23, 1282–1288 (2007). [DOI] [PubMed] [Google Scholar]
  • 112.Chen, L.-X., Anantharaman, K., Shaiber, A., Eren, A. M. & Banfield, J. F. Accurate and complete genomes from metagenomes. Genome Res.30, 315–333 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Valero-Mora, P. M. ggplot2: elegant graphics for data analysis. J. Stat. Softw.35, 1–3 (2010). [Google Scholar]
  • 114.Pan, S., Zhao, X.-M. & Coelho, L. P. SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencing. Bioinformatics39, I21–I29 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Grant, J. R. et al. Proksee: in-depth characterization and visualization of bacterial genomes. Nucleic Acids Res.51, W484–W492 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Minch, B. IslaMine. Zenodo 10.5281/zenodo.21708729 (2026). [DOI]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

41467_2026_77295_MOESM2_ESM.docx (9.2KB, docx)

Description of Additional Supplementary Files

Reporting Summary (1.6MB, pdf)

Data Availability Statement

All genomes used in the analysis as well as raw files for phylogenetic trees, alignments, genomic islands, and protein annotations have been deposited on FigShare (https://figshare.com/projects/_b_Widespread_genomic_islands_in_giant_viruses_shape_genome_plasticity_and_mosaicism_b_/273924).

The code for recovering genomic islands is publicly available on GitHub (https://github.com/BenMinch/IslaMine)116.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES